AI PM LAB

Interactive lessons to build your AI product sense. Run the models yourself, push them until they break, and learn what it takes to ship.

  • Live models on every run
  • See every step as it happens
  • Change a setting, run it again, compare
  • Immersive 3D journeys inside a model
Start with the foundations

Step inside the model

Immersive, full-screen journeys through what happens inside an AI model

Foundations

How models behave: tokens, prompts, cost, and when they make things up

3DBeginner
Follow a prompt through tokenization, the context window, and sampling, and see why temperature and limits matter.
TokensContext WindowTemperatureMax Tokens
Step inside
Beginner
Build a prompt from ingredients (role, categories, format, examples, rules) and score it on 10 labeled items. See which ones move the score.
Prompt DesignFew-Shot ExamplesOutput FormatAblation
Beginner
Ask 12 questions, some of which nobody can answer, and count how often the model makes something up. Then try the fixes and see which ones work.
HallucinationGroundingAbstainingTemperature
3DBeginner
Search 25 support tickets by keyword and by meaning, and watch your query land on a map where similar tickets sit together.
EmbeddingsSemantic SearchKeyword SearchSimilarity
Step inside
Beginner
Run the same 12 support jobs with a better prompt, with retrieval, and with a fine-tuning stand-in. Each fixes a different kind of problem, and one bakes in last year's prices.
PromptingRAGFine-tuningTrade-offs
Intermediate
Chat about a long contract and watch where every token and dollar goes. See how prompt caching makes repeat requests cheaper and faster.
TokensPricingPrompt CachingUnit Economics

Building

Putting models to work in a product: tools, retrieval, agents

Beginner
Turn 12 messy sales emails into orders your app can save. Ask for JSON four ways and see which replies your code can read, and which ones are right.
Structured OutputJSON SchemaParsingReliability
3DIntermediate
Watch the agent loop: the model decides to call tools, reads their results, and keeps going until it can answer.
Tool UseAgent LoopGuardrailsMemory
Step inside
3DIntermediate
Ask questions about a fictional product the model has never seen. Turn retrieval on and off to see grounding vs. guessing.
EmbeddingsVector SearchTop-KReranking
Step inside
Intermediate
Give a model 12 support requests and a set of tools. Make the tool names vague or add look-alikes from extra MCP servers, and watch which tool it picks and what each request costs.
Tool UseMCPTool DescriptionsToken Cost
Advanced
A lead agent splits a product brief across specialists, each with their own data. Compare it to a single agent on quality, time, and cost.
OrchestrationSpecialistsHandoffsTradeoffs
Advanced
Run 10 support jobs through fixed steps, through an agent that picks its own, and through both. See which handles the unexpected, which stays reliable, and what each costs.
WorkflowsAgentsReliabilityCost

Shipping

Making it measurable, safe and affordable

3DIntermediate
Run a prompt against a labeled test set and let graders score it. Change the prompt or model, run again, and see whether the score went up.
Test SetsLLM-as-JudgePass RateRegressions
Step inside
Intermediate
Run the same 12 questions at four thinking levels. See where a little thinking fixes multi-step problems, and where more just costs time and money.
ReasoningThinking EffortLatencyCost
Intermediate
A ticket router aces its test set, then meets a real week of customer messages. Find what the test set missed, add it, fix it, and check both scores again.
MonitoringLive TrafficRegressionsTest Sets
3DAdvanced
An email assistant reads a poisoned newsletter and leaks internal numbers. Turn on defenses one at a time and see which ones stop it.
Prompt InjectionGuardrailsUntrusted DataLeast Privilege
Step inside
Advanced
Run 12 support questions through a small model, a large one, a router, and a cascade. Plot each run on accuracy vs. cost and find the sweet spot.
Model ChoiceRoutingCascadesUnit Economics
Advanced
A support agent works through 12 requests and takes real actions. Choose which ones wait for a person, and trade bad actions against review time.
ApprovalsAutonomyRisk PolicyHuman Review