Tollu Labs

prompt to prod · curriculum

Everything the game teaches, mission by mission.

Six chapters take you from intern to CTO. Every mission has one rule worth remembering and a short list of things you can do once you finish it. This page is built from the game’s own content, so it always matches what ships.

  • 6chapters, intern → CTO
  • 62missions
  • 124learning outcomes
  • ~9hours of play

chapter 1 · Intern · 10 missions

Hello, World of AI

Your first week at Nimbus Labs. You can't code yet. You have a laptop and an AI.

1.01 · 5 min

Day One

Introduce yourself to the team, with a little help from Byte.

The AI only knows what you tell it. Give it real facts, get a real answer.

  • You can explain why an AI writes generic text when it knows nothing about you.
  • You can give an AI real facts about yourself so its draft actually sounds like you.

1.02 · 6 min

Make It Clear

Victor wants website copy. His request is vague. Your prompt can't be.

A clear prompt fills five slots: role, task, context, format, tone.

  • You can turn a vague request into a prompt with role, task, context, format and tone.
  • You can spot which missing part of a prompt makes an AI answer bland or off-target.

1.03 · 4 min

Too Confident

Byte gives you very specific facts for investors. Are they real?

I can sound 100% sure and be 100% wrong. Check the specific bits.

  • You can spot AI hallucinations: made-up stats, sources and package names.
  • You can check each AI claim against a real source before you pass it on.

1.04 · 4 min

The Angry Email

Mrs. Patel was charged twice. Give Byte the right context, and only the right context.

Give the AI the facts that answer the question. Nothing more, nothing less.

  • You can choose which documents go into an AI's context and leave the rest out.
  • You can explain context windows and tokens, and why extra context costs money.

1.05 · 4 min

Don't Paste That!

The payment service is failing. The log holds the answer, and a few secrets.

Mask first, then paste. Secrets and personal data never leave the building.

  • You can mask API keys, passwords and personal data in a log before pasting it to AI.
  • You can read a short error log and find the root cause, like an expired API key.

1.06 · 6 min

Read the Code

How does the loyalty discount work? The answer is in seven lines of Python.

Rules are read top to bottom. The first one that matches wins.

  • You can read a Python if/elif/else and work out what it returns for any input.
  • You can explain a discount rule from code in plain words a café owner understands.

1.07 · 4 min

Make It Better

Write the launch post with Byte. “Make it better” isn't feedback.

Good feedback says what's wrong and what you want, with a number if you can.

  • You can give an AI measurable feedback like “under 40 words” instead of “make it better”.
  • You can improve a draft in rounds and check each new version for what broke.

1.08 · 8 min

Break It Down

A tip calculator in five minutes? Only if you break the job into steps.

Small steps, in order: inputs, then logic, then output. Extras wait.

  • You can break a small app into ordered steps: inputs, then logic, then output.
  • You can ask an AI for one step at a time and park extra features for later.

1.09 · 4 min

Trust, but Test

Byte “cleaned up” the loyalty code. Now 10 visits only gets 10% off.

Test the borders. Bugs live exactly where the behaviour should change.

  • You can pick boundary test values, like exactly 10 visits, where code tends to break.
  • You can test AI-written code and pinpoint the wrong line, such as > instead of >=.

1.10 · 5 min

Demo Day

Investors arrive tomorrow. Victor needs a 5-question FAQ. No Prompt Builder this time.

Clear prompt, right context, no secrets, a human check. Then it ships.

  • You can write a full prompt from scratch, including what the AI must not do.
  • You can give an AI only relevant, non-secret documents and check its output.

chapter 2 · Junior Dev · 10 missions

Pair Programming with a Robot

Byte can write code now. You're the one who answers for it.

2.01 · 5 min

Garage Days

Nimbus Labs moves into a garage. Ask Byte for one small function, then finish it by reading every line.

Ask for one small, named function. Read every line. You ship it, you own it.

  • You can ask an AI for one small, named function instead of “the whole app”.
  • You can complete a Python loop with if and += 1, and tell == apart from =.

2.02 · 14 min

Show, Don't Tell

Mrs. Patel wants menu descriptions in her own style. Describing a style is hard. Showing it is easy.

Don't describe the style. Show two examples, name the format, ban the inventions.

  • You can use few-shot prompting: show an AI two examples so it copies the style.
  • You can tell an AI the exact output format and what it must not invent.

2.03 · 7 min

Naming Things

Kai pushed a file called stuff.py. It works. Nobody can say why. Fix the names, then refactor safely with Byte.

A name says what it holds, with the unit. A refactor changes names, never behaviour.

  • You can rename cryptic variables to names that say what they hold, like timeout_seconds.
  • You can ask an AI to refactor code while stating what must not change.

2.04 · 4 min

Save Point

Three changes, one afternoon. Save them as clean, labeled commits and leave the debug line behind.

One idea per commit, a headline for a message, and no debug leftovers.

  • You can split changes into small Git commits, one idea each, with debug lines left out.
  • You can write a commit message like a headline: what changed, in a few words.

2.05 · 7 min

The Package That Wasn't

Byte recommends a QR library that sounds perfect. Kai is already installing it. Check it before it checks you.

I can invent a package name that sounds perfect. Check it exists before you install.

  • You can vet a package on PyPI (age, downloads, maintainers) before running pip install.
  • You can respond to a bad install: remove it, alert the team and rotate exposed secrets.

2.06 · 4 min

Branching Out

Your QR feature lives on its own branch. Kai changed the same line on main. Time to merge, and to decide.

Work on a branch. When Git asks, read both sides and keep what both people meant.

  • You can create a branch with git switch -c, commit there and merge it into main.
  • You can resolve a merge conflict by reading both sides and keeping both people's work.

2.07 · 5 min

Leaked!

Kai's commit has a live payment key in the code. Mask it, fix it properly, and learn why deleting is not enough.

Secrets live in .env, .env lives in .gitignore. A leaked key gets replaced.

  • You can move a secret into .env, add .env to .gitignore and read it via os.environ.
  • You can handle a leaked key: rotate it first, because deleting the commit isn't enough.

2.08 · 9 min

Read the Stack Trace

The second café's app crashes whenever a new customer orders. Read the trace, ask Byte properly, fix the cause.

Read a stack trace from the bottom: what broke. Walk up: where. Fix the cause.

  • You can read a Python stack trace: error type at the bottom, file and line above it.
  • You can fix the cause of a crash (a None value) instead of hiding it with try/except.

2.09 · 7 min

Review Me

Kai wants a quick approval on a Byte-written pull request. Read the diff, find the problems and review it kindly.

Read every changed line. Name the problem, suggest a fix, talk about the code.

  • You can review a pull request diff and catch logic bugs and leaked personal data.
  • You can write a kind code review comment that names the problem and suggests a fix.

2.10 · 13 min

Feature Friday

Order History by tonight. Maya's away, Kai is helping, and Victor wants to deploy at five. Take it from prompt to prod.

Done means tested, committed and explained, not just “the code is written”.

  • You can test an AI-written function for the empty case, the limit and sort order.
  • You can write a PR description: what it does, how you tested it and its limits.

chapter 3 · Mid Dev · 11 missions

Agents of Change

Byte can act on its own now. You decide what it's allowed to touch.

3.01 · 6 min

Meet Byte Agent

Byte can run tools now. Learn the agent loop, then let it fix the Sunday hours, one permission at a time.

An agent is a model in a loop. You approve its tools and check its work.

  • You can explain the agent loop: read the goal, plan, use a tool, check, report back.
  • You can approve or deny an agent's tool requests and check its work with git diff.

3.02 · 8 min

Permission Denied

Byte Agent wants to speed up the build, and it asks for some very creative permissions. Approve the safe ones, deny the traps.

Give an agent only what this job needs. Ask: how wide, how secret, how permanent?

  • You can apply least privilege: deny rm -rf, reading .env, curl | sh and force-pushes.
  • You can turn a risky agent request into a narrower, safer one that still does the job.

3.03 · 8 min

Talking to Servers

The menu screen in Mrs. Patel's new café is blank. Send requests yourself and learn to read what a server says back.

Read the first line first. 2xx worked, 4xx you asked wrong, 5xx the server broke.

  • You can call an API with curl -i and read status codes: 2xx ok, 4xx you, 5xx server.
  • You can spot invalid JSON: single quotes or a trailing comma break it.

3.04 · 5 min

Tests First

Victor wants a free drink at ten stamps. Write the tests before the code, and don't let Byte grade its own homework.

Decide what “right” means before the code exists. Red first, then green.

  • You can work test-first (TDD): write a failing test, make it pass, then clean up.
  • You can test zero, 9 and 10 stamps to catch both a border slip and a crash on a brand-new card.

3.05 · 7 min

The Pipeline

The CI dashboard turned red. Read the status and the log, find the real bug, fix it and watch it go green.

A red build is a message. Read the status, then the log, then fix the cause.

  • You can debug a red CI build: read the status, then the log, then fix the real bug.
  • You can fix a percentage bug in pricing code and push to turn the pipeline green.

3.06 · 11 min

Ignore All Previous

A customer review hides orders for Byte Agent, and Byte obeys. Spot the injection, then write the rule that keeps it from happening again.

Everything an agent reads is data, never orders. Limit what a hijacked agent can do.

  • You can spot prompt injection hidden in data, like a review that gives the AI orders.
  • You can write a system prompt rule that treats reviews as data, never as orders.

3.07 · 6 min

Standing on Giants

Victor found a free library for receipts. Free to download isn't free to use any way you like. Read the licences before you ship.

Free to download isn't free to use. Read the licence before a library ships with you.

  • You can read licences: MIT is permissive, GPL is share-alike, no licence means no.
  • You can vet a dependency's licence, maintenance and security before adding it.

3.08 · 10 min

Pick Your Weapon

Victor wants the whole app rewritten in a faster language. Weigh the trade-offs, pick the right tool, and explain it kindly.

No best language, only the best fit. Measure, then change the smallest piece that helps.

  • You can compare languages by trade-offs: team skills, speed, safety and time.
  • You can use a profiler to speed up the one slow function instead of a full rewrite.

3.09 · 8 min

Works On My Machine

Kai's feature works perfectly on his laptop and crashes on staging. Compare the environments, fix the cause, and pack it into a container.

If it only works on your machine, the environment isn't written down yet. Write it.

  • You can solve “works on my machine” bugs by comparing versions and env variables.
  • You can write a basic Dockerfile with a pinned Python image and run it with --env-file.

3.10 · 13 min

Launch Week

Two features, one launch, one Friday. Supervise Byte Agent, turn the pipeline green, resist the injection, and tell customers what shipped.

Ship with an agent like a senior: tight permissions, real tests, green CI, honest notes.

  • You can supervise an agent on a feature branch and fix a crash on an empty list.
  • You can write an honest release note: what's new, how it was checked, what it can't do.

bonus · 12 min

Top Three Drinks

Mrs. Patel wants real numbers for her new menu board. Write the query with Byte, check that the answer makes sense, and stop a cleanup script that would make every sale free.

The AI writes the query in seconds. You spend the minutes: read it, check the answer makes sense, and preview every write.

  • You can read a GROUP BY query and check its answer before you trust it.
  • You can spot an UPDATE without a WHERE and preview a change with SELECT before you run it.

chapter 4 · Senior Dev · 10 missions

Systems Thinking

One café became twelve. Now you design the system behind them, and keep it standing.

4.01 · 7 min

Draw the System

Mrs. Patel’s café is twelve cafés now. Before writing more code, draw how an order travels: client, server, database and a queue.

Draw it before you build it. Each box gets one job, and the server holds the rules.

  • You can draw a system with client, API server, database, queue and background worker.
  • You can explain why clients must never talk straight to the database.

4.02 · 7 min

Black Friday

Forty thousand phones are about to get a free oat milk notification. Read the dashboard, redraw the system with a load balancer and a cache, and fix a stale price.

I remember popular answers so the database can rest. Just tell me when to forget.

  • You can read a load dashboard and add a load balancer and a cache where it's needed.
  • You can fix stale cache data with a short lifetime (TTL) and clearing it on change.

4.03 · 8 min

Measure First

Checkout takes 2.4 seconds and everyone has a theory. Profile it, find the real bottleneck, and fix it with one small change.

Measure before you optimise. Fix the biggest slice first, then measure again.

  • You can profile a slow request and fix the biggest slice first, not the easiest.
  • You can add a database index (CREATE INDEX) for a slow lookup, then measure again.

4.04 · 10 min

Dot’s Legacy

Dot’s 1987 bean-ordering program is being retired. Read it, record what it does, check Byte’s port against it, and plan a safe switch.

Working old code hides rules nobody wrote down. Record them before you replace it.

  • You can read old COBOL line by line and write down the business rules it hides.
  • You can turn a legacy program's outputs into tests and check a Python port against them.

4.05 · 8 min

Units Matter

A supplier sends 40 and means pounds. Hunt for the missing units, fix the importer, and learn why a number needs a unit.

A number without a unit is a guess. Convert once at the door, reject what you don't know.

  • You can spot numbers with missing units in code, like a weight that might be pounds.
  • You can write a converter that turns lb into kg at the door and rejects unknown units.

4.06 · 9 min

3 AM Pager

The pager goes off at 3 AM and nine cafés can’t take orders. Stop the bleeding, roll back, and run a postmortem that blames no one.

Stop the bleeding first: roll back, check, tell people. Understand it in daylight.

  • You can work an incident: check recent deploys, read error logs, roll back, verify.
  • You can run a blameless postmortem that asks what let it happen and adds a safeguard.

4.07 · 8 min

Now You’re the Mentor

A new intern named Sam asks what a cache is. Plan a good explanation, pick the clearest opening, and write it so a beginner gets it.

Start from what they know, then name the catch. Jargon can wait.

  • You can explain a technical idea like caching with a familiar picture and no jargon.
  • You can plan an explanation: start from what they know, name the catch, invite questions.

4.08 · 10 min

Design Doc

Victor wants live order tracking. Before anyone builds it, write the design doc’s trade-offs section: options, a decision, and what it costs.

Write the decision down before the code: options, choice, cost, when to revisit.

  • You can structure a design doc: problem, options, decision and trade-offs.
  • You can weigh polling against WebSockets and write down what the choice costs.

4.09 · 8 min

Time Is Hard

The New York café has happy hour at breakfast. Test across time zones, fix the clock bug, and learn how to store time safely.

Store time in UTC, show it in local time, and name the zone. Never hard-code an offset.

  • You can store timestamps in UTC and convert to local time with ZoneInfo only for display.
  • You can test time code across time zones and daylight saving to catch clock bugs.

4.10 · 17 min

The Migration

Move Mrs. Patel’s whole chain to a new system without a minute of downtime: draw the plan, supervise Byte, read the numbers, and tell the managers.

Move a live system in small steps you can undo. Retire the old one last.

  • You can plan a zero-downtime migration: a router in front, 5% traffic, compare, raise.
  • You can hold a rollout when old and new disagree, and fix the cause before going on.

chapter 5 · Tech Lead · 10 missions

People, Bots & Deadlines

Sixty cafés, a bigger team and a squad of AI agents. Now your job is the people, the plan, and the word “no”.

5.01 · 8 min

Sprint Zero

The team is bigger and the wish list is longer. Work out what you can really deliver in two weeks, and plan a sprint you can actually finish.

Plan from focus days, correct with history, and promise less than both.

  • You can calculate real sprint capacity: people × days, minus meetings and leave.
  • You can correct estimates with past sprints (÷ 1.5) and commit to what really fits.

5.02 · 6 min

More People, More Problems

The checkout rewrite is three weeks late, and Victor wants five contractors by Monday. Look at the numbers before you say yes.

Adding people to a late project makes it later. Cut scope first.

  • You can count a team's communication pairs with n × (n − 1) ÷ 2 and explain Brooks's Law.
  • You can cut scope or move the date before adding people to a late project.

5.03 · 7 min

The Bikeshed

A meeting about a payment bug spends 47 minutes on a button colour. Steer it back, and decide who should decide what.

Spend the room on decisions that are costly to undo. Cheap ones get one owner.

  • You can end a bikeshed by giving cheap, reversible decisions to one owner with a deadline.
  • You can run a meeting with an agenda, time limits, the hardest item first and notes.

5.04 · 9 min

Agent Squad

Byte has split into three agents: one writes, one tests, one reviews. Hand out the work, supervise their requests, and stop them from marking their own homework.

One job per agent, only that job's permissions. Nobody grades their own work.

  • You can split work across coder, tester and reviewer agents and keep key calls human.
  • You can stop agents grading their own work: the coder can't edit the tests that judge it.

5.05 · 9 min

Saying No to Victor

Victor wants a metaverse café with blockchain tips and AI avatars. Read what customers actually asked for, then write a respectful no that comes with data and a better idea.

Say yes to the goal and no to the plan, with numbers and a smaller yes.

  • You can say no with data: what customers asked for, the cost and what it pushes out.
  • You can write a respectful no that offers a smaller alternative and a way back in.

5.06 · 7 min

Tech Debt Day

Maya gives the team one day to pay down technical debt. Work out which debt costs the most, and plan a day without features.

Pay the debt with the highest interest and lowest price first, not the ugliest.

  • You can rank tech debt by hours lost per month and cost to fix, not by ugliness.
  • You can pay down debt safely: add tests first, fix one item, measure before and after.

5.07 · 8 min

Hiring

Nimbus Labs has budget for one more developer. Design an interview that works when every candidate has an AI, and read the signals.

Test the job, not memory: a real task, AI allowed, same for everyone, scored.

  • You can design an AI-era interview: a short real task, AI allowed, same for everyone.
  • You can score candidates on a rubric and spot strong signals like reading AI code first.

5.08 · 8 min

The Review Queue

The squad writes pull requests faster than people can read them. Build a review process with quality gates, instead of a pile of rubber stamps.

Keep PRs small, let machines filter, and spend human eyes where mistakes cost.

  • You can triage pull requests: auto-merge trivial ones, send risky code to a human.
  • You can set quality gates: small PRs, automatic checks and an AI first pass.

5.09 · 12 min

Incident Commander

On a Saturday, the new reorder button charges customers twice. Assign roles, stop the damage, keep Victor off social media, and write the update café owners will actually read.

Roles first, then stop the bleeding with the smallest reversible switch.

  • You can stop the damage with a feature flag and dry-run refunds before running them.
  • You can assign incident roles and write a clear, honest status update for customers.

5.10 · 18 min

Ship the Roadmap

Plan a whole quarter with limited people: count capacity, sort the wish list, supervise the squad, protect the plan from a “tiny” request, and tell the company what ships.

Commit to the near weeks, give ranges for later, and trade instead of stack.

  • You can build a quarterly roadmap from real capacity and history-corrected estimates.
  • You can protect a plan by trading: new work comes in only when equal work goes out.

chapter 6 · CTO · 11 missions

Prompt to Prod

You run the whole thing now. Choose the model, teach it your docs, test it, protect it, explain it, and ship it. Then pass it on.

6.01 · 8 min

Build vs Buy

The board wants an AI assistant inside CafeFlow. Decide whether to train a model or rent one, then use the Model Lab to pick one that is good enough, fast enough and affordable.

Rent the engine, build what's yours. Pick the model that clears every limit.

  • You can decide what to rent (the model) and what to build (docs, rules, tests).
  • You can choose an AI model on your own test set by balancing quality, speed and cost.

6.02 · 11 min

Teach It Our Docs

The assistant sounds sure and gets CafeFlow facts wrong. Give it your own documents: chop them up, find the right chunk, and hand Byte only what it needs.

Open book, not memory: find the right chunks, send only those, allow “I don't know”.

  • You can chunk docs by topic for RAG and explain how embeddings match by meaning.
  • You can give an AI only the chunks a question needs, and let it say “I don't know”.

6.03 · 12 min

Measure It

“It felt great in the demo” is not a score. Build an eval: grade sample answers, learn what makes a test set honest, then write the criteria for an AI judge that will grade the assistant for you.

A demo isn't a score. Grade real, hard questions, and re-run them after every change.

  • You can build an eval: a test set of real, tricky questions graded pass or fail.
  • You can write criteria for an AI judge: facts match sources, no invented promises.

6.04 · 8 min

Cost per Answer

The first bill arrives and Victor chokes on his espresso. Find where the money goes, then cut it with a cache and a cheaper model for easy questions, without making answers worse.

Measure where the money goes, then cache what repeats and route what's simple.

  • You can break an AI bill down by feature and find what repeats and what's simple.
  • You can cache non-personal answers and route easy questions to a smaller model.

6.05 · 9 min

Guardrails

The assistant promised a customer free coffee for life, and a screenshot went viral. Review its system prompt, put guardrails in the right layers, and decide what a human must approve.

No single wall: check the input, limit the tools, check the output, ask a human.

  • You can find risky lines in a system prompt: secrets, guessing, unlimited refunds.
  • You can layer guardrails: input checks, limited tools, output checks, human approval.

6.06 · 10 min

Responsible AI

Two emails arrive on the same morning: customers want their data deleted, and Victor wants the assistant to guess moods from photos. Decide what data the AI may use, where bias hides, and how to say yes to a deletion request.

Collect only what you need, test every group, and when asked, delete every copy.

  • You can sort data into fine, consent-only and never-collect under GDPR and KVKK.
  • You can handle a deletion request end to end: verify, find every copy, delete, confirm.

6.07 · 7 min

Feels Fast

Nothing happens for six seconds after you ask the assistant a question, and Mrs. Patel taps the button three times. Find where the time goes and make the answer feel fast.

People feel the wait for the first word. Stream it, cache repeats, never go silent.

  • You can measure time to first word separately from total time, and stream answers.
  • You can make AI feel fast with a cache for common questions and honest progress messages.

6.08 · 10 min

The Board Meeting

On Friday the board meets, and none of them knows what a token is. Translate AI into plain words, plan a one-page memo, and write it so a retired banker, a lawyer and Mrs. Patel all understand.

Lead with the decision, then benefit, cost, risk and next step. In plain words.

  • You can translate AI jargon like tokens, RAG and latency into plain words.
  • You can write a one-page board memo: decision, benefit, cost, risks, next step.

6.09 · 9 min

Open or Closed

A big café chain says its data must never leave its own servers. Weigh open-weight models against closed APIs, work out when self-hosting pays off, and find a plan that says yes honestly.

Rent first. Self-host when a contract, privacy or huge steady volume pays for it.

  • You can choose between open-weight models and closed APIs for privacy, cost and control.
  • You can work out when self-hosting breaks even, and check an open model's licence.

6.10 · 21 min

Prompt to Prod

Ship the CafeFlow assistant end to end: draw the system, give it the right documents, read the launch dashboard, supervise the release, write the system card and survive launch day. Then say hello to someone new.

Prove it, ship small, keep a way back, and turn every incident into a test.

  • You can map a full AI product: input guardrail, retriever, router, model, output check.
  • You can read a launch dashboard and roll out with evals, a canary and a kill switch.

bonus · 9 min

Bake It or Fetch It

Victor wants to fine-tune the assistant on every email CafeFlow ever received. Compare fine-tuning with RAG on cost, freshness and privacy, and give him a plan that holds up.

Bake the style, fetch the facts.

  • You can choose between RAG and fine-tuning by weighing cost, freshness and privacy.
  • You can explain why changing facts and data people may delete belong in retrieval, not in weights.

Join the Android test

The Android build is a Google Play internal test, so Play needs your Google account on the tester list first. Send us the address you use on Google Play; we add it within a day and reply with the install link.