You run the whole thing now. Choose the model, teach it your docs, test it, protect it, explain it, and ship it. Then pass it on.
6.01 · 8 min
Build vs Buy
The board wants an AI assistant inside CafeFlow. Decide whether to train a model or rent one, then use the Model Lab to pick one that is good enough, fast enough and affordable.
Rent the engine, build what's yours. Pick the model that clears every limit.
- You can decide what to rent (the model) and what to build (docs, rules, tests).
- You can choose an AI model on your own test set by balancing quality, speed and cost.
6.02 · 11 min
Teach It Our Docs
The assistant sounds sure and gets CafeFlow facts wrong. Give it your own documents: chop them up, find the right chunk, and hand Byte only what it needs.
Open book, not memory: find the right chunks, send only those, allow “I don't know”.
- You can chunk docs by topic for RAG and explain how embeddings match by meaning.
- You can give an AI only the chunks a question needs, and let it say “I don't know”.
6.03 · 12 min
Measure It
“It felt great in the demo” is not a score. Build an eval: grade sample answers, learn what makes a test set honest, then write the criteria for an AI judge that will grade the assistant for you.
A demo isn't a score. Grade real, hard questions, and re-run them after every change.
- You can build an eval: a test set of real, tricky questions graded pass or fail.
- You can write criteria for an AI judge: facts match sources, no invented promises.
6.04 · 8 min
Cost per Answer
The first bill arrives and Victor chokes on his espresso. Find where the money goes, then cut it with a cache and a cheaper model for easy questions, without making answers worse.
Measure where the money goes, then cache what repeats and route what's simple.
- You can break an AI bill down by feature and find what repeats and what's simple.
- You can cache non-personal answers and route easy questions to a smaller model.
6.05 · 9 min
Guardrails
The assistant promised a customer free coffee for life, and a screenshot went viral. Review its system prompt, put guardrails in the right layers, and decide what a human must approve.
No single wall: check the input, limit the tools, check the output, ask a human.
- You can find risky lines in a system prompt: secrets, guessing, unlimited refunds.
- You can layer guardrails: input checks, limited tools, output checks, human approval.
6.06 · 10 min
Responsible AI
Two emails arrive on the same morning: customers want their data deleted, and Victor wants the assistant to guess moods from photos. Decide what data the AI may use, where bias hides, and how to say yes to a deletion request.
Collect only what you need, test every group, and when asked, delete every copy.
- You can sort data into fine, consent-only and never-collect under GDPR and KVKK.
- You can handle a deletion request end to end: verify, find every copy, delete, confirm.
6.07 · 7 min
Feels Fast
Nothing happens for six seconds after you ask the assistant a question, and Mrs. Patel taps the button three times. Find where the time goes and make the answer feel fast.
People feel the wait for the first word. Stream it, cache repeats, never go silent.
- You can measure time to first word separately from total time, and stream answers.
- You can make AI feel fast with a cache for common questions and honest progress messages.
6.08 · 10 min
The Board Meeting
On Friday the board meets, and none of them knows what a token is. Translate AI into plain words, plan a one-page memo, and write it so a retired banker, a lawyer and Mrs. Patel all understand.
Lead with the decision, then benefit, cost, risk and next step. In plain words.
- You can translate AI jargon like tokens, RAG and latency into plain words.
- You can write a one-page board memo: decision, benefit, cost, risks, next step.
6.09 · 9 min
Open or Closed
A big café chain says its data must never leave its own servers. Weigh open-weight models against closed APIs, work out when self-hosting pays off, and find a plan that says yes honestly.
Rent first. Self-host when a contract, privacy or huge steady volume pays for it.
- You can choose between open-weight models and closed APIs for privacy, cost and control.
- You can work out when self-hosting breaks even, and check an open model's licence.
6.10 · 21 min
Prompt to Prod
Ship the CafeFlow assistant end to end: draw the system, give it the right documents, read the launch dashboard, supervise the release, write the system card and survive launch day. Then say hello to someone new.
Prove it, ship small, keep a way back, and turn every incident into a test.
- You can map a full AI product: input guardrail, retriever, router, model, output check.
- You can read a launch dashboard and roll out with evals, a canary and a kill switch.
bonus · 9 min
Bake It or Fetch It
Victor wants to fine-tune the assistant on every email CafeFlow ever received. Compare fine-tuning with RAG on cost, freshness and privacy, and give him a plan that holds up.
Bake the style, fetch the facts.
- You can choose between RAG and fine-tuning by weighing cost, freshness and privacy.
- You can explain why changing facts and data people may delete belong in retrieval, not in weights.