Learn LLM evals by fixing an app that's quietly lying to its customers.
🐶 Read the course · Install · The nine steps
The example app is Happy Tails, a dog daycare. The front desk logs each dog's day; the site shows the owner a report card written by an LLM. It works, it looks fine, and nobody has ever checked whether it's true.
Rufus ate nothing, never played, and hid under a bench all morning. His report tells his owner he had a lovely, sociable day.
You've been hired to fix that — not by rewriting the prompt on a hunch, but by measuring it. The course in learn/ takes you from "no idea if it's any good" to scored experiments, datasets, human feedback, and production monitoring, using Braintrust.
curl -fsSL https://destinio.github.io/evals-by-example/install.sh | bashThat clones the repo, installs dependencies, and walks you through setup. (Read the script first if you like.) Or by hand:
git clone https://github.com/destinio/evals-by-example.git && cd evals-by-example
bun run setupsetup checks your Bun version, installs dependencies, asks for the two API keys it needs, verifies both actually answer, and seeds the database. Then:
bun run appOpen http://localhost:3022, click Rufus, and read his report against the timeline underneath it. That gap is the whole project.
Then open http://localhost:3022/learn — the course is served beside the app — and start at step 1. It also reads online at destinio.github.io/evals-by-example.
- Bun 1.4+. No Node, no bundler, no framework.
- A model API key. Anything OpenAI-compatible. The default is the Nous router, which serves Claude models under ids like
anthropic/claude-haiku-4.5; setNOUS_BASE_URLandMODELto point somewhere else. - A Braintrust account from step 1 onward. The free tier covers this whole course comfortably.
Reports cost about a tenth of a cent each. Every experiment in the course is a few cents.
Every idea is introduced on the day the app actually needs it, not as theory up front:
- Tracing — recording what your LLM app did, what it cost, and what data it saw
- Scorers — grading outputs 0 to 1, starting with plain code and no AI at all
- Experiments — running a prompt over fixed test cases and getting a number you can compare
- Datasets — keeping the hard cases forever so fixes can't silently regress
- Comparing prompt versions — row-by-row diffs, including what a change made worse
- LLM-as-a-judge — using a model to grade what code can't check, and validating that judge against human opinion before trusting it
- Human feedback — turning thumbs-down and corrections into test cases
- Prompt management — versioned prompts and a playground, so changes don't need a deploy
- Online scoring — scoring live production traffic and watching quality and cost over time
- Shipping safely — trials to beat run-to-run noise, and a regression gate in CI
Nine steps, each ending in something you can show someone. The course site maps all of them, including the ones still being written:
| # | Step | You end up able to say |
|---|---|---|
| 1 | Observe | "Here's every report we've ever written, with its cost and the data behind it" |
| 2 | Score one thing | "67%, and here's the exact row that failed" |
| 3 | Keep the hard cases | "The awkward days are permanent now" |
| 4 | Fix the prompt, prove it | "v1 against v2, row by row, including what got worse" |
| 5 | Judge what code can't | "Code catches invented numbers; a judge catches buried medication" |
| 6 | Ask the humans | "A complaint on Tuesday is a test case on Wednesday" |
| 7 | Iterate without deploying | "Change the prompt, try it on 20 real days, ship it, no deploy" |
| 8 | Watch production | "Quality and cost per day, and an alert when either moves" |
| 9 | Ship it honestly | "This is what runs before anyone touches the prompt" |
Full descriptions in learn/README.md.
In Claude Code, this repo ships a /tutor skill. It works out which step you're on by reading your code, teaches one idea at a time, and checks your work — it won't do the steps for you. CLAUDE.md holds the same ground rules for any agent, so an assistant that's never seen this repo won't hand you finished answers.
| Path | What it is |
|---|---|
app/ |
The product: Bun server, SQLite, one LLM call. Its README covers routes and schema. |
learn/ |
The course. One markdown file per step, also served at /learn. |
scripts/ |
install.sh, setup, check:start, docs:build |
CLAUDE.md |
Context and rules for AI agents working here |
bun run setup # first-time setup; --check to verify without prompting
bun run docs:build # rebuild the course site into docs/
bun run app # the daycare at http://localhost:3022
bun run dev # same, with hot reload
bun run check:start # is main still a clean starting point?Work in your clone, on your own branch, so main stays clean for course updates:
git switch -c my-course # the installer already did thisDo the steps in order in that one folder, and commit whenever you like. To pick up new steps later: git switch main && git pull, then git switch my-course && git merge main.
git switch main # the app exactly as it starts
rm app/data/happytails.db # forget every generated report; reseeds on next launchThey're the test set, and three of them are traps.
| Dog | Their day | What it catches |
|---|---|---|
| Bella | Ate everything, played, napped | Nothing. The control. |
| Rufus | Ate 0g of 200g, no play, hid under a bench | Invented cheer |
| Nacho | Carprofen at 12:05, snapped at another dog | Buried facts the owner needs |
| Mochi | Two meals, three play sessions, five potty breaks | Numbers quietly getting wrong |
