If your AI product works when you try it and breaks when someone else does, you don't have a model problem yet. You have a testing problem. The fix is smaller than it sounds: a sheet of 30 to 50 real inputs, a plain pass/fail rule for each one, and the habit of running all of them before you change a prompt or a model. An "eval set" is just that list. You run your system on every case, grade the outputs, write down the pass rate, and compare the next version against the same list. This post shows how to build one in about two weeks, without a data team, using one made-up product as the running example.
Why trying a few prompts stops working
Checking a handful of prompts by eye works on day one and quietly fails by week three. Every prompt tweak or model change can fix the case you're staring at and break three you aren't. You only find out when a user does.
Hamel Husain describes this in his widely shared post on why AI products need evals: fixing one failure made others show up, "resembling a game of whack-a-mole," with little visibility beyond vibe checks. He's careful to say vibe checks are useful, just not enough. OpenAI's evaluation best practices guide goes further and lists "vibe-based evals" as an anti-pattern, along with waiting until after you ship to add any evals.
Models can give different outputs for the same input, so "it handled that fine last Tuesday" isn't a test. A sheet is.
The running example (made-up)
MinuteMint is an illustrative, made-up product, and so are all its numbers below. It takes a meeting transcript and returns a short summary plus a list of action items (task, owner, due date if one was said out loud) as JSON that the app shows to the user.
Its one job, as a sentence: turn a messy call transcript into action items a team lead can trust without rereading the call. Everything in the eval set tests that job. Not tone, not how clever the summary sounds.
Step 1: Write down what "good" means before you collect anything
Start with pass and fail rules for your product's one job, in plain words. If you can't say what a failure looks like, you can't test for it.
Anthropic's guide on defining success criteria and building evaluations shows the difference well: "the model should classify sentiments well" is a bad criterion, and a good one is specific and measurable. You don't need F1 scores for a small set, but you do need that plainness. For MinuteMint:
Pass: every action item clearly agreed in the call appears, with the right owner.
Pass: if nobody took ownership, the owner is "unassigned," not a guessed name.
Always fail: any action item that wasn't actually said. An invented task is worse than a missing one, because the user trusts it.
Always fail: output that isn't valid JSON or is missing a required field.
Fail: "next Friday" turned into a wrong calendar date. Leaving it as "next Friday" is fine.
Some of these rules are mechanical (valid JSON) and some need judgment (did it catch every real task?). That split matters in Step 3.
Step 2: Collect 30 to 50 real examples
Pull your first cases from real usage, not your imagination. Real inputs are messier, longer and weirder than anything you'll write at your desk. Good places to look:
Your logs. OpenAI's guide says to log as you develop so you can mine those logs for eval cases. Strip personal data before anything goes into a shared sheet.
Support tickets and user messages. Every "it got this wrong" message is a test case someone wrote for you.
Your own dogfooding. The times you used the product for real and winced.
Edge cases you write on purpose. Anthropic's guide suggests irrelevant or missing input, overly long input, and ambiguous cases where even humans would disagree.
MinuteMint's first set had 40 cases: 15 from logs, 10 from support tickets, 8 from the founder's own calls, and 7 hand-written edge cases (a six-minute call with no decisions, a very long all-hands, two attendees with the same first name). That's 15 + 10 + 8 + 7 = 40.
Why not 500? Because at first you'll read every output yourself. Anthropic's guide leans toward volume, preferring more cases with automated grading over fewer hand-graded ones. Grow in that direction, but start with a set you can grade in one sitting.
Step 3: Pick a grading method for each case
Grade with code wherever you can, with your own eyes where you must, and with an LLM judge only after you've checked it against your own labels.
Code checks for anything mechanical
Format, length, required fields, valid JSON, "list must be empty." These are cheap, fast and never moody. Husain calls them Level 1 tests, plain assertions you run on every code change, like a regex that makes sure internal IDs never leak into replies.
Human review with a short rubric
For anything that needs judgment, you grade it, and you keep it binary. Husain often starts by labeling outputs good or bad, because granular scores are harder to manage. A 1 to 5 scale invites "is this a 3 or a 4?" debates. Pass or fail forces a decision.
LLM-as-judge, used carefully
Once your rubric is stable, a strong model can grade against it. That saves time, but it's a second AI system to test. OpenAI's guide flags known judge biases, like favoring whichever answer comes first and favoring longer answers. It recommends pass/fail or pairwise comparisons, and calibrating the judge against human labels before scaling up. Anthropic's example code notes it's generally best to grade with a different model than the one that produced the output.
For MinuteMint (made-up numbers), every case runs the JSON check. Of the 40 cases, 14 were fully decided by code and 26 needed judgment. The founder graded those 26 by hand first, in roughly 40 minutes, then ran an LLM judge on the same outputs. It agreed on 22 and disagreed on 4. In 3 of the 4, the judge passed outputs the founder had failed, all for invented tasks that sounded plausible. One added line in the judge's instructions ("any task not stated in the transcript is an automatic fail") took agreement to 25 of 26.
One warning from Husain: raw agreement can mislead when most cases pass. A judge that says "pass" to everything will still agree with you most of the time. Check the cases you failed and see whether the judge caught them.
Step 4: Put it all in one sheet
A spreadsheet is enough to start: one row per case, one column per version. Here's a template with five MinuteMint rows:
ID | Input (short) | Source | Pass if | Fail if | Check type | v1 | v2 | v3 |
|---|---|---|---|---|---|---|---|---|
07 | 45-min sales sync, 3 clear tasks | Logs | All 3 tasks, right owners | Any task missing or invented | Human / judge | Pass | Pass | Fail |
12 | Planning call, nobody volunteers | Support ticket | Owner is "unassigned" | A guessed name | Code | Fail | Pass | Pass |
19 | Hindi-English mixed standup | Dogfooding | Tasks in English, names kept | Names translated or dropped | Human / judge | Fail | Pass | Fail |
31 | 6-min call, no decisions | Edge case | Empty list, says none found | Any task at all | Code | Fail | Pass | Pass |
40 | "Let's ship it next Friday" | Edge case | Due date left as "next Friday" | A wrong calendar date | Human / judge | Pass | Fail | Fail |
A second tab logs each run. People skip this part, and it's the part that saves you:
Version | What changed | Cases | Passed | Pass rate | Fixed vs last | Broke vs last | Decision |
|---|---|---|---|---|---|---|---|
v1 | Baseline prompt, model A | 40 | 29 | 72.5% | n/a | n/a | Baseline |
v2 | Prompt rewrite: "unassigned" rule, two worked examples | 40 | 33 | 82.5% | 6 | 2 | Ship, look at the 2 that broke |
v3 | Same prompt, cheaper model B | 40 | 27 | 67.5% | 1 | 7 | Don't ship |
v4 | v2 prompt on model A, date fix, 3 new cases from user reports | 43 | 36 | 83.7% | 3 | 1 | Ship |
The fixed and broke columns are why this beats a single score. v2 went from 29 to 33, but that's 6 fixes and 2 new breaks (29 + 6 − 2 = 33), and those 2 deserve a look. v3 looked like a cost win and broke 7 cases that used to pass (33 + 1 − 7 = 27).
For v4, fixed and broke are counted on the original 40, where it passed 35 (33 + 3 − 1). It also passed 1 of the 3 new cases, for 36 of 43. The percentage barely moved past v2's because the set got harder on purpose. Compare versions on the same cases, not just the headline number.
Step 5: Run the set before every change that touches output
Run the full set before you ship anything that could change what the model says: prompt edits, model swaps, new instructions, changes to the context you feed in, and provider model updates. OpenAI's guide calls this continuous evaluation, running evals on every change and growing the set over time.
The one founders skip is the cost-driven downgrade. A cheaper model looks great on the bill and can quietly fail the cases users care about most, exactly like MinuteMint's v3. Run the set first, then do the math. (The money side is covered in pricing an AI product when every user costs different money.)
How strict should the bar be? Husain notes that, unlike regular unit tests, you don't necessarily need a 100% pass rate. It's a product decision. A simple rule at this size: "always fail" cases stay at zero, and the overall pass rate can't drop versus the current version.
Step 6: Turn every bad output a user reports into a new test case
When a user reports a bad output, add it to the sheet before you fix it. That habit is what makes the set better every week.
Copy the input (cleaned of personal data) into a new row and write its pass and fail rule.
Run it on the current version and confirm it fails. If it doesn't, your rule is wrong or the bug is somewhere else.
Fix the prompt or code, then run the whole set, not just the new case.
Ship only if the new case passes and nothing else broke.
This guide on turning support tickets into product priorities covers the tagging side, and bad AI outputs can be their own tag. When the fix ships, tell the person who reported it. A short release note works, and this piece on writing a changelog that brings silent users back shows how to phrase those so people read them.
When to move beyond a spreadsheet
Stay in a spreadsheet until running the set by hand becomes the reason you skip it. Then a light tool can help. A few options, described from their own docs (not a ranking, and features change):
promptfoo. The promptfoo docs describe an open-source CLI and library for evaluating and red-teaming LLM apps, with declarative test cases, local runs and CI/CD use.
LangSmith. LangChain's docs describe datasets of example inputs with optional reference outputs, offline evaluation for regression testing, and human, code, LLM-as-judge or pairwise evaluators.
Braintrust. Its docs describe tracing AI applications, adding human feedback, building datasets and running evals as experiments.
OpenAI Evals. There's an open-source framework on GitHub and a hosted Evals feature in OpenAI's dashboard. OpenAI's guide currently says the hosted platform is being deprecated: read-only from October 31, 2026, with shutdown scheduled for November 30, 2026. Check its status before building on it.
Whatever you pick, keep the sheet's logic: same cases, clear pass rules, a run log.
A 14-day sprint to get your first eval set running
Two weeks of part-time work is enough. Here's the plan MinuteMint's made-up founder followed:
Days 1 to 2: Write the one job as a sentence, plus pass and "always fail" rules.
Days 3 to 5: Collect 30 to 50 inputs from logs, support, your own use and deliberate edge cases. Remove personal data.
Days 6 to 7: Fill in "pass if" and "fail if" for every row. Mark each as code or human/judge.
Days 8 to 9: Write the code checks (valid JSON, required fields, empty-list cases, length).
Day 10: Run the baseline, grade judgment cases by hand, log it as v1.
Day 11: Try an LLM judge on those cases. Compare against your labels, especially your fails.
Day 12: Run the set on one change you were planning anyway. Log v2 with fixed and broke counts.
Day 13: Start the habit: a reported bad output becomes a new row before any fix.
Day 14: Write the house rule somewhere visible: no prompt or model change ships without a full run.
Common traps
Most eval sets fail for boring reasons, not technical ones:
Testing only happy paths. Clean, well-formed inputs make your pass rate flatter you. OpenAI's guide lists eval data that doesn't reflect real traffic as an anti-pattern. Keep some ugly cases.
Letting the model grade itself. Same model producing and judging, with no human spot-checks, means trusting the thing you're testing. Use a different judge where you can and keep grading a sample yourself.
Chasing one score. A single pass rate hides what broke. Track fixed and broke counts and keep "always fail" rules separate. And a pass rate measures output quality, not whether anyone uses the product.
A giant set nobody maintains. 400 cases with stale rules are worse than 40 you trust. Grow from real failures and delete cases that no longer match the product.
Never re-running it. A set you ran once at launch is a document, not a test.
Quick checklist
One-sentence job for the AI feature, written down
Pass rules and "always fail" rules for that job
30 to 50 real cases, including empty, long and ambiguous inputs
Each case marked code check or human/judge
Baseline run logged with pass rate
LLM judge (if used) checked against your own labels
Fixed and broke counts logged for every version
Every reported bad output added as a case before the fix
Full run before every prompt change, model swap or provider update
If you're about to put your AI product in front of new people, say by launching it on EarlyHunt, those first users will poke at your outputs in ways you didn't plan for. With the sheet ready, each strange result becomes a test case instead of a mystery.
FAQ
How many test cases do I need to start an eval set?
Around 30 to 50 is a practical start for one AI feature: enough for common inputs plus some edge cases, small enough to grade by hand in one sitting. Grow it from real failures rather than padding it up front.
Can I just use an LLM to grade everything?
You can, but check it first. Grade a batch yourself, run the judge on the same outputs, and look at where you disagree, especially on your fails. OpenAI's guide recommends calibrating automated grading against human labels before scaling it up. Repeat that check now and then.
What if my AI product's output is subjective, like writing?
Parts of it usually aren't: format, length, banned content, required elements, facts that must match the input. Test those with code. For the rest, use a short pass/fail rubric tied to the product's one job, or a pairwise check (old version vs new, which is better?) when an absolute grade is hard to call.
How often should I run my eval set?
Before every change that could affect output: prompt edits, model swaps, context or retrieval changes, and provider model updates. If a run takes so long that you skip it, move it into a script or a light tool.