You built a model. Now: is it actually good, or does it just look good on the one number you checked?
Lesson 8 was about training a model — minimizing loss, watching for overfitting. This lesson is about the step that comes after training, and the one that decides whether your model ships: evaluation. Specifically, evaluation for classifiers — the yes/no, spam/not-spam, hallucination/not-hallucination decisions that a huge share of practical AI systems boil down to. This is the most directly useful lesson in the course for anyone building and shipping AI products, because a wrong evaluation doesn't just mislead you in a notebook — it ships a broken model to real users with a green checkmark on it.
Say you build a hallucination detector: given an LLM's answer, it predicts "hallucinated" or "grounded." You test it against 100 answers a human has already labeled correctly. Every prediction lands in exactly one of four buckets, depending on what the model predicted and what was actually true:
The 2×2 confusion matrix. Rows = what the model predicted. Columns = what was actually true. Green cells are correct; red cells are the two different ways a classifier can be wrong.
Two rows, two columns, four outcomes — that's the whole confusion matrix. Every binary classifier you evaluate, from a content moderation filter to a spam detector to this hallucination checker, produces exactly this table when you run it against labeled test data. The two error types are not the same mistake with two names — they have very different costs:
A grounded, correct LLM answer gets flagged as a hallucination. You annoy users with unnecessary warnings, or block content that was actually fine.
A real hallucination slips through unflagged. A user trusts a fabricated fact. This is usually the more dangerous failure.
The most obvious metric is accuracy: what fraction of predictions were correct?
Remember the 95/5 spam class-imbalance example from Lesson 2? Here it is again, in its native habitat: evaluation. Suppose your content moderation filter runs against 1,000 posts, of which 950 are clean and only 50 are actually policy-violating. A lazy model that predicts "clean" for every single post, never once looking at the content, produces this confusion matrix:
| Actual: Violating | Actual: Clean | |
|---|---|---|
| Predicted: Violating | TP = 0 | FP = 0 |
| Predicted: Clean | FN = 50 | TN = 950 |
A model that does zero moderation — that never once catches a single violation — reports 95% accuracy. That's the exact same trap as Lesson 2's spam classifier that predicts "not spam" every time. Accuracy rewards agreeing with whichever class is bigger, and on imbalanced data one class is almost always much bigger. This is the single most common way an AI eval misleads a team into shipping a useless model with an impressive-looking number attached.
Precision and recall are both built from the confusion matrix, but each asks a different, specific question — and unlike accuracy, each is defined relative to only one row or one column, so a lazy "always predict the majority class" model cannot hide inside it.
Precision asks: when your hallucination detector raises a flag, how often is it right? Low precision means the model cries wolf — users learn to ignore its warnings.
Recall asks: of all the real hallucinations that existed in the data, what fraction did the model actually catch? Low recall means real problems are slipping through undetected.
Go back to the "predict clean every time" moderation model. Its precision and recall for the violating class:
Recall of 0% instantly exposes what accuracy hid: this model catches nothing. That's the whole reason precision and recall exist — they refuse to be flattered by an imbalanced dataset the way accuracy does.
Precision and recall trade off against each other (more on that next), so comparing two models on both numbers separately gets awkward. F1 score combines them into a single number — the harmonic mean, which punishes a big gap between the two more than a plain average would:
A model with precision 0.9 and recall 0.1 gets an F1 around 0.18 — nowhere near the 0.5 a plain average would suggest, because F1 refuses to reward a model that's great at one and terrible at the other. Use F1 as a quick single-number comparison; use precision and recall separately when you need to know which kind of error a model makes.
Most classifiers don't output a hard "yes/no" — they output a probability score (e.g., "73% likely this is a hallucination"), and you pick a threshold above which you call it a positive. Move that threshold, and you trade precision for recall or vice versa.
Same 10 scored predictions, three threshold cutoffs. Sliding the threshold left (lower) flags more items — catches more real hallucinations (higher recall) but also flags more grounded answers by mistake (lower precision). Sliding right does the opposite.
Translate the two ends into a product decision:
Flags anything even slightly suspicious. Catches nearly every real hallucination (high recall) but also flags a lot of fine answers (low precision) — users see frequent, sometimes unnecessary warnings.
Only flags what it's very confident about. Rarely wrong when it does flag something (high precision), but lets more real hallucinations slip through unflagged (low recall).
There's no universally "correct" threshold — it's a product decision, not a math problem. A hallucination detector for a medical-advice chatbot should probably sit at a low, cautious threshold: missing a real hallucination is much worse than an occasional false alarm. A detector that just adds a soft "this might be wrong" badge in a casual chat app can afford a higher threshold and fewer interruptions.
You'll frequently see "ROC curve" or "AUC" in papers and eval dashboards. All it is: instead of picking one threshold and reporting one precision/recall pair, you sweep the threshold across its whole range and plot the tradeoff as a curve. AUC ("area under the curve") compresses that whole curve into one number between 0 and 1 — closer to 1 means the model separates the two classes well across every possible threshold, not just the one you happened to pick. You don't need to derive it by hand; just recognize it as "a way to judge a model's overall discriminating power, independent of any single threshold choice."
You've swapped in a new prompt and re-run your eval set. The new prompt scores higher. Before you ship it, there's a question the score alone can't answer: is that difference real, or could it have happened by chance even if the two prompts were equally good? This is exactly the sampling problem from Lesson 5, showing up again in a new costume.
A p-value answers one specific question: if there were actually no real difference between Prompt A and Prompt B, how surprising would a result this large be, just from random luck in which examples happened to be in your eval set?
A p-value is not "the probability my result is wrong," and it's not "how much better Model B is." It's purely a measure of how surprising your observed gap would be under the assumption that there's no real difference. That's a subtle distinction, but it's the one that keeps you from over-trusting a lucky-looking eval run.
Lesson 5 showed that small samples give you noisy, untrustworthy estimates of a population, and the Central Limit Theorem is exactly why: your eval score is itself just an estimate, built from a sample of test examples, and estimates from small samples swing wildly by chance alone. A model's "true" score against every possible input it might ever see is unknowable — you're always estimating it from a limited eval set, and a small eval set is a small, noisy sample.
Two prompts, same eval set size, tested on your hallucination detector's downstream task (call it "answer quality: pass/fail" per example, graded by a human or an LLM judge):
| Scenario | N (eval examples) | Prompt A score | Prompt B score | Raw gap |
|---|---|---|---|---|
| Small eval set | 20 | 75% (15/20 passed) | 80% (16/20 passed) | +5 points |
| Large eval set | 500 | 75% (375/500 passed) | 80% (400/500 passed) | +5 points |
Same 5-point gap, same percentages — but they tell you very different things.
With only 20 examples, the difference between 75% and 80% is exactly one example — 15 correct vs. 16 correct. Flip a single borderline case (maybe the LLM judge was inconsistent, maybe one test example was ambiguously worded) and the "win" disappears or reverses entirely. A formal significance test on this gap (a two-proportion test) typically returns a p-value well above 0.05 here — nowhere near enough evidence to call Prompt B better. This is indistinguishable from noise.
With 500 examples, that same 5-point gap represents 25 more passing examples (400 vs. 375) — a much larger, harder-to-explain-by-luck difference, and the estimate itself is far less noisy because it's built from a much bigger sample (exactly the CLT logic from Lesson 5: more data narrows the spread of your estimate around the true value). A significance test on this version of the gap typically returns a small p-value — this is a real, reproducible difference, not sampling luck.
There's no universal magic number, because it depends on how big a difference you're trying to detect and how much the individual examples vary. But as a working habit: treat comparisons on eval sets under ~30-50 examples as directional hints, not conclusions. For a difference you actually plan to act on — swapping a production prompt, choosing between two fine-tuned models — aim for enough examples that a couple of flipped grades wouldn't change the headline result, and ideally run a real significance test (a two-proportion z-test or a bootstrap resample) rather than eyeballing the percentages.
1. Your hallucination detector predicts "grounded" for every single answer and scores 92% accuracy on a test set that's 92% grounded, 8% hallucinated. What does this tell you?
2. A content moderation filter has high precision but low recall. What does that mean in practice?
3. You lower the decision threshold on a hallucination filter (flag more things as suspicious). What's the expected effect?
4. Prompt B scores 5 percentage points higher than Prompt A. Why does it matter whether that gap was measured on 20 examples or 500?
5. What does a small p-value (e.g., 0.02) tell you when comparing two prompts' eval scores?