Lesson 9 of 10 · Statistics for AI, GenAI & LLMs

Evaluating AI Models Rigorously

You built a model. Now: is it actually good, or does it just look good on the one number you checked?

Lesson 8 was about training a model — minimizing loss, watching for overfitting. This lesson is about the step that comes after training, and the one that decides whether your model ships: evaluation. Specifically, evaluation for classifiers — the yes/no, spam/not-spam, hallucination/not-hallucination decisions that a huge share of practical AI systems boil down to. This is the most directly useful lesson in the course for anyone building and shipping AI products, because a wrong evaluation doesn't just mislead you in a notebook — it ships a broken model to real users with a green checkmark on it.

1. The confusion matrix: your model's report card

Say you build a hallucination detector: given an LLM's answer, it predicts "hallucinated" or "grounded." You test it against 100 answers a human has already labeled correctly. Every prediction lands in exactly one of four buckets, depending on what the model predicted and what was actually true:

Actual: Hallucinated
Actual: Grounded
Predicted:
Hallucinated
TRUE POSITIVE (TP)
Model flagged it, and it really was a hallucination. Caught it.
FALSE POSITIVE (FP)
Model flagged it, but it was actually fine. A false alarm.
Predicted:
Grounded
FALSE NEGATIVE (FN)
Model said it was fine, but it was actually a hallucination. Missed it.
TRUE NEGATIVE (TN)
Model said it was fine, and it really was fine. Correctly passed.

The 2×2 confusion matrix. Rows = what the model predicted. Columns = what was actually true. Green cells are correct; red cells are the two different ways a classifier can be wrong.

Two rows, two columns, four outcomes — that's the whole confusion matrix. Every binary classifier you evaluate, from a content moderation filter to a spam detector to this hallucination checker, produces exactly this table when you run it against labeled test data. The two error types are not the same mistake with two names — they have very different costs:

False Positive cost

A grounded, correct LLM answer gets flagged as a hallucination. You annoy users with unnecessary warnings, or block content that was actually fine.

False Negative cost

A real hallucination slips through unflagged. A user trusts a fabricated fact. This is usually the more dangerous failure.

Key Insight "The model was wrong" isn't specific enough to act on. "The model has a false-negative problem — it's letting real hallucinations through" tells you exactly what to fix, and whether that's even the failure mode you should be worried about for your product.

2. Why accuracy lies — the callback to Lesson 2

The most obvious metric is accuracy: what fraction of predictions were correct?

accuracy = (TP + TN) / (TP + TN + FP + FN)

Remember the 95/5 spam class-imbalance example from Lesson 2? Here it is again, in its native habitat: evaluation. Suppose your content moderation filter runs against 1,000 posts, of which 950 are clean and only 50 are actually policy-violating. A lazy model that predicts "clean" for every single post, never once looking at the content, produces this confusion matrix:

Actual: ViolatingActual: Clean
Predicted: ViolatingTP = 0FP = 0
Predicted: CleanFN = 50TN = 950
accuracy = (0 + 950) / (0 + 950 + 0 + 50)
= 950 / 1000
= 95% accuracy

A model that does zero moderation — that never once catches a single violation — reports 95% accuracy. That's the exact same trap as Lesson 2's spam classifier that predicts "not spam" every time. Accuracy rewards agreeing with whichever class is bigger, and on imbalanced data one class is almost always much bigger. This is the single most common way an AI eval misleads a team into shipping a useless model with an impressive-looking number attached.

Pitfall Never trust a bare accuracy number without first checking the class balance of the eval set. If one class is rare — hallucinations, fraud, policy violations, defects — accuracy can stay high while the model completely fails at the thing you actually built it to catch.

3. Precision and recall: metrics that can't hide behind class imbalance

Precision and recall are both built from the confusion matrix, but each asks a different, specific question — and unlike accuracy, each is defined relative to only one row or one column, so a lazy "always predict the majority class" model cannot hide inside it.

Precision: of everything I flagged, how much was actually real?

precision = TP / (TP + FP)

Precision asks: when your hallucination detector raises a flag, how often is it right? Low precision means the model cries wolf — users learn to ignore its warnings.

Recall: of everything that was actually real, how much did I catch?

recall = TP / (TP + FN)

Recall asks: of all the real hallucinations that existed in the data, what fraction did the model actually catch? Low recall means real problems are slipping through undetected.

Go back to the "predict clean every time" moderation model. Its precision and recall for the violating class:

precision = TP / (TP + FP) = 0 / (0 + 0) → undefined / 0%, no violations ever flagged
recall = TP / (TP + FN) = 0 / (0 + 50) = 0%

Recall of 0% instantly exposes what accuracy hid: this model catches nothing. That's the whole reason precision and recall exist — they refuse to be flattered by an imbalanced dataset the way accuracy does.

F1: one number when you need to balance both

Precision and recall trade off against each other (more on that next), so comparing two models on both numbers separately gets awkward. F1 score combines them into a single number — the harmonic mean, which punishes a big gap between the two more than a plain average would:

F1 = 2 × (precision × recall) / (precision + recall)

A model with precision 0.9 and recall 0.1 gets an F1 around 0.18 — nowhere near the 0.5 a plain average would suggest, because F1 refuses to reward a model that's great at one and terrible at the other. Use F1 as a quick single-number comparison; use precision and recall separately when you need to know which kind of error a model makes.

Key Insight Accuracy asks "how often was I right overall?" — a question a lazy model can game on imbalanced data. Precision and recall each ask a narrower, harder-to-fake question about one specific type of mistake. That's why they're the default metrics for hallucination detectors, spam filters, fraud detection, and content moderation — anywhere the class you care about is the rare one.

4. The precision/recall tradeoff: how cautious should the filter be?

Most classifiers don't output a hard "yes/no" — they output a probability score (e.g., "73% likely this is a hallucination"), and you pick a threshold above which you call it a positive. Move that threshold, and you trade precision for recall or vice versa.

Same 10 predictions, three threshold choices 0.0 (confident: fine) 1.0 (confident: hallucination) ● green = actually grounded   ● red = actually hallucinated low threshold (0.3) flags 8 of 10 → high recall, lower precision medium (0.5) high threshold (0.8) flags only 2 → high precision, lower recall ← below threshold: predicted "fine" above threshold: predicted "hallucination" →

Same 10 scored predictions, three threshold cutoffs. Sliding the threshold left (lower) flags more items — catches more real hallucinations (higher recall) but also flags more grounded answers by mistake (lower precision). Sliding right does the opposite.

Translate the two ends into a product decision:

Low threshold (cautious filter)

Flags anything even slightly suspicious. Catches nearly every real hallucination (high recall) but also flags a lot of fine answers (low precision) — users see frequent, sometimes unnecessary warnings.

High threshold (permissive filter)

Only flags what it's very confident about. Rarely wrong when it does flag something (high precision), but lets more real hallucinations slip through unflagged (low recall).

There's no universally "correct" threshold — it's a product decision, not a math problem. A hallucination detector for a medical-advice chatbot should probably sit at a low, cautious threshold: missing a real hallucination is much worse than an occasional false alarm. A detector that just adds a soft "this might be wrong" badge in a casual chat app can afford a higher threshold and fewer interruptions.

ROC/AUC, briefly

You'll frequently see "ROC curve" or "AUC" in papers and eval dashboards. All it is: instead of picking one threshold and reporting one precision/recall pair, you sweep the threshold across its whole range and plot the tradeoff as a curve. AUC ("area under the curve") compresses that whole curve into one number between 0 and 1 — closer to 1 means the model separates the two classes well across every possible threshold, not just the one you happened to pick. You don't need to derive it by hand; just recognize it as "a way to judge a model's overall discriminating power, independent of any single threshold choice."

5. Is Model B actually better, or is that just noise?

You've swapped in a new prompt and re-run your eval set. The new prompt scores higher. Before you ship it, there's a question the score alone can't answer: is that difference real, or could it have happened by chance even if the two prompts were equally good? This is exactly the sampling problem from Lesson 5, showing up again in a new costume.

The core intuition: p-values, in plain language

A p-value answers one specific question: if there were actually no real difference between Prompt A and Prompt B, how surprising would a result this large be, just from random luck in which examples happened to be in your eval set?

A p-value is not "the probability my result is wrong," and it's not "how much better Model B is." It's purely a measure of how surprising your observed gap would be under the assumption that there's no real difference. That's a subtle distinction, but it's the one that keeps you from over-trusting a lucky-looking eval run.

Why sample size is the whole game — callback to Lesson 5

Lesson 5 showed that small samples give you noisy, untrustworthy estimates of a population, and the Central Limit Theorem is exactly why: your eval score is itself just an estimate, built from a sample of test examples, and estimates from small samples swing wildly by chance alone. A model's "true" score against every possible input it might ever see is unknowable — you're always estimating it from a limited eval set, and a small eval set is a small, noisy sample.

6. Worked example: 75% vs. 80% — real win or noise?

Two prompts, same eval set size, tested on your hallucination detector's downstream task (call it "answer quality: pass/fail" per example, graded by a human or an LLM judge):

ScenarioN (eval examples)Prompt A scorePrompt B scoreRaw gap
Small eval set2075% (15/20 passed)80% (16/20 passed)+5 points
Large eval set50075% (375/500 passed)80% (400/500 passed)+5 points

Same 5-point gap, same percentages — but they tell you very different things.

N = 20: the gap is one flipped example away from vanishing

With only 20 examples, the difference between 75% and 80% is exactly one example — 15 correct vs. 16 correct. Flip a single borderline case (maybe the LLM judge was inconsistent, maybe one test example was ambiguously worded) and the "win" disappears or reverses entirely. A formal significance test on this gap (a two-proportion test) typically returns a p-value well above 0.05 here — nowhere near enough evidence to call Prompt B better. This is indistinguishable from noise.

Pitfall "Prompt B beat Prompt A on my 20-example eval set" is one of the most common false claims in prompt engineering. With N=20, a single flipped example moves the score by 5 full percentage points. You have not learned that B is better — you've learned that B happened to get lucky on this particular small sample.

N = 500: the same gap is now solid evidence

With 500 examples, that same 5-point gap represents 25 more passing examples (400 vs. 375) — a much larger, harder-to-explain-by-luck difference, and the estimate itself is far less noisy because it's built from a much bigger sample (exactly the CLT logic from Lesson 5: more data narrows the spread of your estimate around the true value). A significance test on this version of the gap typically returns a small p-value — this is a real, reproducible difference, not sampling luck.

Why This Matters Same headline number, opposite conclusion, purely because of sample size. This is the single most actionable fact in this lesson: before you trust any "Model/Prompt B beat A" claim — including your own — check how many eval examples it's based on.

Rule of thumb

There's no universal magic number, because it depends on how big a difference you're trying to detect and how much the individual examples vary. But as a working habit: treat comparisons on eval sets under ~30-50 examples as directional hints, not conclusions. For a difference you actually plan to act on — swapping a production prompt, choosing between two fine-tuned models — aim for enough examples that a couple of flipped grades wouldn't change the headline result, and ideally run a real significance test (a two-proportion z-test or a bootstrap resample) rather than eyeballing the percentages.

Checklist: before you believe an eval result

Run through this before trusting any eval score you or your AI teacher hands you:

Recap

✅ Check Yourself

1. Your hallucination detector predicts "grounded" for every single answer and scores 92% accuracy on a test set that's 92% grounded, 8% hallucinated. What does this tell you?

The model is excellent and ready to ship
Nothing useful — this is the class-imbalance trap; the model has 0% recall on hallucinations and catches none of them
92% accuracy always means the model is well-calibrated

2. A content moderation filter has high precision but low recall. What does that mean in practice?

When it flags something, it's usually right — but it's missing a lot of real violations that slip through unflagged
It flags almost everything, including plenty of false alarms
Precision and recall always move together, so this combination is impossible

3. You lower the decision threshold on a hallucination filter (flag more things as suspicious). What's the expected effect?

Both precision and recall go up
Nothing changes — threshold only affects speed, not accuracy metrics
Recall goes up (catches more real hallucinations) but precision goes down (more false alarms)

4. Prompt B scores 5 percentage points higher than Prompt A. Why does it matter whether that gap was measured on 20 examples or 500?

It doesn't matter — a percentage gap means the same thing regardless of sample size
With only 20 examples, a single flipped example can produce the same 5-point gap by pure chance; with 500 examples the same gap requires 25 more correct answers, which is much harder to explain as noise
Larger eval sets always produce bigger score gaps

5. What does a small p-value (e.g., 0.02) tell you when comparing two prompts' eval scores?

There's a 2% chance the model is wrong
If the two prompts were truly equally good, seeing a gap this large by chance alone would be rare — so the observed difference is more likely to be real
Prompt B is exactly 2% better than Prompt A