You never have all the data. Here's the math for how much that should worry you — and how models turn a pile of observations into a fitted distribution.
Lesson 4 gave you the named distributions — Normal, Bernoulli, Categorical, Poisson — as if you could just look at a system and know its parameters. You can't. In real AI work you never see "the true distribution of all possible user prompts" or "the true distribution of all text a model could ever generate." You see a sample: one eval set, one training corpus, one afternoon's worth of logged requests. This lesson is about the gap between what you measured and what's actually true — and about the one tool (Maximum Likelihood Estimation) that turns a sample into a fitted distribution, which is what "training a model" means at a statistical level.
A population is the complete set of everything you'd ideally want to know about. A sample is the subset you actually have. Every AI system you build is estimating something about a population using only a sample:
Population: all text a model could ever see, across every domain, language, and time period. Sample: your training corpus — a few trillion tokens scraped, licensed, or curated, which is a vanishingly small and non-random slice of "all possible text."
Population: all possible user queries your product will ever receive. Sample: the 200 or 2,000 test prompts your eval set actually contains.
Population: every user who will ever hit your product. Sample: the users who happened to be active during the two weeks you ran the test.
Population: what "all humans" would prefer, given a pair of model outputs. Sample: the handful of raters who actually labeled your RLHF data.
This isn't a footnote — it's the condition every number in AI is produced under. "This model scores 87% on our eval" is not a fact about the population of all possible queries. It's a fact about one sample of queries, and samples carry uncertainty that the raw number hides.
Say your product's true, population-level accuracy — if you could somehow test it on every query it will ever receive — is exactly 80%. You don't get to know that number. You only get to run an eval set and compute the sample accuracy. Watch what happens with different eval set sizes, assuming the true rate is 80% and each query is independently right or wrong (a Bernoulli trial, straight out of Lesson 4):
| Eval set size | Plausible range of observed accuracy | What it looks like |
|---|---|---|
| 10 | ~55% – 100% | A single unlucky (or lucky) batch swings the number wildly |
| 30 | ~65% – 93% | Still noisy enough to mislead a launch decision |
| 200 | ~74% – 86% | Tightening, but a ±3-4 point swing is still plausible from chance alone |
| 2,000 | ~78% – 82% | Now a 3-point difference between two models is probably real |
Nothing about the model changed across these rows — the true accuracy is 80% in every one. What changed is how much random noise the sample size lets through. A 10-example eval set can easily hand you "70% accuracy" or "100% accuracy" from a model that's really sitting at 80%, purely because of which 10 queries you happened to draw.
Here's the question that matters: if you keep drawing samples and averaging them, what happens to those averages? The answer is one of the most useful facts in all of statistics.
Imagine repeating an experiment many times: draw an eval set of size N, compute the average score, write it down, repeat. Do that hundreds of times and look at the distribution of those averages. The Central Limit Theorem (CLT) says:
This is genuinely surprising the first time you see it. Individual eval scores are often 0-or-1 (a Bernoulli distribution, not remotely bell-shaped). But the average score over N examples behaves like a Normal distribution once N is reasonably large. This is why the Normal distribution "shows up everywhere" in statistics and AI reporting — it's not that raw data is usually Normal, it's that averages of almost anything become Normal, and averages are what you report.
All three panels are sampling the same underlying system. At N=5 an average eval run could plausibly land anywhere from ~0.5 to ~1.0. By N=200 the averages are tightly packed around the true 0.80 — the shape is Normal in all three cases, just progressively narrower.
Notice what does not change across the three panels: the center. The true score is 0.80 in all three. What shrinks is the spread — how far a single measured average is likely to land from that true center. That shrinking spread is exactly what the next section puts a number on.
Lesson 2 gave you standard deviation: how spread out individual data points are. Standard error is a close cousin, but it answers a different question — not "how spread out are individual scores?" but "how spread out would the average be if I reran this whole eval many times?" The formula is simple:
The √N in the denominator is the CLT made concrete: as N grows, standard error shrinks — but not linearly. To cut your uncertainty in half, you need 4x the data, because of the square root. This is a genuinely important, slightly annoying fact about statistics: early data is cheap to learn from, and each additional point of precision gets more expensive.
Quadrupling N from 25 to 100 halves the standard error. Quadrupling again, from 100 to 400, halves it again. That's the √N tax in action.
A confidence interval takes standard error and turns it into a plain-language range: "the true score is plausibly within about ±2 standard errors of what I measured." (The "2" here is a common rough rule for a 95%-style interval — the precise mechanics of confidence levels are Lesson 9's territory; for now, treat it as an intuitive width, not an exact computation.)
Lesson 4 handed you distributions with parameters already filled in: "a Bernoulli with p = 0.5," "a Normal with mean 100, std dev 15." In the real world, nobody hands you those parameters — you have to estimate them from data. Maximum Likelihood Estimation (MLE) is the single most important idea for how that's done, and it's also, in disguise, what "training a model" means.
Suppose you're checking whether an LLM's binary classifier (say, "is this email spam?") behaves like a fair coin flip or is biased toward one answer. This is a Bernoulli distribution from Lesson 4 — you just don't know its parameter p (the true probability of "spam"). You observe 20 real classification outcomes from a labeled sample: 14 spam, 6 not-spam.
MLE asks: of all possible values of p between 0 and 1, which one makes "14 spam out of 20" the most probable result? Intuitively — and it turns out, mathematically exactly — the answer is just the observed proportion:
That feels almost too obvious to call a "theorem" — and that's exactly the point. MLE formalizes the obvious move (use the observed proportion) and, crucially, generalizes it to cases where the "obvious" answer isn't obvious at all: fitting a Normal distribution's mean and variance simultaneously, fitting dozens of categorical probabilities at once, or fitting the millions of parameters inside a neural network.
Recall from Lesson 4 that an LLM's next-token output is a categorical distribution over the vocabulary. Training data works the same way in miniature. Say you're estimating a simple unigram language model — the probability of each word — from a tiny corpus of 50 observed words, where "the" appeared 8 times, "a" appeared 5 times, and so on for the rest of the vocabulary:
Do this for every word in the vocabulary and you've fit an entire categorical distribution by MLE — each probability is just "how often did I observe this, out of everything I observed." Fitting a Normal distribution's mean and variance to a batch of continuous scores (like the token-count or latency data from Lesson 2) follows the identical logic: the MLE estimate of the mean is the sample mean, and the MLE estimate of the variance is (almost exactly) the sample variance you already know how to compute.
1. Your eval set has only 15 examples, and a new prompt version scores 6 points higher than the old one. What should you conclude?
2. According to the Central Limit Theorem, what happens to the distribution of sample averages as N grows, even if the underlying data isn't Normal?
3. Standard error at N=100 is 0.04. Roughly what happens to standard error if you quadruple your sample to N=400?
4. You observe 30 binary classification outcomes: 21 labeled "positive," 9 labeled "negative." What's the Maximum Likelihood Estimate of the true positive rate p?
5. Why does this lesson connect Maximum Likelihood Estimation to "training a model"?