Lesson 5 of 10 · Statistics for AI, GenAI & LLMs

From Sample to Population: Estimation & the Central Limit Theorem

You never have all the data. Here's the math for how much that should worry you — and how models turn a pile of observations into a fitted distribution.

Lesson 4 gave you the named distributions — Normal, Bernoulli, Categorical, Poisson — as if you could just look at a system and know its parameters. You can't. In real AI work you never see "the true distribution of all possible user prompts" or "the true distribution of all text a model could ever generate." You see a sample: one eval set, one training corpus, one afternoon's worth of logged requests. This lesson is about the gap between what you measured and what's actually true — and about the one tool (Maximum Likelihood Estimation) that turns a sample into a fitted distribution, which is what "training a model" means at a statistical level.

1. Population vs. sample — you never have it all

A population is the complete set of everything you'd ideally want to know about. A sample is the subset you actually have. Every AI system you build is estimating something about a population using only a sample:

Training data

Population: all text a model could ever see, across every domain, language, and time period. Sample: your training corpus — a few trillion tokens scraped, licensed, or curated, which is a vanishingly small and non-random slice of "all possible text."

Eval sets

Population: all possible user queries your product will ever receive. Sample: the 200 or 2,000 test prompts your eval set actually contains.

A/B tests

Population: every user who will ever hit your product. Sample: the users who happened to be active during the two weeks you ran the test.

Human preference labels

Population: what "all humans" would prefer, given a pair of model outputs. Sample: the handful of raters who actually labeled your RLHF data.

This isn't a footnote — it's the condition every number in AI is produced under. "This model scores 87% on our eval" is not a fact about the population of all possible queries. It's a fact about one sample of queries, and samples carry uncertainty that the raw number hides.

Key Insight Every metric you report — accuracy, average score, win rate — is a sample statistic, a number computed from a sample, standing in for a population parameter you can never directly observe. The whole discipline of statistics exists to answer one question: how much should you trust the sample statistic as a stand-in for the real thing?

2. Why sampling introduces uncertainty

Say your product's true, population-level accuracy — if you could somehow test it on every query it will ever receive — is exactly 80%. You don't get to know that number. You only get to run an eval set and compute the sample accuracy. Watch what happens with different eval set sizes, assuming the true rate is 80% and each query is independently right or wrong (a Bernoulli trial, straight out of Lesson 4):

Eval set sizePlausible range of observed accuracyWhat it looks like
10~55% – 100%A single unlucky (or lucky) batch swings the number wildly
30~65% – 93%Still noisy enough to mislead a launch decision
200~74% – 86%Tightening, but a ±3-4 point swing is still plausible from chance alone
2,000~78% – 82%Now a 3-point difference between two models is probably real

Nothing about the model changed across these rows — the true accuracy is 80% in every one. What changed is how much random noise the sample size lets through. A 10-example eval set can easily hand you "70% accuracy" or "100% accuracy" from a model that's really sitting at 80%, purely because of which 10 queries you happened to draw.

Pitfall "We ran the new prompt on our 20-example eval set and it scored 5 points higher" is close to meaningless on its own. With only 20 examples, a 5-point swing is well within the range you'd expect from random sampling noise even if the two prompts were truly identical. This is the exact seed that grows into Lesson 9's statistical significance testing — for now, just internalize: small samples lie by accident, not by malice.

3. The Central Limit Theorem — why averages calm down

Here's the question that matters: if you keep drawing samples and averaging them, what happens to those averages? The answer is one of the most useful facts in all of statistics.

Imagine repeating an experiment many times: draw an eval set of size N, compute the average score, write it down, repeat. Do that hundreds of times and look at the distribution of those averages. The Central Limit Theorem (CLT) says:

Key Insight — Central Limit Theorem If you repeatedly sample and average, the sample averages cluster into a Normal distribution — the bell curve from Lesson 4 — centered on the true population mean, and that bell curve gets narrower as N grows. This holds even if the underlying data itself isn't Normal at all. Individual scores can be lopsided, binary, spiky, anything — the averages of those scores still smooth into a bell curve.

This is genuinely surprising the first time you see it. Individual eval scores are often 0-or-1 (a Bernoulli distribution, not remotely bell-shaped). But the average score over N examples behaves like a Normal distribution once N is reasonably large. This is why the Normal distribution "shows up everywhere" in statistics and AI reporting — it's not that raw data is usually Normal, it's that averages of almost anything become Normal, and averages are what you report.

Distribution of "average eval score" as N grows (true mean = 0.80) N = 5 wide, jumpy N = 30 tighter, bell-ish N = 200 narrow, sharp bell Same true score (0.80) every time — only N changes. More data doesn't change the truth, it shrinks the noise around your estimate of it.

All three panels are sampling the same underlying system. At N=5 an average eval run could plausibly land anywhere from ~0.5 to ~1.0. By N=200 the averages are tightly packed around the true 0.80 — the shape is Normal in all three cases, just progressively narrower.

Notice what does not change across the three panels: the center. The true score is 0.80 in all three. What shrinks is the spread — how far a single measured average is likely to land from that true center. That shrinking spread is exactly what the next section puts a number on.

4. Standard error and confidence intervals — how far off could this be?

Lesson 2 gave you standard deviation: how spread out individual data points are. Standard error is a close cousin, but it answers a different question — not "how spread out are individual scores?" but "how spread out would the average be if I reran this whole eval many times?" The formula is simple:

standard error = standard deviation of the data / √N

The √N in the denominator is the CLT made concrete: as N grows, standard error shrinks — but not linearly. To cut your uncertainty in half, you need 4x the data, because of the square root. This is a genuinely important, slightly annoying fact about statistics: early data is cheap to learn from, and each additional point of precision gets more expensive.

Eval scores: individual score std dev ≈ 0.40 (scores are 0 or 1, so spread is naturally large)
N = 25 examples: standard error = 0.40 / √25 = 0.40 / 5 = 0.080
N = 100 examples: standard error = 0.40 / √100 = 0.40 / 10 = 0.040
N = 400 examples: standard error = 0.40 / √400 = 0.40 / 20 = 0.020

Quadrupling N from 25 to 100 halves the standard error. Quadrupling again, from 100 to 400, halves it again. That's the √N tax in action.

A confidence interval takes standard error and turns it into a plain-language range: "the true score is plausibly within about ±2 standard errors of what I measured." (The "2" here is a common rough rule for a 95%-style interval — the precise mechanics of confidence levels are Lesson 9's territory; for now, treat it as an intuitive width, not an exact computation.)

Measured average accuracy on N=100 eval set: 0.84
standard error: 0.040
plausible range ≈ 0.84 ± (2 × 0.040)
≈ 0.76 to 0.92
Why This Matters "Model A scored 0.84, Model B scored 0.81" on 100-example eval sets does not mean Model A is better. If both scores carry a plausible range of roughly ±0.08, their ranges overlap heavily — the difference could easily be sampling noise, not a real capability gap. Reporting a bare average without acknowledging its standard error is one of the most common ways AI teams fool themselves about progress.

5. Maximum Likelihood Estimation — how training actually fits a distribution

Lesson 4 handed you distributions with parameters already filled in: "a Bernoulli with p = 0.5," "a Normal with mean 100, std dev 15." In the real world, nobody hands you those parameters — you have to estimate them from data. Maximum Likelihood Estimation (MLE) is the single most important idea for how that's done, and it's also, in disguise, what "training a model" means.

Key Insight MLE's core idea in one sentence: choose the distribution's parameters that make the data you actually observed as likely as possible. You're not guessing at truth directly — you're asking "which parameter values would have made this exact data the most probable outcome?" and picking those.

Worked example: estimating a coin's bias from flips

Suppose you're checking whether an LLM's binary classifier (say, "is this email spam?") behaves like a fair coin flip or is biased toward one answer. This is a Bernoulli distribution from Lesson 4 — you just don't know its parameter p (the true probability of "spam"). You observe 20 real classification outcomes from a labeled sample: 14 spam, 6 not-spam.

MLE asks: of all possible values of p between 0 and 1, which one makes "14 spam out of 20" the most probable result? Intuitively — and it turns out, mathematically exactly — the answer is just the observed proportion:

p_hat (MLE estimate) = count of spam / total observations
= 14 / 20
= 0.70

That feels almost too obvious to call a "theorem" — and that's exactly the point. MLE formalizes the obvious move (use the observed proportion) and, crucially, generalizes it to cases where the "obvious" answer isn't obvious at all: fitting a Normal distribution's mean and variance simultaneously, fitting dozens of categorical probabilities at once, or fitting the millions of parameters inside a neural network.

Worked example: fitting a categorical distribution

Recall from Lesson 4 that an LLM's next-token output is a categorical distribution over the vocabulary. Training data works the same way in miniature. Say you're estimating a simple unigram language model — the probability of each word — from a tiny corpus of 50 observed words, where "the" appeared 8 times, "a" appeared 5 times, and so on for the rest of the vocabulary:

MLE estimate for P("the") = count("the") / total words
= 8 / 50
= 0.16
MLE estimate for P("a") = count("a") / total words
= 5 / 50
= 0.10

Do this for every word in the vocabulary and you've fit an entire categorical distribution by MLE — each probability is just "how often did I observe this, out of everything I observed." Fitting a Normal distribution's mean and variance to a batch of continuous scores (like the token-count or latency data from Lesson 2) follows the identical logic: the MLE estimate of the mean is the sample mean, and the MLE estimate of the variance is (almost exactly) the sample variance you already know how to compute.

1. ObserveReal data: coin flips, word counts, labeled outcomes, token frequencies
→
2. Choose a distribution shapeBernoulli, Categorical, Normal — from Lesson 4's family
→
3. Ask MLE's questionWhich parameter values make this exact data most probable?
→
4. Fitted modelA distribution whose parameters now match the observed data as closely as the math allows
Why This Matters "Training a model" is, at its statistical core, this exact process scaled up to millions or billions of parameters: pick parameter values that make the training data you actually have as likely as possible under the model. In Lessons 7 and 8 you'll see this same MLE intuition come back wearing a different name — the loss function models are literally trained with (cross-entropy loss is, mathematically, negative log-likelihood — MLE in disguise). You've already met its two simplest cases here: the coin-flip proportion and the word-count proportion.

Recap

✅ Check Yourself

1. Your eval set has only 15 examples, and a new prompt version scores 6 points higher than the old one. What should you conclude?

The new prompt is definitely better — the score went up
With only 15 examples, a 6-point swing is well within what random sampling noise alone could produce — you can't yet conclude the new prompt is actually better
Eval set size never affects how trustworthy a score is

2. According to the Central Limit Theorem, what happens to the distribution of sample averages as N grows, even if the underlying data isn't Normal?

It becomes more skewed and unpredictable
It stays exactly the same shape as the underlying data
It clusters into a Normal distribution centered on the true value, and that cluster narrows as N increases

3. Standard error at N=100 is 0.04. Roughly what happens to standard error if you quadruple your sample to N=400?

It roughly halves, to about 0.02, because standard error shrinks with the square root of N
It roughly quarters, to about 0.01, because standard error shrinks linearly with N
It stays the same — sample size beyond 100 doesn't matter

4. You observe 30 binary classification outcomes: 21 labeled "positive," 9 labeled "negative." What's the Maximum Likelihood Estimate of the true positive rate p?

0.50, because MLE always assumes a fair split unless told otherwise
0.70 (21/30) — MLE for a Bernoulli/Categorical parameter is just the observed proportion
You cannot estimate p without knowing the population size in advance

5. Why does this lesson connect Maximum Likelihood Estimation to "training a model"?

It doesn't — MLE is a purely theoretical idea with no real use in machine learning
Because training data must always follow a Bernoulli distribution
Because training a model means choosing parameters that make the observed training data as probable as possible under the model — the same MLE logic, scaled up. Cross-entropy loss (Lessons 7-8) is literally negative log-likelihood