Lesson 10 of 10 · Capstone · Statistics for AI, GenAI & LLMs

The Statistics Unique to LLMs

Everything so far — distributions, entropy, cross-entropy, overfitting, evaluation rigor — generalizes to any predictive model. This lesson is where it all converges on the thing that made you start this course: the statistics that specifically make an LLM's next-token machinery tick.

An LLM, at every single step of generating text, does one thing: it produces a probability distribution over its entire vocabulary — tens of thousands of candidate next tokens, each assigned a probability, all summing to 1. That's the categorical distribution from Lesson 4. Training that model means minimizing cross-entropy loss — Lesson 7's "how surprised was the model by the true next token, on average." Everything in this lesson is what happens once you have that trained distribution in hand: how good was the training (perplexity), how do you turn a distribution into an actual chosen word (sampling strategies), can you trust what the model tells you about its own confidence (calibration), and why does it sometimes say things that are fluent but false (hallucination).

1. Perplexity — cross-entropy loss, made interpretable

Lesson 7 gave you cross-entropy loss: at each position in a sequence, take the negative log of the probability the model assigned to the true next token, then average that across every position. That number is what an LLM is actually trained to minimize. It's mathematically correct and it's the real training signal — but as a number, "1.15 nats" tells a human almost nothing on its own.

Perplexity solves that. It's defined as:

Perplexity = exp( cross-entropy loss )

That's it — perplexity is cross-entropy loss put through exp() and placed on a scale you can actually reason about: roughly "how many words was the model effectively choosing between, on average, at each step." A perplexity of 1 means the model put 100% probability on the correct token every time — perfect prediction. A perplexity equal to the size of the vocabulary means the model was, on average, no better than picking uniformly at random. Everything in between tells you the model's effective "branching factor" at decision time.

Worked example

Say a model is scored on a 4-token stretch of text, and here's the probability it assigned to the actual next token at each position:

PositionP(true token)−ln(p)
10.500.693
20.251.386
30.800.223
40.102.303
Average cross-entropy loss = (0.693 + 1.386 + 0.223 + 2.303) / 4
= 4.605 / 4 = 1.151 nats
Perplexity = exp(1.151)
≈ 3.16

The model was, on average, choosing among roughly 3.16 equally-plausible options at each step — even though its confidence swung wildly, from 80% on one token down to 10% on another. Perplexity smooths that per-token variance into a single, comparable number. (Note: training loss is conventionally reported in nats — natural log, matching PyTorch's and most frameworks' cross-entropy loss. If you're thinking in bits, log base 2, from Lesson 7's entropy examples, the idea is identical — just swap exp() for 2^().)

Key Insight When you see a model leaderboard reporting "Model A: perplexity 12.3 vs. Model B: perplexity 18.7" on the same held-out text, that's not an abstract score — it means Model A was, on average, choosing among about 12 plausible next tokens at each position, while Model B was choosing among about 19. Lower perplexity is a real, interpretable claim about how sharply the model narrows down the vocabulary.

2. Sampling strategies — reshaping the categorical distribution before you pick

Lesson 4 established that an LLM's next-token step is a categorical distribution produced by softmax — a probability for every token in the vocabulary, summing to 1. Once you have that distribution, you still need a rule for turning it into one chosen token. That rule is a sampling strategy, and the three most common ones all work by reshaping the distribution before drawing from it.

Picture the distribution after the prompt "The cat ___" — six plausible next tokens and their base probabilities:

TokenBase probability
sat42%
ran23%
jumped15%
slept10%
meowed6%
flew4%

That base distribution has entropy ≈ 2.19 bits (Lesson 7: "average surprise" — higher means the model is more spread out/uncertain across options). Here's what each strategy does to it.

Temperature

Temperature rescales the logits (the pre-softmax scores) by dividing by a value T before applying softmax again. Low temperature (T < 1) sharpens the distribution — the already-likely token gets pushed even higher, entropy drops, and sampling becomes more confident and deterministic. High temperature (T > 1) flattens the distribution — probabilities move closer together, entropy rises, and sampling becomes more random and "creative." As T → 0, sampling converges to always picking the single highest-probability token (greedy decoding, zero entropy).

Top-k sampling

Keep only the k highest-probability tokens, discard everything else, and renormalize the remaining probabilities so they sum back to 1 before sampling. Simple and effective, but k is fixed regardless of context — the same k=3 might needlessly restrict a very open-ended continuation, or fail to trim a genuinely narrow one.

Top-p (nucleus) sampling

Instead of a fixed count, keep the smallest set of tokens whose cumulative probability exceeds a threshold p, then renormalize and sample from just that set. Because the set size adapts to how peaked or flat the distribution already is, nucleus sampling tends to keep few tokens when the model is confident and more tokens when it's genuinely uncertain — the fix for top-k's fixed-count blind spot. This method comes from Holtzman et al.'s 2019 paper "The Curious Case of Neural Text Degeneration," cited below.

Low Temperature T = 0.5 — sharper 66% 20% 8% sat ran jump slpt meow fly H ≈ 1.46 bits High Temperature T = 2.0 — flatter 28% 21% 17% sat ran jump slpt meow fly H ≈ 2.48 bits Top-k=3 / Top-p=0.8 tail truncated 42% 23% 15% sat ran jump slpt× meow× fly× H ≈ 1.46 bits (kept set)

Same base distribution ("The cat ___"), reshaped three ways. Left: low temperature sharpens it and lowers entropy. Middle: high temperature flattens it and raises entropy. Right: top-k=3 and top-p=0.8 both happen to keep the same three tokens here (cumulative probability hits exactly 0.80 at "jumped") and discard the rest (×) before renormalizing — a hard cutoff rather than a continuous reshape, but it also lands at a similarly low entropy.

ConditionEntropy (bits)Effect
Base (T = 1, no truncation)2.19Original model output
Low temperature (T = 0.5)1.46Sharper, more deterministic
High temperature (T = 2.0)2.48Flatter, more random
Top-k=3 / top-p=0.8 (renormalized)1.46Tail deleted outright, not reshaped
Pitfall Temperature and top-k/top-p can both lower entropy, but they don't do it the same way. Temperature continuously reshapes every probability in the distribution — nothing is ever truly zeroed out. Top-k and top-p draw a hard line and delete the tail entirely, no matter how low you set the temperature elsewhere. They're often combined in production (e.g. temperature 0.7 with top-p 0.9) rather than used alone — check the docs for the API or library you're calling, since the order these are applied in (temperature first vs. truncation first) can change the result.

3. Calibration — does stated confidence match reality?

A model is well-calibrated if, across many predictions where it says "I'm 80% confident," it's actually correct about 80% of the time. This is Lesson 3's Bayesian instinct turned into a checkable empirical claim, and it's measured with exactly Lesson 9's evaluation discipline: bucket predictions by stated confidence, compare each bucket's average confidence to its actual accuracy on held-out data.

Stated confidence bucketn predictionsActual accuracyRead
90–100%20072%Overconfident by ~23 points
70–80%18068%Roughly calibrated
50–60%15055%Well calibrated

The 90–100% bucket is the dangerous one: the model is loudly confident and wrong nearly 3 times in 10. If a downstream system auto-approves anything the model reports above 90% confidence, that gap between stated and actual accuracy becomes a silent, systematic error rate.

Why This Matters Calibration is what tells you whether a model's confidence score is safe to use as a decision gate — auto-approve above X%, route to human review below it. An accurate but poorly-calibrated model (Lesson 9's precision/recall can look fine in aggregate) can still hand you confidence numbers you cannot trust individually. This is a separate failure mode from "the model is often wrong" — calibration is about whether "wrong" and "confident" are correlated the way they should be.

4. Hallucination, framed statistically

"Hallucination" sounds like a mysterious glitch. Statistically, it's the predictable consequence of two things you've now already learned, working together.

1. Sampling from a real distribution

Go back to the Section 2 chart. "Flew" had only 4% probability under the base distribution — but at high temperature it rose to 8.8%. Sample enough times at that temperature, and eventually you draw "flew": "The cat flew across the room." Fluent, grammatical, and false. That's not a malfunction — it's a low-probability, high-entropy token that the sampling procedure was always going to pick sometimes, by design.

2. Poor calibration in low-data regions

Lesson 8's overfitting/generalization problem resurfaces here: in a region of the training distribution the model saw rarely (an obscure fact, a niche API, a rare drug interaction), the model may still confidently assign high probability to one specific, wrong continuation — it's extrapolating into territory it wasn't well-trained on, and nothing in training taught it to say "I don't know" there instead.

The first failure mode is about the sampling step — even a perfectly calibrated model will occasionally sample its way into a false statement, because "occasionally sample the unlikely thing" is exactly what temperature > 0 sampling means. The second is about the model's underlying distribution itself being wrong in sparse regions — badly calibrated, confidently peaked on the wrong answer, no useful signal from entropy or probability to warn you. Reducing hallucination in practice means attacking both: lower temperature or tighter top-p for factual tasks (attacks failure mode 1), and better calibration training or retrieval-grounding so the model has real information to condition on instead of extrapolating (attacks failure mode 2).

Pitfall "The model hallucinated" is often said like it's a categorically different kind of error from "the model was wrong." Statistically, it isn't — it's what happens when you sample from a probability distribution over token sequences and either (a) an unlikely-but-nonzero continuation gets drawn, or (b) the distribution itself was miscalibrated going in. Both are describable, measurable, and — to a degree — controllable through exactly the levers this lesson covers: temperature, truncation, and calibration.

5. A glimpse of RLHF

You'll see "RLHF" (Reinforcement Learning from Human Feedback) mentioned constantly in model release notes, so it's worth recognizing what it is in one paragraph. After the pretraining that minimizes cross-entropy loss (Section 1), models are typically fine-tuned using comparisons — humans or other AI models judging "which of these two responses is better" — turned into a reward signal via reinforcement learning. Left unconstrained, that fine-tuning could drift the model's output distribution very far from its original, broadly-trained distribution in pursuit of the reward. To prevent that, the fine-tuning objective usually includes a KL-divergence penalty — Lesson 7's "distance between two distributions" — that keeps the fine-tuned model's next-token distribution close to the original pretrained one, penalizing it for straying too far even while it optimizes for the preference reward.

RLHF objective ≈ reward(preference) − β · KL(fine-tuned || original)
That's the whole shape of it: chase the reward, but don't let KL divergence from the original distribution grow unchecked.

Recap

🎓 Course Complete

You've reached the end of Statistics for AI, GenAI & LLMs — ten lessons from "why does statistics matter at all" to the exact mechanics running inside the model you talk to every day.

What you now know

Here's how the ten lessons built toward this one:

  1. Lesson 1
    Why Statistics Powers AI — set the premise that every AI system, from a classifier to an LLM, is fundamentally a statistical machine reasoning under uncertainty. Everything since has been filling in that claim.
  2. Lesson 2
    Describing Data — mean, variance, std dev, and distribution shape gave you the vocabulary to describe any pile of numbers, including the probability numbers an LLM outputs.
  3. Lesson 3
    Probability Foundations & Bayes' Theorem — conditional probability and the base-rate fallacy became the backbone of calibration in Section 3 above: confidence is a probability claim, checkable exactly the way you checked a detector's posterior.
  4. Lesson 4
    The Distributions That Run Machine Learning — named the categorical distribution that Section 2 above spent the whole section reshaping with temperature, top-k, and top-p.
  5. Lesson 5
    From Sample to Population — sampling and estimation gave you the lens for treating "the model's true accuracy" and "the model's true calibration" as things you estimate from a finite evaluation set, not facts you simply know.
  6. Lesson 6
    Correlation, Regression & Why Models "Fit" Data — the mechanics of fitting a model to data underlie every distribution this lesson manipulated; a model's next-token probabilities are the direct output of that fitting process.
  7. Lesson 7
    Information Theory for AI — entropy, cross-entropy, and KL divergence are the three ideas this entire lesson is built directly on top of: perplexity is exp(cross-entropy), sampling strategies are entropy management, and RLHF leans on KL divergence.
  8. Lesson 8
    The Statistics Behind Training a Model — overfitting and generalization explained exactly why a model hallucinates with confidence in sparse regions of its training distribution: it's extrapolating into territory it never learned well.
  9. Lesson 9
    Evaluating AI Models Rigorously — precision, recall, and the discipline of measuring on held-out data is exactly the method Section 3 above used to check calibration bucket by bucket.
  10. Lesson 10 (this one)
    The Statistics Unique to LLMs — perplexity, sampling strategies, calibration, and hallucination, all standing on the shoulders of the nine lessons before it.

You started this course as a total beginner in statistics. You now have the tools to read a model card, question a reported accuracy number, reason about why turning up temperature changes what comes out, and explain — with actual math, not a shrug — why a language model sometimes says something confidently false. That's fluency, earned one distribution at a time.

✅ Check Yourself

1. Perplexity is defined as:

The variance of the model's predicted probabilities
The KL divergence between the model and a uniform distribution
exp(cross-entropy loss) — cross-entropy put on an interpretable "effective number of choices" scale

2. Lowering the temperature before sampling a next token does what to the distribution?

Sharpens it — the already-likely token gets pushed higher and entropy decreases
Flattens it — all tokens become closer to equally likely
Has no effect on entropy, only on which token is technically "correct"

3. What's the key difference between top-k and top-p (nucleus) sampling?

Top-k rescales probabilities continuously; top-p deletes tokens outright
Top-k keeps a fixed count of tokens regardless of context; top-p keeps however many tokens are needed to cross a cumulative-probability threshold, so the count adapts to how peaked or flat the distribution is
Top-k only works with temperature 1.0; top-p works at any temperature

4. A model states "85% confident" on many separate predictions. If the model is well-calibrated, what should you observe?

The model should be correct 100% of the time on those predictions
The model should be correct about 85% of the time across those predictions
Calibration is unrelated to how often the model is actually correct

5. Which best describes hallucination in statistical terms?

A random hardware or software bug unrelated to how the model was trained
Proof that the model's training data was too small to matter
Either sampling a legitimately low-probability, high-entropy continuation that happens to be false, or the model being poorly calibrated (confidently wrong) in a sparse region of its training distribution