Everything so far — distributions, entropy, cross-entropy, overfitting, evaluation rigor — generalizes to any predictive model. This lesson is where it all converges on the thing that made you start this course: the statistics that specifically make an LLM's next-token machinery tick.
An LLM, at every single step of generating text, does one thing: it produces a probability distribution over its entire vocabulary — tens of thousands of candidate next tokens, each assigned a probability, all summing to 1. That's the categorical distribution from Lesson 4. Training that model means minimizing cross-entropy loss — Lesson 7's "how surprised was the model by the true next token, on average." Everything in this lesson is what happens once you have that trained distribution in hand: how good was the training (perplexity), how do you turn a distribution into an actual chosen word (sampling strategies), can you trust what the model tells you about its own confidence (calibration), and why does it sometimes say things that are fluent but false (hallucination).
Lesson 7 gave you cross-entropy loss: at each position in a sequence, take the negative log of the probability the model assigned to the true next token, then average that across every position. That number is what an LLM is actually trained to minimize. It's mathematically correct and it's the real training signal — but as a number, "1.15 nats" tells a human almost nothing on its own.
Perplexity solves that. It's defined as:
That's it — perplexity is cross-entropy loss put through exp() and placed on a scale you can actually reason about: roughly "how many words was the model effectively choosing between, on average, at each step." A perplexity of 1 means the model put 100% probability on the correct token every time — perfect prediction. A perplexity equal to the size of the vocabulary means the model was, on average, no better than picking uniformly at random. Everything in between tells you the model's effective "branching factor" at decision time.
Say a model is scored on a 4-token stretch of text, and here's the probability it assigned to the actual next token at each position:
| Position | P(true token) | −ln(p) |
|---|---|---|
| 1 | 0.50 | 0.693 |
| 2 | 0.25 | 1.386 |
| 3 | 0.80 | 0.223 |
| 4 | 0.10 | 2.303 |
The model was, on average, choosing among roughly 3.16 equally-plausible options at each step — even though its confidence swung wildly, from 80% on one token down to 10% on another. Perplexity smooths that per-token variance into a single, comparable number. (Note: training loss is conventionally reported in nats — natural log, matching PyTorch's and most frameworks' cross-entropy loss. If you're thinking in bits, log base 2, from Lesson 7's entropy examples, the idea is identical — just swap exp() for 2^().)
Lesson 4 established that an LLM's next-token step is a categorical distribution produced by softmax — a probability for every token in the vocabulary, summing to 1. Once you have that distribution, you still need a rule for turning it into one chosen token. That rule is a sampling strategy, and the three most common ones all work by reshaping the distribution before drawing from it.
Picture the distribution after the prompt "The cat ___" — six plausible next tokens and their base probabilities:
| Token | Base probability |
|---|---|
| sat | 42% |
| ran | 23% |
| jumped | 15% |
| slept | 10% |
| meowed | 6% |
| flew | 4% |
That base distribution has entropy ≈ 2.19 bits (Lesson 7: "average surprise" — higher means the model is more spread out/uncertain across options). Here's what each strategy does to it.
Temperature rescales the logits (the pre-softmax scores) by dividing by a value T before applying softmax again. Low temperature (T < 1) sharpens the distribution — the already-likely token gets pushed even higher, entropy drops, and sampling becomes more confident and deterministic. High temperature (T > 1) flattens the distribution — probabilities move closer together, entropy rises, and sampling becomes more random and "creative." As T → 0, sampling converges to always picking the single highest-probability token (greedy decoding, zero entropy).
Keep only the k highest-probability tokens, discard everything else, and renormalize the remaining probabilities so they sum back to 1 before sampling. Simple and effective, but k is fixed regardless of context — the same k=3 might needlessly restrict a very open-ended continuation, or fail to trim a genuinely narrow one.
Instead of a fixed count, keep the smallest set of tokens whose cumulative probability exceeds a threshold p, then renormalize and sample from just that set. Because the set size adapts to how peaked or flat the distribution already is, nucleus sampling tends to keep few tokens when the model is confident and more tokens when it's genuinely uncertain — the fix for top-k's fixed-count blind spot. This method comes from Holtzman et al.'s 2019 paper "The Curious Case of Neural Text Degeneration," cited below.
Same base distribution ("The cat ___"), reshaped three ways. Left: low temperature sharpens it and lowers entropy. Middle: high temperature flattens it and raises entropy. Right: top-k=3 and top-p=0.8 both happen to keep the same three tokens here (cumulative probability hits exactly 0.80 at "jumped") and discard the rest (×) before renormalizing — a hard cutoff rather than a continuous reshape, but it also lands at a similarly low entropy.
| Condition | Entropy (bits) | Effect |
|---|---|---|
| Base (T = 1, no truncation) | 2.19 | Original model output |
| Low temperature (T = 0.5) | 1.46 | Sharper, more deterministic |
| High temperature (T = 2.0) | 2.48 | Flatter, more random |
| Top-k=3 / top-p=0.8 (renormalized) | 1.46 | Tail deleted outright, not reshaped |
A model is well-calibrated if, across many predictions where it says "I'm 80% confident," it's actually correct about 80% of the time. This is Lesson 3's Bayesian instinct turned into a checkable empirical claim, and it's measured with exactly Lesson 9's evaluation discipline: bucket predictions by stated confidence, compare each bucket's average confidence to its actual accuracy on held-out data.
| Stated confidence bucket | n predictions | Actual accuracy | Read |
|---|---|---|---|
| 90–100% | 200 | 72% | Overconfident by ~23 points |
| 70–80% | 180 | 68% | Roughly calibrated |
| 50–60% | 150 | 55% | Well calibrated |
The 90–100% bucket is the dangerous one: the model is loudly confident and wrong nearly 3 times in 10. If a downstream system auto-approves anything the model reports above 90% confidence, that gap between stated and actual accuracy becomes a silent, systematic error rate.
"Hallucination" sounds like a mysterious glitch. Statistically, it's the predictable consequence of two things you've now already learned, working together.
Go back to the Section 2 chart. "Flew" had only 4% probability under the base distribution — but at high temperature it rose to 8.8%. Sample enough times at that temperature, and eventually you draw "flew": "The cat flew across the room." Fluent, grammatical, and false. That's not a malfunction — it's a low-probability, high-entropy token that the sampling procedure was always going to pick sometimes, by design.
Lesson 8's overfitting/generalization problem resurfaces here: in a region of the training distribution the model saw rarely (an obscure fact, a niche API, a rare drug interaction), the model may still confidently assign high probability to one specific, wrong continuation — it's extrapolating into territory it wasn't well-trained on, and nothing in training taught it to say "I don't know" there instead.
The first failure mode is about the sampling step — even a perfectly calibrated model will occasionally sample its way into a false statement, because "occasionally sample the unlikely thing" is exactly what temperature > 0 sampling means. The second is about the model's underlying distribution itself being wrong in sparse regions — badly calibrated, confidently peaked on the wrong answer, no useful signal from entropy or probability to warn you. Reducing hallucination in practice means attacking both: lower temperature or tighter top-p for factual tasks (attacks failure mode 1), and better calibration training or retrieval-grounding so the model has real information to condition on instead of extrapolating (attacks failure mode 2).
You'll see "RLHF" (Reinforcement Learning from Human Feedback) mentioned constantly in model release notes, so it's worth recognizing what it is in one paragraph. After the pretraining that minimizes cross-entropy loss (Section 1), models are typically fine-tuned using comparisons — humans or other AI models judging "which of these two responses is better" — turned into a reward signal via reinforcement learning. Left unconstrained, that fine-tuning could drift the model's output distribution very far from its original, broadly-trained distribution in pursuit of the reward. To prevent that, the fine-tuning objective usually includes a KL-divergence penalty — Lesson 7's "distance between two distributions" — that keeps the fine-tuned model's next-token distribution close to the original pretrained one, penalizing it for straying too far even while it optimizes for the preference reward.
You've reached the end of Statistics for AI, GenAI & LLMs — ten lessons from "why does statistics matter at all" to the exact mechanics running inside the model you talk to every day.
Here's how the ten lessons built toward this one:
You started this course as a total beginner in statistics. You now have the tools to read a model card, question a reported accuracy number, reason about why turning up temperature changes what comes out, and explain — with actual math, not a shrug — why a language model sometimes says something confidently false. That's fluency, earned one distribution at a time.
1. Perplexity is defined as:
2. Lowering the temperature before sampling a next token does what to the distribution?
3. What's the key difference between top-k and top-p (nucleus) sampling?
4. A model states "85% confident" on many separate predictions. If the model is well-calibrated, what should you observe?
5. Which best describes hallucination in statistical terms?