The three quantities that measure "how surprised" a model is — and the actual loss function running underneath every LLM's training loop.
In Lesson 4 you met the categorical distribution: an LLM's next-token prediction is a probability spread across its entire vocabulary, produced by softmax. In Lesson 6 you saw that every model "fits" data by minimizing some loss. This lesson connects the two. The loss that trains classifiers and LLMs is built directly out of three ideas from information theory — entropy, cross-entropy, and KL divergence. Once you can read these three words, you can read almost any ML paper's methods section.
Entropy measures how uncertain a probability distribution is. Formally, for a distribution P over outcomes with probabilities p_i:
The intuition: an outcome with probability p_i carries −log₂(p_i) bits of "surprise" — a near-certain event (p close to 1) is barely surprising (near 0 bits); a rare event (p close to 0) is very surprising (many bits). Entropy is just the probability-weighted average surprise across every possible outcome. A distribution that's spread evenly across many outcomes has high entropy (you're consistently unsure what happens next). A distribution concentrated on one outcome has low entropy (you already basically know the answer).
An LLM's next-token distribution is exactly this kind of distribution — a categorical distribution over every token in the vocabulary (Lesson 4). Its entropy tells you, at that specific point in the sentence, how confident the model is:
Same vertical scale in both charts (bar height = probability). Left: one token dominates → low entropy, high confidence. Right: probability is spread across several plausible words → high entropy, low confidence.
Simplified to a handful of candidate words for readability — a real LLM's vocabulary has 50,000+ tokens, and every one of them gets some probability. The shape (spiky vs. flat) is what matters, not the exact count of bars.
Training needs a second ingredient: comparing the model's predicted distribution against what actually happened. That's cross-entropy — it measures the mismatch between a true distribution P (reality) and a predicted distribution Q (the model's guess):
In next-token prediction, the true distribution P is almost always one-hot: probability 1 on the word that actually came next, probability 0 on every other word in the vocabulary. That's a huge simplification — every term in the sum except the true word's term multiplies by zero and vanishes. Cross-entropy collapses to one number:
This is why cross-entropy loss is sometimes just called negative log-likelihood — it is literally minus the log of the probability the model assigned to the answer that turned out to be true. Low probability on the right answer → large loss. High probability on the right answer → small loss.
The red dashed segment is the mismatch: the model put only 70% on the word that actually came next, so it "owes" −log₂(0.70) ≈ 0.51 bits of loss. A perfect prediction (100% on Paris) would close that gap entirely and drive cross-entropy to 0.
The gap matters a lot more than it looks. Compare two models predicting the same true word (Paris), one reasonably confident and correct, one confidently wrong:
| Model's prediction | Probability on "Paris" | Cross-entropy loss |
|---|---|---|
| Confident & correct | 70% | −log₂(0.70) ≈ 0.51 bits |
| Confident & wrong (put 65% on "London" instead) | 5% | −log₂(0.05) ≈ 4.32 bits |
Being confidently wrong costs roughly 8× more loss than being reasonably confident and right. That asymmetry is the entire training signal: gradient descent (Lesson 8) nudges the model's parameters to push probability mass toward whatever word actually showed up next, over and over, across billions of tokens.
CrossEntropyLoss, TensorFlow's categorical_crossentropy) compute the natural log (base e), giving loss in nats, not bits. The shape of the math and the training behavior are identical either way — only the numeric scale changes (1 nat ≈ 1.44 bits). Don't be surprised if a loss curve you see reported doesn't match a bits-based hand calculation exactly.
Cross-entropy actually bundles two things together: how uncertain the true distribution inherently is, plus how much extra "surprise" comes specifically from the model's predictions being wrong. KL (Kullback–Leibler) divergence isolates that second piece — the pure mismatch between two distributions, with the true distribution's own baseline uncertainty subtracted out:
Because next-token labels are one-hot, H(P) = 0 — a certain event has zero entropy, nothing left to be uncertain about. So for LLM training specifically, cross-entropy loss and KL divergence are the exact same number: minimizing cross-entropy loss is minimizing KL divergence between the model's predictions and reality. The KL-divergence framing becomes genuinely useful in situations where neither distribution is one-hot — comparing two full, "spread-out" probability distributions against each other.
H(P) measures the built-in uncertainty of one distribution — high when probability is spread thin across many outcomes, low when it's concentrated on one. A next-token distribution's entropy is a direct readout of how confident the model is at that point in the sentence.H(P,Q) measures the mismatch between a true distribution and a predicted one — low when the model puts high probability on what actually happened, high when it's confidently wrong. Because true labels are one-hot, it collapses to −log(probability assigned to the correct answer) — this is the loss function training both ordinary classifiers and LLM next-token prediction.D_KL(P‖Q) is the pure mismatch piece of cross-entropy, with the true distribution's own baseline entropy subtracted out: cross-entropy = entropy + KL divergence. For one-hot labels the two are numerically identical; KL becomes its own tool when comparing two genuinely spread-out distributions, like a fine-tuned model against its reference model in RLHF.1. Which next-token situation has HIGHER entropy?
2. What does cross-entropy actually measure?
3. Why is cross-entropy the natural loss function for LLM next-token prediction?
4. What is the relationship between entropy, cross-entropy, and KL divergence?
5. In RLHF fine-tuning, why is a KL-divergence penalty applied between the fine-tuned model and a frozen reference model?