Lesson 7 of 10 · Statistics for AI, GenAI & LLMs

Information Theory for AI: Entropy, Cross-Entropy & KL Divergence

The three quantities that measure "how surprised" a model is — and the actual loss function running underneath every LLM's training loop.

In Lesson 4 you met the categorical distribution: an LLM's next-token prediction is a probability spread across its entire vocabulary, produced by softmax. In Lesson 6 you saw that every model "fits" data by minimizing some loss. This lesson connects the two. The loss that trains classifiers and LLMs is built directly out of three ideas from information theory — entropy, cross-entropy, and KL divergence. Once you can read these three words, you can read almost any ML paper's methods section.

1. Entropy — average surprise

Entropy measures how uncertain a probability distribution is. Formally, for a distribution P over outcomes with probabilities p_i:

H(P) = − ∑i pi · log₂(pi) ← measured in bits

The intuition: an outcome with probability p_i carries −log₂(p_i) bits of "surprise" — a near-certain event (p close to 1) is barely surprising (near 0 bits); a rare event (p close to 0) is very surprising (many bits). Entropy is just the probability-weighted average surprise across every possible outcome. A distribution that's spread evenly across many outcomes has high entropy (you're consistently unsure what happens next). A distribution concentrated on one outcome has low entropy (you already basically know the answer).

An LLM's next-token distribution is exactly this kind of distribution — a categorical distribution over every token in the vocabulary (Lesson 4). Its entropy tells you, at that specific point in the sentence, how confident the model is:

"The capital of France is ___"H ≈ 0.55 bits — low entropy
92% Paris France the a other
"My favorite color is ___"H ≈ 2.56 bits — high entropy
22% blue 19% red 17% green 14% purple 12% black 16% other

Same vertical scale in both charts (bar height = probability). Left: one token dominates → low entropy, high confidence. Right: probability is spread across several plausible words → high entropy, low confidence.

Simplified to a handful of candidate words for readability — a real LLM's vocabulary has 50,000+ tokens, and every one of them gets some probability. The shape (spiky vs. flat) is what matters, not the exact count of bars.

Key insight Entropy is a property of a single distribution — it doesn't need a "right answer" to compare against. A next-token distribution can have low entropy and still be wrong (the model is confidently wrong), or high entropy and still contain the right answer buried in the spread. Entropy tells you how certain the model is, not how correct it is.

2. Cross-entropy — the mismatch between what happened and what the model predicted

Training needs a second ingredient: comparing the model's predicted distribution against what actually happened. That's cross-entropy — it measures the mismatch between a true distribution P (reality) and a predicted distribution Q (the model's guess):

H(P, Q) = − ∑i pi · log₂(qi)

In next-token prediction, the true distribution P is almost always one-hot: probability 1 on the word that actually came next, probability 0 on every other word in the vocabulary. That's a huge simplification — every term in the sum except the true word's term multiplies by zero and vanishes. Cross-entropy collapses to one number:

H(P, Q) = − log₂(qcorrect) ← "surprise" of the true next word, under the model's own distribution

This is why cross-entropy loss is sometimes just called negative log-likelihood — it is literally minus the log of the probability the model assigned to the answer that turned out to be true. Low probability on the right answer → large loss. High probability on the right answer → small loss.

True distribution (reality)one-hot: Paris = 100%
100% Paris France the a other
Model's predicted distributioncross-entropy ≈ 0.51 bits
target: 100% 70% Paris France the a other

The red dashed segment is the mismatch: the model put only 70% on the word that actually came next, so it "owes" −log₂(0.70) ≈ 0.51 bits of loss. A perfect prediction (100% on Paris) would close that gap entirely and drive cross-entropy to 0.

The gap matters a lot more than it looks. Compare two models predicting the same true word (Paris), one reasonably confident and correct, one confidently wrong:

Model's predictionProbability on "Paris"Cross-entropy loss
Confident & correct70%−log₂(0.70) ≈ 0.51 bits
Confident & wrong (put 65% on "London" instead)5%−log₂(0.05) ≈ 4.32 bits

Being confidently wrong costs roughly 8× more loss than being reasonably confident and right. That asymmetry is the entire training signal: gradient descent (Lesson 8) nudges the model's parameters to push probability mass toward whatever word actually showed up next, over and over, across billions of tokens.

Why this matters Next-token prediction is a classification problem — "which of these 50,000+ vocabulary tokens comes next?" — over the categorical/softmax distribution from Lesson 4. Cross-entropy is the standard loss for comparing a predicted distribution to a true label in any classifier, image or text. LLM pretraining is, mechanically, the exact same loss function used to train a spam filter or an image classifier — just applied token by token, at enormous scale. That's the "minimizing a loss = fitting a model" idea from Lesson 6, made concrete.
Pitfall This lesson uses log base 2 (bits) because it's the most intuitive unit for "surprise." In practice, deep learning frameworks (PyTorch's CrossEntropyLoss, TensorFlow's categorical_crossentropy) compute the natural log (base e), giving loss in nats, not bits. The shape of the math and the training behavior are identical either way — only the numeric scale changes (1 nat ≈ 1.44 bits). Don't be surprised if a loss curve you see reported doesn't match a bits-based hand calculation exactly.

3. KL divergence — cross-entropy's more precise cousin

Cross-entropy actually bundles two things together: how uncertain the true distribution inherently is, plus how much extra "surprise" comes specifically from the model's predictions being wrong. KL (Kullback–Leibler) divergence isolates that second piece — the pure mismatch between two distributions, with the true distribution's own baseline uncertainty subtracted out:

H(P, Q) = H(P) + DKL(P ‖ Q)
cross-entropy = entropy of the truth + KL divergence (truth → prediction)

Because next-token labels are one-hot, H(P) = 0 — a certain event has zero entropy, nothing left to be uncertain about. So for LLM training specifically, cross-entropy loss and KL divergence are the exact same number: minimizing cross-entropy loss is minimizing KL divergence between the model's predictions and reality. The KL-divergence framing becomes genuinely useful in situations where neither distribution is one-hot — comparing two full, "spread-out" probability distributions against each other.

Key insight One place this shows up directly in modern AI: RLHF (Reinforcement Learning from Human Feedback), used to fine-tune models like ChatGPT toward helpful, human-preferred responses (previewed further in Lesson 10). While the model is being pushed toward a reward signal, a KL-divergence penalty term keeps its output distribution from straying too far from a frozen reference model's distribution. Without it, the model can "reward hack" — drifting into strange, repetitive, or degenerate text that scores well on the reward signal but no longer reads like coherent language. KL divergence is the leash.

Recap

✅ Check Yourself

1. Which next-token situation has HIGHER entropy?

"The capital of France is ___" (one word dominates the distribution)
"My favorite color is ___" (probability spread across many plausible words)

2. What does cross-entropy actually measure?

How spread out a single probability distribution is, with no reference to reality
The total number of parameters in a model
The mismatch between a true distribution and a predicted distribution — low when the model puts high probability on what actually happened, high when it's confidently wrong

3. Why is cross-entropy the natural loss function for LLM next-token prediction?

Predicting the next token is a classification problem over the vocabulary (a categorical/softmax distribution), and cross-entropy is the standard way to score a predicted distribution against a true one-hot label
Cross-entropy is the only loss function that exists in machine learning
Cross-entropy only works for image models, not text

4. What is the relationship between entropy, cross-entropy, and KL divergence?

They're unrelated quantities that happen to share the word "entropy"
Cross-entropy = entropy of the true distribution + KL divergence between the true and predicted distributions
KL divergence = cross-entropy × entropy

5. In RLHF fine-tuning, why is a KL-divergence penalty applied between the fine-tuned model and a frozen reference model?

To make training run faster on GPUs
To force the fine-tuned model to output the exact same text as the reference model every time
To stop the model from drifting too far from sensible, human-like language while it's being optimized toward a reward signal — preventing "reward hacking" into degenerate outputs