Lesson 8 of 10 · Statistics for AI, GenAI & LLMs

The Statistics Behind Training a Model

Lesson 6 gave you squared-error loss. Lesson 7 gave you cross-entropy loss. This lesson is about what actually happens when a model tries to shrink that number over millions of steps — and the two ways that process quietly goes wrong.

1. A loss function is just "how wrong am I, right now"

Strip away the formulas and every loss function you've met so far is doing the same job: it looks at what the model predicted, compares it to the truth, and returns a single number that says how bad that prediction was.

Lesson 6 · Squared-error loss

Used for regression — predicting a continuous number, like an estimated latency or a predicted price. Loss = average squared difference between predicted and actual value. Bigger miss, disproportionately bigger penalty.

Lesson 7 · Cross-entropy loss

Used for classification and next-token prediction — anywhere the model outputs a probability distribution over categories (including "which token comes next"). Punishes confident wrong answers hardest. This is the loss every LLM is trained on.

Different formulas, same role: a loss function takes the model's current parameters (its weights), runs them against some data, and outputs one number — high means "very wrong," low means "pretty close." Training is nothing more than the repeated process of nudging those weights to make that number smaller. That's the entire chapter, at the highest level. Everything below is about how that nudging happens, and the two ways it can go off the rails.

Key Insight Every "the model is training" progress bar you've ever watched — a fine-tuning job, a from-scratch pretraining run, a small classifier's `.fit()` call — is that same loop: compute the loss, nudge the weights a little, recompute the loss, nudge again. Millions of tiny corrections, guided entirely by whichever loss function was chosen for the job.

2. Gradient descent: finding your way downhill in fog

Here's the intuition, with zero calculus required. Imagine every possible setting of a model's weights as a point on a vast, hilly landscape, and the height of the land at that point is the loss — how wrong the model is with those particular weights. High hills are bad settings; low valleys are good ones. This landscape is called the loss landscape, and it has far more dimensions than three, but the picture still works.

Now imagine you're standing somewhere on that landscape, in thick fog. You can't see the whole terrain — no map, no bird's-eye view. All you can do is feel the ground right under your feet and sense which direction slopes downward. So you take one small step in that downhill direction. Then you feel the ground again, from your new spot, and take another small step downhill. Repeat that thousands or millions of times, and you gradually work your way down into a valley — a region of low loss.

That's gradient descent, the algorithm behind almost all model training:

Where you're standing

The model's current weight values.

The height at that spot

The loss for those weights, measured on a batch of training data.

"Feeling the slope"

Computing which direction each weight should move to reduce loss (the technical name for this direction is the gradient — you don't need the calculus behind it, just the idea "which way is downhill from here").

"Taking a step"

Nudging every weight slightly in that downhill direction. How big a step is called the learning rate.

A real training run is that loop repeated over and over, each time on a fresh slice of data. The loss number you'd see in a training log is watching the elevation drop as the model feels its way downhill:

Step 1: loss = 2.41 (a bad, high-elevation start)
Step 50: loss = 1.02
Step 200: loss = 0.31
Step 1000: loss = 0.09
Each step: feel the local slope → nudge every weight a little downhill → repeat.
Pitfall The step size (learning rate) matters more than beginners expect. Too large, and you overshoot the valley entirely — bouncing wildly from one hillside to another, loss refusing to settle. Too small, and you creep downhill so slowly that training that should take hours takes weeks. Neither failure is about the data or the model architecture — it's purely about how big a step gradient descent is allowed to take.

3. Training loss vs. validation loss — two different measuring sticks

Gradient descent only ever looks at one thing: the loss on the data it's currently being nudged by. That data is called the training set, and the loss computed on it is training loss. But training loss alone can't tell you whether the model actually learned something useful — it only tells you how well the model fits the exact examples it's staring at right now.

So every serious training run also holds back a slice of data the model never trains on — the validation set — and periodically checks the loss on that instead. That's validation loss. Because the model never adjusts its weights based on validation data, validation loss is your proxy for "how well does this generalize to examples it hasn't seen," which is the only thing you actually care about in production.

Loss (lower = better) Training Steps → high 0 best checkpoint ↓ overfitting begins
Training loss Validation loss

Both losses fall together through the marked point — healthy learning. Past it, training loss keeps dropping (the model fits its training batches better and better) while validation loss turns back upward. That divergence is the overfitting signature, and the marked point is usually the best checkpoint to actually keep and ship — not the one with the lowest training loss.

Key Insight Training loss going down is not, by itself, good news. It only tells you the model is getting better at the exact examples it's staring at. Validation loss is what tells you whether that improvement is transferring to anything new. Track both, every time — a training log with only one loss curve is only telling you half the story.

4. Overfitting — memorizing instead of learning

Overfitting is what the divergence in the chart above is showing you: the model has started memorizing the training data — including its noise and idiosyncrasies — instead of learning the general pattern underneath it. Training loss keeps dropping because memorization always improves training-set performance. Validation loss stalls or climbs because memorized quirks don't transfer to examples the model hasn't seen.

What it looks like in practice

A model fine-tuned on a support-ticket dataset answers the exact training questions perfectly — but ask the same question with different wording, and accuracy falls off a cliff. It memorized specific phrasings, not the underlying task.

Catastrophic forgetting

A sharper version of the same failure. Fine-tune a capable general-purpose LLM on a small, narrow dataset for too long, and it can overfit so hard to that narrow style or domain that it visibly loses general abilities it had before — it "forgets" broad pretrained knowledge in exchange for excelling at the tiny thing it just memorized.

Pitfall A model that scores near-perfectly on its own training set is not automatically a good model — it may just be an expensive lookup table. Always judge a model by validation (or held-out test) performance, never by training performance alone. This is the exact trap the base-rate fallacy warned about back in Lesson 3: a great-looking number, measured the wrong way, tells you nothing about what you actually need to know.

5. Underfitting — too simple, or undertrained, to learn the pattern at all

Underfitting is the opposite failure: the model is too simple, or hasn't trained long enough, to capture the pattern in the data in the first place. Unlike overfitting, there's no divergence to spot — both training loss and validation loss stay high, and they stay close together, because the model is failing uniformly. It isn't memorizing the training set; it can't even fit the training set well.

Overfitting (recap)

Training loss: low. Validation loss: high and rising. The gap between them is the tell. Fix: more data, regularization, or stop training earlier.

Underfitting

Training loss: high. Validation loss: also high, close to training loss. No gap — the model just hasn't learned the pattern. Fix: more capacity (a bigger model), more training steps, or better features.

A tiny model asked to handle a genuinely complex task, or an LLM fine-tune cut off after only a handful of steps, will both show this signature: flat, stubbornly high loss on everything, training and validation alike.

6. Bias-variance tradeoff

Overfitting and underfitting are two ends of the same dial, and statisticians have names for the failure mode at each end: bias and variance.

High Bias — underfittingHigh Variance — overfitting
What it looks likeConsistently wrong in the same systematic wayWildly different results depending on exactly which training examples it happened to see
Training lossHighLow
Validation lossHigh, close to training lossHigh, far above training loss
Typical causeModel too simple, or undertrained, for the patternModel too flexible for the amount of data, or trained too long

Bias is error from a model that's too rigid to capture the real pattern — it makes the same kind of mistake over and over, regardless of which training data you gave it. Variance is error from a model that's too sensitive — retrain it on a slightly different sample of data and you'd get a noticeably different model, chasing noise instead of signal. As model complexity increases, bias tends to fall (a more flexible model can capture more of the real pattern) while variance tends to rise (a more flexible model has more room to latch onto noise). Total error is the sum of both, and it traces the classic U-shape:

Error Model Complexity → high low sweet spot — best generalization
Bias (underfit error) Variance (overfit error) Total error

Bias falls as the model gets more flexible; variance rises. Total error — the one you actually care about — is their sum, and it's minimized at a middle complexity, not at either extreme. Too simple (left) and bias dominates: underfitting. Too flexible (right) and variance dominates: overfitting. The marked point is the sweet spot you're aiming for.

Key Insight There is no such thing as a model with zero bias and zero variance at the same time — pushing one down tends to push the other up. Training a model well isn't about eliminating error; it's about landing near the bottom of that U, where the two error sources balance out.

7. Regularization: deliberately holding the model back

Regularization is the umbrella term for techniques that deliberately stop a model from fitting its training data too perfectly, in exchange for better generalization — trading a little training accuracy for a lot less overfitting. You don't need the formulas behind any of these to understand what they're for:

Early stopping

Stop training at the checkpoint where validation loss is lowest — exactly the marked point in the chart from section 3 — instead of training until training loss bottoms out.

Weight penalties (L1/L2)

Discourage the model from relying on any single weight becoming huge, which tends to be a sign of memorizing noise.

Dropout

Randomly disable a fraction of neurons during each training step, so the model can't over-rely on any one narrow path through the network.

More / cleaner data

A bigger, more varied training set is simply harder to memorize — the most direct fix for high variance.

Why This Matters Whenever you fine-tune a model, ship an internal classifier, or read someone else's training report, the questions to ask are always the same ones this lesson built: is the loss function the right one for the task? Are training and validation loss tracked separately, and where do they diverge? Is the gap a bias problem or a variance problem — and does the fix look like "give it more capacity" or "hold it back more"? Lesson 9 hands you the tools to score a finished model rigorously; this lesson is about judging whether the process that produced it was sound in the first place.

Recap

✅ Check Yourself

1. During training, what does a loss function actually give you?

The model's final accuracy score
A single number that says how wrong the model's current parameters are, right now
A list of exactly which training examples were misclassified

2. In the "standing on a hill in fog" analogy for gradient descent, what does "taking a small step downhill" correspond to?

Nudging the model's weights slightly in whatever direction reduces the loss
Collecting a fresh batch of training data
Increasing the number of parameters in the model

3. Training loss keeps dropping smoothly, but validation loss bottoms out around step 4000 and then starts climbing. What does this pattern usually signal?

Underfitting — the model hasn't learned enough yet
Overfitting — the model is starting to memorize training data instead of learning the general pattern
A bug in the loss function, since loss should never rise once it starts falling

4. A small, narrow fine-tuning dataset makes a language model excellent at the fine-tuning task but noticeably worse at general tasks it used to handle fine. What's this specific failure called?

Underfitting
Regularization
Catastrophic forgetting

5. On the bias-variance tradeoff curve, high bias corresponds to which failure mode, and high variance to which?

High bias = overfitting; high variance = underfitting
High bias = underfitting (consistently wrong in the same way); high variance = overfitting (wildly sensitive to which training examples it saw)
Bias and variance measure the same thing, just at different points in training