Lesson 6 gave you squared-error loss. Lesson 7 gave you cross-entropy loss. This lesson is about what actually happens when a model tries to shrink that number over millions of steps — and the two ways that process quietly goes wrong.
Strip away the formulas and every loss function you've met so far is doing the same job: it looks at what the model predicted, compares it to the truth, and returns a single number that says how bad that prediction was.
Used for regression — predicting a continuous number, like an estimated latency or a predicted price. Loss = average squared difference between predicted and actual value. Bigger miss, disproportionately bigger penalty.
Used for classification and next-token prediction — anywhere the model outputs a probability distribution over categories (including "which token comes next"). Punishes confident wrong answers hardest. This is the loss every LLM is trained on.
Different formulas, same role: a loss function takes the model's current parameters (its weights), runs them against some data, and outputs one number — high means "very wrong," low means "pretty close." Training is nothing more than the repeated process of nudging those weights to make that number smaller. That's the entire chapter, at the highest level. Everything below is about how that nudging happens, and the two ways it can go off the rails.
Here's the intuition, with zero calculus required. Imagine every possible setting of a model's weights as a point on a vast, hilly landscape, and the height of the land at that point is the loss — how wrong the model is with those particular weights. High hills are bad settings; low valleys are good ones. This landscape is called the loss landscape, and it has far more dimensions than three, but the picture still works.
Now imagine you're standing somewhere on that landscape, in thick fog. You can't see the whole terrain — no map, no bird's-eye view. All you can do is feel the ground right under your feet and sense which direction slopes downward. So you take one small step in that downhill direction. Then you feel the ground again, from your new spot, and take another small step downhill. Repeat that thousands or millions of times, and you gradually work your way down into a valley — a region of low loss.
That's gradient descent, the algorithm behind almost all model training:
The model's current weight values.
The loss for those weights, measured on a batch of training data.
Computing which direction each weight should move to reduce loss (the technical name for this direction is the gradient — you don't need the calculus behind it, just the idea "which way is downhill from here").
Nudging every weight slightly in that downhill direction. How big a step is called the learning rate.
A real training run is that loop repeated over and over, each time on a fresh slice of data. The loss number you'd see in a training log is watching the elevation drop as the model feels its way downhill:
Gradient descent only ever looks at one thing: the loss on the data it's currently being nudged by. That data is called the training set, and the loss computed on it is training loss. But training loss alone can't tell you whether the model actually learned something useful — it only tells you how well the model fits the exact examples it's staring at right now.
So every serious training run also holds back a slice of data the model never trains on — the validation set — and periodically checks the loss on that instead. That's validation loss. Because the model never adjusts its weights based on validation data, validation loss is your proxy for "how well does this generalize to examples it hasn't seen," which is the only thing you actually care about in production.
Both losses fall together through the marked point — healthy learning. Past it, training loss keeps dropping (the model fits its training batches better and better) while validation loss turns back upward. That divergence is the overfitting signature, and the marked point is usually the best checkpoint to actually keep and ship — not the one with the lowest training loss.
Overfitting is what the divergence in the chart above is showing you: the model has started memorizing the training data — including its noise and idiosyncrasies — instead of learning the general pattern underneath it. Training loss keeps dropping because memorization always improves training-set performance. Validation loss stalls or climbs because memorized quirks don't transfer to examples the model hasn't seen.
A model fine-tuned on a support-ticket dataset answers the exact training questions perfectly — but ask the same question with different wording, and accuracy falls off a cliff. It memorized specific phrasings, not the underlying task.
A sharper version of the same failure. Fine-tune a capable general-purpose LLM on a small, narrow dataset for too long, and it can overfit so hard to that narrow style or domain that it visibly loses general abilities it had before — it "forgets" broad pretrained knowledge in exchange for excelling at the tiny thing it just memorized.
Underfitting is the opposite failure: the model is too simple, or hasn't trained long enough, to capture the pattern in the data in the first place. Unlike overfitting, there's no divergence to spot — both training loss and validation loss stay high, and they stay close together, because the model is failing uniformly. It isn't memorizing the training set; it can't even fit the training set well.
Training loss: low. Validation loss: high and rising. The gap between them is the tell. Fix: more data, regularization, or stop training earlier.
Training loss: high. Validation loss: also high, close to training loss. No gap — the model just hasn't learned the pattern. Fix: more capacity (a bigger model), more training steps, or better features.
A tiny model asked to handle a genuinely complex task, or an LLM fine-tune cut off after only a handful of steps, will both show this signature: flat, stubbornly high loss on everything, training and validation alike.
Overfitting and underfitting are two ends of the same dial, and statisticians have names for the failure mode at each end: bias and variance.
| High Bias — underfitting | High Variance — overfitting | |
|---|---|---|
| What it looks like | Consistently wrong in the same systematic way | Wildly different results depending on exactly which training examples it happened to see |
| Training loss | High | Low |
| Validation loss | High, close to training loss | High, far above training loss |
| Typical cause | Model too simple, or undertrained, for the pattern | Model too flexible for the amount of data, or trained too long |
Bias is error from a model that's too rigid to capture the real pattern — it makes the same kind of mistake over and over, regardless of which training data you gave it. Variance is error from a model that's too sensitive — retrain it on a slightly different sample of data and you'd get a noticeably different model, chasing noise instead of signal. As model complexity increases, bias tends to fall (a more flexible model can capture more of the real pattern) while variance tends to rise (a more flexible model has more room to latch onto noise). Total error is the sum of both, and it traces the classic U-shape:
Bias falls as the model gets more flexible; variance rises. Total error — the one you actually care about — is their sum, and it's minimized at a middle complexity, not at either extreme. Too simple (left) and bias dominates: underfitting. Too flexible (right) and variance dominates: overfitting. The marked point is the sweet spot you're aiming for.
Regularization is the umbrella term for techniques that deliberately stop a model from fitting its training data too perfectly, in exchange for better generalization — trading a little training accuracy for a lot less overfitting. You don't need the formulas behind any of these to understand what they're for:
Stop training at the checkpoint where validation loss is lowest — exactly the marked point in the chart from section 3 — instead of training until training loss bottoms out.
Discourage the model from relying on any single weight becoming huge, which tends to be a sign of memorizing noise.
Randomly disable a fraction of neurons during each training step, so the model can't over-rely on any one narrow path through the network.
A bigger, more varied training set is simply harder to memorize — the most direct fix for high variance.
1. During training, what does a loss function actually give you?
2. In the "standing on a hill in fog" analogy for gradient descent, what does "taking a small step downhill" correspond to?
3. Training loss keeps dropping smoothly, but validation loss bottoms out around step 4000 and then starts climbing. What does this pattern usually signal?
4. A small, narrow fine-tuning dataset makes a language model excellent at the fine-tuning task but noticeably worse at general tasks it used to handle fine. What's this specific failure called?
5. On the bias-variance tradeoff curve, high bias corresponds to which failure mode, and high variance to which?