Two variables that move together, a line that summarizes the pattern, and the realization that fitting that line is a miniature version of everything Lesson 8 will call "training."
Lesson 5 gave you the core idea behind training: a model has parameters, and training means choosing the parameters that make the observed data most likely (or, put simply, that fit the data best). That idea was abstract there. This lesson makes it concrete and visual. You're going to fit an actual line to actual points, watch the fit improve as you choose better numbers, and see — directly, not by analogy — that this is what "training a model" means at the smallest possible scale.
Before you can fit a line, you need a way to ask whether two variables are related at all. That's what the Pearson correlation coefficient, almost always written r, measures: the strength and direction of a linear relationship between two numeric variables.
r is always between -1 and +1, no matter what units your data is in. Here are two AI examples that sit at opposite ends of that scale:
Across a batch of logged conversations, word count of the prompt vs. a 1–10 human quality rating of the response comes out around r ≈ 0.15. Barely any linear relationship — long prompts aren't reliably better or worse than short ones.
Across a family of model checkpoints at different parameter counts, size vs. benchmark accuracy comes out around r ≈ 0.99. Strong positive relationship — bigger checkpoints in this family score higher, almost every time.
Here's the sticky one. Say your moderation team notices something in the data: conversations that use more emoji get flagged as toxic more often. The correlation is real and measurable. It is tempting — and wrong — to conclude that emoji use causes toxicity, or to build a moderation rule that penalizes emoji-heavy messages.
What's actually happening is almost certainly this: a third factor, casual, informal conversational register, drives both. People writing casually use more emoji and are more likely to use blunt, aggressive, or crude language that trips a toxicity classifier. Neither emoji nor toxicity causes the other — they're both downstream of the same underlying cause.
The correlation people notice sits along the bottom — but the real arrows point down from the shared cause above it. This is called a confounding variable.
Correlation tells you a relationship exists and roughly how strong it is. Regression goes further: it finds the actual line that best summarizes that relationship, so you can make predictions from it. A line is defined by two numbers you already know from algebra:
"Fitting" a regression line means choosing m and b so the line sits as close as possible to every point in your dataset. Suppose you evaluate a family of model checkpoints at 8 different sizes on a benchmark, and log these illustrative results:
| Model size (B params) | Benchmark score (%) |
|---|---|
| 1 | 52 |
| 2 | 58 |
| 3 | 61 |
| 4 | 68 |
| 5 | 70 |
| 6 | 75 |
| 7 | 79 |
| 8 | 83 |
The trend is obviously upward, but no line passes through all 8 points exactly. Every candidate line will miss some points by some amount. The question is: which line minimizes the total miss?
For any candidate line, the gap between an actual point and the line's prediction at that x-value is called a residual. This is exactly the "deviation from a central value" idea from Lesson 2, just applied to a line instead of a single mean. And the fix for negative-and-positive deviations canceling out is the same fix as before: square them, then find the m and b that make the sum of squared residuals as small as possible. This method is called ordinary least squares.
You don't need to memorize that formula — the point is that it's doing exactly what "minimize squared error" says: it's the algebraic shortcut for the m and b that make the sum of squared residuals smallest. Here's what that line looks like against the actual data, with each residual drawn in as a dashed segment:
Blue dots are the actual benchmark scores. The green line is the least-squares fit. Each dashed red segment is a residual — the gap between what the line predicts and what actually happened. The fit here is unusually clean; real evaluation data is noisier than this idealized example, because architecture, training data quality, and post-training choices all move the score too — not size alone.
Squared error also gives you a second, very readable number: R² (the coefficient of determination). It compares two quantities: how much the actual scores vary around their own mean (total variation), and how much is left over after your line's predictions (unexplained variation).
Read that as: about 99% of the variation in benchmark score, across this set of checkpoints, is explained by model size alone. Only about 1% is left unexplained — noise, or factors the line doesn't capture. R² always falls between 0 (the line explains nothing) and 1 (the line explains everything), and for a simple line like this one it's just r² — the correlation coefficient you met in Section 1, squared.
Here's the moment this lesson has been building toward. Go back to what you just did: you had a dataset, you defined two parameters (m and b), and you searched for the values of those parameters that made a loss — the sum of squared residuals — as small as possible.
This also closes the loop back to Lesson 5's MLE intuition: "training = choosing parameters that make the observed data most likely." Minimizing squared error and maximizing likelihood turn out to be the same search, under very common assumptions about how noise behaves. You didn't need that fact to fit the line above — you just needed to minimize squared error directly — but it's why the two ideas rhyme so closely.
One more thread to pull, and it's exactly where this course goes next: squared error is a perfectly good loss when your model outputs a number (a benchmark score, a price, a latency). But a huge share of AI systems — classifiers, language models choosing the next token — output probabilities instead of numbers. Squared error isn't the natural loss for that case. Lesson 7 introduces the loss built for exactly this situation: cross-entropy, grounded in information theory. Lesson 8 then formalizes "loss function" in general, and covers what happens when a line — or any model — fits its training data too well.
1. Pearson's r for prompt length vs. response quality across 500 conversations comes out to r = 0.03. What's the most accurate conclusion?
2. Conversations using more emoji get flagged as toxic more often. What's the most likely explanation, per this lesson?
3. When you fit a regression line y = mx + b to data, what are you choosing m and b to minimize?
4. A regression line fitting model size to benchmark score has R² = 0.99. What does that mean?
5. How does fitting a regression line connect to what Lesson 5 called "training" (choosing parameters that make observed data most likely)?