Lesson 6 of 10 · Statistics for AI, GenAI & LLMs

Correlation, Regression & Why Models "Fit" Data

Two variables that move together, a line that summarizes the pattern, and the realization that fitting that line is a miniature version of everything Lesson 8 will call "training."

Lesson 5 gave you the core idea behind training: a model has parameters, and training means choosing the parameters that make the observed data most likely (or, put simply, that fit the data best). That idea was abstract there. This lesson makes it concrete and visual. You're going to fit an actual line to actual points, watch the fit improve as you choose better numbers, and see — directly, not by analogy — that this is what "training a model" means at the smallest possible scale.

1. Correlation: does one thing move with another?

Before you can fit a line, you need a way to ask whether two variables are related at all. That's what the Pearson correlation coefficient, almost always written r, measures: the strength and direction of a linear relationship between two numeric variables.

r is always between -1 and +1, no matter what units your data is in. Here are two AI examples that sit at opposite ends of that scale:

Prompt length vs. response quality

Across a batch of logged conversations, word count of the prompt vs. a 1–10 human quality rating of the response comes out around r ≈ 0.15. Barely any linear relationship — long prompts aren't reliably better or worse than short ones.

Model size vs. benchmark score

Across a family of model checkpoints at different parameter counts, size vs. benchmark accuracy comes out around r ≈ 0.99. Strong positive relationship — bigger checkpoints in this family score higher, almost every time.

prompt length vs. quality
r ≈ 0.15
model size vs. benchmark
r ≈ 0.99
−1 (strong negative)0 (no linear relationship)+1 (strong positive)
Key Insight r only measures linear relationships. Two variables can have a strong, obvious, real relationship — a U-shape, a threshold effect, a curve — and still produce r ≈ 0 because that shape isn't a straight line. r ≈ 0 means "no linear pattern," not "no pattern at all." Always look at the scatter plot, not just the number.

2. The trap: correlation is not causation

Here's the sticky one. Say your moderation team notices something in the data: conversations that use more emoji get flagged as toxic more often. The correlation is real and measurable. It is tempting — and wrong — to conclude that emoji use causes toxicity, or to build a moderation rule that penalizes emoji-heavy messages.

What's actually happening is almost certainly this: a third factor, casual, informal conversational register, drives both. People writing casually use more emoji and are more likely to use blunt, aggressive, or crude language that trips a toxicity classifier. Neither emoji nor toxicity causes the other — they're both downstream of the same underlying cause.

Casual, informal register
(the actual common cause)
↙    ↘
More emoji use
↔correlated,
not causal
Higher toxicity-flag rate

The correlation people notice sits along the bottom — but the real arrows point down from the shared cause above it. This is called a confounding variable.

Pitfall If you'd shipped a rule that suppressed emoji-heavy messages, you'd have degraded the product for every casual-but-friendly user while barely touching actual toxic content — because emoji was never the cause. This exact trap shows up constantly in AI systems: "users who ask more follow-up questions churn less" (maybe engagement itself drives both), "longer conversations correlate with lower satisfaction" (maybe frustrated users just send more messages trying to get unstuck). Before you act on a correlation, ask: is there a third factor that could explain both sides?

3. Simple linear regression: fitting a line to data

Correlation tells you a relationship exists and roughly how strong it is. Regression goes further: it finds the actual line that best summarizes that relationship, so you can make predictions from it. A line is defined by two numbers you already know from algebra:

y = mx + b
m = slope (how much y changes per unit of x)
b = intercept (the value of y when x = 0)

"Fitting" a regression line means choosing m and b so the line sits as close as possible to every point in your dataset. Suppose you evaluate a family of model checkpoints at 8 different sizes on a benchmark, and log these illustrative results:

Model size (B params)Benchmark score (%)
152
258
361
468
570
675
779
883

The trend is obviously upward, but no line passes through all 8 points exactly. Every candidate line will miss some points by some amount. The question is: which line minimizes the total miss?

Minimizing squared error — the same idea as Lesson 2's variance

For any candidate line, the gap between an actual point and the line's prediction at that x-value is called a residual. This is exactly the "deviation from a central value" idea from Lesson 2, just applied to a line instead of a single mean. And the fix for negative-and-positive deviations canceling out is the same fix as before: square them, then find the m and b that make the sum of squared residuals as small as possible. This method is called ordinary least squares.

mean size (x̄) = 4.5, mean score (ȳ) = 68.25
slope m = Σ(x-x̄)(y-ȳ) / Σ(x-x̄)² = 183 / 42
≈ 4.36
intercept b = ȳ - m·x̄ = 68.25 - (4.357 × 4.5)
≈ 48.64
fitted line: score ≈ 4.36 × size_B + 48.64

You don't need to memorize that formula — the point is that it's doing exactly what "minimize squared error" says: it's the algebraic shortcut for the m and b that make the sum of squared residuals smallest. Here's what that line looks like against the actual data, with each residual drawn in as a dashed segment:

45 55 65 75 85 Benchmark accuracy (%) 0 2 4 6 8 Model size (billions of parameters) actual score fitted line residual

Blue dots are the actual benchmark scores. The green line is the least-squares fit. Each dashed red segment is a residual — the gap between what the line predicts and what actually happened. The fit here is unusually clean; real evaluation data is noisier than this idealized example, because architecture, training data quality, and post-training choices all move the score too — not size alone.

4. R²: how much of the variation does the line explain?

Squared error also gives you a second, very readable number: R² (the coefficient of determination). It compares two quantities: how much the actual scores vary around their own mean (total variation), and how much is left over after your line's predictions (unexplained variation).

R² = 1 - (sum of squared residuals) / (total variation around the mean)
= 1 - 6.14 / 803.5
≈ 0.99

Read that as: about 99% of the variation in benchmark score, across this set of checkpoints, is explained by model size alone. Only about 1% is left unexplained — noise, or factors the line doesn't capture. R² always falls between 0 (the line explains nothing) and 1 (the line explains everything), and for a simple line like this one it's just r² — the correlation coefficient you met in Section 1, squared.

Key Insight High R² means "this line is a good summary of this data." It does not mean "this line will always be the best possible model," and it does not mean the relationship is causal, or that the line will still fit well on data very different from what it was fit on. A straight line might not always be the right shape for the pattern — that formal question of good fit vs. too good a fit on the training data comes in Lesson 8.

5. The payoff: fitting a line IS training a (tiny) model

Here's the moment this lesson has been building toward. Go back to what you just did: you had a dataset, you defined two parameters (m and b), and you searched for the values of those parameters that made a loss — the sum of squared residuals — as small as possible.

Why This Matters That is not similar to training a machine learning model. It is training a machine learning model — the smallest, simplest one that exists. Linear regression has exactly two learnable parameters. A modern LLM has billions. But the core loop is identical: define parameters, define a loss that measures how wrong the current parameters are, search for the parameter values that make that loss smallest. Every model, no matter how complex, is doing some version of this same thing.

This also closes the loop back to Lesson 5's MLE intuition: "training = choosing parameters that make the observed data most likely." Minimizing squared error and maximizing likelihood turn out to be the same search, under very common assumptions about how noise behaves. You didn't need that fact to fit the line above — you just needed to minimize squared error directly — but it's why the two ideas rhyme so closely.

One more thread to pull, and it's exactly where this course goes next: squared error is a perfectly good loss when your model outputs a number (a benchmark score, a price, a latency). But a huge share of AI systems — classifiers, language models choosing the next token — output probabilities instead of numbers. Squared error isn't the natural loss for that case. Lesson 7 introduces the loss built for exactly this situation: cross-entropy, grounded in information theory. Lesson 8 then formalizes "loss function" in general, and covers what happens when a line — or any model — fits its training data too well.

Recap

✅ Check Yourself

1. Pearson's r for prompt length vs. response quality across 500 conversations comes out to r = 0.03. What's the most accurate conclusion?

There's no relationship at all between prompt length and quality
There's no meaningful linear relationship — but r = 0.03 can't rule out a nonlinear pattern
The data must contain an error, since r should never be that close to zero

2. Conversations using more emoji get flagged as toxic more often. What's the most likely explanation, per this lesson?

Emoji use directly causes toxic language
Toxic language causes people to use more emoji
Both are driven by a third factor — casual, informal conversational register — so the correlation doesn't mean either one causes the other

3. When you fit a regression line y = mx + b to data, what are you choosing m and b to minimize?

The number of data points that fall above the line
The sum of the squared residuals — the squared vertical distances between actual points and the line
The correlation coefficient r

4. A regression line fitting model size to benchmark score has R² = 0.99. What does that mean?

99% of the individual predictions are exactly correct
The correlation coefficient r must be negative
About 99% of the variation in benchmark score is explained by model size — the line fits the data very well

5. How does fitting a regression line connect to what Lesson 5 called "training" (choosing parameters that make observed data most likely)?

It doesn't — regression is a completely different kind of math from ML training
Choosing the slope and intercept that minimize squared error IS a tiny version of training — picking parameters that best match observed data, exactly like a neural network's much larger set of weights
Regression only applies to categorical data, not to anything a model learns