The four numbers you compute first on any dataset — and why AI datasets lie to you if you only look at one of them.
In Lesson 1 you saw the map: data description is where every AI pipeline starts, before probability, before loss functions, before evaluation. This lesson is where you actually do it. No formal probability theory yet — that's Lesson 3. Right now you're just looking at a pile of numbers (token counts, latencies, dataset labels) and learning to describe it honestly.
The mean (average) is what most people reach for first: add everything up, divide by how many there are. The median is the middle value when you sort the data. In a lot of real-world AI data, these two numbers disagree — sharply — and the disagreement itself is a signal worth reading.
Here's a dataset of token counts from 10 responses your LLM generated:
| Response | Tokens |
|---|---|
| 1 | 115 |
| 2 | 118 |
| 3 | 119 |
| 4 | 120 |
| 5 | 121 |
| 6 | 122 |
| 7 | 125 |
| 8 | 128 |
| 9 | 130 |
| 10 | 980 |
Nine ordinary responses, and one that ran away — a generation loop that didn't stop cleanly and produced 980 tokens. Watch what that single row does to the two summaries:
The mean says "typical response length is about 208 tokens." That's not true of a single response in the dataset — nine of the ten cluster tightly around 120. The median, 121.5, is what a typical response actually looks like. One outlier dragged the mean up by nearly 90 tokens.
This is not a toy problem. It's exactly what your data looks like in production:
Most completions are short; a rare few loop, repeat, or hit the max-token cutoff and balloon in length.
Most requests return in milliseconds; a rare few hit a cold start, a retry, or a GPU queue and take seconds.
Most users send a few messages; a rare power user sends hundreds in one sitting.
In every one of these, "average latency" or "average tokens per response" is a number nobody's actual request matches. That's why production dashboards report median and percentiles (like p95 — more on that below) instead of, or alongside, the mean.
Mean and median tell you where the center of your data is. They tell you nothing about how spread out it is. Two datasets can have the same mean and look completely different:
| Dataset | Values (latency, ms) | Mean |
|---|---|---|
| Model A | 198, 199, 200, 201, 202 | 200 |
| Model B | 50, 120, 200, 280, 350 | 200 |
Same mean, wildly different reliability. Model A is consistent; Model B is all over the place — sometimes fast, sometimes almost 2x slower than average. You need a number that captures spread. That's what variance and standard deviation measure.
Take a small, clean dataset — latency in milliseconds for 5 API calls: 200, 210, 195, 205, 190.
The obvious next move is to average those deviations to get a single "typical distance from the mean" number. Try it:
Zero. Always zero, for any dataset — it's a mathematical guarantee, because the mean is defined as the balancing point of the data. Positive and negative deviations cancel perfectly, no matter how spread out the values are. The "average deviation" is useless as a spread measure because it always reports the same answer: nothing.
Squaring a negative number makes it positive, same as taking the absolute value — but squaring has two extra properties that matter enormously for AI:
Variance is now in squared units — "50 milliseconds-squared" isn't something you can picture. Take the square root to bring it back to the original unit:
That's the number people actually mean when they say "the spread of the data": on average, a request's latency sits about 7 ms away from the 200 ms mean. Compact definition:
Mean, median, and standard deviation are three numbers. But two datasets can share all three and still look nothing alike, because they differ in shape — how the values are arranged around the center. The shape that matters most for AI work is whether a distribution is symmetric or has a long tail (skewed).
Same idea, two shapes. In the symmetric case the mean and median land on the same point. In the long-tail case, the rare extreme values on the right pull the mean away from the median — exactly what happened with the 980-token outlier above.
"Skew" just describes which direction the tail stretches. A right-skewed (positively skewed) distribution has a long tail of rare, large values — this is the shape of almost every duration or count you'll measure in AI systems: token counts, latencies, session lengths, file sizes. A left-skewed distribution has a long tail of rare, small values — less common in AI data, but shows up in things like "accuracy scores on an easy benchmark," where most models do well and a few do very badly.
Skew isn't only about continuous numbers like latency — it shows up in categorical data too, as class imbalance. Say you're training a spam classifier and your labeled dataset looks like this:
5,000 "not spam" examples for every 250-ish "spam" examples — the same "long tail" idea applied to a category instead of a number.
This is a distribution shape problem in disguise. A model can score 95% accuracy by predicting "not spam" every single time and never learning what spam looks like. Just like a skewed latency distribution makes "average latency" misleading, a skewed class distribution makes "accuracy" misleading — you'll need better metrics for this in Lesson 9 (precision, recall, confusion matrices).
Before feeding numeric features into most models, you typically standardize them: subtract the mean, divide by the standard deviation, for every value. This is the "z-score" transform:
This rescales every feature so it has mean 0 and standard deviation 1 — putting features that were originally on wildly different scales (say, "user age" 18–90 and "account balance" $0–$5,000,000) onto comparable footing so no single feature dominates just because its raw numbers happen to be bigger. You now know exactly what those two ingredients (mean, standard deviation) mean and where they come from.
A value several standard deviations away from the mean is, by definition, unusual. This is one of the simplest ways to flag bad or contaminated rows before training — a scraped web-text sample that's 40,000 tokens long when the rest of your corpus averages 400, or a labeled "response quality" score of 1 out of 5 when a human rater's other scores cluster at 4–5. Outliers aren't always errors, but they're always worth a look.
Go back to the skewed latency histogram above. If your team reports "average response time: 220ms" and ships, that number is dominated by the fast majority and hides the slow tail — the actual experience of your unluckiest users. Production teams instead report the median (typical experience) alongside a high percentile like p95 ("95% of requests finished faster than this") or p99. The gap between the median and p95 is the long tail, made visible as a single number.
1. Your training dataset's "average session length" is 40 messages, but most users only send 3–5 messages. What's the most likely explanation?
2. Why do we square deviations from the mean instead of just averaging the raw deviations directly?
3. A distribution has a long tail stretching to the right (a few very large values). What's true of its mean vs. median?
4. Why do standard deviation instead of stopping at variance?
5. A spam classifier's training set is 95% "not spam" and 5% "spam." Why is this a distribution-shape problem, similar to a skewed latency dataset?