Lesson 2 of 10 · Statistics for AI, GenAI & LLMs

Describing Data: Mean, Variance, Standard Deviation & Distribution Shape

The four numbers you compute first on any dataset — and why AI datasets lie to you if you only look at one of them.

In Lesson 1 you saw the map: data description is where every AI pipeline starts, before probability, before loss functions, before evaluation. This lesson is where you actually do it. No formal probability theory yet — that's Lesson 3. Right now you're just looking at a pile of numbers (token counts, latencies, dataset labels) and learning to describe it honestly.

1. Mean vs. median — and why one of them lies

The mean (average) is what most people reach for first: add everything up, divide by how many there are. The median is the middle value when you sort the data. In a lot of real-world AI data, these two numbers disagree — sharply — and the disagreement itself is a signal worth reading.

Here's a dataset of token counts from 10 responses your LLM generated:

ResponseTokens
1115
2118
3119
4120
5121
6122
7125
8128
9130
10980

Nine ordinary responses, and one that ran away — a generation loop that didn't stop cleanly and produced 980 tokens. Watch what that single row does to the two summaries:

mean = (115+118+119+120+121+122+125+128+130+980) / 10
= 2078 / 10
= 207.8 tokens
sorted values: 115,118,119,120,121,122,125,128,130,980
median = average of the 5th and 6th values = (121+122)/2
= 121.5 tokens

The mean says "typical response length is about 208 tokens." That's not true of a single response in the dataset — nine of the ten cluster tightly around 120. The median, 121.5, is what a typical response actually looks like. One outlier dragged the mean up by nearly 90 tokens.

Key Insight The mean is sensitive to every value, including extreme ones. The median only cares about what's in the middle — it doesn't move much even if the largest value in your dataset were 980 or 9,800. Statisticians call the median robust to outliers.

This is not a toy problem. It's exactly what your data looks like in production:

Token counts

Most completions are short; a rare few loop, repeat, or hit the max-token cutoff and balloon in length.

API latency

Most requests return in milliseconds; a rare few hit a cold start, a retry, or a GPU queue and take seconds.

User session length

Most users send a few messages; a rare power user sends hundreds in one sitting.

In every one of these, "average latency" or "average tokens per response" is a number nobody's actual request matches. That's why production dashboards report median and percentiles (like p95 — more on that below) instead of, or alongside, the mean.

2. Variance and standard deviation, built from scratch

Mean and median tell you where the center of your data is. They tell you nothing about how spread out it is. Two datasets can have the same mean and look completely different:

DatasetValues (latency, ms)Mean
Model A198, 199, 200, 201, 202200
Model B50, 120, 200, 280, 350200

Same mean, wildly different reliability. Model A is consistent; Model B is all over the place — sometimes fast, sometimes almost 2x slower than average. You need a number that captures spread. That's what variance and standard deviation measure.

Step 1: measure each point's distance from the mean

Take a small, clean dataset — latency in milliseconds for 5 API calls: 200, 210, 195, 205, 190.

mean = (200+210+195+205+190) / 5 = 1000 / 5 = 200 ms
deviation from mean for each point:
200 - 200 = 0
210 - 200 = +10
195 - 200 = -5
205 - 200 = +5
190 - 200 = -10

Step 2: why you can't just average the deviations

The obvious next move is to average those deviations to get a single "typical distance from the mean" number. Try it:

sum of deviations = 0 + 10 + (-5) + 5 + (-10)
= 0

Zero. Always zero, for any dataset — it's a mathematical guarantee, because the mean is defined as the balancing point of the data. Positive and negative deviations cancel perfectly, no matter how spread out the values are. The "average deviation" is useless as a spread measure because it always reports the same answer: nothing.

Pitfall You need to get rid of the signs before averaging, or the spread cancels itself out. There are two honest ways to do it: take the absolute value of each deviation, or square each deviation. Statistics almost always chooses squaring — here's why.

Step 3: square the deviations

Squaring a negative number makes it positive, same as taking the absolute value — but squaring has two extra properties that matter enormously for AI:

squared deviations:
0² = 0
10² = 100
(-5)² = 25
5² = 25
(-10)² = 100
sum of squared deviations = 0+100+25+25+100 = 250
variance = 250 / 5
= 50 ms²

Step 4: undo the squaring to get standard deviation

Variance is now in squared units — "50 milliseconds-squared" isn't something you can picture. Take the square root to bring it back to the original unit:

standard deviation = √variance = √50
≈ 7.07 ms

That's the number people actually mean when they say "the spread of the data": on average, a request's latency sits about 7 ms away from the 200 ms mean. Compact definition:

Key Insight Variance = the average of the squared distances from the mean. Standard deviation = the square root of variance, which puts the spread back into the same units as your original data. Small standard deviation → tightly clustered, predictable data. Large standard deviation → spread out, unpredictable data.

3. Distribution shape: symmetric vs. skewed

Mean, median, and standard deviation are three numbers. But two datasets can share all three and still look nothing alike, because they differ in shape — how the values are arranged around the center. The shape that matters most for AI work is whether a distribution is symmetric or has a long tail (skewed).

Symmetric e.g. model confidence scores around 0.5 mean = median value → Right-skewed (long tail) e.g. API latency — most fast, few very slow median mean value → rare, extreme values stretch the tail right

Same idea, two shapes. In the symmetric case the mean and median land on the same point. In the long-tail case, the rare extreme values on the right pull the mean away from the median — exactly what happened with the 980-token outlier above.

"Skew" just describes which direction the tail stretches. A right-skewed (positively skewed) distribution has a long tail of rare, large values — this is the shape of almost every duration or count you'll measure in AI systems: token counts, latencies, session lengths, file sizes. A left-skewed distribution has a long tail of rare, small values — less common in AI data, but shows up in things like "accuracy scores on an easy benchmark," where most models do well and a few do very badly.

A different kind of shape problem: class imbalance

Skew isn't only about continuous numbers like latency — it shows up in categorical data too, as class imbalance. Say you're training a spam classifier and your labeled dataset looks like this:

95% not spam
5%

5,000 "not spam" examples for every 250-ish "spam" examples — the same "long tail" idea applied to a category instead of a number.

This is a distribution shape problem in disguise. A model can score 95% accuracy by predicting "not spam" every single time and never learning what spam looks like. Just like a skewed latency distribution makes "average latency" misleading, a skewed class distribution makes "accuracy" misleading — you'll need better metrics for this in Lesson 9 (precision, recall, confusion matrices).

4. Why this matters for building AI systems

Why This Matters Every one of these four numbers — mean, median, variance, standard deviation — is something you will compute on real project data, not just on a quiz. Here's where each one shows up.

Normalizing features before training

Before feeding numeric features into most models, you typically standardize them: subtract the mean, divide by the standard deviation, for every value. This is the "z-score" transform:

standardized value = (x - mean) / standard_deviation

This rescales every feature so it has mean 0 and standard deviation 1 — putting features that were originally on wildly different scales (say, "user age" 18–90 and "account balance" $0–$5,000,000) onto comparable footing so no single feature dominates just because its raw numbers happen to be bigger. You now know exactly what those two ingredients (mean, standard deviation) mean and where they come from.

Spotting outliers in training data

A value several standard deviations away from the mean is, by definition, unusual. This is one of the simplest ways to flag bad or contaminated rows before training — a scraped web-text sample that's 40,000 tokens long when the rest of your corpus averages 400, or a labeled "response quality" score of 1 out of 5 when a human rater's other scores cluster at 4–5. Outliers aren't always errors, but they're always worth a look.

"Average latency" can flat-out mislead you

Go back to the skewed latency histogram above. If your team reports "average response time: 220ms" and ships, that number is dominated by the fast majority and hides the slow tail — the actual experience of your unluckiest users. Production teams instead report the median (typical experience) alongside a high percentile like p95 ("95% of requests finished faster than this") or p99. The gap between the median and p95 is the long tail, made visible as a single number.

Pitfall If someone hands you a single "average" number for latency, cost, or token usage with no mention of spread or shape, treat it as incomplete. Ask for the median and p95 too — the mean alone can't tell you whether the data is tightly clustered or has a nasty tail hiding worst-case behavior.

Recap

✅ Check Yourself

1. Your training dataset's "average session length" is 40 messages, but most users only send 3–5 messages. What's the most likely explanation?

The dataset is symmetric and the average is trustworthy
There's a calculation error — mean and median should always match
The distribution is right-skewed — a few power users with very long sessions are pulling the mean up

2. Why do we square deviations from the mean instead of just averaging the raw deviations directly?

Squaring makes the numbers smaller and easier to work with
Raw deviations from the mean always sum to zero, so averaging them directly always gives zero — squaring removes the sign so spread doesn't cancel out
Squaring is only used for negative numbers

3. A distribution has a long tail stretching to the right (a few very large values). What's true of its mean vs. median?

The mean is pulled higher than the median, toward the tail
The mean and median are always equal, regardless of shape
The median is pulled higher than the mean

4. Why do standard deviation instead of stopping at variance?

Variance is always negative and needs correcting
Standard deviation is easier to compute than variance
Variance is in squared units (like ms²); taking the square root brings the spread measure back into the original, interpretable units

5. A spam classifier's training set is 95% "not spam" and 5% "spam." Why is this a distribution-shape problem, similar to a skewed latency dataset?

It isn't related — class imbalance is a completely different topic from numeric skew
Just like a skewed mean hides the true "typical" value, an imbalanced class distribution lets a model score high "accuracy" while never actually learning the minority class
Class imbalance only matters for image models, not text classifiers