Lesson 4 of 10 · Statistics for AI, GenAI & LLMs

The Distributions That Run Machine Learning

Four named shapes of uncertainty — and the one that's quietly running every single token your LLM ever generates.

In Lesson 2 you learned to describe a pile of numbers: mean, variance, standard deviation, shape. In Lesson 3 you learned to reason about uncertain events with probability and Bayes' theorem. This lesson connects them: a probability distribution is a named, reusable "shape" that assigns a probability to every possible outcome of some random process — and it's fully described by a small handful of numbers called parameters, often exactly the mean and variance you already know how to compute.

Instead of re-deriving the shape of your data from scratch every time, you recognize which named distribution fits, plug in its parameters, and immediately know a lot about how it behaves. Four distributions below show up constantly in AI systems — one of them, in particular, is the mathematical object sitting at the output of every LLM.

Key Insight A distribution is a rule, not a dataset. "Normal distribution" doesn't mean one specific set of numbers — it means a whole family of bell-shaped rules, one for every choice of mean and variance. Naming the family and stating the parameters is enough to describe the whole thing.

1. Normal (Gaussian) — the bell curve

The Normal distribution, also called Gaussian, models a continuous quantity that clusters symmetrically around a central value, with values far from the center becoming rapidly less likely. It's the classic bell curve you've probably seen before — now you know exactly what draws it.

Models
A continuous value that clusters around a center with symmetric, tapering spread
Parameters
mean (μ) — the center — and variance (σ²) — the spread. Exactly the mean and variance from Lesson 2.
Shape
Symmetric bell curve. Taller and narrower for small variance; flatter and wider for large variance.

Two AI examples where this exact shape shows up by name:

Weight initialization

Before a neural network sees any training data, its weights need starting values. The standard approach draws each weight from a Normal distribution centered at 0 with a small, carefully chosen variance (schemes like Xavier/Glorot or He initialization tune that variance to the layer size, keeping signals from exploding or vanishing as they pass through the network).

Diffusion model noise

Image-generating diffusion models (the family behind Stable Diffusion, DALL-E, Midjourney) are trained by repeatedly adding Gaussian noise to an image until it's pure static, then training a model to reverse that process one small denoising step at a time. The noise added at every step is drawn from a Normal distribution — it's the literal engine of the "diffusion" in the name.

Why does the Normal distribution show up so often? A rough intuition for now, formalized in Lesson 5: when a quantity is the sum of many small, independent random effects, the result tends toward a bell shape — regardless of what the individual effects looked like. That's a preview of the Central Limit Theorem, not something you need to prove yet.

2. Bernoulli — a single yes/no trial

The Bernoulli distribution models the simplest possible random process: one trial, two outcomes. Success or failure. Yes or no. Spam or not-spam. It's named after Jacob Bernoulli and it's the atomic building block that several other distributions (including the next one) are built from.

Models
A single trial with exactly two outcomes, conventionally labeled 1 ("success") and 0 ("failure")
Parameter
p — the probability of success. The probability of failure is automatically 1 − p, since the two outcomes must add to 1.
Shape
Just two bars: height p at outcome 1, height 1−p at outcome 0. No curve, no spread — one number decides everything.

In AI systems, Bernoulli shows up wherever a label or decision is binary:

Pitfall A classifier's predicted probability (say, 0.83 for "this email is spam") is not itself Bernoulli-distributed — it's a single number between 0 and 1. What's Bernoulli-distributed is the underlying true label, which the model is trying to predict. Don't confuse the parameter p with the model's output score, even though they're often compared directly.

3. Categorical / Multinomial — the one running your LLM right now

The Categorical distribution generalizes Bernoulli from two outcomes to K outcomes — any fixed number of categories, each with its own probability, all summing to exactly 1. Roll a K-sided die once, and the result follows a categorical distribution. (Do it N times and tally how many times each side came up, and the tally follows a related distribution called Multinomial — categorical is one draw, multinomial is the count across many draws. The single-draw case is the one that matters most for what's next.)

Models
One outcome chosen from K distinct categories, each with its own probability
Parameters
K probabilities, p₁, p₂, … , p_K, one per category, constrained to sum to exactly 1
Shape
A bar per category, heights equal to each category's probability. No inherent order — categories are just labels.
Why This Matters Every time a large language model generates a token, its final layer produces a softmax vector — one number per word/subword in the entire vocabulary (often 30,000 to 100,000+ entries), all positive, all summing to exactly 1. That vector is not "sort of like" a categorical distribution. It is a categorical distribution: K = vocabulary size, and p₁ … p_K are its parameters, recomputed fresh at every single token position. Generating text means sampling one outcome from this categorical distribution, over and over, one token at a time.

A tiny worked example makes this concrete. Say a toy vocabulary has only 5 tokens, and the model has just generated "The cat sat on the" — here's what its softmax output over that vocabulary might look like for the next token:

Next tokenProbability (p)
"mat"0.42
"floor"0.31
"chair"0.15
"roof"0.08
"moon"0.04

Five numbers, one per category, summing to 1.00 — a categorical distribution with K = 5. "Greedy" generation always picks the highest-probability token ("mat"). Actual sampling draws randomly according to these probabilities — "mat" is most likely to come up, but "floor" or even "moon" can still be sampled, exactly as a weighted die roll would produce the less-likely face sometimes. Lesson 10 comes back to exactly this mechanism — temperature, top-k, and top-p sampling are all just different ways of reshaping this categorical distribution before drawing from it.

4. Poisson — counting rare events in a fixed window

The Poisson distribution models how many times a rare, random event happens in a fixed interval of time or space — given only the average rate at which it happens. It answers questions shaped like "given that this happens about λ times per interval on average, what's the probability it happens exactly k times this particular interval?"

Models
A count of events (0, 1, 2, 3, …) occurring in a fixed interval, when events are rare and happen independently
Parameter
λ (lambda) — the average rate of events per interval. This single number is both the distribution's mean and its variance.
Shape
Discrete spikes at whole numbers only (you can't have 2.5 events). Right-skewed for small λ — the long-tail shape from Lesson 2 shows up here as a named distribution.

Two places Poisson shows up in AI systems:

API request volume

"We get about 40 requests per minute" is a Poisson rate, λ = 40. Capacity planning uses this to estimate how often you'll see spikes well above 40 in any given minute — critical for setting autoscaling thresholds and avoiding both wasted capacity and outages.

Rare-error counting

"Our model hallucinates in roughly 3 out of every 100 generations" sets a Poisson-style rate for hallucinations per batch. Poisson-based reasoning lets you estimate how unusual it would be to see, say, 10 hallucinations in the next 100 generations if the true rate really is 3 — a first step toward statistical significance testing in Lesson 9.

Worked mini-example: if API requests hit an endpoint at an average rate of λ = 2 per second, the Poisson distribution tells you the probability of seeing exactly k requests in the next second:

P(0 requests) = 0.135 (13.5%)
P(1 request) = 0.271 (27.1%)
P(2 requests) = 0.271 (27.1%) <- matches the average, most likely single outcome
P(3 requests) = 0.180 (18.0%)
P(4 requests) = 0.090 ( 9.0%)
P(5 requests) = 0.036 ( 3.6%)
P(6+ requests) < 2% — increasingly rare, but never zero

Notice the shape: it peaks near λ, then tapers off — right-skewed, exactly like the token-count and latency distributions from Lesson 2. That's not a coincidence; Poisson is often the named distribution sitting underneath a real-world right-skewed count.

5. Shapes, side by side

Here's what these four distributions actually look like, drawn to the same visual language. Notice which are smooth curves over continuous values and which are bars over discrete, separate outcomes.

Normal — weight initialization
μ (mean) σ (spread) ← σ → weight value →
Bernoulli — spam / not-spam label
1.0 0 0.7 0 (not spam) 0.3 1 (spam) probability ↓
Categorical — LLM next-token softmax
.42 .31 .15 .08 .04 mat floor chair roof moon next-token candidates (K = 5)
Poisson — requests per second, λ=2
0 1 2 3 4 5 6 7 requests in the next second →

Same underlying idea — a rule assigning probability to outcomes — four different shapes depending on what's being modeled. Normal is smooth and continuous; the other three are bars over discrete, separate outcomes.

6. Quick-reference comparison

DistributionModels…Parameter(s)AI example
Normal / Gaussian A continuous value clustered symmetrically around a center mean (μ), variance (σ²) Weight initialization; noise in diffusion models
Bernoulli One yes/no trial p (success probability) A binary classifier's true label (spam / not-spam)
Categorical / Multinomial One outcome among K categories (categorical); tallies across N draws (multinomial) p₁ … p_K, summing to 1 An LLM's softmax output over the entire vocabulary at each token step
Poisson A count of rare events in a fixed interval λ (average rate per interval) API requests per minute; hallucinations per N generations
Pitfall This lesson only describes the shapes — it does not show you how to figure out a real dataset's parameters (what's the actual λ for your API traffic? what's the actual μ and σ for your weight distribution?). That's estimation, and it's the entire subject of Lesson 5. Naming the right distribution is step one; fitting it to real data is step two.

Recap

✅ Check Yourself

1. What are the two parameters of a Normal (Gaussian) distribution?

Minimum and maximum value
Mean (the center) and variance (the spread) — the same two numbers from Lesson 2
Median and mode

2. A binary spam classifier's true label (spam = 1, not spam = 0) is best modeled by which distribution?

Bernoulli — a single trial with two outcomes, parameter p
Normal — because probabilities are continuous
Poisson — because spam is a rare event

3. An LLM's softmax output at each token-generation step — a vector of probabilities, one per vocabulary word, summing to 1 — is an example of which distribution?

Normal, because it's a smooth probability curve
Bernoulli, because the model is choosing yes or no for each word
Categorical, because it's one outcome chosen among K categories (the vocabulary), each with its own probability

4. Your team says "we average 40 API requests per minute." Which distribution helps you estimate how likely it is to see, say, 60 requests in a given minute?

Bernoulli, since a request either happens or it doesn't
Poisson, since it models counts of events in a fixed interval given an average rate (λ)
Categorical, since there are many possible request counts

5. True or false: this lesson taught you how to calculate the actual parameters (like λ or σ²) from your own real dataset.

True — you now know how to fit any distribution to data
False — this lesson only described the distributions' shapes and parameters; fitting parameters to real data (estimation) is Lesson 5