Four named shapes of uncertainty — and the one that's quietly running every single token your LLM ever generates.
In Lesson 2 you learned to describe a pile of numbers: mean, variance, standard deviation, shape. In Lesson 3 you learned to reason about uncertain events with probability and Bayes' theorem. This lesson connects them: a probability distribution is a named, reusable "shape" that assigns a probability to every possible outcome of some random process — and it's fully described by a small handful of numbers called parameters, often exactly the mean and variance you already know how to compute.
Instead of re-deriving the shape of your data from scratch every time, you recognize which named distribution fits, plug in its parameters, and immediately know a lot about how it behaves. Four distributions below show up constantly in AI systems — one of them, in particular, is the mathematical object sitting at the output of every LLM.
The Normal distribution, also called Gaussian, models a continuous quantity that clusters symmetrically around a central value, with values far from the center becoming rapidly less likely. It's the classic bell curve you've probably seen before — now you know exactly what draws it.
Two AI examples where this exact shape shows up by name:
Before a neural network sees any training data, its weights need starting values. The standard approach draws each weight from a Normal distribution centered at 0 with a small, carefully chosen variance (schemes like Xavier/Glorot or He initialization tune that variance to the layer size, keeping signals from exploding or vanishing as they pass through the network).
Image-generating diffusion models (the family behind Stable Diffusion, DALL-E, Midjourney) are trained by repeatedly adding Gaussian noise to an image until it's pure static, then training a model to reverse that process one small denoising step at a time. The noise added at every step is drawn from a Normal distribution — it's the literal engine of the "diffusion" in the name.
Why does the Normal distribution show up so often? A rough intuition for now, formalized in Lesson 5: when a quantity is the sum of many small, independent random effects, the result tends toward a bell shape — regardless of what the individual effects looked like. That's a preview of the Central Limit Theorem, not something you need to prove yet.
The Bernoulli distribution models the simplest possible random process: one trial, two outcomes. Success or failure. Yes or no. Spam or not-spam. It's named after Jacob Bernoulli and it's the atomic building block that several other distributions (including the next one) are built from.
In AI systems, Bernoulli shows up wherever a label or decision is binary:
The Categorical distribution generalizes Bernoulli from two outcomes to K outcomes — any fixed number of categories, each with its own probability, all summing to exactly 1. Roll a K-sided die once, and the result follows a categorical distribution. (Do it N times and tally how many times each side came up, and the tally follows a related distribution called Multinomial — categorical is one draw, multinomial is the count across many draws. The single-draw case is the one that matters most for what's next.)
A tiny worked example makes this concrete. Say a toy vocabulary has only 5 tokens, and the model has just generated "The cat sat on the" — here's what its softmax output over that vocabulary might look like for the next token:
| Next token | Probability (p) |
|---|---|
| "mat" | 0.42 |
| "floor" | 0.31 |
| "chair" | 0.15 |
| "roof" | 0.08 |
| "moon" | 0.04 |
Five numbers, one per category, summing to 1.00 — a categorical distribution with K = 5. "Greedy" generation always picks the highest-probability token ("mat"). Actual sampling draws randomly according to these probabilities — "mat" is most likely to come up, but "floor" or even "moon" can still be sampled, exactly as a weighted die roll would produce the less-likely face sometimes. Lesson 10 comes back to exactly this mechanism — temperature, top-k, and top-p sampling are all just different ways of reshaping this categorical distribution before drawing from it.
The Poisson distribution models how many times a rare, random event happens in a fixed interval of time or space — given only the average rate at which it happens. It answers questions shaped like "given that this happens about λ times per interval on average, what's the probability it happens exactly k times this particular interval?"
Two places Poisson shows up in AI systems:
"We get about 40 requests per minute" is a Poisson rate, λ = 40. Capacity planning uses this to estimate how often you'll see spikes well above 40 in any given minute — critical for setting autoscaling thresholds and avoiding both wasted capacity and outages.
"Our model hallucinates in roughly 3 out of every 100 generations" sets a Poisson-style rate for hallucinations per batch. Poisson-based reasoning lets you estimate how unusual it would be to see, say, 10 hallucinations in the next 100 generations if the true rate really is 3 — a first step toward statistical significance testing in Lesson 9.
Worked mini-example: if API requests hit an endpoint at an average rate of λ = 2 per second, the Poisson distribution tells you the probability of seeing exactly k requests in the next second:
Notice the shape: it peaks near λ, then tapers off — right-skewed, exactly like the token-count and latency distributions from Lesson 2. That's not a coincidence; Poisson is often the named distribution sitting underneath a real-world right-skewed count.
Here's what these four distributions actually look like, drawn to the same visual language. Notice which are smooth curves over continuous values and which are bars over discrete, separate outcomes.
Same underlying idea — a rule assigning probability to outcomes — four different shapes depending on what's being modeled. Normal is smooth and continuous; the other three are bars over discrete, separate outcomes.
| Distribution | Models… | Parameter(s) | AI example |
|---|---|---|---|
| Normal / Gaussian | A continuous value clustered symmetrically around a center | mean (μ), variance (σ²) | Weight initialization; noise in diffusion models |
| Bernoulli | One yes/no trial | p (success probability) | A binary classifier's true label (spam / not-spam) |
| Categorical / Multinomial | One outcome among K categories (categorical); tallies across N draws (multinomial) | p₁ … p_K, summing to 1 | An LLM's softmax output over the entire vocabulary at each token step |
| Poisson | A count of rare events in a fixed interval | λ (average rate per interval) | API requests per minute; hallucinations per N generations |
1. What are the two parameters of a Normal (Gaussian) distribution?
2. A binary spam classifier's true label (spam = 1, not spam = 0) is best modeled by which distribution?
3. An LLM's softmax output at each token-generation step — a vector of probabilities, one per vocabulary word, summing to 1 — is an example of which distribution?
4. Your team says "we average 40 API requests per minute." Which distribution helps you estimate how likely it is to see, say, 60 requests in a given minute?
5. True or false: this lesson taught you how to calculate the actual parameters (like λ or σ²) from your own real dataset.