You can already describe a dataset you can see. This lesson is about reasoning over outcomes you haven't seen yet — the math every classifier, spam filter, and hallucination detector runs on.
In Lesson 2 you described the shape of data sitting in front of you — mean, median, spread. That's looking backward at what happened. Probability is looking forward: given what you know, how likely is an outcome you haven't observed? Every AI classifier — spam filter, content moderator, fraud detector, hallucination checker — is a machine that outputs a probability and then someone (or something) decides what to do with it. This lesson gives you the vocabulary and the one theorem that makes that reasoning rigorous: Bayes' Theorem.
Probability is a number between 0 and 1 assigned to an outcome, where 0 means "never happens" and 1 means "always happens." Across every possible outcome of a situation, the probabilities must sum to exactly 1 — something has to happen.
Say a content-moderation model classifies a post into exactly one of three categories. A well-formed model output might look like this:
| Outcome | Probability |
|---|---|
| Safe | 0.82 |
| Needs review | 0.15 |
| Violates policy | 0.03 |
That's not a coincidence — it's a requirement. If you've ever looked at the raw output of a classification model and seen a row of numbers that add to 1.0, you were looking at a probability distribution over outcomes. This is the same "distribution" idea from Lesson 2, just now describing uncertain future outcomes instead of a pile of numbers you already collected. (Lesson 4 gives these shapes formal names — Normal, Bernoulli, Categorical. For now, just recognize the pattern: a set of numbers, each between 0 and 1, summing to 1 across every possible outcome.)
Most useful probabilities in AI aren't standalone — they're conditional. You rarely ask "what's the probability an email is spam?" in a vacuum. You ask "what's the probability an email is spam given that it contains the word 'free'?" That's a completely different, usually much higher, number.
Notation: P(A|B) reads as "the probability of A, given that B is true." The bar isn't division — it means "restrict your attention to the world where B already happened, then ask about A."
Say you run a search engine and you're evaluating retrieval quality. Two different questions:
Across every document in your entire index, what fraction are relevant to a random query? Almost certainly tiny — most documents have nothing to do with any given search.
Restrict to just the documents your search engine returned for a query. What fraction of those are actually relevant? This is what you actually care about — it's your retrieval precision.
Conditioning narrows the population you're reasoning about. "P(spam)" looks at all email ever sent. "P(spam | contains 'free')" looks only at the slice of email containing that word — a much smaller, much spammier slice. This narrowing is the entire mechanism behind Bayes' Theorem, which you'll build in section 4.
Two events are independent if knowing one happened tells you nothing about whether the other happened. They're dependent if knowing one changes your estimate of the other.
Whether an LLM API call lands on a Tuesday, and whether that response happens to contain a hallucinated fact. Day of the week tells you essentially nothing about hallucination risk.
Whether an email contains the word "free," and whether it's spam. Knowing the word appeared substantially raises your estimate that the email is spam. These two facts are entangled.
Formally: A and B are independent exactly when P(A|B) = P(A) — conditioning on B doesn't move the probability of A at all. The moment P(A|B) ≠ P(A), the events are dependent, and that gap between the two numbers is precisely the signal a classifier is built to exploit. A spam filter is, in essence, a machine for finding words and patterns where P(spam | pattern) is far higher than the baseline P(spam).
Here's the problem Bayes' Theorem solves. You usually know P(evidence | hypothesis) — that's easy to measure from data you've already labeled. Example: "given that an email really is spam, what fraction of the time does it contain 'free'?" You can just count that up from a labeled dataset.
But that's backwards from what you actually want to know. You want P(hypothesis | evidence) — "given that this new email contains 'free', what's the probability it's spam?" That's the question you can't directly count, because you don't yet know whether this new email is spam. That's exactly what you're trying to find out.
Bayes' Theorem is the formula that flips one into the other:
Each piece has a name, and each name matters because you'll see them again anywhere probabilistic reasoning shows up in AI:
| Term | Name | Meaning |
|---|---|---|
| P(hypothesis) | Prior | What you believed before seeing this evidence — the baseline rate. |
| P(evidence | hypothesis) | Likelihood | How probable this evidence is, assuming the hypothesis is true. |
| P(evidence) | Evidence (or normalizer) | How probable this evidence is overall, across every hypothesis — it rescales the answer back into a valid 0–1 probability. |
| P(hypothesis | evidence) | Posterior | What you believe after seeing the evidence — the updated, more informed answer. |
In one sentence: Bayes' Theorem is how you update a belief when new evidence arrives. Start with a prior, weigh it by how well the evidence fits, and land on a posterior. This is the mathematical backbone of every classifier that reasons under uncertainty — it's quite literally how a spam filter, a fraud detector, or a hallucination checker turns "I saw this pattern" into "here's my updated probability that something bad is going on."
Say your team ships a lightweight classifier that flags LLM responses likely to contain a hallucinated fact. It's been evaluated and you know two things about it from a labeled test set:
The detector flags a response. What's the probability it's actually a hallucination? Most people's gut says "90%, since that's how good the detector is." Let's compute it properly.
Imagine 10,000 responses, and split them by the real numbers above:
Now apply Bayes' Theorem directly — the posterior is the true positives divided by all flags, true and false:
Here's the flow of the calculation, laid out the way the pieces move through the formula:
A detector with a 90% catch rate and only a 5% false-flag rate still turns out to be wrong about 73% of the time when it raises a flag. That's not a bad detector — it's the base rate working against you.
See it laid out as a population split — 10,000 responses, divided first by ground truth (hallucinated vs. clean) and then by what the detector said:
Read the "Detector flagged" row: 180 real catches against 490 false alarms. The false-positive column is 49x wider than the true-condition column is rare — that imbalance, not the detector's quality, is what drags the posterior down to ~27%.
Why does this happen? Because the clean population (9,800) is so much larger than the hallucinated population (200) that even a small false-positive rate (5%) generates a large false-positive count (490) — more than double the true positives (180). The detector's 90%/5% numbers describe performance within each true class. They say nothing on their own about what a flag means until you weight them by how common each class actually is. That weighting is exactly what the prior does in Bayes' Theorem.
This isn't a hallucination-detector-only problem. The identical arithmetic shows up in:
Fraud might be 0.1% of transactions. A "99% accurate" fraud model can still flag mostly-legitimate transactions, because legitimate transactions vastly outnumber fraudulent ones.
The textbook version of this problem — a 99%-accurate test for a disease affecting 1 in 1,000 people is still wrong most of the time when it comes back positive.
If policy violations are rare relative to total posts, even a strong classifier generates a flood of false flags for human reviewers to wade through.
Many classifiers — spam filters, content moderators, intent classifiers — are literally computing (an approximation of) P(class | features) using exactly this structure: a prior over classes, a likelihood of the observed features given each class, normalized by the evidence. "Naive Bayes" classifiers do this almost exactly as shown above, just with more features multiplied together under the independence assumption from section 3.
Once you know a 90%/5% detector produces a 27% posterior on a 2%-prevalence problem, you can make an informed decision: route flagged responses to human review instead of auto-rejecting them, or raise the detector's precision, or accept the false-positive cost because catching 90% of real hallucinations is still worth 490 wasted reviews. Without knowing the base rate, you can't reason about this tradeoff at all — you'd just be guessing.
When someone hands you a model eval and reports "95% accuracy," your first question should be: accurate on what distribution, and what's the base rate of the thing that matters? A model can post a great accuracy number on an imbalanced test set (echoing the class-imbalance pitfall from Lesson 2) while still being close to useless at the one job you actually need it to do. Lesson 9 builds the full toolkit — precision, recall, confusion matrices — for auditing exactly this.
1. A model outputs P(safe) = 0.7, P(needs review) = 0.2, P(violates policy) = 0.1 for a post. What must be true of any valid probability distribution over outcomes?
2. What does P(relevant | matched query) mean, in plain English?
3. Which pair of events is closest to independent?
4. In Bayes' Theorem, P(hypothesis) — your belief before seeing new evidence — is called the:
5. A hallucination detector catches 90% of real hallucinations and false-flags only 5% of clean responses — but real hallucinations are rare (2% of all responses). Why can P(real hallucination | flagged) still be as low as ~27%?