๐Ÿ“Š Self-paced course

Statistics for AI, GenAI & LLMs

The statistics you actually need to build, evaluate, and reason about AI, generative models, and LLMs โ€” grounded in AI examples throughout, not dice and coins.

10 lessons Sequential โ€” each builds on the last No prerequisites Quizzes included
  1. Why Statistics Powers AI

    The tour before the road trip: where statistics actually shows up when you build, train, and evaluate AI systems โ€” no math yet, just the map.

  2. Describing Data: Mean, Variance & Distribution Shape

    The four numbers you compute first on any dataset โ€” and why AI datasets lie to you if you only look at one of them.

  3. Probability Foundations & Bayes' Theorem

    Reasoning over outcomes you haven't seen yet โ€” the math every classifier, spam filter, and hallucination detector runs on.

  4. The Distributions That Run Machine Learning

    Four named shapes of uncertainty โ€” and the one that's quietly running every single token your LLM ever generates.

  5. From Sample to Population: Estimation & the CLT

    You never have all the data. Here's the math for how much that should worry you โ€” and how models turn a pile of observations into a fitted distribution.

  6. Correlation, Regression & Why Models "Fit" Data

    Two variables that move together, a line that summarizes the pattern, and the realization that fitting that line is a miniature version of training.

  7. Information Theory: Entropy, Cross-Entropy & KL Divergence

    The three quantities that measure "how surprised" a model is โ€” and the actual loss function running underneath every LLM's training loop.

  8. The Statistics Behind Training a Model

    What actually happens when a model tries to shrink its loss over millions of steps โ€” and the two ways that process quietly goes wrong.

  9. Evaluating AI Models Rigorously

    You built a model. Now: is it actually good, or does it just look good on the one number you checked?

  10. The Statistics Unique to LLMs

    Where it all converges: perplexity, sampling and temperature, calibration, and the statistical shape of hallucination.

Also in this repo

40 flashcards โ€” four per lesson, written for active recall rather than definition-matching. Anki-importable, readable as plain Markdown either way.

A memory map โ€” a 54-node JSON Canvas laying out how the ten lessons connect. Renders in Obsidian.

See the README for how to use each.