Source: Comprehensive Guide to Probability Theory and Applications
Tags: uniform distribution, binomial distribution, Bernoulli trials, conditional probability, Bayes' theorem, posterior probability, prior probability, likelihood, total probability, partitioning
Difficulty: Intermediate Prerequisites: Probability foundations (sample spaces, events, inclusion-exclusion, complement rule). Familiarity with combinatorics ("n choose k") is helpful for the binomial section.
Once you can assign probabilities to events (covered in the foundations notes), the next step is learning the standard distributions that model common experiments and, crucially, how to update probabilities when new information arrives. This set of notes covers the uniform and binomial distributions, conditional probability, Bayes' Theorem, and the law of total probability. These tools are the workhorses of applied probability: they show up in medical diagnosis, spam filtering, quality control, and almost every exam in this course.
The uniform distribution gives every outcome equal weight. The binomial distribution counts successes in repeated independent trials. Conditional probability tells you how the chance of one event changes when you know another event has occurred, and Bayes' Theorem lets you reverse that conditioning to go from "effect given cause" to "cause given effect." The law of total probability ties everything together by breaking a hard probability into manageable pieces.
Uniform distribution
A probability distribution where every outcome in the sample space is equally likely. For n outcomes, each has probability 1/n. Think of it as the "perfectly fair" scenario: a fair die, a fair coin, drawing a name from a hat.
Bernoulli trial
A single experiment with exactly two outcomes, usually called "success" (probability p) and "failure" (probability 1 − p). In simple terms, any yes/no, pass/fail, heads/tails experiment.
Binomial distribution
The distribution of the number of successes in n independent Bernoulli trials, each with the same success probability p. Think of it as: "I flip a coin n times; how likely is it I get exactly k heads?"
Conditional probability, P(A|B)
The probability of event A occurring given that event B has already occurred. Calculated as P(A ∩ B) / P(B). In simple terms, it is the probability of A in the smaller universe where B is known to be true.
Bayes' Theorem
A formula that lets you "flip" a conditional probability: given P(B|A) and the individual probabilities of A and B, you can find P(A|B). Think of it as: "I know how likely the symptom is given the disease; now I want to know how likely the disease is given the symptom."
Prior probability
The initial estimate of the probability of an event before new evidence is considered. In simple terms, your starting belief before you learn anything new.
Posterior probability
The updated probability of an event after incorporating new evidence via Bayes' Theorem. In simple terms, your revised belief after the new data arrives.
Law of total probability
A rule that computes the probability of an event by summing over all the ways it can happen, weighted by the probability of each scenario. Think of it as breaking a hard problem into easier slices and adding them up.
Partition
A collection of mutually exclusive, collectively exhaustive events (they do not overlap, and together they cover the entire sample space). In simple terms, a clean way to split the sample space into non-overlapping pieces that account for everything.
Every outcome has probability 1/n where n = |S|.
Most "fair" scenarios default to this: fair dice, fair coins, random selection from a list.
Computing event probability reduces to counting: P(A) = |A| / |S|.
Setup: n independent Bernoulli trials, each with success probability p.
The probability of exactly k successes:
P(X = k) = C(n, k) · pᵏ · (1 − p)ⁿ⁻ᵏ
C(n, k) is "n choose k" = n! / (k!(n − k)!)
Worked example (biased coin):
Coin has P(Heads) = 2/3, P(Tails) = 1/3.
Flip 7 times. Probability of exactly 4 heads:
P(X = 4) = C(7, 4) · (2/3)⁴ · (1/3)³ = 35 · 16/81 · 1/27 = 560/2187 ≈ 0.256.
Definition: P(A|B) = P(A ∩ B) / P(B), provided P(B) > 0.
Interpretation: restrict the sample space to only those outcomes where B occurred, then ask how much of that restricted space also belongs to A.
This is the gateway to every "given that" question on an exam.
Formula:
P(A|B) = [P(B|A) · P(A)] / P(B)
The denominator P(B) is often computed via the law of total probability (see below).
Worked example (medical diagnosis):
P(Meningitis) = 0.0001 (prior, very rare disease).
P(Stiff neck | Meningitis) = 0.8 (likelihood, most patients with meningitis have a stiff neck).
P(Stiff neck) = 0.1 (stiff necks are fairly common in the general population).
P(Meningitis | Stiff neck) = (0.8 × 0.0001) / 0.1 = 0.0008.
Even with a strong symptom, the posterior is tiny because the disease is so rare. This is the base-rate effect, one of the most important lessons in applied probability.
If b₁, b₂, …, bₙ form a partition of S, then for any event a:
P(a) = Σᵢ P(a | bᵢ) · P(bᵢ)
This lets you decompose a tricky probability into conditional slices, compute each one, and add them up.
Worked example (pets):
Population: 60% cats, 40% dogs.
P(friendly | cat) = 0.5, P(friendly | dog) = 0.9.
P(friendly) = 0.5 × 0.6 + 0.9 × 0.4 = 0.30 + 0.36 = 0.66.
Name | Formula |
|---|---|
Uniform probability | P(outcome) = 1/n |
Binomial probability | P(X = k) = C(n, k) · pᵏ · (1 − p)ⁿ⁻ᵏ |
Conditional probability | P(A|B) = P(A ∩ B) / P(B) |
Bayes' Theorem | P(A|B) = P(B|A) · P(A) / P(B) |
Law of total probability | P(A) = Σᵢ P(A|Bᵢ) · P(Bᵢ) |
Bayes' Theorem is the engine behind spam filters (what is the probability this email is spam, given the words it contains?) and medical screening (what is the probability the patient has the disease, given a positive test?). The binomial distribution models quality control: if a factory's defect rate is 2%, what is the probability that a batch of 50 items contains more than 3 defectives? The law of total probability is how actuaries price insurance across different risk segments.
Students often confuse P(A|B) with P(B|A). These are not the same. P(disease | symptom) and P(symptom | disease) can be wildly different, as the meningitis example shows.
Forgetting the base rate (prior probability) when applying Bayes' Theorem. A high likelihood does not guarantee a high posterior if the prior is very small.
Assuming that the binomial formula applies when trials are not independent. If drawing balls from an urn without replacement, the trials are dependent, and you need the hypergeometric distribution instead.
Mixing up the law of total probability with Bayes' Theorem. Total probability computes P(A) by summing over conditions. Bayes' Theorem flips a conditional. They work together, but they answer different questions.
⚠️ Bayes' Theorem problems are near-guaranteed on exams. Practise setting up the numerator (likelihood × prior) and the denominator (total probability) separately before dividing.
⚠️ Binomial questions often ask for "at least" or "at most" k successes. Remember the complement shortcut: P(X ≥ 1) = 1 − P(X = 0).
⚠️ Conditional probability is the most common source of errors in multi-step probability problems. Always check: are you conditioning on the right event?
⚠️ Know when to use the law of total probability to compute P(B) inside Bayes' Theorem. This two-step process (total probability for the denominator, then Bayes for the answer) is a standard exam pattern.
True or false: In a uniform distribution with 8 outcomes, each outcome has probability 1/8.
Fill in the blank: P(A|B) = P(A ∩ B) / ______.
True or false: P(A|B) always equals P(B|A).
A fair coin is flipped 5 times. How many terms does the binomial formula need to compute P(X = 3)?
True or false: The law of total probability requires the conditioning events to form a partition of the sample space.
Answers: 1. True. 2. P(B). 3. False. 4. One term: C(5,3) · (0.5)³ · (0.5)² = 10/32 = 5/16. 5. True.
Q: A biased coin with P(Heads) = 2/3 is flipped 7 times. What is the probability of getting exactly 4 heads?
A: P(X = 4) = C(7,4) · (2/3)⁴ · (1/3)³ = 35 · (16/81) · (1/27) = 560/2187 ≈ 0.256.
Q: P(Meningitis) = 0.0001, P(Stiff neck | Meningitis) = 0.8, P(Stiff neck) = 0.1. What is P(Meningitis | Stiff neck)?
A: By Bayes' Theorem: (0.8 × 0.0001) / 0.1 = 0.00008 / 0.1 = 0.0008, or 0.08%.
Q: In a population of 60% cats and 40% dogs, 50% of cats and 90% of dogs are friendly. What is the probability a randomly chosen pet is friendly?
A: P(friendly) = 0.5 × 0.6 + 0.9 × 0.4 = 0.30 + 0.36 = 0.66, or 66%.
Q: What is the key difference between P(A|B) and P(B|A)?
A: P(A|B) is the probability of A in the restricted universe where B has occurred. P(B|A) is the probability of B in the restricted universe where A has occurred. They use different denominators and can have very different values.
Q: When can you not use the binomial distribution?
A: When trials are not independent or when the probability of success changes between trials. For example, drawing cards from a deck without replacement violates the independence assumption.
Conditional probability and Bayes' Theorem connect directly to machine learning, where classifiers like Naive Bayes are built on these principles.
The binomial distribution is a special case of the more general multinomial distribution (more than two outcome categories) and approaches the normal distribution for large n (the Central Limit Theorem).
The law of total probability is the foundation of Markov chains and hidden Markov models in computer science.
conditional probability formula, Bayes' theorem example, binomial distribution formula, Bernoulli trials, n choose k, prior and posterior probability, law of total probability, partition of sample space, base rate fallacy, medical diagnosis probability, uniform distribution, likelihood, total probability rule, STAT 101, CS foundations Purdue, probability distributions