Difficulty: Intermediate | Prerequisites: Sample spaces, events, set operations, probability axioms (see Part 1 notes).
Once you have the language of sample spaces and events, the next step is learning how to calculate probabilities of combined and conditional events. This material covers the complement rule, the addition rule, conditional probability, the multiplication rule, Bayes' rule, and independence. These tools let you solve multi-step problems: "given that B has happened, how likely is A?" and "are these two events related or not?" Nearly every applied probability question in the course depends on these rules.
The complement rule, addition rule, and multiplication rule let you compute probabilities of combined events. Conditional probability answers "how likely is A, given that B occurred?" Bayes' rule flips a conditional around, and independence tells you when knowing one event gives you no information about another.
Complement rule
P(E′) = 1 − P(E). The probability that an event does not occur equals one minus the probability that it does.
In simple terms: "the chance of not-E" = 1 minus "the chance of E."
Addition rule (general)
P(A ∪ B) = P(A) + P(B) − P(A ∩ B). Accounts for the overlap so you do not double-count outcomes that belong to both events.
Think of it as: add the two circles, then subtract the overlap you counted twice.
Inclusion-exclusion principle
The three-event extension of the addition rule: P(A ∪ B ∪ C) = P(A) + P(B) + P(C) − P(A ∩ B) − P(A ∩ C) − P(B ∩ C) + P(A ∩ B ∩ C).
In simple terms: add, subtract pairwise overlaps, add back the triple overlap.
Conditional probability
P(A | B) = P(A ∩ B) / P(B), provided P(B) > 0. The probability of A occurring given that B has already occurred.
Think of it as: shrink the sample space to just B, then see how much of A sits inside it.
General multiplication rule
P(A ∩ B) = P(A | B) × P(B). The joint probability of A and B equals the conditional times the probability of the condition.
Independence (of two events)
Events A and B are independent if P(A ∩ B) = P(A) × P(B). Equivalently, P(A | B) = P(A): knowing B happened does not change the probability of A.
In simple terms: one event has no bearing on the other.
Pairwise independence
A collection of events where every pair is independent. Every two-event combination satisfies the independence condition.
Mutual independence
A stronger condition: every sub-collection of events (not just pairs) satisfies the independence condition simultaneously.
Think of it as: pairwise independence checks pairs; mutual independence checks every possible grouping.
Law of total probability
If {A₁, A₂, ..., Aₙ} partition the sample space (mutually exclusive and exhaustive), then P(B) = Σ P(B | Aᵢ) × P(Aᵢ).
In simple terms: break the sample space into slices and add up the probability of B within each slice.
Bayes' rule
P(Aᵢ | B) = [P(B | Aᵢ) × P(Aᵢ)] / P(B). Lets you reverse a conditional: given that B occurred, how likely is each possible cause Aᵢ?
Think of it as: "flip the conditional" using what you know about causes and evidence.
Prior probability
The initial belief about how likely an event is before observing new evidence.
Likelihood
P(B | A): the probability of the observed evidence B under a particular hypothesis A.
Posterior probability
The updated probability of a hypothesis after incorporating new evidence, computed via Bayes' rule.
Partition
A set of events that are mutually exclusive (no overlap) and collectively exhaustive (they cover the entire sample space).
In simple terms: a way of slicing up all possibilities so every outcome lands in exactly one slice.
Complement rule
P(E′) = 1 − P(E).
Useful when computing the probability of "at least one" occurrence. Instead of adding up many cases, compute 1 minus the probability of none.
Example: probability of rolling at least one six in four rolls of a die = 1 − P(no sixes in four rolls).
Addition rule
General form: P(A ∪ B) = P(A) + P(B) − P(A ∩ B).
If A and B are mutually exclusive (A ∩ B = ∅), the intersection term drops out: P(A ∪ B) = P(A) + P(B).
The subtraction corrects for outcomes you counted in both P(A) and P(B).
Inclusion-exclusion for three events
P(A ∪ B ∪ C) = P(A) + P(B) + P(C) − P(A ∩ B) − P(A ∩ C) − P(B ∩ C) + P(A ∩ B ∩ C).
The pattern extends to any number of events: add singles, subtract pairs, add triples, subtract quadruples, and so on.
Conditional probability narrows the sample space. Once you know B has occurred, you ignore every outcome outside B and look at how much of A sits within it.
Formula: P(A | B) = P(A ∩ B) / P(B), provided P(B) > 0.
Reading it: "the probability of A given B."
The vertical bar "|" means "given that" or "conditional on."
General multiplication rule
Rearranging the conditional probability formula gives the joint probability:
P(A ∩ B) = P(A | B) × P(B).
Equivalently, P(A ∩ B) = P(B | A) × P(A).
For independent events, P(A | B) = P(A), so this simplifies to P(A ∩ B) = P(A) × P(B).
The multiplication rule is how you handle "and" questions: the probability that both A and B occur.
Law of total probability
When you can split the sample space into non-overlapping slices (a partition {A₁, A₂, ..., Aₙ}), the probability of any event B is:
P(B) = Σ P(B | Aᵢ) × P(Aᵢ).
Each slice contributes its own share of B's total probability.
This is essential whenever you know conditional probabilities within each slice but need the overall probability.
Bayes' rule
Bayes' rule reverses a conditional. You know P(B | Aᵢ) (the likelihood of the evidence given a cause) and P(Aᵢ) (the prior probability of each cause), and you want P(Aᵢ | B) (the posterior probability of the cause given the evidence).
P(Aᵢ | B) = [P(B | Aᵢ) × P(Aᵢ)] / P(B).
P(B) in the denominator is computed using the law of total probability.
The classic example: a medical test. You know the test's sensitivity (P(positive | disease)) and the disease's prevalence (P(disease)). Bayes' rule tells you P(disease | positive), which is what you and your doctor care about.
The denominator P(B) acts as a normalising constant. It ensures the posterior probabilities across all causes sum to 1.
Two events A and B are independent if knowing that one occurred does not change the probability of the other.
Test for independence: P(A ∩ B) = P(A) × P(B). If this holds, the events are independent.
Equivalently: P(A | B) = P(A) and P(B | A) = P(B).
Mutually exclusive vs independent
These are two different concepts, and they are almost opposites.
Mutually exclusive events cannot occur together: P(A ∩ B) = 0.
Independent events can occur together, and their joint probability is the product of their individual probabilities.
If A and B are mutually exclusive and both have non-zero probability, they are not independent. Knowing A occurred tells you B definitely did not, which is information, so they are dependent.
Pairwise vs mutual independence
Pairwise independence: every pair of events in a collection is independent.
Mutual independence: every sub-collection (pairs, triples, and so on) satisfies the product rule.
Mutual independence is the stronger condition. Pairwise independence does not guarantee mutual independence.
Tree diagrams
A tree diagram maps out a sequence of events as branches. Each branch carries the (conditional) probability for that step, and you read off joint probabilities by multiplying along a path from root to leaf.
Multiply along branches to get the joint probability of a complete path.
Sum across leaves to get a marginal probability (useful with the law of total probability).
Particularly helpful for two-stage or three-stage problems where the second stage depends on the first.
Bayesian updating example
Suppose you have two hypotheses about a coin: it is fair (P(heads) = 0.5) or biased (P(heads) = 0.8), and your prior gives equal weight to each.
Start with P(fair) = 0.5 and P(biased) = 0.5.
Observe the data (say, three heads in a row).
Compute the likelihood of the data under each hypothesis:
P(HHH | fair) = 0.5³ = 0.125.
P(HHH | biased) = 0.8³ = 0.512.
Apply Bayes' rule to get the posterior:
P(biased | HHH) = (0.512 × 0.5) / [(0.512 × 0.5) + (0.125 × 0.5)] = 0.256 / 0.3185 ≈ 0.804.
After seeing three heads, the probability that the coin is biased jumps from 0.5 to roughly 0.8. Each new observation shifts the posterior further.
When a system has multiple components that function independently, you can compute overall reliability using the multiplication rule and complements.
Series system (all components must work): P(system works) = P(A works) × P(B works) × ... Each component failing independently means you multiply their individual reliabilities.
Parallel system (at least one component must work): P(system works) = 1 − P(all fail) = 1 − P(A fails) × P(B fails) × ... The complement approach is the clean way to handle this.
Example: a circuit with two independent switches in parallel, each with 0.9 reliability. P(circuit works) = 1 − (0.1)(0.1) = 0.99.
Complement rule
P(E′) = 1 − P(E)
Addition rule
P(A ∪ B) = P(A) + P(B) − P(A ∩ B)
Inclusion-exclusion (three events)
P(A ∪ B ∪ C) = P(A) + P(B) + P(C) − P(A ∩ B) − P(A ∩ C) − P(B ∩ C) + P(A ∩ B ∩ C)
Conditional probability
P(A | B) = P(A ∩ B) / P(B)
Multiplication rule
P(A ∩ B) = P(A | B) × P(B)
Independence test
P(A ∩ B) = P(A) × P(B)
Law of total probability
P(B) = Σ P(B | Aᵢ) × P(Aᵢ)
Bayes' rule
P(Aᵢ | B) = [P(B | Aᵢ) × P(Aᵢ)] / P(B)
Medical diagnostic testing is the textbook application of Bayes' rule. A test with 99% sensitivity can still produce a low positive predictive value if the disease is rare, because the base rate (prior) matters enormously.
Spam filters use Bayesian updating: each word in an email shifts the probability that the message is spam, combining prior frequencies with new evidence.
Students often assume mutually exclusive events are independent. They are not (unless one has probability zero). Mutual exclusivity means dependence: if A happens, you know B did not.
A common error in Bayes' rule problems is forgetting to compute P(B) via the law of total probability. Students sometimes use only one branch of the partition in the denominator.
Students frequently reverse conditionals without Bayes' rule, treating P(A | B) as if it were P(B | A). These are generally not the same.
"Independent" does not mean "the events have nothing to do with each other in everyday life." It is a precise mathematical condition about probabilities. Two events can seem related and still be independent, or seem unrelated and be dependent.
⚠️ Bayes' rule problems are near-certain exam material. Be able to set up the prior, the likelihoods, compute P(B) via total probability, and divide.
⚠️ Know the distinction between mutually exclusive and independent cold. This is a favourite exam question phrased as "can two events be both mutually exclusive and independent?"
⚠️ Conditional probability questions often test whether you can correctly identify which event is "given." Read the problem carefully to decide what goes in the numerator and what defines the reduced sample space.
⚠️ Tree diagram questions test whether you can multiply along branches and sum across paths. Practise drawing them for two-stage and three-stage problems.
True or false: If P(A) = 0.3 and P(B) = 0.4 and A, B are independent, then P(A ∩ B) = 0.12.
True or false: P(A | B) is always equal to P(B | A).
Fill in the blank: The law of total probability requires that the events in the sum form a ___ of the sample space.
True or false: If A and B are mutually exclusive with P(A) = 0.2 and P(B) = 0.3, they are independent.
Fill in the blank: In Bayes' rule, the denominator P(B) serves as a ___ constant.
Answers: 1. True (0.3 × 0.4 = 0.12). 2. False (they are generally different). 3. Partition. 4. False (mutually exclusive events with non-zero probabilities are dependent). 5. Normalising.
Q: A disease affects 1% of a population. A test has 95% sensitivity (P(positive | disease) = 0.95) and 90% specificity (P(negative | no disease) = 0.90). What is P(disease | positive)?
A: P(positive) = P(pos | disease) × P(disease) + P(pos | no disease) × P(no disease) = (0.95)(0.01) + (0.10)(0.99) = 0.0095 + 0.099 = 0.1085. Then P(disease | positive) = (0.95 × 0.01) / 0.1085 ≈ 0.0876, or about 8.8%. Even with a good test, a rare disease means most positives are false positives.
Q: Events A and B satisfy P(A) = 0.5, P(B) = 0.4, P(A ∩ B) = 0.2. Are A and B independent?
A: Check: P(A) × P(B) = 0.5 × 0.4 = 0.2 = P(A ∩ B). Yes, A and B are independent.
Q: A system has three independent components in series with reliabilities 0.95, 0.90, and 0.85. What is the system reliability?
A: P(system works) = 0.95 × 0.90 × 0.85 = 0.72675, or about 72.7%.
Q: In a two-stage experiment, P(A) = 0.6 and P(B | A) = 0.3. What is P(A ∩ B)?
A: P(A ∩ B) = P(B | A) × P(A) = 0.3 × 0.6 = 0.18.
Q: Can two events with P(A) > 0 and P(B) > 0 be both mutually exclusive and independent? Explain.
A: No. Mutually exclusive means P(A ∩ B) = 0. Independence requires P(A ∩ B) = P(A) × P(B), which is positive when both probabilities are positive. Zero cannot equal a positive number, so the two conditions are incompatible.
Conditional probability and the multiplication rule lead directly into discrete and continuous random variables, where you compute probabilities of outcomes using probability mass functions and density functions.
Bayes' rule is the foundation of Bayesian statistics, which reappears in machine learning (naive Bayes classifiers, for instance) and in any field that updates beliefs with data.
Independence is the key assumption behind many statistical tests and models. When you later study the binomial distribution, the assumption that each trial is independent is what makes the formula work.
conditional probability, multiplication rule, general multiplication rule, addition rule, complement rule, inclusion-exclusion principle, Bayes' rule, Bayes' theorem, Bayesian updating, law of total probability, total probability theorem, independence, independent events, mutually exclusive vs independent, disjoint vs independent, pairwise independence, mutual independence, prior probability, posterior probability, likelihood, partition, tree diagram, probability tree, system reliability, series system, parallel system, STAT 101, introduction to statistics, conditional given, P(A given B), P(A|B)