Source: STAT 350 Exam 1, Purdue University
Tags: probability, independence, conditional probability, Bayes' theorem, Bayes rule, law of total probability, tree diagram, mutually exclusive, joint probability, sample space
Difficulty: Intermediate | Prerequisites: Set notation (union, intersection, complement), basic probability axioms.
Probability is the language that connects data to inference. This set of notes covers how to tell whether events are independent, how to compute probabilities when they are not (conditional probability), and how to reverse the conditioning direction using Bayes' theorem. These tools appear in almost every applied statistics problem: medical testing, quality control, machine learning classifiers, and (on this exam) training a cat to high-five. If you are coming in cold, make sure you are comfortable with the idea that P(A ∩ B) means "the probability that both A and B happen" and that conditional probability P(A|B) means "the probability of A, given that B has already occurred."
Independence means knowing one event tells you nothing about the other. When events are not independent, you need conditional probabilities. Bayes' theorem lets you flip the direction of a conditional: if you know P(B|A), you can find P(A|B) by combining it with the law of total probability.
Independence (of events)
Two events A and B are independent if and only if P(A ∩ B) = P(A) · P(B) for every pair of outcomes, or equivalently P(A|B) = P(A). Think of it as: learning that B happened does not change your belief about A.
Independence (of random variables)
Two discrete random variables X and Y are independent if and only if P(X = x, Y = y) = P(X = x) · P(Y = y) for every combination (x, y) in their joint support. Checking a single pair is not enough; the condition must hold for all pairs.
Conditional probability
P(A|B) = P(A ∩ B) / P(B), defined when P(B) > 0. In simple terms, it is the probability of A in the reduced universe where B is known to have happened.
Law of total probability
If B₁, B₂, ..., Bₙ partition the sample space, then P(A) = Σ P(A|Bᵢ) · P(Bᵢ). Think of it as computing an overall probability by weighing each scenario by how likely that scenario is.
Bayes' theorem
P(Bⱼ|A) = P(A|Bⱼ) · P(Bⱼ) / P(A). This flips the conditioning direction. The denominator is usually expanded using the law of total probability.
Mutually exclusive (disjoint)
Two events are mutually exclusive if P(A ∩ B) = 0, meaning they cannot both happen. Mutually exclusive events with positive probability are never independent (a common source of confusion).
Tree diagram
A visual tool for organising conditional probabilities. The first set of branches represents the initial partition (e.g. treat offered vs. not offered), and the second set represents the outcomes conditional on each branch. Multiply along branches, add across branches.
Two events: verify P(A ∩ B) = P(A) · P(B), or equivalently P(A|B) = P(A).
If P(A|B) ≠ P(A|Bᶜ), the events are dependent. This is often the quickest check when conditional probabilities are given directly.
Exam example (Q4a): P(High-five | Treat) = 0.8 ≠ P(High-five | No treat) = 0.35, so "Treat offered" and "High-five" are not independent.
For discrete random variables, the joint PMF must factorise into the product of the marginals for every (x, y) pair.
Checking a single cell (e.g. P(X=1, Y=1) = P(X=1)·P(Y=1)) is necessary but not sufficient. Independence requires the equality to hold everywhere.
Exam reference (Q1.1): The statement "if P(X=1,Y=1) = P(X=1)·P(Y=1), then X and Y are independent" is false, because one matching cell does not guarantee all cells match.
P(A ∩ B ∩ C) = 0 does not force P(A ∩ C) = 0 or P(B ∩ C) = 0.
Counterexample: A and C could overlap outside of B. The three-way intersection being empty only means all three cannot happen simultaneously.
Exam reference (Q1.2): The statement is false.
Step 1: Draw the first branches for the partition. Label each with its marginal probability.
Step 2: For each first-level branch, draw second-level branches for each possible outcome. Label with conditional probabilities.
Step 3: The probability of any leaf is the product of the probabilities along the path to it.
Exam example (Q4b):
Branch A (Ignore | Treat) = 1 - 0.8 = 0.2
Branch B (Ignore | No treat) = 0.6
Branch C (Nag | No treat) = 1 - 0.35 - 0.6 = 0.05
P(Nag) = P(Nag | Treat) · P(Treat) + P(Nag | No treat) · P(No treat)
= 0 · 0.7 + 0.05 · 0.3 = 0.015
P(High-five) = P(High-five | Treat) · P(Treat) + P(High-five | No treat) · P(No treat)
= 0.8 · 0.7 + 0.35 · 0.3 = 0.665
Goal: find P(No treat | High-five).
Apply Bayes':
P(No treat | High-five) = P(High-five | No treat) · P(No treat) / P(High-five)
= (0.35 · 0.3) / 0.665
= 0.105 / 0.665
≈ 0.1579
Since 0.1579 < 0.5, the condition for restarting training is not met.
Item | Formula |
|---|---|
Conditional probability | P(A|B) = P(A ∩ B) / P(B) |
Multiplication rule | P(A ∩ B) = P(A|B) · P(B) |
Independence test | P(A ∩ B) = P(A) · P(B) |
Law of total probability | P(A) = Σ P(A|Bᵢ) · P(Bᵢ) |
Bayes' theorem | P(B|A) = P(A|B) · P(B) / P(A) |
Bayes' theorem is the engine behind medical diagnostic testing. When a screening test comes back positive, doctors use the test's sensitivity (true positive rate), the specificity (true negative rate), and the prevalence (base rate of disease) to compute the probability that the patient truly has the condition. The same logic applies to spam filters, fraud detection, and any situation where you observe an outcome and want to reason backward to its cause.
Students often think that checking independence at one point (one pair of values) is sufficient for random variables. Independence requires the factorisation to hold at every point in the joint support.
Mutually exclusive events are frequently confused with independent events. If A and B are mutually exclusive and both have positive probability, they are dependent: knowing A occurred tells you B did not.
P(A ∩ B ∩ C) = 0 does not imply all pairwise intersections are zero. Two of the three events can still overlap.
Forgetting to use the law of total probability for the denominator in Bayes' theorem. The denominator P(A) must account for every way A can happen, not just the branch you care about.
⚠️ Q1.1 tested whether a single-cell match proves independence of random variables (it does not, answer is false).
⚠️ Q1.2 tested whether P(A ∩ B ∩ C) = 0 forces pairwise intersections to be zero (it does not, answer is false).
⚠️ Problem 4 was worth 27 points and required building a tree diagram, using the law of total probability twice, and applying Bayes' theorem. Showing all multiplication and addition steps is essential for full credit.
⚠️ The exam required a final interpretation step: compare the Bayes result to a threshold (0.5) and state a conclusion in context.
True or false: If P(A|B) = P(A), then A and B are independent.
Fill in the blank: Bayes' theorem reverses the direction of ___ probability.
True or false: If events A and B are mutually exclusive with P(A) > 0 and P(B) > 0, they are independent.
Fill in the blank: The law of total probability computes P(A) by summing across a ___ of the sample space.
True or false: P(A ∩ B ∩ C) = 0 implies P(A ∩ C) = 0.
Answers: 1. True. 2. Conditional. 3. False (they are dependent). 4. Partition. 5. False.
Q: A factory has two machines. Machine 1 produces 60% of items and has a 5% defect rate. Machine 2 produces 40% of items and has a 8% defect rate. What is the probability that a randomly selected item is defective?
A: By the law of total probability: P(Defective) = 0.05 · 0.6 + 0.08 · 0.4 = 0.03 + 0.032 = 0.062.
Q: In the factory problem above, given that an item is defective, what is the probability it came from Machine 2?
A: By Bayes' theorem: P(M2 | Defective) = (0.08 · 0.4) / 0.062 = 0.032 / 0.062 ≈ 0.5161.
Q: Events A and B satisfy P(A) = 0.4, P(B) = 0.3, P(A ∩ B) = 0.12. Are A and B independent?
A: Check: P(A) · P(B) = 0.4 · 0.3 = 0.12 = P(A ∩ B). Yes, they are independent.
Q: Heekyung offers a treat 70% of the time. P(High-five | Treat) = 0.8, P(High-five | No treat) = 0.35. Find P(No treat | High-five).
A: P(High-five) = 0.8(0.7) + 0.35(0.3) = 0.665. By Bayes': P(No treat | High-five) = 0.35(0.3) / 0.665 ≈ 0.1579.
Conditional probability is the foundation for understanding conditional distributions of random variables, which appears in the discrete random variables notes (the Binomial-Poisson conditional problem from Q2.5). Bayes' theorem will return when you study hypothesis testing later in the course, where the "prior" and "posterior" language comes from Bayesian reasoning. The tree diagram technique also connects to the law of total expectation, which you may encounter in more advanced probability courses.
probability, conditional probability, Bayes' theorem, Bayes rule, independence, independent events, independent random variables, joint probability, marginal probability, law of total probability, tree diagram, partition, mutually exclusive, disjoint, sample space, STAT 350 Purdue, Exam 1