Difficulty: Intermediate | Prerequisites: Descriptive Statistics and Data Fundamentals notes
Big picture: Once you can describe a dataset, the next step is quantifying uncertainty. Probability gives you the language and the rules for doing that. This unit covers how to assign probabilities to events, how to combine them with addition and multiplication rules, and how to work with discrete random variables (binomial and Poisson). You will use these tools constantly in hypothesis testing later in the course, so the mechanics need to become automatic.
Probability measures how likely an event is, on a scale from 0 to 1. You combine probabilities using addition rules (for "or" questions) and multiplication rules (for "and" questions). A random variable assigns a number to each outcome; when its possible values are countable, it is discrete, and you describe it with a probability mass function. The binomial and Poisson distributions are the two discrete models you need for this exam.
Sample space (S or Ω)
The set of all possible outcomes of an experiment. Think of it as the complete menu of things that could happen.
Subjective probability
A probability based on personal judgement or experience, not on data or equally likely outcomes. In simple terms, it is an educated guess.
Empirical probability
P(A) = (number of times A occurs) / (total number of trials). You run the experiment many times and count. Also called relative frequency probability.
Theoretical probability
P(A) = (number of outcomes in A) / (number of outcomes in S). This assumes all outcomes are equally likely.
Conditional probability
P(A | B) = P(A ∩ B) / P(B). The probability of A given that B has already occurred. Think of it as narrowing your sample space to only those outcomes where B is true, then asking how often A also happens.
Independence
Two events A and B are independent if P(A | B) = P(A), meaning knowing B happened tells you nothing new about A. Equivalently, P(A ∩ B) = P(A) × P(B).
Random variable
A function that assigns a unique numerical value to each outcome in the sample space. Listed as capital letters (X, Y). Can be discrete (countable values: 0, 1, 2, ...) or continuous (any value in an interval).
Probability mass function (PMF)
For a discrete random variable, p(x) = P(X = x). It gives the probability of each specific value. All probabilities must be between 0 and 1, and they must sum to 1.
Binomial distribution
Models the number of successes in n independent trials, each with the same probability of success p. Written as X ~ Bin(n, p).
Poisson distribution
Models the count of events occurring in a fixed interval of time or space, where events happen independently at a constant average rate λ. Written as X ~ Poisson(λ).
DeMorgan's Laws
P(A ∪ B)' = P(A' ∩ B') and P(A ∩ B)' = P(A' ∪ B'). The complement of a union is the intersection of the complements, and vice versa. These are useful for rewriting complex probability expressions into simpler forms.
Bayes' Rule
P(A | B) = [P(B | A) × P(A)] / [P(B | A) × P(A) + P(B | A') × P(A')]. It lets you reverse the direction of a conditional probability. In simple terms, if you know how likely the evidence is given a hypothesis, Bayes' rule tells you how likely the hypothesis is given the evidence.
Subjective: personal belief, not derived from data. Useful but not testable.
Empirical: based on observed frequencies from repeated trials.
Theoretical: based on a model where all outcomes are equally likely (a fair coin, a balanced die).
P(A ∪ B) = P(A) + P(B) – P(A ∩ B)
You subtract P(A ∩ B) to avoid double-counting outcomes that belong to both events.
If A and B cannot happen at the same time, P(A ∩ B) = 0, so:
P(A ∪ B) = P(A) + P(B)
P(A | B) = P(A ∩ B) / P(B)
Rearranging: P(A ∩ B) = P(A | B) × P(B)
P(A ∩ B) = P(A | B) × P(B)
For three events: P(A ∩ B ∩ C) = P(A) × P(B | A) × P(C | A ∩ B)
P(A | B) = P(A), which means P(A ∩ B) = P(A) × P(B).
Independence is a property you check or assume; it does not hold by default.
P(A | B) = [P(B | A) × P(A)] / [P(B | A) × P(A) + P(B | A') × P(A')]
The denominator is just P(B) expanded using the law of total probability.
(A ∪ B)' = A' ∩ B' (not in either = not in A AND not in B)
(A ∩ B)' = A' ∪ B' (not in both = not in A OR not in B)
These work both ways (the "vice versa" note in the cheat sheet).
E(X) = μ_X = Σ x · p(x)
This is a weighted average, where each value is weighted by its probability.
E(a + bX) = a + b · μ_X (linear transformation)
E(X + Y) = μ_X + μ_Y (always true, independence not required)
E(g(X)) = Σ g(xᵢ) · pᵢ
You apply the function to each value first, then take the weighted average.
Var(X) = E[(X – μ_X)²] = Σ(xᵢ – μ_X)² · pᵢ
Shortcut: Var(X) = E(X²) – [E(X)]²
Var(a + bX) = b² · σ²_X (constants added drop out; constants multiplied get squared)
If X and Y are independent: Var(X + Y) = σ²_X + σ²_Y
If X and Y are independent: Var(X – Y) = σ²_X + σ²_Y (note: still addition)
General case (correlated): σ²_{X±Y} = σ²_X + σ²_Y ± 2ρ · σ_X · σ_Y, where ρ is the correlation
Conditions: n independent trials, each with exactly two outcomes (success/failure), same probability p on every trial.
P(X = x) = C(n, x) · pˣ · (1 – p)^(n–x), where C(n, x) = n! / [x!(n–x)!]
Mean: E(X) = np
Standard deviation: SD = √[np(1 – p)]
Skewness cue: p < 0.5 gives a right-skewed distribution; p = 0.5 is symmetric; p > 0.5 is left-skewed.
Conditions: counts of events in a fixed interval, events occur independently at a constant average rate λ.
P(X = x) = (e^(–λ) · λˣ) / x!, for x = 0, 1, 2, ...
Mean: μ = λ
Variance: σ² = λ (mean and variance are equal, which is a defining feature)
Standard deviation: σ = √λ
If the rate is given for one interval but you need a different interval, scale λ proportionally. For example, if λ = 5 calls/hour and you want a 2-hour window, use λ = 10.
Item | Formula |
|---|---|
Addition rule | P(A ∪ B) = P(A) + P(B) – P(A ∩ B) |
Conditional probability | P(A | B) = P(A ∩ B) / P(B) |
Bayes' rule | P(A | B) = P(B | A)P(A) / [P(B | A)P(A) + P(B | A')P(A')] |
Binomial PMF | P(X = x) = C(n,x) pˣ (1–p)^(n–x) |
Binomial mean | np |
Binomial SD | √[np(1–p)] |
Poisson PMF | P(X = x) = e^(–λ) λˣ / x! |
Poisson mean & variance | Both equal λ |
Variance shortcut | Var(X) = E(X²) – [E(X)]² |
The binomial distribution is the model behind quality-control sampling: pull 50 items off the line and count how many are defective. The Poisson distribution models arrivals, such as the number of customers who walk into a shop per hour or the number of server errors per day. If you have ever seen a call centre staffing model, it almost certainly uses Poisson assumptions underneath.
"Independent and mutually exclusive mean the same thing." They do not. Mutually exclusive events cannot happen together (P(A ∩ B) = 0). Independent events can happen together; knowing one occurred just does not change the probability of the other. In fact, if two events with nonzero probability are mutually exclusive, they cannot be independent.
"Var(X – Y) = Var(X) – Var(Y)." Variance always adds, whether you are adding or subtracting the random variables (assuming independence). The formula is Var(X – Y) = Var(X) + Var(Y).
"I can use the binomial distribution whenever there are two outcomes." You also need a fixed number of independent trials with the same probability on each trial. If the probability changes from trial to trial, the binomial does not apply.
"For the Poisson distribution, mean and standard deviation are both λ." The mean is λ and the variance is λ, but the standard deviation is √λ. Do not confuse variance with standard deviation here.
⚠️ Be able to identify whether a problem is binomial or Poisson from the wording alone. Key cue: "number of successes in n trials" signals binomial; "count of events in an interval" signals Poisson.
⚠️ Bayes' rule problems almost always appear on the exam. Set up the numerator and denominator carefully; most errors come from getting the denominator wrong.
⚠️ Know when to add and when to subtract P(A ∩ B) in the addition rule.
⚠️ The variance shortcut E(X²) – [E(X)]² is faster than computing from the definition. Practise it.
⚠️ For Poisson problems, check whether the interval in the question matches the interval for the given λ. Rescale if needed.
True or false: P(A ∪ B) = P(A) + P(B) is always valid.
A: False. It is only valid when A and B are mutually exclusive (disjoint). Otherwise you must subtract P(A ∩ B).
Fill in the blank: For a binomial distribution with n = 20 and p = 0.3, the expected number of successes is ___.
A: 6 (since E(X) = np = 20 × 0.3)
True or false: If X ~ Poisson(4), then Var(X) = 4.
A: True. For a Poisson distribution, variance equals λ.
Fill in the blank: DeMorgan's law says (A ∪ B)' = ___.
A: A' ∩ B'
Q: Events A and B are independent with P(A) = 0.4 and P(B) = 0.3. What is P(A ∩ B)?
A: P(A ∩ B) = P(A) × P(B) = 0.4 × 0.3 = 0.12.
Q: A fair coin is flipped 10 times. What is the probability of getting exactly 7 heads?
A: X ~ Bin(10, 0.5). P(X = 7) = C(10,7) × 0.5⁷ × 0.5³ = 120 × (1/1024) ≈ 0.1172.
Q: A call centre receives an average of 3 calls per minute. What is the probability of receiving exactly 5 calls in one minute?
A: X ~ Poisson(3). P(X = 5) = e^(–3) × 3⁵ / 5! = e^(–3) × 243 / 120 ≈ 0.1008.
Q: You know P(B | A) = 0.7, P(A) = 0.4, and P(B | A') = 0.2. Find P(A | B).
A: By Bayes' rule: P(A | B) = (0.7 × 0.4) / (0.7 × 0.4 + 0.2 × 0.6) = 0.28 / (0.28 + 0.12) = 0.28 / 0.40 = 0.70.
Q: Var(X) = 9, Var(Y) = 16, and X and Y are independent. What is Var(2X – 3Y)?
A: Var(2X – 3Y) = 4 × Var(X) + 9 × Var(Y) = 4(9) + 9(16) = 36 + 144 = 180. (Constants are squared; variances add even when subtracting the variables.)
Probability rules are the machinery behind every hypothesis test and confidence interval you will encounter later. The binomial distribution connects to the normal distribution through the Central Limit Theorem: when n is large enough, Bin(n, p) is well-approximated by a normal distribution. The Poisson distribution also has a normal approximation for large λ. These bridges matter in later chapters.
probability rules, addition rule, multiplication rule, conditional probability, Bayes theorem, Bayes rule, independence, mutually exclusive, disjoint events, DeMorgan's law, complement rule, sample space, random variable, PMF, probability mass function, expected value, variance rules, binomial distribution, Poisson distribution, discrete random variable, counting rule, combination, Purdue STAT 301, intro to statistics midterm 1