Difficulty: Intermediate | Prerequisites: Chapters 1-3 study notes (descriptive statistics, z-scores, empirical rule)
This block introduces the mathematical machinery that sits beneath statistical inference. Chapter 4 covers probability rules, the language of events and the tools for combining them (addition rule, multiplication rule, conditional probability, Bayes' Rule). Chapter 5 introduces discrete random variables, the binomial and Poisson distributions, and expected value calculations. Chapter 6 moves to continuous distributions: the uniform, exponential and, most importantly, the normal distribution and how to use the z-table. Everything here is prerequisite for sampling distributions and inference in Chapters 7-12.
Probability is the study of randomness and gives you the rules for quantifying how likely events are. Discrete random variables (binomial, Poisson) model counts; continuous random variables (uniform, exponential, normal) model measurements. The normal distribution is the single most important distribution in this course: learning to convert to z-scores and read the z-table is essential.
Experiment (random)
An activity with at least two possible outcomes whose result cannot be predicted with certainty.
Sample space (S or Ω)
The complete listing of all possible outcomes of an experiment.
Event
Any collection of outcomes from an experiment, denoted by a capital Latin letter. A simple event contains exactly one outcome.
Complement (A')
All outcomes in S that are not in A. P(A') = 1 - P(A).
Union (A ∪ B)
All outcomes in A or B or both. Read as "A or B."
Intersection (A ∩ B)
All outcomes in both A and B simultaneously. Read as "A and B."
Disjoint (mutually exclusive)
Two events with no outcomes in common: A ∩ B = {}. In simple terms, they cannot both happen at the same time.
Conditional probability P(A|B)
The probability of A occurring given that B has occurred: P(A|B) = P(A ∩ B) / P(B). Think of it as narrowing your universe to only the cases where B happened, then asking how often A also happened.
Independence
Two events are independent if knowing one occurred does not change the probability of the other: P(A|B) = P(A), or equivalently P(A ∩ B) = P(A) × P(B).
Bayes' Rule
A formula for reversing the direction of a conditional probability: P(A|B) = P(B|A) × P(A) / P(B). Use it when you know P(B|A) but need P(A|B).
Random variable (r.v.)
A function that assigns a numerical value to each outcome in a sample space. Discrete r.v.s take countable values; continuous r.v.s take values in one or more intervals.
Probability mass function (pmf)
For a discrete r.v., the function p(x) = P(X = x) giving the probability of each value. All probabilities must be between 0 and 1 and must sum to 1.
Expected value E(X)
The long-run average value of a random variable: E(X) = Σ xᵢ pᵢ for discrete variables. Think of it as the theoretical mean.
Variance of a random variable Var(X)
A measure of spread: Var(X) = E(X²) - [E(X)]². The standard deviation is σ = √Var(X).
Binomial distribution X ~ B(n, p)
Models the number of successes in n identical, independent trials, each with success probability p. Must satisfy BInS: Binary outcomes, Independent trials, fixed n, same probability of Success.
Poisson distribution X ~ Poisson(λ)
Models the count of events in a fixed interval of time, length or area. The parameter λ is both the mean and the variance.
Probability density function (pdf)
For a continuous r.v., the function f(x) such that the area under the curve between two values gives the probability. f(x) itself is not a probability.
Cumulative distribution function (cdf)
F(x) = P(X ≤ x). For continuous distributions, this is the integral of the pdf from -∞ to x.
Uniform distribution
A continuous distribution where every value between a and b is equally likely. pdf = 1/(b - a) for a < x < b. Mean = (a + b)/2.
Exponential distribution
Models time between events. pdf = λe^(-λx) for x ≥ 0. Mean = 1/λ. The exponential is the "inverse" of the Poisson: Poisson counts events per interval, exponential measures time between events.
Normal distribution X ~ N(μ, σ²)
The bell curve. Symmetric, unimodal, completely described by μ and σ². The standard normal has μ = 0 and σ² = 1.
Standard normal (Z)
A normal distribution with mean 0 and variance 1. Any normal variable X can be converted to Z via z = (x - μ) / σ.
Types of probability:
Subjective: based on personal judgement, used when the event can only occur once or a few times (e.g. predicting election outcomes)
Empirical: determined from data using the frequentist approach with a large sample
Theoretical (equally likely): calculated without experiments, by dividing favourable outcomes by total outcomes
Properties of probability:
0 ≤ P(A) ≤ 1 for any event A
P(S) = 1 (something must happen)
P({}) = 0 (the impossible event has probability zero)
P(A) = sum of probabilities of all outcomes in A
Key rules:
Complement rule: P(A') = 1 - P(A)
General addition rule: P(A ∪ B) = P(A) + P(B) - P(A ∩ B)
Addition rule for disjoint events: P(A ∪ B) = P(A) + P(B)
General multiplication rule: P(A ∩ B) = P(A) × P(B|A)
If independent: P(A ∩ B) = P(A) × P(B)
Conditional probability and Bayes' Rule:
P(A|B) = P(A ∩ B) / P(B)
Keywords for conditional probability: "given", "assume", "suppose", "if"
Keyword for intersection: "and"
Use Bayes' Rule when you know the conditional probability in one direction but need it in the other direction. Do not use it when you already know the intersection directly
P(A|B) = P(B|A) × P(A) / [P(B|A) × P(A) + P(B|A') × P(A')]
Disjoint vs. independent:
Disjoint means no common outcomes. Can be shown with Venn diagrams
Independent means one event does not affect the other's probability. Can only be verified mathematically
These concepts are not related. In fact, if two events with non-zero probabilities are disjoint, they usually cannot be independent
Rules for expected value:
Rule 1: E(a + bX) = a + b × E(X)
Rule 2: E(X ± Y) = E(X) ± E(Y)
Rule 3: E(g(X)) = Σ g(xᵢ) pᵢ
Rules for variance:
Rule 1: Var(a + bX) = b² × Var(X)
Rule 2 (independent): Var(X ± Y) = Var(X) + Var(Y)
Note: variances always add, even when subtracting random variables
Standard deviations cannot be added or subtracted directly
Rule 3 (correlated): Var(X ± Y) = Var(X) + Var(Y) ± 2ρσ_X σ_Y
Variance is never negative. Always calculate variance first, then take the square root for standard deviation
Binomial distribution X ~ B(n, p):
P(X = x) = C(n, x) × p^x × (1 - p)^(n - x), for x = 0, 1, ..., n
Mean: E(X) = np
Variance: Var(X) = np(1 - p)
Standard deviation: σ = √(np(1 - p))
If p < 0.5 the distribution is right-skewed. If p = 0.5 it is symmetric. If p > 0.5 it is left-skewed
"At least one" = complement of zero: P(X ≥ 1) = 1 - P(X = 0)
Poisson distribution X ~ Poisson(λ):
P(X = x) = e^(-λ) × λ^x / x!, for x = 0, 1, 2, ...
Mean = variance = λ, standard deviation = √λ
If the units of λ do not match the question's time frame, scale accordingly: λ' = (rate)(new interval)
The Poisson counts events per interval; the exponential measures time between events
General continuous distribution properties:
f(x) ≥ 0 everywhere
The total area under the pdf equals 1
P(X = a) = 0 for any single point a (probability is area, and a single point has zero width)
For continuous distributions, P(X ≤ x) and P(X < x) are the same
Uniform distribution X ~ Uniform(a, b):
pdf: f(x) = 1/(b - a) for a < x < b
Mean: (a + b) / 2
Variance: (b - a)² / 12
Standard deviation: (b - a) / √12
Exponential distribution:
pdf: f(x) = λe^(-λx) for x ≥ 0
cdf: F(x) = 1 - e^(-λx)
Mean = standard deviation = 1/λ
Variance = 1/λ²
Median: set F(x) = 0.5 and solve
Normal distribution X ~ N(μ, σ²):
Bell-shaped, symmetric, unimodal
Mean = median (because of symmetry)
Inflection points at μ ± σ
To find probabilities: standardise to Z using z = (x - μ) / σ, then use the z-table
The z-table gives P(Z ≤ z), so:
P(Z > z) = 1 - P(Z ≤ z)
P(a < Z < b) = P(Z < b) - P(Z < a)
By symmetry: P(Z ≥ z) = P(Z ≤ -z)
To find a value from a probability (percentile): look up the probability in the body of the z-table, read off the z-value, then convert back: x = μ + σz
Do not interpolate; choose the closest table value
Checking normality (five methods):
Histogram: should look symmetric and bell-shaped
Backward empirical rule: check if proportions match 68/95/99.7
IQR/s ratio: should be close to 1.4 for normal data
Normal probability plot (QQ plot): data should fall roughly on a straight line
Shapiro-Wilk test: a formal inference test for normality
Quantity | Formula |
|---|---|
Complement rule | P(A') = 1 - P(A) |
General addition rule | P(A ∪ B) = P(A) + P(B) - P(A ∩ B) |
Conditional probability | P(A|B) = P(A ∩ B) / P(B) |
Multiplication rule | P(A ∩ B) = P(A) × P(B|A) |
Bayes' Rule | P(A|B) = P(B|A)P(A) / [P(B|A)P(A) + P(B|A')P(A')] |
Binomial P(X = x) | C(n,x) p^x (1-p)^(n-x) |
Binomial mean | np |
Binomial variance | np(1-p) |
Poisson P(X = x) | e^(-λ) λ^x / x! |
Poisson mean/variance | λ |
Uniform mean | (a+b)/2 |
Uniform variance | (b-a)²/12 |
Exponential mean | 1/λ |
Exponential cdf | 1 - e^(-λx) |
Standard normal conversion | z = (x - μ) / σ |
Back-conversion | x = μ + σz |
qnorm(p) – z-value for a given cumulative probability p
pnorm(z) – cumulative probability P(Z ≤ z)
dbinom(x, n, p) – P(X = x) for binomial
pbinom(x, n, p) – P(X ≤ x) for binomial
dpois(x, lambda) – P(X = x) for Poisson
ppois(x, lambda) – P(X ≤ x) for Poisson
Bayes' Rule is the foundation of medical diagnostic testing. A test with 98% sensitivity and 95% specificity sounds almost perfect, but if the disease prevalence is only 1%, a positive result means you actually have only about a 16.5% chance of having the disease. This counter-intuitive result is why screening tests produce so many false positives. The binomial distribution models quality control (what fraction of items are defective), and the Poisson models rare events per interval (customer arrivals, system failures, radioactive decay).
Students often confuse disjoint and independent. They are completely separate concepts. If two events with non-zero probability are disjoint, they cannot be independent, because knowing one occurred tells you the other definitely did not.
Students frequently forget that P(A|B) ≠ P(B|A). The conditional probability changes direction depending on what is being conditioned on.
When computing variance of a difference, students subtract variances instead of adding them. Var(X - Y) = Var(X) + Var(Y) when X and Y are independent. Variances always add.
Students sometimes treat the pdf value f(x) as a probability. It is not. Only the area under the curve between two points gives a probability.
⚠️ "At least one" problems almost always use the complement: P(X ≥ 1) = 1 - P(X = 0).
⚠️ Check BInS before using the binomial distribution. If any condition fails, it is not binomial.
⚠️ For the normal distribution, the z-table gives P(Z ≤ z). All other probabilities require manipulation (subtraction, complement).
⚠️ Bayes' Rule is commonly tested using tree diagrams. Draw the tree, label branches with conditional probabilities, multiply along paths.
⚠️ When scaling the Poisson, make sure units match: λ' = (rate per original interval) × (new interval length).
⚠️ Var(X - Y) = Var(X) + Var(Y) for independent variables. The minus does not carry through.
True or false: If A and B are disjoint, then P(A ∪ B) = P(A) × P(B). False. For disjoint events, P(A ∪ B) = P(A) + P(B). The multiplication rule P(A) × P(B) applies to independent events, not disjoint ones.
Fill in the blank: For a binomial distribution, the mean is ___ and the variance is ___. Mean = np, Variance = np(1 - p).
True or false: For a continuous random variable, P(X = 3) = 0. True. A single point has zero width, so the area under the curve is zero.
Fill in the blank: The exponential distribution with λ = 0.5 has a mean of ___. 1/0.5 = 2.
True or false: The Poisson variance equals its mean. True. Both equal λ.
Q: A medical test has 95% sensitivity (P(+|D) = 0.95) and 90% specificity (P(-|D') = 0.90). The disease prevalence is 2%. What is the probability a person who tests positive actually has the disease?
A: Using Bayes' Rule: P(D|+) = P(+|D)P(D) / [P(+|D)P(D) + P(+|D')P(D')] = (0.95)(0.02) / [(0.95)(0.02) + (0.10)(0.98)] = 0.019 / (0.019 + 0.098) = 0.019 / 0.117 ≈ 0.162, so about 16.2%.
Q: X ~ B(20, 0.3). Find the mean and standard deviation.
A: Mean = np = 20(0.3) = 6. Variance = np(1-p) = 20(0.3)(0.7) = 4.2. Standard deviation = √4.2 ≈ 2.049.
Q: An IT helpdesk receives an average of 4 calls per hour (Poisson). What is the probability of receiving exactly 2 calls in the next hour?
A: P(X = 2) = e^(-4) × 4² / 2! = e^(-4) × 16 / 2 = e^(-4) × 8 ≈ 0.0183 × 8 ≈ 0.1465.
Q: X ~ N(100, σ² = 225). What is the probability that X is between 85 and 115?
A: σ = 15. z₁ = (85 - 100)/15 = -1.0, z₂ = (115 - 100)/15 = 1.0. P(-1.0 < Z < 1.0) = P(Z < 1.0) - P(Z < -1.0) = 0.8413 - 0.1587 = 0.6826.
The normal distribution is the centrepiece of Chapter 7 (sampling distributions) via the Central Limit Theorem. Conditional probability and Bayes' Rule appear again in regression diagnostics. The binomial distribution is the basis for proportion tests in Chapters 8-10. The Poisson and exponential distributions appear in reliability engineering and queueing theory.
probability, sample space, event, complement, union, intersection, disjoint, mutually exclusive, conditional probability, independence, Bayes' Rule, Bayes theorem, random variable, pmf, probability mass function, expected value, variance, standard deviation, binomial distribution, BInS, Poisson distribution, pdf, probability density function, cdf, cumulative distribution function, uniform distribution, exponential distribution, normal distribution, bell curve, z-score, z-table, standard normal, percentile, STAT 350, Purdue, introductory statistics