Difficulty: Introductory to Intermediate | Prerequisites: basic probability (Ch. 1–4), summation notation, comfort with simple algebra.
This is where the course shifts from describing data you already have to modelling outcomes you have not yet observed. Random variables give you a formal way to attach numbers to the results of an experiment, and probability mass functions tell you how likely each of those numbers is. Once you can do that, you can compute an expected value (the long-run average) and a variance (how spread out the outcomes are). These tools underpin every distribution you will meet for the rest of the course: binomial, Poisson, normal, and beyond.
A random variable is a function that turns each outcome of an experiment into a number. A probability mass function lists every possible value of a discrete random variable alongside its probability. From the PMF you can calculate the mean (expected value) and variance, and there are shortcut rules for linear transformations and combinations of random variables.
Random variable (RV)
A function that assigns a unique numerical value to each outcome in a sample space. Think of it as a rule that converts experimental results into numbers you can do maths with.
Discrete random variable
A random variable whose possible values are countable (finite or countably infinite). In simple terms, you could list them out, even if the list is very long.
Continuous random variable
A random variable that can take any value within an interval. You will work with these in later chapters; for now, know the distinction.
Probability mass function (PMF)
A function p(x) that gives the probability of each value of a discrete random variable. Two rules govern it: every individual probability lies between 0 and 1 inclusive, and the probabilities across all values sum to exactly 1.
Expected value / mean (E(X), µ)
The long-run average of a random variable, calculated as the sum of each value multiplied by its probability. Think of it as the centre of gravity of the distribution.
Variance (Var(X), σ²)
A measure of how spread out the values of a random variable are around the mean. Computed as E(X²) minus [E(X)]², or equivalently as the sum of (x – µ)² × p(x) over all values.
Standard deviation (σ)
The square root of the variance. In simple terms, it puts the spread back into the same units as the original variable.
Cumulative distribution function (CDF)
P(X ≤ x), the probability that the random variable takes a value less than or equal to x. Used more extensively with continuous distributions, but the concept applies to discrete ones as well.
A random variable is the bridge between a sample space (outcomes described in words) and numbers you can calculate with.
Discrete RVs have countable outcomes: number of heads in 10 flips, number of defective items in a batch.
Continuous RVs span intervals: time until a light bulb fails, exact weight of a package.
Every probability must satisfy 0 ≤ p(x) ≤ 1.
The sum of all p(x) values equals 1: Σ p(x) = 1.
Example: toss a coin 3 times, let X = number of successes (S). The sample space includes SSS, SSN, SNS, NSS, NSN, NNS, SNN, NNN. From this you build the PMF by counting how many outcomes correspond to each value of X.
E(X) = µ = Σ xᵢ · pᵢ, the sum of each value times its probability.
Worked example: if X takes values 0, 1, 2, 3 with probabilities 0.6, 0.25, 0.1, 0.05, then E(X) = 0(0.6) + 1(0.25) + 2(0.1) + 3(0.05) = 0.6.
Rule 1 (linear transformation): E(a + bX) = a + b · E(X). Shifting and scaling a random variable shifts and scales its mean the same way.
Rule 2 (sum/difference): E(X ± Y) = E(X) ± E(Y). The expected value of a sum is the sum of the expected values, regardless of dependence.
Rule 3 (function of X): E(g(X)) = Σ g(xᵢ) · pᵢ. Apply the function to each value first, then take the weighted average.
Worked example: g(X) = 400 + 100X – 15. E(g(X)) = 400 + 100 · E(X) – 15 = 400 + 100(0.6) – 15 = 445.
Worked example: 2E(X) + 3E(Y) = 2(0.6) + 3(0.95) = 4.05.
Var(X) = σ² = E[(X – µ)²] = Σ (xᵢ – µ)² · pᵢ.
Shortcut formula: Var(X) = E(X²) – [E(X)]². Compute E(X²) first, then subtract the square of the mean. This is almost always the faster route.
Worked example: E(X²) = 0²(0.6) + 1²(0.25) + 2²(0.1) + 3²(0.05) = 1.1. Then Var(X) = 1.1 – (0.6)² = 1.1 – 0.36 = 0.74.
Standard deviation: σ = √Var(X) = √0.74 ≈ 0.86.
Rule 1 (linear transformation): Var(a + bX) = b² · Var(X). The additive constant drops out; only the scaling factor matters, and it gets squared.
Rule 2 (independent sum/difference): Var(X ± Y) = Var(X) + Var(Y), when X and Y are independent. Variances always add, even for differences.
σ(X±Y) = √[σ²(X) + σ²(Y)].
Rule 3 (general, correlated case): Var(X ± Y) = σ²(X) + σ²(Y) + 2ρσ(X)σ(Y), where ρ is the correlation. When independent, ρ = 0 and this reduces to Rule 2.
Worked example: g(X) = 400 + 100X – 15. Var(g(X)) = (100)² · Var(X) = 10 000 × 0.74 = 7400. σ(g(X)) = √7400 ≈ 86.02.
Worked example: σ of 2X + 3Y = √[4σ²(X) + 9σ²(Y)] = √[4(0.74) + 9(0.9975)] = √11.49 ≈ 3.389.
F(x) = P(X ≤ x), the running total of probabilities up to and including x.
Primarily used with continuous distributions, but the concept applies to discrete cases as well.
For a discrete RV, the CDF is a step function that jumps at each possible value.
E(X) = \mu = \sum x_i \, p_iE(a + bX) = a + b \, E(X)E(X \pm Y) = E(X) \pm E(Y)E(g(X)) = \sum g(x_i) \, p_i\text{Var}(X) = \sigma^2 = E(X^2) - [E(X)]^2\text{Var}(a + bX) = b^2 \, \text{Var}(X)\text{Var}(X \pm Y) = \sigma_X^2 + \sigma_Y^2 \quad \text{(when independent)}\sigma_X = \sqrt{\text{Var}(X)}Expected value is how insurance companies set premiums: they multiply each possible payout by its probability and charge enough to cover the long-run average plus a margin. Variance is why a diversified stock portfolio is less risky than a single stock, even if the expected returns are identical: combining independent assets reduces the overall variance.
Students often assume E(X) must be a value that X can actually take. It does not have to be. If a die has faces 1 through 6, E(X) = 3.5, which is not a possible roll.
Confusing sample variance (s², divides by n – 1) with population variance (σ², divides by N or uses the PMF directly). In this chapter you are working with the population/theoretical variance from the PMF.
Forgetting that Var(X – Y) = Var(X) + Var(Y) when independent, not Var(X) – Var(Y). Variances add for both sums and differences.
Applying the linear-transformation variance rule without squaring the coefficient: Var(3X) = 9 Var(X), not 3 Var(X).
⚠️ The shortcut Var(X) = E(X²) – [E(X)]² is the most efficient route to variance on timed exams. Know it cold.
⚠️ Expect at least one question that asks you to compute E(g(X)) or Var(a + bX) using the rules, not from scratch.
⚠️ PMF validity checks (do probabilities sum to 1? are all between 0 and 1?) appear as quick marks. Do not skip the check.
⚠️ Be ready to find a missing probability from the fact that all probabilities must sum to 1.
True or false: the expected value of a discrete random variable must be one of its possible values.
False. E(X) can fall between possible values.
Fill in the blank: if all probabilities in a PMF are valid, they must each be between ___ and ___, and they must sum to ___.
0 and 1; sum to 1.
True or false: Var(X – Y) = Var(X) – Var(Y) when X and Y are independent.
False. Variances add: Var(X – Y) = Var(X) + Var(Y).
Fill in the blank: Var(5X) = ___ Var(X).
True or false: E(X + Y) = E(X) + E(Y) holds only when X and Y are independent.
False. It holds regardless of dependence.
Q: A discrete random variable X has the following PMF: P(0) = 0.4, P(1) = 0.3, P(2) = 0.2, P(3) = ?. What is P(3)?
A: P(3) = 1 – (0.4 + 0.3 + 0.2) = 0.1.
Q: Using the PMF above, calculate E(X).
A: E(X) = 0(0.4) + 1(0.3) + 2(0.2) + 3(0.1) = 0 + 0.3 + 0.4 + 0.3 = 1.0.
Q: Using the same PMF, compute Var(X) via the shortcut formula.
A: E(X²) = 0²(0.4) + 1²(0.3) + 2²(0.2) + 3²(0.1) = 0 + 0.3 + 0.8 + 0.9 = 2.0. Var(X) = 2.0 – (1.0)² = 1.0.
Q: If Y = 10 + 5X, what are E(Y) and Var(Y)?
A: E(Y) = 10 + 5(1.0) = 15. Var(Y) = 5² × 1.0 = 25.
Q: If X and Y are independent with Var(X) = 4 and Var(Y) = 9, what is the standard deviation of X + Y?
A: Var(X + Y) = 4 + 9 = 13. σ(X + Y) = √13 ≈ 3.606.
The mean and variance rules here carry directly into the binomial and Poisson distributions (Ch. 5 continued), where the PMF has a named formula and the mean and variance have their own shortcuts. The CDF concept becomes central in the continuous (normal) distribution chapters. The variance rules for sums of independent random variables form the basis of sampling distributions and confidence intervals later in the course.
Random variable, RV, discrete random variable, continuous random variable, probability mass function, PMF, probability distribution, expected value, mean of a random variable, E(X), mu, variance, Var(X), sigma squared, standard deviation, sigma, cumulative distribution function, CDF, linear transformation rules, sum of random variables, independence, Purdue STAT, intro statistics Chapter 5