Random Variables and Probability Distributions, STAT 101 – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Descriptive statistics notes and probability foundations notes (both earlier documents in this set).

This is where probability becomes a modelling tool. Instead of asking "what is the chance of this event?" you start asking "what does the full picture of possible outcomes look like, and what can I expect on average?" Random variables connect the abstract rules of probability to the concrete distributions (binomial, Poisson, normal) that you will use for inference later in the course. If you are comfortable with conditional probability and the multiplication rule, you are ready for this material.

TL;DR

Random variables assign numbers to outcomes. Discrete ones have a probability mass function (pmf); continuous ones have a probability density function (pdf). Expected value is the long-run average; variance measures spread. The binomial models fixed-trial success counts, the Poisson models event rates, and the normal distribution is the bell curve that underpins most of statistical inference.

Key Terms

Random variable

A function that assigns a numerical value to each outcome in a sample space. Think of it as a rule that turns the result of an experiment into a number you can do maths with.

Discrete random variable

A random variable that takes countable values (often whole numbers). Example: number of heads in 10 coin flips. You can list all possible values, even if the list is long.

Continuous random variable

A random variable that can take any value within an interval. Example: the exact time a customer waits in a queue. You cannot list every possible value because there are infinitely many.

Probability mass function (pmf)

For a discrete random variable, the function that gives the probability of each possible value. P(X = x) for every x. All values are non-negative and sum to 1. In simple terms, it is the complete table of "value, probability" pairs.

Probability density function (pdf)

For a continuous random variable, the function whose area under the curve over an interval gives the probability of falling in that interval. The probability at any single exact point is 0; only intervals have non-zero probability.

Expected value (E(X), mean of a random variable)

The long-run average value of the random variable if the experiment were repeated many times. For discrete variables: E(X) = sum of [x times P(X = x)]. Think of it as a weighted average, where the weights are probabilities.

Variance of a random variable (Var(X))

Measures how spread out the values of the random variable are around the expected value. Var(X) = E[(X - E(X))squared]. Larger variance means more spread.

Binomial distribution

Models the number of successes in a fixed number (n) of independent trials, each with the same probability of success (p). Example: number of heads in 20 coin flips. In simple terms, it counts how many times something happens in a set of identical, independent attempts.

Poisson distribution

Models the number of events occurring in a fixed interval of time or space, when events happen independently at a constant average rate (lambda). Example: number of emails received per hour. In simple terms, it counts how many times something happens in a window.

Normal distribution (Gaussian distribution)

A symmetric, bell-shaped continuous distribution defined by its mean (mu) and standard deviation (sigma). The most important distribution in statistics because of the Central Limit Theorem. In simple terms, it is the classic bell curve.

Standard normal distribution

A normal distribution with mean 0 and standard deviation 1. Any normal variable can be converted to a standard normal variable using the z-score formula. Z-tables are based on this distribution.

Exponential distribution

Models the waiting time between independent events occurring at a constant rate. The continuous counterpart to the Poisson. Example: time between successive customer arrivals.

Normal probability plot (Q-Q plot)

A graph that plots observed data against expected z-scores. If the data is approximately normal, the points fall roughly along a straight line. Deviations from the line suggest skewness, outliers or other departures from normality.

Random Variables

A random variable bridges events and numbers. Once you define a random variable, you can talk about its distribution, its average and its spread using precise mathematical tools.

  • Discrete random variables have a probability mass function (pmf) that lists every possible value alongside its probability. Two rules must hold: every probability is between 0 and 1, and all probabilities sum to exactly 1.

  • Continuous random variables have a probability density function (pdf). The probability that X falls in an interval [a, b] equals the area under the pdf curve between a and b. The total area under the entire curve is 1. A key consequence: P(X = any single value) = 0 for continuous variables, because a single point has no width and therefore no area.

The distinction matters for how you calculate probabilities. For discrete variables, you sum individual probabilities. For continuous variables, you integrate (or look up areas in a table).

Expected Value and Variance

Expected value and variance summarise a random variable's distribution the same way mean and standard deviation summarise a dataset.

Expected value is the long-run average. If you repeated the experiment thousands of times and averaged the results, that average would converge to E(X).

  • Discrete: E(X) = sum of [x times P(X = x)] across all values of x.

  • Continuous: E(X) = integral of [x times f(x)] dx.

Linearity of expectation is one of the most useful rules in probability. E(aX + bY) = aE(X) + bE(Y). This holds whether or not X and Y are independent.

Variance measures spread around the expected value.

  • Var(X) = E[(X - E(X)) squared] = E(X squared) - [E(X)] squared.

  • For independent random variables: Var(X + Y) = Var(X) + Var(Y). Independence is required here, unlike for expected value.

Standard deviation of a random variable is the square root of its variance, putting the measure of spread back into the original units.

Binomial Distribution

The binomial distribution models the count of successes in a fixed number of independent, identical trials.

Four conditions must all hold for a binomial model to apply.

  1. Fixed number of trials (n).

  1. Each trial has exactly two outcomes (success or failure).

  1. The probability of success (p) is the same on every trial.

  1. Trials are independent.

The probability of exactly k successes:

P(X = k) = C(n, k) x p^k x (1 - p)^(n - k)

where C(n, k) is the binomial coefficient "n choose k."

  • Mean: E(X) = np

  • Variance: Var(X) = np(1 - p)

  • Standard deviation: SD(X) = square root of np(1 - p)

As a quick check: the mean should feel sensible. If you flip a fair coin 100 times, you expect 100 x 0.5 = 50 heads. The variance tells you how much that count typically varies from 50.

Poisson Distribution

The Poisson distribution models the count of events in a fixed interval (of time, space, volume, etc.) when events occur independently at a constant average rate.

Conditions for the Poisson model:

  1. Events occur independently of each other.

  1. The average rate (lambda) is constant across the interval.

  1. Two events cannot occur at exactly the same instant.

The probability of exactly k events:

P(X = k) = (lambda^k x e^(-lambda)) / k!

  • Mean: E(X) = lambda

  • Variance: Var(X) = lambda

A distinctive feature: the mean and variance are equal. If you are told the mean number of calls per hour is 5, the variance is also 5.

The Poisson can approximate the binomial when n is large and p is small (so that np = lambda is moderate). This is sometimes called the "law of rare events."

Normal Distribution and Standardisation

The normal distribution is a symmetric, bell-shaped curve defined entirely by two parameters: the mean (mu) and the standard deviation (sigma).

Properties worth knowing:

  • The curve is symmetric about the mean.

  • The mean, median and mode are all equal.

  • About 68% of values fall within 1 standard deviation of the mean, about 95% within 2, and about 99.7% within 3 (the empirical rule, or 68-95-99.7 rule).

  • The tails extend infinitely in both directions but never touch zero.

Standardisation converts any normal variable X into the standard normal variable Z (mean 0, standard deviation 1):

z = (x - mu) / sigma

Once standardised, you use z-tables to look up probabilities. The table gives the area to the left of a given z-value, which equals P(Z is less than or equal to z).

Common tasks:

  • Finding P(X < some value): standardise, then look up the z-score.

  • Finding a percentile: look up the area in the z-table to find the z-score, then convert back: x = mu + z times sigma.

  • Finding P(a < X < b): compute P(X < b) - P(X < a).

Assessing Normality

Many statistical methods assume the data is approximately normal, so checking that assumption matters.

Normal probability plots (Q-Q plots) are the standard tool. They plot the ordered data values on the y-axis against the z-scores those values would have if the data were perfectly normal on the x-axis.

  • If the points fall roughly along a straight line, the data is approximately normal.

  • A curve bowing upward suggests right skew.

  • A curve bowing downward suggests left skew.

  • Points peeling away from the line at the ends suggest heavy tails (more extreme values than a normal distribution would produce) or outliers.

Histograms and boxplots provide supporting evidence, but the normal probability plot is the most sensitive single check.

Exponential and Other Continuous Distributions

Exponential distribution models the waiting time between independent events occurring at a constant rate lambda.

  • pdf: f(x) = lambda x e^(-lambda x), for x >= 0.

  • Mean: 1 / lambda.

  • Variance: 1 / lambda squared.

  • The exponential is memoryless: the probability of waiting another t minutes is the same regardless of how long you have already waited.

Other continuous distributions covered at an introductory level:

  • Gamma distribution: generalises the exponential. Models waiting times for multiple events (the time until the k-th event). Has shape and scale parameters.

  • Beta distribution: models probabilities or proportions bounded between 0 and 1. Useful for modelling rates, percentages and Bayesian priors.

  • Weibull distribution: used in reliability engineering and failure analysis. Can model increasing, decreasing or constant hazard rates depending on its shape parameter.

  • Lognormal distribution: applies when the logarithm of the variable is normally distributed. Common for income data, stock prices and survival times, all of which are positively skewed.

Formulas

Expected value (discrete):

E(X) = \sum_{x} x \cdot P(X = x)

Variance:

\text{Var}(X) = E[(X - E(X))^2] = E(X^2) - [E(X)]^2

Linearity of expectation: E(aX + bY) = aE(X) + bE(Y)

Variance of independent sum: Var(X + Y) = Var(X) + Var(Y)

Binomial probability:

P(X = k) = \binom{n}{k} p^k (1-p)^{n-k}

Binomial mean: E(X) = np. Binomial variance: Var(X) = np(1 - p).

Poisson probability:

P(X = k) = \frac{\lambda^k e^{-\lambda}}{k!}

Poisson mean and variance: E(X) = Var(X) = lambda.

Standardisation: z = (x - mu) / sigma

Exponential pdf: f(x) = lambda x e^(-lambda x), for x >= 0. Mean = 1 / lambda.

Real-World Applications

The binomial distribution models quality control inspections: out of 200 manufactured parts, how many are defective if each has a 2% defect rate? The Poisson distribution is used in traffic engineering to model the number of cars arriving at an intersection per minute, and in biology to count mutations per DNA segment. The normal distribution underpins most standardised testing (SAT, GRE), where raw scores are converted to a bell-curve scale. The exponential distribution models the time between server requests in web infrastructure, helping engineers plan capacity.

Common Misconceptions

  • Students often confuse the pmf with the pdf. The pmf gives actual probabilities for discrete variables; the pdf gives density for continuous variables, and you need the area under the curve (not the height) to get a probability.

  • A common error is thinking P(X = 5) is meaningful for a continuous random variable. For continuous distributions, only interval probabilities are non-zero.

  • Students sometimes apply the binomial formula when trials are not independent (e.g. drawing cards without replacement). If the population is large relative to the sample, independence is approximately satisfied, but for small populations it is not.

  • Many students forget that the Poisson mean and variance are equal. If you are told the mean is 7, the standard deviation is the square root of 7, not 7.

Why It Matters / Exam Flags

  • Be ready to verify whether a situation meets the four binomial conditions before applying the formula. Exams often include a scenario where one condition fails.

  • Know how to use the z-table in both directions: given a z-score, find a probability; given a probability, find a z-score (for percentile problems).

  • Expect at least one problem asking you to compute expected value from a probability table. Multiply each value by its probability, then sum.

  • The 68-95-99.7 rule for the normal distribution is tested frequently. Know it cold.

  • Be prepared to interpret a normal probability plot. "Roughly straight" means approximately normal; systematic curvature means it is not.

  • For the Poisson, know that mean equals variance. If a problem states one, you have both.

Quick Self-Test

  1. True or false: for a continuous random variable, P(X = 3) = 0. (True)

  1. Fill in the blank: the mean of a binomial distribution with n = 50 and p = 0.4 is ____. (20)

  1. True or false: the Poisson variance equals lambda squared. (False, it equals lambda)

  1. Fill in the blank: approximately ____% of values in a normal distribution fall within 2 standard deviations of the mean. (95)

  1. True or false: E(X + Y) = E(X) + E(Y) only when X and Y are independent. (False, linearity holds regardless of independence)

Practice Q&A

Q: A random variable X has the following distribution: P(X = 1) = 0.2, P(X = 2) = 0.5, P(X = 3) = 0.3. What is E(X)?

A: E(X) = (1)(0.2) + (2)(0.5) + (3)(0.3) = 0.2 + 1.0 + 0.9 = 2.1.

Q: A factory produces light bulbs with a 5% defect rate. In a batch of 20 bulbs, what is the probability that exactly 2 are defective?

A: This is binomial with n = 20, p = 0.05, k = 2. P(X = 2) = C(20, 2) x (0.05)^2 x (0.95)^18 = 190 x 0.0025 x 0.3972 = 0.1887, or about 18.9%.

Q: A call centre receives an average of 3 calls per minute. What is the probability of receiving exactly 5 calls in a given minute?

A: Poisson with lambda = 3. P(X = 5) = (3^5 x e^(-3)) / 5! = (243 x 0.0498) / 120 = 0.1008, or about 10.1%.

Q: Scores on an exam are normally distributed with mean 72 and standard deviation 8. What proportion of students scored above 88?

A: z = (88 - 72) / 8 = 2.0. From the z-table, P(Z < 2.0) = 0.9772. So P(Z > 2.0) = 1 - 0.9772 = 0.0228. About 2.3% of students scored above 88.

Q: For a binomial distribution with n = 100 and p = 0.3, what are the mean and standard deviation?

A: Mean = np = 100 x 0.3 = 30. Variance = np(1 - p) = 100 x 0.3 x 0.7 = 21. Standard deviation = square root of 21 = 4.58.

Connections to Other Topics

The normal distribution connects directly to the Central Limit Theorem, which states that the sampling distribution of the sample mean approaches normal as sample size grows, regardless of the population shape. This is the foundation for confidence intervals and hypothesis tests. The binomial distribution becomes approximately normal when np and n(1 - p) are both at least 10, which is why the normal approximation to the binomial is a standard exam topic. Expected value and variance reappear in regression, where you model the expected value of a response variable as a function of predictors.

Related Terms / Search Tags

random variable, discrete, continuous, probability mass function, pmf, probability density function, pdf, expected value, mean of a random variable, variance, standard deviation, binomial distribution, Bernoulli trial, Poisson distribution, lambda, normal distribution, Gaussian, bell curve, standard normal, z-table, z-score, standardisation, empirical rule, 68-95-99.7, normal probability plot, Q-Q plot, exponential distribution, memoryless property, gamma distribution, beta distribution, Weibull distribution, lognormal distribution, Central Limit Theorem, STAT 101, introduction to statistics, Purdue