Probability Distributions: Discrete Random Variables – STAT, Introduction to Statistics – Study Notes
offline

Source: Probability Distributions, Key Concepts and Applications (Purdue University)

Tags: probability distribution, discrete distribution, random variable, binomial, Poisson, hypergeometric, uniform distribution, PMF, PDF, CDF, expected value, variance, intro stats

Difficulty: Introductory Prerequisites: Basic probability rules (addition rule, multiplication rule, complement rule). Comfortable with summation notation and combinatorics (factorials, "n choose k").


Big Picture

Probability distributions are the bridge between raw probability and applied statistics. Once you can describe how a random variable behaves, you can predict outcomes, estimate parameters, and run hypothesis tests. This first set of notes covers the discrete side: variables that take countable values, and the four named distributions you will use most often. If you are coming in cold, make sure you are solid on basic probability rules and counting methods before diving in.


TL;DR

A probability distribution maps every possible outcome of a random variable to its probability. Discrete distributions handle countable outcomes (coin flips, defect counts, survey responses). The four you need to know are the discrete uniform, binomial, Poisson, and hypergeometric, each suited to a different type of experiment.


Key Terms

Random variable

A numerical outcome of a random process. In simple terms, it is the number you get when you run an experiment.

Discrete random variable

A random variable that takes on countable values (0, 1, 2, …). Think of it as anything you can list out: number of heads, number of students, number of defective parts.

Continuous random variable

A random variable that can take any value in an interval. Think of it as anything you measure on a scale: height, temperature, time.

Probability distribution function (PDF, discrete case)

A function that gives P(X = x) for every possible value x. In simple terms, it is the rule that tells you how likely each outcome is.

Probability mass function (PMF)

Another name for the discrete PDF. Each value gets a specific probability, and all probabilities sum to 1.

Cumulative distribution function (CDF)

F(x) = P(X ≤ x). It tells you the probability of getting a value at or below x. Think of it as a running total of probability as you move along the number line.

Expected value (mean), E(X)

The long-run average of the random variable if you repeated the experiment many times. Formally: E(X) = Σ xᵢ P(xᵢ).

Variance, Var(X)

A measure of how spread out the distribution is around the mean. Formally: Var(X) = Σ (xᵢ − μ)² P(xᵢ).


Core Content

Foundations of Discrete Distributions

  • Every valid discrete distribution satisfies two rules:

    • Each probability is between 0 and 1 inclusive

    • All probabilities sum to exactly 1

  • The CDF is non-decreasing and approaches 1 as x grows without bound

Discrete Uniform Distribution

  • Every outcome is equally likely (rolling a fair die, drawing a random digit)

  • If the possible values run from a to b:

    • Mean: μ = (a + b) / 2

    • Variance: σ² = ((b − a + 1)² − 1) / 12

  • Straightforward to work with because each outcome has probability 1 / (b − a + 1)

Binomial Distribution

  • Models the number of successes in n independent trials, each with the same probability of success p

  • Conditions (all four must hold):

    • Fixed number of trials, n

    • Two outcomes per trial (success / failure)

    • Constant probability p on every trial

    • Trials are independent

  • Mean: μ = np

  • Variance: σ² = np(1 − p)

  • PMF: P(X = k) = C(n, k) · pᵏ · (1 − p)ⁿ⁻ᵏ

  • Classic examples: number of heads in 10 coin flips, number of defective items in a production batch of 50

Poisson Distribution

  • Counts the number of events in a fixed interval of time or space, where events occur independently at a constant average rate λ

  • Conditions:

    • Events are independent of one another

    • The average rate λ is constant across the interval

    • Two events cannot occur at exactly the same instant

  • Mean: μ = λ

  • Variance: σ² = λ (note: mean equals variance, a signature property)

  • PMF: P(X = k) = (λᵏ · e⁻λ) / k!

  • Classic examples: number of calls to a helpdesk per hour, number of cars through a toll booth in 10 minutes

Hypergeometric Distribution

  • Models the number of successes in n draws from a finite population of size N containing K successes, drawn without replacement

  • This is the distribution to reach for when sampling without replacement matters (small population, large sample fraction)

  • Mean: μ = nK / N

  • Variance: σ² = n · (K/N) · ((N − K)/N) · ((N − n)/(N − 1))

  • PMF: P(X = k) = [C(K, k) · C(N − K, n − k)] / C(N, n)

  • Classic example: drawing 5 cards from a 52-card deck and counting how many are hearts (K = 13, N = 52, n = 5)


Formulas and Diagrams

Distribution

Mean (μ)

Variance (σ²)

PMF

Discrete uniform (a to b)

(a + b) / 2

((b − a + 1)² − 1) / 12

1 / (b − a + 1)

Binomial (n, p)

np

np(1 − p)

C(n,k) pᵏ (1−p)ⁿ⁻ᵏ

Poisson (λ)

λ

λ

(λᵏ e⁻λ) / k!

Hypergeometric (N, K, n)

nK/N

n(K/N)((N−K)/N)((N−n)/(N−1))

C(K,k)C(N−K,n−k) / C(N,n)

Expected value (discrete): E(X) = Σ xᵢ P(xᵢ)

Variance (discrete): Var(X) = Σ (xᵢ − μ)² P(xᵢ)


Real-World Applications

  • Binomial in quality control: a manufacturer inspects 100 items from a production line and counts defectives. If the defect rate is stable at 3%, the count of defective items follows a Binomial(100, 0.03) distribution.

  • Poisson in operations: a call centre models incoming calls per 15-minute window to decide how many agents to staff. If calls arrive at an average rate of 8 per window, a Poisson(8) model helps estimate the probability of being overwhelmed.


Common Misconceptions

  • Students often confuse the binomial and hypergeometric. Use the binomial when trials are independent (sampling with replacement, or from a very large population). Use the hypergeometric when you sample without replacement from a smallish population.

  • Students sometimes assume Poisson requires events to be rare. It does not. λ can be large. The requirement is independence and a constant rate, not rarity.

  • A common error is forgetting that variance for the Poisson equals its mean. If you calculate them separately and they do not match, check your working.

  • Students occasionally apply the binomial formula when p changes from trial to trial. If the probability of success shifts between trials, the binomial does not apply.


Why It Matters / Exam Flags

⚠️ You will almost certainly be asked to identify which distribution fits a given scenario. Read the conditions carefully: fixed n with constant p = binomial; events per interval at constant rate = Poisson; without replacement from finite population = hypergeometric.

⚠️ Expect at least one question requiring you to compute a binomial or Poisson probability by hand using the PMF formula.

⚠️ The relationship "Poisson mean = Poisson variance" is a common true/false or short-answer target.

⚠️ Know when the binomial can approximate the hypergeometric (when the population is much larger than the sample).


Quick Self-Test

  1. True or false: For a valid probability distribution, the sum of all probabilities can be less than 1 if some outcomes are impossible.

    False. The sum must equal exactly 1.

  1. Fill in the blank: The binomial distribution requires that the probability of success is ______ on every trial.

    Constant (the same).

  1. True or false: The Poisson distribution's mean and variance are always equal.

    True.

  1. Fill in the blank: Sampling without replacement from a finite population calls for the ______ distribution.

    Hypergeometric.

  1. True or false: A random variable that counts the number of emails you receive per day is best modelled as continuous.

    False. Counts are discrete.


Practice Q&A

Q: A factory produces light bulbs with a 5% defect rate. In a random sample of 20 bulbs, what distribution models the number of defective bulbs, and what are its mean and variance?

A: Binomial with n = 20, p = 0.05. Mean = 20 × 0.05 = 1. Variance = 20 × 0.05 × 0.95 = 0.95.

Q: A hospital emergency department sees an average of 4 trauma cases per night. What is the probability of seeing exactly 6 trauma cases on a given night?

A: Use Poisson with λ = 4. P(X = 6) = (4⁶ · e⁻⁴) / 6! = 4096 · 0.0183 / 720 ≈ 0.1042, or about 10.4%.

Q: You draw 5 cards from a standard 52-card deck without replacement. What distribution models the number of aces drawn, and what is the expected number of aces?

A: Hypergeometric with N = 52, K = 4, n = 5. Expected number of aces = nK/N = 5 × 4 / 52 ≈ 0.385.

Q: Explain why sampling 10 items from a warehouse of 1,000,000 can reasonably use the binomial instead of the hypergeometric.

A: When the population is vastly larger than the sample, removing one item barely changes the composition. The probability of success stays nearly constant across draws, so the independence assumption of the binomial holds as a practical approximation.

Q: A fair six-sided die is rolled once. What are the mean and variance of the outcome?

A: Discrete uniform with a = 1, b = 6. Mean = (1 + 6)/2 = 3.5. Variance = ((6 − 1 + 1)² − 1)/12 = 35/12 ≈ 2.917.


Connections to Other Topics

  • Discrete distributions connect directly to continuous distributions (covered in Part 2). The binomial, for instance, can be approximated by the normal distribution when n is large, which is the foundation of many hypothesis tests.

  • The Poisson distribution links to the exponential distribution: if events arrive at rate λ per unit time (Poisson), the time between consecutive events follows an Exponential(λ) distribution.

  • Expected value and variance reappear throughout inferential statistics, particularly in confidence intervals and the Central Limit Theorem.


Related Terms / Search Tags

probability distribution, discrete probability, random variable, PMF, probability mass function, CDF, cumulative distribution function, expected value, mean, variance, binomial distribution, Poisson distribution, hypergeometric distribution, discrete uniform, sampling with replacement, sampling without replacement, n choose k, combinatorics, intro to statistics, Purdue statistics, STAT, probability theory, defect rate, count data, independent trials