Probability Distributions, PDFs/CDFs, Sampling Distributions, and the CLT, STAT 350 Midterm 1 – Study Notes
offline

Source: STAT 350 Practice Exam 1, Purdue University

Tags: PDF, CDF, normal distribution, binomial distribution, Poisson distribution, Central Limit Theorem, CLT, sampling distribution, expected value, variance, standard deviation, percentile, z-score, continuous random variable

Difficulty: Intermediate Prerequisites: Probability rules, basic integration, comfort with summation notation. Review the Probability Rules study notes before tackling this material.


Big Picture

This is the largest and most heavily tested block of content on STAT 350 Midterm 1. It covers how to work with probability distributions (both continuous and discrete), how to validate a PDF, build a CDF, compute probabilities and percentiles, and find expected values and variances. It then connects individual distributions to what happens when you take samples: the sampling distribution of the mean and the Central Limit Theorem. The Poisson distribution rounds out the discrete side. If you can work a piecewise PDF/CDF problem from start to finish and apply the CLT to find probabilities about sample means, you are well prepared for the exam.


TL;DR

A valid PDF integrates to 1 over its support. The CDF is the running integral of the PDF. Probabilities come from the CDF. The normal distribution is central because the CLT guarantees that sample means are approximately normal for large samples, regardless of the population shape. The Poisson distribution models counts of rare events in fixed intervals and scales linearly with the interval length.


Key Terms

Probability density function (PDF)

For a continuous random variable X, the PDF f(x) describes the relative likelihood of X near any value. The total area under f(x) must equal 1, and f(x) ≥ 0 everywhere. Think of it as the "shape" of the distribution. The area under the curve between two values gives the probability of landing between them.

Cumulative distribution function (CDF)

F(x) = P(X ≤ x). The CDF is the running total of probability up to x. For continuous variables, it is the integral of the PDF from negative infinity to x. In simple terms, it answers "what is the probability of getting a value at most x?"

Valid PDF conditions

Two conditions: (1) f(x) ≥ 0 for all x, and (2) the integral of f(x) over all x equals 1.

Expected value (mean)

E[X] = integral of x · f(x) dx over the support. The "centre of mass" of the distribution. Think of it as the long-run average if you repeated the experiment many times.

Variance and standard deviation

Var(X) = E[X^2] - (E[X])^2. The standard deviation is SD(X) = sqrt(Var(X)). Variance measures how spread out the distribution is around its mean.

Percentile (quantile)

The p-th percentile is the value x_p such that F(x_p) = p. In other words, p proportion of the distribution falls at or below x_p. In simple terms, "the value below which p fraction of the data lies."

Normal distribution

X ~ N(mu, sigma^2). Symmetric, bell-shaped. Fully determined by its mean mu and variance sigma^2. The standard normal has mu = 0 and sigma = 1.

Standard normal (Z)

Z ~ N(0, 1). Any normal variable X can be standardised: Z = (X - mu) / sigma. Probabilities are read from the Z-table.

Binomial distribution

X ~ Binomial(n, p). Counts the number of successes in n independent Bernoulli trials, each with success probability p. E[X] = np, Var(X) = np(1-p).

Poisson distribution

X ~ Poisson(lambda). Models the count of events in a fixed interval when events occur independently at a constant average rate lambda. E[X] = lambda, Var(X) = lambda. Think of it as the distribution for "how many rare events happen in a given time window."

Central Limit Theorem (CLT)

If X_1, ..., X_n are independent and identically distributed with mean mu and variance sigma^2, then the sampling distribution of the sample mean x-bar is approximately N(mu, sigma^2 / n) for sufficiently large n. In simple terms, averages of large samples are approximately normally distributed regardless of the population's shape.

Sampling distribution of the sample mean

The probability distribution of x-bar across all possible samples of size n from a population. Its mean is mu, and its variance is sigma^2 / n (standard error = sigma / sqrt(n)).


Core Content

Working with Piecewise PDFs

Many exam problems give a PDF defined in pieces over different intervals, with an unknown constant k. The strategy:

  • Write the total-area condition: the sum of the integrals of f(x) over each piece equals 1.

  • Solve for k.

  • For the CDF, integrate the PDF piece by piece, carrying forward the accumulated probability from previous intervals.

  • Check your CDF: it should be 0 at the left boundary of the support, 1 at the right boundary, and non-decreasing throughout.

Worked example (from the practice exam)

The PDF is:

  • f(x) = k(x + 1) for -1 < x < 0

  • f(x) = (12/211)(-x^2 + 2x + 8) for 0 < x < 2

  • f(x) = k(3 - x) for 2 < x < 3

  • f(x) = 0 otherwise

To find k, integrate each piece:

  • Integral from -1 to 0 of k(x+1) dx = k · [x^2/2 + x] from -1 to 0 = k · [0 - (1/2 - 1)] = k/2

  • Integral from 0 to 2 of (12/211)(-x^2 + 2x + 8) dx = (12/211) · [-x^3/3 + x^2 + 8x] from 0 to 2 = (12/211) · (-8/3 + 4 + 16) = (12/211)(52/3) = 208/211

  • Integral from 2 to 3 of k(3-x) dx = k · [3x - x^2/2] from 2 to 3 = k · [(9 - 9/2) - (6 - 2)] = k/2

Setting k/2 + 208/211 + k/2 = 1 gives k + 208/211 = 1, so k = 3/211.

Building the CDF from a Piecewise PDF

For the same example, the CDF for each interval is the accumulated probability up to x:

  • F(x) = 0 for x < -1

  • F(x) = (3/422)(x + 1)^2 for -1 ≤ x < 0 (integrate k(t+1) from -1 to x with k = 3/211)

  • F(x) = (1/422)(-8x^3 + 24x^2 + 192x + 3) for 0 ≤ x < 2 (given in the exam)

  • F(x) = (1/422)(-3x^2 + 18x + 395) for 2 ≤ x < 3

  • F(x) = 1 for x ≥ 3

The missing piece (labelled A in the exam) for 2 ≤ x < 3 is found by:

  • Starting with F(2) = (1/422)(419) from the previous piece

  • Adding the integral of k(3 - t) from 2 to x: (3/211)[3t - t^2/2] from 2 to x = (3/422)(6x - x^2 - 8)

  • Combining: F(x) = (1/422)(419 + 18x - 3x^2 - 24) = (1/422)(-3x^2 + 18x + 395)

Verify: F(3) = (1/422)(-27 + 54 + 395) = 422/422 = 1. Correct.

Computing Probabilities from the CDF

  • P(X > a) = 1 - F(a)

  • P(a < X < b) = F(b) - F(a)

Example: P(X > 0.5) = 1 - F(0.5). Using the piece for 0 ≤ x < 2:

F(0.5) = (1/422)(-8(0.125) + 24(0.25) + 192(0.5) + 3) = (1/422)(-1 + 6 + 96 + 3) = 104/422

P(X > 0.5) = 1 - 104/422 = 318/422 = 159/211 ≈ 0.7536

Example: P(0.5 < X < 2.5) = F(2.5) - F(0.5).

F(2.5) = (1/422)(-3(6.25) + 18(2.5) + 395) = (1/422)(421.25) = 1685/1688

P(0.5 < X < 2.5) = 1685/1688 - 104/422 = 1685/1688 - 416/1688 = 1269/1688 ≈ 0.7517

Finding Percentiles from the CDF

To find the p-th percentile, set F(x) = p and solve for x. You must first determine which piece of the CDF contains the answer.

Example: Find the 0.1 percentile (p = 0.001).

F(-1) = 0, F(0) = 3/422 ≈ 0.0071. Since 0.001 < 0.0071, the answer lies in -1 ≤ x < 0.

Set (3/422)(x + 1)^2 = 0.001. Then (x+1)^2 = 0.001 · 422/3 = 0.14067. So x + 1 = 0.3751, giving x ≈ -0.62.

Expected Value, Variance, and Standard Deviation

For the practice exam distribution, E[X] = 1 (the distribution is symmetric about x = 1 on the support [-1, 3]).

Given E[X^2] = 2836/2110:

  • Var(X) = E[X^2] - (E[X])^2 = 2836/2110 - 1 = 726/2110 = 363/1055 ≈ 0.3441

  • SD(X) = sqrt(363/1055) ≈ 0.5866

The Normal Distribution and the Z-Table

  • If X ~ N(mu, sigma^2), then Z = (X - mu)/sigma ~ N(0,1).

  • P(mu - 2sigma < X < mu + 2sigma) = P(-2 < Z < 2) = P(Z < 2) - P(Z < -2).

  • By symmetry, P(Z < -2) = 1 - P(Z < 2), so this equals 2P(Z < 2) - 1. This identity is tested as a true/false question.

The Binomial Distribution Variance

For X ~ Binomial(n, p):

  • E[X] = np

  • Var(X) = np(1 - p)

Rewriting: Var(X) = np(1-p) = (1/n) · np · (n - np) = (1/n) · E[X] · (n - E[X]).

This algebraic identity is true and is tested. It is just a rearrangement, not a separate formula.

Variance of a Sum vs Variance of a Mean

This distinction causes many errors:

  • If Y = X_1 + X_2 + ... + X_n (the sum) and the X_i are independent with variance sigma^2, then Var(Y) = n · sigma^2. The variances add.

  • If X-bar = Y/n (the mean), then Var(X-bar) = sigma^2 / n.

The statement "Var(sum) = sigma^2 / n" is false. That is the variance of the mean, not the sum. Var(sum) = n · sigma^2.

The Central Limit Theorem

The CLT says: if X_1, ..., X_n are iid with mean mu and variance sigma^2, then for large n:

X-bar approximately ~ N(mu, sigma^2 / n)

Key conditions and caveats:

  • If the population is already normal, then X-bar is exactly normal for any n, even n = 1. No minimum sample size is needed.

  • The "n ≥ 30" guideline is a rule of thumb for non-normal populations. It is not a universal requirement.

  • For extremely skewed populations or populations with outliers, n = 30 may not be sufficient. You may need a "sufficiently large" sample size, and there should be no real outliers.

  • The statement "the CLT requires n ≥ 30 for the sampling distribution of the mean to be normal" is false in general. It is false specifically when the population is already normal.

Sampling Distribution of the Mean

Once you know E[X] and Var(X), and the CLT (or normality of the population) applies:

X-bar ~ N(mu, sigma^2/n)

Standard error: SE = sigma / sqrt(n)

To find probabilities: standardise as Z = (X-bar - mu) / (sigma / sqrt(n)) and use the Z-table.

Example: With E[X] = 1, SD(X) = sqrt(363/1055) ≈ 0.5866, and n = 42:

SE = 0.5866 / sqrt(42) ≈ 0.0905

P(X-bar > 1.35) = P(Z > (1.35 - 1) / 0.0905) = P(Z > 3.87) ≈ 0.0001

For the 91.92nd percentile of X-bar: find z such that P(Z < z) = 0.9192. From the Z-table, z = 1.40 (since P(Z < 1.40) = 0.9192).

X-bar = mu + z · SE = 1 + 1.40 · 0.0905 = 1 + 0.1267 = 1.1267

The Poisson Distribution

X ~ Poisson(lambda). Used for counts of events in a fixed interval.

Core properties:

  • E[X] = lambda

  • Var(X) = lambda (mean equals variance, a signature property)

  • SD(X) = sqrt(lambda)

  • P(X = k) = (lambda^k · e^(-lambda)) / k!

Scaling: if the rate is lambda for a given interval length, then for an interval of different length, scale proportionally. If lambda = 3.5 per 12 hours, then lambda = 3.5/12 per 1 hour ≈ 0.2917 per hour.

Independence: counts in non-overlapping intervals are independent. So P(X_1 = a and X_2 = b) = P(X_1 = a) · P(X_2 = b) for non-overlapping time periods.

Worked example (from the practice exam)

A call centre averages 3.5 unsatisfactory calls per 12-hour period (Poisson).

(a) Average per 1 hour: lambda_1 = 3.5 / 12 = 0.2917

(b) P(at least 1 unsatisfactory in 1 hour): P(X ≥ 1) = 1 - P(X = 0) = 1 - e^(-0.2917) ≈ 1 - 0.7469 = 0.2531

(c) P(exceeds average in 12 hours): P(X > 3.5) = P(X ≥ 4), since X is discrete.

P(X ≤ 3) = P(X=0) + P(X=1) + P(X=2) + P(X=3)

With lambda = 3.5:

  • P(X=0) = e^(-3.5) ≈ 0.0302

  • P(X=1) = 3.5 · e^(-3.5) ≈ 0.1057

  • P(X=2) = 3.5^2 / 2 · e^(-3.5) ≈ 0.1850

  • P(X=3) = 3.5^3 / 6 · e^(-3.5) ≈ 0.2158

P(X ≤ 3) ≈ 0.5367, so P(X ≥ 4) ≈ 0.4634

(d) P(exactly 1 in each of two separate 1-hour periods): Since non-overlapping intervals are independent: P(X_1 = 1) · P(X_2 = 1).

P(X = 1) = lambda_1 · e^(-lambda_1) = 0.2917 · 0.7469 ≈ 0.2179

P(both) = 0.2179^2 ≈ 0.0475

(e) SD for 1-hour period: SD = sqrt(lambda_1) = sqrt(0.2917) ≈ 0.5401


Formulas / Diagrams

  • PDF validity: integral of f(x) dx = 1, and f(x) ≥ 0

  • CDF: F(x) = P(X ≤ x) = integral from -inf to x of f(t) dt

  • P(a < X < b) = F(b) - F(a)

  • E[X] = integral of x · f(x) dx

  • Var(X) = E[X^2] - (E[X])^2

  • SD(X) = sqrt(Var(X))

  • Percentile: solve F(x_p) = p for x_p

  • Z-score: Z = (X - mu) / sigma

  • P(mu - k·sigma < X < mu + k·sigma) = 2P(Z < k) - 1

  • Binomial: E[X] = np, Var(X) = np(1-p)

  • Poisson: P(X=k) = lambda^k · e^(-lambda) / k!, E[X] = Var(X) = lambda

  • CLT: X-bar ~ N(mu, sigma^2/n) for large n

  • Standard error: SE = sigma / sqrt(n)

  • Var(sum of n independent) = n · sigma^2 (not sigma^2/n)


Real-World Applications

The CLT is the reason polling companies can estimate the preferences of millions of voters from a sample of a thousand. As long as the sample is large enough and properly random, the sample mean's distribution is approximately normal, which lets you put a margin of error on the estimate. The Poisson distribution models events like the number of server crashes per month, customer arrivals per hour, or defective items per production batch, anywhere you are counting occurrences of something relatively rare over a fixed window.


Common Misconceptions

  • Students often believe the CLT always requires n ≥ 30. If the population is normal, the sampling distribution of the mean is exactly normal for any sample size, even n = 5 or n = 2. The n ≥ 30 rule of thumb is for non-normal populations.

  • Students confuse Var(sum) with Var(mean). For n independent observations with variance sigma^2: Var(sum) = n · sigma^2 and Var(mean) = sigma^2/n. These differ by a factor of n^2.

  • Students sometimes forget that the Poisson rate must be scaled to match the interval. If the rate is per 12 hours, you must divide by 12 to get the rate per hour.

  • When finding P(X > lambda) for a Poisson, students sometimes forget that X is discrete. "Exceeds 3.5" means X ≥ 4, not X > 3.


Why It Matters / Exam Flags

⚠️ For a piecewise PDF: the integral of all pieces must sum to 1. Check that the CDF is 0 at the left boundary and 1 at the right.

⚠️ The CDF piece for 2 ≤ x < 3 must be continuous with the CDF piece for 0 ≤ x < 2 at x = 2. Use F(2) from the previous piece as your starting value.

⚠️ Var(Y = sum of X_i) = n · sigma^2, not sigma^2/n. The formula sigma^2/n is for the variance of the mean, not the sum.

⚠️ If the population is normal, X-bar is exactly normal for any n. The CLT's large-sample requirement is irrelevant.

⚠️ For the Poisson, P(X > 3.5) = P(X ≥ 4) = 1 - P(X ≤ 3). Compute the individual terms P(X=0) through P(X=3).

⚠️ P(mu - 2sigma < X < mu + 2sigma) = 2P(Z < 2) - 1 is true for any normal X.

⚠️ Var(X) = (1/n) · E[X] · (n - E[X]) is a valid identity for the binomial, just a rearrangement of np(1-p).

⚠️ All numeric answers should have four decimal places unless stated otherwise.


Quick Self-Test

  1. True or false: If the population is normally distributed, you need n ≥ 30 for the sampling distribution of the mean to be normal.

  1. Fill in the blank: If Y is the sum of n independent variables each with variance sigma^2, then Var(Y) = ________.

  1. True or false: For a Poisson random variable, the mean equals the variance.

  1. Fill in the blank: To find the 91.92nd percentile using the Z-table, look up the z-value where P(Z < z) = ________.

  1. True or false: P(mu - 2sigma < X < mu + 2sigma) = 2P(Z < 2) - 1 when X is normal.

Answers: 1. False (exact normality holds for any n). 2. n · sigma^2. 3. True. 4. 0.9192. 5. True.


Practice Q&A

Q: A population is known to be normally distributed. A research group can only afford a sample of size 5. Is the sampling distribution of the sample mean normally distributed?

A: Yes. When the population itself is normal, the sampling distribution of the sample mean is exactly normal for any sample size, including n = 5. The CLT's large-sample condition is not needed here because normality is inherited directly from the population.

Q: Let X_1, ..., X_n be iid with mean mu and variance sigma^2. If Y = sum of X_i, what is Var(Y)?

A: Var(Y) = n · sigma^2. Because the variables are independent, variances add. This is distinct from Var(X-bar) = sigma^2/n.

Q: X follows a Binomial(n, p) distribution. Show that Var(X) = (1/n) · E[X] · (n - E[X]).

A: E[X] = np, so (1/n) · E[X] · (n - E[X]) = (1/n)(np)(n - np) = p(n - np) = np - np^2 = np(1 - p) = Var(X). The identity holds.

Q: A Poisson process has an average of 3.5 events per 12 hours. What is the average rate per hour, and what is the probability of at least one event in a 1-hour period?

A: Rate per hour: lambda = 3.5/12 ≈ 0.2917. P(X ≥ 1) = 1 - P(X = 0) = 1 - e^(-0.2917) ≈ 1 - 0.7469 = 0.2531.

Q: For a distribution with E[X] = 1, E[X^2] = 2836/2110, and n = 42, find the 91.92nd percentile of X-bar.

A: Var(X) = 2836/2110 - 1 = 726/2110. SE = sqrt(726/2110) / sqrt(42) ≈ 0.5866 / 6.4807 ≈ 0.0905. The z-value for 91.92% is 1.40 (from the Z-table). X-bar = 1 + 1.40(0.0905) = 1.1267.

Q: Two non-overlapping 1-hour periods each have a Poisson rate of 0.2917. What is the probability of exactly 1 event in each period?

A: Non-overlapping intervals are independent. P(X = 1) = 0.2917 · e^(-0.2917) ≈ 0.2179. The probability of exactly 1 in each = 0.2179^2 ≈ 0.0475.


Connections to Other Topics

  • The normal distribution reappears in confidence intervals and hypothesis testing later in STAT 350. The Z-table skills practised here carry directly forward.

  • The CLT is the reason confidence intervals for the mean work: the sampling distribution being approximately normal allows us to quantify uncertainty with a margin of error.

  • The Poisson distribution connects to the exponential distribution (the time between Poisson events is exponentially distributed), which may appear in later coursework.


Related Terms / Search Tags

PDF, probability density function, CDF, cumulative distribution function, piecewise PDF, valid PDF, integration, expected value, E[X], variance, Var(X), standard deviation, percentile, quantile, z-score, z-table, normal distribution, standard normal, N(0 1), binomial distribution, Poisson distribution, lambda, Central Limit Theorem, CLT, sampling distribution, sample mean, x-bar, standard error, sigma over root n, variance of sum, variance of mean, STAT 350 Purdue, intro statistics distributions