Continuous Distributions, Sampling, and the Central Limit Theorem, STAT 301 Midterm 1 – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Descriptive Statistics notes; Probability Rules and Discrete Distributions notes

Big picture: This unit moves from discrete random variables (which take countable values) to continuous random variables (which can take any value in an interval). The key tools change from PMFs to PDFs and CDFs, but the logic stays the same: you describe the distribution, compute probabilities, and find means and variances. The unit also introduces the normal distribution (the most important distribution in statistics) and the Central Limit Theorem, which explains why sample means behave predictably even when the underlying population does not. This material is the direct on-ramp to confidence intervals and hypothesis testing.


TL;DR

Continuous random variables use probability density functions (PDFs) instead of mass functions. You find probabilities by integrating. The three continuous distributions on this exam are the uniform, exponential, and normal. The normal distribution is central because the Central Limit Theorem guarantees that sample means are approximately normal for large samples, regardless of the population shape.


Key Terms

Probability density function (PDF)

f(x), a function where the area under the curve between two points gives the probability that the random variable falls in that interval. The PDF itself is not a probability; only areas under it are.

Cumulative distribution function (CDF)

F(x) = P(X ≤ x) = ∫ from –∞ to x of f(t) dt. It gives the accumulated probability up to a value x. Think of it as a running total of probability from left to right.

Uniform distribution

A continuous distribution where every value in the interval [a, b] is equally likely. The PDF is a flat rectangle at height 1/(b – a).

Exponential distribution

Models the time between events in a Poisson process. If events arrive at rate λ per unit time, the waiting time between them follows Exp(λ). Mean = 1/λ, standard deviation = 1/λ.

Normal distribution

The bell-shaped distribution defined by mean μ and standard deviation σ. Written as X ~ N(μ, σ²). Its PDF is f(x) = (1 / (σ√(2π))) × e^(–(x–μ)² / (2σ²)).

Z-score

z = (x – μ) / σ. It measures how many standard deviations a value is from the mean. Converting to z-scores lets you use the standard normal table.

Parameter

A numerical descriptive measure of a population (e.g., the true population mean μ). Fixed but usually unknown.

Statistic

A quantity computed from sample values (e.g., the sample mean x̄). It varies from sample to sample.

Sampling distribution

The probability distribution of a statistic over all possible samples of the same size from the same population. In simple terms, it tells you how much a statistic (like x̄) would bounce around if you kept taking new samples.

Central Limit Theorem (CLT)

For a sufficiently large sample size, the sampling distribution of the sample mean is approximately normal, regardless of the shape of the population distribution. The mean of the sampling distribution equals the population mean, and the standard deviation equals σ / √n.


Core Content

Continuous Probability Distributions: General Rules

  • P(a < X < b) = ∫ from a to b of f(x) dx (area under the PDF).

  • You cannot find the probability at a single exact point; P(X = c) = 0 for any specific value c.

  • The CDF is F(x) = P(X ≤ x). To find P(X > x), compute 1 – F(x).

  • To find a percentile, solve F(y) = p for y. This is the reverse process of finding a probability.

Building a CDF from a Piecewise PDF

When the PDF is defined in pieces:

  • Integrate each piece separately.

  • For any x below the lower bound, F(x) = 0.

  • For any x above the upper bound, F(x) = 1.

  • At each boundary, add the maximum CDF value of the previous piece to the current piece (the CDF is cumulative, so it must be continuous and non-decreasing).

Uniform Distribution: X ~ Uniform(a, b)

  • PDF: f(x) = 1/(b – a) for a ≤ x ≤ b, and 0 otherwise.

  • CDF: F(x) = (x – a)/(b – a) for a ≤ x < b.

  • Mean: E(X) = (a + b) / 2

  • Standard deviation: σ = √[(b – a)² / 12]

  • Probabilities reduce to length ratios, so geometry works here: P(c < X < d) = (d – c)/(b – a).

Worked Example: Uniform

A packaging line produces cartons with weights uniformly distributed between 18.2 and 20.4 lbs.

  • P(X < 20) = (20 – 18.2) / (20.4 – 18.2) = 1.8 / 2.2 ≈ 0.818

  • Mean = (18.2 + 20.4) / 2 = 19.3

  • σ = (20.4 – 18.2) / √12 ≈ 0.635

Exponential Distribution: X ~ Exp(λ)

  • PDF: f(x) = λe^(–λx) for x ≥ 0.

  • CDF: F(x) = 1 – e^(–λx) for x ≥ 0.

  • Mean: E(X) = 1/λ

  • Standard deviation: σ = 1/λ (mean and SD are equal)

  • Memoryless property: P(X > s + t | X > s) = P(X > t). The remaining wait time does not depend on how long you have already waited.

Worked Example: Exponential

A machine repair takes an average of 2 hours. Repair time is exponential with λ = 0.5.

  • P(X > 1.5) = 1 – F(1.5) = 1 – (1 – e^(–0.5 × 1.5)) = e^(–0.75) ≈ 0.4724

Normal Distribution: X ~ N(μ, σ²)

  • The bell curve: symmetric about μ, with spread controlled by σ.

  • To find probabilities, convert to z-scores: z = (x – μ) / σ, then use the standard normal (Z) table.

  • Z-table values give the area to the left of the z-score: P(Z ≤ z).

  • For P(Z > z), compute 1 – P(Z ≤ z).

  • For P(a < Z < b), compute P(Z ≤ b) – P(Z ≤ a).

  • Always include the area "all the way until z" when reading the table; the table gives cumulative left-tail area.

Parameters vs. Statistics

  • A parameter describes a population (μ, σ, p). It is fixed.

  • A statistic describes a sample (x̄, s, p̂). It varies from sample to sample.

  • The goal of inferential statistics is to use statistics to estimate parameters.

Population Distribution vs. Sampling Distribution

  • The population distribution describes the values of the variable across all individuals in the population.

  • The sampling distribution describes how a statistic (like x̄) behaves across all possible samples of size n.

  • These are different things. Even if the population is skewed, the sampling distribution of x̄ can be approximately normal (thanks to the CLT).

Sampling Distribution of the Sample Mean

  • Mean of x̄: μ_{x̄} = μ (the sampling distribution is centred on the true population mean).

  • Standard deviation of x̄: σ_{x̄} = σ / √n (larger samples produce less variability in x̄).

  • If the population is normal, x̄ is exactly normal for any sample size.

  • If the population is not normal, x̄ is approximately normal for large n (the Central Limit Theorem).

Central Limit Theorem (CLT)

The CLT says: regardless of the population's shape, if n is "large enough," the distribution of x̄ is approximately N(μ, σ²/n).

  • A common rule of thumb is n ≥ 30, but this depends on how non-normal the population is. Highly skewed populations need larger samples.

  • The CLT is the reason the normal distribution is so dominant in statistics. Most of the inference you do later relies on it.

Worked Example: CLT

An electronics company makes resistors with μ = 100 ohms and σ = 10 ohms. For a sample of n = 25:

  • σ_{x̄} = 10 / √25 = 2

  • P(x̄ < 95) = P(Z < (95 – 100)/2) = P(Z < –2.5) ≈ 0.0062

Worked Example: Uniform → CLT

A bus arrives every 10 minutes. Wait time is Uniform(0, 10).

  • For one student: P(X < 6) = 6/10 = 0.6

  • For 40 students (their average wait): μ = 5, σ = 10/√12 ≈ 2.887, σ_{x̄} = 2.887/√40 ≈ 0.4564

  • P(x̄ < 6) = P(Z < (6 – 5)/0.4564) = P(Z < 2.19) ≈ 0.9857


Formulas

Item

Formula

PDF to probability

P(a < X < b) = ∫ₐᵇ f(x) dx

CDF

F(x) = P(X ≤ x) = ∫₋∞ˣ f(t) dt

Uniform PDF

f(x) = 1/(b–a), a ≤ x ≤ b

Uniform mean

(a+b)/2

Uniform SD

(b–a)/√12

Exponential PDF

f(x) = λe^(–λx), x ≥ 0

Exponential CDF

F(x) = 1 – e^(–λx)

Exponential mean & SD

Both 1/λ

Normal PDF

f(x) = (1/(σ√(2π))) e^(–(x–μ)²/(2σ²))

Z-score

z = (x – μ)/σ

Sampling dist. of x̄: mean

μ_{x̄} = μ

Sampling dist. of x̄: SD

σ_{x̄} = σ/√n


Real-World Applications

The exponential distribution is the standard model for "time until the next event" in reliability engineering: how long until a server fails, how long until the next customer arrives. The CLT is why opinion polls work. Even though individual voter preferences are binary (not normal at all), the average across a large sample is approximately normal, which lets you build confidence intervals and margins of error from a single poll.


Common Misconceptions

  • "The PDF gives the probability at a point." It does not. For continuous distributions, P(X = c) = 0. The PDF gives probability density; only the area under the curve (an integral) gives an actual probability.

  • "P(X ≤ 5) and P(X < 5) are different for continuous variables." They are equal. Since P(X = 5) = 0 for a continuous variable, including or excluding the endpoint changes nothing.

  • "The CLT says the population becomes normal." It does not. The population distribution stays whatever it is. The CLT says the sampling distribution of the sample mean becomes approximately normal.

  • "A larger sample makes each individual observation more normal." No. A larger sample makes the distribution of the sample mean closer to normal. Individual observations still follow the population distribution.


Why It Matters / Exam Flags

⚠️ Know the PDF, CDF, mean, and standard deviation for all three continuous distributions (uniform, exponential, normal). These are not always given on the formula sheet.

⚠️ For exponential problems, check whether the rate λ or the mean 1/λ is given. Mixing these up is one of the most common errors.

⚠️ When computing probabilities with the normal distribution, always convert to z-scores first and read the table from the correct direction.

⚠️ CLT problems require you to use σ/√n, not σ, as the standard deviation. Forgetting to divide by √n is the single most frequent exam mistake on sampling-distribution questions.

⚠️ Be comfortable going from a piecewise PDF to a CDF by integrating each piece and chaining the boundary values.


Quick Self-Test

True or false: For a continuous random variable, P(X = 3) could equal 0.2.

A: False. For continuous variables, the probability at any single point is always 0.

Fill in the blank: The standard deviation of the sampling distribution of x̄ is σ divided by ___.

A: √n

True or false: The Central Limit Theorem requires the population to be normally distributed.

A: False. The CLT works for any population shape, as long as the sample size is large enough.

Fill in the blank: For X ~ Exp(λ), the mean is ___ and the variance is ___.

A: 1/λ and 1/λ² (equivalently, the standard deviation is 1/λ)


Practice Q&A

Q: X is uniformly distributed on [2, 10]. What is P(X > 7)?

A: P(X > 7) = (10 – 7)/(10 – 2) = 3/8 = 0.375.

Q: The time between arrivals at a ticket counter follows an exponential distribution with a mean of 4 minutes. What is the probability that the next arrival takes more than 5 minutes?

A: λ = 1/4 = 0.25. P(X > 5) = e^(–0.25 × 5) = e^(–1.25) ≈ 0.2865.

Q: A population has μ = 50 and σ = 12. A random sample of n = 36 is drawn. What is the probability that the sample mean is between 48 and 52?

A: σ_{x̄} = 12/√36 = 2. P(48 < x̄ < 52) = P((48–50)/2 < Z < (52–50)/2) = P(–1 < Z < 1) ≈ 0.6827.

Q: Why can you use the normal distribution to approximate the sampling distribution of x̄ even when the population is skewed?

A: The Central Limit Theorem guarantees that for sufficiently large n, the sampling distribution of x̄ is approximately normal regardless of the population's shape.

Q: A continuous random variable has PDF f(x) = x/2 for 0 ≤ x ≤ 2 and 0 elsewhere. Find E(X).

A: E(X) = ∫₀² x · (x/2) dx = ∫₀² x²/2 dx = [x³/6]₀² = 8/6 = 4/3 ≈ 1.333.


Connections to Other Topics

The normal distribution and the CLT are the foundation for confidence intervals and hypothesis tests, which make up most of the second half of an intro stats course. The exponential distribution connects back to the Poisson: if events occur at rate λ per unit time (Poisson process), the time between events is Exp(λ). Understanding CDFs also sets you up for working with p-values, which are just tail areas under a distribution.


Related Terms / Search Tags

continuous random variable, PDF, probability density function, CDF, cumulative distribution function, uniform distribution, exponential distribution, normal distribution, bell curve, z-score, standard normal, z-table, percentile, parameter, statistic, sampling distribution, standard error, Central Limit Theorem, CLT, sigma over root n, piecewise PDF, memoryless property, Purdue STAT 301, intro to statistics midterm 1