Difficulty: Intermediate | Prerequisites: Basic probability, normal distribution (z-scores), expected value and variance rules.
This chapter bridges the gap between describing a single sample and making inferences about an entire population. You need a solid grasp of normal distributions and z-scores before diving in. The core idea: when you repeatedly draw samples of the same size from a population and compute their means, those means form their own distribution with predictable properties. The Central Limit Theorem then guarantees that this distribution of sample means is approximately normal for large samples, regardless of the population's shape. This is the theoretical engine behind confidence intervals and hypothesis tests in later chapters.
A sampling distribution tracks how a statistic (like the sample mean) varies across every possible same-sized sample from a population. Its mean equals the population mean, and its spread shrinks as sample size grows. The Central Limit Theorem says this distribution is approximately normal when n is large enough, even if the population itself is skewed or unusual.
Parameter
A number that describes a characteristic of the entire population. Denoted by Greek letters (e.g. μ for the population mean, σ for the population standard deviation). Parameters are usually unknown constants.
In simple terms, this is the "true answer" you are trying to estimate. You almost never know it exactly.
Statistic
Any number calculated from a sample. Denoted by Latin letters (e.g. x̄ for the sample mean, s for the sample standard deviation).
Think of it as your best guess at the parameter, computed from the data you collected.
Sampling distribution
The distribution of a statistic's values across all possible samples of the same size drawn from the same population. It describes the long-term behaviour of the statistic.
In simple terms, if you could repeat your study infinitely many times (same sample size each time), the histogram of all those sample means is the sampling distribution.
Standard error (of the mean)
The standard deviation of the sampling distribution of x̄, equal to σ / √n. It measures how much the sample mean typically varies from sample to sample.
Think of it as the "typical miss" when you use a single sample mean to estimate the population mean.
Central Limit Theorem (CLT)
A theorem stating that, for a sufficiently large sample size n, the sampling distribution of the sample mean is approximately normal, regardless of the shape of the population distribution.
In simple terms, even if the population is skewed, uniform or otherwise non-normal, average enough observations together and the distribution of those averages will look like a bell curve.
Parameters use Greek letters (μ, σ) and describe the population. They are fixed but usually unknown.
Statistics use Latin letters (x̄, s) and are calculated from sample data. They vary from sample to sample.
The goal of inference is to use a statistic to learn about the parameter.
Pick a sample of size n from a population. Compute your statistic (e.g. the mean). Repeat for every possible sample of that same size. The collection of all those statistic values is the sampling distribution.
"Same size" matters because the value of the statistic can depend on n.
This is a theoretical idea: you do not literally draw every possible sample, but the concept underpins all of inference.
The mean of the sampling distribution equals the population mean: μ(x̄) = μ.
This means x̄ is an unbiased estimator of μ.
The standard deviation of the sampling distribution (standard error) is σ(x̄) = σ / √n.
As n increases, the standard error decreases, so the sample mean clusters more tightly around μ.
If the population itself is normal, then the sampling distribution of x̄ is exactly normal for any n.
Given a population with mean μ and standard deviation σ, for a sample of size n:
Standard error of x̄ = σ / √n
Example: If μ = 1.5 min, σ = 0.35 min and n = 5, then σ(x̄) = 0.35 / √5 = 0.1565 min.
When the population is normal (or n is large enough for the CLT), standardise and use the z-table:
z = (x̄ – μ) / (σ / √n)
Worked example (normal population, n = 5, μ = 1.5, σ = 0.35):
P(x̄ ≤ 2.0) = P(Z ≤ (2.0 – 1.5) / 0.1565) = P(Z ≤ 3.19) = 0.9993
P(x̄ within 0.3 of the mean) = P(1.2 ≤ x̄ ≤ 1.8) = P(−1.92 ≤ Z ≤ 1.92) = 0.9726 – 0.0274 = 0.9452
A common exam set-up gives you both scenarios:
One resistor: P(X < 95) with μ = 100, σ = 10 → z = (95 – 100) / 10 = −0.50, P = 0.3085. This is NOT a sampling distribution problem.
Sample of 25 resistors: P(x̄ < 95) → z = (95 – 100) / (10 / √25) = −2.5, P = 0.0062. Now the denominator is the standard error.
The probability for the sample mean is much smaller because the standard error (10 / 5 = 2) is much tighter than the population SD (10).
If the population is normal, the sampling distribution of x̄ is exactly normal for any n. You do not need the CLT in that case.
If the population is NOT normal (skewed, uniform, discrete, etc.), the CLT tells you the sampling distribution of x̄ is approximately normal, provided n is large enough.
There is no single magic number. A common guideline is n ≥ 30, but the more skewed the population, the larger n needs to be.
The course shows histograms of sampling distributions at n = 2, 5, 10, 20 for normal, uniform and exponential populations. Even a strongly skewed exponential population produces a roughly bell-shaped sampling distribution by about n = 20.
The mean of the sampling distribution is still μ (unchanged).
The standard error is still σ / √n (unchanged).
The shape is what changes: it becomes approximately normal.
The CLT also applies to discrete populations. When n is large enough, the sampling distribution of the mean looks continuous and approximately normal.
Any linear combination of independent normal random variables is exactly normal, for any n.
Multiplying a normal distribution by a constant or adding normal distributions together does not change the distribution type.
A bus arrives every 10 minutes. Wait time is Uniform(0, 10), so μ = 5 and σ = √(8.333) = 2.887.
One student: P(X < 6) = 6/10 = 0.6 (straight from the uniform density).
40 students (sample mean): σ(x̄) = 2.887 / √40 = 0.4564. By the CLT, P(x̄ < 6) = P(Z < (6 – 5) / 0.4564) = P(Z < 2.19) = 0.9857.
The CLT made the uniform population irrelevant once n = 40 was large enough to justify normal approximation.
\mu_{\bar{X}} = \mu\sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}}Z = \frac{\bar{X} - \mu}{\sigma / \sqrt{n}}For a Uniform(a, b) population:
\mu = \frac{a+b}{2}, \quad \sigma = \frac{b-a}{\sqrt{12}}Students often confuse the population standard deviation (σ) with the standard error (σ / √n). The standard error is always smaller than σ (for n > 1) and applies to the distribution of the sample mean, not to individual observations.
Students sometimes apply the CLT when the population is already normal. If the population is normal, the sampling distribution of x̄ is exactly normal for every n. The CLT is needed only when the population is non-normal.
A common mistake is forgetting to check whether a problem asks about one observation (use σ) or a sample mean (use σ / √n). The wording "a randomly selected item" vs "the average of n items" is the signal.
Students often think the CLT makes the population itself become normal. It does not. The CLT only says the distribution of the sample mean approaches normality.
⚠️ Know when to use σ vs σ / √n. If the question says "one observation," use σ. If it says "sample mean" or "average of n," use σ / √n.
⚠️ Be prepared to state whether a problem requires the CLT or not. If the population is normal, the sampling distribution is exact, no CLT needed. If the population is non-normal and n is large, cite the CLT.
⚠️ The standard error formula σ / √n appears in confidence intervals and hypothesis tests in Chapters 9–11. Master it here.
⚠️ Exam questions often give a non-normal population (uniform, exponential) and ask you to compute a probability for a sample mean. You must recognise that the CLT applies and then standardise using the standard error.
True or False: The mean of the sampling distribution of x̄ is always equal to the population mean, regardless of sample size. (True)
True or False: The CLT says the population distribution becomes normal when n is large. (False, it says the sampling distribution of the mean becomes approximately normal.)
Fill in the blank: The standard error of the mean equals _____ divided by _____. (σ divided by √n)
True or False: If the population is already normal, you need the CLT to justify using z-scores for the sample mean. (False, the sampling distribution is exactly normal for any n.)
Fill in the blank: Increasing the sample size makes the standard error _____. (Smaller)
Q: A population has μ = 50 and σ = 12. If you draw samples of size 36, what is the standard error of the sampling distribution of x̄?
A: σ / √n = 12 / √36 = 12 / 6 = 2.
Q: Using the same population (μ = 50, σ = 12, n = 36), what is the probability that the sample mean exceeds 53?
A: z = (53 – 50) / 2 = 1.5. P(Z > 1.5) = 1 – 0.9332 = 0.0668.
Q: A population distribution is strongly right-skewed. A researcher takes a sample of n = 4. Can she assume the sampling distribution of x̄ is normal? Why or why not?
A: No. With n = 4 and a strongly skewed population, the sample size is too small for the CLT to apply. The sampling distribution will still reflect the skewness of the population.
Q: Waiting time at a petrol station is Uniform(0, 8) minutes. For a sample of 50 customers, find P(x̄ > 4.5).
A: μ = 4, σ = (8 – 0) / √12 = 2.309. Standard error = 2.309 / √50 = 0.3266. By the CLT, z = (4.5 – 4) / 0.3266 = 1.53. P(Z > 1.53) = 1 – 0.9370 = 0.0630.
Q: Why is the probability of a sample mean being far from μ much smaller than the probability of a single observation being far from μ?
A: The sample mean has a smaller spread (standard error = σ / √n) than individual observations (spread = σ). Averaging reduces variability, so extreme values of x̄ are rarer than extreme individual values.
This connects to confidence intervals (Ch. 9) because the standard error σ / √n appears directly in the margin of error formula. Without the sampling distribution, there is no basis for constructing an interval estimate.
This connects to hypothesis testing (Ch. 10) because the test statistic z = (x̄ – μ₀) / (σ / √n) is just a standardised value from the sampling distribution. The entire logic of p-values depends on knowing what the sampling distribution looks like under H₀.
The CLT also explains why the t-distribution (used when σ is unknown) works: for large n the t-distribution converges to the standard normal.
sampling distribution, standard error, standard error of the mean, central limit theorem, CLT, sample mean distribution, σ / √n, sigma over root n, law of large numbers, normal approximation, z-score for sample mean, population parameter vs sample statistic, unbiased estimator, expected value of x-bar, variance of x-bar, STAT 302, Purdue statistics