Difficulty: Intermediate-Advanced | Prerequisites: Chapters 1-6 study notes (descriptive statistics, probability, normal distribution, z-scores)
This is where the course shifts from describing data and calculating probabilities to making inferences about populations using samples. Chapter 7 introduces sampling distributions and the Central Limit Theorem, which explains why the normal distribution dominates inference. Chapter 8 uses sampling distributions to build confidence intervals (plausible ranges for a population mean). Chapter 9 uses them for hypothesis testing (deciding whether data supports or contradicts a claim). Together, these chapters form the core of inferential statistics and account for a large portion of the exam.
A sampling distribution describes how a statistic (like x-bar) varies across all possible samples. The Central Limit Theorem says this distribution is approximately normal for large n, regardless of the population shape. Confidence intervals give a range of plausible values for the population mean. Hypothesis tests check whether the data contradicts a specific claim about the population mean. Both rely on the same machinery: sampling distributions, z or t statistics, and assumptions about normality.
Parameter
A numerical descriptive measure of a population, indicated by Greek letters (μ, σ). Usually a fixed constant you do not know.
Statistic
Any quantity computed from sample values, indicated by Latin letters (x-bar, s). Think of it as your estimate of the parameter.
Sampling distribution
The probability distribution of a statistic across all possible samples of the same size from the same population. In simple terms, it answers the question: "If I repeated this sampling process infinitely many times, what would the distribution of my statistic look like?"
Central Limit Theorem (CLT)
If you draw samples of size n from any population with mean μ and finite variance σ², then for large enough n, the sampling distribution of x-bar is approximately N(μ, σ²/n). The rule of thumb is n ≥ 30, though this depends on how skewed the population is.
Standard error
The standard deviation of the sampling distribution: σ/√n. It measures how much the sample mean varies from sample to sample.
Point estimate
A single number computed from a sample that serves as a best guess for a population parameter.
Unbiased estimator
A statistic whose expected value equals the parameter it estimates. x-bar is an unbiased estimator of μ because E(x-bar) = μ.
Minimum Variance Unbiased Estimator (MVUE)
Among all unbiased estimators, the one with the smallest variance. x-bar is the MVUE for μ.
Confidence interval (CI)
An interval of plausible values for a population parameter, constructed so that in repeated sampling, a specified percentage of such intervals capture the true parameter.
Confidence coefficient (C)
The probability that the CI captures the true parameter in repeated sampling. Common values: 0.90, 0.95, 0.99.
Margin of error (ME)
The half-width of a confidence interval: z_(α/2) × σ/√n.
Critical value (z_(α/2))
The z-score that cuts off an area of α/2 in the right tail of the standard normal. For 95% confidence, z_(α/2) = 1.96.
t-distribution
Used in place of the z-distribution when the population standard deviation σ is unknown and estimated by s. It has heavier tails than the standard normal. The shape depends on degrees of freedom (df = n - 1). As df increases, the t-distribution approaches the standard normal.
Degrees of freedom (df)
For a single-sample t-test or CI: df = n - 1.
Null hypothesis (H₀)
The initial claim assumed to be true, the "status quo." It always contains an equals sign: H₀: μ = μ₀.
Alternative hypothesis (Hₐ)
The claim that contradicts H₀. It represents what you are trying to show. It can be one-sided (μ > μ₀ or μ < μ₀) or two-sided (μ ≠ μ₀).
Test statistic
A value calculated from sample data that measures how far the data diverge from what H₀ predicts. For the z-test: z_ts = (x-bar - μ₀) / (σ/√n). For the t-test: t_ts = (x-bar - μ₀) / (s/√n).
p-value
The probability, assuming H₀ is true, of observing a test statistic as extreme as or more extreme than the one actually observed. A small p-value provides evidence against H₀.
Significance level (α)
The maximum probability of a Type I error you are willing to accept. Set before looking at the data. Common value: 0.05.
Type I error
Rejecting H₀ when it is actually true. Probability = α.
Type II error
Failing to reject H₀ when it is actually false. Probability = β.
Power
The probability of correctly rejecting H₀ when it is false: Power = 1 - β. Higher power is better. Power increases as α increases, σ decreases, n increases, or the true mean moves further from μ₀.
Key results for x-bar:
Mean of the sampling distribution: μ_(x-bar) = μ (the sample mean is unbiased)
Standard deviation of the sampling distribution: σ_(x-bar) = σ/√n
As n increases, the sampling distribution becomes narrower, meaning sample means are less variable than individual observations
Shape of the sampling distribution:
If the population is normal, then x-bar is exactly normal: X-bar ~ N(μ, σ²/n)
If the population is not normal but n is large enough (typically n ≥ 30), the CLT says x-bar is approximately normal
Any linear combination of independent normal random variables is also normal
The CLT also applies to discrete random variables
Spotting a sampling distribution problem:
Look for words like "average", "mean", "sampling distribution of the mean"
In addition to the average, the problem must state the number of objects being averaged over (the sample size n)
Key difference from population distribution problems:
Individual observation: use σ in the denominator when standardising
Sample mean: use σ/√n in the denominator
Assumptions for inference:
You have a simple random sample (SRS) from the population of interest
The variable is normally distributed or approximately normal (if n ≥ 30, CLT applies)
CI when σ is known (z-interval):
Formula: x-bar ± z_(α/2) × σ/√n
The margin of error is ME = z_(α/2) × σ/√n
Common critical values: 90% CI uses z = 1.645, 95% CI uses z = 1.96, 99% CI uses z = 2.576
R code: z <- qnorm((1-C)/2, lower.tail=FALSE) then c(xbar - z*sigma/sqrt(n), xbar + z*sigma/sqrt(n))
CI when σ is unknown (t-interval):
Replace σ with s and z with t: x-bar ± t_(α/2, df) × s/√n
df = n - 1
R code: t <- qt((1-C)/2, df=n-1, lower.tail=FALSE) then c(xbar - t*s/sqrt(n), xbar + t*s/sqrt(n))
Or use t.test(data, conf.level=C)$conf.int
Interpreting a CI:
Correct: "We are 95% confident that the population mean is captured by the interval (a, b)"
Incorrect: "There is a 95% probability that μ lies in this interval" (μ is fixed, not random; the interval is the random part)
Each individual CI is either 100% correct or 0% correct. The 95% refers to the long-run success rate of the method
Precision of confidence intervals (three ways to reduce ME):
Lower the confidence level (reduce z_(α/2)), but this means less confidence
Reduce σ (improve experimental design), but this is often not under your control
Increase n (take a larger sample), the most practical option
Sample size determination:
To achieve a desired margin of error ME: n = (z_(α/2) × σ / ME)²
Always round up to the next whole number
Procedure for hypothesis testing (four steps):
Step 1: Identify the parameter and describe it in context
Step 2: State H₀ and Hₐ in symbols (do this before looking at the data)
Step 3: Calculate the test statistic and find the p-value
Step 4: Make the decision with reason and state the conclusion in context
Formulating hypotheses:
What you want to prove goes in the alternative hypothesis
H₀ always contains "="
One-sided upper tail: H₀: μ = μ₀, Hₐ: μ > μ₀
One-sided lower tail: H₀: μ = μ₀, Hₐ: μ < μ₀
Two-sided: H₀: μ = μ₀, Hₐ: μ ≠ μ₀
Use one-sided only if you believe before looking at the data that only one direction matters
Test statistic (z-test, σ known):
z_ts = (x-bar - μ₀) / (σ/√n)
Test statistic (t-test, σ unknown):
t_ts = (x-bar - μ₀) / (s/√n), df = n - 1
Calculating the p-value:
Right-tailed (Hₐ: μ > μ₀): p = P(Z ≥ z_ts) = pnorm(zts, lower.tail=FALSE)
Left-tailed (Hₐ: μ < μ₀): p = P(Z ≤ z_ts) = pnorm(zts)
Two-tailed (Hₐ: μ ≠ μ₀): p = 2 × P(Z ≤ -|z_ts|) = 2*pnorm(-abs(zts))
Decision rule:
p-value ≤ α: reject H₀, conclude Hₐ in context
p-value > α: fail to reject H₀, cannot conclude Hₐ in context
Never say "accept H₀." You fail to reject it
Writing the conclusion:
Must include: the decision (reject or fail to reject), the p-value, what the alternative hypothesis means in plain language
Template: "The data [does/does not] give [strong] support (p-value = [value]) to the claim that [statement of Hₐ in words]."
Do not use mathematical symbols in the written conclusion
Errors and power:
Type I error (α): rejecting H₀ when it is true
Type II error (β): failing to reject H₀ when it is false
Power = 1 - β: probability of correctly rejecting a false H₀
Power increases as: α increases, σ decreases, n increases, the distance between μ₀ and μₐ increases
To calculate power: (1) find the cutoff using α and the null distribution, (2) calculate β using the alternative mean and the cutoff, (3) power = 1 - β
Connection between CIs and hypothesis tests:
A two-sided hypothesis test at level α rejects H₀: μ = μ₀ if and only if μ₀ falls outside the (1 - α) confidence interval
Both use the same assumptions and the same sampling distribution machinery
Quantity | Formula |
|---|---|
Standard error of x-bar | σ/√n |
z-interval CI | x-bar ± z_(α/2) × σ/√n |
t-interval CI | x-bar ± t_(α/2, n-1) × s/√n |
Sample size for desired ME | n = (z_(α/2) × σ / ME)² |
z test statistic | z_ts = (x-bar - μ₀) / (σ/√n) |
t test statistic | t_ts = (x-bar - μ₀) / (s/√n), df = n - 1 |
Power | 1 - β, where β = P(fail to reject H₀ when H₀ is false) |
qnorm(alpha/2, lower.tail=FALSE) – z critical value
qt(alpha/2, df=n-1, lower.tail=FALSE) – t critical value
pnorm(zts, lower.tail=FALSE) – right-tail p-value for z
2*pnorm(-abs(zts)) – two-tail p-value for z
pt(tts, df=n-1, lower.tail=FALSE) – right-tail p-value for t
2*pt(-abs(tts), df=n-1) – two-tail p-value for t
t.test(data, mu=mu0, alternative="two.sided") – complete t-test
t.test(data, conf.level=0.95)$conf.int – t confidence interval
Confidence intervals and hypothesis tests are how pharmaceutical companies demonstrate that a new drug works (clinical trials), how manufacturers verify that products meet specifications (quality control), and how polling organisations estimate election outcomes (margin of error). The p-value approach has come under scrutiny: the American Statistical Association issued a statement in 2016 warning that p-values should not be the sole basis for scientific conclusions and are not the probability that the hypothesis is true.
Students often interpret a 95% CI as "there is a 95% chance μ is in this interval." That is wrong. μ is fixed. The interval is what varies from sample to sample. The correct interpretation is about the method's long-run success rate.
Students confuse "fail to reject H₀" with "accept H₀." Failing to reject does not prove H₀ is true. It means the evidence was insufficient to conclude otherwise.
Students sometimes use σ when they should use s (or vice versa). If the population standard deviation is unknown, you must use s and the t-distribution.
Students forget that the hypothesis test decision (reject or fail to reject) must be made by comparing the p-value to α, not by looking at the test statistic alone (in this course).
⚠️ Know when to use z vs. t. If σ is known, use z. If σ is unknown and estimated by s, use t with df = n - 1.
⚠️ Be able to check assumptions: SRS and normality (either stated, or n ≥ 30 for CLT).
⚠️ For CIs, know the correct interpretation. The interval is random, μ is fixed.
⚠️ For hypothesis tests, always write the conclusion in context and include the p-value.
⚠️ Never "accept H₀." The correct phrasing is "fail to reject H₀."
⚠️ Know how to calculate power: find the cutoff from α and the null distribution, then compute β under the alternative mean.
⚠️ The CLT is about the sampling distribution of x-bar, not about the population itself. The population distribution does not become normal as n increases.
True or false: The Central Limit Theorem says that the population distribution becomes normal as the sample size increases. False. The CLT says the sampling distribution of x-bar becomes approximately normal. The population distribution does not change.
Fill in the blank: For a 99% confidence interval using a z-distribution, z_(α/2) = ___. 2.576.
True or false: If the p-value is 0.03 and α = 0.05, we reject H₀. True. The p-value (0.03) is less than α (0.05).
Fill in the blank: The standard error of the sample mean is σ divided by ___. √n.
True or false: Increasing the sample size increases the power of a hypothesis test. True.
Q: A population has μ = 50 and σ = 10. A sample of n = 64 is taken. What is the probability that the sample mean exceeds 52?
A: σ_(x-bar) = 10/√64 = 1.25. z = (52 - 50)/1.25 = 1.60. P(Z > 1.60) = 1 - 0.9452 = 0.0548.
Q: A sample of 25 observations has x-bar = 34.2 and s = 4.1. Construct a 95% confidence interval for μ.
A: Since σ is unknown, use t with df = 24. t_(0.025, 24) ≈ 2.064. ME = 2.064 × 4.1/√25 = 2.064 × 0.82 = 1.69. CI = (34.2 - 1.69, 34.2 + 1.69) = (32.51, 35.89). We are 95% confident that the true population mean is captured by (32.51, 35.89).
Q: A manufacturer claims the mean weight of its bags of flour is 5 kg. A sample of 36 bags gives x-bar = 4.85 kg with a known σ = 0.5 kg. Test at α = 0.05 whether the mean differs from 5 kg.
A: H₀: μ = 5, Hₐ: μ ≠ 5. z_ts = (4.85 - 5)/(0.5/√36) = -0.15/0.0833 = -1.80. p-value = 2 × P(Z ≤ -1.80) = 2 × 0.0359 = 0.0718. Since 0.0718 > 0.05, fail to reject H₀. The data does not give sufficient support (p-value = 0.0718) to the claim that the mean weight of flour bags differs from 5 kg.
Q: With α = 0.05, σ = 10, n = 100, μ₀ = 50, and μₐ = 52, calculate the power of a one-sided upper-tail test.
A: σ_(x-bar) = 10/√100 = 1. z_(0.05) = 1.645. Cutoff = 50 + 1.645(1) = 51.645. β = P(X-bar ≤ 51.645 | μ = 52) = P(Z ≤ (51.645 - 52)/1) = P(Z ≤ -0.355) = 0.3613. Power = 1 - 0.3613 = 0.6387 (about 64%).
Confidence intervals and hypothesis tests extend to two-sample problems in Chapter 10 and to multiple groups in Chapter 11 (ANOVA). The t-distribution reappears in regression (Chapter 12) for testing whether slope coefficients are significant. The concept of power is important for study design: before collecting data, researchers calculate the sample size needed to achieve adequate power.
sampling distribution, Central Limit Theorem, CLT, standard error, point estimate, unbiased estimator, MVUE, confidence interval, CI, margin of error, critical value, z-interval, t-interval, t-distribution, degrees of freedom, null hypothesis, alternative hypothesis, test statistic, p-value, significance level, alpha, Type I error, Type II error, power, beta, reject, fail to reject, statistically significant, one-tailed test, two-tailed test, z-test, t-test, STAT 350, Purdue, introductory statistics