Difficulty: Intermediate to Advanced | Prerequisites: Ch. 6–7 (normal distribution, sampling distributions, CLT)
Confidence intervals and hypothesis tests are the two main tools of statistical inference. A confidence interval gives a range of plausible values for a population parameter. A hypothesis test asks whether the data provide enough evidence to reject a specific claim about that parameter. Together they form the backbone of Chapters 10–12. If you are shaky on the logic of p-values, the difference between Type I and Type II errors, or when to use a z-test vs. a t-test, everything that follows will be harder than it needs to be.
A confidence interval is estimate ± margin of error. Use z when σ is known; use t when σ is unknown. A hypothesis test compares sample data to a null hypothesis H₀ using a test statistic and p-value. If p ≤ α, reject H₀. Type I error is a false rejection; Type II error is a failure to reject a false null.
Point estimate
A single number computed from a sample that serves as a best guess for the population parameter.
Estimator
A statistic used to estimate a parameter. It is a random variable with its own distribution, mean and variance. An estimate is one specific value of the estimator.
Unbiased estimator
An estimator whose expected value equals the parameter it estimates: E(θ̂) = θ. Among all unbiased estimators, the one with the smallest variance is the minimum variance unbiased estimator (MVUE).
Confidence interval (CI)
An interval of values, estimate ± margin of error, constructed so that in repeated sampling it captures the true parameter a specified percentage (C) of the time.
Confidence coefficient (C)
The probability that the CI captures the true parameter in repeated sampling. Expressed as a percentage, this is the confidence level.
Margin of error
The ± part of a confidence interval. For a z-interval: z_{α/2} · σ/√n.
t-distribution
A symmetric, bell-shaped distribution with heavier tails than the standard normal. Used when σ is unknown and replaced by s. Specified by degrees of freedom ν = n – 1.
Null hypothesis (H₀)
The default claim assumed to be true until evidence says otherwise. Always contains an equals sign.
Alternative hypothesis (Hₐ)
The claim that contradicts H₀. Can be one-sided (upper tail: μ > μ₀, or lower tail: μ < μ₀) or two-sided (μ ≠ μ₀).
Test statistic
A number computed from sample data that measures how far the data diverge from what H₀ predicts.
p-value
The smallest significance level at which H₀ can be rejected. Equivalently, the probability of observing a test statistic as extreme as, or more extreme than, the one observed, assuming H₀ is true.
Significance level (α)
The threshold for rejection. If p ≤ α, the result is statistically significant and H₀ is rejected.
Type I error
Rejecting H₀ when it is actually true. P(Type I error) = α.
Type II error
Failing to reject H₀ when it is actually false. P(Type II error) = β.
Power
1 – β. The probability of correctly rejecting a false H₀. Higher power means the test is better at detecting a real effect.
A point estimate is a single number; a confidence interval adds a range of uncertainty
An unbiased estimator has E(θ̂) = θ. Among unbiased estimators, prefer the one with the smallest variance (MVUE)
Assumptions: SRS, population is normal (or approximately normal), σ is known, μ is unknown
CI: x̄ ± z_{α/2} · σ/√n
Required sample size for a given margin of error: n = (z_{α/2} · σ / ME)²
Higher confidence level means wider interval (lower precision). To increase precision: lower C, reduce σ, or increase n
One-sided bounds: upper bound μ < x̄ + z_α · σ/√n; lower bound μ > x̄ – z_α · σ/√n
Replace σ with s and z with t: CI = x̄ ± t_{α/2, n–1} · s/√n
Degrees of freedom: ν = n – 1. Always round down
The t procedure is robust against non-normality:
n < 15: population should be close to normal
15 < n < 40: mild skewness is acceptable
n > 40: procedure is usually valid regardless
Four parts: state H₀ and Hₐ, compute the test statistic, find the p-value, make a decision
H₀ always has an equals sign. Hₐ defines the direction of the test
Decision rule: if p ≤ α, reject H₀ and conclude Hₐ in context. If p > α, fail to reject H₀ (you do not "accept" H₀)
A non-significant result means the data are consistent with the null, not that the null is proven true
Type I error (α): false positive, rejecting a true H₀
Type II error (β): false negative, failing to reject a false H₀
Power = 1 – β. Power increases when: α increases, the true parameter is farther from μ₀, σ decreases, n increases
To find β: find the rejection boundary using z_α, then compute the probability of falling in the non-rejection region under the true parameter value
z-test (when σ is known): z_ts = (x̄ – μ₀) / (σ/√n)
t-test (when σ is unknown): t_ts = (x̄ – μ₀) / (s/√n), df = n – 1
p-value depends on the direction of Hₐ:
Upper-tailed: P(Z ≥ z_ts) or P(T ≥ t_ts)
Lower-tailed: P(Z ≤ z_ts) or P(T ≤ t_ts)
Two-tailed: 2P(Z ≥ |z_ts|) or 2P(T ≥ |t_ts|)
Procedure: identify the parameter, state hypotheses, compute test statistic, find p-value, state decision and conclusion in context
\text{CI (z): } \bar{x} \pm z_{\alpha/2}\,\frac{\sigma}{\sqrt{n}}n = \left(\frac{z_{\alpha/2}\,\sigma}{\text{ME}}\right)^2\text{CI (t): } \bar{x} \pm t_{\alpha/2,\,n-1}\,\frac{s}{\sqrt{n}}z_{ts} = \frac{\bar{x} - \mu_0}{\sigma / \sqrt{n}}t_{ts} = \frac{\bar{x} - \mu_0}{s / \sqrt{n}}, \quad df = n - 1"95% confidence" does not mean there is a 95% probability that the true parameter is inside this particular interval. It means that 95% of intervals constructed this way, across repeated samples, would capture the true parameter
Students often say "accept H₀" when they fail to reject. The correct language is "fail to reject H₀" or "the data are consistent with H₀." Failing to reject does not prove H₀ is true
A small p-value does not measure the size of the effect. It measures the strength of evidence against H₀. A tiny p-value with a trivially small effect can occur in a large sample
The t-distribution and the z-distribution are not interchangeable. Use t when σ is estimated by s; use z when σ is known. As n grows large, t approaches z
⚠️ Know which test to use: z-test (σ known) vs. t-test (σ unknown). This is a common decision point.
⚠️ Be able to state conclusions in context: "The data (do/do not) provide sufficient evidence at the α level to conclude that [Hₐ in words]."
⚠️ Understand the relationship between CI width, confidence level, sample size and σ.
⚠️ Power and β calculations come up. Know which factors increase power.
True or False: A 99% CI is narrower than a 95% CI for the same data. (False – it is wider)
Fill in the blank: The degrees of freedom for a one-sample t-test are ___. (n – 1)
True or False: If p = 0.03 and α = 0.05, you reject H₀. (True)
True or False: Type II error is rejecting a true null hypothesis. (False – that is Type I)
Fill in the blank: Power = ___. (1 – β)
Q: A sample of 25 has x̄ = 82 and s = 10. Construct a 95% CI for μ.
A: df = 24. t_{0.025, 24} ≈ 2.064. CI = 82 ± 2.064(10/√25) = 82 ± 2.064(2) = 82 ± 4.128 = (77.87, 86.13).
Q: H₀: μ = 50, Hₐ: μ > 50. Sample gives x̄ = 53, σ = 6, n = 36. Find the test statistic and p-value.
A: z = (53 – 50)/(6/√36) = 3/1 = 3.0. p-value = P(Z ≥ 3.0) ≈ 0.0013. Reject H₀ at any common α.
Q: What happens to the width of a CI if you double the sample size?
A: The margin of error is proportional to 1/√n. Doubling n reduces the margin of error by a factor of √2 ≈ 1.41, so the interval gets narrower.
Q: Explain Type I and Type II error in the context of a court trial.
A: Type I error: convicting an innocent person (rejecting a true H₀). Type II error: acquitting a guilty person (failing to reject a false H₀).
The z-CI and z-test from Ch. 8–9 extend to two-sample problems in Ch. 10 and to proportions. The t-test reappears in paired data (Ch. 10.3) and regression (Ch. 12). ANOVA (Ch. 11) generalises the two-sample t-test to k groups. The logic of hypothesis testing (H₀, Hₐ, p-value, decision) is the same framework used in every inferential chapter.
Point estimate, estimator, unbiased, MVUE, confidence interval, CI, margin of error, confidence level, confidence coefficient, z-interval, t-interval, t-distribution, degrees of freedom, null hypothesis, alternative hypothesis, test statistic, p-value, significance level, alpha, Type I error, Type II error, power, beta, z-test, t-test, one-tailed, two-tailed, robust, STAT 35000, Purdue statistics