Difficulty: Intermediate | Prerequisites: Confidence intervals (Ch. 8–9), sampling distributions, z and t tables
Hypothesis testing is a four-step procedure for using sample data to decide whether there is enough evidence to reject a claim about a population parameter. You state hypotheses, compute a test statistic, find a p-value, and make a decision. The same one-sample/two-sample/paired framework from confidence intervals applies here, but instead of estimating a range, you are testing a specific claim.
Null hypothesis (H₀)
The default claim about the population parameter, typically a statement of "no effect" or "no difference." This is what you assume is true until the data provide strong evidence against it.
Alternative hypothesis (Hₐ)
The claim you are trying to find evidence for. It can be one-sided (greater than or less than) or two-sided (not equal to).
Test statistic
A standardised value that measures how far your sample result is from the null hypothesis value, in units of standard error. Think of it as: how many standard errors is my sample from what H₀ claims?
P-value
The probability of observing a test statistic as extreme as (or more extreme than) the one calculated, assuming H₀ is true. A small p-value means your data are unlikely under H₀. In simple terms, it measures the strength of the evidence against H₀.
Significance level (α)
The threshold you set before collecting data for deciding when to reject H₀. Common choices are 0.05, 0.01, and 0.10. If the p-value ≤ α, you reject H₀.
Type I error
Rejecting H₀ when it is true (a false positive). The probability of a Type I error is α.
Type II error (β)
Failing to reject H₀ when it is false (a false negative). The probability depends on the true parameter value, sample size, and α.
Power (1 – β)
The probability of correctly rejecting H₀ when it is false. Think of it as the test's ability to detect a real effect.
Statistical significance
A result is statistically significant if the p-value is less than or equal to α. It means the observed effect is unlikely to have occurred by chance alone, under H₀.
Practical significance
Whether the size of the observed effect is large enough to matter in the real world. A result can be statistically significant but practically unimportant, especially with very large samples.
Δ₀ (Delta-naught)
The hypothesised difference in the two-sample cases. It appears in the null hypothesis for the difference between two means.
Tags: null hypothesis, alternative hypothesis, test statistic, p-value, alpha, significance level, Type I error, Type II error, power, statistical significance, practical significance, delta-naught
Name the population parameter you are testing (e.g. μ = the true mean weight of cereal boxes). Describe it in the context of the problem, not just as a symbol.
Write H₀ and Hₐ using proper notation.
Two-tailed: H₀: μ = μ₀ vs. Hₐ: μ ≠ μ₀
Upper-tailed: H₀: μ = μ₀ vs. Hₐ: μ > μ₀
Lower-tailed: H₀: μ = μ₀ vs. Hₐ: μ < μ₀
For two-sample problems, replace μ with μ₁ – μ₂ and μ₀ with Δ₀.
Pick the right formula (see test statistics section below), compute df where needed, and find the p-value.
For z tests, you calculate the p-value yourself using the z table.
For t tests, the p-value will be given to you on the exam.
Compare the p-value to α.
If p-value ≤ α: reject H₀. State the reason ("Since p-value = ___ ≤ α = ___, we reject H₀") and write the conclusion in the context of the problem.
If p-value > α: fail to reject H₀. State the reason and conclude that there is not sufficient evidence to support the alternative.
The wording matters. You never "accept H₀." You either reject it or fail to reject it.
Tags: four steps, hypothesis testing procedure, null hypothesis, alternative hypothesis, test statistic, p-value, decision, conclusion, reject, fail to reject
z_ts = (x̄ – μ₀) / (σ / √n)
Degrees of freedom: N/A.
t_ts = (x̄ – μ₀) / (s / √n)
Degrees of freedom: df = n – 1.
t'_ts = (x̄₁ – x̄₂ – Δ₀) / √(s₁²/n₁ + s₂²/n₂)
Degrees of freedom: will be given to you.
t_ts = (d̄ – Δ₀) / (s_D / √n)
Degrees of freedom: df = n – 1 (where n = number of pairs).
The logic is identical to choosing the right CI formula. Ask: Is σ known? How many samples? Are they paired?
Tags: z test statistic, t test statistic, one-sample, two-sample, independent, matched pairs, paired t test
Alternative | P-value formula |
|---|---|
Upper-tailed: Hₐ: μ > μ₀ | P(Z > z_ts) = 1 – P(Z ≤ z_ts) |
Lower-tailed: Hₐ: μ < μ₀ | P(Z < z_ts) |
Two-tailed: Hₐ: μ ≠ μ₀ | 2 · P(Z > |
For t tests, you need to know the equations but the actual p-value will be provided on the exam.
The p-value is the probability of getting data as extreme as (or more extreme than) what you observed, assuming H₀ is true. It is not the probability that H₀ is true.
Type I error: Rejecting H₀ when H₀ is true. This can only happen when you reject. Its probability is α.
Type II error: Failing to reject H₀ when H₀ is false. This can only happen when you fail to reject. Its probability is β.
You can only make one type of error per test. If you rejected, the only possible error is Type I. If you failed to reject, the only possible error is Type II.
Power = 1 – β. It is the probability of correctly rejecting H₀ when a specific alternative is true.
Factors that increase power:
Larger sample size (n)
Larger true effect size (the farther μ_a is from μ₀)
Larger α (but this also increases Type I error risk)
Smaller variability (σ)
To calculate power, you need Hₐ, α, and the true value of the mean μ_a.
A two-sided CI and a two-sided test at the same α give consistent results. If μ₀ falls outside the (1 – α) CI, the test rejects H₀. If μ₀ falls inside the CI, the test fails to reject.
The CI gives more information: it shows you the range of plausible values, not just a yes/no decision.
Statistical significance means the result is unlikely under H₀ (p-value ≤ α). Practical significance asks whether the observed effect is large enough to matter.
With a very large sample, even a tiny, unimportant difference can be statistically significant. Always consider the actual size of the effect, not just whether the p-value crossed the threshold.
Tags: p-value calculation, Type I error, Type II error, power, alpha, beta, statistical significance, practical significance, CI vs hypothesis test
All test statistics follow the same pattern:
test statistic = (estimate – null value) / standard error
Case | Estimate | Null value | Standard error |
|---|---|---|---|
One-sample z | x̄ | μ₀ | σ/√n |
One-sample t | x̄ | μ₀ | s/√n |
Two-sample independent t | x̄₁ – x̄₂ | Δ₀ | √(s₁²/n₁ + s₂²/n₂) |
Matched pairs t | d̄ | Δ₀ | s_D/√n |
P-value directions:
Upper-tailed: area to the right of the test statistic
Lower-tailed: area to the left of the test statistic
Two-tailed: double the area in one tail
"The p-value is the probability that H₀ is true." No. It is the probability of observing data this extreme assuming H₀ is true. Those are very different statements.
"Failing to reject H₀ means H₀ is true." Failing to reject simply means you did not find enough evidence against H₀. Absence of evidence is not evidence of absence.
"A smaller p-value means a larger effect." Not necessarily. P-values depend on both the effect size and the sample size. A large sample can produce a tiny p-value for a trivially small effect.
Confusing Type I and Type II errors. Remember: Type I = false alarm (you rejected but should not have). Type II = missed detection (you failed to reject but should have). Only one error type is possible per test outcome.
Thinking statistical significance automatically implies practical importance. Always evaluate whether the magnitude of the effect matters in context.
⚠️ All four steps must appear in your answer. Missing any step loses marks.
⚠️ Step 4 requires three things: the decision (reject or fail to reject), the reason with numbers (p-value compared to α), and the conclusion in context.
⚠️ For z tests, you must calculate the p-value yourself. For t tests, it will be given.
⚠️ Know how to state Type I and Type II errors in context. "A Type I error would mean we concluded [specific claim] when in reality [null was true]." Use the problem's words, not generic symbols.
⚠️ Be ready to predict: given a CI result, what would the hypothesis test conclude, and vice versa?
⚠️ Power questions: know which factors increase or decrease power, and be able to calculate it given Hₐ, α, and μ_a.
True or False: If p-value = 0.03 and α = 0.05, you reject H₀. (True.)
True or False: A Type II error can occur when you reject H₀. (False. Type II can only occur when you fail to reject.)
Fill in the blank: Power increases as the true mean moves ___ from the null value. (farther)
True or False: "Fail to reject H₀" is the same as "accept H₀." (False. We never accept H₀.)
Fill in the blank: For a two-tailed z test, the p-value is ___ times the one-tail area. (2)
Q: A cereal company claims the mean box weight is 500g. A sample of 40 boxes gives x̄ = 497g and s = 10g. Test at α = 0.05 whether the mean differs from 500g. What are the hypotheses?
A: H₀: μ = 500 vs. Hₐ: μ ≠ 500 (two-tailed). The parameter μ is the true mean weight of all cereal boxes.
Q: Using the cereal example above, calculate the test statistic.
A: This is a one-sample t test (σ unknown). t_ts = (497 – 500) / (10/√40) = –3 / 1.581 = –1.897, with df = 39.
Q: If the p-value for the cereal test is 0.065, what is your decision at α = 0.05?
A: Since p-value = 0.065 > α = 0.05, we fail to reject H₀. There is not sufficient evidence to conclude that the mean box weight differs from 500g.
Q: In the cereal example, which error could you have made, and what would it mean in context?
A: Since we failed to reject H₀, the only possible error is Type II: failing to detect that the true mean weight differs from 500g when it does.
Q: A 95% CI for μ₁ – μ₂ is (–2.3, 5.1). Would a two-sided test of H₀: μ₁ – μ₂ = 0 reject at α = 0.05?
A: No. Zero is inside the interval, so the test would fail to reject H₀ at α = 0.05.
Hypothesis testing builds directly on the confidence interval material. The same assumptions (normality, independence, random sampling) apply, and the same formula-selection logic (z vs. t, one-sample vs. two-sample vs. paired) carries over. If you are comfortable with CIs, the mechanics of testing are largely the same, just with a different question being asked.
The concepts of Type I error, power, and significance level will reappear in the ANOVA chapter, where the test is extended to compare more than two groups at once.
hypothesis testing, null hypothesis, alternative hypothesis, H0, Ha, test statistic, p-value, significance level, alpha, Type I error, Type II error, power, beta, z test, t test, one-sample, two-sample, independent, matched pairs, paired t, reject, fail to reject, statistical significance, practical significance, delta-naught, four-step procedure, decision rule, CI vs hypothesis test