Difficulty: Intermediate | Prerequisites: Hypothesis testing (Ch. 9–10), confidence intervals, t and F distributions
One-way ANOVA lets you test whether the means of three or more groups are all equal, using a single F test instead of multiple pairwise t tests. If the F test rejects, you follow up with Tukey's HSD to identify which specific pairs of means differ. The mechanics revolve around partitioning total variability into between-group and within-group components.
ANOVA (Analysis of Variance)
A method for testing whether the means of three or more populations are equal by comparing the variability between groups to the variability within groups. In simple terms, it asks: is the spread between group averages larger than you would expect from random chance alone?
Factor
The categorical variable that defines the groups being compared (e.g. drug dosage, teaching method, fertiliser type).
Levels of the factor
The specific categories or values the factor takes (e.g. 0 mg, 20 mg, 40 mg). The number of levels is k.
Response variable
The quantitative outcome being measured (e.g. test score, blood pressure, plant height).
Sum of Squares for Factor A (SSA)
The variation in the data explained by differences between the group means. Also called the between-group sum of squares.
Sum of Squares for Error (SSE)
The variation within groups, not explained by the factor. Also called the within-group sum of squares.
Total Sum of Squares (SST)
The total variation in the data. SST = SSA + SSE.
Mean Square for Factor A (MSA)
SSA divided by its degrees of freedom (k – 1). It estimates the between-group variability.
Mean Square for Error (MSE)
SSE divided by its degrees of freedom (n – k). It estimates the within-group variability.
F statistic
F = MSA / MSE. If the group means are truly equal, F should be near 1. Large values of F provide evidence that at least one mean differs.
Tukey's HSD (Honestly Significant Difference)
A post-hoc multiple comparison procedure used after a significant ANOVA result to determine which specific pairs of group means differ. It controls the overall Type I error rate.
Tags: ANOVA, analysis of variance, factor, levels, response variable, SSA, SSE, SST, MSA, MSE, F statistic, F test, Tukey HSD, between-group, within-group, sum of squares, mean square
If you wanted to test whether three group means differ, you could run three separate t tests (group 1 vs. 2, 1 vs. 3, 2 vs. 3). The problem is that each test has its own Type I error rate, and the combined chance of at least one false positive grows quickly. With k groups, you would need k(k–1)/2 pairwise tests, and the overall error rate inflates well beyond α.
ANOVA solves this by using a single test. It compares the variability between groups to the variability within groups. If the between-group variability is much larger than the within-group variability, at least one group mean is different.
The F distribution lets you test all k means simultaneously while controlling the overall Type I error at α. A t test compares only two means. When k = 2, the F test and the two-sample t test give equivalent results (F = t²).
When reading a problem, identify:
Factor: what categorical variable defines the groups
Levels: the specific groups (k = number of levels)
Response variable: what is being measured
Populations: the k groups being compared
Total observations: n = n₁ + n₂ + ... + n_k
Independence: Observations are independent, both within and between groups.
Normality: Each population is approximately normally distributed. Check with QQ plots or histograms (you will not have to draw them; they will be provided).
Equal variances (homogeneity): The population standard deviations are approximately equal. Rule of thumb: the ratio of the largest sample standard deviation to the smallest should be less than 2.
Tags: ANOVA rationale, F vs t, factor, levels, response, assumptions, normality, equal variances, homogeneity, inflated Type I error, multiple comparisons problem
Source | df | SS | MS |
|---|---|---|---|
Factor A | k – 1 | SSA = Σ nᵢ(x̄ᵢ. – x̄..)² | MSA = SSA / (k – 1) |
Error | n – k | SSE = Σ(nᵢ – 1)sᵢ² | MSE = SSE / (n – k) |
Total | n – 1 | SST = SSA + SSE |
Key relationships you can use to fill in missing values:
SST = SSA + SSE
df_total = df_A + df_E, i.e. (n – 1) = (k – 1) + (n – k)
MS = SS / df for each row
On the exam, you will not compute raw summations. You will be given enough values to fill the rest in using addition, subtraction, multiplication, or division.
F = MSA / MSE
Degrees of freedom: df1 = k – 1 (numerator), df2 = n – k (denominator).
The F distribution is right-skewed and always positive. Large F values suggest group means differ.
Step 1: Parameter: the means μ₁, μ₂, ..., μ_k for all k populations.
Step 2: H₀: μ₁ = μ₂ = ... = μ_k vs. Hₐ: at least one mean is different.
Step 3: Calculate F = MSA / MSE with df1 = k – 1 and df2 = n – k. The p-value will be given.
Step 4: If p-value ≤ α, reject H₀ and conclude that at least one population mean differs. State this in context.
Note: rejecting H₀ tells you that the means are not all equal, but it does not tell you which ones differ. That is what Tukey's method is for.
Tags: ANOVA table, SSA, SSE, SST, MSA, MSE, F statistic, F test, df1, df2, one-way ANOVA hypothesis test
Use Tukey's method only after the ANOVA F test has rejected H₀. If the F test does not reject, you stop: there is no evidence that any means differ, so there is nothing to follow up on.
Running all possible pairwise t tests after ANOVA inflates the overall Type I error rate. Tukey's procedure adjusts the critical value so that the family-wise error rate stays at α across all comparisons.
For each pair of group means i and j:
x̄ᵢ. – x̄ⱼ. ± t**_column,df · SE
where:
t**column,df = q(α, k, n–k) / √2 (the Tukey critical value, derived from the Studentized Range Distribution)
SE = √(MSE · (1/nᵢ + 1/nⱼ)), which simplifies to √(2·MSE / nᵢ) when all groups have equal size
After computing all pairwise intervals, you present the results as a visual display: list the group means in ascending order on a line. Groups whose confidence interval for the difference includes zero are underlined together, meaning they are not significantly different.
Example from the review: 0 mg (57.60), 20 mg (69.28), 40 mg (75.70). If the interval for 20 mg vs. 40 mg contains zero, those two are underlined together. If the interval for 0 mg vs. 20 mg does not contain zero, they are not connected.
State the final results in plain language: "The 40 mg and 20 mg groups have significantly higher means than the control, but the 40 mg and 20 mg groups do not significantly differ from each other."
Tags: Tukey HSD, multiple comparisons, Studentized Range, q value, pairwise comparison, visual display, post-hoc test, family-wise error rate
Source | df | SS | MS |
|---|---|---|---|
Factor A | k – 1 | SSA | MSA = SSA/(k–1) |
Error | n – k | SSE | MSE = SSE/(n–k) |
Total | n – 1 | SST = SSA + SSE |
F = MSA / MSE, with df1 = k – 1, df2 = n – k
Critical value: t** = q_(α, k, n–k) / √2
Standard error: SE = √(MSE · (1/nᵢ + 1/nⱼ))
Confidence interval: (x̄ᵢ. – x̄ⱼ.) ± t** · SE
"ANOVA tells you which means differ." ANOVA only tells you that at least one mean is different. You need a follow-up procedure (Tukey's) to identify which pairs differ.
"You can skip ANOVA and go straight to pairwise t tests." Doing this inflates the overall Type I error rate. ANOVA controls the error for the omnibus test; Tukey's controls it for the pairwise comparisons.
"A large F always means practically important differences." As with any hypothesis test, statistical significance does not guarantee practical significance. Consider the actual sizes of the group mean differences.
"Equal variances means the standard deviations must be identical." The assumption is approximate. The rule of thumb is that the ratio of the largest to smallest sample standard deviation should be less than 2.
⚠️ You will need to fill in a partially completed ANOVA table using the relationships SST = SSA + SSE, df_total = df_A + df_E, and MS = SS / df.
⚠️ Know how to identify the factor, levels, response variable, and total observations from a word problem.
⚠️ The p-value for the F test will be given. You still need to state all four steps.
⚠️ Tukey's analysis: you may be asked to compute t** and SE, or you may be given the confidence intervals and asked to produce the visual display and state conclusions in plain language.
⚠️ Check assumptions: be ready to interpret a QQ plot or histogram for normality, and to compute the ratio of max to min standard deviations for the equal-variance check.
⚠️ Remember: Tukey's only comes after a significant F test. If F does not reject, you stop.
True or False: ANOVA can be used to compare the means of two groups. (True, but it is equivalent to a two-sample t test in that case.)
Fill in the blank: SST = ___ + ___. (SSA + SSE)
True or False: If the F test fails to reject H₀, you should proceed with Tukey's analysis. (False. You stop.)
Fill in the blank: The degrees of freedom for the error term in ANOVA are ___. (n – k)
True or False: MSA = SSA · (k – 1). (False. MSA = SSA / (k – 1).)
Q: A study compares the effect of three fertilisers on plant height. There are 10 plants per group (30 total). What are the factor, levels, response variable, and total observations?
A: Factor: fertiliser type. Levels: the three fertilisers (k = 3). Response variable: plant height. Total observations: n = 30.
Q: Given SSA = 120, SSE = 280, k = 3, and n = 30, complete the ANOVA table and compute F.
A: SST = 120 + 280 = 400. df_A = 3 – 1 = 2. df_E = 30 – 3 = 27. df_T = 29. MSA = 120/2 = 60. MSE = 280/27 = 10.37. F = 60/10.37 = 5.79.
Q: In the fertiliser example, the p-value for F = 5.79 is 0.008. What is your conclusion at α = 0.05?
A: Since p-value = 0.008 ≤ 0.05, we reject H₀. There is sufficient evidence to conclude that at least one fertiliser produces a different mean plant height.
Q: After rejecting H₀ in ANOVA, the Tukey interval for Fertiliser A vs. Fertiliser B is (–1.2, 4.8), and for A vs. C is (2.1, 8.3). Which pairs differ significantly?
A: The interval for A vs. B contains zero, so those two means do not significantly differ. The interval for A vs. C does not contain zero, so A and C have significantly different mean plant heights.
Q: Why is running three separate t tests (one for each pair) a problem?
A: Each t test carries its own α risk. The overall probability of at least one Type I error rises above the nominal α, because the errors compound across all tests. ANOVA and Tukey's procedure control the family-wise error rate.
ANOVA extends the two-sample t test to k groups. When k = 2, the F statistic equals t², so the results are identical. The assumptions (normality, independence, equal variances) echo the assumptions for t-based inference from Chapters 8 and 9.
The concept of partitioning variability (SST = SSA + SSE) appears throughout statistics, including in regression analysis, where total variability is split into explained and unexplained components in a very similar way.
Tukey's method is one of several multiple comparison procedures. Others (Bonferroni, Scheffé) exist but are not covered in this exam.
ANOVA, analysis of variance, one-way ANOVA, factor, levels, response variable, SSA, SSE, SST, MSA, MSE, F statistic, F test, F distribution, degrees of freedom, df1, df2, Tukey HSD, honestly significant difference, multiple comparisons, Studentized Range, q value, post-hoc test, pairwise comparison, visual display, family-wise error rate, equal variances, homogeneity, normality assumption, QQ plot