Difficulty: Intermediate | Prerequisites: ANOVA fundamentals (see companion notes), hypothesis testing, confidence intervals.
A significant ANOVA result tells you that at least one group mean differs, but it stops there. To find out which specific pairs of means are different, you need a post-hoc (multiple comparison) procedure. The central challenge is controlling the family-wise error rate: every additional pairwise test you run increases the chance of a false positive, so these methods apply corrections that keep the overall false-positive rate at your chosen α. This material follows directly from ANOVA fundamentals and is essential for interpreting real experimental results.
Post-hoc methods let you pinpoint which group means differ after ANOVA rejects H₀. The main options are Fisher's LSD (no correction, generally not recommended), Bonferroni (simple but conservative), Tukey's HSD (best all-round for all pairwise comparisons) and Dunnett's (best when comparing treatments to a single control). If the confidence interval for a pair of means does not contain zero, those means are significantly different.
LSD (Fisher's Least Significant Difference)
A post-hoc method that performs individual t-tests between every pair of group means without adjusting for multiple comparisons. In simple terms, it is the "do nothing extra" approach, and it does not protect you against inflated false positives.
Bonferroni correction
A method that divides the desired family-wise significance level (α) by the number of comparisons (c), using α/c as the threshold for each individual test. Think of it as splitting your error budget equally across all the tests you plan to run.
Tukey's Honest Significant Difference (HSD)
A post-hoc method that uses the Studentized range distribution (Q distribution) to compare all pairs of means simultaneously, controlling the family-wise error rate. In simple terms, this is the go-to method when you want to compare every group to every other group.
Dunnett's method
A post-hoc method designed for comparing multiple treatment groups against a single control group, running only k − 1 comparisons rather than all possible pairs. Think of it as the method you use when one group is the baseline and you only care whether the treatments differ from it.
Studentized range distribution (Q distribution)
The sampling distribution of the range (largest minus smallest) of k independent observations from a normal distribution, divided by an estimate of the standard deviation. It is the basis for Tukey's HSD critical values.
Post-hoc test (post-hoc comparison)
Any statistical test performed after ANOVA to identify which specific group means differ. "Post hoc" is Latin for "after this."
Confidence interval for the difference of means
An interval estimate for (μᵢ − μⱼ). If the interval does not contain zero, the two means are significantly different at the chosen confidence level.
Critical value
The threshold from the relevant distribution (t, Q, or Dunnett's t) against which the test statistic is compared. If the observed statistic exceeds it, the difference is significant.
Runs standard t-tests for every pair of means with no correction for the number of comparisons.
Does not control the family-wise error rate. The more groups you have, the worse the inflation.
Generally not recommended for serious analysis. It remains in textbooks mainly as a historical baseline and to illustrate why corrections matter.
The simplest correction: divide α by the total number of comparisons c, and use α/c as the significance threshold for each individual test.
Example: with α = 0.05 and c = 10 comparisons, each test uses α_each = 0.005.
The critical t-value is looked up at α/c (two-tailed), making each test harder to pass.
Strengths. Easy to understand and apply. Works for any type of test, not just ANOVA follow-ups.
Weaknesses. Becomes very conservative as c grows. With many comparisons, the adjusted threshold is so strict that real differences may go undetected (increased Type II error). Confidence intervals widen accordingly.
The standard choice when you want to compare every group to every other group.
Uses the Studentized range distribution (Q distribution) instead of the t-distribution. The critical value q is looked up for k groups and (n − k) degrees of freedom.
The confidence interval for the difference between means i and j is: (x̄ᵢ − x̄ⱼ) ± q(α, k, n−k) × √(MSE/nᵢ + MSE/nⱼ).
If zero falls outside this interval, the pair is significantly different.
Strengths. Controls the family-wise error rate across all pairwise comparisons simultaneously. More powerful than Bonferroni (narrower intervals, more likely to detect true differences).
Weaknesses. Designed for all pairwise comparisons. If you only care about comparisons to a control, Dunnett's method is more efficient.
In R, the function TukeyHSD() computes these intervals directly from an aov() object.
Designed for the specific case where you compare each of several treatment groups to one control group.
Runs only k − 1 comparisons (each treatment vs. control) rather than k(k−1)/2.
Uses its own critical value from the Dunnett distribution: t(α, k, n−k).
Strengths. More powerful than Tukey or Bonferroni for this particular design because it makes fewer comparisons and tailors the critical value to them.
Weaknesses. Cannot be used for all-pairs comparisons. You must designate a control group.
All pairwise comparisons needed → Tukey's HSD (first choice) or Bonferroni (simpler but more conservative).
Comparisons to a single control only → Dunnett's method.
Fisher's LSD → avoid unless the ANOVA has only three groups and a significant overall F-test (some textbooks allow it in that narrow case).
For each pair (i, j), the post-hoc method produces a confidence interval for (μᵢ − μⱼ).
If the interval contains zero, you conclude the two means are not significantly different at the chosen α.
If the interval does not contain zero, the difference is significant. The sign tells you which mean is larger.
Side-by-side boxplots. Show the distribution and median of each group. Useful for a quick visual check of overlap.
Effects plots. Plot each group mean with its confidence interval. Non-overlapping intervals suggest significant differences.
Significance lines (underlining). Order the group means from lowest to highest and draw a line under groups that are not significantly different. Groups not connected by the same line differ significantly.
Graphical displays are helpful for communication but do not replace the formal tests, because they cannot directly account for error-rate control.
Bonferroni adjusted significance level:
\alpha_{\text{each}} = \frac{\alpha_{\text{overall}}}{c}where c = number of comparisons = k(k−1)/2 for all pairwise tests.
Tukey's HSD confidence interval:
(\bar{x}_i - \bar{x}_j) \pm q_{\alpha,\, k,\, n-k} \times \sqrt{\frac{\text{MSE}}{n_i} + \frac{\text{MSE}}{n_j}}where q is the critical value from the Studentized range (Q) distribution, k is the number of groups, and n is the total number of observations.
Dunnett's critical value:
t_{\alpha,\, k,\, n-k}from the Dunnett distribution. The confidence interval has the same form as Tukey's but uses this critical value and is computed only for treatment-vs.-control pairs.
In a clinical trial comparing three drug dosages against a placebo, Dunnett's method is the natural choice because the placebo is the control and the question is simply whether each dosage outperforms it. In agricultural research testing the yield of several fertiliser blends, Tukey's HSD is used because every blend-to-blend comparison matters for choosing the best product.
Students often think Bonferroni is always the best correction because it is the simplest. In practice, with many comparisons it becomes so conservative that it misses real differences. Tukey's HSD is usually preferable for all-pairs work.
A frequent error is applying Tukey's HSD when only comparisons to a control are needed. That wastes statistical power. Use Dunnett's instead.
Some students believe that if a confidence interval for (μᵢ − μⱼ) is entirely positive, the result is inconclusive. It is not: it means μᵢ is significantly larger than μⱼ.
Students sometimes skip the ANOVA step and jump straight to pairwise comparisons. The standard workflow runs ANOVA first; post-hoc tests follow only if the overall F-test is significant.
⚠️ Know which method to use and when. A common exam question gives a scenario and asks you to pick the appropriate post-hoc procedure (all pairs → Tukey; treatments vs. control → Dunnett).
⚠️ Be able to compute the Bonferroni adjusted α for a given number of groups. Example: 4 groups → 6 comparisons → α_each = 0.05/6 ≈ 0.0083.
⚠️ Understand how to interpret a confidence interval for (μᵢ − μⱼ): contains zero = not significant; does not contain zero = significant.
⚠️ Expect a question on why Fisher's LSD is inadequate for multiple comparisons.
⚠️ Know the general 8-step workflow: ANOVA → choose method → calculate critical values → compute CIs → check for zero → graph → conclude.
True or false: Fisher's LSD controls the family-wise error rate. (False. It applies no correction.)
Fill in the blank: The Bonferroni correction sets each individual test's significance level to ______. (α / c, where c is the number of comparisons.)
True or false: Tukey's HSD is more powerful than the Bonferroni correction for all-pairs comparisons. (True.)
When should you use Dunnett's method instead of Tukey's HSD? (When you are comparing multiple treatments against a single control group.)
True or false: If a confidence interval for (μ₁ − μ₂) is (2.3, 8.7), the two means are not significantly different. (False. Zero is not in the interval, so the difference is significant.)
Q: A study has five treatment groups. Using the Bonferroni correction at α = 0.05, what is the adjusted significance level for each pairwise comparison?
A: There are 5(4)/2 = 10 comparisons. α_each = 0.05 / 10 = 0.005.
Q: Explain why Tukey's HSD is generally preferred over the Bonferroni correction for all-pairs comparisons.
A: Tukey's HSD uses the Studentized range distribution, which is specifically designed for simultaneous pairwise comparisons. It produces narrower confidence intervals than Bonferroni, giving it greater statistical power to detect real differences while still controlling the family-wise error rate.
Q: A Tukey's HSD confidence interval for (μ_A − μ_B) is (−1.2, 3.8). Are groups A and B significantly different? Explain.
A: No. The interval contains zero, so we cannot conclude that the means differ at the chosen significance level.
Q: You are testing a new fertiliser against an existing standard and an untreated control. Which post-hoc method is most appropriate, and why?
A: Dunnett's method, because the question is whether each treatment differs from the control, not whether the treatments differ from each other. Dunnett's runs only k − 1 comparisons (each treatment vs. control), giving more power than Tukey's HSD for this design.
Q: Describe the general procedure for multiple comparisons after ANOVA, starting from a significant F-test.
A: (1) Confirm the ANOVA F-test is significant. (2) Choose α (usually 0.05). (3) Select the appropriate post-hoc method (Tukey for all pairs, Dunnett for treatment vs. control). (4) Calculate the critical value from the relevant distribution. (5) Compute confidence intervals for each pair of interest. (6) Check whether zero is inside or outside each interval. (7) Graph the results (e.g. effects plot or significance lines). (8) State which pairs differ in plain language.
This material follows directly from ANOVA fundamentals (see the companion notes). The Bonferroni correction is a general principle that appears beyond ANOVA wherever multiple hypothesis tests are run, including in genomics and neuroimaging. Confidence interval interpretation connects back to the broader confidence-interval material from earlier in the course. If you continue into experimental design, the choice of post-hoc method is tightly linked to the structure of the experiment (factorial designs, repeated measures).
Related Terms / Search Tags
multiple comparisons, post-hoc test, post hoc analysis, pairwise comparisons, Fisher's LSD, least significant difference, Bonferroni correction, Bonferroni adjustment, Tukey's HSD, Tukey's honest significant difference, Studentized range distribution, Q distribution, Dunnett's method, Dunnett's test, family-wise error rate, FWER, Type I error inflation, confidence interval for difference of means, significance lines, effects plot, boxplot comparison, critical value, adjusted alpha, multiple testing correction, introduction to statistics, Purdue statistics