Difficulty: Intermediate | Prerequisites: ANOVA fundamentals (first half of Chapter 12), confidence intervals, hypothesis testing.
Once ANOVA tells you that at least one population mean is different, the natural follow-up is: which ones? You cannot simply run all possible two-sample t-tests because that inflates your Type I error rate far beyond alpha. This section covers the methods that correct for that problem by adjusting the critical value used in pairwise confidence intervals. Choosing the right method depends on whether you are comparing every pair of means (use Tukey) or only comparing each treatment to a control (use Dunnett).
After a significant ANOVA result, multiple comparison methods let you identify which specific pairs of means differ while controlling the overall (familywise) Type I error rate. Tukey's HSD is the go-to method for all pairwise comparisons; Dunnett's method is preferred when comparing treatments to a single control. Both work by widening the confidence interval (larger critical value) so the chance of a false positive across all tests stays at alpha.
Familywise error rate (alpha_overall)
The probability of making at least one Type I error across the entire set of pairwise comparisons. This is what multiple comparison methods are designed to control.
Think of it as: the chance that anywhere in your batch of tests, you wrongly declare a difference that does not exist.
Type I error inflation
The increase in the overall probability of a false positive that occurs when you perform multiple hypothesis tests, each at alpha. With k = 10 groups and 45 pairwise t-tests at alpha = 0.05, the overall risk rises to about 0.63.
In simple terms, the more tests you run, the more likely you are to get a false alarm somewhere.
Pairwise comparison
A test or confidence interval comparing the means of two specific groups. With k groups, the number of all possible pairwise comparisons is c = k(k - 1) / 2.
LSD (Fisher's Least Significant Difference)
The simplest method: it uses the ordinary two-sample t-test critical value with no adjustment. Because it does not correct for multiple comparisons, it is generally not appropriate and is not recommended.
Bonferroni correction
Divides alpha_overall by c (the number of comparisons) to get alpha_each for each individual test. Conservative: the resulting confidence intervals are often wider than necessary, especially when c is large.
Think of it as: spreading your error budget evenly across every test, which is safe but wasteful.
Tukey's HSD (Honestly Significant Difference)
The preferred method when comparing all pairs of means. Uses a critical value from the Studentized range distribution, Q, divided by the square root of 2. Produces narrower intervals than Bonferroni, so it detects more real differences.
In simple terms, Tukey is the standard choice for "which groups differ from which?"
Studentized range distribution
The probability distribution used to derive the Tukey critical value. Denoted Q with parameters alpha, k (number of means), and n - k (error degrees of freedom). Calculated in R with qtukey().
Tukey parameter
The value Q_alpha,k,n-k from the Studentized range distribution. The actual critical value used in the confidence interval is Q / sqrt(2).
Dunnett's method
Used when one group is a control and you only want to compare every other group to that control. Makes k - 1 comparisons instead of c = k(k-1)/2. Has its own critical value, d_alpha,k,n-k.
Think of it as: the efficient choice when one group is the baseline.
Graphical display (underline method)
A visual summary of Tukey results. Order the group means from lowest to highest and draw a line under groups whose difference is not statistically significant. Groups connected by the same line are considered statistically the same.
Margin of error (for pairwise CI)
The product of the adjusted critical value and the standard error. The different methods (LSD, Bonferroni, Tukey, Dunnett) change the critical value, which changes the margin of error.
When you run multiple t-tests at alpha = 0.05 each, each test controls its own Type I error, but the overall probability of at least one false positive across all tests climbs rapidly.
The overall risk depends on the number of groups k (and therefore the number of pairs c):
k = 2: overall risk = 0.05
k = 3: overall risk = 0.12
k = 4: overall risk = 0.20
k = 6: overall risk = 0.37
k = 10: overall risk = 0.63
This is the primary reason for using multiple comparison procedures. Two secondary reasons also apply:
Pooling all groups gives a better estimate of the standard deviation (using MSE from the ANOVA table).
The framework extends more naturally to two-way ANOVA and other designs.
Graphically: side-by-side boxplots or effects plots (plots of group means). Useful as a visual check, but they do not show the error or tell you whether differences are statistically significant.
Multiple comparisons: the analytical approach. Builds confidence intervals for each pair of means and checks whether zero falls inside each interval.
The general form is: (x-bar_i - x-bar_j) +/- t**_column,df * SE
The standard error is: SE = sqrt( MSE * (1/n_i + 1/n_j) )
When all group sizes are equal (n_i = n_j = n_i for all groups), this simplifies to: SE = sqrt( 2 * MSE / n_i )
The different methods (LSD, Bonferroni, Tukey, Dunnett) all use the same standard error. What changes is the critical value t**.
Decision rule: if zero is in the interval, the two means are not significantly different (fail to reject H_0 for that pair). If zero is not in the interval, the two means are significantly different (reject H_0 for that pair).
LSD (Fisher's Least Significant Difference)
Uses the ordinary t critical value with no adjustment.
Does not control the familywise error rate.
Not recommended. Included here only for completeness.
Bonferroni
Divides alpha_overall by c to get alpha_each = alpha_overall / c.
The critical value becomes t at alpha/(2c) with n - k degrees of freedom.
Conservative: alpha_overall is usually much smaller than c * alpha_each, so the intervals are wider than they need to be.
When c is large, alpha_each becomes tiny, making it hard to detect real differences.
Not the preferred method for all-pairs comparisons, but still used in specialised circumstances.
Tukey's HSD (the preferred method for all pairwise comparisons)
Uses the Studentized range distribution to derive its critical value.
The Tukey parameter Q is obtained from: qtukey(p = 1 - alpha, nmeans = k, df = n - k) in R.
The critical value for the confidence interval is: t** = Q / sqrt(2).
Produces narrower intervals than Bonferroni, so it is more powerful (detects more real differences).
R command to get all Tukey intervals at once: TukeyHSD(fit, conf.level = 1 - alpha), where "fit" is the aov() object from the ANOVA.
Requires the ANOVA results (specifically MSE and degrees of freedom).
Dunnett's Method (the preferred method when comparing to a control)
Used when one group is a control and you compare every treatment to that control.
Makes only k - 1 comparisons instead of k(k-1)/2.
Has its own critical value: t** = d_alpha,k,n-k.
For the Paxil example: the Dunnett critical value (2.50) is smaller than the Tukey critical value (2.668), so Dunnett intervals are narrower. This means Dunnett can detect differences that Tukey misses, but only for comparisons against the control.
Used after Tukey's method to summarise results visually.
Step 1: Order the group means from lowest to highest. Include the mean values under the group labels.
Step 2: Draw a line under groups that are NOT statistically significantly different (i.e., whose confidence interval contains zero).
Groups sharing a line are statistically equivalent. Groups that do not share any line are significantly different.
This is always done by hand; R does not produce this display automatically.
Important: the relationship is not transitive. B can be the same as C, and C can be the same as D, but B and D can still be different. Each pair of adjacent means may be close enough to overlap, while the two endpoints are too far apart.
Perform the ANOVA test and confirm a significant result. If the result is not significant, stop: there is no reason to look for pairwise differences.
Select a familywise significance level alpha (usually the same alpha used in the ANOVA test).
Choose the method: Tukey for all pairwise comparisons, Dunnett for comparisons to a control.
Calculate the critical value (Tukey parameter / sqrt(2), or Dunnett d value).
Calculate all required confidence intervals.
Determine which intervals contain zero (not significant) and which do not (significant).
(Tukey only) Create the graphical display.
Write a conclusion in context, answering the original question in plain English.
A sample of 15 men split into three groups (0 mg, 20 mg, 40 mg of Paxil). After one week, serotonin levels measured. The ANOVA table:
Source | df | SS | MS | F | p-value |
|---|---|---|---|---|---|
Dose | 2 | 841.88 | 420.94 | 8.36 | 0.005 |
Error | 12 | 604.34 | 50.36 | ||
Total | 14 | 1446.23 |
Group means: 0 mg = 57.60, 20 mg = 69.28, 40 mg = 75.70.
Part (b), Dunnett (comparing to 0 mg control):
Critical value d = 2.50. SE = 4.488.
D20 vs D0: 11.68, interval (0.46, 22.9), significant.
D40 vs D0: 18.1, interval (6.88, 29.32), significant.
Conclusion: both 20 mg and 40 mg raise serotonin levels compared to the control.
Part (c), Tukey (all pairwise):
Critical value = Q / sqrt(2) = 3.773 / 1.414 = 2.668. SE = 4.488.
D20 vs D0: 11.68, interval (-0.30, 23.65), not significant.
D40 vs D0: 18.1, interval (6.12, 30.07), significant.
D40 vs D20: 6.42, interval (-5.55, 18.39), not significant.
Graphical display: 0 mg and 20 mg share a line; 20 mg and 40 mg share a line. Only 0 mg vs 40 mg is significant.
Conclusion: only the 40 mg dose significantly raises serotonin levels compared to the control. The 20 mg dose is not significantly different from the control under Tukey.
The Dunnett result differs from the Tukey result because Tukey's critical value (2.668) is larger than Dunnett's (2.50), making Tukey intervals wider and harder to reach significance.
c = \binom{k}{2} = \frac{k(k-1)}{2}\bar{x}_{i.} - \bar{x}_{j.} \pm t^{**}_{\text{column}, df} \cdot SESE = \sqrt{MSE \left( \frac{1}{n_i} + \frac{1}{n_j} \right)}When all group sizes are equal (n_i = n_j):
SE = \sqrt{\frac{2 \cdot MSE}{n_i}}t^{**}_{\text{column}, dfe} = t_{\alpha / 2c, \, n-k}where c = k(k-1)/2 is the number of comparisons.
t^{**}_{\text{columns}, df} = \frac{Q_{\alpha, k, n-k}}{\sqrt{2}}In R: Q = qtukey(p = 1 - alpha, nmeans = k, df = n - k). Remember to divide by sqrt(2).
t^{**}_{\text{column}, df} = d_{\alpha, k, n-k}\alpha_{\text{overall}} \leq \alpha_{\text{each}} + \alpha_{\text{each}} + \cdots + \alpha_{\text{each}} = c \cdot \alpha_{\text{each}}Therefore:
\alpha_{\text{each}} = \frac{\alpha_{\text{overall}}}{c}In clinical trials, researchers routinely use Dunnett's method to compare each treatment arm to a placebo group, while controlling the overall false-positive rate. In quality engineering, Tukey's HSD is used to compare product measurements across multiple suppliers or production shifts to determine which, if any, differ from the rest.
Students often think you should always perform multiple comparisons after ANOVA. You should not. If the ANOVA result is not significant (fail to reject H_0), there is no evidence of any difference, and post-hoc tests are not warranted.
Students sometimes use Tukey when comparing to a control. Dunnett is the correct choice in that scenario because it makes fewer comparisons (k - 1 instead of k(k-1)/2) and has a smaller critical value, giving narrower intervals and more power.
Students forget to divide the Tukey parameter Q by sqrt(2) when asked for the critical value. The critical value is Q / sqrt(2), not Q itself.
Students assume that "not significantly different" is transitive. It is not. Group B can be statistically the same as C, and C the same as D, while B and D are significantly different. Each comparison is evaluated on its own interval.
Expect to be asked to choose the correct method for a given scenario: Tukey for all pairs, Dunnett for comparisons to a control.
You may be given Tukey output from R and asked to interpret which pairs are significant. Check whether each interval contains zero.
The graphical display (underline method) is commonly tested. Practise ordering means and drawing lines under non-significant pairs.
Know the relationship between Tukey and Dunnett critical values: Tukey's is always larger (wider intervals), so a pair that is significant under Dunnett may not be significant under Tukey.
Be prepared to explain why the Bonferroni method is too conservative for all-pairs comparisons.
True or False: You should perform multiple comparisons even if the ANOVA result is not significant.
False. Only proceed with post-hoc comparisons after a significant ANOVA result.
Fill in the blank: The number of pairwise comparisons for k = 5 groups is ______.
c = 5(4)/2 = 10.
True or False: The difference between LSD, Bonferroni, Tukey, and Dunnett is the standard error used.
False. The standard error is the same. What differs is the critical value.
Fill in the blank: If zero is contained in a pairwise confidence interval, the two means are ______ (significantly different / not significantly different).
Not significantly different.
True or False: Tukey's critical value is always smaller than Dunnett's critical value.
False. Tukey's critical value is always larger than Dunnett's, which is why Tukey intervals are wider.
Q: You have k = 4 groups. How many pairwise comparisons does Tukey's method require? How many does Dunnett's method require (assuming group 1 is the control)?
A: Tukey requires c = 4(3)/2 = 6 comparisons. Dunnett requires k - 1 = 3 comparisons (each treatment vs. the control).
Q: A Tukey confidence interval for mu_2 - mu_1 is (-3.4, 8.2). Is this pair significantly different at the chosen alpha? Explain.
A: No. The interval contains zero, so there is not enough evidence to conclude that mu_2 and mu_1 are different.
Q: In the Paxil example, the Dunnett method found 20 mg vs 0 mg significant, but Tukey did not. Explain why the results differ.
A: Tukey's critical value (2.668) is larger than Dunnett's (2.50) because Tukey must control the error rate across all k(k-1)/2 pairs, while Dunnett only controls it across k - 1 comparisons. The larger critical value widens the Tukey interval enough for zero to fall inside it, making the 20 mg vs 0 mg comparison non-significant under Tukey.
Q: After Tukey's method, you find: A vs B not significant, B vs C not significant, A vs C significant. Draw the graphical display (means ordered A < B < C).
A: A and B share a line. B and C share a line. But A and C do not share a line. The display looks like: A B C, with one underline connecting A and B, and a separate underline connecting B and C. This is an example of non-transitivity.
Q: Why is Bonferroni not preferred for all-pairs comparisons when k is large?
A: Because alpha_each = alpha_overall / c becomes very small as c grows, making the critical value very large and the intervals very wide. The result is that practically meaningful differences may fail to reach statistical significance. Tukey's method controls the familywise error rate more efficiently.
This material is the direct continuation of the ANOVA fundamentals from the first half of Chapter 12. The confidence interval mechanics (point estimate +/- critical value * SE) are the same framework from Chapters 9 and 11, just with an adjusted critical value. The idea of controlling error rates across multiple tests also appears in more advanced settings such as multiple regression and experimental design.
The Bonferroni correction is a general-purpose tool that surfaces in many areas of statistics (not just ANOVA), so understanding why it is conservative here helps in those contexts too.
multiple comparisons, post-hoc tests, Tukey HSD, Tukey's honestly significant difference, Bonferroni correction, Bonferroni adjustment, Dunnett test, Dunnett method, Fisher LSD, least significant difference, familywise error rate, experimentwise error rate, Type I error inflation, pairwise confidence intervals, Studentized range distribution, qtukey, graphical display underline method, ANOVA follow-up, which means are different, STAT 101, Purdue statistics, Chapter 12 section 2, post-hoc ANOVA, multiple testing correction, alpha adjustment