One-Way ANOVA and Tukey's HSD, STAT 350 – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Basic hypothesis testing, confidence intervals, familiarity with R.

This topic sits in the comparison-of-means portion of an introductory statistics course. You have already learnt how to compare two groups (two-sample t-test). ANOVA extends that logic to three or more groups at once, and Tukey's HSD tells you which specific pairs differ after ANOVA finds a significant result. If you are comfortable with the idea of a null hypothesis, a p-value, and a confidence interval, you are ready for this material.

TL;DR

One-way ANOVA tests whether the means of three or more groups are all equal. When it rejects, Tukey's Honestly Significant Difference (HSD) test compares every pair of groups and tells you which specific pairs differ, while controlling the family-wise error rate. The R output here compares log-transformed oil prices across four continents and finds that most pairs are significantly different, with Europe highest and South America lowest.

Key Terms

One-way ANOVA (analysis of variance)

A hypothesis test that compares the means of three or more independent groups by partitioning total variability into between-group and within-group components. The null hypothesis is that all group means are equal.

Think of it as asking: "Are any of these groups noticeably different from each other, or could the variation I see just be noise?"

F-test / F-statistic

The ratio of between-group variance to within-group variance. A large F means the groups differ more than you would expect from random variation alone.

In simple terms, a big F says the group labels actually matter.

Tukey's HSD (Honestly Significant Difference)

A post-hoc pairwise comparison procedure run after ANOVA rejects. It compares every pair of group means and produces a confidence interval and adjusted p-value for each difference, while controlling the overall Type I error rate.

Think of it as the follow-up question: "ANOVA said something differs, but which pairs?"

Pairwise comparison

A test of the difference between two specific group means. With k groups there are k(k-1)/2 pairs. For 4 groups that gives 6 pairs.

Adjusted p-value (p adj)

A p-value corrected for the fact that you are running multiple comparisons simultaneously. Tukey's method builds this correction into its procedure so that the family-wise error rate stays at the chosen alpha.

In simple terms, it accounts for the increased chance of a false positive when you test many pairs at once.

Confidence interval (lwr, upr)

In the Tukey HSD output, each pair gets a confidence interval for the true difference in means. If the interval does not contain zero, the difference is statistically significant at the chosen confidence level.

Log transformation (lOilPrice)

Taking the natural logarithm of a variable before analysis. This is common when the original data are right-skewed or when variances differ across groups. The "l" prefix in lOilPrice signals that the response has been log-transformed.

Group means (sample means per level)

The average value of the response variable within each group. In this output: South America 3.12, Asia 3.15, North America 3.56, Europe 3.69 (all on the log scale).

Family-wise error rate

The probability of making at least one Type I error (false positive) across all comparisons. Tukey's HSD keeps this at or below alpha even when many pairs are tested.

qtukey (studentised range distribution)

The distribution used to derive Tukey's critical values. The R function qtukey(p, k, df) returns quantiles of this distribution, where k is the number of groups and df is the within-group degrees of freedom.

Core Content

Setting Up a One-Way ANOVA in R

  • The model formula lOilPrice ~ ContName tells R: "explain variation in log oil price by continent name."

  • aov() fits the model. The data frame is OilPrice.sub, a subset of a larger dataset.

  • The response variable (lOilPrice) has been log-transformed before analysis, likely to stabilise variance and reduce skewness.

  • There are 4 groups (continents): South America, Asia, North America, Europe.

  • With 4 groups and a total of n = 668 observations (inferred from df = 664 within groups, plus k = 4 groups), the between-group df = k - 1 = 3 and the within-group df = n - k = 664.

Reading the Tukey HSD Output

The Tukey HSD table has one row per pair. Each row shows:

  • diff: the difference in sample means (group named first minus group named second)

  • lwr: lower bound of the confidence interval for the true difference

  • upr: upper bound of the confidence interval for the true difference

  • p adj: the adjusted p-value for that pairwise comparison

Interpreting Pairwise Differences

  • Europe vs Asia (diff = 0.542): significantly different (p adj = 0.000). Europe's mean log price is about 0.54 higher than Asia's.

  • North America vs Asia (diff = 0.407): significantly different (p adj = 0.000).

  • South America vs Asia (diff = -0.027): not significantly different (p adj = 0.928). The confidence interval (-0.166 to 0.112) contains zero.

  • North America vs Europe (diff = -0.134): significantly different (p adj = 0.002). Europe's mean is higher.

  • South America vs Europe (diff = -0.569): significantly different (p adj = 0.000). Europe's mean is much higher.

  • South America vs North America (diff = -0.434): significantly different (p adj = 0.000). North America's mean is higher.

Five of the six pairs show a statistically significant difference. The only pair that does not differ significantly is South America and Asia.

Ordering the Group Means

The group means on the log scale, from lowest to highest:

  • South America: 3.125

  • Asia: 3.152

  • North America: 3.559

  • Europe: 3.694

South America and Asia cluster together at the lower end. North America and Europe are both higher, with Europe highest. The underline beneath South America and Asia in the output is a visual grouping: those two are not significantly different from each other.

The qtukey Critical Value

The command qtukey(0.99, 3, 664) / sqrt(2) computes a critical value from the studentised range distribution. Here 0.99 is the confidence level, 3 is the number of groups minus one (or the number of means being compared, depending on the parameterisation), and 664 is the within-group degrees of freedom. The result, 2.924, is used to construct simultaneous confidence intervals at the 99% level.

Formulas and R Commands

F = \frac{\text{MS}_{\text{between}}}{\text{MS}_{\text{within}}} = \frac{\text{SS}_{\text{between}} / (k - 1)}{\text{SS}_{\text{within}} / (n - k)}

where k = number of groups, n = total observations.

\text{Tukey HSD interval for } \mu_i - \mu_j = (\bar{x}_i - \bar{x}_j) \pm q_{\alpha, k, \text{df}_W} \cdot \frac{s_W}{\sqrt{n_{\text{per group}}}}

where q is the critical value from the studentised range distribution and s_W is the square root of the within-group mean square.

Key R Functions

  • aov(response ~ factor, data = df): fits the one-way ANOVA model.

  • TukeyHSD(aov_model): runs Tukey's HSD on the fitted ANOVA object. Returns diff, lwr, upr, and p adj for every pair.

  • qtukey(p, nmeans, df): returns quantiles of the studentised range distribution. Dividing by sqrt(2) converts it to the form used for pairwise confidence intervals.

  • summary(aov_model): prints the ANOVA table with the F-statistic and overall p-value (not shown in this output, but typically run first).

Real-World Application

This dataset compares petrol or crude oil prices across continents. Energy analysts and economists use exactly this kind of comparison to understand geographic pricing disparities. The log transformation is standard practice because commodity prices tend to be right-skewed, and working on the log scale stabilises variance and makes the ANOVA assumptions more plausible.

The finding that Europe has significantly higher log oil prices than all other continents, while South America and Asia do not differ from each other, could inform discussions about taxation policy, refining capacity, or supply chain logistics.

Common Misconceptions

  • Students often think a significant ANOVA result means every pair of groups differs. It does not. ANOVA only tells you that at least one pair differs. You need a post-hoc test like Tukey's HSD to find out which pairs.

  • Students sometimes confuse "diff" in the Tukey output with an absolute difference. The sign matters: a negative diff means the second-named group has the higher mean.

  • A common error is interpreting results on the log scale as though they are on the original scale. A difference of 0.54 in log prices does not mean prices differ by 0.54 dollars. On the original scale, it corresponds to a multiplicative factor of e^0.54, which is roughly 1.72.

  • Students occasionally assume that a large adjusted p-value (like 0.928 for South America vs Asia) means the two group means are equal. It means there is insufficient evidence to conclude they differ. That is a weaker statement.

Why It Matters / Exam Flags

  • ⚠️ You will likely be asked to read Tukey HSD output and state which pairs are significantly different. Practice identifying this from the p adj column and from whether the confidence interval contains zero.

  • ⚠️ Expect questions about why a post-hoc test is needed after ANOVA, rather than just running many t-tests. The answer centres on controlling the family-wise error rate.

  • ⚠️ Know how to interpret the sign of the "diff" column. The convention is first group minus second group in the row label.

  • ⚠️ Be ready to explain why a log transformation might be applied before running ANOVA (stabilise variance, reduce skewness, meet normality assumptions).

  • ⚠️ Ordering the group means and identifying which groups can be grouped together (not significantly different) is a classic exam task. Here, South America and Asia form one group; North America and Europe each stand apart.

Quick Self-Test

  1. True or False: A significant ANOVA F-test tells you exactly which group means differ. (False. It only tells you at least one pair differs.)

  1. True or False: In the Tukey HSD output, a confidence interval that contains zero means the pair is significantly different. (False. Containing zero means the difference is not significant.)

  1. Fill in the blank: With 4 groups, Tukey's HSD compares ______ pairs. (6)

  1. True or False: The pair South America and Asia shows a significant difference in this output. (False. p adj = 0.928, well above any conventional alpha.)

  1. Fill in the blank: The continent with the highest mean log oil price is ______. (Europe, at 3.694)

Practice Q&A

Q: Based on the Tukey HSD output, which pairs of continents do not show a statistically significant difference in mean log oil price? Explain how you can tell.

A: South America and Asia. Their adjusted p-value is 0.928 (far above 0.05), and their 95% confidence interval for the difference (-0.166 to 0.112) contains zero.

Q: The diff for Europe minus Asia is 0.542. What does this value represent on the original (non-log) scale?

A: On the log scale, 0.542 is the difference in means. On the original scale, this corresponds to a ratio of e^0.542, which is approximately 1.72. So European oil prices are roughly 1.72 times Asian oil prices, on average.

Q: Why is Tukey's HSD preferred over running six separate two-sample t-tests?

A: Running six t-tests each at alpha = 0.05 inflates the probability of at least one false positive well beyond 5%. Tukey's HSD adjusts for multiple comparisons so the family-wise error rate is controlled at the stated alpha level.

Q: A student claims that because the ANOVA was run on log-transformed prices, the results are invalid. Is the student correct?

A: No. Log-transforming is a standard and valid technique to meet ANOVA assumptions (normality and equal variance). The results are valid on the log scale. Interpretation simply requires back-transforming (exponentiating) differences to understand them as multiplicative factors on the original scale.

Q: Arrange the four continents from lowest to highest mean log oil price. Which continents can be grouped together as "not significantly different"?

A: South America (3.125), Asia (3.152), North America (3.559), Europe (3.694). South America and Asia can be grouped together (p adj = 0.928). All other pairs differ significantly.

Connections to Other Topics

This connects to the two-sample t-test because ANOVA generalises it to more than two groups. If you only had two continents, the F-test would give the same p-value as a two-sided t-test.

It also connects to regression: one-way ANOVA is equivalent to a linear regression with dummy variables for each group. The aov() output in R can be rewritten as lm(lOilPrice ~ ContName) and the F-test is the same.

The concept of controlling error rates across multiple tests reappears in later topics such as Bonferroni correction, Scheffé's method, and false discovery rate (FDR) control.


Related Terms / Search Tags

One-way ANOVA, analysis of variance, Tukey HSD, Tukey's Honestly Significant Difference, post-hoc test, pairwise comparison, multiple comparisons, family-wise error rate, FWER, adjusted p-value, studentised range distribution, qtukey, log transformation, group means, F-test, F-statistic, between-group variance, within-group variance, oil prices by continent, STAT 350, Purdue, introduction to statistics, R aov function, TukeyHSD R, confidence interval for difference in means.