Source: Objectives for Final Exam, Introduction to Statistics (Purdue University)
Difficulty: Intermediate | Prerequisites: Two-sample t-tests, hypothesis testing (four-step procedure from Midterm 2), basic R syntax.
ANOVA (Analysis of Variance) is the method you use when you need to compare the means of three or more populations at once. You already know how to compare two means with a t-test, but running multiple t-tests inflates your Type I error rate. ANOVA solves this by using a single F-test that compares variation between groups to variation within groups. If you are comfortable with hypothesis testing and the idea of sampling distributions, you have the foundation you need here.
ANOVA tests whether the means of three or more groups differ by comparing how much the group means spread out (between-group variance) to how much individual observations spread within each group (within-group variance). A large F statistic means the groups probably do not all share the same mean. When ANOVA rejects the null, Tukey's HSD tells you which specific pairs of groups differ.
Factor
The categorical explanatory variable whose effect you are testing. Think of it as the "grouping variable" that splits your data into populations.
Levels of the factor
The individual categories or groups within the factor. If the factor is "drug dosage" with values 0 mg, 20 mg, and 40 mg, there are three levels.
Response variable
The quantitative outcome you are measuring across the groups.
k
The number of groups (levels of the factor).
n
The total number of observations across all groups.
SSA (Sum of Squares for Factor A)
Measures the variation between group means. In simple terms, this captures how far apart the group averages are from the overall average.
SSE (Sum of Squares for Error)
Measures the variation within groups. This is the "noise" or natural spread of observations around their own group mean.
SST (Total Sum of Squares)
The total variation in the data. SST = SSA + SSE.
MSA (Mean Square for Factor)
SSA divided by its degrees of freedom (k - 1). Think of it as the average between-group variance.
MSE (Mean Square for Error)
SSE divided by its degrees of freedom (n - k). Think of it as the average within-group variance.
F statistic
The test statistic for ANOVA, calculated as MSA / MSE. A large F means the between-group differences are big relative to the noise.
Tukey's HSD (Honestly Significant Difference)
A post-hoc multiple comparison method that builds confidence intervals for every pair of group means, controlling the overall Type I error rate.
Tukey parameter (Q)
The critical value from the Studentised Range distribution used in Tukey's procedure. In R: qtukey(C, k, dfe).
ANOVA compares means indirectly by looking at variances. If the group means are truly equal, the spread of the group averages around the overall mean should be small relative to the spread of individual observations within each group.
When between-group variance is large relative to within-group variance, the evidence points to at least one group mean being different.
For any ANOVA problem, identify:
Factor: the categorical variable (e.g. "teaching method")
Levels: the specific groups (e.g. lecture, online, hybrid)
Response variable: the quantitative outcome (e.g. exam score)
Populations: the groups being compared (one per level)
Total observations (n): the count of all data points across all groups
Three conditions must hold:
Independence / SRS: Each sample is a simple random sample from its population, and the samples are independent of one another.
Normality: Each population is normally distributed. Check with a normal probability plot or histogram for each group.
Equal (constant) standard deviations: All populations share the same standard deviation. Check by computing the ratio of the largest sample standard deviation to the smallest. A ratio under 2 is generally acceptable.
The table has three rows: Factor, Error, and Total.
Factor row: df = k - 1, SSA = between-group sum of squares, MSA = SSA / (k - 1)
Error row: df = n - k, SSE = within-group sum of squares, MSE = SSE / (n - k)
Total row: df = n - 1, SST = total sum of squares
Key relationships:
SST = SSA + SSE
dft = dfa + dfe
You can find any missing piece by addition or subtraction
F = MSA / MSE
Numerator degrees of freedom: df1 = k - 1
Denominator degrees of freedom: df2 = n - k
The F distribution is right-skewed and always positive. Large values of F provide evidence against the null.
Hypotheses: H0: all population means are equal. Ha: not all population means are equal (written in words, not symbols).
Test statistic: F = MSA / MSE, with df1 = k - 1 and df2 = n - k.
P-value: Found from the F distribution. In R: pf(fts, df1, df2, lower.tail = FALSE).
Conclusion: Compare the p-value to your significance level. State the conclusion in context.
Fit the model: fit <- aov(response ~ factor, data = dataset)
View results: summary(fit)
The summary() step is essential. Running aov() alone does not display the ANOVA table.
When comparing exactly two groups, you can use either a two-sample t-test or ANOVA. The results are equivalent: F = t squared.
The t-test is preferred for two groups because it allows one-sided alternatives and gives a confidence interval for the difference in means directly.
ANOVA (F-test) is required when comparing three or more groups, because it controls the Type I error rate across all comparisons simultaneously.
After ANOVA rejects the null, you know at least one pair of means differs, but not which pair.
Running all pairwise t-tests inflates the overall Type I error rate. With k groups, there are k(k-1)/2 pairs, and each test carries its own chance of a false positive.
A multiple comparison procedure like Tukey's HSD controls the family-wise error rate.
When to use it: After ANOVA rejects the null hypothesis (p-value < alpha).
What it does: Builds a confidence interval for every pairwise difference in means. If an interval does not contain zero, those two groups are significantly different.
The interval formula:
Difference: x-bar_i - x-bar_j
Margin: t** times SE, where t** = Q / sqrt(2) and SE = sqrt(MSE * (1/n_i + 1/n_j))
For balanced designs (equal group sizes): SE = sqrt(2 * MSE / n_i)
R code:
Get the Tukey parameter: qtukey(C, k, dfe)
Get all intervals: TukeyHSD(fit, conf.level = 0.95)
Visual display: Groups are listed in order of their means, and a line is drawn under groups that are not significantly different from each other. Groups not connected by a line are significantly different.
Interpreting results in plain language: State which groups differ and which do not. For example: "The 40 mg group has a significantly higher mean response than the control, but the 20 mg group is not significantly different from either."
SSA = \sum_{i=1}^{k} n_i (\bar{x}_{i.} - \bar{x}_{..})^2SSE = \sum_{i=1}^{k} (n_i - 1) s_i^2SST = SSA + SSEF = \frac{MSA}{MSE} = \frac{SSA / (k-1)}{SSE / (n-k)}\text{Tukey interval: } \bar{x}_{i.} - \bar{x}_{j.} \pm t^{**}_{\text{column}, df} \cdot SEt^{**} = \frac{Q_{\alpha, k, n-k}}{\sqrt{2}}, \quad SE = \sqrt{MSE \left(\frac{1}{n_i} + \frac{1}{n_j}\right)}ANOVA is widely used in clinical trials to compare the effectiveness of three or more drug dosages, in agriculture to test whether different fertiliser treatments produce different crop yields, and in manufacturing to determine whether several machines produce parts with the same average dimension. Any time you have one categorical factor and one quantitative response across more than two groups, ANOVA is the standard tool.
Students often think ANOVA tests whether all means are different from each other. It does not. ANOVA only tests whether at least one mean differs from the rest.
Students sometimes write H_a using symbols (e.g. mu_1 =/= mu_2 =/= mu_3). The alternative hypothesis for ANOVA is stated in words: "not all population means are equal."
A common error is running multiple pairwise t-tests instead of using Tukey's HSD after rejecting H_0. This inflates the Type I error rate and can produce misleading results.
Students sometimes forget that ANOVA requires the summary() call after aov() in R. Without it, you see the model object, not the ANOVA table.
You will need to fill in a one-way ANOVA table by hand. Practise working backwards from partial information (e.g. given SST and SSA, find SSE; given df and SS, find MS).
The alternative hypothesis must be written in words, not symbols. This is different from the one-sample and two-sample tests.
There are two degrees of freedom for the F statistic (df1 = k - 1, df2 = n - k). Know both.
You will be given extra R output on the exam. Be able to identify which output corresponds to the ANOVA you need and which is irrelevant.
For Tukey's HSD, you will not do hand calculations, but you must understand the formulas and be able to interpret the visual display and output.
True or false: The null hypothesis for one-way ANOVA states that all population means are different. (False, it states they are all equal.)
Fill in the blank: SST = ______ + ______. (SSA + SSE)
True or false: If k = 4 and n = 40, the error degrees of freedom is 36. (True: n - k = 40 - 4 = 36.)
True or false: A Tukey confidence interval that contains zero means those two group means are significantly different. (False, it means they are not significantly different.)
Fill in the blank: In R, the p-value for an ANOVA F-test is calculated with pf(fts, df1, df2, lower.tail = ______). (FALSE)
Q: A researcher tests three brands of fertiliser on plant growth. What is the factor, what are the levels, and what is the response variable?
A: The factor is fertiliser brand. The levels are the three specific brands. The response variable is plant growth (a quantitative measure such as height or yield).
Q: You are given SSA = 120, SSE = 480, k = 4, and n = 40. Calculate MSA, MSE, and the F statistic.
A: MSA = 120 / (4 - 1) = 40. MSE = 480 / (40 - 4) = 13.33. F = 40 / 13.33 = 3.00.
Q: An ANOVA test returns a p-value of 0.002 at alpha = 0.05. What do you conclude, and what is the next step?
A: Reject H_0. There is sufficient evidence that not all population means are equal. The next step is to perform Tukey's HSD to determine which specific pairs of means differ.
Q: A Tukey interval for the difference between Group A and Group B is (2.1, 8.7). What does this tell you?
A: Because the interval does not contain zero, the means of Group A and Group B are significantly different. Specifically, Group A's mean is estimated to be between 2.1 and 8.7 units higher than Group B's mean.
Q: Why is it inappropriate to perform six separate two-sample t-tests when comparing four group means?
A: Each t-test carries an alpha-level chance of a Type I error. Running six tests without adjustment inflates the overall probability of at least one false positive well beyond the nominal alpha level. Tukey's HSD controls the family-wise error rate.
ANOVA extends the two-sample t-test to more than two groups. When k = 2, the F statistic equals the square of the t statistic, and the conclusions are identical. This connects directly to Chapter 12 (Linear Regression), which uses its own ANOVA table to test whether the regression model explains a significant amount of variation in Y. The logic is the same: partition total variability into explained and unexplained portions, then compare them with an F ratio.
ANOVA, analysis of variance, one-way ANOVA, F-test, F statistic, F distribution, between-group variance, within-group variance, sum of squares, SSA, SSE, SST, mean square, MSA, MSE, degrees of freedom, factor, levels, response variable, Tukey, Tukey's HSD, honestly significant difference, multiple comparisons, pairwise comparisons, post-hoc test, family-wise error rate, Type I error inflation, aov(), summary(), pf(), qtukey(), TukeyHSD(), STAT 301, Purdue, intro stats