Difficulty: Intermediate | Prerequisites: Two-sample independent t-test (Chapter 11), basic hypothesis testing, confidence intervals.
Chapter 12 picks up where the two-sample t-test left off. When you have three or more groups and want to know whether their population means differ, you cannot simply run multiple t-tests without inflating your Type I error rate. ANOVA gives you a single test that compares all group means at once by splitting the total variability in the data into "between-group" and "within-group" components. If you are comfortable with the two-sample independent t-test from Chapter 11 and with basic hypothesis testing, this chapter is the natural next step.
ANOVA (Analysis of Variance) tests whether the means of three or more populations are all equal by comparing the variance between groups to the variance within groups. A large F statistic means the group means are more spread out than you would expect from random noise alone, so you reject the null hypothesis that all means are equal. ANOVA tells you that at least one mean differs, but it does not tell you which one.
Factor
The categorical variable that defines the groups being compared. In a drug trial comparing three dosages, the factor is "dosage."
In simple terms, the factor is whatever divides your data into groups.
Level (group)
Each distinct value or category of the factor. If the factor is dosage with values 0 mg, 20 mg, and 40 mg, there are three levels. The number of levels is denoted k.
Think of it as: each level is one of the groups you are comparing.
One-way ANOVA
An ANOVA with a single factor. This is the only type covered in this chapter.
In simple terms, you are testing one grouping variable at a time.
Two-way ANOVA
An ANOVA with two factors. Not covered in this course but mentioned for context.
Sum of Squares, Between (SSA)
The total squared deviation of each group mean from the overall (grand) mean, weighted by group size. Measures how much the group means vary from each other.
Think of it as: the "signal" in your data, the part of the variability explained by group membership.
Sum of Squares, Error (SSE)
The total squared deviation of individual observations from their own group mean. Measures the variability within each group.
Think of it as: the "noise," the natural spread that has nothing to do with which group a subject belongs to.
Sum of Squares, Total (SST)
SSA + SSE. The total variability in the entire data set.
Mean Square (MS)
A sum of squares divided by its degrees of freedom: MS = SS / df. You will work with MSA (between) and MSE (within).
In simple terms, a mean square is just a variance estimate.
F Statistic (F_ts)
The test statistic for ANOVA, calculated as MSA / MSE. A ratio of between-group variance to within-group variance.
Think of it as: signal divided by noise. A large F means the group differences are big relative to the random scatter.
ANOVA Table
A structured summary that organises the degrees of freedom, sums of squares, mean squares, F statistic, and p-value in one place.
Residual (error term)
The difference between an observed value and its group mean. Denoted as the epsilon term in the ANOVA model. Residuals are assumed to be independent, normally distributed, and to have the same variance across all groups.
Degrees of Freedom (df)
For ANOVA: df_A = k - 1 (between groups), df_E = n - k (within groups), df_T = n - 1 (total). Here k is the number of groups and n is the total sample size.
The model is X_ij = mu_i + epsilon_ij, where X_ij is the observed data point, mu_i is the population mean for group i, and epsilon_ij is the residual (error).
Residuals are assumed to be independent and identically distributed as Normal with mean 0 and variance sigma-squared.
This single variance assumption is what drives the equal-variance requirement across all groups.
k independent simple random samples, one drawn from each population.
Each population follows a normal distribution with unknown mean mu_i.
All populations share the same variance sigma-squared (unknown). In practice, check this with the rule: s_max / s_min must be less than or equal to 2.
When sample sizes are equal across groups, ANOVA is more robust to violations of normality and equal variance.
Null hypothesis H_0: mu_1 = mu_2 = ... = mu_k (all population means are equal).
Alternative hypothesis H_a: at least one mu_i is different.
The alternative does NOT claim that all means differ from each other. It only says that at least one is different from the rest.
ANOVA works by comparing two sources of variability:
Between-group variance (MSA): how much the group means vary around the grand mean. This is the "signal."
Within-group variance (MSE): how much individual observations vary within their own group. This is the "noise."
If the group means are truly equal, MSA and MSE should be roughly the same, giving an F ratio near 1.
If at least one mean differs, MSA will be inflated relative to MSE, producing a large F.
The ANOVA table organises everything you need:
Source column: Factor A (between), Error (within), Total.
df column: k - 1, n - k, n - 1.
SS column: SSA, SSE, SST (and SST = SSA + SSE).
MS column: MSA = SSA / (k - 1), MSE = SSE / (n - k). No MS for Total.
F column: F_ts = MSA / MSE. Only one F value, in the Factor row.
p-value column: P(F >= F_ts) from an F distribution with df1 = k - 1 and df2 = n - k.
If the p-value is less than alpha, reject H_0. Conclude that at least one population mean is different.
If the p-value is greater than or equal to alpha, fail to reject H_0. There is not enough evidence to say the means differ.
ANOVA does not tell you which means differ. That requires a follow-up procedure (covered in Part 2 of these notes).
When k = 2, the F-test and the two-sample independent t-test give the same p-value. In fact, t_ts squared = F_ts.
With only two groups, the t-test is preferred because it is less restrictive (it does not require equal variances if you use the Welch version).
With three or more groups, ANOVA is the correct approach.
X_{ij} = \mu_i + \varepsilon_{ij}where epsilon_ij are i.i.d. Normal(0, sigma-squared).
MS = \frac{SS}{df}F_{ts} = \frac{MSA}{MSE} = \frac{\text{between-group variance}}{\text{within-group variance}}with numerator df = k - 1 and denominator df = n - k.
\frac{s_{\max}}{s_{\min}} \leq 2Source | df | SS | MS | F | p-value |
|---|---|---|---|---|---|
Factor A (between) | k - 1 | SSA | SSA / (k - 1) | MSA / MSE | P(F >= F_ts) |
Error (within) | n - k | SSE | SSE / (n - k) | ||
Total | n - 1 | SSA + SSE |
When k = 2:
t_{ts}^2 = F_{ts}and the p-values are identical.
ANOVA is used whenever researchers need to compare outcomes across three or more conditions. A pharmaceutical company testing three dosages of a drug on serotonin levels (exactly the Paxil example in this chapter) uses one-way ANOVA to determine whether dosage matters before investigating which specific dose is best. Agricultural scientists use it to compare crop yields across different fertiliser treatments, and manufacturers use it to compare product quality across multiple production lines.
Students often think the alternative hypothesis states that all means are different from each other. It does not. H_a only claims that at least one mean differs from the rest.
Students sometimes believe a significant ANOVA result tells you which groups are different. It does not. ANOVA only tells you that a difference exists somewhere. You need a post-hoc procedure (Tukey, Dunnett, etc.) to identify which pairs differ.
Students confuse SSA and SSE. Remember: SSA measures variability between group means (signal), while SSE measures variability within groups (noise). A common memory aid: A = Among groups, E = Error within groups.
Students occasionally forget that the equal-variance assumption applies to all k populations, not just any two. The check s_max / s_min <= 2 must hold across all groups simultaneously.
You will almost certainly be asked to state the hypotheses for ANOVA. Write H_0 in symbols (mu_1 = mu_2 = ... = mu_k) and H_a in words ("at least one mu_i is different").
Expect to be given partial ANOVA tables and asked to fill in the missing values. Know the relationships: SST = SSA + SSE, df_T = df_A + df_E, and MS = SS / df.
The F statistic is always MSA / MSE, never the other way round.
Know the three assumptions cold: independent SRSs, normality, equal variances.
Be prepared to explain why you cannot simply run repeated two-sample t-tests. The answer is Type I error inflation, which is the bridge to Chapter 12's multiple comparison methods.
True or False: The alternative hypothesis in ANOVA states that all population means are different from each other.
False. It states that at least one mean is different.
Fill in the blank: The F statistic is calculated as ______ divided by ______.
MSA divided by MSE (between-group variance divided by within-group variance).
True or False: If k = 2, the ANOVA F-test and the two-sample independent t-test produce different p-values.
False. They produce the same p-value, and t_ts squared = F_ts.
Fill in the blank: The degrees of freedom for the error term in a one-way ANOVA are ______.
n - k (total observations minus number of groups).
True or False: ANOVA requires that all populations have the same variance.
True. Check with s_max / s_min <= 2.
Q: An ANOVA test produces F_ts = 1.02 with a p-value of 0.38. At alpha = 0.05, what is your conclusion?
A: Fail to reject H_0. There is not sufficient evidence to conclude that the population means differ. Since the p-value (0.38) exceeds alpha (0.05), the differences in sample means could be due to chance alone.
Q: You have four treatment groups (k = 4) and a total of 20 observations (n = 20). What are the degrees of freedom for the between-group and within-group sources?
A: df_A = k - 1 = 3. df_E = n - k = 16.
Q: In a completed ANOVA table, SSA = 841.88 and SSE = 604.34. What is SST?
A: SST = SSA + SSE = 841.88 + 604.34 = 1446.22.
Q: A researcher wants to compare the means of five populations. Why is it inappropriate to run 10 separate two-sample t-tests at alpha = 0.05?
A: Running 10 separate tests inflates the overall Type I error rate well beyond 0.05. With k = 5 groups, the overall risk of at least one false rejection rises to roughly 0.40 (40%), making it far too likely that you will declare a difference that does not exist.
Q: An ANOVA has k = 3, n = 15, SSA = 841.88, SSE = 604.34. Complete the ANOVA table and state whether the result is significant at alpha = 0.05.
A: df_A = 2, df_E = 12. MSA = 841.88 / 2 = 420.94. MSE = 604.34 / 12 = 50.36. F_ts = 420.94 / 50.36 = 8.36. The critical F value at alpha = 0.05 with df (2, 12) is approximately 3.89, and the p-value is about 0.005. Since 8.36 > 3.89 (or equivalently, p < 0.05), reject H_0. At least one population mean is different.
This material builds directly on the two-sample independent t-test from Chapter 11. The pooled standard deviation concept carries over, and the equal-variance assumption is the same idea extended to k groups. ANOVA also sets the stage for the multiple comparison procedures (Tukey, Bonferroni, Dunnett) covered in the second half of Chapter 12, which answer the question ANOVA leaves open: which specific means differ?
In a broader statistics curriculum, one-way ANOVA leads to two-way ANOVA (multiple factors), repeated-measures ANOVA (dependent samples), and eventually to regression, which can be viewed as a generalisation of the same variance-partitioning idea.
ANOVA, analysis of variance, one-way ANOVA, F-test, F statistic, F distribution, between-group variance, within-group variance, MSA, MSE, SSA, SSE, SST, sum of squares, mean square, degrees of freedom, factor, level, group, treatment, ANOVA table, null hypothesis all means equal, Type I error, pooled variance, signal to noise, STAT 101, Purdue statistics, Chapter 12, comparing multiple means, one-factor ANOVA