One-Way ANOVA, STAT – Ch. 12 – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Two-sample t-tests, basic hypothesis testing framework

Big Picture

ANOVA extends the two-sample t-test to k groups. Instead of running many pairwise t-tests (which inflates Type I error), you partition the total variability into "between groups" and "within groups" and compare them with a single F-test. This chapter covers one-way ANOVA only, meaning one categorical factor. You should already be comfortable with hypothesis testing, pooled standard deviations, and the logic of p-values.

TL;DR

One-way ANOVA tests whether the means of k populations are all equal by comparing between-group variation to within-group variation. If the F-statistic is large (p-value small), at least two group means differ. Follow-up multiple comparison procedures (Tukey, Dunnett) then pinpoint which pairs are different.


Key Terms

One-way ANOVA

A hypothesis test that compares the means of k populations using a single categorical factor. In simple terms, it asks: "Are any of these group averages different from each other?"

Factor

The categorical variable that defines the groups being compared (e.g. gasoline brand, drug dosage). Think of it as the thing you are changing or grouping by.

Levels

The distinct categories or values of the factor. If you test three dosages (0 mg, 20 mg, 40 mg), k = 3 levels.

Response variable

The quantitative outcome you measure (e.g. serotonin level, fuel efficiency).

SSA (Sum of Squares for Factor A, between groups)

Measures how much the group means differ from the overall mean. Large SSA suggests the factor matters.

SSE (Sum of Squares for Error, within groups)

Measures variability of individual observations around their own group mean. This is the noise.

SST (Total Sum of Squares)

The total variability in the data. SST = SSA + SSE.

MSA (Mean Square for Factor)

SSA divided by its degrees of freedom (k - 1). An estimate of variance that reflects group differences.

MSE (Mean Square for Error)

SSE divided by its degrees of freedom (n - k). Always an unbiased estimator of the common variance, regardless of whether H0 is true.

F-statistic (F_ts)

MSA / MSE. Under H0 it sits near 1. Large values signal that at least two means differ.

Family-Wise Error Rate (FWER)

The probability that at least one of a set of pairwise comparisons produces a Type I error. This is what multiple comparison methods control.

Tukey's HSD (Honestly Significant Difference)

A multiple comparison method that compares all pairs of means while controlling the FWER. The most commonly used post-hoc procedure.

Dunnett's method

A multiple comparison method for comparing each treatment group to a single control group. Uses k - 1 comparisons instead of k-choose-2.

Bonferroni correction

A conservative multiple comparison adjustment that divides the significance level by the number of comparisons. Rarely preferred because Tukey and Dunnett are more powerful.


Core Content

ANOVA Setup

  • Identify the factor (the categorical grouping variable) and its levels (how many groups).

  • Identify the response variable (the quantitative outcome being measured).

  • Each level of the factor corresponds to a population whose mean you want to compare.

  • Example: "Do five brands of gasoline affect fuel efficiency?" The factor is brand (k = 5 levels), the response is miles per gallon.

Three Assumptions

  1. Simple random samples (SRS): k independent random samples, one from each population.

  1. Normality: Each population is normally distributed. Check with histograms and normal probability plots. The F-test is robust to mild departures from normality.

  1. Equal variances: All k populations share the same variance. Check by computing the ratio of the largest sample standard deviation to the smallest. If s_max / s_min is 2 or less, the assumption is considered satisfied.

The ANOVA Table

The one-way ANOVA table breaks total variability into two pieces:

Source

df

SS

MS

F

Factor A (between)

k - 1

SSA

MSA = SSA / (k - 1)

MSA / MSE

Error (within)

n - k

SSE

MSE = SSE / (n - k)

Total

n - 1

SST = SSA + SSE

s = sqrt(MSE)

  • Between-group variation (SSA) captures how far each group mean sits from the overall mean.

  • Within-group variation (SSE) captures how far individual observations sit from their own group mean.

  • If H0 is true, both MSA and MSE estimate the same common variance, so F sits near 1.

  • If H0 is false, MSA overestimates the variance, pushing F well above 1.

The Four-Step Hypothesis Test

Step 1 - Parameters: Define each population mean in context (e.g. "let mu_0 = population mean serotonin for the 0 mg group").

Step 2 - Hypotheses:

  • H0: mu_1 = mu_2 = ... = mu_k (all population means are equal)

  • Ha: At least two of the mu_i are different

  • The alternative is always two-sided. You never specify which means differ.

Step 3 - Test statistic and p-value:

  • F_ts = MSA / MSE

  • The p-value comes from an F-distribution with df1 = k - 1 (numerator) and df2 = n - k (denominator).

  • In R: pf(fts, df1, df2, lower.tail = FALSE)

  • In R with raw data: fit <- aov(response ~ factor, data = df) then summary(fit)

Step 4 - Decision and conclusion:

  • If p-value < alpha, reject H0. State the conclusion in context, including the p-value.

  • If p-value >= alpha, fail to reject H0.

t-Test vs F-Test When k = 2

When there are only two groups, the pooled two-sample t-test and one-way ANOVA give the same result: t_ts squared = F_ts.

Feature

Two-sample t

One-way ANOVA

Variance assumption

Same or different

Same (pooled)

Tails

One or two

Always two-tailed

Null value of difference

Any real number

Always 0

Number of groups

2 only

2 or more

The t-test is preferred for two groups because it allows one-tailed tests and non-zero null differences.

Multiple Comparison Procedures

Once ANOVA rejects H0, you know at least two means differ, but not which ones. Multiple comparisons identify the specific pairs.

Why not just run many two-sample t-tests?

  • Doing so inflates the FWER. Each individual test controls its own alpha, but the probability that at least one test commits a Type I error grows with the number of comparisons.

  • ANOVA pools all groups for a better variance estimate.

Number of pairwise comparisons: C(k, 2) = k(k - 1) / 2

Choosing a method:

  • Comparing all pairs with no prior expectation: use Tukey's HSD.

  • Comparing each treatment to a single control: use Dunnett's method.

Tukey's procedure:

  • Compute the Tukey critical value: Q = qtukey(1 - alpha, k, n - k), then t** = Q / sqrt(2).

  • For each pair (i, j), build the confidence interval: (x-bar_i - x-bar_j) +/- t** * SE, where SE = sqrt(MSE * (1/n_i + 1/n_j)).

  • If 0 is inside the interval, those two means are not significantly different. If 0 is outside, they are.

  • In R: TukeyHSD(fit, conf.level = 0.95)

Dunnett's procedure:

  • Uses a smaller critical value than Tukey (fewer comparisons to protect against), so its intervals are narrower.

  • Only k - 1 comparisons instead of k(k - 1) / 2.

  • Tukey's critical value > Dunnett's critical value, which means Tukey's intervals are wider. A pair declared significant by Dunnett may not be significant by Tukey.

Graphical display of results (Tukey):

  • Order the group means from lowest to highest.

  • Draw a line underneath groups whose means are not significantly different.

  • Groups connected by the same line are statistically indistinguishable.


Formulas

SSA = \sum_{i=1}^{k} n_i (\bar{x}_{i.} - \bar{x}_{..})^2
SSE = \sum_{i=1}^{k} \sum_{j=1}^{n_i} (x_{ij} - \bar{x}_{i.})^2 = \sum_{i=1}^{k} (n_i - 1) s_i^2
SST = SSA + SSE = \sum_{i=1}^{k} \sum_{j=1}^{n_i} (x_{ij} - \bar{x}_{..})^2
MSA = \frac{SSA}{k - 1}, \quad MSE = \frac{SSE}{n - k}, \quad F_{ts} = \frac{MSA}{MSE}
s = \sqrt{MSE}

Tukey confidence interval for comparing means i and j:

(\bar{x}_{i.} - \bar{x}_{j.}) \pm \frac{Q_{\alpha, k, n-k}}{\sqrt{2}} \cdot \sqrt{MSE \left(\frac{1}{n_i} + \frac{1}{n_j}\right)}

R code for the p-value (given the test statistic):

pf(fts, df1, df2, lower.tail = FALSE)

R code for the full ANOVA + Tukey procedure:

fit <- aov(response ~ factor, data = dataframe) summary(fit) TukeyHSD(fit, ordered = TRUE, conf.level = C)


Common Misconceptions

  • "Rejecting H0 tells us which means differ." It does not. ANOVA only says at least two means are different. You need a follow-up multiple comparison procedure (Tukey, Dunnett) to identify the specific pairs.

  • "We can just run many two-sample t-tests instead of ANOVA." This inflates the family-wise error rate. Each t-test controls its own alpha, but the probability of at least one false positive across all tests grows quickly. ANOVA controls this in a single test.

  • "If the standard deviations are not identical, ANOVA cannot be used." The assumption is equal population variances, and we check it with the ratio s_max / s_min. If that ratio is 2 or less, the assumption is considered met. Exact equality is not required.

  • "The alternative hypothesis says all means are different from each other." The alternative only states that at least two are different. Some means may still be equal under Ha.


Why It Matters / Exam Flags

  • You will be asked to complete an ANOVA table by hand. Know how to compute df, SS, MS, and F from partial information and use the relationships (SST = SSA + SSE, df_T = df_A + df_E) to cross-check.

  • Expect to write out all four steps of the hypothesis test in context: define parameters, state hypotheses, compute the test statistic and p-value, then state the conclusion with the p-value included.

  • You must know when to use Tukey (all pairwise) vs Dunnett (comparisons to a control). Picking the wrong method on the exam is a common deduction.

  • Interpreting TukeyHSD output from R is frequently tested: if zero is inside a confidence interval, those two means are not significantly different.

  • Creating the graphical display (ordered means with underlines connecting non-significant pairs) is a required skill for Tukey problems.

  • Remember that Tukey's critical value is larger than Dunnett's, so Tukey intervals are wider. A pair significant under Dunnett may lose significance under Tukey.

  • The relationship t_ts squared = F_ts when k = 2 is a conceptual point that appears on exams.


Quick Self-Test

  1. True or False: The alternative hypothesis in one-way ANOVA specifies which means are different. (False, it only says at least two differ.)

  1. Fill in the blank: df_A = ____, df_E = ____, df_T = ____ for k = 4 groups and n = 28 total observations. (3, 24, 27)

  1. True or False: MSE changes depending on whether H0 is true or false. (False, MSE is always an unbiased estimator of the common variance.)

  1. Fill in the blank: If the ratio s_max / s_min is ____ or less, the equal-variance assumption is considered valid. (2)

  1. True or False: Tukey's method has a smaller critical value than Dunnett's. (False, Tukey's is larger because it protects against more comparisons.)


Practice Q&A

Q: A study tests four fertilisers on crop yield with n_i = 10 per group (n = 40). The ANOVA table shows SSA = 120 and SSE = 360. Compute MSA, MSE, and the F-statistic.

A: df_A = 4 - 1 = 3, df_E = 40 - 4 = 36. MSA = 120 / 3 = 40. MSE = 360 / 36 = 10. F_ts = 40 / 10 = 4.0.

Q: Using the result above, the p-value is 0.015. At alpha = 0.05, what is your conclusion? Write it in context.

A: We reject H0 because 0.015 < 0.05. The data provides strong support (p = 0.015) for the claim that the population mean crop yield differs for at least two of the four fertilisers.

Q: After rejecting H0, you want to compare all pairs. Which multiple comparison method do you use, and why?

A: Tukey's HSD, because we have no designated control group and we want to compare every pair of means.

Q: A Tukey confidence interval for fertiliser A minus fertiliser B is (-3.2, 1.8). Are these two means significantly different?

A: No. Zero is inside the interval, so we fail to reject H0 for this pair. Fertilisers A and B are not significantly different.

Q: In the same study, suppose fertiliser A is a control (no fertiliser). You want to compare each treatment to the control only. Which method is appropriate?

A: Dunnett's method, because there is a designated control and we only need k - 1 = 3 comparisons rather than all 6 pairwise comparisons.

Q: Why does Tukey sometimes fail to detect a difference that Dunnett detects?

A: Tukey protects against more comparisons (all pairs), so its critical value is larger, making its confidence intervals wider. A difference that falls outside Dunnett's narrower interval may fall inside Tukey's wider one.


Connections to Other Topics

ANOVA builds directly on the two-sample t-test from earlier in the course. The pooled variance idea is the same; ANOVA just extends it to k groups. The F-distribution here reappears in linear regression (Chapter 13), where the regression ANOVA table tests whether the slope is significantly different from zero. Understanding SSA, SSE, and SST now will make the regression ANOVA table feel familiar.

Related Terms / Search Tags

One-way ANOVA, analysis of variance, F-test, F-statistic, between-group variation, within-group variation, SSA, SSE, SST, MSA, MSE, degrees of freedom, Tukey HSD, Tukey's honestly significant difference, Dunnett's test, Bonferroni, multiple comparisons, family-wise error rate, FWER, post-hoc test, pairwise comparisons, completely randomised design, pooled variance, aov() in R, summary(), TukeyHSD(), pf(), qtukey()