One-Way ANOVA Foundations and Assumptions, STAT 101 – Study Notes
offline

Source: Comprehensive Guide to One-Way ANOVA Analysis Techniques, STAT 101 (Purdue University)

Difficulty: Intermediate | Prerequisites: Basic probability, descriptive statistics, two-sample t-tests.


Big Picture

One-way ANOVA sits at the point in an introductory statistics course where you move beyond comparing two groups and start handling three or more at once. If you are comfortable with two-sample t-tests, you already have the intuition: ANOVA generalises that logic by comparing variation between groups to variation within groups, producing a single test statistic (the F-statistic) that tells you whether any of the group means differ. You should already be familiar with means, standard deviations and the idea of a sampling distribution before reading these notes.


TL;DR

One-way ANOVA tests whether the means of three or more independent groups are all equal. It does this by partitioning total variability into between-group and within-group components, then comparing them with an F-statistic. A small p-value means at least one group mean is different, but not which one.

Key Terms

One-way ANOVA (Analysis of Variance)

A statistical method for comparing the means of three or more independent populations using a single categorical factor. In simple terms, it is the tool you reach for when a t-test is not enough because you have more than two groups.

Factor

The categorical independent variable that defines the groups being compared (e.g. gasoline brand, drug dosage, teaching method). Think of it as the "thing you are changing" across groups.

Levels

The distinct categories or groups within a factor, denoted by k. If the factor is "gasoline brand" and there are five brands, k = 5. Levels can be qualitative or quantitative values treated as categories.

Between-group variation (SSA / MSA)

The portion of total variability that comes from differences among the group means. A large between-group variation relative to within-group variation suggests the factor matters.

Within-group variation (SSE / MSE)

The portion of total variability that comes from individual differences inside each group. This is the "noise" or random scatter that would exist even if the factor had no effect.

Null hypothesis (H₀)

The claim that all population means are equal: μ₁ = μ₂ = ... = μₖ. In plain terms, "the factor makes no difference."

Alternative hypothesis (H₁ / Hₐ)

The claim that at least one population mean differs from the others. It does not say which one or by how much.

Independence (assumption)

Each observation is drawn as a simple random sample, and the samples from different groups do not influence one another.

Normality (assumption)

The data within each population are approximately normally distributed. Checked with histograms or normal probability plots.

Homogeneity of variances / homoscedasticity (assumption)

The population variances are equal across all groups. A common rule of thumb: the ratio of the largest to the smallest standard deviation should be less than 2. Levene's test is a formal check.

Core Content

What ANOVA Does and Why It Exists

  • A two-sample t-test compares two group means. With three or more groups, running every pairwise t-test inflates the Type I error rate (the chance of a false positive). ANOVA avoids this by testing all groups in a single step.

  • The core logic: if the group means really are equal, the spread of those means should be no larger than what random sampling noise alone would produce. ANOVA formalises that comparison.

Factors and Levels

  • The factor is the categorical independent variable (e.g. "brand of fertiliser").

  • The levels are the specific categories within the factor (e.g. Brand A, Brand B, Brand C). The number of levels is k.

  • Levels can be naturally categorical or can be quantitative values grouped into categories (e.g. low / medium / high dosage).

Why Variance, Not Just Means?

  • Sample means will almost always differ a little, even when the population means are identical, because of random variation.

  • ANOVA asks whether the observed differences among sample means are larger than what within-group noise would predict.

  • It partitions total variability into two sources:

    • Between-group variation (SSA): how far each group mean sits from the grand mean.

    • Within-group variation (SSE): how far individual observations sit from their own group mean.

  • If between-group variation is large relative to within-group variation, the factor is likely having a real effect.

Hypotheses

  • H₀: μ₁ = μ₂ = ... = μₖ (all population means are equal).

  • H₁: At least one μᵢ ≠ μⱼ for some i ≠ j.

  • Rejecting H₀ tells you that at least one group differs, but does not tell you which group or groups. You need post-hoc tests (e.g. Tukey's HSD) for that.

The Three Assumptions

  • Independence: observations are simple random samples, and the groups do not influence each other. This is a study-design issue, not something you test statistically.

  • Normality: data within each group are approximately normally distributed. Check with histograms or Q-Q plots. ANOVA is fairly robust to mild departures from normality, especially with larger sample sizes.

  • Homogeneity of variances (homoscedasticity): all groups share the same population variance. Rule of thumb: the largest group standard deviation divided by the smallest should be less than 2. Levene's test provides a formal check. Balanced designs (equal sample sizes per group) make ANOVA more robust to violations here.

The ANOVA Model

  • Each observation is modelled as: Xᵢⱼ = μᵢ + εᵢⱼ

    • Xᵢⱼ is the j-th observation in group i.

    • μᵢ is the true population mean of group i.

    • εᵢⱼ is random error, assumed to be independent and normally distributed with mean 0 and constant variance σ².

  • In words: every data point equals its group's true mean plus some random noise.

Formulas

The ANOVA model:

X_{ij} = \mu_i + \varepsilon_{ij}

where εᵢⱼ ~ N(0, σ²), independently.

The key relationship linking the three sums of squares:

SST = SSA + SSE

In words: total variation = between-group variation + within-group variation.

Real-World Applications

ANOVA is how a pharmaceutical company tests whether three different dosages of a drug produce different average blood-pressure reductions, or how an agricultural researcher compares crop yields across four fertiliser types. Any time you need a single, controlled test across multiple groups, ANOVA is the standard tool.


Common Misconceptions

  • "A significant ANOVA result tells you which groups differ." It does not. ANOVA only says that at least one mean is different. You need a follow-up procedure (Tukey's HSD, Bonferroni correction, etc.) to identify which pairs differ.

  • "Failing to reject H₀ proves the means are equal." It does not prove equality. It may simply reflect low statistical power, small sample sizes or high within-group variability.

  • "You can just run multiple t-tests instead." You can, but the familywise error rate inflates quickly. With five groups there are 10 pairwise comparisons; at α = 0.05 each, the chance of at least one false positive climbs well above 5%.

  • "ANOVA requires perfectly normal data." The test is reasonably robust to moderate departures from normality, especially with balanced designs and larger samples. Severe skew or heavy outliers are more of a concern.


Why It Matters / Exam Flags

⚠️ Be able to state the null and alternative hypotheses in words and in notation.

⚠️ Know the three assumptions and how to check each one (independence by design, normality by plots, equal variances by the s_max / s_min < 2 rule or Levene's test).

⚠️ Understand that ANOVA partitions total variation: SST = SSA + SSE. Exams often give two of three and ask for the third.

⚠️ A significant F-test does not identify which groups differ. Expect a question that tests this distinction.

Quick Self-Test

  1. True or false: ANOVA can only compare exactly three groups. False. It works for three or more.

  1. Fill in the blank: The total sum of squares equals ______ + ______. SSA + SSE.

  1. True or false: If the F-test is significant, you immediately know which group mean is different. False. You need post-hoc comparisons.

  1. True or false: ANOVA assumes the population variances are equal across groups. True.

  1. Fill in the blank: In the ANOVA model, εᵢⱼ is assumed to follow a ______ distribution with mean ______ and variance ______. Normal, 0, σ².


Practice Q&A

Q: A researcher compares the mean reaction times of participants under four different lighting conditions. State the null and alternative hypotheses.

A: H₀: μ₁ = μ₂ = μ₃ = μ₄ (mean reaction time is the same under all four conditions). H₁: At least one μᵢ differs from another.

Q: Why is it problematic to run six separate t-tests to compare four group means instead of using ANOVA?

A: Each t-test carries a probability of a Type I error (α). Running multiple tests inflates the overall (familywise) error rate well beyond the intended α = 0.05, making it more likely you will find a "significant" difference by chance alone.

Q: Name the three assumptions of one-way ANOVA and briefly describe how each is checked.

A: (1) Independence, ensured by study design (random sampling, no group influences another). (2) Normality, checked via histograms or Q-Q plots for each group. (3) Homogeneity of variances, checked by the s_max / s_min < 2 rule or Levene's test.

Q: A study finds a non-significant F-test result. A student concludes that all group means are proven equal. What is wrong with this conclusion?

A: Failing to reject H₀ does not prove H₀ is true. The study may have had insufficient power (too few observations or too much within-group variability) to detect a real difference.


Connections to Other Topics

This material connects directly to two-sample t-tests: when k = 2, the F-statistic equals the square of the t-statistic, and the p-values are identical. It also sets the stage for post-hoc multiple comparisons (Tukey's HSD, Bonferroni) and for two-way ANOVA, which introduces a second factor. Understanding how variance is partitioned here is foundational for regression analysis, where a similar decomposition (SSR + SSE = SST) appears.


Related Terms / Search Tags

One-way ANOVA, analysis of variance, F-test, F-statistic, between-group variation, within-group variation, SSA, SSE, SST, MSA, MSE, factor, levels, null hypothesis, alternative hypothesis, homoscedasticity, homogeneity of variances, Levene's test, independence assumption, normality assumption, ANOVA model, comparing multiple means, Type I error inflation, familywise error rate, STAT 101, intro statistics, Purdue