Difficulty: Intermediate to Advanced | Prerequisites: Descriptive Statistics, Probability, and Inference Foundations notes
This final set of notes covers the more advanced inference topics on the STAT 101 final: ANOVA for comparing three or more group means, the Tukey method for figuring out which groups differ, linear regression for modelling relationships, and the decision framework for choosing the right test in a given scenario. Master this material and you can handle the hardest questions on the exam.
ANOVA (Analysis of Variance)
A method for testing whether the means of three or more groups are all equal. Despite its name, it works by comparing variances. Think of it as the multi-group version of the two-sample t-test.
F-statistic (ANOVA)
The ratio of between-group variability to within-group variability. A large F means the group means differ more than you would expect from random variation alone.
Between-group variability
How much the group means differ from the overall mean. If this is large relative to within-group variability, at least one group is likely different.
Within-group variability
How much individual observations within each group differ from their own group mean. This is the baseline noise.
Multiple comparisons problem
When you perform many pairwise tests simultaneously, the probability of at least one false positive (Type I error) grows rapidly. With k groups and all pairwise comparisons, the chance of at least one false alarm is much higher than alpha.
Tukey's Honest Significant Difference (HSD)
A post-hoc method that adjusts for the multiple comparisons problem. It constructs simultaneous confidence intervals for every pairwise difference in means, controlling the family-wise error rate.
Linear regression
A method for modelling the relationship between a quantitative response variable and one or more explanatory variables using a straight line. The fitted line minimises the sum of squared residuals.
Regression line (least-squares line)
The line y-hat = b0 + b1 × x, where b0 is the y-intercept and b1 is the slope. Each point on the line represents the predicted mean of y for a given value of x.
R-squared (coefficient of determination)
The proportion of variability in the response variable that is explained by the linear relationship with the explanatory variable. Ranges from 0 to 1. Think of it as the fraction of the outcome that the model accounts for.
Residual
The difference between an observed value and the value predicted by the regression line: y minus y-hat. Residuals are used to check model assumptions.
Confidence interval for the mean response
An interval estimating the average y-value for all individuals at a specific x. It is narrower than a prediction interval because it estimates a mean, not an individual.
Prediction interval
An interval estimating a single new individual y-value at a specific x. It is wider than a confidence interval for the mean because individual observations vary more than averages.
Two-sample independent
Comparing means from two separate, unrelated groups (e.g. treatment vs. control with different people in each). Observations in one group have no natural pairing with observations in the other.
Two-sample paired
Comparing means where each observation in one group is naturally paired with one in the other (e.g. before and after measurements on the same person, or twins). The analysis uses the differences within each pair.
This question trips students up because the name "Analysis of Variance" seems misleading when the goal is to compare means.
The logic: if all group means are truly equal, then the variation between group means should be no larger than what you would expect from the random variation within groups. ANOVA compares these two sources of variation through the F-statistic.
If the group means are all the same, the between-group variability should be small relative to within-group variability, and F will be close to 1.
If at least one group mean is different, the between-group variability will be inflated, and F will be large.
Comparing means directly, one pair at a time, would not account for the overall pattern of variation, and it would create a multiple comparisons problem.
Suppose you have 4 groups and you test all 6 pairwise comparisons using two-sample t-tests at alpha = 0.05. Each individual test has a 5% chance of a false positive, but across 6 tests the probability of at least one false positive is much higher (roughly 26% if the tests were independent).
This is why ANOVA exists: the F-test provides a single overall test of whether any differences exist, at the stated alpha level. Only if the F-test rejects do you follow up with pairwise comparisons.
For those pairwise comparisons, you need a method that controls the family-wise error rate, such as Tukey's HSD. Using ordinary two-sample t-tests inflates the overall Type I error rate.
After ANOVA rejects the null, Tukey's method tells you which specific pairs of groups differ.
The interval: for each pair of groups, Tukey constructs a confidence interval for the difference in means. If the interval does not contain zero, the two groups are significantly different.
The visual representation: Tukey output is often shown as a set of intervals plotted on a number line. Each interval represents a pairwise difference. Intervals that cross the zero line indicate no significant difference. Intervals entirely to one side of zero indicate a significant difference.
The Tukey method controls the overall Type I error rate across all comparisons simultaneously, so you can trust the conclusions even when many pairs are tested.
For a given value of x, the point on the regression line (y-hat) is the estimated mean of y for all individuals with that x-value. It is a prediction of the average response, not a prediction of any single individual's response.
For example, if the regression equation for exam score based on study hours is y-hat = 40 + 5x, then at x = 6 hours, y-hat = 70. This means the estimated average exam score for all students who study 6 hours is 70, not that every such student will score exactly 70.
What R-squared tells you:
The proportion of variability in y that is explained by the linear relationship with x.
An R-squared of 0.72 means 72% of the variation in the response is accounted for by the model.
Higher R-squared means a tighter fit; the data points cluster more closely around the regression line.
What R-squared does not tell you:
Whether the relationship is causal. A high R-squared says the two variables move together linearly; it says nothing about why.
Whether a linear model is appropriate. You can have a high R-squared even when a curved model would fit better. Always check residual plots.
The direction or size of the effect. The slope (b1) tells you the direction and rate of change; R-squared does not.
Both are centred at y-hat for a given x, but they answer different questions and have different widths.
Confidence interval for the mean response at x: estimates where the average y falls for all individuals with that x-value. It accounts for uncertainty in the position of the regression line.
Prediction interval for an individual new observation at x: estimates where a single new y-value will fall. It accounts for the same uncertainty in the line plus the additional scatter of individual observations around the line.
The prediction interval is always wider because individual observations are more variable than averages. As n grows, the confidence interval for the mean narrows toward zero width, but the prediction interval remains wide because individual variability does not shrink with more data.
A common analogy: the confidence interval is like estimating the average temperature in July. The prediction interval is like predicting the temperature on one specific July day. The day-to-day variation makes the second task harder.
The exam will give you a scenario and ask which test to use. Work through these questions in order:
1. What type of data do you have?
Quantitative response (means): use t-tests, ANOVA, or regression.
Categorical response (proportions): use z-tests for proportions or chi-squared tests.
2. How many groups or variables?
One group: one-sample t-test (for a mean) or one-sample z-test (for a proportion).
Two groups: two-sample t-test or paired t-test.
Three or more groups: ANOVA.
Relationship between two quantitative variables: linear regression.
3. If two groups, are the samples independent or paired?
Independent: the subjects in one group have no connection to the subjects in the other. Different people, different items, different experimental units. Use a two-sample independent t-test.
Paired: each observation in one group has a natural partner in the other. Before/after on the same person, left eye vs. right eye, twins. Use a paired t-test (which is really just a one-sample t-test on the differences).
How to tell: ask whether you can match up the observations one-to-one in a way that is meaningful. If yes, it is paired. If you can rearrange the observations in one group without losing information, it is independent.
Comparing a sample mean to a known value: one-sample t-test.
Comparing two means, separate groups: two-sample independent t-test.
Comparing two means, same subjects measured twice: paired t-test.
Comparing three or more group means: one-way ANOVA, then Tukey if ANOVA rejects.
Modelling a linear relationship between two quantitative variables: simple linear regression.
Comparing a sample proportion to a known value: one-sample z-test for proportions.
Comparing two proportions from independent groups: two-sample z-test for proportions.
Students often think a high R-squared means the linear model is correct. It does not. A parabolic relationship can produce a deceptively high R-squared when fitted with a straight line. Always check residual plots.
Students sometimes confuse the confidence interval for the mean response with the prediction interval. The prediction interval is always wider. If you are predicting where a single new observation will land, you need the prediction interval.
Students assume that if ANOVA does not reject, all groups have exactly the same mean. It only means there is not enough evidence to conclude the means differ. The true means could still be somewhat different.
Students try to use multiple two-sample t-tests instead of ANOVA when comparing three or more groups. This inflates the Type I error rate because it ignores the multiple comparisons problem.
⚠️ Expect a scenario question asking you to choose between a paired t-test and an independent two-sample t-test. The key is whether observations are naturally matched.
⚠️ Know why ANOVA uses variances to compare means. Be able to explain the logic of the F-statistic in plain English.
⚠️ Be prepared to interpret Tukey output: which pairs differ (interval does not contain zero) and which do not.
⚠️ Know the distinction between R-squared and causation. A question will likely test whether you over-interpret R-squared.
⚠️ Be able to explain in words why a prediction interval is wider than a confidence interval for the mean at the same x-value.
True or false: if a Tukey confidence interval for the difference between groups A and B contains zero, we conclude the means are significantly different. (False, it means they are NOT significantly different.)
Fill in the blank: the prediction interval is always ______ than the confidence interval for the mean response at the same x-value. (Wider)
True or false: R-squared = 0.85 means the explanatory variable causes 85% of the change in the response. (False, R-squared measures explained variation, not causation.)
Fill in the blank: a study measures each participant's blood pressure before and after a medication. This is a ______ design. (Paired)
True or false: running six separate two-sample t-tests across four groups at alpha = 0.05 keeps the overall Type I error rate at 5%. (False, it inflates the rate well above 5%.)
Q: An ANOVA test for four fertiliser types yields a p-value of 0.001. A student concludes that all four fertiliser types produce different mean yields. Is this conclusion correct?
A: No. The ANOVA F-test only tells you that at least one group mean differs from the others. It does not tell you which pairs are different. A Tukey post-hoc analysis is needed to identify the specific pairs.
Q: A Tukey analysis produces the following 95% confidence interval for the difference in means between Group A and Group B: (minus 2.1, 3.4). Are Groups A and B significantly different?
A: No. The interval contains zero, which means the difference in means is not statistically significant at the 0.05 level.
Q: In a regression of salary on years of experience, R-squared = 0.64. A student says "years of experience causes 64% of salary variation." What is wrong?
A: R-squared measures explained variation, not causation. The correct interpretation is that 64% of the variability in salary is explained by the linear relationship with years of experience. Other factors (industry, education, location) may be the true drivers.
Q: A researcher wants to predict the exam score of one specific student who studies 8 hours. Should she use a confidence interval for the mean or a prediction interval?
A: A prediction interval, because she is predicting a single individual's score, not the average score of all students who study 8 hours.
Q: Ten runners have their 5K times measured before and after a training programme. Is this a two-sample independent or paired scenario?
A: Paired. The same runners are measured twice (before and after), so each "before" observation is naturally paired with its "after" observation. A paired t-test on the differences is appropriate.
Q: Why can you not simply run three separate two-sample t-tests to compare three group means at alpha = 0.05?
A: Because each test has a 5% chance of a false positive, and running three tests raises the overall probability of at least one false positive well above 5%. ANOVA handles this by performing one overall test, and Tukey's method adjusts the pairwise comparisons to control the family-wise error rate.
ANOVA builds directly on the inference foundations from the previous notes: it uses the same logic of null hypotheses, p-values, and Type I errors, extended to multiple groups. Regression connects to confidence intervals (the confidence interval for the mean response) and hypothesis testing (testing whether the slope is zero). The choice between paired and independent designs ties back to experimental design principles: pairing reduces variability by controlling for subject-level differences, which increases power.
ANOVA, analysis of variance, F-test, F-statistic, between-group variability, within-group variability, multiple comparisons, family-wise error rate, Tukey HSD, Tukey's honest significant difference, post-hoc test, pairwise comparison, linear regression, least squares, regression line, slope, intercept, R-squared, coefficient of determination, residual, residual plot, confidence interval for the mean, prediction interval, two-sample independent t-test, paired t-test, matched pairs, one-sample t-test, one-way ANOVA, choosing inference method