Difficulty: Advanced | Prerequisites: Chapters 1 to 10 study notes
Chapter 11 extends hypothesis testing to compare the means of three or more groups at once using one-way ANOVA. Chapter 12 introduces simple linear regression, where you model the relationship between two quantitative variables with a straight line. Together, these chapters are the culmination of the course: you will need everything you have learnt about descriptive statistics, probability, and inference. Regression in particular is often the centrepiece of the final exam.
ANOVA tests whether the means of several groups differ by comparing between-group variation (SSA) to within-group variation (SSE) via an F-statistic. Linear regression fits a line y-hat = b0 + b1*x to data, measures fit with r and r-squared, and allows you to build confidence intervals for the slope, the mean response, and individual predictions.
One-way ANOVA (Analysis of Variance)
A method for testing whether the means of three or more groups are all equal. It splits total variation into a between-group component and a within-group component.
Think of it as asking: is the spread between group averages large enough, relative to the noise within groups, to conclude the groups differ?
SSA (Sum of Squares for Factor A)
The between-group variation. Measures how much the group means differ from the overall mean.
SSE (Sum of Squares for Error)
The within-group variation. Measures how much individual observations differ from their own group mean.
SST (Sum of Squares Total)
SSA + SSE. The total variation in the data.
MSA and MSE (Mean Squares)
SSA divided by its degrees of freedom (k minus 1) gives MSA. SSE divided by its degrees of freedom (N minus k) gives MSE.
F-statistic
MSA / MSE. A large F suggests the group means are not all equal.
Tukey method
A post-hoc comparison procedure used after a significant ANOVA to determine which specific pairs of group means differ.
Response variable (Y)
The outcome you are trying to explain or predict. Always plotted on the vertical axis.
Explanatory variable (X)
The variable you believe explains or predicts changes in Y. Always plotted on the horizontal axis.
Scatterplot
A graph of paired (x, y) observations. Used to assess the direction (positive/negative), form (linear/nonlinear), and strength of an association.
Linear regression model
Y = beta_0 + beta_1 * X + epsilon. The population model with an error term epsilon.
Least squares regression line
The fitted line y-hat = b0 + b1 * x that minimises the sum of squared residuals.
Residual
e_i = y_i minus y-hat_i. The vertical distance between an observed point and the fitted line.
Sample correlation coefficient (r)
Measures the strength and direction of a linear relationship. Ranges from minus 1 to plus 1.
Coefficient of determination (r-squared)
The proportion of variation in Y explained by the regression on X. Equals SSR / SST.
In simple terms, it tells you what percentage of the outcome's variability your model accounts for.
Prediction interval
A range for a single new observation at a given x-value. Wider than a confidence interval for the mean response because it also accounts for individual variability.
Source | df | SS | MS |
|---|---|---|---|
Factor A | k minus 1 | SSA = sum of n_i * (x-bar_i minus x-bar..)^2 | MSA = SSA / dfa |
Error | N minus k | SSE = sum of (n_i minus 1) * s_i^2 | MSE = SSE / dfe |
Total | N minus 1 | SST = SSA + SSE |
SST = SSA + SSE
dft = dfa + dfe
F = MSA / MSE, with df1 = dfa, df2 = dfe
Compare pairs of group means: x-bar_i minus x-bar_j, plus or minus t** * SE
t** = Q_(alpha, k, N minus k) / sqrt(2)
SE = sqrt(MSE * (1/n_i + 1/n_j)), which simplifies to sqrt(2 * MSE / n_i) when group sizes are equal
Groups whose intervals overlap are not significantly different
Direction: positive association (upward trend) or negative association (downward trend)
Form: linear or nonlinear
Strength: how tightly the points cluster around the trend
Outliers: points that fall away from the overall pattern
No association: a cloud of points with no discernible pattern
Population model: Y = beta_0 + beta_1 * X + epsilon
Fitted (estimated) line: y-hat = b0 + b1 * x
Slope: b1 = S_XY / S_XX
Intercept: b0 = y-bar minus b1 * x-bar
Residual: e_i = y_i minus y-hat_i
Source | df | SS | MS |
|---|---|---|---|
Regression | 1 | SSR = sum of (y-hat_i minus y-bar)^2 | MSR = SSR / dfr |
Error | n minus 2 | SSE = sum of (y_i minus y-hat_i)^2 | MSE = SSE / dfe |
Total | n minus 1 | SST = sum of (y_i minus y-bar)^2 |
SST = SSR + SSE
SSR = b1 * S_XY
Sample correlation: r = S_XY / sqrt(S_XX * S_YY)
r ranges from minus 1 (perfect negative) to plus 1 (perfect positive)
Coefficient of determination: r^2 = SSR / SST
r^2 tells you the fraction of variation in Y explained by the linear relationship with X
Standard deviation about the regression line: s = sqrt(MSE)
Model utility test: F_ts = MSR / MSE (tests whether beta_1 = 0)
Confidence interval for beta_1: b1 +/- t_(alpha/2, n minus 2) * sqrt(MSE / S_XX)
df = n minus 2 for all regression inference
Estimated value: y-hat at x*
SE = sqrt(MSE * (1/n + (x* minus x-bar)^2 / S_XX))
Use t_(alpha/2, n minus 2)
Same estimated value: y-hat at x*
SE = sqrt(MSE * (1 + 1/n + (x* minus x-bar)^2 / S_XX))
The "1 +" accounts for individual variability, making this interval wider than the CI for the mean
The data come from an SRS
The relationship between x and y is linear (check the scatterplot and residual plot)
The residuals are normally distributed (check a QQ plot or histogram)
The residuals have constant variance, or equal spread (check the scatterplot and residual plot)
F = \frac{MSA}{MSE}, \quad df_1 = k-1, \quad df_2 = N-k\text{Tukey SE} = \sqrt{MSE \left(\frac{1}{n_i} + \frac{1}{n_j}\right)}\hat{\beta}_1 = b_1 = \frac{S_{XY}}{S_{XX}}, \quad \hat{\beta}_0 = b_0 = \bar{y} - b_1 \bar{x}r = \frac{S_{XY}}{\sqrt{S_{XX} \cdot S_{YY}}}, \quad r^2 = \frac{SSR}{SST}\hat{\sigma} = s = \sqrt{MSE} = \sqrt{\frac{SSE}{n-2}}F_{ts} = \frac{MSR}{MSE}b_1 \pm t_{\alpha/2,\, n-2} \sqrt{\frac{MSE}{S_{XX}}}SE_{\hat{\mu}^*} = \sqrt{MSE \left[\frac{1}{n} + \frac{(x^* - \bar{x})^2}{S_{XX}}\right]}SE_{\hat{y}^*} = \sqrt{MSE \left[1 + \frac{1}{n} + \frac{(x^* - \bar{x})^2}{S_{XX}}\right]}ANOVA is used in agriculture to compare crop yields across different fertiliser treatments, and in medicine to compare the effectiveness of several drugs. Linear regression is everywhere: predicting house prices from square footage, estimating exam scores from study hours, modelling the relationship between advertising spend and sales.
Students often interpret a significant ANOVA result as meaning every group mean differs from every other. It only tells you that at least one pair differs. You need Tukey's method (or a similar post-hoc test) to identify which pairs.
Students confuse correlation with causation. A strong r does not mean X causes Y. There may be lurking or confounding variables.
Students forget that r-squared describes explained variation, not correctness. An r-squared of 0.67 means 67% of the variation in Y is accounted for by the linear relationship with X, not that the model is "67% correct."
Students mix up the confidence interval for the mean response and the prediction interval. The prediction interval is always wider because it accounts for the variability of a single new observation on top of the uncertainty in the mean.
Be able to fill in a one-way ANOVA table given partial information (e.g. given SSA and SST, compute SSE).
The exam practice questions from the crib sheet ask you to compute r, r-squared, the regression equation, and a 95% CI for the slope. Practise the full calculation chain: S_XY, S_XX, b1, b0, r, MSE, SE(b1), and the interval.
Know the four regression assumptions and which plot checks each one (scatterplot for linearity and constant variance, residual plot, QQ plot for normality).
When asked whether there is an association between X and Y, check whether 0 is in the confidence interval for the slope. If 0 is not in the interval, there is a significant linear association.
Correlation does not imply causation. The exam will ask you to state this and identify the lurking variable.
True or False: SST = SSA + SSE in a one-way ANOVA. (Answer: True.)
Fill in the blank: The degrees of freedom for error in simple linear regression are ______. (Answer: n minus 2.)
True or False: A prediction interval for an individual response is narrower than a confidence interval for the mean response. (Answer: False, it is wider.)
Fill in the blank: r^2 = SSR / ______. (Answer: SST.)
True or False: A significant ANOVA F-test tells you which specific group means differ. (Answer: False, you need a post-hoc test like Tukey.)
Q: In a one-way ANOVA with k = 4 groups and N = 24 total observations, what are the degrees of freedom for Factor A and for Error?
A: dfa = k minus 1 = 3. dfe = N minus k = 20.
Q: Given S_XY = 90.14 and S_XX = 3797.04, compute the slope b1 of the least squares regression line.
A: b1 = S_XY / S_XX = 90.14 / 3797.04 = 0.0237.
Q: Given S_XY = 90.14, S_XX = 3797.04, and S_YY = 3.208, compute the sample correlation r and the coefficient of determination r-squared.
A: r = 90.14 / sqrt(3797.04 * 3.208) = 90.14 / sqrt(12,180.70) = 90.14 / 110.37 = 0.82 (approximately). r^2 = (0.82)^2 = 0.67. About 67% of the variation in Y is explained by the linear relationship with X.
Q: You run a regression with n = 60. S_YY = 3.288, b1 = 0.0237, S_XY = 90.14. Compute MSE.
A: SSR = b1 * S_XY = 0.0237 * 90.14 = 2.136. SSE = S_YY minus SSR = 3.288 minus 2.136 = 1.152. MSE = SSE / (n minus 2) = 1.152 / 58 = 0.01986.
Q: After running an ANOVA and finding a significant F-test, a Tukey comparison shows the 95% confidence interval for (mu_1 minus mu_2) is (minus 5.2, 3.1). Are groups 1 and 2 significantly different?
A: No. The interval contains 0, so there is no significant difference between the means of groups 1 and 2 at the 95% confidence level.
ANOVA is an extension of the two-sample t-test (Chapters 9 and 10) to more than two groups. If you run ANOVA with only two groups, the F-statistic equals the square of the t-statistic. Linear regression builds on the concepts of variability (Chapter 3) and hypothesis testing (Chapters 8 to 10). The regression F-test is structurally the same test you learnt in ANOVA, just applied to the regression model. Beyond this course, multiple regression extends the same ideas to several explanatory variables.
STAT 101, Purdue, ANOVA, analysis of variance, one-way ANOVA, F-test, F-statistic, SSA, SSE, SST, MSA, MSE, between-group, within-group, Tukey, post-hoc, pairwise comparison, linear regression, least squares, OLS, slope, intercept, residual, scatterplot, correlation, r, r-squared, coefficient of determination, response variable, explanatory variable, prediction interval, confidence interval for mean response, model utility test, regression assumptions, normality, constant variance, linearity, SRS