Difficulty: Intermediate to Advanced | Prerequisites: Chapter 13a (regression line, ANOVA table), Parts 1 and 2 of these notes (correlation and diagnostics), familiarity with confidence intervals and hypothesis testing from earlier chapters.
Once you have verified the assumptions (Part 2), you can make formal statistical claims about the regression. This section covers four inference procedures: the F-test for overall association, the t-test and confidence interval for the population slope, the confidence interval for the mean response at a given x value, and the prediction interval for a single new observation at a given x value. The F-test and t-test for the slope answer the same question ("is there a linear association?") and give identical p-values in simple linear regression, but they generalise differently to multiple regression. The two intervals at a point look similar but answer different questions, and the prediction interval is always wider.
The F-test checks whether x and y are linearly associated (H0: no association). The t-test on the slope does the same thing, and in simple regression F = t-squared with the same p-value. The confidence interval for the slope tells you the plausible range of the population slope. At a specific x value, you can build a confidence interval for the mean response (narrower) or a prediction interval for the next individual observation (wider, because it adds the point-level variance). All four procedures require the regression assumptions to hold.
F-test (model utility test)
A hypothesis test that compares the variance explained by the regression model (MSR) to the unexplained variance (MSE). It determines whether there is any linear association between x and y.
In simple terms, it asks: "Does the regression line do a better job than just using the mean of y?"
MSR (mean square regression)
SSR divided by its degrees of freedom (always 1 in simple linear regression). It measures how much variance the model explains.
MSE (mean square error)
SSE divided by (n - 2). It estimates the variance of the residuals, sigma-squared. Its square root, s, estimates sigma.
Standard error of b1
The estimated standard deviation of the sampling distribution of the slope. Calculated as sqrt(MSE / S_XX). Used to build confidence intervals and test statistics for the slope.
Think of it as how much the estimated slope would bounce around if you repeated the study many times.
Confidence interval for the slope
An interval estimate for the true population slope beta-1. Centred on b1 with a margin of error based on the t-distribution with n - 2 degrees of freedom.
Confidence interval for the mean at a point (mu)*
An interval for the average value of y across all observations at a specific x = x*. Its width depends on how far x* is from the mean of x.
Prediction interval (Y)*
An interval for the next single observation of y at x = x*. Always wider than the confidence interval for the mean because it includes the individual observation's variability on top of the estimation error.
Think of it as: the CI for the mean tells you where the average sits; the prediction interval tells you where a single new data point might land.
Confidence bands
The curves traced out when you compute the confidence interval for the mean at every x value along the regression line. They form hyperbolic curves that are narrowest at x-bar and widen as you move away.
Interpolation
Predicting within the range of observed x values. Generally safe.
Extrapolation
Predicting outside the range of observed x values. Risky, because you have no data to confirm the relationship continues linearly beyond the observed range.
The F-test asks whether the regression model explains a significant amount of the variation in y.
Step 1: There are no parameters to define (the hypotheses are stated in words).
Step 2: Hypotheses (always in words, always two-sided):
H0: there is no association between x and y
Ha: there is an association between x and y
Step 3: Test statistic and p-value:
F = MSR / MSE, from the ANOVA table
df1 = 1 (always, for simple linear regression), df2 = n - 2
P-value = P(F > F_ts), right-tail only
Step 4: Decision and conclusion in context. If p < alpha, reject H0 and conclude the data provides support for a linear relationship.
A large F means most of the variance in y is explained by the model (large slope relative to noise). A small F means the model does little better than guessing the mean.
The F-test is called the model utility test. A significant F-test does not guarantee a large R-squared; in noisy systems (ecology, for example), R-squared can be below 0.2 and the F-test still significant.
The standard error of b1 is:
SE(b1) = sqrt(MSE / S_XX)
The confidence interval for the population slope beta-1 is:
b1 plus or minus t_(alpha/2, n-2) * SE(b1)
Interpretation: "We are C% confident that the population slope between y and x is covered by the interval (lower, upper)."
By convention, state y first and x second when describing the slope. If 0 is not in the interval, there is a significant linear association. The sign of the entire interval tells you the direction of the relationship.
The hypotheses can be two-sided (beta-1 not equal to beta-1_0) or one-sided (beta-1 < beta-1_0 or beta-1 > beta-1_0). Usually beta-1_0 = 0, testing for association.
Step 1: Define beta-1 = the population slope of y versus x.
Step 2: H0: beta-1 = 0, Ha: beta-1 not equal to 0 (or one-sided).
Step 3: t_ts = b1 / SE(b1) = b1 / sqrt(MSE / S_XX). Degrees of freedom = n - 2.
Step 4: Decision and conclusion.
When beta-1_0 = 0 and the test is two-sided, this t-test and the F-test give exactly the same p-value, and F_ts = t_ts-squared. This is because df1 = 1 for simple regression. In multiple regression, df1 > 1 and the two tests no longer coincide.
This answers: "What is the average value of y for all observations at x = x*?"
The estimated value is y-hat* = b0 + b1 * x*.
The standard error is:
SE(mu-hat*) = sqrt( MSE * [1/n + (x* - x-bar)^2 / S_XX] )
The confidence interval is:
y-hat* plus or minus t_(alpha/2, n-2) * SE(mu-hat*)
The width of this interval depends on how far x* is from x-bar. The interval is narrowest at x-bar and grows wider as you move away. This produces the hyperbolic confidence bands you see plotted around the regression line.
Interpretation: "We are C% confident that the population mean [y variable] is covered by the interval (lower, upper) when [x variable] equals [x*]."
Always state the x* value in the interpretation.
This answers: "What range will the next individual y value fall in when x = x*?"
The standard error adds the individual observation variance (one extra MSE term):
SE(Y-hat*) = sqrt( MSE * [1 + 1/n + (x* - x-bar)^2 / S_XX] )
The prediction interval is:
y-hat* plus or minus t_(alpha/2, n-2) * SE(Y-hat*)
The prediction interval is always wider than the confidence interval for the mean, because it accounts for both the uncertainty in estimating the line and the natural scatter of individual points around it. The extra width comes from that leading "1" inside the square root.
In practice, the prediction interval is used more often. You usually want to know where a single future measurement will land, not the average of many measurements.
Good experimental design is essential. Without it, the data and all inferences are unreliable.
Both correlation and regression describe linear relationships only.
Both are affected by outliers and influential points, which invalidate the assumptions and therefore the inference.
Always plot the data before running inference.
Beware of extrapolation. Predictions outside the observed range of x are unreliable because the linear relationship may not continue.
Beware of lurking variables. A hidden variable may drive both x and y, producing a misleading association.
Association does not imply causation. Statistics can establish association; establishing causation requires additional reasoning and evidence beyond the regression.
Source | df | SS | MS |
|---|---|---|---|
Regression | 1 | SSR = b1 * S_XY | MSR = SSR / 1 |
Error | n - 2 | SSE = SST - SSR | MSE = SSE / (n - 2) |
Total | n - 1 | SST = S_YY | SST / (n - 1) |
F = \frac{MSR}{MSE}, \quad df_1 = 1, \quad df_2 = n - 2SE_{b_1} = \sqrt{\frac{MSE}{S_{XX}}}b_1 \pm t_{\alpha/2,\, n-2} \cdot \sqrt{\frac{MSE}{S_{XX}}}t_{ts} = \frac{b_1 - \beta_{1,0}}{\sqrt{MSE / S_{XX}}}Usually beta-1,0 = 0. Degrees of freedom = n - 2.
SE_{\hat{\mu}^*} = \sqrt{MSE \left[\frac{1}{n} + \frac{(x^* - \bar{x})^2}{S_{XX}}\right]}\hat{y}^* \pm t_{\alpha/2,\, n-2} \cdot SE_{\hat{\mu}^*}SE_{\hat{Y}^*} = \sqrt{MSE \left[1 + \frac{1}{n} + \frac{(x^* - \bar{x})^2}{S_{XX}}\right]}Note the extra "1" compared with the mean-at-a-point formula.
\hat{y}^* \pm t_{\alpha/2,\, n-2} \cdot SE_{\hat{Y}^*}In the biodiesel example from the course, the prediction interval tells a fuel engineer the range of cetane numbers they should expect the next time they measure a biofuel with a given iodine value. The confidence interval for the mean would instead estimate the long-run average cetane number for fuels at that iodine value. The prediction interval is wider because any single measurement also carries its own random scatter.
Students often write the F-test hypotheses using symbols (H0: beta-1 = 0). In this course, the F-test hypotheses are stated in words: "there is no association" and "there is an association." Save the symbolic hypotheses for the t-test on the slope.
Students frequently confuse the CI for the mean at a point with the prediction interval. The CI for the mean estimates where the average sits; the prediction interval estimates where the next single data point will fall. The prediction interval is always wider.
Students sometimes report R-squared as the p-value or vice versa. They measure different things. A significant F-test (small p-value) can occur with a low R-squared in noisy data.
Students occasionally forget to include x* in the interpretation of the CI for the mean or the prediction interval. You must state the specific x value at which the interval was calculated.
Students sometimes assume that a significant test for association proves causation. It does not. Association and causation are separate claims.
Be ready to calculate the F-test from an ANOVA table (F = MSR / MSE) and to state the correct degrees of freedom (1, n - 2).
The relationship F_ts = t_ts-squared is commonly tested. Know why it holds (df1 = 1 in simple regression) and why it breaks in multiple regression.
Expect to be asked to build a confidence interval for the slope from b1, MSE, S_XX, and a t critical value. Know the formula cold.
When computing the CI for the mean or the prediction interval, the difference is one extra "1" under the square root. Exams frequently present both and ask which is wider.
Interpretation phrasing matters. For the slope: "the population slope between y and x." For the mean at a point: "the population mean y when x = [value]." For the prediction interval: "the predicted y for the next observation when x = [value]."
The cautions (extrapolation, lurking variables, causation) appear regularly. Know concrete examples.
True or False: The F-test hypotheses for association are written in symbols (H0: beta-1 = 0). ___
Fill in the blank: In simple linear regression, F_ts = ___-squared.
True or False: The prediction interval is always narrower than the confidence interval for the mean at the same x*. ___
Fill in the blank: The degrees of freedom for all t-based inference in simple linear regression are n - ___.
True or False: The confidence bands around a regression line are straight lines parallel to the fitted line. ___
Answers: 1. False (they are stated in words). 2. t_ts. 3. False (it is always wider). 4. 2. 5. False (they are hyperbolic curves, narrowest at x-bar).
Q: From the cetane example, the ANOVA table gives MSR = 298.254 and MSE = 6.577 with n = 14. Calculate the F-test statistic and state the degrees of freedom.
A: F = 298.254 / 6.577 = 45.35. df1 = 1, df2 = 14 - 2 = 12.
Q: Using b1 = -0.209, MSE = 6.577, and S_XX = 6802.77, calculate the standard error of b1.
A: SE(b1) = sqrt(6.577 / 6802.77) = sqrt(0.000967) = 0.0311.
Q: Build a 95% confidence interval for the population slope given b1 = -0.209, SE = 0.0311, and t_(0.025, 12) = 2.1788.
A: -0.209 plus or minus 2.1788 * 0.0311 = -0.209 plus or minus 0.0678 = (-0.277, -0.141). Since the interval contains only negative values, the slope is significantly negative.
Q: Explain the difference between the confidence interval for the mean at x = 100 and the prediction interval at x = 100.**
A: The CI for the mean estimates the average cetane number across all biofuels with iodine value 100. The prediction interval estimates the range for a single new biofuel measurement at iodine value 100. The prediction interval is wider because it adds the individual observation's variability (an extra MSE term under the square root).
Q: A researcher uses the cetane regression to predict the cetane number at an iodine value of 200, well beyond the observed data range of 59 to 132. Is this appropriate?
A: No. This is extrapolation. There is no data to confirm the linear relationship holds at iodine value 200. The relationship may curve, flatten, or reverse outside the observed range.
The F-test here is structurally the same as the ANOVA F-test from earlier in the course; the linear regression model maps directly onto the one-way ANOVA model with beta-0 + beta-1 * x playing the role of mu-i. The confidence interval and hypothesis test for the slope follow the same estimate-plus-or-minus-critical-value-times-SE pattern used for means throughout the course. In multiple regression (beyond this course), the F-test generalises to test whether any of the predictors matter, while the t-test on each slope tests one predictor at a time.
F-test, model utility test, MSR, MSE, ANOVA table for regression, t-test for slope, standard error of slope, confidence interval for slope, confidence interval for mean at a point, prediction interval, confidence bands, interpolation, extrapolation, lurking variable, association vs causation, simple linear regression inference, degrees of freedom n-2, beta-1, b1, population slope