Regression Inference and Prediction, STAT – Ch. 13 – Study Notes
offline

Difficulty: Intermediate-Advanced | Prerequisites: Linear regression model and calculations (Part 1 of these notes)

Big Picture

Once you have a fitted regression line, the natural next questions are: "Is this relationship real or just noise?" and "How precisely can I predict new values?" This set of notes covers the inferential side of regression, including hypothesis tests for association (F-test and t-test on the slope), confidence intervals for the slope, confidence intervals for the mean response at a given X, and prediction intervals for individual new observations. The distinction between the last two is one of the most commonly tested points on the exam.

TL;DR

Two tests check whether X and Y are linearly related: an F-test (using the regression ANOVA table) and a t-test on the slope. Both give the same p-value in simple regression. Confidence intervals for the slope quantify the precision of b1. At a specific x*, a confidence interval estimates the population mean response, while a prediction interval estimates a single new observation. The prediction interval is always wider because it includes the extra uncertainty of an individual point.


Key Terms

F-test for association (model utility test)

Tests H0: there is no linear association between X and Y vs Ha: there is a linear association. Uses F_ts = MSR / MSE from the regression ANOVA table. Think of it as asking: "Does the regression line explain a significant amount of Y's variation?"

t-test for the slope

Tests H0: beta_1 = beta_1_0 (often 0) vs Ha: beta_1 is not equal to (or greater/less than) beta_1_0. In simple regression with beta_1_0 = 0 and a two-sided alternative, this gives the same p-value as the F-test, and t_ts squared = F_ts.

Standard error of the slope (SE of b1)

sqrt(MSE / S_XX). Measures how precisely the slope is estimated. Smaller S_XX or larger MSE means less precision.

Confidence interval for the slope

b1 +/- t_(alpha/2, n-2) * sqrt(MSE / S_XX). Gives a range of plausible values for the true population slope beta_1.

Confidence interval for the mean response at x*

Estimates the average Y across all observations at a specific X = x*. Uses SE = sqrt(MSE * [1/n + (x* - x-bar) squared / S_XX]). The interval is narrowest at x-bar and widens as x* moves away from the centre.

Prediction interval for a new observation at x*

Estimates the Y value of a single new observation at X = x*. Uses SE = sqrt(MSE * [1 + 1/n + (x* - x-bar) squared / S_XX]). Always wider than the confidence interval for the mean because it includes the extra variability of an individual point (sigma squared).

Extrapolation

Predicting Y at an X outside the observed data range. Both confidence and prediction intervals become unreliable here because the linear relationship may not extend.


Core Content

F-Test for Association (Model Utility Test)

This tests whether the regression model as a whole is useful.

Step 1: Can be skipped (no parameters to define beyond the model).

Step 2: H0: there is no linear association between X and Y. Ha: there is a linear association between X and Y.

Step 3: F_ts = MSR / MSE. The numerator df = 1 (regression), denominator df = n - 2 (error). P-value from pf(Fts, 1, n-2, lower.tail = FALSE).

Step 4: Compare p-value to alpha. If p < alpha, reject H0. Conclude that the data does (or does not) provide support for a linear association.

R code: fit <- lm(YVar ~ XVar, data = Table) then summary(fit). The F-statistic and its p-value appear at the bottom of the output.

t-Test for the Slope

This tests a specific claim about the population slope beta_1.

The null value is typically beta_1_0 = 0 (no linear relationship), but it can be any value.

Hypotheses come in three forms:

Lower-tail

Two-tailed

Upper-tail

H0

beta_1 >= beta_1_0

beta_1 = beta_1_0

beta_1 <= beta_1_0

Ha

beta_1 < beta_1_0

beta_1 != beta_1_0

beta_1 > beta_1_0

Test statistic: t_ts = b1 / sqrt(MSE / S_XX) = b1 / SE(b1). Degrees of freedom = n - 2.

When beta_1_0 = 0 and the test is two-sided, t_ts squared = F_ts, and the p-values are identical.

R code: tts <- b1 / SE then 2 * pt(abs(tts), n-2, lower.tail = FALSE) for two-tailed.

Confidence Interval for the Slope

b1 +/- t_(alpha/2, n-2) * sqrt(MSE / S_XX).

Interpretation: "We are C% confident that the true population slope is covered by the interval (lower, upper)."

The critical value comes from qt(alpha/2, n-2, lower.tail = FALSE) in R.

Confidence Interval for the Mean Response at x*

This answers: "What is the average Y for all members of the population at X = x*?"

The point estimate is y-hat = b0 + b1 * x*.

The standard error is SE_mu = sqrt(MSE * [1/n + (x* - x-bar) squared / S_XX]).

The interval is y-hat +/- t_(alpha/2, n-2) * SE_mu.

Interpretation: "We are C% confident that the population mean response is covered by this interval when X = x*."

R code: newdata <- data.frame(XVar = x_star) then predict(fit, newdata, interval = "confidence", level = C).

Prediction Interval for a New Observation at x*

This answers: "What will the next individual Y value be at X = x*?"

The point estimate is the same: y-hat = b0 + b1 * x*.

The standard error is SE_Y = sqrt(MSE * [1 + 1/n + (x* - x-bar) squared / S_XX]).

The interval is y-hat +/- t_(alpha/2, n-2) * SE_Y.

Interpretation: "We are C% confident that the next observation's Y value is covered by this interval when X = x*."

R code: same as above, but interval = "prediction".

Why the Prediction Interval Is Always Wider

The confidence interval for the mean estimates an average, which has only the sampling uncertainty of the line. The prediction interval estimates an individual point, which has the sampling uncertainty of the line plus the random scatter of individuals around the line (sigma squared). The extra "1" inside the square root of SE_Y is what makes the prediction interval wider.


Formulas

Standard error of the slope:

SE_{b_1} = \sqrt{\frac{MSE}{S_{XX}}}

t-test statistic for the slope:

t_{ts} = \frac{b_1 - \beta_{1_0}}{\sqrt{\frac{MSE}{S_{XX}}}}

Confidence interval for the slope:

b_1 \pm t_{\alpha/2,\, n-2} \cdot \sqrt{\frac{MSE}{S_{XX}}}

Standard error for the mean response at x*:

SE_{\hat{\mu}*} = \sqrt{MSE \left[\frac{1}{n} + \frac{(x^* - \bar{x})^2}{S_{XX}}\right]}

Confidence interval for the mean response:

\hat{y}^* \pm t_{\alpha/2,\, n-2} \cdot SE_{\hat{\mu}*}

Standard error for the prediction at x*:

SE_{\hat{Y}*} = \sqrt{MSE \left[1 + \frac{1}{n} + \frac{(x^* - \bar{x})^2}{S_{XX}}\right]}

Prediction interval:

\hat{y}^* \pm t_{\alpha/2,\, n-2} \cdot SE_{\hat{Y}*}

R code for intervals (with raw data):

predict(fit, newdata, interval = "confidence", level = 0.95) predict(fit, newdata, interval = "prediction", level = 0.95)

R code for intervals (without raw data):

SE_mu <- sqrt(MSE * (1/n + (xstar - xbar)^2 / Sxx)) SE_Y <- sqrt(MSE * (1 + 1/n + (xstar - xbar)^2 / Sxx))


Common Misconceptions

  • "The confidence interval for the mean and the prediction interval are the same thing." They are not. The confidence interval targets the population average at x*; the prediction interval targets a single new observation. The prediction interval is always wider.

  • "A wider prediction interval means a worse model." The prediction interval is wider by construction (it includes individual variability). A wider interval relative to the confidence interval is expected and does not reflect model quality.

  • "The F-test and the t-test for the slope can give different conclusions." In simple linear regression with a two-sided test and beta_1_0 = 0, they always give the same p-value. They can differ only in multiple regression or with one-sided slope tests.

  • "I can use the regression line to predict outside the data range and the interval will protect me." The interval formulas assume the linear relationship holds at x*. If x* is outside the range of observed X values, the relationship may not be linear there, and both intervals become unreliable.


Why It Matters / Exam Flags

  • Expect to be asked whether the y-intercept has a meaningful interpretation in a given scenario. If X = 0 is outside the data range or makes no practical sense, say so.

  • You must know the four-step procedure for both the F-test and the t-test on the slope, including writing the conclusion in context.

  • Computing confidence and prediction intervals by hand (plugging into the SE formulas, looking up the t-critical value, building the interval) is a core exam skill.

  • The most commonly tested conceptual question: "Why is the prediction interval wider than the confidence interval?" Know the answer cold.

  • Interpreting R output from predict() is tested. Know which word to change ("confidence" vs "prediction") and what the three columns (fit, lwr, upr) represent.

  • Extrapolation is a favourite trick question. If x* is outside the data range, state that the prediction is unreliable regardless of what the interval says.

  • The relationship t_ts squared = F_ts appears in questions that ask you to verify one test using the other.


Quick Self-Test

  1. True or False: The prediction interval for a new Y at x* is always wider than the confidence interval for the mean Y at the same x*. (True)

  1. Fill in the blank: The degrees of freedom for the t-test on the slope in simple regression is ____. (n - 2)

  1. True or False: In simple linear regression, the F-test p-value and the two-sided t-test p-value (when testing beta_1 = 0) are different. (False, they are the same.)

  1. Fill in the blank: The extra term inside the square root that makes the prediction SE larger than the mean-response SE is ____. (1, representing the individual observation variance)

  1. True or False: If zero is inside the 95% confidence interval for the slope, you would reject H0: beta_1 = 0 at alpha = 0.05. (False, you would fail to reject.)


Practice Q&A

Q: Given b1 = -0.209, MSE = 6.577, S_XX = 6802.77, and n = 14, compute the 95% confidence interval for the slope. The critical value from R is t = 2.1788.

A: SE = sqrt(6.577 / 6802.77) = sqrt(0.000967) = 0.03109. The interval is -0.209 +/- 2.1788 * 0.03109 = -0.209 +/- 0.0677 = (-0.277, -0.141). We are 95% confident the population slope is between -0.277 and -0.141.

Q: Using the same data, is there a useful linear relationship at alpha = 0.05? Conduct the t-test.

A: H0: beta_1 = 0, Ha: beta_1 != 0. t_ts = -0.209 / 0.03109 = -6.722. df = 12. P-value = 2 * P(T > 6.722) = 2.13 * 10^-5. Since 2.13 * 10^-5 < 0.05, reject H0. The data provides strong support for a linear relationship.

Q: For the cetane/iodine example, the 95% confidence interval for the mean cetane number at iodine = 100 is (52.936, 55.690) and the 95% prediction interval is (48.512, 60.114). Explain why the prediction interval is wider.

A: The confidence interval estimates the population average cetane number at iodine = 100 and only captures the uncertainty in estimating the line. The prediction interval estimates a single new observation's cetane number and must also account for the natural scatter of individual points around the line (sigma squared). The extra variability makes it wider.

Q: A 95% CI for the slope is (0.0903, 0.2045). Does this support the claim that there is a positive linear association?

A: Yes. Zero is not in the interval, and both endpoints are positive. This is consistent with rejecting H0: beta_1 = 0 and concluding a positive association.

Q: You want to predict Y at x = 250, but the observed X values range from 50 to 200. Should you use the regression line?*

A: No. x* = 250 is outside the observed range, so this would be extrapolation. The linear relationship may not hold there, and any interval would be unreliable.


Connections to Other Topics

The F-test here is the regression version of the same F-test from Chapter 12 (ANOVA). In one-way ANOVA, F compares group means; in regression, F asks whether the slope is non-zero. The t-test for the slope parallels the one-sample and two-sample t-tests from earlier chapters but applied to the regression coefficient. Confidence and prediction intervals extend the interval-estimation logic from one-sample and two-sample settings to the regression context.

Related Terms / Search Tags

F-test for association, model utility test, t-test for slope, beta_1, hypothesis test for slope, confidence interval for slope, standard error of slope, mean response, confidence interval at a point, prediction interval, SE of prediction, SE of mean response, extrapolation, interpolation, t-critical value, qt(), predict() in R, lm(), summary(), pf(), pt(), regression inference, simple linear regression inference