Simple Linear Regression: Model and Calculations, STAT – Ch. 13 – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Chapter 12 (ANOVA concepts), basic algebra

Big Picture

Linear regression models the relationship between two quantitative variables. It sits at the heart of predictive statistics: given an explanatory variable X, can you predict a response variable Y? This chapter covers simple (one-predictor) linear regression, including fitting the line, measuring how well it fits, and testing whether the relationship is statistically significant. The ANOVA table reappears here in a regression context, so familiarity with Chapter 12 helps.

TL;DR

Simple linear regression fits a straight line (y-hat = b0 + b1 * x) to paired data by minimising the sum of squared residuals. R-squared tells you how much of Y's variation the line explains; the Pearson correlation r tells you the strength and direction. Hypothesis tests (F-test or t-test on the slope) determine whether the linear association is statistically significant.


Key Terms

Response variable (Y)

The outcome you are trying to predict or explain. Also called the dependent variable in maths (though in statistics, "dependent" strictly implies causation).

Explanatory variable (X)

The predictor that you believe explains or influences Y. Also called the independent variable. This is the variable treated as fixed.

Population regression model

Y = beta_0 + beta_1 * X + epsilon, where beta_0 is the true intercept, beta_1 is the true slope, and epsilon is the random error term.

Least squares regression line

y-hat = b0 + b1 * x, the fitted line that minimises the sum of squared residuals. b0 and b1 are estimates of the population parameters.

Slope (b1)

The predicted change in Y for a one-unit increase in X. Think of it as the rate: rise over run.

Intercept (b0)

The predicted value of Y when X = 0. Often has no practical meaning if X = 0 is outside the data range.

Residual (e_i)

y_i - y-hat_i, the vertical distance between an observed point and the fitted line. Positive means the point is above the line.

S_XX

The sum of squared deviations of the X values from their mean. Used in computing the slope.

S_YY

The sum of squared deviations of the Y values from their mean. This equals SST.

S_XY

The sum of the products of the deviations of X and Y from their respective means. The sign of S_XY determines the sign of the slope.

Coefficient of determination (R-squared)

SSR / SST. The proportion of Y's total variability explained by the regression line. Ranges from 0 to 1. In simple terms: "What fraction of the ups and downs in Y does the line account for?"

Pearson correlation coefficient (r)

S_XY / sqrt(S_XX * S_YY), or equivalently sign(b1) * sqrt(R-squared). Measures the strength and direction of the linear association. Ranges from -1 to +1.

Common standard deviation (s)

sqrt(MSE). Estimates the standard deviation of the error terms. Think of it as the typical size of a residual.

SSR (Regression Sum of Squares)

The variation in Y explained by the regression line. SSR = b1 * S_XY.

SSE (Error Sum of Squares)

The variation in Y not explained by the line. SSE = SST - SSR.

Extrapolation

Using the regression line to predict Y at an X value outside the range of observed data. Risky, because the linear relationship may not hold outside that range.

Interpolation

Predicting Y at an X value within the range of observed data. This is the safe use of the model.


Core Content

Identifying Variables

  • The explanatory variable (X) is the one you believe does the explaining or influencing.

  • The response variable (Y) is the outcome being measured.

  • Swapping X and Y changes the slope (the two slopes are not inverses of each other) and changes which variable is assumed fixed.

  • Association does not imply causation. Good experimental design is required to establish causality.

Interpreting Scatterplots

When you look at a scatterplot, assess four things:

  • Form: Linear, curved, clusters, or no pattern.

  • Direction: Positive (upward trend), negative (downward), or none (horizontal).

  • Strength: How tightly the points cluster around the pattern. Hard to judge by eye because it depends on axis scaling.

  • Outliers: Points far from the pattern. In regression, the key question is whether an outlier is influential (changes the line) rather than just far from the line. Y-direction outliers usually do not change the regression line much, but X-direction outliers can be highly influential.

The Regression Model

The population model is Y = beta_0 + beta_1 * X + epsilon.

  • beta_0 and beta_1 are fixed (unknown) parameters.

  • epsilon is a random error term.

  • The fitted (estimated) line is y-hat = b0 + b1 * x.

  • The line always passes through the point (x-bar, y-bar).

Four Assumptions for Linear Regression

  1. SRS / Independence: Each pair of observations is from a simple random sample and the pairs are independent of each other.

  1. Linearity: The relationship between X and Y in the population is linear. Check with the scatterplot and the residual plot.

  1. Constant variance (homoscedasticity): The standard deviation of the residuals is the same for all values of X. Check with the residual plot (look for a fan or funnel shape).

  1. Normality: The residuals are normally distributed. Check with a histogram of residuals and a normal probability plot of residuals.

Which Plot Checks Which Assumption

Assumption

Diagnostic plots

SRS

None (assumed by design)

Linearity

Scatterplot, residual plot

Constant standard deviation

Scatterplot, residual plot

Normality of residuals

Histogram of residuals, normal probability plot of residuals

A common exam mistake is confusing these. The normal probability plot checks normality, not linearity. A straight pattern on the normal probability plot means the residuals are normal; it does not mean the data relationship is linear.

Residual Plots

A residual plot graphs the residuals (y - y-hat) against X (or against y-hat). Two advantages over the original scatterplot:

  • Comparing to a horizontal line at zero is easier than comparing to a slanted regression line.

  • The vertical scale is larger, making patterns and outliers more visible.

What to look for: random scatter around zero is good. A curve means the linearity assumption fails. A fan shape means the constant-variance assumption fails.

Least Squares Estimates

The slope is b1 = S_XY / S_XX.

The intercept is b0 = y-bar - b1 * x-bar.

The ANOVA Table for Regression

Source

df

SS

MS

F

Regression

1

SSR = b1 * S_XY

MSR = SSR / 1

MSR / MSE

Error

n - 2

SSE = SST - SSR

MSE = SSE / (n - 2)

Total

n - 1

SST = S_YY

Note that for simple linear regression, df_regression is always 1, and MSR = SSR.

The common standard deviation is s = sqrt(MSE).

Coefficient of Determination (R-squared)

R-squared = SSR / SST.

Interpretation: "R-squared percent of the variation in Y is explained by the linear regression on X."

A high R-squared does not guarantee a good model. Watch for nonlinearity, outliers, unexplained variance from lurking variables, small sample sizes, and extrapolation.

Pearson Correlation (r)

r = S_XY / sqrt(S_XX * S_YY) = sign(b1) * sqrt(R-squared).

Strength guidelines:

  • |r| from 0 to 0.5: weak

  • |r| from 0.5 to 0.8: moderate

  • |r| from 0.8 to 1: strong

The sign of r matches the sign of the slope. r = 0 does not mean X and Y are unassociated; it means there is no linear association. There could still be a strong curved relationship.

Swapping X and Y does not change r.


Formulas

S_{XX} = \sum x_i^2 - \frac{1}{n}\left(\sum x_i\right)^2
S_{YY} = \sum y_i^2 - \frac{1}{n}\left(\sum y_i\right)^2
S_{XY} = \sum x_i y_i - \frac{1}{n}\left(\sum x_i\right)\left(\sum y_i\right)

Slope and intercept:

b_1 = \frac{S_{XY}}{S_{XX}}, \quad b_0 = \bar{y} - b_1 \bar{x}

Sums of squares:

SSR = b_1 \cdot S_{XY}, \quad SST = S_{YY}, \quad SSE = SST - SSR

Mean squares:

MSR = \frac{SSR}{1}, \quad MSE = \frac{SSE}{n-2}

Common standard deviation, R-squared, and correlation:

s = \sqrt{MSE}, \quad R^2 = \frac{SSR}{SST}, \quad r = \frac{S_{XY}}{\sqrt{S_{XX} \cdot S_{YY}}} = \text{sign}(b_1)\sqrt{R^2}

Residual:

e_i = y_i - \hat{y}_i

R code:

fit <- lm(YVar ~ XVar, data = Table) summary(fit)


Common Misconceptions

  • "A high R-squared means the model is good." Not necessarily. The true relationship could be curved (nonlinearity), there could be influential outliers, lurking variables may explain the real pattern, or the sample size may be too small for R-squared to be meaningful.

  • "A small r means X and Y are not associated." A small r means weak linear association. X and Y could have a strong curved, circular, or clustered relationship that r completely misses. Always look at the scatterplot.

  • "The normal probability plot being a straight line means the data are linear." The normal probability plot checks whether the residuals are normally distributed. It says nothing about whether X and Y have a linear relationship. That is checked with the scatterplot and the residual plot.

  • "No outliers on the scatterplot means the residuals are normal." Normality is checked with histograms and normal probability plots of the residuals. A scatterplot without outliers does not prove normality.

  • "Correlation implies causation." It does not. Association is necessary for causation but not sufficient. Only a well-designed experiment can establish causality.

  • "The intercept always has a meaningful interpretation." Most of the time X = 0 is outside the range of the data, so the intercept is just a mathematical feature of the line, not a meaningful prediction.


Why It Matters / Exam Flags

  • You will be asked to compute S_XX, S_YY, and S_XY from given summations and then derive b1, b0, SSR, SSE, SST, R-squared, and r by hand. Know the calculation chain cold.

  • Building the regression ANOVA table by hand is frequently tested. Remember that R does not print SST or df_T, so you must calculate them yourself.

  • Interpreting the slope in context ("for every one-unit increase in X, Y is predicted to change by b1 units") is a standard exam question. Always include units and direction.

  • The intercept interpretation question often has a trick: if X = 0 is not meaningful in context, say so.

  • Know which plots check which assumptions. Confusing the normal probability plot (normality) with linearity is a guaranteed mark deduction.

  • Expect questions on what R-squared does and does not tell you: it measures explained variation but does not guarantee the model is appropriate.

  • The relationship r = sign(b1) * sqrt(R-squared) is tested both as a calculation and as a conceptual question.

  • Beware of extrapolation questions. If x* is outside the range of the data, state that prediction is unreliable.


Quick Self-Test

  1. True or False: Swapping X and Y in a regression gives you the same slope but with the opposite sign. (False. Swapping changes the slope to a different value entirely; the two slopes are not inverses.)

  1. Fill in the blank: b1 = S_XY / ____. (S_XX)

  1. True or False: R-squared = 0.80 means the regression line explains 80% of the variation in Y. (True)

  1. Fill in the blank: The degrees of freedom for error in simple linear regression is ____. (n - 2)

  1. True or False: If r is close to zero, there is no association between X and Y. (False. There may be a nonlinear association that r does not capture.)


Practice Q&A

Q: You are given n = 11, sum of x_i = 580, sum of y_i = -84, sum of x_i squared = 32588, sum of x_i * y_i = -5485. Compute S_XX, S_XY, b1, and b0 (given x-bar = 52.727, y-bar = -7.636).

A: S_XX = 32588 - (580 squared / 11) = 32588 - 30581.82 = 2006.18. S_XY = -5485 - (580 * -84 / 11) = -5485 - (-4429.09) = -1055.91. b1 = -1055.91 / 2006.18 = -0.526. b0 = -7.636 - (-0.526)(52.727) = 20.11. The line is y-hat = 20.11 - 0.526x.

Q: Using the regression line y-hat = 20.11 - 0.526x, predict the response for someone who is 51 years old.

A: y-hat = 20.11 - 0.526(51) = 20.11 - 26.826 = -6.716.

Q: Given SSR = 555.41 and SST = 938.55, compute R-squared and interpret it.

A: R-squared = 555.41 / 938.55 = 0.5918. Approximately 59% of the observed variation in the response is explained by the linear regression on the explanatory variable.

Q: If R-squared = 0.7908 and the slope is negative, what is the Pearson correlation r?

A: r = -sqrt(0.7908) = -0.8893. The negative sign comes from the slope being negative.

Q: A scatterplot shows a strong U-shaped curve. The computed r is 0.02. Does this mean X and Y are not associated?

A: No. r measures linear association only. A strong curved pattern with r near zero means there is a nonlinear association that r misses. Always look at the scatterplot.

Q: What is the residual if the observed y = 58.7 and y-hat = 54.322?

A: e = y - y-hat = 58.7 - 54.322 = 4.378.


Connections to Other Topics

The regression ANOVA table mirrors the one-way ANOVA table from Chapter 12: both partition total variation into explained and unexplained pieces. The F-test for association in regression tests whether the slope is non-zero, which parallels the F-test in ANOVA asking whether group means differ. Residual analysis here connects to the assumptions and diagnostics that underpin all of inferential statistics.

Related Terms / Search Tags

Simple linear regression, least squares, regression line, slope, intercept, residual, scatterplot, residual plot, R-squared, coefficient of determination, Pearson correlation, r, S_XX, S_YY, S_XY, SSR, SSE, SST, MSR, MSE, F-test for regression, ANOVA table for regression, explained variation, unexplained variation, extrapolation, interpolation, lurking variable, influential point, homoscedasticity, constant variance, normality of residuals, lm() in R, summary(), ggplot scatterplot