Linear Regression Basics, STAT 301 Ch. 12 Part 1 – Study Notes
offline

Source: Objectives for Final Exam, Introduction to Statistics (Purdue University)

Difficulty: Intermediate | Prerequisites: ANOVA concepts (Chapter 11), hypothesis testing fundamentals, basic algebra.

Big Picture

Linear regression models the relationship between a quantitative explanatory variable (X) and a quantitative response variable (Y) using a straight line. Where ANOVA asks "do these group means differ?", regression asks "as X changes, does Y change in a predictable, linear way?" This chapter covers building the model, checking whether it works, measuring how well it fits, and using it to make predictions. The ANOVA table reappears here, but now it partitions variation explained by the regression line versus unexplained error.


TL;DR

Linear regression fits a straight line (Y-hat = b0 + b1 * X) to data using the least squares method. R-squared tells you how much of the variation in Y the line explains. The correlation coefficient r measures the strength and direction of the linear relationship. Before trusting any of this, you check four assumptions: SRS, linearity, normality of residuals, and constant variance of residuals.


Key Terms

Explanatory variable (X)

The independent variable, the predictor. This is the variable you believe influences or predicts the response.

Response variable (Y)

The dependent variable, the outcome you are measuring or trying to predict.

Scatterplot

A graph with X on the horizontal axis and Y on the vertical axis. Each point is one observation. You interpret it for pattern (linear or curved), direction (positive or negative), strength (tight or loose), outliers, and constant variance.

Linear regression model

Y = beta_0 + beta_1 * X + epsilon, where beta_0 is the population y-intercept, beta_1 is the population slope, and epsilon is the random error term.

Least squares regression line

Y-hat = b0 + b1 * X. The line that minimises the sum of squared residuals. b0 and b1 are point estimates of beta_0 and beta_1.

Residual

The difference between an observed value and its predicted value: e_i = y_i - y-hat_i. Think of it as the "leftover" the model did not explain.

S_XX, S_XY, S_YY

Summation quantities used to calculate the slope, intercept, correlation, and ANOVA table entries. They are corrected sums of squares and cross-products.

R-squared (coefficient of determination)

The proportion of the total variation in Y that is explained by the linear relationship with X. Calculated as SSR / SST. In simple terms, it tells you what percentage of Y's behaviour the line captures.

r (sample correlation coefficient)

Measures the strength and direction of the linear association between X and Y. Ranges from -1 to +1. Calculated as S_XY / sqrt(S_XX * S_YY), or as the square root of R-squared with the sign of b1.

Standard deviation about the least squares line (s)

The point estimate of sigma, calculated as sqrt(MSE). It measures the typical size of a residual, in the units of Y.

SSR (Sum of Squares for Regression)

The variation in Y explained by the regression line. SSR = b1 * S_XY.

SSE (Sum of Squares for Error)

The variation in Y not explained by the regression line. SSE = SST - SSR.

SST (Total Sum of Squares)

The total variation in Y. SST = S_YY.

Extrapolation

Using the regression line to predict Y at an X value outside the range of the observed data. This is unreliable because you have no evidence the linear pattern continues beyond your data.


Core Content

Explanatory vs Response Variables

  • The explanatory variable (X) is the one you think explains or predicts the outcome. The response variable (Y) is the outcome being predicted.

  • On a scatterplot, X goes on the horizontal axis and Y on the vertical axis.

Interpreting a Scatterplot

Describe five features:

  • Pattern: Is the relationship linear, curved, or no pattern at all?

  • Direction: Positive (Y increases as X increases) or negative (Y decreases as X increases)?

  • Strength: How tightly do the points follow the pattern? Strong, moderate, or weak?

  • Outliers: Are there any points far from the overall pattern?

  • Constant variance: Does the vertical spread of points stay roughly the same across the range of X, or does it fan out?

The Linear Regression Model

The population model is Y = beta_0 + beta_1 * X + epsilon.

  • beta_0 is the population y-intercept (the mean value of Y when X = 0).

  • beta_1 is the population slope (the change in the mean of Y for a one-unit increase in X).

  • epsilon is the random error, assumed to be normally distributed with mean 0 and constant standard deviation sigma.

Four Assumptions for Linear Regression

  1. Simple random sample (SRS): The data are collected via a random process.

  1. Linearity: The relationship between X and Y is linear. Check with a scatterplot of X vs Y and a residual plot (residuals vs X or residuals vs fitted values).

  1. Normality of residuals: The residuals follow a normal distribution. Check with a normal probability plot of residuals or a histogram of residuals.

  1. Constant variance (homoscedasticity): The spread of residuals is the same for all values of X. Check with a residual plot.

For the exam, you must name all graphs that can check each assumption, even if you only use one.

Calculating S_XX, S_XY, S_YY

These are computed from raw summations:

  • S_XX = sum(x squared) - (sum x) squared / n

  • S_XY = sum(xy) - (sum x)(sum y) / n

  • S_YY = sum(y squared) - (sum y) squared / n

The Least Squares Regression Line

  • Slope: b1 = S_XY / S_XX

  • Intercept: b0 = y-bar - b1 * x-bar

  • The equation: Y-hat = b0 + b1 * X

  • From computer output, read the coefficients from the "Estimate" column: the intercept row gives b0, and the X row gives b1.

Point Prediction

To predict Y for a given X value, substitute that value into the regression equation: Y-hat = b0 + b1 * X. This works only for X values within the range of the observed data (no extrapolation).

ANOVA Table for Linear Regression

The regression ANOVA table partitions variation in Y:

  • Regression row: df = 1, SSR = b1 * S_XY, MSR = SSR / 1

  • Error row: df = n - 2, SSE = SST - SSR, MSE = SSE / (n - 2)

  • Total row: df = n - 1, SST = S_YY

Key relationships:

  • SST = SSR + SSE

  • dft = dfr + dfe

  • If given the ANOVA table from R, you may need to calculate dft and SST by hand (add the df and SS values from the regression and error rows).

Standard Deviation About the Least Squares Line

s = sqrt(MSE). This estimates sigma, the typical distance of an observation from the regression line in the units of Y.

Coefficient of Determination (R-squared)

R-squared = SSR / SST. It tells you the proportion of total variation in Y explained by the linear model.

What R-squared does not tell you:

  • It does not confirm linearity (a curved pattern can still produce a moderate R-squared).

  • It does not detect outliers.

  • A high R-squared does not guarantee good prediction for individual observations.

Sample Correlation (r)

Two ways to calculate:

  • r = S_XY / sqrt(S_XX * S_YY)

  • r = sqrt(R-squared) with the sign of b1

Properties of r:

  • Ranges from -1 to +1.

  • The sign matches the direction of the linear relationship.

  • r = 0 means no linear association ("uncorrelated"), but there could still be a non-linear pattern.

  • Switching X and Y does not change r.

  • Correlation does not tell you about the form of the relationship (it only measures linear association).

Interpreting the Y-Intercept

The y-intercept (b0) is the predicted value of Y when X = 0. Whether it makes sense to interpret depends on whether X = 0 is within the range of the data and whether X = 0 has a meaningful real-world interpretation. If X = 0 is outside the data range or is not physically meaningful, the intercept should not be interpreted on its own.


Formulas

S_{XX} = \sum x^2 - \frac{(\sum x)^2}{n}
S_{XY} = \sum xy - \frac{(\sum x)(\sum y)}{n}
S_{YY} = \sum y^2 - \frac{(\sum y)^2}{n}
b_1 = \frac{S_{XY}}{S_{XX}}, \quad b_0 = \bar{y} - b_1 \bar{x}
\hat{Y} = b_0 + b_1 X
SSR = b_1 \cdot S_{XY}, \quad SST = S_{YY}, \quad SSE = SST - SSR
s = \sqrt{MSE} = \sqrt{\frac{SSE}{n-2}}
R^2 = \frac{SSR}{SST}
r = \frac{S_{XY}}{\sqrt{S_{XX} \cdot S_{YY}}} = \pm\sqrt{R^2} \text{ (sign of } b_1\text{)}

Real-World Applications

Linear regression is used to predict house prices from square footage, estimate fuel consumption from vehicle weight, and model how study hours relate to exam scores. Any time you have two quantitative variables and want to quantify their linear relationship or make predictions, regression is the standard approach.


Common Misconceptions

  • Students often confuse R-squared with correlation (r). R-squared is the proportion of variance explained (always between 0 and 1). Correlation is the strength and direction of linear association (between -1 and +1). They are related: R-squared = r squared.

  • A high R-squared does not mean the model is correct. The relationship could be curved, and the line could still explain a decent chunk of variance.

  • Students sometimes interpret the y-intercept even when X = 0 is meaningless or outside the data range. Always check whether X = 0 makes sense in context before interpreting b0.

  • Correlation does not imply causation. Even a strong linear association does not prove that X causes Y. Confounding variables, reverse causation, or coincidence could be at play.


Why It Matters / Exam Flags

  • You will need to calculate b0 and b1 by hand from summation quantities. Practise the S_XX, S_XY, S_YY formulas.

  • You will need to fill in a regression ANOVA table by hand. If given R output, you must add the regression and error rows yourself to get dft and SST.

  • Know the difference between R-squared and r. The exam may ask you to calculate one from the other.

  • Be prepared to identify which graphs check which assumptions. You must list all applicable graphs for each assumption.

  • Extrapolation questions are common. If X = x* is outside the observed X range, state that prediction is unreliable and explain why.


Quick Self-Test

  1. True or false: R-squared = 0.81 means r = 0.81. (False. r = +/- 0.9, with the sign matching the slope.)

  1. Fill in the blank: The regression ANOVA table has ______ degree(s) of freedom for regression. (1)

  1. True or false: If X = 0 is outside the data range, you should still interpret the y-intercept. (False.)

  1. Fill in the blank: SST = SSR + ______. (SSE)

  1. True or false: Switching X and Y changes the value of r. (False. Correlation is symmetric.)


Practice Q&A

Q: Given sum(x) = 100, sum(y) = 200, sum(x squared) = 2500, sum(xy) = 5200, n = 10, calculate S_XX and S_XY.

A: S_XX = 2500 - (100 squared / 10) = 2500 - 1000 = 1500. S_XY = 5200 - (100 * 200 / 10) = 5200 - 2000 = 3200.

Q: From the previous question, what is b1?

A: b1 = S_XY / S_XX = 3200 / 1500 = 2.133.

Q: An R output shows SSR = 450, SSE = 150, dfr = 1, dfe = 18. Calculate SST, dft, R-squared, and s.

A: SST = 450 + 150 = 600. dft = 1 + 18 = 19. R-squared = 450 / 600 = 0.75. MSE = 150 / 18 = 8.33. s = sqrt(8.33) = 2.887.

Q: A regression of exam score (Y) on hours studied (X) gives Y-hat = 40 + 5X. Predict the score for a student who studies 8 hours. The data range for X was 2 to 12 hours.

A: Y-hat = 40 + 5(8) = 80. This is a valid prediction because X = 8 is within the observed range (2 to 12).

Q: In the previous model, should you interpret the y-intercept of 40?

A: X = 0 (zero hours of study) is outside the observed data range of 2 to 12 hours, so interpreting the intercept as a predicted exam score for zero study hours is not appropriate.


Connections to Other Topics

The regression ANOVA table uses the same logic as the one-way ANOVA from Chapter 11: partition total variability into explained and unexplained portions and compare them with an F ratio. The F-test for the regression model (covered in Part 2) is directly analogous to the one-way ANOVA F-test. The residuals and assumptions here also connect to the broader theme of checking model validity before drawing conclusions.


Related Terms / Search Tags

linear regression, simple linear regression, least squares, least squares line, regression line, slope, intercept, y-intercept, b0, b1, beta, residual, residual plot, scatterplot, explanatory variable, response variable, predictor, independent variable, dependent variable, S_XX, S_XY, S_YY, sum of squares, SSR, SSE, SST, ANOVA table for regression, R-squared, coefficient of determination, correlation, r, sample correlation, standard deviation about the line, MSE, extrapolation, homoscedasticity, constant variance, normality of residuals, STAT 301, Purdue, intro stats