Difficulty: Intermediate | Prerequisites: Correlation analysis, scatterplots, basic hypothesis testing
Simple linear regression takes the correlation concept one step further: instead of just measuring whether two variables are linearly related, it builds a predictive model. You fit a straight line to the data so you can estimate the value of a response variable (Y) from a predictor variable (X). This is one of the most fundamental tools in statistics and appears across nearly every applied field. You should be comfortable with scatterplots, correlation, and basic hypothesis testing before diving in.
Simple linear regression fits a straight line (Y = b₀ + b₁X) to paired data, minimising the sum of squared residuals. The slope tells you how much Y changes, on average, for each one-unit increase in X. You can test whether the slope is significantly different from zero, check whether the model's assumptions hold, and use the line to make predictions, including interval estimates that quantify your uncertainty.
Least squares regression line
The straight line that minimises the sum of the squared vertical distances (residuals) between each observed data point and the line itself.
Think of it as the single best-fitting straight line through your scatterplot, where "best" means the total squared error is as small as possible.
Slope (b₁)
The change in the predicted value of Y for each one-unit increase in X.
In simple terms, it tells you how steep the line is and in which direction. A slope of 0.5 means that for every additional unit of X, Y goes up by 0.5 on average.
Intercept (b₀)
The predicted value of Y when X equals zero.
Think of it as where the regression line crosses the Y-axis. It may or may not have a meaningful real-world interpretation, depending on whether X = 0 is within the range of your data.
Residual (eᵢ)
The difference between an observed value and its predicted (fitted) value: eᵢ = yᵢ - ŷᵢ.
In simple terms, a residual is how far off the prediction was for a particular data point. Positive means the actual value was above the line; negative means below.
Coefficient of determination (R²)
The proportion of the total variability in Y that is explained by the linear relationship with X.
Think of it as a percentage: R² = 0.76 means 76% of the variation in Y can be accounted for by X through the regression model.
Residual standard error (RSE or sₑ)
An estimate of the typical size of a residual, measured in the same units as Y.
In simple terms, it tells you roughly how far a typical observation falls from the regression line.
Fitted value (ŷᵢ)
The value of Y predicted by the regression line for a given X value.
Think of it as the point on the line directly above or below each observed data point.
Extrapolation
Using the regression model to predict Y for an X value outside the range of the data used to build the model.
In simple terms, you are extending the line beyond where you have evidence. This is risky because the linear pattern may not hold outside the observed range.
Prediction interval
An interval estimate for a single new observation's Y value at a given X. It is wider than a confidence interval because it accounts for both the uncertainty in the line's position and the natural scatter of individual points around the line.
Confidence interval for the mean response
An interval estimate for the average value of Y across all observations at a given X. It is narrower than a prediction interval because it only accounts for uncertainty in the line's position, not individual scatter.
The model form is: ŷ = b₀ + b₁x, where ŷ is the predicted response and x is the predictor.
The slope b₁ and intercept b₀ are chosen to minimise the sum of squared residuals (SSE).
In R, lm(Y ~ X, data = dataset) fits the model. summary() on the result gives you the coefficients, standard errors, t-statistics, p-values, R², and residual standard error.
The slope b₁ is the estimated change in Y for each one-unit increase in X.
Example: if predicting Sepal.Length from Petal.Length (both in cm) and the slope is 0.41, then for each additional centimetre of petal length, sepal length increases by approximately 0.41 cm on average.
Always state the units and the direction.
The intercept b₀ is the predicted value of Y when X = 0.
It may not be meaningful if X = 0 falls outside the range of the data. In the iris example, a petal length of 0 cm is not biologically meaningful, so the intercept serves a mathematical role rather than a practical one.
A residual is eᵢ = yᵢ - ŷᵢ. Positive residuals sit above the line; negative residuals sit below.
The largest positive residual identifies the observation most above the line. The most negative residual identifies the observation most below the line.
In R, resid(model) gives all residuals. which.max(resid(model)) and which.min(resid(model)) find the extreme cases.
You want to test whether the predictor has a statistically significant linear relationship with the response.
t-test for the slope:
H₀: β₁ = 0 (no linear relationship)
H₁: β₁ ≠ 0
Test statistic: t = b₁ / SE(b₁), with n - 2 degrees of freedom.
If the p-value < α, reject H₀ and conclude the slope is significantly different from zero.
F-test for the slope (in simple regression):
Uses the ANOVA table from the regression output.
F = MSR / MSE, where MSR is the regression mean square and MSE is the error mean square.
In simple linear regression (one predictor), the F-test and the t-test are equivalent: F = t², and both give the same p-value.
In R, anova(model) gives the F-statistic.
Normality of residuals:
Checked visually with a Q-Q plot (Normal probability plot) of the residuals.
In R: qqnorm(resid(model)); qqline(resid(model)).
If points roughly follow the diagonal line, the normality assumption is reasonable.
Constant variance (homoscedasticity):
Checked with a residuals-vs-fitted-values plot.
In R: plot(fitted(model), resid(model)).
You want to see a random scatter with no fanning, funnelling, or curved patterns.
If the spread of residuals changes noticeably across the range of fitted values, constant variance is violated.
R² is the proportion of the total variability in Y explained by the regression model.
Interpretation template: "Approximately [R² × 100]% of the variability in [Y] is explained by the linear relationship with [X]."
In simple linear regression, R² = r² (the square of the Pearson correlation coefficient).
Approximately 95% of observations fall within ±2sₑ (two residual standard errors) of their predicted values.
This gives you a practical sense of the model's precision. The fill-in-the-blank version: 95% of observations have a Sepal.Length within ± 2sₑ cm of the regression line's prediction.
Prediction interval: for a single new observation at a given X. Wider, because it includes the scatter of individual points.
Confidence interval for the mean response: for the average Y at a given X. Narrower, because averaging reduces variability.
In R: predict(model, newdata, interval = "prediction", level = 0.80) for an 80% prediction interval; interval = "confidence" for the confidence interval.
The level parameter sets the confidence percentage (e.g. 0.80 for 80%).
Using the model to predict at X values outside the range of the training data is called extrapolation.
The model has no evidence about what happens beyond the observed range, so such predictions are unreliable.
Example: the iris petal lengths in the dataset range from about 1 cm to 6.9 cm. Predicting sepal length for a petal length of 0.2 cm is extrapolation and should be treated with heavy scepticism.
Slope:
b_1 = r \cdot \frac{s_y}{s_x} = \frac{\sum(x_i - \bar{x})(y_i - \bar{y})}{\sum(x_i - \bar{x})^2}Intercept:
b_0 = \bar{y} - b_1 \bar{x}Residual:
e_i = y_i - \hat{y}_i = y_i - (b_0 + b_1 x_i)Residual standard error:
s_e = \sqrt{\frac{\sum e_i^2}{n - 2}} = \sqrt{\frac{SSE}{n - 2}}Coefficient of determination:
R^2 = 1 - \frac{SSE}{SST} = \frac{SSR}{SST}where SST = total sum of squares, SSR = regression sum of squares, SSE = error (residual) sum of squares.
t-statistic for testing β₁ = 0:
t = \frac{b_1}{SE(b_1)}with n - 2 degrees of freedom.
F-statistic (simple regression):
F = \frac{MSR}{MSE} = t^2with 1 and n - 2 degrees of freedom.
Prediction interval for a new observation at x₀:
\hat{y}_0 \pm t_{\alpha/2,\, n-2} \cdot s_e \sqrt{1 + \frac{1}{n} + \frac{(x_0 - \bar{x})^2}{\sum(x_i - \bar{x})^2}}Confidence interval for the mean response at x₀:
\hat{y}_0 \pm t_{\alpha/2,\, n-2} \cdot s_e \sqrt{\frac{1}{n} + \frac{(x_0 - \bar{x})^2}{\sum(x_i - \bar{x})^2}}The only difference is the "1 +" inside the square root for the prediction interval. That extra 1 accounts for the variability of an individual observation around the mean.
Regression models are used everywhere: predicting house prices from square footage, estimating crop yield from rainfall, modelling the relationship between advertising spend and revenue. In biology, the iris example illustrates how physical measurements of one plant feature can predict another, which is useful for species classification.
Students often interpret the intercept literally, even when X = 0 is outside the data range. If no observation has an X value near zero, the intercept is just a mathematical anchor for the line, not a meaningful prediction.
Students confuse prediction intervals with confidence intervals for the mean. The prediction interval is always wider because it covers where a single new point might land, not just where the average sits.
Students sometimes think a significant slope means the model fits well. Significance only tells you the slope is different from zero. R² tells you how well the model fits. You can have a statistically significant but practically weak relationship.
Students assume the regression line will hold outside the observed data range. Extrapolation is risky. The relationship may change entirely beyond the data you collected.
⚠️ You will be asked to interpret the slope in context, with units. "For each additional cm of petal length, sepal length increases by approximately b₁ cm, on average."
⚠️ Know both the t-test and the F-test for testing the slope. In simple regression they give the same p-value, but you should be able to report both test statistics.
⚠️ R² interpretation is a staple exam question. State it as a percentage of variability explained.
⚠️ The 95% rule (within ±2sₑ) is a common fill-in-the-blank question.
⚠️ Know the difference between a prediction interval and a confidence interval for the mean, and when to use each. "Sepal length of an iris with a 4.8 cm petal" calls for a prediction interval. "Average sepal length of all irises with 4.8 cm petals" calls for a confidence interval for the mean.
⚠️ Extrapolation questions ("can you" vs "should you" predict) test whether you understand the limits of the model.
True or False: In simple linear regression, the F-test and the t-test for the slope always give the same conclusion.
Answer: True. In simple regression (one predictor), F = t², and the p-values are identical.
Fill in the blank: Approximately 95% of observations fall within ± ______ residual standard errors of their predicted values.
Answer: 2 (i.e. ±2sₑ).
True or False: A prediction interval is narrower than a confidence interval for the mean response at the same X value.
Answer: False. A prediction interval is always wider.
Fill in the blank: R² = 0.60 means that ______% of the variability in Y is explained by the model.
Answer: 60%.
True or False: If the regression slope is statistically significant at α = 0.05, the model necessarily fits the data well.
Answer: False. Significance tells you the slope is non-zero. R² and residual plots tell you about fit quality.
Q: A regression model predicts Sepal.Length = 4.31 + 0.41 × Petal.Length. Interpret the slope.
A: For each additional centimetre of petal length, the predicted sepal length increases by approximately 0.41 cm, on average.
Q: The residual standard error of a model is 0.47 cm. Roughly what range captures 95% of the observed values around the line?
A: Within ±0.94 cm (2 × 0.47) of the regression line's predictions.
Q: A model has R² = 0.76. What does this tell you?
A: Approximately 76% of the variability in the response variable is explained by the linear relationship with the predictor.
Q: You are asked to predict sepal length for a petal length of 0.2 cm, but the data only include petal lengths from 1.0 cm to 6.9 cm. What is the concern?
A: This is extrapolation. The model has no evidence about petal lengths as small as 0.2 cm, so the prediction is unreliable. The linear relationship may not hold outside the observed range.
Q: What is the difference between a prediction interval and a confidence interval for the mean response?
A: A prediction interval covers where a single new observation might fall. A confidence interval for the mean response covers where the average of all Y values at that X might be. The prediction interval is wider because individual observations are more variable than averages.
Q: In a residuals-vs-fitted-values plot, what pattern would suggest the constant variance assumption is violated?
A: A fan or funnel shape, where the spread of residuals increases (or decreases) as fitted values increase, indicates non-constant variance (heteroscedasticity).
Simple linear regression is the foundation for multiple regression, where you predict Y from several predictors at once. The assumptions you check here (normality, constant variance) carry forward into every regression model you will build. ANOVA, which you may have already encountered, is a special case of regression where the predictor is categorical rather than continuous.
Simple linear regression, least squares, OLS, regression line, slope, intercept, residuals, fitted values, R-squared, coefficient of determination, residual standard error, t-test for slope, F-test for regression, ANOVA table, normality assumption, Q-Q plot, constant variance, homoscedasticity, heteroscedasticity, prediction interval, confidence interval for mean response, extrapolation, iris dataset, STAT 355