Two-Sample Tests, ANOVA and Linear Regression, STAT 350 Chapters 10-12 – Study Notes
offline

Difficulty: Advanced | Prerequisites: Chapters 1-9 study notes (especially confidence intervals, hypothesis testing, t-distribution)


Big Picture

This final block extends inference from one sample to two samples (Chapter 10), then to multiple groups (Chapter 11), and finally to the relationship between two continuous variables (Chapter 12). Chapter 10 introduces two-sample independent and matched-pair tests. Chapter 11 (ANOVA) tests whether the means of several groups are equal by comparing variances. Chapter 12 (correlation and regression) models a linear relationship between an explanatory and response variable. These chapters pull together everything from earlier in the course.


TL;DR

Two-sample tests compare means of two populations using either independent samples or matched pairs. ANOVA extends this to more than two groups by comparing between-group variability to within-group variability using the F-distribution. Linear regression fits a line y = b₀ + b₁x to bivariate data, and the correlation coefficient r measures the strength of the linear relationship. All three methods rely on the same inference framework: assumptions, test statistics, p-values and confidence intervals.


Key Terms

Two-sample independent

Two samples drawn from distinct populations with no pairing or matching between observations. The samples have no effect on each other.

Two-sample matched pair (paired data)

Each observation in sample 1 is matched with a specific observation in sample 2 (e.g. before/after measurements on the same subject). Analysis is done on the differences within each pair.

Satterthwaite approximation

A formula for estimating the degrees of freedom in a two-sample t-test when population variances are not assumed equal. It usually gives a non-integer df.

Pooling

Combining two sample variances into a single estimate under the assumption that both populations have the same variance. Not used in this course because the consequences of being wrong (unequal variances) are severe.

Factor

In ANOVA, the variable that differentiates the populations. In simple terms, the thing you are testing the effect of.

Level (group)

The individual categories of the factor. If the factor is "drug dose" with values 0, 20 and 40 mg, then k = 3 levels.

One-way ANOVA

An analysis of variance with a single factor. Tests whether k population means are all equal by comparing between-group variation to within-group variation.

Sum of Squares (SS)

  • SSA (factor/between groups): variation explained by differences between group means

  • SSE (error/within groups): variation within each group, unexplained by the factor

  • SST (total): total variation in the data. SST = SSA + SSE

Mean Square (MS)

Sum of squares divided by its degrees of freedom: MSA = SSA/(k-1), MSE = SSE/(n-k).

F-distribution

A right-skewed, always-positive distribution used for the ANOVA test statistic. It has two degrees of freedom parameters: df₁ = k - 1 (numerator) and df₂ = n - k (denominator).

F test statistic

F_ts = MSA / MSE. When H₀ is true, F_ts ≈ 1. When H₀ is false, F_ts is much greater than 1.

Response variable (Y)

The outcome variable in regression, also called the dependent variable. It goes on the y-axis.

Explanatory variable (X)

The variable that predicts or explains Y, also called the independent variable. It goes on the x-axis and is treated as fixed (no error term).

Simple linear regression

A model of the form Y = β₀ + β₁X + ε, where ε ~ N(0, σ²). It fits a straight line to bivariate data.

Least squares

The method for finding the regression line by minimising the sum of squared vertical distances (residuals) between data points and the line.

Residual

The difference between an observed value and the predicted value: eᵢ = yᵢ - ŷᵢ.

Coefficient of determination (R²)

The proportion of total variability in Y explained by the regression line: R² = SSR/SST. A higher R² means a better fit. For simple linear regression, R² = r².

Sample correlation coefficient (r)

A unitless measure of the strength and direction of the linear relationship between two variables. r ranges from -1 to 1. The sign of r matches the sign of the slope.


Core Content

Chapter 10 – Two-Sample Inference

Two-sample independent (σ known, z-test):

  • Test statistic: z_ts = [(x̄₁ - x̄₂) - Δ₀] / √(σ₁²/n₁ + σ₂²/n₂)

  • CI: (x̄₁ - x̄₂) ± z_(α/2) × √(σ₁²/n₁ + σ₂²/n₂)

  • Normally Δ₀ = 0 (testing whether means are equal)

Two-sample independent (σ unknown, t-test):

  • Test statistic: t'_ts = [(x̄₁ - x̄₂) - Δ₀] / √(s₁²/n₁ + s₂²/n₂)

  • Degrees of freedom: Satterthwaite approximation (usually non-integer)

  • CI: (x̄₁ - x̄₂) ± t_(α/2, ν) × √(s₁²/n₁ + s₂²/n₂)

  • R code: t.test(quantVar ~ catVar, mu=0, conf.level=0.95, paired=FALSE, var.equal=FALSE)

  • We do not assume equal variances in this course (var.equal = FALSE always)

Robustness of the two-sample t-procedure:

  • n < 15: population distribution should be close to normal

  • 15 < n < 40: mild skewness is acceptable

  • n > 40: procedure is usually valid

  • Works best when n₁ ≈ n₂ and the distributions have similar shapes

Two-sample matched pair:

  • Compute differences: Dᵢ = X₁ᵢ - X₂ᵢ

  • This reduces to a one-sample t-test on the differences

  • Test statistic: t_ts = (D-bar - Δ₀) / (s_D / √n), df = n - 1

  • R code: t.test(before, after, paired=TRUE, alternative="two.sided")

How to tell independent from paired:

  • If a lurking variable links individual members of the two samples (same person measured twice, twins, matched partners), use paired

  • If there is no natural pairing, use independent

Chapter 11 – ANOVA

When to use ANOVA:

  • You have more than two groups and want to test whether their means are all equal

  • There is one factor with k levels

  • Each observation falls into exactly one group

ANOVA hypotheses (always the same form):

  • H₀: μ₁ = μ₂ = ... = μ_k

  • Hₐ: at least two μᵢ are different

  • Rejecting H₀ does not tell you which means differ, only that at least two do

ANOVA assumptions:

  • k independent SRSs, one from each population

  • Each population is normally distributed

  • All populations have the same variance σ²

  • Check constant variance: s_max / s_min ≤ 2

ANOVA model:

  • X_ij = μᵢ + ε_ij, where ε_ij ~ N(0, σ²) iid

  • DATA = FIT + RESIDUAL

ANOVA table:

Source

SS

df

MS

F

Factor (between)

SSA

k - 1

MSA = SSA/(k-1)

MSA/MSE

Error (within)

SSE

n - k

MSE = SSE/(n-k)

Total

SST

n - 1

  • SST = SSA + SSE

  • dft = dfa + dfe

  • √MSE is the estimate of σ

Decision rule:

  • p-value = P(F ≥ F_ts) using the F-distribution with df₁ = k-1, df₂ = n-k

  • The p-value is always from the right tail (F is always positive)

  • R code: fit <- aov(Quant ~ Qual, data=TableName); summary(fit)

t-test vs. F-test (when k = 2):

  • When k = 2, t²_ts = F_ts and the p-values are the same

  • The t-test is more flexible (allows one-sided tests, non-zero Δ₀)

  • For k > 2, ANOVA is the correct approach

After rejecting ANOVA H₀:

  • ANOVA tells you "something differs" but not what

  • Post-hoc tests (Tukey, Bonferroni) identify which specific pairs of means differ (these methods may or may not be covered on your exam)

Chapter 12 – Correlation and Linear Regression

Scatterplot interpretation:

  • Form: linear, curved, clusters, no pattern

  • Direction: positive (direct), negative (inverse), horizontal (no association)

  • Strength: how close points fall to the line

  • Outliers: x-outliers can drastically change the line; y-outliers usually have less effect

Assumptions for linear regression:

  • SRS with pairs independent of each other

  • The relationship between X and Y is linear in the population

  • Residuals are normally distributed

  • Residuals have constant variance (homoscedasticity)

  • X is fixed (no error in X); all randomness comes from ε

Least squares regression line:

  • ŷ = b₀ + b₁x

  • Slope: b₁ = S_XY / S_XX

  • Intercept: b₀ = ȳ - b₁x̄

  • S_XY = Σxᵢyᵢ - (Σxᵢ)(Σyᵢ)/n

  • S_XX = Σxᵢ² - (Σxᵢ)²/n

  • The line always passes through the point (x̄, ȳ)

Interpretation of slope and intercept:

  • Slope b₁: for each one-unit increase in X, Y changes by b₁ units on average

  • Intercept b₀: the predicted value of Y when X = 0 (may not have a meaningful interpretation if X = 0 is outside the data range)

Residual analysis:

  • eᵢ = yᵢ - ŷᵢ

  • SSE = Σ(yᵢ - ŷᵢ)², df = n - 2

  • MSE = SSE / (n - 2)

  • √MSE = s, the estimate of σ

Coefficient of determination R²:

  • R² = SSR / SST = 1 - SSE/SST

  • Interpretation: "R² × 100% of the variation in Y is explained by the linear relationship with X"

  • R² is not resistant to outliers

  • A high R² does not guarantee linearity, always check the scatterplot

  • For simple linear regression, r² = R²

Sample correlation coefficient r:

  • r = S_XY / √(S_XX × S_YY)

  • r is unitless and does not change when units change

  • -1 ≤ r ≤ 1

  • |r| < 0.5: weak linear relationship

  • 0.5 ≤ |r| ≤ 0.8: moderate

  • |r| > 0.8: strong

  • r = 0 means X and Y are uncorrelated (no linear relationship, though a nonlinear one may exist)

  • Sign of r matches sign of slope b₁

  • Correlation does not imply causation

Properties of correlation:

  • Switching X and Y does not change r (but does change the regression line)

  • r makes no distinction between explanatory and response variables

  • r measures only linear association; a strong curved relationship can produce r ≈ 0

Regression hypothesis test (is the slope significant?):

  • H₀: β₁ = 0 (no linear relationship)

  • Hₐ: β₁ ≠ 0 (or > 0 or < 0)

  • t_ts = b₁ / SE(b₁), df = n - 2

  • R code: summary(lm(Y ~ X, data=TableName))


Formulas

Quantity

Formula

Two-sample z statistic

z = [(x̄₁ - x̄₂) - Δ₀] / √(σ₁²/n₁ + σ₂²/n₂)

Two-sample t statistic

t = [(x̄₁ - x̄₂) - Δ₀] / √(s₁²/n₁ + s₂²/n₂)

Paired t statistic

t = (D-bar - Δ₀) / (s_D / √n), df = n - 1

F test statistic (ANOVA)

F = MSA / MSE

ANOVA df

dfa = k - 1, dfe = n - k, dft = n - 1

Slope b₁

S_XY / S_XX

Intercept b₀

ȳ - b₁x̄

R²

SSR / SST

Correlation r

S_XY / √(S_XX × S_YY)

MSE (regression)

SSE / (n - 2)


R Commands Reference

Two-sample tests:

  • t.test(quantVar ~ catVar, mu=0, paired=FALSE, var.equal=FALSE, alternative="two.sided")

  • t.test(before, after, paired=TRUE)

ANOVA:

  • fit <- aov(Quant ~ Qual, data=TableName); summary(fit)

  • tapply(Quant, Qual, sd) to check constant variance assumption

  • pf(fts, df1, df2, lower.tail=FALSE) for manual p-value

Regression:

  • lm(Y ~ X, data=TableName) to fit the model

  • summary(lm(Y ~ X)) for coefficients, R², t-tests

  • cor(X, Y) for the correlation coefficient

  • plot(fit) for residual diagnostics


Real-World Applications

Two-sample tests are used in clinical trials to compare a treatment group against a control group. Matched-pair designs are used in before/after studies (e.g. blood pressure before and after medication) to eliminate person-to-person variability. ANOVA is used in agriculture (comparing crop yields across fertiliser types), manufacturing (comparing defect rates across production lines) and pharmaceuticals (comparing dosage effects). Linear regression is used in economics (modelling relationships between variables like income and spending), engineering (calibrating instruments) and medicine (predicting patient outcomes from clinical measurements).


Common Misconceptions

  • Students confuse independent and paired designs. The key question is whether there is a natural pairing between individual observations across the two samples. If so, use paired. If not, use independent.

  • Students sometimes pool variances (assume σ₁ = σ₂) without justification. In this course, do not pool. Always use var.equal = FALSE.

  • Students interpret a significant ANOVA result as meaning all means differ. It means at least two differ. You need post-hoc tests to identify which ones.

  • Students confuse correlation and causation. A high r only means X and Y move together in a linear pattern. It says nothing about whether X causes Y.

  • Students sometimes interpret R² as a measure of linearity. A high R² means the line explains much of the variation, but the data could still be nonlinear. A low R² could mean poor linearity or high scatter.


Why It Matters / Exam Flags

⚠️ Know how to tell independent from paired. If the problem mentions "same subjects measured twice", "before and after", or any natural pairing, it is matched pair.

⚠️ In ANOVA, the alternative hypothesis is always "at least two μᵢ are different." Never write Hₐ: μ₁ ≠ μ₂ ≠ ... ≠ μ_k.

⚠️ Check the constant variance assumption in ANOVA: s_max / s_min ≤ 2.

⚠️ For regression, be able to interpret slope, intercept, R² and r in context.

⚠️ Know that the regression line always passes through (x̄, ȳ).

⚠️ Be able to use R output: read the ANOVA table (F, p-value) and regression summary (coefficients, R², t-tests for significance of β₁).

⚠️ SST = SSA + SSE in ANOVA, and SST = SSR + SSE in regression. The decomposition is the same idea.


Quick Self-Test

True or false: In a two-sample independent t-test, we always assume the population variances are equal. False. In this course we do not assume equal variances (no pooling).

Fill in the blank: In one-way ANOVA with k = 4 groups and n = 40 total observations, dfa = ___ and dfe = ___. dfa = 3, dfe = 36.

True or false: A correlation coefficient of r = -0.9 indicates a weak linear relationship. False. |r| = 0.9 > 0.8 indicates a strong linear relationship. The negative sign indicates direction, not strength.

Fill in the blank: In simple linear regression, the residual degrees of freedom is n - ___. n - 2.

True or false: If the ANOVA F-test is significant, we know which specific group means differ. False. The F-test only tells you that at least two differ. Post-hoc tests are needed to identify which pairs.


Practice Q&A

Q: Two independent samples give x̄₁ = 45, s₁ = 6, n₁ = 30 and x̄₂ = 42, s₂ = 5, n₂ = 35. Test whether the means differ at α = 0.05.

A: H₀: μ₁ - μ₂ = 0, Hₐ: μ₁ - μ₂ ≠ 0. SE = √(36/30 + 25/35) = √(1.2 + 0.714) = √1.914 = 1.384. t_ts = (45 - 42)/1.384 = 2.168. Using Satterthwaite df ≈ 55.6, the two-sided p-value ≈ 0.035. Since 0.035 < 0.05, reject H₀. The data gives support (p = 0.035) to the claim that the two population means differ.

Q: An ANOVA table shows SSA = 120, SSE = 480, k = 4, n = 40. Calculate the F statistic and state the decision at α = 0.05.

A: dfa = 3, dfe = 36. MSA = 120/3 = 40. MSE = 480/36 = 13.33. F_ts = 40/13.33 = 3.00. Using an F-table or R, P(F ≥ 3.00) with df₁ = 3, df₂ = 36 ≈ 0.044. Since 0.044 < 0.05, reject H₀. The data gives support (p = 0.044) to the claim that at least two population means differ.

Q: A regression analysis gives b₁ = -0.526, b₀ = 20.11, SSR = 555.41, SST = 938.55. What is R² and what does it mean?

A: R² = 555.41 / 938.55 = 0.592. About 59.2% of the variation in the response variable is explained by the linear relationship with the explanatory variable. The remaining 40.8% is due to other sources of variability.

Q: If r = -0.85 for a data set, what can you conclude about the linear relationship?

A: There is a strong negative linear association between the two variables. As one increases, the other tends to decrease. Since r² = 0.7225, about 72.25% of the variability in one variable is explained by its linear relationship with the other.


Connections to Other Topics

Two-sample tests are a special case of ANOVA with k = 2 (and when using pooled variance, t² = F). ANOVA's decomposition of total variation into explained and unexplained components is the same idea that appears in regression (SST = SSR + SSE). The regression t-test for β₁ = 0 is equivalent to testing whether the population correlation ρ = 0. These connections illustrate that many seemingly different methods are variations of the same underlying framework: comparing signal (explained variation) to noise (unexplained variation).


Related Terms / Search Tags

two-sample t-test, independent samples, matched pair, paired t-test, Satterthwaite approximation, pooling, ANOVA, analysis of variance, one-way ANOVA, factor, level, F-distribution, F-test, MSA, MSE, SSA, SSE, SST, between-group variation, within-group variation, post-hoc test, Tukey, scatterplot, explanatory variable, response variable, simple linear regression, least squares, slope, intercept, residual, coefficient of determination, R-squared, correlation, Pearson correlation, r, regression line, prediction, STAT 350, Purdue, introductory statistics