Difficulty: Advanced | Prerequisites: Chapters 1-9 study notes (especially confidence intervals, hypothesis testing, t-distribution)
This final block extends inference from one sample to two samples (Chapter 10), then to multiple groups (Chapter 11), and finally to the relationship between two continuous variables (Chapter 12). Chapter 10 introduces two-sample independent and matched-pair tests. Chapter 11 (ANOVA) tests whether the means of several groups are equal by comparing variances. Chapter 12 (correlation and regression) models a linear relationship between an explanatory and response variable. These chapters pull together everything from earlier in the course.
Two-sample tests compare means of two populations using either independent samples or matched pairs. ANOVA extends this to more than two groups by comparing between-group variability to within-group variability using the F-distribution. Linear regression fits a line y = b₀ + b₁x to bivariate data, and the correlation coefficient r measures the strength of the linear relationship. All three methods rely on the same inference framework: assumptions, test statistics, p-values and confidence intervals.
Two-sample independent
Two samples drawn from distinct populations with no pairing or matching between observations. The samples have no effect on each other.
Two-sample matched pair (paired data)
Each observation in sample 1 is matched with a specific observation in sample 2 (e.g. before/after measurements on the same subject). Analysis is done on the differences within each pair.
Satterthwaite approximation
A formula for estimating the degrees of freedom in a two-sample t-test when population variances are not assumed equal. It usually gives a non-integer df.
Pooling
Combining two sample variances into a single estimate under the assumption that both populations have the same variance. Not used in this course because the consequences of being wrong (unequal variances) are severe.
Factor
In ANOVA, the variable that differentiates the populations. In simple terms, the thing you are testing the effect of.
Level (group)
The individual categories of the factor. If the factor is "drug dose" with values 0, 20 and 40 mg, then k = 3 levels.
One-way ANOVA
An analysis of variance with a single factor. Tests whether k population means are all equal by comparing between-group variation to within-group variation.
Sum of Squares (SS)
SSA (factor/between groups): variation explained by differences between group means
SSE (error/within groups): variation within each group, unexplained by the factor
SST (total): total variation in the data. SST = SSA + SSE
Mean Square (MS)
Sum of squares divided by its degrees of freedom: MSA = SSA/(k-1), MSE = SSE/(n-k).
F-distribution
A right-skewed, always-positive distribution used for the ANOVA test statistic. It has two degrees of freedom parameters: df₁ = k - 1 (numerator) and df₂ = n - k (denominator).
F test statistic
F_ts = MSA / MSE. When H₀ is true, F_ts ≈ 1. When H₀ is false, F_ts is much greater than 1.
Response variable (Y)
The outcome variable in regression, also called the dependent variable. It goes on the y-axis.
Explanatory variable (X)
The variable that predicts or explains Y, also called the independent variable. It goes on the x-axis and is treated as fixed (no error term).
Simple linear regression
A model of the form Y = β₀ + β₁X + ε, where ε ~ N(0, σ²). It fits a straight line to bivariate data.
Least squares
The method for finding the regression line by minimising the sum of squared vertical distances (residuals) between data points and the line.
Residual
The difference between an observed value and the predicted value: eᵢ = yᵢ - ŷᵢ.
Coefficient of determination (R²)
The proportion of total variability in Y explained by the regression line: R² = SSR/SST. A higher R² means a better fit. For simple linear regression, R² = r².
Sample correlation coefficient (r)
A unitless measure of the strength and direction of the linear relationship between two variables. r ranges from -1 to 1. The sign of r matches the sign of the slope.
Two-sample independent (σ known, z-test):
Test statistic: z_ts = [(x̄₁ - x̄₂) - Δ₀] / √(σ₁²/n₁ + σ₂²/n₂)
CI: (x̄₁ - x̄₂) ± z_(α/2) × √(σ₁²/n₁ + σ₂²/n₂)
Normally Δ₀ = 0 (testing whether means are equal)
Two-sample independent (σ unknown, t-test):
Test statistic: t'_ts = [(x̄₁ - x̄₂) - Δ₀] / √(s₁²/n₁ + s₂²/n₂)
Degrees of freedom: Satterthwaite approximation (usually non-integer)
CI: (x̄₁ - x̄₂) ± t_(α/2, ν) × √(s₁²/n₁ + s₂²/n₂)
R code: t.test(quantVar ~ catVar, mu=0, conf.level=0.95, paired=FALSE, var.equal=FALSE)
We do not assume equal variances in this course (var.equal = FALSE always)
Robustness of the two-sample t-procedure:
n < 15: population distribution should be close to normal
15 < n < 40: mild skewness is acceptable
n > 40: procedure is usually valid
Works best when n₁ ≈ n₂ and the distributions have similar shapes
Two-sample matched pair:
Compute differences: Dᵢ = X₁ᵢ - X₂ᵢ
This reduces to a one-sample t-test on the differences
Test statistic: t_ts = (D-bar - Δ₀) / (s_D / √n), df = n - 1
R code: t.test(before, after, paired=TRUE, alternative="two.sided")
How to tell independent from paired:
If a lurking variable links individual members of the two samples (same person measured twice, twins, matched partners), use paired
If there is no natural pairing, use independent
When to use ANOVA:
You have more than two groups and want to test whether their means are all equal
There is one factor with k levels
Each observation falls into exactly one group
ANOVA hypotheses (always the same form):
H₀: μ₁ = μ₂ = ... = μ_k
Hₐ: at least two μᵢ are different
Rejecting H₀ does not tell you which means differ, only that at least two do
ANOVA assumptions:
k independent SRSs, one from each population
Each population is normally distributed
All populations have the same variance σ²
Check constant variance: s_max / s_min ≤ 2
ANOVA model:
X_ij = μᵢ + ε_ij, where ε_ij ~ N(0, σ²) iid
DATA = FIT + RESIDUAL
ANOVA table:
Source | SS | df | MS | F |
|---|---|---|---|---|
Factor (between) | SSA | k - 1 | MSA = SSA/(k-1) | MSA/MSE |
Error (within) | SSE | n - k | MSE = SSE/(n-k) | |
Total | SST | n - 1 |
SST = SSA + SSE
dft = dfa + dfe
√MSE is the estimate of σ
Decision rule:
p-value = P(F ≥ F_ts) using the F-distribution with df₁ = k-1, df₂ = n-k
The p-value is always from the right tail (F is always positive)
R code: fit <- aov(Quant ~ Qual, data=TableName); summary(fit)
t-test vs. F-test (when k = 2):
When k = 2, t²_ts = F_ts and the p-values are the same
The t-test is more flexible (allows one-sided tests, non-zero Δ₀)
For k > 2, ANOVA is the correct approach
After rejecting ANOVA H₀:
ANOVA tells you "something differs" but not what
Post-hoc tests (Tukey, Bonferroni) identify which specific pairs of means differ (these methods may or may not be covered on your exam)
Scatterplot interpretation:
Form: linear, curved, clusters, no pattern
Direction: positive (direct), negative (inverse), horizontal (no association)
Strength: how close points fall to the line
Outliers: x-outliers can drastically change the line; y-outliers usually have less effect
Assumptions for linear regression:
SRS with pairs independent of each other
The relationship between X and Y is linear in the population
Residuals are normally distributed
Residuals have constant variance (homoscedasticity)
X is fixed (no error in X); all randomness comes from ε
Least squares regression line:
ŷ = b₀ + b₁x
Slope: b₁ = S_XY / S_XX
Intercept: b₀ = ȳ - b₁x̄
S_XY = Σxᵢyᵢ - (Σxᵢ)(Σyᵢ)/n
S_XX = Σxᵢ² - (Σxᵢ)²/n
The line always passes through the point (x̄, ȳ)
Interpretation of slope and intercept:
Slope b₁: for each one-unit increase in X, Y changes by b₁ units on average
Intercept b₀: the predicted value of Y when X = 0 (may not have a meaningful interpretation if X = 0 is outside the data range)
Residual analysis:
eᵢ = yᵢ - ŷᵢ
SSE = Σ(yᵢ - ŷᵢ)², df = n - 2
MSE = SSE / (n - 2)
√MSE = s, the estimate of σ
Coefficient of determination R²:
R² = SSR / SST = 1 - SSE/SST
Interpretation: "R² × 100% of the variation in Y is explained by the linear relationship with X"
R² is not resistant to outliers
A high R² does not guarantee linearity, always check the scatterplot
For simple linear regression, r² = R²
Sample correlation coefficient r:
r = S_XY / √(S_XX × S_YY)
r is unitless and does not change when units change
-1 ≤ r ≤ 1
|r| < 0.5: weak linear relationship
0.5 ≤ |r| ≤ 0.8: moderate
|r| > 0.8: strong
r = 0 means X and Y are uncorrelated (no linear relationship, though a nonlinear one may exist)
Sign of r matches sign of slope b₁
Correlation does not imply causation
Properties of correlation:
Switching X and Y does not change r (but does change the regression line)
r makes no distinction between explanatory and response variables
r measures only linear association; a strong curved relationship can produce r ≈ 0
Regression hypothesis test (is the slope significant?):
H₀: β₁ = 0 (no linear relationship)
Hₐ: β₁ ≠ 0 (or > 0 or < 0)
t_ts = b₁ / SE(b₁), df = n - 2
R code: summary(lm(Y ~ X, data=TableName))
Quantity | Formula |
|---|---|
Two-sample z statistic | z = [(x̄₁ - x̄₂) - Δ₀] / √(σ₁²/n₁ + σ₂²/n₂) |
Two-sample t statistic | t = [(x̄₁ - x̄₂) - Δ₀] / √(s₁²/n₁ + s₂²/n₂) |
Paired t statistic | t = (D-bar - Δ₀) / (s_D / √n), df = n - 1 |
F test statistic (ANOVA) | F = MSA / MSE |
ANOVA df | dfa = k - 1, dfe = n - k, dft = n - 1 |
Slope b₁ | S_XY / S_XX |
Intercept b₀ | ȳ - b₁x̄ |
R² | SSR / SST |
Correlation r | S_XY / √(S_XX × S_YY) |
MSE (regression) | SSE / (n - 2) |
Two-sample tests:
t.test(quantVar ~ catVar, mu=0, paired=FALSE, var.equal=FALSE, alternative="two.sided")
t.test(before, after, paired=TRUE)
ANOVA:
fit <- aov(Quant ~ Qual, data=TableName); summary(fit)
tapply(Quant, Qual, sd) to check constant variance assumption
pf(fts, df1, df2, lower.tail=FALSE) for manual p-value
Regression:
lm(Y ~ X, data=TableName) to fit the model
summary(lm(Y ~ X)) for coefficients, R², t-tests
cor(X, Y) for the correlation coefficient
plot(fit) for residual diagnostics
Two-sample tests are used in clinical trials to compare a treatment group against a control group. Matched-pair designs are used in before/after studies (e.g. blood pressure before and after medication) to eliminate person-to-person variability. ANOVA is used in agriculture (comparing crop yields across fertiliser types), manufacturing (comparing defect rates across production lines) and pharmaceuticals (comparing dosage effects). Linear regression is used in economics (modelling relationships between variables like income and spending), engineering (calibrating instruments) and medicine (predicting patient outcomes from clinical measurements).
Students confuse independent and paired designs. The key question is whether there is a natural pairing between individual observations across the two samples. If so, use paired. If not, use independent.
Students sometimes pool variances (assume σ₁ = σ₂) without justification. In this course, do not pool. Always use var.equal = FALSE.
Students interpret a significant ANOVA result as meaning all means differ. It means at least two differ. You need post-hoc tests to identify which ones.
Students confuse correlation and causation. A high r only means X and Y move together in a linear pattern. It says nothing about whether X causes Y.
Students sometimes interpret R² as a measure of linearity. A high R² means the line explains much of the variation, but the data could still be nonlinear. A low R² could mean poor linearity or high scatter.
⚠️ Know how to tell independent from paired. If the problem mentions "same subjects measured twice", "before and after", or any natural pairing, it is matched pair.
⚠️ In ANOVA, the alternative hypothesis is always "at least two μᵢ are different." Never write Hₐ: μ₁ ≠ μ₂ ≠ ... ≠ μ_k.
⚠️ Check the constant variance assumption in ANOVA: s_max / s_min ≤ 2.
⚠️ For regression, be able to interpret slope, intercept, R² and r in context.
⚠️ Know that the regression line always passes through (x̄, ȳ).
⚠️ Be able to use R output: read the ANOVA table (F, p-value) and regression summary (coefficients, R², t-tests for significance of β₁).
⚠️ SST = SSA + SSE in ANOVA, and SST = SSR + SSE in regression. The decomposition is the same idea.
True or false: In a two-sample independent t-test, we always assume the population variances are equal. False. In this course we do not assume equal variances (no pooling).
Fill in the blank: In one-way ANOVA with k = 4 groups and n = 40 total observations, dfa = ___ and dfe = ___. dfa = 3, dfe = 36.
True or false: A correlation coefficient of r = -0.9 indicates a weak linear relationship. False. |r| = 0.9 > 0.8 indicates a strong linear relationship. The negative sign indicates direction, not strength.
Fill in the blank: In simple linear regression, the residual degrees of freedom is n - ___. n - 2.
True or false: If the ANOVA F-test is significant, we know which specific group means differ. False. The F-test only tells you that at least two differ. Post-hoc tests are needed to identify which pairs.
Q: Two independent samples give x̄₁ = 45, s₁ = 6, n₁ = 30 and x̄₂ = 42, s₂ = 5, n₂ = 35. Test whether the means differ at α = 0.05.
A: H₀: μ₁ - μ₂ = 0, Hₐ: μ₁ - μ₂ ≠ 0. SE = √(36/30 + 25/35) = √(1.2 + 0.714) = √1.914 = 1.384. t_ts = (45 - 42)/1.384 = 2.168. Using Satterthwaite df ≈ 55.6, the two-sided p-value ≈ 0.035. Since 0.035 < 0.05, reject H₀. The data gives support (p = 0.035) to the claim that the two population means differ.
Q: An ANOVA table shows SSA = 120, SSE = 480, k = 4, n = 40. Calculate the F statistic and state the decision at α = 0.05.
A: dfa = 3, dfe = 36. MSA = 120/3 = 40. MSE = 480/36 = 13.33. F_ts = 40/13.33 = 3.00. Using an F-table or R, P(F ≥ 3.00) with df₁ = 3, df₂ = 36 ≈ 0.044. Since 0.044 < 0.05, reject H₀. The data gives support (p = 0.044) to the claim that at least two population means differ.
Q: A regression analysis gives b₁ = -0.526, b₀ = 20.11, SSR = 555.41, SST = 938.55. What is R² and what does it mean?
A: R² = 555.41 / 938.55 = 0.592. About 59.2% of the variation in the response variable is explained by the linear relationship with the explanatory variable. The remaining 40.8% is due to other sources of variability.
Q: If r = -0.85 for a data set, what can you conclude about the linear relationship?
A: There is a strong negative linear association between the two variables. As one increases, the other tends to decrease. Since r² = 0.7225, about 72.25% of the variability in one variable is explained by its linear relationship with the other.
Two-sample tests are a special case of ANOVA with k = 2 (and when using pooled variance, t² = F). ANOVA's decomposition of total variation into explained and unexplained components is the same idea that appears in regression (SST = SSR + SSE). The regression t-test for β₁ = 0 is equivalent to testing whether the population correlation ρ = 0. These connections illustrate that many seemingly different methods are variations of the same underlying framework: comparing signal (explained variation) to noise (unexplained variation).
two-sample t-test, independent samples, matched pair, paired t-test, Satterthwaite approximation, pooling, ANOVA, analysis of variance, one-way ANOVA, factor, level, F-distribution, F-test, MSA, MSE, SSA, SSE, SST, between-group variation, within-group variation, post-hoc test, Tukey, scatterplot, explanatory variable, response variable, simple linear regression, least squares, slope, intercept, residual, coefficient of determination, R-squared, correlation, Pearson correlation, r, regression line, prediction, STAT 350, Purdue, introductory statistics