Difficulty: Intermediate to Advanced | Prerequisites: Probability and Continuous Random Variables notes, Special Distributions notes
This is where the course comes together. Hypothesis testing is the formal framework for deciding whether observed data provide enough evidence to reject a claim about a population. ANOVA extends that framework to comparing three or more groups at once. Two-sample tests handle the common scenario of comparing exactly two groups, and confidence intervals give you a range of plausible values rather than a binary yes/no decision. Experimental design tells you whether the data you collected can actually support a causal conclusion. You need everything from the probability and distributions notes to follow the logic here.
Hypothesis testing follows a four-step procedure: state hypotheses, compute a test statistic, find the p-value, and draw a conclusion. ANOVA tests whether at least one group mean differs from the others. Two-sample t-tests compare two specific groups, either independent or paired. Confidence intervals estimate the true difference between means. Experimental design determines whether you can claim causation or only association.
Null hypothesis (H0)
The default claim, typically that there is no effect or no difference. You assume it is true until the data give you strong enough evidence to reject it.
Alternative hypothesis (Ha or H1)
The claim you are testing for. It is what you conclude if you reject H0. Can be one-sided (< or >) or two-sided (not equal).
Test statistic
A number computed from the sample data that measures how far the observed result is from what H0 predicts. For means, this is usually a t-statistic or z-statistic.
P-value
The probability of observing a test statistic as extreme as (or more extreme than) the one you calculated, assuming H0 is true. A small p-value means the data are unlikely under H0.
Significance level (alpha)
The threshold you set before the test (commonly 0.05). If the p-value is less than alpha, you reject H0.
Type I error
Rejecting H0 when it is actually true. The probability of a Type I error equals alpha. Think of it as a false alarm.
Type II error
Failing to reject H0 when it is actually false. Think of it as a missed detection.
ANOVA (Analysis of Variance)
A method for testing whether the means of three or more groups are all equal. It uses the F-statistic to compare between-group variation to within-group variation.
F-statistic
The ratio of the mean square between groups to the mean square within groups. A large F suggests the group means are not all equal.
Tukey's HSD (Honestly Significant Difference)
A post-hoc test used after ANOVA rejects H0. It identifies which specific pairs of group means are significantly different.
Confidence interval
A range of values that is likely to contain the true population parameter. A 95% confidence interval means that if you repeated the sampling procedure many times, about 95% of the resulting intervals would contain the true value.
Independent two-sample t-test
Compares the means of two unrelated groups. The subjects in one group have no connection to the subjects in the other.
Paired t-test
Compares two related measurements on the same subjects (before/after, or matched pairs). The key feature is that each observation in one group is naturally linked to one in the other.
Observational study
The researcher observes and records data without manipulating any variables. Cannot establish causation.
Experiment
The researcher randomly assigns subjects to treatment groups and controls variables. Can establish causation.
Lurking variable
A variable not included in the study that can affect the response. It is a hidden factor you did not measure.
Confounding variable
A variable that is related to both the treatment and the response, making it impossible to tell which one is actually causing the effect.
The four-step procedure
Step 1: Specify the parameter of interest. State what mu represents (e.g. "the true mean verbal SAT score of students in 2003").
Step 2: Write the hypotheses. H0: mu = value (or mu >= value, or mu <= value). Ha: the alternative claim.
Step 3: Compute the test statistic, degrees of freedom, and p-value. t = (x-bar - mu_0) / (s / sqrt(n)). df = n - 1.
Step 4: Draw a conclusion. If p-value < alpha, reject H0. State the conclusion in context.
Worked example: SAT scores
In 2001 the average verbal SAT was 605. A sample of 20 students from 2003 is taken.
H0: mu >= 605 (or mu = 605). Ha: mu < 605 (one-sided, lower tail).
From the R output: t = -3.778, df = 19, p = 0.001259.
Since 0.001259 < 0.05, reject H0.
Conclusion: the data provide strong evidence that the true mean verbal SAT score in 2003 is less than the 2001 value of 605.
Choosing the correct tail
"Has decreased" or "is less than" indicates a lower-tail test (Ha: mu < value).
"Has increased" or "is greater than" indicates an upper-tail test (Ha: mu > value).
"Is different from" indicates a two-tailed test (Ha: mu not equal to value).
One-way ANOVA setup
Tests H0: mu_1 = mu_2 = ... = mu_k (all group means are equal) against Ha: at least one mu_i is different.
Uses the F-statistic = Mean Square Between / Mean Square Within.
Degrees of freedom: df1 = k - 1 (between groups), df2 = n - k (within groups), where k is the number of groups and n is the total sample size.
Reading the ANOVA table
The table has rows for Factor (between) and Error (within), and columns for Degrees of Freedom, Sum of Squares, Mean Square, F, and P-value.
Mean Square = Sum of Squares / Degrees of Freedom.
If the p-value is small (less than alpha), reject H0 and conclude at least one mean is different.
Example: vitamin C and odontoblasts (guinea pig study)
Three dosage levels (0.5, 1, and 2 mg/day). k = 3, n = 30.
df1 = 2, df2 = 27. F = 27.44. p = 3.13e-7.
Since p < 0.01, reject H0. At least one dosage has a different mean length of odontoblasts.
Tukey's HSD (post-hoc comparisons)
After rejecting H0 in ANOVA, Tukey's test tells you which pairs of means differ significantly.
Each pair is compared. If the confidence interval for the difference does not contain 0, that pair is significantly different.
Example: comparing dosages 1 vs 0.5, 2 vs 0.5, and 2 vs 1. If all intervals exclude 0, all pairs are significantly different.
Critical true/false facts about ANOVA
The null hypothesis is that the population means are equal, not the sample means. (Students mix this up.)
ANOVA tests the means of the populations, not the variances (despite the name "analysis of variance").
A strong case for causation is best made in an experiment, not an observational study.
In rejecting H0, you can conclude at least two means differ, but you cannot say all means are different from each other without a post-hoc test.
One-way ANOVA can be used with two or more groups (three or more is typical, but two is allowed; with two groups a two-sample t-test is preferred).
The F-statistic is large when the between-group variation is large relative to the within-group variation.
When to use the independent two-sample t-test
Use when the two groups are unrelated (different people, different items, no matching).
Example: comparing the running speeds of two groups who wore different brands of compression tights. Different people in each group, no pairing.
When to use the paired t-test
Use when each observation in one group is naturally linked to one in the other.
Common scenarios: the same person measured before and after treatment, or the same person tested under two conditions.
Example: the FitnessCo compression tights study. Each participant ran 2.5 miles with Group A's tights, then again with Group B's tights. The person is the common characteristic, so the data are paired.
Independent two-sample t-test procedure
H0: mu_1 - mu_2 = 0. Ha: mu_1 - mu_2 is not equal to 0 (or > 0 or < 0).
t = (x-bar_1 - x-bar_2) / sqrt(s_1^2/n_1 + s_2^2/n_2).
Degrees of freedom: use the Satterthwaite approximation (typically given or computed by software).
Compare p-value to alpha.
Paired t-test procedure
Compute the differences d_i = observation_1 - observation_2 for each pair.
Treat the differences as a single sample and run a one-sample t-test on d-bar.
H0: mu_d = 0. Ha: mu_d is not equal to 0.
t = d-bar / (s_d / sqrt(n)).
Worked example: FitnessCo compression tights
41 participants, each ran with both brands. Group A mean: 35.9 min (s = 3.8). Group B mean: 36.5 min (s = 4.5). Mean difference (B - A): 3.6 min (s = 3.9).
Satterthwaite df = 77.87.
The two-sample matched pairs test is appropriate because the same person runs under both conditions.
95% CI for the difference: from the output, the interval is (approximately 1.026 to 2.574), confirming that Group B is significantly slower than Group A.
Worked example: diabetes blood sugar study
Two drugs tested on patients. Patients randomly assigned to drug A or drug B for a month, then fasting blood sugar measured. This is an experiment (random assignment exists), so the two-sample independent procedure applies.
H0: mu_A - mu_B = 0. Ha: mu_A - mu_B is not equal to 0.
The pairing would require matching patients, but here there is nothing linking a patient in group A to a specific patient in group B.
General form
Point estimate +/- (critical value) x (standard error).
For a mean: x-bar +/- t* x (s / sqrt(n)).
For a difference of means: (x-bar_1 - x-bar_2) +/- t* x sqrt(s_1^2/n_1 + s_2^2/n_2).
Interpreting a confidence interval
If the interval for mu_1 - mu_2 does not contain 0, the two means are significantly different at that confidence level.
If the interval is entirely positive, mu_1 > mu_2. If entirely negative, mu_1 < mu_2.
Example: a 99% CI of (0.99168, 7.35182) for a drug dosage difference means the true population mean difference is between roughly 1.0 and 7.4 units, and since 0 is not in the interval, the difference is significant.
Upper vs. lower confidence bounds
An upper confidence bound is used when "everything is less than a value" is the concern. You want to know the maximum plausible value.
A lower confidence bound is used when "everything is greater than a value" is the concern.
One-sided bounds use a one-sided critical value (e.g. t with alpha rather than alpha/2).
Practical significance vs. statistical significance
A statistically significant result (p < alpha) may not be practically meaningful if the difference is very small.
Example: if the blood sugar reduction is less than 8 units, it may not be clinically meaningful even if the test gives p < 0.05. The researcher should tell the doctors there is no practical difference.
Observational study vs. experiment
Observational study: no random assignment to groups. The researcher observes what naturally occurs. Cannot prove causation.
Experiment: the researcher randomly assigns subjects to treatment groups. Can prove causation (if properly designed).
Key question to ask: was there random assignment?
If yes, it is an experiment. If no, it is an observational study.
The diabetes drug study: two drugs tested, patients randomly assigned. This is an experiment.
Lurking and confounding variables
Lurking variable: the severity of the disease can affect which drugs work or not. This was not directly controlled in the study.
Confounding variable: the person who was on drugs before the study may respond differently. The drug and the amount of drug would also have to be accounted for.
Physiological factors such as stress can affect blood sugar, and if not controlled for, they confound the results.
Problems specific to the diabetes study
There are two types of diabetes (type I and type II), and drugs that work for one may not work for the other.
Lurking variable: the person taking insulin to control their disease. Short-acting and long-acting insulin have different effects.
A blood sugar reading would not be the appropriate response variable if the person changed their insulin dose. A more relevant response variable would be how much insulin was used.
The behaviorist did not remove the individual's response (physiological factors), although it did not remove any variable or change between periods.
Design improvements
Control for diabetes type.
Record and standardise insulin use.
Account for dietary and exercise differences.
Use a crossover design or matched pairs where possible.
t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} \quad \text{(one-sample t-test)}t = \frac{(\bar{x}_1 - \bar{x}_2) - 0}{\sqrt{s_1^2/n_1 + s_2^2/n_2}} \quad \text{(two-sample independent t-test)}t = \frac{\bar{d}}{s_d / \sqrt{n}} \quad \text{(paired t-test)}F = \frac{\text{MS}_{\text{between}}}{\text{MS}_{\text{within}}} \quad \text{(ANOVA F-statistic)}\text{CI} = \bar{x} \pm t^* \cdot \frac{s}{\sqrt{n}} \quad \text{(confidence interval for a mean)}\text{CI}_{\text{diff}} = (\bar{x}_1 - \bar{x}_2) \pm t^* \cdot \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}df_1 = k - 1, \quad df_2 = n - k \quad \text{(ANOVA degrees of freedom)}Hypothesis tests are used in clinical trials to decide whether a new drug is better than a placebo, in manufacturing to check whether a process meets quality standards, and in education to evaluate whether a new teaching method changes outcomes. ANOVA is the standard tool for comparing multiple treatments in agricultural experiments, A/B/C testing in tech companies, and comparing lab conditions in pharmaceutical research.
Students often confuse the null hypothesis with the alternative. The null is the status quo; the alternative is what you hope to show.
A p-value of 0.03 does not mean "there is a 3% chance H0 is true." It means there is a 3% chance of seeing data this extreme if H0 were true.
ANOVA tests the means, not the variances. The name "analysis of variance" refers to the method (partitioning variance), not the parameter being tested.
Students mix up paired and independent tests. If the same person appears in both groups, it is paired. If different people are in each group, it is independent.
⚠️ Always show all four steps of the hypothesis test. Skipping a step loses marks even if your final answer is correct.
⚠️ When reading R output for ANOVA, identify which output corresponds to which question. The exam often presents multiple outputs on one page.
⚠️ The paired vs. independent decision is a frequent exam question. State the common characteristic that makes the data paired.
⚠️ Confidence interval interpretation: if the interval for the difference includes 0, the result is not significant. If it does not include 0, it is significant.
⚠️ For the experimental design section, always identify the lurking and confounding variables by name and explain how they affect the study.
⚠️ Practical significance questions: even if p < 0.05, if the effect size is too small to matter in practice, say so. The exam tests whether you understand this distinction.
True or False: The null hypothesis in ANOVA is that the sample means are all equal. (False: it is the population means)
True or False: If a 95% confidence interval for mu_1 - mu_2 is (1.2, 4.8), the difference is statistically significant. (True: 0 is not in the interval)
Fill in the blank: In a paired t-test, you first compute the ______ for each pair, then run a one-sample t-test on those values. (differences)
True or False: An observational study can establish causation. (False: only an experiment with random assignment can)
True or False: A small p-value means the effect is practically important. (False: statistical significance and practical significance are different things)
Q: A sample of 20 students has a mean verbal SAT score of 542.62. The 2001 population average was 605. The t-statistic is -3.778 with df = 19 and p = 0.001259. At alpha = 0.05, what do you conclude?
A: Since p = 0.001259 < 0.05, reject H0. There is strong evidence that the true mean verbal SAT score in 2003 is less than 605.
Q: An ANOVA on three vitamin C dosage groups gives F = 27.44 and p = 3.13e-7. What is your conclusion?
A: Reject H0. At least one dosage level has a different mean odontoblast length than the others.
Q: Runners test two brands of compression tights by each running 2.5 miles with brand A and then 2.5 miles with brand B. Is this a paired or independent design? Why?
A: Paired, because each participant runs under both conditions. The person is the common characteristic linking the two measurements.
Q: A diabetes study randomly assigns patients to Drug A or Drug B. Is this an observational study or an experiment?
A: An experiment, because the researcher randomly assigned patients to treatment groups.
Q: A 95% CI for the mean difference in running speed between two groups is (1.026, 2.574). What does this tell you?
A: The true population mean difference is between approximately 1.0 and 2.6 minutes. Since the interval does not contain 0, the difference is statistically significant at the 5% level. Group B is significantly slower.
Q: The blood sugar difference between two drugs is statistically significant (p < 0.05) but the effect is less than 8 units. What should the researcher tell the doctors?
A: The difference is statistically significant but not practically significant. A difference of less than 8 units is too small to matter clinically, so there is no practical difference between the drugs.
Hypothesis testing builds directly on the sampling distributions and CLT from the Probability notes. The t-distribution is a modification of the normal that accounts for estimating sigma from the sample. ANOVA generalises the two-sample t-test to k groups. Confidence intervals are the dual of hypothesis tests: if a value falls outside the CI, the corresponding test would reject at that confidence level. Experimental design principles connect to research methods courses across the sciences.
Hypothesis testing, null hypothesis, alternative hypothesis, p-value, significance level, alpha, Type I error, Type II error, t-test, one-sample t-test, two-sample t-test, paired t-test, independent samples, ANOVA, one-way ANOVA, F-statistic, F-test, degrees of freedom, mean square, sum of squares, Tukey HSD, post-hoc test, confidence interval, margin of error, critical value, upper bound, lower bound, practical significance, statistical significance, observational study, experiment, random assignment, lurking variable, confounding variable, Satterthwaite approximation, SAT scores, compression tights, diabetes, blood sugar, STAT 101, Purdue, introduction to statistics