Source: Introduction to Statistics, Purdue University Tags: t-test, z-test, paired t-test, two-sample t-test, proportion test, Wilcoxon test, Mann-Whitney U, non-parametric test, test statistic formula, one-sample t-test
Difficulty: Intermediate Prerequisites: Hypothesis testing fundamentals (null and alternative hypotheses, p-values, significance level, Type I and Type II errors). If those concepts are not solid, review the Hypothesis Testing Fundamentals notes first.
Once you understand the logic of hypothesis testing, the next question is: which test do I use? The answer depends on what you are comparing (means vs proportions), how many groups you have (one sample, two samples, or paired observations), and whether your data meet the assumptions for parametric tests. This set of notes covers the most common tests you will encounter in an introductory course: t-tests for means, z-tests for proportions, and Wilcoxon tests as non-parametric fallbacks. Each follows the same reject-or-fail-to-reject framework; the formulas simply change to suit the data structure.
Use a one-sample t-test to compare a sample mean to a known value, a paired t-test for before-and-after designs, a two-sample t-test to compare two independent group means, and z-tests for the same logic applied to proportions. When normality assumptions fail, Wilcoxon tests provide rank-based alternatives that do not require a normal distribution.
One-sample t-test
A test that compares the mean of a single sample to a hypothesised population mean (μ0), using the t-distribution.
In simple terms, you are asking: "Is this sample's average different enough from the target value that the difference is unlikely to be due to chance?"
Paired-sample t-test (dependent t-test)
A test that compares two related measurements (e.g., before and after) by analysing the differences within each pair.
Think of it as a one-sample t-test performed on the column of differences. You are testing whether the average difference is zero.
Two-sample t-test (independent t-test)
A test that compares the means of two independent groups to determine whether they differ.
In simple terms, you are asking: "Do these two groups come from populations with the same mean, or is the gap between their averages too large to be explained by sampling variability alone?"
One-sample z-test for a proportion
A test that compares an observed sample proportion (p-hat) to a hypothesised population proportion (p0), using the standard normal distribution.
Think of it as the proportion equivalent of the one-sample t-test, but because proportions have a known variance structure, you use z instead of t.
Two-sample z-test for proportions
A test that compares the proportions from two independent groups using a pooled estimate of the common proportion.
In simple terms, you are checking whether the difference in success rates between two groups is large enough to be real rather than random noise.
Wilcoxon signed rank test
A non-parametric test that replaces the paired t-test when the normality assumption is not met. It uses the ranks of the absolute differences rather than the raw values.
Think of it as a distribution-free version of the paired t-test. It tests whether a distribution is centred at a particular value.
Wilcoxon rank sum test (Mann-Whitney U test)
A non-parametric test that replaces the two-sample t-test. It pools both samples, ranks all observations, and tests whether one group tends to produce higher or lower ranks.
In simple terms, it asks: "Do the values from one group tend to be larger than those from the other?" without assuming anything about the shape of the distributions.
Pooled proportion (p-hat)
In a two-sample z-test for proportions, the combined proportion from both samples, calculated as the total number of successes divided by the total sample size. It is used in the denominator of the test statistic under the assumption that H0 is true (i.e., the two population proportions are equal).
Degrees of freedom
A parameter of the t-distribution that depends on the sample size. For a one-sample t-test, df = n − 1. For a two-sample t-test with unequal variances, the degrees of freedom are approximated by the Welch-Satterthwaite formula.
In simple terms, degrees of freedom control the shape of the t-distribution. Fewer degrees of freedom produce heavier tails, reflecting greater uncertainty in small samples.
When to use: You have one sample and want to test whether the population mean equals a specific value.
Hypotheses:
H0: μ = μ0
H1: μ ≠ μ0 (two-sided), μ > μ0 (right-tailed), or μ < μ0 (left-tailed)
Assumptions: The data are approximately normally distributed, or the sample size is large enough for the Central Limit Theorem to apply (roughly n ≥ 30).
Decision: Reject H0 if p-value < α. Otherwise, fail to reject.
When to use: You have two related measurements per subject (e.g., pre-treatment and post-treatment scores).
Hypotheses:
H0: μd = 0 (the mean difference is zero)
H1: μd ≠ 0, μd > 0, or μd < 0
Key idea: Compute the difference d for each pair, then run a one-sample t-test on the differences.
Assumptions: The differences are approximately normally distributed.
When to use: You have two independent groups and want to test whether their population means are equal.
Hypotheses:
H0: μ1 = μ2
H1: μ1 ≠ μ2, μ1 > μ2, or μ1 < μ2
Assumptions: Both samples are approximately normal (or large), and the samples are independent. If variances are unequal, use the Welch version (which adjusts the degrees of freedom).
When to use: You have categorical data (success/failure) from one sample and want to test whether the population proportion equals a specific value.
Hypotheses:
H0: p = p0
H1: p ≠ p0, p > p0, or p < p0
Assumptions: np0 ≥ 10 and n(1 − p0) ≥ 10 (the normal approximation to the binomial is reasonable).
Note: The standard error uses p0 (the hypothesised value), not p-hat, because you are computing the distribution of the test statistic under H0.
When to use: You have two independent samples with categorical outcomes and want to compare their population proportions.
Hypotheses:
H0: p1 = p2
H1: p1 ≠ p2, p1 > p2, or p1 < p2
Key idea: Under H0, the two proportions are equal, so you pool the data to estimate a single common proportion (p-hat pooled) and use that in the standard error.
When to use: The normality assumption is clearly violated (small sample, heavy skew, outliers), and you cannot rely on the Central Limit Theorem.
These tests work with ranks instead of raw values, which makes them robust to outliers and non-normal distributions.
They test the location of a distribution (its centre) rather than the mean specifically.
Compute the difference for each pair.
Take the absolute values of the differences and rank them.
Sum the ranks that came from positive differences and the ranks from negative differences separately.
Compare the smaller rank sum to critical values (or compute a p-value) to decide whether the distribution of differences is centred at zero.
Pool all observations from both groups into one list and rank them from smallest to largest.
Sum the ranks for one of the groups.
Compare that rank sum to its expected value under H0 (which assumes both groups have the same distribution).
A rank sum that is much higher or lower than expected suggests one group's values tend to be systematically larger.
t = (x̄ − μ0) / (s / √n)
Where x̄ is the sample mean, μ0 is the hypothesised mean, s is the sample standard deviation, and n is the sample size. Degrees of freedom: df = n − 1.
t = d̄ / (s_d / √n)
Where d̄ is the mean of the paired differences, s_d is the standard deviation of the differences, and n is the number of pairs. Degrees of freedom: df = n − 1.
t = (x̄1 − x̄2) / √(s1²/n1 + s2²/n2)
Where x̄1, x̄2 are the sample means, s1, s2 are the sample standard deviations, and n1, n2 are the sample sizes. Degrees of freedom are approximated by the Welch-Satterthwaite formula.
z = (p̂ − p0) / √(p0(1 − p0) / n)
Where p̂ is the sample proportion, p0 is the hypothesised proportion, and n is the sample size.
z = (p̂1 − p̂2) / √(p̂(1 − p̂)(1/n1 + 1/n2))
Where p̂ is the pooled proportion: p̂ = (x1 + x2) / (n1 + n2), and x1, x2 are the number of successes in each sample.
Two-sided test: p-value = 2 × P(T > |t|) or 2 × P(Z > |z|)
Right-tailed test: p-value = P(T > t) or P(Z > z)
Left-tailed test: p-value = P(T < t) or P(Z < z)
The two-sample t-test is how A/B testing works in technology companies: you split users into two groups, expose them to different versions of a product, and test whether the difference in outcomes (click rates, revenue, time on page) is statistically significant. The z-test for proportions is the standard tool for comparing survey results across demographic groups, such as checking whether voter support for a policy differs between two regions.
"You always use a z-test for means if the sample is large." In practice, most software uses the t-test regardless of sample size. With large n, the t-distribution closely approximates the z-distribution anyway, so it makes no practical difference, but the t-test is the safer default.
"A paired t-test and a two-sample t-test are interchangeable." They are not. Using a two-sample test on paired data ignores the within-pair correlation and typically produces a less powerful test. If the data are naturally paired, use the paired test.
"Non-parametric tests are always worse than parametric tests." When assumptions hold, parametric tests are more powerful. When assumptions are violated, non-parametric tests can be more reliable. The right test is the one whose assumptions fit your data.
"The z-test for proportions uses p-hat in the standard error." Under H0, you use the hypothesised p0 for a one-sample test. For a two-sample test, you use the pooled proportion. Using p-hat in the wrong place is a common exam error.
⚠️ Be able to identify which test to use from a word problem. The key questions are: means or proportions? One sample, two independent samples, or paired samples? Normal data or not?
⚠️ Know every formula on this page. Exams typically give you the scenario and expect you to write down the correct test statistic and compute it.
⚠️ For the proportion z-test, remember that the standard error under H0 uses p0 (one-sample) or the pooled proportion (two-sample), not the individual sample proportions.
⚠️ Understand when to use a Wilcoxon test instead of a t-test. The trigger is a clear violation of the normality assumption with a sample too small for the CLT to rescue you.
⚠️ For paired tests, always compute the differences first and then test. Do not test the two columns separately.
Fill in the blank: A paired t-test is essentially a one-sample t-test performed on the ______. Differences.
True or false: In a two-sample z-test for proportions, you use the individual sample proportions in the standard error formula. False. You use the pooled proportion, because H0 assumes the two population proportions are equal.
True or false: The Wilcoxon rank sum test requires normally distributed data. False. It is a non-parametric test that does not assume normality.
Fill in the blank: For a one-sample t-test with 25 observations, the degrees of freedom equal ______. 24.
True or false: A two-tailed p-value is always double the one-tailed p-value for the same test statistic. True (for symmetric distributions like t and z).
Q: A company claims its light bulbs last 1,000 hours on average. You test 36 bulbs and find a sample mean of 985 hours with a standard deviation of 60 hours. Compute the test statistic for a one-sample t-test.
A: t = (985 − 1000) / (60 / √36) = −15 / 10 = −1.5.
Q: In a before-and-after study with 20 subjects, the mean difference is 3.2 and the standard deviation of the differences is 4.0. Compute the test statistic for a paired t-test.
A: t = 3.2 / (4.0 / √20) = 3.2 / 0.894 ≈ 3.58.
Q: A survey finds that 120 out of 400 respondents in City A support a policy, compared to 90 out of 300 in City B. What is the pooled proportion for a two-sample z-test?
A: p-hat pooled = (120 + 90) / (400 + 300) = 210 / 700 = 0.30.
Q: When would you choose a Wilcoxon signed rank test over a paired t-test?
A: When the differences between pairs are clearly non-normal (e.g., heavily skewed or containing extreme outliers) and the sample size is too small for the Central Limit Theorem to make the t-test reliable.
Q: You are comparing the average test scores of students taught by two different methods. The students were randomly assigned to each method. Which test is appropriate?
A: A two-sample (independent) t-test, because the two groups are independent and you are comparing means.
These tests are all special cases of the general linear model. ANOVA extends the two-sample t-test to three or more groups. Regression analysis uses a t-test on each coefficient to determine whether predictors are significantly related to the outcome.
The z-test for proportions connects to the chi-square test of independence: for a 2×2 contingency table, the two-sample z-test and the chi-square test give identical p-values (since z² = χ² with 1 degree of freedom).
Non-parametric methods extend well beyond the Wilcoxon tests. The Kruskal-Wallis test is the non-parametric analogue of one-way ANOVA, and Spearman's rank correlation is the non-parametric version of Pearson's correlation.
one-sample t-test, paired t-test, dependent t-test, two-sample t-test, independent t-test, Welch's t-test, z-test for proportion, two-proportion z-test, Wilcoxon signed rank, Wilcoxon rank sum, Mann-Whitney U, non-parametric test, rank-based test, test statistic, t-distribution, standard normal, pooled proportion, degrees of freedom, Welch-Satterthwaite, before and after test, A/B testing, STAT 101, intro stats, Purdue