Difficulty: Intermediate | Prerequisites: Basic probability, sampling distributions, one-sample t-tests, R fundamentals.
Big picture: The two-sample t-test is how you compare means from two independent groups when you do not know the population standard deviations. It sits at the heart of introductory inference and shows up in nearly every applied statistics course. You need to be comfortable with null and alternative hypotheses, p-values, and the logic of hypothesis testing before this will click. If you can run and interpret a one-sample t-test, this is the natural next step.
A two-sample independent t-test (Welch's version) compares the means of two groups that have no natural pairing. You set up null and alternative hypotheses, check assumptions (SRS, approximate normality), compute a t-statistic, and compare the p-value to your significance level. If the p-value is below alpha, you reject the null; if not, you fail to reject it.
Two-sample independent t-test (Welch's t-test)
A hypothesis test that compares the means of two independent groups without assuming equal population variances. Uses the Welch-Satterthwaite approximation for degrees of freedom.
Think of it as: "Are these two groups actually different on average, or could the gap just be random noise?"
Null hypothesis (H₀)
The default claim that there is no difference between the two population means. Formally: μ₁ - μ₂ = 0.
In simple terms, this means you are assuming the two groups come from populations with the same average until the data convince you otherwise.
Alternative hypothesis (Hₐ)
The claim you are testing for, which contradicts the null. Can be two-sided (μ₁ ≠ μ₂), or one-sided (μ₁ > μ₂ or μ₁ < μ₂).
Think of it as: the direction or difference you suspect is real.
p-value
The probability of observing a test statistic as extreme as (or more extreme than) the one you got, assuming the null hypothesis is true.
In simple terms, a small p-value means your data would be unlikely if the null were true, so you have evidence against it.
Significance level (α)
The threshold you set before testing. If the p-value falls below α, you reject H₀. Common values: 0.05, 0.01, 0.10.
Think of it as: how much risk of a false alarm you are willing to accept.
Degrees of freedom (df)
In a Welch t-test, the df is calculated from the sample sizes and variances of both groups (Welch-Satterthwaite formula). It is not simply n₁ + n₂ - 2.
In simple terms, df controls the shape of the t-distribution you compare your test statistic against.
Confidence interval
A range of plausible values for the true difference in population means, at a given confidence level. A 99% CI means: if you repeated the study many times, about 99% of the intervals constructed this way would contain the true difference.
SRS (Simple Random Sample)
A sampling method where every member of the population has an equal chance of being selected, and selections are independent. This is the baseline assumption that makes the inference valid.
You have two separate groups with no natural pairing or matching between individuals.
You want to compare the population means of some quantitative variable across those two groups.
You do not know the population standard deviations (you estimate them from the data).
There is no before/after or matched-pairs structure. If subjects were matched or measured twice, you would use a paired t-test instead.
Welch's t-test does not assume equal variances across the two groups. It adjusts the degrees of freedom using the Welch-Satterthwaite formula. This is the safer default.
Pooled (equal-variance) t-test assumes σ₁ = σ₂. Only appropriate when you have good reason to believe the population variances are equal. In R, this corresponds to var.equal = TRUE in t.test().
When in doubt, use Welch's. It is more robust and the default in R's t.test() function.
Two-sided (H₀: μ₁ - μ₂ = 0, Hₐ: μ₁ - μ₂ ≠ 0): Use when you have no prior directional expectation. You are simply asking whether the means differ.
One-sided (e.g., Hₐ: μ₁ - μ₂ > 0): Use when you have a specific directional claim to test, and you set this direction before looking at the data.
The direction must be decided before analysis. Choosing a one-sided test after seeing which sample mean is larger is a form of data snooping and invalidates the test.
Research question: Do students who exercise frequently have higher mean Cold_Latency_Swear times than students who do not exercise frequently?
Let μ_freq = population mean Cold_Latency_Swear for students who exercise frequently.
Let μ_notfreq = population mean Cold_Latency_Swear for students who do not exercise frequently.
H₀: μ_freq - μ_notfreq = 0 (no difference)
Hₐ: μ_freq - μ_notfreq > 0 (frequent exercisers have a higher mean)
This is a one-sided test because the research claim has a specific direction.
Significance level: α = 0.01.
Independence / SRS: Both samples are simple random samples drawn independently from their respective populations. Even when this is given (as in many course projects), you must state it explicitly.
Approximate normality: The distribution of the response variable in each population should be approximately normal, or the sample sizes should be large enough for the Central Limit Theorem to compensate. Check this with histograms and QQ plots for each group separately.
Histograms (one per group, with density overlay):
Overlay the kernel density curve (red in the project) and the fitted normal curve (blue) on each histogram.
Look for approximate symmetry, a single peak (unimodal), and reasonable agreement between the kernel density and the normal curve.
Strong skewness, multiple modes, or heavy tails are warning signs.
Boxplots (side by side):
Compare medians, spreads (IQR), and the presence of outliers across groups.
Outliers appear as individual points beyond the whiskers. A few mild outliers in a moderate sample are usually tolerable.
In the project example, both groups showed similar centres and spreads, with no extreme outliers relative to the scale of the data.
QQ (normal probability) plots (one per group):
Points should fall roughly along the reference line if the data are approximately normal.
Systematic curvature suggests skewness; heavy tails produce points that flare away from the line at both ends.
In the project, both groups' QQ plots followed the reference line reasonably well, supporting the normality assumption.
If the data are clearly skewed or have heavy tails, consider a transformation (commonly log or square root) and re-check the diagnostic plots on the transformed data.
If you apply a transformation, you must show the histograms of the original (untransformed) data as justification, then provide all diagnostic graphs for the transformed variable.
If the sample sizes are large (commonly n > 30 per group), moderate departures from normality are less concerning thanks to the CLT.
Assumption | How it was checked | Verdict |
|---|---|---|
SRS | Stated as given in the project description | Satisfied |
Approximate normality | Histograms (symmetric, unimodal), QQ plots (points near the line) | Satisfied, no transformation needed |
t.test(response ~ group, mu = 0, conf.level = 0.99, paired = FALSE, alternative = "greater", var.equal = FALSE)
response ~ group: formula syntax, where the quantitative variable is on the left and the grouping factor is on the right.
mu = 0: the hypothesised difference in means under H₀.
conf.level = 0.99: produces a 99% confidence interval (matches α = 0.01).
paired = FALSE: independent samples, not matched pairs.
alternative = "greater": one-sided test (first group's mean > second group's mean). Use "less" for the other direction, or "two.sided" for a two-sided test.
var.equal = FALSE: Welch's t-test (does not assume equal variances).
Test statistic: t = 1.425
Degrees of freedom: df = 119 (Welch-Satterthwaite approximation)
p-value: 0.07838
Sample means: Frequently = 120.55, NotFrequently = 108.64
99% confidence interval for (μ_freq - μ_notfreq): (-7.80, ∞)
Compare the p-value to α. Here: 0.07838 > 0.01.
Since the p-value exceeds the significance level, we fail to reject H₀.
Conclusion in context: At the 0.01 significance level, there is not sufficient evidence to conclude that students who exercise frequently have a higher mean Cold_Latency_Swear time than students who do not exercise frequently.
The 99% one-sided confidence interval is (-7.80, ∞). This means we are 99% confident the true difference (μ_freq - μ_notfreq) is greater than -7.80. Because this interval includes zero, it is consistent with no real difference between the groups.
Failing to reject H₀ does not prove the null is true. It means the data did not provide strong enough evidence, at the chosen α level, to conclude the alternative is true. The observed difference (about 12 points) was small relative to the variability within each group.
Welch's t-statistic:
t = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}}Where x̄₁ and x̄₂ are the sample means, s₁ and s₂ are the sample standard deviations, and n₁ and n₂ are the sample sizes.
Welch-Satterthwaite degrees of freedom:
df = \frac{\left(\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}\right)^2}{\frac{\left(\frac{s_1^2}{n_1}\right)^2}{n_1 - 1} + \frac{\left(\frac{s_2^2}{n_2}\right)^2}{n_2 - 1}}This is not a whole number in general. R computes it automatically.
Confidence interval for the difference in means:
(\bar{x}_1 - \bar{x}_2) \pm t^* \cdot \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}Where t* is the critical value from the t-distribution at the chosen confidence level with df from the Welch-Satterthwaite formula. For a one-sided interval, the lower (or upper) bound uses the one-tailed critical value, and the other bound is -∞ (or +∞).
Two-sample t-tests appear constantly in clinical trials (does a drug group improve more than a placebo group?), A/B testing in tech (does the new checkout flow lead to higher average spend?), and quality control in manufacturing (does a new process produce parts with a different mean dimension than the old one?). The core logic is always the same: two independent groups, one quantitative outcome, and the question of whether the population means genuinely differ.
Students often think "fail to reject H₀" means "H₀ is true." It does not. It means the evidence was too weak, at the chosen α, to support the alternative. The null may still be false.
Students sometimes choose one-sided vs. two-sided after seeing the data. The direction of the test must be decided before looking at the results. Picking the direction post hoc inflates the actual Type I error rate.
Welch's t-test and the pooled t-test are not interchangeable. Using var.equal = TRUE when the variances differ can produce misleading results. Welch's is the safer default.
A large p-value does not mean "no effect." It means the study could not detect one at that sample size and significance level. There may be a real but small difference that the test lacked the power to find.
⚠️ You will be expected to state all assumptions explicitly, even the ones that are "given." Missing the SRS statement can cost marks.
⚠️ Choosing between paired and independent t-tests is a common exam question. The key question: are the observations in the two groups naturally linked (same subject measured twice, or matched pairs)? If yes, paired. If no, independent.
⚠️ You must justify one-sided vs. two-sided before performing the test. Exams often ask why the choice was made.
⚠️ Interpreting the conclusion in context (not just "reject" or "fail to reject") is typically required for full marks. State what the result means for the original research question.
⚠️ Be ready to interpret a confidence interval alongside the hypothesis test. They should tell a consistent story: if zero is inside the CI, you should not be rejecting H₀.
⚠️ If the diagnostic plots show non-normality, you must apply a transformation before proceeding. Ignoring violated assumptions can result in a large penalty (25 points in this course's project).
True or False: Welch's t-test assumes the two populations have equal variances.
False. That is the pooled t-test. Welch's relaxes this assumption.
True or False: If the p-value is 0.03 and α = 0.01, you reject H₀.
False. 0.03 > 0.01, so you fail to reject.
Fill in the blank: The direction of a one-sided hypothesis test must be chosen ______ analysing the data.
Before.
True or False: A 99% confidence interval that contains zero is consistent with failing to reject H₀ at α = 0.01.
True.
Fill in the blank: If the QQ plot shows systematic curvature away from the reference line, the normality assumption is likely ______.
Violated (and a transformation may be needed).
Q: A researcher wants to compare the average test scores of students taught by Method A versus Method B. The students are different individuals in each group. What test should be used, and why?
A: A two-sample independent t-test (Welch's), because the two groups are separate (no pairing or matching), and the researcher is comparing means of a quantitative variable.
Q: You run a two-sample t-test and get t = 2.31, df = 45, p-value = 0.025. Your significance level is α = 0.05. State your conclusion.
A: Since 0.025 < 0.05, we reject H₀. There is sufficient evidence at the 0.05 level to conclude that the population means differ (if two-sided) or that the mean is greater/less in the hypothesised direction (if one-sided).
Q: Why is Welch's t-test generally preferred over the pooled t-test?
A: Welch's does not assume equal population variances. It adjusts the degrees of freedom to account for unequal variances, making it more robust. When variances happen to be equal, Welch's gives results very close to the pooled test, so there is little cost to using it as the default.
Q: A 95% confidence interval for μ₁ - μ₂ is (-3.2, 8.7). What does this tell you about a two-sided test at α = 0.05?
A: Because the interval contains zero, we would fail to reject H₀ at α = 0.05. The data are consistent with no difference between the population means.
Q: You find the histograms for one group are strongly right-skewed. What should you do before running the t-test?
A: Apply a transformation (e.g., log or square root) to the response variable, then re-check the diagnostic plots (histograms, QQ plots) on the transformed data. If the transformed data look approximately normal, run the t-test on the transformed variable. Show the original histogram as justification for the transformation.
This connects to one-sample t-tests because the mechanics are nearly identical; the two-sample version simply extends the logic to a difference of means rather than a single mean. It also connects to ANOVA: when you have more than two groups, ANOVA generalises the two-sample comparison into an F-test. The assumption-checking workflow here (histograms, QQ plots, stating SRS) carries directly into ANOVA and linear regression diagnostics.
Two-sample t-test, independent samples t-test, Welch's t-test, Welch-Satterthwaite, unpaired t-test, comparing two means, hypothesis testing, null hypothesis, alternative hypothesis, one-sided test, two-sided test, p-value, significance level, alpha, confidence interval, degrees of freedom, t-statistic, fail to reject, normality assumption, QQ plot, normal probability plot, histogram, boxplot, SRS, simple random sample, data transformation, log transformation, R t.test function, STAT 350, introduction to statistics, Purdue, Cold_Latency_Swear, exercise frequency