Difficulty: Intermediate | Prerequisites: One-sample t-test, basic hypothesis testing, matched pairs t-test (companion notes).
The two-sample (independent samples) t-test compares the means of two separate populations using data from two independent groups. Unlike matched pairs, there is no natural link between individual observations in group 1 and group 2. This lab uses gender as the grouping variable and weekday sleep time as the measured variable, asking whether male and female college students differ in average weekday sleep. The Welch (unequal variances) version of the test is the default here.
The two-sample independent t-test checks whether two separate groups have different population means. You collect one measurement per subject, group them by a categorical variable (e.g. gender), and test whether the difference in group means is statistically significant. When you cannot assume the two groups have equal variances, use the Welch (unequal variances) version, which adjusts the degrees of freedom.
Independent samples (two-sample design)
Two groups drawn from two separate populations with no pairing or matching between individuals. Each subject belongs to exactly one group. In the lab, males form one group and females form the other.
In simple terms, think of it as "different people in each group, measured once each."
Two-sample t-test (independent samples t-test)
A hypothesis test that compares the means of two independent groups. The test statistic measures how far apart the two sample means are, relative to the variability within each group.
Grouping variable
The categorical variable that defines the two groups. In the lab, Gender is the grouping variable (male vs. female). In SPSS, you specify it as the "Grouping Variable" and define the two group codes.
Levene's test for equality of variances
A preliminary test that checks whether the two groups have equal population variances. SPSS runs it automatically. If Levene's test is significant (p < 0.05), the equal-variances assumption is questionable, and you should read the "Equal variances not assumed" row.
Welch's t-test (equal variances not assumed)
The version of the two-sample t-test that does not require equal variances. It uses a modified formula for degrees of freedom (often a non-integer). This lab instructs you to use this version throughout.
Pooled t-test (equal variances assumed)
The version that assumes both populations have the same variance and pools the two sample variances into one estimate. Use this only when Levene's test is non-significant and you have reason to believe the variances are equal. The lab says not to use it.
Mean difference (x-bar_1 - x-bar_2)
The difference between the two sample means. In the lab: 7.0582 - 6.9486 = 0.1096 hours (male mean minus female mean).
Standard error of the difference
Estimates how much the difference in sample means would fluctuate from sample to sample. For unequal variances, it is computed from each group's variance and sample size separately.
Degrees of freedom (Welch)
A corrected df that accounts for unequal variances and unequal sample sizes. It is usually not an integer. In the lab, df = 243.282.
Two-tailed (two-sided) test
A hypothesis test where the alternative hypothesis says the means are simply different (not specifying which is larger). H0: mu_1 = mu_2 vs. Ha: mu_1 is not equal to mu_2. The lab's gender comparison uses a two-tailed test.
You have two separate populations (e.g. males and females, treatment and control).
Each subject is measured once and belongs to exactly one group.
There is no valid reason to pair an individual in one group with an individual in the other.
You are comparing two means, one per group.
Ask: "Are these two columns of data from the same subjects measured twice, or from entirely different groups of subjects?" If different groups with no pairing, use the two-sample test.
In the lab, Gender splits the 250 students into 116 males and 134 females. Each student appears in only one group, and there is no natural pairing between a specific male and a specific female.
Define which group is "1" and which is "2." In the lab (male - female order):
H0: mu_male = mu_female (equivalently, mu_male - mu_female = 0)
Ha: mu_male is not equal to mu_female
This is a two-tailed test because the research question asks whether sleep differs by gender, without specifying a direction.
SPSS produces two tables for an independent-samples t-test:
Group Statistics table: shows each group's sample size (N), mean, standard deviation, and standard error. From the lab:
Males: N = 116, mean = 7.0582, SD = 0.49576
Females: N = 134, mean = 6.9486, SD = 0.49860
Independent Samples Test table: shows Levene's test results and two rows of t-test output:
"Equal variances assumed" (pooled): uses a single pooled variance estimate.
"Equal variances not assumed" (Welch): uses separate variances. The lab instructs you to use this row.
From the Welch row:
t = 1.739
df = 243.282
Two-sided p = 0.083
Mean difference = 0.10961
95% CI for the difference: (-0.01457, 0.23378)
Levene's test in the lab gives F = 0.007, p = 0.931. Because p is large (well above 0.05), the variances appear similar. In this case, both rows give nearly identical results.
However, the lab instruction says "do not assume equal variances," so you read the Welch row regardless. In practice, many statisticians default to Welch because it is safe whether variances are equal or not.
The 95% CI for (mu_male - mu_female) is (-0.015, 0.234). Because this interval contains zero, it is consistent with no difference between the groups. The data do not provide enough evidence to conclude the means differ.
Compare the p-value to alpha (0.05):
Two-sided p = 0.083, which is greater than 0.05.
Fail to reject H0.
Conclusion in context: at the 5% significance level, there is not sufficient evidence that the population mean weekday sleep time differs between male and female college students.
Difference of sample means:
x-bar_1 - x-bar_2
Standard error (unequal variances, Welch):
SE = sqrt( s_1^2 / n_1 + s_2^2 / n_2 )
where s_1 and s_2 are the sample standard deviations and n_1, n_2 the sample sizes.
t-statistic (Welch):
t = (x-bar_1 - x-bar_2) / SE
From the lab: t = 0.10961 / 0.06304 = 1.739.
Welch degrees of freedom:
df = ( s_1^2/n_1 + s_2^2/n_2 )^2 / [ (s_1^2/n_1)^2/(n_1 - 1) + (s_2^2/n_2)^2/(n_2 - 1) ]
This often produces a non-integer. In the lab, df = 243.282. You do not need to compute this by hand; SPSS does it.
95% confidence interval for the difference of population means:
(x-bar_1 - x-bar_2) +/- t*(alpha/2, df) x SE
From the lab: 0.10961 +/- (1.970)(0.06304), giving (-0.01457, 0.23378).
Two-sample t-tests are among the most common inferential tools in practice. Medical researchers compare treatment vs. placebo groups. Marketers compare conversion rates between two ad variants (when summarised as means). HR analysts compare performance scores across office locations. Any time you have two independent groups and a continuous outcome, this is the starting framework.
Students often confuse "two columns of data" with "two independent samples." Two columns can also represent matched pairs (same subjects, two measurements). The deciding factor is whether there is a natural pairing.
Reading the wrong row of the SPSS output is a frequent mistake. When told not to assume equal variances, always read the "Equal variances not assumed" row. The two rows can give different p-values, especially with unequal sample sizes.
A p-value of 0.083 does not mean "the means are equal." It means the data do not provide strong enough evidence to conclude they differ. Failing to reject H0 is not the same as proving H0 true.
Students sometimes report the conclusion as applying to the sample. It does not. The sample means are what they are (you can see they differ by 0.1096 hours). The test is about whether that difference is large enough to infer a population-level difference.
For a two-tailed test, use the two-sided p-value. Do not halve it or double it. SPSS labels the columns clearly.
The most common exam setup: you are given data from two groups and asked to decide whether it is matched pairs or two-sample. If there is no pairing, it is two-sample.
State which group is "1" and which is "2" when writing hypotheses. If you write mu_male - mu_female, the sign of your mean difference must match that order.
Know the difference between "equal variances assumed" and "not assumed." If the question says "do not assume equal variances," you must use the Welch row.
A confidence interval that includes zero is consistent with failing to reject H0. Exams often test this link.
"Fail to reject H0" is the correct phrasing, never "accept H0." The test cannot prove the null is true.
True or False: A two-sample independent t-test requires the same subjects to appear in both groups. (False. Each subject belongs to exactly one group.)
Fill in the blank: When the instruction says "do not assume equal variances," you read the ____ row of the SPSS output. (Equal variances not assumed / Welch)
True or False: If the two-sided p-value is 0.083, you reject H0 at alpha = 0.05. (False. 0.083 > 0.05, so you fail to reject.)
Fill in the blank: Levene's test checks whether the two groups have equal ____. (variances)
True or False: "Fail to reject H0" is equivalent to "accept H0." (False. Failing to reject means the evidence is insufficient to conclude a difference; it does not prove the null is true.)
Q: A company wants to know whether employees in Office A have different average commute times from employees in Office B. 40 employees are sampled from each office. Is this a matched pairs or two-sample problem?
A: Two-sample. The employees in Office A are entirely different people from those in Office B, with no pairing.
Q: In the sleep study, the sample mean difference (male - female) is 0.1096 hours and the Welch two-sided p-value is 0.083. What is your conclusion at alpha = 0.05?
A: Fail to reject H0. At the 5% significance level, there is not sufficient evidence that the population mean weekday sleep time differs between male and female college students.
Q: The 95% CI for (mu_male - mu_female) is (-0.015, 0.234). A classmate says this proves males and females sleep the same amount. Is the classmate correct?
A: No. The CI includes zero, which is consistent with no difference, but it does not prove the means are equal. It simply means we cannot rule out zero as a plausible value for the true difference.
Q: Levene's test gives F = 0.007 and p = 0.931. What does this tell you?
A: The p-value is very large, so there is no evidence the two groups have different variances. The equal-variances assumption is reasonable here, though the lab still asks you to use the unequal-variances (Welch) row.
Q: Why does the Welch test often produce non-integer degrees of freedom?
A: The Welch df formula is a weighted combination of the two groups' variances and sample sizes. Because it divides and squares these quantities, the result is typically not a whole number. SPSS handles this automatically.
The two-sample t-test is the independent-groups counterpart to the matched pairs t-test (see the companion study notes). When you move to more than two groups, the natural extension is one-way ANOVA, which generalises the comparison to k groups.
Levene's test connects to the broader topic of checking assumptions before running inferential procedures. Similar assumption checks appear in regression (homoscedasticity) and ANOVA (equal variances across groups).
The Welch correction is an example of a robust procedure: it relaxes an assumption (equal variances) at a small cost in power, making it safer for general use. This idea of trading a bit of efficiency for broader applicability comes up again in nonparametric methods.
Two-sample t-test, independent samples t-test, Welch t-test, unpooled t-test, pooled t-test, equal variances assumed, equal variances not assumed, Levene's test, grouping variable, SPSS independent samples, comparison of means, gender differences in sleep, STAT 350 Lab 4, Purdue STAT 350, confidence interval for difference of means, two-tailed hypothesis test, fail to reject, p-value interpretation, Welch degrees of freedom, Satterthwaite approximation