Confidence Intervals, Hypothesis Testing and Two-Sample Inference, STAT 302 Ch. 9–11 – Study Notes
offline

Difficulty: Intermediate–Advanced | Prerequisites: Ch. 7 (sampling distributions, standard error, CLT), basic z-table and t-table usage.

Big Picture

These three chapters form the heart of introductory inference. Chapter 9 introduces confidence intervals: given sample data, what range of values is plausible for the true population parameter? Chapter 10 introduces hypothesis testing: given a claim about the population, does the data provide enough evidence to reject it? Chapter 11 extends both tools to comparing two populations. You need the sampling distribution ideas from Chapter 7 (standard error, CLT) as the foundation for everything here. This material accounts for the largest portion of the final exam.

TL;DR

A confidence interval gives a range of plausible values for a population parameter, built from the sample estimate plus or minus a margin of error. A hypothesis test formalises the question "is this result surprising enough to reject the default assumption?" using a test statistic and p-value. Both techniques come in z-versions (σ known) and t-versions (σ unknown), and both extend to two-sample problems (independent samples and matched pairs). Practical significance asks whether a statistically significant result matters in the real world.


Key Terms

Confidence interval (CI)

An interval of the form estimate ± margin of error that, at a stated confidence level, is expected to capture the true population parameter.

Think of it as a net you cast around your sample result. A 95% CI means that if you repeated the study many times, about 95% of those nets would contain the true value.

Margin of error (ME)

The half-width of a confidence interval. ME = critical value × standard error.

Critical value

The z or t value that marks the boundary of the middle C% of the sampling distribution. For a 95% CI with known σ, the critical value is z(α/2) = 1.96.

Confidence level (C)

The proportion of all possible confidence intervals (from repeated sampling) that would contain the true parameter. Common levels: 90%, 95%, 99%. Related to α by C = 1 – α.

Null hypothesis (H₀)

The default claim, assumed true until evidence says otherwise. Always contains an equals sign. Represents "no effect" or "no difference."

Think of it as the "innocent until proven guilty" assumption.

Alternative hypothesis (Hₐ)

The claim you are trying to find evidence for. Contradicts H₀. Can be one-sided (< or >) or two-sided (≠).

Test statistic

A standardised value that measures how far the sample result is from the null hypothesis value, in units of standard error. z(ts) = (x̄ – μ₀) / (σ / √n) when σ is known; t(ts) = (x̄ – μ₀) / (s / √n) when σ is unknown.

P-value

The probability of observing a test statistic as extreme as (or more extreme than) the one calculated, assuming H₀ is true. A small p-value means the data is inconsistent with H₀.

Significance level (α)

The threshold for rejecting H₀. If p-value ≤ α, reject H₀. Common values: 0.01, 0.05, 0.10.

Type I error

Rejecting H₀ when it is actually true (false positive). Probability = α.

Type II error

Failing to reject H₀ when it is actually false (false negative). Probability = β.

Power

The probability of correctly rejecting H₀ when it is false. Power = 1 – β.

Think of it as the sensitivity of your test: how good is it at detecting a real effect?

Degrees of freedom (df)

For a one-sample t-test: df = n – 1. For two-sample pooled t: df = n₁ + n₂ – 2. For Welch (unpooled): df is computed via the Satterthwaite formula.

Effect size

The difference between the experimental value and the reference value, divided by the standard deviation. Measures the practical magnitude of an effect. Rule of thumb: < 0.2 small, 0.2–0.8 moderate, > 0.8 large.

Practical significance

Whether a statistically significant result matters in the real world. A result can be statistically significant but practically meaningless (e.g. a drug lowers temperature by 0.4°C when the clinical threshold is 1°C).


Core Content: Confidence Intervals

Assumptions

  • The data must come from a properly randomised sample (SRS).

  • The sampling distribution of x̄ must be (approximately) normal, either because the population is normal or because n is large enough for the CLT.

CI When σ Is Known (z-interval)

Formula: x̄ ± z(α/2) × (σ / √n)

  • For 95% confidence: z(α/2) = 1.96

  • For 99% confidence: z(α/2) = 2.576

  • R code to find the critical value: z <- qnorm((1-C)/2, lower.tail = FALSE)

Worked example: 100 corn plots, x̄ = 123.8 bushels, σ = 12.3. Standard error = 12.3 / √100 = 1.23. The 95% CI is 123.8 ± 1.96(1.23) = (121.4, 126.2).

CI When σ Is Unknown (t-interval)

Replace σ with the sample standard deviation s, and use the t-distribution with df = n – 1.

Formula: x̄ ± t(α/2, n−1) × (s / √n)

  • R code: t <- qt((1-C)/2, df, lower.tail = FALSE)

  • The t-distribution has heavier tails than the z-distribution, so t critical values are larger, producing wider intervals.

Worked example: 41 students, x̄ = 7.16 hours, s = 3.56, df = 40. t(0.025, 40) = 2.021. The 95% CI is 7.16 ± 2.021(3.56 / √41) = (6.036, 8.284).

Interpreting a CI Correctly

  • Correct: "We are 95% confident that the interval captures the true mean μ."

  • Incorrect: "We are 95% confident that μ lies in the interval." (Subtle but important: μ is fixed, the interval is random.)

  • Incorrect: "95% of the data falls in this interval." (The CI is about μ, not individual data points.)

Making CIs Narrower

  • Reduce the confidence level (smaller critical value).

  • Increase n (smaller standard error).

  • Reduce σ (tighter measurements, though often not under the researcher's control).

Confidence Bounds (One-Sided)

  • Upper bound: μ < x̄ + z(α) × (σ / √n). Use z(α) instead of z(α/2) because all the error is on one side.

  • Lower bound: μ > x̄ – z(α) × (σ / √n).

  • R code for bound critical value: z <- qnorm(1-C, lower.tail = FALSE)

Sample Size for a Target Margin of Error

  • z-interval: n = (z(α/2) × σ / ME)²

  • t-interval: n = (t(α/2, n'−1) × s / ME)², using a preliminary estimate of s. Always round up.


Core Content: Hypothesis Testing

The Four-Step Framework

  1. Identify the parameter and describe it in context. Example: μ = the true average weight of cherry tomato boxes.

  1. State the hypotheses in symbols only. H₀ always contains =. Hₐ uses ≠, <, or >.

  1. Calculate the test statistic and p-value.

  1. Make a decision and state the conclusion in context (plain English, no symbols).

Setting Up Hypotheses

  • Two-sided: H₀: μ = μ₀, Hₐ: μ ≠ μ₀. Use when you have no prior reason to expect one direction.

  • One-sided upper: H₀: μ ≤ μ₀, Hₐ: μ > μ₀. The alternative is what you want to show.

  • One-sided lower: H₀: μ ≥ μ₀, Hₐ: μ < μ₀.

  • Decide the direction before looking at the data. If unsure, default to two-sided.

Test Statistics

  • σ known: z(ts) = (x̄ – μ₀) / (σ / √n)

  • σ unknown: t(ts) = (x̄ – μ₀) / (s / √n), df = n – 1

P-Value Computation

Test type

P-value formula

R code (σ known)

R code (σ unknown)

Upper-tailed

P(Z > z(ts))

pnorm(zts, lower.tail=FALSE)

pt(tts, df, lower.tail=FALSE)

Lower-tailed

P(Z < z(ts))

pnorm(zts, lower.tail=TRUE)

pt(tts, df, lower.tail=TRUE)

Two-tailed

2P(Z < -

z(ts)

)

Decision Rule

  • If p-value ≤ α: reject H₀. Conclusion: "The data provides (strong) support for Hₐ [stated in context]."

  • If p-value > α: fail to reject H₀. Conclusion: "The data does not provide sufficient support for Hₐ [stated in context]."

  • Never say "accept H₀." You can only fail to reject it.

Worked Example (z-test)

Cherry tomatoes labelled 227 g. Sample of 4 boxes: x̄ = 222 g, σ = 5 g, α = 0.05.

  • H₀: μ = 227, Hₐ: μ ≠ 227

  • z(ts) = (222 – 227) / (5 / √4) = −2

  • p-value = 2 × P(Z ≤ −2) = 2 × 0.0228 = 0.0455

  • 0.0455 < 0.05 → reject H₀ (barely). The data gives some support to the claim that the true weight differs from 227 g.

Connection Between CI and Hypothesis Test

When C + α = 1 (e.g. 95% CI and α = 0.05), the CI and two-sided test give consistent answers. If the null value falls outside the CI, the test rejects H₀, and vice versa.


Core Content: Type I and Type II Errors and Power

Error Types

Reject H₀

Fail to reject H₀

H₀ is true

Type I error (prob = α)

Correct decision

H₀ is false

Correct decision (Power = 1 – β)

Type II error (prob = β)

Examples

  • Judicial system: H₀: innocent. Type I = convicting an innocent person. Type II = acquitting a guilty person. In the US system, Type I is considered worse.

  • Grading: H₀: answer is correct. Type I = marking a correct answer wrong. Type II = marking a wrong answer correct.

  • Process improvement (paint drying): H₀: μ ≥ 75 min. Type I = adopting the new method when it is no better. Type II = keeping the old method when the new one is better.

Power

  • Power = 1 – β = probability of correctly rejecting H₀ when it is false.

  • You never say "accept H₀" because when β is large (power is small), you cannot be confident that failing to reject was the right call.

How to Increase Power

  • Increase n: this is the practical lever. Larger samples narrow both distributions, reducing both α and β.

  • Increase α: this decreases β but is usually not an option because α is already set at the maximum acceptable level.

  • Decrease σ: narrows both curves, but σ is usually already minimised by good experimental design.

  • Larger true effect (|μₐ – μ₀|): the further the real mean is from the null value, the easier it is to detect. But you cannot control the true mean.


Core Content: Two-Sample Inference

Independent Samples

Two groups drawn independently from two distinct populations.

Assumptions:

  • Each group is an SRS from its population.

  • The statistic in each group has (approximately) a normal distribution.

Known σ:

  • Test statistic: z(ts) = [(x̄₁ – x̄₂) – Δ₀] / √(σ₁²/n₁ + σ₂²/n₂)

  • CI: (x̄₁ – x̄₂) ± z(α/2) × √(σ₁²/n₁ + σ₂²/n₂)

Unknown σ, two approaches:

  • Pooled (equal variances assumed): combine both sample variances into a pooled estimate S²(p). df = n₁ + n₂ – 2.

  • Unpooled / Welch (unequal variances): keep the sample variances separate. The degrees of freedom are computed via the Satterthwaite approximation (a complex formula, usually done in software).

Test statistic (unpooled): t'(ts) = [(x̄₁ – x̄₂) – Δ₀] / √(s₁²/n₁ + s₂²/n₂)

CI (unpooled): (x̄₁ – x̄₂) ± t(α/2, ν) × √(s₁²/n₁ + s₂²/n₂)

Robustness Guidelines for Two-Sample t

  • n₁ + n₂ < 15: population distributions should be close to normal.

  • 15 < n₁ + n₂ < 40: mild skewness is acceptable.

  • n₁ + n₂ > 40: the procedure is usually valid.

Matched Pairs

The same subjects measured twice (before/after) or naturally paired subjects. You compute the difference d = observation₁ – observation₂ for each pair, then do a one-sample t-test on the differences.

Assumptions:

  • Each pair is from a population of pairs (SRS of pairs).

  • The pairs are independent of each other.

  • The differences d(i) have a normal distribution with mean μ(D) and standard deviation σ(D).

Test statistic: t(ts) = (d̄ – Δ₀) / (s(D) / √n), df = n – 1

Worked example (nursing scores): 8 nurses rated before and after training. d̄ = −1.30 (pre – post), s(D) = 0.861, α = 0.01. H₀: μ(D) ≥ 0, Hₐ: μ(D) < 0.

  • t(ts) = (−1.30) / (0.861 / √8) = −4.27, df = 7

  • p-value = P(T < −4.27) = 0.00185

  • 0.00185 < 0.01 → reject H₀. The data provides strong support that the training improved scores.

Choosing Between Independent and Paired

  • Use paired when there is a large variance within each population and a strong correlation linking the two measurements (a lurking variable that pairs the observations).

  • Use independent when each population has a small variance and the two groups are truly unrelated.


Core Content: Statistical vs Practical Significance and Effect Size

Statistical Significance Is Not the Whole Story

A result can be statistically significant (p-value ≤ α) yet practically meaningless. Three examples from the course:

  • A drug lowers temperature by 0.4°C (p < 0.01), but clinical benefit requires at least 1°C. Statistically significant, practically useless.

  • A process improvement is real (p < 0.01), but the cost of implementing it exceeds the savings. Statistically significant, not worth doing.

  • New truck brakes reduce stopping distance from 512 feet to 510 feet. Statistically significant, but 2 feet is negligible when a car is in the way.

Effect Size

Effect size = (experimental value – reference value) / standard deviation

This standardises the difference so you can judge its magnitude regardless of the units.

Rule-of-thumb thresholds (when no domain-specific guidance is available):

  • < 0.2: small (practically indistinguishable)

  • 0.2 to 0.8: moderate

  • > 0.8: large (practically distinct)

How to Assess Practical Significance

  • If you fail to reject H₀, practical significance does not arise.

  • If you reject H₀, compute the difference (between the estimate and the null value, or the closest limit of the CI) and the effect size.

  • Both the difference and the effect size must be large for the result to be practically significant. Stakeholders determine what counts as "large" in context; on exams, the question will tell you or you use the rule-of-thumb thresholds.


Formulas

One-Sample Confidence Intervals

\text{z-interval: } \bar{x} \pm z_{\alpha/2} \frac{\sigma}{\sqrt{n}}
\text{t-interval: } \bar{x} \pm t_{\alpha/2,\, n-1} \frac{s}{\sqrt{n}}

One-Sample Test Statistics

z_{ts} = \frac{\bar{x} - \mu_0}{\sigma / \sqrt{n}}
t_{ts} = \frac{\bar{x} - \mu_0}{s / \sqrt{n}}, \quad df = n - 1

Two-Sample Independent (unknown σ, unpooled)

t'_{ts} = \frac{(\bar{x}_1 - \bar{x}_2) - \Delta_0}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}}
\text{CI: } (\bar{x}_1 - \bar{x}_2) \pm t_{\alpha/2,\,\nu} \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}

Matched Pairs

t_{ts} = \frac{\bar{d} - \Delta_0}{s_D / \sqrt{n}}, \quad df = n - 1

Sample Size

n = \left( \frac{z_{\alpha/2} \cdot \sigma}{ME} \right)^2 \quad \text{(z)} \qquad n = \left( \frac{t_{\alpha/2,\,n'-1} \cdot s}{ME} \right)^2 \quad \text{(t)}

Effect Size

\text{effect size} = \frac{\text{difference}}{\text{standard deviation}}

Common Misconceptions

  • Students often say "we are 95% confident that μ lies in the interval." The correct phrasing is "we are 95% confident that the interval captures μ." μ is fixed; the interval is what varies from sample to sample.

  • Students often say "accept H₀" when the p-value is large. You can only fail to reject H₀. Failing to reject does not mean H₀ is true; it means the data did not provide strong enough evidence against it.

  • Students confuse the confidence level with the probability that μ is in this particular interval. Once the interval is computed, μ is either in it or not. The 95% refers to the long-run success rate of the method.

  • Students sometimes pick one-sided vs two-sided after seeing the data. The direction must be chosen before you look at the results, based on the research question.

  • Students confuse statistical significance with practical significance. A tiny effect can be statistically significant if n is large enough. Always check whether the result matters in context.


Why It Matters / Exam Flags

⚠️ Know when to use z vs t. If σ is given, use z. If only s is given, use t with df = n – 1.

⚠️ The four-step hypothesis test framework (parameter, hypotheses, test statistic + p-value, decision + context) is expected in full on the exam. Missing any step costs marks.

⚠️ Be able to write the conclusion in plain English, in context. No symbols, no jargon. "The data provides strong support (p = 0.002) that the average number of credit cards is greater than 2."

⚠️ Know the relationship: C + α = 1. A 95% CI corresponds to a two-sided test at α = 0.05. An upper-tailed test at α corresponds to a lower confidence bound at level C = 1 – α.

⚠️ For two-sample problems, identify whether the design is independent or paired before choosing a formula. The signal for paired data: the same subjects measured twice, or subjects matched in pairs.

⚠️ Practical significance questions appear on the exam. If you reject H₀, compute the difference and effect size and state whether the result matters in context.


Quick Self-Test

  1. True or False: If σ is unknown, you should use the z-distribution. (False, use the t-distribution.)

  1. Fill in the blank: The degrees of freedom for a one-sample t-test are _____. (n – 1)

  1. True or False: A p-value of 0.03 means there is a 3% chance that H₀ is true. (False. The p-value is the probability of the observed data, assuming H₀ is true.)

  1. True or False: Increasing the sample size increases the power of a test. (True)

  1. Fill in the blank: In a matched pair design, you compute _____ for each pair and then perform a one-sample t-test on those values. (The difference d)


Practice Q&A

Q: A sample of 64 light bulbs has a mean lifetime of 1,200 hours. The population standard deviation is 100 hours. Construct a 99% confidence interval for the true mean lifetime.

A: z(0.005) = 2.576. SE = 100 / √64 = 12.5. CI = 1200 ± 2.576(12.5) = (1167.8, 1232.2). We are 99% confident the true mean lifetime is captured by (1167.8, 1232.2) hours.

Q: A company claims its cereal boxes contain 500 g on average. A sample of 25 boxes yields x̄ = 495 g and s = 10 g. Test at α = 0.05 whether the true mean is less than 500 g.

A: Step 1: μ = true mean weight of cereal boxes. Step 2: H₀: μ ≥ 500, Hₐ: μ < 500. Step 3: t(ts) = (495 – 500) / (10 / √25) = −2.5, df = 24. p-value = P(T < −2.5) ≈ 0.0098. Step 4: 0.0098 < 0.05 → reject H₀. The data provides strong support that the true mean weight is less than 500 g.

Q: What is the Type I error in the cereal box problem above?

A: Concluding the boxes are underfilled (rejecting H₀) when in fact they contain 500 g on average.

Q: Two groups of students take different prep courses. Group A (n = 30): x̄ = 78, s = 8. Group B (n = 35): x̄ = 82, s = 10. Is there a significant difference at α = 0.05?

A: H₀: μ₁ – μ₂ = 0, Hₐ: μ₁ – μ₂ ≠ 0. SE = √(64/30 + 100/35) = √(2.133 + 2.857) = √4.99 = 2.234. t'(ts) = (78 – 82) / 2.234 = −1.79. Using the Welch approximation for df (computed by software), the two-tailed p-value is approximately 0.078. Since 0.078 > 0.05, we fail to reject H₀. There is not sufficient evidence to conclude the prep courses produce different mean scores.

Q: A study finds that a new training programme raises nurse competency scores by an average of 1.30 points (s(D) = 0.861, n = 8 pairs). The effect size is 0.39 / 0.861 = 0.453. Is this practically significant?

A: The effect size of 0.453 is moderate (between 0.2 and 0.8). Whether it is practically significant depends on context: if the measurement precision is coarser than 0.39 points, the improvement may not be detectable in practice.


Connections to Other Topics

This connects to sampling distributions (Ch. 7) because every CI and test statistic formula uses the standard error σ / √n (or s / √n). The CLT justifies treating the sampling distribution as normal when n is large.

This connects to experimental design (Ch. 8) because whether you can draw a causal conclusion from a significant result depends on whether the data came from a randomised experiment or an observational study.

The matched pair t-test is conceptually a one-sample t-test applied to the differences, so mastering the one-sample case carries directly over.


Related Terms / Search Tags

confidence interval, margin of error, critical value, z-interval, t-interval, hypothesis test, null hypothesis, alternative hypothesis, p-value, significance level, alpha, test statistic, Type I error, Type II error, power, beta, degrees of freedom, one-sample z-test, one-sample t-test, two-sample t-test, independent samples, matched pairs, paired t-test, pooled variance, Welch t-test, Satterthwaite approximation, effect size, practical significance, statistical significance, confidence bound, upper bound, lower bound, sample size formula, STAT 302, Purdue statistics