Confidence Intervals, STAT Intro Ch. 8–9 – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Basic probability, sampling distributions, normal and t distributions

TL;DR

Confidence intervals give you a range of plausible values for an unknown population parameter (usually μ) based on sample data. The interval itself is random, not the parameter. You need to know which formula to use (one-sample z, one-sample t, two-sample independent, or matched pairs), how to check assumptions, how to interpret the result in context, and how to calculate the sample size needed for a desired margin of error.


Key Terms

Confidence interval (CI)

A range of values, computed from sample data, that is likely to contain the true population parameter. In simple terms, it is your best estimate plus or minus a margin of error.

Confidence level (1 – α)

The proportion of all possible samples for which the CI procedure would capture the true parameter. A 95% confidence level means that if you repeated the sampling process many times, about 95% of those intervals would contain μ.

Margin of error (ME)

The half-width of a confidence interval. It measures how far the interval extends above and below the point estimate. In simple terms, it is the "plus or minus" part of your estimate.

Critical value (z or t)**

The number of standard errors you go out from the point estimate to achieve the desired confidence level. For z, you look it up from the standard normal table; for t, you need the degrees of freedom as well.

Degrees of freedom (df)

A parameter of the t-distribution that depends on sample size. For one-sample t problems, df = n – 1. Think of it as the number of independent pieces of information in your sample after estimating the mean.

Point estimate

A single value used to estimate the population parameter, such as the sample mean x̄. It sits at the centre of the confidence interval.

Standard error (SE)

The standard deviation of the sampling distribution of your estimator. For a sample mean with known σ, SE = σ/√n. For unknown σ, SE = s/√n. In simple terms, it tells you how much your sample mean tends to vary from sample to sample.

Matched pairs (paired samples)

A design where each observation in one group is naturally linked to an observation in the other (e.g. before/after measurements on the same subjects). You analyse the differences d = x₁ – x₂ rather than the raw values.

Independent samples

Two samples drawn separately with no pairing between individual observations. Each group's data is collected from a distinct set of subjects.

Tags: confidence interval, CI, margin of error, ME, critical value, z-star, t-star, degrees of freedom, df, standard error, point estimate, matched pairs, paired samples, independent samples


Core Content: CI Width, Meaning, and Assumptions

What Affects the Width of a Confidence Interval

  • Confidence level: Increasing the confidence level (e.g. from 90% to 95%) makes the interval wider, because a larger critical value is needed to be "more confident."

  • Sample size (n): Increasing n makes the interval narrower. The standard error shrinks as n grows, since SE is proportional to 1/√n.

  • Variability (σ or s): Greater spread in the data produces a wider interval. More variability means less precision in your estimate.

  • Summary: To get a narrower CI without sacrificing confidence, increase your sample size.

What Is Random About a Confidence Interval

The population parameter μ is fixed (not random). What is random is the interval itself, because it depends on the sample you happened to draw. Different samples produce different intervals. The confidence level describes the long-run success rate of the procedure, not the probability that one particular interval contains μ.

Assumptions for Inference

Before constructing a CI, you need to verify:

  • Random sampling: The data must come from a random sample (or randomised experiment). This cannot be checked with a graph; it is a design requirement.

  • Independence: Observations must be independent of each other. For sampling without replacement, the population should be at least 10 times the sample size.

  • Normality of the sampling distribution: For small samples, the underlying population should be approximately normal. Check with a QQ plot or histogram. For large samples (roughly n ≥ 30), the Central Limit Theorem makes this less critical.

  • Known vs. unknown σ: If σ is known, use the z procedure. If σ is unknown and estimated by s, use the t procedure.

Tags: CI width, confidence level effect, sample size effect, variability, assumptions, normality, independence, random sampling, CLT


Core Content: CI Formulas for All Four Cases

Choosing the Right Formula

Ask two questions: (1) How many samples? (2) Is σ known or unknown?

  • One sample, σ known → one-sample z

  • One sample, σ unknown → one-sample t

  • Two independent samples, σ unknown → two-sample t (independent)

  • Two related/paired samples → matched-pairs t

One-Sample z Interval

Used when the population standard deviation σ is known.

Formula: x̄ ± z_(α/2) · (σ / √n)

Degrees of freedom: not applicable (the z-distribution has no df parameter).

One-Sample t Interval

Used when σ is unknown and you estimate it with the sample standard deviation s.

Formula: x̄ ± t_(α/2, n−1) · (s / √n)

Degrees of freedom: df = n – 1.

Two-Sample t Interval (Independent)

Used to estimate the difference μ₁ – μ₂ when the two groups are unrelated.

Formula: (x̄₁ – x̄₂) ± t_(α/2, ν) · √(s₁²/n₁ + s₂²/n₂)

Degrees of freedom: will be given to you on the exam (the Welch approximation formula is complex).

Matched-Pairs t Interval

Used when observations are paired. Work with the differences dᵢ = x₁ᵢ – x₂ᵢ.

Formula: d̄ ± t_(α/2, n−1) · (s_D / √n)

Degrees of freedom: df = n – 1, where n is the number of pairs.

Interpreting a CI

When a question asks you to "interpret" the interval, you must state the conclusion in words in context. For example: "We are 95% confident that the true mean [context] is between [lower bound] and [upper bound]."

If it only asks you to "calculate," the numeric interval alone is sufficient.

Tags: one-sample z, one-sample t, two-sample t, independent samples CI, matched pairs CI, Welch, degrees of freedom, interpret CI


Core Content: One-Sided Bounds and Sample Size

One-Sided Confidence Bounds (One-Sample z)

Sometimes you only need a bound in one direction.

  • Lower bound: μ > x̄ – z_α · (σ / √n)

  • Upper bound: μ < x̄ + z_α · (σ / √n)

Notice the critical value is z_α, not z_(α/2). Because you are putting all of the α area into one tail, you use a smaller critical value than the two-sided case.

The same logic extends to t-based bounds, with t_(α, n−1) replacing z_α.

Calculating Required Sample Size

If you want your margin of error to be no larger than a specified value ME, you can solve for the necessary n before collecting data.

One-sample z case:

n = (z_(α/2) · σ / ME)²

You need to know (or assume) σ. Always round n up to the next whole number.

One-sample t case:

n = (t'_(α/2, n−1) · s / m)²

Here t'_(α/2, n−1) comes from a preliminary study that produced the estimate s. The degrees of freedom for the critical value are based on that preliminary study's sample size, not the new n you are solving for. Again, round up.

Tags: one-sided bound, lower bound, upper bound, sample size calculation, margin of error, required sample size, round up


Formulas and Diagrams

All four CI formulas share the same structure:

Point estimate ± (critical value) · (standard error)

The differences between cases come down to which point estimate, which critical value, and which standard error you plug in.

Case

Point estimate

Critical value

Standard error

One-sample z

x̄

z_(α/2)

σ/√n

One-sample t

x̄

t_(α/2, n−1)

s/√n

Two-sample independent t

x̄₁ – x̄₂

t_(α/2, ν)

√(s₁²/n₁ + s₂²/n₂)

Matched pairs t

d̄

t_(α/2, n−1)

s_D/√n

Sample size formulas:

  • z case: n = (z_(α/2) · σ / ME)²

  • t case: n = (t'_(α/2, n−1) · s / m)²


Common Misconceptions

  • "There is a 95% probability that μ is in this interval." Incorrect. The parameter is fixed. The 95% refers to the procedure's long-run success rate across many samples, not the probability for one specific interval.

  • Confusing z_(α/2) with z_α. For a two-sided CI, you split α across both tails, so you use z_(α/2). For a one-sided bound, you use z_α (all of α in one tail). Mixing these up changes your critical value.

  • Using z when σ is unknown. If you only have the sample standard deviation s, you must use the t-distribution. Using z underestimates the width of the interval.

  • Forgetting to round up sample size. The formula for n almost always gives a non-integer. You must round up to the next whole number, never down, to ensure the margin of error requirement is met.


Why It Matters / Exam Flags

⚠️ You will need to pick the correct column (z vs. t, one-sample vs. two-sample vs. paired). Read the problem carefully for clues: is σ given? Are observations naturally paired?

⚠️ "Interpret" on the exam means you must write a sentence in context ("We are 95% confident that..."). If the question only says "calculate," the interval alone is fine.

⚠️ Know the difference between one-sided bounds and two-sided intervals, especially the change in critical value (z_α vs. z_(α/2)).

⚠️ Sample size problems: always round up, and for the t case, use the critical value from the preliminary study's degrees of freedom.

⚠️ Be ready to explain what makes a CI wider or narrower. This is a common conceptual question.


Quick Self-Test

  1. True or False: Increasing the sample size makes a confidence interval wider. (False, it makes it narrower.)

  1. True or False: A 99% confidence interval is wider than a 95% confidence interval, all else equal. (True.)

  1. Fill in the blank: When σ is unknown, you use the ___ distribution instead of the z distribution. (t)

  1. True or False: For a one-sided lower bound, you use z_(α/2) as the critical value. (False, you use z_α.)

  1. Fill in the blank: When calculating required sample size, you always round ___. (up)


Practice Q&A

Q: A random sample of 36 students has a mean test score of 72 with a known population standard deviation of 12. Construct a 95% confidence interval for the population mean.

A: This is a one-sample z problem (σ known). z_(0.025) = 1.96. ME = 1.96 · (12/√36) = 1.96 · 2 = 3.92. The interval is 72 ± 3.92, or (68.08, 75.92).

Q: A researcher measures the blood pressure of 25 patients before and after a treatment. The mean difference is 5.2 mmHg with s_D = 8.1. Which CI formula should be used?

A: Matched-pairs t interval. The measurements are paired (same patients, before vs. after). The formula is d̄ ± t_(α/2, n−1) · (s_D / √n), with df = 24.

Q: Interpret the interval (3.1, 7.3) for the mean difference in blood pressure (before minus after treatment) at the 95% confidence level.

A: We are 95% confident that the true mean decrease in blood pressure after treatment is between 3.1 and 7.3 mmHg. Since the entire interval is positive, there is evidence that the treatment reduces blood pressure.

Q: How large a sample is needed to estimate a population mean within 2 units at 95% confidence, if σ = 10?

A: n = (z_(α/2) · σ / ME)² = (1.96 · 10 / 2)² = (9.8)² = 96.04. Round up: n = 97.

Q: What happens to the width of a 90% CI if you increase the confidence level to 99%, keeping everything else the same?

A: The interval becomes wider. A higher confidence level requires a larger critical value, which increases the margin of error.


Connections to Other Topics

Confidence intervals and hypothesis tests are two sides of the same coin. A two-sided CI at confidence level (1 – α) will exclude a hypothesised value μ₀ exactly when the corresponding hypothesis test rejects H₀ at significance level α. This connection is covered in detail in the hypothesis testing notes.

The assumptions here (normality, independence, random sampling) carry directly into every inference procedure you will meet for the rest of the course, including ANOVA.

The concept of standard error reappears whenever you do inference on any parameter, not just the mean.


Related Terms / Search Tags

confidence interval, CI, margin of error, ME, critical value, z-star, t-star, z-interval, t-interval, one-sample z, one-sample t, two-sample t, independent samples, matched pairs, paired t, degrees of freedom, df, standard error, SE, point estimate, sample size calculation, confidence level, alpha, one-sided bound, upper bound, lower bound, interpret CI, width of CI, assumptions for inference, normality, CLT, central limit theorem, sampling distribution