Difficulty: Intermediate | Prerequisites: Sampling distributions, normal distribution basics, concept of a parameter vs. a statistic.
Big picture: This topic covers how to estimate and test claims about population proportions using sample data. Proportions come up whenever you are working with categorical data where the outcome is "yes/no," "success/failure," or any two-category split. The underlying distribution is the normal approximation to the binomial, which is why every procedure here uses a z-statistic rather than a t-statistic. If you are comfortable with the idea that a sample proportion is a point estimate of a population proportion, you are ready for this material.
When your data is categorical (proportions, percentages, rates), you use z-based procedures. A confidence interval gives you a plausible range for the true proportion. A hypothesis test tells you whether your sample data provide enough evidence to reject a claim about that proportion. One-sample procedures compare a single proportion to a hypothesised value; two-sample procedures compare two groups to each other.
Population proportion (p)
The true fraction of the population that has the characteristic of interest. This is a parameter, meaning you almost never know it exactly. You estimate it from the sample.
In simple terms, this is the "real answer" you are trying to find out.
Sample proportion (p-hat, p̂)
The fraction of successes observed in your sample: p̂ = x / n, where x is the number of successes and n is the sample size.
Think of it as your best estimate of p based on the data you collected.
Standard error (SE)
A measure of how much p̂ is expected to vary from sample to sample. For a single proportion: SE = √[p̂(1 − p̂) / n]. For a hypothesis test, you substitute the hypothesised value p₀ in place of p̂.
In simple terms, it tells you how "noisy" your estimate is. Larger samples produce smaller standard errors.
Confidence level (C)
The probability that the interval procedure will capture the true parameter if repeated many times. Common values are 90%, 95%, and 99%. A higher confidence level produces a wider interval.
Think of it as the long-run success rate of the method, not the probability that this particular interval is correct.
Null hypothesis (H₀)
The default claim about the population proportion, typically that p equals some specific value (p₀). You assume it is true unless the data give strong evidence against it.
Alternative hypothesis (Hₐ)
The claim you are testing for. It can be one-sided (p < p₀ or p > p₀) or two-sided (p ≠ p₀).
p-value
The probability of observing a result at least as extreme as the one you got, assuming the null hypothesis is true. A small p-value (typically below your significance level α) means the data are unlikely under H₀.
Think of it as: "If nothing interesting is going on, how surprised should I be by this result?"
z-statistic (z-score for a test)
The number of standard errors your sample proportion sits away from the hypothesised proportion: z = (p̂ − p₀) / √[p₀(1 − p₀) / n].
Purpose: Estimate the true population proportion p with a confidence interval.
When to use: You have one sample from one population and want a range of plausible values for p.
Conditions to check:
Random sample (or random assignment)
Independence: n ≤ 10% of the population (the 10% condition)
Large enough sample: np̂ ≥ 10 and n(1 − p̂) ≥ 10
Calculator function: 1-PropZInt
Inputs required: x (number of successes), n (sample size), C-Level (confidence level)
Purpose: Test a claim about the value of a single population proportion.
When to use: You want to know whether the data provide sufficient evidence that p differs from (or is greater/less than) a specific value p₀.
Conditions to check:
Random sample
Independence: n ≤ 10% of the population
Large enough sample: np₀ ≥ 10 and n(1 − p₀) ≥ 10 (note: you use p₀ here, not p̂)
Calculator function: 1-PropZTest
Inputs required: p₀ (hypothesised proportion), x, n, direction of Hₐ (≠, <, or >)
Purpose: Estimate the difference between two population proportions (p₁ − p₂) with a confidence interval.
When to use: You have independent samples from two populations and want a range for how much their proportions differ.
Conditions to check:
Both samples are random and independent of each other
Independence within each sample: n₁ ≤ 10% of population 1, n₂ ≤ 10% of population 2
Large enough samples: n₁p̂₁ ≥ 10, n₁(1 − p̂₁) ≥ 10, n₂p̂₂ ≥ 10, n₂(1 − p̂₂) ≥ 10
Calculator function: 2-PropZInt
Inputs required: x₁, n₁, x₂, n₂, C-Level
Purpose: Test whether two population proportions are equal (or whether one is greater/less than the other).
When to use: You want to know whether there is a statistically significant difference between two groups.
Conditions to check: Same as the two-proportion interval, but for the test you use the pooled proportion p̂c = (x₁ + x₂) / (n₁ + n₂) when checking the large-sample condition.
Calculator function: 2-PropZTest
Inputs required: x₁, n₁, x₂, n₂, direction of Hₐ (≠, <, or >)
One-proportion confidence interval:
p̂ ± z* × √[p̂(1 − p̂) / n]
where z* is the critical value for your confidence level (1.645 for 90%, 1.96 for 95%, 2.576 for 99%).
One-proportion z-test statistic:
z = (p̂ − p₀) / √[p₀(1 − p₀) / n]
Two-proportion confidence interval:
(p̂₁ − p̂₂) ± z* × √[p̂₁(1 − p̂₁)/n₁ + p̂₂(1 − p̂₂)/n₂]
Two-proportion z-test statistic:
z = (p̂₁ − p̂₂) / √[p̂c(1 − p̂c)(1/n₁ + 1/n₂)]
where p̂c = (x₁ + x₂) / (n₁ + n₂) is the pooled proportion.
Polling organisations use one-proportion confidence intervals every time they report a candidate's support with a "margin of error." Medical trials comparing the recovery rate of a treatment group against a control group rely on two-proportion z-tests to decide whether the treatment works.
"A 95% confidence interval means there is a 95% probability the true proportion is inside this interval." This is incorrect. The 95% refers to the method: if you repeated the sampling process many times, about 95% of the intervals you construct would capture the true proportion. Any single interval either contains p or it does not.
"You always use p̂ in the standard error formula." For a hypothesis test, you use the hypothesised value p₀ in the denominator, because you are assuming H₀ is true. You only use p̂ for confidence intervals.
"Failing to reject H₀ means H₀ is true." It means the data did not provide strong enough evidence to reject it. That is not the same as proving it correct.
"A small p-value tells you the effect is large." A small p-value means the result is unlikely under H₀. The actual size of the difference could be trivially small, especially with a very large sample.
⚠️ You will be asked to choose between a z-procedure and a t-procedure. The key: proportions always use z; means use t (in intro stats).
⚠️ Checking conditions is often worth marks on its own. Write them out explicitly, with numbers, every time.
⚠️ Know when to use p₀ vs. p̂ in the standard error. Tests use p₀. Intervals use p̂.
⚠️ For two-proportion tests, you must use the pooled proportion in the standard error. For two-proportion intervals, you do not pool.
True or False: A one-proportion z-test uses the sample proportion p̂ in the denominator of the test statistic.
Fill in the blank: The calculator function for a confidence interval comparing two proportions is __________.
True or False: If a 95% confidence interval for p₁ − p₂ contains 0, you would fail to reject H₀: p₁ = p₂ at the 5% significance level.
Fill in the blank: The condition np₀ ≥ 10 and n(1 − p₀) ≥ 10 is called the __________ condition.
True or False: Increasing the confidence level from 95% to 99% makes the interval narrower.
Answers: 1. False (it uses p₀). 2. 2-PropZInt. 3. True. 4. Large sample (or "success/failure" or "normal approximation") condition. 5. False (it makes the interval wider).
Q: A survey finds that 162 out of 400 adults support a new policy. Construct a 95% confidence interval for the true proportion of adults who support the policy. Which calculator function do you use, and what do you enter?
A: Use 1-PropZInt. Enter x = 162, n = 400, C-Level = 0.95. The sample proportion is 162/400 = 0.405. The interval is 0.405 ± 1.96 × √(0.405 × 0.595 / 400), which gives approximately (0.357, 0.453).
Q: A company claims that 70% of customers are satisfied. You survey 200 customers and find 126 are satisfied. At the 5% significance level, is there evidence that the true proportion differs from 0.70?
A: Use 1-PropZTest. H₀: p = 0.70, Hₐ: p ≠ 0.70. Enter p₀ = 0.70, x = 126, n = 200. The sample proportion is 0.63. The test statistic is z = (0.63 − 0.70) / √(0.70 × 0.30 / 200) ≈ −2.16. The p-value is approximately 0.031. Since 0.031 < 0.05, reject H₀. There is sufficient evidence that the true proportion differs from 0.70.
Q: In a study, 45 out of 150 patients in Group A recovered, and 60 out of 160 patients in Group B recovered. Is there a significant difference at α = 0.05?
A: Use 2-PropZTest. H₀: p₁ = p₂, Hₐ: p₁ ≠ p₂. Enter x₁ = 45, n₁ = 150, x₂ = 60, n₂ = 160. The pooled proportion is (45 + 60)/(150 + 160) = 105/310 ≈ 0.339. Compute the z-statistic and compare the resulting p-value to 0.05.
Q: What is the difference between when you pool proportions and when you do not?
A: You pool when performing a two-proportion hypothesis test, because under H₀ you assume the two proportions are equal, so you estimate that common proportion using the combined data. You do not pool when constructing a two-proportion confidence interval, because you are not assuming the proportions are equal.
This material connects directly to inference for means (the next topic), which uses t-procedures instead of z-procedures. The logic of confidence intervals and hypothesis tests is identical; only the distribution and the type of data change. Chi-squared tests for categorical data extend these ideas to situations with more than two categories.
proportion, population proportion, sample proportion, p-hat, z-test, z-interval, 1-PropZInt, 1-PropZTest, 2-PropZInt, 2-PropZTest, confidence interval for proportion, hypothesis test for proportion, two-proportion z-test, pooled proportion, significance test, p-value, null hypothesis, alternative hypothesis, standard error, margin of error, categorical data, binomial, normal approximation, success-failure condition, intro stats, AP Statistics, Purdue STAT