Difficulty: Intermediate to Advanced | Prerequisites: Basic probability concepts (events, outcomes, probability as a fraction)
This material covers the statistical backbone of data science: how to model random events with probability distributions, how sample averages behave as you collect more data (the Central Limit Theorem), and how to formally test whether observed data contradicts a claim (hypothesis testing). These three ideas thread through nearly every analysis you will do, from A/B testing in tech to quality control in manufacturing. You should be comfortable with basic probability (what a probability is, how to compute a mean) before working through this.
The binomial distribution models the number of successes in repeated trials. The Central Limit Theorem says sample means become normally distributed as sample size grows. Hypothesis testing uses z-scores and p-values to decide whether data is consistent with a claimed population parameter.
Probability Mass Function (PMF)
A function that gives the probability of each possible value of a discrete random variable. The PMF histogram shows these probabilities as bar heights. In simple terms, it answers "how likely is each outcome?"
Binomial distribution
Models the number of successes in n independent trials, each with the same probability p of success. Written as B(n, p). Think of it as counting how many times something works out of n attempts.
Bernoulli probability (p)
The probability of success on a single trial. Estimated from data as the sample mean number of successes divided by the number of trials per experiment. In simple terms, it is the per-attempt success rate.
Cumulative Distribution Function (CDF)
The probability that a random variable takes a value less than or equal to k: P(X ≤ k) = P(X = 0) + P(X = 1) + ... + P(X = k). In simple terms, it is the running total of probabilities from the left.
Sample mean (x-bar)
The average of observed values in a sample. Used to estimate the population mean.
Central Limit Theorem (CLT)
States that the sampling distribution of the sample mean approaches a normal distribution as the sample size grows, regardless of the shape of the original population distribution. This is why normal distributions appear everywhere in statistics.
Standard error (SE)
The standard deviation of the sampling distribution of the sample mean: SE = σ / √n. It measures how much sample means typically vary from the true population mean. Larger samples produce smaller standard errors.
Z-score
The number of standard deviations a value is from the mean: z = (x - μ) / σ. For sample means, z = (x-bar - μ) / (σ / √n). In simple terms, it converts any value to a position on the standard normal curve.
Null hypothesis (H₀)
The default claim being tested, usually that a parameter equals a specific value (e.g., μ = 250). You assume it is true until the data provides strong evidence against it.
Alternative hypothesis (H₁)
The claim you are trying to find evidence for. It contradicts H₀. If H₁ is μ ≠ 250, you are testing whether the true mean differs from 250 in either direction (two-tailed).
P-value
The probability of observing data as extreme as (or more extreme than) what you got, assuming H₀ is true. A small p-value means your data would be very unusual if H₀ were correct. In simple terms, it is how surprised you should be if the null hypothesis is right.
Significance level (α)
The threshold for rejecting H₀. Common values are 0.05 and 0.01. If the p-value is less than α, you reject H₀.
68-95-99.7 rule (empirical rule)
For a normal distribution: about 68% of values fall within 1 standard deviation of the mean, 95% within 2, and 99.7% within 3. This lets you estimate tail probabilities without a z-table.
A probability mass histogram shows the probability of each outcome on the y-axis and the outcome value on the x-axis.
From frequency to probability: divide each frequency by the total number of trials. If 11 out of 20 trials had 9 successes, P(X = 9) = 11/20 = 0.55.
All bars must sum to 1. This is a fundamental check: if your probabilities do not add to 1, something is wrong.
Outcomes with zero frequency still exist on the x-axis but have a bar height of zero.
The binomial distribution applies when you have n independent trials, each with the same probability p of success.
Estimating p from data: compute the sample mean number of successes (weighted average of outcomes), then divide by n (the number of trials per experiment). Example: mean = (7×1 + 8×3 + 9×11 + 10×5) / 20 = 9, so p = 9/10 = 0.9.
PMF formula: P(X = k) = C(n, k) × p^k × (1 - p)^(n - k), where C(n, k) is "n choose k".
Useful identities: C(n, 0) = C(n, n) = 1, and C(n, 1) = C(n, n-1) = n.
The CDF gives P(X ≤ k). To find P(X ≥ k), use the complement: P(X ≥ k) = 1 - P(X ≤ k - 1), or simply add the individual probabilities for values k, k+1, ..., n.
Example: P(X ≥ 9) = P(X = 9) + P(X = 10) = 0.39 + 0.35 = 0.74. This means there is a 74% chance of transmitting at least 9 out of 10 packets successfully.
If X has mean μ and standard deviation σ, the sample mean of n observations is approximately normally distributed:
Distribution of the sample mean: N(μ, σ² / n), meaning the mean stays the same but the variance shrinks.
Standard error: SE = σ / √n.
Example: μ = 10, σ = 4, n = 64. Then the sample mean follows N(10, 0.5²), so SE = 4 / √64 = 0.5.
The CLT works regardless of the original distribution's shape, provided n is large enough (typically n ≥ 30 is sufficient).
To find probabilities involving a sample mean, convert to a z-score: z = (x-bar - μ) / SE.
z = 1 means the value is 1 standard error above the mean. By the 68-95-99.7 rule, 68% of values fall within ±1 SE, so 32% are outside, and 16% are above +1 SE.
z = 2: about 2.5% in each tail.
z = 3: about 0.15% in each tail.
Example: Pr(sample mean > 10.5) when μ = 10, SE = 0.5. z = (10.5 - 10) / 0.5 = 1. Pr(Z > 1) = (1 - 0.68) / 2 = 0.16.
A z-test follows four steps:
Step 1, state hypotheses: H₀: μ = μ₀ (the claimed value). H₁: μ ≠ μ₀ (two-tailed) or μ > μ₀ / μ < μ₀ (one-tailed).
Step 2, compute the z-statistic: z = (x-bar - μ₀) / (σ / √n).
Step 3, find the p-value: for a two-tailed test, p = 2 × P(Z > |z|). Use the 68-95-99.7 rule or a z-table.
Step 4, compare p to α and decide: if p < α, reject H₀. State the conclusion in context.
Example: μ₀ = 250, x-bar = 253, σ = 6, n = 36. z = (253 - 250) / (6/√36) = 3/1 = 3. Two-tailed p = 2 × P(Z > 3) = 2 × 0.0015 = 0.003. Since 0.003 < 0.01 = α, reject H₀. There is strong evidence the true mean differs from 250 mL.
Binomial PMF
P(X = k) = C(n, k) × p^k × (1 - p)^(n - k)
where C(n, k) = n! / (k! × (n - k)!)
Sample mean (x-bar)
x-bar = (Σ xᵢ) / n
Bernoulli probability estimate
p = x-bar / (number of trials per experiment)
Standard error of the sample mean
SE = σ / √n
Z-score for a sample mean
z = (x-bar - μ) / (σ / √n)
CDF
P(X ≤ k) = Σ P(X = i) for i = 0 to k
Two-tailed p-value
p-value = 2 × P(Z > |z|)
68-95-99.7 rule tail probabilities
P(Z > 1) ≈ 0.16
P(Z > 2) ≈ 0.025
P(Z > 3) ≈ 0.0015
Confusing the PMF with the CDF. The PMF gives the probability of exactly one value. The CDF gives the probability of that value or anything below it. They are related but answer different questions.
Forgetting to double the tail for a two-tailed test. If H₁ is μ ≠ μ₀, the p-value is 2 × P(Z > |z|), not just P(Z > |z|). Students lose marks by computing only one tail.
Using σ instead of σ/√n for the sample mean. The standard deviation of the sample mean (the standard error) is σ/√n, not σ. The sample mean is less variable than individual observations.
Thinking the CLT only works for normal populations. The whole point of the CLT is that it works for any distribution, provided n is large enough.
Interpreting the p-value as the probability that H₀ is true. The p-value is the probability of the observed data given H₀, not the probability of H₀ given the data. This is a subtle but important distinction.
⚠️ The exam asks you to compute binomial probabilities by hand using C(n, k) identities. Memorise C(n, 0) = C(n, n) = 1 and C(n, 1) = C(n, n-1) = n.
⚠️ You will be asked to estimate p from data. The method is: sample mean of successes ÷ number per trial.
⚠️ CLT questions give you μ, σ, and n, and ask you to write the distribution of the sample mean. The answer is always N(μ, (σ/√n)²).
⚠️ The 68-95-99.7 rule is the primary tool for estimating probabilities on this exam (no z-table is provided).
⚠️ Hypothesis testing questions walk through all four steps. You must state H₀ and H₁, compute z, find the p-value, and state your decision in context.
⚠️ You must justify whether the test is one-tailed or two-tailed based on H₁.
If P(X = 7) = 0.05, P(X = 8) = 0.15, P(X = 9) = 0.55, P(X = 10) = 0.25, what is P(X ≥ 9)?
True or False: the standard error of the sample mean increases as n increases.
If z = 2, what is the approximate probability P(Z > 2) using the 68-95-99.7 rule?
Fill in the blank: for a two-tailed test, p-value = ___ × P(Z > |z|).
True or False: rejecting H₀ at α = 0.01 requires stronger evidence than rejecting at α = 0.05.
Answers: 1. 0.55 + 0.25 = 0.80. 2. False (it decreases, because SE = σ/√n). 3. Approximately 0.025. 4. 2. 5. True (α = 0.01 is a stricter threshold).
Q: A wireless network sends 10 packets per trial. In 20 trials, the results are: 7 successes (1 trial), 8 (3 trials), 9 (11 trials), 10 (5 trials). Estimate the Bernoulli probability p.
A: Sample mean = (7×1 + 8×3 + 9×11 + 10×5) / 20 = 180/20 = 9. p = 9/10 = 0.9.
Q: Using p = 0.9 and n = 10, compute P(X = 9).
A: P(X = 9) = C(10, 9) × 0.9⁹ × 0.1¹ = 10 × 0.39 × 0.1 = 0.39.
Q: If μ = 10, σ = 4, and n = 64, what is the standard error of the sample mean?
A: SE = 4 / √64 = 4 / 8 = 0.5.
Q: A coffee machine claims μ₀ = 250 mL. You observe x-bar = 253 with σ = 6 and n = 36. Compute z and state whether you reject H₀ at α = 0.01.
A: z = (253 - 250) / (6/√36) = 3/1 = 3. Two-tailed p = 2 × 0.0015 = 0.003. Since 0.003 < 0.01, reject H₀.
Q: Why is the test in the previous question two-tailed?
A: Because H₁ is μ ≠ 250, which considers deviations in both directions (too high or too low).
The binomial distribution connects to the Bernoulli distribution (a single trial) and extends to the normal distribution via the CLT when n is large. Hypothesis testing is the formal version of what data scientists do informally in every A/B test, and the z-test pattern generalises to t-tests (when σ is unknown) and chi-squared tests (for categorical data) later in the course.
PMF, probability mass function, binomial distribution, Bernoulli probability, CDF, cumulative distribution function, Central Limit Theorem, CLT, standard error, z-score, z-test, null hypothesis, alternative hypothesis, p-value, significance level, 68-95-99.7 rule, empirical rule, hypothesis testing, two-tailed test, sample mean, ECE 20875, Python for Data Science, Purdue