Source: STAT 350 Exam 1, Purdue University
Tags: normal distribution, z-score, standard normal, percentile, empirical rule, 68-95-99.7, forward problem, backward problem, IQR, bell curve, Gaussian
Difficulty: Introductory to Intermediate | Prerequisites: Basic probability concepts, familiarity with integration or table lookup.
The normal distribution is the single most important continuous distribution in introductory statistics. Nearly every inferential method you will encounter later in this course (confidence intervals, hypothesis tests) leans on it. If you understand how to convert between raw values and z-scores, read the standard normal table in both directions, and compute percentiles and IQR from a normal model, you have the mechanical backbone for most of what follows. You should already be comfortable with the idea of a probability density function (PDF) and know that area under a curve represents probability.
A normal distribution is fully described by its mean and standard deviation. Converting to z-scores lets you use one universal table to find any probability or percentile. "Forward" problems go from an x-value to a probability; "backward" problems go from a probability to an x-value.
Normal distribution (Gaussian distribution)
A continuous, symmetric, bell-shaped distribution parameterised by mean μ and standard deviation σ. Written X ~ N(μ, σ²). In simple terms, it is the classic "bell curve" where most values cluster near the centre and thin out symmetrically toward the tails.
Z-score (standard score)
The number of standard deviations a value sits above or below the mean: z = (x - μ) / σ. Think of it as a universal ruler that lets you compare values from different normal distributions on the same scale.
Standard normal distribution
The special case N(0, 1), with mean 0 and standard deviation 1. Every normal distribution can be converted to this one via the z-score formula, which is why there is only one z-table.
Forward problem
A problem where you are given an x-value (or z-value) and asked to find the corresponding probability (area under the curve). In simple terms, "I know the score, what proportion falls below it?"
Backward problem
A problem where you are given a probability (area) and asked to find the x-value or z-value that produces it. In simple terms, "I know the percentile, what score cuts it off?"
Percentile (quantile)
The value below which a given percentage of observations fall. The 50th percentile is the median. For a normal distribution, percentiles are found by working backward from the z-table.
Interquartile range (IQR)
Q₃ - Q₁, the width of the middle 50% of the distribution. For any normal distribution, IQR = 2 × 0.67 × σ ≈ 1.34σ (using the z-table approximation z₀.₂₅ ≈ -0.67). Think of it as a robust measure of spread that ignores the tails.
Empirical rule (68-95-99.7 rule)
For a normal distribution, approximately 68% of values lie within 1σ of the mean, 95% within 2σ, and 99.7% within 3σ. This is a rough guide, not exact, and crucially, "within 3 standard deviations" does not mean every single observation must fall there.
For any normal distribution, the mean, median, and mode are all equal to μ.
The PDF is symmetric about μ, so the 25th and 75th percentiles are equidistant from the mean.
If a continuous PDF is an even function over a symmetric interval [-c, c], the median (50th percentile) is 0. This follows because equal area accumulates on each side of 0.
Formula: z = (x - μ) / σ
Reverse formula: x = μ + z · σ
A linear transformation Y = a · X + b applied to X ~ N(μ, σ) gives Y ~ N(aμ + b, |a|σ).
E[Y] = a · E[X] + b
SD(Y) = |a| · SD(X)
Converting units (e.g. centimetres to inches) is a linear transformation, so z-scores and percentiles are preserved.
Standardise: compute z = (x - μ) / σ.
Look up P(Z < z) in the standard normal table.
For a range: P(a < X < b) = P(Z < z_b) - P(Z < z_a).
Example from the exam: X ~ N(118, 10). Find P(μ - 1.5σ < X < μ + 1.5σ).
This is P(-1.5 < Z < 1.5) = 2 · P(Z < 1.5) - 1 = 2(0.9332) - 1 = 0.8664.
Start with the desired area (probability).
Find the corresponding z-value from the table (reading it in reverse).
Transform back: x = μ + z · σ.
Example from the exam: find Q₁ and Q₃ for X ~ N(118, 10).
z₀.₂₅ = -0.67, z₀.₇₅ = +0.67
Q₁ = 118 + (-0.67)(10) = 111.3
Q₃ = 118 + (0.67)(10) = 124.7
IQR = 124.7 - 111.3 = 13.4
If E[X] = -2 and SD(X) = 5, and Y = 5X + 35:
E[Y] = 5(-2) + 35 = 25
SD(Y) = 5 · 5 = 25
So E[Y] = SD(Y) = 25. (This is a true statement about this specific transformation.)
Item | Formula |
|---|---|
Z-score | z = (x - μ) / σ |
Reverse z-score | x = μ + z · σ |
Probability within k SDs | P(μ - kσ < X < μ + kσ) = 2·Φ(k) - 1 |
IQR (normal) | IQR = 2 · z₀.₇₅ · σ ≈ 1.34σ |
Linear transformation mean | E[aX + b] = a·E[X] + b |
Linear transformation SD | SD(aX + b) = |
Broadcasting networks use the normal model of college basketball game lengths (as in the exam, μ = 118 min, σ = 10 min) to schedule programming and advertising slots. Knowing the IQR and outlier fences tells producers how much buffer time to build in before the next show.
Students sometimes believe that "within 3 standard deviations" means every individual observation must fall there. The empirical rule says approximately 99.7% do, but a normal distribution has infinite tails, so extreme values are rare but possible. The statement "every woman's height lies within 3 SDs" is false.
Confusing "forward" and "backward" is common. If you start with a number and want a probability, that is forward. If you start with a probability and want a number, that is backward.
Students assume IQR depends on the mean for a normal distribution. It does not. IQR = 1.34σ regardless of μ, because shifting the centre does not change the spread.
Forgetting that SD(aX + b) = |a| · SD(X) (the additive constant drops out) is a frequent calculation error.
⚠️ The exam tested whether you know that not every value must lie within 3 SDs (Q2.2, answer E is the false statement).
⚠️ Forward vs. backward terminology appeared as a true/false item (Q1.5, true).
⚠️ Computing IQR from a normal model was a 10-point free-response sub-part (Q3b). You must show the z-table lookup and the reverse transformation.
⚠️ Linear transformation properties (E[Y] and SD(Y)) were tested in Q1.4. Remember: constants add to the mean but not to the SD; multipliers scale both.
⚠️ The symmetry property of even PDFs implying median = 0 was tested in Q1.3 (true).
True or false: For X ~ N(50, 8), the median is 50.
Fill in the blank: The z-score for a value equal to the mean is ___.
True or false: Converting from kilograms to pounds changes a person's percentile in a normally distributed population.
Fill in the blank: For a standard normal, the 75th percentile is approximately z = ___.
True or false: P(-2 < Z < 2) ≈ 0.95.
Answers: 1. True. 2. Zero. 3. False (linear transformation preserves percentiles). 4. 0.67. 5. True (empirical rule).
Q: A distribution is N(200, 15). What is the probability a randomly selected value falls within 1.5 standard deviations of the mean?
A: P(-1.5 < Z < 1.5) = 2(0.9332) - 1 = 0.8664. The actual mean and SD do not matter here because "within 1.5 SDs" standardises directly.
Q: For X ~ N(80, 12), find the IQR.
A: z₀.₂₅ ≈ -0.67, z₀.₇₅ ≈ 0.67. Q₁ = 80 - 0.67(12) = 71.96. Q₃ = 80 + 0.67(12) = 88.04. IQR = 88.04 - 71.96 = 16.08.
Q: If E[X] = 10 and SD(X) = 3, and Y = 4X - 5, find E[Y] and SD(Y).
A: E[Y] = 4(10) - 5 = 35. SD(Y) = 4(3) = 12.
Q: Which is a backward normal problem: (a) "What percentage of light bulbs last more than 1,200 hours?" or (b) "How many hours must a bulb last to be in the top 5%?"
A: (b) is backward. You start with a probability (5%) and solve for the x-value.
Q: True or false: For a symmetric, even PDF on [-c, c], the 50th percentile must be 0.
A: True. Symmetry of the PDF about 0 means equal area on each side, so half the distribution falls below 0.
The z-score transformation connects directly to the topic of continuous PDFs and CDFs (covered in the piecewise PDF notes). Any continuous distribution can be standardised, but the normal is the only one where a single table suffices. The IQR and outlier fence calculations connect to descriptive statistics and boxplots (covered in the boxplot and outlier detection notes). Later in the course, the Central Limit Theorem will bring you back to the normal distribution even when the underlying data are not normal.
normal distribution, Gaussian, bell curve, z-score, standard score, standard normal, z-table, cumulative distribution, percentile, quantile, median, IQR, interquartile range, forward problem, backward problem, empirical rule, 68-95-99.7 rule, linear transformation, expected value, standard deviation, STAT 350 Purdue, Exam 1