Normal Distribution and Applications, STAT 350 Exam 1 – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Z-scores, basic probability, standard normal table


Big Picture

The normal distribution is the single most important continuous distribution in introductory statistics. It appears in modelling real-world measurements (heights, speeds, errors), in the Central Limit Theorem, and as an approximation to other distributions. STAT 350 Exam 1 tests your ability to standardise, look up probabilities, compute conditional probabilities, use the empirical rule, find percentiles, and combine normal random variables. If you can handle these tasks fluently, you are in good shape for roughly a third of the exam.


TL;DR

A normal random variable X ~ N(μ, σ) is fully described by its mean μ and standard deviation σ. To find probabilities, convert to the standard normal Z = (X - μ)/σ and use the Z-table. The distribution is symmetric about μ, so P(-z < Z < 0) = P(0 < Z < z). The empirical rule (68-95-99.7) gives quick approximations for intervals within 1, 2, or 3 standard deviations of the mean.


Key Terms

Normal distribution, N(μ, σ)

A continuous distribution with a symmetric, bell-shaped density curve centred at μ. The parameter σ controls the spread.

  • Mean: E[X] = μ

  • Variance: Var(X) = σ²

  • Support: all real numbers (-∞, +∞)

  • μ can be any real number; σ must be positive. There is no requirement that μ/σ > 1.

Standard normal distribution, N(0, 1)

The special case with μ = 0 and σ = 1. Denoted Z. The Z-table gives cumulative probabilities P(Z ≤ z).

Z-score (standardisation)

Z = (X - μ) / σ. Converts any normal random variable to the standard normal scale.

In simple terms, a Z-score tells you how many standard deviations a value is from the mean.

Empirical rule (68-95-99.7 rule)

For a normal distribution:

  • About 68% of values fall within 1σ of the mean

  • About 95% fall within 2σ of the mean

  • About 99.7% fall within 3σ of the mean

This applies specifically to normal distributions. P(μ - 2σ < X < μ + 2σ) ≈ 0.9544 to be precise, which is approximately 0.95.

Percentile / quantile

The value below which a given percentage of observations fall. The 97th percentile is the value exceeded by only 3% of the distribution. To find it, look up the Z-value corresponding to cumulative probability 0.97, then convert back: X = μ + Zσ.

Cumulative distribution function (CDF)

F(x) = P(X ≤ x). For the standard normal, the Z-table gives this directly.


Core Content

Standardising and Using the Z-Table

  • Given X ~ N(μ, σ), to find P(X > a):

    • Compute z = (a - μ) / σ

    • Look up P(Z ≤ z) in the table

    • P(X > a) = 1 - P(Z ≤ z)

Worked example: X ~ N(2, 0.19). Find P(X > 2.3).

  • z = (2.3 - 2) / 0.19 = 0.3 / 0.19 ≈ 1.5789 ≈ 1.58

  • P(Z ≤ 1.58) = 0.9429 (from the table: row 1.5, column 0.08)

  • P(X > 2.3) = 1 - 0.9429 = 0.0571

Symmetry of the Standard Normal

The standard normal distribution is symmetric about 0. This means:

  • P(Z < -z) = P(Z > z) for any z > 0

  • P(-z < Z < 0) = P(0 < Z < z) for any z > 0

These two regions have exactly the same area under the curve. The correct symbol connecting them is =.

The Empirical Rule and the 2σ Interval

For X ~ N(μ, σ²):

  • P(μ - 2σ < X < μ + 2σ) ≈ 0.9544

More precisely, from the Z-table: P(-2 < Z < 2) = P(Z < 2) - P(Z < -2) = 0.9772 - 0.0228 = 0.9544.

This is approximately 0.95, so the statement "P(E[X] - 2√Var(X) < X < E[X] + 2√Var(X)) ≈ 0.95" is true for any normal random variable, regardless of μ and σ. (Note: 2√Var(X) = 2σ.)

Conditional Probability with the Normal Distribution

P(X > a | X > b) = P(X > a) / P(X > b) when a > b.

This works because {X > a} is a subset of {X > b} when a > b, so P(X > a ∩ X > b) = P(X > a).

Worked example: X ~ N(2, 0.19). Find P(X > 2.3 | X > 2).

  • P(X > 2.3) = 0.0571 (computed above)

  • P(X > 2) = 0.5 (since 2 is the mean of a symmetric distribution)

  • P(X > 2.3 | X > 2) = 0.0571 / 0.5 = 0.1142

Combining Normal Distribution with Binomial

When a normal probability gives you p, and you repeat the experiment n times independently, the count of "successes" follows a binomial distribution.

Worked example: P(one zombie > 2.3 mph) = 0.0571. If 10 zombies are selected independently, what is P(at least one > 2.3)?

  • Let Y = number with speed > 2.3. Then Y ~ Bin(10, 0.0571).

  • P(Y ≥ 1) = 1 - P(Y = 0) = 1 - (1 - 0.0571)^10 = 1 - 0.9429^10

  • 0.9429^10 ≈ 0.5765 (compute step by step or use a calculator)

  • P(Y ≥ 1) ≈ 1 - 0.5565 ≈ 0.4435

(Note: the exact value depends on rounding. With p = 0.0571, (0.9429)^10 ≈ 0.5565, giving P ≈ 0.4435.)

Expected Value of Sums of Normal Random Variables

If X1 and X2 are independent, each ~ N(μ, σ):

  • E[X1 + X2] = 2μ

  • After t hours of travel at speed Xi, distance = Xi * t.

  • Expected total distance for two zombies over 3 hours: E[3X1 + 3X2] = 3E[X1] + 3E[X2] = 3μ + 3μ = 6μ.

Worked example: μ = 2 mph, 2 zombies, 3 hours: expected combined distance = 6 * 2 = 12 miles.

Finding Percentiles (Inverse Normal)

To find the value x such that P(X ≤ x) = p:

  1. Find the Z-value z such that P(Z ≤ z) = p from the table.

  1. Convert back: x = μ + zσ.

Worked example: "Top 3%" means P(X > x) = 0.03, so P(X ≤ x) = 0.97.

  • From the Z-table, P(Z ≤ 1.88) = 0.9699 ≈ 0.97, so z ≈ 1.88.

  • x = μ + zσ = 2 + 1.88 * 0.19 = 2 + 0.3572 = 2.3572 mph.


Formulas / Diagrams

Standardisation: Z = (X - μ) / σ

Inverse standardisation (finding X from Z): X = μ + Zσ

Conditional probability (subset case): P(X > a | X > b) = P(X > a) / P(X > b), for a > b

Sum of independent normals: If X1 ~ N(μ1, σ1) and X2 ~ N(μ2, σ2) are independent, then X1 + X2 ~ N(μ1 + μ2, √(σ1² + σ2²))

Linear transformation: If X ~ N(μ, σ), then aX + b ~ N(aμ + b, |a|σ)


Common Misconceptions

  • Students sometimes think μ/σ must be greater than 1 for a normal distribution. There is no such restriction. μ can be any real number and σ any positive real number.

  • A common error is computing P(X > a | X > b) as P(X > a) - P(X > b). The correct formula uses division, not subtraction.

  • Students often forget that P(X > μ) = 0.5 for any normal distribution, because of symmetry. This simplifies many conditional probability problems.

  • When finding percentiles, students sometimes use the wrong tail. "Top 3%" means the area to the right is 0.03, so the area to the left is 0.97.


Why It Matters / Exam Flags

⚠️ Z-table lookups with standardisation appear on nearly every exam. Practise until the conversion is automatic.

⚠️ Conditional probability with normal distributions (P(X > a | X > b)) is a multi-part free response favourite.

⚠️ The "at least one" binomial follow-up to a normal probability calculation is a common exam pattern.

⚠️ Percentile questions require working backwards from the Z-table. Know how to read the table in reverse.

⚠️ The symmetry property P(-z < Z < 0) = P(0 < Z < z) appears as a quick multiple choice question.


Quick Self-Test

  1. True or false: P(μ - 2σ < X < μ + 2σ) ≈ 0.95 for any normal random variable X.

  1. Fill in the blank: P(-z < Z < 0) ______ P(0 < Z < z) for any z > 0, where Z is standard normal.

  1. True or false: If X ~ N(5, 2), then P(X > 5) = 0.5.

  1. Fill in the blank: To find the 90th percentile of X ~ N(μ, σ), find z such that P(Z ≤ z) = ______, then compute x = ______.

  1. True or false: P(X > 3 | X > 1) = P(X > 3) - P(X > 1).


Practice Q&A

Q: X ~ N(10, 3). Find P(X > 14.5).

A: z = (14.5 - 10)/3 = 1.50. P(Z ≤ 1.50) = 0.9332. P(X > 14.5) = 1 - 0.9332 = 0.0668.

Q: X ~ N(10, 3). Find P(X > 14.5 | X > 10).

A: P(X > 14.5 | X > 10) = P(X > 14.5)/P(X > 10) = 0.0668/0.5 = 0.1336.

Q: The top 5% of a N(100, 15) distribution starts at what value?

A: P(Z ≤ z) = 0.95, so z ≈ 1.645. x = 100 + 1.645 * 15 = 100 + 24.675 = 124.675.

Q: Each of 8 independent items has a 0.10 probability of being defective. What is the probability that at least one is defective?

A: Y ~ Bin(8, 0.10). P(Y ≥ 1) = 1 - (0.90)^8 = 1 - 0.4305 = 0.5695.


Connections to Other Topics

The normal distribution connects to the binomial distribution: when you compute P(X > a) for a normal variable and then count how many of n independent items satisfy that condition, the count follows Bin(n, p) where p = P(X > a). This two-step pattern is a staple of STAT 350.

It connects to the exponential distribution through contrast: the exponential is right-skewed and models waiting times, while the normal is symmetric and models measurements. Different shapes, different applications.


Related Terms / Search Tags

normal distribution, Gaussian, bell curve, standard normal, Z-score, Z-table, standardisation, empirical rule, 68-95-99.7, percentile, quantile, inverse normal, conditional probability, CDF, cumulative distribution function, symmetry, μ, σ, mean, standard deviation, variance, STAT 350, Purdue, continuous distribution