Probability Distributions: Continuous Random Variables – STAT, Introduction to Statistics – Study Notes
offline

Source: Probability Distributions, Key Concepts and Applications (Purdue University)

Tags: continuous distribution, normal distribution, bell curve, exponential distribution, continuous uniform, probability density function, area under curve, z-score, 68-95-99.7 rule, intro stats

Difficulty: Introductory to Intermediate Prerequisites: Part 1 (Discrete Random Variables) notes. Familiarity with integration concepts is helpful but not essential for an intro course; you mainly need to understand that area under a curve represents probability.


Big Picture

Where discrete distributions count outcomes, continuous distributions measure them. Height, time, temperature, and weight all live on a continuous scale, and the tools change accordingly: probabilities come from areas under curves rather than from summing individual values. This set of notes covers three continuous distributions (uniform, normal, exponential) plus the key formulas and approximation ideas that tie discrete and continuous worlds together. The normal distribution alone underpins most of inferential statistics, so this material carries forward into nearly every subsequent topic in the course.


TL;DR

Continuous random variables can take any value in an interval, so the probability of any single exact value is zero; instead you work with probabilities over ranges (areas under the density curve). The three distributions to know are the continuous uniform (equal likelihood across an interval), the normal (the bell curve that drives most of statistics), and the exponential (time between events in a Poisson process).


Key Terms

Probability density function (PDF)

For continuous variables, f(x) describes the relative likelihood of the variable near x. The probability that X falls between a and b is the integral of f(x) from a to b. In simple terms, probability = area under the curve between two points.

Area under the curve

The total area under any valid PDF equals 1. A slice of that area between two values gives you the probability of landing in that range.

Normal distribution (Gaussian distribution)

The symmetric, bell-shaped distribution defined by its mean μ and standard deviation σ. In simple terms, it is the "bell curve" that shows up whenever many small, independent effects add together.

Standard normal distribution

A normal distribution with μ = 0 and σ = 1. Any normal variable can be converted to the standard normal using a z-score.

Z-score (standard score)

z = (x − μ) / σ. It tells you how many standard deviations a value sits from the mean. Think of it as re-expressing any normal value on the same universal scale.

68-95-99.7 rule (empirical rule)

For any normal distribution, roughly 68% of values fall within 1σ of the mean, 95% within 2σ, and 99.7% within 3σ. This is the quick mental shortcut for "how unusual is this value?"

Exponential distribution

Models the time (or distance) between events in a Poisson process. If events happen at rate λ, the waiting time to the next event is Exponential(λ).

Memoryless property

A defining feature of the exponential distribution: the probability of waiting at least t more units does not depend on how long you have already waited. Think of it this way: the light bulb does not "remember" that it has been on for 500 hours.

Continuous uniform distribution

Every value in the interval [a, b] is equally likely. The density is a flat rectangle at height 1/(b − a).

Expected value (continuous)

E(X) = ∫ x f(x) dx over the entire range. The centre of mass of the density curve.

Variance (continuous)

Var(X) = ∫ (x − μ)² f(x) dx over the entire range. How spread out the density is.


Core Content

Continuous Uniform Distribution

  • Every value in the interval [a, b] is equally likely

  • PDF: f(x) = 1 / (b − a) for a ≤ x ≤ b, and 0 elsewhere

  • Mean: μ = (a + b) / 2

  • Variance: σ² = (b − a)² / 12

  • Probabilities are calculated by simple geometry: P(c ≤ X ≤ d) = (d − c) / (b − a)

  • Example: if a bus arrives at a stop at a uniformly random time between 0 and 30 minutes, the wait time follows Uniform(0, 30)

Normal Distribution

  • The most important distribution in statistics, defined entirely by two parameters: mean μ (centre) and standard deviation σ (spread)

  • PDF: f(x) = (1 / √(2πσ²)) · exp(−(x − μ)² / (2σ²))

    • You will rarely compute this by hand; the formula is mainly for reference

  • Key properties:

    • Symmetric around μ

    • The mean, median, and mode are all equal

    • Tails extend to −∞ and +∞ but probabilities drop off rapidly

  • The 68-95-99.7 rule:

    • About 68% of data within μ ± 1σ

    • About 95% within μ ± 2σ

    • About 99.7% within μ ± 3σ

  • Z-scores and the standard normal:

    • Convert any normal variable to standard normal: z = (x − μ) / σ

    • Use z-tables (or software) to find probabilities for any normal distribution

    • Example: if exam scores are Normal(μ = 70, σ = 10), a score of 85 has z = (85 − 70)/10 = 1.5, meaning it is 1.5 standard deviations above the mean

  • The normal distribution arises naturally whenever you sum many small, independent effects (this is the Central Limit Theorem at work)

Exponential Distribution

  • Models waiting times between independent events that occur at a constant average rate λ

  • PDF: f(x) = λ e⁻λˣ for x ≥ 0

  • Mean: μ = 1/λ (average waiting time is the reciprocal of the rate)

  • Variance: σ² = 1/λ²

  • The distribution is right-skewed: short waits are most common, long waits are possible but rare

  • Memoryless property: P(X > s + t | X > s) = P(X > t). No matter how long you have already waited, the probability of waiting t more units remains the same

  • Direct link to Poisson: if arrivals per hour are Poisson(λ), the time between arrivals is Exponential(λ)

  • Example: if an average of 3 earthquakes per year occur in a region, the time between consecutive earthquakes follows Exponential(3), with mean waiting time of 1/3 year ≈ 4 months

Key Formulas and Approximations

  • Total area under any PDF equals 1. This is the continuous analogue of "all probabilities sum to 1" for discrete distributions.

  • P(X = exact value) = 0 for continuous variables. Probability only comes from intervals: P(a ≤ X ≤ b) = ∫ₐᵇ f(x) dx. A consequence is that P(X ≤ b) = P(X < b) for continuous variables.

  • CDF: F(x) = P(X ≤ x) = ∫ from −∞ to x of f(t) dt. The CDF is the area under the curve up to x.

  • Expected value (continuous): E(X) = ∫ x f(x) dx

  • Variance (continuous): Var(X) = ∫ (x − μ)² f(x) dx

  • Normal approximation to the binomial: when n is large and p is not too close to 0 or 1 (common rule of thumb: np ≥ 5 and n(1−p) ≥ 5), the Binomial(n, p) distribution is well approximated by Normal(np, np(1−p)). This is the practical link between Part 1 and Part 2.


Formulas and Diagrams

Distribution

PDF

Mean (μ)

Variance (σ²)

Continuous uniform [a, b]

1/(b − a)

(a + b)/2

(b − a)²/12

Normal (μ, σ²)

(1/√(2πσ²)) e^(−(x−μ)²/(2σ²))

μ

σ²

Exponential (λ)

λ e⁻λˣ, x ≥ 0

1/λ

1/λ²

Z-score conversion: z = (x − μ) / σ

CDF (general): F(x) = P(X ≤ x)


Real-World Applications

  • Normal in manufacturing: component tolerances (bolt diameters, resistor values) follow normal distributions. Quality control limits (±3σ from target) come directly from the 68-95-99.7 rule.

  • Exponential in reliability engineering: the time until a component fails (assuming a constant hazard rate) is modelled as exponential. This is how engineers estimate warranty periods and maintenance schedules.


Common Misconceptions

  • Students often think the height of the PDF at a point is the probability at that point. It is not. For continuous variables, probability comes only from areas (integrals), not from single-point heights.

  • A common mistake is treating P(X = 5) as non-zero for a continuous distribution. For continuous variables, P(X = any exact value) = 0. Always work with intervals.

  • Students sometimes confuse the standard deviation σ with the variance σ². The normal distribution's formula uses σ², but the 68-95-99.7 rule uses σ. Keep track of which one a question is asking for.

  • The memoryless property of the exponential trips people up. It feels counterintuitive, but it is a mathematical consequence of the constant hazard rate, not a statement about physical wear and tear.


Why It Matters / Exam Flags

⚠️ The normal distribution will appear on virtually every exam in this course. Know how to convert to z-scores and read a z-table both ways (given z, find area; given area, find z).

⚠️ Expect a question testing whether you understand that P(X = exact value) = 0 for continuous distributions.

⚠️ The 68-95-99.7 rule is a frequent short-answer or multiple-choice target. Be able to apply it quickly without a table.

⚠️ Know the conditions for the normal approximation to the binomial (np ≥ 5 and n(1−p) ≥ 5) and be ready to use it in a problem.

⚠️ The exponential-Poisson link is a classic "connect-the-concepts" question: if you are given a Poisson rate, be prepared to switch to the exponential for waiting-time questions.


Quick Self-Test

  1. True or false: For a continuous random variable, P(X = 3) could be 0.15.

    False. P(X = exact value) = 0 for any continuous distribution.

  1. Fill in the blank: About ___% of values in a normal distribution fall within two standard deviations of the mean.

    95%.

  1. True or false: The exponential distribution is symmetric.

    False. It is right-skewed.

  1. Fill in the blank: If events arrive at rate λ = 5 per hour, the mean waiting time between events is ______ hours.

    1/5 = 0.2 hours (12 minutes).

  1. True or false: For a continuous uniform distribution on [0, 10], P(3 ≤ X ≤ 7) = 0.4.

    True. (7 − 3)/(10 − 0) = 0.4.


Practice Q&A

Q: Exam scores are normally distributed with μ = 72 and σ = 8. What proportion of students scored above 88?

A: z = (88 − 72)/8 = 2.0. From the z-table, P(Z ≤ 2.0) ≈ 0.9772. So P(X > 88) ≈ 1 − 0.9772 = 0.0228, or about 2.3%.

Q: A continuous uniform random variable ranges from 2 to 10. What is its mean and variance?

A: Mean = (2 + 10)/2 = 6. Variance = (10 − 2)²/12 = 64/12 ≈ 5.333.

Q: Customers arrive at a shop at an average rate of 6 per hour. What is the probability that the time between two consecutive arrivals exceeds 20 minutes?

A: Rate λ = 6 per hour. Convert 20 minutes to 1/3 hour. P(X > 1/3) = e⁻⁶ˣ⁽¹/³⁾ = e⁻² ≈ 0.1353, or about 13.5%.

Q: Why does P(X ≤ 5) = P(X < 5) for a continuous distribution, but not for a discrete distribution?

A: For continuous distributions, P(X = 5) = 0, so including or excluding the single point makes no difference. For discrete distributions, P(X = 5) can be positive, so the two expressions may differ.

Q: A Binomial(200, 0.4) distribution is to be approximated by a normal. What are the parameters of the approximating normal, and are the conditions met?

A: np = 80, n(1−p) = 120. Both exceed 5, so the conditions are met. The approximating normal has μ = 80 and σ² = np(1−p) = 48 (σ ≈ 6.93).


Connections to Other Topics

  • The normal distribution is the foundation of confidence intervals and hypothesis testing, which are typically the next major units in an introductory statistics course. Z-scores and z-tables will reappear constantly.

  • The Central Limit Theorem (CLT) explains why the normal distribution is so dominant: the sampling distribution of the sample mean approaches normality regardless of the population shape, given a large enough sample. This connects directly to the normal approximation to the binomial covered above.

  • The exponential distribution links back to Part 1's Poisson distribution: one describes counts per interval, the other describes time between events. Mastering both sides of that relationship is essential for applied probability problems.


Related Terms / Search Tags

continuous probability distribution, probability density function, PDF, area under curve, normal distribution, Gaussian, bell curve, z-score, standard score, standard normal, 68-95-99.7 rule, empirical rule, exponential distribution, memoryless property, continuous uniform, CDF, cumulative distribution function, expected value continuous, variance continuous, normal approximation to binomial, Central Limit Theorem, CLT, intro statistics, Purdue statistics, STAT, waiting time, hazard rate, reliability