Continuous Distributions and Sampling, STAT 35000 Ch. 6–7 – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Ch. 4–5 (probability, discrete random variables)

Big Picture

This section moves from discrete random variables (countable outcomes) to continuous ones (any value in an interval). The normal distribution is the star here: it underpins nearly every inferential method in the course. The Central Limit Theorem then explains why the normal distribution shows up everywhere, even when the population itself is not normal. Sampling distributions connect the individual-level data (Ch. 2–3) to the sample-level summaries that drive inference (Ch. 8–12). The study design material (producing data) rounds out the picture by explaining how the sample is actually collected.


TL;DR

Continuous random variables have probability density functions (pdfs) where probability equals area under the curve. The normal distribution is the most important continuous distribution in this course. The Central Limit Theorem says the sampling distribution of the sample mean is approximately normal for large n, regardless of the population shape. Good study design (randomisation, replication, control) is what makes inference valid.


Key Terms

Probability density function (pdf)

A smooth curve f(x) describing the distribution of a continuous random variable. Probability is the area under the curve between two points, not the height. f(x) ≥ 0 everywhere, and the total area under the curve is 1.

Cumulative distribution function (cdf)

F(x) = P(X ≤ x). The area under the pdf from –∞ to x.

Normal distribution

X ~ N(μ, σ²). Symmetric, bell-shaped, unimodal. Fully specified by its mean μ and standard deviation σ. Inflection points sit at one SD from the mean in each direction.

Standard normal distribution

Z ~ N(0, 1). The special case where μ = 0 and σ = 1. Any normal variable can be standardised to Z using z = (x – μ) / σ.

Uniform distribution

Probability is spread evenly over the interval (a, b). f(x) = 1/(b – a). Mean = (a + b)/2. SD = (b – a)/√12.

Exponential distribution

Models the waiting time until a specific event occurs. f(x) = λe⁻ᵞˣ for x ≥ 0. Mean = SD = 1/λ. Var = 1/λ².

Normal approximation to the binomial

When np ≥ 10 and n(1 – p) ≥ 10, a binomial can be approximated by a normal distribution. The continuity correction adds or subtracts 0.5 from the boundary.

Parameter vs. statistic

A parameter is a numerical measure of a population (Greek letters). A statistic is a quantity computed from a sample (Latin letters).

Sampling distribution

The probability distribution of a statistic across all possible samples of the same size from the same population.

Central Limit Theorem (CLT)

For large enough n, the sampling distribution of the sample mean is approximately N(μ, σ²/n), regardless of the shape of the population. In simple terms, averages of large samples are roughly normal even when individual observations are not.

Standard error

σ/√n. The standard deviation of the sampling distribution of the sample mean. Gets smaller as n increases.

Simple random sample (SRS)

A sample chosen so that every possible sample of size n has the same chance of being selected.

Stratified random sample

Divide the population into groups (strata) and take an SRS from each group.

Confounding

Occurs when two variables are associated such that their effects on the response cannot be separated. In simple terms, you cannot tell which variable is driving the result.

Lurking variable

A variable not among the explanatory or response variables that may influence both.

Simpson's Paradox

An association that holds within every subgroup can reverse direction when the subgroups are combined.


Core Content

Continuous Random Variables (Ch. 6.1)

  • Probability for a continuous RV is area under the pdf, not the height at a point. P(X = any specific value) = 0

  • Properties: f(x) ≥ 0 everywhere; total area = 1

  • Mean: E(X) = ∫ x f(x) dx. Variance: Var(X) = ∫ (x – μ)² f(x) dx

  • cdf: F(x) = P(X ≤ x) = ∫ from –∞ to x of f(s) ds

  • The 100p-th percentile is the value x where F(x) = p. The median is the 50th percentile

The Normal Distribution (Ch. 6.2, 6.5)

  • X ~ N(μ, σ²). The curve is symmetric and bell-shaped with inflection points at μ ± σ

  • Standardisation: z = (x – μ) / σ, converting any normal to Z ~ N(0, 1)

  • Z-table gives P(Z ≤ z). For P(Z > z), use 1 – P(Z ≤ z). For P(z₁ < Z < z₂), subtract: P(Z < z₂) – P(Z < z₁)

  • Procedure for normal problems: sketch and shade, standardise to Z, look up the table, compute, state conclusion in context

  • Normal approximation to binomial: when np ≥ 10 and n(1 – p) ≥ 10, use X ≈ N(np, np(1 – p)). Apply the continuity correction: P(X ≤ b) ≈ P(X < b + 0.5)

Checking Normality (Ch. 6.3)

  • Four methods: visual inspection of graphs, backward empirical rule, IQR/s ratio, normal probability (QQ) plot

  • QQ plot procedure: sort the data, record corresponding percentiles, find matching z-values, plot data vs. z-values. Points roughly on a straight line indicate normality

Other Continuous Distributions (Ch. 6.4)

  • Uniform: constant density over (a, b). E(X) = (a + b)/2. σ = (b – a)/√12

  • Exponential: models time until an event. f(x) = λe⁻ᵞˣ. E(X) = 1/λ. Var(X) = 1/λ². F(x) = 1 – e⁻ᵞˣ for x ≥ 0

  • Gamma: generalisation of the exponential. Used in theoretical statistics and actuarial science

  • Beta: defined on [0, 1]. Models proportions and percentages. The uniform is a special case

  • Weibull: used in lifetime modelling. Lognormal: log of a normal, used in products of distributions. Cauchy: symmetric with long, heavy tails

Sampling Distributions and the CLT (Ch. 7)

  • A parameter describes a population (μ, σ). A statistic describes a sample (x̄, s)

  • The sampling distribution of a statistic is the distribution of that statistic across all possible samples of the same size

  • Key results for the sample mean:

    • μ(x̄) = μ (the sampling distribution is centred at the population mean)

    • σ(x̄) = σ/√n (the spread shrinks as n grows)

    • If the population is normal, x̄ is exactly normal for any n

    • If the population is not normal, the CLT says x̄ is approximately normal for large n

  • Any linear combination of independent normal RVs is also normal

  • iid = independent and identically distributed

Producing Data (Ch. 1.3)

  • Experimental study: investigator applies treatments. Observational study: investigator observes without intervening

  • Three principles of good experiments: control, randomisation, replication

  • Bias is to accuracy as variability is to precision. Random sampling reduces bias; larger samples reduce variability

  • Key study designs:

    • Completely randomised: treatments assigned entirely by chance

    • Matched pair: each unit paired with a similar unit

    • Block design: units grouped into blocks of similar individuals, randomisation done within each block

  • Sampling methods: SRS, stratified random sample, convenience sample (biased)

  • Sources of bias: undercoverage, nonresponse, response bias

  • Confounding: when two variables’ effects cannot be separated. Lurking variables can cause confounding

  • Causation requires: strong association, consistency across studies, temporal ordering (cause before effect), plausibility

  • Simpson’s Paradox: a trend in subgroups can reverse when the groups are combined


Formulas Reference

f(x) = \frac{1}{\sigma\sqrt{2\pi}}\, e^{-\frac{(x-\mu)^2}{2\sigma^2}}
z = \frac{x - \mu}{\sigma} \quad \Longleftrightarrow \quad x = \mu + \sigma z
\sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}}
f(x) = \frac{1}{b-a}, \quad E(X) = \frac{a+b}{2}, \quad \sigma_X = \frac{b-a}{\sqrt{12}}
f(x) = \lambda e^{-\lambda x}, \quad E(X) = \frac{1}{\lambda}, \quad \text{Var}(X) = \frac{1}{\lambda^2}
\text{Continuity correction: } P(X \le b) \approx P\!\left(Z < \frac{b + 0.5 - np}{\sqrt{np(1-p)}}\right)

Common Misconceptions

  • Students often think P(X = 3) for a continuous RV gives a meaningful number. It does not. For continuous variables, P(X = any exact value) = 0. You must compute probability over an interval

  • The CLT does not say the population becomes normal. It says the sampling distribution of the sample mean becomes approximately normal. The population stays whatever shape it is

  • Confounding and lurking variables are not the same thing, though they are related. A lurking variable is hidden; confounding is what happens when a lurking (or other) variable is entangled with the explanatory variable

  • The continuity correction (adding or subtracting 0.5) applies only when using a continuous distribution to approximate a discrete one. Students often forget it or apply it in the wrong direction


Why It Matters / Exam Flags

⚠️ Z-table problems are a staple. Be comfortable going both directions: given x find probability, and given probability find x.

⚠️ Know the CLT conditions and what it guarantees. Be ready to state why x̄ is approximately normal for a given problem.

⚠️ Expect at least one question on study design: identify whether a study is experimental or observational, spot sources of bias, or explain why correlation does not imply causation.

⚠️ The normal approximation to the binomial requires np ≥ 10 and n(1 – p) ≥ 10. Check both before applying.


Quick Self-Test

  1. True or False: The standard error of the sample mean decreases as n increases. (True)

  1. Fill in the blank: For X ~ N(100, 25), the standard deviation is ___. (5, since σ² = 25)

  1. True or False: The CLT requires the population to be normally distributed. (False)

  1. True or False: In an observational study, the investigator assigns treatments to subjects. (False – that is an experiment)

  1. Fill in the blank: For a Uniform(0, 10) distribution, E(X) = ___. (5)


Practice Q&A

Q: A rash lasts a normally distributed number of days with mean 6 and SD 1.5. What symmetric interval about the mean captures 95% of durations?

A: 95% corresponds to z = ±1.96. Interval = 6 ± 1.96(1.5) = 6 ± 2.94 = (3.06, 8.94).

Q: A population has mean 50 and SD 12. For a sample of size 36, what is the standard error of x̄?

A: σ/√n = 12/√36 = 12/6 = 2.

Q: Why can you not conclude causation from an observational study?

A: Because treatments are not randomly assigned, lurking and confounding variables may explain the observed association.

Q: X ~ Exponential(λ = 0.5). What is P(X > 4)?

A: P(X > 4) = 1 – F(4) = 1 – (1 – e⁻²) = e⁻² ≈ 0.1353.


Connections to Other Topics

The normal distribution is the foundation for confidence intervals (Ch. 8) and hypothesis tests (Ch. 9). The sampling distribution of x̄ is the reason we can build z-tests and t-tests. Study design principles from Ch. 1.3 explain when inference results are trustworthy and when they are not. The exponential distribution connects back to the Poisson (Ch. 5.5): if events arrive at a Poisson rate λ, the time between events is Exponential(λ).


Related Terms / Search Tags

Continuous random variable, pdf, probability density function, cdf, cumulative distribution function, normal distribution, standard normal, z-score, z-table, uniform distribution, exponential distribution, gamma, beta, Weibull, lognormal, Cauchy, normal approximation, continuity correction, QQ plot, normality check, sampling distribution, Central Limit Theorem, CLT, standard error, parameter, statistic, SRS, simple random sample, stratified sample, experimental study, observational study, randomisation, replication, control, confounding, lurking variable, Simpson’s Paradox, STAT 35000, Purdue statistics