Difficulty: Intermediate | Prerequisites: Ch. 4–5 (probability, discrete random variables)
This section moves from discrete random variables (countable outcomes) to continuous ones (any value in an interval). The normal distribution is the star here: it underpins nearly every inferential method in the course. The Central Limit Theorem then explains why the normal distribution shows up everywhere, even when the population itself is not normal. Sampling distributions connect the individual-level data (Ch. 2–3) to the sample-level summaries that drive inference (Ch. 8–12). The study design material (producing data) rounds out the picture by explaining how the sample is actually collected.
Continuous random variables have probability density functions (pdfs) where probability equals area under the curve. The normal distribution is the most important continuous distribution in this course. The Central Limit Theorem says the sampling distribution of the sample mean is approximately normal for large n, regardless of the population shape. Good study design (randomisation, replication, control) is what makes inference valid.
Probability density function (pdf)
A smooth curve f(x) describing the distribution of a continuous random variable. Probability is the area under the curve between two points, not the height. f(x) ≥ 0 everywhere, and the total area under the curve is 1.
Cumulative distribution function (cdf)
F(x) = P(X ≤ x). The area under the pdf from –∞ to x.
Normal distribution
X ~ N(μ, σ²). Symmetric, bell-shaped, unimodal. Fully specified by its mean μ and standard deviation σ. Inflection points sit at one SD from the mean in each direction.
Standard normal distribution
Z ~ N(0, 1). The special case where μ = 0 and σ = 1. Any normal variable can be standardised to Z using z = (x – μ) / σ.
Uniform distribution
Probability is spread evenly over the interval (a, b). f(x) = 1/(b – a). Mean = (a + b)/2. SD = (b – a)/√12.
Exponential distribution
Models the waiting time until a specific event occurs. f(x) = λe⁻ᵞˣ for x ≥ 0. Mean = SD = 1/λ. Var = 1/λ².
Normal approximation to the binomial
When np ≥ 10 and n(1 – p) ≥ 10, a binomial can be approximated by a normal distribution. The continuity correction adds or subtracts 0.5 from the boundary.
Parameter vs. statistic
A parameter is a numerical measure of a population (Greek letters). A statistic is a quantity computed from a sample (Latin letters).
Sampling distribution
The probability distribution of a statistic across all possible samples of the same size from the same population.
Central Limit Theorem (CLT)
For large enough n, the sampling distribution of the sample mean is approximately N(μ, σ²/n), regardless of the shape of the population. In simple terms, averages of large samples are roughly normal even when individual observations are not.
Standard error
σ/√n. The standard deviation of the sampling distribution of the sample mean. Gets smaller as n increases.
Simple random sample (SRS)
A sample chosen so that every possible sample of size n has the same chance of being selected.
Stratified random sample
Divide the population into groups (strata) and take an SRS from each group.
Confounding
Occurs when two variables are associated such that their effects on the response cannot be separated. In simple terms, you cannot tell which variable is driving the result.
Lurking variable
A variable not among the explanatory or response variables that may influence both.
Simpson's Paradox
An association that holds within every subgroup can reverse direction when the subgroups are combined.
Probability for a continuous RV is area under the pdf, not the height at a point. P(X = any specific value) = 0
Properties: f(x) ≥ 0 everywhere; total area = 1
Mean: E(X) = ∫ x f(x) dx. Variance: Var(X) = ∫ (x – μ)² f(x) dx
cdf: F(x) = P(X ≤ x) = ∫ from –∞ to x of f(s) ds
The 100p-th percentile is the value x where F(x) = p. The median is the 50th percentile
X ~ N(μ, σ²). The curve is symmetric and bell-shaped with inflection points at μ ± σ
Standardisation: z = (x – μ) / σ, converting any normal to Z ~ N(0, 1)
Z-table gives P(Z ≤ z). For P(Z > z), use 1 – P(Z ≤ z). For P(z₁ < Z < z₂), subtract: P(Z < z₂) – P(Z < z₁)
Procedure for normal problems: sketch and shade, standardise to Z, look up the table, compute, state conclusion in context
Normal approximation to binomial: when np ≥ 10 and n(1 – p) ≥ 10, use X ≈ N(np, np(1 – p)). Apply the continuity correction: P(X ≤ b) ≈ P(X < b + 0.5)
Four methods: visual inspection of graphs, backward empirical rule, IQR/s ratio, normal probability (QQ) plot
QQ plot procedure: sort the data, record corresponding percentiles, find matching z-values, plot data vs. z-values. Points roughly on a straight line indicate normality
Uniform: constant density over (a, b). E(X) = (a + b)/2. σ = (b – a)/√12
Exponential: models time until an event. f(x) = λe⁻ᵞˣ. E(X) = 1/λ. Var(X) = 1/λ². F(x) = 1 – e⁻ᵞˣ for x ≥ 0
Gamma: generalisation of the exponential. Used in theoretical statistics and actuarial science
Beta: defined on [0, 1]. Models proportions and percentages. The uniform is a special case
Weibull: used in lifetime modelling. Lognormal: log of a normal, used in products of distributions. Cauchy: symmetric with long, heavy tails
A parameter describes a population (μ, σ). A statistic describes a sample (x̄, s)
The sampling distribution of a statistic is the distribution of that statistic across all possible samples of the same size
Key results for the sample mean:
μ(x̄) = μ (the sampling distribution is centred at the population mean)
σ(x̄) = σ/√n (the spread shrinks as n grows)
If the population is normal, x̄ is exactly normal for any n
If the population is not normal, the CLT says x̄ is approximately normal for large n
Any linear combination of independent normal RVs is also normal
iid = independent and identically distributed
Experimental study: investigator applies treatments. Observational study: investigator observes without intervening
Three principles of good experiments: control, randomisation, replication
Bias is to accuracy as variability is to precision. Random sampling reduces bias; larger samples reduce variability
Key study designs:
Completely randomised: treatments assigned entirely by chance
Matched pair: each unit paired with a similar unit
Block design: units grouped into blocks of similar individuals, randomisation done within each block
Sampling methods: SRS, stratified random sample, convenience sample (biased)
Sources of bias: undercoverage, nonresponse, response bias
Confounding: when two variables’ effects cannot be separated. Lurking variables can cause confounding
Causation requires: strong association, consistency across studies, temporal ordering (cause before effect), plausibility
Simpson’s Paradox: a trend in subgroups can reverse when the groups are combined
f(x) = \frac{1}{\sigma\sqrt{2\pi}}\, e^{-\frac{(x-\mu)^2}{2\sigma^2}}z = \frac{x - \mu}{\sigma} \quad \Longleftrightarrow \quad x = \mu + \sigma z\sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}}f(x) = \frac{1}{b-a}, \quad E(X) = \frac{a+b}{2}, \quad \sigma_X = \frac{b-a}{\sqrt{12}}f(x) = \lambda e^{-\lambda x}, \quad E(X) = \frac{1}{\lambda}, \quad \text{Var}(X) = \frac{1}{\lambda^2}\text{Continuity correction: } P(X \le b) \approx P\!\left(Z < \frac{b + 0.5 - np}{\sqrt{np(1-p)}}\right)Students often think P(X = 3) for a continuous RV gives a meaningful number. It does not. For continuous variables, P(X = any exact value) = 0. You must compute probability over an interval
The CLT does not say the population becomes normal. It says the sampling distribution of the sample mean becomes approximately normal. The population stays whatever shape it is
Confounding and lurking variables are not the same thing, though they are related. A lurking variable is hidden; confounding is what happens when a lurking (or other) variable is entangled with the explanatory variable
The continuity correction (adding or subtracting 0.5) applies only when using a continuous distribution to approximate a discrete one. Students often forget it or apply it in the wrong direction
⚠️ Z-table problems are a staple. Be comfortable going both directions: given x find probability, and given probability find x.
⚠️ Know the CLT conditions and what it guarantees. Be ready to state why x̄ is approximately normal for a given problem.
⚠️ Expect at least one question on study design: identify whether a study is experimental or observational, spot sources of bias, or explain why correlation does not imply causation.
⚠️ The normal approximation to the binomial requires np ≥ 10 and n(1 – p) ≥ 10. Check both before applying.
True or False: The standard error of the sample mean decreases as n increases. (True)
Fill in the blank: For X ~ N(100, 25), the standard deviation is ___. (5, since σ² = 25)
True or False: The CLT requires the population to be normally distributed. (False)
True or False: In an observational study, the investigator assigns treatments to subjects. (False – that is an experiment)
Fill in the blank: For a Uniform(0, 10) distribution, E(X) = ___. (5)
Q: A rash lasts a normally distributed number of days with mean 6 and SD 1.5. What symmetric interval about the mean captures 95% of durations?
A: 95% corresponds to z = ±1.96. Interval = 6 ± 1.96(1.5) = 6 ± 2.94 = (3.06, 8.94).
Q: A population has mean 50 and SD 12. For a sample of size 36, what is the standard error of x̄?
A: σ/√n = 12/√36 = 12/6 = 2.
Q: Why can you not conclude causation from an observational study?
A: Because treatments are not randomly assigned, lurking and confounding variables may explain the observed association.
Q: X ~ Exponential(λ = 0.5). What is P(X > 4)?
A: P(X > 4) = 1 – F(4) = 1 – (1 – e⁻²) = e⁻² ≈ 0.1353.
The normal distribution is the foundation for confidence intervals (Ch. 8) and hypothesis tests (Ch. 9). The sampling distribution of x̄ is the reason we can build z-tests and t-tests. Study design principles from Ch. 1.3 explain when inference results are trustworthy and when they are not. The exponential distribution connects back to the Poisson (Ch. 5.5): if events arrive at a Poisson rate λ, the time between events is Exponential(λ).
Continuous random variable, pdf, probability density function, cdf, cumulative distribution function, normal distribution, standard normal, z-score, z-table, uniform distribution, exponential distribution, gamma, beta, Weibull, lognormal, Cauchy, normal approximation, continuity correction, QQ plot, normality check, sampling distribution, Central Limit Theorem, CLT, standard error, parameter, statistic, SRS, simple random sample, stratified sample, experimental study, observational study, randomisation, replication, control, confounding, lurking variable, Simpson’s Paradox, STAT 35000, Purdue statistics