Source: Purdue University, Introduction to Statistics
Tags: mean, median, mode, sample variance, standard deviation, quartiles, Q1, Q3, IQR, interquartile range, five-number summary, boxplot, modified boxplot, outliers, inner fences, measures of centre, measures of spread
Difficulty: Introductory Prerequisites: Chapter 1–2 concepts (variable types, distribution shapes). You need to know the difference between categorical and quantitative data and what "skewed" means.
Once you know what type of data you have, the next step is to summarise it numerically. Chapter 3 gives you the toolkit: measures of centre (where the data sits) and measures of spread (how much the data varies). The choice between mean/standard deviation and median/IQR depends on the shape of the distribution and the presence of outliers. This chapter also introduces the five-number summary and boxplots, which are visual and numerical tools for comparing distributions and spotting outliers. Everything here feeds directly into probability and inference later in the course.
The mean, median, and mode describe the centre of a dataset. Variance and standard deviation measure spread around the mean, while the IQR measures spread around the median. Use mean and standard deviation for roughly symmetric data without outliers; use median and IQR when the data are skewed or contain outliers.
Mean (x̄)
The arithmetic average of all observations: x̄ = (1/n) Σ xᵢ.
Think of it as the balance point of the data. It is sensitive to extreme values.
Median (x̃)
The middle value when the data are sorted from smallest to largest. If n is even, it is the average of the two middle values.
In simple terms, half the data fall below and half above.
Mode
The value that occurs most frequently. No calculation is required.
Think of it as the most popular value. A dataset can have more than one mode.
Sample variance (s²)
The average of the squared deviations from the mean, divided by n − 1: s² = Σ(xᵢ − x̄)² / (n − 1)
In simple terms, it measures how spread out the data are. The n − 1 divisor (rather than n) corrects for the fact that you are estimating from a sample.
Sample standard deviation (s)
The square root of the sample variance: s = √(s²).
Think of it as the "typical" distance of a data point from the mean, measured in the original units.
First quartile (Q₁)
The value below which roughly 25% of the data fall. Found using the depth d₁ = n/4 on the sorted data.
Third quartile (Q₃)
The value below which roughly 75% of the data fall. Found using the depth d₃ = 3n/4 on the sorted data.
Interquartile range (IQR)
IQR = Q₃ − Q₁. It captures the spread of the middle 50% of the data.
In simple terms, the IQR tells you how wide the "middle chunk" of your data is, ignoring the tails.
Five-number summary
Minimum, Q₁, Median, Q₃, Maximum. This set of five values gives a concise picture of the distribution.
Modified boxplot
A boxplot that marks outliers individually and extends the whiskers only to the most extreme non-outlier values (the inner fences).
Outlier (by the fence rule)
An observation is an outlier if it falls below Q₁ − 1.5 × IQR or above Q₃ + 1.5 × IQR.
Mean (x̄ = (1/n) Σ xᵢ): add all values, divide by the count. Pulled towards extreme values.
Median (x̃): sort the data, find the middle position. Resistant to outliers.
Mode: the most frequently occurring value. Useful for categorical data as well.
For a symmetric distribution, mean ≈ median. For a right-skewed distribution, the mean is typically greater than the median. For a left-skewed distribution, the mean is typically less than the median.
Range: maximum − minimum. Simple but heavily influenced by outliers.
Sample variance (s²): uses the squared deviations from the mean. The computational formula is often more convenient:
s² = [Σ xᵢ² − (1/n)(Σ xᵢ)²] / (n − 1)
The exam will provide the summations if needed.
Sample standard deviation (s): s = √(s²). Same units as the data.
IQR: Q₃ − Q₁. Resistant to outliers, so it pairs naturally with the median.
Sort the data from smallest to largest.
Compute d₁ = n/4 and d₃ = 3n/4.
If dₖ is a whole number, the quartile is the average of the observation at position dₖ and the observation at position dₖ + 1.
If dₖ is not a whole number, round up to the next integer and take that observation.
Lower inner fence: Q₁ − 1.5 × IQR
Upper inner fence: Q₃ + 1.5 × IQR
Any observation below the lower fence or above the upper fence is an outlier.
You do not need to distinguish between mild and extreme outliers for this exam.
The five-number summary is: Minimum, Q₁, Median, Q₃, Maximum.
A modified boxplot draws the box from Q₁ to Q₃ with a line at the median, extends whiskers to the most extreme non-outlier values, and plots outliers as individual points.
Interpreting a boxplot: describe shape (symmetric if the median is centred in the box and whiskers are similar length; skewed if not) and note any outliers.
Side-by-side boxplots let you compare distributions across groups by placing boxplots on the same scale.
Symmetric, no outliers: use mean and standard deviation.
Skewed or has outliers: use median and IQR.
The exam will ask you to justify your choice based on the distribution's shape.
Mean: x̄ = (1/n) Σᵢ xᵢ
Sample variance: s² = Σ(xᵢ − x̄)² / (n − 1) = [Σ xᵢ² − (1/n)(Σ xᵢ)²] / (n − 1)
Sample standard deviation: s = √(s²)
Quartile depths: d₁ = n/4, d₃ = 3n/4
IQR: IQR = Q₃ − Q₁
Outlier fences: Lower fence = Q₁ − 1.5 × IQR Upper fence = Q₃ + 1.5 × IQR
Salary data is a classic case where the choice of centre matters. The mean salary at a company might be £85,000, but if one executive earns £2 million, the median might be £55,000. Reporting the mean without context gives a misleading picture. The median and IQR are the standard tools when data are skewed, which is why government agencies typically report "median household income" rather than "mean household income."
Students often divide by n instead of n − 1 when computing sample variance. The n − 1 denominator (called Bessel's correction) is required because you are estimating from a sample.
Students sometimes think the median is always one of the data values. When n is even, the median is the average of the two middle values and may not appear in the dataset at all.
Forgetting to sort the data before finding the median or quartiles is a frequent calculation error.
Students sometimes assume "outlier" means "mistake." Outliers are simply values far from the bulk of the data. They might be errors, but they might also be legitimate and interesting.
⚠️ You will be asked to calculate the mean, median, variance, standard deviation, quartiles, and IQR by hand. Show sorted data and quartile depth calculations for full marks.
⚠️ Expect a question where you must decide whether to use mean/standard deviation or median/IQR, and justify why based on the distribution.
⚠️ Outlier detection using the 1.5 × IQR fence rule will be tested. You must show the fence calculation and compare each data point against it.
⚠️ Interpreting and comparing boxplots (including side-by-side) is a likely question.
True or False: The sample variance uses n in the denominator.
Fill in the blank: The IQR is calculated as __________.
True or False: For right-skewed data, the mean is typically less than the median.
Fill in the blank: An observation is an outlier if it is greater than __________ or less than __________.
True or False: The five-number summary includes the mean.
Answers: 1. False (n − 1). 2. Q₃ − Q₁. 3. False (mean is typically greater). 4. Q₃ + 1.5 × IQR; Q₁ − 1.5 × IQR. 5. False (it includes min, Q₁, median, Q₃, max).
Q: Given the sorted data 2, 5, 7, 8, 12, 15, 18, 20, and n = 8, find Q₁ and Q₃.
A: d₁ = 8/4 = 2 (whole number), so Q₁ = average of 2nd and 3rd values = (5 + 7)/2 = 6. d₃ = 3(8)/4 = 6 (whole number), so Q₃ = average of 6th and 7th values = (15 + 18)/2 = 16.5.
Q: Using Q₁ = 6 and Q₃ = 16.5 from above, determine the outlier fences and state whether any data points are outliers.
A: IQR = 16.5 − 6 = 10.5. Lower fence = 6 − 1.5(10.5) = −9.75. Upper fence = 16.5 + 1.5(10.5) = 32.25. All data values (2 through 20) fall within these fences, so there are no outliers.
Q: A dataset has a distribution that is strongly right-skewed. Should you report the mean or the median as the measure of centre? Why?
A: Report the median. The mean is pulled towards the long right tail and overstates where most of the data sit. The median is resistant to skewness and outliers.
Q: What are the components of a five-number summary?
A: Minimum, Q₁, Median, Q₃, Maximum.
Q: In a modified boxplot, what do the whiskers extend to?
A: The whiskers extend to the most extreme data values that are not outliers (i.e. the values closest to but still within the inner fences). Any values beyond the fences are plotted as individual points.
The mean and variance formulas here are the sample versions; Chapter 5 introduces the population-level expected value E(X) and variance Var(X) for random variables, using the same logic but with probability weights.
The standard deviation reappears in Chapter 6 when you standardise to z-scores: z = (x − μ) / σ.
Choosing between mean/SD and median/IQR connects to interpreting distribution shapes from Chapter 2.
mean, sample mean, x-bar, median, mode, sample variance, sample standard deviation, s-squared, quartile, first quartile, third quartile, Q1, Q3, interquartile range, IQR, five-number summary, boxplot, modified boxplot, whiskers, inner fence, outlier detection, 1.5 IQR rule, measures of centre, measures of spread, resistant measure, Bessel's correction, STAT 101, intro to statistics, Purdue