Difficulty: Introductory | Prerequisites: Chapter 3 Part 1 (mean, median, variance, standard deviation, IQR)
Part 1 of this chapter gave you the numbers. This part gives you the pictures and the rules that make those numbers useful. Boxplots visualise the five-number summary. Fences give you a formal way to flag outliers. The empirical rule tells you what percentage of data falls within one, two and three standard deviations of the mean in a normal distribution. Z-scores let you standardise any observation so you can compare values from different datasets on the same scale. All four tools come up repeatedly in later chapters.
Boxplots show a dataset's centre, spread and skewness in one picture. Fences detect outliers using the IQR. The empirical rule (68-95-99.7) applies to normal distributions. Z-scores convert any data point into a number of standard deviations from the mean.
Five-number summary
The set of five values that describe a dataset's distribution: minimum, Q1, median (Q2), Q3, maximum. Think of it as the skeleton of a boxplot.
Boxplot (box-and-whisker plot)
A graphical display built from the five-number summary. The box spans Q1 to Q3 (the IQR), with a line at the median. Whiskers extend to the smallest and largest non-outlier values. In simple terms, it is a picture that shows centre, spread and skewness in one glance.
Modified boxplot
A boxplot where the whiskers stop at the inner fences and outliers are plotted individually as dots. This is the version you will use most often.
Standard boxplot
A boxplot where the whiskers extend all the way to the minimum and maximum, hiding any outliers.
Inner fence (IF)
The boundary beyond which a data point is considered a mild outlier. Lower inner fence = Q1 − 1.5 × IQR. Upper inner fence = Q3 + 1.5 × IQR.
Outer fence (OF)
The boundary beyond which a data point is considered an extreme outlier. Lower outer fence = Q1 − 3 × IQR. Upper outer fence = Q3 + 3 × IQR.
Empirical rule (68-95-99.7 rule)
For any normal (bell-shaped) distribution: roughly 68% of data falls within 1 standard deviation of the mean, 95% within 2, and 99.7% within 3. In simple terms, almost all the data in a normal distribution lives within three standard deviations of the centre.
Z-score (standard score)
The number of standard deviations a data point sits above or below the mean. Formula: zᵢ = (xᵢ − x̄) / s. Think of it as a universal ruler that lets you compare apples to oranges across different datasets.
A boxplot compresses a dataset into five key values and draws them as a picture.
The five-number summary
Minimum
Q1 (25th percentile)
Median / Q2 (50th percentile)
Q3 (75th percentile)
Maximum
Reading a boxplot
The box itself spans from Q1 to Q3. Its width is the IQR.
The line inside the box marks the median.
Whiskers extend from each end of the box toward the minimum and maximum (in a standard boxplot) or toward the inner fences (in a modified boxplot).
What the shape tells you
Uniform distribution: the box is centred and the whiskers are roughly equal length.
Bell-shaped (symmetric): the median line sits near the middle of the box, and whiskers are roughly equal.
Right-skewed: the median line is closer to Q1, and the right whisker is longer.
Left-skewed: the median line is closer to Q3, and the left whisker is longer.
Boxplots are especially useful for comparing two or more groups side by side. You can instantly see which group has a higher centre, which has more spread, and whether either has outliers.
Outliers are data points that fall unusually far from the rest of the dataset. The fence method uses the IQR to define "unusually far" in a precise way.
Inner fences (mild outlier boundaries)
Lower inner fence: IF_L = Q1 − 1.5 × IQR
Upper inner fence: IF_H = Q3 + 1.5 × IQR
Any point between the inner and outer fences is a mild outlier.
Outer fences (extreme outlier boundaries)
Lower outer fence: OF_L = Q1 − 3 × IQR
Upper outer fence: OF_H = Q3 + 3 × IQR
Any point beyond an outer fence is an extreme outlier.
How fences relate to boxplots
In a modified boxplot, whiskers stop at the inner fences (or at the most extreme non-outlier value). Mild outliers are plotted as individual dots. Extreme outliers may be marked with a different symbol. In a standard boxplot, there are no fences shown and outliers are hidden inside the whiskers.
Why it matters
Outliers inflate the mean and the standard deviation. Identifying them lets you decide whether to investigate them further, keep them, or (with justification) remove them.
The empirical rule applies only to normal (bell-shaped, symmetric) distributions. It gives you a quick way to estimate how much data falls within a certain distance of the mean.
The three bands
About 68% of the data falls within 1 standard deviation of the mean: between μ − σ and μ + σ.
About 95% falls within 2 standard deviations: between μ − 2σ and μ + 2σ.
About 99.7% falls within 3 standard deviations: between μ − 3σ and μ + 3σ.
What this means in practice
If you know the mean and standard deviation of a normally distributed variable, you can immediately say where most of the data lives.
Finding a data point more than 2 standard deviations from the mean is uncommon (only about 5% of the time). Finding one more than 3 standard deviations away is very rare (about 0.3%).
This is why values far from the mean are suspicious and worth investigating as possible outliers.
When it does not apply
The rule is only an approximation, and it works well only for distributions that are roughly bell-shaped. For skewed or multimodal distributions, these percentages will not hold.
Standardisation converts raw data values into z-scores so that different datasets become comparable on a single scale.
The formula
zᵢ = (xᵢ − x̄) / s
Subtract the mean from the observation, then divide by the standard deviation.
Interpreting a z-score
A z-score of 0 means the observation equals the mean.
A positive z-score means the observation is above the mean. For example, z = 1.5 means the value sits 1.5 standard deviations above the mean.
A negative z-score means the observation is below the mean.
Using the empirical rule as a guide: z-scores beyond ±2 are unusual, and beyond ±3 are very rare in normal data.
Why standardise?
Suppose you scored 80 on a maths exam (mean 70, s = 5) and 85 on an English exam (mean 80, s = 10). Raw scores suggest English was better, but the z-scores tell a different story: z_maths = (80 − 70)/5 = 2.0, z_english = (85 − 80)/10 = 0.5. You performed much better relative to the class in maths. This is the power of standardisation: it strips away the original units and lets you compare on equal footing.
Measure | Formula |
|---|---|
Lower inner fence | Q1 − 1.5 × IQR |
Upper inner fence | Q3 + 1.5 × IQR |
Lower outer fence | Q1 − 3 × IQR |
Upper outer fence | Q3 + 3 × IQR |
Z-score | zᵢ = (xᵢ − x̄) / s |
Z-scores are used in finance to measure how unusual a stock's return is compared to its historical average. In medicine, growth charts for children are built on percentiles and the empirical rule: a paediatrician flags a child whose weight z-score drops below −2. Boxplots are a standard tool in quality control dashboards for comparing batch-to-batch variation in manufacturing.
Students assume the empirical rule works for any distribution. It does not. It is reliable only for distributions that are approximately normal (bell-shaped).
Students mix up inner and outer fences. Inner fences use 1.5 × IQR, outer fences use 3 × IQR. A point between the inner and outer fence is mild; beyond the outer fence is extreme.
Students think a z-score tells you whether a value is "good" or "bad." It only tells you how far the value is from the mean in standard-deviation units. The interpretation depends on context.
Students forget that a modified boxplot and a standard boxplot look different. The modified version shows outliers as individual points; the standard version hides them.
⚠️ Be ready to compute inner and outer fences from Q1, Q3 and IQR, then identify which data points are outliers.
⚠️ Know the empirical rule percentages cold: 68%, 95%, 99.7%. These come up in both multiple-choice and short-answer formats.
⚠️ Expect a question asking you to compute a z-score and interpret it in context (e.g., "is this exam score unusually high?").
⚠️ Be able to identify skewness from a boxplot: which way the median shifts inside the box, and which whisker is longer.
True or false: The empirical rule applies to all distributions. False. Only to approximately normal (bell-shaped) distributions.
Fill in the blank: A z-score of 0 means the observation equals the ______. Mean.
True or false: In a modified boxplot, outliers are shown as individual points. True.
Fill in the blank: To find the upper inner fence, compute Q3 + ______ × IQR. 1.5.
True or false: If a boxplot's median line is closer to Q1 and the right whisker is longer, the data are left-skewed. False. That pattern indicates right skew.
Q: A dataset has Q1 = 10 and Q3 = 30. Compute the inner and outer fences.
A: IQR = 30 − 10 = 20. Lower inner fence = 10 − 1.5(20) = −20. Upper inner fence = 30 + 1.5(20) = 60. Lower outer fence = 10 − 3(20) = −50. Upper outer fence = 30 + 3(20) = 90.
Q: Using the fences from the previous question, classify a data point with value 75.
A: 75 is above the upper inner fence (60) but below the upper outer fence (90), so it is a mild outlier.
Q: A normally distributed variable has mean 50 and standard deviation 8. Between what two values do approximately 95% of the data fall?
A: 50 − 2(8) = 34 and 50 + 2(8) = 66. About 95% of the data falls between 34 and 66.
Q: You scored 72 on a test with mean 64 and standard deviation 5. What is your z-score, and is it unusual?
A: z = (72 − 64) / 5 = 1.6. A z-score of 1.6 is above average but not unusual by the empirical rule (you would need z beyond ±2 to be in the outer 5%).
Q: Looking at a boxplot, how can you tell whether data are right-skewed?
A: The median line inside the box will be closer to Q1 (the left edge of the box), and the right whisker will be noticeably longer than the left whisker.
Z-scores are the gateway to the standard normal distribution (Chapter 4 or 5, depending on your syllabus), where you will use z-tables to find exact probabilities instead of the rough 68-95-99.7 approximations. The fence-based outlier detection here connects to later discussions of influential points in regression (Chapter 12+). Boxplots reappear whenever you compare groups, including in ANOVA and non-parametric tests.
Boxplot, box-and-whisker plot, five-number summary, modified boxplot, standard boxplot, inner fence, outer fence, mild outlier, extreme outlier, 1.5 IQR rule, 3 IQR rule, empirical rule, 68-95-99.7 rule, normal distribution, bell curve, z-score, standard score, standardisation, standardization, z-value, number of standard deviations, skewness from boxplot, right skewed, left skewed, Purdue STAT, intro statistics chapter 3