Numerical Measures: Central Tendency and Variability – STAT, Ch. 3 – Study Notes
offline

Difficulty: Introductory | Prerequisites: Basic algebra, Chapter 2 (data types and distributions)

This is where statistics moves from pictures to numbers. Chapters 1 and 2 covered how to collect and display data; this chapter gives you the actual quantities that summarise a dataset: where its centre sits and how spread out the values are. These measures are the vocabulary for every later topic in the course, from probability to hypothesis testing. If you skipped the earlier chapters, you can still follow this material, but you should know the difference between a population and a sample.


TL;DR

Central tendency (mean, median) tells you where the middle of your data is. Variability (range, variance, standard deviation, IQR) tells you how spread out the values are around that centre. Together they give you a two-number snapshot of any dataset.

Key Terms

Sample mean (x̄)

The arithmetic average of all observed values in a sample. Add them up, divide by how many there are. Think of it as the balance point of the data.

Sample median (x̃)

The middle value when the data are sorted from smallest to largest. If the sample size is odd, it is the central number. If even, it is the average of the two central numbers. In simple terms, this is the value that splits the dataset in half.

Sample size (n)

The total number of observations in a sample.

Sample range

The difference between the largest and smallest values: Max minus Min. Think of it as the total span of the data from end to end.

Sample variance (s²)

The average of the squared deviations from the mean, using n minus 1 in the denominator. It measures how far, on average, each data point sits from the mean. In simple terms, a larger variance means data points are more spread out.

Sample standard deviation (s)

The square root of the sample variance. It is in the same units as the original data, which makes it more interpretable than variance. Think of it as a typical distance from the mean.

Percentile

The value below which a given percentage of the data falls when the data are divided into 100 equal parts.

Quartiles (Q1, Q2, Q3)

Three values that split a sorted dataset into four equal parts. Q1 marks the lowest 25%, Q2 is the median (lowest 50%), and Q3 marks the lowest 75%.

Interquartile range (IQR)

Q3 minus Q1. It captures the spread of the middle 50% of the data. In simple terms, this is the range that ignores the extreme values at both ends.

Core Content: Measures of Central Tendency

Central tendency answers one question: where is the centre of this distribution? The two main measures are the mean and the median.

Notation convention

  • Sample statistics use Latin/Greek letters: x̄ for the sample mean, s for the sample standard deviation.

  • Population parameters use Greek letters: μ for the population mean, σ for the population standard deviation.

  • Keep these straight. Exams love asking which symbol goes with which.

Sample mean (x̄)

  • Formula: x̄ = (1/n) × Σxᵢ

  • Add every observation, then divide by the number of observations.

  • The mean uses every data point, which makes it sensitive to extreme values (outliers). One very large or very small value can pull the mean significantly in that direction.

Sample median (x̃)

  • Sort the data from smallest to largest.

  • If n is odd: the median is the single middle value.

  • If n is even: the median is the average of the two middle values.

  • The median is resistant to outliers. It does not budge much when an extreme value enters the dataset.

When to prefer which

  • Symmetric data: mean and median will be close together, and either is a good summary.

  • Skewed data or data with outliers: the median is usually the better measure of centre, because the mean gets dragged toward the tail.

Core Content: Measures of Variability

Central tendency alone is not enough. Two datasets can have the same mean but look completely different if one is tightly clustered and the other is widely spread. Variability measures tell you how spread out the data are.

Sample range

  • Range = Max − Min

  • Simple to compute but uses only two data points, so a single outlier can inflate it dramatically.

Sample variance (s²)

  • Formula: s² = [1 / (n − 1)] × Σ(xᵢ − x̄)²

  • Measures the average squared distance of each observation from the sample mean.

  • The denominator is n − 1 (not n) because this corrects for the bias that comes from estimating the population variance with a sample. You will sometimes hear this called Bessel's correction.

  • If the variance is close to 0, the data points are tightly clustered around the mean. If it is large, they are spread out.

Sample standard deviation (s)

  • Formula: s = √s² = √{ [1 / (n − 1)] × Σ(xᵢ − x̄)² }

  • It is simply the square root of the variance.

  • Because it is in the same units as the data (not squared units), it is much easier to interpret. You can think of it as a typical deviation from the mean.

  • Outliers inflate the standard deviation, just as they inflate the mean.

Interquartile range (IQR)

  • IQR = Q3 − Q1

  • It measures the spread of only the middle 50% of the data.

  • Because it ignores the top 25% and bottom 25%, it is resistant to outliers, just like the median.

  • The IQR is the variability measure paired with the median, the same way the standard deviation is paired with the mean.

Choosing the right pair

  • Symmetric data, no outliers: report the mean and standard deviation.

  • Skewed data or outliers present: report the median and IQR.

Formulas at a Glance

Measure

Formula

Sample mean

x̄ = (1/n) × Σxᵢ

Sample range

Max − Min

Sample variance

s² = [1/(n−1)] × Σ(xᵢ − x̄)²

Sample std dev

s = √s²

IQR

Q3 − Q1

Real-World Applications

The mean and standard deviation are the backbone of quality control in manufacturing: a factory monitors the mean diameter of bolts and flags any batch whose standard deviation crosses a threshold. The median household income is the standard measure in economics precisely because a few billionaires would skew the mean upwards and misrepresent most people's earnings.


Common Misconceptions

  • Students often think the mean is always the best measure of centre. It is not. When outliers or skewness are present, the median is more representative.

  • Students confuse variance and standard deviation. Variance is in squared units; standard deviation is in the original units. If your data are in kilograms, variance is in kg², which is not directly interpretable.

  • Students forget that the sample variance formula divides by n − 1, not n. On a calculator-heavy exam, using n instead of n − 1 gives the wrong answer.

  • Students assume the range is a reliable measure of spread. It uses only two data points, so a single outlier can make it misleading.


Why It Matters / Exam Flags

⚠️ Know which symbol is which: x̄ and s for samples, μ and σ for populations. This is tested constantly.

⚠️ Be ready to compute a sample variance by hand: square each deviation, sum them, divide by n − 1.

⚠️ Expect a question asking which summary (mean vs. median, std dev vs. IQR) is appropriate for a skewed dataset.

⚠️ Percentile and quartile definitions are common multiple-choice targets.

Quick Self-Test

  1. True or false: The median is always equal to the mean. False. They are equal only when the distribution is perfectly symmetric.

  1. Fill in the blank: The sample variance divides by ______, not n. n − 1.

  1. True or false: The IQR is affected by outliers. False. It measures only the middle 50%.

  1. Fill in the blank: Q2 is another name for the ______. Median.

  1. True or false: Standard deviation is measured in the same units as the original data. True.

Practice Q&A

Q: A dataset has values 3, 7, 7, 12, 21. What is the sample mean?

A: (3 + 7 + 7 + 12 + 21) / 5 = 50 / 5 = 10.

Q: For the same dataset (3, 7, 7, 12, 21), what is the median?

A: The data are already sorted. With n = 5 (odd), the median is the 3rd value: 7.

Q: Compute the sample variance for the dataset 4, 8, 6, 2, 10.

A: Mean = 30/5 = 6. Deviations: −2, 2, 0, −4, 4. Squared: 4, 4, 0, 16, 16. Sum = 40. Variance = 40 / (5 − 1) = 10.

Q: A dataset has Q1 = 20 and Q3 = 45. What is the IQR?

A: IQR = Q3 − Q1 = 45 − 20 = 25.

Q: Your dataset is heavily right-skewed. Should you report the mean or the median as the measure of centre? Why?

A: The median, because it is resistant to the extreme values in the right tail that would inflate the mean.

Connections to Other Topics

This material feeds directly into Chapter 3 Part 2 (boxplots, outlier detection, the empirical rule, and z-scores), which uses the mean, standard deviation and IQR as inputs. Later in the course, the sample mean and standard deviation become the building blocks for confidence intervals and hypothesis tests. The distinction between sample statistics and population parameters introduced here is the foundation of all inferential statistics.


Related Terms / Search Tags

Sample mean, arithmetic mean, average, x-bar, sample median, central tendency, measures of centre, variability, spread, dispersion, sample range, sample variance, sample standard deviation, s-squared, Bessel's correction, n minus 1, percentile, quartile, Q1, Q2, Q3, interquartile range, IQR, resistant measure, skewness, outlier effect on mean, Purdue STAT, intro statistics chapter 3