Difficulty: Introductory | Prerequisites: None
This covers the foundational layer of the entire course: what statistics is, how data is collected and displayed, and how to summarise a dataset with a handful of numbers. Everything from confidence intervals to regression later in the course assumes you can already describe a distribution by its shape, centre and spread, compute means, medians, standard deviations and z-scores, and read a boxplot. If you are behind on anything here, the later chapters will feel harder than they need to.
Statistics is the science of collecting, organising, analysing and interpreting data. Descriptive statistics use graphs (histograms, boxplots) and numerical summaries (mean, median, standard deviation, quartiles) to describe a dataset. Inferential statistics use a sample to draw conclusions about a larger population.
Statistics
The science of collecting, organising, analysing and interpreting data. In simple terms, it is the toolkit for turning raw numbers into useful conclusions.
Descriptive statistics
Graphical and numerical methods used to describe, organise and summarise data. Think of it as the "here is what the data looks like" step.
Inferential statistics
Techniques for analysing a small, specific dataset in order to draw conclusions about a larger collection of data. In simple terms, this means using a sample to make claims about a whole population.
Population
The entire collection of individuals or objects to be studied.
Sample
A subset of the population, a small selection taken from the entire collection.
Variable
A characteristic of an individual or object in a population. Can be quantitative (numerical) or qualitative (categorical).
Frequency distribution
A table showing each category (label or class) of data alongside its count (frequency) or relative frequency (frequency divided by total count).
Histogram
A bar chart showing the distribution of a quantitative variable. For discrete data, each bar sits above a single value. For continuous data, the x-axis is divided into equal class intervals (number of classes is roughly the square root of the number of observations).
Sample mean (x̄)
The arithmetic average of the observations: x̄ = (1/n) Σxᵢ. The population equivalent is μ.
Sample median (x̃)
The middle value when observations are sorted smallest to largest. If n is odd, the median is the centre observation. If n is even, it is the average of the two centre observations.
Mode (M)
The value with the greatest frequency.
Variance (s²)
s² = [1/(n–1)] Σ(xᵢ – x̄)². Measures how spread out the data is. A variance of zero means all observations are the same. The population variance is σ².
Standard deviation (s)
The square root of the variance. Has the same units as the original observations, which makes it more interpretable than variance for comparing spread.
Quartiles (Q1, Q3)
Q1 is the median of the lower half of the data; Q3 is the median of the upper half. Together with the median they divide the sorted data into four equal parts.
Interquartile range (IQR)
IQR = Q3 – Q1. Measures the spread of the middle 50% of the data.
Outlier
An observation that falls outside the overall pattern. Mild outlier fences: Q1 – 1.5(IQR) and Q3 + 1.5(IQR). Extreme outlier fences: Q1 – 3(IQR) and Q3 + 3(IQR).
Five-number summary
Minimum, Q1, median, Q3, maximum. The backbone of a boxplot.
Boxplot
A graphical display built from the five-number summary. The box spans Q1 to Q3 with a line at the median. Whiskers extend to the most extreme non-outlier values. Mild outliers get closed circles; extreme outliers get open circles.
Empirical rule (68–95–99.7 rule)
For roughly normal data: about 68% of observations fall within 1 SD of the mean, 95% within 2 SDs, and 99.7% within 3 SDs.
z-score
zᵢ = (xᵢ – x̄) / s. A measure of relative standing that tells you how many standard deviations an observation is from the mean. The sum of all z-scores in a dataset is zero.
Statistics has four components: collection, organisation, analysis, interpretation
Two main branches:
Descriptive statistics: summarise and display data (graphs, tables, summary numbers)
Inferential statistics: use a sample to draw conclusions about a population
The inferential workflow follows a pattern: state a claim (status quo), run an experiment (check the claim), assess likelihood (is the result consistent with the claim?), draw a conclusion (reasonable or rare?)
A population is the whole group you care about. A sample is the subset you actually measure. A variable is the characteristic you record for each individual
Variables are classified by number (univariate, bivariate, multivariate) and by type (numerical or categorical)
When looking at any graph, describe three things: shape, centre, spread
Shapes of distributions:
Symmetric: mirror image about the centre
Positively (right) skewed: long tail to the right
Negatively (left) skewed: long tail to the left
Normal distribution: symmetric, bell-shaped. Heavy tails or light tails are departures from normality
Relative frequency = frequency / total count
Histogram construction for continuous data: divide the x-axis into roughly √(n) equal intervals, then draw bars whose heights are the frequency or relative frequency for each interval
An outlier is an individual observation that falls outside the overall pattern
Measures of central tendency: mean, median, mode. These tell you where the majority of the data sits
Sample notation uses lowercase Latin letters (x̄, s). Population notation uses Greek letters (μ, σ)
The mean is sensitive to outliers and skewness. The median is resistant to both
If data is roughly symmetric, report the mean and standard deviation
If data is skewed, report the median and IQR
Variance and standard deviation measure spread:
s² = 0 means every observation is identical
s is always ≥ 0 and shares the units of the original data
When n = 1, standard deviation is undefined
Quartile procedure: sort data, find the median, then find the median of the lower half (Q1) and upper half (Q3). Use the d = n/4 rule: if d is an integer, Q1 is the average of observations at positions d and d+1; otherwise Q1 is the observation at position ⌈d⌉
Outlier detection uses inner and outer fences:
Inner fences: Q1 – 1.5(IQR) and Q3 + 1.5(IQR). Points beyond these are mild outliers
Outer fences: Q1 – 3(IQR) and Q3 + 3(IQR). Points beyond these are extreme outliers
Boxplot construction: draw a box from Q1 to Q3, mark the median, extend whiskers to the most extreme non-outlier values, plot outliers as individual points
Empirical rule (for approximately normal data):
68% within 1 SD of the mean
95% within 2 SDs of the mean
99.7% within 3 SDs of the mean
z-score: zᵢ = (xᵢ – x̄) / s. Tells you how many standard deviations an observation is from the mean. The sum of all z-scores in a dataset is zero
z-scores are useful for comparing observations from different distributions (different units or scales)
\bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_is^2 = \frac{1}{n-1} \sum_{i=1}^{n} (x_i - \bar{x})^2s = \sqrt{s^2} = \sqrt{\frac{1}{n-1} \sum (x_i - \bar{x})^2}IQR = Q_3 - Q_1z_i = \frac{x_i - \bar{x}}{s}\text{Relative frequency} = \frac{\text{frequency}}{\text{total count}}\text{Number of classes} \approx \sqrt{n}Students often confuse the mean and the median, or assume they are interchangeable. They are not. The mean is pulled toward outliers and skewness; the median resists both. If the data is skewed, the median is the better measure of centre
Variance is not in the same units as the data. Standard deviation is. If you are asked to compare spread using the same units as the original observations, use standard deviation, not variance
The empirical rule only applies to distributions that are approximately normal (bell-shaped). Applying it to heavily skewed data will give misleading percentages
A z-score of zero does not mean the observation is unusual. It means the observation is exactly at the mean
Students sometimes forget that n – 1 (not n) is used in the denominator when computing sample variance. This is the degrees-of-freedom correction for estimating a population parameter from a sample
⚠️ Know when to use mean + SD vs. median + IQR. Symmetric data: mean and SD. Skewed data: median and IQR.
⚠️ Be able to construct and read a boxplot from a five-number summary, including identifying mild and extreme outliers.
⚠️ The empirical rule percentages (68, 95, 99.7) come up frequently. Know which SD boundaries they correspond to.
⚠️ Expect z-score calculations. Remember that the sum of all z-scores is zero and that z-scores let you compare across different distributions.
True or False: The median is always equal to the mean in a symmetric distribution. (True)
Fill in the blank: About ___% of data falls within two standard deviations of the mean in a normal distribution. (95%)
True or False: A z-score of –2 means the observation is two standard deviations below the mean. (True)
True or False: IQR measures the spread of all the data. (False – it measures the spread of the middle 50%)
Fill in the blank: The denominator in the sample variance formula is ___. (n – 1)
Q: A dataset has Q1 = 20, Q3 = 40. What are the inner fences for outlier detection?
A: IQR = 40 – 20 = 20. Lower inner fence = 20 – 1.5(20) = –10. Upper inner fence = 40 + 1.5(20) = 70.
Q: A student scores 78 on an exam where the class mean is 70 and the SD is 5. What is the student's z-score?
A: z = (78 – 70) / 5 = 1.6. The student scored 1.6 standard deviations above the mean.
Q: You have a heavily right-skewed dataset. Should you report the mean or the median as the measure of centre? Why?
A: The median, because it is resistant to the extreme values in the right tail that pull the mean upward.
Q: What does a z-score of 0 tell you about an observation?
A: The observation is exactly at the sample mean.
Q: In a boxplot, the whiskers extend to which values?
A: The most extreme observations that are not outliers (i.e. that fall within the inner fences).
This material connects directly to Ch. 6 (normal distribution), where z-scores become the primary tool for finding probabilities. The concept of spread (SD) reappears in Ch. 7 (sampling distributions) as the standard error of the sample mean. Boxplots and normality checks are used again in Ch. 11 (ANOVA) to verify assumptions.
Descriptive statistics, inferential statistics, population, sample, variable, frequency distribution, histogram, mean, median, mode, variance, standard deviation, quartiles, IQR, interquartile range, boxplot, box-and-whisker plot, five-number summary, empirical rule, 68-95-99.7 rule, z-score, standard score, outlier, inner fence, outer fence, skewness, symmetric distribution, central tendency, spread, STAT 35000, Purdue statistics