Difficulty: Beginner | Prerequisites: Basic algebra
Statistics is the science of learning from data. This first block of the course covers the foundations: what data looks like, how to summarise it visually and numerically, and how to describe its shape and spread. Everything you learn here feeds directly into probability (Chapter 4 onward) and inference (Chapters 7-12). If you are comfortable with means, medians, standard deviations and z-scores, the rest of the course is substantially easier.
Statistics splits into descriptive (summarising what you have) and inferential (drawing conclusions about a larger population from a sample). You summarise categorical data with bar graphs and frequency tables, and numerical data with histograms, boxplots, means, medians, standard deviations and z-scores. The choice between mean/standard deviation and median/IQR depends on whether the distribution is symmetric or skewed.
Statistics
The science of collecting, organising, analysing and interpreting data. Think of it as the toolkit for turning raw numbers into useful conclusions.
Descriptive statistics
Graphical and numerical methods used to describe, organise and summarise data. In simple terms, this is everything you do to understand the data you already have, before making any wider claims.
Inferential statistics
Techniques used to analyse a small, specific data set in order to draw conclusions about a larger collection of data. In simple terms, this is how you go from "what my sample shows" to "what is probably true for the whole population."
Population
The entire collection of individuals or objects you want to study.
Sample
A subset of the population. Think of it as the slice of reality you actually measure.
Variable
A characteristic of an individual or object in the population. It is a column in a data table. Variables are either quantitative (numerical) or qualitative (categorical).
Discrete data
Numerical data that is finite or countable. You do not measure it, you count it (e.g. number of apps on a phone).
Continuous data
Numerical data that falls on an interval and is measured rather than counted (e.g. time spent playing games).
Frequency distribution
A listing of possible values and how often each occurs. Frequency simply means the count.
Relative frequency
Frequency divided by the total count. It expresses the proportion as a decimal or percentage.
Histogram
A graph for quantitative variables where bars touch (no gaps unless a class has zero observations). The number of classes (bins) is approximately the square root of the number of observations for small data sets, or 20-30 classes for large ones.
Bar graph
A graph for qualitative (categorical) variables. Bars do not touch.
Unimodal / bimodal / multimodal
A distribution with one peak, two peaks, or more than two peaks respectively. Each peak often represents a different sub-population.
Symmetric distribution
A distribution where the left and right sides are roughly mirror images.
Positively skewed (right-skewed)
The tail stretches to the right. The mean is pulled above the median.
Negatively skewed (left-skewed)
The tail stretches to the left. The mean is pulled below the median.
Outlier
A data point that is far from the rest of the data. In simple terms, a value that does not fit the general pattern. Never delete outliers without justification.
Sample mean (x-bar)
The arithmetic average: the sum of all observations divided by n. It is sensitive (not resistant) to outliers.
Population mean (μ)
The mean of the entire population. Greek letters indicate population parameters, Latin letters indicate sample statistics.
Sample median (x-tilde)
The middle value of sorted data. It is resistant to outliers, making it a better measure of centre for skewed distributions.
Mode
The value with the greatest frequency. It is insensitive to outliers.
Sample variance (s²)
The average squared deviation from the mean, using n-1 in the denominator.
Sample standard deviation (s)
The square root of the variance. It is in the same units as the original data and is not resistant to outliers.
Interquartile range (IQR)
Q3 minus Q1. It captures the central 50% of the data and is resistant to outliers.
Quartiles (Q1, Q2, Q3)
Q1 is the 25th percentile, Q2 is the median (50th percentile), Q3 is the 75th percentile.
Five-number summary
Minimum, Q1, median, Q3, maximum. These five values define a boxplot.
Boxplot (modified)
A graphical representation of the five-number summary with outliers explicitly shown as individual points. Whiskers extend from Q1 to the smallest non-outlier and from Q3 to the largest non-outlier.
z-score
A standardised value showing how many standard deviations a data point is from the mean: z = (xᵢ - x-bar) / s. Positive z-scores sit above the mean, negative below.
Empirical rule (68-95-99.7 rule)
For normal distributions: 68% of data falls within 1 standard deviation of the mean, 95% within 2, and 99.7% within 3.
Statistics has three main branches: data collection, descriptive statistics and inferential statistics
The logic of inference follows a four-step chain: make a claim (status quo), run an experiment to check the claim, evaluate the likelihood of the result under the claim, then draw a conclusion (the outcome is either reasonable or rare)
Probability is the bridge between descriptive and inferential statistics, it quantifies "how likely"
Variables split into quantitative (numerical) and qualitative (categorical)
Solution trail for statistics problems:
Find keywords in the problem statement
Translate those words into statistical terminology
Determine which concepts apply
Develop a strategy
Solve
Data sets are classified by the number of variables observed: univariate (one variable), bivariate (two), multivariate (three or more)
Histograms are for quantitative variables (bars touch). Bar graphs are for qualitative variables (bars do not touch)
When building a histogram, use approximately √n classes for small data sets or 20-30 classes for large ones
Too many bins makes it hard to spot peaks
Too few bins makes it hard to see the distribution shape
When examining any graph, look for three things: shape, centre and variability (spread)
Shapes of distributions: symmetric, positively skewed (right tail), negatively skewed (left tail), heavy-tailed, light-tailed
Never delete outliers without a substantive reason
Measures of centre:
Mean (x-bar): sum of observations divided by n. Not resistant to outliers
Median (x-tilde): middle of sorted data. Resistant to outliers
Mode: most frequent value. Insensitive to outliers
If the distribution is symmetric, use the mean. If highly skewed, use the median
Measures of spread:
Range: maximum minus minimum. Heavily influenced by outliers
Variance (s²) and standard deviation (s): measure how far observations are from the mean. Neither is resistant to outliers. s² = 0 means every observation is identical
IQR (Q3 - Q1): the spread of the central 50%. Resistant to outliers
Choosing centre and spread together:
Symmetric data: use mean and standard deviation
Skewed data: use median and IQR
Outlier detection (1.5 IQR rule):
Inner fences (mild outliers): Q1 - 1.5(IQR) and Q3 + 1.5(IQR)
Outer fences (extreme outliers): Q1 - 3(IQR) and Q3 + 3(IQR)
A calculated outlier that sits very close to the whisker may not be a real outlier in practice
z-scores:
z = (xᵢ - x-bar) / s
The mean of all z-scores is always 0, the standard deviation is always 1
The z-score distribution has the same shape as the original distribution
The sum of the squared z-scores equals the number of observations
Quantity | Formula |
|---|---|
Sample mean | x-bar = (Σ xᵢ) / n |
Sample variance | s² = Σ(xᵢ - x-bar)² / (n - 1) |
Sample standard deviation | s = √(s²) |
IQR | Q3 - Q1 |
Inner fences | Q1 - 1.5(IQR), Q3 + 1.5(IQR) |
Outer fences | Q1 - 3(IQR), Q3 + 3(IQR) |
z-score | z = (xᵢ - x-bar) / s |
mean(VariableName) – sample mean
median(VariableName) – sample median
var(VariableName) – sample variance
sd(VariableName) – sample standard deviation
quantile(VariableName) – five-number summary
as.numeric() – convert quantile output to a number before arithmetic
max(VariableName) - min(VariableName) – range
The mean and standard deviation are used constantly in quality control to determine whether a manufacturing process is producing items within acceptable limits. When a process produces a skewed distribution (for example, insurance claim amounts, where most are small but a few are very large), the median and IQR give a more honest picture of "typical" than the mean does.
Students often think the mean and median are interchangeable. They are not. In a skewed distribution the mean is pulled toward the tail, making it a poor indicator of the typical value.
Students sometimes delete outliers to "clean" the data. Never remove an outlier unless you have a substantive reason (such as a known data entry error). Outliers may carry real information.
Students confuse histograms and bar graphs. Histograms are for quantitative data (bars touch); bar graphs are for categorical data (bars do not touch).
Students occasionally compute IQR as Q1 - Q3 instead of Q3 - Q1. The IQR is always a positive number.
⚠️ Know when to use mean/standard deviation vs. median/IQR. This decision depends on the shape of the distribution.
⚠️ Be able to calculate and interpret z-scores. A z-score of 2.0 means the observation is 2 standard deviations above the mean.
⚠️ Know the empirical rule (68-95-99.7) and when it applies (normal distributions only).
⚠️ Be able to identify outliers using the 1.5 IQR rule and state whether each is mild or extreme.
⚠️ Remember: population parameters use Greek letters (μ, σ), sample statistics use Latin letters (x-bar, s).
True or false: The median is always equal to the mean in any data set. False. They are equal only when the distribution is perfectly symmetric.
Fill in the blank: If a data point has a z-score of -1.5, it lies ___ standard deviations ___ the mean. 1.5 standard deviations below the mean.
True or false: The IQR is resistant to outliers. True.
Fill in the blank: In the empirical rule, ___% of data falls within two standard deviations of the mean. 95%.
True or false: A histogram is appropriate for displaying categorical data. False. Histograms are for quantitative data; bar graphs are for categorical data.
Q: A data set has Q1 = 20, Q3 = 40, and a data point at 75. Is 75 an outlier? If so, is it mild or extreme?
A: IQR = 40 - 20 = 20. Upper inner fence = 40 + 1.5(20) = 70. Upper outer fence = 40 + 3(20) = 100. Since 75 is above the inner fence (70) but below the outer fence (100), it is a mild outlier.
Q: A symmetric distribution has a mean of 50 and a standard deviation of 5. Using the empirical rule, what range captures approximately 95% of the data?
A: 50 - 2(5) to 50 + 2(5) = 40 to 60.
Q: A student scores 78 on an exam where the mean is 70 and the standard deviation is 4. What is the student's z-score, and what does it mean?
A: z = (78 - 70) / 4 = 2.0. The student scored 2 standard deviations above the mean.
Q: You have a heavily right-skewed income distribution. Which measures of centre and spread should you report?
A: Median and IQR, because both are resistant to the influence of the extreme high-income outliers that cause the skew.
This material connects directly to Chapter 6 (normal distributions), where the empirical rule and z-scores become essential. The z-score reappears as the test statistic in hypothesis testing (Chapter 9). Understanding variance here is foundational for ANOVA (Chapter 11), which compares variances between groups.
descriptive statistics, inferential statistics, population, sample, variable, quantitative, qualitative, categorical, numerical, discrete, continuous, frequency distribution, relative frequency, histogram, bar graph, unimodal, bimodal, skewness, symmetric, outlier, mean, median, mode, variance, standard deviation, IQR, interquartile range, quartile, five-number summary, boxplot, z-score, empirical rule, 68-95-99.7, STAT 350, Purdue, introductory statistics