Difficulty: Beginner | Prerequisites: None. This is foundational material for the entire STAT 101 course.
This material forms the bedrock of introductory statistics. Before you can run any hypothesis test or build any model, you need to know what kind of data you are looking at, how to summarise it, and how to spot when something looks unusual. Everything here, from classifying variables to reading a boxplot to calculating a z-score, feeds directly into the probability and inference topics that follow. If you are joining the course late or picking it up mid-semester, start here.
Variables are either categorical (groups) or numeric (quantities), and the type determines which graphs and summary statistics to use. Centre is captured by mean, median or mode; spread by range, variance, standard deviation or IQR. Z-scores let you compare values across different distributions by measuring how many standard deviations a point sits from the mean.
Categorical variable (qualitative variable)
A variable that places observations into groups or categories, such as gender, colour or brand. You cannot do arithmetic on these values. In simple terms, it answers "which group?" not "how much?"
Numeric variable (quantitative variable)
A variable that measures a quantity. Can be discrete (countable, whole numbers, e.g. number of children) or continuous (measurable on a scale, e.g. height, temperature). Think of it as anything you could meaningfully average.
Discrete variable
A numeric variable that takes countable values, often whole numbers. In simple terms, you can list all the possible values.
Continuous variable
A numeric variable that can take any value within a range, including fractions and decimals. Think of it as measurements on a smooth scale.
Histogram
A graph that shows the frequency distribution of numeric data by grouping values into bins (classes) and displaying the count or proportion in each bin as a bar. Unlike a bar graph, the bars touch because the data is continuous.
Mean (arithmetic average)
The sum of all values divided by the number of values. Sensitive to outliers. In simple terms, this is what most people think of when they hear "average."
Median
The middle value when data are arranged in order. If there is an even number of observations, it is the average of the two middle values. Think of it as the value that splits the dataset in half. More resistant to outliers than the mean.
Mode
The most frequently occurring value in a dataset. Useful for categorical data or for spotting peaks in numeric distributions. A dataset can be unimodal, bimodal or multimodal.
Range
Maximum value minus minimum value. The simplest measure of spread, but easily distorted by a single outlier.
Variance
The average of the squared deviations from the mean. Measures overall spread, but is expressed in squared units, which can be hard to interpret directly.
Standard deviation
The square root of the variance. Expressed in the same units as the original data, making it more interpretable than variance. Think of it as the typical distance of a data point from the mean.
Interquartile range (IQR)
Q3 minus Q1. Captures the spread of the middle 50% of the data. Resistant to outliers, making it a better choice than range or standard deviation for skewed distributions.
Five-number summary
Minimum, Q1, Median, Q3, Maximum. These five values give a quick portrait of a distribution's centre, spread and range.
Boxplot (box-and-whisker plot)
A graph built from the five-number summary. The box spans Q1 to Q3, a line marks the median, whiskers extend to the most extreme points within 1.5 x IQR, and anything beyond is plotted individually as a potential outlier.
Outlier
A data point that falls far from the general pattern. In a boxplot, any point beyond the whiskers (more than 1.5 x IQR from Q1 or Q3). Outliers can heavily influence the mean and standard deviation.
Z-score (standard score)
The number of standard deviations a data point sits from the mean: z = (x - mean) / standard deviation. Allows comparison across distributions with different units or scales. In simple terms, it answers "how unusual is this value?"
Skewness
A measure of asymmetry in a distribution. Right-skewed (positive skew) means the tail stretches to the right; left-skewed (negative skew) means the tail stretches to the left. Think of it as which direction the long tail points.
Every statistical analysis begins by identifying what type of data you have. The type determines which summary statistics are valid and which graphs to use.
Categorical (qualitative): places observations into groups. Examples: gender, eye colour, brand preference. Cannot be averaged or summed meaningfully.
Visualised with pie charts (proportions) or bar graphs (counts or percentages).
Numeric (quantitative): measures a quantity on a number line.
Discrete: countable, often whole numbers. Example: number of siblings, number of defective items in a batch.
Continuous: can take any value in a range. Example: weight, time, temperature.
Visualised with histograms, boxplots and scatterplots.
Before running any analysis, check that your variable is appropriate for the research question. Ask: is the variable measured correctly? Is it relevant? Does its type match the statistical method you plan to use?
Good graphs reveal patterns that raw numbers hide. When reading any graph, look for four things: shape, centre, spread and outliers.
Shape: is the distribution symmetric, left-skewed or right-skewed? Is it unimodal, bimodal or multimodal?
Centre: where does the "typical" value fall?
Spread: how much variability is there from the smallest to the largest values?
Outliers: are there data points sitting far from the rest?
Matching the graph to the variable type is essential.
Categorical data: use pie charts to show proportions or bar graphs to show counts/percentages. Bars do not touch because categories are distinct.
Numeric data: use histograms to show the frequency distribution across bins. Bars touch because the data is continuous. Choosing the right number of bins matters: too few oversimplify, too many create noise. Sturges' rule and the square-root method are common guides for selecting bin count.
Histograms are the primary tool for understanding the distribution of numeric data.
Number of classes (bins): the bin count shapes what you see. Too few bins can mask important features; too many can make random noise look like real patterns. Use Sturges' rule (k = 1 + 3.322 log n) or the square-root rule (k = square root of n) as starting points, then adjust.
Reading the shape: a histogram tells you whether data are symmetric (bell-shaped), right-skewed (tail to the right), left-skewed (tail to the left), or multimodal (multiple peaks).
Spotting outliers: look for isolated bars separated from the bulk of the data by gaps. Outliers pull the mean and inflate the standard deviation, so identifying them early matters for choosing the right summary statistics.
These three statistics each answer the question "what is a typical value?" in a different way.
Mean: sum all values and divide by the count. The mean uses every data point, which makes it informative for symmetric data but vulnerable to outliers. One extreme value can drag the mean a long way from the bulk of the data.
Median: sort the data and take the middle value (or the average of the two middle values for an even count). The median is resistant to outliers and skewness, making it the better choice when distributions are not symmetric.
Mode: the most frequently occurring value. The only measure of centre that works for categorical data. Also useful for identifying peaks in numeric distributions.
In a symmetric distribution, mean, median and mode are roughly equal. In a skewed distribution, the mean gets pulled toward the tail. The direction of the pull tells you the direction of the skew: if the mean is greater than the median, the data is right-skewed.
Centre alone does not describe a distribution. Two datasets can have the same mean but wildly different variability.
Range: maximum minus minimum. Quick to calculate but tells you nothing about how data is distributed between the extremes. A single outlier can make the range misleading.
Variance: the average of the squared deviations from the mean. Squaring ensures that positive and negative deviations do not cancel out. The downside: variance is in squared units, which is not intuitive.
Standard deviation: the square root of the variance. Back in the original units, so it is directly interpretable. For roughly bell-shaped data, about 68% of values fall within one standard deviation of the mean, about 95% within two.
Interquartile range (IQR): Q3 minus Q1, capturing the middle 50% of data. Resistant to outliers. Use IQR (paired with the median) when data is skewed; use standard deviation (paired with the mean) when data is symmetric.
A boxplot is a compact visual built from five values: minimum, Q1, median, Q3 and maximum.
The box spans Q1 to Q3, showing the middle 50% of the data (the IQR).
A line inside the box marks the median.
Whiskers extend from each end of the box to the most extreme data point that is still within 1.5 x IQR of the box edge.
Any data point beyond the whiskers is plotted as an individual dot and flagged as a potential outlier.
Boxplots are especially useful for comparing distributions side by side. You can quickly see differences in centre, spread, symmetry and the presence of outliers across groups.
The shape and nature of your data should drive your choice of summary statistic and graph.
Situation | Centre | Spread | Graph |
|---|---|---|---|
Symmetric, no outliers | Mean | Standard deviation | Histogram |
Skewed or outliers present | Median | IQR | Boxplot |
Categorical data | Mode | n/a | Bar graph or pie chart |
Always visualise the data before choosing a summary. A histogram or boxplot will reveal skewness and outliers that raw numbers alone might hide.
A z-score tells you how many standard deviations a value sits from the mean, and in which direction.
z = (x - mean) / standard deviation
A z-score of 0 means the value equals the mean.
A positive z-score means the value is above the mean.
A negative z-score means the value is below the mean.
Z-scores beyond +2 or -2 are often considered unusual; beyond +3 or -3, very unusual.
The real power of z-scores is comparability. If a student scores 85 on a maths exam (class mean 70, SD 10) and 80 on an English exam (class mean 65, SD 5), which performance was stronger relative to the class? The maths z-score is (85 - 70) / 10 = 1.5. The English z-score is (80 - 65) / 5 = 3.0. The English performance was further above the class, despite the lower raw score.
\bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_is^2 = \frac{1}{n-1} \sum_{i=1}^{n} (x_i - \bar{x})^2s = \sqrt{s^2}\text{IQR} = Q_3 - Q_1z = \frac{x - \mu}{\sigma}Outlier fences for boxplots: lower fence = Q1 - 1.5 x IQR, upper fence = Q3 + 1.5 x IQR. Any value outside these fences is a potential outlier.
Median household income is reported instead of mean income because income distributions are heavily right-skewed; a small number of very high earners would drag the mean well above what most people earn. Standard deviation is used in manufacturing quality control to set tolerance limits: a part whose measurement falls more than 3 standard deviations from the target mean is flagged as defective. Z-scores underpin grading on a curve, where raw scores are converted to a common scale so that performance across different exams or sections can be compared fairly.
Students often think the mean is always the best measure of centre. It is not. When data is skewed or contains outliers, the median gives a more representative picture.
Students confuse bar graphs and histograms. Bar graphs are for categorical data and have gaps between bars. Histograms are for numeric data and have no gaps.
A common error is reporting the range as a pair of numbers ("the range is 10 to 50"). The range is a single number: 50 - 10 = 40.
Students sometimes assume that an outlier is always a mistake or should be removed. An outlier may be a legitimate, important data point. Investigate before discarding.
Expect questions asking you to pick the appropriate measure of centre or spread for a given distribution. The answer nearly always hinges on whether the data is symmetric or skewed.
Know how to read and interpret a boxplot: identify median, IQR, whisker endpoints, and outliers from a diagram.
Be ready to compute and interpret z-scores. A question might give you two test scores from different distributions and ask which is more impressive.
Questions about histogram bin count tend to test whether you understand how too few or too many bins distort the picture.
The relationship between mean and median in skewed distributions is a perennial exam topic.
True or false: the median is always equal to the mean in a symmetric distribution. (True)
Fill in the blank: a z-score of -2 means the value is ____ standard deviations below the mean. (2)
True or false: the IQR is affected by outliers. (False)
Fill in the blank: in a boxplot, any point beyond ____ x IQR from Q1 or Q3 is flagged as a potential outlier. (1.5)
True or false: a histogram is appropriate for displaying categorical data. (False)
Q: A dataset has a mean of 50 and a standard deviation of 8. What is the z-score for a value of 66?
A: z = (66 - 50) / 8 = 2.0. The value is exactly 2 standard deviations above the mean.
Q: A distribution is right-skewed with several high outliers. Should you report the mean or the median as the measure of centre? Why?
A: Report the median. In a right-skewed distribution, the mean is pulled toward the high outliers and overstates the typical value. The median is resistant to this pull.
Q: A boxplot shows Q1 = 20, Q3 = 40. What are the outlier fences?
A: IQR = 40 - 20 = 20. Lower fence = 20 - 1.5(20) = -10. Upper fence = 40 + 1.5(20) = 70. Any value below -10 or above 70 is a potential outlier.
Q: You are given two histograms with the same mean but different shapes. One is symmetric and one is right-skewed. Which measure of spread would you use for each, and why?
A: For the symmetric histogram, use standard deviation (paired with the mean). For the right-skewed histogram, use IQR (paired with the median), because the skewness and possible outliers inflate the standard deviation.
Q: A student scored 72 on Exam A (mean 60, SD 6) and 85 on Exam B (mean 75, SD 10). On which exam did the student perform better relative to the class?
A: Exam A z-score = (72 - 60) / 6 = 2.0. Exam B z-score = (85 - 75) / 10 = 1.0. The student performed better relative to the class on Exam A.
Z-scores lead directly into the normal distribution and standardisation, covered in the probability and distributions notes. The idea of "how many standard deviations from the mean" becomes the basis for finding probabilities using z-tables. Measures of spread (variance, standard deviation) reappear when you study the variance of random variables and the expected value formulas for binomial and Poisson distributions. Boxplots and the five-number summary connect to the concept of percentiles and quantiles, which are used in hypothesis testing and confidence intervals later in the course.
descriptive statistics, categorical variable, qualitative variable, numeric variable, quantitative variable, discrete variable, continuous variable, histogram, bar graph, pie chart, mean, median, mode, average, central tendency, range, variance, standard deviation, IQR, interquartile range, five-number summary, boxplot, box-and-whisker plot, outlier, z-score, standard score, skewness, right-skewed, left-skewed, symmetric distribution, Sturges' rule, data visualisation, STAT 101, introduction to statistics