Difficulty: Introductory | Prerequisites: None
Statistics splits into three jobs: collecting data, describing it, and drawing conclusions from it. Chapters 1–3 cover the tools for the first two: how to classify variables, read the shape of a distribution, and summarise a dataset with a handful of numbers (mean, median, standard deviation, quartiles). Master these and you have the vocabulary and toolkit the rest of the course builds on.
Population
The entire group you want to draw conclusions about. Denoted with Greek letters (μ, σ). In simple terms, it is every single case that fits your research question, whether or not you can actually measure them all.
Sample
A subset of the population that you actually collect data from. Denoted with Latin letters (x̅, s). Think of it as the slice of reality you can get your hands on.
Descriptive statistics
Methods for organising and summarising data you already have: charts, averages, spreads. No guessing beyond the dataset.
Inferential statistics
Methods for using sample data to draw conclusions about a larger population. This is where probability enters the picture.
Categorical variable
A variable whose values are labels or categories (colour, type, yes/no). You can count them but not meaningfully average them.
Quantitative variable
A variable whose values are numbers that represent amounts or measurements. These can be discrete (countable points, e.g. 12 occurrences) or continuous (any value in an interval, e.g. time).
Unimodal / bimodal / multimodal
Describes how many peaks (humps) a distribution has. One peak is unimodal, two is bimodal, more than two is multimodal.
Skewness
Whether a distribution's tail stretches further to the right (right-skewed) or to the left (left-skewed). A symmetric distribution has roughly equal tails on both sides.
Mean
The arithmetic average: sum all values, divide by the count. Sensitive to outliers.
Median
The middle value when data are sorted. Resistant to outliers, which makes it the better centre measure for skewed data.
Standard deviation (s)
The typical distance of data points from the mean. It is the square root of the variance. A larger s means more spread.
Quartiles (Q1, Q2, Q3)
Values that split sorted data into four equal parts. Q1 is the 25th percentile, Q2 is the median, Q3 is the 75th percentile.
Interquartile range (IQR)
Q3 minus Q1. It measures the spread of the middle 50% of the data and is the basis for the outlier rule.
Outlier
A data point that falls below Q1 – 1.5 × IQR or above Q3 + 1.5 × IQR. Think of it as a value suspiciously far from the bulk of the data.
Five-number summary
Min, Q1, Median, Q3, Max. These five values are enough to sketch a boxplot and give you the shape and spread of a dataset at a glance.
Data collection – designing studies and surveys to gather useful data
Descriptive statistics – summarising and visualising what you collected
Inferential statistics – using sample results to say something about the population
Population parameters use Greek letters (μ for mean, σ for standard deviation)
Sample statistics use Latin letters (x̅ for mean, s for standard deviation)
Lurking variable – a variable not included as explanatory or response, but it may influence both
Confounding variable – two variables whose effects on the response cannot be separated
Three causal pictures to keep straight:
Causation: x directly causes y
Common response: a third variable z drives both x and y
Confounding: z is tangled with x so you cannot tell which one affects y
Sampling methods
Probability sample: every member of the population has a known chance of selection
Simple random sample (SRS): choose n individuals entirely at random from the whole population
Stratified random sample: divide the population into groups (strata), then take a random sample within each group
Convenience sample: pick whoever is easiest to reach (not random, results are unreliable for inference)
Bias to watch for
Undercoverage: some groups in the population are systematically left out
Non-response: some selected individuals never complete the study
By number of variables
Univariate: one variable
Bivariate: two variables
Multivariate: more than two
By type
Categorical: labels or categories (colour, type, yes/no)
Quantitative: numerical values
Discrete: countable (e.g. 12 occurrences)
Continuous: any value in an interval (e.g. time)
Distribution shape – described by the number of peaks (unimodal, bimodal, multimodal) and by symmetry or skew (symmetric, right-skewed, left-skewed)
Relative frequency = frequency of a category / total count
Measures of centre
Mean: x̅ = (1/n) × Σx
Median: middle value of the sorted data
Mode: most frequently occurring value
Measures of spread
Variance: s²
Standard deviation: s = √(s²)
Quartiles
Sort data from least to greatest
Q1 = value at position n/4
Q2 = median
Q3 = value at position 3n/4
IQR = Q3 – Q1
Outlier rule
Below Q1 – 1.5 × IQR = outlier
Above Q3 + 1.5 × IQR = outlier
Five-number summary: Min, Q1, Median, Q3, Max – the skeleton of a boxplot
Formula | Expression |
|---|---|
Mean | x̅ = (1/n) × Σx |
Standard deviation | s = √(variance) |
Quartile positions | Q1 at n/4, Q2 at n/2 (median), Q3 at 3n/4 |
IQR | Q3 – Q1 |
Outlier boundaries | Lower: Q1 – 1.5 × IQR, Upper: Q3 + 1.5 × IQR |
Relative frequency | frequency / total count |
The five-number summary and boxplot are the standard first look at any dataset in science, business, and engineering. When a hospital reviews patient wait times or a factory checks part dimensions, they start with exactly these summaries.
Understanding sampling bias matters every time you read a poll or a product review. If the sample is convenience-based, the conclusions may not apply to the wider population.
Students often confuse population parameters (μ, σ) with sample statistics (x̅, s). The symbols are not interchangeable; they refer to different things.
The mean is not always the best measure of centre. For skewed data or data with outliers, the median is more representative.
A bimodal distribution does not mean the data is wrong. It often signals two distinct groups within the dataset.
The IQR outlier rule is a guideline, not an absolute law. A point flagged as an outlier still deserves investigation before removal.
⚠️ Know when to use mean vs. median. If the exam gives you a skewed dataset and asks for the "best" measure of centre, the answer is median.
⚠️ Be able to compute all five numbers of the five-number summary from raw data.
⚠️ Expect a question asking you to classify a variable as categorical vs. quantitative, and discrete vs. continuous.
⚠️ Know the difference between SRS, stratified, and convenience sampling, and be able to identify which method a scenario describes.
⚠️ Understand lurking vs. confounding variables and be able to distinguish causation from common response and confounding.
True or false: A sample statistic is written with a Greek letter.
Fill in the blank: IQR = ___ – ___.
True or false: A right-skewed distribution has a longer tail on the left side.
A dataset has Q1 = 20 and Q3 = 40. Any value below ___ or above ___ is an outlier.
True or false: A convenience sample is a type of probability sample.
Answers: 1. False (Latin letters). 2. Q3 – Q1. 3. False (longer tail on the right). 4. Below −10, above 70. 5. False.
Q: A researcher surveys only the students in her 9 a.m. lecture to estimate campus-wide opinion on dining hall food. What type of sample is this, and what bias might result?
A: This is a convenience sample. It may produce undercoverage bias because students who do not take that class are excluded.
Q: A dataset has values 2, 5, 7, 8, 12, 15, 40. Calculate the five-number summary and determine whether 40 is an outlier.
A: Min = 2, Q1 = 5, Median = 8, Q3 = 15, Max = 40. IQR = 15 – 5 = 10. Upper fence = 15 + 1.5(10) = 30. Since 40 > 30, it is an outlier.
Q: Explain the difference between a lurking variable and a confounding variable.
A: A lurking variable is one that is not included in the study but may influence the results. A confounding variable is one whose effect on the response variable cannot be separated from the effect of the explanatory variable. A confounding variable is measured but entangled; a lurking variable is unmeasured entirely.
Q: A distribution is unimodal and left-skewed. Which is larger, the mean or the median?
A: The median is larger. In a left-skewed distribution, the tail of low values pulls the mean downward.
These descriptive tools reappear in every later chapter. The mean and variance formulas from Ch. 3 generalise directly into the expected value and variance of random variables in Ch. 5. Sampling methods from Ch. 1.3 set up the logic of inference you will meet after the midterm, where the quality of your sample determines how trustworthy your conclusions are.
STA 101, Purdue, intro to statistics, descriptive statistics, inferential statistics, population vs sample, Greek letters vs Latin letters, categorical vs quantitative, discrete vs continuous, unimodal, bimodal, multimodal, symmetric, right skew, left skew, mean, median, mode, standard deviation, variance, quartile, Q1, Q2, Q3, IQR, interquartile range, outlier, five-number summary, boxplot, relative frequency, SRS, simple random sample, stratified sample, convenience sample, undercoverage, non-response, lurking variable, confounding variable, causation vs correlation