Difficulty: Introductory | Prerequisites: None
These three chapters lay the groundwork for the entire STAT 101 course. Chapter 1 introduces what statistics is and how data is collected (sampling). Chapter 2 covers the types of data you will encounter and how distributions are described. Chapter 3 gives you the numerical tools to summarise a dataset: centre, spread, and position. If you do not have these foundations solid, every later chapter (probability, inference, regression) will feel harder than it needs to.
Statistics has three branches: data collection, description, and inference. Populations are the full group of interest; samples are subsets we actually measure. To describe a sample numerically, you need the mean, standard deviation, quartiles, and IQR, and you need to know how to spot outliers with the 1.5 * IQR rule.
Population (denoted with Greek letters)
The total set of observations of interest. In a study on the weight of adult women, the population is the weight of all adult women.
In simple terms, this is the entire group you want to draw conclusions about.
Sample (denoted with Latin letters)
A subset of the population that is actually measured. In the same study, the sample is the group of adult women whose weights you record.
Think of it as the slice of the population you can get your hands on.
Lurking variable
A variable that is not among the explanatory or response variables in a study but may influence the results.
In simple terms, it is a hidden factor you did not account for in your study design.
Confounding variable
Two variables associated in such a way that their effects on a response variable cannot be distinguished from each other.
Think of it as two possible causes tangled together so you cannot tell which one is doing the work.
Simple Random Sample (SRS)
A sample of size n chosen so that every possible sample of that size has an equal chance of being selected.
In simple terms, every member of the population has the same shot at being picked.
Stratified random sample
The population is divided into groups (strata), and a random sample is drawn from each group. For example, choosing 10 random women and 10 random men from a database.
Convenience sample
A sample chosen because it is easy to collect, not because it is random. Surveying only the people in your class is a convenience sample.
Categorical variable
A variable that records a category or label, such as colour or type. You cannot do arithmetic on it.
Quantitative variable
A variable that records a numerical value. It can be discrete (countable points, e.g. 12 occurrences) or continuous (any value in an interval, e.g. time).
Unimodal, bimodal, multimodal
Descriptions of a distribution's shape by the number of peaks (humps). One peak is unimodal, two is bimodal, more than two is multimodal.
Mean (x-bar)
The arithmetic average of a dataset. Sum all values and divide by n.
Standard deviation (s)
A measure of how spread out the data values are from the mean. It is the square root of the variance.
Think of it as the typical distance a data point sits from the average.
Quartiles (Q1, Q2, Q3)
Values that split a sorted dataset into four equal parts. Q1 is the 25th percentile, Q2 is the median (50th percentile), Q3 is the 75th percentile.
Interquartile Range (IQR)
Q3 minus Q1. It measures the spread of the middle 50% of the data.
Outlier
A data point that falls below Q1 minus 1.5 * IQR or above Q3 plus 1.5 * IQR.
Five-number summary
Min, Q1, Median, Q3, Max. The five values that define a box plot.
Data collection -- designing studies and surveys to gather information
Descriptive statistics -- organising and summarising data (tables, graphs, numerical summaries)
Inferential statistics -- drawing conclusions about a population from a sample
The population is the full group you care about. Parameters (population summaries) use Greek letters (e.g. mu, sigma).
The sample is the subset you measure. Statistics (sample summaries) use Latin letters (e.g. x-bar, s).
Probability sample -- every member has a known chance of selection
Simple Random Sample (SRS) -- every possible sample of size n is equally likely
Stratified random sample -- divide the population into strata, then SRS within each stratum
Convenience sample -- chosen for ease, not randomness; results are not generalisable
Undercoverage -- some groups in the population are left out of the sampling frame
Non-response -- selected individuals do not complete the study
A lurking variable is not measured in the study but can influence the relationship you observe.
A confounding variable is entangled with the explanatory variable so you cannot separate their effects.
Three causal diagrams to know: direct causation (x causes y), a common response (z causes both x and y), and confounding (z is associated with both x and y and you cannot tell them apart).
Univariate -- one variable
Bivariate -- two variables
Multivariate -- more than two variables
Categorical -- labels or categories (colour, type)
Quantitative -- numerical; split into discrete (countable) and continuous (any value in a range)
Unimodal -- one peak
Bimodal -- two peaks
Multimodal -- more than two peaks
Relative frequency = frequency / total count
The mean is the arithmetic average: sum all values, divide by n.
The median (Q2) is the middle value of the sorted data.
Standard deviation (s) is the square root of the variance.
IQR = Q3 minus Q1, covering the middle 50% of the data.
Sort the data from least to greatest.
Q1 is at position n/4, Q2 is the median, Q3 is at position 3n/4.
Outlier rule: any point below Q1 minus 1.5 * IQR or above Q3 plus 1.5 * IQR is flagged as an outlier.
The five-number summary (Min, Q1, Median, Q3, Max) is the basis for a box plot.
\bar{x} = \frac{1}{n} \sum xs = \sqrt{\text{variance}} = \sqrt{s^2}Q1 = \text{value at position } \frac{n}{4}, \quad Q3 = \text{value at position } \frac{3n}{4}IQR = Q3 - Q1\text{Lower outlier bound} = Q1 - 1.5 \times IQR\text{Upper outlier bound} = Q3 + 1.5 \times IQR\text{Relative frequency} = \frac{\text{frequency}}{\text{total count}}Polling organisations use stratified random sampling to ensure they hear from every demographic group, not just the easiest people to reach. The five-number summary and box plots are used in quality control to quickly spot whether a manufacturing process is producing outliers.
Students often confuse lurking variables with confounding variables. A lurking variable is simply unmeasured; a confounding variable is measured but entangled with the explanatory variable so their effects cannot be separated.
Students sometimes think the mean is always the best measure of centre. For skewed data, the median is more representative because extreme values pull the mean toward the tail.
Students assume convenience samples are fine for inference. They are not, because the lack of randomness means you cannot generalise to the population.
Students forget to sort the data before computing quartiles. The quartile positions (n/4, 3n/4) only make sense on sorted data.
Expect questions asking you to identify the population vs. the sample in a given study.
Expect questions naming a sampling method and asking you to classify it (SRS, stratified, convenience).
Outlier detection using the 1.5 * IQR rule is a routine exam calculation. Know the steps: sort, find Q1 and Q3, compute IQR, compute bounds, check each point.
The five-number summary is the basis of box plot interpretation questions.
True or False: A sample statistic is denoted with a Greek letter. (Answer: False, Greek letters are for population parameters.)
Fill in the blank: IQR = ______ minus ______. (Answer: Q3 minus Q1.)
True or False: A convenience sample can be used to make valid inferences about a population. (Answer: False.)
Fill in the blank: A data point is an outlier if it falls below Q1 minus ______ or above Q3 plus ______. (Answer: 1.5 * IQR.)
True or False: A bimodal distribution has exactly one peak. (Answer: False, it has two.)
Q: A researcher surveys every third person leaving a library to ask about reading habits. What sampling method is this, and can the results be generalised to the entire city's population?
A: This is a convenience sample (systematic, but only from library visitors). It cannot be generalised to the city because non-library visitors are excluded (undercoverage).
Q: Given the sorted data set 0, 15, 20, 20, 35, 60, 60, 125 (n = 8), find Q1, Q3, the IQR, and identify any outliers.
A: Q1 position = 8/4 = 2, so Q1 = 15 (interpolation may give 17.5 depending on method). Q3 position = 3(8)/4 = 6, so Q3 = 60. IQR = 60 minus 15 = 45. Lower bound = 15 minus 1.5(45) = minus 52.5. Upper bound = 60 plus 1.5(45) = 127.5. No values fall outside these bounds, so there are no outliers.
Q: In a study examining whether a new fertiliser increases crop yield, the researcher notices that fields receiving more sunlight also received more fertiliser. Is sunlight a lurking variable or a confounding variable?
A: Sunlight is a confounding variable. It is associated with both the explanatory variable (fertiliser amount) and the response variable (crop yield), and its effect cannot be separated from the fertiliser's effect.
Q: What is the five-number summary for the data set 3, 7, 8, 12, 15, 18, 22?
A: Min = 3, Q1 = 7, Median = 12, Q3 = 18, Max = 22.
The population-vs.-sample distinction carries directly into Chapters 8 to 10, where you build confidence intervals and run hypothesis tests: you use sample statistics to estimate population parameters. The mean and standard deviation from Chapter 3 feed into every formula for z-tests and t-tests. Understanding data types (categorical vs. quantitative) determines which statistical test you can apply later in the course.
STAT 101, introduction to statistics, Purdue, descriptive statistics, inferential statistics, population parameter, sample statistic, SRS, simple random sample, stratified sample, convenience sample, undercoverage, non-response, lurking variable, confounding variable, categorical variable, quantitative variable, discrete variable, continuous variable, unimodal, bimodal, multimodal, mean, average, x-bar, standard deviation, variance, quartile, Q1, Q2, Q3, median, IQR, interquartile range, outlier, 1.5 IQR rule, five-number summary, box plot, relative frequency