Data, Sampling, and Descriptive Statistics, STAT 101 Chs. 1–3 – Study Notes
offline

Difficulty: Introductory | Prerequisites: None

Big Picture

These three chapters lay the groundwork for the entire STAT 101 course. Chapter 1 introduces what statistics is and how data is collected (sampling). Chapter 2 covers the types of data you will encounter and how distributions are described. Chapter 3 gives you the numerical tools to summarise a dataset: centre, spread, and position. If you do not have these foundations solid, every later chapter (probability, inference, regression) will feel harder than it needs to.


TL;DR

Statistics has three branches: data collection, description, and inference. Populations are the full group of interest; samples are subsets we actually measure. To describe a sample numerically, you need the mean, standard deviation, quartiles, and IQR, and you need to know how to spot outliers with the 1.5 * IQR rule.


Key Terms

Population (denoted with Greek letters)

The total set of observations of interest. In a study on the weight of adult women, the population is the weight of all adult women.

In simple terms, this is the entire group you want to draw conclusions about.

Sample (denoted with Latin letters)

A subset of the population that is actually measured. In the same study, the sample is the group of adult women whose weights you record.

Think of it as the slice of the population you can get your hands on.

Lurking variable

A variable that is not among the explanatory or response variables in a study but may influence the results.

In simple terms, it is a hidden factor you did not account for in your study design.

Confounding variable

Two variables associated in such a way that their effects on a response variable cannot be distinguished from each other.

Think of it as two possible causes tangled together so you cannot tell which one is doing the work.

Simple Random Sample (SRS)

A sample of size n chosen so that every possible sample of that size has an equal chance of being selected.

In simple terms, every member of the population has the same shot at being picked.

Stratified random sample

The population is divided into groups (strata), and a random sample is drawn from each group. For example, choosing 10 random women and 10 random men from a database.

Convenience sample

A sample chosen because it is easy to collect, not because it is random. Surveying only the people in your class is a convenience sample.

Categorical variable

A variable that records a category or label, such as colour or type. You cannot do arithmetic on it.

Quantitative variable

A variable that records a numerical value. It can be discrete (countable points, e.g. 12 occurrences) or continuous (any value in an interval, e.g. time).

Unimodal, bimodal, multimodal

Descriptions of a distribution's shape by the number of peaks (humps). One peak is unimodal, two is bimodal, more than two is multimodal.

Mean (x-bar)

The arithmetic average of a dataset. Sum all values and divide by n.

Standard deviation (s)

A measure of how spread out the data values are from the mean. It is the square root of the variance.

Think of it as the typical distance a data point sits from the average.

Quartiles (Q1, Q2, Q3)

Values that split a sorted dataset into four equal parts. Q1 is the 25th percentile, Q2 is the median (50th percentile), Q3 is the 75th percentile.

Interquartile Range (IQR)

Q3 minus Q1. It measures the spread of the middle 50% of the data.

Outlier

A data point that falls below Q1 minus 1.5 * IQR or above Q3 plus 1.5 * IQR.

Five-number summary

Min, Q1, Median, Q3, Max. The five values that define a box plot.


Core Content

Three Branches of Statistics (Ch. 1)

  • Data collection -- designing studies and surveys to gather information

  • Descriptive statistics -- organising and summarising data (tables, graphs, numerical summaries)

  • Inferential statistics -- drawing conclusions about a population from a sample

Population vs. Sample (Ch. 1)

  • The population is the full group you care about. Parameters (population summaries) use Greek letters (e.g. mu, sigma).

  • The sample is the subset you measure. Statistics (sample summaries) use Latin letters (e.g. x-bar, s).

Sampling Methods (Ch. 1)

  • Probability sample -- every member has a known chance of selection

  • Simple Random Sample (SRS) -- every possible sample of size n is equally likely

  • Stratified random sample -- divide the population into strata, then SRS within each stratum

  • Convenience sample -- chosen for ease, not randomness; results are not generalisable

Sampling Problems (Ch. 1)

  • Undercoverage -- some groups in the population are left out of the sampling frame

  • Non-response -- selected individuals do not complete the study

Lurking and Confounding Variables (Ch. 1)

  • A lurking variable is not measured in the study but can influence the relationship you observe.

  • A confounding variable is entangled with the explanatory variable so you cannot separate their effects.

  • Three causal diagrams to know: direct causation (x causes y), a common response (z causes both x and y), and confounding (z is associated with both x and y and you cannot tell them apart).

Data Types (Ch. 2)

  • Univariate -- one variable

  • Bivariate -- two variables

  • Multivariate -- more than two variables

  • Categorical -- labels or categories (colour, type)

  • Quantitative -- numerical; split into discrete (countable) and continuous (any value in a range)

Distribution Shape (Ch. 2)

  • Unimodal -- one peak

  • Bimodal -- two peaks

  • Multimodal -- more than two peaks

  • Relative frequency = frequency / total count

Measures of Centre (Ch. 3)

  • The mean is the arithmetic average: sum all values, divide by n.

  • The median (Q2) is the middle value of the sorted data.

Measures of Spread (Ch. 3)

  • Standard deviation (s) is the square root of the variance.

  • IQR = Q3 minus Q1, covering the middle 50% of the data.

Quartiles and Outliers (Ch. 3)

  • Sort the data from least to greatest.

  • Q1 is at position n/4, Q2 is the median, Q3 is at position 3n/4.

  • Outlier rule: any point below Q1 minus 1.5 * IQR or above Q3 plus 1.5 * IQR is flagged as an outlier.

  • The five-number summary (Min, Q1, Median, Q3, Max) is the basis for a box plot.


Formulas

\bar{x} = \frac{1}{n} \sum x
s = \sqrt{\text{variance}} = \sqrt{s^2}
Q1 = \text{value at position } \frac{n}{4}, \quad Q3 = \text{value at position } \frac{3n}{4}
IQR = Q3 - Q1
\text{Lower outlier bound} = Q1 - 1.5 \times IQR
\text{Upper outlier bound} = Q3 + 1.5 \times IQR
\text{Relative frequency} = \frac{\text{frequency}}{\text{total count}}

Real-World Applications

Polling organisations use stratified random sampling to ensure they hear from every demographic group, not just the easiest people to reach. The five-number summary and box plots are used in quality control to quickly spot whether a manufacturing process is producing outliers.


Common Misconceptions

  • Students often confuse lurking variables with confounding variables. A lurking variable is simply unmeasured; a confounding variable is measured but entangled with the explanatory variable so their effects cannot be separated.

  • Students sometimes think the mean is always the best measure of centre. For skewed data, the median is more representative because extreme values pull the mean toward the tail.

  • Students assume convenience samples are fine for inference. They are not, because the lack of randomness means you cannot generalise to the population.

  • Students forget to sort the data before computing quartiles. The quartile positions (n/4, 3n/4) only make sense on sorted data.


Why It Matters / Exam Flags

  • Expect questions asking you to identify the population vs. the sample in a given study.

  • Expect questions naming a sampling method and asking you to classify it (SRS, stratified, convenience).

  • Outlier detection using the 1.5 * IQR rule is a routine exam calculation. Know the steps: sort, find Q1 and Q3, compute IQR, compute bounds, check each point.

  • The five-number summary is the basis of box plot interpretation questions.


Quick Self-Test

  1. True or False: A sample statistic is denoted with a Greek letter. (Answer: False, Greek letters are for population parameters.)

  1. Fill in the blank: IQR = ______ minus ______. (Answer: Q3 minus Q1.)

  1. True or False: A convenience sample can be used to make valid inferences about a population. (Answer: False.)

  1. Fill in the blank: A data point is an outlier if it falls below Q1 minus ______ or above Q3 plus ______. (Answer: 1.5 * IQR.)

  1. True or False: A bimodal distribution has exactly one peak. (Answer: False, it has two.)


Practice Q&A

Q: A researcher surveys every third person leaving a library to ask about reading habits. What sampling method is this, and can the results be generalised to the entire city's population?

A: This is a convenience sample (systematic, but only from library visitors). It cannot be generalised to the city because non-library visitors are excluded (undercoverage).

Q: Given the sorted data set 0, 15, 20, 20, 35, 60, 60, 125 (n = 8), find Q1, Q3, the IQR, and identify any outliers.

A: Q1 position = 8/4 = 2, so Q1 = 15 (interpolation may give 17.5 depending on method). Q3 position = 3(8)/4 = 6, so Q3 = 60. IQR = 60 minus 15 = 45. Lower bound = 15 minus 1.5(45) = minus 52.5. Upper bound = 60 plus 1.5(45) = 127.5. No values fall outside these bounds, so there are no outliers.

Q: In a study examining whether a new fertiliser increases crop yield, the researcher notices that fields receiving more sunlight also received more fertiliser. Is sunlight a lurking variable or a confounding variable?

A: Sunlight is a confounding variable. It is associated with both the explanatory variable (fertiliser amount) and the response variable (crop yield), and its effect cannot be separated from the fertiliser's effect.

Q: What is the five-number summary for the data set 3, 7, 8, 12, 15, 18, 22?

A: Min = 3, Q1 = 7, Median = 12, Q3 = 18, Max = 22.


Connections to Other Topics

The population-vs.-sample distinction carries directly into Chapters 8 to 10, where you build confidence intervals and run hypothesis tests: you use sample statistics to estimate population parameters. The mean and standard deviation from Chapter 3 feed into every formula for z-tests and t-tests. Understanding data types (categorical vs. quantitative) determines which statistical test you can apply later in the course.


Related Terms / Search Tags

STAT 101, introduction to statistics, Purdue, descriptive statistics, inferential statistics, population parameter, sample statistic, SRS, simple random sample, stratified sample, convenience sample, undercoverage, non-response, lurking variable, confounding variable, categorical variable, quantitative variable, discrete variable, continuous variable, unimodal, bimodal, multimodal, mean, average, x-bar, standard deviation, variance, quartile, Q1, Q2, Q3, median, IQR, interquartile range, outlier, 1.5 IQR rule, five-number summary, box plot, relative frequency