Descriptive Statistics and Data Types, STA 101 Ch. 1–3 – Study Notes
offline

Difficulty: Introductory | Prerequisites: None

TL;DR

Statistics splits into three jobs: collecting data, describing it, and drawing conclusions from it. Chapters 1–3 cover the tools for the first two: how to classify variables, read the shape of a distribution, and summarise a dataset with a handful of numbers (mean, median, standard deviation, quartiles). Master these and you have the vocabulary and toolkit the rest of the course builds on.

Key Terms

Population

The entire group you want to draw conclusions about. Denoted with Greek letters (μ, σ). In simple terms, it is every single case that fits your research question, whether or not you can actually measure them all.

Sample

A subset of the population that you actually collect data from. Denoted with Latin letters (x̅, s). Think of it as the slice of reality you can get your hands on.

Descriptive statistics

Methods for organising and summarising data you already have: charts, averages, spreads. No guessing beyond the dataset.

Inferential statistics

Methods for using sample data to draw conclusions about a larger population. This is where probability enters the picture.

Categorical variable

A variable whose values are labels or categories (colour, type, yes/no). You can count them but not meaningfully average them.

Quantitative variable

A variable whose values are numbers that represent amounts or measurements. These can be discrete (countable points, e.g. 12 occurrences) or continuous (any value in an interval, e.g. time).

Unimodal / bimodal / multimodal

Describes how many peaks (humps) a distribution has. One peak is unimodal, two is bimodal, more than two is multimodal.

Skewness

Whether a distribution's tail stretches further to the right (right-skewed) or to the left (left-skewed). A symmetric distribution has roughly equal tails on both sides.

Mean

The arithmetic average: sum all values, divide by the count. Sensitive to outliers.

Median

The middle value when data are sorted. Resistant to outliers, which makes it the better centre measure for skewed data.

Standard deviation (s)

The typical distance of data points from the mean. It is the square root of the variance. A larger s means more spread.

Quartiles (Q1, Q2, Q3)

Values that split sorted data into four equal parts. Q1 is the 25th percentile, Q2 is the median, Q3 is the 75th percentile.

Interquartile range (IQR)

Q3 minus Q1. It measures the spread of the middle 50% of the data and is the basis for the outlier rule.

Outlier

A data point that falls below Q1 – 1.5 × IQR or above Q3 + 1.5 × IQR. Think of it as a value suspiciously far from the bulk of the data.

Five-number summary

Min, Q1, Median, Q3, Max. These five values are enough to sketch a boxplot and give you the shape and spread of a dataset at a glance.

Core Content

Ch. 1 – The Three Branches of Statistics

  • Data collection – designing studies and surveys to gather useful data

  • Descriptive statistics – summarising and visualising what you collected

  • Inferential statistics – using sample results to say something about the population

  • Population parameters use Greek letters (μ for mean, σ for standard deviation)

  • Sample statistics use Latin letters (x̅ for mean, s for standard deviation)

Ch. 1.3 – Study Design and Sampling

  • Lurking variable – a variable not included as explanatory or response, but it may influence both

  • Confounding variable – two variables whose effects on the response cannot be separated

  • Three causal pictures to keep straight:

    • Causation: x directly causes y

    • Common response: a third variable z drives both x and y

    • Confounding: z is tangled with x so you cannot tell which one affects y

  • Sampling methods

    • Probability sample: every member of the population has a known chance of selection

    • Simple random sample (SRS): choose n individuals entirely at random from the whole population

    • Stratified random sample: divide the population into groups (strata), then take a random sample within each group

    • Convenience sample: pick whoever is easiest to reach (not random, results are unreliable for inference)

  • Bias to watch for

    • Undercoverage: some groups in the population are systematically left out

    • Non-response: some selected individuals never complete the study

Ch. 2 – Classifying Variables and Distribution Shape

  • By number of variables

    • Univariate: one variable

    • Bivariate: two variables

    • Multivariate: more than two

  • By type

    • Categorical: labels or categories (colour, type, yes/no)

    • Quantitative: numerical values

      • Discrete: countable (e.g. 12 occurrences)

      • Continuous: any value in an interval (e.g. time)

  • Distribution shape – described by the number of peaks (unimodal, bimodal, multimodal) and by symmetry or skew (symmetric, right-skewed, left-skewed)

  • Relative frequency = frequency of a category / total count

Ch. 3 – Numerical Summaries

  • Measures of centre

    • Mean: x̅ = (1/n) × Σx

    • Median: middle value of the sorted data

    • Mode: most frequently occurring value

  • Measures of spread

    • Variance: s²

    • Standard deviation: s = √(s²)

  • Quartiles

    • Sort data from least to greatest

    • Q1 = value at position n/4

    • Q2 = median

    • Q3 = value at position 3n/4

    • IQR = Q3 – Q1

  • Outlier rule

    • Below Q1 – 1.5 × IQR = outlier

    • Above Q3 + 1.5 × IQR = outlier

  • Five-number summary: Min, Q1, Median, Q3, Max – the skeleton of a boxplot

Formulas

Formula

Expression

Mean

x̅ = (1/n) × Σx

Standard deviation

s = √(variance)

Quartile positions

Q1 at n/4, Q2 at n/2 (median), Q3 at 3n/4

IQR

Q3 – Q1

Outlier boundaries

Lower: Q1 – 1.5 × IQR, Upper: Q3 + 1.5 × IQR

Relative frequency

frequency / total count

Real-World Applications

The five-number summary and boxplot are the standard first look at any dataset in science, business, and engineering. When a hospital reviews patient wait times or a factory checks part dimensions, they start with exactly these summaries.

Understanding sampling bias matters every time you read a poll or a product review. If the sample is convenience-based, the conclusions may not apply to the wider population.

Common Misconceptions

  • Students often confuse population parameters (μ, σ) with sample statistics (x̅, s). The symbols are not interchangeable; they refer to different things.

  • The mean is not always the best measure of centre. For skewed data or data with outliers, the median is more representative.

  • A bimodal distribution does not mean the data is wrong. It often signals two distinct groups within the dataset.

  • The IQR outlier rule is a guideline, not an absolute law. A point flagged as an outlier still deserves investigation before removal.

Why It Matters / Exam Flags

⚠️ Know when to use mean vs. median. If the exam gives you a skewed dataset and asks for the "best" measure of centre, the answer is median.

⚠️ Be able to compute all five numbers of the five-number summary from raw data.

⚠️ Expect a question asking you to classify a variable as categorical vs. quantitative, and discrete vs. continuous.

⚠️ Know the difference between SRS, stratified, and convenience sampling, and be able to identify which method a scenario describes.

⚠️ Understand lurking vs. confounding variables and be able to distinguish causation from common response and confounding.

Quick Self-Test

  1. True or false: A sample statistic is written with a Greek letter.

  1. Fill in the blank: IQR = ___ – ___.

  1. True or false: A right-skewed distribution has a longer tail on the left side.

  1. A dataset has Q1 = 20 and Q3 = 40. Any value below ___ or above ___ is an outlier.

  1. True or false: A convenience sample is a type of probability sample.

Answers: 1. False (Latin letters). 2. Q3 – Q1. 3. False (longer tail on the right). 4. Below −10, above 70. 5. False.

Practice Q&A

Q: A researcher surveys only the students in her 9 a.m. lecture to estimate campus-wide opinion on dining hall food. What type of sample is this, and what bias might result?

A: This is a convenience sample. It may produce undercoverage bias because students who do not take that class are excluded.

Q: A dataset has values 2, 5, 7, 8, 12, 15, 40. Calculate the five-number summary and determine whether 40 is an outlier.

A: Min = 2, Q1 = 5, Median = 8, Q3 = 15, Max = 40. IQR = 15 – 5 = 10. Upper fence = 15 + 1.5(10) = 30. Since 40 > 30, it is an outlier.

Q: Explain the difference between a lurking variable and a confounding variable.

A: A lurking variable is one that is not included in the study but may influence the results. A confounding variable is one whose effect on the response variable cannot be separated from the effect of the explanatory variable. A confounding variable is measured but entangled; a lurking variable is unmeasured entirely.

Q: A distribution is unimodal and left-skewed. Which is larger, the mean or the median?

A: The median is larger. In a left-skewed distribution, the tail of low values pulls the mean downward.

Connections to Other Topics

These descriptive tools reappear in every later chapter. The mean and variance formulas from Ch. 3 generalise directly into the expected value and variance of random variables in Ch. 5. Sampling methods from Ch. 1.3 set up the logic of inference you will meet after the midterm, where the quality of your sample determines how trustworthy your conclusions are.

Related Terms / Search Tags

STA 101, Purdue, intro to statistics, descriptive statistics, inferential statistics, population vs sample, Greek letters vs Latin letters, categorical vs quantitative, discrete vs continuous, unimodal, bimodal, multimodal, symmetric, right skew, left skew, mean, median, mode, standard deviation, variance, quartile, Q1, Q2, Q3, IQR, interquartile range, outlier, five-number summary, boxplot, relative frequency, SRS, simple random sample, stratified sample, convenience sample, undercoverage, non-response, lurking variable, confounding variable, causation vs correlation