Descriptive Statistics and Data Fundamentals, STAT 301 Midterm 1 – Study Notes
offline

Difficulty: Beginner | Prerequisites: None

Big picture: This is the foundation for everything else in an introductory statistics course. Before you can run any tests or build any models, you need to know what kind of data you are working with, how to summarise it numerically, and how to spot outliers. These concepts reappear in every subsequent unit, so getting comfortable here saves you time later. If you have never taken a stats course before, this is the right place to start.


TL;DR

Statistics splits into two camps: descriptive (summarise what you have) and inferential (draw conclusions about a larger group). You describe data using measures of centre and spread, including mean, variance, and the five-number summary. The empirical rule and IQR fences help you flag unusual values.


Key Terms

Descriptive statistics

Graphical and numerical methods used to describe, organise, and summarise data. Think of it as the "what does my data look like?" step before you start drawing conclusions.

Inferential statistics

Techniques and methods used to analyse a small, specific set of data in order to draw a conclusion about a larger, more general collection of data. In simple terms, you are using a sample to make educated guesses about an entire population.

Univariate data

A dataset consisting of a single observation per subject. Think of it as one column in a spreadsheet.

Bivariate data

A dataset with two observations per subject (two columns). Used when you want to explore the relationship between two variables.

Multivariate data

Three or more observations per subject. Most real-world datasets are multivariate.

Numerical variable

A variable whose values are numbers that represent quantities (height, temperature, income).

Categorical variable

A variable whose values represent groups or labels (colour, gender, department).

Population variance (σ²)

The average of the squared deviations from the population mean. It measures how spread out the entire population's values are.

Sample variance (s²)

An estimate of population variance calculated from a sample: s² = (1 / (n – 1)) × Σ(xᵢ – x̄)². The denominator is n – 1 (not n) to correct for the bias introduced by estimating from a sample.

Standard deviation

The square root of the variance. It brings the spread measure back into the same units as the original data, which makes it easier to interpret.

Mean (x̄)

The arithmetic average: x̄ = (1/n) × Σxᵢ. The balancing point of the dataset.


Core Content

Branches of Statistics

  • Data collection comes first: you gather observations before anything else.

  • Descriptive statistics summarises what you collected, using graphs, tables, and numerical measures.

  • Inferential statistics uses the summary to make claims about the wider population.

Types of Variables

  • Numerical variables carry quantitative meaning (you can add or average them).

  • Categorical variables carry qualitative meaning (you count how many fall in each group).

  • The distinction matters because it determines which summary measures and tests you can use.

Five-Number Summary

The five-number summary consists of:

  • Minimum

  • First quartile (Q₁)

  • Median

  • Third quartile (Q₃)

  • Maximum

This gives you a quick picture of the centre and spread of the data, and it is the basis for box plots.

Outlier Detection: The 1.5 × IQR Rule

  • Compute IQR = Q₃ – Q₁.

  • Mild outliers lie beyond Q₃ + 1.5 × IQR or below Q₁ – 1.5 × IQR.

  • Extreme outliers lie beyond Q₃ + 3 × IQR or below Q₁ – 3 × IQR.

  • This is a mechanical rule, not a judgement call. Flag first, investigate second.

Mean and Variance

  • The mean is sensitive to outliers; a single extreme value can shift it substantially.

  • Variance quantifies spread by averaging the squared distances from the mean.

  • Standard deviation (the square root of variance) is in the same units as the data, making it more interpretable than variance alone.

  • Population variance uses N in the denominator. Sample variance uses n – 1.

The Empirical Rule (68-95-99.7 Rule)

For data that is approximately normal (bell-shaped):

  • About 68% of observations fall within 1 standard deviation of the mean.

  • About 95% fall within 2 standard deviations.

  • About 99.7% fall within 3 standard deviations.

This is a quick-check tool. If your data is heavily skewed, the empirical rule does not apply cleanly.


Formulas

Measure

Formula

Mean

x̄ = (1/n) Σxᵢ

Sample variance

s² = [1/(n–1)] Σ(xᵢ – x̄)²

Population variance

σ² = (1/N) Σ(xᵢ – μ)²

Standard deviation

s = √(s²)

IQR

Q₃ – Q₁

Mild outlier fences

Q₁ – 1.5·IQR and Q₃ + 1.5·IQR


Real-World Applications

The five-number summary and IQR rule are exactly how quality-control engineers decide whether a manufacturing process is producing items within acceptable tolerances. If package weights start drifting outside the fences, the line gets flagged for recalibration.


Common Misconceptions

  • "Variance and standard deviation measure the same thing, so it doesn't matter which I report." Variance is in squared units; standard deviation is in the original units. Reporting variance for heights would give you "square centimetres," which is not useful for interpretation.

  • "I divide by n for sample variance." You divide by n – 1. The n – 1 correction (Bessel's correction) accounts for the fact that a sample underestimates the true spread.

  • "The empirical rule works for any dataset." It requires an approximately normal (bell-shaped) distribution. For skewed data, the percentages can be quite different.

  • "An outlier is always an error." The IQR rule flags unusual values mechanically. They may be legitimate observations that carry important information.


Why It Matters / Exam Flags

⚠️ Know when to use n vs. n – 1 in the denominator. This comes up in nearly every variance calculation on the exam.

⚠️ Be able to compute a five-number summary by hand and identify outliers using the 1.5 × IQR rule.

⚠️ The empirical rule percentages (68, 95, 99.7) are frequently tested as quick-recall questions.

⚠️ Make sure you can distinguish descriptive from inferential statistics with a concrete example.


Quick Self-Test

True or false: The standard deviation can be negative.

A: False. It is the square root of variance, which is always non-negative.

Fill in the blank: About ___% of data in a normal distribution falls within 2 standard deviations of the mean.

A: 95%

True or false: If you have two observations per subject, you have multivariate data.

A: False. Two observations per subject is bivariate. Multivariate means three or more.


Practice Q&A

Q: What is the difference between a parameter and a statistic?

A: A parameter is a numerical measure of a population. A statistic is a numerical measure computed from a sample. You use statistics to estimate parameters.

Q: A dataset has Q₁ = 20, Q₃ = 40. A value of 75 is observed. Is it a mild outlier, extreme outlier, or neither?

A: IQR = 40 – 20 = 20. Upper mild fence = 40 + 1.5(20) = 70. Upper extreme fence = 40 + 3(20) = 100. Since 75 > 70 but 75 < 100, it is a mild outlier.

Q: Why does sample variance divide by n – 1 instead of n?

A: Dividing by n systematically underestimates the true population variance. Using n – 1 corrects for this bias (Bessel's correction), giving an unbiased estimator of σ².

Q: State the empirical rule.

A: For a normal distribution, approximately 68% of data falls within 1 SD of the mean, 95% within 2 SDs, and 99.7% within 3 SDs.


Connections to Other Topics

This material connects directly to probability distributions (the next unit), because variance and standard deviation become the parameters of those distributions. It also connects to sampling distributions and the Central Limit Theorem, where the standard deviation of the sample mean depends on the population standard deviation divided by √n.


Related Terms / Search Tags

descriptive statistics, inferential statistics, mean, average, sample variance, population variance, standard deviation, five number summary, box plot, IQR, interquartile range, outlier detection, 1.5 IQR rule, empirical rule, 68-95-99.7 rule, numerical variable, categorical variable, univariate, bivariate, multivariate, Bessel's correction, measures of centre, measures of spread, Purdue STAT 301, intro to statistics midterm 1