Descriptive Statistics and Data Summaries, STAT 350 – Study Notes
offline

Source: Practice Exam 1, Spring 2016 | Purdue University

Difficulty: Introductory | Prerequisites: Basic algebra, familiarity with summation notation.

Big Picture

This material covers the foundational tools you use to describe and summarise a dataset before doing any formal inference. Medians, quartiles, the IQR and outlier detection sit at the heart of exploratory data analysis, and exam questions on them are essentially free marks once you have the procedure down. You should already be comfortable sorting data and reading a histogram. Everything here feeds directly into the probability and normal-distribution topics that follow.


TL;DR

Descriptive statistics condense raw data into a handful of numbers (centre, spread, shape) so you can see what is going on before running any formal test. The five-number summary, IQR and inner-fence rule for outliers are the core toolkit. Know the formulas for quartile depth, and practise spotting skewness from a histogram.


Key Terms

Inferential statistics

Techniques and methods used to analyse a small, specific set of data (a sample) in order to draw a conclusion about a large, more general collection of data (a population). In simple terms, you use what you can measure to make educated guesses about what you cannot.

Descriptive statistics

Graphical and numerical methods used to describe, organise and summarise data. Think of it as the "here is what the data look like" step, before any generalisation.

Skewness (negative / positive)

A measure of asymmetry in a distribution. A negatively skewed (left-skewed) distribution has a longer tail stretching to the left. A positively skewed (right-skewed) distribution has a longer tail to the right. In simple terms, the tail points in the direction of the skew's name.

Median (x̃)

The middle value when data are sorted in ascending order. For an odd number of observations n, it is the value at position (n + 1) / 2. It is resistant to outliers, unlike the mean.

Quartiles (Q₁, Q₃)

Values that split sorted data into four equal parts. Q₁ (the 25th percentile) separates the lower quarter; Q₃ (the 75th percentile) separates the upper quarter. You find their positions using the quartile depth formula.

Interquartile range (IQR)

Q₃ minus Q₁. It measures the spread of the middle 50% of the data and is the basis for the outlier rule.

Outlier

An observation that falls below the lower inner fence or above the upper inner fence (Q₁ - 1.5 × IQR or Q₃ + 1.5 × IQR). In simple terms, a data point that is unusually far from the bulk of the data.

Five-number summary

Minimum, Q₁, median, Q₃, maximum. These five values give you a quick snapshot of the centre, spread and range, and are the basis for a boxplot.


Core Content

Inferential vs Descriptive Statistics

  • Descriptive statistics summarise a dataset you already have: graphs, averages, spreads.

  • Inferential statistics use a sample to draw conclusions about a larger population.

  • The exam tests whether you can distinguish the two. If the question involves generalising from a sample to a population, it is inferential.

Reading Histogram Shape

  • Look at where the tail stretches. The tail direction gives the skew its name.

  • A histogram with a long left tail (values trailing off toward lower numbers) is negatively skewed.

  • A histogram with a long right tail is positively skewed.

  • Symmetric means roughly equal tails on both sides.

  • For negatively skewed data: mean < median < mode (the mean gets pulled toward the tail).

Finding the Median, Q₁ and Q₃

  • Sort the data in ascending order.

  • The median position for n observations: the value at position (n + 1) / 2.

    • For n = 11, the median is the 6th value.

  • Quartile depth formula:

    • d₁ = n / 4. Round up to the next whole number to get the position of Q₁.

    • d₃ = 3n / 4. Round up to get the position of Q₃.

    • Example with n = 11: d₁ = 11/4 = 2.75, round up to 3, so Q₁ is the 3rd value. d₃ = 33/4 = 8.25, round up to 9, so Q₃ is the 9th value.

Outlier Detection Using Inner Fences

  • Calculate IQR = Q₃ - Q₁.

  • Lower inner fence (IF_L) = Q₁ - 1.5 × IQR.

  • Upper inner fence (IF_H) = Q₃ + 1.5 × IQR.

  • Any data point below IF_L or above IF_H is an outlier.

  • You must state both fences and check both directions, even if outliers exist on only one side.

  • Example: Q₁ = 12, Q₃ = 20, IQR = 8. IF_L = 12 - 12 = 0 (no low outlier). IF_H = 20 + 12 = 32 (33 is above 32, so 33 is an outlier).

Five-Number Summary

  • Report all five values: minimum, Q₁, median, Q₃, maximum.

  • These are the building blocks of a boxplot.

  • Example: 5, 12, 14, 20, 33.


Formulas and Diagrams

\text{Median position} = \frac{n+1}{2}
d_1 = \frac{n}{4} \quad \text{(round up for } Q_1 \text{ position)}
d_3 = \frac{3n}{4} \quad \text{(round up for } Q_3 \text{ position)}
\text{IQR} = Q_3 - Q_1
\text{IF}_L = Q_1 - 1.5 \times \text{IQR}
\text{IF}_H = Q_3 + 1.5 \times \text{IQR}

Common Misconceptions

  • Students often confuse the direction of skewness. The skew is named for where the tail goes, not where the bulk of the data sits. A long left tail means negatively skewed.

  • Students sometimes forget to round up the quartile depth. d₁ = 2.75 means you go to the 3rd observation, not the 2nd or an interpolation between the 2nd and 3rd.

  • When checking for outliers, students frequently compute only one fence. You must check both the lower and upper inner fences, even when outliers appear on only one side.

  • Confusing descriptive and inferential statistics. If the question involves only summarising the data at hand, it is descriptive. If it involves drawing a conclusion about a broader population, it is inferential.


Why It Matters / Exam Flags

⚠️ The definition of inferential statistics is a common multiple-choice question. Know the exact wording: analysing a small set to draw conclusions about a larger collection.

⚠️ You will be asked to compute Q₁, Q₃, IQR and identify outliers. Show every step: the quartile depth calculation, both inner fences, and a clear statement of which values (if any) are outliers.

⚠️ The five-number summary is often a quick 2-point question. List all five values; do not skip any.

⚠️ Histogram shape questions test whether you can identify skewness by looking at the tail direction.


Quick Self-Test

  1. True or false: Inferential statistics are used only to summarise the data you already have. (False, that is descriptive statistics.)

  1. A histogram has a long tail stretching to the left. The distribution is ______ skewed. (Negatively.)

  1. True or false: If Q₁ = 10 and Q₃ = 30, the IQR is 30. (False, IQR = Q₃ - Q₁ = 20.)

  1. Fill in the blank: The upper inner fence equals Q₃ + ______ × IQR. (1.5.)

  1. True or false: For a dataset of 11 sorted values, the median is the 5th value. (False, it is the 6th value.)


Practice Q&A

Q: The following sorted dataset has 11 values: 5, 8, 12, 12, 14, 14, 16, 16, 20, 25, 33. What are Q₁ and Q₃?

A: d₁ = 11/4 = 2.75, round up to 3, so Q₁ is the 3rd value = 12. d₃ = 3(11)/4 = 8.25, round up to 9, so Q₃ is the 9th value = 20.

Q: Using Q₁ = 12 and Q₃ = 20, determine whether there are any outliers in the dataset above.

A: IQR = 20 - 12 = 8. Lower fence = 12 - 1.5(8) = 0. No values fall below 0, so no low outliers. Upper fence = 20 + 1.5(8) = 32. The value 33 exceeds 32, so 33 is an outlier.

Q: A histogram shows frequencies concentrated on the right with a long tail extending to the left. What is the shape of this distribution?

A: Negatively skewed (left-skewed).

Q: Which of the following best defines inferential statistics? (a) Methods to summarise data, (b) Methods to analyse a sample and draw conclusions about a population, (c) Methods to organise data graphically.

A: (b). Inferential statistics use sample data to draw conclusions about a larger population.


Connections to Other Topics

The five-number summary and boxplot connect directly to the normal distribution: when data are roughly symmetric, the mean and median are close, and the empirical rule (68-95-99.7) applies. Outlier detection matters again when you study regression later in the course, because influential points can distort a fitted line. Understanding skewness prepares you for the Central Limit Theorem, which explains why sampling distributions become approximately normal even when the underlying population is skewed.


Related Terms / Search Tags

Descriptive statistics, inferential statistics, histogram shape, left-skewed, right-skewed, negatively skewed, positively skewed, symmetric distribution, median, quartiles, Q1, Q3, quartile depth, interquartile range, IQR, inner fences, outlier detection, 1.5 IQR rule, five-number summary, boxplot, box-and-whisker plot, STAT 350, Purdue, introductory statistics, exploratory data analysis, EDA