Descriptive Statistics, Boxplots, and Data Visualization, STAT 350 Midterm 1 – Study Notes
offline

Source: STAT 350 Practice Exam 1, Purdue University

Tags: descriptive statistics, boxplot, modified boxplot, inner fence, outer fence, outlier, IQR, median, mean, sample statistic, population parameter, measures of center, measures of spread, skewness, inferential statistics, probability

Difficulty: Foundational Prerequisites: None. This is typically the first block of content in an introductory statistics course.


Big Picture

Before you can model data or test hypotheses, you need to describe it. Descriptive statistics gives you the vocabulary and the tools: numbers that summarise centre and spread, and plots that show shape, outliers, and comparisons at a glance. The boxplot is the single most exam-tested visualisation in STAT 350 because it packages five-number summaries, fences, and outlier detection into one picture. Understanding what each part of the boxplot means, and how to read it, is essential. This topic also lays down the critical distinction between a sample statistic and a population parameter, which runs through every later topic in the course.


TL;DR

Descriptive statistics summarises a dataset using measures of centre (mean, median), measures of spread (standard deviation, IQR, range), and visualisations such as boxplots. The modified boxplot is the standard exam format: it uses inner and outer fences to classify outliers, and reading it correctly requires knowing how fences are calculated from Q1, Q3, and the IQR.


Key Terms

Sample statistic

A numerical summary computed from sample data. Examples: sample mean (x-bar), sample median, sample standard deviation (s). Think of it as a number you calculate from the data you collected.

Population parameter

A numerical summary that describes the entire population. Examples: population mean (mu), population standard deviation (sigma). Usually unknown and estimated using sample statistics. In simple terms, the "true" value you are trying to learn about.

x-bar (sample mean)

The arithmetic average of the sample observations: x-bar = (sum of all values) / n. It is a sample statistic and a measure of centre. Think of it as the "balance point" of the data.

Sample median

The middle value when data are sorted. For an even number of observations, it is the average of the two middle values. It is a measure of centre that is resistant to outliers. In simple terms, the value that splits the sorted data in half.

Interquartile range (IQR)

IQR = Q3 - Q1, where Q1 is the 25th percentile and Q3 is the 75th percentile. It measures the spread of the middle 50% of the data. Think of it as how wide the "box" is in a boxplot. Note: IQR is a measure of spread, not a measure of central tendency.

Five-number summary

Minimum, Q1, Median, Q3, Maximum. The foundation of a boxplot.

Modified boxplot

A boxplot that uses inner fences to identify outliers. Whiskers extend to the most extreme data points within the inner fences (not to the min and max). Points beyond the inner fences are plotted individually. In simple terms, a boxplot that flags unusual values instead of hiding them in the whiskers.

Inner fences

  • Lower inner fence (IFL): Q1 - 1.5 · IQR

  • Upper inner fence (IFU): Q3 + 1.5 · IQR

Points outside the inner fences but inside the outer fences are "mild" outliers (often marked with asterisks).

Outer fences

  • Lower outer fence (OFL): Q1 - 3 · IQR

  • Upper outer fence (OFU): Q3 + 3 · IQR

Points beyond the outer fences are "extreme" outliers (often marked with solid dots).

Skewness

A description of asymmetry in a distribution.

  • Right-skewed (positive skew): tail extends to the right, mean > median.

  • Left-skewed (negative skew): tail extends to the left, mean < median.

  • Symmetric: roughly mirror-image around the centre, mean ≈ median.

Descriptive statistics vs inferential statistics vs probability

Descriptive statistics summarises observed data. Inferential statistics uses sample data to draw conclusions about populations. Probability starts from a known population model and predicts what samples will look like.


Core Content

Measures of Centre and Their Properties

  • The sample mean (x-bar) is sensitive to outliers because every data point contributes to the sum.

  • The sample median is resistant to outliers because it depends only on the position of the middle values, not on extreme values.

  • When asked "which measure of central tendency is least influenced by outliers," the answer is the median.

  • IQR and standard deviation are measures of spread, not central tendency. This distinction is tested.

What x-bar Is (and Is Not)

  • x-bar is a sample statistic, not a population parameter (mu is the population mean).

  • x-bar is a measure of centre (location), not a measure of spread.

  • x-bar is calculated from observed data. It estimates mu but is not itself a population quantity.

Reading a Modified Boxplot

  • The box spans from Q1 to Q3. The line inside the box is the median.

  • Whiskers extend from the box to the most extreme data point that is still within the inner fences. They do not necessarily reach the inner fence values.

  • Mild outliers (between inner and outer fences) are typically shown as asterisks (*).

  • Extreme outliers (beyond outer fences) are typically shown as solid dots.

Calculating Inner Fences from a Boxplot

  • Step 1: Read Q1 and Q3 from the edges of the box.

  • Step 2: Compute IQR = Q3 - Q1.

  • Step 3: IFL = Q1 - 1.5 · IQR. IFU = Q3 + 1.5 · IQR.

  • The inner fence is a calculated boundary. It may not correspond to any actual data point.

  • On the exam, you may need to estimate Q1 and Q3 from the boxplot and then compute the fence. The answer may be a range rather than an exact value.

Identifying "Real" Outliers in Boxplots

  • In a modified boxplot, mild outliers are plotted as individual symbols outside the whiskers.

  • "Real" outliers typically refer to extreme outliers, those beyond the outer fences, shown as solid dots.

  • An algorithm that has no individually plotted points outside its whiskers has no outliers.

  • If a boxplot shows only asterisks, those are mild outliers. A solid dot indicates an extreme outlier.

Describing the Shape of a Distribution from a Boxplot

  • Compare the median line's position within the box. If it is closer to Q1, the data may be right-skewed. If closer to Q3, left-skewed.

  • Compare whisker lengths. A much longer upper whisker suggests right skew.

  • Outliers on one side also suggest skew in that direction.

  • Symmetric: median is roughly centred in the box, whiskers are roughly equal, outliers (if any) appear on both sides.

Comparing Boxplots Side by Side

  • Lower median and lower box position indicate generally smaller values.

  • A shorter box (smaller IQR) indicates less variability in the middle 50%.

  • An algorithm with a lower median and shorter box is generally "better" if smaller values are preferred (e.g., faster running times).

  • When comparing algorithms: look at both centre (which is typical?) and spread (how consistent?).

Descriptive vs Inferential Statistics vs Probability

  • Descriptive statistics: summarise the data you have (compute means, draw histograms, make tables).

  • Inferential statistics: use sample data to make conclusions about the population (confidence intervals, hypothesis tests).

  • Probability: given a known population model and all its parameters, predict what samples will look like. This goes from population to sample, the reverse direction of inference.

Estimating Counts from Boxplots

  • In a boxplot with n observations, approximately 25% of data lies below Q1, 50% below the median, and 75% below Q3.

  • To estimate how many observations exceed a given value, determine where that value falls relative to Q1, median, Q3, and use the corresponding percentage.

  • For example, if a value falls at approximately Q3, then roughly 25% of observations exceed it. With n = 64, that is about 16 observations.


Formulas / Diagrams

  • IQR = Q3 - Q1

  • Lower inner fence: IFL = Q1 - 1.5 · IQR

  • Upper inner fence: IFU = Q3 + 1.5 · IQR

  • Lower outer fence: OFL = Q1 - 3 · IQR

  • Upper outer fence: OFU = Q3 + 3 · IQR

  • Sample mean: x-bar = (1/n) · sum(x_i)


Real-World Applications

Boxplots are used in software engineering to compare the performance of different algorithms or systems. A data centre might plot response times for several server configurations side by side: a configuration with a lower median and tighter box is more consistently fast, even if another configuration occasionally produces a very fast response. Outliers in this context represent unusually slow requests that could indicate a bug or resource contention.


Common Misconceptions

  • Students often list IQR as a "measure of central tendency." IQR measures spread (the width of the middle 50%), not centre. The median is the measure of central tendency that is resistant to outliers.

  • Students sometimes think whiskers in a modified boxplot always extend to the minimum and maximum. They extend only to the most extreme data point within the inner fences. Points beyond the fences are plotted separately.

  • Students confuse "the inner fence value" with "the whisker endpoint." The whisker reaches the nearest actual data point inside the fence, which may be well inside the calculated fence value.

  • Students sometimes think descriptive statistics and inferential statistics are the same. Descriptive summarises observed data; inferential draws conclusions about the population. Probability goes in the other direction: known population to predicted samples.


Why It Matters / Exam Flags

⚠️ The median is the measure of central tendency least influenced by outliers, not the IQR (which is a measure of spread).

⚠️ x-bar is a sample statistic. It is not a population average and it is not a measure of spread.

⚠️ Inner fences are computed, not read directly from the boxplot. You must calculate them from Q1 and Q3.

⚠️ In a modified boxplot, asterisks typically denote mild outliers and solid dots denote extreme ("real") outliers.

⚠️ Probability is the correct tool when you know the population model and want to predict sample behaviour. Inferential statistics goes the other direction.


Quick Self-Test

  1. True or false: IQR is a measure of central tendency.

  1. Fill in the blank: The lower inner fence equals Q1 minus ________ times the IQR.

  1. True or false: x-bar is a population parameter.

  1. If a distribution is right-skewed, which is larger: the mean or the median?

  1. True or false: In a modified boxplot, the whiskers always extend to the minimum and maximum.

Answers: 1. False (it measures spread). 2. 1.5. 3. False (it is a sample statistic). 4. The mean. 5. False (they extend to the most extreme values within the inner fences).


Practice Q&A

Q: Which measure of central tendency is least likely to be influenced by outliers: sample median, sample mean, IQR, or sample standard deviation?

A: Sample median. It depends only on the middle position of the sorted data, not on extreme values. IQR and standard deviation are measures of spread, not central tendency.

Q: Which of the following is a correct description of x-bar? (a) sample statistic, (b) measure of spread, (c) population average, (d) more than one, (e) none.

A: (a). x-bar is a sample statistic. It is a measure of centre (not spread) computed from sample data (not the population).

Q: In a modified boxplot with 64 observations, approximately how many observations lie above Q3?

A: About 25% of 64 = 16 observations.

Q: A boxplot has its median line noticeably closer to Q1 than to Q3, and the upper whisker is much longer than the lower whisker. What is the shape?

A: Right-skewed (positively skewed). The longer upper whisker and the median's position closer to Q1 both indicate that the distribution stretches further to the right.

Q: Which tool do you use when you know the population model and all parameters and want to gain insight into possible samples: descriptive statistics, data collection, probability, or inferential statistics?

A: Probability. Starting from a known population and predicting sample behaviour is the domain of probability. Inferential statistics does the reverse, going from sample to population.

Q: A software developer compares four algorithms using boxplots of running time. One algorithm (new) has a noticeably lower median and a shorter box than the other three. Why might it be advantageous?

A: The lower median indicates it typically runs faster, and the shorter box (smaller IQR) indicates its running times are more consistent. Together, these suggest the new algorithm is both faster on average and more predictable.


Connections to Other Topics

  • Boxplots and five-number summaries connect directly to percentiles and quantiles, which appear in distribution problems (e.g., "find the 0.1 percentile" in the continuous PDF/CDF topic).

  • The distinction between sample statistics and population parameters is the foundation of sampling distributions and the Central Limit Theorem, the next major exam topic.

  • Understanding skewness is essential for knowing when the CLT requires a larger sample size to produce an approximately normal sampling distribution.


Related Terms / Search Tags

boxplot, modified boxplot, box-and-whisker plot, five-number summary, Q1, Q3, IQR, interquartile range, inner fence, outer fence, mild outlier, extreme outlier, median, mean, sample statistic, population parameter, x-bar, measures of centre, measures of spread, skewness, right-skewed, left-skewed, descriptive statistics, inferential statistics, probability, STAT 350 Purdue, intro statistics