Source: Introduction to Statistics, Purdue University, Chapters 1 & 2
Difficulty: Introductory | Prerequisites: None
These two chapters lay the groundwork for everything else in the course. Chapter 1 introduces the language of statistics: what populations and samples are, what variables look like, and how to classify data. Chapter 2 moves into how you organise and display that data using frequency distributions, charts and histograms. If you can nail these definitions and classification rules now, the rest of the course builds on them cleanly.
Statistics starts by defining what you are studying (the population) and what slice of it you actually measure (the sample). The characteristics you measure are variables, and they split into two families: categorical (qualitative) and numerical (quantitative), with numerical variables further divided into discrete and continuous. Once you have data, you summarise it with frequency distributions and display it with pie charts, bar graphs or histograms, each suited to a different variable type.
Population
The entire collection of individuals or objects to be considered or studied. Think of it as the full group you care about, even if you never measure every member.
Sample
A subset of the entire population, a small selection of individuals or objects taken from the entire collection. In simple terms, this is the slice of the population you actually collect data from.
Variable
A characteristic or attribute that can take on different values across the individuals in a study. Think of it as the thing you are measuring or recording about each subject.
Quantitative variable (numerical)
A variable whose values are numbers that represent measurable quantities. In simple terms, you can do arithmetic with these values (add, average, etc.).
Qualitative variable (categorical)
A variable whose values represent categories or labels rather than quantities. Think of it as data that sorts subjects into groups, not onto a number line.
Discrete data
Numerical data that is finite and countable, typically the result of counting. In simple terms, you can list out every possible value (0, 1, 2, 3...).
Continuous data
Numerical data that falls on intervals and is the result of measuring. Think of it as values that can take any number within a range, including decimals.
Univariate data
Data that involves only one variable per observation.
Frequency distribution
A summary that organises data by listing each category (or class) alongside how often it occurs.
Relative frequency
The frequency of a class divided by the total count. It tells you what proportion of the dataset falls into each category.
Class (or label)
The category or bin into which data values are grouped within a frequency distribution.
Histogram
A bar-style chart for quantitative variables where bars touch (no gaps), representing either discrete counts or continuous intervals.
Unimodal
A distribution with a single peak.
Bimodal
A distribution with two distinct peaks.
Multimodal
A distribution with three or more peaks.
Symmetric distribution
A distribution where the left and right sides are roughly mirror images.
Positively skewed (right-skewed)
A distribution with a longer tail stretching to the right. Think of it as the bulk of the data sitting on the left, with outliers pulling the tail rightward.
Negatively skewed (left-skewed)
A distribution with a longer tail stretching to the left. The bulk of the data sits on the right.
Population: the entire group you want to draw conclusions about.
Example: all undergraduate students at Purdue University.
Sample: a manageable subset drawn from the population.
Example: 200 randomly selected Purdue undergraduates.
The whole point of inferential statistics (later in the course) is to use sample data to make claims about the population, so keeping these two concepts distinct is essential from day one.
Categorical (qualitative): values are labels or categories, not numbers you can average.
Examples: eye colour, major, yes/no responses.
Numerical (quantitative): values are numbers that represent measurable quantities.
Discrete: finite, countable values. You count them.
Examples: number of siblings, number of courses this semester.
Continuous: values fall within intervals. You measure them.
Examples: height, weight, temperature.
The quick test: if "average" makes sense for the values, the variable is quantitative. If it does not (you cannot average "blue" and "green"), it is categorical.
A frequency distribution organises raw data into classes (categories or bins) and records how many observations fall into each.
Class (label): the category or interval grouping the data.
Frequency: the count of observations in that class.
Relative frequency: the frequency divided by the total count. Gives you a proportion (or percentage) rather than a raw count.
Relative frequency = frequency / total count
Frequency distributions are the bridge between raw data and visual displays. You build the table first, then choose the right chart.
Pie charts are for categorical variables.
Slices sized by counts or percentages.
Best when you have a small number of categories and want to show parts of a whole.
Bar graphs are for categorical (qualitative) variables.
Bar heights represent counts or percentages per category.
Bars do not touch (gaps between them signal discrete categories).
Histograms are for quantitative variables.
Bars do touch (no gaps), because the data is on a continuous or countable number line.
Used for both discrete (finite data) and continuous distributions.
For continuous data, the recommended number of classes is the square root of the number of observations.
When you look at a histogram or a density curve, you are describing two things: modality (how many peaks) and skewness (where the tail stretches).
Modality
Unimodal: one peak. The most common shape you will encounter.
Bimodal: two distinct peaks. Often signals two subgroups in the data.
Multimodal: three or more peaks.
Skewness
Symmetric: left and right sides mirror each other. The mean and median are roughly equal.
Positively skewed (right-skewed): tail extends to the right. The mean is pulled higher than the median.
Negatively skewed (left-skewed): tail extends to the left. The mean is pulled lower than the median.
Tail weight
A light-tailed distribution has values that cluster tightly around the centre.
A heavy-tailed distribution has more extreme values, further from the centre.
Relative frequency
Relative frequency = frequency / total count
Number of classes for a continuous histogram
Number of classes = square root of the number of observations
Variable classification tree (from the source notes)
Univariate data splits into two branches:
Categorical data (labels, groups)
Numerical data, which splits again into:
Discrete data (finite, countable, you count)
Continuous data (intervals, you measure)
Population vs. sample: polling companies survey a sample of voters (a few thousand) to predict election outcomes for an entire country (the population). The quality of the sample determines the quality of the prediction.
Variable types: a hospital records both categorical variables (blood type, diagnosis) and numerical variables (heart rate, blood pressure). Knowing the type determines which statistical test you run later.
Histograms and skewness: income distributions are classically right-skewed, meaning most people earn near the median while a small number earn far more. This is why median household income is reported more often than mean income.
Students often think that any variable with numbers is quantitative. It is not. Postcodes, phone numbers and student ID numbers are categorical, because averaging them is meaningless.
Students confuse discrete and continuous by focusing on whether the number has decimals. The real distinction is count vs. measure. A shoe size of 10.5 is still discrete (finite set of possible values); temperature of 10.5 degrees is continuous.
Students assume bar graphs and histograms are the same chart. They are not. Bar graphs have gaps (categorical data); histograms have no gaps (numerical data on a continuous scale).
"Right-skewed" trips people up because the bulk of data is on the left. The name refers to where the tail goes, not where the data clusters.
⚠️ Expect classification questions: "Is this variable quantitative or qualitative? Discrete or continuous?" These appear on nearly every intro stats exam.
⚠️ Know which chart goes with which variable type. Matching pie charts to categorical data and histograms to quantitative data is a staple exam question.
⚠️ Be ready to compute relative frequency from a frequency table. It is a quick calculation but easy marks if you remember the formula.
⚠️ Distribution shape descriptions (unimodal, bimodal, skewed left/right) are commonly tested with a histogram image. Practise identifying the shape at a glance.
True or false: a sample is always smaller than the population. (True)
Fill in the blank: a variable whose values are labels or categories is called a ______ variable. (qualitative / categorical)
True or false: in a histogram, bars have gaps between them. (False, that is a bar graph)
Fill in the blank: relative frequency equals ______ divided by ______. (frequency divided by total count)
True or false: a positively skewed distribution has its tail on the left side. (False, the tail extends to the right)
Q: A researcher records the hair colour of 500 university students. Is hair colour a quantitative or qualitative variable?
A: Qualitative (categorical). Hair colour is a label, not a number you can average.
Q: You roll a die 60 times and record the outcome each time. Is the outcome discrete or continuous?
A: Discrete. The possible values are 1, 2, 3, 4, 5, 6, a finite countable set.
Q: In a frequency table, one class has a frequency of 12 and the total number of observations is 80. What is the relative frequency for that class?
A: 12 / 80 = 0.15, or 15%.
Q: You are given a histogram that has a single peak with a long tail stretching to the right. Describe the distribution.
A: The distribution is unimodal and positively skewed (right-skewed).
Q: Which type of chart would you use to display the distribution of a continuous quantitative variable: a pie chart, a bar graph, or a histogram?
A: A histogram. Pie charts and bar graphs are for categorical variables.
Q: A dataset of exam scores has two clear peaks. What term describes this distribution shape?
A: Bimodal. The two peaks suggest there may be two distinct subgroups in the data (for example, students who studied and students who did not).
This connects to descriptive statistics (Chapter 3 onwards) because measures like the mean, median and standard deviation only make sense once you know whether you are working with quantitative data, and whether that data is discrete or continuous.
Understanding skewness here prepares you for later discussions about when the mean is a reliable measure of centre versus when the median is more appropriate. In a skewed distribution, the mean gets pulled toward the tail, which is why you will see the median used for things like household income.
Frequency distributions and histograms are also the foundation for probability distributions you will encounter later, particularly the normal distribution, which is symmetric and unimodal.
population vs sample, qualitative vs quantitative, categorical variable, numerical variable, discrete vs continuous, univariate data, frequency distribution, relative frequency, class interval, pie chart, bar graph, histogram, number of classes formula, unimodal, bimodal, multimodal, symmetric distribution, positively skewed, negatively skewed, right-skewed, left-skewed, light tail, heavy tail, intro to statistics, STAT Purdue, Chapter 1 statistics, Chapter 2 statistics, types of variables, data classification, frequency table