Foundations of Statistics, Data Types, and Distributions – STAT 101 Ch. 1–2 – Study Notes
offline

Source: Purdue University, Introduction to Statistics

Tags: statistics branches, data collection, descriptive statistics, inferential statistics, population, sample, parameter, statistic, Greek letter, Latin letter, categorical, quantitative, discrete, continuous, histogram, distribution shape, skewness, unimodal, bimodal

Difficulty: Introductory Prerequisites: None. This is the starting point for the entire course.


Big Picture

Statistics is the science of collecting, organising, analysing, and interpreting data to make decisions. This opening material lays the groundwork for everything that follows: you need to know what kind of data you are working with before you can choose the right tools to summarise or draw conclusions from it. Chapters 1 and 2 answer two foundational questions: "What are we doing with data?" and "What kind of data do we have?" If you are behind, start here, because every later chapter assumes you can identify variable types and understand the difference between a population and a sample.


TL;DR

Statistics splits into three branches: collecting data, describing data, and making inferences from data. The data itself can be categorical or quantitative, and quantitative data can be discrete or continuous. Knowing which type you have determines which summary measures and visualisations you should use.


Key Terms

Data collection

The branch of statistics concerned with obtaining data through surveys, experiments, or observational studies.

In simple terms, this is the "gathering" phase, before any analysis happens.

Descriptive statistics

Methods for organising, summarising, and presenting data in an informative way (tables, graphs, numerical summaries).

Think of it as taking a pile of raw numbers and turning them into something a human can read at a glance.

Inferential statistics

Methods that use sample data to draw conclusions or make predictions about a larger population.

In simple terms, this means using what you know from a subset to say something about the whole.

Population

The complete collection of all individuals or items of interest. Population parameters are denoted with Greek letters (e.g. μ for mean, σ for standard deviation).

Think of it as "everyone or everything you care about" in a study.

Sample

A subset of the population that is actually observed or measured. Sample statistics are denoted with Latin (Roman) letters (e.g. x̄ for sample mean, s for sample standard deviation).

In simple terms, the sample is who or what you could practically get data from.

Parameter

A numerical summary of a population (Greek letter).

Statistic

A numerical summary of a sample (Latin letter).

Probability

A branch of mathematics that measures the likelihood of events occurring. In this course, probability serves as the bridge from descriptive statistics to inferential statistics: you use known population information to predict what samples might look like.

Inferential statistics (in contrast to probability)

The reverse direction: using sample data to make claims about an unknown population.

Univariate data

Data that involves a single variable per observation.

Bivariate data

Data that involves two variables per observation (e.g. height and weight for each person).

Multivariate data

Data that involves three or more variables per observation.

Categorical variable (qualitative)

A variable whose values are labels or categories with no natural numerical ordering. Examples: eye colour, car brand, pass/fail.

In simple terms, you can name it but you cannot meaningfully average it.

Quantitative variable

A variable whose values are numbers on which arithmetic operations make sense. Examples: height, temperature, number of siblings.

Discrete variable

A quantitative variable that takes a countable number of values, typically whole numbers. Example: number of children in a household.

Think of it as something you count.

Continuous variable

A quantitative variable that can take any value within an interval. Example: time, weight, temperature.

Think of it as something you measure.

Histogram

A bar chart for quantitative data where each bar covers an interval (bin) and its height represents frequency or relative frequency. Bars touch because the variable is continuous or treated as continuous.

Unimodal

A distribution with one clear peak.

Bimodal

A distribution with two prominent peaks.

Multimodal

A distribution with three or more peaks.

Symmetric distribution

A distribution where the left and right sides are roughly mirror images.

Right-skewed (positively skewed)

The right tail is longer. Most data cluster on the left.

Think of it as being "pulled" towards large values.

Left-skewed (negatively skewed)

The left tail is longer. Most data cluster on the right.

Think of it as being "pulled" towards small values.


Core Content

Branches of Statistics

  • Data collection covers how data are obtained, including survey design and sampling.

  • Descriptive statistics covers summarising and displaying data (charts, averages, spreads).

  • Inferential statistics covers using sample results to make generalisations about the population.

  • A given scenario belongs to exactly one branch. Ask: "Are we gathering data? Summarising data? Or drawing a conclusion beyond the data?"

Population vs. Sample and Notation

  • The population is the entire group of interest; the sample is the observed portion.

  • Greek letters denote population values: μ (mean), σ (standard deviation), p (proportion in some contexts).

  • Latin letters denote sample values: x̄ (mean), s (standard deviation).

  • When a question lists a group of objects or people, decide: is this every member of the group of interest (population), or a subset chosen for study (sample)?

Probability vs. Inferential Statistics

  • Probability starts with a known population model and asks, "What will the sample look like?"

  • Inferential statistics starts with sample data and asks, "What can we say about the population?"

  • They are two sides of the same coin, running in opposite directions.

Identifying Variable Types

  • First ask: are the values categories/labels, or numbers you could do maths with?

    • Categories → categorical.

    • Numbers with meaningful arithmetic → quantitative.

  • If quantitative, ask: can the variable take only isolated values (counting), or any value in a range (measuring)?

    • Counting → discrete.

    • Measuring → continuous.

Interpreting Histograms

  • You will not be asked to draw one, but you will be asked to describe one.

  • Describe the shape: number of peaks (unimodal, bimodal, multimodal) and symmetry (symmetric, right-skewed, left-skewed).

  • Look for outliers: values that stand apart from the main body of data.

  • Note the centre (where the data cluster) and the spread (how far the data extend).


Real-World Applications

Knowing the variable type is not academic busywork. If a hospital records patient blood types, that is categorical data, and computing a mean blood type is meaningless. If they record patient ages, that is quantitative continuous data, and a histogram quickly reveals whether the patient population skews young or old. Choosing the wrong summary for the wrong data type produces nonsense.


Common Misconceptions

  • Students often think any variable with numbers is quantitative. It is not. Postcodes, phone numbers, and jersey numbers are categorical even though they look numerical, because arithmetic on them is meaningless.

  • Students sometimes confuse "discrete" with "categorical." Discrete variables are still quantitative; you can average the number of pets people own.

  • Calling a distribution "skewed right" when the long tail points left is a common mix-up. The skew is named for the direction of the tail, not the peak.

  • Students sometimes think probability and inferential statistics are the same thing. They go in opposite directions: probability predicts samples from known populations; inference estimates populations from observed samples.


Why It Matters / Exam Flags

⚠️ You will be asked to classify a scenario into one of the three branches of statistics. Read the scenario carefully: is it about gathering, summarising, or concluding?

⚠️ Expect a question asking whether a given list refers to a population or a sample. The entire group of interest is the population; a subset is the sample.

⚠️ Distinguishing categorical vs. quantitative and discrete vs. continuous is tested directly and also underpins later chapters on choosing summary measures and distributions.

⚠️ Interpreting histogram shape (peaks, symmetry, skewness, outliers) will appear. Practise describing shapes in words.


Quick Self-Test

  1. True or False: A sample statistic is denoted with a Greek letter.

  1. Fill in the blank: A variable that records the colour of cars in a car park is __________ (categorical / quantitative).

  1. True or False: A right-skewed distribution has a longer tail on the left side.

  1. Fill in the blank: Data that involves exactly two variables per observation is called __________ data.

  1. True or False: Inferential statistics uses sample data to make conclusions about a population.

Answers: 1. False (Latin letter). 2. Categorical. 3. False (longer tail on the right). 4. Bivariate. 5. True.


Practice Q&A

Q: A researcher records the heights and weights of 200 university students. Is this data univariate, bivariate, or multivariate?

A: Bivariate, because two variables (height and weight) are measured for each student.

Q: A news report states, "Based on a poll of 1,200 likely voters, the candidate leads by 5 points." Is this descriptive or inferential statistics?

A: Inferential statistics, because the poll (sample) is used to draw a conclusion about all likely voters (population).

Q: The number of apps installed on students' phones is recorded. Is this variable discrete or continuous?

A: Discrete, because the count of apps takes only whole-number values.

Q: A histogram of exam scores shows one peak near 75 and a long tail stretching towards lower scores. Describe the shape.

A: Unimodal and left-skewed (the long tail points towards the lower values).

Q: A survey records each respondent's favourite streaming service. What type of variable is this?

A: Categorical, because the values are labels (Netflix, Hulu, etc.) with no meaningful numerical ordering.


Connections to Other Topics

  • The population-vs.-sample distinction is the backbone of Chapter 7 (sampling distributions) and all of inferential statistics.

  • Variable types determine which summary statistics to use in Chapter 3 (mean and standard deviation for symmetric quantitative data; median and IQR for skewed or outlier-heavy data).

  • Distribution shape (skewness) reappears in Chapter 5 when you study the binomial distribution's skewness based on the value of p.


Related Terms / Search Tags

statistics branches, data collection, descriptive statistics, inferential statistics, population vs sample, parameter vs statistic, Greek letter notation, Latin letter notation, categorical variable, qualitative variable, quantitative variable, discrete variable, continuous variable, univariate, bivariate, multivariate, histogram interpretation, distribution shape, unimodal, bimodal, multimodal, symmetric, right skewed, left skewed, positively skewed, negatively skewed, outlier detection, STAT 101, intro to statistics, Purdue