Difficulty: Introductory–Intermediate | Prerequisites: None beyond general familiarity with data collection.
Big Picture
Before you can do any inference, you need data, and the way you collect data determines what conclusions are valid. This chapter covers the difference between experiments and observational studies, how to design experiments properly, how to sample a population and what can go wrong. It also introduces lurking variables and Simpson's paradox, which explain why correlation does not imply causation. You do not need heavy maths here, but the concepts are tested constantly on exams.
Experiments impose conditions on subjects and can establish causation. Observational studies just measure what is already happening and cannot prove causation. Good experiments use control, randomisation and replication. Good sampling avoids bias through methods like simple random sampling (SRS) or stratified sampling. Lurking variables can create misleading associations.
Anecdote / anecdotal evidence
A short story about a single incident or person. Anecdotes draw attention because they are unusual, but they cannot serve as a basis for decisions without formal study.
Think of it as one data point that got famous because it was interesting, not because it was representative.
Experiment
A study in which the researcher deliberately imposes conditions (treatments) on subjects and measures the response. Because conditions are controlled, experiments are the strongest way to establish cause and effect.
Observational study
A study in which the researcher measures variables of interest without attempting to influence them. Often used when experiments would be unethical or impractical (e.g. studying the effect of smoking on lung cancer).
Experimental unit / subject
The individual person, animal or object on which a treatment is applied.
Factor
The explanatory variable whose effect the experiment is designed to test. In simple terms, it is what differentiates the groups.
Level
A specific value or category of a factor. For instance, if the factor is "type of drug," the levels might be "drug A" and "placebo."
Response variable
The outcome that is measured in the experiment. This is what you are interested in seeing change.
Completely randomised design
All subjects are assigned to treatments entirely at random.
Matched pair design
Each subject is paired with another similar subject, and one member of each pair receives each treatment. This removes the effect of extraneous variables that the pair shares.
Block design
An extension of matched pairs. Subjects are grouped into blocks of similar individuals, and treatments are randomly assigned within each block. The only systematic difference between blocks is the variable of interest.
Simple random sampling (SRS)
A sampling method in which every possible sample of size n has an equal chance of being selected. The foundational method for probability-based sampling.
Think of it as drawing names from a hat where every combination of n names is equally likely.
Stratified random sampling
The population is divided into groups (strata), and an SRS is drawn from each group. The sample sizes from each stratum need not be equal.
Convenience sample
A sample chosen because it is easy to obtain, not because it is representative. A major source of bias.
Self-selection bias
Occurs when subjects volunteer to participate. Volunteers tend to hold stronger opinions and are not representative of the whole population.
Undercoverage
Occurs when some groups in the population are left out of the sampling frame entirely or are underrepresented.
Nonresponse
Occurs when selected participants do not complete the study or provide incomplete measurements, even though the original sample was properly randomised.
Lurking variable (extraneous variable / confounding variable)
A variable not included in the study that affects both the explanatory and response variables. It can create a misleading appearance of a direct relationship.
Simpson's paradox
A situation in which an association that holds within every subgroup reverses direction when the subgroups are combined. Caused by a lurking variable that is unevenly distributed across groups.
Double-blind experiment
Neither the subjects nor the experimenters know which treatment each subject receives. This reduces both placebo effects and experimenter bias.
In an experiment, the researcher imposes treatments. In an observational study, the researcher only measures what already exists.
Only a well-designed experiment can establish causation. Observational studies can identify associations but cannot rule out lurking variables.
Ethical constraints: you cannot assign people to smoke for 20 years to study lung cancer.
Practical constraints: you cannot control someone's height to study its relationship with weight.
Anecdotal evidence (e.g. the autism/MMR scare) can raise questions worth studying formally, but anecdotes alone are never sufficient grounds for a conclusion.
Every experiment has four components to identify:
Experimental unit (subject): who or what is being studied.
Factor: the explanatory variable being tested.
Levels: the specific values or categories of the factor (e.g. drug vs placebo).
Response variable: the outcome being measured.
Example (sickle cell study): 150 patients get hydroxyurea, 150 get placebo. Units = patients, factor = type of medicine, levels = drug or placebo, response = episodes of pain.
Completely randomised design: all subjects are allocated to treatments at random. Simplest approach, works well when the subjects are fairly homogeneous.
Matched pair design: each unit is paired with a similar unit, and one from each pair gets each treatment. Removes the effect of variables shared within the pair.
Block design: an extension of matched pairs to groups (blocks) of more than two. Within each block subjects are similar on key extraneous variables. Treatments are randomly assigned within blocks.
Control: hold everything constant except the treatment itself. Use a placebo or standard treatment for comparison.
Randomisation: assign subjects to treatments randomly to avoid systematic bias.
Replication: use enough subjects so that results are not driven by a few unusual individuals.
Blinding: single-blind (subjects do not know their treatment) or double-blind (neither subjects nor experimenters know) to reduce bias.
Label each of the N individuals.
Use a random number generator (or draw numbers from a hat).
Select n individuals without replacement.
Be careful about generalising beyond your experimental conditions. If your subjects are all 20-year-old university students, you cannot claim the result applies to all adults.
"Control what you can, block what you cannot control, and randomise to create comparable groups."
Simple random sampling (SRS): every sample of size n is equally likely. The baseline for all probability sampling. Method: label every unit, generate random integers, select the smallest n.
Stratified random sampling: divide the population into strata (groups), then draw an SRS within each stratum. Useful when subgroups differ meaningfully. Sample sizes from each stratum need not be equal.
Convenience sampling: pick whoever is easiest to reach. Fast but rarely representative.
Self-selection: subjects choose to participate (e.g. a voluntary online survey). Respondents tend to hold extreme views.
Convenience bias: sample is chosen for ease, not representativeness.
Self-selection bias: volunteers differ systematically from the population.
Undercoverage: some groups are missing from or underrepresented in the sample.
Nonresponse: selected participants fail to complete the study, which may distort results if nonrespondents differ from respondents.
Random sampling reduces bias.
Increasing sample size reduces the variability of a statistic from an SRS.
Experiments can establish causation because the researcher controls the variables.
Observational studies cannot establish causation because lurking variables may explain the observed association.
A lurking variable is not measured in the study but influences both the explanatory and response variables. Classic examples:
Shoe size and maths scores in children. Lurking variable: age. Older children have bigger feet and better maths skills.
Ice cream sales and drowning deaths. Lurking variable: temperature/season. Hot weather drives both.
Number of churches and number of bars in a town. Lurking variable: town size. More people means more of both.
Global temperature and the number of pirates. The lurking variable is simply time: temperatures have risen while piracy has declined.
Causation: x directly causes y.
Common response: a lurking variable z causes both x and y.
Confounding: z may affect both x and y, but you cannot untangle x's direct effect from z's.
An association that holds in every subgroup reverses when the subgroups are combined, because a lurking variable is unevenly distributed.
Example: Women may have a higher acceptance rate than men in every individual department, yet a lower overall acceptance rate. This happens when women disproportionately apply to the more competitive departments.
Students often assume that a strong correlation means one variable causes the other. Correlation does not imply causation, particularly in observational studies where lurking variables may be at work.
Students sometimes think a large sample fixes bad sampling. A convenience sample of 10,000 people is still biased; size reduces variability, not bias. Only randomisation reduces bias.
Students confuse "randomised experiment" with "random sample." A randomised experiment randomly assigns treatments to subjects. A random sample randomly selects subjects from a population. Both use randomness, but for different purposes.
Students often mix up the terms "factor" and "level." The factor is the explanatory variable itself (e.g. drug type), and the levels are its categories (e.g. aspirin, ibuprofen, placebo).
⚠️ Be able to classify any study as an experiment or an observational study and explain why.
⚠️ Given a study description, identify the experimental units, factor(s), levels and response variable.
⚠️ Know the three experimental designs (completely randomised, matched pair, block) and when each is preferred.
⚠️ Identify the type of bias in a sampling scenario (convenience, self-selection, undercoverage, nonresponse).
⚠️ Given a correlation, suggest a plausible lurking variable and explain the relationship structure (causation, common response, confounding).
⚠️ Simpson's paradox examples are a favourite exam question. Be ready to explain how aggregating data can reverse a trend.
True or False: An observational study can establish causation if the sample is large enough. (False)
Fill in the blank: In an experiment, the _____ is what differentiates the treatment groups. (Factor)
True or False: Stratified random sampling gives every individual the same probability of selection. (False, probabilities can differ across strata.)
True or False: Double-blinding means neither the subjects nor the researchers know who receives which treatment. (True)
Fill in the blank: A variable not in the study that affects both the explanatory and response variables is called a _____ variable. (Lurking / confounding / extraneous)
Q: A researcher wants to know whether a new fertiliser increases tomato yield. She applies the fertiliser to 30 randomly chosen plots and leaves 30 other plots untreated. Is this an experiment or an observational study?
A: Experiment. The researcher deliberately imposed the treatment (fertiliser vs no fertiliser) on the plots.
Q: A survey finds that people who drink red wine tend to have lower rates of heart disease. Can you conclude that red wine prevents heart disease?
A: No. This is an observational study. Lurking variables (such as overall diet, exercise, income) may explain the association.
Q: In a clinical trial, patients are divided into age groups (under 40, 40–60, over 60), and within each group, half are randomly assigned to the drug and half to placebo. What experimental design is this?
A: Block design. The blocks are the age groups, and treatments are randomly assigned within each block.
Q: A university emails a survey to all 20,000 students but only 800 respond. Name the source of bias.
A: This is self-selection bias (and nonresponse). The 800 who responded chose to participate and are likely not representative of the full student body.
Q: In a certain city, neighbourhoods with more police officers also have more crime. Does this mean police cause crime?
A: No. A lurking variable (existing crime rate) is likely at work. More officers are deployed to areas that already have higher crime.
This connects to sampling distributions (Ch. 7) because the validity of a sampling distribution depends on proper random sampling. Bias in data collection undermines every inference technique that follows.
This connects to hypothesis testing (Ch. 10) because the logic of a hypothesis test mirrors the structure of an experiment: H₀ is the "no effect" baseline, and Hₐ is the treatment effect you are trying to detect.
Lurking variables and confounding are revisited whenever you interpret a two-sample comparison (Ch. 11) or regression in later courses.
experiment vs observational study, randomised controlled trial, RCT, experimental design, completely randomised design, matched pairs, block design, SRS, simple random sampling, stratified sampling, convenience sampling, self-selection, undercoverage, nonresponse bias, lurking variable, confounding variable, extraneous variable, Simpson's paradox, causation vs correlation, double-blind, single-blind, placebo, control group, STAT 302, Purdue statistics