Sampling Distributions and Study Design – STAT 101 Ch. 7 and 1.3 – Study Notes
offline

Source: Purdue University, Introduction to Statistics

Tags: sampling distribution, sample mean, standard error, Central Limit Theorem, CLT, parameter, statistic, observational study, experiment, experimental design, randomization, replication, control, confounding variable, lurking variable, simple random sampling, stratified sampling, matched pair, block design, bias, convenience sample, undercoverage, nonresponse

Difficulty: Intermediate Prerequisites: Chapter 6 (normal distribution, z-table). You need to be comfortable standardising with z = (x − μ) / σ and looking up probabilities.


Big Picture

Chapter 7 answers the question: if I take many samples from the same population and compute the sample mean each time, what does the distribution of those means look like? The answer, given by the Central Limit Theorem, is "approximately normal," and that single fact is what makes most of inferential statistics possible. Chapter 1.3 steps back to ask how data should be collected in the first place, covering the design principles that determine whether your conclusions are trustworthy. Together, these sections connect the "probability machinery" of Chapters 4–6 to the real practice of drawing conclusions from data.


TL;DR

The sampling distribution of the sample mean has the same mean as the population but a smaller standard deviation (σ/√n). The Central Limit Theorem says this distribution is approximately normal for large n, regardless of the population shape. Meanwhile, good study design (randomisation, control, replication) ensures that the data feeding into these methods are worth analysing.


Key Terms

Parameter

A numerical value describing a characteristic of the population (e.g. μ, σ, p).

Statistic

A numerical value computed from a sample that estimates a parameter (e.g. x̄, s, p̂).

Sampling distribution

The probability distribution of a statistic across all possible samples of a given size from a population.

In simple terms, it is what you would see if you could take every possible sample and plot the statistic each time.

Standard error of the sample mean

σ / √n. This is the standard deviation of the sampling distribution of x̄. It measures how much x̄ varies from sample to sample.

Think of it as the "typical" distance of a sample mean from the true population mean.

Central Limit Theorem (CLT)

For a sufficiently large sample size n, the sampling distribution of x̄ is approximately normal with mean μ and standard deviation σ/√n, regardless of the shape of the population distribution.

In simple terms, averages of large samples behave like a normal distribution, even when the underlying data do not.

Observational study

A study where the researcher observes and records data without imposing any treatment or intervention. Can identify associations but cannot establish causation.

Experiment

A study where the researcher deliberately imposes a treatment on subjects and observes the response. When properly designed, experiments can establish cause-and-effect relationships.

Anecdotal study/evidence

Evidence based on individual cases or personal stories rather than systematic data collection. It is not reliable for drawing general conclusions.

Experimental units (subjects)

The individuals or items on which the experiment is performed.

Response variable

The outcome variable being measured or observed in the experiment.

Explanatory variable (factor/treatment)

The variable that the researcher manipulates or categorises to observe its effect on the response variable.

Levels

The specific values or categories of a factor in an experiment.

Control

Holding other variables constant or using a comparison group to isolate the effect of the treatment.

Randomisation

Randomly assigning experimental units to treatment groups to eliminate systematic differences between groups.

Replication

Applying each treatment to enough experimental units to detect meaningful differences and reduce the effect of chance variation.

Simple random sampling (SRS)

Every possible sample of size n has an equal chance of being selected.

Stratified sampling

The population is divided into homogeneous subgroups (strata), and a random sample is drawn from each stratum.

Matched pair design

A design where subjects are paired based on some characteristic, and one member of each pair receives each treatment. Alternatively, each subject receives both treatments in random order.

Block design

Subjects are grouped into blocks of similar individuals, and treatments are randomly assigned within each block. Reduces variability due to known sources of difference.

Bias

Systematic error that consistently pushes results in one direction, making the sample unrepresentative of the population.

Convenience sample

A sample chosen because it was easy to obtain, not because it was representative. Prone to bias.

Undercoverage

When some members of the population have no chance of being included in the sample.

Nonresponse

When selected individuals do not participate, and those who do not respond differ systematically from those who do.

Lurking variable

A variable not included in the study that affects both the explanatory and response variables. Can create a misleading association.

Confounding variable

A variable whose effect on the response variable cannot be separated from the effect of the explanatory variable. In simple terms, you cannot tell which variable is responsible for the observed result.


Core Content

Sampling Distributions (Chapter 7)

Parameters vs. Statistics

  • A parameter is a fixed (often unknown) number describing the population.

  • A statistic is a number computed from a sample; it changes from sample to sample.

  • The sampling distribution shows how the statistic behaves across repeated sampling.

Sampling Distribution of the Sample Mean

  • Mean of x̄: E(x̄) = μ (the sample mean is an unbiased estimator of the population mean).

  • Standard deviation of x̄: σ_x̄ = σ / √n.

  • As n increases, σ_x̄ decreases, meaning sample means cluster more tightly around μ.

When the Population Is Normal

  • If the population distribution is normal, then the sampling distribution of x̄ is exactly normal for any sample size: x̄ ~ N(μ, σ/√n).

  • You can compute probabilities by standardising: z = (x̄ − μ) / (σ/√n).

Central Limit Theorem

  • When the population is not normal (or its shape is unknown), the CLT says the sampling distribution of x̄ is still approximately normal, provided n is "large enough."

  • A common rule of thumb: n ≥ 30 is usually sufficient, though the more skewed the population, the larger n needs to be.

  • Once the CLT applies, use z = (x̄ − μ) / (σ/√n) to compute probabilities, just as in the normal-population case.

Study Design (Chapter 1.3)

Observational Studies vs. Experiments

  • Observational study: no intervention; the researcher merely observes existing conditions. Can show association but not causation.

  • Experiment: the researcher assigns treatments to subjects. With proper design, experiments can demonstrate causation.

  • Anecdotal evidence: individual stories or examples. Not a basis for statistical conclusions.

Identifying Components of an Experiment

  • Experimental units/subjects: who or what is being studied.

  • Response variable: the outcome you measure.

  • Explanatory variable/factor: what you change or compare.

  • Levels: the specific values of the factor.

  • Treatment: a specific combination of factor levels applied to an experimental unit.

Principles of Experimental Design

  • Control: use a control group or hold extraneous variables constant.

  • Randomisation: randomly assign units to treatments to prevent systematic bias.

  • Replication: use enough units per treatment to ensure results are not due to chance.

Sampling Methods

  • Simple random sampling (SRS): every sample of size n is equally likely. The gold standard.

  • Stratified sampling: divide into strata, then SRS within each. Useful when strata differ from each other but are internally homogeneous.

  • Matched pair: pair similar units and assign one treatment per pair, or use each subject as their own control.

Sampling Problems

  • Bias: systematic error. The sample does not represent the population.

  • Convenience sample: selecting whoever is easiest to reach. Almost always biased.

  • Undercoverage: some population members are excluded from the sampling frame.

  • Nonresponse: selected individuals refuse or fail to participate, and they differ from respondents.

Matched Pair and Block Designs

  • A matched pair design is appropriate when subjects can be meaningfully paired on a relevant characteristic, or when each subject can serve as their own control (e.g. before/after measurements).

  • A block design extends this logic: group subjects into blocks that are similar on some variable, then randomise within each block. It reduces variability from known sources and increases the experiment's sensitivity to the treatment effect.

Lurking and Confounding Variables

  • A lurking variable is not part of the study but influences the relationship you observe. Example: ice cream sales and drowning rates both rise in summer, but heat (the lurking variable) drives both.

  • A confounding variable is entangled with the explanatory variable so that you cannot tell which one causes the response. Example: if all younger patients receive the new drug and all older patients receive the old drug, age confounds the treatment effect.

  • Randomisation is the primary tool for handling confounding.


Formulas / Diagrams

Sampling distribution of x̄: Mean: E(x̄) = μ Standard deviation: σ_x̄ = σ / √n

Standardising the sample mean: z = (x̄ − μ) / (σ / √n)


Real-World Applications

The CLT is why political polls work. A survey of 1,000 people can estimate the opinion of millions, because the sample mean (proportion) has a known, tight distribution around the population proportion. Study design principles matter equally: a drug trial that fails to randomise, or a survey conducted only via social media (convenience sample with undercoverage), produces results that no amount of statistical calculation can rescue. The design determines whether the data are worth analysing at all.


Common Misconceptions

  • Students often confuse the standard deviation of the population (σ) with the standard error of the mean (σ/√n). The standard error is always smaller and shrinks with increasing n.

  • The CLT does not say the data become normal. It says the distribution of the sample mean becomes approximately normal. The raw data keep their original shape.

  • Students sometimes think observational studies are "bad." They are not; they are simply limited in that they cannot establish causation. Many important studies (epidemiology, economics, astronomy) are necessarily observational.

  • Confusing confounding with lurking: a confounding variable is mixed up with the treatment, making it impossible to separate their effects. A lurking variable is simply unobserved. A lurking variable becomes confounding when it is associated with both the explanatory and response variables and is not controlled for.


Why It Matters / Exam Flags

⚠️ Expect a question that asks you to compute a probability involving the sample mean, using z = (x̄ − μ) / (σ/√n).

⚠️ You will need to state whether the CLT applies in a given scenario and explain why (large enough n, or population already normal).

⚠️ Distinguishing between observational studies and experiments is directly tested, including whether causal conclusions are warranted.

⚠️ Identifying sampling method (SRS, stratified, matched pair) and sampling problems (bias, convenience, undercoverage, nonresponse) from a scenario description is likely.

⚠️ You may be asked to draw or describe an experimental design, including identifying experimental units, factors, levels, and response variable.

⚠️ Defining and distinguishing lurking vs. confounding variables is a common exam question.


Quick Self-Test

  1. True or False: The standard error of the sample mean increases as the sample size increases.

  1. Fill in the blank: The Central Limit Theorem says the sampling distribution of x̄ is approximately __________ for large n.

  1. True or False: An observational study can establish a cause-and-effect relationship.

  1. Fill in the blank: In a __________ sample, every possible sample of size n has an equal chance of being selected.

  1. True or False: A confounding variable is one whose effect cannot be separated from the explanatory variable's effect.

Answers: 1. False (it decreases). 2. Normal. 3. False (only association). 4. Simple random. 5. True.


Practice Q&A

Q: A population has μ = 50 and σ = 10. A sample of size n = 25 is drawn. What is the standard deviation of the sampling distribution of x̄?

A: σ_x̄ = 10 / √25 = 10 / 5 = 2.

Q: Using the values above, what is P(x̄ > 53)?

A: z = (53 − 50) / 2 = 1.5. From the z-table, P(Z ≤ 1.5) ≈ 0.9332. So P(x̄ > 53) = 1 − 0.9332 = 0.0668.

Q: A researcher surveys people at a shopping centre about their spending habits. What type of study is this, and what sampling problem is most likely present?

A: This is an observational study. The most likely sampling problem is convenience sampling (and possibly undercoverage), since only people who happen to be at the shopping centre are surveyed.

Q: In a study comparing two fertilisers, one researcher applies Fertiliser A to all plots on a sunny hillside and Fertiliser B to all plots in a shady valley. Identify the confounding variable.

A: Sunlight exposure is the confounding variable. It is entangled with the fertiliser treatment, so you cannot tell whether differences in plant growth are due to the fertiliser or the sunlight.

Q: Why is a block design sometimes preferred over a completely randomised design?

A: A block design groups subjects by a variable known to affect the response (e.g. age, location), then randomises within each block. This reduces variability from that known source, making it easier to detect the treatment effect.


Connections to Other Topics

  • The sampling distribution of x̄ is built directly on the normal distribution from Chapter 6 and the expected-value/variance rules from Chapter 5.

  • Study design (Chapter 1.3) determines whether the probability models in Chapters 4–6 are even applicable: biased data invalidates the assumptions behind inference.

  • The CLT connects to every inferential method you will meet after this exam: confidence intervals and hypothesis tests for means and proportions all rely on the normality of the sampling distribution.


Related Terms / Search Tags

sampling distribution, sample mean, x-bar, standard error, sigma over root n, Central Limit Theorem, CLT, parameter vs statistic, observational study, experiment, causation, association, anecdotal evidence, experimental units, response variable, explanatory variable, factor, levels, treatment, control group, randomisation, replication, simple random sampling, SRS, stratified sampling, matched pair design, block design, bias, convenience sample, undercoverage, nonresponse, lurking variable, confounding variable, common response, experimental design, STAT 101, intro to statistics, Purdue