Experimental Design and Inference Foundations, STAT 101 – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Descriptive Statistics and Probability notes

TL;DR

This set of notes covers how to evaluate whether an experiment is well designed (controls, randomisation, lurking variables) and the core mechanics of statistical inference: confidence intervals, hypothesis tests, p-values, Type I and Type II errors, and how to check whether the assumptions behind your test are valid. If you understand this material, you can read and critique nearly any basic statistical analysis.


Key Terms

Experimental design

The plan for how subjects are assigned to treatments and how data is collected. A good design minimises bias and confounding so that any observed effect can reasonably be attributed to the treatment.

Control group

The group that does not receive the treatment (or receives a placebo). It exists so you have a baseline to compare the treatment group against.

Randomisation

Assigning subjects to treatment and control groups at random. This balances out lurking variables, both known and unknown, across groups. Think of it as insurance against systematic bias.

Lurking variable (confounding variable)

A variable not included in the study that affects both the explanatory variable and the response variable, potentially creating a false appearance of a relationship. In simple terms, it is the hidden factor that could be the real explanation.

Simple random sample (SRS)

Every possible sample of size n has an equal chance of being selected. The gold standard for avoiding selection bias.

Stratified sampling

Divide the population into homogeneous subgroups (strata), then take an SRS from each. Useful when you want to ensure representation of key subgroups.

Cluster sampling

Divide the population into clusters (often geographic), randomly select entire clusters, and sample everyone within the chosen clusters. Cheaper when subjects are spread out, but less precise than stratified sampling.

Systematic sampling

Select every k-th individual from a list after a random start. Quick and simple, but can go wrong if the list has a hidden periodic pattern.

Confidence interval

A range of plausible values for a population parameter, constructed so that if you repeated the sampling process many times, a stated percentage (the confidence level) of those intervals would contain the true parameter. Think of it as a net, not a bullseye.

Confidence level

The long-run proportion of confidence intervals that would capture the true parameter value. A 95% confidence level means 95 out of 100 intervals will contain the truth, not that there is a 95% probability the truth is in this particular interval.

Critical value

The number of standard errors from the centre of the distribution that corresponds to the desired confidence level. For a 95% confidence interval using the normal distribution, the critical value is approximately 1.96. It sets how wide the interval is.

Standard error

The standard deviation of a sampling distribution. It measures how much a statistic (like the sample mean) would vary from sample to sample. In simple terms, it is the typical size of sampling error.

Type I error

Rejecting the null hypothesis when it is actually true. A false positive. The probability of a Type I error is alpha, the significance level.

Type II error

Failing to reject the null hypothesis when it is actually false. A false negative. The probability of a Type II error is beta.

Power

The probability of correctly rejecting a false null hypothesis. Power = 1 minus beta. Higher power means you are more likely to detect a real effect.

P-value

The probability of observing a test statistic as extreme as, or more extreme than, the one actually observed, assuming the null hypothesis is true. In simple terms, it measures how surprising your data would be if nothing were going on.

Null hypothesis (H0)

The default claim that there is no effect, no difference, or no relationship. It is the hypothesis you are trying to find evidence against.

Alternative hypothesis (Ha)

The claim that there is an effect, a difference, or a relationship. It is what you conclude if the evidence against H0 is strong enough.

Statistical significance

A result is statistically significant when the p-value is below the chosen significance level (alpha). It means the data are unlikely under the null hypothesis.

Practical significance

Whether a statistically significant result is large enough to matter in the real world. A drug that lowers blood pressure by 0.1 mmHg may be statistically significant with a large sample, but it is not practically meaningful.


Core Content: Experimental Design

What Makes a Good Experimental Design

A well-designed experiment has four features:

  • Randomisation: subjects are randomly assigned to groups, so differences between groups at the start are due to chance, not systematic bias.

  • Control: there is a comparison group (control or placebo) so that the effect of the treatment can be isolated.

  • Replication: enough subjects in each group that the results are not driven by one or two unusual individuals.

  • Blinding: ideally, neither the subjects nor the researchers measuring outcomes know who is in which group (double-blind). This prevents expectations from influencing results.

If a study lacks any of these, you should be able to explain what could go wrong.

Choosing the Right Sampling Procedure

The sampling method depends on the research question, the population structure, and practical constraints.

  • Simple random sample: best when the population is relatively homogeneous and a list of all members is available.

  • Stratified sample: best when the population has known subgroups (e.g. age brackets, regions) and you need accurate estimates within each subgroup or want to reduce variability.

  • Cluster sample: best when a complete list of individuals is unavailable but natural groupings exist (e.g. schools, city blocks). More practical, but tends to have higher sampling error.

  • Systematic sample: best for large, ordered lists where an SRS is logistically difficult. Risky if the list has a repeating pattern that aligns with the sampling interval.

  • Convenience sample: selecting whoever is easiest to reach. Almost never valid for inference because it introduces selection bias.

Lurking Variables

A lurking variable creates a spurious association between the explanatory and response variables.

How to spot one: ask yourself, "Is there something else that could explain both the explanatory variable and the response?" If yes, that is a lurking variable.

Classic example: ice cream sales and drowning rates are positively correlated. The lurking variable is temperature (hot weather). Ice cream does not cause drowning.

Randomisation is the primary defence against lurking variables in experiments. In observational studies, you cannot randomise, so lurking variables are always a concern.


Core Content: Inference Foundations

Probability vs. Statistical Inference

Probability works forward: given a known population, what is likely to happen in a sample? Inference works backward: given observed sample data, what can we conclude about the unknown population?

Probability is the engine that makes inference possible, because inference relies on probability models to quantify uncertainty.

Confidence Intervals

The general form of a confidence interval is:

point estimate ± (critical value) × (standard error)

  • The point estimate is the statistic from your sample (e.g. the sample mean).

  • The critical value sets the width and corresponds to the confidence level. For 95% confidence using the normal distribution, z* = 1.96. For smaller samples, a t* value from the t-distribution is used.

  • The standard error measures the variability of the estimate across samples.

Why confidence intervals matter: a single point estimate tells you nothing about precision. A confidence interval tells you how much that estimate could plausibly shift if you drew a different sample.

What a 95% confidence interval means: if you were to take 100 different samples and build 100 confidence intervals, about 95 of them would contain the true population parameter. It does not mean there is a 95% probability that the parameter falls in this specific interval.

Hypothesis Testing

The general procedure:

  1. State the null hypothesis (H0) and alternative hypothesis (Ha).

  1. Choose a significance level (alpha), typically 0.05.

  1. Compute a test statistic from the sample data.

  1. Find the p-value (the probability of getting a result this extreme or more, assuming H0 is true).

  1. Compare the p-value to alpha. If p-value < alpha, reject H0. If p-value >= alpha, fail to reject H0.

Why We Never "Accept" the Null

Failing to reject H0 is not the same as proving H0 is true. The data may simply not be strong enough to detect an effect. Just as a jury verdict of "not guilty" does not mean "innocent," failure to reject means insufficient evidence, not confirmation.

Type I and Type II Errors

  • Type I error (false positive): rejecting H0 when H0 is true. Controlled by alpha.

  • Type II error (false negative): failing to reject H0 when H0 is false. Controlled by beta.

  • Power = 1 minus beta: the probability of catching a real effect.

Which error is worse depends on context. In criminal trials, a Type I error (convicting an innocent person) is considered worse. In a medical screening, a Type II error (missing a disease) may be more dangerous.

Factors that increase power: larger sample size, larger true effect, higher alpha (at the cost of more Type I errors), lower variability.

P-Values

A small p-value (e.g. 0.003) means the observed data would be very surprising if the null hypothesis were true. A large p-value (e.g. 0.42) means the data are consistent with the null hypothesis.

The p-value is not the probability that H0 is true. It is a measure of evidence against H0.

Statistical Significance vs. Practical Significance

Statistical significance tells you the effect is unlikely to be zero. Practical significance asks whether the effect is large enough to care about.

With a very large sample, even a tiny, meaningless difference can be statistically significant. Always ask: is this result large enough to matter in context?

Checking Assumptions

Most inference procedures assume:

  • Independence: observations are independent of each other (usually ensured by random sampling or random assignment).

  • Normality: the sampling distribution of the statistic is approximately normal. For means, the Central Limit Theorem helps when n is large (commonly n >= 30). For small samples, check the sample data for strong skewness or outliers.

  • Equal variances (for two-sample t-tests and ANOVA): the spread within each group should be similar. Check by comparing sample standard deviations; a rough rule is the ratio of the largest to smallest should be less than 2.

From output (e.g. residual plots, normal probability plots), you should be able to judge whether assumptions are met. A badly curved normal probability plot or a funnel-shaped residual plot signals trouble.


Common Misconceptions

  • Students often interpret a 95% confidence interval as "there is a 95% probability the true parameter is in this interval." That is incorrect. The 95% refers to the long-run success rate of the method, not to any single interval.

  • Students sometimes say "accept the null hypothesis" when they mean "fail to reject." These are very different statements. You can never prove the null is true from sample data.

  • Students confuse statistical significance with practical importance. A statistically significant result can be trivially small and completely unimportant in practice.

  • Students sometimes think a large p-value proves the null hypothesis is true. It does not. It simply means the data do not provide strong evidence against it.

Why It Matters / Exam Flags

  • ⚠️ Be prepared to evaluate a described study and identify flaws: missing control group, no randomisation, lurking variables, selection bias.

  • ⚠️ Know the correct interpretation of a confidence interval and confidence level. This is tested frequently.

  • ⚠️ Expect a question asking whether a Type I or Type II error is worse in a specific scenario. You need to identify the real-world consequence of each.

  • ⚠️ Know why we say "fail to reject" rather than "accept" the null hypothesis.

  • ⚠️ Be able to look at residual plots or normal probability plots and judge whether assumptions are satisfied.

Quick Self-Test

  1. True or false: a convenience sample is appropriate for making inferences about a population. (False)

  1. Fill in the blank: the probability of a Type I error is called ______. (Alpha, the significance level)

  1. True or false: if the p-value is 0.08 and alpha is 0.05, you reject the null hypothesis. (False)

  1. Fill in the blank: Power = 1 minus ______. (Beta)

  1. True or false: increasing the sample size increases the power of a test. (True)


Practice Q&A

Q: A pharmaceutical company tests a new pain reliever. Patients are randomly assigned to receive either the drug or a placebo, and neither the patients nor the doctors measuring pain levels know who received which. What elements of good experimental design are present?

A: Randomisation (random assignment to groups), control (placebo group), and blinding (double-blind, since neither patients nor doctors know the assignment). Replication would depend on the sample size.

Q: A 95% confidence interval for the mean commute time is (22, 38) minutes. A student says "there is a 95% chance the true mean is between 22 and 38." Is this correct?

A: No. The true mean is fixed; it is either in the interval or it is not. The 95% refers to the procedure: if we repeated the study many times, about 95% of the resulting intervals would contain the true mean.

Q: A study finds that students who eat breakfast score higher on exams. A student concludes that eating breakfast causes higher scores. What is wrong with this reasoning?

A: This is likely an observational study, not an experiment. A lurking variable (e.g. socioeconomic status, general health habits, sleep quality) could explain both breakfast-eating and higher scores. Without random assignment, you cannot establish causation.

Q: In a hypothesis test with alpha = 0.05, the p-value is 0.02. The estimated effect is that a new fertiliser increases crop yield by 0.3 grams per hectare. Is this result practically significant?

A: It is statistically significant (p = 0.02 < 0.05), but a 0.3-gram increase per hectare is negligible in agricultural terms. The result has no practical significance.

Q: A researcher fails to reject the null hypothesis that a new teaching method has no effect. She concludes the method definitely does not work. What is wrong?

A: Failure to reject the null does not prove it is true. It may mean the sample was too small to detect a real effect (low power), or the effect might exist but be small.

Connections to Other Topics

Experimental design connects directly to inference: a well-designed study is what makes confidence intervals and hypothesis tests valid. If the sampling is biased or lurking variables are uncontrolled, no amount of statistical machinery can fix the conclusions. The concepts of Type I error, Type II error, and power reappear throughout ANOVA (covered in the next set of notes), where multiple comparisons raise the risk of false positives.

Related Terms / Search Tags

experimental design, control group, randomisation, random assignment, blinding, double-blind, placebo, replication, lurking variable, confounding variable, simple random sample, SRS, stratified sampling, cluster sampling, systematic sampling, convenience sample, confidence interval, confidence level, critical value, z-star, t-star, margin of error, standard error, hypothesis test, null hypothesis, alternative hypothesis, p-value, significance level, alpha, Type I error, false positive, Type II error, false negative, power, beta, statistical significance, practical significance, residual plot, normal probability plot, assumptions