Hypothesis Testing Fundamentals, STAT 101 – Study Notes
offline

Source: Introduction to Statistics, Purdue University Tags: hypothesis testing, null hypothesis, alternative hypothesis, p-value, significance level, Type I error, Type II error, statistical power, alpha, beta, critical value, test statistic

Difficulty: Intermediate Prerequisites: Basic probability, sampling distributions, standard normal distribution. If you are not comfortable with the idea of a sampling distribution or what a standard deviation measures, review those topics first.


Big Picture

Hypothesis testing is the main framework statistics uses to answer yes-or-no questions about populations when all you have is a sample. It sits at the centre of nearly every empirical discipline, from clinical trials to quality control to social science research. The logic is indirect: you assume nothing interesting is happening (the null hypothesis), collect data, and then ask how surprising your data would be under that assumption. If the answer is "very surprising," you reject the null. Everything else in this topic, every formula and every test, is a variation on that single idea.


TL;DR

You set up two competing claims about a population (null vs alternative), calculate a test statistic from your sample, and use the p-value to decide whether the data are surprising enough to reject the null hypothesis. The significance level (alpha) is the threshold you pick in advance, and two types of error (Type I and Type II) describe the ways the decision can go wrong.


Key Terms

Null hypothesis (H0)

The default claim about a population parameter. It typically states "no effect" or "no difference." You assume it is true unless the evidence says otherwise.

In simple terms, this is the boring explanation: nothing interesting is going on.

Alternative hypothesis (H1 or Ha)

The competing claim you are trying to find evidence for. It represents the effect or difference you suspect exists.

Think of it as the interesting explanation, the one you need strong evidence to support.

Test statistic

A single number calculated from the sample data that measures how far the observed result sits from what the null hypothesis predicts.

In simple terms, it is a standardised score that tells you how unusual your sample is, assuming H0 is true. Larger absolute values mean more evidence against H0.

p-value

The probability of obtaining a test statistic at least as extreme as the one you calculated, assuming the null hypothesis is true.

Think of it as a measure of surprise. A small p-value means your data would be very unlikely if H0 were true, which is grounds for rejecting it.

Significance level (α, alpha)

The threshold you choose before collecting data for how small the p-value must be to reject H0. It is also the maximum probability of a Type I error you are willing to accept.

In simple terms, alpha is the line you draw in the sand. The conventional choice is 0.05, meaning you accept a 5% risk of rejecting a true null hypothesis.

Critical value

A boundary point (or points) on the test statistic's distribution that separates the rejection region from the non-rejection region.

Think of it as the cutoff score. If your test statistic lands beyond the critical value, you reject H0.

Type I error (α)

Rejecting the null hypothesis when it is, in fact, true. A false positive.

In simple terms, you conclude something is happening when it is not. The probability of this equals your significance level.

Type II error (β, beta)

Failing to reject the null hypothesis when the alternative is, in fact, true. A false negative.

Think of it as missing a real effect because your sample was not convincing enough.

Statistical power (1 − β)

The probability of correctly rejecting H0 when H1 is true.

In simple terms, power is your test's ability to detect a real effect. Higher power means you are less likely to miss something that is there. Power increases with larger sample sizes, larger effect sizes, and higher alpha levels.


Core Content

The Logic of Hypothesis Testing

  • You begin by assuming H0 is true. This is not a statement of belief; it is a methodological starting point.

  • You collect sample data and compute a test statistic that summarises how far the observed result is from what H0 predicts.

  • You then calculate the p-value: the probability of seeing a result this extreme or more extreme, given that H0 is true.

  • You compare the p-value to your pre-chosen significance level (α).

The Decision Rule

  • If p-value < α, reject H0. The data are unlikely enough under H0 that you conclude in favour of H1.

  • If p-value ≥ α, fail to reject H0. The data are not surprising enough to abandon the default claim.

  • "Fail to reject" is deliberate language. You have not proven H0 is true; you simply lack sufficient evidence against it.

One-Tailed vs Two-Tailed Tests

  • Two-tailed (two-sided): H1 states the parameter is different from the null value in either direction (e.g., μ ≠ μ0). The rejection region is split across both tails of the distribution.

  • One-tailed (one-sided): H1 specifies a direction (e.g., μ > μ0 or μ < μ0). The entire rejection region sits in one tail, making it easier to reject H0 in that direction but blind to effects in the other.

Understanding Errors

  • Type I error and significance level are the same quantity, just viewed from different angles. Alpha is the rate at which you are willing to make Type I errors.

  • Type II error depends on the true value of the parameter, the sample size, and the variability in the data. It is harder to control directly.

  • There is a trade-off: lowering α (to reduce false positives) raises β (increasing false negatives), all else being equal. The only free way to improve both is to increase the sample size.

The Role of the Critical Value

  • For a given α and a known distribution (z, t, chi-square, etc.), the critical value marks the boundary of the rejection region.

  • You can frame every hypothesis test in two equivalent ways: compare the p-value to α, or compare the test statistic to the critical value. They always give the same answer.


Real-World Applications

Hypothesis testing is the backbone of clinical drug trials: regulators require that a new treatment shows a statistically significant improvement over placebo before it can be approved. It is also how manufacturers decide whether a production line is meeting quality specifications, by testing whether the observed defect rate differs from the target.


Common Misconceptions

  • "The p-value is the probability that H0 is true." It is not. The p-value assumes H0 is true and tells you how likely your data are under that assumption. Those are different statements.

  • "Failing to reject H0 means H0 is true." It means you did not find enough evidence against it. A small or underpowered study can easily miss a real effect.

  • "A statistically significant result is practically important." Significance depends on sample size. With a large enough sample, trivially small effects become statistically significant. Always check the effect size.

  • "α = 0.05 is a universal rule." It is a convention, not a law. Different fields and different stakes call for different thresholds. Particle physics uses α ≈ 0.0000003 (5-sigma).


Why It Matters / Exam Flags

⚠️ Know the precise definitions of p-value, Type I error, and Type II error. These are among the most commonly tested (and most commonly botched) concepts in introductory statistics.

⚠️ Be able to set up H0 and H1 from a word problem. Identify whether the test is one-tailed or two-tailed from the phrasing of the research question.

⚠️ Understand the decision rule in both forms: p-value vs α, and test statistic vs critical value.

⚠️ Do not confuse "fail to reject H0" with "accept H0." Exams will penalise the latter phrasing.


Quick Self-Test

True or false: A p-value of 0.03 means there is a 3% chance the null hypothesis is true. False. The p-value is the probability of observing data this extreme assuming H0 is true, not the probability that H0 is true.

True or false: Increasing the sample size increases the power of a test. True. A larger sample gives a more precise estimate, making it easier to detect a real effect.

Fill in the blank: A Type I error occurs when you ______ H0 and it is actually ______. Reject; true.

Fill in the blank: The probability of a Type II error is denoted by ______, and statistical power equals ______. β; 1 − β.

True or false: If the p-value is greater than α, you accept the null hypothesis. False. You fail to reject it. That is not the same as accepting it.


Practice Q&A

Q: A researcher sets α = 0.05 and obtains a p-value of 0.08. What is the correct conclusion?

A: Fail to reject H0. The p-value exceeds the significance level, so there is insufficient evidence to conclude that the alternative hypothesis is true.

Q: Explain the difference between a Type I error and a Type II error in the context of a drug trial.

A: A Type I error would mean concluding the drug works when it does not (approving an ineffective drug). A Type II error would mean concluding the drug does not work when it does (failing to approve an effective drug).

Q: Why is "accept H0" considered incorrect language?

A: Because failing to reject H0 only means the data were not extreme enough to rule it out. It does not prove H0 is true. There may be an effect that the study was too small or too noisy to detect.

Q: If a researcher changes α from 0.05 to 0.01, what happens to the probability of a Type I error and the probability of a Type II error (assuming sample size stays the same)?

A: The probability of a Type I error decreases (from 5% to 1%). The probability of a Type II error increases, because the test now demands more extreme evidence to reject H0, making it harder to detect a real effect.

Q: A test statistic falls in the rejection region. What can you conclude about the p-value relative to α?

A: The p-value must be less than α. Falling in the rejection region and having a p-value below α are equivalent statements.


Connections to Other Topics

This material connects directly to confidence intervals: a 95% confidence interval contains all parameter values that would not be rejected by a two-sided test at α = 0.05. If you understand one, you understand the other.

Hypothesis testing also lays the groundwork for ANOVA (testing more than two group means), regression analysis (testing whether coefficients are significantly different from zero), and chi-square tests (testing relationships between categorical variables).


Related Terms / Search Tags

hypothesis test, significance test, null hypothesis, alternative hypothesis, H0, H1, Ha, p-value, p value, alpha level, significance level, Type I error, Type II error, false positive, false negative, statistical power, power of a test, critical value, rejection region, one-tailed test, two-tailed test, one-sided test, two-sided test, fail to reject, decision rule, STAT 101, intro stats, Purdue