Why Study Statistics and Common Misconceptions – STAT, Handout 01 – Study Notes

Source: Handout 01, Principles of Statistics I (Texas A&M University) | Textbook: Ott & Longnecker, 7th Ed. | Reference article: Jessica Utts, "What Educated Citizens Should Know About Statistics and Probability," The American Statistician, May 2003

Tags: why study statistics, statistical literacy, misconceptions, correlation vs causation, statistical significance, practical significance, sample size, survey bias, conditional probability, natural variability, false positive, false negative


TL;DR

Statistics matters because nearly everyone encounters data-based claims, and those claims can be misleading. The most common traps include confusing correlation with causation, confusing statistical significance with practical significance, using samples that are too small or too large without adjusting interpretation, and misunderstanding probability and variability. Being statistically literate protects you from bad conclusions in medicine, business, law, and daily life.


Key Terms

Statistical significance

A result is statistically significant when the observed effect is unlikely to have occurred by random chance alone, given the assumptions of the test.

Practical significance

A result is practically significant when the observed effect is large enough to matter in the real world, regardless of whether it cleared a statistical threshold.

Correlation vs. causation

Correlation means two variables move together. Causation means a change in one variable produces a change in the other. Establishing causation requires very restrictive experimental conditions.

Conditional probability

The probability of an event occurring given that another event has already occurred with certainty.

False positive (Type I error context)

A test declares a condition present when it is not.

False negative (Type II error context)

A test declares a condition absent when it is present.

Natural variability

The inherent spread in repeated measurements of a naturally occurring event. Misunderstanding variability leads to incorrect claims about what is "normal" vs. "abnormal."


Core Content

Reasons to Study Statistics

  • Critical reading: you are constantly exposed to data-based claims (product ads, polls, scientific findings). Many are inferences from samples, and some are invalid. Statistical literacy lets you tell the difference.

  • Professional necessity: many careers require interpreting sampling results or running analyses. Physicians evaluate drug trial data, engineers monitor quality, accountants audit by sampling accounts.

  • Legal and policy relevance: courts increasingly rely on probability and statistical inference to evaluate evidence. Statistics appears in salary discrimination suits, product liability cases, and epidemiological testimony.

  • Breadth of application: statistics is essential in the social, biological, and physical sciences, in business forecasting, in engineering and manufacturing quality control, and in government policy.

Common Misconception 1: Correlation Implies Causation

This is the most frequent misinterpretation of statistical findings.

  • A statistically significant relationship between two variables does not mean one causes the other

  • Causation can only be concluded under very restrictive experimental constraints

  • Utts' example: a Newsweek article linked strength of religious beliefs to physical healing, but many other factors affect patient health, and a causal conclusion was not warranted

Common Misconception 2: Statistical Significance vs. Practical Significance

These are two different things, and confusing them is especially easy with large datasets.

  • Utts' example: a study of 507,125 military recruits found a statistically significant difference in average height between those born in spring and those born in autumn

  • The difference was roughly 0.25 inches

  • Statistically significant (because the sample was enormous) but practically meaningless

The takeaway: with a large enough sample, even trivially small effects become statistically significant. Always ask whether the effect size matters in context.

Common Misconception 3: Failure to Find Significance Means No Effect

A study may fail to detect a real difference simply because the sample was too small.

  • Many government-funded studies require researchers to demonstrate that their chosen sample size has adequate power to detect specified differences

  • Methods for determining appropriate sample sizes are covered in later handouts on hypothesis testing

Common Misconception 4: Survey Bias

Surveys are everywhere, especially during election years, and many sources of bias can creep in:

  • Selection bias: how people are chosen for inclusion

  • Question wording bias: how questions are phrased can steer responses

  • Mode bias: how questions are posed to the respondent (in person, phone, online) can affect answers

These issues are covered in more detail in Handout 2.

Common Misconception 5: Conditional Probability Confusion

Many people find conditional probability counterintuitive.

  • Example: a diagnostic test for E. coli in meat has a very low false positive rate and a very low false negative rate

  • Even so, if the particular strain of E. coli occurs very rarely, the probability that E. coli is truly present given a positive test result can still be very low

  • This is the base-rate problem: the rarity of the condition overwhelms the accuracy of the test

  • Covered formally in Handout 13 (conditional probability, Bayes' theorem)

Common Misconception 6: Misunderstanding Natural Variability

People confuse "average" with "normal," ignoring how much natural variation exists.

  • Utts' example: a waste water treatment plant blamed an odour problem on "abnormal" rainfall (170–180% of "normal")

  • Historical rainfall in the region ranged from 6.1 to 37.4 inches, with a median of 16.7 inches

  • The year in question saw 29.7 inches, well within the historical range

  • The company compared to the average but ignored the natural spread. The correct approach: evaluate the percentile of the observed value within the full distribution

  • Covered formally in Handout 6 (data summaries and percentiles)


Why It Matters / Exam Flags

⚠️ "Correlation does not imply causation" is tested constantly. Be ready to identify when a study can and cannot support a causal claim.

⚠️ Know the distinction between statistical significance and practical significance, and be able to give an example of each.

⚠️ Understand why a non-significant result does not prove no effect exists (relates to sample size and power).

⚠️ Be able to list sources of bias in surveys (selection, wording, mode).

⚠️ The base-rate / conditional probability trap (low-prevalence condition + accurate test = still many false alarms) is a classic exam question.

⚠️ "Average" and "normal" are not synonyms. Normal must account for the full range of variability.


Practice Q&A

Q: A study finds a statistically significant correlation between ice cream sales and drowning rates. Can we conclude that ice cream causes drowning?

A: No. Correlation does not imply causation. A confounding variable (hot weather) likely explains both increases. Causal conclusions require controlled experimental conditions.

Q: A study of 500,000 people finds a statistically significant difference of 0.1 beats per minute in resting heart rate between two groups. Is this practically significant?

A: Almost certainly not. The enormous sample size made even a trivially small difference detectable. The effect size (0.1 bpm) has no practical relevance in medicine or health.

Q: A clinical trial with 20 participants finds no significant difference between a new drug and a placebo. Does this prove the drug is ineffective?

A: No. The sample may have been too small to detect a real difference. The study may have lacked sufficient statistical power.

Q: A diagnostic test has a 99% sensitivity and 99% specificity. If the condition affects 1 in 10,000 people, is a positive result likely to be correct?

A: No. Because the condition is so rare, even with high accuracy the vast majority of positive results will be false positives. This is the base-rate (conditional probability) problem.

Q: A company claims rainfall was "180% of normal." Why might this be misleading?

A: "Normal" typically refers to the average, but if the natural variability of rainfall in the region is very wide, 180% of the average may still fall well within the historical range. The correct approach is to evaluate the percentile of the observed rainfall within the full distribution.


Related Terms / Search Tags

why study statistics, statistical literacy, correlation vs causation, confounding variable, statistical significance vs practical significance, effect size, sample size and power, survey bias, selection bias, question wording bias, conditional probability, Bayes' theorem, base rate fallacy, false positive, false negative, natural variability, percentile, average vs normal, Jessica Utts, American Statistician