Source: Handout 01, Principles of Statistics I (Texas A&M University) | Textbook: Ott & Longnecker, An Introduction to Statistical Methods and Data Analysis, 7th Ed.
Tags: statistics definition, scientific method, research process, learning from data, population, sample, hypothesis testing, study design, data collection, data analysis
Statistics is the science of learning from data when information is limited and variable. It follows a structured process (define the problem, collect data, summarise, analyse, interpret, communicate) that mirrors the scientific method. The quality of any study's conclusions depends entirely on how well the sample was selected from the population of interest.
Statistics
The science of designing studies or experiments, collecting data, and modelling/analysing data for the purpose of decision-making and scientific discovery when the available information is both limited and variable. In short: the science of learning from data.
Population
The set of all measurements of interest to the researcher collecting data.
Sample
Any subset of the measurements selected from the population.
Scientific method
A set of principles and procedures used by researchers in their pursuit of knowledge, involving formulation of research goals, modelling/analysing data in context, and testing hypotheses.
Exploratory data analysis (EDA)
Methods used when research is in a new area without existing theories, providing insights that help a researcher formulate theories about the phenomenon under study.
Hypothesis
A proposed explanation or prediction about a population, to be tested using collected data.
Define the problem
Collect the data
Summarise the data
Analyse and model the data
Interpret the analyses and models
Communicate the results
This sequence is the backbone of the entire course. Every example and technique ties back to one or more of these steps.
The scientific method and the "learning from data" process run in parallel. The cycle looks like this:
Formulate a research goal, along with research hypotheses and models
Plan the study: decide sample size, identify variables, define experimental units, choose a sampling mechanism
Collect data and manage it carefully
Draw inferences: produce graphs, estimate parameters, test hypotheses, assess models
Make decisions: write conclusions, present findings
Formulate new models and hypotheses based on what the data revealed, then repeat
The process is iterative. Conclusions from one study often become the starting point for the next.
A more granular view of the same cycle, highlighting the decisions researchers face:
Define research objectives: formulate hypotheses, choose models, specify the population of interest
Design the data collection process: select desired power of tests, confidence level, sample size, experimental units, treatments, sampling design, and variables
Collect data: monitor selection of experimental units, monitor the measurement process, monitor recording of data
Analyse data: produce graphs, tables, and displays; fit models; compute estimates (with standard errors), tests (with p-values), and effect sizes
Make inferences about research objectives: assess validity of model conditions, decide about research hypotheses, distinguish statistical from practical significance
Report results: use non-statistical language, include graphs and tables, write clearly with proper grammar
Develop new studies or address gaps: validate fitted models, explore new models, handle non-response or missing variables
The whole point of statistics is to say something about a population using only a sample.
A sample can be "good", "bad", "biased", or "properly selected"
The challenge is recognising when a sample is bad or biased
To make valid inferences about a larger group, you must carefully define the target population and design a study where the sample is appropriately selected from that population
Four real-world examples show how the learning-from-data process works in practice:
Light bulb quality control: testing 1,000 of 500,000 daily bulbs to estimate the overall defective rate, provided the 1,000 are selected properly
Smoking and weight gain: a random sample of 400 former smokers showed an average weight increase of 5 pounds after one year, suggesting a real relationship (techniques for assessing significance come later)
Nitrogen fertiliser and wheat yield: 30 fields, 5 nitrogen rates, 6 fields per rate, all cultivated identically. The goal was to generalise results beyond the study fields to similar fields
Public opinion polling: as few as 1,500 people can represent 200+ million voters, provided selection is unbiased and questions are nonleading
In every case, the same pattern holds: define the question, design the collection, choose the right variables, analyse, and report clearly.
Results only generalise beyond the study participants when:
The target population is clearly defined
The sample is drawn from that population using appropriate methods
Relevant variables (age, sex, soil type, environmental conditions, etc.) are measured and reported
The report describes the sample characteristics so readers can judge applicability
⚠️ Know the six steps of "learning from data" and be able to map any given study to those steps.
⚠️ Population vs. sample is a foundational distinction. Expect definitions and application questions.
⚠️ The scientific method is iterative, not linear. Conclusions feed back into new hypotheses.
⚠️ A sample's validity depends on how it was selected, not just its size.
⚠️ Be able to identify what makes a study's conclusions generalisable (or not).
Q: What is the definition of statistics as given in this course?
A: The science of designing studies or experiments, collecting data, and modelling/analysing data for the purpose of decision-making and scientific discovery when the available information is both limited and variable.
Q: List the six steps of the "learning from data" process.
A: Define the problem, collect the data, summarise the data, analyse/model the data, interpret the analyses/models, communicate the results.
Q: What is the difference between a population and a sample?
A: A population is the set of all measurements of interest to the researcher. A sample is any subset of measurements selected from the population.
Q: In the light bulb example, why is it acceptable to test only 1,000 of the 500,000 bulbs produced daily?
A: Because, provided the 1,000 bulbs are selected properly, the fraction defective in the sample will with high probability be very close to the fraction defective in the entire day's production.
Q: Why is the scientific method described as iterative rather than linear?
A: Because the conclusions of one study often lead to the formulation of new research goals, new hypotheses, and new models, which then require further data collection and analysis.
Q: In the wheat experiment, what additional information must the report include for the results to be generalisable?
A: Information about soil characteristics and environmental conditions (temperature, rainfall) of the study fields, so that farmers can judge whether the conclusions apply to their own fields.
learning from data, scientific method in statistics, population definition, sample definition, study design, data collection, hypothesis formulation, exploratory data analysis, EDA, generalisability, inference, sampling mechanism, experimental units, research process, statistical discovery, Ott and Longnecker, Principles of Statistics I, TAMU statistics