Source: Principles of Statistics I, Texas A&M / Tamhane & Dunlop Ch. 3
Tags: estimation, population mean, variance of estimator, SRS estimation, stratified estimation, cluster estimation, finite population correction, experimental design, randomisation, replication, blocking, experimental unit, treatment, factor, level, CRD, RCBD, Latin square, split plot, confounding, covariate, Taguchi
Different sampling methods require different formulas for estimating the population mean and its variance. Using the wrong formula (e.g. assuming SRS when stratified sampling was used) can produce grossly incorrect results. Experimental design rests on three pillars: randomisation, replication, and blocking. A well-designed experiment is economical, measures the influence of multiple factors, and allows valid statistical inference.
Population mean (µ)
The true average value of a measurement across all units in the population.
Estimator of the population mean (µ-hat)
A formula applied to sample data to produce an estimate of µ. The formula differs depending on whether SRS, stratified, or cluster sampling was used.
Finite population correction (FPC)
The factor (N – n)/(N – 1) that adjusts the variance estimate when the sample is a non-trivial fraction of the population. When n/N is very small, FPC is approximately 1 and can be ignored.
Experimental unit
The entity to which a treatment is randomly assigned.
Measurement unit
The entity on which a measurement or observation is made. Often the same as the experimental unit, but not always.
Treatment
A specific combination of factor levels applied to an experimental unit.
Factor
A controllable experimental variable thought to influence the response.
Level
A specific value of a factor.
Block
A group of homogeneous experimental units. Blocking groups similar units together to reduce variability within comparisons.
Replication
Observations on two or more experimental units that have been randomly assigned to the same treatment.
Subsampling
Multiple measurements on the same experimental unit under the same treatment, either taken at different locations (spatially) or different times (longitudinally).
Confounding
When one or more effects cannot be unambiguously attributed to a single factor or interaction because they are entangled with other effects.
Interaction
When the effect of one factor on the response depends on the level of another factor.
Covariate
An uncontrollable variable that influences the response but is unaffected by the experimental factors.
The sampling method determines which formula is correct. Using the wrong one can lead to grossly incorrect estimates of both µ-hat and its variance.
Simple Random Sampling
The estimator is the sample mean:
µ-hat = (1/n) × Σ yᵢ
The estimated variance is:
Var(µ-hat)_SRS = (s²/n) × (N – n)/(N – 1)
where s² = [1/(n – 1)] × Σ(yᵢ – y-bar)².
When n/N is very small, the finite population correction factor drops out and Var(µ-hat) simplifies to s²/n. This also applies when sampling is done with replacement.
Stratified Random Sampling
With L strata of sizes N₁, N₂, ..., N_L and sample sizes n₁, n₂, ..., n_L:
µ-hat = (1/N) × Σ Nᵢ × y-barᵢ
This is a weighted average of stratum sample means, weighted by stratum size.
The estimated variance is:
Var(µ-hat)_STRATIFIED = (1/N²) × Σ [Nᵢ² × ((Nᵢ – nᵢ)/(Nᵢ – 1)) × (sᵢ²/nᵢ)]
The key insight: Var(µ-hat)_STRATIFIED is much smaller than Var(µ-hat)_SRS when the within-stratum variances (sᵢ²) are much smaller than the overall variance (s²). Stratification pays off when strata are internally homogeneous.
Single-Stage Cluster Sampling
With n clusters selected from N total clusters, where cluster i has mᵢ elements:
µ-hat = Σ yᵢ / Σ mᵢ = Σ(mᵢ × y-barᵢ) / Σ mᵢ
This is a weighted average of cluster means, weighted by cluster size.
The estimated variance is:
Var(µ-hat) = [(N – n)/(N × n × M-bar²)] × [Σ(mᵢ × y-barᵢ – mᵢ × µ-hat)² / (n – 1)]
where M-bar = M/N is the average cluster size in the population.
The formulas for all three methods differ substantially. If you assume SRS was used when the data actually came from a stratified or cluster design, both your point estimate and your variance estimate may be wrong.
Two foundational quotes frame the discipline:
G. Box, S. Hunter, and W. Hunter: "Whenever possible, experiments should be comparative. If you are testing a modification of a process, the modified and unmodified processes should be run side by side in the same experiment."
R.A. Fisher: "It is possible, and indeed it all too frequent, for an experiment to be so conducted that no valid estimate of error is available. In such a case, the experiment cannot be said, strictly speaking, to be capable of proving anything."
Randomisation: experimental units must be randomly assigned to treatments (or randomly selected from treatment populations). The time order and spatial positioning of experiments must also be randomised. This protects against unknown sources of bias and ensures the validity of statistical procedures.
Replication: multiple experimental units must receive the same treatment. Without replication, there is no way to estimate experimental error, and without an error estimate, no valid inference can be made.
Blocking: grouping homogeneous experimental units together reduces variability within comparisons. When variation in responses is as large as the treatment differences, blocking (along with increased sample size and use of covariates) can rescue the experiment.
Every designed experiment has three structural components:
Method of randomisation (design structure): completely randomised design (CRD), randomised complete block design (RCBD), balanced incomplete block design (BIBD), Latin square, crossover, split plot, and others.
Treatment structure: one-way classification, factorial, fractional factorial, with fixed, random, or mixed factor levels.
Measurement structure: single measurement per unit, repeated measurements under different treatments, longitudinal or spatial repeated measurements, or subsampling.
P1 – Formulate questions and hypotheses before running the experiment. This minimises the number of replications needed and ensures all necessary measurements are taken.
P2 – Critically analyse the research hypotheses. Review the literature, evaluate whether the aim is reasonable, and forecast possible outcomes to check that the resulting data can be analysed properly (watch for too many zeros, categorical data, too few replications, or correlated observations).
P3 – Select procedures for conducting research. Determine treatments, measurements, selection method, sample size, design type, and effects of adjacent units on each other. Outline recording tables, document procedures, and state costs.
P4 – Eliminate personal biases. Never discard the "most discrepant" observation from a set of three. Never place a favoured treatment under the best conditions. Discard values only after critical examination, and always report excluded data with reasons.
P5 – Evaluate statistical assumptions. Check that distributional requirements hold through residual analysis.
P6 – Write a quality report regardless of outcome. Include graphics, describe methods so readers can assess validity, and report results even when the null hypothesis is not rejected (to prevent publication bias and accumulated Type I errors in the literature). Crucially, report both the p-value and the effect size with confidence intervals, distinguishing statistically significant results from practically significant ones.
Masking of factor effects: when variability in responses is as large as the treatment differences, effects go undetected. Solutions: increase sample size, use blocking, add covariates, or combine all three.
Uncontrolled factors: known influential factors that are not included as treatment or blocking variables can severely compromise conclusions. Examples: soil fertility differences between plots, position on greenhouse benches, position in ovens, time of day.
Erroneous principles of efficiency: cutting corners on factors or levels to save time or money can mean important factors are ignored or non-linear effects go undetected because too few levels were used across too narrow a range.
A semiconductor manufacturer wants to evaluate two coating types (C₁, C₂) and three thicknesses (T₁, T₂, T₃) on wafer conductivity. 72 wafers, 12 per treatment combination, conductivity measured before and after coating at five locations per wafer.
This example is used to illustrate four possible designs:
Scenario I (CRD): all 72 wafers on one day, one lab. Treatments randomly assigned.
Scenario II (RCBD): 24 wafers per day over three days. Days serve as blocks.
Scenario III (Latin square variant): 6 wafers per day across 6 labs. Each treatment appears in every day-lab combination.
Scenario IV (Split plot): coating type is hard to change (whole-plot factor), thickness is easy to change (sub-plot factor). On each day, 12 wafers get each coating, then 4 of those 12 get each thickness.
On-line experiments (EVOP): run while the process is in full production, varying two or more factors within a narrow range around normal operations to find the optimum.
Off-line experiments: run in laboratories or pilot plants, allowing broader exploration of the factor space.
Two goals in quality improvement: bring the product on target (mean equals target value) and achieve uniformity (small variability around the target). Together these minimise MSE = Bias² + Variance.
SRS estimator and variance:
µ-hat = (1/n) Σ yᵢ
Var(µ-hat) = (s²/n) × (N – n)/(N – 1)
Stratified estimator and variance:
µ-hat = (1/N) Σ Nᵢ y-barᵢ
Var(µ-hat) = (1/N²) Σ [Nᵢ² × ((Nᵢ – nᵢ)/(Nᵢ – 1)) × sᵢ²/nᵢ]
Cluster estimator and variance:
µ-hat = Σ yᵢ / Σ mᵢ
Var(µ-hat) = [(N – n)/(NnM-bar²)] × [Σ(mᵢy-barᵢ – mᵢµ-hat)² / (n – 1)]
MSE for quality improvement:
MSE = (Bias)² + (StDev)² = (Distance to Target)² + Variance
⚠️ Using the SRS variance formula when data was collected via stratified or cluster sampling will give incorrect results. Always identify the sampling method before computing.
⚠️ Stratified sampling beats SRS when within-stratum variance is much smaller than overall variance.
⚠️ The three pillars (randomisation, replication, blocking) are fundamental. Know why each matters.
⚠️ Know the difference between experimental unit and measurement unit. In the wafer example, the wafer is the experimental unit; the five locations on each wafer are subsamples.
⚠️ Statistical significance (small p-value) is different from practical significance (large effect size). Exams may ask you to distinguish between them.
⚠️ Understand why the split-plot design exists: some factors are harder or more expensive to change than others, so complete randomisation is impractical.
⚠️ Publication bias: if researchers only publish when they reject the null, 5% of studies on a false hypothesis will still appear to support it (Type I errors). This is why reporting non-significant results matters.
Q: Why is it important to know which sampling method was used before computing the estimated variance of µ-hat?
A: Because the formula for Var(µ-hat) differs substantially across SRS, stratified, and cluster sampling. Applying the wrong formula can yield a grossly incorrect estimate of precision.
Q: Under what condition does stratified sampling produce a much smaller variance than SRS?
A: When the within-stratum variances (sᵢ²) are much smaller than the overall population variance (s²), meaning the strata are internally homogeneous but differ from each other.
Q: What are the three pillars of experimental design?
A: Randomisation, replication, and blocking.
Q: What is the difference between an experimental unit and a measurement unit?
A: The experimental unit is the entity to which a treatment is randomly assigned. The measurement unit is the entity on which the observation is made. They are often the same, but when multiple measurements are taken on a single experimental unit (subsampling), they differ.
Q: In the wafer experiment, why does Scenario IV use a split-plot design instead of a CRD?
A: Changing coating type on the machine requires considerable set-up time, while changing thickness does not. The split-plot design keeps the hard-to-change factor (coating type) constant within larger groups and randomises the easy-to-change factor (thickness) within those groups, saving time without sacrificing validity.
Q: An experiment finds a p-value of 0.001 for the difference between two treatments, but the estimated difference is 0.2 units on a scale of 0 to 1000. Is this result practically significant?
A: The result is statistically significant (very small p-value) but likely not practically significant. The effect size of 0.2 on a 1000-point scale is negligible. Both the p-value and the magnitude of the effect should be reported.
Q: Why should researchers report results even when the null hypothesis is not rejected?
A: To avoid publication bias. If only "significant" results are published, Type I errors (false positives at the 5% rate) accumulate in the literature, potentially supporting a false hypothesis while the 95% of studies that correctly found no effect remain unpublished.
population mean estimation, sample mean, variance of estimator, finite population correction, FPC, SRS variance, stratified sampling variance, cluster sampling variance, weighted average, experimental design, randomisation, replication, blocking, experimental unit, measurement unit, treatment, factor, level, interaction, confounding, covariate, CRD, completely randomised design, RCBD, randomised complete block design, Latin square, split plot, crossover design, factorial design, fractional factorial, subsampling, EVOP, evolutionary operation, Taguchi, MSE, statistical significance vs practical significance, effect size, publication bias, Type I error, Federer principles