Difficulty: Intermediate | Prerequisites: Basic descriptive statistics, scatterplots
Correlation measures the strength and direction of the linear relationship between two quantitative variables. It sits at the heart of regression analysis: before you build a model to predict one variable from another, you need to know whether a meaningful linear association exists at all. If you have covered scatterplots and basic summary statistics, you have the tools to follow this material. Correlation is a gateway concept that leads directly into simple linear regression, which is the next major topic.
The Pearson correlation coefficient r tells you how closely two numeric variables follow a straight-line pattern, on a scale from -1 to +1. Values near +1 or -1 indicate strong linear relationships; values near 0 indicate weak or no linear relationship. You can formally test whether the population correlation differs from zero, and when you test many pairs at once, you need to adjust your significance level (Bonferroni correction) to avoid false positives.
Pearson correlation coefficient (r)
A number between -1 and +1 that quantifies the strength and direction of the linear association between two quantitative variables.
In simple terms, r tells you how tightly the points on a scatterplot cluster around a straight line, and whether the line slopes up or down.
Population correlation (ρ, rho)
The true correlation between two variables across the entire population, as opposed to the sample estimate r.
Think of it as the value you would get if you could measure every single member of the population rather than just your sample.
Pairwise correlation
The correlation calculated between one specific pair of variables. With four variables, there are C(4,2) = 6 unique pairs.
In simple terms, you take every possible two-variable combination and compute r for each one.
Bonferroni correction
A method for adjusting the significance level when conducting multiple hypothesis tests simultaneously, to control the overall Type I error rate. The adjusted significance level is α divided by the number of tests.
Think of it as raising the bar for significance to account for the fact that running many tests gives you more chances to get a false positive by luck.
Linear relationship
An association between two variables that can be reasonably described by a straight line on a scatterplot.
In simple terms, as one variable increases, the other tends to increase (or decrease) at a roughly constant rate.
r is calculated from the paired observations of two variables, using their means, standard deviations, and the sum of their cross-products.
In R, cor(x, y) gives you r for two vectors. For a full matrix of pairwise correlations across multiple columns, use cor(dataframe[, numeric_columns]).
With four numeric variables (e.g. the iris dataset's Sepal.Length, Sepal.Width, Petal.Length, Petal.Width), there are 6 unique pairs to examine.
Direction: A positive r means both variables tend to increase together. A negative r means one tends to decrease as the other increases.
Strength: The closer |r| is to 1, the tighter the points cluster around a line. Rough guidelines:
|r| > 0.8: strong linear relationship
0.5 < |r| < 0.8: moderate linear relationship
|r| < 0.5: weak linear relationship
Interpretation template: "There is a [strong/moderate/weak] [positive/negative] linear relationship between [variable 1] and [variable 2], with r = [value]."
Always pair the numeric value of r with a contextual sentence. Saying "r = 0.87" is not enough on its own.
The null hypothesis is H₀: ρ = 0 (no linear relationship in the population).
The alternative hypothesis is H₁: ρ ≠ 0.
In R, cor.test(x, y) returns the test statistic, p-value, and confidence interval.
If the p-value is less than your chosen α, you reject H₀ and conclude the population correlation is significantly different from zero.
When testing all 6 pairs of 4 variables, you are running 6 tests simultaneously.
Without correction, the probability of at least one false positive across all tests exceeds α = 0.05.
Bonferroni correction: divide your overall α by the number of tests. For 6 tests at α = 0.05, the adjusted significance level is 0.05/6 ≈ 0.0083.
A pair's correlation is statistically significant only if its p-value falls below 0.0083, not the usual 0.05.
Pearson correlation coefficient:
r = \frac{\sum_{i=1}^{n}(x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum_{i=1}^{n}(x_i - \bar{x})^2 \cdot \sum_{i=1}^{n}(y_i - \bar{y})^2}}Alternatively, using standard deviations:
r = \frac{1}{n-1}\sum_{i=1}^{n}\left(\frac{x_i - \bar{x}}{s_x}\right)\left(\frac{y_i - \bar{y}}{s_y}\right)Test statistic for H₀: ρ = 0:
t = \frac{r\sqrt{n-2}}{\sqrt{1-r^2}}This follows a t-distribution with n - 2 degrees of freedom under the null hypothesis.
Bonferroni-adjusted significance level:
\alpha_{\text{adj}} = \frac{\alpha}{k}where k is the number of simultaneous tests.
Students often think a high r means one variable causes the other. Correlation measures association, not causation. Two variables can be strongly correlated because of a lurking third variable.
Students sometimes assume r captures all types of relationships. It only measures linear association. Two variables can have a perfect curved relationship and still produce an r near zero.
A correlation of r = 0 does not mean the variables are unrelated. It means there is no linear relationship. There may still be a strong non-linear pattern.
Students often forget that outliers can dramatically inflate or deflate r. A single extreme point can shift the correlation substantially.
⚠️ You will be asked to interpret r in context, not just report the number. Always state direction, strength, and the names of the variables.
⚠️ Know how to identify the strongest and weakest linear relationships from a correlation matrix. Strongest = largest |r|; weakest = smallest |r|.
⚠️ When testing multiple correlations at once, the Bonferroni correction is expected. If the problem gives you α = 0.05/6, that is the corrected threshold.
⚠️ The sign of r matters as much as the magnitude. Be precise about whether the relationship is positive or negative.
True or False: A correlation of r = -0.92 indicates a weaker relationship than r = 0.75.
Answer: False. Strength is based on |r|. |-0.92| = 0.92 > 0.75, so the negative correlation is stronger.
Fill in the blank: The Pearson correlation coefficient can only range from ______ to ______.
Answer: -1 to +1.
True or False: If r = 0 between two variables, they have no relationship whatsoever.
Answer: False. They have no linear relationship, but a non-linear relationship may still exist.
Fill in the blank: When running 6 correlation tests simultaneously at α = 0.05, the Bonferroni-adjusted significance level is ______.
Answer: 0.05/6 ≈ 0.0083.
Q: You compute the correlation between height and weight for a sample of 50 adults and find r = 0.82. Interpret this value.
A: There is a strong positive linear relationship between height and weight. As height increases, weight tends to increase as well.
Q: A researcher tests the correlation between 10 pairs of variables at α = 0.05 and finds three p-values below 0.05. Should all three be considered significant?
A: Not necessarily. With 10 tests, the Bonferroni-adjusted threshold is 0.05/10 = 0.005. Only pairs with p-values below 0.005 should be considered significant.
Q: You compute r = 0.02 between two variables. Does this prove the variables are independent?
A: No. It indicates virtually no linear association, but a non-linear relationship could still exist. Always check a scatterplot.
Q: Why does the test for H₀: ρ = 0 use a t-distribution with n - 2 degrees of freedom?
A: The test statistic is derived under the assumption that the data come from a bivariate normal distribution, and two parameters (the means) are estimated from the data, leaving n - 2 degrees of freedom.
Correlation is the foundation for simple linear regression. The slope of the regression line is directly related to r through the formula b₁ = r(s_y/s_x). Understanding correlation also connects to multiple regression, where partial correlations describe relationships after controlling for other variables.
Pearson r, correlation coefficient, sample correlation, population correlation, rho, pairwise correlation, correlation matrix, scatterplot, linear association, Bonferroni correction, multiple comparisons, Type I error, significance testing, hypothesis test for correlation, iris dataset, STAT 355