Correlation, Introduction to Statistics Ch. 13b – Study Notes (Part 1 of 3)
offline

Difficulty: Intermediate | Prerequisites: Chapter 13a (Simple Linear Regression basics, least-squares line, ANOVA table for regression).

Big Picture

This section builds on the simple linear regression material from Chapter 13a by introducing the Pearson correlation coefficient, the standard measure of the strength and direction of a linear relationship between two quantitative variables. You should already be comfortable with scatterplots, the least-squares regression line, the ANOVA table (SSR, SSE, SST), and the coefficient of determination R-squared. If any of those feel shaky, revisit Chapter 13a first. Correlation ties together the regression slope and R-squared into a single number that is easy to interpret but easy to misuse, so this section spends equal time on calculation and on the situations where correlation is not a valid measure.


TL;DR

The sample correlation r measures how tightly points cluster around a straight line: it ranges from -1 to +1, shares the sign of the regression slope, and equals plus or minus the square root of R-squared in simple linear regression. A value near 0 means no linear association, but there could still be a curved relationship. Always plot your data before trusting r.

Key Terms

Sample correlation (r)

A number between -1 and +1 that measures the direction and strength of the linear relationship between two quantitative variables. Also called Pearson's correlation coefficient.

In simple terms, r tells you how close the data points sit to a straight line and whether the line slopes up or down.

Pearson's correlation coefficient

The formal name for r. It is the most common measure of linear association and is calculated using standardised deviations of x and y from their means.

Think of it as the "industry-standard" way to put a number on how linear your scatterplot looks.

S_XX (sum of squares for x)

The sum of squared deviations of x from its mean. Computed as the sum of x-squared minus (sum of x) squared divided by n. Used in the denominator when calculating both the slope and the correlation.

S_YY (sum of squares for y)

The same idea applied to y. Equal to SST (corrected total sum of squares) from the ANOVA table.

S_XY (sum of cross-products)

The sum of the products of the deviations of x and y from their respective means. This is the numerator of both the slope b1 and the correlation r. Its sign determines whether the association is positive or negative.

Coefficient of determination (R-squared)

The proportion of observed variation in y explained by x. In simple linear regression, R-squared equals r-squared.

In simple terms, R-squared is the percentage of the scatter that the line accounts for. Squaring r removes the sign.

Positive association

When r is greater than 0. As x increases, y tends to increase. On a scatterplot, the cloud of points slopes upward from left to right.

Negative association

When r is less than 0. As x increases, y tends to decrease. The cloud slopes downward.

Linearly uncorrelated

When r equals 0. There is no linear pattern. This does not mean x and y are unrelated; there may be a non-linear association.


Core Content

Calculating Sample Correlation

Three equivalent formulas appear in the course (Eqs. 13b.1a, 1b, 1c). The one you will use most on homework is Eq. 13b.1a:

  • r = S_XY / sqrt(S_XX * S_YY)

You compute S_XX, S_YY, and S_XY from the computing formulas (sums of x, y, x-squared, y-squared, and xy), then plug straight into the formula above.

The other two forms give insight into what correlation means:

  • Eq. 13b.1b divides by (n - 1) * s_x * s_y, showing that r is a kind of average product of standardised values.

  • Eq. 13b.1c writes each observation as a z-score for x times a z-score for y, then averages. This makes it clear that r is unitless.

All three give the same number. If you are given S_XY, S_XX, and S_YY (or the raw sums to compute them), you can find r.

Shortcut via R-squared (Simple Linear Regression Only)

In simple linear regression, r-squared = R-squared. So you can get r by taking the square root of R-squared and applying the sign of the slope (or of S_XY, which is the same thing).

  • If b1 is negative, r is negative.

  • If b1 is positive, r is positive.

If you already have the ANOVA table and the slope, you do not need to calculate r from scratch.

This shortcut does not work in multiple regression.

Properties of Correlation

  • r always falls between -1 and +1.

  • Switching x and y does not change r (unlike regression, where the slope changes).

  • r has no units, because each variable is divided by its own standard deviation.

  • r = +1 or r = -1 only when every point lies exactly on a straight line. This is rare with real data.

  • r = 0 means no linear association. It does not mean "no association at all"; there could be a curve.

Sign of the Correlation

The sign of r equals the sign of S_XY, which also equals the sign of the slope b1.

Positive r: when x is above its mean, y tends to be above its mean too (and below/below). Points fall mostly in Quadrants I and III of the mean-centred scatterplot.

Negative r: when x is above its mean, y tends to be below its mean. Points fall mostly in Quadrants II and IV.

Interpreting the Strength of r

Rule of thumb (these are guidelines, not hard boundaries):

  • |r| < 0.5: weak linear association

  • 0.5 <= |r| < 0.8: moderate linear association

  • |r| >= 0.8: strong linear association

If you are working with R-squared instead, square those cutoffs: R-squared < 0.25 is weak, 0.25 to 0.64 is moderate, above 0.64 is strong.

These thresholds depend on the field. In ecology, r = 0.4 may be noteworthy; in a physics lab, r = 0.9 might be considered poor.

Cautions About Correlation

  • Both variables must be quantitative. Correlation does not apply to categorical variables.

  • Correlation measures only linear relationships. All the scatterplots in Fig. 13b.6 (parabola, sine wave, circle, X-shape, etc.) have r = 0 despite having clear associations.

  • Correlation is not resistant to outliers. A single extreme point can inflate or collapse r.

  • Correlation is not a complete summary of bivariate data. Different datasets can produce the same r (Anscombe's quartet, Fig. 13b.7). Always plot first.

  • Correlation (association) does not imply causation. Additional information beyond statistics is needed to establish a causal link.


Formulas

r = \frac{S_{XY}}{\sqrt{S_{XX} \cdot S_{YY}}}

where:

S_{XX} = \sum x^2 - \frac{(\sum x)^2}{n}
S_{YY} = \sum y^2 - \frac{(\sum y)^2}{n}
S_{XY} = \sum xy - \frac{(\sum x)(\sum y)}{n}

Alternative forms:

r = \frac{\sum[(x_i - \bar{x})(y_i - \bar{y})]}{(n-1)\, s_x \, s_y}
r = \frac{1}{n-1} \sum \left[\left(\frac{x_i - \bar{x}}{s_x}\right)\left(\frac{y_i - \bar{y}}{s_y}\right)\right]

Relationship to R-squared (simple linear regression only):

r = \pm \sqrt{R^2} \quad \text{(sign matches } b_1 \text{)}

Real-World Applications

Correlation is used in any field where two measurements are collected together and you want a quick summary of whether they move in step. In engineering, you might correlate iodine value with cetane number to decide whether a cheap lab test can stand in for an expensive one. In public health, researchers correlate variables like physical activity and blood pressure to identify risk factors worth investigating further.


Common Misconceptions

  • Students often say "r = 0 means x and y are not associated." It does not. It means there is no linear association. A perfect U-shape or circle will also give r = 0.

  • Students sometimes assume a high r proves the data is linear. It does not. A tight curve or a cluster with one far-flung point can also produce a high r (Anscombe's Quartet).

  • Students frequently confuse r with R-squared. If r = -0.8, R-squared = 0.64, not 0.8. And in multiple regression, r is not simply the square root of R-squared.

  • Students occasionally treat correlation as a measure that works on any data type. It requires two quantitative (numerical) variables. It cannot be used if one of the variables is categorical.


Why It Matters / Exam Flags

  • You will be asked to calculate r from provided sums (S_XX, S_YY, S_XY) or from the ANOVA table and slope. Know both routes.

  • Expect at least one question testing whether you understand that r = 0 does not rule out non-linear association.

  • The rule-of-thumb thresholds (0.5, 0.8) for weak, moderate, and strong are commonly tested. Know them, and know that they apply to the absolute value of r.

  • The fact that switching x and y does not change r (but does change the regression equation) is a favourite exam distinction.

  • "Correlation does not imply causation" appears on nearly every intro stats exam. Stating this fact alone, however, does not answer a question that asks whether causation exists in a specific scenario; you need additional reasoning.


Quick Self-Test

  1. True or False: The correlation coefficient can take any value from negative infinity to positive infinity. ___

  1. True or False: If r = 0, then there is no relationship of any kind between x and y. ___

  1. Fill in the blank: For simple linear regression, r = plus or minus the square root of ___.

  1. True or False: Switching the roles of x and y changes the value of the correlation. ___

  1. True or False: The correlation is resistant to outliers. ___

Answers: 1. False (r is always between -1 and +1). 2. False (r = 0 means no linear relationship; a non-linear one may exist). 3. R-squared. 4. False. 5. False.


Practice Q&A

Q: You are given S_XX = 6802.769, S_YY = 377.174, and S_XY = -1424.41. Calculate the sample correlation r.

A: r = -1424.41 / sqrt(6802.769 * 377.174) = -1424.41 / sqrt(2565330.7) = -1424.41 / 1601.66 = -0.8892. The correlation is approximately -0.89.

Q: An ANOVA table gives R-squared = 0.7908 and the slope is b1 = -0.209. What is r?

A: r = -sqrt(0.7908) = -0.8893. The sign is negative because the slope is negative.

Q: Would the correlation between the age of a used car and its resale price be positive or negative? Explain.

A: Negative. As the age of a car increases, its resale price tends to decrease, so y decreases as x increases.

Q: A dataset of seven points produces r = 0.95, but a scatterplot reveals one far-flung point at the upper right and the other six points clustered with no clear trend. Is the correlation meaningful?

A: No. This is a cluster situation (similar to Anscombe's Plot 4). The high r is driven by the outlier pulling the line toward it. You need to look at the scatterplot, not just the number.

Q: True or False: The sample correlation measures the strength of any relationship between two continuous variables.

A: False. It measures the strength of a linear relationship only.


Connections to Other Topics

This connects to Chapter 13a (Simple Linear Regression) because the correlation is derived from the same S_XX, S_YY, and S_XY quantities used to fit the regression line and build the ANOVA table. It also connects to the broader principle of assumption-checking: just as you check normality before a t-test, you must check linearity before using r. The next set of notes (Part 2, Regression Diagnostics) covers exactly how to verify the assumptions graphically.


Related Terms / Search Tags

Pearson correlation coefficient, sample correlation, r value, r-squared, coefficient of determination, linear association, positive correlation, negative correlation, uncorrelated, strength of linear relationship, scatterplot, S_XX, S_YY, S_XY, Anscombe's quartet, outlier effect on correlation, correlation vs causation, simple linear regression, bivariate data, standardised scores, z-scores