Why Spearman exists, what ties do to the ranks, and the justification that carries the mark.
Drag the outlier and read both coefficients at once.
Pearson collapses from 0.9789 to 0.7924 while Spearman does not move off 1. Say the number out loud: one point out of eight has cost Pearson nearly a fifth of its value.
That is the entire argument for Spearman, and it is far more convincing as a live number than as a sentence in a textbook. Keep dragging back and forth; the stillness of the Spearman readout is the thing they remember.
Everyone can rank 3, 7, 9. The examined case is 7, 7, 9, where both sevens take the mean of the ranks they occupy, so 1.5 and 1.5, and the nine takes 3 rather than 2.
Get this wrong and every subsequent difference is wrong. Make them write the ranks under the data in a row, not in their heads.
| 1. rs on y = x² | As x goes up, y goes up every single time, so the ranks match perfectly and Spearman is exactly 1. |
| 2. Rank for each 7 | They occupy ranks 1 and 2, so both get 1.5, and the 9 gets rank 3. |
| 3. Which to quote | B, Spearman. It uses only the order, so one stray value barely shifts it. The justification is where the mark is. |
1 markRanking both variables correctly, including shared ranks for ties.
1 markThe coefficient itself, usually straight from the calculator.
1 markA reason for preferring one coefficient, in the context of the data. “Because of the outlier” alone is thin; name what the outlier does.
| They give | What it means |
|---|---|
| 0.9789 (Q1) | Gave Pearson. On y = x² the relationship is perfect but not linear, which is exactly the case that separates the two coefficients. |
| 1 (Q2) | Gave the first tied rank to both and then used 2 for the nine. The ranks no longer sum correctly, which is the check to show them. |
| 2 (Q2) | Averaged the wrong pair, or ranked from the top. Either is worth distinguishing before correcting. |
| "A, Pearson" (Q3) | They default to the familiar one. Send them back to the widget and make them read the two numbers again. |
| "Neither, remove the outlier" (Q3) | Tempting and sometimes right, but it needs justifying as a data decision, not used as an escape. If they say this, ask what they would write in an IA to defend it. |
"Which one do I use?" Linear and clean: Pearson. Monotonic but curved, ordinal data, or an outlier present: Spearman. Say which and why, because the why is the mark.
"Can Spearman be 1 when Pearson is not?" Yes, and the widget is sitting on that case. Any strictly increasing relationship gives Spearman 1, however curved.
"Does Spearman mean causation?" No, no more than Pearson does. Worth saying, because the unfamiliar name makes it sound stronger than it is.
| What is happening | |
|---|---|
| 1 | Drag the outlier. Read both numbers. Say the fifth-of-its-value line. |
| 2 | Ranking by hand, including the tie case, written under the data. |
| 3 | The three questions. |
| 4 | A real data set with a defensible outlier, and a written justification each. |
| 5 | Flag that the IA rewards exactly this judgement, and that quoting both coefficients with a reason is a strong move. |
Do not say Spearman is “the one for small data sets”. Size is not the criterion; shape and outliers are.
Do not accept “because it is better” as the justification. It earns nothing and it is not true in general.
Three tiers, ramping the way practice should: the method on its own, then the method inside something real, then a challenge. Set the tier the class in front of you needs rather than one undifferentiated sheet. Answers are given so these can go straight onto a board.
The method on its own, with friendly numbers. Set these first and move on quickly once they are secure.
The same skill inside a real situation, where the first job is working out what is being asked.
Reasoning, working backwards, or spotting an error. These are where the top grades are decided.
Works on a phone. Nothing is loaded from any other site, so it runs behind a school firewall, and nothing a student does is saved or sent anywhere.