Pearson's r, the regression line, and the two marks that are not about calculating anything.
Eight students recorded hours revised and the mark they scored. The correlation looks convincing. Drag one point a long way from the others and watch how convincing it stays.
Eight draggable points. Nothing is stored and nothing is sent.
Pearson's r measures how close the points lie to a straight line, and nothing else. It runs from −1 to 1.
It says nothing about the steepness of the relationship. A gentle line with points tightly on it has r near 1. A steep line with scattered points has r near 0. Gradient and r answer different questions: how much, and how tightly.
r only detects linear association. Data lying perfectly on a parabola can have r near zero. A near zero r means "not linear", not "no relationship", which is why you sketch the scatter before you trust the number.
With the eight points as they start, the means are x̅ = 5 and y̅ = 62.
Check that against the readout above before you move anything. If the two ever disagree, it is the hand calculation to distrust.
Interpreting the gradient. Each extra hour of revision is associated with about 4.6 more marks. In context, with units. Writing "the gradient is 4.60" restates the answer and earns nothing.
Knowing the limit. Predicting from 40 hours gives 4.6 × 40 + 39 = 223, a mark above 100. The model was built from data between 1 and 9 hours, and outside that range it is extrapolation. The impossible value is the proof.
Association is not cause. The data cannot tell you that revising causes higher marks. A student who revises nine hours is probably different in other ways too. Saying so is worth a mark; reciting "correlation does not imply causation" without naming an alternative explanation usually is not.
1. A cafe in Bangkok finds n = 8.2t − 118, where t is the maximum temperature in degrees Celsius and n is iced coffees sold. Estimate sales at 31 degrees.
2. Using the same model, the owner asks about 5 degrees. Why should you refuse to answer with this model?
3. Two data sets both have r = 0.95. Set A has a regression gradient of 0.4; set B has a gradient of 12. What does that tell you?
Finding the line is two marks and almost nobody loses them, because the calculator does it. Do not spend your revision there.
One mark is interpretation in context with units. One is knowing when the model stops applying. Those two are where the grades separate, and neither requires a calculation.
Do not round the gradient and then use the rounded value to find the intercept. With real data that loses accuracy marks.
These pages are free and stay free, but they are general and your IA is not. Send me your research question, or whatever exists so far, and I will tell you in writing whether the topic has a ceiling on it, where the marks are going, and what to change first. That costs nothing and it comes back within 24 hours.
Written by a serving IB Diploma and Career-related Programme Coordinator and Head of Mathematics, who reads internal assessments across every subject group every year. If you then want the whole draft reviewed properly against all five criteria, that is the paid one, and it is refunded if it does not name at least three specific things to fix.
Get a free verdict Full written review, $99
I never write any part of it. Not a sentence, not a calculation, not your data. Under 18: a parent buys this and the thread is with them. I do not work with students at my own school.