Applications and Interpretation SL, Topic 4. One lesson, written the way a student working alone actually needs it.
This is a sample. It is one lesson from one topic, built to the standard an entire unit would be written to, so that the standard can be judged rather than described.
Eight students recorded how many hours they revised and the mark they scored. The correlation coefficient comes out high, so the conclusion writes itself: revise more, score more. Except the examiner is not asking whether the numbers agree. They are asking what the model lets you claim.
Move any point and watch both the regression line and r respond. Try dragging one point far away from the others, and watch what a single outlier does to a coefficient that looked convincing.
Eight draggable points. Nothing is stored and nothing is sent.
With the original data, the means are 5 hours and 62 marks. The gradient of the least squares line is the ratio of covariance to the variance in x:
Check it against the widget above before you move on. With the points untouched it reads r = 0.989 and y = 4.60x + 39.0, which is what the hand calculation gives. If the two ever disagree, the hand calculation is the one to distrust.
Two marks for the line, and almost nobody loses those, because the calculator produces them. The marks that actually move a grade are the next two.
One mark is for interpreting the gradient in context and with units: each extra hour of revision is associated with about 4.6 more marks. Writing "the gradient is 4.60" earns nothing, because it restates the answer instead of explaining it.
The last mark is for knowing what the model cannot do. Predicting the mark for 40 hours of revision is extrapolation far outside the data, and the honest answer says so. The model gives 4.6 × 40 + 39 = 223, a mark above 100, which is impossible. A student who writes that down and moves on has lost a mark that a student who notices the absurdity and names extrapolation keeps.
Question. A cafe records the daily maximum temperature, t degrees Celsius, and the number of iced coffees sold, n. For their twelve days of data, r = 0.94 and the least squares line is n = 8.2t − 118.
(a) Interpret the gradient in context. (b) Use the model to estimate sales at 31 degrees. (c) The owner asks what sales would be at 5 degrees. Explain why you would not answer with this model.
For each additional degree of maximum temperature, the model predicts about 8.2 more iced coffees sold per day. Context and units, both needed.
Twelve days of data from a cafe that sells iced coffee will not contain a 5 degree day, so 5 lies outside the range of the data and using the model there is extrapolation. The relationship is not guaranteed to continue, and in fact the model gives 8.2 × 5 − 118 = −77, a negative number of coffees, which is impossible. That absurd answer is the evidence that the model has been pushed somewhere it does not apply.
Part (c) is where the grades separate. Saying "because it is too cold" is reasoning about coffee, not about mathematics, and earns nothing. The mark is for naming extrapolation and, better still, for showing that the model produces an impossible value.
Every lesson in a unit is built like this one: the misconception first, an interactive the student can push until it breaks, the working set out one step per line with the equals signs aligned, division written as a fraction rather than a slash, and commentary on where the marks actually move. Assessment and mark schemes are written alongside, not bolted on.
Keith Spencer. Serving IB Diploma and Career-related Programme Coordinator and Head of Mathematics, Bangkok. Thirty years teaching US, UK and IB curricula.