Topic 4.1 · teacher page

Running sampling, bias and outliers

Answers, what each wrong answer tells you, and the one demonstration that makes the lesson land.

The one thing to do with the widget

Ask them to predict before you move the slider. Most will say both the mean and the median will go up. Then push the ninth student out to 95 and let them watch the median sit perfectly still.

That forty seconds is the argument for reporting a median alongside a mean, and it is far more convincing than the sentence that says so. With these nine values the median can only ever land on 28, 31 or 33, however far out the ninth goes, and asking why is a good two minutes.

Answers

QuestionAnswer
1. Stratified, Year 1130. 450 of 900 is a half, and half of 60 is 30.
2. Upper outlier boundary48. IQR is 12, and 30 + 1.5 × 12 = 48.
3. The 500-person surveyB. A large sample from one place is still only representative of that place.
4. The 17.2 cm heightC. Investigate and correct it. Almost certainly 172 cm mistyped.

Worked example answers on the student page: the stratified sample gives 20 from Year 12, and the outlier test gives an upper boundary of 54, so 58 is flagged.

What each wrong answer tells you

The page responds to the specific answer rather than saying "try again", so you know what a student has already been told before they put a hand up.

They enterWhat it means, and what the page says
20 (Q1)They solved for the wrong year group. Reading error, not a method error. The page tells them so and sends them back to the question rather than the method.
27 (Q1)They used 450 out of 1000 rather than 900. Almost always a student who added the three year groups carelessly. Worth catching because it will recur in every stratified question.
42 (Q2)They added one IQR, not one and a half. The most common slip on this topic by a distance.
36 (Q2)Multiplier of 0.5 rather than 1.5. Usually a half-remembered rule.
0 (Q2)They computed the lower boundary. They know the method and did not read which end was asked for.
"500 is not large enough"The one that matters. This is the central misconception of the whole sub-topic, and the page names it: bias is about who could be asked, never how many were. If a student picks this, stop and deal with it.
"Delete the 17.2"They have treated the outlier rule as permission. Worth a whole-class correction: the rule flags a value for attention and never authorises quiet deletion.

Where the marks go

Stratified sizes are close to free marks, and they are lost to rounding. Insist on the fraction. A student who writes 0.333 and multiplies will meet a case where the answer is not a whole number and will not know what to do.

Outliers are one mark for the calculation and one for the judgement. "58 is an outlier so I removed it" scores worse than "58 is an outlier, so I checked it, and it is genuine".

Bias wants the mechanism: who was more likely to be asked, and who could not be asked at all. "The sample is too small" earns nothing and is the most frequent answer given.

A possible order

 What is happening
1The widget. Predict, then move, then the question about why the median can only take three values.
2Population and sample, the five techniques, and the quota versus stratified distinction. This is the part they will be examined on and the part that reads as dull, so keep it brisk.
3The stratified worked example, then question 1. Circulate: the 450-out-of-1000 slip is visible from across the room.
4Outliers, the rule, the worked example, question 2.
5Questions 3 and 4, which are judgement rather than calculation, then discuss them as a class. Do not set these for homework; the conversation is the lesson.

Two things not to say

Do not say "a bigger sample would fix it" even as a throwaway, because it is the exact misconception the sub-topic exists to remove, and they will remember your version over the page's.

Do not describe an outlier as "a value that is wrong". The whole of question 4 depends on the difference between unusual and incorrect, and once a class has heard outlier and error used as synonyms it is very hard to undo.

Questions to set

Three tiers, ramping the way practice should: the method on its own, then the method inside something real, then a challenge. Set the tier the class in front of you needs rather than one undifferentiated sheet. Answers are given so these can go straight onto a board.

1Fluency

The method on its own, with friendly numbers. Set these first and move on quickly once they are secure.

  1. Stratified sizesA school of 1200 has 500 in Year 10, 400 in Year 11 and 300 in Year 12. Find a stratified sample of 60.
    25, 20 and 15. Each is its share of 1200 times 60.
  2. Systematic intervalFor a systematic sample of 60 from 1200, what is the sampling interval?
    Every 20th student, from a random start in the first 20.

2In context

The same skill inside a real situation, where the first job is working out what is being asked.

  1. Name the methodA teacher samples by taking the first 30 students through the gate one morning. Name the method and the group it misses.
    Convenience sampling. It misses anyone who arrives late or travels differently, and arrival time is probably linked to what is being measured.
  2. Choose a methodYou want opinions on a new canteen menu from a school of 1200 across three year groups with different lunch slots. Name a method and justify it.
    Stratified by year group, because the lunch slot differs by year and so will the experience. A simple random sample could by chance miss a whole slot.

3Challenge

Reasoning, working backwards, or spotting an error. These are where the top grades are decided.

  1. Why size is not enoughA student surveys 500 people outside one Bangkok mall and writes that the sample is representative because it is large. Explain the fault.
    Size fixes random error, not bias. Only people at that mall at that time could be asked, so the sampling frame was already narrow. A biased sample of 500 is more confidently wrong than a biased sample of 50.
  2. Design against a known biasA survey by school email will over-represent students who check email. Suggest a change and say what it costs.
    Sample from the roll and chase non-responders in person, which costs time but removes self-selection. Any fix that keeps volunteering keeps the bias.

Practicalities

Works on a phone. Nothing a student types is saved or sent anywhere, so there is no account, no data to consent to, and nothing to lose if they close the tab except their own place. Equally there is no record: if you want evidence, take their written answers to questions 3 and 4 on paper.