Topic 4.18 · teacher page · Higher Level

Two errors, one line

Why a stricter test is not a better test, and the only thing that improves both error rates at once.

The one thing to do with the animation

Ask them to find a position for the line that makes both errors small.

Let them try. α and β are the two tails on opposite sides of the same line, so every position trades one for the other, and a minute of genuine attempt is worth more than being told.

Then switch to the second control. With α pinned at 5%, increasing n narrows both curves, they overlap less, and β falls on its own. More data is the only thing that improves both, and that is the sentence to leave on the board.

β has no value until the alternative is specific

“H₀ is false” covers a drift to 501 and a drift to 540, and those are missed at completely different rates. So β is not a property of the test alone.

This is why every question asking for β hands you a particular alternative mean. If a student asks “what is β” with no alternative given, the answer is that the question is incomplete, and understanding why is the point.

The answers

1. P(Type I) at the 1% level0.01. It is the significance level, by definition.
2. Power when β = 0.231 − 0.23 = 0.77.
3. Alarm with no fireA, Type I. H₀ is no fire, it is true, and the alarm rejected it.
4. Moving 5% to 1%B, β rises. Demanding more evidence lets more real drifts through.

Where the marks go

1 markNaming the error in the words of the question, not just “Type I”.

1 markGetting the direction right: Type I needs H₀ true, Type II needs it false.

1 markComputing β from the stated alternative with σ/√n, not σ.

What each wrong answer tells you

They giveWhat it means
1 (Q1)Gave 1% as a percentage where a decimal was asked for. Worth a word, not a reteach.
0.99 (Q1)The complement, so they answered “not making the error”.
0.23 (Q2)Gave β back. They have not separated power from the error rate.
B (Q3)Swapped the two. A Type II error is a fire with no alarm, and the fire alarm framing is usually enough to fix it for good.
A (Q4)The assumption that stricter is better in every respect. This is the misconception the whole page targets; send them back to the sliding line.
D (Q4)Thinks β can be zero. It cannot while the two distributions overlap at all.

Other things they will say

"Which error is worse?" It depends on the costs, and that is the honest answer. Screening for an illness and convicting in a trial point in opposite directions, and choosing 5% out of habit is choosing a cost ratio without noticing.

"Why 5%?" Convention, not mathematics. Saying this plainly is better than letting them believe it is derived from something.

"Can I make both zero?" Only with infinite data or no overlap. Every real test accepts both errors at some rate, which is worth knowing before they read a scientific claim.

A possible order

 What is happening
1Let them hunt for a line position that makes both small. Let them fail. Then name the trade.
2The second control, with α pinned, and the more-data conclusion.
3The two-by-two table and the four questions.
4Computing β from a given alternative mean, carefully, with the standard error.
5Two contexts with opposite cost structures, and which error each should protect against.

Two things not to say

Do not describe a 1% test as “more reliable”. It is more reluctant, which is a different thing and the exact confusion being examined.

Do not compute β without stating the alternative mean you used. It is meaningless without it, and the habit matters more than the arithmetic.

Questions to set

Three tiers, ramping the way practice should: the method on its own, then the method inside something real, then a challenge. Set the tier the class in front of you needs rather than one undifferentiated sheet. Answers are given so these can go straight onto a board.

1Fluency

The method on its own, with friendly numbers. Set these first and move on quickly once they are secure.

  1. Type IA test is carried out at the 1% significance level. State the probability of a Type I error.
    0.01, which is the significance level by definition.
  2. PowerFor a particular alternative, β = 0.23. State the power.
    1 − 0.23 = 0.77

2In context

The same skill inside a real situation, where the first job is working out what is being asked.

  1. Classify in contextA fire alarm sounds when there is no fire. In testing language, which error is that?
    A Type I error: rejecting the null, that there is no fire, when it was true.
  2. Move the levelA researcher moves from the 5% level to the 1% level and changes nothing else. State the effect on each error.
    The probability of a Type I error falls to 0.01 and the probability of a Type II error rises. Making one harder to commit always makes the other easier.

3Challenge

Reasoning, working backwards, or spotting an error. These are where the top grades are decided.

  1. Why both cannot shrinkExplain why no position of the decision line makes both errors small, and what does.
    The two distributions overlap, so any line splits that overlap between them; moving it trades one error for the other. Only separating the distributions helps, which means a larger sample or a bigger real effect.
  2. Choose the levelFor a drinking water plant near Bangkok, stopping a working line is expensive and shipping underfilled bottles risks a fine. Explain how you would choose the significance level.
    By the relative cost of the two errors. If a fine and the reputational damage outweigh the stoppage, accept a higher Type I rate to protect against Type II. The level is a business decision, not a statistical constant.

Practicalities

Works on a phone. Nothing is loaded from any other site, so it runs behind a school firewall, and nothing a student does is saved or sent anywhere.