IB Mathematics Internal Assessment
Every year I read explorations where a t-test or a chi-squared test has been bolted on near the end, run once, and never mentioned again. The test is not wrong, mathematically. It is just doing no work. Here is how to tell whether your test belongs in your IA, and how to write it up so it earns marks rather than filling space.
The most reliable sign of padding is a student who decided to "use a t-test" or "do a chi-squared" before they had finished deciding what they were actually asking. A statistical test is an answer to a specific question about data you already have. If you cannot state that question in one sentence before you touch the test, the test is not ready to be chosen yet.
Ask yourself: am I comparing two group means, am I checking whether two categorical variables are related, or am I checking whether my data matches a distribution I expected? Each of those is a different question, and each points to a different tool. If your honest answer is "I wanted some statistics in my IA", stop and choose a real question first.
A t-test compares means. It answers questions like: is the average reaction time under one condition genuinely different from another, or could the difference you saw plausibly be chance, given how much your data varies within each group? That only means something if the comparison itself means something to your research question, and if you have thought about whether your two samples are independent or paired, because that choice changes which version of the test is correct.
It also assumes your data is at least roughly suited to it. A t-test on eight values scraped together at the last minute, with no thought given to how they were collected, will produce a number, but that number will not carry the weight a reader expects a hypothesis test to carry. Small, badly chosen samples are the most common reason a t-test looks like padding even when the mathematics is done correctly.
The exploration is about something else entirely, a t-test appears in one paragraph comparing two small, conveniently available samples, a p-value is quoted, and the narrative moves on without asking what the result changes about the investigation.
The student is genuinely asking whether two conditions differ, collected paired data specifically to answer that, and the result of the test changes what they do next, for instance whether they treat the two groups separately in the model that follows.
Chi-squared has two common jobs in an IA. A test of independence asks whether two categorical variables are related, for example whether preference for something is associated with a group a person belongs to. A goodness of fit test asks whether observed counts match a distribution you expected, for example whether the results of repeated trials look consistent with what a fair or unbiased process would produce.
Both versions need enough data spread sensibly across categories. Chi-squared on a handful of observations crammed into cells with tiny expected counts is a common weakness, because the test's own conditions are not being met, and a moderator reading closely will notice that the result cannot be trusted even before checking whether it was interpreted correctly.
This is a popular way to get chi-squared into an AI exploration, but on its own it is closer to a classroom exercise than an investigation. Nothing about the result feeds into a wider question, and there is nothing to reflect on beyond "the die seemed fair" or "it did not".
The categories come from data the student collected for a reason connected to their research question, the expected counts are checked before the test is trusted, and the conclusion is used to say something specific about the relationship being investigated, not just whether the p-value crossed a threshold.
State the question the test is answering, in your own words, before you run it. State your hypotheses clearly and connect them explicitly to that question. Check whatever conditions the test requires, and say that you have checked them, even briefly. Report the result, then spend at least as many words interpreting what it means for your original question as you spent calculating it. If the result changes what you do next in the exploration, say so. If it does not change anything, that is worth asking yourself honestly before you submit, because it is usually the clearest sign of all that the test was never necessary.
None of this needs to be long. A well-handled test that does real work can be three short paragraphs: the question, the check of conditions, and the interpretation. A badly handled test that is padding can run to a full page and still add nothing, because length was never the problem.
If you answered no to any of those, the test may still belong in your IA, but only once you have gone back and fixed the gap. A statistical test in a Maths IA is judged on whether it was the right tool for a real question, not on whether the arithmetic behind it was correct.
Not sure your test is pulling its weight?
Send me the test you are using and what you are trying to show with it. I read real submitted explorations every year, and I will tell you honestly whether it is doing genuine work or whether it has been bolted on, while there is still time to change it.
Get the IA sorted with meMore like this: all Maths IA guides. Related: using your own data in the Maths IA
A new IA guide goes up most nights. If you would rather not keep checking, leave an email and I will send the one that matters that week. No selling, and leave whenever you like.
If you are under 18, use a parent's email. I do not correspond privately with students.
These pages are free and stay free, but they are general and your IA is not. Send me your research question, or whatever exists so far, and I will tell you in writing whether the topic has a ceiling on it, where the marks are going, and what to change first. That costs nothing and it comes back within 24 hours.
Written by a serving IB Diploma and Career-related Programme Coordinator and Head of Mathematics, who reads internal assessments across every subject group every year. If you then want the whole draft reviewed properly against all five criteria, that is the paid one, and it is refunded if it does not name at least three specific things to fix.
Get a free verdict Full written review, $99
I never write any part of it. Not a sentence, not a calculation, not your data. Under 18: a parent buys this and the thread is with them. I do not work with students at my own school.