Case study · IB Diploma assessment

One rubric, both sides of the desk

American International School Chennai. Biology, Chemistry and Physics, 2025-26 session.
Ran in full

The problem

The internal assessment is the piece of diploma science coursework a school marks itself. Which means the standard is only as consistent as the people applying it, and students only understand it as well as their own teacher happens to explain it.

Across three subjects that produces predictable drift. A student in one class learns that error bars need labelling; a student next door does not hear it until feedback. Two teachers read the same evaluation and land two bands apart, both reasonably, because the strand rewards something neither has pinned down out loud. Nobody is doing anything wrong. The standard is simply living in several heads at once, in slightly different versions.

Part one: teaching the standard

Four workshops for the whole diploma science cohort, across five meetings, before anyone started writing. Biology, Chemistry and Physics in the same room for the same sessions.

Research design

Operationalising variables, choosing a range that will actually produce change, and a feasibility check against the equipment the school owns. Rules of thumb rather than vague advice: spread beats stack, six to eight levels before ten repeats, thirty to sixty data points, pilot three by three first.

Students left with two or three candidate research questions and a one-page design card.

Data analysis

The session where the subjects genuinely differ, and the only one that branches. Physics does propagation and gradient methods; Chemistry does calibration curves and propagated uncertainty; Biology does biological replicates and the difference between standard deviation and standard error. The common spine is what an error bar has to be labelled as, and that a result is an effect size with an uncertainty attached rather than a number on its own.

Conclusion and evaluation

A structure for limitations that removes the guesswork: source, effect with direction and rough size, evidence, and a specific improvement. Worked against real IB-marked exemplars so students could see the difference between a limitation that scores and one that is just an apology.

Feedback and submission

The mechanics that quietly cost students marks. One round of written feedback, per-subject draft and final dates, and the three thousand word limit, said plainly: submit over it and the only useful feedback anyone can give you is to cut, and we cannot tell you where.

The fifth meeting

Students read a complete IA as a group and marked it against the rubric themselves. Then they said which strands were easy to find evidence for and which they struggled with, and only after that did we show them the marks it actually received.

The students did the same exercise the teachers would do in April: mark unfamiliar work against the rubric, then find out how far off you were.

Part two: marking to it

In April the department moderated the session's internal assessments. Every IA was read by two or three teachers, and the assignments crossed subject and division on purpose.

Middle school teachers read diploma IAs. Not as observers. As readers, with marks that counted in the conversation. A colleague who does not teach the course cannot lean on knowing what the student meant, so they are forced to read what the rubric actually says. That is the point, and it is also the fastest way to build a shared understanding of where a grade 6 to 12 science programme is heading.

Marks are entered against strands rather than as a global impression, every band carries a written comment, and readers cannot see each other's scores until both have submitted. The moderator then sees the marks side by side and agrees a final score with a note recording the reasoning.

The tool, from both sides

Built in Apps Script with a spreadsheet behind it and an export for IBIS. Switch between a reviewer marking an IA and the moderator agreeing the final score.

Any teacher in the department. Sees only the IAs assigned to them, and cannot see the other reviewer's marks. Sees every reviewer's marks side by side, agrees the final score, and tracks the department.
moderation / queuemoderation / scoremoderation / comparemoderation / department

My queue

Session 2026 · reviewer

You are one of two or three readers on each of these. You will not see anyone else's marks until moderation.

CandidateSubjectInvestigationStatus
Candidate 01Chemistry HLRate of reaction and catalyst surface areaSubmitted
Candidate 04Chemistry SLEnthalpy change of neutralisationSubmitted
Candidate 09Biology SLEnzyme activity across temperatureIn progress
Candidate 12Physics HLSpring constant from oscillation periodNot started
Candidate 17Physics SLFocal length and lens curvatureNot started

Assignments deliberately cross subject and division. A biology teacher reads a physics IA, and a middle school teacher reads a diploma one. The point is a shared standard, not subject expertise.

Candidate 09 · Biology SL

Reader 2 of 2 · not submitted
Investigation, page 4 of 11
Selected text, with a reader comment attached
Data analysis · uncertaintyError bars are shown but not identified. Standard deviation or standard error is not stated in the caption.
Research design4 / 6
RQ and context5-63-41-2
Methodology considerations5-63-41-2
Methodology detail5-63-41-2
Range is well chosen and justified. Control of pH is asserted rather than described.
Data analysis3 / 6
Conclusion4 / 6
Evaluationnot yet marked
Submit my marks

Marks are made against strands, not against a global impression, and every band carries a comment. The comment is the part that survives into the moderation conversation.

Candidate 09 · Biology SL

Both readers submitted · ready to agree
CriterionReader 1Reader 2AI passAgreed
Research design4444
Data analysisreaders differ by 25344
Conclusion4434
Evaluation3233
Total16131415 / 24
Moderation note

Reader 1 credited the calibration work under data analysis; Reader 2 read it as method. Agreed at 4: the processing is sound but the error bars are unlabelled, which caps the strand. Both readers updated their own notes after the conversation.

The AI pass is a third reader that never sets a mark. It reads the IA against the same strands and its only job is to make disagreement visible: where two humans and a machine all land differently, the strand is being read three ways. The agreed column is always human, always from the conversation.

The disagreement is the useful part. Two readers two bands apart on the same strand says the strand is being read differently across the department, and that is a training conversation rather than a marking one.

Department view

Session 2026
38IAs in the session
2.3readers per IA, average
11flagged for discussion
34agreed and locked
Progress by subject
Biology14 / 14
Chemistry12 / 14
Physics8 / 10
Where readers disagreed most
Evaluation6
Data analysis4
Conclusion1
Research design0
What that told us

Evaluation drew more disagreement than the other three criteria combined. That is not a marking problem, it is a teaching one: the strand asks students to size the effect of a limitation, and readers were split on how much sizing was enough.

It became the first item in the following year's workshop sequence, and the rubric language for that strand was rewritten in the student-facing materials.

Export for IBISExport CSV

The moderation data is a curriculum instrument as much as an assessment one. Where the adults disagree is where the teaching is least settled.

What the disagreement told us

Evaluation drew more reader disagreement than the other three criteria put together. Research design drew almost none.

That is not a marking problem, and treating it as one would have wasted the finding. Research design is concrete, so readers agree. Evaluation asks a student to size the effect of a limitation, and the readers were split on how much sizing was enough, which means the students had no way of knowing either. The strand was under-taught, and the moderation data is what exposed it.

So it became the first item in the next workshop sequence, and the rubric language for that strand was rewritten in the student-facing material. The assessment system fed the teaching, which is the only reason to collect the data at all.

What happened to the marks

Moderated internal assessment grades, Chemistry and Physics
Moderated IABeforeAfter
IA entries2215
Mean grade4.75.6
At grade 5 or above55%87%
At grade 6 or above27%60%
At grade 3 or below18%none

Chemistry and Physics combined, both levels, before and after this programme was introduced. Biology is excluded because the earlier year's data was not available to me. Figures count IA entries rather than students, since some candidates sit both subjects, and each year is a different cohort.

The top band did not move, and I would not expect it to. A seven is scrutinised closely by moderators because it indicates near-perfection, and students who reach one were generally always going to. The change sits in the lower half: the grade four band halved and nothing fell below it. That is where the workshops were aimed. When students understand what the task is actually asking of them, the internal assessment becomes approachable, and that shows up here first.

Where else this fits

Any assessment a school marks itself has this shape. Coursework, portfolios, performance tasks, moderated projects. The mark is only meaningful if the people applying the criteria agree what the criteria mean, and students can only aim at a standard they have been shown rather than told about.

What made this work was doing both halves with the same instrument. The students calibrated against a marked exemplar in September and the teachers calibrated against each other in April, using the same strands and the same language. Neither half would have done much alone. A department that agrees internally but never shows students the rubric just marks its own confusion consistently.