Building AI literacy into a science course
A grade 11 and 12 Earth Science course that teaches AI literacy as a set of measurable competences inside real science work, not alongside it.
The course
A redesigned grade 11 and 12 Earth Science course, taught across three sections, built to teach and measure AI literacy inside real science work rather than alongside it.
Every unit ends in something a real audience could read: a source dossier, a habitability report for a NASA-style evaluation committee, a climate brief for Chennai policymakers, a position paper argued at a stakeholder roundtable. Students defend each one aloud to a panel. The science is anchored in local ground wherever it can be, from the Cooum river to the Pallikaranai marsh to the Cauvery water dispute.
A familiar discipline
None of this is entirely new. For years we taught students to treat Wikipedia the same way: a good place to start research, never a source to cite in a final reference list. AI gets the same treatment, with documentation added. Naming AI as a source that is usually right but never guaranteed is what conditions a student to keep checking.
AI as a source, not an oracle
Students apply the same four-part source-evaluation framework to an AI response that they apply to a peer-reviewed paper. Every substantive AI use is documented on a template whose hardest field asks for one specific thing the AI got wrong. Six competences from the EU-OECD AI Literacy Framework are instrumented across the year through a pre and post diagnostic, per-block self-ratings, and coded student reflections. The instruments are built and the study runs this year.
Every student must name one specific thing the AI got wrong. If you cannot find one, you were not checking carefully enough. That single requirement turns passive users into critical ones.
Try it: spot what the AI got wrong
Here is a confident AI answer about a real earthquake. One specific fact is wrong. Finding it is the move every student makes, every time. Click the claim you would check.
The 1964 Great Alaska earthquake struck near Anchorage on . It registered a moment magnitude of , making it the most powerful earthquake in . The shaking lasted about and triggered a tsunami that reached the coast of California.
That is the whole skill: treat the AI as a source, verify the confident specifics, and document what it got wrong.
Language support, by design
For multilingual learners, the course makes language support visible rather than covert. Vocabulary is pre-taught, sentence stems sit on every task and are offered to everyone rather than assigned by level, and translation is a legitimate, documented step rather than something to hide. The reading tool for the course's hardest primary source carries definitions in six languages, AI-generated and deliberately left unverified, so that the students who speak those languages are the ones who correct them. In that task, the multilingual student is the expert in the room.
The framework it is anchored to
The design is anchored to a published international framework rather than a rubric invented for the occasion. That matters because it makes the claims checkable by someone who has never seen this course.
doi.org/10.1787/65cd27d4-en
Four domains, nineteen competences, each built from knowledge, skills and attitudes. Ethics runs through all four rather than sitting in one.
Six of the nineteen are instrumented. Thirteen are taught opportunistically and not measured. Choosing six is the design decision. Measuring all nineteen would mean measuring none of them properly.
Where the six live
Every measured competence is tied to an artefact the course already produces, so the evidence is a by-product of the work rather than an extra task bolted on. Unit 0 is measured across the board because the pre-diagnostic baselines all six.
| Competence | U0 | U1 | U2 | U3 | U4 | U5 | U6 | U7 |
|---|---|---|---|---|---|---|---|---|
| Engage 3Whether AI output should be accepted, revised or rejected | M | T | T | T | T | M | T | T |
| Engage 5How AI systems consume energy and natural resources | M | · | · | · | M | · | · | P |
| Create 3Directing AI to elicit feedback and refine results | M | R | R | R | R | M | · | · |
| Manage 1Deciding whether to use AI at all, given the task | M | T | T | T | T | M | T | T |
| Manage 3Decomposing a problem to decide where AI belongs | M | · | · | R | · | M | · | · |
| Shape 2Evaluating AI systems against defined criteria and test cases | M | · | · | · | · | M | · | · |
Three worth naming. Manage 1 asks students to defend a decision not to use AI. Shape 2 has them test whether a stakeholder bot fairly represents the person it claims to speak for. Engage 5 turns AI's own energy and water footprint into course content, with the Chennai and Tamil Nadu angle.
What this can and cannot tell you
With three non-randomised sections, this is design-based action research. It is built to inform decisions at one school and to sharpen my own practice. It is not generalisable and I will not present it as though it were.
It is not a head-to-head platform comparison
The two platforms have different purposes and afford different interaction modes, so a winner claim would be confounded. I am not going to make one.
Equity guardrail
All three sections get an equivalent learning experience and identical assessment. Consent affects only whether a student's data enters the analysis. It never affects their activities, their support, or their grade.
Confounds, named before anyone asks
- Existing fluency is uneven, because one platform is already in school-wide use.
- Novelty effect in the section using the newer platform.
- The platforms drive different pedagogy, not just different interfaces.
- Two teachers deliver three sections, not one.
- Teacher-as-researcher bias: the person who designed the study also teaches one of its sections.
- The diagnostic is built for this course, not a validated published scale.
- Both vendors ship changes during the study window.
- Students mature between August and the post-diagnostic for reasons unrelated to either platform.
What is done about the worst of them
For teacher-as-researcher bias, the one with most exposure: a second coder scores a ten to fifteen percent sample blind to section, and inter-coder agreement is reported alongside any result.
Could I explain to my teacher exactly what I did, and would I be comfortable doing so?
Where else this fits
The framework is subject-agnostic. I built it in Earth and Environmental Science, but the source discipline, the documentation habit, and the visible-by-design language support transfer to any course, and I carry them forward into new material. In the end it is less about a tool and more about a stance: AI as a source to be evaluated, never a substitute for judgment.