🧙 Wiz Kids
LearnTeaching Digital Skills

AI-text detectors don't work: the evidence schools need

Evidence grade: STRONG — and unusually one-sided. The claim "AI-text detectors are unreliable for consequential decisions" is supported by benchmark studies, documented bias findings, trivial-evasion demonstrations, and the market leaders' own retreats. Schools keep buying them anyway, which is why this page exists: it's the citation to bring to the meeting.

The market response to the homework problem was inevitable: tools claiming to detect machine-written text, sold with accuracy percentages. Schools should not stake any consequential decision on them. That's not a cautious hedge — it's what the evidence, the theory, and the vendors' own behavior all say.

Why it fails in principle

Detection assumes machine text carries a stable fingerprint. It doesn't, for structural reasons: generative models are trained to write like people and improve at exactly that with every release, so detection chases a target whose whole optimization pressure is escape; "human-sounding" and "machine-sounding" are overlapping distributions, not categories — competent-but-plain human writing (a diligent student following taught structures, a non-native writer using learned formulas) sits precisely where the models live; and trivial transformations (paraphrase, "make it sound younger," light manual editing) collapse whatever signal existed. Watermarking — a proposed cryptographic fix — requires generator cooperation and dies on paraphrase, and years of proposals have shipped nothing schools can rely on.

Why it fails in measurement

What to do instead

The alternatives are the previous page's repertoire — certify in-class, assess process, design AI-assumed tasks — plus one rule for the tools that do exist: a detector score may prompt a conversation, never constitute evidence. "Tell me about this piece" resolves in two minutes what the score can't resolve at all, works at primary age better than anywhere, and treats the child as a person rather than a probability.

What the evidence doesn't say

In the classroom

  1. Don't buy detectors; if the school has one, demote it to conversation-prompt status by written policy.
  2. Protect EAL students explicitly — the bias finding belongs in any staff briefing that mentions detection.
  3. Invest the detector budget in assessment redesign — in-class certification and process evidence, which work regardless of what models ship next year.

How Wiz Kids applies this

Nothing to detect: certification is live in-app performance, so the question "did a machine write this?" never arises for anything we assess. The curriculum's contribution is the literacy — children learn what machine text is and isn't good for, including, honestly, that adults can't reliably spot it either; the Hall of Mirrors teaches checking claims, not vibing authorship.

References


© Glu IO Pty. Ltd. — Wiz Kids (wiz.kids). Link freely; republication requires permission — see terms. Found an error in our reading of the research? We correct fast: tell any teacher piloting Wiz Kids.