AI-text detectors don't work: the evidence schools need
The market response to the homework problem was inevitable: tools claiming to detect machine-written text, sold with accuracy percentages. Schools should not stake any consequential decision on them. That's not a cautious hedge — it's what the evidence, the theory, and the vendors' own behavior all say.
Why it fails in principle
Detection assumes machine text carries a stable fingerprint. It doesn't, for structural reasons: generative models are trained to write like people and improve at exactly that with every release, so detection chases a target whose whole optimization pressure is escape; "human-sounding" and "machine-sounding" are overlapping distributions, not categories — competent-but-plain human writing (a diligent student following taught structures, a non-native writer using learned formulas) sits precisely where the models live; and trivial transformations (paraphrase, "make it sound younger," light manual editing) collapse whatever signal existed. Watermarking — a proposed cryptographic fix — requires generator cooperation and dies on paraphrase, and years of proposals have shipped nothing schools can rely on.
Why it fails in measurement
- False positives at unusable rates: benchmark studies testing commercial detectors against known-human text found error rates wildly inconsistent with marketing claims — and for accusation purposes, the false-positive rate is the only number that matters: a 2% false-positive rate across a school year of submissions manufactures dozens of wrongly-accused children.
- The bias finding: Liang and colleagues (2023) showed detectors flagging non-native English speakers' genuine writing as machine-generated at dramatically elevated rates — formulaic-but-honest prose is what detectors think machines sound like. For classrooms with EAL students this converts an unreliable tool into a discriminatory one.
- The vendor retreat: OpenAI launched its own detector in 2023 and withdrew it within months for low accuracy — the maker of the generator concluding it couldn't reliably detect its own output is the market's most honest data point. Turnitin and peers ship theirs with confidence-eroding disclaimers the sales decks omit.
- The asymmetry of harm: a missed detection costs one assignment's integrity; a false accusation costs a child's trust in school, permanently, with the burden landing (per the bias finding) on predictable children. The same calculus as every screening tool: base rates and error costs, not marketing percentages.
What to do instead
The alternatives are the previous page's repertoire — certify in-class, assess process, design AI-assumed tasks — plus one rule for the tools that do exist: a detector score may prompt a conversation, never constitute evidence. "Tell me about this piece" resolves in two minutes what the score can't resolve at all, works at primary age better than anywhere, and treats the child as a person rather than a probability.
What the evidence doesn't say
- It doesn't say teachers can't notice — a teacher who knows a child's voice noticing a sudden register jump is pattern-recognition with context, legitimately a conversation-starter; the indictment is of scores as evidence, not attention.
- It doesn't predict permanent impossibility — provenance approaches may mature (the same story as screenshots); schools should decide on today's evidence, not roadmaps.
- It doesn't excuse doing nothing — it redirects the effort from detection (dead end) to assessment design (works now).
In the classroom
- Don't buy detectors; if the school has one, demote it to conversation-prompt status by written policy.
- Protect EAL students explicitly — the bias finding belongs in any staff briefing that mentions detection.
- Invest the detector budget in assessment redesign — in-class certification and process evidence, which work regardless of what models ship next year.
How Wiz Kids applies this
Nothing to detect: certification is live in-app performance, so the question "did a machine write this?" never arises for anything we assess. The curriculum's contribution is the literacy — children learn what machine text is and isn't good for, including, honestly, that adults can't reliably spot it either; the Hall of Mirrors teaches checking claims, not vibing authorship.
References
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7).
- Weber-Wulff, D., et al. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19 — the benchmark study of commercial detectors.
- OpenAI (2023). AI classifier withdrawal notice — the vendor retreat, in the vendor's words.
- Sadasivan, V. S., et al. (2023). Can AI-generated text be reliably detected? — the theoretical impossibility results, paraphrase attacks included.
© Glu IO Pty. Ltd. — Wiz Kids (wiz.kids). Link freely; republication requires permission — see terms. Found an error in our reading of the research? We correct fast: tell any teacher piloting Wiz Kids.