🧙 Wiz Kids
LearnResearch Digests & Reference

Effect sizes for teachers: what d actually means

Evidence grade: methodological reference. Effect sizes are arithmetic, not opinion; the contested part is interpretation benchmarks, where this page follows Kraft's (2020) field-realistic analysis over the older lab-derived conventions — and explains why the difference matters for anyone reading vendor decks.

Every serious claim in this library eventually rests on a number like d = 0.4, so five minutes on what that means buys years of better reading. An effect size (Cohen's d) expresses a difference between groups in standard-deviation units — d = 0.5 means the intervention group's average sits half a standard deviation above the control's. Its whole point is comparability: unlike "improved by 12 marks" (12 out of what? on which test?), d travels across studies. That portability is why meta-analyses speak in it — and why it's the number to hunt for whenever "significant" is doing the marketing.

The benchmark problem (and the 0.4 myth)

Cohen's old conventions — 0.2 small, 0.5 medium, 0.8 large — were explicitly rough, derived largely from lab psychology. Education's popular version, Hattie's "0.4 hinge" (effects above 0.4 worth pursuing), has been influential and is methodologically indefensible: it averages across meta-analyses of wildly different quality, outcome types and populations, treating lab vocabulary studies and year-long school reforms as the same currency (the critical literature — Wecker, Slavin and others — has documented the aggregation problems thoroughly). The correction that matters practically is Kraft (2020): for causal field studies with standardized achievement outcomes, his schema reads under 0.05 as small, 0.05–0.20 as medium, and 0.20+ as large — about a third of education RCT effects land below 0.05 — and the interventions education actually celebrates (high-dosage tutoring among them) live in that range. The upshot inverts the folk reading: a "small" effect from an honest field RCT frequently outranks a "large" effect from a lab study — because the field number survived reality.

The four questions that scale any d

  1. Measured on what? Narrow, researcher-designed tests inflate d relative to broad standardized outcomes — same intervention, different number.
  2. Against what control? "Beat doing nothing" and "beat ordinary teaching" are different claims wearing the same d.
  3. At what cost? d per dollar and per hour is the deployable metric — a d = 0.1 that's free and scales (a retrieval-practice tweak) beats a d = 0.3 that consumes a term.
  4. For whom? Averages hide distributions — word processing's d concentrates in struggling writers; an average can be honest and still misdirect your decision.

Worked examples from this library

The spacing and retrieval effects: robust, replicated, lab-and-field — the rare educational findings that hold size across contexts, which is why we build on them structurally. Growth mindset: the Nature-published national experiment found ~0.1 grade-point effects concentrated in lower-achieving students — small-and-real, honestly sized, useful because cheap. And the vendor-deck pattern to recognize: a large d from a small, short, narrow-outcome, vendor-run study is the least impressive large number in education — every scaling question above deflates it.

What this page doesn't claim

References


© Glu IO Pty. Ltd. — Wiz Kids (wiz.kids). Link freely; republication requires permission — see terms. Found an error in our reading of the research? We correct fast: tell any teacher piloting Wiz Kids.