How we grade evidence (and why every article carries a grade)
Every article in this library opens with an evidence grade. This page is the ruleset — published so the grades can be argued with, which is the point of having them. A reference source that can't say "the evidence here is thin" cannot be trusted when it says "the evidence here is strong."
Why grade at all
Education is where findings go to be overclaimed. The path is well-worn: a real result in a specific population becomes a conference keynote, becomes a poster in a staffroom, becomes a mandate stripped of every boundary condition — and a decade later the backlash discards the real finding along with the inflation. (The learning styles myth ran this pipeline without even the real result at the start.) The replication crisis made the stakes explicit: when the Open Science Collaboration (2015) attempted 100 replications of published psychology findings, well under half replicated cleanly. Psychology's response — bigger samples, preregistration, multi-lab replications — has made recent, large, preregistered work more trustworthy, and made "one exciting study" less so. Our grades encode that lesson.
The scale
STRONG — The finding has (a) replication across decades, independent labs, and populations including children where we claim classroom relevance; (b) meta-analytic support or an equivalent weight of converging studies; and (c) no live scientific dispute about existence (disputes about mechanism or optimal parameters don't disqualify — we say so in the article). Structural decisions may be built on STRONG findings. Examples: spacing, retrieval practice.
PROMISING — Good studies point one way, but replication is limited, populations are narrow, or effects are design-dependent. Worth applying where costs are low; worth saying "this might not hold" out loud. Example: broad transfer claims for computational thinking.
MIXED — Credible studies genuinely disagree, or large replications found much smaller effects than the famous originals. We report both sides and what survives. Example: growth-mindset interventions after the large-scale replications.
THIN — Little direct research exists; what we offer is triangulated from adjacent evidence and labelled practitioner guidance. THIN is a grade, not an insult — most specific operational questions in education are THIN, and pretending otherwise is how folklore tables get made. Example: typing benchmarks by age.
Grades can also run against a claim (the myth pages are "STRONG — against").
The rules behind the grades
- Citations must be real and load-bearing. Author-year in text, references at the bottom, characterized conservatively. We never fabricate statistics or effect sizes; famous numbers (Bloom's two sigma) are quoted with their caveats attached.
- Children count double. A finding shown only in undergraduates gets flagged when we apply it to nine-year-olds; classroom studies upgrade a grade faster than lab studies.
- "What the evidence doesn't say" is mandatory. Every article names the overclaim adjacent to its finding — inoculation against our own future misquotation.
- Interest is declared. Where an article describes something we build (and therefore profit from readers believing), a position note says so at the top.
- Grades move. New evidence changes grades, and articles change with them.
- Verification before promotion. Before this library is actively promoted, every citation gets a resolution check — and until a page has had one, we treat challenges as probably-right-until-checked.
What this framework is not
It is not GRADE, Cochrane, or a formal systematic-review methodology — those are the right tools for clinical questions and we borrow their spirit (grade the evidence, not the enthusiasm) at a scale a small team can execute honestly. Where a genuine systematic review exists on a question we cover, it outranks our reading, and we cite it.
References
- Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science, 349(6251).
- Guyatt, G., et al. (2008). GRADE: an emerging consensus on rating quality of evidence. BMJ, 336 — the inspiration for grading-as-discipline.
- Dunlosky, J., et al. (2013). Psychological Science in the Public Interest, 14(1) — the model of utility-grading applied to education.
© Glu IO Pty. Ltd. — Wiz Kids (wiz.kids). Link freely; republication requires permission — see terms. Found an error in our reading of the research? We correct fast: tell any teacher piloting Wiz Kids.