Performance tasks vs multiple choice: what each can actually certify
Ask a class "which shortcut copies text?" and mark the As; or watch each child copy a paragraph and see whose hands know. Both produce scores; they certify different things, and for skills they diverge: children pass recognition items for actions they cannot perform (recognition is easier than recall, which is easier than execution), and — the sneakier direction — fluent performers sometimes fail the verbal item for a thing their hands do perfectly, because automatized skill runs below the vocabulary layer. An MCQ score on a skill is a proxy with a systematic bias toward overrating knowledge and underrating fluency.
The validity frame
Messick's construct-validity question — what is this score evidence of? — settles the design choice per skill type. Motor and procedural skills (typing, shortcuts, file operations): only performance certifies; the construct is execution. Judgment skills (spotting the phish, reply-vs-reply-all): scenario-based selection is legitimate — the construct is choosing, and a well-built scenario is a performance of judgment; this is the honorable use of the selected-response format. Conceptual knowledge (what a URL is): MCQs serve, with the standard craft caveats (plausible distractors, no cue-leakage). The failure mode is the mismatch: certifying doing constructs with knowing items — which is precisely what most digital-skills quizzes do, because quizzes were what paper could afford.
The cost history, and its end
Performance assessment was always valid and always expensive — one assessor watching one child do one task doesn't scale, which is why paper-era certification defaulted to MCQs and why ICILS-style performance studies keep surprising people who trusted quiz-and-survey data. Software dissolved the economics: a simulated inbox can watch every child actually attach the file, every attempt, with per-step evidence — performance assessment at MCQ prices. Post-scarcity, the remaining reasons to certify skills by quiz are inertia and habit.
What the evidence doesn't say
- It doesn't retire selected-response — judgment scenarios and concept checks are legitimate constructs for it; the indictment is skills-by-quiz, not the format.
- It doesn't make all performance tasks good — a performance task can be invalid too (testing window-dressing, tolerating lucky-click success); par and goal design are the craft layer.
- It doesn't cover dispositions — whether a child chooses safe behavior unobserved is beyond both formats; that's culture's territory, not assessment's.
In the classroom
- Sort your constructs: for each reported skill, ask "is this a doing, a judging, or a knowing?" — then match the evidence type; mismatches are where report cards lie.
- Replace skill-quizzes opportunistically — every "which button?" item that becomes a "do it" task upgrades the evidence for free.
- Keep scenario MCQs for judgment, and make the distractors the real temptations (the too-good offer, the reply-all) — distractor quality is the whole game.
How Wiz Kids applies this
The engine is this page implemented: doing-constructs are certified only by doing (17 task engines exist so that attach the file means attaching a file), judgment-constructs run as scenario choices with honest distractors, and the dashboard's claims inherit the validity — a tick means the construct's own evidence type was satisfied, which is what makes it worth a teacher's trust.
References
- Messick, S. (1995). Validity of psychological assessment. American Psychologist, 50(9).
- Fraillon, J., et al. (2019). ICILS 2018 — the performance-assessment tradition exposing what surveys and quizzes miss.
- Shavelson, R. J., Baxter, G. P., & Pine, J. (1992). Performance assessments: Political rhetoric and measurement reality. Educational Researcher, 21(4) — the honest costs-and-validity accounting.
© Glu IO Pty. Ltd. — Wiz Kids (wiz.kids). Link freely; republication requires permission — see terms. Found an error in our reading of the research? We correct fast: tell any teacher piloting Wiz Kids.