Assessing digital skills: performance tasks beat quizzes
Ask a child "Which key combination copies selected text?" and you measure recognition of a fact about a skill. Ask a child to copy this line into that document and you measure the skill. For most school subjects the gap between those two is a nuance; for digital skills — which are overwhelmingly procedural — it is the whole question. A pupil can know Ctrl+C the way one knows a capital city and still hunt through menus for a minute; another can be fluent at the keyboard and fail the quiz item's wording.
The validity argument
Assessment validity means the score reflects the construct you claim to measure (Messick's classic framing, 1995). The construct "can attach a file to an email" is an ability to do; a selected-response item about attachment icons measures, at best, a correlate — contaminated by reading load, test-wiseness, and recognition-vs-recall gaps, and blind to the fluency dimension entirely (knowing of a method ≠ producing it in three seconds mid-task — automaticity is invisible to quizzes). This is why the field's serious measurement efforts went performance-based: ICILS — the international benchmark for exactly these skills — assesses students in live software environments performing authentic tasks, not answering questions about software (Fraillon et al., 2019).
The classic objections to performance assessment (Shavelson and colleagues' work documents them): it's expensive to administer, hard to score consistently, and task-specific — performance on one task generalizes less than test builders hope, so you need many tasks. All true. All largely dissolved for this subject by the medium itself: the computer can present the task, verify the outcome, and log the process — administration is free, "scoring" is outcome-checking (the file either arrived attached or didn't), and many short tasks are cheap, which answers the sampling problem too.
Doing it at primary scale
- Define outcomes, verify outcomes. "The document contains the pasted line," "page 2 only was printed," "the search used quotes." Process variety is fine — outcomes are the check.
- Many small tasks beat one grand project for measurement (the task-sampling result); the grand project earns its place for integration and motivation, not reliability.
- Measure fluency where it matters: for foundational skills (typing, common shortcuts), pace distinguishes automatized from laboriously-recalled — measured kindly and privately, never as a public race (leaderboards).
- Keep it formative. Black and Wiliam's landmark review (1998) locates assessment's leverage in feedback that changes what happens next; every performance check that routes a child to the right next practice is worth ten that produce a grade. Doubly so because a performance check is retrieval practice — the assessment teaches.
- Ten-minute audit for new cohorts: type a paragraph, save it findably, email it attached. Three outcomes, most of the fundamentals, no quiz sheet (the audit that ends "they're digital natives" debates).
What the evidence doesn't say
- It doesn't banish selected-response entirely — judgment constructs (is this a scam? which reply is kind?) are legitimately assessed by choosing among options, because choosing is the performance for judgment skills.
- It doesn't make observation obsolete: posture, eyes-on-keys, mouse grip — some process matters and only a watching teacher sees it.
- It doesn't validate any specific vendor's checks, ours included, without asking the one question: could this be passed without the skill? (could it be passed without remembering? is the same question for reviews).
How Wiz Kids applies this
Every assessment in the product is a verified performance: simulated applications check outcomes ("the gem files are in the chest," "the total uses SUM," "only necessary cookies were accepted"), with process left free and efficiency scored separately as stars. Judgment skills use scenario choices — because there, choosing is the doing. Mastery, placement and review all reuse the same principle: demonstrate, don't describe.
References
- Messick, S. (1995). Validity of psychological assessment. American Psychologist, 50(9).
- Shavelson, R. J., Baxter, G. P., & Gao, X. (1993). Sampling variability of performance assessments. Journal of Educational Measurement, 30(3).
- Black, P., & Wiliam, D. (1998). Assessment and classroom learning / Inside the black box. Assessment in Education, 5(1).
- Fraillon, J., et al. (2019). IEA International Computer and Information Literacy Study 2018 — the performance-based model at international scale.
© Glu IO Pty. Ltd. — Wiz Kids (wiz.kids). Link freely; republication requires permission — see terms. Found an error in our reading of the research? We correct fast: tell any teacher piloting Wiz Kids.