🧙 Wiz Kids
LearnClassroom Implementation

Assessing digital skills: performance tasks beat quizzes

Evidence grade: STRONG on the validity argument, with known costs. That performance assessment measures procedural skill more validly than selected-response items is close to definitional, and it's how the serious international assessments (ICILS) work. The honest trade-offs — cost, scoring consistency, task-sampling — are well documented too, and technology happens to neutralize most of them for this subject.

Ask a child "Which key combination copies selected text?" and you measure recognition of a fact about a skill. Ask a child to copy this line into that document and you measure the skill. For most school subjects the gap between those two is a nuance; for digital skills — which are overwhelmingly procedural — it is the whole question. A pupil can know Ctrl+C the way one knows a capital city and still hunt through menus for a minute; another can be fluent at the keyboard and fail the quiz item's wording.

The validity argument

Assessment validity means the score reflects the construct you claim to measure (Messick's classic framing, 1995). The construct "can attach a file to an email" is an ability to do; a selected-response item about attachment icons measures, at best, a correlate — contaminated by reading load, test-wiseness, and recognition-vs-recall gaps, and blind to the fluency dimension entirely (knowing of a method ≠ producing it in three seconds mid-task — automaticity is invisible to quizzes). This is why the field's serious measurement efforts went performance-based: ICILS — the international benchmark for exactly these skills — assesses students in live software environments performing authentic tasks, not answering questions about software (Fraillon et al., 2019).

The classic objections to performance assessment (Shavelson and colleagues' work documents them): it's expensive to administer, hard to score consistently, and task-specific — performance on one task generalizes less than test builders hope, so you need many tasks. All true. All largely dissolved for this subject by the medium itself: the computer can present the task, verify the outcome, and log the process — administration is free, "scoring" is outcome-checking (the file either arrived attached or didn't), and many short tasks are cheap, which answers the sampling problem too.

Doing it at primary scale

What the evidence doesn't say

How Wiz Kids applies this

Every assessment in the product is a verified performance: simulated applications check outcomes ("the gem files are in the chest," "the total uses SUM," "only necessary cookies were accepted"), with process left free and efficiency scored separately as stars. Judgment skills use scenario choices — because there, choosing is the doing. Mastery, placement and review all reuse the same principle: demonstrate, don't describe.

References


© Glu IO Pty. Ltd. — Wiz Kids (wiz.kids). Link freely; republication requires permission — see terms. Found an error in our reading of the research? We correct fast: tell any teacher piloting Wiz Kids.