Cognitive load theory: designing for the bottleneck
Every idea a child learns passes through working memory — a bottleneck famously capacious enough for only a handful of novel elements at once (Miller's "seven, plus or minus two," 1956; closer to three or four for genuinely novel material in later estimates). Long-term memory, by contrast, is effectively unbounded, and what's stored there — organized as schemas — passes through the bottleneck as single chunks. Cognitive load theory (Sweller, 1988; Sweller, van Merriënboer & Paas, 1998, 2019) is the systematic working-out of one design question: is the instruction spending the bottleneck on learning, or wasting it?
The theory's useful vocabulary: intrinsic load (the task's inherent complexity — how many elements interact), extraneous load (imposed by presentation, not content — the waste), and the processing that builds schemas (historically "germane"). The designer's job: manage intrinsic, minimize extraneous, leave room for the building.
The effects that matter for teaching
The worked-example effect (Sweller & Cooper, 1985, and a large literature since): novices learn more from studying worked solutions than from solving equivalent problems unaided. Problem-solving search consumes the bottleneck; a worked example spends the same capacity on the solution's structure. The practical form is the fading sequence: worked example → completion problems (partial solutions) → full problems.
The expertise reversal effect (Kalyuga et al., 2003): the worked-example advantage reverses as knowledge grows — for learners with established schemas, studying examples becomes redundant and solving becomes superior. No method is best independent of the learner's current knowledge; sequencing is everything.
The split-attention effect: when understanding requires integrating separated sources (diagram here, explanatory text there), the integration itself burns working memory. Physically integrate them — labels on the diagram, instructions beside the thing they describe.
The redundancy effect: presenting the same information in duplicate simultaneous forms (reading text aloud verbatim while displaying it; decorating a clear diagram with restating text) hurts — processing the duplicate is pure extraneous load. More supports are not automatically kinder.
The guidance argument: Kirschner, Sweller and Clark's much-cited 2006 review argues from the same architecture that minimally-guided discovery fails novices — they search where they should be building. Read carefully, this is not an argument against all inquiry; it's an argument about novices, fully consistent with reducing guidance as expertise grows (the reversal effect again).
What the evidence doesn't say
- It doesn't say "make everything easy." Desirable difficulties and CLT divide the territory: CLT eliminates difficulties that are presentational waste; the desirable ones (retrieval, spacing) engage learning processes. A clean instruction card followed by an answer-free retrieval demand is both theories at once.
- It doesn't license dull minimalism. Motivation matters; the theory only insists that decorations not sit between the learner and the structure being learned.
- It doesn't fix the load numbers precisely — treat "3–4 novel elements" as an engineering constant with error bars, not scripture.
In the classroom
- For new material: worked examples first, then fade. "Watch one, complete one, do one" is CLT in a sentence.
- Put words next to their referents — instructions inside the interface being taught, labels on the diagram, the hint beside the field it concerns.
- Kill redundancy: don't read slides aloud verbatim; don't caption what the picture already says (presentation literacy is this effect taught to children).
- Sequence by element interactivity: teach interacting elements (formula + cell references + fill-down) in cumulative steps small enough that each fits the bottleneck with room to spare.
- Automatize the tools: fluent typing and mouse use are schema-side capacity — every skill left unautomatized is a permanent tax on the bottleneck during everything else (why typing fluency matters).
How Wiz Kids applies this
Task instructions live inside the simulated interface, next to what they describe (split-attention); teaching tasks demonstrate and scaffold while review tasks strip guidance (fading + expertise reversal, mechanized); lessons introduce one or two interacting elements at a time along a prerequisite graph that keeps intrinsic load bounded; and the whole typing thread exists because automatized transcription frees the bottleneck for everything downstream.
References
- Miller, G. A. (1956). The magical number seven, plus or minus two. Psychological Review, 63(2).
- Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2).
- Sweller, J., & Cooper, G. A. (1985). The use of worked examples as a substitute for problem solving in learning algebra. Cognition and Instruction, 2(1).
- Sweller, J., van Merriënboer, J. J. G., & Paas, F. (1998; updated 2019). Cognitive architecture and instructional design. Educational Psychology Review, 10(3); 31(2).
- Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1).
- Kirschner, P. A., Sweller, J., & Clark, R. E. (2006). Why minimal guidance during instruction does not work. Educational Psychologist, 41(2).
© Glu IO Pty. Ltd. — Wiz Kids (wiz.kids). Link freely; republication requires permission — see terms. Found an error in our reading of the research? We correct fast: tell any teacher piloting Wiz Kids.