Working memory is not short-term storage. It is a system that does things while you hold things.

The original model came out in 1974, published as "Working Memory" in the Psychology of Learning and Motivation. Before that, short-term memory was treated like a passive holding tank. Baddeley and Hitch showed it actually processes information. That distinction matters if you are building any kind of cognitive assessment tool, working on human-computer interaction, or designing training protocols for people who need to hold information while manipulating it. Most people conflate the two. You will see it come up constantly in papers and product specs. The core architecture has three pieces in the original formulation: the central executive, the phonological loop, and the visuospatial sketchpad. The central executive does not store anything. It allocates attention, switches tasks, suppresses irrelevant information. The phonological loop handles spoken and written language material through subvocal rehearsal. The sketchpad handles visual and spatial information. That is the skeleton. It is also incomplete, and everyone who actually uses this model runs into that incompleteness pretty fast.

Baddeley And Hitch Working Memory Model in practice

I built a cognitive load measurement system for a UX research team a few years back. We were testing whether a particular dashboard interface was overloading users during data entry tasks. The standard approach was to measure response times and error rates while varying working memory load. Here is where it broke down in a way most people do not anticipate. The phonological loop and the visuospatial sketchpad are not independent. They compete for central executive resources when a task requires both simultaneously. Our initial model predicted that dual-modality tasks (verbal instructions plus visual manipulation) would have additive cost. They did not. The cost was multiplicative. Adding a visual reasoning requirement on top of a verbal one blew up reaction times far more than either component alone. The workaround was straightforward once we understood the bottleneck. We stopped treating the two subsystems as separate lanes and started measuring the interference ratio directly. We used a baseline single-modality condition, then a dual-modality condition, and calculated the deviation from additivity. Tasks that showed a deviation above roughly 18 percent were flagged as exceeding working memory capacity for our participant pool. That threshold varied by population, but it gave us a concrete cutoff instead of guessing. There is also a nuance that does not make it into introductory textbooks. The phonological loop is not just about sound. It is about articulatory encoding. When you read silently, you are still running subvocal rehearsal. This means that any task which prevents subvocalization — like articulatory suppression, where someone repeats "the, the, the" aloud — will degrade phonological loop performance dramatically, even if the primary task is visually presented. I learned this the hard way when a colleague ran an experiment without controlling for background speech in the testing room. The results looked like a central executive deficit. It was just phonological interference from people talking in the hallway.

The model got revised in 2000 when Baddeley added the episodic buffer. This was an acknowledgment that the original two-slots-only architecture could not explain how information moved between the subsystems and into long-term memory. The episodic buffer acts as a temporary store that integrates information from all sources. It has a limited capacity of about four chunks, similar to the original estimates for short-term memory, but it is to long-term memory representations, meaning prior knowledge can expand what fits into it through rehearsal and chunking strategies. One thing people miss when they try to apply this model is the asymmetry between the subsystems. The phonological loop is strongly lateralized to the left hemisphere in most right-handed people. The visuospatial sketchpad is more bilaterally distributed. This means that left-hemisphere damage tends to selectively impair the phonological loop while sparing spatial working memory, and vice versa for certain right-hemisphere lesions. If you are working with clinical populations, assuming symmetric degradation across subsystems will give you wrong predictions about task performance. Another practical limitation: the central executive is a theoretical construct, not a localized brain region. You can identify its effects through behavioral paradigms like the dual-task method or the Stroop task, but there is no single neural substrate. This makes it hard to measure directly. Most assessments infer central executive function from task-switching costs or inhibition measures, and those inferences are noisy. The effect sizes in published studies on individual differences in central executive capacity typically range from r = 0.30 to r = 0.50 when predicting real-world performance, which is moderate at best.

Get the Full Details

Working Memory Model (Baddeley and Hitch)
Working Memory Model (Baddeley and Hitch)

If you are designing a study or a product around working memory constraints, the most reliable approach is to measure the specific subsystem your task relies on rather than treating working memory as a unitary capacity. Use the digit span for phonological load, the Corsi block-tapping task for visuospatial load, and dual-task interference for central executive demand. Combining them into a single composite score loses information and tends to underperform compared to targeted measures. The model from 1974 is still the foundation, but treating it as anything other than a starting point will cost you more than it saves.