Language Learning

Does Shadowing Actually Help You Remember Words?

Vocaby Team 10 min read
The word retain on a warm paper card, representing vocabulary held in memory through speech

Shadowing means repeating speech as you hear it, with almost no delay. It is usually sold as accent training. The memory research points somewhere more interesting. Saying a word aloud makes it measurably easier to recall later, and the part of working memory that holds unfamiliar sound turns out to be the same part that builds new vocabulary.

What shadowing actually is

Shadowing is not listen-then-repeat. You do not wait for the sentence to finish. You start speaking roughly a beat behind the speaker and stay there, running your voice alongside theirs while the audio keeps going.

That gap matters more than it sounds. Waiting for a pause lets you rehearse mentally, tidy the sentence up, and produce a cleaned-up version. Staying a beat behind gives you no time to edit. You are forced to process and produce at once, which is uncomfortable and is exactly the point.

The technique did not start as a teaching method. Speech shadowing was a laboratory task in mid-century cognitive psychology, used to study attention. Give someone a stream of speech, make them repeat it continuously, and see what else they manage to notice while doing it. Language learners borrowed it later, largely through Alexander Arguelles, who popularised it as a self-study routine.

Worth keeping in mind that the lab version and the learner version are not the same exercise. Researchers wanted to occupy attention. Learners want to build something. The technique is identical, the goal is not.

Why saying a word out loud makes it stickier

The cleanest finding here has nothing to do with language teaching. It is a plain memory result.

MacLeod and colleagues ran a series of recognition experiments comparing words read aloud against words read silently, and the aloud words were remembered better. They named it the production effect in a 2010 paper in the Journal of Experimental Psychology, and it has held up well since.

The mechanism they proposed is distinctiveness rather than effort. Producing a word lays down extra information about it, a record of having said it, and at test time that extra record is what lets you tell the produced word apart from everything else you saw. The words you spoke are simply easier to pick out of the crowd later.

Two practical consequences follow, and the second one surprises people:

  • The effect is comparative. It shows up strongly when some items are produced and others are not. If you read absolutely everything aloud, you flatten the contrast that makes produced items stand out. Being selective is a feature.
  • Volume is not the ingredient. The advantage extends to mouthing, whispering, writing and singing. Articulation is what counts, so shadowing silently on a crowded train still does most of the work.

One honest caveat before this gets oversold. A 2023 paper in Memory and Cognition found that reading text aloud helps memory but does not help comprehension. That distinction runs through this whole topic: production strengthens your hold on the form of a word. It does nothing to explain what the word means.

The phonological loop, and why sound is the bottleneck

Here is the piece that connects production to vocabulary specifically.

In a 1998 paper in Psychological Review, Baddeley, Gathercole and Papagno reviewed word learning across normal adults, children, and neuropsychological patients, and argued something fairly bold. The phonological loop, the part of working memory that holds sound, evolved primarily to store unfamiliar sound patterns while a permanent memory record is being built.

So on their account the loop is not a general-purpose scratchpad that happens to be useful for language. It is a language-learning device, and holding a new word’s sound shape long enough for it to set is its main job.

That reframes what is hard about learning a word. Meaning is usually the easy part, since a definition or a translation gives it to you in one shot. The difficult part is holding an unfamiliar phonological form steady long enough for it to become permanent. If the sound never stabilises, the word stays a shape you recognise on a page and cannot retrieve by ear.

Shadowing is, mechanically, loading that buffer on purpose and holding it under load. You do not have to accept the full evolutionary claim for the practical read to hold. The sound of a word is not decoration on top of the meaning. It is the part memory struggles with.

What the shadowing research shows, and what it doesn’t

This is where most articles on shadowing get enthusiastic and stop checking. The evidence is real but narrower than the enthusiasm.

Yo Hamada’s work, published in RELC Journal in 2016 and 2019, is the most useful place to look. The consistent finding is that shadowing improves bottom-up processing: phoneme discrimination, word recognition, the automatisation of speech perception. Learners get better at hearing where one word stops and the next begins, which frees up attention for everything else.

The limits are equally clear. In Hamada’s 2016 study the effect on listening was significant but confined to questions involving short conversations. It was not a general listening upgrade, and it certainly was not evidence that shadowing teaches you new words.

So the fair summary is this. Shadowing is well supported as perception training and weakly supported as a vocabulary acquisition method. It will not get words into your head faster than reading or spaced review will.

What it does do is fix a specific and very common failure: you know a word, you would recognise it instantly in writing, and yet it sails past you unrecognised when a native speaker says it at speed. That word is not missing. It is stored in a form your ears cannot reach. Shadowing is what makes it reachable.

The mistake that quietly cancels the benefit

You can only copy a contrast you can hear.

That single sentence accounts for most wasted shadowing. If your first language does not distinguish two sounds, your ear collapses them, and you will shadow along confidently while reproducing your own version. The audio is correct, your output is not, and nothing in the loop tells you so. Twenty repetitions later you have rehearsed the error twenty times.

For Korean speakers the usual suspects are /f/ against /p/, /r/ against /l/, /b/ against /v/, and the tense and lax vowel pairs like the ones in beat and bit. None of these feel ambiguous while you are shadowing. They feel completely fine, which is the problem.

The fix is unglamorous: find out what the target is before you try to copy it. Read the transcription, see that beat is /biːt/ and bit is /bɪt/, and now you know there are two distinct things to aim at rather than one thing you keep hitting. Notation turns an invisible contrast into a visible one.

This is the practical reason every word in Vocaby carries both IPA and audio rather than one or the other. The audio gives you a model to shadow. The IPA tells you what you are supposed to be hearing in it. On their own each is half the tool.

Shadowing compared with the alternatives

These four practices get discussed as if they were interchangeable. They train different things.

PracticeWhat you doWhat it trainsWhere it falls short
Passive listeningPlay audio, follow the meaningComprehension, exposure to natural rhythmNo production, so nothing consolidates in memory
Listen and repeatPause, then reproduce the lineCareful articulation of single wordsThe pause lets you rehearse, so it never reaches real speed
Reading aloudSpeak from a written textMemory for the words on the pageYour own rhythm, not the speaker’s — nothing corrects timing
ShadowingSpeak alongside live audio, a beat behindPerception under time pressure, rhythm, retrieval by earMeaning drops out; comprehension needs a separate pass

The row that matters most is the last column of reading aloud. It gives you the production effect but leaves your prosody untouched, because you are supplying the melody yourself. Shadowing borrows someone else’s timing, and timing is the thing learners almost never fix on their own.

A fifteen-minute loop that survives contact with a real week

Short and repeated beats long and occasional. Sixty to ninety seconds of audio is enough material for the whole session.

  1. Listen once with no speaking. Just find out what is in there.
  2. Read the transcript. Mark every word you could not pick out of the audio — those are the targets, not the words you did not know.
  3. Check the transcription for the marked words. You are looking for contrasts you might be collapsing, not for a full phonetics lesson.
  4. Shadow with the transcript in front of you, four or five passes. Stay a beat behind and do not stop when you stumble.
  5. Shadow without the transcript. This is the pass that shows what actually transferred.
  6. Record yourself once and listen back. Uncomfortable, and the fastest feedback available to you.
  7. Take the words that broke and put them into your review queue.

Step seven is the one people skip, and it is where the session either compounds or evaporates.

Where shadowing fits with spaced repetition

Shadowing and spaced review solve different halves of the same problem, and they are easy to confuse because both involve repeating things.

Shadowing is massed practice. You do it many times inside a few minutes, which is good for encoding a form and unremarkable for long-term retention. Spaced repetition is the opposite: few repetitions, deliberately far apart, timed to catch a word just as it starts to fade. Massed practice builds the trace, spacing keeps it.

The workflow that follows is simple enough. Shadow to encode the sound. Send the words that gave you trouble to a scheduler and let the intervals stretch on their own. Vocaby uses FSRS to time reviews, which means the schedule adapts to how you personally forget instead of assuming a fixed ladder. If you want the underlying logic, we went through how many exposures a word really needs separately.

When it goes wrong

Three failures cover nearly everything.

You cannot keep up. Slow the audio rather than skipping words. Shadowing at 0.75 speed with every syllable intact is real practice; shadowing at full speed while dropping half the sentence is a rhythm you are teaching yourself to keep.

You have no idea what you just said. Expected, and not a failure. Comprehension and shadowing compete for the same attention, which is precisely why the technique was invented as an attention task. Understand the passage first, then shadow it.

You sound robotic. Almost always a sign you are copying words rather than the melody over them. Try humming the intonation of a line with no words at all, then put the words back on top of it. English rhythm carries a surprising amount of intelligibility, and it is the part that transfers to sentences you have never practised.

Shadowing will not fill your vocabulary. It will stop the words you already own from disappearing the moment somebody says them at full speed, which for most intermediate learners is the more painful gap.

Want the IPA and native audio for every word you shadow? Get Vocaby on the App Store.

Frequently asked questions

Does shadowing help you learn new vocabulary?
Not directly. The evidence for shadowing is strongest on the perception side: it sharpens phoneme discrimination and word recognition in running speech, which is what Yo Hamada's research reports. What it does for vocabulary is make words you have already met retrievable by ear, so they stop vanishing the moment a native speaker says them at normal speed. Treat it as a retrieval fix, not an acquisition method.
How long should a shadowing session be?
Fifteen minutes on sixty to ninety seconds of audio beats an hour on a whole podcast episode. Shadowing is dense work, and the returns come from repeating a short passage until the rhythm is automatic rather than from covering more material. Most people can hold real attention for about that long before the mouth keeps moving and the brain checks out.
Do I need to shadow out loud, or is whispering enough?
Whispering and even silent mouthing still work. MacLeod and colleagues found the memory advantage for produced words extends to mouthing, whispering, writing and singing, not just full voice. Volume is not the active ingredient, articulation is. That makes shadowing practical on a train or in an office where speaking aloud is not an option.