Turn them on. A 2024 meta-analysis found a medium benefit for captioned viewing over watching without, and the common advice to switch subtitles off and struggle is not supported by the evidence. The more useful finding is that the setting people argue about matters less than the order you use the settings in.
The advice everyone gives is half right
Somebody will always tell you that subtitles are a crutch. The argument is that your eyes take over, your ears go slack, and you finish the episode having read a screenplay rather than heard a language. There is something to it. Anyone who has watched with native-language subtitles knows the sensation of the audio becoming background noise.
Where the advice goes wrong is the conclusion. If reading is doing the work, the fix is not to remove the text and hope. It is to make the text and the sound arrive in an order where each one helps the other, which turns out to be a thing you can actually organise.
What the numbers say about captions
The clearest measurement comes from a 2024 meta-analysis in Language Learning by Kurokawa, Hein and Uchihara. They pooled 89 effect sizes from 49 separate experiments comparing captioned viewing with uncaptioned viewing, and found a medium effect on incidental vocabulary learning, g = 0.56.
Two things about that number are worth holding onto.
The first is that it is real but not enormous. Medium means captions reliably help; it does not mean an episode with captions is a study session. If you watch two hours of television with captions on, you have not done two hours of vocabulary work, and treating viewing as if it were study is how people end up with a large passive vocabulary and nothing they can produce.
The second is what “captions” meant in those studies: on-screen text in the language you are learning, matched to the audio. That is a narrower thing than “subtitles” in ordinary speech, which often means text in your first language. The evidence for captions is not automatically evidence for the subtitle track you are probably using.
It is also worth knowing what kind of learning was being measured. Incidental means the participants were not told to learn the words; they watched, and the test came afterwards. That is the honest comparison for how anyone actually uses television, and it sets a low ceiling by design. Deliberate study beats incidental exposure on any single word, every time. What viewing offers is volume and context, not efficiency, and a medium effect on top of something you were going to do anyway is a good deal on those terms.
Comprehension is not acquisition
Here is the trap that makes subtitle arguments so heated. Captions improve how much of the show you understand, immediately and dramatically. That improvement is obvious from the inside. Vocabulary gains are slower, smaller and invisible while they happen.
So the person defending subtitles is reporting a real experience, and the person attacking them is describing a real limitation, and both of them are talking about different outcomes. Understanding an episode and learning words from it are separate results, and the setting that maximises the first does not automatically maximise the second.
I would go further: the feeling of comprehension while text is on screen tells you almost nothing about your listening. If you want to know where your listening actually is, watch a clip with the text hidden and see what survives. It is usually a deflating experiment, which is why almost nobody runs it.
The finding that should change what you do
The most practically useful result I have seen on this came out in 2025, in the British Journal of Educational Psychology, from Yuan and Tang. They put 162 upper-intermediate Chinese learners of English through four conditions and measured how many target words survived.
| Condition | What the viewer saw | Result |
|---|---|---|
| L1 then bilingual | Native-language subtitles first, then both languages | Best for word form and meaning recall |
| Bilingual then bilingual | Both languages, twice | Behind the sequential condition |
| L2 then L2 | English captions, twice | Behind the sequential condition |
| No subtitles | Nothing on screen | Worst |
The winner was not a setting. It was a sequence: watch it once with subtitles in your own language, then watch it again with both languages on screen. That condition also produced the lowest cognitive load of the four, which is the part that makes sense of the rest.
The reasoning is straightforward once you see it. The first pass spends your attention on working out what is happening. Once you already know what is happening, the second pass has attention free for the English, and a word you meet while you already understand the situation has somewhere to attach. Doing both jobs at once is what makes single-pass viewing so tiring and so leaky.
The obvious practical objection is that most streaming services will not show you two subtitle tracks at once. That is true, and it does not sink the method, because the part doing the work is the order rather than the dual display. Watch it once in your language, then again with English captions. You keep the sequence, you keep the free attention on the second pass, and you lose only the convenience of having the translation sitting there when a line defeats you.
The cost is that you are watching the same thing twice, which sounds like a lot until you compare it against the alternative. One exhausted pass with nothing retained is not cheaper than two comfortable passes with four words kept. Shorter material makes this easy: a twenty-minute episode twice is more use than a film once, and a single scene you actually like is better than either.
Why “just turn them off” is the worst option
In that study the no-subtitle condition finished last on vocabulary. Not a close last.
This is worth saying plainly because the switch-them-off advice has a moral quality to it that makes it hard to argue with. It sounds like discipline. But an unfamiliar word in fast speech, with no text, in a scene you are already struggling to follow, does not become a learnable item just because you suffered through it. You need to notice a word before you can learn it, and noticing is exactly what gets crowded out when comprehension is consuming everything you have.
There is a real case for subtitle-free listening, and it is not vocabulary. It is training your ear to segment speech, which is a genuine skill and does need practice without a safety net. Just do not confuse that exercise with vocabulary work. We wrote about the listening side of this in learning vocabulary by listening.
What moderated the effect, and why it is interesting
The meta-analysis also reported which variables changed the size of the captioning benefit. Three came out as moderators: the instructional level of the learners, who the video was originally made for, and whether the study gave a vocabulary pretest.
That last one is the one I keep thinking about. A pretest is a research instrument, not a teaching method, but sitting one tells you which words the researchers care about, and you then watch differently. Attention was doing part of the work that the captions were getting credit for.
You cannot pretest yourself before every episode. You can do the cheap version, which is to decide before you press play that you are hunting for something specific: phrasal verbs, or words about work, or whatever the episode is likely to be full of. Arriving with a target changes what you notice, and noticing is the whole bottleneck.
The words that stick are the ones that come back
None of this changes the underlying arithmetic of vocabulary. A word you meet once in an episode and never again is gone by the weekend, subtitles or not. Repeated encounters over spaced intervals are what move a word into memory, and a single viewing supplies exactly one encounter. We went through what the research says about how many it takes in how many encounters to learn a word.
This is the honest limit of learning from video. It is very good at giving you first meetings with words in a context vivid enough that you remember where you met them. It is bad at giving you the fourth and seventh meetings. Something else has to do that part, whether that is a notebook, an app, or the accident of watching the same series twice.
That first meeting is worth more than it sounds, though. A word you met in a scene arrives with a face attached, a tone of voice, and a situation, and that is a far better hook than a line in a word list. The people who learn well from television are not the ones who watch the most. They are the ones who take three words out of an episode and put them somewhere that will hand them back next week.
Vocaby’s entries carry IPA and audio next to the definition, which is the specific gap video leaves: you heard the word, you are not sure what it was, and the spelling you guessed does not find it. Reviews are scheduled with FSRS, so a word you half-remember from Tuesday’s episode comes back before it fades rather than after. The word explorer has more than 29,000 entries if you would rather look something up than drill it.
What to actually do
- Watch it once with subtitles in your language. Get the plot. Stop feeling guilty about this part; it is the condition that won.
- Watch it again with both languages on screen if your player supports it, or with English captions if it does not. This is the pass where words are actually available to you.
- Decide what you are hunting for before you start. One category, not everything.
- Write down three or four words. Not thirty. The ones you write down are the only ones that will get a second encounter, so the number should be one you will genuinely review.
- Keep a separate, text-free listening session for ear training, and judge it on comprehension rather than on vocabulary.
The subtitle argument has run for years because both camps are describing something true. Captions help you understand and they can let you coast; the resolution is not picking a side but sequencing the two passes so each one does the job it is good at.
Learn English vocabulary with Vocaby on the App StoreFrequently asked questions
- Do subtitles help you learn vocabulary?
- Yes, modestly. A 2024 meta-analysis in Language Learning pooled 89 effect sizes from 49 experiments and found a medium effect for captioned viewing over uncaptioned viewing on incidental vocabulary learning. Captions are not a shortcut, but the version of the advice that tells you to switch them off is not supported.
- Should I use English subtitles or subtitles in my own language?
- Neither one wins outright. In a 2025 study of 162 upper-intermediate learners, the condition that beat every single-setting option was watching first with native-language subtitles and then again with both languages on screen. The order mattered more than the choice, and it also produced the lowest reported cognitive load.
- Why do I understand everything with subtitles but nothing without them?
- Because reading and listening are separate skills and subtitles let the stronger one carry the weaker one. That is not a failure, but it does mean comprehension while reading along is a poor measure of your listening. Test yourself on a clip with the text hidden if you want to know where your listening actually is.
Keep reading
Can You Learn Vocabulary Just by Listening?
Learners who only listened to a story could write the meaning of 2 percent of its new words right afterward, and none three months later. Here is what the research says about listening versus reading, the one mode that beats both, and a routine that makes audio actually count.
How to Look Up a Word You Only Heard
You caught a word in a podcast, you know roughly how it sounds, and you cannot find it because you cannot spell it. Dictionaries are sorted by the one part of a spoken word you are least likely to have heard correctly. Here is a method that works from the sound instead.
Does Shadowing Actually Help You Remember Words?
Shadowing is usually sold as accent training. The memory research points somewhere more useful, and there is one mistake that quietly cancels the whole benefit.