Guessing from context works, but only when you already know almost every other word on the page. The most cited threshold is 98% coverage, roughly one unknown word in fifty, and a 2023 replication was unable to fully reproduce it. Below that line you are not inferring anything. You are just guessing.
What guessing from context actually requires
Nearly every guide to vocabulary tells you to work out unfamiliar words from their surroundings instead of reaching for a dictionary. It is good advice with a condition attached, and the condition usually gets left off.
To infer a word from its context you need two things. The context has to be doing enough work to narrow down the meaning, and you have to be able to read the context. Both matter, and the second is where most learners are quietly stuck. You cannot use a clue you cannot read.
Which makes “does guessing work?” the wrong question. The one worth asking is how much of a page you have to own already before the technique switches on at all.
The 98% threshold, and why it is contested
The number that gets quoted is 98%. It comes from Hu and Nation’s 2000 study, which asked how much of a text a second-language reader needs to know for adequate unassisted comprehension. They used a 633-word narrative and varied how much of it was familiar, testing readers at 80%, 90%, 95%, and 100% coverage. The conclusion that something near 98% is required has been cited heavily ever since, and it shaped a good deal of how graded readers are built.
It is worth knowing how that figure was produced. The estimate came from a regression on 66 university students in New Zealand, and the 98% point was interpolated rather than directly tested. That is a small study carrying a very large amount of pedagogical weight.
In 2023, Kremmel and colleagues published a replication in Language Learning, this time with 104 adult learners in Sri Lanka in a non-academic setting. The original results did not fully replicate.
None of that makes the threshold wrong. It makes it a rough boundary rather than a switch, one that shifts with the text, the reader, and what you decide “adequate” means. What survives the replication trouble is the shape of the finding, which is the part that actually matters: comprehension does not decline gently as unknown words accumulate. It falls away faster than intuition suggests.
What those coverage numbers feel like on the page
Coverage percentages sound abstract until you convert them into how often you get stopped.
| Coverage | Unknown words | What reading feels like | Can you guess? |
|---|---|---|---|
| 80% | 1 in 5 | Every sentence breaks down. You are decoding, not reading. | No |
| 90% | 1 in 10 | Roughly one unknown per line. Gist only, often wrong. | Rarely |
| 95% | 1 in 20 | Readable but effortful. Unknowns still cluster. | Sometimes |
| 98% | 1 in 50 | Fluent. Unknown words arrive alone and stand out. | Usually |
| 100% | None | Effortless, and no new vocabulary either. | Not needed |
The arithmetic is harsher than it looks. At 95% coverage, which sounds close to complete, a 500-word page hands you 25 unknown words. They do not distribute themselves politely either. They cluster in the harder sentences, which are exactly the sentences you needed to be intact.
Notice the bottom row too. Perfect coverage is comfortable and teaches you nothing. There is a narrow band, somewhere around the high nineties, where reading is both possible and productive.
A guess that works, and why it works
Take this sentence:
“The apology did nothing to assuage her anger, and if anything she left the room angrier than she came in.”
If you do not know assuage, the sentence still hands it to you. It sits after “did nothing to,” so it is a verb. It takes “her anger” as its object. The second clause tells you the anger increased, and it is framed as a contrast with what the apology failed to achieve. So assuage has to mean something like calm or reduce. No dictionary required, and the guess is close enough to be genuinely useful.
Now change one word:
“The apology did nothing to assuage her rancour.”
If rancour is also unfamiliar, the whole thing collapses. Two unknown words in one short sentence and neither can be used to unlock the other. Nothing about the first example was special except that everything around the target word was solid.
That is the mechanism in miniature. Contextual inference is not a skill you apply to a sentence. It is something a sentence does for you, when you have paid for it in advance.
Four ways a context guess goes wrong
- The context is neutral. Most sentences in the wild are not trying to define their words, and the helpfully explanatory ones are common in textbooks but rare in ordinary writing.
- The context actively misleads. A sentence can comfortably support a plausible wrong reading with nothing in it to flag the error.
- You get the gist without the sense. You infer “something negative,” move on with a blurry approximation, and it feels like comprehension right up until the next context breaks it.
- The guess never gets checked. This is the expensive one. An unverified guess hardens, and a wrong meaning you hold confidently takes far more work to dislodge than a word you know you never learned.
The last two are worth dwelling on because they are invisible while they happen. Nothing tells you that your inference was approximate. You simply carry it forward.
Why your first language misleads you here
In your first language, guessing from context works so reliably that it feels like a general reading skill. It isn’t. It rides on tens of thousands of known words, grammar you never have to think about, and years of accumulated intuition about which words keep company with which. Adults read their own language at effectively full coverage almost all the time, which is precisely the condition under which inference is easy.
Carry that habit into a second language at 90% coverage and it stops paying. The technique has not changed. The ground it was standing on has.
This is also why “just read more” fails so many intermediate learners. The advice is sound at the top of the range and close to useless below it, and it is dispensed without checking which side of the line the learner is on.
What to do if you are below the line
Most learners reading real, ungraded English are somewhere between 90% and 95%. That is the awkward middle, and the answer is not to stop reading. It is to stop reading whatever happens to be in front of you.
The useful insight is that coverage is a property of you and the text, not just of you. Your vocabulary grows slowly. Your coverage on a particular page can be raised immediately, and there are three practical ways to do it.
Read narrowly. Stay with one author, one topic, or one series for a long stretch instead of sampling widely. Vocabulary recycles heavily inside a domain, so the specialised words you meet on page one keep reappearing, and your effective coverage on that material climbs far faster than your general vocabulary does. Someone sitting at 92% across English generally can be well into the high nineties inside a familiar subject. Narrow reading is the cheapest way to buy yourself the conditions that make inference work.
Use graded readers without embarrassment. They exist precisely to hold coverage in the productive band. The objection to them is usually that they feel too easy, but easy is the point. A text you can read at 98% is teaching you more per hour than a text you are grinding through at 90%, even though the second one feels like harder work. Difficulty and learning are not the same variable.
Pre-teach the page. Skim a chapter for unknown words, learn them first, then read it properly. This inverts the usual order and it is the fastest single way to turn an unreadable text into a readable one. It also means the words arrive twice, once in isolation and once in context, which is a better sequence than either alone.
All three do the same thing. Rather than waiting for your vocabulary to rise to the text, they bring the text down to where inference already functions.
Building the coverage that makes guessing work
The implication is slightly uncomfortable: guessing from context is a reward for having studied vocabulary, not an alternative to it. Under the threshold, extensive reading is slow and yields little. Above it, reading becomes the most efficient vocabulary engine available.
So the order matters more than the method.
- Front-load frequency. Coverage is bought most cheaply with the most common words, because they account for a wildly disproportionate share of any text.
- Keep studying explicitly while you read. The two are not rivals. Deliberate study raises coverage, and coverage is what makes reading productive.
- Guess, then verify. Attempt the inference, because the attempt does the memory work, then confirm it. The verification step is what stops approximations from setting.
- Reread things. A second pass at higher coverage extracts what the first pass could not reach.
One more thing worth saying plainly. A successful guess is not the same as a learned word. Inferring a meaning once puts a faint trace down and very little else. Words need repeated meetings before they are reliably yours, which is a separate problem from working out what they mean, and we looked at the numbers on that in how many times you need to see a word. If you want the reading side specifically, building vocabulary through reading covers how to run it once your coverage is high enough to support it.
How Vocaby fits
Coverage is the bottleneck, and coverage is exactly what deliberate study buys you.
Vocaby holds more than 29,000 words, which reaches well past the high-frequency band that ordinary reading already supplies, and it schedules reviews with FSRS spaced repetition so that words come back near the point where you would otherwise lose them. Each review is a retrieval attempt rather than a re-reading, which is the same mechanism that makes guess-then-check work. You can browse the range in the word library.
Every entry carries IPA and real audio, and that matters more than it sounds here. A word you guessed correctly on the page but have never heard is still missing half of itself. You can hold the meaning firmly and fail to recognise the word when someone says it, because it was stored as a shape and never as a sound.
Guess from context by all means. Just build the page you are guessing on first.
Download Vocaby on the App Store to learn words with audio, IPA, examples, and spaced repetition on every entry.
Frequently asked questions
- Can you learn English vocabulary just by reading?
- Eventually, yes, but only once your known-word coverage is high enough that unfamiliar words arrive one at a time instead of in clusters. Below that point reading produces very little vocabulary gain per hour, because the sentences that should be supplying the clues are themselves full of gaps. Deliberate study is what gets you to the level where reading starts paying.
- How much of a text do I need to understand to guess new words?
- The most cited figure is 98% of the running words, which is about one unknown word in fifty. That number comes from Hu and Nation's 2000 study and a 2023 replication was unable to fully reproduce it, so treat it as a rough boundary rather than a hard line. The practical version is simpler. If you are meeting more than one or two unknown words per paragraph, you are past the point where guessing works.
- Should I look words up or guess them first?
- Guess first, then check. The attempt is what makes the meaning stick, because a retrieval attempt leaves a stronger trace than simply being told. Skipping the check is the real mistake. An unverified guess tends to harden into a confident wrong meaning, and those are considerably harder to correct later than a word you know you never learned.
Keep reading
How Many Times Do You Need to See a Word to Learn It?
The research lands around eight to ten encounters, but that number hides the important part. What counts as learned changes the answer completely, and recognition turns out to be far cheaper than recall.
The Best Vocabulary Apps in 2026: An Honest Comparison
A fair, balanced look at the best vocabulary apps in 2026, including Vocaby, Anki, Quizlet, Magoosh, WordUp, and Memrise. We compare spaced repetition, pronunciation, free tiers, and who each app actually suits.
Infer, Imply, Suggest: What Exam Questions Ask
Every reading test hands you two vocabularies. One is in the passage and you study it. The other is in the question, it is much smaller, and it decides which of four true-sounding answers is the one they will mark correct.