Vocabulary Tips

What Your Vocabulary Size Test Score Really Means

Vocaby Team 9 min read
A Vocaby word card showing the headword retain with its IPA transcription

A vocabulary size test gives you a real number. It just measures something narrower than most people assume: written recognition of word families, sampled by frequency, then multiplied. Once you know those three design choices, you can tell which conclusions the number supports and which ones it cannot carry.

The number is an estimate built from a small sample

No test asks you about every word. Nation and Beglar’s Vocabulary Size Test, the standard academic instrument, has 140 multiple-choice items covering the 14,000 most frequent word families of English. Those 140 items are spread across 14 frequency bands, ten items per band, and your total is multiplied by 100.

So the test asks you about ten words to decide what you know about a thousand.

That is a reasonable design. Sampling is how every survey works, and frequency bands are a sensible way to stratify the sample, because the words in the fourth thousand really are learned before the words in the ninth. But it means your result is an estimate with error bars, and nothing in the presentation of the number tells you how wide they are.

One correct answer is worth a hundred words

This is the part that changes how you should read your score.

Because each item stands in for a whole band, getting one more question right moves your reported vocabulary size by 100. A four-option multiple-choice question you guess correctly does the same. Two lucky guesses across a 140-item test and your number moves by 200 without anything about your English changing.

The practical consequence: stop reading small differences. If you scored 8,200 in March and 8,400 in June, you learned nothing about your progress. That gap is two items. It is inside the noise of the instrument, and treating it as growth is the same error as weighing yourself on a bathroom scale twice in one morning and concluding you gained half a kilo.

Differences worth taking seriously start somewhere around a thousand, which is ten items, which is a whole band’s worth of sampling.

A word family is not a word

The unit these tests count is the word family, and this is where the number quietly gets generous.

A word family bundles a base word with its inflections and common derivations. Nation, national, nationalism, nationalize, nationality are one family, one countable item. Answer correctly on nation and the test credits you with the family.

You can see the problem. Knowing nation is not knowing nationalize. Learners routinely control the base and stumble on the derived forms, especially the ones that shift part of speech, and a family-counted total papers over that entirely. This is not a fringe complaint; it is the main argument in the research literature for moving vocabulary tests from word families to lemmas as the unit of counting.

So when a test tells you that you know 8,000 words, it is telling you that you recognised the base member of 8,000 families. The number of distinct forms you can actually handle is smaller, and how much smaller depends on how much derivational morphology you have absorbed.

This also explains the thing that confuses people most about these tests: taking two of them and getting two wildly different numbers. If one counts word families and the other counts lemmas, they are not measuring the same objects, and the family-based one will report a larger figure for the same person. Add in the other variables (how many bands each test covers, whether the format is multiple choice or yes/no, whether there is a correction for guessing) and two results a few thousand apart can both be correct readings of the same vocabulary.

Neither test is lying. A vocabulary size number just has no meaning detached from the instrument that produced it, the way a shoe size means nothing without knowing whether it is UK or US. Quote the test along with the number, or do not quote the number.

Why the tests contain words that do not exist

Some vocabulary tests use a yes/no format: here is a word, do you know it, tick or leave blank. It is fast, which is why online tests favour it. It is also trivially easy to inflate, because nothing forces you to prove the claim.

Anderson and Freebody solved this in 1983 by salting the list with pseudowords, invented items built to look and sound like plausible English. If you tick those, you were not reporting knowledge, you were reporting familiarity or optimism, and the scoring formula adjusts your real-word total downward in proportion.

I find this the most honest feature of the whole enterprise. The test designers built in a mechanism whose only purpose is to catch the test-taker being generous with themselves, which tells you how routine that generosity is. If you have taken a yes/no test and wondered why some entries looked oddly unfamiliar for words supposedly in the common range, some of them were not words.

Recognition is the ceiling, not the floor

Every mainstream size test measures reception: you see a written word and decide whether you know it, or pick its meaning from four options. Nothing asks you to produce anything.

That matters because recognition is the easiest form of knowing and it arrives first. You can recognise a word from context, from having read it twice, from its resemblance to something in another language you speak, without being able to summon it when you need it or use it in a sentence that sounds right to a native speaker.

Your size score is therefore an upper bound on your usable vocabulary, not a measure of it. The gap between the two is the whole subject of active versus passive vocabulary, and for most learners it is large — the passive store is commonly several times the active one.

None of that makes the score useless. It makes it a measure of your reading ceiling, which is a genuinely useful thing to know, as long as you do not quote it as though it described your speaking.

What the number is actually good for

Here is the honest accounting of what a size score supports.

What you want to knowDoes the score answer it?What to use instead
Can I read this novel without a dictionary?Yes, reasonably wellNothing better exists
Am I improving over six months?Only for large movesSame test, same conditions, read direction
How does my English compare to my friend’s?NoDifferent tests are not comparable
Can I speak at this level?NoProduction tasks, timed writing
Which words should I study next?Indirectly, via the bandThe band where your score collapses
Am I ready for an exam?NoThe exam’s own practice materials

The first row is where these tests earn their keep. Reading comprehension is strongly tied to how much of a text you recognise, and a size estimate maps cleanly onto that question in a way nothing else does as cheaply. If you want a sense of the targets involved, how many words you need to be fluent covers the thresholds.

The band where you fall off is the real output

Most people read their total and close the tab. The total is the least interesting thing the test produced.

If the test reports per-band results, look at where your correct answers stop being consistent. Someone scoring near-perfect through the fifth band and then dropping sharply in the sixth has learned something specific and actionable: their study should sit in the sixth and seventh thousand, not in whatever list of “advanced words” a search turns up.

This is more useful than the total for a simple reason. The total describes where you have been; the band describes where the next work is. Vocabulary is learned roughly in frequency order because that is the order you meet words in, and the edge of your knowledge is the place where new study has the highest chance of being immediately useful.

A word two bands past your edge is one you will not meet again for months, and you will forget it before you do.

What to do the week after you take one

A test result is only worth the behaviour change it causes. Concretely:

  • Write down the band, not the total. “I fall off in the seventh thousand” is a study plan. “I know 8,400 words” is a fact about nothing.
  • Do not retake it for at least three months. Anything sooner and you are measuring noise, plus you will remember items from last time.
  • Check the gap between reception and production. Take twenty words from your comfortable range and write a sentence with each without looking anything up. The ones that stall are the shape of your real problem.
  • Pick reading slightly above your edge, not far above. Material where you meet an unknown word every page or two teaches vocabulary. Material where you meet four a paragraph teaches frustration.
  • Fix the forgetting before adding more. If words you learned in spring are gone by autumn, a bigger intake will not help. That is a scheduling problem, and spaced repetition is the thing that addresses it.

The last point deserves emphasis because it is the one most people skip. A size test is a snapshot of what survived, and if the survival rate is the broken part of your system, raising the input rate just raises the loss rate with it.

The thing the number cannot see

There is one more limit worth stating plainly, because no test format currently gets around it.

Knowing a word is not a binary. Somewhere between never having seen a word and using it correctly in a job interview, there are several distinct states: recognising it, guessing its meaning from context, knowing its meaning, knowing its register, knowing which words it habitually appears with. A size test collapses all of that into a yes.

That is why two people with the same score can read the same page and come away with different amounts of it. One knows 8,000 words at the recognition level. The other knows 8,000 words well enough to notice that a writer chose terse instead of concise and meant something by it.

You cannot get that from a number, and you should not expect to. A size test is a cheap reading of one dimension of your English, taken once. It is worth about as much attention as that description suggests.

For what it is worth, this is why Vocaby stores each of its 29,000+ entries with an example sentence, IPA and audio rather than a bare gloss, and schedules reviews with FSRS: the goal is depth on the words you have, not a higher number on a test that only counts recognition. You can browse the entries on the Vocaby word pages.

Take the test, note the band, then go back to the words. Get Vocaby on the App Store

Frequently asked questions

Is my vocabulary size test result accurate?
It is accurate for what it measures, which is written recognition of word families sampled by frequency. It is not a count of words you can use in speech or writing, and it carries a wide margin of error because the score is multiplied. On Nation and Beglar's Vocabulary Size Test each correct item is worth 100, so a two-item swing moves your reported size by 200. Treat the number as a band, not a figure.
Why do vocabulary tests include made-up words?
To catch over-claiming. Anderson and Freebody introduced pseudowords in 1983 for exactly this reason: in a yes/no format, nothing stops you from ticking words you have merely seen before. Claiming to know a word that does not exist is evidence that some of your other yes answers are inflated too, so the scoring adjusts your total downward.
Should I retake the test to track progress?
Yes, but only against itself. Retake the same test, in the same conditions, several months apart, and read the direction rather than the difference. Comparing a result from one test to a result from another tells you almost nothing, because the two may sample different frequency bands and count different lexical units.