Skip to content
wordpiece.org

Counting against a context window

Measured tokens-per-word ratios for plain prose, URLs, code and CJK, and why budgeting a prompt from word count alone is wrong.

Published 7 min read Verified

You have a context window, a document, and a rough deadline. The question is whether it fits. The instinct is to count words, multiply by a constant, and hope. That instinct works for plain English prose and fails badly for almost everything else.

Every ratio below was measured by tokenizing real inputs with the two vocabularies this site pins, and dividing content tokens by whitespace-delimited words. They are not averages from a blog post — they are what these specific tokenizers actually do.

Plain English is the easy case

InputWordsbert-base-uncasedmultilingual-cased
Tokenization is the first thing a model does.810 tokens (1.25×)11 tokens (1.38×)
Café naïve résumé33 tokens (1.00×)8 tokens (2.67×)
BERT base uncased34 tokens (1.33×)6 tokens (2.00×)

For ordinary sentences in the uncased model, roughly 1.25–1.4 tokens per word is a reasonable first guess. That is the one case where a constant works.

Note how quickly it stops working. Café naïve résumé is 1.00× in the uncased model and 2.67× in the cased one. Same three words, same sentence, more than double the cost — because accented characters are outside the English vocabulary and have to be split. If your text contains any accented Latin, the word-count heuristic undercounts by a factor you cannot guess in advance.

Machine-ish text is where it explodes

InputWordsTokens (uncased)Ratio
https://example.com/a-b_c11313.00×
https://example.com/a-b_c/some/path?q=112121.00×
snake_case_identifier and camelCaseIdentifier3144.67×
hello 😀 world382.67×
Hello, world!!! (really?)3103.33×

A URL is the worst case here, and the reason is mechanical: : / . - _ ? = are each isolated into their own token before splitting begins. https://example.com/a-b_c reads as one word to a human and costs thirteen; extend it with a path and a query string and it costs twenty-one. An identifier with two underscores costs five times what its word count suggests. If your workload is logs, diffs, URLs or code, word count is not a budgeting input at all.

Non-Latin scripts have no useful word count either

Chinese and Japanese text has no spaces, so “words” is not even a meaningful unit to divide by:

InputTokens (uncased)Tokens (multilingual-cased)
分词测试4 (1 known, 3 [UNK])4 (all known)
東京駅3 (2 known, 1 [UNK])3 (all known)
中文English混合55

Each CJK character is isolated and looked up individually, so the cost tracks characters, not words. A character-based estimate is roughly right here and a word-based one is meaningless.

Budgeting in practice

  1. Do not multiply words by a constant and ship it. It is only defensible for plain English in the uncased model, and even there it drifts.
  2. Measure the actual text. Paste a representative sample — not a toy sentence — into the tokenizer. A few hundred real characters is enough to get the ratio that matters for your input.
  3. Measure with the tokenizer you will deploy with. The cased multilingual model costs materially more on accents and mixed case. A count measured on the wrong vocabulary is not an estimate of the right one; it is a different fact.
  4. Remember the fixed overhead. Every encoding carries [CLS] and [SEP], so add 2 to whatever you measure. On a short input that overhead is a large fraction of the total.
  5. Watch the failure that does not announce itself. A text full of [UNK] can still have a plausible token count. Count is not coverage — check the [UNK] count as well.

The point is not that token counting is expensive. It is that it is exact and local, which makes it cheap to just measure rather than estimate. Paste the text, read the number, and the budgeting problem disappears.

Sources

Token counts in this article were produced by the tokenizers pinned on the sources page and are checked by the site's test suite. If a count ever stops matching the pinned vocabulary, the build fails.

Check it yourself

Open the tokenizer

Every count in this guide comes from a vocabulary pinned by checksum. Paste the same text into the tool and you will get the same output.

Related guides