How WordPiece splits text
The greedy longest-match rule behind BERT subword tokens, why continuations carry
BERT does not look up whole words. It looks up the longest prefix of the word that exists in a fixed vocabulary, takes it, and repeats on what is left. Getting that rule exactly right is the whole difference between a token count you can reproduce and one you cannot.
The vocabulary is the unit of account
Each tokenizer ships with a fixed list of strings mapped to integer ids. bert-base-uncased has 30,522 of them; bert-base-multilingual-cased has 119,547. A string is one token if and only if it appears in that list. Nothing outside the list can ever become a token in any other form.
This is why the count is reproducible: given the same vocabulary file, the same text, and the same normalization rules, the output is fully determined. It is also why the count is specific — a number measured with a different vocabulary is a different fact, not an approximation of the same one.
Before any splitting: normalization
Splitting happens on normalized text, not on the raw string:
- Invalid characters are dropped. Null bytes, the replacement character
�, and control characters other than tab, newline and carriage return are removed. - Whitespace runs collapse. Any run of whitespace becomes a single space; leading and trailing whitespace is dropped entirely.
- Accents may be stripped. In the uncased model,
éis decomposed with NFD and the combining marks are dropped, socafénormalizes tocafe. - Case may be folded. Only in the uncased model. Lowercasing keeps just the first character of the expansion, matching the reference implementation exactly.
Every one of those steps can change your count, and normalization is not symmetric between the two models. Café and cafe are the same word, but the cased model preserves both the capital and the accent, and they tokenize differently:
cafe -> bert-base-uncased: cafe | bert-base-multilingual-cased: ca ##fe
Café -> bert-base-uncased: cafe | bert-base-multilingual-cased: Café
In the uncased model both strings collapse to the single token cafe. In the cased model they are two tokens apart from each other — so the same sentence can cost more than you expect once accents and capitals are preserved.
Punctuation is isolated, not glued
The pre-tokenizer splits punctuation into standalone units before WordPiece runs. The set is a fixed literal list — -!#$%&'()*+,./:;<=>?@[]\^_{|}~— and it is **not** the Unicode punctuation category.$and_are currency symbols and connector punctuation, not punctuation marks, yet BERT isolates them anyway. Reproduce the list verbatim orprice is $1,234.56` tokenizes differently than you expect.
CJK characters are isolated one code point at a time by the same stage. Each Chinese character becomes its own word, and then each is looked up on its own.
The greedy loop
For each isolated word, WordPiece starts at position 0 and repeatedly tries the longest substring that exists in the vocabulary:
tokenization, in bert-base-uncased:
start=0 try "tokenization" not in vocab
try "tokenizatio" not in vocab
...
try "token" in vocab -> emit "token"
start=5 try "ization" not in vocab
try "##ization" in vocab -> emit "##ization"
Two properties of this loop are worth stating explicitly.
Continuation pieces are prefixed with ##. Once a word has been split, every piece after the first is looked up with the ## prefix attached, and the emitted token keeps it. So token and ##ization are two different vocabulary lookups. This is what makes the split recoverable from the token sequence alone: a ## prefix is the signal that the piece continues the word before it.
The match is greedy, not optimal. It takes the longest available piece at each step and never backtracks to find a segmentation with fewer pieces. For most vocabulary coverage this coincides with the best possible split, but it is not guaranteed to minimize token count.
The same word, two vocabularies
tokenization is in neither BERT vocabulary, so both models have to split it — and they split it differently. These are the actual outputs of the two vocabularies this site pins, not estimates:
bert-base-uncased -> token ##ization (2 tokens)
bert-base-multilingual-cased -> tok ##eni ##zation (3 tokens)
The uncased model gets lucky: token is in its vocabulary, so greedy matching consumes five characters at once and the remainder ##ization also happens to exist as a continuation. Two pieces cover the whole word.
The cased model cannot use token at all: that entry is simply not in its vocabulary. The greedy loop finds the longest available prefix instead — tok — and what remains does not fit in a single continuation piece, so it becomes ##eni plus ##zation.
That third token is avoidable, which makes it a clean demonstration of the greedy rule above. The cased vocabulary does contain ##ization (id 19980). A split of tok + ##ization would cover the word in two pieces instead of three. Greedy matching cannot find it, because by the time the loop is three characters in, it has already committed to ##eni and never backtracks to test the longer continuation from an earlier position. The vocabulary offers a cheaper segmentation; the algorithm does not take it.
This is why a token count measured on an uncased model is not a token count for a cased one. The extra token is not a rounding difference — it is a direct consequence of case-sensitivity and of which prefix the greedy loop happens to reach first.
Run it yourself: paste tokenization into the tool and switch between the two tokenizers to see the pieces and the counts change together.
The general lesson is the one that matters when you are budgeting a context window: you cannot predict the count from the word count. Splitting happens against one fixed vocabulary, so an unfamiliar word, an unusual capitalisation, a run of punctuation or a non-Latin script can each move the count in a direction you would not predict. Measure with the tokenizer your pipeline actually loads, which is what the tool on this site does.
The [UNK] escape hatch
If no prefix at some position can be matched, the algorithm does not skip the remainder or approximate. It discards the entire word and emits a single [UNK] instead. One unmatchable character costs you the whole word, and you still pay a token for it.
There is one more limit: any word longer than max_input_chars_per_word — 100 characters for these two models — is emitted as [UNK] without any attempt at splitting.
Verify it yourself
Every count in this article is a property of a specific vocabulary file at a specific commit. The tokenizer on this site pins both vocabularies by SHA-256 and refuses to run if the bytes do not match, so what you see there is what the pinned model sees. The sources page lists the exact repos, commits and verification dates.
Sources
- Devlin et al., BERT (arXiv:1810.04805), §3.1
Defines WordPiece as a greedy longest-match-first subword algorithm with continuation notation.
- google-bert/bert-base-uncased tokenizer.json
The exact vocabulary this article describes, pinned to commit 86b5e09 on this site.
Token counts in this article were produced by the tokenizers pinned on the sources page and are checked by the site's test suite. If a count ever stops matching the pinned vocabulary, the build fails.