Diagnosing [UNK] tokens
What an [UNK] token actually means, why one unmatchable character costs a whole word, and how to choose a tokenizer that fits your input.
[UNK] is the tokenizer’s way of saying “I have nothing for this.” It is not a best guess and not an approximation — when WordPiece cannot match a word from its vocabulary, it throws the whole word away and emits a single [UNK] in its place. The model then receives the [UNK] id and no information about which characters were actually in the word.
This matters because [UNK] is silent by default. Nothing crashes; the text simply becomes worse input, and quality drops in ways that are hard to trace back. Counting them is the cheapest diagnostic you have.
One bad character costs a whole word
WordPiece does not skip the unmatchable part and tokenize the rest. If no prefix of a word can be matched at some position, the entire word becomes one [UNK]. So a single rare character can take a long, otherwise perfectly tokenizable word down to a single useless token.
There is also a hard length limit. Any word longer than max_input_chars_per_word — 100 characters for both tokenizers here — is emitted as [UNK] without any attempt at splitting. A 150-character run of letters becomes one [UNK], not one long token.
Coverage is per-vocabulary, not per-language
Whether text produces [UNK] depends entirely on which vocabulary you loaded. [UNK] does not mean “this language is unsupported” in the abstract; it means “this specific vocabulary has no piece for this specific word.”
The clearest example on this site is the uncased English vocabulary meeting Chinese text. Its 30,522 entries cover English and very little else, so most Chinese characters fall outside it:
分词测试 -> bert-base-uncased: 分 [UNK] [UNK] [UNK] (1 known, 3 unknown)
bert-base-multilingual-cased: 分 词 测 试 (all four known)
One character (分) happens to be in the English vocabulary; the other three are not. The multilingual model, with 119,547 entries covering 104 languages, resolves all four cleanly. The same input, two tokenizers, opposite outcomes — and the count is 4 either way, which is exactly why you have to look at the tokens, not just the total, to notice.
Text that fragments instead of failing
Not everything exotic becomes [UNK]. Characters that do have a piece in the vocabulary get split into those pieces, even when the split is useless. The ASCII spelling of an emoji codepoint is ordinary text, so it fragments cleanly into sub-pieces:
hello U0001f600 world -> bert-base-uncased: hello u ##00 ##01 ##f ##60 ##0 world
No [UNK] at all — but the emoji has silently become six tokens that carry no emoji meaning. This is a different failure from [UNK], and an equally common one: the count balloons without any token being flagged. This is also exactly what happens to real punctuation-heavy text. Hello, world!!! (really?) costs ten tokens across three words, because each , ! ( ? ) is isolated first.
So a low [UNK] count is not by itself evidence that your input is handled well. It means only that every word had some match. Whether those matches mean anything is a separate question, answered by reading the pieces.
For completeness, the actual emoji character is a different case from its ASCII spelling:
hello 😀 world -> bert-base-uncased: hello [UNK] world
The real emoji has no representation in either vocabulary, so it becomes [UNK] — one token, flagged, and visibly broken. That is the better of the two failures, because at least it tells you the model cannot see it.
How to use this on your own input
- Paste a representative sample of your real text into the tokenizer.
- Read the
[UNK]count in the summary. Any non-zero number is text your model cannot read. - If it is non-zero, switch to the multilingual tokenizer and compare. If the count drops to zero, the text was fine — the vocabulary was the problem.
- If it stays non-zero in both, the text contains characters neither vocabulary covers, and no choice between these two tokenizers will fix it.
Because the two vocabularies are pinned by commit and checksum on this site, the counts you compare are counts for those exact vocabularies — not an approximation of what some other revision would give you.
Sources
- Devlin et al., BERT (arXiv:1810.04805), §3.1
Defines [UNK] as the token used for a word that cannot be tokenized from the vocabulary.
- google-bert/bert-base-uncased tokenizer.json
The 30,522-entry English vocabulary behind the [UNK] examples here.
- google-bert/bert-base-multilingual-cased tokenizer.json
The 119,547-entry multilingual vocabulary compared against it.
Token counts in this article were produced by the tokenizers pinned on the sources page and are checked by the site's test suite. If a count ever stops matching the pinned vocabulary, the build fails.