Which of the two vocabularies to use
Measured side by side on English prose, mixed case, accents, URLs and CJK, so the choice between bert-base-uncased and bert-base-multilingual-cased is made on counts rather than on which model name sounds more capable.
There are two vocabularies on this site and they are not interchangeable. Picking the wrong one does not fail loudly. It returns a plausible token count, and the cost lands later as a shorter context window, a worse [UNK] rate, or text your model was never really able to read.
Every number below was produced by running the same inputs through both pinned vocabularies. Reproduce any of them in the tokenizer by switching the model and pasting the same string.
The short answer
For plain English text, bert-base-uncased is cheaper and loses nothing you wanted. For anything that carries case, accents or non-Latin scripts as meaningful content, bert-base-multilingual-cased is the only one of the two that can actually read it.
The interesting cases are in between, where the counts diverge for reasons that have nothing to do with language.
English prose: uncased wins on cost
The same sentence, both vocabularies:
The quick brown fox jumps over the lazy dog
bert-base-uncased 11 tokens
[CLS] the quick brown fox jumps over the lazy dog [SEP]
bert-base-multilingual-cased 14 tokens
[CLS] The quick brown f ##ox jump ##s over the la ##zy dog [SEP]
That is a 27% difference on text where the cased model is not recovering anything useful. The becomes The in a vocabulary that stores lowercase forms, so the capital no longer matches any entry and the word is rebuilt from sub-pieces. fox, jumps and lazy break for the same reason.
If your input is ordinary English prose and you are budgeting a context window, folding case is strictly better: same meaning, fewer tokens.
Case you meant to preserve
Now the capitals carry information — an acronym, a product name, a mixed-case identifier:
Deploy to AWS in October
bert-base-uncased 8 tokens
[CLS] deploy to aw ##s in october [SEP]
bert-base-multilingual-cased 9 tokens
[CLS] De ##ploy to A ##WS in October [SEP]
AWS is the case that matters. The uncased model folds it to aws, splits it into aw + ##s, and the model now sees a lowercase fragment instead of the initials. The cased model pays one extra token across the sentence and keeps A ##WS intact.
This is the tradeoff stated plainly: one token per field where case is semantic, saved on every ordinary word where it is not.
Accents: the widest gap on this page
Café résumé
bert-base-uncased 4 tokens
[CLS] cafe resume [SEP]
bert-base-multilingual-cased 7 tokens
[CLS] Café r ##és ##um ##é [SEP]
The uncased model strips accents before it looks anything up, so it spends two tokens and produces cafe / resume. The cased model keeps the diacritics and pays nearly double.
Neither answer is wrong. They answer different questions. If you are indexing French text and need to match café against cafe queries, the two tokens you saved are worth the folding. If you are feeding a model text where Café and Cafe are different restaurant names, the fold has destroyed a distinction you needed.
URLs fragment identically
One case where the choice does not matter:
https://example.com/path?q=1
bert-base-uncased 15 tokens
bert-base-multilingual-cased 15 tokens
Punctuation is isolated first in both models, so the URL breaks into the same single-character pieces either way. Note what this costs regardless: a 25-character URL spends 13 content tokens. A long URL in a prompt is not cheap input, and it is not cheaper on the other model.
CJK: the decision is not close
Chinese is where choosing wrong stops being a cost problem and becomes a correctness one.
分词测试
bert-base-uncased 6 tokens, 3 of them [UNK]
[CLS] 分 [UNK] [UNK] [UNK] [SEP]
bert-base-multilingual-cased 6 tokens, 0 [UNK]
[CLS] 分 词 测 试 [SEP]
東京駅
bert-base-uncased 5 tokens, 1 [UNK]
[CLS] 東 京 [UNK] [SEP]
bert-base-multilingual-cased 5 tokens, 0 [UNK]
[CLS] 東 京 駅 [SEP]
Both models spend the same number of tokens on 分词测试. A budget check that compares totals would call this a tie and would be wrong about half the input: the uncased model reads one character and gets nothing for the other three.
This is the practical argument for checking the [UNK] count rather than the token total, and it is the clearest case where there is no third option — if your text is Chinese, one of these two models can read it and the other cannot.
Deciding
| Your input | Use | Because |
|---|---|---|
| Plain English prose | bert-base-uncased |
27% fewer tokens, identical meaning |
| Acronyms, product names, identifiers | bert-base-multilingual-cased |
keeps AWS intact where folding breaks it |
| Accented text where the accent is data | bert-base-multilingual-cased |
the fold is lossy for you |
| Accented text where the accent is noise | bert-base-uncased |
matching accented to unaccented queries is the point |
| Chinese, Japanese, Korean | bert-base-multilingual-cased |
the other model emits [UNK] for most characters |
| URLs, code, punctuation | either | identical fragmentation, pick on the rest of your text |
When you are unsure, run a representative sample of your real text through both and compare the [UNK] count first and the total second. A model that is cheaper but emitting [UNK] is not cheaper.
Sources
- Devlin et al., BERT (arXiv:1810.04805), §3.1 — defines the greedy longest-match rule both vocabularies here implement.
- google-bert/bert-base-uncased tokenizer.json — the 30,522-entry English vocabulary, pinned on this site to commit
86b5e09. - google-bert/bert-base-multilingual-cased tokenizer.json — the 119,547-entry multilingual vocabulary, pinned on this site to commit
3f076fd.
Every count above was measured against those two pinned tokenizer.json files on 4 October 2026 and is checked by this site's test suite; if a count stops matching the pinned vocabulary, the build fails. The pinned commits and their checksums are listed on the sources page.
Sources
- Hugging Face tokenizers · offsets
Offsets are returned as (start, end) character spans into the original input string.
- google-bert/bert-base-multilingual-cased tokenizer.json
The pinned vocabulary whose offsets are shown below.
Token counts in this article were produced by the tokenizers pinned on the sources page and are checked by the site's test suite. If a count ever stops matching the pinned vocabulary, the build fails.