Two vocabularies, two different answers
These two WordPiece tokenizers disagree about almost everything, and both are correct for their own model. The table below is the difference that changes your token count.
Side by side
| Attribute | bert-base-uncased | bert-base-multilingual-cased |
|---|---|---|
| Algorithm | WordPiece (greedy longest match) | WordPiece (greedy longest match) |
| Vocabulary size | 30,522 | 119,547 |
| Lowercasing | Yes — input is lowercased | No — case is preserved |
| Accent handling | Accents stripped (NFD, marks removed) | Accents preserved |
| CJK handling | Characters split per CJK char; most become [UNK] | Characters split per CJK char; 104-script vocabulary |
| Upstream repository | google-bert/bert-base-uncased | google-bert/bert-base-multilingual-cased |
| Pinned commit | 86b5e0934494bd15c9632b12f734a8a67f723594 | 3f076fdb1ab68d5b2880cb87a0886f315b8146f8 |
| License | Apache-2.0 | Apache-2.0 |
bert-base-uncased
English BERT. Lowercases and strips accents, so case and diacritics never affect the count. Chinese characters outside its vocabulary become [UNK], which is why it is a poor choice for CJK text.
- Vocabulary
- 30,522
- Lowercases
- yes
- Strips accents
- yes
- License
- Apache-2.0
bert-base-multilingual-cased
Cased BERT covering 104 languages. Case and diacritics are preserved, so the same sentence can produce different counts from bert-base-uncased.
- Vocabulary
- 119,547
- Lowercases
- no
- Strips accents
- no
- License
- Apache-2.0
The one thing to take away
A token count is not a property of your text. It is a property of your text
and the exact vocabulary it was measured against. Uncased English BERT
will undercount accented text and turn most Chinese characters into
[UNK]; the multilingual model keeps both. Always
measure with the tokenizer your pipeline actually loads.