Skip to content
wordpiece.org

Two vocabularies, two different answers

These two WordPiece tokenizers disagree about almost everything, and both are correct for their own model. The table below is the difference that changes your token count.

Side by side

Attribute bert-base-uncased bert-base-multilingual-cased
Algorithm WordPiece (greedy longest match) WordPiece (greedy longest match)
Vocabulary size 30,522 119,547
Lowercasing Yes — input is lowercased No — case is preserved
Accent handling Accents stripped (NFD, marks removed) Accents preserved
CJK handling Characters split per CJK char; most become [UNK] Characters split per CJK char; 104-script vocabulary
Upstream repository google-bert/bert-base-uncased google-bert/bert-base-multilingual-cased
Pinned commit 86b5e0934494bd15c9632b12f734a8a67f723594 3f076fdb1ab68d5b2880cb87a0886f315b8146f8
License Apache-2.0 Apache-2.0

bert-base-uncased

English BERT. Lowercases and strips accents, so case and diacritics never affect the count. Chinese characters outside its vocabulary become [UNK], which is why it is a poor choice for CJK text.

Vocabulary
30,522
Lowercases
yes
Strips accents
yes
License
Apache-2.0
Diagnose [UNK] coverage

bert-base-multilingual-cased

Cased BERT covering 104 languages. Case and diacritics are preserved, so the same sentence can produce different counts from bert-base-uncased.

Vocabulary
119,547
Lowercases
no
Strips accents
no
License
Apache-2.0
Diagnose [UNK] coverage

The one thing to take away

A token count is not a property of your text. It is a property of your text and the exact vocabulary it was measured against. Uncased English BERT will undercount accented text and turn most Chinese characters into [UNK]; the multilingual model keeps both. Always measure with the tokenizer your pipeline actually loads.

Measuring against a context window