Skip to content
wordpiece.org

Guides

Each guide explains one thing about tokenizer output, using counts measured against the exact vocabularies this site pins. Every number can be checked in the tool.

Why the line number in vocab.txt is the token id

A WordPiece vocabulary is an ordered list, so a token's id is its position. That single fact explains why reordering lines silently corrupts a checkpoint and why adding a word is not an editing task.

Read

Which of the two vocabularies to use

Measured side by side on English prose, mixed case, accents, URLs and CJK, so the choice between bert-base-uncased and bert-base-multilingual-cased is made on counts rather than on which model name sounds more capable.

Read