Why the line number in vocab.txt is the token id
A WordPiece vocabulary is an ordered list, so a token's id is its position. That single fact explains why reordering lines silently corrupts a checkpoint and why adding a word is not an editing task.
In BERT's WordPiece vocabulary there is no lookup table and no hash. A token's id is its position in the file. Get that one fact straight and a whole class of confusing bug becomes obvious; get it wrong and you corrupt a model checkpoint without any error message.
This page describes the two vocabularies this site pins: bert-base-uncased (30,522 entries) and bert-base-multilingual-cased (119,547 entries). Every count below was measured against those exact files on 2026-10-05.
The rule, stated once
A WordPiece vocabulary is an ordered list. Index 0 is the first entry, index 1 is the second, and the index is the token id that goes into the embedding lookup.
In the tokenizer.json files these two repos ship, the model section is literal JSON with exactly this shape:
"model": {
"unk_token": "[UNK]",
"continuing_subword_prefix": "##",
"max_input_chars_per_word": 100,
"vocab": { "[PAD]": 0, "[UNK]": 1, ... }
}
The vocab object maps each token string to its integer id. Read a value out of that map and you have the model's embedding row index. There is no additional layer of indirection.
Why that makes reordering destructive
What the failure looks like
The embedding matrix in a BERT checkpoint has one row per vocabulary entry. Row 7682 and row 7683 are learned parameters, and the checkpoint stores them positionally alongside the vocabulary.
Now suppose you "clean up" the file by moving one line. Both halves move together, so nothing complains:
- The vocabulary still has 30,522 entries, so every id is still in range.
- The embedding matrix is still 30,522 rows, so no shape check fails.
- Every id after the moved line now points at the wrong learned vector.
The result is a model that loads, runs, and produces fluent-looking nonsense. This is why vocabulary files are treated as build artifacts with recorded checksums, not as editable configuration.
This site treats them the same way. Each pinned tokenizer.json has a recorded SHA-256 that is verified before the tokenizer is used in the browser; the values are listed on the sources page.
Adding a word is not an insert
The natural reaction to a missing word is to add a line. For WordPiece that is wrong twice over.
First, position matters, as above. Inserting at the front or middle renumbers everything after it and invalidates the embedding matrix.
Second, and more fundamentally, WordPiece segments greedily. Adding tokenization as a single entry does not help unless every longer word that would otherwise match it first is also absent. The subword pieces it would shadow already exist and are already ordered, so the vocabulary is a compressed representation with a specific optimality property, not a lookup list you extend.
Both vocabularies here already contain a full set of [unused0] … [unused99] placeholder entries, reserved precisely so that people do not insert lines.
What a word that is missing actually looks like
The safe behaviour when no entry matches is [UNK], and that is what you see on the two pinned vocabularies for Chinese text. bert-base-uncased has 30,522 entries and cannot represent most Chinese characters:
bert-base-multilingual-cased spends exactly the same number of tokens and reads all four characters. Same length, completely different outcome. A total-token budget would call this a tie.
For ordinary English words neither model needs a new entry, because WordPiece composes them from pieces that already exist:
Three content tokens out of unaffable, assembled from una, ##ffa and ##ble. No vocabulary change is needed, and none should be made.
Where the special tokens sit
The five special tokens sit at fixed, low ids in both vocabularies, declared in the added_tokens array rather than inside model.vocab:
| Token | id | Declared in |
|---|---|---|
[PAD] | 0 | added_tokens, special: true |
[UNK] | 100 | added_tokens, special: true |
[CLS] | 101 | added_tokens, special: true |
[SEP] | 102 | added_tokens, special: true |
[MASK] | 103 | added_tokens, special: true |
These ids are identical in both files. the is id 1996 in the uncased vocabulary and id 10105 in the cased one — same token, same position in the sentence, completely different embedding row. That is the practical reason a count from one model is not a valid count for the other, and the reason which vocabulary you use is a decision with real consequences rather than a preference.
Checking a vocabulary yourself
Three checks catch the reordering case before it reaches a model:
- Count the entries. It must equal the vocabulary size recorded for the model, 30,522 or 119,547.
- Check the special ids.
[PAD]=0,[UNK]=100,[CLS]=101,[SEP]=102,[MASK]=103. - Hash the file. Compare against the checksum published with the checkpoint.
On this site, the first two are what the tokenizer does on every page load, and the third is recorded on the sources page so any drift is visible rather than silent.
Sources
- Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (arXiv:1810.04805) — §3.1, which introduces WordPiece and its greedy longest-match-first segmentation.
- google-bert/bert-base-uncased,
tokenizer.jsonat commit86b5e09— the 30,522-entry English vocabulary described above. The link points at the full 40-character commit. - google-bert/bert-base-multilingual-cased,
tokenizer.jsonat commit3f076fd— the 119,547-entry multilingual vocabulary described above, also linked at the full commit.
Every count on this page was measured against those two pinned tokenizer.json files on 2026-10-05. This site does not host BPE, SentencePiece or tiktoken vocabularies, so no count here applies to a GPT-family or Llama-family model.
Sources
- Hugging Face tokenizers · offsets
Offsets are returned as (start, end) character spans into the original input string.
- google-bert/bert-base-multilingual-cased tokenizer.json
The pinned vocabulary whose offsets are shown below.
Token counts in this article were produced by the tokenizers pinned on the sources page and are checked by the site's test suite. If a count ever stops matching the pinned vocabulary, the build fails.