Why the special tokens are not in the WordPiece vocabulary file
[CLS], [SEP], [MASK], [PAD] and [UNK] live outside the main vocabulary, which is why pasting them into a BERT tokenizer gives you ordinary subword fragments instead of the control tokens you typed.
If you paste the literal string [CLS] into a BERT tokenizer you will not get the classification token. You will get the word cl followed by ##s, wrapped in real [CLS] and [SEP] markers added by the post-processor. This trips people up regularly, and the reason is a genuine design decision rather than an oversight.
Both vocabularies pinned here, bert-base-uncased and bert-base-multilingual-cased, behave the same way. Every count below was measured on 2026-10-05.
Special tokens live in a separate list
The five control tokens are declared in the added_tokens array at the top level of tokenizer.json, not inside model.vocab:
"added_tokens": [
{ "id": 0, "content": "[PAD]", "special": true },
{ "id": 100, "content": "[UNK]", "special": true },
{ "id": 101, "content": "[CLS]", "special": true },
{ "id": 102, "content": "[SEP]", "special": true },
{ "id": 103, "content": "[MASK]", "special": true }
]
The special: true flag is what separates them. A special token is handled by the post-processor, never by the normalizer, never by the pre-tokenizer, and never by the WordPiece matcher.
The order the stages run in
- Normalizer — sees your text.
[,C,L,S,]are ordinary characters here. - Pre-tokenizer — splits on whitespace and isolates punctuation, so
[CLS]becomes three words:[,CLS,]. - WordPiece — looks each one up.
[and]are punctuation entries;CLSis not, so it fragments. - Post-processor — appends the real
[CLS]and[SEP].
Step 3 is where the illusion breaks. The text you typed is treated as ordinary prose because, as far as the tokenizer is concerned, it is.
What actually happens
Before the output below: the input was five bracketed control tokens, and every character in them is ordinary input as far as the tokenizer is concerned.
| Typed token | Uncased output | Count | Cased output | Count |
|---|---|---|---|---|
[CLS] | [ cl ##s ] | 4 | [ CL ##S ] | 4 |
[SEP] | [ sep ] | 3 | [ SE ##P ] | 4 |
[MASK] | [ mask ] | 3 | [ MA ##S ##K ] | 5 |
[PAD] | [ pad ] | 3 | [ PA ##D ] | 4 |
[UNK] | [ un ##k ] | 4 | [ UN ##K ] | 4 |
The last column is the point: the only [CLS] and [SEP] in the output are the two the post-processor added. None of the ones you typed survived as itself.
Look at the left column. The input was five bracketed control tokens, and the output contains seventeen content tokens made of punctuation fragments and lowercased letters. Not one [CLS] id from your input survived. The two [CLS] and [SEP] at the ends were added by the post-processor regardless of what you typed.
The right column is worse, because the vocabulary is cased. CLS has no entry, so it becomes CL plus ##S, and MASK becomes three tokens rather than two.
The one case where typing works
Text that contains a bracketed token as a literal — because you are reading a log file, a JSON payload, or a tutorial — must survive the round trip. Both vocabularies fragment it, and neither gives you an escape.
This site does not expose a special-token toggle, so the practical answer for text like this is to accept the fragmentation and read the token list as what it is: the literal characters you typed, broken into the vocabulary's own pieces. The ids you need for masking are fixed and documented, so you can construct them yourself rather than hoping the tokenizer recognises the spelling.
Masking is done by id, not by string
The normal path for masked language modelling does not involve the string [MASK] at all. You build the input, then overwrite one position with id 103. That is why the pipeline works even though the tokenizer cannot recognise the text [MASK]:
The tokenizer does not see a mask at all. It sees the words The capital of France is and a full stop, and it spends nine or eleven content tokens on them. The model later reads position 6 as id 103.
The word mask in the left column is lowercase because bert-base-uncased lowercases input. That is the same lowercase mechanism that made [CLS] become cl ##s, and it is why an unmasked [MASK] string is never the token you wanted.
Padding is length, not content
[PAD] has id 0 and is the lowest id in the vocabulary. It exists to make variable-length sequences rectangular, and it is added after the fact to reach a target length:
Three tokens in, three tokens out. Padding is never produced by encoding text — it is applied to the id sequence afterwards, and only if you ask for it. Neither pinned tokenizer.json enables automatic padding ("padding": null), so a short input stays short.
Why the separation exists
Keeping control tokens out of the matchable vocabulary has three concrete payoffs:
- No collision. A user can never exhaust the
[unused]placeholder slots or shadow[MASK]with a word the tokenizer prefers, because these are not candidates for greedy matching. - Stable ids. The ids 0, 100, 101, 102 and 103 are part of the model interface. They are the same in both vocabularies and across the whole family, so code that hard-codes them works across checkpoints.
- Round-trip safety. Decoding an id sequence can tell control tokens apart from content tokens, which is what lets the decoder and the offset reader reconstruct the original text exactly.
The placeholders follow the same logic from the other direction. Both vocabularies reserve a block of [unused0] … [unused99] entries so there is room to add vocabulary later without renumbering, as the line-number rule requires. They too are declared outside model.vocab, so they behave like ordinary bracketed text:
[unused0] typed as literal text costs four tokens in the uncased vocabulary and five in the cased one, and produces no id in the reserved range. These are names, not tokens, unless something explicitly loads them as added_tokens without special: true.
Sources
- Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (arXiv:1810.04805) — §3, which introduces
[MASK]input corruption and the[CLS]sequence representation. - google-bert/bert-base-uncased,
tokenizer.jsonat commit86b5e09— theadded_tokensarray and vocabulary described above. The link points at the full 40-character commit. - Hugging Face
tokenizers, API reference —AddedTokenand thespecialflag.
Every count on this page was measured against those two pinned tokenizer.json files on 2026-10-05. This site does not host BPE, SentencePiece or tiktoken vocabularies, so no count here applies to a GPT-family or Llama-family model.
Sources
- Hugging Face tokenizers · offsets
Offsets are returned as (start, end) character spans into the original input string.
- google-bert/bert-base-multilingual-cased tokenizer.json
The pinned vocabulary whose offsets are shown below.
Token counts in this article were produced by the tokenizers pinned on the sources page and are checked by the site's test suite. If a count ever stops matching the pinned vocabulary, the build fails.