The ids of [CLS], [SEP], [MASK], [PAD] and [UNK], and what each one is for
Five control tokens, five fixed ids, five different jobs in a BERT forward pass. What each one does in the sequence, when it is added, and how to tell them apart from the token you meant to type.
The five control tokens in a BERT vocabulary have fixed ids and fixed jobs. They are worth knowing precisely, because most "my count is off by two" reports trace back to not knowing whether [CLS] and [SEP] were already counted.
Both vocabularies on this site, bert-base-uncased and bert-base-multilingual-cased, assign these five tokens the same ids. Every count below was measured on 2026-10-05.
The five, and their ids
| Token | id | Added by | Purpose |
|---|---|---|---|
[PAD] | 0 | you, after encoding | fill a short sequence to a fixed length |
[UNK] | 100 | the tokenizer | a word that matched nothing in the vocabulary |
[CLS] | 101 | the tokenizer, always | sequence-level representation |
[SEP] | 102 | the tokenizer, always | segment boundary; two for a sentence pair |
[MASK] | 103 | you, after encoding | the position the model predicts |
The split by who adds them is the one that matters for counting. [CLS] and [SEP] are automatic, so they are in every count you will ever see from this site. [PAD] and [MASK] are not automatic at all, and neither pinned tokenizer.json enables them.
[CLS] and [SEP]: the two you always pay for
A single sentence always carries exactly two of these, no matter how short it is:
One content token, three total. The tokenizer returned [CLS] short [SEP], so a naive "how many tokens does short use" answer of 1 is wrong for any budget that includes model input length.
This is the arithmetic behind context window accounting: a sequence of n content tokens occupies n + 2 positions in the model.
What each one carries
[CLS] sits at position 0. In BERT pretraining its final hidden state becomes the pooled vector used for sequence classification, so it is the position a classifier reads.
[SEP] terminates a segment. For a sentence pair — the NLI and QA format BERT was pretrained on — the post-processor emits [CLS] A [SEP] B [SEP], so pairs cost three control tokens rather than two. The tokenizer.json encodes this in its pair template, with segment ids 0 0 0 1 1:
"pair": [
{ "SpecialToken": { "id": "[CLS]", "type_id": 0 } },
{ "Sequence": { "id": "A", "type_id": 0 } },
{ "SpecialToken": { "id": "[SEP]", "type_id": 0 } },
{ "Sequence": { "id": "B", "type_id": 1 } },
{ "SpecialToken": { "id": "[SEP]", "type_id": 1 } }
]
Those type_id values are the segment embeddings. They are how the model knows sentence B begins at a specific position rather than being a continuation of A, and why you cannot simply concatenate two texts and reuse the first sentence's count.
[MASK]: a position, not a string
[MASK] is never produced by encoding text. It is written into an id sequence after the fact, which is the only way to be sure the model sees a real mask token rather than the word mask.
On the pinned vocabularies the token exists at id 103 and has an entry in the vocabulary, but nothing in the encode path produces it. Type the string and you get fragments; assign the id and you get a mask. The guide on why these tokens sit outside the matchable vocabulary covers that mechanism in full.
[PAD]: length, not content
[PAD] has id 0, the lowest in the vocabulary, and does exactly one thing: make a sequence rectangular so it can be batched.
Three consequences that show up in practice:
- A padded position carries no meaning, so attention masks must exclude it or the model attends to padding.
- Padding changes
nTokensbut notnContentTokens. If your budget counts content, padding does not consume it. - Both files here have
"padding": null, so encoding never pads by itself. Batch encoding pads only because you asked for a batch.
That last point is a frequent source of irreproducible counts: one person's tokenizer returns 3 tokens for an input, another returns 512, and both are behaving correctly.
[UNK]: the vocabulary giving up
[UNK] is the only control token the tokenizer can emit on its own, and it appears exactly when WordPiece fails to match any substring of a word.
bert-base-uncased reads one of the four characters and gives up on three, spending six total tokens either way. bert-base-multilingual-cased reads all four. Same cost, completely different meaning, and the only signal that distinguishes them is the [UNK] count.
One detail is easy to miss: [UNK] carries no information about which characters were lost. A single [UNK] could stand for one unreadable character or for a whole word longer than max_input_chars_per_word (100 in both files). Reading the offsets is the only way to recover the span it covers.
Seeing all of this in one output
The tokenizer on this site shows the special-token mask alongside the ids, so the control positions are visible directly. On any normal sentence the pattern is unambiguous:
Seven content tokens, nine total, with [CLS] and [SEP] at the ends and nothing else special. The sentence splits on punctuation — the final . becomes its own token — which is the pre-tokenizer's doing its job.
Sources
- Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (arXiv:1810.04805) — §3.1 for WordPiece and
[CLS]/[SEP], §3.4 for[MASK]input corruption. - google-bert/bert-base-multilingual-cased,
tokenizer.jsonat commit3f076fd— theadded_tokensids and thepairpost-processor template quoted above. - google-bert/bert-base-uncased,
tokenizer.jsonat commit86b5e09— the same ids in the 30,522-entry vocabulary.
Every count on this page was measured against those two pinned tokenizer.json files on 2026-10-05. This site does not host BPE, SentencePiece or tiktoken vocabularies, so no count here applies to a GPT-family or Llama-family model.
Sources
- Hugging Face tokenizers · offsets
Offsets are returned as (start, end) character spans into the original input string.
- google-bert/bert-base-multilingual-cased tokenizer.json
The pinned vocabulary whose offsets are shown below.
Token counts in this article were produced by the tokenizers pinned on the sources page and are checked by the site's test suite. If a count ever stops matching the pinned vocabulary, the build fails.