Counting against a context window
Measured tokens-per-word ratios for plain prose, URLs, code and CJK, and why budgeting a prompt from word count alone is wrong.
ReadEach guide explains one thing about tokenizer output, using counts measured against the exact vocabularies this site pins. Every number can be checked in the tool.
Measured tokens-per-word ratios for plain prose, URLs, code and CJK, and why budgeting a prompt from word count alone is wrong.
ReadWhat a character span in a tokenizer output actually points at, why [CLS] and [SEP] carry [0,0], and how offsets survive accent stripping.
ReadThe greedy longest-match rule behind BERT subword tokens, why continuations carry
ReadWhat an [UNK] token actually means, why one unmatchable character costs a whole word, and how to choose a tokenizer that fits your input.
ReadFive control tokens, five fixed ids, five different jobs in a BERT forward pass. What each one does in the sequence, when it is added, and how to tell them apart from the token you meant to type.
ReadOne file holds the vocabulary and the algorithm, the other holds the switches that decide what happens before lookup. Confusing the two produces a tokenizer that loads cleanly and counts wrong.
ReadA WordPiece vocabulary is an ordered list, so a token's id is its position. That single fact explains why reordering lines silently corrupts a checkpoint and why adding a word is not an editing task.
Read[CLS], [SEP], [MASK], [PAD] and [UNK] live outside the main vocabulary, which is why pasting them into a BERT tokenizer gives you ordinary subword fragments instead of the control tokens you typed.
ReadMeasured side by side on English prose, mixed case, accents, URLs and CJK, so the choice between bert-base-uncased and bert-base-multilingual-cased is made on counts rather than on which model name sounds more capable.
Read