What tokenizer.json and tokenizer_config.json each decide
One file holds the vocabulary and the algorithm, the other holds the switches that decide what happens before lookup. Confusing the two produces a tokenizer that loads cleanly and counts wrong.
Both files matter and they do different jobs. tokenizer.json is the tokenizer itself. tokenizer_config.json is the set of switches that decide what gets done to your text before the tokenizer looks at its own vocabulary. Get them mixed up and you get a tokenizer that loads without error and quietly produces different counts than the model was trained with.
This page describes the two vocabularies pinned here, bert-base-uncased and bert-base-multilingual-cased. Every count was measured against those exact files on 2026-10-05.
The division of labour
tokenizer.json is self-contained. It carries everything needed to go from a string to ids:
{
"version": "1.0",
"truncation": null,
"padding": null,
"added_tokens": [ ... ],
"normalizer": { ... },
"pre_tokenizer": { ... },
"post_processor": { ... },
"decoder": { ... },
"model": {
"unk_token": "[UNK]",
"continuing_subword_prefix": "##",
"max_input_chars_per_word": 100,
"vocab": { ... }
}
}
The stages run in that order, and each one is a separate reason your output can differ from someone else's:
| Stage | What it does | Can change the count? |
|---|---|---|
normalizer | drops invalid chars, collapses whitespace, lowercases, strips accents | yes, often a lot |
pre_tokenizer | splits on whitespace, isolates punctuation and CJK | yes, sets word boundaries |
model | greedy WordPiece over each word | yes |
post_processor | inserts [CLS] and [SEP] | yes, by exactly 2 for a single sequence |
truncation / padding | cut or extend to a fixed length | yes, silently |
tokenizer_config.json holds the model-facing settings: vocabulary file references, do_lower_case, strip_accents, tokenize_chinese_chars, model_max_length, special token ids, and the truncation and padding strategy the training pipeline used. A library that loads only the config file will read those flags and, from them, build a normalizer — which is why it can disagree with a runtime that loaded the serialized tokenizer.json.
Why the two can drift
The config carries the declared switches. The serialized tokenizer.json carries the switches as actually serialized. If someone edits the config after the tokenizer was built, the config is no longer a description of the artifact — it is a wish.
That matters most for strip_accents, which is the field most often edited by hand. When it is set to null the behaviour is inherited from do_lower_case: accents stripped when lowercasing is on, preserved when it is off. Setting it explicitly to true or false overrides that, and the two vocabularies on this site are exactly the two cases that matter.
What a filename costs
The difference is not theoretical. Filenames in real repositories are punctuation-heavy strings, and both vocabularies isolate the punctuation into single-character tokens:
Nine content tokens for a 21-character filename. Both models agree exactly here, which is itself the useful finding: tokenizer.json fragments identically under both, because every interesting character in it is punctuation or an underscore.
Reading the config yourself
The fields worth checking on any BERT checkpoint, in the order they affect output:
do_lower_case— if this disagrees with the built normalizer, every capital in your input is being miscounted.strip_accents—nullmeans "followdo_lower_case"; anything else is an explicit override.tokenize_chinese_chars— controls whether CJK is split one character per token.model_max_length— this is what a library clamps to if you pass no explicit limit.truncationandpaddingintokenizer.json— bothnullin these two files, meaning no automatic truncation.
Note the last row carefully. null truncation means a long input is not silently cut by the tokenizer. If your pipeline truncated, something else did it, and that something else is the reason a reader cannot reproduce your count.
Sources
- Hugging Face
tokenizers, serialization reference — thetokenizer.jsonschema, the normalizer/pre-tokenizer/model/post-processor pipeline, and whattruncationandpaddingmean. - google-bert/bert-base-uncased,
tokenizer_config.jsonat commit86b5e09— the declared switches for the English vocabulary. - google-bert/bert-base-uncased,
tokenizer.jsonat commit86b5e09— the serialized artifact those switches describe.
Counts measured against those two pinned tokenizer.json files on 2026-10-05. This site does not ship BPE, SentencePiece or tiktoken vocabularies, so nothing here applies to a GPT-family or Llama-family model.
Sources
- Hugging Face tokenizers · offsets
Offsets are returned as (start, end) character spans into the original input string.
- google-bert/bert-base-multilingual-cased tokenizer.json
The pinned vocabulary whose offsets are shown below.
Token counts in this article were produced by the tokenizers pinned on the sources page and are checked by the site's test suite. If a count ever stops matching the pinned vocabulary, the build fails.