Skip to content
wordpiece.org

What tokenizer.json and tokenizer_config.json each decide

One file holds the vocabulary and the algorithm, the other holds the switches that decide what happens before lookup. Confusing the two produces a tokenizer that loads cleanly and counts wrong.

Published Verified

Both files matter and they do different jobs. tokenizer.json is the tokenizer itself. tokenizer_config.json is the set of switches that decide what gets done to your text before the tokenizer looks at its own vocabulary. Get them mixed up and you get a tokenizer that loads without error and quietly produces different counts than the model was trained with.

This page describes the two vocabularies pinned here, bert-base-uncased and bert-base-multilingual-cased. Every count was measured against those exact files on 2026-10-05.

The division of labour

tokenizer.json is self-contained. It carries everything needed to go from a string to ids:

{
  "version": "1.0",
  "truncation": null,
  "padding": null,
  "added_tokens": [ ... ],
  "normalizer": { ... },
  "pre_tokenizer": { ... },
  "post_processor": { ... },
  "decoder": { ... },
  "model": {
    "unk_token": "[UNK]",
    "continuing_subword_prefix": "##",
    "max_input_chars_per_word": 100,
    "vocab": { ... }
  }
}

The stages run in that order, and each one is a separate reason your output can differ from someone else's:

StageWhat it doesCan change the count?
normalizerdrops invalid chars, collapses whitespace, lowercases, strips accentsyes, often a lot
pre_tokenizersplits on whitespace, isolates punctuation and CJKyes, sets word boundaries
modelgreedy WordPiece over each wordyes
post_processorinserts [CLS] and [SEP]yes, by exactly 2 for a single sequence
truncation / paddingcut or extend to a fixed lengthyes, silently

tokenizer_config.json holds the model-facing settings: vocabulary file references, do_lower_case, strip_accents, tokenize_chinese_chars, model_max_length, special token ids, and the truncation and padding strategy the training pipeline used. A library that loads only the config file will read those flags and, from them, build a normalizer — which is why it can disagree with a runtime that loaded the serialized tokenizer.json.

Why the two can drift

The config carries the declared switches. The serialized tokenizer.json carries the switches as actually serialized. If someone edits the config after the tokenizer was built, the config is no longer a description of the artifact — it is a wish.

That matters most for strip_accents, which is the field most often edited by hand. When it is set to null the behaviour is inherited from do_lower_case: accents stripped when lowercasing is on, preserved when it is off. Setting it explicitly to true or false overrides that, and the two vocabularies on this site are exactly the two cases that matter.

What a filename costs

The difference is not theoretical. Filenames in real repositories are punctuation-heavy strings, and both vocabularies isolate the punctuation into single-character tokens:

Nine content tokens for a 21-character filename. Both models agree exactly here, which is itself the useful finding: tokenizer.json fragments identically under both, because every interesting character in it is punctuation or an underscore.

Reading the config yourself

The fields worth checking on any BERT checkpoint, in the order they affect output:

  • do_lower_case — if this disagrees with the built normalizer, every capital in your input is being miscounted.
  • strip_accents — null means "follow do_lower_case"; anything else is an explicit override.
  • tokenize_chinese_chars — controls whether CJK is split one character per token.
  • model_max_length — this is what a library clamps to if you pass no explicit limit.
  • truncation and padding in tokenizer.json — both null in these two files, meaning no automatic truncation.

Note the last row carefully. null truncation means a long input is not silently cut by the tokenizer. If your pipeline truncated, something else did it, and that something else is the reason a reader cannot reproduce your count.

Sources

Counts measured against those two pinned tokenizer.json files on 2026-10-05. This site does not ship BPE, SentencePiece or tiktoken vocabularies, so nothing here applies to a GPT-family or Llama-family model.

Sources

Token counts in this article were produced by the tokenizers pinned on the sources page and are checked by the site's test suite. If a count ever stops matching the pinned vocabulary, the build fails.

Check it yourself

Open the tokenizer

Every count in this guide comes from a vocabulary pinned by checksum. Paste the same text into the tool and you will get the same output.