Skip to content
wordpiece.org

Sources

A tokenizer result only means something alongside the exact file that produced it. Here is every source, commit, checksum and verification date this site depends on.

bert-base-uncased

bert-base-uncased (WordPiece, uncased)

Apache-2.0
Pinned commit
86b5e0934494bd15c9632b12f734a8a67f723594
File
tokenizer.json
SHA-256 of the hosted file
ce64fce797c24f68df90b40a3f74f579b336a493db14bd583fd520ea0d8c9a98
Vocabulary size
30,522
Source last modified
2024-02-19

This site hosts its own copy of tokenizer.json at /tokenizers/bert-base-uncased.json. The browser verifies those bytes against the SHA-256 above before building the tokenizer, and refuses to run if they differ. The tool therefore does not depend on Hugging Face being reachable at the time you use it.

bert-base-multilingual-cased

bert-base-multilingual-cased (WordPiece, cased)

Apache-2.0
Pinned commit
3f076fdb1ab68d5b2880cb87a0886f315b8146f8
File
tokenizer.json
SHA-256 of the hosted file
f4a4d5bf7301717e261fafbe26e1eb967f6ba4cb3ae0ab7a29f4642ec229f386
Vocabulary size
119,547
Source last modified
2024-02-19

This site hosts its own copy of tokenizer.json at /tokenizers/bert-base-multilingual-cased.json. The browser verifies those bytes against the SHA-256 above before building the tokenizer, and refuses to run if they differ. The tool therefore does not depend on Hugging Face being reachable at the time you use it.

How to reproduce any number on this site

  1. 1. Download tokenizer.json at the commit listed above.
  2. 2. Check that its SHA-256 matches the value above.
  3. 3. Load it with the reference implementation, for example with AutoTokenizer.from_pretrained pointed at that repo name.
  4. 4. Compare tokens, ids and offsets with what the tokenizer shows.

The equivalence between this site's implementation and the reference implementation is checked on every build by a suite of reference vectors covering tokens, ids, offsets and all three masks across 20 inputs. See the methodology page.