Skip to content
wordpiece.org

Methodology

A number is only worth trusting if you know how it was produced. This page describes exactly how the tokenizer here is implemented and what guarantees are checked on every build.

What is actually running

The tokenizer on this site is a from-scratch implementation of the BERT pipeline, written in TypeScript and executed in your browser. It is not a wrapper around a third-party library. The pipeline is the same three stages the reference implementation uses, in order:

  1. Normalize. Drop invalid characters, collapse whitespace runs, then apply lowercasing and accent stripping according to the tokenizer's own configuration.
  2. Pre-tokenize. Split on whitespace, isolate a fixed punctuation set, and isolate CJK code points individually.
  3. WordPiece. For each resulting word, greedily match the longest prefix present in the vocabulary, emitting a single [UNK] if no prefix matches.

Writing it directly is what makes character offsets and the special-token mask available alongside the tokens and ids, which is the information this site exists to show.

How it is verified

Correctness is pinned to a set of reference vectors. For 20 inputs across both supported tokenizers, the expected tokens, ids, offsets and all three masks were generated with the reference Python implementation and stored in the repository. The test suite runs on every build and compares this implementation against those vectors field by field.

The 20 inputs are chosen to hit the places implementations usually disagree: accented text, case folding, CJK, emoji and control characters, over-long words, punctuation runs, currency formatting, URLs, underscore identifiers and empty or whitespace-only input.

What the numbers in the guides are

Counts quoted in the guides and on the homepage are themselves checked. Each one is declared in a test alongside the tokenizer and input that produces it, so if a vocabulary, a threshold or an implementation change ever makes a published number wrong, the build fails rather than shipping a stale claim.

What is deliberately not claimed

  • Results apply only to the tokenizer and commit named on screen, not to BERT models generally.
  • No claim is made about models using BPE, SentencePiece or tiktoken.
  • No inference is performed and no generated text is produced.
  • No claim is made about ranking, accuracy, speed or cost of any model.

Pinned vocabularies

Each tokenizer file is verified in the browser against a recorded SHA-256 before use, and the tool refuses to run on a mismatch.

Reference vectors

Tokens, ids, offsets and masks are compared against reference output on every build across both vocabularies.

Local processing

Tokenization runs entirely in your browser. There is no endpoint that could receive your text.