See exactly what the model sees
Paste text, pick one of the two pinned WordPiece vocabularies, and read the split, the ids and the character offsets. Everything runs in this tab — the text you paste is never uploaded.
- Nothing uploaded
- Checksum-verified vocabularies
- Offsets and masks included
Input
Runs entirely in this tab. Nothing is uploaded, stored or logged.
No text tokenized yet
Paste text and run the tokenizer to see the exact split, ids and character offsets.
Nothing to tokenize
That input has no non-whitespace characters. WordPiece drops whitespace, so the encoding is only the two special tokens and no content token is charged for it.
What this tool will not do
Not a universal token counter
Results are for the tokenizer you select, at the commit shown. Models using BPE, SentencePiece or tiktoken count differently, and this site will not estimate for them.
Not a model playground
There is no inference, no API call and no generated text. You get the tokenizer's output and nothing else.
Not a mirror of your text
Nothing you paste is stored, logged or sent anywhere. Reload the page and it is gone.