Skip to content
wordpiece.org

About this site

A tokenizer inspector for BERT WordPiece models, built for developers who need the exact split rather than an estimate.

What it does

You paste text. The site runs a WordPiece tokenizer in your browser and shows the subword pieces, the integer ids, the character offset of every token, and the special-token mask. Results are always tied to a specific tokenizer at a specific commit, with a checksum the browser verifies before the tokenizer is allowed to run.

Who it is for

Developers and ML engineers working with BERT-family models who need to answer a concrete question: given the tokenizer my pipeline actually loads, what does the model see? If you are already comfortable with what a token is, this site will not explain attention or transformers to you.

What it does not do

  • It does not run models or generate text.
  • It does not count tokens for BPE, SentencePiece or tiktoken vocabularies.
  • It does not collect accounts, email addresses or contact details. There is no form on this site.
  • It does not host user text. There is nowhere for it to go.

Independence

This is an independent project. It is not affiliated with, endorsed by or sponsored by Google or Hugging Face. The tokenizer vocabularies come from Google's public repositories on the Hugging Face Hub and are used under the Apache License 2.0, with full attribution on the sources page. Model names referenced here, including BERT, are trademarks of their respective owners, used descriptively to identify the tokenizers being inspected.

Corrections

Every technical claim on this site is checked against the pinned vocabularies by an automated suite, and each guide records its sources and the date its content was last verified. If something here is wrong, the methodology page explains exactly how it is checked, which makes it straightforward to confirm or contradict.

2
pinned WordPiece tokenizers
150,069
vocabulary entries vendored
0
data collected about you