Skip to content
wordpiece.org

Reading token offsets

What a character span in a tokenizer output actually points at, why [CLS] and [SEP] carry [0,0], and how offsets survive accent stripping.

Published 5 min read Verified

Every token in a Hugging Face encoding carries an offset: a half-open character span [start, end) into the exact string you passed in. Offsets are what let a model point back at the words it processed — they are how attention heads, highlighters and token-to-text mapping work. They are also easy to misread, because the span points into the original text, not into the normalized text the model actually saw.

The two special tokens are always [0,0]

[CLS] and [SEP] are added by the tokenizer, not taken from your input, so they have no characters to point at. Their offsets are always [0, 0]. Every real token has a span that advances through your string in order.

Offsets point into the raw input, not the normalized one

This is the detail that causes surprise. Normalization may change the characters the model sees — it can lowercase, strip accents, delete control characters and collapse whitespace — but an offset always refers to the original input. So in the uncased model, which strips accents, Café still reports the span [0, 4] covering all four original characters, even though the model received the single token cafe.

Here is Café naïve résumé under both tokenizers, with the exact offsets this site’s pinned vocabularies produce:

bert-base-uncased (accents stripped, lowercased)
  token   offsets
  [CLS]   [0, 0]
  cafe    [0, 4]      <- covers the 4 original chars of "Café"
  naive   [5, 10]
  resume  [11, 17]
  [SEP]   [0, 0]

bert-base-multilingual-cased (case and accents kept)
  token   offsets
  [CLS]   [0, 0]
  Café    [0, 4]
  na      [5, 7]
  ##ï     [7, 8]      <- three separate tokens now cover "naïve"
  ##ve    [8, 10]
  r       [11, 12]
  ##és    [12, 14]
  ##um    [14, 16]
  ##é     [16, 17]
  [SEP]   [0, 0]

Two things to read out of that:

  1. Offsets stay contiguous across subword pieces. na, ##ï, ##ve cover [5,10] between them, with no gap, even though they are three separate tokens. Subword splitting never loses the mapping back to the source text.
  2. The span is about the source, not the token. cafe reports four characters because the word Café was four characters, even though the token itself is a shorter normalized string. Offset length reflects the original text; token length does not.

What offsets are good for

  • Mapping a token’s attention or prediction back to the exact substring of a document that caused it.
  • Highlighting, in a UI, which part of the input each token came from.
  • Reconstructing the original text from tokens, for auditing.

What offsets are not good for

  • Counting characters inside a token. A single token can stand for many characters.
  • Recovering the model’s view of the text. For that you want the token string itself, not the offset.
  • Comparing across tokenizers as if the spans were stable. A different tokenizer will segment the same string into a different number of spans over the same offsets.

The tokenizer on this site shows the offsets for every token and lists them in the “Ids and offsets” view, tied to the exact pinned vocabulary that produced them.

Sources

Token counts in this article were produced by the tokenizers pinned on the sources page and are checked by the site's test suite. If a count ever stops matching the pinned vocabulary, the build fails.

Check it yourself

Open the tokenizer

Every count in this guide comes from a vocabulary pinned by checksum. Paste the same text into the tool and you will get the same output.

Related guides