7 min read
Counting against a context window
Measured tokens-per-word ratios for plain prose, URLs, code and CJK, and why budgeting a prompt from word count alone is wrong.
Read3 guides on fundamentals.
Measured tokens-per-word ratios for plain prose, URLs, code and CJK, and why budgeting a prompt from word count alone is wrong.
ReadWhat a character span in a tokenizer output actually points at, why [CLS] and [SEP] carry [0,0], and how offsets survive accent stripping.
ReadThe greedy longest-match rule behind BERT subword tokens, why continuations carry
Read