You type "cat" and a chatbot answers. It feels like the computer reads letters the way you do. It does not. Before a language model sees anything, your text is chopped into pieces called tokens, and each piece is swapped for a number. The model only ever works with those numbers.
The simplest plan is one token per character. That works, but it is wasteful: a short sentence becomes dozens of tokens, and the model has to learn from scratch that t, h, e in a row means "the". The opposite plan, one token per whole word, has its own problem: there are endless words (names, typos, "unfriendlier") and any word missing from the list cannot be written at all.
Most modern models pick a middle road: subword tokens. Common chunks like "the" become a single token, and rare words are built from smaller pieces. Nobody writes that list by hand. It is learned from a large pile of text by a method called byte pair encoding (BPE).
Start with every character as its own token. Find the pair of neighbouring tokens that appears most often. Glue that pair into a new token. Repeat. Each repeat is called a merge, and the list of merges is the tokenizer's recipe.
Below is a tiny training text. Press Merge next pair and watch the most common neighbours fuse. The small number under each piece is its ID, the number the model would actually receive. The ▁ mark means "a word starts here", so the tokenizer can tell "the" at the start of a word from "the" in the middle of one. Edit the text to see different merges.
Notice what happens. The first merges grab whatever repeats most, here "a" + "t" because it sits inside cat, sat, mat, rat and ate. After a handful of merges, "the" is one token. Every merge makes the same text shorter in tokens but makes the vocabulary longer, so real tokenizers stop after tens of thousands of merges, a size chosen by the people building the model.
This is the whole algorithm, small enough to read. It ran on the same sentence as the demo:
from collections import Counter
text = "the cat sat on the mat the cat ate the rat"
words = ["▁" + w for w in text.split()]
seqs = [list(w) for w in words] # start: one token per character
def best_pair(seqs):
pairs = Counter()
for s in seqs:
for a, b in zip(s, s[1:]):
pairs[(a, b)] += 1
return max(pairs.items(), key=lambda kv: kv[1])[0], pairs
def merge(seqs, pair):
out = []
for s in seqs:
t, i = [], 0
while i < len(s):
if i + 1 < len(s) and (s[i], s[i + 1]) == pair:
t.append(s[i] + s[i + 1]); i += 2
else:
t.append(s[i]); i += 1
out.append(t)
return out
for step in range(1, 6):
pair, pairs = best_pair(seqs)
seqs = merge(seqs, pair)
n = sum(len(s) for s in seqs)
print(f"merge {step}: {pair[0]!r} + {pair[1]!r} (seen {pairs[pair]}x) -> {n} tokens")
Output:
merge 1: 'a' + 't' (seen 6x) -> 37 tokens
merge 2: '▁' + 't' (seen 4x) -> 33 tokens
merge 3: '▁t' + 'h' (seen 4x) -> 29 tokens
merge 4: '▁th' + 'e' (seen 4x) -> 25 tokens
merge 5: '▁' + 'c' (seen 2x) -> 23 tokens
The text started at 43 tokens (one per character, plus the word-start marks). Five merges brought it down to 23.
The IDs under each chip are simply positions in the vocabulary list. In the demo, the single characters get the first numbers and every new merged piece is added to the end, so the first merge always gets the next free number. A real model's vocabulary works the same way, only much bigger. After tokenizing, the sentence "the cat sat" is nothing but a short list of integers, and that list is the only thing the model receives. Everything it "knows" about language comes from patterns among those numbers, learned during training.
Once a tokenizer is trained, it is used the same way every time: take new text, apply the recipe, get pieces, look each piece up to get its ID. A few things follow from this, and they explain behaviour that surprises beginners:
When you use a language model through an app or an API, the limit on how much text it can handle is counted in tokens, so a long answer or a long pasted file uses up that budget faster than a word count suggests. And when a model fumbles something spelling-related, remember the cause is usually the chopping step, not a lack of effort. The model was handed numbers for chunks, not a row of letters.
One honest simplification: real tokenizers usually work on bytes (so any text, including emoji, can be represented) and use extra rules about where pieces may begin. The merge-the-most-common-pair loop you just ran is the core idea they build on.