Byte-pair encoding (BPE)
BPE is the heart of tokenization. It is the gold-standard algorithm trusted by some of the most powerful language models ever built: GPT-2, GPT-3, and most modern LLMs rely on it to convert raw text into tokens. But what makes it special?
To understand that, we need to look at the core problems any tokenization algorithm must solve.
The first is efficiency: how do we convert text into tokens in a way that is fast and practical at scale? The second, and more interesting, challenge is handling unknown words: what happens when the model encounters a word it has never seen during training?
This is where BPE excels. It was designed with both constraints in mind, and the elegance of its solution is what made it the industry standard.
Byte-pair encoding (BPE) has two phases: first it builds a vocabulary from a training corpus, then it uses that vocabulary to encode new text. Let us walk through each.
Building the vocabulary (training)
Section titled “Building the vocabulary (training)”BPE builds its vocabulary by starting from the smallest possible units, individual characters, and learning to merge them upward.
Here is how it works, step by step:
- Split every word in the training corpus into individual characters.
- Scan the entire corpus and ask one question: which two adjacent units appear together most often?
- Merge that most frequent pair into a single new token and add it to the vocabulary.
- Repeat. Every iteration merges the next most frequent pair.
- Stop once the vocabulary reaches a predefined size.
Encoding new text
Section titled “Encoding new text”Once the vocabulary is built, BPE uses it to encode new text, including words it has never seen before.
- It tries to match the longest possible vocabulary entry against the input word.
- If the full word exists in the vocabulary, it uses it directly.
- If not, it breaks the word into the largest matching subword chunks it can find.
- As a last resort, it falls back to individual characters, which are always in the vocabulary.
We will not build a tokenizer from scratch. BPE is well understood, and production-grade implementations already exist, so instead we use gpt-tokenizer: a popular JavaScript/TypeScript library that implements the exact BPE tokenizers used by GPT-2, GPT-3, and GPT-4. It ships with their real vocabularies and merge rules, runs in pure JS (no Python or native bindings), and gives us encode (text to ids) and decode (ids back to text).
This gives us the same tokenization the real models use, out of the box, so we can focus our effort on the rest of the LLM.
import { encode as gptEncode, decode as gptDecode } from "gpt-tokenizer";
const text = "This is a cold day";
const ids = gptEncode(text); // text -> token idsconsole.log("tokens:", ids);
// Encoding then decoding should return the original text exactly.console.log("round-trip ok:", gptDecode(ids) === text);You can run this in the chapter’s practice code at packages/chapters/src/ch02-tokenizer/02-byte-pair-encoding.ts.
PyTorch → JS mapping
Section titled “PyTorch → JS mapping”| Reference (Python) | Here (TypeScript) |
|---|---|
tiktoken (OpenAI BPE) |
gpt-tokenizer (OpenAI BPE) |