Tokens and a simple word-level tokenizer
What is a token?
Section titled “What is a token?”A token is a small chunk of text (a word, part of a word, or a punctuation mark) that a model treats as a single unit. Tokenization is the process of splitting raw text into these tokens and mapping each one to an integer id, because models work with numbers, not characters. Every LLM starts here: text goes in, tokens come out, and those ids become the model’s actual input.

The steps we cover in this chapter are the bottom of the diagram: turning input text into individual words, then into token ids. The embeddings and Transformer come in later chapters.
Why tokenize? Why not just use ASCII or UTF-8?
Section titled “Why tokenize? Why not just use ASCII or UTF-8?”It is a fair question. UTF-8 already turns every character into a number, and the embedding layer just needs numbers, so why not feed UTF-8 codes straight into the model and skip tokenization?
You absolutely could, and nothing breaks. The real question is not whether it works, but which units you hand to the embedding layer, because that choice decides how hard the model’s job is.
An embedding is just a list of numbers that represents a token’s meaning, learned during training. We will explore embeddings in a later chapter.
The difference is what each number represents. A UTF-8 code represents a single character with no regard for meaning. A token id represents a unit chosen to carry meaning. A token like learning already encapsulates a concept, whereas the characters l, e, a, r, n… each mean almost nothing on their own.
This matters because of what an LLM actually does: predict the next token. The unit you predict decides how hard that job is.
- Character level: a tiny vocabulary, but the model must rebuild meaning character by character over very long sequences. It has to learn that
c,a,tcombine into a concept before it can even reason about it. - Word/subword level: each step works with a unit that already carries meaning, so sequences are shorter and every prediction is more informative.
So feeding UTF-8 straight to the embedding layer is not wrong, it just hands the model the least helpful units. Tokenization is the step where we pick better units to embed, so the model starts from meaning instead of rebuilding it from raw characters.