Chapter 2 · Tokenizer & Data
Chapter 2 covers the full input pipeline: turning raw text into token ids, feeding them to the model as batched, sliding-window (input, target) pairs, and finally into the embeddings the Transformer works with.
Lessons
Section titled “Lessons”@node-llm/core→src/tokenizer/bpe.ts,src/data/dataset.ts,src/data/dataloader.ts- Runnable practice code, one script per lesson:
pnpm ch02:1→ word-level tokenizerpnpm ch02:2→ byte-pair encodingpnpm ch02:3→ data samplingpnpm ch02:4→ Dataset & DataLoaderpnpm ch02:5→ embeddingspnpm ch02:6→ positional encoding