Embeddings
Before we wrap up this chapter, there is one more concept we need to cover, and it is a foundational one: the embedding.
An embedding is a multi-dimensional vector representation of a token. That is the technical definition. But what makes it powerful is the core idea it is built on, an idea that becomes clearer as we go deeper into training: words with similar meaning tend to cluster together in vector space.
Think of it this way. When the model is first initialized, the embedding space is nothing but random numerical values. The model has no understanding of language yet, and every token is just a point floating arbitrarily in a high-dimensional space. But as training begins, something remarkable happens. The embedding space is continuously optimized, and words that share similar concepts or meaning start gravitating toward each other. “King” and “Queen” end up closer to each other than either is to “table”. “Happy” and “joyful” become neighbours. The geometry of the space starts to reflect the structure of language itself.
This is not hand-crafted. No one tells the model that “cat” and “kitten” are related. The model learns it through exposure.

Words with similar meaning cluster together in embedding space: emotions, royalty, and animals each form their own group, while unrelated words like “table” and “chair” sit far apart.
The embedding layer is a lookup table
Section titled “The embedding layer is a lookup table”So how do we go from a token id to its vector? The embedding layer is just a big table with one row per token in the vocabulary. Each row is that token’s vector: a fixed-length list of numbers. Embedding a token means nothing more than looking up its row by id.
Here is a tiny version of that table. Real models give each token a much larger vector, with hundreds or even thousands of dimensions, but for readability we will use just four here:
| Token id | Token | Embedding vector |
|---|---|---|
| 27182 | “Life” | [ 0.11, -0.42, 0.87, -0.05 ] |
| 382 | “ is“ | [ 0.33, 0.02, -0.19, 0.44 ] |
| 1299 | “ like“ | [-0.51, 0.28, 0.66, 0.09 ] |
| 5506 | “ box“ | [ 0.74, -0.13, 0.05, 0.21 ] |
So the full table has 50,257 rows, one per token in the GPT-2 vocabulary. Each row is a vector, and all vectors share the same length. That length is the embedding dimension, embDim.
To embed a sequence you pass the token ids to tf.gather, and it pulls the matching rows and stacks them in order. The ids [27182, 382, 1299, 5506] return those four rows, in sequence.
In code:
import * as tf from "@tensorflow/tfjs-node";
const vocabSize = 50257;const embDim = 768; // GPT-2 small uses 768 dimensions per token
// Starts random, then gets optimized during training.const embedding = tf.randomNormal([vocabSize, embDim]);
// The ids from our DataLoader, shape [batchSize, contextSize].const ids = tf.tensor2d([[27182, 382, 1299, 5506]], undefined, "int32");
// Select each token's row: shape [1, 4, 768].const vectors = tf.gather(embedding, ids);In practice the ids arrive already batched from the DataLoader. We feed that batch tensor straight into tf.gather, which embeds every sequence in the batch at once:
import { DataLoader, LlmDataset } from "@node-llm/core";import { encode as gptEncode } from "gpt-tokenizer";
// Starts random, then gets optimized during training.const embedding = tf.randomNormal([vocabSize, embDim]);
const tokenIds = gptEncode("Life is like box of chocolate you never knew what you gonna get");const dataset = new LlmDataset(tokenIds, 4, 4);
// Tensor batches of shape [batchSize, contextSize].const loader = new DataLoader(dataset, { batchSize: 3, asTensors: true });
for (const batch of loader) { // Embed the whole batch in one call: [3, 4] -> [3, 4, 768]. const embedded = tf.gather(embedding, batch.inputs); console.log(embedded.shape); // [3, 4, 768]
tf.dispose([batch.inputs, batch.targets, embedded]);}Each batch of ids, shape [batchSize, contextSize], becomes a tensor of shape [batchSize, contextSize, embDim]: one vector for every token in every sequence.
The embDim is a design choice. Bigger vectors can capture more nuance but cost more to train. GPT-2 small uses 768; larger models use thousands.
Tokens are not enough: positions matter
Section titled “Tokens are not enough: positions matter”A lookup table hands back the same vector for a token no matter where it sits in the sentence. But “dog bites man” and “man bites dog” mean very different things, so the model also needs to know position.
The fix is a second, smaller table of shape [contextSize, embDim], one vector per slot in the window. We add it to the token embedding, position by position:
input = tokenEmbedding + positionalEmbedding
That sum is the vector that actually flows into the Transformer in the next chapter.
PyTorch → JS mapping
Section titled “PyTorch → JS mapping”| Reference (Python) | Here (TypeScript) |
|---|---|
nn.Embedding |
[vocabSize, embDim] matrix |
embedding(ids) |
tf.gather(embedding, ids) |
positional nn.Embedding |
[contextSize, embDim] matrix |
- An embedding turns a token id into a learned vector of length
embDimthat captures meaning, so similar words end up near each other in vector space. - The embedding layer is a
[vocabSize, embDim]lookup table; embedding a batch of ids is a singletf.gather. - Positional embeddings are added on top so the model knows word order.
- That closes the chapter: we have gone from raw text, to tokens, to batched tensors, to meaningful vectors. Next chapter: attention, where these vectors start to influence one another.