Skip to content

Self-attention

Before we go into self-attention, let’s examine the shortcomings of the previous methods.

One of the most popular tasks in AI and natural language processing is text generation and translation. The dominant approach for a long time was the RNN, the Recurrent Neural Network, following an encoder-decoder architecture.

Here is how it works: the encoder takes the input text, processes it through multiple layers of a neural network, and that final layer feeds into the decoder. The decoder then takes that final hidden state and decodes it to produce the output text.

By the time we reach the decoder, the model has no access to the full input context. It only sees that final compressed state.

In natural language especially, this is deeply problematic. Even a single word change can have a massive impact on what comes next, and every sentence in a paragraph influences the one that follows. Keeping the model blind to the full weight of the input is fundamentally detrimental to the quality of the output.

The self-attention mechanism works by considering the relevance of every position in the input sequence relative to every other position when computing the output. This gives the model full visibility over the input, not just a final snapshot.