Self-attention
Before we go into self-attention, let’s examine the shortcomings of the previous methods.
The old approach: encoder-decoder RNNs
Section titled “The old approach: encoder-decoder RNNs”One of the most popular tasks in AI and natural language processing is text generation and translation. The dominant approach for a long time was the RNN, the Recurrent Neural Network, following an encoder-decoder architecture.
Here is how it works: the encoder takes the input text, processes it through multiple layers of a neural network, and that final layer feeds into the decoder. The decoder then takes that final hidden state and decodes it to produce the output text.
Where it breaks down
Section titled “Where it breaks down”By the time we reach the decoder, the model has no access to the full input context. It only sees that final compressed state.
In natural language especially, this is deeply problematic. Even a single word change can have a massive impact on what comes next, and every sentence in a paragraph influences the one that follows. Keeping the model blind to the full weight of the input is fundamentally detrimental to the quality of the output.
Enter self-attention
Section titled “Enter self-attention”The self-attention mechanism works by considering the relevance of every position in the input sequence relative to every other position when computing the output. This gives the model full visibility over the input, not just a final snapshot.