How does a word extract the correct context from its surrounding words?
Current Position: this lesson focuses on the heart of the Transformer Block — the Multi-Head Self-Attention sub-layer (highlighted below). This is where each token ‘looks at’ the others to gather context.
Consider a common English word like bank. See how surrounding words (Context) determine its meaning in the diagram below:
Humans instantly infer word meanings from context, but computers see words only as raw tokens or numbers. Previous models processed words sequentially, causing them to lose context in long sentences. The Transformer's Self-Attention mechanism resolves this by comparing all words simultaneously, allowing each word to derive its true meaning from surrounding Context.
Self-Attention uses three profiles for every word: Query (Q), Key (K), and Value (V). Think of it like looking for a book in a library:
This is what a word is looking for.
Example: You ask the librarian, "I need a book about AI."
This is the label or title of a word.
Example: The title printed on the spine of the book.
This is the actual meaning or content of the word.
Example: The pages inside the book that you actually read.
Suppose we have this sentence:
"The cat drank the milk because it was hungry."
What does it refer to? The cat or the milk? Here's how Self-Attention solves it:
Behind the scenes, the Self-Attention mechanism uses a 4-step math process to mix words together:
The mechanism compares the Query of one word against the Keys of all other words to see how well they match. This creates a "compatibility score."
This comparison is calculated via a Dot Product. The Query (Q) vector of the target word is multiplied with the Key (K) vector of every other word. A higher dot product result signifies greater semantic alignment!
Sometimes the match scores get too large and break the math later on. To fix this, the mechanism scales all the scores down.
To maintain mathematical stability, large dot product values are scaled down by dividing by √dk (the square root of the Key vector dimension). This prevents vanishing gradients during Softmax computation.
The mechanism converts the scaled scores into percentages that add up to 100%. This tells the word exactly how much attention to pay to every other word.
The scaled scores are passed through a Softmax function. Softmax converts raw scores into probability values between 0 and 1 that sum to 100%, producing the final Attention Weights.
Finally, the Self-Attention mechanism gathers the "Values" (meanings) of the other words, taking a lot from words with high percentages and very little from words with low percentages. It mixes them all together into a new, context-rich super-word!
Select a sentence below. Click on any word in the Query (Q) Row to see how the connection lines and Attention Weights dynamically adjust based on context.
Select the correct answer to verify your understanding:
1. What do the Query (Q), Key (K), and Value (V) vectors represent in Self-Attention?
2. What is the primary reason for Scaling after computing the dot product in Self-Attention?
3. In \"The animal didn't cross the street because it was too tired\" and \"...it was too wide\", how does the meaning of 'it' change?
In the next lesson, we will learn how multiple attention streams work in parallel: the Multi-Head Attention Mechanism.