Lesson 3

Self-Attention Mechanism

How does a word extract the correct context from its surrounding words?

Transformer Block Map — Current Position

Focus of this lesson Rest (other lessons)

Current Position: this lesson focuses on the heart of the Transformer Block — the Multi-Head Self-Attention sub-layer (highlighted below). This is where each token ‘looks at’ the others to gather context.

Linear + Softmax → Next Token Transformer Block × N Add & Norm Feed-Forward Network Add & Norm Multi-Head Self-Attention Causal Mask + Positional Encoding Token Embedding Input Tokens

The Context Problem

Consider a common English word like bank. See how surrounding words (Context) determine its meaning in the diagram below:

Sentence 1 Financial Context
"I deposited my money 💰 at the bank."
Context Keyword: money ➔ Resolves bank
Resolved Meaning 🏦 Financial Institution
Sentence 2 Natural Context
"We sat on the muddy bank of the river 🌊."
Context Keyword: river ➔ Resolves bank
Resolved Meaning 🌊 Riverbank

Humans instantly infer word meanings from context, but computers see words only as raw tokens or numbers. Previous models processed words sequentially, causing them to lose context in long sentences. The Transformer's Self-Attention mechanism resolves this by comparing all words simultaneously, allowing each word to derive its true meaning from surrounding Context.

Query, Key, and Value: The Library Analogy

Self-Attention uses three profiles for every word: Query (Q), Key (K), and Value (V). Think of it like looking for a book in a library:

Q
Query (What I want)

This is what a word is looking for.

Example: You ask the librarian, "I need a book about AI."

K
Key (What I have)

This is the label or title of a word.

Example: The title printed on the spine of the book.

V
Value (The actual content)

This is the actual meaning or content of the word.

Example: The pages inside the book that you actually read.

Let's See a Real Example

Suppose we have this sentence:

"The cat drank the milk because it was hungry."

What does it refer to? The cat or the milk? Here's how Self-Attention solves it:

Query: The word "it" asks — "Who is hungry?"
Key: We check the labels of all other words. The word "cat" has a label that matches living things that get hungry.
Value: Because there's a strong match, "it" absorbs the meaning (Value) of "cat". The mathematical relationship it = cat is now established in the model!

How Self-Attention Works

Behind the scenes, the Self-Attention mechanism uses a 4-step math process to mix words together:

Step 1: The Matchmaker (Dot Product)

The mechanism compares the Query of one word against the Keys of all other words to see how well they match. This creates a "compatibility score."

Example: For the word "it", the Query asks, "What am I referring to?". It checks the other words (Keys) in the sentence "The cat drank the milk because it was hungry." Words like "cat" and "milk" will get a high match score.
Q: "it"
×
K: "cat"
→
Score: 50
Q: "it"
×
K: "milk"
→
Score: 40
📐 Mathematical Explanation (Dot Product):

This comparison is calculated via a Dot Product. The Query (Q) vector of the target word is multiplied with the Key (K) vector of every other word. A higher dot product result signifies greater semantic alignment!

Step 2: Shrinking the Scores (Scaling)

Sometimes the match scores get too large and break the math later on. To fix this, the mechanism scales all the scores down.

Example: Suppose "cat" gets a score of 50, and "milk" gets 40. These numbers are too large and might break the next math step. The model scales them down (e.g., to 5.0 and 4.0) to keep things stable.
50
40
→
÷ √dk
5.0
4.0
📐 Mathematical Explanation (Scaling Trick):

To maintain mathematical stability, large dot product values are scaled down by dividing by √dk (the square root of the Key vector dimension). This prevents vanishing gradients during Softmax computation.

Step 3: Turning Scores into Percentages (Softmax)

The mechanism converts the scaled scores into percentages that add up to 100%. This tells the word exactly how much attention to pay to every other word.

Example: Now the mechanism turns the scaled scores into percentages. The word "it" might pay 80% attention to "cat", 15% to "milk", and 5% to other words. The total is always exactly 100%.
5.0
4.0
Softmax
80%
cat
15%
milk
5%
...
📐 Mathematical Explanation (Softmax Output):

The scaled scores are passed through a Softmax function. Softmax converts raw scores into probability values between 0 and 1 that sum to 100%, producing the final Attention Weights.

Step 4: Combining Meanings (Weighted Sum)

Finally, the Self-Attention mechanism gathers the "Values" (meanings) of the other words, taking a lot from words with high percentages and very little from words with low percentages. It mixes them all together into a new, context-rich super-word!

Example: To build the final meaning for "it", the mechanism takes a large piece (80%) of the "cat" Value, and a small piece (15%) of the "milk" Value. Because the sentence mentioned "hungry", the mathematical calculation resolves "it" to the cat! The word "it" now contains the context of a hungry cat.
80%
×
V: "cat"
+
15%
×
V: "milk"
+ ...
↓
New Context: "it" (Hungry Cat)

Interactive Attention Simulator

Select a sentence below. Click on any word in the Query (Q) Row to see how the connection lines and Attention Weights dynamically adjust based on context.

Query Row (Q) - Click a word
Key/Value Row (K/V)

Quick Assessment (Quiz)

Select the correct answer to verify your understanding:

1. What do the Query (Q), Key (K), and Value (V) vectors represent in Self-Attention?

A) Query is the search for word relationships, Key is the label/index of a word, and Value is its actual semantic content.
B) Query is the next token ID, Key is the embedding matrix, and Value is the Softmax value.
C) Query is the dictionary database, Key is the grammar rules, and Value is the final output word.

2. What is the primary reason for Scaling after computing the dot product in Self-Attention?

A) To shrink vector sizes in memory and accelerate GPU speeds.
B) Large dot products can saturate the Softmax function, leading to vanishing gradients and hindered learning.
C) It is impossible to count the total number of tokens in the sentence without it.

3. In \"The animal didn't cross the street because it was too tired\" and \"...it was too wide\", how does the meaning of 'it' change?

A) The two sentences use different rule-based grammar code.
B) Through Self-Attention, 'it' connects with the descriptive adjectives ('tired' or 'wide') to shift attention to 'animal' or 'street' respectively.
C) The computer sees a spelling difference in the embedding of 'it' in both sentences.

Helpful and Authoritative References (Citations):

In the next lesson, we will learn how multiple attention streams work in parallel: the Multi-Head Attention Mechanism.