Quest 9 - Self-Attention Implementation
Quest 9: Self-Attention Implementation
medium 25 minutes🎯 Learning Objectives
- How self-attention works in transformers (Q, K, V calculation)
- Why scaling by sqrt(d_k) is critical for stable training
- How to implement matrix operations for attention
- The engineering habit: UNDERSTAND THE MATH
📖 Concept: Self-Attention
Self-attention เป็นหัวใจของ Transformer architecture — มันช่วยให้ model รู้ว่า token ไหนควร “attention” กับ token ไหน
Think of it like this: เมื่อคุณอ่านประโยค “The cat sat on the mat” สมองของคุณจะรู้ว่า “cat” เกี่ยวข้องกับ “sat” มากกว่า “the” — self-attention ทำแบบเดียวกันแต่ในระดับ numbers
⚙️ How It Works
The QKV Formula
Q = tokens x Wq (Query — ถามว่า "ฉันกำลังมองหาอะไร?")K = tokens x Wk (Key — บอกว่า "ฉันมีอะไรให้?")V = tokens x Wv (Value — ข้อมูลจริงที่ส่งต่อ)
scores = Q x K^T / sqrt(d_k) (คำนวณความคล้ายคลึง)weights = softmax(scores) (แปลงเป็น probabilities)output = weights x V (รวมข้อมูลตาม attention weights)ตัวอย่าง Matrix Operations
// tokens: [3 tokens, 4 dimensions]const tokens = [ [1, 0, 1, 0], // token 1 [0, 1, 0, 1], // token 2 [1, 1, 0, 0] // token 3];
// Wq, Wk, Wv: [4 dimensions, 3 query/key/value dimensions]const Wq = [[1, 0, 0], [0, 1, 0], [0, 0, 1], [1, 1, 0]];const Wk = [[0, 1, 0], [1, 0, 0], [0, 0, 1], [0, 1, 1]];const Wv = [[1, 0, 1], [0, 1, 0], [1, 1, 0], [0, 0, 1]];
// Q = tokens x Wq// scores = Q x K^T / sqrt(3)// weights = softmax(scores)// output = weights x V💡 Example: Scaling Factor
// ทำไมต้อง sqrt(d_k)?
// ถ้าไม่ scale: scores จะมีค่าใหญ่ -> softmax saturation -> gradients หาย// ถ้า scale: scores อยู่ใน range ที่เหมาะสม -> softmax ทำงานดี
const d_k = 4; // dimension ของ keyconst scalingFactor = Math.sqrt(d_k); // = 2
// scores ที่ไม่ scale: [10, 20, 30] -> softmax ~ [0, 0, 1] (saturation)// scores ที่ scale แล้ว: [5, 10, 15] -> softmax ~ [0.01, 0.04, 0.95] (ดีกว่า)⚠️ Common Mistakes
Mistake 1: ลืม scaling factor
คำนวณ attention โดยไม่หารด้วย sqrt(d_k) -> gradients จะ explode ทำให้ training ไม่ stable
Mistake 2: คำนวณ matrix multiplication ผิด
คูณ matrix ผิด dimension -> ต้อง check dimensions ตลอด: (seq_len x d_k) x (d_k x seq_len)
Mistake 3: ไม่ implement softmax
ใช้ raw scores แทน probabilities -> ต้อง softmax เพื่อให้ weights รวมเป็น 1
Mistake 4: ไม่ handle edge cases
ไม่ test กับ sequence length 1, empty tokens -> ต้อง test edge cases
📝 Knowledge Check
📝 Knowledge Check
Q1:Why must attention scores be divided by sqrt(d_k)?
Q2:What does the 'Q' in QKV stand for?
Q3:What does softmax do in the attention mechanism?
🏋️ Quest: Self-Attention Implementation
เขียน function ที่คำนวณ self-attention ด้วย Q, K, V
-
Download ไฟล์เริ่มต้นของ quest:
Terminal window npx bluebeltdojo download quest-09-self-attentioncd quest-09-self-attention -
เปิด
problem.jsใน editor ของคุณพร้อม AI tool -
ให้ AI ช่วย implement
selfAttention(tokens, Wq, Wk, Wv)ตาม formula -
ตรวจสอบ:
Terminal window node test.js -
ทดสอบ: scaling factor, matrix dimensions, softmax
-
ส่งคำตอบ:
Terminal window npx bluebeltdojo submit
💡 Tip: ถ้า AI ไม่ include scaling factor ให้บอกว่า “Don’t forget to divide by sqrt(d_k)!”
คำใบ้
- อ่าน instructions ใน
problem.jsอย่างละเอียด — มี formula ระบุไว้ - ต้อง implement: matrix multiplication, scaling, softmax
- ทดสอบกับ scaling factor — ถ้าไม่ scale training จะไม่ stable
- ถ้าติดขัด ลอง search “scaled dot-product attention”