skipLink.label

Quest 9 - Self-Attention Implementation

Quest 9: Self-Attention Implementation

medium 25 minutes

🎯 Learning Objectives

  • ✅ How self-attention works in transformers (Q, K, V calculation)
  • ✅ Why scaling by sqrt(d_k) is critical for stable training
  • ✅ How to implement matrix operations for attention
  • ✅ The engineering habit: UNDERSTAND THE MATH

📖 Concept: Self-Attention

Self-attention เป็นหัวใจของ Transformer architecture — มันช่วยให้ model รู้ว่า token ไหนควร “attention” กับ token ไหน

Think of it like this: เมื่อคุณอ่านประโยค “The cat sat on the mat” สมองของคุณจะรู้ว่า “cat” เกี่ยวข้องกับ “sat” มากกว่า “the” — self-attention ทำแบบเดียวกันแต่ในระดับ numbers


⚙️ How It Works

The QKV Formula

Q = tokens x Wq (Query — ถามว่า "ฉันกำลังมองหาอะไร?")
K = tokens x Wk (Key — บอกว่า "ฉันมีอะไรให้?")
V = tokens x Wv (Value — ข้อมูลจริงที่ส่งต่อ)
scores = Q x K^T / sqrt(d_k) (คำนวณความคล้ายคลึง)
weights = softmax(scores) (แปลงเป็น probabilities)
output = weights x V (รวมข้อมูลตาม attention weights)

ตัวอย่าง Matrix Operations

// tokens: [3 tokens, 4 dimensions]
const tokens = [
[1, 0, 1, 0], // token 1
[0, 1, 0, 1], // token 2
[1, 1, 0, 0] // token 3
];
// Wq, Wk, Wv: [4 dimensions, 3 query/key/value dimensions]
const Wq = [[1, 0, 0], [0, 1, 0], [0, 0, 1], [1, 1, 0]];
const Wk = [[0, 1, 0], [1, 0, 0], [0, 0, 1], [0, 1, 1]];
const Wv = [[1, 0, 1], [0, 1, 0], [1, 1, 0], [0, 0, 1]];
// Q = tokens x Wq
// scores = Q x K^T / sqrt(3)
// weights = softmax(scores)
// output = weights x V

💡 Example: Scaling Factor

// ทำไมต้อง sqrt(d_k)?
// ถ้าไม่ scale: scores จะมีค่าใหญ่ -> softmax saturation -> gradients หาย
// ถ้า scale: scores อยู่ใน range ที่เหมาะสม -> softmax ทำงานดี
const d_k = 4; // dimension ของ key
const scalingFactor = Math.sqrt(d_k); // = 2
// scores ที่ไม่ scale: [10, 20, 30] -> softmax ~ [0, 0, 1] (saturation)
// scores ที่ scale แล้ว: [5, 10, 15] -> softmax ~ [0.01, 0.04, 0.95] (ดีกว่า)

⚠️ Common Mistakes

Mistake 1: ลืม scaling factor

คำนวณ attention โดยไม่หารด้วย sqrt(d_k) -> gradients จะ explode ทำให้ training ไม่ stable

Mistake 2: คำนวณ matrix multiplication ผิด

คูณ matrix ผิด dimension -> ต้อง check dimensions ตลอด: (seq_len x d_k) x (d_k x seq_len)

Mistake 3: ไม่ implement softmax

ใช้ raw scores แทน probabilities -> ต้อง softmax เพื่อให้ weights รวมเป็น 1

Mistake 4: ไม่ handle edge cases

ไม่ test กับ sequence length 1, empty tokens -> ต้อง test edge cases


📝 Knowledge Check

📝 Knowledge Check

Q1:Why must attention scores be divided by sqrt(d_k)?

Q2:What does the 'Q' in QKV stand for?

Q3:What does softmax do in the attention mechanism?


🏋️ Quest: Self-Attention Implementation

เขียน function ที่คำนวณ self-attention ด้วย Q, K, V

  1. Download ไฟล์เริ่มต้นของ quest:

    Terminal window
    npx bluebeltdojo download quest-09-self-attention
    cd quest-09-self-attention
  2. เปิด problem.js ใน editor ของคุณพร้อม AI tool

  3. ให้ AI ช่วย implement selfAttention(tokens, Wq, Wk, Wv) ตาม formula

  4. ตรวจสอบ:

    Terminal window
    node test.js
  5. ทดสอบ: scaling factor, matrix dimensions, softmax

  6. ส่งคำตอบ:

    Terminal window
    npx bluebeltdojo submit

💡 Tip: ถ้า AI ไม่ include scaling factor ให้บอกว่า “Don’t forget to divide by sqrt(d_k)!”


คำใบ้

  • อ่าน instructions ใน problem.js อย่างละเอียด — มี formula ระบุไว้
  • ต้อง implement: matrix multiplication, scaling, softmax
  • ทดสอบกับ scaling factor — ถ้าไม่ scale training จะไม่ stable
  • ถ้าติดขัด ลอง search “scaled dot-product attention”