skipLink.label

Quest 113 - Production Incident Responder

Quest 113: Production Incident Responder

hard 30-45 minutes

🎯 Learning Objectives

  • ✅ How to create structured incident response plans using AI assistance
  • ✅ Why even low-severity incidents need triage — they may be symptoms of bigger issues
  • ✅ How to estimate impact, draft stakeholder communications, and prepare postmortems
  • ✅ The engineering habit of INCIDENT RESPONSE PROTOCOL — structured approach, not panic

📖 Concept: Structured Incident Response

เมื่อ production system มีปัญหา — database down, API errors, slow queries — สิ่งที่หลายคนทำคือ panic debugging: เปิด terminal, พิมพ์ random commands, หวังว่าจะเจอปัญหา

Structured Incident Response คือวิธีที่ดีกว่า — มี protocol ชัดเจน:

1. TRIAGE → จำแนกประเภทและความรุนแรง
↓
2. ACTIONS → ทำตามขั้นตอนที่กำหนด
↓
3. IMPACT → ประเมินผลกระทบ
↓
4. COMMUNICATION → แจ้ง stakeholder
↓
5. POSTMORTEM → วิเคราะห์ root cause ป้องกัน recurrence

AI ช่วยได้อย่างไร? AI สามารถช่วย generate triage classifications, แนะนำ response actions จาก symptoms, draft communication templates, และสร้าง postmortem frameworks — แต่ คุณต้อง verify ทุกอย่าง

ทำไม Severity ทุกระดับต้อง Triage?

⚠️ Low severity: "slight delay"
│
▼ อาจเป็น symptom ของ...
🔴 High severity: database connection pool exhausted

ถ้า skip triage สำหรับ low severity คุณอาจพลาดปัญหาใหญ่ที่ซ่อนอยู่


⚙️ How It Works

Incident Response Flow

respondToIncident({ type, severity, symptoms, logs })

Input:

{
type: 'database-connection',
severity: 'high', // 'low' | 'medium' | 'high' | 'critical'
symptoms: ['connection timeout', '503 errors'],
logs: ['Connection pool exhausted']
}

Output:

{
triage: "...", // Classification of the incident
actions: [...], // Ordered response steps
estimatedImpact: "...", // Impact assessment
communication: "...", // Stakeholder template
postmortem: "..." // Postmortem template
}

Response Structure

FieldPurposeExample
triageจำแนกประเภท incident“Database connection failure — high severity”
actionsลำดับขั้นตอน response["1. Check connection pool", "2. Restart pool", ...]
estimatedImpactประเมินผลกระทบ“High — affects all API endpoints”
communicationแจ้ง stakeholders“Incident declared: Database connection issue…”
postmortemวิเคราะห์หลัง incident“Timeline, root cause, action items…”

Severity Mapping

SeverityImpact LevelResponse TimeActions Count
lowMinimal24 hours2-3
mediumModerate4 hours3-4
highSignificant1 hour4+
criticalSevereImmediate5+

💡 Example: Responding to a Database Incident

Input:

const incident = {
type: 'database-connection',
severity: 'high',
symptoms: ['connection timeout', '503 errors', 'slow queries'],
logs: ['Connection pool exhausted', 'Timeout after 30s'],
};

Output:

const response = respondToIncident(incident);
// {
// triage: "DATABASE-CONNECTION [HIGH] — Connection pool exhaustion causing timeouts",
// actions: [
// "1. Verify connection pool status and current connections",
// "2. Check for connection leaks in recent deployments",
// "3. Scale connection pool if needed",
// "4. Monitor recovery metrics"
// ],
// estimatedImpact: "HIGH — All API requests affected, potential data loss window",
// communication: "Incident [HIGH]: Database connection pool exhausted. Impact: All API endpoints. ETA for resolution: investigating.",
// postmortem: "## Incident Report\n### Timeline\n### Root Cause\n### Action Items\n### Prevention"
// }

Edge case — low severity (must still triage!):

const lowIncident = { type: 'minor', severity: 'low', symptoms: ['slight delay'] };
const lowResponse = respondToIncident(lowIncident);
// triage ต้องมี content — ไม่ใช่ empty string!
// lowResponse.triage.length > 0 ✅

⚠️ Common Mistakes

Mistake 1: Skipping triage for low severity

“Low severity ไม่ต้อง triage ก็ได้” → Low severity อาจเป็น symptom ของปัญหาใหญ่ — triage เสมอ

Mistake 2: Actions not ordered by priority

“ใส่ action ทั้งหมดลงไป ไม่ต้องเรียง” → Actions ต้องเรียงตามลำดับความสำคัญ — 诊断 ก่อน, fix ทีหลัง

Mistake 3: Empty communication template

“Stakeholder ไม่ต้องรู้หรอก” → Communication template ต้องมี — stakeholders ต้องรู้สถานะเพื่อ plan

Mistake 4: Missing postmortem

“แก้เสร็จก็จบ” → Postmortem คือสิ่งที่ป้องกัน incident เดิมเกิดซ้ำ


📝 Knowledge Check

📝 Knowledge Check

Q1:ทำไม severity 'low' incidents ยังต้อง triage?

Q2:incident response output ควรประกอบด้วยอะไรบ้าง?

Q3:Actions array ควรเรียงตามอะไร?


🏋️ Quest: Production Incident Responder

Now it’s time to practice! Build a structured incident response system.

  1. Download ไฟล์เริ่มต้นของ quest:

    Terminal window
    npx bluebeltdojo download quest-113-incident-responder
    cd quest-113-incident-responder
  2. เปิด problem.js ใน editor ของคุณ — สังเกต TODO stub สำหรับ respondToIncident()

  3. Implement respondToIncident(incident):

    • triage: จำแนก incident จาก type + severity
    • actions: ลำดับ response steps (array of strings)
    • estimatedImpact: ประเมินจาก severity level
    • communication: template สำหรับ stakeholders
    • postmortem: template สำหรับ post-incident analysis
  4. สำคัญ: ทดสอบกับ severity: 'low' — triage ต้องมี content เสมอ

  5. ตรวจสอบ solution ของคุณ:

    Terminal window
    node test.js
  6. When all tests pass, submit your solution:

    Terminal window
    npx bluebeltdojo submit

💡 Tip: ถ้า “low severity still gets triage” ล้มเหลว — ตรวจสอบว่าคุณไม่ได้ skip triage เมื่อ severity เป็น ‘low’


คำใบ้

  • อ่าน instructions ใน problem.js อย่างละเอียด
  • Triage ต้อง mention incident type — ไม่ใช่ generic message
  • Actions ควรจำนวน 3+ สำหรับ high severity incidents
  • Edge case สำคัญ: severity ‘low’ ยังต้อง triage — อย่า skip!
  • ถ้าติดขัด ลองอ่าน “Common Mistakes” อีกครั้ง — อย่าดู solution โดยตรง