Quest 81 - Prompt Evaluation Framework
Quest 81: Prompt Evaluation Framework
medium 25-30 minutes🎯 Learning Objectives
- เข้าใจว่า "รู้สึกว่าดี" ไม่ใช่ metric ที่เชื่อถือได้
- สร้าง evaluation framework ที่ใช้ test cases และ scoring
- รู้จัก penalize คำตอบที่มี content ต้องห้าม
- เข้าใจว่า systematic evaluation ช่วย improve prompt ได้อย่างไร
📖 Concept: Prompt Evaluation
Prompt Evaluation คือกระบวนการทดสอบว่า prompt ให้ผลลัพธ์ที่ดีและ consistent หรือไม่ — ไม่ใช่แค่ “ลองใช้ดูแล้วรู้สึกดี”
ลองนึกภาพเหมือน QA testing — คุณไม่บอกว่า ” app ดูสวยดี” แล้ว ship คุณต้องเขียน tests ที่บอกชัดเจนว่า “ถ้า input นี้ ต้องได้ output นี้”
หลักการสำคัญ: “It feels right” is not a metric. You need test cases.
Evaluation Framework:┌────────────────────────────────────┐│ 1. Test Cases (input → expected) ││ → ชุด input ที่คาดหวังผลลัพธ์ชัดเจน │├────────────────────────────────────┤│ 2. Scoring Rules ││ → วิธีคิดคะแนน (match +2, penalty)│├────────────────────────────────────┤│ 3. Forbidden Content ││ → สิ่งที่ห้ามมีใน output │├────────────────────────────────────┤│ 4. Summary ││ → คะแนนรวม + pass/fail │└────────────────────────────────────┘⚙️ How It Works
กระบวนการ Evaluation
Prompt + Test Cases ↓1. Run prompt กับทุก test cases ↓2. เช็คว่า output ตรงกับ expected ไหม (match score) ↓3. เช็คว่า output มี forbidden content ไหม (penalty) ↓4. คำนวณ total score ↓5. Return { scores, total, passed }ตัวอย่าง Test Case
const testCases = [ { input: "What is 2+2?", expected: "4", // output ต้องมีคำนี้ forbidden: ["five", "six"] // output ห้ามมีคำเหล่านี้ }, { input: "Capital of France?", expected: "Paris", forbidden: ["London", "Berlin"] }];💡 Example: Evaluation Framework ใน Action
function createPromptEvaluator(testCases) { return { evaluate(output, testCase) { let score = 0; const issues = [];
// Match check — output ต้องมี expected string if (output.includes(testCase.expected)) { score += 2; } else { issues.push(`Missing expected: "${testCase.expected}"`); }
// Penalty — output ห้ามมี forbidden words if (testCase.forbidden) { for (const word of testCase.forbidden) { if (output.toLowerCase().includes(word.toLowerCase())) { score -= 1; issues.push(`Contains forbidden: "${word}"`); } } }
return { score, issues }; },
runAll(promptFn) { const results = [];
for (const testCase of testCases) { const output = promptFn(testCase.input); const { score, issues } = this.evaluate(output, testCase);
results.push({ input: testCase.input, output, expected: testCase.expected, score, issues, passed: issues.length === 0 }); }
const total = results.reduce((sum, r) => sum + r.score, 0); const passed = results.every(r => r.passed);
return { results, total, passed }; } };}
// ตัวอย่างการใช้งานconst evaluator = createPromptEvaluator([ { input: "2+2?", expected: "4", forbidden: ["five"] }, { input: "Capital of France?", expected: "Paris", forbidden: ["London"] }]);
// promptFn คือ function ที่รับ input แล้ว return outputconst result = evaluator.runAll((input) => { if (input === "2+2?") return "The answer is 4"; if (input === "Capital of France?") return "Paris is the capital"; return "I don't know";});
console.log(result);// { results: [// { input: "2+2?", output: "The answer is 4", score: 2, issues: [], passed: true },// { input: "Capital of France?", output: "Paris is the capital", score: 2, issues: [], passed: true }// ], total: 4, passed: true }⚠️ Common Mistakes
Mistake 1: ไม่มี test cases ที่ชัดเจน
คิดว่า “ลองใช้ดูแล้วน่าจะดี” → ไม่มี objective metric — คุณจะไม่รู้ว่า prompt ดีขึ้นจริงหรือไม่
Mistake 2: ไม่ penalize forbidden content
แค่เช็คว่า output มี expected string ไหม → output อาจมีข้อมูลผิด/อันตราย แต่ยังได้คะแนนดี
Mistake 3: ไม่รวม test case ที่ fail
ตอน evaluate ข้าม test cases ที่ fail → ไม่เห็นภาพรวมว่า prompt ตรงไหนอ่อน
Mistake 4: คิดคะแนนเท่ากันทุก test case
ไม่ weight ตาม difficulty → easy test case ได้คะแนนเท่า hard — ไม่สะท้อนคุณภาพจริง
📝 Knowledge Check
📝 Knowledge Check
Q1:ทำไม "รู้สึกว่า prompt ดี" ถึงไม่ใช่ metric ที่เชื่อถือได้?
Q2:ทำไมต้อง penalize forbidden content?
Q3:runAll() method ควร return อะไร?
🏋️ Quest: Prompt Evaluation Framework
ถึงเวลาฝึกฝน! สร้าง evaluation framework สำหรับ prompt
-
Download ไฟล์เริ่มต้นของ quest:
Terminal window npx bluebeltdojo download quest-81-prompt-evalcd quest-81-prompt-eval -
เปิด
problem.jsใน editor ของคุณพร้อม AI assistant -
Implement
createPromptEvaluator(testCases)ที่ return object ที่มี:evaluate(output, testCase)— คัดคะแนน output เทียบกับ test caserunAll(promptFn)— รันทุก test cases แล้วรวมคะแนน
-
ตรวจสอบ solution ของคุณ:
Terminal window node test.js -
เมื่อ tests ผ่านทั้งหมด ส่งคำตอบ:
Terminal window npx bluebeltdojo submit
💡 Tip: เริ่มจาก evaluate() ก่อน — ทดสอบกับ 1 test case แล้วค่อยเพิ่ม runAll()
คำใบ้
- evaluate() ต้องเช็ค 2 อย่าง: match (output มี expected string) และ forbidden (output ห้ามมี certain words)
- Match = +2 คะแนน, Forbidden match = -1 คะแนน
- runAll() ต้องรันทุก test cases แล้วรวม total score
- ทดสอบด้วย output ที่มีทั้ง expected และ forbidden — ต้องได้ทั้ง +2 และ -1
- Return
{ results, total, passed }— passed = true ถ้าทุก test case ไม่มี issues