skipLink.label

Quest 16 - Fine-Tuning Pipeline Design

Quest 16: Fine-Tuning Pipeline Design

medium 25 minutes

🎯 Learning Objectives

  • ✅ Understand the three stages of LLM fine-tuning: SFT → RLHF → DPO
  • ✅ Design data requirements and evaluation metrics for each stage
  • ✅ Recognize when to use RLHF vs. DPO for alignment
  • ✅ Create a complete fine-tuning pipeline design document

📖 Concept: Fine-Tuning Pipeline

การ train LLM ไม่ใช่แค่ feed ข้อมูลเข้าไปแล้วหวังผล — มันเป็น pipeline ที่มี three distinct stages แต่ละ stage มีเป้าหมายและ data requirements ต่างกัน

SFT (Supervised Fine-Tuning): สอน model ให้ทำ task เฉพาะ จากนั้น RLHF (Reinforcement Learning from Human Feedback): ปรับ model ให้ align กับความต้องการของมนุษย์ สุดท้าย DPO (Direct Preference Optimization): ปรับละเอียดโดยไม่ต้องมี reward model

คิดเหมือน martial arts training: SFT คือท่าพื้นฐาน, RLHF คือ sparring กับคู่ซ้อม, DPO คือปรับ technique ให้สมบูรณ์แบบ


⚙️ How It Works

The Fine-Tuning Pipeline

Stage 1: SFT (Supervised Fine-Tuning)
├── Data: Instruction-response pairs
├── Goal: Teach model task-specific behavior
└── Output: Fine-tuned model
Stage 2: RLHF (Reinforcement Learning from Human Feedback)
├── Data: Human preference rankings
├── Goal: Align model with human values
└── Output: Reward model + aligned model
Stage 3: DPO (Direct Preference Optimization)
├── Data: Chosen vs rejected responses
├── Goal: Fine-tune without reward model
└── Output: Final aligned model

Data Requirements per Stage

StageData TypeVolumeExample
SFTInstruction → Response10K-100K pairs“Summarize this” → summary
RLHFPrompt → Response A > Response B50K-200K rankingsHuman rankings
DPOPrompt → Chosen + Rejected10K-50K pairsBest vs worst responses

Evaluation Metrics

StageMetricsWhat It Measures
SFTAccuracy, ROUGE, BLEUTask performance
RLHFWin rate, reward scoreHuman preference alignment
DPODPO loss, preference accuracyDirect preference optimization

💡 Example: Designing the Pipeline

Step 1: Define SFT stage

## SFT Stage
- **Data format:** {"instruction": "...", "response": "..."}
- **Volume:** 50,000 instruction-response pairs
- **Training:** Standard supervised learning, 3 epochs
- **Evaluation:** Task accuracy on held-out set

Step 2: Define RLHF stage

## RLHF Stage
- **Data format:** {"prompt": "...", "response_a": "...", "response_b": "...", "ranking": [0, 1]}
- **Volume:** 100,000 human preference rankings
- **Training:** PPO with reward model
- **Evaluation:** Win rate against SFT baseline

Step 3: Define DPO stage

## DPO Stage
- **Data format:** {"prompt": "...", "chosen": "...", "rejected": "..."}
- **Volume:** 25,000 chosen-rejected pairs
- **Training:** DPO loss (no reward model needed)
- **Evaluation:** Preference accuracy + safety metrics

⚠️ Common Mistakes

Mistake 1: Skipping SFT

“I’ll go straight to RLHF” → SFT provides the foundation. Without it, RLHF has nothing to align.

Mistake 2: Using wrong data format per stage

“I’ll use the same data format for all stages” → Each stage needs specific data format. SFT needs instruction-response, RLHF needs rankings, DPO needs chosen-rejected pairs.

Mistake 3: No evaluation plan

“I’ll just train and hope it works” → Define metrics for each stage. Without evaluation, you can’t measure improvement.

Mistake 4: Overlooking data quality

“More data is always better” → Quality > quantity. 10K high-quality pairs beat 1M noisy ones.


📝 Knowledge Check

📝 Knowledge Check

Q1:What are the three stages of LLM fine-tuning in order?

Q2:What data format does the SFT stage require?

Q3:Why is DPO preferred over RLHF in some cases?


🏋️ Quest: Fine-Tuning Pipeline Design

Now it’s time to design a complete fine-tuning pipeline!

  1. Download ไฟล์เริ่มต้นของ quest:

    Terminal window
    npx bluebeltdojo download quest-16-finetuning-pipeline
    cd quest-16-finetuning-pipeline
  2. เปิด problem.js ใน editor ของคุณพร้อมความช่วยเหลือของ AI

  3. สร้างไฟล์ finetuning-design.md ที่ครอบคลุม SFT, RLHF, DPO stages

  4. ระบุ data requirements และ evaluation metrics สำหรับแต่ละ stage

  5. ตรวจสอบ solution ของคุณ:

    Terminal window
    node test.js
  6. When all tests pass, submit your solution:

    Terminal window
    npx bluebeltdojo submit

💡 Tip: Fine-tuning คือ investment — ออกแบบ pipeline ให้ดีก่อน train จะประหยัดเวลาและ compute ได้มหาศาล


คำใบ้

  • อ่าน instructions ใน problem.js อย่างละเอียด
  • ออกแบบ pipeline ครบทั้ง 3 stages: SFT → RLHF → DPO
  • ระบุ data format, volume, และ metrics สำหรับแต่ละ stage
  • ถ้าติดขัด ลองอ่าน “Common Mistakes” อีกครั้ง — อย่าดู solution โดยตรง