Quest 14 - Scaling Laws Calculator
Quest 14: Scaling Laws Calculator
medium 25 minutes🎯 Learning Objectives
- Understand Chinchilla scaling laws and compute-optimal training
- Calculate optimal parameter and data allocation from a compute budget
- Recognize when a model is over-parameterized or under-trained
- Make informed decisions about model size vs. training data tradeoffs
📖 Concept: Scaling Laws
ก่อนจะ train model ขนาดใหญ่ เราต้องถามก่อนว่า — ควรจัดสรร compute budget ให้กับ model size หรือ training data มากกว่ากัน? Scaling Laws ให้คำตอบ
Chinchilla scaling laws (DeepMind, 2022) ค้นพบว่า optimal allocation คือ model parameters (N) และ training tokens (D) ควรขยายเท่ากันตาม compute budget (C):
N ≈ 0.3 × C^0.5 (parameters)D ≈ 0.3 × C^0.5 (tokens)นี่คือเหตุผลที่ GPT-4 ไม่ได้ใหญ่ที่สุด — มัน balanced ระหว่าง model size กับ data quantity อย่างชาญฉลาด
คิดเหมือน martial arts: คุณไม่ได้ฝึกแต่ท่าเตะหนัก ๆ โดยไม่คิดเรื่องความอึด Scaling laws คือสมดุลระหว่าง power กับ endurance
⚙️ How It Works
The Scaling Law Workflow
1. Receive compute budget C (in FLOPs) ↓2. Compute optimal parameters: N ≈ 0.3 × C^0.5 ↓3. Compute optimal tokens: D ≈ 0.3 × C^0.5 ↓4. Compare N vs D to determine ratio ↓5. Return: { parameters, tokens, ratio }Understanding the Ratio
The ratio field tells you whether the allocation is balanced:
| Ratio | Meaning | Problem |
|---|---|---|
compute-optimal | N and D are balanced | Ideal allocation |
over-parameterized | Too many params, not enough data | Model memorizes instead of learning |
under-trained | Too much data, too few params | Model can’t capture patterns |
Why This Matters in Practice
Budget: 1e18 FLOPs
Optimal: parameters ≈ 9.5M, tokens ≈ 9.5MOver-param: 50M params, 1M tokens → memorizationUnder-trained: 1M params, 50M tokens → underfitting💡 Example: Computing Optimal Allocation
Step 1: Implement the scaling law
function computeOptimal(computeBudget) { const N = 0.3 * Math.pow(computeBudget, 0.5); const D = 0.3 * Math.pow(computeBudget, 0.5);
return { parameters: N, tokens: D, ratio: 'compute-optimal' };}Step 2: Verify with different budgets
// Large budgetcomputeOptimal(1e18) // ~9.5M params, ~9.5M tokens
// Small budgetcomputeOptimal(1e15) // ~949K params, ~949K tokens// Smaller budget → smaller model (sanity check)Step 3: Run the tests
node test.js# Verifies scaling law formula# Checks smaller budget → smaller model# Validates ratio classification⚠️ Common Mistakes
Mistake 1: Using wrong exponent
“C^0.5 looks arbitrary, maybe it’s C^0.3?” → The 0.5 exponent comes from Chinchilla research. Don’t guess — use the empirical formula.
Mistake 2: Not returning all three fields
“Just returning parameters is enough” → Return
{ parameters, tokens, ratio }— the ratio field is essential for evaluation.
Mistake 3: Ignoring the ratio classification
“Parameters and tokens are calculated, that’s enough” → Classify whether the allocation is compute-optimal, over-parameterized, or under-trained.
Mistake 4: Integer-only allocation
“I’ll round everything to integers” → Keep floating-point precision for the parameters and tokens. Rounding too early loses accuracy.
📝 Knowledge Check
📝 Knowledge Check
Q1:According to Chinchilla scaling laws, how should parameters (N) and tokens (D) scale with compute budget (C)?
Q2:What does an 'over-parameterized' ratio indicate?
Q3:Why should you keep floating-point precision when computing optimal allocation?
🏋️ Quest: Scaling Laws Calculator
Now it’s time to implement Chinchilla scaling laws!
-
Download ไฟล์เริ่มต้นของ quest:
Terminal window npx bluebeltdojo download quest-141-scaling-lawscd quest-141-scaling-laws -
เปิด
problem.jsใน editor ของคุณพร้อมความช่วยเหลือของ AI -
Implement
computeOptimal(computeBudget)— คำนวณ optimal parameter/data allocation -
ทดสอบกับ budget หลายขนาด — ต้องมี scaling ที่ถูกต้อง
-
ตรวจสอบ solution ของคุณ:
Terminal window node test.js -
When all tests pass, submit your solution:
Terminal window npx bluebeltdojo submit
💡 Tip: Scaling laws ไม่ใช่แค่ theory — มันกำหนดว่า company ไหนจะ train model ได้คุ้มค่าที่สุด
คำใบ้
- อ่าน instructions ใน
problem.jsอย่างละเอียด - ใช้สูตร Chinchilla: N ≈ 0.3 × C^0.5, D ≈ 0.3 × C^0.5
- ตรวจสอบ edge cases: budget ต่างขนาด, ratio classification
- ถ้าติดขัด ลองอ่าน “Common Mistakes” อีกครั้ง — อย่าดู solution โดยตรง