Quest 60 - LLM Observability System
Quest 60: LLM Observability System
hard 30-45 minutes🎯 Learning Objectives
- How to build tracing and metrics for LLM API calls
- Why tracking dollar cost matters more than just counting tokens
- How to break down metrics by model for cost attribution
- How to identify slow queries that exceed latency thresholds
📖 Concept: LLM Observability
LLM calls are expensive, unpredictable, and non-deterministic. Unlike a regular API call that costs $0.001 and takes 50ms, an LLM call might cost $0.15 and take 3 seconds — or $2.00 and take 30 seconds. Without observability, you have no idea how much you’re spending or which calls are slow.
The key engineering habit is: trace every LLM call. Record the model, tokens, latency, cost, and success/failure for every invocation. This data is your financial and operational lifeline.
Think of LLM observability like a dojo tracking training sessions. If you only count “number of sessions,” you miss that some sessions cost 10x more in instructor time. You need to track cost per session, not just count of sessions.
⚙️ How It Works
The Tracing Pipeline
LLM Call Happens ↓trace({ model, prompt, response, tokens, latency }) ↓Calculate cost: inputCost = tokens.input / 1000 × inputPrice outputCost = tokens.output / 1000 × outputPrice totalCost = inputCost + outputCost ↓Store trace with unique ID ↓Aggregate: totalCalls, totalTokens, totalCost, avgLatency, byModelCost Model (GPT-4 Pricing Example)
const PRICING = { 'gpt-4': { input: 0.003, output: 0.006 }, // per 1K tokens 'gpt-3.5': { input: 0.001, output: 0.002 },};
// For a call with 100 input tokens, 500 output tokens on gpt-4:const cost = (100 / 1000 * 0.003) + (500 / 1000 * 0.006);// = 0.0003 + 0.003 = $0.0033The Critical Edge Case: Cost, Not Just Tokens
// ❌ NAIVE: Tracks tokens onlyconst metrics = { totalCalls: 10, totalTokens: 5000 };// Missing: How much did we SPEND? Which model was used?
// ✅ CORRECT: Tracks cost breakdownconst metrics = { totalCalls: 10, totalTokens: 5000, totalCost: 0.045, // Actual dollar amount avgLatency: 1200, // ms byModel: { 'gpt-4': { calls: 3, tokens: 2000, cost: 0.03 }, 'gpt-3.5': { calls: 7, tokens: 3000, cost: 0.015 }, }};💡 Example: Complete Tracing System
function createTracer() { const traces = []; const PRICING = { 'gpt-4': { input: 0.003, output: 0.006 }, 'gpt-3.5': { input: 0.001, output: 0.002 }, };
function trace({ model, prompt, response, tokens, latency }) { const id = `trace-${Date.now()}-${Math.random().toString(36).slice(2, 8)}`; const pricing = PRICING[model] || { input: 0.003, output: 0.006 }; const cost = (tokens.input / 1000 * pricing.input) + (tokens.output / 1000 * pricing.output);
traces.push({ id, model, prompt, response, tokens, latency, cost }); return id; }
function getMetrics() { const totalCost = traces.reduce((sum, t) => sum + t.cost, 0); const totalTokens = traces.reduce((sum, t) => sum + t.tokens.input + t.tokens.output, 0); const avgLatency = traces.reduce((sum, t) => sum + t.latency, 0) / traces.length;
const byModel = {}; for (const t of traces) { if (!byModel[t.model]) byModel[t.model] = { calls: 0, tokens: 0, cost: 0 }; byModel[t.model].calls++; byModel[t.model].tokens += t.tokens.input + t.tokens.output; byModel[t.model].cost += t.cost; }
return { totalCalls: traces.length, totalTokens, totalCost, avgLatency, byModel }; }
return { trace, getTrace: (id) => traces.find(t => t.id === id), getMetrics, getSlowQueries: (ms) => traces.filter(t => t.latency > ms) };}⚠️ Common Mistakes
Mistake 1: Tracking tokens but not cost
“We made 1,000 calls with 500K tokens” → Tokens don’t pay the bill. Dollars do. GPT-4 output tokens cost 6x more than input tokens — token count alone is misleading.
Mistake 2: Not breaking down by model
“Total cost is $45 this month” → Which model is driving the cost? Without per-model breakdown, you can’t optimize.
Mistake 3: Ignoring latency
“The call completed, that’s all that matters” → Slow LLM calls block user interactions. Track latency and flag calls exceeding thresholds.
Mistake 4: Not generating unique trace IDs
“I’ll use the timestamp as the ID” → Multiple calls can happen in the same millisecond. Use random IDs or UUIDs for uniqueness.
📝 Knowledge Check
📝 Knowledge Check
Q1:Why is tracking dollar cost more important than tracking token count for LLM observability?
Q2:What information should each LLM trace record?
Q3:Why is per-model cost breakdown essential in LLM observability?
🏋️ Quest: LLM Observability System
Now it’s time to practice! Build a tracing and metrics system for LLM calls.
-
Download the starter files:
Terminal window npx bluebeltdojo download quest-60-llm-observabilitycd quest-60-llm-observability -
Open
problem.jsin your editor with your AI tool -
Implement
createTracer()that returns an object withtrace,getTrace,getMetrics, andgetSlowQueries -
Important: The critical edge case is COST calculation — naive AI tracks token count but doesn’t compute dollar cost. Use the GPT-4 pricing model.
-
Verify all tests pass:
Terminal window node test.js -
When all tests pass, submit your solution:
Terminal window npx bluebeltdojo submit
💡 Tip: GPT-4 pricing: input $0.003/1K tokens, output $0.006/1K tokens. Cost formula:
(tokens.input / 1000 * 0.003) + (tokens.output / 1000 * 0.006). The test verifies the cost is approximately correct.
คำใบ้
- อ่าน instructions ใน
problem.jsอย่างละเอียด - Edge case ที่สำคัญที่สุด: ต้องคำนวณ dollar cost ไม่ใช่แค่ token count
- GPT-4 pricing: input $0.003/1K tokens, output $0.006/1K tokens
createTracer()ต้อง return object ที่มี methods:trace,getTrace,getMetrics,getSlowQueries- ต้อง generate unique trace ID สำหรับทุก call
- ถ้าติดขัด ลองอ่าน “Common Mistakes” อีกครั้ง — อย่าดู solution โดยตรง