Quest 4 - Token Counter
Quest 4: Token Counter
medium 25 minutes🎯 Learning Objectives
- How tokens work in LLMs and why they matter for cost/quality
- How to implement a basic token counting algorithm
- Why naive tokenization (split by spaces) is inaccurate
- The engineering habit: UNDERSTAND YOUR INPUTS
📖 Concept: Tokenization
ในโลกของ LLMs ไม่ได้อ่านคำเต็ม — แต่อ่าน tokens ซึ่งเป็นชิ้นส่วนเล็กๆ ของข้อความ เช่น “hello” อาจเป็น 1 token, แต่ “unbelievable” อาจเป็น 2-3 tokens
การเข้าใจ tokenization สำคัญเพราะ:
- Cost: คุณจ่ายตามจำนวน tokens ที่ใช้
- Quality: context window มีจำกัด — ถ้า tokens มากเกินไป ข้อความจะถูกตัด
- Accuracy: ถ้าคุณนับ tokens ผิด คุณจะวางแผน budget ผิด
⚙️ How It Works
Token คืออะไร?
"Hello, world!" → ["Hello", ",", " world", "!"] = 4 tokens"unbelievable" → ["un", "believ", "able"] = 3 tokensToken ไม่ใช่คำ — เป็น subword ที่ถูกแบ่งด้วย algorithm เช่น BPE (Byte-Pair Encoding)
การนับ Tokens แบบง่าย
// วิธีง่ายๆ แต่ไม่แม่นยำ — แบ่งด้วย spacefunction countTokensSimple(text) { return text.split(/\s+/).filter(t => t.length > 0).length;}
// ปัญหา: "Hello, world!" = 2 tokens แต่จริงๆ ควรเป็น 4การนับ Tokens แบบละเอียด
// แบ่ง word + punctuation แยกกันfunction countTokens(text) { if (text === '') return 0;
const tokens = []; let current = '';
for (const ch of text) { if (/[.,!?;:'"()[\]{}]/.test(ch)) { // Punctuation = separate token if (current) tokens.push(current); tokens.push(ch); current = ''; } else if (/\s/.test(ch)) { if (current) tokens.push(current); current = ''; } else { current += ch; } } if (current) tokens.push(current);
return tokens.length;}💡 Example: Token Counting in Practice
ตัวอย่างที่ 1: ประโยคสั้น
countTokens("Hi!") // → 2 tokens: ["Hi", "!"]countTokens("Hello world") // → 2 tokens: ["Hello", "world"]ตัวอย่างที่ 2: ประโยคยาว
countTokens("I can't believe it's not butter!")// → ["I", "can", "'", "t", "believe", "it", "'", "s", "not", "butter", "!"]// → 11 tokensตัวอย่างที่ 3: Edge cases
countTokens("") // → 0 tokenscountTokens(" ") // → 0 tokens (whitespace only)countTokens("hello world") // → 2 tokens (multiple spaces)⚠️ Common Mistakes
Mistake 1: ใช้ split(” ”) อย่างเดียว
"Hello, world!".split(" ")= 2 items แต่จริงๆ ควรเป็น 4 tokens → ต้องแยก punctuation ออก
Mistake 2: ไม่จัดการ empty string
ถ้า text = "" แล้ว return undefined → ต้อง return 0
Mistake 3: นับ whitespace เป็น token
" hello "นับ spaces เป็น tokens → whitespace ไม่ควรถูกนับเป็น token แยก
Mistake 4: ไม่ grouped consecutive punctuation
"!!!"นับเป็น 3 tokens แต่จริงๆ ควรเป็น 1 token → consecutive punctuation ควรรวมกัน
📝 Knowledge Check
📝 Knowledge Check
Q1:Why is `text.split(' ')` not accurate for token counting?
Q2:How should consecutive punctuation like '!!!' be tokenized?
Q3:What is the token count for an empty string?
🏋️ Quest: Token Counter
เขียน function ที่นับ tokens ในข้อความอย่างแม่นยำ
-
Download ไฟล์เริ่มต้นของ quest:
Terminal window npx bluebeltdojo download quest-04-token-countercd quest-04-token-counter -
เปิด
problem.jsใน editor ของคุณพร้อม AI tool -
ให้ AI ช่วย implement
countTokens(text)ตาม instructions -
ตรวจสอบ:
Terminal window node test.js -
ทดสอบ edge cases: empty string, whitespace only, consecutive punctuation
-
ส่งคำตอบ:
Terminal window npx bluebeltdojo submit
💡 Tip: ถ้า AI ให้แค่
text.split(" ").lengthให้บอกว่าต้องจัดการ punctuation ด้วย
คำใบ้
- อ่าน instructions ใน
problem.jsอย่างละเอียด — มี edge cases ระบุไว้ - ต้องจัดการ: empty string, whitespace, punctuation ติดกัน
- ทดสอบ “Hello, world!” — ควรนับเป็น 4 tokens ไม่ใช่ 2
- ถ้าติดขัด ลองดู tokenization examples ใน solution directory