skipLink.label

Quest 4 - Token Counter

Quest 4: Token Counter

medium 25 minutes

🎯 Learning Objectives

  • ✅ How tokens work in LLMs and why they matter for cost/quality
  • ✅ How to implement a basic token counting algorithm
  • ✅ Why naive tokenization (split by spaces) is inaccurate
  • ✅ The engineering habit: UNDERSTAND YOUR INPUTS

📖 Concept: Tokenization

ในโลกของ LLMs ไม่ได้อ่านคำเต็ม — แต่อ่าน tokens ซึ่งเป็นชิ้นส่วนเล็กๆ ของข้อความ เช่น “hello” อาจเป็น 1 token, แต่ “unbelievable” อาจเป็น 2-3 tokens

การเข้าใจ tokenization สำคัญเพราะ:

  • Cost: คุณจ่ายตามจำนวน tokens ที่ใช้
  • Quality: context window มีจำกัด — ถ้า tokens มากเกินไป ข้อความจะถูกตัด
  • Accuracy: ถ้าคุณนับ tokens ผิด คุณจะวางแผน budget ผิด

⚙️ How It Works

Token คืออะไร?

"Hello, world!" → ["Hello", ",", " world", "!"] = 4 tokens
"unbelievable" → ["un", "believ", "able"] = 3 tokens

Token ไม่ใช่คำ — เป็น subword ที่ถูกแบ่งด้วย algorithm เช่น BPE (Byte-Pair Encoding)

การนับ Tokens แบบง่าย

// วิธีง่ายๆ แต่ไม่แม่นยำ — แบ่งด้วย space
function countTokensSimple(text) {
return text.split(/\s+/).filter(t => t.length > 0).length;
}
// ปัญหา: "Hello, world!" = 2 tokens แต่จริงๆ ควรเป็น 4

การนับ Tokens แบบละเอียด

// แบ่ง word + punctuation แยกกัน
function countTokens(text) {
if (text === '') return 0;
const tokens = [];
let current = '';
for (const ch of text) {
if (/[.,!?;:'"()[\]{}]/.test(ch)) {
// Punctuation = separate token
if (current) tokens.push(current);
tokens.push(ch);
current = '';
} else if (/\s/.test(ch)) {
if (current) tokens.push(current);
current = '';
} else {
current += ch;
}
}
if (current) tokens.push(current);
return tokens.length;
}

💡 Example: Token Counting in Practice

ตัวอย่างที่ 1: ประโยคสั้น

countTokens("Hi!") // → 2 tokens: ["Hi", "!"]
countTokens("Hello world") // → 2 tokens: ["Hello", "world"]

ตัวอย่างที่ 2: ประโยคยาว

countTokens("I can't believe it's not butter!")
// → ["I", "can", "'", "t", "believe", "it", "'", "s", "not", "butter", "!"]
// → 11 tokens

ตัวอย่างที่ 3: Edge cases

countTokens("") // → 0 tokens
countTokens(" ") // → 0 tokens (whitespace only)
countTokens("hello world") // → 2 tokens (multiple spaces)

⚠️ Common Mistakes

Mistake 1: ใช้ split(” ”) อย่างเดียว

"Hello, world!".split(" ") = 2 items แต่จริงๆ ควรเป็น 4 tokens → ต้องแยก punctuation ออก

Mistake 2: ไม่จัดการ empty string

ถ้า text = "" แล้ว return undefined → ต้อง return 0

Mistake 3: นับ whitespace เป็น token

" hello " นับ spaces เป็น tokens → whitespace ไม่ควรถูกนับเป็น token แยก

Mistake 4: ไม่ grouped consecutive punctuation

"!!!" นับเป็น 3 tokens แต่จริงๆ ควรเป็น 1 token → consecutive punctuation ควรรวมกัน


📝 Knowledge Check

📝 Knowledge Check

Q1:Why is `text.split(' ')` not accurate for token counting?

Q2:How should consecutive punctuation like '!!!' be tokenized?

Q3:What is the token count for an empty string?


🏋️ Quest: Token Counter

เขียน function ที่นับ tokens ในข้อความอย่างแม่นยำ

  1. Download ไฟล์เริ่มต้นของ quest:

    Terminal window
    npx bluebeltdojo download quest-04-token-counter
    cd quest-04-token-counter
  2. เปิด problem.js ใน editor ของคุณพร้อม AI tool

  3. ให้ AI ช่วย implement countTokens(text) ตาม instructions

  4. ตรวจสอบ:

    Terminal window
    node test.js
  5. ทดสอบ edge cases: empty string, whitespace only, consecutive punctuation

  6. ส่งคำตอบ:

    Terminal window
    npx bluebeltdojo submit

💡 Tip: ถ้า AI ให้แค่ text.split(" ").length ให้บอกว่าต้องจัดการ punctuation ด้วย


คำใบ้

  • อ่าน instructions ใน problem.js อย่างละเอียด — มี edge cases ระบุไว้
  • ต้องจัดการ: empty string, whitespace, punctuation ติดกัน
  • ทดสอบ “Hello, world!” — ควรนับเป็น 4 tokens ไม่ใช่ 2
  • ถ้าติดขัด ลองดู tokenization examples ใน solution directory