Quest 108 - Multimodal Input Processor
Quest 108: Multimodal Input Processor
medium 25-30 minutes🎯 Learning Objectives
- How to process combined text and image inputs into a unified format
- Why images have a fixed high token cost regardless of content length
- How to estimate token usage across different modalities
- The importance of token limits when building multimodal prompts
📖 Concept: Handling Multiple Modalities
Modern AI doesn’t just process text — it handles images, audio, and video together in a single request. This is called multimodal processing. When you send a request to GPT-4V or Claude with both text and images, the system must handle each input type differently while producing a unified output.
The key insight is that different modalities have different token costs. Text tokens scale with content length (~4 characters per token), but images have a fixed cost (typically 1000 tokens per image in simplified models). This means two images cost the same whether they’re simple icons or complex photographs.
A multimodal input processor separates text and image inputs, estimates token usage, and generates a combined prompt with [IMAGE] placeholders where images should appear.
⚙️ How It Works
The Multimodal Processing Pipeline
1. Receive mixed input array ↓2. Separate text parts from image parts ↓3. Count tokens: - Text: ceil(content.length / 4) - Images: 1000 tokens each (fixed) ↓4. Build combined prompt with [IMAGE] placeholders ↓5. Check against token limit (default: 4000) ↓6. Return unified output structureWhy Images Have Fixed Token Cost
Text tokens depend on content:
"Hello" → 1-2 tokens"Hello world" → 2-3 tokens"A 500-word essay..." → ~125 tokensImage tokens are fixed because the model processes images through a vision encoder that produces a standard number of tokens, regardless of image complexity. This is why you must budget carefully when mixing text and images.
💡 Example: Processing Multimodal Inputs
Here’s how to implement a multimodal processor:
Step 1: Separate text and images
const textParts = [];const imageRefs = [];const combinedParts = [];
for (const input of inputs) { if (input.type === 'text') { textParts.push(input.content); combinedParts.push(input.content); } else if (input.type === 'image') { imageRefs.push({ content: input.content, metadata: input.metadata || {} }); combinedParts.push('[IMAGE]'); }}Step 2: Estimate token count
const IMAGE_TOKEN_COST = 1000;const CHARS_PER_TOKEN = 4;const MAX_TOKENS = 4000;
let tokens = 0;for (const input of inputs) { if (input.type === 'text') { tokens += Math.ceil(input.content.length / CHARS_PER_TOKEN); } else if (input.type === 'image') { tokens += IMAGE_TOKEN_COST; }}Step 3: Build the unified output
return { textParts, imageRefs, combined: combinedParts.join(' '), tokens: Math.min(tokens, MAX_TOKENS),};⚠️ Common Mistakes
Mistake 1: Counting image tokens based on content length
“This image description is 20 tokens, so the image is 20 tokens” → Images have a fixed token cost (1000 tokens each), regardless of their content or description length.
Mistake 2: Ignoring the token limit
“Just send everything, the API will handle it” → Token limits exist for a reason. Exceeding them causes truncation or errors. Always check and warn.
Mistake 3: Not generating [IMAGE] placeholders
“I separated text and images, that’s enough” → The combined prompt needs
[IMAGE]markers so the model knows where to expect visual input.
Mistake 4: Returning empty arrays for no input
“If there’s no input, just return nothing” → Always return the full structure with empty arrays and 0 tokens. Callers expect a consistent shape.
📝 Knowledge Check
📝 Knowledge Check
Q1:Why do images have a fixed token cost in multimodal models?
Q2:What is the correct way to estimate text tokens?
Q3:Why should you return a consistent structure even when there's no input?
🏋️ Quest: Multimodal Input Processor
Now it’s time to practice! Build a processor that handles mixed text and image inputs.
-
Download the starter files:
Terminal window npx bluebeltdojo download quest-108-multimodal-processorcd quest-108-multimodal-processor -
Open
problem.jsin your editor with your AI tool -
Implement
processMultimodal(inputs)that:- Separates text and image inputs into their respective arrays
- Estimates token count (text: ~4 chars/token, images: 1000 tokens each)
- Generates combined prompt with
[IMAGE]placeholders - Respects the default 4000 token limit
-
Run
node test.jsand read the failures carefully -
Fix any edge cases the AI missed
-
Verify all tests pass:
Terminal window node test.js -
When all tests pass, submit your solution:
Terminal window npx bluebeltdojo submit
💡 Tip: The most common mistake is treating images like text for token counting. Images always cost 1000 tokens each, regardless of content.
คำใบ้
- อ่าน instructions ใน
problem.jsอย่างละเอียด - Images ราคา 1000 tokens ต่อภาพเสมอ — ไม่เกี่ยวกับ content length
- Text ราคา ~4 ตัวอักษรต่อ token
- สร้าง combined prompt ด้วย
[IMAGE]placeholders - ตรวจสอบ token limit (default 4000)
- ถ้าติดขัด ลองอ่าน “Common Mistakes” อีกครั้ง — อย่าดู solution โดยตรง