Course Finale

Summary (Part B) · Cost Optimization + PM Perspective

Chapter 2 Review · Page 2 of 2: The five-layer cost optimization system, and the complete engineering perspective an AI PM takes away.

Part 3 · Cost Optimization
💰
III. Five-Layer Cost Optimization System
Combined use can reduce Token costs by 70–90%
Model Routing Use lightweight models for simple tasks (classification, extraction); reserve flagship models for complex reasoning Save 60-80%
Syntax Layer YAML instead of JSON (save 15-30%), CSV instead of arrays (save 40%+), remove Markdown decoration Save 15-40%
Semantic Layer Dynamic Few-Shot vector matching (save 87%), LLMLingua-2 compression for long documents (save 60%), place key info at beginning/end Save 60-87%
Output Layer Negative constraints (precise elimination of filler), Diff editing (output only changes), Stop Sequence (timely cutoff) Save 20-50%
KV Cache Fixed System Prompt prefix → high cache hit rate; avoid dynamic timestamps (different every time = always MISS) Save 30-60%
❌ Dynamic Timestamp
Different prefix every time
System Prompt contains "current time: {time}", prefix changes every request → KV Cache always MISS
✅ Static Prefix
Prefix unchanged, cache stable
Put time in user message, not system; keep System Prompt fixed → continuous HIT
Part 4 · AI PM Integrated Perspective
🎯
IV. AI PM Integrated Perspective
From fundamentals to engineering — a complete perspective
🖼️
Image Token Cost
Resolution ≠ quality. 512px is sufficient for coarse classification; use 1K–2K for OCR; reserve 4K for high-precision detection. Task-matched resolution can reduce image costs by 60–90%.
🛡️
Prompt Security
4 attack types: privilege injection / role escape / Few-Shot poisoning / symbol injection. 4 defense layers: input filtering + permission tiering + output validation + audit logging.
📊
Multi-Turn Conversation Cost
Carrying the full history into every turn → exponential cost growth. Active context compression and a history management strategy are essential.
🔄
Streaming Output
Markdown and XML are streaming-friendly; JSON/YAML require waiting for the complete response before parsing. Format choice affects user experience and engineering complexity.
The Ultimate Conclusion of Both Chapters
The essence of an AI Harness is carefully designing every Message. From context management to cost optimization, everything comes down to designing the Message List sent to the model.
Capabilities You Have Mastered
Understand LLM Principles Identify & Mitigate Hallucinations Design Prompts Evaluate Output Formats Plan Agent Architecture Systematically Reduce Costs Defend Against Prompt Attacks
Chapter 2 Complete · Congratulations on Finishing the AI Harness Curriculum · From foundational understanding to engineering implementation, you now have the core perspective of an AI PM.