Course Finale
Summary (Part B) · Cost Optimization + PM Perspective
Chapter 2 Review · Page 2 of 2: The five-layer cost optimization system, and the complete engineering perspective an AI PM takes away.
Part 3 · Cost Optimization
III. Five-Layer Cost Optimization System
Combined use can reduce Token costs by 70–90%
Model Routing
Use lightweight models for simple tasks (classification, extraction); reserve flagship models for complex reasoning
Save 60-80%
Syntax Layer
YAML instead of JSON (save 15-30%), CSV instead of arrays (save 40%+), remove Markdown decoration
Save 15-40%
Semantic Layer
Dynamic Few-Shot vector matching (save 87%), LLMLingua-2 compression for long documents (save 60%), place key info at beginning/end
Save 60-87%
Output Layer
Negative constraints (precise elimination of filler), Diff editing (output only changes), Stop Sequence (timely cutoff)
Save 20-50%
KV Cache
Fixed System Prompt prefix → high cache hit rate; avoid dynamic timestamps (different every time = always MISS)
Save 30-60%
❌ Dynamic Timestamp
Different prefix every time
System Prompt contains "current time: {time}", prefix changes every request → KV Cache always MISS
✅ Static Prefix
Prefix unchanged, cache stable
Put time in user message, not system; keep System Prompt fixed → continuous HIT
Part 4 · AI PM Integrated Perspective
IV. AI PM Integrated Perspective
From fundamentals to engineering — a complete perspective
Image Token Cost
Resolution ≠ quality. 512px is sufficient for coarse classification; use 1K–2K for OCR; reserve 4K for high-precision detection. Task-matched resolution can reduce image costs by 60–90%.
Prompt Security
4 attack types: privilege injection / role escape / Few-Shot poisoning / symbol injection. 4 defense layers: input filtering + permission tiering + output validation + audit logging.
Multi-Turn Conversation Cost
Carrying the full history into every turn → exponential cost growth. Active context compression and a history management strategy are essential.
Streaming Output
Markdown and XML are streaming-friendly; JSON/YAML require waiting for the complete response before parsing. Format choice affects user experience and engineering complexity.
The Ultimate Conclusion of Both Chapters
The essence of an AI Harness is carefully designing every Message. From context management to cost optimization, everything comes down to designing the Message List sent to the model.
Capabilities You Have Mastered
Understand LLM Principles
Identify & Mitigate Hallucinations
Design Prompts
Evaluate Output Formats
Plan Agent Architecture
Systematically Reduce Costs
Defend Against Prompt Attacks
Chapter 2 Complete · Congratulations on Finishing the AI Harness Curriculum · From foundational understanding to engineering implementation, you now have the core perspective of an AI PM.