
Playbooks 12 min read· 17 Aug 2026· By Playbooks
12 tactics to cut your LLM bill in half
Real, measured cost reductions from teams shipping in production.
A 50% cost reduction rarely comes from one heroic change. It's twelve 5% wins compounding.
The list#
- Downgrade to the smallest model that passes your evals
- Prompt cache the system message
- Shorten the system prompt (measure eval delta)
- Cap
max_tokensaggressively - Use
response_format: json_schemato eliminate retries - Deduplicate identical requests with a cache
- Batch embeddings 100 at a time
- Use TokenBaazar's cheaper model routing for non-critical paths
- Stream to abort early on user cancel
- Compress RAG context (extractive summarization pass)
- Move cold prompts to Gemini Flash
- Audit your top-10 highest-cost endpoints monthly
Every one of these is a 5-15% saving. Stack them and you're at 50%+ without touching quality.