agents/prod
09
Scaling & cost

LLM calls dominate the bill

Reported  Cost figures are vendor/blog-sourced — directional. But the shape is consistent: inference is the line item, and routing + caching are the levers.

share of total agent operating cost
LLM API · 70–85%
infra · everything else
  • Multi-model routing −55–65%
  • Semantic caching −30–50%
  • Batch API pricing −~50%
  • Checkpointing pay once