AI / ML · Operating LLMs
LLM cost & latency: math before scaling
Back-of-the-envelope math, the levers that matter (caching, batching, model choice), and the failure mode that bankrupts new LLM features.
The most common reason LLM features die at scale isn't quality. It's economics. A side project that worked beautifully at 100 calls/day melts the credit card at 100K calls/day. A latency that was acceptable for a power user becomes catastrophic when you put it on a checkout page.
This article is the napkin you should have run before announcing the feature.
The cost formula#
For each call:
cost = (input_tokens × input_price) + (output_tokens × output_price)Anthropic's mid-2026 pricing (Claude Sonnet 4.6):
- Input: $3 per million tokens
- Output: $15 per million tokens
- Cached input (5min): $0.30 per million tokens (90% discount)
- Cached input (1hr): $0.60 per million (80%)
Multiply by traffic to get the bill.
A worked example#
Suppose your customer-support copilot:
- Average user query: 50 tokens
- System prompt + retrieved chunks: 4000 tokens (mostly cacheable)
- Average response: 300 tokens
- Volume: 50,000 queries/day
Naive (no caching):
input = 4050 × 50000 × $3 / 1M = $607.50/day = $18,225/mo
output = 300 × 50000 × $15 / 1M = $225.00/day = $6,750/mo
─────────
$24,975/moWith prompt caching (system + chunks cached 5min, ~95% hit rate during peak):
cached input = 4000 × 50000 × $0.30 / 1M × 0.95 = $57.00/day
fresh input = (4000 × 50000 × $3 / 1M × 0.05) + (50 × 50000 × $3 / 1M)
= $30.00/day + $7.50/day = $37.50/day
output = 300 × 50000 × $15 / 1M = $225.00/day
─────────
total = $319.50/day = $9,585/moThat's 62% off, just from caching. Caching is the single biggest cost lever.
- No cache24,975
- Caching9,585
- Caching+Sonnet→Haiku for routing4,200
- Caching+routing+batching2,900
Latency budget#
For an interactive product, the user's patience budget is roughly:
| Use case | p50 budget | p99 budget |
|---|---|---|
| Background async (email reply) | 30s | 5m |
| Async chat (typing indicator OK) | 5s | 15s |
| Live chat (must feel instant) | 1.5s | 4s |
| Inline UI (autocomplete, suggestions) | 250ms | 800ms |
The numbers below are typical for current cloud LLMs:
- Time-to-first-token (TTFT): ~400ms (cached prompt), ~1s (fresh)
- Per-output-token: ~30ms (Sonnet), ~15ms (Haiku), ~50ms (Opus)
So a 300-token Sonnet response with prompt caching: 400ms + 9000ms = 9.4s. That's outside the live-chat budget. Either go to Haiku (300 × 15 = 4.5s, just inside), use streaming so the user sees output progressively, or shrink the output.
The five levers#
Ordered by impact:
1. Prompt caching. Always on, for any prompt with stable preludes. 80–90% input cost reduction. No quality cost. The bug is forgetting to enable it.
2. Right-sizing the model. Sonnet for primary tasks, Haiku for routing/classification/extraction, Opus only when the smaller models can't. A two-stage architecture (Haiku to triage, Sonnet to answer) often costs 1/4 of "Sonnet does everything."
3. Output token discipline. The model loves to write paragraphs when a sentence would do. Set max_tokens aggressively. Tell the model "respond in ≤2 sentences." The cost reduction is real and users prefer the shorter answers.
4. Batching. When you need to run inference on a lot of inputs (eval, classification jobs, content moderation), use the batch API. Anthropic's batch API is 50% off, with 24-hour SLA.
5. Streaming + early stop. For chat, stream and let the user interrupt. Implement client-side stop on navigation. Wasted output is wasted money.
The hidden cost: retries#
Always include retries in your math. Real production sees:
- 0.5–1% transient errors (5xx, rate limit, network)
- Some inputs that produce malformed JSON requiring re-prompt
- Some calls that get truncated and need continuation
A 1% retry rate costs you 1% extra. A 1% retry rate where retries average 1.5× the original size (because you added "earlier you returned X, please fix") is 1.5% extra. It adds up.
Set up real observability:
- Log every call with input/output token counts
- Track p50/p95/p99 latency per model and prompt
- Alert on cost per request crossing a threshold
- Sample 1% of completions for quality review
Sizing for scale#
Before launching a new feature, ask:
- What's the expected QPS at launch? At 10× launch?
- What's the per-call cost? (Calculate it.)
- What's the marginal cost per 100K users?
- At what point does the feature pay for itself?
Numbers like "$24K/month for 50K queries/day" sound abstract until they show up on a quarterly review. Ship the math, not just the feature.
A production checklist#
Before any LLM feature goes prod:
- Cost-per-call documented in the PR
- Prompt caching enabled where possible
-
max_tokensset explicitly - Model size justified (smaller would fail)
- Streaming where UX benefits
- Observability: token counts logged, latency traced
- Rate limiting at the application layer (not just provider)
- Backup model configured (provider goes down)
- Eval set in CI
That's the bar. If a PR can't tick those boxes, it's not ready for production traffic.
Further reading#
- Anthropic's prompt caching docs — the single most ROI-positive feature.
- "The Hidden Cost of Streaming" — Simon Willison.
- OpenAI's batch API docs and Anthropic's batch docs — 50% off for non-realtime work.
- LangSmith / Helicone for cost-per-tenant tracking when you need it.