How to Cut LLM Costs: A Practical Guide
A practical guide to cutting LLM costs: attribute every token, route to cheaper models, trim prompts, cache, and cap retries — without hurting output quality.
To cut LLM costs without hurting quality, start by attributing every token to the model, agent, and workflow that spent it — then pull four levers in order: route easy work to a cheaper model, trim prompt and context bloat, cache repeated calls, and cap runaway retries. Most teams recover 20–40% of spend this way, because the expensive traffic is rarely what they guessed.
Most teams meet their LLM bill the same way: a finance ping about a number that doubled. The provider dashboard shows a total, maybe a split by API key, but not the thing you need — which workflow, agent, or feature spent the money, and whether that spend bought anything. This guide is the practical sequel to attribution: once you can see where the money goes, these are the highest-leverage ways to bring it down without degrading output. If you have not set up measurement yet, start with the companion piece on how to track and cut LLM token costs; if you want a quick planning number first, the free LLM cost estimator models your monthly spend, retry tax included, in a minute.
Step 1: Attribute every token before you cut anything
You cannot optimize what you cannot see, and cutting blind is how teams break quality to save pennies. Capture input, output, and cached tokens on every call, and tag each call with the model, the agent or workflow, and ideally the team or feature. Once spend is attributed this way, the expensive 10% of traffic usually becomes obvious — and it is rarely what anyone guessed. A background classification job quietly running on a frontier model, or one chatty agent that re-sends its whole history on every turn, routinely outspends the feature everyone assumed was the budget.
This is exactly what cost intelligence does: it prices each call and rolls spend up per model, agent, and workflow, so cost stops being an invoice-level abstraction and becomes an engineering signal you can act on. The rest of this guide assumes you have that view — because every lever below is only safe to pull once you can measure whether it actually saved money and left quality intact.
Step 2: Route each job to the cheapest model that passes
The single biggest lever is model selection, and the mistake is picking one model for everything. Most AI estates run a frontier model on work a far cheaper model handles identically — classification, extraction, short rewrites, routing decisions. The fix is to grade each workload on your own outputs, not on a public benchmark: take real production requests, run them through a cheaper model, and compare against your actual quality bar. Where the cheaper model ties, downgrade it; where it slips, keep the stronger model. That per-workload routing is where the 20–40% usually comes from.
Keeping that flexible is easier when your providers are interchangeable. Many teams front their traffic with a router like OpenRouter so they can move a workload between OpenAI and Anthropic's Claude models without a code change, then let the cost and quality data decide which one each job runs on. The goal is not the cheapest model everywhere — it is the cheapest model that still passes, chosen per job and re-checked as providers ship new tiers.
Step 3: Trim prompts, cache, and cap retries
After model choice, three mechanical levers recover the rest. Each is measurable once you have per-call attribution, and together they attack the waste that produces nothing:
| Lever | What it fixes | How to apply it |
|---|---|---|
| Trim prompt & context | System prompts, retrieved chunks, and chat history that grow over time | Cut context that does not change the answer; truncate history; re-rank retrieval to fewer, better chunks |
| Cache repeated calls | Identical or near-identical requests paid for twice | Cache deterministic calls; use provider prompt caching for stable system prompts |
| Cap runaway retries | A failing call retried until it bills several times for nothing | Set a max-retry ceiling, make writes idempotent, and alert on retry spikes |
| Kill zombie workflows | Scheduled jobs still running after anyone consumes their output | Audit what each workflow feeds; retire the ones nothing reads |
Why retries and prompt bloat hide the most waste
Two of those levers deserve emphasis because they are where the quiet money lives. Retries are the biggest: a tool call times out, the agent retries, and you pay for every attempt — at a 15% failure rate that is real spend producing nothing, and it never shows up as a separate line on the invoice. Prompt bloat is next: system prompts, retrieved context, and chat history accrete over months, so the average call costs more this quarter than last with no single change to point at. Token-per-execution is the leading indicator for both — it moves before the invoice does, which is why watching it per workflow catches creep early. For the full picture of how unmonitored spend compounds, see the true cost of unmonitored AI.
A worked example: the expensive 10%
The payoff of attribution is concrete, so here is the shape it usually takes. A team pulls spend per workflow for the first time and finds the top line is not the customer-facing chatbot everyone worried about — it is a nightly enrichment job running a frontier model over thousands of records, most of which a small model would classify identically. Second on the list is an agent that re-sends its entire conversation history on every turn, so a ten-turn chat pays for the first message ten times over. Third is a workflow nobody owns anymore, still firing on schedule, feeding a dashboard that was deprecated last quarter.
None of those are visible on a provider invoice, which shows one total per key. Attributed per workflow, each has an obvious fix: downgrade the enrichment model after validating it on real records, truncate or summarize the agent's history, and retire the zombie workflow outright. Together they are routinely a quarter to a third of the bill, recovered without touching anything a user can see. That is why attribution comes first and the levers come second — the levers only pay off when you aim them at the traffic that is actually expensive, which is almost never the traffic people assumed it was.
Step 4: Make it continuous, not a one-off
Cost creep comes back the moment you stop watching, because the things that cause it — growing context, new consumers, a model quietly swapped for a bigger one — are normal product evolution, not bugs. So treat cost like a production metric: set a budget per workflow, alert on spend anomalies, and re-run your model comparisons periodically as providers change pricing and ship cheaper tiers. The aim is a loop, not a cleanup.
Obsivara automates that loop — per-model, per-agent, and per-workflow spend; waste and retry detection; model comparison on your live traffic; and anomaly alerts when cost drifts from a workflow's baseline — all ingested out of band with no proxy in your request path. You can start on the free tier and move to a paid plan as your estate grows (see pricing), or run a free 48-hour audit to see your real spend and the recoverable waste in it before you change a thing.
Frequently asked questions
How do I cut LLM costs without hurting output quality?
Attribute every token first, then pull levers in order: route each workload to the cheapest model that still passes your own quality bar (validated on real requests, not benchmarks), trim prompt and context that does not change the answer, cache repeated calls, and cap runaway retries. Validate each change against production traffic. Done this way, most teams cut 20–40% of spend with no measurable quality loss.
What is the biggest source of wasted LLM spend?
Usually a frontier model doing work a cheaper model handles identically, closely followed by retry loops that pay for the same call several times and prompts bloated with context the model never uses. All three are invisible on a provider invoice, which shows only a total — you need per-call attribution by model, agent, and workflow to see them, after which the expensive 10% of traffic is usually obvious.
How do I stop LLM costs from creeping back up?
Treat cost as a continuous metric, not a one-time cleanup. Set a budget per workflow, alert on spend anomalies against each workflow's baseline, watch token-per-execution as a leading indicator, and re-run model comparisons as providers change pricing. Cost creep returns because growing context and new consumers are normal evolution, so the fix is an ongoing loop rather than a single audit.
See your AI operations clearly.
Connect your stack in minutes — Obsivara discovers every workflow, scores its health, and attributes every dollar.