Skip to main content

Cutting an agent platform's LLM spend by 79%


A multi-agent software development platform was spending about $107 a day on model calls before it had any product traffic, and nobody could say which agent was spending it. The fix was attribution first, then caching, routing and hard caps, each verified against a full day of real traffic rather than a projection.

The platform runs a team of AI agents with distinct roles (tech lead, backend developer, QA and so on). Each wakes on a schedule, checks for work and acts, and every model call passes through a metering proxy. In under three weeks, daily spend went from about $107 to about $22, a 79% reduction measured over a full 24-hour window.

Measure first, and close the attribution gap

A meaningful share of calls carried no agent role at all. At one point 10.7% of calls, about $15 a day, were billed to nobody. You cannot tune what you cannot attribute.

We made the role a required field at the metering layer and turned “a call with no role” into an alert condition. Unattributed spend fell to 0.4% of calls, about $0.04 a day. Every later decision depended on answering “what is this agent costing us?” with one query.

An early round of fixes was verified on a 10-minute sample and projected to a 43% reduction. Over the same days, full-day spend rose to about $70, driven by an attribution leak and one agent’s volume regression that a short sample could not show. After that we reported only full 24-hour windows.

Three cache layers

Agent workloads are unusually cacheable. Each agent sends a large, stable system prompt (around 30,000 tokens) and often the same short “check for work” message. On the baseline day, tokens written into the prompt cache outnumbered output tokens roughly 100 to 1. Cache effectiveness was the bill.

LayerWhat it matchesLifetimeCost on a hit
Exact matchSHA-256 of model, system prompt, messages and tools24 hoursNothing billed
Semantic matchEmbedding similarity, cosine 0.97 or higher (0.99 on one model family)StoredNothing billed
Provider-side context cacheThe stable system and tool prefix1 hourAbout 10% of the input rate

The quiet win was the provider-side cache lifetime. The default prompt-cache TTL was 5 minutes, shorter than most agents’ schedules, so nearly every call paid full price for its system prompt. Extending it to one hour helped every call at once. On the busiest model path the hit rate reached 98.8%.

Route by task, not by agent

Most scheduled agent work is observation: read some state, post a one-line summary, or confirm nothing changed. At list prices, Gemini Pro cost about 16 times as much per token as Gemini Flash.

We moved model choice to the task level. Routine sweeps and classifications went to the smallest model that handled them (Flash or Flash-Lite); larger models stayed only where the output needed them. Pro spend fell from about $26 a day, 40% of the total, to $0.36 a day. The biggest lever in that phase was a per-agent model setting for the heartbeat itself, the “do I have work?” check, which had been inheriting each agent’s expensive default.

Stop idle agents from thinking

At baseline, one agent spent 12 model round-trips on a single scheduled prompt, only to conclude there was nothing to do. Each round-trip paid for the full system prompt.

The fix was a cheap check before the model is called. An API looks at messages, recent actions, open directives and assigned issues. If all four are silent, the agent exits. A skipped firing costs about $0.0001 against about $0.003 for a full heartbeat. The first version never fired, because it was scoped to the whole deployment instead of each role. Count the skips over 24 hours before you believe a short-circuit works.

Caps for runaway loops

Once caching and routing were in place, cost spikes were volume spikes, and nearly every one was a loop. We catalogued seven classes and capped the main ones twice, in the prompt and on the server, because prompt rules drift silently.

  • Duplicate posts: server-side dedup on near-identical messages, with the window widened from 5 minutes to 2 hours.
  • Probe loops: status checks capped at two per heartbeat.
  • Fix-push loops: at most three CI-fix attempts per change per agent per hour.
  • Stale directives: old instructions expire automatically instead of re-triggering work.

The loop that taught us the most never produced text. The model emitted only tool-use blocks for built-in runtime actions, never wrote a final reply, and was eventually aborted. None of those actions went through the platform’s command line or tool router, so the loop appeared in no audit table. It burned 79 million tokens in 30 minutes, about $1.40 an hour per agent even on the cheapest model, and hit three agents in one day. The detector we built watches the ratio instead: more than 50 model calls with fewer than 5 tool calls in a 15-minute window flags the agent.

CI gates that fail on cost

Caps decay when someone edits a prompt, so we made them part of the build. One gate fails the pipeline if a canonical hard stop or a cap’s marker string disappears from the agent prompts. A cost-regression gate blocks any new scheduled task that pairs an expensive model with a schedule of 120 minutes or less and synchronous inference. The way out is a smaller model, the batch lane, or a slower schedule.

The unit metric: cost per merged change

Daily spend is the bill; cost per merged change is whether the bill is worth it. It went from about $1.78 at baseline to about $0.30. The caveat: the final figure assumes 74 merged changes a day, which we had not re-counted for the final day. If throughput fell, the gain is smaller than it looks.

The cache-poisoning lesson

One evening, six agents began issuing the same shell command in tight loops. They had different system prompts and schedules, but all were receiving the same upstream response ID.

The semantic cache grouped requests by a structural key of model, tools, max tokens and sampling parameters. The system prompt was left out deliberately, on the theory that embedding similarity would tell agents apart. It did not. The embedding captured only about the first kilobyte of the system prompt, and agents with similar shared boilerplate cleared the 0.97 threshold. One agent’s cached response, tool calls included, was served to another, which executed them, and each replay kept the poisoned entry live.

We muted first and traced second. All 48 scheduled heartbeat jobs were disabled and the platform ran chat-only for about a day, then jobs came back one at a time. Turning the semantic cache off for one model family stopped the loop, but hid the bug for four more days. That was a workaround, not a fix.

The real fix was to put identity in the hash, not in the embedding. A semantic cache key must include:

  • the model, tools and sampling parameters
  • a fingerprint of the system prompt (we hashed its opening bytes; hashing all of it is safer)
  • the tenant and environment namespace
  • the agent’s identity

Similarity should only choose between entries that already match on all of those. We also keep a standing query for any response ID served to more than one agent role; any row it returns is an incident.

How we can help

If your AI platform’s model bill is growing faster than its usage, or you cannot say which agent or tenant is spending it, we start by measuring, then work through attribution, cache design and loop caps. The related case study covers the attribution side, and you can contact us to talk through your own estate.

Back to blog

Not sure where to start? Start with an assessment.

A senior review of your app, cloud estate, or AI platform, scoped and quoted before work starts, that ends in a prioritized plan, so you decide what to fix and when.

Talk to an engineer