Skip to main content
~100%
of real LLM spend was going unrecorded
~75%
of lakehouse spend eliminated
Per tenant
cost and margin now computable
4
reliability failure classes closed

Client: A multi-tenant, AI-powered media- and market-intelligence platform. It ingests public content, extracts source-backed facts with large language models, and delivers corroborated intelligence to executives through a six-stage LLM pipeline on Google Kubernetes Engine. Its pricing depends on sharing per-entity processing cost across every tenant that tracks that entity.

The challenge

Margins that improve as tenants are added only work if cost can be traced back to tenants. It couldn’t:

  • About 100% of real spend was unrecorded. In one month the AI provider billed about $86, while the platform’s own telemetry had logged roughly $1.70 all-time in development — and nothing, ever, in staging or production.
  • About 73% of the cloud bill could not be grouped by product. Compute was the largest line and carried no attribution.
  • One tenant’s runaway workload could use up the whole daily budget and stop scheduling for everyone else.
  • Reliability hazards made it worse: liveness probes killing workers mid-computation, a message-queue deadlock, and infrastructure lock races failing unrelated pipelines.

What we did

  • Moved cost recording to the provider layer, so any call path that forgets to record fails loudly instead of silently logging nothing.
  • Fixed attribution at the source. Two live defects were billing spend to the wrong tenant or to no one; both now use explicit, leak-proof tenant scoping.
  • Turned on per-tenant daily budgets, so a runaway tenant can no longer stop the shared scheduler.
  • Cut idle infrastructure. We stopped writes to a lakehouse tier that billed per active table in environments that had never had a tenant — about 75% of that bill, with no loss of function.
  • Restored reliability: a file-based worker heartbeat ended probe kills, a dedicated queue drain fixed the deadlock, lock timeouts stopped false failures, and a corrected CI capacity fallback ended org-wide build outages.

Results

OutcomeResult
LLM cost observabilityFrom ~0% captured to full provider-layer capture with tenant attribution
Lakehouse spend~75% eliminated
AttributionProduct-level cost allocation; per-tenant margin now computable
ReliabilityPlatform-wide outage class closed; probe-kill, deadlock, and CI-outage classes fixed

The business gained what its pricing model needed and never had: the ability to see, per tenant and per entity, what it costs to serve them — with a lower, cleaner bill along the way.

On production agent-platform work, including for Baseline, the same discipline took daily model spend from about $107 to $22 — a 79% reduction — through a three-layer cache (exact match, semantic match, and provider-side caching) with a 98.8% hit rate on the busiest path, routing routine summarization and classification to smaller models while keeping frontier models for code review and human-facing replies, compressing tool output before it reaches the model, and CI checks that fail on cost regressions.

The client's identity is anonymized. The environment, findings, and results are drawn from a real Hat Boy Software engagement; figures are representative and rounded.

More case studies

Not sure where to start? Start with an assessment.

A senior review of your app, cloud estate, or AI platform, scoped and quoted before work starts, that ends in a prioritized plan, so you decide what to fix and when.

Talk to an engineer