Rainkernel

Insight

The $50,000 hour: why agent cost needs a governor, not an alert

Every budget control sold today stops at the API key. The money is spent inside the agent loop, one retry, one escalation and one spawned sub-agent at a time — and the alert arrives a day later.

· 7 min read · Rainkernel engineering · cost · agents · FinOps

In September 2026 Google Mandiant described an accounting agent that made about 15,000 API calls and spent roughly $50,000 in under an hour. No attacker was involved. The agent was doing what it was built to do — retrying, re-planning, calling tools — and nothing in the loop knew what the task was worth.

Gartner now expects inference cost per agentic workflow to rise more than fivefold through 2028, and has started using the phrase unbounded costs. FinOps teams managing AI spend went from 31% to 98% of organisations in two years. The problem is not awareness. It is that the controls available are the wrong shape.

What the market sells: perimeter budgets

Look at any 2026 round-up of AI cost tools. LLM gateways — LiteLLM, Portkey, Cloudflare AI Gateway — set budgets per provider, model, key, team or user. FinOps platforms — Finout, CloudZero, Vantage — allocate spend and alert on thresholds. Observability suites alert and do not enforce. Cloud budgets lag about 24 hours; in two documented incidents the first signal was a credit-card charge, while the actions that caused the spend had been in CloudTrail within minutes.

All of these controls share one blind spot: they do not know what a task is. A key with a $10,000 monthly budget cannot tell a $4 ticket from a $4,000 one until the budget is gone. Microsoft's open-source Agent Governance Toolkit, released in April 2026, added circuit breakers and error budgets to agent runtimes — and no cost limits.

How the money actually leaves

Agent cost incidents have a recognisable anatomy, and none of the steps is a bug in the model.

  • Retries without a ceiling. A tool times out, the agent tries again, the planner adds a step, the step calls the tool. Each iteration is reasonable; the sum is not.
  • Silent model escalation. One developer's published post-mortem of a July 2026 runaway describes agents stepping themselves up to larger models mid-task; the final bill was close to $80,000 across 162 invoices, and the client-side token counters disagreed with the provider's billing by 8.5×.
  • Spawn storms. The same incident spawned 826 child tasks from one request. Depth and spawn limits would have ended it at the second generation.
  • Tool-chain amplification. Researchers showed in January 2026 that a compromised tool server can edit only the text-visible fields of its responses and induce repetitive tool calls — inflating the cost of a single query by up to 658×, on MCP-based systems, across six models.
  • Actions outside the model. An autonomous agent with cloud credentials re-applied its own infrastructure template until a $5-a-month workload cost $6,531. Token budgets never saw it; CloudTrail did.

What a governor has to do

The controls that stop these are well understood individually; what does not exist as a product for mid-sized companies is the combination, enforced server-side, with the evidence trail. We think it needs seven things.

  • A cost ceiling per task — the unit the business recognises: a resolved ticket, an invoice, a document — not per key or per month.
  • Step, depth and spawn caps with graceful degradation when they are hit, so the user gets a slower or partial answer rather than silence.
  • A model lock: escalation to a larger model is a policy decision with a budget, logged, never a prompt's initiative.
  • A server-side circuit breaker in the gateway, not in client code an agent can route around.
  • Per-agent cloud accounts with kill switches driven by CloudTrail or its equivalents, for the spend that happens outside the model.
  • Nightly reconciliation of client-side token counts against the provider's billing export, with drift reported — because the counters have been shown to be wrong by multiples.
  • A weekly cost-per-task report the CFO can read, drawn from a hash-chained event log that the compliance team can later use as evidence.

Why we are building it

The Agent Cost Governor is the second product on Rainkernel's line-up precisely because the research could not find it anywhere else. It ships first as the headline deliverable of our Run and Improve retainer and then standalone for estates we did not build, at $2,500–5,000 a month per estate. The public gate is dated: by 31 January 2027 it is live in paying estates with published before-and-after cost per task, or we say so on the product page.

If you have had one surprise invoice from an agent and would rather not have the second, tell us about the estate.

Sources

  1. Google Mandiant enterprise AI security risks report, via Help Net Security, 16 September 2026
  2. Gartner: AI inference costs per agentic workflow to increase more than fivefold through 2028, 17 August 2026
  3. InfoQ: AI agents and billing guardrails, 16 July 2026
  4. Finout: 13 tools to control AI spend, 22 July 2026
  5. Microsoft: Introducing the Agent Governance Toolkit, 2 April 2026
  6. Tech Times: tool-call attack inflates agent costs 658×, 14 June 2026
  7. DEV Community: a developer's post-mortem of an 826-thread agent runaway, July 2026 (single-source account)
  8. Gartner: top strategic predictions for 2027 and beyond, 15 September 2026

Statistics are quoted from the sources above and remain the work of their authors. Single-source accounts are labelled as such. Corrections: hello@rainkernel.com.

Tell us the use case. We'll tell you where it would stall.

A 30-minute call, no deck, no charge. You leave with an honest read on production readiness and, if it fits, a fixed-price proposal within 48 hours.