Agent Cost Governor
Per-task cost ceilings enforced inside the agent loop, with a server-side circuit breaker and nightly reconciliation against the provider bill.
Product details →NewReadiness Kit v0.1 — our open-source agent attack pack and evaluation harness — ships on GitHub on 18 October 2026.Readiness Kit v0.1 — 18 Oct 2026.Open source →
Product 01 · Available now
Production evidence for an AI agent, produced inside your own cloud, in two weeks.

The problem
Agents now complete real computer tasks at near-human rates, and still most pilots stall before production — not because the model is weak, but because nobody can show that the agent is reliable, affordable and safe on the company's own data. Observability is everywhere; systematic evaluation is not. The Gate replaces the opinion with a report.
66% vs 72%
agents complete real computer tasks at 66% against a 72% human baseline — the capability gap is closing fast.
37%
of teams run online evaluations, although 89% have observability — the gap the Gate closes.
36.5%
average success rate of tool-description poisoning attacks on live MCP servers in the MCPTox benchmark.
The place nobody else holds
A test of one deployed agent before launch with a published scope and a published price. Of the providers we reviewed in October 2026, none publishes a fixed price for this; the certificate for AI vendors, ISO/IEC 42001 certification and red-teaming for the model makers are taken, this is not. Read the sample Evidence Report to see exactly what you get.
From the Lab's October 2026 review of the market — the six open places, and how we earn them. The sample Evidence Report is the deliverable in full.
What it does
Mapped to the OWASP Top 10 for Agentic Applications, ASI01 goal hijack through ASI10 rogue agents: tool-description poisoning, SKILL.md and CLAUDE.md context-file poisoning, PipeLeak-style exfiltration through untrusted form input, tool-chain cost amplification — the classes generic LLM red-teaming under-covers.
20–50 tasks drawn from the agent's real failures and your real traffic, scored against expected outcomes, with a held-out set so the number cannot be gamed. The harness stays with you and gates every later release.
Tokens, retries and tool calls per completed task, measured on your traffic — and the projected ceiling at production volume, so the finance conversation happens before go-live, not after the first invoice.
Over-access, shared accounts, credential rotation, decommissioning, and what the agent can reach that it never needs. Written against ASI03 identity and privilege abuse and the NIST autonomy-inventory expectation.
A first CycloneDX-shaped inventory of models, prompts, tools, MCP servers, skills and data sources the agent depends on — the artefact procurement and insurers have started asking for.
Engineering gets the failing cases and the fixes; the CISO gets the mapped findings and blast radius; procurement and audit get the inventory and the controls statement. Produced inside your cloud; the evidence never leaves your estate.
What you receive
Who buys it
Built on open engines: snyk-agent-scan, Cisco mcp-scanner, Agentic Radar, Promptfoo — under Rainkernel's corpus, rubric, report generator and BOM emitter.
Pricing
Sprint — one agent
$7,500
Two weeks, fixed price. ₹4 lakh for Indian clients. Credited against a Build signed within 30 days.
Ask about this tierFull Gate — tool-using agent
$25,000
The complete attack pack, harness, cost baseline, permission audit and BOM for an agent that calls tools and systems.
Ask about this tierFull Gate — multi-agent or MCP estate
$45,000
Multi-agent workflows or an MCP-connected estate; inter-agent and supply-chain classes included.
Ask about this tierPrices in USD and exclusive of applicable taxes; Indian clients are invoiced in INR with GST. Window: Sprint bookings open for October 2026 · Full Gate from November 2026. Our public commitment for this product: Published prices, fixed scope, acceptance criteria written in numbers. The harness and the report are yours; we keep nothing but the invoice.
FAQ
Partly. The attack pack is adversarial and agent-specific, but the Gate also measures reliability (does the agent resolve the task?), cost (at what price per task?) and governance (who can it act as, what does it depend on?). A pen test answers one of those four questions.
No. The engines run inside your cloud account or VPC, on infrastructure you control. The report is generated there and handed over there. We work under NDA with named engineers and delete working copies within 30 days of a written request.
Most do the first time — that is the point of running it before go-live. The report comes with a fixed-price plan to close the gaps, and the sprint fee is credited against a Build signed within 30 days.
Yes. Partners can issue the Evidence Report under their own brand with Rainkernel as the named assessor, or unbranded. Ask about partner pricing.
Works with
A 30-minute call with the engineers who build it, no deck, no charge. If it fits, a fixed-price proposal within 48 hours.