Rainkernel

Rainkernel Lab

Public work anyone can check. That is how a small lab earns a top place.

A top place for an independent lab comes from public work anyone can check, not from hours or claims. The Lab is where Rainkernel does that work in the open: one checkable item every week, security findings reported under our name, benchmarks with published methods, participation in the standards our reports map to, and an evidence pack for every domain we test in.

Where a top place is still open

Six places with no clear leader. Four are products we already build.

From a 3 October 2026 review of about 35 AI testing and assurance providers. The certificate for AI vendors, ISO/IEC 42001 certification and red-teaming for the model makers are taken. These are not.

  1. 01

    A gate for one agent, with a published scope and a published price

    The certificate for AI vendors (AIUC-1), ISO/IEC 42001 certification and red-teaming for the model makers are taken. None of the providers we reviewed publishes a fixed price for testing one deployed agent before launch. The Agent Readiness Gate does: $7,500, two weeks, scope on the page.

    Open — we hold it from launch · Agent Readiness Gate →

  2. 02

    A neutral assessment of agent tools with public, named results

    MCP scanners are free and vendor-backed, but every one scans point-in-time and none publishes a neutral, named assessment of the servers and skills agents actually load. MCPSentry's public findings log and continuous attestation are built for that place.

    Open — scanner Q1 2027, first public findings from the Readiness Kit corpus in October · MCPSentry →

  3. 03

    Cost per finished task as test evidence

    Of the roughly 35 testing and assurance providers reviewed, one names runaway cost as something it tests. The Gate measures cost per completed task on real traffic and the Agent Cost Governor enforces the ceiling — a reliability result, not a FinOps afterthought.

    Open — measured in every Gate; enforced from December 2026 · Agent Cost Governor →

  4. 04

    One test-and-evidence pack per domain

    Rules push buyers hardest in insurance, banking, hiring, manufacturing, pharma and the public sector, and each asks for different evidence. The same test engine with a pack per domain — the questions, the attack cases and the records that regulator expects — is a position no generalist tool holds.

    Open — packs published domain by domain from the first sprint in each · Compliance Evidence Pack →

  5. 05

    Testing agents in languages other than English

    Public agent benchmarks are overwhelmingly English — the multilingual OmnilingualGAIA2 benchmark only appeared in August 2026 — and agents serving customers in Hindi, Telugu, Tamil, Spanish, German or Japanese are shipped on English test results.

    Open — multilingual test set on the Readiness Kit roadmap, Q1 2027

  6. 06

    A public benchmark in a domain that has none

    Support, coding and browsing have public agent benchmarks; claims triage, prior authorisation, invoice exceptions and customs documents do not. The first public, reproducible benchmark in one of these domains sets the reference others are measured against.

    Open — first candidate: a public support-ticket set with the Readiness Kit v0.1, 18 October 2026

How we get there

Four routes that cost nothing but work, and leave a public record.

Security advisories with our name on them

Flaws found in public agent tools, MCP servers and skills are reported through the maintainers' security advisories and, where they qualify, credited CVEs. The advisory log below lists every one we publish, with the maintainer's acknowledgement.

Proof
Advisory IDs and CVE credits, verifiable on GitHub · reference
When
First advisory expected October 2026, from the Readiness Kit corpus work

OWASP and MCP community work

Our attack pack is mapped to the OWASP Top 10 for Agentic Applications; we contribute cases and findings back to the OWASP GenAI project and take part in the Model Context Protocol registry and security discussions, where the signing and provenance gaps our products address are being worked out.

Proof
Contributions and discussions under our name in public repositories · reference
When
From October 2026

Singapore's AI assurance sandbox and tester accreditation

The AI Verify Foundation runs a Global AI Assurance Sandbox pairing testers with deployers, and is establishing an accreditation programme for AI testing firms — the first government-backed route to being a recognised tester. Rainkernel is applying to both.

Proof
Published sandbox case and, when the programme opens, accreditation status · reference
When
Applications in October 2026

One checkable item a week

Every week the Lab publishes one thing a buyer can verify without talking to us: a benchmark run, a finding, a sample report section, a corpus release, a reproduction of a published agent failure. The log is the record.

Proof
The weekly log below, dated
When
Every week from 5 October 2026

How the lab's time is split

  • Building Product 1 — the Agent Readiness Gate corpus, harness, report generator and the first paid sprints
  • Publishing one checkable item a week: findings, benchmarks, sample reports, corpus releases
  • Talking to buyers: five conversations a week with heads of engineering, AI and security, and with integrators

Weekly log

One checkable item a week.

Dated, with a link once it exists. A planned item that slips stays on the list as slipped; nothing is back-dated.

Weekly published items
WeekItemStatus
3 Oct 2026Sample Evidence Report published (fictional claims-triage agent) — the format every Gate deliverspublished
12–18 Oct 2026Readiness Kit v0.1 on GitHub: attack pack mapped to OWASP ASI01–ASI10, harness seeding, cost baseline, BOM emitter; benchmark run on a public support-ticket setplanned
19–25 Oct 2026First teardown: a published agent failure reproduced with the corpus, with the control that catches itplanned
26 Oct–1 Nov 2026Permission-audit checklist published; dry run of the Full Gate on a synthetic multi-agent estateplanned

Advisories

Security findings published under our name.

Each entry is added only after the maintainer has acknowledged it. The list starts empty because we have published none yet; the first is expected in October 2026 from the Readiness Kit corpus work.

No advisories published yet. Reports go to maintainers first; the entry appears here when they acknowledge it, with the advisory ID and, where issued, the CVE.

Found something in one of our tools? security@rainkernel.com · security.txt. We credit reporters.

Domain packs

One test engine. One evidence pack per domain, in the order the rules push.

The engine — attack pack, harness, cost meter, permission audit, BOM, recorder — is the same everywhere. What changes is the regulator's question, the attack cases that matter and the record that has to exist. Each pack is published after the first Gate in that domain.

Rules push hardest

Insurance and health insurance

What makes buyers prove it. State AI bulletins adopted across most US states; Colorado's rules on unfair discrimination in insurance; CMS-0057-F prior-authorisation APIs for US payers from January 2027; IRDAI outsourcing and governance expectations in India; EU AI Act high-risk pricing and claims categories.

  • Claims-triage and FNOL agents: decision records, adverse-action explanations
  • Prior-authorisation agents: human review trail, turnaround evidence
  • Underwriting copilots: disparate-impact probes on protected classes

Pack output. Decision record with explanation and human-review trail; bias-probe results; model and prompt version inventory

Rules push hard

Banking and fintech

What makes buyers prove it. Model-risk management expectations (SR 11-7 in the US, RBI's FREE-AI framework in India); EU AI Act high-risk credit scoring; AML and KYC record-keeping; consumer-protection rules on automated decisions.

  • KYC and onboarding document agents: injection through uploaded documents
  • Dispute and collections agents: policy adherence, escalation
  • Credit copilots: explanation quality, drift

Pack output. Model-risk validation pack: evaluation set, challenger results, monitoring plan, change records

Rules push hard

HR, hiring and workforce

What makes buyers prove it. New York City Local Law 144 bias audits; Illinois's AI-in-employment amendment from 1 January 2026; Colorado SB 26-189 from 1 January 2027; the EU AI Act's employment high-risk category (December 2027).

  • Screening and ranking agents: bias audit by protected class
  • Interview and assessment agents: explanation, candidate notice
  • HR help-desk agents: policy accuracy, PII handling

Pack output. Bias-audit report in the published format; notice and explanation records; evaluation set and results

Rules push hard

Manufacturing and industrial

What makes buyers prove it. The EU Machinery Regulation from January 2027 (AI in safety functions); the Cyber Resilience Act for products with digital elements (reporting from September 2026); ISO 9001 change control; functional-safety standards.

  • Maintenance and quality-report agents: action safety, authority limits
  • Supplier and procurement agents: document exceptions, cost per document
  • Product-embedded agents: BOM, vulnerability handling

Pack output. Agent BOM; change and incident records; authority-limit test results

Rules push hard

Pharma and life sciences

What makes buyers prove it. GxP computerised-system validation (EU Annex 11, 21 CFR Part 11); FDA and EMA guidance on AI in regulatory decision-making; pharmacovigilance reporting timelines.

  • Batch-record and deviation agents: validation evidence, audit trail
  • Pharmacovigilance intake agents: case completeness, timeliness
  • Medical-information agents: source grounding, off-label guardrails

Pack output. Validation pack (IQ/OQ/PQ-style evidence for the agent), tamper-evident audit trail, periodic review records

Rules push hard

Public sector and education

What makes buyers prove it. Public-procurement rules asking for inventories and impact assessments; records and freedom-of-information laws; the EU AI Act's public-service high-risk category; US federal and state AI-use policies.

  • Citizen-service agents: explainability on request, accessibility
  • Benefits and case agents: decision records, appeal trail
  • Admissions and grants assistants: bias probes

Pack output. AI inventory entry, impact-assessment inputs, decision records, accessibility results

Rules push hard

Healthcare providers

What makes buyers prove it. ONC HTI-1 transparency for decision-support software; state laws on generative AI in patient communications; HIPAA; EU MDR where software is a medical device.

  • Scheduling and intake agents: PHI boundaries, consent
  • Documentation and coding agents: accuracy on held-out charts
  • Patient-facing agents: disclosure, escalation to clinicians

Pack output. PHI data-flow map, evaluation results a clinical reviewer signs, disclosure records

Rules push moderately

Customer service, BPO and GCCs

What makes buyers prove it. EU AI Act transparency duties (Article 50, from August 2026); the FCC's ruling that AI voices in calls fall under the TCPA; client contracts' data-residency and audit clauses; SOC 1 and 2 for service organisations.

  • Support agents at peak: surge, rate limits, fallback
  • Voice agents: disclosure, consent, containment
  • Shared-mailbox triage: routing accuracy, PII

Pack output. Containment and resolution QA score, surge test report, cost per resolution

Rules push moderately

Software and SaaS

What makes buyers prove it. SOC 2 and ISO 27001 questionnaires with AI sections; the Cyber Resilience Act for products sold in the EU; EU AI Act GPAI and transparency duties; customers' procurement asking for an agent inventory.

  • In-product copilots: injection via user content, data boundaries
  • Coding and DevOps agents: credential scope, change control
  • Published MCP servers: description integrity, permissions

Pack output. Agent BOM for the security questionnaire; attack-pack results; cost per task

Rules push least

Supply chain and logistics

What makes buyers prove it. Customs and trade documentation rules; SOX controls around purchase-to-pay; contractual SLAs with carriers and customers — little AI-specific regulation.

  • Document-exception agents: accuracy and cost per document
  • Carrier-communication agents: commitment limits
  • Planning copilots: forecast evaluation

Pack output. Cost-per-document baseline, exception-handling evaluation, change records

Rules push least

Retail and e-commerce

What makes buyers prove it. Consumer-protection and advertising rules; PCI DSS where agents touch payments; DPDP and GDPR for customer data — little AI-specific regulation.

  • Returns and refund agents: action limits, fraud probes
  • Catalogue and search copilots: accuracy, brand-safety
  • Sale-day support agents: surge and fallback

Pack output. Action-limit test results, surge report, cost per resolution

Languages: every pack runs in English and in the languages the agent serves customers in; the multilingual test set is on the Readiness Kit roadmap for Q1 2027. Sector detail on the industries page; the evidence exports on the Compliance Evidence Pack page.

Have an agent we could test in public?

Design partners get the Gate at cost in exchange for a published, anonymised case. Integrators get the toolkit. Either way the result is a report anyone can check.