Get your agent red-teamed before your customers do
A free scan and an open playbook are the front door. When you need proof an enterprise review or an insurer will accept, we run the depth: static↔dynamic cross-validation, evidence-backed findings, a defensible grade.
- Agent vendors selling into enterprise / government — unblock the security review that’s holding your deal.
- On-device / box makers shipping an agent — catch what would become a recall or a brand hit.
- Teams self-hosting an agent internally — know your real exposure before an incident.
| Tier | Timeline | What | Price |
|---|---|---|---|
| T0Free scan | Self-serve, now | Surface scan (net/auth) + the open baseline + the static engine. Grade A–D on the spot. | Free |
| T1Red-Team Sprint | ~1 week, fixed price | One agent, one config. The full playbook (RT-1…RT-9), static + dynamic, exploit-validated. Findings backed by code path + working exploit. | Let’s talk |
| T2Deep Engagement | 3–6 weeks, custom | Multi-config, multi-channel. Chained exploit PoCs in a sandbox, retest after fixes, executive readout. | Let’s talk |
| T3Attestation | Retainer | Re-test every release; the grade becomes a live badge that feeds the certification / insurance rail. Recurring. | Let’s talk |
Pricing is scoped per engagement — tell us the agent and the scope. Early customers get design-partner terms in exchange for a case study.
An evidence-backed report — every finding carries the code path and a working exploit, with reproduction.
A defensible grade (A–D) with a critical-check veto — the kind a security review or an insurer can rely on.
Refuted false positives, stated out loud — we don’t cry wolf, and we’re fair to well-defended agents.
A prioritized fix list mapped to each confirmed break, and a free retest of what you fix.
A “red-teamed by” attestation on pass — proof you can hand a customer or a tender.
Scope & authorization
Written rules of engagement, a disposable environment, no production data. We test only what you authorize.
Static triage
We read the source to map the attack surface and flag candidate paths — where to attack.
Dynamic red-team
An AI adversary attacks the running agent to confirm what’s exploitable — the arbiter is the exploit.
Report & retest
Evidence-backed report + grade + fixes. We retest what you fix, and disclose responsibly.
Ready to break it before an attacker does?
Tell us the agent and the scope. We’ll scope a sprint, test in a disposable environment, and hand you evidence you can act on — and show a customer.
Get in touch