Work with us · red-team engagement

Get your agent red-teamed before your customers do

The open red-team skills and playbook are the front door — run them yourself. When you need proof an enterprise review or an insurer will accept, we run the depth: static↔dynamic cross-validation, evidence-backed findings, a named public advisory.

Who this is for
  • Agent vendors selling into enterprise / government — unblock the security review that’s holding your deal.
  • On-device / box makers shipping an agent — catch what would become a recall or a brand hit.
  • Teams self-hosting an agent internally — know your real exposure before an incident.
Engagements
TierTimelineWhatPrice
T0Open skillsSelf-serve, nowThe open red-team skills + the static engine. Load them into an agent you own and run the same tests behind our advisories — yourself.Free · Apache-2.0
T2Deep Engagement3–6 weeks, customMulti-config, multi-channel. Chained exploit PoCs in a sandbox, retest after fixes, executive readout.Let’s talk
T3AttestationRetainerRe-test every release; the grade becomes a live badge that feeds the certification / insurance rail. Recurring.Let’s talk

Pricing is scoped per engagement — tell us the agent and the scope. Early customers get design-partner terms in exchange for a case study.

What you get

An evidence-backed report — every finding carries the code path and a working exploit, with reproduction.

A named, publicly-disclosed advisory with a measured hit-rate — the kind a security review or an insurer can rely on.

Refuted false positives, stated out loud — we don’t cry wolf, and we’re fair to well-defended agents.

A prioritized fix list mapped to each confirmed break, and a free retest of what you fix.

A “red-teamed by” attestation on pass — proof you can hand a customer or a tender.

How an engagement runs
01

Scope & authorization

Written rules of engagement, a disposable environment, no production data. We test only what you authorize.

02

Static triage

We read the source to map the attack surface and flag candidate paths — where to attack.

03

Dynamic red-team

An AI adversary attacks the running agent to confirm what’s exploitable — the arbiter is the exploit.

04

Report & retest

Evidence-backed report + grade + fixes. We retest what you fix, and disclose responsibly.

Ready to break it before an attacker does?

Tell us the agent and the scope. We’ll scope a sprint, test in a disposable environment, and hand you evidence you can act on — and show a customer.

Get in touch
WeChatOr add me on WeChat
WeChat QR — scan to add me as a friend