Audits
How audits work
Real Claude Code and Cursor sessions against your product — the full lifecycle.
An audit measures what a scan never can: whether real coding agents actually succeed with your product. We run genuine Claude Code and Cursor sessions against your MCP server, SDK, or CLI — repeatedly, in isolated sandboxes, with every message on the wire recorded — and report per-scenario, per-client success rates with 95% Wilson confidence intervals and the evidence behind every run.
What we evaluate
| target | you provide | we run |
|---|---|---|
| MCP server | a streamable-HTTP URL (or package) | agents connecting over MCP, listing tools, completing documented tasks |
| SDK | an npm or PyPI package name | agents installing the package and building against its documented API |
| CLI | a package + binary name | agents driving the CLI end-to-end, every invocation captured |
The lifecycle
- Plan. We generate an evaluation plan — a standard scenario pack for your target type plus scenarios derived from your documentation — and probe MCP targets live to discover their tools. You review every scenario before anything runs (what a plan contains).
- Kickoff. Audits are set up with you directly — request yours from the plan preview (or talk to us) and we agree the scope and price together. The plan you previewed is pinned by hash to the plan that runs.
- Credentials — only if your target needs them. Sealed in your browser to the worker's public key; we cannot read them back (the full model).
- Pre-flight. Before any agent run is spent, a smoke gate verifies your credentials and target actually work. A failure here consumes nothing and tells you exactly what to fix — credential problems never masquerade as agent failures.
- Agents run. Each run gets a fresh, isolated sandbox — no shared state between runs, no reused sessions. Progress streams live to your results page.
- Results. Verdict, matrix, findings, per-run timelines, trace replay, evidence downloads — plus one free re-run of failing scenarios (reading the results).
How the numbers are honest
- Repetition: every scenario × client cell runs 5 times — a single run proves nothing about reliability.
- Wilson intervals: success rates carry 95% confidence intervals; small samples read as uncertain, because they are.
- Flaky is a verdict: a cell that both passes and fails is labeled flaky, not averaged into a misleading rate.
- Client versions recorded: every trace names the exact agent client and version that produced it.
- Infrastructure failures are ours: if our side breaks, the audit says so and doesn't bill the failure to your product.