Self-serve audit — no calls, no scheduling
Real agents. Your codesandbox.io. Measured.
Point us at your MCP server, SDK, or CLI. We generate a test plan from a standard pack plus your own docs, then run it with real Claude Code and Cursor sessions — repeatedly — and hand you pass rates with 95% confidence intervals, every trace, and findings drafted from the evidence.
A plan you can read
Standard scenario packs for your target type, plus 3–6 scenarios derived from your documentation. Every scenario says exactly what a pass proves — before you pay.
Statistically honest
Repeated runs per scenario per agent. Wilson confidence intervals, solid/flaky/broken labels, and client-infrastructure failures excluded from denominators instead of counted against you.
Evidence, not vibes
Watch runs live. Download every trace, wire log, and timeline. Findings are AI-drafted from failing runs and labeled as such — each claim links to the artifact it cites.
what a result looks like
Questions first? audits@agentlens.dev