Scanners score your public website. Vorza verifies your actual product. →

The audit

Real agents. Your buffer.com. Measured.

Point us at your MCP server, SDK, or CLI. We generate an evaluation plan from a standard pack plus your own docs, then run it with real Claude Code and Cursor sessions — repeatedly — and hand you pass rates with 95% confidence intervals, every trace, and findings drafted from the evidence.

npm package or streamable-HTTP URL

Credentials the tests will need (names only — values are entered later, sealed):

Free — no checkout; we set up the run with you.

A plan you can read

Standard scenario packs for your target type, plus 3–6 scenarios derived from your documentation. Every scenario says exactly what a pass proves — before anything runs.

Statistically honest

Repeated runs per scenario per agent. Wilson confidence intervals, solid/flaky/broken labels, and client-infrastructure failures excluded from denominators instead of counted against you.

Evidence, not vibes

Watch runs live. Download every trace, wire log, and timeline. Findings are AI-drafted from failing runs and labeled as such — each claim links to the artifact it cites.

Then keep it green: monitoring re-runs your plan on your schedule

Daily to monthly cadence or on client releases, alerts only on regressions, one click from any completed audit.

what a result looks like

vorza audit — your-mcp-serverillustrative output
connect + discover tools6/6
core tool: search_docs6/6
error recovery4/6
cursor recovers; claude retries the same bad call — divergence flagged
realistic mode: agent chooses your server1/6
agents used built-in fetch instead — your tool descriptions never mention the docs domain

Questions first? Talk to us.