Real Claude Code + Cursor runs
Real Claude Code and Cursor sessions run your MCP server, SDK, or CLI. You get measured pass rates, recorded sessions, and findings with fixes.
They install the wrong thing, misread a schema, claim success, and move on. No error reaches you.
They install the wrong thing, misread a schema, claim success, and move on. No error reaches you.
To find one, an engineer replays every session by hand, on an interface nobody at your company has ever looked at.
Real agents run your product; every failure comes back with the recording and the fix.
What you get
01 · Findings
The most expensive failures never throw an error. Vorza drafts findings for what agents quietly do wrong, evidence attached: click through to the exact wire message.
agents never call your search tool
search_records' description reads “exact match only”, so agents skip it and paginate the full dataset instead, on every session. Nothing errors. Nothing shows in your logs.
definition of fixed · agents reach for search_records first · ≤3 tool calls per session
02 · Run matrix
Agents are non-deterministic, so one run proves nothing. Every scenario runs repeatedly on every client; one matrix says what's solid, what's flaky, and where clients diverge.
| scenario | claude | cursor | verdict |
|---|---|---|---|
| install-and-discover | 5/557–100% | 5/557–100% | solid |
| machine-readable-output | 4/538–96% | 3/523–88% | flaky |
| init-scaffolds-config | 0/50–43% | 0/50–43% | broken · finding drafted |
| error-recovery | 5/557–100% | 2/512–77% | clients diverge |
03 · Fix loop
Claude Code or Cursor connects straight to your audit, pulls each finding's evidence and its definition of fixed, patches your code, then has us re-run the failing scenarios to prove it.
04 · Re-verify
Fixes are proven by fresh runs, not claimed. If a verified fix regresses later, the finding reopens with new evidence. Verified stays verified.
init reports success but writes nothing
If a verified fix regresses in a later release, the finding reopens. It never silently stays “fixed”.
Works with
Setup
01
Target type, source, docs URL. The test plan generates from your own docs; real agents run it repeatedly in fresh sandboxes.
target type
source
@yourorg/your-mcp-server
docs
https://docs.yourproduct.com
02
Findings come back where you work: connect Claude Code or Cursor to your audit over MCP and fix with the evidence in hand.
claude code
claude mcp add --transport http vorza-audit \ https://www.vorza.dev/api/mcp/audit \ --header "Authorization: Bearer <credential>"
cursor · ~/.cursor/mcp.json
{ "mcpServers": { "vorza-audit": {
"url": "https://www.vorza.dev/api/mcp/audit",
"headers": { "Authorization": "Bearer <credential>" }
} } }then: /fix_finding f-001
Free scan
Ten seconds, no signup: a scored pre-check of your public surface.
42 checks · every failing check ships a fix · predicted, until real agents run
Want more? Watch one real Claude Code session run against your site, free. →
Credentials
Sealed in your browser
Credentials are encrypted in your browser to the audit worker's public key. We can't read them back.
Decrypted once, at run time
Decryption happens on the isolated worker; values are injected into a fresh, disposable sandbox.
Scrubbed from evidence
Never logged, never included in prompts, and scrubbed from all artifacts before upload.