How we measure
Every verdict on an Vorza audit traces back to a recorded run. This page explains how the numbers are made.
Real agents, exact versions
We run the actual clients — Claude Code and cursor-agent — not simulations or prompted imitations. The exact client version is recorded in every trace, so a result is always attributable to the software that produced it.
Isolated runs
Every run starts in a fresh, isolated sandbox. No shared state between runs, no warm caches, no leftover config: what the agent finds is exactly what a new user of your product would find. Wire logs rotate per run.
Every message recorded
A transparent proxy sits between the agent and your product and records every JSON-RPC message, CLI invocation, and HTTP request — the wire truth, not a summary. When a scenario fails, you can replay exactly what the agent saw and sent at the moment it broke.
Graded mechanically
Verdicts come from a closed set of mechanical assertions declared before the run: a file exists, output contains a value, a tool call succeeded. A scenario passes because its assertions passed — never because the output looked right to a model or a human skimming logs.
Repeated runs, honest statistics
Agents are non-deterministic, so a single run proves nothing. Every scenario runs repeatedly per client, and pass rates ship with 95% Wilson confidence intervals — the interval is the honest width of what we know, and it narrows as runs accumulate.
Failures that aren't yours don't count
When a run dies of client-side or infrastructure causes — the agent crashed, the sandbox lost networking — it leaves the denominator entirely instead of counting against your product. Credential problems are caught by a pre-flight gate before any agent runs, so a bad key never masquerades as a product failure.
Findings are drafted, evidence is primary
Failures become findings drafted by AI and labelled as exactly that. Every claim in a finding links to the trace it cites; the recording is the source of truth, the prose is the summary.
How severity is decided
Findings carry a five-level severity, prioritized by what it costs you when an agent hits it — task impact first, frequency second:
- critical — an agent cannot complete a core task at all, or the failure corrupts data or leaks something it shouldn't.
- high — a core task fails in most runs, or agents silently do the wrong thing (worse than an error, because nobody notices).
- medium — the task completes but unreliably (flaky), slowly, or only by working around your documented path.
- low — friction that costs turns or tokens without changing the outcome: confusing errors, redundant calls, unclear descriptions.
- info — an observation worth knowing, not a defect.
The results verdict compresses these to three buckets — critical/high → high, medium → moderate, low/info → minor — so the one-line summary stays readable. Frequency weights within a level: a critical that hits 1 run in 5 still outranks a medium that hits all 5, because you can retry flakiness but not data loss. Run an audit to see it on your own product →