Audits
Results, findings & evidence
The verdict, the matrix, AI-drafted findings with evidence, and your free re-run.
Results are layered so the first thing you read is the conclusion and everything under it is the proof. Each layer cites the one below; nothing asks to be taken on faith.
The layers, top to bottom
| layer | what it tells you |
|---|---|
| Verdict | One plain sentence about what agents can and cannot do with your product, with severity counts. |
| Executive summary | One reviewed paragraph plus the 1–3 moves that turn the matrix green. AI-drafted, human-reviewed, evidence-cited. |
| The matrix | Every scenario × agent cell: pass rate over 5 runs with a 95% Wilson interval, labeled solid / flaky / broken. Click a cell for the runs behind it. |
| Findings | The diagnosed issues — each severity-ranked, evidence-linked, with repro steps and a suggested fix. |
| Per-run detail | Timelines and trace replay for every individual run; divergences between agents highlighted. |
| Evidence | Downloadable artifacts: traces, wire logs, timelines, the report, the plan. Every claim above resolves to something here. |
Findings — AI-drafted, human-reviewed, evidence-bound
Findings are drafted by a model from the recorded evidence, then reviewed by us before you see them: your page shows published findings plus an honest count of any still in review. Three properties hold for every finding:
- Labeled AI-drafted — always, on every surface that shows them.
- Evidence-linked — each finding cites specific events in specific runs; the evidence chips jump you into the trace replay at the exact moment.
- Honest about confidence — a finding without resolvable evidence is demoted to a hypothesis and displayed as one, never dressed up as a fact.
Flaky is a first-class verdict
A scenario that passes 3 of 5 runs is not a 60% success story — it's unreliable behaviour your users will hit. Cells that both pass and fail are labeled flaky and diagnosed as such; cross-agent divergences (works in Claude Code, fails in Cursor) are called out separately, because they usually indicate ambiguity in your interface rather than a bug in one client.
Deliverables
Alongside the page itself: report.md (the full written report), summary.json (machine-readable totals and cells), plan.json (exactly what ran), per-run traces and timelines, and the live event log. All downloads are credential-scrubbed before they reach storage (how).
The free re-run
Every audit includes one free re-run of its failing scenarios — after you've shipped fixes, run exactly the failed subset again from the results page, or let your coding agent do it through the Fix Loop. Beyond the free re-run, monitoring re-runs the plan on a schedule.