Audits
Plans, scenarios & assertions
How an evaluation plan is generated, what a scenario contains, and the assertion catalog.
Every audit runs a plan: a versioned document listing exactly which scenarios will run, with which agents, and what each one proves. You see the whole plan before anything runs, and the report you receive prints the same plan hash — preview and measurement can never silently diverge.
Where scenarios come from
- Packs: a standard scenario set for your target type — the tasks every MCP server / SDK / CLI should survive (connect, discover, complete the happy path, recover from errors).
- Your documentation: we read the docs you point us at and derive scenarios from what they promise. If the docs say an agent can do it, the plan verifies it.
- A live probe (MCP targets): we handshake with your server and list its tools, so scenarios name real tools rather than guesses.
Scenarios that can't be validated against the harness are rejected and listed as such in the plan — you see what we chose not to evaluate, too. A coverage table maps plan dimensions to the scenarios that cover them, with honest missing / not applicable entries.
What a scenario contains
The prompt the agent gets, what the scenario proves, which clients run it, its sandbox fixtures, time and turn budgets, repetitions, and — the part verdicts are built from — typed assertions checked against the recorded wire log and sandbox afterwards.
The assertion catalog
Assertion types are a closed, validated set: every verdict traces to one of these, each evaluated mechanically against captured evidence (never against vibes).
| assertion | checks that |
|---|---|
tool_called | the wire log shows the named MCP tool being called |
tool_called_with | a call to the named tool carried the expected arguments |
tool_not_called | no call to the named tool appears (the bypass class, inverted) |
tools_listed | the client listed tools; optionally that specific tools were present |
output_contains | the agent's final answer contains the expected text |
max_turns | the client finished within a turn budget |
file_exists | a file exists in the run's sandbox afterwards |
file_contains | a sandbox file contains the expected text |
file_not_contains | a sandbox file does NOT contain the text (guards derived-view corruption) |
command_succeeds | a command exits 0 in the sandbox after the run |
request_received | the capture server logged a matching HTTP request |
cli_called_with | the CLI was invoked with the given command and arguments |
stderr_matches | a CLI invocation's stderr matches a pattern |
exit_code | a CLI invocation exited with the expected code |
event_order | one recorded event precedes another (setup before use, list before call) |
no_human_intervention | the run completed without ending on a question back to the user |
The customer-readable plan ships as plan.json alongside your results, and the exact plan is always visible on the results page under “what we evaluated” (results).