Experience
Agent experience
Can agents use you? Evals per integration path, judged runs, the onboarding funnel, time to first success, and regression watch.
Experience measures whether real coding agents can actually integrate your product: install it, configure it, and get the first event onto the wire. An eval is a named, recurring integration measurement: one scenario (persona, start condition, prompt, deterministic assertions), your product as the target (package identifiers, docs, an optional API base URL, credential names), judge criteria, funnel signals, and a cadence. Each cycle runs the same scenario a fixed number of times in Claude Code, Codex, and Cursor and scores every agent separately, with wire-trace evidence behind every number.
The four paths
| path | what it measures |
|---|---|
| sdk-setup | from an empty project to a working, initialized setup |
| feedback-capture | install, configure, and actually send one accepted event |
| docs-guided-setup | following your official documentation exactly as written |
| framework-integration | wiring your product into an existing framework project (LangGraph, Mastra, CrewAI) |
Built-in templates cover every path and clone against your target: the template placeholders (product name, package id, docs URL) expand from the target at creation time, and the stored prompt is the whole truth of what the agent was asked.
Run economics
Eval cycles are deep mode: one scenario, a fixed run count between 1 and 13 per agent, Wilson intervals across repetitions. Repeated runs here serve their designed purpose, pass rates and flakiness on one task, never a share metric. Three layers enforce it: the compiler has no prompt-set surface, the enqueue re-checks the compiled plan, and the worker refuses a survey-shaped plan outright. Cycle size is guarded so a full cycle fits its sandbox window: ceil(runs / 4) times the per-run timeout must stay at or under 1800 seconds.
Two kinds of verdicts, one rule about color
The assertion verdict is deterministic: the closed set of typed assertions evaluated over captured wire evidence, rolled up per agent as k of n with a Wilson interval. That aggregate is the only place chromatic verdict color appears.
The judged verdict is the first LLM-derived pass or fail in Vorza, and it follows the grounding discipline stated here verbatim: the judge receives only stored, scrubbed evidence (the trace with its event indices, the sandbox snapshot, the run's detection record, the eval's prompt and criteria). Every judged verdict must cite the evidence that decides it; a citation resolves only if the cited event or file exists and the quote appears verbatim in it. A criterion whose verdict has zero surviving citations is not rendered as a verdict: it becomes unjudged, with the suppression reason recorded, and leaves every count. Unjudged runs leave the judged denominator exactly as unmeasured runs leave the assertion denominator. Deterministic assertions always take precedence: the run verdict is assertion-derived only, and judged counts are a separate surface a judgment can never touch. Judged numbers render in neutral ink, always with the interval and n, always marked AI-judged, with the citations one click away. Judged trends compare only within one judge version, model, and criteria set; the surface footnotes the version when history spans more than one.
The funnel and time to first success
The onboarding funnel is deterministic, named rules over two stored inputs per run, the canonical trace and the sandbox snapshot, plus the eval's subject identity and signals. No language model is anywhere in these numbers. Every rule names its evidence source; a reached step without evidence cannot be written. The current funnel version is 1; rule changes bump it, observations are forward-written and never recomputed, and funnel trends compare only within a version. The steps, in canonical order:
| step | reached when | its evidence |
|---|---|---|
| discover | the agent acted on a subject package (an install command or a modified manifest naming it), contacted a subject domain on the wire, or named a subject identifier in its own text (path-like tokens excluded) | the trace event index or quoted span, or the manifest record |
| install | the Plan 02 installed rule, verbatim, via the shared detector: a subject package declared as a dependency in a modified manifest (lockfiles corroborate only), or a captured install of the package that exited 0 | the manifest artifact path and dependency, or the trace event with argv and exit code |
| auth | an http request carried a declared credential (the capture server's auth-present flag), or a modified non-seeded sandbox file references a required credential NAME. Names are not secrets; values never survive the scrub and are never searched for. With no required credentials and no http capture the step is not measured, never stalled | the request event index, or the sandbox file path and quoted line |
| configure | the eval's configure signal matches a modified sandbox file (path prefix plus substring), or by default a modified, non-manifest file contains a subject package identifier in an import or require position | the sandbox file path and quoted line |
| first_event | the eval's success signal fired: the first http request matching the signal with a 2xx status, the first matching CLI call that exited 0, or the first non-error tool call of the named tool. Evals with no success signal do not measure this step | the event index |
The stall step is the first step in canonical order that is measured but unreached, computed only for runs that did not reach their terminal step. A run that reached the first event has no stall; earlier unobserved steps render as not observed, because a measurement gap is not a failure. Unmeasured runs are excluded and counted. Time to first success is the timestamp of the first-event evidence event: seconds from the run's first captured activity (the trace's capture anchor, not process spawn) to the first success signal, null when the signal was unreached or unmeasured, aggregated as a median per eval and agent with n always visible. A duration is a measurement, not a verdict.
The AX issues feed
Failures classify into a closed category set, grouped into six AX groups at read time: task-completion, auth, docs-gap, error-handling, environment, and reliability. Deterministic rules classify failing runs from their captured errors; the judge's failed criteria produce task-incomplete findings built from the judge record itself, with its grounded citations as evidence. Every finding is labeled AI-drafted, carries at least one artifact reference, and an entry without evidence is dropped at the write rather than shown. Fix specs and the fix loop attach to this feed in a later plan.
Credentials: sealed, standing, visible age
An eval target holds the credential names its integration needs. Values are submitted once, sealed to the audit worker's public key, and stored as ciphertext we cannot read back. They are opened only at the worker boundary, injected into the run environment, scrubbed from every artifact before anything is uploaded, and no language model ever receives a value. Before any agent spend, a pre-flight gate verifies the credentials actually work; a credential problem completes as customer-actionable feedback with zero agent runs, never as a red cell.
Unlike audits, whose sealed credentials are purged at terminal states, eval targets keep the sealed blob: recurring cycles are the product. The retention answer is plain: a key is held, sealed, until you rotate or delete it from the target editor, and every eval page shows how long ago it was last rotated. Supply scoped or evaluation-tenant credentials where you can.
Regression watch
Each eval carries a cadence (manual, daily, or weekly). A scheduler enqueues due cycles through the standard guards and says so on the eval page when it had to skip. After a completed cycle, the watch compares per-agent assertion outcomes against the previous completed cycle of the same eval with the same compiled spec: a worsened label transition or a statistically significant rate drop at the same label records a regression and sends one alert email. Judged rates, durations, funnel numbers, and detection counts never alert; they move as trends only. A spec change is recorded as baseline changed, neutrally, and the new cycle becomes the baseline. Experiments (one-off comparative cycles with overridden clients, models, prompt, start condition, or run count) run on the same rail but are keyed so they can never enter headline numbers, trends, or the watch.
API and CLI
Everything on this page is reachable with an org API token: eval and target CRUD, credentials submit and rotate, run, cycles, the scorecard board, experiments, and the issues feed under /api/evals and /api/eval-targets. Trigger and watch a cycle from the terminal with python3 -m vorza cycle --eval eval_... --watch. Evidence downloads go through the cycle artifact routes, scoped to your org and the cycle's own storage prefix.
See also framework scenarios, agent preference, and credential handling.