Scenarios
Scenarios & the framework matrix
Framework starting environments, the scenario builder, cycles, and the verified framework-by-agent matrix.
Scenarios measure how coding agents build with a framework: the same task, run repeatedly by real Claude Code, Codex, and Cursor sessions inside a sandboxed starting environment, evaluated against a closed set of assertions over captured evidence. The result is the framework-by-agent matrix: pass rates with Wilson confidence intervals, colored only when verified runs back them.
The framework catalog
Each catalog framework ships a starting environment. Tier 1 today: LangGraph, Mastra, and CrewAI. A scenario picks one of three start conditions:
| start condition | the sandbox before the agent's first turn |
|---|---|
| empty | a fresh directory, nothing seeded |
| scaffold | the framework's own starter project (dependency-declared, not pre-installed) |
| template | a small real repo: a Django starter or a Next.js app |
Environments are dependency-declared, never pre-warmed: installation behavior is part of what a run measures.
The builder
A scenario is five fields plus assertions: a persona (folded verbatim into the prompt), a start condition, the task prompt, the agents (and optional model per agent; Cursor pins its own model so cross-agent cells compare agents, not models), and runs per cycle (1 to 13, default 5, the smallest n where the reliability label means something). Built-in templates are read-only clone sources; cloning one into your org makes every field editable. Saving compiles the scenario into the exact harness plan a cycle will run and stamps its hash.
Cycles
A cycle is one execution of one scenario: every agent runs the task n times in its own fresh sandbox, every session's activity is captured, and the closed-set assertions are evaluated against that captured evidence. The compiled plan is snapshotted per cycle, so editing a scenario never changes what a past cycle ran. Cycles carry no credentials by design.
The matrix
The matrix aggregates each scenario's latest completed cycle per framework and agent: passed over measured, the Wilson interval, and a reliability label (solid, flaky, broken). Unmeasurable runs leave the denominator.
Preview sandboxes
Before spending an agent run, provision a preview: an on-demand sandbox seeded with a start condition that reports the exact file tree every run begins from, then dies at its TTL. Available from the framework catalog page or the CLI.
CLI and API
| command | what it does |
|---|---|
| python3 -m vorza sandbox --list | print the framework catalog |
| python3 -m vorza sandbox --framework langgraph | materialize a start condition locally (no web dependency) |
| python3 -m vorza sandbox --framework langgraph --e2b | boot a hosted preview via the API (org token) |
| python3 -m vorza cycle --scenario scn_... --watch | trigger a cycle and watch it finish |
Programmatic access uses org API tokens, minted once under Settings. The scenario endpoints live under /api/scenarios; previews under /api/sandbox.