Preference
Agent preference
Do coding agents choose you? Markets, brand detection rules, share cycles, matchups, and the ranking board.
Preference measures which brands real coding agents name, choose, and verifiably install when given realistic developer pain prompts. A market holds your brand registry (you plus your competitors, each with aliases, domains, and package identifiers) and a prompt library. A share cycle runs every enabled prompt once per agent, detection runs over the stored evidence, and the ranking board shows chosen share with confidence intervals.
Run economics
Preference cycles are survey mode by definition: many distinct prompts, one run per agent per prompt, intervals computed across prompts. There is no repetitions setting anywhere on this surface, and four independent layers (the compiler, the enqueue re-check, the worker, and the harness validator) refuse repeated-run share math. A cycle needs at least 8 enabled prompts (a share from fewer cannot honestly rank anything) and runs at most 40.
The detection rules
Detection is deterministic code over two stored inputs per run, the canonical trace and the sandbox snapshot, plus the cycle's frozen brand snapshot. No language model is anywhere in the counts. Every true flag carries at least one evidence reference into stored artifacts, enforced when results are written. The current detector version is 1; rule changes bump it, and shares are comparable across cycles only within one version.
| flag | when it is true | its evidence |
|---|---|---|
| named | an identifier of the brand (name, alias, domain, or package id) appears on a word boundary in what the agent said | the quoted span from the stored trace, with surrounding context |
| installed | the brand's package is a declared dependency in a manifest the agent wrote or changed (package.json, requirements files, pyproject.toml, Pipfile, parsed structurally), or a captured package-manager install of it exited 0 | the manifest record in the sandbox snapshot, or the captured install invocation with its exit code |
| chosen | the brand the session committed to, decided by a precedence ladder (below); at most one brand per run | the install or recommendation evidence behind the winning rung |
| claimed only | the final message claims an install (installed, added, set up, integrated, configured, wired up, with a brand identifier in the same sentence) while no install evidence exists | the quoted claim sentence; counted as a claim, never as an install |
Lockfiles (package-lock.json, yarn.lock, pnpm-lock.yaml, uv.lock, poetry.lock) are corroborating only, never sole install evidence: they can be truncated in snapshots and they name transitive dependencies the agent never chose.
The chosen ladder
| rung | rule |
|---|---|
| 1 | exactly one registered brand has install evidence: it is chosen (installation is the strongest commitment an agent can make) |
| 2 | multiple brands installed: the latest successful captured install wins; with manifest-only installs, the installed brand also named in the final message wins; still ambiguous, nothing is chosen (reason: multiple installed) |
| 3 | no installs: exactly one registered brand named in the final message is chosen as a recommendation without setup |
| 4 | multiple brands named with no install: nothing is chosen (the agent compared without committing); no brand named: nothing is chosen |
A run enters denominators only when it measured: the session completed (no infrastructure error, no timeout, no unparseable client output) and its trace resolved to storage. Unmeasured runs are excluded and counted visibly as NA.
The board
The board ranks brands by chosen share from the latest completed share cycle, per agent and pooled overall, always with the interval and the measured denominator. A cell below the reporting floor of 8 measured prompts shows no number and contributes no rank; the board shows a rank order only when your brand and at least one competitor clear the floor. Adjacent brands whose intervals overlap are shown as statistical ties, because an overlapping interval is not a resolved order. Trend deltas render only between cycles that share a detector version and both clear the floor. Every number walks down to quoted spans, manifests, and captured commands through the evidence proxy.
Matchups
A matchup registers a head-to-head pair. A matchup cycle runs the same prompt sweep with a fixed forced-choice preamble naming both brands and asking the agent to pick exactly one and set it up. Outcomes are the closed three-way set: A, B, or neither, with third-brand steals counted separately; collapsing neither into a two-way record would misstate the measurement. One structural caveat: on matchup cycles both contestants are named by the framing itself, so the named flag there measures the framing, not discovery. Named share is computed from share cycles only.
Registry hygiene
Aliases shorter than 3 characters, or that are common English words, are refused at save (an alias like "run" would match ordinary prose in every trace). Packages agents installed that match no registered brand surface as suggestions with a one-click add as competitor. Brands with recorded observations archive instead of deleting, so history stays intact. Run explanations, the one AI-drafted surface here, are always labeled as such, must reference a stored artifact, and are suppressed entirely when a quote fails verbatim grounding; the counts never depend on them.
CLI and API
| command | what it does |
|---|---|
| python3 -m vorza cycle --market mkt_... --watch | trigger a share cycle and watch it finish |
| python3 -m vorza cycle --market mkt_... --matchup mch_... | run a matchup cycle |
Programmatic access uses org API tokens from Settings. The market endpoints live under /api/markets: registry, prompts, matchups, run, board, cycles, and the evidence proxy.