# Vorza — full documentation Vorza grades how ready a website is for AI agents, and audits products with real agent runs. ## The scan and the audit Vorza offers two different things and refuses to blur them. The free scan is a static pre-check: 42 checks against the surface a site presents to an agent, scored in seconds — evidence about the surface, not a measurement of behaviour. The paid audit (https://www.vorza.dev/audit) measures behaviour: real coding agents (Claude Code, cursor-agent) run repeatedly against a product — its MCP server, SDK, or CLI — with measured success rates and 95% Wilson confidence intervals, and downloadable evidence for every run. ## The four layers Checks are grouped into four layers, weighted 20/30/40/10. Checks that do not apply to a site are excluded from the denominator entirely. ### Discovery (20 points) Can an agent find you and your documentation at all? Homepage title and meta description quality, robots.txt existence and whether it blocks agent crawlers (ClaudeBot, GPTBot, PerplexityBot, Google-Extended and friends), sitemap presence and validity, docs and API links findable from the homepage. ### Accessibility (30 points) Once an agent asks for a page, does it get usable content? Whether the homepage answers agent user-agents with the same status it gives browsers, how much readable text is served without JavaScript, whether llms.txt exists, parses as markdown and its links resolve, llms-full.txt, JSON-LD structured data, markdown content negotiation, and whether docs render without a JS runtime. ### Usability (40 points) Can an agent actually operate the product? A published and valid OpenAPI document with declared auth schemes and typed error responses, machine-readable errors on nonexistent paths, agents.md, and — where advertised — a working MCP server: initialize handshake, listed tools, complete input schemas. ### Payments (10 points) Can an agent transact? Published x402 / AP2 / ACP markers and 402 challenge behaviour. Nearly every site scores n/a here today; the layer can only raise a score, never sink one by default. ## Scoring rules - A passing check earns its full weight; a warning earns half; a failure earns nothing. - N/A checks leave the denominator — they neither punish nor credit. - Layer scores renormalise over the layers that had anything applicable. - Grades: A+ >= 95, A >= 86, B >= 70, C >= 48, D >= 28, F below. - An unreachable site scores 0/F, with a note that this is a scan failure. - Every failing or warning check ships a concrete, standalone fix. ## The API - POST https://www.vorza.dev/api/scan — body {"url": "https://example.com"} — run a pre-check (~10s, rate limit 10/min/IP, no auth) - GET https://www.vorza.dev/api/score/{domain} — stored result; 404 replies with code DOMAIN_NOT_SCANNED and a next_action containing the exact POST that triggers the scan - GET https://www.vorza.dev/api/badge/{domain} — SVG badge with the scan score and grade - POST https://www.vorza.dev/api/waitlist — body {"email", "source", "detail?"} — reach the team (sources: contact, audit-request with detail.audit_id) or join the journey/monitor waitlist - GET https://www.vorza.dev/api/openapi.json — the full OpenAPI 3.1 contract All errors are JSON: {"code", "message", "next_action?"}. Branch on code; next_action is a request that can be made verbatim. ## What the audit adds The pre-check grades a surface. The audit measures outcomes: real agent sessions in isolated sandboxes — every JSON-RPC message recorded, every run independently asserted — reported as per-scenario, per-client success rates with downloadable evidence. It routinely finds what no static scan can: tasks that succeed in one agent and fail in another, servers whose tools list cleanly but reject every call, error responses that send agents into retry loops. Start at https://www.vorza.dev/audit. Audit facts: every scenario × client cell runs 5 times; rates carry 95% Wilson intervals; cells that both pass and fail are labeled flaky, never averaged away. Findings are AI-drafted from recorded evidence, human-reviewed before customers see them, always labeled AI-drafted, and always evidence-linked. Credentials are sealed in the customer's browser to the audit worker's key (the web app cannot read them), verified by a pre-flight gate before any agent run is spent, scrubbed from every artifact in every encoding, and purged at completion. One free re-run of failing scenarios is included. ## The fix loop (findings to verified fixes) Every published finding in a workspace becomes a durable fix thread with a closed lifecycle: open, acknowledged, fix submitted, re-verifying, verified fixed, regressed, won't fix. verified_fixed and regressed are SYSTEM-ONLY resolutions of ingested re-run wire evidence; no human or agent input can set them. Threads carry grounded fix specs (the exact change required, every change citing stored run evidence through a verbatim quote gate; suppressed entirely when ungrounded), and acceptance criteria derive at read time from what re-verification actually runs. Re-verification is a standard cycle of the bound source (an identical eval cycle, or the audit's affected scenarios); "verified fixed" means the failure did not reproduce across that cycle's n fresh verified runs, with per-client k/n on the event. A recurrence in any later completed cycle flips a verified thread to regressed automatically. Every transition is an append-only, actor-attributed event. The org-scoped MCP server at https://www.vorza.dev/api/mcp/fix-loop (streamable HTTP; bearer: an org API token vrz_…; initialize and tools/list work unauthenticated) serves the loop to the customer's coding agent: tools list_findings, get_finding, get_fix_spec, get_evidence, list_agent_readiness_specs, get_agent_readiness_spec, submit_fix_note, request_reverification, get_reverification_status, plus the fix_finding prompt. The boundary, stated in every tool description: Vorza never reads or writes customer code or repositories; findings, evidence, and fix specs flow out; plain-text fix notes and re-verification requests flow in; nothing sent there reaches any system outside the customer's Vorza workspace. A verified fix also leaves an org-private changelog draft. ## The Fix Loop (audit-scoped MCP) Every completed audit additionally exposes a customer-scoped MCP server at https://www.vorza.dev/api/mcp/audit (streamable HTTP; bearer credential from the audit's results page; initialize and tools/list work unauthenticated). The customer's coding agent pulls each published finding — evidence, repro steps, suggested fix, and acceptance criteria (the failing scenarios' prompts + assertions) — fixes the code in the customer's own repo, then calls request_fix_verification: Vorza re-runs exactly the affected failing scenarios with real agents and reports the new pass rate. Tools: get_audit_status, list_findings, get_finding, get_scenario, get_evidence, get_report, get_fix_brief, request_fix_verification, get_verification_status, plus the fix_finding prompt. A public no-auth MCP server also exists at https://www.vorza.dev/api/mcp (scan_domain, get_score, get_leaderboard, get_skill, create_audit_plan). ## Agent-readiness specs (the reference library) Versioned reference content on building products agents can use, served at https://www.vorza.dev/docs/agent-readiness and over the fix-loop MCP server: MCP server design, in-band error remediation, agent-readable docs, one-shot quickstarts, self-correcting error messages, agent-completable auth, demo mode, and CLI surface design. ## Scenarios (framework starting environments) Scenarios measure how coding agents build with a framework: one task run repeatedly by real Claude Code, Codex, and Cursor sessions inside a sandboxed starting environment (an empty project, the framework's scaffold, or a real repo template), evaluated against closed-set assertions over captured evidence. Tier 1 frameworks: LangGraph, Mastra, CrewAI. The framework-by-agent matrix aggregates each scenario's latest completed cycle as passed/measured with a Wilson interval; verdict colors appear only on verified runs, and every colored cell links to stored trace evidence. Preview sandboxes report the exact tree a run starts from before any agent spend. CLI: python3 -m vorza sandbox (catalog + local/hosted previews) and python3 -m vorza cycle (trigger and watch a cycle) with org API tokens from https://www.vorza.dev/dashboard/settings. ## Agent preference (markets, brands, share cycles) Preference measures which brands real coding agents name, choose, and verifiably install when given realistic developer pain prompts. A market holds a brand registry (the customer's brand plus competitors, each with aliases, domains, and package identifiers) and a customer-extendable prompt library; built-in market templates ship with seed prompt packs. A share cycle is survey mode by definition: 8 to 40 distinct prompts, one run per agent per prompt, Wilson intervals computed across prompts. Detection is deterministic code over the stored trace and sandbox snapshot (no language model in the counts): named requires a quoted trace span; installed requires a declared dependency in an agent-written manifest or a captured package-manager install that exited 0 (lockfiles corroborate, never prove); chosen follows a precedence ladder (single install, then latest install, then sole final-message recommendation, ambiguity means nothing is chosen); an install claim without evidence is counted as exactly that, never as an install. The board ranks brands by chosen share with intervals, a reporting floor of 8 measured prompts below which no number renders, ties on overlapping intervals, and trend only across cycles of one detector version. Matchups run the sweep with a forced-choice preamble and record the closed three-way outcome A / B / neither with third-brand steals counted (named flags there measure the framing, not discovery). A share is a measurement, never a verdict: preference numbers render in neutral ink, and every number walks down to stored evidence. CLI: python3 -m vorza cycle --market mkt_... [--matchup mch_...] --watch. ## AEO (retrievability checks, source attribution, gaps, actions) AEO measures how retrievable and citable a brand is to agents, from its public surfaces: the primary domain, the docs URL, llms.txt, the npm and PyPI registries, and the GitHub presence. A crawl runs a deterministic, versioned check set (18 checks in 4 groups at AEO version 1: discovery, docs, packaging, github) over every brand in a market with an honest bot identity (VorzaBot, documented at https://www.vorza.dev/bot), obeying robots.txt, pacing to at most one request per second per domain, and storing every fetch a result depends on as evidence. The score is the weighted pass fraction over measured checks, rendered 0 to 100 with the check counts, crawl age, and version always visible: a measurement in neutral ink, never a verdict, with no grade letters. Not-measured results carry a closed reason code and always leave the denominator; a check that cannot observe reliably across brands is marked degraded for the crawl and opens no actions. Source attribution derives, for every recommendation run, which hosts the agent verifiably contacted and which registries it verifiably installed from (wire evidence only, host granularity); a host merely mentioned in the final answer is counted as cited only, never as retrieval, and a run without capture is stamped honestly. Gap analysis annotates each missed prompt-and-agent cell with an evidence-backed reason (not_retrieved, retrieved_not_named, named_not_chosen, competitor_default, no_capture, or unattributed). Failed checks on the owned brand become fix threads on the standard loop with rule-drafted fix specs citing the stored fetch artifacts; an owned-scope re-crawl closes an action only when the bound check passes affirmatively, and market impact is tracked separately and never gates closure. CLI: python3 -m vorza cycle --aeo mkt_... --watch. ## Agent experience (evals, judged runs, the funnel, regression watch) Experience measures whether coding agents can actually integrate a product: install it, configure it, and get the first event onto the wire. An eval is a named, recurring integration measurement on one of four paths (sdk-setup, feedback-capture, docs-guided-setup, framework-integration): one scenario run a fixed 1..13 times per agent (deep mode; never a share metric), scored separately in Claude Code, Codex, and Cursor with wire-trace evidence on every run. Two verdicts, one color rule: the deterministic assertion aggregate (k/n with a Wilson interval) is the only chromatic surface; the LLM judge's per-criterion pass/fail is citation-grounded (a quote that does not appear verbatim in the cited event or file voids the citation; zero surviving citations means unjudged and out of every count), always labeled AI-judged, always neutral ink, and never touches the run verdict. Installed-vs-claimed reuses the preference detector: installed needs a manifest dependency or an exit-0 captured install; a claim without evidence is counted as exactly that. The onboarding funnel is deterministic named rules (discover, install, auth, configure, first_event) with a stall step and time-to-first-success (seconds from first captured activity to the success signal, median per agent); a step that could not be measured renders as not measured, never stalled. Failures classify into an AX issues feed (six groups, AI-drafted, evidence-linked). Target credentials are sealed to the worker key, verified by a pre-flight gate before any agent spend (credential problems complete with zero runs), held standing until rotated or deleted, with the rotation age visible on every eval page. Regression watch: daily/weekly cadences re-run the eval and alert exactly once when a verified assertion drop lands between comparable cycles; judged rates and measurements never alert; a spec change is a baseline change, not a regression. CLI: python3 -m vorza cycle --eval eval_... --watch. ## Published verified scorecards Customers can opt a subject into a public page: https://www.vorza.dev/verified/{slug} is a brand's verified scorecard (per-agent verdicts with intervals, preference share, AEO readiness, a verified-fix changelog, and the run descriptors behind every number), https://www.vorza.dev/verified/{slug}/{framework} the per-framework page, and https://www.vorza.dev/compare/{slug} a published head-to-head. Each serves JSON at https://www.vorza.dev/api/public/verified/{slug} and markdown via Accept: text/markdown. ## Monitoring Included in Vorza plans: the completed audit's exact plan re-runs on a schedule (daily / weekly / biweekly / monthly / on client releases) at reduced volume, emailing only when something that passed regresses. Threshold, recipient, pause/resume/cancel, and an on-demand run are customer-controlled. ## How to buy Audits and monitoring are set up directly with the team: generate the free plan preview at https://www.vorza.dev/audit and request it in one click, or start at https://www.vorza.dev/contact. Pricing is agreed per product — a reply comes within a business day. ## The full documentation map - https://www.vorza.dev/docs — overview - https://www.vorza.dev/docs/quickstart — scan to verified fix in five steps - https://www.vorza.dev/docs/scan — how the scan works - https://www.vorza.dev/docs/checks — every check, by layer, with weights and fixes - https://www.vorza.dev/docs/audits — the audit lifecycle - https://www.vorza.dev/docs/plans — plans, scenarios, and the assertion catalog - https://www.vorza.dev/docs/credentials — credential handling and scrubbing - https://www.vorza.dev/docs/results — reading results, findings, evidence - https://www.vorza.dev/docs/scenarios — framework environments, cycles, the verified matrix - https://www.vorza.dev/docs/preference — markets, brand detection, share cycles, matchups - https://www.vorza.dev/docs/aeo — the AEO check catalog, the crawler contract, source attribution, gaps, closure - https://www.vorza.dev/bot — the VorzaBot crawler identity, limits, and opt-out - https://www.vorza.dev/docs/experience — evals, judged runs, the funnel, regression watch - https://www.vorza.dev/docs/fix-loop — the fix loop: threads, specs, re-verification - https://www.vorza.dev/docs/agent-readiness — the agent-readiness spec library - https://www.vorza.dev/docs/monitoring — monitoring - https://www.vorza.dev/docs/mcp — both MCP servers - https://www.vorza.dev/docs/api — the HTTP API tour - https://www.vorza.dev/docs/agent-surfaces — machine-readable surfaces