AEO
AEO
How retrievable and citable is your brand to agents? The deterministic check catalog, the honest crawler, source attribution, gaps, and the closure contract.
AEO measures how retrievable and citable a brand is to agents, computed from its public surfaces: the primary domain, the docs URL, llms.txt, the package registries, and the GitHub presence. A crawl runs the whole check set over every brand in a market (you and your competitors, same checks, same version), stores every fetch it depended on as evidence, and turns each failed check on your own brand into a fix thread on the standard loop.
The check set (v1)
Checks are deterministic and versioned. Any change to a rule, threshold, weight, or the scoring math bumps the version; results are forward-written and never recomputed, and scores compare only within a version. This table is generated from the same catalog the crawler runs, so what you audit here is what executed.
| check | group | subject | what passes | weight | severity |
|---|---|---|---|---|---|
| site-reachable | discovery | domain | The primary domain answers with a 200 within five redirects and the fetch budget. An unreachable site is invisible to every agent. | 4 | high |
| robots-allows-agents | discovery | domain | The robots.txt rules permit the named agent crawlers (ClaudeBot, GPTBot, PerplexityBot and peers) on the site root. An absent robots.txt passes. | 5 | high |
| llms-txt-present | discovery | domain | /llms.txt answers 200 with non-empty text within the size cap. It is the one file agents fetch first to orient on a product. | 4 | medium |
| llms-txt-structured | discovery | domain | The llms.txt parses as the convention, one H1 and at least one markdown link section, so agents can navigate it mechanically. | 2 | low |
| llms-txt-links-resolve | discovery | domain | The first sampled markdown links in llms.txt resolve. A file of dead links sends agents somewhere that does not exist. | 2 | low |
| llms-full-present | discovery | domain | /llms-full.txt answers 200 with non-empty text, letting an agent read the whole documentation in a single fetch. | 1 | info |
| docs-reachable | docs | docs | The registered docs URL answers with a 200 within the fetch budget. | 5 | high |
| docs-static-text | docs | docs | The raw docs HTML carries the content without JavaScript, enough visible text at a reasonable content-to-markup ratio. Agents do not run JavaScript; a client-rendered shell is empty to them. | 4 | medium |
| docs-agent-allowed | docs | docs | The docs origin's robots.txt rules permit the named agent crawlers on the docs path. An absent robots.txt passes. | 4 | high |
| docs-quickstart-linked | docs | docs | The docs page links a quickstart, getting-started, or install page and that link resolves, giving agents a one-hop path to the first working setup. | 2 | low |
| package-published | packaging | package | The registry metadata endpoint answers 200 JSON for every registered package identifier. | 5 | high |
| package-description | packaging | package | The registry metadata carries a non-empty description, the first line agents read when ranking install candidates. | 2 | low |
| package-repository-linked | packaging | package | The registry metadata names a repository URL and it resolves, connecting the package to its source and README surface. | 3 | medium |
| package-homepage-linked | packaging | package | The registry metadata names a homepage or docs URL and it resolves, giving agents the one-hop path from install candidate to documentation. | 2 | low |
| package-readme | packaging | package | The registry readme or long description carries enough content for an agent to understand the package without leaving the registry page. | 2 | low |
| github-repo-resolves | github | repo | The public repo page answers 200. Agents use the repo as a primary discovery and trust surface. | 3 | medium |
| github-readme-present | github | repo | The repository serves a non-empty README, the first document agents read about a codebase. | 2 | low |
| github-description-present | github | repo | The repository's description is non-empty, the one line shown everywhere the repo is listed or previewed. | 1 | info |
Statuses are pass, fail, or not measured. Not-measured results carry a closed reason code (no registered surface, robots disallowed, fetch error, rate limited, budget exhausted, dependency unmeasured, or check degraded) and always leave the denominator, visibly. An absent surface is never half credit and never a fail. When the same check cannot observe reliably across two or more brands in one crawl, it is marked degraded for the whole crawl: every one of its results reports not measured, it opens no fix threads, and the surface says so. Our parser going blind must never become your defect.
The score
Per group, the score is the weighted pass fraction over measured checks; the overall score is the same fraction over all measured checks, rendered 0 to 100 with one decimal. No grade letters, no bands, no warn tier. Every rendering shows how many checks passed, how many were measured, how many were not measured, the crawl age, and the check version.
The crawler
Crawls are honest, bounded, and evidence-preserving. VorzaBot identifies itself with its real user agent (documented at /bot), obeys robots.txt for its own fetches (a path your robots rules disallow for VorzaBot is not fetched, and the dependent checks report not measured with the robots reason), serializes fetches to one domain at one request per second or slower, caps response bodies at 2 MB, and stops at 40 fetches per brand and 300 per crawl. It fetches nothing beyond the registered surfaces, the two package registries, and GitHub. Every fetch a result depends on is stored, and every number on the surface walks down to those stored fetch documents.
Competitors are re-crawled at most once per freshness window (default seven days); your own brand is always due. Each brand's crawl age is always visible beside its results.
Source attribution
For every recommendation run in a share cycle, attribution derives which hosts the agent verifiably contacted (from the captured wire), which registries it verifiably installed from (from captured install commands with exit code zero), and which hosts it merely mentioned in its final answer. A mention is a claim and is counted as exactly that (cited only), never merged into retrieval. Attribution is hosts, not pages: HTTPS is observed at host granularity by design, with no TLS interception, ever. A run without capture gets no retrieval reason, honestly, and leaves every retrieval denominator as a visible n_na.
Gap analysis
A gap is a prompt and agent cell of your latest completed share cycle where your brand missed. Each gap carries a reason only where the stored evidence supports one:
| reason | assigned when |
|---|---|
| not_retrieved | capture active; no verifiable contact with your surfaces; your brand not named. Links your currently failing retrievability checks. |
| retrieved_not_named | the agent fetched your surfaces but never named you |
| named_not_chosen | named, but a competitor was chosen (a preference problem, not a retrieval one) |
| competitor_default | a competitor chosen with zero external web contact in the run: the agent acted from prior knowledge, and in-run retrieval could not have intervened |
| no_capture | the run carried no wire capture, so no reason is claimed |
| unattributed | the evidence supports none of the above |
Gaps are never stored: they recompute against the current check state, and every gap links its run's stored evidence.
Actions and the closure contract
Every failed check on your own brand becomes a finding (competitor results are comparison data, never findings) with a rule-drafted fix spec citing the stored fetch artifacts, threaded through the standard fix loop. The action list orders open threads by severity, gap pressure (how many gap cells implicate the check), and competitor differential (how many competitors ranked above you pass it). The three inputs render on every row; there is no composite score.
Run it
python3 -m vorza cycle --aeo mkt_... --watch # or on the dashboard: AEO -> your market -> Run crawl
Your coding agent works the actions over the fix loop MCP server: it pulls the finding, the grounded fix spec, and the stored crawl evidence, implements the change in your own repository, and requests re-verification. The re-crawl flips the thread only on the check's affirmative pass.