Fix & monitor
Agent-readiness specs
The reference library: MCP server design, in-band error remediation, agent-readable docs, one-shot quickstarts, and more.
The reference library behind our findings: how to build a product that coding agents can discover, evaluate, integrate, and recover with. Each entry is versioned catalog content, served here, over the fix-loop MCP server (list_agent_readiness_specs / get_agent_readiness_spec), and linked from every finding whose category it applies to.
- MCP server design: Tool naming and granularity, schemas as documentation, list-before-auth, pagination, and structured errors that keep an agent on your interface instead of around it.
- In-band error remediation: The fix travels inside the error: next_action shapes, recoverable versus terminal codes, and no out-of-band docs pointers.
- Agent-readable docs: llms.txt, one canonical path per task, verbatim-runnable snippets, and docs that agree with the shipped surface.
- One-shot quickstarts: A cold start to a working integration in one uninterrupted pass, with no hidden prerequisites and verifiable output at the end.
- Self-correcting error messages: Errors that name the exact wrong input and the exact correction, in a machine-parseable shape.
- Agent-completable auth: Auth an agent can finish without a human: env var contracts, token minting, and no interactive-only flows on the critical path.
- Demo mode: A zero-credential evaluation tenant: fake data, real wire shapes, clearly labeled, so agents can try the product before auth exists.
- CLI surface design: Non-interactive by default, stable exit codes, --json output, and parity with the API so every client succeeds the same way.
MCP server design
v1 · applies to: tool-desc-ambiguity, agent-bypasses-interface
An MCP server is an interface read by a model under time pressure. The agent sees your tool list once, picks a tool from its name and description, builds arguments from the schema, and acts. Every ambiguity at that moment costs you the call: the agent falls back to its built-ins, shells out, or guesses.
Tool naming and granularity
Name tools by the user-visible verb, not the product noun: `create_invoice`, not `acme_api_v2`. One tool per task the user would name; not one tool per REST route. A tool that multiplexes many operations behind an `action` argument forces the agent to learn a second interface inside the first.
Say in the description when the tool should be used and what it does NOT do. Neighboring tools compete for the same request; the description is where you win or lose the selection. Lead with the verb: "Create a draft invoice for a customer. Does not send it."
Schemas are documentation
The argument schema is read as prose. Give every argument a description with its expected format and one example value. Mark optionals honestly. An agent that constructs a wrong argument from an undocumented format will often retry the identical call, and each retry is your latency and their tokens.
List before auth
Let `initialize` and `tools/list` succeed without credentials. An agent that cannot see your tools cannot decide to use them, and the decision moment usually happens before auth is configured. Gate every `tools/call` instead, and return a structured error that names the exact credential and where to set it.
Pagination and result size
Cap list results and give the agent a cursor. A tool that returns an unbounded payload blows the context window and reads as broken. State the page size in the description so the agent plans for it.
Structured errors
Every error is a machine-readable object with a stable `code` and a human sentence that names the fix. Agents recover from errors they can parse; they loop on errors they cannot. See the in-band error remediation spec for the shapes.
In-band error remediation
v1 · applies to: agent-blocked, error-message-unactionable
When an agent hits an error, the only material it has is the error itself. A human opens a browser and searches; an agent re-reads the message. If the fix is not in the message, the fix does not happen.
The fix travels inside the error
Every error your surface returns should answer three questions in band: what was wrong, what exact change fixes it, and whether retrying can ever work. The strongest shape is a structured object:
{
"code": "invalid_currency",
"message": "currency 'usdollar' is not a supported code",
"next_action": "use an ISO 4217 code; for US dollars pass 'usd'",
"recoverable": true
}Recoverable versus terminal
Mark whether the same call can succeed after a correction. An agent that cannot tell a bad argument from a hard outage will either give up on a fixable call or hammer an unfixable one. `recoverable: true` plus a `next_action` produces one corrected retry; `recoverable: false` plus a reason produces a clean stop and an honest report.
No out-of-band pointers
"See the documentation" is a dead end in a headless session. If the remediation needs more than a sentence, inline the sentence that matters and keep the URL as a secondary reference, never the primary one. The same rule applies to interactive escapes: an error that says "run our setup wizard" blocks every agent that cannot answer a prompt. Name the non-interactive alternative first.
Precision beats politeness
Quote the offending input back verbatim. "Invalid value" forces the agent to diff its own call against the schema; "'2024-13-01' is not a valid date; expected YYYY-MM-DD with month 01-12" is a one-shot correction. The error message is the highest-leverage documentation surface you own, because it arrives exactly when it is needed.
Agent-readable docs
v2 · applies to: docs-faithful-path-fails, doc-server-mismatch, undocumented-dependency, aeo-discovery, aeo-docs
Agents follow documentation literally. They copy the snippet, run it, and treat the result as ground truth about your product. Docs drift that a human reader routes around becomes a hard failure in an agent session, and the failure gets attributed to your product, not to your docs.
Publish llms.txt
Put an `llms.txt` at your docs root: the canonical API host, the exact package name per registry, the auth env var names, and links to the few pages that matter for integration. This is the first file a cold-starting agent looks for, and it is where you disambiguate your name from every similarly named product an agent might find in a web search.
One canonical path per task
For each core task, keep exactly one documented way to do it. Two quickstarts that diverge produce agents that interleave them. Alternatives belong on clearly separated pages ("advanced", "migrating from X"), not inline as "or you can also".
Verbatim-runnable snippets
Every code block should run as pasted in a fresh environment. Placeholders are declared immediately above the block with the exact env var name. Hidden prerequisites are the classic failure: the snippet assumes a dependency, a config file, or a prior step that the page never states. Re-run your own quickstart from a clean machine on every release; agents do exactly that, cold, every day.
Docs and server parity
Everything the docs advertise must exist on the live surface under the same name: tools in your MCP tools/list, routes in your API, commands in your CLI. A documented tool the server does not expose sends every reader on a hunt that ends in guesswork. Wire a parity check into your release process: diff the documented names against the live listing and fail the release on a gap.
One-shot quickstarts
v1 · applies to: task-incomplete, undocumented-dependency, agent-stalls
The quickstart is the path an agent actually takes. It gets one session, a time budget, and no ability to come back tomorrow. Design the quickstart as a single uninterrupted pass from empty directory to verified working integration.
No hidden prerequisites
List every prerequisite at the top: runtime version, package manager, account state, credentials. Each one discovered mid-path costs a stall or a dead end. If a step needs a dependency, the install command for that dependency appears before the step, in the same page, copy-pasteable.
Every step produces observable progress
An agent that runs a step and sees nothing cannot tell success from a hang. Each step should end in output the agent can check: a printed confirmation, a file that now exists, an HTTP 200. Steps that silently succeed invite re-running and improvisation.
End in verification, not vibes
Close the quickstart with a step that proves the integration works: a request that returns real data, a snippet whose output is shown in the docs for comparison. This is what separates a genuinely complete integration from code that merely exists. Sessions that end with the agent claiming success are common; sessions that end with the claimed result verified on the wire are what your users experience as "it works".
Budget the path
Count the minutes and the steps. A quickstart that needs twenty minutes of wall clock will time out inside most agent budgets, and the stall reads as your product hanging. Move everything optional out of the critical path; the fastest honest path to a verified hello-world is the one to optimize.
Self-correcting error messages
v1 · applies to: error-message-unactionable
The signature failure this spec addresses: an agent makes a call, it fails, and the agent retries the identical call. That loop means your error named the problem but not the fix. The message was written for a human who would sigh and open the docs; the agent has only the message.
Name the exact wrong input
Echo the offending value back, verbatim, with the field it arrived in. The agent is diffing its call against your expectation; do the diff for it. "amount must be a positive integer of cents; received '10.50'" converts a retry loop into a one-shot fix.
Name the exact correction
State the valid alternative, the flag to change, or the precondition to satisfy, as literal text where possible. If only three values are legal, list all three in the message. If the fix is an account action, name the action and the non-interactive way to do it.
Machine-parseable first
A stable error `code`, the offending `field`, the received `value`, and a `next_action` string. Prose alone forces the agent to parse English; structure makes the correction mechanical. Keep codes stable across releases: agents and SDKs pattern-match on them, and a renamed code is a silent behavior change.
Write the message from the caller's side
"Internal validation failed in InvoiceService" describes your implementation. "currency is required when amount is set" describes the caller's next edit. Every error message should be reviewable by one standard: could a competent stranger, seeing only this message and their own call, fix the call without opening anything else?
Agent-completable auth
v1 · applies to: auth-setup-friction
Auth is where agent sessions go to die. A headless agent cannot click a consent screen, answer a device prompt, or copy a token out of a dashboard mid-session. If the only documented path to a working credential is interactive, every agent integration stalls at step one.
Publish the env var contract
One documented environment variable name per credential, stated in the quickstart before the first authenticated call: the exact name, the expected format, and where the value comes from. Agents look for this contract by convention; naming it removes a whole class of guessing. Accept the same credential via a flag for CLI surfaces.
Give agents a way to verify the credential
A whoami-style call that costs nothing and confirms the credential works. The agent runs it first, and an auth failure surfaces as a clean, diagnosable step instead of poisoning the first real request. Return the same structured error shape as everywhere else, naming the env var to fix.
Keep interactive flows off the critical path
OAuth consent, device codes, and browser dashboards can exist, but the quickstart's primary path must work with a pre-provisioned API key or token. Document the interactive flow as the way a human provisions the credential once, and the env var as the way every session uses it after that.
Scope tokens for delegation
Users hand agents credentials. Make that safe to do: scoped keys, short-lived tokens, a revocation surface, and clear naming so a leaked agent key is a contained event. "Create a restricted key for automation" in your dashboard is an agent-readiness feature, not just a security one.
Demo mode
v1 · applies to: auth-setup-friction, agent-blocked
The highest-friction moment in any integration is before credentials exist. A developer (or their agent) wants to know what your product does on the wire before creating an account. If the answer requires a signup, a verification email, and a dashboard visit, the evaluation happens on a competitor with a lower wall, or it happens as guesswork.
Real shapes, fake data
A demo mode serves the real API surface with synthetic data: the same routes, the same schemas, the same error shapes, deterministic fake records. The agent integrates against reality; only the data is staged. A demo that simplifies the wire format sets up an integration that breaks on the first real credential.
Zero-credential entry
A documented public demo key, or an unauthenticated demo base URL. State it in the quickstart and in llms.txt. The demo credential should be rate-limited and unable to touch anything real, which makes it safe to publish verbatim where agents will find it.
Clearly labeled
Every demo response carries an unmistakable marker (a header, a `demo: true` field, ids with a demo prefix) so no session can mistake staged data for production state, and so support can spot a user who shipped against the demo tenant. The label protects you and the user equally.
The graduation path
One documented step from demo to real: swap the env var, nothing else changes. If graduating requires code changes beyond the credential, the demo taught the wrong integration. The demo tenant is also where your own release checks can run cold-start flows without burning real accounts.
CLI surface design
v1 · applies to: agent-bypasses-interface, cross-client-divergence
A CLI is the interface agents reach for first: it is already in their toolbox, it composes, and its output lands directly in the transcript. A CLI designed for interactive humans, though, is a trap for headless sessions.
Non-interactive by default
No prompt without a flag that answers it. Confirmations get `--yes`, choices get flags with defaults, credentials come from env vars or flags. A command that stops to ask a question hangs a headless session until its timeout, and the hang reads as your product being broken. If a prompt is unavoidable, detect the non-TTY case and fail fast with the flag that would have answered it.
Stable exit codes and parseable errors
Zero for success, distinct stable codes for distinct failure classes, errors on stderr in one line the agent can parse. `--json` on every command that outputs data turns your CLI into an API; agents use it heavily when it exists and improvise brittle text parsing when it does not.
--help is the discovery surface
Agents learn a CLI from `--help` output. Every command and flag gets one line that says what it does and when to use it, with the same care as an MCP tool description. Undocumented flags might as well not exist; misdocumented ones are worse.
Parity with the API
Every task the API can do, the CLI can do under a predictable name, and both produce the same result. Divergence between surfaces is why one agent succeeds where another fails on the same task: each client discovered a different subset of your product. Generate the CLI from the same definitions as the API where you can, and diff the two surfaces in your release checks where you cannot.