The short answer
An agent-ready website does three things. AI agents can reach it. They can read it without wading through menus, banners and scripts. And, where it makes sense, they can call a few declared tools instead of guessing which button to press.
We built all three into dardo.studio in early October 2026. Here is what we shipped and how much of it the major AI platforms document using, checked against the live site on 8 October 2026.
The summary is less exciting than most "AI-ready" checklists. The oldest layers carry the most weight: crawler access and clean, semantic HTML. Markdown copies and llms.txt are cheap conveniences. MCP, A2A and WebMCP are real protocols with working clients, but none of the crawler documentation we cite from OpenAI, Anthropic, Perplexity or Google describes their assistants finding a site's tools on their own.
Start with access: robots.txt and the edge
Our robots.txt repeats one group for the default agent (*) and for each AI search crawler and user fetcher we name:
User-agent: OAI-SearchBot
Allow: /
Disallow: /api/
Content-Signal: search=yes, ai-input=yesThe same rules apply to ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot, Claude-User and Bingbot. Only /api/ is closed.
Search crawlers and training crawlers are separate
The major companies now document separate tokens for search and for training:
- OpenAI. OAI-SearchBot surfaces sites in ChatGPT search. Sites that block it are left out of ChatGPT search answers, though they can still appear as plain navigational links. GPTBot collects training data. (OpenAI crawlers)
- Anthropic. Claude-SearchBot indexes for search, ClaudeBot collects training data, and Claude-User fetches pages when a person asks Claude something. (Anthropic crawlers)
- Perplexity. PerplexityBot surfaces and links sites in Perplexity's results and is not used to train foundation models. (Perplexity crawlers)
- Google. Google-Extended is a control token for Gemini training and grounding in other Google products. It does not affect inclusion or ranking in Google Search. (Google common crawlers)
User-triggered fetchers are different. OpenAI says robots.txt may not apply to ChatGPT-User, because a person starts those requests. Perplexity-User and Google's user-triggered fetchers generally ignore it. These fetchers act for one person in real time; where a block works, it mostly stops that person's assistant from reading your page.
Our file doesn't name the training crawlers, so they fall under * and are allowed. That is a business decision, and the vendors document it as a separate control: blocking GPTBot or ClaudeBot is not the same as leaving their search indexes.
Content Signals: one line, no commitments
The Content-Signal line comes from Cloudflare's Content Signals Policy, which names three uses: search, ai-input (feeding content to a model at answer time) and ai-train. We leave ai-train out, which under the policy neither grants nor restricts that use. Cloudflare says signals express preferences, block nothing, and may be ignored. None of the crawler pages cited here mentions them. It costs one line; expect nothing yet.
Check the edge, not just the file
robots.txt states a policy; your CDN decides what happens. When we tested on 2 October, Cloudflare's Browser Integrity Check answered Python's default HTTP client (Python-urllib) and libwww-perl with a 403, while robots.txt allowed everything. Scripts written by coding agents often use Python's standard library unchanged.
We added a Cloudflare configuration rule that exempts GET and HEAD requests for public content; /api/ still answers 403. The same review found our .txt files had no charset, so some clients read "Bogotá" as "Bogotá". The fix was one header: charset=utf-8.
Test with real requests under each user agent. That shows nothing blocks the name; it doesn't prove the real crawler visited.
Plain, semantic HTML: the layer every agent depends on
Google's AI optimization guide describes browser agents that analyze screenshots, inspect the DOM and interpret the accessibility tree. It points site owners to web.dev's agent-friendly guidance, which is mostly accessibility work: use <button> and <a> instead of styled <div>s, connect every label to its input, and keep the layout from shifting under a screenshot.
On dardo.studio, each page has one <main> and labelled <nav> landmarks. Menu and theme toggles are real buttons that report their state with aria-expanded and aria-pressed, and closed menus are inert. Every contact form field sits inside its <label>, which gives it an accessible name. web.dev suggests the for attribute; wrapping the input does the same job.
This helps screen reader users today, which is reason enough.
A clean markdown copy of every page
Agents pay for every token they read, and a rendered page carries navigation, a cookie banner, scripts and decorative graphics. Before this work, requesting our pages with Accept: text/markdown returned all of that as HTML.
Now a build step writes an index.md beside every indexable page. It opens with front matter (title, description, canonical URL, language, the other-language version and the update date), then the page's <main> content without scripts, buttons, decorative images or the in-page table of contents. FAQ answers stay. Forms become a list of their fields and choices, so an agent can tell a person what our contact form asks without touching it.
There are three ways to get the copy:
- Send
Accept: text/markdownto the normal URL. Those responses carryVary: Accept, so caches keep the versions apart. - Request the file:
/en/services/seo/index.md. - Append
.mdto the page path (/en/services/seo.md). For URLs that end in a slash, the llms.txt proposal usesindex.md, the form above.
Every HTML page also points to its copy with <link rel="alternate" type="text/markdown">.
Point search engines back at the HTML
Google's guide notes it can crawl and index many file types besides HTML, without treating them specially. A markdown copy could compete with its own page, so every markdown response names the HTML page as canonical:
$ curl -sI https://dardo.studio/en/services/seo/index.md
content-type: text/markdown; charset=utf-8
link: <https://dardo.studio/en/services/seo/>; rel="canonical", ...Who reads these copies? Cloudflare built Markdown for Agents to convert HTML at the edge for requests that prefer markdown, which suggests agents ask. It doesn't name which clients send the header, and we have no verified list either. If you use Cloudflare's feature, it adds Content-Signal: ai-train=yes, search=yes, ai-input=yes unless your origin sets its own. We generate our copies at build time so they match the page exactly.
llms.txt: a useful index with no search effect
llms.txt is a proposal by Jeremy Howard, first published in September 2024 and still open for community input: a markdown file at /llms.txt with the site's name, a short summary and lists of links an agent may want.
Ours, at /llms.txt and /es/llms.txt, is generated from the same data as the pages, so it can't drift. It states the studio facts (Bogotá, founded 2026, a team of three, how projects are priced, contact routes), lists services and work, and explains how to fetch markdown copies. llms-full.txt holds the full text of the studio, service, work and contact pages.
One line names similarly named businesses that are not us. When we checked on 2 October, they led the search results for "dardo studio". That line may be the most useful one in the file.
The status, plainly:
- Google says you don't need llms.txt to appear in Search or its AI features, that Search ignores it, and that keeping one neither helps nor hurts.
- OpenAI, Anthropic and Perplexity don't say in their crawler documentation that their bots read other sites' llms.txt. OpenAI's, Anthropic's and Perplexity's own documentation sites do publish one, for agents reading their docs.
Keep one if it's generated and accurate. It helps coding agents and tools that look for it. It is not a visibility lever.
Tools agents can call: MCP, A2A and the API catalog
We published two read-only tools:
list_servicesreturns our published services, scope, deliverables and source URLs in English or Spanish, filtered by an optional keyword.get_project_briefreturns the questions to answer before contacting us and the localized contact link for that service.
One implementation sits behind several entry points:
| Entry point | Address on dardo.studio | Standard and status |
|---|---|---|
| MCP server | /mcp, card at /.well-known/mcp/server-card.json | MCP Streamable HTTP; the server card is a draft proposal |
| A2A agent | /a2a, card at /.well-known/agent-card.json | A2A 1.0, JSON-RPC |
| JSON endpoint | /agent/services.json, described in OpenAPI | Plain HTTP |
| API catalog | /.well-known/api-catalog | RFC 9727, IETF Standards Track |
Every HTML and markdown response sends a Link header pointing to the catalog, agent skills index and both cards, so any page leads to the rest.
Design rules we would repeat
- Read-only and public. The MCP tools declare
readOnlyHint: trueand read the same published catalog as the HTML, with no database behind them.get_project_briefdoesn't submit, book or quote; a person reviews and sends. - Bounded input. Request bodies are capped at 8 KiB, browser
Originheaders are checked (the MCP specification requires it), and calls have their own rate limit. - No state. The A2A agent answers immediately, keeps no tasks, has streaming off and doesn't fetch files or URLs sent to it.
- Clear terms. /auth.md says no credentials are needed and that reading public data doesn't authorize sending a message or making a payment.
We also left things out. Readiness scanners check for commerce protocols and OAuth discovery. We sell nothing through a checkout and protect no resources, so publishing those would describe capabilities that don't exist.
Who uses these tools today: MCP clients someone has connected to /mcp, and A2A clients given our card. For a studio, the value is modest: a precise answer to "what does Dardo do, and what should I send them?" The case is stronger for a site with live data people ask about, such as stock, availability or product documentation. Agents that write data need authentication and review steps; that is AI automation work.
WebMCP: the same tools inside the browser
WebMCP lets a page register tools that an AI agent in the browser can call. It is a Draft Community Group Report of the W3C Web Machine Learning Community Group and states that it is not a W3C Standard. web.dev says it is in active development, may change, and can be tried in Chrome through an origin trial.
Our pages register the same two tools through document.modelContext (or navigator.modelContext in older previews). Without the API, the browser runs nothing extra. Because the tools existed, this took about 40 lines. Treat it as an experiment.
Every layer, and whether it's worth it
| Layer | What it is | Who reads it today | Worth it? |
|---|---|---|---|
| Semantic HTML and labelled forms | Real buttons, links, landmarks and labels | Browsers, assistive technology and browser agents | Yes. Build it first |
| robots.txt for search and user agents | Per-crawler rules that separate search from training | OpenAI, Anthropic, Perplexity and Google document their tokens | Yes. Then test at the CDN |
| Content Signals | search, ai-input, ai-train preferences in robots.txt | No AI company we checked documents honoring them | One line. Expect nothing |
| Markdown copies with canonical headers | A clean text version of each page | Agents that request markdown; no published list of which ones | Yes, with the canonical header |
| llms.txt and llms-full.txt | A curated index and full text for agents | Google Search ignores it; no crawler documentation claims to read it | Keep it if generated. Not a visibility lever |
| MCP server (read-only) | Declared tools agents can call | MCP clients a person connects | Only with data or actions worth calling |
| A2A agent card | A machine-readable description of an agent | A2A clients pointed at it | Speculative for most websites |
| API catalog (RFC 9727) | One well-known list of your public APIs | Tools that look for it | Cheap if you already have APIs |
| WebMCP | Tools the page registers in the browser | Chrome, through an origin trial | Experiment |
What we would build again
In order: semantic HTML, tested crawler access, markdown copies with canonical headers, a generated llms.txt, and tools only when there is something worth calling. Then measure. Server logs show which agents fetch markdown copies or call /mcp. A fetch is not a citation, and none of this guarantees an AI system will mention you.
Agent readiness is part of our AI search optimization work, alongside the content and measurement that decide whether AI answers cite you. To build it into your site, tell us about the project.
