---
title: "Agent-Ready Website: What We Built and What Matters — Dardo"
description: "What an agent-ready website needs, from building dardo.studio: markdown copies, llms.txt, AI crawler rules, an MCP server, an A2A card and what to skip."
url: "https://dardo.studio/en/blog/agent-ready-website/"
language: "en"
translation: "https://dardo.studio/es/blog/sitio-web-preparado-para-agentes-de-ia/"
updated: "2026-10-09T20:37:46.127Z"
---

[Blog](https://dardo.studio/en/blog/)

# What an agent-ready website needs: notes from building dardo.studio

We made dardo.studio readable and callable by AI agents, one layer at a time. Here is what each layer does, who documents reading it today, and which ones we would build again.

By [Nicolás Cerón](https://dardo.studio/en/studio/) ·October 9, 2026 · [Leer en español](https://dardo.studio/es/blog/sitio-web-preparado-para-agentes-de-ia/)

![A small wheeled robot following a crimson guide line across a museum gallery at night towards a lit sculpture.](https://dardo.studio/_astro/01M4GYDFZJP0ZV0S7JSNMWYADG_ZB09f9.webp)

## The short answer

An agent-ready website does three things. AI agents can reach it. They can read it without wading through menus, banners and scripts. And, where it makes sense, they can call a few declared tools instead of guessing which button to press.

We built all three into dardo.studio in early October 2026. Here is what we shipped and how much of it the major AI platforms document using, checked against the live site on 8 October 2026.

The summary is less exciting than most "AI-ready" checklists. The oldest layers carry the most weight: crawler access and clean, semantic HTML. Markdown copies and llms.txt are cheap conveniences. MCP, A2A and WebMCP are real protocols with working clients, but none of the crawler documentation we cite from OpenAI, Anthropic, Perplexity or Google describes their assistants finding a site's tools on their own.

## Start with access: robots.txt and the edge

Our [robots.txt](https://dardo.studio/robots.txt) repeats one group for the default agent (`*`) and for each AI search crawler and user fetcher we name:

```
User-agent: OAI-SearchBot
Allow: /
Disallow: /api/
Content-Signal: search=yes, ai-input=yes
```

The same rules apply to ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot, Claude-User and Bingbot. Only `/api/` is closed.

### Search crawlers and training crawlers are separate

The major companies now document separate tokens for search and for training:

- **OpenAI.** OAI-SearchBot surfaces sites in ChatGPT search. Sites that block it are left out of ChatGPT search answers, though they can still appear as plain navigational links. GPTBot collects training data. ([OpenAI crawlers](https://developers.openai.com/api/docs/bots))
- **Anthropic.** Claude-SearchBot indexes for search, ClaudeBot collects training data, and Claude-User fetches pages when a person asks Claude something. ([Anthropic crawlers](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler))
- **Perplexity.** PerplexityBot surfaces and links sites in Perplexity's results and is not used to train foundation models. ([Perplexity crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers))
- **Google.** Google-Extended is a control token for Gemini training and grounding in other Google products. It does not affect inclusion or ranking in Google Search. ([Google common crawlers](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers))

User-triggered fetchers are different. OpenAI says robots.txt may not apply to ChatGPT-User, because a person starts those requests. Perplexity-User and Google's [user-triggered fetchers](https://developers.google.com/crawling/docs/crawlers-fetchers/google-user-triggered-fetchers) generally ignore it. These fetchers act for one person in real time; where a block works, it mostly stops that person's assistant from reading your page.

Our file doesn't name the training crawlers, so they fall under `*` and are allowed. That is a business decision, and the vendors document it as a separate control: blocking GPTBot or ClaudeBot is not the same as leaving their search indexes.

### Content Signals: one line, no commitments

The `Content-Signal` line comes from Cloudflare's [Content Signals Policy](https://blog.cloudflare.com/content-signals-policy/), which names three uses: `search`, `ai-input` (feeding content to a model at answer time) and `ai-train`. We leave `ai-train` out, which under the policy neither grants nor restricts that use. Cloudflare says signals express preferences, block nothing, and may be ignored. None of the crawler pages cited here mentions them. It costs one line; expect nothing yet.

### Check the edge, not just the file

robots.txt states a policy; your CDN decides what happens. When we tested on 2 October, Cloudflare's Browser Integrity Check answered Python's default HTTP client (`Python-urllib`) and `libwww-perl` with a 403, while robots.txt allowed everything. Scripts written by coding agents often use Python's standard library unchanged.

We added a Cloudflare configuration rule that exempts GET and HEAD requests for public content; `/api/` still answers 403. The same review found our `.txt` files had no charset, so some clients read "Bogotá" as "BogotÃ¡". The fix was one header: `charset=utf-8`.

Test with real requests under each user agent. That shows nothing blocks the name; it doesn't prove the real crawler visited.

## Plain, semantic HTML: the layer every agent depends on

Google's [AI optimization guide](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) describes browser agents that analyze screenshots, inspect the DOM and interpret the accessibility tree. It points site owners to web.dev's [agent-friendly guidance](https://web.dev/articles/ai-agent-site-ux), which is mostly [accessibility work](https://dardo.studio/en/services/web-accessibility/): use `<button>` and `<a>` instead of styled `<div>`s, connect every label to its input, and keep the layout from shifting under a screenshot.

On dardo.studio, each page has one `<main>` and labelled `<nav>` landmarks. Menu and theme toggles are real buttons that report their state with `aria-expanded` and `aria-pressed`, and closed menus are `inert`. Every contact form field sits inside its `<label>`, which gives it an accessible name. web.dev suggests the `for` attribute; wrapping the input does the same job.

This helps screen reader users today, which is reason enough.

## A clean markdown copy of every page

Agents pay for every token they read, and a rendered page carries navigation, a cookie banner, scripts and decorative graphics. Before this work, requesting our pages with `Accept: text/markdown` returned all of that as HTML.

Now a build step writes an `index.md` beside every indexable page. It opens with front matter (title, description, canonical URL, language, the other-language version and the update date), then the page's `<main>` content without scripts, buttons, decorative images or the in-page table of contents. FAQ answers stay. Forms become a list of their fields and choices, so an agent can tell a person what our contact form asks without touching it.

There are three ways to get the copy:

- Send `Accept: text/markdown` to the normal URL. Those responses carry `Vary: Accept`, so caches keep the versions apart.
- Request the file: `/en/services/seo/index.md`.
- Append `.md` to the page path (`/en/services/seo.md`). For URLs that end in a slash, the [llms.txt proposal](https://llmstxt.org/) uses `index.md`, the form above.

Every HTML page also points to its copy with `<link rel="alternate" type="text/markdown">`.

### Point search engines back at the HTML

Google's guide notes it can crawl and index many file types besides HTML, without treating them specially. A markdown copy could compete with its own page, so every markdown response names the HTML page as canonical:

```
$ curl -sI https://dardo.studio/en/services/seo/index.md
content-type: text/markdown; charset=utf-8
link: <https://dardo.studio/en/services/seo/>; rel="canonical", ...
```

Who reads these copies? Cloudflare built [Markdown for Agents](https://developers.cloudflare.com/fundamentals/reference/markdown-for-agents/) to convert HTML at the edge for requests that prefer markdown, which suggests agents ask. It doesn't name which clients send the header, and we have no verified list either. If you use Cloudflare's feature, it adds `Content-Signal: ai-train=yes, search=yes, ai-input=yes` unless your origin sets its own. We generate our copies at build time so they match the page exactly.

## llms.txt: a useful index with no search effect

llms.txt is a [proposal by Jeremy Howard](https://llmstxt.org/), first published in September 2024 and still open for community input: a markdown file at `/llms.txt` with the site's name, a short summary and lists of links an agent may want.

Ours, at [/llms.txt](https://dardo.studio/llms.txt) and [/es/llms.txt](https://dardo.studio/es/llms.txt), is generated from the same data as the pages, so it can't drift. It states the studio facts (Bogotá, founded 2026, a team of three, how projects are priced, contact routes), lists services and work, and explains how to fetch markdown copies. `llms-full.txt` holds the full text of the studio, service, work and contact pages.

One line names similarly named businesses that are not us. When we checked on 2 October, they led the search results for "dardo studio". That line may be the most useful one in the file.

The status, plainly:

- **Google** says you don't need llms.txt to appear in Search or its AI features, that Search ignores it, and that keeping one neither helps nor hurts.
- **OpenAI, Anthropic and Perplexity** don't say in their crawler documentation that their bots read other sites' llms.txt. OpenAI's, Anthropic's and Perplexity's own documentation sites do publish one, for agents reading their docs.

Keep one if it's generated and accurate. It helps coding agents and tools that look for it. It is not a visibility lever.

## Tools agents can call: MCP, A2A and the API catalog

We published two read-only tools:

- `**list_services**` returns our published services, scope, deliverables and source URLs in English or Spanish, filtered by an optional keyword.
- `**get_project_brief**` returns the questions to answer before contacting us and the localized contact link for that service.

One implementation sits behind several entry points:

| Entry point   | Address on dardo.studio                         | Standard and status                                                                                                                                                                                           |
| ------------- | ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| MCP server    | /mcp, card at /.well-known/mcp/server-card.json | [MCP](https://modelcontextprotocol.io/specification/2025-11-25/basic/transports) Streamable HTTP; the server card is a [draft proposal](https://modelcontextprotocol.io/community/working-groups/server-card) |
| A2A agent     | /a2a, card at /.well-known/agent-card.json      | [A2A 1.0](https://a2a-protocol.org/latest/specification/), JSON-RPC                                                                                                                                           |
| JSON endpoint | /agent/services.json, described in OpenAPI      | Plain HTTP                                                                                                                                                                                                    |
| API catalog   | /.well-known/api-catalog                        | [RFC 9727](https://www.rfc-editor.org/rfc/rfc9727.html), IETF Standards Track                                                                                                                                 |

Every HTML and markdown response sends a `Link` header pointing to the catalog, agent skills index and both cards, so any page leads to the rest.

### Design rules we would repeat

- **Read-only and public.** The MCP tools declare `readOnlyHint: true` and read the same published catalog as the HTML, with no database behind them. `get_project_brief` doesn't submit, book or quote; a person reviews and sends.
- **Bounded input.** Request bodies are capped at 8 KiB, browser `Origin` headers are checked (the MCP specification requires it), and calls have their own rate limit.
- **No state.** The A2A agent answers immediately, keeps no tasks, has streaming off and doesn't fetch files or URLs sent to it.
- **Clear terms.** [/auth.md](https://dardo.studio/auth.md) says no credentials are needed and that reading public data doesn't authorize sending a message or making a payment.

We also left things out. Readiness scanners check for commerce protocols and OAuth discovery. We sell nothing through a checkout and protect no resources, so publishing those would describe capabilities that don't exist.

Who uses these tools today: MCP clients someone has connected to `/mcp`, and A2A clients given our card. For a studio, the value is modest: a precise answer to "what does Dardo do, and what should I send them?" The case is stronger for a site with live data people ask about, such as stock, availability or product documentation. Agents that write data need authentication and review steps; that is [AI automation](https://dardo.studio/en/services/ai-automation/) work.

## WebMCP: the same tools inside the browser

[WebMCP](https://webmachinelearning.github.io/webmcp/) lets a page register tools that an AI agent in the browser can call. It is a Draft Community Group Report of the W3C Web Machine Learning Community Group and states that it is not a W3C Standard. web.dev says it is in active development, may change, and can be tried in Chrome through an origin trial.

Our pages register the same two tools through `document.modelContext` (or `navigator.modelContext` in older previews). Without the API, the browser runs nothing extra. Because the tools existed, this took about 40 lines. Treat it as an experiment.

## Every layer, and whether it's worth it

| Layer                                  | What it is                                           | Who reads it today                                                   | Worth it?                                    |
| -------------------------------------- | ---------------------------------------------------- | -------------------------------------------------------------------- | -------------------------------------------- |
| Semantic HTML and labelled forms       | Real buttons, links, landmarks and labels            | Browsers, assistive technology and browser agents                    | Yes. Build it first                          |
| robots.txt for search and user agents  | Per-crawler rules that separate search from training | OpenAI, Anthropic, Perplexity and Google document their tokens       | Yes. Then test at the CDN                    |
| Content Signals                        | search, ai-input, ai-train preferences in robots.txt | No AI company we checked documents honoring them                     | One line. Expect nothing                     |
| Markdown copies with canonical headers | A clean text version of each page                    | Agents that request markdown; no published list of which ones        | Yes, with the canonical header               |
| llms.txt and llms-full.txt             | A curated index and full text for agents             | Google Search ignores it; no crawler documentation claims to read it | Keep it if generated. Not a visibility lever |
| MCP server (read-only)                 | Declared tools agents can call                       | MCP clients a person connects                                        | Only with data or actions worth calling      |
| A2A agent card                         | A machine-readable description of an agent           | A2A clients pointed at it                                            | Speculative for most websites                |
| API catalog (RFC 9727)                 | One well-known list of your public APIs              | Tools that look for it                                               | Cheap if you already have APIs               |
| WebMCP                                 | Tools the page registers in the browser              | Chrome, through an origin trial                                      | Experiment                                   |

## What we would build again

In order: semantic HTML, tested crawler access, markdown copies with canonical headers, a generated llms.txt, and tools only when there is something worth calling. Then measure. Server logs show which agents fetch markdown copies or call `/mcp`. A fetch is not a citation, and none of this guarantees [an AI system will mention you](https://dardo.studio/en/blog/appear-in-ai-answers/).

Agent readiness is part of our [AI search optimization](https://dardo.studio/en/services/ai-search-optimization/) work, alongside the content and measurement that decide whether AI answers cite you. To build it into your site, [tell us about the project](https://dardo.studio/en/contact/).

[Nicolás Cerón](https://dardo.studio/en/studio/)

Nicolás Cerón is the founder of Dardo, a brand, web design and development studio in Bogotá, Colombia.

## Keep exploring.

- [Service · **AI search optimization (GEO): get cited by ChatGPT and Google AI**](https://dardo.studio/en/services/ai-search-optimization/)

- [Your project · **Start a conversation**](https://dardo.studio/en/contact/)
