Your developer asks whether to add an llms.txt file to the site. For most UAE businesses the honest answer is yes, eventually, and it should take an afternoon. What follows puts the file next to schema markup, robots.txt and your XML sitemap, then gives you an order of operations based on what AI crawlers actually do today rather than what the proposal hopes they will do.
Key Takeaways
- A proposal, not a guarantee: the big training crawlers largely don't read it yet, per independent testing and practitioner consensus around GPTBot and ClaudeBot.
- The format has a fixed order: an H1 project name, a blockquote summary, then H2-delimited link sections, hosted at your root domain.
- Schema markup still does more work: structured data is the more established, more widely parsed signal for AI visibility.
- Treat it as low-effort hygiene: add it after your content and schema are solid, never instead of them.
- Bilingual UAE sites need a wider stack: neither file settles which language version an answer engine quotes.
What Is an llms.txt File (And What It Isn't)
It's a proposed markdown file at your site's root that hands language models a structured map of your content. Not a standard. Not enforced.
Publishing one says nothing about whether any AI crawler reads it.
The proposal was first written in 2024 and is now on v2, revised after two years of adoption. Thousands of sites publish one, and documentation platforms generate them automatically.
It's voluntary, and it's designed to coexist with current web standards rather than replace them. robots.txt still controls access and your sitemap still lists URLs. The llms.txt file adds a curated map on top, aimed at agents that fetch pages to answer a user's question.
The Required Format and Structure of an llms.txt File
Markdown, in a fixed order: an H1 with your project or site name, a blockquote summary, optional detail paragraphs, then link sections under H2 headers. That is the whole spec.

Markdown is used because it's currently the most widely and easily understood format for language models. The detail paragraphs can hold prose or lists but no further headings. Each H2 section then carries a file list of markdown links.
The file sits at /llms.txt on your root domain. The spec names the Well-Known URIs standard (RFC 8615), which reserves the /.well-known/ prefix for metadata files like this one, as a discussed alternative location.
The spec's own worked example is the FastHTML project's file: an H1 reading # FastHTML, a blockquote describing the python library, then a ## Docs section listing links such as the FastHTML quick start and the HTMX reference, each with a one-line description. Copy that shape and you have a valid llms.txt file.
llms.txt vs Schema Markup vs robots.txt vs XML Sitemap: What Each One Actually Does
Three of these four are established and parsed at scale. One is new and mostly unread. Knowing which is which stops you swapping a working signal for an experimental one.
Here's how the four machine-readable files compare on what they target, whether AI crawlers honour them today, and what they cost to set up.
| File | What it targets | AI crawler support today | Setup effort |
|---|---|---|---|
| Schema.org structured data | Search engines and answer engines | Established, widely parsed | Medium, template by template |
| robots.txt | Access rules for named crawlers | Honoured, including GPTBot | Low |
| XML sitemap | Full URL inventory for indexing | Honoured by search crawlers | Low, usually automatic |
| llms.txt | Curated content map for agents | Limited and largely unconfirmed | Low, roughly an afternoon |
robots.txt already governs whether named AI crawlers such as GPTBot may fetch your pages at all. An llms.txt file grants nothing and blocks nothing; it's additive, a map for agents that have already been let in. Treat it as housekeeping, not as a substitute for structured data or a sitemap.
Do AI Crawlers Actually Use llms.txt Today?
Mostly not. In one independent test on a large publisher's site, the file received zero visits from Google-Extended, GPTbot, PerplexityBot or ClaudeBot between mid-August and late October 2025.
Publishing side adoption is thin too. One industry crawl counted only 951 domains with a published llms.txt file as of July 2025, a tiny fraction of the web.
The practitioner read is more nuanced than a flat dismissal. It's not widely getting honoured by the big training crawlers like GPTBot and ClaudeBot yet, though some documentation stacks ingest their own /llms-full.txt directly to power their docs agents.
Tooling is running ahead of confirmed crawler support. Chrome's Lighthouse now audits sites for an llms.txt file, which is why the topic keeps landing on developer to-do lists before the evidence justifies the priority.
How to Create and Publish an llms.txt File
Write four things: an H1 with your site name, a one-line blockquote summary, optional context paragraphs, then your key pages grouped under H2 headers as markdown links. Save it at your root.
Group pages the way a customer would ask, not the way your CMS files them: services, pricing, documentation, policies. Add a short description after each link. Check your documentation platform first, because many now generate the file for you.
The map is only worth as much as the pages it points at. A tidy index over unquotable content earns nothing, which is why structuring pages so an AI engine can quote them matters more than the index does.
Then read your server logs. Whether GPTBot, ClaudeBot or PerplexityBot ever request /llms.txt is the only honest test you have, and it belongs inside a wider routine for tracking AI search visibility.
Where Schema and Structured Data Do the Heavier Lifting
Structured data is already parsed by search engines and increasingly referenced by AI answer engines. The newer proposal is not. That gap is the entire argument for fixing schema first.

There's a technical reason schema carries the weight. Most AI crawlers can only read a page's basic HTML, not content that gets loaded by JavaScript, so server-rendered markup is what actually reaches the model.
Picture two layers of one stack: schema describes what a page is, and the content map tells an agent which pages exist and why. Our guide to getting cited by ChatGPT, Perplexity and AI Overviews covers the rest of that stack.
Fix your page-level markup before you spend another hour on llms.txt.
Building an AI Visibility Stack for a Dubai Business
Priority order: quotable, well-structured content and schema markup first. An llms.txt file comes third, once the first two hold up. Adding it is a low-effort approach that gives agents a canonical map.
The wrinkle for UAE companies is language. Plenty of Dubai business sites run Arabic and English side by side, and neither schema nor a content map decides which version an answer engine quotes, which is its own discipline covered in AEO for Arabic content.
Not sure which layer your site is weakest on? A free 30-minute consultation gets you an honest read, including a "don't build this" when that's the right answer.
Resist the urge to over-invest here. Some practitioners argue the file currently does nothing at all, which is harsher than the evidence supports, but it's a useful brake. An afternoon of work, not a site rebuild.
Is llms.txt Worth It? The Honest Verdict
Yes, but late in the queue. An llms.txt file is documentation hygiene worth adding once your core content and schema are solid, and it isn't an AI-visibility lever on its own.
Schema markup and clear, quotable pages remain the higher-leverage moves given what crawlers demonstrably fetch today. The file costs almost nothing, so publish it. Just don't expect it to move anything by itself.
Expect this guidance to keep shifting. The spec is already on v2 after two years of change, and crawler behaviour could turn the file from optional to expected inside a single release cycle.
Want a straight answer on where your site is losing AI citations? Book a free 30-minute consultation. We'll listen first, name the two or three fixes that matter most, and tell you plainly if llms.txt isn't one of them.
FAQ
Does adding an llms.txt file improve my Google ranking?
No. There's no evidence it affects ranking, and in independent testing Google's AI crawler never fetched the file at all. Ranking still comes from useful content, technical health and links.
Do ChatGPT, Perplexity and Google's AI Overviews actually read llms.txt files?
Not reliably. In one test running from mid-August to late October 2025, the file received zero visits from GPTbot, PerplexityBot, ClaudeBot or Google-Extended. Some documentation agents do ingest /llms-full.txt directly.
What's the difference between llms.txt and llms-full.txt?
llms.txt is the short curated index of links described in the spec. llms-full.txt is the expanded companion file, referenced in the spec as the full version, and it's what some documentation agents read directly.
Do I need both an llms.txt file and schema markup?
Prioritise schema markup, since it's already parsed at scale by search engines and answer engines. Add the llms.txt file afterwards as cheap insurance, because the two do different jobs and don't overlap.
Where should the llms.txt file be placed on my website?
At /llms.txt on your root domain, alongside robots.txt. The spec names the Well-Known URIs standard (RFC 8615), which reserves the /.well-known/ prefix for metadata files, as a discussed alternative.
Is llms.txt an official web standard?
No. It's a voluntary proposal, now on v2, designed to coexist with existing web standards rather than replace them. Nothing obliges a crawler to honour it.