AI crawler directives: what they are, how they work, and…
AI crawler directives: what the term covers
AI crawler directives are the emerging machine-readable rules, files, headers, and metadata that websites use to signal what AI crawlers and AI systems may access, train on, cite, or reuse. The key caveat is that this is not one settled standard; it is a fragmented landscape of competing proposals and drafts.
AI crawler directives is best understood as an umbrella term, not the name of a single protocol. In practice, it refers to several machine-readable approaches that let publishers communicate preferences to AI crawlers and downstream AI systems about access, training use, citation, reuse, and related policy choices. Today’s landscape spans file-based proposals, robots-style rule files, JSON policy files, and inline mechanisms attached at delivery time through headers or page metadata.
So the useful working definition is simple: AI crawler directives are the emerging controls websites publish to express AI-specific permissions and preferences, but the ecosystem is still early, fragmented, and unevenly adopted.
The main types of AI crawler directives in use and in draft
The main types of AI crawler directives today fall into four buckets: standalone policy files such as ai.txt, robots-style files such as robots-ai.txt, JSON policy files such as ai-policy.json, and inline attachment methods delivered through HTTP headers or HTML meta tags. That framing matters because people often search for “the” AI crawler directive when the reality is a small family of overlapping proposals.
The best-known file-style proposal is ai.txt. In the current draft form, ai.txt is meant to live at a well-known URI and give site owners a standardized, machine-readable place to declare AI usage preferences, policies, and licensing terms. The licensing part is notable because it goes beyond a basic allow-or-block model and tries to express commercial conditions for AI use of site content.
A second branch is robots-ai.txt. This specification uses a robots-like file format to let website owners express preferences specifically about AI crawler access and usage. Its appeal is familiarity: teams already understand the idea of a crawler control file, and the spec adds AI-focused instructions on top of that pattern. It also supports bot-specific rules, so a publisher can allow or block named crawlers such as GPTBot or Claude-Web instead of treating all AI crawlers the same way.
A third approach is ai-policy.json. This proposal defines a JSON-based format for declaring AI training and crawling policies. Because it is structured JSON rather than a line-oriented crawler file, it is often described as a better fit for richer policy expression. In the current spec language, it can combine crawl permissions, training permissions, and citation preferences in one file.
The fourth bucket is inline attachment. Instead of asking crawlers to fetch a separate policy file, the IETF AI preference attachment draft defines a mechanism for attaching preferences directly to content through HTTP headers or HTML meta tags. In particular, it introduces an AI-Preferences header that can declare AI usage policies inline with content delivery. That is useful when a publisher wants preferences to travel with the specific response rather than live only in a sitewide document.
Put together, these formats show why the category is still unsettled. There are multiple serious attempts to solve the same policy problem from different angles: file discovery, robots-style crawler controls, structured JSON policy declaration, and inline preference attachment.
How these directives differ from robots.txt and when they work together
AI crawler directives differ from robots.txt by trying to express richer AI-specific permissions, not just broad crawl access. In practice, they often work alongside robots.txt rather than replacing it, because publishers may need one layer for conventional crawler control and another for AI-specific usage preferences.
Classic robots.txt is mainly about whether a crawler may fetch a URL. The newer AI-focused proposals try to go further. For example, robots-ai.txt is explicitly designed to work alongside standard robots.txt by extending directives for AI crawlers and AI-related usage questions. That makes it less a replacement for robots.txt than a companion layer aimed at use cases robots.txt was never built to describe.
You can see the same pattern in Content Intent Signaling. It has been proposed as a robots.txt directive for controlling how AI models use website content, not merely whether they crawl it. More specifically, it is meant to distinguish whether content is intended for training, retrieval, or no AI use at all. That is a meaningful shift: the question is no longer just access, but purpose.
aI-policy.json points in a similar direction from a different format. Its core claim is that websites can declare AI training and crawling policies in JSON, and it has also been described as supporting integration with existing robots.txt and sitemap.xml workflows. That integration point is less certain than the base format itself, but it reflects the broader reality that site owners may end up publishing multiple machine-readable signals at once rather than betting on one file to do everything.
So if you are comparing AI crawler directives with robots.txt, the practical answer is: robots.txt still handles generic crawler access, while newer AI-oriented directives try to add nuance around training, retrieval, citation, and usage intent.
What site owners can realistically control with AI directives
What site owners can realistically control today is limited but meaningful: they can publish preferences about crawler access, training permission, usage intent, citation or retrieval treatment, licensing, and sometimes per-bot rules. What they cannot assume is universal enforcement, because the ecosystem is still fragmented and draft-heavy.
Across current proposals, several control levers show up repeatedly. ai.txt is positioned as a place to declare general AI usage preferences and policies, including licensing terms for AI use of site content. That makes it useful for publishers who want to signal not only access preferences but also conditions for reuse.
If your main concern is bot-by-bot treatment, robots-ai.txt and similar bot-specific systems are more relevant. The robots-ai.txt specification supports allowing or blocking named AI crawlers such as GPTBot or Claude-Web. Separate bot-specific control systems also describe granular rules for individual crawlers, including GPTBot, Claude, and Google-Extended. Some of those systems go further and support different crawl rates and access windows for different crawler types.
If your concern is use case rather than crawler identity, Content Intent Signaling is the more interesting idea. It is aimed at separating training from retrieval and from blanket disallowance, which is a much more nuanced policy model than a simple yes-or-no crawl rule.
And if you want one structured policy artifact, ai-policy.json is described as combining crawl permissions, training permissions, and citation preferences in a single JSON file.
The realistic caveat is that these directives are signals and policy expressions within an evolving ecosystem, not a magic switch that guarantees every AI system will behave the same way.
FAQ about AI crawler directives
Is there one standard for AI crawler control yet?
No. There is no single settled standard yet. The current landscape includes competing or parallel approaches such as ai.txt, robots-ai.txt, ai-policy.json, and header-based attachment methods, with ongoing draft activity showing that standardization is still in progress.
Do AI crawler directives replace robots.txt?
Usually no. The practical pattern is coexistence: robots.txt still handles broad crawler access, while AI-specific proposals try to express richer permissions around training, retrieval, citation, or bot-specific treatment.
What formats are people talking about when they say AI crawler directives?
Most discussions refer to standalone files like ai.txt and robots-ai.txt, structured JSON files like ai-policy.json, or inline methods such as HTTP headers and HTML meta tags for attaching AI preferences.
Can I block GPTBot and Claude separately?
In some proposals, yes. The robots-ai.txt specification allows website owners to block or allow specific AI crawlers such as GPTBot or Claude-Web, and other bot-specific control systems describe similarly granular per-crawler rules.
Do I need llms.txt as well?
Possibly, but it solves a different problem. Some guidance recommends a /llms.txt file as a curated summary for LLM discovery and transparency rather than as an access-control directive.
Are there practical guides for implementing these files?
Yes. Guidance aimed at AI startups includes setup steps for LLM.txt files and crawler controls designed to improve AI visibility, though that is implementation advice rather than a formal standard.
Turn AI directive research into publishable policy pages
If your team wants to cover evolving AI-search topics without hand-writing every glossary page, the product is built to turn grounded research into publishable pages at scale. It generates eight page archetypes — glossary, how-to, comparison, integration, alternatives, listicle, persona, and free-tool — and does it on a recurring cadence rather than as one-off drafts.
The core pitch is reliability. It grounds content in customer-provided product and competitor truth files, fact-checks claims, and blocks pages that cannot be grounded. It can also queue refreshes automatically when citation counts or rankings slip past plan thresholds, then publish static pages with sitemap support, canonical tags, internal links, and schema.org JSON-LD attached.
In other words: if AI directive standards keep changing, you do not need a brittle manual content workflow to keep up.