Limit generative AI access to your website with the LLMs.txt file

Limit generative AI access to your website with the LLMs.txt file

With the meteoric rise of generative AI tools like ChatGPT, Gemini or Claude, a new question is stirring up the web: who controls access to the content used to train these AIs? Every day, millions of web pages are explored by automated agents that capture text, data and articles without publishers having given their explicit consent. The result: many sites see their content reused by AI, with no attribution, no traffic in return, and sometimes even for commercial purposes.

Faced with this situation, an initiative is starting to emerge: the LLMs.txt file. Inspired by the well-known robots.txt, it lets site owners express their wish to allow or refuse access to generative AI. Still little known, but backed by a growing part of the web community, this file could become an essential tool in the fight to protect online content.

What is the LLMs.txt file?

The LLMs.txt file (for Large Language Models) is a text file, much like the well-known robots.txt, that is placed at the root of a website. Its purpose is to communicate to the bots of generative AI the site owner's wishes regarding access to their content. Where robots.txt is aimed mainly at search engines like Googlebot, llms.txt targets a new category of crawlers: those used by AIs such as ChatGPT (OpenAI), Claude (Anthropic), Google-Extended (Gemini), or Common Crawl.

Main purpose

The aim of this file is to give an explicit instruction to the AI models that scan the web to train their algorithms or enrich their answers. It is not yet an official or binding standard, but rather a voluntary approach that relies on the goodwill of AI providers to respect the instructions addressed to them. It is a form of technical ethics that lets sites say: "I do not want my content to be scraped by generative AI."

A response to today's web challenges

The creation of this file responds to a growing concern: many publishers are finding that their content is being used without permission to train AIs, with no credit, no traffic in return, and sometimes for commercial purposes. LLMs.txt thus offers a first line of defence while we wait for stricter legal frameworks to potentially be introduced. Even though it is not an impassable technical barrier, it represents a clear stance in the debate on online data governance in the age of artificial intelligence.

Why set up an LLMs.txt on your site?

A website's content often represents a significant investment in time, expertise and resources. When a generative AI draws on that content to enrich its answers without mentioning the source or redirecting to the original site, it deprives the publisher of potential traffic and legitimate visibility. By explicitly prohibiting this use through the llms.txt file, a site can protect the strategic value of its content, particularly when it comes to original articles, in-depth research or monetised content.

Protecting your business model

Some sectors — notably media outlets, specialist blogs, educational sites or documentation platforms — rely heavily on their content to generate organic traffic, sell services or monetise their audience. If AIs absorb this content and reproduce the information without encouraging users to visit the source, it endangers their economic support. The llms.txt file then makes it possible to set a clear limit on this non-consensual capture, particularly for players who do not want to take part, even indirectly, in training competing or automated tools.

Reasserting your digital sovereignty

Beyond the commercial considerations, publishing an llms.txt is a symbolic and strategic act. It lets each site publisher reassert their right to control their data and their output. At a time when AIs are redrawing how the web is used, this tool offers sites a way to mark out their digital territory and demand more transparency in the collection of data. It is also a way to take an active part in an ethical and technological debate that is rapidly evolving.

Setting up an llms.txt therefore means taking back the initiative in the face of a technological revolution that, until now, has often taken place without consulting content creators.

How the LLMs.txt file works and how it is structured

The llms.txt file relies on a structure very close to that of the robots.txt file, which makes it easy to understand and adopt. It takes the form of a plain text file in which you specify, for each AI user-agent, the permissions or prohibitions on accessing the site's content. The syntax follows a clear logic: you first name the agent concerned, then the rule that applies.

Which agents can you target?

To date, several generative AIs or organisations linked to the training of language models have publicly declared the names of their bots, which makes it possible to identify them and include them in the llms.txt file. Among the best known are:

  • ChatGPT: OpenAI's crawler (ChatGPT)
  • ai
  • Google-Extended: corresponding to the indexing of data for Gemini
  • CC: the Common Crawl bot, whose data feeds several AIs

The list is not fixed, and new bots may appear as other AI players roll out their technologies.

A reach that is still limited, but significant

It is essential to understand that the llms.txt file is not a technical barrier like a firewall or a captcha. It works on a declarative principle that relies on the good faith of AI companies. In other words, only those that choose to respect these instructions will take the file into account. Nevertheless, in a context where some AIs are seeking to establish ethical legitimacy, this kind of signal is becoming an indicator of transparency and respect for web data.

Even if its technical effectiveness can still be improved, llms.txt is a concrete first step towards a web where content publishers can better manage access to their output in a digital environment increasingly dominated by artificial intelligence.

How to set up an LLMs.txt file on your site?

Setting up an llms.txt file is a simple process, accessible to anyone with access to the site's files. All you have to do is create a text file with the .txt extension and give it the exact name: llms.txt. Inside, you write the instructions intended for the generative AI crawlers, specifying the user-agents concerned and the associated rules, following the standard syntax (User-Agent / Allow or Disallow). It is possible to block certain agents while allowing others, depending on your editorial or commercial policy.

Location on the site

The llms.txt file must be placed at the root of the domain, like robots.txt. This allows the AI agents to detect it automatically when crawling the site. It is important to check that the file is publicly accessible via a simple browser or an HTTP request.

Monitoring, updates and watch

Once in place, the file must be updated regularly to remain effective. AI user-agents evolve quickly, new ones can appear, and some change their name or behaviour. It is therefore advisable to keep an active watch on the players in the sector, notably through official announcements, specialist blogs or lists maintained by the SEO community.

This monitoring lets you adjust the rules over time, according to the way AI practices evolve, but also to any legal or regulatory positions that might strengthen the file's reach.

In short, setting up the llms.txt file is quick and technically simple, but it is part of a proactive editorial management approach, at the crossroads of SEO, content protection and the challenges linked to artificial intelligence.

Current limits and prospects for the future

The llms.txt file, although inspired by the way robots.txt works, currently rests on no official technical standard or universal legal framework. Its effectiveness depends exclusively on the goodwill of the artificial intelligence companies that choose, or not, to respect the instructions given. Some, such as OpenAI or Anthropic, have publicly announced that they take this file into account, while others remain silent or have not communicated their position. As things stand, llms.txt therefore does not guarantee systematic protection against data collection by AIs.

A declarative, non-binding tool

Unlike technical measures such as a firewall, authentication or access-rights management via an API, llms.txt does not technically prevent access to the content. It is an ethical signal or a marker of intent that fits into a logic of web self-regulation. This declarative nature limits its reach, particularly against players who do not play by the rules or against AIs that use intermediary sources (such as Common Crawl) to get around the restriction.

Towards broader recognition?

Despite its current limits, llms.txt could in the medium term gain legitimacy, particularly if the major players in the web and AI agree on a common charter for respecting public data. It could also become part of future legislative measures, in Europe and elsewhere, aimed at better regulating the use of web content to train models. Discussions are already under way, notably in the wake of the European AI Act, to redefine the rules on transparency, copyright and access to data.

While we wait for future legal or technical developments, llms.txt is a concrete first initiative, at once symbolic and structuring, to begin building a more balanced power dynamic between content publishers and the players in generative AI.