For publishers

Set clear boundaries for AI crawlers.

Crawler instructions and access controls have different jobs. Start by deciding which content a bot may request and how you will enforce that decision.

Browser ad filtering acts in a visitor’s browser. Publisher bot controls act on requests arriving at your own site.

Know which direction you are controlling

A browser ad blocker operates for the person browsing a site. A publisher’s bot policy concerns automated clients visiting that publisher’s server. Calling both a “block list” does not make them interchangeable.

Write a policy that distinguishes ordinary search discovery, automated retrieval, and access to restricted content. Consider whether the same policy should apply to all public sections. A clear scope is easier to communicate and evaluate than a catch-all label such as “AI traffic.”

Use robots.txt for cooperative instructions

The robots exclusion protocol gives cooperating crawlers a way to read access instructions. It does not authenticate a visitor or prevent a client from sending a request. Do not put sensitive content in a public location and rely on a disallow rule to keep it private.

The bot and robots.txt guide explains this boundary and links to the protocol itself. Changes in crawler names or product policies should be verified against the operator’s current documentation.

Treat identity as a separate question

A user-agent string is a claim made by a client. When an operator documents a verification method, follow that method and track what was actually verified. Avoid assuming that a bot is genuine simply because its name looks familiar in a log.

When evidence is uncertain, choose a proportionate response. Review the affected paths and volume before adopting broad network blocks. Shared infrastructure can make sweeping decisions costly for legitimate visitors.

Protect resources at the appropriate layer

Private material needs appropriate access controls. Public resources may need rate policies, request limits, caching, or other protections according to their function. Those controls require hosting capabilities outside a static article or a client-side list.

See the scraper control guide for a practical evaluation workflow. Keep the objective specific: reduce a measured burden or enforce a defined access boundary, rather than promise that every automated request can be recognized.

Common questions

Does robots.txt enforce a block?

No. It communicates crawler instructions; it is not an authorization or authentication mechanism.

Will blocking AI bots remove existing copies of content?

A control on new requests does not itself delete copies already held elsewhere. Avoid assuming a forward-looking rule changes past collection.

From Ad Block List Lab

Go a little deeper.

Related guide

Scraper Block List

A list of names or addresses is only one input. A useful scraper policy starts with the resource you need to protect and the behavior causing a problem.

Explore this topic
Related guide

AI Score

Treat a score as one piece of evidence. The important questions are what it measures, how it was evaluated, and when a person should review it.

Explore this topic