Publisher controls

Control scraping with evidence and context.

A list of names or addresses is only one input. A useful scraper policy starts with the resource you need to protect and the behavior causing a problem.

This guide covers defensive controls for sites you operate. It does not provide a live scraper feed or a promise to identify every automated client.

Name the resource and the problem

Begin with a concrete observation: a public catalog endpoint receiving excessive requests, repeated retrieval of an expensive page, or an automated client reaching a restricted area. Each situation can require a different response.

Measure behavior using the information you are authorized to inspect. Keep a baseline for legitimate use, including accessibility tools, integrations, and search crawlers. Automation alone does not establish that a request is harmful or unwanted.

Use multiple signals cautiously

Request frequency, paths, response patterns, and documented identity checks can contribute context. None should be treated as a universal fingerprint. Shared IP addresses and network changes make permanent address-only decisions particularly difficult to interpret.

A practical review asks what rule matched, what action followed, and whether a legitimate task was interrupted. Keep enough information to answer those questions without unnecessarily retaining full URLs or visitor identifiers.

Prefer a scoped, reversible response

Choose a response around the measured problem. Options might include a resource-specific rate limit, cached responses, authentication for private content, or a temporary block. Availability depends on your hosting environment and application.

Document the owner, reason, duration, and rollback condition for each rule. When a policy is temporary, give it a review date. A rule added during an incident can otherwise outlive both the problem and the evidence supporting it.

Check the effect on real visitors

Repeat representative tasks after a change and inspect both accepted and rejected requests. Fewer requests are not automatically evidence of a better result if important users were excluded. Provide a route for reporting access problems.

For cooperative crawler instructions, read AI bot policy basics. For the difference between these instructions and enforcement, see the complete bot-control fieldnote.

Common questions

Is a user-agent block enough?

A client controls its user-agent claim. Treat it as a signal and verify identity where an operator provides a documented method.

Should scraper rules be permanent?

Set a review condition appropriate to the evidence and the problem. Temporary controls deserve explicit expiry or reassessment.

From Ad Block List Lab

Go a little deeper.

Related guide

AI Bot Block List

Crawler instructions and access controls have different jobs. Start by deciding which content a bot may request and how you will enforce that decision.

Explore this topic
Related guide

Web Analytics Block List

A blocked script, an empty dashboard, and reduced data collection are different observations. Use a test that explains what happened.

Explore this topic