Skip to content
TopicTracker
From HackerNewsView original
TranslationTranslation

Cloudflare sets AI crawler deadline: separate search or be blocked

Cloudflare announced new rules requiring AI crawlers to follow a specific protocol for data collection or risk being blocked. The company aims to give website owners more control over how their content is used by AI models, requiring crawlers to identify themselves clearly and respect site policies.

Background

- **Cloudflare** is a major internet infrastructure company that runs a global content delivery network (CDN) and offers services like DDoS protection, website security, and bot management. Many websites route their traffic through Cloudflare. - The company is now taking a stance on **AI crawlers** — automated bots that scrape website content to train large language models (LLMs) like ChatGPT, Google's Gemini, and others. Publishers have increasingly complained that these crawlers use their content without permission or compensation. - Cloudflare already gave site owners tools to block such bots. The new policy goes further: any AI crawler that wants to access Cloudflare-protected sites must publicly identify itself via a standard "user-agent" string (like "GPTBot" or "Google-Extended") and obey the site's robots.txt rules. If it doesn't, Cloudflare will block it automatically by default. - This matters because many AI crawlers have been disguising their identity or ignoring opt-out signals. It also pressures AI companies to negotiate licensing deals with publishers rather than taking data without permission.

Related stories