The nuclear option is gaining traction as web traffic collapses and Google refuses to negotiate with content creators. […] Last week, the content delivery network Cloudflare, which hosts roughly one-fifth of the websites in the world, gave Google an ultimatum.
Beginning Sept. 15, all new websites signing up for Cloudflare, as well as all the customers on its free tier, will have the default settings in their bot management protocol set to block “multi-purpose crawlers” on any webpage that has ads. This means that any crawler that scrapes for both search indexing and AI training will be turned away at the door, unless the site owner decides otherwise.
[…]
While a handful of crawlers fit this description—Apple and Bing, among others—the primary, unnamed target of this action is Google, which infamously uses one crawler to both index sites and train its AI models.
I wonder how long it will be until Alphabet buys Cloudflare and reverses this.
Edit: Cloudflare just announced a deal with OpenAI. Looks like this move wasn’t altruistic, it was just about keeping their data set to themselves.
This is like connecting the NSA prism program to OpenAI.

Cloudflare blocking crawlers by default is pretty funny. Probably just means fewer small websites in the search results.




