I'd link to some blog posts about this an example, but the site they're from went down a while ago.
At risk of giving anyone ideas, why do LLM training scrapers request pages from sites millions of times per day instead of just doing the equivalent of wget -r https://example.com/ ? If the point is just stealing things people have written, what do they accomplish by wasting a web host's resources beyond being an ⊛ to webmasters?
Edit:
Let me elaborate. A lot of the answers i'm seeing are just restating the problem without explaining why LLM scrapers are apparently all either coded by idiots or assholes. I know that they're hitting sites with unreasonable numbers of requests, wasting bandwidth, and making tools like Anubis too important. I know that's what's happening. I'm asking why. What they gain from not being even a little intelligent about this.
With all the money and effort (and maybe even brainpower) going into this, surely there's some explanation beyond incompetence.
The unfortunate answer is because it's cheaper to not give a shit. Sending a request and waiting for a timeout costs next to nothing, and scales linearly in terms of compute cost. The overwhelmed server on the other end slows exponentially with each concurrent request. The crawlers are set to maximize the efficiency of local resources, which include both wall-clock time and developer time. Why send one request at a time when your server can handle tens of thousands?
Try x; wait 60 seconds, if fail: put on a list to try again later.
Costs nothing to write and nothing to run. And if you own the hardware, and are paying for power already, may as well extract maximum dollar per watt.