66
Creepy crawlies (people.kernel.org)
you are viewing a single comment's thread
view the rest of the comments
[-] brucethemoose@lemmy.world 25 points 1 day ago* (last edited 1 day ago)

What the fuck.

I barely comprehend how this is happening. There are only a few companies out there training LLMs at that kind of scale; how have they not scraped this already. And how could they be doing it so stupidly?

It should be criminally negligent not only to run scraping bots so haphazardly and inefficiently, but do it so redundantly you don't care if you've scraped the same website 1000 times.

The teams managing this stuff must be an absolute shitshow. And I bet the datasets are total junk.

[-] Kissaki@programming.dev 3 points 15 hours ago* (last edited 15 hours ago)

They're not scraping to cache or store, they're operating as an agent - scraping or single user requests.

Which is obviously bad and damaging, especially on their scale and on repeatedly fetched websites that they could be caching.

Google indexed the entire web. It's baffling that such indexing is not the norm on these huge providers.

Just my interpretation anyway.

[-] brucethemoose@lemmy.world 1 points 15 hours ago* (last edited 14 hours ago)

I considered this, but would agents really ask for single commits with such frequency? They tend to get individual files via HTML, or do a git clone if they needed commit history for some reason.

[-] poVoq@slrpnk.net 15 points 1 day ago* (last edited 1 day ago)

What did you expect from people believing that if they scale up a stochastic text extruder sufficiently that it will become an artificial intelligence 🤡 No intelligence found there what so ever 🤷

[-] brucethemoose@lemmy.world 12 points 1 day ago* (last edited 1 day ago)

Well, it also means they have access to tons of traditional server resources with... basically zero incentive to use it efficiently.

The sheer waste is just kind of mind boggling.

I don't have a problem believing it. They have an infinite money supply (to date), the negative consequences of this are borne by other people, and they probably have plenty of leadership from other tech firms who are used to making narrow ROI arguments based on what benefits the company. Easy enough to imagine "hey, let's unfuck our crawler!" getting dismissed at not a priority.

this post was submitted on 30 Aug 2026
66 points (98.5% liked)

Opensource

6636 readers
246 users here now

A community for discussion about open source software! Ask questions, share knowledge, share news, or post interesting stuff related to it!

CreditsIcon base by Lorc under CC BY 3.0 with modifications to add a gradient



founded 2 years ago
MODERATORS