994
submitted 4 days ago* (last edited 4 days ago) by tonytins@pawb.social to c/technology@lemmy.world

Executives working on AI at Microsoft and OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing theft of unprecedented proportions,” and the “largest theft of labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire web.”

Those statements and a series of other mask-off moments feature heavily in an unredacted court filing that was unsealed Thursday in the behemoth New York Times vs OpenAI copyright lawsuit that has been winding its way through the court system for years. In a filing asking for summary judgment (basically, a filing with the court asking it to rule), lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained either sealed or redacted at the request of Microsoft and OpenAI.

It’s easy to see why the AI companies wanted to hide this from the public. The statements, taken together, are some of the most damning indictments of the ways LLMs were trained, how they worked, and the immediate threat they pose to human labor. It is a reminder that even as AI becomes more powerful and companies try to shift the narrative to the supposed existential risk of “superintelligent” AI, the tools they have already built were created by stealing from human creativity and labor and are by definition existential threats to the human labor market.


you are viewing a single comment's thread
view the rest of the comments
[-] kablez@lemmy.world 51 points 4 days ago

Gonna get real weird soon when they run out of rich new training data and they begin to consume their own shit. When that happens their entire model will collapse and if the bubble hasn't popped already that may be what causes it.

[-] RepleteLocum@lemmy.blahaj.zone 33 points 4 days ago

They're already doing it. They call it distilling when they take it from another llm. Pretty sure most content was already stolen in the early days and they now rely on distillation and stealing new content.

[-] kablez@lemmy.world 18 points 4 days ago

So it's like if everyone combined the backwash from their water bottles into a new drink...

Sounds super appealing!

[-] SaharaMaleikuhm@feddit.org 2 points 4 days ago

Don't get high off your own supply?

That's my thought. So long as someone inputs anything new to the internet they will be able to scrape it and sell it as their product.

[-] MalReynolds@slrpnk.net 11 points 4 days ago

Pretty sure they mostly use the pre-AI internet (that they scraped and kept) and synthetic data currently. Probably trying (and failing so far or we'd have heard) to adapt to using video as training material at the moment, but developments there will likely apply to robotics at some point. Here's hoping the current chuds have crashed and burned before then and that some sanity has taken over from unfettered capitalist oligarchs dreams of computer slavery.

[-] ScoffingLizard@lemmy.dbzer0.com 2 points 4 days ago

Why don't you thinj they have not adapted to video as training. Is that not what Flock does?

[-] MalReynolds@slrpnk.net 3 points 4 days ago

Flock does pattern recognition, a quite old piece of machine learning. Nothing to do with training a large language model or other 'AI' model.

[-] ScoffingLizard@lemmy.dbzer0.com 1 points 14 hours ago

They almost certainly use that to train AI. Big tech is aggregating the data from many places.

[-] MalReynolds@slrpnk.net 1 points 13 hours ago

Nah, it's in the name, Large Language Models, what gets marketed as 'AI', are trained on text. That's why they're ripping up secondhand books (and copyright law, but that never stopped them) at the moment. No one has really cracked Large Vision Model yet, although there are some primitive versions being attempted in robotics, mostly they translate to text, which is why general purpose robots aren't a thing yet.

[-] Thorry@feddit.org 4 points 4 days ago* (last edited 2 days ago)

Why do you think there has been such an emphasis on hacking with LLMs lately (especially by OpenAI). They figured out all those vulnerability databases were an excellent source for training material. In the past they scraped those, but just for general language training. Now they've specifically trained the models on the information within. Some team figured out how to use that data to train a model and have testing scenarios automated so they could write a good reward function. It wasn't that they figured out the models are good at hacking, they ran out of content and found a new source of good data.

With all the books they've been scanning I wonder if the next thing is going to be a writing assistant or editor or something like that. Even though writing good books is an art form and the actual writing down of the words is the easiest part (still not easy tho).

These companies are starving for content and they've not just poisoned but absolutely destroyed the content well that is the internet. Given they were already hitting diminishing returns hard, it doesn't matter too much to them probably. But more compute and storage has also been hitting diminishing returns hard and customers are complaining about the cost. So they are getting a bit desperate on how to improve these things at all.

[-] WorldsDumbestMan@lemmy.today 2 points 4 days ago

Or they can just keep a database of actual data on their servers, instead of getting new data every single time for some fucking reason.

[-] BilSabab@lemmy.world 1 points 4 days ago

model collapse is the endgame. that's the whole point of LLM.

[-] ScoffingLizard@lemmy.dbzer0.com 3 points 4 days ago

What do you mean? Why would collapse be the endgame?

[-] DeadDigger@lemmy.zip 2 points 4 days ago

It's either agi or model collapse. For every AI system actually. If you have a high enough adoption you start to muss original data so if your AI is not self sufficient in time it will collapse, because it will be trained on its own data, which just is an incentive loop

[-] BilSabab@lemmy.world 1 points 4 days ago

in a manner of speaking - you always end up there. not by design though. models operate via continuous refinement and you can only optimize a model so much until it is a mess and you need to figure out where to roll back. so you either get shit like semantic drift or variance decay and you can whack a mole it to an extent but then you hit the rlhf wall when the model starts gaming its reinforcement framework and the fat lady sings.

[-] ScoffingLizard@lemmy.dbzer0.com 1 points 14 hours ago

gaming its reinforcement framework

Well put. Never thought of it that way.

[-] BilSabab@lemmy.world 1 points 5 hours ago

the only more or less workable way to keep it under control is maintaining a closed loop small-scale environment - kinda like NotebookLM where you upload documents and that's all there is to work with - outside of that it is a mess.

this post was submitted on 18 Sep 2026
994 points (98.6% liked)

Technology

88185 readers
3155 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS