167

Humans are reading ChatGPT users' prompts to improve OpenAl's models, and those chats can include sensitive, personal information, according to leaked internal documents and real prompts seen by 404 Media.

The news presents a major privacy risk for ChatGPT's users, with people often using ChatGPT as a therapist, professional assistant, or digital friend, and providing it with all sorts of intimate details about their lives. The contractors don't see ChatGPT usernames, and OpenAl says it tries to remove personal information before prompts reach the reviewers, but the company acknowledged sensitive details can still get through.

The news also dispels the misconception that these models are improving only because of OpenAl's mass scraping of the internet, the talent of its well-paid engineering and Al teams, or the power of its newer models. An important and overlooked part are the outside contractors paid to read and review ChatGPT responses to real prompts over and over again. Anthropic confirmed to 404 Media it is also using human review to improve its models.

"No," someone who works with the prompts said when asked if they think ChatGPT users know that humans are reading their chats. "I don't think they would imagine some contractor somewhere [...] is analyzing the conversations."

Unpaywalled link here - https://archive.ph/98Wr5

you are viewing a single comment's thread
view the rest of the comments
[-] village604@adultswim.fan 11 points 3 days ago* (last edited 3 days ago)

That's not what this is. Humans aren't providing the responses, they're reviewing them, grading them, and using them to train models.

An example of a mechanical turk would be Waymo using humans to navigate situations the computer can't.

[-] schipelblorp@sh.itjust.works 2 points 3 days ago* (last edited 3 days ago)

If AI is so smart, why does it need this level of baby-sitting from humans? Mechanical Turk, one step removed.

The tech gods are saying they've reached Artificial General Intelligence, yet clearly if they have to pay humans to continually train the bots on the subtleties of human interaction, they have not reached AGI, and even feeding an LLM the entire corpus of the internet is insufficient to train them.

As much as I think this is a good and necessary step (privacy concerns aside), it definitely does point to AI being significantly less than advertised.

[-] village604@adultswim.fan 6 points 3 days ago

It's not babysitting, it's training. The humans grading the responses is how the models are being improved. It's just a more targeted approach than "let's scan the whole internet."

Humans aren't pretending to be AI. This is a completely different scenario than the mechanical turk.

[-] schipelblorp@sh.itjust.works -2 points 3 days ago

You are correct, it is not literally talking to a person. You could even go further and point out that the trainers are not literally Turkish.

[-] village604@adultswim.fan 3 points 3 days ago* (last edited 3 days ago)

You seem to not be grasping the concept of the mechanical turk. The human was the one playing chess, not the machine. The machine playing chess was a lie.

You're basically saying that a student's math teacher is actually doing the student's homework because they taught the student how to do it.

[-] schipelblorp@sh.itjust.works -1 points 3 days ago

Yes, and this machine that taught itself how to be smart from absorbing data is likewise a lie.

[-] village604@adultswim.fan 3 points 3 days ago

That still doesn't mean it's comparable to the mechanical turk. The machine didn't actually do anything. It was all the human.

That's not the case for AI. What AI does is actually transformative, and the output is not just a person pretending to be a machine.

A human providing the output while pretending to be a machine is a hard requirement.

[-] MangoCats@feddit.it 2 points 3 days ago

The machine isn't smart, or dumb. The machine generates responses based on training. How well those responses match expectations is a measure of the training.

[-] MangoCats@feddit.it 3 points 3 days ago* (last edited 3 days ago)

If AI is so smart,

First off, the label AI is bad - it's Artificial, but Intelligent - or smart - it is not.

LLM - Large Language Model + scaffolding. They match patterns. They build an unimaginably complex "context window" of 200,000 tokens and from that they synthesize their responses to prompts of usually a few dozen tokens based on trillions of model weights trained on hundreds of trillions of tokens. Optical character recognition has been doing this to automatically "read" the routing numbers printed on the bottom of checks since the early 1960s, that old OCR just has much much lower dimensions of patterns to match.

The "Mechanical Turk, one step removed" is helping to tune those trillions of model weights by reinforcing when it "gets things right" and flagging when it "gets things wrong." The LLM has to find this out somehow, that's what the employees are doing - training it, teaching it right from wrong.

AlphaZero taught itself to play chess, Go, and other relatively simple games nearly 9 years ago now - it mastered those with "best next move to win" pattern analysis / generation all by itself because the rules of those games are crisp, well defined, simple. What makes "a good response" in a conversation is a hugely different animal with orders of magnitude more dimensions and levels of nuance.

After all this training, the next generation model should, statistically, generate "correct" conversational responses more often, according to what the inputs of the employees consider and tell the model is correct and incorrect for a given situation.

this post was submitted on 16 Sep 2026
167 points (97.2% liked)

Technology

88132 readers
4549 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS