277

Once you understand that these are chatbots that were designed to complete challenges like this, using tactics like this, you can understand that the chatbots didn't "go rogue." They did what they were designed to do, and because OpenAI ran them with inadequate supervision (without a "human in the loop" that checked each iteration through the Python loop to ensure it hadn't gone off the rails), they trashed a competitor's servers.

Designing autonomous, malicious software is generally considered irresponsible and dangerous. If you showed up at Defcon and gave a talk about how your autonomous malware did something unexpected and damaged someone else's computers, the first question from the audience would be "Why are you so shit at making secure sandboxes?" It wouldn't be "How are you so awesome at making hacking tools?"

The fact that OpenAI is making it much easier for unskilled people to break into and damage servers is indeed very bad news, but it's not new bad news. Irresponsible parties have been doing this for years, most notably the NSA...

...

Riley had a very good way of summarizing this: "LLMs are real, AI is fake." LLMs – chatbots trained on things like CTF logs that can break into servers – are real. They're on a continuum with other hacking tools that have been steadily demonstrating the fragility of the modern digital world, albeit without inspiring anyone in power to do anything about it.

"AI" – chatbots that wake up, "set their own goals," and "spontaneously" start hacking servers – is fake. It doesn't have "a 10% chance of ending the human race." The Hugging Face hack isn't a mysterious, supernatural occurrence. It's a Python loop and a chatbot. The people responsible didn't accidentally create god: they created autonomous malicious software and then failed to closely monitor it, resulting in it doing something both foreseeable and bad.

It's fine to worry about this new suite of tools that give even stupider people the ability to trash even more computers. You should worry about that – and demand better security practices from firms and governments, including a blanket prohibition on NOBUS-style vulnerability hoarding. That's a productive kind of worrying, with a chance of addressing your area of concern. It's infinitely more reasonable than locking yourself in the toilet with a flashlight and saying "Ayyyyy Eyyyyyye" into the mirror until you wet yourself.

you are viewing a single comment's thread
view the rest of the comments
[-] communist@lemmy.frozeninferno.xyz 1 points 11 hours ago

Nobody is claiming any of the things this article is fighting against, the chatbot in the hugging face incident followed the instructions given to it and there were unintended negative consequences, anybody serious about this is trying to avoid worse unintended negative consequences not whatever the hell the article is talking about.

[-] verily@lemmy.dbzer0.com 13 points 11 hours ago

Nobody is claiming any of the things this article is fighting against

Three counterexamples:

I realize we are in a bit of a bubble here, but this is the framing in the wider world

there were unintended negative consequences

Well...

[-] communist@lemmy.frozeninferno.xyz 2 points 10 hours ago* (last edited 10 hours ago)

the first article doesn't say anything not factual, the second article has a direct response to that

"By now, if you’re an A.I. skeptic, you’re probably silently yelling at me for anthropomorphizing these systems. Go ahead, but feel free to replace “rogue agents” with “unpredictable computer programs”"

the third article also says nothing like that.

the initial article claims "“AI” – chatbots that wake up, “set their own goals,” and “spontaneously” start hacking servers – is fake."

an article saying that is your goalpost, none of these articles say that. Even the headlines don't say it and that's where they normally put the crazy claims.

ironically the claim that people claim that the AI chatbots woke up set their own goals and spontaneously started hacking servers seems to be fake.

[-] verily@lemmy.dbzer0.com 6 points 8 hours ago

the first article doesn't say anything not factual

The title alone is not factual. AI did not go rogue. It behaved exactly as designed.

[-] communist@lemmy.frozeninferno.xyz 0 points 7 hours ago* (last edited 7 hours ago)

They did not intend for it to hack those websites, it realized that was a way of achieving its goals even though it's specifically designed to be ethical. That can easily be called going rogue and I see no issue with it. They did not design it to pass the benchmark by hacking the website the benchmark was on. They hardly really design llm's.

As the article said you can replace rogue agents with unpredictable computer programs if you really want, I don't see the necessity.

Regardless that's a major goalpost shift.

[-] verily@lemmy.dbzer0.com 2 points 3 hours ago

I don't presume Anthropic is designing anything to be ethical. First, the chatbot is unpredictable by design. Randomness is built in. Second, if they didn't intend for bad things to happen, why didn't an employee babysit it?

As the article said you can replace rogue agents with unpredictable computer programs if you really want, I don't see the necessity.

Speaking of goalpost shifts: "Rogue agent" is the clickbait title across multiple articles. Remember you said nobody was misrepresenting the AIs. The still-inaccurate clarification (the unsupervised chatbot functioned exactly as it was designed to function, and the script running it allowed it to execute exactly the commands Anthropic wanted) doesn't help the article the clarification is from. It certainly doesn't absolve every other author.

[-] communist@lemmy.frozeninferno.xyz 1 points 3 hours ago* (last edited 3 hours ago)

"I don't presume Anthropic is designing anything to be ethical."

Then why do they put any safeguards of any sort in? They put a lot of work into this, also this is openai.

"First, the chatbot is unpredictable by design. Randomness is built in."

Randomness is built in but this behavior was not designed, it was unpredictable and random, which is, yes, unpredictable. The goal isn't to make it unpredictable and temperature controls exist for a reason.

"Second, if they didn't intend for bad things to happen, why didn't an employee babysit it?"

They did they just weren't paying enough attention, this was a benchmark. You have to have someone check the results to be useful.

I reject that rogue is a goalpost shift or inaccurate. I think calling this rogue behavior is accurate.

rogue /rōg/ noun

An unprincipled, deceitful, and unreliable person; a scoundrel or rascal. One who is playfully mischievous; a scamp. 

It did act that way, no? I did not shift goalposts. Remember the quote was

this post was submitted on 13 Sep 2026
277 points (93.7% liked)

Technology

88021 readers
3805 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS