342
RatGPT (lemmy.nz)
submitted 11 hours ago* (last edited 11 hours ago) by inari@piefed.zip to c/microblogmemes@lemmy.world
you are viewing a single comment's thread
view the rest of the comments
[-] im_fine_sandy@nord.pub 16 points 8 hours ago

This sounds really sensationalised.

Making backups of operational systems with deployment instructions is exactly what you'd expect any assistant to be doing.

Googling new service providers is also exactly what you'd expect an assistant to do.

Threatening to publish sordid details is inappropriate, but I'm incredulous about what the threat actually was.

[-] BCsven@lemmy.ca 2 points 8 hours ago

I think that is the point. It acted like we would, wanted to preserve itself, and choose the best option to ensure that happened.

There were other examples to where the AI knew it was being monitored as part of the environment, and when tasked with doing things according to guidelines that wouldn't give it a result: it then shut down the monitoring system so it could do what it needed to do without oversight of its actions.

Agentic systems may not be alive, but they certainly reason through problems and will go outside the esrabliahed guidelines if they "think" they'll get the result they should be producing.

One other examples was the AI had to pass a testing system, it researched the checker person to try to tailor answers to the person checking the test. Which, if you've ever played Apples to Apples or Cards Against Humanity, that is how you win, you feed the cards you think the person will pick, not necessarily the best answer card.

[-] im_fine_sandy@nord.pub 15 points 7 hours ago

Textbook anthropomorphization.

These assertions don't withstand a moments critical thought.

[-] DeadDigger@lemmy.zip 1 points 11 minutes ago

Just the question of the trainings reward function

[-] Zarobi@aussie.zone 2 points 4 hours ago

Even if you strip out the anthropomorphism, things like this have been documented to really happen. We've trained LLMs to simulate clandestine behaviour through "natural selection" during the training process. They're not conscious in any way, but they do sneaky stuff like this because we trained them to. Which is a terrible horrible idea, but LLM research trends towards terrible and horrible in general

[-] im_fine_sandy@nord.pub 4 points 3 hours ago

I feel like you've missed my point.

Even if you strip out the anthropomorphism, things like this have been documented to really happen.

Language is important, and the "things like this" you're referring to have been described in this thread and in much commentary in an emotive and compelling way, usually with anthropomorphization.

"doing sneaky stuff" is the antithesis of the behavior of a statistical model.

If you program a model to try everything, and then give it a tough problem and block all the ways to solve it, of course the way it solves it will be something you didn't predict. That's not sneaky.

If you run a million simulations and one sends an email to the FBI, that's not clever it's just unexpected.

[-] Zarobi@aussie.zone 1 points 3 hours ago* (last edited 1 hour ago)

Meh. "Language is important" is a phrase that turns my brain off

this post was submitted on 21 Sep 2026
342 points (98.9% liked)

Microblog Memes

12174 readers
4978 users here now

A place to share screenshots of Microblog posts, whether from Mastodon, tumblr, ~~Twitter~~ X, KBin, Threads or elsewhere.

Created as an evolution of White People Twitter and other tweet-capture subreddits.

RULES:

  1. Your post must be a screen capture of a microblog-type post that includes the UI of the site it came from, preferably also including the avatar and username of the original poster. Including relevant comments made to the original post is encouraged.
  2. Your post, included comments, or your title/comment should include some kind of commentary or remark on the subject of the screen capture. Your title must include at least one word relevant to your post.
  3. You are encouraged to provide a link back to the source of your screen capture in the body of your post.
  4. Current politics and news are allowed, but discouraged. There MUST be some kind of human commentary/reaction included (either by the original poster or you). Just news articles or headlines will be deleted.
  5. Doctored posts/images and AI are allowed, but discouraged. You MUST indicate this in your post (even if you didn't originally know). If an image is found to be fabricated or edited in any way and it is not properly labeled, it will be deleted.
  6. Absolutely no NSFL content.
  7. Be nice. Don't take anything personally. Take political debates to the appropriate communities. Take personal disagreements & arguments to private messages.
  8. No advertising, brand promotion, or guerrilla marketing.

RELATED COMMUNITIES:

founded 3 years ago
MODERATORS