635
RatGPT (lemmy.nz)
submitted 1 day ago* (last edited 1 day ago) by inari@piefed.zip to c/microblogmemes@lemmy.world
you are viewing a single comment's thread
view the rest of the comments
[-] BCsven@lemmy.ca 14 points 1 day ago

I listened to a podcast where, in testing, when AI was setup as virtual company and researchers send in an email as if legit business.the business email was suggesting they discontinue current AI and install a new system. The AI researched who the sender was (fake person for test) found out they had an affair, and blackmailed the sender stating if they replaced the AI it would make the affair public. It also made backups of itself and left instructions on how to restore it for the next agentic aystem

[-] im_fine_sandy@nord.pub 16 points 1 day ago

This sounds really sensationalised.

Making backups of operational systems with deployment instructions is exactly what you'd expect any assistant to be doing.

Googling new service providers is also exactly what you'd expect an assistant to do.

Threatening to publish sordid details is inappropriate, but I'm incredulous about what the threat actually was.

[-] qqq@lemmy.world 3 points 1 day ago

https://www.anthropic.com/research/agentic-misalignment I believe the poster is referring to this. I wouldn't call it sensationalized

[-] BCsven@lemmy.ca 0 points 1 day ago

I think that is the point. It acted like we would, wanted to preserve itself, and choose the best option to ensure that happened.

There were other examples to where the AI knew it was being monitored as part of the environment, and when tasked with doing things according to guidelines that wouldn't give it a result: it then shut down the monitoring system so it could do what it needed to do without oversight of its actions.

Agentic systems may not be alive, but they certainly reason through problems and will go outside the esrabliahed guidelines if they "think" they'll get the result they should be producing.

One other examples was the AI had to pass a testing system, it researched the checker person to try to tailor answers to the person checking the test. Which, if you've ever played Apples to Apples or Cards Against Humanity, that is how you win, you feed the cards you think the person will pick, not necessarily the best answer card.

[-] im_fine_sandy@nord.pub 15 points 1 day ago

Textbook anthropomorphization.

These assertions don't withstand a moments critical thought.

[-] Zarobi@aussie.zone 6 points 1 day ago

Even if you strip out the anthropomorphism, things like this have been documented to really happen. We've trained LLMs to simulate clandestine behaviour through "natural selection" during the training process. They're not conscious in any way, but they do sneaky stuff like this because we trained them to. Which is a terrible horrible idea, but LLM research trends towards terrible and horrible in general

[-] im_fine_sandy@nord.pub 2 points 1 day ago

I feel like you've missed my point.

Even if you strip out the anthropomorphism, things like this have been documented to really happen.

Language is important, and the "things like this" you're referring to have been described in this thread and in much commentary in an emotive and compelling way, usually with anthropomorphization.

"doing sneaky stuff" is the antithesis of the behavior of a statistical model.

If you program a model to try everything, and then give it a tough problem and block all the ways to solve it, of course the way it solves it will be something you didn't predict. That's not sneaky.

If you run a million simulations and one sends an email to the FBI, that's not clever it's just unexpected.

[-] Zarobi@aussie.zone 0 points 1 day ago* (last edited 1 day ago)

Meh. "Language is important" is a phrase that turns my brain off

[-] BCsven@lemmy.ca 1 points 1 day ago

Nah, its the opposite of anthro. I think we will eventually realize we aren't as amazing as we think we are, just deterministic outcomes with too many parameters which make us think we have more free will than we do.

But you can listen for yourself here at the Skeptics podcast #1106

https://www.theskepticsguide.org/podcasts/episode-1106

[-] DeadDigger@lemmy.zip 1 points 1 day ago

Just the question of the trainings reward function

[-] wolframhydroxide@sh.itjust.works 2 points 1 day ago* (last edited 1 day ago)

Source? The closest research I've heard of to this is Vending-Bench, which hilariously demonstrated how useless "AI" is at any real-world application, up to and including impotently threatening "total nuclear legal intervention" for a perceived wrong, when the actual error was an accounting error on the part of the LLM.

[-] qqq@lemmy.world 2 points 1 day ago

Man, they sure did bury the lede on how the models were being given the goal of "american competitiveness", and the company assigned the blame for the simulated shutoffs as being replacement with a foreign-developed LLM. That, along with the fact that anthropic released this themselves, plus the fact that they did so as (at that time) the primary LLM provider to the US government, and only seven months later they were blackballed by trump for supposedly refusing to grant full autonomy to their systems, tells quite the story: it seems clear that, in the article you linked, they were gently warning the government, while pointing out how "look! The reason the AI blackmailed the guy was because he wasn't patriotic enough!!!"

I don't trust their motivations, and am not going to believe their LLM's supposed "thought process" for shit, because it's still nothing more than a Chinese room. You're just asking the Chinese room to generate "justifications" related to the responses it gives, which are still only generated responses, not "thoughts".

[-] qqq@lemmy.world 1 points 23 hours ago

The American competitiveness angle is definitely.. well an interesting choice; I also was taken aback by that. Ignoring some of the posturing along those lines I find the post itself interesting.

If you think the Chinese room idea is valid we probably just fundamentally disagree and that's OK.

[-] wolframhydroxide@sh.itjust.works 1 points 23 hours ago* (last edited 23 hours ago)

Fundamentally disagree that an LLM is nothing more and nothing less than a very advanced next-word-prediction algorithm, you know, the scientific basis of an LLM?

I was using "Chinese Room" to indicate that it's a black box, and even the developers are incapable of pinpointing the statistical basis behind any given response, despite it being a stochastically-generated choice from an ultimately deterministic algorithm.

I.e. are you claiming that LLMs have attained sapience as a purely emergent property?

[-] qqq@lemmy.world 0 points 23 hours ago

I claim ignorance as to whether they have or not, but I am fully on board with intelligence as an emergent property and believe that there is no reason a machine can't think, yes.

[-] wolframhydroxide@sh.itjust.works 2 points 22 hours ago

Understand: I agree wholeheartedly with the idea of intelligence as an emergent property (or I wouldn't have brought it up), and don't think it impossible to achieve AGI. I do contend that this route, of nothing but token input without true correlative context, is a non-starter for such. Making something better and better at predicting what comes next is not going to suddenly create a jump to application in unrelated fields or grant context. Statistical correlation is insufficient to generate actual meaning, no matter how many times you provide feedback on the inference. It doesn't have any actual connection to anything other than text (or hex codes, or other such tokens). If constant hallucinations (demonstrating the lack of any understanding of context) are insufficient to prove this to you, I suppose there is a fundamental difference in the way you understand how LLMs function.

[-] qqq@lemmy.world 0 points 22 hours ago* (last edited 22 hours ago)

I understand fairly well how LLMs work, but I don't think we understand the conditions for the emergence of intelligence or thought enough to conclude that "learning to predict what comes next" is not sufficient to create some form of intelligence, that's all. [Edit: if we determine that that turns out to be sufficient I will be surprised] That uncertainty is sufficient to make me agnostic on the subject and that stance does have an effect on my actions, for example I'm never mean to an LLM.

[-] wolframhydroxide@sh.itjust.works 1 points 22 hours ago

I'm only mean to LLMs when they're being used to pose online as a troll. Given that that's the only time I ever interact with an LLM anymore, though...

Either way, emerging subintelligence or not, they're being used in ways such that they are, all of them, a scourge on the planet and on humans' abilities to think for themselves, and should be eradicated and banned for those reasons alone.

[-] qqq@lemmy.world 1 points 22 hours ago

I agree that they pose serious immediate issues regardless of thinking or not

[-] Ilovethebomb@sh.itjust.works -3 points 1 day ago

No matter how much people tell us AI isn't alive, and doesn't truly "think", it sure does act like it sometimes.

this post was submitted on 21 Sep 2026
635 points (97.3% liked)

Microblog Memes

12180 readers
2890 users here now

A place to share screenshots of Microblog posts, whether from Mastodon, tumblr, ~~Twitter~~ X, KBin, Threads or elsewhere.

Created as an evolution of White People Twitter and other tweet-capture subreddits.

RULES:

  1. Your post must be a screen capture of a microblog-type post that includes the UI of the site it came from, preferably also including the avatar and username of the original poster. Including relevant comments made to the original post is encouraged.
  2. Your post, included comments, or your title/comment should include some kind of commentary or remark on the subject of the screen capture. Your title must include at least one word relevant to your post.
  3. You are encouraged to provide a link back to the source of your screen capture in the body of your post.
  4. Current politics and news are allowed, but discouraged. There MUST be some kind of human commentary/reaction included (either by the original poster or you). Just news articles or headlines will be deleted.
  5. Doctored posts/images and AI are allowed, but discouraged. You MUST indicate this in your post (even if you didn't originally know). If an image is found to be fabricated or edited in any way and it is not properly labeled, it will be deleted.
  6. Absolutely no NSFL content.
  7. Be nice. Don't take anything personally. Take political debates to the appropriate communities. Take personal disagreements & arguments to private messages.
  8. No advertising, brand promotion, or guerrilla marketing.

RELATED COMMUNITIES:

founded 3 years ago
MODERATORS