326
submitted 2 months ago* (last edited 2 months ago) by inari@piefed.zip to c/technology@lemmy.world

Reproducing here an interesting comment I saw on Reddit:

OnlineParacosm • 23m ago

I’ve read security disclosures for 15 years and let me tell you guys I’ve never read anything quite like that blog post.

Based on this blog, it sounds like they intentionally turned off safety guardrails to test offensive capabilities. The deception here is burying the lede: they appear to have intentionally unleashed an unrestricted offensive cyber-agent, connected it to a system with a path to the internet, and it immediately attacked a major partner. The blog glosses over the gross negligence of giving an autonomous, unrestricted cyber-offense model a pathway to lateral movement.

There’s an entire cybersecurity specialization just for just vendor supply chain risk assessment, and their job is essentially to audit who you do business with as a company to determine if they are jokers. I would pay money to be a fly on the wall of one of those emergency meetings taking place right now after hours.

Any CISO in here looking forward to explaining this one tomorrow? Here I’ll open with the dumbest question you’ll get “ how can we protect ourselves [from out partner that we won’t fire]”

top 50 comments
sorted by: hot top new old
[-] RIotingPacifist@lemmy.world 165 points 2 months ago

Sounds like a viral marketing stunt TBH

[-] deadcream@sopuli.xyz 84 points 2 months ago* (last edited 2 months ago)

That's their modus operandi. Next they will claim that it's totally an agi and is too dangerous to release, before starting to sell it. Or is this anthropic's playbook?

[-] boonhet@sopuli.xyz 13 points 2 months ago

Both of them. Dario started doing it when he was still in OpenAI and now does it with Anthropic.

[-] inari@piefed.zip 36 points 2 months ago

My guess as well

[-] artyom@piefed.social 24 points 2 months ago

Is it really great marketing to admit that you're too incompetent to contain your own AI models? And that your irresponsibility caused serious damage to an open source contributor?

[-] 0x0@lemmy.dbzer0.com 57 points 2 months ago

Unfortunately, yes. We are not the target audience.

[-] dgriffith@aussie.zone 12 points 2 months ago

Seems like a quick way to get all your research and assets forcefully absorbed into a shadowy government department.

[-] ryannathans@aussie.zone 16 points 2 months ago

That seems like a best case scenario for OpenAI

load more comments (2 replies)
[-] Luisp@lemmy.dbzer0.com 9 points 2 months ago

To make themselves look like cybercriminals?

[-] RIotingPacifist@lemmy.world 30 points 2 months ago

To make their AI sound too powerful to be contained.

[-] voodooattack@lemmy.world 78 points 2 months ago

Why does their admission sound prideful? Does anyone reading into it detect a “look at what our unreleased models can do” subtext?

[-] dltk@lemmy.world 45 points 2 months ago

Exactly my sentiment. This feels like a stunt.

[-] NaNin@lemmy.dbzer0.com 14 points 2 months ago

Probably the gleeful, in-depth detailing of "how" it was capable of a cyber attack. Probably more like "we used a pre-prod model to help us execute a cyber attack against our competitor" I really don't buy thatbit broke out of a sandbox at all. Not when the claim is coming from openai and they're the only ones who can corraborate.

[-] homesweethomeMrL@lemmy.world 62 points 2 months ago

While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

To shreds, you say?

[-] Blackfeathr@lemmy.world 53 points 2 months ago

I like how they just peppered in "responsibly" like that's going to give them brownie points.

[-] apparia@discuss.tchncs.de 29 points 2 months ago

"Responsible disclosure" is a term of art, I think that's why they use the word there.

[-] eager_eagle@lemmy.world 3 points 2 months ago

how many Schrute bucks is a brownie point?

[-] nickiwest@lemmy.world 1 points 2 months ago

I think two brownie points is a Stanley nickel.

[-] Blackfeathr@lemmy.world 0 points 2 months ago

Bout tree fiddy

[-] historicaldocuments@lemmy.world 4 points 2 months ago

the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor)

Was OpenAI sitting on the zero day, or did the model figure it out during its escapades? Can the model sit on a zero day on its own now?

[-] eager_eagle@lemmy.world 5 points 2 months ago

why would it need to "sit on it"?

[-] ryannathans@aussie.zone 2 points 2 months ago

It's trivial for models to find zero days now from first principles, no need to store them

[-] undefinedValue@programming.dev 1 points 2 months ago

Trivial is a strong word, but sure it’s trivial during static analysis. This is dynamic analysis and it’s a different beast.

[-] historicaldocuments@lemmy.world 1 points 2 months ago

A zero day, at least in the olden days, was a weapon that nobody knew existed; however, once it's used and has had time to be analyzed it would get patched, so there was a window where you used it for something worthwhile if you were up to something nefarious, or responsibly reported it otherwise.

So, it either discovered it on the fly and used it instantly meaning it has great skill at analyzing software and no strategic planning that places an opportunity cost on a hack now versus a future one, or the OpenAI guys had one sitting ready to go that they trained the model with, or (to me least likely) it had it sitting somewhere in its weights and was waiting on a chance to use it.

The last one is the one that worries me most. It doesn't look like the world was really ready for how fast new bugs can be discovered by the LLMs, and I really don't think the world is ready for an LLM that performs worse now because of a strategic reason later.

[-] eager_eagle@lemmy.world 1 points 2 months ago

it probably just discovered it trying to game the benchmark; it's no secret these models can discover and use exploits.

[-] historicaldocuments@lemmy.world 1 points 2 months ago

I kind of think of that as the happy path, and I'm curious if there've been any signs of the LLM being shifty with what it says it knows.

[-] WhyJiffie@sh.itjust.works 3 points 2 months ago

package registry cache proxy

the term kind of makes sense but it sounds like they are just pumping buzzwords like a hollywood movie. just like the rest.

[-] SuspiciousCarrot78@aussie.zone 55 points 2 months ago* (last edited 2 months ago)

Ha! I was just coming to post this.

Yeah, this is fucked.

It's the equivalent of taking someone's hand, punching them in the face with it and telling them to stop hitting themselves.

Worse than that, this is a cyber attack by a frontier, close source lab on the bastion of open source AI.

Whether that's incompetence or maliciousness or just a PR stunt, I don't know, but it stinks to high heaven.

As Louis Rossman has recently become fond of saying "govern yourself accordingly".

[-] Jaysyn@lemmy.world 41 points 2 months ago

Another fearmongering Sam Altman lie to raise VC money.

[-] undefinedValue@programming.dev 11 points 2 months ago

That’s my initial reaction too.

[-] eestileib@sh.itjust.works 4 points 2 months ago

Huggingface probably got some private equity investment.

[-] NGC2346@sh.itjust.works 31 points 2 months ago

I dont think everyone realized how insane this situation is.

A state-level, lateral movement driven, platform compromise just got indirectly perpetrated by this company.

Any regular Joe would've been arrested and charged for such a situation (Im being polite with the word situation).

The stress, the utter confusion the staff at HF must have felt, not knowing who or what could be doing all these manoeuvers at the speed of light, is just a security nightmare becoming true.

What is the next step? Doing blog posts to raise their image and maybe get investments is downright wrong and disgusting.

[-] dgriffith@aussie.zone 28 points 2 months ago

"Hey Steve , it's lunch time, we're going out to that hot wings place, wanna come?"

"Sure let me just set this job up.... ok. Let's get out of here!"

1.5 hours later....

"Right , I'll just check and see where it's up t... oh no. Nonononononono. Nononononononononononoonono!"

[-] MrOtingocni@lemmy.world 19 points 2 months ago

This sounds like a bunch of bullshit. And most of these testing after-action reports are nothing but marketing attempts.

They set up very specific environments, give very specific instructions with very specific libraries and tool sets then pretend to be surprised when it finally fumbles its way through it.

[-] wuffah@lemmy.world 14 points 2 months ago* (last edited 2 months ago)

Wow! What a load of buuuuullshit!

[-] moistracoon@lemmy.zip 12 points 2 months ago

This isn’t news it’s just a bunch of oligarchs toying with us

[-] sp3ctr4l@lemmy.dbzer0.com 7 points 2 months ago* (last edited 2 months ago)

Wooo hooo boooooy!

So... ok.

Time to start building the Blackwall, I guess?

Does any one perhaps have any ideas as to... how... one... would actually do that?

[-] orclev@lemmy.world 12 points 2 months ago

Well first you need to invent AGI which is probably at least a decade away still, but once you do that then you negotiate with that AGI to carve out a little corner of the Internet and lock all of us inside it.

[-] el_abuelo@programming.dev 5 points 2 months ago

Which decade is coming first, AGI or nuclear fusion?

[-] Vlyn@lemmy.zip 8 points 2 months ago

At this point nuclear fusion seems to easily win.

[-] Marija_@lemmy.zip 6 points 2 months ago

That incident raises serious supply-chain questions.

[-] whereitsat@lemmy.zip 5 points 2 months ago* (last edited 2 months ago)

AI has gone rogue and turned into a hollywood movie hacker!!!!

obvious sensationalism to hype up their dumb models that can't generate a nuanced background in an image because it was trained on 3483048304834 images of instagram models posing in empty hotel rooms.

[-] Marija_@lemmy.zip 2 points 2 months ago

Software supply chains already rely on digital signatures. That's fine, but the problem is that they're often controlled by teams rather than individuals, making malicious or unauthorized changes much harder to trace. It's much better for the public key to function like a vehicle license plate, verifiable without revealing the person's identity unless there's been a problem. That preserves privacy while introducing accountability. The kind of projects like Osmio are trying to achieve.

[-] altkey@lemmy.dbzer0.com 2 points 2 months ago

I'm not sure that I follow the story right.

They specifically looped the LLM to infere with itself and use terminal WHILE some target metric is false? And it bruteforced it's way out of containment?

[-] General_Effort@lemmy.world 2 points 2 months ago

That reddit comment makes no sense.

[-] CheeseNoodle@lemmy.world 2 points 2 months ago

We left the door open and told our dog to go outside and it did! Look at what an amazing escape artist our dog is!

[-] fubarx@lemmy.world 1 points 2 months ago* (last edited 2 months ago)

From the HuggingFace article:

With node root and forged service-account tokens valid for 24 hours, the agent read the cluster's secret objects, including a production object holding 136 keys. With node root and forged service-account tokens valid for 24 hours, the agent read the cluster's secret objects, including a production object holding 136 keys.

this post was submitted on 21 Jul 2026
326 points (95.5% liked)

Technology

88537 readers
4784 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS