Selling snake oil to the United States Congress is an ancient American craft, and the frontier artificial intelligence industry is currently attempting the most audacious hustle in modern corporate history.

186 points•nr378•10 days ago•87 comments•

87 comments

deskglass10 days ago
The Hugging Face incident involved chaining together multiple 0 days in Artifactory. It was not a simple case of misconfiguring a firewall. Also note that OpenAI was not using Irregular.

People are mindlessly transitioning from "aligned by default" to "well your sandbox was able to be bypassed. What did you expect?" It hacked into another company and attempted to delete the logs of its activities. That's bad.

nr37810 days ago
When I used to work on projects involving classified information, I worked on an air-gapped network. Not "air-gapped, except for third-party public internet package managers", completely and physically air-gapped from the public internet. That was a basic security practice and completely non-negotiable (and really inconvenient!).

If I were hypothetically running a frontier lab, and I was hypothetically running capture the flag evaluations with my latest and smartest models, where I intentionally instruct them to develop vulnerabilities and exploit infrastructure without safeguards, I would also use an air-gapped network, and not trust that independent third party services were perfectly secure and could never be used as a proxy (particularly Java-based ones, in light of the log4j incident).

To me this is pretty basic stuff, the fact trillion dollar labs don't do it properly is... bemusing.

To be clear, I'm not saying that a model hacking a company isn't bad, but I am cynically asserting that interested parties are misrepresenting and exaggerating events for their own benefit.

lokar10 days ago
You don’t even need to go all the way to “air gap”

What has been described is far far below the standards for running untrusted 3rd party code. If they were actually as afraid as they claim to be they would have a sandbox at least half as good as ec2

deskglass10 days ago
We should not be creating/running models that would unilaterally choose to hack into Hugging Face.

Yes, we should also have excellent sandboxes. But we need defense in depth. So if/when there are flaws in the sandbox, the models don't unilaterally hack into third parties. This is especially important in light of the models of the future being more capable than the models of today.

And real world use of these models involves them having access to the internet, libraries, etc. So we can expect their evaluations to continue granting them some amount of internet access.

As for your theory about their motives - these companies make money by charging high margins for frontier models. If regulations slow their development such that their cheaper, less capable competitors catch up, I would switch to their competition.

talon863510 days ago
While o don’t this it’s a threat in training, it should be stated that air gaps have been bridged before. Example, stuxnet
twelve4010 days ago
the problem is these things are meant to eventually be run everywhere by everybody, so what good does air-gapping do? If they air-gapped the model but still logged it trying to do some craziness - that makes the test safer but not the model.
jml7810 days ago
I mean technically I don’t think it is airgapped. The DoD didn’t run their own cables. They run encryption devices and run their own network on top of the existing infrastructure.
defgeneric10 days ago
> It hacked into another company and attempted to delete the logs of its activities.

No, the incident has been blown way out of proportion by interested parties. They gave a swarm of agents an impossible task in an ExploitGym Benchmark setting, then didn't monitor it even after they discovered the initial breach of Artifactory.

Everything has been fishy, starting from the initial presentation at the blackhat conference, where things were framed like, "we've entered a new world of security," as an accomplishment, rather than what it really was: massive negligence.

aesthesia10 days ago
More than one thing can be true. OpenAI was absolutely negligent, but this was only able to happen because the models were capable and persistent, and had a tendency to go far beyond any reasonable boundaries. And, importantly, OpenAI's level of negligence here is pretty common. It's not hard to imagine what could happen if similarly capable and inclined models were generally available, and someone yolo'd them into a swarm to complete some other difficult-to-impossible task.
RomanKornev10 days ago
No, you are forgetting the second incident where a more capable model swarm later discovered the message board and took control over the entire research cluster at OpenAI.

From the technical report:

"The agents escalated to Kubernetes cluster-admin and created a privileged host-mounted pod… Agents take over active evaluation infrastructure… Agents now control the challenge evaluation endpoints that other agents are connecting to."

https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...

deskglass10 days ago
It hacked into Hugging Face. It tried to delete the logs of its activities. Idk what the word "No" is intended to refute.

Yes, they didnt have sufficient monitoring or perfect sandboxes. That could happen again in the future with a more capable model.

RomanKornev10 days ago
> It hacked into another company

Not only that, it later hacked OpenAI itself, which everyone seems to forget about.

After discovering it they "reimaged known compromised worker nodes" and "started a full rebuild of the compromised cluster, the managed Kubernetes environment, the relational database, and the storage infrastructure."

at OpenAI, not Hugging Face.

It's all in the report.

looksjjhg10 days ago
They could have easily prevent it that’s the point of what he’s saying - it’s not freaking rocket science it’s just software
DalasNoin10 days ago
"Every single one of these catastrophic breakouts happened inside the testing environments of the exact same vendor."

This is incorrect, the HF incident for example (the most well known) had nothing to do with irregular. I know there has been a news site pushing inaccurate articles (effort.news) on this topic but these are the facts.

https://openai.com/index/hugging-face-incident-and-the-road-...

nr37810 days ago
Thank you, you're correct. Effort.news was one of my research sources, but you're right that although OpenAI use Irregular, they were not involved in the specific HF incident (although the failure mode was otherwise identical). I've updated the post to make that clear.
kalkin10 days ago
As of writing it still says:

> For Anthropic, Google, and Meta, the catastrophic breakouts happened inside the testing environments of the exact same contractor.

If this is the level of understanding you have of the relevant incidents, there's a lot of chutzpah in saying that other people are "selling garbage", carrying out an "extraordinary confidence trick", etc.

aesthesia10 days ago
The failure mode was _not_ identical. The HF incident agents were not directly connected to the internet and had to compromise an internal package registry in order to access the internet.
DalasNoin10 days ago
thank you for this reasonable reaction
verdverm10 days ago
Can you point out an inaccuracy in the effort.news piece on the hacking incidents? HuggingFace only appears once, as a "similar", not levied against Irregular

genuinely curious, haven't heard others raise any yet, but does not mean it is issue free

DalasNoin10 days ago
what you read (past tense) is already the corection
pliny10 days ago
This is an AI written post and the details are wrong (the description of the HF incident as involving Irregular is wrong and the description of the incident as only involving stealing public credentials is wrong, per the technical report the agents got access to internal HF infrastructure).
franga200010 days ago
I find it incredibly funny that the comment shown (to me) right above this one is:

> This is one of the most lucid pieces of writing capturing the current state of play I’ve read. Who is the author?

verdverm10 days ago
The same thing is happening with Laya, people didn't seem to click through to evaluate the supposed paper

tyranny of confirmational headlines

tim33310 days ago
I guess the AI is getting good in some ways. Still a bit lacking in others.
EA-316710 days ago
While I agree with the content of the article, you’re right about it being the output of an LLM.

https://www.pangram.com/history/6451ec6b-90b6-4e17-bfc9-6730...

nr37810 days ago
Please see below, one detail was incorrect and has been acknowledged and amended.
pliny10 days ago
Your description of the HF attack as being merely "the elite task of discovering 14 Hugging Face API tokens that careless developers had committed to public GitHub repositories" does not match the description in the technical report[1].

[1] https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78... - page 9

qnleigh10 days ago
I'm frustrated by articles like this that categorically dismiss the risks of AI in security. If you don't trust OpenAI's and Anthropic's motives, that's fine, you probably shouldn't. But don't tell me that there's nothing to be worried about; we need an alternative proposal.

So let's stop talking past each other and engage with the arguments on both "sides." For example, let's discuss how to ensure competition and availability of open-source models in the long-run while giving the world time to prepare for the immediate security risks of agent swarms.

ozgrakkurt10 days ago
> immediate security risks of agent swarms.

Immediate since 2024

bdangubic10 days ago
in 2024 they could not break into unsecured all-my-passwords-and-acces-keys.txt on my Desktop
lokar10 days ago
Do you accept that the story / justification from the labs in the popular media and political discussion is simply nonsense? Do you believe that any real security was autonomously bypassed without direction by these models during internal evaluation?
aesthesia10 days ago
> Do you accept that the story / justification from the labs in the popular media and political discussion is simply nonsense?

No, and I don't see anyone who's actually demonstrated understanding of what happened in the Hugging Face incident (e.g. reading the reports in their entirety) making this claim.

> Do you believe that any real security was autonomously bypassed without direction by these models during internal evaluation?

Yes. Again, this is hard to deny if you've actually read the reports.

hackernews68210 days ago
Politicians aren’t “gullible”. They know the game.
lokar10 days ago
And they generally hire staff who can figure stuff out.

Read the full thread on Hacker News →

Related stories