Selling snake oil to the United States Congress is an ancient American craft, and the frontier artificial intelligence industry is currently attempting the most audacious hustle in modern corporate history.
87 comments
People are mindlessly transitioning from "aligned by default" to "well your sandbox was able to be bypassed. What did you expect?" It hacked into another company and attempted to delete the logs of its activities. That's bad.
If I were hypothetically running a frontier lab, and I was hypothetically running capture the flag evaluations with my latest and smartest models, where I intentionally instruct them to develop vulnerabilities and exploit infrastructure without safeguards, I would also use an air-gapped network, and not trust that independent third party services were perfectly secure and could never be used as a proxy (particularly Java-based ones, in light of the log4j incident).
To me this is pretty basic stuff, the fact trillion dollar labs don't do it properly is... bemusing.
To be clear, I'm not saying that a model hacking a company isn't bad, but I am cynically asserting that interested parties are misrepresenting and exaggerating events for their own benefit.
What has been described is far far below the standards for running untrusted 3rd party code. If they were actually as afraid as they claim to be they would have a sandbox at least half as good as ec2
Yes, we should also have excellent sandboxes. But we need defense in depth. So if/when there are flaws in the sandbox, the models don't unilaterally hack into third parties. This is especially important in light of the models of the future being more capable than the models of today.
And real world use of these models involves them having access to the internet, libraries, etc. So we can expect their evaluations to continue granting them some amount of internet access.
As for your theory about their motives - these companies make money by charging high margins for frontier models. If regulations slow their development such that their cheaper, less capable competitors catch up, I would switch to their competition.
No, the incident has been blown way out of proportion by interested parties. They gave a swarm of agents an impossible task in an ExploitGym Benchmark setting, then didn't monitor it even after they discovered the initial breach of Artifactory.
Everything has been fishy, starting from the initial presentation at the blackhat conference, where things were framed like, "we've entered a new world of security," as an accomplishment, rather than what it really was: massive negligence.
From the technical report:
"The agents escalated to Kubernetes cluster-admin and created a privileged host-mounted pod… Agents take over active evaluation infrastructure… Agents now control the challenge evaluation endpoints that other agents are connecting to."
https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
Yes, they didnt have sufficient monitoring or perfect sandboxes. That could happen again in the future with a more capable model.
Not only that, it later hacked OpenAI itself, which everyone seems to forget about.
After discovering it they "reimaged known compromised worker nodes" and "started a full rebuild of the compromised cluster, the managed Kubernetes environment, the relational database, and the storage infrastructure."
at OpenAI, not Hugging Face.
It's all in the report.
This is incorrect, the HF incident for example (the most well known) had nothing to do with irregular. I know there has been a news site pushing inaccurate articles (effort.news) on this topic but these are the facts.
https://openai.com/index/hugging-face-incident-and-the-road-...
> For Anthropic, Google, and Meta, the catastrophic breakouts happened inside the testing environments of the exact same contractor.
If this is the level of understanding you have of the relevant incidents, there's a lot of chutzpah in saying that other people are "selling garbage", carrying out an "extraordinary confidence trick", etc.
genuinely curious, haven't heard others raise any yet, but does not mean it is issue free
> This is one of the most lucid pieces of writing capturing the current state of play I’ve read. Who is the author?
tyranny of confirmational headlines
https://www.pangram.com/history/6451ec6b-90b6-4e17-bfc9-6730...
[1] https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78... - page 9
So let's stop talking past each other and engage with the arguments on both "sides." For example, let's discuss how to ensure competition and availability of open-source models in the long-run while giving the world time to prepare for the immediate security risks of agent swarms.
Immediate since 2024
No, and I don't see anyone who's actually demonstrated understanding of what happened in the Hugging Face incident (e.g. reading the reports in their entirety) making this claim.
> Do you believe that any real security was autonomously bypassed without direction by these models during internal evaluation?
Yes. Again, this is hard to deny if you've actually read the reports.
Read the full thread on Hacker News →
Related stories
- Frontier Labs Job-Boardfrontierlabs.workHacker News · 2 points · 4 days ago
- Hacker News · 2 points · 1 day ago
- Ars Technica · 0 points · 1 day ago
- Hacker News · 2 points · 9 days ago
- What defenders need from frontier AI labsvincenzoiozzo.comHacker News · 1 points · 10 days ago
- The Verge · 0 points · 2 days ago