Nvidia says its new software could have prevented OpenAI's Hugging Face incident.
292 comments
The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis
That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.
People keep trying to frame this as
OpenAI: "Hack things, just really go for it"
Agent: hacks
OpenAI: shocked pikachu how could it hack?!?
But the reality is far from this.
Read the MTER report, it's fascinating. https://metr.org/hugging-face-incident-report-aug-2026.pdf
In my reading, people aren't really saying "the AI is at fault", they are saying "hey look here's proof that this is dangerous". Like pointing at all the dead bodies caused by the virus and saying hey maybe we should stop making this virus.
Actually, the metaphor doesn't work at all because there are innumerable ways to shut down the entire thing during all phases including the made up "killing us all" bullshit scenario whereas with a virus there aren't any once a virus escapes containment.
If you can't build it and test it securely, you should not be building it at all. To do it anyway is criminally psychopathic.
This is a very confident statement in the face of a purported non-0% chance of human extinction.
For what reasons do you disagree with the dangers of an intelligence explosion, e.g. Geoffrey Hinton and other experts in the field? https://www.theguardian.com/technology/2026/sep/28/ai-godfat...
I'm curious why you and others seem to write off the possibility so strongly. I would love to feel more confident.
There are clear procedures for dealing with the immediate risk that have been known to the software industry for a long time. Don't let the companies use hypothetical risks as a smokescreen to hide their negligence.
Do you think the developers at Anthropic, OpenAI and Google who were so sloppy as to not put a good sandbox on their cybersecurity tests before will use this technology correctly? They are supposed to be the experts and they couldn't come up with something similar to this? I am not convinced this voluntary tool will change much of anything.
They are rightfully framing the problem as solvable. And this is one option.
What could possibly go wrong?
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 3 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 5 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 10 days ago
- Hacker News · 73 points · 9 days ago
- The Verge · 0 points · about 11 hours ago