Nvidia says its new software could have prevented OpenAI's Hugging Face incident.

223 points•jonbaer•2 days ago•292 comments•

292 comments

Good4boothee1 day ago
Maybe we will skip few steps and add RGB lights to all GPU/NPU devices, that turn red when running "unaligned" code/model.
jameshart1 day ago
We can then just put those in the eyes of the humanoid robot models.
nwhnwh1 day ago
And it reports any incident to another robot that would search for you and put you in prison.
rf151 day ago
you mean the creator hasn't paid Nvidia for the green light?
functionmouse1 day ago
that's a good one
matja1 day ago
"It appears you've loaded weights into your GPU that have not been signed/approved by the government of the country your GPU is registered to..."
21asdffdsa121 day ago
You wouldn't download the worlds stolen knowledge..
prymitive1 day ago
Oh stop, microslop is probably already working on SecureTokenBoot or token2token encryption
alphawhisky1 day ago
I want R2D2 style blinkenlighten!
wavewrangler2 days ago
Did they try just properly sandboxing them first? Or are they still learning how to configure a firewall over there?

The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis

KingOfCoders2 days ago
Like in the Hugging Face hack. They deployed big surface, insecure app and gave AI access to it, then told AI do whatever it takes to fulfill this list. AI hacks insecure service, gets out, "the AI is at fault!" - no it's like running a bio lab with no protections and a virus gets out, then blame the virus for escaping.
js81 day ago
And HF actually tried to use AI to understand what's going on, but they had to use "unsafe" Chinese models since the "safe" ones have been castrated and refused to help. Great plan with the watchdog chip!
IanCal1 day ago
> then told AI do whatever it takes to fulfill this list.

That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.

People keep trying to frame this as

OpenAI: "Hack things, just really go for it"

Agent: hacks

OpenAI: shocked pikachu how could it hack?!?

But the reality is far from this.

Read the MTER report, it's fascinating. https://metr.org/hugging-face-incident-report-aug-2026.pdf

RataNova1 day ago
The application security really should be better across all levels. However the fact does not negate that the agent is already capable of spontaneously generating complex hacking chains without human involvement
radarsat11 day ago
I mean.. in this analogy, I'd both be blaming the company behind the virus and be trying to warn everyone about the danger of the escaped virus itself. So, it kind of fits.

In my reading, people aren't really saying "the AI is at fault", they are saying "hey look here's proof that this is dangerous". Like pointing at all the dead bodies caused by the virus and saying hey maybe we should stop making this virus.

copperx1 day ago
Then go on the news and spread panic that the virus is going to kill us all because it's sentient and impossible to contain.

Actually, the metaphor doesn't work at all because there are innumerable ways to shut down the entire thing during all phases including the made up "killing us all" bullshit scenario whereas with a virus there aren't any once a virus escapes containment.

Symmetry1 day ago
Stronger sandboxes trade off against how well they can trade the models, though. If you want your models to be looking things up and downloading tools from the internet when they're doing their job you need to provide at least a credible facsimile of the internet for their training environment and you can't fit something like that on a single airgapped server's storage.
If you genuinely believe there is even a 1% chance that your creation could destroy the planet or civilization, there is no excuse that is not fundamentally deranged and psychotic.

If you can't build it and test it securely, you should not be building it at all. To do it anyway is criminally psychopathic.

chaoz_1 day ago
pushing for chip-agenda as the best-isolation-layer immediately makes sense given their business
cpburns20091 day ago
Yes, Nvidia is proposing a two pronged approach. OpenShell is the software level sandbox. Sentry is the hardware level monitor.
pyronite1 day ago
> The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis

This is a very confident statement in the face of a purported non-0% chance of human extinction.

For what reasons do you disagree with the dangers of an intelligence explosion, e.g. Geoffrey Hinton and other experts in the field? https://www.theguardian.com/technology/2026/sep/28/ai-godfat...

I'm curious why you and others seem to write off the possibility so strongly. I would love to feel more confident.

voidhorse1 day ago
There's a difference between the current material risks (which OP correctly identifies reduce down to basic human incompetence) and the long term hypothetical risks (which is what Hinton is concerned about).

There are clear procedures for dealing with the immediate risk that have been known to the software industry for a long time. Don't let the companies use hypothetical risks as a smokescreen to hide their negligence.

saturn_vk1 day ago
A chip manufacturer proposes to sell more chips? Who would've guessed
olejorgenb1 day ago
Nvidia OpenShell is a (software) sandbox unless I'm mistaken.
ValueTheory2 days ago
Does this actually do anything other than give a permissions framework for developers who actually want to try to secure their systems?

Do you think the developers at Anthropic, OpenAI and Google who were so sloppy as to not put a good sandbox on their cybersecurity tests before will use this technology correctly? They are supposed to be the experts and they couldn't come up with something similar to this? I am not convinced this voluntary tool will change much of anything.

swozey2 days ago
Google actually practices zero-trust networks. Would love to see what they're seeing, or not seeing.
narrator1 day ago
Beyond Corp was and still is ahead of its time. No trusted internal network: access is granted per user, device, and service based on identity and policy, regardless of network location. Being in the office at Google is the same as being in a cybercafe anywhere on the planet.
jbs7891 day ago
Makes sense strategically for NVDA.

They are rightfully framing the problem as solvable. And this is one option.

hedora2 days ago
So, basically, the government (and, now Nvidia) wants to be able to kill switch all computers moving forward? (including stuff like vehicle and aeronautic control systems, cell phones, and cameras)

What could possibly go wrong?

gattr1 day ago
It might take a few more decades, but eventually we'll get to the point when you can fab fast enough general-purpose chips at home (or at local municipal makerspace), based off free designs.

Read the full thread on Hacker News →

Related stories