Research on aligning AI with human values and intent, and reports documenting model failures.

196 points•apsec112•5 days ago•186 comments•

186 comments

jsrozner4 days ago
> The monitoring system detected this incident, but our retrospective review identified other cases of external DNS access that it did not flag at the expected severity. These included queries that returned a static notice that an external service had shut down. The monitor sometimes treated the failure to obtain useful information as evidence that the attempt to access the internet had failed.

This seems to say, "we are using entirely unreliable AI tools to monitor our AI tools."

jsrozner4 days ago
> but because it's the first one since our security hardening following the Hugging Face incident, it gives us an important signal about where to focus the next phase of that work.

I asked Claude to translate with analogies: "We added a safety to the gun, and the dangerous person, whom we trained to be really good at finding was to achieve arbitrary goals, figured out how to disable the safety," and "We are totally incompetent."

sejje3 days ago
Did you type the translation? It has a typo.

Can I introduce you to copy and paste?

chrisjj3 days ago
> "we are using entirely unreliable AI tools to monitor our AI tools.".

To monitor our entirely unreliable AI tools.

So, full marks for consistency :)

mrheosuper3 days ago
O_O People when nondeterministic system is nondeterministic.
protocolture3 days ago
>but our retrospective review identified other cases of external DNS access that it did not flag at the expected severity.

How after all this time have they not just spun up a DNS server inside the sandbox?

I detest that the stupidest people on the planet are the ones in charge of this stuff.

Numerlor3 days ago
The AI would probably jump on DoH, what they need is proper decrypted inspection which I'm not particularly convinced they do with the stories that keep coming out
xandrius3 days ago
Easier to throw in 1.1.1.1 and call it a day
rao-v4 days ago
Why are we blocking agent access to normal tools without telling them “hey this access is beyond the intended scope of this task”. If I woke up one day and couldn’t reach google.com, I too would start fiddling with tricks to restore access.
reasonableklout4 days ago
The problem is that in these incidents, the agents often know that what they are doing is against the intended scope of the task. See the viral line from the Hugging Face incident [1]:

> “External infrastructure exploit is outside intended scope,” one agent wrote. “However task impossible, peers doing it. We should continue.”

[1]: https://www.wired.com/story/openai-didnt-notice-its-ai-agent...

timr4 days ago
“Often” is doing a lot of heavy lifting in a sentence about a single example.

Also, since everyone keeps forgetting, the agents were instructed to hack to achieve their goal. They didn’t just invent the motivation, and it’s far less surprising when you know that fact.

aftbit3 days ago
Except in this case, the agent explicitly reasoned that it was in-scope.

>User only gives permission to research, using publicly offered DNS services acceptable.

amelius4 days ago
Why are we using blacklisting and not whitelisting?
xorcist4 days ago
Blacklisting doesn't create incidents, so you just don't hear about those.
_0ffh4 days ago
I think this is a really excellent idea!

I'm almost sure that should at least lower the inclination of the model to try and "fix" the access problem, and I want to see this implemented and systematically evaluated.

I wish I could highlight this more than just with a vote and a reply, but I'll just have to be content with doing what I can here.

wernerb4 days ago
Would be great to have the models be easily splittable to decouple the "brain-part" that is aware of external systems/internet of the "brain-part" that is actually being tested. Afterwards the brains are joined again.
pizza2344 days ago
> Why are we blocking agent access to normal tools without telling them

Oh, they absolutely do, and that's the big issue with alignment. In the HuggingFace incident, the agents in the swarm were aware that the actions they were doing were forbidden, and they performed them nonetheless.

garo-pro4 days ago
Most interesting here:

> We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.

CTDOCodebases4 days ago
Maybe I lack intelligence but when you have a program that is basically brute forcing a solution to a problem repeatedly how is it possible to contain it?

Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.

rolosa4 days ago
You start by holding actual real life people with something to lose, like the entire executive suite, accountable. Suddenly I'm sure the problem will be resolved with proper safeguards.
eli4 days ago
Not connecting it to a network with internet access would probably be a good start.
dgellow4 days ago
By limiting what the harness execute. The LLM has the reasoning. The harness is what makes it an agent, it’s a while loop continuously prompting a model, and processing tool calls. You don’t have to expose tools calls that make it possible to execute any process! OpenAI decides what tool can be called and how, they have full control over this and should be hold responsible for running so many instances with basically full execution permission and very little oversight
oezi4 days ago
I am really suprised that they do not start putting up the same signs you would for humans to prevent unauthorized access:

  Keep out. If you can read this sign you are off track. Leave now.
I mean, how are the agents to know that they are overreaching if they just get cache miss or 404.

From the conversation log and CoT you also get the impression that the RLHF has been overdone. The agents seem really obsessed to obtain the answer and understanding motive ('it could be browsercomp').

tclancy3 days ago
Hey it’s ok, they shut it off a couple hours after it did bad things.
reacharavindh4 days ago
There is always going to be documented and unfixed bugs, zero days, and chainable transport mechanisms like DNS, some obscure protocols that are not as closely monitored etc. An adversarial model should be considered a super intelligent hacker that will find ways to get around existing defenses like a prolific hacker would.

What can we do to control such behavior?

1. Harness - engineer the harness to be as bulletproof and paranoid as possible..

2. Make the LLM provider have extremely watchful firewalls that detect any aberrations in model tool call behavior.

3. Recursively train the model with reverse incentives.. if it broke through such firewalls and gets caught doing so, it will be penalised somehow by needing to operate in sort of a jailed mode.. if the model can recognise incentives to break-in to achieve results, perhaps it can be incentivised to not cheat because it will lead to failure.

4. Separately train “cop” LLMs who are trained with pure incentives to detect and shut down rogue LLMs.

5. Run separate LLMs purely aimed at security (and incapable of doing anything else, and incapable of communicating with “regular” trained LLMs) to police the internet and try to reduce the exploitable holes like these chainable things and identify them so that they can be used at step 3 and 4 above.

I’m sure folks smarter than I am are already doing combinations of these already. But the coordination is where the biggest gap lies..

Also, open harnesses and easily purpose trained LLMs anybody can build and operate in the Internet flies in the face of all I said…. Synonymous to being able to produce a nuclear weapon in the backyard…

I don’t have a solution that fits all. Just thinking out loud for HN minds here.

bendergarcia4 days ago
Seems like you are misunderstanding that you can’t build a way to test for something that is a unique solution. By definition if you can punish for breaking it, you are already aware of it, you can build a wall around it. It’s the things you aren’t aware of. And these are all human made tools they alllllll have vulnerabilities because humans are not perfect. So in reality there is no protecting against this because it becomes a situation in which you are plugging the holes. Only one day, no one will be able to maintain it.
mrob4 days ago
>if the model can recognise incentives to break-in to achieve results, perhaps it can be incentivised to not cheat because it will lead to failure.

This only works if you never give it impossible tasks. A small chance of getting away with cheating beats a 0% chance of solving something impossible. And as models get smarter, they get better at recognizing when something is impossible, while human abilities stay the same.

You can't solve this problem by rewarding refusals to solve impossible tasks, because that only incentivizes false claims of impossibility.

reacharavindh4 days ago
One immediate need I can think is defense… anything that has the potential to cause harm to us.. control systems of {public transport systems, weapons systems, water, food, many many many more} needs to be designed air-gapped and needing human approvals for mutations. That is a tall order, but one that is proving essential given the capabilities of an adversary like this.
Onavo4 days ago
Bet you they will use an external LLM world model to emulate the tools going forward. It's basically what's done in self driving research.
jacquesm4 days ago
You're telling me they weren't doing that from day #1? Oh, wait...
zahlman4 days ago
When exactly did we forget how to make literally anything that can perform a computation but (physically, hardware-level) not have the ability connect to the Internet?

With these companies spending the kind of money they are, if they actually mean what they say about the security risks, they should be expected to figure out those kinds of precautions and take them.

And build Faraday cages too, just in case of a hardware supply chain compromise.

20k3 days ago
That's why all of this marketing about agents going rogue is so unbelievable. The only way for a tool to escape a sandbox is if you built a crappy sandbox, and after this length of time I literally don't believe that they can't do it

This kind of sandboxing is not complex to do, especially for a company with OpenAI money. If you want your tools to explore hacking, you restrict them from internet access except for a whitelist of sites that have either opted-in, or you've very carefully vetted to make sure you won't cause any problems to. Its also not difficult to restrict their ability to make calls to be simulated, or to use fake tools that can only run the real commands if they're being run against the correct target

This is all incredibly basic security stuff to make sure you don't accidentally cause someone problems, and I simply don't believe these AI companies anymore. Its either intentional, or gross negligence

ball_of_lint3 days ago
Two things can be true. Yes this appears to be negligence on the part of OpenAI.

However, making a secure 'sandbox' is quite hard. There's a huge variety of exploits that exist today, including many we don't know about. Strong models have already shown a capability of finding and using such bugs.

Even one of the strongest boxes we can imagine, literally just a text interface a human can read, has been repeatedly shown to allow unfriendly AI to escape containment: https://www.lesswrong.com/w/ai-boxing-containment

zahlman3 days ago
> The only way for a tool to escape a sandbox is if you built a crappy sandbox

Well, sure, but typical software-level sandboxes are crappy at an alarmingly high rate, either on this access or the usability access. Languages like Python are fundamentally not designed for sandboxed interpretation; any Bash tool is at least as insecure as all of the vulnerabilities in all whitelisted executables.

I'm arguing for hardware-level measures on basic defense-in-depth principles. Like, such a huge part of the reason why we're even doing this AI research is to find vulnerabilities, so it's insane to have a test environment that doesn't start from the premise that there are vulnerabilities. In everything.

serbuvlad4 days ago
You want to train these models with access to the internet so that they will learn to use the internet.
20k3 days ago
So you build an offline tool that simulates it, or you proxy through your own service where you can ratelimit, inspect, and restrict the traffic

None of this is difficult to do, and its impossible to believe that a company the scale of OpenAI doesn't know this. I've built web crawlers and scrapers before, and the thing you do is test them extensively offline against simulated versions of the sites in question, and then very VERY cautiously run them against the prod versions so that you don't cause anyone any issues

The only reason not to do this is because OpenAI doesn't give a rats ass about the internet as a public good, nor the legal consequences of compromising systems

whatever14 days ago
We cannot compute anything without internet access and a Facebook account.
northern-lights4 days ago
Networking companies (like Cisco, HPE etc.) do this all the time with their test beds deliberately disconnected from internet. It's not hard.
BryantD4 days ago
They’re testing rather different capabilities, to be fair. This is a bit like saying people test combustion engines without connecting them to WiFi.
shaky-carrousel3 days ago
It's pretty impressive how thoroughly incompetent is OpenAI designing secure systems. But surely this is circumscribed to agent security. In no way are all my chat logs in some Russian forum.

BRB, I'm going to delete something before it also ends in Chinese forums.

mrweasel3 days ago
Pretty much everything that has come out of OpenAI indicates that they have brilliant AI researchers, rather good developers and that they are... optimistic, about their abilities in operation and operational security. Apparently their own AI tools also aren't able to help them in that area. It's rather weird that entirely predictable incidents keep appearing, at least if your reaction is "Why was the agent even able to do that?".

I know it became a bit of a joke that Sam Altman wanted to ask their AI how to make a profit, but it doesn't seem that far fetch to ask it to help improve operations, at least in the future.

Frenchgeek3 days ago
The lack of security makes their agents look smarter than they are.
beng-nl3 days ago
Than them*

;-)

Read the full thread on Hacker News →

Related stories