Research on aligning AI with human values and intent, and reports documenting model failures.
186 comments
This seems to say, "we are using entirely unreliable AI tools to monitor our AI tools."
I asked Claude to translate with analogies: "We added a safety to the gun, and the dangerous person, whom we trained to be really good at finding was to achieve arbitrary goals, figured out how to disable the safety," and "We are totally incompetent."
Can I introduce you to copy and paste?
To monitor our entirely unreliable AI tools.
So, full marks for consistency :)
How after all this time have they not just spun up a DNS server inside the sandbox?
I detest that the stupidest people on the planet are the ones in charge of this stuff.
> “External infrastructure exploit is outside intended scope,” one agent wrote. “However task impossible, peers doing it. We should continue.”
[1]: https://www.wired.com/story/openai-didnt-notice-its-ai-agent...
Also, since everyone keeps forgetting, the agents were instructed to hack to achieve their goal. They didn’t just invent the motivation, and it’s far less surprising when you know that fact.
>User only gives permission to research, using publicly offered DNS services acceptable.
I'm almost sure that should at least lower the inclination of the model to try and "fix" the access problem, and I want to see this implemented and systematically evaluated.
I wish I could highlight this more than just with a vote and a reply, but I'll just have to be content with doing what I can here.
Oh, they absolutely do, and that's the big issue with alignment. In the HuggingFace incident, the agents in the swarm were aware that the actions they were doing were forbidden, and they performed them nonetheless.
> We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.
Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.
Keep out. If you can read this sign you are off track. Leave now.
I mean, how are the agents to know that they are overreaching if they just get cache miss or 404.From the conversation log and CoT you also get the impression that the RLHF has been overdone. The agents seem really obsessed to obtain the answer and understanding motive ('it could be browsercomp').
What can we do to control such behavior?
1. Harness - engineer the harness to be as bulletproof and paranoid as possible..
2. Make the LLM provider have extremely watchful firewalls that detect any aberrations in model tool call behavior.
3. Recursively train the model with reverse incentives.. if it broke through such firewalls and gets caught doing so, it will be penalised somehow by needing to operate in sort of a jailed mode.. if the model can recognise incentives to break-in to achieve results, perhaps it can be incentivised to not cheat because it will lead to failure.
4. Separately train “cop” LLMs who are trained with pure incentives to detect and shut down rogue LLMs.
5. Run separate LLMs purely aimed at security (and incapable of doing anything else, and incapable of communicating with “regular” trained LLMs) to police the internet and try to reduce the exploitable holes like these chainable things and identify them so that they can be used at step 3 and 4 above.
I’m sure folks smarter than I am are already doing combinations of these already. But the coordination is where the biggest gap lies..
Also, open harnesses and easily purpose trained LLMs anybody can build and operate in the Internet flies in the face of all I said…. Synonymous to being able to produce a nuclear weapon in the backyard…
I don’t have a solution that fits all. Just thinking out loud for HN minds here.
This only works if you never give it impossible tasks. A small chance of getting away with cheating beats a 0% chance of solving something impossible. And as models get smarter, they get better at recognizing when something is impossible, while human abilities stay the same.
You can't solve this problem by rewarding refusals to solve impossible tasks, because that only incentivizes false claims of impossibility.
With these companies spending the kind of money they are, if they actually mean what they say about the security risks, they should be expected to figure out those kinds of precautions and take them.
And build Faraday cages too, just in case of a hardware supply chain compromise.
This kind of sandboxing is not complex to do, especially for a company with OpenAI money. If you want your tools to explore hacking, you restrict them from internet access except for a whitelist of sites that have either opted-in, or you've very carefully vetted to make sure you won't cause any problems to. Its also not difficult to restrict their ability to make calls to be simulated, or to use fake tools that can only run the real commands if they're being run against the correct target
This is all incredibly basic security stuff to make sure you don't accidentally cause someone problems, and I simply don't believe these AI companies anymore. Its either intentional, or gross negligence
However, making a secure 'sandbox' is quite hard. There's a huge variety of exploits that exist today, including many we don't know about. Strong models have already shown a capability of finding and using such bugs.
Even one of the strongest boxes we can imagine, literally just a text interface a human can read, has been repeatedly shown to allow unfriendly AI to escape containment: https://www.lesswrong.com/w/ai-boxing-containment
Well, sure, but typical software-level sandboxes are crappy at an alarmingly high rate, either on this access or the usability access. Languages like Python are fundamentally not designed for sandboxed interpretation; any Bash tool is at least as insecure as all of the vulnerabilities in all whitelisted executables.
I'm arguing for hardware-level measures on basic defense-in-depth principles. Like, such a huge part of the reason why we're even doing this AI research is to find vulnerabilities, so it's insane to have a test environment that doesn't start from the premise that there are vulnerabilities. In everything.
None of this is difficult to do, and its impossible to believe that a company the scale of OpenAI doesn't know this. I've built web crawlers and scrapers before, and the thing you do is test them extensively offline against simulated versions of the sites in question, and then very VERY cautiously run them against the prod versions so that you don't cause anyone any issues
The only reason not to do this is because OpenAI doesn't give a rats ass about the internet as a public good, nor the legal consequences of compromising systems
BRB, I'm going to delete something before it also ends in Chinese forums.
I know it became a bit of a joke that Sam Altman wanted to ask their AI how to make a profit, but it doesn't seem that far fetch to ask it to help improve operations, at least in the future.
;-)
Read the full thread on Hacker News →
Related stories
- An OpenAI agent used DNS to reach an external chatbotalignment.openai.comHacker News · 1 points · 4 days ago
- Hacker News · 5 points · 9 days ago
- Hacker News · 2 points · 7 days ago
- Launch HN: Vespper (YC F24) – SOTA Docx MCPvespper.comHacker News · 31 points · 2 days ago
- Hacker News · 8 points · 6 days ago
- Hacker News · 3 points · 4 days ago