Decision follows disclosures that OpenAI agents searching government websites had acted in unexpected ways

59 points•smb06•3 days ago•118 comments•

118 comments

dmix3 days ago
AFAIK all of these incidents happened when OpenAI contracted out to a company called Irregular (https://www.irregular.com/) to run these sandboxed CyberGym tests. They all happened around Mar-June and seem to be from the same collection of agent trials. Since then they already released Astra. Halting now is likely just a way to manage blowback.
johntb863 days ago
https://alignment.openai.com/misalignment-reports/an-agent-u... happened very recently, and that seems like it might be the reason they halted everything.
HDThoreaun3 days ago
Kinda weird that theyre testing its ability to dox people
verdverm3 days ago
I believe this is the case for the incidents minus HF, another HNer informed me as such when I made the same claim, that HF incident was wholly inhouse
bigyabai3 days ago
This should be the top comment on every one of these godforsaken posts. I don't want to see a single report about OpenAI hacking the UN until Sam Altman addresses the role Irregular played in these attacks. If he can't provide an honest postmortum concerning their business partners, then he's proving why nobody trusts him.
Topfi3 days ago
But it wasn’t just Irregular. Hugging Face, Medicare, etc. were OpenAI internal.
digitaltrees3 days ago
I think any argument that this is a cynical attempt at regulatory capture is destroyed by this; the economic incentives of releasing more capable models are too large. I might be persuaded that they are actually running out of money, and this is really just a cover for reducing burn..

I welcome this though, I think the models are smart enough for broad economic activity and we could spend a few years simply working to integrate them into workflows and letting society adjust. More intelligence isn't necessary for meaningful impact and the risks that are obvious and present and unsolved aren't worth the cost benefit analysis.

grogenaut3 days ago
I don't really get how you stop that though. How do you even quantify the intelligence of these models? Right now that's mostly by benchmarks. You could just make it fail benchmarks while do well at internal benchmarks.

You could also very likely optimize the current models to be a lot more energy efficient, that'd be a win even if they didn't become more intelligent.

But none of that stops anyone else from researching or improving their models. Or a nation you've told to fuck off from doing so as well.

As has been said in the past, one doesn't put the genie back in the bottle. We only really "stopped" researching nukes because there was diminishing returns. Though one could posit we stopped because simulation became good enough or a myriad of other reasons.

When you're talking about nation state weapons capabilities I don't think you stop, but ai is even easier to share than nukes, it's just a few gigs of numbers versus heavy dangerous materials that are very difficult to source and make. Every gamer, mac owner, etc has the equiv lent of a centrifuge on their desk. Not everyone has a centrifuge on their desk.

ncr1003 days ago
Yes, this is currently a crisis we are in. (At this point in time.)

Everybody remember MAD, mutually assured destruction? That too is a crisis.

We are in a crisis because we are having uncontrolled development and continuous rollout of a dangerous technology.

Well, relatively uncontrolled: the kind of lack of control is part of the crisis -- loss of gross human consensus around the progression of the technology. The corporate angle.

This is a crisis, folks.

prymitive3 days ago
That’s assuming nothing else stops them from deploying more capable models. What if they’ve got scaling issues and simply cannot deliver anything better? Saying that would be disastrous, saying instead “we’re choosing not to deliver” doesn’t trigger investors panic.
digitaltrees3 days ago
But if anyone else releases something they will lose market share. They may be lying but this is a public disclosure to investors. They risk securities fraud if they are manipulating the market with false information. This could block their IPO if they build a public record inconsistent with private actions. There are ways they could mitigate it by using cautious language like “we may reduce” but “stop all” is categorical and unequivocal.
swat5353 days ago
Really? You don’t think achieving regulatory capture far exceeds the gains you would get from releasing a newer model?

Their objective is to ban foreign and domestic competition (including open source) via heavy regulation and achieving a monopoly status.

Perhaps you are under the naive assumption that they are willing to compete in good faith?

digitaltrees1 day ago
Not in a global market.
verdverm3 days ago
are there still financial incentives to releasing mega models?

they cost a lot to run and people are picking smaller models more often because of bill blowouts

> the models are smart enough for broad economic activity and we could spend a few years simply working to integrate them into workflows and letting society adjus

this is the real Ai distillation, with Chinese characteristics (their playbook is broad deployment across their economy over having the best model)

mikert893 days ago
It seems like anthropic is far ahead of openai, and has no reports like this. We have to conclude this is a skill issue/engineering quality problem inside openai.

just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so

rolisz3 days ago
What do you mean? https://www.felonybench.com/

They're almost tied for felonies.

theptip3 days ago
Less bad, but https://www.anthropic.com/news/investigating-incidents-cyber...

In some sense though, sure, skill issue explains the gap vs. Anthropic’s much less severe alignment issues.

mikert893 days ago
im not sure why more people arent calling it out.
andsoitis3 days ago
> We have to conclude

That’s not the most parsimonious explanation even if the assumption it rests on (anthropic ahead of OpenAI) is true, which we don’t have proof of.

mikert893 days ago
i would say its industry consensus at this point. the creative output of the anthropic models is far ahead of openai. the benchmarks cannot capture the difference
ramraj073 days ago
Ive anecdotally heard that openai is far more chaotic, which includes not having a central infra team for example (or at least some teams not counting on depending on them). At least the previous hacks in openai were mainly due to bad infra architecture design.
mikert893 days ago
its likely they are trying to catch up to anthropic, and in doing so are trying riskier training runs.
dumberquestions3 days ago
>It seems like anthropic is far ahead of openai

We don't know what internal models look like, and any guesses about it are just speculation.

mikert893 days ago
the creative and "big picture understanding" of anthropic models are noticeably ahead of openai. external models are distilled representations of internal models, its clear who is ahead
hbarka3 days ago
digitaltrees3 days ago
The article you link to is wrong, it states "AI cannot think for itself, nor can it take independent actions." This is flawed reasoning, AI doesn't need to "think" in the way humans do to have autonomy. Go to codex or claude code or any harness right now, type a prompt and see if it executes a bash command or web search or file edit that you didn't tell it to, that is an autonomous plan and execution. If anything it's even more dangerous that thet can call drop db or kill pid without a user giving instructions.
macNchz3 days ago
> Go to codex or claude code or any harness right now

The harness is the whole thing here. AI generates text. Everything else is undertaken by harnesses and infrastructure humans provide, have control over, and therefore responsibility for.

Every action AI takes is fundamentally not independent, it requires an explicit choice to let the AI write code, have a physical machine to run it on, to have network access, etc. The concept that these things are "rogue" ignores the role humans play in giving them goals and tools to pursue those goals, and makes it seem like it’s a self-determined force, over which humans cannot exercise control at all.

jpnc3 days ago
>AI autonomy

>'type a prompt'

Which is it?

pizza2343 days ago
The article is dangerously misinformed. Detailed explanation here: https://news.ycombinator.com/item?id=49868681.
verdverm3 days ago
I would be skeptical of METR conclusions, they are within the circular funding loop of US Ai
m-s-y3 days ago
I firmly believe that this is just the public-facing story here.

Stopping AI development and research, even slowing it, would be a disaster for the SOTA companies and their first-mover advantage.

There’s almost no way to coordinate this across the world. Zero chance that everyone stops. We can’t even agree to coordinate on weapons tech that’s decades old with zero “everyday joe” impact.

ncr1003 days ago
There's a non-zero possibility that multiple concerns can be seemingly contradictory and simultaneously true, here.

Bad actors can be running the AI companies. Bad actors can also try to do good things. Good things can come from bad things. Bad things can come from good things. All of that's happening right now.

- It's very likely that we are not going to solve this complex crisis (compelling and harmful/deadly (still on target for 2030 AGI) AI Tech development) issues if we avoid trying to solve the challenge of building consensus across humanity around what kind of technology is too dangerous to uncontrollably develop plus roll out continuously

- The companies will be fine. Life matters more than business. I urge focusing on the life angle: regulation, political messaging, consensus building, looking for the best in humanity, protecting intelligent life.

Read the full thread on Hacker News →

Related stories