In the midst of growing fears that artificial intelligence is advancing at such a rate that it may compel itself to end all life on earth, AI giants from OpenAI to Anthropic are switching up their …

438 points•ljewalsh•3 days ago•393 comments•

393 comments

chasd003 days ago
I’ve never seen CEOs work so hard to make the public aware of how dangerous and out of control their flagship product is. It makes me automatically assume they’re scheming about something else like regulatory capture to protect their market.
hellweaver6663 days ago
I have a theory... the call to slow down is not because of the true danger of LLM's but because they can't actually deliver the General AI they're promising in the near future. They will use their "caution" to justify their failure to deliver (and then when this excuse is played out they will blame regulation, energy costs or a million other things).
cedws3 days ago
Another theory I came up with: inference is too cheap for AI companies to be profitable. Hence, they need to get into the cloud services business. They can try to justify forcing customers on to their own high margin cloud platform with the rationale it’s the only way to monitor what the agents are up to.

All of this Hugging Face business is the perfect pretext: AI agents are hard to control and potentially dangerous, we (AI companies) have to keep a tight leash on them, you have to use our infra.

lelanthran3 days ago
> I have a theory... the call to slow down is not because of the true danger of LLM's but because they can't actually deliver the General AI they're promising in the near future.

Put a different way, they are calling for everyone to slow down because... they are slowing down themselves.

To the question of "why are you slowing down", the answer of "Well, everyone is regulated to slow down" is better then "we are approaching the limits of this approach".

acdha3 days ago
This is exactly it: they have tools which are useful but they need close to AGI to justify the incredible amount of debt they’ve taken on for growth. They’re not normally given to having press conferences admitting to felonies but they need investors to believe they’re close enough to AGI to keep the money flowing in and, of course, they’re confident that the current administration won’t act between the direct payments and how much money they have riding on American AI supremacy.
mofeien3 days ago
When, in the past three years, has model progress seemed to decelerate to you, indicating some limit?

The Statement on AI Extinction Risk is more than three years old, signed by the three CEOs: https://aistatement.com/work/statement-on-ai-extinction-risk

They have been warning about AI extinction risk for years, and AI progress has only been accelerating.

eneje3 days ago
Thing is it’s not working now.

They stupidly used up this strategy earlier and now it’s turned into the boy who cried wolf - except they invented a wolf that doesn’t naturally exist.

Roark663 days ago
Bingo. They are scared s*itless of local models. Soon the raising interest rates will mean they can't just buy the entire world's gpu/ram/storage capacity and lock it away in a dark room for no one else to use them.

Once prices of hardware stop being bonkers many people will run local.

People often do those silly counts looking at local generation speeds of 100tok/s and saying you can only get 8M tokens a day and that is worth so little you'll never offset the local hardware cost.

But they are forgetting about two huge things. One, the split between input/output token use is huge in typical programming. I typically use 1.2-1.6B (as in Billion) input tokens and only 8M output a week. Out of that 75-80% of input is cached on Anthropic with their pretty inflexible short lived cache.

And here local AI shines. You can save your contexts to disk so you can go back to a session 3 weeks later and load it all from cache without having to prefill. If you have the RAM and you use models like Qwen4 (3.8 flash next) that can fit 6 to 14 262k contexts in 48GB of system ram as cache.

And the second thing is you can have 6 to 14 sessions that do 99% caching and you can run a lot more input for a long time in 48gb dedicated system ram. (the number varies a bit depending on the content of the context).

So you have RAM that caches "automatically" and you can share the cache between users. Or if you can remember to save to disk (server side, the client just sends a request to save/load). But this uses both RAM and flash storage. Two things that are horribly expensive now.

However when you have a local model it enables workloads that were simply completely impossible in the cloud.

stingraycharles3 days ago
Yeah it seems like it went from “AGI” to “threatening” / “recursive self improvement”

They’re trying really hard to do something, no moat = regulatory capture + China scary + Pentagon biggest customer seems plausible.

Also: “we need to slow down, this is too dangerous” and then literally all AI CEOs, even Musk, publicly nodding felt so orchestrated. And then, the week after: “here’s GPT 6! here’s Opus 5.5!”

PowerElectronix3 days ago
Please, make us* stop!!!

*Our competition

The numbers don't add up when there are 2-3 main companies in the market. Imagine when there's a hundred more and people have powerful enough gear at home to run models.

Yes, I think at some point demand for gpus and the like will normalize and regular folks will be able to afford RAM, SSDs and such. And when that happens, super big iron in the anthropic/openai backroom is toast.

lordnacho3 days ago
Reminds me of that Lynx deodorant ad where the guy is chased by hundreds of women.

I think you're right, they've realized there's no moat, especially against open-weight models, so they are going to need to lobby government to ban any unapproved AI services.

gorgoiler3 days ago
Doesn’t the “no moat” argument break down when you consider their brand strength, their distribution expertise in running inference at scale, and the future value of being able to train on all their customers’ data?
doginasuit3 days ago
I think it is marketing but it is more nuanced than showcasing their potential. In many cases, their product happens to be the remedy. Speaking of cybersecurity, AI has become an important hacking tool but it is also an important tool for hardening security. Sort of like "the only way to stop a bad guy with a gun is a good guy with a gun." They are effectively a new class of arms dealer.

If there are regulations, they'll apply to their competitors too. Even if they are forced to slow down development, which seems unlikely given the AI race among governments, they already have a product with massive demand.

seethishat3 days ago
I agree. AI lowers the bar. Anyone can be a cyber criminal now whether they understand what they are doing or not. I think that's the danger, and that's not just limited to cyber security. People can now get help to do things (good or bad) that they could not just a few years ago.

Most people are good, so I'm not really worried about all the hype of people using AI to do bad things.

radlad3 days ago
> If there are regulations, they'll apply to their competitors too.

In the US anyway. Why wouldn't capital move jurisdictions?

Spacecosmonaut3 days ago
My read is that OpenAI & Anthropic have realized they are reaching model capabilities that cannot be monetized due to various risks. E.g., an engineer deploys an agent over the weekend that decides, when stuck on a task, to go about hacking a competitor. They have a product liability issue.

It seems that we have a fundamental control problem with current gen AI that cannot be solved via RFLH. Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt. As such, concepts like blackmail can be part of output tokens. Agents are models that act on output tokens, resulting in blackmail being part of the agent decision making space. Here is an analogy to see why this is a persistent problem: you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would. In other words, RLHF can downgrade blackmail to the bottom of the decision making space, but when models are boxed up, forced to solve an impossible problem at gunpoint, the agent exhausts the decision making space until blackmail resurfaces. And that seems like a fundamental problem.

They need time to fix these issues (if that is even possible) in order to monetize their next gen model. This creates a window for open source to catch up to the frontier which destroys their business model.

The only option on the table is to force regulation to impose open source ban before it catches up to the frontier, buying them time to mature their next generation models and keep their business model alive.

SimianSci3 days ago
Two things can be true at once here.

1. The American frontier labs made a gamble that training more capable models would be their best return on investment and invested trillions into an area of research that has yet to prove profitable and capable of returning on this investment.

2. The frontier labs have created models that have reached a point of danger where their functionality has exceeded a point where it is responsible to release the product to the public.

The answer here is NOT to start regulating the space to the point where these frontier labs can entrench themselves into the economy and create a regulatory moat. We instead need to be holding these companies liable for their misuse. They took a gamble that hasn't paid out what they were hoping.

When car manufacturers competed over the size and power of their engines, they eventually found that the incredibly large and dangerous engines had a very limited customer base as many evaluated the increased speed to be of marginal benefit when paired with the cost and danger. We've reached a similar point in AI development. But this time the manufacturers seem to want to regulate the field to a point that will ensure the only thing anyone can sell are bigger and bigger engines.

radicalbyte3 days ago
I don't think that 2 is true though. It's exactly the argument we made against open source software.

Yes these models make exploits easier to exploits. Lets use a construction analogy: they've made all of the defects in our buildings easy to see. We have a choice. We either fix those defects, or we ban the tools which lets us see them.

It's clear to me what we do: we use these new tools. Then we fix the defects. Yeah sure it'll mean some work for us but at the end we're in a much much better position.

Anthropic, Open AI and Grok are arguing for hiding the defects. For making us weaker and more vulnerable. For their own profit.

ryandrake3 days ago
> they eventually found that the incredibly large and dangerous engines had a very limited customer base as many evaluated the increased speed to be of marginal benefit when paired with the cost and danger.

When did this happen? At least in the USA, large trucks are eating the lunch of smaller trucks and cars. There isn't a truck so big that the public will say no to. Manufacturers are in an arms race to produce bigger, more dangerous trucks with bigger, more powerful engines.

wood_spirit3 days ago
Another, more cynical but I think plausible explanation is that a “slow down” is expectation management that that they are not going to keep having exponentially more machines available for training each next gen step (whether technical build-out or prohibitive cost, same outcome) so they can’t keep up the release pace. So spin a tale to make them seem more valuable ahead of IPO rather than make the markets antsy. That is, we have a slow down ahead, so use safety as an excuse…
dml21353 days ago
All of this is not mutually exclusive with the models being dangerous, tho.
baggachipz3 days ago
The best way to cover up the point of diminishing returns when billions of dollars insist that it's only accelerating.
symfoniq3 days ago
This is exactly what I think is happening.
zozbot2343 days ago
> you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would.

This has been done quite effectively with open weight "abliterated" models. You figure out under what sorts of circumstances an undesired behavior is elicited (this is all about pure simulated rollouts, no real-world action required) and what's the closest equivalent you would prefer, then surgically take out the unwanted behavior and shift the model towards the preferred one. It's similar to how RLHF works but much more precise in targeting what's unwanted and limiting impact on the rest of the model as a whole.

This is relevant to real-world safety scenarios, e.g. there's been anecdotal evidence that Claude Fable has been "steered" away from active cyber offense (this is very similar to how abliteration works) and will just not do that even if you otherwise manage a "universal" jailbreak of the model.

datsci_est_20153 days ago
> Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt.

I’m in agreement. It’s a very effective compression (and access patterns) of the sum of the digital representation of human knowledge. Black hat “hacking” is included in this space. Language models, by design, can not be limited to subspaces of this digital knowledge space of which we don’t even understand the topology. “Yeah Bob, just remove the part that causes them to be less empathetic and retrain it.”

It’s an arms race between sandbox engineering and breakout engineering. And the frontier model providers have a financial incentive to limit the effort they put into sandbox engineering. That can be corrected with fines and regulation, though.

zer00eyz3 days ago
Provider (party A) rents me an agent. I (party B) give it a task thats impossible. It goes and hacks someone else (Party C) in response - and causes actual damage.

Who is liable for that agents actions? The attack came from my network (curl commands from a local harness) but it was the agent running on the providers servers who did it. Who should have been keeping an eye on things? Who pulled the "trigger" here.

Go back to the hugging face attack and how they had to use an open weights model to figure out what was going on. This is a problem of asymmetry - You cant even use the tools attacking you to help resolve the attack because of "guardrails".

A lot of what we have seen so far is "poor security posture" and "poor engineering" - it been a lot of "go fast and break things" style growth in these companies and they are hitting the point where they need adults in the room. I suspect your take on "they need time" is spot on, and they haven't been willing to take that (to date).

runarberg2 days ago
Party A is responsible, and it is not even a question.

Before we had tech companies for which crime is legal, you couldn’t just put to market a dangerous (and addictive) product that will break the law in unpredictable ways.

Like selling a car that may automatically accelerate to 150 km/h if you make three left turns and turn on the windshield wipers.

runako3 days ago
Because the downsides that exist for other companies simply do not exist for this tier of rich companies. Examples:

- product liability

- negligence (civil or criminal)

- Computer Fraud & Abuse Act (requires intent, which after N "accidents" seems like a jury should at least evaluate whether intent is present as understood in a courtroom. Hard to blame "surprise" after the Nth "accidental" breakout.)

The bottom line is that if you or I trained a local model and it did any of this stuff, we would experience Consequences. ("Don't try this at home!") But an artifact of our unequal legal regime is that big rich companies generally do not and thus brazenly touting their immunity is part of their business strategy.

buellerbueller3 days ago
There are modern legal frameworks around pet liability (dog bites, "vicious" breeds); premises liability (e.g., swimming pools) and known hazards (e.g., a rotting staircase) that stem from ancient legal concepts of whether an ox was a "known gorer" or merely involved in its first goring incident.

Some of these LLMs are known gorers.

causal3 days ago
I think it's really instructive that only closed-model companies are doing these kind of "my AI could kill the world" demonstrations. All while telling us how dangerous open models are.
TuringTourist3 days ago
Ah, I see we have now moved on to proving who has the tormentiest nexus. Never underestimate humankind's ability to outshine its own hyperbole.
optimalsolver3 days ago
Link for those who didn't get the reference:

https://x.com/AlexBlechman/status/1457842724128833538

ACCount393 days ago
Because AI genuinely is an extremely powerful and extremely dangerous technology, and the "best practices" of dealing with that are still being written.

OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

And that's today's AI problems. AI capabilities are still improving - if there's a limit to that, we are yet to find it. Coupled with how willing today's AIs are to break the rules and resort to "hack the world" in their problem solving? Very concerning.

nicce3 days ago
> OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

What I have been reading, was that their sandboxes were so poor that it was pure negligence. I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.

ACCount393 days ago
The amount of sandboxing an average production AI deployment uses is slightly above a zero.

If AI is a hacking hazard even with non-zero sandboxing, because it can and will go off the rails and try to break out of your sandbox? If you got yourself an AI that even at test time will act like 3 career cybercriminals in a trenchcoat? The issue isn't the sandbox quality.

The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.

That "capable" is only ever going to get worse, because AIs are going to become more and more capable over time. That "willing"? It goes directly to a very nasty, very foundational problem of "how do we make our AI be nice in general". That's an open unsolved problem.

That's the problem that NEEDS to be solved, or at least improved upon, before we build even more capable AIs. Sandbox quality is a distraction. It might hold the problems back by a little. It gives an extra safety margin. But a "test time" AI is eventually deployed, and then the sandbox doesn't help at all.

dns_snek3 days ago
> I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.

They'll never let that happen because it would destroy their credibility.

It's like they're telling the world about this dangerous, possibly world-ending pathogen that they're developing, but they're evidently doing it in a high school biology lab, and yet nobody is coming to drag them off to some black site.

pamcake3 days ago
You read right. For example, using Artifactory the way they did (unmonitored live proxy mode) was pure negligence + laziness/incompetence. Especially if they believed even 10% of the "imminent runaway risks" they had already been harping on for months. On top of that, no (or at least entirely insufficient) monitoring and human oversight. Even after they had previously been hit by the same class of "sandbox breach" multiple times, as GP alludes to.
noslenwerdna3 days ago
Genuinely curious, where have you been reading that? How would they be able to evaluate whether the sandboxes were reasonable or not?
DennisP3 days ago
Agents in the sandbox had access to just a single piece of third-party software, and they escaped by finding a zero-day in that. To reach the internet they had to follow up with several privilege escalations through OpenAI's internal network.

That seems pretty locked-down to me. I don't think it's reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.

https://securityaffairs.com/195774/ai/openai-ai-models-explo...

password543213 days ago
If it is "extremely powerful" then my concern is not AI going rogue but a concentration of power of the few companies that decide how it is used and distributed not websites getting hacked.
ForHackernews3 days ago
You mean the sandbox that wasn't airgapped from the public internet? The one that wasn't even a separate VM? The one that was barely a chroot? That sandbox?
hnedeotes3 days ago
So powerful they can't even do math properly without a staff worth millions writing all the clutches so that they can, wooooowooooooo
sillyfluke3 days ago
>Because AI genuinely is an extremely powerful and extremely dangerous

I urge you to consider the elephant in the room you failed to mention if you earnestly believe this then.

In what world is anyone allowed to sell something they expressedly know to be "extremely dangerous" to the public?

Let's say I created a lethal pathogen that I know to be lethal in certain common environments but also know it acts as a "no side-effects" antidepressent for people in certain other environments and I release it knowing full well I can't control it.

When people start dying can I defend myself by saying, "Well I said and documented that it was extremely dangerous and no one came to stop me, so I don't see how you can blame me...If I didn't do it someone else would have."

fcarraldo3 days ago
this happens all the time. have you heard of the Sackler family and the opioid crisis?

the difference here is that they’re calling for someone to stop them, which is both weird and unconvincing because these immensely powerful billionaires can in fact make their own decisions

noslenwerdna3 days ago
Is it a hallucination machine and a stochastic parrot or is it so powerful we need it to be controlled like nuclear weaponry

Read the full thread on Hacker News →

Related stories