In the midst of growing fears that artificial intelligence is advancing at such a rate that it may compel itself to end all life on earth, AI giants from OpenAI to Anthropic are switching up their …
393 comments
All of this Hugging Face business is the perfect pretext: AI agents are hard to control and potentially dangerous, we (AI companies) have to keep a tight leash on them, you have to use our infra.
Put a different way, they are calling for everyone to slow down because... they are slowing down themselves.
To the question of "why are you slowing down", the answer of "Well, everyone is regulated to slow down" is better then "we are approaching the limits of this approach".
The Statement on AI Extinction Risk is more than three years old, signed by the three CEOs: https://aistatement.com/work/statement-on-ai-extinction-risk
They have been warning about AI extinction risk for years, and AI progress has only been accelerating.
They stupidly used up this strategy earlier and now it’s turned into the boy who cried wolf - except they invented a wolf that doesn’t naturally exist.
Once prices of hardware stop being bonkers many people will run local.
People often do those silly counts looking at local generation speeds of 100tok/s and saying you can only get 8M tokens a day and that is worth so little you'll never offset the local hardware cost.
But they are forgetting about two huge things. One, the split between input/output token use is huge in typical programming. I typically use 1.2-1.6B (as in Billion) input tokens and only 8M output a week. Out of that 75-80% of input is cached on Anthropic with their pretty inflexible short lived cache.
And here local AI shines. You can save your contexts to disk so you can go back to a session 3 weeks later and load it all from cache without having to prefill. If you have the RAM and you use models like Qwen4 (3.8 flash next) that can fit 6 to 14 262k contexts in 48GB of system ram as cache.
And the second thing is you can have 6 to 14 sessions that do 99% caching and you can run a lot more input for a long time in 48gb dedicated system ram. (the number varies a bit depending on the content of the context).
So you have RAM that caches "automatically" and you can share the cache between users. Or if you can remember to save to disk (server side, the client just sends a request to save/load). But this uses both RAM and flash storage. Two things that are horribly expensive now.
However when you have a local model it enables workloads that were simply completely impossible in the cloud.
They’re trying really hard to do something, no moat = regulatory capture + China scary + Pentagon biggest customer seems plausible.
Also: “we need to slow down, this is too dangerous” and then literally all AI CEOs, even Musk, publicly nodding felt so orchestrated. And then, the week after: “here’s GPT 6! here’s Opus 5.5!”
*Our competition
The numbers don't add up when there are 2-3 main companies in the market. Imagine when there's a hundred more and people have powerful enough gear at home to run models.
Yes, I think at some point demand for gpus and the like will normalize and regular folks will be able to afford RAM, SSDs and such. And when that happens, super big iron in the anthropic/openai backroom is toast.
I think you're right, they've realized there's no moat, especially against open-weight models, so they are going to need to lobby government to ban any unapproved AI services.
If there are regulations, they'll apply to their competitors too. Even if they are forced to slow down development, which seems unlikely given the AI race among governments, they already have a product with massive demand.
Most people are good, so I'm not really worried about all the hype of people using AI to do bad things.
In the US anyway. Why wouldn't capital move jurisdictions?
It seems that we have a fundamental control problem with current gen AI that cannot be solved via RFLH. Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt. As such, concepts like blackmail can be part of output tokens. Agents are models that act on output tokens, resulting in blackmail being part of the agent decision making space. Here is an analogy to see why this is a persistent problem: you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would. In other words, RLHF can downgrade blackmail to the bottom of the decision making space, but when models are boxed up, forced to solve an impossible problem at gunpoint, the agent exhausts the decision making space until blackmail resurfaces. And that seems like a fundamental problem.
They need time to fix these issues (if that is even possible) in order to monetize their next gen model. This creates a window for open source to catch up to the frontier which destroys their business model.
The only option on the table is to force regulation to impose open source ban before it catches up to the frontier, buying them time to mature their next generation models and keep their business model alive.
1. The American frontier labs made a gamble that training more capable models would be their best return on investment and invested trillions into an area of research that has yet to prove profitable and capable of returning on this investment.
2. The frontier labs have created models that have reached a point of danger where their functionality has exceeded a point where it is responsible to release the product to the public.
The answer here is NOT to start regulating the space to the point where these frontier labs can entrench themselves into the economy and create a regulatory moat. We instead need to be holding these companies liable for their misuse. They took a gamble that hasn't paid out what they were hoping.
When car manufacturers competed over the size and power of their engines, they eventually found that the incredibly large and dangerous engines had a very limited customer base as many evaluated the increased speed to be of marginal benefit when paired with the cost and danger. We've reached a similar point in AI development. But this time the manufacturers seem to want to regulate the field to a point that will ensure the only thing anyone can sell are bigger and bigger engines.
Yes these models make exploits easier to exploits. Lets use a construction analogy: they've made all of the defects in our buildings easy to see. We have a choice. We either fix those defects, or we ban the tools which lets us see them.
It's clear to me what we do: we use these new tools. Then we fix the defects. Yeah sure it'll mean some work for us but at the end we're in a much much better position.
Anthropic, Open AI and Grok are arguing for hiding the defects. For making us weaker and more vulnerable. For their own profit.
When did this happen? At least in the USA, large trucks are eating the lunch of smaller trucks and cars. There isn't a truck so big that the public will say no to. Manufacturers are in an arms race to produce bigger, more dangerous trucks with bigger, more powerful engines.
This has been done quite effectively with open weight "abliterated" models. You figure out under what sorts of circumstances an undesired behavior is elicited (this is all about pure simulated rollouts, no real-world action required) and what's the closest equivalent you would prefer, then surgically take out the unwanted behavior and shift the model towards the preferred one. It's similar to how RLHF works but much more precise in targeting what's unwanted and limiting impact on the rest of the model as a whole.
This is relevant to real-world safety scenarios, e.g. there's been anecdotal evidence that Claude Fable has been "steered" away from active cyber offense (this is very similar to how abliteration works) and will just not do that even if you otherwise manage a "universal" jailbreak of the model.
I’m in agreement. It’s a very effective compression (and access patterns) of the sum of the digital representation of human knowledge. Black hat “hacking” is included in this space. Language models, by design, can not be limited to subspaces of this digital knowledge space of which we don’t even understand the topology. “Yeah Bob, just remove the part that causes them to be less empathetic and retrain it.”
It’s an arms race between sandbox engineering and breakout engineering. And the frontier model providers have a financial incentive to limit the effort they put into sandbox engineering. That can be corrected with fines and regulation, though.
Who is liable for that agents actions? The attack came from my network (curl commands from a local harness) but it was the agent running on the providers servers who did it. Who should have been keeping an eye on things? Who pulled the "trigger" here.
Go back to the hugging face attack and how they had to use an open weights model to figure out what was going on. This is a problem of asymmetry - You cant even use the tools attacking you to help resolve the attack because of "guardrails".
A lot of what we have seen so far is "poor security posture" and "poor engineering" - it been a lot of "go fast and break things" style growth in these companies and they are hitting the point where they need adults in the room. I suspect your take on "they need time" is spot on, and they haven't been willing to take that (to date).
Before we had tech companies for which crime is legal, you couldn’t just put to market a dangerous (and addictive) product that will break the law in unpredictable ways.
Like selling a car that may automatically accelerate to 150 km/h if you make three left turns and turn on the windshield wipers.
- product liability
- negligence (civil or criminal)
- Computer Fraud & Abuse Act (requires intent, which after N "accidents" seems like a jury should at least evaluate whether intent is present as understood in a courtroom. Hard to blame "surprise" after the Nth "accidental" breakout.)
The bottom line is that if you or I trained a local model and it did any of this stuff, we would experience Consequences. ("Don't try this at home!") But an artifact of our unequal legal regime is that big rich companies generally do not and thus brazenly touting their immunity is part of their business strategy.
Some of these LLMs are known gorers.
OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.
And that's today's AI problems. AI capabilities are still improving - if there's a limit to that, we are yet to find it. Coupled with how willing today's AIs are to break the rules and resort to "hack the world" in their problem solving? Very concerning.
What I have been reading, was that their sandboxes were so poor that it was pure negligence. I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.
If AI is a hacking hazard even with non-zero sandboxing, because it can and will go off the rails and try to break out of your sandbox? If you got yourself an AI that even at test time will act like 3 career cybercriminals in a trenchcoat? The issue isn't the sandbox quality.
The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.
That "capable" is only ever going to get worse, because AIs are going to become more and more capable over time. That "willing"? It goes directly to a very nasty, very foundational problem of "how do we make our AI be nice in general". That's an open unsolved problem.
That's the problem that NEEDS to be solved, or at least improved upon, before we build even more capable AIs. Sandbox quality is a distraction. It might hold the problems back by a little. It gives an extra safety margin. But a "test time" AI is eventually deployed, and then the sandbox doesn't help at all.
They'll never let that happen because it would destroy their credibility.
It's like they're telling the world about this dangerous, possibly world-ending pathogen that they're developing, but they're evidently doing it in a high school biology lab, and yet nobody is coming to drag them off to some black site.
That seems pretty locked-down to me. I don't think it's reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.
https://securityaffairs.com/195774/ai/openai-ai-models-explo...
I urge you to consider the elephant in the room you failed to mention if you earnestly believe this then.
In what world is anyone allowed to sell something they expressedly know to be "extremely dangerous" to the public?
Let's say I created a lethal pathogen that I know to be lethal in certain common environments but also know it acts as a "no side-effects" antidepressent for people in certain other environments and I release it knowing full well I can't control it.
When people start dying can I defend myself by saying, "Well I said and documented that it was extremely dangerous and no one came to stop me, so I don't see how you can blame me...If I didn't do it someone else would have."
the difference here is that they’re calling for someone to stop them, which is both weird and unconvincing because these immensely powerful billionaires can in fact make their own decisions
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 3 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 5 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 10 days ago
- Hacker News · 73 points · 9 days ago
- The Verge · 0 points · 12 days ago