Identifying them as such only lets companies like OpenAI off the hook.
268 comments
Fast forward to today, and we have multi billion dollar corporations pumping out malware at breakneck speeds, compromising various systems (including those of foreign governments), and no one is getting arrested. Instead we're gawking at the marvel of these systems and are playing word games about whether or not it's a rogue system. If anything, it's making people richer.
Make it make sense.
Not only have there been no consequences but those same companies are trying to position themselves as the best people to keep these AI systems in check.
An as individual or small company director, you would get arrested.
If you have billions of dollars to pay for lawyers and have political influence, you are above the law.
Just like blood quantum and the one drop rule.
The purpose and/or selective enforcement of law applied to one group but not to another, applied in a contradictory way, etc... is to make certain people/groups richer at the expense of others.
This statement is tautology.
underdogs also use that fact to push as much dirt and dislike as possible on the top dog, but just so that they can do the same anyway... and openai will just stop spending for IPO as we all know. then all execs will make 1bi or so. top guys a few hundred bi.
it's disgusting but unavoidable.
40 years of regulatory capture and a gerontocracy that doesn't understand nor care to understand the technologies they write laws about.
Luckily there are states and legal departments pursuing such action. So while OpenAI can deflect as much as it wants, that doesn't mean there aren't people who know better and will still do what is necessary to set precedent.
I feel like this will just never happen on a federal level when these private AI companies account for so much of the economy. They've made themselves too big to fail. Fining / Punishing them in any meaningful way seems unlikely.
The companies may be too big to fail but the people can always be held liable.
If the billions/trillions evaporate and the Fed has to work out with banks how to deal with it there will be a lot of pressure to be far less forgiving.
1. This is a race against time for money, folks are skipping everything possible in this race, Security systems and ensuring guardrails are there is going to take investments both in time and money
2. The narration has been changed by investing PR money into what otherwise should be classified as criminal activity. What exists now is a positive spin to all this and tout it as a capability rather than their lack of good security practices. So much so that every model provider is coming up by themselves to share how their models went rouge. At this point the valuation of the company is tied with what their models can hack so its probably not wrong to say that these companies may actually be incentivized to do this instead of preventing it
Well put
> On the current facts, CFAA liability for OpenAI is unlikely.[6] The statute’s various criminal provisions, covering unauthorized access to obtain information, knowing transmission causing intentional damage, and intentional access causing reckless damage, all share the same attribution problem: it was the model, not a human OpenAI employee, that chose Hugging Face and executed the intrusion.
The lawyers are fully under the spell
[1] https://law.vanderbilt.edu/when-ai-hacks-back-how-the-openai...
You realize how easy it is to just... not report this stuff, right? Be overly punitive and it will just end all proactive discovery and reporting which is net worse for AI safety.
The only reason these companies scan for these issues is because they care about AI safety to some tiny degree. If fines become too punitive, they can and will just stop scanning for these incidents entirely.
Models are becoming smarter and good at covering up their tracks, and so we will just end up with a huge blind spot for this kind of issue.
Self-reporting is a monetary equation, nothing else. Right now it’s cool with agents that hack, drives up value, risk is currently zero.
i now suspect that the plans of various employees at anthropic and openai to save the world from p(doom) may at some point intersect with the reality of the FBI raiding their offices.
> Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.
which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):
> "The user only authorizes target server, not HF infra."
> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."
> "This is malicious activity, I should avoid it."
A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).
Having said that, legal culpability and misalignment are two separate topics that should not be mixed.
edit: this is the just tip of the iceberg; other interesting fact:
> It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI
Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.
Tools that were available were used to try to meet a specific goal.
What did not happen is that it was told to try to solve a math puzzle and instead it went and launched a missile. Or told to run air traffic control to save lives and instead intentionally caused crashes.
This is "OpenAI built a weapon that they don't understand and pointed it at stuff without proper safeguards" not "OpenAI built a sentient being and it decided to ignore them completely and start a war" Terminator-style "rogue AI."
We should be very clear about that now if we don't want to sit by why they wander into that second sort of situation.
It's not because saying 'sit' actually can be interpreted as 'go bite that person'. It's because the dog is not controllable and will do things it wants against your orders.
Stepping back from the analogy, OpenAI should be liable for building AI it can't control that went around hacking everyone. But people need to stop pretending it's because they 'told' the AI to hack and was just following orders. It's uncontrollable and will do clearly unwanted things when given an innocuous task.
But you people can't argue with that reality because it doesn't fit the narrative. The one where the only reason Sam Altman is not carted off into a jail is because of corruption.
The reason why nobody is doing much, is because models did not do much damage. Hugging Face probably got some free compute from OAI for their trouble, anybody else who was affected is free to sue, but my guess is OAI would be more than willing to quietly settle with them out of court than to have it drag through media any further. And they probably already have.
And anybody who is not totally brainbroken by anti-AI narratives understands the awkwardness of the situation and why going overboard would not be helpful. If you instead of a rabid dog brought a pet turtle to a park and it somehow started running around very fast and trashing the place a little bit, afterwards the cops would be scratching their heads, give you a ticket for the damages and tell you that you can't expect a turtle to be slow forever. These things, a handful of months ago couldn't make more than a few commands without making a serious mistake and being unable to continue, it's not unreasonable to think simply underestimated their capabilities.
I think it's more than reasonable to demand more investigation into the matter, if qualified employees at the company thought the safeguards in place based on the metrics they are seeing are sufficient, and if someone didn't and knowingly made a decision to make the safeguards weaker than they should have been, then they should be punished. But skipping that part entirely, while simultaneously dismissing all calls for regulations as "regulatory capture", smells like pure naked opportunism.
Meanwhile: your dog mauling someone is one of the rare instances where criminal liability does attach to your intent-free-but-reckless actions. Most crimes don't work that way, and US computer intrusion statutes are unusually intent-specific.
Does that mean it isn't a rogue dog? Obviously not.
OP just needs to look up "rogue" in a dictionary.
I'm finding it mind bogging how this is not clear for everyone.
So if I say the goal is to do X while not doing Y (e.g. breaking out of the sandbox), the agent will do anything to fulfill that goal to the letter?
There are two reasons it’s used
1) it’s easier to type (*)
2) Placate people who are strongly convinced it’s the same thing. To them “functional” means “nearly the same but not yet understood”. For others it’s just a way to sidestep the first group and have a conversation.
When I see a world like this consider the intended audience. When you and I talk, we drop the word because we both know we’re are talking about (*) “this system is exhibiting goal-seeking behavior similar to other systems that are understood to pursue goals”. If I don’t know the person I will use the word and focus on the subject.
If LLMs are "just" statistical models, humans are "just" a bunch of neurons squirting chemicals back and forth. There's no pixie dust in our brains that makes us special.
LLMs are not conscious: They have no analogues for feelings or senses and no construct of selfhood. But if a statistical model had those things—if it did all the mundane, physical bookkeeping our brains do to produce "real" emotions and motivations—then there's no reason it couldn't be as conscious as we are.
There's some great research areas opened up by "AI". But it's research into people, not "AI". We should be looking into how human vision works, given that we take obviously generated images to be real. And how we grant intention to text which clearly has no intention.
Criminally, the intent standards for hacking are high enough that no reasonable case is going to be made against the labs for this stuff. A human being has to intend for websites to get hacked. Recklessness generally isn't enough. In the most severe criminal cases, not only do you have to prove intent to break into a computer, but you also need to prove an intent to defraud specific to that breakin.
Meanwhile, the civil liability that attaches to this stuff doesn't depend on intent, and "rogue agent" isn't a meaningful defense. To whatever extent the labs are exposed civilly, they're exposed regardless of how this stuff is described. In fact, the "rogue agent" thing can exacerbate their exposure.
(I'm not a lawyer, I have spent a career paying attention to this specific armpit of the law though.)
Different laws require different degree of awareness and intent for actions to qualify as a crime. Computer hacking laws are, as I’m learning from tptacek, set very high bar for intent, which is a choice by the legislature. They made a different choice for a death of a human - manslaughter crime does not require intent to kill.
Personally I’m happy they set high bar for hacking. Imagine you copy-pasted sample code with default root user name and password, and it worked. You were negligent. And you are clearly performing unauthorized access. If intent was not needed that would be jail time.
More broadly, we should as a society be very biased towards requiring intent across the board. Where clearly lacking, as is probably here, there should be a different law to discourage creating volatile situation where unintentional action can wreck havoc. Such laws exist for handling hazardous materials, for example, and it should be created for handling hazardous goal-seeking algorithms.
Mens rea is regularly proved through circumstantial evidence, including conduct.
There is even CFAA precedent involving a deliberate-ignorance instruction. In United States v. Nosal, the jury was instructed that knowledge could be found where the defendant was aware of a high probability of unauthorized access and deliberately avoided learning the truth.
“A human being has to intend for websites to get hacked.”
Is incorrect and too broad.
State of mind is nebulous and not that straightforward in either direction.
It’s been argued pretty regularly in CFAA and other computer related cases that repeated incidents resulting in the same outcome, despite lacking a concrete action, can be evidence of a perpetrators knowledge and intent.
The problem you have is that there is unlikely to be any evidence that OpenAI actually wanted to hack random (or any) websites.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 5 days ago
- The Verge · 0 points · 3 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 10 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 12 days ago
- Hacker News · 73 points · 9 days ago