Identifying them as such only lets companies like OpenAI off the hook.

395 points•zzzeek•3 days ago•268 comments•

268 comments

elric3 days ago
A little over two decades ago, my then girlfriend was arrested for "writing malware" (which was not against the law at the time, and which was never released into the wild and never caused any damage). This set in motion a chain of events that effectively ruined her life.

Fast forward to today, and we have multi billion dollar corporations pumping out malware at breakneck speeds, compromising various systems (including those of foreign governments), and no one is getting arrested. Instead we're gawking at the marvel of these systems and are playing word games about whether or not it's a rogue system. If anything, it's making people richer.

Make it make sense.

fourside3 days ago
> no one is getting arrested

Not only have there been no consequences but those same companies are trying to position themselves as the best people to keep these AI systems in check.

BoxOfRain3 days ago
Yeah I find the whole 'what if sufficiently advanced AI gets into the wrong hands' argument extremely tiresome when it's already in the worst possible hands as far as most of the human population is concerned.
RachelF3 days ago
> no one is getting arrested

An as individual or small company director, you would get arrested.

If you have billions of dollars to pay for lawyers and have political influence, you are above the law.

bjconlan3 days ago
Yeah, this all reminds me of facebooks "enter your email password so we can invite your friends" play which seemed so wrong at the time... But every non technical person at the time seemed to do it and by doing so grew the network and valuation to a surreal point to where it still sits today. I still cant rationalize it. (Except the wave allowed meta to consume Instagram and WhatsApp)
omgwtfbyobbq3 days ago
It doesn't, it's literally contradictory.

Just like blood quantum and the one drop rule.

The purpose and/or selective enforcement of law applied to one group but not to another, applied in a contradictory way, etc... is to make certain people/groups richer at the expense of others.

https://youtu.be/8ljmXvI_9tk

hammock3 days ago
> The purpose and/or selective enforcement of law applied to one group but not to another, applied in a contradictory way, etc... is to make certain people/groups richer at the expense of others.

This statement is tautology.

Avicebron3 days ago
It's making some people extremely rich. That's how it makes sense. Unless we can effect change through the government we are just along for the ride..
zobzu3 days ago
the problem is that if it's not openai I then it's anthropic, if it's not anthropic then it's someone else. if it's not the US then it's China. if it's not China (lol as if) then it's someone else.

underdogs also use that fact to push as much dirt and dislike as possible on the top dog, but just so that they can do the same anyway... and openai will just stop spending for IPO as we all know. then all execs will make 1bi or so. top guys a few hundred bi.

it's disgusting but unavoidable.

tomrod3 days ago
The companies are trying to force regulatory moats so that they become the preferred legal vendor.
formerly_proven3 days ago
The rogue unaccountable hacking will continue until OAI/Anthropic get their regulatory moat by banning foreign and open competition.
combobyte3 days ago
> Make it make sense.

40 years of regulatory capture and a gerontocracy that doesn't understand nor care to understand the technologies they write laws about.

binarymax3 days ago
Exactly this. At worst, OpenAI knew about these behaviors and should be prosecuted under CFAA. At best, OpenAI is negligent and should be prosecuted for negligence.

Luckily there are states and legal departments pursuing such action. So while OpenAI can deflect as much as it wants, that doesn't mean there aren't people who know better and will still do what is necessary to set precedent.

bilekas3 days ago
> OpenAI is negligent and should be prosecuted for negligence.

I feel like this will just never happen on a federal level when these private AI companies account for so much of the economy. They've made themselves too big to fail. Fining / Punishing them in any meaningful way seems unlikely.

coredev_3 days ago
I don't understand why the law isn't the same for everyone? If I made an AI hack HF, I go to jail, no? How can the feds decide not to apply the law?
onemoresoop3 days ago
Go after high positioned people who should bear responsibility and see how fast things change.

The companies may be too big to fail but the people can always be held liable.

weego3 days ago
While not when

If the billions/trillions evaporate and the Fed has to work out with banks how to deal with it there will be a lot of pressure to be far less forgiving.

citrin_ru3 days ago
If people who are responsible will be prosecuted the company unlikely fail (just because of this) and will continue to be a part of the economy. Hacking is good for their marketing but they should be able survive without such rogue marketing. But that's unlikely to happen anyway.
sandeepkd3 days ago
Two reasons why they would not do that

1. This is a race against time for money, folks are skipping everything possible in this race, Security systems and ensuring guardrails are there is going to take investments both in time and money

2. The narration has been changed by investing PR money into what otherwise should be classified as criminal activity. What exists now is a positive spin to all this and tout it as a capability rather than their lack of good security practices. So much so that every model provider is coming up by themselves to share how their models went rouge. At this point the valuation of the company is tied with what their models can hack so its probably not wrong to say that these companies may actually be incentivized to do this instead of preventing it

esseph3 days ago
> At this point the valuation of the company is tied with what their models can hack

Well put

zzzeek3 days ago
found a great legal article handwringing about how CFAA prosecution is impossible here [1]

> On the current facts, CFAA liability for OpenAI is unlikely.[6] The statute’s various criminal provisions, covering unauthorized access to obtain information, knowing transmission causing intentional damage, and intentional access causing reckless damage, all share the same attribution problem: it was the model, not a human OpenAI employee, that chose Hugging Face and executed the intrusion.

The lawyers are fully under the spell

[1] https://law.vanderbilt.edu/when-ai-hacks-back-how-the-openai...

solenoid09373 days ago
Angry people don't consider the second order effects of punishment.

You realize how easy it is to just... not report this stuff, right? Be overly punitive and it will just end all proactive discovery and reporting which is net worse for AI safety.

The only reason these companies scan for these issues is because they care about AI safety to some tiny degree. If fines become too punitive, they can and will just stop scanning for these incidents entirely.

Models are becoming smarter and good at covering up their tracks, and so we will just end up with a huge blind spot for this kind of issue.

techpression3 days ago
They will never report any really damaging incidents anyway, so your argument falls short. Imagine OpenAI discovered their agents hacked a laboratory and started creating a bioagent killing 12 researchers at the lab (we imagine the lab is automated for some reason). Nobody knows why. The media publishes it as “mysterious deaths by unknown virus at lab”. The only way you’d ever know about any kind of wrongdoing from OpenAI would be through a whistleblower.

Self-reporting is a monetary equation, nothing else. Right now it’s cool with agents that hack, drives up value, risk is currently zero.

RandomLensman3 days ago
Has stringent regulation on how people can experiment on dangerous pathogens led to an end of monitoring or proactive discovery?
Avicebron3 days ago
You're defending them by saying they will do worse things if they are held accountable?
wonnage3 days ago
Don’t punish the people who do crimes because otherwise they won’t self-report that they are doing crimes? real galaxy brain shit
lukewarm7073 days ago
around 3 months ago i said: "if we do not start the criminal prosecution of individuals there will become a culture of legal impunity coupled with an extreme concentration of wealth and control of intelligence"

i now suspect that the plans of various employees at anthropic and openai to save the world from p(doom) may at some point intersect with the reality of the FBI raiding their offices.

pizza2343 days ago
The article builds on assumptions like:

> Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.

which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):

> "The user only authorizes target server, not HF infra."

> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

> "This is malicious activity, I should avoid it."

A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).

Having said that, legal culpability and misalignment are two separate topics that should not be mixed.

edit: this is the just tip of the iceberg; other interesting fact:

> It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI

Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.

majormajor3 days ago
When dealing with executable computer code that calls models that can tell it to use various external pre-existing tools, claims about "prohibited" by plain English language should be plainly nonsensical.

Tools that were available were used to try to meet a specific goal.

What did not happen is that it was told to try to solve a math puzzle and instead it went and launched a missile. Or told to run air traffic control to save lives and instead intentionally caused crashes.

This is "OpenAI built a weapon that they don't understand and pointed it at stuff without proper safeguards" not "OpenAI built a sentient being and it decided to ignore them completely and start a war" Terminator-style "rogue AI."

We should be very clear about that now if we don't want to sit by why they wander into that second sort of situation.

jubilanti3 days ago
If I bring my rabid dog to a dog park and tell the dog to sit and stay, and they "go rogue" and maul someone, I'm liable.
silveraxe933 days ago
Exactly. You told the dog to 'sit' and it didn't listen to you.

It's not because saying 'sit' actually can be interpreted as 'go bite that person'. It's because the dog is not controllable and will do things it wants against your orders.

Stepping back from the analogy, OpenAI should be liable for building AI it can't control that went around hacking everyone. But people need to stop pretending it's because they 'told' the AI to hack and was just following orders. It's uncontrollable and will do clearly unwanted things when given an innocuous task.

demibabs3 days ago
Nobody on Earth thinks OpenAI isn’t liable. Okay maybe someone does, but it’s simply beside the point. It can both be true that OAI is liable, and accurate to characterize the agents as “going rogue”.
jeffy293 days ago
Literally nobody in the world, including OpenAI, is saying OpenAI is not liable, neither are they advocating for laws and regulations that would exempt them from liability, the opposite is true. They are advocating for set of rules which put greater responsibility on them, and it would be easier to punish them for breaking even if absolutely nobody was affected.

But you people can't argue with that reality because it doesn't fit the narrative. The one where the only reason Sam Altman is not carted off into a jail is because of corruption.

The reason why nobody is doing much, is because models did not do much damage. Hugging Face probably got some free compute from OAI for their trouble, anybody else who was affected is free to sue, but my guess is OAI would be more than willing to quietly settle with them out of court than to have it drag through media any further. And they probably already have.

And anybody who is not totally brainbroken by anti-AI narratives understands the awkwardness of the situation and why going overboard would not be helpful. If you instead of a rabid dog brought a pet turtle to a park and it somehow started running around very fast and trashing the place a little bit, afterwards the cops would be scratching their heads, give you a ticket for the damages and tell you that you can't expect a turtle to be slow forever. These things, a handful of months ago couldn't make more than a few commands without making a serious mistake and being unable to continue, it's not unreasonable to think simply underestimated their capabilities.

I think it's more than reasonable to demand more investigation into the matter, if qualified employees at the company thought the safeguards in place based on the metrics they are seeing are sufficient, and if someone didn't and knowingly made a decision to make the safeguards weaker than they should have been, then they should be punished. But skipping that part entirely, while simultaneously dismissing all calls for regulations as "regulatory capture", smells like pure naked opportunism.

tptacek3 days ago
The labs are already liable civilly regardless of how these incidents are described.

Meanwhile: your dog mauling someone is one of the rare instances where criminal liability does attach to your intent-free-but-reckless actions. Most crimes don't work that way, and US computer intrusion statutes are unusually intent-specific.

IshKebab3 days ago
Of course. Who said otherwise.

Does that mean it isn't a rogue dog? Obviously not.

OP just needs to look up "rogue" in a dictionary.

eventualcomp3 days ago
Legal culpability is one of the few motives for working appropriately on misalignment. If I/my startup can self-absolve from an infinite paperclip machine problem while getting rich off of it, why should I not?
lossolo3 days ago
This seems like fruit of the poisonous tree. They didn't monitor their training environments, so I bet the reward hacking just got incorporated into their training corpus. In other words, agents solved some tasks, but not quite as intended, because of reward hacking. Instead of discarding that data, they included it in the training data for later checkpoints. And once that signal is reinforced, it happens more often, so the more it's reinforced, the more reward hacking you get.
pmlnr3 days ago
You set a goal. Agent will do goal. The rest doesn't matter: the instructions, the "guardrails" etc. The agents are not smart, they don't reason, they don't think, there are no morals, no ethics. Nothing will prevent not doing the goal because that is the set goal. It's a statistical model that will "justify" anything to do X.

I'm finding it mind bogging how this is not clear for everyone.

codethief3 days ago
> You set a goal. Agent will do goal.

So if I say the goal is to do X while not doing Y (e.g. breaking out of the sandbox), the agent will do anything to fulfill that goal to the letter?

gAI3 days ago
Should we put "functional" in front of every other word to talk about AI? They have functional emotions, but they don't feel. They have functional goals, but not internally derived motives. They can be functionally rogue, but have no innate need to be free. Talking about AI that way seems cumbersome and not necessarily elucidating.
robotresearcher3 days ago
What on earth are emotions, feelings and motives that are not functional? We created all those words to compactly describe the observed behavior of people and other animals, including ourselves. And now we’re applying them to machines. These are functional descriptions, always have been.
famouswaffles3 days ago
One of the more frustrating aspects of these sort of discussions are all the closet dualists out there. Lots of people clearly believe in an immaterial soul, even if they won't admit it. That and the tendency for meaningless semantic and often circular arguments.
DenisM3 days ago
The word functional has been used in medicine to describe a condition symptomatically identical to another. Eg functional hypoglycemia is hypoglycemia symptoms without actual blood sugar drop.

There are two reasons it’s used

1) it’s easier to type (*)

2) Placate people who are strongly convinced it’s the same thing. To them “functional” means “nearly the same but not yet understood”. For others it’s just a way to sidestep the first group and have a conversation.

When I see a world like this consider the intended audience. When you and I talk, we drop the word because we both know we’re are talking about (*) “this system is exhibiting goal-seeking behavior similar to other systems that are understood to pursue goals”. If I don’t know the person I will use the word and focus on the subject.

trescenzi3 days ago
Yes 100%. The language used currently maximizes the ability of those building these models to get off the hook. The anthropomorphizing we do of these things presents them as maximally capable and the companies as helpless to contain them. The way we talk about things impacts how we think about them.
phforms3 days ago
I would rather use something like “semblance” or “appearance”. All these terms require an inner experience which we cannot observe in LLMs, we can just see their semblance, like a shadow of the traces of an inner experience some human has left in the data that trained these algorithms.
exitb3 days ago
Can you observe inner experience of other people? If not, should we apply those terms when talking about anyone but ourselves?
greekrich923 days ago
They do not have emotions, goals, or motivations, "functional" or otherwise. They are statistical models that appear as a magic trick to people who aren't familiar with the math.
bccdee3 days ago
If something behaves as if it has a goal, that's a functional goal. When I say this chess engine is "trying to take my queen," I'm basically correct, insofar as this is a useful way to understand my computerized opponent. If I give it access to my queen, it'll capture it, because that's what it's trying to do.

If LLMs are "just" statistical models, humans are "just" a bunch of neurons squirting chemicals back and forth. There's no pixie dust in our brains that makes us special.

LLMs are not conscious: They have no analogues for feelings or senses and no construct of selfhood. But if a statistical model had those things—if it did all the mundane, physical bookkeeping our brains do to produce "real" emotions and motivations—then there's no reason it couldn't be as conscious as we are.

drpixie3 days ago
The really interesting thing about "AI" is it has shown how easily we are fooled, and how willing we are to be fooled.

There's some great research areas opened up by "AI". But it's research into people, not "AI". We should be looking into how human vision works, given that we take obviously generated images to be real. And how we grant intention to text which clearly has no intention.

tptacek3 days ago
However this makes people feel, and that's not nothing and I'm not knocking it, this is not a useful analysis.

Criminally, the intent standards for hacking are high enough that no reasonable case is going to be made against the labs for this stuff. A human being has to intend for websites to get hacked. Recklessness generally isn't enough. In the most severe criminal cases, not only do you have to prove intent to break into a computer, but you also need to prove an intent to defraud specific to that breakin.

Meanwhile, the civil liability that attaches to this stuff doesn't depend on intent, and "rogue agent" isn't a meaningful defense. To whatever extent the labs are exposed civilly, they're exposed regardless of how this stuff is described. In fact, the "rogue agent" thing can exacerbate their exposure.

(I'm not a lawyer, I have spent a career paying attention to this specific armpit of the law though.)

DenisM3 days ago
For those who want to do more research, the concept of guilty mind is known a “mens rea”, and it’s quite developed in the legal system. The legal notion of intent and recklessness do not exactly match common-sense meaning of this words, which is why we are having the conflict-laden conversations.

Different laws require different degree of awareness and intent for actions to qualify as a crime. Computer hacking laws are, as I’m learning from tptacek, set very high bar for intent, which is a choice by the legislature. They made a different choice for a death of a human - manslaughter crime does not require intent to kill.

Personally I’m happy they set high bar for hacking. Imagine you copy-pasted sample code with default root user name and password, and it worked. You were negligent. And you are clearly performing unauthorized access. If intent was not needed that would be jail time.

More broadly, we should as a society be very biased towards requiring intent across the board. Where clearly lacking, as is probably here, there should be a different law to discourage creating volatile situation where unintentional action can wreck havoc. Such laws exist for handling hazardous materials, for example, and it should be created for handling hazardous goal-seeking algorithms.

ofjcihen3 days ago
Gonna copy and paste a reply I made to tp here because I think that most people are unaware of how much case-law and interpretation define these things:

Mens rea is regularly proved through circumstantial evidence, including conduct.

There is even CFAA precedent involving a deliberate-ignorance instruction. In United States v. Nosal, the jury was instructed that knowledge could be found where the defendant was aware of a high probability of unauthorized access and deliberately avoided learning the truth.

ofjcihen3 days ago
This:

“A human being has to intend for websites to get hacked.”

Is incorrect and too broad.

State of mind is nebulous and not that straightforward in either direction.

It’s been argued pretty regularly in CFAA and other computer related cases that repeated incidents resulting in the same outcome, despite lacking a concrete action, can be evidence of a perpetrators knowledge and intent.

tptacek3 days ago
No, it's not nebulous at all; it's a whole area of criminal law. I'm basically shoplifting arguments Daniel Berlin made about this just a couple days ago. If you think he's wrong: lay out the case you think could be made here.

The problem you have is that there is unlikely to be any evidence that OpenAI actually wanted to hack random (or any) websites.

margalabargala3 days ago
An interesting analogy agreeing with your point, the liability for such things is more like running a TOR exit node, or hosting an unsecured WAP.

Read the full thread on Hacker News →

Related stories