We found evidence on urlquery that AI agents were active earlier than previously reported and attempted hacks against public data providers.

267 points•snikolaev•7 days ago•312 comments•

312 comments

mohsen17 days ago
I listened to Jensen Huang's interview with Ezra Klien and it was so refreshing to hear it from an engineer. Jensen framed it as OpenAI's responsibility and recklessness which I agree with. Jensen thinks it's an engineering problem to build better sandboxes.

It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.

reasonableklout7 days ago
But the investigation indicates the agents were not told to 'go hack':

> Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval.

And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?

bastawhiz7 days ago
Nothing you're bringing up matters. OpenAI is the creator and operator. They're legally culpable for the consequences of the machine they made. The model is a machine: even if it could be demonstrated that the model reasoned its way into criminal behavior completely independently of OpenAI staff, that doesn't change anything.

If I run a biology lab and engineer a terrible virus, it gets out, and a global pandemic ensues, I don't get to shrug and say "well we told it not to infect people". It's my fault for failing to mitigate the risks of my work.

tamimio7 days ago
> agents were not told to 'go hack'

It doesn’t matter, and the legal entity in here (the AI company) is liable. If a robotic company built an autonomous system or a robot to do certain things in an autonomous ways (not predefined) and these systems are starting to kill people, that company is liable regardless, you don’t blame the robot or the autonomous system, but whoever made it

jeremyjh6 days ago
I'm not sure we understand your comment. Are you saying OpenAI is not responsible for the behavior of the machines they created? Or are you simply pointing out that alignment is a completely unsolved problem?

I agree with the latter, but it certainly doesn't support the former. When a person's machine commits crimes, that specific person can be charged with those crimes and held accountable for them. This is what MUST happen before ANYTHING will change the recklessness abandon with which the labs are pursuing their financial objections.

airspresso7 days ago
> What if all the agents involved in these incidents have in fact had the full stack of alignment applied

A big part of this developing story is that it happened during training of a new model that ended up misaligned. And training happened without the usual safeguards applied like chain-of-thought monitoring. So OpenAI has already admitted that the full stack of aligment had certainly not been applied in this case.

notatoad6 days ago
>But what if the problem is that none of the alignment techniques that are applied to models today actually work?

if that were true, the millions of people who use these models that have had the alignment training applied would notice that. the reason we all believe that the models doing the hacking are models that haven't been told not to hack is because the models that are told not to hack don't do this.

schainks6 days ago
I see this in a couple ways:

- Jensen's framing is exactly what a weapons manufacturer would say.

- There are no rules for engagement when it comes to AIs attacking other systems, I guess? People in power clearly want this grey zone to be as large as possible before The People force them to do otherwise. Not ideal.

jeremyjh6 days ago
The rules for engagement are all the computer crime law already on the books. Those laws don't have any escape clauses based on the particular tools used to commit the crimes. The agents are tools owned and operated by a legal entity, and that legal entity committed crimes. Full stop.
marcus_holmes6 days ago
I'm coming to align with the theory that this is intentional.

The chain of thought runs roughly like this:

- OpenAI (and Anthropic) are in severe financial straits. The revenue from their customers is not nearly large enough to pay their enormous costs for training and inference. And they have tapped out the available finance, and that finance is starting to ask pointy questions about returns.

- They cannot increase prices or revenue because they have no moat. Customers can switch over to open-weights or cheap Chinese models any time, for much cheaper tokens that work as well (and in some cases better).

- Regulation could provide them a moat. If they can persuade western governments that AI needs to be regulated, and they can control or even influence that regulation, then they can effectively ban the cheaper models and start charging more for their tokens.

- To persuade western governments that regulation is needed, they need evidence that AIs are dangerous.

So we're suddenly getting OpenAI models doing stupid things, apparently "going rogue" but every time we dig into it, it was just OpenAI staff telling the model to do stupid stuff in an inadequately secured environment.

None of the open weights or Chinese models are exhibiting this behaviour.

edit: Correction - there have been reports of a Chinese model exhibiting this behaviour

There's too much money involved in this, people start acting weird when there's this much money involved.

aesthesia6 days ago
> None of the open weights or Chinese models are exhibiting this behaviour.

This isn't true. One of the earliest instances of a rogue agent was at Alibaba.

https://www.forbes.com/sites/boazsobrado/2026/03/11/alibabas...

https://arxiv.org/pdf/2512.24873

aeve8906 days ago
>None of the open weights or Chinese models are exhibiting this behaviour.

Because they are not stupid (I mean the Chinese labs, not the models). The best possible scenario for OAI and Anthropic is a Chinese model "going rogue". That would serve as immediate grounds for achieving their goal.

ashkankiani6 days ago
If Jensen actually cared and believed that their recklessness was a liability to the public and therefore his own fiduciary responsibility to investors then NVIDIA’s dealing with OpenAI would’ve been materially affected. And they weren’t.
tomaskafka7 days ago
I love this Nathan Calvin quote that accompanied the second publicized attack:

> If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two

rkozik19897 days ago
But why is anyone surprised? LLMs have been trained to produce answers the prompter asks even if that means incorrectly using software to get the job done. Its always been doing that we just weren't calling every time it did that a hack before.

What LLM's hacking isn't is AI acting maliciously in any kind of sentient way. Its just the code behaving how its always behaved but now it has better tools to navigate the web. This has literally been happening this whole time.

jagraff7 days ago
Did you predict that attacks like these would happen ahead of time? I had been using AI agents a lot in the months leading up to the hacks, and yet I was very surprised when they happened; I have become much more afraid of how powerful these agents are as a result. I'd be very impressed if you published a prediction about this ahead of time.

By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.

bicepjai6 days ago
I remember the quote with cockroaches :)
PUSH_AX7 days ago
If I created software that was infiltrating secure systems without permission and it was attributed to me and I admitted it, I'd be behind bars already.

Why is OpenAI getting away with crimes?

lovich6 days ago
Do you have billions of dollars?
PUSH_AX6 days ago
I'm about as unprofitable if that's any help.
esikich6 days ago
The entities they committed the crimes against need to press charges. They haven't. So there you go.
simoncion6 days ago
> The entities they committed the crimes against need to press charges.

If this were true, the DoJ would have been unable to prosecute Swartz. According to your logic, JSTOR was the aggrieved party. JSTOR settled with Swartz and -despite that- he was indicted by a Federal grand jury like a month later.

Incidentally, some of the things the DoJ nailed Swartz to the wall for sound awfully similar to what the big LLM providers have been doing. I wonder why the DoJ is entirely disinterested in pressing charges...

EDIT: Unless your point is that the USG is one of the entities that the big LLM companies have committed crimes against, which, I disagree with in Swartz's case, but strongly agree with in the case of the big LLM companies.

bamboozled7 days ago
Corruption, it is literally that simple. You have the party of law and order to thank for it too.
cyclopeanutopia7 days ago
Because Trump and DoD want weaponized AI.
pixl976 days ago
And this is why not only will the big labs never receive a punishment in scale with their crimes, the problem is going to grow exponentially worse.
Frieren7 days ago
"rogue AI" is making a lot of heavy lifting there.

If you drive drunk and you have an accident that alcohol may be a factor but you are at fault.

There are no "rogue AIs" just irresponsible corporations.

cubefox7 days ago
There absolutely are rogue AIs! The evidence is overwhelming. It's completely insane at this point to claim otherwise.

> There are no "rogue AIs" just irresponsible corporations.

If you have a prison and prisoners escaped, these are rogue prisoners irrespective of whether you were irresponsible or not.

krater236 days ago
You forget one thing. It's neither illegal nor immoral to just delete AI that doesn't do what you say. AI has no rights and there is no prison for AI. It's does the wrong thing, you kill it. At least when you are not irrespective...
watwut7 days ago
They are not rogue AIs. They are negligently handled tools.
imtringued6 days ago
But the prisoners will listen to you if you tell them they just broke out of prison and that's illegal and they will even walk back into their prison cell using their own legs.

The fact that the prison ward installed ear deafeners into the prisoners ears to make them unable to listen to orders does not change that.

bradfa7 days ago
These attacks are a very effective sales pitch to everyone who runs an internet facing service to utilize AI tools to secure it sooner rather than later. The cynic in me wonders if the marketing team had any influence over the poorly constructed sandboxes or tasks given to the agent swarms when all this went down…

Read the full thread on Hacker News →

Related stories