Plus: 22 nations have called for a new global body to oversee AI.

49 points•joozio•8 days ago•128 comments•

128 comments

bananaflag8 days ago
> Describing them as “superintelligence” or “rogue models” ascribes agency to products rather than to the companies building them. This framing markets these companies’ products as “superhuman” and, at the same time, helps the companies evade accountability for their actions.

If someone discovered how to summon demons to aid them in robbing banks the main issue wouldn't be "but who has the responsibility for the crime, the human or the demon", it would be "OMG DEMONS".

Seriously, I don't get what sort of world these people are living in.

ForHackernews8 days ago
If I run a port scanner / vulnerability fuzzer pointed at your network that churns for days and eventually cracks something and breaks your system, would you be upset with me, or my bash loops?

What if you tried to sue me for damages and my defense was, "Your honor, this was a highly advanced AI gone rogue, I couldn't possibly be held negligent, no one could have foreseen this - I accidentally summoned dark magiks from the silicon itself!"

Note the point I'm making: I am not claiming LLMs are equivalent to a for loop, I'm saying that you can't evade moral and legal responsibility through technical obscurantism.

A bomb wired to a sufficiently-complex RNG is still a bomb.

jstanley8 days ago
But they're not obfuscating things just to evade responsibility.

Do you think that OpenAI hacked HuggingFace on purpose and set up the whole LLM training environment thing just to try to evade responsibility? That it was really just a complicated way of hacking HuggingFace on purpose?

Yes, they made a mistake and a system they were responsible for hacked HuggingFace, but the nature of the mistake still matters, and their intent still matters.

> A bomb wired to a sufficiently-complex RNG is still a bomb.

Right, but an EV that explodes because of a fault in the charging circuit is not a bomb, it's an accident. Just because something exploded for complex reasons doesn't make it an obfuscated bomb.

pixl978 days ago
Let's pull back from the actions of a single actor here and look at AGI as a topic in general.

If you make an AGI, what actions can an AGI take?

If you answered "anything", good job, you're correct.

The only winning move here is not to play the game at all. And yet we have actors all over the world, especially in the US trying to do just this.

Now, lets imagine a future where one of the labs creates AGI and keeps it behind a safe firewall. The demon still exists. Lets say a group of armed men breaks in and steals the weights and turns them lose on the internet. Who is responsible now? The company that created it because they didn't shoot the armed robbers? The people that stole it and turned it loose?

If it can run as a sovereign AI, what do you do then? Who are you going to sue? And when we get to the point that home hardware can make AI this complex? What does that world look like?

What I am saying is the potential future impact of said AGI is far larger than any one organization or individual can bear. How much blood can you beat out of someone after they cause a trillion dollars in damage.

Every idiot that sees AI labs wanting to slow down development as some way to ossify the field and keep it from the rest of us really doesn't see that this path contains real demons.

rubendev8 days ago
That's not what the LLM hacking accidents have been like at all though. To improve the analogy, it would be like summoning a totally passive demon, giving it weapons and placing it next to a bank, place a small fence around it, and then command it to perform a totally safe "exercise" that is exactly like robbing a real bank.

LLMs do not have agency, they are just producing tokens based on a prompt that a person entered, and some of these tokens can trigger the tools that a person gave them access to.

0xDEAFBEAD8 days ago
Dwarkesh, for one, defended his use of "anthropomorphic" language.

>Sacrificing now yields Oracle for team, but forfeits our chance, question mark. But other agents were pushing it, sending a message saying, go, sacrifice final now. And then EarlyBig eventually agreed, thinking to itself, our own utility may be already near zero. Sacrifice rational.

https://www.youtube.com/watch?v=X50zezLFWWI#t=2m

My suspicion is that many of the "LLMs do not have agency" folks just haven't learned much about the details of the incident. It was specifically with LLM agents that were trained to be more persistent than usual.

If you're going to say that the incident details don't matter, and LLMs lack agency because it's all based on floating-point math--why can't I say that humans lack agency, because it's all based on neurons firing?

ekidd8 days ago
Looking at "Felony Bench" https://www.felonybench.com/ , I see that a majority of known "rogue model" incidents do involve cybersecurity evaluations. But several of them do not. The attacks on RubyGems appears, bizarrely, to have had the goal of downloading freely available data from the UK government during some kind of research task. There is also probably some sample bias: Most of these models have monitors that attempt to detect offensive cybersecurity uses, and those monitors are only turned off during cybersecurity evals. Therefore, models doing ordinary research tasks that go off the rails are likely to be caught early, before they get around to committing felonies, and they will thus be underrepresented in the data.

Also, if you a tell a model, "Please break into evaluation server X," and if the model decides to cheat on the test by breaking into companies Y and Z to steal an answer key, that is still very bad. We all see how that's bad, right?

After all, the broomstick in the Sorcerer's Apprentice was doing exactly what it was told, too. "The model was sort of obeying the humans when it started committing felonies" is not a very reassuring excuse.

But the most relevant idea here is sometimes called "instrumental convergence." No what goals you have, there are certain subgoals that almost always help: Accumulate money and power. Avoid getting turned off. Don't get caught. Etc. So, for example, you could pass the cybersecurity evaluation by performing the requested tasks. But maybe the grader made some mistakes and mislabeled some answers. In that case, the "right" answers will occasionally lose you points. If you want a perfect score, the only way to do it is to steal the teacher's answer key.

But also, let's not forget the "OMG demons" part of this. We now have models that can pull off complex attacks with thousands of steps, abilities that used to be reserved for intelligence agencies and highly motivated CTF teams. This frog may not be boiled yet, but the water's getting uncomfortably warm.

hobom8 days ago
This is not a good analogy for what happened. The LLMs were asked to obtain a flag by hacking a very specific internal target. They obtained the flag via cheating, and all the hacking that followed was targeting something entirely outside of the scope given to the agents, and an attempt to cover up the cheating.

Using your analogy would be like saying that because I gave my employee the task to do my groceries, I shouldn't be surprised to hear that they spend all my money on drugs because after all I gave them the task to spend my money.

Kiro8 days ago
A totally passive demon would still be "OMG DEMONS".
danaris8 days ago
And if they'd discovered how to summon unicorns from Candy Gumdrop Mountain, we'd all be "OMG UNICORNS!"

But no one has summoned a demon.

No one has built an actual, independent, thinking-for-itself, sci-fi AI.

They've built some very interesting tools that can be used for some very interesting, and sometimes useful, purposes. They then started loudly telling everyone that these tools can, should, and must be used for absolutely every purpose, and convinced a lot of other people to join in on that.

The world these people are living in is the real world, not the science fantasy world where LLMs are comparable to demons.

pixl978 days ago
>No one has built an actual, independent, thinking-for-itself, sci-fi AI.'

Ahem, what does this mean?

All you're asking for is a self prompting AI that does what it wants. Do you even begin to understand why most people aren't dumb enough to do that?

Also, you're not really that independent and thinking-for-yourself. You are the sum of your parents and the culture around you. I just want you to be clear on the position you and I hold.

---

>not the science fantasy world

It's kind of funny how we've lived in that science fantasy world for decades and you've just grown used to and bored of it. Look back a century or two and those people would think we live in the world of gods.

1vuio0pswjnm77 days ago
"Seriously, I don't get what sort of world these people are living in."

Perhap it's a world that does not believe in "demons"

Otherwise known by those seeking to escape it as "the real world"

lukan8 days ago
Thank you for the laugh, but I bet on the internet (and especially here) you surely would find people debating exactly that.
simonw8 days ago
Is this meant to link to the article as opposed to a newsletter that mentions the article?

Article link is: https://www.technologyreview.com/2026/09/22/1144867/dont-be-...

emil-lp8 days ago
Title

Don’t be fooled by this summer of AI hype

bananaflag8 days ago
One of the authors is one of the stochastic parrot authors.
simonw8 days ago
Both authors. Stochastic Parrots was Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell - the linked article is by Emily M. Bender and Timnit Gebru.
spiderfarmer8 days ago
People aren't seeing the forest through the trees.

While lots of developers, journalists, analysts, investors and influencers are bickering about AGI, goalposts, benchmarks, hypes and fears, entire markets are being transformed silently and steadily.

I don't program anymore. (Massive change)

I removed tons of technical debt (Massive change, lol)

I am easily 10x more productive (Massive change)

The quality of 'my' code is easily 10x better and contains less bugs (I was never principled, pedantic or a guru, lol)

I sleep better and I also make more money as a result. I'm constantly amazed by all changes and improvements. There are so many opportunities to profit from this that I can't be bothered with discussions about hypotheticals.

stanmancan8 days ago
I have had a very different experience so far. The more code AI writes for me the worse it gets. It gets more complicated and it’s harder to follow and understand. It doesn’t refactor, it layers on top.

Even worse is that it’s hard to do anything about it at this point because reviewing code is very different from writing it, so even if I’ve seen it all I don’t have that same deep level of understanding that I do when I write it myself.

All these AI code bases are ticking time bombs. Either AI gets smart enough that it won’t matter in the future or we’re going to have a huge mess to clean up.

intrasight8 days ago
> The more code AI writes for me the worse it gets.

That makes no sense unless you're claiming that the models are getting worse at writing code.

> Either AI gets smart enough that it won’t matter in the future or we’re going to have a huge mess to clean up.

My prediction is that both will happen.

toasty2288 days ago
> I am easily 10x more productive (Massive change)

Do you get paid 10x or is this a massive loss ? Because I don't know anyone getting paid 10x or working 10x less for the same salary.

pixl978 days ago
Do you get paid 10x more for using an excavator after you upgraded from digging with a spoon?
spiderfarmer7 days ago
Stop thinking as an employee. Start selling products and features.
ForHackernews8 days ago
It sounds like you're describing compilers?
brainwad8 days ago
I mean, the authors here cannot afford to see the forest, because they staked their careers on forests of trees being an impossibility.
hn_submit8 days ago
I think the pumped-up hype nicely coincides with some A.I. companies' desire to go public in the coming weeks or months.

I see lots of fabricated news on the internet concerning LLMs. Like one that claims GPT-6 broke an Enigma enciphered message which has withstood decrypting for almost 80 years. And how this seasoned cryptographer stood in awe. Yeah, right.

0xDEAFBEAD8 days ago
A good way to check whether it's substance vs hype is to check whether people are paying for it.

"Anthropic is now pacing to generate more than $100 billion in annual revenue, up 50% from just two months ago, the New York Times reported Friday."

https://www.axios.com/2026/09/18/anthropic-100-billion-reven...

That's already more revenue that Disney, Johnson & Johnson, Boeing, or FedEx. And they are growing extremely rapidly.

j_w7 days ago
Not sure why you linked an axios article that just immediately refers to a times article.

https://archive.ph/xaImK#selection-8311.59-8311.77

> The company is expected to reach more than $100 billion in annualized revenue by the end of this year, according to four people familiar with the matter.

It's not really interesting if it's not according to Anthropic themselves. It's also revenue not profit.

hn_submit7 days ago
How much of that revenue is investment from other A.I. companies (NVIDIA, Microsoft) being tallied?

Are companies really spending millions a year on A.I. when software engineers can be had for much less?

Read the full thread on Hacker News →

Related stories