519 points•realsarm•12 days ago•396 comments•

396 comments

drtgh12 days ago
> relatively poorly understood technology

Poorly understood? how convenient...

LLMs are vectorial databases with losses that index statistically filled data, which uses a text interface to query such statistically filled data. The output is a string concatenation (statistically concatenated bit by bit).

When the LLMs are queried (prompted), you can get random mixed data as output, ERRORS, due to undesired indexes getting closer at one point while the string was being concatenated for the output, what affects the rest of the indexed content that will be concatenated.

It is intrinsic to this tech. The larger the context, the greater the probability of get mixed data. And if the provider lowers the precision of those indexes -in order to decrease hardware resources and energy consumption- such probability increases to the point where those errors are granted.

Even knowing that the queries can return wrong/mixed data in the responses, errors, the companies developing this, decided to introduce a new product, that connects such LLMs outputs to the command console, latter connected to internet, raw 'eval' running commands from such outputs witch obviously can contain whatever mixed random. Then we started to hear "oh, it deleted my directory", etc, and it seems the next one will be "a missile killed my wife", because it is a text concatenation engine with errors.

To name it "hallucination" is an euphemism... those are errors, and they are granted to happen at one moment. If they do not know this, then they ate too much marketing without doing their job, or it was a convenient contract for the pocket$ of someone.

theptip12 days ago
> LLMs are vectorial databases

You use a bunch of technical-sounding words here to make it sound like you understand. But to be clear, nobody understands why the evolved weights of a NN make the decisions that they do.

Almost nothing is understood about the actual representations used for nontrivial concepts, decision algorithms, etc.

If you look at the field of mechanistic interpretability, compared to “GOFAI” like learned decision trees, an LLM is completely opaque.

semiquaver12 days ago
I’m shocked how many otherwise well-informed people don’t understand or agree with this very fundamental fact of just how little we actually understand about why LLMs work as well as they do. They figure “it’s science, of course there’s math and theory behind it.”

AI research is almost as purely empirical as the gradient descent loops its practitioners use to optimize their models. “Why” anything at all works is barely an afterthought.

strangecasts12 days ago
> If you look at the field of mechanistic interpretability, compared to “GOFAI” like learned decision trees, an LLM is completely opaque.

I think the field deserves more credit than that, there are plenty of interpretability tools like

* natural language autoencoders for explanations of activations: https://transformer-circuits.pub/2026/nla/index.html (demo at https://www.neuronpedia.org/llama3.3-70b-it/nla )

* easier-to-interpret language model families like Backpack models: https://aclanthology.org/2023.acl-long.506/

* attribution graphs to trace internal reasoning steps: https://www.anthropic.com/research/open-source-circuit-traci... (demo at https://www.neuronpedia.org/gemma-2-2b/graph)

* functional analyses which have identified how LLMs do arithmetic - https://arxiv.org/html/2502.00873v1 - and how refusal happens: https://arxiv.org/abs/2406.11717

* data attribution methods linking training data to specific attention heads https://arxiv.org/abs/2601.21996

If we could give a comprehensive and global explanation of an LLM's behavior in a single paragraph, we wouldn't need the model to begin with, but that doesn't mean there's absolutely no understanding of the model internals whatsoever

coldtea12 days ago
>But to be clear, nobody understands why the evolved weights of a NN make the decisions that they do.

We might not understand particular "emergent" capabilities, but the low level mechanism is not just understood, but a deterministic algorithm with a handful of basic componets, that are well understood themselves.

tantalor12 days ago
They don't "make decisions".

That's like saying "my d20 decided to roll a 17"

gizajob12 days ago
Please go on, else you risk sounding like the person you’re criticising. The structure of the neural network is somewhat opaque because it’s hard to understand as the individual weights can’t be usefully interrogated, and naturally, it comes from big datasets which a human brain can’t really absorb in toto. Your comment was interesting so I’d like more of it.
margalabargala12 days ago
I agree with most of your comment, but...

> To name it "hallucination" is an euphemism... those are errors

I find this and other "don't anthropomorphize the computer" statements incredibly unconvincing.

People develop terms for things and language has always contained overloaded or "literally inaccurate" terms.

An LLM can have "hallucinations" in the same way a modern computer program can have "bugs".

usernomdeguerre12 days ago
I disagree, I think 'Hallucination' is a risk-shedding weasel-word. It's meant to shift blame away from the technology and its creator (multibillion dollar AI companies etc) in a way that doesn't hold those actors accountable or responsible for the outcomes.

In any other software it would be an error, regression, bug. And in a human process it would be at ~least something someone would call 'bullshit'.

piker12 days ago
I also agree with the parent, and I would also suggest "hallucination" is better than "error" which might imply an available deterministic correction. Hallucination makes it clear we're dealing with something different than an "error" or "bug".
bix612 days ago
Knowingly causing errors is not forgivable whereas hallucinations sounds esoteric and moves blame away from the people who are knowingly causing errors. It’s marketing speak.
DanHulton12 days ago
I’ll even go one step further - I don’t even like saying “Artificial Intelligence”. I think even that anthropomorphizes the machine too much. I prefer “Simulated Intelligence”, and I feel like that describes what is going on much better.

We are, through this process, simulating intelligence. These models aren’t intelligent, but they can simulate it. Every simulation has a degree of fidelity, and we’re not at 100%, not even with the top models. When you think about it in those terms, I find it becomes a lot easier to keep their limitations in mind. Additionally, it becomes easier to remember that this is an algorithm that you are running, and are responsible for, not another being that you can ascribe blame to.

john_strinlai12 days ago
>language has always contained overloaded or "literally inaccurate" terms.

"literally" is a great example of this, because it can also mean "not literally, but with emphasis".

pftburger12 days ago
It's not the first time we are encountering this issue. We've seen it in other autonomous systems. Trains are an older one, cars are a newer one. As you move out of the lower levels, the operator has a tendency to assume the system is increasingly more capable than it is. In trains, its so bad that they generate fake signals that the operator needs to respond to within a timeframe. I'd love to see this with implementations of other critical autonomous systems like this. Occasionally inject known errors into the system and expect the operator to catch them. If they don't, well... If it was a train driver I think we would fire them. If its an intelligence operative ordering a strike? :shrugs wearliy:
pocksuppet12 days ago
note the airline industry has moved past firing pilots who make mistakes, since that turned out to be a recipe for more plane crashes, not less. Instead, they find out why the mistake happened, and fix it. In some cases, this involves firing the pilot. They do not do that by default.
gamblor95612 days ago
There's an HN thread from yesterday in which people are extolling the ability of these vectorial databases to practice law because most of them don't understand how LLMs work. They assume that LLMs "understand" what they're being asked and what they're regurgitating.

Lane Kiffin almost destroyed LSU's football program acting on legal advice from ChatGPT. A video game publisher owes the former owners of a studio it acquired $200+ million because he based his actions on legal advice from ChatGPT. In the past week alone, California has disciplined over a dozen attorneys for LLM hallucinations because they used LLMs (mostly ChatGPT) to produce their legal pleadings.

And that's in an area where there are multiple safeguards to catch the issues before they become permanent problems. There's absolutely no justification for using AI in warfare, where mistakes tend to be pretty final.

thayne12 days ago
I think it is accurate to say that it is poorly understood by the general population, and probably the majority of operators using LLMs. Although I agree that is partly the fault of the companies making LLMs and related products.
jmward0112 days ago
History shows the US has a lot of hallucinated intelligence leading to war. WMD in Iraq comes to mind. I personally don't believe US intelligence on practically anything. It is all tainted. The pressure to 'find targets' to justify a political objective is overwhelming and putting it behind a black box that refuses to show its homework to those in ops using it, and ultimately the US people to judge decisions, is a cancer that leads to epic mistakes. Everything hidden in a dark 'need to know don't question it' box is bound to end up corrupt since there are no checks on that system. We have been building systems and processes for a long time that tell us what we want to hear, not what is real and not what we need to know. AI hasn't changed this, it has just made it even harder to realize since the product seems more polished.
beloch12 days ago
Sometimes intelligence "finds" evidence that suits political goals. e.g. You want to invade a country but are having a hard time convincing allies, getting UN approval, etc.. So, you let it be known to your spooks that you're not going to look too closely at their sources if they could just, pretty please, find something/anything juicy right bloody quick. The WMD evidence for the second U.S. invasion of Iraq was likely a case of this.

Then there's old-fashioned F'ups that don't fit your political agenda and are often quite damaging and embarrassing, not to mention lethal for people who don't deserve it. e.g. The U.S. used AI tools meant for rapidly picking targets in the middle of a war to plan their initial strikes on Iran. They had time to double check everything and do their due diligence before striking, but they didn't. So, a school next to a military base was targeted and a lot of kids died. This was a genuine F'up resulting from relying on a tool meant to give rapid but merely okay target selection under time pressure when there was no time pressure. The real mistake was made by humans.

The current case of the mistaken nuclear weapon parts shipment seems like an old-fashioned F'up, updated for the times. The people who didn't simply trust the tools and actually double checked should be commended. Others in their situation wouldn't have. I fully expect AI will be scapegoated for a lot of similar F'ups in the future even though it's still the responsibility of human beings to use ethics, caution, and restraint. AI doesn't get fired. Doesn't sue. It's actually pretty awesome for taking the blame.

elil1712 days ago
Was WMD "hallucinated" or a lie?
Towaway6912 days ago
Back then it was called “the truth” only later did it become something else. Perhaps an untruth.
partyficial12 days ago
both. they wanted to deceive people (lie). so they made up a reason (hallucinate).
smurf985212 days ago
Agreed, but They were accurate on the Russia full scale invasion, months before it happened
naasking12 days ago
Individual incidents aren't representative of accuracy or well-tuned truth seeking processes, that can only be assessed over time.
jameson12 days ago
It reminds me of the an Soviet officer who disobeyed early warning system's alert that US had launched four ICBMs and did not immediately relay the issue up to the chain of command.

https://en.wikipedia.org/wiki/Stanislav_Petrov

https://en.wikipedia.org/wiki/1983_Soviet_nuclear_false_alar...

jordanb12 days ago
Also 99 luftbaloons which was about a kid releasing some party balloons in Germany which confuses the EWS and causes WWIII.

Or the War Games movie and the Norad training mistake that inspired it.

nrr12 days ago
Pedantic clarification: The (original German) song itself didn't mention an actor in particular who had released the balloons, just that there were 99 balloons that flew to the horizon and jet fighters being scrambled in response. The epilogue tells of the consequent whole bunch of lasting destruction as the narrator talks about their patrols.

The English translation "99 Red Balloons" is considerably different as far as the details go.

JumpCrisscross12 days ago
It reminds me of laughing at the stupid old Europeans starting a war because an Austrian numpty got himself shot in Serbia. Like, we almost sunk a Chinese ship because an AI thought the sour candies they were transporting were nukes is in the same category of shitheadedness. (I'm embellishing–we don't know what was on board.)
pocksuppet12 days ago
We did sink many non-Chinese ships because a human thought the fish they were transporting were drugs.
pipes12 days ago
That's a very limited view of why ww1 started.
afavour12 days ago
That is a below high school level understanding of what caused WW1, for what it’s worth.
stefs12 days ago
the assassination of archduke franz ferdinand was just the trigger but not the reason. the war would have started anyway.
kakacik12 days ago
Not sure what is there to laugh at when millions died in pretty horrible ways including americans, anyway that was classical war mongering where one of the parties got exactly what they wanted (prussians/germans). That they got more than they wanted and how it folded them is part of history.

This case, its bullshit machine bullshitting randomly in between specs of stolen wisdom. Nobody asked for that, nobody is in control. We all humans lose in all cases. Quite different scenarios if you asked me.

Jordan-11712 days ago
"You maniacs! You blew it up! God damn you all to hell!"

"You're absolutely right, and that's on me. That's not just a mistake — it's a failure."

bwat4912 days ago
Why the shipping manifest details were load bearing:
paimapi12 days ago
The user was absolutely right. There were over 100 children in that elementary school [0]. I should not have recommended the double-tap.

Checking to see if there are better targets...

Clauding...

[0] https://en.wikipedia.org/wiki/2026_Minab_school_attack#:~:te...

ikrenji12 days ago
— a nice touch
jawiggins12 days ago
> The US military swung into action with plans to intercept the vessel, ... Military planes were in the air

A few months ago I listened to a talk a General (Admiral?) gave at CSIS where he said that the US purposefully announced their drone-hellscape plan for a Taiwanese invasion in order to force the PLA to reconsider their options/success-likelihood. I wonder if something similar could be coming of this reporting, on the face it looks like an embarrassing fumble, but it implies:

a) the US is able to, and regularly is, tracking and analyzing the manifests of ships between Iran and China.

b) the US is ready and willing to interdict and board vessels even from the PLA.

That these facts are now public might deter the Chinese leadership from attempting to share nuclear tech with Iran or other countries in the future.

IndeanCondor12 days ago
The Chinese haven't shared nuclear tech with the DPRK, an explicit PRC ally, whose advancements are instead based on Russian designs. PRC leadership are quite miffed with the North Koreans and view proliferation there as pushing RoK to manufacture weapons as well.

The PRC has been hard against nuclear proliferation as a policy over decades, it is highly compliant with IAEA inspection norms, despite the NPT not making it mandatory to be under those inspections. This policy is not something the US has in the past or will in the future engender into it through force.

jawiggins12 days ago
Just because they haven't done something in the past doesn't mean they won't in the future. China has been providing material support to Iran both in terms of intel and military hardward, so it's not outlandish to be monitoring and planning for it to continue happening into the future.

Read the full thread on Hacker News →

Related stories