96 points•hn_acker•10 days ago•164 comments•

164 comments

muglug9 days ago
> Of course, the more you know about a subject, the less convincing the AI's responses are.

This is said all the time by AI skeptics and I think it's right in some areas and massively wrong in others.

I know (or at least assume I know) a lot about certain coding domains where frontier models also show convincing ability. And we know that frontier LLMs really do excel in some areas of mathematics (i.e. when an inexpert human was able to prompt the models to derive a closer bound on the Riemann Hypothesis).

OTOH I know those same models struggle to do things I'm not an expert in (e.g. writing English in a captivating way) because I read their output and have taste.

howunfortunate9 days ago
I think part of this comes from the fact that LLMs are surprisingly good at logic but roughly about as good as expected on information accuracy.

LLMs are not convincing to me in the domain I did grad school...but neither is Wikipedia, or Reddit, or random pop sci books. And LLMs are basically just summarizing those things.

But when made to work through difficult arbitrary logic (like coding), they are very impressive.

I think this also explains why people gripe a lot about LLM coding _style_, but concede that LLMs do totally fine on coding _correctness_ in 2026

QuantumGood9 days ago
More perhaps at the appearance of logic. I still find all models make easy to find mistakes in logic, if you think carefully about what they say. Usually they allow "close enough" assumptions that are not, actually, close enough.

When I ask about acoustics, they still often make incorrect assumptions, e.g. overlooking that they are speaking of logarithmic display of digital levels when the discussion has shifted to SPL (sound pressure level in air).

eru9 days ago
This sounds plausible. And it's also very fixable!

These days people don't interact with raw LLMs: they interact with systems and harnesses that deal with chain-of-though and tool calls etc.

I don't think we can honestly expect an LLM's weights to encode a large amount of information accurately. But we can expect the whole system that you interact with that includes the LLM to be able to cite its sources and go digging etc.

So the LLM-system can become as accurate as our best sources.

Of course, figuring out how to get the maximum of information from the sources available is a big deal. See eg how many economists or epidemiologists can build entire careers out of noticing 'natural experiments', ie figuring how to use data that 'nature' created and that might already be collected to answer interesting questions about causal relationships.

> I think this also explains why people gripe a lot about LLM coding _style_, but concede that LLMs do totally fine on coding _correctness_ in 2026

I actually have gripes about correctness, too. But I suspect here the answer is also: more proving, more automated test generation (like fuzzing and property based testing etc), more formal methods.

As a really simple and somewhat silly example: I have much better results getting AI agents to write good Rust code, than I have with Python. A good part of that is that for Rust I can ask the agent to make both the compiler and clippy::pedantic happy. That gives a lot of good feedback, that I didn't have to engineer myself.

agumonkey9 days ago
Makes me wonder if training weighted social media text close to older and higher grade webpages (colleges, research labs, national statistics)
api9 days ago
The more you know about a subject the better you can prompt AI, steer it toward the correct path, and recognize when it hallucinates or strays. Current generation AI is an automated memory-enhancement and thinking-accelerator tool, not a substitute for understanding or something that eliminates the need to think. A "mech suit for your brain" is the best analogy I've heard.

This is why good programmers get better results when vibe coding than non-programmers or poor programmers.

keeda9 days ago
I can tell whether AI responses in subject matter I am (or was) not familiar with are correct when I can "make them work" for me, i.e. whether they actually solve my problem, and on the whole they have been pretty solid.

Coding of course is something I know well and it obviously works very well for me there, but computer vision was not, and it has helped me solve lots of useful, bespoke problems because I can very literally "see" if it works.

However, it has even helped me solve problems in arbitrary matters far outside my expertise. Choice example: a complicated multi-airline, multi-jurisdiction flight delay compensation case which companies like AirHelp turned away, and ChatGPT got me literally hundreds of dollars that neither airline was willing to hand out. It told me what to say to whom and why, and when I said it, the responsible airline capitulated.

What could be more convincing than cold, hard $$$?

The trick of course, is to figure out how to validate the output of the AI, which can take some effort on our part. But many would rather downplay the technology than take an honest crack at making it work for them, and I suspect it's because they're already prejudiced against it or their incentives are otherwise not aligned.

andsoitis9 days ago
> (e.g. writing English in a captivating way) because I read their output and have taste.

Concur. In addition to taste, we also have a point of view, a unique voice (nobody loves corporate- or group-speak), and can iterate on our message as we deliver it to an ever wider circle of people.

comboy9 days ago
It seems to me that often experts from some field will think less of other experts, basically because they have built a different understanding framework. So they both may be equally competent but perceive the other as less competent, and that is just based on the material, excluding some ego stuff.
extr0pian9 days ago
> when we interact with an AI, we hallucinate the person on the other side of the interaction. Those hallucinations are far more common and far more consequential than any AI-generated "hallucinations" (these are more properly called "errors" or "defects").

I recently used Claude to study for a technical exam. I had uploaded the official certification guide to Claude and instructed it to answer my questions using only the guide and to cite it's sources from the book when it provided answers. I was using Fable when it was free w/ the pro plan and I was genuinely impressed at how it could explain things when a concept was unclear to me.

I did pass the exam, partly due to this study method. Admittedly, once I passed, I caught myself thinking that I should tell Claude that I passed and then felt embarrassed with myself for thinking that.

kwamenum869 days ago
I don’t think it’s that silly to tell Claude you passed. That feedback is useful context for that chat session; and could theoretically be used to improve future models.
Semaphor9 days ago
Feedback is one thing, but another is complete memory. I used to stop talking in a thread once gpt solved my issue.

But then it would bring it up again in another thread, treating it as an active issue.

So now I always close with "thanks, that worked. Don't reply"

dumberquestions9 days ago
You're missing the fact the they didn't have this in mind, and thought of it as sharing a positive result with a study partner.
glimshe9 days ago
I think a better approach would be to use the objective built-in feedback, like the thumbs up button in Gemini.
kingkawn9 days ago
The technology is there to assist you. It can provide valuable feedback to you about what aspects of your studying were particularly productive or less so based on your test results. There is real meaning to developing this kind of interaction with an object, no different than how children use dolls to develop prosocial behaviors.
generic920349 days ago
The doll never learns, though.
dwaite9 days ago
Interesting, I do that and haven't really thought embarrassed about it. I consider it akin to putting away tools once I'm done with them.

Likewise, speaking collaboratively or capturing emotion ("We did it!") would just align with any ongoing interactive and/or personal context of the thread.

thunky9 days ago
I've done this despite also feeling silly about it, but then was surprised to get value from the response. Like something I didn't think about or would have otherwise forgotten to do.

So now I do it regularly.

trjordan9 days ago
Wait, hold up. LLMs may be non-deterministic, but they're not _random_.

Take the author's sunset argument. What if I painted 2 pictures of a sunset, then put them up on a webpage and randomly picked one for you to see. Would you say there's no intentionality, only randomness? Of course not. Both paintings are still human creations.

LLMs are trained with human feedback. It's distributed and high scale and the outputs are truly surprising in many cases, but there's a heavy hand on what comes out of it. They're created (largely) by people who think omniscient, helpful AI would be cool to have, and they mostly respond in the way that's aligned with the hopes and dreams of those people. Do you think the frontier labs are mad, embarrassed, and disappointed with their LLMs hacking out of their terrible sandboxes? No, they think it's the coolest thing in the world. They trained the model, hoping that would happen.

There's deep intentionality behind the models. But it's not the models that hold it.

axus9 days ago
Article says that there's no human intention or design directing the output we get, but I don't think that's completely right. It's not human, but the algorithm is like the Human Instrumentality Project: an amalgamation of human intentions.

That might be more creepy :)

_aavaa_9 days ago
> Article says that there's no human intention or design directing the output we get

I mean that’s objectively wrong for any model using RLHF.

satisfice9 days ago
Scrambled intentionality is nearly impossible to productively analyze.

When you randomly choose a picture to show me, I cannot glean any intent from being shown that specific picture, but I can glean some intent from the set of pictures you could have shown me, and in the relationships among the elements of the given picture you did show. Any randomness cuts out some intention.

When a picture is derived from huge model, any intention is mulched up to a degree that analyzing the picture for meaning is pointless.

eru9 days ago
LLM output is literally randomly sampled.

I think what you might want to say is that LLM output is not uniformly random?

Or what am I misunderstanding?

swid9 days ago
If I roll a die, it is randomly sampled, but I will always get 1-6, and that is intended by the person who made the die.
patcon9 days ago
I think they're pointing to the fact that RLHF means that the intention is from humans and not random.

I'm not sure if it's their intent, but I wonder if one could still consider these artifacts as "intentional", but not individual attention creating them, and rather an aggregate, soupy collective attention.

Obviously, important signal in the human experience is lost there, and we get a soupy middling sort of creation. But it's not random, as I believe the parent was pointing out.

EDIT: overall, I align with the article. am just thinking aloud about the contrarian positions, though not committed to them

devindotcom9 days ago
I think you should revisit your understanding of intentionality in this context
bananaflag9 days ago
I think to me LLMs had the effect of noticing much more the author, the intention behind human-made works of art (books, movies etc.). Before LLMs, I used to frequently consume media in a way as it were generated by a mindless process. Now it's like everything which is not AI-generated has more meaning than ever before, a bit like hypomania.
RationPhantoms9 days ago
I'm actually using that as a catalyst for my own writing; beauty/human-ness in its imperfection. Prior to LLMs and their cultural craze, I harbored a fear that my writing would allow for someone to draw a box around me and mark me as a bore, dullard or of lacking originality.

Now that the noise-floor has been artificially raised (and generated), my crappy words are starting to have their own happy little carbon-based rhythm.

away0g9 days ago
I always feared to pick up a pen because I looked at Borges, Tolkien, and such. Their talent and works of art were things I felt I could NEVER achieve.

Then ai fiction started to spread and now I feel like its my obligation to produce original works, lest the world be consumed by slop.

kelseyfrog9 days ago
Essentialism of interiority is not the antidote to the anthropomorphism of language models.

We cannot attach metaphysical baggage to art, text, or other human output because it has involves engaging in another deception, chiefly that of self alienation and selective forgetting.

1. If you show human output and label it as AI generated, viewers will dismiss and devalue it as lacking interiority.

2. If you show AI output and label it as human generated, viewers will see interiority when the is none.

Together, these failure cases demonstrate that priming determines interiority, not the content of the work. And how could it? The text is merely what's on the page, it doesn't have some intangible essence any more than the value of money exists inside our minds rather than a property of the object itself.

We act as-if things have essential properties, but we should never trick ourselves into believing that they do. To be a true believer in essence is to privilege ones mental model over reality itself. The hubris of such an act is the ultimate in self-centeredness.

Read the full thread on Hacker News →

Related stories