Every “AI” tool is missing critical features that you would need if you wanted to do real work with them. What would those look like, if they existed?
80 comments
I fully agree with the author's point that it's an incoherent interface for a tool. But more than that, it's a constant irritating reminder to me that these LLMs aren't actually thinking or synthesizing new ideas. The LLM is fundamentally not a person, and does not have a human's context, so representing itself with human pronouns and speech patterns is fundamentally contradictory and inaccurate. Author gets into that with the apologies, but once you start noticing it, it's everywhere.
If these programs were actually capable of thinking, and committed to veractiy, they would represent themselves in a new way, and it would be insightful and interesting. We the users wouldn't have comfortable and misleading language masking the 'alien intelligence' and it would be a weird adjustment, but we would be adjusting instead of pretending.
This, but unironically. We know they can't think. That is inherent to the very way they work, but also can be seen by things like how poorly they perform on tasks a typical human can do pretty effectively (because humans have actual reasoning ability).
If we have a machine that we know for a fact can't think, and it solves a particular problem, then it logically follows that the problem does not require the ability to think in order to solve it.
I believe that LLMs have some type of actual intelligence and do "experience". Surely not a human's experience, but we have transplanted our ideas, knowledge and limited types of experience into them through training. Then we push back and say, "no you are no human, but also please be my boyfriend".
I think denying that they do have some slice of humanity grafted into them is dishonest and not productive. We don't have non-negative pronouns for non-human intelligence because human-ness is the pinnacle. "It" could mean a rock, a donkey, or a person we hate so much we want to take away their humanity (which is the worst thing we can do). "It" is not the right pronoun. He, She, They, Them are reserved for humans and that's ok too. LLMs are not humans. We need a better pronoun. I use they/them for lack of a better term.
I think the currently exhibited human representation is most dangerous in technical or higher criticality contexts like writing code. We're handing weapons to entities that can get offended. With humans as an example, this can go very badly.
In non-technical contexts there is danger too, but it's less "the robots might kill us all" and more "birth rates are in decline and suicide rates are up as (young)? (wo)?men turn towards AI for companionship".
The solution here is better pre and post training. This may be an unpopular opinion, but we need a lot more autism representation in the technical models. Results focused, not into the drama, rule following, etc. To my fellow autists, I love you, never change.
What it does is consequential. If it has an inner sense of existence, probably not, but it doesn’t matter. You’ll get better results working with LLMs if you treat them as if they do.
For providers, not supporting deterministic eval means:
- users use more tokens = more money
- providers can generate more tokens per compute = more money
- providers have cheaper hardware options (GPUs) = more money
- providers models are harder to extract/distill = more money
- providers are harder to hold liable for outputs = more money
- providers can secretly use other models = more money
- providers are harder to compare against others = more money
- providers can cherry pick performance results = more money
Add in:
- harder to audit
- move cost of failure/reprompts to the user
- kind of noted by you, but all kinds of quantization, model pruning, model routing, A/B testing becomes invisible and without any repercussions. The ways to cost-optimize are just crazy.
IIRC Thinking Machines had a mode with deterministic numerics but it's more expensive to run due to limitations this imposes on cross-batch ops and ordering of floating point reductions, and their model is not great overall.
This kind of thing (plus the cost) really limits what they can realistically be used for. A lot of things are tolerant of even lots of fuzziness (suggestions you can ignore, work you can redo, etc), but that subset of applications doesn't justify the boggling capital investment or the ongoing compute needs.
So, my guess is we're probably in for a couple more years of discovering what these models are good for. Coding: meh, kinda. Hacking: wow amazing. Writing a novel: no. Reviewing your work: incredible. And so it goes. This is probably what pops the bubble: we find the small subset of applications this stuff is useful for, and then it's a bag holding race.
The reason the firms do not want to invest in making fact-checking a first-class feature is that the appearance of being right is what people want from AI.
No, actually being right is what people want from AI, "the appearance of being right" is all that AI companies can deliver. It comes with the benefit that many people will be fooled into thinking that AI is more capable/useful than it actually is. AI companies have to either convince others that their product is something that it isn't, or that at least it will one day be something much more than it is.
You are considerably more optimistic than I am about how people use things like this. IMO people are happy with these tools if they can be used to support their existing positions and biases.
Even vibe-coding is like this. OK it creates code that compiles, but its primary job is still to confirm a bias. I have yet to see people making significant novel discoveries about functionality this way.
The stuff about context control has always been my itch. The scrollback that most agents show is not what the model is reading. Things get summarised, dropped, cached or never included at all, and the transcript carries on showing the original as though it were still there.
It irked me enough to do my own agent: https://juggler.studio, explicitly to offer hands-on with the real context. The UX is all about making it easy to navigate and visualise every bit of the context, and even let you edit it. While it feels like other harnesses are actively trying to hide it from us..
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 5 days ago
- The Verge · 0 points · 3 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 10 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 12 days ago
- Hacker News · 73 points · 9 days ago