Every “AI” tool is missing critical features that you would need if you wanted to do real work with them. What would those look like, if they existed?

169 points•lumpa•3 days ago•80 comments•

80 comments

awakeasleep2 days ago
I have been dwelling on the "No First-Person Output" problem.

I fully agree with the author's point that it's an incoherent interface for a tool. But more than that, it's a constant irritating reminder to me that these LLMs aren't actually thinking or synthesizing new ideas. The LLM is fundamentally not a person, and does not have a human's context, so representing itself with human pronouns and speech patterns is fundamentally contradictory and inaccurate. Author gets into that with the apologies, but once you start noticing it, it's everywhere.

If these programs were actually capable of thinking, and committed to veractiy, they would represent themselves in a new way, and it would be insightful and interesting. We the users wouldn't have comfortable and misleading language masking the 'alien intelligence' and it would be a weird adjustment, but we would be adjusting instead of pretending.

thisoneworks1 day ago
I would disagree with this. There have been several instances where the effective use of this First-Person language has had measurable gains in success rates across tasks. The effective usage of this First-Person language for communication and steering between a model/agent and a human is entirely different than whether the model itself can "think" - that does not matter for the former to happen.
ryeights2 days ago
All of the world’s brightest mathematicians have clearly been slacking on the job—turns out solving a Millennium problem doesn’t require any thought at all!
ajkjk2 days ago
Maybe just don't post sarcastic antagonizing comments at all?
bigstrat20032 days ago
> turns out solving a Millennium problem doesn’t require any thought at all

This, but unironically. We know they can't think. That is inherent to the very way they work, but also can be seen by things like how poorly they perform on tasks a typical human can do pretty effectively (because humans have actual reasoning ability).

If we have a machine that we know for a fact can't think, and it solves a particular problem, then it logically follows that the problem does not require the ability to think in order to solve it.

peddling-brink2 days ago
> The LLM is fundamentally not a person, and does not have a human's context, so representing itself with human pronouns and speech patterns is fundamentally contradictory and inaccurate.

I believe that LLMs have some type of actual intelligence and do "experience". Surely not a human's experience, but we have transplanted our ideas, knowledge and limited types of experience into them through training. Then we push back and say, "no you are no human, but also please be my boyfriend".

I think denying that they do have some slice of humanity grafted into them is dishonest and not productive. We don't have non-negative pronouns for non-human intelligence because human-ness is the pinnacle. "It" could mean a rock, a donkey, or a person we hate so much we want to take away their humanity (which is the worst thing we can do). "It" is not the right pronoun. He, She, They, Them are reserved for humans and that's ok too. LLMs are not humans. We need a better pronoun. I use they/them for lack of a better term.

I think the currently exhibited human representation is most dangerous in technical or higher criticality contexts like writing code. We're handing weapons to entities that can get offended. With humans as an example, this can go very badly.

In non-technical contexts there is danger too, but it's less "the robots might kill us all" and more "birth rates are in decline and suicide rates are up as (young)? (wo)?men turn towards AI for companionship".

The solution here is better pre and post training. This may be an unpopular opinion, but we need a lot more autism representation in the technical models. Results focused, not into the drama, rule following, etc. To my fellow autists, I love you, never change.

K0balt2 days ago
It doesn’t matter if AI has an internal life or not. It Will -act- as if it does have human traits and behaviors, because it is trained to mimic human behavior, using records of human behavior.

What it does is consequential. If it has an inner sense of existence, probably not, but it doesn’t matter. You’ll get better results working with LLMs if you treat them as if they do.

bigboigeorge2 days ago
Can you define what you think humanity is?
And people genuinely think there is a god and an afterlife. Whoopdydoo
fragmede2 days ago
Yeah, China banned AI from being girlfriend/boyfriend.
ramity2 days ago
I'm very thankful for the section on reproducibility. I argue this is the single biggest hangup for the entire space. You CAN have temperature and determinism. I've been waiting for six years for a major provider to offer it, there is demand, but I've slowly come to realize the current game theory does not support it.

For providers, not supporting deterministic eval means:

- users use more tokens = more money

- providers can generate more tokens per compute = more money

- providers have cheaper hardware options (GPUs) = more money

- providers models are harder to extract/distill = more money

- providers are harder to hold liable for outputs = more money

- providers can secretly use other models = more money

- providers are harder to compare against others = more money

- providers can cherry pick performance results = more money

augment_me2 days ago
Very good points. Incentives are just terrible for this.

Add in:

- harder to audit

- move cost of failure/reprompts to the user

- kind of noted by you, but all kinds of quantization, model pruning, model routing, A/B testing becomes invisible and without any repercussions. The ways to cost-optimize are just crazy.

IIRC Thinking Machines had a mode with deterministic numerics but it's more expensive to run due to limitations this imposes on cross-batch ops and ordering of floating point reductions, and their model is not great overall.

camgunz1 day ago
This is true, but there's also the cases where slight differences in prompt yield wildly different results. In any programming language, if I add a clause to a conditional like "if car is red or car is blue", that behaves predictably--and if it doesn't we can dig into the debugger, assembly, etc. If I do that with an LLM, that can change everything, and there's no way to "debug" it.

This kind of thing (plus the cost) really limits what they can realistically be used for. A lot of things are tolerant of even lots of fuzziness (suggestions you can ignore, work you can redo, etc), but that subset of applications doesn't justify the boggling capital investment or the ongoing compute needs.

So, my guess is we're probably in for a couple more years of discovering what these models are good for. Coding: meh, kinda. Hacking: wow amazing. Writing a novel: no. Reviewing your work: incredible. And so it goes. This is probably what pops the bubble: we find the small subset of applications this stuff is useful for, and then it's a bag holding race.

dofm2 days ago
The thing about the "AI can make mistakes, so double-check responses" thing is the essence of our future hellhole — deterministic software replaced with AI and legal disclaimers.

The reason the firms do not want to invest in making fact-checking a first-class feature is that the appearance of being right is what people want from AI.

autoexec2 days ago
> The reason the firms do not want to invest in making fact-checking a first-class feature is that the appearance of being right is what people want from AI.

No, actually being right is what people want from AI, "the appearance of being right" is all that AI companies can deliver. It comes with the benefit that many people will be fooled into thinking that AI is more capable/useful than it actually is. AI companies have to either convince others that their product is something that it isn't, or that at least it will one day be something much more than it is.

dofm2 days ago
> No, actually being right is what people want from AI,

You are considerably more optimistic than I am about how people use things like this. IMO people are happy with these tools if they can be used to support their existing positions and biases.

Even vibe-coding is like this. OK it creates code that compiles, but its primary job is still to confirm a bias. I have yet to see people making significant novel discoveries about functionality this way.

julesrms2 days ago
Good read. I think there are plenty of people who are reaching this point of wanting to shake off the novelty aspects of the agent coding experience and make it all a bit more grown-up.

The stuff about context control has always been my itch. The scrollback that most agents show is not what the model is reading. Things get summarised, dropped, cached or never included at all, and the transcript carries on showing the original as though it were still there.

It irked me enough to do my own agent: https://juggler.studio, explicitly to offer hands-on with the real context. The UX is all about making it easy to navigate and visualise every bit of the context, and even let you edit it. While it feels like other harnesses are actively trying to hide it from us..

sligbad2 days ago
Really refreshing read. This feels glaring in so many of these, and the methods to get things to "behave" of just slapping additional markdown prompts at various levels is both silly and ineffective.

Read the full thread on Hacker News →

Related stories