41 comments
"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."
The vote is currently 64% yes, 18% no.
Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...
Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).
"It’s ASCII art of a bare foot and lower leg, with the toes pointing to the right."
No tool calling, just an immediate reply with the correct answer.
> A bare foot and ankle, pointing right, with three little toes.
I wonder how much of the wide variation in perceptions of LLM capabilities is driven by the gulf between free models and frontier models. Luna getting something wrong is not always great evidence for LLMs be unable to do that thing.
Edit: for curious skeptics without access to 6.1 Sol, I tried 3 times and it got it all 3 times. Convo share link: https://chatgpt.com/share/e/6abeb955-7614-832e-a5e1-b1bd134f...
Like, is this an ice-cream? A tooth?
Because "for me DeepSeek Flash 4.1 nailed it immediately", trust me bro.
(_)(_)(_) represents the wheels
They do look rather wheel-like; I have to assume you see them as toes though?It's like the duck-bunny picture to me. If I focus on the "wheels", I see a steam train locomotive (but perhaps I'm only seeing that because I read your comment?); if I look at the ankle I see a foot.
I think I would have failed this test!
And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.
But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.
That alone tells you a lot IMO
Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?
More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you'll trip them up in a way that actually will not work on ESL students and children. Certainly there's no way to find a person that struggles with that and is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it's getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn't be and that's a signal too.
The best progress we've made is that most people do agree that this doesn't practically matter very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that "of course it's a solved problem!" always feels naive or fake.
jerf, 2024: "If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn't what I was talking about as "high level math".
"I also am not surprised by "Consider a generation function" coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit "have you considered using wood?" is not a system that can build a house autonomously.
"It especially won't seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs."
The voting gloss: "An AI fully solves a research-level math problem on its own, not just suggesting an approach."
Yes, I'm satisfied. I don't even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from "lol, can't add two six-digit numbers" to research-math level almost overnight in comparison.
To your point, this example. The issue expressed here is with humans, not AI. We are still pretty terrible at writing specs. TBF, the AIs are too but that wasn’t being voted on.
Its fun. Can you add a sort by controversial? I'd like to know where people disagree the most between yes and no.
Sure, sure, what LLMs make still isn't "efficient bug-free code": my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.
In your case, the comment you link to says „business tasks“ and you expanded it now to „arbitrary new tasks“. Those are not the same. An LLM today sure can do many many many business-speak conversion tasks.
Not reliably, and not without supervision. That's the main point. I'm trying really hard to figure out a workflow that doesn't require me to review the code and I just don't see how it's possible (yet)
You either need a comprehensive test suite (which requires understanding the code in order to create) or you need to review the actual implementation code to make sure it does the right thing
Consider I was replying to this:
> So are we all going to be out of a job?
While your boss now has the capacity to ask Claude to train a new AI model to auto-balance a tower defence game's mob, cost, and tower parameters (I know because I've done it), this only matters if you and your boss are working in a video games company.
If you and your boss are actually florists, you care if your boss can get Claude to automate a rose pruning, dead-heading, and fertilising robot.
People are trying, but I don't think they'd be happy with 91.5% success rate: https://www.emerald.com/ir/article-abstract/doi/10.1108/IR-0...
This means it has to handle basically all business tasks, so "arbitrary". I'm not sure what percent you have in mind by "many many many" but I would say it can't code half the things you need in an efficient and minimally buggy way.
I actually think this would take AGI to solve, which makes me optimistic about the future of software development.
All the benchmarks are currently testing against automated tests the AI can use as an oracle
if/when you can tell a model to do a thing and be confident that it did the thing, it's joever for 90% of knowledge workers.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 3 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 10 days ago
- The Verge · 0 points · 5 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 10 days ago
- Hacker News · 73 points · 9 days ago
- The Verge · 0 points · about 8 hours ago