The new era of tech seems to be built on superstitious behaviour
309 comments
His “imitation game” had three participants: a human participant, a computer participant, and an interrogator. The interrogator’s job was to talk to the participants and try to determine which participant is human and which is a computer.
He wasn’t interested in computers being able to fool the interrogator on occasion. The point where he thought the question of whether machines can think becomes moot is when the interrogator is unable to do much better than chance over many trials.
That’s a pretty high bar, and I don’t actually believe that LLMs have closed the gap with it by all that much. They still have so many obvious tells. And those tells are something Turing anticipated and accounted for. He explicitly considered deliberate deception as an essential part of the test, right there on the second page of a 30-odd page paper.
Frontier Labs are not interested in having LLMs being able to pass as humans. If anything, they explicitly train them not to. In many ways, this ability has regressed severely since the original GPT-3 with no instruct tuning or RL. How many 'tells' would there be really if a frontier model trained with frontier techniques is optimized to pass this test? I think this was something Turing did not quite forsee. That such machines might be created but not really care about this specific shape of the test. Regardless, i think his broader point about functional equivalence is spot on.
He proposes his game grounded on functional equivalence, then goes through a slew of objections on the question of 'Can Machines think?'. It's a terrific, very prescient read, and there's no objection you hear today (and in the last few years) concerning LLMs he didn't address.
I don't think Turing intended the judges in the test to be completely arbitrary people.
Oh but we have. Claude "How can I help you today?" etc. Undoubtedly there are users whothink this is a sign of intelligence.
Situations like this are precisely why academics tend to avoid the spotlight. You say one slightly off thing and your perceived authority echoes forever with the intellectually lazy.
citation needed. It has been used as a rubicon for a long time. Ever since Eliza, at least. And there were big headlines and lots of talk around the time LMs became "good enough". I specifically remember when someone had a test done around "a teenager talking in a different language" or somesuch, claiming it was the first time the test was passed.
It is pretty normal that once it was unquestionably "passed", lots of people started claiming it wasn't even that big of a deal. Tesler's theorem and all that.
And even if you think the specific formulation of Turing isn't that important (and I'd somewhat agree), you can still use the concept to look at other things. Imagine asking a mathematician 5 years ago the chances of a Erdos problem being solved by a computer end to end. Or a millennium prize. Or ask a swe if a repo could be generated by a computer from the input "write a mario style game", or any other examples of proven expertise.
The one thing this 2023 article gets partially correct imo is that any intelligence we see in AI (as of 2026) is our own - not that it’s a mirror but that the intelligence comes from the way that the words are put together, which comes from written human language created by (allegedly) intelligent creatures put in as input in both the training and prompt, among other places.
Rearranging and repeating the words, even in context, does not intelligence make. I’m not even convinced that you’re intelligent, dear reader.
Most people on Earth try to put what they're seeing into the context of what they understand; mental gymnastics to try and understand what is happening based on their prior experience. They have absolutely 0 understanding of how it works under the hood so, to them, it must be alive.
An LLM with CoT is Turing-complete. Training is, basically, compression (the training data gets lossily compressed into the model's weights). The information-theoretic limit of compression is an algorithm that reproduces functionality of a system that produced the training data.
No "magic" is required to get to a system that reproduces at least some facets of the human brain functionality.
Three years ago I was skeptical that stochastic gradient descent (and other known techniques) are the way. But evidence kept piling up.
I can't quite put my finger on it, but aren't these two statements add odds with each other? Intelligence is hard to define, consciousness even more so, but wouldn't "intelligence" imply some sort of agency? If not, I'd argue computers were intelligent long before the age of LLMs. And likewise, doesn't a tool imply the lack of intelligence and agency, even if the tool's function is very elaborate?
I got the impression that both these statements are made by the same people, or at least people with similar takes on AI. Is that wrong and there are "intelligence" and "tool" factions? Or do people disagree with my assumption and there's nothing wrong with the concept of "intelligent tools"?
Kinda refreshing this discussion, compared to the builder vs. tinkerer debates, imo.
Given a certain world state (including a tool's internal state), its effects back on the world state (as initiated by me) are at some leve of description understandable, expected and repeatable. Swing hammer, drive nail into wood. Make slicing motion with knife, cut meat. Type ' find /path/to/some/dir -name "keyword"', find files with keyword. Point harness at codebase with prompt 'fix bug X', actually fix bug X.
All these examples are at some level of description incredibly complex (think of all particles interacting at the (sub-)atomic level even when using a hammer to only drive a nail into some wood), and of course all the electrons flowing through the GPUs doing matrix multiplications in order to fix bug X, but at some level of description (the one I just used) they are also incredibly simple and understandable.
Intelligence is rather nebulous (and as used by OpenAI/Anthropic, quite threatening), but I don't think this definition of a tool precludes it to be "intelligent". They feel more orthogonal. The intelligence (or perhaps capability) feels like it is related to the size of the chunk of the world state that it can take into account and affect, while still resulting in understandable, expected and repeatable effects. LLMs, when properly harnessed, are pretty great at this currently and we are still discovering what they are consistently capable of.
Calling harnessed LLMs tools is perhaps also a more grounding frame specifically to counter-act the anthropomorphizing framing that OpenAI and Anthropic consistently go for in their game of AI-doom-chicken talk. The tool framing is in that sense maybe a (self-)jedi-mind-trick.
None of this matters for the practical outcome.
You'd think that this has been understood over the last 4 years, but apparently it keeps circling back to this.
[Edit: I see that it was written back in 2023. Then (2023) should be added in the submission title]
If it generates functional output that works, then it works. And it works. It's not a psychic's con when it outputs Lean-verified proofs. It isn't a con when it can find and exploit zero-days.
The OP is still in the "denial" phase. Most I see are already in "anger" (a blurry fury against everything AI-shaped, from vague reasons piling on all "bad stuff" political reasons they already hated before) or "bargaining" (mathematicians scrambling to come up with a new definition of their job and retcon that it was always the main part anyway). A few are already in "depression" and feel like spectators on the Titanic, and the tiniest sliver is at "acceptance" with some kind of well-informed plan for their future.
If you ask a common question to an LLM with unusual qualifiers, it tends to ignore the qualifiers and give you the typical answer. I saw a demonstration of this with the whole "the surgeon is my mother" "puzzle" that people use to expose implicit gender bias (ie where they assume the surgeon is a man). Ask variations of this and it'll keep going back to the standard form.
Another one I saw was multiplying large numbers. The starting and ending digits tended to be correct but the middle digits were wrong. Why? Because it's really not doing multiplication at all. It's looking for statistical answers. It's unlikely to have met the exact pair of very large numbers you're multiplying before.
Now pundits will argue that all of these are solvable problems and individually they are. But my suspicion is that there will be a neverending stream of such edge cases and it'll be impossible to trust an LLM's output unless you are knowledgeable enough to fact check it yourself.
Now if your example of identifying zero days, this comes up with what I can only describe as "light positives", meaning it's technically a bug but essentially impossible to exploit. IIRC this came up with the demonstration where someone pointed Fable at some BSD code. I'm not sure if there have been any true false positives and obviously false negatives are impossible to know.
I guess my point is that I think LLMs are way more limited than a lot of people think.
I don't know about that. We're going through maybe the biggest wave of electrification and growth demand in history. Electricians are still quite expensive for regular people to hire. There's lots of room for growth.
I think the posts mental model of stastically likely prompt completions is spot on.
This was written in July 2023. ChatGPT was released November 2022. No matter your views on AI, surely you can't blame the OP for writing this after a few months ChatGPT was released.
https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-g...
Of course people were skeptical:
https://www.reddit.com/r/slatestarcodex/comments/aslze7/gpt2...
Bad takes are bad takes. The author was perplexed by the "many people" convinced that models are intelligent and is argued against opinions/arguments he's been exposed to. Called proposed use-cases "borderline fraudulent pseudoscience."
He also published a second edition of "The Intelligence Illusion" in Sep 2025, so seemingly still stands by (some variant of) this belief.
It's an interesting reminder of how much general discourse has shifted since 2023 (I haven't heard of stochastic parrots in months!) but being wrong early doesn't change that.
There is a difference between you not liking something and it being wrong. Maybe think about that with your grey matter.
Read the full thread on Hacker News →
Related stories
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- How Big Tech Runs Tech Projects and the Curious Absence of Scrum (2021)newsletter.pragmaticengineer.comHacker News · 1 points · 10 days ago
- How Big Tech Runs Tech Projects and the Curious Absence of Scrum (2021)newsletter.pragmaticengineer.comHacker News · 1 points · 10 days ago
- Hacker News · 1 points · 8 days ago
- Hacker News · 5 points · 6 days ago
- Become Worthless (To Tech Companies)coryd.devHacker News · 4 points · 4 days ago