Tell me you don’t understand software without literally using those words!!!

565 points•firstSpeaker•3 days ago•540 comments•

540 comments

efficax2 days ago
Reading the code does not mean you understand the code. One lesson that experience in software gave me: I never understood the code. You think it works a certain way, until you find out that it doesn't.

What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.

If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.

Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.

layer82 days ago
> I never understood the code. You think it works a certain way, until you find out that it doesn't. What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.

Testing isn’t the same as understanding the code, or proving (even informally) that it is correct. Having the LLM do all these things above doesn’t lead you or the LLM to understand the code, to logically reason about its behavior over all possible states and inputs.

“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.

emtel2 days ago
You’re technically correct, but the vast majority of software has never been built to the kinds of standards you are describing. LLMs are not displacing that kind of work!
ModernMech2 days ago
Your code is only as good as what you can prove. Understanding the code is not the goal, it’s only important insofar as it helps you evolve the codebase predictably and without bugs or regressions, and understanding is not easily measurable or transferable.

Moreover, when your codebase is hundreds of thousands to millions LOC, I question how much you can ever truly understand it at the level you’re saying.

coldtea2 days ago
>“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.

We're not writing theorems, dude.

Except in the equally pedantic sense that every program is a proof to a theorem...

We're writing plain enterprise and web software, closer to CRUD than NASA.

If you said that even before LLMs 0.1% of teams "checked all assumptions against what the code and underlying systems are actually guaranteeing" in any kind of formal way, you'd be overestimating it.

jdkoeck2 days ago
> Reading the code does not mean you understand the code.

Reading the code may not be enough to understand the behaviour of your program, but believing you can understand the behaviour of a program without at least reading the high level code is truly silly.

(by high level, I mean the code living in the higher layers - of course we don't often read the code of the generated assembly, or the interpreter, or the browser, but that's because they're reliable abstractions, unlike prompts!)

rco87862 days ago
> believing you can understand the behaviour of a program without at least reading the high level code is truly silly.

have you ever used a library after only reading the README and documentation, or do you always pull the source and read through it before you think you understand it?

12ag5a2 days ago
Strange that the world worked before 2024 and software gets worse now. Your debit card transactions for example worked.

This sounds like a typical testimonial whose mind has become captive to Claude. It is like Scientology.

bananaflag2 days ago
Before 2024, I once went to an ATM to retrieve money and selected 50. Note that I selected it from a menu, not typed it. The ATM then told me that it cannot give me 50 because it is not a multiple of 5.
TaLiTrabout 13 hours ago
Strange that GP never claimed software doesn't function without LLMs and yet you strawman-man them then make some bad-faith dismissal of their experience by accusing them of having psychosis.
MattDamonSpace2 days ago
Yeah no one wrote buggy code before 2024 right
paimapi2 days ago
I mean, I think the problem isn't that the LLM doesn't know how to code, it's that companies are expecting 3-5x velocity with the bottleneck of code review and testing becoming much more severe than before

if you're an MBA-brained exec who doesn't actively use LLMs to code and you just believe whatever slop it outputs at first without checking it, you're not going to realize how recklessly it can be used, how you need to be critical and skeptical of its outputs, that you need to explore it's reasoning and logic (which is still really easy compared to understanding legacy code and barely takes any time!)

say you also believe all this marketing hype about 'how dangerous (ie capable) AI agents are.' LLMs can do anything you think so you just say 'ship it' without building out the tooling and capabilities to enable faster code review and better tests. and to keep the shareholders happy, you start cutting jobs that you can't directly connect to a KPI (ie the platform/SRE team who would be the ones who can trial, onboard, and maintain those capabilities for your teams)

and from this, suddenly a lot of debit card stops working and the only one getting the blame are individual SWEs trying to hit their sprint velocity. the fact that you fucked up the whole SDLC real bad with your incompetence gets you a golden parachute and you job hop to a better paycheck. rinse and repeat

jstummbillig2 days ago
I am not sure that is true but I don't think it would be strange or incompatible if you consider base rates.

Coding can have become cheaper and better. A lot of people who would not have been able to write software before now are. Add the ones that already have been able to write software, who will now use AI to do it better, and you might net "software is getting worse".

OGWhales1 day ago
> What LLMs make possible is for me to say: find out all the ways this thing works.

They are not good at that. The space of possibilities can be massive and LLMs are terrible at exploring such space because they predict from the prior tokens they made. They are inherently bad at exploring new space because it is antithetical to how they work.

To me it's deeply concerning that so many people are getting fooled into thinking that LLMs are actually good at covering their bases like you're describing here. It's one of their weakest qualities.

> Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible.

It usually does quite a bad job at this too, often the tests it wrote feel like that of a lazy student that didn't really want to do the task and just sort of cheats at it or does a really shallow job. It certainly cannot run the software in every scenario possible.

> Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while

This is something that is said by everyone who contests anyone pointing out the risks in being overly trusting of AI or otherwise points out their flaws. I use the latest and greatest all the time and all the time I'll point out something that it totally overlooked and get hit with the classic "you're absolutely right". This occurs because I actually read the code and can see the myriad of blatant issues that still occur when using LLMs and know better than to trust them. You will find so many issue by delving into the details.

zero_shift1 day ago
It's a cognitive bias.

People ask AI to do things either they haven't done, cannot do, or don't want to.

So when the AI produces something that basically works and looks like a plausibly good attempt, they jump to the conclusion that AI must be excellent at that task.

flatline2 days ago
I don't think the discrepancy is in LLM capability improvements over the past year.

Correctness has never been a priority across an industry where rapid iteration and feature delivery drive sales. There's always some opportunity cost to doing things right, at the price of technical debt down the road. If AI is primarily used to produce fragile code, people will be wary of AI solutions. There's also ongoing public debate about AI safety and alignment. Deploying AI in safety critical applications feels riskier than ever in the current environment, even though it doesn't have to be.

ThrowawayR22 days ago
It would be deliciously ironic if AI was the straw that broke the camel's back where a deluge of bugs and anti-AI sentiment caused the public to vote for legal liability for software defects and licensure of developers. No more of this "no warranty, express or implied" business for us.
askonomm2 days ago
What I've found is that AI allows lazy and incompetent developers to be more lazy and more incompetent. This then has the effect that product quality suffers more, faster. As a result of the sheer amount of code now being pushed out, code reviews, a thing that previously somewhat prevented lazy and incompetent developers from pushing out horrible code, is effectively dead in the water since no human can actually review such amounts of code realistically anymore. Some companies have adopted AI to review code, which, well ... you have AI make code, AI review code ... I hope you can see the stupidity here if you expect to see any deterministic results at all.

I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers.

Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.

rgoulter2 days ago
> a thing that previously somewhat prevented lazy and incompetent developers from pushing out horrible code

Brings to mind this classification https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Cl...

"""I distinguish four types. There are clever, hardworking, stupid, and lazy officers. Usually two characteristics are combined. Some are clever and hardworking; their place is the General Staff. The next ones are stupid and lazy; they make up 90 percent of every army and are suited to routine duties. Anyone who is both clever and lazy is qualified for the highest leadership duties, because he possesses the mental clarity and strength of nerve necessary for difficult decisions. One must beware of anyone who is both stupid and hardworking; he must not be entrusted with any responsibility because he will always only cause damage"""

banannaise2 days ago
The problem here is that AI is consistently one of the four things: hardworking. This makes it very efficient at transforming "stupid and lazy" inputs into "stupid and hardworking" outputs.

Now instead of 90% stupid and lazy (harmless, useful for grunt work) you have 90% stupid and hardworking (aggressively causing damage).

Melkazt2 days ago
I'm both clever and stupid, depends on the day.
whatever12 days ago
Even if you are competent I cannot review your 5,000 lines of code you produce per day vs the 100 you were producing before the LLM apocalypse.
rfgplk2 days ago
5,000 is the output velocity of someone not fully immersed in agentic coding. I've seen repos do ~100k to ~250k loc changes per week.
bitwize2 days ago
That's okay. Reviewing the code will become the agents' job as well.

A couple more step functions in model capability of the type we've seen in the past year, and there will pretty much be no reason for humans to be involved in the development process at all. All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.

patorjk2 days ago
I'm seeing this too. I've worked with devs that would previously push PRs that wouldn't work or run correctly. Those PRs wouldn't get merged in. Now they're putting up PRs which seem to work at first glance, but have hidden problems. For example, one guy introduced a huge PR for a visualization and it seemed to work fine, though another dev mentioned to me that we already use recharts and it does 90% of what this guy's PR does (his code does all the drawing logic itself). Maybe AI will get good enough to clean up these kinds of messes, but in the near term I imagine there will be a lot of code bases that will be filling up with dragons.
bwfan1232 days ago
At a startup I worked, there was an engineer whose code was incoherent and buggy. So, we were literally better off if that engineer did nothing because their net output was negative. Engineers like that become weaponized with LLMs, and negative numbers become larger negative numbers when scaled up.
icedchai2 days ago
I've seen similar. They wasted weeks of senior engineering time, between reviews, meetings, and follow up in Slack, only to have the PR closed without merge. The offending individual was eventually moved to another project.
ilaksh2 days ago
Is that the fault of AI or management for not firing them?
ben_w2 days ago
Limitations of AI are a thing; but one rhetorical point keeps coming up (I don't think it's just you) and confusing me:

> I hope you can see the stupidity here if you expect to see any deterministic results at all.

Are you expecting humans to be deterministic in the code they produce?

Thanemate2 days ago
Someone who knows that 1 + 1 = 2 will not decide that it's suddenly 3 unless we start accounting for health problems. Making mistakes is not the same as non-deterministic.
rhdunn2 days ago
By not reviewing, reading, or understanding the code generated by agentic LLMs the output is effectively like a compiler. However, a compiler has deterministic behaviour that can be repeated and verified.

The behaviour/output of an LLM is not like that. Ask an LLM to create a dashboard to show games by genre and it will generate different results with each run, and each model/model version produces wildly different results.

ex-aws-dude2 days ago
I've seen LLMs do something correct 98% of the time then randomly do something crazy that a human would never do because we have continual learning

As humans we don't have our memory reset multiple times per day

lkjdsklf2 days ago
The difference is that with llms you have multiple levels of nondeterminism compounding each other
temp003452 days ago
I read such articles more or less every day. This article would be 100% correct if it came out 1 year ago, 75% correct 9 months ago, 50% correct 3 months ago and it's probably 25% correct now if not less.

I totally understand where this is coming from. I too am struggling with accepting that my 30+ years of programming experience is quickly becoming obsolete. I'm losing sleep about this, it's tough.

But just go ahead and give the latest models (Opus 5.5 / Astra 6 as of today) another try. See what they are capable of and read the code which they produce. Any problem area, low level C++ or high level Typescript or Clojure or a weird combination of these..

Don't be shy, give them a big task, let them build an entire app, UI and all..

Now compare the output to Opus 4 or gpt-5 from 1 year ago - when they couldn't put together a single function without it being weird and buggy.

This is exactly my problem, not that the models are very good already, but how fast they got so good. So if coding is not solved yet, it'll get there very soon.

o_nate2 days ago
We also read these kinds of rebuttals almost every day. The original article made a number of substantive critiques about where AI falls short in the actual requirements for building maintainable, reliable, business-critical software. So its not enough just to point to newer models without explaining how the newer models solve these problems. Can the new models take full end-to-end ownership of a system? If not, how do they solve the problem of humans taking ownership of AI generated code?
aunty_helen2 days ago
Yes they can and tomorrow the models will be better. By the time code hits prod and is used in anger the frontier will have moved.

It’s really not an acceptable opinion to have, that one day we’re going to be developers drowning in ai generated code. That’s a problem for opus 6 in 4-5 weeks time.

Developers are still very much needed, business people just don’t have the chops to wire a deterministic system together. But acting like some poor class naming or 2k loc file is the disaster it once was is dated.

NCFZ2 days ago
But most software isn’t critical and doesn’t need the level of reliability that’s sacrificed when using LLMs.

I don’t see the ops comment as a rebuttal. He agrees coding is not completely solved. However, it’s getting closer to being solved.

an0malous2 days ago
I use the latest models every day for app development, it does still make many mistakes and poor decisions. Just today I dealt with an issue where it allowed a silent failure that would delete all of the users data without anyone noticing. The code it writes for one task is usually pretty good, but it still doesn’t always follow conventions well even if you’ve documented them. It also has no sense for long-term architecture or organization, or when to make tradeoffs for less complexity because it’s a temporary or prototype feature that needs to stay easy to change. I’ve found that if I do even 3 months of pure agentic coding with no code reviews, it’s aged 10x faster than a human coded codebase so it’s like a 30 month legacy codebase now. I know people who are obsessed with AI coding and they’re throwing away projects they’ve built over a year and starting over because it’s become too slow to make changes.

What kind of coding are you using these models for? Most of the people I know who share your perspective never go beyond the prototyping stage. I’d be curious to hear from anyone who’s been AI coding for more than six months, shipping it to real users, and isn’t looking at their code at all.

ysajang2 days ago
I'd be curious to hear from anyone who's been AI coding for more than six months, shipping it to real users, and isn't looking at their code at all.
lnkl2 days ago
>This article would be 100% correct if it came out 1 year ago, 75% correct 9 months ago, 50% correct 3 months ago and it's probably 25% correct now if not less.

Feels like I read comment similar to this one each year since 2023.

throwawayffffas2 days ago
I can tell you it was not correct in January of 2026. The real swift started happening with the latest models opus 4.7, fable 5, kimi k3, glm 5.2.

That's when the models started to be coherent enough for real work.

They still fuck up, but it does not feel the code was written by drunk interns anymore.

lbrito2 days ago
I've been reading such arguments more or less every day for years now: dude, you need to try the latest model. Forget about last week's model, it didn't work. This week's model is the real deal.
Tade02 days ago
> But just go ahead and give the latest models (Opus 5.5 / Astra 6 as of today) another try. See what they are capable of and read the code which they produce.

Well, for one, they're capable of draining our (or companies') wallets.

I resolved a huge merge conflict for $60 today. Opus 5.5 did a great job and spent just 1h 16min on this. I could probably run six such sessions today, if I disregarded the need to read and understand the code.

This money has to come from somewhere and my concern is that it will be from decreasing the number of people hired and/or their salaries.

At the same time I firmly believe people who had a tendency to produce tech debt will keep doing that, regardless how brilliant LLMs will become. Unscrewing this is going to cost a lot of money.

hajile2 days ago
These companies are operating at huge losses, but are still cutting prices. I don't think companies are ready for what happens when the bills come due and they have to charge enough to not only be profitable, but pay back the years they've been blowing hundreds of billions of dollars.

Solving a merge conflict for $60/hr is not so bad, but what if it's $600/hr?

hibikir2 days ago
> You cannot be responsible for what you can’t control either. That understanding is key to reasoning about system behavior and fixing it when the AI inevitably fails.

This is not a good premise. All over law, you will find people made responsible for what they don't control and they kind of own. Unleash a dog that harms a child, or just have it in an environment where it can escape, and see what happens.

There is such things as unpredictable situations where one might not be held responsible, as a problem might occur well past reasonable guidelines.

So of course you can be held accountable for what an AI that uou supposedly cannot quite control does, or for the AI-written code you deliver. Treat it like the releasing a wolf pack, or selling an unsafe toy that can maim children. There's precedent everywhere.

hanifbbz2 days ago
Came here to give an answer but your last sentence kinda made the point I was gonna make. If one is legally in control, then one is accountable (the dog or unsafe toy example in reality is OpenAI's agents hacking huggingface for example).

The difference seems to be that some companies are above the law apparently.

thesumofall2 days ago
I think the author underestimates how boring and simple 90% of enterprise software is. The part that isn’t powering aircraft and power plants. So much of it originates from one-nighters, badly managed subcontractors, and requirements that are of low quality to begin with (because they are written by people who have very different day jobs). And you know what? Most of that runs 24/7 without a glitch. LLMs just gives us more of that. And maybe it’s even better
cautiouscat2 days ago
I’ve been working in enterprise for ten years and I wouldn’t call any of the services I’ve worked on simple.
thesumofall2 days ago
There is probably a marvelous software core in many large enterprises but at least in my experience it’s surrounded by layers and layers of very basic stuff. Read from a database, write to a database (often not even with any logic dealing with parallel read/writes). Excel Macros. Basic forms to submit data. Scripts to print some stuff from a database. …
balderdash2 days ago
I wish that were the case, my multiple experiences are quite different. its not that the software doesn't work - but the software doesn't actually capture the process, or properly talk with other systems, or isn't set up for a new business model or product line...then you have layers and layers of "tools" created to make it work, and a bunch of people using spreadsheets and csv files to do workarounds, and then the guy that wrote a bunch of the tool's 10 years ago leaves and there is no documentation or it needs to be re written in order to upgrade some other piece of the system.

i guess another way of saying this is that on the micro level a lot of this stuff is not rocket science, but at the macro level it becomes hugely complex.

mxey2 days ago
The majority of existing software runs without a glitch? Are you being serious?
thesumofall2 days ago
Yes. Bad UI, cumbersome flows, too little automation, … plenty of flaws, but the business just keeps on running. Orders received, invoices written, documents shared, …

Read the full thread on Hacker News →

Related stories