Tell me you don’t understand software without literally using those words!!!
540 comments
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.
Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.
Testing isn’t the same as understanding the code, or proving (even informally) that it is correct. Having the LLM do all these things above doesn’t lead you or the LLM to understand the code, to logically reason about its behavior over all possible states and inputs.
“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.
Moreover, when your codebase is hundreds of thousands to millions LOC, I question how much you can ever truly understand it at the level you’re saying.
We're not writing theorems, dude.
Except in the equally pedantic sense that every program is a proof to a theorem...
We're writing plain enterprise and web software, closer to CRUD than NASA.
If you said that even before LLMs 0.1% of teams "checked all assumptions against what the code and underlying systems are actually guaranteeing" in any kind of formal way, you'd be overestimating it.
Reading the code may not be enough to understand the behaviour of your program, but believing you can understand the behaviour of a program without at least reading the high level code is truly silly.
(by high level, I mean the code living in the higher layers - of course we don't often read the code of the generated assembly, or the interpreter, or the browser, but that's because they're reliable abstractions, unlike prompts!)
have you ever used a library after only reading the README and documentation, or do you always pull the source and read through it before you think you understand it?
This sounds like a typical testimonial whose mind has become captive to Claude. It is like Scientology.
if you're an MBA-brained exec who doesn't actively use LLMs to code and you just believe whatever slop it outputs at first without checking it, you're not going to realize how recklessly it can be used, how you need to be critical and skeptical of its outputs, that you need to explore it's reasoning and logic (which is still really easy compared to understanding legacy code and barely takes any time!)
say you also believe all this marketing hype about 'how dangerous (ie capable) AI agents are.' LLMs can do anything you think so you just say 'ship it' without building out the tooling and capabilities to enable faster code review and better tests. and to keep the shareholders happy, you start cutting jobs that you can't directly connect to a KPI (ie the platform/SRE team who would be the ones who can trial, onboard, and maintain those capabilities for your teams)
and from this, suddenly a lot of debit card stops working and the only one getting the blame are individual SWEs trying to hit their sprint velocity. the fact that you fucked up the whole SDLC real bad with your incompetence gets you a golden parachute and you job hop to a better paycheck. rinse and repeat
Coding can have become cheaper and better. A lot of people who would not have been able to write software before now are. Add the ones that already have been able to write software, who will now use AI to do it better, and you might net "software is getting worse".
They are not good at that. The space of possibilities can be massive and LLMs are terrible at exploring such space because they predict from the prior tokens they made. They are inherently bad at exploring new space because it is antithetical to how they work.
To me it's deeply concerning that so many people are getting fooled into thinking that LLMs are actually good at covering their bases like you're describing here. It's one of their weakest qualities.
> Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible.
It usually does quite a bad job at this too, often the tests it wrote feel like that of a lazy student that didn't really want to do the task and just sort of cheats at it or does a really shallow job. It certainly cannot run the software in every scenario possible.
> Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while
This is something that is said by everyone who contests anyone pointing out the risks in being overly trusting of AI or otherwise points out their flaws. I use the latest and greatest all the time and all the time I'll point out something that it totally overlooked and get hit with the classic "you're absolutely right". This occurs because I actually read the code and can see the myriad of blatant issues that still occur when using LLMs and know better than to trust them. You will find so many issue by delving into the details.
People ask AI to do things either they haven't done, cannot do, or don't want to.
So when the AI produces something that basically works and looks like a plausibly good attempt, they jump to the conclusion that AI must be excellent at that task.
Correctness has never been a priority across an industry where rapid iteration and feature delivery drive sales. There's always some opportunity cost to doing things right, at the price of technical debt down the road. If AI is primarily used to produce fragile code, people will be wary of AI solutions. There's also ongoing public debate about AI safety and alignment. Deploying AI in safety critical applications feels riskier than ever in the current environment, even though it doesn't have to be.
I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers.
Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.
Brings to mind this classification https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Cl...
"""I distinguish four types. There are clever, hardworking, stupid, and lazy officers. Usually two characteristics are combined. Some are clever and hardworking; their place is the General Staff. The next ones are stupid and lazy; they make up 90 percent of every army and are suited to routine duties. Anyone who is both clever and lazy is qualified for the highest leadership duties, because he possesses the mental clarity and strength of nerve necessary for difficult decisions. One must beware of anyone who is both stupid and hardworking; he must not be entrusted with any responsibility because he will always only cause damage"""
Now instead of 90% stupid and lazy (harmless, useful for grunt work) you have 90% stupid and hardworking (aggressively causing damage).
A couple more step functions in model capability of the type we've seen in the past year, and there will pretty much be no reason for humans to be involved in the development process at all. All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.
> I hope you can see the stupidity here if you expect to see any deterministic results at all.
Are you expecting humans to be deterministic in the code they produce?
The behaviour/output of an LLM is not like that. Ask an LLM to create a dashboard to show games by genre and it will generate different results with each run, and each model/model version produces wildly different results.
As humans we don't have our memory reset multiple times per day
I totally understand where this is coming from. I too am struggling with accepting that my 30+ years of programming experience is quickly becoming obsolete. I'm losing sleep about this, it's tough.
But just go ahead and give the latest models (Opus 5.5 / Astra 6 as of today) another try. See what they are capable of and read the code which they produce. Any problem area, low level C++ or high level Typescript or Clojure or a weird combination of these..
Don't be shy, give them a big task, let them build an entire app, UI and all..
Now compare the output to Opus 4 or gpt-5 from 1 year ago - when they couldn't put together a single function without it being weird and buggy.
This is exactly my problem, not that the models are very good already, but how fast they got so good. So if coding is not solved yet, it'll get there very soon.
It’s really not an acceptable opinion to have, that one day we’re going to be developers drowning in ai generated code. That’s a problem for opus 6 in 4-5 weeks time.
Developers are still very much needed, business people just don’t have the chops to wire a deterministic system together. But acting like some poor class naming or 2k loc file is the disaster it once was is dated.
I don’t see the ops comment as a rebuttal. He agrees coding is not completely solved. However, it’s getting closer to being solved.
What kind of coding are you using these models for? Most of the people I know who share your perspective never go beyond the prototyping stage. I’d be curious to hear from anyone who’s been AI coding for more than six months, shipping it to real users, and isn’t looking at their code at all.
Feels like I read comment similar to this one each year since 2023.
That's when the models started to be coherent enough for real work.
They still fuck up, but it does not feel the code was written by drunk interns anymore.
Well, for one, they're capable of draining our (or companies') wallets.
I resolved a huge merge conflict for $60 today. Opus 5.5 did a great job and spent just 1h 16min on this. I could probably run six such sessions today, if I disregarded the need to read and understand the code.
This money has to come from somewhere and my concern is that it will be from decreasing the number of people hired and/or their salaries.
At the same time I firmly believe people who had a tendency to produce tech debt will keep doing that, regardless how brilliant LLMs will become. Unscrewing this is going to cost a lot of money.
Solving a merge conflict for $60/hr is not so bad, but what if it's $600/hr?
This is not a good premise. All over law, you will find people made responsible for what they don't control and they kind of own. Unleash a dog that harms a child, or just have it in an environment where it can escape, and see what happens.
There is such things as unpredictable situations where one might not be held responsible, as a problem might occur well past reasonable guidelines.
So of course you can be held accountable for what an AI that uou supposedly cannot quite control does, or for the AI-written code you deliver. Treat it like the releasing a wolf pack, or selling an unsafe toy that can maim children. There's precedent everywhere.
The difference seems to be that some companies are above the law apparently.
i guess another way of saying this is that on the micro level a lot of this stuff is not rocket science, but at the macro level it becomes hugely complex.
Read the full thread on Hacker News →
Related stories
- Lobsters · 86 points · about 1 year ago
- Hacker News · 1 points · 7 days ago
- Coding is not solvedblog.alexewerlof.comLobsters · 25 points · about 16 hours ago
- Hacker News · 1 points · 7 days ago
- Hacker News · 3 points · 1 day ago
- DEV Community · 22 points · 10 days ago