One common take on the coding agents that I see goes something like this: “Sure, AI helps you output more code, but won’t the quality suffer?”
169 comments
However, I just don't think that's realistic. It's asking an author to suddenly become an editor. It's asking somebody who writes code to now read and debug others code.
It can actually be harder to find the the bug in a tricky piece of code than it can be to write your own correct code from scratch. I see AI introduce all sorts of bugs all the time in my personal projects that I would never introduce, and would never think to test for, especially around anything graphical.
This has been a big part of the job for anyone on a team for at least 20 years. I do agree that it’s the hardest and worst part of the job, and has now become the majority of the job for anyone who isn’t vibe coding. So, that sucks.
The largest problem these days is the volume of code developers are expected to review. The volume went up significantly.
Anybody who has reviewed pull requests can tell you that sooner or later you approve a PR after many rounds of changes because it's finally "good enough".
Fighting with a robot to just do the damned thing is less fraught because they don't get offended by critiques but it takes more round trips to get them pointed in the direction you want.
That's exactly right. Open source projects are currently drowning under LLM generated PRs, where those who used to write code are simply punting that work to AI, but still expecting others to review it. It's not okay to expect such a free lunch. If you moved the labour of writing code one step away, then you are yourself the first line of defence now, so you better start reviewing code that you claim to be yours.
I expect the same from colleagues, I'm not interested in treating them as a middle man between me and Claude.
Writing code has always involved reading and debugging your own code, at an absolute minimum, even if you did everything solo. In any remotely serious collaborative effort, it also involved code review and collaborative debugging; people use issue trackers and assign themselves and each other "tickets", which often involve fixing issues that are ultimately caused by someone else's code.
> It can actually be harder to find the the bug in a tricky piece of code than it can be to write your own correct code from scratch.
Part of the point is to reject tricky code exactly because it is tricky (as this is rarely actually necessary).
I think the bigger issue (like many things in software engineering) is a management issue. Once upon a time, I could take a look at the final output of a project and if it looked like a Ferrari on the outside, I could have some confidence that there was a good engine under the hood. OF COURSE THIS WASNT ALWAYS TRUE, but something that looked good, or was performant, or whatever, was a decent proxy for the code underneath being good. And with a smart human, there were ancillary things. Having spent 20 hours coding something, they probably thought through the edge cases that their manager, or product team hadn’t considered.
With AI, everyone’s output looks like a Ferrari, so it is hard to know what the internals are like.
A lot of people will probably look at this and say “well you need better management”, but better management has always been elusive in software engineering. Furthermore, reviewing AI generated code is soul crushing work and I don’t know who wants to do it.
In my guesstimate the number of good engineering managers out there is actually very very small and in practice, the best managers that I’ve seen are the ones who don’t think they are good managers, so they just set a very high hiring bar and hire people who don’t need much management.
Based on my experience, this is a significant issue with AI generated code. You wouldn't expect a real Ferrari supercar to have random internal mechanical components that are, for no reason, completely inappropriate for a high-speed car design. With AI generated code, such inappropriate components can appear randomly at any point in the implementation stack. And very often, they are deeply buried under non-trivial algorithms and are thus not easy to spot.
I’m doubtful that the author’s recommendation always work, but I do some similar things and they do seem to help.
I'd like to think the time and practice I've put into software engineering has made me better at it. If that's not true, then there's no reason to prefer senior or principal engineers with years of experience over newcomers.
You've missed the point. Nobody doubts writing code well or badly is indeed a skill issue.
The question is that "once you account for all of the things you need to do to make the code very high quality, did vibe coding actually provide any real value?"
I'm certain there are guardrails that help bolster vibe coding but I'm equally certain that when ive prompted something important I usually have to redo it enough times that just writing it manually myself usually would have been quicker.
Then I watch other people who code who dump on that opinion and I see total slop. They just can't tell the difference.
That wasn't the case with 00s agile / Uncle Bob stuff since proving that any of it was helpful was impossible - you just had to believe (and if you didn't believe there was something wrong with you!).
Though arguably more of a process and judgement issue than skill.
What makes LLM-generated code a bit special there is that misjudging how to deal with it seems to be what most people do. So the default is broken.
Whereas in prior iterations of "skill issue", the default was working.
1. they don't care
2. the rest of the team doesn't care
3. the powers that be actively discourage it because velocity.
I think where things get dicey is being able to write in any language. I write and review code in many languages and frameworks I'm not fluent in, so it's hard for me to distinguish between working code and great code. I can spot when the fundamental logic is wrong, but when it comes to "best fit" choices I'm clueless.
The thing is, if you follow SWE best practices indiscriminately, then you ll have a shit code base in no time.
There is no silver bullet, and no replacement for experience and mindfulness.
At work, with a team and code reviews, the 50%-100% figure seems much more likely.
This can probably move towards the more spectacular productivity gains as the AI's output becomes more reliable, people realize this, and less time is spend on code review and cleaning up the output.
This is something I have been trying to get right as well. I've attempted to use lots of linting and things like strong typing, duplicate checks, cyclomatic complexity, and robust tests. However, I still happen to find issues, which requires me to look at the code (at least at a high level)
For example, I can say "Don't repeat yourself, and don't re-write helper functions" and I will even have a duplicate linter check, but inevitably the LLM will always want to re-write a similar yet slightly different helper function. Like it will always want to re-write something small like a trim() or a toString() function in every file.
But if I add a separate post-implementation pass to "find and fix X" by the agent, it'll usually find and fix the issues.
So I've started doing it for everything from naming conventions to duplicate code to other problems. It does cost more tokens, but now I get less frustrated at having to fix basic issues in the PRs.
I have better luck telling it positive things rather than lots of "never do X" style things.
It's worth reflecting on why these things are important to you and whether they remain important in an agent-developed codebase.
When implementing new features or making large refactoring changes; I use the superpowers:brainstorming skill. That has consistent process which has worked really well. I alway review the code before merging, but most of the time there are few issues to correct.
I don't do 95% coverage, but I have increased it from 65% to about +80% and that is sufficient.
Thanks for adding such a thoughtful and level-headed comment to the discussion.
Kindly share the metrics by which you evaluate the changes.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 2 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 4 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 11 days ago
- Hacker News · 73 points · 9 days ago