One common take on the coding agents that I see goes something like this: “Sure, AI helps you output more code, but won’t the quality suffer?”

129 points•bucket2015•11 days ago•169 comments•

169 comments

zug_zug11 days ago
I think this is a bit of a simplistic mental approach. I've certainly seen a lot of "The engineer owns the outcome, AI is just a tool, don't release anything you don't vouch for."

However, I just don't think that's realistic. It's asking an author to suddenly become an editor. It's asking somebody who writes code to now read and debug others code.

It can actually be harder to find the the bug in a tricky piece of code than it can be to write your own correct code from scratch. I see AI introduce all sorts of bugs all the time in my personal projects that I would never introduce, and would never think to test for, especially around anything graphical.

christophilus11 days ago
> It's asking somebody who writes code to now read and debug others code.

This has been a big part of the job for anyone on a team for at least 20 years. I do agree that it’s the hardest and worst part of the job, and has now become the majority of the job for anyone who isn’t vibe coding. So, that sucks.

OptionOfT10 days ago
Disagree. At least back in the day there weren't endless comments about how this widget is load-bearing, and how honestly the other widget carries the derived widget, referencing decision ADR-100 that is nowhere to be found. All these comments matter because once accepted as part of the codebase the next LLM takes these comments as canonical.

The largest problem these days is the volume of code developers are expected to review. The volume went up significantly.

phrotoma10 days ago
It's a different of degree, not kind.

Anybody who has reviewed pull requests can tell you that sooner or later you approve a PR after many rounds of changes because it's finally "good enough".

Fighting with a robot to just do the damned thing is less fraught because they don't get offended by critiques but it takes more round trips to get them pointed in the direction you want.

sameerds10 days ago
> It's asking somebody who writes code to now read and debug others code.

That's exactly right. Open source projects are currently drowning under LLM generated PRs, where those who used to write code are simply punting that work to AI, but still expecting others to review it. It's not okay to expect such a free lunch. If you moved the labour of writing code one step away, then you are yourself the first line of defence now, so you better start reviewing code that you claim to be yours.

sfn4210 days ago
That's what I do and expect my colleagues to do. Even before LLMs I was reviewing my own PRs before submitting them to others. I still do that. I work closely with Claude to create something good that I'm happy with, then I review it and test it to ensure it's good. And only then do I submit the PR to colleagues for final review.

I expect the same from colleagues, I'm not interested in treating them as a middle man between me and Claude.

CoolestBeans10 days ago
I agree. When you write your own code, you know what your intention was when writing it. Furthermore, as you gain experience and mature you know in the back of your mind that every mistake during code writing costs disproportionately more to fix later on. You only get that feeling by owning the code. AI cannot do that. It can't have skin in the game in that way.
zahlman10 days ago
> It's asking somebody who writes code to now read and debug others code.

Writing code has always involved reading and debugging your own code, at an absolute minimum, even if you did everything solo. In any remotely serious collaborative effort, it also involved code review and collaborative debugging; people use issue trackers and assign themselves and each other "tickets", which often involve fixing issues that are ultimately caused by someone else's code.

> It can actually be harder to find the the bug in a tricky piece of code than it can be to write your own correct code from scratch.

Part of the point is to reject tricky code exactly because it is tricky (as this is rarely actually necessary).

skybrian10 days ago
If you can explain how to reproduce a bug, you can ask the AI to debug the code and it usually works, in my experience. If not, you can ask it to add logging or other tools for better observability.
chadash11 days ago
I agree that agents can produce decent code. In general, I don’t find agentic code beautiful but neither is most of the code I write. The code for ingesting CSV files into my ETL pipeline doesn’t have to be beautiful, it just has to work.

I think the bigger issue (like many things in software engineering) is a management issue. Once upon a time, I could take a look at the final output of a project and if it looked like a Ferrari on the outside, I could have some confidence that there was a good engine under the hood. OF COURSE THIS WASNT ALWAYS TRUE, but something that looked good, or was performant, or whatever, was a decent proxy for the code underneath being good. And with a smart human, there were ancillary things. Having spent 20 hours coding something, they probably thought through the edge cases that their manager, or product team hadn’t considered.

With AI, everyone’s output looks like a Ferrari, so it is hard to know what the internals are like.

A lot of people will probably look at this and say “well you need better management”, but better management has always been elusive in software engineering. Furthermore, reviewing AI generated code is soul crushing work and I don’t know who wants to do it.

In my guesstimate the number of good engineering managers out there is actually very very small and in practice, the best managers that I’ve seen are the ones who don’t think they are good managers, so they just set a very high hiring bar and hire people who don’t need much management.

thw_9a83c10 days ago
> With AI, everyone’s output looks like a Ferrari, so it is hard to know what the internals are like.

Based on my experience, this is a significant issue with AI generated code. You wouldn't expect a real Ferrari supercar to have random internal mechanical components that are, for no reason, completely inappropriate for a high-speed car design. With AI generated code, such inappropriate components can appear randomly at any point in the implementation stack. And very often, they are deeply buried under non-trivial algorithms and are thus not easy to spot.

axegon_11 days ago
Ah, the "skill issue" argument again. Same crap aswhen everyonewas worshiping Musk 5-6 years ago, this time it's dario and altman with a claude/chatgpt mask. Crash can't come soon enough.
skybrian11 days ago
And why wouldn’t writing software be a skill issue? Yes, it’s an annoying meme, but we should expect that there are better and worse ways to write software. It would be weird if everyone got the same results regardless of experience.

I’m doubtful that the author’s recommendation always work, but I do some similar things and they do seem to help.

tyleo11 days ago
Not only that, but you really want it to be a skill. The book _Making Software_ describes skills as things you can get better at through practice, and talents as things you're born with.

I'd like to think the time and practice I've put into software engineering has made me better at it. If that's not true, then there's no reason to prefer senior or principal engineers with years of experience over newcomers.

malfist11 days ago
Anyone who thinks they can produce high quality code from an LLM is mistaken about how to judge code. Trust me, I've seen enough PRs to last a life time. A lot of professionals wouldn't know good code if it slapped them in the face.
pydry11 days ago
>And why wouldn’t writing software be a skill issue

You've missed the point. Nobody doubts writing code well or badly is indeed a skill issue.

The question is that "once you account for all of the things you need to do to make the code very high quality, did vibe coding actually provide any real value?"

I'm certain there are guardrails that help bolster vibe coding but I'm equally certain that when ive prompted something important I usually have to redo it enough times that just writing it manually myself usually would have been quicker.

Then I watch other people who code who dump on that opinion and I see total slop. They just can't tell the difference.

Sharlin11 days ago
In a way it reminds me of the good old "if agile doesn't work for you, you're not doing agile right".
osigurdson10 days ago
Agree. These days, if you think you have a methodology that works better than others, you can actually try it / compare it and publish it so that others can replicate and critique your work. Articles like this one, that merely claim they've found the secret sauce, therefore should not carry much weight.

That wasn't the case with 00s agile / Uncle Bob stuff since proving that any of it was helpful was impossible - you just had to believe (and if you didn't believe there was something wrong with you!).

hypfer11 days ago
It actually is though?

Though arguably more of a process and judgement issue than skill.

What makes LLM-generated code a bit special there is that misjudging how to deal with it seems to be what most people do. So the default is broken.

Whereas in prior iterations of "skill issue", the default was working.

post-it11 days ago
What's a crash going to do? The internet didn't disappear after the dot com bubble popped.
ModernMech11 days ago
I think the point is just it doesn’t have to get worse, so there are things you can do to prevent / change it if it is deteriorating.
Sharlin11 days ago
Yes, but it doesn't matter if nobody actually does that. Either because

1. they don't care

2. the rest of the team doesn't care

3. the powers that be actively discourage it because velocity.

fishfasell11 days ago
I think there's a lot of setup and context required for an AI agent to consistently write good code. Once the agent has these guard rails in place I usually get great quality- far better than what I would write in most cases.

I think where things get dicey is being able to write in any language. I write and review code in many languages and frameworks I'm not fluent in, so it's hard for me to distinguish between working code and great code. I can spot when the fundamental logic is wrong, but when it comes to "best fit" choices I'm clueless.

this_user11 days ago
The issue is that in order to have the agent write good code, you need to implement standard SWE best practices. But that also means a lot of manual intervention in terms of writing specs, checking acceptance criteria, and reviewing code. So you end up spending a lot of time on managing your agent, which means you won't get a 1000% productivity gain, you get maybe 50 or 100, possible less in some areas and with some issues.
lolakutty11 days ago
> implement standard SWE best practices

The thing is, if you follow SWE best practices indiscriminately, then you ll have a shit code base in no time.

There is no silver bullet, and no replacement for experience and mindfulness.

user4392811 days ago
A 1000% productivity gain is quite possible on solo greenfield projects.

At work, with a team and code reviews, the 50%-100% figure seems much more likely.

This can probably move towards the more spectacular productivity gains as the AI's output becomes more reliable, people realize this, and less time is spend on code review and cleaning up the output.

beezlewax11 days ago
50 or 100 seems unlikely. Even with all these improvements, custom setups and guardrails it just isn't that much faster for me.
bigstrat200310 days ago
You get 0% productivity gains if you are careful and actually reviewing the code the LLM produces. The only way to actually get the massive productivity gains that AI bros claim is to throw quality out the window.
kuczmama11 days ago
I'm curious as to what guardrails you've tried.

This is something I have been trying to get right as well. I've attempted to use lots of linting and things like strong typing, duplicate checks, cyclomatic complexity, and robust tests. However, I still happen to find issues, which requires me to look at the code (at least at a high level)

For example, I can say "Don't repeat yourself, and don't re-write helper functions" and I will even have a duplicate linter check, but inevitably the LLM will always want to re-write a similar yet slightly different helper function. Like it will always want to re-write something small like a trim() or a toString() function in every file.

bucket201511 days ago
I find that if I leave an instruction in AGENTS.md to "do not do X", there's a good chance the agent will forget it.

But if I add a separate post-implementation pass to "find and fix X" by the agent, it'll usually find and fix the issues.

So I've started doing it for everything from naming conventions to duplicate code to other problems. It does cost more tokens, but now I get less frustrated at having to fix basic issues in the PRs.

esprehn11 days ago
Have you tried something like "Always consult the utils/ package before writing helper functions. When adding a new generic helper function justify it in your design or PR description."

I have better luck telling it positive things rather than lots of "never do X" style things.

ytoawwhra9210 days ago
> Don't repeat yourself, and don't re-write helper functions

It's worth reflecting on why these things are important to you and whether they remain important in an agent-developed codebase.

nicce11 days ago
I would say that it is like gardening. If you let them go havoc from the start, the weed will take over. If you keep focusing on removing the weed and enforce specific standards and practices over the code base and it keeps growing, over time LLMs start to suddenly follow that and they don't make so much slop anymore. At least that is my experience. But I force specific audit agent after every added feature which says them to force compliance with AGENTS.md and check the consistency with the code base.
teliskr11 days ago
I am getting really good results from claude. We have a 22-year old legacy system. The system is stable, but had issues as all legacy systems do. Claude has been great for modernizing the codebase, updating dependencies, auditing security, and rapidly adding new features. It has worked well with existing code style and patterns. Sometimes it is a little off-track, but overall it is pretty amazing.

When implementing new features or making large refactoring changes; I use the superpowers:brainstorming skill. That has consistent process which has worked really well. I alway review the code before merging, but most of the time there are few issues to correct.

I don't do 95% coverage, but I have increased it from 65% to about +80% and that is sufficient.

deterministic10 days ago
My experience as well, working on very large-scale C++ code.

Thanks for adding such a thoughtful and level-headed comment to the discussion.

lolakutty11 days ago
>most of the time there are few issues to correct.

Kindly share the metrics by which you evaluate the changes.

teliskr11 days ago
Lately, I've seen a couple responses to my comments which inquire about metrics. They seem strange and I wonder if they are bots. I just noticed this inquiry is from an account that is 11 days old. How does one benefit by adding bot comments in a forum like this?
teliskr11 days ago
I don't require metrics in these instances. I review the code and test the functionality. That is sufficient for my needs.

Read the full thread on Hacker News →

Related stories