In a recent episode of President Curtis , the President struggles with opening a door on two separate occasions. These doors don't work beca...

277 points•pxx•3 days ago•123 comments•

123 comments

pmarreck3 days ago
I am big on reproducibility (nix aficionado) and determinism (flagging test failures are a red-alert, all-hands-on-deck situation in my world) and correctness.

I am also big on testing (the correct things). And nine-nines (big on Elixir).

And... I'm also big on agent-assisted dev. Which requires pretty much every check in the book to stay productive in. And that's fine to me. I've seen bugs that I wouldn't have made myself. And I've also seen my own bugs fixed. They've all gotten fixed in short order. I don't see why this is a problem.

Raise your personal standards.

Thing is, the unreliable-software situation was already untenable before agents (in poor hands) made it worse.

lokar3 days ago
I don’t think the author (or many people) doubt that one can (and some will) find a way that does not “suck”

But it’s pretty clear that most people are not. For whatever reasons (mgmt pressure, trying to get ahead, skill issues, etc) they half ass it, accept the 10% (silent) fail rate and blame the bad outcomes on the AI as if that absolves them. Or, adopt the attitude that 10% fail is fine, and people who say otherwise are being picky, or are anti-ai luddites or whatever. You should accept that things will suck.

gchamonlive3 days ago
If you took a bad but functional AI generated service and transported it back to 2018 it would have been at worst just mediocre. People do seem forget how dreadful devslop was in the past. I'd take an AI generated mess to disentangle every time over a spaghetti codebase that grew organically in the hands of careless managers.
add-sub-mul-div3 days ago
Yeah, a good reason to be touchy about AI is that it tips the balance of power to lazy people who don't want to work or think. On the scale of our whole society. Imagining the ideal responsible use by most others is folly. No matter how responsible and conscientious you, the reader, are with your use of it.
a34729t2 days ago
I see lots of people having totally checked out, since management cares less about nines and more about tokens consumed. Plus everybody figures theyll be laid off if not tomorrow next quarter or two, and their career will be unviable within 5 years.

I don't condone this view, but i understand where they are coming from!

majormajor3 days ago
> Thing is, the unreliable-software situation was already untenable before agents (in poor hands) made it worse.

I have definitely seen more bafflingly-poor OSS software that just plain doesn't work frequently now than before.

But it's mostly software that wouldn't have existed before because it's trying to do super-niche things. So on the "hobby" side of things, whatever.

But from a "trying to develop software as a business that you want to be a going concern," quality from people who should know better is less tenable than it used to be.

joe_the_user3 days ago
Thing is, the unreliable-software situation was already untenable before agents (in poor hands) made it worse.

Yes but that's the big thing, now isn't it? These are nice tools, used wisely. But their unwise use, oh boy...

The problem is one needs to be in a situation where the incentive is towards quality rather than speed. But that situation rather rare now - thirty years ago, Microsoft won the office wars with crap that had features. And nothing has fundamentally changed in web development since the LPad crisis.

The problem is those companies whose incentive is to allow bugs where it's the involuntary users who suffer will bite you no matter what quality you make your own software.

sscaryterry3 days ago
This, 100% I've seen far far worse produced by humans.
Tanjreeve3 days ago
If the anti-AI people have skill issues because they're holding it wrong then when the output is crap it's the fault of the person with no skill issues who is using It correctly. It can't be both ways.
adamddev13 days ago
Excellent post. People always defend agentic/LLM-driven development by saying, "Well it's good enough", or "It works most of the time."

That may be tolerable for some user-facing app. But what if we start normalizing failures in the libraries, the infrastructure, and the compilers? Everything descends into a mess of unreliability, and that slows EVERYTHING and EVERYONE down.

Root_Denied3 days ago
Banking/Finance is the one industry I've seen push back against this type of thinking. Transactions must be handled in a perfect and repeatable way, or the system is unusable as far as the company is concerned.

There's definitely still AI/LLM integration happening, but is kept out of specific areas of the business.

ryandrake3 days ago
Same with aviation and safety. I honestly believe all programmers should, early in their careers, do a brief "tour of duty" in an industry where the stakes are high and "good enough" isn't good enough. You might not choose to make it your entire career, but at least you're exposed to the discipline, however briefly. Most software developers today have never in their lives worked on a project where defects were taken seriously and where there was process and documentation designed to reduce their occurrence.
geraneum3 days ago
I think the failures matter in non-sensitive environments as well albeit with a different threshold.

If you randomly screw up customer orders (think of DoorDash or an online shop or Airbnb). They lose trust in you and you lose your business to the competition. Going happy go lucky and being irresponsible in the business can bankrupt most* businesses.

* well, of course except the criminal empires which are bailed out by our tax money.

torben-friis3 days ago
I've been seeing the very opposite. Fintech companies treating design of financial systems with the casualness of a frontend aesthetic change. And tons of business people integrating their vibecoded POCs with financially sensitive data sources.
paulddraper3 days ago
I’ve had banking transactions fail a number of times for unknowable or ill-defined reasons. One just last week in fact.

So that is a strange choice for repeatable, understandable operations. Might as well use Jev.

hammock3 days ago
This is the difference between engineering and knowledge work.

Manufacturing lines have tight tolerances. Science has 95% confidence intervals (or greater). HFT has fractional pennies to steamroll up. But “business” (broadly), leadership, macro decisions 3+ steps removed from the coal face can safely operate at wider tolerances.

I cringe whenever I see “xx.xx% growth” on a report as if the value in that hundredth of a percent place is going to sway anyone’s opinion one way or the other. It’s superfluous, wasteful and I would argue, harmful.

The U.S. Marines teach the “70% solution” which says that making a decision that is 70% correct now is better than making a 100% correct decision later.

The speed of your OODA loops is critically important, and cannot be overlooked or expensed in favor of determinism, predictability etc for its own sake. (After all “no plan survives first contact”)

gr_norm3 days ago
Exactly. Reliable abstractions are more important than ever. They're the dues the rest of us must pay to support vibe coding.
Terr_3 days ago
A new flavor of tragedy of the commons and free-rider problems.

Everybody wants to be the ones building cheap slop, hoping somebody else makes protections and repairs...

grumbel3 days ago
> People always defend agentic/LLM-driven development by saying, "Well it's good enough", or "It works most of the time."

The main argument for LLM-driven development is much simpler: "It will get better".

The current state of LLM coding is about a year old. Imagine if we dismissed human coding efforts after a year. Rust, Python2 -> Python3 transition, Python type checking, Windows, C++, … nothing of that was done in a year and emerged in perfection in the first year. Everything takes ages to mature into a usable product. LLM coding is still in the "throw mud at the wall and see what sticks" stage, give it some more years and see how it will develop and what approaches actually work at. For the time being, LLMs are just the most useful development tool in the history of development tools, that's a pretty solid start in such a short time.

joaogui13 days ago
No one moved everything to Rust on the first year the language was released either, and you can get analogies for all of your examples. But it does look like the majority of people have jumped into agentic coding
someonebaggy3 days ago
If it's not good now but it will get better, why should we use it now, instead of waiting until it's better?
psadri3 days ago
Write tests first. Have agent iterate until they are satisfied.

The point is that it boils down to writing the tests correctly, regardless of who is implementing the actual code. Hand-written code without test coverage has the same problems as AI generated code.

wartywhoa233 days ago
> Have agent iterate until they are satisfied.

Or have them rig the tests so that they always pass.

theamk3 days ago
> When a button breaks on a website, I have a model about what should have happened. Somewhere a contract got broken. [...] I might not have access to debug just an HTTP status 500, but I expect there to be somebody whose job is to understand why the endpoint is 500ing. The ownership is well-defined albeit opaque³.

> For many users, however, the actual experience is roughly just "stupid thing sucks." Software already feels capricious; more failures just change the rate of frustration.

I am betting author does not use cloud services much. It is not just "users", it's developers as well. Github is returning 5xx? AWS service does not work? Your email did not get delivered? Nothing we (developers) can do, "stupid thing sucks".

ryandrake3 days ago
We see this fatalistic attitude all the time in software. "Bugs are inevitable." No, they aren't! Bugs are a choice. Almost all companies choose bugs because "no bugs" is too expensive. It's sometimes the right choice but we need to acknowledge that it's a choice and not some natural property of software. Imagine if people who built bridges or airplanes thought "bridge collapses and airplane crashes are inevitable, no way to solve it."
parpfish3 days ago
I’ve had several occasions where I would build a feature “properly” with defensive guardrails and handling of edge cases “that will never happen (but are technically possible)” and then you run into the coworker that says “it’s overengineered. YAGNI. Just do the one-line fix”
simoncion3 days ago
> It is not just "users", it's developers as well.

One can be simultaneously a developer and a user. Distributed systems [0] weren't invented five years ago, after all. ;)

"Github owns this part that we rely on for correct operation and we can do fuckall about it when it fails." is a well-defined ownership model.

[0] ...implying the existence of distinct parts that can be independently developed and independently fail...

layer83 days ago
> This leads to a normalization of inexplicability.

It’s also tightly connected to a normalization of lack of accountability.

> This isn't "getting an FTP account, mounting it locally with curlftpfs, and then using SVN or CVS on the mounted filesystem" -- you still have to do the hard part.

This is probably losing the younger portion of the audience by now. ;)

measurablefunc3 days ago
The people accountable are those hosting & writing the algorithms, e.g. Therac-25.
WorldMaker3 days ago
"Confidence scores" have always implied an anthopocentric meaning that doesn't exist. An algorithm doesn't have "confidence" in the way that a person has confidence, but as soon you put something with that name in front of a business person they assume the number is always a meaningful "letter grade curve" or "universal percentage". I still believe so much that the old quote to "there's lies, damned lies, and then statistics" remains a key to understanding so much why ML is leading to dumb outcomes versus hype. People don't understand statistics, so machines that produce nothing but statistics especially confuse people. (I feel this applies to LLMs as well.)

Read the full thread on Hacker News →

Related stories