Several months ago, I decided that AI contributions were no longer welcome in a FOSS project I am building and maintaining - LibreWeddingPlanner. It’s not that it got a lot of contributions with AI — actually all…

188 points•saibotk•5 days ago•227 comments•

227 comments

assimpleaspossi5 days ago
I find it strange to see people writing articles like this as if everyone has used AI for decades. I've programmed for decades. I thought I retired three years ago but got an offer I couldn't refuse. Already there were little things I'd forgotten how to use.

Over the past six months I tried using Claude, chatgpt, Grok and Gemini. At best I got reminders of how things worked. People online say they use them to write their code. The code they supplied to me has NEVER worked or was so convoluted that I threw it away and did it myself.

At most, I use these tools as search engines. Even then some references are poor.

I'm starting to think this is becoming a sad, sad world and AI is just the new TV of the programming world.

lelanthran4 days ago
> The code they supplied to me has NEVER worked or was so convoluted that I threw it away and did it myself.

For me, it's like reading prose with "Not X, not Y, just Z": it's technically correct, but grates like fingernails on a chalkboard.

I have real trouble sometimes, reading what SOTA (Fable, etc) generate - no isolation or partitioning at all.

The worst was the planning an AI does. When I plan something, it'll be split according to data structures "An object to hold this, an intermediary for the obejct to talk to ORM, a serialiser for it that does this", etc.

The "plans" from SOTA are sometimes just hilarious. It'll go "phase one, implement these user-facing features. Phase two, implement those user-facing features", etc.

That's not a plan, it's an aspiration! A roadmap maybe. A plan, in my way of working, is a blueprint of where all the data goes, with algorithms connecting them. With AI, the data is incidental, the algorithms are incidental, only the goal (in the form of tests) remain. It'll work out some spur-of-the-moment idea around data at the time of writing.

So yeah, I do what you do and throw their stuff away. Currently having more success laying down a skeleton manually and then asking them to add a single feature at a time.

That's with SOTA models as of September-26-2026.

jpleyden984 days ago
Just reading this there appears to be 2 problems.

1) The size of task the AI has been given to do appears to be too big, which is why it looks like a roadmap/aspiration. You can ask it to implement a single feature or even a single part of a feature. Just keep cutting the size of the tasks until you become comfortable with it.

2) You like plans in a particular way following data structures/data etc. have you actually told the models this. It doesn't magically know. For the record the fact that the models focus on the end behaviour covered with tests is the way to to it imo. The actual implementation is less important and can be refactored as you wish fairly easily with the AI with the tests ensuring the feature still works.

I do agree though that the current SOTA models are very keen to just implement absolutely everything straight away without explaining/exploring properly. You can customise it fairly easily by using the various skills/agent/claude files to remember your preferred workflow, imo the agents adhere to theses better than they used to even just a few months ago.

saurik4 days ago
I think part of the problem -- answering all of why people are somehow OK with this, and even why the AI does this in the first place -- is that the majority of software developers never got to the point of understanding any of this: the lack of understanding how to carefully structure data surrounds a question they don't even know how to pose, much less answer, and so "write some tests then incrementally try to make them work without breaking any of the existing tests" is the only way they know how to develop at all. In the end, that means that, with the current state of the art (which might change, of course... potentially quickly), your strategy of treating the AI as a junior engineer who fundamentally isn't ready to do your senior-level architecture job makes a lot of sense.
seunosewa4 days ago
With Claude and Codex, you can specify exactly how you want the plan and code to be researched and written. That goes into your rules file. (Claude.md etc). Also, make them read the existing code so they can follow the patterns.
fernandotakai4 days ago
i'm using LLMs/GenAI to do one thing: write unit tests.

since i really don't follow the idea of "writing unit tests first", i implement the feature, test as a user, and then use LLMLs to write the basic unit test. then, i will write more tests to make sure i'm covering everything.

feels like an ok-ish compromise because LLMs can do some ok job with defensive code, while i maintain the main implementation and more advanced test scenarios.

black_knight5 days ago
My experience was similar to yours, upon till earlier this year. Now the code which comes out of Claude code is acceptable most of the time.

It usually takes me two or three iterations to get there though. Discussing design and principles before writing the bulk of the code is a must. And then a pass or two of review to weed out ugliness.

Still saves time compared to writing the code by hand. Especially for tricky things, where type checking and tests can verify correctness.

ddb261fcf6a6e4 days ago
>It usually takes me two or three iterations to get there though.

That's the whole problem with AI imo and I think the slot machine analogy is mostly right. It's just not predictable whatsoever and then you won't even be able to review all of the thousands of lines of code that you generate. You never know what you get and this has some serious safety implications that are not acceptable. Yes, it's fine as chat to just generate some snippets here and there that can be easily reviewed. Agentic coding is horrible imo.

pessimizer4 days ago
> It usually takes me two or three iterations to get there though. Discussing design and principles before writing the bulk of the code is a must. And then a pass or two of review to weed out ugliness.

Am I crazy, or hasn't it been this good for a very long time? The ability to get code out of it after correcting it, correcting it, specifying and respecifying, instrumenting and reinstrumenting, reviewing and demanding refactoring, "no not like that", etc. has been there (for me) nearly from the start. They're great when you're working on something you're not an expert at, and fine if you're working on something that you are pretty good at (if you like to have a cheerleader that sometimes trips and falls on her face.)

My problem is that they don't understand some things that are very clear, and after you've corrected them to get them on track, you're exhausted. You put all of those corrections into a file for them so that when they make the same mistake in the next session you won't have to wrack your brain correcting them, then they a) ignore the file, or b) make a bunch of spurious objections because they were all ready to object and the saved response killed all of the content of those objections. They still seem remarkably dumb.

squeegeeninja4 days ago
The phrasing makes it sound like you are using them through a chat interface. Have you tried something like Claude Code? The real value only starts materialising once it has sufficient access to your environment.
crnkofe4 days ago
I find LLMs are a hit&miss. Can be great at rewriting a function from lang A to lang B. Fails completely at a refactor. Great at coming up with a bunch of networking hosts/IPs to test a particular func. Fails when doing simple validation. Great for prototyping an alternative UI but not even remotely production ready. Always confident regardless of whether the result is correct or false. Run it three times with same query get 3 different results etc. Plenty of claims online how people have become a 1000x developer but 0 actual examples of working code in production. Given they're stochastic by design all of this makes sense in a way.

There's no clear path moving forward. Overreliance on LLMs means your knowledge will exponentially decay and you will absolutely crash any future tech interviews becoming unemployable. Not using it for some quick wins feels wasteful. Finding balance between the two extremes in addition to all existing software development woes is really hard.

kamranjon4 days ago
You would probably be surprised by how many jobs require that you use AI - I even had an interview where it was strongly encouraged to use it during the technical phase. I guess what I’m getting at is, for many people this isn’t really a choice.
stack_framer4 days ago
My employer pays for Claude, and my approach is to use it as a better Google search. It's often not better.

Just today it made three glaring mistakes in one session:

1. It read a file in the wrong directory, because that file had the same name as the file in the right directory. It apologized when I challenged it, promising me that it would remember to "read import statements" in the future.

2. It miscounted the number of times a function was called in my repo. It said 20, while my built-in IDE search accurately showed 17. Again, it apologized when I corrected it.

3. It referred to a variable by name that does not exist anywhere in my code. It apologized, and said it was referring to a variable used internally by one of the third-party packages installed in my repo.

So many apologies.

It's the little things like this that remind me on a regular basis just how little I can trust artificial "intelligence."

dkn4 days ago
I have had a ton of success in exposing AST-based tools to agents when working in large, old codebases.

Without them, not only do agents get simple things like function call counts wrong, they tend to return different results. I use this as an example when showing people how the tooling works.

Grep is fine for simple use cases. A step up from that is ast-grep and I need to explore this tool more. But I had the most success building a small pipeline that reads the old code base, parses it file by file using tree sitter, and then loads it into a SQLite database for querying. For example, I have it capture construct definitions and usages and represent those as directed edges and nodes in a single table depending on the node type. The agent is instructed on how to query it and perform interesting queries like build call graphs, or determine dependencies between domains (modularity is not great in this codebase) which is helpful for us to extract around capability lines.

I also calculate fitness statistics, and have some code to capture specific details and knowledge about this very old framework that short circuits agent work in the future. We have some “interesting” magical libraries and functions that block static analyzers from going beyond the call site. This is mitigated, and means agents don’t have to “guess”.

Making all of this available to the different team members at my work has been pretty helpful. It’s faster (fewer tool calls), cheaper (fewer tokens), and accurate.

jiehong4 days ago
I wish LSP servers would fill that gap, but they tend to work based on a cursor position.

Otherwise, that’s exactly the tool to help those kind of queries IMO

pydry4 days ago
I'm amused at how many hacker news accounts saw this and immediately jumped to the conclusion that you either were using the model wrong or using the wrong model.

ive been regularly seeing this exact reaction online since November 2025, sadly :/

dumberquestions4 days ago
It's just a little baffling to see someone describe a level of performance I haven't experienced since 2025, despite frequently using the tech, as being a frequent concern.
Tadpole91814 days ago
I still do upfront planning and then do careful review of all AI content. The behavior the parent described above hasn't happened for me since around Opus 4.6.

The most telling one is counting function invocations wrong, because that's simply not how models work anymore. They use terminal commands and Python scripts for research like that (if not an LSP, if one took the time to set up their tools most effectively).

Combined with their attitude, I have little doubt that the parent has disabled tool calls, is working in some janky Harness like chat/Duo/Juno, or is using a severely reduced or outdated model.

stack_framer4 days ago
> I'm amused at how many hacker news accounts saw this and immediately jumped to the conclusion that you either were using the model wrong or using the wrong model.

Yes, so I'll disclose it: I was using Opus 5.5, on medium effort, in Claude desktop, which has full access to my entire repo.

Now everyone can officially lambast me for "using the model wrong or using the wrong model," exactly as you say. But I find it quite interesting that one of the commenters here assumed I was using Opus 4.6, because these mistakes sound like that old version! I'm using the version released just four freaking days ago!

I expect some commenters will now say, "Oh, you should have been using Fable, you old boomer." To them I say: "Yeah, well my employer doesn't allow me to use Fable." And, in jest: "Now get off my lawn."

gered4 days ago
Exactly this. If you were using the exact model today that 6-12 months ago people here were telling you "absolutely does not make this kind of mistake anymore" they'd tell you the same thing again, just replacing $OLDER_MODEL with $NEWER_MODEL. Why did $OLDER_MODEL not make this mistake 6-12 months ago, but it does now? The answer to that question is obvious, but AI-bros cannot understand that.

It's basically impossible at this point to take these people seriously anymore.

garg4 days ago
Is your employer forcing you to use Haiku to save on costs?
beepbooptheory4 days ago
It's a small tangent but I am constantly taken aback by just how much the discourse here has slid from nerdom to dorkeyness. All the talk here used to be pendantic and technical and overly complex nerd speak, always stuck in the process itself, feeling above being a 'user', etc.

Now everything is just like above, dorkspeak. Where it reminds me so much more of kids arguing in the playground about "who would win in a fight Darth Vader or Batman"; or console-vs-PC debates. Everything is about the genuine complexity of navigating certain products, of being first and foremost a consumer of something and putting all your energy into comparing various things you are free to choose from.

Its not even like its less techincal, or more mean now, or anything like that. It's just very different and I know its been a while but it feels like it happened overnight.

Matl4 days ago
Or maybe not everyone's codebases are another TODO web app.
moomoo114 days ago
lol ik right?

i read some of these posts and it feels like the experience with boomers i had to help with their computers at my college job.

they did the least and expected the most.

samrus4 days ago
I am skeptical of letting AI do everything as well but this does seem like your using a less capable model. Fable doesnt really do this. In my experience it does just "get" what to do given a clearly defined and measurable outcome
iloveoof4 days ago
Are you using the best models? I feel like my experience is completely different. I am an expert in a small part of a huge monolithic codebase that I’ve worked in for years. When customers report issues that would take me days or weeks to debug, AI can figure it out on the first try.
Applejinx4 days ago
…as far as you know, at the first glance.

Could you work out that it was wrong, given weeks to go and check its work? If so, your trust is misplaced.

You ARE taking days or weeks to go and check, yes?

austin-cheney5 days ago
Yeah, I have never understood the over reliance on AI. Writing the code is not the challenge. The time it takes to push a new feature and test it out is often trivial, maybe a few hours.

The real challenge is forming the new ideas in the first place and most of those new ideas coming either from using the code as a product or time spent maintaining and refactoring large code.

Anyways, if you want to continue on the path towards regaining control and take it to the next level I wrote something similar here: https://blog.sharefile.systems/be-brave-go-low/

XCSme5 days ago
Wriring the code is not the challenge, but it's what was taking up most of the time. Not the typing itself, but also because I had to think of how to implement it.

Now I can just say "add 2FA" and in 5 minutes, while I test something else, it is done.

It also made iterations a lot faster, you can try something out, see how it feels, if it doesn't work, you can just trash all the code and start again.

anygivnthursday5 days ago
But you still have to review that 2FA code and that involves thinking through the implementation, right?
aytigra4 days ago
I don't know what it built for you in 5 minutes, probably something that "works".

I have spent two weeks using opus just to write a plan/design for 2FA and iron it out until review (about 7 of them) doesn't flag it with 20+ problems (with security holes of various sizes), for which I had to guide it through to not turn it into a mess and whac-a-mole.

AlienRobot4 days ago
>Not the typing itself, but also because I had to think of how to implement it.

I believe "writing the code" means literally just "writing the code" not thinking how to implement it.

automatic61315 days ago
The limiting factor of development (for money) is good ideas, or valuable ideas. If you have 5 features you want, but only time in a month for 2, then that forces you to choose the best 2, which come with the tech debt, support and opportunity cost.

Removing the opportunity cost doesn't eliminate the other two costs of a feature

Mawr4 days ago
> The time it takes to push a new feature and test it out is often trivial, maybe a few hours.

As opposed to writing a prompt for 30s, then doing something else for 15m while the AI works on it? :)

> The real challenge is forming the new ideas in the first place and most of those new ideas coming either from using the code as a product [...]

Exactly.

AI is absolutely excellent at rapid prototyping. Prompt -> result -> use the product -> prompt. Repeat until the right design emerges. Once that happens, the code can be cleaned up, or even rewritten from the grounds up, taking into account what was learned.

joenot4434 days ago
How many lines of code do you write, on a good day?

I'd expect it's in the tens of thousands, given how effortless you find it.

grey-area4 days ago
Would that be a good thing?

Code is a liability.

austin-cheney4 days ago
According to GitHub it was about +2114 / -1519 lines of code from 45 files yesterday. I felt productive yesterday.

https://github.com/prettydiff/aphorio/commit/0c730389c9e4e69...

podgietaru4 days ago
This actually sums up a lot of what I feel about AI. For everything. For flyers, menus, all the rest of it.

It always only took a few hours. And yeah, whatever, sure it adds up. But now you've got this dogshit looking flyer outside your restaurant, and it would have taken you like 45 minutes to adept a Canva template.

For AI, the skill loss and all the rest of it, the lack of control, the lack of knowledge of the codebase. It doesn't feel like saving the few hours is ever worth it.

gizajob5 days ago
“ There were tasks I could have done in 20 minutes easily, that took 5 minutes of an AI agent, and then 2 days for me to review.”

This is using AI for productivity in any domain, in a nutshell. I just wrote a book using Claude as an experiment, and while the thing got done and it was an amazing tool and a great experience, what I’m left with is a book where every line needs rewriting, there are logical inconsistencies throughout, and the style is so bad it should actually just be binned rather than rewritten.

thwarted4 days ago
> I just wrote a book using Claude as an experiment, and while the thing got done and it was an amazing tool and a great experience, what I’m left with is a book where every line needs rewriting, there are logical inconsistencies throughout, and the style is so bad it should actually just be binned rather than rewritten.

How was this a great experience if what was produced needed such extensive changes that your own assessment is that it should be thrown out? At what point is the necessary rework so much that the thing being reworked didn't really contribute much to the end product at all?

gizajob4 days ago
It was really interesting in having an assistant who was completely on the ball and up to speed, never forgot, could pick up from where we left off days ago, could produce and summarise arguments… but the issue was that the amazing assistant could only really work one chapter at a time even with a plan, and changing one chapter would often break things said in other chapters, and their style of writing was such that the resorted to cliches, bad metaphors, and generally bad or repetitive prose. So the grammar was fine, the style was bad to the point where the whole work would not be something I could have my name on or near. I really could come back to it, just having to retype everything in markdown made me see that the project had legs and the topic is sufficient for a book length treatment, but even with the help of Claude I have a few months work to do to get it done and out there. Just it’s a different few months work to the work I would have had without Claude. I could also have attempted to pay Claude thousands of dollars to see if it worked better, but the book is not likely to recoup what that would cost.
pigpop4 days ago
That's not surprising, I've used LLMs to write several chapters of a book and it's many times more difficult than getting them to write working code. It's not a task that they're well optimized for but the biggest hindrance is that there are practically no tools that allow you to validate prose especially at the lengths required for a book.

Getting a good writing style out of them requires careful prompting and many corrections, their default writing style(s) are so highly reinforced by training that they will always tend to drift back to them. Maintaining continuity requires you to create a lot of documentation outside of the text itself. It's a much more manual process than working on a codebase where you've set up a lot of automation and tooling that allows them to check their own work.

Edit: it's also worth noting that many LLMs have gotten much worse at writing prose as they have gotten better at writing code.

abc123abc1234 days ago
This is the way. Basically, what determines when to use AI or not is what you enjoy about the creative process. Do you only care about the result, use AI, and then spend hours (or days) fixing errors, tweaking, rewriting, using other AI:s to validate and improve the first AI etc. You get the result, and your activity consist of arguing with AI:s.

If you enjoy the craft and the creative process of actually coming up with new thoughts, instead of relying on the probabilistic combinations of thoughts of others, then you can just as well do it yourself and have full control of the process.

I use AI for low value work with dead lines, where the customers don't really care about the result either. For the golden services and customer engagements were I can tell the customer cares deeply, I use little to no AI, and then get a deep sense of fulfillment due to a job well done.

abalashov4 days ago
Sadly, a great deal of applied business programming in modern capitalism doesn't give you the option of enjoying the creative process.

In effect, everything is classified as: "low value work with dead lines, where the customers don't really care about the result either."

AnimalMuppet4 days ago
So you use AI when it's all right to produce slop. That actually seems reasonable.
iloveoof4 days ago
I don’t see how that’s possible. When I am doing a comprehensive code review with rigorous functional testing of another developer’s work, it usually takes me around half as long to review as the developer took to write it. That includes the back and forth of MR issues and fixes. If AI writes it why would it take 100x longer to review than write it? At worst case it’s just a draft of something I can develop myself, so it shouldn’t take longer than 20 minutes. At best case it is a code review exercise so it takes me half as long.
gizajob4 days ago
AI can produce for you an amazing jpeg of an oil painting you’ve ideated together. You could print it out on a large format printer but you’re still only holding a printout of an AI oil painting. To actually make a piece of art that’s worth being on a wall and standing as an artwork, then you still have to paint the painting using oil paint.

Maybe robot arms can solve this last part one day.

abc123abc1234 days ago
Depends on how the AI is used. If used intelligently, for short, readable code snippets, reviewing and understand is easy.

If used naively, where AI spits out 1000s of lines of code, that is far from perfect, written in a style that might not be what you are used to, it can take longer to parse, than if your colleague of 2 years wrote it.

manbash5 days ago
I noticed lately that recent LLMs write very short sentences, sprinkling so many periods over one paragraph. I'm pretty sure that this is some new regression that we're having with new models.

Have you noticed it, too?

gizajob5 days ago
Hmmm I was using Claude Opus for the most part. I didn’t notice any particular shortness in the length of the sentences, but that’s also an issue in itself – the whole thing is just average fine-ness: the sentences are fine, the paragraphs are fine, everything is generic and inoffensive and fine. But at the same time it was giving huge amounts of pushback when I was getting it to interrogate logical faults of certain domains, to the point where I was having to argue with and even convince an AI using data and even its own analysis that there was cause to talk about certain things from a certain angle (can appreciate I’m being vague about which things). It can’t write complex, deeply-claused sentences (or won’t). And it also wants to write snappy little concluding sentences that make a fairly well-argued piece of sociological analysis read like Sex and the City or something.

And just like that, the book got abandoned.

N.b. some of its analysis and laying out of faults in arguments was actually pellucid and brilliant, it can’t be denied. Just it comes with prose that can’t really be used for anything. And even on another occasion when I got it to help me redraft and extend a different book of mine, then it randomly and consistently started stripping out all the stylistic flourishes out of my sentences, to the point where it couldn’t notice that word choices were very deliberate and actually set up little punchlines and logical payoffs paragraphs or chapters hence. And even when I explained and showed it what it was doing, it was like “ahh that’s so clever and brilliant” but just continued to do the same thing.

CuriouslyC5 days ago
That's a side effect of people bashing the em-dash and the labs rushing to correct. The period is the most common replacement (along with some minor tweaks to sentence structure), so now that models are being RL'd away from the em-dash the models are overusing a new construction.
delegate5 days ago
I sometimes look at it via the metaphor of music.

Playing an instrument vs electronic/computer music.

We do forget skills we don't practice, especially fine motor skills (like playing the guitar or typing code).

There's inherent pleasure in playing a musical instrument - practicing improves fine motor skills and produces satisfaction.

You can play for yourself and that can be a great experience.

Often people create music for other listeners - and now the satisfaction comes not just from your skill, but from how the music impacts your listeners.

They say you can put more of your 'soul' into music made with an instrument, but I'd say there's quite a bit of electronic music with just as much soul.

People who create electronic music don't generate any of those sounds with their fine motor skills, but they do have a plan about how the song progresses and what emotional state it elicits in users.

That's why you have DJs which are more popular than others.

If you stop playing the guitar for a year, then pick it up and try playing something, you will feel very rusty. But give it a week of practice and most of your skill comes back.. and in 1 month you're back to your peak skill.

I guess my point is - If you go full on agentic, you'll loose some of your coding skill, but you can get it back fairly quickly if you go back to manual coding. On the flip side, you get better at using AI if you use it, so your thinking is at a higher level, but you give up understanding the low level details of how exactly the code works.

Either way you're making 'music', albeit a different kind of music.

Z_I_F_F4 days ago
A more apt comparison would be playing guitar vs. prompting an AI to generate a guitar solo.
vconnor4 days ago
“I am a guitarist and I am glad I don’t have to do the boring part of playing the guitar any more”

90% of Hacker News posers^Wprogrammers

ajconway4 days ago
Engineering is not music though. If you're building a skyscraper, and your tower crane breaks, you could go old-school like they did with the pyramids in ancient Egypt. But why would you?

Now, manual labor does have its place as a form of art--take high-precision hand-built timepieces for example.

user4326784 days ago
Yeah, only if you had coding skills in the first place.
sodapopcan4 days ago
Ha, I could talk about this framing for days as I think about it a lot.

To add some points on he other side of this analogy:

There is not a lot of purely electronic music that has stood the test of time, at least not when it comes to popularity or, more relevant to the metaphor, profitability. There is usually at very least a human voice in the (literal) mix, but more often than not there are also traditional instruments mixed in.

Take this next point as you will as I am being a bit tongue-in-cheek: While making music-making more accessible to more people is totally great, if I could go a week without hearing a variation of the phrase "Check out my dark ambient drone project!" I would feel oddly accomplished.

Most importantly, though:

> But give it a week of practice and most of your skill comes back.. and in 1 month you're back to your peak skill.

This is only true if you had the skill to begin with. For many electronic musicians, by which I mean junior developers, this is not the case. Does it matter? As a 45-year-old traditional musician... er, I mean hand-coder... I think so, but also ¯\_(ツ)_/¯

Read the full thread on Hacker News →

Related stories