Several months ago, I decided that AI contributions were no longer welcome in a FOSS project I am building and maintaining - LibreWeddingPlanner. It’s not that it got a lot of contributions with AI — actually all…
227 comments
Over the past six months I tried using Claude, chatgpt, Grok and Gemini. At best I got reminders of how things worked. People online say they use them to write their code. The code they supplied to me has NEVER worked or was so convoluted that I threw it away and did it myself.
At most, I use these tools as search engines. Even then some references are poor.
I'm starting to think this is becoming a sad, sad world and AI is just the new TV of the programming world.
For me, it's like reading prose with "Not X, not Y, just Z": it's technically correct, but grates like fingernails on a chalkboard.
I have real trouble sometimes, reading what SOTA (Fable, etc) generate - no isolation or partitioning at all.
The worst was the planning an AI does. When I plan something, it'll be split according to data structures "An object to hold this, an intermediary for the obejct to talk to ORM, a serialiser for it that does this", etc.
The "plans" from SOTA are sometimes just hilarious. It'll go "phase one, implement these user-facing features. Phase two, implement those user-facing features", etc.
That's not a plan, it's an aspiration! A roadmap maybe. A plan, in my way of working, is a blueprint of where all the data goes, with algorithms connecting them. With AI, the data is incidental, the algorithms are incidental, only the goal (in the form of tests) remain. It'll work out some spur-of-the-moment idea around data at the time of writing.
So yeah, I do what you do and throw their stuff away. Currently having more success laying down a skeleton manually and then asking them to add a single feature at a time.
That's with SOTA models as of September-26-2026.
1) The size of task the AI has been given to do appears to be too big, which is why it looks like a roadmap/aspiration. You can ask it to implement a single feature or even a single part of a feature. Just keep cutting the size of the tasks until you become comfortable with it.
2) You like plans in a particular way following data structures/data etc. have you actually told the models this. It doesn't magically know. For the record the fact that the models focus on the end behaviour covered with tests is the way to to it imo. The actual implementation is less important and can be refactored as you wish fairly easily with the AI with the tests ensuring the feature still works.
I do agree though that the current SOTA models are very keen to just implement absolutely everything straight away without explaining/exploring properly. You can customise it fairly easily by using the various skills/agent/claude files to remember your preferred workflow, imo the agents adhere to theses better than they used to even just a few months ago.
since i really don't follow the idea of "writing unit tests first", i implement the feature, test as a user, and then use LLMLs to write the basic unit test. then, i will write more tests to make sure i'm covering everything.
feels like an ok-ish compromise because LLMs can do some ok job with defensive code, while i maintain the main implementation and more advanced test scenarios.
It usually takes me two or three iterations to get there though. Discussing design and principles before writing the bulk of the code is a must. And then a pass or two of review to weed out ugliness.
Still saves time compared to writing the code by hand. Especially for tricky things, where type checking and tests can verify correctness.
That's the whole problem with AI imo and I think the slot machine analogy is mostly right. It's just not predictable whatsoever and then you won't even be able to review all of the thousands of lines of code that you generate. You never know what you get and this has some serious safety implications that are not acceptable. Yes, it's fine as chat to just generate some snippets here and there that can be easily reviewed. Agentic coding is horrible imo.
Am I crazy, or hasn't it been this good for a very long time? The ability to get code out of it after correcting it, correcting it, specifying and respecifying, instrumenting and reinstrumenting, reviewing and demanding refactoring, "no not like that", etc. has been there (for me) nearly from the start. They're great when you're working on something you're not an expert at, and fine if you're working on something that you are pretty good at (if you like to have a cheerleader that sometimes trips and falls on her face.)
My problem is that they don't understand some things that are very clear, and after you've corrected them to get them on track, you're exhausted. You put all of those corrections into a file for them so that when they make the same mistake in the next session you won't have to wrack your brain correcting them, then they a) ignore the file, or b) make a bunch of spurious objections because they were all ready to object and the saved response killed all of the content of those objections. They still seem remarkably dumb.
There's no clear path moving forward. Overreliance on LLMs means your knowledge will exponentially decay and you will absolutely crash any future tech interviews becoming unemployable. Not using it for some quick wins feels wasteful. Finding balance between the two extremes in addition to all existing software development woes is really hard.
Just today it made three glaring mistakes in one session:
1. It read a file in the wrong directory, because that file had the same name as the file in the right directory. It apologized when I challenged it, promising me that it would remember to "read import statements" in the future.
2. It miscounted the number of times a function was called in my repo. It said 20, while my built-in IDE search accurately showed 17. Again, it apologized when I corrected it.
3. It referred to a variable by name that does not exist anywhere in my code. It apologized, and said it was referring to a variable used internally by one of the third-party packages installed in my repo.
So many apologies.
It's the little things like this that remind me on a regular basis just how little I can trust artificial "intelligence."
Without them, not only do agents get simple things like function call counts wrong, they tend to return different results. I use this as an example when showing people how the tooling works.
Grep is fine for simple use cases. A step up from that is ast-grep and I need to explore this tool more. But I had the most success building a small pipeline that reads the old code base, parses it file by file using tree sitter, and then loads it into a SQLite database for querying. For example, I have it capture construct definitions and usages and represent those as directed edges and nodes in a single table depending on the node type. The agent is instructed on how to query it and perform interesting queries like build call graphs, or determine dependencies between domains (modularity is not great in this codebase) which is helpful for us to extract around capability lines.
I also calculate fitness statistics, and have some code to capture specific details and knowledge about this very old framework that short circuits agent work in the future. We have some “interesting” magical libraries and functions that block static analyzers from going beyond the call site. This is mitigated, and means agents don’t have to “guess”.
Making all of this available to the different team members at my work has been pretty helpful. It’s faster (fewer tool calls), cheaper (fewer tokens), and accurate.
Otherwise, that’s exactly the tool to help those kind of queries IMO
ive been regularly seeing this exact reaction online since November 2025, sadly :/
The most telling one is counting function invocations wrong, because that's simply not how models work anymore. They use terminal commands and Python scripts for research like that (if not an LSP, if one took the time to set up their tools most effectively).
Combined with their attitude, I have little doubt that the parent has disabled tool calls, is working in some janky Harness like chat/Duo/Juno, or is using a severely reduced or outdated model.
Yes, so I'll disclose it: I was using Opus 5.5, on medium effort, in Claude desktop, which has full access to my entire repo.
Now everyone can officially lambast me for "using the model wrong or using the wrong model," exactly as you say. But I find it quite interesting that one of the commenters here assumed I was using Opus 4.6, because these mistakes sound like that old version! I'm using the version released just four freaking days ago!
I expect some commenters will now say, "Oh, you should have been using Fable, you old boomer." To them I say: "Yeah, well my employer doesn't allow me to use Fable." And, in jest: "Now get off my lawn."
It's basically impossible at this point to take these people seriously anymore.
Now everything is just like above, dorkspeak. Where it reminds me so much more of kids arguing in the playground about "who would win in a fight Darth Vader or Batman"; or console-vs-PC debates. Everything is about the genuine complexity of navigating certain products, of being first and foremost a consumer of something and putting all your energy into comparing various things you are free to choose from.
Its not even like its less techincal, or more mean now, or anything like that. It's just very different and I know its been a while but it feels like it happened overnight.
i read some of these posts and it feels like the experience with boomers i had to help with their computers at my college job.
they did the least and expected the most.
Could you work out that it was wrong, given weeks to go and check its work? If so, your trust is misplaced.
You ARE taking days or weeks to go and check, yes?
The real challenge is forming the new ideas in the first place and most of those new ideas coming either from using the code as a product or time spent maintaining and refactoring large code.
Anyways, if you want to continue on the path towards regaining control and take it to the next level I wrote something similar here: https://blog.sharefile.systems/be-brave-go-low/
Now I can just say "add 2FA" and in 5 minutes, while I test something else, it is done.
It also made iterations a lot faster, you can try something out, see how it feels, if it doesn't work, you can just trash all the code and start again.
I have spent two weeks using opus just to write a plan/design for 2FA and iron it out until review (about 7 of them) doesn't flag it with 20+ problems (with security holes of various sizes), for which I had to guide it through to not turn it into a mess and whac-a-mole.
I believe "writing the code" means literally just "writing the code" not thinking how to implement it.
Removing the opportunity cost doesn't eliminate the other two costs of a feature
As opposed to writing a prompt for 30s, then doing something else for 15m while the AI works on it? :)
> The real challenge is forming the new ideas in the first place and most of those new ideas coming either from using the code as a product [...]
Exactly.
AI is absolutely excellent at rapid prototyping. Prompt -> result -> use the product -> prompt. Repeat until the right design emerges. Once that happens, the code can be cleaned up, or even rewritten from the grounds up, taking into account what was learned.
I'd expect it's in the tens of thousands, given how effortless you find it.
Code is a liability.
https://github.com/prettydiff/aphorio/commit/0c730389c9e4e69...
It always only took a few hours. And yeah, whatever, sure it adds up. But now you've got this dogshit looking flyer outside your restaurant, and it would have taken you like 45 minutes to adept a Canva template.
For AI, the skill loss and all the rest of it, the lack of control, the lack of knowledge of the codebase. It doesn't feel like saving the few hours is ever worth it.
This is using AI for productivity in any domain, in a nutshell. I just wrote a book using Claude as an experiment, and while the thing got done and it was an amazing tool and a great experience, what I’m left with is a book where every line needs rewriting, there are logical inconsistencies throughout, and the style is so bad it should actually just be binned rather than rewritten.
How was this a great experience if what was produced needed such extensive changes that your own assessment is that it should be thrown out? At what point is the necessary rework so much that the thing being reworked didn't really contribute much to the end product at all?
Getting a good writing style out of them requires careful prompting and many corrections, their default writing style(s) are so highly reinforced by training that they will always tend to drift back to them. Maintaining continuity requires you to create a lot of documentation outside of the text itself. It's a much more manual process than working on a codebase where you've set up a lot of automation and tooling that allows them to check their own work.
Edit: it's also worth noting that many LLMs have gotten much worse at writing prose as they have gotten better at writing code.
If you enjoy the craft and the creative process of actually coming up with new thoughts, instead of relying on the probabilistic combinations of thoughts of others, then you can just as well do it yourself and have full control of the process.
I use AI for low value work with dead lines, where the customers don't really care about the result either. For the golden services and customer engagements were I can tell the customer cares deeply, I use little to no AI, and then get a deep sense of fulfillment due to a job well done.
In effect, everything is classified as: "low value work with dead lines, where the customers don't really care about the result either."
Maybe robot arms can solve this last part one day.
If used naively, where AI spits out 1000s of lines of code, that is far from perfect, written in a style that might not be what you are used to, it can take longer to parse, than if your colleague of 2 years wrote it.
Have you noticed it, too?
And just like that, the book got abandoned.
N.b. some of its analysis and laying out of faults in arguments was actually pellucid and brilliant, it can’t be denied. Just it comes with prose that can’t really be used for anything. And even on another occasion when I got it to help me redraft and extend a different book of mine, then it randomly and consistently started stripping out all the stylistic flourishes out of my sentences, to the point where it couldn’t notice that word choices were very deliberate and actually set up little punchlines and logical payoffs paragraphs or chapters hence. And even when I explained and showed it what it was doing, it was like “ahh that’s so clever and brilliant” but just continued to do the same thing.
Playing an instrument vs electronic/computer music.
We do forget skills we don't practice, especially fine motor skills (like playing the guitar or typing code).
There's inherent pleasure in playing a musical instrument - practicing improves fine motor skills and produces satisfaction.
You can play for yourself and that can be a great experience.
Often people create music for other listeners - and now the satisfaction comes not just from your skill, but from how the music impacts your listeners.
They say you can put more of your 'soul' into music made with an instrument, but I'd say there's quite a bit of electronic music with just as much soul.
People who create electronic music don't generate any of those sounds with their fine motor skills, but they do have a plan about how the song progresses and what emotional state it elicits in users.
That's why you have DJs which are more popular than others.
If you stop playing the guitar for a year, then pick it up and try playing something, you will feel very rusty. But give it a week of practice and most of your skill comes back.. and in 1 month you're back to your peak skill.
I guess my point is - If you go full on agentic, you'll loose some of your coding skill, but you can get it back fairly quickly if you go back to manual coding. On the flip side, you get better at using AI if you use it, so your thinking is at a higher level, but you give up understanding the low level details of how exactly the code works.
Either way you're making 'music', albeit a different kind of music.
90% of Hacker News posers^Wprogrammers
Now, manual labor does have its place as a form of art--take high-precision hand-built timepieces for example.
To add some points on he other side of this analogy:
There is not a lot of purely electronic music that has stood the test of time, at least not when it comes to popularity or, more relevant to the metaphor, profitability. There is usually at very least a human voice in the (literal) mix, but more often than not there are also traditional instruments mixed in.
Take this next point as you will as I am being a bit tongue-in-cheek: While making music-making more accessible to more people is totally great, if I could go a week without hearing a variation of the phrase "Check out my dark ambient drone project!" I would feel oddly accomplished.
Most importantly, though:
> But give it a week of practice and most of your skill comes back.. and in 1 month you're back to your peak skill.
This is only true if you had the skill to begin with. For many electronic musicians, by which I mean junior developers, this is not the case. Does it matter? As a 45-year-old traditional musician... er, I mean hand-coder... I think so, but also ¯\_(ツ)_/¯
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 2 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 4 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 9 days ago
- Hacker News · 73 points · 8 days ago
- The Verge · 0 points · 11 days ago