We cut PR wait time and running costs by rethinking CI as a system, from the infrastructure underneath it to how work gets scheduled and tests get parallelized.
407 comments
Everyone's going so fast that they keep hitting walls. Review, CI, product asking for things, whatever.
Why have we not seen an improvements in products?
While every post and thread feels like a 90's wall street office, the new android and iphone ship with fewer features than usual. No indie guys come up with a linux-sized alternative OS. Switch 2 remains unhacked. Windows takes 3 seconds to show the right click menu.
Is everyone just running full speed in circles or something?
Our QA, formerly a fairly frequent blocker of all our releases, are doing more in-depth reviews and catching issues earlier in our release process. They have become unblocked to the point they are actively chasing down work that starts to slip.
We have cleaned up and tuned both our security alerts and operations logs and improved our tenant isolation in our service in a way that makes customer and formal audits SIGNIFICANTLY easier.
We're setting ourselves up for faster human development of the hard-things. Our development environment and infrastructure are faster, cleaner, more auditable processes, and cheaper overall to operate.
These fixes mostly don't show up in our product change logs, and definitely don't fall into "new features". It would largely be invisible to the outside world, but our costs are going down (though to be fair, not offsetting the spend on AI to date), internal productivity has improved, operational incidents are down, and customer satisfaction is up.
Is that a fair summary?
Everyone’s trying to figure out how to convert this speed to product features at scale, but enterprises are like container ships. Lots of might but slow to turn. The littler companies can actually take advantage of this and produce higher quality products at much faster speed. I think you’re expecting too much in the short term and too little in the long term. AI-native companies are gonna eat everyone’s lunch, once they figure out how to actually do it reliably.
There was a video by a layperson 2 years ago where they wanted to know the mechanics of some elusive pokemon pinball spawns where there was a whole bunch of missinformation about it online. The people behind that project had the correct addresses but they were completely unlabeled and the layman had to essentially figure out and do some of the work themselves to sort it out and figure it out. Nowadays that would never happen. An LLM will just do it for you within the day.
Wind Waker in 4K w/official Bluetooth GameCube controller
(Nintendo went out of their way to invent a new protocol so that it didn’t “just work” like OG switch controllers)
Couldn't you say this about any company? What isn't just a matter of "figuring it out"?
Why can't they use AI to crack this nut? :D
> If you’re not familiar with normal emulator development time...
Emulators have historically been a shitshow until someone figures out missing pieces here and there. By their nature, that tide lifts all ships. It has very little to do with using AI. I'd argue the real reason emulators aren't what they used to be is the hardware is more complex now and the scene is no longer attracting the most talented devs. Video games used to be on the cutting edge and carried a lot more cultural weight. It's pretty underwhelming to get cred for working on lame x86 hardware that plays yet another franchise reboot.
For the question where are the alternative OSes? Here is one that I've seen. There's probably more - https://www.reddit.com/r/ClaudeAI/comments/1wfpydl/i_asked_c...
For that other stuff you mentioned like the right click menu. Those huge corporate projects suffer more from layers of institutional dysfunction and will be very very slow to show any improvement. Their dysfunction can't be solved with just faster coding.
Using AI to build more features is easier than using AI to improve existing projects. People will gradually figure out how to do latter too, it'll just take longer.
Do you have actual productive examples? As in, products with a real userbase that couldn't exist or be scaled pre-AI? Genuinely asking, I might have missed some large hits. The closest I can remember was bun rewrite kerfuffle, which seemed more a marketing action than anything.
And a few months ago I went through a random sample of new apps on the App Store and the vast majority were just wrappers for AI APIs. So it was more of a new gold rush situation than a productivity bump.
Scale in headcount was a tactic to get investment, then you had to find stuff for everyone to work on. Suddenly people have the time to engineer so hard that we get runtime JSON defined CSS rendering engines to produce the same buttons we've had since 1995, instead of just writing a stylesheet and html.
Code output velocity from AI threatens to be useful, except it's also prone to over-engineering and burning tokens on the unnecessary. It's learned from the best after all. My suspicion is there is a lot getting done, but it's just not that impactful to flagship products.
As others have noted, there is a lot of new work going into passion projects that would have never happened otherwise, and that is cool. But I wouldn't hold my breath for SV tech to become super pragmatic and effective.
> Probably for the same reason that SV companies hiring thousands of developers struggled to improve their product much past the original product, that was built by a handful of people.
Brooks: "Adding manpower to a late software project makes it later."
> Scale in headcount was a tactic to get investment, then you had to find stuff for everyone to work on. Suddenly people have the time to engineer so hard that we get runtime JSON defined CSS rendering engines to produce the same buttons we've had since 1995, instead of just writing a stylesheet and html.
Brooks: "All repairs tend to destroy the structure, to increase the entropy and disorder of the system. Less and less effort is spent on fixing original design flaws; more and more is spent on fixing flaws introduced by earlier fixes. As time passes, the system becomes less and less well-ordered."
> Code output velocity from AI threatens to be useful, except it's also prone to over-engineering and burning tokens on the unnecessary. It's learned from the best after all. My suspicion is there is a lot getting done, but it's just not that impactful to flagship products.
Brooks: "C. S. Lewis has stated it more perceptively: 'That is the key to history. Terrific energy is expended—civilizations are built up—excellent institutions devised; but each time something goes wrong. Some fatal flaw always brings the selfish and cruel people to the top, and then it all slides back into misery and ruin. In fact, the machine conks. It seems to start up all right and runs a few yards, and then it breaks down.'"
That Santayana quote is a bit worn out but applies beautifully here. It's worth recovering its context:
Santayana (1954, p 82): "Progress, far from consisting in change, depends on retentiveness. When change is absolute there remains no being to improve and no direction is set for possible improvement: and when experience is not retained, as among savages, infancy is perpetual. Those who cannot remember the past are condemned to repeat it"
they're working on the ad tech business and other business-related systems involved including internal tools.
and all the other platform engineering shit under the hood that makes much of the web scaleable... and lots of it is open source and contributed to by various engineers from these companies.
end of the day, these are businesses. they're not charities or whatever casual shit.
nobody is stopping you from building a competing product that's lean or whatever.
why don't you do something like that? i'm sure you're a genius.
The simplest explanation is that they don't give a flying flamingo about what you or I consider "improvements to products".
This report is an example.
There are several changes that modify CI behaviour, where the article gives no corresponding quality measurement.
They replaced type aware custom lint rules with AST-only static analysis. They don't say anything about what those new rules detect, didn't do old-vs-new rule comparison. They switched the TypeScript check from tsc to tsgo. Again, they are very proud of the performance improvement, but don't seem to care about diagnostic equivalence. The list goes on. They don't even report pass/fail agreement between the old and the new CI. They have 4x more tests, but no idea whether this big test suite works any better than the smaller old one, or even whether it works at all.
The idea is to drive actual testing from this, but in this era, I think it's interesting as a way to use AI to generate tests, and to cross-check those tests with the natural language descriptions, in a bit of a cycle that helps refine the highest level definition of the software.
Once that's nailed down, the implementation is just details.. normal engineering concerns like maintainability etc notwithstanding of course, but you can trust more and more the AI agents to get it right. The design specs being natural enough for humans to deal with but interpretable/specific enough to actually generate tests is pretty interesting for the bottleneck you are talking about, I think.
At some point, with all this velocity, human users become the bottleneck, unable to keep up with and learn all the changes and new features. Luckily, there's a simple solution: just replace the human users with agentic AI users.
Because every developer is now a slop cannon by default, by default users will experience churn and whiplash, and things will break all over the place. As you point out, this is bad. It's also impossible to fix without deploying agents on the QA side. Like it or hate it, agentic testing is inevitable to protect users from the churn and noise caused by the slop cannon. I don't think that replaces test engineers at all - if anything it makes the job more fun. If you've ever had to keep playwright tests in sync with the target manually, and kept the CI environment up to speed with toolchain changes, you know what I mean.
Whether the "slop cannon by default" situation could have been avoided in the first place, is another question... But we're here now and there's no going back. Might as well deal with it as best as we can.
TLDR: it's not all bad :)
If you even review PRs still: when was the last time you didn't just skip over tests? And if you ever looked at tests in an LLM-heavy PR, how many of those tests tested something useful, and not just built-ins and trivial behaviours?
There's at least some awareness in the industry of how LLMs generate a lot of boilerplate in business logic. It feels like we're much less aware of how much of it is in tests.
The opposite. Reviewing PRs is now about just reviewing the tests, as the code is very likely to be fine if tests are relevant and they're passing.
The main enabler of agentic coding is heavy end-to-end/characterisation testing.
We can ask AI to do very large-scale refactors, for example, because if tests pass, it's 99% that everything is fine.
It's a practice, it's not the best practice. It depends on the team/org/codebase/feature/etc. It always costs time and money and effort. Lots of it from multiple people.
You can get much better output by shifting that cost into hiring much better professionals, not better developers, but overall professionals.
The kind of people you can blindly trust that the software they are writing will be good, you don't need to get involved.
Of course there are exceptions. The author may actually want a review. Or the piece of code might be touching something extremely critical to the business but also easy to get hard.
But besides that? PRs are just productivity porn, or "we do engineering right because we follow Twitter" porn.
The best performing teams I had you hired individuals that removed work and responsibilities off your shoulders without you ever having to regret it. Never added it.
I laugh off engineers that "no you have to review, because it spreads information, enhances quality" and yada yada yada, while in the real world way more critical decisions are made by a single individual without requiring somebody reviewing their work.
How do you prevent churn because this type of developer is somewhere around 1 in 20 or 1 in 100 in terms of rarity?
"It depends on the team/org/codebase/feature/etc"
Boss: whoa, that's a lot more testing than we'd like to pay for, can you cut it down?
Yeah, was not surprised to read this. Actions is convenient if you already use GitHub, but it can also be pretty slow. Given reliability is also a major issue with GitHub these days I expect to see more orgs moving to different pipelines
I don't think it's out of malice, but I do feel uncomfortable with the fact that the people who sell me the CI coordination software also sell me the minutes for the boxes that run the CI software. There's _some_ alignment of interests but not as much as I would want!
Because they're running away from on-premise and the associated Ops teams.
It is that grey zone where it is kinda worth it to pay someone to do it, but also might not be worth it the headaches of managing that person and the infra (like what happens if they go on vacation). The improved speed is the thing that tilts the balance.
Not saying that they're shooting for either of those segments, but someone should. Plenty of enterprise orgs that need (or are convinced they need) Jira's featureset.
Linear is essentially a small subset of Jira.
It's liked and works because it's simpler and covers most of the needs of teams without huge need of defining their own processes.
But be in business long enough and you find out that appealing to small teams and small orgs that don't require flexible and custom process definition doesn't make enough money and doesn't grow you forever.
So they will keep eventually growing in features and flexibility till they cover the reason Jira is so successful: it is more powerful and adaptable.
I'm no fan of Jira, but I've tried and seen enough of these tools to understand why it has the success that it has.
This is even more clear when you go beyond software development and need processes for teams like sales, HR, legal, etc. Jira can fit all of them.
Linear? It's a joke.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 2 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 4 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 9 days ago
- Hacker News · 73 points · 9 days ago
- The Verge · 0 points · 11 days ago