We cut PR wait time and running costs by rethinking CI as a system, from the infrastructure underneath it to how work gets scheduled and tests get parallelized.

316 points•julian_digital•9 days ago•407 comments•

407 comments

torben-friis9 days ago
Here's my constant question:

Everyone's going so fast that they keep hitting walls. Review, CI, product asking for things, whatever.

Why have we not seen an improvements in products?

While every post and thread feels like a 90's wall street office, the new android and iphone ship with fewer features than usual. No indie guys come up with a linux-sized alternative OS. Switch 2 remains unhacked. Windows takes 3 seconds to show the right click menu.

Is everyone just running full speed in circles or something?

TrueDuality9 days ago
A good chunk of what my company has been doing with AI falls into either burning down our known tech-debt and "easy wins" that no one ever had the bandwidth to approach... And improving / automating our processes. The former is having a direct and meaningful impact on the quality and availability of our services.

Our QA, formerly a fairly frequent blocker of all our releases, are doing more in-depth reviews and catching issues earlier in our release process. They have become unblocked to the point they are actively chasing down work that starts to slip.

We have cleaned up and tuned both our security alerts and operations logs and improved our tenant isolation in our service in a way that makes customer and formal audits SIGNIFICANTLY easier.

We're setting ourselves up for faster human development of the hard-things. Our development environment and infrastructure are faster, cleaner, more auditable processes, and cheaper overall to operate.

These fixes mostly don't show up in our product change logs, and definitely don't fall into "new features". It would largely be invisible to the outside world, but our costs are going down (though to be fair, not offsetting the spend on AI to date), internal productivity has improved, operational incidents are down, and customer satisfaction is up.

jurgenburgen9 days ago
Cost is stable (or rising?), user-facing delivery is sameish, (some) engineers seem happier because they are allowed to gold-plate and prepare for future “human development of the hard-things”.

Is that a fair summary?

lazyasciiart9 days ago
How is dev satisfaction?
verpeteren9 days ago
This is the way!
lukevp9 days ago
PS5 emulation has gone from barely working to running Dark Souls at 10+ FPS with virtually no graphical glitches… in 6 weeks. If you’re not familiar with normal emulator development time, this is… quite extraordinary. There have been insane progress on decompilations and many other things in the emulator space.

Everyone’s trying to figure out how to convert this speed to product features at scale, but enterprises are like container ships. Lots of might but slow to turn. The littler companies can actually take advantage of this and produce higher quality products at much faster speed. I think you’re expecting too much in the short term and too little in the long term. AI-native companies are gonna eat everyone’s lunch, once they figure out how to actually do it reliably.

decompgoldenage9 days ago
Also, in 12 months, we went from seeing game dissassembly and decompilations projects be slow and annoying, to suddenly having enough mapower where people are taking the liberty to CHOOSE which decompilations to support. DK64's release was 80% anti-AI marketing, not because they were strictly anti-AI, but because they were telling the OTHER decompilation to fuck off and learn some goddamned standards. Completely unthinkable before.

There was a video by a layperson 2 years ago where they wanted to know the mechanics of some elusive pokemon pinball spawns where there was a whole bunch of missinformation about it online. The people behind that project had the correct addresses but they were completely unlabeled and the layman had to essentially figure out and do some of the work themselves to sort it out and figure it out. Nowadays that would never happen. An LLM will just do it for you within the day.

thejazzman9 days ago
To further this point, I’ve picked up the torch and got Switch2 controllers working in Dolphin and Cemu

Wind Waker in 4K w/official Bluetooth GameCube controller

(Nintendo went out of their way to invent a new protocol so that it didn’t “just work” like OG switch controllers)

sublinear9 days ago
> AI-native companies are gonna eat everyone’s lunch, once they figure out how to actually do it reliably.

Couldn't you say this about any company? What isn't just a matter of "figuring it out"?

Why can't they use AI to crack this nut? :D

> If you’re not familiar with normal emulator development time...

Emulators have historically been a shitshow until someone figures out missing pieces here and there. By their nature, that tide lifts all ships. It has very little to do with using AI. I'd argue the real reason emulators aren't what they used to be is the hardware is more complex now and the scene is no longer attracting the most talented devs. Video games used to be on the cutting edge and carried a lot more cultural weight. It's pretty underwhelming to get cred for working on lame x86 hardware that plays yet another franchise reboot.

jetbalsa9 days ago
I did the same thing with a old multiplayer game called Wulfram, we got it online and a mostly working server (most of the game logic was server side with no surviving code or binary) with the help of AI, Mostly opus and astra. https://wulfram3.com is our efforts in this.
apf69 days ago
Last year people were asking "if AI is so great then where are the new apps?". Then data for 2026 came out and now the IOS app store has a 84% percent year-over-year increase in new app submissions.

For the question where are the alternative OSes? Here is one that I've seen. There's probably more - https://www.reddit.com/r/ClaudeAI/comments/1wfpydl/i_asked_c...

For that other stuff you mentioned like the right click menu. Those huge corporate projects suffer more from layers of institutional dysfunction and will be very very slow to show any improvement. Their dysfunction can't be solved with just faster coding.

Using AI to build more features is easier than using AI to improve existing projects. People will gradually figure out how to do latter too, it'll just take longer.

torben-friis9 days ago
That's just volume though. I also know, and have no trouble believing, that github is going down partly due to the weight of all the vibecoded pushes. But that does not necessarily translate to people's needs and wants being covered, for all we know iOS just has a plague of unused POCs.

Do you have actual productive examples? As in, products with a real userbase that couldn't exist or be scaled pre-AI? Genuinely asking, I might have missed some large hits. The closest I can remember was bun rewrite kerfuffle, which seemed more a marketing action than anything.

onetimeusename9 days ago
I have this theory that the economics of AI development would make it so it's much more viable to make in-house, custom, apps rather than paying subscriptions for apps. I have even seen this play out on a personal level. A friend of mine who has no dev experience vibe-coded something custom that evite does. So it's possible measuring the number of apps miscounts ones that are not public or shared. I'm not negating anything you said, just adding something about metrics on apps.
matt_heimer9 days ago
The OS one is funny to me, it's been hard to keep osdev.org online due to the insane amount AI bot traffic.
globular-toast9 days ago
We know it's easy to produce working code now, what's missing is useful software that actually solves a problem. What problem does that toy OS solve? The main point of toy OSes has been education, but you learn nothing when using an LLM. So all of this is just pointless energy consumption, it's not solving any real problems that people have.
sarchertech9 days ago
Last time I looked the other app stores didn’t show anywhere near that much of a bump (that was a several months ago though).

And a few months ago I went through a random sample of new apps on the App Store and the vast majority were just wrappers for AI APIs. So it was more of a new gold rush situation than a productivity bump.

ehnto9 days ago
Probably for the same reason that SV companies hiring thousands of developers struggled to improve their product much past the original product, that was built by a handful of people.

Scale in headcount was a tactic to get investment, then you had to find stuff for everyone to work on. Suddenly people have the time to engineer so hard that we get runtime JSON defined CSS rendering engines to produce the same buttons we've had since 1995, instead of just writing a stylesheet and html.

Code output velocity from AI threatens to be useful, except it's also prone to over-engineering and burning tokens on the unnecessary. It's learned from the best after all. My suspicion is there is a lot getting done, but it's just not that impactful to flagship products.

As others have noted, there is a lot of new work going into passion projects that would have never happened otherwise, and that is cool. But I wouldn't hold my breath for SV tech to become super pragmatic and effective.

huurtehoog9 days ago
As usual, Brooks has 50 year old insights about this 'novel problem'

> Probably for the same reason that SV companies hiring thousands of developers struggled to improve their product much past the original product, that was built by a handful of people.

Brooks: "Adding manpower to a late software project makes it later."

> Scale in headcount was a tactic to get investment, then you had to find stuff for everyone to work on. Suddenly people have the time to engineer so hard that we get runtime JSON defined CSS rendering engines to produce the same buttons we've had since 1995, instead of just writing a stylesheet and html.

Brooks: "All repairs tend to destroy the structure, to increase the entropy and disorder of the system. Less and less effort is spent on fixing original design flaws; more and more is spent on fixing flaws introduced by earlier fixes. As time passes, the system becomes less and less well-ordered."

> Code output velocity from AI threatens to be useful, except it's also prone to over-engineering and burning tokens on the unnecessary. It's learned from the best after all. My suspicion is there is a lot getting done, but it's just not that impactful to flagship products.

Brooks: "C. S. Lewis has stated it more perceptively: 'That is the key to history. Terrific energy is expended—civilizations are built up—excellent institutions devised; but each time something goes wrong. Some fatal flaw always brings the selfish and cruel people to the top, and then it all slides back into misery and ruin. In fact, the machine conks. It seems to start up all right and runs a few yards, and then it breaks down.'"

That Santayana quote is a bit worn out but applies beautifully here. It's worth recovering its context:

Santayana (1954, p 82): "Progress, far from consisting in change, depends on retentiveness. When change is absolute there remains no being to improve and no direction is set for possible improvement: and when experience is not retained, as among savages, infancy is perpetual. Those who cannot remember the past are condemned to repeat it"

dmboyd9 days ago
With pay-per-token, there’s also an incentive for over-engineered but functionally harmless architectures. A json based css engine is probably something that can be test-cased really well in the training set.
moomoo119 days ago
those thousands of developers aren't working on the FB feed or whatever...

they're working on the ad tech business and other business-related systems involved including internal tools.

and all the other platform engineering shit under the hood that makes much of the web scaleable... and lots of it is open source and contributed to by various engineers from these companies.

end of the day, these are businesses. they're not charities or whatever casual shit.

nobody is stopping you from building a competing product that's lean or whatever.

why don't you do something like that? i'm sure you're a genius.

saithound9 days ago
> Why have we not seen an improvements in products? [..] Is everyone just running full speed in circles or something?

The simplest explanation is that they don't give a flying flamingo about what you or I consider "improvements to products".

This report is an example.

There are several changes that modify CI behaviour, where the article gives no corresponding quality measurement.

They replaced type aware custom lint rules with AST-only static analysis. They don't say anything about what those new rules detect, didn't do old-vs-new rule comparison. They switched the TypeScript check from tsc to tsgo. Again, they are very proud of the performance improvement, but don't seem to care about diagnostic equivalence. The list goes on. They don't even report pass/fail agreement between the old and the new CI. They have 4x more tests, but no idea whether this big test suite works any better than the smaller old one, or even whether it works at all.

plant-ian9 days ago
Are they really adding 2,000 tests a week to their codebase?
aliclark9 days ago
In my case it's not the CI that's the bottleneck. It's the human testing side. Does it work, sure. But does it actually do the thing we want (and more importantly) does it do it in a way our customers will understand and actually like?
lbrito9 days ago
I feel no one talks about this, and yet it is glaringly obvious. Are we even solving the right problem? No one cares. Push code. Number go up.
dbbk8 days ago
Sounds like you want a product manager
Aeolun9 days ago
That problem isn’t new though. It’s not the biz people that are enthusiastic over AI, it’s the devs.
skort9 days ago
I feel like if anything LLMs have reduced coding/software to "throw as much as possible at the wall and see what sticks". It feels quite shortsighted and wasteful, especially considering that humanity needs to get better about how it produces and consumes energy. It's sort of bleak.
eggplantemoji699 days ago
Which is a bit terrifying and removes the ‘engineering’ component of a critical part of the job, in exchange for slot machine / pray to ai gods…which comes with disconnect from the code, something I assume will adversely compound over time.
pydry9 days ago
hasn't improved software either. it's made the overall quality worse if anything.
shykes9 days ago
Think of it as a layered problem. If the bottom layer (CI) cannot keep up with the output of agents, then solving problems at a higher layer - like user experience checks - will be exponentially slower and less reliable. Kind of like how optimizing tight inner loops makes your whole program faster.
user439289 days ago
It's true because it usually makes no sense to QA test a build where the test suite has not yet passed.
Sharlin9 days ago
Uh, no. If stage X is the bottleneck, it's immaterial how much you speed up stage Y.
radarsat19 days ago
I was just discussing with my team that this exact thing maybe brings back to relevance the idea of "behaviour driven development" and I was reminded of this Cucumber/Gherkin lib & language that defines a kind of executable prose you can use to specify how the software should behave. It's an interpretable programming language but designed to be close to how a human might just write down their specs of what kind of actions and responses are expected from a software system.

The idea is to drive actual testing from this, but in this era, I think it's interesting as a way to use AI to generate tests, and to cross-check those tests with the natural language descriptions, in a bit of a cycle that helps refine the highest level definition of the software.

Once that's nailed down, the implementation is just details.. normal engineering concerns like maintainability etc notwithstanding of course, but you can trust more and more the AI agents to get it right. The design specs being natural enough for humans to deal with but interpretable/specific enough to actually generate tests is pretty interesting for the bottleneck you are talking about, I think.

Sharlin9 days ago
Clearly you have to replace obsolete human testers with agentic AI testers, duh.

At some point, with all this velocity, human users become the bottleneck, unable to keep up with and learn all the changes and new features. Luckily, there's a simple solution: just replace the human users with agentic AI users.

dwaltrip9 days ago
Maybe we can also replace the customers with agentic consumers.
MikeNotThePope9 days ago
That’s what I’m doing! Results yet to be measured.
shykes9 days ago
I think you're being sarcastic, but there is actual truth behind what you're saying.

Because every developer is now a slop cannon by default, by default users will experience churn and whiplash, and things will break all over the place. As you point out, this is bad. It's also impossible to fix without deploying agents on the QA side. Like it or hate it, agentic testing is inevitable to protect users from the churn and noise caused by the slop cannon. I don't think that replaces test engineers at all - if anything it makes the job more fun. If you've ever had to keep playwright tests in sync with the target manually, and kept the CI environment up to speed with toolchain changes, you know what I mean.

Whether the "slop cannon by default" situation could have been avoided in the first place, is another question... But we're here now and there's no going back. Might as well deal with it as best as we can.

TLDR: it's not all bad :)

dgroshev9 days ago
I suspect a substantial part of this is an avalanche of useless testing.

If you even review PRs still: when was the last time you didn't just skip over tests? And if you ever looked at tests in an LLM-heavy PR, how many of those tests tested something useful, and not just built-ins and trivial behaviours?

There's at least some awareness in the industry of how LLMs generate a lot of boilerplate in business logic. It feels like we're much less aware of how much of it is in tests.

sz4kerto9 days ago
> If you even review PRs still: when was the last time you didn't just skip over tests?

The opposite. Reviewing PRs is now about just reviewing the tests, as the code is very likely to be fine if tests are relevant and they're passing.

The main enabler of agentic coding is heavy end-to-end/characterisation testing.

We can ask AI to do very large-scale refactors, for example, because if tests pass, it's 99% that everything is fine.

epolanski9 days ago
I was (luckily) not reviewing PRs even before AI.

It's a practice, it's not the best practice. It depends on the team/org/codebase/feature/etc. It always costs time and money and effort. Lots of it from multiple people.

You can get much better output by shifting that cost into hiring much better professionals, not better developers, but overall professionals.

The kind of people you can blindly trust that the software they are writing will be good, you don't need to get involved.

Of course there are exceptions. The author may actually want a review. Or the piece of code might be touching something extremely critical to the business but also easy to get hard.

But besides that? PRs are just productivity porn, or "we do engineering right because we follow Twitter" porn.

The best performing teams I had you hired individuals that removed work and responsibilities off your shoulders without you ever having to regret it. Never added it.

I laugh off engineers that "no you have to review, because it spreads information, enhances quality" and yada yada yada, while in the real world way more critical decisions are made by a single individual without requiring somebody reviewing their work.

datsci_est_20159 days ago
How much above the local market rate are you paying for this level of qualified developer? 20%? 50%?

How do you prevent churn because this type of developer is somewhere around 1 in 20 or 1 in 100 in terms of rarity?

Rapzid8 days ago
A wise person once said:

"It depends on the team/org/codebase/feature/etc"

_zoltan_9 days ago
I actually look at the tests first. But for this to work, you have to have good hygiene. If the test makes sense, fails without the PR but passed with the PR, I'm almost happy.
cryptonector8 days ago
Boss: full branch code coverage testing is the best way, make it so!

Boss: whoa, that's a lot more testing than we'd like to pay for, can you cut it down?

MikeNotThePope9 days ago
Probably half the generated tests in my app are for markdown files.
classictraffic9 days ago
> Moving our workloads off GitHub Actions to third-party runners with faster CPUs, higher-performance storage, and better cache infrastructure gave us faster machines to run the same pipeline on

Yeah, was not surprised to read this. Actions is convenient if you already use GitHub, but it can also be pretty slow. Given reliability is also a major issue with GitHub these days I expect to see more orgs moving to different pipelines

speedgoose9 days ago
I wonder how much faster GitHub actions could be if they weren’t running on Azure. Because Azure is either slow or very expensive.
chris_money2029 days ago
I think its actually because GitHub isn't solely running on Azure. They are currently using a mix of on prem, AWS, and Azure. The Azure migration has been a challenge with Azure running into capacity issues due to AI load.
rtpg9 days ago
to be fair all CI providers I've had the pleasure of working with have gnarly performance profiles for the boxes they provide.

I don't think it's out of malice, but I do feel uncomfortable with the fact that the people who sell me the CI coordination software also sell me the minutes for the boxes that run the CI software. There's _some_ alignment of interests but not as much as I would want!

stackskipton9 days ago
It's not Azure as much as self-hosted runners are basement bin servers on clearly massively oversubscribed machines.
wereHamster9 days ago
I wonder why larger companies don't use self hosted github runners. You can buy a pretty beefy machine (TB of RAM, 256 cores, fast NVMe disk) and tests will run faster than on any hosted platform. Plus you don't have to shard as aggressively because more fits into one machine, benefit of shared page cache, shared persistent disk, can easily preserve working directory (for example my pnpm install takes 0 seconds, because the node_modules folder is already present from the previous run).
oblio9 days ago
> I wonder why larger companies don't use self hosted github runners.

Because they're running away from on-premise and the associated Ops teams.

Lucasoato9 days ago
Microsoft is pushing so hard to get companies far away from the on-premises world. Once they're in the Cloud, there's too much vendor lock-in for big enterprises to go away.
DanielHB9 days ago
My company did this a couple of years ago to some machines we host ourselves. It is a bit of trouble to set up but it is not that bad and can easily save a 1000+ dollars per month since GHA runners are very expensive.

It is that grey zone where it is kinda worth it to pay someone to do it, but also might not be worth it the headaches of managing that person and the infra (like what happens if they go on vacation). The improved speed is the thing that tilts the balance.

dneri9 days ago
We switched our Actions workload to blacksmith.sh (not affiliated) and have been pretty happy with how fast and inexpensive they are. I wouldn't be surprised to see this trend continue.
idkasam7 days ago
Agree, and it's already happening. Mostly smaller orgs so far though, they can just swap the runner label and move on. Larger orgs are starting to look into it but it's slow, lots of red tape (security review, procurement, "we're already paying GitHub").
algesten9 days ago
Good thing Linear has been finished for a long time doesn't need more features, so AI coding can go slow. Oh damn, it's busy becoming the next Jira :(
reticulates9 days ago
I’m conflicted about Linear’s progression. I dislike some of the features but on the whole they’ve managed to keep the software pleasant to use, it doesn’t feel to me that it is drifting towards Jira territory, rather, it feels like it is losing the carefully considered product design because now code is cheap to generate. I’m not worried about it turning into Jira but it has lost its soul. Still a great product.
mh-9 days ago
I haven't used Linear in any serious capacity, but I feel like there's plenty of market in "Jira that doesn't feel like it hates its users" and "Jira but we care about performance".

Not saying that they're shooting for either of those segments, but someone should. Plenty of enterprise orgs that need (or are convinced they need) Jira's featureset.

stavros9 days ago
All software becomes popular because it's not the old behemoth, but it's new, fast, and simple. All software then grows and grows until it becomes the new behemoth, and the circle of life begins anew.
epolanski9 days ago
How could it be different?

Linear is essentially a small subset of Jira.

It's liked and works because it's simpler and covers most of the needs of teams without huge need of defining their own processes.

But be in business long enough and you find out that appealing to small teams and small orgs that don't require flexible and custom process definition doesn't make enough money and doesn't grow you forever.

So they will keep eventually growing in features and flexibility till they cover the reason Jira is so successful: it is more powerful and adaptable.

I'm no fan of Jira, but I've tried and seen enough of these tools to understand why it has the success that it has.

This is even more clear when you go beyond software development and need processes for teams like sales, HR, legal, etc. Jira can fit all of them.

Linear? It's a joke.

Read the full thread on Hacker News →

Related stories