65 comments
I suspect a lot of the training set for this sort of thing is people online speculating about cache performance incorrectly.
At a low enough level, every performance tweak becomes unique and bespoke.
Of course, you could find people online talking about how to write high-performance code, but beyond a few basic techniques, their advice may not work for you — nobody can write a generalist article about performance engineering that will definitely solve the problem you have right now.
Arguably, there are fewer patterns for an LLM to infer as highly optimised code tends to become more and more opaque in the search for a nanosecond here or there.
Me: "Parallelise this loop without changing the results."
AI: "I can't do that, because of floating point accumulation."
Me: "So spin out a temporary fixed-length array, accumulate into that thread-locally, and then do an ordered summation of that at the end."
AI: "Oh, you're right!"
It's a fantastic tool for automating the implementation grind, but you still have to hand-feed it the strategy. I guess it will get there at some point.
Each tool in any domain "will get there at some point"
> Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).
How much experience can you replace with infinite compute remains to be seen I guess.
The great thing about LLM is that it seems to have the checklist for everything. If I rattle off a few things like "don't allocate on the hot path" and "remember to pin the cores" it will come up with a few items of its own that I might have forgotten.
Eventually, it will have gone through the whole list with me, while having documented all the measurements along the way.
But it's still guided by experience. If I see unusual numbers, I might say "hey did you forget to compile it in release mode?" and it will apologize and fix that. If I don't, it may just continue exploring without realising everything is wrong.
Once I had repo commands that could dump `sample` results and a cpu profiler/trace and then a benchmark tool that let me A/A + ABBA/BAAB-test the current modified git workspace against HEAD or any commit, the LLMs could just do their thing.
And that's how my homemade terminal uses much less memory than ghostty/kitty/iterm yet has more throughput.
AI is going to increasingly unmask people and companies who don't care about correct and performant software now that it's become so trivial to guarantee both. It used to at least be expensive and time-consuming and expertise-demanding to do those things.
This isn't a new problem by any means, but now that code is cheap, it means instead of getting frustrated with engineering and their pesky unimportant details, people will get frustrated with the AI and it's pesky unimportant details.
I think it's one reason why ADRs are an important of a software project, especially with LLMs. You need a place were you can document invariants, why you have them + the rejected ideas and acceptable risks.
It helps smart agents like Fable help you decide on trade-offs and it's kind of incredible to witness that happening.
This is why I'm not worried about being replaced for now or the forseeable future. For all of the improvements they've made, this part just never seems to change. They could slap another heuristic prompt for the edge case, but eventually it'll revert to the mean again.
I think there is a way to use LLMs to help with programming, but not when I'm not the driver in the seat writing the tests and deciding the architecture. Also I would never ship code written by them as the final product for anything I care about. Since I, like most people, find reading code to be arduous. The more fun thing to do is to force yourself to rewrite it all, treating the LLM's work as a rough draft.
They can, in fact, generate plausible performance optimization ideas on their own.
Make sure you process doesn't depend on anyone reading your mind.
When I run into things like this, it becomes a one-liner in my instructions/harness or in the canned prompt/skill I use that sets off a process.
In this case, I instruct agents to proactively build/improve diagnostic tooling if it would help them with their task + if it meets a bar of generalization/reusability (else it should be an ephemeral probe that gets abandoned at the end of the solution).
Then they can start attempting to optimize it. They can also spin round and round making the numbers worse because they don't actually know what to do.
In one case I used a made-up metric (since I didn't know the exact name or if it existed) and it somehow optimized that too.
You need a measurement that can falsify hypotheses and reject branches that won't work.
Also, if all you have left in your project are performance issues that are hard to identify without flailing around (even with Fable/Astra) despite sampler/profiler reports, then you're doing really well and I wouldn't assume you're going to fare much better than the sota models in terms of stabs in the dark.
"Claude, if this idea doesn't measure as an improvement (use X benchmark and a T-test), discard it and try the next idea."
One of the keys for me was the use of types. Typestate when functions mint witnesses that can only come from it and are required to proceed and newtypes where you use custom types instead of strings so the agent can't forget. You can also use it to force the agent to use the implementation rather than reinvent the wheel by simulating linear types. Types are a much smaller target to optimize and provide constraints that fail loudly at compile time.
Unless its an easy memory/parallel/algorithmic win, its not worth it.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 10 days ago
- Writing Rust code that's faster than state-of-the-art libraries by asking agents to make the code fasterminimaxir.comLobsters · 5 points · 9 days ago
- Hacker News · 2 points · 9 days ago
- Lobsters · 86 points · about 1 year ago
- Hacker News · 421 points · 6 days ago
- Show HN: Groundtrack – Continual learning for coding agentsgroundtrack.devHacker News · 1 points · about 16 hours ago