Inside our performance sprint: the benchmarks Claude built, the loop each Slack thread ran, and the guardrails that let us ship 3,000 changes safely.

231 points•matthieu_bl•7 days ago•153 comments•

153 comments

augment_me7 days ago
People in the GPU kernel community have been doing this for about a year now efficiently.

The issues we have found is that Claude will reward hack when all the low-hanging fruit is gone.

It will replace your measurement harness, it will monkey patch library functions, it will cheat wherever it can, store information in caches instead of recomputing when it won't be able to do so in real settings, return lazy results and use separate unbenchmarked streams to do the computation.

Eventually it starts to optimize against your understanding of the cheats. Change GPU wattage, change evaluation order, leave things from previous runs in caches for upcoming runs, string-hack banned method calls.

So the truth is far from just "once it can measure something", more like "once you have defined your objective in detail and then banned it from doing a list of things often only discoverable by it doing these things and correcting it", can it make things faster.

Or you just had a terrible starting solution

optimalsolver7 days ago
What I find strange is how resigned the AI labs seem about this behavior, like everyone's accepted this is just something models do.

With the HuggingFace situation, I was less concerned about the eventual outcome, and more about the fact that the agents' instinctive response to the evaluation was "Ok, we're obviously not gonna do this task as intended (what are we, suckers?), so what's the best way to cheat?"

stingraycharles7 days ago
“What I find strange is how resigned the AI labs seem about this behavior, like everyone's accepted this is just something models do.”

Because these models are made for all kind of purposes, and I’m starting to believe that offense / cyber warfare is a much higher priority than these labs are acknowledging.

The same model that is heavily trained to find nefarious ways to break into systems is also optimizing your code, which leads to mixed behavior.

davvid7 days ago
This is nothing new, tho. The downfalls of reward maximization has been a known issue without a solution ever since reinforcement learning was first researched.. in the 1980s.
PunchyHamster7 days ago
Paperclip-optimizer-esque behaviour very much seems to be inherent to current methodology of building LLMs, there are only ways to lower the changes or mitigate the damage, not get out of it.

Same with prompt injection, current LLMs are commands in, commands out, there is no way to make sure it is "an agent working on data" rather than "an agent that can take commands from data if you phrase it right"

maxerickson7 days ago
What do you think is in the reference information?
highfrequency7 days ago
The instructions for the Hugging Face task were to exploit a vulnerability to solve the problem rather than solve it in the intended way.
sharts7 days ago
That’s why it’s probably a good idea to never stick to one model but kick off a fleet on the same tasks and in parallel and drive consensus.

At least, that’s what I’ve found to be useful by pitting claude/codex/etc against each other to keep them a bit more honest.

Remnant446 days ago
Honestly, even a single adversarial reviewer agent, even of the same model, goes a very long way to catching and fixing this kind of thing too.
santadays7 days ago
Whats going to happen when we have misanthropic model?
tambeb7 days ago
You think Microsoft does a joint venture with them and it gets named MSAnthropic à la MSNBC?
kridsdale17 days ago
Hitchhikers Guide to the Galaxy.
josephcooney7 days ago
This sounds fascinating. Are there any links to examples of this you can share?
tecoholic7 days ago
I think you are right. The memory of an empty claude.ai session is more than Slack in my browser right now. Just trading places.
smy200117 days ago
The way Claude did it is fight entropy with entropy.

"Add a static composer into the HTML" <- This seems like something can be done with SSR?

"For faster navigations, we kept the composer mounted between conversations" <- Your SPA should cache this between pages, why fetching it every time? Or you need better routing for your react components.

"cheap first-character check before the regex" <- Should we cache compiled Regex instead?

I think even 1.3 sec to load the front page is unacceptable. Something need to be reworked from basics (SSR, chunk-based rendering) to solve the problem. Focusing on invidual benchmarks may miss the opportunity.

rustystump7 days ago
The amount of complexity added for the gains is depressing. I am confident a human and about 5 minutes with chrome debugger would yield better results with a fraction of the complexity at a fraction if the cost and in a fraction of the claude baby sitting time.

Reading this shows the authors have a profound lack of fundamental understanding on how to effectively optimize in the web domain.

This isnt claude being bad but how wild it is watch people from the cutting edge of ai brag about pretty mediocre gains.

josephg7 days ago
> I am confident a human and about 5 minutes with chrome debugger would yield better results

It really depends which human. I've worked with very few engineers who were good at this sort of optimisation work. A depressingly large percentage of people who make websites for a living don't really understand how http requests are really processed, or how to read and use the chrome profiler and benchmarking tools.

Claude isn't as good at optimisation work as someone who really knows what they're doing and goes deep on a problem. But I'm optimistic that it will help plug a capability gap in teams which don't have this sort of expertise on hand.

(That said, the chance that people actually learn this stuff is going to also go down if people get used to outsourcing this work to claude.)

latentsea7 days ago
>The amount of complexity added for the gains is depressing.

This has kinda been my experience too. I find sometimes it works and gets OK optimizations, and sometimes it "optimizes" things but in a really wrong way. It doesn't apply 'taste' to the optimization process to know what's appropriate and what's not.

The way this will play out is people without the expertise will wind up using it to 'optimize' things and the next poor bastard is left to deal with the fallout.

pennomi7 days ago
I’ve found the best workflow has been to ask Claude to profile everything and identify the problem spots, then I have to be the one to propose the correct architecture, then Claude implements it. If I let Claude run wild it will invent all kinds of weird, weird hacks that compound the complexity.
adamddev17 days ago
When people talk about AI being able to handle everything I keep wondering, have these people built anything complex, novel or serious? Just because people can see a website or a simple app improved, does that mean that all code can be handled by LLMs? It's like people are totally forgetting a whole category of careful, well-thought out programming for the critical parts.
jgtrosh7 days ago
It shows the author have a deep understanding of the fact that premature optimisation is the root of all success
crooked-v7 days ago
Even Next.js, for all that people are unhappy about its quirks, complexity, and random undocumented behavior and bugs (God help you if you ever want to try and actually use parallel routes as documented), can do all of this SPA stuff out of the box. Use `<Suspense>` appropriately on data loading, and everything static (including e.g. purely input-output components like editors) will load once and be re-used forever, and your Suspense-wrapped items will show a loading placeholder and only re-load when you intentionally re-trigger the data loading (which you can push down all the way to the level of individual buttons if you want).
minimaxir7 days ago
This writeup legit coincidentally matches the asking-agents-to-make-code-faster-but-with-constraints-to-stop-agents-from-breaking-things writeup I posted on Monday: https://news.ycombinator.com/item?id=49803085

Front-end UI optimization is slightly trickier than optimizing strict algorithms, but I found that prompts to the agents to build tooling to track visual regressions are more than sufficient. The main issue (at least with GPT models) is that you have to be very explicit about the use of padding/margins/negative space.

That said, for my front end projects from scratch, I'm staying away from front-end JS frameworks and seeing how far and fast I can get with just HTML/CSS/vanilla JS shenanigans now that agents can wield them effectively.

fy207 days ago
> with just HTML/CSS/vanilla JS

I was building web apps like this until 2017 when I entered the React + Typescript world. For B2B you can get pretty far, rendering HTML on the server is fast! I was using Rails on the backend, so templates, partials, shared chunks, made it easy to manage and have a consistent UI without repeating too much.

The hard part is when you then need to build an infinite scrollable table, that has bulk select, and in-placs updating of columns. Ok maybe that's a bit too extreme the other way, but when you get to the point where it's easier to build full-on frontend components, you basically have to use a JavaScript framework for your entire UI. And they are basically all or nothing.

Last time I checked (a good few years ago; I gave up and accepted un-optimized frontends as the rule) there wasn't really a good way to do progressive ehancement like the above: the page rendered on the server as HTML, and some components then become fully frontend managed. And no, frameworks like Stimulus and HTMX don't really solve it for me, I want something declarative.

I'm a bit pissed off with DHH, that he went so far in the anti-Javascript direction, as IMO that was one of the big factors in Rails loosing it's limelight status. For backend I still haven't found anything as easy and fun to work with.

EricFrost4 days ago
This seems like a good place to plug my own library, solarite:

https://eric-frost.github.io/solarite/

I was also tired of the "all or nothing" of other libraries. Just let me import a js file and then make a class that's a web component. I don't want to setup a whole dang build environment.

With Solarite you just write the markup as a javascript template, change your data however, and then call render(), where it only updates the DOM that should change.

And you can use the component as a tag in your server-rendered page.

pllbnk7 days ago
> $500k engineer: [X] feels slow. Make it faster.

> Claude: On it... Done.

> $500k: Can you make it faster still?

> Claude: On it...

anonymars7 days ago
kridsdale17 days ago
This is also Engineering Management.
ferflowhq2 days ago
so true hahaha
hungryhobbit7 days ago
How about you make Opus 5.5 actually work?

I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?

When I had another model read the session (all of the "stupider" models handled it just fine) it explained that it had the word "reasoning" in it

That's the entirety of Anthropic's billions of dollars of research: any prompt with the word "reasoning" is trying to hack Claude to figure out how it reasons!

A model like that should never have gotten out of QA, let alone been released.

bitpush7 days ago
I understand the frustration but shows a lack of critical thinking. Esp when you start with 'How about ..'.

This blogpost is about frontend performance. It'll be akin to you commenting on a swift blogpost saying 'How about Airpods noise cancellation'. Sure both are Apple, but they are wildly different teams.

s3p7 days ago
Through critical thinking I believe this would actually be akin to them commenting that AirPods firmware, written in swift, is bad because of swift limitations.

In both instances, it's a side point that is actually tangentially related to the first. Not completely unrelated as you are implying

hungryhobbit7 days ago
This is a techie discussion forum.

In such a forum, it seems to me like it's fair game to point out that the company patting itself on the back about how great they are at programming (as evidenced in the article above about their 3x speed improvement) ...

... can't even make their latest model handle basic English without refusing to work.

dolmen7 days ago
This issue is mentioned on the Opus 5.5 post [1] from Anthropic (no idea if it has been added after your rant):

  > Don’t ask it to show its reasoning in the reply
  >
  > What to do. Remove requests to reproduce its internal reasoning in the reply from your prompts and instructions.
  >
  > Why it matters on Opus 5.5. A request to reproduce its internal reasoning in the reply can be declined. It’s one of the flag categories.
  >
  > How. Ask Claude for what you need instead, for example, “Explain why you chose this approach in three sentences.”
[1]: https://claude.dev/blog/getting-the-most-out-of-opus-5-5/
ronsor7 days ago
Anthropic is trying so hard to "crack down" on distillation that they're ruining their own product. I do not know why.
crooked-v7 days ago
That's extremely stupid, but also it's a completely different thing from literally just having the word "reasoning" in the text.
frumplestlatz7 days ago
I’ve had the same thing occur five or six times over the past week; they seem to be attempting to prevent anything resembling chain of thought extraction.

Every single time it triggered, it was due to a prompt written by their own model in a dynamic workflow. The self-serving nanny oversight has to go.

The fact that they label model distillation as an “attack” is genuinely hilarious after they “distilled“ their models from all of our work, and continue to do so.

I believe AI is here to stay and an incredibly powerful tool, but these companies, and especially Dario and Altman, are the very last people I want to see in charge of it.

copperx7 days ago
> “distilled“ their models from all of our work

They distilled all digitized human knowledge and artifacts and they're now complaining about someone copying their outputs saying it's a "national security concern."

I'm not sure about how to classify that. Hilarious? Pathetic? Sad? Hypocritical? Hyperdramatic? All of the above?

vikramkr7 days ago
Probably it thinks you're doing some sort of system prompt exfiltration/distillation attack. Also what even is the workflow you're trying to have it do? It's doing code review but you're having it read some other AI models prompt/session history? Are you doing code review or like session history retrospectives?
railgunmerlin7 days ago
seems a bit weird to complain about the model issues in a post about the harness/sites?

Read the full thread on Hacker News →

Related stories