485 points•ibobev•12 days ago•206 comments•

206 comments

bmenrigh8 days ago
I was recently able to get a workload of mine optimized enough on Zen 5 to hit a sustained 6.0 IPC/core (3.0 / thread) at 5.1 GHz. Seeing the a > 99.8% branch prediction rate and a > 99.99% L2 cache hit rate retiring > 1T instructions every 3 seconds feels amazing.

Zen5 is incredible when you're able to make the most of it. I’m super excited about Zen6.

rithdmc8 days ago
For someone who has never touched workload optimization, can you share a little about how you measure this? Does IPC mean inter-process communication in this context?

I know nothing about workload optimization, and I'd like to know more - right now I feel like the Good Burger gif. "Yeah, I know some of these words".

j4k0bfr8 days ago
I've never done workload optimisation myself, but have looked into it just-in-case. It's a really interesting field and unfortunately, you need to know about both your compiler and target CPU/GPU to get the big gains. I believe most modern CPUs have registers for cache use. You have to do some guesswork to get IPC numbers, since core frequency can vary.

Increasing instructions-per-clock is all about minimising program branching (essentially 'if' statements). Because a CPU core can execute instructions faster than main memory can fetch em. It's a fun game to look at an 'if' statement and figure out how you could instead make it an arithmetic operation :).

Maximising cache hits is all about how you structure and access data. For example, if a cache entry is n bytes long, you want to ensure your struct is smaller than n bytes. Having very consistent access patterns can also help (e.g. arrays-of-structs vs structs-of-arrays).

This field is super deep, it's very fun to learn about!

kristofferc8 days ago
Instructions Per Cycle
MobiusHorizons7 days ago
in the context of CPU performance / architechture `IPC` pretty much always means "Instructions per clock". That's a measure of the internal parallelism a given CPU core is achieving. Most of the time code is not able to get very close to the theoretical maximums a core can achieve for very long, so the numbers the OP is quoting are incredibly impressive.
GrantMoyer8 days ago
I'm not too familiar with the CPU world of performance measurement, but in GPU land, and I suspect for CPUs too, there are a slew of hardware "performance counters" which are registers that increment every time the event they measure happens.

So there could be a "instructions completed" counter and a "cycle" counter, and before starting the benchmark, you record the current value of both counters, then after the benchmark, record the final values, and compute the Instructions Per Cycle (IPC) as the difference in instructions completed over the difference in cycle count.

As for optimization, at a high level it's about minimizing the amount of time any part of the CPU is waiting for other parts of the CPU. The specifics require a lot of background knowledge about how modern CPUs work, more than can fit in a post, but if you're interested, topics to read about include:

### CPU cache hierarchy

CPUs store copies of data from RAM in smaller, faster memory physically closer to where the computation happens, so it's available more quickly. The CPU decides what values to store in the cache, and gives only limited control to the program, so an optimal program needs to be careful to not make the CPU make bad caching decisions (including for synchronizing the cache between multiple threads of execution).

### Instruction pipelining and out-of-order execution (a.k.a. "superscalar" execution)

Modern CPUs operate like an assembly line. A new instruction can start executing before the previous instruction(s) finishes. CPUs also have redundant hardware, so multiple instructions can be in progress at the same step of the pipeline. But there are limitations; sometimes the input to one instruction depends on the output of the previous instruction, so the whole pipeline stalls until the result is ready. An optimal program orders its operations to avoid these stalls as much as possible.

### Branch predicition

When the code to execute depends on the result of a computation, like in an `if` statement, we call it a branch. Branches can stall the pipeline, because the CPU doesn't know what instructions to execute next until the current result is ready. However, to mitigate this, modern CPUs predict which code-path will be taken when a branch is reached and begin executing the associated instructions immediately. If the prediction is right, the pipeline stall is avoided, but if it's wrong the pipeline state has to be restored to what it was before the wrong branch started executing, which is even more expensive than a stall. Usually, the CPU predicts branches correctly, so branch prediction is a net gain. Optimal programs need to understand how the CPU makes these predictions and make their branches as predictable as possible.

### Single Instruction, Multiple Data (SIMD)

Some programs do the same operations to each item of a set of data. CPUs have so-called SIMD instructions to accelerate this by, unsurprisingly, performing the same operation to multiple items at once. For example, they could add 4, 8, or even up to 64 pairs of numbers at once, depending on the instruction set and the range of the inputs. Suitable programs are optimized by arranging their inputs and operations so that SIMD instructions can be used — either directly or by being written in a way that a compiler can translate individual operations to SIMD operations.

adgjlsfhk18 days ago
zen5 is really the first CPU of the avx512 era (since it's the first time normal compiler devs have them to play with)
Sharlin8 days ago
And it’s full width for the first time.
RandomOnyx8 days ago
We had Ice Lake and Rocket Lake too, they seem to have disappeared from the collective memory
stinkbeetle8 days ago
Can you share any more details about the workload? Always interesting to hear of something like that which isn't a useless microbenchmark.

Does it use vector? What can you hit with SMT disabled?

bmenrigh8 days ago
It’s a backtracking search program looking for integer solutions to a specific problem. I’ve tuned a series of bloom filters to fill my L2 cache so that I rarely have to touch main memory (this alone took my IPC from 0.1-0.3 to 3.0 per thread). Without SMT it’s 4.6 IPC/core.

I think it’s only able to exceed 4.0/thread with SMT off because of a uops cache? From what I’ve read the Zen5 front end only had a 4-wide instruction decode per thread.

ultrahax8 days ago
I work on game servers. I can only dream of such an IPC, but it doesn’t stop me trying, heh.
mitxela8 days ago
Array-oriented programming. Add all velocities to all positions, etc
jacquesm8 days ago
I Picked a threadripper for a box that needed a lot of IO and figured that with that many PCI lanes I couldn't go wrong. But I have to admit I've been more than pleasantly surprised by the performance of the CPU as well, it - easily - outperforms all of the XEON and I7 based boxes that I have. The only thing I wished I would have done different is to max it out with 256G DIMMs when they weren't the price of a car.
Gracana8 days ago
Zen 5 really is nice. I swapped my RAM from a Xeon W Sapphire Rapids machine into a Threadripper Pro 9000 series machine and I get almost double the memory read performance, plus it's a heck of a lot faster in single and multi core performance, and it runs cooler and quieter. Huge win all around. Aside from the price... (I went from Xeon W5-3435X to TR Pro 9985WX, eep.)
rubiquity8 days ago
I went the opposite way and went back to Intel after many many years on AMD. I snagged a Xeon 654 Granite Rapids ($850) for a new workstation geared towards local inference. I don’t need a ton of CPU cores and all Granite Rapids have 8 memory channels where as you need very high end Threadrippers (9975wx for $4000) for an equivalent due to their CCD design.

As a bonus Granite Rapids is super power efficient and runs much cooler than my previous TR and Ryzens. Intel seems to be doing good things again! My only minor complaint is that P2P doesn’t work on my multiple GPU setup because every PCIe5x16 lane has its own dedicated root to the CPU, but all that bandwidth is useful for MoE models that are offloaded to RAM.

WD-428 days ago
I have the complete opposite experience. Got a 9550x with 64gb ddr5 and a fairly high end mobo about two year ago. Just running the memory at stock speed. About 50% of the time I’d reboot and one or both of the sticks would only be detected as 2gb. Would need to do a hard shutdown to get it back.

I eventually gave up and turned off memory context restore and now I just deal with the minute plus (!) time to Post.

Not sure if amd memory controllers are just garbo or what but I’m going back to intel next chance I get.

jamesforestwest8 days ago
Threadripper is underrated for IO-heavy boxes. I made the same mistake with RAM prices, though; should have maxed it out when it was cheap
eptcyka8 days ago
Ye, that's like saying one should've bought Amazon, Apple and Google stock when it was low.
louthy8 days ago
I’ve been running the 64 core (128 logical cores) Threadripper PRO 3995WX in my dev machine for 4 years now (with 256gb ram).

Not sure I’ll need to upgrade my computer ever again :D

Implicated8 days ago
I picked up a pre-ai-price-insanity AX162-R at Hetzner a while back and loaded it up on memory to max out the 12 channels the 48c EPYC 9454P as a "this will be the last mysql box I'll need" and have I been _wildly_ impressed with it's performance. The things I throw at it are honestly laughable at times, wildly irresponsible queries against a rather large database, the redis qps metrics are ridiculous and I just load it up with random ggufs since the memory bandwidth is... not terrible and 384GB of it is... useful.

The web app that's hosted on it deals with lots of images and text - over 100mil of each deduplicated, embedded, simhashed - it does the hashing, and the embedding in real time during ingest. Just handles it.

These things are absolutely insane.

nicman238 days ago
i have 2 epyc boards from 2019 and they are still just printing money
Yokolos12 days ago
I suspect the improvements are even more dramatic going from Zen 1 through to Zen 5. AMD has really hit the jackpot with how scalable the Ryzen CPU is considering how they're able to improve the performance from year to year. This is a stark difference to the FX series during the 2010s, which saw very small YoY performance increases by comparison. Ryzen really is AMD's equivalent to what Nehalem/Core was for Intel back in the mid 2000s.
phire8 days ago
Part of the reason we didn't see much in the way of YoY improvements for Bulldozer, is that AMD almost immediately abandoned it and threw resources at Zen after it launched.

Steamroller did see 30% IPC improvements over Bulldozer (all the design work would have been done before they switched to Zen), but AMD canceled the full FX version, and only ever shipped the APU version of Steamroller (with only 2 modules, aka 4 threads).

If they had shipped a Steamroller FX cpu, the generational improvements would have looked similar to many of the generational improvements that Zen received... but didn't really matter as Bulldozer started so far behind.

FlowingRiver8 days ago
That is all true but I will defend the FX series a little. Mostly now that there is a lot of software that scales across cores better now, they haven't aged as terribly as others have. They aren't great but not terrible considering.
PorciiVorbesc8 days ago
>Mostly now that there is a lot of software that scales across cores better now

That's pretty much irrelevant since the AMD's FX arch's issues weren't that SW at the time wasn't using all the 8 cores. Intel dropped the Core 2 Duo and Quad into the era where most SW was still stuck in single threaded for a long time and those CPUs still ripped single-threaded SW tasks regardless.

Here's the big reasons why the FX sucked back then and why they still suck today in the multi-thread SW era:

  Instead of discrete, fully independent cores, AMD grouped processing units into "Modules" where each module contained two integer execution units, but they had to share critical resources like one FPU, the instruction fetch/decode pipeline, and the L2 cache so when both "cores" inside a module were heavily taxed especially with math or physics-heavy calculations (like in videogames), they choked fighting over shared hardware.

  AMD designed Bulldozer with a very long pipeline, betting they could sacrifice efficiency per clock cycle in exchange for extraordinarily high clock speeds(a-la Intel Pentium 4) but the IPC was so bad that an FX core was often slower clock-for-clock than AMD’s previous-generation Phenom II chips and also their power consumption exploded. 

  FX processors were plagued by high cache latencies and an inefficient memory subsystem as another bottleneck.

So unless you're into collecting vintage CPUs as display pieces, this one definitely belongs in the e-waste pile instead of burning electricity, because it did not age like wine with the adoption of SW multi threading like people were hoping.
Zardoz848 days ago
I just was to say the same thing. I had a good experience with a FX-8370E. They go for too many cores to early, and sacrificed some of the CPU performance to do it. The gamble gone wrong...
account428 days ago
Yes, my FX 8350 was doing really well on my Gentoo machine for the price. It really depends on the use case - AMD went hard in on multi core while Intel focused on single thread perf.
jauntywundrkind8 days ago
Compound this with Linux users seeing a ~8% gain (and some very big gains here and there) on performance from ongoing kernel improvements too! https://www.phoronix.com/review/linux-618-73-amd-epyc

This only goes back one year, to Linux 6.18. Wins would be even bigger if we go back another year. Also, this isn't tracking any of the rest of the improvements in userland: it's just the kernel. Some newer GCC, and upcoming new x86-64v3 targets will all have some pretty nice wins too.

Great days to be on open source. And it only ever gets better.

kristianp12 days ago
2 years sounds very fast. It seems to be because the 3d cache SKUs of each generation were released at different stages. The Zen 3 one was a later variant.

First zen 3 November 5, 2020 with desktop processors. First desktop Ryzen 9000 processors on August 8, 2024. So the generations were about 4 years apart.

ksec8 days ago
Yes I stated that as well in the previous comment [1] on how Apple got 50% faster in 3 years.

>Compared this to the article "How did AMD Ryzen get 50% faster in two years?" The Zen 3 uArch being used in the article came out in 2020. So it is more like AMD got 50% faster in 5 years.

[1] https://news.ycombinator.com/item?id=49772502

kristianp7 days ago
I was wondering where you got the info about A20 Pro: "shrink back to 9 decode width". Also I imagine the A20 Pro is very similar to an M6, but downsized.

Not implying anything here, but my comment was actually before yours, this article was reposted via HNs 2nd chance pool, https://news.ycombinator.com/pool

Read the full thread on Hacker News →

Related stories