206 comments
Zen5 is incredible when you're able to make the most of it. I’m super excited about Zen6.
I know nothing about workload optimization, and I'd like to know more - right now I feel like the Good Burger gif. "Yeah, I know some of these words".
Increasing instructions-per-clock is all about minimising program branching (essentially 'if' statements). Because a CPU core can execute instructions faster than main memory can fetch em. It's a fun game to look at an 'if' statement and figure out how you could instead make it an arithmetic operation :).
Maximising cache hits is all about how you structure and access data. For example, if a cache entry is n bytes long, you want to ensure your struct is smaller than n bytes. Having very consistent access patterns can also help (e.g. arrays-of-structs vs structs-of-arrays).
This field is super deep, it's very fun to learn about!
So there could be a "instructions completed" counter and a "cycle" counter, and before starting the benchmark, you record the current value of both counters, then after the benchmark, record the final values, and compute the Instructions Per Cycle (IPC) as the difference in instructions completed over the difference in cycle count.
As for optimization, at a high level it's about minimizing the amount of time any part of the CPU is waiting for other parts of the CPU. The specifics require a lot of background knowledge about how modern CPUs work, more than can fit in a post, but if you're interested, topics to read about include:
### CPU cache hierarchy
CPUs store copies of data from RAM in smaller, faster memory physically closer to where the computation happens, so it's available more quickly. The CPU decides what values to store in the cache, and gives only limited control to the program, so an optimal program needs to be careful to not make the CPU make bad caching decisions (including for synchronizing the cache between multiple threads of execution).
### Instruction pipelining and out-of-order execution (a.k.a. "superscalar" execution)
Modern CPUs operate like an assembly line. A new instruction can start executing before the previous instruction(s) finishes. CPUs also have redundant hardware, so multiple instructions can be in progress at the same step of the pipeline. But there are limitations; sometimes the input to one instruction depends on the output of the previous instruction, so the whole pipeline stalls until the result is ready. An optimal program orders its operations to avoid these stalls as much as possible.
### Branch predicition
When the code to execute depends on the result of a computation, like in an `if` statement, we call it a branch. Branches can stall the pipeline, because the CPU doesn't know what instructions to execute next until the current result is ready. However, to mitigate this, modern CPUs predict which code-path will be taken when a branch is reached and begin executing the associated instructions immediately. If the prediction is right, the pipeline stall is avoided, but if it's wrong the pipeline state has to be restored to what it was before the wrong branch started executing, which is even more expensive than a stall. Usually, the CPU predicts branches correctly, so branch prediction is a net gain. Optimal programs need to understand how the CPU makes these predictions and make their branches as predictable as possible.
### Single Instruction, Multiple Data (SIMD)
Some programs do the same operations to each item of a set of data. CPUs have so-called SIMD instructions to accelerate this by, unsurprisingly, performing the same operation to multiple items at once. For example, they could add 4, 8, or even up to 64 pairs of numbers at once, depending on the instruction set and the range of the inputs. Suitable programs are optimized by arranging their inputs and operations so that SIMD instructions can be used — either directly or by being written in a way that a compiler can translate individual operations to SIMD operations.
Does it use vector? What can you hit with SMT disabled?
I think it’s only able to exceed 4.0/thread with SMT off because of a uops cache? From what I’ve read the Zen5 front end only had a 4-wide instruction decode per thread.
As a bonus Granite Rapids is super power efficient and runs much cooler than my previous TR and Ryzens. Intel seems to be doing good things again! My only minor complaint is that P2P doesn’t work on my multiple GPU setup because every PCIe5x16 lane has its own dedicated root to the CPU, but all that bandwidth is useful for MoE models that are offloaded to RAM.
I eventually gave up and turned off memory context restore and now I just deal with the minute plus (!) time to Post.
Not sure if amd memory controllers are just garbo or what but I’m going back to intel next chance I get.
Not sure I’ll need to upgrade my computer ever again :D
The web app that's hosted on it deals with lots of images and text - over 100mil of each deduplicated, embedded, simhashed - it does the hashing, and the embedding in real time during ingest. Just handles it.
These things are absolutely insane.
Steamroller did see 30% IPC improvements over Bulldozer (all the design work would have been done before they switched to Zen), but AMD canceled the full FX version, and only ever shipped the APU version of Steamroller (with only 2 modules, aka 4 threads).
If they had shipped a Steamroller FX cpu, the generational improvements would have looked similar to many of the generational improvements that Zen received... but didn't really matter as Bulldozer started so far behind.
That's pretty much irrelevant since the AMD's FX arch's issues weren't that SW at the time wasn't using all the 8 cores. Intel dropped the Core 2 Duo and Quad into the era where most SW was still stuck in single threaded for a long time and those CPUs still ripped single-threaded SW tasks regardless.
Here's the big reasons why the FX sucked back then and why they still suck today in the multi-thread SW era:
Instead of discrete, fully independent cores, AMD grouped processing units into "Modules" where each module contained two integer execution units, but they had to share critical resources like one FPU, the instruction fetch/decode pipeline, and the L2 cache so when both "cores" inside a module were heavily taxed especially with math or physics-heavy calculations (like in videogames), they choked fighting over shared hardware.
AMD designed Bulldozer with a very long pipeline, betting they could sacrifice efficiency per clock cycle in exchange for extraordinarily high clock speeds(a-la Intel Pentium 4) but the IPC was so bad that an FX core was often slower clock-for-clock than AMD’s previous-generation Phenom II chips and also their power consumption exploded.
FX processors were plagued by high cache latencies and an inefficient memory subsystem as another bottleneck.
So unless you're into collecting vintage CPUs as display pieces, this one definitely belongs in the e-waste pile instead of burning electricity, because it did not age like wine with the adoption of SW multi threading like people were hoping.This only goes back one year, to Linux 6.18. Wins would be even bigger if we go back another year. Also, this isn't tracking any of the rest of the improvements in userland: it's just the kernel. Some newer GCC, and upcoming new x86-64v3 targets will all have some pretty nice wins too.
Great days to be on open source. And it only ever gets better.
First zen 3 November 5, 2020 with desktop processors. First desktop Ryzen 9000 processors on August 8, 2024. So the generations were about 4 years apart.
>Compared this to the article "How did AMD Ryzen get 50% faster in two years?" The Zen 3 uArch being used in the article came out in 2020. So it is more like AMD got 50% faster in 5 years.
Not implying anything here, but my comment was actually before yours, this article was reposted via HNs 2nd chance pool, https://news.ycombinator.com/pool
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 2 days ago
- Ars Technica · 0 points · 1 day ago
- Hacker News · 1 points · 8 days ago
- DEV Community · 11 points · 13 days ago
- Writing Rust code that's faster than state-of-the-art libraries by asking agents to make the code fasterminimaxir.comLobsters · 5 points · 8 days ago
- Hacker News · 1 points · 9 days ago