295 points•dagmx•13 days ago•100 comments•

100 comments

pdw13 days ago
The intro of this article repeats the common assertion that

> ARM is the most relaxed, allowing significant hardware optimizations; and x86 is the most strict, enforcing a very strong coherency model that doesn’t allow a lot of room for optimization

but I've seen some compelling arguments that a relaxed model doesn't necessarily have much of a benefit, https://fgiesen.wordpress.com/2026/08/25/memory-ordering-in-...

cwzwarich13 days ago
[Disclaimer: I wrote Rosetta 2 and determined the spec for Apple's TSO mode, so I am obviously biased.]

Giesen's article comes off as well-meaning cope from an x86 fan. A relaxed memory model really does give you some performance. Another memory model flaw here in x86 is more architectural, which is that every instruction with the LOCK prefix is essentially a full barrier (of course, x86 could have provided different instructions while still being under TSO). In programs that make heavy usage of atomic reference counting, this actually helps quite a bit.

I would probably put that performance benefit in the single digit percentage range like my sibling comment, which may not seem like much to a SW engineer but is actually pretty serious in CPU microarchitecture. It also helps to be stacked with other architectural advantages over x86, e.g. fixed-length instructions, 32 GPRs (which Intel copied in APX), LDP/STP (which Intel also copied in APX), etc.

One of the old arguments from TSO enjoyers was that TSO helps avoid concurrency bugs that people would accidentally introduce, but this was before the C++ memory model propagated throughout the programming world. Nowadays, I think people generally conceptualize memory consistency in terms of acquire/release anyways, so why not use a CPU architecture that uses the same model?

spijdar13 days ago
Do you think there would be any worthwhile gains from relaxing address-dependent load ordering, like on Alpha/AXP? Or was that just a lot of extra pain for little reward?
adrian_b12 days ago
In my opinion, conceptualizing memory consistency in terms of acquire/release is wrong and it confuses many programmers.

One problem is that 2 kinds of ordered loads and stores are not enough, 4 kinds are needed.

Some algorithms need not only load-acquire and store-release, but also a store that is guaranteed to be executed before all subsequent stores and a load that is guaranteed to be executed after all previous loads. These properties are the opposite of those provided by load-acquire and store-release. Except for x86 where any loads and stores behave like this, the other popular ISAs do not have such loads and stores, so stronger than necessary instructions must be used, i.e. load barriers and store barriers. Moreover, the instruction that is named a load barrier in the Arm ISA is not a load barrier, but a stronger barrier, but it must be used instead of a load barrier as no better alternative exists.

Besides the fact that not only load-acquire and store-release, but also other 2 ordered loads and stores are needed, a much more serious problem is that load-acquire is not the instruction that is really needed.

I have never seen any useful algorithm where load-acquire is the correct instruction to use. In all algorithms, what you want is not an instruction, but a loop that compares memory repeatedly, waiting for some condition to be fulfilled. The 4 most frequent kinds of loops that are needed are wait-for-not-equal, wait-for-equal, wait-for-even and wait-for-zero.

These loops may execute a load-acquire, but that is not the desired behavior. What you really need are 2 kinds of barriers, one inside the loop and one immediately after the loop.

The internal barrier must prevent the CPU from speculatively executing many future loop instances beyond the conditional jump that terminates the loop body, as it normally does. On x86-64, the instruction PAUSE provides such a barrier. On Aarch64, I suspect that a load-acquire instruction does inhibit this kind of speculative execution, despite the fact that this behavior is not documented. Otherwise, a CPU executing this kind of loop would waste a lot of energy and resources.

The barrier after such a loop must prevent speculative memory accesses beyond it. The semantics of load-acquire are not really needed, because all such loads are done in a loop and the memory accesses that follow the loop cannot be executed before such a load, due to the control dependency created by the conditional jump that follows the load. Nonetheless, while normal execution is impossible, the CPU can execute speculatively any loads following the loop and only this speculative execution can break the acquire semantics. Therefore what you really need is a speculation barrier after the loop, not an acquire barrier, whose behavior is provided automatically by the loop, even without any instruction with acquire semantics.

On x86-64, if a loop is terminated by an unconditional jump, it is said that such a jump blocks the speculative accesses beyond it. This is why the examples provided by Intel in its optimization manual about how to write this kind of acquire loop show loops terminated with unconditional jumps, even if this makes the loops longer, as otherwise the unconditional jump could have been eliminated by moving the conditional jump at the end of the loop. On x86-64, an alternative to unconditional jumps is the LFENCE instruction, which is a barrier for speculative memory accesses.

The so-called load-acquire instruction of Arm might also implement the textbook behavior of load-acquire, of ordering the memory accesses, despite the fact that this behavior is always superfluous, but it must also have the undocumented behavior of being a speculation barrier for memory accesses, otherwise the Arm CPUs would have been very inefficient. However, dedicated speculation barriers of the 2 kinds needed inside the loop and outside the loop would have been more efficient than a load-acquire instruction.

crest12 days ago
Does any other architecture allow split cacheline atomic operations across sockets?
Aissen13 days ago
A recent study seemed to support Fabian's well written article: https://dl.acm.org/doi/epdf/10.1145/3779212.3790129
wat1000013 days ago
Is this just a question of what one considers to be “significant”? I’d consider 3% to be significant but the authors apparently don’t.
kccqzy13 days ago
It’s a question of how much resources to allocate to the hardware team, and how much resources to be distributed diffusely to the software engineers but especially to the compiler team.

Even your linked paper contends that the actual observed slowdown is as much as 22% in the Geekbench example, but the thesis is that the slowdown is not inherent to TSO, but merely to the specific hardware implementation. Is it worthwhile for a company to optimize its TSO to chase the final gains, or is it better not to have this feature in the first place and just change the compiler?

Indeed my instinct is that it is better to do this in software, where the programmer clearly communicates which stores are ordered, and which may happen in arbitrary order.

dmitrygr12 days ago

  > but I've seen some compelling arguments that a relaxed model doesn't necessarily have much of a benefit,
"ELI5" oversimplified explanation why this is insane and can be no other way:

You write A (not in cache), you write B, you write C, ... you write Y, you read Z (in L1d).

In ARM, when A misses L1d, L2, and L3, you can issue writes to B..Y, and the read of Z (which hits in 2 cycles) and go on your merry way, your pipeline sure of the value of Z.

With TSO you CANNOT allow the writes to B..Y to be seen by anyone before A, because everyone expects to only see them after A, so you either buffer them or, eventually, sit on your proverbial ass and wait for A to drain. You might have the buffers to put off the consequences for a while, but that is explicitly a cost ARM lacks, AND(this one hurts the most) you can't simply let the load of Z become an architecturally committed observation before A is done. You can speculate it, but you have to preserve the TSO ordering constraints. Eventually you'll run out of write-buffer slots or other resources, or find that your guess of Z's value was ... wrong. This is an example of how TSO loses perf compared to relaxed. The list of how it gains perf compared to relaxed is shorter: { }

So, in some cases TSO is the same perf as relaxed, in no cases it is faster, in some cases it is slower. So it must be slower overall, since "overall" is a weighted mix of those. "By how much" is a question of detail and workload. However, unless your workload miraculously never suffers a cache miss while a later instruction hits, TSO must be slower. :)

dagmx13 days ago
For reference , Fex is a translation framework for x86 to ARM much like Apple’s Rosetta2 and Microsoft’s Prism.

Valve sponsor development as it’s also the way the new Steam Frame supports x86 games. It’s also being used (as a fork) in Crossover Beta to replace the use of Rosetta2.

MiroslavPokorny13 days ago
Why doesnt Stream require their binaries to be compiled to some bytecode and transpiled during the install ?

THen they wouldnt require any emulator for any new compiles.

swiftcoder13 days ago
> Why doesnt Stream require their binaries to be compiled to some bytecode and transpiled during the install ?

They still have to support the entire back-catalog. It's not reasonable to expect thousands of existing games to port to ARM

rwmj13 days ago
I don't know, but assumed that Valve doesn't require studios to recompile their software or use any special tooling, it's basically just packaging of existing executables. This is also why they do Windows on Linux emulation.
gwbas1c13 days ago
A few reasons:

Bytecode can't really abstract the differences in memory model between the two different processors without some kind of consequence. (IE, it would be slower.) I've personally done some high performance multithreaded programming in C# / .Net, but it only "works" because C# / .Net assumes the TSO memory model. (Described in TFA.)

In contrast, games need to squeak every cycle of performance out of their chips, and optimizations can be very CPU specific. When games target bytecode, they either won't be able to take full advantage of the hardware, or otherwise will need a lot of platform-specific fallbacks (that negate the point of bytecode anyway.)

(This is why I prefer console gaming or "simple" games that don't tax the hardware.)

matheusmoreira13 days ago
Existing games will not be recompiled for the new bytecode target, and they want all of those games to work regardless.
wat1000013 days ago
That dream of write-once-run-anywhere has been attempted for decades and is still a massive struggle. And Steam isn’t in a position to mandate that kind of massive change. They’re big, but they still have competition from other stores and from direct sales.
mrpippy12 days ago
> A potential concern is that when jumping between x86 emulation and ARM code, that the ARM code will pay unnecessary overhead due to all its accesses being TSO now. While this is a reasonable concern, the amount of ARM native code executing under emulation approaches 0%.

This may be true when FEX is executing as a usermode whole-process emulator on Linux, but it is not true when FEX is built for Windows(/Wine)'s ARM64EC mode. With ARM64EC a thread could be running a very small amount of emulated code while everything else is native.

I believe Microsoft Office is built as ARM64EC in order to support x86_64 plugins, in this case the entire suite itself (along with all the system DLLs) are native ARM64EC code and the only emulation would be for plugins. Kingdom Come Deliverance 2 has an ARM64EC build where the main game EXE is small and x86_64, but the actual game engine is in an ARM64EC DLL.

I don't know of a good solution for this though, enabling/disabling TSO needs a kernel syscall so is too slow to be doing constantly when entering/leaving emulation. With cases like KCD2 where the game itself is ARM64EC, maybe it could be faster to not use hardware TSO.

asksomeoneelse13 days ago
Great article ! This is the kind of content I always hope to find on HN's front page.

I really wonder how things are organized at Apple to allow for vertical integration to work so well. That feature alone must have involved so many people from so many different teams.

EraYaN13 days ago
The actual product people understood it to be paramount to the success of the product and so I feel like from up high there was actual commitment.
MBCook12 days ago
They’ve been through this a few times before, and they absolutely know how important it is.

If they had tried to move to Apple Silicon and said “but none of your old software will work“ it would’ve been dead in the water. Look at how well the early Windows on ARM efforts went, although they were also hamstrung by hardware.

modeless13 days ago
As noted in the article, Apple solved this problem six years ago by simply adding an x86-compatible memory ordering mode to their chip when x86 emulation became important. Yet another way Apple's chips lead the industry.
saagarjha13 days ago
Well well well “modeless” has decided to finally see the light of modes
tancop13 days ago
Arm was never modeless. Thumb is a separate encoding with different instruction semantics and Jazelle ran Java bytecode. Both of them need a special branch instruction to enter. What they don't have is legacy modes like real, v8086 or native 16/32 protected that have no reason to exist when a CPU in long mode can run 16 and 32 bit code (in compatibility sub mode) just fine.
gavinsyancey13 days ago
And as noted in the article, while that helps a lot with most of the issues, there are some corner-cases they still don't handle.
MBCook12 days ago
Yeah, I thought that was interesting. Knowing Apple they must’ve profiled a ton of code and decided the hit from not “fixing“ that wasn’t worth enough.

The M chips were already so much faster then the Intel chips Apple was using before (except on Mac Pro maybe) that it was probably still a net win.

mitxela13 days ago
A legitimate benefit of vertical integration. They control their own CPUs so they can just do that. Linux has to run on whatever it's given.
karel-3d13 days ago
The word "simply" is doing a lot of work there
astrange12 days ago
It's not the only such CPU, Fujitsu's also have TSO.

Read the full thread on Hacker News →

Related stories