100 comments
> ARM is the most relaxed, allowing significant hardware optimizations; and x86 is the most strict, enforcing a very strong coherency model that doesn’t allow a lot of room for optimization
but I've seen some compelling arguments that a relaxed model doesn't necessarily have much of a benefit, https://fgiesen.wordpress.com/2026/08/25/memory-ordering-in-...
Giesen's article comes off as well-meaning cope from an x86 fan. A relaxed memory model really does give you some performance. Another memory model flaw here in x86 is more architectural, which is that every instruction with the LOCK prefix is essentially a full barrier (of course, x86 could have provided different instructions while still being under TSO). In programs that make heavy usage of atomic reference counting, this actually helps quite a bit.
I would probably put that performance benefit in the single digit percentage range like my sibling comment, which may not seem like much to a SW engineer but is actually pretty serious in CPU microarchitecture. It also helps to be stacked with other architectural advantages over x86, e.g. fixed-length instructions, 32 GPRs (which Intel copied in APX), LDP/STP (which Intel also copied in APX), etc.
One of the old arguments from TSO enjoyers was that TSO helps avoid concurrency bugs that people would accidentally introduce, but this was before the C++ memory model propagated throughout the programming world. Nowadays, I think people generally conceptualize memory consistency in terms of acquire/release anyways, so why not use a CPU architecture that uses the same model?
One problem is that 2 kinds of ordered loads and stores are not enough, 4 kinds are needed.
Some algorithms need not only load-acquire and store-release, but also a store that is guaranteed to be executed before all subsequent stores and a load that is guaranteed to be executed after all previous loads. These properties are the opposite of those provided by load-acquire and store-release. Except for x86 where any loads and stores behave like this, the other popular ISAs do not have such loads and stores, so stronger than necessary instructions must be used, i.e. load barriers and store barriers. Moreover, the instruction that is named a load barrier in the Arm ISA is not a load barrier, but a stronger barrier, but it must be used instead of a load barrier as no better alternative exists.
Besides the fact that not only load-acquire and store-release, but also other 2 ordered loads and stores are needed, a much more serious problem is that load-acquire is not the instruction that is really needed.
I have never seen any useful algorithm where load-acquire is the correct instruction to use. In all algorithms, what you want is not an instruction, but a loop that compares memory repeatedly, waiting for some condition to be fulfilled. The 4 most frequent kinds of loops that are needed are wait-for-not-equal, wait-for-equal, wait-for-even and wait-for-zero.
These loops may execute a load-acquire, but that is not the desired behavior. What you really need are 2 kinds of barriers, one inside the loop and one immediately after the loop.
The internal barrier must prevent the CPU from speculatively executing many future loop instances beyond the conditional jump that terminates the loop body, as it normally does. On x86-64, the instruction PAUSE provides such a barrier. On Aarch64, I suspect that a load-acquire instruction does inhibit this kind of speculative execution, despite the fact that this behavior is not documented. Otherwise, a CPU executing this kind of loop would waste a lot of energy and resources.
The barrier after such a loop must prevent speculative memory accesses beyond it. The semantics of load-acquire are not really needed, because all such loads are done in a loop and the memory accesses that follow the loop cannot be executed before such a load, due to the control dependency created by the conditional jump that follows the load. Nonetheless, while normal execution is impossible, the CPU can execute speculatively any loads following the loop and only this speculative execution can break the acquire semantics. Therefore what you really need is a speculation barrier after the loop, not an acquire barrier, whose behavior is provided automatically by the loop, even without any instruction with acquire semantics.
On x86-64, if a loop is terminated by an unconditional jump, it is said that such a jump blocks the speculative accesses beyond it. This is why the examples provided by Intel in its optimization manual about how to write this kind of acquire loop show loops terminated with unconditional jumps, even if this makes the loops longer, as otherwise the unconditional jump could have been eliminated by moving the conditional jump at the end of the loop. On x86-64, an alternative to unconditional jumps is the LFENCE instruction, which is a barrier for speculative memory accesses.
The so-called load-acquire instruction of Arm might also implement the textbook behavior of load-acquire, of ordering the memory accesses, despite the fact that this behavior is always superfluous, but it must also have the undocumented behavior of being a speculation barrier for memory accesses, otherwise the Arm CPUs would have been very inefficient. However, dedicated speculation barriers of the 2 kinds needed inside the loop and outside the loop would have been more efficient than a load-acquire instruction.
Even your linked paper contends that the actual observed slowdown is as much as 22% in the Geekbench example, but the thesis is that the slowdown is not inherent to TSO, but merely to the specific hardware implementation. Is it worthwhile for a company to optimize its TSO to chase the final gains, or is it better not to have this feature in the first place and just change the compiler?
Indeed my instinct is that it is better to do this in software, where the programmer clearly communicates which stores are ordered, and which may happen in arbitrary order.
> but I've seen some compelling arguments that a relaxed model doesn't necessarily have much of a benefit,
"ELI5" oversimplified explanation why this is insane and can be no other way:You write A (not in cache), you write B, you write C, ... you write Y, you read Z (in L1d).
In ARM, when A misses L1d, L2, and L3, you can issue writes to B..Y, and the read of Z (which hits in 2 cycles) and go on your merry way, your pipeline sure of the value of Z.
With TSO you CANNOT allow the writes to B..Y to be seen by anyone before A, because everyone expects to only see them after A, so you either buffer them or, eventually, sit on your proverbial ass and wait for A to drain. You might have the buffers to put off the consequences for a while, but that is explicitly a cost ARM lacks, AND(this one hurts the most) you can't simply let the load of Z become an architecturally committed observation before A is done. You can speculate it, but you have to preserve the TSO ordering constraints. Eventually you'll run out of write-buffer slots or other resources, or find that your guess of Z's value was ... wrong. This is an example of how TSO loses perf compared to relaxed. The list of how it gains perf compared to relaxed is shorter: { }
So, in some cases TSO is the same perf as relaxed, in no cases it is faster, in some cases it is slower. So it must be slower overall, since "overall" is a weighted mix of those. "By how much" is a question of detail and workload. However, unless your workload miraculously never suffers a cache miss while a later instruction hits, TSO must be slower. :)
Valve sponsor development as it’s also the way the new Steam Frame supports x86 games. It’s also being used (as a fork) in Crossover Beta to replace the use of Rosetta2.
THen they wouldnt require any emulator for any new compiles.
They still have to support the entire back-catalog. It's not reasonable to expect thousands of existing games to port to ARM
Bytecode can't really abstract the differences in memory model between the two different processors without some kind of consequence. (IE, it would be slower.) I've personally done some high performance multithreaded programming in C# / .Net, but it only "works" because C# / .Net assumes the TSO memory model. (Described in TFA.)
In contrast, games need to squeak every cycle of performance out of their chips, and optimizations can be very CPU specific. When games target bytecode, they either won't be able to take full advantage of the hardware, or otherwise will need a lot of platform-specific fallbacks (that negate the point of bytecode anyway.)
(This is why I prefer console gaming or "simple" games that don't tax the hardware.)
This may be true when FEX is executing as a usermode whole-process emulator on Linux, but it is not true when FEX is built for Windows(/Wine)'s ARM64EC mode. With ARM64EC a thread could be running a very small amount of emulated code while everything else is native.
I believe Microsoft Office is built as ARM64EC in order to support x86_64 plugins, in this case the entire suite itself (along with all the system DLLs) are native ARM64EC code and the only emulation would be for plugins. Kingdom Come Deliverance 2 has an ARM64EC build where the main game EXE is small and x86_64, but the actual game engine is in an ARM64EC DLL.
I don't know of a good solution for this though, enabling/disabling TSO needs a kernel syscall so is too slow to be doing constantly when entering/leaving emulation. With cases like KCD2 where the game itself is ARM64EC, maybe it could be faster to not use hardware TSO.
I really wonder how things are organized at Apple to allow for vertical integration to work so well. That feature alone must have involved so many people from so many different teams.
If they had tried to move to Apple Silicon and said “but none of your old software will work“ it would’ve been dead in the water. Look at how well the early Windows on ARM efforts went, although they were also hamstrung by hardware.
The M chips were already so much faster then the Intel chips Apple was using before (except on Mac Pro maybe) that it was probably still a net win.
Read the full thread on Hacker News →
Related stories
- The scourge of x86 emulationfex-emu.comLobsters · 77 points · 12 days ago
- Hacker News · 1 points · 3 days ago
- Hacker News · 2 points · 3 days ago
- Hacker News · 1 points · 4 days ago
- Hacker News · 1 points · 4 days ago
- Refract – A Quest Emulation Toolgithub.comHacker News · 1 points · about 4 hours ago