200 points•verdagon•5 days ago•52 comments•

52 comments

ack_complete3 days ago
AArch64 definitely has a much more comprehensive baseline than x86-64, but there are some optional extensions that are situationally impactful, including the Crypto extension and some of the newer accumulation / dot product instructions. And unlike Intel, ARM has no portable equivalent to CPUID for querying feature flags and is terrible at documenting which intrinsics require specific FEAT_* flags.

The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs. For basically any mainstream x86 CPU, it's trivial to find documentation listing what ISA level it supports and general execution widths and latencies for common operations. For the majority of ARM CPUs, there's absolutely nothing. ARM only has optimization guides for selected Cortex cores, and NVIDIA published info for their Olympus core. But execution details had to be reverse engineered for Apple M1, and there is nothing for Oryon. This is especially bad for in-order cores, which unfortunately is still relevant because new CPUs are still being shipped with in-order efficiency cores.

officialchicken3 days ago
I really hope this is my last x86-64 CPU, Intel has become an incompetent steward. The "killer feature" for AVX-2(56) at the time of original release was basically lag/jitter-free video playback. IMO, 512 should have never been released for desktop CPUs and restricted to servers. One day I will to switch to a mainstream Neoverse dev box running linux. I also target Cortex-M in rust, so it's got a lot of the typical issues related to missing docs (e.g. bringup of non-heterogenous cores, meaning that M3/M4 still can't be used in a big.little chip)
nixon_why693 days ago
> IMO, 512 should have never been released for desktop CPUs and restricted to servers.

Why not? To save die space?

tancop3 days ago
> ... you can just assert that they’re all recent enough to at least have AVX2 that was introduced over 10 years ago, and have the program crash or misbehave if it ever runs on anything without AVX2

> However, if you are distributing the binaries for other people to run, that’s not really an option.

This all depends on what kind of software you're making. A lot of games set their requirements about 5 generations back, like FC 27 where the minimum is a Ryzen 1600. That lets them use AVX2 unconditionally and prevent complaints from users who tried to run it with a super old CPU.

Then you get whole Linux distros like CachyOS and Clear (RIP) that rebuild the world for each architecture level and have them as separate variants. I think it still counts as binaries for other people.

sharktheone3 days ago
I am hoping for portable SIMD so much. But I still think that often a manually rolled SIMD will be faster.

Also the state of SIMD in Cranelift is also very WIP. They pretty much just support a subset of 128bit vectors with some rare exceptions.

dwattttt3 days ago
I guarantee you with my lack of skill, my attempt at using portable simd will exist while using manual intrinsics I doubt I'd get there.

The question for me is whether portable simd will result in faster code than plain auto-vectorisation; for the simplest loops auto has me beat (the few times I've tried it), but I imagine as the complexity grows I'll be more likely to try do something that breaks auto-vectorisation, and it'll be more obvious to me when I do that in portable simd.

cbolton3 days ago
The handwavy section on RISC-V is underwhelming. Yes the ecosystem is still nascent but something like the SpacemiT K3 is far from "abysmal" for vector operations and would make a nice test case with its two core types both supporting RVV 1.0 (one with vlen 256, the other with vlen 1024). On the more industrial side, for example the SiFive Intelligence X280 has been used in TPU by Google for years already and I doubt they are the only ones so calling the ISA completely irrelevant is a bit of a stretch.

Just for the sake of curiosity it would be nice to have a peek at what SIMD in Rust looks like on RISC-V today. Yes, even if it requires some "obscure compiler flags" for now (while we wait for the Oilsm extension).

camel-cdr2 days ago
Also, rust could just be based and make -mno-strict-align the default.

Read the full thread on Hacker News →

Related stories