Go 1.27 adds an experimental platform-agnostic SIMD API

412 points•yurivish•6 days ago•152 comments•

152 comments

ImJasonH5 days ago
https://imjasonh.github.io/playground/palette-swap/ swaps colors in a provided image in wasm, entirely locally in your browser, to benchmark portable SIMD vs non-portable archsimd vs non-SIMD.

Portable SIMD is ~11% slower than non-portable SIMD in this case, but both are ~5x faster than non-SIMD.

qprofyeh5 days ago
This feature opens many doors for optimizing low-level performance in Go projects, that are already running multicore. IIRC there aren’t a lot of languages with built-in std lib support for SIMD and variants. Love the way Go is trying new stuff lately.
pjmlp5 days ago
Besides the usual C and C++, we have Java, .NET, D, Zig, Julia, Swift, Rust.

So yeah, also appreciate having Go in the group instead of manually having to write Assembly.

However not many languages adopt ways to manually write SIMD, because most of us have no idea how to write good SIMD code in first place, I surely don't.

vlod5 days ago
You probably weren't looking for a tutorial about SIMD, but just in case you were interested, Mitchell [0] did one recently that got on HN [1]

[0] Mitchell Hashimoto: "Everyone Should Know SIMD" https://mitchellh.com/writing/everyone-should-know-simd

[1] https://news.ycombinator.com/item?id=49010648

stingraycharles5 days ago
Even with languages that adopt ways to manually write SIMD, it’s mostly left to library maintainers rather than application developers.

I work for a C++ timeseries database startup that leverages SIMD about as much as we possibly can, and except for some extremely rare places we just use libraries.

Thaxll5 days ago
With AI I'm pretty sure SIMD will be easier to integrate when necessary.
abirch5 days ago
Vectorizing computations has been Matlabs secret sauce.
KeplerBoy5 days ago
Does matlab these days do stuff like JIT operator fusing to avoid memory roundtrips and take advantage of FMAs?
mastermage5 days ago
Julia does that too.
mshockwave5 days ago
Just want to say among many portable SIMD solutions I’ve seen recently (e.g. Fearless SIMD), this is the first that makes non-fixed vectors like SVE and RISC-V vector (RVV) easier to support. Glad to see they made this decision
janwas5 days ago
We pioneered this in Highway and shared some advice on the API. Great to see this decision taken :D
melodyogonna5 days ago
How so? I imagine you'd still want to constrain the length to the maximum vector size supported by the lowest platform you want to support or you lose the portability and actually end up with code that performs much worse than the scalar alternative on some platforms.

Mojo has an even more portable simd[1] type that isn't just generic over length but also over type. In my opinion it is almost always better to specialize for each platform and use portable implementation as fallback. It's a shame that just very few languages support Zig-like comptime, because it would be excellent for specializations without introducing runtime penalties.

1. https://mojolang.org/docs/std/simd/SIMD/

mshockwave5 days ago
> constrain the length to the maximum vector size supported by the lowest platform you want to support or you lose the portability and actually end up with code that performs much worse than the scalar alternative on some platforms.

Or, put a dynamic factor into your vector size and design everything around it. Such that every platforms can plug in their own factor and _scale_ the size of vectors. This is basically what LLVM IR does for SVE and RVV: `<vscale x 4 x i32>` where vscale is the said dynamic factor. Though the exact value of vscale is only known during runtime, it doesn't matter -- we still can design compiler optimizations and lowering around it. The generated binaries can then be portable across platforms with different vscale values.

beached_whale5 days ago
C++ is getting std::simd in the latest version and I am all aboard writing the vectorization with the least amount of intrinsic builtins I am able to. Even if not optimal, it's far better than the scalar ops.
reactordev5 days ago
Seconded!! This doesn’t really help the well established codebases much that are already doing this on a platform specific path but in general this is much appreciated for the future.
beached_whale5 days ago
Write it once with N errors, not N*M errors :)
sixdimensional5 days ago
I did some testing with the experimental SIMD on a project I was doing to make speech-to-text and text-to-speech models run natively in Go (with CGO_ENABLED=0, so no C depenencies), and testing non-SIMD w/ SIMD.

I don't have formal benchmarks for that, but I can anecdotally say the SIMD work made a measurable improvement in the performance of the calculations vs. just plain Go. I'm very optimistic about how these improvements will help make the Go runtime an even better target for more of these types of work going forward, especially since it is cross-platform.

Read the full thread on Hacker News →

Related stories