HN.zip

Everyone Should Know SIMD

65 points by WadeGrimridge - 14 comments
eska [3 hidden]5 mins ago
I don’t know zig syntax, but wouldn’t it be possible to put this common pattern into a macro and simplify it to mostly a lambda on V?
wrl [3 hidden]5 mins ago
i was having a conversation with a friend recently about simd in zig (which i have recently picked up and been having a pretty good time with). i find that simd writes decently well, though there's a few weird things:

- some builtins purport to work on simd vectors but actually just unpack the vectors and do their work per-element (e.g. running `@sin()` on a `@Vector(4, f32)` will unpack the vector, run `@sin()` 4 times, and then pack it back into a vector).

- a lot of `std.math` is scalar-only (some functions support vectors, though, and i've got a pr open for one of them and plan to do more).

- i'm certainly missing some intrinsics that i get from xmmintrin.h (rcp, rsqrt, few others).

in general though i'm finding it pretty capable.

mitchell, i know you hang around some of these comments sometimes – i noticed that in ghostty you bring in some c++ libs to do the simd heavy lifting for you. any plans to port that to zig? anything missing from the language or libs that's preventing it?

qurren [3 hidden]5 mins ago
I just do gcc -O3 and get SIMD without having to learn it
ashton314 [3 hidden]5 mins ago
In the article, Mitchel mentions how this doesn’t always work. In fact, as someone who’s worked in compiler development, I can say it’s a small miracle when it does work.
mitchellh [3 hidden]5 mins ago
Case-in-point, the example in my own post doesn't auto-vectorize with LLVM or GCC at highest optimization levels. Basically, compilers will never auto-vectorize loops with an early loop break afaik.
forrestthewoods [3 hidden]5 mins ago
auto-vectorization is not nearly as good as you would hope it to be.

The best SIMD optimizations likely require changing your data format from AoS to SoA.

ethin [3 hidden]5 mins ago
Either this or you have to do special tricks like pairwise tree reductions and hand-unroll certain portions of loops.
nylonstrung [3 hidden]5 mins ago
The one feature in Jonathan Blow's Jai language I really envy is a a single keyword to switch AoS to SoA and visa-versa at comptime
mbStavola [3 hidden]5 mins ago
Didn't he drop this feature years ago?
raegis [3 hidden]5 mins ago
What are AoS and SoA?
Georgelemental [3 hidden]5 mins ago
Array of Structs and Struct of Arrays https://en.wikipedia.org/wiki/AoS_and_SoA
nylonstrung [3 hidden]5 mins ago
Array of Structs and Struct of Arrays
formerly_proven [3 hidden]5 mins ago
And -march=native or at least -march=x86-64-v3 or similar, alternatively identifying relevant functions and manually invoking FMV and uarch specialization via target_clones. Plus non-integer code can generally not be autovectorized in normal-math mode since FP is non-commutative.
Joker_vD [3 hidden]5 mins ago
Well, then I just prompt Claude and get SIMD without having to learn it /s