The author presents a novel vectorizable polyfill for clz and ctz bit-counting operations using floating-point arithmetic tricks, achieving significant speedups over scalar hardware instructions on vectorized workloads. The clz implementation uses double-precision exponent extraction via bitwise manipulation, running at 0.45 ns/iteration on Haswell versus 1 ns for the scalar version.
Background
CLZ (count leading zeros) and CTZ (count trailing zeros) are common bit-manipulation instructions found in modern CPU instruction sets like x86 BMI and ARM NEON, widely used in compilers, cryptography, and data structures.
- Source
- Lobsters
- Published
- Oct 10, 2026 at 01:45 AM
- Score
- 6.0 / 10