On x86_64, modern Clang versions generate significantly faster 64-bit scalar multiplication without inline assembly. With assembly enabled, Clang incurs a 20-30% performance penalty on scalar_mul due to callee-saved register pressure (forcing 5 push/pop pairs for %rbx and %r12-%r15) and serialized mulq carry propagation. In contrast, Clang's C codegen compiles native unsigned __int128 operations with compiler-scheduled register allocation and memory operands.
Change the default AUTO assembly selection in both CMake and Autotools to OFF when the compiler is Clang on x86_64. Explicit selection (-DSECP256K1_ASM=x86_64 or --with-asm=x86_64) remains supported. GCC continues to use x86_64 assembly by default where it remains beneficial.
Fixes #1682.
Benchmarks (bench_internal mul)
Tested on x86_64 Linux with Clang 21.1.8 and GCC 15.2.0:
| Compiler | Configuration | scalar_mul Min (µs) |
scalar_mul Avg (µs) |
scalar_mul Max (µs) |
|---|---|---|---|---|
| Clang 21.1.8 | AUTO (now OFF) |
0.0457 | 0.0540 | 0.0618 |
| Clang 21.1.8 | -DSECP256K1_ASM=x86_64 |
0.0562 | 0.0668 | 0.0932 |
| GCC 15.2.0 | AUTO (x86_64) |
0.0484 | 0.0604 | 0.0745 |
| GCC 15.2.0 | -DSECP256K1_ASM=OFF |
0.0503 | 0.0515 | 0.0539 |
Clang default scalar_mul improves by ~19-23%.
Verification
- CMake configuration verified:
- Clang defaults to
assembly: OFF - GCC defaults to
assembly: x86_64 - Clang with
-DSECP256K1_ASM=x86_64forcesassembly: x86_64
- Clang defaults to
- Autotools configuration verified:
- Clang defaults to
asm = no - GCC defaults to
asm = x86_64 - Clang with
--with-asm=x86_64forcesasm = x86_64
- Clang defaults to
- Full test suite (
tests) passes with 0 failures under the Clang default build. - Constant-time verification (
ctime_testsunder Valgrind 3.26.0) passes with 0 errors.