build: disable x86_64 assembly by default on Clang #1945

pull Yudis-bit wants to merge 1 commits into bitcoin-core:master from Yudis-bit:build-clang-auto-asm-off changing 2 files +16 −3
  1. Yudis-bit commented at 4:27 PM on September 27, 2026: contributor

    On x86_64, modern Clang versions generate significantly faster 64-bit scalar multiplication without inline assembly. With assembly enabled, Clang incurs a 20-30% performance penalty on scalar_mul due to callee-saved register pressure (forcing 5 push/pop pairs for %rbx and %r12-%r15) and serialized mulq carry propagation. In contrast, Clang's C codegen compiles native unsigned __int128 operations with compiler-scheduled register allocation and memory operands.

    Change the default AUTO assembly selection in both CMake and Autotools to OFF when the compiler is Clang on x86_64. Explicit selection (-DSECP256K1_ASM=x86_64 or --with-asm=x86_64) remains supported. GCC continues to use x86_64 assembly by default where it remains beneficial.

    Fixes #1682.

    Benchmarks (bench_internal mul)

    Tested on x86_64 Linux with Clang 21.1.8 and GCC 15.2.0:

    Compiler Configuration scalar_mul Min (µs) scalar_mul Avg (µs) scalar_mul Max (µs)
    Clang 21.1.8 AUTO (now OFF) 0.0457 0.0540 0.0618
    Clang 21.1.8 -DSECP256K1_ASM=x86_64 0.0562 0.0668 0.0932
    GCC 15.2.0 AUTO (x86_64) 0.0484 0.0604 0.0745
    GCC 15.2.0 -DSECP256K1_ASM=OFF 0.0503 0.0515 0.0539

    Clang default scalar_mul improves by ~19-23%.

    Verification

    • CMake configuration verified:
      • Clang defaults to assembly: OFF
      • GCC defaults to assembly: x86_64
      • Clang with -DSECP256K1_ASM=x86_64 forces assembly: x86_64
    • Autotools configuration verified:
      • Clang defaults to asm = no
      • GCC defaults to asm = x86_64
      • Clang with --with-asm=x86_64 forces asm = x86_64
    • Full test suite (tests) passes with 0 failures under the Clang default build.
    • Constant-time verification (ctime_tests under Valgrind 3.26.0) passes with 0 errors.
  2. build: disable x86_64 assembly by default on Clang
    Modern Clang compiles native 128-bit scalar multiplication without inline assembly significantly faster (~20-25% lower latency on scalar_mul) due to improved register allocation, avoidance of callee-saved register pressure (%rbx, %r12-%r15), and superscalar scheduling. Set the default AUTO assembly selection in both CMake and Autotools to OFF when compiling with Clang on x86_64. Explicit selection (-DSECP256K1_ASM=x86_64 or --with-asm=x86_64) continues to be honored. GCC retains x86_64 assembly by default. Fixes #1682.
    0a0ca71569
  3. Yudis-bit requested review from Copilot on Sep 27, 2026
  4. Copilot commented at 4:27 PM on September 27, 2026: none

    Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.


github-metadata-mirror

This is a metadata mirror of the GitHub repository bitcoin-core/secp256k1. This site is not affiliated with GitHub. Content is generated from a GitHub metadata backup.
generated: 2026-09-28 09:33 UTC