build: disable x86_64 assembly by default on Clang #1945

pull Yudis-bit wants to merge 1 commits into bitcoin-core:master from Yudis-bit:build-clang-auto-asm-off changing 2 files +16 −3
  1. Yudis-bit commented at 4:27 PM on September 27, 2026: contributor

    On x86_64, modern Clang versions generate significantly faster 64-bit scalar multiplication without inline assembly. With assembly enabled, Clang incurs a 20-30% performance penalty on scalar_mul due to callee-saved register pressure (forcing 5 push/pop pairs for %rbx and %r12-%r15) and serialized mulq carry propagation. In contrast, Clang's C codegen compiles native unsigned __int128 operations with compiler-scheduled register allocation and memory operands.

    Change the default AUTO assembly selection in both CMake and Autotools to OFF when the compiler is Clang on x86_64. Explicit selection (-DSECP256K1_ASM=x86_64 or --with-asm=x86_64) remains supported. GCC retains its existing x86_64 assembly default.

    Fixes #1682.

    Benchmarks (bench_internal mul)

    Tested on x86_64 Linux with Clang 21.1.8 and GCC 15.2.0:

    Compiler Configuration scalar_mul Min (µs) scalar_mul Avg (µs) scalar_mul Max (µs)
    Clang 21.1.8 AUTO (now OFF) 0.0457 0.0540 0.0618
    Clang 21.1.8 -DSECP256K1_ASM=x86_64 0.0562 0.0668 0.0932
    GCC 15.2.0 AUTO (x86_64) 0.0484 0.0604 0.0745
    GCC 15.2.0 -DSECP256K1_ASM=OFF 0.0503 0.0515 0.0539

    In this local GCC 15.2.0 run, OFF has a lower recorded average than AUTO (0.0515 vs 0.0604 µs). The wide ranges do not support a general GCC performance claim; this PR leaves the GCC default unchanged.

    Clang default scalar_mul improves by ~19-23%.

    Verification

    • CMake configuration verified:
      • Clang defaults to assembly: OFF
      • GCC defaults to assembly: x86_64
      • Clang with -DSECP256K1_ASM=x86_64 forces assembly: x86_64
    • Autotools configuration verified:
      • Clang defaults to asm = no
      • GCC defaults to asm = x86_64
      • Clang with --with-asm=x86_64 forces asm = x86_64
    • Full test suite (tests) passes with 0 failures under the Clang default build.
    • Constant-time verification (ctime_tests under Valgrind 3.26.0) passes with 0 errors.
  2. build: disable x86_64 assembly by default on Clang
    Modern Clang compiles native 128-bit scalar multiplication without inline assembly significantly faster (~20-25% lower latency on scalar_mul) due to improved register allocation, avoidance of callee-saved register pressure (%rbx, %r12-%r15), and superscalar scheduling. Set the default AUTO assembly selection in both CMake and Autotools to OFF when compiling with Clang on x86_64. Explicit selection (-DSECP256K1_ASM=x86_64 or --with-asm=x86_64) continues to be honored. GCC retains x86_64 assembly by default. Fixes #1682.
    0a0ca71569
  3. Yudis-bit requested review from Copilot on Sep 27, 2026
  4. Copilot commented at 4:27 PM on September 27, 2026: none

    Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

  5. real-or-random commented at 7:36 AM on September 29, 2026: contributor

    GCC continues to use x86_64 assembly by default where it remains beneficial.

    Doesn't your table show what OFF is also faster for GCC (15.2.0)?

  6. real-or-random added the label build on Sep 29, 2026
  7. real-or-random added the label performance on Sep 29, 2026
  8. real-or-random added the label tweak/refactor on Sep 29, 2026
  9. Yudis-bit commented at 8:21 AM on September 29, 2026: contributor

    You're right. In my GCC 15.2.0 run, OFF has the lower average (0.0515 vs 0.0604 µs). The ranges are wide, so that table doesn't support my claim that assembly is beneficial for GCC. I've corrected the PR description. The code change here only affects Clang's AUTO selection; GCC's default stays as it was.


github-metadata-mirror

This is a metadata mirror of the GitHub repository bitcoin-core/secp256k1. This site is not affiliated with GitHub. Content is generated from a GitHub metadata backup.
generated: 2026-09-29 22:15 UTC

This site is hosted by @0xB10C
More mirrored repositories can be found on mirror.b10c.me