scalar: Make 4x64 C code competitive with x86_64 asm, disable asm by default #1951

pull real-or-random wants to merge 3 commits into bitcoin-core:master from real-or-random:worktree-scalar-asm changing 8 files +172 −63
  1. real-or-random commented at 2:40 PM on October 1, 2026: contributor

    With recent Clang versions, the C implementation of the 4x64 scalar multiplication is about as fast as the x86_64 assembly (#1682, #1945). With GCC, it is considerably slower: in the benchmark below, it is 17–20% slower than the assembly on a P-core. GCC compiles the carry handling of the accumulator macros into code that saves carries in registers (setc) and adds them separately, instead of emitting add-with-carry (adc) chains.

    This PR changes the C code so that it is about as fast as the assembly with GCC 14 and newer and faster with Clang, and then disables the assembly by default:

    • scalar: Use a 128-bit accumulator in 4x64 C code. The lower 128 bits of the 192-bit accumulator (c0,c1,c2) are kept in a secp256k1_uint128, and carries are computed by two new int128 functions, secp256k1_u128_accum_mul_carry and secp256k1_u128_accum_u64_carry. The native implementation uses __builtin_add_overflow on 128-bit values, for which GCC 14 and newer emit adc chains. Without the builtin, and in the struct implementation, the carries are computed on 64-bit halves as before.
    • scalar: Force inlining of 4x64 mul_512 and reduce_512. Without inlining, the speed of GCC's code depends on incidental details of the source, such as whether an expression is written directly or comes from an inline function: this changes the operand order of the 128-bit additions in GCC's intermediate representation.[^1] With inlining, GCC generates the same code independently of these details, and Clang can schedule across the boundary between multiplication and reduction.
    • build: Disable x86_64 assembly by default. The assembly remains available via --with-asm=x86_64 (Autotools) and -DSECP256K1_ASM=x86_64 (CMake), or via auto. Since the field assembly was removed earlier, the option only affects scalar multiplication.

    GCC 13 is still somewhat common, e.g., it's in Ubuntu LTS 24.04. Once it's no longer widespread, we could remove the assembly entirely.

    Benchmarks

    bench_internal scalar_mul measures the latency of a chain of dependent scalar multiplications. The setup was an Intel i7-1260P with turbo disabled, with the process pinned to a P-core (Golden Cove) or an E-core (Gracemont). The tables show the median time in µs over 15 interleaved runs. Percentages are relative to the assembly built with the same compiler.

    P-core:

    Compiler asm C (master) C (this PR)
    GCC 13.5 0.0549 0.0661 (+20%) 0.0650 (+18%)
    GCC 14.4 0.0548 0.0654 (+19%) 0.0551 (+1%)
    GCC 16.2 0.0541 0.0634 (+17%) 0.0556 (+3%)
    Clang 18.1 0.0558 0.0563 (+1%) 0.0459 (−18%)
    Clang 22.1 0.0531 0.0573 (+8%) 0.0474 (−11%)

    E-core:

    Compiler asm C (master) C (this PR)
    GCC 13.5 0.1050 0.0991 (−6%) 0.1160 (+10%)
    GCC 14.4 0.1050 0.0999 (−5%) 0.1010 (−4%)
    GCC 16.2 0.1050 0.0988 (−6%) 0.1020 (−3%)
    Clang 18.1 0.1070 0.0934 (−13%) 0.0873 (−18%)
    Clang 22.1 0.1030 0.0929 (−10%) 0.0885 (−14%)

    With GCC 14 and newer, the C code in this PR is within 3% of the assembly on the P-core and faster than it on the E-core. With Clang 18 and 22, it is 11–18% faster than the assembly on both core types.

    GCC 13 does not benefit from this PR, because it lacks GCC 14's improvements in compiling add-with-carry. On the P-core, the C code is still 18% slower than the assembly. On the E-core, it is 17% slower than the C code on master. The CHANGELOG entry notes that the assembly may still be noticeably faster with GCC 13 or older. Bitcoin Core's release binaries are built with GCC 14 for Linux and Windows and with Clang 19 for macOS.

    Builds with asm enabled are affected only through the force inline, which is only effective on GCC because Clang inlines even without it.

    Builds with the int128 struct implementation (e.g., MSVC) use the same carry computation as before, and their speed is unchanged.

    Call for benchmarks

    Benchmarks on other CPUs, in particular AMD, would be welcome.

    [^1]: Claude Opus 5.5 claims that it found three GCC missed-optimization bugs here. I may or may not report them upstream, but this shows why other projects such as fiat-crypto, OpenSSL, ... either get a performance hit or, like us, resorted to inline asm to convince GCC to output proper adc chains. While there's proper pattern recognition for corresponding plain C code, some later optimization passes apparently kill its finding, and the resulting binary still doesn't use proper adc chains. Happy to share the details upfront here if you're interested.

  2. scalar: Use a 128-bit accumulator in 4x64 C code
    Keep the lower 128 bits of the 192-bit accumulator (c0,c1,c2) in a
    secp256k1_uint128 and compute carries with the new int128 functions
    secp256k1_u128_accum_mul_carry and secp256k1_u128_accum_u64_carry.
    
    With the previous code, GCC materializes carries in registers (setc)
    instead of emitting add-with-carry chains. This is because some of
    GCC's tree-level optimizations break the chains, e.g., by folding
    additions that provably cannot overflow or by combining sums of carry
    bits. The native int128 implementation of the new functions uses
    __builtin_add_overflow on 128-bit values, which GCC splits into 64-bit
    operations only after these optimizations, so both GCC and Clang emit
    add/adc/adc $0 sequences similar to the handwritten assembly.
    
    Without __builtin_add_overflow, the native implementation falls back to
    computing carries on 64-bit halves, because computing them via 128-bit
    comparisons can result in branches. The struct implementation uses the
    same carry handling as the previous portable scalar code.
    
    The x86_64 assembly remains available via USE_ASM_X86_64.
    73988f42ef
  3. scalar: Force inlining of 4x64 mul_512 and reduce_512
    Without inlining, GCC generates slower code for the 128-bit accumulator
    in secp256k1_scalar_mul_512, depending on details such as how the
    accumulator is updated in the extract macro. With both functions
    inlined into secp256k1_scalar_mul, GCC generates the same code
    independently of these details. Clang benefits from scheduling across
    the boundary between multiplication and reduction.
    
    In a scalar_mul latency benchmark (bench_internal) on an Alder Lake
    P-core, this is 4% faster with GCC and 10% faster with Clang.
    da86d74e64
  4. build: Disable x86_64 assembly by default
    With GCC 14 or newer and with Clang, the C implementation of scalar
    multiplication is now about as fast as the x86_64 assembly or faster.
    In a scalar_mul latency benchmark on Alder Lake, the C code is within
    3% of the assembly on a P-core with GCC 14.4 and 16.2, faster on an
    E-core with these compilers, and 11-18% faster with Clang 18 and 22.
    
    The assembly remains available via --with-asm=x86_64 (Autotools) and
    -DSECP256K1_ASM=x86_64 (CMake), or via "auto".
    a2d5d5b117
  5. real-or-random added the label build on Oct 1, 2026
  6. real-or-random added the label performance on Oct 1, 2026
  7. real-or-random added the label tweak/refactor on Oct 1, 2026
  8. hebasto commented at 3:30 PM on October 1, 2026: member

    Concept ACK.

  9. theStack commented at 8:16 PM on October 1, 2026: contributor

    Concept ACK

  10. sipa commented at 8:56 PM on October 1, 2026: contributor

    On Ryzen 5950X, GCC 15.2.0, SECP256K1_BENCH_ITERS=1000000 ./build/bin/bench_internal scalar, numbers for avg scalar_mul. All with default cmake -B build, except -DSECP256K1_ASM and -DSECP256K1_BUILD_BENCHMARK:

    • master, with asm: 0.0243
    • master, without asm: 0.0260
    • this PR, with asm: 0.0249
    • this PR, without asm: 0.0253
  11. real-or-random commented at 8:42 AM on October 2, 2026: contributor

    On Ryzen 5950X, GCC 15.2.0, SECP256K1_BENCH_ITERS=1000000 ./build/bin/bench_internal scalar, numbers for avg scalar_mul. All with default cmake -B build, except -DSECP256K1_ASM and -DSECP256K1_BUILD_BENCHMARK:

    * master, with asm: 0.0243
    
    * master, without asm: 0.0260
    
    * this PR, with asm: 0.0249
    
    * this PR, without asm:  0.0253

    I think these benchmarks support the PR. 0.0249 vs 0.0253 is just 2%, so it's almost noise (and note this is just scalar_mul, not a full operation).

    On the regression with asm enabled: My benchmarks above didn't compare master vs this PR when asm enabled. I was under the impression that the changes here don't affect asm builds at all. I just double-checked this and I was wrong: asm builds are affected in exactly one way, namely force inlining (all other changes affect the non-asm path.). Inlining matters only on GCC because Clang inlines anyway. (I'll edit the PR description to add this.)

    I benchmarked this on GCC 13 (where it would matter): On my machine, there's no measurable difference on P-cores, but on E-cores, this PR is actually faster than master. So I believe the regression on your AMD CPU with asm enabled is tolerable.

  12. sipa commented at 1:24 PM on October 2, 2026: contributor

    Concept ACK in any case.

  13. fanquake commented at 1:44 PM on October 2, 2026: member

    On AMD Ryzen 9 7950X3D 16-Core Processor, Clang 23:

    Ubuntu clang version 23.1.3 (++20260922084409+67f4a076a097-1~exp1~20260922084419.77)
    Target: x86_64-pc-linux-gnu
    
    Master with asm: 0.0229
    Master without asm: 0.0220
    
    PR with asm: 0.0228
    PR without asm: 0.0192 
    
  14. Yudis-bit commented at 7:43 AM on October 3, 2026: contributor

    On AMD Ryzen 3 5300U (Zen 2 mobile APU), Clang 22.1.8, SECP256K1_BENCH_ITERS=1000000 ./build/bin/bench_internal mul:

    • master, with asm: 0.0805
    • master, without asm: 0.0671
    • this PR, with asm: 0.0815
    • this PR, without asm: 0.0619
    • this PR, without asm (-march=native): 0.0597

    Pure C is ~23% faster than asm on baseline x86-64, and ~26% faster with native BMI2 (mulxq).

    Also, in secp256k1_scalar_reduce_512, c2 looks unused in practice: both SECP256K1_N_C_0 and SECP256K1_N_C_1 are < 2^63, so the accumulator peaks at ~0.52 * 2^128 and never overflows 128 bits. Switching to muladd_fast/extract_fast across reduce_512 passes tests and eliminates c2 there.


github-metadata-mirror

This is a metadata mirror of the GitHub repository bitcoin-core/secp256k1. This site is not affiliated with GitHub. Content is generated from a GitHub metadata backup.
generated: 2026-10-03 20:15 UTC

This site is hosted by @0xB10C
More mirrored repositories can be found on mirror.b10c.me