With recent Clang versions, the C implementation of the 4x64 scalar multiplication is about as fast as the x86_64 assembly (#1682, #1945). With GCC, it is considerably slower: in the benchmark below, it is 17–20% slower than the assembly on a P-core. GCC compiles the carry handling of the accumulator macros into code that saves carries in registers (setc) and adds them separately, instead of emitting add-with-carry (adc) chains.
This PR changes the C code so that it is about as fast as the assembly with GCC 14 and newer and faster with Clang, and then disables the assembly by default:
- scalar: Use a 128-bit accumulator in 4x64 C code. The lower 128 bits of the 192-bit accumulator
(c0,c1,c2)are kept in asecp256k1_uint128, and carries are computed by two new int128 functions,secp256k1_u128_accum_mul_carryandsecp256k1_u128_accum_u64_carry. The native implementation uses__builtin_add_overflowon 128-bit values, for which GCC 14 and newer emitadcchains. Without the builtin, and in the struct implementation, the carries are computed on 64-bit halves as before. - scalar: Force inlining of 4x64 mul_512 and reduce_512. Without inlining, the speed of GCC's code depends on incidental details of the source, such as whether an expression is written directly or comes from an inline function: this changes the operand order of the 128-bit additions in GCC's intermediate representation.[^1] With inlining, GCC generates the same code independently of these details, and Clang can schedule across the boundary between multiplication and reduction.
- build: Disable x86_64 assembly by default. The assembly remains available via
--with-asm=x86_64(Autotools) and-DSECP256K1_ASM=x86_64(CMake), or viaauto. Since the field assembly was removed earlier, the option only affects scalar multiplication.
GCC 13 is still somewhat common, e.g., it's in Ubuntu LTS 24.04. Once it's no longer widespread, we could remove the assembly entirely.
Benchmarks
bench_internal scalar_mul measures the latency of a chain of dependent scalar multiplications. The setup was an Intel i7-1260P with turbo disabled, with the process pinned to a P-core (Golden Cove) or an E-core (Gracemont). The tables show the median time in µs over 15 interleaved runs. Percentages are relative to the assembly built with the same compiler.
P-core:
| Compiler | asm | C (master) | C (this PR) |
|---|---|---|---|
| GCC 13.5 | 0.0549 | 0.0661 (+20%) | 0.0650 (+18%) |
| GCC 14.4 | 0.0548 | 0.0654 (+19%) | 0.0551 (+1%) |
| GCC 16.2 | 0.0541 | 0.0634 (+17%) | 0.0556 (+3%) |
| Clang 18.1 | 0.0558 | 0.0563 (+1%) | 0.0459 (−18%) |
| Clang 22.1 | 0.0531 | 0.0573 (+8%) | 0.0474 (−11%) |
E-core:
| Compiler | asm | C (master) | C (this PR) |
|---|---|---|---|
| GCC 13.5 | 0.1050 | 0.0991 (−6%) | 0.1160 (+10%) |
| GCC 14.4 | 0.1050 | 0.0999 (−5%) | 0.1010 (−4%) |
| GCC 16.2 | 0.1050 | 0.0988 (−6%) | 0.1020 (−3%) |
| Clang 18.1 | 0.1070 | 0.0934 (−13%) | 0.0873 (−18%) |
| Clang 22.1 | 0.1030 | 0.0929 (−10%) | 0.0885 (−14%) |
With GCC 14 and newer, the C code in this PR is within 3% of the assembly on the P-core and faster than it on the E-core. With Clang 18 and 22, it is 11–18% faster than the assembly on both core types.
GCC 13 does not benefit from this PR, because it lacks GCC 14's improvements in compiling add-with-carry. On the P-core, the C code is still 18% slower than the assembly. On the E-core, it is 17% slower than the C code on master. The CHANGELOG entry notes that the assembly may still be noticeably faster with GCC 13 or older. Bitcoin Core's release binaries are built with GCC 14 for Linux and Windows and with Clang 19 for macOS.
Builds with asm enabled are affected only through the force inline, which is only effective on GCC because Clang inlines even without it.
Builds with the int128 struct implementation (e.g., MSVC) use the same carry computation as before, and their speed is unchanged.
Call for benchmarks
Benchmarks on other CPUs, in particular AMD, would be welcome.
[^1]: Claude Opus 5.5 claims that it found three GCC missed-optimization bugs here. I may or may not report them upstream, but this shows why other projects such as fiat-crypto, OpenSSL, ... either get a performance hit or, like us, resorted to inline asm to convince GCC to output proper adc chains. While there's proper pattern recognition for corresponding plain C code, some later optimization passes apparently kill its finding, and the resulting binary still doesn't use proper adc chains. Happy to share the details upfront here if you're interested.