Clang beats handwritten `x86_64` assembly for `scalar_mul` #1682

issue hebasto opened this issue on June 5, 2025
  1. hebasto commented at 2:18 PM on June 5, 2025: member

    On Ubuntu 25.04:

    $ ./build_clang17_0_6_asm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0319    ,     0.0345    ,     0.0569 
    field_mul                     ,     0.0173    ,     0.0173    ,     0.0175
    $ ./build_clang17_0_6_noasm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0266    ,     0.0268    ,     0.0275 
    field_mul                     ,     0.0173    ,     0.0174    ,     0.0177
    
    $ ./build_clang18_1_8_asm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0336    ,     0.0345    ,     0.0357 
    field_mul                     ,     0.0170    ,     0.0171    ,     0.0172
    $ ./build_clang18_1_8_noasm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0266    ,     0.0267    ,     0.0271 
    field_mul                     ,     0.0170    ,     0.0171    ,     0.0174
    
    $ ./build_clang19_1_7_asm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0329    ,     0.0330    ,     0.0331 
    field_mul                     ,     0.0166    ,     0.0167    ,     0.0171 
    $ ./build_clang19_1_7_noasm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0270    ,     0.0270    ,     0.0271 
    field_mul                     ,     0.0167    ,     0.0167    ,     0.0169 
    
    $ ./build_clang20_1_2_asm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0330    ,     0.0343    ,     0.0447 
    field_mul                     ,     0.0164    ,     0.0170    ,     0.0214 
    $ ./build_clang20_1_2_noasm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0269    ,     0.0270    ,     0.0271 
    field_mul                     ,     0.0165    ,     0.0165    ,     0.0166 
    
  2. sipa commented at 2:19 PM on June 5, 2025: contributor

    Can you post benchmarks with GCC on the same machine?

  3. hebasto commented at 2:27 PM on June 5, 2025: member

    Can you post benchmarks with GCC on the same machine?

    Sure!

    $ ./build_gcc11_5_0_asm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0289    ,     0.0291    ,     0.0301 
    field_mul                     ,     0.0151    ,     0.0152    ,     0.0154 
    $ ./build_gcc11_5_0_noasm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0313    ,     0.0314    ,     0.0316 
    field_mul                     ,     0.0151    ,     0.0151    ,     0.0153 
    
    $ ./build_gcc15_0_1_asm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0292    ,     0.0293    ,     0.0295 
    field_mul                     ,     0.0170    ,     0.0171    ,     0.0172
    $ ./build_gcc15_0_1_noasm/bin/bench_internal mul
    Benchmark                     ,    Min(us)    ,    Avg(us)    ,    Max(us)    
    
    scalar_mul                    ,     0.0327    ,     0.0328    ,     0.0338 
    field_mul                     ,     0.0171    ,     0.0171    ,     0.0172
    
  4. real-or-random added the label performance on Jun 5, 2025
  5. real-or-random commented at 8:26 PM on June 5, 2025: contributor

    Fwiw, I've hacked together a compiler explorer instance where you can compare clang's output (inline asm disabled) on different versions. I don't see a big change, it's probably just a lot of incremental improvements that add up.

  6. real-or-random commented at 7:35 AM on September 29, 2026: contributor

    Am I right that there's no "regression" in the normal meaning of the word here (i.e., something which was fast on previous library or compiler versions became slower)? It's just that the Clang's code became faster, and it's faster than the ASM. (Let me edit the title.)

  7. real-or-random renamed this:
    Performance regression in `scalar_mul` with Clang and `x86_64` assembly enabled
    Clang beats handwritten `x86_64` assembly for `scalar_mul`
    on Sep 29, 2026

github-metadata-mirror

This is a metadata mirror of the GitHub repository bitcoin-core/secp256k1. This site is not affiliated with GitHub. Content is generated from a GitHub metadata backup.
generated: 2026-09-29 22:15 UTC

This site is hosted by @0xB10C
More mirrored repositories can be found on mirror.b10c.me