pruneassumevalid: avoid downloading witnesses and writing blocks/undos during pruned assumevalid IBD #36476

issue l0rinc opened this issue on October 8, 2026
  1. l0rinc commented at 2:45 PM on October 8, 2026: contributor

    Problem

    During pruned assumevalid, most downloaded blocks must wait their turn to connect, so the node deserializes and reserializes them to disk, only to read and deserialize them again later. The node also builds and writes undo records, even though pruning will delete most of these blocks and undo records before synchronization finishes. During assumevalid, we also download and hash witnesses to check the witness commitment, then discard them without executing their scripts or using them to build the UTXO set. We still count signature operations, even though their limits are meant to protect against expensive script checks that we already skip. With headers-first sync, we don't realistically expect reorgs during assumevalid, so we only store these blocks and undo data to recover if the node crashes mid-sync.

    This work is especially expensive on small machines with slow storage: in my testing, ordinary pruned IBD on a Raspberry Pi 4 with 2 GB of RAM has taken about 22 days, and early benchmarks of the approach presented here bring this down to about 3–4 days with the same resulting disk end-state.

    If a node rejected a block covered by assumevalid with thousands of blocks built on top of it, the most likely explanation would be a fault in its software or hardware rather than the whole network being wrong. Therefore, I would like to offer a faster way to sync for users willing to extend that trust to witness data, with a separate, default-off option that documents exactly which additional checks it skips, while full validation remains available. Full validation is not going away, this is just a faster way to get to a usable state.

    Proposed approach

    During assumevalid, skipping witness downloads, block and undo writes/serializations could make pruned IBD practical on HDD/SD cards again. Processing blocks in memory without temporary block and undo files could also lay the groundwork for combining this approach with SwiftSync and Utreexo.

    I am benchmarking prototypes of the following steps on desktops, small servers, and Raspberry Pis with different storage devices, and plan to publish the remaining proposals and fuller results in the following weeks.

    1. Prepare upcoming blocks while connecting the current one

    Similarly to parallel prevout fetching, the block loader proposed in #36000 moves computations onto worker threads, reading and deserializing upcoming blocks while the node connects the current one.

    The same workers could also run the context-independent checks, so each block is ready to connect while UTXO checks and updates still happen in chain order.

    Once block preparation runs in parallel, the next step is to avoid the block and undo writes.

    2. Build the UTXO set without writing disposable history

    Since assumevalid already skips witness script checks, -pruneassumevalid could skip downloading witnesses during that period by requesting blocks with the existing MSG_BLOCK inventory type. The block loader's workers could deserialize and check these blocks directly from memory, and the node would then use them to update the UTXO set and discard them without building undo records or writing block and undo files. The measured prototype skips the witness-dependent checks and retains all others - including sigop counting, which I will separately propose removing during assumevalid regardless of pruning status. If people feel the need to redo the utxo-independent checks later (which we have already done), we can add a separate background RCP for that.

    The usual 1,024-block download window can hold up to about 1 GB of serialized data at the maximum stripped block size of 1 MB, and only a few blocks are deserialized at a time. With validation this fast, the download window should become configurable: a 2,048-block window has already been measurably faster in some runs, and users could use more memory to download further ahead.

    To keep corner-case handling simple, the node falls back to ordinary pruning whenever the optimization cannot apply (including after assumevalid), and the block loader continues preparing blocks from disk. Indexes and wallets may need updates to support -pruneassumevalid, similarly to the changes for pruned txindex support.

    3. Resume from persisted chainstate

    A crash during normal block connection leaves the saved UTXO set intact, so the node can resume from the last completed chainstate flush and download the later blocks again. If the crash interrupts a flush between database batches or leaves a wallet needing omitted blocks, the user would need to redo IBD, which I think is a reasonable tradeoff given the speedup and how rare crashes should be. Larger write batches or a smaller -dbcache could let each flush fit in a single atomic batch, avoiding the interrupted-flush case.

    Downloading a pruned block again with getblockfrompeer doesn't restore its undo data, so #36452 makes startup verification stop at that missing data instead of reporting database corruption.

    4. Make serving blocks without witnesses cheap too

    A node already sends full blocks (with witnesses) over P2P directly from local storage, whereas sending a block without witnesses currently requires deserializing the full block and reserializing it without witnesses. A small parser could prepare the message in place by skipping the witness data along with the segwit marker and flag bytes, without constructing transaction objects or reserializing them. Prototype benchmarks showed that reading a block and preparing a witnessless message was almost 10× faster with the in-place parser.

    Early measurements

    These are full mainnet IBD runs from an empty datadir, downloading blocks from real peers and validating through height 968,869 with the v32 assumevalid hash.

    Machine Storage -dbcache Elapsed time
    Ryzen 7 3700X SSD 2,000 MiB 3h 04m 45s
    Core i9-9900K SSD 2,000 MiB 3h 21m 32s
    Umbrel, Intel N150 SSD 2,000 MiB 5h 41m 38s
    Raspberry Pi 5, 16 GB RAM SSD 2,000 MiB 9h 59m 31s

    Each machine ran the same prototype once, with -prune=10000 and -blocksonly. The prototype combines parallel block loading with skipping block and undo writes. I am also testing preparation queue sizes, worker counts, download windows, and database batch sizes before choosing defaults.

    Previous work

  2. andrewtoth commented at 12:16 PM on October 9, 2026: contributor

    Interesting proposal. The benchmarks look promising. I would be curious to see a breakdown of each step and how much speedup each represents.

    There are some ideas here that can be applied independently and not require any new trust assumptions.

    Running context-independent checks in parallel could be useful for all IBD, pruned or not, assumevalid or not.

    Not writing the block and undo files could be useful for all pruned IBD, assumevalid or not. One issue though is if a block in the best headers chain turns out to be invalid. In that case we would always want to retain the last ~288 blocks and undo we validated in memory (maybe we could have a smaller recovery window). We could persist this rollback window before every chainstate flush so we could correctly verify on startup if a crash occurs.


github-metadata-mirror

This is a metadata mirror of the GitHub repository bitcoin/bitcoin. This site is not affiliated with GitHub. Content is generated from a GitHub metadata backup.
generated: 2026-10-11 08:51 UTC

This site is hosted by @0xB10C
More mirrored repositories can be found on mirror.b10c.me