Problem
During pruned assumevalid, most downloaded blocks must wait their turn to connect, so the node deserializes and reserializes them to disk, only to read and deserialize them again later. The node also builds and writes undo records, even though pruning will delete most of these blocks and undo records before synchronization finishes. During assumevalid, we also download and hash witnesses to check the witness commitment, then discard them without executing their scripts or using them to build the UTXO set. We still count signature operations, even though their limits are meant to protect against expensive script checks that we already skip. With headers-first sync, we don't realistically expect reorgs during assumevalid, so we only store these blocks and undo data to recover if the node crashes mid-sync.
This work is especially expensive on small machines with slow storage: in my testing, ordinary pruned IBD on a Raspberry Pi 4 with 2 GB of RAM has taken about 22 days, and early benchmarks of the approach presented here bring this down to about 3–4 days with the same resulting disk end-state.
If a node rejected a block covered by assumevalid with thousands of blocks built on top of it, the most likely explanation would be a fault in its software or hardware rather than the whole network being wrong. Therefore, I would like to offer a faster way to sync for users willing to extend that trust to witness data, with a separate, default-off option that documents exactly which additional checks it skips, while full validation remains available. Full validation is not going away, this is just a faster way to get to a usable state.
Proposed approach
During assumevalid, skipping witness downloads, block and undo writes/serializations could make pruned IBD practical on HDD/SD cards again. Processing blocks in memory without temporary block and undo files could also lay the groundwork for combining this approach with SwiftSync and Utreexo.
I am benchmarking prototypes of the following steps on desktops, small servers, and Raspberry Pis with different storage devices, and plan to publish the remaining proposals and fuller results in the following weeks.
1. Prepare upcoming blocks while connecting the current one
Similarly to parallel prevout fetching, the block loader proposed in #36000 moves computations onto worker threads, reading and deserializing upcoming blocks while the node connects the current one.
The same workers could also run the context-independent checks, so each block is ready to connect while UTXO checks and updates still happen in chain order.
Once block preparation runs in parallel, the next step is to avoid the block and undo writes.
2. Build the UTXO set without writing disposable history
Since assumevalid already skips witness script checks, -pruneassumevalid could skip downloading witnesses during that period by requesting blocks with the existing MSG_BLOCK inventory type.
The block loader's workers could deserialize and check these blocks directly from memory, and the node would then use them to update the UTXO set and discard them without building undo records or writing block and undo files.
The measured prototype skips the witness-dependent checks and retains all others - including sigop counting, which I will separately propose removing during assumevalid regardless of pruning status. If people feel the need to redo the utxo-independent checks later (which we have already done), we can add a separate background RCP for that.
The usual 1,024-block download window can hold up to about 1 GB of serialized data at the maximum stripped block size of 1 MB, and only a few blocks are deserialized at a time. With validation this fast, the download window should become configurable: a 2,048-block window has already been measurably faster in some runs, and users could use more memory to download further ahead.
To keep corner-case handling simple, the node falls back to ordinary pruning whenever the optimization cannot apply (including after assumevalid), and the block loader continues preparing blocks from disk.
Indexes and wallets may need updates to support -pruneassumevalid, similarly to the changes for pruned txindex support.
3. Resume from persisted chainstate
A crash during normal block connection leaves the saved UTXO set intact, so the node can resume from the last completed chainstate flush and download the later blocks again.
If the crash interrupts a flush between database batches or leaves a wallet needing omitted blocks, the user would need to redo IBD, which I think is a reasonable tradeoff given the speedup and how rare crashes should be.
Larger write batches or a smaller -dbcache could let each flush fit in a single atomic batch, avoiding the interrupted-flush case.
Downloading a pruned block again with getblockfrompeer doesn't restore its undo data, so #36452 makes startup verification stop at that missing data instead of reporting database corruption.
4. Make serving blocks without witnesses cheap too
A node already sends full blocks (with witnesses) over P2P directly from local storage, whereas sending a block without witnesses currently requires deserializing the full block and reserializing it without witnesses. A small parser could prepare the message in place by skipping the witness data along with the segwit marker and flag bytes, without constructing transaction objects or reserializing them. Prototype benchmarks showed that reading a block and preparing a witnessless message was almost 10× faster with the in-place parser.
Early measurements
These are full mainnet IBD runs from an empty datadir, downloading blocks from real peers and validating through height 968,869 with the v32 assumevalid hash.
| Machine | Storage | -dbcache |
Elapsed time |
|---|---|---|---|
| Ryzen 7 3700X | SSD | 2,000 MiB | 3h 04m 45s |
| Core i9-9900K | SSD | 2,000 MiB | 3h 21m 32s |
| Umbrel, Intel N150 | SSD | 2,000 MiB | 5h 41m 38s |
| Raspberry Pi 5, 16 GB RAM | SSD | 2,000 MiB | 9h 59m 31s |
Each machine ran the same prototype once, with -prune=10000 and -blocksonly. The prototype combines parallel block loading with skipping block and undo writes.
I am also testing preparation queue sizes, worker counts, download windows, and database batch sizes before choosing defaults.
Previous work
- #9484: Introduce assumevalid setting to skip validation presumed valid scripts introduced skipping historical script checks based on a configured block hash and chain work conditions.
- #13151: net: Serve blocks directly from disk when possible introduced sending the stored bytes directly for witness-inclusive requests. #13098: Skip tx-rehashing on historic blocks was an earlier, closed proposal to skip transaction hashing when serving historical blocks.
- #27050: p2p, validation: Don't download witnesses for assumed-valid blocks when running in prune mode is the closest earlier proposal. It also requested blocks without witnesses during pruned assumevalid, but stored them with a separate witness-pruned status alongside ordinary undo data. This proposal writes neither blocks nor undo data during that period, avoiding the extra state and the corner cases around serving stored stripped blocks or reading them after a restart with different settings. There are no stripped block files to serve, and after assumevalid the node uses ordinary pruning. The Review Club discussion covers the validation and trust implications.
- #35295: validation: fetch block input prevouts in parallel during ConnectBlock fetches a block's input prevouts from chainstate on worker threads and was merged as a follow-up to the earlier, closed #31132.
- #35662: script: prevent partial reinitialization of sighash data also stops allocating unused signature-check data when assumevalid skips scripts.
- #36000: validation: prefetch blocks while connecting is the pending block-loader proposal described above.
- #36002: txindex: allow running in pruned mode lets lookups report missing blocks so users can fetch them again on demand. Its merged prerequisite, #35531, reduced the index size and replaced stored block-file positions with block references, so a transaction can still be located after its block is downloaded again.
- #36244: validation, net: Process blocks asynchronously and reduce cs_main contention is related work on separating network handling from block connection, a different scheduling layer from block preparation.
- #36452: validation: fix false corruption errors after refetching pruned blocks handles verification when block and undo availability differ.