kernel: can't work with pruned blocksdir #35627

issue edilmedeiros opened this issue on July 1, 2026
  1. edilmedeiros commented at 1:20 AM on July 1, 2026: contributor

    Bitcoinkernel is not able to work datadirs created by pruned nodes (see log snippet below) Nor the api exposes a way to prune the blocksdir it manages.

    I checked that the internal pruning machinery already exists and did some experimentation to add this feature, but wanted to get some feedback about the api before working on a PR (or review if someone is willing to take this).

    A first step would be to inform the chainstate_manager about pruning by setting an option.

    BITCOINKERNEL_API int btck_chainstate_manager_options_set_prune(
        btck_ChainstateManagerOptions* chainstate_manager_options,
        uint64_t prune_mode);
    

    First question would be the semantics of the prune_mode argument. In a PoC, I followed Core's -prune cli option semantics (0 = unpruned, 1 = pruned with no target, >= 550 means pruned with a target of that many MiB). Seems simple, but feels awkward.

    I would expect the library to automatically prune the blocksdir (e.g. during btck_chainstate_manager_process_block) if the target is set. Still, I think there should be an api for triggering pruning. I thought of:

    BITCOINKERNEL_API int btck_chainstate_manager_prune(
        btck_ChainstateManager* chainstate_manager);
    
    BITCOINKERNEL_API int btck_chainstate_manager_prune_to_height(
        btck_ChainstateManager* chainstate_manager, 
        int32_t height);
    

    The former would prune to target (if set) and the latter would prune to height, no matter the target. Maybe an even better functionality for library consumers would be a function to prune specific entries, e.g.:

    BITCOINKERNEL_API int btck_chainstate_manager_prune_block_entry_entry(
        btck_ChainstateManager* chainstate_manager,
        const btck_BlockTreeEntry* block_tree_entry);
    

    Finally, it's still not clear to me how to properly test this since there are no functional tests for bitcoinkernel.

    <details><summary>Log snippet</summary> <p>

    2026-07-01T00:49:34Z Using the 'arm_shani(1way;2way)' SHA256 implementation
    2026-07-01T00:49:34Z Script verification uses 0 additional threads
    2026-07-01T00:49:34Z Using obfuscation key for blocksdir *.dat files (../datadir-pruned/signet/blocks): '6d7ae47f0b8b49e1'
    2026-07-01T00:49:34Z Opening LevelDB in ../datadir-pruned/signet/blocks/index
    2026-07-01T00:49:34Z Opened LevelDB successfully
    2026-07-01T00:49:34Z Using obfuscation key for ../datadir-pruned/signet/blocks/index: 0000000000000000
    2026-07-01T00:49:34Z Using 16 MiB out of 16 MiB requested for signature cache, able to store 524288 elements
    2026-07-01T00:49:34Z Using 16 MiB out of 16 MiB requested for script execution cache, able to store 524288 elements
    2026-07-01T00:49:34Z Assuming ancestors of block 00000008414aab61092ef93f1aacc54cf9e9f16af29ddad493b908a01ff5c329 have valid signatures.
    2026-07-01T00:49:34Z Setting nMinimumChainWork=00000000000000000000000000000000000000000000000000000b463ea0a4b8
    2026-07-01T00:49:34Z Loading block index db: last block file = 128
    2026-07-01T00:49:34Z Loading block index db: last block file info: CBlockFileInfo(blocks=2009, size=127357639, heights=309098...311171, time=2026-06-16...2026-06-30)
    2026-07-01T00:49:34Z Checking all blk files are present...
    2026-07-01T00:49:34Z Loading block index db: Block files have previously been pruned
    2026-07-01T00:49:34Z [error] Failed to load chain state from your data directory: You need to rebuild the database using -reindex to go back to unpruned mode.  This will redownload the entire blockchain
    chainstate_manager_create failed
    

    </p> </details>

  2. stickies-v commented at 2:07 PM on July 1, 2026: contributor

    I think giving the user tools to manage pruning is the way to go, and would allow them to e.g. implement prune locks (which Bitcoin Core also has) instead of having to rely on bitcoinkernel to do the tracking. Automatic pruning makes that cumbersome.

    However, I'm not sure we're currently in a good spot to offer a high quality pruning API. Block storage is not a chainstate concern, so exposing pruning through ChainstateManager does not get me excited. I think we first need to think about and work on separating block storage and chainstate management, and then revisit pruning once we have that separation? Of course, the need for pruning can be a good motivation for moving that work forward.

  3. edilmedeiros commented at 10:41 PM on July 1, 2026: contributor

    Of course, the need for pruning can be a good motivation for moving that work forward.

    I'm planning to use bitcoinkernel mostly as an (one-shot) ingestor for a database indexer, so pruning is not a concern currently. But if we want to go live, this would be desirable. Moreover, I was exploring the api with @oleonardolima in the context of having a BDK block provider based on bitcoinkernel. A wallet like this probably would benefit from heavily pruning the blocksdir.

  4. willcl-ark added the label Kernel on Jul 8, 2026
  5. musaHaruna commented at 10:03 PM on September 21, 2026: contributor

    I have been looking at and exploring this issue for some time now, and I think there are two broad questions here: what pruning capabilities the kernel should expose through its public API, and how block storage should be separated from ChainstateManager internally.

    Firstly, I think we would like to be able to open an already-pruned datadir and continue validation using its existing chainstate and available block data. Then there is manual pruning through the API, and finally automatic pruning, where the library decides when to prune toward a configured storage target while respecting its internal constraints. Another useful capability would be opening block storage and reading retained blocks without initializing the UTXO chainstate. I think that is related to this discussion, although it does not need to be a prerequisite for adding basic pruning support.

    On automatic pruning, I can see why @stickies-v describes it as cumbersome. At first it looks like the user only needs to provide a disk target and the kernel handles the rest, much like Bitcoin Core does. But an external application may still need blocks that the kernel no longer needs for validation.

    For example, the kernel could have validated up to height 900,000 while an application’s indexer has only committed its progress through 899,600. The blocks the kernel keeps for validation may not include all the blocks the indexer still needs. If pruning is triggered, it could delete a file containing blocks the indexer has not yet processed. The disk target alone does not tell the kernel that those blocks are still needed.

    I think we can reuse Bitcoin Core’s existing prune-lock machinery here, such as PruneLockInfo, BlockManager::UpdatePruneLock() and DeletePruneLock(). A public API could let applications tell the kernel which blocks they still need, update that after safely committing their progress, and remove the protection when it is no longer needed, alongside setting the pruning target.

    The target would therefore be a soft storage target. If an application still needs more history than the target allows, storage may remain above the target for as long as those blocks remain protected by a prune lock.

    This would allow applications to build their own indexes without using Bitcoin Core’s index builder. They would not need to reproduce the kernel’s pruning constants, such as MIN_BLOCKS_TO_KEEP or the internal prune-lock buffer. They would express what history they need, and the kernel would apply its own rules.

    However, I think the API needs to make the startup order clear. After a restart, the application must be able to tell the kernel which blocks it still needs before automatic pruning can begin. It should not have to create a manager and then rush to restore its prune locks before those blocks are deleted. This could mean supplying the initial prune locks when configuring the manager, or explicitly enabling automatic pruning after the application has set its locks.

    So I think the cumbersome part is defining and supporting these guarantees for external users. With manual pruning, an application can finish its work, safely commit its database and then request pruning, while the kernel still enforces its own pruning rules and respects any prune locks. This makes me think opening an already-pruned datadir and supporting manual pruning could be a reasonable first step, while the automatic-pruning interface is discussed further.

    On pruning a specific block entry, I do not think that should be part of the API under the current storage model. Bitcoin Core prunes whole block files, so pruning the file containing one block could also affect other blocks that an application or validation still needs. Only marking that individual block as unavailable would not itself reclaim the disk space.

    I think a request to prune through a height or toward a storage target would fit the existing machinery better, with the kernel deciding which whole files are eligible. The result should also make clear what actually happened, since the requested height or target may not be reached.

    On the second broad question, I think the separation should build on the boundary that already exists.BlockManageralready has its own constructor, owns the block-tree database and provides loading, lookup and read operations. Existing tests also construct it independently. But ChainstateManager currently owns its lifetime, and the public kernel API exposes storage access through chainstate management.

    I think the ownership should eventually look roughly like:

    Common owner
      owns BlockManager
      owns ChainstateManager
    ChainstateManager
      borrows BlockManager
    

    For Bitcoin Core, that common owner could be NodeContext, while the kernel may have a different owner. What’s important is that BlockManager, and the block-index entries validation refers to, remain valid for as long as those references are used.

    This would separate ownership of block storage from chainstate management, while allowing a chainstate manager to use the same storage instance for validation.

    This would also provide a clearer basis for exposing storage access independently through the public kernel API. For example, an indexer should eventually be able to open block storage, load its block-tree metadata and read retained blocks without initializing the UTXO chainstate or the rest of the validation machinery. A later API could have a shape like:

    btck_BlockStorage* btck_block_storage_open(const btck_BlockStorageOptions* options);
    btck_Block* btck_block_storage_read_block(const btck_BlockStorage* storage, const btck_BlockHash* hash);
    

    I am only exploring a possible direction here. I do not think this public API needs to be part of the first separation work.

    For pruning, I think validation could provide already-decided validation constraints to storage instead of BlockManager querying live chainstate state while choosing files. It should calculate the storage target to use and, during IBD, how much extra space pruning should leave below that target so incoming blocks do not immediately trigger another pruning event. If automatic pruning is not currently allowed, validation simply would not invoke the automatic pruning operation.

    BlockManager would then apply any prune locks to further limit which blocks can be pruned. Within those limits, it would choose eligible block and undo files, update and persist the metadata marking their data as unavailable, and then delete the files. Conceptually:

    Validation / ChainstateManager:
      determines which block heights validation can safely allow pruning
      decides whether to request automatic pruning
      calculates the storage target to use
      calculates how much extra space to leave below that target during IBD
      passes those limits to storage
    
    BlockManager:
      applies prune locks to further restrict pruning
      selects whole blk/rev file pairs within the allowed limits
      updates block-index and file metadata
      persists those metadata changes
      deletes the selected files
    

    I think this is a better boundary because changes to how validation calculates pruning limits for snapshots, IBD or multiple chainstates could stay in validation, without changing block-file selection code. Storage still receives the information it needs, but it does not need to understand why validation arrived at those limits.

    The order in which these changes are saved is also important. I think BlockManager should own the storage steps, including persisting the metadata before deleting the corresponding block and undo files. Otherwise, a crash could leave the database claiming that data is available after its file has been deleted.

    I am thinking of expressing this through a PrunePlan and a BlockManager::FlushBlockStorage() operation. Planning would select the files without changing their metadata or deleting them. The storage operation would then follow this order:

    ApplyPrunePlan(plan); // Update in-memory availability metadata.
    FlushCurrentBlockFile(request.retention.chain_tip_height);
    WriteBlockIndexDB();
    UnlinkPrunedFiles(plan.files);
    

    This would not mean moving the whole FlushStateToDisk() workflow into BlockManager. Validation would still coordinate this operation with saving the UTXO chainstate, managing its cache and sending notifications, while preserving the existing order.

    I also considered callbacks or interfaces from storage into validation. These could remove the direct dependency on the chainstate classes, but if they only expose questions such as whether the node is in IBD or whether a historical chainstate exists, storage would still be making the same validation-dependent decisions. For this operation, I think it is simpler for validation to calculate the pruning limits and pass them directly to storage.

    A separate coordinator could be useful if the pruning and persistence workflow becomes difficult to manage in the existing flush code, or if several callers need to share it. But I do not think we need to introduce one just to establish this boundary. We can keep the existing flush code coordinating the operation while giving BlockManager responsibility for the storage steps and their ordering.

    A broader block-index abstraction could also reduce direct access to CBlockIndex state. However, changing its representation or ownership would affect validation, fork handling, snapshots, persistence and the pointers held by other code. I think that can be separate work. Small helper methods may be useful during this refactor, but a broader redesign should not be required before separating storage ownership and pruning responsibilities.

    So I think the natural progression is to first move BlockManager ownership above ChainstateManager and make the lifetime relationship explicit. Then remove Chainstate and ChainstateManager dependencies from block-file selection by passing value-like pruning constraints instead. After that, make sure the storage-side pruning operation preserves the existing persistence ordering.

    Apologies for the length of this comment. I wanted to be as explicit as possible and cover my thoughts on the issue. I may have overlooked some details, and there may be a better approach than what I have suggested, I would appreciate any corrections or other perspectives.

  6. pzafonte commented at 2:42 PM on October 5, 2026: contributor

    I'm looking into adding pruning support to kernel-node and would like feedback on my approach before opening anything. (I opened one by mistake and closed it right away, sorry for the noise.) Following the suggestions above, I first separated block storage from chainstate management, then added pruning to a block manager handle in the kernel API. The work is on this branch and I'd also like to know if there's a better way to scope it for review here. Thanks!

  7. KY-U commented at 8:55 PM on October 7, 2026: contributor

    I think it is useful to distinguish the immediate compatibility problem from the broader API question:

    1. Should bitcoinkernel be able to open and continue using an already pruned datadir?
    2. Should bitcoinkernel expose an API that lets applications configure or trigger pruning?

    The first seems like a compatibility problem with the existing storage backend and may not inherently require exposing operations for managing or triggering pruning.

    The second is an API design question. I agree with @stickies-v that putting pruning operations on ChainstateManager would be inadequate as it would make core's storage management part of the validation API, exposing concerns that validation consumers ideally shouldn't need to handle.

    My concern goes a little further. Because pruning manages persisted block and undo data, rather than determining block validity, I'm inclined to think an API for triggering pruning does not naturally belong to the kernel interface. The need to open an already pruned datadir does not seem to establish by itself the need for such an API.

    The ownership change discused above by @musaHaruna, moving ownership of BlockManager out of ChainstateManager, may be valuable independently. However, I'm not sure it answers by itself whether a public pruning API should exist. Supporting an already pruned datadir, changing BlockManager ownership, and exposing pruning controls solve different problems, so I think these changes can be motivated and evaluated separately.

  8. pzafonte commented at 5:16 AM on October 8, 2026: contributor

    @KY-U thanks for suggesting how to split these out.

    The first seems like a compatibility problem with the existing storage backend and may not inherently require exposing operations for managing or triggering pruning.

    Agreed, in principle. kernel-node panics on a datadir pruned by Core, because the kernel fails at the m_have_pruned && !options.prune check in CompleteChainstateInitialization(). The kernel could instead detect the pruned flag and open the datadir in manual prune mode, like -prune=1, without any API change. I wouldn't do that though. An indexer or anything else that needs every block gets an error at startup today, but with that change it would start normally and only run into the missing blocks later. Maybe that's enough to open a pruned datadir, it may not be enough to keep it pruned safely.

    I'm inclined to think an API for triggering pruning does not naturally belong to the kernel interface.

    Though, I agree pruning shouldn't be on ChainstateManager. My branch still configures it through the chainstate manager options and gets the block manager from the chainstate manager, so I'm not proposing it as it is. I disagree with this statement overall. Some applications need a trigger or prune locks in the kernel interface, and today it has neither. A node built on the kernel can only prune through the kernel. When I deleted block files from outside, the next load failed with "Error loading block database", because the block index still says those blocks have data. Without a trigger, the only way to run pruned is for the kernel to prune on its own with a configured target.

    Pruning on its own like that isn't safe for applications that read block data after it's connected. kernel-node's wallet reads each block's undo data on a separate thread, and if it starts up behind the tip, it catches up from its saved height, which can be far behind. If pruning got there first, the wallet couldn't scan those blocks and the node would shut down. Core has the same problem with its indexes. An index that is behind catches up from disk on its own thread, and BaseIndex sets a prune lock at its best block so pruning can't get ahead of it. An application on the kernel needs the same protection. block_connected doesn't include the spent outputs, so it has to read them from disk, and it has no prune lock to set.

    In my opinion, applications would need a way to tell the kernel which blocks to keep. Either a trigger, so the application prunes only after committing its own progress (the ordering @musaHaruna described above), or prune locks, which already limit how far FlushStateToDisk() prunes. The kernel API needs at least one of them.

    I'm fine with the trigger or the locks living on a separate block storage interface instead of the validation API. But without either of them in the kernel interface, I don't think such applications would have a way to keep the blocks they still need.

  9. musaHaruna commented at 5:21 PM on October 9, 2026: contributor

    @KY-U thanks for suggesting how to split this out, and @pzafonte thanks for explaining the wallet use case and why applications need a way to protect the block and undo data they still need.

    Supporting an already pruned datadir, changing BlockManager ownership, and exposing pruning controls solve different problems, so I think these changes can be motivated and evaluated separately.

    I agree with this distinction. Supporting an already-pruned datadir can be a separate compatibility change and proceed independently of the ownership refactor. The broader pruning use case also motivates separating storage from chainstate management, giving us a clearer foundation for exposing pruning controls through a separate block-storage API.

    An indexer or anything else that needs every block gets an error at startup today, but with that change it would start normally and only run into the missing blocks later.

    I agree that silently accepting a pruned datadir could hide this mismatch. Could accepting pruned storage be an explicit choice, while preserving the startup failure for callers that have not opted in?

    I think we could keep the compatibility change focused by clearly documenting what that choice means: the kernel can open the datadir, but this does not restore deleted history or guarantee that the remaining data is sufficient for the application. An application building an index that requires missing history would need to obtain or reconstruct that data, or use an unpruned source. This depends on the indexer's requirements; downloading raw blocks alone may also be insufficient if it needs undo data.

    I do not think the compatibility change needs to implement that recovery or determine every indexer's requirements. Explicit acceptance, clear documentation, and clear errors when requested data is unavailable could give us a useful first step, leaving recovery to the application.

    Either a trigger, so the application prunes only after committing its own progress [...] or prune locks, which already limit how far FlushStateToDisk() prunes.

    I agree that applications need one of these mechanisms to protect the history they still require.

    Following @KY-U’s suggested split, the compatibility work could proceed alongside the internal separation work. The refactor would move BlockManager ownership out of ChainstateManager and separate storage responsibilities from validation while preserving the existing safety constraints and persistence ordering.

    Once that boundary is established, my preference would be to expose a small manual pruning API through a separate block-storage interface. Applications could process the required history, durably commit their progress, and request pruning through a safe height. The application would ensure that the requested pruning does not remove history still needed by its wallets, indexers, or other components, while storage would enforce validation constraints and report what was actually pruned. Automatic pruning would remain disabled for this workflow.

    We could then discuss automatic pruning and public prune locks, including how applications establish their retention requirements before automatic pruning begins after a restart. @stickies-v, @KY-U, @pzafonte, what do you think about an explicit opt-in for opening already-pruned datadirs as a separate compatibility patch? That patch could proceed independently of the storage refactor, with manual pruning controls considered afterward. @pzafonte, I will take a closer look at the refactor on your branch. If there is agreement on the compatibility approach, I would be happy to start working on that patch.


github-metadata-mirror

This is a metadata mirror of the GitHub repository bitcoin/bitcoin. This site is not affiliated with GitHub. Content is generated from a GitHub metadata backup.
generated: 2026-10-11 12:51 UTC

This site is hosted by @0xB10C
More mirrored repositories can be found on mirror.b10c.me