txindex: allow running in pruned mode #36002

pull andrewtoth wants to merge 7 commits into bitcoin:master from andrewtoth:txindex-prune changing 30 files +420 −93
  1. andrewtoth commented at 12:01 AM on August 18, 2026: contributor

    After #35531, each txindex entry now points to the block hash containing the tx instead of the file location of the block. The block hash is then used to find the file position of the block to read the tx from. This allows us to determine the hash of the containing block even if the block is later pruned, and also determine the location of the block on disk if it is later downloaded again via getblockfrompeer.

    This PR enables maintaining a txindex while pruned. If a txid lookup tries to read a tx from a pruned block from after the index was started, the hash of the missing block is surfaced to the caller. Otherwise, a warning is returned that the tx might be in a pruned block that was not indexed.

    If desired, a user may call getblockfrompeer using the returned hash to download the block, then try again successfully. The downloaded block will be added to the front of the block queue, so won't be pruned again for at least 288 blocks. Of course a larger prune setting is desired if lots of historical blocks will be read, and there are privacy issues if the user is looking up their own transactions. Ideally the blocks can be fetched from a trusted node or from random peers with decoy blocks fetched occasionally.

  2. DrahtBot commented at 12:01 AM on August 18, 2026: contributor

    <!--e57a25ab6845829454e8d69fc972939a-->

    The following sections might be updated with supplementary metadata relevant to reviewers and maintainers.

    <!--006a51241073e994b41acfe9ec718e94-->

    Code Coverage & Benchmarks

    For details see: https://corecheck.dev/bitcoin/bitcoin/pulls/36002.

    <!--021abf342d371248e50ceaed478a90ca-->

    Reviews

    See the guideline and AI policy for information on the review process.

    Type Reviewers
    Concept ACK l0rinc, ajtowns, arejula27

    If your review is incorrectly listed, please copy-paste <code>&lt;!--meta-tag:bot-skip--&gt;</code> into the comment that the bot should ignore.

    <!--174a7506f384e20aa4161008e828411d-->

    Conflicts

    Reviewers, this pull request conflicts with the following ones:

    • #35474 (node: move index ownership to NodeContext by w0xlt)
    • #34729 (Reduce log noise by ajtowns)
    • #34132 (coins, dbwrapper: remove error catcher, make point-read failures fatal by l0rinc)

    If you consider this pull request important, please also help to review the conflicting pull requests. Ideally, start with the one that should be merged first.

    <!--5faf32d7da4f0f540f40219e4f7537a3-->

    LLM Linter (✨ experimental)

    Possible places where named args for integral literals may be used (e.g. func(x, /*named_arg=*/0) in C++, and func(x, named_arg=0) in Python):

    • std::make_unique<TxIndex>(interfaces::MakeChain(node), index_cache_sizes.tx_index, false, do_reindex) in src/init.cpp

    <sup>2026-09-18 03:54:15</sup>

  3. andrewtoth force-pushed on Aug 18, 2026
  4. DrahtBot added the label CI failed on Aug 18, 2026
  5. DrahtBot commented at 12:14 AM on August 18, 2026: contributor

    <!--85328a0da195eb286784d51f73fa0af9-->

    🚧 At least one of the CI tasks failed. <sub>Task iwyu: https://github.com/bitcoin/bitcoin/actions/runs/32082864144/job/95549090994</sub> <sub>LLM reason (✨ experimental): CI failed because IWYU reported an include issue (“Failure generated from IWYU”) in src/index/txindex.cpp.</sub>

    <details><summary>Hints</summary>

    Try to run the tests locally, according to the documentation. However, a CI failure may still happen due to a number of reasons, for example:

    • Possibly due to a silent merge conflict (the changes in this pull request being incompatible with the current code in the target branch). If so, make sure to rebase on the latest commit of the target branch.

    • A sanitizer issue, which can only be found by compiling with the sanitizer and running the affected test.

    • An intermittent issue.

    Leave a comment here, if you need help tracking down a confusing failure.

    </details>

  6. l0rinc commented at 12:20 AM on August 18, 2026: contributor

    Big concept ACK \:D/

  7. DrahtBot removed the label CI failed on Aug 18, 2026
  8. ajtowns commented at 5:24 AM on August 25, 2026: contributor

    This allows us to determine the hash of the containing block even if the block is later pruned,

    Hmm, I think that means the txindex starts off at ~30GB and grows at ~2.5GB/year indefinitely, so in 40 years' time might be 130GB. I guess that makes sense for nodes wanting to run a local block-explorer type setup; particularly if they keep a significant range of real blocks around; I think expected block usage is about 100GB per year, so 5 years worth of history would be half a terabyte, and probably cover most lookups.

    I'd definitely prefer to have a -prunetxindex=1 option I could set, so that my prune=2000 (~1 week worth of blocks) node could lookup txs without needing to be reprovisioned with extra disk though. (With such an option enabled, should be feasible to setup a txindex on an already pruned node without redownloading blocks, I think)

    Concept ACK, looks okay to me at first glance.

    What's the (intended) behaviour if you enable txindex on a node that is already pruned? I think it will just hit FatalErrorf in BaseIndex::ProcessBlock and shut the node down?

  9. andrewtoth commented at 8:03 PM on August 29, 2026: contributor

    What's the (intended) behaviour if you enable txindex on a node that is already pruned?

    It's the same behavior if you try and start coinstatsindex or blockfilterindex after you've pruned past their last indexed block - the startup error: Error: <index> best block of the index goes beyond pruned data (including undo data). Please disable the index or reindex (which will download the whole blockchain again)

    I'd definitely prefer to have a -prunetxindex=1 option I could set

    I think that could be added as a follow-up, and would not conflict with this approach. We could just delete entries by sequence until the last sequence maps to a block we have not yet pruned. We would have to compact the db afterwards though. That way though you will just get a "tx not found" response for a historical lookup where the index has also been pruned past, instead of the block hash.

  10. mzumsande commented at 3:04 PM on September 7, 2026: contributor

    I wonder if the current design would be more helpful for users - or a genuine partial index that would only report on data of the non-pruned range and clean up old entries every now and then.

    The alternative design would have the downside that it wouldn't report the possible blocks pruned txns might be in. Plus, it would need changes to baseindex, so more complicated.

    However, it would have the big upside that it could be enabled on an already pruned node.

  11. andrewtoth commented at 3:28 PM on September 7, 2026: contributor

    should be feasible to setup a txindex on an already pruned node without redownloading blocks

    it would need changes to baseindex, so more complicated.

    I think we can come up with something that would be interesting for both use cases, but without the complexity if we removed the "cleanup" part.

    If we start with an empty index and are pruned, we set the locator to the last pruned block we have and sync. This way we get all the blocks indexed that we have, and future blocks that are pruned we can still return the block hash. When we want to "cleanup", we can just delete the txindex folder and restart.

    the downside that it wouldn't report the possible blocks pruned txns might be in

    Indeed, this has the downside that some existing txs will return not found, and some will return the pruned block hash. So it introduces some ambiguity into the result if you create the index while already pruned. It also does not bound the txindex disk size, but that has already been reduced significantly via #35531.

  12. andrewtoth force-pushed on Sep 9, 2026
  13. andrewtoth commented at 12:34 AM on September 9, 2026: contributor

    Thanks @ajtowns @mzumsande for your thoughts.

    I've updated this change to also allow starting an index with an already pruned node. For any tx lookup miss where we don't get a block hash and started indexing after genesis, the response returns context in the error message that the tx may have been in a block that we never indexed.

    We also return the starting block of the index in getindexinfo.

    Regarding cleaning up entries in the txindex that are pruned, I don't think the complexity of this is worth it. If you are running pruned, you need at least 11 GB, and the txindex grows at about 2.5 GB per year. So every few months you can just nuke your index directory and rebuild. If you're running with a very low prune target, this shouldn't take more than a few minutes max (likely under a minute for -prune=2000).

    Wdyt?

  14. DrahtBot added the label Needs rebase on Sep 9, 2026
  15. andrewtoth force-pushed on Sep 10, 2026
  16. andrewtoth commented at 1:50 AM on September 10, 2026: contributor

    Rebased due to include conflicts with #36116.

  17. andrewtoth force-pushed on Sep 10, 2026
  18. DrahtBot added the label CI failed on Sep 10, 2026
  19. DrahtBot commented at 2:41 AM on September 10, 2026: contributor

    <!--85328a0da195eb286784d51f73fa0af9-->

    🚧 At least one of the CI tasks failed. <sub>Task iwyu: https://github.com/bitcoin/bitcoin/actions/runs/34426977165/job/102714188149</sub> <sub>LLM reason (✨ experimental): CI failed because IWYU reported include/header issues (generated “Failure generated from IWYU” and exited non-zero after applying include changes).</sub>

    <details><summary>Hints</summary>

    Try to run the tests locally, according to the documentation. However, a CI failure may still happen due to a number of reasons, for example:

    • Possibly due to a silent merge conflict (the changes in this pull request being incompatible with the current code in the target branch). If so, make sure to rebase on the latest commit of the target branch.

    • A sanitizer issue, which can only be found by compiling with the sanitizer and running the affected test.

    • An intermittent issue.

    Leave a comment here, if you need help tracking down a confusing failure.

    </details>

  20. DrahtBot removed the label Needs rebase on Sep 10, 2026
  21. DrahtBot removed the label CI failed on Sep 10, 2026
  22. arejula27 commented at 1:59 PM on September 13, 2026: contributor

    Concept ACK.

    I'm excited about this PR and I love the idea and the approach.

    I have to go deeper into the code (just did a first review to see the idea), but meanwhile I have some questions.

    BaseIndex now has two prune-related flags, AllowPrune() and AllowPartialHistory(), and they aren't really independent. Partial history is only ever set up on a pruned node, so (as I understand it) an index with AllowPartialHistory() but not AllowPrune() can't work. Would a single enum value express that relationship better, so the invariant lives in the type instead of a convention that a future index could break? Something like an AllowPrune enum with values Disallowed, FullHistory and PartialHistory.

    Another option would be to move AllowPartialHistory() to TxIndex instead. As far as I can understand no other current index could set it to true anyway. Maybe your idea was for future ones?

    Finally, I think one big addition with this is that we can even access txs on pruned nodes, since #35531 the index stores the block "virtual" reference rather than a file position , so the node knows exactly which block is missing. From there the caller can fetch that block with getblockfrompeer and retry the lookup. Right now those hashes only reach the caller inside the error string, so automating the fetch and retry means parsing the message. Would returning them as structured data be worth it? Or even an optional flag on getrawtransaction that requests the block for us in the same rpc call?

  23. txindex: allow prune mode when there are no legacy entries
    Replace AllowPrune()'s bool with an IndexPrunePolicy enum.
    4caa72b094
  24. txindex: allow creating index from an already pruned node
    Add IndexPrunePolicy::PartialHistory so a fresh txindex can start from the
    earliest unpruned block
    ae29aa404d
  25. rpc: return index starting height in getindexinfo
    Expose first_block_height so callers can see where a partial index begins.
    4385da9a45
  26. txindex: return pruned block hashes from FindTx
    FindTx now returns a TxLookupResult so callers can distinguish a
    missing transaction from one that may only be in pruned blocks.
    8301a14ab1
  27. rpc: report pruned blocks from GetTransaction
    getrawtransaction, gettxoutproof, and REST /tx now return potential pruned block hashes in the error message if the transaction is not found.
    RPC errors also include those hashes in as structured data in the json error.data field.
    getrawtransaction verbosity 2 omits fee and prevout when undo
    data is unavailable instead of throwing.
    
    A caller can call getblockfrompeer with the candidate block hashes and try again.
    cb4f95eddd
  28. rpc,rest: return tx miss error text about unknown pruned blocks
    Return extra context in the error message when we miss fetching a tx via the txindex.
    If we are still syncing, or if we only have a partial index and the tx could be in a block we don't know about.
    adc3480e4a
  29. andrewtoth force-pushed on Sep 18, 2026
  30. andrewtoth force-pushed on Sep 18, 2026
  31. DrahtBot added the label CI failed on Sep 18, 2026
  32. andrewtoth commented at 3:39 AM on September 18, 2026: contributor

    Thanks @arejula27. Those are good suggestions.

    I renamed AllowPrune to GetPrunePolicy which returns an enum.

    Another option would be to move AllowPartialHistory() to TxIndex instead. As far as I can understand no other current index could set it to true anyway. Maybe your idea was for future ones?

    Yes it is for now, but if these changes are accepted we could do the same exercise for txospenderindex.

    I added returning the pruned block hashes in an optional data field of the JSON error response. I kept the REST error response as a string. All REST errors are strings, and I didn't want to introduce a special case. That would be a more invasive change, since all existing clients likely always expect plain/text string error responses. Of course that can be done as a follow-up if there is interest.

  33. test: cover txindex lookups of pruned blocks
    Extend the prune+txindex functional test to check that
    getrawtransaction, gettxoutproof, and REST /tx return the pruned block hash,
    and that utxoupdatepsbt skips those previous transactions until the
    block is fetched.
    
    Also test partial indexes when starting from a pruned node, and the text responses when we don't know the pruned block hash.
    94a7beb9cf
  34. andrewtoth force-pushed on Sep 18, 2026
  35. DrahtBot removed the label CI failed on Sep 18, 2026
  36. arejula27 commented at 7:24 AM on September 18, 2026: contributor

    Thx! Will try the exercise! I am finishing the (deeper) code review, hope to post it soon :D

  37. in src/index/tx_lookup_result.h:22 in 94a7beb9cf
      17 | +    /// Hash of the block containing the transaction. null if the transaction
      18 | +    /// was found in the mempool or was not found.
      19 | +    uint256 block_hash;
      20 | +    /// Only set when the transaction was not found but may exist in a pruned block.
      21 | +    /// Note that these blocks may be false positives and not contain the transaction.
      22 | +    std::vector<uint256> pruned_block_hashes;
    


    arejula27 commented at 10:00 AM on September 24, 2026:

    This struct represents two possible outcomes, the tx was found or it wasn't, but the type doesn't keep its fields consistent: it can hold a tx together with pruned_block_hashes, or a block_hash without a tx. My proposal is to use std::variant<TxFound, TxMiss> rather than one struct where the outcome is implied by which fields are set.

    For example, wallet::TxState in src/wallet/transaction.h models a transaction's state the same way, as a variant of structs.

    I suggest something like:

    /// A transaction found in the mempool or in a block.
    struct TxFound {
        CTransactionRef tx;
        /// Null if the transaction was found in the mempool.
        uint256 block_hash;
    };
    
    /// A lookup that did not find the transaction.
    struct TxMiss {
        /// Blocks that may contain the transaction but were pruned. They may be false positives.
        std::vector<uint256> pruned_block_hashes;
    };
    

    arejula27 commented at 9:45 PM on September 24, 2026:

    I would personally prefer use util::expected, but i think the project approach is to use it only over real errors (like BlockReadFailed kind of errors)

  38. in src/test/txindex_tests.cpp:365 in 94a7beb9cf
     361 | @@ -352,4 +362,28 @@ BOOST_FIXTURE_TEST_CASE(txindex_reorg_keeps_stale_entries, TestChain100Setup)
     362 |      txindex.Stop();
     363 |  }
     364 |  
     365 | +BOOST_FIXTURE_TEST_CASE(txindex_pruned_lookup, TestChain100Setup)
    


    arejula27 commented at 10:04 AM on September 24, 2026:

    I would add two more scenarios, as I have tested them and they are not covered:

    • A pruned block outside the active chain must not be reported (dropping the if (in_active_chain) condition leaves txindex_tests and feature_txindex_prune.py passing). The test mines a tx, invalidates its block so it becomes stale, prunes that block and checks the lookup returns no candidates.

    • Several pruned candidates must all be returned, which is what makes pruned_block_hashes a vector (keeping only the first hash also leaves both suites passing). The test forges an entry with a colliding prefix pointing at another block, the way txindex_collision_scan_path does, prunes both blocks and checks that both hashes come back.

    I checked both cases with mutation testing, and can share the diff if it helps.

  39. in src/node/transaction.cpp:158 in 94a7beb9cf
     159 | -                hashBlock = result->block_hash;
     160 | -                return result->tx;
     161 | +                return result;
     162 |              }
     163 | +        } else if (!block_index) {
     164 | +            return TxLookupResult{std::move(result.pruned_block_hashes)};
    


    arejula27 commented at 10:09 AM on September 24, 2026:

    nit: result has no tx and no block_hash in this branch, so this builds the same object again. Is there a reason not to return result; directly?

  40. in src/rpc/txoutproof.cpp:112 in 94a7beb9cf
     110 | +                    throw JSONRPCError(RPC_MISC_ERROR, PrunedBlocksErrorMessage(result.pruned_block_hashes),
     111 | +                                       PrunedBlocksErrorData(result.pruned_block_hashes));
     112 |                  }
     113 | +                if (!result.tx || result.block_hash.IsNull()) {
     114 | +                    std::string message{"Transaction not yet in block"};
     115 | +                    if (g_txindex) message += TxIndexMissErrorDetails(g_txindex->GetSummary());
    


    arejula27 commented at 10:10 AM on September 24, 2026:

    nit: with a partial index the full error reads "Transaction not yet in block. The transaction may be in an earlier block not covered by this index, which starts at height N". "Not yet in block" suggests the tx is unconfirmed, while the added sentence suggests it may be confirmed in an old block. It sounds weird. Maybe use a neutral base message here too when the partial history detail is appended, e.g. "Transaction not found"?

  41. arejula27 commented at 11:01 AM on September 24, 2026: contributor

    Thanks for applying the changes I asked for in the first review. I've done a deeper pass now, built it and ran txindex_tests and feature_txindex_prune.py, and have a few more notes.

    gettxoutproof only reports the pruned block hash when the tx is fully spent: with an unspent output the block comes from the UTXO set and the error is still "Block not available (pruned data)", even though pblockindex is right there. Would it make sense to report that hash too?

    When an existing txindex falls behind pruned data, the error still offers only "disable the index or reindex (which will download the whole blockchain again)", while deleting the index directory now gives a working partial index without downloading anything. Would it be worth mentioning that for indexes that allow partial history? Something like "%s best block of the index goes beyond pruned data. Please disable the index, delete its directory to rebuild it from the earliest block still on disk, or reindex (which will download the whole blockchain again)".

    On partial history ( which i really like the idea), some users prune only to save block storage but could afford a full txindex (around 30 GB plus ~2.5 GB per year, as mentioned earlier in this PR). Today they would need a full reindex to get the whole index (no partial), which on a pruned node means downloading and validating the chain again. Would a follow-up similar to the background chainstate in assumeutxo make sense? Serving lookups from the partial index while a background task fetches the missing historical blocks, indexes them and discards them, moving first_block_height down until it reaches genesis.

  42. in src/rpc/node.cpp:368 in 94a7beb9cf
     364 | @@ -365,6 +365,7 @@ static UniValue SummaryToJSON(const IndexSummary&& summary, std::string index_na
     365 |      UniValue entry(UniValue::VOBJ);
     366 |      entry.pushKV("synced", summary.synced);
     367 |      entry.pushKV("best_block_height", summary.best_block_height);
     368 | +    entry.pushKV("first_block_height", summary.first_block_height);
    


    fjahr commented at 10:42 AM on September 28, 2026:

    synced: true doesn't really mean the same thing as before. So some clients which are not checking the new first_block_height may not get what they think they are getting.

  43. in src/index/base.cpp:133 in 94a7beb9cf
     129 | +        const CBlockIndex* start{nullptr};
     130 | +        if (GetPrunePolicy() == IndexPrunePolicy::PartialHistory && m_chainstate->m_blockman.IsPruneMode() && index_chain.Tip()) {
     131 | +            const auto& first{m_chainstate->m_blockman.GetFirstBlock(*index_chain.Tip(), BLOCK_HAVE_DATA)};
     132 | +            start = first.pprev;
     133 | +            first_block_height = first.nHeight;
     134 | +            if (start) LogInfo("%s starting at height %d; earlier history will not be indexed", GetName(), first.nHeight);
    


    fjahr commented at 10:45 AM on September 28, 2026:

    I am not sure if this is enough to communicate what is happening to the user. Starting with -prune=XXX -txindex=1 may give you a partial or full txindex depending on if you had the index on before. Possibly the partial index could use an explicit opt-in? Not sure if that is needed but it seems to be closer to what @ajtowns had in mind originally as well when he asked for -prunetxindex=1.

  44. in src/rpc/rawtransaction.cpp:340 in 94a7beb9cf
     336 | @@ -341,10 +337,12 @@ static RPCMethod getrawtransaction()
     337 |          } else if (!f_txindex_ready) {
     338 |              errmsg = "No such mempool transaction. Blockchain transactions are still in the process of being indexed";
     339 |          } else {
     340 | -            errmsg = "No such mempool or blockchain transaction";
     341 | +            errmsg = "No such mempool or blockchain transaction" + TxIndexMissErrorDetails(g_txindex->GetSummary());
    


    fjahr commented at 11:01 AM on September 28, 2026:

    Because of what we do with f_txindex_ready above I am not sure all the infos in TxIndexMissErrorDetails would ever be relevant or possible to hit, e.g. at least the synced check and added message.

  45. fjahr commented at 11:01 AM on September 28, 2026: contributor

    Would you mind splitting the PR into enabling plain prune mode and keep the partial index and all the related refactorings for a follow-up to that? I see that you added it as response to feedback but the plain prune mode should be a much simpler change, pretty much uncontroversial and relatively quick to merge. For the partial index there are a lot of edge cases imaginable, and it needs to be explained to the user pretty well. I added some comments with examples (from a very shallow first pass) that indicate to me that the partial may require more discussion and shouldn't block the plain prune mode. @arejula27 also has raised some good, related questions as well.


github-metadata-mirror

This is a metadata mirror of the GitHub repository bitcoin/bitcoin. This site is not affiliated with GitHub. Content is generated from a GitHub metadata backup.
generated: 2026-09-28 19:51 UTC

This site is hosted by @0xB10C
More mirrored repositories can be found on mirror.b10c.me