This is a full implementation of Erlay. Its purpose is to check the integrity and correctness of the implementation against changes/additions that may originate from the review process and/or rebases on top of newer functionality.
This is not to be merged. Functionality will be spread across multiple smaller PRs to ease the review process.
Approach
This approach uses Erlay as a fallback mechanism for transaction propagation. Instead of mixing fanout and reconciliation into a single connection type, the current approach leaves the existing connections as they are, and adds additional low-bandwidth connections to be used in case the node is being eclipsed. This connections should have minimal cost under normal circumstances, and only undergo real traffic in case other 8 full-outbound connections are being captured.
outbound-full-reconciliation connections
For now, we are adding 4 additional reconciliation-only connections to the node while we test it's impact on real node running the approach. Further analysis may be needed to pick a meaningful value for this. The number of inbound connections should also be scaled based on how many connections we are adding.
DrahtBot
commented at 12:40 PM on June 23, 2026:
contributor
<!--e57a25ab6845829454e8d69fc972939a-->
The following sections might be updated with supplementary metadata relevant to reviewers and maintainers.
See the guideline and AI policy for information on the review process.
A summary of reviews will appear here.
<!--174a7506f384e20aa4161008e828411d-->
Conflicts
Reviewers, this pull request conflicts with the following ones:
#35561 (net: move some CNodeState fields to Peer by Crypt-iQ)
#35522 (refactor: Extract per-message helpers from SendMessages() (move-only) by pablomartin4btc)
#34743 (p2p: don't disconnect manual peers for block stalling by willcl-ark)
#32554 (bench: replace embedded raw block with configurable block generator by l0rinc)
#31260 (scripted-diff: Type-safe settings retrieval by ryanofsky)
If you consider this pull request important, please also help to review the conflicting pull requests. Ideally, start with the one that should be merged first.
<!--5faf32d7da4f0f540f40219e4f7537a3-->
LLM Linter (✨ experimental)
Possible typos and grammar issues:
Giving it a 6x marging to prevent flakiness -> Giving it a 6x margin to prevent flakiness [“marging” is misspelled]
peer1 only has one one tx in the set, which matches out sketch som no diff -> peer1 only has one tx in the set, which matches our sketch so no diff [multiple typos/broken words: “one one”, “out”, “som”]
Possible places where named args for integral literals may be used (e.g. func(x, /*named_arg=*/0) in C++, and func(x, named_arg=0) in Python):
AddTxsToReconSet(tracker, peer_id2, 5) in src/test/txreconciliation_tests.cpp
AddTxsToReconSet(tracker, peer_id0, 5) in src/test/txreconciliation_tests.cpp
AddTxsToReconSet(tracker, peer_id0, 10) in src/test/txreconciliation_tests.cpp
AddTxsToReconSet(tracker, peer_id1, 10) in src/test/txreconciliation_tests.cpp
AddTxsToReconSet(tracker, peer_id2, 10) in src/test/txreconciliation_tests.cpp
self.generate_txs(self.wallet, 0, 10, 0) in test/functional/p2p_txrecon_initiator.py
self.test_reconciliation_initiator_no_extension(3, 3, 20) in test/functional/p2p_txrecon_initiator.py
self.generate_txs(self.wallet, 0, 1, 0) in test/functional/p2p_txrecon_responder.py
self.generate_txs(self.wallet, 0, 5, 0) in test/functional/p2p_txrecon_responder.py
self.test_reconciliation_responder_flow_no_extension(3, 3, 20) in test/functional/p2p_txrecon_responder.py
Possible places where comparison-specific test macros should replace generic comparisons:
sr-gi
commented at 1:13 PM on June 23, 2026:
member
Rebased on master.
Opening this so it can be tested against CI and to make sure nothing obvious is missing. Next step will be testing with real nodes (most likely in Warnet).
Things to consider/add:
This is currently using sendtxrcncl as a way, for outbound nodes, to signal they want to establish a full-reconciliation connection, plus adding a limit on the number of full-reconcilition inbounds that will be accepted by a node. This is not the originally intended way of using sendtxrcncl. It may be worth considering an approach based on BIP-434.
The extension phase was designed so peers with an ongoing reconciliation that had underpredicted the sketch capacity could still reconcile without having to directly fallback to fanout. Given reconciliation is now used as fallback, and the extra bandwidth may not be as problematic, it may be worth considering overshooting the initial sketch capacity and getting rid of the extension phase.
Fuzz tests are missing
DrahtBot added the label CI failed on Jun 23, 2026
sr-gi force-pushed on Jun 23, 2026
sr-gi force-pushed on Jun 23, 2026
DrahtBot removed the label CI failed on Jun 23, 2026
sr-gi force-pushed on Jun 23, 2026
sr-gi
commented at 11:04 PM on June 23, 2026:
member
Fixed some typos and addressed some of the linter suggestions
brunoerg
commented at 4:23 PM on June 29, 2026:
contributor
**Testing report (LLM/experimental) based on the results of an incremental mutation testing run for this PR - full result is avaliable at: https://bitcoincore.space**
The overall result is mixed. The PR already has useful coverage for the happy-path reconciliation flow, basic protocol violations, and some queue/timer behavior. However, the survivors show that the tests still leave several important behaviors only weakly specified, especially where reconciliation falls back, extends, or interacts with INV relay bookkeeping.
What the current tests do well
The existing tests cover the basic handshake and some normal flows reasonably well:
unit coverage for peer registration, queue rotation, set insertion/removal, and non-extension reconciliation in src/test/txreconciliation_tests.cpp
protocol-violation checks for clearly invalid message ordering and unsupported peers
That baseline is good enough to kill most straightforward mutations. The survivors are mostly about what happens around the edges of the protocol, not the central flow.
Main gaps exposed by the surviving mutants
1. Extension handling is under-tested
This is the clearest gap. Several survivors change extension behavior without being detected:
not sending REQSKETCHEXT after an undecodable sketch in net_processing
add end-to-end functional tests that force an extension round and assert the exact message sequence: REQTXRCNCL -> SKETCH -> REQSKETCHEXT -> SKETCH -> RECONCILDIFF
assert both extension success and extension failure behavior
verify that extension failure falls back to announcing the snapshotted set, not the live set
verify that state is cleared after extension completion and that a second reconciliation starts cleanly
2. Boundary conditions are not pinned down tightly enough
Several survivors change strict inequalities to inclusive ones and still pass:
remote_sketch_capacity > MAX_SKETCH_CAPACITY changed to >=
extended_capacity > MAX_SKETCH_CAPACITY * 2 changed to >=
peer_q > Q_PRECISION changed to >=
This means the tests exercise invalid-above-limit cases, but not the exact boundary values that should remain valid. There is already a good example of this style for rounded q formatting in src/test/txreconciliation_tests.cpp, but the same precision is missing for protocol acceptance thresholds.
What should be improved:
add exact-boundary tests for MAX_SKETCH_CAPACITY, 2 * MAX_SKETCH_CAPACITY, and Q_PRECISION
check both sides of each threshold: limit must pass, limit + 1 must fail
3. INV relay side effects are only partially asserted
Many survivors in PeerManagerImpl::AnnounceTxs() and the send path in SendMessages() remove or alter important side effects without breaking tests:
skipping heap construction/pop order
changing continue to break when one candidate should not be sent
not inserting into m_tx_inventory_known_filter
not erasing from m_tx_inventory_to_send
not flushing INV batches at MAX_INV_SZ
not clearing the batch after sending
The current tests mostly check that transactions eventually arrive. They do not strongly check how they are batched, filtered, or suppressed from future re-announcement. That leaves a lot of bookkeeping mutations alive.
What should be improved:
add tests with a mix of sendable and unsendable transactions and assert that later eligible transactions are still announced
assert that transactions already known to the peer are not re-announced in the next round
add a case above MAX_INV_SZ and verify the number and sizes of INV messages
verify that reconciliation and fanout paths do not duplicate announcements after internal state should have been cleared
4. Some state-management behaviors are observable in principle, but not asserted
Survivors also show weak checking around reconciliation state transitions:
not clearing m_short_id_mapping
not clearing peer state after handling a result
always taking the “removed” branch in TryRemovingFromSet
not recording m_announced_while_reconciling
returning success from queue-selection paths that should be false
These are not all equally severe, but together they indicate that the tests often validate the final external effect of a single round without checking the follow-up round that would expose stale state.
What should be improved:
add two-step tests that perform one reconciliation round and then immediately start another to detect stale snapshots, stale mappings, or stale queue state
specifically verify the “received while reconciling” behavior across the extension path, not just the non-extension path
Bottom line
The current tests are good at proving that the basic txreconciliation flow works. They are not yet strong enough to fully specify the extension path, the exact protocol boundaries, or the internal bookkeeping that prevents duplicate, truncated, or stale announcements.
If I had to prioritize follow-up work, I would do it in this order:
Add full initiator/responder extension-path functional tests.
Add exact boundary tests for sketch capacity and q.
Add stronger assertions around INV batching, filtering, and duplicate suppression.
Add second-round/state-cleanup tests to catch stale reconciliation state.
That would likely eliminate most of the meaningful survivors from this run and materially improve confidence in the PR’s tests.
DrahtBot added the label Needs rebase on Jul 24, 2026
sr-gi force-pushed on Aug 13, 2026
DrahtBot removed the label Needs rebase on Aug 13, 2026
sr-gi
commented at 8:05 PM on August 13, 2026:
member
sr-gi
commented at 7:23 PM on August 19, 2026:
member
Added benches for how long it takes to build and decode sketches based on their capacity. This may affect the max sketch capacity that we may accept to work with.
ns/element
element/s
err%
total
benchmark
9,992.50
100,075.03
0.6%
0.33
ReconcileSketchConstructReconSetMax
536,306.23
1,864.61
0.1%
96.69
ReconcileSketchDecodeExtensionMax
281,615.18
3,550.94
0.1%
25.38
ReconcileSketchDecodeMaxCapacity
39,963.18
25,023.03
1.1%
1.30
ReconcileSketchDecodeReconSetMax
2,379.92
420,182.37
3.5%
0.01
ReconcileSketchDecodeTypical
This results in the following:
benchmark
capacity
per-decode
ReconcileSketchDecodeTypical
54
0.13 ms
ReconcileSketchDecodeReconSetMax
3001
120 ms
ReconcileSketchDecodeMaxCapacity
8192
2.3 s
ReconcileSketchDecodeExtensionMax
16384
8.8 s
The decode for a typical/realistic round (7 tx/s, 30s, q=0.25) is in the order of tens of microseconds.
A round between two Core nodes (limiting their set size to MAX_RECONSET_SIZE) takes ~120ms
A peer that advertises sketches as big as they get (MAX_SKETCH_CAPACITY) can make us take ~2.3s to decode
This is amplified to 8.8s in the extension case, as we allow up to double the max capacity
This applies only to outbound peers, as those are the ones that makes us decode. This bears the question, should we limit the maximum sketch size that we accept far bellow MAX_SKETCH_CAPACITY?
In the case of Core nodes, sets cannot have more than MAX_RECONSET_SIZE elements. So the maximum sketch sizes are 3001 for the initial sketch, 6002 for an extension.
DrahtBot added the label CI failed on Aug 19, 2026
DrahtBot
commented at 7:24 PM on August 19, 2026:
contributor
<!--85328a0da195eb286784d51f73fa0af9-->
🚧 At least one of the CI tasks failed.
<sub>Task iwyu: https://github.com/bitcoin/bitcoin/actions/runs/32289116534/job/96185584884</sub>
<sub>LLM reason (✨ experimental): CI failed because IWYU detected and rejected incorrect/missing #include ordering/content (generated “Failure generated from IWYU” for src/bench/txreconciliation.cpp).</sub>
<details><summary>Hints</summary>
Try to run the tests locally, according to the documentation. However, a CI failure may still
happen due to a number of reasons, for example:
Possibly due to a silent merge conflict (the changes in this pull request being
incompatible with the current code in the target branch). If so, make sure to rebase on the latest
commit of the target branch.
A sanitizer issue, which can only be found by compiling with the sanitizer and running the
affected test.
An intermittent issue.
Leave a comment here, if you need help tracking down a confusing failure.
</details>
sr-gi force-pushed on Aug 19, 2026
DrahtBot removed the label CI failed on Aug 19, 2026
sr-gi
commented at 10:22 AM on August 20, 2026:
member
This should be ready to review at this point, even though some details still need deciding. Happy to split the PR into smaller chunks if it makes it easier.
Open questions (that may arise from reviewing parts of the code):
Should we negotiate reconciliation using BIP-434?
Should the new connection type be considered a full-outbound? If so, the eviction logic needs to be patched to include these.
Are extensions worth if we will be using reconciliation as a fallback, or should we just slightly overestimate the sketch capacity and default to fanout on failure?
Should we cap the maximum capacity sketch we accept closer to MAX_RECONSET_SIZE than to MAX_SKETCH_CAPACITY?
DrahtBot added the label Needs rebase on Aug 24, 2026
refactor: redesigns txreconciliation file split and namespace
Splits the txreconciliation logic in three files instead of two, allowing the
TxreconciliationState to be properly tested, instead of being internal to
txreconciliation.cpp.
Also includes everything in the node namespace, instead of being part
of an anonymous one.
7bd378c211
refactor: remove legacy comments
These comments became irrelevant in one of the previous code changes.
They simply don't make sense anymore.
61ddf8b3f9
refactor: Defines generic error to be used in several reconciliation methods8001a217b2
refactor: add full stop in existing txreconciliation LogDebug lines8c6d5b8626
p2p: Allows inbound reconciliation connections up to a limit
Set the current limit to 32.
5cc559b300
net, gui, test: adds new connection type (OUTBOUND_FULL_RECONCILIATION)
Adds a new connection type that will be used for reconciliation only.
Defines the default max number of this type of connections to 4.
6d4c8958e5
p2p: send SENDTXRCNCL messages only over OUTBOUND_FULL_RECONCILIATION connection4e934e9482
p2p: Functions to add/remove wtxids to tx reconciliation sets
They will be used later on.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
07bc2023ab
p2p: Add transactions to reconciliation sets
Transactions are added to the reconciliation sets of reconciling peers, and processed normally for fanout peers.
c7d795891b
p2p: Add helper to compute reconciliation tx short ids and a cache of short ids to wtxidsf70bbe45ff
p2p: Deal with shortid collisions for reconciliation sets
If a transaction to be added to a peer's recon set has a shot id collisions (a previously
added wtxid maps to the same short id), both transaction should be fanout, given
our peer may have added the opposite transaction to our recon set, and these two
transaction won't be reconciled.
0ad81eb3fc
p2p: Add peers to reconciliation queue on negotiation
When we're finalizing negotiation, we should add the peers
for which we will initiate reconciliations to the queue.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
b7db597f8d
p2p: Track reconciliation requests schedule
We initiate reconciliation by looking at the queue periodically
with equal intervals between peers to achieve efficiency.
This will be later used to see whether it's time to initiate.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
eade3c245e
p2p: Initiate reconciliation round
When the time comes for the peer, we send a
reconciliation request with the parameters which
will help the peer to construct a (hopefully) sufficient
reconciliation sketch for us. We will then use that
sketch to find missing transactions.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
When the time comes, we should send a sketch of our
local reconciliation set to the reconciliation initiator.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
00a61e5ced
p2p: Add a function to identify local/remote missing txs
When the sketches from both sides are combined successfully,
the diff is produced. Then this diff can (together with the local txs)
be used to identified which transactions are missing locally and remotely.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
4c017ecf66
refactor: Extract AnnounceTxs out of SendMessages
Move the transaction announcement loop from SendMessages into its own
method, so it can be reused to announce transactions after a
reconciliation round. No behaviour change.
4aaf75c678
p2p: Add a function to announce transactions after reconciliation
Transactions the peer is found to be missing during a reconciliation round
need to be announced right away, rather than waiting for the next trickle
interval and being added back to the reconciliation set.
Add AnnounceReconciliationTxs on top of AnnounceTxs.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
ac702e5d0b
p2p: Handle reconciliation sketch and successful decoding
If after decoding a reconciliation sketch it turned out
to be insufficient to find set difference, request extension.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
357d928df3
p2p: Be ready to receive sketch extension
Store the initial sketches so that we are able to process
extension sketch while avoiding transmitting the same data.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
4e012c64e5
p2p: Prepare for sketch extension request
To be ready to respond to a sketch extension request
from our peer, we should store a snapshot of our state
and capacity of the initial sketch, so that we compute
extension of the same size and over the exact same
transactions.
Transactions arriving during this reconciliation will
be instead stored in the regular set.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
2ffc8711c1
p2p: Keep track of announcements during txrcncl extension
If peer failed to reconcile based on our initial response sketch,
they will ask us for a sketch extension. Store this request to respond later.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
0cc7a189c0
p2p: Respond to sketch extension request
Sending an extension may allow the peer to reconcile
transactions, because now the full sketch has twice
as much capacity.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
3765ea63ca
p2p: Handle sketch extension
If a peer sent us an extension sketch, we should
reconstruct a full sketch from it with the snapshot
we stored initially, and attempt to decode the difference.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
c46c20de67
p2p: Add a finalize incoming reconciliation function
This currently unused function is supposed to be used once
a reconciliation round is done. It cleans the state corresponding
to the passed reconciliation.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
e614d19949
p2p: Handle reconciliation finalization message
Once a peer tells us reconciliation is done, we should behave as follows:
- if it was successful, just respond them with the transactions they asked
by short ID.
- if it was a full failure, respond with all local transactions from the reconciliation
set snapshot
- if it was a partial failure (only low or high part was failed after a bisection),
respond with all transactions which were asked for by short id,
and announce local txs which belong to the failed chunk.
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
2352f568a9
p2p, test: Add tx reconciliation functional tests
We may still need to add more tests, specially around extensions (if we keep them)
Co-authored-by: Gleb Naumenko <naumenko.gs@gmail.com>
fe1a5c5156
p2p: Remove transactions from reconciliation sets when removed from the mempool
A transaction that has left our mempool can no longer be served, but nothing
removed it from the reconciliation sets it had been added to. It would still be
sketched, requested by the peer, and then dropped, wasting sketch capacity and a
round trip after every block.
Prune on both removal signals: removeUnchecked deliberately skips
TransactionRemovedFromMempool for MemPoolRemovalReason::BLOCK, so mined
transactions, which are the bulk of the case, only arrive via
MempoolTransactionsRemovedForBlock.
Snapshots of an in-flight round are left untouched so that a sketch extension
still describes the same elements as the sketch already sent. Flagging the wtxid
instead keeps it out of the announcement.
a101971de3
bench: Adds txreconciliation benches
Raises the question of whether the current sketch capacity limits are too high
534a89069d
sr-gi force-pushed on Aug 24, 2026
DrahtBot removed the label Needs rebase on Aug 24, 2026
sr-gi
commented at 12:34 PM on August 24, 2026:
member
This is a metadata mirror of the GitHub repository
bitcoin/bitcoin.
This site is not affiliated with GitHub.
Content is generated from a GitHub metadata backup.
generated: 2026-08-26 05:51 UTC
This site is hosted by @0xB10C More mirrored repositories can be found on mirror.b10c.me