Two bugs found by simulation testing:
1. Suspect nodes could never recover after dissemination budget expired.
When probing a target, the Suspect state was not re-enqueued into the
dissemination queue (only Dead was). After budget exhaustion, the
Suspect node never received a piggyback telling it it was suspected,
so it could never refute via incarnation bump. Fix: re-enqueue both
Suspect and Dead state on outgoing probes.
2. Registry entries merged before membership changes in same piggyback.
When a message carried both a death notification and registry entries,
the entries were merged first, then immediately tombstoned. Fix: split
extract_registry_piggyback into unpack + deferred merge, processing
membership changes before merging registry entries.
Authored by Claude, lovingly guided by Zachery Aaron Shores-Chmielewski
Bug 1 (swim/node.rs): SwimProbe::check_suspicion_timeouts() calls
members.declare_dead() before translate_probe_actions() processes the
DeclareDead action. The second declare_dead() returned false (already dead),
so MembershipChanged{Dead} was never emitted and the death was never
enqueued for dissemination. Fix: remove the redundant declare_dead() call
in translate_probe_actions since the probe already performed the mutation.
Bug 2 (swim/node.rs): apply_membership_update() — which processes piggyback
on every ping/ack/ping_req — updated the internal member list but never
emitted NodeAction::MembershipChanged. This meant DistributedNode was blind
to all state transitions learned via gossip piggyback (e.g., a dead node
refuting via incarnation bump). Fix: return MembershipChanged actions from
apply_piggyback and propagate through handle_ping/handle_ack/handle_ping_req.
Bug 3 (node.rs): DistributedNode::handle_ping/handle_ack/handle_ping_req
never processed MembershipChanged actions from SwimNode — only tick() did.
Fix: extract process_membership_changes() helper and call it from all four
message paths (tick, handle_ping, handle_ack, handle_ping_req).
Additional fixes:
- registry.rs: add re_disseminate_all() for anti-entropy on partition heal
- node.rs: call re_disseminate_all on MemberState::Alive transitions so
registry state accumulated during partition reaches recovering nodes
- cluster_scenarios: enable dead_reprobe in 10% message loss test, since
correct death dissemination (now working) causes cascading false deaths
without a recovery mechanism
Authored by Claude, lovingly guided by Zachery Aaron Shores-Chmielewski