feat(cluster): auto-down downing strategy — either-node crash now fails over (owner decision 2026-07-21: availability over partition-safety)

Two-node keep-oldest could NEVER survive a crash of the oldest/active node:
Akka.NET 1.5.62 KeepOldest.OldestDecision only lets down-if-alone rescue a
side with >= 2 members, so the 1-vs-1 survivor takes DownReachable and downs
ITSELF — proven live on the rig ('SBR took decision ... including myself')
before this change. static-quorum(1) is worse (IsTooManyMembers -> DownAll);
keep-majority just re-keys the fatal crash to the lowest address.

SplitBrainResolverStrategy gains 'auto-down' (new default): BuildHocon emits
Akka's AutoDowning provider with auto-down-unreachable-after = StableAfter.
The leader among the REACHABLE members downs the unreachable peer, so the
survivor takes over singletons and /health/active in ~25s regardless of which
node died. Accepted trade (explicit owner decision): a real network partition
runs dual-active until an operator restarts one side. keep-oldest remains
supported; DownIfAlone validation is now scoped to it.

Live drill on the rebuilt rig: active-crash TAKEOVER in 28s (victim still
down; all 7 singletons Younger->Oldest), standby-crash removal 27s with 0
routing blips; victims rejoin as standby in 2s. New real-cluster tests pin
both directions (SbrFailoverTests.AutoDown_*); TwoNodeClusterFixture gains a
strategy knob. All 16 appsettings flipped (src, docker, docker-env2, and the
gitignored wonder-app-vd03 overlay on disk — owner must sync to the host).
Docs: decision record docs/plans/2026-07-21-auto-down-availability-decision.md,
Component-ClusterInfrastructure downing section rewritten, drill + README
reworked (active mode now asserts takeover), deferred-work SBR row resolved.
This commit is contained in:
Joseph Doherty
2026-07-21 10:53:40 -04:00
parent dced0d2794
commit cf3bd52f93
29 changed files with 479 additions and 135 deletions
@@ -38,12 +38,27 @@ public class ClusterOptions
public List<string> SeedNodes { get; set; } = new();
/// <summary>
/// Split-brain resolver strategy. Must be <c>keep-oldest</c> for the two-node
/// clusters ScadaBridge uses: quorum strategies (<c>keep-majority</c>,
/// <c>static-quorum</c>) cannot distinguish a crash from a partition with only
/// two nodes and would shut down the whole cluster.
/// Downing strategy for unreachable members. Two supported values:
/// <list type="bullet">
/// <item><c>auto-down</c> (default, decision 2026-07-21) — availability-first: each
/// side downs the unreachable peer after <see cref="StableAfter"/>, so a hard crash
/// of EITHER node (oldest included) fails over to the survivor. The accepted trade:
/// a true network partition produces two live one-node clusters (dual-active) until
/// an operator restarts one side. Chosen because ScadaBridge pairs run one node per
/// VM with no shared lease infrastructure, and a stalled system is a bigger risk
/// than a rare partition.</item>
/// <item><c>keep-oldest</c> — partition-safe SBR: downs the side without the oldest
/// member. In a TWO-node cluster this makes a crash of the oldest/active node a
/// total outage: Akka's <c>down-if-alone</c> only rescues the survivor when its own
/// side has ≥2 members (verified against Akka.NET 1.5.62 <c>KeepOldest.Decide</c>
/// and live on the docker rig, 2026-07-21).</item>
/// </list>
/// Other SBR strategies are rejected: <c>static-quorum</c> with quorum 1 hits Akka's
/// <c>IsTooManyMembers</c> guard (2 &gt; 2*1-1) and downs ALL on any unreachability;
/// <c>keep-majority</c> just moves the fatal crash from the oldest to the
/// lowest-address node.
/// </summary>
public string SplitBrainResolverStrategy { get; set; } = "keep-oldest";
public string SplitBrainResolverStrategy { get; set; } = "auto-down";
/// <summary>
/// Time the cluster membership must remain stable before the split-brain
@@ -71,9 +86,12 @@ public class ClusterOptions
public int MinNrOfMembers { get; set; } = 1;
/// <summary>
/// The keep-oldest resolver's <c>down-if-alone</c> flag. When <c>true</c> (the
/// design-doc requirement), the oldest node downs itself if it finds it has no
/// other reachable members, rather than running as an isolated single-node cluster.
/// The keep-oldest resolver's <c>down-if-alone</c> flag; only consulted when
/// <see cref="SplitBrainResolverStrategy"/> is <c>keep-oldest</c>. When <c>true</c>,
/// the oldest node downs itself if it finds it has no other reachable members,
/// rather than running as an isolated single-node cluster. Note that in a two-node
/// cluster this does NOT let the younger survivor take over from a crashed oldest —
/// Akka's alone-check requires the surviving side to have ≥2 members.
/// </summary>
public bool DownIfAlone { get; set; } = true;