feat(cluster): auto-down downing strategy — either-node crash now fails over (owner decision 2026-07-21: availability over partition-safety)
Two-node keep-oldest could NEVER survive a crash of the oldest/active node:
Akka.NET 1.5.62 KeepOldest.OldestDecision only lets down-if-alone rescue a
side with >= 2 members, so the 1-vs-1 survivor takes DownReachable and downs
ITSELF — proven live on the rig ('SBR took decision ... including myself')
before this change. static-quorum(1) is worse (IsTooManyMembers -> DownAll);
keep-majority just re-keys the fatal crash to the lowest address.
SplitBrainResolverStrategy gains 'auto-down' (new default): BuildHocon emits
Akka's AutoDowning provider with auto-down-unreachable-after = StableAfter.
The leader among the REACHABLE members downs the unreachable peer, so the
survivor takes over singletons and /health/active in ~25s regardless of which
node died. Accepted trade (explicit owner decision): a real network partition
runs dual-active until an operator restarts one side. keep-oldest remains
supported; DownIfAlone validation is now scoped to it.
Live drill on the rebuilt rig: active-crash TAKEOVER in 28s (victim still
down; all 7 singletons Younger->Oldest), standby-crash removal 27s with 0
routing blips; victims rejoin as standby in 2s. New real-cluster tests pin
both directions (SbrFailoverTests.AutoDown_*); TwoNodeClusterFixture gains a
strategy knob. All 16 appsettings flipped (src, docker, docker-env2, and the
gitignored wonder-app-vd03 overlay on disk — owner must sync to the host).
Docs: decision record docs/plans/2026-07-21-auto-down-availability-decision.md,
Component-ClusterInfrastructure downing section rewritten, drill + README
reworked (active mode now asserts takeover), deferred-work SBR row resolved.
This commit is contained in:
@@ -38,12 +38,27 @@ public class ClusterOptions
|
||||
public List<string> SeedNodes { get; set; } = new();
|
||||
|
||||
/// <summary>
|
||||
/// Split-brain resolver strategy. Must be <c>keep-oldest</c> for the two-node
|
||||
/// clusters ScadaBridge uses: quorum strategies (<c>keep-majority</c>,
|
||||
/// <c>static-quorum</c>) cannot distinguish a crash from a partition with only
|
||||
/// two nodes and would shut down the whole cluster.
|
||||
/// Downing strategy for unreachable members. Two supported values:
|
||||
/// <list type="bullet">
|
||||
/// <item><c>auto-down</c> (default, decision 2026-07-21) — availability-first: each
|
||||
/// side downs the unreachable peer after <see cref="StableAfter"/>, so a hard crash
|
||||
/// of EITHER node (oldest included) fails over to the survivor. The accepted trade:
|
||||
/// a true network partition produces two live one-node clusters (dual-active) until
|
||||
/// an operator restarts one side. Chosen because ScadaBridge pairs run one node per
|
||||
/// VM with no shared lease infrastructure, and a stalled system is a bigger risk
|
||||
/// than a rare partition.</item>
|
||||
/// <item><c>keep-oldest</c> — partition-safe SBR: downs the side without the oldest
|
||||
/// member. In a TWO-node cluster this makes a crash of the oldest/active node a
|
||||
/// total outage: Akka's <c>down-if-alone</c> only rescues the survivor when its own
|
||||
/// side has ≥2 members (verified against Akka.NET 1.5.62 <c>KeepOldest.Decide</c>
|
||||
/// and live on the docker rig, 2026-07-21).</item>
|
||||
/// </list>
|
||||
/// Other SBR strategies are rejected: <c>static-quorum</c> with quorum 1 hits Akka's
|
||||
/// <c>IsTooManyMembers</c> guard (2 > 2*1-1) and downs ALL on any unreachability;
|
||||
/// <c>keep-majority</c> just moves the fatal crash from the oldest to the
|
||||
/// lowest-address node.
|
||||
/// </summary>
|
||||
public string SplitBrainResolverStrategy { get; set; } = "keep-oldest";
|
||||
public string SplitBrainResolverStrategy { get; set; } = "auto-down";
|
||||
|
||||
/// <summary>
|
||||
/// Time the cluster membership must remain stable before the split-brain
|
||||
@@ -71,9 +86,12 @@ public class ClusterOptions
|
||||
public int MinNrOfMembers { get; set; } = 1;
|
||||
|
||||
/// <summary>
|
||||
/// The keep-oldest resolver's <c>down-if-alone</c> flag. When <c>true</c> (the
|
||||
/// design-doc requirement), the oldest node downs itself if it finds it has no
|
||||
/// other reachable members, rather than running as an isolated single-node cluster.
|
||||
/// The keep-oldest resolver's <c>down-if-alone</c> flag; only consulted when
|
||||
/// <see cref="SplitBrainResolverStrategy"/> is <c>keep-oldest</c>. When <c>true</c>,
|
||||
/// the oldest node downs itself if it finds it has no other reachable members,
|
||||
/// rather than running as an isolated single-node cluster. Note that in a two-node
|
||||
/// cluster this does NOT let the younger survivor take over from a crashed oldest —
|
||||
/// Akka's alone-check requires the surviving side to have ≥2 members.
|
||||
/// </summary>
|
||||
public bool DownIfAlone { get; set; } = true;
|
||||
|
||||
|
||||
Reference in New Issue
Block a user