feat(cluster): auto-down downing strategy — either-node crash now fails over (owner decision 2026-07-21: availability over partition-safety)

Two-node keep-oldest could NEVER survive a crash of the oldest/active node:
Akka.NET 1.5.62 KeepOldest.OldestDecision only lets down-if-alone rescue a
side with >= 2 members, so the 1-vs-1 survivor takes DownReachable and downs
ITSELF — proven live on the rig ('SBR took decision ... including myself')
before this change. static-quorum(1) is worse (IsTooManyMembers -> DownAll);
keep-majority just re-keys the fatal crash to the lowest address.

SplitBrainResolverStrategy gains 'auto-down' (new default): BuildHocon emits
Akka's AutoDowning provider with auto-down-unreachable-after = StableAfter.
The leader among the REACHABLE members downs the unreachable peer, so the
survivor takes over singletons and /health/active in ~25s regardless of which
node died. Accepted trade (explicit owner decision): a real network partition
runs dual-active until an operator restarts one side. keep-oldest remains
supported; DownIfAlone validation is now scoped to it.

Live drill on the rebuilt rig: active-crash TAKEOVER in 28s (victim still
down; all 7 singletons Younger->Oldest), standby-crash removal 27s with 0
routing blips; victims rejoin as standby in 2s. New real-cluster tests pin
both directions (SbrFailoverTests.AutoDown_*); TwoNodeClusterFixture gains a
strategy knob. All 16 appsettings flipped (src, docker, docker-env2, and the
gitignored wonder-app-vd03 overlay on disk — owner must sync to the host).
Docs: decision record docs/plans/2026-07-21-auto-down-availability-decision.md,
Component-ClusterInfrastructure downing section rewritten, drill + README
reworked (active mode now asserts takeover), deferred-work SBR row resolved.
This commit is contained in:
Joseph Doherty
2026-07-21 10:53:40 -04:00
parent dced0d2794
commit cf3bd52f93
29 changed files with 479 additions and 135 deletions
@@ -195,4 +195,49 @@ public class HoconBuilderTests
"Akka.Cluster.SBR.SplitBrainResolverProvider, Akka.Cluster",
config.GetString("akka.cluster.downing-provider-class"));
}
[Fact]
public void BuildHocon_AutoDownStrategy_EmitsAutoDowningProvider()
{
// Decision 2026-07-21 (availability over partition-safety): 'auto-down' must
// swap the downing provider to Akka's AutoDowning so the survivor downs a
// crashed peer — including a crashed OLDEST, which two-node keep-oldest
// cannot survive — after StableAfter.
var cluster = DefaultCluster();
cluster.SplitBrainResolverStrategy = "auto-down";
cluster.StableAfter = TimeSpan.FromSeconds(15);
var hocon = AkkaHostedService.BuildHocon(
DefaultNode(), cluster, new[] { "Central" },
TimeSpan.FromSeconds(5), TimeSpan.FromSeconds(15));
var config = ConfigurationFactory.ParseString(hocon);
Assert.Equal(
"Akka.Cluster.AutoDowning, Akka.Cluster",
config.GetString("akka.cluster.downing-provider-class"));
Assert.Equal(
TimeSpan.FromSeconds(15),
config.GetTimeSpan("akka.cluster.auto-down-unreachable-after"));
// The SBR section must NOT be active alongside AutoDowning.
Assert.False(config.HasPath("akka.cluster.split-brain-resolver.active-strategy"));
}
[Fact]
public void BuildHocon_AutoDownStrategy_IsCaseInsensitive_AndDocumentStaysIntact()
{
var cluster = DefaultCluster();
cluster.SplitBrainResolverStrategy = "Auto-Down";
var hocon = AkkaHostedService.BuildHocon(
DefaultNode(), cluster, new[] { "Central" },
TimeSpan.FromSeconds(5), TimeSpan.FromSeconds(15));
var config = ConfigurationFactory.ParseString(hocon);
Assert.Equal(
"Akka.Cluster.AutoDowning, Akka.Cluster",
config.GetString("akka.cluster.downing-provider-class"));
// Keys after the downing block must remain intact (document not corrupted).
Assert.Equal(1, config.GetInt("akka.cluster.min-nr-of-members"));
Assert.True(config.GetBoolean("akka.cluster.run-coordinated-shutdown-when-down"));
}
}