feat(cluster): auto-down downing strategy — either-node crash now fails over (owner decision 2026-07-21: availability over partition-safety)

Two-node keep-oldest could NEVER survive a crash of the oldest/active node:
Akka.NET 1.5.62 KeepOldest.OldestDecision only lets down-if-alone rescue a
side with >= 2 members, so the 1-vs-1 survivor takes DownReachable and downs
ITSELF — proven live on the rig ('SBR took decision ... including myself')
before this change. static-quorum(1) is worse (IsTooManyMembers -> DownAll);
keep-majority just re-keys the fatal crash to the lowest address.

SplitBrainResolverStrategy gains 'auto-down' (new default): BuildHocon emits
Akka's AutoDowning provider with auto-down-unreachable-after = StableAfter.
The leader among the REACHABLE members downs the unreachable peer, so the
survivor takes over singletons and /health/active in ~25s regardless of which
node died. Accepted trade (explicit owner decision): a real network partition
runs dual-active until an operator restarts one side. keep-oldest remains
supported; DownIfAlone validation is now scoped to it.

Live drill on the rebuilt rig: active-crash TAKEOVER in 28s (victim still
down; all 7 singletons Younger->Oldest), standby-crash removal 27s with 0
routing blips; victims rejoin as standby in 2s. New real-cluster tests pin
both directions (SbrFailoverTests.AutoDown_*); TwoNodeClusterFixture gains a
strategy knob. All 16 appsettings flipped (src, docker, docker-env2, and the
gitignored wonder-app-vd03 overlay on disk — owner must sync to the host).
Docs: decision record docs/plans/2026-07-21-auto-down-availability-decision.md,
Component-ClusterInfrastructure downing section rewritten, drill + README
reworked (active mode now asserts takeover), deferred-work SBR row resolved.
This commit is contained in:
Joseph Doherty
2026-07-21 10:53:40 -04:00
parent dced0d2794
commit cf3bd52f93
29 changed files with 479 additions and 135 deletions
+3 -3
View File
@@ -213,13 +213,13 @@ Other peers in the `scadaproj` family (see `scadaproj/CLAUDE.md` for details): `
- Two-person MxGateway secured writes (M7): two new global roles — `Operator` (initiates) + `Verifier` (approves) — added alongside the canonical `Administrator`/`Designer`/`Deployer`/`Viewer`, with `RequireOperator`/`RequireVerifier` policies. An Operator submits a secured write from the Central UI Secured Writes page (`/operations/secured-writes`); it stays a `Pending` `PendingSecuredWrite` row until a *distinct* Verifier approves it (no-self-approval enforced server-side in the ManagementActor, plus a compare-and-swap race guard). Approval relays a `WriteTagRequest` to the site MxGateway; MxGateway-protocol connections only; each lifecycle event (submit/approve/reject/execute) emits a best-effort `AuditChannel.SecuredWrite` / `AuditKind.SecuredWrite*` central-direct-write row sharing the row id as `CorrelationId`. (SecuredWrite audit rows stamp `SourceNode` via `ICentralAuditWriter`/`INodeIdentityProvider`.) Pending secured writes expire server-side after a configurable TTL (`ManagementServiceOptions.SecuredWritePendingTtl`, default 24 h): an overdue `Pending` row is CAS'd to `Expired` (never relayed) — enforced at approve/reject and swept opportunistically on list (arch-review S2, `AuditKind.SecuredWriteExpire`).
### Cluster & Failover
- Keep-oldest split-brain resolver with `down-if-alone = on`, 15s stable-after.
- **`auto-down` downing strategy (decision 2026-07-21 — availability over partition-safety).** Akka's `AutoDowning` provider, `auto-down-unreachable-after` = 15s: the leader among the REACHABLE members downs the unreachable peer, so a hard crash of EITHER node (active/oldest included) fails over to the survivor in ~25s. Accepted trade: a real partition → dual-active until an operator restarts one side. `keep-oldest` remains a supported `SplitBrainResolverStrategy` value (partition-safe, but an oldest-crash is a total outage — Akka's `down-if-alone` only rescues a side with ≥2 members, proven live + in 1.5.62 source). Decision record: `docs/plans/2026-07-21-auto-down-availability-decision.md`.
- Both nodes are seed nodes. `min-nr-of-members = 1`.
- Failure detection: 2s heartbeat, 10s threshold. Total failover ~25s (drill-measured 27s, 0 routing blips — `docker/failover-drill.sh`).
- Failure detection: 2s heartbeat, 10s threshold. Total failover ~25s (drill-measured 2026-07-21 under auto-down: active-crash TAKEOVER in 28s, standby-crash removal in 27s with 0 routing blips — `docker/failover-drill.sh`).
- CoordinatedShutdown for graceful singleton handover.
- Automatic dual-node recovery from persistent storage.
- **Active/standby is decided by `ActiveNodeEvaluator.SelfIsOldestUp`, never by cluster leadership** — see the Architecture note above. `/health/active` is **central-only** (site nodes map no `/health/*` at all) and backs both Traefik's active-node routing and `IActiveNodeGate`, so the proxy and the Inbound API always agree on which node is active. Central never needs to know which *site* node is active: ClusterClient contact rotation reaches either receptionist and the site-internal `ClusterSingletonProxy` lands the work on the active node for free. The **exception is gRPC**, which picks `GrpcNodeAAddress`/`GrpcNodeBAddress` explicitly and flips on error.
- **KNOWN GAP — two-node keep-oldest has a total-outage hole.** Only the *first* seed listed in `Cluster:SeedNodes` may self-join to form a new cluster, so a lone restarted non-first-seed node (first seed still down) loops on `InitJoin` forever — never `Up`, never routable. After the oldest/active node crashes, the younger survivor self-downs and **cannot re-bootstrap alone**; recovery is operator-driven. `failover-drill.sh` has `DRILL_MODE=active` specifically *"to make the registered gap observable — not to pretend it is covered."* Closing it is the registered deferred keep-oldest topology/strategy decision. See `docs/requirements/Component-ClusterInfrastructure.md:113`.
- **Seed-node boot-alone constraint (still real under auto-down).** Only the *first* seed listed in `Cluster:SeedNodes` may self-join to form a new cluster, so a node that must BOOT alone while its peer is dead (cold start of only the non-first-seed VM, or the survivor crashing while the peer is still down) loops on `InitJoin` forever — never `Up`, never routable; recovery is operator-driven (restart the first seed, or a self-first seed override). The former keep-oldest active-crash total outage is CLOSED by the auto-down decision — the survivor keeps running and takes over; `failover-drill.sh DRILL_MODE=active` now asserts that takeover. See `docs/requirements/Component-ClusterInfrastructure.md` → Downing Strategy.
### UI & Monitoring
- Central UI: Blazor Server (ASP.NET Core + SignalR) with Bootstrap CSS. No third-party component frameworks (no Blazorise, MudBlazor, Radzen, etc.). Build custom Blazor components for tables, grids, forms, etc.
@@ -11,7 +11,7 @@
"akka.tcp://scadabridge@scadabridge-env2-central-a:8081",
"akka.tcp://scadabridge@scadabridge-env2-central-b:8081"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
@@ -11,7 +11,7 @@
"akka.tcp://scadabridge@scadabridge-env2-central-a:8081",
"akka.tcp://scadabridge@scadabridge-env2-central-b:8081"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
@@ -13,7 +13,7 @@
"akka.tcp://scadabridge@scadabridge-env2-site-x-a:8082",
"akka.tcp://scadabridge@scadabridge-env2-site-x-b:8082"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
@@ -13,7 +13,7 @@
"akka.tcp://scadabridge@scadabridge-env2-site-x-a:8082",
"akka.tcp://scadabridge@scadabridge-env2-site-x-b:8082"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
+14 -12
View File
@@ -273,29 +273,31 @@ All test passwords are `password`. See `infra/glauth/config.toml` for the full l
### Automated failover drill (`failover-drill.sh`)
```bash
DRILL_MODE=standby bash docker/failover-drill.sh # default — survivable younger-node crash
DRILL_MODE=active bash docker/failover-drill.sh # oldest-node crash — measures the registered outage gap
DRILL_MODE=standby bash docker/failover-drill.sh # default — younger-node crash, active untouched
DRILL_MODE=active bash docker/failover-drill.sh # oldest-node crash — survivor must TAKE OVER
```
The scripted drill (`docker kill` = SIGKILL, the hard-crash path — a `docker stop` would take the graceful `CoordinatedShutdown` path and would not prove crash recovery) has **two modes**, because under the unified oldest-member semantics the *active* node IS the oldest, i.e. the one crash two-node keep-oldest cannot survive:
The scripted drill (`docker kill` = SIGKILL, the hard-crash path — a `docker stop` would take the graceful `CoordinatedShutdown` path and would not prove crash recovery) has **two modes**, and since the **auto-down decision (2026-07-21)** both expect recovery — the cluster runs Akka's `AutoDowning` provider (`auto-down-unreachable-after` = 15s), under which the leader among the *reachable* members downs the unreachable peer, so a crash of either node fails over:
- **`DRILL_MODE=standby` (default) — kills the STANDBY (younger) central node.** The survivable direction: SBR downs the crashed member and the active node keeps its singletons. Expected result: **no routing outage at all** (the active node is never touched, so `/health/active` blips = 0) and member removal on the survivor within **~25s** (10s failure-detection threshold + 15s stable-after; the 2s heartbeat interval is not additive). PASS = the survivor logs the member removal within `TIMEOUT_S` (default 90s) while routing stays up.
- **`DRILL_MODE=active` — kills the ACTIVE (oldest) central node.** Expected result: a **total central outage** until the victim container is restarted — this is the registered deferred keep-oldest decision (master tracker 2026-07-08): keep-oldest downs the partition *without* the oldest, so the younger survivor downs itself, and it cannot re-form a cluster alone (see the seed-node constraint below). The drill confirms the dark window, then recovery within ~2 min of restarting the victim. The mode exists to make the registered gap *observable*, not to pretend it is covered.
- **`DRILL_MODE=standby` (default) — kills the STANDBY (younger) central node.** The active node is untouched: expected result is **no routing outage at all** (`/health/active` blips = 0) and member removal on the survivor within **~25s** (10s failure-detection threshold + 15s auto-down window; the 2s heartbeat interval is not additive). PASS = the survivor logs the downing/removal within `TIMEOUT_S` (default 90s) while routing stays up.
- **`DRILL_MODE=active` — kills the ACTIVE (oldest) central node.** The survivor must **take over while the victim is still down**: it auto-downs the dead oldest, becomes the oldest member itself, re-hosts all singletons, and its `/health/active` goes 200. PASS = survivor active within `TIMEOUT_S`, then Traefik routing to it. (Under the pre-2026-07-21 `keep-oldest` strategy this direction was a proven total outage — the younger survivor took `DownReachable` and downed itself, because Akka's `down-if-alone` only rescues a side with ≥ 2 members.)
The drill exercises S1 (SBR downing on hard crash), S3 (single active node routed through Traefik), and the Task 20 restart/rejoin contract. Requires a running cluster (`bash docker/deploy.sh`) and `curl` + `docker` on the host.
Both modes finish by restarting the victim and confirming it rejoins as a ready standby. The drill exercises downing-on-hard-crash, S3 (single active node routed through Traefik), and the Task 20 restart/rejoin contract. Requires a running cluster (`bash docker/deploy.sh`) and `curl` + `docker` on the host.
**Seed-node bootstrap constraint.** Only the FIRST seed in `Cluster:SeedNodes` may self-join to form a *new* cluster. Both central nodes list `scadabridge-central-a` first (`docker/central-node-a/appsettings.Central.json`, `docker/central-node-b/appsettings.Central.json`), so a lone restarted `central-b` (with `central-a` still down) loops on `InitJoin` forever — it never reaches `Up`, and `/health/active` never returns 200. Operator recovery actions: **(1)** restart the dead first-seed node (`central-a`) — preferred; or **(2)** restart the survivor with a self-first seed override (env `ScadaBridge__Cluster__SeedNodes__0=akka.tcp://scadabridge@<self-host>:8081`, `ScadaBridge__Cluster__SeedNodes__1=<peer>`). The repo deliberately does NOT ship self-first ordering per node: with *both* nodes self-first, a simultaneous cold start can let each self-join independently → two one-node clusters that never merge (the cold-start split-brain the identical-seed-order convention exists to prevent). The real remedy is the pending keep-oldest topology/strategy decision (deferred, owner: user).
**Partition trade (accepted).** Auto-down is availability-first: in a *real network partition* (both nodes alive, link cut) each side downs the other and both run active — dual-active until an operator restarts one side after the partition heals. This was an explicit owner decision (2026-07-21): site pairs have no shared lease infrastructure to arbitrate, and a stalled system is a bigger risk than a rare partition. See `docs/plans/2026-07-21-auto-down-availability-decision.md`.
> **Observed results** (plan R2-01 T3):
**Seed-node bootstrap constraint (still applies to boot-alone).** Only the FIRST seed in `Cluster:SeedNodes` may self-join to form a *new* cluster. Both central nodes list `scadabridge-central-a` first (`docker/central-node-a/appsettings.Central.json`, `docker/central-node-b/appsettings.Central.json`), so a lone *restarted* `central-b` (with `central-a` still down) loops on `InitJoin` forever. Under auto-down this no longer causes the active-crash outage (the survivor keeps running — it never restarts), but it still bites when a node must boot alone (cold start of only the non-first-seed VM, or the survivor crashing while its peer is still dead). Operator recovery: **(1)** restart the first-seed node (`central-a`) — preferred; or **(2)** restart the survivor with a self-first seed override (env `ScadaBridge__Cluster__SeedNodes__0=akka.tcp://scadabridge@<self-host>:8081`, `ScadaBridge__Cluster__SeedNodes__1=<peer>`). The repo deliberately does NOT ship self-first ordering per node: with *both* nodes self-first, a simultaneous cold start can let each self-join independently → two one-node clusters that never merge.
> **Observed results** (auto-down decision verification):
>
> **Run 2026-07-13** against a freshly-deployed cluster on `main` @ `99544985` (round-2 merged image; `active=central-a`). Both directions behaved exactly as the design predicts.
> **Run 2026-07-21** against a freshly-deployed cluster with `SplitBrainResolverStrategy: auto-down` (first drill: `active=central-a`). Both directions recovered.
>
> | Direction (`DRILL_MODE`) | Outcome | Measured |
> |--------------------------|---------|----------|
> | `standby` (younger-node crash) | **PASS** — SBR downed+removed the crashed `central-b`; active `central-a` kept all 7 singletons; recovered on restart. | Member removed in **27s** (budget ~25s: 10s detection + 15s stable-after); **0** `/health/active` routing blips (active node never touched); routable **0s** after victim restart. |
> | `active` (oldest-node crash) | **Outage as designed** — killing the oldest/active `central-a` made the younger `central-b` self-down (total central outage — the registered keep-oldest gap); recovered after restarting the victim, `central-b` then assuming Oldest and re-hosting all singletons. | Outage confirmed at **9s**; central routable again **4s** after restarting `central-a`. |
> | `active` (oldest-node crash) | **PASS — TAKEOVER** — `central-b` auto-downed the dead oldest, went `Younger -> Oldest` on all 7 singletons, and served `/health/active` **while the victim was still down**; restarted victim rejoined as standby. | Survivor active + Traefik routing in **28s** (budget ~25s: 10s detection + 15s auto-down + hand-over); victim ready **2s** after restart. |
> | `standby` (younger-node crash) | **PASS** — active node untouched; survivor downed+removed the crashed member; restarted victim rejoined as standby. | Member removed in **27s**; **0** `/health/active` routing blips; victim ready **2s** after restart. |
>
> Notes: the `standby` PASS shows the survivable direction is clean end-to-end (SBR `DownUnreachable` decision + per-singleton "Member removed" in the survivor log, zero routing interruption). The `active` result **empirically confirms the deferred keep-oldest topology gap** (master tracker 2026-07-08 / `docs/plans/2026-07-08-deferred-work-register.md`): a hard crash of the active/oldest central node is a total outage until that node (the first seed) is restarted the remedy remains the pending topology/strategy decision. In-process envelope (`FailoverTimingTests`, plan R2-01 T4) independently measured full failover at **33.7s**.
> Historical baseline (keep-oldest, run 2026-07-13 on `99544985`): `standby` PASS with member removal in 27s / 0 routing blips; `active` was a **total outage** — `central-b` self-downed ~20s after the kill (live SBR log 2026-07-21: `SBR took decision Akka.Cluster.SBR.DownReachable … including myself`) and could not re-bootstrap until `central-a` returned. That result is what motivated the auto-down decision. In-process envelope (`FailoverTimingTests`) measured full failover at **33.7s**.
### Central Failover
@@ -11,7 +11,7 @@
"akka.tcp://scadabridge@scadabridge-central-a:8081",
"akka.tcp://scadabridge@scadabridge-central-b:8081"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
@@ -11,7 +11,7 @@
"akka.tcp://scadabridge@scadabridge-central-a:8081",
"akka.tcp://scadabridge@scadabridge-central-b:8081"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
+52 -39
View File
@@ -1,33 +1,33 @@
#!/usr/bin/env bash
# Failover drill against the running docker cluster (bash docker/deploy.sh first).
#
# ROUND-2 REWRITE (arch-review 01 round 2, N1). The original drill killed the
# ACTIVE central node — but under the unified oldest-member semantics the
# active node IS the oldest, i.e. the one crash two-node keep-oldest CANNOT
# survive (registered deferred user decision, master tracker 2026-07-08;
# SbrFailoverTests.cs XML doc). Two modes:
# AUTO-DOWN REWRITE (decision 2026-07-21). The cluster now runs the 'auto-down'
# downing strategy (availability-first): the leader among the REACHABLE members
# downs the unreachable peer after StableAfter, so a hard crash of EITHER
# central node — the active/oldest included — fails over to the survivor. The
# accepted trade (made explicitly by the owner) is dual-active during a real
# network partition. Both drill directions therefore expect RECOVERY:
#
# DRILL_MODE=standby (default) — kills the STANDBY (younger) central node.
# The survivable direction: SBR downs the crashed member, the active node
# keeps its singletons, and Traefik routing never goes dark. PASS = the
# survivor logs the member removal within TIMEOUT_S (budget ~25s+: 10s
# failure detection + 15s stable-after) while /health/active stays up.
# The active node is untouched: expect zero /health/active routing blips
# and member removal on the survivor within ~25s (10s failure detection +
# 15s auto-down-unreachable-after).
#
# DRILL_MODE=active — kills the ACTIVE (oldest) central node. THE EXPECTED
# OUTCOME IS A TOTAL CENTRAL OUTAGE: keep-oldest downs the partition
# without the oldest, so the younger survivor downs ITSELF (down-if-alone
# cannot help — the alone-oldest is dead and cannot down itself), and the
# self-downed survivor cannot re-form a cluster alone unless it is the
# FIRST seed (both nodes list central-a first; only the first seed may
# self-join). This mode measures the dark window and PASSes only when
# central recovers AFTER the victim container is restarted. It exists to
# make the registered gap observable — not to pretend it is covered.
# DRILL_MODE=active — kills the ACTIVE (oldest) central node. THE SURVIVOR
# MUST TAKE OVER: it downs the dead oldest, becomes oldest itself, hosts
# the singletons, and /health/active goes 200 on the survivor WHILE THE
# VICTIM IS STILL DOWN. Budget ~25s + singleton hand-over + health-probe
# margin. (Under the pre-2026-07-21 keep-oldest strategy this direction
# was a total outage — the younger survivor downed ITSELF, verified live;
# Akka's down-if-alone only rescues a side with >= 2 members.)
#
# Both modes finish by restarting the victim and confirming it rejoins as a
# fresh incarnation (standby).
set -euo pipefail
TRAEFIK_URL="${TRAEFIK_URL:-http://localhost:9000}"
TIMEOUT_S="${TIMEOUT_S:-90}"
DRILL_MODE="${DRILL_MODE:-standby}"
OUTAGE_CONFIRM_S="${OUTAGE_CONFIRM_S:-60}"
active_container() {
if curl -sf -o /dev/null "http://localhost:9001/health/active"; then echo scadabridge-central-a
@@ -35,6 +35,7 @@ active_container() {
else echo "ERROR: no active central node found" >&2; exit 1; fi
}
peer_of() { [ "$1" = scadabridge-central-a ] && echo scadabridge-central-b || echo scadabridge-central-a; }
port_of() { [ "$1" = scadabridge-central-a ] && echo 9001 || echo 9002; }
case "$DRILL_MODE" in
standby|active) ;;
@@ -47,6 +48,7 @@ if [ "$DRILL_MODE" = standby ]; then
else
VICTIM="$ACTIVE"; SURVIVOR=$(peer_of "$ACTIVE")
fi
SURVIVOR_PORT=$(port_of "$SURVIVOR")
echo "mode=${DRILL_MODE} active=${ACTIVE} victim=${VICTIM} survivor=${SURVIVOR}"
KILL_AT=$(date -u +%Y-%m-%dT%H:%M:%SZ)
@@ -54,60 +56,71 @@ docker kill "${VICTIM}" > /dev/null
START=$(date +%s)
if [ "$DRILL_MODE" = standby ]; then
echo "Standby crash: waiting for ${SURVIVOR} to DOWN+REMOVE the dead member (SBR budget ~25s)..."
echo "Standby crash: waiting for ${SURVIVOR} to DOWN+REMOVE the dead member (budget ~25s)..."
BLIPS=0
while true; do
ELAPSED=$(( $(date +%s) - START ))
curl -sf -o /dev/null "${TRAEFIK_URL}/health/active" || BLIPS=$((BLIPS + 1))
if docker logs --since "${KILL_AT}" "${SURVIVOR}" 2>&1 | grep -Eiq "marking.*node.*down|member removed|is removed"; then
echo "PASS: survivor removed the crashed member in ${ELAPSED}s (budget ~25s: 10s detection + 15s stable-after)."
if docker logs --since "${KILL_AT}" "${SURVIVOR}" 2>&1 | grep -Eiq "auto-downing|marking.*node.*down|member removed|is removed"; then
echo "PASS: survivor downed/removed the crashed member in ${ELAPSED}s (budget ~25s: 10s detection + 15s auto-down)."
echo "Active-node routing blips during the drill: ${BLIPS} (expected 0 — the active node was never touched)."
break
fi
if (( ELAPSED > TIMEOUT_S )); then
echo "FAIL: no downing/removal evidence on ${SURVIVOR} after ${ELAPSED}s — SBR did not act" >&2
echo "FAIL: no downing/removal evidence on ${SURVIVOR} after ${ELAPSED}s — auto-down did not act" >&2
docker start "${VICTIM}" > /dev/null
exit 1
fi
sleep 1
done
else
echo "Active crash: EXPECTING a central outage (registered keep-oldest gap). Watching /health/active..."
DARK_STREAK=0
echo "Active crash: waiting for ${SURVIVOR} to take over as the active node (victim stays DOWN; budget ~25s + hand-over)..."
while true; do
ELAPSED=$(( $(date +%s) - START ))
if curl -sf -o /dev/null "${TRAEFIK_URL}/health/active"; then DARK_STREAK=0; else DARK_STREAK=$((DARK_STREAK + 1)); fi
if (( DARK_STREAK >= 10 )); then
echo "Outage confirmed at ${ELAPSED}s: no active central node — the younger survivor self-downed"
echo "(keep-oldest downs the partition WITHOUT the oldest; this is the registered deferred gap)."
if curl -sf -o /dev/null "http://localhost:${SURVIVOR_PORT}/health/active"; then
echo "PASS: ${SURVIVOR} took over as active in ${ELAPSED}s with the victim still down"
echo "(downed the dead oldest via auto-down, assumed Oldest, re-hosted the singletons)."
break
fi
if (( ELAPSED > OUTAGE_CONFIRM_S )); then
echo "NOTE: /health/active stayed reachable ${ELAPSED}s after killing the oldest — better than the"
echo "registered gap predicts. Do NOT celebrate: capture both nodes' logs and investigate before trusting it."
break
if (( ELAPSED > TIMEOUT_S )); then
echo "FAIL: ${SURVIVOR} never became active within ${ELAPSED}s of killing the oldest — takeover did not happen." >&2
docker logs --since "${KILL_AT}" "${SURVIVOR}" 2>&1 | grep -Ei "sbr|downing|oldest|shutting down|terminated" | tail -20 >&2 || true
docker start "${VICTIM}" > /dev/null
exit 1
fi
sleep 1
done
echo "Confirming Traefik routes to the new active node..."
TR_START=$(date +%s)
while ! curl -sf -o /dev/null "${TRAEFIK_URL}/health/active"; do
if (( $(date +%s) - TR_START > 60 )); then
echo "FAIL: survivor is active but not routable through Traefik after 60s" >&2
docker start "${VICTIM}" > /dev/null
exit 1
fi
sleep 1
done
echo "Traefik routing recovered $(( $(date +%s) - START ))s after the kill."
fi
echo "Restarting ${VICTIM}..."
docker start "${VICTIM}" > /dev/null
RESTART_AT=$(date +%s)
echo "Waiting for central to be routable again through Traefik (${TRAEFIK_URL}/health/active)..."
echo "Waiting for the restarted victim to rejoin as a ready standby (${VICTIM} /health/ready)..."
VICTIM_PORT=$(port_of "$VICTIM")
while true; do
ELAPSED=$(( $(date +%s) - RESTART_AT ))
if curl -sf -o /dev/null "${TRAEFIK_URL}/health/active"; then
echo "Recovered: an active central node is routable ${ELAPSED}s after the victim restart."
if curl -sf -o /dev/null "http://localhost:${VICTIM_PORT}/health/ready"; then
echo "Recovered: ${VICTIM} is ready (rejoined as a fresh incarnation) ${ELAPSED}s after restart."
break
fi
if (( ELAPSED > 120 )); then
echo "FAIL: central not routable 120s after restarting ${VICTIM}" >&2
echo "FAIL: ${VICTIM} not ready 120s after restart" >&2
exit 1
fi
sleep 1
done
echo "Survivor singleton/downing evidence (last 20 matching log lines from ${SURVIVOR}):"
docker logs "${SURVIVOR}" 2>&1 | grep -Ei "singleton|oldest|downing|removed" | tail -20 || true
echo "Survivor downing/singleton evidence (last 20 matching log lines from ${SURVIVOR}):"
docker logs "${SURVIVOR}" 2>&1 | grep -Ei "auto-downing|singleton|oldest|downing|removed" | tail -20 || true
echo "Drill complete (${DRILL_MODE}). Verify on the Health dashboard that both nodes show Up and exactly one is Primary."
+1 -1
View File
@@ -14,7 +14,7 @@
"akka.tcp://scadabridge@scadabridge-site-a-a:8082",
"akka.tcp://scadabridge@scadabridge-site-a-b:8082"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
+1 -1
View File
@@ -14,7 +14,7 @@
"akka.tcp://scadabridge@scadabridge-site-a-a:8082",
"akka.tcp://scadabridge@scadabridge-site-a-b:8082"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
+1 -1
View File
@@ -14,7 +14,7 @@
"akka.tcp://scadabridge@scadabridge-site-b-a:8082",
"akka.tcp://scadabridge@scadabridge-site-b-b:8082"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
+1 -1
View File
@@ -14,7 +14,7 @@
"akka.tcp://scadabridge@scadabridge-site-b-a:8082",
"akka.tcp://scadabridge@scadabridge-site-b-b:8082"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
+1 -1
View File
@@ -14,7 +14,7 @@
"akka.tcp://scadabridge@scadabridge-site-c-a:8082",
"akka.tcp://scadabridge@scadabridge-site-c-b:8082"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
+1 -1
View File
@@ -14,7 +14,7 @@
"akka.tcp://scadabridge@scadabridge-site-c-a:8082",
"akka.tcp://scadabridge@scadabridge-site-c-b:8082"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
@@ -48,5 +48,5 @@ Two live items previously tracked ONLY in `archreview/plans/00-MASTER-TRACKER.md
| # | Item | Where noted | Rationale for deferral | Revisit trigger |
|---|------|-------------|------------------------|-----------------|
| SBR | **SBR oldest-crash total-outage gap** — 2-node `keep-oldest` downs the partition *without* the oldest, so a hard crash of the ACTIVE (oldest) central node makes the standby self-down (~10s) → total central outage until the crashed node restarts; only a younger-node crash fails over. | `archreview/plans/00-MASTER-TRACKER.md:194` + auto-memory `sbr-keep-oldest-2node-active-crash-gap` | Remedy is a production SBR topology/strategy decision (keep-majority + a 3rd/lighthouse seed node, static-quorum, or an accepted-risk note) — **owner: user decision**, not silently changeable. | Before the next production deployment that adds a central node, or the first real active-node crash. |
| SBR | ~~**SBR oldest-crash total-outage gap**~~ **RESOLVED 2026-07-21 (owner decision — availability over partition-safety).** All clusters switched from `keep-oldest` to the `auto-down` downing strategy (Akka `AutoDowning`, `auto-down-unreachable-after` = 15s): a hard crash of EITHER node — active/oldest included — now fails over to the survivor in ~25s with no operator action. Accepted trade: a real network partition produces dual-active until an operator restarts one side. Decision record + evidence (live keep-oldest `DownReachable … including myself` log, Akka.NET 1.5.62 `KeepOldest.OldestDecision` source, rejected alternatives incl. the static-quorum-1 `DownAll` trap): `docs/plans/2026-07-21-auto-down-availability-decision.md`. | `archreview/plans/00-MASTER-TRACKER.md:194` + auto-memory `sbr-keep-oldest-2node-active-crash-gap` (both now historical) | — | Closed. Residual: seed-node boot-alone constraint (unchanged, documented in `Component-ClusterInfrastructure.md`); dual-active recovery is operator-driven. |
| vd03 | **`deploy/wonder-app-vd03/` overlay edits unapplied** — `appsettings.Central.json` needs `AllowSingleNodeCluster: true` + phantom-seed removal + `NodeName: central-a`; `install.ps1` needs `sc.exe failure` recovery actions. The `deploy/wonder-app-vd03/` artifact directory is intentionally untracked (production config out of source control), so the repo cannot ship the fix. | `archreview/plans/00-MASTER-TRACKER.md:198` (PLAN-01 T16/T20/T23) | Needs on-host access; without `NodeName` that deployment's audit rows stamp NULL `SourceNode`**partially mitigated once PLAN-R2-08 Task 7 lands: the host now FAILS AT BOOT with a key-naming error instead of silently NULLing, so applying the overlay becomes mandatory at the next upgrade.** Owner: whoever maintains the host (user). | Next wonder-app-vd03 deployment/upgrade — **the Task 7 validator makes this row unskippable then.** |
@@ -0,0 +1,108 @@
# Auto-Down Downing Strategy — Availability Over Partition-Safety (Decision, 2026-07-21)
**Status: DECIDED and implemented (owner decision, 2026-07-21).** Resolves the registered
deferred "keep-oldest topology/strategy" question (master tracker 2026-07-08;
`docs/plans/2026-07-08-deferred-work-register.md` → SBR row).
## The decision
All two-node ScadaBridge clusters (central and every site pair) switch their downing
strategy from the SBR **keep-oldest** resolver to Akka's **`AutoDowning`** provider
(`ClusterOptions.SplitBrainResolverStrategy: "auto-down"`, now the default):
- `downing-provider-class = "Akka.Cluster.AutoDowning, Akka.Cluster"`
- `auto-down-unreachable-after` = `ClusterOptions.StableAfter` (15s production)
The leader among the **reachable** members downs the unreachable peer after the
stability window. Consequence: a hard crash of **either** node — the active/oldest
included — fails over to the survivor in ~25s (10s failure detection + 15s window),
with no operator action and no victim restart required.
**The owner's stated rationale, verbatim in effect:** the pairs run one node per VM at
each site with no Kubernetes and no SQL available site-side, and "network partitions are
less of a risk than if this stops working." Availability wins.
## The accepted trade (read this before debugging a dual-active)
In a **real network partition** (both nodes alive, link cut) each side downs the other
and continues as a one-node cluster: **both run active** — two oldest-Up members, two
sets of singletons, `/health/active` = 200 on both. The pre-decision keep-oldest
resolver would instead have sacrificed the younger side. Recovery from dual-active is
operator-driven: after the partition heals, restart ONE side; the restarted node rejoins
its peer as a fresh incarnation and becomes standby. (The two sides do not merge on
their own — the mutual downing quarantines the association.)
## Why the crashed-oldest direction was unsurvivable before (evidence)
Live drill on the docker rig, 2026-07-21, `keep-oldest` + `down-if-alone = on`
(config verified live): killing the active/oldest `central-a` produced, on `central-b`:
```
SBR took decision Akka.Cluster.SBR.DownReachable and is downing
[akka.tcp://scadabridge@scadabridge-central-b:8081] including myself,
[1] unreachable of [2] members
```
The survivor downed ITSELF, exited (`run-coordinated-shutdown-when-down`), and its
restarted incarnation looped on `InitJoin` (non-first-seed cannot self-form) until the
victim returned. Root cause in Akka.NET 1.5.62 `KeepOldest.OldestDecision`
(`src/core/Akka.Cluster/SBR/DowningStrategy.cs`):
```csharp
// oldest is on the OTHER (unreachable) side:
if (DownIfAlone && otherSide == 1 && thisSide >= 2) // survivor side must be >= 2
return DownUnreachable.Instance;
return DownReachable.Instance; // 1-vs-1 → down MYSELF
```
`down-if-alone` is designed for ≥3-node clusters; with 1-vs-1 it deliberately keeps the
oldest side ("the node on the other side is no better" — upstream comment). So two-node
keep-oldest can never survive an oldest crash. This corrected an earlier
mis-explanation in the repo ("the alone-oldest is dead and cannot down itself").
## Alternatives rejected
| Option | Why not |
|---|---|
| keep-oldest (status quo) | Oldest crash = total outage (proven above). Remains a supported `SplitBrainResolverStrategy` value for deployments preferring partition-safety. |
| static-quorum, quorum 1 | Akka's `IsTooManyMembers` guard (`2 > 2*1-1`) returns **DownAll** on any unreachability — total shutdown, strictly worse. |
| static-quorum, quorum 2 | Survivor (1 < 2) downs itself on any crash. |
| keep-majority | 1-vs-1 tie keeps the lowest-address side — moves the fatal crash from "oldest" to "lowest address", same hole. |
| lease-majority | Needs a shared lease store (K8s API, SQL, …) reachable by both nodes — not available at sites. |
| third arbiter node | Would make `down-if-alone` work, but there is no third VM at sites. |
| custom downing provider | Would reimplement exactly what `AutoDowning` already does, tested upstream. If a future Akka.NET release removes `AutoDowning`, port it then. |
## What changed (implementation slice, same session)
- `ClusterOptions`: `SplitBrainResolverStrategy` default → `"auto-down"`; docs rewritten.
`DownIfAlone` kept (keep-oldest-only knob, validated only under keep-oldest).
- `ClusterOptionsValidator`: allows `auto-down` | `keep-oldest`; `DownIfAlone` requirement
scoped to keep-oldest.
- `AkkaHostedService.BuildHocon`: downing block branches on the strategy (AutoDowning
provider + `auto-down-unreachable-after` vs the SBR block).
- All 16 `appsettings` (src Host ×2, `docker/` ×8, `docker-env2/` ×4, and the gitignored
`deploy/wonder-app-vd03/` ×2 on-disk overlay) flipped to `auto-down`.
**Owner action: sync the wonder-app-vd03 overlay to the host and restart both services
together.**
- `docker/failover-drill.sh`: `active` mode now asserts the survivor TAKES OVER while
the victim is down (previously it asserted the outage).
- Tests: HOCON emission (`HoconBuilderTests`), validator/default tests, and two new
real-cluster tests in `SbrFailoverTests``AutoDown_HardCrashOfOldestNode_
YoungerSurvivorTakesOverSingleton` (the direction keep-oldest could never pass) and
`AutoDown_HardCrashOfYoungerNode_OldestKeepsSingleton`. `TwoNodeClusterFixture` gained
a `strategy` parameter (default `auto-down`).
- Docs: `Component-ClusterInfrastructure.md` (Downing Strategy section rewritten),
`docker/README.md` (drill docs + results), `CLAUDE.md`, deferred-work register entry
resolved.
## Residual operational notes
- **Seed-node bootstrap constraint still applies to boot-alone**: only the first seed
may self-form a cluster. Auto-down removes the active-crash outage (the survivor never
restarts), but a node that must BOOT alone while its peer is dead (cold start of only
the non-first-seed VM, or the survivor crashing while the peer is still down) still
waits in `InitJoin` for its peer. Operator recovery unchanged (restart first seed, or
self-first seed override).
- Monitoring already surfaces dual-active if it ever happens: both nodes report
`IsActive` in heartbeats / both `/health/active` = 200 — the Health dashboard shows
two Primaries.
@@ -92,27 +92,33 @@ Akka.NET cluster singletons run on the active node of their cluster and migrate
- Health reporting resumes from the new active node.
- Alarm states are re-evaluated from incoming values (alarm state is in-memory only).
## Split-Brain Resolution
## Downing Strategy (auto-down — availability-first)
The system uses the Akka.NET **keep-oldest** split-brain resolver strategy:
**Decision 2026-07-21 (owner decision, resolves the deferred keep-oldest topology/strategy question):** the clusters run the **`auto-down`** downing strategy — Akka's `AutoDowning` provider with `auto-down-unreachable-after` = `StableAfter` (15s). The leader among the *reachable* members downs the unreachable peer after the stability window:
- On a network partition, the node that has been in the cluster longest remains active. The younger node downs itself.
- **Stable-after duration**: 15 seconds. The cluster membership must remain stable (no changes) for 15 seconds before the resolver acts to down unreachable nodes. This prevents premature downing during startup or rolling restarts.
- **`down-if-alone = on`**: The keep-oldest resolver is configured with `down-if-alone` enabled. If the oldest node finds itself alone (no other reachable members), it downs itself rather than continuing as a single-node cluster. This prevents the oldest node from running in isolation during a network partition while the younger node also forms its own cluster.
- **Why keep-oldest**: With only two nodes, quorum-based strategies (static-quorum, keep-majority) cannot distinguish "one node crashed" from "network partition" — both sides see fewer than quorum and both would down themselves, resulting in total cluster shutdown. Keep-oldest with `down-if-alone` provides safe singleton ownership — at most one node runs the cluster singleton at any time.
- **Either-node crash is survivable.** If the standby crashes, the active node downs it and continues (as before). If the **active/oldest** node crashes, the younger survivor becomes leader among the reachable members, downs the dead oldest, becomes the oldest itself, re-hosts every cluster singleton, and `/health/active` flips to it — **no operator action and no victim restart required**. This closes the keep-oldest total-outage gap.
- **The accepted trade: dual-active during a real network partition.** With both nodes alive but partitioned, each side downs the other and continues as a one-node cluster — both claim active until the partition heals and an operator restarts one side (the restarted node rejoins the other as a fresh incarnation). This trade was chosen deliberately: site pairs run one node per VM with no shared lease infrastructure (no Kubernetes, no SQL at sites) to arbitrate, and a stalled system is a bigger operational risk than a rare LAN partition.
- **Stable-after duration**: 15 seconds of sustained unreachability before downing. This prevents premature downing during startup, rolling restarts, or transient network blips.
### Down-if-alone recovery
### Why not the alternatives (all verified against Akka.NET 1.5.62 source, 2026-07-21)
When a node downs itself (via `down-if-alone`, or any other SBR decision), the resolver is configured with `run-coordinated-shutdown-when-down = on`, so the self-down runs `CoordinatedShutdown` and **terminates that node's `ActorSystem`**. The Host process must not keep running with a dead actor system — it would serve nothing and be restarted by nobody. The recovery contract is:
- **keep-oldest** (used until 2026-07-21): partition-safe, but in a two-node cluster a crash of the **oldest** is a total outage. `KeepOldest.OldestDecision` only lets `down-if-alone` rescue the survivor when the surviving side has **≥ 2 members** (`otherSide == 1 && thisSide >= 2`); with 1-vs-1 the younger survivor takes `DownReachable` — it downs *itself*. Verified live on the docker rig: the survivor logged `SBR took decision Akka.Cluster.SBR.DownReachable … including myself`, exited, and looped on `InitJoin`. The strategy remains supported in `ClusterOptions` (`SplitBrainResolverStrategy: keep-oldest`) for deployments that prefer partition-safety over availability.
- **static-quorum (quorum-size 1)**: Akka's `IsTooManyMembers` guard (`members > quorum*2-1`, i.e. `2 > 1`) returns **DownAll** on any unreachability — total shutdown, strictly worse.
- **keep-majority**: a 1-vs-1 split keeps the side with the lowest address, which just moves the fatal crash from "the oldest" to "the lowest-address node".
- **lease-majority**: needs a shared lease store (Kubernetes API, SQL, …) reachable from both nodes — not available at sites (one node per VM, no shared infrastructure).
1. Self-down ⇒ `CoordinatedShutdown``ActorSystem` termination.
### Downed-node recovery
When a node is downed (auto-downed by its peer after a partition heals, or a keep-oldest self-down where that strategy is configured), `run-coordinated-shutdown-when-down = on` runs `CoordinatedShutdown` and **terminates that node's `ActorSystem`**. The Host process must not keep running with a dead actor system — it would serve nothing and be restarted by nobody. The recovery contract is:
1. Down ⇒ `CoordinatedShutdown``ActorSystem` termination.
2. The Host watches `ActorSystem.WhenTerminated`; a termination that is **not** the host's own `StopAsync` triggers `IHostApplicationLifetime.StopApplication()`, so **the process exits**.
3. The service supervisor restarts it — docker `restart: unless-stopped`, or Windows service recovery actions (`sc.exe failure … restart/…`).
4. The restarted process rejoins as a **fresh incarnation** (the keep-oldest resolver handles the rejoin cleanly; there is no stale membership to reconcile)**but only while a peer still holding cluster state is reachable**. A lone restarted node that is *not* the first seed cannot re-form a cluster on its own (see the seed-node bootstrap constraint below); it waits for its peer.
4. The restarted process rejoins as a **fresh incarnation****but only while a peer still holding cluster state is reachable**. A lone restarted node that is *not* the first seed cannot re-form a cluster on its own (see the seed-node bootstrap constraint below); it waits for its peer.
**Seed-node bootstrap constraint.** Only the FIRST seed listed in `Cluster:SeedNodes` may self-join to form a *new* cluster. All nodes list the same first seed (e.g. `scadabridge-central-a`), so a lone restarted non-first-seed node (with the first seed still down) loops on `InitJoin` forever — never `Up`, never routable. This is why the two-node keep-oldest **oldest/active-node crash is a total-outage gap**: after the oldest dies the younger survivor self-downs, and it cannot re-bootstrap alone. Recovery is operator-driven — either restart the dead first-seed node (preferred) or restart the survivor with a self-first seed override (`ScadaBridge__Cluster__SeedNodes__0` = self, `__1` = peer). The repo does not ship self-first ordering per node: with both nodes self-first a simultaneous cold start risks two independent one-node clusters that never merge. Removing the gap itself is the **registered deferred keep-oldest topology/strategy decision** (master tracker 2026-07-08, owner: user).
**Seed-node bootstrap constraint (still applies).** Only the FIRST seed listed in `Cluster:SeedNodes` may self-join to form a *new* cluster. All nodes list the same first seed (e.g. `scadabridge-central-a`), so a lone restarted non-first-seed node (with the first seed still down) loops on `InitJoin` forever — never `Up`, never routable. Under auto-down this constraint no longer causes the active-crash outage (the survivor keeps running and never restarts), but it still bites when a node must **boot alone** — e.g. a cold start of only the non-first-seed VM, or the survivor itself crashing while its peer is still dead. Recovery is operator-driven — either restart the first-seed node (preferred) or restart the survivor with a self-first seed override (`ScadaBridge__Cluster__SeedNodes__0` = self, `__1` = peer). The repo does not ship self-first ordering per node: with both nodes self-first a simultaneous cold start risks two independent one-node clusters that never merge.
The docker failover drill (`docker/failover-drill.sh`) exercises this per direction: `standby` mode proves SBR downing + singleton continuity on the oldest; `active` mode measures the registered total-outage gap and the recovery-on-restart path.
The docker failover drill (`docker/failover-drill.sh`) proves both directions: `standby` mode kills the younger node (active untouched, zero routing blips); `active` mode kills the active/oldest node and asserts the survivor **takes over while the victim is still down**.
## Single-Node Operation
@@ -38,12 +38,27 @@ public class ClusterOptions
public List<string> SeedNodes { get; set; } = new();
/// <summary>
/// Split-brain resolver strategy. Must be <c>keep-oldest</c> for the two-node
/// clusters ScadaBridge uses: quorum strategies (<c>keep-majority</c>,
/// <c>static-quorum</c>) cannot distinguish a crash from a partition with only
/// two nodes and would shut down the whole cluster.
/// Downing strategy for unreachable members. Two supported values:
/// <list type="bullet">
/// <item><c>auto-down</c> (default, decision 2026-07-21) — availability-first: each
/// side downs the unreachable peer after <see cref="StableAfter"/>, so a hard crash
/// of EITHER node (oldest included) fails over to the survivor. The accepted trade:
/// a true network partition produces two live one-node clusters (dual-active) until
/// an operator restarts one side. Chosen because ScadaBridge pairs run one node per
/// VM with no shared lease infrastructure, and a stalled system is a bigger risk
/// than a rare partition.</item>
/// <item><c>keep-oldest</c> — partition-safe SBR: downs the side without the oldest
/// member. In a TWO-node cluster this makes a crash of the oldest/active node a
/// total outage: Akka's <c>down-if-alone</c> only rescues the survivor when its own
/// side has ≥2 members (verified against Akka.NET 1.5.62 <c>KeepOldest.Decide</c>
/// and live on the docker rig, 2026-07-21).</item>
/// </list>
/// Other SBR strategies are rejected: <c>static-quorum</c> with quorum 1 hits Akka's
/// <c>IsTooManyMembers</c> guard (2 &gt; 2*1-1) and downs ALL on any unreachability;
/// <c>keep-majority</c> just moves the fatal crash from the oldest to the
/// lowest-address node.
/// </summary>
public string SplitBrainResolverStrategy { get; set; } = "keep-oldest";
public string SplitBrainResolverStrategy { get; set; } = "auto-down";
/// <summary>
/// Time the cluster membership must remain stable before the split-brain
@@ -71,9 +86,12 @@ public class ClusterOptions
public int MinNrOfMembers { get; set; } = 1;
/// <summary>
/// The keep-oldest resolver's <c>down-if-alone</c> flag. When <c>true</c> (the
/// design-doc requirement), the oldest node downs itself if it finds it has no
/// other reachable members, rather than running as an isolated single-node cluster.
/// The keep-oldest resolver's <c>down-if-alone</c> flag; only consulted when
/// <see cref="SplitBrainResolverStrategy"/> is <c>keep-oldest</c>. When <c>true</c>,
/// the oldest node downs itself if it finds it has no other reachable members,
/// rather than running as an isolated single-node cluster. Note that in a two-node
/// cluster this does NOT let the younger survivor take over from a crashed oldest —
/// Akka's alone-check requires the surviving side to have ≥2 members.
/// </summary>
public bool DownIfAlone { get; set; } = true;
@@ -12,9 +12,18 @@ namespace ZB.MOM.WW.ScadaBridge.ClusterInfrastructure;
/// </summary>
public sealed class ClusterOptionsValidator : OptionsValidatorBase<ClusterOptions>
{
/// <summary>Split-brain resolver strategies safe for ScadaBridge's two-node clusters.</summary>
/// <summary>
/// Downing strategies supported for ScadaBridge's two-node clusters.
/// <c>auto-down</c> (default) survives a crash of either node at the accepted cost
/// of dual-active during a real partition; <c>keep-oldest</c> is partition-safe but
/// cannot survive a crash of the oldest node. Quorum strategies are rejected:
/// <c>static-quorum</c> quorum-size 1 trips Akka's IsTooManyMembers guard (DownAll
/// on any unreachability in a 2-node cluster) and <c>keep-majority</c> keys the
/// fatal crash to the lowest-address node instead of the oldest.
/// </summary>
private static readonly HashSet<string> AllowedStrategies = new(StringComparer.OrdinalIgnoreCase)
{
"auto-down",
"keep-oldest"
};
@@ -37,8 +46,9 @@ public sealed class ClusterOptionsValidator : OptionsValidatorBase<ClusterOption
builder.RequireThat(
!string.IsNullOrWhiteSpace(options.SplitBrainResolverStrategy)
&& AllowedStrategies.Contains(options.SplitBrainResolverStrategy),
$"ClusterOptions.SplitBrainResolverStrategy must be 'keep-oldest' for a two-node cluster; " +
$"'{options.SplitBrainResolverStrategy}' would risk a total cluster shutdown on a partition.");
$"ClusterOptions.SplitBrainResolverStrategy must be 'auto-down' or 'keep-oldest' for a " +
$"two-node cluster; '{options.SplitBrainResolverStrategy}' would risk a total cluster " +
"shutdown on a partition or an unreachability event.");
builder.RequireThat(options.MinNrOfMembers == 1,
$"ClusterOptions.MinNrOfMembers must be 1 (was {options.MinNrOfMembers}); " +
@@ -58,7 +68,11 @@ public sealed class ClusterOptionsValidator : OptionsValidatorBase<ClusterOption
$"FailureDetectionThreshold ({options.FailureDetectionThreshold}); otherwise nodes are " +
"declared unreachable before a heartbeat can arrive.");
builder.RequireThat(options.DownIfAlone,
// DownIfAlone is a keep-oldest knob; under auto-down each side downs the
// unreachable peer regardless, so the flag is inert and any value is fine.
var isKeepOldest = string.Equals(
options.SplitBrainResolverStrategy, "keep-oldest", StringComparison.OrdinalIgnoreCase);
builder.RequireThat(!isKeepOldest || options.DownIfAlone,
"ClusterOptions.DownIfAlone must be true for the keep-oldest resolver "
+ "(Component-ClusterInfrastructure.md → Split-Brain Resolution); with it false the "
+ "oldest node can run as an isolated single-node cluster during a partition while the "
@@ -200,8 +200,9 @@ public class AkkaHostedService : IHostedService
_communicationOptions.TransportHeartbeatInterval.TotalSeconds,
_communicationOptions.TransportFailureThreshold.TotalSeconds);
// Down-if-alone recovery watchdog: SBR's keep-oldest down-if-alone plus
// run-coordinated-shutdown-when-down means a self-downed node terminates
// Downed-node recovery watchdog: any downing decision against this node
// (auto-down by the peer, or an SBR self-down under keep-oldest) plus
// run-coordinated-shutdown-when-down means a downed node terminates
// its own ActorSystem. If that happens outside our StopAsync, the Host
// process must exit so the service supervisor (docker
// `restart: unless-stopped` / Windows service recovery) restarts it and
@@ -228,14 +229,21 @@ public class AkkaHostedService : IHostedService
/// seed-node URI, role or split-brain strategy containing a quote, backslash or
/// whitespace cannot corrupt the document or be silently misparsed.
///
/// The <c>keep-oldest down-if-alone</c> flag is emitted from
/// The downing block branches on <see cref="ClusterOptions.SplitBrainResolverStrategy"/>:
/// <c>auto-down</c> (default; decision 2026-07-21) installs Akka's
/// <c>AutoDowning</c> provider with <c>auto-down-unreachable-after</c> =
/// <see cref="ClusterOptions.StableAfter"/> — the leader among the REACHABLE members
/// downs the unreachable peer, so a crash of either node (oldest included) fails
/// over to the survivor; the accepted trade is dual-active during a real network
/// partition. Any other value takes the SBR path, where the
/// <c>keep-oldest down-if-alone</c> flag is emitted from
/// <see cref="ClusterOptions.DownIfAlone"/> rather than hard-coded, so the bound
/// configuration value is actually consumed.
///
/// The split-brain-resolver <c>downing-provider-class</c> is installed
/// explicitly: Akka defaults to <c>NoDowning</c>, under which the entire
/// split-brain-resolver section is inert and singletons never migrate on a hard
/// crash or partition. Naming the SBR provider is what activates automatic downing.
/// A <c>downing-provider-class</c> is always installed explicitly: Akka defaults
/// to <c>NoDowning</c>, under which the downing configuration is inert and
/// singletons never migrate on a hard crash or partition. Naming the provider is
/// what activates automatic downing.
///
/// Every duration is rendered via <see cref="DurationHocon"/> in
/// milliseconds, so sub-second cluster timing values (e.g. a 750ms heartbeat) are
@@ -258,6 +266,25 @@ public class AkkaHostedService : IHostedService
clusterOptions.SeedNodes.Select(QuoteHocon));
var rolesStr = string.Join(",", roles.Select(QuoteHocon));
// auto-down (default): AutoDowning provider — the leader among the reachable
// members downs the unreachable peer after StableAfter, so a crash of EITHER
// node fails over to the survivor (dual-active during a real partition is the
// accepted trade — decision 2026-07-21). Anything else: the SBR provider with
// the configured active-strategy (keep-oldest), which is partition-safe but
// cannot survive a crash of the oldest node in a two-node cluster.
var downingBlock = string.Equals(
clusterOptions.SplitBrainResolverStrategy, "auto-down", StringComparison.OrdinalIgnoreCase)
? $@"downing-provider-class = ""Akka.Cluster.AutoDowning, Akka.Cluster""
auto-down-unreachable-after = {DurationHocon(clusterOptions.StableAfter)}"
: $@"downing-provider-class = ""Akka.Cluster.SBR.SplitBrainResolverProvider, Akka.Cluster""
split-brain-resolver {{
active-strategy = {QuoteHocon(clusterOptions.SplitBrainResolverStrategy)}
stable-after = {DurationHocon(clusterOptions.StableAfter)}
keep-oldest {{
down-if-alone = {(clusterOptions.DownIfAlone ? "on" : "off")}
}}
}}";
return $@"
audit-telemetry-dispatcher {{
type = ForkJoinDispatcher
@@ -287,14 +314,7 @@ akka {{
seed-nodes = [{seedNodesStr}]
roles = [{rolesStr}]
min-nr-of-members = {clusterOptions.MinNrOfMembers}
downing-provider-class = ""Akka.Cluster.SBR.SplitBrainResolverProvider, Akka.Cluster""
split-brain-resolver {{
active-strategy = {QuoteHocon(clusterOptions.SplitBrainResolverStrategy)}
stable-after = {DurationHocon(clusterOptions.StableAfter)}
keep-oldest {{
down-if-alone = {(clusterOptions.DownIfAlone ? "on" : "off")}
}}
}}
{downingBlock}
failure-detector {{
heartbeat-interval = {DurationHocon(clusterOptions.HeartbeatInterval)}
acceptable-heartbeat-pause = {DurationHocon(clusterOptions.FailureDetectionThreshold)}
@@ -13,7 +13,7 @@
"akka.tcp://scadabridge@localhost:8081",
"akka.tcp://scadabridge@localhost:8082"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
@@ -16,7 +16,7 @@
"akka.tcp://scadabridge@localhost:8082",
"akka.tcp://scadabridge@localhost:8085"
],
"SplitBrainResolverStrategy": "keep-oldest",
"SplitBrainResolverStrategy": "auto-down",
"StableAfter": "00:00:15",
"HeartbeatInterval": "00:00:02",
"FailureDetectionThreshold": "00:00:10",
@@ -10,7 +10,10 @@ public class ClusterOptionsTests
{
var options = new ClusterOptions();
Assert.Equal("keep-oldest", options.SplitBrainResolverStrategy);
// 'auto-down' is the default posture (decision 2026-07-21): a crash of either
// node fails over to the survivor; dual-active during a real partition is the
// accepted trade for pairs with no shared lease infrastructure.
Assert.Equal("auto-down", options.SplitBrainResolverStrategy);
Assert.Equal(TimeSpan.FromSeconds(15), options.StableAfter);
Assert.Equal(TimeSpan.FromSeconds(2), options.HeartbeatInterval);
Assert.Equal(TimeSpan.FromSeconds(10), options.FailureDetectionThreshold);
@@ -153,9 +153,10 @@ public class ClusterOptionsValidatorTests
}
[Fact]
public void DownIfAloneFalse_FailsValidation()
public void DownIfAloneFalse_UnderKeepOldest_FailsValidation()
{
var options = ValidOptions();
options.SplitBrainResolverStrategy = "keep-oldest";
options.DownIfAlone = false;
var result = new ClusterOptionsValidator().Validate(null, options);
@@ -164,6 +165,34 @@ public class ClusterOptionsValidatorTests
Assert.Contains("DownIfAlone", result.FailureMessage);
}
[Fact]
public void AutoDownStrategy_Passes()
{
// Decision 2026-07-21: 'auto-down' is the supported availability-first
// posture — either-node crash fails over; dual-active on a real partition
// is the accepted trade.
var options = ValidOptions();
options.SplitBrainResolverStrategy = "auto-down";
var result = new ClusterOptionsValidator().Validate(null, options);
Assert.True(result.Succeeded, result.FailureMessage);
}
[Fact]
public void DownIfAloneFalse_UnderAutoDown_Passes()
{
// DownIfAlone is a keep-oldest knob; under auto-down it is inert and must
// not block startup.
var options = ValidOptions();
options.SplitBrainResolverStrategy = "auto-down";
options.DownIfAlone = false;
var result = new ClusterOptionsValidator().Validate(null, options);
Assert.True(result.Succeeded, result.FailureMessage);
}
[Fact]
public void Validate_AccumulatesAllFailures()
{
@@ -195,4 +195,49 @@ public class HoconBuilderTests
"Akka.Cluster.SBR.SplitBrainResolverProvider, Akka.Cluster",
config.GetString("akka.cluster.downing-provider-class"));
}
[Fact]
public void BuildHocon_AutoDownStrategy_EmitsAutoDowningProvider()
{
// Decision 2026-07-21 (availability over partition-safety): 'auto-down' must
// swap the downing provider to Akka's AutoDowning so the survivor downs a
// crashed peer — including a crashed OLDEST, which two-node keep-oldest
// cannot survive — after StableAfter.
var cluster = DefaultCluster();
cluster.SplitBrainResolverStrategy = "auto-down";
cluster.StableAfter = TimeSpan.FromSeconds(15);
var hocon = AkkaHostedService.BuildHocon(
DefaultNode(), cluster, new[] { "Central" },
TimeSpan.FromSeconds(5), TimeSpan.FromSeconds(15));
var config = ConfigurationFactory.ParseString(hocon);
Assert.Equal(
"Akka.Cluster.AutoDowning, Akka.Cluster",
config.GetString("akka.cluster.downing-provider-class"));
Assert.Equal(
TimeSpan.FromSeconds(15),
config.GetTimeSpan("akka.cluster.auto-down-unreachable-after"));
// The SBR section must NOT be active alongside AutoDowning.
Assert.False(config.HasPath("akka.cluster.split-brain-resolver.active-strategy"));
}
[Fact]
public void BuildHocon_AutoDownStrategy_IsCaseInsensitive_AndDocumentStaysIntact()
{
var cluster = DefaultCluster();
cluster.SplitBrainResolverStrategy = "Auto-Down";
var hocon = AkkaHostedService.BuildHocon(
DefaultNode(), cluster, new[] { "Central" },
TimeSpan.FromSeconds(5), TimeSpan.FromSeconds(15));
var config = ConfigurationFactory.ParseString(hocon);
Assert.Equal(
"Akka.Cluster.AutoDowning, Akka.Cluster",
config.GetString("akka.cluster.downing-provider-class"));
// Keys after the downing block must remain intact (document not corrupted).
Assert.Equal(1, config.GetInt("akka.cluster.min-nr-of-members"));
Assert.True(config.GetBoolean("akka.cluster.run-coordinated-shutdown-when-down"));
}
}
@@ -5,22 +5,26 @@ using Xunit;
namespace ZB.MOM.WW.ScadaBridge.IntegrationTests.Cluster;
/// <summary>
/// Behavioral proof that the SBR downing provider enabled in
/// Behavioral proof that the downing provider enabled in
/// <c>AkkaHostedService.BuildHocon</c> (arch-review 01 Critical) is actually active:
/// after a hard crash, the surviving node DOWNS and REMOVES the crashed member. Under
/// the pre-fix Akka default (NoDowning) the crashed member lingers <c>Unreachable</c>
/// forever, so the member-removal assertions here are impossible to satisfy without the
/// fix — that is what gives the test teeth.
/// fix — that is what gives the tests teeth.
///
/// IMPORTANT — <c>keep-oldest</c> two-node semantics (verified empirically, Akka 1.5.62):
/// SBR downs the partition that does NOT contain the oldest member. So the crash that
/// SBR can recover from in a two-node cluster is the crash of the YOUNGER node — the
/// oldest survives and keeps its singletons. Crashing the OLDEST node instead makes the
/// younger survivor down ITSELF (total cluster loss); <c>down-if-alone=on</c> does not
/// change this on a hard crash because the alone-oldest is no longer running to down
/// itself. That asymmetry (active/oldest-node crash is NOT covered by two-node
/// keep-oldest) is a design-level gap tracked separately, not something this test can
/// assert as a success path.
/// TWO-NODE SEMANTICS (verified against Akka.NET 1.5.62 source + live on the docker
/// rig, 2026-07-21):
/// <list type="bullet">
/// <item><c>keep-oldest</c> — downs the side that does NOT contain the oldest member,
/// and its <c>down-if-alone</c> escape only fires when the surviving side has ≥2
/// members (<c>KeepOldest.OldestDecision</c>: <c>otherSide == 1 &amp;&amp; thisSide &gt;= 2</c>).
/// With 1-vs-1 the younger survivor therefore takes <c>DownReachable</c> — it downs
/// ITSELF — so only a YOUNGER-node crash is survivable.</item>
/// <item><c>auto-down</c> (production default, decision 2026-07-21) — the leader among
/// the reachable members downs the unreachable peer after the stability window, so a
/// crash of EITHER node fails over to the survivor; the accepted trade is dual-active
/// during a real network partition.</item>
/// </list>
/// </summary>
public class SbrFailoverTests
{
@@ -46,7 +50,8 @@ public class SbrFailoverTests
[Fact]
public async Task HardCrashOfYoungerNode_SbrDownsIt_AndOldestKeepsSingleton()
{
await using var cluster = await TwoNodeClusterFixture.StartAsync();
// Pinned to keep-oldest: this is the SBR path's (only) survivable direction.
await using var cluster = await TwoNodeClusterFixture.StartAsync(strategy: "keep-oldest");
var (_, proxyA) = StartSingleton(cluster.NodeA); // oldest hosts the singleton
StartSingleton(cluster.NodeB);
@@ -78,4 +83,80 @@ public class SbrFailoverTests
}
throw new Xunit.Sdk.XunitException($"Singleton stopped answering on the surviving oldest node after SBR downing: {last}");
}
[Fact]
public async Task AutoDown_HardCrashOfOldestNode_YoungerSurvivorTakesOverSingleton()
{
// Decision 2026-07-21: the direction two-node keep-oldest can NEVER survive
// (proven live on the docker rig — the younger survivor took DownReachable and
// self-downed). Under auto-down the survivor must instead down the crashed
// oldest and TAKE OVER its singleton.
await using var cluster = await TwoNodeClusterFixture.StartAsync(strategy: "auto-down");
StartSingleton(cluster.NodeA); // oldest hosts the singleton initially
var (_, proxyB) = StartSingleton(cluster.NodeB);
// Singleton reachable from B while A is alive (proxy routes to the oldest).
var echo = await proxyB.Ask<string>("ping", TimeSpan.FromSeconds(20));
Assert.Equal("ping", echo);
var victimAddress = Akka.Cluster.Cluster.Get(cluster.NodeA).SelfAddress;
await TwoNodeClusterFixture.CrashNode(cluster.NodeA);
// 1) The younger survivor must DOWN and REMOVE the crashed OLDEST member —
// the exact step keep-oldest refuses (it downs itself instead).
await TwoNodeClusterFixture.WaitForMemberRemoved(
cluster.NodeB, victimAddress, TimeSpan.FromSeconds(30));
// 2) B must still be a functioning cluster member (not self-downed) …
var clusterB = Akka.Cluster.Cluster.Get(cluster.NodeB);
Assert.False(clusterB.IsTerminated, "survivor's Cluster extension terminated — it downed itself");
// 3) … and the singleton must migrate to B and answer again.
var deadline = DateTime.UtcNow + TimeSpan.FromSeconds(30);
Exception? last = null;
while (DateTime.UtcNow < deadline)
{
try
{
var echo2 = await proxyB.Ask<string>("ping-after-oldest-crash", TimeSpan.FromSeconds(3));
Assert.Equal("ping-after-oldest-crash", echo2);
return;
}
catch (Exception ex) { last = ex; }
}
throw new Xunit.Sdk.XunitException(
$"Singleton never migrated to the younger survivor after the oldest crashed under auto-down: {last}");
}
[Fact]
public async Task AutoDown_HardCrashOfYoungerNode_OldestKeepsSingleton()
{
// The previously-survivable direction must STAY survivable under auto-down.
await using var cluster = await TwoNodeClusterFixture.StartAsync(strategy: "auto-down");
var (_, proxyA) = StartSingleton(cluster.NodeA);
StartSingleton(cluster.NodeB);
Assert.Equal("ping", await proxyA.Ask<string>("ping", TimeSpan.FromSeconds(20)));
var victimAddress = Akka.Cluster.Cluster.Get(cluster.NodeB).SelfAddress;
await TwoNodeClusterFixture.CrashNode(cluster.NodeB);
await TwoNodeClusterFixture.WaitForMemberRemoved(
cluster.NodeA, victimAddress, TimeSpan.FromSeconds(30));
var deadline = DateTime.UtcNow + TimeSpan.FromSeconds(30);
Exception? last = null;
while (DateTime.UtcNow < deadline)
{
try
{
var echo2 = await proxyA.Ask<string>("ping-after-crash", TimeSpan.FromSeconds(3));
Assert.Equal("ping-after-crash", echo2);
return;
}
catch (Exception ex) { last = ex; }
}
throw new Xunit.Sdk.XunitException(
$"Singleton stopped answering on the surviving oldest node under auto-down: {last}");
}
}
@@ -28,22 +28,26 @@ public sealed class TwoNodeClusterFixture : IAsyncDisposable
public static async Task<TwoNodeClusterFixture> StartAsync(
string role = "Central", TimeSpan? stableAfter = null,
int? portA = null, int? portB = null,
TimeSpan? heartbeatInterval = null, TimeSpan? failureDetectionThreshold = null)
TimeSpan? heartbeatInterval = null, TimeSpan? failureDetectionThreshold = null,
string strategy = "auto-down")
{
var f = new TwoNodeClusterFixture();
f.PortA = portA ?? GetFreeTcpPort();
f.PortB = portB ?? GetFreeTcpPort();
f.NodeA = f.StartNode(f.PortA, role, stableAfter, heartbeatInterval, failureDetectionThreshold);
f.NodeA = f.StartNode(f.PortA, role, stableAfter, heartbeatInterval, failureDetectionThreshold, strategy);
await WaitForMembersUp(f.NodeA, 1, TimeSpan.FromSeconds(20));
f.NodeB = f.StartNode(f.PortB, role, stableAfter, heartbeatInterval, failureDetectionThreshold);
f.NodeB = f.StartNode(f.PortB, role, stableAfter, heartbeatInterval, failureDetectionThreshold, strategy);
await WaitForMembersUp(f.NodeA, 2, TimeSpan.FromSeconds(20));
await WaitForMembersUp(f.NodeB, 2, TimeSpan.FromSeconds(20));
return f;
}
/// <summary>Starts a node from production HOCON; used by StartAsync and by restart-scenarios.</summary>
/// <summary>Starts a node from production HOCON; used by StartAsync and by restart-scenarios.
/// <paramref name="strategy"/> defaults to the production posture (auto-down, decision
/// 2026-07-21); pass "keep-oldest" to exercise the legacy SBR path.</summary>
public ActorSystem StartNode(int port, string role, TimeSpan? stableAfter = null,
TimeSpan? heartbeatInterval = null, TimeSpan? failureDetectionThreshold = null)
TimeSpan? heartbeatInterval = null, TimeSpan? failureDetectionThreshold = null,
string strategy = "auto-down")
{
var nodeOptions = new NodeOptions { Role = role, NodeHostname = "127.0.0.1", RemotingPort = port };
var clusterOptions = new ClusterOptions
@@ -53,7 +57,7 @@ public sealed class TwoNodeClusterFixture : IAsyncDisposable
$"akka.tcp://scadabridge@127.0.0.1:{PortA}",
$"akka.tcp://scadabridge@127.0.0.1:{PortB}",
},
SplitBrainResolverStrategy = "keep-oldest",
SplitBrainResolverStrategy = strategy,
StableAfter = stableAfter ?? TimeSpan.FromSeconds(3),
HeartbeatInterval = heartbeatInterval ?? TimeSpan.FromMilliseconds(500),
FailureDetectionThreshold = failureDetectionThreshold ?? TimeSpan.FromSeconds(2),
@@ -12,13 +12,14 @@ namespace ZB.MOM.WW.ScadaBridge.PerformanceTests.Failover;
/// 2s heartbeat / 10s failure-detection threshold / 15s SBR stable-after —
/// the CLAUDE.md "total failover ~25s" design envelope.
///
/// Measures the SURVIVABLE direction only: hard-crash of the YOUNGER node,
/// timed to the survivor's member REMOVAL (detection + stable-after + gossip)
/// with singleton continuity asserted on the oldest. The oldest/active-node
/// crash is NOT a recovery to time — two-node keep-oldest makes the younger
/// survivor down itself (total outage; registered deferred user decision,
/// see SbrFailoverTests XML doc + master tracker 2026-07-08). Covers overall
/// review P2-10 / report-08 NF2 and report-01 round-2 N1's measurement ask.
/// Runs under the production default downing strategy (auto-down, decision
/// 2026-07-21 — either-direction crash fails over; see SbrFailoverTests XML
/// doc). Measures a hard-crash of the YOUNGER node, timed to the survivor's
/// member REMOVAL (detection + stability window + gossip) with singleton
/// continuity asserted on the oldest; the oldest-crash direction is covered
/// behaviorally by SbrFailoverTests.AutoDown_HardCrashOfOldestNode_* and by
/// the docker failover drill. Covers overall review P2-10 / report-08 NF2
/// and report-01 round-2 N1's measurement ask.
/// </summary>
public class FailoverTimingTests(ITestOutputHelper output)
{