feat(health)!: active node is the oldest Up member, not the leader (0.3.0)

Promotes the two apps' private replacements for ActiveNodeHealthCheck into one
shared primitive, and retires the leader/RoleLeader selection they were written
to avoid.

Both Akka consumers had already hand-rolled a replacement rather than use this
package (ScadaBridge OldestNodeActiveHealthCheck, OtOpcUa
ClusterPrimaryHealthCheck), so the entire active-node surface here —
ActiveNodeHealthCheck, AkkaActiveNodeGate — had ZERO consumers. It was not
merely unused: it was avoided, twice, for the same reason.

Leadership is address-ordered (host, then port) and has no relationship to time;
singleton placement is age-ordered. The two agree on a freshly-formed cluster,
which is why single-node and happy-path tests never caught it. They diverge
permanently after any restart: the restarted node rejoins as the youngest but
keeps its address, so if it holds the lower address it becomes leader while the
singletons — and all the work they own — stay on the other node.

New ClusterActiveNode is the single implementation: oldest Up member, optional
role scope, plus role-preference resolution for a fused node that must answer
for the role its singletons are pinned to. ActiveNodeHealthCheck and
AkkaActiveNodeGate both delegate to it, so an endpoint gate and the
/health/active probe an orchestrator routes by cannot disagree.

BREAKING (behaviour, not signature):
- Selection is by age, not leadership.
- The role-filtered mode no longer reports Healthy for a node LACKING the role.
  That "not applicable => Healthy" made the tier answer 200 on every node and
  silently broke leader-pinning (lmxopcua#494); it is now Unhealthy by default,
  overridable via NoActiveRoleStatus. This case is reachable, not defensive —
  OtOpcUa's RoleParser admits dev-only and cluster-role-only nodes.
- Identity compares UniqueAddress, so a node restarted on the same host:port is
  correctly a different member during the overlap.

Results now carry activeRole/selfAddress/activeNode in the 0.2.0 per-entry data
object, so a standby reports WHO is active and a dashboard can render a pair
from either node's payload alone.

Tests: 45 Akka (was 39), 76 total. The age-vs-address divergence is pinned
against a real two-node cluster built so the oldest member is not the
lowest-addressed one, with a fixture sanity check so it cannot pass for the
wrong reason.
This commit is contained in:
Joseph Doherty
2026-07-24 13:07:56 -04:00
parent c81a6615ba
commit a61e041e58
11 changed files with 1332 additions and 263 deletions
+531
View File
@@ -0,0 +1,531 @@
# Akka.NET HA, Failover & Singletons — ScadaBridge + OtOpcUa Deep Dive
*Written 2026-07-22 against ScadaBridge `main @ 69b3ccfc` and OtOpcUa (`lmxopcua`) `master @ 0de1352e`.
Both trees contain the 2026-07-21 downing/election corrections (ScadaBridge `cf3bd52f`; OtOpcUa `2964361a` auto-down + `50b55ee4` oldest-driver election — note `0de1352e` itself is the adjacent NU1903 dependency pin, not the redundancy fix).*
**Design priority (stated 2026-07-22):** for the 2-node site pairs, a network-partition
split-brain is *less* risky than failing to bring the second node up quickly. HA goal:
either node can be launched first, automatic failover, no manual intervention. The
architecture below is evaluated against that priority throughout.
---
## 1. How Akka.NET handles failover — the fundamentals both apps build on
The deployment unit at every site is a **2-node Akka.NET cluster** (one node per Windows
VM). Central is also a 2-node cluster, with SQL Server; sites have no SQL Server.
Everything HA-related decomposes into five Akka mechanisms:
### 1.1 Membership and seed nodes
A cluster forms by gossip. Each node lists `seed-nodes`; a booting node sends `InitJoin`
to the seeds. **Only the FIRST seed in the list may join itself to form a *new* cluster**
any other node waits until some existing member answers. Both apps set
`min-nr-of-members = 1`, so a single node (if it's allowed to form) goes `Up` alone and
hosts all singletons — the pair does not need both members to become operational.
### 1.2 Failure detection (the heartbeat)
Every node heartbeats its peers and feeds the intervals into a **Phi Accrual failure
detector**. When the suspicion level crosses the threshold, the peer is marked
**Unreachable** — a suspicion, not a removal. Both apps run the same values:
| Setting | Value (both apps) | Meaning |
|---|---|---|
| `failure-detector.heartbeat-interval` | **2 s** | Cluster-membership heartbeat rate |
| `failure-detector.threshold` | 10.0 | Phi threshold (suspicion sensitivity) |
| `failure-detector.acceptable-heartbeat-pause` | **10 s** | Tolerated silence before Unreachable |
| `transport-failure-detector` (remoting layer) | 2 s / 10 s pause | Detects dead TCP associations |
So a hard-killed peer is declared Unreachable roughly **1012 s** after death. These are
exactly the "heartbeat with a small adjustable timeout" — in ScadaBridge they are bound
from config (`Cluster:HeartbeatInterval`, `Cluster:FailureDetectionThreshold`
`ClusterOptions`, rendered into HOCON by `AkkaHostedService.BuildHocon`); in OtOpcUa they
live in `akka.conf` (lines 6872) in the Cluster library.
### 1.3 Downing (the decision to act)
Unreachable alone changes nothing — the leader cannot make membership decisions while any
member is unreachable, singletons do not move, the cluster is wedged. Something must
**Down** the unreachable member. This is the strategy choice, and it is the whole
availability-vs-partition-safety trade:
- **Auto-downing** (`Akka.Cluster.AutoDowning` + `auto-down-unreachable-after`): after the
window, the leader *among the reachable members* downs the unreachable one. In a 1-vs-1
split, **each side downs the other** — the crash case fails over perfectly; a genuine
partition produces dual-active.
- **Split Brain Resolver `keep-oldest` (+ `down-if-alone`)**: keeps the side containing
the oldest member. Partition-safe, but **fatally broken for a 2-node pair**: in
Akka.NET 1.5.62, `KeepOldest.OldestDecision` only lets `down-if-alone` rescue a side
holding ≥ 2 members (`otherSide == 1 && thisSide >= 2`). A 1-vs-1 survivor whose peer
(the oldest) crashed takes `DownReachable` and **downs itself** — proven live:
`SBR took decision Akka.Cluster.SBR.DownReachable … including myself, [1] unreachable of [2] members`.
Net effect: the pair could not survive a crash of the oldest node — total outage.
**Both apps switched their default from `keep-oldest` to `auto-down` on 2026-07-21**
(ScadaBridge `cf3bd52f`, decision record
`ScadaBridge/docs/plans/2026-07-21-auto-down-availability-decision.md`; OtOpcUa
`2964361a`, same fix ported). The owner rationale recorded in the decision doc is
verbatim the priority stated above: *"network partitions are less of a risk than if this
stops working."*
Rejected alternatives (from the decision record): `static-quorum(1)``DownAll` (worse);
`static-quorum(2)` → survivor self-downs; `keep-majority` → keeps the lowest-address side,
same fatal crash re-keyed; `lease-majority` → needs a shared lease store (no SQL/K8s at
sites); a third arbiter node → no third VM at sites; a custom downing provider → would
reimplement AutoDowning.
### 1.4 Cluster singletons (what actually "fails over")
`ClusterSingletonManager` runs on every node of the (optionally role-filtered) cluster
and guarantees exactly one instance of the actor, hosted on the **oldest member** of that
role. `ClusterSingletonProxy` routes to wherever it currently lives. Failover semantics:
- **Graceful shutdown** (service stop, redeploy): hand-over protocol — the old host
transfers to the new oldest with no dead window beyond hand-over latency. Both apps
hook `CoordinatedShutdown` (`run-by-clr-shutdown-hook = on`; ScadaBridge additionally
drains each singleton via a `PhaseClusterLeave` `GracefulStop` task, 10 s per singleton,
`phases.cluster-leave.timeout = 15s`).
- **Hard crash**: the singleton is simply gone until the dead node is **Downed**; then the
new oldest member's manager instantiates a fresh instance. This is why downing is the
linchpin — no downing decision, no singleton recovery, ever.
### 1.5 Oldest-member semantics — the family-wide rule
Both codebases converged on the same insight: **Akka's "leader" (lowest address) and
"oldest member" (singleton placement) diverge permanently after any node restart** — the
restarted node rejoins as a *new incarnation*, becoming the youngest, but may still have
the lowest address and thus be leader. Any "active node" decision keyed on
leader/lowest-address can therefore point at a node that is *not* hosting the singletons.
Both apps now derive "active/primary" strictly from **oldest Up member** (ScadaBridge
`ActiveNodeEvaluator.SelfIsOldestUp`; OtOpcUa `RedundancyStateActor.SelectDriverPrimary`
with `Member.AgeOrdering`), which by construction matches singleton placement.
---
## 2. The failover timeline (and the knobs that tune it)
End-to-end failover after a hard node death, with current production values:
```
t=0 node dies (process crash / VM loss)
t≈10-12s failure detector: peer marked Unreachable ← acceptable-heartbeat-pause (10s)
t≈25-27s auto-down window expires; survivor downs peer ← auto-down-unreachable-after (15s)
t≈25-30s survivor is now oldest; singletons instantiate;
primary/active gates flip; ServiceLevel 250 moves
```
Measured, not theoretical: ScadaBridge's 2026-07-21 docker drill measured **28 s**
(active-node crash → standby takeover) and **27 s** (standby crash → removal, zero routing
blips); the in-process `FailoverTimingTests` measured 33.7 s for a full cycle.
**Tuning for faster failover** (per the availability-first priority):
| Knob | Where | Current | Effect of lowering |
|---|---|---|---|
| `acceptable-heartbeat-pause` | SB: `Cluster:FailureDetectionThreshold`; OtOpcUa: `akka.conf` | 10 s | Faster Unreachable, but < ~5 s risks false positives from GC pauses / Windows VM scheduling stalls / snapshot freezes |
| `heartbeat-interval` | SB: `Cluster:HeartbeatInterval`; OtOpcUa: `akka.conf` | 2 s | More samples → detector converges faster; cheap to lower to 1 s |
| Auto-down window | SB: `Cluster:StableAfter`; OtOpcUa: `DowningStableAfter` const (15 s) + emitted HOCON | 15 s | Direct 1:1 reduction of failover time. Every second cut is a second of dual-active risk during a *transient* blip (link flap, VM pause) — the window exists to let transients heal before the irreversible Down |
| `down-removal-margin` (OtOpcUa akka.conf) | 15 s | Delay before rebalancing after removal | Keep ≈ downing window |
Constraint pinned by an OtOpcUa test (`Downing_window_exceeds_the_acceptable_heartbeat_pause`):
the downing window **must exceed** `acceptable-heartbeat-pause`, otherwise a node can be
downed on a pause the detector was configured to tolerate. A realistic aggressive profile
for the site pairs would be ~5 s pause + ~78 s window ≈ **1213 s total failover**,
at the cost of downing a peer that stalls (not dies) for > ~12 s — acceptable per the
stated priority, but worth validating against real Windows VM GC/backup/snapshot pauses
before rollout. In ScadaBridge these are plain per-deployment config values; in OtOpcUa
the 15 s window is currently a constant in `ServiceCollectionExtensions.cs` (only the
*strategy* is config-selectable), so making the window configurable would be a small
change there.
**What happens to the dead node's side:** both apps set
`run-coordinated-shutdown-when-down = on` — a node that gets Downed (e.g. it was merely
partitioned and the other side downed it) terminates its own ActorSystem. Both apps run a
termination watchdog (ScadaBridge in `AkkaHostedService`; OtOpcUa
`ActorSystemTerminationWatchdog`) that distinguishes intentional stop from self-down and
calls `StopApplication()`, so the **process exits and the service supervisor restarts it**
(Windows: `sc.exe failure … restart/5000/restart/30000/restart/60000`; Docker:
`restart: unless-stopped`). The restarted process rejoins as a fresh incarnation —
youngest member, i.e. it comes back as the **standby**, never usurping the survivor.
---
## 3. ScadaBridge
### 3.1 Cluster formation
All HOCON is generated in code — `AkkaHostedService.BuildHocon(...)`
(`src/ZB.MOM.WW.ScadaBridge.Host/Actors/AkkaHostedService.cs:258329`) from
`ClusterOptions`/`NodeOptions` (`src/ZB.MOM.WW.ScadaBridge.ClusterInfrastructure/ClusterOptions.cs`).
Key facts:
- **Seed nodes**: central docker pair → `scadabridge-central-a:8081` then
`…-central-b:8081` (both nodes list `central-a` first). Site pairs analogous on 8082.
- **`MinNrOfMembers = 1`** (validator *requires* exactly 1) — one node forms the cluster
and brings up all singletons.
- **Roles**: central node = `["Central"]`; site node = `["Site", "site-{SiteId}"]`
(`BuildRoles`, line 406). The role drives the entire composition branch in `Program.cs`.
- **Ports**: central remoting 8081; site remoting 8082, gRPC 8083, metrics 8084.
### 3.2 Downing configuration
Configurable via **`Cluster:SplitBrainResolverStrategy`**, allowed values `auto-down` |
`keep-oldest` (case-insensitive; validator `ClusterOptionsValidator`). **Default and all
16 appsettings files: `auto-down`.** The HOCON branch (`AkkaHostedService.cs:275286`):
```csharp
var downingBlock = string.Equals(strategy, "auto-down", OrdinalIgnoreCase)
? $@"downing-provider-class = ""Akka.Cluster.AutoDowning, Akka.Cluster""
auto-down-unreachable-after = {DurationHocon(clusterOptions.StableAfter)}"
: /* SBR provider + split-brain-resolver { active-strategy, stable-after,
keep-oldest { down-if-alone } } */;
```
`StableAfter = 00:00:15` doubles as the SBR `stable-after` and the auto-down window.
`DownIfAlone` is now validated only under `keep-oldest` (inert under auto-down).
`keep-oldest` remains supported and SBR-path tests still pin it
(`SbrFailoverTests.HardCrashOfYoungerNode_SbrDownsIt_AndOldestKeepsSingleton`).
### 3.3 Cluster singletons
All registered through `SingletonRegistrar.Start(...)` (names `{name}-singleton` /
`{name}-proxy`, PoisonPill termination, optional role, per-singleton
`PhaseClusterLeave` drain task of 10 s).
**Central — 7 singletons** (role not further scoped; central nodes are all `Central`):
| Singleton | Actor | Purpose |
|---|---|---|
| `notification-outbox` | `NotificationOutboxActor` | Notification ingest/dispatch/purge |
| `audit-log-ingest` | `AuditLogIngestActor` | Central audit ingest from sites |
| `site-call-audit` | `SiteCallAuditActor` | SiteCalls upserts + central→site retry/discard relay |
| `audit-log-purge` | `AuditLogPurgeActor` | Daily partition-switch purge |
| `site-audit-reconciliation` | `SiteAuditReconciliationActor` | Per-site audit reconciliation pull |
| `kpi-history-recorder` | `KpiHistoryRecorderActor` | KPI sampling (not readiness-gated) |
| `pending-deployment-purge` | `PendingDeploymentPurgeActor` | Expired staging-row reclaim (not readiness-gated) |
`RequiredSingletonsHealthCheck` probes the first five via bounded `Identify` for
`/health/ready`; the last two are deliberately best-effort.
**Site — 2 singletons**, role-scoped to `site-{SiteId}` so each site pair elects its own:
| Singleton | Actor | Purpose |
|---|---|---|
| `deployment-manager` | `DeploymentManagerActor` | Deployment lifecycle; drains in-flight SQLite writes on hand-over |
| `event-log-handler` | `EventLogHandlerActor` | Event-log queries always served by the active node (event log is node-local, not replicated) |
Deliberately **not** singletons: `dcl-manager` (DataConnectionManagerActor) and
`CertStoreActor` run on **every** site node — the deployment singleton fans PKI trust
changes to both nodes' cert stores so a failover never finds a stale trust store.
### 3.4 Active-node gating
One definition of "active": `ActiveNodeEvaluator.SelfIsOldestUp`
(`src/ZB.MOM.WW.ScadaBridge.Communication/ClusterState/ActiveNodeEvaluator.cs`) — oldest
Up member, explicitly **not** `cluster.State.Leader` (the XML docs spell out the
leader-vs-oldest divergence after restarts). Consumers:
- **Central inbound API**: `IActiveNodeGate``InboundApiEndpointFilter` returns
**HTTP 503 `{code:"STANDBY_NODE"}`** on the standby; Traefik routes on
`/health/active` (`OldestNodeActiveHealthCheck` + `DatabaseHealthCheck`), so the proxy
and the API agree on which node is active.
- **Site store-and-forward**: `storeAndForwardService.SetDeliveryGate(() =>
clusterNodeProvider.SelfIsPrimary)` — only the active node runs the delivery sweep;
without the gate both nodes would deliver the same replicated `sf_messages` rows
(systematic duplicate external calls). Re-evaluated per sweep tick, so delivery resumes
within one `RetryTimerInterval` after failover.
- **Heartbeat to central**: `SiteCommunicationActor` stamps `HeartbeatMessage.IsActive`
from the same evaluator (exception-safe, defaults to `false`).
Takeover is purely membership-driven: peer Downed → survivor becomes oldest →
`SelfIsOldestUp` flips → singletons, S&F gate, `/health/active`, heartbeat all follow.
### 3.5 Site state without SQL Server — LocalDb
Sites persist everything in **one consolidated SQLite file** (`LocalDb:Path`, required,
on the data volume) via `ZB.MOM.WW.LocalDb` — 10 replicated tables: `OperationTracking`,
`site_events`, `sf_messages` (the store-and-forward buffer), and 7 config tables
(deployed configurations, static overrides, shared scripts, external systems, DB
connections, DCL definitions, native alarm state).
**Pair replication** (default OFF; opt-in `LocalDb:Replication:PeerAddress` + matching
`ApiKey` both sides; fail-closed auth interceptor; rides the existing gRPC 8083 listener):
CDC triggers capture row mutations, shipped bidirectionally under HLC last-writer-wins;
snapshot resync merges per row and never deletes. **This is what makes site failover
real**: the standby already holds the S&F buffer and the full config, so on takeover it
just opens its delivery gate over data it already has — zero cross-node fetch (the old
notify-and-fetch deploy path is deleted).
Not replicated (by design): the event log (hence the `event-log-handler` singleton),
in-memory alarm state (re-evaluated from incoming values on the survivor), and
`notification_lists`/`smtp_configurations` (would ship plaintext SMTP passwords).
**Operational constraints:**
- **Stop/start a site pair TOGETHER.** Rolling upgrades are unsupported: Phase 2 deleted
the legacy replicator and its `SfBufferSnapshot` compat handler, so a mixed-version
pair runs but silently does not converge.
- A node offline longer than `TombstoneRetention` (7 d) can **resurrect deleted rows** on
rejoin; within the window rejoin is safe and self-correcting.
- `MaxBatchSize` counts **rows, not bytes** — rig pins 16 (the 500 default could build a
~35 MB batch against the 4 MB gRPC cap).
### 3.6 Verified behavior + remaining gaps
- Auto-down failover drill (docker `failover-drill.sh`, SIGKILL): active-crash takeover
28 s, standby-crash 27 s, 0 routing blips; both drill modes now *assert* recovery.
- After a real partition the two mutually-downed halves **do not merge** (quarantined
association) — an operator restarts one side, which rejoins as standby. This is the
accepted dual-active trade.
- **The old "restart central-a and central-b together (SBR self-down)" gotcha is closed**
by auto-down — but a related constraint survives with a different mechanism: the
**seed-node bootstrap constraint** (§6.1).
- ⚠ Stale doc: `docs/deployment/topology-guide.md:101` still says keep-oldest +
down-if-alone. Authoritative: `Component-ClusterInfrastructure.md`, the auto-down
decision record, `docker/README.md`.
---
## 4. OtOpcUa
### 4.1 Cluster formation
Base HOCON in the library (`src/Core/ZB.MOM.WW.OtOpcUa.Cluster/Resources/akka.conf`):
system `otopcua`, remoting port 4053, `min-nr-of-members = 1`, empty seeds/roles overlaid
at runtime from `AkkaClusterOptions` (`Cluster:*` config; roles restricted to
`admin` | `driver` | `dev` by `RoleParser`). Bootstrap via Akka.Hosting:
`WithOtOpcUaClusterBootstrap` (remoting + clustering + the downing HOCON, Prepended).
Note the dual role plumbing: `OTOPCUA_ROLES` drives DI composition (`hasAdmin`/`hasDriver`
branches in `Program.cs`), while `Cluster__Roles__*` drives the actual Akka member roles —
docker-dev sets both (a past production incident came from setting only `OTOPCUA_ROLES`,
leaving the Akka member role-less).
**Current mesh shape (important):** today OtOpcUa runs **one Akka mesh for the whole
fleet** — docker-dev is six nodes (central-1/2 = `admin,driver`; site-a-1/2, site-b-1/2 =
`driver`) all seeded by `central-1`, with logical MAIN/SITE-A/SITE-B separation done by
`ServerCluster.ClusterId` rows in the shared ConfigDb, not by separate meshes. The
per-cluster mesh design (§4.4) will change this to one mesh per 2-node pair — the shape
that matches your site deployment.
### 4.2 Downing configuration
Same key as ScadaBridge: **`Cluster:SplitBrainResolverStrategy`**, default **`auto-down`**,
allowed `auto-down` | `keep-oldest`, anything else **fails the host at startup**
(`ArgumentOutOfRangeException` — no silent fallback; pinned by
`Unknown_strategy_throws_rather_than_silently_falling_back`). `BuildDowningHocon` emits:
```
akka.cluster {
downing-provider-class = "Akka.Cluster.AutoDowning, Akka.Cluster"
auto-down-unreachable-after = 15000ms
}
```
(the comment notes `auto-down-unreachable-after` defaults to `off` upstream and is
load-bearing — AutoDowning without it never downs anything). Under `keep-oldest` the
typed `KeepOldestOption { DownIfAlone = true }` installs the SBR instead; the
`split-brain-resolver` block in akka.conf is inert under the default. Tests read the
**effective config off a running ActorSystem** (not the input HOCON) — provider class,
15 s window, case-insensitivity, and the window-exceeds-heartbeat-pause invariant.
**Live gate still open:** the 1-vs-1 oldest-crash pathology cannot be exercised on the
current six-node docker-dev mesh; the crash-the-oldest drill is deferred to the
per-cluster mesh work (Phase 6/7). Supporting live evidence exists from a 2-container
kill test (`archreview/plans/artifacts/459-oldest-crash-live-finding-2026-07-15.md`) and
from ScadaBridge's identical fix. Cautionary note recorded in `docs/Redundancy.md`:
the pre-existing `HardKillFailoverTests` stayed green under broken keep-oldest because it
shut down the *transport* (ActorSystem alive) rather than killing the process — "a
standing example of a green test over a fatal defect."
### 4.3 Redundancy primary election + ServiceLevel
- **`RedundancyStateActor`** (an `admin`-role singleton) computes the fleet snapshot:
`SelectDriverPrimary` = oldest **Up** member carrying the `driver` role
(`Member.AgeOrdering`; `Leaving` excluded — a leaving node is handing its singletons
over and must not be Primary). Debounced 250 ms, republished on a 10 s heartbeat over
distributed pub/sub topic `redundancy-state` (DPS doesn't replay to late subscribers).
- **`NodeRedundancyState`** carries `RedundancyRole { Primary, Secondary, Detached }` and
`IsDriverPrimary` (renamed from `IsRoleLeaderForDriver` — "the name now describes what
the value means rather than how it used to be computed").
- **ServiceLevel** (`ServiceLevelCalculator`): health basis 240 (healthy) / 200 (stale) /
100 (critically degraded) / 0 (down), **+10 if driver Primary**, so a healthy pair
reads **250 / 240**. Published into the OPC UA `Server.ServiceLevel` variable via
`SdkServiceLevelPublisher` (deferred publisher until the SDK server exists; first
computed value always published so no node lingers at the SDK default 255).
- **Client-visible failover is OPC UA non-transparent redundancy**: each node has its own
`ApplicationUri`; clients read `Server.ServerArray` (self + peers, from
`PeerApplicationUris`) and pick the endpoint with the highest ServiceLevel. The shipped
client (`OpcUaClientService`) does this: failover URL list, KeepAlive-failure trigger,
single-flight failover guard, data + alarm subscription replay after reconnect.
**Consumer note:** because Primary = oldest (not lowest address), after any node
restart a ServiceLevel-selecting client may legitimately prefer a different node than
before — that's correct behavior, not a bug.
**Primary gating of the data plane (`PrimaryGatePolicy`)** — the S4 default-deny fix:
```csharp
localRole switch {
Primary =&gt; true,
Secondary or Detached =&gt; false,
_ /* unknown */ =&gt; driverMemberCount &lt;= 1, // multi-driver: DENY until proven
};
```
During the boot window (role not yet known) a multi-driver node **rejects device writes**
with `not primary (role unknown)` (denial meter `reason=role-unknown`); a single-driver
node stays default-allow. Gated surfaces: device writes, native-alarm acks, alarm-alert
fan-out. A deliberately *different* policy governs alarm-history drain: unknown → drain
anyway ("a duplicate row beats a silent gap"), keyed on whether the Primary shares this
node's queue (pair-local). Live-proven on a 2-node docker rig (R2-04 T15): ServiceLevel
250/240, boot-window write rejected, steady-state secondary rejected + reverted, primary
write reaches the device.
Failover sequence when the primary dies: peer downed (auto-down) → survivor is oldest
driver → `SelectDriverPrimary` flips → snapshot republished → gates open on the survivor →
ServiceLevel 250 moves → OPC UA clients see the old endpoint die (KeepAlive) and re-select
by ServiceLevel. Singletons migrate to the same node the election names, because both
derive from member age.
### 4.4 Singletons and the per-cluster mesh gap
Five control-plane singletons, all role-pinned to `admin`:
| Singleton | Actor | Purpose |
|---|---|---|
| `config-publish` | `ConfigPublishCoordinator` | Persists per-node deploy ACKs |
| `admin-operations` | `AdminOperationsActor` | Operator ack/shelve; TagConfig deploy gate |
| `audit-writer` | `AuditWriterActor` | Audit persistence |
| `fleet-status` | `FleetStatusBroadcaster` | AdminUI fleet panel |
| `redundancy-state` | `RedundancyStateActor` | Primary election + DPS publisher |
**KNOWN LIMITATION:** the Primary election is scoped **per Akka cluster, not per
application Cluster**. On the one-mesh fleet, `SelectDriverPrimary` names exactly ONE
Primary among *all* driver nodes (MAIN + SITE-A + SITE-B), while every gated resource is
pair-local. The per-cluster mesh design
(`OtOpcUa/docs/plans/2026-07-21-per-cluster-mesh-design.md`) fixes this by construction:
one Akka mesh per application Cluster (two nodes max), central↔cluster joined by explicit
transports (modeled on ScadaBridge's ClusterClient/gRPC/HTTP trio), so "am I the driver
Primary?" and "the resource I'm gating" have the same scope. Status: Phases 0a (auto-down)
and 0b (oldest-Up election) DONE and merged; Phases 17 (ClusterNode columns, comm actors,
config fetch-and-cache, cut driver-side ConfigDb, gRPC stream contract, mesh partition,
failover drill/live gate) not started — **but now sequenced by a program plan
(`OtOpcUa/docs/plans/2026-07-22-per-cluster-mesh-program.md`, 2026-07-22), driven by the
decision that OtOpcUa pairs will run on the SAME Windows VMs as the ScadaBridge nodes** —
two independent, identically-postured 2-node clusters per site, and driver nodes off the
ConfigDb entirely (sites have no SQL Server).
### 4.5 Surviving central-SQL loss — availability guarantees
Two mechanisms make a driver pair independent of central SQL Server availability:
- **Address-space availability guarantee (#485/#486):** an artifact the node cannot read
is *no answer*, never *an empty configuration*. Previously a transient ConfigDb error
mid-deploy parsed zero bytes into an empty composition → the planner diffed it as a
`PureRemove` of every node → served address space emptied, drivers stopped,
subscriptions cleared, deploy still reported Applied. Now `OpcUaPublishActor` /
`DriverHostActor` hold last-known-good address space, drivers, and subscriptions, and
report the deploy **Failed** without advancing the revision so the retry lands.
- **LocalDb Phase 1 boot-from-cache:** every driver node caches the chunked, SHA-256
verified deployed-configuration artifact in a pair-local SQLite DB; a node restarting
into a central-SQL outage boots from the last artifact it applied. With opt-in pair
replication (`deployment_artifacts` + `deployment_pointer`, HLC last-writer-wins,
fail-closed auth) either node holds a byte-identical copy even if it never applied that
deploy itself. Scope: config artifact only — no live tags/alarms/historian data.
---
## 5. Dual-active: what it concretely means here
Under auto-down, a *genuine* partition (both nodes alive, link cut) ends with each side
downing the other and running active until an operator restarts one side (the mutual
downing quarantines the association — the sides will not self-merge). Accepted per the
availability-first decision; concrete blast radius:
**ScadaBridge site pair:** both nodes' `SelfIsOldestUp` = true → both open the S&F
delivery gate → **duplicate external deliveries** of the same replicated `sf_messages`
rows for the partition's duration; both stamp `IsActive` heartbeats. LocalDb replication
itself is partition-tolerant (HLC last-writer-wins reconverges on heal, after the
operator restart). Central pair: both serve `/health/active` = 200 and run duplicate
singleton sets against the same SQL Server — duplicated sweeps/purges (mostly idempotent
upserts), and Traefik may alternate between two "active" backends.
**OtOpcUa pair:** both nodes elect themselves driver Primary → both publish
**ServiceLevel 250** → clients on either side keep their session; both sides accept
device writes → **two masters against the same field devices** for the duration. This is
the classic OT dual-primary risk — mitigated by rarity (2 VMs, one LAN segment, real
partitions ≈ switch failure) and by the alternative (keep-oldest) having been proven to
turn a plain crash into a total outage.
Detection/ops: the survivor logs the downing decision; OtOpcUa emits
`otopcua.redundancy.service_level_change` and `primary_gate_denied` meters; two
simultaneous 250s / two 200-responses on `/health/active` are the alarm signal. Recovery
is always "restart one side; it rejoins as youngest/standby."
---
## 6. Gaps vs the stated HA goal ("either node first, automatic failover")
### 6.1 The seed-node bootstrap constraint — CLOSED 2026-07-22 (differently per repo)
**Automatic failover: achieved.** Either node crashing is survivable, unattended, in
~2730 s. The remaining gap — a non-first seed cold-starting alone loops on `InitJoin`
forever (the "registered outage gap") — was **closed in both repos on 2026-07-22**, but
execution overturned the original design and the two repos legitimately diverged:
**⚠️ The original design's safety claim was FALSE — proven independently in both repos.**
The planned `SelfFormAfter` watchdog assumed "window expires mid-join-handshake → benign,
Akka ignores `Join` once joined." Wrong: the watchdog sits *outside* Akka's join
handshake and cannot distinguish *no seed answered* from *a seed answered and the join is
in flight*. On a routine restart into a live peer, the join stalls briefly behind removal
of the node's own stale incarnation; a `Join(self)` fired in that window **abandons the
in-flight join and forms a second cluster at the same address — a permanent island**.
ScadaBridge proved it with a written test (split still unhealed after 90 s); OtOpcUa hit
it live when its manual-failover drill bounced a node (InitJoinAck received, no Welcome
inside the window, fallback fired → second cluster).
**ScadaBridge's fix — self-first seed ordering (no runtime code).** Every node lists
ITSELF as `seed-nodes[0]`, partner second. Akka's own `FirstSeedNodeProcess` then
implements exactly the intended semantics *inside* the handshake (InitJoin the other
seeds; self-join only if nobody answers) — race-free by construction. A `StartupValidator`
rule hard-gates the ordering at boot. Live-proven behaviors: lone cold-start self-joins
after ~5 s (`seed-node-timeout`); restart into a live peer rejoins (never islands);
simultaneous reachable cold-start converges to ONE cluster (the old "two clusters that
never merge" objection was disproved by test); a genuine boot partition dual-forms — the
accepted auto-down trade. There is **no `SelfFormAfter` option in ScadaBridge** — the
whole watchdog was rejected in code review.
**OtOpcUa's fix — converged on self-first ordering the same day.** OtOpcUa first shipped
the watchdog and hit the race live (its manual-failover drill bounced a node; the fallback
islanded it), patched it with a TCP reachability guard — and then **retired the whole
watchdog hours later** (`lmxopcua master @ 3f24d4d6`, a breaking change):
`ClusterBootstrapFallback` and `Cluster:SelfFormAfter` are **deleted**, docker-dev
`central-2` now lists itself as `SeedNodes__0`, and a new `AkkaClusterOptionsValidator`
enforces the ordering at boot. Two subtleties its rule handles that ScadaBridge's didn't
need: the invariant is **conditional** (self must be `seed[0]` only IF self appears in
the list at all — site nodes seed off central only and are deliberately not seeds), and
identity compares on **`PublicHostname` + port** (the address Akka puts in `SelfAddress`),
because matching the `0.0.0.0` bind address would leave the rule silently inert on the
whole rig. The retirement was proven by a falsifiability control: the old peer-first
ordering test failed at 11 s against the watchdog — demonstrating the watchdog could
never usefully fire for the nodes that needed it.
**Result: both repos now share the identical mechanism** — self-first seed ordering,
validator-enforced, no bootstrap runtime code. The mesh program's Phase 6 convergence
item is already done; what remains there is only the per-pair seed topology itself.
**Ops action (outstanding):** the gitignored `ScadaBridge/deploy/wonder-app-vd03/`
overlay must have its `SeedNodes` reordered self-first before its next deploy — the new
validator rule is a deliberate hard boot gate.
### 6.2 Other open items
- **OtOpcUa auto-down live gate** still open (needs a real 1-vs-1 rig; deferred to
per-cluster mesh Phase 6/7). ScadaBridge's identical mechanism *is* drill-proven.
- **OtOpcUa election scope** is cluster-wide until per-cluster mesh Phases 17 land
(§4.4) — only relevant while multiple pairs share one mesh.
- **OtOpcUa's 15 s downing window is a code constant**, not config — worth making
tunable alongside any heartbeat-timeout reduction (ScadaBridge's already is:
`Cluster:StableAfter`).
- **ScadaBridge site pairs: stop/start together** for upgrades (no mixed-version LocalDb
replication); ≥ 7-day-offline nodes risk tombstone resurrection.
- **Stale doc**: `ScadaBridge/docs/deployment/topology-guide.md:101` still describes
keep-oldest.
---
## 7. Quick reference — file map
| Concern | ScadaBridge | OtOpcUa |
|---|---|---|
| HOCON / downing emit | `Host/Actors/AkkaHostedService.cs:258329` | `Cluster/ServiceCollectionExtensions.cs:119146` + `Cluster/Resources/akka.conf` |
| Options + validation | `ClusterInfrastructure/ClusterOptions.cs`, `ClusterOptionsValidator.cs` | `Cluster/AkkaClusterOptions.cs`, `RoleParser.cs` |
| Strategy key | `Cluster:SplitBrainResolverStrategy` (default `auto-down`) | same key, same default; unknown value = startup failure |
| Active/primary definition | `Communication/ClusterState/ActiveNodeEvaluator.cs` (oldest Up) | `ControlPlane/Redundancy/RedundancyStateActor.cs` (oldest Up `driver`) |
| Data-plane gating | `InboundApiEndpointFilter` (503 STANDBY_NODE), S&F `SetDeliveryGate` | `Runtime/Drivers/PrimaryGatePolicy.cs` + `DriverHostActor` gates |
| Singleton registration | `Host/Actors/SingletonRegistrar.cs` (7 central + 2 site) | `ControlPlane/ServiceCollectionExtensions.cs:4494` (5, role `admin`) |
| Client-visible failover | Traefik on `/health/active`; gRPC A/B flip | OPC UA ServiceLevel 250/240 + `Server.ServerArray`; client KeepAlive failover |
| Self-down recovery | watchdog in `AkkaHostedService` → process exit → supervisor | `Host/ActorSystemTerminationWatchdog.cs` → same |
| Decision records | `docs/plans/2026-07-21-auto-down-availability-decision.md`, `docs/requirements/Component-ClusterInfrastructure.md` | `docs/Redundancy.md` §"Split-brain / downing", `docs/plans/2026-07-21-per-cluster-mesh-design.md` |
| Failover drills / tests | `docker/failover-drill.sh` (28 s/27 s), `SbrFailoverTests`, `FailoverTimingTests` | `SplitBrainResolverActivationTests`, `RedundancyPrimaryElectionTests`, R2-04 T13/T15 live gate |