diff --git a/docs/plans/2026-07-21-per-cluster-mesh-design.md b/docs/plans/2026-07-21-per-cluster-mesh-design.md new file mode 100644 index 00000000..f19b7ebf --- /dev/null +++ b/docs/plans/2026-07-21-per-cluster-mesh-design.md @@ -0,0 +1,222 @@ +# Per-Cluster Akka Mesh — design + +> **Status:** design for review. No implementation planned yet. +> **Decision requested:** the three open questions in §6, before any per-phase plan is written. + +## 1. What this changes + +Today OtOpcUa runs **one Akka mesh containing every node of every application Cluster**. That was a +deliberate choice — `docs/plans/2026-06-07-per-cluster-scoping.md`, "Per-ClusterId Scoping +(hub-and-spoke **single mesh**)" — so the deploy channel could stay in-mesh: AdminUI → +`admin-operations` singleton → `deployments` topic → every node filters to its own `ClusterId`. + +The proposal is to move to the sister project's shape: **one Akka mesh per application Cluster, +two nodes maximum**, with central and clusters joined by explicit transports instead of cluster +gossip. + +The motivation is that redundancy becomes correct *by construction*. Every Primary-gated decision +(inbound device writes, native alarm acks, fleet alerts, ServiceLevel, and — since Phase 2 — the +alarm-history drain) currently asks "am I the driver Primary?" of a **cluster-wide** election, while +every resource being gated is **pair-local**. Splitting the mesh collapses the two scopes into one +and the question stops being ambiguous. + +## 2. What ScadaBridge actually does + +Researched directly from `/Users/dohertj2/Desktop/ScadaBridge`. Five things differ from the +summary in its own CLAUDE.md, and they matter: + +**Three transports cross the boundary, not two.** + +| Transport | Direction | Carries | +|---|---|---| +| Akka **ClusterClient** | central → site (+ replies) | command/control: deploy notify, lifecycle, queries, subscribe handshakes | +| **gRPC** server-streaming | **central dials into the site** | real-time attribute + alarm events | +| Plain **HTTP**, token-gated | site → central | the deployment config itself (notify-and-fetch) | + +**The gRPC direction is inverted from the obvious guess.** Data flows site→central, but *each site +node hosts the gRPC server* (Kestrel h2c :8083) and *central is the client*. There is no gRPC server +on central at all. + +**Active node = OLDEST Up member, explicitly not the cluster leader.** +`ActiveNodeEvaluator.SelfIsOldestUp` is documented as *"THE single definition of 'active node'"*: + +> *"Cluster LEADERSHIP (lowest address) is an Akka-internal concept that diverges from singleton +> placement permanently once the original first node restarts and rejoins; every product-level +> active/standby decision must use this evaluator, never `cluster.State.Leader`."* + +The equivalence **oldest-Up == where `ClusterSingletonManager` places singletons** *is* the design. + +**All clusters share one ActorSystem name** (`"scadabridge"`, hardcoded). They are separate clusters +only by seed-node partitioning — required, because Akka.Remote address matching means ClusterClient +could not reach a differently-named system. + +**Site nodes carry two roles:** `"Site"` and `"site-{SiteId}"`. Singletons scope to the +**site-specific** role. + +Other load-bearing details: + +- **Central discovers sites from the database** (a `Site` entity with `NodeAAddress`/`NodeBAddress` + + `GrpcNodeAAddress`/`GrpcNodeBAddress`), refreshed every 60 s and on admin change. Sites discover + central from **appsettings** — static, restart required. The asymmetry is deliberate. +- **Exactly one actor is exposed per side** via `ClusterClientReceptionist.RegisterService` — + `/user/central-communication` and `/user/site-communication` — registered **per node, not as a + singleton**, so contact rotation reaches whichever node answers. +- **No central buffering when a site is unreachable.** The send is dropped with a warning and the + caller's Ask times out. *"It keeps the central coordinator stateless with respect to site + availability."* A `ConnectionStateChanged` mechanism was built for this and **deleted as dead + code**. +- **Every unhandled message replies with a typed failure rather than dropping**, so a central Ask + can never stall on a missing handler. +- **Sender preservation is the whole Ask idiom**: `_client.Tell(new ClusterClient.Send(path, msg), + Sender)` routes the reply straight back to the waiting Ask, bypassing the comm actor. + +### The registered outage gap — read before copying + +> `Component-ClusterInfrastructure.md:113`: *"Only the FIRST seed listed in `Cluster:SeedNodes` may +> self-join to form a *new* cluster… a lone restarted non-first-seed node (with the first seed still +> down) loops on `InitJoin` forever — never `Up`, never routable. This is why the two-node +> keep-oldest **oldest/active-node crash is a total-outage gap**: after the oldest dies the younger +> survivor self-downs, and it cannot re-bootstrap alone. Recovery is operator-driven."* + +Their failover drill has a mode that exists *"to make the registered gap observable — not to pretend +it is covered."* **A two-node keep-oldest mesh has an acknowledged total-outage hole.** If OtOpcUa +adopts 2-node meshes, this must be decided deliberately, not inherited. + +### Both inter-cluster transports are unauthenticated + +Akka remoting has no TLS, no secure cookie, no `trusted-selection-paths`. The gRPC listener is h2c +with no auth covering `SiteStreamService` — including two Pull RPCs that return audit rows. The only +authenticated pieces are the HTTP fetch token and `LocalDbSyncAuthInterceptor`, which gates *only* +`/localdb_sync.v1.LocalDbSync/`. Their boundary assumes a trusted network. + +**OtOpcUa already ships that same fail-closed interceptor** (LocalDb Phase 1) — it is a ready-made +template if we choose to authenticate more than they did. + +## 3. What OtOpcUa has today + +**Nine cross-node DPS topics plus a singleton**, all currently relying on one mesh: + +| Channel | Direction | Purpose | +|---|---|---| +| `deployments` / `deployment-acks` | central ↔ nodes | deploy dispatch + acks | +| `driver-control` | central → nodes | AdminUI Reconnect/Restart | +| `redundancy-state` | central → nodes | roles + peer probe results | +| `alerts` | nodes → central | fleet alarms → `/alerts` | +| `driver-health`, `driver-resilience-status`, `fleet-status`, `script-logs` | nodes → central | AdminUI live panels | +| `admin-operations` singleton | central | operator ack/shelve into engines | + +Three facts make the split cheaper than that table suggests: + +1. **The deploy path is already notify-and-fetch.** `DispatchDeployment` carries only + `DeploymentId` + `RevisionHash` + `CorrelationId`; the node fetches the artifact itself. We + independently arrived at the pattern ScadaBridge had to retrofit, and we do **not** carry its + 128 KB Akka frame exposure on this path. +2. **Deploy state already has a DB substrate.** `ConfigPublishCoordinator` persists per-node ACKs to + `NodeDeploymentState` so a singleton failover recovers in-flight state from the DB. +3. **`ClusterNode` rows already enumerate every node per application Cluster** — the same table + `AdminOperationsActor` groups by cluster. This replaces the coordinator's one genuinely + mesh-bound dependency: it derives its expected-ack set from `Akka.Cluster.State.Members` filtered + by role. + +So the deploy channel needs only a small notify + ack transport. **The nine live-telemetry channels +are the real work**, and they are exactly what ScadaBridge built ClusterClient + gRPC for. + +## 4. A defect to fix regardless of this design + +OtOpcUa currently uses **two different node-selection rules at once**: + +- Split-brain resolution is **keep-oldest** (`Core/ZB.MOM.WW.OtOpcUa.Cluster/Resources/akka.conf:49`). +- `ClusterSingletonManager` places singletons on the **oldest** member. +- But `RedundancyStateActor.BuildSnapshot` elects Primary from **`RoleLeader("driver")`** — lowest + address among driver members. + +Those select the same node only by coincidence. After a restart-and-rejoin the role leader and the +oldest member diverge permanently, which means the node the SBR protects and hosts every singleton +on need not be the node the data-plane gates consider Primary. This is precisely the rule ScadaBridge +forbids by name, and for the reason it gives: *"both sides claim leadership during a partition"* — +the dual-primary shape archreview 03/S4 exists to prevent. + +**Switching the role derivation to oldest-Up-with-role is small, is correct under either topology, +and should not wait for this design.** It also happens to be exactly what the separate-mesh model +needs, so it is not throwaway work. + +## 5. Target architecture + +Mirror ScadaBridge's shape, adapted to OtOpcUa's existing substrate: + +- **One Akka mesh per application `Cluster`, two nodes max.** Same ActorSystem name (`otopcua`) + everywhere; separation by seed-node partitioning only. +- **Roles per node:** `driver` plus a cluster-specific `cluster-{ClusterId}`, with singletons scoped + to the cluster-specific role. +- **Active node = oldest Up member with role** (§4), one definition, used by the data-plane gates, + ServiceLevel, and the alarm drain alike. +- **Central ↔ cluster command/control over ClusterClient**, with exactly one receptionist-registered + actor per side and central discovering cluster node addresses from `ClusterNode` rows (extended + with Akka + gRPC addresses, mirroring their `Site` entity). +- **Deploy stays notify-and-fetch.** The notify crosses via ClusterClient; the artifact continues to + come from the shared ConfigDb (see §6.1). The coordinator's expected-ack set comes from + `ClusterNode` rows instead of cluster membership. +- **Live telemetry over gRPC**, central dialling into each cluster node, replacing the seven + observability topics with one stream contract carrying a `oneof` event — additive-only field + evolution, contract-locked by test. + +Deliberately **not** copied: their transient-only central alarm cache and site-local-primary storage +(see §6.1), and their unauthenticated transports (§6.3). + +## 6. Open questions — these need answering before a plan + +### 6.1 Shared ConfigDb, or autonomous clusters? + +"Mimic ScadaBridge" splits into two very different commitments: + +- **Topology** — separate mesh per cluster. Contained, given §3. +- **Data architecture** — ScadaBridge sites are *autonomous*: their own SQLite, config fetched from + central over token-gated HTTP, because a site is not assumed to reach central's database at all. + +OtOpcUa's premise is the opposite: one shared ConfigDb every node reads directly, with LocalDb as an +**outage cache** — the entire basis of Phase 1's boot-from-cache. + +**Recommendation: take the topology, keep the shared ConfigDb.** Sites are autonomous in ScadaBridge +because they genuinely cannot rely on reaching central. If OtOpcUa's nodes can — and Phase 1 assumes +they normally can — then site-local-primary buys a large amount of machinery (staged pending rows, +fetch tokens, TTL purge singleton, reconciliation service) for a problem we do not have. Adopting it +would also reverse the premise of the LocalDb work just shipped. + +### 6.2 The two-node keep-oldest outage gap + +Inheriting 2-node keep-oldest inherits the registered hole in §2: the younger survivor cannot +re-bootstrap alone after the oldest dies. Options: accept it with an operator runbook (their choice); +add a third lightweight seed/witness per cluster; or use a different downing strategy. **This needs +an explicit decision — it is a total-outage hole, not a degradation.** + +### 6.3 Authentication on the new transports + +Their boundary assumes a trusted network. OtOpcUa's clusters may span less trusted links. If so we +need what they did not build: Akka.Remote TLS and a gRPC interceptor covering the whole service +surface. `LocalDbSyncAuthInterceptor` is already in our tree as a fail-closed template. + +## 7. Sequencing sketch + +Deliberately not a task plan — per-phase plans follow once §6 is settled. + +| Phase | Content | Independent of the split? | +|---|---|---| +| 0 | Oldest-Up role derivation (§4) | **Yes** — ship first | +| 1 | `ClusterNode` gains Akka + gRPC address columns; coordinator sources its expected-ack set from the DB | Yes | +| 2 | Comm actors + receptionist registration; ClusterClient transport; deploy notify + acks across the boundary | No | +| 3 | gRPC stream contract; migrate the seven observability topics | No | +| 4 | Mesh partition: per-cluster seed nodes, cluster-scoped roles + singletons; rig models the real topology | No | +| 5 | Auth per §6.3; failover drill; live gate | No | + +## 8. Risks + +- **The rig currently models the topology that would be abolished.** Phase 4 rewrites + `docker-dev/docker-compose.yml` substantially, and every live gate that depends on it. +- **Losing gossip loses free fleet-wide observability.** Seven AdminUI panels are fed by DPS today; + each needs an explicit stream and a reconnect story. +- **Akka frame size.** We are clean on the deploy path, but anything new that carries payload over + ClusterClient inherits the 128 KB default with `log-frame-size-exceeding` off — a silent + single-message drop that leaves the association healthy. If we send payload at all, set both. +- **Do not stack application-level last-write-wins on LocalDb's HLC.** Their `SiteStorageService` + carries this warning in code. Verified: our `deployment_pointer` upsert is unconditional and the + alarm sink is idempotent on a payload hash, so we are currently clean — keep it that way.