From 741ab97b896ddba4199a950764dab14ef7396d1b Mon Sep 17 00:00:00 2001 From: Joseph Doherty Date: Fri, 24 Jul 2026 02:38:26 -0400 Subject: [PATCH] feat(mesh-phase6): split docker-dev into three independent 2-node Akka meshes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Rewrite docker-dev/docker-compose.yml from one shared six-node gossip ring into three self-contained pairs (MAIN: central-1/2, SITE-A: site-a-1/2, SITE-B: site-b-1/2), all still on the ActorSystem name `otopcua` and the one shared docker network — separation is purely by per-pair self-first Cluster:SeedNodes, faithful to how production splits over a WAN rather than a firewall. Per node: - Cluster__Roles / OTOPCUA_ROLES gain a `cluster-` role (cluster-MAIN / cluster-SITE-A / cluster-SITE-B), the marker SplitTopologyTransportValidator and ClusterRoleInfo key off. - Cluster__SeedNodes now lists only the node's own pair (self first, partner second); site nodes no longer seed off central-1. - MeshTransport__Mode default flips Dps -> ClusterClient (still overridable via OTOPCUA_MESH_MODE) on all six nodes — SplitTopologyTransportValidator now refuses to boot a cluster-role node left on Dps. - Telemetry__Mode=Grpc added on all six driver nodes; TelemetryDial__Mode=Grpc added on central-1/central-2 only (the admin pair). - ConfigSource/ConfigServe modes and MeshTransport contact points are unchanged (already split-correct since Phases 3/4/2). seed-clusters.sql needed no changes — the ClusterNode rows already carry the correct ClusterId/Host/AkkaPort/GrpcPort per node. Validated with `docker compose -f docker-dev/docker-compose.yml config`. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW --- docker-dev/docker-compose.yml | 262 +++++++++++++++++++++------------- 1 file changed, 166 insertions(+), 96 deletions(-) diff --git a/docker-dev/docker-compose.yml b/docker-dev/docker-compose.yml index 5330a628..4d26abed 100644 --- a/docker-dev/docker-compose.yml +++ b/docker-dev/docker-compose.yml @@ -1,25 +1,58 @@ -# docker-dev/ — Mac-friendly single-mesh hub-and-spoke fleet for v2 development + manual UI exercise. +# docker-dev/ — Mac-friendly THREE-mesh hub-and-spoke fleet for v2 development + manual UI exercise. # -# Topology: ONE Akka mesh seeded by `central-1`. Logical separation between -# tenants is by ServerCluster.ClusterId rows (MAIN / SITE-A / SITE-B) in the one -# shared `OtOpcUa` ConfigDb — NOT by separate meshes. All six host nodes join the -# same gossip ring and the central UI deploys to every cluster over it. +# Topology (per-cluster mesh Phase 6, 2026-07-24): THREE independent 2-node Akka +# meshes, not one six-node mesh. Every node still runs the same ActorSystem name +# `otopcua` and all six containers still sit on the ONE shared docker-dev network +# — separation is purely by per-pair self-first Cluster__SeedNodes partitioning, +# never by network isolation. This is deliberately faithful to production, where +# the three meshes are separated by a WAN, not a firewall: nothing here stops +# central-1 opening a bare TCP connection to site-a-1's Akka port, it just never +# tries, because no seed list and no gossip ever names it. +# +# MAIN central-1 + central-2 (fused admin+driver, roles admin,driver,cluster-MAIN) +# SITE-A site-a-1 + site-a-2 (driver-only, roles driver,cluster-SITE-A) +# SITE-B site-b-1 + site-b-2 (driver-only, roles driver,cluster-SITE-B) +# +# Each pair seeds ONLY itself (self-first, partner-second) — no node lists a node +# outside its own mesh as a seed. Logical/tenant separation (ServerCluster.ClusterId +# rows MAIN / SITE-A / SITE-B in the one shared `OtOpcUa` ConfigDb) now COINCIDES +# with mesh separation: a `cluster-` Akka role marks which mesh + which +# ServerCluster row a node belongs to (RoleParser.ClusterRolePrefix = "cluster-"). +# +# Since central no longer shares a gossip ring with the site meshes, it reaches +# them over the three explicit split-topology transports instead of gossip: +# - Command plane: MeshTransport__Mode=ClusterClient (central dials each site +# mesh's ClusterClientReceptionist; DPS is gossip-only and +# would not cross the mesh boundary at all now) +# - Live telemetry: Telemetry__Mode=Grpc on every driver node (all six) + +# TelemetryDial__Mode=Grpc on the admin pair (central-1/2) +# — the site node hosts the gRPC server, central dials in +# - Deployed config: ConfigSource__Mode=FetchAndCache on the site nodes (already +# true since Phase 3/4 — driver-only nodes have no ConfigDb +# connection at all); central stays ConfigSource-less/Direct +# and keeps serving via ConfigServe. +# SplitTopologyTransportValidator (Cluster project) REFUSES to start any node +# carrying a `cluster-*` role that is left on a DPS-shaped transport — these +# three flips are therefore mandatory here, not optional dark switches. # # Stack: # sql SQL Server 2022 — hosts the one ConfigDb every node uses # cluster-seed one-shot mssql-tools job that INSERTs the ServerCluster + # ClusterNode rows scoping each tenant, then exits (idempotent) # -# central-1, central-2 OTOPCUA_ROLES=admin,driver — the ONLY UI + deploy -# singleton, plus the MAIN cluster's OPC UA publishers. -# Reachable at http://localhost:9200 (via Traefik). -# Both are Akka seed nodes, each listing ITSELF first -# (2026-07-22) so either can cold-start while the other -# is down; the site nodes seed off central-1 only. -# site-a-1, site-a-2 OTOPCUA_ROLES=driver — driver-only members of the same -# site-b-1, site-b-2 mesh, scoped to SITE-A / SITE-B by ClusterId. They -# serve no UI and authenticate no users; the central -# cluster manages and deploys to them over the mesh. +# central-1, central-2 OTOPCUA_ROLES=admin,driver,cluster-MAIN — the ONLY UI + +# deploy singleton, plus the MAIN cluster's OPC UA +# publishers. Reachable at http://localhost:9200 (via +# Traefik). Both are Akka seed nodes for the MAIN mesh +# ONLY, each listing ITSELF first (2026-07-22) so either +# can cold-start while the other is down. +# site-a-1, site-a-2 OTOPCUA_ROLES=driver,cluster-SITE-A — driver-only +# site-b-1, site-b-2 OTOPCUA_ROLES=driver,cluster-SITE-B — members of their +# OWN 2-node mesh (self-first seeded within the pair, NOT +# off central). They serve no UI and authenticate no +# users; central manages + deploys to them over the +# ClusterClient/gRPC/FetchAndCache transports above, not +# gossip. # # Auth is real LDAP against the shared GLAuth on the Linux Docker host # (10.100.0.35:3893, dc=zb,dc=local) — there is no LDAP container here. @@ -45,17 +78,12 @@ # listener, verified free against every other per-node port: Akka 4053, ConfigServe 4055 on # central, LocalDb sync 9001 on site-a, OPC UA 4840, HTTP 9000) + Telemetry__ApiKey # (committed dev key, matches the SQL password / secrets KEK / ConfigServe key exception). -# central-1/central-2 also carry the matching TelemetryDial__ApiKey. Both Telemetry__Mode and -# TelemetryDial__Mode are left UNSET here, so every node defaults to "Dps" (dark — telemetry -# keeps riding the existing DistributedPubSub topic). The ClusterNode.GrpcPort column (seeded -# by docker-dev/seed/seed-clusters.sql) points central's dial supervisor at each node's :4056. -# Unlike MeshTransport/ConfigSource, no `${VAR:-Dps}` interpolation is wired for either Mode key -# (a bare shell export does nothing here — these keys aren't referenced anywhere in this file, so -# there is nothing for it to substitute into). To run the Phase 5 live gate, manually add -# `Telemetry__Mode: "Grpc"` to the `environment:` block of all six driver nodes (central-1, -# central-2, site-a-1, site-a-2, site-b-1, site-b-2) and `TelemetryDial__Mode: "Grpc"` to -# central-1's and central-2's blocks — or supply both via a docker-compose.override.yml carrying -# the same keys — then `docker compose -f docker-dev/docker-compose.yml up -d` to recreate. +# central-1/central-2 also carry the matching TelemetryDial__ApiKey. As of Phase 6, Telemetry__Mode +# is set to "Grpc" on all six nodes and TelemetryDial__Mode is set to "Grpc" on central-1/central-2 +# ONLY — SplitTopologyTransportValidator requires it once a node carries a cluster-* role, so this +# is no longer an optional live-gate step (see the Phase 5 comment history in git blame for the +# prior dark-switch shape). The ClusterNode.GrpcPort column (seeded by +# docker-dev/seed/seed-clusters.sql) points central's dial supervisor at each node's :4056. # # Usage: # docker compose -f docker-dev/docker-compose.yml up -d --build @@ -84,6 +112,15 @@ name: otopcua-dev # AnnounceInterval is shortened to 5s (default 30s) so anti-entropy convergence is # observable quickly during dev exercise. Each node keeps its own local SQLite store # (Secrets:SqlitePath=otopcua-secrets.db, per-container); replication syncs them. +# +# Per-cluster mesh Phase 6: this DPS topic now rides each pair's OWN 2-node mesh, not one +# shared six-node mesh — so secret replication is PAIR-LOCAL. central-1/central-2 replicate +# with each other; site-a-1/site-a-2 replicate with each other; site-b-1/site-b-2 replicate +# with each other. There is no cross-cluster secret sync after the split (by design — the site +# meshes don't share a gossip ring with central or with each other, consistent with Phase 4's +# "driver-only nodes have no SQL, hence no SQL-backed cross-mesh hub" posture). No code change +# was needed for this — Enabled=true + the same KEK on every node is still correct; it just now +# converges within three independent 2-node groups instead of one six-node group. x-secrets-env: &secrets-env Secrets__Replication__Enabled: "true" Secrets__Replication__AnnounceInterval: "00:00:05" @@ -155,8 +192,9 @@ services: # ── Central cluster (2-node fused admin+driver) ───────────────────────────── # The only UI + deploy singleton; also the MAIN cluster's OPC UA publishers. - # central-1 seeds the single Akka mesh that every other node joins; central-2 is a seed too - # (listing itself first), so it can form the mesh alone if central-1 is down. + # Per-cluster mesh Phase 6: central-1 and central-2 form their OWN 2-node "MAIN" Akka mesh, + # NOT a shared mesh with the site nodes — each lists itself first as seed, so either can + # cold-start while the other is down, and the site meshes are never in this pair's seed list. central-1: &otopcua-host build: @@ -189,7 +227,7 @@ services: migrator: { condition: service_completed_successfully } environment: <<: *secrets-env - OTOPCUA_ROLES: "admin,driver" + OTOPCUA_ROLES: "admin,driver,cluster-MAIN" ASPNETCORE_URLS: "http://+:9000" ConnectionStrings__ConfigDb: "Server=sql,1433;Database=OtOpcUa;User Id=sa;Password=OtOpcUa!Dev123;TrustServerCertificate=True;" Cluster__Hostname: "0.0.0.0" @@ -203,18 +241,22 @@ services: ConfigServe__GrpcListenPort: "4055" ConfigServe__ApiKey: "configserve-docker-dev-key" ConfigSource__Mode: "Direct" - # Per-cluster mesh Phase 5 (gRPC live-telemetry stream). DARK SWITCH — Telemetry:Mode is left - # unset here (defaults to "Dps"), so the node keeps publishing telemetry on the existing - # DistributedPubSub topic; the dedicated gRPC listener below binds regardless (the node "always - # hosts" per docs/Telemetry.md) so flipping the mode later is a config change, not a redeploy. - # Port 4056 checked free against every other port this node binds: Akka 4053, ConfigServe 4055, - # OPC UA 4840 (container-internal), HTTP 9000. TelemetryDial__ApiKey (below) is central's - # dial-side key and must equal every node's Telemetry:ApiKey (fail-closed interceptor). + # Per-cluster mesh Phase 5/6 (gRPC live-telemetry stream). central-1 carries a cluster-* + # role + admin, so SplitTopologyTransportValidator requires BOTH Telemetry:Mode=Grpc + # (it is also a driver node, publishing MAIN's own telemetry) AND TelemetryDial:Mode=Grpc + # (it is the admin-role dialer for every site node). The dedicated gRPC listener below binds + # regardless of mode (the node "always hosts" per docs/Telemetry.md). Port 4056 checked free + # against every other port this node binds: Akka 4053, ConfigServe 4055, OPC UA 4840 + # (container-internal), HTTP 9000. TelemetryDial__ApiKey (below) is central's dial-side key + # and must equal every node's Telemetry:ApiKey (fail-closed interceptor). Telemetry__GrpcListenPort: "4056" + Telemetry__Mode: "Grpc" Telemetry__ApiKey: "telemetry-docker-dev-key" + TelemetryDial__Mode: "Grpc" TelemetryDial__ApiKey: "telemetry-docker-dev-key" # Both redundancy peers are seeds (#459 / ScadaBridge parity) so a restarted node can - # re-join the mesh via EITHER peer, not only central-1. + # re-join the MAIN mesh via EITHER peer — NOT via a site node; the site meshes are never in + # this pair's seed list (per-cluster mesh Phase 6). # # ORDER IS LOAD-BEARING (2026-07-22): each node lists ITSELF first. Akka runs # FirstSeedNodeProcess — the only path that can form a NEW cluster when no peer answers @@ -223,19 +265,25 @@ services: # boot by AkkaClusterOptionsValidator; see docs/Redundancy.md. Cluster__SeedNodes__0: "akka.tcp://otopcua@central-1:4053" Cluster__SeedNodes__1: "akka.tcp://otopcua@central-2:4053" + # Per-cluster mesh Phase 6: cluster-MAIN marks this node as a member of the MAIN mesh + # (RoleParser.ClusterRolePrefix = "cluster-"), read by SplitTopologyTransportValidator to + # require ClusterClient/Grpc transports, and by ClusterRoleInfo to scope redundancy + + # ClusterNodeAddressReconcilerActor to ClusterId=MAIN. Cluster__Roles__0: "admin" Cluster__Roles__1: "driver" - # Mesh command transport (per-cluster mesh Phase 2). DARK SWITCH: the rig comes up on "Dps", - # the transport this repo has always used, and every node keeps its ClusterClient wiring built - # and its comm actor registered with the receptionist either way. Flip the whole rig with - # OTOPCUA_MESH_MODE=ClusterClient at `docker compose up` — no compose edit, and no rebuild. + Cluster__Roles__2: "cluster-MAIN" + # Mesh command transport (per-cluster mesh Phase 2, made mandatory by Phase 6). Central no + # longer shares a gossip ring with the site meshes, so DPS (gossip-only) cannot reach them — + # ClusterClient is the only transport that can. SplitTopologyTransportValidator refuses to + # start any node carrying a cluster-* role that is left on Dps, so this is fixed, not an + # env-var-overridable default, on every node in this file. # # Contact points are node ADDRESSES ONLY. "/system/receptionist" is appended at construction # time, and MeshTransportOptionsValidator refuses to start a node whose contact carries an # actor-path suffix, because a doubled suffix resolves to nothing and drops every ack in # silence. Both central nodes are listed so a driver node's client can rotate to whichever one # is up — that rotation is why the comm actors are per-node and NOT cluster singletons. - MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-Dps}" + MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-ClusterClient}" MeshTransport__CentralContactPoints__0: "akka.tcp://otopcua@central-1:4053" MeshTransport__CentralContactPoints__1: "akka.tcp://otopcua@central-2:4053" Security__Jwt__SigningKey: "docker-dev-signing-key-with-at-least-32-bytes-of-utf8-content-12345" @@ -298,7 +346,7 @@ services: migrator: { condition: service_completed_successfully } environment: <<: *secrets-env - OTOPCUA_ROLES: "admin,driver" + OTOPCUA_ROLES: "admin,driver,cluster-MAIN" ASPNETCORE_URLS: "http://+:9000" ConnectionStrings__ConfigDb: "Server=sql,1433;Database=OtOpcUa;User Id=sa;Password=OtOpcUa!Dev123;TrustServerCertificate=True;" Cluster__Hostname: "0.0.0.0" @@ -309,29 +357,33 @@ services: ConfigServe__GrpcListenPort: "4055" ConfigServe__ApiKey: "configserve-docker-dev-key" ConfigSource__Mode: "Direct" - # Per-cluster mesh Phase 5 (gRPC live-telemetry stream) — see central-1. DARK SWITCH: Mode - # stays unset (⇒ Dps); the dedicated :4056 listener binds regardless. + # Per-cluster mesh Phase 5/6 (gRPC live-telemetry stream) — see central-1. central-2 carries + # a cluster-* role + admin, so both Telemetry:Mode=Grpc and TelemetryDial:Mode=Grpc are + # required (SplitTopologyTransportValidator); the dedicated :4056 listener binds regardless. Telemetry__GrpcListenPort: "4056" + Telemetry__Mode: "Grpc" Telemetry__ApiKey: "telemetry-docker-dev-key" + TelemetryDial__Mode: "Grpc" TelemetryDial__ApiKey: "telemetry-docker-dev-key" # Both redundancy peers are seeds (#459 / ScadaBridge parity) — see central-1. SELF FIRST: # central-2 lists itself as seed-nodes[0], which is what lets it cold-start while central-1 - # is down (it previously listed central-1 first and simply never came Up in that case). + # is down (it previously listed central-1 first and simply never came Up in that case). Per + # per-cluster mesh Phase 6, central-2's seed list contains ONLY the MAIN pair — no site node. Cluster__SeedNodes__0: "akka.tcp://otopcua@central-2:4053" Cluster__SeedNodes__1: "akka.tcp://otopcua@central-1:4053" + # Per-cluster mesh Phase 6: cluster-MAIN marks this node as a member of the MAIN mesh — see + # central-1's Cluster__Roles comment. Cluster__Roles__0: "admin" Cluster__Roles__1: "driver" - # Mesh command transport (per-cluster mesh Phase 2). DARK SWITCH: the rig comes up on "Dps", - # the transport this repo has always used, and every node keeps its ClusterClient wiring built - # and its comm actor registered with the receptionist either way. Flip the whole rig with - # OTOPCUA_MESH_MODE=ClusterClient at `docker compose up` — no compose edit, and no rebuild. + Cluster__Roles__2: "cluster-MAIN" + # Mesh command transport (per-cluster mesh Phase 2, made mandatory by Phase 6) — see central-1. # # Contact points are node ADDRESSES ONLY. "/system/receptionist" is appended at construction # time, and MeshTransportOptionsValidator refuses to start a node whose contact carries an # actor-path suffix, because a doubled suffix resolves to nothing and drops every ack in # silence. Both central nodes are listed so a driver node's client can rotate to whichever one # is up — that rotation is why the comm actors are per-node and NOT cluster singletons. - MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-Dps}" + MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-ClusterClient}" MeshTransport__CentralContactPoints__0: "akka.tcp://otopcua@central-1:4053" MeshTransport__CentralContactPoints__1: "akka.tcp://otopcua@central-2:4053" Security__Jwt__SigningKey: "docker-dev-signing-key-with-at-least-32-bytes-of-utf8-content-12345" @@ -379,9 +431,10 @@ services: - otopcua-localdb-central-2:/app/data # ── Site A cluster (2-node driver-only) ───────────────────────────────────── - # Driver-only members of the single mesh, scoped to SITE-A by ClusterId. No UI, - # no user auth — managed + deployed to by the central cluster over the mesh. - # All site nodes seed central-1. + # Per-cluster mesh Phase 6: site-a-1 and site-a-2 form their OWN 2-node "SITE-A" Akka mesh, + # separate from MAIN and SITE-B — self-first seeded WITHIN the pair, never off central. No UI, + # no user auth; central manages + deploys to this mesh over ClusterClient/gRPC/FetchAndCache, + # not gossip (central and site-a share no seed list, so they never form one ring). site-a-1: <<: *otopcua-host @@ -390,7 +443,7 @@ services: central-1: { condition: service_started } migrator: { condition: service_completed_successfully } environment: - OTOPCUA_ROLES: "driver" + OTOPCUA_ROLES: "driver,cluster-SITE-A" # Per-cluster mesh Phase 4: driver-only site nodes hold NO ConfigDb connection string. Their # config comes from central via ConfigSource:Mode=FetchAndCache (gRPC fetch → LocalDb cache), their # scripted-alarm condition state lives in LocalDb, and central persists their deploy acks. Direct @@ -406,29 +459,32 @@ services: ConfigSource__CentralFetchEndpoints__0: "http://central-1:4055" ConfigSource__CentralFetchEndpoints__1: "http://central-2:4055" ConfigSource__ApiKey: "configserve-docker-dev-key" - # Per-cluster mesh Phase 5 (gRPC live-telemetry stream) — see central-1/central-2. DARK - # SWITCH: Telemetry:Mode stays unset (⇒ Dps); the dedicated :4056 listener binds regardless - # (the node "always hosts" per docs/Telemetry.md). Port checked free against this node's - # other ports: Akka 4053, OPC UA 4840, HTTP 9000, LocalDb sync 9001 (below). ApiKey must - # equal central's TelemetryDial:ApiKey (fail-closed interceptor). + # Per-cluster mesh Phase 5/6 (gRPC live-telemetry stream) — see central-1/central-2. This + # node carries cluster-SITE-A + driver, so SplitTopologyTransportValidator requires + # Telemetry:Mode=Grpc (mandatory, not a dark switch, as of Phase 6). The dedicated :4056 + # listener binds regardless (the node "always hosts" per docs/Telemetry.md). Port checked + # free against this node's other ports: Akka 4053, OPC UA 4840, HTTP 9000, LocalDb sync 9001 + # (below). ApiKey must equal central's TelemetryDial:ApiKey (fail-closed interceptor). Telemetry__GrpcListenPort: "4056" + Telemetry__Mode: "Grpc" Telemetry__ApiKey: "telemetry-docker-dev-key" - # Site nodes are deliberately NOT seeds — they join the central pair's mesh. The self-first - # seed rule is conditional for exactly this reason (it binds only when a node's own address - # is in its own seed list), so these configs are exempt rather than broken. - Cluster__SeedNodes__0: "akka.tcp://otopcua@central-1:4053" + # Per-cluster mesh Phase 6: site-a-1 seeds its OWN "SITE-A" mesh (itself first, site-a-2 + # second) — it no longer seeds off central-1. Central is reached over ClusterClient/gRPC/ + # FetchAndCache, never gossip; central and site-a share no seed list. + Cluster__SeedNodes__0: "akka.tcp://otopcua@site-a-1:4053" + Cluster__SeedNodes__1: "akka.tcp://otopcua@site-a-2:4053" Cluster__Roles__0: "driver" - # Mesh command transport (per-cluster mesh Phase 2). DARK SWITCH: the rig comes up on "Dps", - # the transport this repo has always used, and every node keeps its ClusterClient wiring built - # and its comm actor registered with the receptionist either way. Flip the whole rig with - # OTOPCUA_MESH_MODE=ClusterClient at `docker compose up` — no compose edit, and no rebuild. + Cluster__Roles__1: "cluster-SITE-A" + # Mesh command transport (per-cluster mesh Phase 2, made mandatory by Phase 6) — see + # central-1/central-2. Central dials this mesh's ClusterClientReceptionist via the contact + # points below; there is no gossip path anymore. # # Contact points are node ADDRESSES ONLY. "/system/receptionist" is appended at construction # time, and MeshTransportOptionsValidator refuses to start a node whose contact carries an # actor-path suffix, because a doubled suffix resolves to nothing and drops every ack in # silence. Both central nodes are listed so a driver node's client can rotate to whichever one # is up — that rotation is why the comm actors are per-node and NOT cluster singletons. - MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-Dps}" + MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-ClusterClient}" MeshTransport__CentralContactPoints__0: "akka.tcp://otopcua@central-1:4053" MeshTransport__CentralContactPoints__1: "akka.tcp://otopcua@central-2:4053" <<: *secrets-env @@ -490,7 +546,7 @@ services: central-1: { condition: service_started } migrator: { condition: service_completed_successfully } environment: - OTOPCUA_ROLES: "driver" + OTOPCUA_ROLES: "driver,cluster-SITE-A" # Per-cluster mesh Phase 4: no ConfigDb string on driver-only site nodes (see site-a-1). Cluster__Hostname: "0.0.0.0" Cluster__Port: "4053" @@ -501,23 +557,27 @@ services: ConfigSource__CentralFetchEndpoints__0: "http://central-1:4055" ConfigSource__CentralFetchEndpoints__1: "http://central-2:4055" ConfigSource__ApiKey: "configserve-docker-dev-key" - # Per-cluster mesh Phase 5 (gRPC live-telemetry stream) — see central-1/site-a-1. DARK - # SWITCH: Telemetry:Mode stays unset (⇒ Dps); the dedicated :4056 listener binds regardless. + # Per-cluster mesh Phase 5/6 (gRPC live-telemetry stream) — see central-1/site-a-1. + # Telemetry:Mode=Grpc is mandatory as of Phase 6 (cluster-SITE-A + driver); the dedicated + # :4056 listener binds regardless. Telemetry__GrpcListenPort: "4056" + Telemetry__Mode: "Grpc" Telemetry__ApiKey: "telemetry-docker-dev-key" - Cluster__SeedNodes__0: "akka.tcp://otopcua@central-1:4053" + # Per-cluster mesh Phase 6: site-a-2 seeds its OWN "SITE-A" mesh (itself first, site-a-1 + # second) — it no longer seeds off central-1. + Cluster__SeedNodes__0: "akka.tcp://otopcua@site-a-2:4053" + Cluster__SeedNodes__1: "akka.tcp://otopcua@site-a-1:4053" Cluster__Roles__0: "driver" - # Mesh command transport (per-cluster mesh Phase 2). DARK SWITCH: the rig comes up on "Dps", - # the transport this repo has always used, and every node keeps its ClusterClient wiring built - # and its comm actor registered with the receptionist either way. Flip the whole rig with - # OTOPCUA_MESH_MODE=ClusterClient at `docker compose up` — no compose edit, and no rebuild. + Cluster__Roles__1: "cluster-SITE-A" + # Mesh command transport (per-cluster mesh Phase 2, made mandatory by Phase 6) — see + # central-1/site-a-1. # # Contact points are node ADDRESSES ONLY. "/system/receptionist" is appended at construction # time, and MeshTransportOptionsValidator refuses to start a node whose contact carries an # actor-path suffix, because a doubled suffix resolves to nothing and drops every ack in # silence. Both central nodes are listed so a driver node's client can rotate to whichever one # is up — that rotation is why the comm actors are per-node and NOT cluster singletons. - MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-Dps}" + MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-ClusterClient}" MeshTransport__CentralContactPoints__0: "akka.tcp://otopcua@central-1:4053" MeshTransport__CentralContactPoints__1: "akka.tcp://otopcua@central-2:4053" <<: *secrets-env @@ -563,6 +623,8 @@ services: - otopcua-localdb-site-a-2:/app/data # ── Site B cluster (2-node driver-only) ───────────────────────────────────── + # Per-cluster mesh Phase 6: site-b-1 and site-b-2 form their OWN 2-node "SITE-B" Akka mesh, + # separate from MAIN and SITE-A — same shape as site-a-1/site-a-2, see that pair's header. site-b-1: <<: *otopcua-host @@ -571,7 +633,7 @@ services: central-1: { condition: service_started } migrator: { condition: service_completed_successfully } environment: - OTOPCUA_ROLES: "driver" + OTOPCUA_ROLES: "driver,cluster-SITE-B" # Per-cluster mesh Phase 4: no ConfigDb string on driver-only site nodes (see site-a-1). Cluster__Hostname: "0.0.0.0" Cluster__Port: "4053" @@ -582,23 +644,27 @@ services: ConfigSource__CentralFetchEndpoints__0: "http://central-1:4055" ConfigSource__CentralFetchEndpoints__1: "http://central-2:4055" ConfigSource__ApiKey: "configserve-docker-dev-key" - # Per-cluster mesh Phase 5 (gRPC live-telemetry stream) — see central-1/site-a-1. DARK - # SWITCH: Telemetry:Mode stays unset (⇒ Dps); the dedicated :4056 listener binds regardless. + # Per-cluster mesh Phase 5/6 (gRPC live-telemetry stream) — see central-1/site-a-1. + # Telemetry:Mode=Grpc is mandatory as of Phase 6 (cluster-SITE-B + driver); the dedicated + # :4056 listener binds regardless. Telemetry__GrpcListenPort: "4056" + Telemetry__Mode: "Grpc" Telemetry__ApiKey: "telemetry-docker-dev-key" - Cluster__SeedNodes__0: "akka.tcp://otopcua@central-1:4053" + # Per-cluster mesh Phase 6: site-b-1 seeds its OWN "SITE-B" mesh (itself first, site-b-2 + # second) — it no longer seeds off central-1. + Cluster__SeedNodes__0: "akka.tcp://otopcua@site-b-1:4053" + Cluster__SeedNodes__1: "akka.tcp://otopcua@site-b-2:4053" Cluster__Roles__0: "driver" - # Mesh command transport (per-cluster mesh Phase 2). DARK SWITCH: the rig comes up on "Dps", - # the transport this repo has always used, and every node keeps its ClusterClient wiring built - # and its comm actor registered with the receptionist either way. Flip the whole rig with - # OTOPCUA_MESH_MODE=ClusterClient at `docker compose up` — no compose edit, and no rebuild. + Cluster__Roles__1: "cluster-SITE-B" + # Mesh command transport (per-cluster mesh Phase 2, made mandatory by Phase 6) — see + # central-1/site-a-1. # # Contact points are node ADDRESSES ONLY. "/system/receptionist" is appended at construction # time, and MeshTransportOptionsValidator refuses to start a node whose contact carries an # actor-path suffix, because a doubled suffix resolves to nothing and drops every ack in # silence. Both central nodes are listed so a driver node's client can rotate to whichever one # is up — that rotation is why the comm actors are per-node and NOT cluster singletons. - MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-Dps}" + MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-ClusterClient}" MeshTransport__CentralContactPoints__0: "akka.tcp://otopcua@central-1:4053" MeshTransport__CentralContactPoints__1: "akka.tcp://otopcua@central-2:4053" <<: *secrets-env @@ -635,7 +701,7 @@ services: central-1: { condition: service_started } migrator: { condition: service_completed_successfully } environment: - OTOPCUA_ROLES: "driver" + OTOPCUA_ROLES: "driver,cluster-SITE-B" # Per-cluster mesh Phase 4: no ConfigDb string on driver-only site nodes (see site-a-1). Cluster__Hostname: "0.0.0.0" Cluster__Port: "4053" @@ -646,23 +712,27 @@ services: ConfigSource__CentralFetchEndpoints__0: "http://central-1:4055" ConfigSource__CentralFetchEndpoints__1: "http://central-2:4055" ConfigSource__ApiKey: "configserve-docker-dev-key" - # Per-cluster mesh Phase 5 (gRPC live-telemetry stream) — see central-1/site-a-1. DARK - # SWITCH: Telemetry:Mode stays unset (⇒ Dps); the dedicated :4056 listener binds regardless. + # Per-cluster mesh Phase 5/6 (gRPC live-telemetry stream) — see central-1/site-a-1. + # Telemetry:Mode=Grpc is mandatory as of Phase 6 (cluster-SITE-B + driver); the dedicated + # :4056 listener binds regardless. Telemetry__GrpcListenPort: "4056" + Telemetry__Mode: "Grpc" Telemetry__ApiKey: "telemetry-docker-dev-key" - Cluster__SeedNodes__0: "akka.tcp://otopcua@central-1:4053" + # Per-cluster mesh Phase 6: site-b-2 seeds its OWN "SITE-B" mesh (itself first, site-b-1 + # second) — it no longer seeds off central-1. + Cluster__SeedNodes__0: "akka.tcp://otopcua@site-b-2:4053" + Cluster__SeedNodes__1: "akka.tcp://otopcua@site-b-1:4053" Cluster__Roles__0: "driver" - # Mesh command transport (per-cluster mesh Phase 2). DARK SWITCH: the rig comes up on "Dps", - # the transport this repo has always used, and every node keeps its ClusterClient wiring built - # and its comm actor registered with the receptionist either way. Flip the whole rig with - # OTOPCUA_MESH_MODE=ClusterClient at `docker compose up` — no compose edit, and no rebuild. + Cluster__Roles__1: "cluster-SITE-B" + # Mesh command transport (per-cluster mesh Phase 2, made mandatory by Phase 6) — see + # central-1/site-a-1. # # Contact points are node ADDRESSES ONLY. "/system/receptionist" is appended at construction # time, and MeshTransportOptionsValidator refuses to start a node whose contact carries an # actor-path suffix, because a doubled suffix resolves to nothing and drops every ack in # silence. Both central nodes are listed so a driver node's client can rotate to whichever one # is up — that rotation is why the comm actors are per-node and NOT cluster singletons. - MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-Dps}" + MeshTransport__Mode: "${OTOPCUA_MESH_MODE:-ClusterClient}" MeshTransport__CentralContactPoints__0: "akka.tcp://otopcua@central-1:4053" MeshTransport__CentralContactPoints__1: "akka.tcp://otopcua@central-2:4053" <<: *secrets-env