Files
ScadaBridge/docker
Joseph Doherty 4a6341d871 feat(cluster): self-first seed ordering closes the boot-alone outage gap
Every node now lists ITSELF as seed-nodes[0] and its partner second. Akka runs
FirstSeedNodeProcess -- the only bootstrap path that can form a NEW cluster when
no peer answers InitJoin -- exclusively for seed-nodes[0]; every other node runs
JoinSeedNodeProcess and retries InitJoin forever. That is why a lone cold-starting
central-b never came Up (the "registered outage gap"), and self-first ordering
closes it using Akka's own protocol.

- 6 node appsettings swapped (the *-node-b configs; the -a nodes were already
  self-first). All 14 shipped node configs now satisfy the invariant.
- StartupValidator enforces it at boot, comparing host AND port -- the invariant
  fails silently when broken, so it is enforced loudly. NOTE: the gitignored
  deploy/wonder-app-vd03/ overlay must be reordered before its next deploy or
  that node will refuse to boot.
- SelfFirstSeedBootstrapTests: real in-process clusters at production
  failure-detection timings, incl. a falsifiability control proving the OLD
  peer-first ordering never forms.

Rejected alternative (implemented, measured, discarded): an external self-form
timer calling Cluster.Join(SelfAddress) after a window. It sits outside Akka's
join handshake and so cannot tell "no seed answered" from "a seed answered and
the join is in flight". On a routine standby restart the peer is alive but the
join stalls behind removal of the node's own stale incarnation; a Join(self)
during TryingToJoin abandons the in-flight join and forms a second cluster at
the same address -- still split after 90s. Docs that claimed self-first ordering
was unsafe for simultaneous cold start are corrected: while mutually reachable
the InitJoin handshake converges them to one cluster (measured).
2026-07-22 06:32:00 -04:00
..

ScadaBridge Docker Infrastructure

Local Docker deployment of the full ScadaBridge cluster topology: a 2-node central cluster and three 2-node site clusters.

Cluster Topology

              ┌───────────────────┐
              │  Traefik LB :9000 │  ◄── CLI / Browser
              │  Dashboard :8180  │
              └────────┬──────────┘
                       │ routes to active node
┌──────────────────────┼──────────────────────────────┐
│                  Central Cluster                    │
│                                                     │
│  ┌─────────────────┐     ┌─────────────────┐        │
│  │  central-node-a  │◄──►│  central-node-b  │       │
│  │  (leader/oldest) │     │  (standby)       │       │
│  │  Web UI :9001    │     │  Web UI :9002    │       │
│  │  Akka   :9011    │     │  Akka   :9012    │       │
│  └────────┬─────────┘     └─────────────────┘       │
│           │                                         │
└───────────┼─────────────────────────────────────────┘
            │ Akka.NET Remoting (hub-and-spoke)
            ├──────────────────┬──────────────────┐
            ▼                  ▼                  ▼
┌────────────────────┐ ┌────────────────────┐ ┌────────────────────┐
│  Site-A Cluster    │ │  Site-B Cluster    │ │  Site-C Cluster    │
│  (Test Plant A)    │ │  (Test Plant B)    │ │  (Test Plant C)    │
│                    │ │                    │ │                    │
│  node-a ◄──► node-b│ │  node-a ◄──► node-b│ │  node-a ◄──► node-b│
│  Akka :9021 :9022  │ │  Akka :9031 :9032  │ │  Akka :9041 :9042  │
│  gRPC :9023 :9024  │ │  gRPC :9033 :9034  │ │  gRPC :9043 :9044  │
└────────────────────┘ └────────────────────┘ └────────────────────┘

Central Cluster (active/standby)

Runs the web UI (Blazor Server), Template Engine, Deployment Manager, Security, Inbound API, Management Service, and Health Monitoring. Connects to MS SQL for configuration and machine data, LDAP for authentication, and SMTP for notifications.

Site Clusters (active/standby each)

Each site cluster runs Site Runtime, Data Connection Layer, Store-and-Forward, and Site Event Logging. Sites connect to OPC UA for device data and to the central cluster via Akka.NET remoting. Each site node also hosts a gRPC streaming server (port 8083) that central nodes connect to for real-time attribute value and alarm state streams. Deployed configurations and S&F buffers are stored in local SQLite databases per node.

Site Cluster Site Identifier Central UI Name
Site-A site-a Test Plant A
Site-B site-b Test Plant B
Site-C site-c Test Plant C

Port Allocation

Application Nodes

Node Container Name Host Web Port Host Akka Port Host gRPC Port Internal Ports
Traefik LB scadabridge-traefik 9000 80 (proxy), 8080 (dashboard)
Central A scadabridge-central-a 9001 9011 5000 (web), 8081 (Akka)
Central B scadabridge-central-b 9002 9012 5000 (web), 8081 (Akka)
Site-A A scadabridge-site-a-a 9021 9023 8082 (Akka), 8083 (gRPC)
Site-A B scadabridge-site-a-b 9022 9024 8082 (Akka), 8083 (gRPC)
Site-B A scadabridge-site-b-a 9031 9033 8082 (Akka), 8083 (gRPC)
Site-B B scadabridge-site-b-b 9032 9034 8082 (Akka), 8083 (gRPC)
Site-C A scadabridge-site-c-a 9041 9043 8082 (Akka), 8083 (gRPC)
Site-C B scadabridge-site-c-b 9042 9044 8082 (Akka), 8083 (gRPC)

Port block pattern: 90X1/90X2 (Akka), 90X3/90X4 (gRPC) where X = 0 (central), 2 (site-a), 3 (site-b), 4 (site-c). gRPC streaming ports are used by central nodes to subscribe to real-time site data streams.

Infrastructure Services (from infra/docker-compose.yml)

Service Container Name Host Port Purpose
MS SQL 2022 scadabridge-mssql 1433 Configuration and machine data databases
LDAP (GLAuth) scadabridge-ldap 3893 Authentication with test users
SMTP (Mailpit) scadabridge-smtp 1025 / 8025 Email capture (SMTP / web UI)
OPC UA scadabridge-opcua 50000 / 8080 Simulated OPC UA server (protocol / web UI)
REST API scadabridge-restapi 5200 External REST API for integration testing

All containers communicate over the shared scadabridge-net Docker bridge network using container names as hostnames.

Directory Structure

docker/
├── Dockerfile                          # Multi-stage build (shared by all nodes)
├── docker-compose.yml                  # 8-node application stack
├── build.sh                            # Build Docker image
├── deploy.sh                           # Build + deploy all containers
├── seed-sites.sh                       # Create test sites with Akka + gRPC addresses
├── teardown.sh                         # Stop and remove containers
├── central-node-a/
│   ├── appsettings.Central.json        # Central node A configuration
│   └── logs/                           # Serilog file output (gitignored)
├── central-node-b/
│   ├── appsettings.Central.json
│   └── logs/
├── site-a-node-a/
│   ├── appsettings.Site.json           # Site-A node A configuration
│   ├── data/                           # SQLite databases (gitignored)
│   └── logs/
├── site-a-node-b/
│   ├── appsettings.Site.json
│   ├── data/
│   └── logs/
├── site-b-node-a/
│   ├── appsettings.Site.json           # Site-B node A configuration
│   ├── data/
│   └── logs/
├── site-b-node-b/
│   ├── appsettings.Site.json
│   ├── data/
│   └── logs/
├── site-c-node-a/
│   ├── appsettings.Site.json           # Site-C node A configuration
│   ├── data/
│   └── logs/
└── site-c-node-b/
    ├── appsettings.Site.json
    ├── data/
    └── logs/

Commands

Initial Setup

Start infrastructure services first, then build and deploy the application:

# 1. Start test infrastructure (MS SQL, LDAP, SMTP, OPC UA)
cd infra && docker compose up -d && cd ..

# 2. Build and deploy all 8 ScadaBridge nodes
docker/deploy.sh

# 3. Seed test sites (first-time only, after cluster is healthy)
docker/seed-sites.sh

After Code Changes

Rebuild and redeploy. The Docker build cache skips NuGet restore when only source files change:

docker/deploy.sh

Stop Application Nodes

Stops and removes all 8 application containers. Site SQLite databases and log files are preserved in node directories:

docker/teardown.sh

Stop Everything

docker/teardown.sh
cd infra && docker compose down && cd ..

View Logs

# All nodes (follow mode)
docker compose -f docker/docker-compose.yml logs -f

# Single node
docker logs -f scadabridge-central-a

# Filter by site cluster
docker compose -f docker/docker-compose.yml logs -f site-a-a site-a-b
docker compose -f docker/docker-compose.yml logs -f site-b-a site-b-b
docker compose -f docker/docker-compose.yml logs -f site-c-a site-c-b

# Persisted log files
ls docker/central-node-a/logs/

Restart a Single Node

docker restart scadabridge-central-a

Check Cluster Health

# Central node A health check
curl -s http://localhost:9001/health/ready | python3 -m json.tool

# Central node B health check
curl -s http://localhost:9002/health/ready | python3 -m json.tool

CLI Access

The CLI connects to the Central Host's HTTP management API via the Traefik load balancer at http://localhost:9000, which routes to the active central node:

dotnet run --project src/ZB.MOM.WW.ScadaBridge.CLI -- \
    --url http://localhost:9000 \
    --username multi-role --password password \
    template list

Direct access to individual nodes is also available at http://localhost:9001 (central-a) and http://localhost:9002 (central-b).

Note: The multi-role test user has Admin, Design, and Deployment roles. The admin user only has the Admin role and cannot perform design or deployment operations. See infra/glauth/config.toml for all test users and their group memberships.

A recommended ~/.scadabridge/config.json for the Docker test environment:

{
  "managementUrl": "http://localhost:9000"
}

With this config file in place, the URL is automatic:

dotnet run --project src/ZB.MOM.WW.ScadaBridge.CLI -- \
    --username multi-role --password password \
    template list

Clear Site Data

Remove SQLite databases to reset site state (deployed configs, S&F buffers):

# Single site
rm -rf docker/site-a-node-a/data docker/site-a-node-b/data
docker restart scadabridge-site-a-a scadabridge-site-a-b

# All sites
rm -rf docker/site-*/data
docker restart scadabridge-site-a-a scadabridge-site-a-b \
    scadabridge-site-b-a scadabridge-site-b-b \
    scadabridge-site-c-a scadabridge-site-c-b

Rebuild Image From Scratch (no cache)

docker build --no-cache -t scadabridge:latest -f docker/Dockerfile .

Build Cache

The Dockerfile uses a multi-stage build optimized for fast rebuilds:

  1. Restore stage: Copies only .csproj files and runs dotnet restore. This layer is cached as long as no project file changes.
  2. Build stage: Copies source code and runs dotnet publish --no-restore. Re-runs on any source change but skips restore.
  3. Runtime stage: Uses the slim aspnet:10.0 base image with only the published output.

Typical rebuild after a source-only change takes ~5 seconds (restore cached, only build + publish runs).

Test Users

All test passwords are password. See infra/glauth/config.toml for the full list.

Username Roles Use Case
admin Admin System administration
designer Design Template authoring
deployer Deployment Instance deployment (all sites)
multi-role Admin, Design, Deployment Full access for testing

Failover Testing

Automated failover drill (failover-drill.sh)

DRILL_MODE=standby bash docker/failover-drill.sh   # default — younger-node crash, active untouched
DRILL_MODE=active  bash docker/failover-drill.sh    # oldest-node crash — survivor must TAKE OVER

The scripted drill (docker kill = SIGKILL, the hard-crash path — a docker stop would take the graceful CoordinatedShutdown path and would not prove crash recovery) has two modes, and since the auto-down decision (2026-07-21) both expect recovery — the cluster runs Akka's AutoDowning provider (auto-down-unreachable-after = 15s), under which the leader among the reachable members downs the unreachable peer, so a crash of either node fails over:

  • DRILL_MODE=standby (default) — kills the STANDBY (younger) central node. The active node is untouched: expected result is no routing outage at all (/health/active blips = 0) and member removal on the survivor within ~25s (10s failure-detection threshold + 15s auto-down window; the 2s heartbeat interval is not additive). PASS = the survivor logs the downing/removal within TIMEOUT_S (default 90s) while routing stays up.
  • DRILL_MODE=active — kills the ACTIVE (oldest) central node. The survivor must take over while the victim is still down: it auto-downs the dead oldest, becomes the oldest member itself, re-hosts all singletons, and its /health/active goes 200. PASS = survivor active within TIMEOUT_S, then Traefik routing to it. (Under the pre-2026-07-21 keep-oldest strategy this direction was a proven total outage — the younger survivor took DownReachable and downed itself, because Akka's down-if-alone only rescues a side with ≥ 2 members.)

Both modes finish by restarting the victim and confirming it rejoins as a ready standby. The drill exercises downing-on-hard-crash, S3 (single active node routed through Traefik), and the Task 20 restart/rejoin contract. Requires a running cluster (bash docker/deploy.sh) and curl + docker on the host.

Partition trade (accepted). Auto-down is availability-first: in a real network partition (both nodes alive, link cut) each side downs the other and both run active — dual-active until an operator restarts one side after the partition heals. This was an explicit owner decision (2026-07-21): site pairs have no shared lease infrastructure to arbitrate, and a stalled system is a bigger risk than a rare partition. See docs/plans/2026-07-21-auto-down-availability-decision.md.

Seed-node ordering — every node lists ITSELF first (decision 2026-07-22). Akka runs FirstSeedNodeProcess — the only bootstrap path that can form a new cluster when no peer answers InitJoin — exclusively when seed-nodes[0] is the node's own address; every other node runs JoinSeedNodeProcess, which retries InitJoin forever and can never form a cluster. Each shipped node config therefore lists itself first and its partner second (docker/central-node-b/appsettings.Central.json leads with scadabridge-central-b), and StartupValidator fails the boot if that ordering is ever broken. This closes the former registered outage gap, where a lone cold-starting central-b (with central-a down) never came Up and recovery was operator-driven.

Self-first ordering is safe, and the three interesting cases are covered by SelfFirstSeedBootstrapTests (real in-process clusters at production failure-detection timings):

Scenario Behavior
Lone cold-start, peer dead Forms alone in ~5s (seed-node-timeout) — operational, unattended
Restart into a live peer InitJoinAck answers, node rejoins; never islands
Both cold-start simultaneously (mutually reachable) The InitJoin handshake resolves it before either self-joins → one 2-member cluster

An earlier revision of this README claimed the repo deliberately avoided self-first ordering because simultaneous cold start would produce "two one-node clusters that never merge". That is not what happens while the nodes are mutually reachable — the handshake converges them (measured, row 3 above). Only a genuine boot-time partition splits them, which is the same class auto-down already accepts.

Rejected alternative — an external self-form timer. A watchdog that waits N seconds for membership and then calls Cluster.Join(SelfAddress) was implemented and discarded: it cannot see Akka's join handshake, so it cannot distinguish "no seed answered" from "a seed answered and the join is in flight". On a routine standby restart the peer is alive but the join stalls behind removal of the restarting node's own stale incarnation; a Join(self) issued during TryingToJoin abandons the in-flight join and forms a second cluster at the same address — a permanent split (measured: still split after 90s). Akka's own first-seed process has no such race because it is part of the handshake.

Observed results (auto-down decision verification):

Run 2026-07-21 against a freshly-deployed cluster with SplitBrainResolverStrategy: auto-down (first drill: active=central-a). Both directions recovered.

Direction (DRILL_MODE) Outcome Measured
active (oldest-node crash) PASS — TAKEOVERcentral-b auto-downed the dead oldest, went Younger -> Oldest on all 7 singletons, and served /health/active while the victim was still down; restarted victim rejoined as standby. Survivor active + Traefik routing in 28s (budget ~25s: 10s detection + 15s auto-down + hand-over); victim ready 2s after restart.
standby (younger-node crash) PASS — active node untouched; survivor downed+removed the crashed member; restarted victim rejoined as standby. Member removed in 27s; 0 /health/active routing blips; victim ready 2s after restart.

Historical baseline (keep-oldest, run 2026-07-13 on 99544985): standby PASS with member removal in 27s / 0 routing blips; active was a total outagecentral-b self-downed ~20s after the kill (live SBR log 2026-07-21: SBR took decision Akka.Cluster.SBR.DownReachable … including myself) and could not re-bootstrap until central-a returned. That result is what motivated the auto-down decision. In-process envelope (FailoverTimingTests) measured full failover at 33.7s.

Central Failover

# Stop the active central node
docker stop scadabridge-central-a

# Verify central-b takes over (check logs for leader election)
docker logs -f scadabridge-central-b

# Access UI on standby node
open http://localhost:9002

# Restore the original node
docker start scadabridge-central-a

Site Failover

# Stop the active site-a node
docker stop scadabridge-site-a-a

# Verify site-a-b takes over singleton (DeploymentManager)
docker logs -f scadabridge-site-a-b

# Restore
docker start scadabridge-site-a-a

Same pattern applies for site-b (scadabridge-site-b-a/scadabridge-site-b-b) and site-c (scadabridge-site-c-a/scadabridge-site-c-b).

Failover takes approximately 25 seconds (2s heartbeat + 10s detection threshold + 15s stable-after for split-brain resolver).