# ScadaBridge Docker Infrastructure Local Docker deployment of the full ScadaBridge cluster topology: a 2-node central cluster and three 2-node site clusters. ## Cluster Topology ``` ┌───────────────────┐ │ Traefik LB :9000 │ ◄── CLI / Browser │ Dashboard :8180 │ └────────┬──────────┘ │ routes to active node ┌──────────────────────┼──────────────────────────────┐ │ Central Cluster │ │ │ │ ┌─────────────────┐ ┌─────────────────┐ │ │ │ central-node-a │◄──►│ central-node-b │ │ │ │ (leader/oldest) │ │ (standby) │ │ │ │ Web UI :9001 │ │ Web UI :9002 │ │ │ │ Akka :9011 │ │ Akka :9012 │ │ │ └────────┬─────────┘ └─────────────────┘ │ │ │ │ └───────────┼─────────────────────────────────────────┘ │ Akka.NET Remoting (hub-and-spoke) ├──────────────────┬──────────────────┐ ▼ ▼ ▼ ┌────────────────────┐ ┌────────────────────┐ ┌────────────────────┐ │ Site-A Cluster │ │ Site-B Cluster │ │ Site-C Cluster │ │ (Test Plant A) │ │ (Test Plant B) │ │ (Test Plant C) │ │ │ │ │ │ │ │ node-a ◄──► node-b│ │ node-a ◄──► node-b│ │ node-a ◄──► node-b│ │ Akka :9021 :9022 │ │ Akka :9031 :9032 │ │ Akka :9041 :9042 │ │ gRPC :9023 :9024 │ │ gRPC :9033 :9034 │ │ gRPC :9043 :9044 │ └────────────────────┘ └────────────────────┘ └────────────────────┘ ``` ### Central Cluster (active/standby) Runs the web UI (Blazor Server), Template Engine, Deployment Manager, Security, Inbound API, Management Service, and Health Monitoring. Connects to MS SQL for configuration and machine data, LDAP for authentication, and SMTP for notifications. ### Site Clusters (active/standby each) Each site cluster runs Site Runtime, Data Connection Layer, Store-and-Forward, and Site Event Logging. Sites connect to OPC UA for device data and to the central cluster via Akka.NET remoting. Each site node also hosts a gRPC streaming server (port 8083) that central nodes connect to for real-time attribute value and alarm state streams. Deployed configurations and S&F buffers are stored in local SQLite databases per node. | Site Cluster | Site Identifier | Central UI Name | |-------------|-----------------|-----------------| | Site-A | `site-a` | Test Plant A | | Site-B | `site-b` | Test Plant B | | Site-C | `site-c` | Test Plant C | ## Port Allocation ### Application Nodes | Node | Container Name | Host Web Port | Host Akka Port | Host gRPC Port | Internal Ports | |------|---------------|---------------|----------------|----------------|----------------| | Traefik LB | `scadabridge-traefik` | 9000 | — | — | 80 (proxy), 8080 (dashboard) | | Central A | `scadabridge-central-a` | 9001 | 9011 | — | 5000 (web), 8081 (Akka) | | Central B | `scadabridge-central-b` | 9002 | 9012 | — | 5000 (web), 8081 (Akka) | | Site-A A | `scadabridge-site-a-a` | — | 9021 | 9023 | 8082 (Akka), 8083 (gRPC) | | Site-A B | `scadabridge-site-a-b` | — | 9022 | 9024 | 8082 (Akka), 8083 (gRPC) | | Site-B A | `scadabridge-site-b-a` | — | 9031 | 9033 | 8082 (Akka), 8083 (gRPC) | | Site-B B | `scadabridge-site-b-b` | — | 9032 | 9034 | 8082 (Akka), 8083 (gRPC) | | Site-C A | `scadabridge-site-c-a` | — | 9041 | 9043 | 8082 (Akka), 8083 (gRPC) | | Site-C B | `scadabridge-site-c-b` | — | 9042 | 9044 | 8082 (Akka), 8083 (gRPC) | Port block pattern: `90X1`/`90X2` (Akka), `90X3`/`90X4` (gRPC) where X = 0 (central), 2 (site-a), 3 (site-b), 4 (site-c). gRPC streaming ports are used by central nodes to subscribe to real-time site data streams. ### Infrastructure Services (from `infra/docker-compose.yml`) | Service | Container Name | Host Port | Purpose | |---------|---------------|-----------|---------| | MS SQL 2022 | `scadabridge-mssql` | 1433 | Configuration and machine data databases | | LDAP (GLAuth) | `scadabridge-ldap` | 3893 | Authentication with test users | | SMTP (Mailpit) | `scadabridge-smtp` | 1025 / 8025 | Email capture (SMTP / web UI) | | OPC UA | `scadabridge-opcua` | 50000 / 8080 | Simulated OPC UA server (protocol / web UI) | | REST API | `scadabridge-restapi` | 5200 | External REST API for integration testing | All containers communicate over the shared `scadabridge-net` Docker bridge network using container names as hostnames. ## Directory Structure ``` docker/ ├── Dockerfile # Multi-stage build (shared by all nodes) ├── docker-compose.yml # 8-node application stack ├── build.sh # Build Docker image ├── deploy.sh # Build + deploy all containers ├── seed-sites.sh # Create test sites with Akka + gRPC addresses ├── teardown.sh # Stop and remove containers ├── central-node-a/ │ ├── appsettings.Central.json # Central node A configuration │ └── logs/ # Serilog file output (gitignored) ├── central-node-b/ │ ├── appsettings.Central.json │ └── logs/ ├── site-a-node-a/ │ ├── appsettings.Site.json # Site-A node A configuration │ ├── data/ # SQLite databases (gitignored) │ └── logs/ ├── site-a-node-b/ │ ├── appsettings.Site.json │ ├── data/ │ └── logs/ ├── site-b-node-a/ │ ├── appsettings.Site.json # Site-B node A configuration │ ├── data/ │ └── logs/ ├── site-b-node-b/ │ ├── appsettings.Site.json │ ├── data/ │ └── logs/ ├── site-c-node-a/ │ ├── appsettings.Site.json # Site-C node A configuration │ ├── data/ │ └── logs/ └── site-c-node-b/ ├── appsettings.Site.json ├── data/ └── logs/ ``` ## Commands ### Initial Setup Start infrastructure services first, then build and deploy the application: ```bash # 1. Start test infrastructure (MS SQL, LDAP, SMTP, OPC UA) cd infra && docker compose up -d && cd .. # 2. Build and deploy all 8 ScadaBridge nodes docker/deploy.sh # 3. Seed test sites (first-time only, after cluster is healthy) docker/seed-sites.sh ``` ### After Code Changes Rebuild and redeploy. The Docker build cache skips NuGet restore when only source files change: ```bash docker/deploy.sh ``` ### Stop Application Nodes Stops and removes all 8 application containers. Site SQLite databases and log files are preserved in node directories: ```bash docker/teardown.sh ``` ### Stop Everything ```bash docker/teardown.sh cd infra && docker compose down && cd .. ``` ### View Logs ```bash # All nodes (follow mode) docker compose -f docker/docker-compose.yml logs -f # Single node docker logs -f scadabridge-central-a # Filter by site cluster docker compose -f docker/docker-compose.yml logs -f site-a-a site-a-b docker compose -f docker/docker-compose.yml logs -f site-b-a site-b-b docker compose -f docker/docker-compose.yml logs -f site-c-a site-c-b # Persisted log files ls docker/central-node-a/logs/ ``` ### Restart a Single Node ```bash docker restart scadabridge-central-a ``` ### Check Cluster Health ```bash # Central node A health check curl -s http://localhost:9001/health/ready | python3 -m json.tool # Central node B health check curl -s http://localhost:9002/health/ready | python3 -m json.tool ``` ### CLI Access The CLI connects to the Central Host's HTTP management API via the Traefik load balancer at `http://localhost:9000`, which routes to the active central node: ```bash dotnet run --project src/ZB.MOM.WW.ScadaBridge.CLI -- \ --url http://localhost:9000 \ --username multi-role --password password \ template list ``` Direct access to individual nodes is also available at `http://localhost:9001` (central-a) and `http://localhost:9002` (central-b). > **Note:** The `multi-role` test user has Admin, Design, and Deployment roles. The `admin` user only has the Admin role and cannot perform design or deployment operations. See `infra/glauth/config.toml` for all test users and their group memberships. A recommended `~/.scadabridge/config.json` for the Docker test environment: ```json { "managementUrl": "http://localhost:9000" } ``` With this config file in place, the URL is automatic: ```bash dotnet run --project src/ZB.MOM.WW.ScadaBridge.CLI -- \ --username multi-role --password password \ template list ``` ### Clear Site Data Remove SQLite databases to reset site state (deployed configs, S&F buffers): ```bash # Single site rm -rf docker/site-a-node-a/data docker/site-a-node-b/data docker restart scadabridge-site-a-a scadabridge-site-a-b # All sites rm -rf docker/site-*/data docker restart scadabridge-site-a-a scadabridge-site-a-b \ scadabridge-site-b-a scadabridge-site-b-b \ scadabridge-site-c-a scadabridge-site-c-b ``` ### Rebuild Image From Scratch (no cache) ```bash docker build --no-cache -t scadabridge:latest -f docker/Dockerfile . ``` ## Build Cache The Dockerfile uses a multi-stage build optimized for fast rebuilds: 1. **Restore stage**: Copies only `.csproj` files and runs `dotnet restore`. This layer is cached as long as no project file changes. 2. **Build stage**: Copies source code and runs `dotnet publish --no-restore`. Re-runs on any source change but skips restore. 3. **Runtime stage**: Uses the slim `aspnet:10.0` base image with only the published output. Typical rebuild after a source-only change takes ~5 seconds (restore cached, only build + publish runs). ## Test Users All test passwords are `password`. See `infra/glauth/config.toml` for the full list. | Username | Roles | Use Case | |----------|-------|----------| | `admin` | Admin | System administration | | `designer` | Design | Template authoring | | `deployer` | Deployment | Instance deployment (all sites) | | `multi-role` | Admin, Design, Deployment | Full access for testing | ## Failover Testing ### Automated failover drill (`failover-drill.sh`) ```bash DRILL_MODE=standby bash docker/failover-drill.sh # default — younger-node crash, active untouched DRILL_MODE=active bash docker/failover-drill.sh # oldest-node crash — survivor must TAKE OVER ``` The scripted drill (`docker kill` = SIGKILL, the hard-crash path — a `docker stop` would take the graceful `CoordinatedShutdown` path and would not prove crash recovery) has **two modes**, and since the **auto-down decision (2026-07-21)** both expect recovery — the cluster runs Akka's `AutoDowning` provider (`auto-down-unreachable-after` = 15s), under which the leader among the *reachable* members downs the unreachable peer, so a crash of either node fails over: - **`DRILL_MODE=standby` (default) — kills the STANDBY (younger) central node.** The active node is untouched: expected result is **no routing outage at all** (`/health/active` blips = 0) and member removal on the survivor within **~25s** (10s failure-detection threshold + 15s auto-down window; the 2s heartbeat interval is not additive). PASS = the survivor logs the downing/removal within `TIMEOUT_S` (default 90s) while routing stays up. - **`DRILL_MODE=active` — kills the ACTIVE (oldest) central node.** The survivor must **take over while the victim is still down**: it auto-downs the dead oldest, becomes the oldest member itself, re-hosts all singletons, and its `/health/active` goes 200. PASS = survivor active within `TIMEOUT_S`, then Traefik routing to it. (Under the pre-2026-07-21 `keep-oldest` strategy this direction was a proven total outage — the younger survivor took `DownReachable` and downed itself, because Akka's `down-if-alone` only rescues a side with ≥ 2 members.) Both modes finish by restarting the victim and confirming it rejoins as a ready standby. The drill exercises downing-on-hard-crash, S3 (single active node routed through Traefik), and the Task 20 restart/rejoin contract. Requires a running cluster (`bash docker/deploy.sh`) and `curl` + `docker` on the host. **Partition trade (accepted).** Auto-down is availability-first: in a *real network partition* (both nodes alive, link cut) each side downs the other and both run active — dual-active until an operator restarts one side after the partition heals. This was an explicit owner decision (2026-07-21): site pairs have no shared lease infrastructure to arbitrate, and a stalled system is a bigger risk than a rare partition. See `docs/plans/2026-07-21-auto-down-availability-decision.md`. **Seed-node bootstrap constraint (still applies to boot-alone).** Only the FIRST seed in `Cluster:SeedNodes` may self-join to form a *new* cluster. Both central nodes list `scadabridge-central-a` first (`docker/central-node-a/appsettings.Central.json`, `docker/central-node-b/appsettings.Central.json`), so a lone *restarted* `central-b` (with `central-a` still down) loops on `InitJoin` forever. Under auto-down this no longer causes the active-crash outage (the survivor keeps running — it never restarts), but it still bites when a node must boot alone (cold start of only the non-first-seed VM, or the survivor crashing while its peer is still dead). Operator recovery: **(1)** restart the first-seed node (`central-a`) — preferred; or **(2)** restart the survivor with a self-first seed override (env `ScadaBridge__Cluster__SeedNodes__0=akka.tcp://scadabridge@:8081`, `ScadaBridge__Cluster__SeedNodes__1=`). The repo deliberately does NOT ship self-first ordering per node: with *both* nodes self-first, a simultaneous cold start can let each self-join independently → two one-node clusters that never merge. > **Observed results** (auto-down decision verification): > > **Run 2026-07-21** against a freshly-deployed cluster with `SplitBrainResolverStrategy: auto-down` (first drill: `active=central-a`). Both directions recovered. > > | Direction (`DRILL_MODE`) | Outcome | Measured | > |--------------------------|---------|----------| > | `active` (oldest-node crash) | **PASS — TAKEOVER** — `central-b` auto-downed the dead oldest, went `Younger -> Oldest` on all 7 singletons, and served `/health/active` **while the victim was still down**; restarted victim rejoined as standby. | Survivor active + Traefik routing in **28s** (budget ~25s: 10s detection + 15s auto-down + hand-over); victim ready **2s** after restart. | > | `standby` (younger-node crash) | **PASS** — active node untouched; survivor downed+removed the crashed member; restarted victim rejoined as standby. | Member removed in **27s**; **0** `/health/active` routing blips; victim ready **2s** after restart. | > > Historical baseline (keep-oldest, run 2026-07-13 on `99544985`): `standby` PASS with member removal in 27s / 0 routing blips; `active` was a **total outage** — `central-b` self-downed ~20s after the kill (live SBR log 2026-07-21: `SBR took decision Akka.Cluster.SBR.DownReachable … including myself`) and could not re-bootstrap until `central-a` returned. That result is what motivated the auto-down decision. In-process envelope (`FailoverTimingTests`) measured full failover at **33.7s**. ### Central Failover ```bash # Stop the active central node docker stop scadabridge-central-a # Verify central-b takes over (check logs for leader election) docker logs -f scadabridge-central-b # Access UI on standby node open http://localhost:9002 # Restore the original node docker start scadabridge-central-a ``` ### Site Failover ```bash # Stop the active site-a node docker stop scadabridge-site-a-a # Verify site-a-b takes over singleton (DeploymentManager) docker logs -f scadabridge-site-a-b # Restore docker start scadabridge-site-a-a ``` Same pattern applies for site-b (`scadabridge-site-b-a`/`scadabridge-site-b-b`) and site-c (`scadabridge-site-c-a`/`scadabridge-site-c-b`). Failover takes approximately 25 seconds (2s heartbeat + 10s detection threshold + 15s stable-after for split-brain resolver).