7fd5cb2b56
ClusterClient→gRPC migration Phase 4 (docs/plans/2026-07-22-clusterclient-to-grpc-plan.md). Phases 2/3 proved both directions on gRPC; this removes the Akka transport underneath. Deleted: - AkkaCentralTransport, AkkaSiteTransport (+ their dedicated tests) - ISiteClientFactory + DefaultSiteClientFactory; CentralCommunicationActor legacy ctor + SelectTransport (Host now builds GrpcSiteTransport and injects it) - ClusterClient creation + both ClusterClientReceptionist.RegisterService calls in AkkaHostedService; the RegisterCentralClient message + receive block - CommunicationOptions.CentralContactPoints; the CentralTransport/SiteTransport coexistence flags; the CentralTransportMode/SiteTransportKind enums gRPC is now the only site↔central transport (site→central CentralControlService via GrpcCentralTransport; central→site SiteCommandService via GrpcSiteTransport), both built unconditionally by the Host. NoOpCentralTransport is the fail-loud null-default so TestKit command-dispatch suites still construct the site actor without a wired transport; production always injects GrpcCentralTransport. Config: CentralGrpcEndpoints is now unconditional — CommunicationOptionsValidator rejects blank entries (role-agnostic), and StartupValidator requires a Site node to list >=1 endpoint (fail-fast, mirrors GrpcPsk). Rig configs moved CentralContactPoints -> CentralGrpcEndpoints (docker x6, docker-env2 x2, Host default, deploy/wonder-app-vd03). Kept Akka.Cluster.Tools (ClusterSingleton still used). Tests: build 0/0; Communication.Tests 640, Host.Tests 421 green. Removed the ClusterClient.Send per-site-routing tests (covered by the transport suites), swapped the ISiteClientFactory-based ctors to a substitute ISiteCommandTransport, converted the audit-push integration relay to an in-process bridge transport. Docs: Component-Communication/Host/StoreAndForward, components/Communication, topology-guide, grpc_streams (SUPERSEDED note), the frame-size known-issue (retired amendment), and CLAUDE.md transport decisions. Not included: the dead IntegrationCallRequest path (#32) is a separate user-owned behavioral decision — SiteEnvelope routing is transport-agnostic so it still compiles.
268 lines
12 KiB
Markdown
268 lines
12 KiB
Markdown
# ScadaBridge Cluster Topology Guide
|
|
|
|
## Architecture Overview
|
|
|
|
ScadaBridge uses a hub-and-spoke architecture:
|
|
- **Central Cluster**: Two-node active/standby Akka.NET cluster for management, UI, and coordination.
|
|
- **Site Clusters**: Two-node active/standby Akka.NET clusters at each remote site for data collection and local processing.
|
|
|
|
```mermaid
|
|
%%{init: {'theme':'base', 'themeVariables': {'textColor':'#111111','lineColor':'#555555','edgeLabelBackground':'#ffffff','fontSize':'15px'}}}%%
|
|
flowchart TD
|
|
USERS["Users<br/>(HTTPS / LB)"]
|
|
|
|
subgraph CENTRAL["Central Cluster"]
|
|
NA["Node A<br/>Active"]
|
|
NB["Node B<br/>Standby"]
|
|
NA <--> NB
|
|
end
|
|
|
|
USERS --> NA
|
|
CENTRAL --> SITE01
|
|
CENTRAL --> SITE02
|
|
CENTRAL --> SITE03
|
|
CENTRAL --> SITEN
|
|
|
|
subgraph SITE01["Site 01"]
|
|
S01A["A<br/>Active"]
|
|
S01B["B<br/>Standby"]
|
|
end
|
|
subgraph SITE02["Site 02"]
|
|
S02A["A<br/>Active"]
|
|
S02B["B<br/>Standby"]
|
|
end
|
|
subgraph SITE03["Site 03"]
|
|
S03A["A<br/>Active"]
|
|
S03B["B<br/>Standby"]
|
|
end
|
|
subgraph SITEN["Site N"]
|
|
SNA["A<br/>Active"]
|
|
SNB["B<br/>Standby"]
|
|
end
|
|
|
|
classDef start fill:#d5e8d4,stroke:#82b366,color:#111111;
|
|
classDef proc fill:#dae8fc,stroke:#6c8ebf,color:#111111;
|
|
classDef dec fill:#fff2cc,stroke:#d6b656,color:#111111;
|
|
classDef warn fill:#ffe6cc,stroke:#d79b00,color:#111111;
|
|
classDef muted fill:#f5f5f5,stroke:#999999,color:#666666;
|
|
class USERS dec
|
|
class CENTRAL proc
|
|
class NA,S01A,S02A,S03A,SNA start
|
|
class NB,S01B,S02B,S03B,SNB muted
|
|
class SITE01,SITE02,SITE03,SITEN warn
|
|
```
|
|
|
|
## Central Cluster Setup
|
|
|
|
### Cluster Configuration
|
|
|
|
Both central nodes must be configured as seed nodes for each other:
|
|
|
|
**Node A** (`central-01.example.com`):
|
|
```json
|
|
{
|
|
"ScadaBridge": {
|
|
"Node": {
|
|
"Role": "Central",
|
|
"NodeHostname": "central-01.example.com",
|
|
"RemotingPort": 8081
|
|
},
|
|
"Cluster": {
|
|
"SeedNodes": [
|
|
"akka.tcp://scadabridge@central-01.example.com:8081",
|
|
"akka.tcp://scadabridge@central-02.example.com:8081"
|
|
]
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
**Node B** (`central-02.example.com`):
|
|
```json
|
|
{
|
|
"ScadaBridge": {
|
|
"Node": {
|
|
"Role": "Central",
|
|
"NodeHostname": "central-02.example.com",
|
|
"RemotingPort": 8081
|
|
},
|
|
"Cluster": {
|
|
"SeedNodes": [
|
|
"akka.tcp://scadabridge@central-02.example.com:8081",
|
|
"akka.tcp://scadabridge@central-01.example.com:8081"
|
|
]
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
> **Seed order is load-bearing — each node lists ITSELF first** (decision 2026-07-22). Note Node B's list is the reverse of Node A's. Akka only lets `seed-nodes[0]` form a *new* cluster, so a node listing its partner first can never boot while that partner is down. `StartupValidator` rejects the boot if the ordering is wrong, comparing host **and** port; use the same spelling of the hostname in `NodeHostname` and in the seed URI, since Akka does no DNS canonicalisation (`central-02` and `central-02.example.com` are different seed identities). See `docs/requirements/Component-ClusterInfrastructure.md` → Seed Node Ordering.
|
|
|
|
### Cluster Behavior
|
|
|
|
- **Split-brain resolver**: `auto-down` (`AutoDowning` provider, `auto-down-unreachable-after` = 15s) since the 2026-07-21 availability-over-partition-safety decision — the leader among the *reachable* members downs the unreachable peer, so a hard crash of **either** node fails over. Accepted trade: a real partition leaves both sides active until an operator restarts one. `keep-oldest` (with `down-if-alone = on`) remains a supported `SplitBrainResolverStrategy` value, but in a two-node cluster it cannot survive a crash of the oldest node. See `docs/plans/2026-07-21-auto-down-availability-decision.md`.
|
|
- **Minimum members**: `min-nr-of-members = 1` — a single node can form a cluster.
|
|
- **Failure detection**: 2-second heartbeat interval, 10-second threshold.
|
|
- **Total failover time**: ~25 seconds from node failure to singleton migration.
|
|
- **Singleton handover**: Uses CoordinatedShutdown for graceful migration.
|
|
|
|
### Shared State
|
|
|
|
Both central nodes share state through:
|
|
- **SQL Server**: All configuration, deployment records, templates, and audit logs.
|
|
- **JWT signing key**: Same `JwtSigningKey` in both nodes' configuration.
|
|
- **Data Protection keys**: Shared key ring (stored in SQL Server or shared file path).
|
|
|
|
### Load Balancer
|
|
|
|
A load balancer sits in front of both central nodes for the Blazor Server UI:
|
|
- Health check: `GET /health/ready`
|
|
- Protocol: HTTPS (TLS termination at LB or pass-through)
|
|
- Sticky sessions: Not required (JWT + shared Data Protection keys)
|
|
- If the active node fails, the LB routes to the standby (which becomes active after singleton migration).
|
|
|
|
## Site Cluster Setup
|
|
|
|
### Cluster Configuration
|
|
|
|
Each site has its own two-node cluster:
|
|
|
|
**Site Node A** (`site-01-a.example.com`):
|
|
```json
|
|
{
|
|
"ScadaBridge": {
|
|
"Node": {
|
|
"Role": "Site",
|
|
"NodeHostname": "site-01-a.example.com",
|
|
"SiteId": "plant-north",
|
|
"RemotingPort": 8081
|
|
},
|
|
"Cluster": {
|
|
"SeedNodes": [
|
|
"akka.tcp://scadabridge@site-01-a.example.com:8081",
|
|
"akka.tcp://scadabridge@site-01-b.example.com:8081"
|
|
]
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
> **Site Node B reverses this list** — `site-01-b` first, `site-01-a` second — per the self-first seed rule above. It applies to site pairs exactly as it does to the central pair: without it, `site-01-b` cannot boot while `site-01-a` is down.
|
|
|
|
### Site Cluster Behavior
|
|
|
|
- Same split-brain resolver as central (keep-oldest).
|
|
- Singleton actors: Site Deployment Manager migrates on failover.
|
|
- Staggered instance startup: 50ms delay between Instance Actor creation to prevent reconnection storms.
|
|
- SQLite persistence: each node owns its own consolidated LocalDb database, kept in step by
|
|
asynchronous CDC replication over a gRPC sync stream (LocalDb Phase 1 + 2). The nodes do NOT
|
|
share a SQLite file.
|
|
|
|
### Site Pair Upgrades — stop and start BOTH nodes together
|
|
|
|
**A rolling upgrade of a site pair, one node at a time, is no longer supported.** It worked while
|
|
the bespoke replicator kept a legacy `SfBufferSnapshot` compatibility handler so a new standby
|
|
could still apply an old active node's monolithic snapshot. LocalDb Phase 2 deleted that handler
|
|
along with the replicator, so a mixed-version pair has no common replication path: the two nodes
|
|
will run, but they will not converge, and the divergence is silent.
|
|
|
|
Stop both nodes of a site pair, upgrade both, then start both.
|
|
|
|
**Related bound — do not leave one node of a pair offline for long.** A node absent for longer than
|
|
`LocalDb:Replication:TombstoneRetention` (default **7 days**) can **resurrect deleted rows** when
|
|
it rejoins: deletes replicate as HLC-ordered tombstones, and once a tombstone is pruned there is
|
|
nothing left to suppress the stale row the returning node still holds. Within the retention window
|
|
a rejoin is safe and self-correcting (verified live: a node stopped and restarted mid-load rejoined
|
|
with both nodes byte-identical and zero duplicates). Beyond it, rebuild the returning node's
|
|
database from its peer rather than letting it rejoin.
|
|
|
|
### Central-Site Communication
|
|
|
|
Three transports cross the boundary, not one — **all now gRPC or HTTP; Akka ClusterClient was removed
|
|
in Phase 4 of the ClusterClient→gRPC migration (2026-07-23), and Akka remoting no longer crosses the
|
|
boundary at all:**
|
|
|
|
- **gRPC command/control** — both directions, on sticky-failover channel pairs, dialled directly (no
|
|
receptionist, no "active central" to identify — each side dials both of the peer's node endpoints):
|
|
- *Site → central* to the central-hosted **`CentralControlService`** (`GrpcCentralTransport`): the
|
|
site lists the central nodes' gRPC endpoints in `ScadaBridge:Communication:CentralGrpcEndpoints`
|
|
(e.g. `http://scadabridge-central-a:8083`, the central's `CentralGrpcPort`, default 8083 — direct
|
|
h2c, **not** via Traefik, which is HTTP/1 only). A Site node must list at least one; central nodes
|
|
leave it empty.
|
|
- *Central → site* to the site-hosted **`SiteCommandService`** (`GrpcSiteTransport`): central dials
|
|
the site's `GrpcNodeAAddress` / `GrpcNodeBAddress` (from the Site entity), NodeA→NodeB failover.
|
|
- **gRPC streaming + audit pull** — real-time data and audit/telemetry pull on the site-hosted
|
|
**`SiteStreamService`**. Note the direction is inverted from the data flow: each **site node hosts
|
|
the server** on `GrpcPort` (default 8083, h2c) and central dials in.
|
|
- **Plain HTTP** — the deploy config itself, fetched by the site with a per-deployment token.
|
|
|
|
#### gRPC control-plane preshared key (required)
|
|
|
|
Every site node must set `ScadaBridge:Communication:GrpcPsk`, and central must hold the same
|
|
value for that site. **`StartupValidator` refuses to boot a site node without it**, deliberately:
|
|
the gate is fail-closed, so an unset key would leave the node joined, healthy-looking and
|
|
answering heartbeats while refusing every gRPC call — no live subscriptions, no audit pull, no
|
|
cached-telemetry ingest.
|
|
|
|
| Side | Where the key lives |
|
|
|---|---|
|
|
| Site node (both nodes of the pair, identical) | `ScadaBridge:Communication:GrpcPsk`, in production `${secret:SB-GRPC-PSK-<siteId>}` |
|
|
| Central | secret `SB-GRPC-PSK-<siteId>` in its store — **or** `ScadaBridge:Communication:SitePsks:<siteId>` |
|
|
|
|
The store is the source that matters in production, because sites are added at runtime and their
|
|
keys cannot be enumerated in configuration at boot; `SitePsks` covers a host running without a
|
|
master key (the docker rig) and one-off pins.
|
|
|
|
One key **per site**, never one for the fleet: a compromised site must not yield another site's
|
|
key. And never share it with `LocalDb:Replication:ApiKey` — that authenticates the *pair partner*
|
|
for database replication, a different trust relationship on the same listener.
|
|
|
|
**Rotation:** set the new value on both sides, then restart the pair (pairs restart together
|
|
anyway — see above). **Upgrading to a build that has this gate requires seeding the key first**,
|
|
including in the on-host `deploy/` overlays.
|
|
|
|
The key is a bearer token over plaintext h2c, so it is readable and replayable by anyone on the
|
|
path. That is the accepted posture today — the same trusted-network assumption the boundary
|
|
already made, now with authentication rather than none. TLS on these listeners is follow-on
|
|
hardening and needs no change to the key design.
|
|
|
|
## Scaling Guidelines
|
|
|
|
### Target Scale
|
|
|
|
- 10 sites maximum per central cluster
|
|
- 500 machines (instances) total across all sites
|
|
- 75 tags per machine (37,500 total tag subscriptions)
|
|
|
|
### Resource Requirements
|
|
|
|
| Component | CPU | RAM | Disk | Notes |
|
|
|-----------|-----|-----|------|-------|
|
|
| Central node | 4 cores | 8 GB | 50 GB | SQL Server is separate |
|
|
| Site node | 2 cores | 4 GB | 20 GB | SQLite databases grow with S&F |
|
|
| SQL Server | 4 cores | 16 GB | 100 GB | Shared across central cluster |
|
|
|
|
### Network Bandwidth
|
|
|
|
- Health reports: ~1 KB per site per 30 seconds = negligible
|
|
- Tag value updates: Depends on data change rate; OPC UA subscription-based
|
|
- Deployment artifacts: One-time burst per deployment (varies by config size)
|
|
- Debug view streaming: ~500 bytes per attribute change per subscriber
|
|
|
|
## Dual-Node Failure Recovery
|
|
|
|
### Scenario: Both Nodes Down
|
|
|
|
1. **First node starts**: Forms a single-node cluster (`min-nr-of-members = 1`).
|
|
2. **Central**: Reconnects to SQL Server, reads deployment state, becomes operational.
|
|
3. **Site**: Opens SQLite databases, rebuilds Instance Actors from persisted configs, resumes S&F retries.
|
|
4. **Second node starts**: Joins the existing cluster as standby.
|
|
|
|
### Automatic Recovery
|
|
|
|
No manual intervention required for dual-node failure. The first node to start will:
|
|
- Form the cluster
|
|
- Take over all singletons
|
|
- Begin processing immediately
|
|
- Accept the second node when it joins
|