Phase 0 of the ClusterClient→gRPC migration
(docs/plans/2026-07-22-clusterclient-to-grpc-plan.md). Standalone hardening: it
closes a gap that exists today and is a precondition for moving command/control
onto gRPC in later phases.
T0.1 — delete the ManagementActor ClusterClientReceptionist registration.
It was built for an out-of-cluster CLI that was never written: the shipped CLI
speaks HTTP Basic to /management, which asks the actor in-process through
ManagementActorHolder. Nothing in the repo ever sent to /user/management. The
actor still runs there; only the cross-boundary advertisement is gone. Six
documents claimed the CLI used ClusterClient — including the CLI's own README
"Architecture Notes" — and are corrected here rather than left to rot.
T0.2 — record, do not port, the dead integration-routing path.
IntegrationCallRequest is unwired at BOTH ends: RouteIntegrationCallAsync has
zero callers anywhere, and RegisterLocalHandler(Integration, …) appears only in
a test, so production always answers "Integration handler not available". It is
excluded from the gRPC contract (28 of 29 commands migrate) rather than
enshrined on an additive-only wire format, and deleting it during a
transport migration would mix a behavioural change into a change whose whole
value is that behaviour is identical. See
docs/known-issues/2026-07-22-integration-call-routing-is-dead-code.md.
T0.3 — preshared-key authentication on SiteStreamService.
The service shipped with no auth at all: plaintext h2c, no interceptor, so
anything that could reach a site node's :8083 could open a live data stream or
read audit rows back via PullAuditEvents/PullSiteCalls. ControlPlaneAuthInterceptor
now gates /sitestream.SiteStreamService/ — modeled on LocalDbSyncAuthInterceptor
(constant-time compare, fail-closed, PermissionDenied) but gating a SET of
service prefixes so phases 1A/1B add services rather than interceptors. LocalDb
sync keeps its own separate key: it authenticates the pair partner, not central,
and collapsing the two would make a site's central-facing key also admit writes
into its database.
Keys are per site (SB-GRPC-PSK-<siteId>), never fleet-wide, so a compromised
site yields only its own. Central attaches them through ControlPlaneCredentials,
which binds CallCredentials to the channel — covering unary and streaming
uniformly, and letting the key resolve asynchronously, which a client
interceptor could not do without blocking. All three central→site channel
creation sites go through it (SiteStreamGrpcClient and both audit pull invokers);
the pull invokers' channel caches are re-keyed by (site, endpoint) because
credentials are per-site and bound to the channel.
Two decisions beyond the plan:
* StartupValidator now requires GrpcPsk on Site nodes. The plan specified only
the runtime gate, but fail-closed with no boot check produces a node that
joins, answers heartbeats and reports healthy while refusing every stream,
audit pull and telemetry ingest — silent and total. Same reasoning as the
existing inbound API-key pepper rule.
* Added Communication:SitePsks as a central-side key map. The plan assumed
central would read the store, seeded via a dev KEK; the docker rig
deliberately boots with no master key, so store-only resolution would leave
it unable to dial its own sites. The store stays primary — it is the only
source that can serve a site added at runtime — with the map covering
key-less hosts and one-off pins. Neither source falling back to
"unauthenticated" is the invariant.
T0.4 — dev keys on both rigs and tests.
34 tests. The seven that matter most exercise a real in-process gRPC stack over
TestServer: the unit tests on either side of the wire would both stay green if
the halves disagreed, and gRPC refuses call credentials on a plaintext channel
by default — the UnsafeUseInsecureChannelCallCredentials opt-in is only provable
by making a real call. They confirm correct key passes on unary AND streaming,
wrong key and no-credentials both get PermissionDenied, and an unresolvable key
fails the call with nothing reaching the service.
OPERATIONAL: a site node upgraded to this build without a key will not boot.
That includes the gitignored deploy/wonder-app-vd03/ overlay.
11 KiB
ScadaBridge Cluster Topology Guide
Architecture Overview
ScadaBridge uses a hub-and-spoke architecture:
- Central Cluster: Two-node active/standby Akka.NET cluster for management, UI, and coordination.
- Site Clusters: Two-node active/standby Akka.NET clusters at each remote site for data collection and local processing.
%%{init: {'theme':'base', 'themeVariables': {'textColor':'#111111','lineColor':'#555555','edgeLabelBackground':'#ffffff','fontSize':'15px'}}}%%
flowchart TD
USERS["Users<br/>(HTTPS / LB)"]
subgraph CENTRAL["Central Cluster"]
NA["Node A<br/>Active"]
NB["Node B<br/>Standby"]
NA <--> NB
end
USERS --> NA
CENTRAL --> SITE01
CENTRAL --> SITE02
CENTRAL --> SITE03
CENTRAL --> SITEN
subgraph SITE01["Site 01"]
S01A["A<br/>Active"]
S01B["B<br/>Standby"]
end
subgraph SITE02["Site 02"]
S02A["A<br/>Active"]
S02B["B<br/>Standby"]
end
subgraph SITE03["Site 03"]
S03A["A<br/>Active"]
S03B["B<br/>Standby"]
end
subgraph SITEN["Site N"]
SNA["A<br/>Active"]
SNB["B<br/>Standby"]
end
classDef start fill:#d5e8d4,stroke:#82b366,color:#111111;
classDef proc fill:#dae8fc,stroke:#6c8ebf,color:#111111;
classDef dec fill:#fff2cc,stroke:#d6b656,color:#111111;
classDef warn fill:#ffe6cc,stroke:#d79b00,color:#111111;
classDef muted fill:#f5f5f5,stroke:#999999,color:#666666;
class USERS dec
class CENTRAL proc
class NA,S01A,S02A,S03A,SNA start
class NB,S01B,S02B,S03B,SNB muted
class SITE01,SITE02,SITE03,SITEN warn
Central Cluster Setup
Cluster Configuration
Both central nodes must be configured as seed nodes for each other:
Node A (central-01.example.com):
{
"ScadaBridge": {
"Node": {
"Role": "Central",
"NodeHostname": "central-01.example.com",
"RemotingPort": 8081
},
"Cluster": {
"SeedNodes": [
"akka.tcp://scadabridge@central-01.example.com:8081",
"akka.tcp://scadabridge@central-02.example.com:8081"
]
}
}
}
Node B (central-02.example.com):
{
"ScadaBridge": {
"Node": {
"Role": "Central",
"NodeHostname": "central-02.example.com",
"RemotingPort": 8081
},
"Cluster": {
"SeedNodes": [
"akka.tcp://scadabridge@central-02.example.com:8081",
"akka.tcp://scadabridge@central-01.example.com:8081"
]
}
}
}
Seed order is load-bearing — each node lists ITSELF first (decision 2026-07-22). Note Node B's list is the reverse of Node A's. Akka only lets
seed-nodes[0]form a new cluster, so a node listing its partner first can never boot while that partner is down.StartupValidatorrejects the boot if the ordering is wrong, comparing host and port; use the same spelling of the hostname inNodeHostnameand in the seed URI, since Akka does no DNS canonicalisation (central-02andcentral-02.example.comare different seed identities). Seedocs/requirements/Component-ClusterInfrastructure.md→ Seed Node Ordering.
Cluster Behavior
- Split-brain resolver:
auto-down(AutoDowningprovider,auto-down-unreachable-after= 15s) since the 2026-07-21 availability-over-partition-safety decision — the leader among the reachable members downs the unreachable peer, so a hard crash of either node fails over. Accepted trade: a real partition leaves both sides active until an operator restarts one.keep-oldest(withdown-if-alone = on) remains a supportedSplitBrainResolverStrategyvalue, but in a two-node cluster it cannot survive a crash of the oldest node. Seedocs/plans/2026-07-21-auto-down-availability-decision.md. - Minimum members:
min-nr-of-members = 1— a single node can form a cluster. - Failure detection: 2-second heartbeat interval, 10-second threshold.
- Total failover time: ~25 seconds from node failure to singleton migration.
- Singleton handover: Uses CoordinatedShutdown for graceful migration.
Shared State
Both central nodes share state through:
- SQL Server: All configuration, deployment records, templates, and audit logs.
- JWT signing key: Same
JwtSigningKeyin both nodes' configuration. - Data Protection keys: Shared key ring (stored in SQL Server or shared file path).
Load Balancer
A load balancer sits in front of both central nodes for the Blazor Server UI:
- Health check:
GET /health/ready - Protocol: HTTPS (TLS termination at LB or pass-through)
- Sticky sessions: Not required (JWT + shared Data Protection keys)
- If the active node fails, the LB routes to the standby (which becomes active after singleton migration).
Site Cluster Setup
Cluster Configuration
Each site has its own two-node cluster:
Site Node A (site-01-a.example.com):
{
"ScadaBridge": {
"Node": {
"Role": "Site",
"NodeHostname": "site-01-a.example.com",
"SiteId": "plant-north",
"RemotingPort": 8081
},
"Cluster": {
"SeedNodes": [
"akka.tcp://scadabridge@site-01-a.example.com:8081",
"akka.tcp://scadabridge@site-01-b.example.com:8081"
]
}
}
}
Site Node B reverses this list —
site-01-bfirst,site-01-asecond — per the self-first seed rule above. It applies to site pairs exactly as it does to the central pair: without it,site-01-bcannot boot whilesite-01-ais down.
Site Cluster Behavior
- Same split-brain resolver as central (keep-oldest).
- Singleton actors: Site Deployment Manager migrates on failover.
- Staggered instance startup: 50ms delay between Instance Actor creation to prevent reconnection storms.
- SQLite persistence: each node owns its own consolidated LocalDb database, kept in step by asynchronous CDC replication over a gRPC sync stream (LocalDb Phase 1 + 2). The nodes do NOT share a SQLite file.
Site Pair Upgrades — stop and start BOTH nodes together
A rolling upgrade of a site pair, one node at a time, is no longer supported. It worked while
the bespoke replicator kept a legacy SfBufferSnapshot compatibility handler so a new standby
could still apply an old active node's monolithic snapshot. LocalDb Phase 2 deleted that handler
along with the replicator, so a mixed-version pair has no common replication path: the two nodes
will run, but they will not converge, and the divergence is silent.
Stop both nodes of a site pair, upgrade both, then start both.
Related bound — do not leave one node of a pair offline for long. A node absent for longer than
LocalDb:Replication:TombstoneRetention (default 7 days) can resurrect deleted rows when
it rejoins: deletes replicate as HLC-ordered tombstones, and once a tombstone is pruned there is
nothing left to suppress the stale row the returning node still holds. Within the retention window
a rejoin is safe and self-correcting (verified live: a node stopped and restarted mid-load rejoined
with both nodes byte-identical and zero duplicates). Beyond it, rebuild the returning node's
database from its peer rather than letting it rejoin.
Central-Site Communication
Three transports cross the boundary, not one:
- Akka ClusterClient — command/control. Sites list every central node in
ScadaBridge:Communication:CentralContactPoints; contact rotation reaches whichever node answers, so no "active central" needs to be identified. (There is noCommunication:CentralSeedNodesetting — earlier revisions of this guide named one that never existed in the code.) - gRPC — real-time data and audit pull. Note the direction is inverted from the data flow:
each site node hosts the gRPC server on
GrpcPort(default 8083, h2c) and central dials in. - Plain HTTP — the deploy config itself, fetched by the site with a per-deployment token.
gRPC control-plane preshared key (required)
Every site node must set ScadaBridge:Communication:GrpcPsk, and central must hold the same
value for that site. StartupValidator refuses to boot a site node without it, deliberately:
the gate is fail-closed, so an unset key would leave the node joined, healthy-looking and
answering heartbeats while refusing every gRPC call — no live subscriptions, no audit pull, no
cached-telemetry ingest.
| Side | Where the key lives |
|---|---|
| Site node (both nodes of the pair, identical) | ScadaBridge:Communication:GrpcPsk, in production ${secret:SB-GRPC-PSK-<siteId>} |
| Central | secret SB-GRPC-PSK-<siteId> in its store — or ScadaBridge:Communication:SitePsks:<siteId> |
The store is the source that matters in production, because sites are added at runtime and their
keys cannot be enumerated in configuration at boot; SitePsks covers a host running without a
master key (the docker rig) and one-off pins.
One key per site, never one for the fleet: a compromised site must not yield another site's
key. And never share it with LocalDb:Replication:ApiKey — that authenticates the pair partner
for database replication, a different trust relationship on the same listener.
Rotation: set the new value on both sides, then restart the pair (pairs restart together
anyway — see above). Upgrading to a build that has this gate requires seeding the key first,
including in the on-host deploy/ overlays.
The key is a bearer token over plaintext h2c, so it is readable and replayable by anyone on the path. That is the accepted posture today — the same trusted-network assumption the boundary already made, now with authentication rather than none. TLS on these listeners is follow-on hardening and needs no change to the key design.
Scaling Guidelines
Target Scale
- 10 sites maximum per central cluster
- 500 machines (instances) total across all sites
- 75 tags per machine (37,500 total tag subscriptions)
Resource Requirements
| Component | CPU | RAM | Disk | Notes |
|---|---|---|---|---|
| Central node | 4 cores | 8 GB | 50 GB | SQL Server is separate |
| Site node | 2 cores | 4 GB | 20 GB | SQLite databases grow with S&F |
| SQL Server | 4 cores | 16 GB | 100 GB | Shared across central cluster |
Network Bandwidth
- Health reports: ~1 KB per site per 30 seconds = negligible
- Tag value updates: Depends on data change rate; OPC UA subscription-based
- Deployment artifacts: One-time burst per deployment (varies by config size)
- Debug view streaming: ~500 bytes per attribute change per subscriber
Dual-Node Failure Recovery
Scenario: Both Nodes Down
- First node starts: Forms a single-node cluster (
min-nr-of-members = 1). - Central: Reconnects to SQL Server, reads deployment state, becomes operational.
- Site: Opens SQLite databases, rebuilds Instance Actors from persisted configs, resumes S&F retries.
- Second node starts: Joins the existing cluster as standby.
Automatic Recovery
No manual intervention required for dual-node failure. The first node to start will:
- Form the cluster
- Take over all singletons
- Begin processing immediately
- Accept the second node when it joins