Files
ScadaBridge/docs/deployment/topology-guide.md
T
Joseph Doherty 9d2834e30a docs+log(siteeventlogging): explain the site_events purge oplog-backlog burst (R7)
The daily site_events retention purge (and the storage-cap trim) is CDC-captured
on a replication-enabled site node exactly like any other write — correct by
design, since LocalDb Phase 2 deliberately has no purge-exemption path — so the
backlog jumps by the deleted batch size at purge time. LocalDbOplogBacklog /
localdb_oplog_depth spike, drain, and an operator watching the gauge with no
context reads it as a replication fault.

Documentation + one log line, no behaviour change:

- topology-guide.md gains "Reading the replication backlog — the daily
  site_events purge burst": when it fires (PurgeInterval 24h, anchored to the
  active node's PROCESS START, not a wall-clock hour, so it moves after every
  failover), where it shows (replicated nodes only — not rig site-b/site-c),
  the healthy signature (LocalDbReplicationConnected stays true, backlog
  returns to ~0) and what a genuine fault looks like instead.
- Component-SiteEventLogging.md Storage records the same under retention/purge;
  Component-HealthMonitoring.md gains the two previously-undocumented
  LocalDbReplicationConnected / LocalDbOplogBacklog metric rows carrying the
  caveat, with cross-references both ways.
- EventLogPurgeService emits one Information line naming the row count and the
  expected transient backlog when a purge deleted rows on a replication-enabled
  node, so the spike is correlatable in the log. Replication-awareness comes in
  as a Host-supplied SiteEventLogReplicationCheck delegate, mirroring the
  existing SiteEventLogActiveNodeCheck seam: SiteLocalDbSetup.ReplicationIsConfigured
  goes internal so the PeerAddress-OR-ApiKey rule stays in one place and
  SiteEventLogging never learns to read LocalDb config. Unregistered ⇒ no note,
  matching the default that replication is opt-in and off.

Both delete paths carry the note (a cap trim is usually the larger burst); the
predicate is try/caught since a log-wording check must never break the purge.

Tests: 5 new EventLogPurgeServiceTests cases (replicated logs it, unreplicated
does not, zero-rows does not, cap purge logs it, throwing predicate still purges
and swallows) via a local capturing ILogger. SiteEventLogging 81/81 green,
Host 490/490 green, full solution build clean (0 warnings).
2026-08-15 03:26:30 -04:00

18 KiB

ScadaBridge Cluster Topology Guide

Architecture Overview

ScadaBridge uses a hub-and-spoke architecture:

  • Central Cluster: Two-node active/standby Akka.NET cluster for management, UI, and coordination.
  • Site Clusters: Two-node active/standby Akka.NET clusters at each remote site for data collection and local processing.
%%{init: {'theme':'base', 'themeVariables': {'textColor':'#111111','lineColor':'#555555','edgeLabelBackground':'#ffffff','fontSize':'15px'}}}%%
flowchart TD
    USERS["Users<br/>(HTTPS / LB)"]

    subgraph CENTRAL["Central Cluster"]
        NA["Node A<br/>Active"]
        NB["Node B<br/>Standby"]
        NA <--> NB
    end

    USERS --> NA
    CENTRAL --> SITE01
    CENTRAL --> SITE02
    CENTRAL --> SITE03
    CENTRAL --> SITEN

    subgraph SITE01["Site 01"]
        S01A["A<br/>Active"]
        S01B["B<br/>Standby"]
    end
    subgraph SITE02["Site 02"]
        S02A["A<br/>Active"]
        S02B["B<br/>Standby"]
    end
    subgraph SITE03["Site 03"]
        S03A["A<br/>Active"]
        S03B["B<br/>Standby"]
    end
    subgraph SITEN["Site N"]
        SNA["A<br/>Active"]
        SNB["B<br/>Standby"]
    end

    classDef start fill:#d5e8d4,stroke:#82b366,color:#111111;
    classDef proc fill:#dae8fc,stroke:#6c8ebf,color:#111111;
    classDef dec fill:#fff2cc,stroke:#d6b656,color:#111111;
    classDef warn fill:#ffe6cc,stroke:#d79b00,color:#111111;
    classDef muted fill:#f5f5f5,stroke:#999999,color:#666666;
    class USERS dec
    class CENTRAL proc
    class NA,S01A,S02A,S03A,SNA start
    class NB,S01B,S02B,S03B,SNB muted
    class SITE01,SITE02,SITE03,SITEN warn

Central Cluster Setup

Cluster Configuration

Both central nodes must be configured as seed nodes for each other:

Node A (central-01.example.com):

{
  "ScadaBridge": {
    "Node": {
      "Role": "Central",
      "NodeHostname": "central-01.example.com",
      "RemotingPort": 8081
    },
    "Cluster": {
      "SeedNodes": [
        "akka.tcp://scadabridge@central-01.example.com:8081",
        "akka.tcp://scadabridge@central-02.example.com:8081"
      ]
    }
  }
}

Node B (central-02.example.com):

{
  "ScadaBridge": {
    "Node": {
      "Role": "Central",
      "NodeHostname": "central-02.example.com",
      "RemotingPort": 8081
    },
    "Cluster": {
      "SeedNodes": [
        "akka.tcp://scadabridge@central-02.example.com:8081",
        "akka.tcp://scadabridge@central-01.example.com:8081"
      ]
    }
  }
}

Seed order is load-bearing — each node lists ITSELF first (decision 2026-07-22). Note Node B's list is the reverse of Node A's. Akka only lets seed-nodes[0] form a new cluster, so a node listing its partner first can never boot while that partner is down. StartupValidator rejects the boot if the ordering is wrong, comparing host and port; use the same spelling of the hostname in NodeHostname and in the seed URI, since Akka does no DNS canonicalisation (central-02 and central-02.example.com are different seed identities). See docs/requirements/Component-ClusterInfrastructure.md → Seed Node Ordering.

Cluster Behavior

  • Split-brain resolver: auto-down (AutoDowning provider, auto-down-unreachable-after = 15s) since the 2026-07-21 availability-over-partition-safety decision — the leader among the reachable members downs the unreachable peer, so a hard crash of either node fails over. Accepted trade: a real partition leaves both sides active until an operator restarts one. keep-oldest (with down-if-alone = on) remains a supported SplitBrainResolverStrategy value, but in a two-node cluster it cannot survive a crash of the oldest node. See docs/plans/2026-07-21-auto-down-availability-decision.md.
  • Minimum members: min-nr-of-members = 1 — a single node can form a cluster.
  • Failure detection: 2-second heartbeat interval, 10-second threshold.
  • Total failover time: ~25 seconds from node failure to singleton migration.
  • Singleton handover: Uses CoordinatedShutdown for graceful migration.

Shared State

Both central nodes share state through:

  • SQL Server: All configuration, deployment records, templates, and audit logs.
  • JWT signing key: Same JwtSigningKey in both nodes' configuration.
  • Data Protection keys: Shared key ring (stored in SQL Server or shared file path).

Load Balancer

A load balancer sits in front of both central nodes for the Blazor Server UI:

  • Health check: GET /health/ready
  • Protocol: HTTPS (TLS termination at LB or pass-through)
  • Sticky sessions: Not required (JWT + shared Data Protection keys)
  • If the active node fails, the LB routes to the standby (which becomes active after singleton migration).

Site Cluster Setup

Cluster Configuration

Each site has its own two-node cluster:

Site Node A (site-01-a.example.com):

{
  "ScadaBridge": {
    "Node": {
      "Role": "Site",
      "NodeHostname": "site-01-a.example.com",
      "SiteId": "plant-north",
      "RemotingPort": 8081
    },
    "Cluster": {
      "SeedNodes": [
        "akka.tcp://scadabridge@site-01-a.example.com:8081",
        "akka.tcp://scadabridge@site-01-b.example.com:8081"
      ]
    }
  }
}

Site Node B reverses this listsite-01-b first, site-01-a second — per the self-first seed rule above. It applies to site pairs exactly as it does to the central pair: without it, site-01-b cannot boot while site-01-a is down.

Site Cluster Behavior

  • Same split-brain resolver as central (auto-down, per the 2026-07-21 decision — see the Central Cluster Behavior note above).
  • Singleton actors: Site Deployment Manager migrates on failover.
  • Staggered instance startup: 50ms delay between Instance Actor creation to prevent reconnection storms.
  • SQLite persistence: each node owns its own consolidated LocalDb database, kept in step by asynchronous CDC replication over a gRPC sync stream (LocalDb Phase 1 + 2). The nodes do NOT share a SQLite file.
  • CDC capture triggers are installed only on a node that has replication configuredLocalDb:Replication:PeerAddress or LocalDb:Replication:ApiKey. Either key counts, because only the initiating half of a pair sets PeerAddress (one bidirectional stream, dialled by one side); the passive half carries the key alone. A deliberately unreplicated node — site-b and site-c on the rig — runs with no triggers at all and stops paying the per-write capture cost.
  • Stale-trigger cleanup is automatic (LocalDb 0.2.0). A node with no replication configured does not merely skip registration — at boot it calls DeregisterReplicated on all ten tables, dropping any capture triggers an earlier build installed and pruning those tables' oplog and row-version rows. It is idempotent, so a file that was never registered reports nothing to clean; when something was cleaned the node logs it once at Information. Recreating the data volume is no longer required to stop an in-place-upgraded node from capturing.

Turning replication ON for a site that has been running without it

Supported as of LocalDb 0.2.0. Set the keys on both nodes and restart both (see the stop-and-start-together rule below — this is a pair-wide change, not a rolling one). Existing rows are carried across:

  • Pre-existing rows are baselined automatically. ScadaBridge registers every replicated table with baselineExistingRows: true. Capture is change-data-capture, so rows written while the node had no triggers appear in neither the oplog nor __localdb_row_version — and LocalDb's snapshot resync streams from that ledger. Baselining seeds the ledger for those rows at the LWW floor (HLC 0, stamped with the node's own id) and flags a snapshot resync, so the peer actually receives them. Copying one node's database onto the other beforehand is no longer necessary.
  • What the floor means for conflicts. Every genuine HLC is a UTC millisecond shifted left 16 bits, so it is strictly greater than 0: a baselined row loses to any real remote write of the same key and wins only where the peer holds no version of that key at all. The one ambiguous case is both nodes baselining the same key (e.g. both were restored from the same legacy file) — both hold HLC 0 and the node-id tie-break decides. That is convergent but arbitrary as to which content survives, so if the two files may disagree on a key, start both nodes from one node's database.
  • Seeding is idempotent (ON CONFLICT DO NOTHING), so a row that already has a genuine version keeps it and no snapshot is flagged. Booting with baselining on every start is free after the first.

Turning replication back OFF is likewise a both-nodes change: deregistration must be symmetric, because the sync handshake compares the two nodes' registered-table digests fail-closed — a node that drops a table its peer still replicates stops syncing with a schema-mismatch error rather than diverging silently. Turning it on again later re-baselines, which is what makes the ledger prune on deregistration safe.

Reading the replication backlog — the daily site_events purge burst

The site health report carries LocalDbReplicationConnected and LocalDbOplogBacklog (nullable — null means "no reading", not "disconnected with an empty backlog"), and the same numbers export as the localdb_* Prometheus series (localdb_oplog_depth is the backlog gauge). A healthy pair sits at a backlog of roughly zero, so a sudden spike naturally reads as a replication problem.

One expected spike is not a problem: the daily site_events retention purge. site_events is one of the ten replicated tables, and its retention DELETE is captured by CDC exactly like an ordinary write — by design, since LocalDb Phase 2 there is deliberately no purge-exemption path (the same property that makes a mass DELETE dangerous, which is why ReplaceAllAsync was deleted rather than reinstated). One oplog row is therefore queued per deleted event, and the backlog jumps by the size of the day's expired batch.

  • When. Every ScadaBridge:SiteEventLog:PurgeInterval (default 24 h), plus once at startup. The timer is anchored to the active node's process start, not to a wall-clock hour, so the burst lands at a different time of day after each failover or restart — do not expect it at a fixed hour. The storage-cap trim (default 1 GB) can produce the same shape off-schedule, and is usually the larger of the two.
  • Where it shows. Only on a node with replication configured — the rig's site-a. site-b/site-c have no capture triggers at all and report no backlog for a purge.
  • What healthy looks like. LocalDbReplicationConnected stays true across the spike, and the backlog drains back to ~0 as the peer acks the batch — within seconds to a couple of minutes depending on batch size (delta messages are bounded by LocalDb:Replication:MaxBatchBytes, default 2 MB, and secondarily by MaxBatchSize). No dead letters, no schema-mismatch errors.
  • What is actually wrong. LocalDbReplicationConnected false while the backlog climbs, a backlog that keeps rising across successive readings rather than draining, or a backlog that never returns near zero between bursts. Those point at the sync stream — an ApiKey mismatch (fail-closed: the pair simply stops converging), an unreachable peer, or an asymmetric registered-table set.

Correlating it in the log. When a purge on a replication-enabled node actually deletes rows it logs an Information line next to the purge count:

Purged 41230 events older than 30 days
Purged 41230 site_events rows on a replication-enabled node — a transient LocalDb oplog backlog
is expected while the deletes replicate to the peer. It drains on its own;
LocalDbReplicationConnected staying true with the backlog returning to ~0 is the healthy signature.

An unreplicated node logs only the first line. If a backlog spike has no such line near it in the active node's log, the purge is not the explanation and the spike is worth investigating.

Site Pair Upgrades — stop and start BOTH nodes together

A rolling upgrade of a site pair, one node at a time, is no longer supported. It worked while the bespoke replicator kept a legacy SfBufferSnapshot compatibility handler so a new standby could still apply an old active node's monolithic snapshot. LocalDb Phase 2 deleted that handler along with the replicator, so a mixed-version pair has no common replication path: the two nodes will run, but they will not converge, and the divergence is silent.

Stop both nodes of a site pair, upgrade both, then start both.

Related bound — do not leave one node of a pair offline for long. A node absent for longer than LocalDb:Replication:TombstoneRetention (default 7 days) can resurrect deleted rows when it rejoins: deletes replicate as HLC-ordered tombstones, and once a tombstone is pruned there is nothing left to suppress the stale row the returning node still holds. Within the retention window a rejoin is safe and self-correcting (verified live: a node stopped and restarted mid-load rejoined with both nodes byte-identical and zero duplicates). Beyond it, rebuild the returning node's database from its peer rather than letting it rejoin.

Central-Site Communication

Three transports cross the boundary, not one — all now gRPC or HTTP; Akka ClusterClient was removed in Phase 4 of the ClusterClient→gRPC migration (2026-07-23), and Akka remoting no longer crosses the boundary at all:

  • gRPC command/control — both directions, on sticky-failover channel pairs, dialled directly (no receptionist, no "active central" to identify — each side dials both of the peer's node endpoints):
    • Site → central to the central-hosted CentralControlService (GrpcCentralTransport): the site lists the central nodes' gRPC endpoints in ScadaBridge:Communication:CentralGrpcEndpoints (e.g. http://scadabridge-central-a:8083, the central's CentralGrpcPort, default 8083 — direct h2c, not via Traefik, which is HTTP/1 only). A Site node must list at least one; central nodes leave it empty.
    • Central → site to the site-hosted SiteCommandService (GrpcSiteTransport): central dials the site's GrpcNodeAAddress / GrpcNodeBAddress (from the Site entity), NodeA→NodeB failover.
  • gRPC streaming + audit pull — real-time data and audit/telemetry pull on the site-hosted SiteStreamService. Note the direction is inverted from the data flow: each site node hosts the server on GrpcPort (default 8083, h2c) and central dials in.
  • Plain HTTP — the deploy config itself, fetched by the site with a per-deployment token.

gRPC control-plane preshared key (required)

Every site node must set ScadaBridge:Communication:GrpcPsk, and central must hold the same value for that site. StartupValidator refuses to boot a site node without it, deliberately: the gate is fail-closed, so an unset key would leave the node joined, healthy-looking and answering heartbeats while refusing every gRPC call — no live subscriptions, no audit pull, no cached-telemetry ingest.

Side Where the key lives
Site node (both nodes of the pair, identical) ScadaBridge:Communication:GrpcPsk, in production ${secret:SB-GRPC-PSK-<siteId>}
Central secret SB-GRPC-PSK-<siteId> in its store — or ScadaBridge:Communication:SitePsks:<siteId>

The store is the source that matters in production, because sites are added at runtime and their keys cannot be enumerated in configuration at boot; SitePsks covers a host running without a master key (the docker rig) and one-off pins.

One key per site, never one for the fleet: a compromised site must not yield another site's key. And never share it with LocalDb:Replication:ApiKey — that authenticates the pair partner for database replication, a different trust relationship on the same listener.

Rotation: set the new value on both sides, then restart the pair (pairs restart together anyway — see above). Upgrading to a build that has this gate requires seeding the key first, including in the on-host deploy/ overlays.

The key is a bearer token over plaintext h2c, so it is readable and replayable by anyone on the path. That is the accepted posture today — the same trusted-network assumption the boundary already made, now with authentication rather than none. TLS on these listeners is follow-on hardening and needs no change to the key design.

Scaling Guidelines

Target Scale

  • 10 sites maximum per central cluster
  • 500 machines (instances) total across all sites
  • 75 tags per machine (37,500 total tag subscriptions)

Resource Requirements

Component CPU RAM Disk Notes
Central node 4 cores 8 GB 50 GB SQL Server is separate
Site node 2 cores 4 GB 20 GB SQLite databases grow with S&F
SQL Server 4 cores 16 GB 100 GB Shared across central cluster

Network Bandwidth

  • Health reports: ~1 KB per site per 30 seconds = negligible
  • Tag value updates: Depends on data change rate; OPC UA subscription-based
  • Deployment artifacts: One-time burst per deployment (varies by config size)
  • Debug view streaming: ~500 bytes per attribute change per subscriber

Dual-Node Failure Recovery

Scenario: Both Nodes Down

  1. First node starts: Forms a single-node cluster (min-nr-of-members = 1).
  2. Central: Reconnects to SQL Server, reads deployment state, becomes operational.
  3. Site: Opens SQLite databases, rebuilds Instance Actors from persisted configs, resumes S&F retries.
  4. Second node starts: Joins the existing cluster as standby.

Automatic Recovery

No manual intervention required for dual-node failure. The first node to start will:

  • Form the cluster
  • Take over all singletons
  • Begin processing immediately
  • Accept the second node when it joins