Rebuilt scadabridge:latest from main @ 8524a7f7 and recreated only the
env2 containers. All three gate checks pass on site-x:
1. both site nodes boot with the key (StartupValidator fail-closed →
reaching 'Application started' proves the key present);
2. control-plane PSK auth: no-header / wrong-key ⇒ PermissionDenied,
correct key ⇒ success, on both nodes (:9123, :9124);
3. LocalDb unaffected (local-only; 0 errors, healthy boot).
Bonus: central registers site-x online via gRPC heartbeat; no real
ClusterClient/receptionist (only the benign ClusterClientSiteAuditClient
label, same as the primary rig). Noted a seed-data gap (ScadaBridgeConfig2
dbo.Sites is empty) — orthogonal to the transport.
35 KiB
ClusterClient → gRPC migration — live gate results
Rig: docker/ (2 central + 3×2 site + traefik), rebuilt from the branch under test via
bash docker/deploy.sh. Recorded check-by-check in the family's live-gate format. Phase 5's
full eight-check gate is recorded further down as those phases land; this file starts with
Phase 0, whose DoD has its own smaller gate.
Phase 0 — PSK auth + dead-code removal — PASS (2026-07-22)
Branch feat/grpc-phase0-psk @ 228ff8b4. Image rebuilt, all 9 containers recreated.
Baseline (pre-change build, same rig)
An unauthenticated call from the host to a site's audit-pull RPC was accepted:
$ grpcurl -plaintext -d '{"batch_size":1}' localhost:9023 sitestream.SiteStreamService/PullAuditEvents
{}
That is the gap Phase 0 closes, reproduced rather than assumed.
Checks
| # | Check | Result |
|---|---|---|
| 1 | All 8 nodes boot with keys configured (new StartupValidator rule) |
PASS — all recreated and reached ready |
| 2 | Unauthenticated PullAuditEvents ⇒ PermissionDenied, all 3 sites |
PASS |
| 3 | Wrong key (site-b's key presented to site-a) ⇒ PermissionDenied |
PASS — per-site scoping is real, not decorative |
| 4 | Correct key ⇒ success, all 3 sites | PASS |
| 5 | Central's own authenticated paths still work | PASS — 14 successful PullAuditEvents from central to site-a; 0 auth failures in either central's log |
| 6 | LocalDb sync unaffected by the new interceptor | PASS — 0 control-plane rejections and 0 sync auth failures on the passive peer; session connected after one boot-order retry |
| 7 | No interceptor-activation errors | PASS — 0 (see the defect below) |
Evidence for 2–4:
=== NO CREDENTIALS ===
:9023 -> ERROR: Code: PermissionDenied Message: Control plane authentication failed.
:9033 -> ERROR: Code: PermissionDenied Message: Control plane authentication failed.
:9043 -> ERROR: Code: PermissionDenied Message: Control plane authentication failed.
=== WRONG KEY (site-b's key against site-a) ===
:9023 -> ERROR: Code: PermissionDenied Message: Control plane authentication failed.
=== CORRECT KEY ===
site-a :9023 -> {}
site-b :9033 -> {}
site-c :9043 -> {}
Site-a rejected exactly 2 calls — the two deliberate probes above — and nothing else.
Defect the gate caught that the test suite did not
First run of this gate FAILED, and is worth recording because the failure mode is deceptive.
Grpc.AspNetCore activates a type-registered interceptor through
InterceptorRegistration.GetFactory(), which throws when more than one public constructor is
applicable. ControlPlaneAuthInterceptor shipped with two — the DI one and a prefix-set
overload intended for later phases.
The throw happens inside the pipeline, per call, so:
- nothing failed at startup; the node booted, joined its pair and reported healthy;
- every gated call died with
Unknown / "Exception was thrown by handler", which reads as a handler bug rather than an auth bug; - correct key, wrong key and no key produced identical errors — the tell. A gate that cannot distinguish those is not authenticating anything.
Site-a's log at the time: three PullAuditEvents calls, three identical
System.InvalidOperationException: Multiple constructors accepting all given argument types have been found in type 'ControlPlaneAuthInterceptor'.
The full suite was green when this shipped — 29 suites, 6,872 tests, 0 failures. The
in-process end-to-end test missed it because it registered the interceptor with
AddSingleton alongside AddGrpc, so DI returned the instance and gRPC's activation path
never ran.
Fixed in 228ff8b4: the prefix-set constructor is internal; the end-to-end harness now
registers exactly as Program.cs does (by type, not in DI); and a reflection assertion pins
"exactly one public constructor", since that is the actual invariant.
Lesson for phases 1A/1B, which both add services to this interceptor: extend
DefaultGatedPrefixes; do not add a second public constructor. And any in-process harness for
a DI-activated component must mirror the production registration shape or it proves less than
it appears to.
Test suite alongside the gate
Non-Playwright: 29 suites, 6,872 tests, 0 failures.
Playwright (against this rig): 170 passed, 2 failed, 1 skipped of 173. Both failures were
run down to root cause and both are pre-existing on main, unrelated to Phase 0 — this
branch touches no EF, CentralUI, Transport or ManagementService file (git diff --stat main...HEAD -- src/ is 15 files, all Communication/Host/AuditLog gRPC plumbing).
An earlier run of this suite reported 44 failures. That run is void: a docker/deploy.sh
was recreating the cluster underneath it, so the fast LoginTests/NavigationTests failures
were "app unreachable", not defects.
1. TransportImportTests.ImportSyntheticBundle_AppliesAndShowsAuditDrillIn — a real
production bug, not a test defect. Central's log during the failure:
[ERR] An exception occurred while iterating over the results of a query ...
System.InvalidOperationException: The configured execution strategy
'SqlServerRetryingExecutionStrategy' does not support user-initiated transactions.
at Microsoft.EntityFrameworkCore.Query.Internal.SplitQueryingEnumerable`1.AsyncEnumerator.MoveNextAsync()
BundleImporter.cs:1298 opens a user-initiated transaction; the central context is configured
with EnableRetryOnFailure (ConfigurationDatabase/ServiceCollectionExtensions.cs:33). SQL
Server's retrying strategy refuses to run a split query inside a caller's transaction, so
bundle import fails against real MS SQL. The fix is the one the exception names: wrap the
transaction in Database.CreateExecutionStrategy().ExecuteAsync(...).
Why the whole unit/integration suite is green on it: those tests use the in-memory EF provider,
which has no retrying execution strategy — and BeginTransactionAsync is a no-op there. The
comment directly above line 1298 documents that divergence without drawing the conclusion. Only
a rig-backed test can see this.
2. SmsNotificationE2ETests.SmsConfigPage_CreateOrRender_NeverLeaksAuthToken — a stale test
fixture. No server-side error at all: the page renders 200, and no INSERT INTO SmsConfigurations is ever issued. The test's fixture SID is ACtest123 (d6ead8ae,
2026-06-19). SmsConfiguration.razor:231 rejects anything not matching ^AC[0-9a-fA-F]{32}$,
added by 40088a21 (2026-07-10) to close an un-escaped URI-interpolation hole. Save() sets
_formError and returns — no toast, exactly as observed. The fixture was never updated.
This has been failing since 2026-07-10, and it matters more than a red line: everything after the toast assertion — including the secret-non-leak assertion that the Auth Token value never reaches the page HTML — has not executed since. Fix is a valid 32-hex SID in the fixture.
Not covered by this gate
- Streaming subscriptions were exercised in-process (TestServer), not over the rig. The
interceptor is path-scoped, not method-scoped, so the rig's
PullAuditEventsevidence covers the same code path — but a liveSubscribeInstanceunder load is untested here. - Key rotation on a live pair.
docker-env2was updated with its own key but not redeployed or gated.
Phase 1A — central control plane (site→central over gRPC) — PASS (2026-07-22)
Branch feat/grpc-central-control @ 0e162cb2. Rig rebuilt; site-a flipped to
CentralTransport=Grpc with CentralGrpcEndpoints=[central-a:8083, central-b:8083],
site-b/c left on Akka (default) to prove coexistence. The site-a flag flip was a
DoD-test-only rig edit — reverted from the branch, never committed (the plan keeps the default
Akka until Phase 4).
The defect this gate caught (T1A.2 shipped it; fixed in 0e162cb2)
First rebuild: central's entire HTTP surface was gone. central-a logged only
Now listening on: http://[::]:8083 — no :5000. Central UI, the Management + Inbound API,
and every /health/* endpoint (Traefik routing + IActiveNodeGate both depend on them) were
dead. The node booted, joined the cluster and served gRPC fine; no startup error.
Cause: builder.WebHost.ConfigureKestrel(o => o.ListenAnyIP(8083, Http2)) puts Kestrel into
explicit-endpoints mode, which suppresses the URLs from ASPNETCORE_URLS/--urls — it is
not additive, contrary to the comment T1A.2 shipped. Central's whole HTTP/1 surface lives on
that URL (http://+:5000 on the rig; a different port in production). The site branch has the
same ConfigureKestrel shape but nothing on ASPNETCORE_URLS to lose — it binds every port it
needs explicitly — which is why the pattern looked safe.
Every unit + E2E test uses TestServer, which never binds real Kestrel, so the whole suite
(6,872) stayed green. Only a live node exposes a missing listener. Fix: parse the port(s) from
the configured URLs and re-declare them (Http1AndHttp2) alongside the gRPC port (Http2) in
the one ConfigureKestrel call — Program.ParseHttpBindPorts + CentralHttpBindPortsTests
(14 cases). After the fix: Now listening on: http://[::]:5000 and :8083; 9001 ready
200/active Healthy, 9002 standby, Traefik LB 200, CLI over the LB works.
Checks (post-fix rebuild)
| # | Check | Result |
|---|---|---|
| 1 | site-a rides authenticated gRPC to CentralControlService |
PASS — Heartbeat, ReportSiteHealth, ReconcileSite all HTTP/2 → 200; 0 auth failures on either central |
| 2 | Heartbeat drives the active flag | PASS — 144+ heartbeats over gRPC; site-a online=True at central |
| 3 | Health page live | PASS — ReportSiteHealth lands; central shows site-a online=True, sequence advancing, alongside Akka site-b/c |
| 4 | Reconcile works after site restart | PASS — restarted site-a-a; it logged Site→central transport: gRPC to 2 central endpoint(s), then Reconcile pass … complete: 0 fetched, 0 failed, 0 orphan(s) |
| 5 | Coexistence | PASS — site-b/c log Created ClusterClient to central; both online=True — Akka and gRPC sites side by side |
| 6 | Central HTTP surface intact under the new gRPC listener | PASS — after the fix (see above) |
Not independently exercised on this gate
- Notification e2e (
SubmitNotification/QueryNotificationStatus) and audit ingest (IngestAuditEvents/IngestCachedTelemetry) were NOT driven live: the rig has templates but no deployed instance, so nothing emits site→central notifications or audit rows on its own. These four RPCs traverse the identicalCentralControlGrpcService→CentralCommunicationActorAsk path that checks 1–3 proved live under real auth, and their payload mappers carry 32 round-trip goldens — but the payloads themselves were not put over the wire here. Phase 2's central-kill S&F soak is where they get their live workout; flag for a fuller 1A proof if a deployed-instance rig is set up before then. - Cross-node failover/failback of the site→central channel under a central-node kill (unit-proven via TestServer; not exercised on the rig at 1A).
Phase 1B — site command plane (central→site over gRPC) — PASS (proportionate) (2026-07-23)
Branch feat/grpc-sitecommand (rebased onto 1A-merged main). Rig rebuilt with central
flipped to SiteTransport=Grpc — a DoD-test-only edit to both central appsettings, reverted
from the branch (default stays Akka).
SiteTransport is a central-wide flag (CentralCommunicationActor.SelectTransport picks one
transport for all sites), so the plan's "flip for site-a only" is not achievable — the flip
routes central→site commands for all three sites to gRPC. Command-plane per-site coexistence
therefore cannot be shown (unlike the site→central plane in 1A, which is per-site). This is a
plan-vs-code finding, recorded rather than worked around.
Checks
| # | Check | Result |
|---|---|---|
| 1 | central→site commands ride authenticated gRPC SiteCommandService |
PASS — ExecuteQuery (event-log) and ExecuteParked (parked query) HTTP/2 → 200 |
| 2 | Per-site PSK resolution across all sites | PASS — site-a, site-b, site-c each answered ExecuteQuery → 200 under its own SB-GRPC-PSK-{site}; 0 auth failures on any site node |
| 3 | Central HTTP surface intact under the central-wide gRPC flip | PASS — central :5000 and :8083 both listening; 9001 ready 200, LB 200 (the 1A Kestrel fix carried through the merge) |
| 4 | Query round-trips return correct data | PASS — every health event-log/parked-messages returned success:true with the right siteId and empty result sets (bare rig) |
Not driven on this proportionate gate
TriggerSiteFailover— unit-proven (two ordering tests pin ack-before-Leavevia the dispatcher's dry-run resolve + deferredCommitLeave), but not live-driven here: it is destructive (forces the active node to leave) and has no CLI verb (UI/management-only).- Tag commands (
BrowseNode/ReadTagValues/WriteTag) and the lifecycle enable/disable matrix — need a deployed instance + data connection the bare rig lacks (same blocker as 1A). - Parked retry against the STANDBY node — needs a parked operation to exist, which needs a deployed instance.
- Command-plane coexistence (site-b/c on Akka while site-a on gRPC) — not expressible; the flag is central-wide (check-1/2 instead prove all three sites over gRPC with distinct keys).
The instance-dependent matrix (tag ops, lifecycle, standby parked retry) and TriggerSiteFailover
get their live exercise in Phase 3's full UI command matrix; the transport itself is proven here.
Phase 2 — full site→central cutover + S&F soak — PASS (2026-07-23)
All three sites (all 6 nodes) flipped to CentralTransport=Grpc (edit to
docker/site-*/appsettings.Site.json, reverted in git after the gate — defaults stay Akka),
rebuilt from main + --force-recreate. Central left on SiteTransport=Akka (Phase 3 owns
that direction). The gate was driven by a live S&F workload, closing the notification/audit
path that 1A/1B could only unit-prove.
S&F driver
A minimal dependency-free template SoakNotify (one 5 s Interval script:
Notify.To("Engineering Alerts").Send(...)), 3 instances deployed+enabled on site-a → a steady
3 notifications per 5 s bucket (the pre-existing Motor Controller "soak-motor" instances
need 30 OPC UA bindings and were unusable). The central dbo.Notifications table (one row per
NotificationId, insert-if-not-exists) is the no-loss/no-dupes source of truth.
Checks
| # | Check | Result |
|---|---|---|
| 1 | All 6 site nodes on CentralTransport=Grpc |
PASS — each logs Site→central transport: gRPC to 2 central endpoint(s); 0 PermissionDenied across all six for the whole run |
| 2 | Full control plane rides gRPC CentralControlService |
PASS — central sees Heartbeat, ReportSiteHealth, SubmitNotification (the S&F path), IngestAuditEvents — 196 RPCs / 90 s, 0 non-200 |
| 3 | Health/heartbeat cadence unchanged, no sequence regressions | PASS — CentralHealthAggregator logged 0 sequence-regression/out-of-order lines; heartbeat steady at ~144/30 s across 3 sites |
| 4 | Single-node failover — hard-kill the active central (central-a) | PASS — CentralChannelProvider logged sticky failover central-a:8083 → central-b:8083 at the instant of kill; central-b active in 29 s (auto-down); notif count froze at 72 during the gap |
| 5 | Buffer drains, no loss/dupes (single-node) | PASS — count resumed 72→101; every 5 s bucket through the outage = exactly 3, no gap; 101 total == 101 distinct |
| 6 | Failback — restart central-a | PASS — rejoined ready in ~5 s as standby (active=503); central-b retained active (oldest-Up, no role flap); traffic uninterrupted (uniform 3/bucket across failback); central-a singletons → Younger |
| 7 | Full central outage — hard-kill both central (~59 s) | PASS — count frozen at 155 for the entire outage; sites buffered continuously |
| 8 | Cold re-form + drain, no loss/dupes (both-down) | PASS — cold cluster re-formed, central-b active in ~14 s; ~42 buffered notifications drained; every 5 s bucket across the whole ~59 s both-dead window = exactly 3, no gap; 216 total == 216 distinct |
Checks 4–8 also stand as a live preview of Phase 5 checks 4 (failover/failback) and 5 (mid-drain kill, zero loss/zero dupes).
Notes
- Audit telemetry also rides the new gRPC
CentralControlService(IngestAuditEvents) — theSiteAuditTelemetryActor"ClusterClientSiteAuditClient" label is legacy naming, not the wire path. SoCentralTransport=Grpcmoves heartbeat, health, notification S&F and audit off ClusterClient in one flip. - Sites settled on central-a as the gRPC endpoint after the cold both-restart while central-b
held the active/singleton role — the gRPC endpoint node and the singleton host can differ; central-a
receives the forward and Akka-routes to the
NotificationOutboxActorsingleton on central-b. Both are correct and independent. - Some notifications land
Parkedat central (no SMTP config on the bare rig) — irrelevant to the transport proof: theNotificationsrow is written on forward regardless of downstream SMTP delivery, so the count is a faithful no-loss/no-dupes measure. - Rig left running on the gRPC build; git config reverted to Akka default (a redeploy from
mainresets to all-Akka).
Phase 3 — full central→site cutover + command matrix — PASS (2026-07-23)
Central flipped to SiteTransport=Grpc (both central nodes; the flag is central-wide —
CentralCommunicationActor.SelectTransport). Sites kept on CentralTransport=Grpc from Phase 2,
so the rig ran both directions on gRPC simultaneously — the eventual all-gRPC end state.
Config bind-mounted (:ro), so a container --force-recreate (no image rebuild — code unchanged
since Phase 2) applied it. Central logged the selection at startup:
central→site command transport: gRPC (SiteCommandService).
The Phase 2 SoakNotify instances (#95–97) survived the recreate as Enabled (site volume
persisted deployed_configurations this time), so the notification workload kept flowing —
and lifecycle commands became drivable, closing the 1B/Phase-2 gap.
Checks
| # | Check | Result |
|---|---|---|
| 1 | Central selects the gRPC command transport | PASS — central→site command transport: gRPC (SiteCommandService); 0 Created ClusterClient to site lines on either central |
| 2 | ExecuteQuery (event-log) over gRPC, all 3 sites |
PASS — all 3 returned correlationIds; site logs SiteCommandService/ExecuteQuery - 200 (site-a-a/b-a/c-a each served the calls; standby nodes 0 — work lands on the active site node via the singleton proxy) |
| 3 | ExecuteParked (parked query) over gRPC, all 3 sites |
PASS — all 3 returned success payloads over SiteCommandService/ExecuteParked |
| 4 | ExecuteLifecycle (disable→enable) over gRPC, site-a |
PASS — instance disable/enable #95 → success:true; site served SiteCommandService/ExecuteLifecycle (×4), instance state toggled — NEW live coverage vs 1B/Phase-2 (needed a deployed instance) |
| 5 | Site-node kill mid-command → clean error, no hang | PASS — killed the active site-a node (site-a-a) during a 1 s query loop: the in-flight call returned TIMEOUT at dur=30.1 s = the QueryTimeout deadline — bounded, not an indefinite hang (see note) |
| 6 | Site-pair failover mid-stream | PASS — the very next query (~32 s after kill) and all subsequent ones succeeded automatically via site-a-b; SitePairChannelProvider failed the gRPC channel NodeA→NodeB and the site singleton migrated; site→central S&F never stopped (notif count climbed 619→715 through the kill, still no dupes) |
| 7 | No PSK drift | PASS — 0 PermissionDenied/Unauthenticated across all 8 nodes for the whole run |
| 8 | Zero ClusterClient activity on the flipped path | PASS — central built no site ClusterClients; command routing is entirely SiteCommandService gRPC |
Note on check 5 (deadline vs fast-fail)
The command in flight when site-a-a was hard-killed (SIGKILL) waited the full 30 s
QueryTimeout rather than failing fast on connect-refused: an already-dispatched gRPC call on a
dropped connection isn't observed as unsent, so it correctly cannot be auto-retried on the peer
node (it might have executed) and returns the deadline error to the caller — exactly the plan's
"deadline ≠ retry" rule. The deadline is the backstop; "no hang beyond deadline" is satisfied
(30.1 s). Only provably-unsent connect failures fail over fast, which is why every subsequent
call recovered immediately via site-a-b.
Not driven on this gate (unchanged from 1B/Phase 2)
ExecuteOpcUa(BrowseNode/ReadTagValues/WriteTag) — needs an OPC-bound deployed instance; the bare rig's only deployable template (SoakNotify) has no data connection, andMotor Controllerneeds 30 OPC UA bindings. Unit-proven (dispatcher routing ×28).ExecuteRoute(inbound-API → routed site script) — needs an inbound method + routing target the rig lacks.TriggerFailover— no CLI verb (UI/management-only), destructive; unit-proven (ack-before-Leaveordering tests). The hard-kill in check 6 is the live equivalent of a site-pair failover.
Rig left running with both transports on gRPC; git config reverted to Akka default (a redeploy
from main resets to all-Akka).
Phase 5 — full eight-check gate on the deletion build — PASS (2026-07-23)
The migration's terminal gate, run on main after Phase 4 (7fd5cb2b) — the build where the
Akka ClusterClient/ClusterClientReceptionist transport is physically deleted, the
CentralTransport/SiteTransport flags are gone, and gRPC is the only site↔central transport
(no flags to flip). Rig rebuilt from main via bash docker/deploy.sh, all 9 containers recreated
from the new image; central MS SQL and the S&F driver (SoakNotify template #2147 / instances
#95–97 / list "Engineering Alerts" #28) persisted from prior phases.
Provenance guard: the pre-existing rig image was built 11:17, before the Phase 4 commit
(12:54); it still carried ClusterClient code and would have invalidated checks 7–8. Confirmed
by rebuild timestamp, then by 0 ClusterClient/receptionist log lines across all 8 nodes on the
fresh boot. (The lone ClusterClientReceptionist symbol still present is inside
Akka.Cluster.Tools.dll — the library we deliberately keep for ClusterSingleton — not our code.)
Header names for the negative probes: authorization: Bearer <psk> + x-scadabridge-site: <siteId>
(ISitePskProvider.cs). Docker rig PSKs: dev-grpc-psk-docker-site-{a,b,c}
(ScadaBridge__Communication__SitePsks__site-*, central override; sites read their own
GrpcPsk from mounted appsettings.Site.json).
Checks
| # | Check | Result |
|---|---|---|
| 1 | PSK negatives | PASS (with a contract clarification) — see below |
| 2 | Site→central matrix | PASS — notifications, both audit paths, health, heartbeat, reconcile all live over gRPC |
| 3 | Central→site matrix | PASS (proportionate) — ExecuteQuery/ExecuteParked/ExecuteLifecycle live; instance-dependent RPCs carried forward (see below) |
| 4 | Failover / failback | PASS — active-central kill → sticky flip in 1 s, central-b active in 26 s; failback kept central-b active (oldest-Up, no flap) |
| 5 | Mid-drain kill, zero loss/dupes | PASS — buffered notifications flushed; total == distinct across the outage |
| 6 | Frame-class retirement (>128 KB) | PASS — a single 297,574-byte gRPC reply succeeded (Akka's 128 KB frame would have dropped it) |
| 7 | No cross-boundary Akka association | PASS — each Akka cluster's membership is strictly its own pair; 0 cross-boundary association lines both directions |
| 8 | Full-rig restart discipline | PASS — all 9 restarted together came up clean; 0 receptionist/ClusterClient lines on any node; sites reconnected over gRPC; S&F resumed with no dupes |
Check 1 — PSK negatives (and the contract clarification)
Against the two gated services on site-a (:9023), using the local proto descriptors
(-proto sitestream.proto / site_command.proto; server reflection is off in this build):
SiteStreamService/PullAuditEvents:
no credentials -> PermissionDenied "Control plane authentication failed."
wrong key (site-b's key) -> PermissionDenied
correct key -> success (audit events returned)
SiteCommandService/ExecuteQuery:
no credentials -> PermissionDenied
wrong key (site-b's key) -> PermissionDenied
site-a-a's log recorded exactly 4 control-plane rejections — the four deliberate negative probes above, nothing else.
Clarification the gate produced: the plan listed "missing x-scadabridge-site ⇒
PermissionDenied" as a negative case. It does not hold on the site side, and that is
correct by design. ControlPlaneAuthInterceptor.Authorize compares the presented bearer against
the node's own single GrpcPsk (_options.Value.GrpcPsk) and never reads
x-scadabridge-site. That header is a central-side routing hint — SitePskProvider uses it to
pick which site's key to expect, because central holds many. A single-key site node needs no such
hint. Per-site key isolation — the actual security property — is proven by the wrong-key
rejection (site-b's key refused at site-a), not by the header. Recorded rather than "fixed": adding
a header requirement to the site interceptor would be theatre.
LocalDb sync unaffected: the replicated node (site-a) logged one boot-order
faulted; reconnecting (peer node-b not yet up), the passive peer accepted the
/localdb_sync.v1.LocalDbSync/Sync POST 1 s later, and the health report settled to
localDbReplicationConnected:true, localDbOplogBacklog:0. 0 LocalDb sync auth failures — the
control-plane interceptor and the separate LocalDbSyncAuthInterceptor don't interfere. site-b/c
report localDbReplicationConnected:false — the intended rig posture (only site-a replicated).
Check 2 — site→central matrix
- Notifications e2e:
dbo.Notificationsclimbed continuously at the driver's ~3 rows / 5 s, total == distinct at every sample (insert-if-not-exists no-loss/no-dupe truth). - Both audit paths live over gRPC: in a 3-minute window,
dbo.AuditLogcarriednode-a/node-brows (105 each — site-originated, forwarded over the gRPCIngestAuditEventsRPC) andcentral-arows (420 — central-direct-write from the outbox dispatcher). Row count climbing live. - Health live per site:
health summaryshows all three sitesisOnline:true, sequence numbers advancing, fresh heartbeats; site-aenabledInstanceCount:3with a liveNotificationS&F buffer. - Heartbeat → active flag: central-a
active=200, central-bactive=503; each site's report drives its online flag. - Reconcile self-heal: restarting
site-a-aproduced a fresh (17:11:24)Site→central transport: gRPC to 2 central endpoint(s)thenReconcile pass … complete: 0 fetched, 0 failed, 0 orphan(s); S&F never dropped across the restart (active node kept emitting), no dupes. - Not driven (carry-forward, bare rig): cached-call telemetry (
SoakNotifyonly callsNotify.Send, noCachedCall/CachedWrite); a notification reaching Delivered rather thanParked(no SMTP on the rig — the transport truth is the row-count, per Phase 2).
Check 3 — central→site matrix
ExecuteQuery (event-log, all 3 sites → correlationIds + entries), ExecuteParked (parked-messages,
all 3 sites → empty, no parked ops on a bare rig), and ExecuteLifecycle (disable→enable #95 →
success:true both) all rode SiteCommandService gRPC; site active nodes logged them
(site-a-a: 8 ExecuteLifecycle + 4 ExecuteParked + 4 ExecuteQuery; site-b-a/c-a:
ExecuteParked + ExecuteQuery). Carried forward, unchanged from Phase 3 (needs a deployed
OPC-bound instance / inbound method / a parked op / the destructive UI-only failover verb the bare
rig lacks): ExecuteOpcUa, ExecuteRoute, standby parked retry/discard, TriggerFailover
(the hard-kill in Check 4 is its live equivalent). All four are unit-proven (dispatcher routing ×28,
ack-before-Leave ordering).
Checks 4 & 5 — failover / failback / mid-drain kill
Hard-killed the active central (central-a) mid-drain: site-a-b logged the sticky endpoint flip
central-a:8083 → central-b:8083 1 s after the kill; central-b reached active in 26 s
(auto-down). The notification count froze during the gap, then flushed the buffered rows and
resumed at the steady 3 / 5 s. Post-drain total == distinct (4395 == 4395) — zero loss, zero
duplicates. Failback: restarting central-a, it rejoined and central-b retained active
(a=503 standby / b=200 active) — oldest-Up, no role flap.
Check 6 — frame-class retirement
PullAuditEvents{since_utc: 2020-01-01, batch_size: 5000} against the active site node returned a
single 297,574-byte reply successfully. Over Akka's default 128 KB frame
(log-frame-size-exceeding off) this class of payload was silently dropped and the caller's Ask
timed out (docs/known-issues/2026-06-26-…). On gRPC (4 MB default cap) it is a normal reply.
Check 7 — no cross-boundary Akka association
Each Akka cluster's membership is strictly its own pair: central [central-a, central-b], site-a
[site-a-a, site-a-b], site-c [site-c-a, site-c-b] — no site node ever appears in central's
membership and vice versa. Central's log holds 0 Association with remote system …site-*
lines and 0 references to any site address; site-a's log holds 0 references to any central
address. The two boundaries are wired only by per-pair Akka remoting (8082, internal) and the gRPC
control/data planes (8083); there is no Akka association across the site↔central line.
Check 8 — full-rig restart discipline
Restarted all 9 rig containers together (pairs together, honouring the LocalDb-replication constraint).
Central re-formed with central-a active; 0 receptionist/ClusterClient lines across all 8 nodes;
all three sites logged Site→central transport: gRPC to 2 central endpoint(s); S&F resumed
(4407 → 4483, distinct == total, no dupes).
End state
The migration is complete and proven on the deletion build. The rig runs the committed all-gRPC
default (Phase 4 flipped the shipped appsettings to CentralGrpcEndpoints — unlike Phases 2/3, a
redeploy from main now yields all-gRPC, not all-Akka). No ClusterClient remains anywhere in the
running system or the source tree. Deferred, unchanged from earlier phases: the instance-dependent
central→site RPC matrix (ExecuteOpcUa/ExecuteRoute/standby parked retry) and live
TriggerFailover, all unit-proven; PSK rotation on a live pair. docker-env2 is now gated —
see below.
Env2 gate — docker-env2 on the gRPC PSK build — PASS (2026-07-23, Gitea #31)
The secondary Transport-testing topology (docker-env2/: 2 central + 1 site-x × 2 nodes, host
ports 91XX, shared infra) carried its GrpcPsk in config since Phase 0 but had never been
redeployed or gated against the migration build. Rebuilt scadabridge:latest from main @
8524a7f7 (includes the #28 Transport execution-strategy fix) and recreated only the env2
containers (bash docker-env2/deploy.sh; the primary rig's running image was untouched). Site-x
PSK dev-grpc-psk-docker-env2-site-x on both site nodes matches central's
ScadaBridge__Communication__SitePsks__site-x.
| # | Check | Result |
|---|---|---|
| 1 | Both site nodes boot with the key (StartupValidator passes) | PASS |
| 2 | Unauthenticated control-plane call ⇒ PermissionDenied; correct key ⇒ success | PASS |
| 3 | LocalDb unaffected | PASS |
Check 1 — both site nodes boot
scadabridge-env2-site-x-a and -b both reached Application started with Now listening on: http://[::]:8083 (the gRPC control/data listener). StartupValidator is fail-closed on a missing
ScadaBridge:Communication:GrpcPsk — a site node without it refuses to boot — so reaching
Application started is itself the proof the key is present and valid. Each node also logged a clean
Reconcile pass for site site-x node node-{a,b} complete: 0 fetched, 0 failed, 0 orphan(s), i.e.
the site→central gRPC path was already carrying traffic.
Check 2 — PSK auth on the control plane
grpcurl against the site-hosted SiteStreamService/PullAuditEvents on both nodes
(localhost:9123 = site-x-a, localhost:9124 = site-x-b):
- No auth header ⇒
PermissionDenied: Control plane authentication failed. - Wrong bearer key (
Bearer WRONG-KEY,x-scadabridge-site: site-x) ⇒PermissionDenied. - Correct key (
Bearer dev-grpc-psk-docker-env2-site-x) ⇒{}(empty audit set — the env has no S&F driver seeded — but not a permission error: auth accepted). Identical on both nodes.
(Same contract clarification as the primary rig: the site interceptor authenticates the node's single
GrpcPsk and does not require x-scadabridge-site; per-site isolation is proven by the wrong-key
rejection.)
Check 3 — LocalDb unaffected
Env2 site-x runs LocalDb local-only (no LocalDb:Replication:PeerAddress configured — the rig's
replication posture is proven on the primary docker/ stack, not here), so "unaffected" means the
consolidated site DB stays operational on the gRPC build. Site logs hold 0 LocalDb
errors/exceptions and both nodes boot healthy with the LocalDb at /app/data/site-localdb.db; the
clean reconcile passes above confirm the site DB + site→central path work end to end.
Bonus / observations
- No real ClusterClient. Central logs hold 0 ClusterClient/receptionist references; the site
logs' only match is the benign string label
client=ClusterClientSiteAuditClientin aSiteAuditTelemetryActor createdline (a legacy identifier, not an AkkaClusterClient) — identical to the primary rig's site nodes. No real receptionist/ClusterClient actor path exists. - Central registers site-x over gRPC.
Site site-x registered online via heartbeat— the site→central gRPC command/control path is live. - Seed-data note (not a migration issue).
ScadaBridgeConfig2.dbo.Siteshas 0 rows, so central logsReconcile request from unknown site 'site-x' … replying with empty gapand evicts it from the health aggregator as "no longer configured." This is a first-time-setup seeding gap (docker-env2/seed-sites.shnot yet run on the fresh DB), orthogonal to the transport — which demonstrably carries the heartbeat and reconcile regardless.