Full eight-check gate on the Phase 4 deletion build (main @ 7fd5cb2b),
rig rebuilt from main. All 8 checks PASS: PSK negatives; site->central
matrix (notif no-loss/dupe, both audit paths, reconcile self-heal);
central->site matrix (Query/Parked/Lifecycle); active-central kill ->
sticky flip 1s + central-b active 26s; mid-drain total==distinct;
297,574-byte gRPC reply (128KB frame-class retired); cluster membership
pair-only (no cross-boundary Akka association); full-rig restart 0
receptionist/ClusterClient lines on any of 8 nodes.
Check-1 clarification recorded: site-side interceptor uses the node's
single GrpcPsk and ignores x-scadabridge-site (a central-side routing
hint) -- per-site isolation proven by wrong-key reject. Instance-
dependent central->site RPCs (OpcUa/Route/standby-parked/TriggerFailover)
carried forward, unit-proven.
28 KiB
ClusterClient → gRPC-only cross-cluster transport — implementation plan
Date: 2026-07-22. Design: ~/Desktop/scadaproj/scadabridge_clusterclient_to_grpc.md
(read it first, §7 especially — it contains the deep-dive corrections this plan builds on).
Goal: all site↔central traffic rides gRPC with PSK auth from ZB.MOM.WW.Secrets;
ClusterClient/ClusterClientReceptionist are deleted; Akka remoting never crosses the
site↔central boundary.
Everything below was code-verified 2026-07-22. If a cited line has drifted, re-locate by the quoted identifier, never by line number.
How to execute this plan (read before starting)
- Phase order: 0 → (1A ∥ 1B) → (2 ∥ 3) → 4 → 5. Phases marked ∥ are parallel-safe only in
separate git worktrees (
git worktree add ../ScadaBridge-1B feat/grpc-sitecommand) — never run two agents against one working tree (destructive git races are a known family incident). Designated merge order when tracks meet: 1A lands first, 1B rebases (expected conflicts are confined toZB.MOM.WW.ScadaBridge.Communication.csprojproto ItemGroups andProgram.csservice/Map blocks — resolve as union). - Branches:
feat/grpc-phase0-psk,feat/grpc-central-control(1A),feat/grpc-sitecommand(1B), then per-phase branches; PR each phase tomainwith its DoD met. - Build/test:
dotnet build ZB.MOM.WW.ScadaBridge.slnx(0 warnings — TreatWarningsAsErrors),dotnet test ZB.MOM.WW.ScadaBridge.slnx. Tests are xunit v2 (2.9.3) +Akka.TestKit.Xunit21.5.62 + NSubstitute. TestKit tests inheritTestKit— never hand-rollActorSystem.Create+join. - Proto codegen is CHECKED-IN, not build-time (protoc segfaults in the linux_arm64 image).
For every new/changed proto follow the sitestream recipe documented in
src/ZB.MOM.WW.ScadaBridge.Communication/ZB.MOM.WW.ScadaBridge.Communication.csproj: temporarily uncomment/add the<Protobuf Include=... GrpcServices="Both" />item, delete stale generated files,dotnet buildon macOS, copyobj/**/Protos/*.csinto a committed folder (mirrorSiteStreamGrpc/), re-comment the item.docker/regen-proto.shexists. - Rig:
docker/deploy.shrebuilds the 2-central + 3×2-site cluster (+ traefik). Central UI9001/9002:5000, site gRPC90x3/90x4:8083. External sites↔central rules: never host-sqlite3a live WAL DB;aspnet:10.0has nocurl; seeding viadocker/seed-sites.sh. - Coexistence rule: every migrated path sits behind a config flag with the Akka implementation as default until Phase 4. Rollback at any point = flip the flag.
The two choke points (where all transport code changes)
- Site→central:
SiteCommunicationActor(src/ZB.MOM.WW.ScadaBridge.Communication/Actors/SiteCommunicationActor.cs) — 7 message types, allClusterClient.Sendto/user/central-communication(Ask, exceptHeartbeatMessageTell). - Central→site:
CentralCommunicationActor(same dir) — theSiteEnvelopehandler (:452-467) unwraps and routes via per-site ClusterClient. ALL producers (CommunicationService's 27 +SiteCallAuditActor's 2 relays) go through it.CommunicationServicegets NO interface extraction; its ~20 consumers are untouched.
Phase 0 — PSK auth + dead-code removal (standalone hardening; parallel-safe with 1A/1B prep)
T0.1 Delete the vestigial /user/management receptionist registration
AkkaHostedService.cs:459 (RegisterService of ManagementActor) — delete the registration
only (the actor stays; ManagementEndpoints.cs:117 asks it in-process via
ManagementActorHolder). Grep-verify nothing sends to /user/management. Update
docs/requirements/Component-Host.md REQ-HOST-6a to record the removal.
T0.2 File the dead-integration issue
IntegrationCallRequest (CommunicationService.cs:238) is unwired in production —
RegisterLocalHandler(Integration, …) exists only in SiteCommunicationActorTests.cs:102, so
production always replies "Integration handler not available" (SiteCommunicationActor.cs:134-136).
File a Gitea issue (decide: wire or delete); exclude it from the gRPC contract (28 of 29
commands migrate). Do not change its behavior in this program.
T0.3 PSK interceptor + options
New src/ZB.MOM.WW.ScadaBridge.Host/ControlPlaneAuthInterceptor.cs, copied from
LocalDbSyncAuthInterceptor.cs (same file layout: seal, 4 server handlers → shared
Authorize, authorization: Bearer metadata extraction, CryptographicOperations.FixedTimeEquals
over UTF-8, fail-closed when the expected key is unset, reject with
StatusCode.PermissionDenied). Differences from the template:
- Path scope: gates a set of service prefixes (constructor-provided), initially the
sitestream service prefix (read the real package/service name from
sitestream.proto—/…SiteStreamService/); later phases add the new services. LocalDb sync keeps its own interceptor + key untouched. - Key source, site side (server on
:8083): new optionCommunicationOptions.GrpcPsk(ScadaBridge:Communication:GrpcPsk), production value${secret:SB-GRPC-PSK-<siteId>}resolved by the existing pre-hostSecretReferenceExpander(Program.cs:56-65) — zero new resolution code. - Key source, central side (clients today; server in 1A): central's site set is dynamic, so
central resolves secret
SB-GRPC-PSK-{siteId}at channel-build time via the runtimeISecretResolver(copy the fail-closed lazy pattern fromDataConnectionLayer/Adapters/MxGatewayDataConnection.cs:82-98), cached per site, cache invalidated on site remove. New helperSitePskProviderin Communication (interface) + Host (implementation overISecretResolver).
Wire: add the interceptor to the site's existing AddGrpc (Program.cs:526-527, alongside the
LocalDb one). Attach the PSK on central's existing site-dialing clients as call-level
Metadata Authorization: Bearer <psk> (the LocalDb sync client pattern,
SyncBackgroundService.cs:82-86): SiteStreamGrpcClient (subscribe calls),
GrpcPullAuditEventsClient, GrpcPullSiteCallsClient — each already flows through a channel/
invoker creation point where the site id is known.
T0.4 Rig + tests
- Rig: mirror the LocalDb dev-key pattern — literal
ScadaBridge:Communication:GrpcPsk: "dev-grpc-psk-site-a"in BOTHdocker/site-a-node-*/appsettings.Site.json(and site-b/c with their own keys, since unlike LocalDb this is not optional), plus central-side dev secrets: seedSB-GRPC-PSK-site-{a,b,c}into both centrals' secret stores (secret CLI seed flow, or dev-KEK env as the rig's Secrets setup already does). - Tests: interceptor unit tests (wrong key / missing header / unset expected key ⇒
PermissionDenied; non-gated service path passes; constant-time compare exercised);SitePskProviderfail-closed test; one Host wiring test asserting the interceptor is registered.
Phase 0 DoD: suite green; on the rig, an unauthenticated grpcurl/test client gets
PermissionDenied on SiteStream, authenticated streaming + pull still work end-to-end. PR merged.
Phase 1A — Central control plane (site→central) — worktree A
T1A.1 central_control.proto
New Protos/central_control.proto in the Communication project (checked-in codegen per recipe),
package scadabridge.centralcontrol.v1, service CentralControlService:
| RPC | Wraps DTO (source file under Commons/Messages/) |
Notes |
|---|---|---|
SubmitNotification |
NotificationSubmit/NotificationSubmitAck (Notification/NotificationMessages.cs:30,47) |
11 fields incl. Guid? execution ids → string |
QueryNotificationStatus |
NotificationStatusQuery/Response (:55,62) |
|
IngestAuditEvents |
reuse existing AuditEventBatch/IngestAck from sitestream.proto |
import, don't duplicate; ForwardState/IngestedAtUtc stay off-wire (AuditEventDtoMapper.cs:23-27) |
IngestCachedTelemetry |
reuse CachedTelemetryBatch/IngestAck |
same |
ReconcileSite |
ReconcileSiteRequest/Response (Deployment/ReconcileSiteRequest.cs, ReconcileSiteResponse.cs incl. ReconcileGapItem) |
map<string,string> for name→hash |
ReportSiteHealth |
SiteHealthReport/SiteHealthReportAck (Health/SiteHealthReport.cs) |
the big one: ~30 fields incl. maps of ConnectionHealth enum, TagResolutionStatus, TagQualityCounts, NodeStatus list, SiteAuditBacklogSnapshot; model nullable ints/doubles with wrappers; keep SequenceNumber |
Heartbeat |
HeartbeatMessage (Health/HeartbeatMessage.cs) → google.protobuf.Empty reply |
fire-and-forget semantics preserved client-side (don't await failure into caller) |
Mappers in Communication/Grpc/CentralControlDtoMapper.cs with round-trip golden tests
(construct DTO → proto → DTO, assert deep-equal; include null-optional cases).
T1A.2 Central hosting
Central branch of Program.cs (~:262 services, :449-469 Map block):
builder.Services.AddGrpc(o => o.Interceptors.Add<ControlPlaneAuthInterceptor>()) — central's
interceptor variant verifies Bearer against the per-site PSK looked up by the required
x-scadabridge-site metadata header via SitePskProvider (fail-closed: missing header ⇒
PermissionDenied). Add an explicit Kestrel listener for h2c gRPC: new option
ScadaBridge:Node:CentralGrpcPort (default 8083, symmetric with sites), configured like the
site branch does (Program.cs:507-512, HttpProtocols.Http2); central's :5000 stays as-is
(Traefik is HTTP/1 — gRPC does NOT go through traefik). MapGrpcService<CentralControlGrpcService>().
CentralControlGrpcService (Host or Communication): decode proto → the SAME message types →
Ask the existing CentralCommunicationActor handlers (they already handle all 7 — zero
handler logic changes) → encode reply. Reuse the readiness convention: reject Unavailable
until the central actor system is up (mirror SiteStreamGrpcServer.SetReady).
T1A.3 Site-side client + transport seam
- New
ICentralTransport(Communication): one method per the 7 sends. Implementations:AkkaCentralTransport(extracted verbatim from today'sSiteCommunicationActorsend blocks) andGrpcCentralTransport(new). GrpcCentralTransport: channel pair with sticky failover + failback per design §3.5 — new sharedCentralChannelProvider: endpoints from new optionScadaBridge:Communication:CentralGrpcEndpoints(List, e.g.["http://scadabridge-central-a:8083","http://scadabridge-central-b:8083"]; validator: required when transport=Grpc); sticky-until-failure; flip on connect-fail/Unavailable; background failback probe every 30–60 s (gRPC health or aHeartbeatping); reconnect backoff copied fromSyncBackgroundService.cs:151(1 s doubling, cap 60 s). Attach PSK (GrpcPskoption) +x-scadabridge-siteon every call; per-call deadlines from the matchingCommunicationOptionstimeout (NotificationForwardTimeout,HealthReportTimeout, etc. — today's Ask timeouts, unchanged values). Cross-node auto-retry only on connect-fail/Unavailable— never onDeadlineExceeded.SiteCommunicationActorselects the implementation from new optionScadaBridge:Communication:CentralTransport(Akka|Grpc, defaultAkka). The 7 handler bodies delegate to the injected transport; reply/fault semantics identical (timeout or non-OK status ⇒ sameStatus.Failurethe S&F/audit layers already treat as transient).
T1A.4 Tests
Extend Communication.Tests (TestKit): SiteCommunicationActor with an NSubstitute
ICentralTransport — all 7 paths, fault propagation (transport throw ⇒ same failure the S&F
tests expect). GrpcCentralTransport unit tests with an in-process TestServer gRPC host:
failover flip, sticky behavior, failback probe, PSK attached, deadline set, no-retry-on-deadline.
Reuse/extend DirectActorSiteStreamAuditClient for the ingest integration harness.
NotificationForwarderTests/SiteAuditTelemetryActorTests/HealthReportSenderTests must pass
unmodified (they sit above the seam — if they need edits, the seam is wrong).
1A DoD: suite green; on the rig with site-a flipped to CentralTransport=Grpc
(central gRPC port published, e.g. 9013/9014:8083): notification e2e, audit rows land, health
page live, heartbeat drives active flag, reconcile works after site restart — while site-b/c
still run Akka (coexistence proven).
Phase 1B — Site command plane (central→site) — worktree B
T1B.1 site_command.proto
Package scadabridge.sitecommand.v1, service SiteCommandService — 28 commands (29 minus
dead IntegrationCallRequest), grouped into domain RPCs with oneof request/response
envelopes (full command list + reply types + CommunicationService.cs line refs in the design
doc §7 / recon inventory):
| RPC | Commands (count) | Deadline source |
|---|---|---|
ExecuteLifecycle |
RefreshDeployment, Enable/Disable/DeleteInstance, DeploymentStateQuery, DeployArtifacts (6) | DeploymentTimeout/LifecycleTimeout/ArtifactDeploymentTimeout |
ExecuteOpcUa |
BrowseNode, SearchAddressSpace, ReadTagValues, VerifyEndpoint, Trust/List/RemoveServerCert, WriteTag (8) | QueryTimeout (browse/search per existing Ask usage) |
ExecuteQuery |
EventLogQuery, DebugSnapshot, Subscribe/UnsubscribeDebugView (4) | QueryTimeout/DebugViewTimeout |
ExecuteParked |
ParkedMessageQuery/Retry/Discard, RetryParkedOperation, DiscardParkedOperation (5) | QueryTimeout; relay callers keep RelayTimeout(10s) < QueryTimeout(30s) ordering |
ExecuteRoute |
RouteToCall/GetAttributes/SetAttributes/WaitForAttribute (4) | IntegrationTimeout; WaitForAttribute uses its dynamic timeout |
TriggerFailover |
TriggerSiteFailover (1) | LifecycleTimeout |
Mappers SiteCommandDtoMapper.cs + round-trip golden tests for every command/reply (the bulk of
this track — budget accordingly; enums, TrackedOperationId struct → string guid, nullable
wrappers).
T1B.2 Site server: shared dispatcher
Refactor SiteCommunicationActor's receive table into SiteCommandDispatcher (pure routing:
message → _deploymentManagerProxy / _artifactHandler / _eventLogHandler /
_parkedMessageHandler / failover handler — preserving EXACTLY today's targets, including the
node-local parked-message handler; see design §7.3, the replicated-store semantics are
deliberate). The actor and a new SiteCommandGrpcService (mapped in the site branch next to
SiteStreamGrpcServer, gated by ControlPlaneAuthInterceptor + the readiness flag) both call
the dispatcher — one routing truth, both transports.
T1B.3 Central client + transport seam
ISiteCommandTransport (send SiteEnvelope-equivalent, Ask or Tell) injected into
CentralCommunicationActor; implementations AkkaSiteTransport (today's per-site
ClusterClient path, extracted) and GrpcSiteTransport. GrpcSiteTransport uses a new shared
SitePairChannelProvider: addresses from the Site entity's existing
GrpcNodeAAddress/GrpcNodeBAddress (the streaming path's columns — do NOT invent
LoadSiteAddressesFromDb, it doesn't exist; reuse the ISiteRepository reads +
CentralCommunicationActor's existing DB-driven cache-refresh loop :532-598 to build/refresh
channels instead of ClusterClients), sticky failover/failback per §3.5, PSK from
SitePskProvider + deadlines per the table above. Config flag
ScadaBridge:Communication:SiteTransport (Akka | Grpc, default Akka) selected inside
CentralCommunicationActor — CommunicationService and SiteCallAuditActor unchanged.
T1B.4 Tests
TestKit: CentralCommunicationActor with substitute ISiteCommandTransport (envelope routing,
Ask-sender reply plumbing, per-site transport lifecycle on site add/remove/change);
SiteCommandDispatcher unit tests (every command → correct target, incl. parked→local handler
and failover→local); SiteCommandGrpcService via TestServer (auth, readiness, one command per
oneof group); existing CommunicationServiceTests/CentralCommunicationActor*Tests pass with
the Akka implementation as default.
1B DoD: suite green; rig central flipped to SiteTransport=Grpc for site-a only: from
CentralUI — deploy refresh, enable/disable instance, browse, read tag, write tag, event-log
query, parked query/retry/discard (run retry against the STANDBY site node explicitly —
replicated-store semantics), TriggerSiteFailover — all work; site-b/c untouched on Akka.
Phase 2 (after 1A) ∥ Phase 3 (after 1B) — full cutover on the rig + hardening
- Phase 2: flip all sites to
CentralTransport=Grpc. Soak: S&F drain under central-a kill (rows stay Pending, resume without loss or duplicates — sequence/dedup layers unchanged), failback observed when central-a returns, health/heartbeat cadence unchanged inCentralHealthAggregator(no sequence regressions logged). - Phase 3: flip central to
SiteTransport=Grpcfor all sites. Soak: full CentralUI command matrix against each site; site-node kill mid-command returns a clean error (no hang beyond deadline); site pair failover mid-stream of commands. - Both phases: watch for
PermissionDeniednoise (would indicate PSK drift), and confirm zero ClusterClient log activity on flipped paths.
Phase 4 — deletion + config cutover (sequential, after 2+3)
- Flip both flag defaults to
Grpc; rig + docs updated; one soak cycle. - Delete:
AkkaCentralTransport/AkkaSiteTransport, ClusterClient creation (AkkaHostedService.cs:942-953),DefaultSiteClientFactory(+ its tests), per-site ClusterClient cache inCentralCommunicationActor(keep the DB refresh loop — it now feedsSitePairChannelProvider), receptionist registrations:436and:935(T0.1 already removed:459),CommunicationOptions.CentralContactPoints(+ validator + rig configs), then the flags themselves. - Grep-gates:
rg -i "clusterclient|receptionist" src tests docker docs→ only historical docs;rg "CentralContactPoints"→ empty. - Docs: update
grpc_streams.md(its "ClusterClient keeps command/control" split is superseded; also fix itsLoadSiteAddressesFromDbdoc-vs-code gap),Component-Host.md,Component-StoreAndForward.md:137, and add adocs/known-issuescross-ref note that the frame-size class is retired. KeepAkka.Cluster.Tools(ClusterSingleton still used).
Phase 5 — live gate (rig; sequential; every check PASS required)
- PSK negative: no key / wrong key / missing
x-scadabridge-site⇒PermissionDenied; unset server key ⇒ all rejected (fail-closed); LocalDb sync key unaffected. - Site→central matrix: notification e2e (delivered + status query), audit + cached-telemetry
rows in
dbo.AuditLog/site-calls,/monitoring/healthlive per site, heartbeat → active flag, reconcile self-heal after site-node restart. - Central→site matrix: all 6 RPC groups exercised from CentralUI, parked retry/discard on the standby node, failover command drains cleanly.
- Failover/failback: kill central-a → sites flip to central-b sticky (S&F uninterrupted);
restart central-a → failback within probe cadence; same for a site node from central's side;
booting node rejects
Unavailableuntil ready and the client fails over. - Mid-drain kill: kill central during an S&F drain burst — zero loss, zero duplicates.
- Frame-class retirement: issue a command/reply > 128 KB (large browse/event-log result) — succeeds over gRPC (impossible before).
- Boundary check: with everything on gRPC, verify NO Akka association exists between any
site container and central (
netstat/Akka logs) — remoting is pair-internal only. - Restart discipline: full-rig restart (pairs together) comes up clean; no receptionist/ ClusterClient log lines anywhere.
Record results in docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md (check-by-check,
the family's live-gate format).
Effort & parallelization summary
| Track | Est. | Parallel with |
|---|---|---|
| Phase 0 | 2–3 d | 1A/1B proto authoring |
| 1A | 1.5–2 wk | 1B (separate worktrees; 1A merges first) |
| 1B | 2–3 wk (mapper-heavy) | 1A |
| 2, 3 | 2–4 d each | each other (independent flags/paths) |
| 4 | 2–3 d | — |
| 5 | 2–3 d | — |
Critical path ≈ 1B: ~4–6 weeks total, matching the design estimate.
Task checklist (tick as you go; IDs reference the sections above)
Phase 0 — PSK + dead code (branch feat/grpc-phase0-psk)
- T0.1 Delete
/user/managementreceptionist registration (AkkaHostedService.cs:459) + Component-Host.md update - T0.2 File dead-
IntegrationCallRequestissue; record exclusion (28 of 29 migrate) - T0.3
ControlPlaneAuthInterceptor+CommunicationOptions.GrpcPsk+SitePskProvider; gate SiteStream; PSK attached on central's streaming + pull clients - T0.4 Rig dev keys (all 3 sites + central store seeds) + interceptor/provider/wiring tests
- Phase 0 DoD: suite green; rig unauthenticated ⇒
PermissionDenied, authenticated paths work; PR merged (#25, ff tomain@3fa95555; gate PASS in2026-07-22-clusterclient-to-grpc-live-gate.md)
Phase 1A — central control plane (worktree, feat/grpc-central-control)
- T1A.1
central_control.proto(7 RPCs; checked-in codegen) +CentralControlDtoMapper+ round-trip golden tests - T1A.2 Central hosting:
AddGrpc+ per-site-PSK interceptor (x-scadabridge-site),CentralGrpcPorth2c listener (8083),CentralControlGrpcService(Ask existing handlers), readiness gate - T1A.3
ICentralTransport(Akka extract + Grpc impl),CentralChannelProvider(sticky failover/failback, backoff, deadlines, PSK),CentralTransportflag defaultAkka,CentralGrpcEndpointsoption + validator - T1A.4 Tests: actor-with-fake-transport ×7, TestServer transport tests, S&F/audit/health suites pass unmodified
- 1A DoD: rig site-a on
Grpcproves site→central paths (heartbeat/health/reconcile + coexistence) while site-b/c stay Akka; PR #26 merged (aa60f438). Notification/audit deferred to Phase 2 soak (no deployed instance); rig caught + fixed a central:5000HTTP-drop regression (0e162cb2)
Phase 1B — site command plane (worktree, feat/grpc-sitecommand)
- T1B.1
site_command.proto(6 oneof RPCs / 28 commands) +SiteCommandDtoMapper+ round-trip golden tests (all 28 commands + 22 reply shapes + 18 nested types; reflection-driven coverage guard over the mapper surface and the generated oneof descriptors) - T1B.2
SiteCommandDispatcherrefactor (actor + newSiteCommandGrpcServiceshare it; parked stays node-local; failover ack-before-Leave via dry-run resolve + deferredCommitLeave; interceptorDefaultGatedPrefixesextended toSiteCommandService) - T1B.3
ISiteCommandTransportinCentralCommunicationActor(Akka extract + Grpc impl),SitePairChannelProvider(Site entity Grpc columns + DB refresh loop),SiteTransportflag defaultAkka - T1B.4 Tests: dispatcher routing ×28, actor envelope/reply plumbing, TestServer service tests, existing Communication suites green
- 1B DoD (proportionate): rig central on
SiteTransport=Grpcproves central→site rides authenticated gRPCSiteCommandServicefor all 3 sites (ExecuteQuery/ExecuteParked→ 200, per-site PSK, 0 auth failures); rebased on 1A. Instance-dependent commands (tag ops/lifecycle/standby parked retry) +TriggerSiteFailoverdeferred to Phase 3 (no deployed instance / destructive UI-only); command-plane coexistence not expressible (SiteTransportis central-wide). Gate:2026-07-22-clusterclient-to-grpc-live-gate.md
Phase 2 ∥ 3 — cutover + soak
- P2 All sites
CentralTransport=Grpc; central-kill S&F soak (no loss/dupes), failback observed, health sequences clean — PASS 2026-07-23 (live gate: single-node active-kill failover 29s + full-outage ~59s both-down drain; every 5s bucket = exactly 3 through both outages, 216 total == 216 distinct; 0 auth failures / 0 seq regressions; whole control plane — heartbeat/health/notification/audit — on gRPCCentralControlService) - P3 Central
SiteTransport=Grpcall sites; full UI command matrix per site; site-kill mid-command clean; no PSK noise; zero ClusterClient log activity on flipped paths — PASS 2026-07-23 (live gate:ExecuteQuery/ExecuteParkedall 3 sites +ExecuteLifecycledisable/enable on site-a, all 200 overSiteCommandService; active-site-node hard-kill mid-command → cleanTIMEOUTat the 30s deadline, next call failed over to site-a-b automatically; 0 PermissionDenied / 0 ClusterClient-to-site on both central.ExecuteOpcUa/ExecuteRoute/TriggerFailoverdeferred — no OPC-bound instance / no CLI verb, unit-proven)
Phase 4 — deletion
- gRPC made the only transport (flags DELETED, not flipped — end state is identical) + deleted
AkkaCentralTransport/AkkaSiteTransport, ClusterClient creation (AkkaHostedService),DefaultSiteClientFactory+ISiteClientFactory, both receptionist registrations (:436/:1001),CentralContactPoints+ theCentralTransport/SiteTransportflags +CentralTransportMode/SiteTransportKindenums.NoOpCentralTransportadded as the fail-loud null-default;CentralGrpcEndpointsnow unconditional (StartupValidator requires ≥1 on Site). KeptAkka.Cluster.Tools(ClusterSingleton). Rig configs (docker ×6, docker-env2 ×2, Host default, wonder-app-vd03) movedCentralContactPoints→CentralGrpcEndpoints. Full solution build 0/0; Communication.Tests 640, Host.Tests + StartupValidator + SiteActorPath green, audit-push integration green (2026-07-23) - Grep-gates pass (
CentralContactPoints→ only "replaces the former" doc refs + plan trackers; deleted symbols → 0 live refs; remainingclusterclientin src = the misleadingly-namedClusterClientSiteAuditClient[transport-agnostic, works unchanged] + stale inline doc-comments, noted as follow-up) - Docs updated:
Component-Communication.md,components/Communication.md,Component-Host.md,Component-StoreAndForward.md,topology-guide.md,grpc_streams.md, known-issues frame-size amendment, CLAUDE.md transport decisions
Phase 5 — live gate (record in 2026-07-22-clusterclient-to-grpc-live-gate.md) — PASS 2026-07-23 (deletion build, main @ 7fd5cb2b)
- 1 PSK negatives · [x] 2 site→central matrix · [x] 3 central→site matrix · [x] 4 failover/failback both directions · [x] 5 mid-drain kill · [x] 6 >128 KB frame-class proof · [x] 7 no cross-boundary Akka association · [x] 8 full-rig restart clean
- All 8 PASS. Notables: check 1 clarified the site-side gate uses the node's single
GrpcPskand ignoresx-scadabridge-site(central-side routing hint) — per-site isolation proven by wrong-key reject, not the header; check 6 = a 297,574-byte single gRPC reply (a payload Akka's 128 KB frame dropped); check 7 = Akka cluster membership strictly pair-only, 0 cross-boundary association; check 8 = full-rig restart, 0 receptionist/ClusterClient lines on any of the 8 nodes. Instance-dependent central→site RPCs (ExecuteOpcUa/ExecuteRoute/standby parked retry/TriggerFailover) carried forward unit-proven — bare rig has no OPC-bound instance / inbound method / parked op / CLI failover verb.
Gotchas for the executor (will bite; read twice)
- Generated proto C# is committed; never add an active
<Protobuf>item to the csproj in a final commit (linux_arm64 protoc segfault breaks the Docker build). HeartbeatMessagemust stay fire-and-forget end-to-end — don't let a gRPC failure surface as a fault to the heartbeat timer path.- Deadline ≠ retry: no automatic cross-node retry on
DeadlineExceededfor WriteTag/Deploy/ Failover; only on provably-unsent failures. - The parked-message handler is node-local on purpose (replicated store); do not "fix" it onto the singleton proxy.
- Ack-before-Leave on
TriggerSiteFailover(SiteCommunicationActor.cs:569): the gRPC reply must be written before the node leaves — verify the response completes under failover. - Inner-before-outer timeouts:
RelayTimeout(10 s) <QueryTimeout(30 s) must survive the deadline mapping (CommunicationService.cs:779-786documents why). SiteStreamGrpcServer.AuditIngestAskTimeout(30 s) is "one source of truth" shared withCentralCommunicationActor— keep the new central service on the same constant.- Rig sites reach central by container name (
scadabridge-central-{a,b}:8083), NOT via traefik (HTTP/1 only). - xunit v2: use
Xunit.SkippableFactfor env-gated tests, notAssert.Skip.