Files
ScadaBridge/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md
T
Joseph Doherty 2ee84af1c0 feat(grpc): PSK-authenticate the site gRPC control plane; drop the vestigial management receptionist registration
Phase 0 of the ClusterClient→gRPC migration
(docs/plans/2026-07-22-clusterclient-to-grpc-plan.md). Standalone hardening: it
closes a gap that exists today and is a precondition for moving command/control
onto gRPC in later phases.

T0.1 — delete the ManagementActor ClusterClientReceptionist registration.
It was built for an out-of-cluster CLI that was never written: the shipped CLI
speaks HTTP Basic to /management, which asks the actor in-process through
ManagementActorHolder. Nothing in the repo ever sent to /user/management. The
actor still runs there; only the cross-boundary advertisement is gone. Six
documents claimed the CLI used ClusterClient — including the CLI's own README
"Architecture Notes" — and are corrected here rather than left to rot.

T0.2 — record, do not port, the dead integration-routing path.
IntegrationCallRequest is unwired at BOTH ends: RouteIntegrationCallAsync has
zero callers anywhere, and RegisterLocalHandler(Integration, …) appears only in
a test, so production always answers "Integration handler not available". It is
excluded from the gRPC contract (28 of 29 commands migrate) rather than
enshrined on an additive-only wire format, and deleting it during a
transport migration would mix a behavioural change into a change whose whole
value is that behaviour is identical. See
docs/known-issues/2026-07-22-integration-call-routing-is-dead-code.md.

T0.3 — preshared-key authentication on SiteStreamService.
The service shipped with no auth at all: plaintext h2c, no interceptor, so
anything that could reach a site node's :8083 could open a live data stream or
read audit rows back via PullAuditEvents/PullSiteCalls. ControlPlaneAuthInterceptor
now gates /sitestream.SiteStreamService/ — modeled on LocalDbSyncAuthInterceptor
(constant-time compare, fail-closed, PermissionDenied) but gating a SET of
service prefixes so phases 1A/1B add services rather than interceptors. LocalDb
sync keeps its own separate key: it authenticates the pair partner, not central,
and collapsing the two would make a site's central-facing key also admit writes
into its database.

Keys are per site (SB-GRPC-PSK-<siteId>), never fleet-wide, so a compromised
site yields only its own. Central attaches them through ControlPlaneCredentials,
which binds CallCredentials to the channel — covering unary and streaming
uniformly, and letting the key resolve asynchronously, which a client
interceptor could not do without blocking. All three central→site channel
creation sites go through it (SiteStreamGrpcClient and both audit pull invokers);
the pull invokers' channel caches are re-keyed by (site, endpoint) because
credentials are per-site and bound to the channel.

Two decisions beyond the plan:

  * StartupValidator now requires GrpcPsk on Site nodes. The plan specified only
    the runtime gate, but fail-closed with no boot check produces a node that
    joins, answers heartbeats and reports healthy while refusing every stream,
    audit pull and telemetry ingest — silent and total. Same reasoning as the
    existing inbound API-key pepper rule.

  * Added Communication:SitePsks as a central-side key map. The plan assumed
    central would read the store, seeded via a dev KEK; the docker rig
    deliberately boots with no master key, so store-only resolution would leave
    it unable to dial its own sites. The store stays primary — it is the only
    source that can serve a site added at runtime — with the map covering
    key-less hosts and one-off pins. Neither source falling back to
    "unauthenticated" is the invariant.

T0.4 — dev keys on both rigs and tests.
34 tests. The seven that matter most exercise a real in-process gRPC stack over
TestServer: the unit tests on either side of the wire would both stay green if
the halves disagreed, and gRPC refuses call credentials on a plaintext channel
by default — the UnsafeUseInsecureChannelCallCredentials opt-in is only provable
by making a real call. They confirm correct key passes on unary AND streaming,
wrong key and no-credentials both get PermissionDenied, and an unresolvable key
fails the call with nothing reaching the service.

OPERATIONAL: a site node upgraded to this build without a key will not boot.
That includes the gitignored deploy/wonder-app-vd03/ overlay.
2026-07-22 17:51:09 -04:00

25 KiB
Raw Blame History

ClusterClient → gRPC-only cross-cluster transport — implementation plan

Date: 2026-07-22. Design: ~/Desktop/scadaproj/scadabridge_clusterclient_to_grpc.md (read it first, §7 especially — it contains the deep-dive corrections this plan builds on). Goal: all site↔central traffic rides gRPC with PSK auth from ZB.MOM.WW.Secrets; ClusterClient/ClusterClientReceptionist are deleted; Akka remoting never crosses the site↔central boundary.

Everything below was code-verified 2026-07-22. If a cited line has drifted, re-locate by the quoted identifier, never by line number.

How to execute this plan (read before starting)

  • Phase order: 0 → (1A ∥ 1B) → (2 ∥ 3) → 4 → 5. Phases marked ∥ are parallel-safe only in separate git worktrees (git worktree add ../ScadaBridge-1B feat/grpc-sitecommand) — never run two agents against one working tree (destructive git races are a known family incident). Designated merge order when tracks meet: 1A lands first, 1B rebases (expected conflicts are confined to ZB.MOM.WW.ScadaBridge.Communication.csproj proto ItemGroups and Program.cs service/Map blocks — resolve as union).
  • Branches: feat/grpc-phase0-psk, feat/grpc-central-control (1A), feat/grpc-sitecommand (1B), then per-phase branches; PR each phase to main with its DoD met.
  • Build/test: dotnet build ZB.MOM.WW.ScadaBridge.slnx (0 warnings — TreatWarningsAsErrors), dotnet test ZB.MOM.WW.ScadaBridge.slnx. Tests are xunit v2 (2.9.3) + Akka.TestKit.Xunit2 1.5.62 + NSubstitute. TestKit tests inherit TestKit — never hand-roll ActorSystem.Create+join.
  • Proto codegen is CHECKED-IN, not build-time (protoc segfaults in the linux_arm64 image). For every new/changed proto follow the sitestream recipe documented in src/ZB.MOM.WW.ScadaBridge.Communication/ZB.MOM.WW.ScadaBridge.Communication.csproj: temporarily uncomment/add the <Protobuf Include=... GrpcServices="Both" /> item, delete stale generated files, dotnet build on macOS, copy obj/**/Protos/*.cs into a committed folder (mirror SiteStreamGrpc/), re-comment the item. docker/regen-proto.sh exists.
  • Rig: docker/deploy.sh rebuilds the 2-central + 3×2-site cluster (+ traefik). Central UI 9001/9002:5000, site gRPC 90x3/90x4:8083. External sites↔central rules: never host-sqlite3 a live WAL DB; aspnet:10.0 has no curl; seeding via docker/seed-sites.sh.
  • Coexistence rule: every migrated path sits behind a config flag with the Akka implementation as default until Phase 4. Rollback at any point = flip the flag.

The two choke points (where all transport code changes)

  • Site→central: SiteCommunicationActor (src/ZB.MOM.WW.ScadaBridge.Communication/Actors/SiteCommunicationActor.cs) — 7 message types, all ClusterClient.Send to /user/central-communication (Ask, except HeartbeatMessage Tell).
  • Central→site: CentralCommunicationActor (same dir) — the SiteEnvelope handler (:452-467) unwraps and routes via per-site ClusterClient. ALL producers (CommunicationService's 27 + SiteCallAuditActor's 2 relays) go through it. CommunicationService gets NO interface extraction; its ~20 consumers are untouched.

Phase 0 — PSK auth + dead-code removal (standalone hardening; parallel-safe with 1A/1B prep)

T0.1 Delete the vestigial /user/management receptionist registration

AkkaHostedService.cs:459 (RegisterService of ManagementActor) — delete the registration only (the actor stays; ManagementEndpoints.cs:117 asks it in-process via ManagementActorHolder). Grep-verify nothing sends to /user/management. Update docs/requirements/Component-Host.md REQ-HOST-6a to record the removal.

T0.2 File the dead-integration issue

IntegrationCallRequest (CommunicationService.cs:238) is unwired in production — RegisterLocalHandler(Integration, …) exists only in SiteCommunicationActorTests.cs:102, so production always replies "Integration handler not available" (SiteCommunicationActor.cs:134-136). File a Gitea issue (decide: wire or delete); exclude it from the gRPC contract (28 of 29 commands migrate). Do not change its behavior in this program.

T0.3 PSK interceptor + options

New src/ZB.MOM.WW.ScadaBridge.Host/ControlPlaneAuthInterceptor.cs, copied from LocalDbSyncAuthInterceptor.cs (same file layout: seal, 4 server handlers → shared Authorize, authorization: Bearer metadata extraction, CryptographicOperations.FixedTimeEquals over UTF-8, fail-closed when the expected key is unset, reject with StatusCode.PermissionDenied). Differences from the template:

  • Path scope: gates a set of service prefixes (constructor-provided), initially the sitestream service prefix (read the real package/service name from sitestream.proto/…SiteStreamService/); later phases add the new services. LocalDb sync keeps its own interceptor + key untouched.
  • Key source, site side (server on :8083): new option CommunicationOptions.GrpcPsk (ScadaBridge:Communication:GrpcPsk), production value ${secret:SB-GRPC-PSK-<siteId>} resolved by the existing pre-host SecretReferenceExpander (Program.cs:56-65) — zero new resolution code.
  • Key source, central side (clients today; server in 1A): central's site set is dynamic, so central resolves secret SB-GRPC-PSK-{siteId} at channel-build time via the runtime ISecretResolver (copy the fail-closed lazy pattern from DataConnectionLayer/Adapters/MxGatewayDataConnection.cs:82-98), cached per site, cache invalidated on site remove. New helper SitePskProvider in Communication (interface) + Host (implementation over ISecretResolver).

Wire: add the interceptor to the site's existing AddGrpc (Program.cs:526-527, alongside the LocalDb one). Attach the PSK on central's existing site-dialing clients as call-level Metadata Authorization: Bearer <psk> (the LocalDb sync client pattern, SyncBackgroundService.cs:82-86): SiteStreamGrpcClient (subscribe calls), GrpcPullAuditEventsClient, GrpcPullSiteCallsClient — each already flows through a channel/ invoker creation point where the site id is known.

T0.4 Rig + tests

  • Rig: mirror the LocalDb dev-key pattern — literal ScadaBridge:Communication:GrpcPsk: "dev-grpc-psk-site-a" in BOTH docker/site-a-node-*/appsettings.Site.json (and site-b/c with their own keys, since unlike LocalDb this is not optional), plus central-side dev secrets: seed SB-GRPC-PSK-site-{a,b,c} into both centrals' secret stores (secret CLI seed flow, or dev-KEK env as the rig's Secrets setup already does).
  • Tests: interceptor unit tests (wrong key / missing header / unset expected key ⇒ PermissionDenied; non-gated service path passes; constant-time compare exercised); SitePskProvider fail-closed test; one Host wiring test asserting the interceptor is registered.

Phase 0 DoD: suite green; on the rig, an unauthenticated grpcurl/test client gets PermissionDenied on SiteStream, authenticated streaming + pull still work end-to-end. PR merged.


Phase 1A — Central control plane (site→central) — worktree A

T1A.1 central_control.proto

New Protos/central_control.proto in the Communication project (checked-in codegen per recipe), package scadabridge.centralcontrol.v1, service CentralControlService:

RPC Wraps DTO (source file under Commons/Messages/) Notes
SubmitNotification NotificationSubmit/NotificationSubmitAck (Notification/NotificationMessages.cs:30,47) 11 fields incl. Guid? execution ids → string
QueryNotificationStatus NotificationStatusQuery/Response (:55,62)
IngestAuditEvents reuse existing AuditEventBatch/IngestAck from sitestream.proto import, don't duplicate; ForwardState/IngestedAtUtc stay off-wire (AuditEventDtoMapper.cs:23-27)
IngestCachedTelemetry reuse CachedTelemetryBatch/IngestAck same
ReconcileSite ReconcileSiteRequest/Response (Deployment/ReconcileSiteRequest.cs, ReconcileSiteResponse.cs incl. ReconcileGapItem) map<string,string> for name→hash
ReportSiteHealth SiteHealthReport/SiteHealthReportAck (Health/SiteHealthReport.cs) the big one: ~30 fields incl. maps of ConnectionHealth enum, TagResolutionStatus, TagQualityCounts, NodeStatus list, SiteAuditBacklogSnapshot; model nullable ints/doubles with wrappers; keep SequenceNumber
Heartbeat HeartbeatMessage (Health/HeartbeatMessage.cs) → google.protobuf.Empty reply fire-and-forget semantics preserved client-side (don't await failure into caller)

Mappers in Communication/Grpc/CentralControlDtoMapper.cs with round-trip golden tests (construct DTO → proto → DTO, assert deep-equal; include null-optional cases).

T1A.2 Central hosting

Central branch of Program.cs (~:262 services, :449-469 Map block): builder.Services.AddGrpc(o => o.Interceptors.Add<ControlPlaneAuthInterceptor>()) — central's interceptor variant verifies Bearer against the per-site PSK looked up by the required x-scadabridge-site metadata header via SitePskProvider (fail-closed: missing header ⇒ PermissionDenied). Add an explicit Kestrel listener for h2c gRPC: new option ScadaBridge:Node:CentralGrpcPort (default 8083, symmetric with sites), configured like the site branch does (Program.cs:507-512, HttpProtocols.Http2); central's :5000 stays as-is (Traefik is HTTP/1 — gRPC does NOT go through traefik). MapGrpcService<CentralControlGrpcService>().

CentralControlGrpcService (Host or Communication): decode proto → the SAME message types → Ask the existing CentralCommunicationActor handlers (they already handle all 7 — zero handler logic changes) → encode reply. Reuse the readiness convention: reject Unavailable until the central actor system is up (mirror SiteStreamGrpcServer.SetReady).

T1A.3 Site-side client + transport seam

  • New ICentralTransport (Communication): one method per the 7 sends. Implementations: AkkaCentralTransport (extracted verbatim from today's SiteCommunicationActor send blocks) and GrpcCentralTransport (new).
  • GrpcCentralTransport: channel pair with sticky failover + failback per design §3.5 — new shared CentralChannelProvider: endpoints from new option ScadaBridge:Communication:CentralGrpcEndpoints (List, e.g. ["http://scadabridge-central-a:8083","http://scadabridge-central-b:8083"]; validator: required when transport=Grpc); sticky-until-failure; flip on connect-fail/Unavailable; background failback probe every 3060 s (gRPC health or a Heartbeat ping); reconnect backoff copied from SyncBackgroundService.cs:151 (1 s doubling, cap 60 s). Attach PSK (GrpcPsk option) + x-scadabridge-site on every call; per-call deadlines from the matching CommunicationOptions timeout (NotificationForwardTimeout, HealthReportTimeout, etc. — today's Ask timeouts, unchanged values). Cross-node auto-retry only on connect-fail/Unavailable — never on DeadlineExceeded.
  • SiteCommunicationActor selects the implementation from new option ScadaBridge:Communication:CentralTransport (Akka | Grpc, default Akka). The 7 handler bodies delegate to the injected transport; reply/fault semantics identical (timeout or non-OK status ⇒ same Status.Failure the S&F/audit layers already treat as transient).

T1A.4 Tests

Extend Communication.Tests (TestKit): SiteCommunicationActor with an NSubstitute ICentralTransport — all 7 paths, fault propagation (transport throw ⇒ same failure the S&F tests expect). GrpcCentralTransport unit tests with an in-process TestServer gRPC host: failover flip, sticky behavior, failback probe, PSK attached, deadline set, no-retry-on-deadline. Reuse/extend DirectActorSiteStreamAuditClient for the ingest integration harness. NotificationForwarderTests/SiteAuditTelemetryActorTests/HealthReportSenderTests must pass unmodified (they sit above the seam — if they need edits, the seam is wrong).

1A DoD: suite green; on the rig with site-a flipped to CentralTransport=Grpc (central gRPC port published, e.g. 9013/9014:8083): notification e2e, audit rows land, health page live, heartbeat drives active flag, reconcile works after site restart — while site-b/c still run Akka (coexistence proven).


Phase 1B — Site command plane (central→site) — worktree B

T1B.1 site_command.proto

Package scadabridge.sitecommand.v1, service SiteCommandService28 commands (29 minus dead IntegrationCallRequest), grouped into domain RPCs with oneof request/response envelopes (full command list + reply types + CommunicationService.cs line refs in the design doc §7 / recon inventory):

RPC Commands (count) Deadline source
ExecuteLifecycle RefreshDeployment, Enable/Disable/DeleteInstance, DeploymentStateQuery, DeployArtifacts (6) DeploymentTimeout/LifecycleTimeout/ArtifactDeploymentTimeout
ExecuteOpcUa BrowseNode, SearchAddressSpace, ReadTagValues, VerifyEndpoint, Trust/List/RemoveServerCert, WriteTag (8) QueryTimeout (browse/search per existing Ask usage)
ExecuteQuery EventLogQuery, DebugSnapshot, Subscribe/UnsubscribeDebugView (4) QueryTimeout/DebugViewTimeout
ExecuteParked ParkedMessageQuery/Retry/Discard, RetryParkedOperation, DiscardParkedOperation (5) QueryTimeout; relay callers keep RelayTimeout(10s) < QueryTimeout(30s) ordering
ExecuteRoute RouteToCall/GetAttributes/SetAttributes/WaitForAttribute (4) IntegrationTimeout; WaitForAttribute uses its dynamic timeout
TriggerFailover TriggerSiteFailover (1) LifecycleTimeout

Mappers SiteCommandDtoMapper.cs + round-trip golden tests for every command/reply (the bulk of this track — budget accordingly; enums, TrackedOperationId struct → string guid, nullable wrappers).

T1B.2 Site server: shared dispatcher

Refactor SiteCommunicationActor's receive table into SiteCommandDispatcher (pure routing: message → _deploymentManagerProxy / _artifactHandler / _eventLogHandler / _parkedMessageHandler / failover handler — preserving EXACTLY today's targets, including the node-local parked-message handler; see design §7.3, the replicated-store semantics are deliberate). The actor and a new SiteCommandGrpcService (mapped in the site branch next to SiteStreamGrpcServer, gated by ControlPlaneAuthInterceptor + the readiness flag) both call the dispatcher — one routing truth, both transports.

T1B.3 Central client + transport seam

ISiteCommandTransport (send SiteEnvelope-equivalent, Ask or Tell) injected into CentralCommunicationActor; implementations AkkaSiteTransport (today's per-site ClusterClient path, extracted) and GrpcSiteTransport. GrpcSiteTransport uses a new shared SitePairChannelProvider: addresses from the Site entity's existing GrpcNodeAAddress/GrpcNodeBAddress (the streaming path's columns — do NOT invent LoadSiteAddressesFromDb, it doesn't exist; reuse the ISiteRepository reads + CentralCommunicationActor's existing DB-driven cache-refresh loop :532-598 to build/refresh channels instead of ClusterClients), sticky failover/failback per §3.5, PSK from SitePskProvider + deadlines per the table above. Config flag ScadaBridge:Communication:SiteTransport (Akka | Grpc, default Akka) selected inside CentralCommunicationActorCommunicationService and SiteCallAuditActor unchanged.

T1B.4 Tests

TestKit: CentralCommunicationActor with substitute ISiteCommandTransport (envelope routing, Ask-sender reply plumbing, per-site transport lifecycle on site add/remove/change); SiteCommandDispatcher unit tests (every command → correct target, incl. parked→local handler and failover→local); SiteCommandGrpcService via TestServer (auth, readiness, one command per oneof group); existing CommunicationServiceTests/CentralCommunicationActor*Tests pass with the Akka implementation as default.

1B DoD: suite green; rig central flipped to SiteTransport=Grpc for site-a only: from CentralUI — deploy refresh, enable/disable instance, browse, read tag, write tag, event-log query, parked query/retry/discard (run retry against the STANDBY site node explicitly — replicated-store semantics), TriggerSiteFailover — all work; site-b/c untouched on Akka.


Phase 2 (after 1A) ∥ Phase 3 (after 1B) — full cutover on the rig + hardening

  • Phase 2: flip all sites to CentralTransport=Grpc. Soak: S&F drain under central-a kill (rows stay Pending, resume without loss or duplicates — sequence/dedup layers unchanged), failback observed when central-a returns, health/heartbeat cadence unchanged in CentralHealthAggregator (no sequence regressions logged).
  • Phase 3: flip central to SiteTransport=Grpc for all sites. Soak: full CentralUI command matrix against each site; site-node kill mid-command returns a clean error (no hang beyond deadline); site pair failover mid-stream of commands.
  • Both phases: watch for PermissionDenied noise (would indicate PSK drift), and confirm zero ClusterClient log activity on flipped paths.

Phase 4 — deletion + config cutover (sequential, after 2+3)

  1. Flip both flag defaults to Grpc; rig + docs updated; one soak cycle.
  2. Delete: AkkaCentralTransport/AkkaSiteTransport, ClusterClient creation (AkkaHostedService.cs:942-953), DefaultSiteClientFactory (+ its tests), per-site ClusterClient cache in CentralCommunicationActor (keep the DB refresh loop — it now feeds SitePairChannelProvider), receptionist registrations :436 and :935 (T0.1 already removed :459), CommunicationOptions.CentralContactPoints (+ validator + rig configs), then the flags themselves.
  3. Grep-gates: rg -i "clusterclient|receptionist" src tests docker docs → only historical docs; rg "CentralContactPoints" → empty.
  4. Docs: update grpc_streams.md (its "ClusterClient keeps command/control" split is superseded; also fix its LoadSiteAddressesFromDb doc-vs-code gap), Component-Host.md, Component-StoreAndForward.md:137, and add a docs/known-issues cross-ref note that the frame-size class is retired. Keep Akka.Cluster.Tools (ClusterSingleton still used).

Phase 5 — live gate (rig; sequential; every check PASS required)

  1. PSK negative: no key / wrong key / missing x-scadabridge-sitePermissionDenied; unset server key ⇒ all rejected (fail-closed); LocalDb sync key unaffected.
  2. Site→central matrix: notification e2e (delivered + status query), audit + cached-telemetry rows in dbo.AuditLog/site-calls, /monitoring/health live per site, heartbeat → active flag, reconcile self-heal after site-node restart.
  3. Central→site matrix: all 6 RPC groups exercised from CentralUI, parked retry/discard on the standby node, failover command drains cleanly.
  4. Failover/failback: kill central-a → sites flip to central-b sticky (S&F uninterrupted); restart central-a → failback within probe cadence; same for a site node from central's side; booting node rejects Unavailable until ready and the client fails over.
  5. Mid-drain kill: kill central during an S&F drain burst — zero loss, zero duplicates.
  6. Frame-class retirement: issue a command/reply > 128 KB (large browse/event-log result) — succeeds over gRPC (impossible before).
  7. Boundary check: with everything on gRPC, verify NO Akka association exists between any site container and central (netstat/Akka logs) — remoting is pair-internal only.
  8. Restart discipline: full-rig restart (pairs together) comes up clean; no receptionist/ ClusterClient log lines anywhere.

Record results in docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md (check-by-check, the family's live-gate format).

Effort & parallelization summary

Track Est. Parallel with
Phase 0 23 d 1A/1B proto authoring
1A 1.52 wk 1B (separate worktrees; 1A merges first)
1B 23 wk (mapper-heavy) 1A
2, 3 24 d each each other (independent flags/paths)
4 23 d
5 23 d

Critical path ≈ 1B: ~46 weeks total, matching the design estimate.

Task checklist (tick as you go; IDs reference the sections above)

Phase 0 — PSK + dead code (branch feat/grpc-phase0-psk)

  • T0.1 Delete /user/management receptionist registration (AkkaHostedService.cs:459) + Component-Host.md update
  • T0.2 File dead-IntegrationCallRequest issue; record exclusion (28 of 29 migrate)
  • T0.3 ControlPlaneAuthInterceptor + CommunicationOptions.GrpcPsk + SitePskProvider; gate SiteStream; PSK attached on central's streaming + pull clients
  • T0.4 Rig dev keys (all 3 sites + central store seeds) + interceptor/provider/wiring tests
  • Phase 0 DoD: suite green; rig unauthenticated ⇒ PermissionDenied, authenticated paths work; PR merged

Phase 1A — central control plane (worktree, feat/grpc-central-control)

  • T1A.1 central_control.proto (7 RPCs; checked-in codegen) + CentralControlDtoMapper + round-trip golden tests
  • T1A.2 Central hosting: AddGrpc + per-site-PSK interceptor (x-scadabridge-site), CentralGrpcPort h2c listener (8083), CentralControlGrpcService (Ask existing handlers), readiness gate
  • T1A.3 ICentralTransport (Akka extract + Grpc impl), CentralChannelProvider (sticky failover/failback, backoff, deadlines, PSK), CentralTransport flag default Akka, CentralGrpcEndpoints option + validator
  • T1A.4 Tests: actor-with-fake-transport ×7, TestServer transport tests, S&F/audit/health suites pass unmodified
  • 1A DoD: rig site-a on Grpc proves all 5 site→central paths while site-b/c stay Akka; PR merged (before 1B)

Phase 1B — site command plane (worktree, feat/grpc-sitecommand)

  • T1B.1 site_command.proto (6 oneof RPCs / 28 commands) + SiteCommandDtoMapper + round-trip golden tests (all 28 + replies)
  • T1B.2 SiteCommandDispatcher refactor (actor + new SiteCommandGrpcService share it; parked stays node-local)
  • T1B.3 ISiteCommandTransport in CentralCommunicationActor (Akka extract + Grpc impl), SitePairChannelProvider (Site entity Grpc columns + DB refresh loop), SiteTransport flag default Akka
  • T1B.4 Tests: dispatcher routing ×28, actor envelope/reply plumbing, TestServer service tests, existing Communication suites green
  • 1B DoD: rig central on Grpc for site-a proves full command matrix incl. standby parked retry; rebased on 1A; PR merged

Phase 2 ∥ 3 — cutover + soak

  • P2 All sites CentralTransport=Grpc; central-kill S&F soak (no loss/dupes), failback observed, health sequences clean
  • P3 Central SiteTransport=Grpc all sites; full UI command matrix per site; site-kill mid-command clean; no PSK noise; zero ClusterClient log activity on flipped paths

Phase 4 — deletion

  • Defaults flip to Grpc + soak; then delete Akka transports, ClusterClient creation, DefaultSiteClientFactory, receptionist registrations, CentralContactPoints, then the flags
  • Grep-gates pass (clusterclient|receptionist → historical docs only; CentralContactPoints → empty)
  • Docs updated: grpc_streams.md, Component-Host.md, Component-StoreAndForward.md:137, known-issues cross-ref

Phase 5 — live gate (record in 2026-07-22-clusterclient-to-grpc-live-gate.md)

  • 1 PSK negatives · [ ] 2 site→central matrix · [ ] 3 central→site matrix · [ ] 4 failover/failback both directions · [ ] 5 mid-drain kill · [ ] 6 >128 KB frame-class proof · [ ] 7 no cross-boundary Akka association · [ ] 8 full-rig restart clean

Gotchas for the executor (will bite; read twice)

  • Generated proto C# is committed; never add an active <Protobuf> item to the csproj in a final commit (linux_arm64 protoc segfault breaks the Docker build).
  • HeartbeatMessage must stay fire-and-forget end-to-end — don't let a gRPC failure surface as a fault to the heartbeat timer path.
  • Deadline ≠ retry: no automatic cross-node retry on DeadlineExceeded for WriteTag/Deploy/ Failover; only on provably-unsent failures.
  • The parked-message handler is node-local on purpose (replicated store); do not "fix" it onto the singleton proxy.
  • Ack-before-Leave on TriggerSiteFailover (SiteCommunicationActor.cs:569): the gRPC reply must be written before the node leaves — verify the response completes under failover.
  • Inner-before-outer timeouts: RelayTimeout(10 s) < QueryTimeout(30 s) must survive the deadline mapping (CommunicationService.cs:779-786 documents why).
  • SiteStreamGrpcServer.AuditIngestAskTimeout (30 s) is "one source of truth" shared with CentralCommunicationActor — keep the new central service on the same constant.
  • Rig sites reach central by container name (scadabridge-central-{a,b}:8083), NOT via traefik (HTTP/1 only).
  • xunit v2: use Xunit.SkippableFact for env-gated tests, not Assert.Skip.