From f1ad96708316eb508cf634cdbbfeaf96eae50abe Mon Sep 17 00:00:00 2001 From: Joseph Doherty Date: Wed, 22 Jul 2026 17:10:33 -0400 Subject: [PATCH] =?UTF-8?q?docs(plans):=20ClusterClient=E2=86=92gRPC-only?= =?UTF-8?q?=20migration=20plan=20(phases=200=E2=80=935=20+=20tasklist)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Complete executable plan: PSK-from-Secrets auth (Phase 0), central_control + site_command proto contracts behind transport seams inside the two communication actors (1A ∥ 1B in worktrees), per-direction cutover flags, ClusterClient/receptionist deletion, 8-check live gate. Design doc: scadaproj/scadabridge_clusterclient_to_grpc.md (§7 = deep-dive corrections). --- .../2026-07-22-clusterclient-to-grpc-plan.md | 355 ++++++++++++++++++ 1 file changed, 355 insertions(+) create mode 100644 docs/plans/2026-07-22-clusterclient-to-grpc-plan.md diff --git a/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md b/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md new file mode 100644 index 00000000..d91322fb --- /dev/null +++ b/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md @@ -0,0 +1,355 @@ +# ClusterClient → gRPC-only cross-cluster transport — implementation plan + +**Date:** 2026-07-22. **Design:** `~/Desktop/scadaproj/scadabridge_clusterclient_to_grpc.md` +(read it first, §7 especially — it contains the deep-dive corrections this plan builds on). +**Goal:** all site↔central traffic rides gRPC with PSK auth from `ZB.MOM.WW.Secrets`; +`ClusterClient`/`ClusterClientReceptionist` are deleted; Akka remoting never crosses the +site↔central boundary. + +Everything below was code-verified 2026-07-22. If a cited line has drifted, re-locate by the +quoted identifier, never by line number. + +## How to execute this plan (read before starting) + +- **Phase order:** 0 → (1A ∥ 1B) → (2 ∥ 3) → 4 → 5. Phases marked ∥ are parallel-safe **only in + separate git worktrees** (`git worktree add ../ScadaBridge-1B feat/grpc-sitecommand`) — never + run two agents against one working tree (destructive git races are a known family incident). + Designated merge order when tracks meet: 1A lands first, 1B rebases (expected conflicts are + confined to `ZB.MOM.WW.ScadaBridge.Communication.csproj` proto ItemGroups and `Program.cs` + service/Map blocks — resolve as union). +- **Branches:** `feat/grpc-phase0-psk`, `feat/grpc-central-control` (1A), + `feat/grpc-sitecommand` (1B), then per-phase branches; PR each phase to `main` with its DoD met. +- **Build/test:** `dotnet build ZB.MOM.WW.ScadaBridge.slnx` (0 warnings — TreatWarningsAsErrors), + `dotnet test ZB.MOM.WW.ScadaBridge.slnx`. Tests are xunit **v2** (2.9.3) + `Akka.TestKit.Xunit2` + 1.5.62 + NSubstitute. TestKit tests inherit `TestKit` — never hand-roll + `ActorSystem.Create`+join. +- **Proto codegen is CHECKED-IN, not build-time** (protoc segfaults in the linux_arm64 image). + For every new/changed proto follow the sitestream recipe documented in + `src/ZB.MOM.WW.ScadaBridge.Communication/ZB.MOM.WW.ScadaBridge.Communication.csproj`: + temporarily uncomment/add the `` item, delete stale + generated files, `dotnet build` on macOS, copy `obj/**/Protos/*.cs` into a committed folder + (mirror `SiteStreamGrpc/`), re-comment the item. `docker/regen-proto.sh` exists. +- **Rig:** `docker/deploy.sh` rebuilds the 2-central + 3×2-site cluster (+ traefik). Central UI + `9001/9002:5000`, site gRPC `90x3/90x4:8083`. **External sites↔central rules:** never + host-`sqlite3` a live WAL DB; `aspnet:10.0` has no `curl`; seeding via `docker/seed-sites.sh`. +- **Coexistence rule:** every migrated path sits behind a config flag with the Akka + implementation as default until Phase 4. Rollback at any point = flip the flag. + +## The two choke points (where all transport code changes) + +- **Site→central:** `SiteCommunicationActor` + (`src/ZB.MOM.WW.ScadaBridge.Communication/Actors/SiteCommunicationActor.cs`) — 7 message types, + all `ClusterClient.Send` to `/user/central-communication` (Ask, except `HeartbeatMessage` Tell). +- **Central→site:** `CentralCommunicationActor` (same dir) — the `SiteEnvelope` handler + (`:452-467`) unwraps and routes via per-site ClusterClient. ALL producers + (`CommunicationService`'s 27 + `SiteCallAuditActor`'s 2 relays) go through it. + **`CommunicationService` gets NO interface extraction; its ~20 consumers are untouched.** + +--- + +## Phase 0 — PSK auth + dead-code removal (standalone hardening; parallel-safe with 1A/1B prep) + +### T0.1 Delete the vestigial `/user/management` receptionist registration +`AkkaHostedService.cs:459` (`RegisterService` of `ManagementActor`) — delete the registration +only (the actor stays; `ManagementEndpoints.cs:117` asks it in-process via +`ManagementActorHolder`). Grep-verify nothing sends to `/user/management`. Update +`docs/requirements/Component-Host.md` REQ-HOST-6a to record the removal. + +### T0.2 File the dead-integration issue +`IntegrationCallRequest` (`CommunicationService.cs:238`) is unwired in production — +`RegisterLocalHandler(Integration, …)` exists only in `SiteCommunicationActorTests.cs:102`, so +production always replies "Integration handler not available" (`SiteCommunicationActor.cs:134-136`). +File a Gitea issue (decide: wire or delete); **exclude it from the gRPC contract** (28 of 29 +commands migrate). Do not change its behavior in this program. + +### T0.3 PSK interceptor + options +New `src/ZB.MOM.WW.ScadaBridge.Host/ControlPlaneAuthInterceptor.cs`, copied from +`LocalDbSyncAuthInterceptor.cs` (same file layout: seal, 4 server handlers → shared +`Authorize`, `authorization: Bearer` metadata extraction, `CryptographicOperations.FixedTimeEquals` +over UTF-8, **fail-closed when the expected key is unset**, reject with +`StatusCode.PermissionDenied`). Differences from the template: + +- **Path scope:** gates a *set* of service prefixes (constructor-provided), initially the + sitestream service prefix (read the real package/service name from `sitestream.proto` — + `/…SiteStreamService/`); later phases add the new services. LocalDb sync keeps its own + interceptor + key untouched. +- **Key source, site side (server on `:8083`):** new option + `CommunicationOptions.GrpcPsk` (`ScadaBridge:Communication:GrpcPsk`), production value + `${secret:SB-GRPC-PSK-}` resolved by the existing pre-host `SecretReferenceExpander` + (`Program.cs:56-65`) — zero new resolution code. +- **Key source, central side (clients today; server in 1A):** central's site set is dynamic, so + central resolves secret **`SB-GRPC-PSK-{siteId}`** at channel-build time via the runtime + `ISecretResolver` (copy the fail-closed lazy pattern from + `DataConnectionLayer/Adapters/MxGatewayDataConnection.cs:82-98`), cached per site, cache + invalidated on site remove. New helper `SitePskProvider` in Communication (interface) + + Host (implementation over `ISecretResolver`). + +Wire: add the interceptor to the site's existing `AddGrpc` (`Program.cs:526-527`, alongside the +LocalDb one). Attach the PSK on central's existing site-dialing clients as call-level +`Metadata` `Authorization: Bearer ` (the LocalDb sync client pattern, +`SyncBackgroundService.cs:82-86`): `SiteStreamGrpcClient` (subscribe calls), +`GrpcPullAuditEventsClient`, `GrpcPullSiteCallsClient` — each already flows through a channel/ +invoker creation point where the site id is known. + +### T0.4 Rig + tests +- Rig: mirror the LocalDb dev-key pattern — literal `ScadaBridge:Communication:GrpcPsk: + "dev-grpc-psk-site-a"` in BOTH `docker/site-a-node-*/appsettings.Site.json` (and site-b/c with + their own keys, since unlike LocalDb this is not optional), plus central-side dev secrets: seed + `SB-GRPC-PSK-site-{a,b,c}` into both centrals' secret stores (secret CLI seed flow, or dev-KEK + env as the rig's Secrets setup already does). +- Tests: interceptor unit tests (wrong key / missing header / unset expected key ⇒ + `PermissionDenied`; non-gated service path passes; constant-time compare exercised); + `SitePskProvider` fail-closed test; one Host wiring test asserting the interceptor is + registered. + +**Phase 0 DoD:** suite green; on the rig, an unauthenticated `grpcurl`/test client gets +`PermissionDenied` on SiteStream, authenticated streaming + pull still work end-to-end. PR merged. + +--- + +## Phase 1A — Central control plane (site→central) — worktree A + +### T1A.1 `central_control.proto` +New `Protos/central_control.proto` in the Communication project (checked-in codegen per recipe), +package `scadabridge.centralcontrol.v1`, service `CentralControlService`: + +| RPC | Wraps DTO (source file under `Commons/Messages/`) | Notes | +|---|---|---| +| `SubmitNotification` | `NotificationSubmit`/`NotificationSubmitAck` (`Notification/NotificationMessages.cs:30,47`) | 11 fields incl. `Guid?` execution ids → string | +| `QueryNotificationStatus` | `NotificationStatusQuery`/`Response` (`:55,62`) | | +| `IngestAuditEvents` | reuse existing `AuditEventBatch`/`IngestAck` from sitestream.proto | import, don't duplicate; `ForwardState`/`IngestedAtUtc` stay off-wire (`AuditEventDtoMapper.cs:23-27`) | +| `IngestCachedTelemetry` | reuse `CachedTelemetryBatch`/`IngestAck` | same | +| `ReconcileSite` | `ReconcileSiteRequest`/`Response` (`Deployment/ReconcileSiteRequest.cs`, `ReconcileSiteResponse.cs` incl. `ReconcileGapItem`) | map for name→hash | +| `ReportSiteHealth` | `SiteHealthReport`/`SiteHealthReportAck` (`Health/SiteHealthReport.cs`) | the big one: ~30 fields incl. maps of `ConnectionHealth` enum, `TagResolutionStatus`, `TagQualityCounts`, `NodeStatus` list, `SiteAuditBacklogSnapshot`; model nullable ints/doubles with wrappers; keep `SequenceNumber` | +| `Heartbeat` | `HeartbeatMessage` (`Health/HeartbeatMessage.cs`) → `google.protobuf.Empty` reply | fire-and-forget semantics preserved client-side (don't await failure into caller) | + +Mappers in `Communication/Grpc/CentralControlDtoMapper.cs` with **round-trip golden tests** +(construct DTO → proto → DTO, assert deep-equal; include null-optional cases). + +### T1A.2 Central hosting +Central branch of `Program.cs` (~`:262` services, `:449-469` Map block): +`builder.Services.AddGrpc(o => o.Interceptors.Add())` — central's +interceptor variant verifies `Bearer` against the per-site PSK looked up by the **required +`x-scadabridge-site` metadata header** via `SitePskProvider` (fail-closed: missing header ⇒ +`PermissionDenied`). Add an **explicit Kestrel listener** for h2c gRPC: new option +`ScadaBridge:Node:CentralGrpcPort` (default **8083**, symmetric with sites), configured like the +site branch does (`Program.cs:507-512`, `HttpProtocols.Http2`); central's `:5000` stays as-is +(Traefik is HTTP/1 — gRPC does NOT go through traefik). `MapGrpcService()`. + +`CentralControlGrpcService` (Host or Communication): decode proto → the SAME message types → +`Ask` the existing `CentralCommunicationActor` handlers (they already handle all 7 — zero +handler logic changes) → encode reply. Reuse the readiness convention: reject `Unavailable` +until the central actor system is up (mirror `SiteStreamGrpcServer.SetReady`). + +### T1A.3 Site-side client + transport seam +- New `ICentralTransport` (Communication): one method per the 7 sends. Implementations: + `AkkaCentralTransport` (extracted verbatim from today's `SiteCommunicationActor` send blocks) + and `GrpcCentralTransport` (new). +- `GrpcCentralTransport`: **channel pair with sticky failover + failback** per design §3.5 — new + shared `CentralChannelProvider`: endpoints from new option + `ScadaBridge:Communication:CentralGrpcEndpoints` (List, e.g. + `["http://scadabridge-central-a:8083","http://scadabridge-central-b:8083"]`; validator: required + when transport=Grpc); sticky-until-failure; flip on connect-fail/`Unavailable`; background + failback probe every 30–60 s (gRPC health or a `Heartbeat` ping); reconnect backoff copied from + `SyncBackgroundService.cs:151` (1 s doubling, cap 60 s). Attach PSK (`GrpcPsk` option) + + `x-scadabridge-site` on every call; per-call deadlines from the matching `CommunicationOptions` + timeout (`NotificationForwardTimeout`, `HealthReportTimeout`, etc. — today's Ask timeouts, + unchanged values). **Cross-node auto-retry only on connect-fail/`Unavailable`** — never on + `DeadlineExceeded`. +- `SiteCommunicationActor` selects the implementation from new option + `ScadaBridge:Communication:CentralTransport` (`Akka` | `Grpc`, **default `Akka`**). The 7 + handler bodies delegate to the injected transport; reply/fault semantics identical (timeout or + non-OK status ⇒ same `Status.Failure` the S&F/audit layers already treat as transient). + +### T1A.4 Tests +Extend `Communication.Tests` (TestKit): `SiteCommunicationActor` with an NSubstitute +`ICentralTransport` — all 7 paths, fault propagation (transport throw ⇒ same failure the S&F +tests expect). `GrpcCentralTransport` unit tests with an in-process `TestServer` gRPC host: +failover flip, sticky behavior, failback probe, PSK attached, deadline set, no-retry-on-deadline. +Reuse/extend `DirectActorSiteStreamAuditClient` for the ingest integration harness. +`NotificationForwarderTests`/`SiteAuditTelemetryActorTests`/`HealthReportSenderTests` must pass +unmodified (they sit above the seam — if they need edits, the seam is wrong). + +**1A DoD:** suite green; on the rig with site-a flipped to `CentralTransport=Grpc` +(central gRPC port published, e.g. `9013/9014:8083`): notification e2e, audit rows land, health +page live, heartbeat drives active flag, reconcile works after site restart — while site-b/c +still run Akka (coexistence proven). + +--- + +## Phase 1B — Site command plane (central→site) — worktree B + +### T1B.1 `site_command.proto` +Package `scadabridge.sitecommand.v1`, service `SiteCommandService` — **28 commands** (29 minus +dead `IntegrationCallRequest`), grouped into domain RPCs with `oneof` request/response +envelopes (full command list + reply types + `CommunicationService.cs` line refs in the design +doc §7 / recon inventory): + +| RPC | Commands (count) | Deadline source | +|---|---|---| +| `ExecuteLifecycle` | RefreshDeployment, Enable/Disable/DeleteInstance, DeploymentStateQuery, DeployArtifacts (6) | `DeploymentTimeout`/`LifecycleTimeout`/`ArtifactDeploymentTimeout` | +| `ExecuteOpcUa` | BrowseNode, SearchAddressSpace, ReadTagValues, VerifyEndpoint, Trust/List/RemoveServerCert, WriteTag (8) | `QueryTimeout` (browse/search per existing Ask usage) | +| `ExecuteQuery` | EventLogQuery, DebugSnapshot, Subscribe/UnsubscribeDebugView (4) | `QueryTimeout`/`DebugViewTimeout` | +| `ExecuteParked` | ParkedMessageQuery/Retry/Discard, RetryParkedOperation, DiscardParkedOperation (5) | `QueryTimeout`; relay callers keep `RelayTimeout`(10s) < `QueryTimeout`(30s) ordering | +| `ExecuteRoute` | RouteToCall/GetAttributes/SetAttributes/WaitForAttribute (4) | `IntegrationTimeout`; WaitForAttribute uses its dynamic timeout | +| `TriggerFailover` | TriggerSiteFailover (1) | `LifecycleTimeout` | + +Mappers `SiteCommandDtoMapper.cs` + round-trip golden tests for every command/reply (the bulk of +this track — budget accordingly; enums, `TrackedOperationId` struct → string guid, nullable +wrappers). + +### T1B.2 Site server: shared dispatcher +Refactor `SiteCommunicationActor`'s receive table into `SiteCommandDispatcher` (pure routing: +message → `_deploymentManagerProxy` / `_artifactHandler` / `_eventLogHandler` / +`_parkedMessageHandler` / failover handler — preserving EXACTLY today's targets, including the +node-local parked-message handler; see design §7.3, the replicated-store semantics are +deliberate). The actor and a new `SiteCommandGrpcService` (mapped in the site branch next to +`SiteStreamGrpcServer`, gated by `ControlPlaneAuthInterceptor` + the readiness flag) both call +the dispatcher — one routing truth, both transports. + +### T1B.3 Central client + transport seam +`ISiteCommandTransport` (send `SiteEnvelope`-equivalent, Ask or Tell) injected into +**`CentralCommunicationActor`**; implementations `AkkaSiteTransport` (today's per-site +ClusterClient path, extracted) and `GrpcSiteTransport`. `GrpcSiteTransport` uses a new shared +**`SitePairChannelProvider`**: addresses from the `Site` entity's existing +`GrpcNodeAAddress`/`GrpcNodeBAddress` (the streaming path's columns — do NOT invent +`LoadSiteAddressesFromDb`, it doesn't exist; reuse the `ISiteRepository` reads + +`CentralCommunicationActor`'s existing DB-driven cache-refresh loop `:532-598` to build/refresh +channels instead of ClusterClients), sticky failover/failback per §3.5, PSK from +`SitePskProvider` + deadlines per the table above. Config flag +`ScadaBridge:Communication:SiteTransport` (`Akka` | `Grpc`, default `Akka`) selected inside +`CentralCommunicationActor` — `CommunicationService` and `SiteCallAuditActor` unchanged. + +### T1B.4 Tests +TestKit: `CentralCommunicationActor` with substitute `ISiteCommandTransport` (envelope routing, +Ask-sender reply plumbing, per-site transport lifecycle on site add/remove/change); +`SiteCommandDispatcher` unit tests (every command → correct target, incl. parked→local handler +and failover→local); `SiteCommandGrpcService` via TestServer (auth, readiness, one command per +oneof group); existing `CommunicationServiceTests`/`CentralCommunicationActor*Tests` pass with +the Akka implementation as default. + +**1B DoD:** suite green; rig central flipped to `SiteTransport=Grpc` for site-a only: from +CentralUI — deploy refresh, enable/disable instance, browse, read tag, write tag, event-log +query, parked query/retry/discard (run retry against the STANDBY site node explicitly — +replicated-store semantics), `TriggerSiteFailover` — all work; site-b/c untouched on Akka. + +--- + +## Phase 2 (after 1A) ∥ Phase 3 (after 1B) — full cutover on the rig + hardening + +- **Phase 2:** flip all sites to `CentralTransport=Grpc`. Soak: S&F drain under central-a kill + (rows stay Pending, resume without loss or duplicates — sequence/dedup layers unchanged), + failback observed when central-a returns, health/heartbeat cadence unchanged in + `CentralHealthAggregator` (no sequence regressions logged). +- **Phase 3:** flip central to `SiteTransport=Grpc` for all sites. Soak: full CentralUI command + matrix against each site; site-node kill mid-command returns a clean error (no hang beyond + deadline); site pair failover mid-stream of commands. +- Both phases: watch for `PermissionDenied` noise (would indicate PSK drift), and confirm zero + ClusterClient log activity on flipped paths. + +## Phase 4 — deletion + config cutover (sequential, after 2+3) + +1. Flip both flag defaults to `Grpc`; rig + docs updated; one soak cycle. +2. Delete: `AkkaCentralTransport`/`AkkaSiteTransport`, ClusterClient creation + (`AkkaHostedService.cs:942-953`), `DefaultSiteClientFactory` (+ its tests), per-site + ClusterClient cache in `CentralCommunicationActor` (keep the DB refresh loop — it now feeds + `SitePairChannelProvider`), receptionist registrations `:436` and `:935` (T0.1 already removed + `:459`), `CommunicationOptions.CentralContactPoints` (+ validator + rig configs), then the + flags themselves. +3. Grep-gates: `rg -i "clusterclient|receptionist" src tests docker docs` → only historical docs; + `rg "CentralContactPoints"` → empty. +4. Docs: update `grpc_streams.md` (its "ClusterClient keeps command/control" split is + superseded; also fix its `LoadSiteAddressesFromDb` doc-vs-code gap), `Component-Host.md`, + `Component-StoreAndForward.md:137`, and add a `docs/known-issues` cross-ref note that the + frame-size class is retired. Keep `Akka.Cluster.Tools` (ClusterSingleton still used). + +## Phase 5 — live gate (rig; sequential; every check PASS required) + +1. **PSK negative:** no key / wrong key / missing `x-scadabridge-site` ⇒ `PermissionDenied`; + unset server key ⇒ all rejected (fail-closed); LocalDb sync key unaffected. +2. **Site→central matrix:** notification e2e (delivered + status query), audit + cached-telemetry + rows in `dbo.AuditLog`/site-calls, `/monitoring/health` live per site, heartbeat → active + flag, reconcile self-heal after site-node restart. +3. **Central→site matrix:** all 6 RPC groups exercised from CentralUI, parked retry/discard on + the standby node, failover command drains cleanly. +4. **Failover/failback:** kill central-a → sites flip to central-b sticky (S&F uninterrupted); + restart central-a → failback within probe cadence; same for a site node from central's side; + booting node rejects `Unavailable` until ready and the client fails over. +5. **Mid-drain kill:** kill central during an S&F drain burst — zero loss, zero duplicates. +6. **Frame-class retirement:** issue a command/reply > 128 KB (large browse/event-log result) — + succeeds over gRPC (impossible before). +7. **Boundary check:** with everything on gRPC, verify NO Akka association exists between any + site container and central (`netstat`/Akka logs) — remoting is pair-internal only. +8. **Restart discipline:** full-rig restart (pairs together) comes up clean; no receptionist/ + ClusterClient log lines anywhere. + +Record results in `docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md` (check-by-check, +the family's live-gate format). + +## Effort & parallelization summary + +| Track | Est. | Parallel with | +|---|---|---| +| Phase 0 | 2–3 d | 1A/1B proto authoring | +| 1A | 1.5–2 wk | 1B (separate worktrees; 1A merges first) | +| 1B | 2–3 wk (mapper-heavy) | 1A | +| 2, 3 | 2–4 d each | each other (independent flags/paths) | +| 4 | 2–3 d | — | +| 5 | 2–3 d | — | + +Critical path ≈ 1B: **~4–6 weeks total**, matching the design estimate. + +## Task checklist (tick as you go; IDs reference the sections above) + +**Phase 0 — PSK + dead code** (branch `feat/grpc-phase0-psk`) +- [ ] T0.1 Delete `/user/management` receptionist registration (`AkkaHostedService.cs:459`) + Component-Host.md update +- [ ] T0.2 File dead-`IntegrationCallRequest` issue; record exclusion (28 of 29 migrate) +- [ ] T0.3 `ControlPlaneAuthInterceptor` + `CommunicationOptions.GrpcPsk` + `SitePskProvider`; gate SiteStream; PSK attached on central's streaming + pull clients +- [ ] T0.4 Rig dev keys (all 3 sites + central store seeds) + interceptor/provider/wiring tests +- [ ] Phase 0 DoD: suite green; rig unauthenticated ⇒ `PermissionDenied`, authenticated paths work; PR merged + +**Phase 1A — central control plane** (worktree, `feat/grpc-central-control`) +- [ ] T1A.1 `central_control.proto` (7 RPCs; checked-in codegen) + `CentralControlDtoMapper` + round-trip golden tests +- [ ] T1A.2 Central hosting: `AddGrpc` + per-site-PSK interceptor (`x-scadabridge-site`), `CentralGrpcPort` h2c listener (8083), `CentralControlGrpcService` (Ask existing handlers), readiness gate +- [ ] T1A.3 `ICentralTransport` (Akka extract + Grpc impl), `CentralChannelProvider` (sticky failover/failback, backoff, deadlines, PSK), `CentralTransport` flag default `Akka`, `CentralGrpcEndpoints` option + validator +- [ ] T1A.4 Tests: actor-with-fake-transport ×7, TestServer transport tests, S&F/audit/health suites pass unmodified +- [ ] 1A DoD: rig site-a on `Grpc` proves all 5 site→central paths while site-b/c stay Akka; PR merged (before 1B) + +**Phase 1B — site command plane** (worktree, `feat/grpc-sitecommand`) +- [ ] T1B.1 `site_command.proto` (6 oneof RPCs / 28 commands) + `SiteCommandDtoMapper` + round-trip golden tests (all 28 + replies) +- [ ] T1B.2 `SiteCommandDispatcher` refactor (actor + new `SiteCommandGrpcService` share it; parked stays node-local) +- [ ] T1B.3 `ISiteCommandTransport` in `CentralCommunicationActor` (Akka extract + Grpc impl), `SitePairChannelProvider` (Site entity Grpc columns + DB refresh loop), `SiteTransport` flag default `Akka` +- [ ] T1B.4 Tests: dispatcher routing ×28, actor envelope/reply plumbing, TestServer service tests, existing Communication suites green +- [ ] 1B DoD: rig central on `Grpc` for site-a proves full command matrix incl. standby parked retry; rebased on 1A; PR merged + +**Phase 2 ∥ 3 — cutover + soak** +- [ ] P2 All sites `CentralTransport=Grpc`; central-kill S&F soak (no loss/dupes), failback observed, health sequences clean +- [ ] P3 Central `SiteTransport=Grpc` all sites; full UI command matrix per site; site-kill mid-command clean; no PSK noise; zero ClusterClient log activity on flipped paths + +**Phase 4 — deletion** +- [ ] Defaults flip to `Grpc` + soak; then delete Akka transports, ClusterClient creation, `DefaultSiteClientFactory`, receptionist registrations, `CentralContactPoints`, then the flags +- [ ] Grep-gates pass (`clusterclient|receptionist` → historical docs only; `CentralContactPoints` → empty) +- [ ] Docs updated: `grpc_streams.md`, `Component-Host.md`, `Component-StoreAndForward.md:137`, known-issues cross-ref + +**Phase 5 — live gate** (record in `2026-07-22-clusterclient-to-grpc-live-gate.md`) +- [ ] 1 PSK negatives · [ ] 2 site→central matrix · [ ] 3 central→site matrix · [ ] 4 failover/failback both directions · [ ] 5 mid-drain kill · [ ] 6 >128 KB frame-class proof · [ ] 7 no cross-boundary Akka association · [ ] 8 full-rig restart clean + +## Gotchas for the executor (will bite; read twice) + +- Generated proto C# is committed; never add an active `` item to the csproj in a + final commit (linux_arm64 protoc segfault breaks the Docker build). +- `HeartbeatMessage` must stay fire-and-forget end-to-end — don't let a gRPC failure surface as + a fault to the heartbeat timer path. +- Deadline ≠ retry: no automatic cross-node retry on `DeadlineExceeded` for WriteTag/Deploy/ + Failover; only on provably-unsent failures. +- The parked-message handler is node-local **on purpose** (replicated store); do not "fix" it + onto the singleton proxy. +- Ack-before-Leave on `TriggerSiteFailover` (`SiteCommunicationActor.cs:569`): the gRPC reply + must be written before the node leaves — verify the response completes under failover. +- Inner-before-outer timeouts: `RelayTimeout`(10 s) < `QueryTimeout`(30 s) must survive the + deadline mapping (`CommunicationService.cs:779-786` documents why). +- `SiteStreamGrpcServer.AuditIngestAskTimeout` (30 s) is "one source of truth" shared with + `CentralCommunicationActor` — keep the new central service on the same constant. +- Rig sites reach central by container name (`scadabridge-central-{a,b}:8083`), NOT via traefik + (HTTP/1 only). +- xunit v2: use `Xunit.SkippableFact` for env-gated tests, not `Assert.Skip`.