Files
ScadaBridge/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md
T

356 lines
25 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ClusterClient → gRPC-only cross-cluster transport — implementation plan
**Date:** 2026-07-22. **Design:** `~/Desktop/scadaproj/scadabridge_clusterclient_to_grpc.md`
(read it first, §7 especially — it contains the deep-dive corrections this plan builds on).
**Goal:** all site↔central traffic rides gRPC with PSK auth from `ZB.MOM.WW.Secrets`;
`ClusterClient`/`ClusterClientReceptionist` are deleted; Akka remoting never crosses the
site↔central boundary.
Everything below was code-verified 2026-07-22. If a cited line has drifted, re-locate by the
quoted identifier, never by line number.
## How to execute this plan (read before starting)
- **Phase order:** 0 → (1A ∥ 1B) → (2 ∥ 3) → 4 → 5. Phases marked ∥ are parallel-safe **only in
separate git worktrees** (`git worktree add ../ScadaBridge-1B feat/grpc-sitecommand`) — never
run two agents against one working tree (destructive git races are a known family incident).
Designated merge order when tracks meet: 1A lands first, 1B rebases (expected conflicts are
confined to `ZB.MOM.WW.ScadaBridge.Communication.csproj` proto ItemGroups and `Program.cs`
service/Map blocks — resolve as union).
- **Branches:** `feat/grpc-phase0-psk`, `feat/grpc-central-control` (1A),
`feat/grpc-sitecommand` (1B), then per-phase branches; PR each phase to `main` with its DoD met.
- **Build/test:** `dotnet build ZB.MOM.WW.ScadaBridge.slnx` (0 warnings — TreatWarningsAsErrors),
`dotnet test ZB.MOM.WW.ScadaBridge.slnx`. Tests are xunit **v2** (2.9.3) + `Akka.TestKit.Xunit2`
1.5.62 + NSubstitute. TestKit tests inherit `TestKit` — never hand-roll
`ActorSystem.Create`+join.
- **Proto codegen is CHECKED-IN, not build-time** (protoc segfaults in the linux_arm64 image).
For every new/changed proto follow the sitestream recipe documented in
`src/ZB.MOM.WW.ScadaBridge.Communication/ZB.MOM.WW.ScadaBridge.Communication.csproj`:
temporarily uncomment/add the `<Protobuf Include=... GrpcServices="Both" />` item, delete stale
generated files, `dotnet build` on macOS, copy `obj/**/Protos/*.cs` into a committed folder
(mirror `SiteStreamGrpc/`), re-comment the item. `docker/regen-proto.sh` exists.
- **Rig:** `docker/deploy.sh` rebuilds the 2-central + 3×2-site cluster (+ traefik). Central UI
`9001/9002:5000`, site gRPC `90x3/90x4:8083`. **External sites↔central rules:** never
host-`sqlite3` a live WAL DB; `aspnet:10.0` has no `curl`; seeding via `docker/seed-sites.sh`.
- **Coexistence rule:** every migrated path sits behind a config flag with the Akka
implementation as default until Phase 4. Rollback at any point = flip the flag.
## The two choke points (where all transport code changes)
- **Site→central:** `SiteCommunicationActor`
(`src/ZB.MOM.WW.ScadaBridge.Communication/Actors/SiteCommunicationActor.cs`) — 7 message types,
all `ClusterClient.Send` to `/user/central-communication` (Ask, except `HeartbeatMessage` Tell).
- **Central→site:** `CentralCommunicationActor` (same dir) — the `SiteEnvelope` handler
(`:452-467`) unwraps and routes via per-site ClusterClient. ALL producers
(`CommunicationService`'s 27 + `SiteCallAuditActor`'s 2 relays) go through it.
**`CommunicationService` gets NO interface extraction; its ~20 consumers are untouched.**
---
## Phase 0 — PSK auth + dead-code removal (standalone hardening; parallel-safe with 1A/1B prep)
### T0.1 Delete the vestigial `/user/management` receptionist registration
`AkkaHostedService.cs:459` (`RegisterService` of `ManagementActor`) — delete the registration
only (the actor stays; `ManagementEndpoints.cs:117` asks it in-process via
`ManagementActorHolder`). Grep-verify nothing sends to `/user/management`. Update
`docs/requirements/Component-Host.md` REQ-HOST-6a to record the removal.
### T0.2 File the dead-integration issue
`IntegrationCallRequest` (`CommunicationService.cs:238`) is unwired in production —
`RegisterLocalHandler(Integration, …)` exists only in `SiteCommunicationActorTests.cs:102`, so
production always replies "Integration handler not available" (`SiteCommunicationActor.cs:134-136`).
File a Gitea issue (decide: wire or delete); **exclude it from the gRPC contract** (28 of 29
commands migrate). Do not change its behavior in this program.
### T0.3 PSK interceptor + options
New `src/ZB.MOM.WW.ScadaBridge.Host/ControlPlaneAuthInterceptor.cs`, copied from
`LocalDbSyncAuthInterceptor.cs` (same file layout: seal, 4 server handlers → shared
`Authorize`, `authorization: Bearer` metadata extraction, `CryptographicOperations.FixedTimeEquals`
over UTF-8, **fail-closed when the expected key is unset**, reject with
`StatusCode.PermissionDenied`). Differences from the template:
- **Path scope:** gates a *set* of service prefixes (constructor-provided), initially the
sitestream service prefix (read the real package/service name from `sitestream.proto`
`/…SiteStreamService/`); later phases add the new services. LocalDb sync keeps its own
interceptor + key untouched.
- **Key source, site side (server on `:8083`):** new option
`CommunicationOptions.GrpcPsk` (`ScadaBridge:Communication:GrpcPsk`), production value
`${secret:SB-GRPC-PSK-<siteId>}` resolved by the existing pre-host `SecretReferenceExpander`
(`Program.cs:56-65`) — zero new resolution code.
- **Key source, central side (clients today; server in 1A):** central's site set is dynamic, so
central resolves secret **`SB-GRPC-PSK-{siteId}`** at channel-build time via the runtime
`ISecretResolver` (copy the fail-closed lazy pattern from
`DataConnectionLayer/Adapters/MxGatewayDataConnection.cs:82-98`), cached per site, cache
invalidated on site remove. New helper `SitePskProvider` in Communication (interface) +
Host (implementation over `ISecretResolver`).
Wire: add the interceptor to the site's existing `AddGrpc` (`Program.cs:526-527`, alongside the
LocalDb one). Attach the PSK on central's existing site-dialing clients as call-level
`Metadata` `Authorization: Bearer <psk>` (the LocalDb sync client pattern,
`SyncBackgroundService.cs:82-86`): `SiteStreamGrpcClient` (subscribe calls),
`GrpcPullAuditEventsClient`, `GrpcPullSiteCallsClient` — each already flows through a channel/
invoker creation point where the site id is known.
### T0.4 Rig + tests
- Rig: mirror the LocalDb dev-key pattern — literal `ScadaBridge:Communication:GrpcPsk:
"dev-grpc-psk-site-a"` in BOTH `docker/site-a-node-*/appsettings.Site.json` (and site-b/c with
their own keys, since unlike LocalDb this is not optional), plus central-side dev secrets: seed
`SB-GRPC-PSK-site-{a,b,c}` into both centrals' secret stores (secret CLI seed flow, or dev-KEK
env as the rig's Secrets setup already does).
- Tests: interceptor unit tests (wrong key / missing header / unset expected key ⇒
`PermissionDenied`; non-gated service path passes; constant-time compare exercised);
`SitePskProvider` fail-closed test; one Host wiring test asserting the interceptor is
registered.
**Phase 0 DoD:** suite green; on the rig, an unauthenticated `grpcurl`/test client gets
`PermissionDenied` on SiteStream, authenticated streaming + pull still work end-to-end. PR merged.
---
## Phase 1A — Central control plane (site→central) — worktree A
### T1A.1 `central_control.proto`
New `Protos/central_control.proto` in the Communication project (checked-in codegen per recipe),
package `scadabridge.centralcontrol.v1`, service `CentralControlService`:
| RPC | Wraps DTO (source file under `Commons/Messages/`) | Notes |
|---|---|---|
| `SubmitNotification` | `NotificationSubmit`/`NotificationSubmitAck` (`Notification/NotificationMessages.cs:30,47`) | 11 fields incl. `Guid?` execution ids → string |
| `QueryNotificationStatus` | `NotificationStatusQuery`/`Response` (`:55,62`) | |
| `IngestAuditEvents` | reuse existing `AuditEventBatch`/`IngestAck` from sitestream.proto | import, don't duplicate; `ForwardState`/`IngestedAtUtc` stay off-wire (`AuditEventDtoMapper.cs:23-27`) |
| `IngestCachedTelemetry` | reuse `CachedTelemetryBatch`/`IngestAck` | same |
| `ReconcileSite` | `ReconcileSiteRequest`/`Response` (`Deployment/ReconcileSiteRequest.cs`, `ReconcileSiteResponse.cs` incl. `ReconcileGapItem`) | map<string,string> for name→hash |
| `ReportSiteHealth` | `SiteHealthReport`/`SiteHealthReportAck` (`Health/SiteHealthReport.cs`) | the big one: ~30 fields incl. maps of `ConnectionHealth` enum, `TagResolutionStatus`, `TagQualityCounts`, `NodeStatus` list, `SiteAuditBacklogSnapshot`; model nullable ints/doubles with wrappers; keep `SequenceNumber` |
| `Heartbeat` | `HeartbeatMessage` (`Health/HeartbeatMessage.cs`) → `google.protobuf.Empty` reply | fire-and-forget semantics preserved client-side (don't await failure into caller) |
Mappers in `Communication/Grpc/CentralControlDtoMapper.cs` with **round-trip golden tests**
(construct DTO → proto → DTO, assert deep-equal; include null-optional cases).
### T1A.2 Central hosting
Central branch of `Program.cs` (~`:262` services, `:449-469` Map block):
`builder.Services.AddGrpc(o => o.Interceptors.Add<ControlPlaneAuthInterceptor>())` — central's
interceptor variant verifies `Bearer` against the per-site PSK looked up by the **required
`x-scadabridge-site` metadata header** via `SitePskProvider` (fail-closed: missing header ⇒
`PermissionDenied`). Add an **explicit Kestrel listener** for h2c gRPC: new option
`ScadaBridge:Node:CentralGrpcPort` (default **8083**, symmetric with sites), configured like the
site branch does (`Program.cs:507-512`, `HttpProtocols.Http2`); central's `:5000` stays as-is
(Traefik is HTTP/1 — gRPC does NOT go through traefik). `MapGrpcService<CentralControlGrpcService>()`.
`CentralControlGrpcService` (Host or Communication): decode proto → the SAME message types →
`Ask` the existing `CentralCommunicationActor` handlers (they already handle all 7 — zero
handler logic changes) → encode reply. Reuse the readiness convention: reject `Unavailable`
until the central actor system is up (mirror `SiteStreamGrpcServer.SetReady`).
### T1A.3 Site-side client + transport seam
- New `ICentralTransport` (Communication): one method per the 7 sends. Implementations:
`AkkaCentralTransport` (extracted verbatim from today's `SiteCommunicationActor` send blocks)
and `GrpcCentralTransport` (new).
- `GrpcCentralTransport`: **channel pair with sticky failover + failback** per design §3.5 — new
shared `CentralChannelProvider`: endpoints from new option
`ScadaBridge:Communication:CentralGrpcEndpoints` (List<string>, e.g.
`["http://scadabridge-central-a:8083","http://scadabridge-central-b:8083"]`; validator: required
when transport=Grpc); sticky-until-failure; flip on connect-fail/`Unavailable`; background
failback probe every 3060 s (gRPC health or a `Heartbeat` ping); reconnect backoff copied from
`SyncBackgroundService.cs:151` (1 s doubling, cap 60 s). Attach PSK (`GrpcPsk` option) +
`x-scadabridge-site` on every call; per-call deadlines from the matching `CommunicationOptions`
timeout (`NotificationForwardTimeout`, `HealthReportTimeout`, etc. — today's Ask timeouts,
unchanged values). **Cross-node auto-retry only on connect-fail/`Unavailable`** — never on
`DeadlineExceeded`.
- `SiteCommunicationActor` selects the implementation from new option
`ScadaBridge:Communication:CentralTransport` (`Akka` | `Grpc`, **default `Akka`**). The 7
handler bodies delegate to the injected transport; reply/fault semantics identical (timeout or
non-OK status ⇒ same `Status.Failure` the S&F/audit layers already treat as transient).
### T1A.4 Tests
Extend `Communication.Tests` (TestKit): `SiteCommunicationActor` with an NSubstitute
`ICentralTransport` — all 7 paths, fault propagation (transport throw ⇒ same failure the S&F
tests expect). `GrpcCentralTransport` unit tests with an in-process `TestServer` gRPC host:
failover flip, sticky behavior, failback probe, PSK attached, deadline set, no-retry-on-deadline.
Reuse/extend `DirectActorSiteStreamAuditClient` for the ingest integration harness.
`NotificationForwarderTests`/`SiteAuditTelemetryActorTests`/`HealthReportSenderTests` must pass
unmodified (they sit above the seam — if they need edits, the seam is wrong).
**1A DoD:** suite green; on the rig with site-a flipped to `CentralTransport=Grpc`
(central gRPC port published, e.g. `9013/9014:8083`): notification e2e, audit rows land, health
page live, heartbeat drives active flag, reconcile works after site restart — while site-b/c
still run Akka (coexistence proven).
---
## Phase 1B — Site command plane (central→site) — worktree B
### T1B.1 `site_command.proto`
Package `scadabridge.sitecommand.v1`, service `SiteCommandService` — **28 commands** (29 minus
dead `IntegrationCallRequest`), grouped into domain RPCs with `oneof` request/response
envelopes (full command list + reply types + `CommunicationService.cs` line refs in the design
doc §7 / recon inventory):
| RPC | Commands (count) | Deadline source |
|---|---|---|
| `ExecuteLifecycle` | RefreshDeployment, Enable/Disable/DeleteInstance, DeploymentStateQuery, DeployArtifacts (6) | `DeploymentTimeout`/`LifecycleTimeout`/`ArtifactDeploymentTimeout` |
| `ExecuteOpcUa` | BrowseNode, SearchAddressSpace, ReadTagValues, VerifyEndpoint, Trust/List/RemoveServerCert, WriteTag (8) | `QueryTimeout` (browse/search per existing Ask usage) |
| `ExecuteQuery` | EventLogQuery, DebugSnapshot, Subscribe/UnsubscribeDebugView (4) | `QueryTimeout`/`DebugViewTimeout` |
| `ExecuteParked` | ParkedMessageQuery/Retry/Discard, RetryParkedOperation, DiscardParkedOperation (5) | `QueryTimeout`; relay callers keep `RelayTimeout`(10s) < `QueryTimeout`(30s) ordering |
| `ExecuteRoute` | RouteToCall/GetAttributes/SetAttributes/WaitForAttribute (4) | `IntegrationTimeout`; WaitForAttribute uses its dynamic timeout |
| `TriggerFailover` | TriggerSiteFailover (1) | `LifecycleTimeout` |
Mappers `SiteCommandDtoMapper.cs` + round-trip golden tests for every command/reply (the bulk of
this track — budget accordingly; enums, `TrackedOperationId` struct → string guid, nullable
wrappers).
### T1B.2 Site server: shared dispatcher
Refactor `SiteCommunicationActor`'s receive table into `SiteCommandDispatcher` (pure routing:
message → `_deploymentManagerProxy` / `_artifactHandler` / `_eventLogHandler` /
`_parkedMessageHandler` / failover handler — preserving EXACTLY today's targets, including the
node-local parked-message handler; see design §7.3, the replicated-store semantics are
deliberate). The actor and a new `SiteCommandGrpcService` (mapped in the site branch next to
`SiteStreamGrpcServer`, gated by `ControlPlaneAuthInterceptor` + the readiness flag) both call
the dispatcher — one routing truth, both transports.
### T1B.3 Central client + transport seam
`ISiteCommandTransport` (send `SiteEnvelope`-equivalent, Ask or Tell) injected into
**`CentralCommunicationActor`**; implementations `AkkaSiteTransport` (today's per-site
ClusterClient path, extracted) and `GrpcSiteTransport`. `GrpcSiteTransport` uses a new shared
**`SitePairChannelProvider`**: addresses from the `Site` entity's existing
`GrpcNodeAAddress`/`GrpcNodeBAddress` (the streaming path's columns — do NOT invent
`LoadSiteAddressesFromDb`, it doesn't exist; reuse the `ISiteRepository` reads +
`CentralCommunicationActor`'s existing DB-driven cache-refresh loop `:532-598` to build/refresh
channels instead of ClusterClients), sticky failover/failback per §3.5, PSK from
`SitePskProvider` + deadlines per the table above. Config flag
`ScadaBridge:Communication:SiteTransport` (`Akka` | `Grpc`, default `Akka`) selected inside
`CentralCommunicationActor` — `CommunicationService` and `SiteCallAuditActor` unchanged.
### T1B.4 Tests
TestKit: `CentralCommunicationActor` with substitute `ISiteCommandTransport` (envelope routing,
Ask-sender reply plumbing, per-site transport lifecycle on site add/remove/change);
`SiteCommandDispatcher` unit tests (every command → correct target, incl. parked→local handler
and failover→local); `SiteCommandGrpcService` via TestServer (auth, readiness, one command per
oneof group); existing `CommunicationServiceTests`/`CentralCommunicationActor*Tests` pass with
the Akka implementation as default.
**1B DoD:** suite green; rig central flipped to `SiteTransport=Grpc` for site-a only: from
CentralUI — deploy refresh, enable/disable instance, browse, read tag, write tag, event-log
query, parked query/retry/discard (run retry against the STANDBY site node explicitly —
replicated-store semantics), `TriggerSiteFailover` — all work; site-b/c untouched on Akka.
---
## Phase 2 (after 1A) ∥ Phase 3 (after 1B) — full cutover on the rig + hardening
- **Phase 2:** flip all sites to `CentralTransport=Grpc`. Soak: S&F drain under central-a kill
(rows stay Pending, resume without loss or duplicates — sequence/dedup layers unchanged),
failback observed when central-a returns, health/heartbeat cadence unchanged in
`CentralHealthAggregator` (no sequence regressions logged).
- **Phase 3:** flip central to `SiteTransport=Grpc` for all sites. Soak: full CentralUI command
matrix against each site; site-node kill mid-command returns a clean error (no hang beyond
deadline); site pair failover mid-stream of commands.
- Both phases: watch for `PermissionDenied` noise (would indicate PSK drift), and confirm zero
ClusterClient log activity on flipped paths.
## Phase 4 — deletion + config cutover (sequential, after 2+3)
1. Flip both flag defaults to `Grpc`; rig + docs updated; one soak cycle.
2. Delete: `AkkaCentralTransport`/`AkkaSiteTransport`, ClusterClient creation
(`AkkaHostedService.cs:942-953`), `DefaultSiteClientFactory` (+ its tests), per-site
ClusterClient cache in `CentralCommunicationActor` (keep the DB refresh loop — it now feeds
`SitePairChannelProvider`), receptionist registrations `:436` and `:935` (T0.1 already removed
`:459`), `CommunicationOptions.CentralContactPoints` (+ validator + rig configs), then the
flags themselves.
3. Grep-gates: `rg -i "clusterclient|receptionist" src tests docker docs` → only historical docs;
`rg "CentralContactPoints"` → empty.
4. Docs: update `grpc_streams.md` (its "ClusterClient keeps command/control" split is
superseded; also fix its `LoadSiteAddressesFromDb` doc-vs-code gap), `Component-Host.md`,
`Component-StoreAndForward.md:137`, and add a `docs/known-issues` cross-ref note that the
frame-size class is retired. Keep `Akka.Cluster.Tools` (ClusterSingleton still used).
## Phase 5 — live gate (rig; sequential; every check PASS required)
1. **PSK negative:** no key / wrong key / missing `x-scadabridge-site` ⇒ `PermissionDenied`;
unset server key ⇒ all rejected (fail-closed); LocalDb sync key unaffected.
2. **Site→central matrix:** notification e2e (delivered + status query), audit + cached-telemetry
rows in `dbo.AuditLog`/site-calls, `/monitoring/health` live per site, heartbeat → active
flag, reconcile self-heal after site-node restart.
3. **Central→site matrix:** all 6 RPC groups exercised from CentralUI, parked retry/discard on
the standby node, failover command drains cleanly.
4. **Failover/failback:** kill central-a → sites flip to central-b sticky (S&F uninterrupted);
restart central-a → failback within probe cadence; same for a site node from central's side;
booting node rejects `Unavailable` until ready and the client fails over.
5. **Mid-drain kill:** kill central during an S&F drain burst — zero loss, zero duplicates.
6. **Frame-class retirement:** issue a command/reply > 128 KB (large browse/event-log result) —
succeeds over gRPC (impossible before).
7. **Boundary check:** with everything on gRPC, verify NO Akka association exists between any
site container and central (`netstat`/Akka logs) — remoting is pair-internal only.
8. **Restart discipline:** full-rig restart (pairs together) comes up clean; no receptionist/
ClusterClient log lines anywhere.
Record results in `docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md` (check-by-check,
the family's live-gate format).
## Effort & parallelization summary
| Track | Est. | Parallel with |
|---|---|---|
| Phase 0 | 23 d | 1A/1B proto authoring |
| 1A | 1.52 wk | 1B (separate worktrees; 1A merges first) |
| 1B | 23 wk (mapper-heavy) | 1A |
| 2, 3 | 24 d each | each other (independent flags/paths) |
| 4 | 23 d | — |
| 5 | 23 d | — |
Critical path ≈ 1B: **~46 weeks total**, matching the design estimate.
## Task checklist (tick as you go; IDs reference the sections above)
**Phase 0 — PSK + dead code** (branch `feat/grpc-phase0-psk`)
- [x] T0.1 Delete `/user/management` receptionist registration (`AkkaHostedService.cs:459`) + Component-Host.md update
- [x] T0.2 File dead-`IntegrationCallRequest` issue; record exclusion (28 of 29 migrate)
- [x] T0.3 `ControlPlaneAuthInterceptor` + `CommunicationOptions.GrpcPsk` + `SitePskProvider`; gate SiteStream; PSK attached on central's streaming + pull clients
- [x] T0.4 Rig dev keys (all 3 sites + central store seeds) + interceptor/provider/wiring tests
- [x] Phase 0 DoD: suite green; rig unauthenticated ⇒ `PermissionDenied`, authenticated paths work; PR merged (#25, ff to `main` @ `3fa95555`; gate PASS in `2026-07-22-clusterclient-to-grpc-live-gate.md`)
**Phase 1A — central control plane** (worktree, `feat/grpc-central-control`)
- [x] T1A.1 `central_control.proto` (7 RPCs; checked-in codegen) + `CentralControlDtoMapper` + round-trip golden tests
- [x] T1A.2 Central hosting: `AddGrpc` + per-site-PSK interceptor (`x-scadabridge-site`), `CentralGrpcPort` h2c listener (8083), `CentralControlGrpcService` (Ask existing handlers), readiness gate
- [ ] T1A.3 `ICentralTransport` (Akka extract + Grpc impl), `CentralChannelProvider` (sticky failover/failback, backoff, deadlines, PSK), `CentralTransport` flag default `Akka`, `CentralGrpcEndpoints` option + validator
- [ ] T1A.4 Tests: actor-with-fake-transport ×7, TestServer transport tests, S&F/audit/health suites pass unmodified
- [ ] 1A DoD: rig site-a on `Grpc` proves all 5 site→central paths while site-b/c stay Akka; PR merged (before 1B)
**Phase 1B — site command plane** (worktree, `feat/grpc-sitecommand`)
- [x] T1B.1 `site_command.proto` (6 oneof RPCs / 28 commands) + `SiteCommandDtoMapper` + round-trip golden tests (all 28 + replies)
- [ ] T1B.2 `SiteCommandDispatcher` refactor (actor + new `SiteCommandGrpcService` share it; parked stays node-local)
- [ ] T1B.3 `ISiteCommandTransport` in `CentralCommunicationActor` (Akka extract + Grpc impl), `SitePairChannelProvider` (Site entity Grpc columns + DB refresh loop), `SiteTransport` flag default `Akka`
- [ ] T1B.4 Tests: dispatcher routing ×28, actor envelope/reply plumbing, TestServer service tests, existing Communication suites green
- [ ] 1B DoD: rig central on `Grpc` for site-a proves full command matrix incl. standby parked retry; rebased on 1A; PR merged
**Phase 2 ∥ 3 — cutover + soak**
- [ ] P2 All sites `CentralTransport=Grpc`; central-kill S&F soak (no loss/dupes), failback observed, health sequences clean
- [ ] P3 Central `SiteTransport=Grpc` all sites; full UI command matrix per site; site-kill mid-command clean; no PSK noise; zero ClusterClient log activity on flipped paths
**Phase 4 — deletion**
- [ ] Defaults flip to `Grpc` + soak; then delete Akka transports, ClusterClient creation, `DefaultSiteClientFactory`, receptionist registrations, `CentralContactPoints`, then the flags
- [ ] Grep-gates pass (`clusterclient|receptionist` → historical docs only; `CentralContactPoints` → empty)
- [ ] Docs updated: `grpc_streams.md`, `Component-Host.md`, `Component-StoreAndForward.md:137`, known-issues cross-ref
**Phase 5 — live gate** (record in `2026-07-22-clusterclient-to-grpc-live-gate.md`)
- [ ] 1 PSK negatives · [ ] 2 site→central matrix · [ ] 3 central→site matrix · [ ] 4 failover/failback both directions · [ ] 5 mid-drain kill · [ ] 6 >128 KB frame-class proof · [ ] 7 no cross-boundary Akka association · [ ] 8 full-rig restart clean
## Gotchas for the executor (will bite; read twice)
- Generated proto C# is committed; never add an active `<Protobuf>` item to the csproj in a
final commit (linux_arm64 protoc segfault breaks the Docker build).
- `HeartbeatMessage` must stay fire-and-forget end-to-end — don't let a gRPC failure surface as
a fault to the heartbeat timer path.
- Deadline ≠ retry: no automatic cross-node retry on `DeadlineExceeded` for WriteTag/Deploy/
Failover; only on provably-unsent failures.
- The parked-message handler is node-local **on purpose** (replicated store); do not "fix" it
onto the singleton proxy.
- Ack-before-Leave on `TriggerSiteFailover` (`SiteCommunicationActor.cs:569`): the gRPC reply
must be written before the node leaves — verify the response completes under failover.
- Inner-before-outer timeouts: `RelayTimeout`(10 s) < `QueryTimeout`(30 s) must survive the
deadline mapping (`CommunicationService.cs:779-786` documents why).
- `SiteStreamGrpcServer.AuditIngestAskTimeout` (30 s) is "one source of truth" shared with
`CentralCommunicationActor` — keep the new central service on the same constant.
- Rig sites reach central by container name (`scadabridge-central-{a,b}:8083`), NOT via traefik
(HTTP/1 only).
- xunit v2: use `Xunit.SkippableFact` for env-gated tests, not `Assert.Skip`.