docs(plans): ClusterClient→gRPC-only migration plan (phases 0–5 + tasklist)
Complete executable plan: PSK-from-Secrets auth (Phase 0), central_control + site_command proto contracts behind transport seams inside the two communication actors (1A ∥ 1B in worktrees), per-direction cutover flags, ClusterClient/receptionist deletion, 8-check live gate. Design doc: scadaproj/scadabridge_clusterclient_to_grpc.md (§7 = deep-dive corrections).
This commit is contained in:
@@ -0,0 +1,355 @@
|
||||
# ClusterClient → gRPC-only cross-cluster transport — implementation plan
|
||||
|
||||
**Date:** 2026-07-22. **Design:** `~/Desktop/scadaproj/scadabridge_clusterclient_to_grpc.md`
|
||||
(read it first, §7 especially — it contains the deep-dive corrections this plan builds on).
|
||||
**Goal:** all site↔central traffic rides gRPC with PSK auth from `ZB.MOM.WW.Secrets`;
|
||||
`ClusterClient`/`ClusterClientReceptionist` are deleted; Akka remoting never crosses the
|
||||
site↔central boundary.
|
||||
|
||||
Everything below was code-verified 2026-07-22. If a cited line has drifted, re-locate by the
|
||||
quoted identifier, never by line number.
|
||||
|
||||
## How to execute this plan (read before starting)
|
||||
|
||||
- **Phase order:** 0 → (1A ∥ 1B) → (2 ∥ 3) → 4 → 5. Phases marked ∥ are parallel-safe **only in
|
||||
separate git worktrees** (`git worktree add ../ScadaBridge-1B feat/grpc-sitecommand`) — never
|
||||
run two agents against one working tree (destructive git races are a known family incident).
|
||||
Designated merge order when tracks meet: 1A lands first, 1B rebases (expected conflicts are
|
||||
confined to `ZB.MOM.WW.ScadaBridge.Communication.csproj` proto ItemGroups and `Program.cs`
|
||||
service/Map blocks — resolve as union).
|
||||
- **Branches:** `feat/grpc-phase0-psk`, `feat/grpc-central-control` (1A),
|
||||
`feat/grpc-sitecommand` (1B), then per-phase branches; PR each phase to `main` with its DoD met.
|
||||
- **Build/test:** `dotnet build ZB.MOM.WW.ScadaBridge.slnx` (0 warnings — TreatWarningsAsErrors),
|
||||
`dotnet test ZB.MOM.WW.ScadaBridge.slnx`. Tests are xunit **v2** (2.9.3) + `Akka.TestKit.Xunit2`
|
||||
1.5.62 + NSubstitute. TestKit tests inherit `TestKit` — never hand-roll
|
||||
`ActorSystem.Create`+join.
|
||||
- **Proto codegen is CHECKED-IN, not build-time** (protoc segfaults in the linux_arm64 image).
|
||||
For every new/changed proto follow the sitestream recipe documented in
|
||||
`src/ZB.MOM.WW.ScadaBridge.Communication/ZB.MOM.WW.ScadaBridge.Communication.csproj`:
|
||||
temporarily uncomment/add the `<Protobuf Include=... GrpcServices="Both" />` item, delete stale
|
||||
generated files, `dotnet build` on macOS, copy `obj/**/Protos/*.cs` into a committed folder
|
||||
(mirror `SiteStreamGrpc/`), re-comment the item. `docker/regen-proto.sh` exists.
|
||||
- **Rig:** `docker/deploy.sh` rebuilds the 2-central + 3×2-site cluster (+ traefik). Central UI
|
||||
`9001/9002:5000`, site gRPC `90x3/90x4:8083`. **External sites↔central rules:** never
|
||||
host-`sqlite3` a live WAL DB; `aspnet:10.0` has no `curl`; seeding via `docker/seed-sites.sh`.
|
||||
- **Coexistence rule:** every migrated path sits behind a config flag with the Akka
|
||||
implementation as default until Phase 4. Rollback at any point = flip the flag.
|
||||
|
||||
## The two choke points (where all transport code changes)
|
||||
|
||||
- **Site→central:** `SiteCommunicationActor`
|
||||
(`src/ZB.MOM.WW.ScadaBridge.Communication/Actors/SiteCommunicationActor.cs`) — 7 message types,
|
||||
all `ClusterClient.Send` to `/user/central-communication` (Ask, except `HeartbeatMessage` Tell).
|
||||
- **Central→site:** `CentralCommunicationActor` (same dir) — the `SiteEnvelope` handler
|
||||
(`:452-467`) unwraps and routes via per-site ClusterClient. ALL producers
|
||||
(`CommunicationService`'s 27 + `SiteCallAuditActor`'s 2 relays) go through it.
|
||||
**`CommunicationService` gets NO interface extraction; its ~20 consumers are untouched.**
|
||||
|
||||
---
|
||||
|
||||
## Phase 0 — PSK auth + dead-code removal (standalone hardening; parallel-safe with 1A/1B prep)
|
||||
|
||||
### T0.1 Delete the vestigial `/user/management` receptionist registration
|
||||
`AkkaHostedService.cs:459` (`RegisterService` of `ManagementActor`) — delete the registration
|
||||
only (the actor stays; `ManagementEndpoints.cs:117` asks it in-process via
|
||||
`ManagementActorHolder`). Grep-verify nothing sends to `/user/management`. Update
|
||||
`docs/requirements/Component-Host.md` REQ-HOST-6a to record the removal.
|
||||
|
||||
### T0.2 File the dead-integration issue
|
||||
`IntegrationCallRequest` (`CommunicationService.cs:238`) is unwired in production —
|
||||
`RegisterLocalHandler(Integration, …)` exists only in `SiteCommunicationActorTests.cs:102`, so
|
||||
production always replies "Integration handler not available" (`SiteCommunicationActor.cs:134-136`).
|
||||
File a Gitea issue (decide: wire or delete); **exclude it from the gRPC contract** (28 of 29
|
||||
commands migrate). Do not change its behavior in this program.
|
||||
|
||||
### T0.3 PSK interceptor + options
|
||||
New `src/ZB.MOM.WW.ScadaBridge.Host/ControlPlaneAuthInterceptor.cs`, copied from
|
||||
`LocalDbSyncAuthInterceptor.cs` (same file layout: seal, 4 server handlers → shared
|
||||
`Authorize`, `authorization: Bearer` metadata extraction, `CryptographicOperations.FixedTimeEquals`
|
||||
over UTF-8, **fail-closed when the expected key is unset**, reject with
|
||||
`StatusCode.PermissionDenied`). Differences from the template:
|
||||
|
||||
- **Path scope:** gates a *set* of service prefixes (constructor-provided), initially the
|
||||
sitestream service prefix (read the real package/service name from `sitestream.proto` —
|
||||
`/…SiteStreamService/`); later phases add the new services. LocalDb sync keeps its own
|
||||
interceptor + key untouched.
|
||||
- **Key source, site side (server on `:8083`):** new option
|
||||
`CommunicationOptions.GrpcPsk` (`ScadaBridge:Communication:GrpcPsk`), production value
|
||||
`${secret:SB-GRPC-PSK-<siteId>}` resolved by the existing pre-host `SecretReferenceExpander`
|
||||
(`Program.cs:56-65`) — zero new resolution code.
|
||||
- **Key source, central side (clients today; server in 1A):** central's site set is dynamic, so
|
||||
central resolves secret **`SB-GRPC-PSK-{siteId}`** at channel-build time via the runtime
|
||||
`ISecretResolver` (copy the fail-closed lazy pattern from
|
||||
`DataConnectionLayer/Adapters/MxGatewayDataConnection.cs:82-98`), cached per site, cache
|
||||
invalidated on site remove. New helper `SitePskProvider` in Communication (interface) +
|
||||
Host (implementation over `ISecretResolver`).
|
||||
|
||||
Wire: add the interceptor to the site's existing `AddGrpc` (`Program.cs:526-527`, alongside the
|
||||
LocalDb one). Attach the PSK on central's existing site-dialing clients as call-level
|
||||
`Metadata` `Authorization: Bearer <psk>` (the LocalDb sync client pattern,
|
||||
`SyncBackgroundService.cs:82-86`): `SiteStreamGrpcClient` (subscribe calls),
|
||||
`GrpcPullAuditEventsClient`, `GrpcPullSiteCallsClient` — each already flows through a channel/
|
||||
invoker creation point where the site id is known.
|
||||
|
||||
### T0.4 Rig + tests
|
||||
- Rig: mirror the LocalDb dev-key pattern — literal `ScadaBridge:Communication:GrpcPsk:
|
||||
"dev-grpc-psk-site-a"` in BOTH `docker/site-a-node-*/appsettings.Site.json` (and site-b/c with
|
||||
their own keys, since unlike LocalDb this is not optional), plus central-side dev secrets: seed
|
||||
`SB-GRPC-PSK-site-{a,b,c}` into both centrals' secret stores (secret CLI seed flow, or dev-KEK
|
||||
env as the rig's Secrets setup already does).
|
||||
- Tests: interceptor unit tests (wrong key / missing header / unset expected key ⇒
|
||||
`PermissionDenied`; non-gated service path passes; constant-time compare exercised);
|
||||
`SitePskProvider` fail-closed test; one Host wiring test asserting the interceptor is
|
||||
registered.
|
||||
|
||||
**Phase 0 DoD:** suite green; on the rig, an unauthenticated `grpcurl`/test client gets
|
||||
`PermissionDenied` on SiteStream, authenticated streaming + pull still work end-to-end. PR merged.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1A — Central control plane (site→central) — worktree A
|
||||
|
||||
### T1A.1 `central_control.proto`
|
||||
New `Protos/central_control.proto` in the Communication project (checked-in codegen per recipe),
|
||||
package `scadabridge.centralcontrol.v1`, service `CentralControlService`:
|
||||
|
||||
| RPC | Wraps DTO (source file under `Commons/Messages/`) | Notes |
|
||||
|---|---|---|
|
||||
| `SubmitNotification` | `NotificationSubmit`/`NotificationSubmitAck` (`Notification/NotificationMessages.cs:30,47`) | 11 fields incl. `Guid?` execution ids → string |
|
||||
| `QueryNotificationStatus` | `NotificationStatusQuery`/`Response` (`:55,62`) | |
|
||||
| `IngestAuditEvents` | reuse existing `AuditEventBatch`/`IngestAck` from sitestream.proto | import, don't duplicate; `ForwardState`/`IngestedAtUtc` stay off-wire (`AuditEventDtoMapper.cs:23-27`) |
|
||||
| `IngestCachedTelemetry` | reuse `CachedTelemetryBatch`/`IngestAck` | same |
|
||||
| `ReconcileSite` | `ReconcileSiteRequest`/`Response` (`Deployment/ReconcileSiteRequest.cs`, `ReconcileSiteResponse.cs` incl. `ReconcileGapItem`) | map<string,string> for name→hash |
|
||||
| `ReportSiteHealth` | `SiteHealthReport`/`SiteHealthReportAck` (`Health/SiteHealthReport.cs`) | the big one: ~30 fields incl. maps of `ConnectionHealth` enum, `TagResolutionStatus`, `TagQualityCounts`, `NodeStatus` list, `SiteAuditBacklogSnapshot`; model nullable ints/doubles with wrappers; keep `SequenceNumber` |
|
||||
| `Heartbeat` | `HeartbeatMessage` (`Health/HeartbeatMessage.cs`) → `google.protobuf.Empty` reply | fire-and-forget semantics preserved client-side (don't await failure into caller) |
|
||||
|
||||
Mappers in `Communication/Grpc/CentralControlDtoMapper.cs` with **round-trip golden tests**
|
||||
(construct DTO → proto → DTO, assert deep-equal; include null-optional cases).
|
||||
|
||||
### T1A.2 Central hosting
|
||||
Central branch of `Program.cs` (~`:262` services, `:449-469` Map block):
|
||||
`builder.Services.AddGrpc(o => o.Interceptors.Add<ControlPlaneAuthInterceptor>())` — central's
|
||||
interceptor variant verifies `Bearer` against the per-site PSK looked up by the **required
|
||||
`x-scadabridge-site` metadata header** via `SitePskProvider` (fail-closed: missing header ⇒
|
||||
`PermissionDenied`). Add an **explicit Kestrel listener** for h2c gRPC: new option
|
||||
`ScadaBridge:Node:CentralGrpcPort` (default **8083**, symmetric with sites), configured like the
|
||||
site branch does (`Program.cs:507-512`, `HttpProtocols.Http2`); central's `:5000` stays as-is
|
||||
(Traefik is HTTP/1 — gRPC does NOT go through traefik). `MapGrpcService<CentralControlGrpcService>()`.
|
||||
|
||||
`CentralControlGrpcService` (Host or Communication): decode proto → the SAME message types →
|
||||
`Ask` the existing `CentralCommunicationActor` handlers (they already handle all 7 — zero
|
||||
handler logic changes) → encode reply. Reuse the readiness convention: reject `Unavailable`
|
||||
until the central actor system is up (mirror `SiteStreamGrpcServer.SetReady`).
|
||||
|
||||
### T1A.3 Site-side client + transport seam
|
||||
- New `ICentralTransport` (Communication): one method per the 7 sends. Implementations:
|
||||
`AkkaCentralTransport` (extracted verbatim from today's `SiteCommunicationActor` send blocks)
|
||||
and `GrpcCentralTransport` (new).
|
||||
- `GrpcCentralTransport`: **channel pair with sticky failover + failback** per design §3.5 — new
|
||||
shared `CentralChannelProvider`: endpoints from new option
|
||||
`ScadaBridge:Communication:CentralGrpcEndpoints` (List<string>, e.g.
|
||||
`["http://scadabridge-central-a:8083","http://scadabridge-central-b:8083"]`; validator: required
|
||||
when transport=Grpc); sticky-until-failure; flip on connect-fail/`Unavailable`; background
|
||||
failback probe every 30–60 s (gRPC health or a `Heartbeat` ping); reconnect backoff copied from
|
||||
`SyncBackgroundService.cs:151` (1 s doubling, cap 60 s). Attach PSK (`GrpcPsk` option) +
|
||||
`x-scadabridge-site` on every call; per-call deadlines from the matching `CommunicationOptions`
|
||||
timeout (`NotificationForwardTimeout`, `HealthReportTimeout`, etc. — today's Ask timeouts,
|
||||
unchanged values). **Cross-node auto-retry only on connect-fail/`Unavailable`** — never on
|
||||
`DeadlineExceeded`.
|
||||
- `SiteCommunicationActor` selects the implementation from new option
|
||||
`ScadaBridge:Communication:CentralTransport` (`Akka` | `Grpc`, **default `Akka`**). The 7
|
||||
handler bodies delegate to the injected transport; reply/fault semantics identical (timeout or
|
||||
non-OK status ⇒ same `Status.Failure` the S&F/audit layers already treat as transient).
|
||||
|
||||
### T1A.4 Tests
|
||||
Extend `Communication.Tests` (TestKit): `SiteCommunicationActor` with an NSubstitute
|
||||
`ICentralTransport` — all 7 paths, fault propagation (transport throw ⇒ same failure the S&F
|
||||
tests expect). `GrpcCentralTransport` unit tests with an in-process `TestServer` gRPC host:
|
||||
failover flip, sticky behavior, failback probe, PSK attached, deadline set, no-retry-on-deadline.
|
||||
Reuse/extend `DirectActorSiteStreamAuditClient` for the ingest integration harness.
|
||||
`NotificationForwarderTests`/`SiteAuditTelemetryActorTests`/`HealthReportSenderTests` must pass
|
||||
unmodified (they sit above the seam — if they need edits, the seam is wrong).
|
||||
|
||||
**1A DoD:** suite green; on the rig with site-a flipped to `CentralTransport=Grpc`
|
||||
(central gRPC port published, e.g. `9013/9014:8083`): notification e2e, audit rows land, health
|
||||
page live, heartbeat drives active flag, reconcile works after site restart — while site-b/c
|
||||
still run Akka (coexistence proven).
|
||||
|
||||
---
|
||||
|
||||
## Phase 1B — Site command plane (central→site) — worktree B
|
||||
|
||||
### T1B.1 `site_command.proto`
|
||||
Package `scadabridge.sitecommand.v1`, service `SiteCommandService` — **28 commands** (29 minus
|
||||
dead `IntegrationCallRequest`), grouped into domain RPCs with `oneof` request/response
|
||||
envelopes (full command list + reply types + `CommunicationService.cs` line refs in the design
|
||||
doc §7 / recon inventory):
|
||||
|
||||
| RPC | Commands (count) | Deadline source |
|
||||
|---|---|---|
|
||||
| `ExecuteLifecycle` | RefreshDeployment, Enable/Disable/DeleteInstance, DeploymentStateQuery, DeployArtifacts (6) | `DeploymentTimeout`/`LifecycleTimeout`/`ArtifactDeploymentTimeout` |
|
||||
| `ExecuteOpcUa` | BrowseNode, SearchAddressSpace, ReadTagValues, VerifyEndpoint, Trust/List/RemoveServerCert, WriteTag (8) | `QueryTimeout` (browse/search per existing Ask usage) |
|
||||
| `ExecuteQuery` | EventLogQuery, DebugSnapshot, Subscribe/UnsubscribeDebugView (4) | `QueryTimeout`/`DebugViewTimeout` |
|
||||
| `ExecuteParked` | ParkedMessageQuery/Retry/Discard, RetryParkedOperation, DiscardParkedOperation (5) | `QueryTimeout`; relay callers keep `RelayTimeout`(10s) < `QueryTimeout`(30s) ordering |
|
||||
| `ExecuteRoute` | RouteToCall/GetAttributes/SetAttributes/WaitForAttribute (4) | `IntegrationTimeout`; WaitForAttribute uses its dynamic timeout |
|
||||
| `TriggerFailover` | TriggerSiteFailover (1) | `LifecycleTimeout` |
|
||||
|
||||
Mappers `SiteCommandDtoMapper.cs` + round-trip golden tests for every command/reply (the bulk of
|
||||
this track — budget accordingly; enums, `TrackedOperationId` struct → string guid, nullable
|
||||
wrappers).
|
||||
|
||||
### T1B.2 Site server: shared dispatcher
|
||||
Refactor `SiteCommunicationActor`'s receive table into `SiteCommandDispatcher` (pure routing:
|
||||
message → `_deploymentManagerProxy` / `_artifactHandler` / `_eventLogHandler` /
|
||||
`_parkedMessageHandler` / failover handler — preserving EXACTLY today's targets, including the
|
||||
node-local parked-message handler; see design §7.3, the replicated-store semantics are
|
||||
deliberate). The actor and a new `SiteCommandGrpcService` (mapped in the site branch next to
|
||||
`SiteStreamGrpcServer`, gated by `ControlPlaneAuthInterceptor` + the readiness flag) both call
|
||||
the dispatcher — one routing truth, both transports.
|
||||
|
||||
### T1B.3 Central client + transport seam
|
||||
`ISiteCommandTransport` (send `SiteEnvelope`-equivalent, Ask or Tell) injected into
|
||||
**`CentralCommunicationActor`**; implementations `AkkaSiteTransport` (today's per-site
|
||||
ClusterClient path, extracted) and `GrpcSiteTransport`. `GrpcSiteTransport` uses a new shared
|
||||
**`SitePairChannelProvider`**: addresses from the `Site` entity's existing
|
||||
`GrpcNodeAAddress`/`GrpcNodeBAddress` (the streaming path's columns — do NOT invent
|
||||
`LoadSiteAddressesFromDb`, it doesn't exist; reuse the `ISiteRepository` reads +
|
||||
`CentralCommunicationActor`'s existing DB-driven cache-refresh loop `:532-598` to build/refresh
|
||||
channels instead of ClusterClients), sticky failover/failback per §3.5, PSK from
|
||||
`SitePskProvider` + deadlines per the table above. Config flag
|
||||
`ScadaBridge:Communication:SiteTransport` (`Akka` | `Grpc`, default `Akka`) selected inside
|
||||
`CentralCommunicationActor` — `CommunicationService` and `SiteCallAuditActor` unchanged.
|
||||
|
||||
### T1B.4 Tests
|
||||
TestKit: `CentralCommunicationActor` with substitute `ISiteCommandTransport` (envelope routing,
|
||||
Ask-sender reply plumbing, per-site transport lifecycle on site add/remove/change);
|
||||
`SiteCommandDispatcher` unit tests (every command → correct target, incl. parked→local handler
|
||||
and failover→local); `SiteCommandGrpcService` via TestServer (auth, readiness, one command per
|
||||
oneof group); existing `CommunicationServiceTests`/`CentralCommunicationActor*Tests` pass with
|
||||
the Akka implementation as default.
|
||||
|
||||
**1B DoD:** suite green; rig central flipped to `SiteTransport=Grpc` for site-a only: from
|
||||
CentralUI — deploy refresh, enable/disable instance, browse, read tag, write tag, event-log
|
||||
query, parked query/retry/discard (run retry against the STANDBY site node explicitly —
|
||||
replicated-store semantics), `TriggerSiteFailover` — all work; site-b/c untouched on Akka.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 (after 1A) ∥ Phase 3 (after 1B) — full cutover on the rig + hardening
|
||||
|
||||
- **Phase 2:** flip all sites to `CentralTransport=Grpc`. Soak: S&F drain under central-a kill
|
||||
(rows stay Pending, resume without loss or duplicates — sequence/dedup layers unchanged),
|
||||
failback observed when central-a returns, health/heartbeat cadence unchanged in
|
||||
`CentralHealthAggregator` (no sequence regressions logged).
|
||||
- **Phase 3:** flip central to `SiteTransport=Grpc` for all sites. Soak: full CentralUI command
|
||||
matrix against each site; site-node kill mid-command returns a clean error (no hang beyond
|
||||
deadline); site pair failover mid-stream of commands.
|
||||
- Both phases: watch for `PermissionDenied` noise (would indicate PSK drift), and confirm zero
|
||||
ClusterClient log activity on flipped paths.
|
||||
|
||||
## Phase 4 — deletion + config cutover (sequential, after 2+3)
|
||||
|
||||
1. Flip both flag defaults to `Grpc`; rig + docs updated; one soak cycle.
|
||||
2. Delete: `AkkaCentralTransport`/`AkkaSiteTransport`, ClusterClient creation
|
||||
(`AkkaHostedService.cs:942-953`), `DefaultSiteClientFactory` (+ its tests), per-site
|
||||
ClusterClient cache in `CentralCommunicationActor` (keep the DB refresh loop — it now feeds
|
||||
`SitePairChannelProvider`), receptionist registrations `:436` and `:935` (T0.1 already removed
|
||||
`:459`), `CommunicationOptions.CentralContactPoints` (+ validator + rig configs), then the
|
||||
flags themselves.
|
||||
3. Grep-gates: `rg -i "clusterclient|receptionist" src tests docker docs` → only historical docs;
|
||||
`rg "CentralContactPoints"` → empty.
|
||||
4. Docs: update `grpc_streams.md` (its "ClusterClient keeps command/control" split is
|
||||
superseded; also fix its `LoadSiteAddressesFromDb` doc-vs-code gap), `Component-Host.md`,
|
||||
`Component-StoreAndForward.md:137`, and add a `docs/known-issues` cross-ref note that the
|
||||
frame-size class is retired. Keep `Akka.Cluster.Tools` (ClusterSingleton still used).
|
||||
|
||||
## Phase 5 — live gate (rig; sequential; every check PASS required)
|
||||
|
||||
1. **PSK negative:** no key / wrong key / missing `x-scadabridge-site` ⇒ `PermissionDenied`;
|
||||
unset server key ⇒ all rejected (fail-closed); LocalDb sync key unaffected.
|
||||
2. **Site→central matrix:** notification e2e (delivered + status query), audit + cached-telemetry
|
||||
rows in `dbo.AuditLog`/site-calls, `/monitoring/health` live per site, heartbeat → active
|
||||
flag, reconcile self-heal after site-node restart.
|
||||
3. **Central→site matrix:** all 6 RPC groups exercised from CentralUI, parked retry/discard on
|
||||
the standby node, failover command drains cleanly.
|
||||
4. **Failover/failback:** kill central-a → sites flip to central-b sticky (S&F uninterrupted);
|
||||
restart central-a → failback within probe cadence; same for a site node from central's side;
|
||||
booting node rejects `Unavailable` until ready and the client fails over.
|
||||
5. **Mid-drain kill:** kill central during an S&F drain burst — zero loss, zero duplicates.
|
||||
6. **Frame-class retirement:** issue a command/reply > 128 KB (large browse/event-log result) —
|
||||
succeeds over gRPC (impossible before).
|
||||
7. **Boundary check:** with everything on gRPC, verify NO Akka association exists between any
|
||||
site container and central (`netstat`/Akka logs) — remoting is pair-internal only.
|
||||
8. **Restart discipline:** full-rig restart (pairs together) comes up clean; no receptionist/
|
||||
ClusterClient log lines anywhere.
|
||||
|
||||
Record results in `docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md` (check-by-check,
|
||||
the family's live-gate format).
|
||||
|
||||
## Effort & parallelization summary
|
||||
|
||||
| Track | Est. | Parallel with |
|
||||
|---|---|---|
|
||||
| Phase 0 | 2–3 d | 1A/1B proto authoring |
|
||||
| 1A | 1.5–2 wk | 1B (separate worktrees; 1A merges first) |
|
||||
| 1B | 2–3 wk (mapper-heavy) | 1A |
|
||||
| 2, 3 | 2–4 d each | each other (independent flags/paths) |
|
||||
| 4 | 2–3 d | — |
|
||||
| 5 | 2–3 d | — |
|
||||
|
||||
Critical path ≈ 1B: **~4–6 weeks total**, matching the design estimate.
|
||||
|
||||
## Task checklist (tick as you go; IDs reference the sections above)
|
||||
|
||||
**Phase 0 — PSK + dead code** (branch `feat/grpc-phase0-psk`)
|
||||
- [ ] T0.1 Delete `/user/management` receptionist registration (`AkkaHostedService.cs:459`) + Component-Host.md update
|
||||
- [ ] T0.2 File dead-`IntegrationCallRequest` issue; record exclusion (28 of 29 migrate)
|
||||
- [ ] T0.3 `ControlPlaneAuthInterceptor` + `CommunicationOptions.GrpcPsk` + `SitePskProvider`; gate SiteStream; PSK attached on central's streaming + pull clients
|
||||
- [ ] T0.4 Rig dev keys (all 3 sites + central store seeds) + interceptor/provider/wiring tests
|
||||
- [ ] Phase 0 DoD: suite green; rig unauthenticated ⇒ `PermissionDenied`, authenticated paths work; PR merged
|
||||
|
||||
**Phase 1A — central control plane** (worktree, `feat/grpc-central-control`)
|
||||
- [ ] T1A.1 `central_control.proto` (7 RPCs; checked-in codegen) + `CentralControlDtoMapper` + round-trip golden tests
|
||||
- [ ] T1A.2 Central hosting: `AddGrpc` + per-site-PSK interceptor (`x-scadabridge-site`), `CentralGrpcPort` h2c listener (8083), `CentralControlGrpcService` (Ask existing handlers), readiness gate
|
||||
- [ ] T1A.3 `ICentralTransport` (Akka extract + Grpc impl), `CentralChannelProvider` (sticky failover/failback, backoff, deadlines, PSK), `CentralTransport` flag default `Akka`, `CentralGrpcEndpoints` option + validator
|
||||
- [ ] T1A.4 Tests: actor-with-fake-transport ×7, TestServer transport tests, S&F/audit/health suites pass unmodified
|
||||
- [ ] 1A DoD: rig site-a on `Grpc` proves all 5 site→central paths while site-b/c stay Akka; PR merged (before 1B)
|
||||
|
||||
**Phase 1B — site command plane** (worktree, `feat/grpc-sitecommand`)
|
||||
- [ ] T1B.1 `site_command.proto` (6 oneof RPCs / 28 commands) + `SiteCommandDtoMapper` + round-trip golden tests (all 28 + replies)
|
||||
- [ ] T1B.2 `SiteCommandDispatcher` refactor (actor + new `SiteCommandGrpcService` share it; parked stays node-local)
|
||||
- [ ] T1B.3 `ISiteCommandTransport` in `CentralCommunicationActor` (Akka extract + Grpc impl), `SitePairChannelProvider` (Site entity Grpc columns + DB refresh loop), `SiteTransport` flag default `Akka`
|
||||
- [ ] T1B.4 Tests: dispatcher routing ×28, actor envelope/reply plumbing, TestServer service tests, existing Communication suites green
|
||||
- [ ] 1B DoD: rig central on `Grpc` for site-a proves full command matrix incl. standby parked retry; rebased on 1A; PR merged
|
||||
|
||||
**Phase 2 ∥ 3 — cutover + soak**
|
||||
- [ ] P2 All sites `CentralTransport=Grpc`; central-kill S&F soak (no loss/dupes), failback observed, health sequences clean
|
||||
- [ ] P3 Central `SiteTransport=Grpc` all sites; full UI command matrix per site; site-kill mid-command clean; no PSK noise; zero ClusterClient log activity on flipped paths
|
||||
|
||||
**Phase 4 — deletion**
|
||||
- [ ] Defaults flip to `Grpc` + soak; then delete Akka transports, ClusterClient creation, `DefaultSiteClientFactory`, receptionist registrations, `CentralContactPoints`, then the flags
|
||||
- [ ] Grep-gates pass (`clusterclient|receptionist` → historical docs only; `CentralContactPoints` → empty)
|
||||
- [ ] Docs updated: `grpc_streams.md`, `Component-Host.md`, `Component-StoreAndForward.md:137`, known-issues cross-ref
|
||||
|
||||
**Phase 5 — live gate** (record in `2026-07-22-clusterclient-to-grpc-live-gate.md`)
|
||||
- [ ] 1 PSK negatives · [ ] 2 site→central matrix · [ ] 3 central→site matrix · [ ] 4 failover/failback both directions · [ ] 5 mid-drain kill · [ ] 6 >128 KB frame-class proof · [ ] 7 no cross-boundary Akka association · [ ] 8 full-rig restart clean
|
||||
|
||||
## Gotchas for the executor (will bite; read twice)
|
||||
|
||||
- Generated proto C# is committed; never add an active `<Protobuf>` item to the csproj in a
|
||||
final commit (linux_arm64 protoc segfault breaks the Docker build).
|
||||
- `HeartbeatMessage` must stay fire-and-forget end-to-end — don't let a gRPC failure surface as
|
||||
a fault to the heartbeat timer path.
|
||||
- Deadline ≠ retry: no automatic cross-node retry on `DeadlineExceeded` for WriteTag/Deploy/
|
||||
Failover; only on provably-unsent failures.
|
||||
- The parked-message handler is node-local **on purpose** (replicated store); do not "fix" it
|
||||
onto the singleton proxy.
|
||||
- Ack-before-Leave on `TriggerSiteFailover` (`SiteCommunicationActor.cs:569`): the gRPC reply
|
||||
must be written before the node leaves — verify the response completes under failover.
|
||||
- Inner-before-outer timeouts: `RelayTimeout`(10 s) < `QueryTimeout`(30 s) must survive the
|
||||
deadline mapping (`CommunicationService.cs:779-786` documents why).
|
||||
- `SiteStreamGrpcServer.AuditIngestAskTimeout` (30 s) is "one source of truth" shared with
|
||||
`CentralCommunicationActor` — keep the new central service on the same constant.
|
||||
- Rig sites reach central by container name (`scadabridge-central-{a,b}:8083`), NOT via traefik
|
||||
(HTTP/1 only).
|
||||
- xunit v2: use `Xunit.SkippableFact` for env-gated tests, not `Assert.Skip`.
|
||||
Reference in New Issue
Block a user