From 54da10dc005f2ad169fce9802f3e5d21b514aaf9 Mon Sep 17 00:00:00 2001 From: Joseph Doherty Date: Thu, 23 Jul 2026 11:45:12 -0400 Subject: [PATCH] =?UTF-8?q?docs(grpc):=20Phase=203=20live=20gate=20?= =?UTF-8?q?=E2=80=94=20PASS;=20central=E2=86=92site=20command=20cutover?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Central flipped to SiteTransport=Grpc (both nodes, central-wide flag), sites kept on gRPC from Phase 2 → both directions on gRPC at once. Central logs 'central→site command transport: gRPC (SiteCommandService)'; 0 ClusterClient-to- site on either central. Command matrix over SiteCommandService (all 200): - ExecuteQuery (event-log) + ExecuteParked (parked) — all 3 sites - ExecuteLifecycle disable/enable #95 on site-a — NEW live coverage vs 1B/P2 (SoakNotify instances survived the recreate as Enabled) Resilience: - hard-kill the ACTIVE site-a node mid-query-loop → in-flight call returned TIMEOUT at exactly the 30s QueryTimeout deadline (bounded, no hang; deadline≠ retry — an in-flight call can't be safely re-sent) - next + all subsequent queries recovered automatically via site-a-b (SitePairChannelProvider NodeA→NodeB); site→central S&F never stopped (notif 619→715, still no dupes) - 0 PermissionDenied across all 8 nodes ExecuteOpcUa/ExecuteRoute/TriggerFailover deferred (no OPC-bound instance / no CLI verb; unit-proven). Rig left both-gRPC; git reverted to Akka default. --- ...6-07-22-clusterclient-to-grpc-live-gate.md | 51 +++++++++++++++++++ .../2026-07-22-clusterclient-to-grpc-plan.md | 2 +- ...2-clusterclient-to-grpc-plan.md.tasks.json | 2 +- 3 files changed, 53 insertions(+), 2 deletions(-) diff --git a/docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md b/docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md index c20e74a4..329e24ea 100644 --- a/docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md +++ b/docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md @@ -284,3 +284,54 @@ kill, zero loss/zero dupes). delivery, so the count is a faithful no-loss/no-dupes measure. - Rig left running on the gRPC build; git config reverted to Akka default (a redeploy from `main` resets to all-Akka). + +--- + +## Phase 3 — full central→site cutover + command matrix — **PASS** (2026-07-23) + +Central flipped to `SiteTransport=Grpc` (both central nodes; the flag is **central-wide** — +`CentralCommunicationActor.SelectTransport`). Sites kept on `CentralTransport=Grpc` from Phase 2, +so the rig ran **both directions on gRPC simultaneously** — the eventual all-gRPC end state. +Config bind-mounted (`:ro`), so a container `--force-recreate` (no image rebuild — code unchanged +since Phase 2) applied it. Central logged the selection at startup: +`central→site command transport: gRPC (SiteCommandService)`. + +The Phase 2 `SoakNotify` instances (#95–97) **survived the recreate as Enabled** (site volume +persisted `deployed_configurations` this time), so the notification workload kept flowing — +and lifecycle commands became drivable, closing the 1B/Phase-2 gap. + +### Checks + +| # | Check | Result | +|---|---|---| +| 1 | Central selects the gRPC command transport | **PASS** — `central→site command transport: gRPC (SiteCommandService)`; **0** `Created ClusterClient to` site lines on either central | +| 2 | `ExecuteQuery` (event-log) over gRPC, all 3 sites | **PASS** — all 3 returned correlationIds; site logs `SiteCommandService/ExecuteQuery - 200` (site-a-a/b-a/c-a each served the calls; standby nodes 0 — work lands on the active site node via the singleton proxy) | +| 3 | `ExecuteParked` (parked query) over gRPC, all 3 sites | **PASS** — all 3 returned `success` payloads over `SiteCommandService/ExecuteParked` | +| 4 | `ExecuteLifecycle` (disable→enable) over gRPC, site-a | **PASS** — `instance disable`/`enable #95` → `success:true`; site served `SiteCommandService/ExecuteLifecycle` (×4), instance state toggled — **NEW live coverage vs 1B/Phase-2 (needed a deployed instance)** | +| 5 | **Site-node kill mid-command → clean error, no hang** | **PASS** — killed the **active** site-a node (site-a-a) during a 1 s query loop: the in-flight call returned `TIMEOUT` at **dur=30.1 s = the `QueryTimeout` deadline** — bounded, not an indefinite hang (see note) | +| 6 | **Site-pair failover mid-stream** | **PASS** — the very next query (~32 s after kill) and all subsequent ones succeeded automatically via **site-a-b**; `SitePairChannelProvider` failed the gRPC channel NodeA→NodeB and the site singleton migrated; site→central S&F never stopped (notif count climbed 619→715 through the kill, still no dupes) | +| 7 | No PSK drift | **PASS** — **0** `PermissionDenied`/`Unauthenticated` across all 8 nodes for the whole run | +| 8 | Zero ClusterClient activity on the flipped path | **PASS** — central built no site ClusterClients; command routing is entirely `SiteCommandService` gRPC | + +### Note on check 5 (deadline vs fast-fail) + +The command in flight when site-a-a was **hard-killed** (SIGKILL) waited the full 30 s +`QueryTimeout` rather than failing fast on connect-refused: an already-dispatched gRPC call on a +dropped connection isn't observed as unsent, so it correctly cannot be auto-retried on the peer +node (it might have executed) and returns the deadline error to the caller — exactly the plan's +"deadline ≠ retry" rule. The deadline is the backstop; "no hang beyond deadline" is satisfied +(30.1 s). Only **provably-unsent** connect failures fail over fast, which is why every *subsequent* +call recovered immediately via site-a-b. + +### Not driven on this gate (unchanged from 1B/Phase 2) + +- **`ExecuteOpcUa`** (BrowseNode/ReadTagValues/WriteTag) — needs an OPC-bound deployed instance; + the bare rig's only deployable template (`SoakNotify`) has no data connection, and `Motor + Controller` needs 30 OPC UA bindings. Unit-proven (dispatcher routing ×28). +- **`ExecuteRoute`** (inbound-API → routed site script) — needs an inbound method + routing target + the rig lacks. +- **`TriggerFailover`** — no CLI verb (UI/management-only), destructive; unit-proven (ack-before- + `Leave` ordering tests). The hard-kill in check 6 is the live equivalent of a site-pair failover. + +Rig left running with **both** transports on gRPC; git config reverted to Akka default (a redeploy +from `main` resets to all-Akka). diff --git a/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md b/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md index 2f07a06d..7b976557 100644 --- a/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md +++ b/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md @@ -324,7 +324,7 @@ Critical path ≈ 1B: **~4–6 weeks total**, matching the design estimate. **Phase 2 ∥ 3 — cutover + soak** - [x] P2 All sites `CentralTransport=Grpc`; central-kill S&F soak (no loss/dupes), failback observed, health sequences clean — **PASS 2026-07-23** (live gate: single-node active-kill failover 29s + full-outage ~59s both-down drain; every 5s bucket = exactly 3 through both outages, 216 total == 216 distinct; 0 auth failures / 0 seq regressions; whole control plane — heartbeat/health/notification/audit — on gRPC `CentralControlService`) -- [ ] P3 Central `SiteTransport=Grpc` all sites; full UI command matrix per site; site-kill mid-command clean; no PSK noise; zero ClusterClient log activity on flipped paths +- [x] P3 Central `SiteTransport=Grpc` all sites; full UI command matrix per site; site-kill mid-command clean; no PSK noise; zero ClusterClient log activity on flipped paths — **PASS 2026-07-23** (live gate: `ExecuteQuery`/`ExecuteParked` all 3 sites + `ExecuteLifecycle` disable/enable on site-a, all 200 over `SiteCommandService`; active-site-node hard-kill mid-command → clean `TIMEOUT` at the 30s deadline, next call failed over to site-a-b automatically; 0 PermissionDenied / 0 ClusterClient-to-site on both central. `ExecuteOpcUa`/`ExecuteRoute`/`TriggerFailover` deferred — no OPC-bound instance / no CLI verb, unit-proven) **Phase 4 — deletion** - [ ] Defaults flip to `Grpc` + soak; then delete Akka transports, ClusterClient creation, `DefaultSiteClientFactory`, receptionist registrations, `CentralContactPoints`, then the flags diff --git a/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md.tasks.json b/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md.tasks.json index adf4308c..20a75e83 100644 --- a/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md.tasks.json +++ b/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md.tasks.json @@ -208,7 +208,7 @@ "id": "P3", "phase": "3", "subject": "Central SiteTransport=Grpc all sites; full UI command matrix; site-kill mid-command clean; zero ClusterClient activity", - "status": "pending", + "status": "completed", "activeForm": "Running the central->site cutover soak", "blockedBy": [ "P1B.DoD"