diff --git a/docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md b/docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md index ff510384..c20e74a4 100644 --- a/docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md +++ b/docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md @@ -234,3 +234,53 @@ plan-vs-code finding, recorded rather than worked around. The instance-dependent matrix (tag ops, lifecycle, standby parked retry) and `TriggerSiteFailover` get their live exercise in Phase 3's full UI command matrix; the transport itself is proven here. + +--- + +## Phase 2 — full site→central cutover + S&F soak — **PASS** (2026-07-23) + +All three sites (**all 6 nodes**) flipped to `CentralTransport=Grpc` (edit to +`docker/site-*/appsettings.Site.json`, reverted in git after the gate — defaults stay Akka), +rebuilt from `main` + `--force-recreate`. Central left on `SiteTransport=Akka` (Phase 3 owns +that direction). The gate was driven by a **live S&F workload**, closing the notification/audit +path that 1A/1B could only unit-prove. + +### S&F driver + +A minimal dependency-free template `SoakNotify` (one 5 s `Interval` script: +`Notify.To("Engineering Alerts").Send(...)`), 3 instances deployed+enabled on site-a → a steady +**3 notifications per 5 s bucket** (the pre-existing `Motor Controller` "soak-motor" instances +need 30 OPC UA bindings and were unusable). The central `dbo.Notifications` table (one row per +`NotificationId`, insert-if-not-exists) is the no-loss/no-dupes source of truth. + +### Checks + +| # | Check | Result | +|---|---|---| +| 1 | All 6 site nodes on `CentralTransport=Grpc` | **PASS** — each logs `Site→central transport: gRPC to 2 central endpoint(s)`; **0** `PermissionDenied` across all six for the whole run | +| 2 | Full control plane rides gRPC `CentralControlService` | **PASS** — central sees `Heartbeat`, `ReportSiteHealth`, `SubmitNotification` (the S&F path), `IngestAuditEvents` — **196 RPCs / 90 s, 0 non-200** | +| 3 | Health/heartbeat cadence unchanged, no sequence regressions | **PASS** — `CentralHealthAggregator` logged **0** sequence-regression/out-of-order lines; heartbeat steady at ~144/30 s across 3 sites | +| 4 | **Single-node failover** — hard-kill the **active** central (central-a) | **PASS** — `CentralChannelProvider` logged sticky failover `central-a:8083 → central-b:8083` at the instant of kill; central-b active in **29 s** (auto-down); notif count froze at 72 during the gap | +| 5 | Buffer drains, no loss/dupes (single-node) | **PASS** — count resumed 72→101; **every 5 s bucket through the outage = exactly 3**, no gap; 101 total == 101 distinct | +| 6 | **Failback** — restart central-a | **PASS** — rejoined **ready in ~5 s as standby** (`active=503`); central-b **retained active** (oldest-Up, no role flap); traffic uninterrupted (uniform 3/bucket across failback); central-a singletons → `Younger` | +| 7 | **Full central outage** — hard-kill **both** central (~59 s) | **PASS** — count frozen at 155 for the entire outage; sites buffered continuously | +| 8 | Cold re-form + drain, no loss/dupes (both-down) | **PASS** — cold cluster re-formed, central-b active in **~14 s**; ~42 buffered notifications drained; **every 5 s bucket across the whole ~59 s both-dead window = exactly 3**, no gap; **216 total == 216 distinct** | + +Checks 4–8 also stand as a live preview of Phase 5 checks 4 (failover/failback) and 5 (mid-drain +kill, zero loss/zero dupes). + +### Notes + +- **Audit telemetry also rides the new gRPC `CentralControlService`** (`IngestAuditEvents`) — the + `SiteAuditTelemetryActor` "ClusterClientSiteAuditClient" label is legacy naming, not the wire + path. So `CentralTransport=Grpc` moves heartbeat, health, notification S&F **and** audit off + ClusterClient in one flip. +- Sites settled on **central-a** as the gRPC endpoint after the cold both-restart while **central-b** + held the active/singleton role — the gRPC endpoint node and the singleton host can differ; central-a + receives the forward and Akka-routes to the `NotificationOutboxActor` singleton on central-b. Both + are correct and independent. +- Some notifications land `Parked` at central (no SMTP config on the bare rig) — irrelevant to the + transport proof: the `Notifications` **row** is written on forward regardless of downstream SMTP + delivery, so the count is a faithful no-loss/no-dupes measure. +- Rig left running on the gRPC build; git config reverted to Akka default (a redeploy from `main` + resets to all-Akka). diff --git a/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md b/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md index a555a786..2f07a06d 100644 --- a/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md +++ b/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md @@ -323,7 +323,7 @@ Critical path ≈ 1B: **~4–6 weeks total**, matching the design estimate. - [x] 1B DoD (proportionate): rig central on `SiteTransport=Grpc` proves central→site rides authenticated gRPC `SiteCommandService` for all 3 sites (`ExecuteQuery`/`ExecuteParked` → 200, per-site PSK, 0 auth failures); rebased on 1A. Instance-dependent commands (tag ops/lifecycle/standby parked retry) + `TriggerSiteFailover` deferred to Phase 3 (no deployed instance / destructive UI-only); command-plane coexistence not expressible (`SiteTransport` is central-wide). Gate: `2026-07-22-clusterclient-to-grpc-live-gate.md` **Phase 2 ∥ 3 — cutover + soak** -- [ ] P2 All sites `CentralTransport=Grpc`; central-kill S&F soak (no loss/dupes), failback observed, health sequences clean +- [x] P2 All sites `CentralTransport=Grpc`; central-kill S&F soak (no loss/dupes), failback observed, health sequences clean — **PASS 2026-07-23** (live gate: single-node active-kill failover 29s + full-outage ~59s both-down drain; every 5s bucket = exactly 3 through both outages, 216 total == 216 distinct; 0 auth failures / 0 seq regressions; whole control plane — heartbeat/health/notification/audit — on gRPC `CentralControlService`) - [ ] P3 Central `SiteTransport=Grpc` all sites; full UI command matrix per site; site-kill mid-command clean; no PSK noise; zero ClusterClient log activity on flipped paths **Phase 4 — deletion** diff --git a/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md.tasks.json b/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md.tasks.json index 43218e43..adf4308c 100644 --- a/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md.tasks.json +++ b/docs/plans/2026-07-22-clusterclient-to-grpc-plan.md.tasks.json @@ -198,7 +198,7 @@ "id": "P2", "phase": "2", "subject": "All sites CentralTransport=Grpc; central-kill S&F soak, failback observed, health sequences clean", - "status": "pending", + "status": "completed", "activeForm": "Running the site->central cutover soak", "blockedBy": [ "P1A.DoD"