docs(grpc): Phase 3 live gate — PASS; central→site command cutover
Central flipped to SiteTransport=Grpc (both nodes, central-wide flag), sites kept on gRPC from Phase 2 → both directions on gRPC at once. Central logs 'central→site command transport: gRPC (SiteCommandService)'; 0 ClusterClient-to- site on either central. Command matrix over SiteCommandService (all 200): - ExecuteQuery (event-log) + ExecuteParked (parked) — all 3 sites - ExecuteLifecycle disable/enable #95 on site-a — NEW live coverage vs 1B/P2 (SoakNotify instances survived the recreate as Enabled) Resilience: - hard-kill the ACTIVE site-a node mid-query-loop → in-flight call returned TIMEOUT at exactly the 30s QueryTimeout deadline (bounded, no hang; deadline≠ retry — an in-flight call can't be safely re-sent) - next + all subsequent queries recovered automatically via site-a-b (SitePairChannelProvider NodeA→NodeB); site→central S&F never stopped (notif 619→715, still no dupes) - 0 PermissionDenied across all 8 nodes ExecuteOpcUa/ExecuteRoute/TriggerFailover deferred (no OPC-bound instance / no CLI verb; unit-proven). Rig left both-gRPC; git reverted to Akka default.
This commit is contained in:
@@ -324,7 +324,7 @@ Critical path ≈ 1B: **~4–6 weeks total**, matching the design estimate.
|
||||
|
||||
**Phase 2 ∥ 3 — cutover + soak**
|
||||
- [x] P2 All sites `CentralTransport=Grpc`; central-kill S&F soak (no loss/dupes), failback observed, health sequences clean — **PASS 2026-07-23** (live gate: single-node active-kill failover 29s + full-outage ~59s both-down drain; every 5s bucket = exactly 3 through both outages, 216 total == 216 distinct; 0 auth failures / 0 seq regressions; whole control plane — heartbeat/health/notification/audit — on gRPC `CentralControlService`)
|
||||
- [ ] P3 Central `SiteTransport=Grpc` all sites; full UI command matrix per site; site-kill mid-command clean; no PSK noise; zero ClusterClient log activity on flipped paths
|
||||
- [x] P3 Central `SiteTransport=Grpc` all sites; full UI command matrix per site; site-kill mid-command clean; no PSK noise; zero ClusterClient log activity on flipped paths — **PASS 2026-07-23** (live gate: `ExecuteQuery`/`ExecuteParked` all 3 sites + `ExecuteLifecycle` disable/enable on site-a, all 200 over `SiteCommandService`; active-site-node hard-kill mid-command → clean `TIMEOUT` at the 30s deadline, next call failed over to site-a-b automatically; 0 PermissionDenied / 0 ClusterClient-to-site on both central. `ExecuteOpcUa`/`ExecuteRoute`/`TriggerFailover` deferred — no OPC-bound instance / no CLI verb, unit-proven)
|
||||
|
||||
**Phase 4 — deletion**
|
||||
- [ ] Defaults flip to `Grpc` + soak; then delete Akka transports, ClusterClient creation, `DefaultSiteClientFactory`, receptionist registrations, `CentralContactPoints`, then the flags
|
||||
|
||||
Reference in New Issue
Block a user