docs(grpc): Phase 2 live gate — PASS; site→central S&F cutover soak
All 6 site nodes flipped to CentralTransport=Grpc; whole control plane (heartbeat/health/notification S&F/audit) rides gRPC CentralControlService, 196 RPCs/90s 0 non-200, 0 PermissionDenied, 0 health-sequence regressions. Soak driven by a live SoakNotify workload (3 instances, 3 notifs/5s): - single-node active-kill (central-a): sticky failover central-a:8083→b:8083 logged instantly, central-b active in 29s, buffer drained 72→101, every 5s bucket = exactly 3 through the gap, 101==101 distinct - failback: central-a rejoined ready ~5s as standby, central-b kept active (oldest-Up, no flap), traffic uninterrupted - full outage (both central down ~59s): count frozen, cold re-form central-b active ~14s, ~42 buffered drained, every 5s bucket = exactly 3 across the whole dead window, 216==216 distinct — zero loss, zero dupes Also previews Phase 5 checks 4 (failover/failback) and 5 (mid-drain kill). Rig config reverted to Akka default (defaults stay Akka until Phase 4).
This commit is contained in:
@@ -323,7 +323,7 @@ Critical path ≈ 1B: **~4–6 weeks total**, matching the design estimate.
|
||||
- [x] 1B DoD (proportionate): rig central on `SiteTransport=Grpc` proves central→site rides authenticated gRPC `SiteCommandService` for all 3 sites (`ExecuteQuery`/`ExecuteParked` → 200, per-site PSK, 0 auth failures); rebased on 1A. Instance-dependent commands (tag ops/lifecycle/standby parked retry) + `TriggerSiteFailover` deferred to Phase 3 (no deployed instance / destructive UI-only); command-plane coexistence not expressible (`SiteTransport` is central-wide). Gate: `2026-07-22-clusterclient-to-grpc-live-gate.md`
|
||||
|
||||
**Phase 2 ∥ 3 — cutover + soak**
|
||||
- [ ] P2 All sites `CentralTransport=Grpc`; central-kill S&F soak (no loss/dupes), failback observed, health sequences clean
|
||||
- [x] P2 All sites `CentralTransport=Grpc`; central-kill S&F soak (no loss/dupes), failback observed, health sequences clean — **PASS 2026-07-23** (live gate: single-node active-kill failover 29s + full-outage ~59s both-down drain; every 5s bucket = exactly 3 through both outages, 216 total == 216 distinct; 0 auth failures / 0 seq regressions; whole control plane — heartbeat/health/notification/audit — on gRPC `CentralControlService`)
|
||||
- [ ] P3 Central `SiteTransport=Grpc` all sites; full UI command matrix per site; site-kill mid-command clean; no PSK noise; zero ClusterClient log activity on flipped paths
|
||||
|
||||
**Phase 4 — deletion**
|
||||
|
||||
Reference in New Issue
Block a user