Files
ScadaBridge/docs/plans/2026-07-22-clusterclient-to-grpc-live-gate.md
T

199 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ClusterClient → gRPC migration — live gate results
Rig: `docker/` (2 central + 3×2 site + traefik), rebuilt from the branch under test via
`bash docker/deploy.sh`. Recorded check-by-check in the family's live-gate format. Phase 5's
full eight-check gate is recorded further down as those phases land; this file starts with
Phase 0, whose DoD has its own smaller gate.
---
## Phase 0 — PSK auth + dead-code removal — **PASS** (2026-07-22)
Branch `feat/grpc-phase0-psk` @ `228ff8b4`. Image rebuilt, all 9 containers recreated.
### Baseline (pre-change build, same rig)
An unauthenticated call from the host to a site's audit-pull RPC was **accepted**:
```
$ grpcurl -plaintext -d '{"batch_size":1}' localhost:9023 sitestream.SiteStreamService/PullAuditEvents
{}
```
That is the gap Phase 0 closes, reproduced rather than assumed.
### Checks
| # | Check | Result |
|---|---|---|
| 1 | All 8 nodes boot with keys configured (new `StartupValidator` rule) | **PASS** — all recreated and reached ready |
| 2 | Unauthenticated `PullAuditEvents``PermissionDenied`, all 3 sites | **PASS** |
| 3 | Wrong key (site-b's key presented to site-a) ⇒ `PermissionDenied` | **PASS** — per-site scoping is real, not decorative |
| 4 | Correct key ⇒ success, all 3 sites | **PASS** |
| 5 | Central's own authenticated paths still work | **PASS** — 14 successful `PullAuditEvents` from central to site-a; **0** auth failures in either central's log |
| 6 | LocalDb sync unaffected by the new interceptor | **PASS****0** control-plane rejections and **0** sync auth failures on the passive peer; session connected after one boot-order retry |
| 7 | No interceptor-activation errors | **PASS** — 0 (see the defect below) |
Evidence for 24:
```
=== NO CREDENTIALS ===
:9023 -> ERROR: Code: PermissionDenied Message: Control plane authentication failed.
:9033 -> ERROR: Code: PermissionDenied Message: Control plane authentication failed.
:9043 -> ERROR: Code: PermissionDenied Message: Control plane authentication failed.
=== WRONG KEY (site-b's key against site-a) ===
:9023 -> ERROR: Code: PermissionDenied Message: Control plane authentication failed.
=== CORRECT KEY ===
site-a :9023 -> {}
site-b :9033 -> {}
site-c :9043 -> {}
```
Site-a rejected exactly **2** calls — the two deliberate probes above — and nothing else.
### Defect the gate caught that the test suite did not
**First run of this gate FAILED**, and is worth recording because the failure mode is
deceptive.
`Grpc.AspNetCore` activates a type-registered interceptor through
`InterceptorRegistration.GetFactory()`, which throws when more than one public constructor is
applicable. `ControlPlaneAuthInterceptor` shipped with two — the DI one and a prefix-set
overload intended for later phases.
The throw happens **inside the pipeline, per call**, so:
- nothing failed at startup; the node booted, joined its pair and reported healthy;
- every gated call died with `Unknown / "Exception was thrown by handler"`, which reads as a
handler bug rather than an auth bug;
- **correct key, wrong key and no key produced identical errors** — the tell. A gate that
cannot distinguish those is not authenticating anything.
Site-a's log at the time: three `PullAuditEvents` calls, three identical
`System.InvalidOperationException: Multiple constructors accepting all given argument types
have been found in type 'ControlPlaneAuthInterceptor'`.
The full suite was green when this shipped — **29 suites, 6,872 tests, 0 failures**. The
in-process end-to-end test missed it because it registered the interceptor with
`AddSingleton` alongside `AddGrpc`, so DI returned the instance and gRPC's activation path
never ran.
Fixed in `228ff8b4`: the prefix-set constructor is `internal`; the end-to-end harness now
registers exactly as `Program.cs` does (by type, not in DI); and a reflection assertion pins
"exactly one public constructor", since that is the actual invariant.
**Lesson for phases 1A/1B, which both add services to this interceptor:** extend
`DefaultGatedPrefixes`; do not add a second public constructor. And any in-process harness for
a DI-activated component must mirror the production registration shape or it proves less than
it appears to.
### Test suite alongside the gate
Non-Playwright: **29 suites, 6,872 tests, 0 failures**.
Playwright (against this rig): **170 passed, 2 failed, 1 skipped** of 173. Both failures were
run down to root cause and **both are pre-existing on `main`, unrelated to Phase 0** — this
branch touches no EF, CentralUI, Transport or ManagementService file (`git diff --stat
main...HEAD -- src/` is 15 files, all Communication/Host/AuditLog gRPC plumbing).
An earlier run of this suite reported 44 failures. That run is **void**: a `docker/deploy.sh`
was recreating the cluster underneath it, so the fast `LoginTests`/`NavigationTests` failures
were "app unreachable", not defects.
**1. `TransportImportTests.ImportSyntheticBundle_AppliesAndShowsAuditDrillIn` — a real
production bug, not a test defect.** Central's log during the failure:
```
[ERR] An exception occurred while iterating over the results of a query ...
System.InvalidOperationException: The configured execution strategy
'SqlServerRetryingExecutionStrategy' does not support user-initiated transactions.
at Microsoft.EntityFrameworkCore.Query.Internal.SplitQueryingEnumerable`1.AsyncEnumerator.MoveNextAsync()
```
`BundleImporter.cs:1298` opens a user-initiated transaction; the central context is configured
with `EnableRetryOnFailure` (`ConfigurationDatabase/ServiceCollectionExtensions.cs:33`). SQL
Server's retrying strategy refuses to run a split query inside a caller's transaction, so
**bundle import fails against real MS SQL**. The fix is the one the exception names: wrap the
transaction in `Database.CreateExecutionStrategy().ExecuteAsync(...)`.
Why the whole unit/integration suite is green on it: those tests use the in-memory EF provider,
which has no retrying execution strategy — and `BeginTransactionAsync` is a no-op there. The
comment directly above line 1298 documents that divergence without drawing the conclusion. Only
a rig-backed test can see this.
**2. `SmsNotificationE2ETests.SmsConfigPage_CreateOrRender_NeverLeaksAuthToken` — a stale test
fixture.** No server-side error at all: the page renders 200, and no `INSERT INTO
SmsConfigurations` is ever issued. The test's fixture SID is `ACtest123` (`d6ead8ae`,
2026-06-19). `SmsConfiguration.razor:231` rejects anything not matching `^AC[0-9a-fA-F]{32}$`,
added by `40088a21` (2026-07-10) to close an un-escaped URI-interpolation hole. `Save()` sets
`_formError` and returns — no toast, exactly as observed. The fixture was never updated.
This has been failing since 2026-07-10, and it matters more than a red line: everything after
the toast assertion — including **the secret-non-leak assertion that the Auth Token value never
reaches the page HTML** — has not executed since. Fix is a valid 32-hex SID in the fixture.
### Not covered by this gate
- Streaming subscriptions were exercised in-process (TestServer), not over the rig. The
interceptor is path-scoped, not method-scoped, so the rig's `PullAuditEvents` evidence covers
the same code path — but a live `SubscribeInstance` under load is untested here.
- Key **rotation** on a live pair.
- `docker-env2` was updated with its own key but not redeployed or gated.
---
## Phase 1A — central control plane (site→central over gRPC) — **PASS** (2026-07-22)
Branch `feat/grpc-central-control` @ `0e162cb2`. Rig rebuilt; **site-a flipped to
`CentralTransport=Grpc`** with `CentralGrpcEndpoints=[central-a:8083, central-b:8083]`,
**site-b/c left on Akka** (default) to prove coexistence. The site-a flag flip was a
DoD-test-only rig edit — reverted from the branch, never committed (the plan keeps the default
`Akka` until Phase 4).
### The defect this gate caught (T1A.2 shipped it; fixed in `0e162cb2`)
**First rebuild: central's entire HTTP surface was gone.** central-a logged only
`Now listening on: http://[::]:8083` — no `:5000`. Central UI, the Management + Inbound API,
and every `/health/*` endpoint (Traefik routing + `IActiveNodeGate` both depend on them) were
dead. The node booted, joined the cluster and served gRPC fine; **no startup error.**
Cause: `builder.WebHost.ConfigureKestrel(o => o.ListenAnyIP(8083, Http2))` puts Kestrel into
explicit-endpoints mode, which **suppresses the URLs from `ASPNETCORE_URLS`/`--urls`** — it is
not additive, contrary to the comment T1A.2 shipped. Central's whole HTTP/1 surface lives on
that URL (`http://+:5000` on the rig; a different port in production). The site branch has the
same `ConfigureKestrel` shape but nothing on `ASPNETCORE_URLS` to lose — it binds every port it
needs explicitly — which is why the pattern looked safe.
Every unit + E2E test uses `TestServer`, which never binds real Kestrel, so the whole suite
(6,872) stayed green. Only a live node exposes a missing listener. Fix: parse the port(s) from
the configured URLs and re-declare them (`Http1AndHttp2`) alongside the gRPC port (`Http2`) in
the one `ConfigureKestrel` call — `Program.ParseHttpBindPorts` + `CentralHttpBindPortsTests`
(14 cases). After the fix: `Now listening on: http://[::]:5000` **and** `:8083`; 9001 ready
`200`/active Healthy, 9002 standby, Traefik LB `200`, CLI over the LB works.
### Checks (post-fix rebuild)
| # | Check | Result |
|---|---|---|
| 1 | site-a rides authenticated gRPC to `CentralControlService` | **PASS** — Heartbeat, `ReportSiteHealth`, `ReconcileSite` all HTTP/2 → 200; **0** auth failures on either central |
| 2 | Heartbeat drives the active flag | **PASS** — 144+ heartbeats over gRPC; site-a `online=True` at central |
| 3 | Health page live | **PASS**`ReportSiteHealth` lands; central shows site-a `online=True`, sequence advancing, alongside Akka site-b/c |
| 4 | Reconcile works after site restart | **PASS** — restarted `site-a-a`; it logged `Site→central transport: gRPC to 2 central endpoint(s)`, then `Reconcile pass … complete: 0 fetched, 0 failed, 0 orphan(s)` |
| 5 | Coexistence | **PASS** — site-b/c log `Created ClusterClient to central`; both `online=True` — Akka and gRPC sites side by side |
| 6 | Central HTTP surface intact under the new gRPC listener | **PASS** — after the fix (see above) |
### Not independently exercised on this gate
- **Notification e2e (`SubmitNotification`/`QueryNotificationStatus`) and audit ingest
(`IngestAuditEvents`/`IngestCachedTelemetry`)** were NOT driven live: the rig has templates
but **no deployed instance**, so nothing emits site→central notifications or audit rows on its
own. These four RPCs traverse the identical `CentralControlGrpcService`
`CentralCommunicationActor` Ask path that checks 13 proved live under real auth, and their
payload mappers carry 32 round-trip goldens — but the payloads themselves were not put over
the wire here. **Phase 2's central-kill S&F soak is where they get their live workout;** flag
for a fuller 1A proof if a deployed-instance rig is set up before then.
- Cross-node failover/failback of the site→central channel under a central-node kill (unit-proven
via TestServer; not exercised on the rig at 1A).