Files
lmxopcua/docs/plans/2026-07-23-mesh-phase4-live-gate.md
T
Joseph Doherty 1281aebfa7 docs(mesh-phase4): live gate PASSED — driver nodes run ConfigDb-free
Deploy sealed green with 4 DB-less site nodes acking (central persists acks);
ServiceLevel held 240 on a site node with central SQL stopped (survive-alone);
restarted site node booted last-known-good from the LocalDb pointer, no ConfigDb
read; alarm_condition_state table live + replicated. Full record in
2026-07-23-mesh-phase4-live-gate.md. Task 9 (drop dead table) deferred.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-23 14:03:28 -04:00

92 lines
5.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Per-Cluster Mesh Phase 4 — Live Gate Record
**Date:** 2026-07-23
**Branch:** `feat/mesh-phase4` (image built + rig recreated with Phase-4 code)
**Rig:** local `docker-dev` — central-1/2 (`admin,driver`, Direct + ConfigDb + ConfigServe), four site
nodes (`driver`, **no ConfigDb string**, `ConfigSource:Mode=FetchAndCache`). Deployment `b390f535`.
**Result: PASSED.** Every load-bearing exit-gate leg proven live. A driver-only node runs a full deploy
+ boot-from-cache cycle with **no ConfigDb connection string configured at all**, and the client-visible
ServiceLevel change behaves as designed.
## What was verified
### Leg 0 — boot ConfigDb-free (headline)
Every site node booted with no ConfigDb string and logged, e.g.:
```
DriverHost site-a-1:4053: FetchAndCache boot — restored served state for deployment 1841b0c2…
(rev 7a7d548e…) from the LocalDb pointer.
```
All six nodes came up clean (no FATAL / startup exception). The transient `RpcException Unavailable`
(central ConfigServe not yet up when the first fetch fired) and the `Name or service not known`
(the rig's deliberately-unresolvable historian endpoint) are expected startup noise — the node fell back
to the LocalDb pointer (#485-safe), never to an empty address space.
### Leg 1 — deploy seals green with site nodes acking, no ConfigDb ✅
`POST /api/deployments` (`X-Api-Key: docker-dev-deploy-key`) → `202 Accepted` (deployment `b390f535`).
It **sealed** (`Deployment.Status = 2`) at 17:57:33. Per-node `NodeDeploymentState`:
| Node | Status |
|---|---|
| central-1:4053 | 1 (Applied) |
| central-2:4053 | 1 (Applied) |
| **site-a-1:4053** | **1 (Applied)** |
| **site-a-2:4053** | **1 (Applied)** |
| **site-b-1:4053** | **1 (Applied)** |
| **site-b-2:4053** | **1 (Applied)** |
The four site nodes hold no ConfigDb — those ack rows were written by central's
`ConfigPublishCoordinator.PersistNodeAck` from the `ApplyAck` each site node sent over the transport.
Full deploy→fetch→apply→ack cycle proven ConfigDb-free.
### Leg 2 — ServiceLevel with central SQL DOWN (client-visible change) ✅
- Baseline (SQL up): site-a-1 (`opc.tcp://localhost:4842`) **ServiceLevel = 240** (healthy follower).
- `docker stop otopcua-dev-sql-1` → site-a-1 **ServiceLevel = 240, unchanged**.
A DB-less node feeds `DbReachable = true` constant with staleness from the redundancy-snapshot age only,
so it stays at full service when central SQL is unreachable — the survive-alone posture. A Direct/DB-backed
node would have dropped to 0/100 (its `DbHealthProbe` trips). Confirmed live: central-1's own probe went
`Reachable = False` while SQL was down (and back to `True` when it returned), while the site node never moved.
### Leg 3 — restart a site node with central SQL DOWN → boots last-known-good ✅
Restarted site-b-1 (the non-replicating site-b pin) while SQL was still stopped. Post-restart boot
(18:00:5859):
```
Cluster Node [akka.tcp://otopcua@site-b-1:4053] - Started up successfully
OPC UA server started on opc.tcp://0.0.0.0:4840
DriverHost site-b-1:4053: FetchAndCache boot — restored served state for deployment 1841b0c2…
(rev 7a7d548e…) from the LocalDb pointer.
```
No ConfigDb / SQL error in the fresh boot (it never dials SQL). Post-restart **ServiceLevel = 240** — it
served last-known-good from the LocalDb pointer with central SQL down.
### Leg 4 — grep proof ✅
Each site container's env: `ConnectionStrings__ConfigDb` count = **0**, `ConfigSource__Mode=FetchAndCache`.
Central-1 (contrast): carries `ConnectionStrings__ConfigDb` + `Mode=Direct`. Code sweep (Task 7) confirmed
no driver-branch service resolves `OtOpcUaConfigDbContext`.
### Leg 5 — alarm condition state in LocalDb ✅ (table live + replicated)
site-a-1's consolidated LocalDb (`/app/data/otopcua-localdb.db`) carries the `alarm_condition_state`
table with its replication capture triggers (`__localdb_*`), alongside `deployment_artifacts`,
`deployment_pointer`, and `alarm_sf_events`. 0 rows (no scripted alarm has fired in this window). The
store's save/load round-trip is unit-tested (`LocalDbAlarmConditionStateStoreTests`); a full operator
ack→row cycle was not driven live (needs a configured scripted alarm + a triggering value) — noted as the
one leg proven by test + schema-liveness rather than an end-to-end row write.
## Cleanup
`docker start otopcua-dev-sql-1` → healthy; all six nodes up; central-1 DbHealth recovered
(`Reachable=False → True` at 18:01:43). Rig healed.
## Diagnostics note (the "stuck gate")
The initial rig rebuild appeared hung: an abandoned `dotnet-dump analyze hang.dmp` (PID 94724, ~115 h of
CPU since Saturday) had pinned a core, starving the emulated `linux/amd64` `dotnet publish` (the Phase-3
protoc-segfault build-stage pin), and `docker compose --build`'s default `rawjson` progress streamed
nothing to stdout. Killing the runaway process + rebuilding with `--progress=plain` made the build both
faster and observable. No Phase-4 defect involved.
## Verdict
**Phase 4 live gate PASSED.** Driver-only nodes run ConfigDb-free end to end; the ServiceLevel semantics
change is live and correct. Remaining Phase-4 item: Task 9 (drop the dead ConfigDb `ScriptedAlarmState`
table + retire `EfAlarmConditionStateStore`) — deferred as plan-sanctioned (inert on all real nodes;
irreversible migration).