Files
lmxopcua/docs/plans/2026-07-23-mesh-phase4-cut-driver-configdb.md
T

32 KiB
Raw Blame History

Per-Cluster Mesh Phase 4 — Cut the Driver-Side ConfigDb Connection

For Claude: REQUIRED SUB-SKILL: Use superpowers-extended-cc:executing-plans to implement this plan task-by-task.

Goal: A driver-only OtOpcUa node runs a full deploy + alarm + historian cycle with no ConfigDb connection string configured at all — LocalDb is its steady-state config + alarm-state store, and central persists its deploy acks. Grep-level proof: no driver-branch service can resolve OtOpcUaConfigDbContext.

Architecture: Phase 3 already gave driver nodes their config from central (fetch-and-cache into LocalDb) behind the ConfigSource:Mode=FetchAndCache dark switch. Phase 4 removes the remaining four driver-side ConfigDb consumers and the registration itself, so a driver-only host boots with the ConfigDb connection string absent. The fused central node keeps ConfigDb via its admin role — this phase gates the whole registration on hasAdmin, never on hasDriver.

Tech stack: .NET 10, Akka.NET, EF Core (ConfigDb, SQL Server — central only after this phase), ZB.MOM.WW.LocalDb (SQLite, replicated), xUnit + Shouldly.


Design decisions (settled — do not relitigate)

  1. ConfigDb is registered iff hasAdmin. Program.cs:126's unconditional AddOtOpcUaConfigDb moves inside a hasAdmin guard. A driver-only node requires no ConfigDb connection string and must not throw for its absence. Fused admin,driver central nodes are unaffected (admin owns the DB; driver actors on that node may still be handed a non-null factory).
  2. Driver-only ⇒ FetchAndCache is mandatory. Direct mode reads central SQL, which a driver-only node no longer has. ConfigSourceOptionsValidator fails start for a driver-role, non-admin node whose Mode is Direct. (A fused admin,driver node may stay Direct — it has the DB.)
  3. Db-reachability retires from ServiceLevel on driver-only nodes (user decision 2026-07-23): LocalDb is the config store and is in-process, so DbReachable is treated as a constant true on nodes with no DbHealthProbeActor. Staleness already carries "running behind on config." A healthy site node serving last-known-good with central down publishes 240/250 — the survive-alone posture. This is a client-visible ServiceLevel semantics change — documented in docs/Redundancy.md + the interop note.
  4. Central is the system of record for deploy acks. ConfigPublishCoordinator.PersistNodeAck already writes NodeDeploymentState from every ApplyAck it receives over the transport, and seeds the Applying rows at dispatch. So DriverHostActor's direct UpsertNodeDeploymentState SQL writes are redundant and are removed. The bootstrap NodeDeploymentState read is already replaced by the LocalDb deployment_pointer for FetchAndCache (Phase 3).
  5. ScriptedAlarmState (Part 9 condition state) → LocalDb, wholesale for driver-role nodes. It is read only by EfAlarmConditionStateStore; AdminUI /alerts uses live SignalR, not the table. A new replicated alarm_condition_state LocalDb table + LocalDbAlarmConditionStateStore replaces the EF store on every driver-role node (fused central included), same journey Phase 2 made for the alarm S&F buffer. The dead ConfigDb ScriptedAlarmState table is dropped in a final trivial migration.
  6. OpcUaPublishActor._dbFactory becomes null on driver-only nodes. Its only ConfigDb use is LoadArtifact/LoadLatestArtifact (the #485 artifact-blob reads); FetchAndCache already passes msg.Artifact in-hand so the fallback never fires. Phase 4 asserts the null-factory path never reads ConfigDb and every driver-only rebuild carries bytes.

The five consumers (design §6.1 audit) and their Phase-4 disposition

Consumer Current ConfigDb use Phase-4 disposition Task
DriverHostActor UpsertNodeDeploymentState writes (×9), bootstrap NodeDeploymentState read, EfAlarmConditionStateStore ctor drop SQL ack-writes (central persists); nullable factory; LocalDb condition store 4, 56
EfAlarmConditionStateStore Part 9 condition state in ConfigDb ScriptedAlarmState LocalDbAlarmConditionStateStore (new replicated table) 5, 6
DbHealthProbeActor SELECT 1 vs ConfigDb → DbReachable ServiceLevel input not spawned on driver-only; DbReachable=true constant 2, 3
OpcUaPublishActor LoadArtifact/LoadLatestArtifact nullable factory; rebuild always carries bytes 3
ServiceCollectionExtensions (Runtime + Configuration) + Program.cs registration hasAdmin-gated; nullable threading 1, 7

Task 0: ConfigSourceOptionsValidator — driver-only ⇒ FetchAndCache

Classification: small Estimated implement time: ~4 min Parallelizable with: Task 5

Files:

  • Modify: src/Core/ZB.MOM.WW.OtOpcUa.Cluster/ConfigSourceOptionsValidator.cs
  • Test: tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/Configuration/ConfigSourceOptionsValidatorTests.cs (or the existing validator test file — locate with grep -rl ConfigSourceOptionsValidator tests)

Context: The validator today checks the FetchAndCache fields (endpoints/key). Phase 4 adds the role-conditional rule. The validator must be given the node's roles — check how it reads them today (likely IConfiguration["Cluster:Roles"] or an injected role view); mirror AkkaClusterOptionsValidator, which already matches on roles at boot.

Step 1: Write the failing test — a driver-role, non-admin node with Mode=Direct → validation fails with a message naming the flag; a fused admin,driver node with Mode=Direct → passes; a driver-only node with Mode=FetchAndCache → passes.

Step 2: Run it, confirm it fails (rule not yet present).

Step 3: Implement the rule: roles.Contains("driver") && !roles.Contains("admin") && Mode==DirectValidateOptionsResult.Fail(...). Read roles the same way the validator/AddValidatedOptions wiring already does; do not invent a new source.

Step 4: Run tests, confirm pass.

Step 5: Commitfeat(mesh-phase4): driver-only node must be FetchAndCache (validator).


Task 1: Program.cs — gate ConfigDb registration + health check on hasAdmin

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: none (boot path; other tasks build on the nullable factory it introduces)

Files:

  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Host/Program.cs:125-126 (the unconditional AddOtOpcUaConfigDb)
  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Host/Health/HealthEndpoints.cs:27-31 (DatabaseHealthCheck<OtOpcUaConfigDbContext> — driver-only has no ConfigDb to probe)
  • Test: tests/Server/ZB.MOM.WW.OtOpcUa.Host.IntegrationTests/DriverOnlyNoConfigDbBootTests.cs (create)

Context: AddOtOpcUaConfigDb throws if the ConfigDb connection string is absent. Line 126 comment says "always registered regardless of role. ConfigDb is required for everything" — that statement is what Phase 4 falsifies. Gate the call on hasAdmin. Anything downstream that resolves IDbContextFactory<OtOpcUaConfigDbContext> non-optionally must already be hasAdmin-only or become nullable (Tasks 27 handle the driver-branch resolvers). Grep GetRequiredService<...ConfigDbContext...> and GetRequiredService<IDbContextFactory<OtOpcUaConfigDbContext>> across the Host to find any that would now throw on a driver-only node; each must move under hasAdmin or switch to GetService.

Step 1: Write the failing test — build a Host WebApplication with roles driver only and no ConfigDb connection string; assert it builds + starts without throwing, and that services.GetService<IDbContextFactory<OtOpcUaConfigDbContext>>() is null. (Model the harness on DeploymentArtifactFetchBoundaryTests.StartServerAsync — minimal builder, in-memory where needed.)

Step 2: Run it — fails today (AddOtOpcUaConfigDb throws "Connection string 'ConfigDb' is required").

Step 3: Implement — wrap AddOtOpcUaConfigDb in if (hasAdmin); wrap the ConfigDb health check in if (hasAdmin); sweep the GetRequiredService sites found above. Add a boot log line on a driver-only node: "driver-only node — no ConfigDb registered; config via ConfigSource:Mode=FetchAndCache".

Step 4: Run the test + dotnet build — confirm pass and no unresolved-service throw.

Step 5: Commitfeat(mesh-phase4): register ConfigDb only on admin-role nodes.


Task 1b: driver-only ILdapGroupRoleMappingService — empty (appsettings-only group→role) — DISCOVERED during Task 1

Classification: small Estimated implement time: ~4 min Parallelizable with: Task 5

Files:

  • Create: src/Server/ZB.MOM.WW.OtOpcUa.Security/Ldap/NullLdapGroupRoleMappingService.cs (or a Host-side driver-only impl — place it beside OtOpcUaGroupRoleMapper)
  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Host/Program.cs:317-327 (the hasDriver auth block — register the empty service with TryAddScoped so a fused node keeps the real EF one; fix the now-false comment at 322-323)
  • Test: tests/Server/ZB.MOM.WW.OtOpcUa.Security.Tests/ (or wherever OtOpcUaGroupRoleMapper is tested)

Context — a 6th ConfigDb consumer the §6.1 audit missed. OtOpcUaGroupRoleMapper (the OPC UA data-path group→role mapper, IGroupRoleMapper<string>, registered at Program.cs:327) merges the appsettings Security:Ldap:GroupToRole baseline (always available) with system-wide DB grants from ILdapGroupRoleMappingService — which is registered only inside AddOtOpcUaConfigDb. Task 1 gated that on hasAdmin, so a driver-only node now cannot resolve IGroupRoleMapper<string> at OPC UA login (lazy scoped resolve — Build() does not surface it; the auth path does). LdapGroupRoleMapping rows are not in the deployment artifact (ConfigComposer never composes them), so driver nodes never received DB grants via config in the first place. The honest fix: a driver-only ILdapGroupRoleMappingService whose GetByGroupsAsync returns empty (CRUD methods throw NotSupportedException — never called with no AdminUI), registered via TryAddScoped in the hasDriver block so the fused node's real EF impl still wins. Behavior reduction to document (Task 8): DB-backed LDAP group→role grants (the AdminUI RoleGrants page) apply only where ConfigDb is present (admin/fused nodes); driver-only site nodes map groups→roles from Security:Ldap:GroupToRole appsettings only.

Step 1: Write the failing test — resolve IGroupRoleMapper<string> in a driver-only service graph and call MapAsync(["some-group"]) → returns the appsettings baseline with no throw (today it throws: ILdapGroupRoleMappingService unresolvable).

Step 2: Run — fails.

Step 3: Implement the empty service + TryAddScoped registration in the hasDriver block; correct the Program.cs:322-323 comment to state the driver-only node uses the empty impl and why.

Step 4: Run tests + build.

Step 5: Commitfix(mesh-phase4): driver-only nodes map LDAP groups from appsettings (no ConfigDb grants).


Task 2: DbHealthProbeActor — do not spawn on driver-only; DbReachable=true constant

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: none (ServiceLevel; Task 3 depends on its message contract)

Files:

  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ServiceCollectionExtensions.cs:329-332 (spawn site)
  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Runtime/OpcUa/OpcUaPublishActor.cs:641-655 (the _dbHealthProbe is null seam + the Stale derivation)
  • Test: tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/OpcUa/OpcUaPublishActorServiceLevelTests.cs (extend the existing ServiceLevel test file)

Context: ServiceLevelCalculator.Compute collapses (DbReachable=false, probe ok, not stale) → 0. The retired-input rule (decision 3): on a node with _dbHealthProbe is null, feed DbReachable: true and derive Stale without the DB-health term — Stale = (now - snapshotEntry.AsOfUtc) > _staleWindow || runningBehindConfig. Do NOT fall through to LegacyRoleOnly for driver-only nodes: that seam exists for the pre-first-sample bootstrap window on a DB-backed node, and would peg a healthy site node at the role-only mapping forever. Introduce an explicit "DB-less node" branch that computes the full NodeHealthInputs with DbReachable=true and the DB-free staleness. runningBehindConfig = the Phase-3 running-from-cache / last-apply-failed signal if reachable from the publish actor; if it is not currently threaded here, use the snapshot-age term alone and record a TODO(Phase 4 follow-up) — do not thread a new dependency just for this in-scope.

Step 1: Write failing tests — (a) DB-less node, member Up, probe ok, snapshot fresh → 240 (+10 if Primary); (b) DB-less node, snapshot stale (>window) → 200; (c) DB-backed node unchanged (regression). Assert against OpcUaPublishActor's computed ServiceLevelChanged, not the calculator in isolation.

Step 2: Run — (a)/(b) fail (today a DB-less node hits LegacyRoleOnly or 0).

Step 3: Implement — in ServiceCollectionExtensions, spawn DbHealthProbeActor only when dbFactory is not null; register DbHealthProbeActorKey only then; pass dbHealthProbe: null into OpcUaPublishActor.Props on a DB-less node. In OpcUaPublishActor, add the DB-less branch described above (guard on "no probe wired AND this is a driver-only/DB-less node" — a clean signal is _dbHealthProbe is null combined with a ctor flag dbLess, threaded from the spawn site, so the legacy-bootstrap seam stays reachable only on DB-backed nodes before their first sample).

Step 4: Run tests + build — confirm pass.

Step 5: Commitfeat(mesh-phase4): retire DB-reachability from ServiceLevel on DB-less nodes.


Task 3: OpcUaPublishActor — nullable factory; driver-only rebuild always carries bytes

Classification: standard Estimated implement time: ~4 min Parallelizable with: none (same file as Task 2 — sequence after it)

Files:

  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Runtime/OpcUa/OpcUaPublishActor.cs
  • Test: tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/OpcUa/OpcUaPublishActorTests.cs

Context: _dbFactory is already nullable on this actor; HandleRebuild's _dbFactory is null || _applier is null guard falls back to a raw _sink.RebuildAddressSpace(). That fallback is the wrong path for a driver-only node that HAS an applier but no factory — it would skip the diff-and-apply. The correct driver-only invariant: _applier is wired, _dbFactory is null, and every RebuildAddressSpace carries msg.Artifact in-hand (Phase-3 DriverHostActor already does this). Split the guard so a null factory with an applier still does the diff-and-apply pass, using msg.Artifact, and asserts (logs an Error + abandons per #485) if msg.Artifact is null on such a node — a null artifact with no factory to fall back to is "no answer," never an empty rebuild.

Step 1: Write failing tests — (a) applier wired, factory null, RebuildAddressSpace with Artifact → diff-and-apply runs (assert applier invoked, not the raw-sink fallback); (b) applier wired, factory null, RebuildAddressSpace with null Artifact → no rebuild, Error logged (#485 negative).

Step 2: Run — (a) fails (null factory currently routes to the raw-sink fallback).

Step 3: Implement — restructure HandleRebuild: if (_applier is null) { raw-sink fallback; return; } then the diff-and-apply block, where artifact = msg.Artifact ?? (_dbFactory is not null ? (depId path else latest) : Array.Empty<byte>()); the existing #485 zero-length guard then abandons the null-factory

  • null-artifact case cleanly.

Step 4: Run tests + build.

Step 5: Commitfeat(mesh-phase4): OpcUaPublish diff-applies from in-hand bytes with no ConfigDb.


Task 4: DriverHostActor — nullable factory; drop redundant SQL ack-writes on DB-less nodes

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: none (core actor; Tasks 56 build the store it constructs)

Files:

  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Runtime/Drivers/DriverHostActor.cs (_dbFactory field/ctor/Props → nullable; the ~9 UpsertNodeDeploymentState call sites; the bootstrap NodeDeploymentState read at 606-612; the EfAlarmConditionStateStore ctor at 588; the CreateDbContext sites at 1920, 2096, 2588)
  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ServiceCollectionExtensions.cs:455 (Props call)
  • Test: tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/Drivers/DriverHostActorDbLessTests.cs (create)

Context: Decision 4 — central's ConfigPublishCoordinator.PersistNodeAck is the ack system of record. On a DB-less node UpsertNodeDeploymentState has no ConfigDb to write and its row is written by central anyway. Guard every UpsertNodeDeploymentState body on _dbFactory is not null (make the method a no-op when null) rather than deleting call sites — keeps the fused-node behaviour identical and the diff small. Same for the bootstrap read (606-612) and the Direct-only CreateDbContext sites (1920/2096/2588): each is already unreachable on a FetchAndCache node (Phase 3 routes bootstrap through the LocalDb pointer and apply through the fetcher), so the change is defensive nullability, not new control flow. Verify each of 1920/2096/2588 is inside a Direct-mode-only path before guarding — if any is on the shared apply path, that is a Phase-3 gap to surface, not silently guard.

Two specific sites the Task-2 code review surfaced (MUST be handled here):

  • Remove the interim dbFactory! at ServiceCollectionExtensions.cs:~470 (added in Task 2 to preserve compilation) — once DriverHostActor._dbFactory is nullable, pass the plain (nullable) dbFactory.
  • UpsertNodeDeploymentState (2613-2643) is called UNCONDITIONALLY on the FetchAndCache apply path (BeginFetchAndCacheApply/HandleFetchedForApply, ~1716/1723/1774/1785/1801). Its try/catch swallows the null-factory NRE and logs a misleading "transient DB" warning on every deploy on a driver-only node. The _dbFactory is not null guard fixes this.
  • SpawnScriptedAlarmHost (called unconditionally from PreStart ~524) constructs new EfAlarmConditionStateStore(_dbFactory, ...) at ~587-588 — with a null factory every alarm Load/Save NREs (swallowed → scripted-alarm state silently dead on driver-only nodes). This is fixed by Task 6 (which swaps in the LocalDb-backed store), NOT by a null guard — do NOT leave the Ef store constructed with a null factory. If Task 6 lands after this task, add a temporary _dbFactory is not null guard around the Ef-store construction so a DB-less node skips it cleanly until Task 6 wires the LocalDb store; the two tasks together must leave a DB-less node with a WORKING (LocalDb) condition store, not a skipped one.

Step 1: Write failing tests — construct a DriverHostActor with dbFactory: null, fetchAndCacheMode: true, a fake fetcher + LocalDb cache (mirror DriverHostActorFetchAndCacheTests): (a) a full dispatch → Applied ack, no NullReferenceException from a factory deref; (b) bootstrap with a seeded LocalDb pointer → revision restored with no ConfigDb read.

Step 2: Run — fails today (_dbFactory non-nullable; UpsertNodeDeploymentState derefs it).

Step 3: Implement — make _dbFactory nullable through field/ctor/both Props overloads (nullable param last, matching the Phase-3 additions); UpsertNodeDeploymentState early-returns when null (log Debug once); guard the read sites. Update ServiceCollectionExtensions:455 to pass the (already nullable) dbFactory.

Step 4: Run tests + build.

Step 5: Commitfeat(mesh-phase4): DriverHostActor runs with no ConfigDb (central persists acks).


Task 5: alarm_condition_state LocalDb table + LocalDbAlarmConditionStateStore

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 0

Files:

  • Create: src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ScriptedAlarms/LocalDbAlarmConditionStateStore.cs
  • Modify: the LocalDb schema/DDL + RegisterReplicated wiring (find via grep -rn "alarm_sf_events" src — mirror that table's registration exactly, in the same LocalDbSetup.OnReady DDL→RegisterReplicated order which is load-bearing)
  • Test: tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/ScriptedAlarms/LocalDbAlarmConditionStateStoreTests.cs (port EfAlarmConditionStateStoreTests — same IAlarmStateStore contract, so the assertion bodies transfer verbatim; only the store construction changes)

Context: LocalDbAlarmConditionStateStore : IAlarmStateStore must round-trip the same fields EfAlarmConditionStateStore does (Enabled/Acked/Confirmed/Shelving + ack/confirm audit + CommentsJson; Active/LastActive/LastCleared transient; LastTransitionUtc ↔ UpdatedAtUtc). Columns mirror ScriptedAlarmState. Replicated like alarm_sf_events so a failover peer holds the state — but unlike the S&F buffer there is no drain gate: condition state is not delivered-once, it is the current snapshot, idempotent last-write-wins on UpdatedAtUtc (the Ef store already tolerates the write race that way). Do NOT stack app-level LWW on LocalDb's HLC (design §8) — the row upsert keyed by ScriptedAlarmId is a single unconditional write, which is HLC-safe; keep it that way.

Step 1: Write failing tests — the ported IAlarmStateStore round-trip suite against the LocalDb store over a temp SQLite DB (the Runtime.Tests harness already opens LocalDb elsewhere — reuse it).

Step 2: Run — fails (store + table absent).

Step 3: Implement — the store + the alarm_condition_state DDL + RegisterReplicated registration.

Step 4: Run tests + build.

Step 5: Commitfeat(mesh-phase4): LocalDb alarm-condition-state store (replicated, pair-local).


Task 6: Wire the LocalDb condition store into DriverHostActor for driver-role nodes

Classification: standard Estimated implement time: ~4 min Parallelizable with: none (depends on Task 5's type + Task 4's nullable factory)

Files:

  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Runtime/Drivers/DriverHostActor.cs:588 (construct LocalDbAlarmConditionStateStore instead of EfAlarmConditionStateStore)
  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ServiceCollectionExtensions.cs (resolve the LocalDb store dependency; thread it into DriverHostActor.Props)
  • Test: extend DriverHostActorDbLessTests — the ScriptedAlarm host constructs against LocalDb with a null ConfigDb factory.

Context: The store is constructed inside DriverHostActor.BuildScriptedAlarmHost (line ~588). It takes the LocalDb connection the node already opened (Phase 1/2). Prefer resolving an IAlarmStateStore from DI (registered in the Host's hasDriver LocalDb block) over new-ing it in the actor, so the actor stays factory-agnostic. Confirm the registration lands in the same AddOtOpcUaLocalDb hasDriver branch as the S&F sink.

Step 1: Write failing test — DB-less DriverHostActor boots its ScriptedAlarm host and a condition save/load round-trips through LocalDb (no ConfigDb).

Step 2: Run — fails (still constructs the Ef store needing a factory).

Step 3: Implement — register IAlarmStateStore → LocalDbAlarmConditionStateStore in the Host hasDriver LocalDb block; resolve + thread it; delete the EfAlarmConditionStateStore construction from the actor. Leave the EfAlarmConditionStateStore class in the tree (removed in Task 9's migration cleanup) or delete it now if nothing else references it (grep first).

Step 4: Run tests + build.

Step 5: Commitfeat(mesh-phase4): driver-role nodes store alarm condition state in LocalDb.


Task 7: Registration cleanup + full-graph sweep

Classification: standard Estimated implement time: ~5 min Parallelizable with: none (integration seam over everything above)

Files:

  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ServiceCollectionExtensions.cs (final nullable threading)
  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Host/Program.cs (any remaining driver-branch ConfigDb resolves)
  • Test: tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/ServiceCollectionExtensionsTests.cs

Context: The exit-gate proof. After Tasks 16, sweep for any service the driver-only graph resolves that still needs OtOpcUaConfigDbContext. Grep both projects for OtOpcUaConfigDbContext, IDbContextFactory<OtOpcUaConfigDbContext>, AddOtOpcUaConfigDb, ScriptedAlarmStates, and GetRequiredService of any of them; each hit must be hasAdmin-gated or nullable-tolerant. LdapGroupRoleMappingService / ILdapGroupRoleMappingService (registered inside AddOtOpcUaConfigDb) is admin/auth-side — confirm the driver-only login path does not resolve it, or that the driver-only node never serves login (it hosts no AdminUI — see the topology table).

Step 1: Write the failing test — a "driver-only DI graph resolves nothing ConfigDb-bound" test: build the hasDriver && !hasAdmin service provider and assert GetService<IDbContextFactory<OtOpcUaConfigDbContext>>() is null AND that spawning the runtime actors does not throw. Also pin the Task-1b fused-node ordering (Task 1b review gap): build the hasAdmin && hasDriver graph through the same AddOtOpcUaConfigDb + hasDriver TryAdd sequence and assert GetService<ILdapGroupRoleMappingService>() resolves the real LdapGroupRoleMappingService, not the Null… — so a later reorder of the role blocks (or switching AddOtOpcUaConfigDb to a TryAdd) fails a red test instead of silently stripping DB-backed role grants from central. Also resolve IGroupRoleMapper<string> in a driver-only scope to catch the lazy auth-path break Task 1b fixed.

Step 2: Run — fix whatever it catches.

Step 3: Implement the remaining guards.

Step 4: Run the full Runtime.Tests + Host.IntegrationTests suites + dotnet build.

Step 5: Commitfeat(mesh-phase4): driver-only DI graph resolves nothing ConfigDb-bound.


Task 8: Docs — Redundancy.md (ServiceLevel change) + interop note + Configuration.md + CLAUDE.md

Classification: small Estimated implement time: ~5 min Parallelizable with: Task 9

Files:

  • Modify: docs/Redundancy.md (ServiceLevel: DB-less node table + the DbReachable=true constant + the 240-with-central-down behaviour; the DbHealthProbeActor row becomes admin-only)
  • Modify: docs/Configuration.md (driver-only nodes: no ConfigDb string; ConfigSource:Mode must be FetchAndCache; validator rule)
  • Modify: CLAUDE.md (extend the mesh section — Phase 4: driver ConfigDb cut; ServiceLevel semantics; alarm condition state in LocalDb)
  • Modify: docs/plans/2026-07-22-per-cluster-mesh-program.md (Phase 4 status + tracking row)

Step 14: Prose only — state the client-visible ServiceLevel change explicitly (a healthy DB-less node now publishes 240/250 with central unreachable, where a Direct node would have dropped to 0/100).

Step 5: Commitdocs(mesh-phase4): ServiceLevel semantics + driver ConfigDb cut.


Task 9: Drop the dead ConfigDb ScriptedAlarmState table (migration) + retire EfAlarmConditionStateStore

Classification: standard Estimated implement time: ~4 min Parallelizable with: none (after Task 6 proves LocalDb store is the live path)

Files:

  • Create: EF migration dropping ScriptedAlarmState (find the migrations dir via grep -rl "class OtOpcUaConfigDbContext" → its Migrations/ sibling; use dotnet ef migrations add DropScriptedAlarmState -c OtOpcUaConfigDbContext)
  • Modify: src/Core/ZB.MOM.WW.OtOpcUa.Configuration/OtOpcUaConfigDbContext.cs, .../Entities/ScriptedAlarmState.cs (remove the DbSet + entity)
  • Delete: src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ScriptedAlarms/EfAlarmConditionStateStore.cs + its test (if nothing else references them — grep first)

Context: Only do the delete/drop once Task 6 is green and grep confirms no remaining reference. If the migration surface is risky (central deployments with existing rows), the drop can be deferred to a follow-up — mark it clearly and stop, rather than shipping a risky migration under time pressure. Keeping a dead-but-harmless table is an acceptable interim.

Step 14: Add migration, remove entity, delete dead store, dotnet build + migration-applies check against a scratch DB.

Step 5: Commitrefactor(mesh-phase4): drop dead ConfigDb ScriptedAlarmState table.


Task 10: Rig config — driver-only site nodes lose their ConfigDb string

Classification: standard Estimated implement time: ~4 min Parallelizable with: Task 8

Files:

  • Modify: docker-dev/docker-compose.yml (site-a-1/2, site-b-1/2: remove ConnectionStrings__ConfigDb; set ConfigSource__Mode: "FetchAndCache" as the default, not behind OTOPCUA_CONFIG_MODE — Phase 4 makes Direct invalid for these nodes; central-1/2 keep ConfigDb
    • Direct)

Context: The rig currently defaults site nodes to Direct and flips to FetchAndCache via OTOPCUA_CONFIG_MODE. Phase 4 makes FetchAndCache the only valid mode for driver-only site nodes, so it becomes their committed default and the ConfigDb string is removed to prove the cut. Verify each site service still boots (the validator from Task 0 now requires this shape). This is a rig change only — the live gate (below) runs it.

Step 14: Edit compose; lmxopcua-fix-equivalent local docker compose up; confirm the four site nodes boot with no ConfigDb string.

Step 5: Commitchore(mesh-phase4): rig site nodes run with no ConfigDb string.


Task 11: Live gate

Classification: high-risk Estimated implement time: live verification (not a code task) Parallelizable with: none (final)

Files:

  • Create: docs/plans/2026-07-23-mesh-phase4-live-gate.md (record the run, as Phases 13 did)

Exit gate (design §6.1 + program Phase 4):

  1. Bring up the rig (central pair Direct + ConfigDb; four site nodes FetchAndCache, no ConfigDb string). All six healthy.
  2. Deploy a config via POST :9200/api/deployments (X-Api-Key: docker-dev-deploy-key, Content-Type: application/json); confirm it seals green with the site nodes acking — proving central persists the site-node acks with the site nodes holding no ConfigDb.
  3. Alarm cycle — drive a scripted alarm ack/shelve on a site node; confirm the condition state round-trips through LocalDb (inspect alarm_condition_state via the docker-cp SQLite recipe from the Phase-1 follow-ups memory) and replicates to its pair peer.
  4. Historian — confirm the alarm S&F drain + (if wired) continuous historization still run on the Primary site node with no ConfigDb.
  5. ServiceLevel — with central SQL stopped, a healthy site node still publishes 240/250 (Client.CLI redundancy), proving the retired DB-reachability input (a Direct node would have shown 0/100). Restart a site node → it boots last-known-good from the LocalDb pointer, no ConfigDb read.
  6. Grep proofgrep -rn "GetRequiredService<IDbContextFactory<OtOpcUaConfigDbContext>>\|AddOtOpcUaConfigDb" shows every hit hasAdmin-gated; a driver-only container's env has no ConnectionStrings__ConfigDb.

Record findings verbatim (the Phase 13 gates each caught a defect offline tests missed — expect the same and do not close the gate on a green-by-luck run).


Batching for execution

  • Batch 1 (mechanics): Task 0, Task 1 — options rule + registration gate.
  • Batch 2 (ServiceLevel): Task 2, Task 3 — the client-visible change (sequential, same file).
  • Batch 3 (actor + store): Task 4, Task 5 (‖), then Task 6.
  • Batch 4 (integration): Task 7 — the sweep + exit-gate DI proof.
  • Batch 5 (finish): Task 8 (‖ Task 10), Task 9, then Task 11 live gate.