Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
32 KiB
Per-Cluster Mesh Phase 4 — Cut the Driver-Side ConfigDb Connection
For Claude: REQUIRED SUB-SKILL: Use superpowers-extended-cc:executing-plans to implement this plan task-by-task.
Goal: A driver-only OtOpcUa node runs a full deploy + alarm + historian cycle with no
ConfigDb connection string configured at all — LocalDb is its steady-state config + alarm-state
store, and central persists its deploy acks. Grep-level proof: no driver-branch service can resolve
OtOpcUaConfigDbContext.
Architecture: Phase 3 already gave driver nodes their config from central (fetch-and-cache into
LocalDb) behind the ConfigSource:Mode=FetchAndCache dark switch. Phase 4 removes the remaining
four driver-side ConfigDb consumers and the registration itself, so a driver-only host boots with the
ConfigDb connection string absent. The fused central node keeps ConfigDb via its admin role —
this phase gates the whole registration on hasAdmin, never on hasDriver.
Tech stack: .NET 10, Akka.NET, EF Core (ConfigDb, SQL Server — central only after this phase),
ZB.MOM.WW.LocalDb (SQLite, replicated), xUnit + Shouldly.
Design decisions (settled — do not relitigate)
- ConfigDb is registered iff
hasAdmin.Program.cs:126's unconditionalAddOtOpcUaConfigDbmoves inside ahasAdminguard. A driver-only node requires noConfigDbconnection string and must not throw for its absence. Fusedadmin,drivercentral nodes are unaffected (admin owns the DB; driver actors on that node may still be handed a non-null factory). - Driver-only ⇒
FetchAndCacheis mandatory.Directmode reads central SQL, which a driver-only node no longer has.ConfigSourceOptionsValidatorfails start for adriver-role, non-adminnode whoseModeisDirect. (A fusedadmin,drivernode may stayDirect— it has the DB.) - Db-reachability retires from ServiceLevel on driver-only nodes (user decision 2026-07-23):
LocalDb is the config store and is in-process, so
DbReachableis treated as a constanttrueon nodes with noDbHealthProbeActor. Staleness already carries "running behind on config." A healthy site node serving last-known-good with central down publishes 240/250 — the survive-alone posture. This is a client-visible ServiceLevel semantics change — documented indocs/Redundancy.md+ the interop note. - Central is the system of record for deploy acks.
ConfigPublishCoordinator.PersistNodeAckalready writesNodeDeploymentStatefrom everyApplyAckit receives over the transport, and seeds theApplyingrows at dispatch. SoDriverHostActor's directUpsertNodeDeploymentStateSQL writes are redundant and are removed. The bootstrapNodeDeploymentStateread is already replaced by the LocalDbdeployment_pointerforFetchAndCache(Phase 3). ScriptedAlarmState(Part 9 condition state) → LocalDb, wholesale for driver-role nodes. It is read only byEfAlarmConditionStateStore; AdminUI/alertsuses live SignalR, not the table. A new replicatedalarm_condition_stateLocalDb table +LocalDbAlarmConditionStateStorereplaces the EF store on every driver-role node (fused central included), same journey Phase 2 made for the alarm S&F buffer. The dead ConfigDbScriptedAlarmStatetable is dropped in a final trivial migration.OpcUaPublishActor._dbFactorybecomes null on driver-only nodes. Its only ConfigDb use isLoadArtifact/LoadLatestArtifact(the #485 artifact-blob reads);FetchAndCachealready passesmsg.Artifactin-hand so the fallback never fires. Phase 4 asserts the null-factory path never reads ConfigDb and every driver-only rebuild carries bytes.
The five consumers (design §6.1 audit) and their Phase-4 disposition
| Consumer | Current ConfigDb use | Phase-4 disposition | Task |
|---|---|---|---|
DriverHostActor |
UpsertNodeDeploymentState writes (×9), bootstrap NodeDeploymentState read, EfAlarmConditionStateStore ctor |
drop SQL ack-writes (central persists); nullable factory; LocalDb condition store | 4, 5–6 |
EfAlarmConditionStateStore |
Part 9 condition state in ConfigDb ScriptedAlarmState |
→ LocalDbAlarmConditionStateStore (new replicated table) |
5, 6 |
DbHealthProbeActor |
SELECT 1 vs ConfigDb → DbReachable ServiceLevel input |
not spawned on driver-only; DbReachable=true constant |
2, 3 |
OpcUaPublishActor |
LoadArtifact/LoadLatestArtifact |
nullable factory; rebuild always carries bytes | 3 |
ServiceCollectionExtensions (Runtime + Configuration) + Program.cs |
registration | hasAdmin-gated; nullable threading |
1, 7 |
Task 0: ConfigSourceOptionsValidator — driver-only ⇒ FetchAndCache
Classification: small Estimated implement time: ~4 min Parallelizable with: Task 5
Files:
- Modify:
src/Core/ZB.MOM.WW.OtOpcUa.Cluster/ConfigSourceOptionsValidator.cs - Test:
tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/Configuration/ConfigSourceOptionsValidatorTests.cs(or the existing validator test file — locate withgrep -rl ConfigSourceOptionsValidator tests)
Context: The validator today checks the FetchAndCache fields (endpoints/key). Phase 4 adds the
role-conditional rule. The validator must be given the node's roles — check how it reads them today
(likely IConfiguration["Cluster:Roles"] or an injected role view); mirror AkkaClusterOptionsValidator,
which already matches on roles at boot.
Step 1: Write the failing test — a driver-role, non-admin node with Mode=Direct → validation
fails with a message naming the flag; a fused admin,driver node with Mode=Direct → passes; a
driver-only node with Mode=FetchAndCache → passes.
Step 2: Run it, confirm it fails (rule not yet present).
Step 3: Implement the rule: roles.Contains("driver") && !roles.Contains("admin") && Mode==Direct
→ ValidateOptionsResult.Fail(...). Read roles the same way the validator/AddValidatedOptions wiring
already does; do not invent a new source.
Step 4: Run tests, confirm pass.
Step 5: Commit — feat(mesh-phase4): driver-only node must be FetchAndCache (validator).
Task 1: Program.cs — gate ConfigDb registration + health check on hasAdmin
Classification: high-risk Estimated implement time: ~5 min Parallelizable with: none (boot path; other tasks build on the nullable factory it introduces)
Files:
- Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Host/Program.cs:125-126(the unconditionalAddOtOpcUaConfigDb) - Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Host/Health/HealthEndpoints.cs:27-31(DatabaseHealthCheck<OtOpcUaConfigDbContext>— driver-only has no ConfigDb to probe) - Test:
tests/Server/ZB.MOM.WW.OtOpcUa.Host.IntegrationTests/DriverOnlyNoConfigDbBootTests.cs(create)
Context: AddOtOpcUaConfigDb throws if the ConfigDb connection string is absent. Line 126 comment
says "always registered regardless of role. ConfigDb is required for everything" — that statement is
what Phase 4 falsifies. Gate the call on hasAdmin. Anything downstream that resolves
IDbContextFactory<OtOpcUaConfigDbContext> non-optionally must already be hasAdmin-only or become
nullable (Tasks 2–7 handle the driver-branch resolvers). Grep GetRequiredService<...ConfigDbContext...>
and GetRequiredService<IDbContextFactory<OtOpcUaConfigDbContext>> across the Host to find any that
would now throw on a driver-only node; each must move under hasAdmin or switch to GetService.
Step 1: Write the failing test — build a Host WebApplication with roles driver only and no
ConfigDb connection string; assert it builds + starts without throwing, and that
services.GetService<IDbContextFactory<OtOpcUaConfigDbContext>>() is null. (Model the harness on
DeploymentArtifactFetchBoundaryTests.StartServerAsync — minimal builder, in-memory where needed.)
Step 2: Run it — fails today (AddOtOpcUaConfigDb throws "Connection string 'ConfigDb' is required").
Step 3: Implement — wrap AddOtOpcUaConfigDb in if (hasAdmin); wrap the ConfigDb health check in
if (hasAdmin); sweep the GetRequiredService sites found above. Add a boot log line on a driver-only
node: "driver-only node — no ConfigDb registered; config via ConfigSource:Mode=FetchAndCache".
Step 4: Run the test + dotnet build — confirm pass and no unresolved-service throw.
Step 5: Commit — feat(mesh-phase4): register ConfigDb only on admin-role nodes.
Task 1b: driver-only ILdapGroupRoleMappingService — empty (appsettings-only group→role) — DISCOVERED during Task 1
Classification: small Estimated implement time: ~4 min Parallelizable with: Task 5
Files:
- Create:
src/Server/ZB.MOM.WW.OtOpcUa.Security/Ldap/NullLdapGroupRoleMappingService.cs(or a Host-side driver-only impl — place it besideOtOpcUaGroupRoleMapper) - Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Host/Program.cs:317-327(thehasDriverauth block — register the empty service withTryAddScopedso a fused node keeps the real EF one; fix the now-false comment at 322-323) - Test:
tests/Server/ZB.MOM.WW.OtOpcUa.Security.Tests/(or whereverOtOpcUaGroupRoleMapperis tested)
Context — a 6th ConfigDb consumer the §6.1 audit missed. OtOpcUaGroupRoleMapper (the OPC UA
data-path group→role mapper, IGroupRoleMapper<string>, registered at Program.cs:327) merges the
appsettings Security:Ldap:GroupToRole baseline (always available) with system-wide DB grants from
ILdapGroupRoleMappingService — which is registered only inside AddOtOpcUaConfigDb. Task 1 gated
that on hasAdmin, so a driver-only node now cannot resolve IGroupRoleMapper<string> at OPC UA login
(lazy scoped resolve — Build() does not surface it; the auth path does). LdapGroupRoleMapping rows are
not in the deployment artifact (ConfigComposer never composes them), so driver nodes never received
DB grants via config in the first place. The honest fix: a driver-only ILdapGroupRoleMappingService whose
GetByGroupsAsync returns empty (CRUD methods throw NotSupportedException — never called with no AdminUI),
registered via TryAddScoped in the hasDriver block so the fused node's real EF impl still wins.
Behavior reduction to document (Task 8): DB-backed LDAP group→role grants (the AdminUI RoleGrants page)
apply only where ConfigDb is present (admin/fused nodes); driver-only site nodes map groups→roles from
Security:Ldap:GroupToRole appsettings only.
Step 1: Write the failing test — resolve IGroupRoleMapper<string> in a driver-only service graph and
call MapAsync(["some-group"]) → returns the appsettings baseline with no throw (today it throws:
ILdapGroupRoleMappingService unresolvable).
Step 2: Run — fails.
Step 3: Implement the empty service + TryAddScoped registration in the hasDriver block; correct the
Program.cs:322-323 comment to state the driver-only node uses the empty impl and why.
Step 4: Run tests + build.
Step 5: Commit — fix(mesh-phase4): driver-only nodes map LDAP groups from appsettings (no ConfigDb grants).
Task 2: DbHealthProbeActor — do not spawn on driver-only; DbReachable=true constant
Classification: high-risk Estimated implement time: ~5 min Parallelizable with: none (ServiceLevel; Task 3 depends on its message contract)
Files:
- Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ServiceCollectionExtensions.cs:329-332(spawn site) - Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Runtime/OpcUa/OpcUaPublishActor.cs:641-655(the_dbHealthProbe is nullseam + theStalederivation) - Test:
tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/OpcUa/OpcUaPublishActorServiceLevelTests.cs(extend the existing ServiceLevel test file)
Context: ServiceLevelCalculator.Compute collapses (DbReachable=false, probe ok, not stale) → 0.
The retired-input rule (decision 3): on a node with _dbHealthProbe is null, feed
DbReachable: true and derive Stale without the DB-health term —
Stale = (now - snapshotEntry.AsOfUtc) > _staleWindow || runningBehindConfig. Do NOT fall through to
LegacyRoleOnly for driver-only nodes: that seam exists for the pre-first-sample bootstrap window on a
DB-backed node, and would peg a healthy site node at the role-only mapping forever. Introduce an explicit
"DB-less node" branch that computes the full NodeHealthInputs with DbReachable=true and the DB-free
staleness. runningBehindConfig = the Phase-3 running-from-cache / last-apply-failed signal if reachable
from the publish actor; if it is not currently threaded here, use the snapshot-age term alone and record
a TODO(Phase 4 follow-up) — do not thread a new dependency just for this in-scope.
Step 1: Write failing tests — (a) DB-less node, member Up, probe ok, snapshot fresh → 240
(+10 if Primary); (b) DB-less node, snapshot stale (>window) → 200; (c) DB-backed node unchanged
(regression). Assert against OpcUaPublishActor's computed ServiceLevelChanged, not the calculator
in isolation.
Step 2: Run — (a)/(b) fail (today a DB-less node hits LegacyRoleOnly or 0).
Step 3: Implement — in ServiceCollectionExtensions, spawn DbHealthProbeActor only when
dbFactory is not null; register DbHealthProbeActorKey only then; pass dbHealthProbe: null into
OpcUaPublishActor.Props on a DB-less node. In OpcUaPublishActor, add the DB-less branch described
above (guard on "no probe wired AND this is a driver-only/DB-less node" — a clean signal is
_dbHealthProbe is null combined with a ctor flag dbLess, threaded from the spawn site, so the
legacy-bootstrap seam stays reachable only on DB-backed nodes before their first sample).
Step 4: Run tests + build — confirm pass.
Step 5: Commit — feat(mesh-phase4): retire DB-reachability from ServiceLevel on DB-less nodes.
Task 3: OpcUaPublishActor — nullable factory; driver-only rebuild always carries bytes
Classification: standard Estimated implement time: ~4 min Parallelizable with: none (same file as Task 2 — sequence after it)
Files:
- Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Runtime/OpcUa/OpcUaPublishActor.cs - Test:
tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/OpcUa/OpcUaPublishActorTests.cs
Context: _dbFactory is already nullable on this actor; HandleRebuild's _dbFactory is null || _applier is null guard falls back to a raw _sink.RebuildAddressSpace(). That fallback is the
wrong path for a driver-only node that HAS an applier but no factory — it would skip the
diff-and-apply. The correct driver-only invariant: _applier is wired, _dbFactory is null, and every
RebuildAddressSpace carries msg.Artifact in-hand (Phase-3 DriverHostActor already does this). Split
the guard so a null factory with an applier still does the diff-and-apply pass, using msg.Artifact,
and asserts (logs an Error + abandons per #485) if msg.Artifact is null on such a node — a null artifact
with no factory to fall back to is "no answer," never an empty rebuild.
Step 1: Write failing tests — (a) applier wired, factory null, RebuildAddressSpace with
Artifact → diff-and-apply runs (assert applier invoked, not the raw-sink fallback); (b) applier wired,
factory null, RebuildAddressSpace with null Artifact → no rebuild, Error logged (#485 negative).
Step 2: Run — (a) fails (null factory currently routes to the raw-sink fallback).
Step 3: Implement — restructure HandleRebuild: if (_applier is null) { raw-sink fallback; return; }
then the diff-and-apply block, where artifact = msg.Artifact ?? (_dbFactory is not null ? (depId path else latest) : Array.Empty<byte>()); the existing #485 zero-length guard then abandons the null-factory
- null-artifact case cleanly.
Step 4: Run tests + build.
Step 5: Commit — feat(mesh-phase4): OpcUaPublish diff-applies from in-hand bytes with no ConfigDb.
Task 4: DriverHostActor — nullable factory; drop redundant SQL ack-writes on DB-less nodes
Classification: high-risk Estimated implement time: ~5 min Parallelizable with: none (core actor; Tasks 5–6 build the store it constructs)
Files:
- Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Runtime/Drivers/DriverHostActor.cs(_dbFactoryfield/ctor/Props → nullable; the ~9UpsertNodeDeploymentStatecall sites; the bootstrapNodeDeploymentStateread at 606-612; theEfAlarmConditionStateStorector at 588; the CreateDbContext sites at 1920, 2096, 2588) - Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ServiceCollectionExtensions.cs:455(Props call) - Test:
tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/Drivers/DriverHostActorDbLessTests.cs(create)
Context: Decision 4 — central's ConfigPublishCoordinator.PersistNodeAck is the ack system of record.
On a DB-less node UpsertNodeDeploymentState has no ConfigDb to write and its row is written by central
anyway. Guard every UpsertNodeDeploymentState body on _dbFactory is not null (make the method a no-op
when null) rather than deleting call sites — keeps the fused-node behaviour identical and the diff small.
Same for the bootstrap read (606-612) and the Direct-only CreateDbContext sites (1920/2096/2588): each is
already unreachable on a FetchAndCache node (Phase 3 routes bootstrap through the LocalDb pointer and
apply through the fetcher), so the change is defensive nullability, not new control flow. Verify each of
1920/2096/2588 is inside a Direct-mode-only path before guarding — if any is on the shared apply path,
that is a Phase-3 gap to surface, not silently guard.
Two specific sites the Task-2 code review surfaced (MUST be handled here):
- Remove the interim
dbFactory!atServiceCollectionExtensions.cs:~470(added in Task 2 to preserve compilation) — onceDriverHostActor._dbFactoryis nullable, pass the plain (nullable)dbFactory. UpsertNodeDeploymentState(2613-2643) is called UNCONDITIONALLY on the FetchAndCache apply path (BeginFetchAndCacheApply/HandleFetchedForApply, ~1716/1723/1774/1785/1801). Its try/catch swallows the null-factory NRE and logs a misleading "transient DB" warning on every deploy on a driver-only node. The_dbFactory is not nullguard fixes this.SpawnScriptedAlarmHost(called unconditionally from PreStart ~524) constructsnew EfAlarmConditionStateStore(_dbFactory, ...)at ~587-588 — with a null factory every alarm Load/Save NREs (swallowed → scripted-alarm state silently dead on driver-only nodes). This is fixed by Task 6 (which swaps in the LocalDb-backed store), NOT by a null guard — do NOT leave the Ef store constructed with a null factory. If Task 6 lands after this task, add a temporary_dbFactory is not nullguard around the Ef-store construction so a DB-less node skips it cleanly until Task 6 wires the LocalDb store; the two tasks together must leave a DB-less node with a WORKING (LocalDb) condition store, not a skipped one.
Step 1: Write failing tests — construct a DriverHostActor with dbFactory: null,
fetchAndCacheMode: true, a fake fetcher + LocalDb cache (mirror DriverHostActorFetchAndCacheTests):
(a) a full dispatch → Applied ack, no NullReferenceException from a factory deref; (b) bootstrap
with a seeded LocalDb pointer → revision restored with no ConfigDb read.
Step 2: Run — fails today (_dbFactory non-nullable; UpsertNodeDeploymentState derefs it).
Step 3: Implement — make _dbFactory nullable through field/ctor/both Props overloads (nullable
param last, matching the Phase-3 additions); UpsertNodeDeploymentState early-returns when null (log
Debug once); guard the read sites. Update ServiceCollectionExtensions:455 to pass the (already nullable)
dbFactory.
Step 4: Run tests + build.
Step 5: Commit — feat(mesh-phase4): DriverHostActor runs with no ConfigDb (central persists acks).
Task 5: alarm_condition_state LocalDb table + LocalDbAlarmConditionStateStore
Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 0
Files:
- Create:
src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ScriptedAlarms/LocalDbAlarmConditionStateStore.cs - Modify: the LocalDb schema/DDL +
RegisterReplicatedwiring (find viagrep -rn "alarm_sf_events" src— mirror that table's registration exactly, in the sameLocalDbSetup.OnReadyDDL→RegisterReplicatedorder which is load-bearing) - Test:
tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/ScriptedAlarms/LocalDbAlarmConditionStateStoreTests.cs(portEfAlarmConditionStateStoreTests— sameIAlarmStateStorecontract, so the assertion bodies transfer verbatim; only the store construction changes)
Context: LocalDbAlarmConditionStateStore : IAlarmStateStore must round-trip the same fields
EfAlarmConditionStateStore does (Enabled/Acked/Confirmed/Shelving + ack/confirm audit + CommentsJson;
Active/LastActive/LastCleared transient; LastTransitionUtc ↔ UpdatedAtUtc). Columns mirror
ScriptedAlarmState. Replicated like alarm_sf_events so a failover peer holds the state — but unlike
the S&F buffer there is no drain gate: condition state is not delivered-once, it is the current snapshot,
idempotent last-write-wins on UpdatedAtUtc (the Ef store already tolerates the write race that way).
Do NOT stack app-level LWW on LocalDb's HLC (design §8) — the row upsert keyed by ScriptedAlarmId
is a single unconditional write, which is HLC-safe; keep it that way.
Step 1: Write failing tests — the ported IAlarmStateStore round-trip suite against the LocalDb
store over a temp SQLite DB (the Runtime.Tests harness already opens LocalDb elsewhere — reuse it).
Step 2: Run — fails (store + table absent).
Step 3: Implement — the store + the alarm_condition_state DDL + RegisterReplicated registration.
Step 4: Run tests + build.
Step 5: Commit — feat(mesh-phase4): LocalDb alarm-condition-state store (replicated, pair-local).
Task 6: Wire the LocalDb condition store into DriverHostActor for driver-role nodes
Classification: standard Estimated implement time: ~4 min Parallelizable with: none (depends on Task 5's type + Task 4's nullable factory)
Files:
- Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Runtime/Drivers/DriverHostActor.cs:588(constructLocalDbAlarmConditionStateStoreinstead ofEfAlarmConditionStateStore) - Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ServiceCollectionExtensions.cs(resolve the LocalDb store dependency; thread it intoDriverHostActor.Props) - Test: extend
DriverHostActorDbLessTests— the ScriptedAlarm host constructs against LocalDb with a null ConfigDb factory.
Context: The store is constructed inside DriverHostActor.BuildScriptedAlarmHost (line ~588). It
takes the LocalDb connection the node already opened (Phase 1/2). Prefer resolving an
IAlarmStateStore from DI (registered in the Host's hasDriver LocalDb block) over new-ing it in the
actor, so the actor stays factory-agnostic. Confirm the registration lands in the same AddOtOpcUaLocalDb
hasDriver branch as the S&F sink.
Step 1: Write failing test — DB-less DriverHostActor boots its ScriptedAlarm host and a condition
save/load round-trips through LocalDb (no ConfigDb).
Step 2: Run — fails (still constructs the Ef store needing a factory).
Step 3: Implement — register IAlarmStateStore → LocalDbAlarmConditionStateStore in the Host
hasDriver LocalDb block; resolve + thread it; delete the EfAlarmConditionStateStore construction from
the actor. Leave the EfAlarmConditionStateStore class in the tree (removed in Task 9's migration cleanup)
or delete it now if nothing else references it (grep first).
Step 4: Run tests + build.
Step 5: Commit — feat(mesh-phase4): driver-role nodes store alarm condition state in LocalDb.
Task 7: Registration cleanup + full-graph sweep
Classification: standard Estimated implement time: ~5 min Parallelizable with: none (integration seam over everything above)
Files:
- Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ServiceCollectionExtensions.cs(final nullable threading) - Modify:
src/Server/ZB.MOM.WW.OtOpcUa.Host/Program.cs(any remaining driver-branch ConfigDb resolves) - Test:
tests/Server/ZB.MOM.WW.OtOpcUa.Runtime.Tests/ServiceCollectionExtensionsTests.cs
Context: The exit-gate proof. After Tasks 1–6, sweep for any service the driver-only graph
resolves that still needs OtOpcUaConfigDbContext. Grep both projects for
OtOpcUaConfigDbContext, IDbContextFactory<OtOpcUaConfigDbContext>, AddOtOpcUaConfigDb,
ScriptedAlarmStates, and GetRequiredService of any of them; each hit must be hasAdmin-gated or
nullable-tolerant. LdapGroupRoleMappingService / ILdapGroupRoleMappingService (registered inside
AddOtOpcUaConfigDb) is admin/auth-side — confirm the driver-only login path does not resolve it, or
that the driver-only node never serves login (it hosts no AdminUI — see the topology table).
Step 1: Write the failing test — a "driver-only DI graph resolves nothing ConfigDb-bound" test:
build the hasDriver && !hasAdmin service provider and assert
GetService<IDbContextFactory<OtOpcUaConfigDbContext>>() is null AND that spawning the runtime actors
does not throw. Also pin the Task-1b fused-node ordering (Task 1b review gap): build the
hasAdmin && hasDriver graph through the same AddOtOpcUaConfigDb + hasDriver TryAdd sequence and assert
GetService<ILdapGroupRoleMappingService>() resolves the real LdapGroupRoleMappingService, not the
Null… — so a later reorder of the role blocks (or switching AddOtOpcUaConfigDb to a TryAdd) fails a
red test instead of silently stripping DB-backed role grants from central. Also resolve
IGroupRoleMapper<string> in a driver-only scope to catch the lazy auth-path break Task 1b fixed.
Step 2: Run — fix whatever it catches.
Step 3: Implement the remaining guards.
Step 4: Run the full Runtime.Tests + Host.IntegrationTests suites + dotnet build.
Step 5: Commit — feat(mesh-phase4): driver-only DI graph resolves nothing ConfigDb-bound.
Task 8: Docs — Redundancy.md (ServiceLevel change) + interop note + Configuration.md + CLAUDE.md
Classification: small Estimated implement time: ~5 min Parallelizable with: Task 9
Files:
- Modify:
docs/Redundancy.md(ServiceLevel: DB-less node table + theDbReachable=trueconstant + the 240-with-central-down behaviour; theDbHealthProbeActorrow becomes admin-only) - Modify:
docs/Configuration.md(driver-only nodes: noConfigDbstring;ConfigSource:Modemust beFetchAndCache; validator rule) - Modify:
CLAUDE.md(extend the mesh section — Phase 4: driver ConfigDb cut; ServiceLevel semantics; alarm condition state in LocalDb) - Modify:
docs/plans/2026-07-22-per-cluster-mesh-program.md(Phase 4 status + tracking row)
Step 1–4: Prose only — state the client-visible ServiceLevel change explicitly (a healthy DB-less node now publishes 240/250 with central unreachable, where a Direct node would have dropped to 0/100).
Step 5: Commit — docs(mesh-phase4): ServiceLevel semantics + driver ConfigDb cut.
Task 9: Drop the dead ConfigDb ScriptedAlarmState table (migration) + retire EfAlarmConditionStateStore
Classification: standard Estimated implement time: ~4 min Parallelizable with: none (after Task 6 proves LocalDb store is the live path)
Files:
- Create: EF migration dropping
ScriptedAlarmState(find the migrations dir viagrep -rl "class OtOpcUaConfigDbContext"→ itsMigrations/sibling; usedotnet ef migrations add DropScriptedAlarmState -c OtOpcUaConfigDbContext) - Modify:
src/Core/ZB.MOM.WW.OtOpcUa.Configuration/OtOpcUaConfigDbContext.cs,.../Entities/ScriptedAlarmState.cs(remove the DbSet + entity) - Delete:
src/Server/ZB.MOM.WW.OtOpcUa.Runtime/ScriptedAlarms/EfAlarmConditionStateStore.cs+ its test (if nothing else references them — grep first)
Context: Only do the delete/drop once Task 6 is green and grep confirms no remaining reference. If the migration surface is risky (central deployments with existing rows), the drop can be deferred to a follow-up — mark it clearly and stop, rather than shipping a risky migration under time pressure. Keeping a dead-but-harmless table is an acceptable interim.
Step 1–4: Add migration, remove entity, delete dead store, dotnet build + migration-applies check
against a scratch DB.
Step 5: Commit — refactor(mesh-phase4): drop dead ConfigDb ScriptedAlarmState table.
Task 10: Rig config — driver-only site nodes lose their ConfigDb string
Classification: standard Estimated implement time: ~4 min Parallelizable with: Task 8
Files:
- Modify:
docker-dev/docker-compose.yml(site-a-1/2, site-b-1/2: removeConnectionStrings__ConfigDb; setConfigSource__Mode: "FetchAndCache"as the default, not behindOTOPCUA_CONFIG_MODE— Phase 4 makes Direct invalid for these nodes; central-1/2 keep ConfigDbDirect)
Context: The rig currently defaults site nodes to Direct and flips to FetchAndCache via
OTOPCUA_CONFIG_MODE. Phase 4 makes FetchAndCache the only valid mode for driver-only site nodes, so
it becomes their committed default and the ConfigDb string is removed to prove the cut. Verify each site
service still boots (the validator from Task 0 now requires this shape). This is a rig change only — the
live gate (below) runs it.
Step 1–4: Edit compose; lmxopcua-fix-equivalent local docker compose up; confirm the four site
nodes boot with no ConfigDb string.
Step 5: Commit — chore(mesh-phase4): rig site nodes run with no ConfigDb string.
Task 11: Live gate
Classification: high-risk Estimated implement time: live verification (not a code task) Parallelizable with: none (final)
Files:
- Create:
docs/plans/2026-07-23-mesh-phase4-live-gate.md(record the run, as Phases 1–3 did)
Exit gate (design §6.1 + program Phase 4):
- Bring up the rig (central pair
Direct+ ConfigDb; four site nodesFetchAndCache, no ConfigDb string). All six healthy. - Deploy a config via
POST :9200/api/deployments(X-Api-Key: docker-dev-deploy-key,Content-Type: application/json); confirm it seals green with the site nodes acking — proving central persists the site-node acks with the site nodes holding no ConfigDb. - Alarm cycle — drive a scripted alarm ack/shelve on a site node; confirm the condition state
round-trips through LocalDb (inspect
alarm_condition_statevia the docker-cp SQLite recipe from the Phase-1 follow-ups memory) and replicates to its pair peer. - Historian — confirm the alarm S&F drain + (if wired) continuous historization still run on the Primary site node with no ConfigDb.
- ServiceLevel — with central SQL stopped, a healthy site node still publishes 240/250
(Client.CLI
redundancy), proving the retired DB-reachability input (a Direct node would have shown 0/100). Restart a site node → it boots last-known-good from the LocalDb pointer, no ConfigDb read. - Grep proof —
grep -rn "GetRequiredService<IDbContextFactory<OtOpcUaConfigDbContext>>\|AddOtOpcUaConfigDb"shows every hithasAdmin-gated; a driver-only container's env has noConnectionStrings__ConfigDb.
Record findings verbatim (the Phase 1–3 gates each caught a defect offline tests missed — expect the same and do not close the gate on a green-by-luck run).
Batching for execution
- Batch 1 (mechanics): Task 0, Task 1 — options rule + registration gate.
- Batch 2 (ServiceLevel): Task 2, Task 3 — the client-visible change (sequential, same file).
- Batch 3 (actor + store): Task 4, Task 5 (‖), then Task 6.
- Batch 4 (integration): Task 7 — the sweep + exit-gate DI proof.
- Batch 5 (finish): Task 8 (‖ Task 10), Task 9, then Task 11 live gate.