Gate record: docs/plans/2026-07-20-localdb-phase2-live-gate.md.
Checks 1, 2, 5, 6 pass. Checks 3 and 4 are NOT satisfied — the defect they
were meant to confirm turned out to be the opposite of what the plan assumed.
Three defects crash-looped every driver node before check 1 could even run:
1. An empty ServerHistorian:ApiKey kills the host. ServerHistorianOptions-
Validator exists to turn exactly that class of failure into a named
OptionsValidationException, but its documented fail tier explicitly excluded
ApiKey on the reasoning that a keyless client "degrades — the gateway rejects
calls". It does not: the client validates its own options at construction, so
the process dies during Akka startup and never makes a call.
2. UseTls disagreeing with the endpoint scheme kills the host too, in both
directions (both messages confirmed in the shipped client assembly). Moving an
endpoint from https to http without clearing UseTls is an ordinary migration
slip.
3. Plaintext h2c was UNREACHABLE. HistorianGatewayClientAdapter forwarded the
TLS-only options unconditionally, and AllowUntrustedServerCertificate defaults
to false, so it always sent RequireCertificateValidation=true — which the
client rejects outright when UseTls=false. Every http:// deployment crashed,
though the scheme is documented as the supported way to select h2c, and the
only workaround was to assert a certificate posture for a connection that has
no certificate.
The fourth was the blocker, and it is Phase 2's own:
4. The drain gate deferred to a Primary that cannot deliver. Redundancy roles are
elected CLUSTER-WIDE; the alarm queue is PAIR-LOCAL. On the rig the elected
driver Primary is central-1 — it carries the driver Akka role, replicates
nobody's LocalDb and does not even run the alarm historian — so every driver
node logged "Historian drain suspended", including the two site-b nodes that
have no peer at all. Nothing drained anywhere, where before Phase 2 it drained
fine. The cost is not a duplicate; it is the buffer growing to the capacity
wall and evicting the audit trail it exists to protect.
Fixed in three layers: a separate ShouldDrainAlarmHistory policy (unknown role
drains; the two gates now deliberately disagree, and a test pins that); peer-
host matching in DriverHostActor so a node stands down only for a Primary
holding its rows; and AddAlarmHistorian short-circuiting the gate when
replication is unconfigured — testing BOTH Replication:PeerAddress and
SyncListenPort, since only the dialing half sets the former while both halves
share the queue.
Every one of these follows from the asymmetry: a false allow costs a duplicate
row, which at-least-once delivery already accepts and payload-hash ids
collapse; a false deny loses data silently.
A third vacuous test, caught by the same delete-the-guard discipline: the
role-view tests stayed green with the guard removed, because AwaitAssert polls
until an assertion passes and the assertion was "reads open" — which is the
SEEDED value, satisfied at the first poll before the actor processed anything.
They now assert the sequence of published values through a recording view; the
control then goes red for exactly the cases that matter.
Migration evidence: 11 legacy rows across two deliberately overlapping files
converged to exactly 9 identical rows on both nodes, proving D-6's payload-hash
identity on real nodes rather than in a fixture.
Open design fork, recorded in the gate doc rather than decided here: a pair
cannot currently identify its own Primary, so both halves drain. Safe in every
topology — nothing loses data — but the gate's de-duplication benefit is
unrealised until roles are scoped per pair.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
Copies the pre-consolidation store-and-forward queue into the consolidated
database on first boot, then renames the legacy file aside. Rows are in that
file precisely because the historian could not be reached, so dropping them
on upgrade would discard exactly the alarm audit trail the queue exists to
protect.
Runs last in OnReady, after every RegisterReplicated call. It is the only
thing in OnReady that writes rows, and capture is trigger-based: a migration
that ran before registration would recover the backlog locally and never
replicate a line of it, silently and permanently.
Ids are derived from the payload rather than the plan's mig-{node}-{legacyId}
scheme. Node-prefixing solves the collision the legacy AUTOINCREMENT key
would cause -- node A's row 7 and node B's row 7 are different alarms -- but
it preserves a duplication that should be collapsed instead. A warm pair's
two legacy files OVERLAP: HistorianAdapterActor default-writes while its
redundancy role is unknown, so both nodes accepted the same transitions
during every boot window. Prefixed ids would carry those duplicates into the
merged buffer forever; equal-payload ids converge them. The same property
makes a crash between commit and rename harmless under INSERT OR IGNORE.
OnReady now takes IConfiguration rather than offering an overload that skips
the migration. A wiring mistake that silently discarded a node's undelivered
alarm history is not a mistake worth making possible.
The copy is restricted to the columns the legacy table actually has. Naming
a column an older build never wrote throws "no such column", which would
discard every row in the table rather than the one field.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
Adds the alarm store-and-forward buffer to the consolidated LocalDb file and
registers it for replication, so a node that dies with undelivered alarm
history no longer takes it to the grave.
The table keeps the legacy queue's shape rather than the status column the
plan sketched. An acknowledged row is deleted, not marked delivered: that
needs no second sweeper to keep the table bounded, leaves the capacity
semantics untouched, and the replication engine carries the delete as a
tombstone so the peer drops its copy anyway -- which is what the status
column was for. last_error is retained because it is the only operator-facing
record of why a row was dead-lettered.
The primary key is app-minted TEXT. The legacy AUTOINCREMENT RowId cannot
replicate under last-writer-wins: two nodes would independently allocate
rowid 7 to different alarms and silently overwrite each other. The drain
index therefore orders by enqueued_at_utc rather than insertion order, with
id as a tiebreak so the ordering is total.
Tables are created unconditionally, independent of AlarmHistorian:Enabled.
An empty registered table costs three triggers; creating it lazily would mean
a node that enables the historian later writes rows before its capture
triggers exist, which is exactly the silent-loss shape OnReady's ordering
comment warns about.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
The Kestrel re-bind read only urls/ASPNETCORE_URLS. The aspnet:8.0+ base images
set ASPNETCORE_HTTP_PORTS=8080 (all interfaces) as the CONTAINER default in place
of ASPNETCORE_URLS, so a driver node that never sets URLS fell straight through to
Kestrel's localhost:5000 default — silently moving its health/metrics surface to
loopback when the sync listener took over Kestrel. Now falls through URLS →
HTTP_PORTS/HTTPS_PORTS (wildcard, all-interfaces) → localhost:5000, mirroring
Kestrel's own source precedence. Found on the docker-dev live gate (central nodes
were fine — they set ASPNETCORE_URLS=9000 explicitly).
New Props/ctor parameter appended LAST at all three positions - the forwarding
list inside Props.Create is positional and compiles into an expression tree, so a
mis-ordered argument is usually type-compatible and binds the wrong dependency at
runtime. All 26 existing test construction sites use named arguments and needed no
change.
ReconcileDrivers now returns the blob it loaded. It catches its own DB failures and
returns without rethrowing, so ApplyAndAck can reach its success path having spawned
zero drivers; caching on 'the apply succeeded' would persist an empty artifact as
last-known-good and boot the node into an empty address space during the next
outage. Verified: weakening the guard to null-only turns EmptyArtifact_IsNotCached
red.
The store runs in its own try/catch because an Applied ACK has already been sent by
then - an escaping exception would land in ApplyAndAck's catch and emit a second,
contradictory Failed ACK for a deployment the fleet believes is live.
The Runtime.Deployment namespace shadowed the Deployment EF entity type for
every file under ZB.MOM.WW.OtOpcUa.Runtime.Tests.*: C# name lookup walks the
enclosing namespace chain and finds the namespace before a using-imported type,
so 19 pre-existing call sites failed with CS0118. Caught only by a full-solution
build - the Host and its own tests compiled fine.
Bumps the four ZB.MOM.WW.Secrets pins 0.2.1 -> 0.2.2 (closes the OtOpcUa side
of scadaproj#1, tracked here as #482). 0.2.1's Akka replicator deadlocked any
hosted process at startup when Secrets:Replication:Enabled was true: the
package's DI graph closed a circular singleton dependency through factory
lambdas (store decorator -> replicator -> actor provider -> cache invalidator
-> resolver -> store), which MS.DI's StackGuard turns into a silent
cross-thread call-site-lock deadlock. 0.2.2 defers the invalidator edge to
first eviction. The flag stays default-false; enabling remains a
per-environment decision.
The_startup_hook_actually_creates_the_replication_actor is now a real test:
SecretReplicationStarter's docs had promised it since the adoption, and the
upstream fix finally makes a provider-based resolve runnable - container built
exactly as the host does, hook started under a watchdog, replication actor
proven to exist by ActorSelection on a self-joined single-node cluster (no
TestKit needed, which matters because Akka.TestKit.Xunit2 is xunit-v2-only and
this project is on xunit.v3). Also corrects the stale rationale that blamed
the old hang on DistributedPubSub needing a joined cluster - the actor
constructor was never reached; it was the DI cycle.
Verified: SecretsReplicationRegistrationTests 8/8 on the 0.2.2 feed packages;
full slnx build 0 errors; the 2-node Akka live convergence gate re-run against
the published 0.2.2 packages passes 6/6 (write->peer, tombstone propagation
without resurrection, delete visibility through the resolver cache, reverse
direction, wrong-KEK fail-closed).
Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts
Routes the host's secrets registration through a new AddOtOpcUaSecrets extension
that gates the ISecretStore implementation on Secrets:Replication:Enabled.
Opt-in gate (default FALSE)
This call decides which ISecretStore every node resolves — including driver-role
nodes with no auth/AdminUI, where a wrong store surfaces as drivers failing to
open sessions rather than as a failing test. With the flag false the wiring is
the pre-existing AddZbSecrets(config, "Secrets") call, unchanged, so current
behavior is byte-identical. With it true, AddZbSecretsAkkaReplication replaces
that call (it invokes AddZbSecrets internally; calling both would double-register).
Extracted to a named extension specifically so the registration is testable:
Program.cs is top-level statements and cannot be exercised by a container test,
which is how a "registered but never resolvable" defect ships unnoticed.
Serializer HOCON
AkkaSecretsReplication.SerializationConfig is merged into the ActorSystem config
inside the AddAkka configurator, conditionally on the same gate — a non-replicating
node carries no bindings for messages it will never see. Merged via
AddHocon(..., HoconAddMode.Append), Akka.Hosting's fallback merge and the same mode
the existing base-config merge uses; a raw Config.WithFallback would fight the
builder's own assembly.
Lazy-actor mitigation
The replication actor is created lazily on first ISecretStore resolution, so a node
that never touches a secret would never announce a manifest and would silently never
converge. SecretReplicationStarter (IHostedService) resolves the store once at
startup to make participation unconditional.
KNOWN BLOCKER — replication is currently NON-FUNCTIONAL; do not enable
ZB.MOM.WW.Secrets.Replicator.AkkaDotNet 0.2.0 never binds its own ISecretReplicator.
AddZbSecretsAkkaReplication calls AddZbSecrets FIRST, which does
TryAddSingleton<ISecretReplicator, NoOpSecretReplicator>(); the package's own
TryAddSingleton<ISecretReplicator>(AkkaSecretReplicator) that follows is therefore
a no-op. Verified empirically in a built container: with Enabled=true,
ISecretReplicator resolves to NoOpSecretReplicator, so ReplicatingSecretStore
publishes into a sink and no actor is ever spawned.
Consequence: the startup hook cannot create the actor, and the test asserting it
does is committed Skipped with the evidence. Not worked around here — the fix
belongs upstream (AddSingleton, or register before calling AddZbSecrets).
Because the flag defaults false, this commit is inert in production.
Tests: SecretsReplicationRegistrationTests (new) — disabled path resolves plain
SqliteSecretStore and needs no ActorSystem; enabled path resolves
ReplicatingSecretStore AND the undecorated concrete SqliteSecretStore the decorator
is built from (the exact registration gap that shipped once); startup hook registered
only when enabled. Red before wiring (4 assertion failures), green after: 6 pass,
1 skipped (blocker above).
Build: 861 warnings / 0 errors, unchanged from baseline (full --no-incremental A/B).
Host.IntegrationTests: 123 pass, 6 skip, 1 fail — AbCip_Green_AgainstSim, verified
pre-existing on the stashed tree (fixture-gated).
Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts