Enables the alarm store-and-forward sink on all four driver nodes of the
docker-dev rig -- the replicating site-a pair and the default-OFF site-b pin
-- so the live gate has a real buffer to watch converge, and can see that
site-b's sink works as a plain node-local queue with no peer traffic.
Two rig details worth stating rather than rediscovering. The endpoint is
deliberately unresolvable: there is no HistorianGateway here, so every drain
attempt fails, which is exactly the historian-outage state the buffer exists
for. But it still has to be a syntactically valid absolute http(s) URI or the
host refuses to start, because ServerHistorianOptionsValidator is
consumer-gated -- AlarmHistorian:Enabled=true makes the endpoint required
even while ServerHistorian:Enabled=false. And MaxAttempts is raised far above
the production default of 10, which against a permanently unreachable gateway
would dead-letter the entire queue about five minutes in, turning a buffering
test into a dead-letter test.
Docs record what an operator now has to know: that a rising queue depth on a
Secondary is correct and on BOTH nodes is not (that shape is a redundancy
snapshot naming neither node -- check node identity before suspecting the
sink), that delivery is at-least-once across a failover by design with no
dedup layer, and that AlarmHistorian:DatabasePath is removed but should be
left in place through the upgrade because the migrator still reads it to find
the file to copy.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
The gate closed both follow-ups on the docker-dev rig and found a third defect
that offline tests could not have found first: it lived behind the one being
fixed. Once 0.1.2 made back-fill work, a rebuilt node could be observed writing
for the first time — and its writes went nowhere, because the healthy node kept
the old peer's seq watermark. Library 0.1.1 -> 0.1.2 -> 0.1.3 over the course
of the gate; both consumers now pinned to 0.1.3.
Check 4 (was PARTIAL, now PASS): with SQL stopped and site-a-2's LocalDb volume
destroyed, a-1 logged "Snapshot sent (as_of_seq 0, 4 rows)" — a line that could
not appear on 0.1.1 — and a-2 came back with a row_version dump byte-identical
to a-1's, carrying a-1's origin node ids and its ORIGINAL applied_at_utc. It
then booted from cache and served 17 ns=2 nodes, diff-identical to its peer,
from a configuration it never applied and could not have fetched.
Check 8 (PASS): two restarts with an unchanged artifact left the oplog at 7 and
the pointer timestamp frozen — the re-cache was skipped outright. Positive
control: a real config change still wrote, 7 -> 10 -> 13.
Also records one unrelated pre-existing finding the gate's SQL flapping
surfaced: a transient ConfigDb error during artifact load empties the served
address space while logging "rebuild becomes no-op". The cache layer correctly
refused to store the bad artifact and the node recovered on restart, but the
two logs disagree and it is the same class as the Phase 1 gate's check-3
defect. Untouched by this work; wants its own issue.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
1. Wiped-node back-fill (live gate check 4). Fixed upstream in
ZB.MOM.WW.LocalDb 0.1.2 and pinned here. The gate's diagnosis was slightly
off: the oplog cap was not the gate. Snapshot detection measured the peer's
gap against the oldest surviving oplog row and read an empty oplog as "no
gap possible" — and an empty oplog is the steady state of a converged pair,
since ack-pruning deletes everything the peer confirmed. The healthy state
was the one state that could not heal a wiped peer. Now measured against
last_acked_seq when the oplog is empty.
Pinned at this level too, not just the library's: the pair harness grows a
WipePassiveAsync (a NEW database — a rebuilt node comes back with a new node
id and a zero watermark), and the new scenario asserts the emptied oplog as
its precondition before wiping, then requires the deployment artifact to
come back byte-identical with no new deploy. Verified RED against the pinned
0.1.1 (times out waiting for back-fill) and green on 0.1.2.
2. Oplog growth on default-OFF nodes (live gate check 8), fixed at the source:
StoreAsync now skips entirely when the pointer already names this
deployment/revision, the SHA matches the bytes, and the expected chunk count
is present. Re-caching an artifact the node already holds — every restart's
boot-from-cache, every RestoreApplied — writes nothing and mints no oplog
rows, where before it cost a delete plus an insert per chunk plus a pointer
update, all identical to what was already there.
Identity is over the bytes and the chunks are counted, both deliberately: a
re-composed artifact can carry the same ids with different bytes, and a
pointer can name a deployment whose chunks are missing, where skipping would
make an unreadable cache permanent. Both guards verified RED against a naive
pointer-only skip.
The gate's stated bound was also wrong and is corrected in the doc: growth
was never headed for the 1M row cap. AddZbLocalDbReplication is registered
unconditionally, so MaintenanceBackgroundService runs on default-OFF nodes
and the 7-day MaxOplogAge cap prunes.
Runtime.Tests 412/0/31, Host.IntegrationTests LocalDb 45/45. Pre-existing and
unrelated: AbCip_Green_AgainstSim (needs the AB CIP docker fixture, red on a
clean master too) and a flaky Roslyn race test (green 2/2 in isolation).
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
New operations runbook (enable/disable, fail-closed ApiKey rule, stop/start-
together, tombstone-retention window, MaxBatchSize row-count-vs-4MB, never
sqlite3-a-live-WAL-DB cp-triplet recipe). Configuration.md gains a LocalDb
section; Redundancy.md gains a pair-local config cache section (what boot-from-
cache does and does not cover); CLAUDE.md gains the Phase 1 scope paragraph.
OtOpcUa is Akka-clustered and AddZbSecrets is registered unconditionally, so every
node (admin, driver, fused) resolves secrets and needs the SAME KEK + same store rows.
Ship G-5 as a production-posture runbook rather than hardcoding Source=File + shared
paths into the committed role-overlay appsettings (which dev + TwoNodeClusterHarness
also consume — that would break every dev/CI boot). Base appsettings stays on
Source=Environment. Documents the interim File-KEK + shared-SQLite posture and the
G-7 hand-off (ConfigDb-backed ISecretStore mirroring the existing DP
PersistKeysToDbContext<OtOpcUaConfigDbContext> key-ring sharing). Cross-linked from
docs/security.md. Mirrors the ScadaBridge G-5 resolution.