Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
6.2 KiB
LocalDb Phase 1 follow-ups — live gate (docker-dev rig)
Closes the two limitations left open by
2026-07-20-localdb-phase1-live-gate.md(checks 4 and 8). Rig: localdocker-dev/docker-compose.yml, 6 nodes rebuilt fromfix/localdb-phase1-followups. site-a pair replicates; central + site-b stay default-OFF.Safety rules followed (same as the Phase 1 gate): all DB inspection via
docker cpof thedb/-wal/-shmtriplet, querying the copy — never hostsqlite3on the live WAL file.
Result
Both follow-ups closed, and the gate found a third defect — one that was unreachable until the
first fix landed. Library ZB.MOM.WW.LocalDb went 0.1.1 → 0.1.2 → 0.1.3 in the course of this
gate; both consumers (OtOpcUa, ScadaBridge) are pinned to 0.1.3.
| # | Check | Result | Evidence |
|---|---|---|---|
| 8 | Oplog stops growing on a default-OFF node | ✅ PASS | site-b-1 restarted twice with an unchanged cached artifact: oplog_depth 7 → 7 → 7, deployment_pointer.applied_at_utc frozen at 05:34:58 across both — the re-cache was skipped outright. Repeated on the 0.1.3 build after a real deploy: 13 → 13. Pre-fix this grew ~2 rows per restart forever, because a default-OFF node has no peer to ack and prune them. |
| 8b | …but a genuinely new deployment still writes (positive control) | ✅ PASS | A real config change (renamed tag SITEA-tag) produced rev 822bdf34, then 002c16b3. site-b-1's oplog 7 → 10 → 13 and its pointer advanced each time. The skip suppresses identical re-writes only; it is not a dead cache-write path. |
| 4 | A wiped node is back-filled by its peer, no new deploy, central down | ✅ PASS (was ⚠️ PARTIAL) | Precondition asserted first — site-a-1's oplog 0, needs_snapshot 0: the exact state that used to make back-fill impossible. Then SQL stopped, site-a-2's LocalDb volume destroyed, container recreated. a-1 logged Snapshot sent (as_of_seq 0, 4 rows) — a line that could not appear on 0.1.1. a-2 came back with pointer SITE-A / 475df87a @ 03:09:12 (a-1's original timestamp, not a re-derived one), chunks 2, and a __localdb_row_version dump byte-identical to a-1's, carrying origin node ids 6812ca9c and 935659a5 — the rebuilt node's own id appears nowhere, so the rows were replicated, not recreated. |
| 4b | …and the back-filled node then boots from cache | ✅ PASS | With SQL still down, site-a-2 logged RUNNING FROM CACHE — … booted deployment 475df87a… cached at 2026-07-21T03:09:12, raw subtree materialised (containers=16, tags=1). OPC UA browse of both nodes: 17 ns=2 nodes each, diff-identical — the rebuilt node serves exactly what its healthy peer does, from a configuration it never applied itself and could not have fetched. |
| 4c | The rebuilt node's OWN writes reach its peer | ✅ PASS after a third fix (see below) | On 0.1.2 this failed: a-1 held last_applied_remote = 10 while the rebuilt a-2's oplog was seqs 1,2,3 and its last_acked stayed 0 — never accepted. On 0.1.3, a-1 logs Replication peer identity changed (ebb10854 -> c85d8b1d): the peer was rebuilt … Resetting the inbound watermark, and after the next deploy a-1 shows last_applied_remote = 3 / a-2 shows last_acked = 3 with a drained oplog and identical row_versions. |
The third defect (the reason this gate was worth running)
0.1.2 made back-fill work, which for the first time let a rebuilt node be observed writing. It
turned out its writes went nowhere.
last_applied_remote_seq is a watermark in the peer's seq space, and a rebuilt peer is a new
database whose oplog numbers from 1. Two independent bugs stacked, and fixing either alone leaves the
data stuck:
- The healthy node kept advertising the old peer's watermark. The sending side seeds its pump
directly from that number (
_sentThruSeq), so the rebuilt node skipped its entire oplog. → the watermark (and the observed peer clock) is now reset when the peer's node id changes.last_acked_seqis deliberately not reset: it describes our own oplog's pruned horizon and is exactly what0.1.2keys the snapshot decision off — clearing it would switch back-fill back off. - A peer claiming to have applied more of our stream than we have ever produced can only be
remembering a previous incarnation of us. → that claim is now rejected in favour of pumping from
the start; LWW makes the re-send harmless. This mirrors the clamp
RecordPeerAckAsyncalready applies to acks.
Both were reproduced offline as a RED test before fixing
(SnapshotResyncTests.RebuiltPeer_OwnWritesReachTheHealthyNode_DespiteTheOldPeersWatermark).
Why offline tests could not have found it first: the wiped-peer scenario had no test because the scenario was believed impossible to recover from — it was a documented limitation, not a bug. The gate is what turned the limitation into a reproducible case, and the second defect only existed behind the first.
Unrelated finding (NOT a follow-up of this work — pre-existing; filed as lmxopcua issue #485)
While SQL was being stopped and started to drive check 4, site-b-1 took a deploy whose artifact load hit a transient ConfigDb error, and emptied its served address space:
[06:11:20 ERR] An error occurred using the connection to database 'OtOpcUa' on server 'sql,1433'.
[06:11:20 WRN] OpcUaPublish: failed to load artifact for deployment 43d54397…; rebuild becomes no-op
[06:11:20 INF] AddressSpaceApplier: applied plan (kind=PureRemove, added=0, removed=16, rebuild=True)
The log says "rebuild becomes no-op" and the applier then removed all 16 nodes — those two disagree.
The cache layer behaved correctly and refused to store the bad artifact ("not caching deployment … its
artifact was empty or could not be loaded, so it is not a usable last-known-good configuration"), and
the node recovered fully on restart (containers=16). But a transient ConfigDb blip emptying the
served address space instead of holding last-known-good is the same failure class as the Phase 1
gate's check-3 defect. Untouched by this work — site-b-1 has replication off entirely — and induced by
this gate's own SQL flapping.