# LocalDb Phase 1 follow-ups — live gate (docker-dev rig) > Closes the two limitations left open by `2026-07-20-localdb-phase1-live-gate.md` (checks 4 and 8). > Rig: local `docker-dev/docker-compose.yml`, 6 nodes rebuilt from `fix/localdb-phase1-followups`. > site-a pair replicates; central + site-b stay default-OFF. > > **Safety rules followed** (same as the Phase 1 gate): all DB inspection via `docker cp` of the > `db`/`-wal`/`-shm` triplet, querying the copy — never host `sqlite3` on the live WAL file. ## Result **Both follow-ups closed, and the gate found a third defect** — one that was *unreachable* until the first fix landed. Library `ZB.MOM.WW.LocalDb` went `0.1.1` → `0.1.2` → `0.1.3` in the course of this gate; both consumers (OtOpcUa, ScadaBridge) are pinned to `0.1.3`. | # | Check | Result | Evidence | |---|---|---|---| | 8 | Oplog stops growing on a default-OFF node | ✅ PASS | site-b-1 restarted twice with an unchanged cached artifact: `oplog_depth` **7 → 7 → 7**, `deployment_pointer.applied_at_utc` frozen at `05:34:58` across both — the re-cache was skipped outright. Repeated on the `0.1.3` build after a real deploy: **13 → 13**. Pre-fix this grew ~2 rows per restart forever, because a default-OFF node has no peer to ack and prune them. | | 8b | …but a genuinely new deployment still writes (positive control) | ✅ PASS | A real config change (renamed tag `SITEA-tag`) produced rev `822bdf34`, then `002c16b3`. site-b-1's oplog **7 → 10 → 13** and its pointer advanced each time. The skip suppresses *identical* re-writes only; it is not a dead cache-write path. | | 4 | A wiped node is back-filled by its peer, no new deploy, central down | ✅ PASS (was ⚠️ PARTIAL) | Precondition asserted first — site-a-1's oplog **0**, `needs_snapshot 0`: the exact state that used to make back-fill impossible. Then SQL stopped, site-a-2's LocalDb **volume destroyed**, container recreated. a-1 logged `Snapshot sent (as_of_seq 0, 4 rows)` — a line that could not appear on `0.1.1`. a-2 came back with pointer `SITE-A / 475df87a @ 03:09:12` (a-1's **original** timestamp, not a re-derived one), chunks 2, and a `__localdb_row_version` dump **byte-identical to a-1's**, carrying origin node ids `6812ca9c` and `935659a5` — the rebuilt node's own id appears nowhere, so the rows were replicated, not recreated. | | 4b | …and the back-filled node then boots from cache | ✅ PASS | With SQL still down, site-a-2 logged `RUNNING FROM CACHE — … booted deployment 475df87a… cached at 2026-07-21T03:09:12`, `raw subtree materialised (containers=16, tags=1)`. OPC UA browse of both nodes: **17 ns=2 nodes each, diff-identical** — the rebuilt node serves exactly what its healthy peer does, from a configuration it never applied itself and could not have fetched. | | 4c | The rebuilt node's OWN writes reach its peer | ✅ PASS **after a third fix** (see below) | On `0.1.2` this failed: a-1 held `last_applied_remote = 10` while the rebuilt a-2's oplog was seqs **1,2,3** and its `last_acked` stayed **0** — never accepted. On `0.1.3`, a-1 logs `Replication peer identity changed (ebb10854 -> c85d8b1d): the peer was rebuilt … Resetting the inbound watermark`, and after the next deploy a-1 shows `last_applied_remote = 3` / a-2 shows `last_acked = 3` with a drained oplog and identical row_versions. | ## The third defect (the reason this gate was worth running) `0.1.2` made back-fill work, which for the first time let a rebuilt node be observed *writing*. It turned out its writes went nowhere. `last_applied_remote_seq` is a watermark in the **peer's** seq space, and a rebuilt peer is a new database whose oplog numbers from 1. Two independent bugs stacked, and fixing either alone leaves the data stuck: 1. The healthy node kept advertising the **old** peer's watermark. The *sending* side seeds its pump directly from that number (`_sentThruSeq`), so the rebuilt node skipped its entire oplog. → the watermark (and the observed peer clock) is now reset when the peer's node id changes. `last_acked_seq` is deliberately **not** reset: it describes our own oplog's pruned horizon and is exactly what `0.1.2` keys the snapshot decision off — clearing it would switch back-fill back off. 2. A peer claiming to have applied more of our stream than we have ever produced can only be remembering a previous incarnation of *us*. → that claim is now rejected in favour of pumping from the start; LWW makes the re-send harmless. This mirrors the clamp `RecordPeerAckAsync` already applies to acks. Both were reproduced offline as a RED test before fixing (`SnapshotResyncTests.RebuiltPeer_OwnWritesReachTheHealthyNode_DespiteTheOldPeersWatermark`). **Why offline tests could not have found it first:** the wiped-peer scenario had no test because the scenario was believed impossible to recover from — it was a *documented limitation*, not a bug. The gate is what turned the limitation into a reproducible case, and the second defect only existed behind the first. ## Unrelated finding (NOT a follow-up of this work — pre-existing; `lmxopcua` issue #485, since FIXED) While SQL was being stopped and started to drive check 4, site-b-1 took a deploy whose artifact load hit a transient ConfigDb error, and **emptied its served address space**: ``` [06:11:20 ERR] An error occurred using the connection to database 'OtOpcUa' on server 'sql,1433'. [06:11:20 WRN] OpcUaPublish: failed to load artifact for deployment 43d54397…; rebuild becomes no-op [06:11:20 INF] AddressSpaceApplier: applied plan (kind=PureRemove, added=0, removed=16, rebuild=True) ``` The log says "rebuild becomes no-op" and the applier then removed all 16 nodes — those two disagree. The cache layer behaved correctly and refused to store the bad artifact ("not caching deployment … its artifact was empty or could not be loaded, so it is not a usable last-known-good configuration"), and the node recovered fully on restart (`containers=16`). But a transient ConfigDb blip emptying the served address space instead of holding last-known-good is the same failure *class* as the Phase 1 gate's check-3 defect. Untouched by this work — site-b-1 has replication off entirely — and induced by this gate's own SQL flapping. **Fixed** (issue #485): `OpcUaPublishActor.HandleRebuild` now treats an artifact it could not obtain as *no answer* rather than *an empty configuration*, and abandons the rebuild with the served address space (and `_lastApplied`) intact. Nothing legitimate produces a zero-length blob — an operator who really deletes everything still deploys a JSON document with empty arrays — so "no bytes" can only mean the read failed. The log excerpt above is now impossible: the `PureRemove` never runs, and the loader's own "the rebuild is abandoned" line is finally true. Covered by `OpcUaPublishActorRebuildTests.Rebuild_keeps_the_last_known_good_address_space_when_the_artifact_load_fails`, with a positive control proving a *readable* artifact that genuinely drops an equipment still tears its subtree down.