Checks 1, 2, 3, 4 and the oplog/dead-letter/metrics half of 8 captured and passing. Check 5 blocked on rig state, checks 6, 7, 9, 10 not run. Task 21 must NOT proceed on this evidence. Strongest result is check 3+4 together: the deployed config row is byte- identical on both site-a nodes with the SAME __localdb_row_version HLC and originating node id, and the standby logged zero config fetches in the deploy window. The standby holds the row node A wrote, obtained purely by CDC — which is the whole point of deleting notify-and-fetch. Check 5 is blocked by a rig-shaping problem, not a product fault: removing the owed soakgen instances orphaned their 11,804 buffered sf_messages, and a replacement probe's external system never reached the site nodes (site config tables come from deployment artifacts, and an external system referenced only by name inside script code is not carried). Same blocker stops check 10. Two method corrections recorded in the doc: - The plan says to run DB checks host-side against the bind mounts. That is WRONG — a host sqlite3 open poisons the container's WAL. All reads here copy the db/-wal/-shm triplet and query the copy. - An apparent "localdb_* metrics missing" finding was an artifact of the aspnet:10.0 image having no curl, with the error swallowed by 2>/dev/null. Re-scraped via a network-sharing sidecar: the metrics are present. Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts
7.5 KiB
LocalDb Phase 2 — live gate evidence (INCOMPLETE, 2026-07-20)
Gate for Task 20. Status: 5 of 10 checks captured and passing, 1 blocked, 4 not run. Do NOT proceed to Task 21 on this evidence.
Build under test: feat/localdb-phase2 @ 166f07fa (LocalDb 0.1.1, confirmed live — the
metrics scrape reports otel_scope_version="0.1.1"). Rig redeployed via docker/deploy.sh,
exit 0, all 8 nodes up, central healthy.
Method note — reading the databases
Every SQLite read below went through a snapshot helper that copies the db/-wal/-shm
triplet and queries the COPY. The plan's instruction to "run DB checks host-side against the
bind mounts" is wrong and must not be followed. macOS↔container locks do not cross virtiofs,
so a host sqlite3 open recovers the WAL out from under the container and triggers a permanent
disk I/O error (SQLITE_IOERR_SHORT_READ 522) storm — the root cause of the 2026-07-20 incident.
cp is a pure byte read with no SQLite involvement, so copying first is safe.
Passing
1. Migration ran — PASS
Both markers present on both site-a nodes, and row counts match the legacy source:
| Node | .migrated markers |
legacy sf_messages |
consolidated sf_messages |
|---|---|---|---|
| site-a-node-a | scadabridge.db.migrated, store-and-forward.db.migrated |
11804 | 11804 |
| site-a-node-b | both | 11804 | 11804 |
deployed_configurations was 0 in both legacy files, so that table's migration is untested by
this check — it is covered instead by check 3, which deploys fresh.
Note the migration itself ran on an earlier deploy of this build (markers dated 05:04), so this is after-the-fact evidence, not an observed run. The markers and counts are exactly what the plan asks for, but a from-scratch re-run was not performed.
2. No SMTP/notification rows migrated — PASS
smtp_configurations and notification_lists are both 0 on both nodes.
That count alone is vacuous — the legacy source also had 0 SMTP rows, so an exclusion that did not work would look identical. The non-vacuous evidence is the registered-table list: the CDC capture triggers exist for exactly these ten tables on both nodes —
OperationTracking, data_connection_definitions, database_connections,
deployed_configurations, external_systems, native_alarm_state,
sf_messages, shared_scripts, site_events, static_attribute_overrides
— and smtp_configurations / notification_lists have no triggers, so no replication
channel exists for them at all. That is the property the security decision actually rests on.
3. Config converges byte-identical — PASS
Deployed gate-sensor-1 (template 3, site 1) to site-a. Both nodes:
- whole-row md5 identical:
0ec0f3cf790bfd65c45edf6c4afae50f deployment_id20b81607…,revision_hashsha256:253e1a56…,deployed_at2026-07-20T09:06:11.2087221+00:00— identical__localdb_row_version: HLC116951506695487488, originating nodecf61c688-a57f-4c5c-ba76-feff6e5679fe,is_tombstone0 — identical on both
Identical HLC and originating node id is the strong form: node B holds the row A wrote, not an independently-derived copy.
4. The standby made no config fetch — PASS
Deploy window from 2026-07-20T09:05:52Z:
scadabridge-site-a-a(active, deploy target): 1 fetch —Fetching config for deployment 20b81607… instance gate-sensor-1 (notify-and-fetch). Legitimate: this is the active node pulling its own artifact.scadabridge-site-a-b(standby): 0 fetches.
Under the old architecture the standby was notified by SiteReplicationActor and fetched
independently. It now receives the config purely by CDC replication — which check 3 corroborates
from the other direction (the row on B carries A's node id).
8 (partial). Oplog drained, no dead letters, localdb_* scraped — PASS
__localdb_oplog= 0 rows on both nodes;localdb_oplog_depth= 0 on both.__localdb_dead_letter= 0 on both nodes.localdb_sync_applied_total= 11808 (node a) / 224 (node b);localdb_sync_reconnects_total= 1.
Correction worth recording: an earlier scrape appeared to show zero localdb_* metrics. That
was an artifact — docker exec … curl fails because the aspnet:10.0 image has no curl, and
the error was being swallowed by 2>/dev/null. The correct method is a network-sharing sidecar:
docker run --rm --network container:scadabridge-site-a-a curlimages/curl:latest -s localhost:8084/metrics
I nearly recorded "metrics not scraped" as a finding on the strength of a missing binary.
Blocked
5. Store-and-forward convergence — NOT DEMONSTRATED
Not a product failure — the rig could not be driven into the required state. What happened:
- Removing the owed
soakgen-1..4instances left their 11,804 bufferedsf_messagesorphaned (11,780 Pending / 24 Parked). The site now logs a continuous flood ofCached-telemetry drain: no tracking snapshot for …warnings for them. - A purpose-built probe (external system
GateDeadTarget→http://127.0.0.1:9, methodPing, plus an interval script on template 3 callingExternalSystem.CachedCall) deployed cleanly and itsScriptActor GateBufferProbe started, but produced zero buffered rows and zero errors. - Root cause identified: the site nodes'
external_systemstable still holds onlyTest REST API -> http://127.0.0.1:9.GateDeadTargetnever reached the site, so the script's call cannot resolve a target. Site config tables are populated from deployment artifacts, and the artifact forgate-sensor-1evidently does not carry an external system that is referenced only by name inside script code.
Consequence: the central-side restore of ExternalSystemDefinitions id 1 to
http://scadabridge-restapi:5200 has not propagated to the site nodes, which still target the
dead 127.0.0.1:9 — visible in the log as POST http://127.0.0.1:9/api/Add →
Connection refused. The backlog therefore cannot drain either.
To finish check 5, someone needs a probe whose external system is genuinely delivered to the site
(most likely bound to an instance, the way the retired SoakGenerator template did it), and the
orphaned backlog should be cleared first so the signal is not buried.
Not run
- 6. Site failover — buffered messages survive a site-a flip and deliver exactly once.
- 7. Cascade delete — no orphan
static_attribute_overrides/native_alarm_stateon the peer. (The offline suite covers this well:LocalDbPhase2ConvergenceTests.RemovingAnInstance_ ConvergesAllThreeCascadeTables, verified non-vacuous. The live run is still owed.) - 9. Both nodes stopped and started together (D5) — clean rejoin.
- 10. Post-cutover drain under churn — oplog rises and drains between bursts. Blocked by the same missing load generator as check 5.
Rig state as left
gate-sensor-1(instance 9, template 3) deployed on site-a, with aGateBufferProbeinterval script (template 3, script 1015) that currently does nothing.- External system
GateDeadTarget(id 2) with methodPing(id 2) — created for the probe. ExternalSystemDefinitionsid 1 restored centrally tohttp://scadabridge-restapi:5200; the site copies are still stale athttp://127.0.0.1:9.SoakGeneratortemplate (2021) andsoakgen-1..4deleted, per the owed cleanup — but their 11,804 orphanedsf_messagesremain and are generating a warning flood.soak-motor-1..4(template 4) untouched; template 4 is not deployable.