Files
ScadaBridge/docs/plans/2026-07-19-localdb-phase2-gate.md
T
Joseph Doherty 12eb97e07a docs(localdb): close the phase 2 gate — questions answered, plan written
Task 2 of the Phase 2 plan, plus the clean re-run of Task 1's cut-short
sampling that Task 2 was waiting on.

Gate doc: status NOT STARTED -> CLOSED. All five §5 open questions answered
inline. The questions are kept, not deleted — several of the answers overturn
a premise the gate itself stated, and that reasoning is the value:

- Oplog sizing is not the binding constraint. Cap overrun prunes and flags
  needs_snapshot (graceful resync), so the caps need no defensive sizing; the
  real ceiling is D6's 4 MB gRPC cap, binding on MaxBatchSize x row-bytes.
- LWW-vs-in-flight-send cannot arise on a correct pair (only the active node
  sweeps). The honest cost is D2's semantic change: the standby is convergent,
  no longer byte-identical.
- Migration-under-load is designed out, not managed — both mechanisms are
  deleted in the same commit and D5 forecloses rolling upgrades.
- Rollback is genuinely "revert the commit": the migration is additive and
  leaves the legacy files intact. Bounded cost is the post-cutover delta.
- One DB per process stays; every write axis is bounded.

§2's requirement to state CDC's duplicate-delivery bound is routed explicitly
to Task 21 rather than left implicit.

Task 1 re-run (safe copy-based snap(), both nodes restarted first, six 60s
intervals): sf_messages 0.80 rows/sec insert, 0/sec retry-UPDATE, oplog 0,
alarms 0, max payload_json 76 B, zero SQLite errors across 30 min, both site-a
nodes converged at 3564 rows. Independently confirms the disk-I/O storm was
observer-induced. Recorded with its limits stated: 0.80/s is ~1.6% of the 50/s
ceiling and the retry path was never exercised — acceptable only because that
ceiling is structural rather than empirical, and the rig's 721 B config rows
are explicitly NOT representative.

GATE VERDICT: PROCEED to Task 3. Binding output is MaxBatchSize 500 -> 16.

Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts
2026-07-20 02:15:53 -04:00

14 KiB
Raw Blame History

LocalDb Phase 2 — Gate Document

Status: CLOSED (2026-07-20). The implementation plan is 2026-07-19-localdb-adoption-phase2.md; its decisions D1D6 answer this gate and its Task 1 supplied the measurements §5 demanded (soak record: 2026-07-19-localdb-phase2-soak.md). Each §5 question is answered inline below. The questions are deliberately not deleted — the reasoning is the value, and several of the answers overturn a premise stated above.

Phase 1 live gate: PASS (2026-07-19, docker rig). See the Task 12 commit for evidence.

Phase 2 goal (from the design doc, scadaproj docs/plans/2026-07-19-scadabridge-localdb-design.md §1): move scadabridge.db's config tables and StoreAndForward's sf_messages into the same consolidated LocalDb file, then delete SiteReplicationActor and StoreAndForward's ReplicationService outright. CDC triggers replace hand-shipped ops; the library's snapshot resync replaces the bespoke chunked anti-entropy. The design explicitly chose replace + delete, no dual-mechanism period — recoverable from git.


1. What Phase 1 actually established

Facts a Phase 2 plan can rely on, all verified rather than assumed:

  • One ILocalDb per site process at LocalDb:Path, required (ValidateOnStart). A site config missing it fails to boot.
  • SiteLocalDbSetup.OnReady is the single place tables are created and registered. Ordering is load-bearing: DDL → RegisterReplicated → writes, then the legacy migrator. Rows written before registration are invisible to the peer forever, silently.
  • Replication is registered in SiteServiceRegistration (not Program.cs) so the composition-root tests cover it; only MapZbLocalDbSync lives in Program.cs.
  • Inbound auth is LocalDbSyncAuthInterceptor, fail-closed, scoped to /localdb_sync.v1.LocalDbSync/, sharing the existing 8083 h2c listener.
  • Default-OFF is real and proven side-by-side: rig site-a replicates, site-b/site-c do not.
  • ZbTelemetryOptions.Meters is a silent allowlist (SiteServiceRegistration.ObservedMeters).
  • Bidirectional sync works with only one node configured as initiator.

2. What Phase 2 must port before deleting anything

The bespoke replicator's tests encode behaviour that the design's CDC replacement must still satisfy. Port the test intents first, delete second — the plan should treat the existing tests as the specification, not as code to be removed alongside the implementation.

  • tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/SiteReplicationActorTests.cs
  • tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationServiceTests.cs

Implementation under deletion:

  • src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Actors/SiteReplicationActor.cs
  • src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ReplicationService.cs

Known accepted behaviour to carry forward, not silently drop: the N5 bounded-duplicate race the bespoke replicator documents and accepts. A Phase 2 plan must state whether CDC inherits the same bound, a tighter one, or a different failure shape.

Routed (2026-07-20). The failure shape changes, so N5's bound does not simply carry over: CDC has no chunked anti-entropy hop to duplicate across, but LWW admits re-delivery of a row whose park loses to a later write (see the LWW answer in §5). Task 21 states the new duplicate-delivery bound explicitly and rewrites the N5 note rather than deleting it.

3. Substantially harder than Phase 1

Phase 1 moved two tables that nothing replicated before, so the worst case was "no worse than today". Phase 2 replaces a working mechanism, which makes it a different risk class:

  • sf_messages is an outbound buffer with side effects. Phase 1's tables are read-mostly records; a store-and-forward row that resurrects or double-delivers sends real traffic to an external system. Park/requeue/remove becoming ordinary captured row updates/deletes needs its own analysis of what LWW does to an in-flight send.
  • Config tables drive deployment. A converged-but-wrong config row is a deployed instance behaving incorrectly, not a missing history entry.
  • No dual-mechanism period. The chosen posture means the cutover is the test. That raises the bar on the offline convergence suite and the live gate considerably.
  • Migration is not a no-op this time. Phase 1's migrator usually finds nothing, because the legacy files sat outside the volume. scadabridge.db and store-and-forward.db are real, populated, on the volume, and actively replicated while the migration runs.

4. Required inputs before writing the plan

  1. The design doc §1 Phase 2 and §2 schema/migration sections, re-read against post-Phase-1 reality (several Phase 1 assumptions changed during execution — see the amendments array in 2026-07-19-localdb-adoption-phase1.md.tasks.json).
  2. The two test files in §2, read as specifications.
  3. Rig soak observations that do not exist yet — see §5.

5. Open questions — answer these first

  • Oplog growth and churn under real config/S&F write rates. Phase 1's tables are low-volume; sf_messages is not. MaxOplogRows / MaxOplogAge / snapshot-resync thresholds cannot be chosen from first principles. This needs a Phase-1 rig soak that has not been run. The live gate proved correctness, not steady-state behaviour over time.

    ANSWERED — the soak ran (Task 1); the caps are not the binding constraint, batch bytes are. Two of this question's own premises turned out to be wrong, and both corrections relax it:

    1. The stop condition was much weaker than written. Exceeding MaxOplogRows / MaxOplogAge does not wedge or lose data — OplogStore prunes and flags the peer needs_snapshot (OplogStore.cs:109-138, MaintenanceBackgroundService.cs:57), which degrades to a graceful snapshot resync. Overrun is a performance event, not a correctness event, so these caps do not need to be sized defensively.
    2. sf_messages has a hard ceiling, not an estimate. SweepBatchLimit (500) ÷ RetryTimerInterval (10 s) = ≤50 row-writes/sec, structurally. Measured rig rate under the purpose-built SoakGenerator load was far below that (see the soak record). Rows are small — max payload_json 76 bytes on the rig.

    native_alarm_state is answered under the last question in this section, not here (and the plan's D4, whose premise it corrects). The real ceiling this question was groping toward is D6's 4 MB gRPC message cap, and it binds on MaxBatchSize × max-row-bytes, not on oplog depth — Task 19 sets LocalDb:Replication:MaxBatchSize = 16, the one firmly evidence-backed number the soak produced. The keyed-instances escape hatch is not needed; this plan proceeds.

    Independently, the soak retired a scare: a disk I/O error storm initially read as a Phase 1 library defect and STOP-gated this plan was root-caused as observer-induced (host-side sqlite3 against live bind-mounted WAL files) — see docs/known-issues/2026-07-20-localdb-disk-io-error-under-load.md §0. The unmodified library sustains the full soak load with zero SQLite errors.

  • LWW semantics for in-flight store-and-forward sends. What does last-writer-wins do when one node parks a message the other is mid-delivery on?

    ANSWERED — the scenario cannot arise on a correctly-behaving pair, and D2 covers what happens if it does. Only the active node sweeps and delivers; the standby holds a convergent copy and sends nothing. So "one node mid-delivery while the other parks" is a split-brain symptom, not steady-state behaviour, and store-and-forward is not where it should be defended against.

    If it does happen, LWW resolves per row by HLC: the later write wins, and the losing node's view converges to it. The concrete risk is at-least-once delivery — a park that loses to a stale in-flight update can be re-swept and re-sent. That is not new; store-and-forward is already at-least-once by construction (a send that succeeds but whose ack is lost is retried). What is new is D2's semantic change, recorded here because it is the honest cost of this gate: the standby is no longer guaranteed byte-identical to the active buffer, only convergent. The bespoke replicator's ReplaceAllAsync bought identity by wiping — and the N1 directional guard existed precisely because that wipe was dangerous. The library's snapshot resync merges per row and never wipes (SnapshotApplier, LwwApplier.cs:69-78), so the failure the guard prevented is structurally impossible rather than merely tested against. Task 10 ports SfBufferResyncPredicateTests as a convergence assertion, not a directional-authority one.

  • Migration-under-load. How does the one-time copy of an actively-replicating scadabridge.db interact with the bespoke replicator still running during cutover?

    ANSWERED — the interaction is designed out, not managed. The two mechanisms never run concurrently. SiteReplicationActor and StoreAndForward's ReplicationService are deleted in the same commit that registers the Phase 2 tables (Tasks 14/15), and D5 forecloses rolling upgrades: both nodes of a site stop and start together. A process that boots with the new code has no bespoke replicator to race, and one running the old code has no CDC triggers. There is no window in which a row is written by one mechanism and read by the other.

    Within a single booting process the ordering is the load-bearing part, and it is the same invariant Phase 1 established: DDL → RegisterReplicated → migrate → writes (SiteLocalDbSetup.OnReady). Rows written before registration are invisible to the peer forever, silently — which is why Tasks 8/9 run the migrator strictly after registration, so every migrated row is captured by the CDC triggers and replicates normally. Both nodes migrating independently is fine: they migrate the same source rows, and LWW converges them.

  • Rollback story. With no dual-mechanism period, what is the recovery path if the cutover fails in production — beyond "revert the commit"?

    ANSWERED — "revert the commit" is genuinely the path, and it is safe because the migration is additive and the legacy files are left intact. Tasks 8/9 copy rows out of scadabridge.db and store-and-forward.db into the consolidated LocalDb file; they do not drop, truncate, or delete the source databases. Rolling back is therefore: stop both nodes of the site, deploy the previous build, start both together (D5). The old code reopens the legacy files and finds them exactly as it left them.

    The bounded, honest cost of a rollback is the delta — rows written into the consolidated DB after cutover do not flow back into the legacy files. For config tables that self-heals: SiteReconciliationActor reports local inventory to central at startup and fetches whatever it lacks, so a rolled-back node re-converges to central's truth without operator action. For sf_messages the delta is lost undelivered buffer — the practical mitigation is to drain the buffer before cutting over, which Task 20's live gate and Task 21's runbook both call for.

    What has no rollback is a site pair split across versions — hence D5. That is a deployment-procedure constraint, and Task 21 puts it in the runbook rather than leaving it as tribal knowledge.

  • Does the consolidated file stay appropriate? One-DB-per-process was chosen partly because Phase 1's tables were small. Adding config + S&F changes the size and write profile of that single file.

    ANSWERED — yes, and the soak is the evidence. The write profile the consolidated file must absorb is bounded on every axis:

    Table Rate bound Row size Basis
    sf_messages ≤50 writes/sec (hard) 76 B max on the rig SweepBatchLimit ÷ RetryTimerInterval
    native_alarm_state distinct_source_refs × 10/sec small per-SourceReference coalescing on a 100 ms flush (NativeAlarmActor.cs:473-502)
    config tables deploy-driven, effectively idle ~721 B max on the rig Task 1 measurement

    These are unremarkable for SQLite in WAL mode, and the soak ran the full generator load against the Phase 1 consolidated file for 30 minutes with zero SQLite errors and no oplog growth pathology. One DB per process stays.

    Two caveats carried into the plan rather than hidden here. First, the rig's config rows are tiny (max 721 B) and cannot be treated as representative — the size ceiling that matters comes from the documented ~6070 KB production config_json, which is what motivates D6 and MaxBatchSize = 16. Second, D4's premise was wrong: native_alarm_state is not "unbounded by design" and not "by a wide margin the highest-volume table" — the coalescing flush bounds it, and this rig has no alarm generator at all (measured 0 rows), so its bound is analytic rather than observed. If a production site ever shows alarm churn that swamps the shared oplog, the keyed-instances hatch named in D4 remains the escape — it is simply not needed to start.

6. Not in scope (unchanged from the design)

auditlog.db (diverges per node by design — central pulls the union; replicating it would double-forward), the secrets store (has its own answer in Secrets.Replicator.SqlServer hub mode), and central nodes (SQL-Server-first, no LocalDb use case).