Files
ScadaBridge/docs/plans/2026-07-19-localdb-adoption-phase2.md
Joseph Doherty 7b5a5a6f34 docs(localdb): phase 2 truth pass across both repos
Normative first: Component-StoreAndForward.md:83 specified the whole chunked,
ack-confirmed SfBufferSnapshotChunk resync protocol, which Phase 2 deleted. It
is rewritten rather than removed — the failure modes it reasoned about still
exist, they are just bounded differently — and it now states the
duplicate-delivery bound explicitly: a message can be delivered twice only when
the OLD primary delivered it and the status change had not yet replicated when
the gate flipped. One flush interval plus the in-flight ack, and unlike the old
model it does NOT grow with backlog depth or with how long a node was absent.
The N1 directional-authority and N5 orphan-row hazards are recorded as
structurally gone, not merely unguarded.

Frame-size known-issue amended: its 2026-06-26 resolution replaced the
intra-site hop with notify-and-fetch; Phase 2 then deleted notify-and-fetch
itself, so the 128 KB Akka frame constraint no longer applies to that hop in
any form. Successor ceiling recorded (4 MB gRPC cap via MaxBatchSize, which
batches by ROW COUNT), including that the failure mode differs — an oversized
gRPC message is rejected, not silently dropped.

Deployment docs gain the two operational constraints that have no home in code:
a site pair must be stopped and started TOGETHER (the SfBufferSnapshot compat
handler that made a mixed-version pair converge went with the replicator, and a
mixed pair now diverges silently), and a node offline beyond TombstoneRetention
can resurrect deleted rows on rejoin.

Both CLAUDE.md files corrected — each still said Phase 2 was NOT started.

Definition of done closed: build 0 warnings, all 10 suites green (3509 tests,
0 failures, 0 skips). Two DoD items needed amending rather than ticking: the
stale-symbol grep still matches 4 lines, all deliberate comment prose recording
what was deleted (a literal zero would delete the explanations that stop the
old design coming back), and the live gate has 10 evidence items, not the 9 the
checklist claimed.

Deletes the phase2 resume-state scratch doc, which said to delete it on landing.

Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts
2026-07-20 05:37:48 -04:00

58 KiB
Raw Permalink Blame History

ScadaBridge LocalDb Phase 2 Implementation Plan

For Claude: REQUIRED SUB-SKILL: Use superpowers-extended-cc:executing-plans to implement this plan task-by-task.

Goal: Move scadabridge.db's 9 config tables and StoreAndForward's sf_messages into the consolidated ZB.MOM.WW.LocalDb database established by Phase 1, then delete SiteReplicationActor and StoreAndForward's ReplicationService outright.

Architecture: Both stores stop owning their own SQLite files and take ILocalDb instead, exactly as OperationTrackingStore and SiteEventLogger did in Phase 1. Their DDL is extracted into *Schema.Apply(SqliteConnection) helpers that SiteLocalDbSetup.OnReady runs before RegisterReplicated. Trigger-based CDC then replicates every row mutation, replacing the bespoke hand-shipped ops; the library's per-row-LWW snapshot resync replaces the chunked anti-entropy.

Tech Stack: .NET 10, Akka.NET, Microsoft.Data.Sqlite, ZB.MOM.WW.LocalDb 0.1.0, gRPC, xunit.

Branch: feat/localdb-phase2, cut from feat/localdb-phase1.

Reviewed and corrected 2026-07-19 against the actual code (three verification sweeps over SiteRuntime, StoreAndForward, and Host/rig/docs, plus the LocalDb library source). Load-bearing corrections: D1 (the guarded write has a second surviving caller — SiteReconciliationActor — so it is NOT deleted), D3 (the active node already purges — Task 12 is now a pin, not a re-home), D6 (new: the 4 MB gRPC message cap × row-count-only batching), Task 1 (the Phase 2 tables aren't in the Phase 1 oplog — the soak measures legacy-DB write rates; port 8084; real metric names; no sqlite3 in containers), and Task 14 (do NOT register the two notification/SMTP tables). Details inline at each site.


Read this before Task 1

Phase 1's record is the prerequisite context. Read, in order:

  1. docs/plans/2026-07-19-localdb-adoption-phase1.md.tasks.json — the amendments, prerequisites and knownFlakes arrays. Several Phase 1 assumptions changed during execution.
  2. docs/plans/2026-07-19-localdb-phase2-gate.md — the gate. §5's open questions are answered in the Decisions section below; the gate itself is superseded by this plan.
  3. src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs — the pattern every new table follows.

Two known flakes — do not chase these

  • SiteRuntime.Tests InstanceActorChildAttributeRaceTests — intermittent ActorNotFoundException under full-suite load. Passes in isolation. Pre-existing.
  • AuditLog.Tests ParentExecutionIdCorrelationTests — ~91 s and times out on a cold MSSQL fixture, ~1 s warm. Re-run before investigating.

Decisions — the gate's open questions, answered

These were resolved from code recon before this plan was written. Each names its evidence. Task 2 records them formally; the rest of the plan assumes them.

D1. Config moves to CDC, and notify-and-fetch is deleted, not preserved

SiteReplicationActor never sends config JSON over the replication hop — it sends a deployment id plus fetch coordinates, and the standby HTTP-fetches from central itself (SiteReplicationActorTests #2, #5). That design exists for exactly one reason: the config blob exceeds Akka's 128 KB frame limit (docs/known-issues/2026-06-26-deploy-config-exceeds-akka-frame-size.md).

Under CDC the row itself replicates over the LocalDb gRPC stream, which has no such limit. The standby therefore never fetches at all.

The version guard, however, does not die with it. StoreDeployedConfigIfNewerAsync (SiteStorageService.cs:301-336) guards its upsert with:

WHERE excluded.deployed_at > deployed_configurations.deployed_at

That predicate protects against a stale fetch landing after a newer one, and the method has two production callers, not one: SiteReplicationActor.HandleApplyConfigDeploy (SiteReplicationActor.cs:375 — deleted in Task 15) and SiteReconciliationActor (SiteReconciliationActor.cs:166) — the per-node startup self-heal that reports local inventory to central and HTTP-fetches gap items. The reconciliation actor survives Phase 2, and its stale-fetch race (a reconcile fetch completing after a newer deploy landed) still exists. The method and its guard stay. Under CDC the reconcile write simply becomes a second CDC writer of deployed_configurations — benign: a gap fetch only fires when the node genuinely lacks the row, the content is deterministic per (deployment id, revision), and the guard blocks local downgrades.

Do not try to reproduce the deployed_at guard on top of LWW. They order by different clocks (central's deploy time vs. the writing node's HLC) and mixing them produces a non-convergent store.

D2. ReplaceAllAsync is deleted, and the N1 directional guard becomes unnecessary

StoreAndForwardStorage.ReplaceAllAsync (StoreAndForwardStorage.cs:284-306) is an unqualified DELETE FROM sf_messages followed by a full re-insert. Its own doc warns "Never call on an active node — it discards every in-flight row." SfBufferResyncPredicateTests exists to enforce that only a standby ever applies it — the N1 Critical regression test.

The LocalDb library's snapshot resync merges per row under LWW and never wipes:

  • SnapshotApplier.OnBeginAsync (SnapshotStreamer.cs:163-170) resets counters only. There is no wholesale DELETE FROM or TRUNCATE anywhere in the class.
  • OnBatchAsync (:172-186) wraps each snapshot row as an ordinary oplog entry and pushes it through the same LwwApplier as a delta.
  • LwwApplier.cs:69-78 discards the incoming row when the local row's HLC is higher.
  • Row-level deletes DO still replicate — the delete trigger captures a tombstone (TriggerSqlGenerator), SnapshotStreamer streams tombstone row-versions, and LwwApplier applies an incoming tombstone as a real DELETE. Only the destructive whole-table replace is gone. Caveat: tombstones are pruned after LocalDb:Replication:TombstoneRetention (default 7 days) — a node offline longer than that can resurrect deleted rows on rejoin; Task 21 puts this in the runbook.

So a node with newer local rows keeps them. The failure the directional guard prevents — a stale peer wiping a live buffer — is structurally impossible, not merely guarded against. Port SfBufferResyncPredicateTests as a convergence assertion (both nodes end with the union, newest per id wins), not as a directional-authority assertion.

Consequence to state plainly: the standby is no longer guaranteed byte-identical to the active node's buffer. It is guaranteed convergent. That is a real semantic change and Task 2 records it.

D3. The SMTP purge already runs on the active node — pin it, don't move it

PurgeCentralOnlyNotificationConfigAsync (SiteStorageService.cs:811-821) deletes notification_lists and smtp_configurations — including plaintext SMTP passwords left by pre-fix deployments. The original recon placed its only caller on the standby's replication path; that was wrong. It has two callers: DeploymentManagerActor.HandleDeployArtifacts (DeploymentManagerActor.cs:1921) — the active node's artifact-apply — and SiteReplicationActor.HandleApplyArtifacts (SiteReplicationActor.cs:456), the standby copy that dies with the actor. The purge therefore never lapses: the active-node call already exists and stays. Task 12 shrinks to pinning that call with a regression test (it sits inside a method Task 16 edits, and nothing today fails if the call is dropped).

Because no site code has written those two tables since 2026-07-10 (writers removed), the migrator skips them (Task 9), and the active purge keeps them empty, they are permanently empty in the consolidated DB — which is why Task 14 does not register them for replication (see the rationale there).

D4. native_alarm_state volume is measured, not assumed

scadabridge.db is not only config. native_alarm_state mirrors live A&C conditions from NativeAlarmActor.cs:504 via batched upserts, and is by a wide margin the highest-volume table in either database. sf_messages worst case is ~50 row-writes/sec (DefaultMaxRetries = 50, 10 s sweep, 4-way target parallelism). Neither number can be turned into an oplog cap from first principles.

Task 1 is a rig soak that measures both. Its output sets MaxOplogRows / MaxOplogAge in Task 19. If the soak shows the shared oplog cannot absorb alarm-storm churn, the escape hatch is keyed instances — named in the adoption design doc (~/Desktop/scadaproj/docs/plans/2026-07-19-scadabridge-localdb-design.md:139-141, "Out of scope"; not the 07-17 library design doc) — that is a library-level effort and would suspend this plan, so Task 1 gates everything after Task 2.

D5. Cutover forecloses rolling site upgrades

SiteReplicationActor retains a legacy monolithic SfBufferSnapshot handler specifically for rolling upgrades. With no dual-mechanism period, a site pair cannot be upgraded one node at a time — one node would speak a protocol the other no longer implements. Both nodes of a site must be stopped and started together. Task 21 puts this in the deployment runbook.

D6. The 4 MB gRPC message cap replaces the 128 KB Akka frame as the size ceiling

The Akka frame limit that motivated notify-and-fetch is gone, but LocalDb sync has its own ceiling the plan must respect: neither side configures gRPC message sizes — ScadaBridge's AddGrpc (Program.cs:519-521) sets only the auth interceptor, and the library's initiator channel is a bare GrpcChannel.ForAddress — so the 4 MB default receive limit applies in both directions. Batching is row-count-only (MaxBatchSize, default 500 — SyncSession.cs:227, SnapshotStreamer.cs:55; there is no byte-aware chunking). deployed_configurations.config_json is documented to exceed 128 KB per row; a few dozen such rows in one batch is over 4 MB, which fails the stream and wedges replication on a poison batch.

Task 1 measures the real max/typical config_json size on the rig's legacy DB; Task 19 sets LocalDb:Replication:MaxBatchSize so max-row-bytes × MaxBatchSize stays comfortably under 4 MB. If any single row approaches 4 MB, stop: that needs a LocalDb library change (byte-aware batching or configurable message sizes) — a scadaproj effort, same class as the keyed-instances hatch.


Execution waves

Wave Tasks Note
0 1, 2 Soak + decision record. Task 1 gates the rest.
1 3, 4 Schema extraction — independent files, dispatch in parallel
2 5, 6 Store rewiring — independent files, dispatch in parallel
3 7, 8, 9 Setup + migrators
4 10, 11 Port test intents as specs — before any deletion
5 12, 13 Pin the active-node purge; notify-and-fetch scope check
6 14, 15, 16, 17 The cutover, serial
7 18, 19 Convergence suite + rig config
8 20, 21 Live gate + docs

Task 1: Rig soak — measure oplog growth under real write rates

Classification: high-risk Estimated implement time: ~40 min wall-clock (mostly waiting) Parallelizable with: none — this gates the plan

This is the gate item from 2026-07-19-localdb-phase2-gate.md §5. It is measurement, not code.

Files:

  • Create: docs/plans/2026-07-19-localdb-phase2-soak.md

Method note (corrected): on this rig the Phase 2 tables are still in the legacy DBs — sf_messages in store-and-forward.db, native_alarm_state in scadabridge.db — and are not registered with LocalDb, so driving S&F/alarm churn moves __localdb_oplog not at all. This soak therefore measures (a) source write rates from the legacy DBs and derives the oplog sizing arithmetically, and (b) that the Phase 1 oplog stays healthy alongside. The empirical post-cutover drain check lives in Task 20 (evidence item 10).

Also: the containers have no sqlite3 binary (bare aspnet:10.0 image), and the data dirs are host bind mounts — but NEVER run sqlite3 on the macOS host against the live files: host↔container POSIX locks do not propagate across the virtiofs boundary, and a single host-side read checkpoints + resets the WAL out from under the container, permanently poisoning its connections (SQLITE_IOERR_SHORT_READdisk I/O error storm until restart). This is the root cause of the 2026-07-20 "disk I/O error under load" incident — see docs/known-issues/2026-07-20-localdb-disk-io-error-under-load.md §0. Safe sampling: copy the file triplet first (cp x.db x.db-wal x.db-shm /tmp/… — plain cp takes no SQLite locks and cannot harm the live DB), then query the copy. The helper used below:

snap() {  # snap <node-dir> <db> "<sql>"  — query a throwaway copy, never the live file
  local d; d=$(mktemp -d)
  cp "$1/$2" "$d/" && cp "$1/$2-wal" "$1/$2-shm" "$d/" 2>/dev/null
  sqlite3 "$d/$2" "$3"; rm -rf "$d"
}

Metrics are on port 8084 inside the container (not published to the host).

Step 1: Bring up the rig with Phase 1 replication enabled

cd ~/Desktop/ScadaBridge     # already on feat/localdb-phase2 (= phase1 + docs)
bash docker/deploy.sh

Site-a's pair replicates; site-b/site-c deliberately do not (the default-OFF pin). Confirm:

docker exec scadabridge-site-a-a curl -s localhost:8084/metrics | grep '^localdb_'

Expected: localdb_* series present (localdb_oplog_depth, localdb_sync_* — the meter is ZB.MOM.WW.LocalDb.Replication). If absent, the meter allowlist regressed — see SiteServiceRegistration.ObservedMeters.

Step 2: Capture baselines (from throwaway copies — never the live files, see the snap helper above)

cd ~/Desktop/ScadaBridge/docker
snap site-a-node-a/data site-localdb.db "SELECT COUNT(*) FROM __localdb_oplog;"
snap site-a-node-a/data store-and-forward.db \
  "SELECT COUNT(*), COALESCE(SUM(retry_count),0) FROM sf_messages;"
snap site-a-node-a/data scadabridge.db "SELECT COUNT(*) FROM native_alarm_state;"

Record all values and the wall-clock time. (COUNT + SUM(retry_count) is the S&F row-write proxy: every failed attempt is one UPDATE that increments retry_count.)

Step 3: Drive sustained load for 30 minutes

Two generators, run concurrently:

  • Store-and-forward churn. Point a store-and-forward target at an address that refuses connections, then enqueue messages. Each failed attempt costs one UPDATE. With the default 10 s sweep this is ~1 write/message/10 s per lane.
  • Alarm churn. Drive A&C conditions on a deployed instance so native_alarm_state upserts fire. This is the number that actually matters — it is unbounded by design.

Step 4: Sample write rates every 5 minutes

cd ~/Desktop/ScadaBridge/docker
for i in $(seq 1 6); do
  echo "=== t+$((i*5))m ==="
  snap site-a-node-a/data store-and-forward.db \
    "SELECT COUNT(*), COALESCE(SUM(retry_count),0) FROM sf_messages;"
  snap site-a-node-a/data scadabridge.db \
    "SELECT COUNT(*), COUNT(CASE WHEN last_transition_at > datetime('now','-5 minutes') THEN 1 END) FROM native_alarm_state;"
  snap site-a-node-a/data site-localdb.db "SELECT COUNT(*) FROM __localdb_oplog;"
  docker exec scadabridge-site-a-a curl -s localhost:8084/metrics \
    | grep -E '^localdb_(oplog_depth|sync)'
  sleep 300
done

Per-interval deltas of COUNT + SUM(retry_count) give S&F row-writes/sec; the last_transition_at-window count approximates alarm upserts/sec (it undercounts refresh-only upserts — say so in the findings if alarm volume is near a threshold).

Step 5: Measure config_json sizes (D6)

snap site-a-node-a/data scadabridge.db \
  "SELECT COUNT(*), MAX(LENGTH(config_json)), AVG(LENGTH(config_json)) FROM deployed_configurations;"

If the rig's configs are trivially small, also run this against a production-representative legacy DB (wonder) before trusting the number.

Step 6: Record findings

Write docs/plans/2026-07-19-localdb-phase2-soak.md with:

  • rows/sec observed for sf_messages and for native_alarm_state, separately
  • max/avg config_json bytes and the derived MaxBatchSize (max-row-bytes × MaxBatchSize comfortably under 4 MB — D6)
  • Phase 1 oplog depth over the run (should stay near zero — it is not driven by this load)
  • dead-letter count (must be 0)
  • the recommended MaxOplogRows and MaxOplogAge, with the arithmetic (sustained rows/sec × retention window, with alarm-storm headroom)

Step 7: The decision gate

If the measured sustained write rate × any reasonable MaxOplogAge exceeds any reasonable MaxOplogRows — i.e. the shared oplog cannot absorb this write profile even on paper — stop. Do not proceed to Task 3. Report to the user: the keyed-instances escape hatch (~/Desktop/scadaproj/docs/plans/2026-07-19-scadabridge-localdb-design.md:139-141) is required first, and that is a library effort in scadaproj. Likewise stop if any single config_json approaches 4 MB (D6 — needs byte-aware batching in the library). The empirical drain-under-churn confirmation happens post-cutover in Task 20; monotonic growth there is the same stop condition.

Step 8: Commit

git add docs/plans/2026-07-19-localdb-phase2-soak.md
git commit -m "docs(localdb): phase 2 rig soak findings — oplog sizing evidence"

(The feat/localdb-phase2 branch already exists — it carries this plan — so commit to it; do not git checkout -b.)


Task 2: Decision record

Classification: trivial Estimated implement time: ~3 min Parallelizable with: none (needs Task 1's numbers)

Files:

  • Modify: docs/plans/2026-07-19-localdb-phase2-gate.md

Update the gate's status header from NOT STARTED to point at this plan, and answer each §5 question inline with a one-paragraph resolution referencing D1D5 above plus Task 1's measured numbers. Do not delete the questions — the reasoning is the value.

Commit: docs(localdb): close the phase 2 gate — questions answered, plan written


Task 3: Extract StoreAndForwardSchema.Apply

Classification: small Estimated implement time: ~4 min Parallelizable with: Task 4

Mirrors OperationTrackingSchema exactly. Read that file first — it is the reference.

Files:

  • Create: src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardSchema.cs
  • Modify: src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:32-117
  • Test: tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardSchemaTests.cs

Step 1: Write the failing test

[Fact]
public void Apply_IsIdempotent_AndCreatesEveryColumn()
{
    var path = Path.Combine(Path.GetTempPath(), $"sf-schema-{Guid.NewGuid():N}.db");
    try
    {
        using var connection = new SqliteConnection($"Data Source={path}");
        connection.Open();

        StoreAndForwardSchema.Apply(connection);
        StoreAndForwardSchema.Apply(connection);   // second run must not throw

        using var cmd = connection.CreateCommand();
        cmd.CommandText = "SELECT name FROM pragma_table_info('sf_messages')";
        var columns = new List<string>();
        using var reader = cmd.ExecuteReader();
        while (reader.Read()) columns.Add(reader.GetString(0));

        // 12 original + 4 additive
        Assert.Equal(16, columns.Count);
        Assert.Contains("last_attempt_at_ms", columns);
        Assert.Contains("parent_execution_id", columns);
    }
    finally
    {
        SqliteConnection.ClearAllPools();
        File.Delete(path);
    }
}

Step 2: Run it — expect FAIL (StoreAndForwardSchema does not exist)

dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj --filter "FullyQualifiedName~StoreAndForwardSchemaTests"

Step 3: Create the schema class

Move, verbatim: the DDL literal from StoreAndForwardStorage.cs:50-68, the four AddColumnIfMissingAsync calls (:79-92) and their helper (:125-142), the last_attempt_at_ms backfill (:98-105), and the due-index (:109-114).

Convert every *Async to sync (ExecuteNonQuery, ExecuteScalar) — OnReady is a synchronous Action<ILocalDb>. Depend only on Microsoft.Data.Sqlite, never on the LocalDb library.

Carry this comment onto the class, because it is the whole reason the file exists:

/// <para>
/// Deliberately depends only on <c>Microsoft.Data.Sqlite</c>, not on the LocalDb library.
/// The Host applies this DDL to a LocalDb-managed connection before
/// <c>RegisterReplicated</c> installs the capture triggers; nothing about the schema
/// itself is LocalDb-specific, and the store still calls it so a directly-constructed
/// store (tests, tooling) remains self-sufficient.
/// </para>

Step 4: Reduce StoreAndForwardStorage.InitializeAsync to a call

public async Task InitializeAsync()
{
    EnsureDatabaseDirectoryExists();
    await using var connection = await OpenConnectionAsync();
    StoreAndForwardSchema.Apply(connection);
}

Step 5: Run the S&F suite — expect PASS

dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj

Step 6: Commit

git add -A && git commit -m "refactor(sf): extract StoreAndForwardSchema.Apply from the storage class"

Task 4: Extract SiteStorageSchema.Apply

Classification: small Estimated implement time: ~5 min Parallelizable with: Task 3

Files:

  • Create: src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageSchema.cs
  • Modify: src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:58-196
  • Test: tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Persistence/SiteStorageSchemaTests.cs

Same shape as Task 3. Move the DDL literal (:79-158, all 9 tables), MigrateSchemaAsync (:167-181) and TryAddColumnAsync (:183-196), converting to sync.

Do not move the PRAGMA journal_mode=WAL at :68-76. LocalDb owns the connection's pragmas.

Test assertion: all 9 tables exist after Apply, and a second Apply does not throw. Name every table explicitly — a loop that asserts "9 tables" would pass if one were renamed:

foreach (var table in new[]
         {
             "deployed_configurations", "static_attribute_overrides", "shared_scripts",
             "external_systems", "database_connections", "notification_lists",
             "data_connection_definitions", "smtp_configurations", "native_alarm_state",
         })
{
    Assert.True(TableExists(connection, table), $"missing table: {table}");
}

Commit: refactor(site): extract SiteStorageSchema.Apply from SiteStorageService


Task 5: Rewire StoreAndForwardStorage onto ILocalDb

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 6

Follow OperationTrackingStore (Phase 1, commit 0b5e9b44) as the reference — it solved this exact problem.

Files:

  • Modify: src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:22-26,177-185
  • Modify: src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ServiceCollectionExtensions.cs:18-25
  • Test: tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ (fixtures)

Step 1: Replace the constructor

public StoreAndForwardStorage(ILocalDb localDb, ILogger<StoreAndForwardStorage> logger)
{
    ArgumentNullException.ThrowIfNull(localDb);
    ArgumentNullException.ThrowIfNull(logger);
    _localDb = localDb;
    _logger = logger;
}

Step 2: Replace OpenConnectionAsync

// CreateConnection returns an ALREADY-OPEN, pragma-configured connection with the
// zb_hlc_next() UDF registered. Calling OpenAsync on it throws, and a raw
// SqliteConnection would lack the UDF, making every capture trigger fail closed.
private SqliteConnection OpenConnection() => _localDb.CreateConnection();

Update all 18 call sites (:36, 238, 258, 286, 321, 378, 412, 449, 482, 503, 548, 572, 589, 619, 638, 659, 677, 694) from await using var connection = await OpenConnectionAsync(); to await using var connection = OpenConnection();.

Step 3: Delete EnsureDatabaseDirectoryExists (:152-168) — LocalDb owns the file.

Step 4: Update the DI registration

services.AddSingleton<StoreAndForwardStorage>(sp => new StoreAndForwardStorage(
    sp.GetRequiredService<ILocalDb>(),
    sp.GetRequiredService<ILogger<StoreAndForwardStorage>>()));

Step 5: Fix the test fixtures

Every test constructing StoreAndForwardStorage with a connection string now needs an ILocalDb. Phase 1 created tests/ZB.MOM.WW.ScadaBridge.SiteEventLogging.Tests/TestLocalDb.cs for exactly this — read it and add an equivalent (or reference it) in the S&F test project. It must use a real temp file (never :memory: — capture triggers need a durable file) and call SqliteConnection.ClearAllPools() before deleting, including the -wal and -shm sidecars.

Step 6: Full S&F suite — expect PASS

dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj

Commit: refactor(sf): StoreAndForwardStorage takes ILocalDb instead of a connection string


Task 6: Rewire SiteStorageService onto ILocalDb

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 5

Files:

  • Modify: src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:13-51
  • Modify: src/ZB.MOM.WW.ScadaBridge.SiteRuntime/ServiceCollectionExtensions.cs:36-42
  • Modify: src/ZB.MOM.WW.ScadaBridge.Host/SiteServiceRegistration.cs:70-71

Same shape as Task 5, but note the differences:

  • There are 21 inline new SqliteConnection(_connectionString) + OpenAsync() pairs, not a helper (lines 60, 206, 248, 311, 345, 386, 414, 443, 470, 498, 538, 583, 608, 640, 663, 692, 730, 769, 813, 837, 868 — counted before Task 4, which removes the :60 one inside InitializeAsync). Add a private SqliteConnection OpenConnection() => _localDb.CreateConnection(); and convert all of them.
  • public SqliteConnection CreateConnection() at :51 is a public escape hatch used by SiteExternalSystemRepository (the only repository consumer). Keep the member; change the body to => _localDb.CreateConnection();.
  • The busy-timeout floor (BusyTimeoutFloorSeconds, :36) and its SqliteConnectionStringBuilder normalization are deleted — LocalDb configures pragmas on every connection it hands out.
  • AddSiteRuntime(string siteDbConnectionString) loses its parameter. Keep the no-arg overload at :23-28 (it already exists) and delete the string one, updating SiteServiceRegistration.cs:70-71 to services.AddSiteRuntime();.

Watch for: transactions at :348 (RemoveDeployedConfigAsync) and :540 (UpsertNativeAlarmsAsync) must keep using connection.BeginTransaction() on the LocalDb connection, not ILocalDb.BeginTransactionAsync() — mixing the two would put the cascade's three deletes in different transactions.

Step: Run both affected suites

dotnet test tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests.csproj
dotnet test tests/ZB.MOM.WW.ScadaBridge.Host.Tests/ZB.MOM.WW.ScadaBridge.Host.Tests.csproj

Host.Tests will surface DI fixtures that now need LocalDb:Path — Phase 1 hit the identical failure (tasks.json Task 4 note). Add the config key to each failing fixture.

Commit: refactor(site): SiteStorageService takes ILocalDb instead of a connection string


Task 7: Extend SiteLocalDbSetup with the new DDL (not yet registered)

Classification: standard Estimated implement time: ~3 min Parallelizable with: none

Deliberately applies DDL without calling RegisterReplicated. The tables live in the consolidated file but are not yet captured, so the bespoke replicator keeps working unchanged and the tree stays green. Task 14 flips them on and deletes the bespoke path in one commit.

Files:

  • Modify: src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs:43-57
  • Test: tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbWiringTests.cs

Step 1: Extend OnReady

using (var connection = db.CreateConnection())
{
    OperationTrackingSchema.Apply(connection);
    SiteEventLogSchema.Apply(connection);
    SiteStorageSchema.Apply(connection);
    StoreAndForwardSchema.Apply(connection);
}

db.RegisterReplicated("OperationTracking");
db.RegisterReplicated("site_events");

// Phase 2 tables are created here but NOT yet registered — the bespoke replicator
// still owns them until the Task 14 cutover. Registering them now would double-
// replicate every row (harmless, both paths upsert, but it would mask a defect in
// either one).

SiteLocalDbLegacyMigrator.Migrate(db, config);

Step 2: Pin it

Add a test asserting all 4 Phase-1-and-2 table sets exist after OnReady, and that db.ReplicatedTables contains exactly OperationTracking and site_events — the "not yet" is the assertion that matters.

Commit: feat(localdb): create config + sf_messages tables in the consolidated DB


Task 8: Extend the legacy migrator for sf_messages

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 9

Read SiteLocalDbLegacyMigrator.cs in full first. MigrateTracking (:116-161) is the closest reference — like sf_messages, it has a native TEXT PK and needs no id synthesis.

Files:

  • Modify: src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbLegacyMigrator.cs:57-66
  • Test: tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbLegacyMigratorTests.cs

Step 1: Write the failing tests

Three, minimum:

[Fact] public void SfMessages_AreCopiedFromTheLegacyFile() { }
[Fact] public void SfMessages_Migration_IsIdempotent_WhenRerunAfterACrashBeforeRename() { }
[Fact] public void MigratedSfMessages_EnterTheOplog_SoTheyActuallyReplicate() { }

The third is the one that catches the ordering defect. Phase 1's equivalent (SiteLocalDbLegacyMigratorTests.cs:244, MigratedRows_EnterTheOplog_SoTheyActuallyReplicate) asserts on __localdb_oplog directly — copy that assertion. The oplog table is __localdb_oplog (LocalDbSchema.cs:19), not zb_oplog.

Harness note: production OnReady does not register sf_messages until the Task 14 cutover, so these tests must build their own ILocalDb (TestLocalDb), call RegisterReplicated("sf_messages") themselves, and then run the migrator — the oplog assertion needs live capture triggers. Correspondingly, Task 14 must keep SiteLocalDbLegacyMigrator.Migrate as the last call in OnReady, after all registrations.

Step 2: Add ResolveStoreAndForwardPath + MigrateStoreAndForward

private const string DefaultStoreAndForwardPath = "./data/store-and-forward.db";

internal static string ResolveStoreAndForwardPath(IConfiguration config) =>
    config["ScadaBridge:StoreAndForward:SqliteDbPath"] ?? DefaultStoreAndForwardPath;

Migrate all 16 columns with INSERT OR IGNORE. The PK is native TEXT — no id synthesis, unlike MigrateEvents.

Step 3: Wire into Migrate — add the call alongside the existing two.

Commit: feat(localdb): migrate legacy store-and-forward.db into the consolidated DB


Task 9: Extend the legacy migrator for the 9 config tables

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 8

Files:

  • Modify: src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbLegacyMigrator.cs
  • Test: tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbLegacyMigratorTests.cs

All 9 tables migrate from one legacy file (ScadaBridge:Database:SiteDbPath), so this is one MigrateSiteStorage that opens the legacy DB once and copies table by table inside a single transaction, then renames once.

Two tables are deliberately skipped: notification_lists and smtp_configurations. They are purged on every deploy and are permanently empty by design (site write paths were removed 2026-07-10). Migrating them would resurrect plaintext SMTP passwords from a pre-fix legacy file into a replicated table. Add this as a comment and as a test:

[Fact]
public void Migration_DoesNotCopyNotificationOrSmtpRows_EvenWhenTheLegacyFileHasThem()
{
    // These carry plaintext SMTP passwords in pre-2026-07-10 files. They are purged
    // on every deploy and must never enter a replicated table.
}

Step: Verify all 9 tables' columns match between SiteStorageSchema and the migrator's INSERT lists. A column-count mismatch here fails at runtime on a real deployment and nowhere else.

Commit: feat(localdb): migrate legacy scadabridge.db config tables into the consolidated DB


Task 10: Port the S&F replication test intents as CDC specs

Classification: standard Estimated implement time: ~5 min Parallelizable with: Task 11

Specifications first, deletion second. These tests encode behaviour the replacement must still satisfy. The gate document named two files; the real specification is five.

Files:

  • Read (as spec): tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationServiceTests.cs
  • Read (as spec): tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardReplicationTests.cs
  • Create: tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbStoreAndForwardConvergenceTests.cs

Port these intents (not the mechanics) as two-node convergence assertions, using Phase 1's LocalDbSitePairConvergenceTests.cs as the harness reference:

Intent Assertion under CDC
Add materialises on the peer enqueue on A → row present on B
Remove deletes on the peer deliver on A → row absent on B
Park sets status Parked on the peer park on A → B's row has status Parked
Requeue resets status + RetryCount requeue on A → B shows Pending, RetryCount 0
Add-then-Remove never converges to present the portable form of the strict-ordering test
Apply Add twice is idempotent, newest wins LWW gives this
Park when Add was lost still materialises upsert semantics

Explicitly do not port: ReplicationServiceTests.ReplicationOperations_AreDispatchedInIssueOrder (:169-195; 200 interleaved ops dispatched synchronously in strict issue order, 400 observed). That asserts the mechanism — inline fire-and-forget dispatch — not the outcome. CDC capture is asynchronous and batched by construction. The portable intent is row 5 above. Record this in the test file as a comment so a future reader does not think it was dropped by accident.

Commit: test(localdb): port store-and-forward replication intents as CDC convergence specs


Task 11: Port the resync + directional-authority tests

Classification: standard Estimated implement time: ~5 min Parallelizable with: Task 10

Files:

  • Read (as spec): tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/SiteReplicationActorTests.cs
  • Read (as spec): tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/Cluster/SfBufferResyncPredicateTests.cs
  • Create: tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbConfigConvergenceTests.cs

SfBufferResyncPredicateTests is the most valuable test in the set — the N1 Critical regression. Per D2 it must be re-expressed, not ported literally:

[Fact]
public async Task ANodeWithNewerLocalRows_KeepsThem_WhenAPeerSnapshotArrives()
{
    // N1 Critical (was: SfBufferResyncPredicateTests). The bespoke replicator needed a
    // directional guard because ReplaceAllAsync was a destructive DELETE-then-INSERT,
    // so a wrong-direction resync wiped a live buffer.
    //
    // LocalDb's snapshot resync merges per row under LWW and never deletes
    // (SnapshotApplier has no DELETE; LwwApplier.cs:69-78 discards an incoming row
    // whose HLC is lower than the local row's). The failure mode is structurally
    // impossible rather than guarded against — so this asserts convergence, NOT
    // directional authority. There is no active/standby asymmetry left to enforce.
}

Also do not port: actor tests #1#5 (notify-and-fetch config fetch/retry/supersede). Per D1 the fetch path is deleted, so these describe code that will not exist. Replace with a single assertion that a config deploy on A converges to B without B making any HTTP call — assert on a fetcher test double that records zero invocations. That is the positive proof that notify-and-fetch is gone.

Scope the zero-fetch assertion to the deploy-replication flow. SiteReconciliationActor survives Phase 2 and legitimately HTTP-fetches at node startup when central reports gaps (D1). The two-node harness has no central, so its pass no-ops there — but do not write the assertion as "the standby never fetches, ever"; assert the fetcher double records zero calls during deploy convergence.

Commit: test(localdb): port resync + config replication intents as CDC convergence specs


Task 12: Pin the active-node SMTP purge

Classification: standard Estimated implement time: ~3 min Parallelizable with: none

Per D3 (corrected): the active node's artifact-apply already calls PurgeCentralOnlyNotificationConfigAsyncDeploymentManagerActor.cs:1921, inside HandleDeployArtifacts (:1864-1963). Nothing is re-homed. What is missing is a pin: Task 16 edits this actor's wiring, and no test today fails if the purge call is dropped (ArtifactStorageTests covers the storage method, not the actor's call site).

Files:

  • Test: tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/DeploymentManagerActorTests.cs

Step 1: Write the pin test (red-first by commenting out the call locally, then restore)

[Fact]
public async Task ApplyingArtifacts_PurgesCentralOnlyNotificationConfig()
{
    // Security cleanup: notification_lists and smtp_configurations can hold plaintext
    // SMTP passwords from pre-2026-07-10 deployments. The ACTIVE node's artifact-apply
    // purges them (DeploymentManagerActor.HandleDeployArtifacts); the standby's copy of
    // this call dies with SiteReplicationActor in Task 15. This test pins the active-node
    // call so Task 16's actor edits cannot silently drop it.
    // Arrange: real temp-file SiteStorageService, seed one row into each table,
    // drive HandleDeployArtifacts; assert both tables are empty afterwards.
}

Step 2: No production change is expected. If recon of the call site shows otherwise, stop and report.

Commit: test(site): pin the active-node notification-config purge


Task 13: Notify-and-fetch scope check — the guarded write stays

Classification: standard Estimated implement time: ~3 min Parallelizable with: none

Per D1 (corrected). The original task deleted StoreDeployedConfigIfNewerAsync here — that was wrong twice over: (a) it has a second surviving caller, and (b) deleting it before Task 15 would not even compile, since SiteReplicationActor still calls it until then.

Files:

  • Modify: src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:301-336 (doc comment only)

Step 1: Keep StoreDeployedConfigIfNewerAsync. Its callers are SiteReplicationActor.HandleApplyConfigDeploy (SiteReplicationActor.cs:375 — dies with the actor in Task 15) and SiteReconciliationActor (SiteReconciliationActor.cs:166) — the per-node startup self-heal against central, which survives Phase 2 and still needs the deployed_at guard against a stale reconcile fetch. Update the method's doc comment to name reconciliation as the remaining caller (and drop any mention of standby replication). The unguarded StoreDeployedConfigAsync (:241-271) stays the active deploy path.

Step 2: ConfigFetchRetryCount — answered, but its removal moves to Task 17.

The grep was run during plan review: the only production reader is SiteReplicationActor.cs:157 (plus its validator rule and tests). SiteReconciliationActor does a single best-effort pass and does not read it; IDeploymentConfigFetcher itself is kept (used by DeploymentManagerActor's refresh path, SiteReconciliationActor, and registered at ServiceCollectionExtensions.cs:84). So the option (SiteRuntimeOptions.cs:68) and its validator rule (SiteRuntimeOptionsValidator.cs:56-58) become dead after Task 15 deletes the actor — remove them in Task 17 (config-key cleanup), not here, or the build breaks. Re-run grep -rn "ConfigFetchRetryCount" src/ tests/ there to confirm.

Commit: docs(site): StoreDeployedConfigIfNewerAsync serves reconciliation — guard stays under CDC


Task 14: Register the Phase 2 tables and delete ReplicationService

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: none

The cutover. Registration and bespoke-path deletion land in one commit so the two mechanisms never both run.

Files:

  • Modify: src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs
  • Delete: src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ReplicationService.cs
  • Modify: src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardService.cs — field :39, ctor param :243, and the 6 emission call sites :654, 805, 836, 872, 1120, 1146
  • Modify: src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ServiceCollectionExtensions.cs:63-68 (the ReplicationService registration) and :32 (GetRequiredService<ReplicationService>() inside the StoreAndForwardService factory at :27-61)
  • Modify: src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:284-306
  • Delete: tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationServiceTests.cs
  • Delete: tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardReplicationTests.cs
  • Delete: tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationWireSerializationPinTests.cs

Step 1: Register 8 tables in SiteLocalDbSetup.OnReady, replacing the Task 7 comment: the 7 config tables (deployed_configurations, static_attribute_overrides, shared_scripts, external_systems, database_connections, data_connection_definitions, native_alarm_state) plus sf_messages. Do NOT register notification_lists or smtp_configurations (this reverses the plan's original instruction): they are permanently empty by design — no site writer since 2026-07-10, the migrator skips them (Task 9), and the active-node purge (DeploymentManagerActor.cs:1921) keeps them empty — so registering them would create a standing replication channel whose only historical payload was plaintext SMTP passwords, for zero benefit. Put that rationale in an OnReady comment so a future writer has to make an explicit decision.

Keep SiteLocalDbLegacyMigrator.Migrate as the last call in OnReady, after all registrations, so migrated rows enter the oplog through live capture triggers (the Phase 1 ordering that is silently load-bearing).

Note the two composite-PK tables are fine: LocalDb's RegisterReplicated supports multi-column PKs (SqliteLocalDb.RegisterReplicated orders pk ordinals), and no Phase 2 table has a BLOB column (which would be rejected).

Step 2: Delete ReplicationService and remove the ctor parameter (:243), the _replication field (:39), and all 6 emission call sites from StoreAndForwardService.

Step 3: Delete ReplaceAllAsync (:284-306). Per D2 the library's resync merges; a destructive delete-all must not survive into a replicated table, where CDC would capture the mass-delete and ship it to the peer.

Step 4: Check GetAllMessagesAsync's Truncated flag — it exists only for the resync path. If nothing else reads it, simplify the return type.

Step 5: CompositionRootTests.cs:474 asserts ReplicationService is a registered site singleton. Delete that assertion.

Step 6: Full S&F + Host suites — expect PASS

Commit: feat(localdb)!: replicate sf_messages via CDC, delete ReplicationService


Task 15: Delete SiteReplicationActor and its messages

Classification: high-risk Estimated implement time: ~4 min Parallelizable with: none

Files:

  • Delete: src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Actors/SiteReplicationActor.cs
  • Delete: src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Messages/ReplicationMessages.cs
  • Delete: tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/SiteReplicationActorTests.cs
  • Delete: tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/ResyncWireSerializationPinTests.cs
  • Delete: tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/Cluster/SfBufferResyncPredicateTests.cs

The whole ReplicationMessages.cs file becomes dead — it contains exactly the 10 records ReplicateConfigDeploy/Remove/SetEnabled, ReplicateArtifacts, ReplicateStoreAndForward, ApplyConfigDeploy/Remove/SetEnabled, ApplyArtifacts, ApplyStoreAndForward, whose only consumers are SiteReplicationActor, the 5 DeploymentManagerActor Tell sites (Task 16), the AkkaHostedService wiring (Task 16), and the deleted tests. The four resync records (RequestSfBufferResync, SfBufferSnapshot, SfBufferSnapshotChunk, SfBufferResyncAck) are not in this file — they are declared at the bottom of SiteReplicationActor.cs (:678-707) and die with the actor file.

Do not delete ActiveNodeEvaluator (Communication/ClusterState/ActiveNodeEvaluator.cs) — the S&F delivery gate still needs it (AkkaHostedService.cs:866SetDeliveryGateSelfIsPrimaryActiveNodeEvaluator.SelfIsOldestUp). Its doc comment at :14 mentions SiteReplicationActor and :16 mentions the resync authority checks; the actor's own direct call at SiteReplicationActor.cs:288 goes away with the file. Update the comment, keep the type.

Confirm the deletions are complete before committing:

grep -rn "SiteReplicationActor\|ReplicateStoreAndForward\|SfBufferSnapshot" src/ tests/

Expected: no matches.

Commit: feat(localdb)!: delete SiteReplicationActor — CDC replaces hand-shipped ops


Task 16: Clean up DeploymentManagerActor and AkkaHostedService

Classification: standard Estimated implement time: ~4 min Parallelizable with: none

Files:

  • Modify: src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Actors/DeploymentManagerActor.cs:55,161-190,794,970,1004,1102,1952
  • Modify: src/ZB.MOM.WW.ScadaBridge.Host/Actors/AkkaHostedService.cs:770-796,806-812,1056

Step 1: Remove the IActorRef? replicationActor = null ctor parameter (:169 — the ctor signature spans :161-175; :184 is the field assignment), the _replicationActor field (:55), and all 5 _replicationActor?.Tell(...) sites (:794, 970, 1004, 1102, 1952).

The parameter is optional and positional — removing it silently shifts every argument after it (healthCollector, serviceProvider, loggerFactory, configFetcher, startupLoadRetryInterval, configLoader). The sharpest hazard is the same-typed IActorRef? optional immediately before itdclManager at :168: a caller passing positionally could silently feed the wrong actor ref with zero compile errors. Update the Props.Create at AkkaHostedService.cs:806-812 (which passes replicationActor positionally at :810) and every test constructing this actor — check each call site by hand.

Step 2: Delete replicationService / replicationLogger resolution (:770-773), the actor creation (:785-789), the handler wiring (:792-796), and the actor-inventory comment at :1056.

Step 3: activeNodeCheck (:781-783) is still used by SiteCommunicationActor (:816-821). Keep it.

Step 4: Full Host + SiteRuntime + IntegrationTests suites.

Commit: refactor(site): drop the replication actor from the deployment + host wiring


Task 17: Config-key cleanup

Classification: standard Estimated implement time: ~4 min Parallelizable with: none

Files:

  • Modify: src/ZB.MOM.WW.ScadaBridge.Host/StartupValidator.cs:113-115
  • Modify: src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardOptions.cs:12
  • Modify: src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardOptionsValidator.cs:19-21
  • Modify: src/ZB.MOM.WW.ScadaBridge.SiteRuntime/SiteRuntimeOptions.cs:68 + SiteRuntimeOptionsValidator.cs:56-58 (moved here from Task 13)
  • Modify: 10 × appsettings.Site.json (Host template + 6 docker/site-* + 2 docker-env2/site-x-*
    • deploy/wonder-app-vd03/appsettings.Site.json — the wonder single-node install, which the original count missed; it sets both paths and ReplicationEnabled: false)
Key Action
ScadaBridge:Database:SiteDbPath → migration-only. Relax the Require at StartupValidator.cs:113 or every site config keeps it mandatory forever.
ScadaBridge:StoreAndForward:SqliteDbPath → migration-only. Relax the non-empty rule at StoreAndForwardOptionsValidator.cs:19-21.
ScadaBridge:StoreAndForward:ReplicationEnabled fully dead. Delete the property and every config entry (including wonder's).
SiteRuntime … ConfigFetchRetryCount dead after Task 15 (only reader was SiteReplicationActor.cs:157; see Task 13). Delete the option, its validator rule, and any config entries — confirm with grep -rn "ConfigFetchRetryCount" src/ tests/ first.

ParkedOperationRelayTests.cs:39 and ParkedMessageHandlerActorTests.cs:34 set ReplicationEnabled = false — they only need the flag to exist, so remove those lines. StoreAndForwardOptionsTests.cs touches it three times: :14 asserts the default, :26 sets and :33 asserts a customized value — remove all three references.

Leave the path keys present in configs with a comment — they are read by the legacy migrator and removing them would strand un-migrated data on a node that has not yet started once.

Commit: chore(config): retire ReplicationEnabled, make legacy db paths migration-only


Task 18: Two-node convergence suite for the Phase 2 tables

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 19

Extend Phase 1's LocalDbSitePairConvergenceTests.cs harness (real loopback Kestrel h2c + the real LocalDbSyncAuthInterceptor, initialized via SiteLocalDbSetup.OnReady).

Files:

  • Modify: tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbSitePairConvergenceTests.cs

Scenarios, beyond Tasks 1011:

  1. A config deploy on A converges to B, and B's deployed_configurations row matches byte for byte.
  2. RemoveDeployedConfigAsync's 3-table cascade converges — static_attribute_overrides and native_alarm_state orphans do not survive on B. There are no foreign keys (SiteStorageService DDL has none), so this is three independent delete streams that LWW may reorder. This is the most likely real defect in the whole plan.
  3. A burst of native_alarm_state upserts converges and the oplog drains.
  4. Node B restarts and catches up via snapshot resync without losing rows written on B while down.

Verify the suite is non-vacuous. (Corrected claim: Phase 1's convergence suite did not do a mismatched-key run — it uses one shared key, LocalDbSitePairConvergenceTests.cs:47; wrong-key denial is unit-covered separately in Host.Tests/LocalDbSyncAuthInterceptorTests.cs.) Prove it directly for the new scenarios: run once with the new table registrations disabled on node B (or a deliberately wrong key) and confirm every new scenario goes red, then restore and record the result in the commit message. A convergence suite that cannot fail proves nothing.

Commit: test(localdb): two-node convergence for config tables + sf_messages


Task 19: Rig configuration

Classification: small Estimated implement time: ~3 min Parallelizable with: Task 18

Files:

  • Modify: docker/site-a-node-a/appsettings.Site.json, docker/site-a-node-b/appsettings.Site.json

Set LocalDb:Replication:MaxOplogRows and LocalDb:Replication:MaxOplogAge (exact key paths — the lib binds the LocalDb:Replication section; defaults 1,000,000 rows / 7 days) from Task 1's measured numbers — not from a guess. Per D6, also set LocalDb:Replication:MaxBatchSize (default 500) so max-config_json-bytes × MaxBatchSize stays comfortably under the 4 MB gRPC receive limit. Keep site-b/site-c unreplicated so default-OFF stays proven side by side on one rig.

Commit: chore(docker): size the site-a oplog caps from the phase 2 soak


Task 20: Live gate on the docker rig

Classification: high-risk Estimated implement time: ~30 min wall-clock Parallelizable with: none

Files:

  • Create: docs/plans/2026-07-19-localdb-phase2-live-gate.md
cd ~/Desktop/ScadaBridge && bash docker/deploy.sh

Tooling: the containers have no sqlite3 — run DB checks host-side against the bind mounts (docker/site-a-node-{a,b}/data/); metrics via docker exec scadabridge-site-a-a curl -s localhost:8084/metrics (port 8084, not published to the host).

Evidence required — each must be captured, not asserted:

  1. Migration ran. store-and-forward.db.migrated and scadabridge.db.migrated exist on both site-a nodes; row counts in the consolidated DB match the legacy files. (Both nodes migrate their own legacy files independently; identical-content rows then LWW-converge — expected, not an anomaly.)
  2. No SMTP/notification rows migratedSELECT COUNT(*) FROM smtp_configurations is 0.
  3. Config converges. Deploy an instance to site-a; both nodes' deployed_configurations rows are byte-identical with identical __localdb_row_version and originating node id.
  4. The standby made no config HTTP fetch during the deploy (both nodes up) — grep the deploy window of its log. This is the positive proof that notify-and-fetch is gone. A SiteReconciliationActor fetch at node startup is legitimate (it survives Phase 2) and does not fail this check.
  5. S&F converges. Enqueue against a dead target; the parked row appears on the peer.
  6. Site failover. Flip the site-a pair and confirm buffered messages survive and deliver exactly once after the flip. docker/failover-drill.sh targets CENTRAL nodes and does not exercise this — run the site flip directly.
  7. Cascade delete. Remove a deployed config; confirm no orphan static_attribute_overrides or native_alarm_state rows on the peer.
  8. Zero dead letters, oplog drained, localdb_* scraped from /metrics.
  9. Both nodes stopped and started together (D5) — confirm a clean rejoin.
  10. Post-cutover drain under churn — the empirical half Task 1 could not measure pre-cutover: re-run Task 1's two load generators against this build and watch localdb_oplog_depth (and SELECT COUNT(*) FROM __localdb_oplog) rise and drain between bursts. Monotonic growth that never drains is the same stop condition as Task 1 step 7 (keyed-instances hatch).

If any check fails, stop and report. Do not proceed to Task 21.

Commit: docs(localdb): phase 2 live gate evidence


Task 21: Documentation truth pass

Classification: standard Estimated implement time: ~5 min Parallelizable with: none

Files:

  • Modify: CLAUDE.md (ScadaBridge), ~/Desktop/scadaproj/CLAUDE.md (LocalDb component row)
  • Modify: docs/requirements/Component-StoreAndForward.md:83
  • Modify: docs/components/StoreAndForward.md, SiteRuntime.md, Host.md
  • Modify: docs/known-issues/2026-06-26-deploy-config-exceeds-akka-frame-size.md
  • Modify: docs/deployment/installation-guide.md + topology-guide.md (+ production-checklist.md if it covers upgrades) — there is no file literally named "runbook"

Specific corrections:

  • Component-StoreAndForward.md:83 is a long normative paragraph specifying the whole chunked resync protocol and the N5 race. Rewrite for CDC — state the new duplicate-delivery bound explicitly rather than deleting the discussion.
  • The frame-size known-issue is already marked "Status: RESOLVED (2026-06-26)" (resolved by the notify-and-fetch rework). Amend it to record that Phase 2 deleted notify-and-fetch itself — CDC ships the row over gRPC, so the frame constraint is gone entirely — and note the successor ceiling: the 4 MB gRPC message cap managed via MaxBatchSize (D6).
  • Add to the deployment docs (D5): a site pair must be stopped and started together. Rolling one node at a time is no longer supported, because the legacy SfBufferSnapshot compatibility handler is gone. Also record the D2 caveat: a node offline longer than LocalDb:Replication:TombstoneRetention (default 7 days) can resurrect deleted rows on rejoin.
  • Update the scadaproj LocalDb component row: Phase 2 complete, what replicates now, what SiteReplicationActor used to do.

Commit: docs(localdb): phase 2 truth pass across both repos


Definition of done

All met, 2026-07-20.

  • Solution builds with 0 warnings (TreatWarningsAsErrors is on)
  • Every suite green — Host 330, CentralUI 925, Commons 684, SiteRuntime 512, AuditLog 355, Communication 312, StoreAndForward 130, HealthMonitoring 97, SiteEventLogging 70, IntegrationTests 94. 0 failures, 0 skips.
  • grep -rn "SiteReplicationActor\|ReplicationService\|SfBufferSnapshot" src/no code references. Four matches remain and are all deliberate: prose inside /// and // comments in ActiveNodeEvaluator, SiteLocalDbSetup and AkkaHostedService that records what was deleted and why. The check as written is too blunt — a literal "no matches" would force deleting the very explanations that stop someone re-introducing the old design. Read it as "no references from code".
  • Live gate PASS — 10 evidence items, not the 9 this checklist says (Task 20 itself lists ten). All captured in docs/plans/2026-07-19-localdb-phase2-live-gate.md, with two scope caveats recorded there rather than glossed.
  • Convergence suite verified non-vacuous — with the eight RegisterReplicated calls commented out, all four Task 18 scenarios go red (4 failed / 0 passed); restored, 20/20 across the three LocalDb suites.
  • Both CLAUDE.md files updated