# ScadaBridge LocalDb Phase 2 Implementation Plan > **For Claude:** REQUIRED SUB-SKILL: Use superpowers-extended-cc:executing-plans to implement this plan task-by-task. **Goal:** Move `scadabridge.db`'s 9 config tables and StoreAndForward's `sf_messages` into the consolidated `ZB.MOM.WW.LocalDb` database established by Phase 1, then delete `SiteReplicationActor` and StoreAndForward's `ReplicationService` outright. **Architecture:** Both stores stop owning their own SQLite files and take `ILocalDb` instead, exactly as `OperationTrackingStore` and `SiteEventLogger` did in Phase 1. Their DDL is extracted into `*Schema.Apply(SqliteConnection)` helpers that `SiteLocalDbSetup.OnReady` runs before `RegisterReplicated`. Trigger-based CDC then replicates every row mutation, replacing the bespoke hand-shipped ops; the library's per-row-LWW snapshot resync replaces the chunked anti-entropy. **Tech Stack:** .NET 10, Akka.NET, Microsoft.Data.Sqlite, ZB.MOM.WW.LocalDb 0.1.0, gRPC, xunit. **Branch:** `feat/localdb-phase2`, cut from `feat/localdb-phase1`. > **Reviewed and corrected 2026-07-19** against the actual code (three verification sweeps over > SiteRuntime, StoreAndForward, and Host/rig/docs, plus the LocalDb library source). Load-bearing > corrections: D1 (the guarded write has a second surviving caller — `SiteReconciliationActor` — > so it is NOT deleted), D3 (the active node already purges — Task 12 is now a pin, not a > re-home), D6 (new: the 4 MB gRPC message cap × row-count-only batching), Task 1 (the Phase 2 > tables aren't in the Phase 1 oplog — the soak measures legacy-DB write rates; port 8084; real > metric names; no `sqlite3` in containers), and Task 14 (do NOT register the two > notification/SMTP tables). Details inline at each site. --- ## Read this before Task 1 Phase 1's record is the prerequisite context. Read, in order: 1. `docs/plans/2026-07-19-localdb-adoption-phase1.md.tasks.json` — the `amendments`, `prerequisites` and `knownFlakes` arrays. Several Phase 1 assumptions changed during execution. 2. `docs/plans/2026-07-19-localdb-phase2-gate.md` — the gate. §5's open questions are answered in the Decisions section below; the gate itself is superseded by this plan. 3. `src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs` — the pattern every new table follows. ### Two known flakes — do not chase these - `SiteRuntime.Tests InstanceActorChildAttributeRaceTests` — intermittent `ActorNotFoundException` under full-suite load. Passes in isolation. Pre-existing. - `AuditLog.Tests ParentExecutionIdCorrelationTests` — ~91 s and times out on a **cold** MSSQL fixture, ~1 s warm. Re-run before investigating. --- ## Decisions — the gate's open questions, answered These were resolved from code recon before this plan was written. Each names its evidence. **Task 2 records them formally; the rest of the plan assumes them.** ### D1. Config moves to CDC, and notify-and-fetch is deleted, not preserved `SiteReplicationActor` never sends config JSON over the replication hop — it sends a deployment id plus fetch coordinates, and the standby HTTP-fetches from central itself (`SiteReplicationActorTests` #2, #5). That design exists for exactly one reason: the config blob exceeds Akka's 128 KB frame limit (`docs/known-issues/2026-06-26-deploy-config-exceeds-akka-frame-size.md`). Under CDC the row itself replicates over the LocalDb gRPC stream, which has no such limit. The standby therefore never fetches at all. The version guard, however, does **not** die with it. `StoreDeployedConfigIfNewerAsync` (`SiteStorageService.cs:301-336`) guards its upsert with: ```sql WHERE excluded.deployed_at > deployed_configurations.deployed_at ``` That predicate protects against **a stale fetch landing after a newer one**, and the method has **two** production callers, not one: `SiteReplicationActor.HandleApplyConfigDeploy` (`SiteReplicationActor.cs:375` — deleted in Task 15) *and* `SiteReconciliationActor` (`SiteReconciliationActor.cs:166`) — the per-node startup self-heal that reports local inventory to central and HTTP-fetches gap items. The reconciliation actor **survives Phase 2**, and its stale-fetch race (a reconcile fetch completing after a newer deploy landed) still exists. **The method and its guard stay.** Under CDC the reconcile write simply becomes a second CDC writer of `deployed_configurations` — benign: a gap fetch only fires when the node genuinely lacks the row, the content is deterministic per (deployment id, revision), and the guard blocks local downgrades. > **Do not** try to reproduce the `deployed_at` guard on top of LWW. They order by different > clocks (central's deploy time vs. the writing node's HLC) and mixing them produces a > non-convergent store. ### D2. `ReplaceAllAsync` is deleted, and the N1 directional guard becomes unnecessary `StoreAndForwardStorage.ReplaceAllAsync` (`StoreAndForwardStorage.cs:284-306`) is an unqualified `DELETE FROM sf_messages` followed by a full re-insert. Its own doc warns "Never call on an active node — it discards every in-flight row." `SfBufferResyncPredicateTests` exists to enforce that only a standby ever applies it — the N1 Critical regression test. The LocalDb library's snapshot resync **merges per row under LWW and never wipes**: - `SnapshotApplier.OnBeginAsync` (`SnapshotStreamer.cs:163-170`) resets counters only. There is no wholesale `DELETE FROM` or `TRUNCATE` anywhere in the class. - `OnBatchAsync` (`:172-186`) wraps each snapshot row as an ordinary oplog entry and pushes it through the same `LwwApplier` as a delta. - `LwwApplier.cs:69-78` discards the incoming row when the local row's HLC is higher. - **Row-level deletes DO still replicate** — the delete trigger captures a tombstone (`TriggerSqlGenerator`), `SnapshotStreamer` streams tombstone row-versions, and `LwwApplier` applies an incoming tombstone as a real `DELETE`. Only the destructive whole-table replace is gone. Caveat: tombstones are pruned after `LocalDb:Replication:TombstoneRetention` (default 7 days) — a node offline longer than that can resurrect deleted rows on rejoin; Task 21 puts this in the runbook. So a node with newer local rows keeps them. The failure the directional guard prevents — a stale peer wiping a live buffer — is structurally impossible, not merely guarded against. **Port `SfBufferResyncPredicateTests` as a convergence assertion (both nodes end with the union, newest per id wins), not as a directional-authority assertion.** Consequence to state plainly: the standby is no longer guaranteed byte-identical to the active node's buffer. It is guaranteed *convergent*. That is a real semantic change and Task 2 records it. ### D3. The SMTP purge already runs on the active node — pin it, don't move it `PurgeCentralOnlyNotificationConfigAsync` (`SiteStorageService.cs:811-821`) deletes `notification_lists` and `smtp_configurations` — including plaintext SMTP passwords left by pre-fix deployments. The original recon placed its only caller on the standby's replication path; **that was wrong.** It has two callers: `DeploymentManagerActor.HandleDeployArtifacts` (`DeploymentManagerActor.cs:1921`) — the **active** node's artifact-apply — and `SiteReplicationActor.HandleApplyArtifacts` (`SiteReplicationActor.cs:456`), the standby copy that dies with the actor. The purge therefore never lapses: the active-node call already exists and stays. Task 12 shrinks to **pinning that call with a regression test** (it sits inside a method Task 16 edits, and nothing today fails if the call is dropped). Because no site code has written those two tables since 2026-07-10 (writers removed), the migrator skips them (Task 9), and the active purge keeps them empty, they are **permanently empty in the consolidated DB** — which is why Task 14 does **not** register them for replication (see the rationale there). ### D4. `native_alarm_state` volume is measured, not assumed `scadabridge.db` is not only config. `native_alarm_state` mirrors live A&C conditions from `NativeAlarmActor.cs:504` via batched upserts, and is by a wide margin the highest-volume table in either database. `sf_messages` worst case is ~50 row-writes/sec (`DefaultMaxRetries = 50`, 10 s sweep, 4-way target parallelism). Neither number can be turned into an oplog cap from first principles. **Task 1 is a rig soak that measures both.** Its output sets `MaxOplogRows` / `MaxOplogAge` in Task 19. If the soak shows the shared oplog cannot absorb alarm-storm churn, the escape hatch is keyed instances — named in the **adoption** design doc (`~/Desktop/scadaproj/docs/plans/2026-07-19-scadabridge-localdb-design.md:139-141`, "Out of scope"; not the 07-17 library design doc) — that is a library-level effort and would suspend this plan, so **Task 1 gates everything after Task 2.** ### D5. Cutover forecloses rolling site upgrades `SiteReplicationActor` retains a legacy monolithic `SfBufferSnapshot` handler specifically for rolling upgrades. With no dual-mechanism period, a site pair cannot be upgraded one node at a time — one node would speak a protocol the other no longer implements. **Both nodes of a site must be stopped and started together.** Task 21 puts this in the deployment runbook. ### D6. The 4 MB gRPC message cap replaces the 128 KB Akka frame as the size ceiling The Akka frame limit that motivated notify-and-fetch is gone, but LocalDb sync has its own ceiling the plan must respect: **neither side configures gRPC message sizes** — ScadaBridge's `AddGrpc` (`Program.cs:519-521`) sets only the auth interceptor, and the library's initiator channel is a bare `GrpcChannel.ForAddress` — so the **4 MB default receive limit** applies in both directions. Batching is **row-count-only** (`MaxBatchSize`, default 500 — `SyncSession.cs:227`, `SnapshotStreamer.cs:55`; there is no byte-aware chunking). `deployed_configurations.config_json` is documented to exceed 128 KB per row; a few dozen such rows in one batch is over 4 MB, which fails the stream and wedges replication on a poison batch. **Task 1 measures the real max/typical `config_json` size** on the rig's legacy DB; **Task 19 sets `LocalDb:Replication:MaxBatchSize`** so `max-row-bytes × MaxBatchSize` stays comfortably under 4 MB. If any single row approaches 4 MB, **stop**: that needs a LocalDb library change (byte-aware batching or configurable message sizes) — a `scadaproj` effort, same class as the keyed-instances hatch. --- ## Execution waves | Wave | Tasks | Note | |---|---|---| | 0 | 1, 2 | Soak + decision record. **Task 1 gates the rest.** | | 1 | 3, 4 | Schema extraction — independent files, dispatch in parallel | | 2 | 5, 6 | Store rewiring — independent files, dispatch in parallel | | 3 | 7, 8, 9 | Setup + migrators | | 4 | 10, 11 | Port test intents as specs — **before** any deletion | | 5 | 12, 13 | Pin the active-node purge; notify-and-fetch scope check | | 6 | 14, 15, 16, 17 | The cutover, serial | | 7 | 18, 19 | Convergence suite + rig config | | 8 | 20, 21 | Live gate + docs | --- ### Task 1: Rig soak — measure oplog growth under real write rates **Classification:** high-risk **Estimated implement time:** ~40 min wall-clock (mostly waiting) **Parallelizable with:** none — this gates the plan This is the gate item from `2026-07-19-localdb-phase2-gate.md` §5. It is measurement, not code. **Files:** - Create: `docs/plans/2026-07-19-localdb-phase2-soak.md` > **Method note (corrected):** on this rig the Phase 2 tables are still in the **legacy** DBs — > `sf_messages` in `store-and-forward.db`, `native_alarm_state` in `scadabridge.db` — and are > **not registered** with LocalDb, so driving S&F/alarm churn moves `__localdb_oplog` **not at > all**. This soak therefore measures (a) **source write rates from the legacy DBs** and derives > the oplog sizing arithmetically, and (b) that the Phase 1 oplog stays healthy alongside. The > *empirical* post-cutover drain check lives in Task 20 (evidence item 10). > > Also: the containers have **no `sqlite3` binary** (bare `aspnet:10.0` image), but the data dirs > are **host bind mounts** — run `sqlite3` on the host against `docker/site-a-node-*/data/`. > Metrics are on port **8084** inside the container (not published to the host). **Step 1: Bring up the rig with Phase 1 replication enabled** ```bash cd ~/Desktop/ScadaBridge # already on feat/localdb-phase2 (= phase1 + docs) bash docker/deploy.sh ``` Site-a's pair replicates; site-b/site-c deliberately do not (the default-OFF pin). Confirm: ```bash docker exec scadabridge-site-a-a curl -s localhost:8084/metrics | grep '^localdb_' ``` Expected: `localdb_*` series present (`localdb_oplog_depth`, `localdb_sync_*` — the meter is `ZB.MOM.WW.LocalDb.Replication`). If absent, the meter allowlist regressed — see `SiteServiceRegistration.ObservedMeters`. **Step 2: Capture baselines (host-side, against the bind mounts)** ```bash cd ~/Desktop/ScadaBridge/docker sqlite3 site-a-node-a/data/site-localdb.db "SELECT COUNT(*) FROM __localdb_oplog;" sqlite3 site-a-node-a/data/store-and-forward.db \ "SELECT COUNT(*), COALESCE(SUM(retry_count),0) FROM sf_messages;" sqlite3 site-a-node-a/data/scadabridge.db "SELECT COUNT(*) FROM native_alarm_state;" ``` Record all values and the wall-clock time. (`COUNT + SUM(retry_count)` is the S&F row-write proxy: every failed attempt is one `UPDATE` that increments `retry_count`.) **Step 3: Drive sustained load for 30 minutes** Two generators, run concurrently: - **Store-and-forward churn.** Point a store-and-forward target at an address that refuses connections, then enqueue messages. Each failed attempt costs one `UPDATE`. With the default 10 s sweep this is ~1 write/message/10 s per lane. - **Alarm churn.** Drive A&C conditions on a deployed instance so `native_alarm_state` upserts fire. This is the number that actually matters — it is unbounded by design. **Step 4: Sample write rates every 5 minutes** ```bash cd ~/Desktop/ScadaBridge/docker for i in $(seq 1 6); do echo "=== t+$((i*5))m ===" sqlite3 site-a-node-a/data/store-and-forward.db \ "SELECT COUNT(*), COALESCE(SUM(retry_count),0) FROM sf_messages;" sqlite3 site-a-node-a/data/scadabridge.db \ "SELECT COUNT(*), COUNT(CASE WHEN last_transition_at > datetime('now','-5 minutes') THEN 1 END) FROM native_alarm_state;" sqlite3 site-a-node-a/data/site-localdb.db "SELECT COUNT(*) FROM __localdb_oplog;" docker exec scadabridge-site-a-a curl -s localhost:8084/metrics \ | grep -E '^localdb_(oplog_depth|sync)' sleep 300 done ``` Per-interval deltas of `COUNT + SUM(retry_count)` give S&F row-writes/sec; the `last_transition_at`-window count approximates alarm upserts/sec (it undercounts refresh-only upserts — say so in the findings if alarm volume is near a threshold). **Step 5: Measure `config_json` sizes (D6)** ```bash sqlite3 site-a-node-a/data/scadabridge.db \ "SELECT COUNT(*), MAX(LENGTH(config_json)), AVG(LENGTH(config_json)) FROM deployed_configurations;" ``` If the rig's configs are trivially small, also run this against a production-representative legacy DB (wonder) before trusting the number. **Step 6: Record findings** Write `docs/plans/2026-07-19-localdb-phase2-soak.md` with: - rows/sec observed for `sf_messages` and for `native_alarm_state`, separately - max/avg `config_json` bytes and the derived **`MaxBatchSize`** (`max-row-bytes × MaxBatchSize` comfortably under 4 MB — D6) - Phase 1 oplog depth over the run (should stay near zero — it is not driven by this load) - dead-letter count (must be 0) - **the recommended `MaxOplogRows` and `MaxOplogAge`**, with the arithmetic (sustained rows/sec × retention window, with alarm-storm headroom) **Step 7: The decision gate** If the measured sustained write rate × any reasonable `MaxOplogAge` exceeds any reasonable `MaxOplogRows` — i.e. the shared oplog cannot absorb this write profile even on paper — **stop. Do not proceed to Task 3.** Report to the user: the keyed-instances escape hatch (`~/Desktop/scadaproj/docs/plans/2026-07-19-scadabridge-localdb-design.md:139-141`) is required first, and that is a library effort in `scadaproj`. Likewise stop if any single `config_json` approaches 4 MB (D6 — needs byte-aware batching in the library). The empirical drain-under-churn confirmation happens post-cutover in Task 20; monotonic growth there is the same stop condition. **Step 8: Commit** ```bash git add docs/plans/2026-07-19-localdb-phase2-soak.md git commit -m "docs(localdb): phase 2 rig soak findings — oplog sizing evidence" ``` (The `feat/localdb-phase2` branch already exists — it carries this plan — so commit to it; do **not** `git checkout -b`.) --- ### Task 2: Decision record **Classification:** trivial **Estimated implement time:** ~3 min **Parallelizable with:** none (needs Task 1's numbers) **Files:** - Modify: `docs/plans/2026-07-19-localdb-phase2-gate.md` Update the gate's status header from `NOT STARTED` to point at this plan, and answer each §5 question inline with a one-paragraph resolution referencing D1–D5 above plus Task 1's measured numbers. Do not delete the questions — the reasoning is the value. **Commit:** `docs(localdb): close the phase 2 gate — questions answered, plan written` --- ### Task 3: Extract `StoreAndForwardSchema.Apply` **Classification:** small **Estimated implement time:** ~4 min **Parallelizable with:** Task 4 Mirrors `OperationTrackingSchema` exactly. Read that file first — it is the reference. **Files:** - Create: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardSchema.cs` - Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:32-117` - Test: `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardSchemaTests.cs` **Step 1: Write the failing test** ```csharp [Fact] public void Apply_IsIdempotent_AndCreatesEveryColumn() { var path = Path.Combine(Path.GetTempPath(), $"sf-schema-{Guid.NewGuid():N}.db"); try { using var connection = new SqliteConnection($"Data Source={path}"); connection.Open(); StoreAndForwardSchema.Apply(connection); StoreAndForwardSchema.Apply(connection); // second run must not throw using var cmd = connection.CreateCommand(); cmd.CommandText = "SELECT name FROM pragma_table_info('sf_messages')"; var columns = new List(); using var reader = cmd.ExecuteReader(); while (reader.Read()) columns.Add(reader.GetString(0)); // 12 original + 4 additive Assert.Equal(16, columns.Count); Assert.Contains("last_attempt_at_ms", columns); Assert.Contains("parent_execution_id", columns); } finally { SqliteConnection.ClearAllPools(); File.Delete(path); } } ``` **Step 2: Run it — expect FAIL** (`StoreAndForwardSchema` does not exist) ```bash dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj --filter "FullyQualifiedName~StoreAndForwardSchemaTests" ``` **Step 3: Create the schema class** Move, verbatim: the DDL literal from `StoreAndForwardStorage.cs:50-68`, the four `AddColumnIfMissingAsync` calls (`:79-92`) and their helper (`:125-142`), the `last_attempt_at_ms` backfill (`:98-105`), and the due-index (`:109-114`). Convert every `*Async` to sync (`ExecuteNonQuery`, `ExecuteScalar`) — `OnReady` is a synchronous `Action`. Depend only on `Microsoft.Data.Sqlite`, never on the LocalDb library. Carry this comment onto the class, because it is the whole reason the file exists: ```csharp /// /// Deliberately depends only on Microsoft.Data.Sqlite, not on the LocalDb library. /// The Host applies this DDL to a LocalDb-managed connection before /// RegisterReplicated installs the capture triggers; nothing about the schema /// itself is LocalDb-specific, and the store still calls it so a directly-constructed /// store (tests, tooling) remains self-sufficient. /// ``` **Step 4: Reduce `StoreAndForwardStorage.InitializeAsync` to a call** ```csharp public async Task InitializeAsync() { EnsureDatabaseDirectoryExists(); await using var connection = await OpenConnectionAsync(); StoreAndForwardSchema.Apply(connection); } ``` **Step 5: Run the S&F suite — expect PASS** ```bash dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj ``` **Step 6: Commit** ```bash git add -A && git commit -m "refactor(sf): extract StoreAndForwardSchema.Apply from the storage class" ``` --- ### Task 4: Extract `SiteStorageSchema.Apply` **Classification:** small **Estimated implement time:** ~5 min **Parallelizable with:** Task 3 **Files:** - Create: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageSchema.cs` - Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:58-196` - Test: `tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Persistence/SiteStorageSchemaTests.cs` Same shape as Task 3. Move the DDL literal (`:79-158`, all 9 tables), `MigrateSchemaAsync` (`:167-181`) and `TryAddColumnAsync` (`:183-196`), converting to sync. **Do not move** the `PRAGMA journal_mode=WAL` at `:68-76`. LocalDb owns the connection's pragmas. **Test assertion:** all 9 tables exist after `Apply`, and a second `Apply` does not throw. Name every table explicitly — a loop that asserts "9 tables" would pass if one were renamed: ```csharp foreach (var table in new[] { "deployed_configurations", "static_attribute_overrides", "shared_scripts", "external_systems", "database_connections", "notification_lists", "data_connection_definitions", "smtp_configurations", "native_alarm_state", }) { Assert.True(TableExists(connection, table), $"missing table: {table}"); } ``` **Commit:** `refactor(site): extract SiteStorageSchema.Apply from SiteStorageService` --- ### Task 5: Rewire `StoreAndForwardStorage` onto `ILocalDb` **Classification:** high-risk **Estimated implement time:** ~5 min **Parallelizable with:** Task 6 Follow `OperationTrackingStore` (Phase 1, commit `0b5e9b44`) as the reference — it solved this exact problem. **Files:** - Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:22-26,177-185` - Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ServiceCollectionExtensions.cs:18-25` - Test: `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/` (fixtures) **Step 1: Replace the constructor** ```csharp public StoreAndForwardStorage(ILocalDb localDb, ILogger logger) { ArgumentNullException.ThrowIfNull(localDb); ArgumentNullException.ThrowIfNull(logger); _localDb = localDb; _logger = logger; } ``` **Step 2: Replace `OpenConnectionAsync`** ```csharp // CreateConnection returns an ALREADY-OPEN, pragma-configured connection with the // zb_hlc_next() UDF registered. Calling OpenAsync on it throws, and a raw // SqliteConnection would lack the UDF, making every capture trigger fail closed. private SqliteConnection OpenConnection() => _localDb.CreateConnection(); ``` Update all 18 call sites (`:36, 238, 258, 286, 321, 378, 412, 449, 482, 503, 548, 572, 589, 619, 638, 659, 677, 694`) from `await using var connection = await OpenConnectionAsync();` to `await using var connection = OpenConnection();`. **Step 3: Delete `EnsureDatabaseDirectoryExists` (`:152-168`)** — LocalDb owns the file. **Step 4: Update the DI registration** ```csharp services.AddSingleton(sp => new StoreAndForwardStorage( sp.GetRequiredService(), sp.GetRequiredService>())); ``` **Step 5: Fix the test fixtures** Every test constructing `StoreAndForwardStorage` with a connection string now needs an `ILocalDb`. Phase 1 created `tests/ZB.MOM.WW.ScadaBridge.SiteEventLogging.Tests/TestLocalDb.cs` for exactly this — read it and add an equivalent (or reference it) in the S&F test project. It must use a real temp **file** (never `:memory:` — capture triggers need a durable file) and call `SqliteConnection.ClearAllPools()` before deleting, including the `-wal` and `-shm` sidecars. **Step 6: Full S&F suite — expect PASS** ```bash dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj ``` **Commit:** `refactor(sf): StoreAndForwardStorage takes ILocalDb instead of a connection string` --- ### Task 6: Rewire `SiteStorageService` onto `ILocalDb` **Classification:** high-risk **Estimated implement time:** ~5 min **Parallelizable with:** Task 5 **Files:** - Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:13-51` - Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/ServiceCollectionExtensions.cs:36-42` - Modify: `src/ZB.MOM.WW.ScadaBridge.Host/SiteServiceRegistration.cs:70-71` Same shape as Task 5, but note the differences: - There are **21** inline `new SqliteConnection(_connectionString)` + `OpenAsync()` pairs, not a helper (lines 60, 206, 248, 311, 345, 386, 414, 443, 470, 498, 538, 583, 608, 640, 663, 692, 730, 769, 813, 837, 868 — counted before Task 4, which removes the :60 one inside `InitializeAsync`). Add a `private SqliteConnection OpenConnection() => _localDb.CreateConnection();` and convert all of them. - `public SqliteConnection CreateConnection()` at `:51` is a public escape hatch used by `SiteExternalSystemRepository` (the only repository consumer). Keep the member; change the body to `=> _localDb.CreateConnection();`. - The busy-timeout floor (`BusyTimeoutFloorSeconds`, `:36`) and its `SqliteConnectionStringBuilder` normalization are **deleted** — LocalDb configures pragmas on every connection it hands out. - `AddSiteRuntime(string siteDbConnectionString)` loses its parameter. Keep the no-arg overload at `:23-28` (it already exists) and delete the string one, updating `SiteServiceRegistration.cs:70-71` to `services.AddSiteRuntime();`. **Watch for:** transactions at `:348` (`RemoveDeployedConfigAsync`) and `:540` (`UpsertNativeAlarmsAsync`) must keep using `connection.BeginTransaction()` on the LocalDb connection, not `ILocalDb.BeginTransactionAsync()` — mixing the two would put the cascade's three deletes in different transactions. **Step: Run both affected suites** ```bash dotnet test tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests.csproj dotnet test tests/ZB.MOM.WW.ScadaBridge.Host.Tests/ZB.MOM.WW.ScadaBridge.Host.Tests.csproj ``` Host.Tests will surface DI fixtures that now need `LocalDb:Path` — Phase 1 hit the identical failure (tasks.json Task 4 note). Add the config key to each failing fixture. **Commit:** `refactor(site): SiteStorageService takes ILocalDb instead of a connection string` --- ### Task 7: Extend `SiteLocalDbSetup` with the new DDL (not yet registered) **Classification:** standard **Estimated implement time:** ~3 min **Parallelizable with:** none Deliberately applies DDL **without** calling `RegisterReplicated`. The tables live in the consolidated file but are not yet captured, so the bespoke replicator keeps working unchanged and the tree stays green. Task 14 flips them on and deletes the bespoke path in one commit. **Files:** - Modify: `src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs:43-57` - Test: `tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbWiringTests.cs` **Step 1: Extend `OnReady`** ```csharp using (var connection = db.CreateConnection()) { OperationTrackingSchema.Apply(connection); SiteEventLogSchema.Apply(connection); SiteStorageSchema.Apply(connection); StoreAndForwardSchema.Apply(connection); } db.RegisterReplicated("OperationTracking"); db.RegisterReplicated("site_events"); // Phase 2 tables are created here but NOT yet registered — the bespoke replicator // still owns them until the Task 14 cutover. Registering them now would double- // replicate every row (harmless, both paths upsert, but it would mask a defect in // either one). SiteLocalDbLegacyMigrator.Migrate(db, config); ``` **Step 2: Pin it** Add a test asserting all 4 Phase-1-and-2 table sets exist after `OnReady`, and that `db.ReplicatedTables` contains **exactly** `OperationTracking` and `site_events` — the "not yet" is the assertion that matters. **Commit:** `feat(localdb): create config + sf_messages tables in the consolidated DB` --- ### Task 8: Extend the legacy migrator for `sf_messages` **Classification:** high-risk **Estimated implement time:** ~5 min **Parallelizable with:** Task 9 Read `SiteLocalDbLegacyMigrator.cs` in full first. `MigrateTracking` (`:116-161`) is the closest reference — like `sf_messages`, it has a native TEXT PK and needs no id synthesis. **Files:** - Modify: `src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbLegacyMigrator.cs:57-66` - Test: `tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbLegacyMigratorTests.cs` **Step 1: Write the failing tests** Three, minimum: ```csharp [Fact] public void SfMessages_AreCopiedFromTheLegacyFile() { } [Fact] public void SfMessages_Migration_IsIdempotent_WhenRerunAfterACrashBeforeRename() { } [Fact] public void MigratedSfMessages_EnterTheOplog_SoTheyActuallyReplicate() { } ``` The third is the one that catches the ordering defect. Phase 1's equivalent (`SiteLocalDbLegacyMigratorTests.cs:244`, `MigratedRows_EnterTheOplog_SoTheyActuallyReplicate`) asserts on `__localdb_oplog` directly — copy that assertion. **The oplog table is `__localdb_oplog`** (`LocalDbSchema.cs:19`), not `zb_oplog`. > **Harness note:** production `OnReady` does not register `sf_messages` until the Task 14 > cutover, so these tests must build their own `ILocalDb` (`TestLocalDb`), call > `RegisterReplicated("sf_messages")` themselves, and then run the migrator — the oplog > assertion needs live capture triggers. Correspondingly, Task 14 must keep > `SiteLocalDbLegacyMigrator.Migrate` as the **last** call in `OnReady`, after all > registrations. **Step 2: Add `ResolveStoreAndForwardPath` + `MigrateStoreAndForward`** ```csharp private const string DefaultStoreAndForwardPath = "./data/store-and-forward.db"; internal static string ResolveStoreAndForwardPath(IConfiguration config) => config["ScadaBridge:StoreAndForward:SqliteDbPath"] ?? DefaultStoreAndForwardPath; ``` Migrate all 16 columns with `INSERT OR IGNORE`. The PK is native TEXT — **no id synthesis**, unlike `MigrateEvents`. **Step 3: Wire into `Migrate`** — add the call alongside the existing two. **Commit:** `feat(localdb): migrate legacy store-and-forward.db into the consolidated DB` --- ### Task 9: Extend the legacy migrator for the 9 config tables **Classification:** high-risk **Estimated implement time:** ~5 min **Parallelizable with:** Task 8 **Files:** - Modify: `src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbLegacyMigrator.cs` - Test: `tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbLegacyMigratorTests.cs` All 9 tables migrate from one legacy file (`ScadaBridge:Database:SiteDbPath`), so this is one `MigrateSiteStorage` that opens the legacy DB once and copies table by table inside a single transaction, then renames once. **Two tables are deliberately skipped:** `notification_lists` and `smtp_configurations`. They are purged on every deploy and are permanently empty by design (site write paths were removed 2026-07-10). Migrating them would resurrect plaintext SMTP passwords from a pre-fix legacy file into a **replicated** table. Add this as a comment and as a test: ```csharp [Fact] public void Migration_DoesNotCopyNotificationOrSmtpRows_EvenWhenTheLegacyFileHasThem() { // These carry plaintext SMTP passwords in pre-2026-07-10 files. They are purged // on every deploy and must never enter a replicated table. } ``` **Step: Verify all 9 tables' columns match** between `SiteStorageSchema` and the migrator's `INSERT` lists. A column-count mismatch here fails at runtime on a real deployment and nowhere else. **Commit:** `feat(localdb): migrate legacy scadabridge.db config tables into the consolidated DB` --- ### Task 10: Port the S&F replication test intents as CDC specs **Classification:** standard **Estimated implement time:** ~5 min **Parallelizable with:** Task 11 **Specifications first, deletion second.** These tests encode behaviour the replacement must still satisfy. The gate document named two files; the real specification is five. **Files:** - Read (as spec): `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationServiceTests.cs` - Read (as spec): `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardReplicationTests.cs` - Create: `tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbStoreAndForwardConvergenceTests.cs` Port these **intents** (not the mechanics) as two-node convergence assertions, using Phase 1's `LocalDbSitePairConvergenceTests.cs` as the harness reference: | Intent | Assertion under CDC | |---|---| | Add materialises on the peer | enqueue on A → row present on B | | Remove deletes on the peer | deliver on A → row absent on B | | Park sets status Parked on the peer | park on A → B's row has status Parked | | Requeue resets status + RetryCount | requeue on A → B shows Pending, RetryCount 0 | | Add-then-Remove never converges to present | **the portable form of the strict-ordering test** | | Apply Add twice is idempotent, newest wins | LWW gives this | | Park when Add was lost still materialises | upsert semantics | **Explicitly do not port:** `ReplicationServiceTests.ReplicationOperations_AreDispatchedInIssueOrder` (`:169-195`; 200 interleaved ops dispatched synchronously in strict issue order, 400 observed). That asserts the *mechanism* — inline fire-and-forget dispatch — not the outcome. CDC capture is asynchronous and batched by construction. The portable intent is row 5 above. Record this in the test file as a comment so a future reader does not think it was dropped by accident. **Commit:** `test(localdb): port store-and-forward replication intents as CDC convergence specs` --- ### Task 11: Port the resync + directional-authority tests **Classification:** standard **Estimated implement time:** ~5 min **Parallelizable with:** Task 10 **Files:** - Read (as spec): `tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/SiteReplicationActorTests.cs` - Read (as spec): `tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/Cluster/SfBufferResyncPredicateTests.cs` - Create: `tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbConfigConvergenceTests.cs` **`SfBufferResyncPredicateTests` is the most valuable test in the set** — the N1 Critical regression. Per D2 it must be re-expressed, not ported literally: ```csharp [Fact] public async Task ANodeWithNewerLocalRows_KeepsThem_WhenAPeerSnapshotArrives() { // N1 Critical (was: SfBufferResyncPredicateTests). The bespoke replicator needed a // directional guard because ReplaceAllAsync was a destructive DELETE-then-INSERT, // so a wrong-direction resync wiped a live buffer. // // LocalDb's snapshot resync merges per row under LWW and never deletes // (SnapshotApplier has no DELETE; LwwApplier.cs:69-78 discards an incoming row // whose HLC is lower than the local row's). The failure mode is structurally // impossible rather than guarded against — so this asserts convergence, NOT // directional authority. There is no active/standby asymmetry left to enforce. } ``` **Also do not port:** actor tests #1–#5 (notify-and-fetch config fetch/retry/supersede). Per D1 the fetch path is deleted, so these describe code that will not exist. Replace with a single assertion that a config deploy on A converges to B **without** B making any HTTP call — assert on a fetcher test double that records zero invocations. That is the positive proof that notify-and-fetch is gone. > **Scope the zero-fetch assertion to the deploy-replication flow.** `SiteReconciliationActor` > survives Phase 2 and legitimately HTTP-fetches at node startup when central reports gaps (D1). > The two-node harness has no central, so its pass no-ops there — but do not write the assertion > as "the standby never fetches, ever"; assert the fetcher double records zero calls **during > deploy convergence**. **Commit:** `test(localdb): port resync + config replication intents as CDC convergence specs` --- ### Task 12: Pin the active-node SMTP purge **Classification:** standard **Estimated implement time:** ~3 min **Parallelizable with:** none Per D3 (corrected): the active node's artifact-apply **already** calls `PurgeCentralOnlyNotificationConfigAsync` — `DeploymentManagerActor.cs:1921`, inside `HandleDeployArtifacts` (`:1864-1963`). Nothing is re-homed. What is missing is a **pin**: Task 16 edits this actor's wiring, and no test today fails if the purge call is dropped (`ArtifactStorageTests` covers the storage method, not the actor's call site). **Files:** - Test: `tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/DeploymentManagerActorTests.cs` **Step 1: Write the pin test (red-first by commenting out the call locally, then restore)** ```csharp [Fact] public async Task ApplyingArtifacts_PurgesCentralOnlyNotificationConfig() { // Security cleanup: notification_lists and smtp_configurations can hold plaintext // SMTP passwords from pre-2026-07-10 deployments. The ACTIVE node's artifact-apply // purges them (DeploymentManagerActor.HandleDeployArtifacts); the standby's copy of // this call dies with SiteReplicationActor in Task 15. This test pins the active-node // call so Task 16's actor edits cannot silently drop it. // Arrange: real temp-file SiteStorageService, seed one row into each table, // drive HandleDeployArtifacts; assert both tables are empty afterwards. } ``` **Step 2: No production change is expected.** If recon of the call site shows otherwise, stop and report. **Commit:** `test(site): pin the active-node notification-config purge` --- ### Task 13: Notify-and-fetch scope check — the guarded write stays **Classification:** standard **Estimated implement time:** ~3 min **Parallelizable with:** none Per D1 (corrected). The original task deleted `StoreDeployedConfigIfNewerAsync` here — that was wrong twice over: (a) it has a **second surviving caller**, and (b) deleting it before Task 15 would not even compile, since `SiteReplicationActor` still calls it until then. **Files:** - Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:301-336` (doc comment only) **Step 1: Keep `StoreDeployedConfigIfNewerAsync`.** Its callers are `SiteReplicationActor.HandleApplyConfigDeploy` (`SiteReplicationActor.cs:375` — dies with the actor in Task 15) **and `SiteReconciliationActor` (`SiteReconciliationActor.cs:166`)** — the per-node startup self-heal against central, which survives Phase 2 and still needs the `deployed_at` guard against a stale reconcile fetch. Update the method's doc comment to name reconciliation as the remaining caller (and drop any mention of standby replication). The unguarded `StoreDeployedConfigAsync` (`:241-271`) stays the active deploy path. **Step 2: `ConfigFetchRetryCount` — answered, but its removal moves to Task 17.** The grep was run during plan review: the only production reader is `SiteReplicationActor.cs:157` (plus its validator rule and tests). `SiteReconciliationActor` does a single best-effort pass and does not read it; `IDeploymentConfigFetcher` itself is kept (used by `DeploymentManagerActor`'s refresh path, `SiteReconciliationActor`, and registered at `ServiceCollectionExtensions.cs:84`). So the option (`SiteRuntimeOptions.cs:68`) and its validator rule (`SiteRuntimeOptionsValidator.cs:56-58`) become dead **after Task 15 deletes the actor** — remove them in Task 17 (config-key cleanup), not here, or the build breaks. Re-run `grep -rn "ConfigFetchRetryCount" src/ tests/` there to confirm. **Commit:** `docs(site): StoreDeployedConfigIfNewerAsync serves reconciliation — guard stays under CDC` --- ### Task 14: Register the Phase 2 tables and delete `ReplicationService` **Classification:** high-risk **Estimated implement time:** ~5 min **Parallelizable with:** none **The cutover.** Registration and bespoke-path deletion land in one commit so the two mechanisms never both run. **Files:** - Modify: `src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs` - Delete: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ReplicationService.cs` - Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardService.cs` — field `:39`, ctor param `:243`, and the 6 emission call sites `:654, 805, 836, 872, 1120, 1146` - Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ServiceCollectionExtensions.cs:63-68` (the `ReplicationService` registration) **and `:32`** (`GetRequiredService()` inside the `StoreAndForwardService` factory at `:27-61`) - Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:284-306` - Delete: `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationServiceTests.cs` - Delete: `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardReplicationTests.cs` - Delete: `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationWireSerializationPinTests.cs` **Step 1: Register 8 tables** in `SiteLocalDbSetup.OnReady`, replacing the Task 7 comment: the 7 config tables (`deployed_configurations`, `static_attribute_overrides`, `shared_scripts`, `external_systems`, `database_connections`, `data_connection_definitions`, `native_alarm_state`) plus `sf_messages`. **Do NOT register `notification_lists` or `smtp_configurations`** (this reverses the plan's original instruction): they are permanently empty by design — no site writer since 2026-07-10, the migrator skips them (Task 9), and the active-node purge (`DeploymentManagerActor.cs:1921`) keeps them empty — so registering them would create a standing replication channel whose only historical payload was plaintext SMTP passwords, for zero benefit. Put that rationale in an `OnReady` comment so a future writer has to make an explicit decision. Keep `SiteLocalDbLegacyMigrator.Migrate` as the **last** call in `OnReady`, after all registrations, so migrated rows enter the oplog through live capture triggers (the Phase 1 ordering that is silently load-bearing). Note the two composite-PK tables are fine: LocalDb's `RegisterReplicated` supports multi-column PKs (`SqliteLocalDb.RegisterReplicated` orders `pk` ordinals), and no Phase 2 table has a BLOB column (which would be rejected). **Step 2: Delete `ReplicationService`** and remove the ctor parameter (`:243`), the `_replication` field (`:39`), and all 6 emission call sites from `StoreAndForwardService`. **Step 3: Delete `ReplaceAllAsync`** (`:284-306`). Per D2 the library's resync merges; a destructive delete-all must not survive into a replicated table, where CDC would capture the mass-delete and ship it to the peer. **Step 4: Check `GetAllMessagesAsync`'s `Truncated` flag** — it exists only for the resync path. If nothing else reads it, simplify the return type. **Step 5: `CompositionRootTests.cs:474`** asserts `ReplicationService` is a registered site singleton. Delete that assertion. **Step 6: Full S&F + Host suites — expect PASS** **Commit:** `feat(localdb)!: replicate sf_messages via CDC, delete ReplicationService` --- ### Task 15: Delete `SiteReplicationActor` and its messages **Classification:** high-risk **Estimated implement time:** ~4 min **Parallelizable with:** none **Files:** - Delete: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Actors/SiteReplicationActor.cs` - Delete: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Messages/ReplicationMessages.cs` - Delete: `tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/SiteReplicationActorTests.cs` - Delete: `tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/ResyncWireSerializationPinTests.cs` - Delete: `tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/Cluster/SfBufferResyncPredicateTests.cs` The whole `ReplicationMessages.cs` file becomes dead — it contains exactly the 10 records `ReplicateConfigDeploy/Remove/SetEnabled`, `ReplicateArtifacts`, `ReplicateStoreAndForward`, `ApplyConfigDeploy/Remove/SetEnabled`, `ApplyArtifacts`, `ApplyStoreAndForward`, whose only consumers are `SiteReplicationActor`, the 5 `DeploymentManagerActor` Tell sites (Task 16), the `AkkaHostedService` wiring (Task 16), and the deleted tests. The four resync records (`RequestSfBufferResync`, `SfBufferSnapshot`, `SfBufferSnapshotChunk`, `SfBufferResyncAck`) are **not** in this file — they are declared at the bottom of `SiteReplicationActor.cs` (`:678-707`) and die with the actor file. **Do not delete** `ActiveNodeEvaluator` (`Communication/ClusterState/ActiveNodeEvaluator.cs`) — the S&F **delivery gate** still needs it (`AkkaHostedService.cs:866` → `SetDeliveryGate` → `SelfIsPrimary` → `ActiveNodeEvaluator.SelfIsOldestUp`). Its doc comment at `:14` mentions `SiteReplicationActor` and `:16` mentions the resync authority checks; the actor's own direct call at `SiteReplicationActor.cs:288` goes away with the file. Update the comment, keep the type. Confirm the deletions are complete before committing: ```bash grep -rn "SiteReplicationActor\|ReplicateStoreAndForward\|SfBufferSnapshot" src/ tests/ ``` Expected: no matches. **Commit:** `feat(localdb)!: delete SiteReplicationActor — CDC replaces hand-shipped ops` --- ### Task 16: Clean up `DeploymentManagerActor` and `AkkaHostedService` **Classification:** standard **Estimated implement time:** ~4 min **Parallelizable with:** none **Files:** - Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Actors/DeploymentManagerActor.cs:55,161-190,794,970,1004,1102,1952` - Modify: `src/ZB.MOM.WW.ScadaBridge.Host/Actors/AkkaHostedService.cs:770-796,806-812,1056` **Step 1:** Remove the `IActorRef? replicationActor = null` ctor parameter (**`:169`** — the ctor signature spans `:161-175`; `:184` is the field *assignment*), the `_replicationActor` field (`:55`), and all 5 `_replicationActor?.Tell(...)` sites (`:794, 970, 1004, 1102, 1952`). The parameter is optional and positional — removing it **silently shifts every argument after it** (`healthCollector`, `serviceProvider`, `loggerFactory`, `configFetcher`, `startupLoadRetryInterval`, `configLoader`). The sharpest hazard is the **same-typed `IActorRef?` optional immediately before it** — `dclManager` at `:168`: a caller passing positionally could silently feed the wrong actor ref with zero compile errors. Update the `Props.Create` at `AkkaHostedService.cs:806-812` (which passes `replicationActor` positionally at `:810`) and every test constructing this actor — check each call site by hand. **Step 2:** Delete `replicationService` / `replicationLogger` resolution (`:770-773`), the actor creation (`:785-789`), the handler wiring (`:792-796`), and the actor-inventory comment at `:1056`. **Step 3:** `activeNodeCheck` (`:781-783`) is still used by `SiteCommunicationActor` (`:816-821`). Keep it. **Step 4:** Full Host + SiteRuntime + IntegrationTests suites. **Commit:** `refactor(site): drop the replication actor from the deployment + host wiring` --- ### Task 17: Config-key cleanup **Classification:** standard **Estimated implement time:** ~4 min **Parallelizable with:** none **Files:** - Modify: `src/ZB.MOM.WW.ScadaBridge.Host/StartupValidator.cs:113-115` - Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardOptions.cs:12` - Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardOptionsValidator.cs:19-21` - Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/SiteRuntimeOptions.cs:68` + `SiteRuntimeOptionsValidator.cs:56-58` (moved here from Task 13) - Modify: **10** × `appsettings.Site.json` (Host template + 6 `docker/site-*` + 2 `docker-env2/site-x-*` + **`deploy/wonder-app-vd03/appsettings.Site.json`** — the wonder single-node install, which the original count missed; it sets both paths *and* `ReplicationEnabled: false`) | Key | Action | |---|---| | `ScadaBridge:Database:SiteDbPath` | → migration-only. **Relax the `Require` at `StartupValidator.cs:113`** or every site config keeps it mandatory forever. | | `ScadaBridge:StoreAndForward:SqliteDbPath` | → migration-only. Relax the non-empty rule at `StoreAndForwardOptionsValidator.cs:19-21`. | | `ScadaBridge:StoreAndForward:ReplicationEnabled` | → **fully dead.** Delete the property and every config entry (including wonder's). | | `SiteRuntime … ConfigFetchRetryCount` | → **dead after Task 15** (only reader was `SiteReplicationActor.cs:157`; see Task 13). Delete the option, its validator rule, and any config entries — confirm with `grep -rn "ConfigFetchRetryCount" src/ tests/` first. | `ParkedOperationRelayTests.cs:39` and `ParkedMessageHandlerActorTests.cs:34` set `ReplicationEnabled = false` — they only need the flag to exist, so remove those lines. `StoreAndForwardOptionsTests.cs` touches it three times: `:14` asserts the default, `:26` sets and `:33` asserts a customized value — remove all three references. **Leave the path keys present in configs** with a comment — they are read by the legacy migrator and removing them would strand un-migrated data on a node that has not yet started once. **Commit:** `chore(config): retire ReplicationEnabled, make legacy db paths migration-only` --- ### Task 18: Two-node convergence suite for the Phase 2 tables **Classification:** high-risk **Estimated implement time:** ~5 min **Parallelizable with:** Task 19 Extend Phase 1's `LocalDbSitePairConvergenceTests.cs` harness (real loopback Kestrel h2c + the real `LocalDbSyncAuthInterceptor`, initialized via `SiteLocalDbSetup.OnReady`). **Files:** - Modify: `tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbSitePairConvergenceTests.cs` Scenarios, beyond Tasks 10–11: 1. A config deploy on A converges to B, and B's `deployed_configurations` row matches byte for byte. 2. `RemoveDeployedConfigAsync`'s **3-table cascade** converges — `static_attribute_overrides` and `native_alarm_state` orphans do not survive on B. There are no foreign keys (`SiteStorageService` DDL has none), so this is three independent delete streams that LWW may reorder. **This is the most likely real defect in the whole plan.** 3. A burst of `native_alarm_state` upserts converges and the oplog drains. 4. Node B restarts and catches up via snapshot resync without losing rows written on B while down. **Verify the suite is non-vacuous.** (Corrected claim: Phase 1's convergence suite did **not** do a mismatched-key run — it uses one shared key, `LocalDbSitePairConvergenceTests.cs:47`; wrong-key *denial* is unit-covered separately in `Host.Tests/LocalDbSyncAuthInterceptorTests.cs`.) Prove it directly for the new scenarios: run once with the new table registrations disabled on node B (or a deliberately wrong key) and confirm every new scenario goes red, then restore and record the result in the commit message. A convergence suite that cannot fail proves nothing. **Commit:** `test(localdb): two-node convergence for config tables + sf_messages` --- ### Task 19: Rig configuration **Classification:** small **Estimated implement time:** ~3 min **Parallelizable with:** Task 18 **Files:** - Modify: `docker/site-a-node-a/appsettings.Site.json`, `docker/site-a-node-b/appsettings.Site.json` Set `LocalDb:Replication:MaxOplogRows` and `LocalDb:Replication:MaxOplogAge` (exact key paths — the lib binds the `LocalDb:Replication` section; defaults 1,000,000 rows / 7 days) from **Task 1's measured numbers** — not from a guess. Per D6, also set `LocalDb:Replication:MaxBatchSize` (default 500) so `max-config_json-bytes × MaxBatchSize` stays comfortably under the 4 MB gRPC receive limit. Keep site-b/site-c unreplicated so default-OFF stays proven side by side on one rig. **Commit:** `chore(docker): size the site-a oplog caps from the phase 2 soak` --- ### Task 20: Live gate on the docker rig **Classification:** high-risk **Estimated implement time:** ~30 min wall-clock **Parallelizable with:** none **Files:** - Create: `docs/plans/2026-07-19-localdb-phase2-live-gate.md` ```bash cd ~/Desktop/ScadaBridge && bash docker/deploy.sh ``` Tooling: the containers have no `sqlite3` — run DB checks host-side against the bind mounts (`docker/site-a-node-{a,b}/data/`); metrics via `docker exec scadabridge-site-a-a curl -s localhost:8084/metrics` (port **8084**, not published to the host). Evidence required — each must be captured, not asserted: 1. **Migration ran.** `store-and-forward.db.migrated` and `scadabridge.db.migrated` exist on both site-a nodes; row counts in the consolidated DB match the legacy files. (Both nodes migrate their own legacy files independently; identical-content rows then LWW-converge — expected, not an anomaly.) 2. **No SMTP/notification rows migrated** — `SELECT COUNT(*) FROM smtp_configurations` is 0. 3. **Config converges.** Deploy an instance to site-a; both nodes' `deployed_configurations` rows are byte-identical with identical `__localdb_row_version` and originating node id. 4. **The standby made no config HTTP fetch during the deploy** (both nodes up) — grep the deploy window of its log. This is the positive proof that notify-and-fetch is gone. A `SiteReconciliationActor` fetch at node *startup* is legitimate (it survives Phase 2) and does not fail this check. 5. **S&F converges.** Enqueue against a dead target; the parked row appears on the peer. 6. **Site failover.** Flip the site-a pair and confirm buffered messages survive and deliver exactly once after the flip. `docker/failover-drill.sh` targets **CENTRAL** nodes and does not exercise this — run the site flip directly. 7. **Cascade delete.** Remove a deployed config; confirm no orphan `static_attribute_overrides` or `native_alarm_state` rows on the peer. 8. **Zero dead letters**, oplog drained, `localdb_*` scraped from `/metrics`. 9. **Both nodes stopped and started together** (D5) — confirm a clean rejoin. 10. **Post-cutover drain under churn** — the empirical half Task 1 could not measure pre-cutover: re-run Task 1's two load generators against this build and watch `localdb_oplog_depth` (and `SELECT COUNT(*) FROM __localdb_oplog`) rise **and drain between bursts**. Monotonic growth that never drains is the same stop condition as Task 1 step 7 (keyed-instances hatch). **If any check fails, stop and report.** Do not proceed to Task 21. **Commit:** `docs(localdb): phase 2 live gate evidence` --- ### Task 21: Documentation truth pass **Classification:** standard **Estimated implement time:** ~5 min **Parallelizable with:** none **Files:** - Modify: `CLAUDE.md` (ScadaBridge), `~/Desktop/scadaproj/CLAUDE.md` (LocalDb component row) - Modify: `docs/requirements/Component-StoreAndForward.md:83` - Modify: `docs/components/StoreAndForward.md`, `SiteRuntime.md`, `Host.md` - Modify: `docs/known-issues/2026-06-26-deploy-config-exceeds-akka-frame-size.md` - Modify: `docs/deployment/installation-guide.md` + `topology-guide.md` (+ `production-checklist.md` if it covers upgrades) — there is no file literally named "runbook" Specific corrections: - `Component-StoreAndForward.md:83` is a long **normative** paragraph specifying the whole chunked resync protocol and the N5 race. Rewrite for CDC — state the new duplicate-delivery bound explicitly rather than deleting the discussion. - The frame-size known-issue is **already marked "Status: RESOLVED (2026-06-26)"** (resolved by the notify-and-fetch rework). Amend it to record that Phase 2 deleted notify-and-fetch itself — CDC ships the row over gRPC, so the frame constraint is gone entirely — and note the successor ceiling: the 4 MB gRPC message cap managed via `MaxBatchSize` (D6). - **Add to the deployment docs (D5):** a site pair must be stopped and started together. Rolling one node at a time is no longer supported, because the legacy `SfBufferSnapshot` compatibility handler is gone. Also record the D2 caveat: a node offline longer than `LocalDb:Replication:TombstoneRetention` (default 7 days) can resurrect deleted rows on rejoin. - Update the scadaproj LocalDb component row: Phase 2 complete, what replicates now, what `SiteReplicationActor` used to do. **Commit:** `docs(localdb): phase 2 truth pass across both repos` --- ## Definition of done - [ ] Solution builds with **0 warnings** (`TreatWarningsAsErrors` is on) - [ ] Every suite green: Host, CentralUI, Commons, SiteRuntime, AuditLog, Communication, StoreAndForward, HealthMonitoring, IntegrationTests, SiteEventLogging - [ ] `grep -rn "SiteReplicationActor\|ReplicationService\|SfBufferSnapshot" src/` → no matches - [ ] Live gate PASS with all 9 evidence items captured - [ ] Convergence suite verified non-vacuous - [ ] Both `CLAUDE.md` files updated