From f6ca82a9e2b7476b3d7dcd7182d8de363cb39949 Mon Sep 17 00:00:00 2001 From: Joseph Doherty Date: Sun, 19 Jul 2026 15:04:37 -0400 Subject: [PATCH] docs(localdb): phase 2 implementation plan (21 tasks, full scope) Moves scadabridge.db's 9 config tables + sf_messages into the consolidated LocalDb file and deletes SiteReplicationActor + StoreAndForward's ReplicationService. Resolves the phase 2 gate's five open questions from code recon (D1-D5): D1 notify-and-fetch is DELETED, not preserved. It exists only because the config blob exceeds Akka's 128KB frame; LocalDb sync is gRPC. The deployed_at version guard protected against a stale fetch racing, so it dies with the fetch. Do not reproduce it on LWW - different clocks. D2 ReplaceAllAsync is deleted and the N1 directional guard becomes unnecessary: LocalDb's snapshot resync merges per-row LWW and never deletes (SnapshotApplier has no DELETE; LwwApplier.cs:69-78 discards a lower-HLC incoming row). Semantic change - the standby is convergent, no longer byte-identical. D3 The SMTP purge (plaintext passwords) rides the replication path and is re-homed to the active node BEFORE any deletion. D4 native_alarm_state volume is measured by a rig soak, not assumed. Task 1 gates the plan and stops it if the oplog cannot absorb the churn. D5 No dual-mechanism period forecloses rolling site upgrades - both nodes must stop and start together. Recon also found the gate doc's "two test files are the spec" undercounts: the real specification is five files, including the N1 Critical regression test and the only Requeue coverage. Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts --- .../2026-07-19-localdb-adoption-phase2.md | 1019 +++++++++++++++++ ...7-19-localdb-adoption-phase2.md.tasks.json | 89 ++ 2 files changed, 1108 insertions(+) create mode 100644 docs/plans/2026-07-19-localdb-adoption-phase2.md create mode 100644 docs/plans/2026-07-19-localdb-adoption-phase2.md.tasks.json diff --git a/docs/plans/2026-07-19-localdb-adoption-phase2.md b/docs/plans/2026-07-19-localdb-adoption-phase2.md new file mode 100644 index 00000000..a4315d6a --- /dev/null +++ b/docs/plans/2026-07-19-localdb-adoption-phase2.md @@ -0,0 +1,1019 @@ +# ScadaBridge LocalDb Phase 2 Implementation Plan + +> **For Claude:** REQUIRED SUB-SKILL: Use superpowers-extended-cc:executing-plans to implement this plan task-by-task. + +**Goal:** Move `scadabridge.db`'s 9 config tables and StoreAndForward's `sf_messages` into the +consolidated `ZB.MOM.WW.LocalDb` database established by Phase 1, then delete `SiteReplicationActor` +and StoreAndForward's `ReplicationService` outright. + +**Architecture:** Both stores stop owning their own SQLite files and take `ILocalDb` instead, exactly +as `OperationTrackingStore` and `SiteEventLogger` did in Phase 1. Their DDL is extracted into +`*Schema.Apply(SqliteConnection)` helpers that `SiteLocalDbSetup.OnReady` runs before +`RegisterReplicated`. Trigger-based CDC then replicates every row mutation, replacing the bespoke +hand-shipped ops; the library's per-row-LWW snapshot resync replaces the chunked anti-entropy. + +**Tech Stack:** .NET 10, Akka.NET, Microsoft.Data.Sqlite, ZB.MOM.WW.LocalDb 0.1.0, gRPC, xunit. + +**Branch:** `feat/localdb-phase2`, cut from `feat/localdb-phase1`. + +--- + +## Read this before Task 1 + +Phase 1's record is the prerequisite context. Read, in order: + +1. `docs/plans/2026-07-19-localdb-adoption-phase1.md.tasks.json` — the `amendments`, + `prerequisites` and `knownFlakes` arrays. Several Phase 1 assumptions changed during execution. +2. `docs/plans/2026-07-19-localdb-phase2-gate.md` — the gate. §5's open questions are answered in + the Decisions section below; the gate itself is superseded by this plan. +3. `src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs` — the pattern every new table follows. + +### Two known flakes — do not chase these + +- `SiteRuntime.Tests InstanceActorChildAttributeRaceTests` — intermittent `ActorNotFoundException` + under full-suite load. Passes in isolation. Pre-existing. +- `AuditLog.Tests ParentExecutionIdCorrelationTests` — ~91 s and times out on a **cold** MSSQL + fixture, ~1 s warm. Re-run before investigating. + +--- + +## Decisions — the gate's open questions, answered + +These were resolved from code recon before this plan was written. Each names its evidence. **Task 2 +records them formally; the rest of the plan assumes them.** + +### D1. Config moves to CDC, and notify-and-fetch is deleted, not preserved + +`SiteReplicationActor` never sends config JSON over the replication hop — it sends a deployment id +plus fetch coordinates, and the standby HTTP-fetches from central itself +(`SiteReplicationActorTests` #2, #5). That design exists for exactly one reason: the config blob +exceeds Akka's 128 KB frame limit (`docs/known-issues/2026-06-26-deploy-config-exceeds-akka-frame-size.md`). + +Under CDC the row itself replicates over the LocalDb gRPC stream, which has no such limit. The +standby therefore never fetches at all. + +This also disposes of the version guard. `StoreDeployedConfigIfNewerAsync` +(`SiteStorageService.cs:301-336`) guards its upsert with: + +```sql +WHERE excluded.deployed_at > deployed_configurations.deployed_at +``` + +That predicate protects against **a stale fetch landing after a newer one** — a race that only +exists because the standby was independently fetching. With the fetch gone, the only writer is the +active node's deploy path, and HLC ordering on a single writer is monotonic. The guard becomes +dead code. + +> **Do not** try to reproduce the `deployed_at` guard on top of LWW. They order by different +> clocks (central's deploy time vs. the writing node's HLC) and mixing them produces a +> non-convergent store. + +### D2. `ReplaceAllAsync` is deleted, and the N1 directional guard becomes unnecessary + +`StoreAndForwardStorage.ReplaceAllAsync` (`StoreAndForwardStorage.cs:284-306`) is an unqualified +`DELETE FROM sf_messages` followed by a full re-insert. Its own doc warns "Never call on an active +node — it discards every in-flight row." `SfBufferResyncPredicateTests` exists to enforce that only +a standby ever applies it — the N1 Critical regression test. + +The LocalDb library's snapshot resync **merges per row under LWW and never deletes**: + +- `SnapshotApplier.OnBeginAsync` (`SnapshotStreamer.cs:163-170`) resets counters only. There is no + `DELETE` or `TRUNCATE` anywhere in the class. +- `OnBatchAsync` (`:172-186`) wraps each snapshot row as an ordinary oplog entry and pushes it + through the same `LwwApplier` as a delta. +- `LwwApplier.cs:69-78` discards the incoming row when the local row's HLC is higher. + +So a node with newer local rows keeps them. The failure the directional guard prevents — a stale +peer wiping a live buffer — is structurally impossible, not merely guarded against. **Port +`SfBufferResyncPredicateTests` as a convergence assertion (both nodes end with the union, newest +per id wins), not as a directional-authority assertion.** + +Consequence to state plainly: the standby is no longer guaranteed byte-identical to the active +node's buffer. It is guaranteed *convergent*. That is a real semantic change and Task 2 records it. + +### D3. The SMTP purge is re-homed before anything is deleted + +`HandleApplyArtifacts` calls `PurgeCentralOnlyNotificationConfigAsync` +(`SiteStorageService.cs:811-821`), which deletes `notification_lists` and `smtp_configurations` — +including plaintext SMTP passwords left by pre-fix deployments. This is a **security cleanup riding +the replication path**, not a table sync. CDC will not reproduce it. + +Under CDC, if the active node purges, the deletes replicate as ordinary tombstones. So the fix is +to ensure the *active* node's deploy path performs the purge. Task 12 does this **before** any +deletion, so the purge never stops happening even for one commit. + +### D4. `native_alarm_state` volume is measured, not assumed + +`scadabridge.db` is not only config. `native_alarm_state` mirrors live A&C conditions from +`NativeAlarmActor.cs:504` via batched upserts, and is by a wide margin the highest-volume table in +either database. `sf_messages` worst case is ~50 row-writes/sec (`DefaultMaxRetries = 50`, 10 s +sweep, 4-way target parallelism). Neither number can be turned into an oplog cap from first +principles. + +**Task 1 is a rig soak that measures both.** Its output sets `MaxOplogRows` / `MaxOplogAge` in +Task 19. If the soak shows the shared oplog cannot absorb alarm-storm churn, the escape hatch named +in the design doc (line 140) is keyed instances — that is a library-level effort and would suspend +this plan, so **Task 1 gates everything after Task 2.** + +### D5. Cutover forecloses rolling site upgrades + +`SiteReplicationActor` retains a legacy monolithic `SfBufferSnapshot` handler specifically for +rolling upgrades. With no dual-mechanism period, a site pair cannot be upgraded one node at a time — +one node would speak a protocol the other no longer implements. **Both nodes of a site must be +stopped and started together.** Task 21 puts this in the deployment runbook. + +--- + +## Execution waves + +| Wave | Tasks | Note | +|---|---|---| +| 0 | 1, 2 | Soak + decision record. **Task 1 gates the rest.** | +| 1 | 3, 4 | Schema extraction — independent files, dispatch in parallel | +| 2 | 5, 6 | Store rewiring — independent files, dispatch in parallel | +| 3 | 7, 8, 9 | Setup + migrators | +| 4 | 10, 11 | Port test intents as specs — **before** any deletion | +| 5 | 12, 13 | Re-home the purge, delete notify-and-fetch | +| 6 | 14, 15, 16, 17 | The cutover, serial | +| 7 | 18, 19 | Convergence suite + rig config | +| 8 | 20, 21 | Live gate + docs | + +--- + +### Task 1: Rig soak — measure oplog growth under real write rates + +**Classification:** high-risk +**Estimated implement time:** ~40 min wall-clock (mostly waiting) +**Parallelizable with:** none — this gates the plan + +This is the gate item from `2026-07-19-localdb-phase2-gate.md` §5. It is measurement, not code. + +**Files:** +- Create: `docs/plans/2026-07-19-localdb-phase2-soak.md` + +**Step 1: Bring up the rig with Phase 1 replication enabled** + +```bash +cd ~/Desktop/ScadaBridge +git checkout feat/localdb-phase1 +bash docker/deploy.sh +``` + +Site-a's pair replicates; site-b/site-c deliberately do not (the default-OFF pin). Confirm: + +```bash +docker exec scadabridge-site-a-a curl -s localhost:8080/metrics | grep '^localdb_' +``` + +Expected: `localdb_*` series present. If absent, the meter allowlist regressed — see +`SiteServiceRegistration.ObservedMeters`. + +**Step 2: Capture a baseline** + +```bash +docker exec scadabridge-site-a-a sqlite3 /app/data/site-localdb.db \ + "SELECT COUNT(*) FROM __localdb_oplog;" +``` + +Record the count and the wall-clock time. + +**Step 3: Drive sustained load for 30 minutes** + +Two generators, run concurrently: + +- **Store-and-forward churn.** Point a store-and-forward target at an address that refuses + connections, then enqueue messages. Each failed attempt costs one `UPDATE`. With the default + 10 s sweep this is ~1 write/message/10 s per lane. +- **Alarm churn.** Drive A&C conditions on a deployed instance so `native_alarm_state` upserts + fire. This is the number that actually matters — it is unbounded by design. + +**Step 4: Sample oplog depth every 5 minutes** + +```bash +for i in $(seq 1 6); do + echo "=== t+$((i*5))m ===" + docker exec scadabridge-site-a-a sqlite3 /app/data/site-localdb.db \ + "SELECT COUNT(*) FROM __localdb_oplog;" + docker exec scadabridge-site-a-a curl -s localhost:8080/metrics \ + | grep -E '^localdb_(oplog_backlog|replication_dead_letters|sync_connected)' + sleep 300 +done +``` + +**Step 5: Record findings** + +Write `docs/plans/2026-07-19-localdb-phase2-soak.md` with: +- rows/sec observed for `sf_messages` and for `native_alarm_state`, separately +- peak `localdb_oplog_backlog` +- whether backlog drained between bursts or grew monotonically +- dead-letter count (must be 0) +- **the recommended `MaxOplogRows` and `MaxOplogAge`**, with the arithmetic + +**Step 6: The decision gate** + +If backlog grew monotonically and never drained, the shared oplog cannot absorb this write profile. +**Stop. Do not proceed to Task 3.** Report to the user: the design doc's keyed-instances escape +hatch (design doc line 140) is required first, and that is a library effort in `scadaproj`. + +**Step 7: Commit** + +```bash +git checkout -b feat/localdb-phase2 +git add docs/plans/2026-07-19-localdb-phase2-soak.md +git commit -m "docs(localdb): phase 2 rig soak findings — oplog sizing evidence" +``` + +--- + +### Task 2: Decision record + +**Classification:** trivial +**Estimated implement time:** ~3 min +**Parallelizable with:** none (needs Task 1's numbers) + +**Files:** +- Modify: `docs/plans/2026-07-19-localdb-phase2-gate.md` + +Update the gate's status header from `NOT STARTED` to point at this plan, and answer each §5 +question inline with a one-paragraph resolution referencing D1–D5 above plus Task 1's measured +numbers. Do not delete the questions — the reasoning is the value. + +**Commit:** `docs(localdb): close the phase 2 gate — questions answered, plan written` + +--- + +### Task 3: Extract `StoreAndForwardSchema.Apply` + +**Classification:** small +**Estimated implement time:** ~4 min +**Parallelizable with:** Task 4 + +Mirrors `OperationTrackingSchema` exactly. Read that file first — it is the reference. + +**Files:** +- Create: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardSchema.cs` +- Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:32-117` +- Test: `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardSchemaTests.cs` + +**Step 1: Write the failing test** + +```csharp +[Fact] +public void Apply_IsIdempotent_AndCreatesEveryColumn() +{ + var path = Path.Combine(Path.GetTempPath(), $"sf-schema-{Guid.NewGuid():N}.db"); + try + { + using var connection = new SqliteConnection($"Data Source={path}"); + connection.Open(); + + StoreAndForwardSchema.Apply(connection); + StoreAndForwardSchema.Apply(connection); // second run must not throw + + using var cmd = connection.CreateCommand(); + cmd.CommandText = "SELECT name FROM pragma_table_info('sf_messages')"; + var columns = new List(); + using var reader = cmd.ExecuteReader(); + while (reader.Read()) columns.Add(reader.GetString(0)); + + // 12 original + 4 additive + Assert.Equal(16, columns.Count); + Assert.Contains("last_attempt_at_ms", columns); + Assert.Contains("parent_execution_id", columns); + } + finally + { + SqliteConnection.ClearAllPools(); + File.Delete(path); + } +} +``` + +**Step 2: Run it — expect FAIL** (`StoreAndForwardSchema` does not exist) + +```bash +dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj --filter "FullyQualifiedName~StoreAndForwardSchemaTests" +``` + +**Step 3: Create the schema class** + +Move, verbatim: the DDL literal from `StoreAndForwardStorage.cs:50-68`, the four +`AddColumnIfMissingAsync` calls (`:79-92`) and their helper (`:125-142`), the `last_attempt_at_ms` +backfill (`:98-105`), and the due-index (`:109-114`). + +Convert every `*Async` to sync (`ExecuteNonQuery`, `ExecuteScalar`) — `OnReady` is a synchronous +`Action`. Depend only on `Microsoft.Data.Sqlite`, never on the LocalDb library. + +Carry this comment onto the class, because it is the whole reason the file exists: + +```csharp +/// +/// Deliberately depends only on Microsoft.Data.Sqlite, not on the LocalDb library. +/// The Host applies this DDL to a LocalDb-managed connection before +/// RegisterReplicated installs the capture triggers; nothing about the schema +/// itself is LocalDb-specific, and the store still calls it so a directly-constructed +/// store (tests, tooling) remains self-sufficient. +/// +``` + +**Step 4: Reduce `StoreAndForwardStorage.InitializeAsync` to a call** + +```csharp +public async Task InitializeAsync() +{ + EnsureDatabaseDirectoryExists(); + await using var connection = await OpenConnectionAsync(); + StoreAndForwardSchema.Apply(connection); +} +``` + +**Step 5: Run the S&F suite — expect PASS** + +```bash +dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj +``` + +**Step 6: Commit** + +```bash +git add -A && git commit -m "refactor(sf): extract StoreAndForwardSchema.Apply from the storage class" +``` + +--- + +### Task 4: Extract `SiteStorageSchema.Apply` + +**Classification:** small +**Estimated implement time:** ~5 min +**Parallelizable with:** Task 3 + +**Files:** +- Create: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageSchema.cs` +- Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:58-196` +- Test: `tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Persistence/SiteStorageSchemaTests.cs` + +Same shape as Task 3. Move the DDL literal (`:79-158`, all 9 tables), `MigrateSchemaAsync` +(`:167-181`) and `TryAddColumnAsync` (`:183-196`), converting to sync. + +**Do not move** the `PRAGMA journal_mode=WAL` at `:68-76`. LocalDb owns the connection's pragmas. + +**Test assertion:** all 9 tables exist after `Apply`, and a second `Apply` does not throw. Name +every table explicitly — a loop that asserts "9 tables" would pass if one were renamed: + +```csharp +foreach (var table in new[] + { + "deployed_configurations", "static_attribute_overrides", "shared_scripts", + "external_systems", "database_connections", "notification_lists", + "data_connection_definitions", "smtp_configurations", "native_alarm_state", + }) +{ + Assert.True(TableExists(connection, table), $"missing table: {table}"); +} +``` + +**Commit:** `refactor(site): extract SiteStorageSchema.Apply from SiteStorageService` + +--- + +### Task 5: Rewire `StoreAndForwardStorage` onto `ILocalDb` + +**Classification:** high-risk +**Estimated implement time:** ~5 min +**Parallelizable with:** Task 6 + +Follow `OperationTrackingStore` (Phase 1, commit `0b5e9b44`) as the reference — it solved this +exact problem. + +**Files:** +- Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:22-26,177-185` +- Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ServiceCollectionExtensions.cs:18-25` +- Test: `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/` (fixtures) + +**Step 1: Replace the constructor** + +```csharp +public StoreAndForwardStorage(ILocalDb localDb, ILogger logger) +{ + ArgumentNullException.ThrowIfNull(localDb); + ArgumentNullException.ThrowIfNull(logger); + _localDb = localDb; + _logger = logger; +} +``` + +**Step 2: Replace `OpenConnectionAsync`** + +```csharp +// CreateConnection returns an ALREADY-OPEN, pragma-configured connection with the +// zb_hlc_next() UDF registered. Calling OpenAsync on it throws, and a raw +// SqliteConnection would lack the UDF, making every capture trigger fail closed. +private SqliteConnection OpenConnection() => _localDb.CreateConnection(); +``` + +Update all 18 call sites (`:36, 238, 258, 286, 321, 378, 412, 449, 482, 503, 548, 572, 589, 619, +638, 659, 677, 694`) from `await using var connection = await OpenConnectionAsync();` to +`await using var connection = OpenConnection();`. + +**Step 3: Delete `EnsureDatabaseDirectoryExists` (`:152-168`)** — LocalDb owns the file. + +**Step 4: Update the DI registration** + +```csharp +services.AddSingleton(sp => new StoreAndForwardStorage( + sp.GetRequiredService(), + sp.GetRequiredService>())); +``` + +**Step 5: Fix the test fixtures** + +Every test constructing `StoreAndForwardStorage` with a connection string now needs an `ILocalDb`. +Phase 1 created `tests/ZB.MOM.WW.ScadaBridge.SiteEventLogging.Tests/TestLocalDb.cs` for exactly +this — read it and add an equivalent (or reference it) in the S&F test project. It must use a real +temp **file** (never `:memory:` — capture triggers need a durable file) and call +`SqliteConnection.ClearAllPools()` before deleting, including the `-wal` and `-shm` sidecars. + +**Step 6: Full S&F suite — expect PASS** + +```bash +dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj +``` + +**Commit:** `refactor(sf): StoreAndForwardStorage takes ILocalDb instead of a connection string` + +--- + +### Task 6: Rewire `SiteStorageService` onto `ILocalDb` + +**Classification:** high-risk +**Estimated implement time:** ~5 min +**Parallelizable with:** Task 5 + +**Files:** +- Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:13-51` +- Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/ServiceCollectionExtensions.cs:36-42` +- Modify: `src/ZB.MOM.WW.ScadaBridge.Host/SiteServiceRegistration.cs:70-71` + +Same shape as Task 5, but note the differences: + +- There are **22** inline `new SqliteConnection(_connectionString)` + `OpenAsync()` pairs, not a + helper. Add a `private SqliteConnection OpenConnection() => _localDb.CreateConnection();` and + convert all of them. +- `public SqliteConnection CreateConnection()` at `:51` is a public escape hatch used by site + repositories. Keep the member; change the body to `=> _localDb.CreateConnection();`. +- The busy-timeout floor (`BusyTimeoutFloorSeconds`, `:36`) and its `SqliteConnectionStringBuilder` + normalization are **deleted** — LocalDb configures pragmas on every connection it hands out. +- `AddSiteRuntime(string siteDbConnectionString)` loses its parameter. Keep the no-arg overload at + `:23-28` (it already exists) and delete the string one, updating + `SiteServiceRegistration.cs:70-71` to `services.AddSiteRuntime();`. + +**Watch for:** transactions at `:348` (`RemoveDeployedConfigAsync`) and `:540` +(`UpsertNativeAlarmsAsync`) must keep using `connection.BeginTransaction()` on the LocalDb +connection, not `ILocalDb.BeginTransactionAsync()` — mixing the two would put the cascade's three +deletes in different transactions. + +**Step: Run both affected suites** + +```bash +dotnet test tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests.csproj +dotnet test tests/ZB.MOM.WW.ScadaBridge.Host.Tests/ZB.MOM.WW.ScadaBridge.Host.Tests.csproj +``` + +Host.Tests will surface DI fixtures that now need `LocalDb:Path` — Phase 1 hit the identical +failure (tasks.json Task 4 note). Add the config key to each failing fixture. + +**Commit:** `refactor(site): SiteStorageService takes ILocalDb instead of a connection string` + +--- + +### Task 7: Extend `SiteLocalDbSetup` with the new DDL (not yet registered) + +**Classification:** standard +**Estimated implement time:** ~3 min +**Parallelizable with:** none + +Deliberately applies DDL **without** calling `RegisterReplicated`. The tables live in the +consolidated file but are not yet captured, so the bespoke replicator keeps working unchanged and +the tree stays green. Task 14 flips them on and deletes the bespoke path in one commit. + +**Files:** +- Modify: `src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs:43-57` +- Test: `tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbWiringTests.cs` + +**Step 1: Extend `OnReady`** + +```csharp +using (var connection = db.CreateConnection()) +{ + OperationTrackingSchema.Apply(connection); + SiteEventLogSchema.Apply(connection); + SiteStorageSchema.Apply(connection); + StoreAndForwardSchema.Apply(connection); +} + +db.RegisterReplicated("OperationTracking"); +db.RegisterReplicated("site_events"); + +// Phase 2 tables are created here but NOT yet registered — the bespoke replicator +// still owns them until the Task 14 cutover. Registering them now would double- +// replicate every row (harmless, both paths upsert, but it would mask a defect in +// either one). + +SiteLocalDbLegacyMigrator.Migrate(db, config); +``` + +**Step 2: Pin it** + +Add a test asserting all 4 Phase-1-and-2 table sets exist after `OnReady`, and that +`db.ReplicatedTables` contains **exactly** `OperationTracking` and `site_events` — the "not yet" +is the assertion that matters. + +**Commit:** `feat(localdb): create config + sf_messages tables in the consolidated DB` + +--- + +### Task 8: Extend the legacy migrator for `sf_messages` + +**Classification:** high-risk +**Estimated implement time:** ~5 min +**Parallelizable with:** Task 9 + +Read `SiteLocalDbLegacyMigrator.cs` in full first. `MigrateTracking` (`:116-161`) is the closest +reference — like `sf_messages`, it has a native TEXT PK and needs no id synthesis. + +**Files:** +- Modify: `src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbLegacyMigrator.cs:57-66` +- Test: `tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbLegacyMigratorTests.cs` + +**Step 1: Write the failing tests** + +Three, minimum: + +```csharp +[Fact] public void SfMessages_AreCopiedFromTheLegacyFile() { } +[Fact] public void SfMessages_Migration_IsIdempotent_WhenRerunAfterACrashBeforeRename() { } +[Fact] public void MigratedSfMessages_EnterTheOplog_SoTheyActuallyReplicate() { } +``` + +The third is the one that catches the ordering defect. Phase 1's equivalent asserts on +`__localdb_oplog` directly — copy that assertion. **The oplog table is `__localdb_oplog`** +(`LocalDbSchema.cs:19`), not `zb_oplog`. + +**Step 2: Add `ResolveStoreAndForwardPath` + `MigrateStoreAndForward`** + +```csharp +private const string DefaultStoreAndForwardPath = "./data/store-and-forward.db"; + +internal static string ResolveStoreAndForwardPath(IConfiguration config) => + config["ScadaBridge:StoreAndForward:SqliteDbPath"] ?? DefaultStoreAndForwardPath; +``` + +Migrate all 16 columns with `INSERT OR IGNORE`. The PK is native TEXT — **no id synthesis**, unlike +`MigrateEvents`. + +**Step 3: Wire into `Migrate`** — add the call alongside the existing two. + +**Commit:** `feat(localdb): migrate legacy store-and-forward.db into the consolidated DB` + +--- + +### Task 9: Extend the legacy migrator for the 9 config tables + +**Classification:** high-risk +**Estimated implement time:** ~5 min +**Parallelizable with:** Task 8 + +**Files:** +- Modify: `src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbLegacyMigrator.cs` +- Test: `tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbLegacyMigratorTests.cs` + +All 9 tables migrate from one legacy file (`ScadaBridge:Database:SiteDbPath`), so this is one +`MigrateSiteStorage` that opens the legacy DB once and copies table by table inside a single +transaction, then renames once. + +**Two tables are deliberately skipped:** `notification_lists` and `smtp_configurations`. They are +purged on every deploy and are permanently empty by design (site write paths were removed +2026-07-10). Migrating them would resurrect plaintext SMTP passwords from a pre-fix legacy file into +a **replicated** table. Add this as a comment and as a test: + +```csharp +[Fact] +public void Migration_DoesNotCopyNotificationOrSmtpRows_EvenWhenTheLegacyFileHasThem() +{ + // These carry plaintext SMTP passwords in pre-2026-07-10 files. They are purged + // on every deploy and must never enter a replicated table. +} +``` + +**Step: Verify all 9 tables' columns match** between `SiteStorageSchema` and the migrator's +`INSERT` lists. A column-count mismatch here fails at runtime on a real deployment and nowhere else. + +**Commit:** `feat(localdb): migrate legacy scadabridge.db config tables into the consolidated DB` + +--- + +### Task 10: Port the S&F replication test intents as CDC specs + +**Classification:** standard +**Estimated implement time:** ~5 min +**Parallelizable with:** Task 11 + +**Specifications first, deletion second.** These tests encode behaviour the replacement must still +satisfy. The gate document named two files; the real specification is five. + +**Files:** +- Read (as spec): `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationServiceTests.cs` +- Read (as spec): `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardReplicationTests.cs` +- Create: `tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbStoreAndForwardConvergenceTests.cs` + +Port these **intents** (not the mechanics) as two-node convergence assertions, using Phase 1's +`LocalDbSitePairConvergenceTests.cs` as the harness reference: + +| Intent | Assertion under CDC | +|---|---| +| Add materialises on the peer | enqueue on A → row present on B | +| Remove deletes on the peer | deliver on A → row absent on B | +| Park sets status Parked on the peer | park on A → B's row has status Parked | +| Requeue resets status + RetryCount | requeue on A → B shows Pending, RetryCount 0 | +| Add-then-Remove never converges to present | **the portable form of the strict-ordering test** | +| Apply Add twice is idempotent, newest wins | LWW gives this | +| Park when Add was lost still materialises | upsert semantics | + +**Explicitly do not port:** `ReplicationServiceTests` #10 (200 interleaved ops dispatched +synchronously in strict issue order, 400 observed). That asserts the *mechanism* — inline +fire-and-forget dispatch — not the outcome. CDC capture is asynchronous and batched by construction. +The portable intent is row 5 above. Record this in the test file as a comment so a future reader +does not think it was dropped by accident. + +**Commit:** `test(localdb): port store-and-forward replication intents as CDC convergence specs` + +--- + +### Task 11: Port the resync + directional-authority tests + +**Classification:** standard +**Estimated implement time:** ~5 min +**Parallelizable with:** Task 10 + +**Files:** +- Read (as spec): `tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/SiteReplicationActorTests.cs` +- Read (as spec): `tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/Cluster/SfBufferResyncPredicateTests.cs` +- Create: `tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbConfigConvergenceTests.cs` + +**`SfBufferResyncPredicateTests` is the most valuable test in the set** — the N1 Critical +regression. Per D2 it must be re-expressed, not ported literally: + +```csharp +[Fact] +public async Task ANodeWithNewerLocalRows_KeepsThem_WhenAPeerSnapshotArrives() +{ + // N1 Critical (was: SfBufferResyncPredicateTests). The bespoke replicator needed a + // directional guard because ReplaceAllAsync was a destructive DELETE-then-INSERT, + // so a wrong-direction resync wiped a live buffer. + // + // LocalDb's snapshot resync merges per row under LWW and never deletes + // (SnapshotApplier has no DELETE; LwwApplier.cs:69-78 discards an incoming row + // whose HLC is lower than the local row's). The failure mode is structurally + // impossible rather than guarded against — so this asserts convergence, NOT + // directional authority. There is no active/standby asymmetry left to enforce. +} +``` + +**Also do not port:** actor tests #1–#5 (notify-and-fetch config fetch/retry/supersede). Per D1 the +fetch path is deleted, so these describe code that will not exist. Replace with a single assertion +that a config deploy on A converges to B **without** B making any HTTP call — assert on a fetcher +test double that records zero invocations. That is the positive proof that notify-and-fetch is gone. + +**Commit:** `test(localdb): port resync + config replication intents as CDC convergence specs` + +--- + +### Task 12: Re-home the SMTP purge off the replication path + +**Classification:** high-risk +**Estimated implement time:** ~4 min +**Parallelizable with:** none + +Per D3. **This must land before any deletion** so the purge never lapses, even for one commit. + +**Files:** +- Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Actors/DeploymentManagerActor.cs` (artifact-apply path) +- Test: `tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/DeploymentManagerActorTests.cs` + +**Step 1: Write the failing test** + +```csharp +[Fact] +public async Task ApplyingArtifacts_PurgesCentralOnlyNotificationConfig_OnTheActiveNode() +{ + // Security cleanup: notification_lists and smtp_configurations can hold plaintext + // SMTP passwords from pre-2026-07-10 deployments. This ran on the standby via + // SiteReplicationActor.HandleApplyArtifacts; with that actor deleted the ACTIVE + // node must do it, and CDC replicates the deletes as ordinary tombstones. +} +``` + +**Step 2: Call `PurgeCentralOnlyNotificationConfigAsync` from the active node's artifact-apply path** + +The existing call site is `SiteReplicationActor.HandleApplyArtifacts`. Find where +`DeploymentManagerActor` applies artifacts locally (the same path that reaches +`_replicationActor?.Tell(new ReplicateArtifacts(command))` at `:1952`) and call it there. + +**Step 3: Verify the deletes replicate** — extend the Task 11 convergence suite with a case where +A purges a seeded row and B's copy disappears. + +**Commit:** `fix(site): purge central-only notification config on the active node` + +--- + +### Task 13: Delete the notify-and-fetch config path + +**Classification:** high-risk +**Estimated implement time:** ~4 min +**Parallelizable with:** none + +Per D1. + +**Files:** +- Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:301-336` +- Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/SiteRuntimeOptions.cs:66-68` +- Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/SiteRuntimeOptionsValidator.cs:56-58` + +**Step 1: Delete `StoreDeployedConfigIfNewerAsync`** — its only caller is +`SiteReplicationActor.HandleApplyConfigDeploy`, deleted in Task 15. The unguarded +`StoreDeployedConfigAsync` (`:241-271`) is the active node's path and stays. + +**Step 2: Check `ConfigFetchRetryCount` before removing it** + +`IDeploymentConfigFetcher` is used by **both** the standby replication path *and* the active +singleton's `RefreshDeploymentCommand` path. Only the first is being deleted. + +```bash +grep -rn "ConfigFetchRetryCount" src/ tests/ +``` + +If the active path reads it, **keep the option and its validator rule** and note that in the commit +message. If nothing outside the deleted actor reads it, remove both. + +**Commit:** `refactor(site): delete notify-and-fetch config replication — CDC ships the row` + +--- + +### Task 14: Register the Phase 2 tables and delete `ReplicationService` + +**Classification:** high-risk +**Estimated implement time:** ~5 min +**Parallelizable with:** none + +**The cutover.** Registration and bespoke-path deletion land in one commit so the two mechanisms +never both run. + +**Files:** +- Modify: `src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs` +- Delete: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ReplicationService.cs` +- Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardService.cs:39,243,654,805,836,872,1120,1146` +- Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ServiceCollectionExtensions.cs:63-68` +- Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:284-306` +- Delete: `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationServiceTests.cs` +- Delete: `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardReplicationTests.cs` +- Delete: `tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationWireSerializationPinTests.cs` + +**Step 1: Register all 11 tables** in `SiteLocalDbSetup.OnReady`, replacing the Task 7 comment. +Skip `notification_lists` and `smtp_configurations`? **No — register them.** They must replicate +their *deletes* (Task 12). Register all 9 config tables plus `sf_messages`. + +**Step 2: Delete `ReplicationService`** and remove the ctor parameter + all 6 emission call sites +from `StoreAndForwardService`. + +**Step 3: Delete `ReplaceAllAsync`** (`:284-306`). Per D2 the library's resync merges; a destructive +delete-all must not survive into a replicated table, where CDC would capture the mass-delete and +ship it to the peer. + +**Step 4: Check `GetAllMessagesAsync`'s `Truncated` flag** — it exists only for the resync path. If +nothing else reads it, simplify the return type. + +**Step 5: `CompositionRootTests.cs:474`** asserts `ReplicationService` is a registered site +singleton. Delete that assertion. + +**Step 6: Full S&F + Host suites — expect PASS** + +**Commit:** `feat(localdb)!: replicate sf_messages via CDC, delete ReplicationService` + +--- + +### Task 15: Delete `SiteReplicationActor` and its messages + +**Classification:** high-risk +**Estimated implement time:** ~4 min +**Parallelizable with:** none + +**Files:** +- Delete: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Actors/SiteReplicationActor.cs` +- Delete: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Messages/ReplicationMessages.cs` +- Delete: `tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/SiteReplicationActorTests.cs` +- Delete: `tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/ResyncWireSerializationPinTests.cs` +- Delete: `tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/Cluster/SfBufferResyncPredicateTests.cs` + +The whole `ReplicationMessages.cs` file becomes dead: `ReplicateConfigDeploy/Remove/SetEnabled`, +`ReplicateArtifacts`, `ReplicateStoreAndForward`, `ApplyConfigDeploy/Remove/SetEnabled`, +`ApplyArtifacts`, `ApplyStoreAndForward`, plus the resync records `RequestSfBufferResync`, +`SfBufferSnapshot`, `SfBufferSnapshotChunk`, `SfBufferResyncAck`. + +**Do not delete** `ActiveNodeEvaluator` (`Communication/ClusterState/ActiveNodeEvaluator.cs`) — the +S&F **delivery gate** still needs it. Only its doc comment at `:14` mentions replication; update +the comment, keep the type. + +Confirm the deletions are complete before committing: + +```bash +grep -rn "SiteReplicationActor\|ReplicateStoreAndForward\|SfBufferSnapshot" src/ tests/ +``` + +Expected: no matches. + +**Commit:** `feat(localdb)!: delete SiteReplicationActor — CDC replaces hand-shipped ops` + +--- + +### Task 16: Clean up `DeploymentManagerActor` and `AkkaHostedService` + +**Classification:** standard +**Estimated implement time:** ~4 min +**Parallelizable with:** none + +**Files:** +- Modify: `src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Actors/DeploymentManagerActor.cs:55,161-190,794,970,1004,1102,1952` +- Modify: `src/ZB.MOM.WW.ScadaBridge.Host/Actors/AkkaHostedService.cs:770-796,806-812,1056` + +**Step 1:** Remove the `IActorRef? replicationActor = null` ctor parameter (`:184`), the +`_replicationActor` field (`:55`), and all 5 `_replicationActor?.Tell(...)` sites. + +The parameter is optional and positional — removing it **silently shifts every argument after it** +(`healthCollector`, `serviceProvider`, `loggerFactory`, `configFetcher`, ...). Update the +`Props.Create` at `AkkaHostedService.cs:806-812` and every test constructing this actor. Compile +errors will not catch a positional shift between two same-typed optional parameters — check each +call site by hand. + +**Step 2:** Delete `replicationService` / `replicationLogger` resolution (`:770-772`), the actor +creation (`:785-789`), the handler wiring (`:792-796`), and the actor-inventory comment at `:1056`. + +**Step 3:** `activeNodeCheck` (`:781-783`) is still used by `SiteCommunicationActor` (`:816-821`). +Keep it. + +**Step 4:** Full Host + SiteRuntime + IntegrationTests suites. + +**Commit:** `refactor(site): drop the replication actor from the deployment + host wiring` + +--- + +### Task 17: Config-key cleanup + +**Classification:** standard +**Estimated implement time:** ~4 min +**Parallelizable with:** none + +**Files:** +- Modify: `src/ZB.MOM.WW.ScadaBridge.Host/StartupValidator.cs:113-115` +- Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardOptions.cs:12` +- Modify: `src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardOptionsValidator.cs:20` +- Modify: 9 × `appsettings.Site.json` (Host template + 6 `docker/site-*` + 2 `docker-env2/site-x-*`) + +| Key | Action | +|---|---| +| `ScadaBridge:Database:SiteDbPath` | → migration-only. **Relax the `Require` at `StartupValidator.cs:113`** or every site config keeps it mandatory forever. | +| `ScadaBridge:StoreAndForward:SqliteDbPath` | → migration-only. Relax the non-empty rule at `StoreAndForwardOptionsValidator.cs:20`. | +| `ScadaBridge:StoreAndForward:ReplicationEnabled` | → **fully dead.** Delete the property and all 9 config entries. | + +`ParkedOperationRelayTests.cs:39` and `ParkedMessageHandlerActorTests.cs:34` set +`ReplicationEnabled = false` — they only need the flag to exist, so remove those lines. +`StoreAndForwardOptionsTests.cs:14,26,33` assert its default; delete those assertions. + +**Leave the path keys present in configs** with a comment — they are read by the legacy migrator and +removing them would strand un-migrated data on a node that has not yet started once. + +**Commit:** `chore(config): retire ReplicationEnabled, make legacy db paths migration-only` + +--- + +### Task 18: Two-node convergence suite for the Phase 2 tables + +**Classification:** high-risk +**Estimated implement time:** ~5 min +**Parallelizable with:** Task 19 + +Extend Phase 1's `LocalDbSitePairConvergenceTests.cs` harness (real loopback Kestrel h2c + the real +`LocalDbSyncAuthInterceptor`, initialized via `SiteLocalDbSetup.OnReady`). + +**Files:** +- Modify: `tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbSitePairConvergenceTests.cs` + +Scenarios, beyond Tasks 10–11: + +1. A config deploy on A converges to B, and B's `deployed_configurations` row matches byte for byte. +2. `RemoveDeployedConfigAsync`'s **3-table cascade** converges — `static_attribute_overrides` and + `native_alarm_state` orphans do not survive on B. There are no foreign keys + (`SiteStorageService` DDL has none), so this is three independent delete streams that LWW may + reorder. **This is the most likely real defect in the whole plan.** +3. A burst of `native_alarm_state` upserts converges and the oplog drains. +4. Node B restarts and catches up via snapshot resync without losing rows written on B while down. + +**Verify the suite is non-vacuous.** Phase 1 did this by mismatching the two API keys and confirming +all scenarios went red. Do the same and record the result in the commit message. A convergence suite +that cannot fail proves nothing. + +**Commit:** `test(localdb): two-node convergence for config tables + sf_messages` + +--- + +### Task 19: Rig configuration + +**Classification:** small +**Estimated implement time:** ~3 min +**Parallelizable with:** Task 18 + +**Files:** +- Modify: `docker/site-a-node-a/appsettings.Site.json`, `docker/site-a-node-b/appsettings.Site.json` + +Set `MaxOplogRows` and `MaxOplogAge` from **Task 1's measured numbers** — not from a guess. Keep +site-b/site-c unreplicated so default-OFF stays proven side by side on one rig. + +**Commit:** `chore(docker): size the site-a oplog caps from the phase 2 soak` + +--- + +### Task 20: Live gate on the docker rig + +**Classification:** high-risk +**Estimated implement time:** ~30 min wall-clock +**Parallelizable with:** none + +**Files:** +- Create: `docs/plans/2026-07-19-localdb-phase2-live-gate.md` + +```bash +cd ~/Desktop/ScadaBridge && bash docker/deploy.sh +``` + +Evidence required — each must be captured, not asserted: + +1. **Migration ran.** `store-and-forward.db.migrated` and `scadabridge.db.migrated` exist on both + site-a nodes; row counts in the consolidated DB match the legacy files. +2. **No SMTP/notification rows migrated** — `SELECT COUNT(*) FROM smtp_configurations` is 0. +3. **Config converges.** Deploy an instance to site-a; both nodes' `deployed_configurations` rows + are byte-identical with identical `__localdb_row_version` and originating node id. +4. **The standby made no HTTP fetch** — nothing in its log resembling a config fetch. This is the + positive proof that notify-and-fetch is gone. +5. **S&F converges.** Enqueue against a dead target; the parked row appears on the peer. +6. **Site failover.** Flip the site-a pair and confirm buffered messages survive and deliver + exactly once after the flip. `docker/failover-drill.sh` targets **CENTRAL** nodes and does not + exercise this — run the site flip directly. +7. **Cascade delete.** Remove a deployed config; confirm no orphan `static_attribute_overrides` or + `native_alarm_state` rows on the peer. +8. **Zero dead letters**, oplog drained, `localdb_*` scraped from `/metrics`. +9. **Both nodes stopped and started together** (D5) — confirm a clean rejoin. + +**If any check fails, stop and report.** Do not proceed to Task 21. + +**Commit:** `docs(localdb): phase 2 live gate evidence` + +--- + +### Task 21: Documentation truth pass + +**Classification:** standard +**Estimated implement time:** ~5 min +**Parallelizable with:** none + +**Files:** +- Modify: `CLAUDE.md` (ScadaBridge), `~/Desktop/scadaproj/CLAUDE.md` (LocalDb component row) +- Modify: `docs/requirements/Component-StoreAndForward.md:83` +- Modify: `docs/components/StoreAndForward.md`, `SiteRuntime.md`, `Host.md` +- Modify: `docs/known-issues/2026-06-26-deploy-config-exceeds-akka-frame-size.md` +- Modify: `docs/deployment/` runbook + +Specific corrections: + +- `Component-StoreAndForward.md:83` is a long **normative** paragraph specifying the whole chunked + resync protocol and the N5 race. Rewrite for CDC — state the new duplicate-delivery bound + explicitly rather than deleting the discussion. +- The frame-size known-issue is **resolved by** this change (D1). Mark it resolved and say why, + rather than leaving a live known-issue describing deleted code. +- **Add to the deployment runbook (D5):** a site pair must be stopped and started together. Rolling + one node at a time is no longer supported, because the legacy `SfBufferSnapshot` compatibility + handler is gone. +- Update the scadaproj LocalDb component row: Phase 2 complete, what replicates now, what + `SiteReplicationActor` used to do. + +**Commit:** `docs(localdb): phase 2 truth pass across both repos` + +--- + +## Definition of done + +- [ ] Solution builds with **0 warnings** (`TreatWarningsAsErrors` is on) +- [ ] Every suite green: Host, CentralUI, Commons, SiteRuntime, AuditLog, Communication, + StoreAndForward, HealthMonitoring, IntegrationTests, SiteEventLogging +- [ ] `grep -rn "SiteReplicationActor\|ReplicationService\|SfBufferSnapshot" src/` → no matches +- [ ] Live gate PASS with all 9 evidence items captured +- [ ] Convergence suite verified non-vacuous +- [ ] Both `CLAUDE.md` files updated diff --git a/docs/plans/2026-07-19-localdb-adoption-phase2.md.tasks.json b/docs/plans/2026-07-19-localdb-adoption-phase2.md.tasks.json new file mode 100644 index 00000000..41bd42b7 --- /dev/null +++ b/docs/plans/2026-07-19-localdb-adoption-phase2.md.tasks.json @@ -0,0 +1,89 @@ +{ + "planPath": "docs/plans/2026-07-19-localdb-adoption-phase2.md", + "execution": { + "mode": "parallel-waves", + "implementerModel": "opus", + "isolation": "worktree", + "branch": "feat/localdb-phase2", + "baseBranch": "feat/localdb-phase1", + "note": "Phase 1's branch is NOT merged/pushed, so phase 2 branches from it. Dispatch every unblocked task concurrently per the wave table in the plan. Parallel implementers MUST use worktree isolation - concurrent git in one worktree races destructively." + }, + "scopeDecision": { + "date": "2026-07-19", + "by": "user", + "choice": "Full scope as designed", + "note": "User chose both surfaces (config tables + sf_messages) in one phase, deleting both bespoke mechanisms together, over the recommended split of S&F-first. The four open design questions are therefore resolved INSIDE this plan as D1-D5 rather than deferred." + }, + "decisions": [ + { + "id": "D1", + "subject": "Config moves to CDC; notify-and-fetch is DELETED, not preserved", + "evidence": "SiteReplicationActor sends id+fetch-coords only because the config blob exceeds Akka's 128KB frame (docs/known-issues/2026-06-26-deploy-config-exceeds-akka-frame-size.md). LocalDb sync is gRPC - no such limit. StoreDeployedConfigIfNewerAsync's guard (SiteStorageService.cs:325, `WHERE excluded.deployed_at > deployed_configurations.deployed_at`) protects only against a stale FETCH racing; with no fetch there is no race. Do NOT reproduce the guard on top of LWW - deployed_at and HLC are different clocks and mixing them is non-convergent." + }, + { + "id": "D2", + "subject": "ReplaceAllAsync deleted; the N1 directional guard becomes unnecessary", + "evidence": "LocalDb's snapshot resync MERGES per-row LWW and never deletes: SnapshotApplier.OnBeginAsync (SnapshotStreamer.cs:163-170) resets counters only, no DELETE anywhere in the class; OnBatchAsync (:172-186) routes snapshot rows through the same LwwApplier as deltas; LwwApplier.cs:69-78 discards an incoming row whose HLC is lower. So a stale peer cannot wipe a live buffer - structurally impossible, not guarded. SEMANTIC CHANGE: the standby is convergent, no longer byte-identical." + }, + { + "id": "D3", + "subject": "The SMTP purge is re-homed BEFORE anything is deleted", + "evidence": "SiteReplicationActor.HandleApplyArtifacts calls PurgeCentralOnlyNotificationConfigAsync (SiteStorageService.cs:811-821), deleting notification_lists + smtp_configurations incl. plaintext SMTP passwords from pre-2026-07-10 deployments. A security cleanup riding the replication path; CDC will not reproduce it. Task 12 moves it to the active node's deploy path before any deletion, so it never lapses." + }, + { + "id": "D4", + "subject": "native_alarm_state volume is MEASURED, not assumed", + "evidence": "scadabridge.db is not only config - native_alarm_state mirrors live A&C conditions (NativeAlarmActor.cs:504) and is the highest-volume table in either DB. sf_messages worst case ~50 row-writes/sec. Task 1 is a rig soak that measures both and sets MaxOplogRows/MaxOplogAge. If backlog grows monotonically, STOP: the design doc's keyed-instances escape hatch (line 140) is a scadaproj library effort that would suspend this plan." + }, + { + "id": "D5", + "subject": "Cutover forecloses rolling site upgrades", + "evidence": "SiteReplicationActor retains a legacy monolithic SfBufferSnapshot handler for rolling upgrades. With no dual-mechanism period, one node would speak a protocol the other no longer implements. Both nodes of a site must be stopped and started together. Task 21 puts this in the deployment runbook." + } + ], + "reconFindings": [ + "The gate doc named 2 test files as the specification; the real spec is 5 - it missed StoreAndForwardReplicationTests.cs (incl. the only Requeue coverage), ReplicationWireSerializationPinTests.cs, ResyncWireSerializationPinTests.cs, and SfBufferResyncPredicateTests.cs (the N1 Critical regression test).", + "sf_messages has NO version column today - ON CONFLICT(id) DO UPDATE has no comparison predicate (StoreAndForwardStorage.cs:331-345). 'Newest wins' is bare arrival order. LWW-by-HLC is an IMPROVEMENT here, not a regression.", + "No autoincrement-integer PKs exist anywhere in Phase 2 scope - all 9 config tables use natural TEXT or composite TEXT keys, sf_messages is TEXT. Phase 1's site_events GUID conversion has no Phase 2 analogue.", + "SiteStorageService has NO foreign keys. RemoveDeployedConfigAsync (:343-376) is a manual 3-statement cascade in one transaction. Under CDC these become three independent delete streams that LWW may reorder - the most likely real defect in the plan (Task 18 scenario 2).", + "DeploymentManagerActor's replicationActor is an OPTIONAL POSITIONAL parameter (:184). Removing it silently shifts every argument after it. Compilers will not catch a shift between two same-typed optional params - check every call site by hand (Task 16).", + "ActiveNodeEvaluator must NOT be deleted - the S&F delivery gate still uses it. Only its doc comment mentions replication.", + "ConfigFetchRetryCount may not be fully dead: IDeploymentConfigFetcher serves BOTH the standby replication path and the active singleton's RefreshDeploymentCommand path. Verify before removing (Task 13 step 2).", + "notification_lists and smtp_configurations are deliberately NOT migrated (Task 9) - migrating them would resurrect plaintext SMTP passwords from a pre-fix legacy file into a replicated table." + ], + "tasks": [ + {"id": 1, "subject": "Task 1: Rig soak - measure oplog growth under real write rates", "status": "pending", "classification": "high-risk", "note": "GATES EVERYTHING AFTER TASK 2. Stop the plan if oplog backlog grows monotonically."}, + {"id": 2, "subject": "Task 2: Decision record - close the phase 2 gate", "status": "pending", "classification": "trivial", "blockedBy": [1]}, + {"id": 3, "subject": "Task 3: Extract StoreAndForwardSchema.Apply", "status": "pending", "classification": "small", "blockedBy": [2]}, + {"id": 4, "subject": "Task 4: Extract SiteStorageSchema.Apply", "status": "pending", "classification": "small", "blockedBy": [2]}, + {"id": 5, "subject": "Task 5: Rewire StoreAndForwardStorage onto ILocalDb", "status": "pending", "classification": "high-risk", "blockedBy": [3]}, + {"id": 6, "subject": "Task 6: Rewire SiteStorageService onto ILocalDb", "status": "pending", "classification": "high-risk", "blockedBy": [4]}, + {"id": 7, "subject": "Task 7: Extend SiteLocalDbSetup with the new DDL (not yet registered)", "status": "pending", "classification": "standard", "blockedBy": [5, 6]}, + {"id": 8, "subject": "Task 8: Extend the legacy migrator for sf_messages", "status": "pending", "classification": "high-risk", "blockedBy": [7]}, + {"id": 9, "subject": "Task 9: Extend the legacy migrator for the 9 config tables", "status": "pending", "classification": "high-risk", "blockedBy": [7]}, + {"id": 10, "subject": "Task 10: Port the S&F replication test intents as CDC specs", "status": "pending", "classification": "standard", "blockedBy": [8]}, + {"id": 11, "subject": "Task 11: Port the resync + directional-authority tests", "status": "pending", "classification": "standard", "blockedBy": [9]}, + {"id": 12, "subject": "Task 12: Re-home the SMTP purge off the replication path", "status": "pending", "classification": "high-risk", "blockedBy": [10, 11], "note": "MUST land before any deletion so the security purge never lapses."}, + {"id": 13, "subject": "Task 13: Delete the notify-and-fetch config path", "status": "pending", "classification": "high-risk", "blockedBy": [12]}, + {"id": 14, "subject": "Task 14: Register the Phase 2 tables and delete ReplicationService", "status": "pending", "classification": "high-risk", "blockedBy": [13], "note": "THE CUTOVER. Registration + bespoke deletion in ONE commit so both mechanisms never run together."}, + {"id": 15, "subject": "Task 15: Delete SiteReplicationActor and its messages", "status": "pending", "classification": "high-risk", "blockedBy": [14]}, + {"id": 16, "subject": "Task 16: Clean up DeploymentManagerActor and AkkaHostedService", "status": "pending", "classification": "standard", "blockedBy": [15]}, + {"id": 17, "subject": "Task 17: Config-key cleanup", "status": "pending", "classification": "standard", "blockedBy": [16]}, + {"id": 18, "subject": "Task 18: Two-node convergence suite for the Phase 2 tables", "status": "pending", "classification": "high-risk", "blockedBy": [17]}, + {"id": 19, "subject": "Task 19: Rig configuration", "status": "pending", "classification": "small", "blockedBy": [17]}, + {"id": 20, "subject": "Task 20: Live gate on the docker rig", "status": "pending", "classification": "high-risk", "blockedBy": [18, 19]}, + {"id": 21, "subject": "Task 21: Documentation truth pass", "status": "pending", "classification": "standard", "blockedBy": [20]} + ], + "knownFlakes": [ + { + "test": "SiteRuntime.Tests InstanceActorChildAttributeRaceTests.ChildActors_AreSeededFromAnIsolatedCopy_NotTheLiveAttributesDictionary", + "note": "Intermittent ActorNotFoundException under full-suite load; passes in isolation. Pre-existing, carried over from Phase 1." + }, + { + "test": "AuditLog.Tests ParentExecutionIdCorrelationTests.InboundRoutedRun_AllRoutedRows_CarryInboundExecutionId_AsParentExecutionId", + "note": "Cold-MSSQL-fixture timing: ~91s and AwaitAssert-times-out cold, ~1s warm. Re-run before investigating." + } + ], + "lastUpdated": "2026-07-19", + "phase2Status": "PLANNED - not started. Task 1 (rig soak) gates the rest." +}