Moves scadabridge.db's 9 config tables + sf_messages into the consolidated LocalDb file and deletes SiteReplicationActor + StoreAndForward's ReplicationService. Resolves the phase 2 gate's five open questions from code recon (D1-D5): D1 notify-and-fetch is DELETED, not preserved. It exists only because the config blob exceeds Akka's 128KB frame; LocalDb sync is gRPC. The deployed_at version guard protected against a stale fetch racing, so it dies with the fetch. Do not reproduce it on LWW - different clocks. D2 ReplaceAllAsync is deleted and the N1 directional guard becomes unnecessary: LocalDb's snapshot resync merges per-row LWW and never deletes (SnapshotApplier has no DELETE; LwwApplier.cs:69-78 discards a lower-HLC incoming row). Semantic change - the standby is convergent, no longer byte-identical. D3 The SMTP purge (plaintext passwords) rides the replication path and is re-homed to the active node BEFORE any deletion. D4 native_alarm_state volume is measured by a rig soak, not assumed. Task 1 gates the plan and stops it if the oplog cannot absorb the churn. D5 No dual-mechanism period forecloses rolling site upgrades - both nodes must stop and start together. Recon also found the gate doc's "two test files are the spec" undercounts: the real specification is five files, including the N1 Critical regression test and the only Requeue coverage. Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts
42 KiB
ScadaBridge LocalDb Phase 2 Implementation Plan
For Claude: REQUIRED SUB-SKILL: Use superpowers-extended-cc:executing-plans to implement this plan task-by-task.
Goal: Move scadabridge.db's 9 config tables and StoreAndForward's sf_messages into the
consolidated ZB.MOM.WW.LocalDb database established by Phase 1, then delete SiteReplicationActor
and StoreAndForward's ReplicationService outright.
Architecture: Both stores stop owning their own SQLite files and take ILocalDb instead, exactly
as OperationTrackingStore and SiteEventLogger did in Phase 1. Their DDL is extracted into
*Schema.Apply(SqliteConnection) helpers that SiteLocalDbSetup.OnReady runs before
RegisterReplicated. Trigger-based CDC then replicates every row mutation, replacing the bespoke
hand-shipped ops; the library's per-row-LWW snapshot resync replaces the chunked anti-entropy.
Tech Stack: .NET 10, Akka.NET, Microsoft.Data.Sqlite, ZB.MOM.WW.LocalDb 0.1.0, gRPC, xunit.
Branch: feat/localdb-phase2, cut from feat/localdb-phase1.
Read this before Task 1
Phase 1's record is the prerequisite context. Read, in order:
docs/plans/2026-07-19-localdb-adoption-phase1.md.tasks.json— theamendments,prerequisitesandknownFlakesarrays. Several Phase 1 assumptions changed during execution.docs/plans/2026-07-19-localdb-phase2-gate.md— the gate. §5's open questions are answered in the Decisions section below; the gate itself is superseded by this plan.src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs— the pattern every new table follows.
Two known flakes — do not chase these
SiteRuntime.Tests InstanceActorChildAttributeRaceTests— intermittentActorNotFoundExceptionunder full-suite load. Passes in isolation. Pre-existing.AuditLog.Tests ParentExecutionIdCorrelationTests— ~91 s and times out on a cold MSSQL fixture, ~1 s warm. Re-run before investigating.
Decisions — the gate's open questions, answered
These were resolved from code recon before this plan was written. Each names its evidence. Task 2 records them formally; the rest of the plan assumes them.
D1. Config moves to CDC, and notify-and-fetch is deleted, not preserved
SiteReplicationActor never sends config JSON over the replication hop — it sends a deployment id
plus fetch coordinates, and the standby HTTP-fetches from central itself
(SiteReplicationActorTests #2, #5). That design exists for exactly one reason: the config blob
exceeds Akka's 128 KB frame limit (docs/known-issues/2026-06-26-deploy-config-exceeds-akka-frame-size.md).
Under CDC the row itself replicates over the LocalDb gRPC stream, which has no such limit. The standby therefore never fetches at all.
This also disposes of the version guard. StoreDeployedConfigIfNewerAsync
(SiteStorageService.cs:301-336) guards its upsert with:
WHERE excluded.deployed_at > deployed_configurations.deployed_at
That predicate protects against a stale fetch landing after a newer one — a race that only exists because the standby was independently fetching. With the fetch gone, the only writer is the active node's deploy path, and HLC ordering on a single writer is monotonic. The guard becomes dead code.
Do not try to reproduce the
deployed_atguard on top of LWW. They order by different clocks (central's deploy time vs. the writing node's HLC) and mixing them produces a non-convergent store.
D2. ReplaceAllAsync is deleted, and the N1 directional guard becomes unnecessary
StoreAndForwardStorage.ReplaceAllAsync (StoreAndForwardStorage.cs:284-306) is an unqualified
DELETE FROM sf_messages followed by a full re-insert. Its own doc warns "Never call on an active
node — it discards every in-flight row." SfBufferResyncPredicateTests exists to enforce that only
a standby ever applies it — the N1 Critical regression test.
The LocalDb library's snapshot resync merges per row under LWW and never deletes:
SnapshotApplier.OnBeginAsync(SnapshotStreamer.cs:163-170) resets counters only. There is noDELETEorTRUNCATEanywhere in the class.OnBatchAsync(:172-186) wraps each snapshot row as an ordinary oplog entry and pushes it through the sameLwwApplieras a delta.LwwApplier.cs:69-78discards the incoming row when the local row's HLC is higher.
So a node with newer local rows keeps them. The failure the directional guard prevents — a stale
peer wiping a live buffer — is structurally impossible, not merely guarded against. Port
SfBufferResyncPredicateTests as a convergence assertion (both nodes end with the union, newest
per id wins), not as a directional-authority assertion.
Consequence to state plainly: the standby is no longer guaranteed byte-identical to the active node's buffer. It is guaranteed convergent. That is a real semantic change and Task 2 records it.
D3. The SMTP purge is re-homed before anything is deleted
HandleApplyArtifacts calls PurgeCentralOnlyNotificationConfigAsync
(SiteStorageService.cs:811-821), which deletes notification_lists and smtp_configurations —
including plaintext SMTP passwords left by pre-fix deployments. This is a security cleanup riding
the replication path, not a table sync. CDC will not reproduce it.
Under CDC, if the active node purges, the deletes replicate as ordinary tombstones. So the fix is to ensure the active node's deploy path performs the purge. Task 12 does this before any deletion, so the purge never stops happening even for one commit.
D4. native_alarm_state volume is measured, not assumed
scadabridge.db is not only config. native_alarm_state mirrors live A&C conditions from
NativeAlarmActor.cs:504 via batched upserts, and is by a wide margin the highest-volume table in
either database. sf_messages worst case is ~50 row-writes/sec (DefaultMaxRetries = 50, 10 s
sweep, 4-way target parallelism). Neither number can be turned into an oplog cap from first
principles.
Task 1 is a rig soak that measures both. Its output sets MaxOplogRows / MaxOplogAge in
Task 19. If the soak shows the shared oplog cannot absorb alarm-storm churn, the escape hatch named
in the design doc (line 140) is keyed instances — that is a library-level effort and would suspend
this plan, so Task 1 gates everything after Task 2.
D5. Cutover forecloses rolling site upgrades
SiteReplicationActor retains a legacy monolithic SfBufferSnapshot handler specifically for
rolling upgrades. With no dual-mechanism period, a site pair cannot be upgraded one node at a time —
one node would speak a protocol the other no longer implements. Both nodes of a site must be
stopped and started together. Task 21 puts this in the deployment runbook.
Execution waves
| Wave | Tasks | Note |
|---|---|---|
| 0 | 1, 2 | Soak + decision record. Task 1 gates the rest. |
| 1 | 3, 4 | Schema extraction — independent files, dispatch in parallel |
| 2 | 5, 6 | Store rewiring — independent files, dispatch in parallel |
| 3 | 7, 8, 9 | Setup + migrators |
| 4 | 10, 11 | Port test intents as specs — before any deletion |
| 5 | 12, 13 | Re-home the purge, delete notify-and-fetch |
| 6 | 14, 15, 16, 17 | The cutover, serial |
| 7 | 18, 19 | Convergence suite + rig config |
| 8 | 20, 21 | Live gate + docs |
Task 1: Rig soak — measure oplog growth under real write rates
Classification: high-risk Estimated implement time: ~40 min wall-clock (mostly waiting) Parallelizable with: none — this gates the plan
This is the gate item from 2026-07-19-localdb-phase2-gate.md §5. It is measurement, not code.
Files:
- Create:
docs/plans/2026-07-19-localdb-phase2-soak.md
Step 1: Bring up the rig with Phase 1 replication enabled
cd ~/Desktop/ScadaBridge
git checkout feat/localdb-phase1
bash docker/deploy.sh
Site-a's pair replicates; site-b/site-c deliberately do not (the default-OFF pin). Confirm:
docker exec scadabridge-site-a-a curl -s localhost:8080/metrics | grep '^localdb_'
Expected: localdb_* series present. If absent, the meter allowlist regressed — see
SiteServiceRegistration.ObservedMeters.
Step 2: Capture a baseline
docker exec scadabridge-site-a-a sqlite3 /app/data/site-localdb.db \
"SELECT COUNT(*) FROM __localdb_oplog;"
Record the count and the wall-clock time.
Step 3: Drive sustained load for 30 minutes
Two generators, run concurrently:
- Store-and-forward churn. Point a store-and-forward target at an address that refuses
connections, then enqueue messages. Each failed attempt costs one
UPDATE. With the default 10 s sweep this is ~1 write/message/10 s per lane. - Alarm churn. Drive A&C conditions on a deployed instance so
native_alarm_stateupserts fire. This is the number that actually matters — it is unbounded by design.
Step 4: Sample oplog depth every 5 minutes
for i in $(seq 1 6); do
echo "=== t+$((i*5))m ==="
docker exec scadabridge-site-a-a sqlite3 /app/data/site-localdb.db \
"SELECT COUNT(*) FROM __localdb_oplog;"
docker exec scadabridge-site-a-a curl -s localhost:8080/metrics \
| grep -E '^localdb_(oplog_backlog|replication_dead_letters|sync_connected)'
sleep 300
done
Step 5: Record findings
Write docs/plans/2026-07-19-localdb-phase2-soak.md with:
- rows/sec observed for
sf_messagesand fornative_alarm_state, separately - peak
localdb_oplog_backlog - whether backlog drained between bursts or grew monotonically
- dead-letter count (must be 0)
- the recommended
MaxOplogRowsandMaxOplogAge, with the arithmetic
Step 6: The decision gate
If backlog grew monotonically and never drained, the shared oplog cannot absorb this write profile.
Stop. Do not proceed to Task 3. Report to the user: the design doc's keyed-instances escape
hatch (design doc line 140) is required first, and that is a library effort in scadaproj.
Step 7: Commit
git checkout -b feat/localdb-phase2
git add docs/plans/2026-07-19-localdb-phase2-soak.md
git commit -m "docs(localdb): phase 2 rig soak findings — oplog sizing evidence"
Task 2: Decision record
Classification: trivial Estimated implement time: ~3 min Parallelizable with: none (needs Task 1's numbers)
Files:
- Modify:
docs/plans/2026-07-19-localdb-phase2-gate.md
Update the gate's status header from NOT STARTED to point at this plan, and answer each §5
question inline with a one-paragraph resolution referencing D1–D5 above plus Task 1's measured
numbers. Do not delete the questions — the reasoning is the value.
Commit: docs(localdb): close the phase 2 gate — questions answered, plan written
Task 3: Extract StoreAndForwardSchema.Apply
Classification: small Estimated implement time: ~4 min Parallelizable with: Task 4
Mirrors OperationTrackingSchema exactly. Read that file first — it is the reference.
Files:
- Create:
src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardSchema.cs - Modify:
src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:32-117 - Test:
tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardSchemaTests.cs
Step 1: Write the failing test
[Fact]
public void Apply_IsIdempotent_AndCreatesEveryColumn()
{
var path = Path.Combine(Path.GetTempPath(), $"sf-schema-{Guid.NewGuid():N}.db");
try
{
using var connection = new SqliteConnection($"Data Source={path}");
connection.Open();
StoreAndForwardSchema.Apply(connection);
StoreAndForwardSchema.Apply(connection); // second run must not throw
using var cmd = connection.CreateCommand();
cmd.CommandText = "SELECT name FROM pragma_table_info('sf_messages')";
var columns = new List<string>();
using var reader = cmd.ExecuteReader();
while (reader.Read()) columns.Add(reader.GetString(0));
// 12 original + 4 additive
Assert.Equal(16, columns.Count);
Assert.Contains("last_attempt_at_ms", columns);
Assert.Contains("parent_execution_id", columns);
}
finally
{
SqliteConnection.ClearAllPools();
File.Delete(path);
}
}
Step 2: Run it — expect FAIL (StoreAndForwardSchema does not exist)
dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj --filter "FullyQualifiedName~StoreAndForwardSchemaTests"
Step 3: Create the schema class
Move, verbatim: the DDL literal from StoreAndForwardStorage.cs:50-68, the four
AddColumnIfMissingAsync calls (:79-92) and their helper (:125-142), the last_attempt_at_ms
backfill (:98-105), and the due-index (:109-114).
Convert every *Async to sync (ExecuteNonQuery, ExecuteScalar) — OnReady is a synchronous
Action<ILocalDb>. Depend only on Microsoft.Data.Sqlite, never on the LocalDb library.
Carry this comment onto the class, because it is the whole reason the file exists:
/// <para>
/// Deliberately depends only on <c>Microsoft.Data.Sqlite</c>, not on the LocalDb library.
/// The Host applies this DDL to a LocalDb-managed connection before
/// <c>RegisterReplicated</c> installs the capture triggers; nothing about the schema
/// itself is LocalDb-specific, and the store still calls it so a directly-constructed
/// store (tests, tooling) remains self-sufficient.
/// </para>
Step 4: Reduce StoreAndForwardStorage.InitializeAsync to a call
public async Task InitializeAsync()
{
EnsureDatabaseDirectoryExists();
await using var connection = await OpenConnectionAsync();
StoreAndForwardSchema.Apply(connection);
}
Step 5: Run the S&F suite — expect PASS
dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj
Step 6: Commit
git add -A && git commit -m "refactor(sf): extract StoreAndForwardSchema.Apply from the storage class"
Task 4: Extract SiteStorageSchema.Apply
Classification: small Estimated implement time: ~5 min Parallelizable with: Task 3
Files:
- Create:
src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageSchema.cs - Modify:
src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:58-196 - Test:
tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Persistence/SiteStorageSchemaTests.cs
Same shape as Task 3. Move the DDL literal (:79-158, all 9 tables), MigrateSchemaAsync
(:167-181) and TryAddColumnAsync (:183-196), converting to sync.
Do not move the PRAGMA journal_mode=WAL at :68-76. LocalDb owns the connection's pragmas.
Test assertion: all 9 tables exist after Apply, and a second Apply does not throw. Name
every table explicitly — a loop that asserts "9 tables" would pass if one were renamed:
foreach (var table in new[]
{
"deployed_configurations", "static_attribute_overrides", "shared_scripts",
"external_systems", "database_connections", "notification_lists",
"data_connection_definitions", "smtp_configurations", "native_alarm_state",
})
{
Assert.True(TableExists(connection, table), $"missing table: {table}");
}
Commit: refactor(site): extract SiteStorageSchema.Apply from SiteStorageService
Task 5: Rewire StoreAndForwardStorage onto ILocalDb
Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 6
Follow OperationTrackingStore (Phase 1, commit 0b5e9b44) as the reference — it solved this
exact problem.
Files:
- Modify:
src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:22-26,177-185 - Modify:
src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ServiceCollectionExtensions.cs:18-25 - Test:
tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/(fixtures)
Step 1: Replace the constructor
public StoreAndForwardStorage(ILocalDb localDb, ILogger<StoreAndForwardStorage> logger)
{
ArgumentNullException.ThrowIfNull(localDb);
ArgumentNullException.ThrowIfNull(logger);
_localDb = localDb;
_logger = logger;
}
Step 2: Replace OpenConnectionAsync
// CreateConnection returns an ALREADY-OPEN, pragma-configured connection with the
// zb_hlc_next() UDF registered. Calling OpenAsync on it throws, and a raw
// SqliteConnection would lack the UDF, making every capture trigger fail closed.
private SqliteConnection OpenConnection() => _localDb.CreateConnection();
Update all 18 call sites (:36, 238, 258, 286, 321, 378, 412, 449, 482, 503, 548, 572, 589, 619, 638, 659, 677, 694) from await using var connection = await OpenConnectionAsync(); to
await using var connection = OpenConnection();.
Step 3: Delete EnsureDatabaseDirectoryExists (:152-168) — LocalDb owns the file.
Step 4: Update the DI registration
services.AddSingleton<StoreAndForwardStorage>(sp => new StoreAndForwardStorage(
sp.GetRequiredService<ILocalDb>(),
sp.GetRequiredService<ILogger<StoreAndForwardStorage>>()));
Step 5: Fix the test fixtures
Every test constructing StoreAndForwardStorage with a connection string now needs an ILocalDb.
Phase 1 created tests/ZB.MOM.WW.ScadaBridge.SiteEventLogging.Tests/TestLocalDb.cs for exactly
this — read it and add an equivalent (or reference it) in the S&F test project. It must use a real
temp file (never :memory: — capture triggers need a durable file) and call
SqliteConnection.ClearAllPools() before deleting, including the -wal and -shm sidecars.
Step 6: Full S&F suite — expect PASS
dotnet test tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests.csproj
Commit: refactor(sf): StoreAndForwardStorage takes ILocalDb instead of a connection string
Task 6: Rewire SiteStorageService onto ILocalDb
Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 5
Files:
- Modify:
src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:13-51 - Modify:
src/ZB.MOM.WW.ScadaBridge.SiteRuntime/ServiceCollectionExtensions.cs:36-42 - Modify:
src/ZB.MOM.WW.ScadaBridge.Host/SiteServiceRegistration.cs:70-71
Same shape as Task 5, but note the differences:
- There are 22 inline
new SqliteConnection(_connectionString)+OpenAsync()pairs, not a helper. Add aprivate SqliteConnection OpenConnection() => _localDb.CreateConnection();and convert all of them. public SqliteConnection CreateConnection()at:51is a public escape hatch used by site repositories. Keep the member; change the body to=> _localDb.CreateConnection();.- The busy-timeout floor (
BusyTimeoutFloorSeconds,:36) and itsSqliteConnectionStringBuildernormalization are deleted — LocalDb configures pragmas on every connection it hands out. AddSiteRuntime(string siteDbConnectionString)loses its parameter. Keep the no-arg overload at:23-28(it already exists) and delete the string one, updatingSiteServiceRegistration.cs:70-71toservices.AddSiteRuntime();.
Watch for: transactions at :348 (RemoveDeployedConfigAsync) and :540
(UpsertNativeAlarmsAsync) must keep using connection.BeginTransaction() on the LocalDb
connection, not ILocalDb.BeginTransactionAsync() — mixing the two would put the cascade's three
deletes in different transactions.
Step: Run both affected suites
dotnet test tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests.csproj
dotnet test tests/ZB.MOM.WW.ScadaBridge.Host.Tests/ZB.MOM.WW.ScadaBridge.Host.Tests.csproj
Host.Tests will surface DI fixtures that now need LocalDb:Path — Phase 1 hit the identical
failure (tasks.json Task 4 note). Add the config key to each failing fixture.
Commit: refactor(site): SiteStorageService takes ILocalDb instead of a connection string
Task 7: Extend SiteLocalDbSetup with the new DDL (not yet registered)
Classification: standard Estimated implement time: ~3 min Parallelizable with: none
Deliberately applies DDL without calling RegisterReplicated. The tables live in the
consolidated file but are not yet captured, so the bespoke replicator keeps working unchanged and
the tree stays green. Task 14 flips them on and deletes the bespoke path in one commit.
Files:
- Modify:
src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs:43-57 - Test:
tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbWiringTests.cs
Step 1: Extend OnReady
using (var connection = db.CreateConnection())
{
OperationTrackingSchema.Apply(connection);
SiteEventLogSchema.Apply(connection);
SiteStorageSchema.Apply(connection);
StoreAndForwardSchema.Apply(connection);
}
db.RegisterReplicated("OperationTracking");
db.RegisterReplicated("site_events");
// Phase 2 tables are created here but NOT yet registered — the bespoke replicator
// still owns them until the Task 14 cutover. Registering them now would double-
// replicate every row (harmless, both paths upsert, but it would mask a defect in
// either one).
SiteLocalDbLegacyMigrator.Migrate(db, config);
Step 2: Pin it
Add a test asserting all 4 Phase-1-and-2 table sets exist after OnReady, and that
db.ReplicatedTables contains exactly OperationTracking and site_events — the "not yet"
is the assertion that matters.
Commit: feat(localdb): create config + sf_messages tables in the consolidated DB
Task 8: Extend the legacy migrator for sf_messages
Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 9
Read SiteLocalDbLegacyMigrator.cs in full first. MigrateTracking (:116-161) is the closest
reference — like sf_messages, it has a native TEXT PK and needs no id synthesis.
Files:
- Modify:
src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbLegacyMigrator.cs:57-66 - Test:
tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbLegacyMigratorTests.cs
Step 1: Write the failing tests
Three, minimum:
[Fact] public void SfMessages_AreCopiedFromTheLegacyFile() { }
[Fact] public void SfMessages_Migration_IsIdempotent_WhenRerunAfterACrashBeforeRename() { }
[Fact] public void MigratedSfMessages_EnterTheOplog_SoTheyActuallyReplicate() { }
The third is the one that catches the ordering defect. Phase 1's equivalent asserts on
__localdb_oplog directly — copy that assertion. The oplog table is __localdb_oplog
(LocalDbSchema.cs:19), not zb_oplog.
Step 2: Add ResolveStoreAndForwardPath + MigrateStoreAndForward
private const string DefaultStoreAndForwardPath = "./data/store-and-forward.db";
internal static string ResolveStoreAndForwardPath(IConfiguration config) =>
config["ScadaBridge:StoreAndForward:SqliteDbPath"] ?? DefaultStoreAndForwardPath;
Migrate all 16 columns with INSERT OR IGNORE. The PK is native TEXT — no id synthesis, unlike
MigrateEvents.
Step 3: Wire into Migrate — add the call alongside the existing two.
Commit: feat(localdb): migrate legacy store-and-forward.db into the consolidated DB
Task 9: Extend the legacy migrator for the 9 config tables
Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 8
Files:
- Modify:
src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbLegacyMigrator.cs - Test:
tests/ZB.MOM.WW.ScadaBridge.Host.Tests/SiteLocalDbLegacyMigratorTests.cs
All 9 tables migrate from one legacy file (ScadaBridge:Database:SiteDbPath), so this is one
MigrateSiteStorage that opens the legacy DB once and copies table by table inside a single
transaction, then renames once.
Two tables are deliberately skipped: notification_lists and smtp_configurations. They are
purged on every deploy and are permanently empty by design (site write paths were removed
2026-07-10). Migrating them would resurrect plaintext SMTP passwords from a pre-fix legacy file into
a replicated table. Add this as a comment and as a test:
[Fact]
public void Migration_DoesNotCopyNotificationOrSmtpRows_EvenWhenTheLegacyFileHasThem()
{
// These carry plaintext SMTP passwords in pre-2026-07-10 files. They are purged
// on every deploy and must never enter a replicated table.
}
Step: Verify all 9 tables' columns match between SiteStorageSchema and the migrator's
INSERT lists. A column-count mismatch here fails at runtime on a real deployment and nowhere else.
Commit: feat(localdb): migrate legacy scadabridge.db config tables into the consolidated DB
Task 10: Port the S&F replication test intents as CDC specs
Classification: standard Estimated implement time: ~5 min Parallelizable with: Task 11
Specifications first, deletion second. These tests encode behaviour the replacement must still satisfy. The gate document named two files; the real specification is five.
Files:
- Read (as spec):
tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationServiceTests.cs - Read (as spec):
tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardReplicationTests.cs - Create:
tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbStoreAndForwardConvergenceTests.cs
Port these intents (not the mechanics) as two-node convergence assertions, using Phase 1's
LocalDbSitePairConvergenceTests.cs as the harness reference:
| Intent | Assertion under CDC |
|---|---|
| Add materialises on the peer | enqueue on A → row present on B |
| Remove deletes on the peer | deliver on A → row absent on B |
| Park sets status Parked on the peer | park on A → B's row has status Parked |
| Requeue resets status + RetryCount | requeue on A → B shows Pending, RetryCount 0 |
| Add-then-Remove never converges to present | the portable form of the strict-ordering test |
| Apply Add twice is idempotent, newest wins | LWW gives this |
| Park when Add was lost still materialises | upsert semantics |
Explicitly do not port: ReplicationServiceTests #10 (200 interleaved ops dispatched
synchronously in strict issue order, 400 observed). That asserts the mechanism — inline
fire-and-forget dispatch — not the outcome. CDC capture is asynchronous and batched by construction.
The portable intent is row 5 above. Record this in the test file as a comment so a future reader
does not think it was dropped by accident.
Commit: test(localdb): port store-and-forward replication intents as CDC convergence specs
Task 11: Port the resync + directional-authority tests
Classification: standard Estimated implement time: ~5 min Parallelizable with: Task 10
Files:
- Read (as spec):
tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/SiteReplicationActorTests.cs - Read (as spec):
tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/Cluster/SfBufferResyncPredicateTests.cs - Create:
tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbConfigConvergenceTests.cs
SfBufferResyncPredicateTests is the most valuable test in the set — the N1 Critical
regression. Per D2 it must be re-expressed, not ported literally:
[Fact]
public async Task ANodeWithNewerLocalRows_KeepsThem_WhenAPeerSnapshotArrives()
{
// N1 Critical (was: SfBufferResyncPredicateTests). The bespoke replicator needed a
// directional guard because ReplaceAllAsync was a destructive DELETE-then-INSERT,
// so a wrong-direction resync wiped a live buffer.
//
// LocalDb's snapshot resync merges per row under LWW and never deletes
// (SnapshotApplier has no DELETE; LwwApplier.cs:69-78 discards an incoming row
// whose HLC is lower than the local row's). The failure mode is structurally
// impossible rather than guarded against — so this asserts convergence, NOT
// directional authority. There is no active/standby asymmetry left to enforce.
}
Also do not port: actor tests #1–#5 (notify-and-fetch config fetch/retry/supersede). Per D1 the fetch path is deleted, so these describe code that will not exist. Replace with a single assertion that a config deploy on A converges to B without B making any HTTP call — assert on a fetcher test double that records zero invocations. That is the positive proof that notify-and-fetch is gone.
Commit: test(localdb): port resync + config replication intents as CDC convergence specs
Task 12: Re-home the SMTP purge off the replication path
Classification: high-risk Estimated implement time: ~4 min Parallelizable with: none
Per D3. This must land before any deletion so the purge never lapses, even for one commit.
Files:
- Modify:
src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Actors/DeploymentManagerActor.cs(artifact-apply path) - Test:
tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/DeploymentManagerActorTests.cs
Step 1: Write the failing test
[Fact]
public async Task ApplyingArtifacts_PurgesCentralOnlyNotificationConfig_OnTheActiveNode()
{
// Security cleanup: notification_lists and smtp_configurations can hold plaintext
// SMTP passwords from pre-2026-07-10 deployments. This ran on the standby via
// SiteReplicationActor.HandleApplyArtifacts; with that actor deleted the ACTIVE
// node must do it, and CDC replicates the deletes as ordinary tombstones.
}
Step 2: Call PurgeCentralOnlyNotificationConfigAsync from the active node's artifact-apply path
The existing call site is SiteReplicationActor.HandleApplyArtifacts. Find where
DeploymentManagerActor applies artifacts locally (the same path that reaches
_replicationActor?.Tell(new ReplicateArtifacts(command)) at :1952) and call it there.
Step 3: Verify the deletes replicate — extend the Task 11 convergence suite with a case where A purges a seeded row and B's copy disappears.
Commit: fix(site): purge central-only notification config on the active node
Task 13: Delete the notify-and-fetch config path
Classification: high-risk Estimated implement time: ~4 min Parallelizable with: none
Per D1.
Files:
- Modify:
src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Persistence/SiteStorageService.cs:301-336 - Modify:
src/ZB.MOM.WW.ScadaBridge.SiteRuntime/SiteRuntimeOptions.cs:66-68 - Modify:
src/ZB.MOM.WW.ScadaBridge.SiteRuntime/SiteRuntimeOptionsValidator.cs:56-58
Step 1: Delete StoreDeployedConfigIfNewerAsync — its only caller is
SiteReplicationActor.HandleApplyConfigDeploy, deleted in Task 15. The unguarded
StoreDeployedConfigAsync (:241-271) is the active node's path and stays.
Step 2: Check ConfigFetchRetryCount before removing it
IDeploymentConfigFetcher is used by both the standby replication path and the active
singleton's RefreshDeploymentCommand path. Only the first is being deleted.
grep -rn "ConfigFetchRetryCount" src/ tests/
If the active path reads it, keep the option and its validator rule and note that in the commit message. If nothing outside the deleted actor reads it, remove both.
Commit: refactor(site): delete notify-and-fetch config replication — CDC ships the row
Task 14: Register the Phase 2 tables and delete ReplicationService
Classification: high-risk Estimated implement time: ~5 min Parallelizable with: none
The cutover. Registration and bespoke-path deletion land in one commit so the two mechanisms never both run.
Files:
- Modify:
src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbSetup.cs - Delete:
src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ReplicationService.cs - Modify:
src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardService.cs:39,243,654,805,836,872,1120,1146 - Modify:
src/ZB.MOM.WW.ScadaBridge.StoreAndForward/ServiceCollectionExtensions.cs:63-68 - Modify:
src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardStorage.cs:284-306 - Delete:
tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationServiceTests.cs - Delete:
tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/StoreAndForwardReplicationTests.cs - Delete:
tests/ZB.MOM.WW.ScadaBridge.StoreAndForward.Tests/ReplicationWireSerializationPinTests.cs
Step 1: Register all 11 tables in SiteLocalDbSetup.OnReady, replacing the Task 7 comment.
Skip notification_lists and smtp_configurations? No — register them. They must replicate
their deletes (Task 12). Register all 9 config tables plus sf_messages.
Step 2: Delete ReplicationService and remove the ctor parameter + all 6 emission call sites
from StoreAndForwardService.
Step 3: Delete ReplaceAllAsync (:284-306). Per D2 the library's resync merges; a destructive
delete-all must not survive into a replicated table, where CDC would capture the mass-delete and
ship it to the peer.
Step 4: Check GetAllMessagesAsync's Truncated flag — it exists only for the resync path. If
nothing else reads it, simplify the return type.
Step 5: CompositionRootTests.cs:474 asserts ReplicationService is a registered site
singleton. Delete that assertion.
Step 6: Full S&F + Host suites — expect PASS
Commit: feat(localdb)!: replicate sf_messages via CDC, delete ReplicationService
Task 15: Delete SiteReplicationActor and its messages
Classification: high-risk Estimated implement time: ~4 min Parallelizable with: none
Files:
- Delete:
src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Actors/SiteReplicationActor.cs - Delete:
src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Messages/ReplicationMessages.cs - Delete:
tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/Actors/SiteReplicationActorTests.cs - Delete:
tests/ZB.MOM.WW.ScadaBridge.SiteRuntime.Tests/ResyncWireSerializationPinTests.cs - Delete:
tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/Cluster/SfBufferResyncPredicateTests.cs
The whole ReplicationMessages.cs file becomes dead: ReplicateConfigDeploy/Remove/SetEnabled,
ReplicateArtifacts, ReplicateStoreAndForward, ApplyConfigDeploy/Remove/SetEnabled,
ApplyArtifacts, ApplyStoreAndForward, plus the resync records RequestSfBufferResync,
SfBufferSnapshot, SfBufferSnapshotChunk, SfBufferResyncAck.
Do not delete ActiveNodeEvaluator (Communication/ClusterState/ActiveNodeEvaluator.cs) — the
S&F delivery gate still needs it. Only its doc comment at :14 mentions replication; update
the comment, keep the type.
Confirm the deletions are complete before committing:
grep -rn "SiteReplicationActor\|ReplicateStoreAndForward\|SfBufferSnapshot" src/ tests/
Expected: no matches.
Commit: feat(localdb)!: delete SiteReplicationActor — CDC replaces hand-shipped ops
Task 16: Clean up DeploymentManagerActor and AkkaHostedService
Classification: standard Estimated implement time: ~4 min Parallelizable with: none
Files:
- Modify:
src/ZB.MOM.WW.ScadaBridge.SiteRuntime/Actors/DeploymentManagerActor.cs:55,161-190,794,970,1004,1102,1952 - Modify:
src/ZB.MOM.WW.ScadaBridge.Host/Actors/AkkaHostedService.cs:770-796,806-812,1056
Step 1: Remove the IActorRef? replicationActor = null ctor parameter (:184), the
_replicationActor field (:55), and all 5 _replicationActor?.Tell(...) sites.
The parameter is optional and positional — removing it silently shifts every argument after it
(healthCollector, serviceProvider, loggerFactory, configFetcher, ...). Update the
Props.Create at AkkaHostedService.cs:806-812 and every test constructing this actor. Compile
errors will not catch a positional shift between two same-typed optional parameters — check each
call site by hand.
Step 2: Delete replicationService / replicationLogger resolution (:770-772), the actor
creation (:785-789), the handler wiring (:792-796), and the actor-inventory comment at :1056.
Step 3: activeNodeCheck (:781-783) is still used by SiteCommunicationActor (:816-821).
Keep it.
Step 4: Full Host + SiteRuntime + IntegrationTests suites.
Commit: refactor(site): drop the replication actor from the deployment + host wiring
Task 17: Config-key cleanup
Classification: standard Estimated implement time: ~4 min Parallelizable with: none
Files:
- Modify:
src/ZB.MOM.WW.ScadaBridge.Host/StartupValidator.cs:113-115 - Modify:
src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardOptions.cs:12 - Modify:
src/ZB.MOM.WW.ScadaBridge.StoreAndForward/StoreAndForwardOptionsValidator.cs:20 - Modify: 9 ×
appsettings.Site.json(Host template + 6docker/site-*+ 2docker-env2/site-x-*)
| Key | Action |
|---|---|
ScadaBridge:Database:SiteDbPath |
→ migration-only. Relax the Require at StartupValidator.cs:113 or every site config keeps it mandatory forever. |
ScadaBridge:StoreAndForward:SqliteDbPath |
→ migration-only. Relax the non-empty rule at StoreAndForwardOptionsValidator.cs:20. |
ScadaBridge:StoreAndForward:ReplicationEnabled |
→ fully dead. Delete the property and all 9 config entries. |
ParkedOperationRelayTests.cs:39 and ParkedMessageHandlerActorTests.cs:34 set
ReplicationEnabled = false — they only need the flag to exist, so remove those lines.
StoreAndForwardOptionsTests.cs:14,26,33 assert its default; delete those assertions.
Leave the path keys present in configs with a comment — they are read by the legacy migrator and removing them would strand un-migrated data on a node that has not yet started once.
Commit: chore(config): retire ReplicationEnabled, make legacy db paths migration-only
Task 18: Two-node convergence suite for the Phase 2 tables
Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 19
Extend Phase 1's LocalDbSitePairConvergenceTests.cs harness (real loopback Kestrel h2c + the real
LocalDbSyncAuthInterceptor, initialized via SiteLocalDbSetup.OnReady).
Files:
- Modify:
tests/ZB.MOM.WW.ScadaBridge.IntegrationTests/LocalDbSitePairConvergenceTests.cs
Scenarios, beyond Tasks 10–11:
- A config deploy on A converges to B, and B's
deployed_configurationsrow matches byte for byte. RemoveDeployedConfigAsync's 3-table cascade converges —static_attribute_overridesandnative_alarm_stateorphans do not survive on B. There are no foreign keys (SiteStorageServiceDDL has none), so this is three independent delete streams that LWW may reorder. This is the most likely real defect in the whole plan.- A burst of
native_alarm_stateupserts converges and the oplog drains. - Node B restarts and catches up via snapshot resync without losing rows written on B while down.
Verify the suite is non-vacuous. Phase 1 did this by mismatching the two API keys and confirming all scenarios went red. Do the same and record the result in the commit message. A convergence suite that cannot fail proves nothing.
Commit: test(localdb): two-node convergence for config tables + sf_messages
Task 19: Rig configuration
Classification: small Estimated implement time: ~3 min Parallelizable with: Task 18
Files:
- Modify:
docker/site-a-node-a/appsettings.Site.json,docker/site-a-node-b/appsettings.Site.json
Set MaxOplogRows and MaxOplogAge from Task 1's measured numbers — not from a guess. Keep
site-b/site-c unreplicated so default-OFF stays proven side by side on one rig.
Commit: chore(docker): size the site-a oplog caps from the phase 2 soak
Task 20: Live gate on the docker rig
Classification: high-risk Estimated implement time: ~30 min wall-clock Parallelizable with: none
Files:
- Create:
docs/plans/2026-07-19-localdb-phase2-live-gate.md
cd ~/Desktop/ScadaBridge && bash docker/deploy.sh
Evidence required — each must be captured, not asserted:
- Migration ran.
store-and-forward.db.migratedandscadabridge.db.migratedexist on both site-a nodes; row counts in the consolidated DB match the legacy files. - No SMTP/notification rows migrated —
SELECT COUNT(*) FROM smtp_configurationsis 0. - Config converges. Deploy an instance to site-a; both nodes'
deployed_configurationsrows are byte-identical with identical__localdb_row_versionand originating node id. - The standby made no HTTP fetch — nothing in its log resembling a config fetch. This is the positive proof that notify-and-fetch is gone.
- S&F converges. Enqueue against a dead target; the parked row appears on the peer.
- Site failover. Flip the site-a pair and confirm buffered messages survive and deliver
exactly once after the flip.
docker/failover-drill.shtargets CENTRAL nodes and does not exercise this — run the site flip directly. - Cascade delete. Remove a deployed config; confirm no orphan
static_attribute_overridesornative_alarm_staterows on the peer. - Zero dead letters, oplog drained,
localdb_*scraped from/metrics. - Both nodes stopped and started together (D5) — confirm a clean rejoin.
If any check fails, stop and report. Do not proceed to Task 21.
Commit: docs(localdb): phase 2 live gate evidence
Task 21: Documentation truth pass
Classification: standard Estimated implement time: ~5 min Parallelizable with: none
Files:
- Modify:
CLAUDE.md(ScadaBridge),~/Desktop/scadaproj/CLAUDE.md(LocalDb component row) - Modify:
docs/requirements/Component-StoreAndForward.md:83 - Modify:
docs/components/StoreAndForward.md,SiteRuntime.md,Host.md - Modify:
docs/known-issues/2026-06-26-deploy-config-exceeds-akka-frame-size.md - Modify:
docs/deployment/runbook
Specific corrections:
Component-StoreAndForward.md:83is a long normative paragraph specifying the whole chunked resync protocol and the N5 race. Rewrite for CDC — state the new duplicate-delivery bound explicitly rather than deleting the discussion.- The frame-size known-issue is resolved by this change (D1). Mark it resolved and say why, rather than leaving a live known-issue describing deleted code.
- Add to the deployment runbook (D5): a site pair must be stopped and started together. Rolling
one node at a time is no longer supported, because the legacy
SfBufferSnapshotcompatibility handler is gone. - Update the scadaproj LocalDb component row: Phase 2 complete, what replicates now, what
SiteReplicationActorused to do.
Commit: docs(localdb): phase 2 truth pass across both repos
Definition of done
- Solution builds with 0 warnings (
TreatWarningsAsErrorsis on) - Every suite green: Host, CentralUI, Commons, SiteRuntime, AuditLog, Communication, StoreAndForward, HealthMonitoring, IntegrationTests, SiteEventLogging
grep -rn "SiteReplicationActor\|ReplicationService\|SfBufferSnapshot" src/→ no matches- Live gate PASS with all 9 evidence items captured
- Convergence suite verified non-vacuous
- Both
CLAUDE.mdfiles updated