Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts
11 KiB
LocalDb Phase 2 — resume state (updated 2026-07-20, pre-compaction #2)
Scratch handoff for picking the plan back up. Delete once Phase 2 lands.
Plan: 2026-07-19-localdb-adoption-phase2.md
(execute with superpowers-extended-cc:executing-plans; task state in
…-phase2.md.tasks.json, which carries a per-task deviation note)
Branch: feat/localdb-phase2 (from feat/localdb-phase1, which is not merged)
Tree: clean @ 0ad11d6b. Build 0 warnings. 25 commits ahead of origin/feat/localdb-phase2.
Where we are
Tasks 1–16 are DONE. The cutover has landed. Next up: Task 17.
| Task | State |
|---|---|
| 1–6 | done (waves 0–2) — see the git log |
7 SiteLocalDbSetup DDL |
done |
8 sf_messages migrator |
done |
| 9 config-table migrator | done |
| 10 S&F CDC convergence specs | done |
| 11 resync + config convergence specs | done |
| 12 active-node purge pin | done |
| 13 guarded-write scope check | done |
| 14 + 15 + 16 the CUTOVER | done — landed as ONE commit 037798b3 |
| 17–21 | not started |
Verification at the cutover boundary: build 0 warnings; SiteRuntime 512, StoreAndForward 130, Host 330, AuditLog 355, ExternalSystemGateway 142, HealthMonitoring 97, LocalDb integration 16 — all pass, 0 failures. (SiteRuntime/StoreAndForward fell from 533/153 as 6 obsolete test files were deleted.)
The bespoke replicators are gone
SiteReplicationActor, ReplicationMessages.cs, ReplicationService and
StoreAndForwardStorage.ReplaceAllAsync are deleted. Ten tables now replicate via LocalDb CDC:
the Phase 1 pair plus sf_messages and the 7 site config tables.
notification_lists and smtp_configurations are created but never registered — permanently
empty by design, and registering them would open a standing replication channel whose only
historical payload was plaintext SMTP passwords. That decision is pinned by its own
security-named test and stated in OnReady.
Next: Task 17 (config-key cleanup)
Partly required, not optional: SiteRuntimeOptions.cs:66 still documents
ConfigFetchRetryCount as "Consumed by SiteReplicationActor", which no longer exists.
Per the plan: relax StartupValidator.cs:113 (SiteDbPath) and
StoreAndForwardOptionsValidator.cs:19-21 (SqliteDbPath) to migration-only; delete
ReplicationEnabled and ConfigFetchRetryCount outright with their validator rules and config
entries; leave the two path keys present in configs with a comment — the legacy migrator reads
them, and removing them would strand un-migrated data on a node that has not yet started once.
10 appsettings.Site.json files, including deploy/wonder-app-vd03/.
Then 18 (two-node convergence suite), 19 (rig config — MaxBatchSize = 16), 20 (live gate),
21 (docs truth pass).
Findings that change the plan (carry these forward)
From Task 1 (unchanged)
- D6's premise is wrong, its risk is real. Largest known production
config_jsonis ~60–70 KB, not >128 KB → no stop condition. ButMaxBatchSize=500× 70 KB ≈ 35 MB vs a 4 MB gRPC cap → Task 19 must setLocalDb:Replication:MaxBatchSize = 16. The one firmly evidence-backed number. - D4 is wrong.
native_alarm_stateis bounded by per-SourceReferencecoalescing on a 100 ms flush (NativeAlarmActor.cs:473-502). sf_messageshas a hard 50 rows/sec ceiling —SweepBatchLimit(500) ÷RetryTimerInterval(10 s).- Task 1 step 7's stop condition is weaker than written. Exceeding the oplog caps prunes and
flags
needs_snapshot→ graceful resync, not data loss.
Measured (six 60 s intervals, copy-based snap()): sf_messages 0.80 rows/sec insert, 0/sec
retry-UPDATE, oplog 0, max payload_json 76 B, max config_json 721 B (rig — not representative),
zero SQLite errors in 30 min.
From waves 1–2
- LATENT PHASE-1 DEFECT, fixed in-app only. The LocalDb library never creates the parent
directory of
LocalDb:Path, andSqliteLocalDbopens the file eagerly in its constructor, so a missing directory is a hard boot failure (SQLite Error 14). The default site config uses the relative./data/site-localdb.db; the docker rig escapes only because its volume mount creates/app/data. Fixed viaSiteLocalDbDirectory.Ensure(config)beforeAddZbLocalDb, pinned byHost.Tests/SiteLocalDbDirectoryTests. OPEN DECISION FOR THE USER: the real fix arguably belongs in theZB.MOM.WW.LocalDblibrary — every other consumer still has this gap. Not done here: a library change means a version bump and is outside this plan's scope. SiteStorageService.CreateConnection()contract flipped — returns an already-open connection; callingOpen/OpenAsyncon one throws.- LocalDb has NO in-memory mode. Every
Mode=Memory;Cache=Sharedfixture became a real temp file. - A store constructor change's test blast radius is ~7× the plan's estimate (Task 5: 40 files, 7 projects).
tests/ZB.MOM.WW.ScadaBridge.TestSupportholdsTestLocalDb(dispose before deleting — the master connection anchors the WAL). Use it; don't copy fixtures.- Where behaviour moved owner, retarget tests — don't delete them with the code. WAL is genuinely LocalDb's; directory creation is not. Both directions occurred.
From waves 3–5 (new)
- PLAN DEFECT — Tasks 14/15/16 cannot compile separately.
SiteReplicationActortakes aReplicationServiceand callsReplaceAllAsync;DeploymentManagerActorTells types fromReplicationMessages.cs;AkkaHostedServiceconstructs the actor. Landed as one commit, which also strengthens Task 14's own invariant (never both mechanisms, never neither). - The legacy-copy column intersection. The plan said "migrate all 16 columns"; that would be a
data-loss bug. An older legacy file lacks
execution_id/parent_execution_id/last_attempt_at_ms, naming a missing column throwsno such column, and the existing reader treats that as an unrecognised shape and silently discards every row. The copy intersects the legacy and current column sets, with a required-column (PK) guard so the tolerance cannot degrade into copying NULL-keyed rows. MigratorColumnLists_MatchTheLiveSchema(added beyond the plan). A column-list mismatch is invisible at runtime in BOTH directions — a typo is silently dropped by the intersection, a missing column silently leaves data behind.SiteStorageTables/LegacyTableareinternalfor it.- Absence assertions need a positive control.
MessageAddedThenRemoved_NeverConvergesToPresentPASSED withsf_messagesunregistered entirely — an absent row is also what a pair replicating nothing looks like. Fixed with a control row that must converge in the same window. - The positional-argument hazard was REAL. Removing
DeploymentManagerActor's optionalIActorRef? replicationActorshifted 6 trailing optionals and 4 test call sites bound the wrong arguments, some with no compile error.Props.Createbuilds an expression tree, which rejects OUT-OF-POSITION named arguments — so named args work only if in position, and the rest must be padded positionally. ReplaceAllAsyncwas unsafe to keep, not merely unused. A mass DELETE on a now-replicated table would be captured and shipped to the peer.LocalDbSitePairHarnessis the shared two-node fixture (Phase 1 + both Phase 2 suites; Task 18 makes four). ItsPhase2ReplicatedTableslist is literal, not derived from production code, so a registration that drifts fails the tests instead of agreeing with itself.- Ordinal sort gotcha:
_(0x5F) sorts beforeb(0x62), sodata_connection_definitionsprecedesdatabase_connections. - Never
git checkout <file>to undo a non-vacuity experiment — it reverts to HEAD and takes in-progress work with it (cost one redo of Task 8). Usecpbackups.
Rig operational gotchas (hard-won; not obvious from the repo)
ExternalSystem.Calldoes NOT buffer to store-and-forward — onlyExternalSystem.CachedCalldoes. The single most important detail for generating S&F load.- Never query the bind-mounted SQLite files from the macOS host while containers run. Use the
plan's
snaphelper, ordocker run --rm -v <dir>:/d alpine/sqlite3 …. - No
curlin theaspnet:10.0image, and metrics port 8084 is not published. Use a sidecar:docker run --rm --network container:scadabridge-site-a-a curlimages/curl:latest -s localhost:8084/metrics - Seeded template 4 (
Motor Controller) is not deployable — 34 validation errors. The soak used a purpose-builtSoakGeneratortemplate. - CLI:
template script updaterequires--nameand--trigger-typeeven to change only--code. In zsh, don't put auth flags in an unquoted variable. - Rig auth:
--url http://localhost:9000 --username multi-role --password password.LdapGroupMappingsroles must be canonicalDesigner/Deployer; central caches them at startup.
Rig state as left
Reseeded and healthy; both site-a nodes restarted after the poisoning incident.
Needs cleanup before the Task 20 live gate:
ExternalSystemDefinitionsid 1 is still repointed tohttp://127.0.0.1:9— restore tohttp://scadabridge-restapi:5200.- Template
SoakGenerator(id 2021) and instancessoakgen-1..4(ids 5–8) are still deployed on site-a and still generating load.
Commits (this branch, newest last)
| Commit | What |
|---|---|
25463d52 … d512572d |
plan review, soak, root-cause + hardening, first resume doc |
12eb97e0 |
Task 2 — gate CLOSED |
9e3239c5 |
Task 3 — StoreAndForwardSchema.Apply |
ac5eb12c |
Task 4 — SiteStorageSchema.Apply |
f2efeb37 |
Tasks 5+6 — both stores take ILocalDb; latent directory defect fixed; TestSupport added |
fefbbb31 |
resume state through wave 2 |
f8aa02e2 |
Task 7 — Phase 2 DDL in the consolidated DB (not yet registered) |
bdc0dffe |
Task 8 — migrate legacy store-and-forward.db |
5ddc7eed |
Task 9 — migrate legacy scadabridge.db config tables |
2bbe6631 |
Task 10 — S&F CDC convergence specs (+ LocalDbSitePairHarness extraction) |
c56bf4ae |
Task 11 — resync + config convergence specs |
79ce5161 |
Tasks 12+13 — purge pin; guarded-write scope |
037798b3 |
Tasks 14+15+16 — THE CUTOVER |
0ad11d6b |
task state for the cutover |
Two known flakes, do not chase: InstanceActorChildAttributeRaceTests (intermittent under
full-suite load) and ParentExecutionIdCorrelationTests (~91 s / times out on a cold MSSQL
fixture, ~1 s warm). Neither reproduced during waves 1–5.
Execution notes
- Waves ran sequentially in the main worktree, not as the plan's parallel worktree agents — the tasks were small enough that worktree setup + merge cost more than it saved, and sequential execution removes the git race the isolation rule exists to prevent.
- Subagents were used only in waves 1–2, for bulk mechanical fixture rewiring. Their reported results were re-verified independently — one agent's partial run had missed 2 real failures that a full-suite run caught. Re-run the suites yourself.
- Every new test was checked for non-vacuity by breaking the thing it claims to pin (removing
DDL, removing the
Migratecall, unregistering a table, commenting out the purge). This caught one genuinely vacuous test. Keep doing it.