docs(localdb): resume state through the cutover (tasks 1-16 done, task 17 next)

Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts
This commit is contained in:
Joseph Doherty
2026-07-20 04:22:30 -04:00
parent 0ad11d6b55
commit 3364145d63
+146 -120
View File
@@ -1,131 +1,154 @@
# LocalDb Phase 2 — resume state (written 2026-07-20, pre-compaction)
# LocalDb Phase 2 — resume state (updated 2026-07-20, pre-compaction #2)
Scratch handoff for picking the plan back up. Delete once Phase 2 lands.
**Plan:** [`2026-07-19-localdb-adoption-phase2.md`](2026-07-19-localdb-adoption-phase2.md)
(execute with `superpowers-extended-cc:executing-plans`; task state in
`…-phase2.md.tasks.json`)
**Branch:** `feat/localdb-phase2` (from `feat/localdb-phase1`, which is **not** merged/pushed)
**Tree:** clean @ `19ab0ac9`. Build 0 warnings. **Nothing pushed to origin yet.**
`…-phase2.md.tasks.json`, which carries a per-task deviation note)
**Branch:** `feat/localdb-phase2` (from `feat/localdb-phase1`, which is **not** merged)
**Tree:** clean @ `0ad11d6b`. Build 0 warnings. 25 commits ahead of `origin/feat/localdb-phase2`.
---
## Where we are
**Tasks 16 are DONE (waves 0, 1, 2). Next up: Wave 3 = Tasks 7, 8, 9.**
Task 1's gate is closed with verdict **PROCEED**; its earlier STOP was superseded (the
`disk I/O error` storm was **observer-induced** — host-side `sqlite3` reads of live
bind-mounted WAL files reset the WAL under the container, not a product defect; write-up in
[`docs/known-issues/2026-07-20-localdb-disk-io-error-under-load.md`](../known-issues/2026-07-20-localdb-disk-io-error-under-load.md)
§0/§11a).
**Tasks 116 are DONE. The cutover has landed. Next up: Task 17.**
| Task | State |
|---|---|
| 1 Rig soak | done — gate PROCEED |
| 2 Decision record | done — gate doc CLOSED, all §5 questions answered inline |
| 3 `StoreAndForwardSchema.Apply` | done |
| 4 `SiteStorageSchema.Apply` | done |
| 5 S&F store → `ILocalDb` | done |
| 6 `SiteStorageService``ILocalDb` | done |
| **721** | **not started, no code written** |
| 16 | done (waves 02) — see the git log |
| 7 `SiteLocalDbSetup` DDL | done |
| 8 `sf_messages` migrator | done |
| 9 config-table migrator | done |
| 10 S&F CDC convergence specs | done |
| 11 resync + config convergence specs | done |
| 12 active-node purge pin | done |
| 13 guarded-write scope check | done |
| **14 + 15 + 16 the CUTOVER** | **done — landed as ONE commit `037798b3`** |
| **1721** | **not started** |
**Verification at the wave-2 boundary:** full solution build 0 warnings; SiteRuntime 532,
Host 318, AuditLog 355, StoreAndForward 153, ExternalSystemGateway 142, HealthMonitoring 97
**1597 passed / 0 failed**.
**Verification at the cutover boundary:** build 0 warnings; SiteRuntime 512, StoreAndForward 130,
Host 330, AuditLog 355, ExternalSystemGateway 142, HealthMonitoring 97, LocalDb integration 16 —
**all pass, 0 failures.** (SiteRuntime/StoreAndForward fell from 533/153 as 6 obsolete test files
were deleted.)
### Task 1's measured numbers (clean re-run, 2026-07-20)
### The bespoke replicators are gone
Six consecutive 60 s intervals, copy-based `snap()` sampling on restarted nodes:
`sf_messages` **0.80 rows/sec** insert, **0/sec** retry-UPDATE, oplog 0, `native_alarm_state` 0,
max `payload_json` **76 B**, max `config_json` **721 B** (rig — *not* representative),
**zero** SQLite errors in 30 min, both site-a nodes converged at 3564 rows. Honest gap: 0.80/s
is ~1.6 % of the 50/s ceiling and the retry path was never exercised — fine only because that
ceiling is structural, not empirical.
`SiteReplicationActor`, `ReplicationMessages.cs`, `ReplicationService` and
`StoreAndForwardStorage.ReplaceAllAsync` are deleted. Ten tables now replicate via LocalDb CDC:
the Phase 1 pair plus `sf_messages` and the 7 site config tables.
### Next: Wave 3 (Tasks 7, 8, 9)
`notification_lists` and `smtp_configurations` are created but **never registered** — permanently
empty by design, and registering them would open a standing replication channel whose only
historical payload was plaintext SMTP passwords. That decision is pinned by its own
security-named test and stated in `OnReady`.
Task 7 extends `SiteLocalDbSetup.OnReady` with the new DDL (**not** registered yet — registration
is Task 14, the cutover). Tasks 8 and 9 are the legacy migrators and are both **high-risk**.
Ordering inside `OnReady` is load-bearing: **DDL → `RegisterReplicated` → migrate**. Rows written
before registration are invisible to the peer forever, silently.
---
## Next: Task 17 (config-key cleanup)
Partly **required**, not optional: `SiteRuntimeOptions.cs:66` still documents
`ConfigFetchRetryCount` as "Consumed by `SiteReplicationActor`", which no longer exists.
Per the plan: relax `StartupValidator.cs:113` (`SiteDbPath`) and
`StoreAndForwardOptionsValidator.cs:19-21` (`SqliteDbPath`) to migration-only; delete
`ReplicationEnabled` and `ConfigFetchRetryCount` outright with their validator rules and config
entries; **leave the two path keys present in configs** with a comment — the legacy migrator reads
them, and removing them would strand un-migrated data on a node that has not yet started once.
**10** `appsettings.Site.json` files, including `deploy/wonder-app-vd03/`.
Then 18 (two-node convergence suite), 19 (rig config — `MaxBatchSize = 16`), 20 (live gate),
21 (docs truth pass).
---
## Findings that change the plan (carry these forward)
From [`2026-07-19-localdb-phase2-soak.md`](2026-07-19-localdb-phase2-soak.md):
### From Task 1 (unchanged)
- **D6's premise is wrong, its risk is real.** The plan says `config_json` is "documented to
exceed 128 KB per row." The source actually says the *escaped Akka envelope* blew the 128 KB
frame and the raw JSON "only needs to be ~60-70 KB". So the largest known production row is
~6070 KB → **no stop condition**. But `MaxBatchSize=500` is unsafe (500 × 70 KB ≈ 35 MB vs a
4 MB gRPC cap) → **Task 19 must set `LocalDb:Replication:MaxBatchSize = 16`.** This is the one
firmly evidence-backed number.
- **D4 is wrong.** `native_alarm_state` is not "unbounded by design" — `NativeAlarmActor`
coalesces per `SourceReference` on a 100 ms flush (`NativeAlarmActor.cs:473-502`), so
worst case is `distinct_source_refs × 10/sec`.
- **`sf_messages` has a hard 50 rows/sec ceiling** — `SweepBatchLimit` (500) ÷
`RetryTimerInterval` (10 s). A ceiling, not an estimate.
- **Task 1 step 7's stop condition is much weaker than written.** Exceeding the oplog caps
prunes and flags `needs_snapshot` → graceful snapshot resync (`OplogStore.cs:109-138`,
`MaintenanceBackgroundService.cs:57`). Not data loss, not a wedge.
- **D6's premise is wrong, its risk is real.** Largest known production `config_json` is ~6070 KB,
not >128 KB → no stop condition. But `MaxBatchSize=500` × 70 KB ≈ 35 MB vs a 4 MB gRPC cap →
**Task 19 must set `LocalDb:Replication:MaxBatchSize = 16`.** The one firmly evidence-backed number.
- **D4 is wrong.** `native_alarm_state` is bounded by per-`SourceReference` coalescing on a 100 ms
flush (`NativeAlarmActor.cs:473-502`).
- **`sf_messages` has a hard 50 rows/sec ceiling** — `SweepBatchLimit` (500) ÷ `RetryTimerInterval` (10 s).
- **Task 1 step 7's stop condition is weaker than written.** Exceeding the oplog caps prunes and
flags `needs_snapshot` → graceful resync, not data loss.
The plan's own Task-1 method also needed four corrections (metrics sidecar, safe sampling,
no alarm generator on this rig, template 4 undeployable) — all recorded in the soak doc §1.
Measured (six 60 s intervals, copy-based `snap()`): `sf_messages` **0.80 rows/sec** insert, **0/sec**
retry-UPDATE, oplog 0, max `payload_json` 76 B, max `config_json` 721 B (rig — *not* representative),
zero SQLite errors in 30 min.
### Found during waves 12 (new)
### From waves 12
- **LATENT PHASE-1 DEFECT, found and fixed in Task 6.** The plan assumed directory creation
moved to LocalDb along with file ownership. **It did not.** The LocalDb library never creates
the parent directory, and `SqliteLocalDb` opens the file **eagerly in its constructor**, so a
missing directory is a **hard boot failure** (`SQLite Error 14: unable to open database file`),
not a degraded start. The default site config uses the **relative** path
`./data/site-localdb.db`, so any site node without a pre-existing `data/` fails to boot; the
docker rig escapes only because its volume mount creates `/app/data`. Latent since Phase 1 made
`LocalDb:Path` required. Fixed in ScadaBridge via `SiteLocalDbDirectory.Ensure(config)` called
before `AddZbLocalDb`, pinned by `Host.Tests/SiteLocalDbDirectoryTests`.
**OPEN DECISION FOR THE USER:** the real fix arguably belongs in the `ZB.MOM.WW.LocalDb`
library (it owns the file) — **every other LocalDb consumer still has this gap.** Not done
here because a library change means a version bump and is outside this plan's scope.
- **Silent contract change on `SiteStorageService.CreateConnection()`.** It used to return an
**unopened** connection that callers opened; it now returns an **already-open** LocalDb
connection, and calling `Open`/`OpenAsync` on one **throws**. `SiteExternalSystemRepository`
dropped 5 `OpenAsync` calls. Anything new that consumes this member must not open it.
- **LocalDb has NO in-memory mode.** `LocalDbOptions.Path` is a filesystem path, so every
fixture that used `Mode=Memory;Cache=Shared` had to move to a real temp file. Expect this for
any remaining store that gets rewired.
- **Test fallout from a store constructor change is ~7× what the plan estimates.** Task 5's plan
named "fixtures" in one project; the real blast radius was **40 files across 7 test projects**.
Budget accordingly for Tasks 14/15.
- **`tests/ZB.MOM.WW.ScadaBridge.TestSupport`** is a NEW shared, non-test library holding
`TestLocalDb` (`Create`/`CreateTemp`/`DeleteFiles`; dispose **before** deleting — the master
connection anchors the WAL). Use it rather than copying fixtures. Wired into AuditLog,
ExternalSystemGateway, HealthMonitoring, IntegrationTests, SiteRuntime, StoreAndForward tests.
- **Where behaviour moved owner, retarget tests — do not delete them with the code.** WAL genuinely
is LocalDb's job (test retargeted in place); directory creation is **not** (test moved to
`Host.Tests`). Both directions occurred, so check which owner actually holds the guarantee.
- **LATENT PHASE-1 DEFECT, fixed in-app only.** The LocalDb library never creates the parent
directory of `LocalDb:Path`, and `SqliteLocalDb` opens the file **eagerly in its constructor**, so
a missing directory is a **hard boot failure** (`SQLite Error 14`). The default site config uses
the **relative** `./data/site-localdb.db`; the docker rig escapes only because its volume mount
creates `/app/data`. Fixed via `SiteLocalDbDirectory.Ensure(config)` before `AddZbLocalDb`, pinned
by `Host.Tests/SiteLocalDbDirectoryTests`.
**OPEN DECISION FOR THE USER:** the real fix arguably belongs in the `ZB.MOM.WW.LocalDb` library
**every other consumer still has this gap.** Not done here: a library change means a version
bump and is outside this plan's scope.
- **`SiteStorageService.CreateConnection()` contract flipped** — returns an **already-open**
connection; calling `Open`/`OpenAsync` on one **throws**.
- **LocalDb has NO in-memory mode.** Every `Mode=Memory;Cache=Shared` fixture became a real temp file.
- **A store constructor change's test blast radius is ~7× the plan's estimate** (Task 5: 40 files,
7 projects).
- **`tests/ZB.MOM.WW.ScadaBridge.TestSupport`** holds `TestLocalDb` (dispose **before** deleting —
the master connection anchors the WAL). Use it; don't copy fixtures.
- **Where behaviour moved owner, retarget tests — don't delete them with the code.** WAL is genuinely
LocalDb's; directory creation is **not**. Both directions occurred.
### From waves 35 (new)
- **PLAN DEFECT — Tasks 14/15/16 cannot compile separately.** `SiteReplicationActor` takes a
`ReplicationService` and calls `ReplaceAllAsync`; `DeploymentManagerActor` Tells types from
`ReplicationMessages.cs`; `AkkaHostedService` constructs the actor. Landed as one commit, which
also strengthens Task 14's own invariant (never both mechanisms, never neither).
- **The legacy-copy column intersection.** The plan said "migrate all 16 columns"; that would be a
data-loss bug. An older legacy file lacks `execution_id`/`parent_execution_id`/
`last_attempt_at_ms`, naming a missing column throws `no such column`, and the existing reader
treats that as an unrecognised shape and **silently discards every row**. The copy intersects the
legacy and current column sets, with a required-column (PK) guard so the tolerance cannot degrade
into copying NULL-keyed rows.
- **`MigratorColumnLists_MatchTheLiveSchema`** (added beyond the plan). A column-list mismatch is
invisible at runtime in BOTH directions — a typo is silently dropped by the intersection, a
missing column silently leaves data behind. `SiteStorageTables`/`LegacyTable` are `internal` for it.
- **Absence assertions need a positive control.** `MessageAddedThenRemoved_NeverConvergesToPresent`
PASSED with `sf_messages` unregistered entirely — an absent row is also what a pair replicating
nothing looks like. Fixed with a control row that must converge in the same window.
- **The positional-argument hazard was REAL.** Removing `DeploymentManagerActor`'s optional
`IActorRef? replicationActor` shifted 6 trailing optionals and 4 test call sites bound the wrong
arguments, some with no compile error. **`Props.Create` builds an expression tree, which rejects
OUT-OF-POSITION named arguments** — so named args work only if in position, and the rest must be
padded positionally.
- **`ReplaceAllAsync` was unsafe to keep, not merely unused.** A mass DELETE on a now-replicated
table would be captured and shipped to the peer.
- **`LocalDbSitePairHarness`** is the shared two-node fixture (Phase 1 + both Phase 2 suites; Task 18
makes four). Its `Phase2ReplicatedTables` list is **literal, not derived from production code**, so
a registration that drifts fails the tests instead of agreeing with itself.
- **Ordinal sort gotcha:** `_` (0x5F) sorts before `b` (0x62), so `data_connection_definitions`
precedes `database_connections`.
- **Never `git checkout <file>` to undo a non-vacuity experiment** — it reverts to HEAD and takes
in-progress work with it (cost one redo of Task 8). Use `cp` backups.
---
## Rig operational gotchas (hard-won; not obvious from the repo)
- **`ExternalSystem.Call` does NOT buffer to store-and-forward — only `ExternalSystem.CachedCall`
does.** Using `Call` gives HTTP traffic and zero `sf_messages` rows. This is the single most
important detail for generating S&F load.
- **Never** query the bind-mounted SQLite files from the macOS host while containers run. Use
the plan's `snap` helper, or `docker run --rm -v <dir>:/d alpine/sqlite3 …`.
- **No `curl` in the `aspnet:10.0` image**, and metrics port 8084 is not published. Use a
sidecar: `docker run --rm --network container:scadabridge-site-a-a curlimages/curl:latest -s localhost:8084/metrics`
- **Seeded template 4 (`Motor Controller`) is not deployable** — 34 pre-deployment validation
errors. The soak used a purpose-built `SoakGenerator` template instead.
- CLI: `template script update` requires `--name` and `--trigger-type` even to change only
`--code`. In zsh, don't put auth flags in an unquoted variable (no word-splitting).
does.** The single most important detail for generating S&F load.
- **Never** query the bind-mounted SQLite files from the macOS host while containers run. Use the
plan's `snap` helper, or `docker run --rm -v <dir>:/d alpine/sqlite3 …`.
- **No `curl` in the `aspnet:10.0` image**, and metrics port 8084 is not published. Use a sidecar:
`docker run --rm --network container:scadabridge-site-a-a curlimages/curl:latest -s localhost:8084/metrics`
- **Seeded template 4 (`Motor Controller`) is not deployable** — 34 validation errors. The soak used
a purpose-built `SoakGenerator` template.
- CLI: `template script update` requires `--name` and `--trigger-type` even to change only `--code`.
In zsh, don't put auth flags in an unquoted variable.
- Rig auth: `--url http://localhost:9000 --username multi-role --password password`.
`LdapGroupMappings` role strings must be canonical `Designer`/`Deployer`; central caches them
at startup, so restart central after seeding them.
---
`LdapGroupMappings` roles must be canonical `Designer`/`Deployer`; central caches them at startup.
## Rig state as left
@@ -134,39 +157,42 @@ Reseeded and healthy; both site-a nodes restarted after the poisoning incident.
**Needs cleanup before the Task 20 live gate:**
- `ExternalSystemDefinitions` id 1 is still repointed to `http://127.0.0.1:9` — restore to
`http://scadabridge-restapi:5200`.
- Template `SoakGenerator` (id 2021) and instances `soakgen-1..4` (ids 58) are still deployed
on site-a and **still generating load**.
- Template `SoakGenerator` (id 2021) and instances `soakgen-1..4` (ids 58) are still deployed on
site-a and **still generating load**.
---
## Commits this session
## Commits (this branch, newest last)
| Commit | What |
|---|---|
| `25463d52` | plan review pass — corrections from code verification (D1/D3/D6, Task 1 method, Task 14) |
| `cf46e596` | fix(rig): `seed-sites.sh` canonical `Designer`/`Deployer` role names |
| `fc553bd9` | soak findings (gate verdict, since superseded) |
| `82c869a1` | task record: gate verdict |
| `e9e11d63` | root-cause brief for the disk-I/O issue |
| `8652eab9` | root cause + hardening: extended SQLite error codes, `SiteAuditTelemetryActor` async-context fix, `reseed.sh` init-script fix |
| `d512572d` | this resume-state doc (first version) |
| `12eb97e0` | **Task 2**gate CLOSED, all §5 questions answered + Task 1's clean re-run |
| `9e3239c5` | **Task 3**extract `StoreAndForwardSchema.Apply` |
| `ac5eb12c` | **Task 4**extract `SiteStorageSchema.Apply` |
| `3dfb288b` | task state: tasks 34 + deviations |
| `f2efeb37` | **Tasks 5+6**both stores take `ILocalDb`; latent directory defect fixed; `TestSupport` lib added |
| `19ab0ac9` | task state: tasks 56, the directory defect, the `CreateConnection` contract change |
| `25463d52` `d512572d` | plan review, soak, root-cause + hardening, first resume doc |
| `12eb97e0` | **Task 2** — gate CLOSED |
| `9e3239c5` | **Task 3**`StoreAndForwardSchema.Apply` |
| `ac5eb12c` | **Task 4**`SiteStorageSchema.Apply` |
| `f2efeb37` | **Tasks 5+6** — both stores take `ILocalDb`; latent directory defect fixed; `TestSupport` added |
| `fefbbb31` | resume state through wave 2 |
| `f8aa02e2` | **Task 7** — Phase 2 DDL in the consolidated DB (not yet registered) |
| `bdc0dffe` | **Task 8**migrate legacy `store-and-forward.db` |
| `5ddc7eed` | **Task 9**migrate legacy `scadabridge.db` config tables |
| `2bbe6631` | **Task 10**S&F CDC convergence specs (+ `LocalDbSitePairHarness` extraction) |
| `c56bf4ae` | **Task 11** — resync + config convergence specs |
| `79ce5161` | **Tasks 12+13**purge pin; guarded-write scope |
| `037798b3` | **Tasks 14+15+16 — THE CUTOVER** |
| `0ad11d6b` | task state for the cutover |
Two known flakes, do not chase: `InstanceActorChildAttributeRaceTests` (intermittent under full-suite
load) and `ParentExecutionIdCorrelationTests` (~91 s / times out on a cold MSSQL fixture, ~1 s warm).
Neither reproduced during waves 12.
Two known flakes, do not chase: `InstanceActorChildAttributeRaceTests` (intermittent under
full-suite load) and `ParentExecutionIdCorrelationTests` (~91 s / times out on a cold MSSQL
fixture, ~1 s warm). Neither reproduced during waves 15.
## Execution notes
- Waves 12 were run **sequentially in the main worktree**, not as the plan's parallel
worktree agents — the tasks were small enough that worktree setup + merge cost more than it
saved, and sequential execution removes the git race the isolation rule exists to prevent.
- Subagents *were* used for the bulk mechanical fixture rewiring (4 of them, disjoint file sets,
explicitly forbidden from running git). **Their reported results were re-verified
independently** — one agent's partial run had missed 2 real failures that a full-suite run
caught. Re-run the suites yourself; do not take an agent's pass at face value.
- Waves ran **sequentially in the main worktree**, not as the plan's parallel worktree agents — the
tasks were small enough that worktree setup + merge cost more than it saved, and sequential
execution removes the git race the isolation rule exists to prevent.
- Subagents were used only in waves 12, for bulk mechanical fixture rewiring. **Their reported
results were re-verified independently** — one agent's partial run had missed 2 real failures that
a full-suite run caught. Re-run the suites yourself.
- Every new test was checked for **non-vacuity** by breaking the thing it claims to pin (removing
DDL, removing the `Migrate` call, unregistering a table, commenting out the purge). This caught one
genuinely vacuous test. Keep doing it.