Records the latent Phase 1 directory defect and its open library-vs-app decision, the CreateConnection contract change, the new TestSupport library, and the verification numbers at the wave-2 boundary. Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts
10 KiB
LocalDb Phase 2 — resume state (written 2026-07-20, pre-compaction)
Scratch handoff for picking the plan back up. Delete once Phase 2 lands.
Plan: 2026-07-19-localdb-adoption-phase2.md
(execute with superpowers-extended-cc:executing-plans; task state in
…-phase2.md.tasks.json)
Branch: feat/localdb-phase2 (from feat/localdb-phase1, which is not merged/pushed)
Tree: clean @ 19ab0ac9. Build 0 warnings. Nothing pushed to origin yet.
Where we are
Tasks 1–6 are DONE (waves 0, 1, 2). Next up: Wave 3 = Tasks 7, 8, 9.
Task 1's gate is closed with verdict PROCEED; its earlier STOP was superseded (the
disk I/O error storm was observer-induced — host-side sqlite3 reads of live
bind-mounted WAL files reset the WAL under the container, not a product defect; write-up in
docs/known-issues/2026-07-20-localdb-disk-io-error-under-load.md
§0/§11a).
| Task | State |
|---|---|
| 1 Rig soak | done — gate PROCEED |
| 2 Decision record | done — gate doc CLOSED, all §5 questions answered inline |
3 StoreAndForwardSchema.Apply |
done |
4 SiteStorageSchema.Apply |
done |
5 S&F store → ILocalDb |
done |
6 SiteStorageService → ILocalDb |
done |
| 7–21 | not started, no code written |
Verification at the wave-2 boundary: full solution build 0 warnings; SiteRuntime 532, Host 318, AuditLog 355, StoreAndForward 153, ExternalSystemGateway 142, HealthMonitoring 97 — 1597 passed / 0 failed.
Task 1's measured numbers (clean re-run, 2026-07-20)
Six consecutive 60 s intervals, copy-based snap() sampling on restarted nodes:
sf_messages 0.80 rows/sec insert, 0/sec retry-UPDATE, oplog 0, native_alarm_state 0,
max payload_json 76 B, max config_json 721 B (rig — not representative),
zero SQLite errors in 30 min, both site-a nodes converged at 3564 rows. Honest gap: 0.80/s
is ~1.6 % of the 50/s ceiling and the retry path was never exercised — fine only because that
ceiling is structural, not empirical.
Next: Wave 3 (Tasks 7, 8, 9)
Task 7 extends SiteLocalDbSetup.OnReady with the new DDL (not registered yet — registration
is Task 14, the cutover). Tasks 8 and 9 are the legacy migrators and are both high-risk.
Ordering inside OnReady is load-bearing: DDL → RegisterReplicated → migrate. Rows written
before registration are invisible to the peer forever, silently.
Findings that change the plan (carry these forward)
From 2026-07-19-localdb-phase2-soak.md:
- D6's premise is wrong, its risk is real. The plan says
config_jsonis "documented to exceed 128 KB per row." The source actually says the escaped Akka envelope blew the 128 KB frame and the raw JSON "only needs to be ~60-70 KB". So the largest known production row is ~60–70 KB → no stop condition. ButMaxBatchSize=500is unsafe (500 × 70 KB ≈ 35 MB vs a 4 MB gRPC cap) → Task 19 must setLocalDb:Replication:MaxBatchSize = 16. This is the one firmly evidence-backed number. - D4 is wrong.
native_alarm_stateis not "unbounded by design" —NativeAlarmActorcoalesces perSourceReferenceon a 100 ms flush (NativeAlarmActor.cs:473-502), so worst case isdistinct_source_refs × 10/sec. sf_messageshas a hard 50 rows/sec ceiling —SweepBatchLimit(500) ÷RetryTimerInterval(10 s). A ceiling, not an estimate.- Task 1 step 7's stop condition is much weaker than written. Exceeding the oplog caps
prunes and flags
needs_snapshot→ graceful snapshot resync (OplogStore.cs:109-138,MaintenanceBackgroundService.cs:57). Not data loss, not a wedge.
The plan's own Task-1 method also needed four corrections (metrics sidecar, safe sampling, no alarm generator on this rig, template 4 undeployable) — all recorded in the soak doc §1.
Found during waves 1–2 (new)
- LATENT PHASE-1 DEFECT, found and fixed in Task 6. The plan assumed directory creation
moved to LocalDb along with file ownership. It did not. The LocalDb library never creates
the parent directory, and
SqliteLocalDbopens the file eagerly in its constructor, so a missing directory is a hard boot failure (SQLite Error 14: unable to open database file), not a degraded start. The default site config uses the relative path./data/site-localdb.db, so any site node without a pre-existingdata/fails to boot; the docker rig escapes only because its volume mount creates/app/data. Latent since Phase 1 madeLocalDb:Pathrequired. Fixed in ScadaBridge viaSiteLocalDbDirectory.Ensure(config)called beforeAddZbLocalDb, pinned byHost.Tests/SiteLocalDbDirectoryTests. OPEN DECISION FOR THE USER: the real fix arguably belongs in theZB.MOM.WW.LocalDblibrary (it owns the file) — every other LocalDb consumer still has this gap. Not done here because a library change means a version bump and is outside this plan's scope. - Silent contract change on
SiteStorageService.CreateConnection(). It used to return an unopened connection that callers opened; it now returns an already-open LocalDb connection, and callingOpen/OpenAsyncon one throws.SiteExternalSystemRepositorydropped 5OpenAsynccalls. Anything new that consumes this member must not open it. - LocalDb has NO in-memory mode.
LocalDbOptions.Pathis a filesystem path, so every fixture that usedMode=Memory;Cache=Sharedhad to move to a real temp file. Expect this for any remaining store that gets rewired. - Test fallout from a store constructor change is ~7× what the plan estimates. Task 5's plan named "fixtures" in one project; the real blast radius was 40 files across 7 test projects. Budget accordingly for Tasks 14/15.
tests/ZB.MOM.WW.ScadaBridge.TestSupportis a NEW shared, non-test library holdingTestLocalDb(Create/CreateTemp/DeleteFiles; dispose before deleting — the master connection anchors the WAL). Use it rather than copying fixtures. Wired into AuditLog, ExternalSystemGateway, HealthMonitoring, IntegrationTests, SiteRuntime, StoreAndForward tests.- Where behaviour moved owner, retarget tests — do not delete them with the code. WAL genuinely
is LocalDb's job (test retargeted in place); directory creation is not (test moved to
Host.Tests). Both directions occurred, so check which owner actually holds the guarantee.
Rig operational gotchas (hard-won; not obvious from the repo)
ExternalSystem.Calldoes NOT buffer to store-and-forward — onlyExternalSystem.CachedCalldoes. UsingCallgives HTTP traffic and zerosf_messagesrows. This is the single most important detail for generating S&F load.- Never query the bind-mounted SQLite files from the macOS host while containers run. Use
the plan's
snaphelper, ordocker run --rm -v <dir>:/d alpine/sqlite3 …. - No
curlin theaspnet:10.0image, and metrics port 8084 is not published. Use a sidecar:docker run --rm --network container:scadabridge-site-a-a curlimages/curl:latest -s localhost:8084/metrics - Seeded template 4 (
Motor Controller) is not deployable — 34 pre-deployment validation errors. The soak used a purpose-builtSoakGeneratortemplate instead. - CLI:
template script updaterequires--nameand--trigger-typeeven to change only--code. In zsh, don't put auth flags in an unquoted variable (no word-splitting). - Rig auth:
--url http://localhost:9000 --username multi-role --password password.LdapGroupMappingsrole strings must be canonicalDesigner/Deployer; central caches them at startup, so restart central after seeding them.
Rig state as left
Reseeded and healthy; both site-a nodes restarted after the poisoning incident.
Needs cleanup before the Task 20 live gate:
ExternalSystemDefinitionsid 1 is still repointed tohttp://127.0.0.1:9— restore tohttp://scadabridge-restapi:5200.- Template
SoakGenerator(id 2021) and instancessoakgen-1..4(ids 5–8) are still deployed on site-a and still generating load.
Commits this session
| Commit | What |
|---|---|
25463d52 |
plan review pass — corrections from code verification (D1/D3/D6, Task 1 method, Task 14) |
cf46e596 |
fix(rig): seed-sites.sh canonical Designer/Deployer role names |
fc553bd9 |
soak findings (gate verdict, since superseded) |
82c869a1 |
task record: gate verdict |
e9e11d63 |
root-cause brief for the disk-I/O issue |
8652eab9 |
root cause + hardening: extended SQLite error codes, SiteAuditTelemetryActor async-context fix, reseed.sh init-script fix |
d512572d |
this resume-state doc (first version) |
12eb97e0 |
Task 2 — gate CLOSED, all §5 questions answered + Task 1's clean re-run |
9e3239c5 |
Task 3 — extract StoreAndForwardSchema.Apply |
ac5eb12c |
Task 4 — extract SiteStorageSchema.Apply |
3dfb288b |
task state: tasks 3–4 + deviations |
f2efeb37 |
Tasks 5+6 — both stores take ILocalDb; latent directory defect fixed; TestSupport lib added |
19ab0ac9 |
task state: tasks 5–6, the directory defect, the CreateConnection contract change |
Two known flakes, do not chase: InstanceActorChildAttributeRaceTests (intermittent under full-suite
load) and ParentExecutionIdCorrelationTests (~91 s / times out on a cold MSSQL fixture, ~1 s warm).
Neither reproduced during waves 1–2.
Execution notes
- Waves 1–2 were run sequentially in the main worktree, not as the plan's parallel worktree agents — the tasks were small enough that worktree setup + merge cost more than it saved, and sequential execution removes the git race the isolation rule exists to prevent.
- Subagents were used for the bulk mechanical fixture rewiring (4 of them, disjoint file sets, explicitly forbidden from running git). Their reported results were re-verified independently — one agent's partial run had missed 2 real failures that a full-suite run caught. Re-run the suites yourself; do not take an agent's pass at face value.