Files
ScadaBridge/docs/plans/2026-07-20-phase2-resume-state.md
T
Joseph Doherty fefbbb31da docs(localdb): resume state through wave 2 (tasks 1-6 done, wave 3 next)
Records the latent Phase 1 directory defect and its open library-vs-app decision,
the CreateConnection contract change, the new TestSupport library, and the
verification numbers at the wave-2 boundary.

Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts
2026-07-20 03:14:18 -04:00

10 KiB
Raw Blame History

LocalDb Phase 2 — resume state (written 2026-07-20, pre-compaction)

Scratch handoff for picking the plan back up. Delete once Phase 2 lands.

Plan: 2026-07-19-localdb-adoption-phase2.md (execute with superpowers-extended-cc:executing-plans; task state in …-phase2.md.tasks.json) Branch: feat/localdb-phase2 (from feat/localdb-phase1, which is not merged/pushed) Tree: clean @ 19ab0ac9. Build 0 warnings. Nothing pushed to origin yet.


Where we are

Tasks 16 are DONE (waves 0, 1, 2). Next up: Wave 3 = Tasks 7, 8, 9.

Task 1's gate is closed with verdict PROCEED; its earlier STOP was superseded (the disk I/O error storm was observer-induced — host-side sqlite3 reads of live bind-mounted WAL files reset the WAL under the container, not a product defect; write-up in docs/known-issues/2026-07-20-localdb-disk-io-error-under-load.md §0/§11a).

Task State
1 Rig soak done — gate PROCEED
2 Decision record done — gate doc CLOSED, all §5 questions answered inline
3 StoreAndForwardSchema.Apply done
4 SiteStorageSchema.Apply done
5 S&F store → ILocalDb done
6 SiteStorageServiceILocalDb done
721 not started, no code written

Verification at the wave-2 boundary: full solution build 0 warnings; SiteRuntime 532, Host 318, AuditLog 355, StoreAndForward 153, ExternalSystemGateway 142, HealthMonitoring 97 — 1597 passed / 0 failed.

Task 1's measured numbers (clean re-run, 2026-07-20)

Six consecutive 60 s intervals, copy-based snap() sampling on restarted nodes: sf_messages 0.80 rows/sec insert, 0/sec retry-UPDATE, oplog 0, native_alarm_state 0, max payload_json 76 B, max config_json 721 B (rig — not representative), zero SQLite errors in 30 min, both site-a nodes converged at 3564 rows. Honest gap: 0.80/s is ~1.6 % of the 50/s ceiling and the retry path was never exercised — fine only because that ceiling is structural, not empirical.

Next: Wave 3 (Tasks 7, 8, 9)

Task 7 extends SiteLocalDbSetup.OnReady with the new DDL (not registered yet — registration is Task 14, the cutover). Tasks 8 and 9 are the legacy migrators and are both high-risk. Ordering inside OnReady is load-bearing: DDL → RegisterReplicated → migrate. Rows written before registration are invisible to the peer forever, silently.


Findings that change the plan (carry these forward)

From 2026-07-19-localdb-phase2-soak.md:

  • D6's premise is wrong, its risk is real. The plan says config_json is "documented to exceed 128 KB per row." The source actually says the escaped Akka envelope blew the 128 KB frame and the raw JSON "only needs to be ~60-70 KB". So the largest known production row is ~6070 KB → no stop condition. But MaxBatchSize=500 is unsafe (500 × 70 KB ≈ 35 MB vs a 4 MB gRPC cap) → Task 19 must set LocalDb:Replication:MaxBatchSize = 16. This is the one firmly evidence-backed number.
  • D4 is wrong. native_alarm_state is not "unbounded by design" — NativeAlarmActor coalesces per SourceReference on a 100 ms flush (NativeAlarmActor.cs:473-502), so worst case is distinct_source_refs × 10/sec.
  • sf_messages has a hard 50 rows/sec ceilingSweepBatchLimit (500) ÷ RetryTimerInterval (10 s). A ceiling, not an estimate.
  • Task 1 step 7's stop condition is much weaker than written. Exceeding the oplog caps prunes and flags needs_snapshot → graceful snapshot resync (OplogStore.cs:109-138, MaintenanceBackgroundService.cs:57). Not data loss, not a wedge.

The plan's own Task-1 method also needed four corrections (metrics sidecar, safe sampling, no alarm generator on this rig, template 4 undeployable) — all recorded in the soak doc §1.

Found during waves 12 (new)

  • LATENT PHASE-1 DEFECT, found and fixed in Task 6. The plan assumed directory creation moved to LocalDb along with file ownership. It did not. The LocalDb library never creates the parent directory, and SqliteLocalDb opens the file eagerly in its constructor, so a missing directory is a hard boot failure (SQLite Error 14: unable to open database file), not a degraded start. The default site config uses the relative path ./data/site-localdb.db, so any site node without a pre-existing data/ fails to boot; the docker rig escapes only because its volume mount creates /app/data. Latent since Phase 1 made LocalDb:Path required. Fixed in ScadaBridge via SiteLocalDbDirectory.Ensure(config) called before AddZbLocalDb, pinned by Host.Tests/SiteLocalDbDirectoryTests. OPEN DECISION FOR THE USER: the real fix arguably belongs in the ZB.MOM.WW.LocalDb library (it owns the file) — every other LocalDb consumer still has this gap. Not done here because a library change means a version bump and is outside this plan's scope.
  • Silent contract change on SiteStorageService.CreateConnection(). It used to return an unopened connection that callers opened; it now returns an already-open LocalDb connection, and calling Open/OpenAsync on one throws. SiteExternalSystemRepository dropped 5 OpenAsync calls. Anything new that consumes this member must not open it.
  • LocalDb has NO in-memory mode. LocalDbOptions.Path is a filesystem path, so every fixture that used Mode=Memory;Cache=Shared had to move to a real temp file. Expect this for any remaining store that gets rewired.
  • Test fallout from a store constructor change is ~7× what the plan estimates. Task 5's plan named "fixtures" in one project; the real blast radius was 40 files across 7 test projects. Budget accordingly for Tasks 14/15.
  • tests/ZB.MOM.WW.ScadaBridge.TestSupport is a NEW shared, non-test library holding TestLocalDb (Create/CreateTemp/DeleteFiles; dispose before deleting — the master connection anchors the WAL). Use it rather than copying fixtures. Wired into AuditLog, ExternalSystemGateway, HealthMonitoring, IntegrationTests, SiteRuntime, StoreAndForward tests.
  • Where behaviour moved owner, retarget tests — do not delete them with the code. WAL genuinely is LocalDb's job (test retargeted in place); directory creation is not (test moved to Host.Tests). Both directions occurred, so check which owner actually holds the guarantee.

Rig operational gotchas (hard-won; not obvious from the repo)

  • ExternalSystem.Call does NOT buffer to store-and-forward — only ExternalSystem.CachedCall does. Using Call gives HTTP traffic and zero sf_messages rows. This is the single most important detail for generating S&F load.
  • Never query the bind-mounted SQLite files from the macOS host while containers run. Use the plan's snap helper, or docker run --rm -v <dir>:/d alpine/sqlite3 ….
  • No curl in the aspnet:10.0 image, and metrics port 8084 is not published. Use a sidecar: docker run --rm --network container:scadabridge-site-a-a curlimages/curl:latest -s localhost:8084/metrics
  • Seeded template 4 (Motor Controller) is not deployable — 34 pre-deployment validation errors. The soak used a purpose-built SoakGenerator template instead.
  • CLI: template script update requires --name and --trigger-type even to change only --code. In zsh, don't put auth flags in an unquoted variable (no word-splitting).
  • Rig auth: --url http://localhost:9000 --username multi-role --password password. LdapGroupMappings role strings must be canonical Designer/Deployer; central caches them at startup, so restart central after seeding them.

Rig state as left

Reseeded and healthy; both site-a nodes restarted after the poisoning incident.

Needs cleanup before the Task 20 live gate:

  • ExternalSystemDefinitions id 1 is still repointed to http://127.0.0.1:9 — restore to http://scadabridge-restapi:5200.
  • Template SoakGenerator (id 2021) and instances soakgen-1..4 (ids 58) are still deployed on site-a and still generating load.

Commits this session

Commit What
25463d52 plan review pass — corrections from code verification (D1/D3/D6, Task 1 method, Task 14)
cf46e596 fix(rig): seed-sites.sh canonical Designer/Deployer role names
fc553bd9 soak findings (gate verdict, since superseded)
82c869a1 task record: gate verdict
e9e11d63 root-cause brief for the disk-I/O issue
8652eab9 root cause + hardening: extended SQLite error codes, SiteAuditTelemetryActor async-context fix, reseed.sh init-script fix
d512572d this resume-state doc (first version)
12eb97e0 Task 2 — gate CLOSED, all §5 questions answered + Task 1's clean re-run
9e3239c5 Task 3 — extract StoreAndForwardSchema.Apply
ac5eb12c Task 4 — extract SiteStorageSchema.Apply
3dfb288b task state: tasks 34 + deviations
f2efeb37 Tasks 5+6 — both stores take ILocalDb; latent directory defect fixed; TestSupport lib added
19ab0ac9 task state: tasks 56, the directory defect, the CreateConnection contract change

Two known flakes, do not chase: InstanceActorChildAttributeRaceTests (intermittent under full-suite load) and ParentExecutionIdCorrelationTests (~91 s / times out on a cold MSSQL fixture, ~1 s warm). Neither reproduced during waves 12.

Execution notes

  • Waves 12 were run sequentially in the main worktree, not as the plan's parallel worktree agents — the tasks were small enough that worktree setup + merge cost more than it saved, and sequential execution removes the git race the isolation rule exists to prevent.
  • Subagents were used for the bulk mechanical fixture rewiring (4 of them, disjoint file sets, explicitly forbidden from running git). Their reported results were re-verified independently — one agent's partial run had missed 2 real failures that a full-suite run caught. Re-run the suites yourself; do not take an agent's pass at face value.