Files
ScadaBridge/docs/plans/2026-07-20-phase2-resume-state.md
T

199 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# LocalDb Phase 2 — resume state (updated 2026-07-20, pre-compaction #2)
Scratch handoff for picking the plan back up. Delete once Phase 2 lands.
**Plan:** [`2026-07-19-localdb-adoption-phase2.md`](2026-07-19-localdb-adoption-phase2.md)
(execute with `superpowers-extended-cc:executing-plans`; task state in
`…-phase2.md.tasks.json`, which carries a per-task deviation note)
**Branch:** `feat/localdb-phase2` (from `feat/localdb-phase1`, which is **not** merged)
**Tree:** clean @ `0ad11d6b`. Build 0 warnings. 25 commits ahead of `origin/feat/localdb-phase2`.
---
## Where we are
**Tasks 116 are DONE. The cutover has landed. Next up: Task 17.**
| Task | State |
|---|---|
| 16 | done (waves 02) — see the git log |
| 7 `SiteLocalDbSetup` DDL | done |
| 8 `sf_messages` migrator | done |
| 9 config-table migrator | done |
| 10 S&F CDC convergence specs | done |
| 11 resync + config convergence specs | done |
| 12 active-node purge pin | done |
| 13 guarded-write scope check | done |
| **14 + 15 + 16 the CUTOVER** | **done — landed as ONE commit `037798b3`** |
| **1721** | **not started** |
**Verification at the cutover boundary:** build 0 warnings; SiteRuntime 512, StoreAndForward 130,
Host 330, AuditLog 355, ExternalSystemGateway 142, HealthMonitoring 97, LocalDb integration 16 —
**all pass, 0 failures.** (SiteRuntime/StoreAndForward fell from 533/153 as 6 obsolete test files
were deleted.)
### The bespoke replicators are gone
`SiteReplicationActor`, `ReplicationMessages.cs`, `ReplicationService` and
`StoreAndForwardStorage.ReplaceAllAsync` are deleted. Ten tables now replicate via LocalDb CDC:
the Phase 1 pair plus `sf_messages` and the 7 site config tables.
`notification_lists` and `smtp_configurations` are created but **never registered** — permanently
empty by design, and registering them would open a standing replication channel whose only
historical payload was plaintext SMTP passwords. That decision is pinned by its own
security-named test and stated in `OnReady`.
---
## Next: Task 17 (config-key cleanup)
Partly **required**, not optional: `SiteRuntimeOptions.cs:66` still documents
`ConfigFetchRetryCount` as "Consumed by `SiteReplicationActor`", which no longer exists.
Per the plan: relax `StartupValidator.cs:113` (`SiteDbPath`) and
`StoreAndForwardOptionsValidator.cs:19-21` (`SqliteDbPath`) to migration-only; delete
`ReplicationEnabled` and `ConfigFetchRetryCount` outright with their validator rules and config
entries; **leave the two path keys present in configs** with a comment — the legacy migrator reads
them, and removing them would strand un-migrated data on a node that has not yet started once.
**10** `appsettings.Site.json` files, including `deploy/wonder-app-vd03/`.
Then 18 (two-node convergence suite), 19 (rig config — `MaxBatchSize = 16`), 20 (live gate),
21 (docs truth pass).
---
## Findings that change the plan (carry these forward)
### From Task 1 (unchanged)
- **D6's premise is wrong, its risk is real.** Largest known production `config_json` is ~6070 KB,
not >128 KB → no stop condition. But `MaxBatchSize=500` × 70 KB ≈ 35 MB vs a 4 MB gRPC cap →
**Task 19 must set `LocalDb:Replication:MaxBatchSize = 16`.** The one firmly evidence-backed number.
- **D4 is wrong.** `native_alarm_state` is bounded by per-`SourceReference` coalescing on a 100 ms
flush (`NativeAlarmActor.cs:473-502`).
- **`sf_messages` has a hard 50 rows/sec ceiling** — `SweepBatchLimit` (500) ÷ `RetryTimerInterval` (10 s).
- **Task 1 step 7's stop condition is weaker than written.** Exceeding the oplog caps prunes and
flags `needs_snapshot` → graceful resync, not data loss.
Measured (six 60 s intervals, copy-based `snap()`): `sf_messages` **0.80 rows/sec** insert, **0/sec**
retry-UPDATE, oplog 0, max `payload_json` 76 B, max `config_json` 721 B (rig — *not* representative),
zero SQLite errors in 30 min.
### From waves 12
- **LATENT PHASE-1 DEFECT, fixed in-app only.** The LocalDb library never creates the parent
directory of `LocalDb:Path`, and `SqliteLocalDb` opens the file **eagerly in its constructor**, so
a missing directory is a **hard boot failure** (`SQLite Error 14`). The default site config uses
the **relative** `./data/site-localdb.db`; the docker rig escapes only because its volume mount
creates `/app/data`. Fixed via `SiteLocalDbDirectory.Ensure(config)` before `AddZbLocalDb`, pinned
by `Host.Tests/SiteLocalDbDirectoryTests`.
**OPEN DECISION FOR THE USER:** the real fix arguably belongs in the `ZB.MOM.WW.LocalDb` library
**every other consumer still has this gap.** Not done here: a library change means a version
bump and is outside this plan's scope.
- **`SiteStorageService.CreateConnection()` contract flipped** — returns an **already-open**
connection; calling `Open`/`OpenAsync` on one **throws**.
- **LocalDb has NO in-memory mode.** Every `Mode=Memory;Cache=Shared` fixture became a real temp file.
- **A store constructor change's test blast radius is ~7× the plan's estimate** (Task 5: 40 files,
7 projects).
- **`tests/ZB.MOM.WW.ScadaBridge.TestSupport`** holds `TestLocalDb` (dispose **before** deleting —
the master connection anchors the WAL). Use it; don't copy fixtures.
- **Where behaviour moved owner, retarget tests — don't delete them with the code.** WAL is genuinely
LocalDb's; directory creation is **not**. Both directions occurred.
### From waves 35 (new)
- **PLAN DEFECT — Tasks 14/15/16 cannot compile separately.** `SiteReplicationActor` takes a
`ReplicationService` and calls `ReplaceAllAsync`; `DeploymentManagerActor` Tells types from
`ReplicationMessages.cs`; `AkkaHostedService` constructs the actor. Landed as one commit, which
also strengthens Task 14's own invariant (never both mechanisms, never neither).
- **The legacy-copy column intersection.** The plan said "migrate all 16 columns"; that would be a
data-loss bug. An older legacy file lacks `execution_id`/`parent_execution_id`/
`last_attempt_at_ms`, naming a missing column throws `no such column`, and the existing reader
treats that as an unrecognised shape and **silently discards every row**. The copy intersects the
legacy and current column sets, with a required-column (PK) guard so the tolerance cannot degrade
into copying NULL-keyed rows.
- **`MigratorColumnLists_MatchTheLiveSchema`** (added beyond the plan). A column-list mismatch is
invisible at runtime in BOTH directions — a typo is silently dropped by the intersection, a
missing column silently leaves data behind. `SiteStorageTables`/`LegacyTable` are `internal` for it.
- **Absence assertions need a positive control.** `MessageAddedThenRemoved_NeverConvergesToPresent`
PASSED with `sf_messages` unregistered entirely — an absent row is also what a pair replicating
nothing looks like. Fixed with a control row that must converge in the same window.
- **The positional-argument hazard was REAL.** Removing `DeploymentManagerActor`'s optional
`IActorRef? replicationActor` shifted 6 trailing optionals and 4 test call sites bound the wrong
arguments, some with no compile error. **`Props.Create` builds an expression tree, which rejects
OUT-OF-POSITION named arguments** — so named args work only if in position, and the rest must be
padded positionally.
- **`ReplaceAllAsync` was unsafe to keep, not merely unused.** A mass DELETE on a now-replicated
table would be captured and shipped to the peer.
- **`LocalDbSitePairHarness`** is the shared two-node fixture (Phase 1 + both Phase 2 suites; Task 18
makes four). Its `Phase2ReplicatedTables` list is **literal, not derived from production code**, so
a registration that drifts fails the tests instead of agreeing with itself.
- **Ordinal sort gotcha:** `_` (0x5F) sorts before `b` (0x62), so `data_connection_definitions`
precedes `database_connections`.
- **Never `git checkout <file>` to undo a non-vacuity experiment** — it reverts to HEAD and takes
in-progress work with it (cost one redo of Task 8). Use `cp` backups.
---
## Rig operational gotchas (hard-won; not obvious from the repo)
- **`ExternalSystem.Call` does NOT buffer to store-and-forward — only `ExternalSystem.CachedCall`
does.** The single most important detail for generating S&F load.
- **Never** query the bind-mounted SQLite files from the macOS host while containers run. Use the
plan's `snap` helper, or `docker run --rm -v <dir>:/d alpine/sqlite3 …`.
- **No `curl` in the `aspnet:10.0` image**, and metrics port 8084 is not published. Use a sidecar:
`docker run --rm --network container:scadabridge-site-a-a curlimages/curl:latest -s localhost:8084/metrics`
- **Seeded template 4 (`Motor Controller`) is not deployable** — 34 validation errors. The soak used
a purpose-built `SoakGenerator` template.
- CLI: `template script update` requires `--name` and `--trigger-type` even to change only `--code`.
In zsh, don't put auth flags in an unquoted variable.
- Rig auth: `--url http://localhost:9000 --username multi-role --password password`.
`LdapGroupMappings` roles must be canonical `Designer`/`Deployer`; central caches them at startup.
## Rig state as left
Reseeded and healthy; both site-a nodes restarted after the poisoning incident.
**Needs cleanup before the Task 20 live gate:**
- `ExternalSystemDefinitions` id 1 is still repointed to `http://127.0.0.1:9` — restore to
`http://scadabridge-restapi:5200`.
- Template `SoakGenerator` (id 2021) and instances `soakgen-1..4` (ids 58) are still deployed on
site-a and **still generating load**.
---
## Commits (this branch, newest last)
| Commit | What |
|---|---|
| `25463d52``d512572d` | plan review, soak, root-cause + hardening, first resume doc |
| `12eb97e0` | **Task 2** — gate CLOSED |
| `9e3239c5` | **Task 3**`StoreAndForwardSchema.Apply` |
| `ac5eb12c` | **Task 4**`SiteStorageSchema.Apply` |
| `f2efeb37` | **Tasks 5+6** — both stores take `ILocalDb`; latent directory defect fixed; `TestSupport` added |
| `fefbbb31` | resume state through wave 2 |
| `f8aa02e2` | **Task 7** — Phase 2 DDL in the consolidated DB (not yet registered) |
| `bdc0dffe` | **Task 8** — migrate legacy `store-and-forward.db` |
| `5ddc7eed` | **Task 9** — migrate legacy `scadabridge.db` config tables |
| `2bbe6631` | **Task 10** — S&F CDC convergence specs (+ `LocalDbSitePairHarness` extraction) |
| `c56bf4ae` | **Task 11** — resync + config convergence specs |
| `79ce5161` | **Tasks 12+13** — purge pin; guarded-write scope |
| `037798b3` | **Tasks 14+15+16 — THE CUTOVER** |
| `0ad11d6b` | task state for the cutover |
Two known flakes, do not chase: `InstanceActorChildAttributeRaceTests` (intermittent under
full-suite load) and `ParentExecutionIdCorrelationTests` (~91 s / times out on a cold MSSQL
fixture, ~1 s warm). Neither reproduced during waves 15.
## Execution notes
- Waves ran **sequentially in the main worktree**, not as the plan's parallel worktree agents — the
tasks were small enough that worktree setup + merge cost more than it saved, and sequential
execution removes the git race the isolation rule exists to prevent.
- Subagents were used only in waves 12, for bulk mechanical fixture rewiring. **Their reported
results were re-verified independently** — one agent's partial run had missed 2 real failures that
a full-suite run caught. Re-run the suites yourself.
- Every new test was checked for **non-vacuity** by breaking the thing it claims to pin (removing
DDL, removing the `Migrate` call, unregistering a table, commenting out the purge). This caught one
genuinely vacuous test. Keep doing it.