Files
lmxopcua/docs/plans/2026-07-20-localdb-adoption-phase2.md
T
Joseph Doherty f2049a31d0 docs(localdb-phase2): pre-answer Task 0's STOP condition — it does not fire
Phase 2's recon carries a STOP condition: a BLOB payload column in the legacy
alarm store-and-forward table could not be registered for replication and would
have to become base64 TEXT, sized against the chunk guidance. Resolved ahead of
execution so the answer survives into the Phase 2 session:

- No BLOB. All 8 columns are TEXT/INTEGER; the payload is PayloadJson TEXT NOT
  NULL. No base64 conversion, no chunk sizing needed.
- The executed DDL (SqliteStoreAndForwardSink.cs:657-667) matches the class
  doc-comment at :17-26 exactly — checked because a doc comment is not evidence.
- Payload is bounded by shape: a serialized AlarmHistorianEvent is 10 scalar
  fields, no collections or nesting, so realistic worst case is low single-digit
  KB against a 171 KB guidance.
- The PK IS autoincrement, confirming the plan's prediction that the migrator
  must mint deterministic mig-{node}-{legacyId} ids and the new table needs a
  TEXT GUID PK.
- Also captured for reuse: the drain's covering index (DeadLettered, RowId), and
  the exact legacy column list the migrator's pragma_table_info intersection
  check compares against.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-21 03:46:56 -04:00

17 KiB
Raw Blame History

OtOpcUa LocalDb Adoption — Phase 2 Implementation Plan (alarm-historian store-and-forward)

For Claude: REQUIRED SUB-SKILL: Use superpowers-extended-cc:executing-plans to implement this plan task-by-task.

Execution model: Optimized for Claude Opus agents (claude --model opus); dispatch high-risk tasks on Opus. Branch feat/localdb-phase2 in ~/Desktop/OtOpcUa (remote lmxopcua). Prerequisite: Phase 1 (docs/plans/2026-07-20-localdb-adoption-phase1.md) is merged (or this branch is stacked on it) and its live gate passed. Do not merge as part of this plan; stop at the DoD task.

Design authority: ~/Desktop/scadaproj/docs/plans/2026-07-20-otopcua-localdb-design.md §3.4, D8, D9. Reference implementation for "replace a bespoke store" phasing: ScadaBridge Phase 2 (~/Desktop/ScadaBridge/docs/plans/2026-07-19-localdb-adoption-phase2.md + its live gate doc).

Goal: Move the alarm-historian store-and-forward buffer (today the standalone alarm-historian.db owned by SqliteStoreAndForwardSink) into the consolidated LocalDb file as a replicated table with a primary-gated drain, so a redundant pair no longer loses buffered alarm history when a node dies — and delete the sink's bespoke file/connection management outright.

Architecture: alarm_sf_events (TEXT GUID PK) registered in LocalDbSetup.OnReady; the sink rewired onto ILocalDb behind its unchanged public seam; drain gated on the delivered-snapshot Primary role via PrimaryGatePolicy (at-least-once across failover, accepted and documented); a one-time idempotent migrator from the legacy file, running after registration.

Risk framing (from ScadaBridge): Phase 2 replaces a working mechanism — a harder risk class than Phase 1's "add where none existed." Cutover (delete + rewire) lands in one commit; there is no dual-mechanism period, and the cutover is the test.

Hard rules: identical to Phase 1's list (OnReady ordering, no autoincrement/BLOB, fail-closed auth already in place, no in-memory SQLite in tests, cp-triplet-only rig inspection), plus:

  • Legacy-copy column lists must INTERSECT with what the legacy file actually has (pragma_table_info probe) — a missing column throws and readers silently discard every row.
  • Absence assertions need a positive control.
  • DoD greps phrased as "no references from code" (explanatory comments may survive).

Task 0: Recon (produces docs/plans/2026-07-20-localdb-phase2-recon.md)

Classification: standard Estimated implement time: ~5 min Parallelizable with: none

Files:

  • Create: docs/plans/2026-07-20-localdb-phase2-recon.md
  • Read-only: src/Core/ZB.MOM.WW.OtOpcUa.Core.AlarmHistorian/SqliteStoreAndForwardSink.cs (and the whole Core.AlarmHistorian project), src/Server/ZB.MOM.WW.OtOpcUa.Runtime/Historian/AlarmHistorianOptions.cs, the drain worker (whatever forwards batches to GatewayHistorian/SendEvent), src/Server/ZB.MOM.WW.OtOpcUa.Runtime/Drivers/PrimaryGatePolicy.cs, src/Server/ZB.MOM.WW.OtOpcUa.Runtime/Drivers/DriverHostActor.cs (how _localRole + driver-member-count reach the gate today)

Record, with file:line citations:

  1. The sink's exact table schema: name, columns, PK type (autoincrement? — decides whether the migrator needs deterministic mig-{node}-{legacyId} ids), indices, and any BLOB columns (STOP condition: a BLOB payload column cannot be registered — it must become base64 TEXT in the new schema, and the recon must size the largest realistic payload against the 171 KB-ish chunk guidance; alarm events are small JSON, so expect this to be fine, but verify).

    PRE-ANSWERED 2026-07-21 (verified against master d218282c) — the STOP condition does NOT fire. Re-verify cheaply during recon, but do not expect a surprise:

    • Executed DDL: SqliteStoreAndForwardSink.cs:657-667. It matches the class doc-comment at :17-26 exactly — no drift between the documented and executed schema.
    • No BLOB. All 8 columns are TEXT/INTEGER; the payload column is PayloadJson TEXT NOT NULL. So no base64 conversion is needed and the chunk-size sizing is moot.
    • Payload is small and bounded by shape. PayloadJson is a serialized AlarmHistorianEvent — 10 scalar fields (AlarmId, EquipmentPath, AlarmName, AlarmTypeName, Severity, EventKind, Message, User, Comment?, TimestampUtc). No collections, no nesting. Realistic worst case is low single-digit KB (operator Comment is the only unbounded-ish field), far under the 171 KB chunk guidance.
    • PK IS autoincrementRowId INTEGER PRIMARY KEY AUTOINCREMENT. This confirms the plan's own prediction: LocalDb cannot replicate an autoincrement key, so the migrator must mint deterministic mig-{node}-{legacyId} ids, and the new table needs a TEXT GUID PK as already specified in the Architecture note.
    • One index to carry across: IX_Queue_Drain ON Queue (DeadLettered, RowId) (:667) — the drain's covering index. The alarm_sf_events equivalent wants the same shape over (dead-lettered flag, insertion order) so the drain query stays index-covered.
    • Legacy column list for the migrator's pragma_table_info intersection check: RowId, AlarmId, EnqueuedUtc, PayloadJson, AttemptCount, LastAttemptUtc, LastError, DeadLettered.
  2. The public seam: the interface the drain worker and producers use (e.g. IAlarmHistorianSink / enqueue+dequeue+markDelivered+deadLetter methods), so the rewire can keep it byte-compatible.

  3. Semantics to preserve: Capacity (1,000,000) enforcement, MaxAttempts (10), dead-letter retention (30 d), BatchSize (100), DrainIntervalSeconds (5) — where each lives.

  4. The drain worker's lifecycle: hosted service or actor? Where a Primary-role check can be injected, and how the delivered-snapshot role (RedundancyStateChanged cache) is accessible from it (via DriverHostActor, a shared status service, or a message). If the role is only available inside DriverHostActor, note the cleanest bridge (e.g. an IRedundancyRoleView singleton the actor updates) — that becomes Task 4's shape.

  5. Whether the sink is constructed per-node config path (AlarmHistorian:DatabasePath) anywhere else (tests, tooling).

  6. How AlarmHistorian:Enabled=false short-circuits (NullAlarmHistorianSink) — the rewire must keep the disabled path allocating no LocalDb tables? No: tables are created unconditionally in OnReady (cheap, empty); only the sink/drain stay Null. Note this in the doc.

Commit: docs(localdb): phase-2 recon findings.


Task 1: alarm_sf_events schema + registration (+ tests)

Classification: standard Estimated implement time: ~4 min Parallelizable with: none (Task 2 depends on it)

Files:

  • Create: src/Core/ZB.MOM.WW.OtOpcUa.Core.AlarmHistorian/AlarmSfSchema.cs (depends only on Microsoft.Data.Sqlite)
  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Host/Configuration/LocalDbSetup.cs
  • Test: extend LocalDbSetupTests

Schema shape (adjust column names to the recon's findings — preserve today's semantics):

CREATE TABLE IF NOT EXISTS alarm_sf_events (
    id               TEXT    NOT NULL PRIMARY KEY,   -- app-minted GUID (never autoincrement)
    payload_json     TEXT    NOT NULL,
    enqueued_at_utc  TEXT    NOT NULL,
    attempts         INTEGER NOT NULL DEFAULT 0,
    status           TEXT    NOT NULL DEFAULT 'pending',  -- pending | delivered | dead
    last_attempt_utc TEXT    NULL,
    dead_at_utc      TEXT    NULL
);
CREATE INDEX IF NOT EXISTS ix_alarm_sf_events_status ON alarm_sf_events(status, enqueued_at_utc);

OnReady order becomes: Phase-1 DDL → AlarmSfSchema.Apply → the two Phase-1 RegisterReplicated calls → RegisterReplicated("alarm_sf_events")migrator (Task 5) last. (All DDL may run before all registrations; the invariant is registration-before-writes.)

TDD: failing test first — exact replicated set becomes ["alarm_sf_events", "deployment_artifacts", "deployment_pointer"] (ordinal-sorted; update the Phase-1 exact-set pins in the same commit — they are supposed to go red here, that's them working). Oplog-capture test for an alarm_sf_events insert. Commit feat(localdb): alarm_sf_events replicated table.


Task 2: Rewire the sink onto ILocalDb + delete bespoke file management (the cutover commit, part 1 of 2 — see Task 3)

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: none

Files:

  • Modify: src/Core/ZB.MOM.WW.OtOpcUa.Core.AlarmHistorian/SqliteStoreAndForwardSink.cs (or replace with LocalDbStoreAndForwardSink.cs — keep the public seam identical either way)
  • Modify: its registration (AddAlarmHistorian) to inject ILocalDb
  • Modify: src/Server/ZB.MOM.WW.OtOpcUa.Runtime/Historian/AlarmHistorianOptions.cs — remove DatabasePath (a breaking config key removal: note it in the runbook/CHANGELOG task)
  • Test: the sink's existing unit tests, rewired to a temp-file ILocalDb via TestLocalDb-style helper (create tests/.../TestSupport helper if none exists — real DB, never a stub: a stubbed bare SqliteConnection lacks zb_hlc_next() and fails closed on registered tables)

Steps:

  1. Write/port failing tests for the seam's semantics: enqueue, drain batch of BatchSize, MaxAttempts → dead-letter, capacity enforcement, dead-letter retention purge.
  2. Implement over ILocalDb.ExecuteAsync/QueryAsync (anonymous-object params). Delete the private connection/pragma/file-open code and any PRAGMA journal_mode calls (LocalDb owns pragmas). GUIDs minted at enqueue (Guid.NewGuid().ToString("N")).
  3. Capacity enforcement: count-based insert guard (preserve today's overflow behavior per recon).
  4. Core.AlarmHistorian gains a package ref on core ZB.MOM.WW.LocalDb (interface only).
  5. Build + project tests green. Commit together with Task 3 if the drain gate can't compile separately (the ScadaBridge tasks-14/15/16 circular-dependency landmine — check before assuming they're independent commits).

Task 3: Primary-gated drain

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: none

Files: (exact shape from recon item 4)

  • Modify: the drain worker
  • Create (if recon says so): src/Server/ZB.MOM.WW.OtOpcUa.Runtime/Redundancy/IRedundancyRoleView.cs — a singleton snapshot (RedundancyRole? LocalRole, int DriverMemberCount) updated by DriverHostActor where it already caches _localRole
  • Test: drain-worker tests + an actor test pinning that DriverHostActor publishes role changes to the view

Semantics:

  • Drain runs only when PrimaryGatePolicy.ShouldServiceAsPrimary(localRole, driverMemberCount) is true — same policy, same boot-window posture (unknown role drains only when the node is alone).
  • Delivered/dead-letter marks are row UPDATEs → they replicate, so the standby's copy tracks drain progress and does not re-deliver already-marked rows after failover.
  • At-least-once across failover is accepted: rows delivered on the old primary whose delivered mark hadn't replicated yet will be re-sent by the new primary. Document in the runbook (Task 7); do NOT build dedup.
  • When replication is OFF (default), the gate still applies but driverMemberCount for a solo node keeps today's behavior — verify with a test: single-node, role unknown → drains (no regression for unpaired deployments).

TDD: failing tests — secondary role does not drain; primary drains; unknown+alone drains; unknown+paired does not. Commit (with Task 2 if coupled): feat(localdb): alarm S&F on LocalDb with primary-gated drain (cutover).


Task 4: One-time legacy migrator (alarm-historian.db → consolidated)

Classification: standard Estimated implement time: ~5 min Parallelizable with: Task 6

Files:

  • Create: src/Server/ZB.MOM.WW.OtOpcUa.Host/Configuration/AlarmSfLegacyMigrator.cs
  • Modify: LocalDbSetup.OnReady — call it last, after all RegisterReplicated calls
  • Test: tests/Server/ZB.MOM.WW.OtOpcUa.Host.Tests/AlarmSfLegacyMigratorTests.cs

Mirror ~/Desktop/ScadaBridge/src/ZB.MOM.WW.ScadaBridge.Host/SiteLocalDbLegacyMigrator.cs:

  • Source path from the pre-removal AlarmHistorian:DatabasePath default (alarm-historian.db) — read the raw config key even though the option property is gone.
  • ShouldMigrate: skip if file missing or <file>.migrated exists. :memory:/file: sources → no-op.
  • If the legacy PK is autoincrement (recon): deterministic ids mig-{NodeName}-{legacyId} (node name from the recon-identified config key) — rerunnable without duplicates under INSERT OR IGNORE, and no cross-node collision. If already GUIDs, copy as-is.
  • One transaction for the whole copy; File.Move(path, path + ".migrated") only after commit; failure throws out of OnReady → boot fails, legacy untouched.
  • Column intersection via pragma_table_info on the legacy table; required-PK guard.
  • Both pair nodes migrate independently; their rows have distinct ids (node-prefixed), so the merged buffer is the union — expected; note that the new primary will drain the standby's migrated rows too.

TDD: failing tests — happy path row counts; idempotent re-run; failure leaves legacy file untouched; migrated rows enter the oplog (assert __localdb_oplog — the registration-order pin); older legacy file missing a column still migrates (intersection). Commit feat(localdb): one-time alarm-historian.db migrator.


Task 5: Convergence + failover scenarios in the pair harness

Classification: high-risk Estimated implement time: ~5 min Parallelizable with: Task 4

Files:

  • Test: tests/Server/ZB.MOM.WW.OtOpcUa.Host.IntegrationTests/LocalDb/AlarmSfConvergenceTests.cs (reuses the Phase-1 LocalDbPairHarness)

Scenarios:

  1. AlarmBurstOnA_ConvergesToB_AndOplogDrains — enqueue N events on A; identical rowset on B; oplog → 0.
  2. DeliveredMarksReplicate — mark rows delivered on A; B's copies show delivered (the no-redeliver-after-failover property, asserted at the data layer).
  3. WritesWhileTransportDown_SurviveRejoin.
  4. Positive control: with RegisterReplicated("alarm_sf_events") commented out, scenarios 13 go red (run locally, record, restore — do not commit red).

Commit test(localdb): alarm S&F convergence scenarios.


Task 6: Rig config + docs

Classification: small Estimated implement time: ~4 min Parallelizable with: Task 4

Files:

  • Modify: docker-dev/docker-compose.yml — enable AlarmHistorian__Enabled=true on the site-a pair if not already (the gate needs a live buffer); remove any AlarmHistorian__DatabasePath env vars (key deleted).
  • Modify: docs/operations/2026-07-20-localdb-pair-replication.md — add: alarm S&F replication semantics, at-least-once-across-failover statement, migrator behavior (.migrated sidecar), DatabasePath key removal.
  • Modify: docs/AlarmHistorian.md + CLAUDE.md — sink now lives in the consolidated LocalDb; drain is primary-gated.

Commit docs+chore(localdb): phase-2 rig config + docs.


Task 7: DoD sweep (offline)

Classification: standard Estimated implement time: ~5 min Parallelizable with: none

  1. Full solution build → 0 warnings; full test suite green (deltas vs pre-branch baseline only).
  2. Greps, phrased as "no references from code": the old bespoke connection management (AlarmHistorian:DatabasePath, direct new SqliteConnection inside Core.AlarmHistorian except via schema helpers/tests) — explanatory comments may remain.
  3. Positive-control evidence from Task 5 recorded.
  4. Exact-set replicated-tables pin = 3 tables, both directions.
  5. Update …phase2.md.tasks.json; commit chore(localdb): phase-2 DoD sweep; STOP and report.

Task 8: Live gate on the docker-dev rig (run only with explicit user go-ahead)

Classification: high-risk (rig) Estimated implement time: ~30 min wall-clock Parallelizable with: none

Same inspection rules as Phase 1's gate (cp-triplet, curl sidecar, observer-suspicion rule). Record in docs/plans/2026-07-20-localdb-phase2-live-gate.md:

  1. Migration ran: .migrated sidecar present on both site-a nodes; row counts match legacy.
  2. Alarm burst (drive a real driver alarm or the historian-gateway-unreachable path) converges: identical rowsets, oplog drains to 0, dead letters 0.
  3. Only the primary drains: stop the historian gateway egress, buffer builds on BOTH nodes' tables (replicated), but only the primary's drain worker logs attempts.
  4. Failover: stop the primary; the standby (new primary) resumes draining the shared buffer; count of double-delivered events observed and recorded (at-least-once evidence, not a failure).
  5. Both-nodes-together restart clean; zero disk I/O error/SQLITE_IOERR in logs.
  6. site-b (default-OFF pin): sink works locally, no sync traffic.

Commit the gate doc; report; do not merge.


Task persistence

Tasks file: docs/plans/2026-07-20-localdb-adoption-phase2.md.tasks.json. Record deviations per task — especially if Tasks 2/3 had to land as one commit (the expected outcome).