71379816e7
Moves the alarm store-and-forward buffer out of its own alarm-historian.db and into the node's consolidated LocalDb, where it replicates to the redundant pair peer. A node that dies holding undelivered alarm history no longer takes it to the grave. Tasks 2 and 3 land together, as the plan anticipated. They are not separable: the gate is a constructor argument of the rewritten sink, and a commit that replicated the queue without gating the drain would be a commit in which both nodes of a pair deliver every alarm event, continuously. That drain gate is the load-bearing part of the change, and the recon explains why it is new work rather than a refinement. Exactly-once delivery across a pair is enforced today on the ENQUEUE side, by HistorianAdapterActor: only the Primary enqueues, so the Secondary's queue is empty and its ungated drain has nothing to send. Replicating the table destroys that invariant. The gate is a Func<bool> the caller supplies, because the drain runs on a timer the sink owns rather than on a mailbox, and Core.AlarmHistorian cannot reference PrimaryGatePolicy in Runtime. Runtime supplies it via a new IRedundancyRoleView singleton that DriverHostActor publishes its Primary-gate verdict to -- the same verdict the inbound-write and native-ack gates use, so there is no second notion of am-I-the-Primary to drift. Two failure modes are deliberately closed: - The view is seeded OPEN, matching the policy's own answer for an unknown role with no driver peer. A deployment that runs no redundancy never publishes to it, and defaulting closed would silently stop its alarm history forever. - A gate that throws is read as not-now, never as permission, and a closed gate reports the new HistorianDrainState.NotPrimary rather than Idle. A Secondary's rising queue is supposed to look different from a stalled drain, and if BOTH nodes report NotPrimary the pair is misconfigured and says so instead of quietly filling toward the capacity ceiling. Row ids are a hash of the payload rather than fresh GUIDs. Both adapters accept the same fanned transition in the window before the first redundancy snapshot arrives, and under last-writer-wins an equal key converges those two accepts into one row instead of duplicating them. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
82 lines
3.2 KiB
JSON
82 lines
3.2 KiB
JSON
{
|
|
"planPath": "docs/plans/2026-07-20-localdb-adoption-phase2.md",
|
|
"tasks": [
|
|
{
|
|
"id": 0,
|
|
"subject": "Task 0: Recon \u2014 sink schema, seam, drain lifecycle, role-view bridge (STOP conditions)",
|
|
"status": "completed",
|
|
"note": "STOP condition does NOT fire (no BLOB; PayloadJson TEXT). Recon doc: docs/plans/2026-07-20-localdb-phase2-recon.md. Key finding: the drain worker is an internal Timer inside the sink (not a hosted service/actor), so the gate is a Func<bool> ctor param; and HistorianAdapterActor already primary-gates ENQUEUE with a different policy (ShouldHistorize) - which is why today's ungated drain is safe and why replication breaks it. Deviations D-1..D-4 recorded."
|
|
},
|
|
{
|
|
"id": 1,
|
|
"subject": "Task 1: alarm_sf_events schema + registration (+ exact-set pin update)",
|
|
"status": "completed",
|
|
"blockedBy": [
|
|
0
|
|
],
|
|
"note": "alarm_sf_events created in AlarmSfSchema (Core.AlarmHistorian) + registered third in LocalDbSetup.OnReady. Exact-set pin updated to 3 tables. DEVIATION D-5: kept the legacy delete-on-ack + dead_lettered flag + last_error rather than the plan's status column (no sweeper for 'delivered' rows; last_error is the only record of why a row died). Drain ORDER BY moves to (enqueued_at_utc, id)."
|
|
},
|
|
{
|
|
"id": 2,
|
|
"subject": "Task 2: Rewire sink onto ILocalDb + delete bespoke file management (cutover 1/2)",
|
|
"status": "completed",
|
|
"blockedBy": [
|
|
1
|
|
],
|
|
"note": "Sink rewritten as LocalDbStoreAndForwardSink over ILocalDb; bespoke file/pragma/schema management deleted with the old class. AlarmHistorian:DatabasePath removed (breaking config key). Ids are a deterministic payload hash (D-1), not GUIDs."
|
|
},
|
|
{
|
|
"id": 3,
|
|
"subject": "Task 3: Primary-gated drain via PrimaryGatePolicy (cutover 2/2 \u2014 may co-commit with Task 2)",
|
|
"status": "completed",
|
|
"blockedBy": [
|
|
2
|
|
],
|
|
"note": "Drain gated on IRedundancyRoleView, a singleton DriverHostActor publishes PrimaryGatePolicy's verdict to on every snapshot. Fails closed on a throwing gate; seeded OPEN so a non-redundant deployment is never silently stopped. New HistorianDrainState.NotPrimary + transition-logged (D-2). Landed with Task 2 as one commit (D-3)."
|
|
},
|
|
{
|
|
"id": 4,
|
|
"subject": "Task 4: One-time alarm-historian.db legacy migrator",
|
|
"status": "pending",
|
|
"blockedBy": [
|
|
3
|
|
]
|
|
},
|
|
{
|
|
"id": 5,
|
|
"subject": "Task 5: Convergence + failover scenarios in the pair harness (+ positive control)",
|
|
"status": "pending",
|
|
"blockedBy": [
|
|
3
|
|
]
|
|
},
|
|
{
|
|
"id": 6,
|
|
"subject": "Task 6: Rig config + docs",
|
|
"status": "pending",
|
|
"blockedBy": [
|
|
3
|
|
]
|
|
},
|
|
{
|
|
"id": 7,
|
|
"subject": "Task 7: DoD sweep (offline) \u2014 STOP and report after this",
|
|
"status": "pending",
|
|
"blockedBy": [
|
|
4,
|
|
5,
|
|
6
|
|
]
|
|
},
|
|
{
|
|
"id": 8,
|
|
"subject": "Task 8: Live gate on the docker-dev rig (needs explicit user go-ahead)",
|
|
"status": "pending",
|
|
"blockedBy": [
|
|
7
|
|
]
|
|
}
|
|
],
|
|
"lastUpdated": "2026-07-21T00:00:00Z"
|
|
}
|