271fcc4e15
Two driver nodes replicating the alarm buffer over the real loopback h2c
transport, through the real fail-closed interceptor. This is the test the
phase exists to pass: schema, registration, sink rewire and drain gate can
all be individually green while a pair still fails to share the undelivered
alarm history that is the whole point of moving the queue.
Rows are written through the production sink rather than hand-rolled
INSERTs, so the id derivation, column set and delete-on-ack semantics under
test are the ones production uses. The scenarios cover a buffered burst
converging with the oplog draining to zero, delivered rows being removed from
BOTH nodes as tombstones (the no-redeliver-after-failover property), writes
during a transport outage surviving the rejoin, and the same event accepted
on both nodes collapsing to one row.
The positive control found a weak assertion. With
RegisterReplicated("alarm_sf_events") commented out, three of the four went
red immediately -- but the same-event-on-both-nodes scenario still PASSED,
because with replication off each node trivially holds its own single copy
and "one row on each" is satisfied by two databases that never spoke. It now
also enqueues a distinct event on B and requires both nodes to hold two rows,
which cannot be satisfied without convergence; it goes red under the control
like the others. Control restored, all four green.
Also updates the second exact-set replicated-tables pin, in LocalDbWiringTests.
The plan predicted these pins would go red here, and that is them working:
one of them is asserted through the full driver DI graph rather than the
OnReady callback, so it catches a registration that exists in the callback
but never reaches a running host.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
84 lines
4.1 KiB
JSON
84 lines
4.1 KiB
JSON
{
|
|
"planPath": "docs/plans/2026-07-20-localdb-adoption-phase2.md",
|
|
"tasks": [
|
|
{
|
|
"id": 0,
|
|
"subject": "Task 0: Recon \u2014 sink schema, seam, drain lifecycle, role-view bridge (STOP conditions)",
|
|
"status": "completed",
|
|
"note": "STOP condition does NOT fire (no BLOB; PayloadJson TEXT). Recon doc: docs/plans/2026-07-20-localdb-phase2-recon.md. Key finding: the drain worker is an internal Timer inside the sink (not a hosted service/actor), so the gate is a Func<bool> ctor param; and HistorianAdapterActor already primary-gates ENQUEUE with a different policy (ShouldHistorize) - which is why today's ungated drain is safe and why replication breaks it. Deviations D-1..D-4 recorded."
|
|
},
|
|
{
|
|
"id": 1,
|
|
"subject": "Task 1: alarm_sf_events schema + registration (+ exact-set pin update)",
|
|
"status": "completed",
|
|
"blockedBy": [
|
|
0
|
|
],
|
|
"note": "alarm_sf_events created in AlarmSfSchema (Core.AlarmHistorian) + registered third in LocalDbSetup.OnReady. Exact-set pin updated to 3 tables. DEVIATION D-5: kept the legacy delete-on-ack + dead_lettered flag + last_error rather than the plan's status column (no sweeper for 'delivered' rows; last_error is the only record of why a row died). Drain ORDER BY moves to (enqueued_at_utc, id)."
|
|
},
|
|
{
|
|
"id": 2,
|
|
"subject": "Task 2: Rewire sink onto ILocalDb + delete bespoke file management (cutover 1/2)",
|
|
"status": "completed",
|
|
"blockedBy": [
|
|
1
|
|
],
|
|
"note": "Sink rewritten as LocalDbStoreAndForwardSink over ILocalDb; bespoke file/pragma/schema management deleted with the old class. AlarmHistorian:DatabasePath removed (breaking config key). Ids are a deterministic payload hash (D-1), not GUIDs."
|
|
},
|
|
{
|
|
"id": 3,
|
|
"subject": "Task 3: Primary-gated drain via PrimaryGatePolicy (cutover 2/2 \u2014 may co-commit with Task 2)",
|
|
"status": "completed",
|
|
"blockedBy": [
|
|
2
|
|
],
|
|
"note": "Drain gated on IRedundancyRoleView, a singleton DriverHostActor publishes PrimaryGatePolicy's verdict to on every snapshot. Fails closed on a throwing gate; seeded OPEN so a non-redundant deployment is never silently stopped. New HistorianDrainState.NotPrimary + transition-logged (D-2). Landed with Task 2 as one commit (D-3)."
|
|
},
|
|
{
|
|
"id": 4,
|
|
"subject": "Task 4: One-time alarm-historian.db legacy migrator",
|
|
"status": "completed",
|
|
"blockedBy": [
|
|
3
|
|
],
|
|
"note": "AlarmSfLegacyMigrator runs LAST in OnReady (which now takes IConfiguration - no skip-migration overload exists). DEVIATION D-6: ids are the payload hash (AlarmSfSchema.DeriveId, lifted out of the sink) rather than mig-{node}-{legacyId} - a warm pair's two legacy files OVERLAP, and node-prefixing would carry that duplication forward forever. DEVIATION: tests live in Host.IntegrationTests; the plan's Host.Tests project does not exist."
|
|
},
|
|
{
|
|
"id": 5,
|
|
"subject": "Task 5: Convergence + failover scenarios in the pair harness (+ positive control)",
|
|
"status": "completed",
|
|
"blockedBy": [
|
|
3
|
|
],
|
|
"note": "4 scenarios green. POSITIVE CONTROL RUN (RegisterReplicated(alarm_sf_events) commented out): 3/4 went red immediately; the 4th (same-event-on-both-nodes) passed VACUOUSLY - with replication off each node trivially held its own single row. Strengthened by enqueuing a second distinct event on B and asserting BOTH nodes hold 2; it then went red under the control too. Control restored, all 4 green. Also fixed a SECOND exact-set pin the plan predicted: LocalDbWiringTests."
|
|
},
|
|
{
|
|
"id": 6,
|
|
"subject": "Task 6: Rig config + docs",
|
|
"status": "pending",
|
|
"blockedBy": [
|
|
3
|
|
]
|
|
},
|
|
{
|
|
"id": 7,
|
|
"subject": "Task 7: DoD sweep (offline) \u2014 STOP and report after this",
|
|
"status": "pending",
|
|
"blockedBy": [
|
|
4,
|
|
5,
|
|
6
|
|
]
|
|
},
|
|
{
|
|
"id": 8,
|
|
"subject": "Task 8: Live gate on the docker-dev rig (needs explicit user go-ahead)",
|
|
"status": "pending",
|
|
"blockedBy": [
|
|
7
|
|
]
|
|
}
|
|
],
|
|
"lastUpdated": "2026-07-21T00:00:00Z"
|
|
}
|