fix(alarms): re-assert non-normal scripted-alarm conditions after a (re)load (#487)
A deploy that triggers a FULL address-space rebuild clears `_alarmConditions` and re-materialises every condition node fresh — inactive, acked, confirmed, unshelved. The engine reloads its persisted state and re-derives Active from the predicate, so it ends up correct; but there is no transition to report, so `LoadAsync` yields `EmissionKind.None` and `OnEngineEmission` drops it. Nothing writes the node. Live-reproduced on the docker-dev rig against the pre-fix image: an alarm that had fired and cleared but was still UNACKNOWLEDGED came back from a rebuild reporting Acknowledged with Retain=false — so it disappeared from ConditionRefresh entirely, while the engine still held Unacknowledged. An operator's outstanding alarm silently vanishes from the alarm list on deploy. `ScriptedAlarmHostActor.ReassertConditionNodes` closes it. After each load it reads the new `ScriptedAlarmEngine.GetProjections()` and sends one `AlarmStateUpdate` per alarm NOT in the Part 9 no-event position. Load-bearing properties: - Node-only. It sends `AlarmStateUpdate` and nothing else — never the `alerts` topic, never the telemetry hub. `alerts` is the historization path, so a row per alarm per deploy would append a duplicate historian/AVEVA record every time anyone deploys. This is why it reads a projection rather than re-emitting: `ScriptedAlarmProjection` carries no `EmissionKind`, so there is nothing for the emission machinery — or a future refactor — to mistake for a transition. The issue's sketch said `GetState(alarmId)` + `LoadedAlarmIds`; that alone would have clobbered the node's severity and message, so the projection also carries the definition's severity, the resolved message template, and the last-observed worst-input quality. - Only non-normal alarms. One in the no-event position already matches what the materialise built, so a never-fired alarm stays byte-identical to a deploy without this pass (including its Message, which the materialise seeds with the display name), and an all-normal deploy writes nothing at all. - Timestamp is the condition's own `LastTransitionUtc`, not the deploy instant. - No duplicate Part 9 events: `WriteAlarmCondition`'s delta-gate compares against the node's CURRENT state, so a re-assert onto a freshly materialised node is a genuine delta (fires once — correct, the rebuild reset what clients see) while one onto a surgically-preserved node is suppressed. Tests: 6 new host tests + 6 new engine tests. Falsifiability checked by disabling the call — 4 of the 6 host tests go red; the other 2 are absence guards against an over-broad fix and pass either way by design. Live gate (docker-dev, central-2, pre-fix image vs. rebuilt image): - pre-fix: condition absent from ConditionRefresh after a rebuild; - post-fix: present, Unacknowledged, Retain=True, stamped with its own LastTransitionUtc rather than the restart instant; - /alerts stayed empty across two re-asserts, with the panel proven live by a real ACTIVATED transition immediately afterwards — no duplicate history.
This commit is contained in:
@@ -102,6 +102,45 @@ Persisted scope per plan decision #14: `Enabled`, `Acked`, `Confirmed`, `Shelvin
|
||||
|
||||
Every mutation the state machine produces is immediately persisted inside the engine's `_evalGate` semaphore, so the store's view is always consistent with the in-memory state.
|
||||
|
||||
### Post-(re)load condition-node re-assert (#487)
|
||||
|
||||
A deploy that triggers a **full address-space rebuild** clears `_alarmConditions` and re-materialises every
|
||||
condition node fresh — inactive, acked, confirmed, unshelved. The engine, reloading in parallel, restores each
|
||||
alarm's persisted state and re-derives `Active` from the predicate, so it ends up *correct*; but a still-active
|
||||
alarm has **no transition to report**, so `LoadAsync` yields `EmissionKind.None` and `OnEngineEmission` drops it.
|
||||
Nothing then writes the node. Before this fix, an active alarm with static dependencies read **normal** to every
|
||||
OPC UA client after such a deploy, until its next real transition — which for a static-dependency alarm may
|
||||
never come.
|
||||
|
||||
`ScriptedAlarmHostActor.ReassertConditionNodes` closes it. After each load completes it reads
|
||||
`ScriptedAlarmEngine.GetProjections()` — a plain read returning `ScriptedAlarmProjection` (condition state +
|
||||
the definition's severity + the resolved message template + last-observed worst-input quality) — and sends one
|
||||
`OpcUaPublishActor.AlarmStateUpdate` per alarm that is **not** in the Part 9 no-event position.
|
||||
|
||||
Four properties are load-bearing:
|
||||
|
||||
- **Node-only.** It sends `AlarmStateUpdate` and nothing else — never the `alerts` topic, never the telemetry
|
||||
hub. `alerts` is the historization path, so an alerts row per alarm per deploy would append a duplicate
|
||||
historian / AVEVA record every time anyone deploys. This is why the re-assert reads a projection instead of
|
||||
re-emitting: `ScriptedAlarmProjection` carries no `EmissionKind`, so there is nothing for the emission
|
||||
machinery — or a future refactor — to mistake for a transition.
|
||||
- **Only non-normal alarms.** An alarm in the no-event position (enabled, inactive, acked, confirmed,
|
||||
unshelved) already matches what the materialise built, so it is skipped. That keeps a never-fired alarm
|
||||
byte-identical to a deploy without this pass — including its Message, which the materialise seeds with the
|
||||
alarm's display name — and writes nothing at all on the common all-normal deploy.
|
||||
- **Ordering.** `DriverHostActor` tells `RebuildAddressSpace` to the publish actor *before* it tells
|
||||
`ApplyScriptedAlarms` to the host, and the re-assert runs later still (after an awaited `LoadAsync` pipes
|
||||
back). Both are local `Tell`s, which enqueue synchronously, so the re-materialise is already queued ahead of
|
||||
these writes.
|
||||
- **No duplicate Part 9 events.** `WriteAlarmCondition`'s delta-gate compares the incoming snapshot against
|
||||
the node's *current* state: a re-assert onto a freshly materialised (normal) node is a genuine delta and
|
||||
fires exactly one condition event — correct, since the rebuild reset what clients see — while a re-assert
|
||||
onto a surgically-preserved node that already holds the same state is suppressed. The write is ungated by
|
||||
redundancy role, like every other node write here; the Secondary keeps its address space warm for failover.
|
||||
|
||||
The timestamp carried is the condition's own `LastTransitionUtc`, **not** the deploy instant, so a client
|
||||
reading `Time` / `ReceiveTime` after a rebuild still sees when the alarm actually went active.
|
||||
|
||||
## Source integration
|
||||
|
||||
`ScriptedAlarmSource` (`ScriptedAlarmSource.cs`) adapts the engine to the driver-agnostic `IAlarmSource` interface. The existing `AlarmSurfaceInvoker` + `GenericDriverNodeManager` fan-out consumes it the same way it consumes Galaxy / AB CIP / FOCAS sources — there is no scripted-alarm-specific code path in the server plumbing. From that point on, the flow into `AlarmConditionState` nodes, the `AlarmAck` session check, and the Historian sink is shared — see [AlarmTracking.md](AlarmTracking.md).
|
||||
|
||||
Reference in New Issue
Block a user