Enables the alarm store-and-forward sink on all four driver nodes of the docker-dev rig -- the replicating site-a pair and the default-OFF site-b pin -- so the live gate has a real buffer to watch converge, and can see that site-b's sink works as a plain node-local queue with no peer traffic. Two rig details worth stating rather than rediscovering. The endpoint is deliberately unresolvable: there is no HistorianGateway here, so every drain attempt fails, which is exactly the historian-outage state the buffer exists for. But it still has to be a syntactically valid absolute http(s) URI or the host refuses to start, because ServerHistorianOptionsValidator is consumer-gated -- AlarmHistorian:Enabled=true makes the endpoint required even while ServerHistorian:Enabled=false. And MaxAttempts is raised far above the production default of 10, which against a permanently unreachable gateway would dead-letter the entire queue about five minutes in, turning a buffering test into a dead-letter test. Docs record what an operator now has to know: that a rising queue depth on a Secondary is correct and on BOTH nodes is not (that shape is a redundancy snapshot naming neither node -- check node identity before suspecting the sink), that delivery is at-least-once across a failover by design with no dedup layer, and that AlarmHistorian:DatabasePath is removed but should be left in place through the upgrade because the migrator still reads it to find the file to copy. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
15 KiB
Alarm Historian — store-and-forward SQLite sink
Reference for ZB.MOM.WW.OtOpcUa.Core.AlarmHistorian
(src/Core/ZB.MOM.WW.OtOpcUa.Core.AlarmHistorian/),
the durable local queue that historizes alarm transitions to AVEVA Historian
without ever blocking the alarm engine or operator actions.
This is the sink mechanics doc. For how the three alarm sources converge on the OPC UA Part 9 surface and which alarms route here, see AlarmTracking.md. For the historian client that drains this queue, see DriverLifecycle.md and ServiceHosting.md.
Why store-and-forward
Scripted alarms (and any future non-Galaxy IAlarmSource, e.g. AB CIP ALMD)
must reach AVEVA Historian, but the historian gateway can be slow, busy, or
disconnected. The sink decouples the alarm engine from historian reachability:
every qualifying transition is committed to a local SQLite queue first, and
a background drain worker forwards rows to the historian on a backoff-aware
cadence. Operator acks and alarm-state transitions are never blocked waiting on
the historian.
Galaxy-native alarms with
$Alarm*extensions reach AVEVA Historian directly via System Platform'sHistorizeToAvevatoggle — they do not flow through this sink. This path is exclusively for non-Galaxy alarm producers.
Contracts
All in
IAlarmHistorianSink.cs
unless noted.
IAlarmHistorianSink— the intake contract.EnqueueAsync(evt, ct)durably enqueues an event and returns as soon as the queue row is committed (fire-and-forget from the engine's perspective; the sink must not block the emitting thread).GetStatus()returns aHistorianSinkStatussnapshot.NullAlarmHistorianSink— the no-op default for tests and deployments that don't historize alarms. It is the default DI binding (registered in the Runtime'sAddOtOpcUaRuntime); production overrides it withLocalDbStoreAndForwardSink.AlarmHistorianEvent(AlarmHistorianEvent.cs) — the source-agnostic event record:AlarmId,EquipmentPath(UNS path, doubles as Historian's SourceNode),AlarmName,AlarmTypeName(Part 9 subtype),Severity,EventKind(free-form transition string — "Activated"/"Cleared"/"Acknowledged"/etc.),Message,User,Comment,TimestampUtc.IAlarmHistorianWriter— what the drain worker delegates writes to.WriteBatchAsync(batch, ct)returns oneHistorianWriteOutcomeper event, in order. Production binds this toGatewayAlarmHistorianWriter(the HistorianGatewaySendEventpath).HistorianWriteOutcome— per-event drain result:Ack(persisted, remove from queue),RetryPlease(transient failure — leave queued, retry after backoff),PermanentFail(malformed/unrecoverable — move to dead-letter).HistorianSinkStatus— diagnostic snapshot surfaced to the AdminUI and/healthz:QueueDepth,DeadLetterDepth,LastDrainUtc,LastSuccessUtc,LastError,DrainState, andEvictedCount.HistorianDrainState—Disabled/Idle/Draining/BackingOff/NotPrimary(ticking, but leaving the replicated queue for the node that holds the Primary role).
LocalDbStoreAndForwardSink
LocalDbStoreAndForwardSink.cs
is the production IAlarmHistorianSink. Construction takes the node's
ILocalDb, an IAlarmHistorianWriter, a logger, and optional batchSize
(default 100), capacity (default 1,000,000), deadLetterRetention (default 30
days), a test clock, and a drainGate.
The drain gate
drainGate is a Func<bool> consulted at the top of every tick; false skips
the tick entirely, leaving the queue untouched and reporting
DrainState = NotPrimary.
It exists because the queue replicates. The Secondary holds a full copy of
the Primary's rows, and an ungated drain there would re-deliver every event,
continuously — not just at a failover. Runtime supplies the gate from
IRedundancyRoleView, a singleton DriverHostActor publishes its
PrimaryGatePolicy verdict to on every redundancy snapshot, so the drain agrees
with the inbound-write and native-ack gates by construction.
Two postures are deliberate:
- Unset (the default) means drain. So does an unpublished view. A deployment that runs no redundancy never publishes a role, and defaulting closed would silently stop its alarm history forever.
- A gate that throws means "not now", never permission. Draining on a gate fault would put a pair into dual delivery exactly when its redundancy signal is sick.
Queue table
The sink owns one table in the node's consolidated LocalDb — created and
registered for replication by LocalDbSetup.OnReady, before any consumer can
resolve the sink:
CREATE TABLE alarm_sf_events (
id TEXT NOT NULL PRIMARY KEY, -- SHA-256 of payload_json
alarm_id TEXT NOT NULL,
enqueued_at_utc TEXT NOT NULL,
payload_json TEXT NOT NULL, -- JSON-serialized AlarmHistorianEvent
attempt_count INTEGER NOT NULL DEFAULT 0,
last_attempt_utc TEXT NULL,
last_error TEXT NULL,
dead_lettered INTEGER NOT NULL DEFAULT 0
);
CREATE INDEX ix_alarm_sf_events_drain
ON alarm_sf_events (dead_lettered, enqueued_at_utc, id);
The primary key is a hash of the payload, not an autoincrement rowid. Two
reasons. Replication converges last-writer-wins per primary key, so two nodes
independently allocating rowid 7 to different alarms would silently overwrite one
another. And an equal-payload key makes the same event arriving twice collapse
into one row — which happens for real, because HistorianAdapterActor
default-writes while its redundancy role is unknown and both nodes of a pair
therefore accept the same fanned transition in every boot window. Two genuinely
distinct events cannot collide: AlarmHistorianEvent carries a full-precision
timestamp alongside the alarm id, kind, message and user.
Ordering follows from that: the drain reads in enqueued_at_utc order (with id
as a tiebreak so the ordering is total), because a hashed key carries no
insertion sequence.
EnqueueAsync does a single INSERT … ON CONFLICT DO NOTHING on the hot path.
To avoid a SELECT COUNT(*) on every enqueue, the sink keeps an in-memory
non-dead-lettered row counter (seeded at startup, kept current by every mutation,
and re-synced from storage every 10,000 enqueues to defend against drift — now
doubly warranted, since replication applies the peer's rows straight into the
table where no local code path can observe them). Pragmas and connection
lifetime belong to LocalDb; the sink no longer manages either.
Drain worker
StartDrainLoop(tickInterval) starts a self-rescheduling one-shot
System.Threading.Timer (not started automatically — tests drive
DrainOnceAsync deterministically). Each tick:
- Consults the drain gate. A closed gate returns immediately, leaving the queue
untouched and
DrainState = NotPrimary. - Purges aged dead-lettered rows past the retention window.
- Reads up to
batchSizenon-dead-lettered rows inenqueued_at_utc, idorder. - Rows with un-deserializable payloads are dead-lettered immediately (by their
own
id) so they can't stall the queue head. - The remaining batch is handed to
IAlarmHistorianWriter.WriteBatchAsync, and each outcome is applied in one transaction:Ackdeletes the row (the delete replicates as a tombstone, so the peer drops its copy and a promoted standby cannot re-send it),PermanentFailflips itsdead_letteredflag,RetryPleasebumps its attempt count and leaves it queued. A row whoseAttemptCounthas reached the configuredMaxAttemptscap (default 10) is dead-lettered automatically on the next drain tick rather than retried — this breaks infinite retry loops for poison events whose payload the historian will always reject (e.g. a malformed alarm record that triggers a permanent SDK error on every attempt). The dead-lettered row remains inspectable viaRetryDeadLettered()for the configured retention window. - The timer re-arms its next due-time to
max(tickInterval, currentBackoff).
Backoff ladder (applied to the timer's next due-time, so a historian outage
genuinely slows the drain cadence): 1s → 2s → 5s → 15s → 60s cap. Any
RetryPlease outcome — or a writer exception, or a writer cardinality violation
(outcome count ≠ event count) — bumps the backoff and sets DrainState = BackingOff; a clean batch resets it. The async-void timer callback is fully
guarded: a fault is logged and recorded into GetStatus() rather than lost as
an unobserved task exception.
Durability bound (important)
The durability guarantee is bounded by capacity (default 1,000,000 rows).
When the non-dead-lettered queue reaches capacity, EnqueueAsync evicts the
oldest non-dead-lettered rows (oldest enqueued_at_utc first) to make room, logs a WARN,
and increments HistorianSinkStatus.EvictedCount. Under a sustained historian
outage, accepted alarm events can therefore be dropped before delivery. A
non-zero EvictedCount is a data-loss signal that requires operator attention —
it surfaces silent loss without log scraping.
Dead-letter + operator recovery
PermanentFail and corrupt-payload rows are retained in-place with
dead_lettered = 1 for the retention window (default 30 days) so operators can
inspect them before the sweeper purges them. RetryDeadLettered() is the
operator action (from the AdminUI) that clears the dead-letter flag and attempt
count on every dead-lettered row, returning them to the regular queue with a
fresh backoff.
Runtime wiring
HistorianAdapterActor
(Runtime/Historian/HistorianAdapterActor.cs)
subscribes to the cluster alerts DPS topic and translates each
AlarmTransitionEvent into an AlarmHistorianEvent, then calls
IAlarmHistorianSink.EnqueueAsync fire-and-forget so the actor loop is never
blocked on historian reachability. The actor is Primary-gated: only the
node whose RedundancyRole is Primary historizes, giving exactly-once
writes across a redundant pair. AlarmTransitionEvent carries AlarmTypeName
(the Part 9 subtype string) and Comment (the operator comment from the
originating ack/shelve command) that populate the corresponding fields of
AlarmHistorianEvent. GatewayAlarmHistorianWriter is the IAlarmHistorianWriter
the drain worker delegates to (the gateway SendEvent path). See
ServiceHosting.md for the (external) HistorianGateway setup.
Scope: scripted alarms only. Galaxy-native alarms historize via System
Platform's HistorizeToAveva toggle (not this actor); AB CIP ALMD is not on
the alerts topic (future).
Configuration
The real sink is opt-in via the AlarmHistorian section of appsettings.json.
When Enabled is false (the default), AddAlarmHistorian registers
NullAlarmHistorianSink and the feature is dormant. When Enabled is true,
AddAlarmHistorian constructs LocalDbStoreAndForwardSink and registers
GatewayAlarmHistorianWriter as the IAlarmHistorianWriter. This section carries
only the Enabled gate + the store-and-forward knobs — the downstream
gateway connection (endpoint / key / TLS) is sourced from the ServerHistorian
section (see Historian.md), and the queue's storage location is the
node's consolidated LocalDb:Path database (see
Configuration.md).
Where the queue lives (changed). The buffer is a table —
alarm_sf_events— in the node's consolidated LocalDb, not a standalonealarm-historian.db. That is what lets it replicate to the redundant pair peer, so undelivered alarm history survives losing the node holding it. Two consequences follow, both covered in the pair-replication runbook: only the Primary drains (the Secondary holds a full replica and would otherwise re-send everything), and delivery is at-least-once across a failover by design. On first boot after the upgrade,AlarmSfLegacyMigratorcopies any existingalarm-historian.dbacross and renames it.migrated.
{
"AlarmHistorian": {
"Enabled": true,
"BatchSize": 100,
"DrainIntervalSeconds": 5,
"Capacity": 1000000,
"DeadLetterRetentionDays": 30
}
}
| Key | Type | Default | Description |
|---|---|---|---|
Enabled |
bool | false |
Enable the durable store-and-forward sink (drains to the HistorianGateway SendEvent path). false → NullAlarmHistorianSink. Requires LocalDb to be registered — an enabled sink with no ILocalDb fails at resolution rather than silently discarding events. |
BatchSize |
int | 100 |
Max rows per drain cycle handed to IAlarmHistorianWriter.WriteBatchAsync. |
DrainIntervalSeconds |
int | 5 |
Seconds between drain-worker ticks. |
Capacity |
long | 1000000 |
Max queued rows before the sink evicts the oldest (data-loss signal via EvictedCount). |
DeadLetterRetentionDays |
int | 30 |
Days to retain dead-lettered rows before purge. |
MaxAttempts |
int | 10 |
Maximum delivery attempts before a poison (perpetually-retrying) row is dead-lettered automatically. Must be > 0. |
The downstream gateway connection lives in
ServerHistorian(Endpoint+ envServerHistorian__ApiKey,UseTls,CaCertificatePath); alarm-historyReadEventsadditionally requires the gateway runningRuntimeDb:EventReadsEnabled=true. The old Wonderware connection keys (SharedSecret/AlarmHistorian:Host/Port/UseTls/ServerCertThumbprint) were pruned.
DatabasePathwas removed. The queue no longer owns a file. The key is still read, byAlarmSfLegacyMigrator, purely to locate a pre-consolidationalarm-historian.dbto migrate — so leave it in place through the upgrade, then delete it once every node has analarm-historian.db.migratedsidecar. A leftover value configures nothing.
DrainStategainedNotPrimary. A gated-off Secondary reports it instead ofIdle, so a rising queue depth there is distinguishable from a stalled drain. Both nodes reporting it means nobody is draining — see the runbook.
Dev and docker-dev deployments leave
Enabledunset (defaults tofalse) so alarm transitions historize to nowhere unless a HistorianGateway is configured.
See also
- AlarmTracking.md — the three alarm sources and the OPC UA Part 9 surface; which alarms route to this sink.
- DriverLifecycle.md —
IHistorianDataSource(the historian read surface; this page covers the write path) and theGatewayHistorianDataSource. - ScriptedAlarms.md — the scripted-alarm engine that emits most events into this sink.
- ServiceHosting.md — the external HistorianGateway backend.
- operations/2026-07-20-localdb-pair-replication.md — where the queue lives, the Primary-only drain, and the one-time legacy migration.