f76d1f91e5
Enables the alarm store-and-forward sink on all four driver nodes of the docker-dev rig -- the replicating site-a pair and the default-OFF site-b pin -- so the live gate has a real buffer to watch converge, and can see that site-b's sink works as a plain node-local queue with no peer traffic. Two rig details worth stating rather than rediscovering. The endpoint is deliberately unresolvable: there is no HistorianGateway here, so every drain attempt fails, which is exactly the historian-outage state the buffer exists for. But it still has to be a syntactically valid absolute http(s) URI or the host refuses to start, because ServerHistorianOptionsValidator is consumer-gated -- AlarmHistorian:Enabled=true makes the endpoint required even while ServerHistorian:Enabled=false. And MaxAttempts is raised far above the production default of 10, which against a permanently unreachable gateway would dead-letter the entire queue about five minutes in, turning a buffering test into a dead-letter test. Docs record what an operator now has to know: that a rising queue depth on a Secondary is correct and on BOTH nodes is not (that shape is a redundancy snapshot naming neither node -- check node identity before suspecting the sink), that delivery is at-least-once across a failover by design with no dedup layer, and that AlarmHistorian:DatabasePath is removed but should be left in place through the upgrade because the migrator still reads it to find the file to copy. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
291 lines
15 KiB
Markdown
291 lines
15 KiB
Markdown
# Alarm Historian — store-and-forward SQLite sink
|
|
|
|
Reference for `ZB.MOM.WW.OtOpcUa.Core.AlarmHistorian`
|
|
([`src/Core/ZB.MOM.WW.OtOpcUa.Core.AlarmHistorian/`](../src/Core/ZB.MOM.WW.OtOpcUa.Core.AlarmHistorian/)),
|
|
the durable local queue that historizes alarm transitions to AVEVA Historian
|
|
without ever blocking the alarm engine or operator actions.
|
|
|
|
This is the *sink mechanics* doc. For how the three alarm sources converge on
|
|
the OPC UA Part 9 surface and which alarms route here, see
|
|
[AlarmTracking.md](AlarmTracking.md). For the historian client that drains this
|
|
queue, see [DriverLifecycle.md](DriverLifecycle.md#ihistoriandatasource--server-side-historian-read-surface)
|
|
and [ServiceHosting.md](ServiceHosting.md).
|
|
|
|
---
|
|
|
|
## Why store-and-forward
|
|
|
|
Scripted alarms (and any future non-Galaxy `IAlarmSource`, e.g. AB CIP ALMD)
|
|
must reach AVEVA Historian, but the historian gateway can be slow, busy, or
|
|
disconnected. The sink decouples the alarm engine from historian reachability:
|
|
every qualifying transition is committed to a **local SQLite queue first**, and
|
|
a background drain worker forwards rows to the historian on a backoff-aware
|
|
cadence. Operator acks and alarm-state transitions are never blocked waiting on
|
|
the historian.
|
|
|
|
> Galaxy-native alarms with `$Alarm*` extensions reach AVEVA Historian directly
|
|
> via System Platform's `HistorizeToAveva` toggle — they do **not** flow through
|
|
> this sink. This path is exclusively for non-Galaxy alarm producers.
|
|
|
|
---
|
|
|
|
## Contracts
|
|
|
|
All in
|
|
[`IAlarmHistorianSink.cs`](../src/Core/ZB.MOM.WW.OtOpcUa.Core.AlarmHistorian/IAlarmHistorianSink.cs)
|
|
unless noted.
|
|
|
|
- **`IAlarmHistorianSink`** — the intake contract. `EnqueueAsync(evt, ct)`
|
|
durably enqueues an event and returns as soon as the queue row is committed
|
|
(fire-and-forget from the engine's perspective; the sink must not block the
|
|
emitting thread). `GetStatus()` returns a `HistorianSinkStatus` snapshot.
|
|
- **`NullAlarmHistorianSink`** — the no-op default for tests and deployments
|
|
that don't historize alarms. It is the default DI binding (registered in the
|
|
Runtime's `AddOtOpcUaRuntime`); production overrides it with
|
|
`LocalDbStoreAndForwardSink`.
|
|
- **`AlarmHistorianEvent`**
|
|
([`AlarmHistorianEvent.cs`](../src/Core/ZB.MOM.WW.OtOpcUa.Core.AlarmHistorian/AlarmHistorianEvent.cs))
|
|
— the source-agnostic event record: `AlarmId`, `EquipmentPath` (UNS path,
|
|
doubles as Historian's SourceNode), `AlarmName`, `AlarmTypeName` (Part 9
|
|
subtype), `Severity`, `EventKind` (free-form transition string —
|
|
"Activated"/"Cleared"/"Acknowledged"/etc.), `Message`, `User`, `Comment`,
|
|
`TimestampUtc`.
|
|
- **`IAlarmHistorianWriter`** — what the drain worker delegates writes to.
|
|
`WriteBatchAsync(batch, ct)` returns one `HistorianWriteOutcome` per event,
|
|
in order. Production binds this to `GatewayAlarmHistorianWriter` (the
|
|
HistorianGateway `SendEvent` path).
|
|
- **`HistorianWriteOutcome`** — per-event drain result: `Ack` (persisted,
|
|
remove from queue), `RetryPlease` (transient failure — leave queued, retry
|
|
after backoff), `PermanentFail` (malformed/unrecoverable — move to
|
|
dead-letter).
|
|
- **`HistorianSinkStatus`** — diagnostic snapshot surfaced to the AdminUI and
|
|
`/healthz`: `QueueDepth`, `DeadLetterDepth`, `LastDrainUtc`, `LastSuccessUtc`,
|
|
`LastError`, `DrainState`, and `EvictedCount`.
|
|
- **`HistorianDrainState`** — `Disabled` / `Idle` / `Draining` / `BackingOff` /
|
|
**`NotPrimary`** (ticking, but leaving the replicated queue for the node that
|
|
holds the Primary role).
|
|
|
|
---
|
|
|
|
## LocalDbStoreAndForwardSink
|
|
|
|
[`LocalDbStoreAndForwardSink.cs`](../src/Core/ZB.MOM.WW.OtOpcUa.Core.AlarmHistorian/LocalDbStoreAndForwardSink.cs)
|
|
is the production `IAlarmHistorianSink`. Construction takes the node's
|
|
`ILocalDb`, an `IAlarmHistorianWriter`, a logger, and optional `batchSize`
|
|
(default 100), `capacity` (default 1,000,000), `deadLetterRetention` (default 30
|
|
days), a test clock, and a **`drainGate`**.
|
|
|
|
### The drain gate
|
|
|
|
`drainGate` is a `Func<bool>` consulted at the top of every tick; `false` skips
|
|
the tick entirely, leaving the queue untouched and reporting
|
|
`DrainState = NotPrimary`.
|
|
|
|
It exists because the queue **replicates**. The Secondary holds a full copy of
|
|
the Primary's rows, and an ungated drain there would re-deliver every event,
|
|
continuously — not just at a failover. Runtime supplies the gate from
|
|
`IRedundancyRoleView`, a singleton `DriverHostActor` publishes its
|
|
`PrimaryGatePolicy` verdict to on every redundancy snapshot, so the drain agrees
|
|
with the inbound-write and native-ack gates by construction.
|
|
|
|
Two postures are deliberate:
|
|
|
|
- **Unset (the default) means drain.** So does an unpublished view. A deployment
|
|
that runs no redundancy never publishes a role, and defaulting closed would
|
|
silently stop its alarm history forever.
|
|
- **A gate that throws means "not now", never permission.** Draining on a gate
|
|
fault would put a pair into dual delivery exactly when its redundancy signal is
|
|
sick.
|
|
|
|
### Queue table
|
|
|
|
The sink owns one table in the node's consolidated LocalDb — created and
|
|
registered for replication by `LocalDbSetup.OnReady`, before any consumer can
|
|
resolve the sink:
|
|
|
|
```sql
|
|
CREATE TABLE alarm_sf_events (
|
|
id TEXT NOT NULL PRIMARY KEY, -- SHA-256 of payload_json
|
|
alarm_id TEXT NOT NULL,
|
|
enqueued_at_utc TEXT NOT NULL,
|
|
payload_json TEXT NOT NULL, -- JSON-serialized AlarmHistorianEvent
|
|
attempt_count INTEGER NOT NULL DEFAULT 0,
|
|
last_attempt_utc TEXT NULL,
|
|
last_error TEXT NULL,
|
|
dead_lettered INTEGER NOT NULL DEFAULT 0
|
|
);
|
|
CREATE INDEX ix_alarm_sf_events_drain
|
|
ON alarm_sf_events (dead_lettered, enqueued_at_utc, id);
|
|
```
|
|
|
|
**The primary key is a hash of the payload, not an autoincrement rowid.** Two
|
|
reasons. Replication converges last-writer-wins *per primary key*, so two nodes
|
|
independently allocating rowid 7 to different alarms would silently overwrite one
|
|
another. And an equal-payload key makes the same event arriving twice collapse
|
|
into one row — which happens for real, because `HistorianAdapterActor`
|
|
default-writes while its redundancy role is unknown and both nodes of a pair
|
|
therefore accept the same fanned transition in every boot window. Two genuinely
|
|
distinct events cannot collide: `AlarmHistorianEvent` carries a full-precision
|
|
timestamp alongside the alarm id, kind, message and user.
|
|
|
|
Ordering follows from that: the drain reads in `enqueued_at_utc` order (with `id`
|
|
as a tiebreak so the ordering is total), because a hashed key carries no
|
|
insertion sequence.
|
|
|
|
`EnqueueAsync` does a single `INSERT … ON CONFLICT DO NOTHING` on the hot path.
|
|
To avoid a `SELECT COUNT(*)` on every enqueue, the sink keeps an in-memory
|
|
non-dead-lettered row counter (seeded at startup, kept current by every mutation,
|
|
and re-synced from storage every 10,000 enqueues to defend against drift — now
|
|
doubly warranted, since replication applies the peer's rows straight into the
|
|
table where no local code path can observe them). Pragmas and connection
|
|
lifetime belong to LocalDb; the sink no longer manages either.
|
|
|
|
### Drain worker
|
|
|
|
`StartDrainLoop(tickInterval)` starts a **self-rescheduling one-shot
|
|
`System.Threading.Timer`** (not started automatically — tests drive
|
|
`DrainOnceAsync` deterministically). Each tick:
|
|
|
|
0. Consults the drain gate. A closed gate returns immediately, leaving the queue
|
|
untouched and `DrainState = NotPrimary`.
|
|
1. Purges aged dead-lettered rows past the retention window.
|
|
2. Reads up to `batchSize` non-dead-lettered rows in `enqueued_at_utc, id` order.
|
|
3. Rows with un-deserializable payloads are dead-lettered immediately (by their
|
|
own `id`) so they can't stall the queue head.
|
|
4. The remaining batch is handed to `IAlarmHistorianWriter.WriteBatchAsync`, and
|
|
each outcome is applied in one transaction: `Ack` deletes the row (the delete
|
|
replicates as a tombstone, so the peer drops its copy and a promoted standby
|
|
cannot re-send it), `PermanentFail` flips its `dead_lettered` flag,
|
|
`RetryPlease` bumps its attempt
|
|
count and leaves it queued. A row whose `AttemptCount` has reached the configured
|
|
**`MaxAttempts`** cap (default 10) is dead-lettered automatically on the next drain
|
|
tick rather than retried — this breaks infinite retry loops for poison events whose
|
|
payload the historian will always reject (e.g. a malformed alarm record that triggers
|
|
a permanent SDK error on every attempt). The dead-lettered row remains inspectable
|
|
via `RetryDeadLettered()` for the configured retention window.
|
|
5. The timer re-arms its next due-time to `max(tickInterval, currentBackoff)`.
|
|
|
|
**Backoff ladder** (applied to the timer's next due-time, so a historian outage
|
|
genuinely slows the drain cadence): 1s → 2s → 5s → 15s → 60s cap. Any
|
|
`RetryPlease` outcome — or a writer exception, or a writer cardinality violation
|
|
(outcome count ≠ event count) — bumps the backoff and sets `DrainState =
|
|
BackingOff`; a clean batch resets it. The async-void timer callback is fully
|
|
guarded: a fault is logged and recorded into `GetStatus()` rather than lost as
|
|
an unobserved task exception.
|
|
|
|
### Durability bound (important)
|
|
|
|
**The durability guarantee is bounded by `capacity` (default 1,000,000 rows).**
|
|
When the non-dead-lettered queue reaches capacity, `EnqueueAsync` evicts the
|
|
oldest non-dead-lettered rows (oldest `enqueued_at_utc` first) to make room, logs a WARN,
|
|
and increments `HistorianSinkStatus.EvictedCount`. Under a sustained historian
|
|
outage, accepted alarm events can therefore be dropped before delivery. A
|
|
non-zero `EvictedCount` is a data-loss signal that requires operator attention —
|
|
it surfaces silent loss without log scraping.
|
|
|
|
### Dead-letter + operator recovery
|
|
|
|
`PermanentFail` and corrupt-payload rows are retained in-place with
|
|
`dead_lettered = 1` for the retention window (default 30 days) so operators can
|
|
inspect them before the sweeper purges them. `RetryDeadLettered()` is the
|
|
operator action (from the AdminUI) that clears the dead-letter flag and attempt
|
|
count on every dead-lettered row, returning them to the regular queue with a
|
|
fresh backoff.
|
|
|
|
---
|
|
|
|
## Runtime wiring
|
|
|
|
`HistorianAdapterActor`
|
|
([`Runtime/Historian/HistorianAdapterActor.cs`](../src/Server/ZB.MOM.WW.OtOpcUa.Runtime/Historian/HistorianAdapterActor.cs))
|
|
subscribes to the cluster **`alerts` DPS topic** and translates each
|
|
`AlarmTransitionEvent` into an `AlarmHistorianEvent`, then calls
|
|
`IAlarmHistorianSink.EnqueueAsync` fire-and-forget so the actor loop is never
|
|
blocked on historian reachability. The actor is **Primary-gated**: only the
|
|
node whose `RedundancyRole` is `Primary` historizes, giving exactly-once
|
|
writes across a redundant pair. `AlarmTransitionEvent` carries `AlarmTypeName`
|
|
(the Part 9 subtype string) and `Comment` (the operator comment from the
|
|
originating ack/shelve command) that populate the corresponding fields of
|
|
`AlarmHistorianEvent`. `GatewayAlarmHistorianWriter` is the `IAlarmHistorianWriter`
|
|
the drain worker delegates to (the gateway `SendEvent` path). See
|
|
[ServiceHosting.md](ServiceHosting.md) for the (external) HistorianGateway setup.
|
|
|
|
**Scope:** scripted alarms only. Galaxy-native alarms historize via System
|
|
Platform's `HistorizeToAveva` toggle (not this actor); AB CIP ALMD is not on
|
|
the `alerts` topic (future).
|
|
|
|
## Configuration
|
|
|
|
The real sink is opt-in via the `AlarmHistorian` section of `appsettings.json`.
|
|
When `Enabled` is `false` (the default), `AddAlarmHistorian` registers
|
|
`NullAlarmHistorianSink` and the feature is dormant. When `Enabled` is `true`,
|
|
`AddAlarmHistorian` constructs `LocalDbStoreAndForwardSink` and registers
|
|
`GatewayAlarmHistorianWriter` as the `IAlarmHistorianWriter`. This section carries
|
|
**only** the `Enabled` gate + the store-and-forward knobs — the downstream
|
|
gateway connection (endpoint / key / TLS) is sourced from the `ServerHistorian`
|
|
section (see [Historian.md](Historian.md)), and the queue's storage location is the
|
|
node's consolidated `LocalDb:Path` database (see
|
|
[Configuration.md](Configuration.md)).
|
|
|
|
> **Where the queue lives (changed).** The buffer is a table — `alarm_sf_events` — in the node's
|
|
> consolidated LocalDb, not a standalone `alarm-historian.db`. That is what lets it **replicate to
|
|
> the redundant pair peer**, so undelivered alarm history survives losing the node holding it.
|
|
> Two consequences follow, both covered in the
|
|
> [pair-replication runbook](operations/2026-07-20-localdb-pair-replication.md):
|
|
> **only the Primary drains** (the Secondary holds a full replica and would otherwise re-send
|
|
> everything), and **delivery is at-least-once across a failover** by design. On first boot after
|
|
> the upgrade, `AlarmSfLegacyMigrator` copies any existing `alarm-historian.db` across and renames
|
|
> it `.migrated`.
|
|
|
|
```json
|
|
{
|
|
"AlarmHistorian": {
|
|
"Enabled": true,
|
|
"BatchSize": 100,
|
|
"DrainIntervalSeconds": 5,
|
|
"Capacity": 1000000,
|
|
"DeadLetterRetentionDays": 30
|
|
}
|
|
}
|
|
```
|
|
|
|
| Key | Type | Default | Description |
|
|
|---|---|---|---|
|
|
| `Enabled` | bool | `false` | Enable the durable store-and-forward sink (drains to the HistorianGateway `SendEvent` path). `false` → `NullAlarmHistorianSink`. Requires LocalDb to be registered — an enabled sink with no `ILocalDb` fails at resolution rather than silently discarding events. |
|
|
| `BatchSize` | int | `100` | Max rows per drain cycle handed to `IAlarmHistorianWriter.WriteBatchAsync`. |
|
|
| `DrainIntervalSeconds` | int | `5` | Seconds between drain-worker ticks. |
|
|
| `Capacity` | long | `1000000` | Max queued rows before the sink evicts the oldest (data-loss signal via `EvictedCount`). |
|
|
| `DeadLetterRetentionDays` | int | `30` | Days to retain dead-lettered rows before purge. |
|
|
| `MaxAttempts` | int | `10` | Maximum delivery attempts before a poison (perpetually-retrying) row is dead-lettered automatically. Must be > 0. |
|
|
|
|
> The downstream gateway connection lives in `ServerHistorian` (`Endpoint` + env `ServerHistorian__ApiKey`,
|
|
> `UseTls`, `CaCertificatePath`); alarm-history `ReadEvents` additionally requires the gateway running
|
|
> `RuntimeDb:EventReadsEnabled=true`. The old Wonderware connection keys (`SharedSecret` /
|
|
> `AlarmHistorian:Host`/`Port`/`UseTls`/`ServerCertThumbprint`) were pruned.
|
|
|
|
> **`DatabasePath` was removed.** The queue no longer owns a file. The key is still *read*, by
|
|
> `AlarmSfLegacyMigrator`, purely to locate a pre-consolidation `alarm-historian.db` to migrate —
|
|
> so leave it in place through the upgrade, then delete it once every node has an
|
|
> `alarm-historian.db.migrated` sidecar. A leftover value configures nothing.
|
|
|
|
> **`DrainState` gained `NotPrimary`.** A gated-off Secondary reports it instead of `Idle`, so a
|
|
> rising queue depth there is distinguishable from a stalled drain. Both nodes reporting it means
|
|
> nobody is draining — see the runbook.
|
|
|
|
> Dev and docker-dev deployments leave `Enabled` unset (defaults to `false`) so alarm transitions historize to nowhere unless a HistorianGateway is configured.
|
|
|
|
---
|
|
|
|
## See also
|
|
|
|
- [AlarmTracking.md](AlarmTracking.md) — the three alarm sources and the OPC UA
|
|
Part 9 surface; which alarms route to this sink.
|
|
- [DriverLifecycle.md](DriverLifecycle.md) — `IHistorianDataSource` (the
|
|
historian *read* surface; this page covers the *write* path) and the
|
|
`GatewayHistorianDataSource`.
|
|
- [ScriptedAlarms.md](ScriptedAlarms.md) — the scripted-alarm engine that emits
|
|
most events into this sink.
|
|
- [ServiceHosting.md](ServiceHosting.md) — the external HistorianGateway backend.
|
|
- [operations/2026-07-20-localdb-pair-replication.md](operations/2026-07-20-localdb-pair-replication.md)
|
|
— where the queue lives, the Primary-only drain, and the one-time legacy migration.
|