feat(alarms): structural degraded-status signal for truncated alarm snapshots
The truncation-cliff fix made alarm transitions truncation-safe but silent:
when GetXmlCurrentAlarms2 returns exactly maxAlmCnt records the worker
suppresses absence-implies-Clear inference and says so only in a rate-limited
stderr warning. No client and no operator could tell a complete active set
from a capped one.
Two additive proto3 booleans carry the verdict out:
- QueryActiveAlarmsReplyPayload.snapshot_truncated = 2 (worker IPC reply)
- ActiveAlarmSnapshot.from_truncated_snapshot = 16 (per record)
The per-record field is not an aesthetic choice. QueryActiveAlarms returns a
bare `stream ActiveAlarmSnapshot` with no envelope, header, or trailer, so a
per-record boolean is the only carrier that stays wire-compatible; an envelope
message would change every existing client's stream element type. The reply
payload states it too because a prefix filter can leave zero records and a
truncated fetch with nothing to report still has to say so. The flag means
"this set may be incomplete", never "this record is unreliable" — it is
independent of the subtag-fallback `degraded` field.
Detection is deliberately UNCHANGED: IsTruncatedFetch remains
`fetchedRecordCount >= maxAlarmsPerFetch`. The live probe (docs/AlarmProbeFindings.md,
ce5d8ae) could not verify whether ALARM_RECORDS/@COUNT reports the total active
count or only the records in the reply, so @COUNT is not parsed for detection;
switching to it stays blocked on probe evidence. The probe's comment
annotations in WnWrapAlarmConsumer.cs are preserved.
Reset semantics: not latched. WnWrapAlarmConsumer.FoldFetch replaces the
verdict on every poll under the same lock as the snapshot merge, so the first
sub-cap fetch clears it; GatewayAlarmMonitor.ClearCache drops it with the cache
generation it describes. A caveat that never turns off is one operators learn
to ignore.
Flow: WnWrapAlarmConsumer.LastSnapshotTruncated -> AlarmDispatcher (stamps every
record) / IAlarmCommandHandler (payload) -> MxAccessCommandExecutor reply ->
GatewayAlarmMonitor._snapshotTruncated -> IGatewayAlarmService.SnapshotTruncated
-> DashboardAlarmQueryResult -> AlarmsPage warning banner (render-side only; the
poll loop and DisposeAsync drain are untouched). The public QueryActiveAlarms
RPC forwards worker snapshots unmodified, so the per-record flag needed no
mapper change — a test pins that.
Parity: this describes OUR fetch mechanics — additive gateway metadata — not
MXAccess provider behavior. No event is synthesized and no MXAccess-observable
semantics change, so it is not a parity deviation.
Tests: worker LastSnapshotTruncated set/reset/consecutive-burst (windev-run);
gateway end-to-end truncated reply -> monitor -> public stream, with the
complete-reply control as the load-bearing assertion; AlarmsPage banner
present/absent. Docs: gateway.md alarm surface, docs/DesignDecisions.md entry.
This commit is contained in:
@@ -726,6 +726,13 @@ message AcknowledgeAlarmReplyPayload {
|
||||
// stream.
|
||||
message QueryActiveAlarmsReplyPayload {
|
||||
repeated ActiveAlarmSnapshot snapshots = 1;
|
||||
// True when the provider fetch backing this reply came back holding the
|
||||
// per-fetch cap (MxGateway:Alarms:MaxAlarmsPerFetch). The reply may then omit
|
||||
// active alarms, and the worker suspends its absence-implies-Clear inference
|
||||
// for that poll — so a reference missing from `snapshots` is not evidence the
|
||||
// alarm cleared. Carried on the payload as well as per-record because a
|
||||
// truncated fetch that filters down to zero records still has to say so.
|
||||
bool snapshot_truncated = 2;
|
||||
}
|
||||
|
||||
message MxEvent {
|
||||
@@ -932,6 +939,16 @@ message ActiveAlarmSnapshot {
|
||||
// OnAlarmTransitionEvent.source_provider; always ALARMMGR or SUBTAG on the
|
||||
// wire (never UNSPECIFIED).
|
||||
AlarmProviderMode source_provider = 15;
|
||||
// True when the provider fetch that produced this snapshot hit the per-fetch
|
||||
// cap: the snapshot set may omit active alarms, and the worker suspended its
|
||||
// absence-implies-Clear inference for that poll. Says nothing about THIS
|
||||
// record's fidelity — the record is as accurate as any other; it flags that
|
||||
// the set it belongs to is possibly incomplete. QueryActiveAlarms returns a
|
||||
// bare `stream ActiveAlarmSnapshot` with no envelope message, so a per-record
|
||||
// boolean is the only additive way to carry set-level degraded status on that
|
||||
// RPC. Distinct from `degraded`, which is about the subtag fallback provider.
|
||||
// Additive (proto3): clients that ignore it deserialize the stream unchanged.
|
||||
bool from_truncated_snapshot = 16;
|
||||
}
|
||||
|
||||
enum AlarmConditionState {
|
||||
|
||||
Reference in New Issue
Block a user