Files
lmxopcua/docs/Telemetry.md
T
Joseph Doherty 30d0697c28 feat(hosts): surface per-host connectivity on /hosts; drop the unwritable table (#521)
`IHostConnectivityProbe` was a dead surface: eleven drivers implement it, `GetHostStatuses()`
had ZERO production call sites, and `OnHostStatusChanged` had no subscriber outside the Galaxy
driver's own aggregator. Per-host connectivity was computed by every driver and read by nobody.

The issue offered "build the publisher or delete it". The publisher as its entity doc described
it — driver nodes upserting `DriverHostStatus` rows — is not buildable: per-cluster mesh Phase 4
gates `AddOtOpcUaConfigDb` on the `admin` role, so a driver-only node has no ConfigDb connection
to write rows with. So the capability is kept and the transport changed.

`DriverHealthChanged.HostStatuses` now carries the probe result to `/hosts` as a Hosts column.
That channel already reached the page, already survives the mesh split via the Phase 5 gRPC
telemetry stream, and already replays a last-value snapshot on re-subscribe — so per-host state
re-primes after a reconnect without a durable store. Both halves of the interface finally do
what they are for: the event triggers a prompt publish, the pull is the source of truth.

The point of the column is the case the driver-level state chip structurally cannot express: a
multi-device driver stays aggregate-Healthy while ONE of its devices is unreachable.

Two traps, both pinned by tests that were falsified against the prod code:

- The host digest MUST be in the publish fingerprint. On a single-host-down transition every
  other fingerprint component is unchanged, so the dedup would swallow exactly the publish
  carrying the news — the trap that already bit the rediscovery signal. Removing it turns the
  guard test red, verified.
- null (no probe) must stay distinct from empty (probe with no hosts). proto3 cannot tell an
  absent repeated field from an empty one, hence the explicit `has_host_statuses` flag;
  collapsing them would render every probe-less driver as one whose devices are all fine.

Dropped: the DriverHostStatus entity, enum, DbSet, model config and table (migration
DropDriverHostStatusTable — empty on every deployment, so the scaffolder's data-loss warning is
moot, and Down() recreates it exactly).

Found en route, NOT fixed here: `DriverInstanceResilienceStatus` is the identical defect — no
writer, no reader, only a DbSet declaration, while the live data rides the
`driver-resilience-status` telemetry channel. Its doc-comment now states that rather than
describing the sampler and AdminUI join that were never built. Filed as #524 rather than widening
this schema change beyond what was asked.

Claude-Session: https://claude.ai/code/session_015p7wGqy3YpZNCpDzTpGMKo
2026-07-30 04:21:08 -04:00

17 KiB

Live Telemetry Transport (v2)

Overview

The AdminUI's live observability panels (/alerts, /script-log, /hosts driver table, driver resilience status) are fed by four node→central event channels. Historically all four rode DistributedPubSub (DPS) on the shared Akka mesh — the same gossip ring that carries redundancy state and the command-plane topics. Per-cluster mesh Phase 5 adds a second transport for these four channels only: one gRPC server-streaming contract, selected per node/central pair by Telemetry:Mode / TelemetryDial:Mode (Dps default | Grpc).

This is the same motivation as Phase 2's MeshTransport and Phase 3's ConfigSource/ConfigServe: DPS only works when central and the node share a gossip ring, and Phase 6 splits the fleet into one Akka mesh per application Cluster. Once that split lands, a driver node in a site's mesh is no longer a cluster member of central's mesh, so DPS can no longer reach it — telemetry needs its own out-of-band transport, same as commands (Phase 2) and config bytes (Phase 3).

Direction: the node hosts, central dials

Load-bearing and easy to get backwards: each driver node hosts the telemetry gRPC server (a dedicated Kestrel h2c listener); central (admin role) is the client and dials in. This mirrors the inversion ScadaBridge already uses for the same problem (its SiteStreamService).

 driver node                          central (admin)
┌─────────────────────┐              ┌──────────────────────────┐
│ 4 publish seams      │              │ TelemetryDialSupervisor   │
│  -> node-local hub    │   Subscribe  │  - discovers ClusterNode   │
│  -> DPS (unchanged)   │◄─────────────│    rows (Host+GrpcPort)   │
│                       │   (gRPC      │  - one reconnecting       │
│ TelemetryStreamGrpc-  │    stream)   │    dialer per node        │
│  Service (h2c server) │─────────────►│  - feeds the SAME sinks   │
│  Telemetry:GrpcListen- │  TelemetryEvent │   the DPS bridges feed   │
│  Port                 │  (oneof, 4 kinds)│   today                  │
└─────────────────────┘              └──────────────────────────┘

A fused admin,driver node both hosts (as driver) and dials (as admin) — it dials itself plus its redundant pair peer, the same way central dials site nodes.

Central discovers driver-node telemetry endpoints from the ClusterNode table (Host + GrpcPort, the latter added in Phase 1 for exactly this purpose) — the same DB-sourced discovery Phase 1's ack set and Phase 2's ClusterClient contact set already use. No shared gossip membership is required to find or dial a node.

The dark switch

Telemetry:Mode (node/serve side) and TelemetryDial:Mode (central/dial side) each default to Dps and can be independently set to Grpc. Both code paths are compiled into every binary — flipping either flag is an appsettings/env change plus a restart, not a redeploy or rebuild — the same discipline as MeshTransport:Mode and ConfigSource:Mode.

What actually changes per mode, precisely:

  • The node ALWAYS does both things, regardless of Telemetry:Mode. Every one of the four publish seams emits into the node-local ITelemetryLocalHub and publishes to DPS, unconditionally. The node additionally always hosts the gRPC server whenever Telemetry:GrpcListenPort > 0 — hosting is not mode-gated at all. Telemetry:Mode only affects central's own bookkeeping about which side it expects to be consulted (it is otherwise inert on the node).
  • Only central's ingest source switches, driven by TelemetryDial:Mode:
    • Dps (default) — today's four DPS SignalR bridges (alert, script-log, driver-status, resilience) subscribe and feed the AdminUI sinks. Unchanged behavior.
    • Grpc — those four DPS bridges are not spawned. Instead the TelemetryDialSupervisor actor dials every driver node's gRPC stream and feeds the identical in-process sinks (IInProcessBroadcaster<AlarmTransitionEvent>, IInProcessBroadcaster<ScriptLogEntry>, IDriverStatusSnapshotStore, IDriverResilienceStatusStore).
  • The AdminUI panels themselves are untouched. /alerts, /script-log, and the /hosts driver table all read from the same sinks in both modes — only the sinks' upstream feed swaps. The fleet-status bridge (out of Phase 5 scope — see below) stays on DPS in both modes.

Because hosting is unconditional and both ingest paths are always compiled in, a fleet can be flipped node-by-node and central-by-central with no coordination window where telemetry is lost — the node is always serving, so central can switch to Grpc whenever it likes.

The four migrated channels

DPS topic (today) Domain record Central sink fed
alerts AlarmTransitionEvent IInProcessBroadcaster<AlarmTransitionEvent> (+ AlertHub); /alerts page; live/disconnected pill
script-logs ScriptLogEntry IInProcessBroadcaster<ScriptLogEntry> (+ ScriptLogHub); /script-log page; live pill
driver-health DriverHealthChanged IDriverStatusSnapshotStore; /hosts driver table
driver-resilience-status DriverResilienceStatusChanged IDriverResilienceStatusStore

These four are carried as oneof event kinds on one TelemetryStreamService.Subscribe RPC — a single stream per (central, driver-node) pair carries all four, rather than one stream per topic. Proto field evolution is additive-only (never renumber/reuse a tag), locked by a contract test that reflects over the oneof cases.

Per-host connectivity rides driver-health (Gitea #521)

DriverHealthChanged.HostStatuses carries each driver's IHostConnectivityProbe.GetHostStatuses() result, rendered as the Hosts column on /hosts. It exists to show the one thing the driver-level state chip structurally cannot: a multi-device driver stays aggregate-Healthy while one of its devices is unreachable. A FOCAS or AbLegacy instance owning several PLCs previously hid that entirely.

Three things to know before touching it:

  • This deliberately does NOT go through SQL. A DriverHostStatus table existed, with an entity doc describing a publisher hosted service that upserted rows from each driver node. That publisher was never written, and per-cluster mesh Phase 4 made it unbuildable as describedProgram.cs gates AddOtOpcUaConfigDb on the admin role, so a driver-only node has no ConfigDb connection to write rows with. The table was dropped (migration DropDriverHostStatusTable); it was empty on every deployment. This channel needs no DB and already survives the mesh split.
  • null ≠ empty, on the wire too. Null means "the driver has no probe"; an empty list means "it has one that currently knows no hosts". proto3 cannot distinguish an absent repeated field from an empty one, so DriverHealth.has_host_statuses carries that bit explicitly. Collapsing the two would render every probe-less driver as one whose devices are all fine.
  • The host digest is part of the publish fingerprint, and must stay there. PublishHealthSnapshot dedups on that fingerprint, and on a single-host-down transition every other component is unchanged — so without it the dedup swallows precisely the publish carrying the news. Same trap that bit the rediscovery signal; pinned by DriverInstanceActorHostStatusTests.A_single_host_going_down_is_not_swallowed_by_the_unchanged_health_dedup. The digest is a flattened, host-name-ordered string for the converse reason: a tuple holding an IReadOnlyList compares by reference and would re-publish on every 30 s heartbeat forever.

⚠️ DriverInstanceResilienceStatus is the same defect, still open. That table also has no writer and no reader — the only reference in the repo is its DbSet declaration — while the live data reaches the AdminUI over the driver-resilience-status channel above. Keep-or-delete is Gitea #524.

The three deferred channels (NOT migrated in Phase 5 — do not read this as "seven done")

The program sketch originally named seven observability topics for Phase 5. Three were scoped out, each for a distinct, settled reason:

  • redundancy-state — bidirectional, built directly from Cluster.State, and pair-local: it drives ServiceLevel and the Primary data-plane gate, consumed in-process by OpcUaPublishActor, ScriptedAlarmHostActor, DriverHostActor, and HistorianAdapterActor. It stays on DPS in both MeshTransport modes today, and it is genuinely mesh-bound by design — under Phase 6 each pair keeps sharing its own small mesh, so DPS keeps working for it in-mesh. Central's display of each pair's redundancy state is a Phase 6 cross-mesh concern (a future observability channel, once central no longer shares gossip with any site pair), not a Phase 5 telemetry-panel migration.
  • fleet-statuscentral-internal, not a node→central stream at all. The admin singleton FleetStatusBroadcaster builds it from the admin node's own cluster membership/reachability/ leader events, and Fleet.razor polls the Config DB and ignores the feed entirely. There is nothing here for a per-node gRPC stream to carry. Revisit in Phase 6, once central genuinely loses gossip visibility of site nodes and needs another way to know a pair's membership.
  • deployment-acks — already migrated, but onto a different transport: it rides the Phase 2 ClusterClient transport (MeshTransport:Mode=ClusterClient) as a command-plane reply, not an observability broadcast. It was never a Phase 5 candidate.

See docs/Redundancy.md for how redundancy-state and the command transports fit together.

Authentication — fail-closed from day one

A shared node bearer key gates the stream: Telemetry:ApiKey (node/serve side) must equal TelemetryDial:ApiKey (central/dial side). TelemetryStreamAuthInterceptor enforces it — path-scoped to /telemetry.v1.TelemetryStreamService/, comparing the Authorization: Bearer token with CryptographicOperations.FixedTimeEquals, and rejecting with PermissionDenied. An unset key rejects every call rather than allowing an open stream — the same fail-closed posture as ConfigServeAuthInterceptor and LocalDbSyncAuthInterceptor.

This supersedes the design doc's §6.3, which had provisionally decided to "match ScadaBridge's unauthenticated posture for now" for inter-cluster transports. ScadaBridge itself has since closed that gap with this identical interceptor pattern, so Phase 5 ships authenticated from the start rather than deferring auth to a later hardening pass. See the design doc's superseded note for the full history.

Upgrading an existing deployment: populate ClusterNode.GrpcPort first (#493)

GrpcPort is a nullable column added by Phase 5, and nothing backfills it onto rows that already existed. On an upgraded deployment every ClusterNode row therefore carries NULL, central finds no dial targets, and flipping TelemetryDial:Mode to Grpc connects to nothing. Fresh installs are unaffected. AkkaPort did not have this problem only because it is NOT NULL with a default of 4053.

Nothing fails loudly, which is what makes this worth stating up front — the code degrades gracefully, once per node, throttled:

ClusterNode <id> has no GrpcPort; it exposes no telemetry stream and is skipped in this dial refresh
(further skips of this node are silent)

Other symptoms: no ESTABLISHED sockets on the telemetry port (a LISTEN but nothing else in docker exec <node> cat /proc/net/tcp6), and the /alerts / /script-log pill never turning live in Grpc mode.

Before flipping any node to Grpc, set each driver node's GrpcPort to that node's own Telemetry:GrpcListenPort — via the AdminUI node editor, or directly:

UPDATE dbo.ClusterNode SET GrpcPort = <that node's Telemetry:GrpcListenPort> WHERE NodeId = '<node>';

There is deliberately no EF data migration backfilling this. The correct value is that node's own listener port, which is per-deployment configuration the database cannot derive; a migration could only invent one, and a wrong dial target is worse than the NULL the dialer already skips explicitly. The docker-dev seed does backfill (GrpcPort IS NULL AND CreatedBy = 'docker-dev-seed' → 4056) because there the port is fixed by docker-compose.yml and therefore actually knowable.

Reconnect and reliability

Central's TelemetryDialSupervisor (an admin-role actor) runs one reconnecting dialer per driver node:

  • Discovery — enabled, non-maintenance ClusterNode rows are read on start and refreshed every TelemetryDial:ContactRefreshSeconds (default 60), plus on an admin-side topology change. A row with a null GrpcPort is skipped (that node exposes no telemetry endpoint — logged, not an error). Dial targets are added/removed as the row set changes.
  • Reconnect backoff — immediate first retry, then a fixed ~5 second backoff, indefinitely. The dialer never gives up permanently: this is an observability channel, not the data plane, so a persistently-unreachable node just means a persistently-stale panel, not a fleet fault.
  • Generation stamping — each dialer carries a monotonically-increasing generation counter, so a late error or late event from a superseded stream (e.g. one that raced a reconnect) is recognized and ignored rather than corrupting the current stream's state.
  • Snapshot replay on reconnect — the node-local hub keeps a last-value cache for driver-health and driver-resilience-status (there is nothing meaningful to replay for alerts/script-logs, which are append logs). On every new Subscribe — including a reconnect — the hub drains the cached snapshots to the new stream before live deltas, so IDriverStatusSnapshotStore / IDriverResilienceStatusStore re-prime immediately rather than sitting stale until the next natural event. alerts/script-logs simply tolerate the gap — a reconnect loses whatever transitions happened while disconnected, same as a DPS resubscribe would.
  • Connection indicator — the /alerts and /script-log live/disconnected pill is driven by aggregate stream health across all dialers: connected when at least one node stream is up, disconnected when all are down. This matches today's DPS SubscribeAck/PostStop pill semantics — the pill has never meant "every node is live," only "the panel has a source." The two store-backed panels (driver-health, driver-resilience-status) carry no fleet-wide connection flag today, on gRPC or on DPS; per-row staleness is unchanged.

Configuration reference

See docs/Configuration.md § Telemetry / TelemetryDial for the full keys table. In brief:

  • Telemetry (node/serve side): Mode (Dps|Grpc, default Dps), GrpcListenPort (0 = disabled — the driver node's dedicated telemetry h2c port), ApiKey (the shared node key; supply via ${secret:}/env, never commit).
  • TelemetryDial (central side): Mode (Dps|Grpc), ApiKey (must equal the nodes' Telemetry:ApiKey), ContactRefreshSeconds (default 60), CallTimeoutSeconds (default 30).

A driver-role node with Telemetry:Mode=Grpc must set GrpcListenPort > 0 (nothing to serve otherwise), and Grpc mode on either side requires a non-empty ApiKey — both are enforced by ValidateOnStart validators, fail-fast at boot.

Relationship to the rest of the mesh program

  • Same dark-switch discipline as Phases 2 and 3. MeshTransport:Mode, ConfigSource:Mode, and Telemetry:Mode/TelemetryDial:Mode are all independently-flippable, restart-only config changes, and in every case the "new" side is always wired so the flip needs no coordinated redeploy.
  • Why the inversion matters for Phase 6. Every other cross-boundary transport in this program (MeshTransport, ConfigSource/ConfigServe) already has central as the addressable, discoverable side and the node as the caller or the callee-by-address. Telemetry is the one channel that is naturally node-sourced and fleet-wide-fanned, so putting central in the client role — dialing out to each node by its ClusterNode-recorded address — is what lets telemetry survive the mesh split cleanly: central never needs cluster membership with a site's mesh to keep watching it.
  • Status: code-complete on feat/mesh-phase5; the live gate (flip the docker-dev rig to Grpc, confirm all panels stay green with the DPS telemetry bridges unspawned, kill-and-reconnect a site node) has not yet run. See docs/plans/2026-07-22-per-cluster-mesh-program.md § Phase 5 and docs/plans/2026-07-23-mesh-phase5-grpc-telemetry-stream.md for the implementation plan.