8b7ea3213d
`GrpcPort` is a nullable column added by Phase 5. Every insert in the docker-dev seed is INSERT-if-not-exists — the right shape for a re-runnable seed, but it means a new nullable column never reaches rows an earlier version of the file created. On a persisted volume all six rows kept NULL, central found no dial targets, and `TelemetryDial:Mode=Grpc` connected to nothing. (`AkkaPort` escaped this only because it is NOT NULL with a default of 4053.) Seed now backfills `GrpcPort IS NULL AND CreatedBy = 'docker-dev-seed'` to 4056. NULL only — it must not overwrite a port an operator set deliberately on a long-lived dev volume, and NULL is the unambiguous "predates the column" marker. Deliberately NOT doing the suggested EF data migration for real deployments. The correct value is each node's own `Telemetry:GrpcListenPort`, which is per-deployment configuration the database cannot derive; a migration could only invent one, and a WRONG dial target is worse than the NULL the dialer already skips explicitly with a throttled warning. The rig can be backfilled because there the port is fixed by docker-compose.yml and therefore actually knowable. `docs/Telemetry.md` gains an upgrade section stating the prerequisite, the symptoms (the throttled skip log, LISTEN with no ESTABLISHED, the pill never going live) and the UPDATE to run before flipping any node to Grpc — plus why there is no migration.
224 lines
15 KiB
Markdown
224 lines
15 KiB
Markdown
# Live Telemetry Transport (v2)
|
|
|
|
## Overview
|
|
|
|
The AdminUI's live observability panels (`/alerts`, `/script-log`, `/hosts` driver table, driver
|
|
resilience status) are fed by four node→central event channels. Historically all four rode
|
|
DistributedPubSub (DPS) on the shared Akka mesh — the same gossip ring that carries redundancy
|
|
state and the command-plane topics. Per-cluster mesh **Phase 5** adds a second transport for these
|
|
four channels only: one gRPC server-streaming contract, selected per node/central pair by
|
|
`Telemetry:Mode` / `TelemetryDial:Mode` (`Dps` default | `Grpc`).
|
|
|
|
This is the same motivation as [Phase 2's `MeshTransport`](Configuration.md#meshtransport-centralnode-command-transport)
|
|
and [Phase 3's `ConfigSource`/`ConfigServe`](Configuration.md#configsource--configserve-config-fetch-and-cache):
|
|
DPS only works when central and the node share a gossip ring, and [Phase 6](plans/2026-07-21-per-cluster-mesh-design.md)
|
|
splits the fleet into one Akka mesh per application `Cluster`. Once that split lands, a driver node
|
|
in a site's mesh is no longer a cluster member of central's mesh, so DPS can no longer reach it —
|
|
telemetry needs its own out-of-band transport, same as commands (Phase 2) and config bytes (Phase 3).
|
|
|
|
## Direction: the node hosts, central dials
|
|
|
|
**Load-bearing and easy to get backwards:** each **driver** node hosts the telemetry gRPC server
|
|
(a dedicated Kestrel h2c listener); **central (admin role) is the client** and dials in. This
|
|
mirrors the inversion ScadaBridge already uses for the same problem (its `SiteStreamService`).
|
|
|
|
```
|
|
driver node central (admin)
|
|
┌─────────────────────┐ ┌──────────────────────────┐
|
|
│ 4 publish seams │ │ TelemetryDialSupervisor │
|
|
│ -> node-local hub │ Subscribe │ - discovers ClusterNode │
|
|
│ -> DPS (unchanged) │◄─────────────│ rows (Host+GrpcPort) │
|
|
│ │ (gRPC │ - one reconnecting │
|
|
│ TelemetryStreamGrpc- │ stream) │ dialer per node │
|
|
│ Service (h2c server) │─────────────►│ - feeds the SAME sinks │
|
|
│ Telemetry:GrpcListen- │ TelemetryEvent │ the DPS bridges feed │
|
|
│ Port │ (oneof, 4 kinds)│ today │
|
|
└─────────────────────┘ └──────────────────────────┘
|
|
```
|
|
|
|
A fused `admin,driver` node both hosts (as driver) and dials (as admin) — it dials itself plus its
|
|
redundant pair peer, the same way central dials site nodes.
|
|
|
|
Central discovers driver-node telemetry endpoints from the `ClusterNode` table (`Host` +
|
|
`GrpcPort`, the latter added in Phase 1 for exactly this purpose) — the same DB-sourced discovery
|
|
Phase 1's ack set and Phase 2's `ClusterClient` contact set already use. No shared gossip
|
|
membership is required to find or dial a node.
|
|
|
|
## The dark switch
|
|
|
|
`Telemetry:Mode` (node/serve side) and `TelemetryDial:Mode` (central/dial side) each default to
|
|
`Dps` and can be independently set to `Grpc`. Both code paths are compiled into every binary —
|
|
flipping either flag is an appsettings/env change plus a restart, **not a redeploy or rebuild** —
|
|
the same discipline as `MeshTransport:Mode` and `ConfigSource:Mode`.
|
|
|
|
What actually changes per mode, precisely:
|
|
|
|
- **The node ALWAYS does both things, regardless of `Telemetry:Mode`.** Every one of the four
|
|
publish seams emits into the node-local `ITelemetryLocalHub` **and** publishes to DPS,
|
|
unconditionally. The node additionally **always** hosts the gRPC server whenever
|
|
`Telemetry:GrpcListenPort > 0` — hosting is not mode-gated at all. `Telemetry:Mode` only affects
|
|
central's own bookkeeping about which side it expects to be consulted (it is otherwise inert on
|
|
the node).
|
|
- **Only central's ingest source switches**, driven by `TelemetryDial:Mode`:
|
|
- `Dps` (default) — today's four DPS SignalR bridges (alert, script-log, driver-status,
|
|
resilience) subscribe and feed the AdminUI sinks. Unchanged behavior.
|
|
- `Grpc` — those four DPS bridges are **not spawned**. Instead the `TelemetryDialSupervisor`
|
|
actor dials every driver node's gRPC stream and feeds the **identical** in-process sinks
|
|
(`IInProcessBroadcaster<AlarmTransitionEvent>`, `IInProcessBroadcaster<ScriptLogEntry>`,
|
|
`IDriverStatusSnapshotStore`, `IDriverResilienceStatusStore`).
|
|
- **The AdminUI panels themselves are untouched.** `/alerts`, `/script-log`, and the `/hosts`
|
|
driver table all read from the same sinks in both modes — only the sinks' upstream feed swaps.
|
|
The `fleet-status` bridge (out of Phase 5 scope — see below) stays on DPS in both modes.
|
|
|
|
Because hosting is unconditional and both ingest paths are always compiled in, a fleet can be
|
|
flipped node-by-node and central-by-central with no coordination window where telemetry is lost —
|
|
the node is always serving, so central can switch to `Grpc` whenever it likes.
|
|
|
|
## The four migrated channels
|
|
|
|
| DPS topic (today) | Domain record | Central sink fed |
|
|
|---|---|---|
|
|
| `alerts` | `AlarmTransitionEvent` | `IInProcessBroadcaster<AlarmTransitionEvent>` (+ `AlertHub`); `/alerts` page; live/disconnected pill |
|
|
| `script-logs` | `ScriptLogEntry` | `IInProcessBroadcaster<ScriptLogEntry>` (+ `ScriptLogHub`); `/script-log` page; live pill |
|
|
| `driver-health` | `DriverHealthChanged` | `IDriverStatusSnapshotStore`; `/hosts` driver table |
|
|
| `driver-resilience-status` | `DriverResilienceStatusChanged` | `IDriverResilienceStatusStore` |
|
|
|
|
These four are carried as `oneof` event kinds on one `TelemetryStreamService.Subscribe` RPC — a
|
|
single stream per (central, driver-node) pair carries all four, rather than one stream per topic.
|
|
Proto field evolution is additive-only (never renumber/reuse a tag), locked by a contract test that
|
|
reflects over the `oneof` cases.
|
|
|
|
## The three deferred channels (NOT migrated in Phase 5 — do not read this as "seven done")
|
|
|
|
The program sketch originally named seven observability topics for Phase 5. Three were scoped out,
|
|
each for a distinct, settled reason:
|
|
|
|
- **`redundancy-state`** — bidirectional, built directly from `Cluster.State`, and **pair-local**:
|
|
it drives ServiceLevel and the Primary data-plane gate, consumed in-process by
|
|
`OpcUaPublishActor`, `ScriptedAlarmHostActor`, `DriverHostActor`, and `HistorianAdapterActor`.
|
|
It stays on DPS in **both** `MeshTransport` modes today, and it is genuinely mesh-bound by
|
|
design — under Phase 6 each pair keeps sharing its own small mesh, so DPS keeps working for it
|
|
in-mesh. Central's *display* of each pair's redundancy state is a Phase 6 cross-mesh concern
|
|
(a future observability channel, once central no longer shares gossip with any site pair), not a
|
|
Phase 5 telemetry-panel migration.
|
|
- **`fleet-status`** — **central-internal**, not a node→central stream at all. The admin singleton
|
|
`FleetStatusBroadcaster` builds it from the admin node's own cluster membership/reachability/
|
|
leader events, and `Fleet.razor` **polls the Config DB** and ignores the feed entirely. There is
|
|
nothing here for a per-node gRPC stream to carry. Revisit in Phase 6, once central genuinely
|
|
loses gossip visibility of site nodes and needs another way to know a pair's membership.
|
|
- **`deployment-acks`** — already migrated, but onto a different transport: it rides the Phase 2
|
|
`ClusterClient` transport (`MeshTransport:Mode=ClusterClient`) as a command-plane reply, not an
|
|
observability broadcast. It was never a Phase 5 candidate.
|
|
|
|
See [`docs/Redundancy.md`](Redundancy.md#command-transport-centralnode) for how `redundancy-state`
|
|
and the command transports fit together.
|
|
|
|
## Authentication — fail-closed from day one
|
|
|
|
A shared node bearer key gates the stream: `Telemetry:ApiKey` (node/serve side) must equal
|
|
`TelemetryDial:ApiKey` (central/dial side). `TelemetryStreamAuthInterceptor` enforces it — path-scoped
|
|
to `/telemetry.v1.TelemetryStreamService/`, comparing the `Authorization: Bearer` token with
|
|
`CryptographicOperations.FixedTimeEquals`, and rejecting with `PermissionDenied`. An **unset** key
|
|
rejects every call rather than allowing an open stream — the same fail-closed posture as
|
|
`ConfigServeAuthInterceptor` and `LocalDbSyncAuthInterceptor`.
|
|
|
|
This **supersedes** [the design doc's §6.3](plans/2026-07-21-per-cluster-mesh-design.md), which had
|
|
provisionally decided to "match ScadaBridge's unauthenticated posture for now" for inter-cluster
|
|
transports. ScadaBridge itself has since closed that gap with this identical
|
|
interceptor pattern, so Phase 5 ships authenticated from the start rather than deferring auth to a
|
|
later hardening pass. See the design doc's superseded note for the full history.
|
|
|
|
## Upgrading an existing deployment: populate `ClusterNode.GrpcPort` first (#493)
|
|
|
|
`GrpcPort` is a **nullable column added by Phase 5**, and nothing backfills it onto rows that
|
|
already existed. On an upgraded deployment every `ClusterNode` row therefore carries `NULL`,
|
|
central finds **no dial targets**, and flipping `TelemetryDial:Mode` to `Grpc` connects to
|
|
nothing. Fresh installs are unaffected. `AkkaPort` did not have this problem only because it is
|
|
`NOT NULL` with a default of 4053.
|
|
|
|
Nothing fails loudly, which is what makes this worth stating up front — the code degrades
|
|
gracefully, once per node, throttled:
|
|
|
|
```
|
|
ClusterNode <id> has no GrpcPort; it exposes no telemetry stream and is skipped in this dial refresh
|
|
(further skips of this node are silent)
|
|
```
|
|
|
|
Other symptoms: no `ESTABLISHED` sockets on the telemetry port (a `LISTEN` but nothing else in
|
|
`docker exec <node> cat /proc/net/tcp6`), and the `/alerts` / `/script-log` pill never turning
|
|
live in `Grpc` mode.
|
|
|
|
**Before flipping any node to `Grpc`**, set each driver node's `GrpcPort` to that node's own
|
|
`Telemetry:GrpcListenPort` — via the AdminUI node editor, or directly:
|
|
|
|
```sql
|
|
UPDATE dbo.ClusterNode SET GrpcPort = <that node's Telemetry:GrpcListenPort> WHERE NodeId = '<node>';
|
|
```
|
|
|
|
There is deliberately **no EF data migration** backfilling this. The correct value is that node's
|
|
own listener port, which is per-deployment configuration the database cannot derive; a migration
|
|
could only invent one, and a *wrong* dial target is worse than the `NULL` the dialer already skips
|
|
explicitly. The docker-dev seed does backfill (`GrpcPort IS NULL AND CreatedBy = 'docker-dev-seed'`
|
|
→ 4056) because there the port is fixed by `docker-compose.yml` and therefore actually knowable.
|
|
|
|
## Reconnect and reliability
|
|
|
|
Central's `TelemetryDialSupervisor` (an admin-role actor) runs **one reconnecting dialer per
|
|
driver node**:
|
|
|
|
- **Discovery** — enabled, non-maintenance `ClusterNode` rows are read on start and refreshed every
|
|
`TelemetryDial:ContactRefreshSeconds` (default 60), plus on an admin-side topology change. A row
|
|
with a null `GrpcPort` is skipped (that node exposes no telemetry endpoint — logged, not an
|
|
error). Dial targets are added/removed as the row set changes.
|
|
- **Reconnect backoff** — immediate first retry, then a fixed ~5 second backoff, indefinitely. The
|
|
dialer never gives up permanently: this is an observability channel, not the data plane, so a
|
|
persistently-unreachable node just means a persistently-stale panel, not a fleet fault.
|
|
- **Generation stamping** — each dialer carries a monotonically-increasing generation counter, so
|
|
a late error or late event from a superseded stream (e.g. one that raced a reconnect) is
|
|
recognized and ignored rather than corrupting the current stream's state.
|
|
- **Snapshot replay on reconnect** — the node-local hub keeps a last-value cache for
|
|
`driver-health` and `driver-resilience-status` (there is nothing meaningful to replay for
|
|
`alerts`/`script-logs`, which are append logs). On every new `Subscribe` — including a
|
|
reconnect — the hub drains the cached snapshots to the new stream before live deltas, so
|
|
`IDriverStatusSnapshotStore` / `IDriverResilienceStatusStore` re-prime immediately rather than
|
|
sitting stale until the next natural event. `alerts`/`script-logs` simply tolerate the gap — a
|
|
reconnect loses whatever transitions happened while disconnected, same as a DPS resubscribe would.
|
|
- **Connection indicator** — the `/alerts` and `/script-log` live/disconnected pill is driven by
|
|
**aggregate** stream health across all dialers: connected when **at least one** node stream is
|
|
up, disconnected when all are down. This matches today's DPS `SubscribeAck`/`PostStop` pill
|
|
semantics — the pill has never meant "every node is live," only "the panel has a source." The
|
|
two store-backed panels (`driver-health`, `driver-resilience-status`) carry no fleet-wide
|
|
connection flag today, on gRPC or on DPS; per-row staleness is unchanged.
|
|
|
|
## Configuration reference
|
|
|
|
See [`docs/Configuration.md` § `Telemetry` / `TelemetryDial`](Configuration.md#telemetry--telemetrydial-live-telemetry-transport)
|
|
for the full keys table. In brief:
|
|
|
|
- **`Telemetry`** (node/serve side): `Mode` (`Dps`|`Grpc`, default `Dps`), `GrpcListenPort` (`0` =
|
|
disabled — the driver node's dedicated telemetry h2c port), `ApiKey` (the shared node key; supply
|
|
via `${secret:}`/env, never commit).
|
|
- **`TelemetryDial`** (central side): `Mode` (`Dps`|`Grpc`), `ApiKey` (must equal the nodes'
|
|
`Telemetry:ApiKey`), `ContactRefreshSeconds` (default `60`), `CallTimeoutSeconds` (default `30`).
|
|
|
|
A driver-role node with `Telemetry:Mode=Grpc` must set `GrpcListenPort > 0` (nothing to serve
|
|
otherwise), and `Grpc` mode on either side requires a non-empty `ApiKey` — both are enforced by
|
|
`ValidateOnStart` validators, fail-fast at boot.
|
|
|
|
## Relationship to the rest of the mesh program
|
|
|
|
- **Same dark-switch discipline as Phases 2 and 3.** `MeshTransport:Mode`, `ConfigSource:Mode`, and
|
|
`Telemetry:Mode`/`TelemetryDial:Mode` are all independently-flippable, restart-only config
|
|
changes, and in every case the "new" side is always wired so the flip needs no coordinated
|
|
redeploy.
|
|
- **Why the inversion matters for Phase 6.** Every other cross-boundary transport in this program
|
|
(`MeshTransport`, `ConfigSource`/`ConfigServe`) already has central as the addressable, discoverable
|
|
side and the node as the caller or the callee-by-address. Telemetry is the one channel that is
|
|
naturally node-sourced and fleet-wide-fanned, so putting central in the client role — dialing out
|
|
to each node by its `ClusterNode`-recorded address — is what lets telemetry survive the mesh split
|
|
cleanly: central never needs cluster membership with a site's mesh to keep watching it.
|
|
- **Status:** code-complete on `feat/mesh-phase5`; the live gate (flip the docker-dev rig to `Grpc`,
|
|
confirm all panels stay green with the DPS telemetry bridges unspawned, kill-and-reconnect a site
|
|
node) has not yet run. See `docs/plans/2026-07-22-per-cluster-mesh-program.md` § Phase 5 and
|
|
`docs/plans/2026-07-23-mesh-phase5-grpc-telemetry-stream.md` for the implementation plan.
|