6dd7858619
Options investigation (Aspire standalone/HealthChecks.UI/Gatus/Homepage/ Grafana vs custom) → decision: custom ZB.MOM.WW.Overview Blazor app in the family look. Design: anonymous read-only single pane, appsettings registry of all four apps' instances, cache-first poller on a configurable timer, configurable staleness timeout, Active/Standby from /health/active, and per-cluster Akka leader chips (needs Health 0.2.0 optional per-entry data). Impl plan: phased + code-verified — Phase 0 Health 0.2.0 (writer data field + AkkaClusterHealthCheck cluster view), Phase 1 ScadaBridge site-node MapZbHealth PR (:8084; role-scoped active check), Phase 2 OtOpcUa bump, Phase 3 the app (:5320), Phase 4 live rig acceptance. Mockup is the visual reference: real Theme tokens + embedded IBM Plex, all card states incl. stale + split-brain. Ready to execute in a fresh session.
154 lines
16 KiB
Markdown
154 lines
16 KiB
Markdown
# ZB.MOM.WW.Overview — family overview dashboard, v1 design
|
||
|
||
**Date:** 2026-07-22
|
||
**Decision context:** `overview_dashboard_options.md` (same date) — Option A chosen: a custom in-house pane, matching the look and feel of the other scadaproj apps. Aspire standalone rejected (resource/status UI disabled outside an AppHost; no persistence; dev-tool positioning).
|
||
**Implementation plan:** `2026-07-22-overview-dashboard-impl-plan.md` (same directory) — phased, code-verified tasks incl. the sister-project changes.
|
||
|
||
## 1. What it is
|
||
|
||
A small Blazor Server app, hosted in scadaproj like the other shared components, that shows **every registered application instance on one page**: Up/Degraded/Down/Unreachable, Active/Standby where applicable, the **current Akka cluster leader** for each cluster group, per-check detail, and a deep link to each instance's own management UI. The page always renders instantly from a cached snapshot — a background poller on a **configurable timer** does all the HTTP work, and a **configurable staleness timeout** marks any data the poller hasn't refreshed as out-of-date/offline. **No login: the dashboard is anonymous and strictly read-only** (it only GETs endpoints that are already anonymous, and has no mutating actions). Configured entirely from an appsettings registry. **Every instance reports through the same channel:** the anonymous `/health/ready` + `/health/active` endpoints of `MapZbHealth`, parsing the canonical `ZbHealthWriter` JSON. Three of the four apps already expose these everywhere; the one gap — ScadaBridge *site* nodes — is closed by a small prerequisite change in ScadaBridge (§4) rather than worked around, so the dashboard has exactly one probe model.
|
||
|
||
Non-goals for v1: history/uptime graphs, alerting/webhooks, telemetry viewing (OTLP/traces), and any push-based ingestion. If ops-grade monitoring is ever needed, that's the Grafana/Prometheus phase in the options doc — this registry ports straight to a scrape config.
|
||
|
||
## 2. Project shape
|
||
|
||
```
|
||
ZB.MOM.WW.Overview/ # plain directory in scadaproj — NOT a nested git repo
|
||
ZB.MOM.WW.Overview.slnx
|
||
src/ZB.MOM.WW.Overview/ # single ASP.NET Core + Blazor Server project
|
||
Program.cs
|
||
Registry/ (options + validator)
|
||
Polling/ (poller BackgroundService, health client, snapshot store)
|
||
Components/ (App.razor, MainLayout → ThemeShell, Pages/Overview.razor, widgets)
|
||
tests/ZB.MOM.WW.Overview.Tests/
|
||
docker/ # HistorianGateway-pattern runtime-only image
|
||
```
|
||
|
||
- .NET 10, global InteractiveServer render mode (same router style as ScadaBridge/HG/mxgw — no per-page render-mode traps).
|
||
- Package references (Gitea feed): `ZB.MOM.WW.Theme` 0.3.1, `ZB.MOM.WW.Health` (0.2.0 — see §4a), `ZB.MOM.WW.Telemetry(.Serilog)` 0.1.0, `ZB.MOM.WW.Configuration` 0.1.0. **No Auth, no Audit, no Secrets**: the dashboard is anonymous read-only by requirement — there is no login, no mutating action to audit, and nothing secret in the registry. (If a login is ever wanted later, `AddZbLdapAuth` + `LoginCard` drop in the family way.)
|
||
- It is an **app**, not a published library — nothing gets packed to the feed.
|
||
|
||
## 3. Registry (appsettings)
|
||
|
||
```jsonc
|
||
"Overview": {
|
||
"PollIntervalSeconds": 10, // global poll cadence (configurable timer)
|
||
"TimeoutSeconds": 3, // per-request timeout
|
||
"StaleAfterSeconds": 45, // snapshot older than this ⇒ instance rendered Stale (out of date/offline)
|
||
"Applications": [
|
||
{
|
||
"Name": "OtOpcUa",
|
||
"ManagementLabel": "AdminUI",
|
||
"Instances": [
|
||
{ "Name": "central-1", "Group": "central",
|
||
"BaseUrl": "http://otopcua-central-1:9000",
|
||
"ManagementUrl": "http://otopcua-central-1:9000/",
|
||
"HasActiveRole": true },
|
||
{ "Name": "site-a-1", "Group": "site-a", "BaseUrl": "http://...", "HasActiveRole": true }
|
||
]
|
||
},
|
||
{
|
||
"Name": "ScadaBridge",
|
||
"Instances": [
|
||
{ "Name": "central-a", "Group": "central", "BaseUrl": "http://sb-central-a:5000",
|
||
"ManagementUrl": "http://sb-central-a:5000/", "HasActiveRole": true },
|
||
{ "Name": "site-a-1", "Group": "site-a",
|
||
"BaseUrl": "http://sb-site-a-1:8084", "HasActiveRole": true } // health on the HTTP/1.1 listener — see §4
|
||
]
|
||
},
|
||
{
|
||
"Name": "MxAccessGateway",
|
||
"ManagementLabel": "Dashboard",
|
||
"Instances": [
|
||
{ "Name": "windev", "BaseUrl": "http://windev:5130",
|
||
"ManagementUrl": "http://windev:5130/" } // single-instance; no active/standby role
|
||
]
|
||
},
|
||
{
|
||
"Name": "HistorianGateway",
|
||
"ManagementLabel": "Dashboard",
|
||
"Instances": [
|
||
{ "Name": "wonder-app-vd03", "BaseUrl": "https://wonder-app-vd03:5222",
|
||
"ManagementUrl": "https://wonder-app-vd03:5222/" } // prod single TLS ALPN port; dev :5220
|
||
]
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
All four family apps are first-class registry citizens; `HasActiveRole` is simply omitted for the two single-instance gateways (mxgw, HistorianGateway), which have no active/standby concept — their cards show status without the Active badge.
|
||
|
||
- Every instance is probed the same way: GET `{BaseUrl}/health/ready`, plus `{BaseUrl}/health/active` when `HasActiveRole`. There is deliberately no per-instance probe-mode switch — one probe model, one status derivation.
|
||
- `PollIntervalSeconds`/`TimeoutSeconds`/`StaleAfterSeconds` are overridable per application and per instance (instance wins) for slow links or chatty pairs; `StaleAfterSeconds` must exceed `PollIntervalSeconds` (validated).
|
||
- Bound + validated with `AddValidatedOptions<OverviewOptions, OverviewOptionsValidator>` (`ZB.MOM.WW.Configuration`): non-empty registry, absolute http(s) URLs, unique app/instance names, positive intervals, stale > poll. Malformed registry fails boot via `ConfigPreflight`, not the first poll.
|
||
|
||
## 4. Prerequisite target-app changes (health from ALL apps)
|
||
|
||
The dashboard's contract is "every registered instance serves `MapZbHealth`". Current reality and the deltas:
|
||
|
||
| App / node kind | Today | Delta needed |
|
||
|---|---|---|
|
||
| OtOpcUa admin + driver nodes | `MapZbHealth` mapped unconditionally | None in code. **Deployment note:** a driver-only Windows-service node with no explicit URLs falls back to Kestrel's `localhost:5000` — every registered node must be given a reachable binding (`ASPNETCORE_URLS`/`HTTP_PORTS`), which the docker rig already does. |
|
||
| ScadaBridge **Central** | `MapZbHealth` (ready: database/akka-cluster/required-singletons; active: active-node) | None. |
|
||
| ScadaBridge **Site** nodes | **No health endpoints** — only the gRPC `:8083` (HTTP/2-only) + `/metrics` on the `:8084` HTTP/1.1 listener | **Small ScadaBridge PR (the one prerequisite):** map `MapZbHealth` on the `:8084` HTTP/1.1 listener alongside `/metrics`. Ready checks: `akka-cluster` (site pair) + site SQLite store (+ `localdb-replication` where enabled); active check: the site pair's oldest-Up active-node rule via `ActiveNodeHealthCheck` — the Host already references all three `ZB.MOM.WW.Health` packages, so this is pure wiring, no new dependencies. Gives site rows full Up/Degraded detail **and** Active/Standby, same as everything else. |
|
||
| MxAccessGateway | `MapZbHealth` (ready: auth-store) | None. |
|
||
| HistorianGateway | `MapZbHealth` (ready: historian-connection, galaxy-sql, conditional runtime/store-forward/pool checks) | None. |
|
||
|
||
Sequencing: the ScadaBridge PR can land any time before the dashboard's acceptance run; the dashboard itself needs no knowledge of it beyond registry entries.
|
||
|
||
### 4a. Shared-lib prerequisite: `ZB.MOM.WW.Health` 0.2.0 — cluster-leader data
|
||
|
||
Requirement: the dashboard must show the **current leader** of each Akka cluster. The canonical health JSON today carries only `status`/`description`/`durationMs` per entry — no structured data — so this needs a small **additive** change to the shared lib (which is exactly where the family wants such things to live):
|
||
|
||
- **Core (`ZbHealthWriter`):** emit an optional `"data": { ... }` object per entry, sourced from `HealthReportEntry.Data`, **only when non-empty** — existing payloads stay byte-identical, so this is non-breaking for every current consumer.
|
||
- **`ZB.MOM.WW.Health.Akka` (`AkkaClusterHealthCheck`):** populate the check's data dictionary from local cluster state: `leader` (address), `selfAddress`, `selfRoles`, `memberCount`, `unreachableCount` (+ `roleLeader:<role>` where the app cares). Each node reports *its own view*; the dashboard reads the leader from any Up node of the group and flags disagreement (a visible split-brain tell — a nice free diagnostic for the 2-node pairs).
|
||
- **Version/rollout:** bump lib to `0.2.0`, publish the 3 packages to the Gitea feed. **Required bumps:** OtOpcUa + ScadaBridge (the Akka apps — leader data appears with the package bump alone, no app code changes, since both register the shared `AkkaClusterHealthCheck`). **Recommended alignment bumps:** mxgw + HistorianGateway (no behavior change for them; keeps the family version matrix aligned).
|
||
|
||
## 5. Polling and status model
|
||
|
||
One `OverviewPollerService : BackgroundService`:
|
||
|
||
- Every cadence tick, polls **all instances in parallel** (`Task.WhenAll`, per-instance `CancellationTokenSource` timeout; named `HttpClient` via `IHttpClientFactory`, no retries — the next tick is the retry).
|
||
- Parses the canonical shape `{ status, totalDurationMs, entries: { name: { status, description, durationMs } } }`. Parsing lives in a small `ZbHealthReportClient` inside this app for v1; promoting it into a `ZB.MOM.WW.Health.Client` package is a deliberate later step, not a v1 dependency (recommendation: promote only when a second consumer appears).
|
||
- Derived per-instance state:
|
||
- **Up** — ready 200 + `"Healthy"`; **Degraded** — 200 + `"Degraded"`; **Down** — 503 / `"Unhealthy"`; **Unreachable** — timeout, refused, DNS, non-JSON (kept distinct from Down: it usually means network/VPN/registry-typo, not the app).
|
||
- **Active/Standby** — `/health/active` 200 → Active, 503 → Standby, error → Unknown (only when `HasActiveRole`).
|
||
- Plus: last poll UTC, latency ms, per-check entries, consecutive-failure count (state flips to Down/Unreachable only after **2** consecutive failures to avoid single-blip flapping; recovery is immediate).
|
||
- **Cache-first rendering:** results land in an in-memory `OverviewSnapshotStore` (single immutable snapshot swapped atomically); an event notifies circuits to re-render. Page loads NEVER trigger a poll — they read the cached snapshot and render instantly; the poller timer is the only thing that does HTTP. No persistence — a restart repolls the world within one cadence.
|
||
- **Staleness:** every instance snapshot carries `LastUpdatedUtc`. When the UI renders (or a periodic sweep runs), any instance whose snapshot age exceeds its effective `StaleAfterSeconds` is shown as **Stale (out of date/offline)** regardless of its last known status — this catches a stuck poller, a paused refresh, or a machine-slept dashboard, not just failed probes.
|
||
- **Leader extraction:** for each `Group` (one Akka cluster = one group, e.g. `OtOpcUa/central`, `ScadaBridge/site-a`), the aggregator reads the `akka-cluster` entry's `data.leader` from each Up member and surfaces the group's leader (resolving the leader address back to a registry instance name where possible). If Up members disagree on the leader, the group is flagged (split-brain indicator).
|
||
- StatusPill mapping: Up→`Ok`, Degraded→`Warn`, Down→`Bad`, Unreachable→`Bad` (distinct label), Stale→`Idle`, Active/Leader badges→`Info`.
|
||
|
||
## 6. UI
|
||
|
||
> **Mockup:** [`docs/mockups/overview-dashboard-mockup.html`](../mockups/overview-dashboard-mockup.html) (static, self-contained; real Theme tokens + embedded IBM Plex) — the visual reference for `InstanceCard`, group headers, leader/split-brain chips, and every status state incl. Stale. Open it in a browser; the implementation should match it.
|
||
|
||
Single page (`/`), `ThemeShell` chassis with its own product name/accent so it reads as a sibling of the other four apps:
|
||
|
||
- One section per application (name + rollup pill = worst instance status), instances as a grid of cards (`TechCard` + `StatusPill` composition — a new `InstanceCard` component, since Theme has no prebuilt tile widget). Cluster groups render a group header with the **Leader: <instance>** chip (and the split-brain warning when member views disagree).
|
||
- Card: instance name + group, status pill, Active badge, **Leader badge** on the leader node's card, latency, relative last-seen (turns into an explicit "stale since …" once past `StaleAfterSeconds`); expandable check list (`entries` with per-check pill + description); a `TechButton` "Open <ManagementLabel>" → `ManagementUrl` (new tab).
|
||
- A header row: total counts (n Up / n Degraded / n Down / n Unreachable / n Stale), last-refresh clock, pause/resume auto-refresh (paused long enough, cards go Stale by design).
|
||
- **No login.** The dashboard is anonymous and read-only end to end; there is nothing to authorize and nothing it can change. It still matches the sister apps' look exactly (ThemeShell + tokens + IBM Plex), just without the login gate.
|
||
|
||
## 7. Self-observability and hosting
|
||
|
||
- The dashboard eats the family dog food: `MapZbHealth` (ready check: "registry loaded" + "last poll cycle completed"), `MapZbMetrics`, `AddZbTelemetry`/`AddZbSerilog`; a `zb_overview_instance_status` gauge (per instance, labeled) so even v1 status is scrapeable — and becomes the Grafana bridge later.
|
||
- Ports: dev `http://localhost:5320` (clear of every family port in use: 5000/5001, 5120/7121, 5220–5222, 8081–8085, 9000, 9200, 4053, 4840); production single TLS ALPN endpoint like HistorianGateway, warn-only if terminated upstream.
|
||
- `docker/` mirrors the HistorianGateway pattern: publish framework-dependent on the host (authed Gitea feed), COPY into `aspnet:10.0`, gitignored env_file. Natural first deployment: the docker-dev rig host, registry pointed at the existing OtOpcUa + ScadaBridge rig nodes — that is also the acceptance test.
|
||
|
||
## 8. Open decisions (settled recommendations)
|
||
|
||
1. **ScadaBridge site rows** — resolved (2026-07-22, user direction): site nodes gain `MapZbHealth` as the one prerequisite PR (§4), so every row is a full health row. *Chosen over* a metrics-liveness fallback (two probe models) and over reading Central's site aggregator (couples the dashboard to ScadaBridge internals and its transport).
|
||
2. **Health-client ownership** — in-app for v1; promote to `ZB.MOM.WW.Health.Client` on second consumer. *(Best practice: contract-owning lib should host the client, but a package bump + publish for one consumer is ceremony without benefit.)*
|
||
3. **History/alerting** — explicitly out of v1; the `zb_overview_instance_status` gauge is the forward hook (Prometheus can scrape the dashboard alone and get fleet status history for free).
|
||
4. **Leader transport** — resolved (2026-07-22, user requirement): via a `data` object in the canonical health JSON (`ZB.MOM.WW.Health` 0.2.0, §4a), populated by the shared `AkkaClusterHealthCheck`. *Chosen over* description-string parsing (fragile) and a separate leader endpoint (second contract for one field).
|
||
5. **Auth** — resolved (2026-07-22, user requirement): none. Anonymous, read-only.
|
||
|
||
## 9. Acceptance (v1 definition of done)
|
||
|
||
- The §4 ScadaBridge site-node health PR is merged and site rows render with full check detail + Active/Standby (no special-cased rows anywhere).
|
||
- Registry pointed at the live docker-dev rigs (OtOpcUa 6-node + ScadaBridge local cluster) renders every instance correctly, including: each Active/Standby pair (central **and** site pairs) showing exactly one Active; **each cluster group showing its current leader, and the leader chip moving when the leader node is stopped**; killing a node flips its card to Unreachable within ~2 cadences and back on restart; a Degraded check (e.g. stop a dependency) shows Warn with the failing entry visible; management links open the right app UIs; the page is reachable with no login.
|
||
- **Cache/staleness proven live:** first page load renders instantly from cache (no poll on request); with auto-refresh paused (or the poller stopped) past `StaleAfterSeconds`, every card flips to Stale, and resumes cleanly.
|
||
- Boot fails with a clear `ConfigPreflight` message on a malformed registry.
|
||
- `dotnet test` green: options validator, status derivation (incl. flap damping, staleness marking, leader extraction + disagreement flag, JSON parsing golden tests against captured real payloads from all four apps — with and without `data`), snapshot store, and a bUnit render test of `InstanceCard`.
|