Closes Phase 3 of the overview-dashboard plan (tasks 3.7-3.9).
Self-observability
- OverviewMetrics: an observable gauge read from the published snapshot at scrape
time, so the metric and the page cannot disagree about what the dashboard
believes, plus a sweep-duration histogram. Instrument names follow the family
METRIC-CONVENTIONS spec (overview.instance.status / overview.poll.duration,
unit s) rather than the plan's literal zb_ names; Prometheus renders them as
overview_instance_status and overview_poll_duration_seconds. Labels are
application/node/group - "node", not "instance", because Prometheus stamps its
own instance label on every scraped series and a colliding metric label is
silently renamed exported_instance. Instances never probed emit nothing: the
enum's zero ordinal is Up, so a placeholder would publish the single most
reassuring answer about a node nobody has heard from.
- Two readiness checks make the dashboard's own /health/ready mean something:
registry-loaded (reports the counts it bound) and last-poll-cycle-completed,
which catches the failure where Kestrel is fine, the page still renders, and
every card on it is quietly frozen. The heartbeat is stamped by the sweep, not
derived from the snapshot's timestamp - the poller publishes the registry
before it probes anything, so a stuck poller would otherwise look healthy.
No cycle yet is Degraded inside the grace window and Unhealthy past it, so a
restart does not fail readiness for its first seconds and a poller that never
started cannot hide behind "still starting".
Configuration
- appsettings.json ships tuning defaults with an EMPTY Applications list, so a
deployment that supplies no registry fails at ConfigPreflight naming the key
instead of booting into a dashboard of nothing.
- Three registries, because reachability genuinely differs: Development is the
host-reachable subset (4 instances), Docker is the full 16-instance fleet by
container DNS name. The rigs publish almost no per-node HTTP port to the host
- OtOpcUa publishes none (only Traefik :9200, load-balanced across the central
pair, so it cannot identify a node) and ScadaBridge keeps site :8084
container-internal. Ports were established by probing the running rigs, not
read off the compose files, because the two disagree: OtOpcUa site nodes serve
health on :8080, not the :9000 their compose anchor implies.
- Real Serilog sinks, not the family's empty "Serilog": {}. AddZbSerilog drives
sinks entirely from configuration and adds none of its own, so the empty
section produces an app that logs NOWHERE - verified against the running
HistorianGateway container, whose log stream is completely empty for exactly
this reason. The HttpClient level override matters too: 16 instances on a 10s
cadence emit several Information lines a second of pure probe noise.
Container
- HistorianGateway-pattern runtime-only image; the compose stack joins the three
rig networks as external, so bringing the dashboard up or down cannot disturb
a rig. The Dockerfile clears ASPNETCORE_HTTP_PORTS as well as ASPNETCORE_URLS:
on .NET 8+ the base image declares its binding via the former, and clearing
only the latter leaves Kestrel logging "Overriding address(es)" every boot.
No healthcheck - aspnet:10.0 has neither curl nor wget, and adding a
--healthcheck entry point to production code to satisfy Compose is not a
trade worth making.
Tests
- BootTests boot the real Program.cs: every endpoint anonymous, the page
rendering from cache while nothing it watches is reachable, the Meters
allowlist proven by an end-to-end /metrics scrape, and both refuse-to-boot
paths (missing registry, stale window inside the poll interval).
The registry is supplied via environment variables, not
ConfigureAppConfiguration, because ConfigPreflight runs before Build() - when
the factory's callbacks are applied - and would not see them.
170 tests, 0 warnings. Verified live against the running rigs: 16 cards, 16
gauge series, page and gauge agreeing exactly.
docker/ — containerised overview dashboard
Runtime-only image, the HistorianGateway pattern: publish on the host (which has the Gitea feed
credentials), copy the output into mcr.microsoft.com/dotnet/aspnet:10.0. Docker never restores.
# from ZB.MOM.WW.Overview/
dotnet publish src/ZB.MOM.WW.Overview/ZB.MOM.WW.Overview.csproj \
-c Release -o docker/publish -p:UseAppHost=false
cd docker && docker compose up -d --build
open http://localhost:5320
docker/publish/ is build output and is gitignored.
Why the container, and not just dotnet run
The dev rigs publish only some node ports to the host:
| Fleet | Health endpoint | Reachable from the Mac host? |
|---|---|---|
| ScadaBridge central-a/b | container :5000 |
yes — localhost:9001 / :9002 |
| ScadaBridge site-* (6) | container :8084 |
no — not published |
| OtOpcUa central-1/2 | container :9000 |
no — only Traefik :9200, load-balanced across the pair |
| OtOpcUa site-* (4) | container :8080 |
no — not published |
| HistorianGateway | container :5220 |
yes — localhost:5220 |
| MxAccessGateway | windev :5130 |
yes, with the VPN up |
So a host-run dashboard can only see part of the fleet. This stack joins the three rig networks
(otopcua-dev_default, scadabridge-net, zb-historiangw_default, all declared external) and
addresses every node by container DNS name, which is how it sees all of them.
Ports above were established by probing the running rigs on 2026-07-24, not read off the
compose files — the two disagree. In particular OtOpcUa's site nodes serve health on :8080, not
the :9000 their compose anchor implies, because their explicit Kestrel listeners for LocalDb sync
re-bind the primary HTTP port.
Start the rigs before this stack. The networks are external, so bringing the dashboard up or
down never disturbs a running rig.
Registry shape
Bound from the Overview section; appsettings.Docker.json is the worked example.
"Overview": {
"PollIntervalSeconds": 10, // sweep cadence; per-app and per-instance overrides allowed
"TimeoutSeconds": 3, // per-request timeout
"StaleAfterSeconds": 45, // must exceed the poll interval, or every card reads Stale
"Applications": [
{
"Name": "ScadaBridge",
"ManagementLabel": "CentralUI", // label on each instance's management link
"Instances": [
{
"Name": "central-a",
"Group": "central", // one Akka cluster = one group; leader + split-brain
// detection are scoped to it. Omit for standalone apps.
"BaseUrl": "http://host:5000", // /health/ready and /health/active are probed under this
"ManagementUrl": "http://localhost:9001/", // followed by the OPERATOR'S browser, so it
// must be host-reachable, not container DNS
"HasActiveRole": true // false for single-instance gateways: no Active badge,
// and /health/active is never probed
}
]
}
]
}
appsettings.json ships the tuning defaults with an empty Applications list, so a deployment
that supplies no registry fails at startup naming the missing key rather than booting into an empty
dashboard.
Configuration traps
- Never set
ASPNETCORE_URLSalongsideKestrel__Endpoints__*. Declaring both puts Kestrel into explicit-endpoints mode and one of them is silently discarded. - No
env_file, no secrets, no login. The dashboard reads anonymous health endpoints and holds no credentials. That is a requirement of the design, not an omission.