Options investigation (Aspire standalone/HealthChecks.UI/Gatus/Homepage/ Grafana vs custom) → decision: custom ZB.MOM.WW.Overview Blazor app in the family look. Design: anonymous read-only single pane, appsettings registry of all four apps' instances, cache-first poller on a configurable timer, configurable staleness timeout, Active/Standby from /health/active, and per-cluster Akka leader chips (needs Health 0.2.0 optional per-entry data). Impl plan: phased + code-verified — Phase 0 Health 0.2.0 (writer data field + AkkaClusterHealthCheck cluster view), Phase 1 ScadaBridge site-node MapZbHealth PR (:8084; role-scoped active check), Phase 2 OtOpcUa bump, Phase 3 the app (:5320), Phase 4 live rig acceptance. Mockup is the visual reference: real Theme tokens + embedded IBM Plex, all card states incl. stale + split-brain. Ready to execute in a fresh session.
24 KiB
ZB.MOM.WW.Overview — implementation plan
Date: 2026-07-22
Design: 2026-07-22-overview-dashboard-design.md (same directory). Read it first; this plan assumes its decisions (custom app, family look/feel, anonymous read-only, cache-first polling, staleness timeout, Akka leader display).
Repos touched: scadaproj/ZB.MOM.WW.Health (shared lib 0.2.0), ~/Desktop/ScadaBridge (site-node health PR), ~/Desktop/OtOpcUa (Health bump + check verification), optional alignment bumps in ~/Desktop/MxAccessGateway + ~/Desktop/HistorianGateway, and the new scadaproj/ZB.MOM.WW.Overview/.
Sequencing: Phase 0 → Phases 1 + 2 (parallel, independent repos) → Phase 3 (app; can start in parallel once 0.2.0 is on the feed) → Phase 4 (live acceptance). Phases 1/2 only gate Phase 4, not Phase 3 development.
All facts below were code-verified 2026-07-22 (file:line cites are from that pass); re-check a cite if the target repo has moved since.
How to run this in a fresh session: read the design doc first, then execute phases in order (0 → 1‖2 → 3 → 4), one phase per sitting with its DoD met before moving on. Work in each repo on a feature branch (feat/health-0.2.0-cluster-data in scadaproj, feat/site-node-health in ScadaBridge, feat/health-0.2.0-bump in OtOpcUa); commit scadaproj work (lib + new app + these docs) on main per family habit only when asked. Nothing in this plan requires user input except merges/pushes and the Phase 4 rig runs; if a cited line has drifted, re-locate by the quoted identifier, not the line number.
Phase 0 — ZB.MOM.WW.Health 0.2.0: entry data + Akka cluster-view
Repo: scadaproj/ZB.MOM.WW.Health/ (plain directory — do NOT git init; commits happen in scadaproj). Akka.Cluster pin 1.5.62 (Directory.Packages.props:9); version is centralized at Directory.Build.props:8 (<Version>0.1.0</Version>).
Task 0.1 — writer: optional per-entry data object
src/ZB.MOM.WW.Health/ZbHealthWriter.cs:
- Add to the private
HealthEntryDto(lines ~76-81):[JsonIgnore(Condition = JsonIgnoreCondition.WhenWritingNull)] public IReadOnlyDictionary<string, object>? Data { get; init; } - In the projection (line ~62):
Data = e.Value.Data.Count > 0 ? e.Value.Data : null(HealthReportEntry.Datais never null; empty when unset). - Invariant: old payloads stay byte-identical. The per-property
JsonIgnoreis essential — the global serializer options deliberately do NOT ignore nulls (the null-descriptionbehavior is pinned byResponseWriterTests.cs:69-85); do not changeSerializerOptions.
Tests (tests/ZB.MOM.WW.Health.Tests/ResponseWriterTests.cs): extend StubHealthCheck (lines 21-31) to accept an optional data dictionary; add (a) non-empty data → "data": {...} emitted with camelCase envelope keys but verbatim data keys, (b) empty data → no data key at all (assert on raw JSON string), (c) existing null-description test still green.
Task 0.2 — AkkaClusterHealthCheck: populate cluster-view data
src/ZB.MOM.WW.Health.Akka/AkkaClusterHealthCheck.cs (currently reads only SelfMember.Status, returns description-only — lines 38-70):
- Add an internal static helper
BuildClusterData(Cluster cluster)returningIReadOnlyDictionary<string, object>with JSON-friendly values only (string/int/string[]):leader—cluster.State.Leader?.ToString()(omit the key when null: pre-join/unknown),selfAddress—cluster.SelfAddress.ToString(),selfRoles—cluster.SelfMember.Roles.OrderBy(r => r).ToArray(),memberCount—cluster.State.Members.Count,unreachableCount—cluster.State.Unreachable.Count. All property names confirmed compiling in this repo (ActiveNodeHealthCheck.cs:116-132,AkkaActiveNodeGate.cs:33-49).
- Pass it through every result path:
HealthCheckResult.Healthy/Degraded/Unhealthy(description, data). The no-ActorSystem / cluster-inaccessible paths keep returning description-only (no data) — pinned byAkkaClusterStatusPolicyTests.cs:63-96. - ASP.NET Core copies
HealthCheckResult.Data→HealthReportEntry.Dataautomatically, andMapZbHealthalready wiresZbHealthWriter(ZbHealthEndpointExtensions.cs:42-54) — no consumer code change needed beyond the package bump, provided the app registers this shared check (verified for ScadaBridge; Task 2.1 verifies OtOpcUa).
Tests (tests/ZB.MOM.WW.Health.Akka.Tests/): the InternalsVisibleTo hook (ActiveNodeHealthCheck.cs:7) exposes the helper. Use Akka.TestKit.Xunit2 (pin 1.5.62; conventions per existing tests — inherit TestKit, never hand-roll ActorSystem.Create + manual join) with a single-node self-joined cluster: assert leader == self address, memberCount == 1, selfRoles round-trip; plus the check-level test that a healthy result now carries data and the degraded no-system path carries none.
Task 0.3 — version, pack, publish
Directory.Build.props:0.1.0→0.2.0. Note the change in the repo's CHANGELOG/README if present.dotnet testfrom the repo root (baseline 58 tests across 3 test projects — all green plus the new ones), thendotnet pack -c Release -o artifacts(3 nupkgs @ 0.2.0).- Publish:
dotnet nuget push artifacts/ZB.MOM.WW.Health*.0.2.0.nupkg --source https://gitea.dohertylan.com/api/packages/dohertj2/nuget/index.json --api-key <from ~/.nuget user config / keychain — same credential every prior publish used>(one command per nupkg or glob). - Verify like every prior lib release (the family gotcha: "published" claims must be proven):
curl -s https://gitea.dohertylan.com/api/packages/dohertj2/nuget/registration/zb.mom.ww.health/index.jsonshows 0.2.0 (NOT/nuget/v3/registration/— that 404s for everything), then restore-verify from a scratch consumer project that pins 0.2.0.
Phase DoD: suite green; 3 packages at 0.2.0 restore-verified from the feed; a manual smoke (any sample host + the check) shows entries["akka-cluster"].data.leader in /health/ready.
Phase 1 — ScadaBridge PR: site-node health endpoints
Repo: ~/Desktop/ScadaBridge (branch off main, e.g. feat/site-node-health). Site branch of src/ZB.MOM.WW.ScadaBridge.Host/Program.cs is lines 485-570: Kestrel :8083 gRPC Http2-only (507-512), :8084 Http1AndHttp2 (516-518); site DI via SiteServiceRegistration.Configure (:532); site pipeline maps only MapZbMetrics (:545), gRPC (:548), MapZbLocalDbSync (:555). No auth middleware and no FallbackPolicy exist on the site pipeline, so anonymous health needs nothing special.
Task 1.1 — bump Health packages
Directory.Packages.props:93-95: all three ZB.MOM.WW.Health* 0.1.0 → 0.2.0.
Task 1.2 — site health checks (register between :532 and var app = builder.Build() at :534)
Mirror the central registration style (Program.cs:277-304, AddTypeActivatedCheck):
akka-cluster(tags: Ready) — the sharedAkkaClusterHealthCheckwithAkkaClusterStatusPolicy.Default, exactly as central does at 282-286. The site ActorSystem→DI bridge already exists (SiteServiceRegistration.cs:168-169). This is also what carriesdata.leaderfor site pairs.localdb(tags: Ready) — new check class in the Host (site namespace): opens the consolidated site SQLite (LocalDb:Path) and runsSELECT 1. There is no EF context on site (central'sDatabaseHealthCheck<ScadaBridgeDbContext>does NOT apply — that context is central-only,Program.cs:266). Optionally enrichdatawithISyncStatus.Connected/OplogBacklog(SiteServiceRegistration.cs:139-140) when replication is enabled — enrich only; replication-off must NOT fail the check (replication is default-OFF by design).active-node(tags: Active) — new thinSitePairActiveNodeHealthCheckdelegating toIClusterNodeProvider.SelfIsPrimary(registered on site,SiteServiceRegistration.cs:172-178; role-scopedsite-{SiteId}oldest-Up). Do NOT reuse central'sOldestNodeActiveHealthCheck— it callsSelfIsOldest(cluster)with no role scope (OldestNodeActiveHealthCheck.cs:30) and would compute oldest across the wrong member set. Follow the existing delegation precedentSiteEventLogActiveNodeCheck(SiteServiceRegistration.cs:187-191). Active semantics: primary → Healthy, standby → Unhealthy (503 = Standby to the dashboard — same convention as central).
Task 1.3 — map endpoints
app.MapZbHealth(); in the site pipeline after UseRouting() (:539), alongside MapZbMetrics() (:545). Served on :8084 (Http1AndHttp2). No auth changes.
Task 1.4 — tests
Extend tests/ZB.MOM.WW.ScadaBridge.Host.Tests/:
- Imitate
HealthCheckTests.cs(WebApplicationFactory<Program>,[Collection(HostBootCollection.Name)],UseSetting("ScadaBridge:Node:Role", ...),SkipMigrations=true, registration assertions viaIOptions<HealthCheckServiceOptions>.Value.Registrations+registration.Factory(services)type checks). Caveat found in recon: every existing factory test boots the Central role; there is no site-role factory bootstrap yet. Build one (Role=Site + minimalScadaBridge:Nodeconfig + tempLocalDb:Path); if the full site host proves too heavy to boot in-test, fall back to theCompositionRootTests.cs/SiteLocalDbWiringTests.csstyle overSiteServiceRegistration.Configurefor registrations, plus a factory test for just the endpoint mapping. Include a positive control for any "central-only check absent on site" assertion. - Assert: site maps
/health/ready+/health/active(non-404); registrations exactly {akka-cluster→Ready, localdb→Ready, active-node→Active} with the intended check types; central registrations unchanged.
Phase DoD: full ScadaBridge suite green; on the local docker rig, curl http://<site-node>:8084/health/ready returns the canonical JSON with akka-cluster carrying data.leader, and /health/active splits 200/503 across the pair; central /health/ready also shows data.leader (free via the bump). PR merged to main.
Phase 2 — OtOpcUa bump (+ optional family alignment)
Task 2.1 — verify the akka check is the shared class
OtOpcUa registers ready checks configdb, akka, admin-leader, localdb-replication (registration reachable from MapOtOpcUaHealth / Program wiring, src/Server/ZB.MOM.WW.OtOpcUa.Host). Verify the akka check is ZB.MOM.WW.Health.Akka.AkkaClusterHealthCheck. If it is → nothing more. If bespoke → either swap to the shared check (preferred, same policy semantics) or populate the same data keys (leader/selfAddress/selfRoles/memberCount/unreachableCount) so the dashboard sees one contract.
Task 2.2 — bump
~/Desktop/OtOpcUa/Directory.Packages.props: ZB.MOM.WW.Health* → 0.2.0. Build + run the Host/AdminUI test suites touched by health. Rig check: any admin node /health/ready shows data.leader.
Task 2.3 (recommended) — alignment bumps, mxgw + HistorianGateway
Neither uses CPM (no Directory.Packages.props — the known trap): bump the inline ZB.MOM.WW.Health pins in MxGateway.Server.csproj and ZB.MOM.WW.HistorianGateway.Server.csproj to 0.2.0. No behavior change (no Akka checks there); keeps the family version matrix aligned. Skippable if either repo has unrelated in-flight work.
Phase DoD: OtOpcUa on 0.2.0 with leader data live on the rig; alignment bumps merged or explicitly deferred with a note.
Phase 3 — the ZB.MOM.WW.Overview app
Location: scadaproj/ZB.MOM.WW.Overview/ — plain directory, NOT a nested git repo (family rule). App conventions (copied from HistorianGateway, the reference implementation — cites from recon):
Task 3.1 — scaffold
ZB.MOM.WW.Overview.slnxwith/src/+/tests/folders (LocalDb's slnx is the minimal template).src/ZB.MOM.WW.Overview/ZB.MOM.WW.Overview.csproj:Sdk="Microsoft.NET.Sdk.Web", net10.0, inline pinned versions, no CPM (the app convention):ZB.MOM.WW.Theme0.3.1,ZB.MOM.WW.Health0.2.0,ZB.MOM.WW.Telemetry+.Serilog0.1.0,ZB.MOM.WW.Configuration0.1.0,Serilog.AspNetCore+ console/file sinks. NO Auth/Audit/Secrets packages.Directory.Build.propswith the family zero-warnings gate (TreatWarningsAsErrors,AnalysisLevel=latest,EnforceCodeStyleInBuild,Deterministic, Nullable/ImplicitUsings) — copy HistorianGateway's.nuget.config: copy HistorianGateway's (nuget.org*+dohertj2-giteasource withZB.MOM.WW.*packageSourceMapping; credentials stay user-level, never committed).
Task 3.2 — registry options + validation
Registry/OverviewOptions.cs (+ ApplicationEntry, InstanceEntry) binding section Overview per the design §3 schema: global PollIntervalSeconds(10)/TimeoutSeconds(3)/StaleAfterSeconds(45); per-application and per-instance overrides (instance wins); instance fields Name, Group?, BaseUrl, ManagementUrl?, HasActiveRole(false).
Registry/OverviewOptionsValidator.cs via OptionsValidatorBase + AddValidatedOptions<OverviewOptions, OverviewOptionsValidator> (ZB.MOM.WW.Configuration): non-empty applications, ≥1 instance each, unique app names + unique instance names within an app, absolute http/https BaseUrl/ManagementUrl, all intervals positive, effective stale > effective poll for every instance. Plus ConfigPreflight.For(builder.Configuration).RequireValue("Overview:Applications:0:Name")…-style presence check pre-Build() (HistorianGateway Program.cs:62-66 is the pattern).
Task 3.3 — health client + status model
Polling/ZbHealthReportClient.cs: GETs a URL, deserializes the canonical shape {status, totalDurationMs, entries:{name:{status, description, durationMs, data?}}} — data values as JsonElement (they arrive as arbitrary JSON; convert to string/int on read). Treat both 200 and 503 as parseable responses (503 carries a body too); classify non-JSON/undeserializable as Unreachable. In-app for v1 (design §8.2).
Polling/InstanceStatus.cs: Up | Degraded | Down | Unreachable (+ separate ActiveState: Active|Standby|Unknown|NotApplicable, IsStale computed at render). Derivation per design §5, including two-strike flap damping (state degrades only after 2 consecutive failures; recovery immediate).
Task 3.4 — poller + snapshot store
Polling/OverviewPollerService.cs : BackgroundService: PeriodicTimer(PollIntervalSeconds); each tick polls all instances in parallel (Task.WhenAll; per-instance CancellationTokenSource with its effective timeout; named HttpClient via IHttpClientFactory; no retries — next tick is the retry). Per instance: /health/ready always; /health/active when HasActiveRole.
Polling/OverviewSnapshotStore.cs: single immutable snapshot object swapped atomically (Volatile.Write/Interlocked.Exchange); each instance snapshot carries LastUpdatedUtc, latency, parsed entries, derived status, consecutive-failure counter. A C# event/Action notifies circuits to re-render on swap. Pages never poll — they read the store only (cache-first requirement).
Staleness: computed at render (now - LastUpdatedUtc > effective StaleAfterSeconds ⇒ render Stale regardless of stored status); the page also runs a small UI tick (~5 s) so ages/staleness advance without a poll event.
Task 3.5 — leader aggregation
Polling/LeaderResolver.cs: for each (Application, Group) with any HasActiveRole instance, collect entries["akka-cluster"].data.leader from members whose status is Up/Degraded; resolve the Akka address (akka.tcp://<system>@host:port) to a registry instance by host match against BaseUrl (fall back to showing the raw address when unmatched — different ports are expected, the Akka port is not the HTTP port); if reporting members disagree → set the group's LeaderDisagreement flag (split-brain tell). Pure static logic → heavily unit-testable.
Task 3.6 — UI
Visual reference (match it): docs/mockups/overview-dashboard-mockup.html — a static, self-contained mockup built from the real Theme tokens + embedded IBM Plex; open it in a browser before styling. It defines: the KPI strip treatment (alert/caution tints), the instance-card anatomy (status stripe on the top edge; dashed border = Unreachable, solid = Down; faded body + "stale since" for Stale), group headers with Leader · <instance> chips and the split-brain warning chip, footer ready/active raw-signal line, and the check-detail <details> layout with the monospace data line. Reuse Theme classes where they exist (chip-*, panel, kv, agg-*); the mockup's extra classes (inst*, group*, check-*) go into the app's own site.css.
Copy HistorianGateway's Blazor skeleton, minus every auth piece:
Components/App.razor(head order matters: Bootstrap CSS →<ThemeHead/>→site.css→HeadOutlet @rendermode=InteractiveServer; body:<Routes @rendermode=InteractiveServer/>→ Bootstrap JS →<ThemeScripts/>→blazor.web.js),Routes.razorwith plainRouteView(noAuthorizeRouteView),_Imports.razorincl.@using ZB.MOM.WW.Theme.Components/Layout/MainLayout.razor→ThemeShell Product="SCADA Overview" Accent="…"(pick an unused accent; HG is#2f855a), Nav = singleNavRailItem Href="/" Match="NavLinkMatch.All"(+ a Health link later if wanted), no AuthorizeView anywhere.Components/Pages/Overview.razor(@page "/"): header strip (counts n Up/Degraded/Down/Unreachable/Stale, last-refresh clock, pause/resume toggle — pause just stops applying poll events + lets staleness paint), then per-application sections: app name + worst-status rollup pill; cluster groups get a sub-header withLeader: <instance>chip (StatusPill Info) and the disagreement warning chip when flagged; instance grid ofInstanceCards.Components/Widgets/InstanceCard.razor:TechCard+StatusPillcomposition — name/group, status pill (mapping per design §5), Active badge, Leader badge, latency, relative last-seen ("stale since …" once stale), native<details>check list (per-check pill + description + data leader line for akka-cluster),TechButton(Ghost/Secondary) →ManagementUrl(target="_blank" rel="noopener").- Program wiring:
AddRazorComponents().AddInteractiveServerComponents();MapRazorComponents<App>().AddInteractiveServerRenderMode()— noRequireAuthorization, no antiforgery/cascading auth state;UseStaticFiles()before endpoints (Theme assets come from_content/ZB.MOM.WW.Theme/).
Task 3.7 — self-observability
AddZbSerilog(o => o.ServiceName = "overview"); AddZbTelemetry(o => { o.ServiceName = "overview"; o.Meters = [OverviewMetrics.MeterName]; }) — Meters is a silent allowlist: forgetting the custom meter name silently drops the gauge (known family gotcha). OverviewMetrics: observable gauge zb_overview_instance_status (labels app/instance/group; value = status enum ordinal) + poll-duration histogram. Register own checks: registry-loaded and last-poll-cycle-completed (Unhealthy if no completed cycle within 2× poll interval — makes the dashboard's own /health/ready honest). MapZbHealth() + MapZbMetrics().
Task 3.8 — tests (tests/ZB.MOM.WW.Overview.Tests)
xunit 2.9.3 + xunit.runner.visualstudio 3.1.4 + Microsoft.NET.Test.Sdk 17.14.1 + bunit 1.40.0 (Theme's own test stack — imitate StatusPillTests.cs: class X : TestContext, RenderComponent<T>, class-list asserts):
- Validator table tests (incl. stale ≤ poll rejected, relative URL rejected).
- Health-client golden tests against captured real payloads from all four apps (capture during Phase 4 prep; until then use handcrafted fixtures matching
ZbHealthWriterincl. entries with and withoutdata). - Status derivation: 200/Healthy→Up, 200/Degraded→Degraded, 503→Down, timeout/refused/garbage→Unreachable, flap damping (1 failure holds, 2 flips, recovery immediate), active 200/503/error → Active/Standby/Unknown.
- Staleness: fresh vs. aged snapshot with per-instance override.
- LeaderResolver: agreement, disagreement flag, address→instance host matching, unmatched fallback.
- Snapshot store: atomic swap + event fired.
- bUnit:
InstanceCardrenders pill class/badges/link per state; Overview page smoke over a stubbed store. WebApplicationFactory<Program>boot tests: page +/health/ready+/metricsreachable with no auth; malformed registry → boot fails with theConfigPreflight/validator message.
Task 3.9 — config + docker
appsettings.json: Logging skeleton, emptySerilog/Telemetrysections (libs default them),Overviewregistry sample;appsettings.Development.json: KestrelHttpendpointhttp://localhost:5320+ a dev registry pointing at the docker rigs. Derive the rig instance URLs from the compose files, don't guess: OtOpcUa docker-dev rig →~/Desktop/OtOpcUadocker-dev compose (6 nodes; AdminUI reachable on host :9200 — per-node published HTTP ports are in the compose port mappings); ScadaBridge local cluster →~/Desktop/ScadaBridge/docker/compose (central pair UI ports + site-node :8084 mappings); mxgw →windev:5130; HistorianGateway → its docker compose (:5220) orwonder-app-vd03:5222(VPN). Production Kestrel via env (Kestrel__Endpoints__Http__Url), the HG pattern; don't setASPNETCORE_URLSalongside Kestrel endpoint config (URLs-override trap).docker/: HG-pattern runtime-only image — host-sidedotnet publish -c Release -o docker/publish -p:UseAppHost=false(host has the Gitea feed creds),FROM mcr.microsoft.com/dotnet/aspnet:10.0,COPY publish/ ./,EXPOSE 5320,ENTRYPOINT ["dotnet","ZB.MOM.WW.Overview.dll"]; compose project namezb-overview, port5320:5320,ASPNETCORE_ENVIRONMENT+ Kestrel env vars. No env_file needed (no secrets). Note:aspnet:10.0has nocurl— healthcheck viawget/dotnet or skip.
Phase DoD: dotnet build 0 warnings; dotnet test green; dotnet run on :5320 renders the page instantly from cache against a dev registry; container builds and runs.
Phase 4 — live acceptance (design §9)
Prep: capture real /health/ready + /health/active payloads from all four apps (post-0.2.0) into the golden-test fixtures. Then run the design §9 checklist end-to-end on the rigs (OtOpcUa 6-node docker-dev + ScadaBridge local cluster + windev mxgw + wonder-app-vd03 HG if reachable/VPN):
- Every registered instance renders correctly; single Active per pair (central AND site).
- Leader chip correct per cluster group and moves when the leader node is stopped; split-brain flag not shown in normal operation.
- Kill a node → Unreachable within ~2 cadences; restart → recovers.
- Force a Degraded check → Warn with failing entry visible.
- Cache-first: first page load instant, no poll triggered by the request.
- Pause auto-refresh past
StaleAfterSeconds→ all cards Stale; resume → clean recovery. Also stop the poller (kill the service loop) → same. - Management links open the right UIs; page reachable with no login.
- Malformed registry → boot fails with a clear preflight message.
curl :5320/health/ready+/metrics(gauge present with per-instance labels).
Program DoD: all 9 pass; ScadaBridge PR merged; OtOpcUa bump merged; Health 0.2.0 on the feed; scadaproj commits for the lib + new app + docs.
Cross-cutting notes for the executor
- Family gotchas that WILL bite:
ZbTelemetryOptions.Meterssilent allowlist; mxgw/HG have noDirectory.Packages.props(inline pins); never host-sqlite3a live container WAL DB during rig checks;Akka.TestKit.Xunit2is xunit-v2-only (fine here — everything in this plan is xunit 2.9.3); Gitea feed verification uses/api/packages/dohertj2/nuget/registration/…. - Do not add auth "while at it" — anonymous read-only is a requirement, not an omission.
- Contract discipline: the
datafield is additive and optional; the dashboard must render fully (minus leader) against 0.1.0-era payloads — that's what keeps Phase 3 unblocked by Phases 1/2 and tolerant of un-bumped deployments. - Commit style: lib change + publish in scadaproj (
feat(health): 0.2.0 …); ScadaBridge and OtOpcUa changes as PRs in their own repos referencing this plan; new app committed in scadaproj.