feat(hosts): surface per-host connectivity on /hosts; drop the unwritable table (#521)

`IHostConnectivityProbe` was a dead surface: eleven drivers implement it, `GetHostStatuses()`
had ZERO production call sites, and `OnHostStatusChanged` had no subscriber outside the Galaxy
driver's own aggregator. Per-host connectivity was computed by every driver and read by nobody.

The issue offered "build the publisher or delete it". The publisher as its entity doc described
it — driver nodes upserting `DriverHostStatus` rows — is not buildable: per-cluster mesh Phase 4
gates `AddOtOpcUaConfigDb` on the `admin` role, so a driver-only node has no ConfigDb connection
to write rows with. So the capability is kept and the transport changed.

`DriverHealthChanged.HostStatuses` now carries the probe result to `/hosts` as a Hosts column.
That channel already reached the page, already survives the mesh split via the Phase 5 gRPC
telemetry stream, and already replays a last-value snapshot on re-subscribe — so per-host state
re-primes after a reconnect without a durable store. Both halves of the interface finally do
what they are for: the event triggers a prompt publish, the pull is the source of truth.

The point of the column is the case the driver-level state chip structurally cannot express: a
multi-device driver stays aggregate-Healthy while ONE of its devices is unreachable.

Two traps, both pinned by tests that were falsified against the prod code:

- The host digest MUST be in the publish fingerprint. On a single-host-down transition every
  other fingerprint component is unchanged, so the dedup would swallow exactly the publish
  carrying the news — the trap that already bit the rediscovery signal. Removing it turns the
  guard test red, verified.
- null (no probe) must stay distinct from empty (probe with no hosts). proto3 cannot tell an
  absent repeated field from an empty one, hence the explicit `has_host_statuses` flag;
  collapsing them would render every probe-less driver as one whose devices are all fine.

Dropped: the DriverHostStatus entity, enum, DbSet, model config and table (migration
DropDriverHostStatusTable — empty on every deployment, so the scaffolder's data-loss warning is
moot, and Down() recreates it exactly).

Found en route, NOT fixed here: `DriverInstanceResilienceStatus` is the identical defect — no
writer, no reader, only a DbSet declaration, while the live data rides the
`driver-resilience-status` telemetry channel. Its doc-comment now states that rather than
describing the sampler and AdminUI join that were never built. Filed as #524 rather than widening
this schema change beyond what was asked.

Claude-Session: https://claude.ai/code/session_015p7wGqy3YpZNCpDzTpGMKo
This commit is contained in:
Joseph Doherty
2026-07-30 04:21:08 -04:00
parent dc9d947bca
commit 30d0697c28
24 changed files with 2411 additions and 305 deletions
@@ -2,6 +2,7 @@ using Shouldly;
using Xunit;
using ZB.MOM.WW.OtOpcUa.AdminUI.Hosts;
using ZB.MOM.WW.OtOpcUa.Commons.Messages.Drivers;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.AdminUI.Tests.Hosts;
@@ -154,4 +155,62 @@ public sealed class HostsDriverViewTests
groups.Select(g => g.ClusterId).ShouldBe(new[] { "Alpha", "Beta", "zeta" });
}
/// <summary>
/// Per-host connectivity flows through to the row (Gitea #521), and <c>DegradedHosts</c> picks out
/// exactly the hosts an operator needs to look at. This is the case the driver-level Status chip
/// cannot express: the driver is Healthy and one of its devices is not.
/// </summary>
[Fact]
public void Build_carries_host_statuses_and_flags_only_the_degraded_ones()
{
var snapshot = Snap("MAIN", "drv-a") with
{
HostStatuses =
[
new HostConnectivityStatus("plc-a", HostState.Running, When),
new HostConnectivityStatus("plc-b", HostState.Stopped, When),
new HostConnectivityStatus("plc-c", HostState.Faulted, When),
],
};
var row = HostsDriverView.Build([snapshot], nodes: null, instances: null).Single().Drivers.Single();
row.State.ShouldBe("Healthy");
row.HostStatuses!.Count.ShouldBe(3);
row.DegradedHosts.Select(h => h.HostName).ShouldBe(["plc-b", "plc-c"]);
}
/// <summary>
/// <see cref="HostState.Unknown"/> counts as degraded. A probe that has not completed its first tick
/// — or one a driver failed to start at all, which AbCip logs explicitly — reports Unknown, and
/// rendering that as healthy is how an unstarted probe stays invisible.
/// </summary>
[Fact]
public void Unknown_host_state_counts_as_degraded()
{
var snapshot = Snap("MAIN", "drv-a") with
{
HostStatuses = [new HostConnectivityStatus("plc-a", HostState.Unknown, When)],
};
var row = HostsDriverView.Build([snapshot], nodes: null, instances: null).Single().Drivers.Single();
row.DegradedHosts.ShouldHaveSingleItem().HostName.ShouldBe("plc-a");
}
/// <summary>
/// A driver with no probe keeps a null host list — distinct from a probe reporting zero hosts. The
/// /hosts column renders "—" for the former and "0 hosts" for the latter, and collapsing them would
/// claim every probe-less driver's devices are fine.
/// </summary>
[Fact]
public void A_driver_without_host_statuses_keeps_null_and_reports_no_degraded_hosts()
{
var row = HostsDriverView.Build([Snap("MAIN", "drv-a")], nodes: null, instances: null)
.Single().Drivers.Single();
row.HostStatuses.ShouldBeNull();
row.DegradedHosts.ShouldBeEmpty();
}
}