feat(hosts): surface per-host connectivity on /hosts; drop the unwritable table (#521)
`IHostConnectivityProbe` was a dead surface: eleven drivers implement it, `GetHostStatuses()` had ZERO production call sites, and `OnHostStatusChanged` had no subscriber outside the Galaxy driver's own aggregator. Per-host connectivity was computed by every driver and read by nobody. The issue offered "build the publisher or delete it". The publisher as its entity doc described it — driver nodes upserting `DriverHostStatus` rows — is not buildable: per-cluster mesh Phase 4 gates `AddOtOpcUaConfigDb` on the `admin` role, so a driver-only node has no ConfigDb connection to write rows with. So the capability is kept and the transport changed. `DriverHealthChanged.HostStatuses` now carries the probe result to `/hosts` as a Hosts column. That channel already reached the page, already survives the mesh split via the Phase 5 gRPC telemetry stream, and already replays a last-value snapshot on re-subscribe — so per-host state re-primes after a reconnect without a durable store. Both halves of the interface finally do what they are for: the event triggers a prompt publish, the pull is the source of truth. The point of the column is the case the driver-level state chip structurally cannot express: a multi-device driver stays aggregate-Healthy while ONE of its devices is unreachable. Two traps, both pinned by tests that were falsified against the prod code: - The host digest MUST be in the publish fingerprint. On a single-host-down transition every other fingerprint component is unchanged, so the dedup would swallow exactly the publish carrying the news — the trap that already bit the rediscovery signal. Removing it turns the guard test red, verified. - null (no probe) must stay distinct from empty (probe with no hosts). proto3 cannot tell an absent repeated field from an empty one, hence the explicit `has_host_statuses` flag; collapsing them would render every probe-less driver as one whose devices are all fine. Dropped: the DriverHostStatus entity, enum, DbSet, model config and table (migration DropDriverHostStatusTable — empty on every deployment, so the scaffolder's data-loss warning is moot, and Down() recreates it exactly). Found en route, NOT fixed here: `DriverInstanceResilienceStatus` is the identical defect — no writer, no reader, only a DbSet declaration, while the live data rides the `driver-resilience-status` telemetry channel. Its doc-comment now states that rather than describing the sampler and AdminUI join that were never built. Filed as #524 rather than widening this schema change beyond what was asked. Claude-Session: https://claude.ai/code/session_015p7wGqy3YpZNCpDzTpGMKo
This commit is contained in:
@@ -8,6 +8,10 @@ using ZB.MOM.WW.OtOpcUa.Commons.Messages.Logging;
|
||||
using ZB.MOM.WW.OtOpcUa.Commons.Protos.Telemetry.V1;
|
||||
using ZB.MOM.WW.OtOpcUa.Host.Grpc;
|
||||
using ZB.MOM.WW.OtOpcUa.Runtime.Telemetry;
|
||||
// Aliased, not imported wholesale: Core.Abstractions also declares a DriverHealth, which would collide
|
||||
// with the proto DriverHealth these tests assert on.
|
||||
using HostConnectivityStatus = ZB.MOM.WW.OtOpcUa.Core.Abstractions.HostConnectivityStatus;
|
||||
using HostState = ZB.MOM.WW.OtOpcUa.Core.Abstractions.HostState;
|
||||
|
||||
namespace ZB.MOM.WW.OtOpcUa.Host.Tests.Grpc;
|
||||
|
||||
@@ -210,6 +214,40 @@ public sealed class TelemetryStreamGrpcServiceTests
|
||||
evt.DriverHealth.LastSuccessfulReadUtc.ToDateTime().ShouldBe(expectedUtc);
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Node side of the #521 host-status carry: the presence flag must be written, so the tri-state
|
||||
/// (no probe / probe with no hosts / probe with hosts) survives a wire proto3 cannot express on the
|
||||
/// repeated field alone. Paired with the central-side decode in <c>TelemetryProtoMapCentralTests</c>.
|
||||
/// </summary>
|
||||
[Fact]
|
||||
public void ToProto_health_writes_the_host_status_presence_flag_and_entries()
|
||||
{
|
||||
var changedAt = new DateTime(2026, 7, 30, 11, 0, 0, DateTimeKind.Utc);
|
||||
var withHosts = new DriverHealthChanged(
|
||||
"cluster-a", "drv-1", "Healthy", null, null, 0, DateTime.UtcNow,
|
||||
HostStatuses: [new HostConnectivityStatus("plc-a", HostState.Stopped, changedAt)]);
|
||||
|
||||
var evt = TelemetryProtoMapNode.ToProto(new TelemetryItem.Health(withHosts), "c");
|
||||
|
||||
evt.DriverHealth.HasHostStatuses.ShouldBeTrue();
|
||||
evt.DriverHealth.HostStatuses.Count.ShouldBe(1);
|
||||
evt.DriverHealth.HostStatuses[0].HostName.ShouldBe("plc-a");
|
||||
evt.DriverHealth.HostStatuses[0].State.ShouldBe("Stopped");
|
||||
evt.DriverHealth.HostStatuses[0].LastChangedUtc.ToDateTime().ShouldBe(changedAt);
|
||||
|
||||
// A probe reporting zero hosts still sets the flag — that is the whole point of having one.
|
||||
var emptyProbe = new DriverHealthChanged(
|
||||
"cluster-a", "drv-1", "Healthy", null, null, 0, DateTime.UtcNow, HostStatuses: []);
|
||||
var emptyEvt = TelemetryProtoMapNode.ToProto(new TelemetryItem.Health(emptyProbe), "c");
|
||||
emptyEvt.DriverHealth.HasHostStatuses.ShouldBeTrue();
|
||||
emptyEvt.DriverHealth.HostStatuses.ShouldBeEmpty();
|
||||
|
||||
// No probe: flag clear.
|
||||
var noProbe = new DriverHealthChanged("cluster-a", "drv-1", "Healthy", null, null, 0, DateTime.UtcNow);
|
||||
TelemetryProtoMapNode.ToProto(new TelemetryItem.Health(noProbe), "c")
|
||||
.DriverHealth.HasHostStatuses.ShouldBeFalse();
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public async Task Client_disconnect_mid_stream_ends_cleanly_without_leaking_a_slot()
|
||||
{
|
||||
|
||||
Reference in New Issue
Block a user