merge: close #521 (host connectivity on telemetry) + #522 (delete inert tier machinery)
v2-ci / build (push) Successful in 4m10s
v2-ci / unit-tests (push) Failing after 17m15s

Claude-Session: https://claude.ai/code/session_015p7wGqy3YpZNCpDzTpGMKo
This commit is contained in:
Joseph Doherty
2026-07-30 04:55:18 -04:00
47 changed files with 2564 additions and 1450 deletions
+33 -25
View File
@@ -106,37 +106,45 @@ calls `AddOtOpcUaDriverProbes` in the `hasAdmin` block, and
---
## IDriverSupervisor — Tier C out-of-process recycle
## IDriverSupervisor / process recycle — REMOVED (Gitea #522)
[`Core.Abstractions/IDriverSupervisor.cs`](../src/Core/ZB.MOM.WW.OtOpcUa.Core.Abstractions/IDriverSupervisor.cs)
`IDriverSupervisor`, `MemoryRecycle`, `MemoryTracking` and
`ScheduledRecycleScheduler` **no longer exist.** This section used to describe
them as the Tier C protection layer.
The process-level supervisor contract a **Tier C** (out-of-process) driver's
topology provides. Its concern is restarting the out-of-process Host when a
hard fault is detected (memory breach, wedge, scheduled recycle window). Tier
A/B drivers run in-process and do **not** have a supervisor — recycling them
would kill every OPC UA session and every co-hosted driver. The Core.Stability
layer only invokes this interface after asserting the tier.
They were removed because the thing they acted on is gone. "Tier C" meant a
driver running **out-of-process behind a supervisor that could restart its Host**
without tearing down the OPC UA session or co-hosted drivers. No such process
remains anywhere: Galaxy reaches MXAccess over gRPC to the external
`mxaccessgw` sidecar (PR 7.2 retired the in-process `Galaxy.Host` / `Proxy` /
`Shared` projects), and FOCAS has run in-process since its managed wire client
landed 2026-04-24. Consistent with that, `IDriverSupervisor` had **zero
implementations**, and `MemoryTracking` / `MemoryRecycle` /
`ScheduledRecycleScheduler` were constructed **only in their own unit tests**
so this was not dormant-but-ready code, it was code with no path to running.
Members:
The operator-facing half went too: `RecycleIntervalSeconds` was authorable
through the AdminUI's driver resilience section and validated by the parser
while configuring nothing. A stored `ResilienceConfig` still carrying the key
parses cleanly (unknown keys are ignored) and the AdminUI preserves it rather
than silently rewriting an operator's config.
- `string DriverInstanceId { get; }` — the driver instance this supervisor
governs.
- `Task RecycleAsync(string reason, CancellationToken cancellationToken)`
request a terminate+restart of the Host process; implementations are
expected to be idempotent under repeat calls during an in-flight recycle.
**`DriverTier` itself survives and is load-bearing** — do not delete it while
following this note. `DriverResilienceOptions.GetTierDefaults(tier)` supplies
the real per-capability timeout / retry / breaker policies, resolved via
`DriverFactoryRegistry.GetTier` and applied by `DriverCapabilityInvokerFactory`.
Every driver runs Tier A because no factory passes a tier, so today that means
one uniform policy set.
Callers (both in
[`Core/Stability/`](../src/Core/ZB.MOM.WW.OtOpcUa.Core/Stability/)):
⚠️ Two `IDriver` members are now **consumerless**: `GetMemoryFootprint()` and
`FlushOptionalCachesAsync()` existed to feed `MemoryTracking`. All 12 drivers
still implement them and nothing calls them. They were left in place rather
than swept, since removing them touches every driver and every test stub —
tracked as Gitea **#525**.
- `ScheduledRecycleScheduler`
([`Core/Stability/ScheduledRecycleScheduler.cs`](../src/Core/ZB.MOM.WW.OtOpcUa.Core/Stability/ScheduledRecycleScheduler.cs))
— opt-in periodic recycle. A `TickAsync` method advanced by the caller's
ambient scheduler decides whether the configured interval has elapsed and, if
so, drives `RecycleAsync`. Its constructor throws unless the tier is C, making
in-process misuse structurally impossible.
- `MemoryRecycle`
([`Core/Stability/MemoryRecycle.cs`](../src/Core/ZB.MOM.WW.OtOpcUa.Core/Stability/MemoryRecycle.cs))
— on a memory hard-breach, calls `RecycleAsync` (when a supervisor is wired).
If per-driver memory protection is wanted again, build it for the architecture
as it is now: a process-wide watchdog, not a per-driver tier whose only remedy
requires a supervisor that cannot exist in-process.
---
+1 -1
View File
@@ -27,7 +27,7 @@ The lifecycle facade `OpcUaApplicationHost` (`src/Server/ZB.MOM.WW.OtOpcUa.OpcUa
## Resilience and capability dispatch
Driver-capability calls (`IReadable.ReadAsync`, `IWritable.WriteAsync`, `ITagDiscovery.DiscoverAsync`, `ISubscribable.SubscribeAsync/UnsubscribeAsync`, the `IHostConnectivityProbe` probe loop, `IAlarmSource` surfaces, and the four `IHistoryProvider` reads) are routed through a `CapabilityInvoker` (`src/Core/ZB.MOM.WW.OtOpcUa.Core/Resilience/CapabilityInvoker.cs`) so the Polly resilience pipeline (retry / timeout / breaker) applies. There is one invoker per `(DriverInstance, IDriver)` pair; all invokers share the process-singleton `DriverResiliencePipelineBuilder`, which keys pipelines on `(DriverInstanceId, hostName, DriverCapability)`. Per-instance resilience options come from `DriverTypeRegistry` (the driver's tier) plus per-instance JSON overrides parsed from `DriverInstance.ResilienceConfig` by `DriverResilienceOptionsParser`.
Driver-capability calls (`IReadable.ReadAsync`, `IWritable.WriteAsync`, `ITagDiscovery.DiscoverAsync`, `ISubscribable.SubscribeAsync/UnsubscribeAsync`, the `IHostConnectivityProbe` probe loop, `IAlarmSource` surfaces, and the four `IHistoryProvider` reads) are routed through a `CapabilityInvoker` (`src/Core/ZB.MOM.WW.OtOpcUa.Core/Resilience/CapabilityInvoker.cs`) so the Polly resilience pipeline (retry / timeout / breaker) applies. There is one invoker per `(DriverInstance, IDriver)` pair; all invokers share the process-singleton `DriverResiliencePipelineBuilder`, which keys pipelines on `(DriverInstanceId, hostName, DriverCapability)`. Per-instance resilience options come from the driver's tier — the `DriverTier` passed at factory registration, resolved via `DriverFactoryRegistry.GetTier` (every driver is Tier A today, since no factory passes one) — plus per-instance JSON overrides parsed from `DriverInstance.ResilienceConfig` by `DriverResilienceOptionsParser`. (It was never `DriverTypeRegistry`; that class was vestigial and was deleted in Gitea #522.)
The `OTOPCUA0001` Roslyn analyzer (`src/Tooling/ZB.MOM.WW.OtOpcUa.Analyzers/UnwrappedCapabilityCallAnalyzer.cs`, category `OtOpcUa.Resilience`, severity Warning) flags direct driver-capability calls that bypass the invoker.
+1 -1
View File
@@ -26,7 +26,7 @@ The project was originally called **LmxOpcUa** (a single-driver Galaxy/MXAccess
| [OpcUaServer.md](OpcUaServer.md) | Top-level server architecture — Core, driver dispatch, Config DB, generations |
| [AddressSpace.md](AddressSpace.md) | `GenericDriverNodeManager` + `ITagDiscovery` + `IAddressSpaceBuilder` |
| [ReadWriteOperations.md](ReadWriteOperations.md) | OPC UA Read/Write → `CapabilityInvoker``IReadable`/`IWritable` |
| [DriverLifecycle.md](DriverLifecycle.md) | Server-side driver lifecycle + infrastructure contracts (`IDriverFactory`, `IDriverProbe`, `IDriverSupervisor`, `IDriverHealthPublisher`, `IDriverConfigEditor`, `IHistorianDataSource`) + the Commons library |
| [DriverLifecycle.md](DriverLifecycle.md) | Server-side driver lifecycle + infrastructure contracts (`IDriverFactory`, `IDriverProbe`, `IDriverHealthPublisher`, `IDriverConfigEditor`, `IHistorianDataSource`) + the Commons library |
| [Subscriptions.md](v1/Subscriptions.md) | Monitored items → `ISubscribable` + per-driver subscription refcount (v1 archive) |
| [AlarmTracking.md](AlarmTracking.md) | `IAlarmSource` + `AlarmSurfaceInvoker` + OPC UA alarm conditions — native Galaxy alarms end-to-end (live) |
| [AlarmTracking.md](v1/AlarmTracking.md) | Original alarm-tracking write-up (v1 archive) |
+32
View File
@@ -88,6 +88,38 @@ single stream per (central, driver-node) pair carries all four, rather than one
Proto field evolution is additive-only (never renumber/reuse a tag), locked by a contract test that
reflects over the `oneof` cases.
### Per-host connectivity rides `driver-health` (Gitea #521)
`DriverHealthChanged.HostStatuses` carries each driver's `IHostConnectivityProbe.GetHostStatuses()`
result, rendered as the **Hosts** column on `/hosts`. It exists to show the one thing the driver-level
state chip structurally cannot: a **multi-device driver stays aggregate-`Healthy` while one of its
devices is unreachable**. A FOCAS or AbLegacy instance owning several PLCs previously hid that
entirely.
Three things to know before touching it:
- **This deliberately does NOT go through SQL.** A `DriverHostStatus` table existed, with an entity
doc describing a publisher hosted service that upserted rows from each driver node. That publisher
was never written, and **per-cluster mesh Phase 4 made it unbuildable as described**`Program.cs`
gates `AddOtOpcUaConfigDb` on the `admin` role, so a driver-only node has no ConfigDb connection to
write rows with. The table was dropped (migration `DropDriverHostStatusTable`); it was empty on
every deployment. This channel needs no DB and already survives the mesh split.
- **null ≠ empty, on the wire too.** Null means "the driver has no probe"; an empty list means "it has
one that currently knows no hosts". proto3 cannot distinguish an absent repeated field from an empty
one, so `DriverHealth.has_host_statuses` carries that bit explicitly. Collapsing the two would render
every probe-less driver as one whose devices are all fine.
- **The host digest is part of the publish fingerprint, and must stay there.** `PublishHealthSnapshot`
dedups on that fingerprint, and on a single-host-down transition *every other component is
unchanged* — so without it the dedup swallows precisely the publish carrying the news. Same trap
that bit the rediscovery signal; pinned by
`DriverInstanceActorHostStatusTests.A_single_host_going_down_is_not_swallowed_by_the_unchanged_health_dedup`.
The digest is a flattened, host-name-ordered **string** for the converse reason: a tuple holding an
`IReadOnlyList` compares by reference and would re-publish on every 30 s heartbeat forever.
⚠️ **`DriverInstanceResilienceStatus` is the same defect, still open.** That table also has no writer
and no reader — the only reference in the repo is its DbSet declaration — while the live data reaches
the AdminUI over the `driver-resilience-status` channel above. Keep-or-delete is Gitea **#524**.
## The three deferred channels (NOT migrated in Phase 5 — do not read this as "seven done")
The program sketch originally named seven observability topics for Phase 5. Three were scoped out,
+2 -2
View File
@@ -17,9 +17,9 @@ Each driver opts into only the capabilities it supports. Every async capability
Driver **factories** are registered at startup in `DriverFactoryRegistry` (`src/Core/ZB.MOM.WW.OtOpcUa.Core/Hosting/DriverFactoryRegistry.cs`) — all 12 in one place, `src/Server/ZB.MOM.WW.OtOpcUa.Host/Drivers/DriverFactoryBootstrap.cs:147-164`. Type names are the `DriverTypeNames` constants (`src/Core/ZB.MOM.WW.OtOpcUa.Core.Abstractions/DriverTypeNames.cs`), guarded bidirectionally by `DriverTypeNamesGuardTests`.
> ⚠️ **`DriverTypeRegistry` / `DriverTypeMetadata` are vestigial.** No production code constructs or reads them — the class is referenced only by its own unit tests. Its `AllowedNamespaceKinds` field was retired along with the v3 `Namespace` entity (`DriverTypeRegistry.cs:97`), and the `SystemPlatform` namespace kind no longer exists.
> ⚠️ **`DriverTypeRegistry` / `DriverTypeMetadata` were DELETED (Gitea #522, 2026-07-30).** They were vestigial — no production code constructed or read them, only their own unit tests did — and `AllowedNamespaceKinds` had already been orphaned by the retirement of the v3 `Namespace` entity. Older docs and plans that name them as the driver-metadata or tier authority describe something that no longer exists.
>
> **Stability tier** likewise does not come from that registry: it is the `DriverTier tier = DriverTier.A` parameter on `DriverFactoryRegistry.cs:46`, and **no factory in `src/Drivers/` passes a tier**, so every shipped driver runs Tier A. The Tier-C-only protections described in [docs/v2/driver-stability.md](../v2/driver-stability.md) — `MemoryRecycle` hard-breach kill (`MemoryRecycle.cs:54`) and `ScheduledRecycleScheduler` (`:43`) — are therefore dormant for every driver.
> **Stability tier** never came from that registry anyway: it is the `DriverTier tier = DriverTier.A` parameter on `DriverFactoryRegistry.cs:46`, and **no factory in `src/Drivers/` passes a tier**, so every shipped driver runs Tier A. The Tier-C recycle protections that [docs/v2/driver-stability.md](../v2/driver-stability.md) describes — `MemoryTracking`, `MemoryRecycle`, `ScheduledRecycleScheduler`, `IDriverSupervisor` — are **gone**, not merely dormant: they required an out-of-process driver Host that no longer exists anywhere (Galaxy is gRPC-to-`mxaccessgw`, FOCAS is in-process), `IDriverSupervisor` had no implementations, and the trackers were only ever constructed in tests. The `Tier` column below is therefore uniformly A, and the tier's remaining effect is `DriverResilienceOptions.GetTierDefaults` — the per-capability timeout / retry / breaker policy set, which **is** live.
## Ground-truth driver list
+21 -16
View File
@@ -1,24 +1,29 @@
# Driver Stability & Isolation — OtOpcUa v2
> ⚠️ **The tier assignments below are NOT what runs (verified against source 2026-07-27).**
> **Every one of the 12 drivers runs as Tier A.** `DriverFactoryRegistry` defaults `DriverTier tier =
> DriverTier.A` and **no factory in `src/Drivers/` passes a tier**, so nothing ever selects B or C. The
> Galaxy and FOCAS Tier-C assignments in the table below are aspirational.
> ⚠️ **HISTORICAL DESIGN RECORD — the isolation model below was never built, and its machinery was
> DELETED on 2026-07-30 (Gitea #522).** Read this document for the reasoning, not for what runs.
>
> Consequences worth knowing:
> **What actually runs:** all 12 drivers are Tier A, in-process. `DriverFactoryRegistry` defaults
> `DriverTier tier = DriverTier.A` and no factory in `src/Drivers/` passes a tier, so nothing ever
> selected B or C. The Galaxy and FOCAS Tier-C assignments in the table below are aspirational.
>
> - The **Tier-C-only protections are dormant for every driver** `MemoryRecycle` gates on
> `when _tier == DriverTier.C`, and `ScheduledRecycleScheduler` refuses unless Tier C. Neither has ever
> engaged in production. `IDriverSupervisor` has zero implementations, yet the config parser still
> validates `RecycleIntervalSeconds`, so an operator can author a recycle interval that does nothing.
> - `DriverTypeRegistry` — which this document's model implies is the tier authority — is **vestigial**:
> referenced only by its own test, with nothing registering metadata at startup.
> - The **separate-Windows-service hosting** described for Tier C does not exist either. Galaxy reaches
> MXAccess over gRPC to the external **mxaccessgw** sidecar; the `Galaxy.Proxy`/`Host`/`Shared` pattern
> named below was retired in PR 7.2, and FOCAS runs in-process like everything else.
> **What was deleted, and why deletion rather than activation:** `MemoryTracking`, `MemoryRecycle`,
> `ScheduledRecycleScheduler`, `IDriverSupervisor`, the vestigial `DriverTypeRegistry`, and the
> operator-authorable `RecycleIntervalSeconds` knob. The premise was gone, not merely unused — Tier C
> meant an **out-of-process Host a supervisor could restart**, and no such process exists: Galaxy
> reaches MXAccess over gRPC to the external **mxaccessgw** sidecar (the `Galaxy.Proxy`/`Host`/`Shared`
> pattern named below was retired in PR 7.2) and FOCAS has run in-process since 2026-04-24.
> Consistently, `IDriverSupervisor` had zero implementations and the trackers were constructed only in
> their own unit tests. Activating the tiers would have armed a recycle that had nothing to recycle.
>
> `docs/drivers/README.md` says Tier A for Galaxy and FOCAS and is the accurate one. Keep-or-delete of the
> tier machinery is tracked as Gitea **#522**; this document is retained as the design record.
> **What survives, and must not be deleted while following this document:** `DriverTier` itself and
> `DriverResilienceOptions.GetTierDefaults(tier)` — the real per-capability timeout / retry / breaker
> policies, resolved via `DriverFactoryRegistry.GetTier` and applied by `DriverCapabilityInvokerFactory`.
> The tier enum is live; only the isolation-and-recycle layer built on top of it is gone.
>
> `docs/drivers/README.md` is the accurate per-driver reference. If per-driver memory protection is
> wanted again, build it for the architecture as it is now — a process-wide watchdog, not a per-driver
> tier whose only remedy needs a supervisor that cannot exist in-process.
>
> **Status**: DRAFT — companion to `plan.md`. Defines the stability tier model, per-driver hosting decisions, cross-cutting protections every driver process must apply, and the canonical worked example (FOCAS) for the high-risk tier.
>
@@ -1,3 +1,5 @@
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Commons.Messages.Drivers;
/// <summary>
@@ -26,6 +28,22 @@ namespace ZB.MOM.WW.OtOpcUa.Commons.Messages.Drivers;
/// The driver-supplied reason string from the same event (e.g. <c>"deploy-time-changed"</c>), shown
/// to the operator alongside the prompt. Null when <paramref name="RediscoveryNeededUtc"/> is null.
/// </param>
/// <param name="HostStatuses">
/// Per-host connectivity as reported by <c>IHostConnectivityProbe.GetHostStatuses()</c>, or null when
/// the driver does not implement that capability. Empty (not null) when it does but knows no hosts yet.
/// <para><b>Why this rides the health snapshot rather than a table.</b> The <c>DriverHostStatus</c>
/// entity used to own this, with a doc-comment describing a publisher hosted service that upserted rows
/// from each driver node. That publisher was never built, and since per-cluster mesh Phase 4 it is
/// unbuildable as described: <c>Program.cs</c> gates <c>AddOtOpcUaConfigDb</c> on the <c>admin</c> role,
/// so a driver-only node has no ConfigDb connection to write rows to. This channel already reaches
/// <c>/hosts</c>, already survives the mesh split via the Phase 5 gRPC telemetry stream, and already
/// replays a last-value snapshot on every (re)subscribe — so per-host state re-primes after a
/// reconnect without a durable store. The table was dropped; see Gitea #521.</para>
/// <para><b>Value-equality caveat.</b> A record's generated <c>Equals</c> compares this list by
/// REFERENCE, so two <c>DriverHealthChanged</c> carrying equal-but-distinct lists are not equal.
/// Nothing dedups on record equality — <c>DriverInstanceActor</c> keeps its own flattened fingerprint
/// precisely because of this — but do not introduce such a comparison without fixing it here first.</para>
/// </param>
public sealed record DriverHealthChanged(
string ClusterId,
string DriverInstanceId,
@@ -35,7 +53,8 @@ public sealed record DriverHealthChanged(
int ErrorCount5Min,
DateTime PublishedUtc,
DateTime? RediscoveryNeededUtc = null,
string? RediscoveryReason = null)
string? RediscoveryReason = null,
IReadOnlyList<HostConnectivityStatus>? HostStatuses = null)
{
/// <summary>
/// DPS topic name. Both the runtime <c>AkkaDriverHealthPublisher</c> and the AdminUI
@@ -70,6 +70,19 @@ message DriverHealth {
google.protobuf.Timestamp published_utc = 7;
google.protobuf.Timestamp rediscovery_needed_utc = 8; // DateTime? absent Timestamp encodes null
optional string rediscovery_reason = 9; // nullable in the record
// Per-host connectivity (Gitea #521). proto3 cannot distinguish an absent repeated field from an
// empty one, and the two mean different things here "driver is not an IHostConnectivityProbe" vs
// "it is one and currently knows no hosts" so the presence flag carries that bit explicitly rather
// than letting an empty list silently claim the driver has no probe.
bool has_host_statuses = 10;
repeated HostConnectivity host_statuses = 11;
}
// Mirrors ZB.MOM.WW.OtOpcUa.Core.Abstractions.HostConnectivityStatus.
message HostConnectivity {
string host_name = 1;
string state = 2; // HostState-as-string, matching DriverHealth.state
google.protobuf.Timestamp last_changed_utc = 3;
}
// Mirrors ZB.MOM.WW.OtOpcUa.Commons.Messages.Drivers.DriverResilienceStatusChanged.
@@ -1,62 +0,0 @@
using ZB.MOM.WW.OtOpcUa.Configuration.Enums;
namespace ZB.MOM.WW.OtOpcUa.Configuration.Entities;
/// <summary>
/// Per-host connectivity snapshot the Server publishes for each driver's
/// <c>IHostConnectivityProbe.GetHostStatuses</c> entry. One row per
/// (<see cref="NodeId"/>, <see cref="DriverInstanceId"/>, <see cref="HostName"/>) triple —
/// a redundant 2-node cluster with one Galaxy driver reporting 3 platforms produces 6
/// rows, not 3, because each server node owns its own runtime view.
/// </summary>
/// <remarks>
/// <para>
/// Supports the per-AppEngine Admin dashboard drill-down. The publisher hosted
/// service on the Server side subscribes to every
/// registered driver's <c>OnHostStatusChanged</c> and upserts rows on transitions +
/// periodic liveness heartbeats. <see cref="LastSeenUtc"/> advances on every
/// heartbeat so the Admin UI can flag stale rows from a crashed Server.
/// </para>
/// <para>
/// No foreign-key to <see cref="ClusterNode"/> — a Server may start reporting host
/// status before its ClusterNode row exists (e.g. first-boot bootstrap), and we'd
/// rather keep the status row than drop it. The Admin-side service left-joins on
/// NodeId when presenting rows.
/// </para>
/// </remarks>
public sealed class DriverHostStatus
{
/// <summary>Server node that's running the driver.</summary>
public required string NodeId { get; set; }
/// <summary>Driver instance's stable id (matches <c>IDriver.DriverInstanceId</c>).</summary>
public required string DriverInstanceId { get; set; }
/// <summary>
/// Driver-side host identifier — Galaxy Platform / AppEngine name, Modbus
/// <c>host:port</c>, whatever the probe returns. Opaque to the Admin UI except as
/// a display string.
/// </summary>
public required string HostName { get; set; }
/// <summary>Gets or sets the current connectivity state of the host.</summary>
public DriverHostState State { get; set; } = DriverHostState.Unknown;
/// <summary>Timestamp of the last state transition (not of the most recent heartbeat).</summary>
public DateTime StateChangedUtc { get; set; }
/// <summary>
/// Advances on every publisher heartbeat — the Admin UI uses
/// <c>now - LastSeenUtc &gt; threshold</c> to flag rows whose owning Server has
/// stopped reporting (crashed, network-partitioned, etc.), independent of
/// <see cref="State"/>.
/// </summary>
public DateTime LastSeenUtc { get; set; }
/// <summary>
/// Optional human-readable detail populated when <see cref="State"/> is
/// <see cref="DriverHostState.Faulted"/> — e.g. the exception message from the
/// driver's probe. Null for Running / Stopped / Unknown transitions.
/// </summary>
public string? Detail { get; set; }
}
@@ -1,17 +1,26 @@
namespace ZB.MOM.WW.OtOpcUa.Configuration.Entities;
/// <summary>
/// Runtime resilience counters the CapabilityInvoker + MemoryTracking + MemoryRecycle
/// surfaces for each <c>(DriverInstanceId, HostName)</c> pair. Separate from
/// <see cref="DriverHostStatus"/> (which owns per-host <i>connectivity</i> state) so a
/// host that's Running but has tripped its breaker or is approaching its memory ceiling
/// shows up distinctly on Admin <c>/hosts</c>.
/// Runtime resilience counters per <c>(DriverInstanceId, HostName)</c> pair.
/// <para><b>⚠️ This table is DEAD: nothing writes it and nothing reads it.</b> The only reference in
/// the repo is the <c>DriverInstanceResilienceStatuses</c> DbSet declaration. Do not treat a query
/// against it as a source of runtime state — it returns empty on every deployment.</para>
/// </summary>
/// <remarks>
/// Per <c>docs/v2/implementation/phase-6-1-resilience-and-observability.md</c> §Stream E.1.
/// The Admin UI left-joins this table on DriverHostStatus for display; rows are written
/// by the runtime via a HostedService that samples the tracker at a configurable
/// interval (default 5 s) — writes are non-critical, a missed sample is tolerated.
/// <para>
/// The original design (<c>docs/v2/implementation/phase-6-1-resilience-and-observability.md</c>
/// §Stream E.1) called for a HostedService sampling the tracker every ~5 s into this table, and
/// an Admin UI that left-joined it on <c>DriverHostStatus</c> for display. Neither was built:
/// there is no sampler, the join exists in no razor file, and <c>DriverHostStatus</c> itself was
/// removed in Gitea #521.
/// </para>
/// <para>
/// The live data does exist — it just never goes through SQL. Resilience state reaches the
/// AdminUI as <c>DriverResilienceStatusChanged</c> over the Phase 5 telemetry stream into the
/// in-memory <c>IDriverResilienceStatusStore</c>, the same shape #521 adopted for per-host
/// connectivity, and for the same reason: per-cluster mesh Phase 4 leaves a driver-only node with
/// no ConfigDb connection to write rows with. Keep-or-delete is tracked as Gitea #524.
/// </para>
/// </remarks>
public sealed class DriverInstanceResilienceStatus
{
@@ -1,21 +0,0 @@
namespace ZB.MOM.WW.OtOpcUa.Configuration.Enums;
/// <summary>
/// Persisted mirror of <c>Core.Abstractions.HostState</c> — the lifecycle state each
/// <c>IHostConnectivityProbe</c>-capable driver reports for its per-host topology
/// (Galaxy Platforms / AppEngines, Modbus PLC endpoints, future OPC UA gateway upstreams).
/// Defined here instead of re-using <c>Core.Abstractions.HostState</c> so the
/// Configuration project stays free of driver-runtime dependencies.
/// </summary>
/// <remarks>
/// The server-side publisher (follow-up PR) translates
/// <c>HostStatusChangedEventArgs.NewState</c> to this enum on every transition and
/// upserts into <see cref="Entities.DriverHostStatus"/>. Admin UI reads from the DB.
/// </remarks>
public enum DriverHostState
{
Unknown,
Running,
Stopped,
Faulted,
}
@@ -0,0 +1,56 @@
using System;
using Microsoft.EntityFrameworkCore.Migrations;
#nullable disable
namespace ZB.MOM.WW.OtOpcUa.Configuration.Migrations
{
/// <summary>
/// Drops the DriverHostStatus table (Gitea #521). The scaffolder warns this "may result in the loss
/// of data"; here it cannot. No code ever inserted a row — the publisher hosted service its entity
/// doc described was never written, and per-cluster mesh Phase 4 made it unbuildable as described,
/// since a driver-only node has no ConfigDb connection. Every deployment's copy of this table is
/// empty. Per-host connectivity now rides DriverHealthChanged.HostStatuses over the Phase 5
/// telemetry stream to /hosts. Down() recreates the table exactly, so the rollback is lossless too.
/// </summary>
public partial class DropDriverHostStatusTable : Migration
{
/// <inheritdoc />
protected override void Up(MigrationBuilder migrationBuilder)
{
migrationBuilder.DropTable(
name: "DriverHostStatus");
}
/// <inheritdoc />
protected override void Down(MigrationBuilder migrationBuilder)
{
migrationBuilder.CreateTable(
name: "DriverHostStatus",
columns: table => new
{
NodeId = table.Column<string>(type: "nvarchar(64)", maxLength: 64, nullable: false),
DriverInstanceId = table.Column<string>(type: "nvarchar(64)", maxLength: 64, nullable: false),
HostName = table.Column<string>(type: "nvarchar(256)", maxLength: 256, nullable: false),
Detail = table.Column<string>(type: "nvarchar(1024)", maxLength: 1024, nullable: true),
LastSeenUtc = table.Column<DateTime>(type: "datetime2(3)", nullable: false),
State = table.Column<string>(type: "nvarchar(16)", maxLength: 16, nullable: false),
StateChangedUtc = table.Column<DateTime>(type: "datetime2(3)", nullable: false)
},
constraints: table =>
{
table.PrimaryKey("PK_DriverHostStatus", x => new { x.NodeId, x.DriverInstanceId, x.HostName });
});
migrationBuilder.CreateIndex(
name: "IX_DriverHostStatus_LastSeen",
table: "DriverHostStatus",
column: "LastSeenUtc");
migrationBuilder.CreateIndex(
name: "IX_DriverHostStatus_Node",
table: "DriverHostStatus",
column: "NodeId");
}
}
}
@@ -394,46 +394,6 @@ namespace ZB.MOM.WW.OtOpcUa.Configuration.Migrations
});
});
modelBuilder.Entity("ZB.MOM.WW.OtOpcUa.Configuration.Entities.DriverHostStatus", b =>
{
b.Property<string>("NodeId")
.HasMaxLength(64)
.HasColumnType("nvarchar(64)");
b.Property<string>("DriverInstanceId")
.HasMaxLength(64)
.HasColumnType("nvarchar(64)");
b.Property<string>("HostName")
.HasMaxLength(256)
.HasColumnType("nvarchar(256)");
b.Property<string>("Detail")
.HasMaxLength(1024)
.HasColumnType("nvarchar(1024)");
b.Property<DateTime>("LastSeenUtc")
.HasColumnType("datetime2(3)");
b.Property<string>("State")
.IsRequired()
.HasMaxLength(16)
.HasColumnType("nvarchar(16)");
b.Property<DateTime>("StateChangedUtc")
.HasColumnType("datetime2(3)");
b.HasKey("NodeId", "DriverInstanceId", "HostName");
b.HasIndex("LastSeenUtc")
.HasDatabaseName("IX_DriverHostStatus_LastSeen");
b.HasIndex("NodeId")
.HasDatabaseName("IX_DriverHostStatus_Node");
b.ToTable("DriverHostStatus", (string)null);
});
modelBuilder.Entity("ZB.MOM.WW.OtOpcUa.Configuration.Entities.DriverInstance", b =>
{
b.Property<Guid>("DriverInstanceRowId")
@@ -44,8 +44,6 @@ public sealed class OtOpcUaConfigDbContext(DbContextOptions<OtOpcUaConfigDbConte
public DbSet<ConfigAuditLog> ConfigAuditLogs => Set<ConfigAuditLog>();
/// <summary>Gets the DbSet of external ID reservations.</summary>
public DbSet<ExternalIdReservation> ExternalIdReservations => Set<ExternalIdReservation>();
/// <summary>Gets the DbSet of driver host statuses.</summary>
public DbSet<DriverHostStatus> DriverHostStatuses => Set<DriverHostStatus>();
/// <summary>Gets the DbSet of driver instance resilience statuses.</summary>
public DbSet<DriverInstanceResilienceStatus> DriverInstanceResilienceStatuses => Set<DriverInstanceResilienceStatus>();
/// <summary>Gets the DbSet of LDAP group role mappings.</summary>
@@ -90,7 +88,6 @@ public sealed class OtOpcUaConfigDbContext(DbContextOptions<OtOpcUaConfigDbConte
ConfigureNodeAcl(modelBuilder);
ConfigureConfigAuditLog(modelBuilder);
ConfigureExternalIdReservation(modelBuilder);
ConfigureDriverHostStatus(modelBuilder);
ConfigureDriverInstanceResilienceStatus(modelBuilder);
ConfigureLdapGroupRoleMapping(modelBuilder);
ConfigureScript(modelBuilder);
@@ -542,31 +539,12 @@ public sealed class OtOpcUaConfigDbContext(DbContextOptions<OtOpcUaConfigDbConte
});
}
private static void ConfigureDriverHostStatus(ModelBuilder modelBuilder)
{
modelBuilder.Entity<DriverHostStatus>(e =>
{
e.ToTable("DriverHostStatus");
// Composite key — one row per (server node, driver instance, probe-reported host).
// A redundant 2-node cluster with one Galaxy driver reporting 3 platforms produces
// 6 rows because each server node owns its own runtime view; the composite key is
// what lets both views coexist without shadowing each other.
e.HasKey(x => new { x.NodeId, x.DriverInstanceId, x.HostName });
e.Property(x => x.NodeId).HasMaxLength(64);
e.Property(x => x.DriverInstanceId).HasMaxLength(64);
e.Property(x => x.HostName).HasMaxLength(256);
e.Property(x => x.State).HasConversion<string>().HasMaxLength(16);
e.Property(x => x.StateChangedUtc).HasColumnType("datetime2(3)");
e.Property(x => x.LastSeenUtc).HasColumnType("datetime2(3)");
e.Property(x => x.Detail).HasMaxLength(1024);
// NodeId-only index drives the Admin UI's per-cluster drill-down (select all host
// statuses for the nodes of a specific cluster via join on ClusterNode.ClusterId).
e.HasIndex(x => x.NodeId).HasDatabaseName("IX_DriverHostStatus_Node");
// LastSeenUtc index powers the Admin UI's stale-row query (now - LastSeen > N).
e.HasIndex(x => x.LastSeenUtc).HasDatabaseName("IX_DriverHostStatus_LastSeen");
});
}
// ConfigureDriverHostStatus is GONE (Gitea #521). The DriverHostStatus table held per-host
// connectivity that a publisher hosted service was supposed to upsert from each driver node. That
// publisher was never written, and per-cluster mesh Phase 4 made it unwritable as designed: ConfigDb
// is registered only on the admin role, so a driver-only node has no connection to write rows with.
// Per-host connectivity now rides DriverHealthChanged.HostStatuses to /hosts over the Phase 5
// telemetry stream, which needs no DB and survives the mesh split.
private static void ConfigureDriverInstanceResilienceStatus(ModelBuilder modelBuilder)
{
@@ -1,104 +0,0 @@
namespace ZB.MOM.WW.OtOpcUa.Core.Abstractions;
/// <summary>
/// Process-singleton registry of driver types known to this OtOpcUa instance.
/// Per-driver assemblies register their type metadata at startup; the Core uses
/// the registry to validate <c>DriverInstance.DriverType</c> values from the central config DB.
/// </summary>
/// <remarks>
/// Per <c>docs/v2/plan.md</c> decisions on JSON content validation happening in the Admin app
/// (not SQL CLR), and on driver type → namespace kind mapping being enforced by
/// <c>sp_ValidateDraft</c>. The registry is the source of truth for both checks.
///
/// Thread-safety: registration is typically single-threaded at startup; lookups happen on
/// every config-apply (multi-threaded). The check-then-act inside <see cref="Register"/> is
/// guarded by a private lock so concurrent registrations are atomic — the "registered only
/// once per process" guarantee holds even if two callers race. Readers operate against the
/// volatile snapshot reference produced by the last successful <see cref="Register"/> and
/// never block.
/// </remarks>
public sealed class DriverTypeRegistry
{
private readonly Lock _writeLock = new();
private IReadOnlyDictionary<string, DriverTypeMetadata> _types =
new Dictionary<string, DriverTypeMetadata>(StringComparer.OrdinalIgnoreCase);
/// <summary>Register a driver type. Throws if the type name is already registered.</summary>
/// <remarks>
/// The check-then-act (duplicate check → copy-on-write rebuild → swap) is performed under
/// <see cref="_writeLock"/> so concurrent <see cref="Register"/> calls cannot silently
/// discard each other's registrations.
/// </remarks>
/// <param name="metadata">The driver type metadata to register.</param>
public void Register(DriverTypeMetadata metadata)
{
ArgumentNullException.ThrowIfNull(metadata);
lock (_writeLock)
{
var snapshot = _types;
if (snapshot.ContainsKey(metadata.TypeName))
{
throw new InvalidOperationException(
$"Driver type '{metadata.TypeName}' is already registered. " +
$"Each driver type may be registered only once per process.");
}
var next = new Dictionary<string, DriverTypeMetadata>(snapshot, StringComparer.OrdinalIgnoreCase)
{
[metadata.TypeName] = metadata,
};
_types = next;
}
}
/// <summary>Look up a driver type by name. Throws if unknown.</summary>
/// <param name="driverType">The driver type name to look up.</param>
/// <returns>The registered metadata for the driver type.</returns>
public DriverTypeMetadata Get(string driverType)
{
ArgumentException.ThrowIfNullOrWhiteSpace(driverType);
if (_types.TryGetValue(driverType, out var metadata))
return metadata;
throw new KeyNotFoundException(
$"Driver type '{driverType}' is not registered. " +
$"Known types: {string.Join(", ", _types.Keys)}.");
}
/// <summary>Try to look up a driver type by name. Returns null if unknown (no exception).</summary>
/// <param name="driverType">The driver type name to look up.</param>
/// <returns>The registered metadata, or <see langword="null"/> if the type is not registered.</returns>
public DriverTypeMetadata? TryGet(string driverType)
{
ArgumentException.ThrowIfNullOrWhiteSpace(driverType);
return _types.GetValueOrDefault(driverType);
}
/// <summary>Snapshot of all registered driver types.</summary>
/// <returns>A snapshot collection of all registered driver type metadata.</returns>
public IReadOnlyCollection<DriverTypeMetadata> All() => _types.Values.ToList();
}
/// <summary>Per-driver-type metadata used by the Core, validator, and Admin UI.</summary>
/// <param name="TypeName">Driver type name (matches <c>DriverInstance.DriverType</c> column values).</param>
/// <param name="DriverConfigJsonSchema">JSON Schema (Draft 2020-12) the driver's <c>DriverConfig</c> column must validate against.</param>
/// <param name="DeviceConfigJsonSchema">JSON Schema for <c>DeviceConfig</c> (multi-device drivers); null if the driver has no device layer.</param>
/// <param name="TagConfigJsonSchema">JSON Schema for <c>TagConfig</c>; required for every driver since every driver has tags.</param>
/// <param name="Tier">
/// Stability tier per <c>docs/v2/driver-stability.md</c> §2-4 and the tiering decisions in
/// <c>docs/v2/plan.md</c>. Drives the shared resilience pipeline defaults
/// (<see cref="Tier"/> × capability → <c>CapabilityPolicy</c>), the <c>MemoryTracking</c>
/// hybrid-formula constants, and whether process-level <c>MemoryRecycle</c> / scheduled-
/// recycle protections apply (Tier C only). Every registered driver type must declare one.
/// </param>
// v3: AllowedNamespaceKinds retired with the Namespace entity — the two OPC UA namespaces (Raw/UNS)
// are implicit, so a driver type no longer declares a namespace-kind compatibility bitmask.
public sealed record DriverTypeMetadata(
string TypeName,
string DriverConfigJsonSchema,
string? DeviceConfigJsonSchema,
string TagConfigJsonSchema,
DriverTier Tier);
@@ -48,20 +48,21 @@ public interface IDriver
/// <summary>
/// Approximate driver-attributable footprint in bytes (caches, queues, symbol tables).
/// Polled every 30s by Core; on cache-budget breach, Core asks the driver to flush via
/// <see cref="FlushOptionalCachesAsync"/>.
/// <para>⚠️ <b>Nothing calls this (Gitea #525).</b> Its only consumer was <c>MemoryTracking</c>,
/// deleted with the rest of the driver-tier recycle machinery in Gitea #522 — the doc-comment here
/// used to claim Core polled it every 30 s, which was never true in v3. All 12 drivers still
/// implement it; most return a real number, and returning <c>0</c> is equally correct today. Do not
/// read a value from it and act on it without first building a consumer that is actually wired.</para>
/// </summary>
/// <remarks>
/// Per <c>docs/v2/driver-stability.md</c> §"In-process only (Tier A/B) — driver-instance
/// allocation tracking". Tier C drivers (process-isolated) report through the same
/// interface but the cache-flush is internal to their host.
/// </remarks>
/// <returns>The approximate memory footprint in bytes.</returns>
long GetMemoryFootprint();
/// <summary>
/// Drop optional caches (symbol cache, browse cache, etc.) to bring footprint back below budget.
/// Required-for-correctness state must NOT be flushed.
/// <para>⚠️ <b>Nothing calls this (Gitea #525)</b> — same reason as
/// <see cref="GetMemoryFootprint"/>: the memory-pressure path that would have invoked it is
/// gone.</para>
/// </summary>
/// <param name="cancellationToken">Cancellation token for the operation.</param>
/// <returns>A task that represents the asynchronous operation.</returns>
@@ -22,13 +22,20 @@ public interface IDriverHealthPublisher
/// </param>
/// <param name="rediscoveryReason">The driver-supplied reason from that event; null when
/// <paramref name="rediscoveryNeededUtc"/> is null.</param>
/// <param name="hostStatuses">
/// Per-host connectivity from <see cref="IHostConnectivityProbe.GetHostStatuses"/>, or null when the
/// driver is not an <see cref="IHostConnectivityProbe"/>. Lets a multi-device driver (a FOCAS or
/// AbLegacy instance owning several PLCs) surface ONE unreachable device that would otherwise be
/// invisible behind an aggregate-Healthy driver row.
/// </param>
void Publish(
string clusterId,
string driverInstanceId,
DriverHealth health,
int errorCount5Min,
DateTime? rediscoveryNeededUtc = null,
string? rediscoveryReason = null);
string? rediscoveryReason = null,
IReadOnlyList<HostConnectivityStatus>? hostStatuses = null);
}
/// <summary>
@@ -49,6 +56,7 @@ public sealed class NullDriverHealthPublisher : IDriverHealthPublisher
DriverHealth health,
int errorCount5Min,
DateTime? rediscoveryNeededUtc = null,
string? rediscoveryReason = null)
string? rediscoveryReason = null,
IReadOnlyList<HostConnectivityStatus>? hostStatuses = null)
{ /* no-op */ }
}
@@ -1,27 +0,0 @@
namespace ZB.MOM.WW.OtOpcUa.Core.Abstractions;
/// <summary>
/// Process-level supervisor contract a Tier C driver's out-of-process topology provides
/// (e.g. <c>Driver.Galaxy.Proxy/Supervisor/</c>). Concerns: restart the Host process when a
/// hard fault is detected (memory breach, wedge, scheduled recycle window).
/// </summary>
/// <remarks>
/// Tier A/B drivers do NOT have a supervisor because they run in-process — recycling would
/// kill every OPC UA session and every co-hosted driver. The Core.Stability layer only invokes
/// this interface for Tier C instances after asserting the tier via
/// <see cref="DriverTypeMetadata.Tier"/>.
/// </remarks>
public interface IDriverSupervisor
{
/// <summary>Driver instance this supervisor governs.</summary>
string DriverInstanceId { get; }
/// <summary>
/// Request the supervisor to recycle (terminate + restart) the Host process. Implementations
/// are expected to be idempotent under repeat calls during an in-flight recycle.
/// </summary>
/// <param name="reason">Human-readable reason — flows into the supervisor's logs.</param>
/// <param name="cancellationToken">Cancels the recycle request; an in-flight restart is not interrupted.</param>
/// <returns>A task that represents the asynchronous operation.</returns>
Task RecycleAsync(string reason, CancellationToken cancellationToken);
}
@@ -19,15 +19,17 @@ public sealed record DriverResilienceOptions
public IReadOnlyDictionary<DriverCapability, CapabilityPolicy> CapabilityPolicies { get; init; }
= new Dictionary<DriverCapability, CapabilityPolicy>();
/// <summary>
/// Periodic scheduled recycle interval for Tier C out-of-process hosts, in seconds.
/// Null (the default) = no scheduled recycle; the driver's Host process keeps running
/// indefinitely unless a memory breach or operator action triggers a recycle. Only
/// respected for <see cref="DriverTier.C"/>; Tier A/B recycle would tear down every
/// OPC UA session, so the loader ignores non-null values for those tiers + logs a
/// warning.
/// </summary>
public int? RecycleIntervalSeconds { get; init; }
// RecycleIntervalSeconds is GONE (Gitea #522). It configured a scheduled recycle of a Tier C
// driver's out-of-process Host — a process that no longer exists anywhere: Galaxy reaches MXAccess
// over gRPC to the external mxaccessgw sidecar (PR 7.2 retired the in-process Host/Proxy/Shared
// projects) and FOCAS has run in-process since its managed wire client landed 2026-04-24. The
// scheduler it fed was constructed only in tests, and IDriverSupervisor — the thing that would
// have performed the recycle — had zero implementations. The parser nevertheless validated the
// knob, so an operator could author a recycle interval through the AdminUI and get nothing.
//
// The JSON key is not an error if present: DriverResilienceOptionsParser ignores unknown keys, so
// a ResilienceConfig blob still carrying "recycleIntervalSeconds" parses cleanly and the value is
// simply dropped, which is what it already did in effect.
/// <summary>
/// Look up the effective policy for a capability, falling back to tier defaults when no
@@ -94,25 +94,15 @@ public static class DriverResilienceOptionsParser
}
}
// Scheduled recycle is Tier C only — reject a configured interval on Tier A/B as a
// misconfiguration surface rather than silently honouring it (recycling an in-process
// driver would kill every OPC UA session + every co-hosted driver).
int? recycleIntervalSeconds = null;
if (shape.RecycleIntervalSeconds is int secs)
{
if (secs <= 0)
AppendDiagnostic(ref parseDiagnostic, $"RecycleIntervalSeconds must be positive; got {secs} — ignoring.");
else if (tier != DriverTier.C)
AppendDiagnostic(ref parseDiagnostic, $"RecycleIntervalSeconds is Tier C only; Tier {tier} driver will not scheduled-recycle.");
else
recycleIntervalSeconds = secs;
}
// No recycle-interval handling: the scheduled-recycle machinery was deleted in Gitea #522
// (nothing constructed it, and IDriverSupervisor — which would have performed the recycle —
// had no implementations). A "recycleIntervalSeconds" key surviving in an existing
// ResilienceConfig blob is harmless: the shape no longer declares it, and unknown JSON
// properties are ignored, so the config still parses and the value is dropped.
return new DriverResilienceOptions
{
Tier = tier,
CapabilityPolicies = merged,
RecycleIntervalSeconds = recycleIntervalSeconds,
};
}
@@ -196,8 +186,8 @@ public static class DriverResilienceOptionsParser
private sealed class ResilienceConfigShape
{
/// <summary>Gets or sets the scheduled recycle interval in seconds.</summary>
public int? RecycleIntervalSeconds { get; set; }
// No RecycleIntervalSeconds — see the note at the return site. Deserialization ignores unknown
// properties, so an existing blob carrying the key still binds.
/// <summary>Gets or sets the per-capability resilience policies.</summary>
public Dictionary<string, CapabilityPolicyShape>? CapabilityPolicies { get; set; }
}
@@ -1,72 +0,0 @@
using Microsoft.Extensions.Logging;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Core.Stability;
/// <summary>
/// Tier C only process-recycle companion to <see cref="MemoryTracking"/>. On a
/// <see cref="MemoryTrackingAction.HardBreach"/> signal, invokes the supplied
/// <see cref="IDriverSupervisor"/> to restart the out-of-process Host.
/// </summary>
/// <remarks>
/// Per <c>docs/v2/plan.md</c>. Tier A/B hard-breach on an in-process
/// driver would kill every OPC UA session and every co-hosted driver, so for Tier A/B this
/// class logs a <b>promotion-to-Tier-C recommendation</b> and does NOT invoke any supervisor.
/// A future tier-migration workflow acts on the recommendation.
/// </remarks>
public sealed class MemoryRecycle
{
private readonly DriverTier _tier;
private readonly IDriverSupervisor? _supervisor;
private readonly ILogger<MemoryRecycle> _logger;
/// <summary>Initializes a new instance of the <see cref="MemoryRecycle"/> class.</summary>
/// <param name="tier">The driver tier.</param>
/// <param name="supervisor">Optional supervisor for process recycling.</param>
/// <param name="logger">Logger for recycling events.</param>
public MemoryRecycle(DriverTier tier, IDriverSupervisor? supervisor, ILogger<MemoryRecycle> logger)
{
_tier = tier;
_supervisor = supervisor;
_logger = logger;
}
/// <summary>
/// Handle a <see cref="MemoryTracking"/> classification for the driver. For Tier C with a
/// wired supervisor, <c>HardBreach</c> triggers <see cref="IDriverSupervisor.RecycleAsync"/>.
/// All other combinations are no-ops with respect to process state (soft breaches + Tier A/B
/// hard breaches just log).
/// </summary>
/// <param name="action">The memory tracking action.</param>
/// <param name="footprintBytes">The current process footprint in bytes.</param>
/// <param name="cancellationToken">Cancellation token for the operation.</param>
/// <returns>True when a recycle was requested; false otherwise.</returns>
public async Task<bool> HandleAsync(MemoryTrackingAction action, long footprintBytes, CancellationToken cancellationToken)
{
switch (action)
{
case MemoryTrackingAction.SoftBreach:
_logger.LogWarning(
"Memory soft-breach on driver {DriverId}: footprint={Footprint:N0} bytes, tier={Tier}. Surfaced to Admin; no action.",
_supervisor?.DriverInstanceId ?? "(unknown)", footprintBytes, _tier);
return false;
case MemoryTrackingAction.HardBreach when _tier == DriverTier.C && _supervisor is not null:
_logger.LogError(
"Memory hard-breach on Tier C driver {DriverId}: footprint={Footprint:N0} bytes. Requesting supervisor recycle.",
_supervisor.DriverInstanceId, footprintBytes);
await _supervisor.RecycleAsync($"Memory hard-breach: {footprintBytes} bytes", cancellationToken).ConfigureAwait(false);
return true;
case MemoryTrackingAction.HardBreach:
_logger.LogError(
"Memory hard-breach on Tier {Tier} in-process driver {DriverId}: footprint={Footprint:N0} bytes. " +
"Recommending promotion to Tier C; NOT auto-killing (decisions #74, #145).",
_tier, _supervisor?.DriverInstanceId ?? "(unknown)", footprintBytes);
return false;
default:
return false;
}
}
}
@@ -1,144 +0,0 @@
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Core.Stability;
/// <summary>
/// Tier-agnostic memory-footprint tracker. Captures the post-initialize <b>baseline</b>
/// from the first samples after <c>IDriver.InitializeAsync</c>, then classifies each
/// subsequent sample against a hybrid soft/hard threshold per
/// <c>docs/v2/plan.md</c> — <c>soft = max(multiplier × baseline, baseline + floor)</c>,
/// <c>hard = 2 × soft</c>.
/// </summary>
/// <remarks>
/// <para>This tracker <b>never kills a process</b>. Soft and hard breaches
/// log + surface to the Admin UI via <c>DriverInstanceResilienceStatus</c>. The matching
/// process-level recycle protection lives in a separate <c>MemoryRecycle</c> that activates
/// for Tier C drivers only (where the driver runs out-of-process behind a supervisor that
/// can safely restart it without tearing down the OPC UA session or co-hosted in-proc
/// drivers).</para>
///
/// <para>Baseline capture: the tracker starts in <see cref="TrackingPhase.WarmingUp"/> for
/// <see cref="BaselineWindow"/> (default 5 min). During that window samples are collected;
/// the baseline is computed as the median once the window elapses. Before that point every
/// classification returns <see cref="MemoryTrackingAction.Warming"/>.</para>
/// </remarks>
public sealed class MemoryTracking
{
private readonly DriverTier _tier;
private readonly TimeSpan _baselineWindow;
private readonly List<long> _warmupSamples = [];
private long _baselineBytes;
private TrackingPhase _phase = TrackingPhase.WarmingUp;
private DateTime? _warmupStartUtc;
/// <summary>Tier-default multiplier/floor constants.</summary>
/// <param name="tier">The driver tier.</param>
/// <returns>The multiplier and floor-bytes constants for the tier.</returns>
public static (int Multiplier, long FloorBytes) GetTierConstants(DriverTier tier) => tier switch
{
DriverTier.A => (Multiplier: 3, FloorBytes: 50L * 1024 * 1024),
DriverTier.B => (Multiplier: 3, FloorBytes: 100L * 1024 * 1024),
DriverTier.C => (Multiplier: 2, FloorBytes: 500L * 1024 * 1024),
_ => throw new ArgumentOutOfRangeException(nameof(tier), tier, $"No memory-tracking constants defined for tier {tier}."),
};
/// <summary>Window over which post-init samples are collected to compute the baseline.</summary>
public TimeSpan BaselineWindow => _baselineWindow;
/// <summary>Current phase: <see cref="TrackingPhase.WarmingUp"/> or <see cref="TrackingPhase.Steady"/>.</summary>
public TrackingPhase Phase => _phase;
/// <summary>Captured baseline; 0 until warmup completes.</summary>
public long BaselineBytes => _baselineBytes;
/// <summary>Effective soft threshold (zero while warming up).</summary>
public long SoftThresholdBytes => _baselineBytes == 0 ? 0 : ComputeSoft(_tier, _baselineBytes);
/// <summary>Effective hard threshold = 2 × soft (zero while warming up).</summary>
public long HardThresholdBytes => _baselineBytes == 0 ? 0 : ComputeSoft(_tier, _baselineBytes) * 2;
/// <summary>Initializes a new instance of the <see cref="MemoryTracking"/> class.</summary>
/// <param name="tier">The driver tier for threshold constants.</param>
/// <param name="baselineWindow">Optional custom baseline window duration (default 5 minutes).</param>
public MemoryTracking(DriverTier tier, TimeSpan? baselineWindow = null)
{
_tier = tier;
_baselineWindow = baselineWindow ?? TimeSpan.FromMinutes(5);
}
/// <summary>
/// Submit a memory-footprint sample. Returns the action the caller should surface.
/// During warmup, always returns <see cref="MemoryTrackingAction.Warming"/> and accumulates
/// samples; once the window elapses the first steady-phase sample triggers baseline capture
/// (median of warmup samples).
/// </summary>
/// <param name="footprintBytes">The current memory footprint in bytes.</param>
/// <param name="utcNow">The current UTC time.</param>
/// <returns>The tracking action the caller should surface for this sample.</returns>
public MemoryTrackingAction Sample(long footprintBytes, DateTime utcNow)
{
if (_phase == TrackingPhase.WarmingUp)
{
_warmupStartUtc ??= utcNow;
_warmupSamples.Add(footprintBytes);
if (utcNow - _warmupStartUtc.Value >= _baselineWindow && _warmupSamples.Count > 0)
{
_baselineBytes = ComputeMedian(_warmupSamples);
_phase = TrackingPhase.Steady;
}
else
{
return MemoryTrackingAction.Warming;
}
}
if (footprintBytes >= HardThresholdBytes) return MemoryTrackingAction.HardBreach;
if (footprintBytes >= SoftThresholdBytes) return MemoryTrackingAction.SoftBreach;
return MemoryTrackingAction.None;
}
private static long ComputeSoft(DriverTier tier, long baseline)
{
var (multiplier, floor) = GetTierConstants(tier);
return Math.Max(multiplier * baseline, baseline + floor);
}
private static long ComputeMedian(List<long> samples)
{
var sorted = samples.Order().ToArray();
var mid = sorted.Length / 2;
return sorted.Length % 2 == 1
? sorted[mid]
: (sorted[mid - 1] + sorted[mid]) / 2;
}
}
/// <summary>Phase of a <see cref="MemoryTracking"/> lifecycle.</summary>
public enum TrackingPhase
{
/// <summary>Collecting post-init samples; baseline not yet computed.</summary>
WarmingUp,
/// <summary>Baseline captured; every sample classified against soft/hard thresholds.</summary>
Steady,
}
/// <summary>Classification the tracker returns per sample.</summary>
public enum MemoryTrackingAction
{
/// <summary>Baseline not yet captured; sample collected, no threshold check.</summary>
Warming,
/// <summary>Below soft threshold.</summary>
None,
/// <summary>Between soft and hard thresholds — log + surface, no action.</summary>
SoftBreach,
/// <summary>
/// ≥ hard threshold. Log + surface + (Tier C only, via <c>MemoryRecycle</c>) request
/// process recycle via the driver supervisor. Tier A/B breach never invokes any
/// kill path.
/// </summary>
HardBreach,
}
@@ -1,92 +0,0 @@
using Microsoft.Extensions.Logging;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Core.Stability;
/// <summary>
/// Tier C opt-in periodic-recycle driver.
/// A tick method advanced by the caller (fed by a background timer in prod; by test clock
/// in unit tests) decides whether the configured interval has elapsed and, if so, drives the
/// supplied <see cref="IDriverSupervisor"/> to recycle the Host.
/// </summary>
/// <remarks>
/// Tier A/B drivers MUST NOT use this class — scheduled recycle for in-process drivers would
/// kill every OPC UA session and every co-hosted driver. The ctor throws when constructed
/// with any tier other than C to make the misuse structurally impossible.
///
/// <para>Keeps no background thread of its own — callers invoke <see cref="TickAsync"/> on
/// their ambient scheduler tick (Phase 6.1 Stream C's health-endpoint host runs one). That
/// decouples the unit under test from wall-clock time and thread-pool scheduling.</para>
/// </remarks>
public sealed class ScheduledRecycleScheduler
{
private readonly TimeSpan _recycleInterval;
private readonly IDriverSupervisor _supervisor;
private readonly ILogger<ScheduledRecycleScheduler> _logger;
private DateTime _nextRecycleUtc;
/// <summary>
/// Construct the scheduler for a Tier C driver. Throws if <paramref name="tier"/> isn't C.
/// </summary>
/// <param name="tier">Driver tier; must be <see cref="DriverTier.C"/>.</param>
/// <param name="recycleInterval">Interval between recycles (e.g. 7 days).</param>
/// <param name="startUtc">Anchor time; next recycle fires at <paramref name="startUtc"/> + <paramref name="recycleInterval"/>.</param>
/// <param name="supervisor">Supervisor that performs the actual recycle.</param>
/// <param name="logger">Diagnostic sink.</param>
public ScheduledRecycleScheduler(
DriverTier tier,
TimeSpan recycleInterval,
DateTime startUtc,
IDriverSupervisor supervisor,
ILogger<ScheduledRecycleScheduler> logger)
{
if (tier != DriverTier.C)
throw new ArgumentException(
$"ScheduledRecycleScheduler is Tier C only (got {tier}). " +
"In-process drivers must not use scheduled recycle; see decisions #74 and #145.",
nameof(tier));
if (recycleInterval <= TimeSpan.Zero)
throw new ArgumentException("RecycleInterval must be positive.", nameof(recycleInterval));
_recycleInterval = recycleInterval;
_supervisor = supervisor;
_logger = logger;
_nextRecycleUtc = startUtc + recycleInterval;
}
/// <summary>Next scheduled recycle UTC. Advances by <see cref="RecycleInterval"/> on each fire.</summary>
public DateTime NextRecycleUtc => _nextRecycleUtc;
/// <summary>Recycle interval this scheduler was constructed with.</summary>
public TimeSpan RecycleInterval => _recycleInterval;
/// <summary>
/// Tick the scheduler forward. If <paramref name="utcNow"/> is past
/// <see cref="NextRecycleUtc"/>, requests a recycle from the supervisor and advances
/// <see cref="NextRecycleUtc"/> by exactly one interval. Returns true when a recycle fired.
/// </summary>
/// <param name="utcNow">The current UTC time.</param>
/// <param name="cancellationToken">The cancellation token.</param>
/// <returns>True if a recycle was triggered; false otherwise.</returns>
public async Task<bool> TickAsync(DateTime utcNow, CancellationToken cancellationToken)
{
if (utcNow < _nextRecycleUtc)
return false;
_logger.LogInformation(
"Scheduled recycle due for Tier C driver {DriverId} at {Now:o}; advancing next to {Next:o}.",
_supervisor.DriverInstanceId, utcNow, _nextRecycleUtc + _recycleInterval);
await _supervisor.RecycleAsync("Scheduled periodic recycle", cancellationToken).ConfigureAwait(false);
_nextRecycleUtc += _recycleInterval;
return true;
}
/// <summary>Request an immediate recycle outside the schedule (e.g. MemoryRecycle hard-breach escalation).</summary>
/// <param name="reason">The reason for requesting an immediate recycle.</param>
/// <param name="cancellationToken">The cancellation token.</param>
/// <returns>A task representing the asynchronous recycle operation.</returns>
public Task RequestRecycleNowAsync(string reason, CancellationToken cancellationToken) =>
_supervisor.RecycleAsync(reason, cancellationToken);
}
@@ -168,6 +168,7 @@ else
<th>Driver</th>
<th>Type</th>
<th>Status</th>
<th>Hosts</th>
<th>Last read</th>
<th>Errors/5 min</th>
<th>Last error</th>
@@ -177,7 +178,7 @@ else
@if (g.Drivers.Count == 0)
{
<tr>
<td colspan="6"><span class="text-muted">No drivers.</span></td>
<td colspan="7"><span class="text-muted">No drivers.</span></td>
</tr>
}
else
@@ -204,6 +205,28 @@ else
</td>
<td>@(d.DriverType ?? "—")</td>
<td><span class="chip @DriverChipClass(d.State)">@d.State</span></td>
<td>
@* Per-host connectivity (Gitea #521). The point of this column is the case
the driver-level Status chip cannot express: a multi-device driver stays
Healthy in aggregate while ONE of its PLCs is unreachable. A driver with
no probe shows "—" — deliberately distinct from a probe reporting zero
hosts, which shows "0 hosts". *@
@if (d.HostStatuses is null)
{
<span class="text-muted">—</span>
}
else if (d.DegradedHosts.Count == 0)
{
<span class="text-muted small">@d.HostStatuses.Count @(d.HostStatuses.Count == 1 ? "host" : "hosts")</span>
}
else
{
<span class="chip chip-warn"
title="@DegradedHostTitle(d)">
@d.DegradedHosts.Count / @d.HostStatuses.Count down
</span>
}
</td>
<td>@(d.LastSuccessfulReadUtc?.ToString("HH:mm:ss 'UTC'") ?? "—")</td>
<td class="numeric">@d.ErrorCount5Min</td>
<td><span class="text-muted small">@(d.LastError ?? "—")</span></td>
@@ -355,6 +378,13 @@ else
_ => "chip-idle",
};
// Tooltip listing each host that is not Running, with its state and when it last changed. Built in
// C# rather than inline in the markup so the string is composed once per render, and so the row can
// never be the thing that 500s the page — string.Join over an already-materialised list does no
// indexing (cf. #504, where slicing a short DB string in Razor took down the whole page).
private static string DegradedHostTitle(HostsDriverRow d) =>
string.Join(" · ", d.DegradedHosts.Select(h => $"{h.HostName}: {h.State} since {h.LastChangedUtc:u}"));
public async ValueTask DisposeAsync()
{
// Unsubscribe first so the singleton store can't invoke a handler on a disposed component.
@@ -20,10 +20,12 @@
</div>
}
<div class="row g-3">
<div class="col-md-4"><label class="form-label">Recycle interval (s, Tier C only)</label>
<input type="number" class="form-control form-control-sm" @bind="_m.RecycleIntervalSeconds" @bind:after="EmitAsync" placeholder="none" disabled="@_m.ParseFailed" /></div>
</div>
@* The "Recycle interval (s, Tier C only)" input was removed in Gitea #522. It authored a knob
with no effect: the scheduled-recycle machinery it fed was constructed only in tests, and
the IDriverSupervisor that would have performed the recycle had no implementations. The
tier model it belonged to assumed an out-of-process driver Host that no longer exists
anywhere. Offering an operator a control that silently does nothing is worse than not
offering it. *@
<div class="table-wrap mt-3">
<table class="data-table">
@@ -23,8 +23,12 @@ public sealed class ResilienceFormModel
public static readonly string[] Capabilities =
["Read", "Write", "Discover", "Subscribe", "Probe", "AlarmSubscribe", "AlarmAcknowledge", "HistoryRead"];
/// <summary>Gets or sets the driver recycle-interval override, in seconds; null = use the tier default.</summary>
public int? RecycleIntervalSeconds { get; set; }
// RecycleIntervalSeconds is GONE (Gitea #522) — the scheduled-recycle machinery it authored was
// deleted, so the field configured nothing. It is deliberately NOT explicitly stripped from stored
// JSON: the model's preserve-unknown-keys contract now covers it, so an existing blob's
// "recycleIntervalSeconds" survives an unrelated edit in _bag rather than being silently dropped,
// and the parser ignores it. Removing the key is a migration decision, not a side effect of
// opening the form.
// capability name -> (timeout, retry, breaker), each nullable.
/// <summary>Gets or sets the per-capability resilience overrides, keyed by capability name.</summary>
@@ -112,7 +116,6 @@ public sealed class ResilienceFormModel
}
model._bag = bag;
model.RecycleIntervalSeconds = GetIntOrNull(bag, "recycleIntervalSeconds");
if (bag.TryGetPropertyValue("capabilityPolicies", out var cpNode) && cpNode is JsonObject cp)
{
@@ -141,8 +144,6 @@ public sealed class ResilienceFormModel
{
if (ParseFailed) return RawStoredJson;
SetOrRemoveInt(_bag, "recycleIntervalSeconds", RecycleIntervalSeconds);
var cp = _bag.TryGetPropertyValue("capabilityPolicies", out var cpNode) && cpNode is JsonObject existing
? existing
: null;
@@ -1,6 +1,7 @@
namespace ZB.MOM.WW.OtOpcUa.AdminUI.Hosts;
using ZB.MOM.WW.OtOpcUa.Commons.Messages.Drivers;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
/// <summary>
/// One configured host node within a cluster, as the <c>/hosts</c> page needs it: the cluster
@@ -40,10 +41,26 @@ public sealed record HostsDriverInstanceInfo(string DriverInstanceId, string Clu
/// is unchanged and an operator must re-browse the device via <c>/raw</c> to pick anything up.</param>
/// <param name="RediscoveryReason">The driver-supplied reason for that report; null when
/// <paramref name="RediscoveryNeededUtc"/> is null.</param>
/// <param name="HostStatuses">Per-host connectivity for a multi-device driver, ordered by host name; null
/// when the driver reports no per-host detail. Lets one unreachable device show even while the driver
/// row itself is Healthy.</param>
public sealed record HostsDriverRow(
string DriverInstanceId, string? Name, string? DriverType, string State,
DateTime? LastSuccessfulReadUtc, string? LastError, int ErrorCount5Min, DateTime PublishedUtc,
DateTime? RediscoveryNeededUtc = null, string? RediscoveryReason = null);
DateTime? RediscoveryNeededUtc = null, string? RediscoveryReason = null,
IReadOnlyList<HostConnectivityStatus>? HostStatuses = null)
{
/// <summary>Hosts that are not <see cref="HostState.Running"/>, ordered by name — the ones worth an
/// operator's attention. Empty when every host is fine or none are reported.
/// <para><see cref="HostState.Unknown"/> counts as degraded on purpose: a probe that has not yet
/// completed its first tick, or one a driver failed to start (AbCip logs exactly this case), reports
/// Unknown — and silently rendering that as healthy is how the gap got missed the first time.</para></summary>
public IReadOnlyList<HostConnectivityStatus> DegradedHosts { get; } =
(HostStatuses ?? [])
.Where(h => h.State != HostState.Running)
.OrderBy(h => h.HostName, StringComparer.OrdinalIgnoreCase)
.ToList();
}
/// <summary>
/// One cluster's section on the <c>/hosts</c> page: its configured nodes plus its enriched
@@ -118,7 +135,8 @@ public static class HostsDriverView
s.ErrorCount5Min,
s.PublishedUtc,
s.RediscoveryNeededUtc,
s.RediscoveryReason);
s.RediscoveryReason,
s.HostStatuses);
})
.OrderBy(d => d.Name ?? d.DriverInstanceId, StringComparer.OrdinalIgnoreCase)
.ThenBy(d => d.DriverInstanceId, StringComparer.OrdinalIgnoreCase)
@@ -3,6 +3,10 @@ using ZB.MOM.WW.OtOpcUa.Commons.Messages.Alerts;
using ZB.MOM.WW.OtOpcUa.Commons.Messages.Drivers;
using ZB.MOM.WW.OtOpcUa.Commons.Messages.Logging;
using ZB.MOM.WW.OtOpcUa.Commons.Protos.Telemetry.V1;
// Aliased, not imported wholesale: Core.Abstractions also declares a DriverHealth, which would collide
// with the proto DriverHealth this file maps.
using HostConnectivityStatus = ZB.MOM.WW.OtOpcUa.Core.Abstractions.HostConnectivityStatus;
using HostState = ZB.MOM.WW.OtOpcUa.Core.Abstractions.HostState;
namespace ZB.MOM.WW.OtOpcUa.ControlPlane.Telemetry;
@@ -120,7 +124,32 @@ public static class TelemetryProtoMapCentral
ErrorCount5Min: msg.ErrorCount5Min,
PublishedUtc: Required(msg.PublishedUtc, "DriverHealth", "published_utc"),
RediscoveryNeededUtc: msg.RediscoveryNeededUtc?.ToDateTime(),
RediscoveryReason: msg.HasRediscoveryReason ? msg.RediscoveryReason : null);
RediscoveryReason: msg.HasRediscoveryReason ? msg.RediscoveryReason : null,
HostStatuses: ToHostStatuses(msg));
}
/// <summary>
/// Projects the repeated host-connectivity field, honouring the explicit presence flag: null when
/// the node said the driver has no probe, an empty list when it has one that knows no hosts.
/// <para>An unparseable state string degrades to <see cref="HostState.Unknown"/> rather than
/// throwing — a node running a newer build that added an enum member must not be able to kill
/// central's telemetry stream, which is observability and has no business failing closed.</para>
/// </summary>
private static IReadOnlyList<HostConnectivityStatus>? ToHostStatuses(DriverHealth msg)
{
if (!msg.HasHostStatuses) return null;
var result = new List<HostConnectivityStatus>(msg.HostStatuses.Count);
foreach (var h in msg.HostStatuses)
{
result.Add(new HostConnectivityStatus(
h.HostName,
// System.Enum qualified: Google.Protobuf.WellKnownTypes also declares an Enum type.
System.Enum.TryParse<HostState>(h.State, ignoreCase: true, out var state) ? state : HostState.Unknown,
h.LastChangedUtc?.ToDateTime() ?? default));
}
return result;
}
/// <summary>Projects a <see cref="DriverResilienceStatus"/> onto a <see cref="DriverResilienceStatusChanged"/>.</summary>
@@ -133,6 +133,21 @@ public static class TelemetryProtoMapNode
if (e.RediscoveryReason is not null)
msg.RediscoveryReason = e.RediscoveryReason;
// Presence flag first — an empty repeated field cannot say whether the driver has a probe at all.
if (e.HostStatuses is not null)
{
msg.HasHostStatuses = true;
foreach (var h in e.HostStatuses)
{
msg.HostStatuses.Add(new HostConnectivity
{
HostName = h.HostName ?? "",
State = h.State.ToString(),
LastChangedUtc = ToUtcTimestamp(h.LastChangedUtc),
});
}
}
return msg;
}
@@ -36,7 +36,8 @@ public sealed class AkkaDriverHealthPublisher : IDriverHealthPublisher
DriverHealth health,
int errorCount5Min,
DateTime? rediscoveryNeededUtc = null,
string? rediscoveryReason = null)
string? rediscoveryReason = null,
IReadOnlyList<HostConnectivityStatus>? hostStatuses = null)
{
var msg = new DriverHealthChanged(
clusterId,
@@ -47,7 +48,8 @@ public sealed class AkkaDriverHealthPublisher : IDriverHealthPublisher
errorCount5Min,
DateTime.UtcNow,
rediscoveryNeededUtc,
rediscoveryReason);
rediscoveryReason,
hostStatuses);
DistributedPubSub.Get(_system).Mediator.Tell(new Publish(TopicName, msg));
// Phase 5: fan the same snapshot into the node-local live-telemetry hub (no-op until a gRPC
// client subscribes). The DPS publish above is unchanged — the hub is a strictly additive tap.
@@ -120,6 +120,14 @@ public sealed class DriverInstanceActor : ReceiveActor, IWithTimers
/// connection affinity (a Galaxy redeploy or a TwinCAT symbol-version bump can land while the driver is
/// between connects), and dropping it in one state would lose the signal silently.</summary>
private sealed record RediscoveryRaised(RediscoveryEventArgs Args);
/// <summary>Self-sent when the wrapped driver raises <see cref="IHostConnectivityProbe.OnHostStatusChanged"/>
/// — one of its hosts went Running ↔ Stopped ↔ Faulted. Marshals the event off the driver's probe thread
/// onto the actor thread. Carries NO payload: the handler re-pulls
/// <see cref="IHostConnectivityProbe.GetHostStatuses"/>, so the driver stays the single source of truth
/// for the host set and a host that appeared since the last publish is picked up too. Handled in every
/// behaviour for the same reason as <see cref="RediscoveryRaised"/> — a probe tick can land while the
/// driver is between connects.</summary>
private sealed record HostStatusRaised;
public sealed class RetryConnect
{
@@ -171,6 +179,7 @@ public sealed class DriverInstanceActor : ReceiveActor, IWithTimers
private EventHandler<DataChangeEventArgs>? _dataChangeHandler;
private EventHandler<AlarmEventArgs>? _alarmEventHandler;
private EventHandler<RediscoveryEventArgs>? _rediscoveryHandler;
private EventHandler<HostStatusChangedEventArgs>? _hostStatusHandler;
/// <summary>When the driver last raised <see cref="IRediscoverable.OnRediscoveryNeeded"/>, and the reason
/// it gave. Null until the first raise. Carried on every subsequent health snapshot so the AdminUI can
@@ -306,6 +315,7 @@ public sealed class DriverInstanceActor : ReceiveActor, IWithTimers
// Attach the rediscovery signal before the first publish. Not per-connect: an IRediscoverable raise
// has no connection affinity, and a driver can observe a remote change while disconnected.
AttachRediscoverySource();
AttachHostStatusSource();
PublishHealthSnapshot();
Timers.StartPeriodicTimer("health-poll", HealthPollTick.Instance, _healthPollInterval);
}
@@ -325,6 +335,7 @@ public sealed class DriverInstanceActor : ReceiveActor, IWithTimers
// Stubbed drivers never enter Connected, so they never kick discovery; swallow defensively in case a
// re-discovery self-tick is ever routed here so it doesn't surface as an Akka Unhandled message.
Receive<RediscoveryRaised>(HandleRediscoveryRaised);
Receive<HostStatusRaised>(_ => PublishHealthSnapshot());
Receive<HealthPollTick>(_ => PublishHealthSnapshot());
}
@@ -378,6 +389,7 @@ public sealed class DriverInstanceActor : ReceiveActor, IWithTimers
// this state; swallow it so it doesn't dead-letter — the next Connected entry re-subscribes.
Receive<SubscribeAlarms>(_ => { });
Receive<RediscoveryRaised>(HandleRediscoveryRaised);
Receive<HostStatusRaised>(_ => PublishHealthSnapshot());
Receive<HealthPollTick>(_ => PublishHealthSnapshot());
}
@@ -438,6 +450,7 @@ public sealed class DriverInstanceActor : ReceiveActor, IWithTimers
Receive<SubscriptionFailed>(msg =>
_log.Debug("DriverInstance {Id}: resubscribe reported failure: {Reason}", _driverInstanceId, msg.Reason));
Receive<RediscoveryRaised>(HandleRediscoveryRaised);
Receive<HostStatusRaised>(_ => PublishHealthSnapshot());
Receive<HealthPollTick>(_ =>
{
PublishHealthSnapshot();
@@ -541,6 +554,7 @@ public sealed class DriverInstanceActor : ReceiveActor, IWithTimers
// this state; swallow it so it doesn't dead-letter — the next Connected entry re-subscribes.
Receive<SubscribeAlarms>(_ => { });
Receive<RediscoveryRaised>(HandleRediscoveryRaised);
Receive<HostStatusRaised>(_ => PublishHealthSnapshot());
Receive<HealthPollTick>(_ => PublishHealthSnapshot());
Timers.StartPeriodicTimer("retry-connect", RetryConnect.Instance, _reconnectInterval);
}
@@ -809,6 +823,48 @@ public sealed class DriverInstanceActor : ReceiveActor, IWithTimers
_rediscoveryHandler = null;
}
/// <summary>Subscribe the driver's <see cref="IHostConnectivityProbe.OnHostStatusChanged"/> (if it is
/// one), marshaling each transition to the actor thread. Idempotent; mirrors
/// <see cref="AttachRediscoverySource"/>, including the PreStart-not-per-connect placement — a probe
/// loop can report a host down while the driver itself is between connects.</summary>
private void AttachHostStatusSource()
{
if (_driver is not IHostConnectivityProbe src || _hostStatusHandler is not null) return;
var self = Self;
_hostStatusHandler = (_, _) => self.Tell(new HostStatusRaised());
src.OnHostStatusChanged += _hostStatusHandler;
}
/// <summary>Symmetric teardown, called from PostStop — same leak argument as
/// <see cref="DetachRediscoverySource"/>.</summary>
private void DetachHostStatusSource()
{
if (_driver is IHostConnectivityProbe src && _hostStatusHandler is not null)
src.OnHostStatusChanged -= _hostStatusHandler;
_hostStatusHandler = null;
}
/// <summary>Current per-host connectivity, or null when the driver is not an
/// <see cref="IHostConnectivityProbe"/>. Pulled fresh rather than cached: the driver owns the host set,
/// and a stale local copy would be a second source of truth to keep in sync.
/// <para><c>GetHostStatuses()</c> is called UNGUARDED (not through <see cref="_invoker"/>) on purpose —
/// it is a pure in-memory snapshot with no I/O, which is exactly why
/// <c>UnwrappedCapabilityCallAnalyzer</c> exempts it. A driver that makes it do I/O breaks that
/// contract; the try/catch below keeps a misbehaving one from killing the health publish.</para></summary>
private IReadOnlyList<HostConnectivityStatus>? CurrentHostStatuses()
{
if (_driver is not IHostConnectivityProbe probe) return null;
try
{
return probe.GetHostStatuses();
}
catch (Exception ex)
{
_log.Warning(ex, "DriverInstance {Id}: GetHostStatuses threw during health publish; omitting host detail", _driverInstanceId);
return null;
}
}
/// <summary>Records the driver's rediscovery raise and re-publishes health so the signal reaches the
/// AdminUI promptly rather than waiting for the next 30 s heartbeat.
/// <para><b>Advisory only.</b> The served address space is deliberately NOT rebuilt: v3 authors raw tags
@@ -961,16 +1017,26 @@ public sealed class DriverInstanceActor : ReceiveActor, IWithTimers
{
var health = _driver.GetHealth();
var errorCount = ErrorCount5Min();
var hostStatuses = CurrentHostStatuses();
// _rediscoveryNeededUtc is PART OF THE FINGERPRINT on purpose. A rediscovery raise on an
// otherwise-unchanged Healthy driver leaves (state, lastSuccess, lastError, errorCount)
// identical, so without it the dedup below would swallow the very publish that carries the
// signal and the operator would never see the prompt.
var fingerprint = (health.State, health.LastSuccessfulRead, health.LastError, errorCount, _rediscoveryNeededUtc);
//
// The host-status digest is in the fingerprint for EXACTLY the same reason, and the failure
// mode is the more likely one of the two: a multi-device driver stays aggregate-Healthy when a
// single device drops, so every other fingerprint component is unchanged and the transition —
// the whole point of this channel — would be deduped away. It is a flattened STRING because a
// tuple holding IReadOnlyList compares by reference, which would never match and so would
// defeat the dedup in the opposite direction (re-publishing every 30 s heartbeat).
var fingerprint = (health.State, health.LastSuccessfulRead, health.LastError, errorCount,
_rediscoveryNeededUtc, HostStatusDigest(hostStatuses));
if (_lastPublishedFingerprint is { } prev && prev.Equals(fingerprint))
return;
_lastPublishedFingerprint = fingerprint;
_healthPublisher.Publish(
_clusterId, _driverInstanceId, health, errorCount, _rediscoveryNeededUtc, _rediscoveryReason);
_clusterId, _driverInstanceId, health, errorCount, _rediscoveryNeededUtc, _rediscoveryReason,
hostStatuses);
}
catch (Exception ex)
{
@@ -978,8 +1044,22 @@ public sealed class DriverInstanceActor : ReceiveActor, IWithTimers
}
}
/// <summary>Order-insensitive value digest of a host-status list for the publish fingerprint. Null in,
/// null out — so "driver is not a probe" and "probe reports no hosts" stay distinguishable rather than
/// both collapsing to the empty string.
/// <para>Ordered by host name because a driver is free to return its hosts in any order (several build
/// the list from a <c>Dictionary</c>), and an order flip would otherwise read as a real transition and
/// re-publish forever.</para></summary>
private static string? HostStatusDigest(IReadOnlyList<HostConnectivityStatus>? statuses) =>
statuses is null
? null
: string.Join('|', statuses
.OrderBy(s => s.HostName, StringComparer.Ordinal)
.Select(s => $"{s.HostName}={s.State}@{s.LastChangedUtc:O}"));
/// <summary>Fingerprint of the last <see cref="PublishHealthSnapshot"/> call; null until first publish.</summary>
private (DriverState State, DateTime? LastSuccess, string? LastError, int ErrorCount, DateTime? RediscoveryNeededUtc)? _lastPublishedFingerprint;
private (DriverState State, DateTime? LastSuccess, string? LastError, int ErrorCount,
DateTime? RediscoveryNeededUtc, string? HostStatusDigest)? _lastPublishedFingerprint;
/// <inheritdoc />
protected override void PostStop()
@@ -988,6 +1068,7 @@ public sealed class DriverInstanceActor : ReceiveActor, IWithTimers
// MUST happen: the IDriver instance can outlive this actor (the host respawns a child around the
// same driver object), so a missing unsubscribe accumulates a handler per respawn holding a dead Self.
DetachRediscoverySource();
DetachHostStatusSource();
try { _driver.ShutdownAsync(CancellationToken.None).GetAwaiter().GetResult(); }
catch (Exception ex) { _log.Warning(ex, "DriverInstance {Id}: ShutdownAsync threw on PostStop", _driverInstanceId); }
OtOpcUaTelemetry.DriverInstanceLifecycle.Add(1,
@@ -1,131 +0,0 @@
using Microsoft.EntityFrameworkCore;
using Shouldly;
using Xunit;
using ZB.MOM.WW.OtOpcUa.Configuration.Entities;
using ZB.MOM.WW.OtOpcUa.Configuration.Enums;
namespace ZB.MOM.WW.OtOpcUa.Configuration.Tests;
/// <summary>
/// End-to-end round-trip through the DB for the <see cref="DriverHostStatus"/> entity
/// added in PR 33 — exercises the composite primary key (NodeId, DriverInstanceId,
/// HostName), string-backed <c>DriverHostState</c> conversion, and the two indexes the
/// Admin UI's drill-down queries will scan (NodeId, LastSeenUtc).
/// </summary>
[Trait("Category", "SchemaCompliance")]
[Collection(nameof(SchemaComplianceCollection))]
public sealed class DriverHostStatusTests(SchemaComplianceFixture fixture)
{
/// <summary>Verifies that the composite key allows the same host across different nodes or drivers.</summary>
[Fact]
public async Task Composite_key_allows_same_host_across_different_nodes_or_drivers()
{
await using var ctx = NewContext();
// Same HostName + DriverInstanceId across two different server nodes — classic 2-node
// redundancy case. Both rows must be insertable because each server node owns its own
// runtime view of the shared host.
var now = DateTime.UtcNow;
ctx.DriverHostStatuses.Add(new DriverHostStatus
{
NodeId = "node-a", DriverInstanceId = "galaxy-1", HostName = "GRPlatform",
State = DriverHostState.Running,
StateChangedUtc = now, LastSeenUtc = now,
});
ctx.DriverHostStatuses.Add(new DriverHostStatus
{
NodeId = "node-b", DriverInstanceId = "galaxy-1", HostName = "GRPlatform",
State = DriverHostState.Stopped,
StateChangedUtc = now, LastSeenUtc = now,
Detail = "secondary hasn't taken over yet",
});
// Same server node + host, different driver instance — second driver doesn't clobber.
ctx.DriverHostStatuses.Add(new DriverHostStatus
{
NodeId = "node-a", DriverInstanceId = "modbus-plc1", HostName = "GRPlatform",
State = DriverHostState.Running,
StateChangedUtc = now, LastSeenUtc = now,
});
await ctx.SaveChangesAsync();
var rows = await ctx.DriverHostStatuses.AsNoTracking()
.Where(r => r.HostName == "GRPlatform").ToListAsync();
rows.Count.ShouldBe(3);
rows.ShouldContain(r => r.NodeId == "node-a" && r.DriverInstanceId == "galaxy-1");
rows.ShouldContain(r => r.NodeId == "node-b" && r.State == DriverHostState.Stopped && r.Detail == "secondary hasn't taken over yet");
rows.ShouldContain(r => r.NodeId == "node-a" && r.DriverInstanceId == "modbus-plc1");
}
/// <summary>Verifies that the upsert pattern updates existing records in place.</summary>
[Fact]
public async Task Upsert_pattern_for_same_key_updates_in_place()
{
// The publisher hosted service (follow-up PR) upserts on every transition +
// heartbeat. This test pins the two-step pattern it will use: check-then-add-or-update
// keyed on the composite PK. If the composite key ever changes, this test breaks
// loudly so the publisher gets a synchronized update.
await using var ctx = NewContext();
var t0 = DateTime.UtcNow;
ctx.DriverHostStatuses.Add(new DriverHostStatus
{
NodeId = "upsert-node", DriverInstanceId = "upsert-driver", HostName = "upsert-host",
State = DriverHostState.Running,
StateChangedUtc = t0, LastSeenUtc = t0,
});
await ctx.SaveChangesAsync();
var t1 = t0.AddSeconds(30);
await using (var ctx2 = NewContext())
{
var existing = await ctx2.DriverHostStatuses.SingleAsync(r =>
r.NodeId == "upsert-node" && r.DriverInstanceId == "upsert-driver" && r.HostName == "upsert-host");
existing.State = DriverHostState.Faulted;
existing.StateChangedUtc = t1;
existing.LastSeenUtc = t1;
existing.Detail = "transport reset by peer";
await ctx2.SaveChangesAsync();
}
await using var ctx3 = NewContext();
var final = await ctx3.DriverHostStatuses.AsNoTracking().SingleAsync(r =>
r.NodeId == "upsert-node" && r.HostName == "upsert-host");
final.State.ShouldBe(DriverHostState.Faulted);
final.Detail.ShouldBe("transport reset by peer");
// Only one row — a naive "always insert" would have created a duplicate PK and thrown.
(await ctx3.DriverHostStatuses.CountAsync(r => r.NodeId == "upsert-node")).ShouldBe(1);
}
/// <summary>Verifies that the State enum is persisted as a string, not an integer.</summary>
[Fact]
public async Task Enum_persists_as_string_not_int()
{
// Fluent config sets HasConversion<string>() on State — the DB stores 'Running' /
// 'Stopped' / 'Faulted' / 'Unknown' as nvarchar(16). Verify by reading the raw
// string back via ADO; if someone drops the conversion the column will contain '1'
// / '2' / '3' and this assertion fails. Matters because DBAs inspecting the table
// directly should see readable state names, not enum ordinals.
await using var ctx = NewContext();
ctx.DriverHostStatuses.Add(new DriverHostStatus
{
NodeId = "enum-node", DriverInstanceId = "enum-driver", HostName = "enum-host",
State = DriverHostState.Faulted,
StateChangedUtc = DateTime.UtcNow, LastSeenUtc = DateTime.UtcNow,
});
await ctx.SaveChangesAsync();
await using var conn = fixture.OpenConnection();
using var cmd = conn.CreateCommand();
cmd.CommandText = "SELECT [State] FROM DriverHostStatus WHERE NodeId = 'enum-node'";
var rawValue = (string?)await cmd.ExecuteScalarAsync();
rawValue.ShouldBe("Faulted");
}
private OtOpcUaConfigDbContext NewContext()
{
var options = new DbContextOptionsBuilder<OtOpcUaConfigDbContext>()
.UseSqlServer(fixture.ConnectionString)
.Options;
return new OtOpcUaConfigDbContext(options);
}
}
@@ -33,7 +33,9 @@ public sealed class SchemaComplianceTests
"RawFolder", "TagGroup", "UnsTagReference",
"DriverInstance", "Device", "Equipment", "Tag", "PollGroup", "VirtualTag",
"NodeAcl", "ExternalIdReservation",
"DriverHostStatus",
// DriverHostStatus was dropped in Gitea #521 — nothing ever wrote a row, and per-cluster
// mesh Phase 4 made the publisher its entity doc described unbuildable (a driver-only node
// has no ConfigDb connection). Per-host connectivity now rides DriverHealthChanged to /hosts.
"DriverInstanceResilienceStatus",
"LdapGroupRoleMapping",
// v3 dropped the orphaned EquipmentImportBatch / EquipmentImportRow tables.
@@ -1,214 +0,0 @@
using Shouldly;
using Xunit;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Core.Abstractions.Tests;
public sealed class DriverTypeRegistryTests
{
private static DriverTypeMetadata SampleMetadata(
string typeName = "Modbus",
DriverTier tier = DriverTier.B) =>
new(typeName,
DriverConfigJsonSchema: "{\"type\": \"object\"}",
DeviceConfigJsonSchema: "{\"type\": \"object\"}",
TagConfigJsonSchema: "{\"type\": \"object\"}",
Tier: tier);
/// <summary>
/// Verifies that Register followed by Get round-trips metadata correctly.
/// </summary>
[Fact]
public void Register_ThenGet_RoundTrips()
{
var registry = new DriverTypeRegistry();
var metadata = SampleMetadata();
registry.Register(metadata);
registry.Get("Modbus").ShouldBe(metadata);
}
/// <summary>
/// Verifies that Register accepts and preserves non-null tier values.
/// </summary>
/// <param name="tier">The driver tier to test.</param>
[Theory]
[InlineData(DriverTier.A)]
[InlineData(DriverTier.B)]
[InlineData(DriverTier.C)]
public void Register_Requires_NonNullTier(DriverTier tier)
{
var registry = new DriverTypeRegistry();
var metadata = SampleMetadata(typeName: $"Driver-{tier}", tier: tier);
registry.Register(metadata);
registry.Get(metadata.TypeName).Tier.ShouldBe(tier);
}
/// <summary>
/// Verifies that Get is case-insensitive when looking up driver types.
/// </summary>
[Fact]
public void Get_IsCaseInsensitive()
{
var registry = new DriverTypeRegistry();
registry.Register(SampleMetadata("Galaxy"));
registry.Get("galaxy").ShouldNotBeNull();
registry.Get("GALAXY").ShouldNotBeNull();
}
/// <summary>
/// Verifies that Get throws when looking up an unknown type.
/// </summary>
[Fact]
public void Get_UnknownType_Throws()
{
var registry = new DriverTypeRegistry();
registry.Register(SampleMetadata("Modbus"));
Should.Throw<KeyNotFoundException>(() => registry.Get("UnregisteredType"));
}
/// <summary>
/// Verifies that TryGet returns null when looking up an unknown type.
/// </summary>
[Fact]
public void TryGet_UnknownType_ReturnsNull()
{
var registry = new DriverTypeRegistry();
registry.Register(SampleMetadata("Modbus"));
registry.TryGet("UnregisteredType").ShouldBeNull();
}
/// <summary>
/// Verifies that Register throws when attempting to register a duplicate type.
/// </summary>
[Fact]
public void Register_DuplicateType_Throws()
{
var registry = new DriverTypeRegistry();
registry.Register(SampleMetadata("Modbus"));
Should.Throw<InvalidOperationException>(() => registry.Register(SampleMetadata("Modbus")));
}
/// <summary>
/// Verifies that duplicate type detection is case-insensitive.
/// </summary>
[Fact]
public void Register_DuplicateTypeIsCaseInsensitive()
{
var registry = new DriverTypeRegistry();
registry.Register(SampleMetadata("Modbus"));
Should.Throw<InvalidOperationException>(() => registry.Register(SampleMetadata("modbus")));
}
/// <summary>
/// Verifies that All returns all registered driver types.
/// </summary>
[Fact]
public void All_ReturnsRegisteredTypes()
{
var registry = new DriverTypeRegistry();
registry.Register(SampleMetadata("Modbus"));
registry.Register(SampleMetadata("S7"));
registry.Register(SampleMetadata("Galaxy"));
var all = registry.All();
all.Count.ShouldBe(3);
all.Select(m => m.TypeName).ShouldBe(new[] { "Modbus", "S7", "Galaxy" }, ignoreOrder: true);
}
/// <summary>
/// Verifies that Get rejects empty or null type names.
/// </summary>
/// <param name="typeName">The type name to test (null, empty, or whitespace).</param>
[Theory]
[InlineData(null)]
[InlineData("")]
[InlineData(" ")]
public void Get_RejectsEmptyTypeName(string? typeName)
{
var registry = new DriverTypeRegistry();
Should.Throw<ArgumentException>(() => registry.Get(typeName!));
}
/// <summary>
/// Core.Abstractions-004: concurrent <see cref="DriverTypeRegistry.Register"/> calls for
/// two different driver types must both succeed and both end up visible to readers. A
/// non-atomic check-then-act would let the second swap silently discard the first
/// registration; this test asserts the "registered only once per process" guarantee
/// holds under concurrent writers too.
/// </summary>
[Fact]
public void Register_ConcurrentDistinctTypes_AllSucceed()
{
var registry = new DriverTypeRegistry();
const int parallelism = 32;
var ready = new ManualResetEventSlim(false);
var threads = Enumerable.Range(0, parallelism)
.Select(i => new Thread(() =>
{
ready.Wait();
registry.Register(SampleMetadata($"Driver-{i}"));
}))
.ToArray();
foreach (var t in threads) t.Start();
ready.Set(); // release all threads simultaneously
foreach (var t in threads) t.Join();
var all = registry.All();
all.Count.ShouldBe(parallelism);
for (var i = 0; i < parallelism; i++)
registry.TryGet($"Driver-{i}").ShouldNotBeNull(
$"Driver-{i} was lost to a race in concurrent Register");
}
/// <summary>
/// Core.Abstractions-004: concurrent <see cref="DriverTypeRegistry.Register"/> calls for
/// the SAME driver type must result in exactly one successful registration and exactly
/// (parallelism - 1) <see cref="InvalidOperationException"/> throws — not a silent
/// last-writer-wins replacement.
/// </summary>
[Fact]
public void Register_ConcurrentDuplicateType_ExactlyOneWins()
{
var registry = new DriverTypeRegistry();
const int parallelism = 16;
var ready = new ManualResetEventSlim(false);
var successes = 0;
var failures = 0;
var threads = Enumerable.Range(0, parallelism)
.Select(_ => new Thread(() =>
{
ready.Wait();
try
{
registry.Register(SampleMetadata("Contended"));
Interlocked.Increment(ref successes);
}
catch (InvalidOperationException)
{
Interlocked.Increment(ref failures);
}
}))
.ToArray();
foreach (var t in threads) t.Start();
ready.Set();
foreach (var t in threads) t.Join();
successes.ShouldBe(1);
failures.ShouldBe(parallelism - 1);
registry.All().Count.ShouldBe(1);
}
}
@@ -143,13 +143,13 @@ public sealed class DriverResilienceOptionsParserTests
public void PropertyNames_AreCaseInsensitive()
{
var json = """
{ "RECYCLEINTERVALSECONDS": 3600 }
{ "CAPABILITYPOLICIES": { "Read": { "retryCount": 7 } } }
""";
var options = DriverResilienceOptionsParser.ParseOrDefaults(DriverTier.C, json, out var diag);
diag.ShouldBeNull();
options.RecycleIntervalSeconds.ShouldBe(3600);
options.Resolve(DriverCapability.Read).RetryCount.ShouldBe(7);
}
/// <summary>Verifies that capability name is case insensitive.</summary>
@@ -182,51 +182,31 @@ public sealed class DriverResilienceOptionsParserTests
options.Resolve(cap).ShouldBe(DriverResilienceOptions.GetTierDefaults(tier)[cap]);
}
/// <summary>Verifies that RecycleIntervalSeconds on Tier C with positive value parses and surfaces.</summary>
[Fact]
public void RecycleIntervalSeconds_TierC_PositiveValue_ParsesAndSurfaces()
{
var options = DriverResilienceOptionsParser.ParseOrDefaults(
DriverTier.C, "{\"recycleIntervalSeconds\":3600}", out var diag);
diag.ShouldBeNull();
options.RecycleIntervalSeconds.ShouldBe(3600);
}
/// <summary>Verifies that RecycleIntervalSeconds when null defaults to null.</summary>
[Fact]
public void RecycleIntervalSeconds_Null_DefaultsToNull()
{
var options = DriverResilienceOptionsParser.ParseOrDefaults(DriverTier.C, "{}", out _);
options.RecycleIntervalSeconds.ShouldBeNull();
}
/// <summary>Verifies that RecycleIntervalSeconds on Tier A or B is rejected with diagnostic.</summary>
/// <param name="tier">The driver tier to test.</param>
/// <summary>
/// <b>Backward-compatibility guard for the #522 deletion.</b> The four
/// <c>RecycleIntervalSeconds</c> tests that lived here are gone with the knob they covered — the
/// scheduled-recycle machinery it configured was constructed only in tests, and the
/// <c>IDriverSupervisor</c> that would have performed the recycle had no implementations.
/// <para>What still matters is that a <b>deployed</b> <c>ResilienceConfig</c> blob written before
/// the removal keeps parsing. The removed property must be ignored as an unknown key, NOT rejected
/// — a driver whose stored config suddenly failed to parse would fall back to tier defaults and
/// silently discard the operator's real timeout/retry overrides alongside the dead one.</para>
/// </summary>
[Theory]
[InlineData(DriverTier.A)]
[InlineData(DriverTier.B)]
public void RecycleIntervalSeconds_OnTierAorB_Rejected_With_Diagnostic(DriverTier tier)
[InlineData(DriverTier.C)]
public void A_legacy_blob_carrying_the_removed_recycle_key_still_parses_cleanly(DriverTier tier)
{
// Decision #74 — in-process drivers must not scheduled-recycle because it would
// tear down every OPC UA session. The parser surfaces a diagnostic rather than
// silently honouring the value.
var options = DriverResilienceOptionsParser.ParseOrDefaults(
tier, "{\"recycleIntervalSeconds\":3600}", out var diag);
var json = """
{ "recycleIntervalSeconds": 3600, "capabilityPolicies": { "Read": { "retryCount": 5 } } }
""";
options.RecycleIntervalSeconds.ShouldBeNull();
diag.ShouldContain("Tier C only");
}
var options = DriverResilienceOptionsParser.ParseOrDefaults(tier, json, out var diag);
/// <summary>Verifies that RecycleIntervalSeconds with non-positive value is rejected with diagnostic.</summary>
[Fact]
public void RecycleIntervalSeconds_NonPositive_Rejected_With_Diagnostic()
{
var options = DriverResilienceOptionsParser.ParseOrDefaults(
DriverTier.C, "{\"recycleIntervalSeconds\":0}", out var diag);
options.RecycleIntervalSeconds.ShouldBeNull();
diag.ShouldContain("must be positive");
diag.ShouldBeNull("an unknown key must be ignored silently, not reported as a misconfiguration");
options.Tier.ShouldBe(tier);
options.Resolve(DriverCapability.Read).RetryCount.ShouldBe(5,
"the operator's real overrides must survive alongside the dead key");
}
// ---- 01/S-6: range clamps (clamp + diagnostic, never brick) ----
@@ -15,14 +15,18 @@ public sealed class DriverResilienceOptionsTests
/// knob (the bulkhead genre this pass deleted) is inertness the interface-forwarding and
/// unwrapped-dispatch guards can't catch. To add a knob: wire it in <c>DriverResiliencePipelineBuilder</c>,
/// add a behavior test that proves it engages, THEN add it to the expected set below.
/// (<c>RecycleIntervalSeconds</c> is itself Tier-C-dormant per 01/U-2 — a documented open question,
/// explicitly out of this pass's scope; it stays in the expected list as the recycle scheduler, not
/// the Polly pipeline, is its consumer.)
/// <para><c>RecycleIntervalSeconds</c> used to sit in the expected set under an explicit carve-out —
/// Tier-C-dormant per 01/U-2, "a documented open question, out of this pass's scope", justified by
/// the recycle scheduler rather than the Polly pipeline being its consumer. Gitea #522 closed that
/// question by deleting the scheduler (it was constructed only in tests, and the
/// <c>IDriverSupervisor</c> that would perform the recycle had no implementations), so the knob went
/// with it and the carve-out is gone. The expected set below is now literally what the name of this
/// test claims.</para>
/// </summary>
[Fact]
public void Options_properties_are_exactly_the_pipeline_wired_set()
{
var expected = new[] { "Tier", "CapabilityPolicies", "RecycleIntervalSeconds" };
var expected = new[] { "Tier", "CapabilityPolicies" };
var actual = typeof(DriverResilienceOptions)
.GetProperties()
@@ -1,105 +0,0 @@
using Microsoft.Extensions.Logging.Abstractions;
using Shouldly;
using Xunit;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
using ZB.MOM.WW.OtOpcUa.Core.Stability;
namespace ZB.MOM.WW.OtOpcUa.Core.Tests.Stability;
[Trait("Category", "Unit")]
public sealed class MemoryRecycleTests
{
/// <summary>Verifies that Tier C hard memory breach requests supervisor recycle.</summary>
[Fact]
public async Task TierC_HardBreach_RequestsSupervisorRecycle()
{
var supervisor = new FakeSupervisor();
var recycle = new MemoryRecycle(DriverTier.C, supervisor, NullLogger<MemoryRecycle>.Instance);
var requested = await recycle.HandleAsync(MemoryTrackingAction.HardBreach, 2_000_000_000, CancellationToken.None);
requested.ShouldBeTrue();
supervisor.RecycleCount.ShouldBe(1);
supervisor.LastReason.ShouldContain("hard-breach");
}
/// <summary>Verifies that Tier A and B hard memory breach never request recycle.</summary>
/// <param name="tier">The driver tier to test.</param>
[Theory]
[InlineData(DriverTier.A)]
[InlineData(DriverTier.B)]
public async Task InProcessTier_HardBreach_NeverRequestsRecycle(DriverTier tier)
{
var supervisor = new FakeSupervisor();
var recycle = new MemoryRecycle(tier, supervisor, NullLogger<MemoryRecycle>.Instance);
var requested = await recycle.HandleAsync(MemoryTrackingAction.HardBreach, 2_000_000_000, CancellationToken.None);
requested.ShouldBeFalse("Tier A/B hard-breach logs a promotion recommendation only (decisions #74, #145)");
supervisor.RecycleCount.ShouldBe(0);
}
/// <summary>Verifies that Tier C without supervisor hard breach is a no-op.</summary>
[Fact]
public async Task TierC_WithoutSupervisor_HardBreach_NoOp()
{
var recycle = new MemoryRecycle(DriverTier.C, supervisor: null, NullLogger<MemoryRecycle>.Instance);
var requested = await recycle.HandleAsync(MemoryTrackingAction.HardBreach, 2_000_000_000, CancellationToken.None);
requested.ShouldBeFalse("no supervisor → no recycle path; action logged only");
}
/// <summary>Verifies that soft memory breach never requests recycle at any tier.</summary>
/// <param name="tier">The driver tier to test.</param>
[Theory]
[InlineData(DriverTier.A)]
[InlineData(DriverTier.B)]
[InlineData(DriverTier.C)]
public async Task SoftBreach_NeverRequestsRecycle(DriverTier tier)
{
var supervisor = new FakeSupervisor();
var recycle = new MemoryRecycle(tier, supervisor, NullLogger<MemoryRecycle>.Instance);
var requested = await recycle.HandleAsync(MemoryTrackingAction.SoftBreach, 1_000_000_000, CancellationToken.None);
requested.ShouldBeFalse("soft-breach is surface-only at every tier");
supervisor.RecycleCount.ShouldBe(0);
}
/// <summary>Verifies that non-breach memory actions are no-ops.</summary>
/// <param name="action">The non-breach memory tracking action to test.</param>
[Theory]
[InlineData(MemoryTrackingAction.None)]
[InlineData(MemoryTrackingAction.Warming)]
public async Task NonBreachActions_NoOp(MemoryTrackingAction action)
{
var supervisor = new FakeSupervisor();
var recycle = new MemoryRecycle(DriverTier.C, supervisor, NullLogger<MemoryRecycle>.Instance);
var requested = await recycle.HandleAsync(action, 100_000_000, CancellationToken.None);
requested.ShouldBeFalse();
supervisor.RecycleCount.ShouldBe(0);
}
private sealed class FakeSupervisor : IDriverSupervisor
{
/// <summary>Gets the driver instance identifier.</summary>
public string DriverInstanceId => "fake-tier-c";
/// <summary>Gets the count of recycle operations.</summary>
public int RecycleCount { get; private set; }
/// <summary>Gets the reason from the last recycle operation.</summary>
public string? LastReason { get; private set; }
/// <summary>Recycles the driver asynchronously.</summary>
/// <param name="reason">The reason for recycling.</param>
/// <param name="cancellationToken">The cancellation token.</param>
public Task RecycleAsync(string reason, CancellationToken cancellationToken)
{
RecycleCount++;
LastReason = reason;
return Task.CompletedTask;
}
}
}
@@ -1,132 +0,0 @@
using Shouldly;
using Xunit;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
using ZB.MOM.WW.OtOpcUa.Core.Stability;
namespace ZB.MOM.WW.OtOpcUa.Core.Tests.Stability;
[Trait("Category", "Unit")]
public sealed class MemoryTrackingTests
{
private static readonly DateTime T0 = new(2026, 4, 19, 12, 0, 0, DateTimeKind.Utc);
/// <summary>Verifies that warming phase returns Warming until the time window elapses.</summary>
[Fact]
public void WarmingUp_Returns_Warming_UntilWindowElapses()
{
var tracker = new MemoryTracking(DriverTier.A, TimeSpan.FromMinutes(5));
tracker.Sample(100_000_000, T0).ShouldBe(MemoryTrackingAction.Warming);
tracker.Sample(105_000_000, T0.AddMinutes(1)).ShouldBe(MemoryTrackingAction.Warming);
tracker.Sample(102_000_000, T0.AddMinutes(4.9)).ShouldBe(MemoryTrackingAction.Warming);
tracker.Phase.ShouldBe(TrackingPhase.WarmingUp);
tracker.BaselineBytes.ShouldBe(0);
}
/// <summary>Verifies that when the window elapses, baseline is captured as median and phase transitions to steady.</summary>
[Fact]
public void WindowElapsed_CapturesBaselineAsMedian_AndTransitionsToSteady()
{
var tracker = new MemoryTracking(DriverTier.A, TimeSpan.FromMinutes(5));
tracker.Sample(100_000_000, T0);
tracker.Sample(200_000_000, T0.AddMinutes(1));
tracker.Sample(150_000_000, T0.AddMinutes(2));
var first = tracker.Sample(150_000_000, T0.AddMinutes(5));
tracker.Phase.ShouldBe(TrackingPhase.Steady);
tracker.BaselineBytes.ShouldBe(150_000_000L, "median of 4 samples [100, 200, 150, 150] = (150+150)/2 = 150");
first.ShouldBe(MemoryTrackingAction.None, "150 MB is the baseline itself, well under soft threshold");
}
/// <summary>Verifies that tier constants match Decision 146 specification.</summary>
/// <param name="tier">The driver tier to test.</param>
/// <param name="expectedMultiplier">Expected growth multiplier.</param>
/// <param name="expectedFloorMB">Expected floor in megabytes.</param>
[Theory]
[InlineData(DriverTier.A, 3, 50)]
[InlineData(DriverTier.B, 3, 100)]
[InlineData(DriverTier.C, 2, 500)]
public void GetTierConstants_MatchesDecision146(DriverTier tier, int expectedMultiplier, long expectedFloorMB)
{
var (multiplier, floor) = MemoryTracking.GetTierConstants(tier);
multiplier.ShouldBe(expectedMultiplier);
floor.ShouldBe(expectedFloorMB * 1024 * 1024);
}
/// <summary>Verifies that soft threshold uses the maximum of multiplier and floor for small baselines.</summary>
[Fact]
public void SoftThreshold_UsesMax_OfMultiplierAndFloor_SmallBaseline()
{
// Tier A: mult=3, floor=50 MB. Baseline 10 MB → 3×10=30 MB < 10+50=60 MB → floor wins.
var tracker = WarmupWithBaseline(DriverTier.A, 10L * 1024 * 1024);
tracker.SoftThresholdBytes.ShouldBe(60L * 1024 * 1024);
}
/// <summary>Verifies that soft threshold uses the maximum of multiplier and floor for large baselines.</summary>
[Fact]
public void SoftThreshold_UsesMax_OfMultiplierAndFloor_LargeBaseline()
{
// Tier A: mult=3, floor=50 MB. Baseline 200 MB → 3×200=600 MB > 200+50=250 MB → multiplier wins.
var tracker = WarmupWithBaseline(DriverTier.A, 200L * 1024 * 1024);
tracker.SoftThresholdBytes.ShouldBe(600L * 1024 * 1024);
}
/// <summary>Verifies that hard threshold is twice the soft threshold.</summary>
[Fact]
public void HardThreshold_IsTwiceSoft()
{
var tracker = WarmupWithBaseline(DriverTier.B, 200L * 1024 * 1024);
tracker.HardThresholdBytes.ShouldBe(tracker.SoftThresholdBytes * 2);
}
/// <summary>Verifies that samples below soft threshold return None.</summary>
[Fact]
public void Sample_Below_Soft_Returns_None()
{
var tracker = WarmupWithBaseline(DriverTier.A, 100L * 1024 * 1024);
tracker.Sample(200L * 1024 * 1024, T0.AddMinutes(10)).ShouldBe(MemoryTrackingAction.None);
}
/// <summary>Verifies that samples at soft threshold return SoftBreach.</summary>
[Fact]
public void Sample_AtSoft_Returns_SoftBreach()
{
// Tier A, baseline 200 MB → soft = 600 MB. Sample exactly at soft.
var tracker = WarmupWithBaseline(DriverTier.A, 200L * 1024 * 1024);
tracker.Sample(tracker.SoftThresholdBytes, T0.AddMinutes(10))
.ShouldBe(MemoryTrackingAction.SoftBreach);
}
/// <summary>Verifies that samples at hard threshold return HardBreach.</summary>
[Fact]
public void Sample_AtHard_Returns_HardBreach()
{
var tracker = WarmupWithBaseline(DriverTier.A, 200L * 1024 * 1024);
tracker.Sample(tracker.HardThresholdBytes, T0.AddMinutes(10))
.ShouldBe(MemoryTrackingAction.HardBreach);
}
/// <summary>Verifies that samples above hard threshold return HardBreach.</summary>
[Fact]
public void Sample_AboveHard_Returns_HardBreach()
{
var tracker = WarmupWithBaseline(DriverTier.A, 200L * 1024 * 1024);
tracker.Sample(tracker.HardThresholdBytes + 100_000_000, T0.AddMinutes(10))
.ShouldBe(MemoryTrackingAction.HardBreach);
}
private static MemoryTracking WarmupWithBaseline(DriverTier tier, long baseline)
{
var tracker = new MemoryTracking(tier, TimeSpan.FromMinutes(5));
tracker.Sample(baseline, T0);
tracker.Sample(baseline, T0.AddMinutes(5));
tracker.BaselineBytes.ShouldBe(baseline);
return tracker;
}
}
@@ -1,118 +0,0 @@
using Microsoft.Extensions.Logging.Abstractions;
using Shouldly;
using Xunit;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
using ZB.MOM.WW.OtOpcUa.Core.Stability;
namespace ZB.MOM.WW.OtOpcUa.Core.Tests.Stability;
[Trait("Category", "Unit")]
public sealed class ScheduledRecycleSchedulerTests
{
private static readonly DateTime T0 = new(2026, 4, 19, 0, 0, 0, DateTimeKind.Utc);
private static readonly TimeSpan Weekly = TimeSpan.FromDays(7);
/// <summary>Verifies constructor throws for Tier A or B.</summary>
/// <param name="tier">The driver tier to test.</param>
[Theory]
[InlineData(DriverTier.A)]
[InlineData(DriverTier.B)]
public void TierAOrB_Ctor_Throws(DriverTier tier)
{
var supervisor = new FakeSupervisor();
Should.Throw<ArgumentException>(() => new ScheduledRecycleScheduler(
tier, Weekly, T0, supervisor, NullLogger<ScheduledRecycleScheduler>.Instance));
}
/// <summary>Verifies constructor throws for zero or negative intervals.</summary>
[Fact]
public void ZeroOrNegativeInterval_Throws()
{
var supervisor = new FakeSupervisor();
Should.Throw<ArgumentException>(() => new ScheduledRecycleScheduler(
DriverTier.C, TimeSpan.Zero, T0, supervisor, NullLogger<ScheduledRecycleScheduler>.Instance));
Should.Throw<ArgumentException>(() => new ScheduledRecycleScheduler(
DriverTier.C, TimeSpan.FromSeconds(-1), T0, supervisor, NullLogger<ScheduledRecycleScheduler>.Instance));
}
/// <summary>Verifies Tick before the next recycle time is a no-op.</summary>
[Fact]
public async Task Tick_BeforeNextRecycle_NoOp()
{
var supervisor = new FakeSupervisor();
var sch = new ScheduledRecycleScheduler(DriverTier.C, Weekly, T0, supervisor, NullLogger<ScheduledRecycleScheduler>.Instance);
var fired = await sch.TickAsync(T0 + TimeSpan.FromDays(6), CancellationToken.None);
fired.ShouldBeFalse();
supervisor.RecycleCount.ShouldBe(0);
}
/// <summary>Verifies Tick at or after the next recycle time fires once and advances.</summary>
[Fact]
public async Task Tick_AtOrAfterNextRecycle_FiresOnce_AndAdvances()
{
var supervisor = new FakeSupervisor();
var sch = new ScheduledRecycleScheduler(DriverTier.C, Weekly, T0, supervisor, NullLogger<ScheduledRecycleScheduler>.Instance);
var fired = await sch.TickAsync(T0 + Weekly + TimeSpan.FromMinutes(1), CancellationToken.None);
fired.ShouldBeTrue();
supervisor.RecycleCount.ShouldBe(1);
sch.NextRecycleUtc.ShouldBe(T0 + Weekly + Weekly);
}
/// <summary>Verifies RequestRecycleNow fires immediately without advancing the schedule.</summary>
[Fact]
public async Task RequestRecycleNow_Fires_Immediately_WithoutAdvancingSchedule()
{
var supervisor = new FakeSupervisor();
var sch = new ScheduledRecycleScheduler(DriverTier.C, Weekly, T0, supervisor, NullLogger<ScheduledRecycleScheduler>.Instance);
var nextBefore = sch.NextRecycleUtc;
await sch.RequestRecycleNowAsync("memory hard-breach", CancellationToken.None);
supervisor.RecycleCount.ShouldBe(1);
supervisor.LastReason.ShouldBe("memory hard-breach");
sch.NextRecycleUtc.ShouldBe(nextBefore, "ad-hoc recycle doesn't shift the cron schedule");
}
/// <summary>Verifies multiple ticks across the recycle interval each advance by one interval.</summary>
[Fact]
public async Task MultipleFires_AcrossTicks_AdvanceOneIntervalEach()
{
var supervisor = new FakeSupervisor();
var sch = new ScheduledRecycleScheduler(DriverTier.C, TimeSpan.FromDays(1), T0, supervisor, NullLogger<ScheduledRecycleScheduler>.Instance);
await sch.TickAsync(T0 + TimeSpan.FromDays(1) + TimeSpan.FromHours(1), CancellationToken.None);
await sch.TickAsync(T0 + TimeSpan.FromDays(2) + TimeSpan.FromHours(1), CancellationToken.None);
await sch.TickAsync(T0 + TimeSpan.FromDays(3) + TimeSpan.FromHours(1), CancellationToken.None);
supervisor.RecycleCount.ShouldBe(3);
sch.NextRecycleUtc.ShouldBe(T0 + TimeSpan.FromDays(4));
}
/// <summary>Fake driver supervisor for testing.</summary>
private sealed class FakeSupervisor : IDriverSupervisor
{
/// <summary>Gets the driver instance ID.</summary>
public string DriverInstanceId => "tier-c-fake";
/// <summary>Gets the number of times RecycleAsync was called.</summary>
public int RecycleCount { get; private set; }
/// <summary>Gets the reason from the most recent recycle call.</summary>
public string? LastReason { get; private set; }
/// <summary>Simulates a driver recycle operation.</summary>
/// <param name="reason">The reason for the recycle.</param>
/// <param name="cancellationToken">Cancellation token.</param>
/// <returns>A completed task.</returns>
public Task RecycleAsync(string reason, CancellationToken cancellationToken)
{
RecycleCount++;
LastReason = reason;
return Task.CompletedTask;
}
}
}
@@ -2,6 +2,7 @@ using Shouldly;
using Xunit;
using ZB.MOM.WW.OtOpcUa.AdminUI.Hosts;
using ZB.MOM.WW.OtOpcUa.Commons.Messages.Drivers;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.AdminUI.Tests.Hosts;
@@ -154,4 +155,62 @@ public sealed class HostsDriverViewTests
groups.Select(g => g.ClusterId).ShouldBe(new[] { "Alpha", "Beta", "zeta" });
}
/// <summary>
/// Per-host connectivity flows through to the row (Gitea #521), and <c>DegradedHosts</c> picks out
/// exactly the hosts an operator needs to look at. This is the case the driver-level Status chip
/// cannot express: the driver is Healthy and one of its devices is not.
/// </summary>
[Fact]
public void Build_carries_host_statuses_and_flags_only_the_degraded_ones()
{
var snapshot = Snap("MAIN", "drv-a") with
{
HostStatuses =
[
new HostConnectivityStatus("plc-a", HostState.Running, When),
new HostConnectivityStatus("plc-b", HostState.Stopped, When),
new HostConnectivityStatus("plc-c", HostState.Faulted, When),
],
};
var row = HostsDriverView.Build([snapshot], nodes: null, instances: null).Single().Drivers.Single();
row.State.ShouldBe("Healthy");
row.HostStatuses!.Count.ShouldBe(3);
row.DegradedHosts.Select(h => h.HostName).ShouldBe(["plc-b", "plc-c"]);
}
/// <summary>
/// <see cref="HostState.Unknown"/> counts as degraded. A probe that has not completed its first tick
/// — or one a driver failed to start at all, which AbCip logs explicitly — reports Unknown, and
/// rendering that as healthy is how an unstarted probe stays invisible.
/// </summary>
[Fact]
public void Unknown_host_state_counts_as_degraded()
{
var snapshot = Snap("MAIN", "drv-a") with
{
HostStatuses = [new HostConnectivityStatus("plc-a", HostState.Unknown, When)],
};
var row = HostsDriverView.Build([snapshot], nodes: null, instances: null).Single().Drivers.Single();
row.DegradedHosts.ShouldHaveSingleItem().HostName.ShouldBe("plc-a");
}
/// <summary>
/// A driver with no probe keeps a null host list — distinct from a probe reporting zero hosts. The
/// /hosts column renders "—" for the former and "0 hosts" for the latter, and collapsing them would
/// claim every probe-less driver's devices are fine.
/// </summary>
[Fact]
public void A_driver_without_host_statuses_keeps_null_and_reports_no_degraded_hosts()
{
var row = HostsDriverView.Build([Snap("MAIN", "drv-a")], nodes: null, instances: null)
.Single().Drivers.Single();
row.HostStatuses.ShouldBeNull();
row.DegradedHosts.ShouldBeEmpty();
}
}
@@ -67,7 +67,7 @@ public class ResilienceFormModelTests
[Fact]
public void Partial_override_round_trips()
{
var m = new ResilienceFormModel { RecycleIntervalSeconds = 3600 };
var m = new ResilienceFormModel();
m.Policies["Read"].TimeoutSeconds = 5;
m.Policies["Read"].RetryCount = 5;
@@ -75,16 +75,37 @@ public class ResilienceFormModelTests
json.ShouldNotBeNull();
var back = ResilienceFormModel.FromJson(json);
back.RecycleIntervalSeconds.ShouldBe(3600);
back.Policies["Read"].TimeoutSeconds.ShouldBe(5);
back.Policies["Write"].IsEmpty.ShouldBeTrue();
}
/// <summary>
/// The removed <c>recycleIntervalSeconds</c> field (Gitea #522) must be PRESERVED in a stored blob
/// rather than stripped, via the model's preserve-unknown-keys contract. Opening a driver's
/// resilience form and saving an unrelated change must not silently rewrite the operator's stored
/// config — dropping the key is a migration decision, not a side effect of a page visit.
/// </summary>
[Fact]
public void The_removed_recycle_key_survives_a_load_save_as_an_unknown_key()
{
const string stored = """
{ "recycleIntervalSeconds": 3600, "capabilityPolicies": { "Read": { "retryCount": 5 } } }
""";
var m = ResilienceFormModel.FromJson(stored);
m.Policies["Read"].TimeoutSeconds = 9; // an unrelated edit
var json = m.ToJson();
json.ShouldNotBeNull();
json.ShouldContain("recycleIntervalSeconds");
json.ShouldContain("3600");
}
[Fact]
public void Malformed_json_yields_empty_model()
{
var m = ResilienceFormModel.FromJson("{ not valid json");
m.RecycleIntervalSeconds.ShouldBeNull();
m.Policies["Read"].IsEmpty.ShouldBeTrue();
}
@@ -3,6 +3,9 @@ using Shouldly;
using ZB.MOM.WW.OtOpcUa.Commons.Protos;
using ZB.MOM.WW.OtOpcUa.Commons.Protos.Telemetry.V1;
using ZB.MOM.WW.OtOpcUa.ControlPlane.Telemetry;
// Aliased, not imported wholesale: Core.Abstractions also declares a DriverHealth, which collides with
// the proto DriverHealth these tests construct.
using HostState = ZB.MOM.WW.OtOpcUa.Core.Abstractions.HostState;
using Xunit;
namespace ZB.MOM.WW.OtOpcUa.ControlPlane.Tests.Telemetry;
@@ -168,6 +171,77 @@ public sealed class TelemetryProtoMapCentralTests
e.PublishedUtc.Kind.ShouldBe(DateTimeKind.Utc);
}
/// <summary>
/// Per-host connectivity (Gitea #521) survives the wire with its tri-state intact:
/// <b>null</b> (driver has no probe) must stay distinguishable from an <b>empty list</b> (it has one
/// that knows no hosts). proto3 cannot tell an absent repeated field from an empty one, which is why
/// <c>has_host_statuses</c> exists — without it a driver with no probe would arrive looking like one
/// whose devices are all fine, and the /hosts column would render "0 hosts" for every driver in the
/// fleet.
/// </summary>
[Fact]
public void ToHealth_host_statuses_round_trip_with_the_null_vs_empty_distinction_intact()
{
var populated = new DriverHealth
{
ClusterId = "c1", DriverInstanceId = "d1", State = "Healthy",
PublishedUtc = Timestamp.FromDateTime(SampleUtc),
HasHostStatuses = true,
HostStatuses =
{
new HostConnectivity { HostName = "plc-a", State = "Running", LastChangedUtc = Timestamp.FromDateTime(OtherUtc) },
new HostConnectivity { HostName = "plc-b", State = "Stopped", LastChangedUtc = Timestamp.FromDateTime(SampleUtc) },
},
};
var mapped = TelemetryProtoMapCentral.ToHealth(populated).HostStatuses;
mapped.ShouldNotBeNull();
mapped!.Count.ShouldBe(2);
mapped[0].HostName.ShouldBe("plc-a");
mapped[0].State.ShouldBe(HostState.Running);
mapped[0].LastChangedUtc.ShouldBe(OtherUtc);
mapped[1].State.ShouldBe(HostState.Stopped);
// A probe that currently knows no hosts: empty, NOT null.
var empty = new DriverHealth
{
ClusterId = "c1", DriverInstanceId = "d1", State = "Healthy",
PublishedUtc = Timestamp.FromDateTime(SampleUtc),
HasHostStatuses = true,
};
TelemetryProtoMapCentral.ToHealth(empty).HostStatuses.ShouldBeEmpty();
// No probe at all: null, NOT empty.
var absent = new DriverHealth
{
ClusterId = "c1", DriverInstanceId = "d1", State = "Healthy",
PublishedUtc = Timestamp.FromDateTime(SampleUtc),
};
TelemetryProtoMapCentral.ToHealth(absent).HostStatuses.ShouldBeNull();
}
/// <summary>
/// An unparseable host state degrades to <see cref="HostState.Unknown"/> rather than throwing. A node
/// running a newer build that added an enum member must not be able to kill central's telemetry
/// stream — this is observability, and it has no business failing closed.
/// </summary>
[Fact]
public void ToHealth_unknown_host_state_string_degrades_instead_of_throwing()
{
var proto = new DriverHealth
{
ClusterId = "c1", DriverInstanceId = "d1", State = "Healthy",
PublishedUtc = Timestamp.FromDateTime(SampleUtc),
HasHostStatuses = true,
HostStatuses = { new HostConnectivity { HostName = "plc-a", State = "Quiescing" } },
};
var mapped = TelemetryProtoMapCentral.ToHealth(proto).HostStatuses;
mapped.ShouldNotBeNull();
mapped![0].State.ShouldBe(HostState.Unknown);
}
[Fact]
public void ToHealth_absent_nullable_timestamp_and_optional_string_map_to_null()
{
@@ -8,6 +8,10 @@ using ZB.MOM.WW.OtOpcUa.Commons.Messages.Logging;
using ZB.MOM.WW.OtOpcUa.Commons.Protos.Telemetry.V1;
using ZB.MOM.WW.OtOpcUa.Host.Grpc;
using ZB.MOM.WW.OtOpcUa.Runtime.Telemetry;
// Aliased, not imported wholesale: Core.Abstractions also declares a DriverHealth, which would collide
// with the proto DriverHealth these tests assert on.
using HostConnectivityStatus = ZB.MOM.WW.OtOpcUa.Core.Abstractions.HostConnectivityStatus;
using HostState = ZB.MOM.WW.OtOpcUa.Core.Abstractions.HostState;
namespace ZB.MOM.WW.OtOpcUa.Host.Tests.Grpc;
@@ -210,6 +214,40 @@ public sealed class TelemetryStreamGrpcServiceTests
evt.DriverHealth.LastSuccessfulReadUtc.ToDateTime().ShouldBe(expectedUtc);
}
/// <summary>
/// Node side of the #521 host-status carry: the presence flag must be written, so the tri-state
/// (no probe / probe with no hosts / probe with hosts) survives a wire proto3 cannot express on the
/// repeated field alone. Paired with the central-side decode in <c>TelemetryProtoMapCentralTests</c>.
/// </summary>
[Fact]
public void ToProto_health_writes_the_host_status_presence_flag_and_entries()
{
var changedAt = new DateTime(2026, 7, 30, 11, 0, 0, DateTimeKind.Utc);
var withHosts = new DriverHealthChanged(
"cluster-a", "drv-1", "Healthy", null, null, 0, DateTime.UtcNow,
HostStatuses: [new HostConnectivityStatus("plc-a", HostState.Stopped, changedAt)]);
var evt = TelemetryProtoMapNode.ToProto(new TelemetryItem.Health(withHosts), "c");
evt.DriverHealth.HasHostStatuses.ShouldBeTrue();
evt.DriverHealth.HostStatuses.Count.ShouldBe(1);
evt.DriverHealth.HostStatuses[0].HostName.ShouldBe("plc-a");
evt.DriverHealth.HostStatuses[0].State.ShouldBe("Stopped");
evt.DriverHealth.HostStatuses[0].LastChangedUtc.ToDateTime().ShouldBe(changedAt);
// A probe reporting zero hosts still sets the flag — that is the whole point of having one.
var emptyProbe = new DriverHealthChanged(
"cluster-a", "drv-1", "Healthy", null, null, 0, DateTime.UtcNow, HostStatuses: []);
var emptyEvt = TelemetryProtoMapNode.ToProto(new TelemetryItem.Health(emptyProbe), "c");
emptyEvt.DriverHealth.HasHostStatuses.ShouldBeTrue();
emptyEvt.DriverHealth.HostStatuses.ShouldBeEmpty();
// No probe: flag clear.
var noProbe = new DriverHealthChanged("cluster-a", "drv-1", "Healthy", null, null, 0, DateTime.UtcNow);
TelemetryProtoMapNode.ToProto(new TelemetryItem.Health(noProbe), "c")
.DriverHealth.HasHostStatuses.ShouldBeFalse();
}
[Fact]
public async Task Client_disconnect_mid_stream_ends_cleanly_without_leaking_a_slot()
{
@@ -0,0 +1,298 @@
using Akka.Actor;
using Shouldly;
using Xunit;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
using ZB.MOM.WW.OtOpcUa.Runtime.Drivers;
using ZB.MOM.WW.OtOpcUa.Runtime.Tests.Harness;
namespace ZB.MOM.WW.OtOpcUa.Runtime.Tests.Drivers;
/// <summary>
/// Covers the <see cref="IHostConnectivityProbe"/> consumer (Gitea #521). Eleven drivers implement the
/// capability; before this wiring <c>GetHostStatuses()</c> had ZERO production call sites and
/// <c>OnHostStatusChanged</c> had no subscriber outside the Galaxy driver's own aggregator, so per-host
/// connectivity was computed by every driver and read by nobody.
/// <para><b>Why it does not go to the DB.</b> The <c>DriverHostStatus</c> table's doc-comment described a
/// publisher hosted service that upserted rows from each driver node. It was never built, and
/// per-cluster mesh Phase 4 made it unbuildable as described — <c>AddOtOpcUaConfigDb</c> is gated on the
/// <c>admin</c> role, so a driver-only node has no ConfigDb connection. The table was dropped; the data
/// rides the driver-health snapshot instead.</para>
/// </summary>
[Trait("Category", "Unit")]
public sealed class DriverInstanceActorHostStatusTests : RuntimeActorTestBase
{
/// <summary>The base case: a probe driver's hosts reach the health publisher.</summary>
[Fact]
public void Host_statuses_reach_the_health_publisher()
{
var driver = new ProbeStubDriver();
driver.SetHost("plc-a", HostState.Running);
driver.SetHost("plc-b", HostState.Running);
var publisher = new RecordingHealthPublisher();
var actor = SpawnDriverActor(driver, publisher);
actor.Tell(new DriverInstanceActor.InitializeRequested("{}"));
AwaitAssert(
() =>
{
var latest = publisher.Published.LastOrDefault();
latest.ShouldNotBeNull();
latest!.HostStatuses.ShouldNotBeNull();
latest.HostStatuses!.Select(h => h.HostName).OrderBy(n => n, StringComparer.Ordinal)
.ShouldBe(["plc-a", "plc-b"]);
},
TimeSpan.FromSeconds(3));
}
/// <summary>
/// <b>The load-bearing case, and the entire reason this channel exists.</b> A multi-device driver
/// stays aggregate-<c>Healthy</c> when ONE of its devices drops — the driver-level state chip cannot
/// express it. So the transition must reach the operator through the per-host detail.
/// <para>This is also the dedup trap that already bit the rediscovery signal once.
/// <c>PublishHealthSnapshot</c> suppresses a publish whose fingerprint repeats, and on this
/// transition (state, lastSuccessfulRead, lastError, errorCount) are ALL unchanged. Unless the
/// host-status digest is part of the fingerprint, the dedup swallows exactly the publish that
/// carries the news.</para>
/// <para><b>Falsifiability:</b> the assertion is that the publish count STRICTLY INCREASES across the
/// transition — an "eventually shows Stopped" assertion would be satisfied by the warm-up publish
/// plus a later 30 s heartbeat and would prove nothing. Drop <c>HostStatusDigest</c> from the
/// fingerprint tuple in <c>DriverInstanceActor</c> and this test must go red. Verified by doing so.</para>
/// </summary>
[Fact]
public void A_single_host_going_down_is_not_swallowed_by_the_unchanged_health_dedup()
{
var driver = new ProbeStubDriver();
driver.SetHost("plc-a", HostState.Running);
driver.SetHost("plc-b", HostState.Running);
var publisher = new RecordingHealthPublisher();
var actor = SpawnDriverActor(driver, publisher);
actor.Tell(new DriverInstanceActor.InitializeRequested("{}"));
AwaitAssert(() => publisher.Published.Count.ShouldBeGreaterThan(0), TimeSpan.FromSeconds(3));
// Settle, so the baseline is a quiet actor: from here only the host state changes. The driver's
// OWN health stays Healthy throughout — that is the point.
ExpectNoMsg(TimeSpan.FromMilliseconds(200));
var before = publisher.Published.Count;
driver.SetHost("plc-b", HostState.Stopped, raise: true);
AwaitAssert(
() => publisher.Published.Count.ShouldBeGreaterThan(before),
TimeSpan.FromSeconds(3));
var latest = publisher.Published[^1];
latest.Health.State.ShouldBe(DriverState.Healthy, "the driver itself never faulted — only one of its devices did");
latest.HostStatuses.ShouldNotBeNull();
latest.HostStatuses!.Single(h => h.HostName == "plc-b").State.ShouldBe(HostState.Stopped);
latest.HostStatuses.Single(h => h.HostName == "plc-a").State.ShouldBe(HostState.Running);
}
/// <summary>
/// The other half of the dedup contract: when nothing changes, the digest must NOT churn. A digest
/// built over an <c>IReadOnlyList</c> by reference, or one sensitive to the order a driver happens to
/// enumerate its hosts in, would differ on every call and re-publish on every 30 s heartbeat forever
/// — turning the dedup off without anyone noticing.
/// </summary>
[Fact]
public void Unchanged_hosts_do_not_defeat_the_dedup()
{
var driver = new ProbeStubDriver();
driver.SetHost("plc-a", HostState.Running);
driver.SetHost("plc-b", HostState.Running);
var publisher = new RecordingHealthPublisher();
var actor = SpawnDriverActor(driver, publisher);
actor.Tell(new DriverInstanceActor.InitializeRequested("{}"));
AwaitAssert(() => publisher.Published.Count.ShouldBeGreaterThan(0), TimeSpan.FromSeconds(3));
ExpectNoMsg(TimeSpan.FromMilliseconds(200));
var before = publisher.Published.Count;
// The driver re-shuffles its host order without changing any state. A real driver builds this list
// from a Dictionary, so enumeration order is not guaranteed stable between calls.
driver.ReverseHostOrder();
// Poke the actor into re-publishing without changing anything material.
driver.RaiseHostStatusChanged();
ExpectNoMsg(TimeSpan.FromMilliseconds(500));
publisher.Published.Count.ShouldBe(before, "an order flip with no state change must be deduped, not re-published");
}
/// <summary>
/// <b>Leak guard.</b> The <see cref="IDriver"/> can OUTLIVE the actor — the host respawns a child
/// around the same driver object — so a missing <c>-=</c> in <c>PostStop</c> accumulates one handler
/// per respawn, each holding a dead <c>Self</c>.
/// </summary>
[Fact]
public void Stopping_the_actor_detaches_the_host_status_handler()
{
var driver = new ProbeStubDriver();
var parent = CreateTestProbe();
parent.IgnoreMessages(_ => true);
var actor = parent.ChildActorOf(DriverInstanceActor.Props(driver));
actor.Tell(new DriverInstanceActor.InitializeRequested("{}"));
AwaitAssert(() => driver.SubscriberCount.ShouldBe(1), TimeSpan.FromSeconds(3));
Watch(actor);
actor.Tell(PoisonPill.Instance);
ExpectTerminated(actor, TimeSpan.FromSeconds(3));
driver.SubscriberCount.ShouldBe(0);
}
/// <summary>
/// A driver with no probe publishes null host statuses — NOT an empty list. The two mean different
/// things at the UI ("no per-host detail available" vs "a probe that currently knows no hosts") and
/// collapsing them would render a driver with no probe as one whose devices are all fine.
/// </summary>
[Fact]
public void A_driver_without_a_probe_publishes_null_host_statuses()
{
var publisher = new RecordingHealthPublisher();
var actor = SpawnDriverActor(new StubDriver(), publisher);
actor.Tell(new DriverInstanceActor.InitializeRequested("{}"));
AwaitAssert(() => publisher.Published.Count.ShouldBeGreaterThan(0), TimeSpan.FromSeconds(3));
publisher.Published.ShouldAllBe(p => p.HostStatuses == null);
}
/// <summary>
/// A probe that throws must not take the health publish down with it. <c>GetHostStatuses()</c> is
/// documented as a pure in-memory snapshot (which is why the capability analyzer exempts it from
/// the guarded-call rule), but a driver is free to violate that, and losing the whole health channel
/// for one misbehaving probe would be a much worse outcome than losing the host detail.
/// </summary>
[Fact]
public void A_throwing_probe_degrades_to_null_without_killing_the_health_publish()
{
var driver = new ProbeStubDriver { ThrowOnGetHostStatuses = true };
var publisher = new RecordingHealthPublisher();
var actor = SpawnDriverActor(driver, publisher);
actor.Tell(new DriverInstanceActor.InitializeRequested("{}"));
AwaitAssert(() => publisher.Published.Count.ShouldBeGreaterThan(0), TimeSpan.FromSeconds(3));
publisher.Published[^1].Health.State.ShouldBe(DriverState.Healthy);
publisher.Published[^1].HostStatuses.ShouldBeNull();
}
private IActorRef SpawnDriverActor(IDriver driver, IDriverHealthPublisher publisher)
{
var parent = CreateTestProbe();
parent.IgnoreMessages(_ => true);
return parent.ChildActorOf(DriverInstanceActor.Props(driver, healthPublisher: publisher));
}
/// <summary>Captures every health publish so a test can assert on the host-status field.</summary>
private sealed record HealthPublish(
DriverHealth Health,
IReadOnlyList<HostConnectivityStatus>? HostStatuses);
private sealed class RecordingHealthPublisher : IDriverHealthPublisher
{
private readonly List<HealthPublish> _published = [];
/// <summary>Thread-safe snapshot — <c>Publish</c> runs on the actor thread while the test asserts
/// from its own.</summary>
public IReadOnlyList<HealthPublish> Published
{
get { lock (_published) return _published.ToArray(); }
}
/// <inheritdoc />
public void Publish(
string clusterId,
string driverInstanceId,
DriverHealth health,
int errorCount5Min,
DateTime? rediscoveryNeededUtc = null,
string? rediscoveryReason = null,
IReadOnlyList<HostConnectivityStatus>? hostStatuses = null)
{
lock (_published) _published.Add(new HealthPublish(health, hostStatuses));
}
}
/// <summary>
/// A stub driver exposing <see cref="IHostConnectivityProbe"/>.
/// <para><b>Its own health is deliberately STABLE</b> (a fixed last-read timestamp), for the same
/// reason <c>RediscoverableStubDriver</c> is: the shared <c>StubDriver</c> returns
/// <c>DateTime.UtcNow</c> from <c>GetHealth()</c>, so its fingerprint differs on every call, the
/// dedup never engages, and any test built on it passes vacuously whether or not the fix is
/// present.</para>
/// </summary>
private sealed class ProbeStubDriver : IDriver, IHostConnectivityProbe
{
private static readonly DateTime FixedLastRead = new(2026, 7, 30, 12, 0, 0, DateTimeKind.Utc);
private static readonly DateTime FixedChangedAt = new(2026, 7, 30, 11, 0, 0, DateTimeKind.Utc);
private readonly List<HostConnectivityStatus> _hosts = [];
/// <summary>When set, <see cref="GetHostStatuses"/> throws — the misbehaving-probe case.</summary>
public bool ThrowOnGetHostStatuses { get; init; }
/// <inheritdoc />
public event EventHandler<HostStatusChangedEventArgs>? OnHostStatusChanged;
/// <inheritdoc />
public string DriverInstanceId => "probe-stub-1";
/// <inheritdoc />
public string DriverType => "Stub";
/// <summary>Number of live subscribers on <see cref="OnHostStatusChanged"/>.</summary>
public int SubscriberCount => OnHostStatusChanged?.GetInvocationList().Length ?? 0;
/// <summary>Adds or updates a host, optionally raising the transition event as a real probe loop does.</summary>
public void SetHost(string hostName, HostState state, bool raise = false)
{
lock (_hosts)
{
var index = _hosts.FindIndex(h => h.HostName == hostName);
var entry = new HostConnectivityStatus(hostName, state, FixedChangedAt);
if (index >= 0) _hosts[index] = entry;
else _hosts.Add(entry);
}
if (raise) RaiseHostStatusChanged();
}
/// <summary>Flips enumeration order without changing any host's state.</summary>
public void ReverseHostOrder()
{
lock (_hosts) _hosts.Reverse();
}
/// <summary>Raises the event exactly as a real probe loop does.</summary>
public void RaiseHostStatusChanged()
=> OnHostStatusChanged?.Invoke(this, new HostStatusChangedEventArgs("plc-b", HostState.Running, HostState.Stopped));
/// <inheritdoc />
public IReadOnlyList<HostConnectivityStatus> GetHostStatuses()
{
if (ThrowOnGetHostStatuses) throw new InvalidOperationException("probe is broken");
lock (_hosts) return _hosts.ToArray();
}
/// <inheritdoc />
public Task InitializeAsync(string driverConfigJson, CancellationToken cancellationToken) => Task.CompletedTask;
/// <inheritdoc />
public Task ReinitializeAsync(string driverConfigJson, CancellationToken cancellationToken) => Task.CompletedTask;
/// <inheritdoc />
public Task ShutdownAsync(CancellationToken cancellationToken) => Task.CompletedTask;
/// <inheritdoc />
public DriverHealth GetHealth() => new(DriverState.Healthy, FixedLastRead, null);
/// <inheritdoc />
public long GetMemoryFootprint() => 0;
/// <inheritdoc />
public Task FlushOptionalCachesAsync(CancellationToken cancellationToken) => Task.CompletedTask;
}
}
@@ -185,7 +185,8 @@ public sealed class DriverInstanceActorRediscoverySignalTests : RuntimeActorTest
DriverHealth health,
int errorCount5Min,
DateTime? rediscoveryNeededUtc = null,
string? rediscoveryReason = null)
string? rediscoveryReason = null,
IReadOnlyList<HostConnectivityStatus>? hostStatuses = null)
{
lock (_published)
{
@@ -122,7 +122,8 @@ public sealed class DriverInstanceActorSubscriptionReconcileTests : RuntimeActor
DriverHealth health,
int errorCount5Min,
DateTime? rediscoveryNeededUtc = null,
string? rediscoveryReason = null)
string? rediscoveryReason = null,
IReadOnlyList<HostConnectivityStatus>? hostStatuses = null)
=> Interlocked.Increment(ref _count);
}
}