fix(health): active tier returns 503 on a role-member-but-not-leader node (0.2.1)

The shared spec states the active tier "Fails (503) on a standby or
role-member-but-not-leader node", and the same spec maps Degraded to HTTP 200.
ActiveNodeDecision.Evaluate returned Degraded for exactly that case, so the tier
answered 200 on every node and could not distinguish an active node from a
standby over HTTP at all - only the response body differed.

The consequence was live, not theoretical: OtOpcUa's Traefik routes the admin UI
by this probe (docker-dev/traefik-dynamic.yml, scripts/install/traefik-dynamic.yml),
so both central nodes always passed and the leader-pinning the design called for
has silently never worked. Found while running the overview dashboard's live
acceptance, which reads the same tier and showed all six rig nodes as Active.
Filed as lmxopcua#494.

Role-member-but-not-leader now returns Unhealthy. The other two branches are
unchanged: a node that lacks the role stays Healthy (the probe is irrelevant to
it), and the startup-safety path still returns Degraded when the ActorSystem or
cluster is not yet available.

Blast radius is one consumer. OtOpcUa is the only user of the role-filtered mode;
ScadaBridge deliberately uses its own OldestNodeActiveHealthCheck (leadership is
address-ordered and diverges from the singleton host), and MxGateway and
HistorianGateway have no Akka cluster.

Version 0.2.1, three packages published to the Gitea feed and restore-verified.
70 tests pass.
This commit is contained in:
Joseph Doherty
2026-07-24 10:46:04 -04:00
parent c22b06deba
commit af1e0a86b2
3 changed files with 21 additions and 9 deletions
+1 -1
View File
@@ -5,7 +5,7 @@
<Nullable>enable</Nullable>
<ImplicitUsings>enable</ImplicitUsings>
<LangVersion>latest</LangVersion>
<Version>0.2.0</Version>
<Version>0.2.1</Version>
<ManagePackageVersionsCentrally>true</ManagePackageVersionsCentrally>
<!-- Emit XML docs so the public API summaries ship inside the packed nupkgs (IntelliSense for
consumers). CS1591 (missing doc on a public member) is suppressed so undocumented test /
@@ -33,8 +33,17 @@ internal static class ActiveNodeDecision
/// <returns>
/// Role-less: Healthy iff the node is Up and the cluster leader, otherwise Unhealthy.
/// Role-filtered: Healthy when the node lacks the role (probe irrelevant) or carries the role and
/// is the role-singleton leader; Degraded when it carries the role but is not the leader.
/// is the role-singleton leader; <strong>Unhealthy</strong> when it carries the role but is not
/// the leader.
/// </returns>
/// <remarks>
/// The role-member-but-not-leader case returns Unhealthy, not Degraded, because the active tier's
/// whole purpose is to be a 503 that an orchestrator can act on — the shared spec states it as
/// "Fails (503) on a standby or role-member-but-not-leader node", and the same spec maps Degraded
/// to 200. Returning Degraded made the tier answer 200 on every node, so a load balancer pointed
/// at it (OtOpcUa's Traefik routes the admin UI this way) kept the standby in the pool and its
/// leader-pinning silently never worked. Found on a live rig; see lmxopcua#494.
/// </remarks>
public static HealthStatus Evaluate(bool selfUp, bool isLeader, bool hasRole, string? requiredRole)
{
if (requiredRole is null)
@@ -43,7 +52,7 @@ internal static class ActiveNodeDecision
if (!hasRole)
return HealthStatus.Healthy;
return isLeader ? HealthStatus.Healthy : HealthStatus.Degraded;
return isLeader ? HealthStatus.Healthy : HealthStatus.Unhealthy;
}
}
@@ -79,8 +88,9 @@ public sealed class ActiveNodeHealthCheck : IHealthCheck
/// <summary>
/// Role-filtered constructor: Healthy when the node lacks <paramref name="role"/> or carries it
/// and is the role-singleton leader; Degraded when it carries the role but is not the leader
/// (OtOpcUa AdminRoleLeader pattern). Degraded when the ActorSystem / cluster is not yet ready.
/// and is the role-singleton leader; Unhealthy (503) when it carries the role but is not the
/// leader (OtOpcUa AdminRoleLeader pattern). Degraded when the ActorSystem / cluster is not yet
/// ready.
/// </summary>
/// <param name="serviceProvider">
/// The application service provider. The <see cref="ActorSystem"/> is resolved lazily so the
@@ -32,7 +32,7 @@ public sealed class ActiveNodeDecisionTests
// Role-filtered: requiredRole != null.
// lacks role -> Healthy (probe irrelevant for this node)
// has role & is leader -> Healthy (selfUp is ignored — role-filtered mode only cares about leadership)
// has role & not leader -> Degraded
// has role & not leader -> Unhealthy (503) — the tier exists to be a 503 an LB can act on
public static IEnumerable<object[]> RoleFilteredCases() => new[]
{
// node lacks the role -> Healthy regardless of selfUp / isLeader
@@ -43,9 +43,11 @@ public sealed class ActiveNodeDecisionTests
new object[] { true, true, true, "admin", HealthStatus.Healthy },
// node carries the role and is leader -> Healthy (selfUp=false: role-filtered mode ignores selfUp)
new object[] { false, true, true, "admin", HealthStatus.Healthy },
// node carries the role but is not leader -> Degraded
new object[] { true, false, true, "admin", HealthStatus.Degraded },
new object[] { false, false, true, "admin", HealthStatus.Degraded },
// Node carries the role but is not leader -> Unhealthy, i.e. HTTP 503.
// Degraded would be a 200 (see the shared health SPEC), which made the active tier answer
// 200 on every node and silently broke OtOpcUa's Traefik leader-pinning. lmxopcua#494.
new object[] { true, false, true, "admin", HealthStatus.Unhealthy },
new object[] { false, false, true, "admin", HealthStatus.Unhealthy },
};
[Theory]