Files
lmxopcua/src/Server/ZB.MOM.WW.OtOpcUa.ControlPlane/Communication/CentralCommunicationActor.cs
T
Joseph Doherty d5b5cb6ede feat(mesh): CentralCommunicationActor — dark-switched fan-out, DB-sourced contacts
Phase 2 Task 3. Central-side end of the boundary: routes every central->node
command out over DPS or ClusterClient per MeshTransport:Mode, and forwards
inbound ApplyAcks to the deploy coordinator singleton (Forward, not Tell, so
the coordinator sees the node rather than this relay).

Three things here are load-bearing and easy to get wrong later:

SendToAll, never Send. Today's DPS publish reaches EVERY DriverHostActor and
the node side has no ClusterId or node filter to compensate -- scoping happens
later, inside the artifact. ClusterClient.Send delivers to exactly ONE
registered actor, so it would deploy to a single node while every other node
silently kept its old config, and the deployment could still seal green.
Sabotage-verified: swapping to Send reddens exactly that one test.

Exactly ONE ClusterClient, fleet-wide. A receptionist serves its whole cluster,
so SendToAll reaches every registered node-comm actor in the mesh regardless of
which contact point was dialled. One client per application Cluster -- the
shape Phase 6 wants -- would fan each command out once per cluster while the
fleet is still one mesh: N x duplicate DispatchDeployment and ApplyAck. Marked
TODO(Phase 6).

The contact set does NOT scope delivery. Excluding a maintenance-mode node from
the contacts does not stop it receiving commands; it still receives them, as it
does under today's broadcast. The filter is about which receptionists are worth
dialling, not about who gets the message. Stated in the code because the
opposite is the natural assumption.

Also carries the sister project's shipped fixes: per-row address parsing so one
malformed ClusterNode cannot abort the refresh and leave the contact set
half-built (regression test orders the bad row FIRST), the cache entry cleared
before recreate so a failed create cannot leave commands routing into a
stopping actor, and a Status.Failure handler so a DB outage is a Warning rather
than a silent debug-level unhandled message.

Two test-harness bugs found and fixed while writing this, both mine: SubscribeAck
goes to the SENDER (so the mediator subscribe needed an explicit sender), and
EventFilter matches case-INSENSITIVELY, so a "was NOT" filter also caught this
actor's own "was not created" warning.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-22 12:10:05 -04:00

309 lines
14 KiB
C#
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
using System.Collections.Immutable;
using Akka.Actor;
using Akka.Cluster.Tools.Client;
using Akka.Cluster.Tools.PublishSubscribe;
using Akka.Event;
using Microsoft.EntityFrameworkCore;
using ZB.MOM.WW.OtOpcUa.Cluster;
using ZB.MOM.WW.OtOpcUa.Commons.Messages.Deploy;
using ZB.MOM.WW.OtOpcUa.Commons.Messages.Mesh;
using ZB.MOM.WW.OtOpcUa.Configuration;
namespace ZB.MOM.WW.OtOpcUa.ControlPlane.Communication;
/// <summary>
/// Central-side end of the central↔node command boundary (per-cluster mesh Phase 2). Routes
/// every central→node command out — over DistributedPubSub or ClusterClient, per
/// <see cref="MeshTransportOptions.Mode"/> — and forwards inbound <see cref="ApplyAck"/>s to the
/// deploy coordinator singleton.
/// </summary>
/// <remarks>
/// <para>
/// <b>Registered per node, NOT as a cluster singleton.</b> A driver node's ClusterClient
/// rotates across contact points and must find a live comm actor at whichever central node
/// answers. A singleton would be reachable only through the one node hosting it, and the
/// rotation would fail silently against the other.
/// </para>
/// <para>
/// <b>Exactly one ClusterClient, fleet-wide.</b> See <see cref="RebuildClient"/> — this is
/// the single most important thing to understand before changing this actor.
/// </para>
/// </remarks>
public sealed class CentralCommunicationActor : ReceiveActor, IWithTimers
{
private const string RefreshTimerKey = "mesh-contact-refresh";
private readonly IDbContextFactory<OtOpcUaConfigDbContext> _dbFactory;
private readonly MeshTransportOptions _options;
private readonly IMeshClusterClientFactory _clientFactory;
private readonly Func<IActorRef?> _coordinator;
private readonly ILoggingAdapter _log = Context.GetLogger();
private readonly bool _useClusterClient;
private IActorRef? _client;
private ImmutableHashSet<string> _currentContacts = ImmutableHashSet<string>.Empty;
private CancellationTokenSource _lifecycle = new();
/// <summary>Gets the timer scheduler driving the periodic contact refresh.</summary>
public ITimerScheduler Timers { get; set; } = null!;
/// <summary>Creates the props for this actor.</summary>
/// <param name="dbFactory">Factory for the config database holding <c>ClusterNode</c> rows.</param>
/// <param name="options">The bound mesh-transport options.</param>
/// <param name="clientFactory">Creates the ClusterClient; substituted in tests.</param>
/// <param name="coordinator">
/// Resolves the deploy-coordinator singleton proxy. A <see cref="Func{TResult}"/> rather than
/// an <see cref="IActorRef"/> so this actor never depends on registration order at wiring
/// time — it resolves on first use and caches.
/// </param>
/// <returns>The props.</returns>
public static Props Props(
IDbContextFactory<OtOpcUaConfigDbContext> dbFactory,
MeshTransportOptions options,
IMeshClusterClientFactory clientFactory,
Func<IActorRef?> coordinator) =>
Akka.Actor.Props.Create(() =>
new CentralCommunicationActor(dbFactory, options, clientFactory, coordinator));
/// <summary>Initializes a new instance of the <see cref="CentralCommunicationActor"/> class.</summary>
/// <param name="dbFactory">Factory for the config database.</param>
/// <param name="options">The bound mesh-transport options.</param>
/// <param name="clientFactory">Creates the ClusterClient.</param>
/// <param name="coordinator">Lazily resolves the deploy-coordinator singleton proxy.</param>
public CentralCommunicationActor(
IDbContextFactory<OtOpcUaConfigDbContext> dbFactory,
MeshTransportOptions options,
IMeshClusterClientFactory clientFactory,
Func<IActorRef?> coordinator)
{
_dbFactory = dbFactory;
_options = options;
_clientFactory = clientFactory;
_coordinator = coordinator;
_useClusterClient = string.Equals(
options.Mode, MeshTransportOptions.ModeClusterClient, StringComparison.OrdinalIgnoreCase);
Receive<MeshCommand>(HandleMeshCommand);
Receive<ApplyAck>(HandleApplyAck);
Receive<RefreshContacts>(_ => LoadContactsFromDb());
Receive<ContactsLoaded>(HandleContactsLoaded);
// A faulted LoadContactsFromDb task is piped here as a Status.Failure. Without this handler
// the failure is an unhandled message (debug level only) and the refresh fails silently —
// an operator cannot distinguish "no nodes configured" from "the database is down", and the
// contact set silently freezes at whatever it last held.
Receive<Status.Failure>(f => _log.Warning(
f.Cause,
"Failed to load ClusterNode contact points; the mesh ClusterClient contact set was NOT "
+ "refreshed and may be stale or empty. Commands may be dropped until the next refresh"));
Receive<SubscribeAck>(_ => { /* DPS subscribe confirmation */ });
}
/// <inheritdoc />
protected override void PreStart()
{
if (!_useClusterClient)
{
_log.Info(
"Mesh transport is {Mode}: commands are published on DistributedPubSub and no "
+ "ClusterClient is created", _options.Mode);
return;
}
Timers.StartPeriodicTimer(
RefreshTimerKey,
new RefreshContacts(),
TimeSpan.Zero,
TimeSpan.FromSeconds(_options.ContactRefreshSeconds));
}
/// <inheritdoc />
protected override void PostStop()
{
_lifecycle.Cancel();
_lifecycle.Dispose();
}
private void HandleMeshCommand(MeshCommand cmd)
{
if (!_useClusterClient)
{
DistributedPubSub.Get(Context.System).Mediator.Tell(new Publish(cmd.Topic, cmd.Message));
return;
}
if (_client is null)
{
// App-level drop, matching the sister project's decision and this repo's buffer-size = 0.
// The command is NOT queued and will NOT be retried: a deploy fails at the coordinator's
// apply deadline naming the silent nodes, which is a better operator signal than a
// command arriving minutes late against a config the DB already recorded as failed.
_log.Warning(
"No mesh ClusterClient — dropping {MessageType}. The last contact refresh produced "
+ "no usable contact point from the enabled, non-maintenance ClusterNode rows; the "
+ "command is not buffered and will not be retried",
cmd.Message.GetType().Name);
return;
}
// SendToAll, NEVER Send. Today's DPS publish reaches EVERY DriverHostActor — there is no
// ClusterId or node filter on the node side at all; scoping happens later, inside the
// artifact (DeploymentArtifact.ResolveClusterScope). ClusterClient.Send delivers to exactly
// ONE registered actor, which would deploy to a single node of the fleet while every other
// node silently kept its old configuration — and the deployment would still seal green if
// the ack set happened to be satisfied.
_client.Tell(new ClusterClient.SendToAll(MeshPaths.NodeCommunication, cmd.Message));
}
private void HandleApplyAck(ApplyAck ack)
{
var coordinator = _coordinator();
if (coordinator is null)
{
_log.Warning(
"Received ApplyAck for {DeploymentId} from {NodeId} but the deploy coordinator is "
+ "not resolvable on this node; the ack is lost and the deployment will time out "
+ "naming a node that applied successfully",
ack.DeploymentId, ack.NodeId);
return;
}
// Forward, not Tell: preserves the original sender across the hop so the coordinator sees
// the node rather than this relay.
coordinator.Forward(ack);
}
private void LoadContactsFromDb()
{
var self = Self;
CancellationToken ct;
try
{
ct = _lifecycle.Token;
}
catch (ObjectDisposedException)
{
return; // Stopping.
}
var systemName = Context.System.Name;
var dbFactory = _dbFactory;
// Off the actor thread, result piped back as a message. A synchronous DB read here would
// block every command dispatch behind a slow or unreachable SQL Server.
Task.Run(async () =>
{
await using var db = await dbFactory.CreateDbContextAsync(ct).ConfigureAwait(false);
var rows = await db.ClusterNodes
.AsNoTracking()
.Where(n => n.Enabled && !n.MaintenanceMode)
.Select(n => new { n.NodeId, n.Host, n.AkkaPort })
.ToListAsync(ct)
.ConfigureAwait(false);
var contacts = new List<string>(rows.Count);
var malformed = new List<string>();
foreach (var row in rows)
{
var address = $"akka.tcp://{systemName}@{row.Host}:{row.AkkaPort}/system/receptionist";
// Parse up front, per row, inside the loop's own guard. A single malformed row must
// not abort the whole refresh and leave the contact set half-built — the sister
// project shipped that bug and its regression test deliberately orders the bad row
// first.
if (ActorPath.TryParse(address, out _)) contacts.Add(address);
else malformed.Add($"{row.NodeId} -> {address}");
}
return new ContactsLoaded(contacts, malformed);
}, ct).PipeTo(self);
}
private void HandleContactsLoaded(ContactsLoaded msg)
{
foreach (var bad in msg.Malformed)
{
_log.Warning(
"ClusterNode {Row} does not yield a usable receptionist address; skipping it in this "
+ "refresh (other nodes are unaffected). Check the row's Host and AkkaPort", bad);
}
var contacts = msg.Contacts.ToImmutableHashSet();
if (contacts.IsEmpty)
{
_log.Warning(
"No usable ClusterNode contact points; the mesh ClusterClient was not created and "
+ "every central→node command will be dropped until a refresh finds one");
return;
}
if (_client is not null && _currentContacts.SetEquals(contacts)) return;
RebuildClient(contacts);
}
/// <summary>Stops any existing client and creates a replacement for the new contact set.</summary>
/// <remarks>
/// <para>
/// <b>TODO(Phase 6): one client per application <c>Cluster</c>.</b> Today this actor
/// creates exactly <b>one</b> client whose contacts are the whole fleet, and that is not
/// a compromise — it is the only correct shape while the fleet is a single Akka mesh.
/// </para>
/// <para>
/// A <c>ClusterClientReceptionist</c> serves its entire cluster, so
/// <c>SendToAll</c> reaches every registered node-comm actor in the mesh <i>regardless of
/// which node's address was used as the contact point</i>. One client per Cluster would
/// therefore fan each command out once per cluster — N× duplicate
/// <c>DispatchDeployment</c> and N× duplicate <c>ApplyAck</c>. Phase 6 splits the meshes,
/// at which point per-cluster clients become both correct and necessary.
/// </para>
/// <para>
/// Corollary worth stating because it is easy to assume otherwise: <b>the contact set
/// does not scope delivery.</b> Excluding a maintenance-mode node from the contacts does
/// not stop it receiving commands — it still receives them, exactly as it does under
/// today's DPS broadcast. The filter is about which receptionists are worth <i>dialling</i>
/// (a node in maintenance may be switched off), not about who gets the message.
/// </para>
/// </remarks>
/// <param name="contacts">The receptionist addresses to dial.</param>
private void RebuildClient(ImmutableHashSet<string> contacts)
{
if (_client is not null)
{
_log.Info("Mesh contact set changed; replacing the ClusterClient");
Context.Stop(_client);
// Cleared now: if the create below throws, a stale ref would route commands into a
// stopping actor and they would vanish without a warning.
_client = null;
_currentContacts = ImmutableHashSet<string>.Empty;
}
try
{
var paths = contacts.Select(ActorPath.Parse).ToImmutableHashSet();
_client = _clientFactory.Create(Context.System, paths);
_currentContacts = contacts;
_log.Info("Mesh ClusterClient created with {Count} contact point(s)", contacts.Count);
}
catch (Exception ex)
{
_log.Error(ex,
"Failed to create the mesh ClusterClient; central→node commands are dropped until "
+ "the next contact refresh succeeds");
}
}
/// <summary>Internal tick asking for a contact-point refresh.</summary>
public sealed record RefreshContacts;
/// <summary>Result of a contact-point load, piped back from the DB read.</summary>
/// <param name="Contacts">Receptionist addresses that parsed cleanly.</param>
/// <param name="Malformed">Rows that did not, described for the log.</param>
public sealed record ContactsLoaded(
IReadOnlyList<string> Contacts,
IReadOnlyList<string> Malformed);
}