2ee84af1c0
Phase 0 of the ClusterClient→gRPC migration
(docs/plans/2026-07-22-clusterclient-to-grpc-plan.md). Standalone hardening: it
closes a gap that exists today and is a precondition for moving command/control
onto gRPC in later phases.
T0.1 — delete the ManagementActor ClusterClientReceptionist registration.
It was built for an out-of-cluster CLI that was never written: the shipped CLI
speaks HTTP Basic to /management, which asks the actor in-process through
ManagementActorHolder. Nothing in the repo ever sent to /user/management. The
actor still runs there; only the cross-boundary advertisement is gone. Six
documents claimed the CLI used ClusterClient — including the CLI's own README
"Architecture Notes" — and are corrected here rather than left to rot.
T0.2 — record, do not port, the dead integration-routing path.
IntegrationCallRequest is unwired at BOTH ends: RouteIntegrationCallAsync has
zero callers anywhere, and RegisterLocalHandler(Integration, …) appears only in
a test, so production always answers "Integration handler not available". It is
excluded from the gRPC contract (28 of 29 commands migrate) rather than
enshrined on an additive-only wire format, and deleting it during a
transport migration would mix a behavioural change into a change whose whole
value is that behaviour is identical. See
docs/known-issues/2026-07-22-integration-call-routing-is-dead-code.md.
T0.3 — preshared-key authentication on SiteStreamService.
The service shipped with no auth at all: plaintext h2c, no interceptor, so
anything that could reach a site node's :8083 could open a live data stream or
read audit rows back via PullAuditEvents/PullSiteCalls. ControlPlaneAuthInterceptor
now gates /sitestream.SiteStreamService/ — modeled on LocalDbSyncAuthInterceptor
(constant-time compare, fail-closed, PermissionDenied) but gating a SET of
service prefixes so phases 1A/1B add services rather than interceptors. LocalDb
sync keeps its own separate key: it authenticates the pair partner, not central,
and collapsing the two would make a site's central-facing key also admit writes
into its database.
Keys are per site (SB-GRPC-PSK-<siteId>), never fleet-wide, so a compromised
site yields only its own. Central attaches them through ControlPlaneCredentials,
which binds CallCredentials to the channel — covering unary and streaming
uniformly, and letting the key resolve asynchronously, which a client
interceptor could not do without blocking. All three central→site channel
creation sites go through it (SiteStreamGrpcClient and both audit pull invokers);
the pull invokers' channel caches are re-keyed by (site, endpoint) because
credentials are per-site and bound to the channel.
Two decisions beyond the plan:
* StartupValidator now requires GrpcPsk on Site nodes. The plan specified only
the runtime gate, but fail-closed with no boot check produces a node that
joins, answers heartbeats and reports healthy while refusing every stream,
audit pull and telemetry ingest — silent and total. Same reasoning as the
existing inbound API-key pepper rule.
* Added Communication:SitePsks as a central-side key map. The plan assumed
central would read the store, seeded via a dev KEK; the docker rig
deliberately boots with no master key, so store-only resolution would leave
it unable to dial its own sites. The store stays primary — it is the only
source that can serve a site added at runtime — with the map covering
key-less hosts and one-off pins. Neither source falling back to
"unauthenticated" is the invariant.
T0.4 — dev keys on both rigs and tests.
34 tests. The seven that matter most exercise a real in-process gRPC stack over
TestServer: the unit tests on either side of the wire would both stay green if
the halves disagreed, and gRPC refuses call credentials on a plaintext channel
by default — the UnsafeUseInsecureChannelCallCredentials opt-in is only provable
by making a real call. They confirm correct key passes on unary AND streaming,
wrong key and no-credentials both get PermissionDenied, and an unresolvable key
fails the call with nothing reaching the service.
OPERATIONAL: a site node upgraded to this build without a key will not boot.
That includes the gitignored deploy/wonder-app-vd03/ overlay.
130 lines
4.4 KiB
C#
130 lines
4.4 KiB
C#
using System.Collections.Concurrent;
|
|
using Microsoft.Extensions.Logging.Abstractions;
|
|
using ZB.MOM.WW.ScadaBridge.Communication.Grpc;
|
|
|
|
namespace ZB.MOM.WW.ScadaBridge.Communication.Tests.Grpc;
|
|
|
|
/// <summary>
|
|
/// Regression tests for Communication-007 — the factory's synchronous
|
|
/// <see cref="SiteStreamGrpcClientFactory.Dispose"/> must not block on the
|
|
/// async disposal path (sync-over-async). It must dispose each client through
|
|
/// the client's synchronous <see cref="SiteStreamGrpcClient.Dispose"/>.
|
|
/// </summary>
|
|
public class SiteStreamGrpcClientFactoryDisposeTests
|
|
{
|
|
/// <summary>
|
|
/// Test client that records whether it was disposed via the sync or async path.
|
|
/// </summary>
|
|
private sealed class TrackingClient : SiteStreamGrpcClient
|
|
{
|
|
public bool SyncDisposeCalled { get; private set; }
|
|
public bool AsyncDisposeCalled { get; private set; }
|
|
|
|
public override void Dispose() => SyncDisposeCalled = true;
|
|
|
|
public override ValueTask DisposeAsync()
|
|
{
|
|
AsyncDisposeCalled = true;
|
|
return ValueTask.CompletedTask;
|
|
}
|
|
}
|
|
|
|
/// <summary>
|
|
/// Test factory that hands out <see cref="TrackingClient"/> instances while
|
|
/// still exercising the base factory's real caching and disposal machinery.
|
|
/// </summary>
|
|
private sealed class TrackingFactory : SiteStreamGrpcClientFactory
|
|
{
|
|
private readonly ConcurrentBag<TrackingClient> _created = new();
|
|
|
|
public TrackingFactory() : base(NullLoggerFactory.Instance) { }
|
|
|
|
public IReadOnlyCollection<TrackingClient> Created => _created.ToList();
|
|
|
|
protected override SiteStreamGrpcClient CreateClient(string siteIdentifier, string grpcEndpoint)
|
|
{
|
|
var client = new TrackingClient();
|
|
_created.Add(client);
|
|
return client;
|
|
}
|
|
}
|
|
|
|
[Fact]
|
|
public void Dispose_DisposesClientsSynchronously_NotViaAsyncPath()
|
|
{
|
|
var factory = new TrackingFactory();
|
|
factory.GetOrCreate("site-a", "http://localhost:5100");
|
|
factory.GetOrCreate("site-b", "http://localhost:5200");
|
|
|
|
factory.Dispose();
|
|
|
|
Assert.NotEmpty(factory.Created);
|
|
Assert.All(factory.Created, c =>
|
|
{
|
|
Assert.True(c.SyncDisposeCalled, "client should be disposed via synchronous Dispose()");
|
|
Assert.False(c.AsyncDisposeCalled, "synchronous Dispose() must not route through DisposeAsync()");
|
|
});
|
|
}
|
|
|
|
[Fact]
|
|
public void Dispose_DoesNotDeadlock_UnderSingleThreadedSynchronizationContext()
|
|
{
|
|
// A strict single-threaded SynchronizationContext: continuations posted to
|
|
// it are only pumped by the worker loop. Sync-over-async (blocking the only
|
|
// thread on an async continuation that needs that same thread) deadlocks here.
|
|
using var ctx = new SingleThreadSyncContext();
|
|
Exception? captured = null;
|
|
var done = new ManualResetEventSlim();
|
|
|
|
ctx.Post(_ =>
|
|
{
|
|
try
|
|
{
|
|
var factory = new SiteStreamGrpcClientFactory(NullLoggerFactory.Instance);
|
|
factory.GetOrCreate("site-a", "http://localhost:5100");
|
|
factory.Dispose();
|
|
}
|
|
catch (Exception ex)
|
|
{
|
|
captured = ex;
|
|
}
|
|
finally
|
|
{
|
|
done.Set();
|
|
}
|
|
}, null);
|
|
|
|
Assert.True(done.Wait(TimeSpan.FromSeconds(5)),
|
|
"factory.Dispose() did not complete — likely a sync-over-async deadlock");
|
|
Assert.Null(captured);
|
|
}
|
|
|
|
/// <summary>Minimal single-threaded synchronization context for the deadlock test.</summary>
|
|
private sealed class SingleThreadSyncContext : SynchronizationContext, IDisposable
|
|
{
|
|
private readonly BlockingCollection<(SendOrPostCallback cb, object? state)> _queue = new();
|
|
private readonly Thread _thread;
|
|
|
|
public SingleThreadSyncContext()
|
|
{
|
|
_thread = new Thread(Run) { IsBackground = true };
|
|
_thread.Start();
|
|
}
|
|
|
|
private void Run()
|
|
{
|
|
SetSynchronizationContext(this);
|
|
foreach (var (cb, state) in _queue.GetConsumingEnumerable())
|
|
cb(state);
|
|
}
|
|
|
|
public override void Post(SendOrPostCallback d, object? state) => _queue.Add((d, state));
|
|
|
|
public void Dispose()
|
|
{
|
|
_queue.CompleteAdding();
|
|
_thread.Join(TimeSpan.FromSeconds(2));
|
|
}
|
|
}
|
|
}
|