2a2e54f0833793d262f9a864d5e52e8418acc3ca
2649 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2a2e54f083 |
feat(mesh): route deploy, driver-control and alarm commands through the mesh router
Phase 2 Task 7. The five central->node publish sites -- ConfigPublishCoordinator's deploy dispatch, and AdminOperationsActor's restart, reconnect, acknowledge and shelve -- now hand a MeshCommand to the transport router instead of telling the DistributedPubSub mediator directly. Neither actor knows which transport is in force any more. _meshRouter is NULLABLE with a mediator fallback, and that is deliberate: about a dozen existing coordinator and admin-operations tests construct these actors without a router, and they must keep exercising the real DPS path rather than a silently-disabled one. The fallback is not a permanent seam -- Phase 6 deletes it with the DPS branch and the single mesh. Wiring note: the comm actor's WithActors block moved ABOVE the singleton registrations so both singletons can resolve it with registry.Get, which is the ordering-by-registration idiom AdminOperationsActor already uses for the coordinator. An earlier version used ActorSelection.ResolveOne().GetAwaiter() .GetResult() inside the props factory -- that blocks inside actor construction and races the very spawn it waits for. Sabotage-verified by inverting the branch so the mediator always wins: all six new tests go red, including the fallback test (which proves it is asserting the mediator path rather than passing by accident). Suites after: ControlPlane 118/118, Runtime 450/450, Host.IntegrationTests 196/202 -- sole failure AbCip_Green_AgainstSim, the fixture baseline that fails on master too. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
16d598560a |
feat(mesh): register comm actors with the receptionist; pin the wire-contract paths
Phase 2 Task 6. Both comm actors are now spawned and receptionist-registered: CentralCommunicationActor from WithOtOpcUaControlPlaneSingletons, and NodeCommunicationActor from WithOtOpcUaRuntimeActors (before DriverHostActor, because it is the host's ackRouter under ClusterClient mode; it has no dependency of its own on the host, since inbound commands travel via the node's EventStream). Both are plain WithActors registrations, NOT singletons -- a driver node's ClusterClient rotates across contact points and must find a live comm actor at whichever node answers. Both register with the receptionist in BOTH transport modes: an idle registration costs nothing, whereas registering only under ClusterClient would mean flipping the flag on a running fleet needs a restart before anything can be reached, turning a config change back into a deployment. The coordinator ref is resolved lazily per message (Func<IActorRef?> + registry.TryGet) so the comm actor never depends on the order in which Akka.Hosting materialises the singleton proxy relative to this block. Adds MeshCommActorPathTests, which boots the real two-node host and Identify- probes both paths. These strings are the wire contract and nothing else asserts they agree: a rename would compile, pass every unit test, deploy, and deliver nothing, because a ClusterClient send to an unregistered path is dropped silently. Sabotage-verified with an actual rename. DELIBERATELY NOT TESTED, with the reason recorded in the file: that the actors are per-node rather than singletons. Two discriminators were tried and both were invalid -- "the path resolves on both nodes" passes under either shape because a ClusterSingletonManager is also created on every node in the role, and "no /singleton child" is null even for the KNOWN singleton /user/config-publish, so Akka.Hosting does not lay them out that way. A sabotage re-registering the actor via WithSingleton left both green, which is what exposed them. Rather than ship an assertion that cannot fail, the property is left to the Task 8 boundary test, where a send only arrives if the target really is registered. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
95e4f97529 |
chore(mesh): mark Phase 2 tasks 3-5 complete
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
5cb0e72166 |
feat(mesh): node subscribers accept EventStream commands; _coordinatorOverride -> _ackRouter
Phase 2 Task 5. DriverHostActor and ScriptedAlarmHostActor now subscribe to the node-local EventStream alongside their existing DPS topic subscriptions, so the transport flag lives ENTIRELY on the publishing side -- exactly one of the two ever publishes, and the switch can be flipped or reverted without touching a single subscriber. Note the alarm case is not symmetric: the DPS subscribe there still carries the node-LOCAL publisher (OtOpcUaServerHostedService's OPC UA Part 9 method calls), which never crosses the boundary and stays on DPS in both modes. Only the central AdminUI leg moves. Renamed DriverHostActor._coordinatorOverride to _ackRouter. Under ClusterClient mode that ref is the NodeCommunicationActor, emphatically not a coordinator, and the old name would send the next reader looking for one. Adds two wiring tests, because the unit tests either side of this seam BOTH pass with the subscription deleted -- the comm actor's test asserts a probe receives the message, and the host tests drive the actor by direct Tell. Neither notices a missing PreStart subscribe. Sabotage-verified: deleting either subscription reddens exactly its own wiring test and nothing else. Three things went wrong while writing those tests, all worth keeping: - The first version raced. ActorOf returns before PreStart runs, and an EventStream.Publish with no subscriber is dropped silently -- no buffering, no dead letter. Fixed with a warm-up whose ack proves PreStart completed; the comment says why it is not ceremony. - The alarm version was VACUOUS: it told the host directly INSIDE the assertion window alongside the relay, so the direct Tell alone produced the log the assertion waited for. Split into warm-up (outside) and relay (inside). - The alarm test needs its own ActorSystem: the shared harness pins loglevel = WARNING and the only signal an unowned alarm gives is a Debug line. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
915beec11a |
feat(mesh): NodeCommunicationActor — inbound commands to the node-local EventStream
Phase 2 Task 4. Node-side end of the boundary, registered with the receptionist per node rather than as a singleton so central's contact rotation reaches whichever node answers. Inbound commands are re-emitted on the node's EventStream, NOT back onto their DistributedPubSub topic. DPS is mesh-wide and central SendToAll-s to every node's comm actor, so a DPS republish would deliver N copies of every command to every subscriber. The EventStream is node-local by construction. It also avoids a registration handshake with actors spawned later: ScriptedAlarmHostActor is a CHILD of DriverHostActor and has no registry key, so there is no ref to hand in at wiring time. Each inbound type is listed explicitly rather than caught by a ReceiveAny, so an unknown type dead-letters loudly instead of being republished blind onto a stream where every subscriber ignores it. Outbound is ApplyAck only, and it uses Send rather than SendToAll -- the deploy coordinator is one singleton behind central's proxy, so SendToAll would deliver one ack per central node and the coordinator would count this node twice. Notably there is NO Ask across this boundary in Phase 2: every migrated command is fire-and-forget, which is why this actor is far smaller than the sister project's equivalent and needs no sender preservation at all. Sabotage-verified twice: deleting the AlarmCommand handler and flipping Send to SendToAll each redden exactly their own test. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
d5b5cb6ede |
feat(mesh): CentralCommunicationActor — dark-switched fan-out, DB-sourced contacts
Phase 2 Task 3. Central-side end of the boundary: routes every central->node command out over DPS or ClusterClient per MeshTransport:Mode, and forwards inbound ApplyAcks to the deploy coordinator singleton (Forward, not Tell, so the coordinator sees the node rather than this relay). Three things here are load-bearing and easy to get wrong later: SendToAll, never Send. Today's DPS publish reaches EVERY DriverHostActor and the node side has no ClusterId or node filter to compensate -- scoping happens later, inside the artifact. ClusterClient.Send delivers to exactly ONE registered actor, so it would deploy to a single node while every other node silently kept its old config, and the deployment could still seal green. Sabotage-verified: swapping to Send reddens exactly that one test. Exactly ONE ClusterClient, fleet-wide. A receptionist serves its whole cluster, so SendToAll reaches every registered node-comm actor in the mesh regardless of which contact point was dialled. One client per application Cluster -- the shape Phase 6 wants -- would fan each command out once per cluster while the fleet is still one mesh: N x duplicate DispatchDeployment and ApplyAck. Marked TODO(Phase 6). The contact set does NOT scope delivery. Excluding a maintenance-mode node from the contacts does not stop it receiving commands; it still receives them, as it does under today's broadcast. The filter is about which receptionists are worth dialling, not about who gets the message. Stated in the code because the opposite is the natural assumption. Also carries the sister project's shipped fixes: per-row address parsing so one malformed ClusterNode cannot abort the refresh and leave the contact set half-built (regression test orders the bad row FIRST), the cache entry cleared before recreate so a failed create cannot leave commands routing into a stopping actor, and a Status.Failure handler so a DB outage is a Warning rather than a silent debug-level unhandled message. Two test-harness bugs found and fixed while writing this, both mine: SubscribeAck goes to the SENDER (so the mediator subscribe needed an explicit sender), and EventFilter matches case-INSENSITIVELY, so a "was NOT" filter also caught this actor's own "was not created" warning. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
e762ae00f2 |
chore(mesh): mark Phase 2 tasks 0-2 complete
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
0d6669d5ff |
feat(mesh): MeshCommand envelope + generation-suffixed ClusterClient factory
Phase 2 Task 2. Two small pieces the comm actors are built on. MeshCommand carries a command plus the DPS topic it would have been published on, so the admin singletons stop knowing which transport is in force. Without it the dark switch would have to be threaded through five publish sites across two actors, each free to drift. Central-side only -- it never crosses the wire, so it carries no serialization contract. DefaultMeshClusterClientFactory appends a generation counter to the actor name. Context.Stop is asynchronous and the name stays reserved until termination completes, so recreating the same name inside one message handler throws InvalidActorNameException -- which is exactly what a contact-set change does (stop the old client, create the replacement, one handler). Ported from the sister project, where it was a shipped bug fix rather than foresight. Sabotage-verified: freezing the name to a constant reddens both name tests with InvalidActorNameException. Dropped a third test I had written -- it asserted ClusterClientSettings directly and never touched the factory, so it was coverage theatre. Replaced with a comment saying where contact propagation IS covered (the Task 8 boundary test, the only place it can actually fail). Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
7b71a6a35d |
feat(mesh): receptionist extension, zero ClusterClient buffering, frame-size logging
Phase 2 Task 1. Three changes to the embedded base config, plus a test that reads every one of them back off a RUNNING ActorSystem rather than off the HOCON text -- the settled idiom here (SplitBrainResolverActivationTests), because several fragments compete and the losing one still reads correctly in the file. - Registers the ClusterClientReceptionist extension. It is not auto-started, so without this the boundary only listens if some code path happens to call ClusterClientReceptionist.Get(system) first: "is the boundary up?" would depend on startup ordering rather than on configuration. - buffer-size = 0. This is the one genuine behaviour change and it is NOT what the sister project runs -- they never wrote this section and inherit 1000. Buffering would replay a DispatchDeployment on reconnect and apply a deployment the coordinator already sealed as TimedOut, leaving the node on a configuration the database says failed. - log-frame-size-exceeding = 32000b. An oversized frame is dropped WITHOUT tearing down the association: heartbeats keep flowing, the node reports healthy, and the only symptom is a command that never arrived. Our deploy notify is payload-free, so this canary should stay quiet. reconnect-timeout and receptionist.role restate Akka defaults; they are stated explicitly and pinned because the behaviour is load-bearing, and the comments say which is which. One test was VACUOUS on first write: Config.GetInt returns 0 for a missing key, so the buffering assertion passed before the feature existed -- it could not distinguish "explicitly 0" from "absent, default 1000 in force". Rewritten to assert through ClusterClientSettings/ClusterReceptionistSettings, the objects the client is actually built from. Sabotage-verified after the fix: setting buffer-size back to 1000 reddens exactly that one test. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
5439f14804 |
feat(mesh): MeshTransportOptions — dark switch between DPS and ClusterClient
Phase 2 Task 0. Binds a validated `MeshTransport` section selecting which transport carries central->node commands, defaulting to Dps so nothing changes until the rig gate passes. The validator exists because every fault it catches otherwise surfaces as an ABSENCE -- a deployment that never arrives, an alarm ack that does nothing -- which has no stack trace and no failing node to point at. Two of the shapes it rejects are ported from the sister project's shipped mistakes: a contact point carrying an actor-path suffix, and (documented in the options XML) a template listing the node's OWN remoting port as a central contact, which is a permanent failure in the initial-contact rotation. Contact points are required only under ClusterClient mode; requiring them under the default would make the section mandatory on admin-only nodes that have no reason to carry them. Sabotage-verified: relaxing the empty-contacts guard and the actor-path-suffix guard turns exactly those two tests red, and nothing else. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
1ec831883c |
docs(mesh): Phase 2 plan — comm actors + ClusterClient transport
Phase 2 of the per-cluster mesh program: move the three central->node command channels (deployments, driver-control, alarm-commands) and the deployment-acks reply channel off DistributedPubSub onto Akka ClusterClient, behind a config flag. Recon of both repos turned up three corrections to the design doc's account of the sister project, all folded into the plan: - Design doc S2 claims "every unhandled message replies with a typed failure". ScadaBridge has no ReceiveAny/Unhandled override in either comm actor; the typed-failure idiom fires only for a MISSING REGISTRATION, per message type. - "No central buffering" was never a setting they wrote. They have zero akka.cluster.client HOCON and run the defaults (buffer-size 1000, reconnect-timeout off). We choose buffer-size = 0 deliberately. - Publish -> Send is the wrong substitution. Today's deploy notify is a broadcast every DriverHostActor receives (no ClusterId filter exists on the node side), so it must become SendToAll. Send would deploy to exactly one node of the fleet and seal green on partial acks. Also records the single-mesh duplicate-delivery trap (one ClusterClient in Phase 2, not one per Cluster -- a receptionist serves its whole mesh) and one deviation from the program plan's exit gate: there is no cross-boundary Ask in Phase 2, because every migrated command is fire-and-forget with a local reply, so "an Ask timing out cleanly" cannot be run as written. Marks the program tracking tables current: prereq + Phase 1 are DONE (they still read "not executed"), Phase 2 in progress. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
7654f24dab |
test(harness): serialize Host.IntegrationTests; drop the reconciler's unused singleton proxy
Closes the last branch-only failures. Measured across five full-suite runs on this branch versus two on master in a clean worktree: master failed only AbCip in this assembly (2/2), while the branch additionally failed one or two E2E deploy tests, rotating between DeployHappyPath, DriverReconnect and EquipmentNamespaceMaterialization (4/4). Each passed in isolation, and the assembly on its own was 195/201 throughout — the failures only appeared under full-suite CPU contention. Cause is contention, not correctness. Every test here builds a real two-node Akka cluster — two Kestrel hosts, two ActorSystems with remoting, six admin singletons, an EF context, a deploy pipeline — and xUnit ran them concurrently with each other and with every other assembly. That was survivable while the coordinator sealed instantly on an empty expected-ack set; now that it waits for a real ack from every configured node, these tests measure an end-to-end round-trip and starve. Serialising this one assembly makes them deterministic: 195/201 with only AbCip_Green_AgainstSim, and the full-solution failure set is now IDENTICAL to master's — 7 shared, zero branch-only. Cost is wall-clock for this assembly, 40s -> 3m35s; it is a handful of heavyweight E2E tests, not a broad unit suite, and a widened timeout alone did not fix it (tried first, at 45s). Also drops createProxyToo for the address reconciler. Nothing resolves ClusterNodeAddressReconcilerKey from the registry — the actor only reads cluster state and logs — so the proxy was a second actor per node hunting for a singleton nobody addresses. Kept because it removes dead machinery, not because it fixed the flakiness; measured on its own, it did not. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
0b7c53f64f |
test(harness): widen the deploy-seal timeout to match a real two-node deploy
Comparing full-suite runs on this branch against master in a clean worktree: both fail 11 tests, 9 identical. Master's two extras are load-flaky unit tests (AbCip Probe_loops, Galaxy EventPumpBoundedChannel); this branch's two extras were both Host.IntegrationTests deploy-path tests — the area this change re-times — so they are attributable here rather than to background flakiness. Cause is margin, not correctness. These waits used to observe a coordinator that sealed instantly on an empty expected-ack set; they now observe a real ApplyAck round-trip from every configured node. 15s was enormous margin against "instant" and thin against the real thing, so they failed only under full-suite CPU contention. Hoisted to TwoNodeClusterHarness.DeploySealTimeout (45s) so the reason is recorded once rather than as four unexplained numbers. DriverReconnectE2eTests is deliberately left alone: it seeds both ClusterNode rows itself and unconditionally, so its expected-ack set is identical before and after this change. Widening its timeout would be papering over a flake this change did not cause. Host.IntegrationTests 195/201; sole failure AbCip_Green_AgainstSim, verified failing on master. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
14746f2995 |
test(harness): seed ClusterNode rows in in-memory mode; retarget the failover deploy test
The full-suite run surfaced two load-dependent failures that were green in isolation. Root cause: SeedDefaultClusterAsync no-opped unless OTOPCUA_HARNESS_USE_SQL=1, on the reasoning that "the in-memory provider ignores FK constraints, so deploy E2E tests pass without it". Phase 1 invalidated that. With no ClusterNode rows the coordinator's expected-ack set is empty, so it seals IMMEDIATELY. "Wait for Sealed" therefore stopped implying "every node has applied" — and because DriverHostActor.UpsertNodeDeploymentState writes each node's row on its own schedule, every test counting those rows after a seal became a race. Reproducibly green alone, red under suite load, and it surfaced in a different test each run, which is what made it look like flakiness rather than a contract change. Seeding in both modes restores the invariant those tests were written against and is closer to production either way: a real fleet always has these rows, because the FK requires them. Deployment_started_with_node_b_down_seals_with_one_node_state asserted the behaviour Phase 1 deliberately removed — its own comment documented membership snapshotting, i.e. sealing green while a configured node never received the deployment. Replaced by two tests covering the new contract: a stopped node is still expected (both state rows exist, B still Applying, deployment does not seal), and MaintenanceMode is what makes it seal with one. The new "does not seal" test asserts both rows exist rather than only the absence of a seal, so it cannot pass against a coordinator that simply died. Host.IntegrationTests 195/201, sole remaining failure AbCip_Green_AgainstSim — verified failing on master in a clean worktree, so pre-existing fixture baseline, not a regression. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
e27b7f43f5 |
chore(mesh): mark Phase 1 tasks complete
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
ee69caa270 |
feat(config): add ClusterNode.MaintenanceMode — the hatch the live gate proved missing
Phase 1 gate step 4 failed: setting Enabled = 0 to take a node out of service returns 422 ClusterEnabledNodeCountMismatch, because DraftValidator.ValidateClusterTopology requires the enabled-node count to equal ServerCluster.NodeCount. On a Warm/Hot pair — every cluster on the rig, and every cluster in the target topology — Enabled can therefore never be the maintenance hatch. The plan called step 4 "the check that makes step 3 acceptable to ship", so Task 3's behaviour change was not shippable as it stood: a node down for maintenance would block every deployment to its cluster. The two rules were each reasonable and contradictory together — the validator reads Enabled as "part of the declared topology", Phase 1 additionally read it as "expect an ack". Split the meanings rather than weaken either rule: Enabled part of the declared topology (validator, untouched) MaintenanceMode expected to participate now (coordinator + reconciler) Rejected alternatives: counting configured rather than enabled nodes (drops the guard against booting a pair into InvalidTopology); downgrading the rule to a warning (weakens a deploy gate for everyone); shipping with no hatch. Sabotage: dropping !n.MaintenanceMode from the coordinator query turns the new test red. It also asserts both nodes remain Enabled, so the fix cannot quietly regress to disabling the row after all. Configuration.Tests 95/95, ControlPlane.Tests 101/101, solution builds clean. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
57e1d01766 |
docs(mesh): rig seed + Phase 1 documentation
- docker-dev seed: AkkaPort = 4053 on all six ClusterNode rows, so a freshly seeded rig matches a migrated one. GrpcPort left null. - config-db-schema.md: both columns, why AkkaPort is NOT NULL/4053 and GrpcPort is nullable, the unenforced duplication + its reconciler, and Enabled's new second meaning as the deploy path's expected-ack set. - Configuration.md: Cluster:Port / PublicHostname now flag that they are stored twice, with the "update the row too" instruction and why the drift is silent. - design doc §7: Phase 1 marked done, plus a "Phase 1 as shipped" note recording both deviations rather than leaving the sketch reading as what happened. - program plan: Phase 1 marked done; AdminUI node edit explicitly deferred. - CLAUDE.md: the deploy-path behaviour change and its three consequences. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
90c1302f71 |
feat(fleet): reconcile ClusterNode dial targets against live membership
ClusterNode.AkkaPort and the node's own Cluster:Port are the same fact in two places and nothing made them agree. Phase 2 dials the row instead of gossiping, so a node binding 4054 while its row says 4053 becomes unreachable from central — and the symptom is a silent absence of acks rather than an error. That is the shape of the Modbus/ModbusTcp and TwinCat/Focas drifts already in this repo. Since Phase 1 also made the rows the deploy path's expected-ack set, drift already costs a failed deployment today. Implemented as the plan's preferred option: an admin-role singleton comparing rows against the membership an admin node can already see, rather than each driver node asserting its own row — Phase 4 removes the driver nodes' ConfigDb connection, so a self-assertion written there would have to be deleted again. Three shapes, split by severity: a row whose dial target disagrees with its own NodeId and a running node with no row are Errors; an enabled row with no matching member is a Warning, because a node down for maintenance is a legitimate state. Findings are logged only when the set changes — a check that reprints the same warning every sweep trains operators to filter it out. Documented limitation: Phase 2 must revisit this. Once the fleet splits into one mesh per cluster an admin node cannot see site members, and every site row would report EnabledRowNotInCluster forever. Sabotage: removing the change-detection guard turns the repeat test red. ControlPlane.Tests 100/100. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
d88e245503 |
feat(deploy): coordinator sources its expected-ack set from ClusterNode rows
Cuts ConfigPublishCoordinator's last genuinely mesh-bound dependency: central must be able to name the nodes a deployment is for without sharing a gossip ring with them, because Phase 2 splits the fleet into one mesh per Cluster. Behaviour change, confirmed before implementing: a configured node that is switched off is now expected, so deploying while it is down fails at the apply deadline instead of sealing green without it. Under the membership rule the operator was told the fleet was deployed when it was not. ClusterNode.Enabled = 0 is the maintenance hatch, and the deadline log now names the silent nodes and points at that hatch — "4/5 acks landed" would leave an operator reading logs on five machines. Two assumptions are now documented on the class rather than left implicit: every ClusterNode row is a driver node (the DB has no role column and deliberately does not gain one — it would drift from Cluster:Roles), and ServerCluster.Enabled is not consulted because nothing else consults it. Dropped from the plan: cluster-scope filtering of the expected set. There is no cluster-scoped deployment to filter on — Deployment has no ClusterId, ConfigComposer always snapshots the whole DB, and ResolveClusterScope is node-side self-scoping of a fleet-wide artifact. Positive control: reverting DiscoverDriverNodes to the membership scan turns three of the four new tests red. The fourth pins the seal-empty branch and is documented as derivation-insensitive rather than left to look like coverage. ControlPlane.Tests 90/90. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
da74ebd696 |
feat(config): migration AddClusterNodeTransportPorts
Additive only — two AddColumn ops, no AlterColumn, so no accumulated model drift is riding along with this change: ALTER TABLE [ClusterNode] ADD [AkkaPort] int NOT NULL DEFAULT 4053; ALTER TABLE [ClusterNode] ADD [GrpcPort] int NULL; Adds a third test covering the DB-side default specifically. The entity round-trip cannot see it: ClusterNode.AkkaPort's CLR initializer is also 4053, so that assertion passes with the mapping default deleted. The new test inserts through raw SQL with the column omitted, so only the schema can supply the value — verified by setting defaultValue: 0, which turns exactly that one test red and leaves the other two green. Configuration.Tests 95/95. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
2bbb02713c |
feat(config): add ClusterNode.AkkaPort/GrpcPort — central's dial targets
Per-cluster mesh Phase 1 groundwork. Once the meshes split (Phase 2) central can no longer see a site node's appsettings, so the transport ports it must dial have to live in a row central can read. AkkaPort is non-nullable with a 4053 default — every node listens on a remoting port, so 0 is never a truthful value and pre-existing rows must migrate to something real. GrpcPort is nullable with no default: nothing listens on it until Phase 5, and a non-null default would assert a port that does not exist. Both are documented as central's dial targets rather than the node's own binding config; the duplication against Cluster:Port is reconciled in Task 4. The new SchemaCompliance test is red until the migration lands (Task 2). Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
e7311f11a6 | docs(mesh): program tracking — Phase 1 plan already exists from the design session; surface its deploy-seal decision gate | ||
|
|
9af935f237 |
docs(cluster): self-first seed ordering replaces the self-form fallback
docs/Redundancy.md's bootstrap section is rewritten around the mechanism that is
actually in force: which of Akka's two bootstrap processes runs is decided by
whether seed-nodes[0] is this node's own address, so the ORDER of Cluster:SeedNodes
is the fix, not a timer. Adds the per-process table, the shipped self-first
example, why the validator is conditional (site nodes are not seeds) and why it
matches PublicHostname rather than the 0.0.0.0 bind address.
The retirement is documented rather than erased: a new subsection explains that a
watchdog outside the join handshake cannot tell "no seed answered" from "a seed
answered and the join is in flight", and the
|
||
|
|
3f24d4d6bf |
feat(cluster)!: self-first seed ordering replaces the InitJoin self-form watchdog
Akka runs FirstSeedNodeProcess -- the only bootstrap path that can form a NEW
cluster when no peer answers InitJoin -- exclusively when seed-nodes[0] is the
node's own address. Every other node runs JoinSeedNodeProcess and retries
InitJoin forever. That is why "both peers are seeds" never meant "either can
cold-start alone", and it is fixed here the way Akka itself intends: each seed
node lists ITSELF first.
- docker-dev central-2 now lists itself as SeedNodes__0 (central-1 was already
self-first). The site nodes are untouched -- they seed off central-1 only and
are deliberately not seeds.
- AkkaClusterOptionsValidator enforces the invariant at boot via the shared
ZB.MOM.WW.Configuration AddValidatedOptions/ValidateOnStart seam. The rule is
CONDITIONAL -- self must be seed[0] only IF self is in the list at all -- or it
would refuse to boot every driver-only site node. Identity is compared on
PublicHostname (falling back to Hostname when blank) AND port, i.e. the address
Akka puts in SelfAddress: matching the 0.0.0.0 bind address would find no seed
anywhere and leave the rule silently inert on the whole docker-dev rig.
- ClusterBootstrapFallback + Cluster:SelfFormAfter are DELETED. The watchdog sat
outside Akka's join handshake, so it could not distinguish "no seed answered"
from "a seed answered and the join is in flight" -- and Cluster.Join(SelfAddress)
is not ignored mid-handshake, it wins. The live gate (
|
||
|
|
a78425ea8f |
docs(plans): record the Task 8 live finding — self-form races the join handshake
Corrects the plan's behavior table (the mid-join-handshake row was FALSE) and notes the mesh-program Phase 6 convergence on ScadaBridge's self-first seed ordering, which retires the watchdog + guard entirely. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
ea45ace1c3 |
fix(cluster): don't self-form while a seed peer is reachable — the fallback islanded a restarting node
The live gate for the manual-failover control caught the self-form fallback forming a SECOND cluster. A node bounced by a failover restarted, received InitJoinAck from its live peer — the join was in flight and healthy — but did not get the Welcome inside the 10s window, because the peer's ring still held the node's previous incarnation (Exiting -> Down -> Removed). The fallback fired on the timer and the node islanded itself until an operator restarted it. The original design assumed Cluster.Join(SelfAddress) would be ignored mid-handshake. It is not — it wins. And since manual failover deliberately produces that restart, every failover could island the node it bounced. Before self-forming, TCP-probe the other seed addresses; a reachable peer means wait another window instead. That is also what the fallback claims to detect: 'no seed answered InitJoin (peer down at boot)'. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
d8a85c3d89 |
feat(adminui): manual failover control on the cluster redundancy page
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
69b697bc3e |
feat(redundancy): manual failover service — graceful Leave of the driver Primary
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
983310edbb |
docs(redundancy): record the self-form fallback live gate as PASSED on docker-dev
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
b853185cfd |
docs(redundancy): SelfFormAfter closes the seed-node bootstrap gap
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
b8ac7e35e6 |
test(cluster): self-form fallback — lone-seed forms, disabled waits, site-node guard
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
9ded86342e |
feat(cluster): InitJoin self-form fallback armed via Akka.Hosting startup task
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
979dc102dd |
feat(cluster): SelfFormAfter option (bootstrap self-form fallback window)
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
8e195263f6 |
docs(claude): record the two redundancy corrections in the project instructions
Both change how someone should reason about this area, and the second changes a public message contract, so they belong in CLAUDE.md rather than only in docs/Redundancy.md. Includes the rule the work established: assert the effective cluster config off a running ActorSystem, never the typed options or the akka.conf text, because several HOCON fragments compete and the losing one still reads correctly in the file. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
0de1352e55 |
fix(deps): pin System.Security.Cryptography.Xml 10.0.10 — fresh restores were red
Four NU1903 high-severity advisories (GHSA-23rf-6693-g89p, GHSA-8q5v-6pqq-x66h, GHSA-cvvh-rhrc-wg4q, GHSA-g8r8-53c2-pm3f) landed in the NuGet audit data against System.Security.Cryptography.Xml 10.0.7, which Microsoft.AspNetCore.DataProtection 10.0.7 pulls in transitively for key storage. Under TreatWarningsAsErrors a fresh restore produces 204 NU1903 errors and the solution does not build at all. This was invisible locally: machines with cached audit data keep building, so the break only shows on a clean clone or a docker image build. It surfaced here by accident — a detached worktree at the pre-change baseline, created to check whether a flaky test predated today's work, could not restore. Worth noting that the diagnostic path was the accident, not the intent: nothing in the normal build or test loop would have reported this before someone else hit it. Same surgical-pin pattern as the SQLitePCLRaw entry directly above it: a direct PackageReference at the one project where the chain enters the repo (Core.Configuration), rather than CentralPackageTransitivePinningEnabled, which this repo already documents as breaking the Roslyn version split. Bumping the DataProtection parent to 10.0.10 instead was rejected for the reason the sister repo recorded when it hit the same advisories: 10.0.10 floors Microsoft.Extensions.* and, via the EFCore adapter, Microsoft.EntityFrameworkCore at 10.0.10, forcing a family-wide servicing bump and an NU1605 downgrade cascade that deserves its own reviewed change. Verified the way the defect demanded — in a fresh worktree, not this one: restore clean (0 errors, down from 204) and full solution build clean. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
84c6524edc |
docs(mesh): record phases 0a/0b as done and plan phase 1
The design doc still described the downing gap and the RoleLeader derivation as open. Both shipped earlier today, so 6.2 and 4 now carry the fix and its pin instead of the diagnosis, and 7 marks 0a/0b done. 0a's live gate stays open and is explicitly deferred to phase 7: the failure only appears in a 1-vs-1 split and docker-dev is a single six-node mesh, so there is no rig on which to drill it until the meshes are actually split. The phase 1 plan lands with one finding that changes its shape. The design pairs "ClusterNode gains address columns" with "coordinator sources its expected-ack set from the DB" as though they were one change; they are independent, and the second is far smaller than it reads. ClusterNode.NodeId is ALREADY in the coordinator's identifier space — the rig seeds central-1:4053 and so on, because NodeDeploymentState.NodeId is FK-bound to it and the coordinator's membership-derived ids already satisfy that FK. That the deploy path works today is standing proof the two sets agree. What is not small is the semantics. Membership-derived and DB-derived sets differ on a configured-but-stopped node: today the deployment seals green without it — telling the operator the fleet is deployed when it is not — and afterwards it would fail at the apply deadline. Failing is the better behaviour and is required once central cannot see cluster membership, but it needs an escape hatch or a node down for maintenance blocks every deploy. ClusterNode.Enabled already is that hatch, so the plan filters on it and gates the phase on a live check that proves both halves: a stopped node fails the deploy legibly, and disabling it lets the deploy seal. The plan also adds a task the design does not have: AkkaPort duplicates the node's own Cluster:Port with nothing making them agree, and an unenforced duplicated address is the same shape as the Modbus/ModbusTcp and TwinCat/Focas drifts already recorded here. Phases 2-7 remain unstarted. They change the running data path and rewrite the rig, and each needs its own plan and live gate. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
b2d19a1730 |
fix(tests): repair the DriverTypeNames guard, and stop MSBuild hoarding worker nodes
Two independent housekeeping defects. DriverTypeNamesGuardTests (3 failures, pre-existing). The guard discovers each driver's *DriverFactoryExtensions.Register by reflection and invokes it, then asserts the registered type-name set matches the DriverTypeNames constants. It built the argument array positionally — registry first, null for everything after — on the assumption that the remaining parameters were optional. Galaxy's are not: its Register takes a REQUIRED ISecretResolver and guards it with ArgumentNullException.ThrowIfNull, so the reflective call threw and all three parity facts failed. OpcUaClient's equivalent parameter is optional, which is why only Galaxy tripped it. Worth stating plainly: the guard was not partially broken, it was completely inert. The throw happened while building the registered set, so NO driver's parity was ever checked — a genuine drift in any of the nine would have been invisible behind the same three red tests. Arguments are now supplied by parameter type, and a parameter the helper cannot satisfy fails with a message naming the factory and the parameter instead of an opaque ArgumentNullException. The stand-in resolver reports every secret absent, so if the "factory func is never invoked" assumption is ever broken, the driver fails closed. Core.Abstractions.Tests: 138 passed, 0 failed. MSBuild node reuse. MSBuild leaves worker nodes resident between builds so the next one starts warm; across this session's repeated solution builds and test runs that reached 55 idle processes holding ~6 GB. That is not just untidy — the integration suites assert on timeouts, so memory pressure produces failures indistinguishable from real regressions, the same failure mode the Workstation-GC fix addressed from a different direction. Directory.Build.rsp now passes -nodeReuse:false. Measured on the same incremental solution build: 15 residual nodes without it, 2 with, and elapsed time 6.12s vs 6.11s — no cost worth the memory on this machine. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
50b55ee4b9 |
fix(redundancy): elect the Primary by cluster age, not by address
RedundancyStateActor derived the driver Primary from ClusterState.RoleLeader("driver").
Akka offers two different notions of a "first" member and they are not the same
one: role leader is the lowest-ADDRESSED Up member (host, then port), while
ClusterSingletonManager places singletons on the OLDEST — lowest up-number.
They agree on a freshly-formed cluster, which is why every existing test passed.
They diverge after any restart: the restarted node re-joins as the youngest while
keeping its address, so if it holds the lower address it becomes role leader while
the singletons stay put. The snapshot would then name a Primary that is not
hosting the work, and every Primary-gated surface follows it — inbound device
writes, native-alarm acks, the fleet-wide alerts emit, and the alarm-history
drain would all enable on the wrong node while the node actually running the
singletons stayed gated off.
BuildSnapshot now selects the oldest Up member carrying the driver role, matching
singleton placement. Leaving members are excluded: a node handing its singletons
over must not be named Primary.
NodeRedundancyState.IsRoleLeaderForDriver is renamed IsDriverPrimary, and
NodeHealthInputs.IsDriverRoleLeader likewise. Keeping the old names would have
left the wire contract asserting a derivation the code no longer uses — the same
drift that made this defect invisible.
Proven by a real two-node cluster rather than a mock. RedundancyPrimaryElectionTests
binds the first-joining node to the HIGHER port, so oldest and lowest-address name
different nodes, and includes a fixture assertion that the divergence actually
occurred — without it the real assertion could pass for the wrong reason. Positive
control: restoring the RoleLeader derivation turns exactly the two election tests
red while the fixture check stays green.
Runtime.Tests 440 passed, ControlPlane.Tests 82, Cluster.Tests 36.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
2964361a6f |
fix(cluster): auto-down downing strategy — a two-node pair could not survive an oldest-node crash
Ports the sister project's live-proven fix (ScadaBridge cf3bd52f). OtOpcUa ran the identical configuration it indicts: keep-oldest with down-if-alone = on. Akka.NET 1.5.62's KeepOldest.OldestDecision only lets the down-if-alone branch rescue a side holding >= 2 members. A two-node cluster losing a peer is always a 1-vs-1 split, so the branch never fires, the lone survivor falls through to DownReachable and downs ITSELF, and run-coordinated-shutdown-when-down then terminates it. down-if-alone is a 3+-node feature and does not do what its name suggests for a pair: a crash of the oldest node is a total outage, which is the exact failure the redundancy pair exists to absorb. Cluster:SplitBrainResolverStrategy now selects the provider, defaulting to auto-down: Akka's AutoDowning with auto-down-unreachable-after = 15s, so the leader among the reachable members downs the unreachable peer and a crash of either node fails over in place. keep-oldest remains available for deployments that would rather take an outage than ever run dual-active during a real partition. An unrecognised value fails the host at startup rather than falling through to the fatal default. The tests assert the EFFECTIVE configuration — they start a real host through WithOtOpcUaClusterBootstrap and read akka.cluster.downing-provider-class back off the running ActorSystem. This is not incidental. The first draft inlined the three calls the bootstrap makes instead of calling it, which pinned the test's own wiring: sabotaging production's HOCON precedence left all 36 green. The prior file had the same shape at a smaller scale, asserting only that BuildClusterOptions returned a KeepOldestOption — true, and true of a configuration that cannot fail over. Positive control: making BuildDowningHocon emit nothing for auto-down turns exactly the two effective-config guards red. Measured while verifying rather than assumed: HoconAddMode.Append also wins here, purely because it is added last, so the mode name is not the guarantee. The comment now says so instead of asserting a precedence rule that does not hold. Not yet live-drilled on OtOpcUa. The docker-dev rig is a single six-node mesh where the 1-vs-1 pathology cannot occur; the kill-the-oldest drill belongs with the per-cluster mesh work that makes every mesh exactly two nodes (design doc 6.2 / Phase 0a). Recorded as an outstanding gate in docs/Redundancy.md. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
c6fb8bb0ad |
docs(design): correct 6.2 — the keep-oldest gap is closed upstream and live here
I reported this gap as still open, citing a requirements doc line and the PR that made the failover drill honest about it. Both were stale: ScadaBridge closed it three hours earlier in cf3bd52f by switching to auto-down. My check missed it because I compared only one direction (main..origin/main) and never asked whether local main was AHEAD of origin, so two unpushed commits were invisible to me. The substance matters more than the correction. Their commit message records a live-proven finding: Akka.NET 1.5.62 KeepOldest.OldestDecision only lets down-if-alone rescue a side with >= 2 members, so in a 1-vs-1 split the survivor takes DownReachable and downs ITSELF. OtOpcUa runs precisely that configuration — akka.conf active-strategy keep-oldest with down-if-alone on, plus the matching typed KeepOldestOption — which means a two-node pair cannot fail over when the oldest node crashes. That is a total-outage defect inside the redundancy feature, independent of the mesh design, and it is now Phase 0a: most urgent, and needing a crash-the-oldest live gate because the failure only appears 1-vs-1. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
8838e92d0a |
fix(tests): StopNodeBAsync also has to stop the node, not just dispose it
Completes
|
||
|
|
835fc08c3f |
fix(tests): stop the integration host from ballooning to 30 GB
User-reported: the integration tests leave instances behind and take 18+ GB.
Measured before changing anything — a SINGLE Host.IntegrationTests process peaked
at 30,153 MB RSS across its ~200 tests.
The cause was not a leak and not parallelism. Referencing the Host drags in the
ASP.NET Core framework reference, which turns Server GC on by default: one heap
per core, tuned for throughput and deliberately reluctant to hand memory back.
Correct for a server, wrong for a test host. Workstation GC takes the same suite
from 30,153 MB to 6,717 MB — 78% less — with no change in runtime (45s).
Two hypotheses were measured and rejected on the way, recorded here so nobody
re-runs the experiment: capping xunit maxParallelThreads to 4 cost 29% wall-clock
and saved nothing (30,153 -> 31,115 MB, i.e. noise), and it turns out a single
process was doing all of it. Reverted.
The harness fix is real but was worth only 4% on its own: TwoNodeClusterHarness
disposed each node without stopping it first, and IHost.DisposeAsync only disposes
the service provider — it never invokes IHostedService.StopAsync, which is where
Akka.Hosting terminates the ActorSystem. So every test left two live ActorSystems,
remoting transports and dispatcher pools included, rooted for the process
lifetime. The class doc-comment asserted the opposite ("DisposeAsync, which runs
CoordinatedShutdown"); it was wrong, and StopNodeBAsync inherited the same bug on
a path failover tests depend on.
The payoff is larger than the memory number. Two failures I had classified as
known-flaky baseline noise now pass consistently:
RoslynVirtualTagEvaluatorTests.Evaluate_racing_ClearCompiledScripts_never_fails_with_disposed
and ContinuousHistorizationRecorderTests.Retry_after_writer_failure_eventually_acks.
They were never flaky — they were starved. Full solution: peak summed test-process
RSS 15.2 GB, zero new failures, two fewer.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
06b5d81bfb |
Merge feat/localdb-phase2: alarm store-and-forward on LocalDb + live-gate fixes
Phase 2 moves the alarm store-and-forward buffer off its standalone alarm-historian.db into the consolidated LocalDb as a replicated alarm_sf_events table, so a node that dies holding undelivered alarm history no longer takes it to the grave. SqliteStoreAndForwardSink becomes LocalDbStoreAndForwardSink; the breaking AlarmHistorian:DatabasePath key is removed, surviving only as the path AlarmSfLegacyMigrator reads to copy a pre-consolidation queue across on first boot. Because the buffer now replicates, exactly-once delivery could no longer rest on the enqueue-side gate alone — the Secondary holds a full copy — so the drain gate had to land in the same commit that registered the table, not as a later refinement. Delivery is at-least-once across a failover by design; row ids are a hash of the payload, so an event both nodes accept in a boot window converges to one row. The live gate on the docker-dev rig found four production defects, three of which crash-looped every driver node before its first check could run: an empty ServerHistorian:ApiKey and a UseTls/scheme mismatch both kill the host (the validator had classified them as degrading, but the gateway client validates its own options at construction), and plaintext h2c was unreachable entirely because the adapter forwarded TLS-only options unconditionally. The fourth was Phase 2's own: the drain deferred to a cluster-wide elected Primary while the queue is pair-local, so on a fleet rig every node — including unpaired ones — suspended its drain and nothing drained anywhere. Also carries the DoD sweep, which found eight live sites still naming the deleted sink including user-visible AdminUI text, and the design work that followed: a per-cluster Akka mesh design mirroring ScadaBridge, and a recorded limitation that the redundancy election is scoped per Akka cluster rather than per application Cluster. Full suite verified twice against a pre-branch baseline worktree: failure sets identical, zero regressions. The first run showed two extra deploy-test timeouts that proved to be resource starvation from a leaked test-host process, not a regression — they pass isolated and did not recur. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
4caaa11e9e |
docs(design): fix two staleness artefacts the settled decisions introduced
Section 5 still listed "site-local-primary storage" among the things deliberately NOT copied from ScadaBridge — which §6.1 now reverses, since driver nodes stop connecting to the ConfigDb and read their configuration from LocalDb. Replaced with an explicit copied/not-copied split, and narrowed the not-copied item to what it should have said all along: their transient-only central alarm cache, which OtOpcUa has no reason to adopt because it persists alarm history through the historian and Phase 2's replicated buffer already survives a failover. Also corrected a renumbering artefact — the risk about rewriting the docker-dev rig cited Phase 4, but the mesh partition moved to Phase 6 when the two data-path phases were inserted. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
fe12ed9d34 |
docs(design): settle the three open questions on the per-cluster mesh
(1) One central DB, driver nodes disconnected from it. Verified in ScadaBridge before adopting: AddConfigurationDatabase is called at Program.cs:264, inside the Central branch (89-478); the Site branch starts at 479, so site nodes never register the central database at all. There is still exactly one authoritative configuration database — it is central-only, and sites fetch from central and cache locally. The consequence worth naming: this promotes LocalDb from an outage cache to the driver node's steady-state configuration store. Boot-from-cache stops being the exceptional path and becomes the normal one. Phase 1's chunking, SHA-256 verification, retention and pair replication all carry over — but a cache defect that used to be a degraded-mode bug becomes a total-availability one, so the #485 unreadable-artifact class now has a larger blast radius. Audited the five driver-side ConfigDb consumers that must be re-homed. Two are more than mechanical: EfAlarmConditionStateStore holds pair-local Part 9 condition state and should follow the alarm S&F buffer into LocalDb, and DbHealthProbeActor feeds ServiceLevel — a driver node with no DB to probe changes what a client-visible value means. (2) The keep-oldest gap: corrected the status rather than the conclusion. What was resolved is the documentation and the drill, not the hole. Gitea PR #12 (closed) carried T1 badc97af, which made the drill "measure the registered keep-oldest outage instead of pretending recovery", and T2 d5364506, which dropped "the undeliverable ~25s active-crash promise". The gap itself is still the registered deferred keep-oldest topology decision on the 2026-07-08 master tracker, with no open issue. Accepted here with the same posture, recorded as a known risk. (3) Auth matches ScadaBridge for now — unauthenticated inter-cluster transports on a trusted network, recorded as an accepted risk with the fail-closed interceptor already in our tree noted as the template for whenever it is revisited. Sequencing grows from six phases to eight; the two that change a running system's data path rather than its wiring each get their own live gate. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
f862804a35 |
docs(design): per-cluster Akka mesh — separate 2-node mesh per application Cluster
Design for review, mirroring the ScadaBridge hub-and-spoke shape. Researched from
the ScadaBridge tree directly rather than from its CLAUDE.md, which understates
the boundary in five ways: there are THREE transports (ClusterClient, gRPC, plus
token-gated HTTP), the gRPC direction is inverted (central dials INTO each site),
active node is the OLDEST Up member and explicitly never the cluster leader, all
clusters share one ActorSystem name, and site nodes carry two roles.
Records a defect worth fixing regardless of the topology decision: OtOpcUa uses
two different node-selection rules at once. SBR is keep-oldest and singletons are
placed on the oldest member, but RedundancyStateActor elects Primary from
RoleLeader("driver") — lowest address. Those diverge permanently after a
restart-and-rejoin, so the node the SBR protects and hosts every singleton on need
not be the node the data-plane gates consider Primary. ScadaBridge forbids that
rule by name, for the reason it gives: both sides claim leadership during a
partition, which is the dual-primary shape archreview 03/S4 exists to prevent.
Three open questions are left explicitly undecided: whether to adopt their
autonomous-site data architecture or keep the shared ConfigDb; what to do about
the registered two-node keep-oldest total-outage gap they acknowledge but have not
closed; and whether to authenticate transports they left unauthenticated.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
ad385eb70f |
docs(redundancy): record the per-Akka-cluster role-election limitation
Corrects two claims made while writing the Phase 2 gate record. The docker-dev rig is NOT misconfigured: central-1/central-2 carry admin,driver deliberately — the compose says so — because they are themselves a pair that runs the MAIN cluster's drivers. And a redundant pair is not a missing concept. The application Cluster (ClusterId) already is one: it owns drivers, devices and ClusterNode rows, and Phase 1's deployment-artifact cache is keyed by it precisely so a pair shares one entry. So the root cause is wider than the alarm drain. RedundancyStateActor elects one Primary per AKKA cluster while every real unit of redundancy is an APPLICATION cluster, and ServiceLevel, the alerts emit gate and the inbound device-write gate all inherit that mis-scoping. On a fleet of pairs in one Akka cluster, exactly one driver node cluster-wide services Primary-gated work. Recommendation revised in the gate doc: fix the election scope (its own change, its own gate, because it touches the safety-critical write gate), and explicitly do NOT derive pair identity from LocalDb:Replication:PeerAddress — that would add a second parallel notion of 'my partner' that the real fix obsoletes, and the library has no initiator flag, so both halves would have to dial just to obtain a config value. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
f9f1b8fcee |
fix(localdb): phase-2 live gate — 4 production defects found and fixed
Gate record: docs/plans/2026-07-20-localdb-phase2-live-gate.md. Checks 1, 2, 5, 6 pass. Checks 3 and 4 are NOT satisfied — the defect they were meant to confirm turned out to be the opposite of what the plan assumed. Three defects crash-looped every driver node before check 1 could even run: 1. An empty ServerHistorian:ApiKey kills the host. ServerHistorianOptions- Validator exists to turn exactly that class of failure into a named OptionsValidationException, but its documented fail tier explicitly excluded ApiKey on the reasoning that a keyless client "degrades — the gateway rejects calls". It does not: the client validates its own options at construction, so the process dies during Akka startup and never makes a call. 2. UseTls disagreeing with the endpoint scheme kills the host too, in both directions (both messages confirmed in the shipped client assembly). Moving an endpoint from https to http without clearing UseTls is an ordinary migration slip. 3. Plaintext h2c was UNREACHABLE. HistorianGatewayClientAdapter forwarded the TLS-only options unconditionally, and AllowUntrustedServerCertificate defaults to false, so it always sent RequireCertificateValidation=true — which the client rejects outright when UseTls=false. Every http:// deployment crashed, though the scheme is documented as the supported way to select h2c, and the only workaround was to assert a certificate posture for a connection that has no certificate. The fourth was the blocker, and it is Phase 2's own: 4. The drain gate deferred to a Primary that cannot deliver. Redundancy roles are elected CLUSTER-WIDE; the alarm queue is PAIR-LOCAL. On the rig the elected driver Primary is central-1 — it carries the driver Akka role, replicates nobody's LocalDb and does not even run the alarm historian — so every driver node logged "Historian drain suspended", including the two site-b nodes that have no peer at all. Nothing drained anywhere, where before Phase 2 it drained fine. The cost is not a duplicate; it is the buffer growing to the capacity wall and evicting the audit trail it exists to protect. Fixed in three layers: a separate ShouldDrainAlarmHistory policy (unknown role drains; the two gates now deliberately disagree, and a test pins that); peer- host matching in DriverHostActor so a node stands down only for a Primary holding its rows; and AddAlarmHistorian short-circuiting the gate when replication is unconfigured — testing BOTH Replication:PeerAddress and SyncListenPort, since only the dialing half sets the former while both halves share the queue. Every one of these follows from the asymmetry: a false allow costs a duplicate row, which at-least-once delivery already accepts and payload-hash ids collapse; a false deny loses data silently. A third vacuous test, caught by the same delete-the-guard discipline: the role-view tests stayed green with the guard removed, because AwaitAssert polls until an assertion passes and the assertion was "reads open" — which is the SEEDED value, satisfied at the first poll before the actor processed anything. They now assert the sequence of published values through a recording view; the control then goes red for exactly the cases that matter. Migration evidence: 11 legacy rows across two deliberately overlapping files converged to exactly 9 identical rows on both nodes, proving D-6's payload-hash identity on real nodes rather than in a fixture. Open design fork, recorded in the gate doc rather than decided here: a pair cannot currently identify its own Primary, so both halves drain. Safe in every topology — nothing loses data — but the gate's de-duplication benefit is unrealised until roles are scoped per pair. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
2e4ccf7fe9 |
chore(localdb): phase-2 DoD sweep
Build: 0 errors solution-wide, and 0 warnings from every project this branch
touches. The ~816 solution-wide warnings are pre-existing xUnit1051 /
OTOPCUA0001 / CS86xx in untouched driver + client test projects.
Tests: full solution run compared against a full run on a detached worktree at
the pre-branch baseline
|
||
|
|
4480a7d755 |
docs(localdb): record deviations D-6 and D-7 in the phase-2 recon
D-6 (payload-hash migrator ids over mig-{node}-{legacyId}) and D-7 (the
plan's Host.Tests project does not exist) were captured only in commit
messages and the tasks file. They belong with the other deviations, where
the next reader looks.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|