1ec831883c
Phase 2 of the per-cluster mesh program: move the three central->node command channels (deployments, driver-control, alarm-commands) and the deployment-acks reply channel off DistributedPubSub onto Akka ClusterClient, behind a config flag. Recon of both repos turned up three corrections to the design doc's account of the sister project, all folded into the plan: - Design doc S2 claims "every unhandled message replies with a typed failure". ScadaBridge has no ReceiveAny/Unhandled override in either comm actor; the typed-failure idiom fires only for a MISSING REGISTRATION, per message type. - "No central buffering" was never a setting they wrote. They have zero akka.cluster.client HOCON and run the defaults (buffer-size 1000, reconnect-timeout off). We choose buffer-size = 0 deliberately. - Publish -> Send is the wrong substitution. Today's deploy notify is a broadcast every DriverHostActor receives (no ClusterId filter exists on the node side), so it must become SendToAll. Send would deploy to exactly one node of the fleet and seal green on partial acks. Also records the single-mesh duplicate-delivery trap (one ClusterClient in Phase 2, not one per Cluster -- a receptionist serves its whole mesh) and one deviation from the program plan's exit gate: there is no cross-boundary Ask in Phase 2, because every migrated command is fire-and-forget with a local reply, so "an Ask timing out cleanly" cannot be run as written. Marks the program tracking tables current: prereq + Phase 1 are DONE (they still read "not executed"), Phase 2 in progress. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
16 lines
2.0 KiB
JSON
16 lines
2.0 KiB
JSON
{
|
|
"planPath": "docs/plans/2026-07-22-per-cluster-mesh-program.md",
|
|
"note": "PROGRAM plan: each task = write that phase's detailed plan (writing-plans), execute it (executing-plans), run its exit gate, update the tracking tables. Prereq task 0 is the separate selfform-fallback plan in this repo.",
|
|
"tasks": [
|
|
{"id": 0, "subject": "Prereq: execute 2026-07-22-selfform-fallback-and-manual-failover.md (7 tasks)", "status": "completed", "note": "Shipped as self-first seed ordering + AkkaClusterOptionsValidator, NOT the planned SelfFormAfter watchdog (live gate caught it islanding a failed-over node). Merged a78425ea."},
|
|
{"id": 1, "subject": "Phase 1: ClusterNode Akka+gRPC address columns; ConfigPublishCoordinator ack set from DB", "status": "completed", "note": "Merged 7654f24d. Escape hatch is ClusterNode.MaintenanceMode, not Enabled=0."},
|
|
{"id": 2, "subject": "Phase 2: comm actors + receptionist + ClusterClient transport (deploy notify/acks, driver-control)", "status": "in_progress", "blockedBy": [1]},
|
|
{"id": 3, "subject": "Phase 3: config fetch-and-cache from central; LocalDb steady-state config store (live gate)", "status": "pending", "blockedBy": [2]},
|
|
{"id": 4, "subject": "Phase 4: cut driver-side ConfigDb — EfAlarmConditionStateStore to LocalDb, DbHealthProbeActor ServiceLevel input, OpcUaPublishActor audit (live gate)", "status": "pending", "blockedBy": [3]},
|
|
{"id": 5, "subject": "Phase 5: gRPC oneof stream contract; migrate 7 observability topics; AdminUI reconnect story", "status": "pending", "blockedBy": [2]},
|
|
{"id": 6, "subject": "Phase 6: mesh partition — per-pair seeds, cluster-scoped roles/singletons, co-location rig rewrite, secrets replication re-scope", "status": "pending", "blockedBy": [0, 4, 5]},
|
|
{"id": 7, "subject": "Phase 7: failover drill per pair (incl. shared-VM failure) + close auto-down 1v1 and self-first cold-start-alone live gates + operator runbook", "status": "pending", "blockedBy": [6]}
|
|
],
|
|
"lastUpdated": "2026-07-22T00:00:00Z"
|
|
}
|