feat(cluster): self-first seed ordering closes the boot-alone outage gap
Every node now lists ITSELF as seed-nodes[0] and its partner second. Akka runs FirstSeedNodeProcess -- the only bootstrap path that can form a NEW cluster when no peer answers InitJoin -- exclusively for seed-nodes[0]; every other node runs JoinSeedNodeProcess and retries InitJoin forever. That is why a lone cold-starting central-b never came Up (the "registered outage gap"), and self-first ordering closes it using Akka's own protocol. - 6 node appsettings swapped (the *-node-b configs; the -a nodes were already self-first). All 14 shipped node configs now satisfy the invariant. - StartupValidator enforces it at boot, comparing host AND port -- the invariant fails silently when broken, so it is enforced loudly. NOTE: the gitignored deploy/wonder-app-vd03/ overlay must be reordered before its next deploy or that node will refuse to boot. - SelfFirstSeedBootstrapTests: real in-process clusters at production failure-detection timings, incl. a falsifiability control proving the OLD peer-first ordering never forms. Rejected alternative (implemented, measured, discarded): an external self-form timer calling Cluster.Join(SelfAddress) after a window. It sits outside Akka's join handshake and so cannot tell "no seed answered" from "a seed answered and the join is in flight". On a routine standby restart the peer is alive but the join stalls behind removal of the node's own stale incarnation; a Join(self) during TryingToJoin abandons the in-flight join and forms a second cluster at the same address -- still split after 90s. Docs that claimed self-first ordering was unsafe for simultaneous cold start are corrected: while mutually reachable the InitJoin handshake converges them to one cluster (measured).
This commit is contained in:
@@ -31,8 +31,20 @@ public class ClusterOptions
|
||||
// when the binding sites can be updated in the same commit.
|
||||
|
||||
/// <summary>
|
||||
/// Akka.NET cluster seed nodes. Both nodes are seed nodes — each node lists
|
||||
/// itself and its partner — so either can start first and form the cluster.
|
||||
/// Akka.NET cluster seed nodes. Both nodes are seed nodes — each node lists itself and its
|
||||
/// partner.
|
||||
/// <para>
|
||||
/// <b>ORDER IS LOAD-BEARING (decision 2026-07-22): every node must list ITSELF first.</b>
|
||||
/// Akka runs <c>FirstSeedNodeProcess</c> — the only bootstrap path that can form a NEW
|
||||
/// cluster when no peer answers <c>InitJoin</c> — exclusively when <c>seed-nodes[0]</c> is
|
||||
/// this node's own address; any other node runs <c>JoinSeedNodeProcess</c> and retries
|
||||
/// <c>InitJoin</c> forever. So merely listing both nodes does NOT mean either can start
|
||||
/// first: a node that lists its partner first can never cold-start while that partner is
|
||||
/// down (the "registered outage gap", <c>docker/README.md</c>). Self-first ordering closes
|
||||
/// it using Akka's own protocol, which — unlike an external self-form timer — is part of
|
||||
/// the join handshake and so cannot mistake an in-flight join for an absent peer.
|
||||
/// Enforced at boot by <c>StartupValidator</c>.
|
||||
/// </para>
|
||||
/// Must contain at least one entry.
|
||||
/// </summary>
|
||||
public List<string> SeedNodes { get; set; } = new();
|
||||
|
||||
Reference in New Issue
Block a user