Joseph Doherty
992f343f11
docs(archreview): PLAN-02 tasks.json — mark T13,T15,T17 complete (batch 5)
2026-07-08 21:02:15 -04:00
Joseph Doherty
499a37ac78
perf(store-and-forward): enable WAL + per-connection busy_timeout on the S&F SQLite store
2026-07-08 21:01:53 -04:00
Joseph Doherty
7626d5ba34
test(communication): pin wire round-trip + type identity of cross-cluster message contracts
2026-07-08 21:00:02 -04:00
Joseph Doherty
4bdd5f0e4a
fix(notifications): Notify.Send enqueues unbounded (maxRetries 0) and defers to sweep — no 30s script-thread block, no stranding
2026-07-08 20:56:09 -04:00
Joseph Doherty
e2944d201c
docs(archreview): PLAN-02 tasks.json — mark T11,T12,T14 complete (batch 4)
2026-07-08 20:48:57 -04:00
Joseph Doherty
c3a8576863
fix(notifications): park corrupt buffered notification payloads instead of silently discarding as delivered (supersedes StoreAndForward-018)
2026-07-08 20:48:16 -04:00
Joseph Doherty
05a71a0260
feat(store-and-forward): deferToSweep enqueue mode with immediate sweep kick for latency-sensitive callers
2026-07-08 20:46:08 -04:00
Joseph Doherty
283ce0647c
fix(store-and-forward): standby applies Add/Park/Requeue as upserts so a lost Add self-heals
2026-07-08 20:42:39 -04:00
Joseph Doherty
9fe3452790
docs(archreview): PLAN-02 tasks.json — mark T8,T9,T16 complete (batch 3)
2026-07-08 20:12:03 -04:00
Joseph Doherty
66cb0febad
fix(communication): replicate heartbeat marks to the peer central node (closes post-failover offline blind window)
2026-07-08 20:11:54 -04:00
Joseph Doherty
59d9cba621
perf(store-and-forward): parallel per-target sweep lanes (cap 4), sequential within lane
2026-07-08 20:07:30 -04:00
Joseph Doherty
1459f9407f
perf(store-and-forward): short-circuit a (category,target) lane after its first transient failure per sweep
2026-07-08 20:06:06 -04:00
Joseph Doherty
03dcd55958
docs(archreview): PLAN-02 tasks.json — mark T4,T6,T7 complete (batch 2)
2026-07-08 17:48:39 -04:00
Joseph Doherty
41ed57421d
fix(host): wire S&F active-node delivery gate on site nodes (SelfIsPrimary); align S&F doc standby-passive wording
2026-07-08 17:48:31 -04:00
Joseph Doherty
1a53b8082a
perf(store-and-forward): bound retry sweep with SweepBatchLimit (default 500) on the due-rows query
2026-07-08 17:45:11 -04:00
Joseph Doherty
e75cd6f3d8
fix(communication): debug-stream teardown/failover unsubscribes via TryGet, never creates or disposes shared channels
2026-07-08 17:43:40 -04:00
Joseph Doherty
80013639ef
docs(archreview): PLAN-02 tasks.json — mark T3,T5,T10 complete (batch 1)
2026-07-08 17:35:39 -04:00
Joseph Doherty
70e7dfd7b9
fix(store-and-forward): inline ordered replication dispatch; Warning + counter on replication failures
2026-07-08 17:35:29 -04:00
Joseph Doherty
ca5e79e922
fix(communication): key gRPC client cache by (site, endpoint) so per-session failover cannot dispose shared channels
2026-07-08 17:33:03 -04:00
Joseph Doherty
76c10c5de6
fix(store-and-forward): gate the retry sweep behind an active-node delivery gate (standby must be passive)
2026-07-08 17:30:21 -04:00
Joseph Doherty
ec47cb5612
docs(archreview): PLAN-01 complete (23/23) — sync tracker + tasks.json; log deploy-artifact deferral
2026-07-08 17:12:46 -04:00
Joseph Doherty
dc7a613dae
docs(cleanup): sync Traefik + Host docs to oldest-member active-node semantics + DB-gated /health/active; TLS roadmap note
...
Renames stale ActiveNodeHealthCheck references to OldestNodeActiveHealthCheck
(Task 7 rename) across the requirements + components doc sets, corrects
leader→oldest-member wording and the embedded IsActiveNode snippet, and adds an
explicit 'production TLS profile not yet implemented' roadmap note.
deploy/wonder-app-vd03 NodeName overlay edits deferred: that production deploy
artifact is not tracked in this repo (same as Tasks 16 & 20).
2026-07-08 17:09:34 -04:00
Joseph Doherty
5117fa97b9
test(docker): failover drill script — SIGKILL the active central node, assert Traefik recovery
2026-07-08 17:05:31 -04:00
Joseph Doherty
1f1dbd916d
docs(host): REQ-HOST-6 documents the hand-rolled HOCON bootstrap; drop unused Akka.Hosting packages
2026-07-08 16:59:19 -04:00
Joseph Doherty
48e97fee01
feat(host): exit process on unexpected ActorSystem termination — completes the down-if-alone recovery loop
...
WhenTerminated watchdog calls IHostApplicationLifetime.StopApplication() when
the ActorSystem dies outside StopAsync (SBR self-down + run-coordinated-shutdown-when-down),
so the service supervisor restarts the node as a fresh incarnation. Optional
ctor param keeps every existing construction site compiling.
Plan's deploy/wonder-app-vd03/install.ps1 service-recovery-actions edit is
skipped: that production deploy artifact is not tracked in this repo (Task 16
skipped its appsettings.Central.json for the same reason).
2026-07-08 16:56:24 -04:00
Joseph Doherty
dea69842d5
fix(host): rate-limit dead-letter warnings (10/min + suppression summary); metric counting unchanged
2026-07-08 16:52:25 -04:00
Joseph Doherty
3ce21734b8
fix(docker): 30s stop_grace_period so graceful redeploys aren't SIGKILLed mid-CoordinatedShutdown
2026-07-08 16:40:17 -04:00
Joseph Doherty
ce66e194a8
feat(cluster): AllowSingleNodeCluster flag — single-seed installs drop the phantom seed
...
Validator lowers the seed-count floor to 1 when AllowSingleNodeCluster is set,
otherwise still requires 2 and points operators at the flag. The plan's
deploy/wonder-app-vd03/appsettings.Central.json edit is skipped: that production
deploy artifact is not tracked in this repo (only its deployment record doc is).
2026-07-08 16:39:12 -04:00
Joseph Doherty
0e3c1df1a7
fix(host): dev site seed no longer targets the metrics port; validator rejects seed-vs-MetricsPort
2026-07-08 16:37:43 -04:00
Joseph Doherty
d962c77bb7
fix(health): evict deleted sites from the aggregator on the periodic site refresh
2026-07-08 16:27:38 -04:00
Joseph Doherty
87c7255912
feat(ui): health dashboard shows metrics-stale badge and offline-since timestamp
2026-07-08 16:23:43 -04:00
Joseph Doherty
c73b7faa11
feat(health): metrics-stale signal + status-transition timestamps; spec now matches heartbeat-liveness code
2026-07-08 16:21:30 -04:00
Joseph Doherty
15d91d760f
docs(archreview): sync PLAN-01 tracker + tasks.json — Wave 1 T2,T5–T12 done (12/23)
2026-07-08 16:14:46 -04:00
Joseph Doherty
6300d6d399
fix(health): acked SendAsync transport — report-loss counter restore is live, not dead code
2026-07-08 16:12:48 -04:00
Joseph Doherty
17af376a8e
feat(comm): ack site health reports end-to-end (SiteHealthReportAck, additive contract)
2026-07-08 16:09:08 -04:00
Joseph Doherty
f7b9d342e4
refactor(host): central singletons via CentralSingletonRegistrar — outbox/audit-ingest gain drain tasks
2026-07-08 16:05:52 -04:00
Joseph Doherty
c255ec31c9
feat(host): CentralSingletonRegistrar — canonical singleton manager+proxy+drain helper
2026-07-08 15:59:01 -04:00
Joseph Doherty
7138d47630
test(cluster): restarted original node stays standby — oldest-member active semantics proven (graceful-leave variant per 2-node keep-oldest gap)
2026-07-08 15:57:45 -04:00
Joseph Doherty
f164f62a8b
fix(cluster): give cluster-leave phase a 15s timeout so 10s singleton drains can complete
2026-07-08 15:55:06 -04:00
Joseph Doherty
ff89e3145c
fix(host): /health/active = oldest member + DB reachable — partition-safe Traefik routing
2026-07-08 15:51:56 -04:00
Joseph Doherty
da6d5ae12c
fix(host): unify active-node on oldest-member semantics — purge/self-report/labels now track the singleton host
2026-07-08 15:48:07 -04:00
Joseph Doherty
6320be9954
feat(host): ClusterActivityEvaluator — oldest-member (singleton-host) active-node definition
2026-07-08 15:43:13 -04:00
Joseph Doherty
6b5c70dd19
docs(archreview): reconcile master tracker + tasks.json with P0 completion (10 tasks done) and SBR follow-up
2026-07-08 15:37:41 -04:00
Joseph Doherty
3f8e18c6f7
docs(archreview): mark all six P0 items delivered in the master tracker
2026-07-08 15:18:55 -04:00
Joseph Doherty
099728f05a
fix(transport): wire template inheritance edges on bundle import (C3) — derived templates no longer land as roots
2026-07-08 15:17:47 -04:00
Joseph Doherty
6ff90702f0
docs(archreview): register active-node-crash SBR gap found during P0 (keep-oldest 2-node limitation)
2026-07-08 15:11:03 -04:00
Joseph Doherty
b8d91dcc5b
test(cluster): prove SBR downs a hard-crashed node and the oldest survivor keeps its singleton
...
Rewritten from the plan's oldest-crash scenario: empirically, 2-node keep-oldest downs
the non-oldest partition, so a hard crash of the OLDEST makes the younger survivor down
itself (total cluster loss) — the survivable case is a crash of the younger node. The
member-removal assertion has teeth (impossible under the pre-fix NoDowning default).
See the test's XML doc for the active/oldest-node-crash gap.
2026-07-08 15:06:54 -04:00
Joseph Doherty
e3a6603c74
test(cluster): add TwoNodeClusterFixture — real two-node in-process cluster rig from production HOCON
2026-07-08 14:58:34 -04:00
Joseph Doherty
f91e75bfe4
fix(communication): guard per-site ClusterClient creation; add real-factory address-edit lifecycle test
2026-07-08 14:57:13 -04:00
Joseph Doherty
75cbac4478
fix(communication): generation-suffixed, sanitized ClusterClient actor names to prevent recreate name collision
2026-07-08 14:55:05 -04:00