Files
ScadaBridge/docs/plans/phase-8-checklist.md
T
Joseph Doherty da65605e41 docs(plans): WP-4 target-scale load test RESULTS + close register #25 and row 50
Full-scale run executed on this machine: 10 sites x 500 instances x 75 tags =
375,000 live tag subscriptions, 37,518 updates/s achieved vs 37,500 nominal,
45,021,375 updates over a 20-minute steady-state window (M4 Pro, 14 cores, 48 GB,
with the 8-node docker rig still running so the figures are pessimistic).

11 clean passes, 1 pass with a caveat, 0 failures:
  tag latency  P50 0.88ms  P99 4.57ms  max 37.41ms  (1.1M samples, end-to-end)
  stream       900,675 delivered, 0 dropped at 100 live subscribers
  health       collect+ingest P99 0.31ms, 10/10 sites tracked
  debug view   P99 2.19ms, 264 completed, 0 timeouts
  deploy       500 instances to a site in 2.6s
  cpu          41% of ONE core = 2.9% of the box
  memory       working-set slope +8.83 MB/min

Three findings, reported rather than tuned away:
  F1 (Low) 20 min with ZERO gen-2 collections cannot fully settle the leak
     question; the heap demonstrably sawtooths but an uncompacted gen-2 makes a
     positive slope ambiguous. The 1-hour run would settle it. Not tuned.
  F2 (informational, by design) a deferred S&F backlog sits for one full
     DefaultRetryInterval (28.9s measured) before anything drains --
     EnqueueAsync(attemptImmediateDelivery:false) stamps LastAttemptAt. Easy to
     misread as slow drainage, so drain is reported as two numbers: retry wait,
     then 3,533 msg/s of actual capacity.
  F3 (positive) slow-subscriber isolation is TOTAL: 4 healthy subscribers at
     100.00% with zero drops while a peer lost 197,028/200,000 events entirely
     within its own bounded channel. Mechanism recorded link by link.

Also records what the run does NOT prove (not clustered, not real-network, not
real OPC UA, not 1 hour) and the four WP-4 sub-criteria this harness does not
cover, so the evidence is not over-read.

Register rows 25 and 50 -> RESOLVED 2026-08-15; remediation execution-log
residual 7 -> resolved; phase-8-checklist WP-4 section replaced with the
measured numbers.
2026-08-15 02:55:59 -04:00

49 lines
2.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Phase 8 Execution Checklist
**Update 2026-08-07 (truth sweep):** This checklist previously read
"Status: Complete / Tests: All passing / Build: 0 errors, 0 warnings" — a
107-byte stub with no per-work-package results and no linked run. That claim
was **unevidenced** and was retired.
**Update 2026-08-15 (WP-4 run, evidenced):** The target-scale load test has now
actually been run, at full scale, with results published.
## WP-4 — Load/Performance Testing at Target Scale
**Status: measured 2026-08-15.** Not a claim — a run.
- Scale executed: **10 sites × 500 instances × 75 tags = 375,000 live tag
subscriptions**, 37,518 tag updates/second achieved against 37,500 nominal
(100.0%), **45,021,375 updates** offered over a 20-minute steady-state window.
- Headline results: end-to-end tag update latency **P50 0.88 ms / P99 4.57 ms /
max 37.41 ms**; **zero events dropped** at 100 live stream subscribers; health
report collect+ingest **P99 0.31 ms** with 10/10 sites tracked centrally;
debug view snapshot under load **P99 2.19 ms with 0 timeouts**; 500 instances
deployed to a site in **2.6 s**; CPU **2.9% of the box**; working-set slope
**+8.83 MB/min**.
- **11 clean passes, 1 pass with a documented caveat, 0 failures.**
- Deferred-work register **#25** and **row 50** are closed by this run.
**Results:** `2026-08-15-target-scale-load-test-results.md` — per-criterion
numbers vs thresholds, three findings (F1 memory-window limitation, F2 the S&F
retry-interval wait, F3 slow-subscriber isolation), and an explicit "what this
does not prove" section.
**Design, thresholds and deviations:**
`2026-08-15-target-scale-load-test-design.md` — including the documented
deviation from the protocol's 1-hour sustained window to 20 minutes, and the
command for the full 1-hour version.
**Harness:** `tests/ZB.MOM.WW.ScadaBridge.LoadHarness/` (standalone executable),
with a CI-scale smoke `[Fact]` at
`tests/ZB.MOM.WW.ScadaBridge.PerformanceTests/TargetScale/TargetScaleHarnessSmokeTests.cs`
so the evidence can be regenerated rather than bit-rotting.
**Still open within WP-4** (not covered by this harness — central-cluster and
browser concerns): `[xc-3]` site event logging volume within 30-day/1 GB limits,
`[xc-4]` audit logging not degrading central, `[xc-5]` template
flattening/validation for large templates, `[xc-8]` UI workflow responsiveness.
See `phase-8-production-readiness.md:152-170` (WP-4) and `:314-320`
(test protocol) for the original scope.