Files
mxaccessgw/archreview/2026-07-12/remediation/90-candidate-findings-next-cycle.md
T

7.1 KiB

Candidate Findings for the Next Review Cycle (surfaced during 2026-07-12 remediation)

These were discovered while remediating the 2026-07-12 backlog but were out of scope for it — each is either pre-existing, by-design residual, or a new observation. They are recorded here (not fixed) so the next review cycle can triage them. None blocks the 2026-07-12 cycle, which is complete.

ID (proposed) Area Severity (est.) Summary
NEXT-01 Testing / macOS Low Fake-worker/e2e gateway tests fail on macOS under the default TMPDIR because the CoreFxPipe_mxaccess-gateway-{pid}-{sessionId} path exceeds the 104-char Unix-domain-socket sun_path limit under /var/folders/…/T/. Workaround today is TMPDIR=/tmp. Fix options: shorten the pipe name, or document the TMPDIR=/tmp requirement in docs/GatewayTesting.md. Surfaced independently by multiple remediation agents.
NEXT-02 Clients (.NET, Java) Low The .NET and Java CLIs render the raw ReplayGap sentinel MxEvent on stream-events instead of a typed gap row — Java text mode prints 0 MX_EVENT_FAMILY_UNSPECIFIED. Same defect class as CLI-36 (Go) / CLI-35 (Python), which were fixed this cycle; the .NET/Java halves were out of scope. The cross-language smoke matrix now records this divergence honestly.
NEXT-03 Gateway alarms Low GatewayAlarmMonitor.ApplyReconcile feed-repair broadcasts (the new acked-delta from GWC-26 and the pre-existing Raise/Clear repair) are at-least-once, not exactly-once: a periodic reconcile can synthesize a transition whose matching live transition is still buffered in the alarm lease, so both broadcast as indistinguishable duplicates on the alarm feed (StreamAlarms + dashboard hub). Pre-existing (the Raise/Clear repair always had it); GWC-26 documented the at-least-once contract rather than closing the race. Closing it needs reconcile/live serialization or a monotonic dedup marker.
NEXT-04 Worker frame writer Low WRK-22/WRK-25 cancellation path: a frame Claimed by a concurrent lock-holder just before its caller's cancellation races in is never awaited by that caller; if the write then faults, TrySetException lands on a Task nobody observes (unobserved-task-exception). By-design residual, non-crash (no UnobservedTaskException handler registered), pre-existing to single-frame WRK-22 and amplified per-batch by WRK-25. Hygiene fix: attach a fault-observing continuation to abandoned/tombstoned frame completions.
NEXT-05 Worker frame writer Info A batch whose remaining frames are tombstoned by cancellation leaves dead PendingFrame entries in _eventFrames/_controlFrames until a future DequeueNext pops and skips them. Same pre-existing behavior as single-frame WRK-22, amplified per-batch; in practice heartbeats purge them promptly, so not a real leak.
NEXT-06 Testing / live LDAP Medium DashboardLdapLiveTests fixtures have drifted from the shared GLAuth directory, leaving the suite with no positive-proof coverage of the service-account bind. Its only success-path test, AuthenticateAsync_AdminInGwAdminGroup_Succeeds, binds admin/admin123, but the directory's admin user carries the standard dev password (scadaproj/infra/glauth/config.toml), so that assertion cannot pass. AuthenticateAsync_ReadOnlyUserMissingGwAdminGroup_Fails binds fixture user readonly, which does not exist in the GLAuth config at all — it passes for the wrong reason (user-not-found rather than the group-missing branch it names; the readonly name is in fact barred by the README's user/group case-collision rule). The three remaining tests are negative assertions that pass whether or not the service account can bind. Net effect: a green DashboardLdapLiveTests run proves nothing about the bind credential — surfaced during SEC-36, where the suite was considered as a substitute for the deferred dashboard-login check and rejected. Fix: realign the fixtures to real directory users (e.g. multi-role/gw-viewer) or add the missing users to the GLAuth config, and add one test that fails when the service-account credential is wrong.
NEXT-07 Deployment / windev High The 10.100.0.48 (windev) gateway deployment is stale and crash-looping, and has been since at least 2026-08-06 (~10k Hosting-failed events/day). The deployed Server binary dates to 2026-06-25 and predates the auth-DB migration of 2026-07-15: it opens a schema-version-3 gateway-auth.db that it supports only at version 2 and aborts at startup, so the MxAccessGw service never reaches a listening state. Not a code defect in the current tree — a deploy-drift/operations gap — but it means the repo's only deployed host has been dark for over a day and any host-level verification (including SEC-36's dashboard-login check) is blocked until it is repaired. Fix: deploy a current Server build to windev, or restore/downgrade the auth DB to schema 2 if the old binary must stand. Worth asking separately why a service in a permanent restart loop raised no alert. Discovered during SEC-36.

Operator actions still pending (from this cycle's runbooks)

These are live-infrastructure actions the operator must execute — the repo-side work is complete and merged:

  • SEC-36 — rotate the dev LDAP service-account credential per docs/runbooks/SEC-36-ldap-credential-rotation.md (generate new secret in scadaproj/infra/glauth, pre-stage the NSSM env var on deployed hosts, rotate GLAuth on 10.100.0.35, verify dashboard login). The committed literal is gone from the working tree but remains recoverable from git history until rotation completes — rotation is the load-bearing half. Executed 2026-08-07: the serviceaccount passsha256 was replaced in scadaproj/infra/glauth/config.toml (commit aada53b) and the shared GLAuth recreated on 10.100.0.35, so the literal recoverable from this repo's history no longer binds dc=zb,dc=local. The new value lives only in the GLAuth hash, windev's NSSM environment, and dev user-secrets. wonder-app-vd03 was out of scope (it binds a different, dc=scadalink/dc=scadabridge directory). Caveat: windev's dashboard /login verification is deferred — that host's gateway is crash-looping on the unrelated stale-deployment fault filed as NEXT-07 — so the bind was verified directly by ldapsearch as cn=serviceaccount,dc=zb,dc=local instead. SEC-36 is now fully Done.
  • TST-30 — register a second Gitea act_runner on 10.100.0.35 per docs/runbooks/TST-30-second-ci-runner.md to relieve the single-shared-runner bottleneck. Executed 2026-08-07: gitea-runner-2 (id 5, capacity 2) is online on 10.100.0.35 via the /opt/gitea compose stack, same container.network: traefik, token from a 0600 file mount; the existing runner (id 1, capacity 4) was untouched. Concurrency verified — jobs from three runs ran simultaneously across both runners, and a gitea-runner-2 job cloned successfully from http://gitea:3000. TST-30 is now fully Done.
  • TST-25 follow-ups — old TST-05 (scheduled live-MXAccess smoke) is now covered by the nightly-windev job; old TST-24 (client wire tests in CI) is unblocked by the working Windows tier.