7.1 KiB
7.1 KiB
Candidate Findings for the Next Review Cycle (surfaced during 2026-07-12 remediation)
These were discovered while remediating the 2026-07-12 backlog but were out of scope for it — each is either pre-existing, by-design residual, or a new observation. They are recorded here (not fixed) so the next review cycle can triage them. None blocks the 2026-07-12 cycle, which is complete.
| ID (proposed) | Area | Severity (est.) | Summary |
|---|---|---|---|
| NEXT-01 | Testing / macOS | Low | Fake-worker/e2e gateway tests fail on macOS under the default TMPDIR because the CoreFxPipe_mxaccess-gateway-{pid}-{sessionId} path exceeds the 104-char Unix-domain-socket sun_path limit under /var/folders/…/T/. Workaround today is TMPDIR=/tmp. Fix options: shorten the pipe name, or document the TMPDIR=/tmp requirement in docs/GatewayTesting.md. Surfaced independently by multiple remediation agents. |
| NEXT-02 | Clients (.NET, Java) | Low | The .NET and Java CLIs render the raw ReplayGap sentinel MxEvent on stream-events instead of a typed gap row — Java text mode prints 0 MX_EVENT_FAMILY_UNSPECIFIED. Same defect class as CLI-36 (Go) / CLI-35 (Python), which were fixed this cycle; the .NET/Java halves were out of scope. The cross-language smoke matrix now records this divergence honestly. |
| NEXT-03 | Gateway alarms | Low | GatewayAlarmMonitor.ApplyReconcile feed-repair broadcasts (the new acked-delta from GWC-26 and the pre-existing Raise/Clear repair) are at-least-once, not exactly-once: a periodic reconcile can synthesize a transition whose matching live transition is still buffered in the alarm lease, so both broadcast as indistinguishable duplicates on the alarm feed (StreamAlarms + dashboard hub). Pre-existing (the Raise/Clear repair always had it); GWC-26 documented the at-least-once contract rather than closing the race. Closing it needs reconcile/live serialization or a monotonic dedup marker. |
| NEXT-04 | Worker frame writer | Low | WRK-22/WRK-25 cancellation path: a frame Claimed by a concurrent lock-holder just before its caller's cancellation races in is never awaited by that caller; if the write then faults, TrySetException lands on a Task nobody observes (unobserved-task-exception). By-design residual, non-crash (no UnobservedTaskException handler registered), pre-existing to single-frame WRK-22 and amplified per-batch by WRK-25. Hygiene fix: attach a fault-observing continuation to abandoned/tombstoned frame completions. |
| NEXT-05 | Worker frame writer | Info | A batch whose remaining frames are tombstoned by cancellation leaves dead PendingFrame entries in _eventFrames/_controlFrames until a future DequeueNext pops and skips them. Same pre-existing behavior as single-frame WRK-22, amplified per-batch; in practice heartbeats purge them promptly, so not a real leak. |
| NEXT-06 | Testing / live LDAP | Medium | DashboardLdapLiveTests fixtures have drifted from the shared GLAuth directory, leaving the suite with no positive-proof coverage of the service-account bind. Its only success-path test, AuthenticateAsync_AdminInGwAdminGroup_Succeeds, binds admin/admin123, but the directory's admin user carries the standard dev password (scadaproj/infra/glauth/config.toml), so that assertion cannot pass. AuthenticateAsync_ReadOnlyUserMissingGwAdminGroup_Fails binds fixture user readonly, which does not exist in the GLAuth config at all — it passes for the wrong reason (user-not-found rather than the group-missing branch it names; the readonly name is in fact barred by the README's user/group case-collision rule). The three remaining tests are negative assertions that pass whether or not the service account can bind. Net effect: a green DashboardLdapLiveTests run proves nothing about the bind credential — surfaced during SEC-36, where the suite was considered as a substitute for the deferred dashboard-login check and rejected. Fix: realign the fixtures to real directory users (e.g. multi-role/gw-viewer) or add the missing users to the GLAuth config, and add one test that fails when the service-account credential is wrong. |
| NEXT-07 | Deployment / windev | High | The 10.100.0.48 (windev) gateway deployment is stale and crash-looping, and has been since at least 2026-08-06 (~10k Hosting-failed events/day). The deployed Server binary dates to 2026-06-25 and predates the auth-DB migration of 2026-07-15: it opens a schema-version-3 gateway-auth.db that it supports only at version 2 and aborts at startup, so the MxAccessGw service never reaches a listening state. Not a code defect in the current tree — a deploy-drift/operations gap — but it means the repo's only deployed host has been dark for over a day and any host-level verification (including SEC-36's dashboard-login check) is blocked until it is repaired. Fix: deploy a current Server build to windev, or restore/downgrade the auth DB to schema 2 if the old binary must stand. Worth asking separately why a service in a permanent restart loop raised no alert. Discovered during SEC-36. |
Operator actions still pending (from this cycle's runbooks)
These are live-infrastructure actions the operator must execute — the repo-side work is complete and merged:
SEC-36 — rotate the dev LDAP service-account credential perExecuted 2026-08-07: thedocs/runbooks/SEC-36-ldap-credential-rotation.md(generate new secret inscadaproj/infra/glauth, pre-stage the NSSM env var on deployed hosts, rotate GLAuth on10.100.0.35, verify dashboard login). The committed literal is gone from the working tree but remains recoverable from git history until rotation completes — rotation is the load-bearing half.serviceaccountpasssha256was replaced inscadaproj/infra/glauth/config.toml(commitaada53b) and the shared GLAuth recreated on10.100.0.35, so the literal recoverable from this repo's history no longer bindsdc=zb,dc=local. The new value lives only in the GLAuth hash, windev's NSSM environment, and dev user-secrets.wonder-app-vd03was out of scope (it binds a different,dc=scadalink/dc=scadabridgedirectory). Caveat: windev's dashboard/loginverification is deferred — that host's gateway is crash-looping on the unrelated stale-deployment fault filed as NEXT-07 — so the bind was verified directly byldapsearchascn=serviceaccount,dc=zb,dc=localinstead. SEC-36 is now fullyDone.TST-30 — register a second GiteaExecuted 2026-08-07:act_runneron10.100.0.35perdocs/runbooks/TST-30-second-ci-runner.mdto relieve the single-shared-runner bottleneck.gitea-runner-2(id 5, capacity 2) is online on10.100.0.35via the/opt/giteacompose stack, samecontainer.network: traefik, token from a0600file mount; the existing runner (id 1, capacity 4) was untouched. Concurrency verified — jobs from three runs ran simultaneously across both runners, and agitea-runner-2job cloned successfully fromhttp://gitea:3000. TST-30 is now fullyDone.- TST-25 follow-ups — old TST-05 (scheduled live-MXAccess smoke) is now covered by the
nightly-windevjob; old TST-24 (client wire tests in CI) is unblocked by the working Windows tier.