Files
mxaccessgw/archreview/2026-07-12/remediation/90-candidate-findings-next-cycle.md
T
Joseph Doherty 129e47e541 docs(tracking): close NEXT-07, file NEXT-08/09/10, record the runner token reset
NEXT-07 is struck: windev was redeployed from origin/main (a346d51) and the service is
healthy, and the root cause the row predicted is confirmed -- the 2026-06-25 build's
Auth.ApiKeys 0.1.2.0 supports auth-DB schema 2 while the database sits at schema 3, which
is the current shared-lib version, so deploying forward was the fix rather than touching
the DB. The original text stays for the triage record.

Three findings surfaced by that work, each deliberately left for the next cycle rather
than patched in passing:

- NEXT-08: the shared GLAuth offers no TLS, so SEC-06 makes GatewayConfiguration.md's
  "deployed hosts must set Ldaps or StartTls" unsatisfiable for anything genuinely
  labelled Production. windev's relabel to Staging is honest for a dev rig but defers
  the posture question rather than answering it.
- NEXT-09: Directory.Build.props:29 quotes a path ending in a backslash, so the SHA-stamp
  git invocation is malformed on Windows and ContinueOnError stamps git's stderr into
  InformationalVersion -- a Windows binary cannot be correlated to a commit, which is what
  TST-11 exists to guarantee.
- NEXT-10: glauth.md's pre-provisioned-user table contradicts both the directory and its
  own dashboard section, and was the root cause of the NEXT-06 fixture drift. Reconciling
  it sweeps the OPC-UA group taxonomy, so it is scoped out here on purpose.

The TST-30 runner work is hygiene, not closure: runner-1 now mounts its registration token
from a 0600 file like runner-2, but both still share one instance-scope token that was
world-readable for months and is provably still live. Gitea 1.26.4 cannot rotate it from
the CLI or API, so the UI reset is recorded as a pending operator action with its
follow-through (refresh the token file, shred the token-bearing compose backups).
2026-08-07 10:31:29 -04:00

14 KiB

Candidate Findings for the Next Review Cycle (surfaced during 2026-07-12 remediation)

These were discovered while remediating the 2026-07-12 backlog but were out of scope for it — each is either pre-existing, by-design residual, or a new observation. They are recorded here so the next review cycle can triage them. None blocks the 2026-07-12 cycle, which is complete. Rows struck through have since been fixed ahead of that cycle; the original finding text is kept so the triage record stays readable.

ID (proposed) Area Severity (est.) Summary
NEXT-01 Testing / macOS Low Fake-worker/e2e gateway tests fail on macOS under the default TMPDIR because the CoreFxPipe_mxaccess-gateway-{pid}-{sessionId} path exceeds the 104-char Unix-domain-socket sun_path limit under /var/folders/…/T/. Workaround today is TMPDIR=/tmp. Fix options: shorten the pipe name, or document the TMPDIR=/tmp requirement in docs/GatewayTesting.md. Surfaced independently by multiple remediation agents.
NEXT-02 Clients (.NET, Java) Low The .NET and Java CLIs render the raw ReplayGap sentinel MxEvent on stream-events instead of a typed gap row — Java text mode prints 0 MX_EVENT_FAMILY_UNSPECIFIED. Same defect class as CLI-36 (Go) / CLI-35 (Python), which were fixed this cycle; the .NET/Java halves were out of scope. The cross-language smoke matrix now records this divergence honestly.
NEXT-03 Gateway alarms Low GatewayAlarmMonitor.ApplyReconcile feed-repair broadcasts (the new acked-delta from GWC-26 and the pre-existing Raise/Clear repair) are at-least-once, not exactly-once: a periodic reconcile can synthesize a transition whose matching live transition is still buffered in the alarm lease, so both broadcast as indistinguishable duplicates on the alarm feed (StreamAlarms + dashboard hub). Pre-existing (the Raise/Clear repair always had it); GWC-26 documented the at-least-once contract rather than closing the race. Closing it needs reconcile/live serialization or a monotonic dedup marker.
NEXT-04 Worker frame writer Low WRK-22/WRK-25 cancellation path: a frame Claimed by a concurrent lock-holder just before its caller's cancellation races in is never awaited by that caller; if the write then faults, TrySetException lands on a Task nobody observes (unobserved-task-exception). By-design residual, non-crash (no UnobservedTaskException handler registered), pre-existing to single-frame WRK-22 and amplified per-batch by WRK-25. Hygiene fix: attach a fault-observing continuation to abandoned/tombstoned frame completions.
NEXT-05 Worker frame writer Info A batch whose remaining frames are tombstoned by cancellation leaves dead PendingFrame entries in _eventFrames/_controlFrames until a future DequeueNext pops and skips them. Same pre-existing behavior as single-frame WRK-22, amplified per-batch; in practice heartbeats purge them promptly, so not a real leak.
NEXT-06 Testing / live LDAP Medium Resolved 2026-08-07 — fixtures realigned to the shared directory (admin/password for the GwAdmin success path, gw-viewer/password for the bind-succeeds-but-no-role path); verified Failed: 0, Passed: 5 live against the shared GLAuth at 10.100.0.35:3893, so the success-path assertion (GwAdmin group claim + Admin role claim) now fails if the service-account credential is wrong. Original finding: DashboardLdapLiveTests fixtures have drifted from the shared GLAuth directory, leaving the suite with no positive-proof coverage of the service-account bind. Its only success-path test, AuthenticateAsync_AdminInGwAdminGroup_Succeeds, binds admin/admin123, but the directory's admin user carries the standard dev password (scadaproj/infra/glauth/config.toml), so that assertion cannot pass. AuthenticateAsync_ReadOnlyUserMissingGwAdminGroup_Fails binds fixture user readonly, which does not exist in the GLAuth config at all — it passes for the wrong reason (user-not-found rather than the group-missing branch it names; the readonly name is in fact barred by the README's user/group case-collision rule). The three remaining tests are negative assertions that pass whether or not the service account can bind. Net effect: a green DashboardLdapLiveTests run proves nothing about the bind credential — surfaced during SEC-36, where the suite was considered as a substitute for the deferred dashboard-login check and rejected. Fix: realign the fixtures to real directory users (e.g. multi-role/gw-viewer) or add the missing users to the GLAuth config, and add one test that fails when the service-account credential is wrong.
NEXT-07 Deployment / windev High Resolved 2026-08-07 — a fresh portable framework-dependent publish of origin/main (a346d51) was built in a clean clone at C:\build\mxgw-redeploy, deployed to C:\publish\mxaccessgw\Server-20260807, and the MxAccessGw NSSM service repointed at it; the service now holds a stable PID with both 5120/5130 listening, a worker spawned, the Galaxy snapshot restored (129 objects / 56,731 attributes) and a clean event log. Root cause confirmed as the version skew this row predicted: the deployed 2026-06-25 build carried ZB.MOM.WW.Auth.ApiKeys 0.1.2.0, which supports auth-DB schema 2, against a gateway-auth.db stamped at schema 3 on 2026-07-15 by an ephemeral run of newer code — schema 3 is the current shared-lib version (SqliteAuthSchema.CurrentVersion=3 in Auth 0.1.5), so the redeploy is the forward fix and the DB was left alone. Rollback artifacts kept: C:\ProgramData\MxGateway\gateway-auth.db.bak-next07 (with -wal/-shm) and the previous C:\publish\mxaccessgw\Server directory. Two side effects worth recording: the old deploy's appsettings.json held the LDAP bind password in plaintext on disk, while the new one keeps the repo's ${secret:ldap/mxgateway/bind} token with the NSSM environment supplying the value, so no plaintext LDAP secret remains on that host; and the redeploy tripped the SEC-06 Ldap:Transport=None production hard-stop (GatewayOptionsValidator.cs:178), resolved by relabelling the host — windev runs Dashboard:DisableLogin=true, which this repo's own docs mark dev/test-only, so its Production label contradicted its configuration and DOTNET_ENVIRONMENT was changed to Staging (that one NSSM environment entry only; the other nine preserved byte-identical). SEC-06 is untouched for genuinely production hosts — see NEXT-08 for the posture problem that relabelling defers. Original finding: The 10.100.0.48 (windev) gateway deployment is stale and crash-looping, and has been since at least 2026-08-06 (~10k Hosting-failed events/day). The deployed Server binary dates to 2026-06-25 and predates the auth-DB migration of 2026-07-15: it opens a schema-version-3 gateway-auth.db that it supports only at version 2 and aborts at startup, so the MxAccessGw service never reaches a listening state. Not a code defect in the current tree — a deploy-drift/operations gap — but it means the repo's only deployed host has been dark for over a day and any host-level verification (including SEC-36's dashboard-login check) is blocked until it is repaired. Fix: deploy a current Server build to windev, or restore/downgrade the auth DB to schema 2 if the old binary must stand. Worth asking separately why a service in a permanent restart loop raised no alert. Discovered during SEC-36.
NEXT-08 Security / LDAP posture Medium The shared GLAuth offers no TLS, so SEC-06 makes it undeployable from a Production-labelled host. GatewayOptionsValidator (src/ZB.MOM.WW.MxGateway.Server/Configuration/GatewayOptionsValidator.cs:178) refuses to start when Ldap:Transport=None in the Production environment, and docs/GatewayConfiguration.md's Transport row states "Deployed hosts must set Ldaps or StartTls" — but the shared instance at 10.100.0.35:3893 has [ldaps] enabled=false, port 3894 closed, and answers StartTLS with protocolError, so neither value can work against it. That instruction is currently unsatisfiable for every host that authenticates there. windev sidestepped it on 2026-08-07 by moving to the Staging environment name (NEXT-07), which is honest for a dev/test rig but is not available to a real production host. Resolution needs either LDAPS/StartTLS on the shared GLAuth (certificate plus a trust story on each gateway host) or an explicit written posture decision that production gateways bind a different, TLS-capable directory. Surfaced during the NEXT-07 redeploy.
NEXT-09 Build / versioning Low Windows builds stamp git's error text into InformationalVersion. src/Directory.Build.props:29 runs git -C "$(MSBuildThisFileDirectory)" …; MSBuild's directory property ends in a backslash, which escapes the closing quote, so the command is malformed on Windows. The target carries ContinueOnError, so the failure is silent and git's stderr is captured as the revision — an observed stamp reads 0.1.2+fatal: cannot change to …. Any Windows build without a preset SourceRevisionId therefore ships a binary that cannot be correlated back to a commit, defeating the point of TST-11. Not reproducible on macOS/Linux, where the separator is /. Fix sketch: append . to the path or trim the trailing separator before quoting. Surfaced while identifying the deployed binary during NEXT-07.
NEXT-10 Docs / glauth Medium glauth.md's "Pre-provisioned users" table contradicts both the directory and the rest of its own file. It documents readonly/readonly123 and admin/admin123, neither of which matches scadaproj/infra/glauth/config.toml (readonly does not exist there; admin carries the standard dev password), and lists the ReadOnly gid as 5501 against an actual 5601. Its dashboard section, by contrast, is correct — so the file is internally inconsistent and a reader cannot tell which half to trust. This table was the root cause of the NEXT-06 fixture drift, and it has propagated further: docs/GatewayTesting.md's MXGATEWAY_LIVE_MXACCESS_WRITE_SECURED_PASSWORD default and the matching literal in WorkerLiveMxAccessSmokeTests both take admin123 from it. Deliberately not fixed in the 2026-08-07 pass: the table is entangled with the OPC-UA group taxonomy (gids, role mapping, and the sister-repo consumers of the same directory), so reconciling it means sweeping that taxonomy as one unit rather than patching two rows.

Operator actions still pending (from this cycle's runbooks)

These are live-infrastructure actions the operator must execute — the repo-side work is complete and merged:

  • SEC-36 — rotate the dev LDAP service-account credential per docs/runbooks/SEC-36-ldap-credential-rotation.md (generate new secret in scadaproj/infra/glauth, pre-stage the NSSM env var on deployed hosts, rotate GLAuth on 10.100.0.35, verify dashboard login). The committed literal is gone from the working tree but remains recoverable from git history until rotation completes — rotation is the load-bearing half. Executed 2026-08-07: the serviceaccount passsha256 was replaced in scadaproj/infra/glauth/config.toml (commit aada53b) and the shared GLAuth recreated on 10.100.0.35, so the literal recoverable from this repo's history no longer binds dc=zb,dc=local. The new value lives only in the GLAuth hash, windev's NSSM environment, and dev user-secrets. wonder-app-vd03 was out of scope (it binds a different, dc=scadalink/dc=scadabridge directory). Caveat, since closed: windev's dashboard /login verification was deferred while that host's gateway was crash-looping on the unrelated stale-deployment fault filed as NEXT-07, so the bind was verified directly by ldapsearch as cn=serviceaccount,dc=zb,dc=local instead. After the 2026-08-07 redeploy the real check ran on windev (login as multi-role → 302 + dashboard cookie, anonymous control → /login), so the rotated credential is now proven through the DashboardAuthenticator path itself; see docs/runbooks/SEC-36-ldap-credential-rotation.md Correction 3. SEC-36 is fully Done.
  • TST-30 — register a second Gitea act_runner on 10.100.0.35 per docs/runbooks/TST-30-second-ci-runner.md to relieve the single-shared-runner bottleneck. Executed 2026-08-07: gitea-runner-2 (id 5, capacity 2) is online on 10.100.0.35 via the /opt/gitea compose stack, same container.network: traefik, token from a 0600 file mount; the existing runner (id 1, capacity 4) was untouched. Concurrency verified — jobs from three runs ran simultaneously across both runners, and a gitea-runner-2 job cloned successfully from http://gitea:3000. TST-30 is now fully Done.
  • TST-30 follow-up — reset the Gitea instance runner registration token. Runner-1's compose block was moved to the same 0600 file-mount pattern as runner-2 on 2026-08-07 (compose and both backups now 0600 root:root, runner-1 recreated with its identity intact), but hygiene alone does not retire the token: both runners share one instance-scope registration token that was world-readable for roughly five months and is still live — a probe registered runner id 6 with it, then deleted it. Gitea 1.26.4 exposes no rotation via CLI or API (both paths are get-or-create and hand back the same value), so the reset must be done in the admin web UI ("Reset registration token"). Afterwards, refresh /opt/gitea/runner_token on 10.100.0.35 and shred the two token-bearing compose backups — they are the last copies of the old value. See docs/runbooks/TST-30-second-ci-runner.md.
  • TST-25 follow-ups — old TST-05 (scheduled live-MXAccess smoke) is now covered by the nightly-windev job; old TST-24 (client wire tests in CI) is unblocked by the working Windows tier.