Compare commits

..

38 Commits

Author SHA1 Message Date
Joseph Doherty 45c530da6e chore(clients): regenerate Go/Java bindings and Rust vendored proto for the statuses contract comments
ci / nightly-windev (push) Has been skipped
ci / java (push) Failing after 7s
ci / windows-x86 (push) Failing after 1m8s
ci / portable (push) Successful in 7m18s
2026-08-09 12:37:35 -04:00
Joseph Doherty c867aca36b test(worker): deterministic pump-wait ordering, env hermeticity, ResolveWriteCompletionTimeout coverage 2026-08-09 12:37:35 -04:00
Joseph Doherty 436ef69f07 fix(worker): thread the completion cache through CreateForTesting
ci / nightly-windev (push) Has been skipped
ci / java (push) Failing after 1m57s
ci / windows-x86 (push) Failing after 1m57s
ci / portable (push) Failing after 8m28s
2026-08-09 12:33:20 -04:00
Joseph Doherty 2b468bd8fc docs: write-completion correlation configuration and semantics
ci / nightly-windev (push) Has been skipped
ci / windows-x86 (push) Failing after 43s
ci / java (push) Failing after 1m56s
ci / portable (push) Failing after 4m24s
2026-08-09 12:28:30 -04:00
Joseph Doherty 431a096cab feat(gateway): configurable worker write-completion wait (MxGateway:Worker:WriteCompletionWaitMilliseconds) 2026-08-09 12:27:30 -04:00
Joseph Doherty b0e2b8ba74 test(worker): write-completion correlation executor coverage 2026-08-09 12:26:23 -04:00
Joseph Doherty 66fe063410 feat(worker): bounded pump-wait correlates OnWriteComplete onto secured-write replies 2026-08-09 12:24:23 -04:00
Joseph Doherty 8de23086d0 feat(worker): share the completion cache between sink and session 2026-08-09 12:23:20 -04:00
Joseph Doherty a76ecdd59c feat(worker): event sink records OnWriteComplete rows into the completion cache 2026-08-09 12:22:54 -04:00
Joseph Doherty fc23a65cca feat(worker): versioned OnWriteComplete completion cache 2026-08-09 12:21:51 -04:00
Joseph Doherty aec95b78c9 docs(proto): document the correlated write-completion statuses contract 2026-08-09 12:20:36 -04:00
Joseph Doherty f9229ee44d docs(plan): write-completion correlation implementation plan 2026-08-09 12:20:00 -04:00
Joseph Doherty 5dbe93d13e docs(design): WriteSecured completion correlation onto the unary reply (OtOpcUa 06/S-1) 2026-08-09 12:15:07 -04:00
Joseph Doherty 129e47e541 docs(tracking): close NEXT-07, file NEXT-08/09/10, record the runner token reset
NEXT-07 is struck: windev was redeployed from origin/main (a346d51) and the service is
healthy, and the root cause the row predicted is confirmed -- the 2026-06-25 build's
Auth.ApiKeys 0.1.2.0 supports auth-DB schema 2 while the database sits at schema 3, which
is the current shared-lib version, so deploying forward was the fix rather than touching
the DB. The original text stays for the triage record.

Three findings surfaced by that work, each deliberately left for the next cycle rather
than patched in passing:

- NEXT-08: the shared GLAuth offers no TLS, so SEC-06 makes GatewayConfiguration.md's
  "deployed hosts must set Ldaps or StartTls" unsatisfiable for anything genuinely
  labelled Production. windev's relabel to Staging is honest for a dev rig but defers
  the posture question rather than answering it.
- NEXT-09: Directory.Build.props:29 quotes a path ending in a backslash, so the SHA-stamp
  git invocation is malformed on Windows and ContinueOnError stamps git's stderr into
  InformationalVersion -- a Windows binary cannot be correlated to a commit, which is what
  TST-11 exists to guarantee.
- NEXT-10: glauth.md's pre-provisioned-user table contradicts both the directory and its
  own dashboard section, and was the root cause of the NEXT-06 fixture drift. Reconciling
  it sweeps the OPC-UA group taxonomy, so it is scoped out here on purpose.

The TST-30 runner work is hygiene, not closure: runner-1 now mounts its registration token
from a 0600 file like runner-2, but both still share one instance-scope token that was
world-readable for months and is provably still live. Gitea 1.26.4 cannot rotate it from
the CLI or API, so the UI reset is recorded as a pending operator action with its
follow-through (refresh the token file, shred the token-bearing compose backups).
2026-08-07 10:31:29 -04:00
Joseph Doherty 1d6858939d docs(sec-36): record the completed windev dashboard verification
SEC-36's primary check -- dashboard /login through the real DashboardAuthenticator
search bind -- was deferred because windev's gateway was crash-looping on the stale
deployment filed as NEXT-07. That host was redeployed 2026-08-07, so the check ran:
login as multi-role returns 302 with the dashboard cookie and the authenticated page
renders the admin nav, while an anonymous control still redirects to /login. The
rotated service-account credential is now proven end-to-end on the deployed host, not
only by the equivalent ldapsearch primitive, and the runbook's Correction 3 is past
tense throughout rather than describing a fault that no longer exists.

Also record why windev runs the Staging environment name. The redeploy tripped SEC-06's
Production hard-stop on Ldap:Transport=None, and windev cannot satisfy it: it binds the
shared GLAuth, which offers no TLS, and runs Dashboard:DisableLogin=true. The Production
label contradicted its own configuration, so the host was relabelled rather than the
guard weakened -- exactly the permissive-staging-rig case the SEC-35 section already
carves out.
2026-08-07 10:31:15 -04:00
Joseph Doherty de67b45d04 test(ldap): align DashboardLdapLiveTests fixtures with the shared directory (NEXT-06)
The suite's fixtures had drifted from the shared GLAuth config, so a green run
proved nothing about the service-account bind: the only success-path test used
admin/admin123, but the directory's admin carries the standard dev password, and
the "not an admin" test used a readonly user that does not exist there at all --
it passed via the user-not-found branch rather than the group-missing branch it
names.

Realign to real users from scadaproj/infra/glauth/config.toml: admin/password
(othergroups include GwAdmin, gid 5610) for the success path, and
gw-viewer/password (GwReader only, gid 5611) for the bind-succeeds-but-no-role
path. Both are published dev credentials documented in glauth.md, not secrets.

The gw-viewer test drops its old no-leak assertion on the credential literal:
the real password is the word "password", which legitimately occurs in the
generic denial text, so the check would fail for the wrong reason. The no-leak
property is still covered with a distinctive literal by the wrong-password test.
In its place the test now asserts the property this fixture is uniquely able to
prove -- an authorization failure must be reported with the same message as an
authentication failure, so it cannot be used to enumerate valid accounts.

appsettings ships Server=localhost, so document the MxGateway__Ldap__Server
override the suite needs to reach the shared GLAuth alongside the existing
MXGATEWAY_RUN_LIVE_LDAP_TESTS and ServiceAccountPassword variables.

Verified live: Failed: 0, Passed: 5 against 10.100.0.35:3893.
2026-08-07 10:03:10 -04:00
Joseph Doherty 3d991d2160 docs(tst-30): remove mislabelled macOS instance runner (id 4)
The local act_runner on this Mac registered as instance runner id 4 with
ubuntu-latest/22.04/20.04 labels, so it competed with the two docker
runners on 10.100.0.35 for Linux jobs it had no Docker daemon to run --
13 of the last 20-run window in historiangw landed on it and all but one
failed. Registration deleted; local config kept disabled for re-use with
mac-specific labels.
2026-08-07 10:02:29 -04:00
Joseph Doherty 05667169eb docs: sync runner-topology and SEC-36 rotation prose with 2026-08-07 executed state 2026-08-07 09:28:49 -04:00
Joseph Doherty 9760497d66 docs(sec-36): rotation executed 2026-08-07; runbook host-path/vd03/verification corrections; new findings (LDAP test fixtures, windev stale deploy) 2026-08-07 09:21:54 -04:00
Joseph Doherty 5b153dac74 docs(clients): record 2026-08-07 publish of 0.2.0 client family (Java 0.2.1); cargo token needs Bearer prefix 2026-08-07 09:15:47 -04:00
Joseph Doherty 41e86481e2 docs(tst-30): second runner gitea-runner-2 live; close operator action 2026-08-07 09:12:00 -04:00
Joseph Doherty a346d514dd test(contracts): scope command-reply fixture invariants past the CLI-40/41 authenticate-user malformed-reply fixtures
ci / portable (push) Successful in 14m0s
ci / java (push) Successful in 6m50s
ci / windows-x86 (push) Failing after 1m21s
ci / nightly-windev (push) Has been skipped
The blanket loop asserted HRESULT/Statuses/ReturnValue on every command_replies fixture, but the authenticate-user.* fixtures added for the malformed-reply and credential-redaction contracts deliberately omit them (NRE on ReturnValue.DataType). Keep universal Kind/ProtocolStatus invariants for all; apply the MXAccess-detail block only to fixtures that carry it. Test-only.
2026-08-07 08:48:49 -04:00
Joseph Doherty a2d3f66b8b docs(archreview): record next-cycle candidate findings + pending operator actions surfaced during remediation
ci / windows-x86 (push) Successful in 1m19s
ci / nightly-windev (push) Has been skipped
ci / java (push) Successful in 2m19s
ci / portable (push) Failing after 4m55s
2026-08-07 08:48:03 -04:00
Joseph Doherty 93d84019b9 docs(tracking): sync IPC-23 domain register to Done (doc wave landed; Grpc.md row intentionally scoped out — DrainEvents is a worker diagnostic)
ci / nightly-windev (push) Has been skipped
ci / windows-x86 (push) Failing after 1m19s
ci / java (push) Successful in 2m4s
ci / portable (push) Failing after 4m43s
2026-08-07 08:47:31 -04:00
Joseph Doherty 6d26ed094c docs(tracking): close old-tracker CLI-24, CLI-34 as Done (2026-07-12 review old-tracker actions; both incidentally fixed)
ci / nightly-windev (push) Has been skipped
ci / java (push) Successful in 2m22s
ci / windows-x86 (push) Successful in 1m20s
ci / portable (push) Failing after 4m26s
2026-08-07 08:11:57 -04:00
Joseph Doherty 4201da63d2 docs(tracking): flip IPC-24/IPC-25 to Done in the Contracts&IPC domain register (missed by the codegen-wave tracker update)
ci / windows-x86 (push) Failing after 1m19s
ci / nightly-windev (push) Has been skipped
ci / java (push) Successful in 2m17s
ci / portable (push) Failing after 4m51s
2026-08-07 08:10:39 -04:00
Joseph Doherty 9c780f8164 Merge branch 'fix/cli-39-version-train'
ci / nightly-windev (push) Has been skipped
ci / windows-x86 (push) Failing after 1m20s
ci / java (push) Successful in 2m12s
ci / portable (push) Failing after 4m44s
# Conflicts:
#	archreview/2026-07-12/remediation/00-tracking.md
2026-08-07 08:09:39 -04:00
Joseph Doherty 440e7cf03d fix(CLI-39): bump Contracts nupkg to 0.2.0; scope pack-clients.ps1 regexes
Code review of the CLI-39 branch caught an Important gap: Contracts.csproj
was left at the already-published 0.1.2 while the .NET Client moved to
0.2.0. Invoke-PackDotnet in scripts/pack-clients.ps1 packs and publishes
both ZB.MOM.WW.MxGateway.Contracts and .Client through the same -Publish
loop, and the new collision guard runs every nupkg it finds through
Assert-GiteaPackageNotPublished. Left as-is, the next real .NET publish
would pack Contracts at 0.1.2, the guard would correctly refuse to
republish it, and the loop would abort mid-way with Client (alphabetically
first) possibly already pushed -- the two packages permanently out of
lockstep.

- src/ZB.MOM.WW.MxGateway.Contracts/ZB.MOM.WW.MxGateway.Contracts.csproj:
  <Version> 0.1.2 -> 0.2.0, matching the .NET Client (they have always
  released together).
- src/Directory.Build.props: corrected a comment that was now stale --
  it claimed the repo-wide 0.1.2 default was kept to match the Contracts
  package, which is no longer true now that Contracts.csproj overrides it.
  The <Version> value itself is unchanged; Server/Worker/Tests staying at
  0.1.2 is a separate, not-yet-made decision, out of scope for CLI-39.
- docs/ClientPackaging.md: Contracts.csproj added as a fifth manifest in
  the Versioning section, with the near-miss recorded.

Also hardened scripts/pack-clients.ps1 per the same review: the Python
(pyproject.toml) and Rust (Cargo.toml) version-extraction regexes now
scope to the [project]/[package] section header instead of matching the
first "version = ..." line anywhere in the file (Cargo.toml has an
identical second one under [workspace.package] -- matching whichever came
first was luck of ordering, not correctness). One-line comment added on
the nuget filename-parse regex.

Verified live against the real Gitea registry: Contracts and Client both
still refuse at 0.1.2 and both now pass at 0.2.0, including running the
actual Invoke-PackDotnet filename-parse-then-guard logic against two
freshly packed real .nupkg files. dotnet build of Contracts.csproj and the
client slnx both clean. No publish performed.
2026-08-07 08:07:40 -04:00
Joseph Doherty ae605d2368 Merge remote-tracking branch 'origin/fix/wrk-22-25-seam'
ci / windows-x86 (push) Failing after 1m19s
ci / nightly-windev (push) Has been skipped
ci / java (push) Successful in 2m10s
ci / portable (push) Failing after 4m27s
# Conflicts:
#	archreview/2026-07-12/remediation/00-tracking.md
#	archreview/2026-07-12/remediation/30-contracts-ipc.md
2026-08-07 08:01:10 -04:00
Joseph Doherty 9b2abef4e1 fix(CLI-39): bump client versions off published 0.1.2; guard the publish pipeline
Converges all five clients on one version after four had drifted onto the
already-published 0.1.2/0.1.1 while their APIs kept changing underneath it:

- Rust Cargo.toml [package] + [workspace.package] -> 0.2.0 (CLIENT_VERSION
  already derives from CARGO_PKG_VERSION, no separate edit).
- Python pyproject.toml + version.py -> 0.2.0; new test asserts __version__
  matches pyproject.toml (closes the CLI-26 residual drift mode).
- Go mxgateway/version.go ClientVersion -> 0.2.0.
- .NET ZB.MOM.WW.MxGateway.Client.csproj <Version> -> 0.2.0.
- Java -> 0.2.1, not 0.2.0: the live Gitea Maven feed already had 0.2.0
  published (2026-06-26), before the CLI-37/38/40/41 conformance fixes
  changed the client's observable behavior, so reusing 0.2.0 would label
  two different APIs identically. Recorded as an exception in
  docs/ClientPackaging.md's new Versioning section.

Publish-pipeline guards:

- scripts/tag-go-module.ps1 implements the CLI-21 guard: after semver
  validation it refuses to tag unless clients/go/mxgateway/version.go's
  ClientVersion already matches the requested tag version.
- scripts/pack-clients.ps1 gains a Gitea package-registry collision guard
  wired into every per-language -Publish step; it aborts if the target
  name+version already exists rather than force-overwriting. Verified live
  against the real Gitea registry (credentials already present in this
  environment) — correctly refuses on every known-published artifact and
  passes on every unpublished target.

Docs updated in the same commit: docs/ClientPackaging.md (new Versioning
section), and the five client READMEs' stale 0.1.1/0.1.2 example versions.

No .proto changes. No publish performed.
2026-08-07 07:58:49 -04:00
Joseph Doherty a55956ffa5 Merge branch 'fix/tst-30-runner-docs'
ci / java (push) Successful in 2m11s
ci / nightly-windev (push) Has been skipped
ci / windows-x86 (push) Failing after 1m38s
ci / portable (push) Failing after 17m43s
# Conflicts:
#	archreview/2026-07-12/remediation/00-tracking.md
2026-08-07 07:50:21 -04:00
Joseph Doherty aba22358f5 Merge branch 'fix/ipc-27-descriptor-test'
ci / nightly-windev (push) Has been skipped
ci / windows-x86 (push) Successful in 1m14s
ci / java (push) Successful in 2m8s
ci / portable (push) Failing after 4m45s
2026-08-07 07:49:33 -04:00
Joseph Doherty b604fed72b docs(TST-30): document shared-runner CI bottleneck + second-runner operator runbook
Doc half of TST-30 (single shared Gitea runner is a CI throughput/availability
bottleneck): docs/GatewayTesting.md's Continuous Integration section gains a
"Runner capacity is shared and finite" subsection covering the maxParallel=1
instance-level runner shared with dohertj2/lmxopcua, the ~20-30 min queue
latency observed under cross-repo contention, and Gitea 1.26's missing run
cancel/delete API. The existing "windev tier down" degraded-mode paragraph now
also covers "runner contended" as a reason to bypass the queue via
CI_SHA=<sha> scripts/ci/run-windev-ci.sh <mode> or the manual windev worktree
flow, generalizing it per the finding's design note.

New operator runbook docs/runbooks/TST-30-second-ci-runner.md carries the
actual runner registration (option a: second act_runner instance on
10.100.0.35 with the same container.network: traefik config, recommended;
option b: dedicated labelled runner, escalation only; option c: runner on
windev, rejected) plus verification steps and the no-cancel caveat. The
optional workflow-level concurrency group is documented as unverified --
framed as "verify before relying on it" -- and left unimplemented in ci.yml,
since registering the runner and any runs-on gating is operator/infra work
outside this repo's tree.

Tracking: TST-30 -> Done (doc half; runner registration operator-pending) in
both registers + change-log row.
2026-08-07 07:47:37 -04:00
Joseph Doherty 6060d21995 fix(IPC-27): close descriptor-freshness blind spots for enums, services, and Galaxy
ClientProtoInputTests.Descriptor_ContainsEveryContractMessageAndField only
compared messages and fields, and only enumerated the gateway/worker
descriptors, so a new enum value, a new RPC, or any galaxy_repository.proto-only
change would not redden the test even though it is documented as the primary
protoc-free CI gate.

Rename to Descriptor_ContainsEveryContractSymbol and extend the reflection walk
on both sides (published protoset and in-process contract) to also collect
enums/enum values ({enumFullName}, {enumFullName}/{valueName}) and
services/methods ({serviceFullName}, {serviceFullName}/{methodName}), and add
GalaxyRepositoryReflection.Descriptor to the enumerated files. The comparison
stays a flat, order-insensitive string-set diff with no protoc dependency.

Update docs/ClientProtoGeneration.md and docs/Contracts.md prose from
"message or field" to the full symbol coverage.

Red-path proof: pointed the test at the pre-IPC-01 stale protoset and confirmed
it failed naming max_frame_bytes, several MxCommandKind/AlarmProviderMode enum
values, MxAccessGateway/StreamAlarms and GalaxyRepository/BrowseChildren, and
the galaxy_repository.v1.* surface; restored the real path and re-ran green.

Flips IPC-27 to Done in the 2026-07-12 remediation tracker and register.
2026-08-07 07:47:18 -04:00
Joseph Doherty 1c2f3a62c1 Merge branch 'fix/sec-36-ldap-secret'
ci / windows-x86 (push) Successful in 1m16s
ci / nightly-windev (push) Has been skipped
ci / java (push) Successful in 2m31s
ci / portable (push) Failing after 4m37s
2026-08-07 07:44:43 -04:00
Joseph Doherty 0646c73e48 Merge branch 'fix/ipc-24-25-codegen'
ci / nightly-windev (push) Has been skipped
ci / windows-x86 (push) Successful in 1m15s
ci / java (push) Successful in 2m5s
ci / portable (push) Failing after 4m32s
# Conflicts:
#	archreview/2026-07-12/remediation/00-tracking.md
2026-08-07 07:42:35 -04:00
Joseph Doherty eacdd2d453 fix(IPC-23,IPC-24,IPC-25,IPC-32): proto-comment regen wave + codegen-freshness guards
Proto comments (comment-only, no wire change):
- mxaccess_worker.proto GatewayHello.max_frame_bytes: every worker->gateway frame
  must serialize within the negotiated max; reply builders truncate (IPC-23).
- mxaccess_gateway.proto DrainEventsReply: count-cap + byte-cap, drain-until-empty
  caller contract (IPC-23).
- mxaccess_gateway.proto ReplayGap.oldest_available_sequence: empty-ring value is
  highest-observed+1, oldest-1 resume formula stays valid (GWC-25 deferred amendment).

Regen wave: Contracts/Generated (C# XML doc), rust vendored protos (byte-copy),
Go bindings (worker binding was genuinely stale - lacked MaxFrameBytes entirely),
Python worker _pb2 (real descriptor delta), Java aggregates (javadoc, zero
protobuf-version churn under the pinned toolchain), client descriptor set.

IPC-24: pinned Java toolchain regenerates with no gencode-version churn, so the
unconditional churn-revert step in ci.yml is a fossil - deleted it; git diff is
now a true message-level drift gate for the single-file Java aggregates.

IPC-25: pin protoc-gen-go v1.36.11 / protoc-gen-go-grpc 1.6.2 in the Go generate
script (+ fix a latent pwsh-7 parse bug); add Check 4 to check-codegen.ps1
(regenerate Go+Python bindings, fail on diff, tool-missing fails not skips); add
the pinned-generator installs to the portable CI job.

IPC-32: relabel check-codegen banners 1/4..4/4 (folded into the Check 4 edit).

Docs: ClientProtoGeneration.md, Contracts.md, GatewayTesting.md, build.gradle
checkGeneratedClean caveat. Tracking: IPC-23/24/25/32 -> Done, GWC-25 proto note
resolved, change-log 2026-08-07.
2026-08-07 07:41:18 -04:00
Joseph Doherty 8c312c717c fix(SEC-36): scrub committed dev LDAP service-account password; add user-secrets channel + rotation runbook
Repo-side half of SEC-36. The appsettings.json plaintext was already discharged
before this branch (HEAD ships the fail-closed ${secret:ldap/mxgateway/bind}
store reference), so the residual leak was the literal value in glauth.md,
docs/GatewayTesting.md, and the historical archreview SEC-06 evidence -- all
scrubbed to <service-account-password> placeholders pointing at the source of
truth scadaproj/infra/glauth/.

- csproj: add <UserSecretsId>mxaccessgw-server</UserSecretsId> (dev channel)
- GatewayOptionsValidator: blank-password message now names both channels
  (dev user-secrets, deployed MxGateway__Ldap__ServiceAccountPassword)
- test: assert the message names both channels
- docs: GatewayConfiguration.md (three channels + rotation note), glauth.md
  (placeholders + rotation-required + runbook pointer), GatewayTesting.md
- new operator runbook docs/runbooks/SEC-36-ldap-credential-rotation.md
  (live rotation + NSSM staging remain operator-pending)
- tracking: SEC-36 -> Done (repo-side) in both registers + change-log

Deviation: kept the ${secret:} reference in appsettings.json rather than
deleting it (spec step 2 assumed the stale plaintext baseline); deleting it
would regress the shipped/documented/tested secret-store channel.

git grep -i for the old value is empty across all tracked files.
2026-08-07 07:40:41 -04:00
77 changed files with 3594 additions and 200 deletions
+22 -11
View File
@@ -60,7 +60,19 @@ jobs:
dotnet tool install --global PowerShell
echo "$HOME/.dotnet/tools" >> "$GITHUB_PATH"
# IPC-01 / IPC-19 / IPC-20: descriptor set + Contracts/Generated must match the current protos.
# IPC-25 Check 4 regenerates the Go and Python client bindings and diffs them, so the pinned
# generators must be present. protoc 34.1 is already installed above; Go and Python are set up
# above. Pin protoc-gen-go / protoc-gen-go-grpc to match the committed header stamps and grpcio
# -tools to match the committed _pb2 stamp, or Check 4 false-fails (or masks drift) under churn.
- name: Install pinned client codegen generators (Check 4)
run: |
go install google.golang.org/protobuf/cmd/protoc-gen-go@v1.36.11
go install google.golang.org/grpc/cmd/protoc-gen-go-grpc@v1.6.2
echo "$(go env GOPATH)/bin" >> "$GITHUB_PATH"
python -m pip install 'grpcio-tools==1.80.0'
# IPC-01 / IPC-19 / IPC-20 / IPC-25: descriptor set + Contracts/Generated + Go/Python bindings
# must match the current protos.
- name: Codegen / descriptor freshness
shell: pwsh
run: ./scripts/check-codegen.ps1
@@ -93,10 +105,12 @@ jobs:
python -m pytest
java:
# Java client runs on a JDK-17 Linux runner (the macOS dev box has no JRE). The protobuf gradle
# plugin rewrites MxaccessGateway.java with spurious protobuf-runtime-version churn on every
# build; when no .proto changed, revert that one file so checkGeneratedClean / a dirty tree does
# not fail the build (repo memory project_java_generated_churn).
# Java client runs on a JDK-17 Linux runner (the macOS dev box has no JRE). The grpc/protobuf
# toolchain is fully pinned (clients/java/build.gradle: grpcVersion 1.76.0 / protobufVersion
# 4.33.1), so a regeneration is byte-identical to the committed aggregates modulo real .proto
# changes — `Verify generated tree is clean` (git diff) is the true drift gate (IPC-24). The
# single-file Java aggregates are where message-level proto drift lands, so this job now catches
# a .proto edited without regenerating and committing the Java client.
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
@@ -115,13 +129,10 @@ jobs:
- name: Gradle test
working-directory: clients/java
run: gradle test
- name: Revert spurious protobuf-version churn (no .proto changed)
# Both generated aggregates can pick up protobuf-runtime-version churn on regen; revert
# both so verify-clean still catches a real, uncommitted proto/codegen change elsewhere.
run: |
git checkout -- clients/java/src/main/generated/main/java/mxaccess_gateway/v1/MxaccessGateway.java || true
git checkout -- clients/java/src/main/generated/main/java/mxaccess_worker/v1/MxaccessWorker.java || true
- name: Verify generated tree is clean
# IPC-24: the pinned grpc/protobuf toolchain regenerates byte-identical output, so this
# git-diff gate now catches message-level proto drift in the single-file Java aggregates
# (the old unconditional churn-revert step masked exactly that class and was deleted).
run: git diff --exit-code -- clients/java/src/main/generated
windows-x86:
File diff suppressed because one or more lines are too long
@@ -61,7 +61,7 @@ Coordinate with (do not block on) open GWC-21: if `EventChannelFullModeTimeout`
The proto comment currently states "`oldest_available_sequence` itself IS still retained", which becomes false in the empty-ring case — per the docs-with-source rule, amend the field comment in the same commit to define the empty-ring value ("when nothing is retained, this is the next sequence that can be delivered — `highest observed + 1` — and the `oldest 1` resume formula remains valid; the interval evicted is unchanged"). This is a comment-only proto change (no descriptor delta), but the repo's codegen rules still apply — see the steps.
**Implementation.**
**Implementation.** (Code + `docs/Sessions.md` landed 2026-08-07 on `fix/gwc-25-replaygap-trio`; the deferred proto-comment amendment below **landed 2026-08-07** with the IPC-23 codegen wave on `fix/ipc-24-25-codegen` — GWC-25 is fully resolved.)
- `Sessions/SessionEventDistributor.cs:463-467`: replace `oldestAvailableSequence = 0;` with `oldestAvailableSequence = gap ? _highestSequenceSeen + 1 : 0;` plus a comment explaining the `oldest 1` client formula this must keep valid (cite this finding).
- `src/ZB.MOM.WW.MxGateway.Contracts/Protos/mxaccess_gateway.proto` (`ReplayGap.oldest_available_sequence`, ~line 759): append the empty-ring sentence above. Then regenerate per repo rules: delete `src/ZB.MOM.WW.MxGateway.Contracts/Generated/*.cs`, `dotnet build src/ZB.MOM.WW.MxGateway.Contracts/ZB.MOM.WW.MxGateway.Contracts.csproj`, and **commit `Generated/`** (net48 worker builds break otherwise). Sync the vendored client copies of the proto byte-identical (`clients/*/`); a comment-only edit changes no descriptor, so: Python `*_pb2*` output is unchanged (comments are not embedded — regenerate with the pinned grpcio-tools only if the files actually differ), Go/C#/Rust generated doc comments will churn — regenerate those per each client README, and revert spurious Java aggregate-file churn if no message-level delta appears (per the established Java convention).
- `docs/Sessions.md` (~lines 228-234, ReplayGap section): document the empty-ring sentinel value and that `after_worker_sequence = oldest_available_sequence 1` is the universal resume formula in both the retained and fully-evicted cases.
@@ -12,16 +12,16 @@ All `path:line` citations were re-verified against the working tree at `4f5371f`
| ID | Sev | Tier | Eff | Dep | Status | Title |
|----|-----|------|-----|-----|--------|-------|
| IPC-23 | Medium | P0 | S¹ | WRK-21 | In progress — mechanics landed with WRK-21; proto-comment/doc wave pending | DrainEvents bound is count-based only; byte-heavy queue still builds a session-killing reply frame (contract requirements here; fix mechanics in WRK-21) |
| IPC-24 | Medium | P0 | S | — | Not started | CI's unconditional Java churn-revert masks real generated-code drift for message-level proto changes |
| IPC-25 | Medium | P0 | M | — | Not started | Committed Go/Python worker bindings are stale at HEAD; no guard covers them |
| IPC-23 | Medium | P0 | S¹ | WRK-21 | Done | DrainEvents bound is count-based only; byte-heavy queue still builds a session-killing reply frame (contract requirements here; fix mechanics in WRK-21) |
| IPC-24 | Medium | P0 | S | — | Done | CI's unconditional Java churn-revert masks real generated-code drift for message-level proto changes |
| IPC-25 | Medium | P0 | M | — | Done | Committed Go/Python worker bindings are stale at HEAD; no guard covers them |
| IPC-26 | Low | P2 | S¹ | WRK-22 | Done (mechanics landed in WRK-22) | Cancelled write leaves a ghost frame that is still written (contract requirement here; fix mechanics in WRK-22) |
| IPC-27 | Low | P2 | S | — | Not started | Descriptor freshness test blind to enums, enum values, services/methods, and the Galaxy contract |
| IPC-27 | Low | P2 | S | — | Done | Descriptor freshness test blind to enums, enum values, services/methods, and the Galaxy contract |
| IPC-28 | Low | — | S | — | Done | `docs/Grpc.md` omits the `CommandTooLarge``ResourceExhausted` mapping |
| IPC-29 | Low | — | S | — | Done (discharged by WRK-26) | Worker writer priority scheduling and write-time sequence stamping undocumented in the frame-protocol doc |
| IPC-30 | Low | P0 | M | WRK-21 (same file/batch) | Done | Oversized worker→gateway event frame is session-fatal — make the death deliberate, structured, and diagnosable |
| IPC-31 | Info | — | — | — | N/A | Gateway stamps sequence at creation, worker at write — accepted divergence; sequence is documented diagnostic-only (`gateway.md:328-330`); revisit only if sequence ever becomes load-bearing |
| IPC-32 | Info | — | S | IPC-25 | Not started | `check-codegen.ps1` check labels miscounted (folded into the IPC-25 script edit) |
| IPC-32 | Info | — | S | IPC-25 | Done | `check-codegen.ps1` check labels miscounted (folded into the IPC-25 script edit) |
¹ Effort for the work owned by *this* plan (proto comments + docs + acceptance criteria). The code mechanics are M and are tracked under WRK-21 / WRK-22 in the worker plan.
@@ -15,7 +15,7 @@ Repo rules that bind every entry: docs change in the same commit as the source (
| SEC-33 | Low | P1 | M | — (co-locate SEC-23) | Done | Any-platform path-rooting acceptance re-opens SEC-01 on Unix; Galaxy `SnapshotCachePath` unvalidated |
| SEC-34 | Low | P2 | S | — | Done | Verification cache: expiry outlives TTL; `Invalidate` races in-flight repopulation |
| SEC-35 | Info | — | S | — | N/A (doc-only note discharged 2026-08-07) | Production hard-stops key on the exact `Production` environment name |
| SEC-36 | Low | P1 | M | cross-repo (`scadaproj/infra/glauth`) | Not started | Committed dev LDAP service-account password: remove from repo and rotate |
| SEC-36 | Low | P1 | M | cross-repo (`scadaproj/infra/glauth`) | Done | Committed dev LDAP service-account password: remove from repo and rotate |
---
@@ -218,3 +218,9 @@ dotnet test src/ZB.MOM.WW.MxGateway.Tests/ZB.MOM.WW.MxGateway.Tests.csproj --fil
dotnet test src/ZB.MOM.WW.MxGateway.Tests/ZB.MOM.WW.MxGateway.Tests.csproj --filter "FullyQualifiedName~GatewayOptionsValidator"
```
(asserts the blank-password validation still fires with the updated message). Manual: with user-secrets set on the dev box, `dotnet run --project src/ZB.MOM.WW.MxGateway.Server/...` and a dashboard `/login` as `multi-role` succeeds against the rotated GLAuth; the deployed-host login re-check from step 1 counts as the production verification. Live-LDAP integration tests (`MXGATEWAY_RUN_LIVE_LDAP_TESTS=1`) only where the GLAuth instance is reachable; otherwise document skipped per the testing matrix.
**Outcome (2026-08-07 — Done, repo-side; live rotation operator-pending).** Landed on `fix/sec-36-ldap-secret`. **The design's baseline had already shifted:** at HEAD `appsettings.json` no longer commits the literal — it ships `"ServiceAccountPassword": "${secret:ldap/mxgateway/bind}"`, a fail-closed encrypted-store reference (documented `GatewayConfiguration.md:252`, tested by `PreHostSecretExpansionTests`) introduced by the Secrets-store adoption after this remediation was written. **Deviation from Implementation step 2:** the `${secret:}` reference was **kept, not deleted** — deleting it regresses the shipped/documented/tested store channel and the committed-plaintext finding is already resolved for `appsettings.json`. The load-bearing residual — the literal value still present in `glauth.md`'s samples (`:33,65,103,136,245`), `docs/GatewayTesting.md`, and the historical `archreview/*` SEC-06 evidence — was scrubbed to `<service-account-password>` placeholders, each with a pointer to the source of truth `scadaproj/infra/glauth/` and a rotation-required note. Steps 36 implemented as designed: `<UserSecretsId>mxaccessgw-server</UserSecretsId>` added (step 3); the `ValidateLdap` blank-password message now names both channels — dev `dotnet user-secrets set "MxGateway:Ldap:ServiceAccountPassword" <value>` and deployed `MxGateway__Ldap__ServiceAccountPassword` — plus a note on the `${secret:}` store default (step 4), asserted by the extended `Validate_Fails_WhenLdapEnabledAndServiceAccountPasswordBlank`; docs updated same commit (step 5); `git grep -i` for the old value is empty across tracked files (step 6). The cross-repo **step 1 (rotate GLAuth on `10.100.0.35`, pre-stage the NSSM env var on `10.100.0.48` and on `wonder-app-vd03` only if `Ldap.Enabled`, verify dashboard login)** is the operator's to execute, captured in the new runbook `docs/runbooks/SEC-36-ldap-credential-rotation.md`. Verification (macOS): `dotnet build …Server` 0 warnings/0 errors; `dotnet test --filter ~GatewayOptionsValidator` green.
**Outcome (2026-08-07 — operator half executed; finding now fully `Done`).** The cross-repo step 1 left open above was executed per `docs/runbooks/SEC-36-ldap-credential-rotation.md`. A new service-account password was generated, the `serviceaccount` `passsha256` in `scadaproj/infra/glauth/config.toml` replaced, and the shared GLAuth recreated on `10.100.0.35` from its actual compose directory — the **load-bearing** half of the finding is now discharged: the value disclosed by this repo's git history (live in the directory since 2026-06-04; not reproduced here) no longer binds `dc=zb,dc=local`. The new value exists only in the three channels the design named — the GLAuth `passsha256` (committed in `scadaproj`, commit `aada53b`), the NSSM service environment on `10.100.0.48`, and each dev box's user-secrets — and in no file of this repo. The retired plaintext was additionally scrubbed from the `scadaproj` glauth comments (`config.toml`, `docker-compose.yml`, `README.md`) and from the docker host's live `docker-compose.yml`; the host's `*.bak-sec36` rollback copies deliberately retain it. **Three runbook facts were wrong and are corrected in a dated block at its top.** (1) Its step 3 said `cd ~/Desktop/scadaproj/infra/glauth` on the docker host; no such path exists there — the stack runs from `/home/dohertj2/zb-glauth` (container `zb-shared-glauth`), fed by the `scp` deploy documented in `scadaproj/infra/glauth/README.md`. (2) `wonder-app-vd03` is **out of scope on documentary evidence**, not merely unchecked: its gateway binds the ScadaBridge/ScadaLink local GLAuth under `dc=scadalink`/`dc=scadabridge`, a different directory that never held this credential (the host is also unreachable from the dev network); no env var was staged there. (3) Its "3-fail / 10-minute per-IP lockout" caution is **inert for this instance**`config.toml:14` sets `LimitFailedBinds = false`. **One Done criterion is met with a caveat:** the new value **is** staged on windev (`10.100.0.48`, 10th `AppEnvironmentExtra` entry on the `MxAccessGw` NSSM service), but the runbook's primary check — dashboard `/login` as `multi-role`**could not run**, because windev's gateway is crash-looping on an unrelated pre-existing fault: the deployed Server binary (2026-06-25) predates the 2026-07-15 auth-DB migration, so it opens a schema-version-3 database it supports only at version 2 and aborts at startup (~10k Hosting-failed events/day since at least 08-06). That is a stale-deployment problem, filed as a next-cycle candidate finding, not a rotation defect. **Verified instead by the equivalent primitive:** a direct `ldapsearch` bind as `cn=serviceaccount,dc=zb,dc=local` with the new value against `10.100.0.35:3893` succeeded and returned the `multi-role` entry — the same search bind the dashboard performs. Also surfaced and filed for next cycle: `DashboardLdapLiveTests` fixture drift leaves the suite with **no positive-proof coverage** of the service-account bind, so it could not have substituted for the dashboard check either. Tracking: both registers' SEC-36 rows, the pending-operator-actions list in `90-candidate-findings-next-cycle.md`, and the `00-tracking.md` progress log.
**Addendum (2026-08-07, later the same day — the caveat is closed).** windev was repaired under NEXT-07 (fresh publish of `origin/main` `a346d51`), and the deferred dashboard check then ran on that host: with `Dashboard:DisableLogin=false` supplied as a process-env-only override on a foreground run, `GET /login` returned 200 with an antiforgery token, `POST /auth/login` as `multi-role`/`password` returned 302 to `/` with a `MxGatewayDashboard` cookie, the authenticated `GET /` rendered the admin nav, and an anonymous control redirected to `/login?ReturnUrl=%2F`. The rotated credential is therefore proven through the real `DashboardAuthenticator` search-bind path on the deployed host, not only by the `ldapsearch` primitive. As deployed windev keeps `DisableLogin=true`, so routine operation there does not exercise LDAP; the standing regression proof is the realigned `DashboardLdapLiveTests` (NEXT-06, commit `de67b45`), 5/5 green against the shared GLAuth. `docs/runbooks/SEC-36-ldap-credential-rotation.md` Correction 3 carries the same record.
@@ -20,7 +20,7 @@ Operating constraints carried from prior work:
| CLI-36 | Medium | P0 | S | — | Done | Go CLI `stream-events` silently destroys the ReplayGap signal |
| CLI-37 | Medium | P1 | M | CLI-38 | Done | Status-array validation must branch on `category` per the proto contract (4-vs-1 divergence) |
| CLI-38 | Medium | P1 | S | — | Done | Align .NET/Go/Java on `hresult < 0` — lands prior CLI-08 and cures the design-doc drift |
| CLI-39 | Medium | P1 | S | CLI-35..38, CLI-45 | Not started | Bump client versions off the already-published 0.1.2 before the next publish; add registry-collision guard |
| CLI-39 | Medium | P1 | S | CLI-35..38, CLI-45 | Done | Bump client versions off the already-published 0.1.2 before the next publish; add registry-collision guard |
| CLI-40 | Low | — | M | — | Done | Port the exact-secret credential scrub to Rust/Java/.NET |
| CLI-41 | Low | — | M | — | Done | Uniform malformed-reply contract for AuthenticateUser/ArchestrAUserToId/AddBufferedItem |
| CLI-42 | Low | P1 | S | — | Done | Document the vendored Rust proto layout (CLI-02's missing doc half) |
@@ -15,7 +15,7 @@ Prior-cycle open findings (TST-05..24 where still open) are tracked in the prior
| TST-27 | Medium | P1 (doc batch) | S | — | Done | `ShowTagValues` config row still says "Reserved" after SEC-25 made the flag live |
| TST-28 | Low | P2 | S | relates IPC-02 | Done | Gateway-side `max_frame_bytes` handshake field untested in the CI-run suite |
| TST-29 | Low | P2 | S | — | Done | Retire `oldtasks.md` after folding the Phase-5 governance record into DesignDecisions.md; delete root docs-review artifacts |
| TST-30 | Low | P2 | M | — | Not started | Single shared Gitea runner is a CI throughput/availability bottleneck (cross-repo contention, no run cancel/delete) |
| TST-30 | Low | P2 | M | — | Done | Single shared Gitea runner is a CI throughput/availability bottleneck (cross-repo contention, no run cancel/delete) |
---
@@ -167,6 +167,10 @@ Independent of the runner count, document the **no-cancel** reality (Gitea 1.26
**Verification.** Push two branches back-to-back and confirm their runs execute concurrently (not serially) once a second runner exists; `GET /repos/dohertj2/mxaccessgw/actions/runners` (or the instance runner list) shows ≥2 runners online; `docs/GatewayTesting.md` describes the shared-runner/no-cancel reality and the bypass. Re-run the TST-25 acceptance push and confirm queue depth is materially lower under a concurrent `lmxopcua` run.
**Outcome (2026-08-07 — Done, doc half; runner registration operator-pending).** Landed on `fix/tst-30-runner-docs`. Implementation step 2 shipped: `docs/GatewayTesting.md`'s Continuous Integration section gained a "Runner capacity is shared and finite" subsection stating the `maxParallel=1` co-located runner is shared with `dohertj2/lmxopcua` at the instance level (not repo-scoped), the ~2030 minute queue latency observed under cross-repo contention, and the Gitea 1.26 no-cancel/no-delete API reality; the existing "windev tier down" degraded-mode paragraph now also covers "runner contended" as a reason to use the bypass, generalized per this finding's design note. New operator runbook `docs/runbooks/TST-30-second-ci-runner.md` carries **step 1** (register a second `act_runner` on `10.100.0.35`, option (a) recommended, same `container.network: traefik` config; option (b) dedicated labelled runner as an escalation; option (c) windev-hosted runner rejected) with the verification checklist (concurrent back-to-back pushes, `GET /repos/dohertj2/mxaccessgw/actions/runners` ≥ 2) and a note that the no-cancel reality persists regardless of runner count. **Step 3 (optional workflow-level `concurrency` group)** is documented in the runbook as unverified — explicitly framed as "verify this Gitea deployment honors it before relying on it" — and left unimplemented in `ci.yml`, since it is a `ci.yml` change out of scope for this doc-only pass. **The actual runner registration (step 1) is infrastructure work outside this repo's tree and remains the operator's to execute**, tracked in the runbook. Verification performed: `grep -n 'maxParallel\|shared\|cancel' docs/GatewayTesting.md` shows the new prose; runbook file exists at the path above; no build required (doc-only change).
**Outcome (2026-08-07 — operator half executed; finding now fully `Done`).** The runner registration left open above was executed per `docs/runbooks/TST-30-second-ci-runner.md` option (a). A second instance-level `act_runner` container, `gitea-runner-2` (runner id 5, capacity 2, labels `ubuntu-latest`/`ubuntu-22.04`), now runs on `10.100.0.35` from the `/opt/gitea` compose stack with the same `container.network: traefik` setting as the original; its registration token is mounted from a `0600` file rather than inlined in compose. The existing `gitea-runner` (id 1, capacity 4) was **not** modified — capacity went 4 → 6 by addition, so the change is reversible by removing one container. Concurrency verified live by pushing HEAD (`a346d51`) to two scratch branches, `scratch/tst30-a` (run 661) and `scratch/tst30-b` (run 662), while an unrelated run (660) was already in flight: at 13:07:53Z jobs from **three** runs were `in_progress` simultaneously — run 660 `portable` and run 662 `portable`/`java` on runner 1, run 661 `portable`/`java` on `gitea-runner-2` — which the pre-change single-runner topology could not have produced. `gitea:3000` resolution holds on the new instance: run 661's `portable` job (task 1145, scheduled on `gitea-runner-2`) logged `git remote add origin http://gitea:3000/dohertj2/mxaccessgw` followed by a successful `fetch … From http://gitea:3000/dohertj2/mxaccessgw`, and its job container's workspace was confirmed checked out at `a346d514dd24e775640e5667aa7cd8e561fec68a`. **Runbook correction:** its verification checklist said `GET /repos/dohertj2/mxaccessgw/actions/runners` should show ≥2 — that endpoint still returns `total_count: 0` because both runners are registered at the **instance** level, exactly as this finding documented; the correct check is `GET /api/v1/admin/actions/runners`, which lists ids 1, 4 (an unrelated local macOS runner), and 5. Recorded as a dated "Executed" note at the top of the runbook. The no-cancel reality is unchanged and the `run-windev-ci.sh` bypass remains valid, so `docs/GatewayTesting.md`'s prose needed no edit.
---
## Cross-domain dependencies
@@ -0,0 +1,25 @@
# Candidate Findings for the Next Review Cycle (surfaced during 2026-07-12 remediation)
These were discovered while remediating the 2026-07-12 backlog but were **out of scope** for it — each is either pre-existing, by-design residual, or a new observation. They are recorded here so the next review cycle can triage them. None blocks the 2026-07-12 cycle, which is complete. Rows struck through have since been fixed ahead of that cycle; the original finding text is kept so the triage record stays readable.
| ID (proposed) | Area | Severity (est.) | Summary |
|---|---|---|---|
| NEXT-01 | Testing / macOS | Low | Fake-worker/e2e gateway tests fail on macOS under the default `TMPDIR` because the `CoreFxPipe_mxaccess-gateway-{pid}-{sessionId}` path exceeds the 104-char Unix-domain-socket `sun_path` limit under `/var/folders/…/T/`. Workaround today is `TMPDIR=/tmp`. Fix options: shorten the pipe name, or document the `TMPDIR=/tmp` requirement in `docs/GatewayTesting.md`. Surfaced independently by multiple remediation agents. |
| NEXT-02 | Clients (.NET, Java) | Low | The .NET and Java CLIs render the raw `ReplayGap` sentinel `MxEvent` on `stream-events` instead of a typed gap row — Java text mode prints `0 MX_EVENT_FAMILY_UNSPECIFIED`. Same defect class as CLI-36 (Go) / CLI-35 (Python), which were fixed this cycle; the .NET/Java halves were out of scope. The cross-language smoke matrix now records this divergence honestly. |
| NEXT-03 | Gateway alarms | Low | `GatewayAlarmMonitor.ApplyReconcile` feed-repair broadcasts (the new acked-delta from GWC-26 **and** the pre-existing Raise/Clear repair) are **at-least-once, not exactly-once**: a periodic reconcile can synthesize a transition whose matching live transition is still buffered in the alarm lease, so both broadcast as indistinguishable duplicates on the alarm feed (StreamAlarms + dashboard hub). Pre-existing (the Raise/Clear repair always had it); GWC-26 documented the at-least-once contract rather than closing the race. Closing it needs reconcile/live serialization or a monotonic dedup marker. |
| NEXT-04 | Worker frame writer | Low | WRK-22/WRK-25 cancellation path: a frame `Claimed` by a concurrent lock-holder just before its caller's cancellation races in is never awaited by that caller; if the write then faults, `TrySetException` lands on a `Task` nobody observes (unobserved-task-exception). By-design residual, non-crash (no `UnobservedTaskException` handler registered), pre-existing to single-frame WRK-22 and amplified per-batch by WRK-25. Hygiene fix: attach a fault-observing continuation to abandoned/tombstoned frame completions. |
| NEXT-05 | Worker frame writer | Info | A batch whose remaining frames are tombstoned by cancellation leaves dead `PendingFrame` entries in `_eventFrames`/`_controlFrames` until a future `DequeueNext` pops and skips them. Same pre-existing behavior as single-frame WRK-22, amplified per-batch; in practice heartbeats purge them promptly, so not a real leak. |
| ~~NEXT-06~~ | Testing / live LDAP | Medium | **Resolved 2026-08-07** — fixtures realigned to the shared directory (`admin`/`password` for the GwAdmin success path, `gw-viewer`/`password` for the bind-succeeds-but-no-role path); verified `Failed: 0, Passed: 5` live against the shared GLAuth at `10.100.0.35:3893`, so the success-path assertion (GwAdmin group claim + Admin role claim) now fails if the service-account credential is wrong. Original finding: `DashboardLdapLiveTests` fixtures have drifted from the shared GLAuth directory, leaving the suite with **no positive-proof coverage of the service-account bind**. Its only success-path test, `AuthenticateAsync_AdminInGwAdminGroup_Succeeds`, binds `admin`/`admin123`, but the directory's `admin` user carries the standard dev password (`scadaproj/infra/glauth/config.toml`), so that assertion cannot pass. `AuthenticateAsync_ReadOnlyUserMissingGwAdminGroup_Fails` binds fixture user `readonly`, which **does not exist** in the GLAuth config at all — it passes for the wrong reason (user-not-found rather than the group-missing branch it names; the `readonly` name is in fact barred by the README's user/group case-collision rule). The three remaining tests are negative assertions that pass whether or not the service account can bind. Net effect: a green `DashboardLdapLiveTests` run proves nothing about the bind credential — surfaced during SEC-36, where the suite was considered as a substitute for the deferred dashboard-login check and rejected. Fix: realign the fixtures to real directory users (e.g. `multi-role`/`gw-viewer`) or add the missing users to the GLAuth config, and add one test that fails when the service-account credential is wrong. |
| ~~NEXT-07~~ | Deployment / windev | High | **Resolved 2026-08-07** — a fresh portable framework-dependent publish of `origin/main` (`a346d51`) was built in a clean clone at `C:\build\mxgw-redeploy`, deployed to `C:\publish\mxaccessgw\Server-20260807`, and the `MxAccessGw` NSSM service repointed at it; the service now holds a stable PID with both `5120`/`5130` listening, a worker spawned, the Galaxy snapshot restored (129 objects / 56,731 attributes) and a clean event log. Root cause confirmed as the version skew this row predicted: the deployed 2026-06-25 build carried `ZB.MOM.WW.Auth.ApiKeys` 0.1.2.0, which supports auth-DB schema 2, against a `gateway-auth.db` stamped at schema 3 on 2026-07-15 by an ephemeral run of newer code — schema 3 is the current shared-lib version (`SqliteAuthSchema.CurrentVersion=3` in Auth 0.1.5), so the redeploy is the forward fix and the DB was left alone. Rollback artifacts kept: `C:\ProgramData\MxGateway\gateway-auth.db.bak-next07` (with `-wal`/`-shm`) and the previous `C:\publish\mxaccessgw\Server` directory. Two side effects worth recording: the old deploy's `appsettings.json` held the LDAP bind password in **plaintext on disk**, while the new one keeps the repo's `${secret:ldap/mxgateway/bind}` token with the NSSM environment supplying the value, so no plaintext LDAP secret remains on that host; and the redeploy tripped the SEC-06 `Ldap:Transport=None` production hard-stop (`GatewayOptionsValidator.cs:178`), resolved by relabelling the host — windev runs `Dashboard:DisableLogin=true`, which this repo's own docs mark dev/test-only, so its `Production` label contradicted its configuration and `DOTNET_ENVIRONMENT` was changed to `Staging` (that one NSSM environment entry only; the other nine preserved byte-identical). SEC-06 is untouched for genuinely production hosts — see NEXT-08 for the posture problem that relabelling defers. Original finding: The `10.100.0.48` (windev) gateway deployment is **stale and crash-looping**, and has been since at least 2026-08-06 (~10k Hosting-failed events/day). The deployed Server binary dates to 2026-06-25 and predates the auth-DB migration of 2026-07-15: it opens a schema-version-3 `gateway-auth.db` that it supports only at version 2 and aborts at startup, so the `MxAccessGw` service never reaches a listening state. Not a code defect in the current tree — a deploy-drift/operations gap — but it means the repo's only deployed host has been dark for over a day and any host-level verification (including SEC-36's dashboard-login check) is blocked until it is repaired. Fix: deploy a current Server build to windev, or restore/downgrade the auth DB to schema 2 if the old binary must stand. Worth asking separately why a service in a permanent restart loop raised no alert. Discovered during SEC-36. |
| NEXT-08 | Security / LDAP posture | Medium | **The shared GLAuth offers no TLS, so SEC-06 makes it undeployable from a `Production`-labelled host.** `GatewayOptionsValidator` (`src/ZB.MOM.WW.MxGateway.Server/Configuration/GatewayOptionsValidator.cs:178`) refuses to start when `Ldap:Transport=None` in the `Production` environment, and `docs/GatewayConfiguration.md`'s `Transport` row states "Deployed hosts must set `Ldaps` or `StartTls`" — but the shared instance at `10.100.0.35:3893` has `[ldaps] enabled=false`, port `3894` closed, and answers StartTLS with `protocolError`, so neither value can work against it. That instruction is currently unsatisfiable for every host that authenticates there. windev sidestepped it on 2026-08-07 by moving to the `Staging` environment name (NEXT-07), which is honest for a dev/test rig but is not available to a real production host. Resolution needs either LDAPS/StartTLS on the shared GLAuth (certificate plus a trust story on each gateway host) or an explicit written posture decision that production gateways bind a different, TLS-capable directory. Surfaced during the NEXT-07 redeploy. |
| NEXT-09 | Build / versioning | Low | **Windows builds stamp git's error text into `InformationalVersion`.** `src/Directory.Build.props:29` runs `git -C "$(MSBuildThisFileDirectory)" …`; MSBuild's directory property ends in a backslash, which escapes the closing quote, so the command is malformed on Windows. The target carries `ContinueOnError`, so the failure is silent and git's stderr is captured as the revision — an observed stamp reads `0.1.2+fatal: cannot change to …`. Any Windows build without a preset `SourceRevisionId` therefore ships a binary that cannot be correlated back to a commit, defeating the point of TST-11. Not reproducible on macOS/Linux, where the separator is `/`. Fix sketch: append `.` to the path or trim the trailing separator before quoting. Surfaced while identifying the deployed binary during NEXT-07. |
| NEXT-10 | Docs / glauth | Medium | **`glauth.md`'s "Pre-provisioned users" table contradicts both the directory and the rest of its own file.** It documents `readonly`/`readonly123` and `admin`/`admin123`, neither of which matches `scadaproj/infra/glauth/config.toml` (`readonly` does not exist there; `admin` carries the standard dev password), and lists the `ReadOnly` gid as `5501` against an actual `5601`. Its dashboard section, by contrast, is correct — so the file is internally inconsistent and a reader cannot tell which half to trust. This table was the **root cause of the NEXT-06 fixture drift**, and it has propagated further: `docs/GatewayTesting.md`'s `MXGATEWAY_LIVE_MXACCESS_WRITE_SECURED_PASSWORD` default and the matching literal in `WorkerLiveMxAccessSmokeTests` both take `admin123` from it. Deliberately **not** fixed in the 2026-08-07 pass: the table is entangled with the OPC-UA group taxonomy (gids, role mapping, and the sister-repo consumers of the same directory), so reconciling it means sweeping that taxonomy as one unit rather than patching two rows. |
## Operator actions still pending (from this cycle's runbooks)
These are **live-infrastructure actions the operator must execute** — the repo-side work is complete and merged:
- ~~**SEC-36** — rotate the dev LDAP service-account credential per `docs/runbooks/SEC-36-ldap-credential-rotation.md` (generate new secret in `scadaproj/infra/glauth`, pre-stage the NSSM env var on deployed hosts, rotate GLAuth on `10.100.0.35`, verify dashboard login). The committed literal is gone from the working tree but remains recoverable from git history until rotation completes — **rotation is the load-bearing half.**~~ **Executed 2026-08-07**: the `serviceaccount` `passsha256` was replaced in `scadaproj/infra/glauth/config.toml` (commit `aada53b`) and the shared GLAuth recreated on `10.100.0.35`, so the literal recoverable from this repo's history no longer binds `dc=zb,dc=local`. The new value lives only in the GLAuth hash, windev's NSSM environment, and dev user-secrets. `wonder-app-vd03` was out of scope (it binds a different, `dc=scadalink`/`dc=scadabridge` directory). **Caveat, since closed:** windev's dashboard `/login` verification was deferred while that host's gateway was crash-looping on the unrelated stale-deployment fault filed as NEXT-07, so the bind was verified directly by `ldapsearch` as `cn=serviceaccount,dc=zb,dc=local` instead. After the 2026-08-07 redeploy the real check ran on windev (login as `multi-role` → 302 + dashboard cookie, anonymous control → `/login`), so the rotated credential is now proven through the `DashboardAuthenticator` path itself; see `docs/runbooks/SEC-36-ldap-credential-rotation.md` Correction 3. SEC-36 is fully `Done`.
- ~~**TST-30** — register a second Gitea `act_runner` on `10.100.0.35` per `docs/runbooks/TST-30-second-ci-runner.md` to relieve the single-shared-runner bottleneck.~~ **Executed 2026-08-07**: `gitea-runner-2` (id 5, capacity 2) is online on `10.100.0.35` via the `/opt/gitea` compose stack, same `container.network: traefik`, token from a `0600` file mount; the existing runner (id 1, capacity 4) was untouched. Concurrency verified — jobs from three runs ran simultaneously across both runners, and a `gitea-runner-2` job cloned successfully from `http://gitea:3000`. TST-30 is now fully `Done`.
- **TST-30 follow-up — reset the Gitea instance runner registration token.** Runner-1's compose block was moved to the same `0600` file-mount pattern as runner-2 on 2026-08-07 (compose and both backups now `0600 root:root`, runner-1 recreated with its identity intact), but hygiene alone does not retire the token: both runners share **one instance-scope registration token** that was world-readable for roughly five months and is still live — a probe registered runner id 6 with it, then deleted it. Gitea 1.26.4 exposes no rotation via CLI or API (both paths are get-or-create and hand back the same value), so the reset must be done in the admin web UI ("Reset registration token"). Afterwards, refresh `/opt/gitea/runner_token` on `10.100.0.35` and shred the two token-bearing compose backups — they are the last copies of the old value. See `docs/runbooks/TST-30-second-ci-runner.md`.
- **TST-25 follow-ups** — old **TST-05** (scheduled live-MXAccess smoke) is now covered by the `nightly-windev` job; old **TST-24** (client wire tests in CI) is unblocked by the working Windows tier.
+1 -1
View File
@@ -49,7 +49,7 @@ Impact: logout (`Dashboard/DashboardEndpointRouteBuilderExtensions.cs:136-155`)
Recommendation: keep the lifetime short (or shorten to ~5 minutes given the factory refreshes per reconnect, `docs/GatewayDashboardDesign.md:497-499`), and confirm no request-path logging captures query strings (Serilog request logging is not currently enabled; keep it that way or scrub `access_token`).
**SEC-6 · Medium — LDAP is plaintext-by-default with a committed service-account password.**
Evidence: `src/ZB.MOM.WW.MxGateway.Server/Configuration/LdapOptions.cs:49-61` (defaults `Transport=None`, `AllowInsecure=true`, `ServiceAccountPassword = "serviceaccount123"`), `appsettings.json:21-33` (same values checked into the repo), `glauth.md:30,327` (dev LDAPS disabled; "binding sends passwords cleartext on the wire").
Evidence: `src/ZB.MOM.WW.MxGateway.Server/Configuration/LdapOptions.cs:49-61` (defaults `Transport=None`, `AllowInsecure=true`, `ServiceAccountPassword = "<service-account-password>"` — value redacted per SEC-36), `appsettings.json:21-33` (same values checked into the repo), `glauth.md:30,327` (dev LDAPS disabled; "binding sends passwords cleartext on the wire").
Impact: every dashboard login sends the operator's password in cleartext to `10.100.0.35:3893`, and the LDAP service-account credential is in source control. This is a documented dev posture (the shadow-options rationale at `LdapOptions.cs:20-28` is explicit that the shared library is secure-by-default), and the validator does enforce the `Transport=None ⇒ AllowInsecure` consistency rule (`GatewayOptionsValidator.cs:82-85`) — but nothing distinguishes dev from prod at runtime.
Recommendation: for production deployment docs, require `Transport=Ldaps`/`StartTls` + `AllowInsecure=false` and move `ServiceAccountPassword` to env-var/secret configuration; consider an `IsProduction` startup check mirroring SEC-4. LDAP injection risk is delegated to the shared `ZB.MOM.WW.Auth.Ldap` provider (bind-then-search per `Dashboard/DashboardAuthenticator.cs:41-47`); its escaping cannot be verified from this repo — flag for review in the donor repo.
+2 -2
View File
@@ -197,7 +197,7 @@ Full design + implementation for each row lives in the linked domain doc under i
| CLI-21 | Low | P2 | S | — | Done | Go `ClientVersion = "0.1.0-dev"` stale vs tagged releases |
| CLI-22 | Low | — | S | — | Not started | Go `newCorrelationID` swallows `crypto/rand` error → empty id |
| CLI-23 | Low | — | S | — | Not started | Go nil-vs-empty bulk short-circuit asymmetry |
| CLI-24 | Low | — | S | — | Not started | Java `MxEventStream` single-consumer constraint undocumented |
| CLI-24 | Low | — | S | — | Done | Java `MxEventStream` single-consumer constraint undocumented (closed 2026-08-07 per 2026-07-12 review old-tracker action; documented at MxEventStream.java:25 "Single consumer") |
| CLI-25 | Low | — | S | — | Not started | Java `close()` does not await channel termination |
| CLI-26 | Low | P2 | S | — | Done | Python `version.py` (0.1.0) ≠ `pyproject.toml` (0.1.2) |
| CLI-27 | Low | — | S | — | Not started | Python `Session.close()` not concurrency-safe; synthesizes reply |
@@ -207,7 +207,7 @@ Full design + implementation for each row lives in the linked domain doc under i
| CLI-31 | Low | — | M | — | Not started | Rust CLI is a single 2,699-line `main.rs` |
| CLI-32 | Low | — | S | — | Not started | Client-side bulk caps differ (.NET/Java unbounded) |
| CLI-33 | Low | — | S | CLI-01,13 | Not started | Per-language event backpressure semantics undocumented |
| CLI-34 | Low | — | S | — | Not started | Python `build/`/`.pytest_cache/` present on disk (untracked) |
| CLI-34 | Low | — | S | — | Done | Python `build/`/`.pytest_cache/` present on disk (untracked) (closed 2026-08-07 per 2026-07-12 review old-tracker action; both gitignored in clients/python/.gitignore) |
### Testing, docs & gaps — [60-testing-docs-gaps.md](60-testing-docs-gaps.md)
@@ -132,7 +132,7 @@ This document turns every finding in the Security/Dashboard/Observability review
## SEC-06 — LDAP plaintext-by-default with a committed service password `Medium` · `P1`
**Finding.** *(review SEC-6)* `Configuration/LdapOptions.cs:49-61` defaults `Transport=None`, `AllowInsecure=true`, `ServiceAccountPassword="serviceaccount123"`; `appsettings.json:21-33` ships the same. `glauth.md:30,327` confirms dev LDAPS is disabled and binds send cleartext. The validator enforces the `Transport=None ⇒ AllowInsecure` consistency rule (`GatewayOptionsValidator.cs:82-85`) but nothing distinguishes dev from prod.
**Finding.** *(review SEC-6)* `Configuration/LdapOptions.cs:49-61` defaults `Transport=None`, `AllowInsecure=true`, `ServiceAccountPassword="<service-account-password>"` (value redacted per SEC-36); `appsettings.json:21-33` ships the same. `glauth.md:30,327` confirms dev LDAPS is disabled and binds send cleartext. The validator enforces the `Transport=None ⇒ AllowInsecure` consistency rule (`GatewayOptionsValidator.cs:82-85`) but nothing distinguishes dev from prod.
**Impact.** Every dashboard login sends the operator's password cleartext to `10.100.0.35:3893`, and a service-account credential is in source control.
@@ -140,7 +140,7 @@ This document turns every finding in the Security/Dashboard/Observability review
**Implementation.**
- `Configuration/GatewayOptionsValidator.cs`: in `ValidateLdap`, when Production and `Transport == None`, emit an error (co-locate with SEC-04's env plumbing).
- Deployment: keep `serviceaccount123` only for local GLAuth dev; document env-var override (`MxGateway__Ldap__ServiceAccountPassword`) for the NSSM-wrapped hosts; rotate the dev credential's reuse.
- Deployment: keep the dev service-account password only for local GLAuth dev; document env-var override (`MxGateway__Ldap__ServiceAccountPassword`) for the NSSM-wrapped hosts; rotate the dev credential's reuse.
- Docs: `docs/GatewayConfiguration.md` Ldap section and a production hardening note referencing `glauth.md`.
- Tests: `GatewayOptionsValidatorTests``Transport=None` + Production → invalid.
+1 -1
View File
@@ -487,7 +487,7 @@ dotnet nuget add source https://gitea.dohertylan.com/api/packages/dohertj2/nuget
Then add the package to your project:
````bash
dotnet add package ZB.MOM.WW.MxGateway.Client --version 0.1.1
dotnet add package ZB.MOM.WW.MxGateway.Client --version 0.2.0
````
The `ZB.MOM.WW.MxGateway.Contracts` package is pulled in transitively.
@@ -19,7 +19,7 @@
<PropertyGroup>
<IsPackable>true</IsPackable>
<PackageId>ZB.MOM.WW.MxGateway.Client</PackageId>
<Version>0.1.2</Version>
<Version>0.2.0</Version>
<Description>.NET 10 gRPC client for the MxAccessGateway service. Provides typed wrappers, retry, and a lazy-browse walker over the Galaxy Repository hierarchy.</Description>
<PackageReadmeFile>README.md</PackageReadmeFile>
<!-- Only the shipped library generates XML docs (matching src/Contracts). The Cli and
+5 -3
View File
@@ -471,7 +471,7 @@ go run ./cmd/mxgw-go smoke -endpoint $env:MXGATEWAY_ENDPOINT -plaintext -api-key
The module is resolved directly from the git repo — no package registry:
````bash
go get gitea.dohertylan.com/dohertj2/mxaccessgw/clients/go@v0.1.1
go get gitea.dohertylan.com/dohertj2/mxaccessgw/clients/go@v0.2.0
````
Then import:
@@ -494,11 +494,13 @@ Go modules in monorepo subdirectories use prefixed tags. To tag a release
from this repo:
````bash
pwsh scripts/tag-go-module.ps1 -Version v0.1.1 -Push
pwsh scripts/tag-go-module.ps1 -Version v0.2.0 -Push
````
The script validates semver, refuses to tag with uncommitted tracked
changes, creates an annotated tag `clients/go/v0.1.1`, and (with `-Push`)
changes, verifies `clients/go/mxgateway/version.go`'s `ClientVersion`
matches the requested tag version (failing the tag otherwise — CLI-21/CLI-39),
creates an annotated tag `clients/go/v0.2.0`, and (with `-Push`)
pushes it to origin.
## Related Documentation
+29 -2
View File
@@ -1,6 +1,16 @@
Set-StrictMode -Version Latest
$ErrorActionPreference = 'Stop'
# Pinned generator baseline. The committed Go bindings stamp these plugin versions in their
# headers (protoc-gen-go v1.36.11 / protoc-gen-go-grpc v1.6.2). Plugin-version drift rewrites
# those header stamps, so a regeneration on an off-pin machine would churn the tree and make
# check-codegen Check 4 false-fail (or mask real drift under churn). Assert the exact versions
# so a regen is deterministic. protoc itself is warn-only (source_code_info is normalized out of
# the committed bindings), matching publish-client-proto-inputs.ps1.
$PinnedProtocGenGoVersion = 'protoc-gen-go v1.36.11'
$PinnedProtocGenGoGrpcVersion = 'protoc-gen-go-grpc 1.6.2'
$PinnedProtocVersion = 'libprotoc 34.1'
$repoRoot = Resolve-Path (Join-Path $PSScriptRoot '..\..')
$protoRoot = Join-Path $repoRoot 'src\ZB.MOM.WW.MxGateway.Contracts\Protos'
$outputRoot = Join-Path $PSScriptRoot 'internal\generated'
@@ -36,8 +46,25 @@ $wingetProtoc = if ($env:LOCALAPPDATA) {
$goBin = if ($env:USERPROFILE) { Join-Path $env:USERPROFILE 'go\bin' } elseif ($env:HOME) { Join-Path $env:HOME 'go/bin' } else { $null }
$protoc = Resolve-Tool -Names @('protoc', 'protoc.exe') -FallbackPaths @($wingetProtoc)
$protocGenGo = Resolve-Tool -Names @('protoc-gen-go', 'protoc-gen-go.exe') -FallbackPaths @((if ($goBin) { Join-Path $goBin 'protoc-gen-go.exe' }), (if ($goBin) { Join-Path $goBin 'protoc-gen-go' }))
$protocGenGoGrpc = Resolve-Tool -Names @('protoc-gen-go-grpc', 'protoc-gen-go-grpc.exe') -FallbackPaths @((if ($goBin) { Join-Path $goBin 'protoc-gen-go-grpc.exe' }), (if ($goBin) { Join-Path $goBin 'protoc-gen-go-grpc' }))
$protocGenGo = Resolve-Tool -Names @('protoc-gen-go', 'protoc-gen-go.exe') -FallbackPaths @(($(if ($goBin) { Join-Path $goBin 'protoc-gen-go.exe' })), ($(if ($goBin) { Join-Path $goBin 'protoc-gen-go' })))
$protocGenGoGrpc = Resolve-Tool -Names @('protoc-gen-go-grpc', 'protoc-gen-go-grpc.exe') -FallbackPaths @(($(if ($goBin) { Join-Path $goBin 'protoc-gen-go-grpc.exe' })), ($(if ($goBin) { Join-Path $goBin 'protoc-gen-go-grpc' })))
# Assert the pinned plugin versions before generating so Check 4 cannot false-fail (or mask drift)
# on an off-pin machine. protoc is warn-only.
$protocGenGoVersion = (& $protocGenGo --version 2>&1 | Out-String).Trim()
if ($protocGenGoVersion -ne $PinnedProtocGenGoVersion) {
throw "protoc-gen-go reports '$protocGenGoVersion', but regeneration is pinned to '$PinnedProtocGenGoVersion'. " +
"Install the pin: go install google.golang.org/protobuf/cmd/protoc-gen-go@v1.36.11"
}
$protocGenGoGrpcVersion = (& $protocGenGoGrpc --version 2>&1 | Out-String).Trim()
if ($protocGenGoGrpcVersion -ne $PinnedProtocGenGoGrpcVersion) {
throw "protoc-gen-go-grpc reports '$protocGenGoGrpcVersion', but regeneration is pinned to '$PinnedProtocGenGoGrpcVersion'. " +
"Install the pin: go install google.golang.org/grpc/cmd/protoc-gen-go-grpc@v1.6.2"
}
$protocVersion = (& $protoc --version 2>&1 | Out-String).Trim()
if ($protocVersion -ne $PinnedProtocVersion) {
Write-Warning "protoc reports '$protocVersion', pin is '$PinnedProtocVersion'. Descriptor comments are normalized out of the committed Go bindings, so patch drift is tolerated; keep CI on the pin."
}
# protoc discovers the plugins on PATH; prepend the directories the resolved plugins live in.
$env:Path = (Split-Path $protocGenGo -Parent) + [System.IO.Path]::PathSeparator + (Split-Path $protocGenGoGrpc -Parent) + [System.IO.Path]::PathSeparator + $env:Path
@@ -2697,6 +2697,9 @@ func (x *Write2Command) GetUserId() int32 {
return 0
}
// The unary reply's statuses field carries the correlated OnWriteComplete
// outcome when it arrives within the worker's bounded wait — see
// MxCommandReply.statuses.
type WriteSecuredCommand struct {
state protoimpl.MessageState `protogen:"open.v1"`
ServerHandle int32 `protobuf:"varint,1,opt,name=server_handle,json=serverHandle,proto3" json:"server_handle,omitempty"`
@@ -2775,6 +2778,9 @@ func (x *WriteSecuredCommand) GetValue() *MxValue {
return nil
}
// The unary reply's statuses field carries the correlated OnWriteComplete
// outcome when it arrives within the worker's bounded wait — see
// MxCommandReply.statuses.
type WriteSecured2Command struct {
state protoimpl.MessageState `protogen:"open.v1"`
ServerHandle int32 `protobuf:"varint,1,opt,name=server_handle,json=serverHandle,proto3" json:"server_handle,omitempty"`
@@ -4575,8 +4581,20 @@ type MxCommandReply struct {
// HRESULT captured from MXAccess or a COM exception. This remains separate
// from gateway protocol status so MXAccess parity details are not hidden by
// transport failures.
Hresult *int32 `protobuf:"varint,5,opt,name=hresult,proto3,oneof" json:"hresult,omitempty"`
ReturnValue *MxValue `protobuf:"bytes,6,opt,name=return_value,json=returnValue,proto3" json:"return_value,omitempty"`
Hresult *int32 `protobuf:"varint,5,opt,name=hresult,proto3,oneof" json:"hresult,omitempty"`
ReturnValue *MxValue `protobuf:"bytes,6,opt,name=return_value,json=returnValue,proto3" json:"return_value,omitempty"`
// Correlated per-item outcome rows. For WRITE_SECURED / WRITE_SECURED2
// replies the worker holds the reply for a bounded window (default 1.5 s,
// MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS) waiting for the matching
// MXAccess OnWriteComplete callback and copies its status rows here, so
// statuses[0] carries the real MXAccess commit outcome (success OR failure)
// while protocol_status/hresult still describe command acceptance only.
// Empty statuses on a write reply means the completion did not arrive
// within the window — the write is unconfirmed, not failed. Correlation is
// best-effort per (server_handle, item_handle): MXAccess's callback carries
// no transaction id, so concurrent writes to the same item within the
// window can swap rows. The OnWriteComplete event still flows on the event
// stream unchanged. Other command kinds leave this field as before.
Statuses []*MxStatusProxy `protobuf:"bytes,7,rep,name=statuses,proto3" json:"statuses,omitempty"`
DiagnosticMessage string `protobuf:"bytes,8,opt,name=diagnostic_message,json=diagnosticMessage,proto3" json:"diagnostic_message,omitempty"`
// Types that are valid to be assigned to Payload:
@@ -5974,8 +5992,12 @@ func (x *WorkerInfoReply) GetMxaccessClsid() string {
}
type DrainEventsReply struct {
state protoimpl.MessageState `protogen:"open.v1"`
Events []*MxEvent `protobuf:"bytes,1,rep,name=events,proto3" json:"events,omitempty"`
state protoimpl.MessageState `protogen:"open.v1"`
// The reply is bounded by both a server-side count cap and the negotiated
// worker-frame byte cap; a reply may therefore carry fewer events than
// `max_events` and fewer than are queued. Callers drain iteratively until an
// empty reply.
Events []*MxEvent `protobuf:"bytes,1,rep,name=events,proto3" json:"events,omitempty"`
unknownFields protoimpl.UnknownFields
sizeCache protoimpl.SizeCache
}
@@ -6411,6 +6433,11 @@ type ReplayGap struct {
// after_worker_sequence = oldest_available_sequence - 1 in the next
// StreamEventsRequest, which will cause the server to replay starting at
// oldest_available_sequence (the first retained event).
// When nothing is retained (the replay ring is empty), this is the next sequence
// that can be delivered — `highest observed + 1` — and the `oldest - 1` resume
// formula remains valid: it resolves to the highest sequence already seen, so the
// follow-up resume replays nothing, reports no gap, and every newer live event
// passes. The interval evicted is unchanged.
OldestAvailableSequence uint64 `protobuf:"varint,2,opt,name=oldest_available_sequence,json=oldestAvailableSequence,proto3" json:"oldest_available_sequence,omitempty"`
unknownFields protoimpl.UnknownFields
sizeCache protoimpl.SizeCache
@@ -431,8 +431,17 @@ type GatewayHello struct {
SupportedProtocolVersion uint32 `protobuf:"varint,1,opt,name=supported_protocol_version,json=supportedProtocolVersion,proto3" json:"supported_protocol_version,omitempty"`
Nonce string `protobuf:"bytes,2,opt,name=nonce,proto3" json:"nonce,omitempty"`
GatewayVersion string `protobuf:"bytes,3,opt,name=gateway_version,json=gatewayVersion,proto3" json:"gateway_version,omitempty"`
unknownFields protoimpl.UnknownFields
sizeCache protoimpl.SizeCache
// Maximum worker-frame payload size, in bytes, negotiated by the gateway from its
// configured pipe limit. The worker adopts this as its frame-protocol MaxMessageBytes
// instead of a hard-coded default; 0 (an older gateway that never set the field) means
// "use the worker's built-in default". Sits above the public gRPC cap by an
// envelope-overhead margin so an accepted gRPC payload always fits one worker frame.
// Every worker->gateway frame — events, heartbeats, faults, and control replies
// including DrainEvents — must serialize within this limit; reply builders truncate
// to fit rather than emit an oversized frame.
MaxFrameBytes uint32 `protobuf:"varint,4,opt,name=max_frame_bytes,json=maxFrameBytes,proto3" json:"max_frame_bytes,omitempty"`
unknownFields protoimpl.UnknownFields
sizeCache protoimpl.SizeCache
}
func (x *GatewayHello) Reset() {
@@ -486,6 +495,13 @@ func (x *GatewayHello) GetGatewayVersion() string {
return ""
}
func (x *GatewayHello) GetMaxFrameBytes() uint32 {
if x != nil {
return x.MaxFrameBytes
}
return 0
}
type WorkerHello struct {
state protoimpl.MessageState `protogen:"open.v1"`
ProtocolVersion uint32 `protobuf:"varint,1,opt,name=protocol_version,json=protocolVersion,proto3" json:"protocol_version,omitempty"`
@@ -1109,11 +1125,12 @@ const file_mxaccess_worker_proto_rawDesc = "" +
"\fworker_event\x18\x12 \x01(\v2\x1f.mxaccess_worker.v1.WorkerEventH\x00R\vworkerEvent\x12P\n" +
"\x10worker_heartbeat\x18\x13 \x01(\v2#.mxaccess_worker.v1.WorkerHeartbeatH\x00R\x0fworkerHeartbeat\x12D\n" +
"\fworker_fault\x18\x14 \x01(\v2\x1f.mxaccess_worker.v1.WorkerFaultH\x00R\vworkerFaultB\x06\n" +
"\x04body\"\x8b\x01\n" +
"\x04body\"\xb3\x01\n" +
"\fGatewayHello\x12<\n" +
"\x1asupported_protocol_version\x18\x01 \x01(\rR\x18supportedProtocolVersion\x12\x14\n" +
"\x05nonce\x18\x02 \x01(\tR\x05nonce\x12'\n" +
"\x0fgateway_version\x18\x03 \x01(\tR\x0egatewayVersion\"\xa1\x01\n" +
"\x0fgateway_version\x18\x03 \x01(\tR\x0egatewayVersion\x12&\n" +
"\x0fmax_frame_bytes\x18\x04 \x01(\rR\rmaxFrameBytes\"\xa1\x01\n" +
"\vWorkerHello\x12)\n" +
"\x10protocol_version\x18\x01 \x01(\rR\x0fprotocolVersion\x12\x14\n" +
"\x05nonce\x18\x02 \x01(\tR\x05nonce\x12*\n" +
+1 -1
View File
@@ -3,7 +3,7 @@ package mxgateway
const (
// ClientVersion is the released semantic version of this Go client module.
// Keep it in sync with the module tag applied by scripts/tag-go-module.ps1.
ClientVersion = "0.1.2"
ClientVersion = "0.2.0"
// GatewayProtocolVersion matches GatewayContractInfo.GatewayProtocolVersion
// in the shared .NET contracts.
+1 -1
View File
@@ -465,7 +465,7 @@ repositories {
}
dependencies {
implementation 'com.zb.mom.ww.mxgateway:zb-mom-ww-mxgateway-client:0.1.2'
implementation 'com.zb.mom.ww.mxgateway:zb-mom-ww-mxgateway-client:0.2.1'
}
````
+8 -1
View File
@@ -13,7 +13,14 @@ ext {
subprojects {
group = 'com.zb.mom.ww.mxgateway'
version = '0.2.0'
// 0.2.0 was already published to the Gitea Maven feed on 2026-06-26,
// before the CLI-37/38/40/41 conformance fixes changed the client's
// observable behavior (status.category-based validation, hresult < 0,
// exact-secret redaction, typed malformed-reply errors). Bump to 0.2.1
// so the published coordinate matches the conformant behavior the other
// four clients ship at 0.2.0 for the first time. See CLI-39 and the
// "Versioning" section of docs/ClientPackaging.md.
version = '0.2.1'
pluginManager.withPlugin('java') {
java {
@@ -3797,6 +3797,9 @@ public final class MxaccessWorker extends com.google.protobuf.GeneratedFile {
* instead of a hard-coded default; 0 (an older gateway that never set the field) means
* "use the worker's built-in default". Sits above the public gRPC cap by an
* envelope-overhead margin so an accepted gRPC payload always fits one worker frame.
* Every worker-&gt;gateway frame events, heartbeats, faults, and control replies
* including DrainEvents must serialize within this limit; reply builders truncate
* to fit rather than emit an oversized frame.
* </pre>
*
* <code>uint32 max_frame_bytes = 4;</code>
@@ -3941,6 +3944,9 @@ public final class MxaccessWorker extends com.google.protobuf.GeneratedFile {
* instead of a hard-coded default; 0 (an older gateway that never set the field) means
* "use the worker's built-in default". Sits above the public gRPC cap by an
* envelope-overhead margin so an accepted gRPC payload always fits one worker frame.
* Every worker-&gt;gateway frame events, heartbeats, faults, and control replies
* including DrainEvents must serialize within this limit; reply builders truncate
* to fit rather than emit an oversized frame.
* </pre>
*
* <code>uint32 max_frame_bytes = 4;</code>
@@ -4499,6 +4505,9 @@ public final class MxaccessWorker extends com.google.protobuf.GeneratedFile {
* instead of a hard-coded default; 0 (an older gateway that never set the field) means
* "use the worker's built-in default". Sits above the public gRPC cap by an
* envelope-overhead margin so an accepted gRPC payload always fits one worker frame.
* Every worker-&gt;gateway frame events, heartbeats, faults, and control replies
* including DrainEvents must serialize within this limit; reply builders truncate
* to fit rather than emit an oversized frame.
* </pre>
*
* <code>uint32 max_frame_bytes = 4;</code>
@@ -4515,6 +4524,9 @@ public final class MxaccessWorker extends com.google.protobuf.GeneratedFile {
* instead of a hard-coded default; 0 (an older gateway that never set the field) means
* "use the worker's built-in default". Sits above the public gRPC cap by an
* envelope-overhead margin so an accepted gRPC payload always fits one worker frame.
* Every worker-&gt;gateway frame events, heartbeats, faults, and control replies
* including DrainEvents must serialize within this limit; reply builders truncate
* to fit rather than emit an oversized frame.
* </pre>
*
* <code>uint32 max_frame_bytes = 4;</code>
@@ -4535,6 +4547,9 @@ public final class MxaccessWorker extends com.google.protobuf.GeneratedFile {
* instead of a hard-coded default; 0 (an older gateway that never set the field) means
* "use the worker's built-in default". Sits above the public gRPC cap by an
* envelope-overhead margin so an accepted gRPC payload always fits one worker frame.
* Every worker-&gt;gateway frame events, heartbeats, faults, and control replies
* including DrainEvents must serialize within this limit; reply builders truncate
* to fit rather than emit an oversized frame.
* </pre>
*
* <code>uint32 max_frame_bytes = 4;</code>
@@ -59,7 +59,7 @@ final class MxGatewayCliTests {
assertEquals(0, run.exitCode());
assertEquals("", run.errors());
assertTrue(run.output().contains("mxgateway-java 0.2.0"));
assertTrue(run.output().contains("mxgateway-java 0.2.1"));
assertTrue(run.output().contains("gatewayProtocolVersion=3"));
assertTrue(run.output().contains("workerProtocolVersion=1"));
}
@@ -89,7 +89,7 @@ final class MxGatewayCliTests {
CliRun run = execute(new FakeClientFactory(), "version", "--json");
assertEquals(0, run.exitCode());
assertTrue(run.output().contains("\"clientVersion\":\"0.2.0\""));
assertTrue(run.output().contains("\"clientVersion\":\"0.2.1\""));
assertTrue(run.output().contains("\"gatewayProtocolVersion\":3"));
}
@@ -63,10 +63,11 @@ protobuf {
// or a plugin/protobuf version bump, silently drifts the committed output. checkGeneratedClean
// fails when the regenerated tree differs from what is committed.
//
// Caveat (repo memory project_java_generated_churn): the protobuf gradle plugin also rewrites
// MxaccessGateway.java with a spurious protobuf-runtime-version delta on every build even when no
// .proto changed. CI reverts that one file (git checkout) before invoking this task; locally, do the
// same when you did not touch a .proto. See docs/GatewayTesting.md "Continuous Integration".
// The grpc/protobuf toolchain is pinned (build.gradle: grpcVersion / protobufVersion), so a
// regeneration is byte-identical to the committed single-file aggregates modulo real .proto
// changes no spurious protobuf-runtime-version churn (IPC-24 verified this and deleted the old
// unconditional CI churn-revert step, which masked message-level drift). Regenerate and commit
// after any .proto change. See docs/GatewayTesting.md "Continuous Integration".
tasks.register('checkGeneratedClean') {
group = 'verification'
description = 'Fails if the committed generated Java tree differs from a fresh regeneration.'
@@ -83,9 +84,9 @@ tasks.register('checkGeneratedClean') {
def dirty = stdout.toString().trim()
if (!dirty.isEmpty()) {
throw new GradleException(
"Generated Java is stale or churned:\n${dirty}\n" +
"Regenerate and commit after a .proto change, or 'git checkout' the spurious " +
"MxaccessGateway.java protobuf-version churn when no .proto changed.")
"Generated Java is stale:\n${dirty}\n" +
"Regenerate and commit the Java client after a .proto change " +
"(gradle :zb-mom-ww-mxgateway-client:generateProto).")
}
}
}
@@ -9,7 +9,7 @@ package com.zb.mom.ww.mxgateway.client;
public final class MxGatewayClientVersion {
private static final int GATEWAY_PROTOCOL_VERSION = 3;
private static final int WORKER_PROTOCOL_VERSION = 1;
private static final String CLIENT_VERSION = "0.2.0";
private static final String CLIENT_VERSION = "0.2.1";
private MxGatewayClientVersion() {
}
+1 -1
View File
@@ -6,7 +6,7 @@ build-backend = "setuptools.build_meta"
[project]
name = "zb-mom-ww-mxaccess-gateway-client"
version = "0.1.2"
version = "0.2.0"
description = "Async Python client for MXAccess Gateway."
readme = "README.md"
requires-python = ">=3.12"
@@ -27,7 +27,7 @@ from google.protobuf import timestamp_pb2 as google_dot_protobuf_dot_timestamp__
import mxaccess_gateway_pb2 as mxaccess__gateway__pb2
DESCRIPTOR = _descriptor_pool.Default().AddSerializedFile(b'\n\x15mxaccess_worker.proto\x12\x12mxaccess_worker.v1\x1a\x1egoogle/protobuf/duration.proto\x1a\x1fgoogle/protobuf/timestamp.proto\x1a\x16mxaccess_gateway.proto\"\x95\x06\n\x0eWorkerEnvelope\x12\x18\n\x10protocol_version\x18\x01 \x01(\r\x12\x12\n\nsession_id\x18\x02 \x01(\t\x12\x10\n\x08sequence\x18\x03 \x01(\x04\x12\x16\n\x0e\x63orrelation_id\x18\x04 \x01(\t\x12\x39\n\rgateway_hello\x18\n \x01(\x0b\x32 .mxaccess_worker.v1.GatewayHelloH\x00\x12\x37\n\x0cworker_hello\x18\x0b \x01(\x0b\x32\x1f.mxaccess_worker.v1.WorkerHelloH\x00\x12\x37\n\x0cworker_ready\x18\x0c \x01(\x0b\x32\x1f.mxaccess_worker.v1.WorkerReadyH\x00\x12;\n\x0eworker_command\x18\r \x01(\x0b\x32!.mxaccess_worker.v1.WorkerCommandH\x00\x12\x46\n\x14worker_command_reply\x18\x0e \x01(\x0b\x32&.mxaccess_worker.v1.WorkerCommandReplyH\x00\x12\x39\n\rworker_cancel\x18\x0f \x01(\x0b\x32 .mxaccess_worker.v1.WorkerCancelH\x00\x12=\n\x0fworker_shutdown\x18\x10 \x01(\x0b\x32\".mxaccess_worker.v1.WorkerShutdownH\x00\x12\x44\n\x13worker_shutdown_ack\x18\x11 \x01(\x0b\x32%.mxaccess_worker.v1.WorkerShutdownAckH\x00\x12\x37\n\x0cworker_event\x18\x12 \x01(\x0b\x32\x1f.mxaccess_worker.v1.WorkerEventH\x00\x12?\n\x10worker_heartbeat\x18\x13 \x01(\x0b\x32#.mxaccess_worker.v1.WorkerHeartbeatH\x00\x12\x37\n\x0cworker_fault\x18\x14 \x01(\x0b\x32\x1f.mxaccess_worker.v1.WorkerFaultH\x00\x42\x06\n\x04\x62ody\"Z\n\x0cGatewayHello\x12\"\n\x1asupported_protocol_version\x18\x01 \x01(\r\x12\r\n\x05nonce\x18\x02 \x01(\t\x12\x17\n\x0fgateway_version\x18\x03 \x01(\t\"i\n\x0bWorkerHello\x12\x18\n\x10protocol_version\x18\x01 \x01(\r\x12\r\n\x05nonce\x18\x02 \x01(\t\x12\x19\n\x11worker_process_id\x18\x03 \x01(\x05\x12\x16\n\x0eworker_version\x18\x04 \x01(\t\"\x8e\x01\n\x0bWorkerReady\x12\x19\n\x11worker_process_id\x18\x01 \x01(\x05\x12\x17\n\x0fmxaccess_progid\x18\x02 \x01(\t\x12\x16\n\x0emxaccess_clsid\x18\x03 \x01(\t\x12\x33\n\x0fready_timestamp\x18\x04 \x01(\x0b\x32\x1a.google.protobuf.Timestamp\"w\n\rWorkerCommand\x12/\n\x07\x63ommand\x18\x01 \x01(\x0b\x32\x1e.mxaccess_gateway.v1.MxCommand\x12\x35\n\x11\x65nqueue_timestamp\x18\x02 \x01(\x0b\x32\x1a.google.protobuf.Timestamp\"\x81\x01\n\x12WorkerCommandReply\x12\x32\n\x05reply\x18\x01 \x01(\x0b\x32#.mxaccess_gateway.v1.MxCommandReply\x12\x37\n\x13\x63ompleted_timestamp\x18\x02 \x01(\x0b\x32\x1a.google.protobuf.Timestamp\"\x1e\n\x0cWorkerCancel\x12\x0e\n\x06reason\x18\x01 \x01(\t\"Q\n\x0eWorkerShutdown\x12/\n\x0cgrace_period\x18\x01 \x01(\x0b\x32\x19.google.protobuf.Duration\x12\x0e\n\x06reason\x18\x02 \x01(\t\"H\n\x11WorkerShutdownAck\x12\x33\n\x06status\x18\x01 \x01(\x0b\x32#.mxaccess_gateway.v1.ProtocolStatus\":\n\x0bWorkerEvent\x12+\n\x05\x65vent\x18\x01 \x01(\x0b\x32\x1c.mxaccess_gateway.v1.MxEvent\"\xa5\x02\n\x0fWorkerHeartbeat\x12\x19\n\x11worker_process_id\x18\x01 \x01(\x05\x12.\n\x05state\x18\x02 \x01(\x0e\x32\x1f.mxaccess_worker.v1.WorkerState\x12?\n\x1blast_sta_activity_timestamp\x18\x03 \x01(\x0b\x32\x1a.google.protobuf.Timestamp\x12\x1d\n\x15pending_command_count\x18\x04 \x01(\r\x12\"\n\x1aoutbound_event_queue_depth\x18\x05 \x01(\r\x12\x1b\n\x13last_event_sequence\x18\x06 \x01(\x04\x12&\n\x1e\x63urrent_command_correlation_id\x18\x07 \x01(\t\"\xf4\x01\n\x0bWorkerFault\x12\x39\n\x08\x63\x61tegory\x18\x01 \x01(\x0e\x32\'.mxaccess_worker.v1.WorkerFaultCategory\x12\x16\n\x0e\x63ommand_method\x18\x02 \x01(\t\x12\x14\n\x07hresult\x18\x03 \x01(\x05H\x00\x88\x01\x01\x12\x16\n\x0e\x65xception_type\x18\x04 \x01(\t\x12\x1a\n\x12\x64iagnostic_message\x18\x05 \x01(\t\x12<\n\x0fprotocol_status\x18\x06 \x01(\x0b\x32#.mxaccess_gateway.v1.ProtocolStatusB\n\n\x08_hresult*\x97\x02\n\x0bWorkerState\x12\x1c\n\x18WORKER_STATE_UNSPECIFIED\x10\x00\x12\x19\n\x15WORKER_STATE_STARTING\x10\x01\x12\x1c\n\x18WORKER_STATE_HANDSHAKING\x10\x02\x12!\n\x1dWORKER_STATE_INITIALIZING_STA\x10\x03\x12\x16\n\x12WORKER_STATE_READY\x10\x04\x12\"\n\x1eWORKER_STATE_EXECUTING_COMMAND\x10\x05\x12\x1e\n\x1aWORKER_STATE_SHUTTING_DOWN\x10\x06\x12\x18\n\x14WORKER_STATE_STOPPED\x10\x07\x12\x18\n\x14WORKER_STATE_FAULTED\x10\x08*\xc7\x04\n\x13WorkerFaultCategory\x12%\n!WORKER_FAULT_CATEGORY_UNSPECIFIED\x10\x00\x12+\n\'WORKER_FAULT_CATEGORY_INVALID_ARGUMENTS\x10\x01\x12\x37\n3WORKER_FAULT_CATEGORY_GATEWAY_AUTHENTICATION_FAILED\x10\x02\x12+\n\'WORKER_FAULT_CATEGORY_PROTOCOL_MISMATCH\x10\x03\x12,\n(WORKER_FAULT_CATEGORY_PROTOCOL_VIOLATION\x10\x04\x12+\n\'WORKER_FAULT_CATEGORY_PIPE_DISCONNECTED\x10\x05\x12\x32\n.WORKER_FAULT_CATEGORY_MXACCESS_CREATION_FAILED\x10\x06\x12\x31\n-WORKER_FAULT_CATEGORY_MXACCESS_COMMAND_FAILED\x10\x07\x12:\n6WORKER_FAULT_CATEGORY_MXACCESS_EVENT_CONVERSION_FAILED\x10\x08\x12\"\n\x1eWORKER_FAULT_CATEGORY_STA_HUNG\x10\t\x12(\n$WORKER_FAULT_CATEGORY_QUEUE_OVERFLOW\x10\n\x12*\n&WORKER_FAULT_CATEGORY_SHUTDOWN_TIMEOUT\x10\x0b\x42&\xaa\x02#ZB.MOM.WW.MxGateway.Contracts.Protob\x06proto3')
DESCRIPTOR = _descriptor_pool.Default().AddSerializedFile(b'\n\x15mxaccess_worker.proto\x12\x12mxaccess_worker.v1\x1a\x1egoogle/protobuf/duration.proto\x1a\x1fgoogle/protobuf/timestamp.proto\x1a\x16mxaccess_gateway.proto\"\x95\x06\n\x0eWorkerEnvelope\x12\x18\n\x10protocol_version\x18\x01 \x01(\r\x12\x12\n\nsession_id\x18\x02 \x01(\t\x12\x10\n\x08sequence\x18\x03 \x01(\x04\x12\x16\n\x0e\x63orrelation_id\x18\x04 \x01(\t\x12\x39\n\rgateway_hello\x18\n \x01(\x0b\x32 .mxaccess_worker.v1.GatewayHelloH\x00\x12\x37\n\x0cworker_hello\x18\x0b \x01(\x0b\x32\x1f.mxaccess_worker.v1.WorkerHelloH\x00\x12\x37\n\x0cworker_ready\x18\x0c \x01(\x0b\x32\x1f.mxaccess_worker.v1.WorkerReadyH\x00\x12;\n\x0eworker_command\x18\r \x01(\x0b\x32!.mxaccess_worker.v1.WorkerCommandH\x00\x12\x46\n\x14worker_command_reply\x18\x0e \x01(\x0b\x32&.mxaccess_worker.v1.WorkerCommandReplyH\x00\x12\x39\n\rworker_cancel\x18\x0f \x01(\x0b\x32 .mxaccess_worker.v1.WorkerCancelH\x00\x12=\n\x0fworker_shutdown\x18\x10 \x01(\x0b\x32\".mxaccess_worker.v1.WorkerShutdownH\x00\x12\x44\n\x13worker_shutdown_ack\x18\x11 \x01(\x0b\x32%.mxaccess_worker.v1.WorkerShutdownAckH\x00\x12\x37\n\x0cworker_event\x18\x12 \x01(\x0b\x32\x1f.mxaccess_worker.v1.WorkerEventH\x00\x12?\n\x10worker_heartbeat\x18\x13 \x01(\x0b\x32#.mxaccess_worker.v1.WorkerHeartbeatH\x00\x12\x37\n\x0cworker_fault\x18\x14 \x01(\x0b\x32\x1f.mxaccess_worker.v1.WorkerFaultH\x00\x42\x06\n\x04\x62ody\"s\n\x0cGatewayHello\x12\"\n\x1asupported_protocol_version\x18\x01 \x01(\r\x12\r\n\x05nonce\x18\x02 \x01(\t\x12\x17\n\x0fgateway_version\x18\x03 \x01(\t\x12\x17\n\x0fmax_frame_bytes\x18\x04 \x01(\r\"i\n\x0bWorkerHello\x12\x18\n\x10protocol_version\x18\x01 \x01(\r\x12\r\n\x05nonce\x18\x02 \x01(\t\x12\x19\n\x11worker_process_id\x18\x03 \x01(\x05\x12\x16\n\x0eworker_version\x18\x04 \x01(\t\"\x8e\x01\n\x0bWorkerReady\x12\x19\n\x11worker_process_id\x18\x01 \x01(\x05\x12\x17\n\x0fmxaccess_progid\x18\x02 \x01(\t\x12\x16\n\x0emxaccess_clsid\x18\x03 \x01(\t\x12\x33\n\x0fready_timestamp\x18\x04 \x01(\x0b\x32\x1a.google.protobuf.Timestamp\"w\n\rWorkerCommand\x12/\n\x07\x63ommand\x18\x01 \x01(\x0b\x32\x1e.mxaccess_gateway.v1.MxCommand\x12\x35\n\x11\x65nqueue_timestamp\x18\x02 \x01(\x0b\x32\x1a.google.protobuf.Timestamp\"\x81\x01\n\x12WorkerCommandReply\x12\x32\n\x05reply\x18\x01 \x01(\x0b\x32#.mxaccess_gateway.v1.MxCommandReply\x12\x37\n\x13\x63ompleted_timestamp\x18\x02 \x01(\x0b\x32\x1a.google.protobuf.Timestamp\"\x1e\n\x0cWorkerCancel\x12\x0e\n\x06reason\x18\x01 \x01(\t\"Q\n\x0eWorkerShutdown\x12/\n\x0cgrace_period\x18\x01 \x01(\x0b\x32\x19.google.protobuf.Duration\x12\x0e\n\x06reason\x18\x02 \x01(\t\"H\n\x11WorkerShutdownAck\x12\x33\n\x06status\x18\x01 \x01(\x0b\x32#.mxaccess_gateway.v1.ProtocolStatus\":\n\x0bWorkerEvent\x12+\n\x05\x65vent\x18\x01 \x01(\x0b\x32\x1c.mxaccess_gateway.v1.MxEvent\"\xa5\x02\n\x0fWorkerHeartbeat\x12\x19\n\x11worker_process_id\x18\x01 \x01(\x05\x12.\n\x05state\x18\x02 \x01(\x0e\x32\x1f.mxaccess_worker.v1.WorkerState\x12?\n\x1blast_sta_activity_timestamp\x18\x03 \x01(\x0b\x32\x1a.google.protobuf.Timestamp\x12\x1d\n\x15pending_command_count\x18\x04 \x01(\r\x12\"\n\x1aoutbound_event_queue_depth\x18\x05 \x01(\r\x12\x1b\n\x13last_event_sequence\x18\x06 \x01(\x04\x12&\n\x1e\x63urrent_command_correlation_id\x18\x07 \x01(\t\"\xf4\x01\n\x0bWorkerFault\x12\x39\n\x08\x63\x61tegory\x18\x01 \x01(\x0e\x32\'.mxaccess_worker.v1.WorkerFaultCategory\x12\x16\n\x0e\x63ommand_method\x18\x02 \x01(\t\x12\x14\n\x07hresult\x18\x03 \x01(\x05H\x00\x88\x01\x01\x12\x16\n\x0e\x65xception_type\x18\x04 \x01(\t\x12\x1a\n\x12\x64iagnostic_message\x18\x05 \x01(\t\x12<\n\x0fprotocol_status\x18\x06 \x01(\x0b\x32#.mxaccess_gateway.v1.ProtocolStatusB\n\n\x08_hresult*\x97\x02\n\x0bWorkerState\x12\x1c\n\x18WORKER_STATE_UNSPECIFIED\x10\x00\x12\x19\n\x15WORKER_STATE_STARTING\x10\x01\x12\x1c\n\x18WORKER_STATE_HANDSHAKING\x10\x02\x12!\n\x1dWORKER_STATE_INITIALIZING_STA\x10\x03\x12\x16\n\x12WORKER_STATE_READY\x10\x04\x12\"\n\x1eWORKER_STATE_EXECUTING_COMMAND\x10\x05\x12\x1e\n\x1aWORKER_STATE_SHUTTING_DOWN\x10\x06\x12\x18\n\x14WORKER_STATE_STOPPED\x10\x07\x12\x18\n\x14WORKER_STATE_FAULTED\x10\x08*\xc7\x04\n\x13WorkerFaultCategory\x12%\n!WORKER_FAULT_CATEGORY_UNSPECIFIED\x10\x00\x12+\n\'WORKER_FAULT_CATEGORY_INVALID_ARGUMENTS\x10\x01\x12\x37\n3WORKER_FAULT_CATEGORY_GATEWAY_AUTHENTICATION_FAILED\x10\x02\x12+\n\'WORKER_FAULT_CATEGORY_PROTOCOL_MISMATCH\x10\x03\x12,\n(WORKER_FAULT_CATEGORY_PROTOCOL_VIOLATION\x10\x04\x12+\n\'WORKER_FAULT_CATEGORY_PIPE_DISCONNECTED\x10\x05\x12\x32\n.WORKER_FAULT_CATEGORY_MXACCESS_CREATION_FAILED\x10\x06\x12\x31\n-WORKER_FAULT_CATEGORY_MXACCESS_COMMAND_FAILED\x10\x07\x12:\n6WORKER_FAULT_CATEGORY_MXACCESS_EVENT_CONVERSION_FAILED\x10\x08\x12\"\n\x1eWORKER_FAULT_CATEGORY_STA_HUNG\x10\t\x12(\n$WORKER_FAULT_CATEGORY_QUEUE_OVERFLOW\x10\n\x12*\n&WORKER_FAULT_CATEGORY_SHUTDOWN_TIMEOUT\x10\x0b\x42&\xaa\x02#ZB.MOM.WW.MxGateway.Contracts.Protob\x06proto3')
_globals = globals()
_builder.BuildMessageAndEnumDescriptors(DESCRIPTOR, _globals)
@@ -35,32 +35,32 @@ _builder.BuildTopDescriptorsAndMessages(DESCRIPTOR, 'mxaccess_worker_pb2', _glob
if not _descriptor._USE_C_DESCRIPTORS:
_globals['DESCRIPTOR']._loaded_options = None
_globals['DESCRIPTOR']._serialized_options = b'\252\002#ZB.MOM.WW.MxGateway.Contracts.Proto'
_globals['_WORKERSTATE']._serialized_start=2316
_globals['_WORKERSTATE']._serialized_end=2595
_globals['_WORKERFAULTCATEGORY']._serialized_start=2598
_globals['_WORKERFAULTCATEGORY']._serialized_end=3181
_globals['_WORKERSTATE']._serialized_start=2341
_globals['_WORKERSTATE']._serialized_end=2620
_globals['_WORKERFAULTCATEGORY']._serialized_start=2623
_globals['_WORKERFAULTCATEGORY']._serialized_end=3206
_globals['_WORKERENVELOPE']._serialized_start=135
_globals['_WORKERENVELOPE']._serialized_end=924
_globals['_GATEWAYHELLO']._serialized_start=926
_globals['_GATEWAYHELLO']._serialized_end=1016
_globals['_WORKERHELLO']._serialized_start=1018
_globals['_WORKERHELLO']._serialized_end=1123
_globals['_WORKERREADY']._serialized_start=1126
_globals['_WORKERREADY']._serialized_end=1268
_globals['_WORKERCOMMAND']._serialized_start=1270
_globals['_WORKERCOMMAND']._serialized_end=1389
_globals['_WORKERCOMMANDREPLY']._serialized_start=1392
_globals['_WORKERCOMMANDREPLY']._serialized_end=1521
_globals['_WORKERCANCEL']._serialized_start=1523
_globals['_WORKERCANCEL']._serialized_end=1553
_globals['_WORKERSHUTDOWN']._serialized_start=1555
_globals['_WORKERSHUTDOWN']._serialized_end=1636
_globals['_WORKERSHUTDOWNACK']._serialized_start=1638
_globals['_WORKERSHUTDOWNACK']._serialized_end=1710
_globals['_WORKEREVENT']._serialized_start=1712
_globals['_WORKEREVENT']._serialized_end=1770
_globals['_WORKERHEARTBEAT']._serialized_start=1773
_globals['_WORKERHEARTBEAT']._serialized_end=2066
_globals['_WORKERFAULT']._serialized_start=2069
_globals['_WORKERFAULT']._serialized_end=2313
_globals['_GATEWAYHELLO']._serialized_end=1041
_globals['_WORKERHELLO']._serialized_start=1043
_globals['_WORKERHELLO']._serialized_end=1148
_globals['_WORKERREADY']._serialized_start=1151
_globals['_WORKERREADY']._serialized_end=1293
_globals['_WORKERCOMMAND']._serialized_start=1295
_globals['_WORKERCOMMAND']._serialized_end=1414
_globals['_WORKERCOMMANDREPLY']._serialized_start=1417
_globals['_WORKERCOMMANDREPLY']._serialized_end=1546
_globals['_WORKERCANCEL']._serialized_start=1548
_globals['_WORKERCANCEL']._serialized_end=1578
_globals['_WORKERSHUTDOWN']._serialized_start=1580
_globals['_WORKERSHUTDOWN']._serialized_end=1661
_globals['_WORKERSHUTDOWNACK']._serialized_start=1663
_globals['_WORKERSHUTDOWNACK']._serialized_end=1735
_globals['_WORKEREVENT']._serialized_start=1737
_globals['_WORKEREVENT']._serialized_end=1795
_globals['_WORKERHEARTBEAT']._serialized_start=1798
_globals['_WORKERHEARTBEAT']._serialized_end=2091
_globals['_WORKERFAULT']._serialized_start=2094
_globals['_WORKERFAULT']._serialized_end=2338
# @@protoc_insertion_point(module_scope)
@@ -1,3 +1,3 @@
"""Package version information."""
__version__ = "0.1.2"
__version__ = "0.2.0"
+17
View File
@@ -1,6 +1,8 @@
"""Tests for the Python CLI."""
import json
import tomllib
from pathlib import Path
import pytest
from click.testing import CliRunner
@@ -12,6 +14,21 @@ from zb_mom_ww_mxgateway_cli.commands import main
_BATCH_EOR = "__MXGW_BATCH_EOR__"
def test_version_matches_pyproject_toml() -> None:
"""`__version__` must track `pyproject.toml`'s `[project].version`.
The existing `version` command tests only assert self-consistency against
`__version__` (the two hardcoded literals could still drift from each
other without either test catching it the CLI-26 residual drift mode).
This test pins `__version__` to the single source of truth instead.
"""
pyproject_path = Path(__file__).resolve().parent.parent / "pyproject.toml"
with pyproject_path.open("rb") as handle:
pyproject = tomllib.load(handle)
assert __version__ == pyproject["project"]["version"]
def test_require_certificate_validation_flag_flows_through_connect(
monkeypatch: pytest.MonkeyPatch,
) -> None:
+2 -2
View File
@@ -590,7 +590,7 @@ checksum = "1d87ecb2933e8aeadb3e3a02b828fed80a7528047e68b4f424523a0981a3a084"
[[package]]
name = "mxgw-cli"
version = "0.1.2"
version = "0.2.0"
dependencies = [
"clap",
"futures-util",
@@ -1490,7 +1490,7 @@ dependencies = [
[[package]]
name = "zb-mom-ww-mxgateway-client"
version = "0.1.2"
version = "0.2.0"
dependencies = [
"futures-core",
"futures-util",
+2 -2
View File
@@ -1,6 +1,6 @@
[package]
name = "zb-mom-ww-mxgateway-client"
version = "0.1.2"
version = "0.2.0"
edition = "2021"
authors = ["Joseph Doherty"]
description = "Async Rust client for the MxAccessGateway gRPC service, including a lazy-browse walker over the Galaxy Repository hierarchy."
@@ -25,7 +25,7 @@ resolver = "2"
[workspace.package]
edition = "2021"
version = "0.1.2"
version = "0.2.0"
authors = ["Joseph Doherty"]
license = "Proprietary"
repository = "https://gitea.dohertylan.com/dohertj2/mxaccessgw"
+1 -1
View File
@@ -436,5 +436,5 @@ Then add the dependency:
```toml
[dependencies]
zb-mom-ww-mxgateway-client = { version = "0.1.1", registry = "dohertj2-gitea" }
zb-mom-ww-mxgateway-client = { version = "0.2.0", registry = "dohertj2-gitea" }
```
@@ -256,6 +256,9 @@ message Write2Command {
int32 user_id = 5;
}
// The unary reply's statuses field carries the correlated OnWriteComplete
// outcome when it arrives within the worker's bounded wait see
// MxCommandReply.statuses.
message WriteSecuredCommand {
int32 server_handle = 1;
int32 item_handle = 2;
@@ -266,6 +269,9 @@ message WriteSecuredCommand {
MxValue value = 5;
}
// The unary reply's statuses field carries the correlated OnWriteComplete
// outcome when it arrives within the worker's bounded wait see
// MxCommandReply.statuses.
message WriteSecured2Command {
int32 server_handle = 1;
int32 item_handle = 2;
@@ -525,6 +531,18 @@ message MxCommandReply {
// transport failures.
optional int32 hresult = 5;
MxValue return_value = 6;
// Correlated per-item outcome rows. For WRITE_SECURED / WRITE_SECURED2
// replies the worker holds the reply for a bounded window (default 1.5 s,
// MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS) waiting for the matching
// MXAccess OnWriteComplete callback and copies its status rows here, so
// statuses[0] carries the real MXAccess commit outcome (success OR failure)
// while protocol_status/hresult still describe command acceptance only.
// Empty statuses on a write reply means the completion did not arrive
// within the window the write is unconfirmed, not failed. Correlation is
// best-effort per (server_handle, item_handle): MXAccess's callback carries
// no transaction id, so concurrent writes to the same item within the
// window can swap rows. The OnWriteComplete event still flows on the event
// stream unchanged. Other command kinds leave this field as before.
repeated MxStatusProxy statuses = 7;
string diagnostic_message = 8;
@@ -676,6 +694,10 @@ message WorkerInfoReply {
}
message DrainEventsReply {
// The reply is bounded by both a server-side count cap and the negotiated
// worker-frame byte cap; a reply may therefore carry fewer events than
// `max_events` and fewer than are queued. Callers drain iteratively until an
// empty reply.
repeated MxEvent events = 1;
}
@@ -760,6 +782,11 @@ message ReplayGap {
// after_worker_sequence = oldest_available_sequence - 1 in the next
// StreamEventsRequest, which will cause the server to replay starting at
// oldest_available_sequence (the first retained event).
// When nothing is retained (the replay ring is empty), this is the next sequence
// that can be delivered `highest observed + 1` and the `oldest - 1` resume
// formula remains valid: it resolves to the highest sequence already seen, so the
// follow-up resume replays nothing, reports no gap, and every newer live event
// passes. The interval evicted is unchanged.
uint64 oldest_available_sequence = 2;
}
@@ -47,6 +47,9 @@ message GatewayHello {
// instead of a hard-coded default; 0 (an older gateway that never set the field) means
// "use the worker's built-in default". Sits above the public gRPC cap by an
// envelope-overhead margin so an accepted gRPC payload always fits one worker frame.
// Every worker->gateway frame events, heartbeats, faults, and control replies
// including DrainEvents must serialize within this limit; reply builders truncate
// to fit rather than emit an oversized frame.
uint32 max_frame_bytes = 4;
}
+85
View File
@@ -32,6 +32,85 @@ $env:MXGATEWAY_TEST_ITEM = 'TestObject.TestInt'
Use plaintext only for a local gateway. Use TLS when the gateway crosses a
machine boundary or uses a production certificate.
## Versioning
Every client's version lives in its own manifest: `clients/rust/Cargo.toml`
(`[package]` and `[workspace.package]`, both must match — `crates/mxgw-cli`
inherits via `version.workspace = true`), `clients/python/pyproject.toml`
(`[project].version`) and `clients/python/src/zb_mom_ww_mxgateway/version.py`
(`__version__`, must match `pyproject.toml`), `clients/go/mxgateway/version.go`
(`ClientVersion`), `clients/dotnet/ZB.MOM.WW.MxGateway.Client/ZB.MOM.WW.MxGateway.Client.csproj`
(`<Version>`), and `clients/java/build.gradle` (`subprojects { version = ... }`,
mirrored by the hand-maintained `MxGatewayClientVersion.CLIENT_VERSION`
constant — the two have drifted before and there is no build-time link
between them, so bump both together).
`src/ZB.MOM.WW.MxGateway.Contracts/ZB.MOM.WW.MxGateway.Contracts.csproj`
(`<Version>`) is a fifth, easy-to-miss manifest: it is not itself a
language client, but `Invoke-PackDotnet` in `scripts/pack-clients.ps1`
packs and publishes it in lockstep with the .NET Client (both
`ZB.MOM.WW.MxGateway.*` nupkgs go through the same `-Publish` loop), and
Contracts and the .NET Client have always released at the same version.
Bump Contracts' `<Version>` alongside the .NET Client's — leaving it behind
means the next `-Publish` packs a stale Contracts version, the collision
guard below correctly refuses to re-publish it, and the loop aborts
mid-way with the Client possibly already pushed (nupkgs are enumerated
alphabetically, and `Client` sorts before `Contracts`). This is distinct
from `src/Directory.Build.props`'s repo-wide `<Version>` default, which
stamps the Server/Worker/test assemblies and is not part of the published
client package set — see the comment there.
**Bump the version before every publish, never after.** A Gitea package feed
rejects re-uploading an existing name+version, and `scripts/pack-clients.ps1`
enforces this before it ever attempts a push: each per-language `-Publish`
step queries the Gitea package API
(`GET /api/v1/packages/dohertj2/{type}/{name}/{version}`) for the version
about to be published and aborts with a clear error if it already exists —
the script never force-overwrites a published artifact. `scripts/tag-go-module.ps1`
carries the equivalent guard for the Go module: it refuses to create a
`clients/go/vX.Y.Z` tag unless `clients/go/mxgateway/version.go`'s
`ClientVersion` already equals `X.Y.Z` (CLI-21/CLI-39), so a forgotten
version bump fails the tag instead of shipping a mismatched module.
As of 2026-08-07 (CLI-39) all five clients — plus `ZB.MOM.WW.MxGateway.Contracts`,
which releases in lockstep with the .NET Client — moved to **0.2.0**,
converging on one number after four of the five had drifted onto the
*already-published* 0.1.2/0.1.1 while their public APIs kept changing
underneath it (see `archreview/2026-07-12/remediation/50-clients.md` CLI-39).
A code-review follow-up on the same branch caught that the initial CLI-39
pass bumped the .NET Client but left `Contracts.csproj` at 0.1.2 — since
both publish through the same `Invoke-PackDotnet` `-Publish` loop, that
would have made the very next `.NET` publish abort on the new collision
guard partway through (Client already pushed, Contracts refused as a
re-publish of the already-published 0.1.2). Fixed in the same branch.
Verified against the live Gitea package API at that time: `nuget` had
`ZB.MOM.WW.MxGateway.Client` and `.Contracts` published through 0.1.2; `pypi` (`zb-mom-ww-mxaccess-gateway-client`)
and `cargo` (`zb-mom-ww-mxgateway-client`) had only reached 0.1.1 despite their
source pinning 0.1.2; **`maven`
(`com.zb.mom.ww.mxgateway:zb-mom-ww-mxgateway-client`) had already published
0.2.0 on 2026-06-26** — before the CLI-37/38/40/41 conformance fixes changed
the client's observable behavior (`category`-based status validation,
`hresult < 0`, exact-secret redaction, typed malformed-reply errors). Reusing
0.2.0 for the conformant Java build would have labeled two different APIs
with the same coordinate, so **Java is the one exception: it shipped as
0.2.1**, not 0.2.0. Operators publishing a future release must re-check the
target version against the live registry before assuming any of these
numbers are still unclaimed — the guards above do this automatically at
publish time, but a version bump in the source is still a manual step per
client.
On 2026-08-07 that release shipped. Published coordinates on
`gitea.dohertylan.com`: `nuget` `ZB.MOM.WW.MxGateway.Client` **0.2.0** and
`ZB.MOM.WW.MxGateway.Contracts` **0.2.0**, `pypi`
`zb-mom-ww-mxaccess-gateway-client` **0.2.0**, `cargo`
`zb-mom-ww-mxgateway-client` **0.2.0**, `maven`
`com.zb.mom.ww.mxgateway:zb-mom-ww-mxgateway-client` **0.2.1** (the Java
exception described above). Go publishes no artifact — it ships as the module
tag `clients/go/v0.2.0`, created at commit `a346d51`. Each coordinate was
confirmed present through the Gitea package API after the push, and
`go list -m` resolves the Go tag. These are the numbers a future release
bumps off.
## .NET
The .NET client uses .NET 10 and references
@@ -133,6 +212,12 @@ a `cargo package` that cannot build from the vendored tree alone would mean
the vendored copies are stale, and verification is what catches that before
publish.
Publishing to the `dohertj2-gitea` alternative registry reads the token from
`CARGO_REGISTRIES_DOHERTJ2_GITEA_TOKEN`, and that variable must hold
`Bearer <token>` — cargo sends the value as the `Authorization` header
verbatim and Gitea's cargo registry rejects a bare token with `401`, unlike
the other feeds, which authenticate with a username/token basic-auth pair.
Regenerate and compile Rust bindings:
```powershell
+21 -11
View File
@@ -74,10 +74,12 @@ protoc-version encoding drift and does not false-fail across protoc releases
(it warns, rather than fails, when protoc is off the pin).
The gateway test project carries an independent, protoc-free freshness guard:
`ClientProtoInputTests.Descriptor_ContainsEveryContractMessageAndField` reflects
over the in-process contract descriptors and fails if any contract message or
field is missing from the committed protoset. This is the primary CI gate for
descriptor staleness; a red test means "regenerate and commit the protoset."
`ClientProtoInputTests.Descriptor_ContainsEveryContractSymbol` reflects over the
in-process contract descriptors `mxaccess_gateway.proto`, `mxaccess_worker.proto`,
and `galaxy_repository.proto` — and fails if any contract message, field, enum,
enum value, service, or method is missing from the committed protoset. This is
the primary CI gate for descriptor staleness; a red test means "regenerate and
commit the protoset."
### Pinned generator versions
@@ -88,15 +90,23 @@ scripts assert the pin and resolve tools from `PATH`:
| Generator | Pinned version | Guard |
|-----------|----------------|-------|
| protoc (descriptor set) | 34.1 | version assertion in `scripts/publish-client-proto-inputs.ps1` |
| `Grpc.Tools` (C# `Generated/`) | 2.80.0 (contracts csproj) | `scripts/check-codegen.ps1` git-diff of `Generated/` |
| `grpcio-tools` (Python) | 1.80.0 (protobuf runtime 6.31.1) | version assertion in `clients/python/generate-proto.ps1` |
| protobuf / grpc-java (Java) | `protobufVersion` / `grpcVersion` in `clients/java/build.gradle` | `checkGeneratedClean` gradle task |
| `Grpc.Tools` (C# `Generated/`) | 2.80.0 (contracts csproj) | `scripts/check-codegen.ps1` git-diff of `Generated/` (Check 2) |
| `protoc-gen-go` (Go) | v1.36.11 | version assertion in `clients/go/generate-proto.ps1`; `check-codegen.ps1` Check 4 |
| `protoc-gen-go-grpc` (Go) | 1.6.2 | version assertion in `clients/go/generate-proto.ps1`; `check-codegen.ps1` Check 4 |
| `grpcio-tools` (Python) | 1.80.0 (protobuf runtime 6.31.1) | version assertion in `clients/python/generate-proto.ps1`; `check-codegen.ps1` Check 4 |
| protobuf / grpc-java (Java) | `protobufVersion` / `grpcVersion` in `clients/java/build.gradle` | CI `git diff --exit-code` over `clients/java/src/main/generated` after `gradle test` (the `checkGeneratedClean` gradle task is the equivalent local check) |
A newer `grpcio-tools` stamps a `GRPC_GENERATED_VERSION` above the pinned grpcio
runtime and breaks Python `pytest`; the Java protobuf plugin rewrites
`MxaccessGateway.java` with spurious protobuf-runtime-version churn on every build
(revert that one file when no `.proto` changed — see
[Gateway Testing](./GatewayTesting.md) "Continuous Integration").
runtime and breaks Python `pytest`, so the Python and Go scripts assert their
generator pins before regenerating. Under the pinned grpc/protobuf toolchain the
Java protobuf plugin regenerates byte-identical output (modulo real `.proto`
changes), so CI enforces Java freshness with a direct `git diff --exit-code -- clients/java/src/main/generated`
step after `gradle test` (which transitively regenerates via `generateProto`); the
`checkGeneratedClean` gradle task is the equivalent check for local/manual use. The
old unconditional churn-revert CI step (which masked message-level drift in the
single-file Java aggregates) was deleted (IPC-24). Go and Python committed
bindings are guarded by `check-codegen.ps1` **Check 4**, which regenerates both
and fails on any diff.
## Output Directories
+10 -5
View File
@@ -114,7 +114,11 @@ dotnet build src/ZB.MOM.WW.MxGateway.Contracts/ZB.MOM.WW.MxGateway.Contracts.csp
`scripts/check-codegen.ps1` enforces this in CI (it force-regenerates and fails on any
`git diff` against the committed `Generated/`) — that regeneration diff in the `portable`
job is the primary guard. The SSH-driven `windows-x86` job's net48 worker build is the
secondary guard (a stale `Generated/` also breaks the x86 build with `CS0246`).
secondary guard (a stale `Generated/` also breaks the x86 build with `CS0246`). The same
script runs four checks in total: the committed client descriptor set (Check 1), the C#
`Generated/` (Check 2), the Rust vendored protos (Check 3), and the Go/Python client bindings
(Check 4) each regenerate and fail on any diff. See
[Client Proto Generation](./ClientProtoGeneration.md) for the pinned generator versions.
Client generation inputs are published through
`clients/proto/proto-inputs.json` and the descriptor set under
@@ -152,10 +156,11 @@ pwsh -File scripts/publish-client-proto-inputs.ps1
Freshness is guarded two ways so a skipped regeneration cannot ship silently:
- `ClientProtoInputTests.Descriptor_ContainsEveryContractMessageAndField` (gateway test project)
reflects over the in-process contract descriptors and fails if any message or field is missing
from the committed protoset. It is semantic (symbol presence), needs no protoc, and runs in the
Linux CI. A red test means "regenerate and commit the protoset."
- `ClientProtoInputTests.Descriptor_ContainsEveryContractSymbol` (gateway test project)
reflects over the in-process contract descriptors — including `galaxy_repository.proto` — and
fails if any message, field, enum, enum value, service, or method is missing from the committed
protoset. It is semantic (symbol presence), needs no protoc, and runs in the Linux CI. A red
test means "regenerate and commit the protoset."
- `pwsh -File scripts/publish-client-proto-inputs.ps1 -Check` rebuilds the descriptor and compares
it to the committed one. The comparison normalizes both sides through the same protoc with
`source_code_info` stripped, so it does not false-fail across protoc releases.
+42
View File
@@ -534,6 +534,48 @@ against the live MXAccess attribute set.
- [Alarm Client Discovery — Subtag provider](./AlarmClientDiscovery.md)
- [gRPC Contract — provider_status and degraded fields](./Grpc.md)
## Secured-Write Completion Correlation
MXAccess writes are fire-and-forget: the toolkit call returns before the
Galaxy commit, and the per-item outcome only exists in the later
`OnWriteComplete` COM callback. The original unary write reply therefore
proved worker-side command acceptance only, forcing consumers (OtOpcUa's
GalaxyDriver) to report every write as provisionally good.
For `WriteSecured`/`WriteSecured2` the worker now holds the unary reply for a
bounded window (`MxGateway:Worker:WriteCompletionWaitMilliseconds`, default
1.5 s, `0` disables; conveyed to the worker via
`MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS`) and copies the matching
callback's status rows onto `MxCommandReply.statuses`. Key choices, argued in
[the design doc](./plans/2026-08-09-write-completion-correlation-design.md):
- **Pump-wait on the STA, not a parked reply.** The executor holds the STA
thread but pumps Windows messages each poll — the shipped ReadBulk pattern —
because commands serialize per session anyway, so freeing the STA during the
wait buys nothing and a parked reply would change the dispatcher/pipe
contracts.
- **Version baseline before the COM call** closes the fast-completion edge: a
callback that dispatches while `WriteSecured` is still on the stack still
correlates.
- **Timeout returns today's shape** (protocol OK, empty statuses):
unconfirmed is honest; a synthesized failure row would trigger consumer-side
write-revert logic on slow-but-successful commits. The 1.5 s default stays
inside OtOpcUa's 2 s Tier A write-resilience budget.
- **Parity preserved.** `protocol_status`/`hresult` keep describing
acceptance; the MXAccess outcome (success or failure) rides only in
`statuses[0]`; the `OnWriteComplete` event still streams unchanged (nothing
swallowed, nothing synthesized).
- **Scope: secured writes only.** Plain `Write`/`Write2` and bulk writes stay
fire-and-forget — the wait would add a device round-trip per write to
high-rate supervisory loops.
- **Best-effort correlation.** The callback carries only
`(hItem, statuses)` — no transaction id — so concurrent writes to the same
item within the window can swap rows; benign for the serialized single-write
consumer contract.
- **Client cancellation needs no special path**: a caller abandoning the RPC
mid-wait leaves the worker to finish its bounded wait and reply; the gateway
discards the reply, the session is never faulted.
## Later Revisit Items
These are explicit post-v1 revisit items, not open blockers:
+10 -1
View File
@@ -114,6 +114,7 @@ launch CWD (SEC-01, SEC-33).
| `MxGateway:Worker:StartupProbeRetryAttempts` | `3` | Number of retry attempts for transient worker startup probe failures before pipe connection and handshake continue. |
| `MxGateway:Worker:StartupProbeRetryDelayMilliseconds` | `250` | Delay between transient startup probe retry attempts. |
| `MxGateway:Worker:PipeConnectAttemptTimeoutMilliseconds` | `2000` | Per-attempt timeout used by the worker named-pipe connect retry path. The overall pipe connection still stays under the startup budget. |
| `MxGateway:Worker:WriteCompletionWaitMilliseconds` | `1500` | Bounded wait the worker holds a `WriteSecured`/`WriteSecured2` reply for the matching MXAccess `OnWriteComplete` callback, so the reply's `statuses` carry the real commit outcome. `0` disables the wait (pure fire-and-forget replies). Must be `>= 0`. The gateway conveys the value to the worker via the `MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS` environment variable. Consumers that time their own writes must budget above this wait: OtOpcUa's GalaxyDriver wraps gateway writes in a 2 s Tier A resilience timeout, so a deployment raising this option past ~2000 must raise that driver `ResilienceConfig` write timeout in step or slow-but-successful commits surface as consumer-side failures. |
| `MxGateway:Worker:ShutdownTimeoutSeconds` | `10` | Grace period for worker shutdown before the gateway treats shutdown as failed and may kill the worker process tree. |
| `MxGateway:Worker:HeartbeatIntervalSeconds` | `5` | Worker heartbeat send interval and gateway heartbeat check cadence input. |
| `MxGateway:Worker:HeartbeatGraceSeconds` | `15` | Maximum age of the last worker heartbeat before the gateway faults the worker. This must be greater than or equal to `HeartbeatIntervalSeconds`. |
@@ -249,7 +250,7 @@ dev/test GLAuth posture (`glauth.md`), not a production posture.
| `MxGateway:Ldap:AllowInsecure` | `true` | Permits a plaintext bind. Must be `true` when `Transport` is `None`; set `false` (with `Ldaps`/`StartTls`) in production. |
| `MxGateway:Ldap:SearchBase` | `dc=zb,dc=local` | Search base DN. |
| `MxGateway:Ldap:ServiceAccountDn` | `cn=serviceaccount,dc=zb,dc=local` | Bind DN for the search account. |
| `MxGateway:Ldap:ServiceAccountPassword` | `${secret:ldap/mxgateway/bind}` | Search-account password. **No longer a committed plaintext value:** `appsettings.json` ships the reference `${secret:ldap/mxgateway/bind}`, which the pre-host `${secret:}` expander resolves at startup from the encrypted secrets store (the code-side design default is now blank, so a missing/unresolved value fails closed rather than falling back to a leaked credential). Seed the value once with `secret set ldap/mxgateway/bind <value>` (the store's master key must be present via `ZB_SECRETS_MASTER_KEY`); startup aborts with `SecretNotFoundException` if the secret is absent. An operator may instead override it directly with the env var `MxGateway__Ldap__ServiceAccountPassword` (double-underscore form) — a plain literal there is used as-is and the secret lookup is skipped. |
| `MxGateway:Ldap:ServiceAccountPassword` | `${secret:ldap/mxgateway/bind}` | Search-account password. **Never a committed plaintext value (SEC-36):** the shared GLAuth bind credential is supplied out-of-band through one of three channels, all binding to this key. **(1) Encrypted secrets store (shipped default):** `appsettings.json` ships the reference `${secret:ldap/mxgateway/bind}`, which the pre-host `${secret:}` expander resolves at startup from the encrypted secrets store (the code-side design default is blank, so a missing/unresolved value fails closed rather than falling back to a leaked credential). Seed it once with `secret set ldap/mxgateway/bind <value>` (the store's master key must be present via `ZB_SECRETS_MASTER_KEY`); startup aborts with `SecretNotFoundException` if the secret is absent. **(2) Deployed hosts — env var:** override directly with `MxGateway__Ldap__ServiceAccountPassword` (double-underscore form) in the NSSM service environment — a plain literal there is used as-is and the store lookup is skipped. **(3) Dev boxes — user-secrets:** `dotnet user-secrets set "MxGateway:Ldap:ServiceAccountPassword" <value>` (the server carries `<UserSecretsId>mxaccessgw-server</UserSecretsId>`; user-secrets load automatically in the Development environment and live under the user profile, outside the tree). The value comes from the GLAuth source of truth `scadaproj/infra/glauth/`, never from a repo file. **Rotation:** because the credential was historically committed, it was rotated in `scadaproj/infra/glauth/` (and the shared GLAuth on `10.100.0.35` redeployed) on 2026-08-07 (SEC-36) — see `docs/runbooks/SEC-36-ldap-credential-rotation.md` for the cutover procedure used, and for future rotations. A blank/unresolved value fails startup validation with a message naming the two supported channels. |
| `MxGateway:Ldap:UserNameAttribute` | `cn` | LDAP attribute holding the login user name. |
| `MxGateway:Ldap:DisplayNameAttribute` | `cn` | LDAP attribute holding the display name. |
| `MxGateway:Ldap:GroupAttribute` | `memberOf` | LDAP attribute enumerating group membership (mapped to dashboard roles via `MxGateway:Dashboard:GroupToRole`). |
@@ -273,6 +274,14 @@ staging rig, e.g. one pointed at the plaintext shared GLAuth). A production-like
deployment must therefore run with the literal `Production` environment name for
the hard-stops to apply.
`windev` (`10.100.0.48`) is deliberately labelled `Staging` (its NSSM
`DOTNET_ENVIRONMENT` entry, set 2026-08-07) rather than left at the `Production`
default. It is the permissive rig the parenthesis above describes: it runs
`Dashboard:DisableLogin=true` and binds the shared GLAuth, which offers no TLS,
so a `Production` label would contradict its own configuration and both
hard-stops would refuse the boot. Label a host `Production` only when its
configuration can satisfy them.
## Secrets Master Key
`${secret:...}` tokens in configuration — currently just
+78 -26
View File
@@ -215,13 +215,21 @@ service described in `glauth.md`.
The suite builds the authenticator with `GatewayOptions.Dashboard.GroupToRole`
set to `{ GwAdmin: Admin }`. `GwAdmin` is the gateway-specific
dashboard-admin role and is **not** part of the five baseline GLAuth role
dashboard-admin role and is **not** part of the baseline GLAuth role
groups — it must be provisioned before the LDAP live tests pass.
`AuthenticateAsync_AdminInGwAdminGroup_Succeeds` fails (rather than skips)
when GLAuth has only the baseline groups, so this is a hard prerequisite
beyond "LDAP is up." See the "Adding a gw-specific group" section of
`glauth.md` for the provisioning step that adds `GwAdmin` and grants it to
`admin`.
beyond "LDAP is up." The shared directory
(`scadaproj/infra/glauth/config.toml`) already provisions `GwAdmin` (gid 5610)
and `GwReader` (gid 5611); see the "Adding a gw-specific group" section of
`glauth.md` for the per-box equivalent.
The fixtures name real users from that shared config, so a run only proves the
service-account bind when it targets the shared directory. `appsettings.json`
ships `Server=localhost` for the local-forward case, so point the suite at the
shared GLAuth with `MxGateway__Ldap__Server=10.100.0.35`; the suite's
`AddEnvironmentVariables()` layer applies the override to the same
`MxGateway:Ldap` section production binds.
`DashboardAuthenticator` delegates the LDAP bind and group search to the shared
`ZB.MOM.WW.Auth.Ldap` provider (`LdapAuthService`) and only maps the resulting
@@ -229,26 +237,33 @@ groups to dashboard roles via `DashboardGroupRoleMapper`; the bind/search
mechanics that decide each outcome live in that shared provider, not in
`DashboardAuthenticator`.
The suite covers both the success path and the failure outcomes: `admin` whose
LDAP groups resolve to the `Admin` role succeeds and emits the role claim;
`readonly` is denied because no group in their `memberOf` appears in
`GroupToRole`; `admin` with a wrong password fails authentication without leaking
the password into `FailureMessage`; an unknown username fails authentication; and
an unreachable LDAP server is absorbed into a failed result rather than throwing.
The suite covers both the success path and the failure outcomes: `admin`, whose
`othergroups` include `GwAdmin`, succeeds and emits the role claim — this is the
one test that proves the service-account bind, because every other outcome below
fails identically whether or not the bind credential is right; `gw-viewer` is
denied because its only group (`GwReader`) is absent from `GroupToRole`, and its
denial message must match the unknown-user denial so an authorization failure
cannot be used to enumerate valid accounts; `admin` with a wrong password fails
authentication without leaking the password into `FailureMessage`; an unknown
username fails authentication; and an unreachable LDAP server is absorbed into a
failed result rather than throwing. Both live users bind with the shared dev
password documented in `glauth.md`.
`appsettings.json` now ships the LDAP bind password as the unexpanded
`${secret:ldap/mxgateway/bind}` token (resolved at gateway startup by the
pre-host secrets expander, which this suite's bare `ConfigurationBuilder`
does not run). Before running the live LDAP suite, set
`MxGateway__Ldap__ServiceAccountPassword` to the real GLAuth service-account
password (dev value `serviceaccount123` for the shared GLAuth) so the suite
binds with the real password instead of the literal token.
password so the suite binds with the real password instead of the literal
token. Obtain the current value from the GLAuth source of truth
`scadaproj/infra/glauth/` (per `glauth.md`); it is not committed here.
Run the LDAP live tests explicitly:
```bash
$env:MXGATEWAY_RUN_LIVE_LDAP_TESTS = "1"
$env:MxGateway__Ldap__ServiceAccountPassword = "serviceaccount123"
$env:MxGateway__Ldap__Server = "10.100.0.35"
$env:MxGateway__Ldap__ServiceAccountPassword = "<service-account-password>"
dotnet test src/ZB.MOM.WW.MxGateway.IntegrationTests/ZB.MOM.WW.MxGateway.IntegrationTests.csproj --filter FullyQualifiedName~DashboardLdapLiveTests
```
@@ -428,10 +443,13 @@ runtime because the x86 Worker cannot build on Linux:
gateway fake-worker tests, and builds/tests the clients that run on Linux: .NET client
build, Go (`gofmt` + `go build` + `go test`), Rust (`cargo fmt --check` + `cargo test`
+ `cargo clippy -D warnings`), and Python (`pytest`).
- **`java`** (Linux, JDK 17) — `gradle test`. The protobuf gradle plugin rewrites
`MxaccessGateway.java` with spurious protobuf-runtime-version churn on every build, so
when no `.proto` changed the job reverts that one file (`git checkout`) before asserting
the generated tree is clean. The dev Mac has no JRE, so Java verification is CI-only.
- **`java`** (Linux, JDK 17) — `gradle test`, then `git diff --exit-code` over the generated
tree. The grpc/protobuf toolchain is pinned (`clients/java/build.gradle`), so a regeneration
is byte-identical to the committed single-file aggregates modulo real `.proto` changes; the
git-diff is therefore a true drift gate that now catches message-level proto drift in the Java
client (IPC-24 deleted the old unconditional churn-revert step, which masked exactly that
class). The dev Mac has a Homebrew JDK 17, so `gradle generateProto` can be run there to refresh
the Java aggregates when a `.proto` changes.
- **`windows-x86`** (Linux runner, per push/PR) — builds the **x86 / net48 Worker and
Worker.Tests**, which are Windows-only and out of scope for the Linux jobs. It runs on a
Linux runner that always schedules and SSHes to windev (`10.100.0.48`), where
@@ -449,21 +467,55 @@ runtime because the x86 Worker cannot build on Linux:
`xUnit1030`). Gated `if: github.event_name == 'schedule'`, so it never gates a push. On
failure it opens a Gitea issue via the Actions token, since nobody watches the Actions page.
The freshness guard `scripts/check-codegen.ps1` fails the build when the committed
client descriptor set or the C# `Generated/` no longer matches the current `.proto`
sources — the codegen drift class this repo has hit repeatedly (stale client
descriptors, net48 `CS0246` on unregenerated protos). The **primary** guard for the
### Runner capacity is shared and finite
CI runs on two co-located runner containers on docker host `10.100.0.35``gitea-runner`
(capacity 4) and `gitea-runner-2` (capacity 2, registered 2026-08-07 per
`docs/runbooks/TST-30-second-ci-runner.md`) — and both runner instances are **shared across
repos**: they interleave `dohertj2/mxaccessgw` and `dohertj2/lmxopcua` jobs across the
combined slots rather than being scoped to this repo (`GET
/repos/dohertj2/mxaccessgw/actions/runners` returns `total_count: 0`; both runners are
registered at the instance level). Every job in a run (`portable`, `java`, `windows-x86`)
still executes serially within that run, so queue latency is additive within a run, but an
active `lmxopcua` run no longer blocks `mxaccessgw` entirely the way a single shared slot
did — the two runners relieve cross-repo contention. This Gitea version (1.26) also exposes
**no run cancel or delete via the API** (`POST .../actions/runs/{id}/cancel` returns 404,
`DELETE .../actions/runs/{id}` returns 400), so a superseded or hung run cannot be cleared
and holds the slot until it finishes or times out — with two runners this means a single
wedged run can still hold slots, because the no-cancel reality is unchanged. See
`docs/runbooks/TST-30-second-ci-runner.md` for the operator runbook that registered the
second runner.
When queue depth (or the missing-cancel reality) makes waiting impractical, verify a
specific commit out of band instead of waiting behind the queue: run
`CI_SHA=<sha> scripts/ci/run-windev-ci.sh <build|test|live>` from a machine with SSH access
to windev (the same script the SSH-driven `windows-x86`/`nightly-windev` jobs use — see
`scripts/ci/README.md`), or fall back to the manual windev worktree procedure below. This is
the same escape hatch used when the windev tier itself is down — TST-30 generalizes it from
"tier down" to "runner contended": either way, a stuck or slow shared runner should not
block verifying a commit.
The freshness guard `scripts/check-codegen.ps1` runs four checks and fails the build when the
committed client descriptor set (Check 1), the C# `Generated/` (Check 2), the Rust vendored
protos (Check 3), or the Go/Python client bindings (Check 4, IPC-25) no longer match the current
`.proto` sources — the codegen drift class this repo has hit repeatedly (stale client
descriptors, net48 `CS0246` on unregenerated protos, silently stale Go/Python worker bindings).
Check 4 regenerates the Go and Python bindings with their pinned generators (`protoc-gen-go`
v1.36.11 / `protoc-gen-go-grpc` 1.6.2, `grpcio-tools` 1.80.0) and fails on any diff; a missing
generator fails the check rather than skipping it. The **primary** guard for the
"regenerate and commit `Generated/`" rule is that regeneration diff in the `portable` job;
the `windows-x86` net48 compile is the **secondary** guard (a stale `Generated/` also breaks
the x86 build with `CS0246`). See [Client Proto Generation](./ClientProtoGeneration.md) and
[Contracts](./Contracts.md).
If the SSH-driven Windows tier is unavailable for infrastructure reasons (windev down, CI
key/secret rotation in flight), fall back to the manual windev worktree procedure as a
degraded mode: on windev, fast-forward an isolated `origin/main` worktree under `C:\build`
(never the dirty Desktop checkout), then run the x86 Worker build and `Worker.Tests`
(`-p:Platform=x86`) there by hand. Do this per merge for worker-touching changes until the
`windows-x86` job is green again.
key/secret rotation in flight) **or** the shared Gitea runner is contended and the queue is
impractical to wait behind (see "Runner capacity is shared and finite" above), fall back to
the manual windev worktree procedure as a degraded mode: on windev, fast-forward an isolated
`origin/main` worktree under `C:\build` (never the dirty Desktop checkout), then run the x86
Worker build and `Worker.Tests` (`-p:Platform=x86`) there by hand. Do this per merge for
worker-touching changes until the `windows-x86` job is green again (tier-down case) or the
queue clears (contention case).
## Related Documentation
@@ -0,0 +1,169 @@
# WriteSecured Completion Correlation — Design
Date: 2026-08-09
Requested by: OtOpcUa (archreview finding 06/S-1, cross-repo residual — gateway half)
Status: approved for implementation (peer contract confirmed over cross-session message)
## Problem
The unary `Invoke` reply for a `WriteSecured` / `WriteSecured2` command proves worker-side
command *acceptance* only. MXAccess writes are fire-and-forget at the toolkit level: the
real per-item outcome arrives later in the `OnWriteComplete` COM callback. Today the worker
returns `CreateOkReply` immediately after the COM call, so `MxCommandReply.statuses` is
always empty and consumers (OtOpcUa's `GatewayGalaxyDataWriter.TranslateReply`) must treat
every write as provisionally good (`galaxy.writes.unconfirmed` meter).
Consumer contract (already shipped OtOpcUa-side): `statuses.Count > 0` → map `statuses[0]`
through the MX status map to a real OPC UA StatusCode; empty → provisional Good +
unconfirmed meter. The `statuses` field already exists on `MxCommandReply` (field 7); no
proto shape change is needed.
## Approaches Considered
1. **Executor pump-wait + versioned completion cache (chosen).** After the `WriteSecured`
COM call, the executor holds the STA thread but explicitly pumps Windows messages each
poll iteration until the matching completion is recorded or a bounded deadline passes.
This is exactly the shipped `ReadBulk` pattern (`MxAccessValueCache.TryWaitForUpdate` +
`pumpStep``StaRuntime.PumpPendingMessages()`), so it adds no new threading model.
2. **Parked asynchronous reply.** Executor returns a "pending" marker; the dispatcher parks
the correlation and the pipe reply is written later from the event dispatch. Rejected:
changes the `IStaCommandExecutor`/dispatcher/pipe contracts for no real gain — commands
serialize per session anyway, so freeing the STA during the wait buys nothing.
3. **Gateway-side correlation.** Gateway watches the session event stream for the
`OnWriteComplete` after the worker reply. Rejected: races the event drain cadence,
couples the gateway to event semantics, still holds the unary RPC, and spreads the
feature across two processes.
## Design
All changes are worker-side (`ZB.MOM.WW.MxGateway.Worker`, net48 x86). The gateway's
`Invoke` already forwards the worker `MxCommandReply` (statuses included) verbatim.
### New: `MxAccessWriteCompletionCache`
Mirror of `MxAccessValueCache`, keyed by `(serverHandle, itemHandle)` (packed long), one
entry per key holding the most recent completion's `RepeatedField<MxStatusProxy>` (cloned)
plus a monotonically increasing per-key `Version`. API:
- `Record(int serverHandle, int itemHandle, RepeatedField<MxStatusProxy> statuses)`
- `ulong CurrentVersion(int serverHandle, int itemHandle)` — 0 when absent
- `bool TryWaitForCompletion(int serverHandle, int itemHandle, ulong sinceVersion,
DateTime deadlineUtc, Action pumpStep, out RepeatedField<MxStatusProxy> statuses,
int pollIntervalMs = 5)` — pump/poll loop identical in shape to
`MxAccessValueCache.TryWaitForUpdate`.
Same locking posture as the value cache: everything runs on the STA thread; a sync root
keeps it nominally thread-safe for tests.
### Sink: record completions
`MxAccessBaseEventSink` owns a `MxAccessWriteCompletionCache` (new optional ctor param,
exposed as a property) and its `OnWriteComplete` handler records into it via the existing
`EnqueueEvent` post-publish hook (same pattern as the value cache on `OnDataChange`):
the streamed `MxEvent` is built exactly once by the mapper, enqueued unchanged for the
event stream, and its `Statuses` are then recorded into the cache. The event stream is
not altered — nothing is swallowed or synthesized.
A new seam interface `IWriteCompletionCacheProvider { MxAccessWriteCompletionCache
WriteCompletionCache { get; } }` is implemented by `MxAccessBaseEventSink` and by test
sinks. `MxAccessSession.Create` pulls the cache from the sink through that interface
(fallback: fresh instance), mirroring the existing `ValueCache` sharing, and exposes it
on the session.
### Executor: bounded pump-wait
In `ExecuteWriteSecured` / `ExecuteWriteSecured2`:
1. Capture `baseline = cache.CurrentVersion(serverHandle, itemHandle)` **before** the COM
call — this closes the fast-completion ordering edge: a callback that dispatches during
or immediately after the COM call bumps the version past the baseline and still
correlates.
2. Call `session.WriteSecured(...)` as today.
3. `cache.TryWaitForCompletion(..., baseline, deadline, pumpStep, out statuses)`; on
success, `reply.Statuses.Add(statuses)`; on timeout, return the reply exactly as today
(protocol OK, empty statuses) — the consumer's honest-unconfirmed path. No invented
failure rows.
The reply's `ProtocolStatus`/`Hresult` stay untouched by the completion outcome: the
command was accepted; the MX outcome (success *or* failure) is carried only in
`statuses[0]`. That preserves MXAccess parity (the native API returns void; the outcome
exists only in the callback) while enriching the reply with information that was already
on the wire.
Timeout default: **1.5 s** (`MxAccessCommandExecutor.DefaultWriteCompletionTimeout`).
The consumer-side budget drives this: OtOpcUa wraps the driver write in a Tier A
resilience policy with a 2 s timeout and a 5-failure breaker, so a worker wait longer
than 2 s would convert slow-but-successful commits into consumer-side false failures
(node revert + breaker pressure). 1.5 s covers the common fast-commit case and degrades
a slow commit to the honest-unconfirmed path instead. It also sits well under the
gateway's 30 s `DefaultCommandTimeoutSeconds` IPC wait and the STA watchdog's 75 s
dispatched-command ceiling.
The wait is deployment-configurable end to end, following the existing
pipe-connect-timeout pattern: a new gateway option
`MxGateway:Worker:WriteCompletionWaitMilliseconds` (default `1500`, validated `>= 0`;
`0` disables the wait and restores pure fire-and-forget replies) is exported by
`WorkerProcessLauncher` as the `MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS` environment
variable, which the worker reads at session construction. A deployment that raises the
gateway wait must raise the OtOpcUa driver's Write `ResilienceConfig` timeout in step
(per-instance operator config on the OtOpcUa side) — documented in
`docs/GatewayConfiguration.md`. `MxAccessStaSession` additionally gets an internal
`WriteCompletionTimeout` seam (read when it constructs the executor in `StartAsync`) so
tests can shorten it without env plumbing.
Client cancellation mid-wait needs no new code path, only verification: when the caller
cancels the unary RPC, the gateway abandons its IPC reply wait, but the worker command
is already in flight — `CancelCommand` only dequeues *queued* commands. The executor
simply finishes its bounded wait and replies; the gateway discards the reply. The
session is never faulted and the `OnWriteComplete` event still flows on the stream.
### Scope
- **In**: `WriteSecured`, `WriteSecured2` (single-item secured writes — inherently
low-rate operator actions, and exactly the OtOpcUa single-write contract). Default-on.
- **Out**: plain `Write`/`Write2` and all bulk write commands stay fire-and-forget —
waiting would add a device round-trip of latency to high-rate supervisory write loops.
- **Correlation fidelity is best-effort**: the MXAccess callback carries only
`(hItem, statuses)` — no transaction id — so a concurrent write to the same item within
the wait window can be attributed to the wrong writer (worst case two writes to the
same item swap status rows — benign for the serialized single-write consumer
contract). Documented on the proto field.
## Error handling
- Completion never arrives (device down): bounded 1.5 s wait, then today's reply shape.
- Event queue overflow during completion: the queue records a fault and the fail-fast
design tears the session down; the post-publish hook not firing in that case is moot.
- COM call throws: unchanged — the dispatcher's existing exception path replies with the
native HResult; no wait is entered.
## Testing (Worker.Tests, x86 — verified on windev)
- `MxAccessWriteCompletionCacheTests`: record/version monotonicity, wait success,
deadline expiry, baseline-before-record fast-completion ordering.
- `MxAccessBaseEventSinkTests`: `OnWriteComplete` both enqueues the event *and* records
the completion; cache instance is the sink-bound one.
- `MxAccessCommandExecutorTests` (via `MxAccessStaSession.DispatchAsync` + fake COM
object + test sink implementing `IWriteCompletionCacheProvider`):
- completion recorded synchronously inside the fake's `WriteSecured` (fast edge) →
reply carries `statuses[0]`;
- completion recorded from the test thread while the executor pump-waits → reply
carries `statuses[0]`;
- no completion + shortened timeout → protocol OK, empty statuses;
- `WriteSecured2` mirrors; plain `Write` does not wait.
- Gateway tests (macOS-runnable): `GatewayOptionsValidator` accepts `>= 0` and rejects
negative `WriteCompletionWaitMilliseconds`; `WorkerProcessLauncher` exports
`MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS` from the option (mirroring the existing
pipe-connect-timeout env assertion).
## Docs in the same change
- `mxaccess_gateway.proto`: comment on `MxCommandReply.statuses` and the
`WriteSecuredCommand`/`WriteSecured2Command` messages describing the correlated
completion contract (populated within the wait window; empty = unconfirmed;
best-effort correlation). Comment-only → wire-identical; regenerate + commit
`Contracts/Generated` (required for the C# build); other clients' generated code is
functionally unchanged.
- `docs/GatewayConfiguration.md`: the new `MxGateway:Worker:WriteCompletionWaitMilliseconds`
option, including the pairing rule with the OtOpcUa driver's Write resilience timeout.
- `gateway.md` command/event surface note; `docs/DesignDecisions.md` entry.
@@ -0,0 +1,575 @@
# WriteSecured Completion Correlation Implementation Plan
> **For Claude:** REQUIRED SUB-SKILL: Use superpowers-extended-cc:executing-plans to implement this plan task-by-task.
**Goal:** Populate `MxCommandReply.statuses[0]` on unary `WriteSecured`/`WriteSecured2` replies with the correlated MXAccess `OnWriteComplete` outcome, bounded by a configurable wait (default 1.5 s), falling back to today's empty-statuses shape on timeout.
**Architecture:** Worker-side only (plus one gateway config option). A versioned per-`(serverHandle, itemHandle)` completion cache (mirror of `MxAccessValueCache`) is populated by the event sink's `OnWriteComplete` post-publish hook; the STA command executor captures a version baseline before the COM call, then pump-waits (ReadBulk precedent) until a newer completion lands or the deadline passes. Design: `docs/plans/2026-08-09-write-completion-correlation-design.md`.
**Tech Stack:** .NET Framework 4.8 x86 worker (no init-only props/positional records!), .NET 10 gateway, protobuf via Grpc.Tools regen. Worker builds/tests run ONLY on windev (10.100.0.48) — local macOS verification covers the gateway + contracts.
---
### Task 0: Create feature branch
**Classification:** trivial
**Estimated implement time:** ~1 min
**Parallelizable with:** none
```bash
cd /Users/dohertj2/Desktop/MxAccessGateway && git checkout -b feat/write-completion-correlation
```
### Task 1: Proto contract comments + regen
**Classification:** small
**Estimated implement time:** ~4 min
**Parallelizable with:** none (everything builds on the regenerated contracts)
**Files:**
- Modify: `src/ZB.MOM.WW.MxGateway.Contracts/Protos/mxaccess_gateway.proto` (~line 528 `statuses`, ~line 259 `WriteSecuredCommand`, ~line 269 `WriteSecured2Command`)
- Regenerate: `src/ZB.MOM.WW.MxGateway.Contracts/Generated/*.cs`
**Step 1:** On `repeated MxStatusProxy statuses = 7;` in `MxCommandReply`, add above the field:
```proto
// Correlated per-item outcome rows. For WRITE_SECURED / WRITE_SECURED2
// replies the worker holds the reply for a bounded window (default 1.5 s,
// MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS) waiting for the matching
// MXAccess OnWriteComplete callback and copies its status rows here, so
// statuses[0] carries the real MXAccess commit outcome (success OR failure)
// while protocol_status/hresult still describe command acceptance only.
// Empty statuses on a write reply means the completion did not arrive
// within the window — the write is unconfirmed, not failed. Correlation is
// best-effort per (server_handle, item_handle): MXAccess's callback carries
// no transaction id, so concurrent writes to the same item within the
// window can swap rows. The OnWriteComplete event still flows on the event
// stream unchanged. Other command kinds leave this field as before.
```
**Step 2:** On `message WriteSecuredCommand` and `message WriteSecured2Command`, append to the existing leading comment (or add one): `// The unary reply's statuses field carries the correlated OnWriteComplete outcome when it arrives within the worker's bounded wait — see MxCommandReply.statuses.`
**Step 3:** Regenerate + verify wire-identical build:
```bash
rm src/ZB.MOM.WW.MxGateway.Contracts/Generated/*.cs
dotnet build src/ZB.MOM.WW.MxGateway.Contracts/ZB.MOM.WW.MxGateway.Contracts.csproj
```
Expected: build succeeds, `git diff --stat` shows only comment-churn in Generated.
**Step 4:** Commit: `git add src/ZB.MOM.WW.MxGateway.Contracts/Protos/mxaccess_gateway.proto src/ZB.MOM.WW.MxGateway.Contracts/Generated && git commit -m "docs(proto): document the correlated write-completion statuses contract"`
(Comment-only proto change is wire-identical; other clients' generated code is intentionally not regenerated — no functional delta.)
### Task 2: MxAccessWriteCompletionCache + tests
**Classification:** standard
**Estimated implement time:** ~5 min
**Parallelizable with:** Task 7
**Files:**
- Create: `src/ZB.MOM.WW.MxGateway.Worker/MxAccess/MxAccessWriteCompletionCache.cs`
- Create: `src/ZB.MOM.WW.MxGateway.Worker.Tests/MxAccess/MxAccessWriteCompletionCacheTests.cs`
**Step 1:** Create the cache — mirror `MxAccessValueCache`'s shape, locking, and net48 constraints (no init-only, plain struct/class):
```csharp
using System;
using System.Collections.Generic;
using System.Threading;
using Google.Protobuf.Collections;
using ZB.MOM.WW.MxGateway.Contracts.Proto;
namespace ZB.MOM.WW.MxGateway.Worker.MxAccess;
/// <summary>
/// Per-session cache of the most recent <c>OnWriteComplete</c> status rows
/// for each (server handle, item handle) pair. Written by the MXAccess
/// event sink as completion callbacks arrive; read by the write command
/// executor so a WriteSecured/WriteSecured2 reply can carry the correlated
/// MXAccess outcome instead of proving command acceptance only.
/// </summary>
/// <remarks>
/// Same threading posture as <see cref="MxAccessValueCache"/>: writers and
/// readers run on the worker's STA thread (COM dispatches events on the
/// apartment thread; commands also execute on the STA), so no internal
/// locking is required. A single sync root keeps it nominally thread-safe
/// for tests that drive it from a non-STA thread.
/// </remarks>
public sealed class MxAccessWriteCompletionCache
{
private readonly Dictionary<long, CompletionEntry> entries = new();
private readonly object syncRoot = new();
/// <summary>Records the status rows of a fresh OnWriteComplete callback for the given handle pair.</summary>
/// <param name="serverHandle">MXAccess server handle.</param>
/// <param name="itemHandle">MXAccess item handle.</param>
/// <param name="statuses">Status rows from the mapped OnWriteComplete event; cloned before storing.</param>
public void Record(
int serverHandle,
int itemHandle,
RepeatedField<MxStatusProxy> statuses)
{
if (statuses is null)
{
throw new ArgumentNullException(nameof(statuses));
}
lock (syncRoot)
{
long key = CreateItemKey(serverHandle, itemHandle);
ulong version = entries.TryGetValue(key, out CompletionEntry existing)
? existing.Version + 1
: 1UL;
entries[key] = new CompletionEntry(version, statuses.Clone());
}
}
/// <summary>Returns the current completion version for a handle pair, or 0 if none was recorded.</summary>
/// <param name="serverHandle">MXAccess server handle.</param>
/// <param name="itemHandle">MXAccess item handle.</param>
/// <returns>The current completion version, or 0 if no completion was recorded.</returns>
public ulong CurrentVersion(
int serverHandle,
int itemHandle)
{
lock (syncRoot)
{
return entries.TryGetValue(CreateItemKey(serverHandle, itemHandle), out CompletionEntry existing)
? existing.Version
: 0UL;
}
}
/// <summary>
/// Polls for a completion newer than <paramref name="sinceVersion"/> until it
/// arrives or the deadline elapses, calling <paramref name="pumpStep"/> on every
/// poll iteration so the worker's STA can dispatch the inbound MXAccess
/// OnWriteComplete message. Same loop shape as
/// <see cref="MxAccessValueCache.TryWaitForUpdate"/>.
/// </summary>
/// <param name="serverHandle">MXAccess server handle.</param>
/// <param name="itemHandle">MXAccess item handle.</param>
/// <param name="sinceVersion">Version snapshot captured before the write COM call.</param>
/// <param name="deadlineUtc">Absolute UTC deadline.</param>
/// <param name="pumpStep">Action that pumps any pending Windows messages.</param>
/// <param name="statuses">The recorded status rows if a completion arrived before the deadline.</param>
/// <param name="pollIntervalMs">How long to sleep between pump cycles. Default 5 ms.</param>
/// <returns><see langword="true"/> if a completion newer than <paramref name="sinceVersion"/> arrived before the deadline; otherwise <see langword="false"/>.</returns>
public bool TryWaitForCompletion(
int serverHandle,
int itemHandle,
ulong sinceVersion,
DateTime deadlineUtc,
Action pumpStep,
out RepeatedField<MxStatusProxy> statuses,
int pollIntervalMs = 5)
{
if (pumpStep is null)
{
throw new ArgumentNullException(nameof(pumpStep));
}
while (true)
{
pumpStep();
lock (syncRoot)
{
if (entries.TryGetValue(CreateItemKey(serverHandle, itemHandle), out CompletionEntry entry)
&& entry.Version > sinceVersion)
{
statuses = entry.Statuses;
return true;
}
}
if (DateTime.UtcNow >= deadlineUtc)
{
statuses = new RepeatedField<MxStatusProxy>();
return false;
}
Thread.Sleep(pollIntervalMs);
}
}
private static long CreateItemKey(
int serverHandle,
int itemHandle)
{
return ((long)serverHandle << 32) | (uint)itemHandle;
}
/// <summary>
/// Snapshot of the most recent OnWriteComplete status rows for a handle
/// pair. <see cref="Version"/> increments by one on every
/// <see cref="Record"/> call so the write executor can detect "a new
/// completion arrived since I captured my baseline".
/// </summary>
/// <remarks>
/// Plain readonly struct (not a record) so this compiles under the
/// worker's net48 target, which lacks <c>IsExternalInit</c>.
/// </remarks>
private readonly struct CompletionEntry
{
public CompletionEntry(
ulong version,
RepeatedField<MxStatusProxy> statuses)
{
Version = version;
Statuses = statuses;
}
public ulong Version { get; }
public RepeatedField<MxStatusProxy> Statuses { get; }
}
}
```
**Step 2:** Tests (mirror `MxAccessValueCacheTests` style; build `MxStatusProxy` rows inline):
- `Record_IncrementsVersionPerKey` — two `Record` calls on the same pair → `CurrentVersion` 1 then 2; a different pair stays independent.
- `TryWaitForCompletion_WhenCompletionNewerThanBaseline_ReturnsStatuses` — record once, wait with `sinceVersion: 0`, deadline in the future → `true`, statuses round-trip (assert an `MxStatusProxy` field value survives the clone).
- `TryWaitForCompletion_WhenOnlyStaleCompletion_TimesOut` — record once, wait with `sinceVersion: CurrentVersion(...)` and a deadline ~50 ms out → `false`, out statuses empty.
- `TryWaitForCompletion_InvokesPumpStepEachIteration` — pumpStep increments a counter; on the counter's second call, `Record` the completion (this proves the pump loop is what lets the callback land); assert `true` and counter >= 2.
- `Record_ClonesStatuses` — mutate the caller's `RepeatedField` after `Record`; waited-out statuses unaffected.
**Step 3:** Cannot compile locally (worker is Windows-only) — defer build/test to Task 9 (windev). Commit: `git add src/ZB.MOM.WW.MxGateway.Worker/MxAccess/MxAccessWriteCompletionCache.cs src/ZB.MOM.WW.MxGateway.Worker.Tests/MxAccess/MxAccessWriteCompletionCacheTests.cs && git commit -m "feat(worker): versioned OnWriteComplete completion cache"`
### Task 3: Sink records completions (+ provider seam)
**Classification:** standard
**Estimated implement time:** ~4 min
**Parallelizable with:** Task 7
**Files:**
- Create: `src/ZB.MOM.WW.MxGateway.Worker/MxAccess/IWriteCompletionCacheProvider.cs`
- Modify: `src/ZB.MOM.WW.MxGateway.Worker/MxAccess/MxAccessBaseEventSink.cs`
- Test: `src/ZB.MOM.WW.MxGateway.Worker.Tests/MxAccess/MxAccessBaseEventSinkTests.cs`
**Step 1:** New seam interface:
```csharp
namespace ZB.MOM.WW.MxGateway.Worker.MxAccess;
/// <summary>
/// Exposes the per-session <see cref="MxAccessWriteCompletionCache"/> an
/// event sink populates from OnWriteComplete callbacks, so
/// <see cref="MxAccessSession.Create"/> can share one instance between the
/// sink (writer) and the write command executor (reader). Implemented by
/// <see cref="MxAccessBaseEventSink"/> and by test sinks that cannot
/// attach to a live MXAccess COM object.
/// </summary>
public interface IWriteCompletionCacheProvider
{
/// <summary>The completion cache bound to this sink.</summary>
MxAccessWriteCompletionCache WriteCompletionCache { get; }
}
```
**Step 2:** `MxAccessBaseEventSink` — declare `IWriteCompletionCacheProvider` on the class; add a `private readonly MxAccessWriteCompletionCache writeCompletionCache;` initialized in the widest ctor (add a new optional-most ctor overload following the existing chain pattern: the 3-arg `(eventQueue, eventMapper, valueCache)` ctor chains to a new 4-arg `(eventQueue, eventMapper, valueCache, writeCompletionCache)` with a fresh cache); expose `public MxAccessWriteCompletionCache WriteCompletionCache => writeCompletionCache;`. Change `OnWriteComplete` to use the post-publish hook (same pattern as `OnDataChange`'s value-cache publish — post-publish only runs after the event cleared the queue, and a queue overflow faults the session anyway):
```csharp
MXSTATUS_PROXY[] statuses = pVars;
EnqueueEvent(
() => eventMapper.CreateOnWriteComplete(
sessionId,
hLMXServerHandle,
phItemHandle,
statuses),
mxEvent => writeCompletionCache.Record(hLMXServerHandle, phItemHandle, mxEvent.Statuses));
```
**Step 3:** Tests in `MxAccessBaseEventSinkTests` (mirror `OnDataChange_ComCallback_PopulatesValueCache` and `ValueCache_ReturnsTheInstanceBoundAtConstruction`):
- `OnWriteComplete_ComCallback_RecordsCompletionAndStillEnqueuesEvent` — drive `sink.OnWriteComplete(7, 21, ref proxies)`; assert the queue got the OnWriteComplete event AND `cache.CurrentVersion(7, 21) == 1`.
- `WriteCompletionCache_ReturnsTheInstanceBoundAtConstruction`.
**Step 4:** Commit: `git commit -m "feat(worker): event sink records OnWriteComplete rows into the completion cache"` (explicit paths).
### Task 4: MxAccessSession plumbing
**Classification:** small
**Estimated implement time:** ~3 min
**Parallelizable with:** Task 7
**Files:**
- Modify: `src/ZB.MOM.WW.MxGateway.Worker/MxAccess/MxAccessSession.cs` (private ctor ~line 18, `Create` ~line 145)
**Step 1:** Add `private readonly MxAccessWriteCompletionCache writeCompletionCache;` + ctor param (after `valueCache`) + null guard; add property mirroring `ValueCache`:
```csharp
/// <summary>
/// Per-session OnWriteComplete completion cache populated by the event
/// sink. The write command executor consults it after a
/// WriteSecured/WriteSecured2 COM call so the unary reply can carry
/// the correlated completion outcome.
/// </summary>
public MxAccessWriteCompletionCache WriteCompletionCache => writeCompletionCache;
```
**Step 2:** In `Create`, next to the value-cache sharing block:
```csharp
// Share the sink's completion cache the same way (production sink
// and completion-aware test sinks implement the provider seam);
// fall back to a fresh cache for other fakes — the write executor
// then simply never observes a completion and replies unconfirmed.
MxAccessWriteCompletionCache writeCompletionCache = eventSink is IWriteCompletionCacheProvider provider
? provider.WriteCompletionCache
: new MxAccessWriteCompletionCache();
```
Pass it to the ctor.
**Step 3:** Commit: `git commit -m "feat(worker): share the completion cache between sink and session"`.
### Task 5: Executor bounded wait + StaSession env plumbing
**Classification:** high-risk (STA/pump semantics)
**Estimated implement time:** ~5 min
**Parallelizable with:** none (touches the same files as 6's tests exercise)
**Files:**
- Modify: `src/ZB.MOM.WW.MxGateway.Worker/MxAccess/MxAccessCommandExecutor.cs` (~lines 14-85 ctors, 447-497 write methods)
- Modify: `src/ZB.MOM.WW.MxGateway.Worker/MxAccess/MxAccessStaSession.cs` (~line 205 executor construction)
**Step 1:** Executor — next to `DefaultReadBulkTimeout`:
```csharp
/// <summary>
/// Default bounded wait for the OnWriteComplete callback after a
/// WriteSecured/WriteSecured2 COM call. 1.5 s keeps the unary reply
/// inside the OtOpcUa driver's 2 s Tier A write-resilience budget (a
/// longer gateway wait must raise that consumer timeout in step) while
/// covering the common fast-commit case; on expiry the reply returns
/// with empty statuses — unconfirmed, not failed.
/// </summary>
internal static readonly TimeSpan DefaultWriteCompletionTimeout = TimeSpan.FromMilliseconds(1500);
private readonly TimeSpan writeCompletionTimeout;
```
Add `TimeSpan? writeCompletionTimeout = null` as a trailing optional parameter on the widest (4-arg) ctor; `this.writeCompletionTimeout = writeCompletionTimeout ?? DefaultWriteCompletionTimeout;`.
**Step 2:** In `ExecuteWriteSecured`, replace the tail (`session.WriteSecured(...); return CreateOkReply(command);`) with:
```csharp
MxAccessWriteCompletionCache completionCache = session.WriteCompletionCache;
// Baseline BEFORE the COM call: a completion that dispatches during or
// immediately after WriteSecured bumps the version past this snapshot,
// so a fast commit still correlates (no missed-callback window).
ulong completionBaseline = completionCache.CurrentVersion(
writeSecuredCommand.ServerHandle,
writeSecuredCommand.ItemHandle);
session.WriteSecured(
writeSecuredCommand.ServerHandle,
writeSecuredCommand.ItemHandle,
writeSecuredCommand.CurrentUserId,
writeSecuredCommand.VerifierUserId,
variantConverter.ConvertToComValue(writeSecuredCommand.Value));
MxCommandReply reply = CreateOkReply(command);
AwaitWriteCompletion(
reply,
completionCache,
writeSecuredCommand.ServerHandle,
writeSecuredCommand.ItemHandle,
completionBaseline);
return reply;
```
Mirror in `ExecuteWriteSecured2`. Shared private helper:
```csharp
/// <summary>
/// Bounded pump-wait for the OnWriteComplete row matching a
/// WriteSecured/WriteSecured2 call, copied onto the reply when it
/// arrives in time. The executor holds the STA thread but pumps
/// Windows messages each poll (ReadBulk precedent) so the COM callback
/// can dispatch re-entrantly; on expiry the reply keeps its empty
/// statuses — the consumer's unconfirmed path, never a synthesized
/// failure. Protocol status/hresult stay acceptance-only either way.
/// </summary>
private void AwaitWriteCompletion(
MxCommandReply reply,
MxAccessWriteCompletionCache completionCache,
int serverHandle,
int itemHandle,
ulong completionBaseline)
{
if (writeCompletionTimeout <= TimeSpan.Zero)
{
return;
}
if (completionCache.TryWaitForCompletion(
serverHandle,
itemHandle,
completionBaseline,
DateTime.UtcNow + writeCompletionTimeout,
pumpStep,
out Google.Protobuf.Collections.RepeatedField<MxStatusProxy> statuses))
{
reply.Statuses.Add(statuses);
}
}
```
**Step 3:** `MxAccessStaSession` — add:
```csharp
/// <summary>
/// Environment variable the gateway's WorkerProcessLauncher sets from
/// MxGateway:Worker:WriteCompletionWaitMilliseconds. 0 disables the
/// write-completion wait (pure fire-and-forget replies).
/// </summary>
internal const string WriteCompletionWaitEnvironmentVariableName =
"MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS";
/// <summary>
/// Bounded WriteSecured/WriteSecured2 completion wait handed to the
/// command executor at StartAsync. Internal-settable as a test seam so
/// Worker.Tests can shorten it without env-var plumbing.
/// </summary>
internal TimeSpan WriteCompletionTimeout { get; set; } = ResolveWriteCompletionTimeout();
internal static TimeSpan ResolveWriteCompletionTimeout()
{
string value = Environment.GetEnvironmentVariable(WriteCompletionWaitEnvironmentVariableName);
return int.TryParse(value, System.Globalization.NumberStyles.Integer, System.Globalization.CultureInfo.InvariantCulture, out int milliseconds)
&& milliseconds >= 0
? TimeSpan.FromMilliseconds(milliseconds)
: MxAccessCommandExecutor.DefaultWriteCompletionTimeout;
}
```
and pass `writeCompletionTimeout: WriteCompletionTimeout` when constructing `MxAccessCommandExecutor` in `StartAsync`. (net48: `Environment.GetEnvironmentVariable` returns `string` — keep nullable annotations consistent with the file.)
**Step 4:** Commit: `git commit -m "feat(worker): bounded pump-wait correlates OnWriteComplete onto secured-write replies"`.
### Task 6: Executor tests
**Classification:** standard
**Estimated implement time:** ~5 min
**Parallelizable with:** none (depends on Tasks 2-5)
**Files:**
- Test: `src/ZB.MOM.WW.MxGateway.Worker.Tests/MxAccess/MxAccessCommandExecutorTests.cs`
**Step 1:** Test support inside the test class file:
- New sink: `private sealed class CompletionCacheEventSink : IMxAccessEventSink, IWriteCompletionCacheProvider { public MxAccessWriteCompletionCache WriteCompletionCache { get; } = new MxAccessWriteCompletionCache(); public void Attach(object mxAccessComObject, string sessionId) { } public void Detach() { } }` (match the exact `IMxAccessEventSink` member list — read the interface first).
- `FakeMxAccessComObject`: add `public Action OnWriteSecuredCallback { get; set; }` (nullable per file convention) invoked at the end of `WriteSecured` and `WriteSecured2`.
**Step 2:** Tests (all through `MxAccessStaSession.DispatchAsync`, constructing the session with the completion sink; set `session.WriteCompletionTimeout` before `StartAsync`):
- `DispatchAsync_WriteSecured_WhenCompletionArrivesDuringComCall_ReturnsStatuses` (fast-completion edge): wire `OnWriteSecuredCallback = () => sink.WriteCompletionCache.Record(82, 821, StatusRows(1))` (helper building a `RepeatedField<MxStatusProxy>` with a recognizable value); dispatch; assert `ProtocolStatus.Code == Ok`, `reply.Statuses.Count == 1`, row round-trips.
- `DispatchAsync_WriteSecured_WhenCompletionArrivesWhileWaiting_ReturnsStatuses`: no fake callback; `WriteCompletionTimeout = TimeSpan.FromSeconds(10)`; start `Task<MxCommandReply> pending = session.DispatchAsync(...)`, then `sink.WriteCompletionCache.Record(...)` from the test thread after a short `Task.Delay(50)`; await; assert statuses present. (No `.ConfigureAwait(false)` in `[Fact]` bodies — xUnit1030 fails the Windows build.)
- `DispatchAsync_WriteSecured_WhenNoCompletion_TimesOutWithEmptyStatusesAndOkProtocol`: `WriteCompletionTimeout = TimeSpan.FromMilliseconds(100)`; assert Ok + `reply.Statuses.Count == 0`.
- `DispatchAsync_WriteSecured2_WhenCompletionArrivesDuringComCall_ReturnsStatuses` (mirror of the fast test).
- `DispatchAsync_Write_DoesNotWaitForCompletion`: `WriteCompletionTimeout = TimeSpan.FromSeconds(30)`, plain `Write` command, no completion recorded; assert `await` completes within a 5 s guard (`Task.WhenAny` with `Task.Delay`) — proving plain writes never enter the wait.
- Baseline test `DispatchAsync_WriteSecured_IgnoresStaleCompletionFromBeforeTheCall`: `Record` once BEFORE dispatch, `WriteCompletionTimeout = 100 ms`, no new completion → empty statuses (stale row not misattributed).
**Step 2b:** Confirm the two existing WriteSecured tests (`DispatchAsync_WriteSecured_ForwardsUserIds`, `..._WriteSecured2_...`) still pass unmodified — they use `NoopEventSink`, so the session falls back to a fresh cache, no completion ever arrives, and the default 1.5 s wait adds latency only; if that latency bothers the suite, switch them to the completion sink with `WriteCompletionTimeout = TimeSpan.Zero`.
**Step 3:** Commit: `git commit -m "test(worker): write-completion correlation executor coverage"`.
### Task 7: Gateway config option + launcher env var
**Classification:** standard
**Estimated implement time:** ~5 min
**Parallelizable with:** Task 2, Task 3, Task 4
**Files:**
- Modify: `src/ZB.MOM.WW.MxGateway.Server/Configuration/WorkerOptions.cs`
- Modify: `src/ZB.MOM.WW.MxGateway.Server/Configuration/GatewayOptionsValidator.cs` (~line 224 block)
- Modify: `src/ZB.MOM.WW.MxGateway.Server/Workers/WorkerProcessLauncher.cs` (~lines 18-22 consts, ~line 175 env block)
- Test: `src/ZB.MOM.WW.MxGateway.Tests/Gateway/Workers/WorkerProcessLauncherTests.cs`, the `GatewayOptionsValidator` test file (find via `grep -rl GatewayOptionsValidatorTests src/ZB.MOM.WW.MxGateway.Tests`)
**Step 1:** `WorkerOptions`:
```csharp
/// <summary>
/// Bounded wait, in milliseconds, the worker holds a WriteSecured/WriteSecured2
/// reply for the matching MXAccess OnWriteComplete callback so the reply's
/// statuses carry the real commit outcome. 0 disables the wait. Deployments
/// raising this above consumer write-timeout budgets (e.g. OtOpcUa's 2 s Tier A
/// write resilience timeout) must raise those in step.
/// </summary>
public int WriteCompletionWaitMilliseconds { get; init; } = 1500;
```
**Step 2:** Validator (>= 0, not the positive helper):
```csharp
if (options.WriteCompletionWaitMilliseconds < 0)
{
builder.Add("MxGateway:Worker:WriteCompletionWaitMilliseconds must be greater than or equal to zero.");
}
```
**Step 3:** Launcher — const `public const string WorkerWriteCompletionWaitEnvironmentVariableName = "MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS";` and in `CreateStartInfo` next to the pipe-connect env line:
```csharp
startInfo.Environment[WorkerWriteCompletionWaitEnvironmentVariableName] =
_workerOptions.WriteCompletionWaitMilliseconds.ToString(System.Globalization.CultureInfo.InvariantCulture);
```
**Step 4:** Tests — mirror the existing pipe-connect-timeout launcher env assertion and an existing validator negative test; add: launcher exports `MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS=1500` by default / custom value when configured; validator rejects `-1`, accepts `0`.
**Step 5:** Run locally:
```bash
dotnet test src/ZB.MOM.WW.MxGateway.Tests/ZB.MOM.WW.MxGateway.Tests.csproj --filter "FullyQualifiedName~WorkerProcessLauncherTests|FullyQualifiedName~GatewayOptionsValidator"
```
Expected: PASS.
**Step 6:** Commit: `git commit -m "feat(gateway): configurable worker write-completion wait (MxGateway:Worker:WriteCompletionWaitMilliseconds)"`.
### Task 8: Docs
**Classification:** small
**Estimated implement time:** ~4 min
**Parallelizable with:** none (write after code settles)
**Files:**
- Modify: `docs/GatewayConfiguration.md` (Worker options table, ~line 113 area)
- Modify: `gateway.md` (command surface / write semantics section)
- Modify: `docs/DesignDecisions.md` (new decision entry)
Content: the option row (default 1500, 0 disables, consumer-budget pairing rule with OtOpcUa's Write resilience timeout); gateway.md note that WriteSecured/WriteSecured2 unary replies now carry correlated completion statuses (bounded wait, empty = unconfirmed, event stream unchanged); DesignDecisions entry summarizing the design doc (link it) including the best-effort per-(hItem) correlation caveat and why plain Write/Write2/bulk stay fire-and-forget. Follow `docs/style-guides/StyleGuide.md` (present tense, why not what).
Commit: `git commit -m "docs: write-completion correlation configuration and semantics"`.
### Task 9: Local + windev verification
**Classification:** standard
**Estimated implement time:** ~10 min (mostly remote build time)
**Parallelizable with:** none
**Step 1 (local):** `dotnet build src/ZB.MOM.WW.MxGateway.NonWindows.slnx` → succeeds; re-run the Task 7 filtered gateway tests.
**Step 2 (push):** `git push -u origin feat/write-completion-correlation`.
**Step 3 (windev):** Use the isolated `C:\build` worktree (NOT the Desktop checkout — dirty feature branch). Remote PS via base64 `-EncodedCommand`; no `$ErrorActionPreference='Stop'` around git. Sequence:
```
git fetch origin && git checkout feat/write-completion-correlation && git pull
dotnet build src/ZB.MOM.WW.MxGateway.Worker/ZB.MOM.WW.MxGateway.Worker.csproj -p:Platform=x86
dotnet test src/ZB.MOM.WW.MxGateway.Worker.Tests/ZB.MOM.WW.MxGateway.Worker.Tests.csproj -p:Platform=x86 --filter "FullyQualifiedName~MxAccessWriteCompletionCacheTests|FullyQualifiedName~MxAccessBaseEventSinkTests|FullyQualifiedName~MxAccessCommandExecutorTests"
```
Expected: build clean (TreatWarningsAsErrors — watch xUnit1030), all filtered tests PASS. Fix-and-push iterations happen from the Mac; windev only builds/tests.
**Step 4:** Full worker test suite once green on the filter (`dotnet test ... -p:Platform=x86`, no filter) — one full pass before merge.
### Task 10: Review, merge, notify
**Classification:** standard
**Estimated implement time:** ~5 min
**Parallelizable with:** none
- Run a code review over the full branch diff (code-reviewer agent or /code-review) and address findings.
- Merge: `git checkout main && git merge --no-ff feat/write-completion-correlation && git push origin main`.
- Notify the OtOpcUa session (SendMessage) that the feature is on `main` and built/tested on windev; flag that the live ArchestrA leg needs the user's Windows deployment (wonder-app-vd03 / 10.100.0.48 redeploy is a separate user-approved step).
- Surface to the user: deployed services (`MxAccessGw` on 10.100.0.48, wonder-app-vd03) do NOT pick this up until redeployed; OtOpcUa's end-to-end verification against a live gateway needs that redeploy.
@@ -0,0 +1,17 @@
{
"planPath": "docs/plans/2026-08-09-write-completion-correlation.md",
"tasks": [
{"id": 0, "subject": "Task 0: Create feature branch", "status": "pending"},
{"id": 1, "subject": "Task 1: Proto contract comments + regen", "status": "pending", "blockedBy": [0]},
{"id": 2, "subject": "Task 2: MxAccessWriteCompletionCache + tests", "status": "pending", "blockedBy": [1]},
{"id": 3, "subject": "Task 3: Sink records completions (+ provider seam)", "status": "pending", "blockedBy": [2]},
{"id": 4, "subject": "Task 4: MxAccessSession plumbing", "status": "pending", "blockedBy": [3]},
{"id": 5, "subject": "Task 5: Executor bounded wait + StaSession env plumbing", "status": "pending", "blockedBy": [4]},
{"id": 6, "subject": "Task 6: Executor tests", "status": "pending", "blockedBy": [5]},
{"id": 7, "subject": "Task 7: Gateway config option + launcher env var", "status": "pending", "blockedBy": [0]},
{"id": 8, "subject": "Task 8: Docs", "status": "pending", "blockedBy": [6, 7]},
{"id": 9, "subject": "Task 9: Local + windev verification", "status": "pending", "blockedBy": [8]},
{"id": 10, "subject": "Task 10: Review, merge, notify", "status": "pending", "blockedBy": [9]}
],
"lastUpdated": "2026-08-09"
}
@@ -0,0 +1,171 @@
# SEC-36 — LDAP Service-Account Credential Rotation (Operator Runbook)
> **Executed 2026-08-07 — the rotation is done; this runbook is now history plus the four
> corrections below.** A new service-account password was generated, `scadaproj/infra/glauth/config.toml`'s
> `serviceaccount` `passsha256` was replaced and the shared GLAuth recreated, and the old value
> (the literal this repo committed, live in the directory since 2026-06-04) no longer binds. The new
> value now exists only in the
> three channels this runbook names: the GLAuth `passsha256` (committed in `scadaproj`), the NSSM
> service environment on `10.100.0.48`, and this dev Mac's user-secrets. The retired plaintext was
> also scrubbed from `scadaproj/infra/glauth/`'s `config.toml`/`docker-compose.yml`/`README.md`
> comments and from the host's live `docker-compose.yml` (the `*.bak-sec36` backups on the host still
> carry it, deliberately — they are the rollback artifacts).
>
> **Correction 1 — host paths in step 3 were stale.** The runbook says
> `cd ~/Desktop/scadaproj/infra/glauth` on `10.100.0.35`. That directory does not exist there:
> `scadaproj` is a dev-workstation checkout, and the docker host runs the stack from
> **`/home/dohertj2/zb-glauth`** (container **`zb-shared-glauth`**, project name `zb-shared-glauth`).
> The repo remains the source of truth; deployment is the `scp` of `config.toml`/`docker-compose.yml`
> into `~/zb-glauth` documented in `scadaproj/infra/glauth/README.md`, followed by
> `docker compose up -d --force-recreate` there.
>
> **Correction 2 — `wonder-app-vd03` is out of scope, on documentary evidence.** The precondition
> above says to check `MxGateway:Ldap:Enabled` on that host. It could not be checked directly (the
> host is unreachable from the dev network), but it is out of scope regardless: its gateway binds a
> **different directory** — the ScadaBridge/ScadaLink local GLAuth under `dc=scadalink`/`dc=scadabridge`,
> not `dc=zb,dc=local` — so this credential is not one it can hold. No env var was staged there and
> none is needed.
>
> **Correction 3 — step 4's dashboard verification is deferred on `10.100.0.48`; a direct bind was
> used instead.** The NEW value **is** staged on windev (added as the 10th `AppEnvironmentExtra`
> entry on the `MxAccessGw` NSSM service), but dashboard `/login` could not exercise it at rotation
> time: windev's gateway was **crash-looping on a pre-existing, unrelated fault** — the deployed Server binary
> (2026-06-25) predates the auth-DB migration of 2026-07-15, so it opens a schema-version-3 database
> it only supports at version 2 and aborts at startup (~10k Hosting-failed events/day since at least
> 08-06). That was a stale-deployment problem, not a rotation problem; it was filed as next-cycle
> finding NEXT-07 and resolved by redeploy on 2026-08-07. **Verification used instead:** a direct `ldapsearch` bind as
> `cn=serviceaccount,dc=zb,dc=local` with the new value against `10.100.0.35:3893` succeeded and
> returned the `multi-role` entry — which is precisely the search bind the dashboard performs, minus
> the HTTP shell. **The deferred check was completed 2026-08-07**, once windev was repaired by the
> redeploy filed under NEXT-07. With `Dashboard:DisableLogin=false` supplied as a process-env-only
> override on a foreground run of the new build, `GET /login` returned 200 with an antiforgery token,
> `POST /auth/login` as `multi-role`/`password` returned 302 to `/` with a `MxGatewayDashboard`
> cookie, the authenticated `GET /` rendered the admin nav, and an anonymous control redirected to
> `/login?ReturnUrl=%2F` — so the rotated credential is proven through the real
> `DashboardAuthenticator` search-bind path on the deployed host. As deployed, windev keeps
> `DisableLogin=true`, so routine operation there does not exercise LDAP; the standing regression
> proof is `DashboardLdapLiveTests` (5/5 green against `10.100.0.35` since commit `de67b45`).
>
> **Correction 4 — the lockout caution under "Verifying the rotation" is inert for this instance.**
> It warns that GLAuth's 3-fail / 10-minute per-IP lockout can lock the whole office when testing
> that the old value is dead. This GLAuth runs `LimitFailedBinds = false` (`config.toml:14`), so no
> failed-bind limiter is active and the caution does not apply here. Keep the caution for any
> instance that enables the limiter.
Operator steps to rotate the shared GLAuth service-account password after the repo-side
removal landed (SEC-36). The repo change (removal of the committed value, the two supported
secret channels, and this runbook) is already merged; the live rotation below is the
load-bearing half and is yours to execute.
> **Never put the old or new password in this repo, in a commit, in a chat, or in this file.**
> The value lives only in the GLAuth source of truth and in each host's out-of-band channel.
## Why
The dev GLAuth service-account password (`cn=serviceaccount,dc=zb,dc=local`) was historically
committed to this repo. Removal alone is insufficient — the old value is permanently recoverable
from git history — so **rotation is required**. Until the shared GLAuth on `10.100.0.35:3893`
stops honoring the old value, the repo history discloses a live directory account with LDAP
search capability over `dc=zb,dc=local`.
## Where the credential lives now (three channels, all bind `MxGateway:Ldap:ServiceAccountPassword`)
- **Source of truth:** `scadaproj/infra/glauth/config.toml` on host `10.100.0.35` (the `serviceaccount`
user's `passsha256`). `scadaproj` is a shared monorepo — stage only the explicit glauth paths.
- **Encrypted secrets store (gateway default):** `appsettings.json` ships `${secret:ldap/mxgateway/bind}`,
resolved from the local encrypted store (seed with `secret set ldap/mxgateway/bind <value>`).
- **Deployed hosts:** env var `MxGateway__Ldap__ServiceAccountPassword` in the NSSM service environment.
- **Dev boxes:** `dotnet user-secrets set "MxGateway:Ldap:ServiceAccountPassword" <value>`
(the server carries `<UserSecretsId>mxaccessgw-server</UserSecretsId>`).
See `docs/GatewayConfiguration.md` (the `ServiceAccountPassword` row) and `glauth.md`.
## Preconditions
- SSH access to the GLAuth docker host `10.100.0.35` and to the deployed gateway host(s).
- Write access to `scadaproj/infra/glauth/`.
- Know which deployed hosts run LDAP-backed dashboard login:
- **`10.100.0.48`** (`windev`) — primary; verify here.
- **`wonder-app-vd03`** — its dashboard is disabled. **Check `MxGateway:Ldap:Enabled` there first.**
If LDAP is disabled (`Enabled=false`), it has nothing to bind and needs no env var — skip it.
- A generated replacement secret (see step 1). Generate the `passsha256` per `glauth.md`
("Generate `passsha256` from a plaintext password").
## Cutover order
Follow this order so no window opens where the deployed dashboard cannot bind. **Do not rotate
GLAuth before the deployed hosts already carry the new value.**
1. **Generate the new secret in `scadaproj/infra/glauth/`.** Pick a new password, compute its
`passsha256`, and stage the change to the `serviceaccount` user in `config.toml` (do not
`docker compose up` yet — the directory must keep honoring the OLD value until the deployed
hosts carry the NEW one).
2. **Pre-stage the NEW value on every LDAP-enabled deployed host** via the env-var channel, so the
host is ready the instant GLAuth flips:
```powershell
nssm get MxAccessGw AppEnvironmentExtra
nssm set MxAccessGw AppEnvironmentExtra MxGateway__Ldap__ServiceAccountPassword=<new-value>
# restart the service so the new environment is picked up
nssm restart MxAccessGw
```
Do this on `10.100.0.48`, and on `wonder-app-vd03` **only if** `MxGateway:Ldap:Enabled=true` there.
(Alternatively seed the encrypted store with `secret set ldap/mxgateway/bind <new-value>`; the
env var overrides the store and is the simplest per-host mechanism.)
At this moment the deployed host holds the NEW value but GLAuth still honors the OLD one — binds
still fail closed against the old directory, which is expected and brief; proceed immediately.
3. **Rotate GLAuth on `10.100.0.35`** to honor the new value:
```bash
ssh 10.100.0.35
cd ~/Desktop/scadaproj/infra/glauth
docker compose up -d --force-recreate
docker compose logs -f # confirm clean startup, no TOML parse error
```
4. **Verify dashboard login on the deployed host(s).** Browse to the gateway dashboard on
`10.100.0.48` and log in as `multi-role` / `password` (Administrator) — a successful login proves
the search bind used the new service-account credential end-to-end. If `wonder-app-vd03` runs
LDAP, verify it too; if its dashboard/LDAP is disabled, no check is needed.
5. **The repo change is already landed** (removal of the committed value, `<UserSecretsId>`, the
validator message naming the two channels, and doc/scrub updates). Nothing more to commit for
the cutover.
6. **Developers set user-secrets on next pull.** After pulling, a dev box with no secret configured
will fail startup with a validation message naming the exact command. One-time per machine:
```bash
dotnet user-secrets set "MxGateway:Ldap:ServiceAccountPassword" <new-value>
```
(value from `scadaproj/infra/glauth/`, never from a repo file).
## Verifying the rotation
- **Primary:** dashboard `/login` as `multi-role` on `10.100.0.48` succeeds (step 4).
- **`wonder-app-vd03`:** only if `MxGateway:Ldap:Enabled=true`; otherwise no action.
- **Live-LDAP integration tests** (opt-in, only where the GLAuth instance is reachable):
```bash
$env:MXGATEWAY_RUN_LIVE_LDAP_TESTS = "1"
$env:MxGateway__Ldap__ServiceAccountPassword = "<new-value>" # shell env only, never committed
dotnet test src/ZB.MOM.WW.MxGateway.IntegrationTests/ZB.MOM.WW.MxGateway.IntegrationTests.csproj `
--filter FullyQualifiedName~DashboardLdapLiveTests
```
A green `DashboardLdapLiveTests` run confirms the new credential binds and searches. Where GLAuth
is unreachable, document the suite as skipped per the `docs/GatewayTesting.md` opt-in matrix.
- **Old value is dead:** after step 3, a bind with the old password must fail. Do not test this from
a shared-NAT box — GLAuth's 3-fail / 10-minute per-IP lockout can lock the whole office.
## Rollback
If dashboard login breaks after step 3, restore the previous `passsha256` in
`scadaproj/infra/glauth/config.toml`, `docker compose up -d --force-recreate`, and re-point the
deployed hosts' env var / store back to the previous value. Because the deployed hosts were
pre-staged in step 2, the exposure window is only steps 2→4.
## Done criteria
- GLAuth on `10.100.0.35` honors only the new value.
- Every LDAP-enabled deployed host binds with the new value (dashboard login verified).
- The source of truth `scadaproj/infra/glauth/config.toml` carries the new `passsha256`.
- No repo file (this one included) contains the old or new value.
- The SEC-36 tracker rows are `Done` with this runbook cited for the operator action.
+148
View File
@@ -0,0 +1,148 @@
# TST-30 — Register A Second CI Runner (Operator Runbook)
> **Executed 2026-08-07 — option (a) shipped; this runbook is now history plus the one
> correction below.** `gitea-runner-2` (runner id 5, capacity 2, labels `ubuntu-latest`/
> `ubuntu-22.04`) runs on `10.100.0.35` from the `/opt/gitea` compose stack with the same
> `container.network: traefik` setting as the original; its registration token is mounted from a
> `0600` file rather than inlined in compose. The existing `gitea-runner` (id 1, capacity 4) was
> left untouched, so capacity went 4 → 6 by addition and the change reverts by removing one
> container. Concurrency was verified by pushing HEAD to two scratch branches while an unrelated
> run was in flight: jobs from three runs ran simultaneously across both runners, and a
> `gitea-runner-2` job cloned successfully from `http://gitea:3000` (the property option (c) was
> rejected for losing).
>
> **Correction to the Verification and Done-criteria sections below:** they expect
> `GET /repos/dohertj2/mxaccessgw/actions/runners` to show ≥2 runners. It does not — it still
> returns `total_count: 0`, correctly, because both runners are registered at the **instance**
> level, which is the very condition the "Why" section describes. Use
> `GET /api/v1/admin/actions/runners` instead (it now lists only id 1 and id 5; id 4 was removed
> 2026-08-07 — a local macOS `act_runner` mislabelled `ubuntu-latest`/`ubuntu-22.04`/`ubuntu-20.04`,
> so it captured Linux-labelled jobs it had no Docker daemon to run and failed them. Its config and
> registration are kept disabled at `~/gitea-act-runner.disabled-2026-08-07` (launchd plist at
> `~/Library/LaunchAgents/com.dohertj2.gitea-act-runner.plist.disabled-2026-08-07`) so it can be
> re-registered with mac-specific labels if a mac-only job ever needs one). For per-job runner
> attribution, `GET /repos/{owner}/{repo}/actions/runs/{id}/jobs`
> exposes `runner_id`/`runner_name` on each job; the `actions/tasks` listing does not.
>
> **Follow-up 2026-08-07 — token hygiene on the host.** `gitea-runner` (id 1) now takes its
> registration token from the same `0600` file mount runner-2 uses instead of an inline plaintext
> value in compose, and `/opt/gitea/docker-compose.yml` plus both `.bak` copies are `0600 root:root`;
> runner-1 was recreated alone and kept its identity (`.runner` byte-identical). Both runners share
> **one instance-scope registration token**, which was world-readable for roughly five months and is
> still live — a probe registered runner id 6 with it, then deleted it. Gitea 1.26.4 cannot rotate
> that token from the CLI or the API (both endpoints are get-or-create and return the same value),
> so **the reset is a pending operator action in the admin web UI** ("Reset registration token").
> After the reset, refresh `/opt/gitea/runner_token` with the new value and shred the two
> token-bearing compose backups, which are the last copies of the old one.
Operator steps to relieve the single shared Gitea Actions runner that CI depends on. The
repo-side half of TST-30 (documenting the shared-runner/no-cancel reality and the
`run-windev-ci.sh` bypass) is already landed in `docs/GatewayTesting.md`; registering the
second runner below is infrastructure work outside this repo's tree and is yours to execute.
## Why
All CI for this repo runs on one co-located `gitea-runner` container on docker host
`10.100.0.35` with `maxParallel=1`. That runner is registered at the **instance** level, not
scoped to this repo (`GET /repos/dohertj2/mxaccessgw/actions/runners` returns
`total_count: 0`), so it is shared with `dohertj2/lmxopcua` and every job in every run across
both repos executes serially on the single slot. A `mxaccessgw` push fans out to `portable`,
`java`, `windows-x86`, and an active `lmxopcua` run blocks all of them — queue depth of
~2030 minutes was observed during TST-25 acceptance under cross-repo contention. Gitea 1.26
also exposes **no run cancel or delete via the API**
(`POST .../actions/runs/{id}/cancel` → 404, `DELETE .../actions/runs/{id}` → 400), so a
superseded or hung run cannot be cleared and holds the slot until it finishes or times out.
This is not a correctness problem — every job still reports accurately — but it undercuts the
fast-feedback purpose of the TST-25 Windows tier and makes CI fragile to a single host: if
`10.100.0.35` wedges or goes down, CI for both repos stops with no failover.
## Options (cheapest first)
- **(a) Register a second `act_runner` instance on `10.100.0.35` — recommended.** The host
already runs `gitea-runner`; add a second `act_runner` container (or raise the existing
runner's `maxParallel` where the docker-in-docker/resource budget allows) so at least two
jobs run concurrently. Cheapest change, and it keeps the runner co-located on the
`container.network: traefik` network that resolves `gitea:3000` — the property TST-03
depended on. **Use the same `container.network: traefik` config as the existing runner.**
- **(b) Dedicate a labelled runner to `mxaccessgw`.** Cleaner isolation — `lmxopcua` load
never blocks this repo — but needs label wiring: register the new runner with a distinct
label (e.g. `mxgw`) and change `.gitea/workflows/ci.yml`'s `runs-on:` for this repo's jobs
to gate on that label (e.g. `runs-on: [ubuntu-latest, mxgw]`). Only do this if (a) proves
insufficient — it is more moving parts for the same throughput gain, and it means `ci.yml`
changes, which is out of scope for the doc-only half of TST-30.
- **(c) Put the runner on windev / a second host — rejected as the primary fix.** windev is
the Windows build target (`10.100.0.48`), not a CI host, and co-locating a Linux runner
there loses the `gitea:3000` name resolution TST-03 relies on. Only consider if
`10.100.0.35` genuinely runs out of capacity for a second instance.
Default to **(a)**. Escalate to (b) only if `lmxopcua` contention persists after a second
instance is online (i.e., (a) is not sufficient because the two repos' combined load exceeds
two slots).
## Preconditions
- SSH/docker access to `10.100.0.35`.
- The existing `gitea-runner` container's compose/run config, to copy its
`container.network: traefik` setting and registration token flow (repo memory
`project_gitea_ci` records this configuration).
- Admin access to Gitea (`gitea.dohertylan.com`) to mint a new runner registration token.
## Steps — option (a): second runner instance
1. On `10.100.0.35`, locate the existing `gitea-runner` container/compose definition and copy
its configuration for a new instance (same `container.network: traefik`, same Docker
socket mount if it uses docker-in-docker, a distinct container name/data volume).
2. In Gitea, generate a new runner registration token (instance-level, since the existing
runner is registered at the instance level too — Admin → Actions → Runners, or
`POST /admin/actions/runners/registration-token`).
3. Register and start the second `act_runner` instance with that token, pointed at the same
Gitea origin.
4. Confirm both runners show online: instance runner list in the Gitea admin UI, or the
equivalent API listing.
## Verification
- Push two branches to `mxaccessgw` back-to-back (or trigger one `mxaccessgw` push while an
`lmxopcua` run is in flight) and confirm both runs execute **concurrently**, not serially —
the second run's jobs should start before the first finishes, not queue behind it.
- `GET /repos/dohertj2/mxaccessgw/actions/runners` (or the instance runner listing) shows
**≥2** runners online.
- Re-run the TST-25 acceptance push (a plain push to a scratch branch) and confirm queue depth
is materially lower than the ~2030 minute baseline observed under a concurrent `lmxopcua`
run.
- Confirm `windows-x86` still resolves `gitea:3000` correctly from a job scheduled on the new
runner instance (the `traefik` network property must hold for both instances).
## The no-cancel reality does not go away
A second runner relieves contention; it does not add a cancel/delete API — Gitea 1.26 still
returns 404/400 for both. A stale or hung run on either runner still holds its slot until it
finishes or times out. Two runners just means one stale run blocks at most half the capacity
instead of all of it. Do not treat the second runner as a substitute for the escape hatch: a
specific commit can still be verified out of band without waiting on either runner via
`CI_SHA=<sha> scripts/ci/run-windev-ci.sh <build|test|live>` (Linux, needs SSH access to
windev) or the manual windev worktree flow — see the "Runner capacity is shared and finite"
section in `docs/GatewayTesting.md`.
## Optional: workflow-level `concurrency` group
As belt-and-suspenders against the missing cancel API, `.gitea/workflows/ci.yml` could add a
top-level `concurrency` group (e.g. keyed on `${{ github.ref }}`) so a newer push to the same
branch automatically supersedes an in-flight run instead of both running to completion.
**Verify this Gitea deployment actually honors `concurrency` and cancels the superseded run
before relying on it** — Gitea Actions' YAML surface does not track GitHub Actions feature
parity release-for-release, and a `concurrency` block that is silently ignored would look like
a working safeguard while doing nothing. If verified working, this is a `ci.yml` change (not
covered by this runbook) and should land as its own small change with its own verification
(push twice to the same branch quickly, confirm the first run's jobs cancel).
## Done criteria
- A second `act_runner` instance (or raised `maxParallel`) is online on `10.100.0.35` with the
same `container.network: traefik` configuration as the existing runner.
- `GET /repos/dohertj2/mxaccessgw/actions/runners` (or the instance listing) shows ≥2 runners.
- Two concurrent runs (one `mxaccessgw`, one `lmxopcua`, or two `mxaccessgw` pushes) execute
in parallel rather than serially.
- `docs/GatewayTesting.md`'s shared-runner/no-cancel prose and the `run-windev-ci.sh` bypass
remain accurate (they describe the bypass as still valid, which it is regardless of runner
count).
+17
View File
@@ -431,6 +431,23 @@ Core commands:
- `AuthenticateUser`
- `ArchestrAUserToId`
**Secured-write completion correlation.** MXAccess writes are fire-and-forget
at the toolkit level — the per-item outcome only exists in the later
`OnWriteComplete` callback. For `WriteSecured` and `WriteSecured2` the worker
therefore holds the unary reply for a bounded window
(`MxGateway:Worker:WriteCompletionWaitMilliseconds`, default 1.5 s, `0`
disables) and, when the matching callback arrives, copies its status rows onto
`MxCommandReply.statuses` — the reply then proves the MXAccess commit, not just
command acceptance. `protocol_status`/`hresult` keep describing acceptance
only; a real MXAccess write failure surfaces in `statuses[0]`, and a reply with
empty `statuses` means unconfirmed (the callback missed the window), never
failed. The `OnWriteComplete` event still flows on the event stream unchanged.
Correlation is best-effort per `(server_handle, item_handle)` — the callback
carries no transaction id, so concurrent writes to the same item within the
window can swap rows. Plain `Write`/`Write2` and the bulk write commands stay
fire-and-forget: waiting there would add a device round-trip of latency to
high-rate supervisory write loops.
Bulk variants (single gRPC round-trip carries the full list, the worker
runs the per-item MXAccess calls sequentially on its STA, and the reply
returns one result per requested entry — per-entry failures populate
+15 -5
View File
@@ -30,10 +30,20 @@ gw-specific role.
| LDAPS | disabled in dev (`Transport=None`, `AllowInsecure=true`) |
| Base DN | `dc=zb,dc=local` |
| Bind DN format | `cn={username},dc=zb,dc=local` |
| Service account DN | `cn=serviceaccount,dc=zb,dc=local` / `serviceaccount123` |
| Service account DN | `cn=serviceaccount,dc=zb,dc=local` (password: `<service-account-password>`) |
| Group OU | `ou=<groupname>,ou=groups,dc=zb,dc=local` |
| Failed-bind throttle | 3 fails → 10-minute IP lockout (per `[behaviors]`) |
> **Service-account password is not committed (SEC-36).** The samples below show
> `<service-account-password>` as a placeholder, not the real value. The single source of
> truth is **`scadaproj/infra/glauth/config.toml`** on host `10.100.0.35`; the gateway consumes
> it out-of-band (encrypted secrets store reference `${secret:ldap/mxgateway/bind}`, the
> `MxGateway__Ldap__ServiceAccountPassword` env var on deployed hosts, or `dotnet user-secrets`
> on dev boxes — see `docs/GatewayConfiguration.md`). The credential was historically committed
> to this repo (and remains recoverable from git history); it **was rotated on 2026-08-07**
> (SEC-36, executed per `docs/runbooks/SEC-36-ldap-credential-rotation.md`) and the old
> committed value no longer binds.
## Pre-existing groups (LmxOpcUa role taxonomy)
These map cleanly onto MxAccess capability boundaries — mxaccessgw
@@ -62,7 +72,7 @@ group below).
| `writeconfig` | `writeconfig123` | 5006 | WriteConfigure | — | + WriteSecured (Configure) |
| `alarmack` | `alarmack123` | 5003 | AlarmAck | — | Alarm acknowledgment |
| `admin` | `admin123` | 5004 | ReadOnly | WriteOperate, AlarmAck, WriteTune, WriteConfigure | All roles |
| `serviceaccount` | `serviceaccount123` | 5999 | ReadOnly | — | LDAP search capability (for bind-then-search) |
| `serviceaccount` | `<service-account-password>` | 5999 | ReadOnly | — | LDAP search capability (for bind-then-search) |
For mxaccessgw dev, `admin` covers every gw-side capability test;
`readonly` is the right "negative" case for proving Browse-OK /
@@ -100,7 +110,7 @@ by `sAMAccountName`, not `cn`. Use this only for dev convenience.
```
1. Bind as the service account (cn=serviceaccount,dc=zb,dc=local
/ serviceaccount123).
/ <service-account-password>).
2. Search under dc=zb,dc=local with filter
(uid=<entered-username>) — or any attribute the deployment
identifies users by. GLAuth populates uid + cn.
@@ -133,7 +143,7 @@ ldap:
allowInsecureLdap: true # dev only
searchBase: "dc=zb,dc=local"
serviceAccountDn: "cn=serviceaccount,dc=zb,dc=local"
serviceAccountPassword: "serviceaccount123"
serviceAccountPassword: "<service-account-password>" # not committed; see source-of-truth note
userNameAttribute: "uid" # GLAuth populates this; AD uses sAMAccountName
displayNameAttribute: "cn"
groupAttribute: "memberOf"
@@ -242,7 +252,7 @@ Or via `ldapsearch` if you have OpenLDAP CLI tools:
```bash
ldapsearch -x -H ldap://10.100.0.35:3893 \
-D "cn=serviceaccount,dc=zb,dc=local" -w serviceaccount123 \
-D "cn=serviceaccount,dc=zb,dc=local" -w '<service-account-password>' \
-b "dc=zb,dc=local" "(uid=multi-role)"
```
+33 -5
View File
@@ -1,7 +1,7 @@
#!/usr/bin/env pwsh
# Codegen freshness guard for CI (IPC-01, IPC-19, IPC-20, CLI-02).
# Codegen freshness guard for CI (IPC-01, IPC-19, IPC-20, IPC-25, CLI-02).
#
# Three checks, all Linux/macOS-runnable (no Server build, no x86 worker):
# Four checks, all Linux/macOS-runnable (no Server build, no x86 worker):
# 1. Published client descriptor set matches the current .proto sources (delegates to
# publish-client-proto-inputs.ps1 -Check, which normalizes source_code_info so it is
# protoc-version tolerant).
@@ -13,6 +13,12 @@
# crate buildable outside the repo, CLI-02) are byte-identical to the canonical Contracts
# protos. A drift means a .proto was edited without refreshing the vendored copies, which would
# publish a stale wire contract to crate consumers while the in-repo build stays correct.
# 4. The committed Go and Python client bindings match a fresh regeneration (IPC-25). The two
# per-client generate-proto.ps1 scripts pin their generators (protoc-gen-go v1.36.11 /
# protoc-gen-go-grpc v1.6.2 for Go; grpcio-tools 1.80.0 for Python), so a clean checkout
# regenerates deterministic output; a non-empty git diff means a .proto was edited without
# regenerating and committing those bindings. A missing generator FAILS the check (a skipped
# guard is the exact silent-drift hole IPC-25 closes), never skips it.
#
# The x86 Worker + Worker.Tests are Windows-only and are guarded by the SSH-driven `windows-x86`
# CI job (see docs/GatewayTesting.md, Continuous Integration), not here.
@@ -28,7 +34,7 @@ $generatedDir = Join-Path $repoRoot 'src/ZB.MOM.WW.MxGateway.Contracts/Generated
$contractsProject = Join-Path $repoRoot 'src/ZB.MOM.WW.MxGateway.Contracts/ZB.MOM.WW.MxGateway.Contracts.csproj'
$failures = New-Object System.Collections.Generic.List[string]
Write-Host '== Check 1/2: client descriptor set freshness =='
Write-Host '== Check 1/4: client descriptor set freshness =='
try {
& (Join-Path $PSScriptRoot 'publish-client-proto-inputs.ps1') -Check
if ($LASTEXITCODE -ne 0) {
@@ -40,7 +46,7 @@ catch {
}
Write-Host ''
Write-Host '== Check 2/2: Contracts/Generated matches a fresh regeneration =='
Write-Host '== Check 2/4: Contracts/Generated matches a fresh regeneration =='
try {
# Force a full regeneration: Grpc.Tools skips regen when the committed .cs look up to date, so
# remove them first (the documented "del Generated/*.cs to force regen" trick).
@@ -66,7 +72,7 @@ catch {
}
Write-Host ''
Write-Host '== Check 3/3: Rust vendored protos match canonical Contracts protos =='
Write-Host '== Check 3/4: Rust vendored protos match canonical Contracts protos =='
try {
$canonicalProtoDir = Join-Path $repoRoot 'src/ZB.MOM.WW.MxGateway.Contracts/Protos'
$vendoredProtoDir = Join-Path $repoRoot 'clients/rust/protos'
@@ -87,6 +93,28 @@ catch {
$failures.Add("Rust vendored proto check failed: $($_.Exception.Message)")
}
Write-Host ''
Write-Host '== Check 4/4: Go and Python client bindings match a fresh regeneration =='
try {
# Regenerate both binding sets with their pinned generators, then diff. The per-client scripts
# throw on a missing or off-pin generator, so any failure here FAILS the check rather than
# skipping it (a skipped guard is exactly the silent-drift hole IPC-25 closes).
$goBindingDir = 'clients/go/internal/generated'
$pyBindingDir = 'clients/python/src/zb_mom_ww_mxgateway/generated'
& (Join-Path $repoRoot 'clients/go/generate-proto.ps1') | Out-Host
& (Join-Path $repoRoot 'clients/python/generate-proto.ps1') | Out-Host
$bindingDiff = (& git -C $repoRoot status --porcelain -- $goBindingDir $pyBindingDir | Out-String).Trim()
if (-not [string]::IsNullOrEmpty($bindingDiff)) {
Write-Host $bindingDiff
$failures.Add("Go/Python client bindings differ from a fresh regeneration. Run clients/go/generate-proto.ps1 and clients/python/generate-proto.ps1 with the pinned generators and commit $goBindingDir and $pyBindingDir.")
}
}
catch {
$failures.Add("Go/Python codegen check failed (tool missing or regeneration error): $($_.Exception.Message)")
}
Write-Host ''
if ($failures.Count -gt 0) {
Write-Host 'Codegen freshness check FAILED:' -ForegroundColor Red
+116
View File
@@ -87,6 +87,12 @@ $GiteaNugetFeed = 'https://gitea.dohertylan.com/api/packages/dohertj2/nuget/inde
$GiteaPypiFeed = 'https://gitea.dohertylan.com/api/packages/dohertj2/pypi'
$JavaHome = '/Users/dohertj2/.local/jdks/jdk-21.0.11+10/Contents/Home'
# Generic Gitea package registry API (https://gitea.dohertylan.com/api/v1/packages/{owner}/{type}/{name}/{version}):
# returns 200 when that exact name+version already exists in the given feed
# type, 404 when it does not. Used as a pre-publish collision guard (CLI-39)
# so a re-run of this script can never silently overwrite a published artifact.
$GiteaPackageApiBase = 'https://gitea.dohertylan.com/api/v1/packages/dohertj2'
function Write-Header {
param([string]$Text)
Write-Host ''
@@ -94,6 +100,64 @@ function Write-Header {
Write-Host $Text -ForegroundColor Cyan
}
function Test-GiteaPackageExists {
<#
.SYNOPSIS
Queries the Gitea package API for an existing name+version in a feed.
.OUTPUTS
$true if the package/version already exists, $false if it does not.
Throws if the registry cannot be reached or returns anything other
than 200/404 callers must treat "cannot verify" as "do not publish".
#>
param(
[Parameter(Mandatory)][string]$Type,
[Parameter(Mandatory)][string]$Name,
[Parameter(Mandatory)][string]$Version
)
$uri = "$GiteaPackageApiBase/$Type/$Name/$Version"
$headers = @{}
if (-not [string]::IsNullOrEmpty($env:GITEA_TOKEN)) {
$user = if ([string]::IsNullOrEmpty($env:GITEA_USERNAME)) { 'dohertj2' } else { $env:GITEA_USERNAME }
$pair = "$($user):$($env:GITEA_TOKEN)"
$basic = [Convert]::ToBase64String([System.Text.Encoding]::ASCII.GetBytes($pair))
$headers['Authorization'] = "Basic $basic"
}
try {
$response = Invoke-WebRequest -Uri $uri -Headers $headers -Method Get -UseBasicParsing -ErrorAction Stop
return ($response.StatusCode -eq 200)
} catch {
$statusCode = $null
if ($_.Exception.PSObject.Properties['Response'] -and $_.Exception.Response) {
$statusCode = [int]$_.Exception.Response.StatusCode
}
if ($statusCode -eq 404) {
return $false
}
throw "Unable to query the Gitea package API for '$Type/$Name/$Version' ($uri): $($_.Exception.Message). Refusing to publish without a collision check — verify manually (or check GITEA_USERNAME/GITEA_TOKEN/network) and retry."
}
}
function Assert-GiteaPackageNotPublished {
<#
.SYNOPSIS
Aborts the script if $Name/$Version already exists in the $Type feed.
Never force-overwrites a published artifact (CLI-39).
#>
param(
[Parameter(Mandatory)][string]$Type,
[Parameter(Mandatory)][string]$Name,
[Parameter(Mandatory)][string]$Version
)
Write-Host "Checking Gitea '$Type' feed for existing '$Name' $Version..."
if (Test-GiteaPackageExists -Type $Type -Name $Name -Version $Version) {
throw "Gitea package '$Name' version '$Version' already exists in the '$Type' feed. Bump the client version before publishing — this script never force-overwrites a published artifact."
}
Write-Host " Not found in the '$Type' feed — safe to publish '$Name' $Version." -ForegroundColor Green
}
# -------- .NET --------
function Invoke-PackDotnet {
@@ -121,6 +185,15 @@ function Invoke-PackDotnet {
if ($Publish) {
Write-Host 'Publishing .NET packages to Gitea...' -ForegroundColor Yellow
Get-ChildItem $OutputDir -Filter 'ZB.MOM.WW.MxGateway.*.nupkg' | ForEach-Object {
# nupkg filenames are '<PackageId>.<Version>.nupkg'; the id itself contains
# dots (e.g. 'ZB.MOM.WW.MxGateway.Client'), so the id capture is lazy and the
# version capture anchors on the leading digit to split at the right dot.
$fileBaseName = [System.IO.Path]::GetFileNameWithoutExtension($_.Name)
if ($fileBaseName -notmatch '^(?<id>.+?)\.(?<version>\d+\.\d+\.\d+(?:[-+][0-9A-Za-z.-]+)?)$') {
throw "Could not parse a NuGet package id/version out of '$($_.Name)'."
}
Assert-GiteaPackageNotPublished -Type 'nuget' -Name $Matches.id -Version $Matches.version
& dotnet nuget push $_.FullName --source $GiteaNugetFeed --api-key $env:GITEA_TOKEN
if ($LASTEXITCODE -ne 0) { throw "dotnet nuget push failed for '$($_.Name)'." }
}
@@ -159,6 +232,20 @@ function Invoke-PackPython {
Write-Host "Packed Python artifacts -> $OutputDir" -ForegroundColor Green
if ($Publish) {
$pyprojectPath = Join-Path $RepoRoot 'clients/python/pyproject.toml'
$pyprojectContent = Get-Content $pyprojectPath -Raw
# Scope to the [project] section (not just the first "version = ..." line
# in the file) — [build-system]/[tool.*] sections can carry their own
# version-shaped keys, and matching the file's first hit would be luck
# of ordering, not correctness.
if ($pyprojectContent -notmatch '(?ms)^\[project\](?<section>.*?)(?=^\[|\z)') {
throw "Could not find a [project] section in '$pyprojectPath'."
}
if ($Matches.section -notmatch '(?m)^\s*version\s*=\s*"([^"]+)"') {
throw "Could not find [project].version in '$pyprojectPath'."
}
Assert-GiteaPackageNotPublished -Type 'pypi' -Name 'zb-mom-ww-mxaccess-gateway-client' -Version $Matches[1]
Write-Host 'Publishing Python distribution to Gitea...' -ForegroundColor Yellow
$wheels = @(Get-ChildItem $OutputDir -Filter 'zb_mom_ww_mxaccess_gateway_client-*.whl')
$sdists = @(Get-ChildItem $OutputDir -Filter 'zb_mom_ww_mxaccess_gateway_client-*.tar.gz')
@@ -206,6 +293,20 @@ function Invoke-PackRust {
Write-Host "Packed Rust artifacts -> $OutputDir" -ForegroundColor Green
if ($Publish) {
$cargoTomlPath = Join-Path $rustDir 'Cargo.toml'
$cargoTomlContent = Get-Content $cargoTomlPath -Raw
# Scope to the [package] section specifically — Cargo.toml also carries a
# [workspace.package] section with its own "version = ..." line (today
# identical, by convention, not by anything this regex can rely on), and
# matching whichever comes first in the file is luck of ordering.
if ($cargoTomlContent -notmatch '(?ms)^\[package\](?<section>.*?)(?=^\[|\z)') {
throw "Could not find a [package] section in '$cargoTomlPath'."
}
if ($Matches.section -notmatch '(?m)^\s*version\s*=\s*"([^"]+)"') {
throw "Could not find [package] version in '$cargoTomlPath'."
}
Assert-GiteaPackageNotPublished -Type 'cargo' -Name 'zb-mom-ww-mxgateway-client' -Version $Matches[1]
Write-Host 'Publishing Rust crate to Gitea...' -ForegroundColor Yellow
Push-Location (Join-Path $RepoRoot 'clients/rust')
try {
@@ -269,6 +370,21 @@ function Invoke-PackJava {
Write-Host "Packed Java artifacts -> $OutputDir" -ForegroundColor Green
if ($Publish) {
$buildGradlePath = Join-Path $javaDir 'build.gradle'
$buildGradleContent = Get-Content $buildGradlePath -Raw
if ($buildGradleContent -notmatch "(?m)^\s*group\s*=\s*'([^']+)'") {
throw "Could not find subprojects { group = '...' } in '$buildGradlePath'."
}
$javaGroup = $Matches[1]
if ($buildGradleContent -notmatch "(?m)^\s*version\s*=\s*'([^']+)'") {
throw "Could not find subprojects { version = '...' } in '$buildGradlePath'."
}
$javaVersion = $Matches[1]
# Gitea's Maven package API identifies the package as "groupId:artifactId",
# not the bare artifact id — passing just the artifact id here would query
# a name that never exists and silently defeat the guard.
Assert-GiteaPackageNotPublished -Type 'maven' -Name "$javaGroup`:zb-mom-ww-mxgateway-client" -Version $javaVersion
Write-Host 'Publishing Java artifacts to Gitea Maven feed...' -ForegroundColor Yellow
Push-Location $javaDir
try {
+17
View File
@@ -36,6 +36,23 @@ if ($Version -notmatch '^v\d+\.\d+\.\d+(-[A-Za-z0-9.-]+)?$') {
throw "Version '$Version' must match semver vX.Y.Z (optionally with -prerelease suffix)."
}
# CLI-21 guard: the tag must match what the module itself reports via
# ClientVersion, or `go get <module>@vX.Y.Z` resolves a tag whose module
# code disagrees with its own version constant.
$versionGoPath = Join-Path $PSScriptRoot '..' 'clients/go/mxgateway/version.go'
if (-not (Test-Path $versionGoPath)) {
throw "Could not find '$versionGoPath' to verify ClientVersion before tagging."
}
$versionGoContent = Get-Content $versionGoPath -Raw
if ($versionGoContent -notmatch 'ClientVersion\s*=\s*"([^"]+)"') {
throw "Could not find a ClientVersion = `"...`" constant in '$versionGoPath'."
}
$clientVersion = $Matches[1]
$tagVersion = $Version.TrimStart('v')
if ($clientVersion -ne $tagVersion) {
throw "clients/go/mxgateway/version.go ClientVersion is '$clientVersion' but the requested tag is '$tagVersion'. Update ClientVersion to match before tagging."
}
$tag = "clients/go/$Version"
Write-Host "Creating Go-module tag: $tag" -ForegroundColor Cyan
+8 -4
View File
@@ -11,10 +11,14 @@
<!-- TST-11: single-source the .NET-side version for Server, Worker, Contracts, and tests
(they otherwise stamp the SDK default 1.0.0, so a deployed gateway cannot be correlated
to a release). Kept at 0.1.2 to match the Contracts package and the aligned Python/Rust/
Go clients; the Java client leads at 0.2.0 after its JDK-17 retarget. The git short SHA is
appended to InformationalVersion (0.1.2+<sha>) so support can map a running binary to a
commit; the query is guarded so a build outside a git checkout still succeeds. -->
to a release). Server/Worker/Tests stay at this default. CLI-39 (2026-08-07) moved the
published `ZB.MOM.WW.MxGateway.Contracts` and `.Client` nuget packages to 0.2.0 via an
explicit <Version> override in Contracts.csproj (MSBuild property-last-write-wins over
this Directory.Build.props default) — Server/Worker assembly stamping and the published
client packages are deliberately decoupled; a broader 0.2.0 alignment for Server/Worker
is a separate, not-yet-made decision. The git short SHA is appended to
InformationalVersion (0.1.2+<sha>) so support can map a running binary to a commit; the
query is guarded so a build outside a git checkout still succeeds. -->
<PropertyGroup>
<Version>0.1.2</Version>
</PropertyGroup>
@@ -8694,6 +8694,11 @@ namespace ZB.MOM.WW.MxGateway.Contracts.Proto {
}
/// <summary>
/// The unary reply's statuses field carries the correlated OnWriteComplete
/// outcome when it arrives within the worker's bounded wait — see
/// MxCommandReply.statuses.
/// </summary>
[global::System.Diagnostics.DebuggerDisplayAttribute("{ToString(),nq}")]
public sealed partial class WriteSecuredCommand : pb::IMessage<WriteSecuredCommand>
#if !GOOGLE_PROTOBUF_REFSTRUCT_COMPATIBILITY_MODE
@@ -9053,6 +9058,11 @@ namespace ZB.MOM.WW.MxGateway.Contracts.Proto {
}
/// <summary>
/// The unary reply's statuses field carries the correlated OnWriteComplete
/// outcome when it arrives within the worker's bounded wait — see
/// MxCommandReply.statuses.
/// </summary>
[global::System.Diagnostics.DebuggerDisplayAttribute("{ToString(),nq}")]
public sealed partial class WriteSecured2Command : pb::IMessage<WriteSecured2Command>
#if !GOOGLE_PROTOBUF_REFSTRUCT_COMPATIBILITY_MODE
@@ -17204,6 +17214,20 @@ namespace ZB.MOM.WW.MxGateway.Contracts.Proto {
private static readonly pb::FieldCodec<global::ZB.MOM.WW.MxGateway.Contracts.Proto.MxStatusProxy> _repeated_statuses_codec
= pb::FieldCodec.ForMessage(58, global::ZB.MOM.WW.MxGateway.Contracts.Proto.MxStatusProxy.Parser);
private readonly pbc::RepeatedField<global::ZB.MOM.WW.MxGateway.Contracts.Proto.MxStatusProxy> statuses_ = new pbc::RepeatedField<global::ZB.MOM.WW.MxGateway.Contracts.Proto.MxStatusProxy>();
/// <summary>
/// Correlated per-item outcome rows. For WRITE_SECURED / WRITE_SECURED2
/// replies the worker holds the reply for a bounded window (default 1.5 s,
/// MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS) waiting for the matching
/// MXAccess OnWriteComplete callback and copies its status rows here, so
/// statuses[0] carries the real MXAccess commit outcome (success OR failure)
/// while protocol_status/hresult still describe command acceptance only.
/// Empty statuses on a write reply means the completion did not arrive
/// within the window — the write is unconfirmed, not failed. Correlation is
/// best-effort per (server_handle, item_handle): MXAccess's callback carries
/// no transaction id, so concurrent writes to the same item within the
/// window can swap rows. The OnWriteComplete event still flows on the event
/// stream unchanged. Other command kinds leave this field as before.
/// </summary>
[global::System.Diagnostics.DebuggerNonUserCodeAttribute]
[global::System.CodeDom.Compiler.GeneratedCode("protoc", null)]
public pbc::RepeatedField<global::ZB.MOM.WW.MxGateway.Contracts.Proto.MxStatusProxy> Statuses {
@@ -22796,6 +22820,12 @@ namespace ZB.MOM.WW.MxGateway.Contracts.Proto {
private static readonly pb::FieldCodec<global::ZB.MOM.WW.MxGateway.Contracts.Proto.MxEvent> _repeated_events_codec
= pb::FieldCodec.ForMessage(10, global::ZB.MOM.WW.MxGateway.Contracts.Proto.MxEvent.Parser);
private readonly pbc::RepeatedField<global::ZB.MOM.WW.MxGateway.Contracts.Proto.MxEvent> events_ = new pbc::RepeatedField<global::ZB.MOM.WW.MxGateway.Contracts.Proto.MxEvent>();
/// <summary>
/// The reply is bounded by both a server-side count cap and the negotiated
/// worker-frame byte cap; a reply may therefore carry fewer events than
/// `max_events` and fewer than are queued. Callers drain iteratively until an
/// empty reply.
/// </summary>
[global::System.Diagnostics.DebuggerNonUserCodeAttribute]
[global::System.CodeDom.Compiler.GeneratedCode("protoc", null)]
public pbc::RepeatedField<global::ZB.MOM.WW.MxGateway.Contracts.Proto.MxEvent> Events {
@@ -24510,6 +24540,11 @@ namespace ZB.MOM.WW.MxGateway.Contracts.Proto {
/// after_worker_sequence = oldest_available_sequence - 1 in the next
/// StreamEventsRequest, which will cause the server to replay starting at
/// oldest_available_sequence (the first retained event).
/// When nothing is retained (the replay ring is empty), this is the next sequence
/// that can be delivered — `highest observed + 1` — and the `oldest - 1` resume
/// formula remains valid: it resolves to the highest sequence already seen, so the
/// follow-up resume replays nothing, reports no gap, and every newer live event
/// passes. The interval evicted is unchanged.
/// </summary>
[global::System.Diagnostics.DebuggerNonUserCodeAttribute]
[global::System.CodeDom.Compiler.GeneratedCode("protoc", null)]
@@ -1164,6 +1164,9 @@ namespace ZB.MOM.WW.MxGateway.Contracts.Proto {
/// instead of a hard-coded default; 0 (an older gateway that never set the field) means
/// "use the worker's built-in default". Sits above the public gRPC cap by an
/// envelope-overhead margin so an accepted gRPC payload always fits one worker frame.
/// Every worker->gateway frame — events, heartbeats, faults, and control replies
/// including DrainEvents — must serialize within this limit; reply builders truncate
/// to fit rather than emit an oversized frame.
/// </summary>
[global::System.Diagnostics.DebuggerNonUserCodeAttribute]
[global::System.CodeDom.Compiler.GeneratedCode("protoc", null)]
@@ -256,6 +256,9 @@ message Write2Command {
int32 user_id = 5;
}
// The unary reply's statuses field carries the correlated OnWriteComplete
// outcome when it arrives within the worker's bounded wait see
// MxCommandReply.statuses.
message WriteSecuredCommand {
int32 server_handle = 1;
int32 item_handle = 2;
@@ -266,6 +269,9 @@ message WriteSecuredCommand {
MxValue value = 5;
}
// The unary reply's statuses field carries the correlated OnWriteComplete
// outcome when it arrives within the worker's bounded wait see
// MxCommandReply.statuses.
message WriteSecured2Command {
int32 server_handle = 1;
int32 item_handle = 2;
@@ -525,6 +531,18 @@ message MxCommandReply {
// transport failures.
optional int32 hresult = 5;
MxValue return_value = 6;
// Correlated per-item outcome rows. For WRITE_SECURED / WRITE_SECURED2
// replies the worker holds the reply for a bounded window (default 1.5 s,
// MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS) waiting for the matching
// MXAccess OnWriteComplete callback and copies its status rows here, so
// statuses[0] carries the real MXAccess commit outcome (success OR failure)
// while protocol_status/hresult still describe command acceptance only.
// Empty statuses on a write reply means the completion did not arrive
// within the window the write is unconfirmed, not failed. Correlation is
// best-effort per (server_handle, item_handle): MXAccess's callback carries
// no transaction id, so concurrent writes to the same item within the
// window can swap rows. The OnWriteComplete event still flows on the event
// stream unchanged. Other command kinds leave this field as before.
repeated MxStatusProxy statuses = 7;
string diagnostic_message = 8;
@@ -676,6 +694,10 @@ message WorkerInfoReply {
}
message DrainEventsReply {
// The reply is bounded by both a server-side count cap and the negotiated
// worker-frame byte cap; a reply may therefore carry fewer events than
// `max_events` and fewer than are queued. Callers drain iteratively until an
// empty reply.
repeated MxEvent events = 1;
}
@@ -760,6 +782,11 @@ message ReplayGap {
// after_worker_sequence = oldest_available_sequence - 1 in the next
// StreamEventsRequest, which will cause the server to replay starting at
// oldest_available_sequence (the first retained event).
// When nothing is retained (the replay ring is empty), this is the next sequence
// that can be delivered `highest observed + 1` and the `oldest - 1` resume
// formula remains valid: it resolves to the highest sequence already seen, so the
// follow-up resume replays nothing, reports no gap, and every newer live event
// passes. The interval evicted is unchanged.
uint64 oldest_available_sequence = 2;
}
@@ -47,6 +47,9 @@ message GatewayHello {
// instead of a hard-coded default; 0 (an older gateway that never set the field) means
// "use the worker's built-in default". Sits above the public gRPC cap by an
// envelope-overhead margin so an accepted gRPC payload always fits one worker frame.
// Every worker->gateway frame events, heartbeats, faults, and control replies
// including DrainEvents must serialize within this limit; reply builders truncate
// to fit rather than emit an oversized frame.
uint32 max_frame_bytes = 4;
}
@@ -7,7 +7,7 @@
<PropertyGroup>
<IsPackable>true</IsPackable>
<PackageId>ZB.MOM.WW.MxGateway.Contracts</PackageId>
<Version>0.1.2</Version>
<Version>0.2.0</Version>
<Authors>Joseph Doherty</Authors>
<Company>ZB MOM WW</Company>
<Copyright>Copyright (c) ZB MOM WW. All rights reserved.</Copyright>
@@ -14,7 +14,19 @@ namespace ZB.MOM.WW.MxGateway.IntegrationTests;
[Trait("Category", "LiveLdap")]
public sealed class DashboardLdapLiveTests
{
/// <summary>Verifies that an admin user in the GwAdmin group authenticates successfully.</summary>
/// <summary>
/// The shared dev/test directory issues every human tester the same well-known password, so
/// the fixtures name it once rather than repeating a literal that drifts per test. This is a
/// published dev credential (see <c>glauth.md</c> and <c>scadaproj/infra/glauth/config.toml</c>),
/// not a secret — unlike the service-account bind password, which is never in source and must
/// arrive via <c>MxGateway__Ldap__ServiceAccountPassword</c>.
/// </summary>
private const string SharedDirectoryPassword = "password";
/// <summary>
/// Verifies that <c>admin</c> — a shared-directory user whose <c>othergroups</c> include
/// GwAdmin (gid 5610) — authenticates successfully and is granted the Admin dashboard role.
/// </summary>
/// <returns>A task that represents the asynchronous operation.</returns>
[LiveLdapFact]
public async Task AuthenticateAsync_AdminInGwAdminGroup_Succeeds()
@@ -23,7 +35,7 @@ public sealed class DashboardLdapLiveTests
DashboardAuthenticationResult result = await authenticator.AuthenticateAsync(
"admin",
"admin123",
SharedDirectoryPassword,
CancellationToken.None);
Assert.True(result.Succeeded);
@@ -38,21 +50,43 @@ public sealed class DashboardLdapLiveTests
&& claim.Value == DashboardRoles.Admin);
}
/// <summary>Verifies that a readonly user without GwAdmin group fails to authenticate.</summary>
/// <summary>
/// Verifies that <c>gw-viewer</c> — a shared-directory user whose only group is GwReader
/// (gid 5611), which this suite's GroupToRole map deliberately leaves unmapped — is denied
/// even though its bind succeeds, and that the denial is indistinguishable from the
/// unknown-user denial.
/// </summary>
/// <returns>A task that represents the asynchronous operation.</returns>
[LiveLdapFact]
public async Task AuthenticateAsync_ReadOnlyUserMissingGwAdminGroup_Fails()
public async Task AuthenticateAsync_ViewerMissingGwAdminGroup_FailsIndistinguishably()
{
DashboardAuthenticator authenticator = CreateAuthenticator();
DashboardAuthenticationResult result = await authenticator.AuthenticateAsync(
"readonly",
"readonly123",
"gw-viewer",
SharedDirectoryPassword,
CancellationToken.None);
Assert.False(result.Succeeded);
Assert.Null(result.Principal);
Assert.DoesNotContain("readonly123", result.FailureMessage, StringComparison.Ordinal);
// This test used to assert the failure message did not echo the credential literal.
// That check cannot survive the move to the shared directory: the real password is the
// word "password", which legitimately occurs in the generic denial text ("The username
// or password is invalid, ..."), so the assertion would fail for the wrong reason. The
// no-leak property is still covered — with a distinctive literal — by
// AuthenticateAsync_AdminWithWrongPassword_FailsWithoutLeakingPassword below. What is
// asserted here instead is the property this fixture is actually uniquely able to prove:
// an authorization failure (valid credentials, no mapped role) must be reported with the
// same message as an authentication failure, so the response cannot be used to enumerate
// valid accounts.
DashboardAuthenticationResult unknownUserResult = await authenticator.AuthenticateAsync(
"no-such-user-9f3c1",
"irrelevant-password",
CancellationToken.None);
Assert.False(string.IsNullOrWhiteSpace(result.FailureMessage));
Assert.Equal(unknownUserResult.FailureMessage, result.FailureMessage);
}
/// <summary>Verifies that authentication with wrong password fails without leaking the password.</summary>
@@ -98,9 +132,11 @@ public sealed class DashboardLdapLiveTests
[LiveLdapFact]
public async Task AuthenticateAsync_ServerUnreachable_FailsWithoutThrowing()
{
// Exercises the connect-failure path: a closed loopback port produces a
// connection error that the shared LdapAuthService must absorb into a Fail
// result rather than propagating an exception to the dashboard.
// Exercises the connect-failure path: overriding only the port keeps whatever host
// the run targets (localhost by default, the shared GLAuth under the
// MxGateway__Ldap__Server override) while pointing at a port nothing listens on, so
// the connection error the shared LdapAuthService must absorb into a Fail result —
// rather than propagate as an exception to the dashboard — is reproduced either way.
DashboardAuthenticator authenticator = CreateAuthenticator(LibraryOptions() with
{
// 1 is a reserved port number that no LDAP server listens on.
@@ -109,7 +145,7 @@ public sealed class DashboardLdapLiveTests
DashboardAuthenticationResult result = await authenticator.AuthenticateAsync(
"admin",
"admin123",
SharedDirectoryPassword,
CancellationToken.None);
Assert.False(result.Succeeded);
@@ -147,8 +183,10 @@ public sealed class DashboardLdapLiveTests
/// <see cref="LibraryLdapOptions.ConnectionTimeoutMs"/>, which governs the
/// unreachable-server test's timing) at whatever value the operator configured, and
/// cannot silently drop a field added to the shared type. The gateway's
/// <c>appsettings.json</c> seeds the dev directory connection (localhost:3893,
/// plaintext, AllowInsecure).
/// <c>appsettings.json</c> seeds the dev directory connection (port 3893, plaintext,
/// AllowInsecure) but ships <c>Server=localhost</c>, so a run against the shared GLAuth
/// needs the <c>MxGateway__Ldap__Server=10.100.0.35</c> environment override that the
/// <c>AddEnvironmentVariables()</c> layer below applies.
/// </summary>
private static LibraryLdapOptions LibraryOptions()
{
@@ -145,7 +145,12 @@ public sealed class GatewayOptionsValidator : OptionsValidatorBase<GatewayOption
builder);
AddIfBlank(
options.ServiceAccountPassword,
"MxGateway:Ldap:ServiceAccountPassword is required when LDAP login is enabled.",
"MxGateway:Ldap:ServiceAccountPassword is required when LDAP login is enabled. "
+ "Never commit it: on dev boxes set user-secrets "
+ "(dotnet user-secrets set \"MxGateway:Ldap:ServiceAccountPassword\" <value>); "
+ "on deployed hosts set the environment variable "
+ "MxGateway__Ldap__ServiceAccountPassword. "
+ "(appsettings.json ships the ${secret:ldap/mxgateway/bind} store reference as the default.)",
builder);
AddIfBlank(
options.UserNameAttribute,
@@ -220,6 +225,10 @@ public sealed class GatewayOptionsValidator : OptionsValidatorBase<GatewayOption
options.PipeConnectAttemptTimeoutMilliseconds,
"MxGateway:Worker:PipeConnectAttemptTimeoutMilliseconds must be greater than zero.",
builder);
if (options.WriteCompletionWaitMilliseconds < 0)
{
builder.Add("MxGateway:Worker:WriteCompletionWaitMilliseconds must be greater than or equal to zero.");
}
AddIfNotPositive(
options.ShutdownTimeoutSeconds,
"MxGateway:Worker:ShutdownTimeoutSeconds must be greater than zero.",
@@ -24,6 +24,15 @@ public sealed class WorkerOptions
/// <summary>The timeout in milliseconds for connecting to the worker pipe.</summary>
public int PipeConnectAttemptTimeoutMilliseconds { get; init; } = 2000;
/// <summary>
/// Bounded wait, in milliseconds, the worker holds a WriteSecured/WriteSecured2
/// reply for the matching MXAccess OnWriteComplete callback so the reply's
/// statuses carry the real commit outcome. 0 disables the wait. Deployments
/// raising this above consumer write-timeout budgets (e.g. OtOpcUa's 2 s Tier A
/// write resilience timeout) must raise those in step.
/// </summary>
public int WriteCompletionWaitMilliseconds { get; init; } = 1500;
/// <summary>The maximum time in seconds for graceful shutdown.</summary>
public int ShutdownTimeoutSeconds { get; init; } = 10;
@@ -21,6 +21,14 @@ public sealed class WorkerProcessLauncher : IWorkerProcessLauncher
public const string WorkerPipeConnectAttemptTimeoutEnvironmentVariableName =
"MXGATEWAY_WORKER_PIPE_CONNECT_ATTEMPT_TIMEOUT_MS";
/// <summary>
/// Conveys MxGateway:Worker:WriteCompletionWaitMilliseconds to the worker:
/// the bounded wait for the OnWriteComplete callback on
/// WriteSecured/WriteSecured2 replies. 0 disables the wait.
/// </summary>
public const string WorkerWriteCompletionWaitEnvironmentVariableName =
"MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS";
private readonly IWorkerProcessFactory _processFactory;
private readonly IWorkerStartupProbe _startupProbe;
private readonly GatewayMetrics _metrics;
@@ -175,6 +183,8 @@ public sealed class WorkerProcessLauncher : IWorkerProcessLauncher
startInfo.Environment[WorkerNonceEnvironmentVariableName] = request.Nonce;
startInfo.Environment[WorkerPipeConnectAttemptTimeoutEnvironmentVariableName] =
_workerOptions.PipeConnectAttemptTimeoutMilliseconds.ToString(System.Globalization.CultureInfo.InvariantCulture);
startInfo.Environment[WorkerWriteCompletionWaitEnvironmentVariableName] =
_workerOptions.WriteCompletionWaitMilliseconds.ToString(System.Globalization.CultureInfo.InvariantCulture);
commandLine = new WorkerProcessCommandLine(executablePath, arguments);
@@ -2,6 +2,10 @@
<PropertyGroup>
<TargetFramework>net10.0</TargetFramework>
<!-- Dev-box channel for MxGateway:Ldap:ServiceAccountPassword (SEC-36): user-secrets
are loaded automatically in the Development environment and live under the user
profile, outside the tree, so the shared GLAuth bind credential is never committed. -->
<UserSecretsId>mxaccessgw-server</UserSecretsId>
</PropertyGroup>
<ItemGroup>
@@ -41,6 +41,7 @@ public sealed class GatewayOptionsTests
Assert.Equal(3, options.Worker.StartupProbeRetryAttempts);
Assert.Equal(250, options.Worker.StartupProbeRetryDelayMilliseconds);
Assert.Equal(2000, options.Worker.PipeConnectAttemptTimeoutMilliseconds);
Assert.Equal(1500, options.Worker.WriteCompletionWaitMilliseconds);
Assert.Equal(10, options.Worker.ShutdownTimeoutSeconds);
Assert.Equal(5, options.Worker.HeartbeatIntervalSeconds);
Assert.Equal(15, options.Worker.HeartbeatGraceSeconds);
@@ -106,6 +107,7 @@ public sealed class GatewayOptionsTests
[InlineData("MxGateway:Worker:ExecutablePath", "worker.dll", "MxGateway:Worker:ExecutablePath must point to a .exe file.")]
[InlineData("MxGateway:Worker:StartupProbeRetryAttempts", "0", "MxGateway:Worker:StartupProbeRetryAttempts must be greater than zero.")]
[InlineData("MxGateway:Worker:PipeConnectAttemptTimeoutMilliseconds", "0", "MxGateway:Worker:PipeConnectAttemptTimeoutMilliseconds must be greater than zero.")]
[InlineData("MxGateway:Worker:WriteCompletionWaitMilliseconds", "-1", "MxGateway:Worker:WriteCompletionWaitMilliseconds must be greater than or equal to zero.")]
[InlineData("MxGateway:Sessions:DefaultLeaseSeconds", "0", "MxGateway:Sessions:DefaultLeaseSeconds must be greater than zero.")]
[InlineData("MxGateway:Sessions:LeaseSweepIntervalSeconds", "0", "MxGateway:Sessions:LeaseSweepIntervalSeconds must be greater than zero.")]
[InlineData("MxGateway:Sessions:DetachGraceSeconds", "-1", "MxGateway:Sessions:DetachGraceSeconds must be zero or greater (0 disables detach-grace retention).")]
@@ -757,9 +757,15 @@ public sealed class GatewayOptionsValidatorTests
new LdapOptions { Enabled = true, ServiceAccountPassword = string.Empty });
ValidateOptionsResult result = new GatewayOptionsValidator().Validate(null, options);
Assert.True(result.Failed);
Assert.Contains(
string failure = Assert.Single(
result.Failures!,
f => f.Contains("MxGateway:Ldap:ServiceAccountPassword is required when LDAP login is enabled."));
// SEC-36: the message must steer the operator to the two supported out-of-band channels
// (dev user-secrets, deployed env var) so a blanked/unresolved credential never gets
// "fixed" by re-committing a value.
Assert.Contains("dotnet user-secrets set", failure);
Assert.Contains("MxGateway__Ldap__ServiceAccountPassword", failure);
}
private static GatewayOptions WithSecurity(SecurityOptions security)
@@ -72,18 +72,33 @@ public sealed class ClientBehaviorFixtureTests
foreach (JsonElement fixture in fixtures)
{
string fixtureId = GetFixtureId(fixture);
MxCommandReply reply = ParseFixture<MxCommandReply>(
fixture,
MxCommandReply.Parser);
// Universal invariants: every command-reply fixture parses to a concrete
// command kind and a concrete protocol status, regardless of what MXAccess
// reply detail (if any) it carries.
Assert.NotEqual(MxCommandKind.Unspecified, reply.Kind);
Assert.NotEqual(ProtocolStatusCode.Unspecified, reply.ProtocolStatus.Code);
Assert.True(reply.HasHresult, $"Fixture '{GetFixtureId(fixture)}' must carry an HRESULT.");
// The malformed-reply and credential-redaction fixtures
// (command-reply.authenticate-user.*) deliberately omit hresult, statuses,
// and/or return_value to exercise the "absent detail" contract paths — see
// docs/ClientBehaviorFixtures.md. The strict MXAccess-reply-detail
// invariants below apply only to fixtures that carry that detail.
if (fixtureId.StartsWith("command-reply.authenticate-user.", StringComparison.Ordinal))
{
continue;
}
Assert.True(reply.HasHresult, $"Fixture '{fixtureId}' must carry an HRESULT.");
Assert.NotEmpty(reply.Statuses);
Assert.NotEqual(MxDataType.Unspecified, reply.ReturnValue.DataType);
Assert.True(
reply.ReturnValue.KindCase != MxValue.KindOneofCase.None || reply.ReturnValue.IsNull,
$"Fixture '{GetFixtureId(fixture)}' must carry a typed value, raw value, or explicit null.");
$"Fixture '{fixtureId}' must carry a typed value, raw value, or explicit null.");
}
MxCommandReply failedWrite = ParseFixture<MxCommandReply>(
@@ -3,22 +3,24 @@ using Google.Protobuf;
using Google.Protobuf.Reflection;
using ZB.MOM.WW.MxGateway.Contracts;
using ZB.MOM.WW.MxGateway.Contracts.Proto;
using ZB.MOM.WW.MxGateway.Contracts.Proto.Galaxy;
namespace ZB.MOM.WW.MxGateway.Tests.Contracts;
public sealed class ClientProtoInputTests
{
/// <summary>
/// Guards the published client descriptor set against silent staleness. Every message
/// and field compiled into the in-process contract (which the build regenerates from the current
/// <c>.proto</c> sources) must appear in the committed protoset. A missing symbol means the
/// descriptor was not regenerated after a proto change; run
/// <c>scripts/publish-client-proto-inputs.ps1</c> and commit the refreshed protoset.
/// Guards the published client descriptor set against silent staleness. Every message,
/// field, enum, enum value, service, and method compiled into the in-process contract
/// (which the build regenerates from the current <c>.proto</c> sources) must appear in the
/// committed protoset. A missing symbol means the descriptor was not regenerated after a
/// proto change; run <c>scripts/publish-client-proto-inputs.ps1</c> and commit the
/// refreshed protoset.
/// The check is semantic (symbol presence) rather than byte-wise, so it is independent of protoc
/// version and does not require protoc on the test runner.
/// </summary>
[Fact]
public void Descriptor_ContainsEveryContractMessageAndField()
public void Descriptor_ContainsEveryContractSymbol()
{
DirectoryInfo repositoryRoot = FindRepositoryRoot();
string descriptorPath = Path.Combine(
@@ -34,20 +36,65 @@ public sealed class ClientProtoInputTests
HashSet<string> publishedMessages = new(StringComparer.Ordinal);
HashSet<string> publishedFields = new(StringComparer.Ordinal);
HashSet<string> publishedEnums = new(StringComparer.Ordinal);
HashSet<string> publishedServices = new(StringComparer.Ordinal);
foreach (FileDescriptorProto file in descriptorSet.File)
{
foreach (DescriptorProto message in file.MessageType)
{
CollectPublishedSymbols(file.Package, message, publishedMessages, publishedFields);
CollectPublishedSymbols(file.Package, message, publishedMessages, publishedFields, publishedEnums);
}
foreach (EnumDescriptorProto enumType in file.EnumType)
{
CollectPublishedEnumSymbols(file.Package, enumType, publishedEnums);
}
foreach (ServiceDescriptorProto service in file.Service)
{
string serviceFullName = string.IsNullOrEmpty(file.Package) ? service.Name : file.Package + "." + service.Name;
publishedServices.Add(serviceFullName);
foreach (MethodDescriptorProto method in service.Method)
{
publishedServices.Add(serviceFullName + "/" + method.Name);
}
}
}
List<string> missing = [];
foreach (FileDescriptor file in new[] { MxaccessGatewayReflection.Descriptor, MxaccessWorkerReflection.Descriptor })
FileDescriptor[] contractFiles =
[
MxaccessGatewayReflection.Descriptor,
MxaccessWorkerReflection.Descriptor,
GalaxyRepositoryReflection.Descriptor,
];
foreach (FileDescriptor file in contractFiles)
{
foreach (MessageDescriptor message in file.MessageTypes)
{
CollectMissingContractSymbols(message, publishedMessages, publishedFields, missing);
CollectMissingContractSymbols(message, publishedMessages, publishedFields, publishedEnums, missing);
}
foreach (EnumDescriptor enumType in file.EnumTypes)
{
CollectMissingEnumSymbols(enumType, publishedEnums, missing);
}
foreach (ServiceDescriptor service in file.Services)
{
if (!publishedServices.Contains(service.FullName))
{
missing.Add(service.FullName);
}
foreach (MethodDescriptor method in service.Methods)
{
string key = service.FullName + "/" + method.Name;
if (!publishedServices.Contains(key))
{
missing.Add(key);
}
}
}
}
@@ -62,7 +109,8 @@ public sealed class ClientProtoInputTests
string package,
DescriptorProto message,
HashSet<string> messages,
HashSet<string> fields)
HashSet<string> fields,
HashSet<string> enums)
{
string fullName = string.IsNullOrEmpty(package) ? message.Name : package + "." + message.Name;
messages.Add(fullName);
@@ -74,7 +122,26 @@ public sealed class ClientProtoInputTests
foreach (DescriptorProto nested in message.NestedType)
{
CollectPublishedSymbols(fullName, nested, messages, fields);
CollectPublishedSymbols(fullName, nested, messages, fields, enums);
}
foreach (EnumDescriptorProto enumType in message.EnumType)
{
CollectPublishedEnumSymbols(fullName, enumType, enums);
}
}
private static void CollectPublishedEnumSymbols(
string containingScope,
EnumDescriptorProto enumType,
HashSet<string> enums)
{
string enumFullName = string.IsNullOrEmpty(containingScope) ? enumType.Name : containingScope + "." + enumType.Name;
enums.Add(enumFullName);
foreach (EnumValueDescriptorProto value in enumType.Value)
{
enums.Add(enumFullName + "/" + value.Name);
}
}
@@ -82,6 +149,7 @@ public sealed class ClientProtoInputTests
MessageDescriptor message,
HashSet<string> publishedMessages,
HashSet<string> publishedFields,
HashSet<string> publishedEnums,
List<string> missing)
{
if (!publishedMessages.Contains(message.FullName))
@@ -100,7 +168,32 @@ public sealed class ClientProtoInputTests
foreach (MessageDescriptor nested in message.NestedTypes)
{
CollectMissingContractSymbols(nested, publishedMessages, publishedFields, missing);
CollectMissingContractSymbols(nested, publishedMessages, publishedFields, publishedEnums, missing);
}
foreach (EnumDescriptor enumType in message.EnumTypes)
{
CollectMissingEnumSymbols(enumType, publishedEnums, missing);
}
}
private static void CollectMissingEnumSymbols(
EnumDescriptor enumType,
HashSet<string> publishedEnums,
List<string> missing)
{
if (!publishedEnums.Contains(enumType.FullName))
{
missing.Add(enumType.FullName);
}
foreach (EnumValueDescriptor value in enumType.Values)
{
string key = enumType.FullName + "/" + value.Name;
if (!publishedEnums.Contains(key))
{
missing.Add(key);
}
}
}
@@ -43,6 +43,10 @@ public sealed class WorkerProcessLauncherTests
"2000",
processFactory.LastStartInfo.Environment[
WorkerProcessLauncher.WorkerPipeConnectAttemptTimeoutEnvironmentVariableName]);
Assert.Equal(
"1500",
processFactory.LastStartInfo.Environment[
WorkerProcessLauncher.WorkerWriteCompletionWaitEnvironmentVariableName]);
Assert.DoesNotContain(Nonce, handle.CommandLine.ToString(), StringComparison.Ordinal);
Assert.DoesNotContain(Nonce, string.Join(" ", handle.CommandLine.Arguments), StringComparison.Ordinal);
Assert.False(pipeReservation.DisposeCalled);
@@ -110,6 +110,47 @@ public sealed class MxAccessBaseEventSinkTests
Assert.Same(cache, sink.ValueCache);
}
/// <summary>
/// Verifies that an OnWriteComplete COM callback records the completion into
/// the per-session write-completion cache for the executor's bounded reply
/// wait AND still enqueues the event unchanged for the outbound stream —
/// correlation observes the event, it never consumes it. The cache update
/// fires only after the event has cleared the queue (post-publish rule).
/// </summary>
[Fact]
public void OnWriteComplete_ComCallback_RecordsCompletionAndStillEnqueuesEvent()
{
MxAccessEventQueue queue = new();
MxAccessWriteCompletionCache completionCache = new();
MxAccessBaseEventSink sink = new(queue, new MxAccessEventMapper(), new MxAccessValueCache(), completionCache);
MXSTATUS_PROXY[] statuses = Array.Empty<MXSTATUS_PROXY>();
sink.OnWriteComplete(hLMXServerHandle: 7, phItemHandle: 21, ref statuses);
Assert.Equal(1, queue.Count);
Assert.True(queue.TryDequeue(out WorkerEvent? workerEvent));
MxEvent mxEvent = workerEvent!.Event;
Assert.Equal(MxEventFamily.OnWriteComplete, mxEvent.Family);
Assert.Equal(7, mxEvent.ServerHandle);
Assert.Equal(21, mxEvent.ItemHandle);
Assert.Equal(1UL, completionCache.CurrentVersion(7, 21));
}
/// <summary>
/// Verifies that the sink-bound write-completion cache is exposed for sharing
/// with the owning <see cref="MxAccessSession"/> so the sink's recordings and
/// the write executor's waits see the same instance.
/// </summary>
[Fact]
public void WriteCompletionCache_ReturnsTheInstanceBoundAtConstruction()
{
MxAccessEventQueue queue = new();
MxAccessWriteCompletionCache completionCache = new();
MxAccessBaseEventSink sink = new(queue, new MxAccessEventMapper(), new MxAccessValueCache(), completionCache);
Assert.Same(completionCache, sink.WriteCompletionCache);
}
/// <summary>
/// Verifies that consecutive OnDataChange callbacks land in the queue with monotonic sequences.
/// </summary>
@@ -897,6 +897,9 @@ public sealed class MxAccessCommandExecutorTests
FakeMxAccessComObjectFactory factory = new(fakeComObject);
using StaRuntime runtime = CreateRuntime();
using MxAccessStaSession session = new(runtime, factory, new NoopEventSink());
// No completion source in this test — disable the bounded reply wait
// so the forwarding assertions don't pay the default 1.5 s timeout.
session.WriteCompletionTimeout = TimeSpan.Zero;
await session.StartAsync(workerProcessId: 1234);
MxCommandReply reply = await session.DispatchAsync(CreateWriteSecuredCommand(
@@ -919,6 +922,8 @@ public sealed class MxAccessCommandExecutorTests
FakeMxAccessComObjectFactory factory = new(fakeComObject);
using StaRuntime runtime = CreateRuntime();
using MxAccessStaSession session = new(runtime, factory, new NoopEventSink());
// Same rationale as the WriteSecured forwarding test above.
session.WriteCompletionTimeout = TimeSpan.Zero;
await session.StartAsync(workerProcessId: 1234);
DateTime timestamp = new(2026, 5, 19, 13, 30, 0, DateTimeKind.Utc);
@@ -934,6 +939,192 @@ public sealed class MxAccessCommandExecutorTests
Assert.Equal(44, fakeComObject.WriteVerifierUserId);
}
/// <summary>
/// Verifies the fast-completion ordering edge: a completion recorded while
/// the WriteSecured COM call is still on the stack (MXAccess committing
/// synchronously) is newer than the pre-call baseline and lands on the
/// reply — the wait never misses a callback that beat it.
/// </summary>
/// <returns>A task that represents the asynchronous operation.</returns>
[Fact]
public async Task DispatchAsync_WriteSecured_WhenCompletionArrivesDuringComCall_ReturnsStatuses()
{
FakeMxAccessComObject fakeComObject = new(registerHandle: 82);
FakeMxAccessComObjectFactory factory = new(fakeComObject);
CompletionCacheEventSink sink = new();
fakeComObject.OnWriteSecuredCallback = () =>
sink.WriteCompletionCache.Record(82, 820, CreateCompletionRows(detail: 4321));
using StaRuntime runtime = CreateRuntime();
using MxAccessStaSession session = new(runtime, factory, sink);
// Hermetic: don't inherit MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS
// from the test runner's environment.
session.WriteCompletionTimeout = TimeSpan.FromSeconds(10);
await session.StartAsync(workerProcessId: 1234);
MxCommandReply reply = await session.DispatchAsync(CreateWriteSecuredCommand(
"write-secured-fast", serverHandle: 82, itemHandle: 820, value: 1, currentUserId: 11, verifierUserId: 22));
Assert.Equal(ProtocolStatusCode.Ok, reply.ProtocolStatus.Code);
Assert.True(reply.HasHresult);
Assert.Equal(0, reply.Hresult);
MxStatusProxy row = Assert.Single(reply.Statuses);
Assert.Equal(4321, row.Detail);
Assert.Equal(MxStatusCategory.Ok, row.Category);
}
/// <summary>
/// Verifies the pump-wait path: the completion arrives after the COM call
/// returned, while the executor is pump-waiting, and still lands on the
/// reply.
/// </summary>
/// <returns>A task that represents the asynchronous operation.</returns>
[Fact]
public async Task DispatchAsync_WriteSecured_WhenCompletionArrivesWhileWaiting_ReturnsStatuses()
{
FakeMxAccessComObject fakeComObject = new(registerHandle: 83);
FakeMxAccessComObjectFactory factory = new(fakeComObject);
CompletionCacheEventSink sink = new();
// Deterministic ordering: the executor captures its version baseline
// BEFORE the COM call, so once the fake's WriteSecured has run the
// baseline is committed and a Record from the test thread is
// guaranteed to be "newer" — no fixed sleep racing the STA thread.
using System.Threading.ManualResetEventSlim comCallReached = new(initialState: false);
fakeComObject.OnWriteSecuredCallback = () => comCallReached.Set();
using StaRuntime runtime = CreateRuntime();
using MxAccessStaSession session = new(runtime, factory, sink);
session.WriteCompletionTimeout = TimeSpan.FromSeconds(10);
await session.StartAsync(workerProcessId: 1234);
Task<MxCommandReply> pending = session.DispatchAsync(CreateWriteSecuredCommand(
"write-secured-waiting", serverHandle: 83, itemHandle: 830, value: 1, currentUserId: 11, verifierUserId: 22));
Assert.True(comCallReached.Wait(TimeSpan.FromSeconds(5)));
sink.WriteCompletionCache.Record(83, 830, CreateCompletionRows(detail: 99));
MxCommandReply reply = await pending;
Assert.Equal(ProtocolStatusCode.Ok, reply.ProtocolStatus.Code);
Assert.Equal(99, Assert.Single(reply.Statuses).Detail);
}
/// <summary>
/// Verifies the timeout fallback: no completion within the bounded wait
/// returns today's reply shape — protocol OK with EMPTY statuses (the
/// consumer's honest-unconfirmed path), never a synthesized failure row.
/// </summary>
/// <returns>A task that represents the asynchronous operation.</returns>
[Fact]
public async Task DispatchAsync_WriteSecured_WhenNoCompletion_TimesOutWithEmptyStatusesAndOkProtocol()
{
FakeMxAccessComObject fakeComObject = new(registerHandle: 84);
FakeMxAccessComObjectFactory factory = new(fakeComObject);
CompletionCacheEventSink sink = new();
using StaRuntime runtime = CreateRuntime();
using MxAccessStaSession session = new(runtime, factory, sink);
session.WriteCompletionTimeout = TimeSpan.FromMilliseconds(100);
await session.StartAsync(workerProcessId: 1234);
MxCommandReply reply = await session.DispatchAsync(CreateWriteSecuredCommand(
"write-secured-timeout", serverHandle: 84, itemHandle: 840, value: 1, currentUserId: 11, verifierUserId: 22));
Assert.Equal(ProtocolStatusCode.Ok, reply.ProtocolStatus.Code);
Assert.True(reply.HasHresult);
Assert.Equal(0, reply.Hresult);
Assert.Empty(reply.Statuses);
}
/// <summary>
/// Verifies the version-baseline rule end to end: a completion recorded
/// BEFORE the write was dispatched is stale and must not be misattributed
/// to this write — the reply times out empty instead.
/// </summary>
/// <returns>A task that represents the asynchronous operation.</returns>
[Fact]
public async Task DispatchAsync_WriteSecured_IgnoresStaleCompletionFromBeforeTheCall()
{
FakeMxAccessComObject fakeComObject = new(registerHandle: 85);
FakeMxAccessComObjectFactory factory = new(fakeComObject);
CompletionCacheEventSink sink = new();
sink.WriteCompletionCache.Record(85, 850, CreateCompletionRows(detail: 1111));
using StaRuntime runtime = CreateRuntime();
using MxAccessStaSession session = new(runtime, factory, sink);
session.WriteCompletionTimeout = TimeSpan.FromMilliseconds(100);
await session.StartAsync(workerProcessId: 1234);
MxCommandReply reply = await session.DispatchAsync(CreateWriteSecuredCommand(
"write-secured-stale", serverHandle: 85, itemHandle: 850, value: 1, currentUserId: 11, verifierUserId: 22));
Assert.Equal(ProtocolStatusCode.Ok, reply.ProtocolStatus.Code);
Assert.Empty(reply.Statuses);
}
/// <summary>
/// Verifies that WriteSecured2 correlates the same way as WriteSecured
/// (fast-completion edge).
/// </summary>
/// <returns>A task that represents the asynchronous operation.</returns>
[Fact]
public async Task DispatchAsync_WriteSecured2_WhenCompletionArrivesDuringComCall_ReturnsStatuses()
{
FakeMxAccessComObject fakeComObject = new(registerHandle: 86);
FakeMxAccessComObjectFactory factory = new(fakeComObject);
CompletionCacheEventSink sink = new();
fakeComObject.OnWriteSecuredCallback = () =>
sink.WriteCompletionCache.Record(86, 860, CreateCompletionRows(detail: 2222));
using StaRuntime runtime = CreateRuntime();
using MxAccessStaSession session = new(runtime, factory, sink);
// Hermetic: same rationale as the WriteSecured fast-completion test.
session.WriteCompletionTimeout = TimeSpan.FromSeconds(10);
await session.StartAsync(workerProcessId: 1234);
MxCommandReply reply = await session.DispatchAsync(CreateWriteSecured2Command(
"write-secured2-fast", serverHandle: 86, itemHandle: 860, value: 1,
timestamp: new DateTime(2026, 8, 9, 12, 0, 0, DateTimeKind.Utc), currentUserId: 33, verifierUserId: 44));
Assert.Equal(ProtocolStatusCode.Ok, reply.ProtocolStatus.Code);
Assert.Equal(2222, Assert.Single(reply.Statuses).Detail);
}
/// <summary>
/// Verifies plain Write stays fire-and-forget: even with a huge completion
/// timeout configured and no completion source, the reply returns
/// immediately (guarded well under the configured wait) with empty
/// statuses — only the secured write kinds enter the bounded wait.
/// </summary>
/// <returns>A task that represents the asynchronous operation.</returns>
[Fact]
public async Task DispatchAsync_Write_DoesNotWaitForCompletion()
{
FakeMxAccessComObject fakeComObject = new(registerHandle: 87);
FakeMxAccessComObjectFactory factory = new(fakeComObject);
CompletionCacheEventSink sink = new();
using StaRuntime runtime = CreateRuntime();
using MxAccessStaSession session = new(runtime, factory, sink);
session.WriteCompletionTimeout = TimeSpan.FromSeconds(30);
await session.StartAsync(workerProcessId: 1234);
Task<MxCommandReply> pending = session.DispatchAsync(CreateWriteCommand(
"plain-write-no-wait", serverHandle: 87, itemHandle: 870, value: 1, userId: 5));
Task completed = await Task.WhenAny(pending, Task.Delay(TimeSpan.FromSeconds(5)));
Assert.Same(pending, completed);
MxCommandReply reply = await pending;
Assert.Equal(ProtocolStatusCode.Ok, reply.ProtocolStatus.Code);
Assert.Empty(reply.Statuses);
}
private static Google.Protobuf.Collections.RepeatedField<MxStatusProxy> CreateCompletionRows(int detail)
{
return new Google.Protobuf.Collections.RepeatedField<MxStatusProxy>
{
new MxStatusProxy
{
Success = 1,
Category = MxStatusCategory.Ok,
Detail = detail,
},
};
}
/// <summary>Verifies that Write without a payload returns an invalid request error.</summary>
/// <returns>A task that represents the asynchronous operation.</returns>
[Fact]
@@ -1775,6 +1966,28 @@ public sealed class MxAccessCommandExecutorTests
TimeSpan.FromMilliseconds(25));
}
/// <summary>
/// Test sink that owns a real write-completion cache without touching the
/// MXAccess COM RCW (Attach is a no-op). Implements the provider seam so
/// <see cref="MxAccessSession.Create"/> shares this cache with the write
/// executor, letting tests record completions the executor's bounded
/// wait then observes.
/// </summary>
private sealed class CompletionCacheEventSink : IMxAccessEventSink, IWriteCompletionCacheProvider
{
public MxAccessWriteCompletionCache WriteCompletionCache { get; } = new MxAccessWriteCompletionCache();
public void Attach(
object mxAccessComObject,
string sessionId)
{
}
public void Detach()
{
}
}
private sealed class FakeMxAccessComObject : IMxAccessServer
{
private readonly int registerHandle;
@@ -1790,6 +2003,14 @@ public sealed class MxAccessCommandExecutorTests
private readonly IReadOnlyDictionary<int, Exception> writeExceptionByItemHandle;
private readonly List<string> operationNames = new();
/// <summary>
/// Invoked at the end of a successful WriteSecured/WriteSecured2 —
/// stands in for MXAccess committing synchronously and delivering
/// OnWriteComplete while the COM call is still on the stack, so
/// tests can exercise the fast-completion ordering edge.
/// </summary>
public Action? OnWriteSecuredCallback { get; set; }
/// <summary>Initializes a fake MXAccess COM object with the given handles and optional exceptions.</summary>
/// <param name="registerHandle">Return value for Register method.</param>
/// <param name="addItemHandle">Return value for AddItem method.</param>
@@ -2100,6 +2321,7 @@ public sealed class MxAccessCommandExecutorTests
WriteValue = value;
WriteThreadId = Environment.CurrentManagedThreadId;
ThrowIfWriteFailureConfigured(itemHandle);
OnWriteSecuredCallback?.Invoke();
}
/// <inheritdoc />
@@ -2120,6 +2342,7 @@ public sealed class MxAccessCommandExecutorTests
WriteTimestamp = timestamp;
WriteThreadId = Environment.CurrentManagedThreadId;
ThrowIfWriteFailureConfigured(itemHandle);
OnWriteSecuredCallback?.Invoke();
}
private void ThrowIfWriteFailureConfigured(int itemHandle)
@@ -15,6 +15,48 @@ namespace ZB.MOM.WW.MxGateway.Worker.Tests.MxAccess;
/// </summary>
public sealed class MxAccessStaSessionTests
{
/// <summary>
/// Verifies the launcher-env-var parse branches of
/// <see cref="MxAccessStaSession.ResolveWriteCompletionTimeout"/>:
/// a valid non-negative value is honored (0 = disabled), while a
/// missing, malformed, or negative value falls back to the executor
/// default. Env mutation is restored in a finally so parallel tests
/// never observe the temporary value.
/// </summary>
/// <param name="rawValue">Raw env-var value, or null for unset.</param>
/// <param name="expectedMilliseconds">Expected resolved wait, or null for the executor default.</param>
[Theory]
[InlineData(null, null)]
[InlineData("", null)]
[InlineData("junk", null)]
[InlineData("-5", null)]
[InlineData("0", 0)]
[InlineData("250", 250)]
public void ResolveWriteCompletionTimeout_ParsesEnvironmentValue(string? rawValue, int? expectedMilliseconds)
{
string? original = Environment.GetEnvironmentVariable(
MxAccessStaSession.WriteCompletionWaitEnvironmentVariableName);
try
{
Environment.SetEnvironmentVariable(
MxAccessStaSession.WriteCompletionWaitEnvironmentVariableName,
rawValue);
TimeSpan resolved = MxAccessStaSession.ResolveWriteCompletionTimeout();
TimeSpan expected = expectedMilliseconds is null
? MxAccessCommandExecutor.DefaultWriteCompletionTimeout
: TimeSpan.FromMilliseconds(expectedMilliseconds.Value);
Assert.Equal(expected, resolved);
}
finally
{
Environment.SetEnvironmentVariable(
MxAccessStaSession.WriteCompletionWaitEnvironmentVariableName,
original);
}
}
/// <summary>
/// Verifies that StartAsync creates the MXAccess COM object and attaches the event sink on the STA thread.
/// </summary>
@@ -0,0 +1,145 @@
using System;
using Google.Protobuf.Collections;
using ZB.MOM.WW.MxGateway.Contracts.Proto;
using ZB.MOM.WW.MxGateway.Worker.MxAccess;
namespace ZB.MOM.WW.MxGateway.Worker.Tests.MxAccess;
/// <summary>
/// Unit tests for <see cref="MxAccessWriteCompletionCache"/>. The cache is
/// consumed by the write command executor's bounded pump-wait so a
/// WriteSecured/WriteSecured2 reply can carry the correlated
/// OnWriteComplete outcome; its version-baseline contract is exercised in
/// isolation here before the STA / COM plumbing gets layered on top.
/// </summary>
public sealed class MxAccessWriteCompletionCacheTests
{
/// <summary>Verifies that Record bumps the version per key and keys stay isolated.</summary>
[Fact]
public void Record_IncrementsVersionPerKey()
{
MxAccessWriteCompletionCache cache = new();
Assert.Equal(0UL, cache.CurrentVersion(7, 21));
cache.Record(7, 21, BuildStatuses(detail: 100));
Assert.Equal(1UL, cache.CurrentVersion(7, 21));
cache.Record(7, 21, BuildStatuses(detail: 200));
Assert.Equal(2UL, cache.CurrentVersion(7, 21));
cache.Record(7, 22, BuildStatuses(detail: 300));
Assert.Equal(1UL, cache.CurrentVersion(7, 22));
Assert.Equal(2UL, cache.CurrentVersion(7, 21));
}
/// <summary>Verifies that a completion newer than the baseline is returned with its status rows.</summary>
[Fact]
public void TryWaitForCompletion_WhenCompletionNewerThanBaseline_ReturnsStatuses()
{
MxAccessWriteCompletionCache cache = new();
cache.Record(7, 21, BuildStatuses(detail: 4321));
bool found = cache.TryWaitForCompletion(
7,
21,
sinceVersion: 0UL,
deadlineUtc: DateTime.UtcNow.AddSeconds(5),
pumpStep: static () => { },
out RepeatedField<MxStatusProxy> statuses);
Assert.True(found);
MxStatusProxy row = Assert.Single(statuses);
Assert.Equal(4321, row.Detail);
Assert.Equal(MxStatusCategory.Ok, row.Category);
}
/// <summary>
/// Verifies that a completion recorded before the baseline was captured is
/// never misattributed to the waiting write: only a strictly newer version
/// satisfies the wait, so a stale row times the wait out.
/// </summary>
[Fact]
public void TryWaitForCompletion_WhenOnlyStaleCompletion_TimesOut()
{
MxAccessWriteCompletionCache cache = new();
cache.Record(7, 21, BuildStatuses(detail: 4321));
ulong baseline = cache.CurrentVersion(7, 21);
bool found = cache.TryWaitForCompletion(
7,
21,
sinceVersion: baseline,
deadlineUtc: DateTime.UtcNow.AddMilliseconds(50),
pumpStep: static () => { },
out RepeatedField<MxStatusProxy> statuses);
Assert.False(found);
Assert.Empty(statuses);
}
/// <summary>
/// Verifies the pump loop is what lets a completion land: the completion is
/// recorded from inside a later pump step (standing in for the STA
/// dispatching the OnWriteComplete message) and the wait then succeeds.
/// </summary>
[Fact]
public void TryWaitForCompletion_InvokesPumpStepEachIteration()
{
MxAccessWriteCompletionCache cache = new();
int pumpCalls = 0;
bool found = cache.TryWaitForCompletion(
7,
21,
sinceVersion: 0UL,
deadlineUtc: DateTime.UtcNow.AddSeconds(5),
pumpStep: () =>
{
pumpCalls++;
if (pumpCalls == 2)
{
cache.Record(7, 21, BuildStatuses(detail: 55));
}
},
out RepeatedField<MxStatusProxy> statuses);
Assert.True(found);
Assert.True(pumpCalls >= 2);
Assert.Equal(55, Assert.Single(statuses).Detail);
}
/// <summary>Verifies that Record stores an independent clone of the caller's rows.</summary>
[Fact]
public void Record_ClonesStatuses()
{
MxAccessWriteCompletionCache cache = new();
RepeatedField<MxStatusProxy> callerRows = BuildStatuses(detail: 77);
cache.Record(7, 21, callerRows);
callerRows[0].Detail = 999;
callerRows.Add(new MxStatusProxy());
Assert.True(cache.TryWaitForCompletion(
7,
21,
sinceVersion: 0UL,
deadlineUtc: DateTime.UtcNow.AddSeconds(5),
pumpStep: static () => { },
out RepeatedField<MxStatusProxy> statuses));
Assert.Equal(77, Assert.Single(statuses).Detail);
}
private static RepeatedField<MxStatusProxy> BuildStatuses(int detail)
{
return new RepeatedField<MxStatusProxy>
{
new MxStatusProxy
{
Success = 1,
Category = MxStatusCategory.Ok,
Detail = detail,
},
};
}
}
@@ -0,0 +1,15 @@
namespace ZB.MOM.WW.MxGateway.Worker.MxAccess;
/// <summary>
/// Exposes the per-session <see cref="MxAccessWriteCompletionCache"/> an
/// event sink populates from OnWriteComplete callbacks, so
/// <see cref="MxAccessSession.Create"/> can share one instance between the
/// sink (writer) and the write command executor (reader). Implemented by
/// <see cref="MxAccessBaseEventSink"/> and by test sinks that cannot
/// attach to a live MXAccess COM object.
/// </summary>
public interface IWriteCompletionCacheProvider
{
/// <summary>The completion cache bound to this sink.</summary>
MxAccessWriteCompletionCache WriteCompletionCache { get; }
}
@@ -5,11 +5,12 @@ using Proto = ZB.MOM.WW.MxGateway.Contracts.Proto;
namespace ZB.MOM.WW.MxGateway.Worker.MxAccess;
/// <summary>Sink for MXAccess COM events that converts them to protobuf format.</summary>
public sealed class MxAccessBaseEventSink : IMxAccessEventSink
public sealed class MxAccessBaseEventSink : IMxAccessEventSink, IWriteCompletionCacheProvider
{
private readonly MxAccessEventMapper eventMapper;
private readonly MxAccessEventQueue eventQueue;
private readonly MxAccessValueCache valueCache;
private readonly MxAccessWriteCompletionCache writeCompletionCache;
private LMXProxyServerClass? server;
private string sessionId = string.Empty;
@@ -50,10 +51,32 @@ public sealed class MxAccessBaseEventSink : IMxAccessEventSink
MxAccessEventQueue eventQueue,
MxAccessEventMapper eventMapper,
MxAccessValueCache valueCache)
: this(eventQueue, eventMapper, valueCache, new MxAccessWriteCompletionCache())
{
}
/// <summary>
/// Initializes a new instance of the MxAccessBaseEventSink class with
/// provided queue, mapper, value cache, and a shared write-completion
/// cache. The completion cache is populated from every successful
/// <c>OnWriteComplete</c> dispatch so the worker's write executor can
/// correlate a WriteSecured/WriteSecured2 reply with the MXAccess
/// completion outcome.
/// </summary>
/// <param name="eventQueue">Queue for buffering converted MXAccess events.</param>
/// <param name="eventMapper">Converter for MXAccess events to protobuf format.</param>
/// <param name="valueCache">Per-session last-value cache shared with the MxAccessSession.</param>
/// <param name="writeCompletionCache">Per-session OnWriteComplete cache shared with the MxAccessSession.</param>
public MxAccessBaseEventSink(
MxAccessEventQueue eventQueue,
MxAccessEventMapper eventMapper,
MxAccessValueCache valueCache,
MxAccessWriteCompletionCache writeCompletionCache)
{
this.eventQueue = eventQueue ?? throw new ArgumentNullException(nameof(eventQueue));
this.eventMapper = eventMapper ?? throw new ArgumentNullException(nameof(eventMapper));
this.valueCache = valueCache ?? throw new ArgumentNullException(nameof(valueCache));
this.writeCompletionCache = writeCompletionCache ?? throw new ArgumentNullException(nameof(writeCompletionCache));
}
/// <summary>
@@ -62,6 +85,14 @@ public sealed class MxAccessBaseEventSink : IMxAccessEventSink
/// </summary>
public MxAccessValueCache ValueCache => valueCache;
/// <summary>
/// The OnWriteComplete completion cache populated by this sink. Exposed
/// via <see cref="IWriteCompletionCacheProvider"/> so the
/// MxAccessSession can share the same instance with the write command
/// executor's bounded completion wait.
/// </summary>
public MxAccessWriteCompletionCache WriteCompletionCache => writeCompletionCache;
/// <inheritdoc />
public void Attach(
object mxAccessComObject,
@@ -143,11 +174,18 @@ public sealed class MxAccessBaseEventSink : IMxAccessEventSink
ref MXSTATUS_PROXY[] pVars)
{
MXSTATUS_PROXY[] statuses = pVars;
EnqueueEvent(() => eventMapper.CreateOnWriteComplete(
sessionId,
hLMXServerHandle,
phItemHandle,
statuses));
// Record the completion for the write executor's bounded reply wait
// only after the event has cleared the queue (same post-publish rule
// as the OnDataChange value cache) — an overflow faults the session,
// so a dropped event never leaves a "fresher" completion behind than
// what shipped to the gateway.
EnqueueEvent(
() => eventMapper.CreateOnWriteComplete(
sessionId,
hLMXServerHandle,
phItemHandle,
statuses),
mxEvent => writeCompletionCache.Record(hLMXServerHandle, phItemHandle, mxEvent.Statuses));
}
/// <summary>
@@ -14,11 +14,22 @@ public sealed class MxAccessCommandExecutor : IStaCommandExecutor
/// <summary>Default per-tag timeout used when <c>ReadBulkCommand.timeout_ms</c> is zero.</summary>
internal static readonly TimeSpan DefaultReadBulkTimeout = TimeSpan.FromMilliseconds(1000);
/// <summary>
/// Default bounded wait for the OnWriteComplete callback after a
/// WriteSecured/WriteSecured2 COM call. 1.5 s keeps the unary reply
/// inside the OtOpcUa driver's 2 s Tier A write-resilience budget (a
/// longer gateway wait must raise that consumer timeout in step) while
/// covering the common fast-commit case; on expiry the reply returns
/// with empty statuses — unconfirmed, not failed.
/// </summary>
internal static readonly TimeSpan DefaultWriteCompletionTimeout = TimeSpan.FromMilliseconds(1500);
private readonly MxAccessSession session;
private readonly VariantConverter variantConverter;
private readonly MxStatusProxyConverter statusProxyConverter;
private readonly IAlarmCommandHandler? alarmCommandHandler;
private readonly Action pumpStep;
private readonly TimeSpan writeCompletionTimeout;
/// <summary>
/// Initializes a command executor with an MXAccess session.
@@ -71,17 +82,26 @@ public sealed class MxAccessCommandExecutor : IStaCommandExecutor
/// <param name="variantConverter">Converter for MXAccess variant values to MxValue protobuf messages.</param>
/// <param name="alarmCommandHandler">Optional handler for alarm-side commands.</param>
/// <param name="pumpStep">Action to pump Windows messages, or null for tests.</param>
/// <param name="writeCompletionTimeout">
/// Bounded wait for the OnWriteComplete callback after a
/// WriteSecured/WriteSecured2 COM call, or null for
/// <see cref="DefaultWriteCompletionTimeout"/>. Zero (or negative)
/// disables the wait entirely — replies keep the pure fire-and-forget
/// shape.
/// </param>
public MxAccessCommandExecutor(
MxAccessSession session,
VariantConverter variantConverter,
IAlarmCommandHandler? alarmCommandHandler,
Action? pumpStep)
Action? pumpStep,
TimeSpan? writeCompletionTimeout = null)
{
this.session = session ?? throw new ArgumentNullException(nameof(session));
this.variantConverter = variantConverter ?? throw new ArgumentNullException(nameof(variantConverter));
this.statusProxyConverter = new MxStatusProxyConverter();
this.alarmCommandHandler = alarmCommandHandler;
this.pumpStep = pumpStep ?? (static () => { });
this.writeCompletionTimeout = writeCompletionTimeout ?? DefaultWriteCompletionTimeout;
}
/// <inheritdoc />
@@ -457,6 +477,14 @@ public sealed class MxAccessCommandExecutor : IStaCommandExecutor
return CreateInvalidRequestReply(command, "WriteSecured command value is required.");
}
// Baseline BEFORE the COM call: a completion that dispatches during or
// immediately after WriteSecured bumps the version past this snapshot,
// so a fast commit still correlates (no missed-callback window).
MxAccessWriteCompletionCache completionCache = session.WriteCompletionCache;
ulong completionBaseline = completionCache.CurrentVersion(
writeSecuredCommand.ServerHandle,
writeSecuredCommand.ItemHandle);
session.WriteSecured(
writeSecuredCommand.ServerHandle,
writeSecuredCommand.ItemHandle,
@@ -464,7 +492,14 @@ public sealed class MxAccessCommandExecutor : IStaCommandExecutor
writeSecuredCommand.VerifierUserId,
variantConverter.ConvertToComValue(writeSecuredCommand.Value));
return CreateOkReply(command);
MxCommandReply reply = CreateOkReply(command);
AwaitWriteCompletion(
reply,
completionCache,
writeSecuredCommand.ServerHandle,
writeSecuredCommand.ItemHandle,
completionBaseline);
return reply;
}
private MxCommandReply ExecuteWriteSecured2(StaCommand command)
@@ -485,6 +520,12 @@ public sealed class MxAccessCommandExecutor : IStaCommandExecutor
return CreateInvalidRequestReply(command, "WriteSecured2 command timestamp value is required.");
}
// Same pre-call baseline rule as ExecuteWriteSecured.
MxAccessWriteCompletionCache completionCache = session.WriteCompletionCache;
ulong completionBaseline = completionCache.CurrentVersion(
writeSecured2Command.ServerHandle,
writeSecured2Command.ItemHandle);
session.WriteSecured2(
writeSecured2Command.ServerHandle,
writeSecured2Command.ItemHandle,
@@ -493,7 +534,14 @@ public sealed class MxAccessCommandExecutor : IStaCommandExecutor
variantConverter.ConvertToComValue(writeSecured2Command.Value),
variantConverter.ConvertToComValue(writeSecured2Command.TimestampValue));
return CreateOkReply(command);
MxCommandReply reply = CreateOkReply(command);
AwaitWriteCompletion(
reply,
completionCache,
writeSecured2Command.ServerHandle,
writeSecured2Command.ItemHandle,
completionBaseline);
return reply;
}
private MxCommandReply ExecuteAddItemBulk(StaCommand command)
@@ -897,6 +945,39 @@ public sealed class MxAccessCommandExecutor : IStaCommandExecutor
}
}
/// <summary>
/// Bounded pump-wait for the OnWriteComplete row matching a
/// WriteSecured/WriteSecured2 call, copied onto the reply when it
/// arrives in time. The executor holds the STA thread but pumps
/// Windows messages each poll (ReadBulk precedent) so the COM callback
/// can dispatch re-entrantly; on expiry the reply keeps its empty
/// statuses — the consumer's unconfirmed path, never a synthesized
/// failure. Protocol status/hresult stay acceptance-only either way.
/// </summary>
private void AwaitWriteCompletion(
MxCommandReply reply,
MxAccessWriteCompletionCache completionCache,
int serverHandle,
int itemHandle,
ulong completionBaseline)
{
if (writeCompletionTimeout <= TimeSpan.Zero)
{
return;
}
if (completionCache.TryWaitForCompletion(
serverHandle,
itemHandle,
completionBaseline,
DateTime.UtcNow + writeCompletionTimeout,
pumpStep,
out Google.Protobuf.Collections.RepeatedField<MxStatusProxy> statuses))
{
reply.Statuses.Add(statuses);
}
}
private static MxCommandReply CreateAlarmFailureReply(StaCommand command, Exception exception)
{
return new MxCommandReply
@@ -13,6 +13,7 @@ public sealed class MxAccessSession : IDisposable
private readonly IMxAccessEventSink eventSink;
private readonly MxAccessHandleRegistry handleRegistry;
private readonly MxAccessValueCache valueCache;
private readonly MxAccessWriteCompletionCache writeCompletionCache;
private bool disposed;
private MxAccessSession(
@@ -21,6 +22,7 @@ public sealed class MxAccessSession : IDisposable
IMxAccessEventSink eventSink,
MxAccessHandleRegistry handleRegistry,
MxAccessValueCache valueCache,
MxAccessWriteCompletionCache writeCompletionCache,
int creationThreadId)
{
this.mxAccessComObject = mxAccessComObject ?? throw new ArgumentNullException(nameof(mxAccessComObject));
@@ -28,6 +30,7 @@ public sealed class MxAccessSession : IDisposable
this.eventSink = eventSink ?? throw new ArgumentNullException(nameof(eventSink));
this.handleRegistry = handleRegistry ?? throw new ArgumentNullException(nameof(handleRegistry));
this.valueCache = valueCache ?? throw new ArgumentNullException(nameof(valueCache));
this.writeCompletionCache = writeCompletionCache ?? throw new ArgumentNullException(nameof(writeCompletionCache));
CreationThreadId = creationThreadId;
}
@@ -45,6 +48,14 @@ public sealed class MxAccessSession : IDisposable
/// </summary>
public MxAccessValueCache ValueCache => valueCache;
/// <summary>
/// Per-session OnWriteComplete completion cache populated by the event
/// sink. The write command executor consults it after a
/// WriteSecured/WriteSecured2 COM call so the unary reply can carry
/// the correlated completion outcome.
/// </summary>
public MxAccessWriteCompletionCache WriteCompletionCache => writeCompletionCache;
/// <summary>Creates a WorkerReady message with session metadata.</summary>
/// <param name="workerProcessId">Process ID of the worker.</param>
/// <returns>The populated <see cref="WorkerReady"/> message.</returns>
@@ -105,6 +116,9 @@ public sealed class MxAccessSession : IDisposable
eventSink,
handleRegistry ?? new MxAccessHandleRegistry(),
valueCache ?? new MxAccessValueCache(),
eventSink is IWriteCompletionCacheProvider provider
? provider.WriteCompletionCache
: new MxAccessWriteCompletionCache(),
creationThreadId ?? Environment.CurrentManagedThreadId);
}
@@ -149,12 +163,22 @@ public sealed class MxAccessSession : IDisposable
? baseSink.ValueCache
: new MxAccessValueCache();
// Share the sink's completion cache the same way (the production
// sink and completion-aware test sinks implement the provider
// seam); fall back to a fresh cache for other fakes — the write
// executor then simply never observes a completion and replies
// unconfirmed.
MxAccessWriteCompletionCache writeCompletionCache = eventSink is IWriteCompletionCacheProvider provider
? provider.WriteCompletionCache
: new MxAccessWriteCompletionCache();
return new MxAccessSession(
mxAccessComObject,
new MxAccessComServer(mxAccessComObject),
eventSink,
new MxAccessHandleRegistry(),
valueCache,
writeCompletionCache,
Environment.CurrentManagedThreadId);
}
catch (Exception exception)
@@ -11,6 +11,14 @@ namespace ZB.MOM.WW.MxGateway.Worker.MxAccess;
public sealed class MxAccessStaSession : IWorkerRuntimeSession
{
/// <summary>
/// Environment variable the gateway's WorkerProcessLauncher sets from
/// MxGateway:Worker:WriteCompletionWaitMilliseconds. 0 disables the
/// write-completion wait (pure fire-and-forget replies).
/// </summary>
internal const string WriteCompletionWaitEnvironmentVariableName =
"MXGATEWAY_WORKER_WRITE_COMPLETION_WAIT_MS";
private static readonly TimeSpan AlarmPollInterval = TimeSpan.FromMilliseconds(500);
private readonly IMxAccessComObjectFactory factory;
@@ -157,6 +165,32 @@ public sealed class MxAccessStaSession : IWorkerRuntimeSession
/// </summary>
public MxAccessEventQueue EventQueue => eventQueue;
/// <summary>
/// Bounded WriteSecured/WriteSecured2 completion wait handed to the
/// command executor at <see cref="StartAsync(string, int, CancellationToken)"/>.
/// Internal-settable as a test seam so Worker.Tests can shorten it
/// without env-var plumbing.
/// </summary>
internal TimeSpan WriteCompletionTimeout { get; set; } = ResolveWriteCompletionTimeout();
/// <summary>
/// Resolves the write-completion wait from the launcher-provided
/// environment variable; a missing or invalid value falls back to
/// <see cref="MxAccessCommandExecutor.DefaultWriteCompletionTimeout"/>.
/// </summary>
internal static TimeSpan ResolveWriteCompletionTimeout()
{
string? value = Environment.GetEnvironmentVariable(WriteCompletionWaitEnvironmentVariableName);
return int.TryParse(
value,
System.Globalization.NumberStyles.Integer,
System.Globalization.CultureInfo.InvariantCulture,
out int milliseconds)
&& milliseconds >= 0
? TimeSpan.FromMilliseconds(milliseconds)
: MxAccessCommandExecutor.DefaultWriteCompletionTimeout;
}
/// <summary>
/// Starts the MXAccess COM session asynchronously.
/// </summary>
@@ -208,12 +242,14 @@ public sealed class MxAccessStaSession : IWorkerRuntimeSession
session,
new VariantConverter(),
alarmCommandHandler,
// ReadBulk needs to pump Windows messages while it waits
// for the first OnDataChange callback so the inbound COM
// event can dispatch on this same STA thread. The pump
// step closes over staRuntime so it always pumps the
// pump tied to the apartment that owns this session.
pumpStep: () => staRuntime.PumpPendingMessages()));
// ReadBulk and the write-completion wait need to pump
// Windows messages while they wait for the inbound COM
// callback (OnDataChange / OnWriteComplete) so it can
// dispatch on this same STA thread. The pump step
// closes over staRuntime so it always pumps the pump
// tied to the apartment that owns this session.
pumpStep: () => staRuntime.PumpPendingMessages(),
writeCompletionTimeout: WriteCompletionTimeout));
return session.CreateWorkerReady(workerProcessId);
},
@@ -0,0 +1,152 @@
using System;
using System.Collections.Generic;
using System.Threading;
using Google.Protobuf.Collections;
using ZB.MOM.WW.MxGateway.Contracts.Proto;
namespace ZB.MOM.WW.MxGateway.Worker.MxAccess;
/// <summary>
/// Per-session cache of the most recent <c>OnWriteComplete</c> status rows
/// for each (server handle, item handle) pair. Written by the MXAccess
/// event sink as completion callbacks arrive; read by the write command
/// executor so a WriteSecured/WriteSecured2 reply can carry the correlated
/// MXAccess outcome instead of proving command acceptance only.
/// </summary>
/// <remarks>
/// Same threading posture as <see cref="MxAccessValueCache"/>: writers and
/// readers run on the worker's STA thread (COM dispatches events on the
/// apartment thread; commands also execute on the STA), so no internal
/// locking is required. A single sync root keeps it nominally thread-safe
/// for tests that drive it from a non-STA thread.
/// </remarks>
public sealed class MxAccessWriteCompletionCache
{
private readonly Dictionary<long, CompletionEntry> entries = new();
private readonly object syncRoot = new();
/// <summary>Records the status rows of a fresh OnWriteComplete callback for the given handle pair.</summary>
/// <param name="serverHandle">MXAccess server handle.</param>
/// <param name="itemHandle">MXAccess item handle.</param>
/// <param name="statuses">Status rows from the mapped OnWriteComplete event; cloned before storing.</param>
public void Record(
int serverHandle,
int itemHandle,
RepeatedField<MxStatusProxy> statuses)
{
if (statuses is null)
{
throw new ArgumentNullException(nameof(statuses));
}
lock (syncRoot)
{
long key = CreateItemKey(serverHandle, itemHandle);
ulong version = entries.TryGetValue(key, out CompletionEntry existing)
? existing.Version + 1
: 1UL;
entries[key] = new CompletionEntry(version, statuses.Clone());
}
}
/// <summary>Returns the current completion version for a handle pair, or 0 if none was recorded.</summary>
/// <param name="serverHandle">MXAccess server handle.</param>
/// <param name="itemHandle">MXAccess item handle.</param>
/// <returns>The current completion version, or 0 if no completion was recorded.</returns>
public ulong CurrentVersion(
int serverHandle,
int itemHandle)
{
lock (syncRoot)
{
return entries.TryGetValue(CreateItemKey(serverHandle, itemHandle), out CompletionEntry existing)
? existing.Version
: 0UL;
}
}
/// <summary>
/// Polls for a completion newer than <paramref name="sinceVersion"/> until it
/// arrives or the deadline elapses, calling <paramref name="pumpStep"/> on every
/// poll iteration so the worker's STA can dispatch the inbound MXAccess
/// OnWriteComplete message. Same loop shape as
/// <see cref="MxAccessValueCache.TryWaitForUpdate"/>.
/// </summary>
/// <param name="serverHandle">MXAccess server handle.</param>
/// <param name="itemHandle">MXAccess item handle.</param>
/// <param name="sinceVersion">Version snapshot captured before the write COM call.</param>
/// <param name="deadlineUtc">Absolute UTC deadline.</param>
/// <param name="pumpStep">Action that pumps any pending Windows messages.</param>
/// <param name="statuses">The recorded status rows if a completion arrived before the deadline; empty otherwise.</param>
/// <param name="pollIntervalMs">How long to sleep between pump cycles. Default 5 ms.</param>
/// <returns><see langword="true"/> if a completion newer than <paramref name="sinceVersion"/> arrived before the deadline; otherwise <see langword="false"/>.</returns>
public bool TryWaitForCompletion(
int serverHandle,
int itemHandle,
ulong sinceVersion,
DateTime deadlineUtc,
Action pumpStep,
out RepeatedField<MxStatusProxy> statuses,
int pollIntervalMs = 5)
{
if (pumpStep is null)
{
throw new ArgumentNullException(nameof(pumpStep));
}
while (true)
{
pumpStep();
lock (syncRoot)
{
if (entries.TryGetValue(CreateItemKey(serverHandle, itemHandle), out CompletionEntry entry)
&& entry.Version > sinceVersion)
{
statuses = entry.Statuses;
return true;
}
}
if (DateTime.UtcNow >= deadlineUtc)
{
statuses = new RepeatedField<MxStatusProxy>();
return false;
}
Thread.Sleep(pollIntervalMs);
}
}
private static long CreateItemKey(
int serverHandle,
int itemHandle)
{
return ((long)serverHandle << 32) | (uint)itemHandle;
}
/// <summary>
/// Snapshot of the most recent OnWriteComplete status rows for a handle
/// pair. <see cref="Version"/> increments by one on every
/// <see cref="Record"/> call so the write executor can detect "a new
/// completion arrived since I captured my baseline".
/// </summary>
/// <remarks>
/// Plain readonly struct (not a record) so this compiles under the
/// worker's net48 target, which lacks <c>IsExternalInit</c>.
/// </remarks>
private readonly struct CompletionEntry
{
public CompletionEntry(
ulong version,
RepeatedField<MxStatusProxy> statuses)
{
Version = version;
Statuses = statuses;
}
public ulong Version { get; }
public RepeatedField<MxStatusProxy> Statuses { get; }
}
}