docs(TST-27,WRK-26,CLI-42,CLI-43,IPC-28): P1 doc-drift batch, discharges IPC-29
ci / nightly-windev (push) Has been skipped
ci / windows-x86 (push) Successful in 1m17s
ci / java (push) Successful in 2m10s
ci / portable (push) Failing after 3m53s

TST-27: docs/GatewayConfiguration.md's ShowTagValues row no longer says
"Reserved" — it now states what false (default) does (DashboardEventBroadcaster
blanks tag values from a deep-cloned MxEvent before the SignalR events-hub
mirror), the security relevance (no per-session hub ACL yet, so this
redaction is the only thing between a low-trust Viewer and other sessions'
tag values), and the honest scope limit (does not cover /browse).

WRK-26 (discharges IPC-29): docs/MxAccessWorkerInstanceDesign.md's "Outbound
Queues" section rewritten from the stale five-level priority list to the
two-class Control/Event scheduler actually shipped, with the collapsed-
decision rationale, and the overflow paragraph rewritten to the implemented
fail-fast. docs/WorkerFrameProtocol.md gained a "Write Scheduling And
Sequencing" section describing HEAD truthfully: WRK-23's peek-stamp-commit
sequencing is live, WRK-25's event-batch flush coalescing is not (the drain
loop still awaits each event write individually), and WRK-22's cancellation
tombstone is not yet defined (noted as pending, not documented as shipped).

CLI-42: clients/rust/README.md and docs/ClientPackaging.md document the
vendored Rust proto layout matching build.rs — repo-path-first resolution
falling back to clients/rust/protos/, the check-codegen.ps1 Check 3 refresh
rule, and why cargo package/publish run without --no-verify.

CLI-43: docs/style-guides/JavaStyleGuide.md now says Java 17 (Ignition 8.3
baseline), mirroring CLI-12's wording, matching the shipped build.gradle.

IPC-28: docs/Grpc.md's exception-mapping prose gained CommandTooLarge ->
ResourceExhausted, and the Invoke section gained the oversized-payload
sentence, cross-referencing GatewayConfiguration.md's headroom rule.

Tracking: TST-27, WRK-26, CLI-42, CLI-43, IPC-28 flipped to Done and IPC-29
marked discharged-by-WRK-26 in 00-tracking.md and the 20/30/50/60 domain
registers, with a 2026-08-07 change-log entry.

Doc-only change; no source, proto, or test edits.
This commit is contained in:
Joseph Doherty
2026-08-07 07:26:58 -04:00
parent 97f79e79ef
commit 10534ec906
12 changed files with 148 additions and 37 deletions
File diff suppressed because one or more lines are too long
@@ -21,7 +21,7 @@ members, no positional records). The worker builds and tests only on the Windows
| WRK-23 | Low | — | S | WRK-21 (rejection path becomes backstop-only) | Done | Rejected frames consume sequence numbers, producing wire gaps |
| WRK-24 | Low | — | S | — | Not started | `AdoptNegotiatedMaxMessageBytes` has no lower-bound sanity check |
| WRK-25 | Low | P2 | S | WRK-22 (both touch enqueue/dequeue) | Not started | WRK-12 flush coalescing never engages on the event hot path |
| WRK-26 | Low | P1 | S | WRK-23 (soft — sequence prose); discharges IPC-29 | Not started | Write-priority and overflow doc drift from the WRK-07 change |
| WRK-26 | Low | P1 | S | WRK-23 (soft — sequence prose); discharges IPC-29 | Done | Write-priority and overflow doc drift from the WRK-07 change |
| WRK-27 | Low | — | S | — | Not started | Alarm poll bypasses the watchdog's in-flight suppression (15 s vs 75 s) |
| WRK-28 | Low | — | S | WRK-21 (land in the same commit cluster) | Done | 10,000 drain cap is a duplicated magic constant with a comment-only sync contract |
@@ -17,8 +17,8 @@ All `path:line` citations were re-verified against the working tree at `4f5371f`
| IPC-25 | Medium | P0 | M | — | Not started | Committed Go/Python worker bindings are stale at HEAD; no guard covers them |
| IPC-26 | Low | P2 | S¹ | WRK-22 | Not started | Cancelled write leaves a ghost frame that is still written (contract requirement here; fix mechanics in WRK-22) |
| IPC-27 | Low | P2 | S | — | Not started | Descriptor freshness test blind to enums, enum values, services/methods, and the Galaxy contract |
| IPC-28 | Low | — | S | — | Not started | `docs/Grpc.md` omits the `CommandTooLarge``ResourceExhausted` mapping |
| IPC-29 | Low | — | S | — | Not started | Worker writer priority scheduling and write-time sequence stamping undocumented in the frame-protocol doc |
| IPC-28 | Low | — | S | — | Done | `docs/Grpc.md` omits the `CommandTooLarge``ResourceExhausted` mapping |
| IPC-29 | Low | — | S | — | Done (discharged by WRK-26) | Worker writer priority scheduling and write-time sequence stamping undocumented in the frame-protocol doc |
| IPC-30 | Low | P0 | M | WRK-21 (same file/batch) | Done | Oversized worker→gateway event frame is session-fatal — make the death deliberate, structured, and diagnosable |
| IPC-31 | Info | — | — | — | N/A | Gateway stamps sequence at creation, worker at write — accepted divergence; sequence is documented diagnostic-only (`gateway.md:328-330`); revisit only if sequence ever becomes load-bearing |
| IPC-32 | Info | — | S | IPC-25 | Not started | `check-codegen.ps1` check labels miscounted (folded into the IPC-25 script edit) |
@@ -23,8 +23,8 @@ Operating constraints carried from prior work:
| CLI-39 | Medium | P1 | S | CLI-35..38, CLI-45 | Not started | Bump client versions off the already-published 0.1.2 before the next publish; add registry-collision guard |
| CLI-40 | Low | — | M | — | Done | Port the exact-secret credential scrub to Rust/Java/.NET |
| CLI-41 | Low | — | M | — | Done | Uniform malformed-reply contract for AuthenticateUser/ArchestrAUserToId/AddBufferedItem |
| CLI-42 | Low | P1 | S | — | Not started | Document the vendored Rust proto layout (CLI-02's missing doc half) |
| CLI-43 | Low | — | S | — | Not started | Java style guide still prescribes "Java 21 preferred" |
| CLI-42 | Low | P1 | S | — | Done | Document the vendored Rust proto layout (CLI-02's missing doc half) |
| CLI-43 | Low | — | S | — | Done | Java style guide still prescribes "Java 21 preferred" |
| CLI-44 | Low | — | S | — | Done | Go event goroutine can mislabel a genuine terminal error as `ErrSlowConsumer` |
| CLI-45 | Low | P1 | M | — | Done | Standardize CLI credential env-var name and fail fast on missing/empty passwords |
@@ -12,7 +12,7 @@ Prior-cycle open findings (TST-05..24 where still open) are tracked in the prior
|----|-----|------|-----|-----|--------|-------|
| TST-25 | High | P1 | M | — (unlocks TST-05, TST-24) | Done | Windows/x86 test tier has zero automation — restore via SSH-driven windev CI job |
| TST-26 | Medium | P1 (folded into TST-25) | S | TST-25 | Done | docs/GatewayTesting.md, check-codegen.ps1, and ci.yml comments describe removed CI jobs |
| TST-27 | Medium | P1 (doc batch) | S | — | Not started | `ShowTagValues` config row still says "Reserved" after SEC-25 made the flag live |
| TST-27 | Medium | P1 (doc batch) | S | — | Done | `ShowTagValues` config row still says "Reserved" after SEC-25 made the flag live |
| TST-28 | Low | P2 | S | relates IPC-02 | Done | Gateway-side `max_frame_bytes` handshake field untested in the CI-run suite |
| TST-29 | Low | P2 | S | — | Done | Retire `oldtasks.md` after folding the Phase-5 governance record into DesignDecisions.md; delete root docs-review artifacts |
| TST-30 | Low | P2 | M | — | Not started | Single shared Gitea runner is a CI throughput/availability bottleneck (cross-repo contention, no run cancel/delete) |
+16 -3
View File
@@ -18,9 +18,22 @@ clients/rust/
crates/mxgw-cli/
```
`build.rs` reads the `.proto` files from
`../../src/ZB.MOM.WW.MxGateway.Contracts/Protos` and generates `tonic`/`prost` bindings
into Cargo build output. `src/generated.rs` declares the Rust modules that
`build.rs` resolves the `.proto` inputs repo-path-first: it prefers the
canonical protos at `../../src/ZB.MOM.WW.MxGateway.Contracts/Protos` (two
levels above `clients/rust`) so a local in-repo `.proto` edit is picked up
live without any extra step, and falls back to the vendored copies checked
into `clients/rust/protos/` only when that canonical directory is absent —
the case for a consumer building the crate unpacked from a published
tarball, where the rest of the mxaccessgw repo does not exist. The vendored
copies are shipped in the published `.crate` via `Cargo.toml`'s `include`
list, which is what makes the crate buildable standalone; they are build
inputs only, never a second source of truth. **Refresh rule:** any commit
that edits a Contracts proto (`mxaccess_gateway.proto`, `mxaccess_worker.proto`,
`galaxy_repository.proto`) must copy the changed file(s) into
`clients/rust/protos/` in that same commit — `scripts/check-codegen.ps1`
Check 3 fails the build on byte drift between the vendored copies and the
canonical Contracts protos. `tonic`/`prost` bindings are generated into
Cargo build output. `src/generated.rs` declares the Rust modules that
include those generated files. `src/generated` remains reserved for checked-in
generator output if the crate later changes to source-tree generation.
+16
View File
@@ -117,6 +117,22 @@ The Rust workspace builds the `mxgateway-client` library crate and the `mxgw`
CLI crate. `build.rs` generates `tonic` and `prost` modules into Cargo build
output on each build that needs updated protobuf output.
`build.rs` resolves its `.proto` inputs repo-path-first, then vendored: it
prefers the canonical protos under `src/ZB.MOM.WW.MxGateway.Contracts/Protos`
so an in-repo edit is live immediately, and falls back to the copies vendored
into `clients/rust/protos/` only when the canonical directory is absent — the
case for a published crate unpacked outside this repo. The vendored copies
are declared in `Cargo.toml`'s `include` list, so `cargo package`/`cargo
publish` ship them inside the `.crate`, making the crate buildable standalone
with no access to the rest of the mxaccessgw repo. Any Contracts proto change
must refresh `clients/rust/protos/` in the same commit; `scripts/check-codegen.ps1`
Check 3 byte-compares the vendored copies against the canonical protos and
fails on drift. Because the vendored protos make a standalone build possible,
`cargo package`/`cargo publish` run **with** verification (no `--no-verify`) —
a `cargo package` that cannot build from the vendored tree alone would mean
the vendored copies are stale, and verification is what catches that before
publish.
Regenerate and compile Rust bindings:
```powershell
+1 -1
View File
@@ -186,7 +186,7 @@ events (a "gap") and must re-snapshot; whatever is still retained is replayed.
| `MxGateway:Dashboard:SnapshotIntervalMilliseconds` | `1000` | Dashboard snapshot refresh interval used by the snapshot SignalR hub and the pages that subscribe to it. |
| `MxGateway:Dashboard:RecentFaultLimit` | `100` | Maximum number of fault summaries projected into each dashboard snapshot. |
| `MxGateway:Dashboard:RecentSessionLimit` | `200` | Maximum number of session summaries projected into each dashboard snapshot. |
| `MxGateway:Dashboard:ShowTagValues` | `false` | Reserved display control for tag values. The dashboard does not show full tag values by default. |
| `MxGateway:Dashboard:ShowTagValues` | `false` | Controls whether tag values reach the dashboard's SignalR events hub mirror. `false` (default): `DashboardEventBroadcaster` blanks tag values from a deep-cloned copy of each `MxEvent` before it reaches any hub subscriber — event metadata (tag reference, quality, status, timestamps) still renders; see `docs/GatewayDashboardDesign.md`'s `EventsHub` row for the mechanism. Security-relevant because the per-session hub ACL that would scope a Viewer to specific sessions does not exist yet: with no per-session scoping, this redaction is currently the only thing standing between a low-trust Viewer and other sessions' tag values, so setting this `true` exposes every session's tag values to every authenticated dashboard viewer. The flag gates only the SignalR hub mirror — it does **not** cover the `/browse` live-value display, which remains a separate, still-open residual. |
| `MxGateway:Dashboard:GroupToRole` | _(empty)_ | LDAP group → dashboard role mapping. Keys are LDAP group names (short CN or full DN — leading-RDN match). Values must be `Admin` (read/write, API-key CRUD) or `Viewer` (read-only). A user whose LDAP groups don't intersect this map cannot sign in; with no mapping at all, only the loopback bypass admits anyone. |
| `MxGateway:Dashboard:DisableLogin` | `false` | Dev/test only. When `true`, replaces the cookie authentication handler with `DashboardAutoLoginAuthenticationHandler`, which auto-authenticates every dashboard request — including requests from remote browsers, not just loopback — as `AutoLoginUser` holding both `Administrator` and `Viewer` roles. No login form, LDAP bind, or cookie is involved. A loud one-time startup warning is logged. Differs from `AllowAnonymousLocalhost`: `DisableLogin` mints a real authenticated principal (so role-gated write affordances appear), whereas `AllowAnonymousLocalhost` satisfies the authorization requirement on loopback only without minting a principal (write affordances stay hidden). Never enable in production. **Production hard-stop (SEC-04):** when the host runs in the `Production` environment and `DisableLogin` is `true`, startup validation fails and the process aborts — the flag is only accepted outside Production, where the one-time startup warning still fires. |
| `MxGateway:Dashboard:AutoLoginUser` | `(null)` | Username stamped on the synthetic principal when `DisableLogin` is `true`. Default `(null)` — a null or blank value falls back to `multi-role`. Has no effect when `DisableLogin` is `false`. |
+3 -1
View File
@@ -82,6 +82,8 @@ return mapper.MapCommandReply(workerReply);
Carrying the enqueue timestamp into the worker layer is what lets queue-wait time be measured separately from worker-side execution time when troubleshooting timeouts.
An accepted gRPC command payload can still be too large for the worker pipe: the envelope built around it must fit `MxGateway:Worker:MaxMessageBytes`, which is validated at startup to sit at least a fixed envelope-overhead reserve above `MaxGrpcMessageBytes` (see the headroom rule in [Gateway Configuration](./GatewayConfiguration.md)) so this should not occur for a conformant payload, but if it does, `WorkerClient` raises `WorkerClientException(CommandTooLarge)` and `Invoke` reports `ResourceExhausted` for that command — the session is not faulted, so a client can retry with a smaller payload without reopening the session.
### `StreamEvents`
`StreamEvents` is a server-streaming RPC. The handler delegates the full pipeline to `IEventStreamService` and just forwards each `MxEvent` onto the response stream. Keeping the channel and producer/consumer machinery out of the handler means cancellation, exception mapping, and metric bookkeeping live in one place.
@@ -247,7 +249,7 @@ StatusCode statusCode = exception.ErrorCode switch
};
```
`WorkerClientException` follows the same pattern: `CommandTimeout` becomes `DeadlineExceeded`, `GatewayShutdown` becomes `Cancelled`, `InvalidState` becomes `FailedPrecondition`, `ProtocolViolation` becomes `Internal`, and unmapped codes fall through to `Unavailable`.
`WorkerClientException` follows the same pattern: `CommandTimeout` becomes `DeadlineExceeded`, `GatewayShutdown` becomes `Cancelled`, `InvalidState` becomes `FailedPrecondition`, `ProtocolViolation` becomes `Internal`, `CommandTooLarge` becomes `ResourceExhausted`, and unmapped codes fall through to `Unavailable`.
## Event Streaming Model
+33 -18
View File
@@ -638,27 +638,42 @@ Rules:
## Outbound Queues
The worker should use bounded outbound queues for replies, events, heartbeats,
and faults.
`WorkerFrameWriter` is a two-class cooperative priority scheduler, not the
five-level queue an earlier design draft called for. Every outbound frame is
enqueued at one of two `WorkerFrameWritePriority` values:
Priority order when writing:
- **Control** — hello, ready, command replies, faults, heartbeats, shutdown
acknowledgements.
- **Event** — MXAccess events drained from the queue.
1. faults,
2. command replies,
3. shutdown acknowledgements,
4. heartbeats,
5. events.
A caller enqueues its frame under a lock, then contends for a single write
lock; whichever caller wins drains **every** currently-queued frame before
releasing the lock, so a reply, fault, or heartbeat enqueued while a large
event backlog is draining still gets written on the very next drain pass
rather than waiting behind it. Within a class the queue is strict FIFO — a
fault does not jump ahead of an already-queued heartbeat or reply. This
collapses the originally-specified five-level order (faults > replies >
shutdown acks > heartbeats > events) into two classes; the decision is
deliberate, not a shortfall: the control queue is shallow in practice (faults
and shutdown acks are rare, replies and heartbeats are the steady traffic), so
the FIFO delay within it is bounded, and a two-class scheduler is simpler to
reason about and test than a five-level one for that same bound. See
`docs/WorkerFrameProtocol.md`'s "Write scheduling and sequencing" section for
the wire-level detail (sequence stamping, per-frame vs. stream-failure
semantics, flush coalescing).
Event overflow policy defaults to fail-fast for parity testing. If the event
queue fills:
1. Capture overflow metrics.
2. Send `WorkerFault` if possible.
3. Stop accepting new commands.
4. Let the gateway close or kill the worker.
Production coalescing may be added later, but it must be explicit and tested.
Do not drop or coalesce events in v1.
Event overflow policy is fail-fast, not "stop accepting new commands and let
the gateway close or kill the worker." When `MxAccessEventQueue.Enqueue` finds
the queue full, it throws and the queue self-records a `WorkerFault` with
category `QueueOverflow`. The event drain loop's next pass observes the fault
through `DrainFault()`, writes the fault frame — a Control-priority frame, so
it is not stuck behind a queued event backlog — and then throws to unwind
`RunAsync`: the worker process exits rather than continuing in a state where
events are being silently lost. The exit currently uses the generic
`WorkerExitCode.UnexpectedFailure` rather than a dedicated overflow code; a
distinct exit code remains open (tracked separately). Do not drop or coalesce
events to avoid this exit — that is explicitly out of scope for v1. Production
coalescing may be added later, but it must be explicit and tested.
The gateway side of the event path is bounded to match. `WorkerClient` buffers
inbound events in a bounded consumer channel plus a staging channel bounded at
+62
View File
@@ -65,6 +65,68 @@ Protocol violations throw `WorkerFrameProtocolException` with a
`WorkerFrameProtocolErrorCode` so callers can distinguish malformed frames,
oversized frames, protocol version mismatches, and session mismatches.
## Write Scheduling And Sequencing
This section covers write scheduling (priority classes, enqueue-then-contend,
flush coalescing) and sequencing (write-time stamping) together, because both
are properties of the same single write lock.
`WorkerFrameWriter` is a two-class cooperative priority scheduler
(`WorkerFrameWritePriority.Control` and `.Event`), not a strict per-kind
priority order. A caller enqueues its frame into the control or event queue
under a lock, then contends for a single write lock; whichever caller wins
drains every frame queued at that moment, control frames first and each class
in FIFO order, so a command reply, fault, heartbeat, or shutdown
acknowledgement is never delayed behind a backlog of queued events. Priority
only reorders *which frame writes next* — it does not affect the sequence
value a frame receives (see below), so a caller cannot infer priority class
from the wire sequence.
The envelope `Sequence` is stamped by the draining lock-holder at the actual
moment of writing, not when the frame is enqueued, so the on-wire order and
the stamped sequence always agree regardless of caller concurrency or
priority reordering. Stamping uses peek-stamp-commit: a candidate sequence is
assigned and the frame is validated (size, non-empty payload) against that
stamped value, but the counter is committed only immediately before the
stream write. A per-frame rejection therefore leaves the counter untouched —
the next accepted frame reuses the candidate number, so the wire sequence
stays contiguous across rejections and an operator reading a pipe capture
never sees a phantom gap from a rejected frame.
Two failure shapes are distinguished during a drain pass:
- **Per-frame rejection** (`InvalidEnvelope`, `MessageTooLarge`,
`ProtocolVersionMismatch`, `SessionMismatch`) is specific to the one frame
that failed validation or sizing. Nothing was written for it, so it fails
only that frame's completion and draining continues with the next queued
frame.
- **Stream failure** (anything else — a broken pipe, an I/O error) means the
underlying stream itself is no longer trustworthy. It fails the frame that
triggered it, every frame already written this batch but not yet flushed,
and every frame still queued, then stops draining entirely so no caller
waits forever on a stream that will not recover.
Flushes are coalesced across a drained batch: each frame in the batch is
written to the stream without an individual flush, then one `FlushAsync`
runs after the whole batch, and only then does every successfully-written
frame's completion resolve — so a caller's `WriteAsync` still does not
complete until its bytes are both written *and* flushed, but a batch that
happened to contain several queued frames pays one flush instead of one per
frame. In practice this coalescing currently engages only when multiple
frames are queued at the moment a lock-holder starts draining. The event
drain loop (`WorkerPipeSession.RunEventDrainLoopAsync`) awaits each drained
event's `WriteAsync` individually before writing the next, so today at most
one event frame is queued per drain pass and each event still costs its own
flush; a dedicated batch write entry point that submits a whole drained
event batch under one lock acquisition is designed but not yet landed, so a
burst of N events currently costs N flushes on the event hot path, not one.
Cancellation semantics for a `WriteAsync` call that is still waiting for the
write lock when its token fires are not yet defined at this layer — pending
a fix that will tombstone the queued frame so a cancelled call is guaranteed
never to reach the wire. Until that lands, a cancelled caller may still see
its frame written by whichever caller next holds the lock.
## Verification
The frame protocol lives in `ZB.MOM.WW.MxGateway.Worker.Ipc` (`WorkerFrameReader`,
+3 -1
View File
@@ -5,7 +5,9 @@ library, CLI, and tests.
## Baseline
- Target the Java version defined by the client build, with Java 21 preferred.
- Target Java 17 (the Ignition 8.3 baseline; the client build enforces
`options.release = 17` with a Gradle toolchain 17). Code must compile and
run on 17; newer JDKs may host the build.
- Use Gradle unless the repository standardizes on Maven.
- Apply a formatter such as Spotless or Google Java Format when configured.
- Keep generated protobuf code separate from handwritten wrappers.