docs(TST-27,WRK-26,CLI-42,CLI-43,IPC-28): P1 doc-drift batch, discharges IPC-29
ci / nightly-windev (push) Has been skipped
ci / windows-x86 (push) Successful in 1m17s
ci / java (push) Successful in 2m10s
ci / portable (push) Failing after 3m53s

TST-27: docs/GatewayConfiguration.md's ShowTagValues row no longer says
"Reserved" — it now states what false (default) does (DashboardEventBroadcaster
blanks tag values from a deep-cloned MxEvent before the SignalR events-hub
mirror), the security relevance (no per-session hub ACL yet, so this
redaction is the only thing between a low-trust Viewer and other sessions'
tag values), and the honest scope limit (does not cover /browse).

WRK-26 (discharges IPC-29): docs/MxAccessWorkerInstanceDesign.md's "Outbound
Queues" section rewritten from the stale five-level priority list to the
two-class Control/Event scheduler actually shipped, with the collapsed-
decision rationale, and the overflow paragraph rewritten to the implemented
fail-fast. docs/WorkerFrameProtocol.md gained a "Write Scheduling And
Sequencing" section describing HEAD truthfully: WRK-23's peek-stamp-commit
sequencing is live, WRK-25's event-batch flush coalescing is not (the drain
loop still awaits each event write individually), and WRK-22's cancellation
tombstone is not yet defined (noted as pending, not documented as shipped).

CLI-42: clients/rust/README.md and docs/ClientPackaging.md document the
vendored Rust proto layout matching build.rs — repo-path-first resolution
falling back to clients/rust/protos/, the check-codegen.ps1 Check 3 refresh
rule, and why cargo package/publish run without --no-verify.

CLI-43: docs/style-guides/JavaStyleGuide.md now says Java 17 (Ignition 8.3
baseline), mirroring CLI-12's wording, matching the shipped build.gradle.

IPC-28: docs/Grpc.md's exception-mapping prose gained CommandTooLarge ->
ResourceExhausted, and the Invoke section gained the oversized-payload
sentence, cross-referencing GatewayConfiguration.md's headroom rule.

Tracking: TST-27, WRK-26, CLI-42, CLI-43, IPC-28 flipped to Done and IPC-29
marked discharged-by-WRK-26 in 00-tracking.md and the 20/30/50/60 domain
registers, with a 2026-08-07 change-log entry.

Doc-only change; no source, proto, or test edits.
This commit is contained in:
Joseph Doherty
2026-08-07 07:26:58 -04:00
parent 97f79e79ef
commit 10534ec906
12 changed files with 148 additions and 37 deletions
+62
View File
@@ -65,6 +65,68 @@ Protocol violations throw `WorkerFrameProtocolException` with a
`WorkerFrameProtocolErrorCode` so callers can distinguish malformed frames,
oversized frames, protocol version mismatches, and session mismatches.
## Write Scheduling And Sequencing
This section covers write scheduling (priority classes, enqueue-then-contend,
flush coalescing) and sequencing (write-time stamping) together, because both
are properties of the same single write lock.
`WorkerFrameWriter` is a two-class cooperative priority scheduler
(`WorkerFrameWritePriority.Control` and `.Event`), not a strict per-kind
priority order. A caller enqueues its frame into the control or event queue
under a lock, then contends for a single write lock; whichever caller wins
drains every frame queued at that moment, control frames first and each class
in FIFO order, so a command reply, fault, heartbeat, or shutdown
acknowledgement is never delayed behind a backlog of queued events. Priority
only reorders *which frame writes next* — it does not affect the sequence
value a frame receives (see below), so a caller cannot infer priority class
from the wire sequence.
The envelope `Sequence` is stamped by the draining lock-holder at the actual
moment of writing, not when the frame is enqueued, so the on-wire order and
the stamped sequence always agree regardless of caller concurrency or
priority reordering. Stamping uses peek-stamp-commit: a candidate sequence is
assigned and the frame is validated (size, non-empty payload) against that
stamped value, but the counter is committed only immediately before the
stream write. A per-frame rejection therefore leaves the counter untouched —
the next accepted frame reuses the candidate number, so the wire sequence
stays contiguous across rejections and an operator reading a pipe capture
never sees a phantom gap from a rejected frame.
Two failure shapes are distinguished during a drain pass:
- **Per-frame rejection** (`InvalidEnvelope`, `MessageTooLarge`,
`ProtocolVersionMismatch`, `SessionMismatch`) is specific to the one frame
that failed validation or sizing. Nothing was written for it, so it fails
only that frame's completion and draining continues with the next queued
frame.
- **Stream failure** (anything else — a broken pipe, an I/O error) means the
underlying stream itself is no longer trustworthy. It fails the frame that
triggered it, every frame already written this batch but not yet flushed,
and every frame still queued, then stops draining entirely so no caller
waits forever on a stream that will not recover.
Flushes are coalesced across a drained batch: each frame in the batch is
written to the stream without an individual flush, then one `FlushAsync`
runs after the whole batch, and only then does every successfully-written
frame's completion resolve — so a caller's `WriteAsync` still does not
complete until its bytes are both written *and* flushed, but a batch that
happened to contain several queued frames pays one flush instead of one per
frame. In practice this coalescing currently engages only when multiple
frames are queued at the moment a lock-holder starts draining. The event
drain loop (`WorkerPipeSession.RunEventDrainLoopAsync`) awaits each drained
event's `WriteAsync` individually before writing the next, so today at most
one event frame is queued per drain pass and each event still costs its own
flush; a dedicated batch write entry point that submits a whole drained
event batch under one lock acquisition is designed but not yet landed, so a
burst of N events currently costs N flushes on the event hot path, not one.
Cancellation semantics for a `WriteAsync` call that is still waiting for the
write lock when its token fires are not yet defined at this layer — pending
a fix that will tombstone the queued frame so a cancelled call is guaranteed
never to reach the wire. Until that lands, a cancelled caller may still see
its frame written by whichever caller next holds the lock.
## Verification
The frame protocol lives in `ZB.MOM.WW.MxGateway.Worker.Ipc` (`WorkerFrameReader`,