Files
mxaccessgw/archreview/2026-07-12/remediation/20-worker.md
T
Joseph Doherty 33ba612ddd
ci / nightly-windev (push) Has been skipped
ci / java (push) Successful in 2m8s
ci / portable (push) Successful in 7m41s
ci / windows-x86 (push) Failing after 12m32s
fix(WRK-21,WRK-28,WRK-23,IPC-30): byte-budget the DrainEvents reply, stop size errors from killing sessions
WRK-21 — DrainEvents was bounded by event count only, so a byte-heavy queue
(large string/array MxValues) built a reply above the negotiated frame maximum:
the writer rejected the frame, the exception unwound the session, and the events
already dequeued were destroyed. The drain is now byte-budgeted inside the queue
lock, so an event is dequeued only once it is known to fit and one that does not
stays at the head. Truncation is reported through the reply's existing
DiagnosticMessage (no contract change); callers drain until an empty reply. Both
reply-write seams — the control-command path and ProcessCommandAsync — now catch
MessageTooLarge and answer the correlation with an InvalidRequest reply instead
of unwinding or faulting the session. Satisfies IPC-23 R1-R3.

WRK-28 — the 10,000 drain ceiling moves to GatewayContractInfo
.MaxDrainEventsPerCommand, referenced by both the gateway request validator and
the worker clamp, replacing a comment-only sync contract. C# const only; no
.proto change.

WRK-23 — WorkerFrameWriter now peek-stamps, validates, then commits the sequence
counter immediately before the stream write, so a per-frame rejection leaves no
phantom gap on the wire.

IPC-30 — an oversized event frame stays session-fatal (it is undeliverable end to
end and neither dropping nor synthesizing a replacement is allowed), but the
death is structured: the event's identity and sizes are logged (never its value),
a WorkerFault with category PROTOCOL_VIOLATION and command method EventDrain is
written, then the session exits as before.

Docs updated in the same change: MxAccessWorkerInstanceDesign.md (drain byte cap,
truncation contract, oversized-head behavior, oversized-event policy, no control
reply is session-fatal on size), WorkerFrameProtocol.md (reply pre-sizing,
non-fatal reply-size rule, oversized-event policy, rejected frames do not consume
sequence numbers), gateway.md (DrainEvents two-axis bound).
2026-08-07 05:38:23 -04:00

540 lines
39 KiB
Markdown

# Worker Process — Remediation Design & Implementation (2026-07-12 cycle)
Source review: [20-worker.md](../20-worker.md) · Roadmap: [00-overall.md](../00-overall.md) · Generated: 2026-07-13 · Baseline: `main` @ `4f5371f`
This cycle's findings are concentrated in the new frame-write scheduler and the DrainEvents bound
that the prior remediation added — the scheduler is fundamentally sound, and every fix below is an
edge-hardening of it rather than a redesign. One finding is P0 (WRK-21, shared with IPC-23: a
diagnostics command can still kill the session); the rest are Lows plus two residuals of prior Done
items (WRK-25, WRK-26). All fixes preserve MXAccess parity (no MXAccess behavior changes — these
are all IPC/liveness-layer changes), preserve the control-before-event write-priority guarantee,
and use net48-safe constructs only (plain constructors and get-only properties, no init-only
members, no positional records). The worker builds and tests only on the Windows x86 host
(windev); verification commands assume that host.
## Finding register
| ID | Sev | Tier | Eff | Dep | Status | Title |
|----|-----|------|-----|-----|--------|-------|
| WRK-21 | Medium | P0 | M | IPC-23 (same defect, fix owned here); WRK-28 (same lines) | Done | DrainEvents bound is count-based only; an oversized reply still kills the session and loses the drained events |
| WRK-22 | Low | — | S | IPC-26 (same defect, fix owned here) | Not started | Cancelled `WriteAsync` leaves its frame queued; it is still written later |
| WRK-23 | Low | — | S | WRK-21 (rejection path becomes backstop-only) | Done | Rejected frames consume sequence numbers, producing wire gaps |
| WRK-24 | Low | — | S | — | Not started | `AdoptNegotiatedMaxMessageBytes` has no lower-bound sanity check |
| WRK-25 | Low | P2 | S | WRK-22 (both touch enqueue/dequeue) | Not started | WRK-12 flush coalescing never engages on the event hot path |
| WRK-26 | Low | P1 | S | WRK-23 (soft — sequence prose); discharges IPC-29 | Not started | Write-priority and overflow doc drift from the WRK-07 change |
| WRK-27 | Low | — | S | — | Not started | Alarm poll bypasses the watchdog's in-flight suppression (15 s vs 75 s) |
| WRK-28 | Low | — | S | WRK-21 (land in the same commit cluster) | Done | 10,000 drain cap is a duplicated magic constant with a comment-only sync contract |
---
## WRK-21 — DrainEvents bound is count-based only; an oversized reply still kills the session and loses the drained events `Medium` · `P0`
**Finding.** `CreateDrainEventsReply` clamps only the event *count* to `MaxDrainEventsPerReply`
(10,000) (`src/ZB.MOM.WW.MxGateway.Worker/Ipc/WorkerPipeSession.cs:545-557`); the reply is written
via `WriteControlReplyAsync` (:481, :485-496) with no catch. A reply whose serialized size exceeds
the negotiated `MaxMessageBytes` throws `WorkerFrameProtocolException(MessageTooLarge)` from
`WorkerFrameWriter.WriteFrameAsync` (`Ipc/WorkerFrameWriter.cs:250-255`) — correctly a *per-frame*
rejection at the writer, but `HandleControlCommandAsync`/`DispatchGatewayEnvelopeAsync` propagate
it out of `RunMessageLoopAsync`, so the session dies and `WorkerApplication.cs:110-119` exits
`ProtocolViolation` (6). The 10,000 drained events were already removed from `MxAccessEventQueue`
by `runtimeSession.DrainEvents(maxEvents)` (:551) and are gone. 10,000 events at ~1.7 KiB each
(large string/array `MxValue`s) exceed the default 16 MiB + 64 KiB frame max, so the comment's
stated goal ("cannot pack the whole queue into one session-killing reply frame", :21-26) is not met
for byte-heavy events. **This is the same defect as IPC-23; this entry owns the single fix** (the
IPC remediation register should point here). Roadmap P0 item 2: byte-aware DrainEvents reply
capping such that **no diagnostics command can be session-fatal**.
**Impact.** A `DrainEvents max_events=0` diagnostic against a byte-heavy queue — the array-heavy
payloads this gateway exists for — terminates a healthy session (exit 6, gateway kills the worker)
*and* silently destroys the drained events. The exact "session-killing reply frame" the prior fix
(IPC-04/WRK slice) was added to prevent, narrowed but not closed.
**Design.** Two layers, both required by the P0 acceptance criterion:
1. **Primary — byte-aware drain that never removes an event it cannot ship.** The reply must be
sized *while draining*, and the size decision must happen inside the queue's lock so events are
only dequeued once they are known to fit. Add a byte-budgeted overload to `MxAccessEventQueue`:
drain from the head while `count < maxEvents` **and** the accumulated serialized size plus the
next event's size stays within a byte budget; an event that does not fit is left at the head for
the next call. The budget is `_options.MaxMessageBytes` (the negotiated value, read at
reply-build time) minus a fixed headroom constant covering the envelope/reply wrapper overhead
(64 KiB — the same envelope-overhead reserve rationale `docs/WorkerFrameProtocol.md` already
documents for the frame max itself). Per-event cost is `WorkerEvent.CalculateSize() + 8` — the
`WorkerEvent` wrapper slightly overestimates the packed `MxEvent`, and +8 conservatively covers
the repeated-field tag and length varint, so the estimate errs strictly on the safe side.
Truncation is reported in the reply's existing `DiagnosticMessage` field ("N events returned,
M remain; repeat DrainEvents for the rest"), so **no proto change** is needed.
Degenerate case: a single head event whose serialized size alone exceeds the budget. It is
**not** drained (draining it would either build an oversized reply or lose it); the reply
returns whatever fit before it (possibly zero events) with a `DiagnosticMessage` naming the
blocked event's `WorkerSequence` so the operator can find the offending tag. Such an event
cannot be shipped on the streaming path either — that policy question is IPC-30 and stays out
of scope here; the WRK-21 guarantee is only that DrainEvents never dies and never loses events.
2. **Backstop — the control-reply write seam must not be session-fatal.** Even with capping (a
future command, a sizing bug), catch `WorkerFrameProtocolException` with
`ErrorCode == MessageTooLarge` at the two reply-write seams and answer the correlation with a
small error reply instead of unwinding the session: `HandleControlCommandAsync`'s
`WriteControlReplyAsync` call (:481) and `ProcessCommandAsync`'s reply write (:630-638), which
today converts *any* write failure into `_state = Faulted` + `MxaccessCommandFailed` — i.e. an
oversized STA command reply also faults the whole session. The error reply uses
`ProtocolStatusCode.InvalidRequest` with an explanatory message ("reply exceeded the negotiated
frame maximum; retry with a smaller request").
Rejected alternatives: (a) adding `truncated`/`remaining_count` fields to `DrainEventsReply` — a
contracts change that regenerates and ripples through all five language clients for a diagnostic
command; `DiagnosticMessage` carries the same information at zero contract cost. (b) Catch-only
(no byte capping) — survives the session but still loses the already-drained events, failing the
"no event loss" half. (c) Draining into the reply and measuring `reply.CalculateSize()`
periodically — protobuf C# does not memoize sizes, making that O(n²), and events over the line
would already be dequeued. (d) A new `ProtocolStatusCode.ReplyTooLarge` enum value — proto change
for a should-be-unreachable backstop; rejected for the same ripple reason as (a).
Parity note: DrainEvents is a gateway-protocol diagnostic, not an MXAccess call — no MXAccess
behavior changes. Control-before-event priority is untouched (the reply stays a Control frame).
**Implementation.**
1. `src/ZB.MOM.WW.MxGateway.Worker/MxAccess/WorkerEventDrainResult.cs` (new): sealed class, plain
constructor + get-only properties (net48-safe): `IReadOnlyList<WorkerEvent> Events`,
`bool TruncatedBySize`, `int RemainingCount`, `ulong OversizedHeadSequence` (0 = none).
2. `MxAccess/MxAccessEventQueue.cs`: add
`public WorkerEventDrainResult Drain(uint maxEvents, int maxTotalBytes)` — under `syncRoot`,
loop `events.Peek()`, compute `workerEvent.CalculateSize() + 8`, dequeue while it fits the
remaining budget and `drained.Count < maxEvents` (0 = existing "all" semantics for the count
dimension); on a head event that alone exceeds `maxTotalBytes`, stop and record its
`WorkerSequence` into `OversizedHeadSequence`; `RemainingCount = events.Count` on exit.
3. `MxAccess/IWorkerRuntimeSession.cs`: add
`WorkerEventDrainResult DrainEvents(uint maxEvents, int maxTotalBytes);` (no default interface
method — the worker is net48, which has no DIM runtime support). Implement in
`MxAccess/MxAccessStaSession.cs` (delegate to the queue) and in every test fake that implements
the interface (the fake runtime session in `Worker.Tests/Ipc/WorkerPipeSessionTests.cs`).
4. `Ipc/WorkerPipeSession.cs`:
- add `private const int DrainReplyFrameHeadroomBytes = 64 * 1024;` next to
`MaxDrainEventsPerReply` (:26) with a comment tying it to the envelope-overhead reserve;
- `CreateDrainEventsReply` (:537-562): call the byte-budgeted overload with
`_options.MaxMessageBytes - DrainReplyFrameHeadroomBytes`; when `TruncatedBySize`, set
`reply.DiagnosticMessage` with returned/remaining counts and, when
`OversizedHeadSequence != 0`, the blocked sequence;
- `HandleControlCommandAsync` (:481): wrap `WriteControlReplyAsync(reply, ...)` in
`try/catch (WorkerFrameProtocolException ex) when (ex.ErrorCode == WorkerFrameProtocolErrorCode.MessageTooLarge)`
→ log `WorkerControlReplyTooLarge` (correlation id, kind, payload size from the exception
message) and write a fallback `MxCommandReply` (same correlation id/kind,
`ProtocolStatusCode.InvalidRequest`, explanatory message) via `WriteControlReplyAsync`;
- `ProcessCommandAsync` (:630-638): put the same catch around the reply write specifically
(inner try), writing the fallback error reply for the correlation instead of falling into the
generic catch that sets `_state = WorkerState.Faulted`.
5. Fold WRK-28 into this commit cluster (same lines — see below).
6. Docs same commit: `docs/MxAccessWorkerInstanceDesign.md` — the control-command/DrainEvents
prose gains the byte cap, the truncation `DiagnosticMessage` contract, and the
oversized-head-event behavior; note in the Outbound Queues section that no control reply is
session-fatal on size. `docs/WorkerFrameProtocol.md` — one sentence in the size-limits section:
the session pre-sizes DrainEvents replies below the negotiated max, and `MessageTooLarge` on a
reply write is answered with an error reply, not session teardown.
**Verification.** On windev:
`dotnet build src/ZB.MOM.WW.MxGateway.Worker/ZB.MOM.WW.MxGateway.Worker.csproj -p:Platform=x86`
then
`dotnet test src/ZB.MOM.WW.MxGateway.Worker.Tests/ZB.MOM.WW.MxGateway.Worker.Tests.csproj -p:Platform=x86 --filter "FullyQualifiedName~WorkerPipeSessionTests|FullyQualifiedName~MxAccessEventQueueTests"`.
New tests:
- `WorkerPipeSessionTests.DrainEvents_ByteHeavyQueue_ReplyIsBoundedAndSessionSurvives` — the
review's named repro: enqueue 10,000 events carrying large string values (≥ ~1.7 KiB each) into
a real `MxAccessEventQueue` behind the fake runtime, send `DrainEvents max_events=0`; assert the
reply envelope's serialized size ≤ `MaxMessageBytes`, `DiagnosticMessage` reports truncation,
and the session still answers a subsequent `Ping` (no fault frame, `RunAsync` still running).
- `WorkerPipeSessionTests.DrainEvents_RepeatedCalls_RecoverAllEventsWithoutLoss` — issue
DrainEvents until empty; assert the union of replies is exactly the 10,000 enqueued events in
order (split/bounded replies, zero loss).
- `WorkerPipeSessionTests.ControlReplyTooLarge_WritesErrorReplyInsteadOfDying` — shrink the
negotiated max via `WorkerFrameProtocolOptions` so the reply write throws `MessageTooLarge`;
assert a fallback `InvalidRequest` reply with the same correlation id is written and the session
survives.
- `WorkerPipeSessionTests.CommandReplyTooLarge_WritesErrorReplyInsteadOfFaulting` — same via the
STA command path (`ProcessCommandAsync`); assert no `WorkerFault`, `_state` stays `Ready`.
- `MxAccessEventQueueTests.Drain_ByteBudget_StopsBeforeBudgetAndLeavesRemainderQueued` — order
preserved, `RemainingCount` exact, undrained events still present.
- `MxAccessEventQueueTests.Drain_ByteBudget_OversizedHead_DrainsNothingAndReportsHeadSequence`.
---
## WRK-22 — Cancelled `WriteAsync` leaves its frame queued; it is still written later `Low` · `—`
**Finding.** `WorkerFrameWriter.WriteAsync` enqueues the `PendingFrame` before the cancellable
lock wait (`Ipc/WorkerFrameWriter.cs:87-103`); if `_writeLock.WaitAsync(cancellationToken)`
throws, nothing removes the frame, so the next lock-holder writes it. Concrete scenario: at
shutdown the event-drain loop's token is cancelled (`Ipc/WorkerPipeSession.cs:298-304`) while its
event frame waits for the lock; the shutdown-ack write then drains queues — ack first (Control),
then the orphaned event — so an event frame is emitted *after* `WorkerShutdownAck`, and a write
reported as cancelled to its caller still reaches the wire. **Same defect as IPC-26** (the IPC
review describes this writer); this entry owns the single fix.
**Impact.** Benign today (the gateway drops post-shutdown/unknown frames), but the "cancelled
means not written" contract is silently violated, an unobserved completion lingers, and any future
gateway-side strictness about frames after the ack turns this into a protocol violation.
**Design.** Tombstone on cancellation, claim on dequeue. `Queue<T>` has no random removal, so the
standard fix is: the cancelled caller marks its frame cancelled under `_gate`; `DequeueNext` skips
completed (cancelled) frames and atomically marks the frame it returns as *claimed*. A frame
already claimed by a draining lock-holder is mid-write and can no longer be recalled — the
documented contract becomes "cancellation guarantees the frame is not written *if it has not yet
been claimed*; a claimed frame may still reach the wire." That residual window is unavoidable
without blocking the canceller behind the very write it is abandoning (rejected: it inverts the
point of cancellation). Rejected alternative: switching to a `LinkedList` for O(1) removal —
tombstoning is simpler, allocation-free, and the queues are short-lived.
**Implementation.**
1. `Ipc/WorkerFrameWriter.cs`, nested `PendingFrame`: add `public bool Claimed;` (plain field,
mutated only under `_gate`).
2. `WriteAsync(envelope, priority, ct)` (:100-113): wrap `await _writeLock.WaitAsync(cancellationToken)`
in `try/catch (OperationCanceledException)``lock (_gate) { if (!frame.Claimed) { frame.Completion.TrySetCanceled(cancellationToken); } }`
then rethrow. (If already claimed, the frame will be written; the caller still observes the
cancellation — documented in the method's XML doc.)
3. `DequeueNext` (:200-216): under `_gate`, loop-dequeue past frames whose
`Completion.Task.IsCanceled`; set `Claimed = true` on the frame actually returned.
`FailAllQueued`/`FailFrames` already use `TrySetException`, which no-ops on cancelled frames.
4. Docs same commit: the "Write scheduling and sequencing" section of
`docs/WorkerFrameProtocol.md` (created by WRK-26 — if WRK-26 has not landed yet, carry the
cancellation paragraph in this commit and let WRK-26 absorb it) states the
cancelled-means-not-written-unless-claimed contract.
**Verification.** On windev:
`dotnet test src/ZB.MOM.WW.MxGateway.Worker.Tests/ZB.MOM.WW.MxGateway.Worker.Tests.csproj -p:Platform=x86 --filter "FullyQualifiedName~WorkerFrameProtocolTests"`.
New tests:
- `WorkerFrameProtocolTests.WriteAsync_CancelledWhileWaitingForLock_FrameIsNeverWritten` — block
the stream (gated fake stream) so writer A holds `_writeLock` mid-write; enqueue an event write
with a CTS; cancel; assert `OperationCanceledException`; unblock A; write a control frame;
assert the captured wire bytes contain A's frame and the control frame only — the cancelled
event envelope never appears.
- `WorkerFrameProtocolTests.WriteAsync_CancelledEventFrame_DoesNotTrailShutdownAck` — the review's
shutdown scenario: cancelled event frame queued, then a Control (ack-shaped) write drains;
assert the event frame is skipped and the ack is the last frame on the wire.
---
## WRK-23 — Rejected frames consume sequence numbers, producing wire gaps `Low` · `—`
**Finding.** `WorkerFrameWriter.WriteFrameAsync` stamps `envelope.Sequence = unchecked(++_nextSequence)`
(`Ipc/WorkerFrameWriter.cs:240`) *before* the empty-payload and `MessageTooLarge` checks
(:243-255), so a per-frame rejection leaves a hole in the on-wire sequence. `gateway.md:328` calls
`sequence` a monotonic *diagnostic* counter, so nothing breaks — but the existing test
`WorkerFrameProtocolTests.WriteAsync_UnderConcurrentCalls_StampsGapFreeMonotonicSequence`
documents a stronger gap-free expectation than the code guarantees once a rejection occurs.
**Impact.** Diagnostics only: an operator correlating wire captures sees phantom gaps after any
rejection and may chase a lost-frame bug that does not exist; the test asserts a guarantee the
code does not hold.
**Design.** Peek-stamp, validate, commit. The sequence value participates in `CalculateSize()`
(varint width), so it cannot simply be stamped after the size checks — the checked size would not
match the written payload. Instead compute a *candidate* (`_nextSequence + 1`), stamp it, run the
size/empty checks against the stamped envelope, and only commit `_nextSequence = candidate`
immediately before the stream write. On a per-frame rejection the counter is untouched, so the
next accepted frame reuses the number and the wire stays contiguous. Safe because `_nextSequence`
is only ever touched by the lock-holder (`:44-48`). The rejected envelope retains its tentative
sequence — a caller-visible mutation on a failed envelope that already exists today. Rejected
alternative: decrement on rejection (`--_nextSequence`) — equivalent effect but leaves a window
where the counter was observably advanced; the candidate/commit form has no intermediate state.
Dependency note: after WRK-21 lands, `MessageTooLarge` at the writer becomes a rare backstop, so
this fix is what makes the existing gap-free test's guarantee genuinely true rather than
usually-true.
**Implementation.**
1. `Ipc/WorkerFrameWriter.cs`, `WriteFrameAsync` (:234-267): keep
`WorkerEnvelopeValidator.Validate` first; replace the stamp with
`ulong candidate = unchecked(_nextSequence + 1); envelope.Sequence = candidate;`; after the
empty-payload and `MaxMessageBytes` checks pass, `_nextSequence = candidate;` immediately
before the buffer build/`WriteTo`. Update the :45-48 comment.
2. Docs same commit: none required beyond WRK-26's scheduling section, which should state
"rejected frames do not consume sequence numbers" once this lands (see WRK-26 ordering note).
**Verification.** On windev:
`dotnet test src/ZB.MOM.WW.MxGateway.Worker.Tests/ZB.MOM.WW.MxGateway.Worker.Tests.csproj -p:Platform=x86 --filter "FullyQualifiedName~WorkerFrameProtocolTests"`.
New test:
- `WorkerFrameProtocolTests.WriteAsync_PerFrameRejection_DoesNotConsumeSequence` — write an
accepted frame, attempt an oversized frame (assert `MessageTooLarge` rejection), write another
accepted frame; assert the two on-wire frames carry sequences 1 and 2.
The existing `WriteAsync_UnderConcurrentCalls_StampsGapFreeMonotonicSequence` must still pass.
---
## WRK-24 — `AdoptNegotiatedMaxMessageBytes` has no lower-bound sanity check `Low` · `—`
**Finding.** `Ipc/WorkerFrameProtocolOptions.cs:124-140` — 0 keeps the default and >256 MiB
throws, but any value from 1 byte up is adopted. A gateway bug or a foreign/old gateway sending
e.g. 512 leaves the session alive but unable to write any frame (every write fails
`MessageTooLarge`) and unable to read most inbound frames — failing confusingly mid-session rather
than at the handshake, contradicting the class's own declared intent (:14-18, "a nonsensical
negotiated value is rejected at the handshake"). The gateway's own validator floors its config at
1024 (`src/ZB.MOM.WW.MxGateway.Server/Configuration/GatewayOptionsValidator.cs:10,230-233`).
**Impact.** A misconfigured or buggy peer produces a session that handshakes cleanly and then
fails every subsequent operation with per-frame size errors — the worst diagnostic shape for an
operator, and exactly what the top-end check was added to prevent.
**Design.** Mirror the gateway's validation floor: reject negotiated values below 1024 with the
same `InvalidConfiguration` error the ceiling uses, so the failure is a handshake-time fault frame
(`ValidateGatewayHello``CompleteStartupHandshakeAsync`'s `WorkerFrameProtocolException` catch →
`TryWriteFaultAsync` + throw) instead of a mid-session mystery. Floor value: **1024**, matching
`GatewayOptionsValidator.MinimumMaxMessageBytes` — the worker must never reject a value the
gateway's own validator accepts as legal configuration. Rejected alternative: a 64 KiB floor
(review's other suggestion) — safer-sounding but would turn a gateway legitimately configured
anywhere in [1024, 65535] into a handshake failure; symmetry of the two processes' validation
bounds is the defensible contract, and 1024 already guarantees hellos, heartbeats, acks, and
faults fit.
**Implementation.**
1. `Ipc/WorkerFrameProtocolOptions.cs`: add
`public const int MinNegotiableFrameBytes = 1024;` (XML doc: matches the gateway's
`MinimumMaxMessageBytes` validation floor); in `AdoptNegotiatedMaxMessageBytes`, after the
`== 0` early return, throw `WorkerFrameProtocolException(InvalidConfiguration)` when
`negotiatedMaxFrameBytes < MinNegotiableFrameBytes`, message symmetrical with the ceiling text.
2. Docs same commit: `docs/WorkerFrameProtocol.md` negotiated-max paragraph (:25-30) — state the
accepted range [1024, 256 MiB], 0 = keep default, out-of-range rejected at the handshake.
**Verification.** On windev:
`dotnet test src/ZB.MOM.WW.MxGateway.Worker.Tests/ZB.MOM.WW.MxGateway.Worker.Tests.csproj -p:Platform=x86 --filter "FullyQualifiedName~WorkerFrameProtocolTests|FullyQualifiedName~WorkerPipeSessionTests"`.
New tests:
- `WorkerFrameProtocolTests.AdoptNegotiatedMaxMessageBytes_BelowFloor_ThrowsInvalidConfiguration`
(512 → throws; 1024 → adopted; boundary pair).
- `WorkerPipeSessionTests.Handshake_GatewayHelloWithTinyMaxFrameBytes_FaultsAtHandshake` — mirror
of the existing >256 MiB handshake test: `GatewayHello.MaxFrameBytes = 512` → fault frame
written, handshake throws, no message loop entered.
---
## WRK-25 — WRK-12 flush coalescing never engages on the event hot path `Low` · `P2`
**Finding.** The event drain loop awaits each event's `WriteAsync` — which completes only after
that frame is written *and flushed* — before writing the next
(`Ipc/WorkerPipeSession.cs:374-382`), so the writer's event queue holds at most one frame from the
drain loop and every event still costs one flush. The batch flush path
(`Ipc/WorkerFrameWriter.cs:120-124, 160-182`) only coalesces when independent producers happen to
queue behind a blocked in-progress write. Residual of prior Done item WRK-12: the tracking-log
claim "a burst of N events costs 1 flush not N" does not describe production behavior.
**Impact.** Performance only: under event bursts (the gateway's core load) each event pays a
semaphore round-trip and a pipe flush; the shipped coalescing mechanism sits idle on the exact
path it was built for.
**Design.** Give the drain loop a real batch entry point: `WriteBatchAsync` on the writer —
enqueue every frame of the drained batch under one `_gate` acquisition (preserving intra-batch
order), take the write lock once, drain (existing logic: control-first, one flush per batch),
then await all completions. The existing single-drainer machinery then does exactly what WRK-12
intended: N events, one flush. Guarantees preserved: control-before-event holds because the batch
still lands in `_eventFrames` and any concurrently queued control frame is drained first by
`DequeueNext`; intra-batch event order holds because the enqueue is atomic under `_gate` and the
event queue is FIFO; the "written and flushed before completion" contract is unchanged. A
per-frame rejection inside the batch (an oversized single event) surfaces from the awaited
completions and terminates the session exactly as today — that policy is IPC-30's, deliberately
not changed here. Rejected alternative (review's other option): fire N `WriteAsync` calls and
`Task.WhenAll` — works (each call enqueues synchronously before its first await) but costs N
semaphore acquisitions and N no-op drains for zero benefit over the explicit batch. Depends on
WRK-22 landing first or together: both touch the enqueue/`DequeueNext` seam, and batch cancellation
reuses WRK-22's tombstoning.
**Implementation.**
1. `Ipc/WorkerFrameWriter.cs`: add
`public async Task WriteBatchAsync(IReadOnlyList<WorkerEnvelope> envelopes, WorkerFrameWritePriority priority, CancellationToken cancellationToken = default)`
— build a `PendingFrame[]`, enqueue all under one `lock (_gate)`, single
`await _writeLock.WaitAsync(cancellationToken)` (on `OperationCanceledException`: tombstone all
unclaimed frames per WRK-22, rethrow), `DrainQueuedFramesAsync()`, release, then await each
completion in order (plain loop over the array — no LINQ needed, net48-safe either way).
2. `Ipc/WorkerPipeSession.cs`, `RunEventDrainLoopAsync` (:374-382): replace the per-event awaited
loop with building the envelope list for the drained batch and one
`WriteBatchAsync(envelopes, WorkerFrameWritePriority.Event, cancellationToken)`. Keep the
priority comment, updated.
3. Docs same commit: `docs/WorkerFrameProtocol.md` scheduling section (WRK-26, which lands earlier
as P1) — amend its flush-coalescing sentence to say the event drain loop submits drained
batches through the batch entry point, so a burst of N events costs one flush.
**Verification.** On windev:
`dotnet test src/ZB.MOM.WW.MxGateway.Worker.Tests/ZB.MOM.WW.MxGateway.Worker.Tests.csproj -p:Platform=x86 --filter "FullyQualifiedName~WorkerFrameProtocolTests|FullyQualifiedName~WorkerPipeSessionTests"`.
New tests:
- `WorkerFrameProtocolTests.WriteBatchAsync_FlushesOnceAndPreservesOrder` — flush-counting fake
stream; N event envelopes → exactly one flush, wire order = batch order, sequences monotonic.
- `WorkerFrameProtocolTests.WriteBatchAsync_ControlFrameQueuedDuringBatch_JumpsRemainingEvents`
control-before-event still holds mid-batch.
- `WorkerPipeSessionTests.EventBurst_DrainLoopCoalescesFlushes` — fake runtime yields a 128-event
batch; assert the capturing stream saw one flush for the batch (this is the assertion the WRK-12
tracking claim needs to actually be true).
---
## WRK-26 — Write-priority and overflow doc drift from the WRK-07 change `Low` · `P1`
**Finding.** `docs/MxAccessWorkerInstanceDesign.md:607-613` still specifies the five-level write
priority (faults > replies > shutdown acks > heartbeats > events) while the implementation is a
two-class scheduler (Control FIFO, then Event — `Ipc/WorkerFrameWritePriority.cs:10-17`,
`Ipc/WorkerFrameWriter.cs:200-216`); within Control a fault does not jump queued
heartbeats/replies. `docs/WorkerFrameProtocol.md` describes the negotiated max (:25-30) but says
nothing about priority classes, flush coalescing, write-time sequence stamping, or the per-frame
vs stream-failure semantics. The adjacent overflow-policy text
(`docs/MxAccessWorkerInstanceDesign.md:615-624`, "stop accepting new commands") is also stale
against the implemented fail-fast (WRK-18's doc half). CLAUDE.md requires affected docs to change
in the same commit as the source; this is the residual of Done item WRK-07. Roadmap P1 doc-drift
batch (item 8); **discharges IPC-29**, which asks for the same `WorkerFrameProtocol.md` section.
**Impact.** The component design doc actively misleads: a maintainer adding a new frame kind will
look for a five-level queue that does not exist, and the accepted FIFO-within-control decision is
recorded nowhere, so it reads as a defect rather than a choice.
**Design.** Doc-only; make the docs describe the implemented scheduler and record the decision.
No code change — the two-class scheduler is correct and accepted (the review's deep-dive found it
deadlock-free and lost-wakeup-free; within-Control FIFO is bounded by the small control queue
depth). Ordering note: this is P1 and will likely land before WRK-23/WRK-25 (P2/untierred), so the
new text must describe HEAD truthfully at commit time — sequence numbers *are* consumed by rejected
frames until WRK-23 lands, and event flushes are *not* yet coalesced on the drain path until
WRK-25 lands; write those two facts as they are, and WRK-23/WRK-25 update the sentences in their
own commits (same-commit docs rule applies to them too).
**Implementation.**
1. `docs/MxAccessWorkerInstanceDesign.md:602-624`: rewrite the "Outbound Queues" section —
two-class cooperative scheduler (`WorkerFrameWriter` + `WorkerFrameWritePriority`): Control
(faults, command replies, shutdown acks, heartbeats, hello/ready) FIFO among themselves and
always ahead of Event frames; enqueue-then-contend, lock-holder drains all queued frames;
record the accepted decision that the original five-level order was collapsed and why (control
queue is shallow, so intra-control FIFO delay is bounded and the machinery stays simple).
Rewrite the overflow paragraph to the implemented fail-fast (fault written, drain loop throws,
process exits — currently the generic exit code; WRK-18 still tracks the dedicated code).
2. `docs/WorkerFrameProtocol.md`: add a "Write scheduling and sequencing" section — priority
classes; `Sequence` stamped at the moment of writing under the write lock so wire order and
stamped sequence always agree; per-frame rejections (`InvalidEnvelope`, `MessageTooLarge`,
version/session mismatch) fail only that frame while stream failures fail the batch and all
queued frames; flush coalescing across a drained batch and the "written and flushed" completion
contract; the rejection/sequence-gap behavior as it stands at commit time (see ordering note);
cancellation semantics once WRK-22 lands (or carried by WRK-22's commit if it lands later).
3. Cross-check `gateway.md:328-330` (sequence prose) still reads true; touch only if wrong.
**Verification.** Doc-only — no build impact. Sanity on any host:
`grep -n "faults" docs/MxAccessWorkerInstanceDesign.md` shows no remaining five-level list;
`grep -n "scheduling" docs/WorkerFrameProtocol.md` finds the new section. Follow
`docs/style-guides/StyleGuide.md` (present tense, explain why). Confirm the IPC remediation
register marks IPC-29 as discharged by this item.
---
## WRK-27 — Alarm poll bypasses the watchdog's in-flight suppression (15 s vs 75 s) `Low` · `—`
**Finding.** `MxAccess/MxAccessStaSession.cs:245-253` invokes `handler.PollOnce()` via
`staRuntime.InvokeAsync` directly, bypassing `StaCommandDispatcher`; `CaptureHeartbeat` takes
`CurrentCommandCorrelationId` only from the dispatcher (:364-381), so during a long `PollOnce` the
heartbeat shows no in-flight command and the watchdog fires `StaHung` as soon as staleness exceeds
`HeartbeatGrace` (15 s default) rather than the 75 s `HeartbeatStuckCeiling` granted to dispatched
commands (`Ipc/WorkerPipeSession.cs:843-861`).
**Impact.** A healthy-but-slow `GetXmlCurrentAlarms2` (large alarm set, busy provider) blocking
the STA for >15 s faults a healthy session. Partially defensible — a blocked STA cannot deliver
data events either — but the 15 s vs 75 s asymmetry between polls and commands is unintentional.
**Design.** Surface poll-in-progress to the watchdog through the heartbeat snapshot as a boolean,
and have the suppression branch honor it alongside `CurrentCommandCorrelationId`. The poll then
gets exactly the command treatment: suppressed up to the ceiling, faulted past it (a poll that
blocks the STA >75 s without pumping *should* fault — that is the ceiling's contract). All types
touched are worker-internal; the heartbeat proto is unchanged, so no contract/client work.
Rejected alternatives: (a) routing the poll through `StaCommandDispatcher` — inflates
`PendingCommandCount`, subjects polls to the 128-entry command backpressure and
shutdown-rejection semantics, and perturbs dispatch ordering for real gateway commands; the poll
is infrastructure, not a command. (b) A synthetic correlation id (e.g. `"alarm-poll"`) in the
heartbeat — visible to the gateway, which treats correlation ids as real command identifiers
(dashboards, per-command timeout logic); misrepresenting a nonexistent command crosses an
observability contract for a purely local suppression need.
**Implementation.**
1. `MxAccess/MxAccessStaSession.cs`: add `private volatile bool staAlarmPollInProgress;`; in the
delegate passed to `staRuntime.InvokeAsync` in `RunAlarmPollLoopAsync` (:247-253), set it
`true` immediately before `EnsureOnAlarmConsumerThread()`/`handler.PollOnce()` and clear it in
a `finally` — the flag flips on the STA thread exactly around the COM call.
2. `MxAccess/WorkerRuntimeHeartbeatSnapshot.cs`: add a `bool StaCallInProgress` get-only property
and constructor parameter (plain ctor assignment — net48-safe; update all construction sites,
including tests). Named generically ("an STA call outside the dispatcher is executing") so any
future non-dispatcher STA work reuses it.
3. `MxAccess/MxAccessStaSession.CaptureHeartbeat` (:364-381): pass `staAlarmPollInProgress`.
4. `Ipc/WorkerPipeSession.ReportWatchdogFaultIfNeededAsync` (:850-861): change the suppression
condition to
`(!string.IsNullOrEmpty(snapshot.CurrentCommandCorrelationId) || snapshot.StaCallInProgress) && staleFor <= _sessionOptions.HeartbeatStuckCeiling`;
extend the branch comment.
5. Docs same commit: `docs/MxAccessWorkerInstanceDesign.md` watchdog section (~:656-702) — one
paragraph: the alarm poll runs outside the dispatcher, advertises itself via the snapshot's
STA-call-in-progress flag, and receives the same grace-to-ceiling suppression as dispatched
commands.
**Verification.** On windev:
`dotnet build src/ZB.MOM.WW.MxGateway.Worker/ZB.MOM.WW.MxGateway.Worker.csproj -p:Platform=x86`
then
`dotnet test src/ZB.MOM.WW.MxGateway.Worker.Tests/ZB.MOM.WW.MxGateway.Worker.Tests.csproj -p:Platform=x86 --filter "FullyQualifiedName~WorkerPipeSessionTests|FullyQualifiedName~MxAccessStaSessionTests"`.
New tests:
- `WorkerPipeSessionTests.Watchdog_StaCallInProgress_SuppressedUntilCeiling` — fake runtime
returns a snapshot with activity stale beyond `HeartbeatGrace` but within the ceiling, empty
correlation id, `StaCallInProgress = true`; assert no `StaHung` fault; then a snapshot stale
beyond the ceiling; assert the fault fires (parity with the existing command-suppression tests).
- `MxAccessStaSessionTests.CaptureHeartbeat_DuringAlarmPoll_ReportsStaCallInProgress` — block
`PollOnce` on a gate in a stub alarm handler, capture a heartbeat mid-poll, assert the flag;
release, assert it clears.
---
## WRK-28 — 10,000 drain cap is a duplicated magic constant with a comment-only sync contract `Low` · `—`
**Finding.** `Ipc/WorkerPipeSession.cs:26` (`MaxDrainEventsPerReply = 10_000`, "Kept in step with
that public ceiling") vs `src/ZB.MOM.WW.MxGateway.Server/Grpc/MxAccessGrpcRequestValidator.cs:13`
(`MaxDrainEventsPerRequest = 10_000`). Both projects already reference the Contracts assembly
(`net10.0;net48`), which could own the constant.
**Impact.** Conventions/maintenance: nothing enforces the sync; the next person to tune one side
silently desynchronizes the gateway's loud rejection threshold from the worker's clamp.
**Design.** Move the ceiling into the Contracts project next to the existing protocol constants in
`GatewayContractInfo` (`src/ZB.MOM.WW.MxGateway.Contracts/GatewayContractInfo.cs`) and reference
it from both sides. A C# `const` only — **no `.proto` change**, so no generated-code regeneration
and no language-client impact (the cap is documented behavior, not wire schema). Land in the same
commit cluster as WRK-21, which rewrites the exact worker lines that consume it. Rejected
alternative: leaving two constants with a unit test asserting equality — works, but the shared
constant removes the failure mode instead of detecting it.
**Implementation.**
1. `src/ZB.MOM.WW.MxGateway.Contracts/GatewayContractInfo.cs`: add
`public const uint MaxDrainEventsPerCommand = 10_000;` with XML doc naming both consumers (the
gateway's request-validation ceiling and the worker's per-reply clamp) and the byte-cap caveat
from WRK-21 (count is necessary, not sufficient).
2. `src/ZB.MOM.WW.MxGateway.Server/Grpc/MxAccessGrpcRequestValidator.cs:13,81-85`: delete the
private const; reference `GatewayContractInfo.MaxDrainEventsPerCommand` (error text unchanged).
3. `src/ZB.MOM.WW.MxGateway.Worker/Ipc/WorkerPipeSession.cs:26,545-557`: delete
`MaxDrainEventsPerReply`; reference the shared constant; trim the "kept in step" comment to
point at the shared home.
4. Docs same commit: none required (no behavior change); if `docs/GatewayConfiguration.md` or the
DrainEvents prose names the literal 10,000, point it at the shared constant's home.
**Verification.** Both sides build: on macOS
`dotnet build src/ZB.MOM.WW.MxGateway.NonWindows.slnx` and
`dotnet test src/ZB.MOM.WW.MxGateway.Tests/ZB.MOM.WW.MxGateway.Tests.csproj --filter "FullyQualifiedName~MxAccessGrpcRequestValidator"`;
on windev
`dotnet build src/ZB.MOM.WW.MxGateway.Worker/ZB.MOM.WW.MxGateway.Worker.csproj -p:Platform=x86`
and
`dotnet test src/ZB.MOM.WW.MxGateway.Worker.Tests/ZB.MOM.WW.MxGateway.Worker.Tests.csproj -p:Platform=x86 --filter "FullyQualifiedName~WorkerPipeSessionTests"`.
Existing drain-bound tests on both sides are the regression guard; no new tests needed beyond
WRK-21's.
---
## Cross-references and sequencing
- **WRK-21 = IPC-23** (single defect, single fix — owned here; the IPC register should link to
this entry). Its `ProcessCommandAsync` backstop narrows command replies only; the oversized
*event-stream* frame policy remains **IPC-30** (out of scope here — a delivery-expectations
decision, not a mechanics fix).
- **WRK-22 = IPC-26** (single defect, single fix — owned here). WRK-25 builds on WRK-22's
claim/tombstone machinery — land WRK-22 first or together.
- **WRK-23** interacts with WRK-21: byte capping makes writer-level `MessageTooLarge` a rare
backstop, and WRK-23 then makes the gap-free sequence guarantee (already asserted by an existing
test) actually hold across that backstop.
- **WRK-26** (P1) lands before the P2/untierred code items — its text must describe HEAD at commit
time; WRK-23 and WRK-25 each update the affected sentences in their own commits per the
same-commit docs rule. WRK-26 also discharges **IPC-29**.
- Suggested commit clusters: (1) WRK-21 + WRK-28 [P0]; (2) WRK-26 [P1 doc batch, with TST-27 /
CLI-42 per roadmap item 8]; (3) WRK-22 + WRK-25 [writer seam]; (4) WRK-23; (5) WRK-24; (6)
WRK-27. Run the full worker suite on windev once per cluster batch, not per task
(`dotnet test src/ZB.MOM.WW.MxGateway.Worker.Tests/ZB.MOM.WW.MxGateway.Worker.Tests.csproj -p:Platform=x86`).
- Prior-cycle open items touching the same code (WRK-02/03/05/08/09/10/13/14/16/17/18/19 —
confirmed still open by this review) remain tracked in `archreview/remediation/20-worker.md`;
WRK-26 deliberately folds in WRK-18's stale overflow *prose* but leaves its exit-code half open.