fix(WRK-21,WRK-28,WRK-23,IPC-30): byte-budget the DrainEvents reply, stop size errors from killing sessions
ci / nightly-windev (push) Has been skipped
ci / java (push) Successful in 2m8s
ci / portable (push) Successful in 7m41s
ci / windows-x86 (push) Failing after 12m32s

WRK-21 — DrainEvents was bounded by event count only, so a byte-heavy queue
(large string/array MxValues) built a reply above the negotiated frame maximum:
the writer rejected the frame, the exception unwound the session, and the events
already dequeued were destroyed. The drain is now byte-budgeted inside the queue
lock, so an event is dequeued only once it is known to fit and one that does not
stays at the head. Truncation is reported through the reply's existing
DiagnosticMessage (no contract change); callers drain until an empty reply. Both
reply-write seams — the control-command path and ProcessCommandAsync — now catch
MessageTooLarge and answer the correlation with an InvalidRequest reply instead
of unwinding or faulting the session. Satisfies IPC-23 R1-R3.

WRK-28 — the 10,000 drain ceiling moves to GatewayContractInfo
.MaxDrainEventsPerCommand, referenced by both the gateway request validator and
the worker clamp, replacing a comment-only sync contract. C# const only; no
.proto change.

WRK-23 — WorkerFrameWriter now peek-stamps, validates, then commits the sequence
counter immediately before the stream write, so a per-frame rejection leaves no
phantom gap on the wire.

IPC-30 — an oversized event frame stays session-fatal (it is undeliverable end to
end and neither dropping nor synthesizing a replacement is allowed), but the
death is structured: the event's identity and sizes are logged (never its value),
a WorkerFault with category PROTOCOL_VIOLATION and command method EventDrain is
written, then the session exits as before.

Docs updated in the same change: MxAccessWorkerInstanceDesign.md (drain byte cap,
truncation contract, oversized-head behavior, oversized-event policy, no control
reply is session-fatal on size), WorkerFrameProtocol.md (reply pre-sizing,
non-fatal reply-size rule, oversized-event policy, rejected frames do not consume
sequence numbers), gateway.md (DrainEvents two-axis bound).
This commit is contained in:
Joseph Doherty
2026-08-07 05:38:23 -04:00
parent ead921cace
commit 33ba612ddd
18 changed files with 1118 additions and 57 deletions
+45
View File
@@ -378,6 +378,20 @@ If event conversion throws, catch it inside the event handler, record a
structured `WorkerFault`, and keep the worker alive only if the fault policy
allows it.
The event drain loop streams queued events as `WorkerEvent` frames. A single
event whose envelope exceeds the negotiated frame maximum is **undeliverable end
to end** — the pipe maximum sits only the envelope-overhead reserve above the
public gRPC cap, so a frame the pipe rejects would also be rejected on the
client-facing stream. The session therefore faults on it rather than dropping it
(a silent drop makes the event stream unfaithful, and a synthesized placeholder
is barred by the no-synthesized-events rule), but the death is structured: the
worker logs the event's identity — family, handles, worker sequence, and sizes,
never the value — writes a `WorkerFault` with category `ProtocolViolation` and
command method `EventDrain` carrying the same identity, and only then exits.
Operator remediation is configuration: raise `MxGateway:Worker:MaxMessageBytes`
for that workload. Other per-frame rejection codes keep their previous behavior
because they indicate worker bugs, not workload size.
## Command Queue
The pipe reader converts `WorkerCommand` messages into `StaCommand` entries.
@@ -440,6 +454,29 @@ Diagnostics:
- `DrainEvents`
- `ShutdownWorker`
`DrainEvents` is answered on the message-loop thread, not the STA, and its reply
is bounded on **two** axes because no diagnostics command may be session-fatal:
- **Count** — `GatewayContractInfo.MaxDrainEventsPerCommand` (10,000) is the
single home of the ceiling, shared by the gateway's request validator (which
rejects a larger `max_events` at the public boundary) and this worker clamp
(which also interprets `max_events = 0`, "as many as available").
- **Bytes** — the count cap alone is not sufficient: byte-heavy events (large
string or array `MxValue`s) overshoot the negotiated frame maximum long before
10,000 events. The drain is therefore byte-budgeted against the negotiated
maximum less a 64 KiB envelope/reply-wrapper reserve, and the size decision
happens inside the event queue's lock, so an event is dequeued only once it is
known to fit. An event that does not fit stays at the head of the queue and is
never lost.
Truncation is reported in the reply's existing `DiagnosticMessage`
("N events returned, M remain; repeat DrainEvents for the rest") rather than in a
new field, so the contract is unchanged and callers drain iteratively until a
reply comes back empty. In the degenerate case where the head event alone exceeds
the budget, the reply returns whatever fit before it (possibly nothing) and names
the blocked event's worker sequence so an operator can find the offending tag;
that event needs a larger `MxGateway:Worker:MaxMessageBytes` to move at all.
Implement method-specific dispatch instead of a generic string method invoker.
Parity tests need stable command-specific request and reply shapes.
@@ -623,6 +660,14 @@ queue fills:
Production coalescing may be added later, but it must be explicit and tested.
Do not drop or coalesce events in v1.
No control reply is session-fatal on size. Reply builders size their payloads
against the negotiated frame maximum, and the two reply-write seams (the
control-command path and the STA command path) additionally catch a
`MessageTooLarge` per-frame rejection and answer that correlation with a small
`InvalidRequest` reply instead of unwinding the session. Oversized *event*
frames keep the opposite policy — see Event Sink — because an event above the
frame maximum cannot be delivered to the client at all.
## Heartbeat And Watchdog
`WorkerPipeSession` starts the heartbeat loop after the gateway validates