Merge remote-tracking branch 'origin/fix/wrk-21-drain-cluster'
ci / windows-x86 (push) Successful in 1m27s
ci / nightly-windev (push) Has been skipped
ci / java (push) Successful in 2m11s
ci / portable (push) Failing after 4m7s

# Conflicts:
#	archreview/2026-07-12/remediation/00-tracking.md
#	docs/MxAccessWorkerInstanceDesign.md
This commit is contained in:
Joseph Doherty
2026-08-07 07:18:58 -04:00
18 changed files with 1299 additions and 57 deletions
+44
View File
@@ -378,6 +378,20 @@ If event conversion throws, catch it inside the event handler, record a
structured `WorkerFault`, and keep the worker alive only if the fault policy
allows it.
The event drain loop streams queued events as `WorkerEvent` frames. A single
event whose envelope exceeds the negotiated frame maximum is **undeliverable end
to end** — the pipe maximum sits only the envelope-overhead reserve above the
public gRPC cap, so a frame the pipe rejects would also be rejected on the
client-facing stream. The session therefore faults on it rather than dropping it
(a silent drop makes the event stream unfaithful, and a synthesized placeholder
is barred by the no-synthesized-events rule), but the death is structured: the
worker logs the event's identity — family, handles, worker sequence, and sizes,
never the value — writes a `WorkerFault` with category `ProtocolViolation` and
command method `EventDrain` carrying the same identity, and only then exits.
Operator remediation is configuration: raise `MxGateway:Worker:MaxMessageBytes`
for that workload. Other per-frame rejection codes keep their previous behavior
because they indicate worker bugs, not workload size.
## Command Queue
The pipe reader converts `WorkerCommand` messages into `StaCommand` entries.
@@ -440,6 +454,29 @@ Diagnostics:
- `DrainEvents`
- `ShutdownWorker`
`DrainEvents` is answered on the message-loop thread, not the STA, and its reply
is bounded on **two** axes because no diagnostics command may be session-fatal:
- **Count** — `GatewayContractInfo.MaxDrainEventsPerCommand` (10,000) is the
single home of the ceiling, shared by the gateway's request validator (which
rejects a larger `max_events` at the public boundary) and this worker clamp
(which also interprets `max_events = 0`, "as many as available").
- **Bytes** — the count cap alone is not sufficient: byte-heavy events (large
string or array `MxValue`s) overshoot the negotiated frame maximum long before
10,000 events. The drain is therefore byte-budgeted against the negotiated
maximum less a 64 KiB envelope/reply-wrapper reserve, and the size decision
happens inside the event queue's lock, so an event is dequeued only once it is
known to fit. An event that does not fit stays at the head of the queue and is
never lost.
Truncation is reported in the reply's existing `DiagnosticMessage`
("N events returned, M remain; repeat DrainEvents for the rest") rather than in a
new field, so the contract is unchanged and callers drain iteratively until a
reply comes back empty. In the degenerate case where the head event alone exceeds
the budget, the reply returns whatever fit before it (possibly nothing) and names
the blocked event's worker sequence so an operator can find the offending tag;
that event needs a larger `MxGateway:Worker:MaxMessageBytes` to move at all.
Implement method-specific dispatch instead of a generic string method invoker.
Parity tests need stable command-specific request and reply shapes.
@@ -634,6 +671,13 @@ consumer that merely drains slower than this worker produces (the staging bound,
metric `QueueOverflow("worker-event-staging")`). See
`docs/GatewayProcessDesign.md`. A worker that outruns its consumer therefore
dies loudly rather than growing gateway memory silently.
No control reply is session-fatal on size. Reply builders size their payloads
against the negotiated frame maximum, and the two reply-write seams (the
control-command path and the STA command path) additionally catch a
`MessageTooLarge` per-frame rejection and answer that correlation with a small
`InvalidRequest` reply instead of unwinding the session. Oversized *event*
frames keep the opposite policy — see Event Sink — because an event above the
frame maximum cannot be delivered to the client at all.
## Heartbeat And Watchdog
+23
View File
@@ -29,6 +29,29 @@ default. A `max_frame_bytes` of 0 (an older gateway that never set the field)
means "use the worker's built-in default". This keeps both ends framing to the
same limit rather than depending on matched compile-time constants.
Every worker-to-gateway frame must serialize within this limit, control replies
included, so reply builders truncate to fit rather than emit a frame the writer
will reject. `WorkerPipeSession` pre-sizes a `DrainEvents` reply below the
negotiated maximum (less a fixed envelope/reply-wrapper reserve) and reports the
truncation in the reply's `DiagnosticMessage`; the caller contract is to repeat
`DrainEvents` until it returns an empty reply. Should a reply still overshoot,
`MessageTooLarge` at the reply-write seam is answered with a small
`InvalidRequest` reply for that correlation, not with session teardown — no
diagnostics command may kill a session.
An oversized *event* frame is the deliberate exception. Such an event is
undeliverable end to end (the pipe maximum sits only the envelope-overhead
reserve above the public gRPC cap), so the session faults: the worker logs the
event's identity and sizes — never its value — writes a `WorkerFault` with
category `ProtocolViolation` and command method `EventDrain`, then exits.
Remediation is raising `MxGateway:Worker:MaxMessageBytes` for that workload.
A per-frame rejection does not consume an envelope `sequence`. The writer stamps
a candidate sequence, runs the empty-payload and size checks against the stamped
envelope, and commits the counter only immediately before the stream write, so
the sequences observed on the wire stay contiguous across rejections and an
operator reading a pipe capture never sees a phantom gap.
## Envelope Validation
`WorkerFrameReader` and `WorkerFrameWriter` validate each envelope against the