Merge remote-tracking branch 'origin/fix/wrk-21-drain-cluster'
# Conflicts: # archreview/2026-07-12/remediation/00-tracking.md # docs/MxAccessWorkerInstanceDesign.md
This commit is contained in:
@@ -378,6 +378,20 @@ If event conversion throws, catch it inside the event handler, record a
|
||||
structured `WorkerFault`, and keep the worker alive only if the fault policy
|
||||
allows it.
|
||||
|
||||
The event drain loop streams queued events as `WorkerEvent` frames. A single
|
||||
event whose envelope exceeds the negotiated frame maximum is **undeliverable end
|
||||
to end** — the pipe maximum sits only the envelope-overhead reserve above the
|
||||
public gRPC cap, so a frame the pipe rejects would also be rejected on the
|
||||
client-facing stream. The session therefore faults on it rather than dropping it
|
||||
(a silent drop makes the event stream unfaithful, and a synthesized placeholder
|
||||
is barred by the no-synthesized-events rule), but the death is structured: the
|
||||
worker logs the event's identity — family, handles, worker sequence, and sizes,
|
||||
never the value — writes a `WorkerFault` with category `ProtocolViolation` and
|
||||
command method `EventDrain` carrying the same identity, and only then exits.
|
||||
Operator remediation is configuration: raise `MxGateway:Worker:MaxMessageBytes`
|
||||
for that workload. Other per-frame rejection codes keep their previous behavior
|
||||
because they indicate worker bugs, not workload size.
|
||||
|
||||
## Command Queue
|
||||
|
||||
The pipe reader converts `WorkerCommand` messages into `StaCommand` entries.
|
||||
@@ -440,6 +454,29 @@ Diagnostics:
|
||||
- `DrainEvents`
|
||||
- `ShutdownWorker`
|
||||
|
||||
`DrainEvents` is answered on the message-loop thread, not the STA, and its reply
|
||||
is bounded on **two** axes because no diagnostics command may be session-fatal:
|
||||
|
||||
- **Count** — `GatewayContractInfo.MaxDrainEventsPerCommand` (10,000) is the
|
||||
single home of the ceiling, shared by the gateway's request validator (which
|
||||
rejects a larger `max_events` at the public boundary) and this worker clamp
|
||||
(which also interprets `max_events = 0`, "as many as available").
|
||||
- **Bytes** — the count cap alone is not sufficient: byte-heavy events (large
|
||||
string or array `MxValue`s) overshoot the negotiated frame maximum long before
|
||||
10,000 events. The drain is therefore byte-budgeted against the negotiated
|
||||
maximum less a 64 KiB envelope/reply-wrapper reserve, and the size decision
|
||||
happens inside the event queue's lock, so an event is dequeued only once it is
|
||||
known to fit. An event that does not fit stays at the head of the queue and is
|
||||
never lost.
|
||||
|
||||
Truncation is reported in the reply's existing `DiagnosticMessage`
|
||||
("N events returned, M remain; repeat DrainEvents for the rest") rather than in a
|
||||
new field, so the contract is unchanged and callers drain iteratively until a
|
||||
reply comes back empty. In the degenerate case where the head event alone exceeds
|
||||
the budget, the reply returns whatever fit before it (possibly nothing) and names
|
||||
the blocked event's worker sequence so an operator can find the offending tag;
|
||||
that event needs a larger `MxGateway:Worker:MaxMessageBytes` to move at all.
|
||||
|
||||
Implement method-specific dispatch instead of a generic string method invoker.
|
||||
Parity tests need stable command-specific request and reply shapes.
|
||||
|
||||
@@ -634,6 +671,13 @@ consumer that merely drains slower than this worker produces (the staging bound,
|
||||
metric `QueueOverflow("worker-event-staging")`). See
|
||||
`docs/GatewayProcessDesign.md`. A worker that outruns its consumer therefore
|
||||
dies loudly rather than growing gateway memory silently.
|
||||
No control reply is session-fatal on size. Reply builders size their payloads
|
||||
against the negotiated frame maximum, and the two reply-write seams (the
|
||||
control-command path and the STA command path) additionally catch a
|
||||
`MessageTooLarge` per-frame rejection and answer that correlation with a small
|
||||
`InvalidRequest` reply instead of unwinding the session. Oversized *event*
|
||||
frames keep the opposite policy — see Event Sink — because an event above the
|
||||
frame maximum cannot be delivered to the client at all.
|
||||
|
||||
## Heartbeat And Watchdog
|
||||
|
||||
|
||||
Reference in New Issue
Block a user