8df35cd63a
WRK-22/IPC-26: tombstone a WriteAsync/WriteBatchAsync cancelled while waiting for the write lock (PendingFrame.Claimed under _gate; DequeueNext skips cancelled, claims the frame it returns) so a cancelled write never reaches the wire unless already claimed mid-write (documented residual). WRK-25: add WriteBatchAsync; RunEventDrainLoopAsync submits the drained event batch through it, so a burst of N events costs one flush not N. IPC-30 oversized-event structured fault preserved via FindOversizedEvent. WRK-24: reject a below-1024 negotiated frame maximum at the handshake (MinNegotiableFrameBytes, matching GatewayOptionsValidator floor). WRK-27: alarm poll advertises StaCallInProgress on the heartbeat snapshot so the watchdog suppresses to the ceiling, not the grace. Docs (WorkerFrameProtocol.md, MxAccessWorkerInstanceDesign.md) and the 2026-07-12 remediation registers/change-log updated in the same commit.
174 lines
9.3 KiB
Markdown
174 lines
9.3 KiB
Markdown
# Worker Frame Protocol
|
|
|
|
The gateway uses the worker frame protocol to move `WorkerEnvelope` protobuf
|
|
messages over a bidirectional named pipe. The frame layer is deliberately small:
|
|
it handles message boundaries, size limits, protobuf parsing, and envelope
|
|
validation before higher-level worker client code routes commands, replies,
|
|
events, and faults.
|
|
|
|
## Frame Format
|
|
|
|
Each frame starts with a four-byte little-endian unsigned payload length,
|
|
followed by the serialized `WorkerEnvelope` payload:
|
|
|
|
```text
|
|
uint32 little-endian payload_length
|
|
payload_length bytes protobuf WorkerEnvelope
|
|
```
|
|
|
|
The reader rejects zero-length payloads and payloads larger than the configured
|
|
maximum before allocating the payload buffer. The default maximum is the 16 MiB
|
|
public gRPC cap plus a 64 KiB envelope-overhead reserve (16842752 bytes) so a
|
|
maximally-sized accepted gRPC payload always fits one worker frame once wrapped
|
|
in a `WorkerEnvelope`.
|
|
|
|
The gateway is the source of truth for this maximum: it conveys the negotiated
|
|
value in the handshake as `GatewayHello.max_frame_bytes`, and the worker adopts
|
|
it as its `WorkerFrameProtocolOptions.MaxMessageBytes` instead of a hard-coded
|
|
default. A `max_frame_bytes` of 0 (an older gateway that never set the field)
|
|
means "use the worker's built-in default". This keeps both ends framing to the
|
|
same limit rather than depending on matched compile-time constants.
|
|
|
|
The worker accepts a negotiated value in the closed range [1024, 256 MiB]
|
|
(`MinNegotiableFrameBytes` .. `MaxNegotiableFrameBytes`); 0 keeps the default.
|
|
A value outside that range is rejected at the handshake with a fault frame
|
|
rather than adopted, because a nonsensical maximum — a gateway bug or a
|
|
foreign/old peer — would otherwise leave a session that handshakes cleanly and
|
|
then fails every subsequent frame with per-frame size errors, the worst
|
|
diagnostic shape for an operator. The 1024-byte floor matches the gateway's own
|
|
`GatewayOptionsValidator.MinimumMaxMessageBytes`, so the worker never rejects a
|
|
value the gateway's validator accepts as legal configuration, and 1024 still
|
|
guarantees hellos, heartbeats, acks, and faults fit.
|
|
|
|
Every worker-to-gateway frame must serialize within this limit, control replies
|
|
included, so reply builders truncate to fit rather than emit a frame the writer
|
|
will reject. `WorkerPipeSession` pre-sizes a `DrainEvents` reply below the
|
|
negotiated maximum (less a fixed envelope/reply-wrapper reserve) and reports the
|
|
truncation in the reply's `DiagnosticMessage`; the caller contract is to repeat
|
|
`DrainEvents` until it returns an empty reply. Should a reply still overshoot,
|
|
`MessageTooLarge` at the reply-write seam is answered with a small
|
|
`InvalidRequest` reply for that correlation, not with session teardown — no
|
|
diagnostics command may kill a session.
|
|
|
|
An oversized *event* frame is the deliberate exception. Such an event is
|
|
undeliverable end to end (the pipe maximum sits only the envelope-overhead
|
|
reserve above the public gRPC cap), so the session faults: the worker logs the
|
|
event's identity and sizes — never its value — writes a `WorkerFault` with
|
|
category `ProtocolViolation` and command method `EventDrain`, then exits.
|
|
Remediation is raising `MxGateway:Worker:MaxMessageBytes` for that workload.
|
|
|
|
A per-frame rejection does not consume an envelope `sequence`. The writer stamps
|
|
a candidate sequence, runs the empty-payload and size checks against the stamped
|
|
envelope, and commits the counter only immediately before the stream write, so
|
|
the sequences observed on the wire stay contiguous across rejections and an
|
|
operator reading a pipe capture never sees a phantom gap.
|
|
|
|
## Envelope Validation
|
|
|
|
`WorkerFrameReader` and `WorkerFrameWriter` validate each envelope against the
|
|
owning session before returning or writing it:
|
|
|
|
- `protocol_version` must match the configured worker protocol version,
|
|
- `session_id` must match the owning gateway session,
|
|
- the envelope must contain one typed `body` value.
|
|
|
|
Protocol violations throw `WorkerFrameProtocolException` with a
|
|
`WorkerFrameProtocolErrorCode` so callers can distinguish malformed frames,
|
|
oversized frames, protocol version mismatches, and session mismatches.
|
|
|
|
## Write Scheduling And Sequencing
|
|
|
|
This section covers write scheduling (priority classes, enqueue-then-contend,
|
|
flush coalescing) and sequencing (write-time stamping) together, because both
|
|
are properties of the same single write lock.
|
|
|
|
`WorkerFrameWriter` is a two-class cooperative priority scheduler
|
|
(`WorkerFrameWritePriority.Control` and `.Event`), not a strict per-kind
|
|
priority order. A caller enqueues its frame into the control or event queue
|
|
under a lock, then contends for a single write lock; whichever caller wins
|
|
drains every frame queued at that moment, control frames first and each class
|
|
in FIFO order, so a command reply, fault, heartbeat, or shutdown
|
|
acknowledgement is never delayed behind a backlog of queued events. Priority
|
|
only reorders *which frame writes next* — it does not affect the sequence
|
|
value a frame receives (see below), so a caller cannot infer priority class
|
|
from the wire sequence.
|
|
|
|
The envelope `Sequence` is stamped by the draining lock-holder at the actual
|
|
moment of writing, not when the frame is enqueued, so the on-wire order and
|
|
the stamped sequence always agree regardless of caller concurrency or
|
|
priority reordering. Stamping uses peek-stamp-commit: a candidate sequence is
|
|
assigned and the frame is validated (size, non-empty payload) against that
|
|
stamped value, but the counter is committed only immediately before the
|
|
stream write. A per-frame rejection therefore leaves the counter untouched —
|
|
the next accepted frame reuses the candidate number, so the wire sequence
|
|
stays contiguous across rejections and an operator reading a pipe capture
|
|
never sees a phantom gap from a rejected frame.
|
|
|
|
Two failure shapes are distinguished during a drain pass:
|
|
|
|
- **Per-frame rejection** (`InvalidEnvelope`, `MessageTooLarge`,
|
|
`ProtocolVersionMismatch`, `SessionMismatch`) is specific to the one frame
|
|
that failed validation or sizing. Nothing was written for it, so it fails
|
|
only that frame's completion and draining continues with the next queued
|
|
frame.
|
|
- **Stream failure** (anything else — a broken pipe, an I/O error) means the
|
|
underlying stream itself is no longer trustworthy. It fails the frame that
|
|
triggered it, every frame already written this batch but not yet flushed,
|
|
and every frame still queued, then stops draining entirely so no caller
|
|
waits forever on a stream that will not recover.
|
|
|
|
Flushes are coalesced across a drained batch: each frame in the batch is
|
|
written to the stream without an individual flush, then one `FlushAsync`
|
|
runs after the whole batch, and only then does every successfully-written
|
|
frame's completion resolve — so a caller's `WriteAsync` still does not
|
|
complete until its bytes are both written *and* flushed, but a batch that
|
|
happened to contain several queued frames pays one flush instead of one per
|
|
frame. The event drain loop (`WorkerPipeSession.RunEventDrainLoopAsync`)
|
|
submits a whole drained event batch through `WriteBatchAsync`, which enqueues
|
|
every frame under one `_gate` acquisition, takes the write lock once, and
|
|
drains them together, so a burst of N events costs one flush rather than N —
|
|
the coalescing the batch machinery was built for now engages on the event hot
|
|
path, not only when independent producers happen to queue behind a blocked
|
|
write. Intra-batch order is preserved (FIFO enqueue under one lock), and a
|
|
concurrently queued control frame is still drained ahead of the batch. A
|
|
per-frame rejection inside a batch (for example one oversized event) surfaces
|
|
from the batch's awaited completions as that frame's
|
|
`WorkerFrameProtocolException`; the remaining completions are still observed
|
|
so none faults unobserved.
|
|
|
|
Cancellation of a `WriteAsync`/`WriteBatchAsync` call that is still waiting
|
|
for the write lock when its token fires tombstones the queued frame: the
|
|
cancelled caller marks its frame under `_gate`, and the draining lock-holder's
|
|
`DequeueNext` skips any tombstoned frame, so a cancelled call is guaranteed
|
|
never to reach the wire — *unless* a lock-holder has already claimed the frame
|
|
to write it. Claiming and cancelling are interlocked under `_gate`, so exactly
|
|
one wins; a frame already claimed is mid-write and can no longer be recalled,
|
|
so the caller observes `OperationCanceledException` while that one frame still
|
|
reaches the wire. That residual window is by design: blocking the canceller
|
|
behind the very write it is abandoning would defeat the point of cancellation.
|
|
|
|
## Verification
|
|
|
|
The frame protocol lives in `ZB.MOM.WW.MxGateway.Worker.Ipc` (`WorkerFrameReader`,
|
|
`WorkerFrameWriter`, `WorkerFrameProtocolOptions`) and is covered by
|
|
`src/ZB.MOM.WW.MxGateway.Worker.Tests/Ipc/WorkerFrameProtocolTests.cs`. The worker is an
|
|
x86 process, so build and test it with `-p:Platform=x86`.
|
|
|
|
Run the focused tests after changing the frame protocol:
|
|
|
|
```powershell
|
|
dotnet test src/ZB.MOM.WW.MxGateway.Worker.Tests/ZB.MOM.WW.MxGateway.Worker.Tests.csproj -p:Platform=x86 --filter WorkerFrameProtocolTests
|
|
```
|
|
|
|
Run the x86 worker build because the frame protocol is part of
|
|
`ZB.MOM.WW.MxGateway.Worker`:
|
|
|
|
```powershell
|
|
dotnet build src/ZB.MOM.WW.MxGateway.Worker/ZB.MOM.WW.MxGateway.Worker.csproj -p:Platform=x86
|
|
```
|
|
|
|
## Related Documentation
|
|
|
|
- [Gateway Process Detailed Design](./GatewayProcessDesign.md)
|
|
- [Protobuf Contracts](./Contracts.md)
|