The BlockDispatch branch waited 5 real seconds and then proceeded regardless — it does not
branch on the wait's result. The long-in-flight test's inspection loop is bounded by
elapsed time and a frame floor, so on a loaded box (the documented 4-5x slowdown class) it
can plausibly outrun that 5 s. When it does, the reply is emitted mid-window,
AssertNotWorkerFault waves it past, and the reply leg then waits for a reply already gone
by — failing at the 20 s cancellation with no message, on exactly the loaded-box run the
widened windows exist to survive.
The wait is a pure safety net: nothing asserts on it firing, and every test that blocks
dispatch releases it explicitly (ReleaseDispatch, or a WorkerShutdown envelope, both of
which Set the event) — none reaches the timeout on a healthy run. Named it
BlockedDispatchSafetyNet and raised it to 30 s, above any window a test opens and above the
20 s cancellation those tests arm, so a wedged test always fails on its own token with its
own message. Dispose still releases the wait, so teardown never waits on it either.
Inline rationale in the test now states the decoupling and what a close pairing would cost,
rather than asserting the window stays inside a 5 s ceiling. Nothing else changed.
Three review findings on 7da52b6:
1. The AssertNotFault doc block landed between the predicate-overload ReadUntilAsync's
doc block and the helper, so the compiler attached the merged block to the helper and
ReadUntilAsync lost its docs entirely. Helper and its docs moved above ReadUntilAsync,
whose docs are back where they belong.
2. The compressed watchdog windows (50 ms grace, 100 ms ceiling) reintroduced the
load-sensitivity the fix removed, one layer down.
ReportWatchdogFaultIfNeededAsync measures staleness AFTER the heartbeat frame is
written and flushed over the real named pipe, so a beat whose pipe I/O outlasts the
ceiling faults a healthy session no matter how fresh the captured activity was. At
100 ms that is a plausible stall on a loaded box, and this is the one test asserting
the watchdog NEVER fires. Widened to a 200 ms grace and a 1 s ceiling — still two
orders of magnitude under the 75 s production default.
The inspection loop is now bounded by a 2 s window (twice the ceiling, so a fake whose
activity stopped advancing still accumulates past it and faults) with a 30-frame floor,
rather than a fixed 30 frames that no longer outran the wider ceiling. The floor keeps
a window that saw almost no beats from passing as a clean one. Two seconds stays well
inside FakeRuntimeSession's 5 s dispatch-block ceiling, so the command is still in
flight for the whole window.
3. AssertNotFault renamed AssertNotWorkerFault, matching the WorkerFault body case it
tests.
Scenario intent unchanged: long in-flight command, pump refreshing, zero fault frames,
reply delivered. Test-only.
The In-process page feeds table still described the alarms page's subscription as
provider status only, contradicting the three passages updated in 7b6dfba.
The loop's catch-all comment (and the cadence bullet that repeated it) claimed the
monitor completes a subscriber's stream on restart. It does not: ClearCache pushes
snapshot_status(false) through the still-open channel, and a subscriber is only
completed-with-error on a failed TryWrite, or removed by its own disposal.
The SnapshotStatus arm's priming claim had no test behind it — every push test
attached to an untruncated feed and pushed the edge itself. ScriptedAlarmFeed now
replays an optional open sequence, and a new test attaches to a feed already
primed provider_status then snapshot_status(truncated) and asserts the banner
comes up with no edge pushed after render.
Audited every Server-0xx finding whose Resolution described a
documentation-only or comment-only change, and spot-checked the doc
sub-claims of otherwise test-backed resolutions. 20 entries annotated
in place (append-only; no historical resolution text rewritten).
Two corrections had not survived and are re-applied:
Server-040: the MapGroupsToRoles lookup-precedence comment moved intact
into DashboardGroupRoleMapping (792e3f9) and was then deleted wholesale
by fca978d, a sweep meant only to strip (Server-NNN) tracking markers.
That also removed a later, substantive paragraph recording that the
shared ZB.MOM.WW.Auth.Ldap provider pre-strips groups to short RDN
names, so a full-DN GroupToRole key is unsupported. Both paragraphs
restored, minus the tracking IDs.
Server-009: the WAL / busy_timeout note vanished when the Storage
section of docs/Authentication.md was rewritten to delegate
connection-factory detail to ZB.MOM.WW.Auth.ApiKeys. The behavior is
still live in the library (confirmed against 0.2.1), so the fix is
prose-only.
Server-011/014/022/023 are annotated as moot rather than regressed:
the IAlarmRpcDispatcher trio was deleted in dc9c0c9 and no stale
'not yet wired' / 'PR A.6/A.7' prose survives in Server source.
Server-038's documented v1 ACL gap was later closed by
IDashboardSessionAcl, so its remarks are current.
Comment/doc-only; no logic changes.
RunAsync_LongInFlightCommandThatKeepsPumping_DoesNotFaultAndDeliversReply failed
deterministically on the Windows box with a StaHung fault whose command_method was
empty and whose staleness was 233 ms — i.e. a fault raised with no command in flight,
before the scenario under test began. FakeRuntimeSession stamps LastStaActivityUtc once,
at construction, and the test only started refreshing it after the blocked dispatch
signalled. Everything between those two points — handshake, STA init, the first
heartbeat — captured a snapshot already stale past the compressed 50 ms grace, with no
correlation id for the watchdog to suppress on, so the watchdog correctly reported the
fake as hung.
The harness, not the product, was wrong: StaRuntime.ThreadMain calls MarkActivity() on
every WaitForWorkOrMessages iteration, so a live worker is never captured stale, idle or
busy. Model that where it belongs — FakeRuntimeSession.RefreshStaActivityOnCapture (opt
in, default off) stamps activity at each CaptureHeartbeat and leaves the rest of the
snapshot alone — and arm it before RunAsync so the first beat is covered. The test-owned
refresh loop goes away with it; a thread-pool loop racing a compressed grace could not
have held the invariant anyway.
Scenario intent is unchanged and slightly stronger: the command still blocks in dispatch
across 30 heartbeats (~600 ms, many multiples of the 100 ms stuck ceiling), no frame may
be a fault, and the reply must still arrive. The reply leg is now fault-checked too
(previously it skipped frames blindly), and the pump keeps running across the release, as
it does in production while the reply is marshalled off the STA. Fault assertions now
report the category and diagnostic message instead of a bare body-case mismatch.
Test-only change; no product code, frame protocol, or STA rule touched.
The truncated-snapshot caveat moved from poll-only to push-driven. The page
already held an in-process alarm-feed subscription for the provider badge; it
now also handles the feed's snapshot_status frame, so a capped provider fetch is
caveated when the monitor decides it rather than up to three seconds later.
The poll's assignment stays as the reconcile baseline — both sources read the
same monitor verdict, and the frame is consumed, never synthesized page-side.
StreamAsync primes every subscriber with a snapshot_status frame at open, so a
page attaching mid-truncation needs no priming logic of its own; the loop is
renamed StatusFeedLoopAsync because it now feeds two indicators, not one.
Both lists predate the scope rename and would mislead anyone creating a key:
CLAUDE.md's Authentication section still named the pre-rename scopes, and
docs/Authentication.md's ops.alice example passed 'read,write', which
GatewayScopes.ValidateScopes rejects outright. Same defect family as the
Build/Test/Run sample fixed in a5f843c.
Recorded as a follow-up: code review finding Server-012 claims it fixed the two
CLAUDE.md lists on 2026-05-18, but neither correction was present — a Resolved
finding is not re-examined, so the sibling Server-0xx doc resolutions want a
spot-check for the same pattern.
The prior plan's "Follow-ups recorded, not started" block described pre-branch
behavior; every item is now closed, narrowed, or restated with its evidence, so
the block no longer misleads a reader who lands on it first. The stale Rust-guard
bullet is corrected in place rather than deleted: Check 3 always existed, and
saying so is the only way the reader learns what the real (one-directional) gap was.
Also fixes the CLAUDE.md apikey sample, which named a verb the parser has never
accepted ('create'; only 'create-key' exists, no alias), omitted the required
--key-id, and listed non-canonical scope strings that GatewayScopes now rejects
at create time — the sample could not have run.
The final integration review's non-blocker reservations, all documentation
or comment truth except one test arm.
The alarm feed opens provider_status -> snapshot_status -> cached
active_alarm -> snapshot_complete, which is what GatewayAlarmMonitor has
done since the snapshot_status frame landed. Two places still described
the old order: docs/Grpc.md said provider_status arrived *after* the
initial snapshot, contradicting its own snapshot_status section two
paragraphs down, and AlarmFeedMessage's leading proto comment named
neither status frame at all. Both now state the sequence the monitor
emits, so a client author reading either one gets the frame order right.
The proto comment change flows through the generated trees (Contracts,
Go, Java) and the client descriptor set; the Rust vendored copy stays
byte-identical to canonical. Python's generator does not carry proto
comments into its output, so it has no delta.
AlarmsHubPublisherTests' valueless-payload case covered snapshot_complete
and provider_status but not snapshot_status, leaving the newest arm
unpinned against the redaction switch that must ignore it. Added.
WnWrapAlarmConsumer's ack comment led with the 2026-05-01 reading that
-55 tracks the 8-arg overload, then refuted itself six lines later with
the 2026-08-18 probe. It now leads with the observation labelled as
narrower than it reads -- mirroring the correction already in
docs/AlarmClientDiscovery.md -- so the block argues one thing: the 6-arg
call site stays for parity, and rc semantics are per the probe. A
paragraph orphaned by an earlier splice is rewrapped. Comment interior
only; the file compiles on Windows.
TST-16 gets a dated closure note rather than a rewrite: the flag it
called dead was implemented 2026-08-18. GatewayDashboardDesign's /browse
paragraph gains the failed-read carve-out GatewayConfiguration already
documented, so the two agree that a failed read keeps its - placeholder.
Follow-up to 90331b6. Three comment/prose corrections, no behaviour change.
AcknowledgeByName's comment still said the 6-arg overload "works and reaches the
alarm-history path correctly", which the same commit's own findings contradict in
three other places. It now says what was observed: rc=0 means accepted, not
applied — the 2026-08-18 probe acked a live alarm six ways and the snapshot, the
OPERATOR_NAME field, and the extension's .Acked attribute all stayed put. The -55
tracks the consumer, not the overload. Subscribe's comment gets the same
treatment: "lets AlarmAckByName succeed" becomes "return rc=0".
AlarmProbeFindings.md said the re-raise arrives as "a separate record" alongside
the returned one, which reads as coexistence and is wrong. The snapshot carries
one record per tag: the cap=1024 replies bracketing the re-raise are both
elementCount=3 (one per TestMachine_00{1,2,3}) at an identical 1613 bytes, and the
old GUID is absent from the later one. The re-raise replaces the record, so a
single poll spanning it sees the Clear and the Raise together.
Worker diff verified strictly comment-only; builds x86 on windev, 0W/0E.
AuthenticateUser("Administrator", "") + WriteSecured raises the alarm UDAs that
plain Write could not touch (SecurityError detail=1008), so the 2026-08-17 blocker
was the verb, exactly as that run's own Unblocking list predicted.
Two of the three open questions are now observed rather than assumed:
- ALARM_RECORDS/@COUNT reports the records in the reply, not the total active
count. With three alarms active it read 1 at cap 1 and 2 at cap 2. There is no
exact truncation signal to switch to, so IsTruncatedFetch's conservative rule is
the design rather than a placeholder — behaviour unchanged, only the comments.
- Clear-then-re-raise mints a new GUID; the ALM->RTN leg keeps its GUID
(reconfirming the 2026-05-01 capture). ComputeTransitions already reads the
re-raise correctly as one instance ending and another beginning.
The acknowledge leg stays unobserved for a narrower reason: every wnwrap ack
surface is inert on this rig. AlarmAckByName returns 0 from the ack-only consumer
and -55 from the SetXmlAlarmQuery-applied one, for both the 6-arg and 8-arg forms,
and neither the snapshot STATE, OPERATOR_NAME, nor the extension's own .Acked
attribute moves. That corrects AlarmClientDiscovery.md, which read the zero return
as a working ack.
Comment- and prose-only; no behaviour change. The three throwaway probes ran from
the windev CI clone and were deleted; that clone is a clean tree at ab3ff16.
The Java import block's AlarmFeedMessage already sorted after
AlarmProviderStatus before c748361; adding AlarmSnapshotStatus widened the
gap. Order all three alphabetically.
The Rust CLI tests the sibling ProviderStatus render path but not the new
SnapshotStatus arm, so the summary string and the JSON shape were both
uncovered. Add the matching test over alarm_feed_message_summary and
alarm_feed_message_to_json.
Task 1 added AlarmSnapshotStatus and AlarmFeedMessage.snapshot_status = 5.
Carry it downstream from the canonical Contracts protos:
- Rust vendored protos under clients/rust/protos, refreshed byte-identical
(build.rs falls back to them for out-of-repo tarball builds)
- client descriptor set (protoc 34.1 pin)
- Go (protoc-gen-go v1.36.11 / protoc-gen-go-grpc 1.6.2)
- Python (grpcio-tools 1.80.0 pin)
- Java (gradle generateProto)
.NET needs no regeneration: the client compiles against the Contracts
Generated/ output committed with the proto change.
The hand-written CLI feed renderers switch on the payload oneof, so codegen
alone does not carry the arm. Add snapshot-status to the .NET, Go, Rust, and
Java renderers; the .NET and Go renderers were also missing provider-status,
which has been on the wire since the provider-mode work, so add it there too.
Java's renderer is an exhaustive switch expression and did not compile until
the new case landed. The Python CLI renders generic protobuf-JSON and needs
no change.
Each client README gains a paragraph on the feed-level frame next to its
existing from_truncated_snapshot paragraph: it arrives at stream open after
provider_status and before the cached active_alarm frames, then on every
verdict change including the clearing frame a monitor restart emits, so a
live consumer can track set completeness without polling QueryActiveAlarms.
GatewayDashboardDesign: list the two new payload cases the AlarmsHub forwards,
and — separately — record the GroupToRole / GroupToTag / UntaggedSessionVisibility
rows the settings page already renders but the bullet list omitted.
The form split tags with the shared ParseList and attached the result verbatim,
so "team-a, TEAM-A" persisted as two entries and the constraints column read
dashboard_tags=[team-a, TEAM-A] — one grant reported as two on the page whose
job is to show what a key was granted. Enforcement never saw it (a session holds
its tags in a case-insensitive set), which is exactly why the display was the
only place it could surface.
De-duplicates ordinal-ignore-case at the attach point only, first spelling
winning, matching ApiKeyAdminCommandLineParser.ParseDashboardTags. ParseList is
untouched: the five glob lists are matched literally, so near-duplicates there
are not necessarily the same rule and must survive verbatim — pinned by a test.
The help text claimed to mirror the CLI flag; it now claims only the shared
separators and the dedupe, since the form still drops an empty segment silently
where the CLI hard-fails. A browser form has no exit code to fail with, so that
difference stays, and Authorization.md now records it.
The constraints column enumerated only the eight positional ApiKeyConstraints
members, so a key whose sole recorded policy was a dashboard tag summarised to
an empty string and rendered as "-" — the same cell a key with no policy at
all gets. ApiKeyConstraints.IsEmpty counts DashboardTags, so that key is not
unconstrained, and the column was quietly telling operators otherwise about a
grant that decides who can watch a session's events.
The create form had no dashboard-tags input either, so tagged keys could only
be minted from the apikey create-key CLI. Adds the field beside the other
constraint lists (same ParseList separators) and attaches it through the
record's init-only member, since it postdates the eight-member constructor.
CreateModel, OpenCreateDialog and TryBuildCreateRequest widen to internal for
the new render tests: the create form is behind a click and static rendering
cannot dispatch one. That is the assembly's existing InternalsVisibleTo seam.
The per-session dashboard event ACL shipped in 693a78d + 7ec0b35 with unit
coverage over a fabricated principal. What a fabricated principal cannot show is
that the group names the shared directory actually returns -- short RDN values,
not DNs -- are the ones Dashboard:GroupToTag keys match. Two [LiveLdapFact]s
close that: gw-viewer binds for real, its GwReader membership grants team-a, and
IDashboardSessionAcl then admits a team-a-tagged session and refuses a
team-b-tagged one; multi-role takes the Administrator bypass. The mapping is
config-side only -- no GLAuth entry, group, or membership was added, and
glauth.md records that explicitly so a future reader does not go looking for a
directory change that never happened.
multi-role is a member of GwReader as well as GwAdmin, so it holds team-a too.
Its bypass is therefore asserted on team-b and on the untagged session -- the two
it would lose if the Administrator branch were ever dropped -- rather than on
team-a, which would pass either way.
One cheap hardening from a prior review: a GatewayOptionsTests case binds
Dashboard:GroupToTag through a real ConfigurationBuilder and looks the group up
mis-cased. The property initializer seeds an OrdinalIgnoreCase dictionary, but
only the binder decides whether that instance survives; if it did not, a
mis-cased group name from the directory would grant no tags and the ACL would
deny with no diagnostic.
Docs follow the shipped shape: docs/Sessions.md gains the session-tag model
(owner-key sourced, immutable, visibility-not-access), gateway.md and CLAUDE.md
gain the ACL in their dashboard-auth paragraphs, and three
GatewayDashboardDesign.md passages that still described the ACL as outstanding
now describe both gated seams and the decision order. GatewayConfiguration.md's
ShowTagValues row no longer claims the redaction is the only thing between a
Viewer and another session's values -- it is now the second of two independent
layers. gateway.md's hub-token lifetime corrected 30 minutes -> 5, matching
HubTokenService. Authentication.md disambiguates --dashboard-tags as the only
constraint flag that splits on commas. The plan doc header is Implemented; its
as-built section 12 already existed and is not duplicated.
Verified: NonWindows.slnx builds clean; GatewayOptions/DashboardSessionAcl/
EventsHub filters 37/37; the live-LDAP suite skips cleanly without the env var
and runs 7/7 green against the shared GLAuth with it.