The truncation-cliff fix made alarm transitions truncation-safe but silent:
when GetXmlCurrentAlarms2 returns exactly maxAlmCnt records the worker
suppresses absence-implies-Clear inference and says so only in a rate-limited
stderr warning. No client and no operator could tell a complete active set
from a capped one.
Two additive proto3 booleans carry the verdict out:
- QueryActiveAlarmsReplyPayload.snapshot_truncated = 2 (worker IPC reply)
- ActiveAlarmSnapshot.from_truncated_snapshot = 16 (per record)
The per-record field is not an aesthetic choice. QueryActiveAlarms returns a
bare `stream ActiveAlarmSnapshot` with no envelope, header, or trailer, so a
per-record boolean is the only carrier that stays wire-compatible; an envelope
message would change every existing client's stream element type. The reply
payload states it too because a prefix filter can leave zero records and a
truncated fetch with nothing to report still has to say so. The flag means
"this set may be incomplete", never "this record is unreliable" — it is
independent of the subtag-fallback `degraded` field.
Detection is deliberately UNCHANGED: IsTruncatedFetch remains
`fetchedRecordCount >= maxAlarmsPerFetch`. The live probe (docs/AlarmProbeFindings.md,
ce5d8ae) could not verify whether ALARM_RECORDS/@COUNT reports the total active
count or only the records in the reply, so @COUNT is not parsed for detection;
switching to it stays blocked on probe evidence. The probe's comment
annotations in WnWrapAlarmConsumer.cs are preserved.
Reset semantics: not latched. WnWrapAlarmConsumer.FoldFetch replaces the
verdict on every poll under the same lock as the snapshot merge, so the first
sub-cap fetch clears it; GatewayAlarmMonitor.ClearCache drops it with the cache
generation it describes. A caveat that never turns off is one operators learn
to ignore.
Flow: WnWrapAlarmConsumer.LastSnapshotTruncated -> AlarmDispatcher (stamps every
record) / IAlarmCommandHandler (payload) -> MxAccessCommandExecutor reply ->
GatewayAlarmMonitor._snapshotTruncated -> IGatewayAlarmService.SnapshotTruncated
-> DashboardAlarmQueryResult -> AlarmsPage warning banner (render-side only; the
poll loop and DisposeAsync drain are untouched). The public QueryActiveAlarms
RPC forwards worker snapshots unmodified, so the per-record flag needed no
mapper change — a test pins that.
Parity: this describes OUR fetch mechanics — additive gateway metadata — not
MXAccess provider behavior. No event is synthesized and no MXAccess-observable
semantics change, so it is not a parity deviation.
Tests: worker LastSnapshotTruncated set/reset/consecutive-burst (windev-run);
gateway end-to-end truncated reply -> monitor -> public stream, with the
complete-reply control as the load-bearing assertion; AlarmsPage banner
present/absent. Docs: gateway.md alarm surface, docs/DesignDecisions.md entry.
Both tasks the detached-lock-wait path starts and discards now carry the
file's fault-observing continuation, factored out of ObserveAbandonedFault as
ObserveFault so the idiom has one definition. DrainDetachedAsync swallows the
drain but its _writeLock.Release() sits in a finally outside that catch, so a
Release that ever throws — a SemaphoreFullException from some future
double-release regression — had no awaiter and would have surfaced on net48 as
TaskScheduler.UnobservedTaskException at finalization instead of an
attributable failure. Same for a throw out of OnDetachedLockWaitSettled.
Review's Minor (distinguishing a cancelled from a faulted lock wait before
tombstoning) is deliberately not taken: a faulted WaitAsync is unreachable
here — nothing disposes _writeLock — so the branch would be untestable new
logic whose only effect is internal state, the caller already receiving the
fault itself from the rethrow. Recorded as a comment at the site instead.
edited on macOS, windev verification pending (plan Task 11). Re-ran the net10
scratch harness over WorkerFrameWriter and the writer suite: 0 warnings,
31/31 pass.
Both questions stay open, and the reason is the finding: the dev rig's alarm
UDAs reject a plain MXAccess Write with SecurityError/detail=1008 from the
responding automation object, so no alarm instance can be created to follow
through an acknowledge and no population can be built to overflow a capped
fetch. The rig is otherwise live — objects deployed and on scan, wnwrap
subscribed, GetXmlCurrentAlarms2 returning well-formed XML — which is what
makes the blocker specific and the unblock (engine-side script, or
AuthenticateUser + WriteSecured, or reclassifying the UDAs) actionable.
Comment-only changes in WnWrapAlarmConsumer: scope the GUID-identity claim to
the leg live capture actually covers, and record that ALARM_RECORDS/@COUNT
exists as a candidate exact truncation signal but is deliberately not trusted
because its semantics under a capped reply are unverified. No behavior change.
Adds a dashboard event-visibility tag to ApiKeyConstraints, riding in the existing
constraints JSON blob so no auth-store schema migration is needed (design
docs/plans/2026-07-10-dashboard-session-acl-tst15.md sections 3/3.1, open call
settled per its own recommendation). The tag is visibility-only: no read, write,
browse, or subscribe path consults it, and HasRead/HasWriteConstraints ignore it.
GatewaySession gains an immutable, ordinal-ignore-case Tags set stamped at
construction from the owning API key, forwarded by MxAccessGatewayService.OpenSession
from the resolved ApiKeyIdentity — never from the wire request, so a client cannot
label its own session with another tenant's tag. ISessionManager gains a tag-carrying
OpenSessionAsync overload whose default implementation forwards to the tagless one, so
an implementation that does not model tags opens an untagged (least visible) session.
apikey create-key gains --dashboard-tags team-a,team-b (repeatable, trimmed,
de-duplicated; an empty segment is rejected rather than dropped) and list-keys prints
the tags column. No enforcement yet — the EventsHub ACL that consumes the tag is a
later change.
WriteAsync enqueued its frame and then contended unconditionally for the write
lock, so a caller that lost the race stayed in WaitAsync until the winning
drainer released — even though that winner writes, flushes, and completes the
loser's control frame at the control-to-event class boundary, part-way through
its pass. The boundary flush made the delivery point honest; the awaited task
was still charged for the whole event backlog it had just been flushed ahead of.
WriteAsync now awaits its own frame's completion racing the lock acquisition.
Completion first: the caller returns at its frame's delivery point and the
outstanding acquisition is detached, not dropped — a continuation drains
whatever is queued and releases, so the lock is never acquired and silently
held and a frame enqueued between the previous drainer's last dequeue and its
release is still written. Lock first: drain as before. Cancellation keeps the
WRK-22 tombstone semantics exactly, and a wait cancelled after the caller has
already detached releases nothing (SemaphoreSlim hands no count to a wait it
cancels), so no count leaks and no queued frame is stranded. A token that fires
after the frame's completion won the race changes nothing — the frame was
delivered. WriteBatchAsync deliberately keeps the plain wait-then-drain shape:
its last completion resolves at the end-of-pass flush anyway.
Three tests: the latency win (a control caller returning while the winning
WriteBatchAsync event burst is demonstrably still blocked mid-pass), a
mixed-priority concurrency soak pinning exactly-once writes and a single
drainer, and the cancel-after-detach corner (a wrongly released count would
surface as the drainer's own Release throwing SemaphoreFullException).
edited on macOS, windev verification pending (plan Task 11). Verified here by
compiling and running WorkerFrameWriter plus the writer suite against net10.0
in a scratch harness: 31/31 pass, and the two behaviour-pinning tests fail
against the pre-change parked implementation.
Groundwork for the per-session dashboard event ACL (docs/plans/2026-07-10-dashboard-session-acl-tst15.md 3.2): a dashboard group can now grant visibility tags, and untagged sessions default to AdminOnly. Enforcement lands with the EventsHub ACL; nothing consumes the grant yet.
GroupToTag is deliberately uncoupled from GroupToRole - a group may appear in either map, both, or neither - and is validated for shape only. Tags gate dashboard event visibility, never data access.
ISessionManager.ReadEventsAsync had zero production call sites: the worker
event channel is drained once by GatewaySession.MapWorkerEventsAsync (the
distributor pump), and every consumer — gRPC subscribers, the dashboard
mirror, the alarm monitor — attaches to the distributor. The interface
member, SessionManager's forwarder, and GatewaySession.ReadEventsAsync are
gone; IWorkerClient/WorkerClient.ReadEventsAsync is untouched, it is the
live worker-channel claim.
No test was removed or rewired: nothing invoked the member through the
interface. Nine ISessionManager test fakes carried a required-member stub
(seven threw NotSupportedException or yielded nothing; EventStreamServiceTests
and GatewaySessionDashboardMirrorTests forwarded to the session; the two
MxAccessGatewayService fakes yielded their Events list) — all nine stubs were
deleted. The MxAccessGatewayService suites' streaming tests already run
through FakeEventStreamService, which reads the same Events list, so their
coverage is unchanged; only the now-inaccurate doc comments on Events /
LastReadEventsSessionId were reworded.
The MapWorkerEventsAsync comment no longer describes a twin to keep in step;
it now states the single-reader claim directly. docs/Sessions.md drops
ReadEventsAsync from the SessionManager member list and from the Run-state
prose. The 2026-08-15 deferred-remediation as-built note records the removal.
The two-class writer already got control bytes out ahead of a queued event
backlog, but a frame counts as delivered only once flushed, and the drain
deferred its single FlushAsync — and every TrySetResult — to the end of the
pass. A heartbeat, command reply, fault, or shutdown ack was therefore written
first and completed last, behind up to a full 128-frame event batch.
The drain now records each frame's priority class on PendingFrame and flushes
at every control-to-event boundary, completing and clearing the written set
there. Cost stays bounded: a pure-event pass still pays exactly one flush, a
run of control frames still pays one for the run, and only a pass that mixes
both classes pays a second — never one flush per control frame, the
syscall-per-heartbeat cost WRK-12 removed.
A boundary flush that itself fails is a new failure window and is handled like
the end-of-pass flush failure, additionally failing the event frame the drain
had already claimed off its queue and every frame still queued. Frames a
boundary flush completed leave the written set, so a later failure in the same
pass can no longer reach back and fail an already-delivered control frame.
The awaited task of a caller that lost the write-lock race is still bounded by
the winning drainer's pass — that enqueue-then-contend parking is unchanged and
now documented on WriteAsync and in docs/WorkerFrameProtocol.md.
MxAccessValueCache.Set deep-copied the Value (recursive for an MxArray), the
SourceTimestamp, and the Statuses RepeatedField (container plus every
MxStatusProxy) on every OnDataChange. The aliasing audit found all three
removable: the sink enqueues the event first — which stamps
WorkerSequence/WorkerTimestamp inside the queue lock — and only then runs the
postPublish hook that reaches Set, so the event is write-once by then and the
queue's ownership invariant forbids later mutation. The producer never reuses
instances (fresh MxEvent per mapper call, fresh MxValue per convert), and the
alias already existed on the read side: MxAccessSession.SucceededRead puts the
cache's own Value/SourceTimestamp/status references on every BulkReadResult,
which the worker only serializes onto the IPC pipe.
Set and CachedValue now carry the ownership contract: the cache holds borrowed
references into an enqueued, write-once MxEvent; consumers may read and
serialize, never mutate. Mutation would corrupt the still-queued event AND
invalidate QueuedEvent.Size — the enqueue-time memoized serialized size the
byte-budgeted Drain charges — so a grown message could overshoot the negotiated
frame max and fault the session with MessageTooLarge. MxAccessEventQueue's
class remark, which claimed the cache keeps an independent snapshot, is
corrected to point at the borrow.
MxAccessWriteCompletionCache.Record keeps its parallel statuses.Clone()
deliberately, with a cross-reference explaining why: it takes a bare
RepeatedField whose provenance its signature cannot constrain, and it is on the
command-rate write path, not the streaming hot path.
Tests: Set_StoresIndependentSnapshot_UnaffectedByLaterEventMutation codified
the invariant being reversed, so it is replaced by
Set_BorrowsTheEventsOwnInstances_ByOwnershipContract (Assert.Same on Value,
SourceTimestamp, and the status row). Adds the missing cached-read-path test to
MxAccessCommandExecutorTests — nothing in the worker exercised was_cached ==
true end to end — asserting the cache hit, reference identity out to the
BulkReadResult, and that no COM call is made for the read.
Not built or tested here: these are net48/x86 worker files that cannot compile
on the macOS tree. Verification is deferred to the windev gate.
The worker-side half of the review tail. Tests and comments only — nothing here
changes worker behavior, and none of it compiles on the macOS tree (net48/x86),
so it was reviewed line by line against the already-windev-validated files.
- MxAccessHandleRegistryTests gains the multi-candidate case behind
MxAccessSession.TryGetCachedReadFor's fall-through: one tag under two item
handles, the lower registered-but-unadvised and the higher advised. Asserted
at the registry rather than the session because the session's read path needs
a live MXAccess COM instance; what the registry owes the scan is the stable
ascending candidate order and a per-item-handle (not per-tag) advice index,
and both are pinned here along with the fall-through contract in prose.
- A single adversarial lifecycle test — register, advise, re-register the same
item handle under a new tag, unadvise, unregister the server — asserting every
index agrees after each step. The individual transitions were already covered;
what was not was that they compose, and a stale entry in any one index
resurrects a handle MXAccess has already retired.
- StaWaitHelperTests.WaitForSignalOrMessages_PreSignalledHandle_ReturnsImmediately
drains pending messages first, like the other two wait tests. Without it a
stale message can end the wait instead of the handle, failing the
signal-consumed post-condition for an unrelated reason.
- GatewayTesting.md records the two findings from the Task 24 windev gate:
SecretsStorePathGuardTests.CreateBuilder_AcceptsSecretsStoreOutsideContentRoot_AndCreatesIt
fails deterministically on Windows on main too (SQLite pooling holds secrets.db
open across the cleanup's recursive delete; pre-existing, tracked separately),
and the StaWaitHelper timing tests' flake signature on a loaded box is a
message wake — the helper working as designed — not a broken wait.
The remediation reviews approved every task but left a tail of small notes.
This lands the gateway-side half of them.
Hardening (behavior changes, all narrow):
- BuildFilteredWriteBulkCommand's unreachable default case failed OPEN: a
fifth bulk-write kind added upstream without a filter case here would have
shipped the DENIED entries to the worker while reporting them denied to the
caller. It now throws UnreachableException.
- SqliteCanonicalAuditStore.ListRecentAsync no longer throws on a row it
cannot date. The retention sweep deliberately preserves such rows (SQLite's
datetime() yields NULL, so the DELETE never matches), which guaranteed the
dashboard's recent-audit view would meet one eventually and lose the whole
page to it. The row is now reported at DateTimeOffset.MinValue with every
other column intact, behind an optional logger.
- The audit drain loop's finally now completes the channel writer alongside
detaching the drain, so a producer that raced past the attached check takes
the write-through branch instead of stranding its event in a buffer nobody
reads until shutdown. TryComplete is idempotent, so StopAsync is unaffected.
Tests:
- MapCommandReply ownership (Assert.Same on the inner reply), mirroring the
existing MapEvent ownership test.
- Redactor key-id length boundary at exactly 64 and 65 characters, pinning
which way it fails. Nothing validates key-id length at creation, so
docs/Diagnostics.md's "which no issued key id does" is now stated as the
heuristic it is.
- ApiKeyFailureLimiter.Reset with a PartitionResolution whose partition was
evicted between the Check and the Reset: inert, and clears nobody else's
block.
- Constraint-cache concurrency stress: the cap is enforced by the inserting
thread, so overshoot must be transient and proportional to the in-flight
inserters, and the cache must settle at or under the cap.
- ListRecentAsync against a raw-SQL undateable row.
Comment/doc accuracy:
- EventsHubViewerRegistry.ReleaseConnection records that it relies on
SignalR's default sequential per-connection dispatch
(MaximumParallelInvocationsPerClient = 1).
- A PERF(followup) note on Invoke's double session resolve and why removing it
needs a SessionManager overload.
- SessionEventDistributor: the volatile-field comment named the pump as the
lock-free reader, but the pump's single capture point is inside _replayLock;
the genuinely lock-free reader is SubscriberCount. OnSubscriberOverflow's
"cannot be observed here" now excepts the DisposeAsync abandon path. The
churn test names its ConcurrentDictionary bucket-order assumption and that a
violation surfaces as a read timeout, not a silent pass.
- The two "restores the sequential drain's behavior" claims (SessionManager,
docs/Sessions.md) were wrong: the sequential drain leaked too, because
KillWorkerAsync's entry ThrowIfCancellationRequested aborted the whole loop
on the first session for zero kills. Reworded to "fixes a leak the
sequential drain also had", with the sweep-bound/shutdown-unbound
ParallelOptions asymmetry explained.
- ISessionManager.ShutdownAsync's token doc: it degrades the drain to a kill
sweep rather than cancelling it, with the bounded overrun stated.
SessionShutdownHostedService.StopAsync records that its cancellation-logging
branch is now unreachable.
Decrement was decrement-first with a single non-retried repair CAS. From zero,
two unmatched decrements (SignalR calls OnDisconnectedAsync for a connection
whose OnConnectedAsync faulted) capture -1 and -2; a real Increment then makes
the count -1, and the first decrementer's stale CompareExchange(0, -1) matches
and resets to zero — erasing a live connection, so the idle gate freezes an open
dashboard. The same lost race also made Decrement report 0 when it had not
written 0.
Clamping now happens inside the compare-and-swap: read, clamp, publish, retry on
loss. A lost race re-reads the fresh value instead of repairing a stale one.
The counter moves to its own file per the one-public-type-per-file convention and
gains direct tests: the zero floor under concurrent unmatched decrements, matched
pairs settling at zero, and an interleaved connect/disconnect stress round. The
stress test asserts the observable invariants only — the specific interleaving
cannot be forced through the public API (verified: the previous implementation
passes it), which its remarks now state rather than implying a reproducer. A hub
wiring test is skipped for the EventsHub reason: driving Hub.OnConnectedAsync
needs caller-clients and connection-context fakes, and the overrides are two
lines of delegation to the tested type.
Also documents that the API-key refresh's pre-gate time check races benignly.