Commit Graph

8 Commits

Author SHA1 Message Date
Joseph Doherty a107a1a8b9 feat(mtconnect): host registration + DriverTypeNames const + guard parity (Task 16)
Register the MTConnect factory + probe with the Host, add the
DriverTypeNames.MTConnect constant, and keep DriverTypeNamesGuardTests green.

- DriverTypeNames: new MTConnect const + appended to All.
- DriverFactoryBootstrap: MTConnectDriverFactoryExtensions.Register in
  Register(...) at the default Tier A (managed HttpClient/XML, in-process;
  Tier C is the only tier arming process-recycle, wrong here), and
  TryAddEnumerable(IDriverProbe -> MTConnectDriverProbe) in
  AddOtOpcUaDriverProbes so the probe reaches admin-only nodes (the
  admin-pinned Test-Connect singleton) without double-registering on a
  fused admin,driver node.
- Host.csproj: ProjectReference to Driver.MTConnect (required to compile
  the Register call).
- Core.Abstractions.Tests.csproj: ProjectReference to Driver.MTConnect so
  the guard's bin scan discovers the factory. This and the constant MUST
  land together -- verified RED (2/4 facts) with the constant alone.
- DriverProbeRegistrationTests: MTConnect added to AdminUiDriverTypeKeys;
  its idempotency fact asserts distinct-probe-count and goes RED without
  this (and RED if TryAddEnumerable is downgraded to AddSingleton).

Reconciles all four "MTConnect" literals onto the one constant: the
factory's DriverTypeName, MTConnectDriver.DriverType, and
MTConnectDriverProbe.DriverType now all alias DriverTypeNames.MTConnect
(no circular reference -- Core.Abstractions is a leaf both driver
projects already reference). This is the repo's TwinCat/Focas anti-drift
convention.
2026-07-27 13:58:06 -04:00
Joseph Doherty dfbcb0e7f0 fix(mtconnect): close the silent-refusal + teardown-wedge defects in the pump (Task 11 review)
C1 (critical): a latched "this endpoint cannot stream" verdict could end with a
GREEN driver holding a permanently silent subscription. DriverInstanceActor
handles a tag-set change as Unsubscribe-then-Subscribe with no re-initialize;
the unsubscribe cleared the stream-degraded flag, the resubscribe was refused by
the latch without a word, and one successful read then reported Healthy.
StartSampleStreamCore now re-asserts the degradation and logs a Warning naming
the remedy, and teardown no longer clears the flag while the latch is set.

NOT promoted to Faulted: DriverHealthReport maps any Faulted driver to a /readyz
503, so that would de-ready the whole node over one misconfigured Agent whose
reads are perfectly healthy. Faulted stays for the case that earns it.

I1: teardown waits are bounded (5 s) and honour the caller's token. They were
unbounded and ignored it, so DriverInstanceActor's 5 s budget was inert and
PostStop — which blocks a dispatcher thread on ShutdownAsync while holding the
lifecycle semaphore — could be wedged forever by one blocking subscriber. The
cancellation source is disposed only when the loop is provably gone. Applied to
StopProbeLoopAsync too: Task 13 reproduced the identical shape, and closing the
class in one method while leaving the other as the pattern to copy is worse than
not closing it.

I2: the reconnect ladder was monotonic for the process lifetime, so ~9 unrelated
drops pinned the driver at MaxBackoffMs forever. A stream that delivered before
dropping now starts a fresh ladder (at attempt 1, not 0, so MinBackoffMs is
still honoured).

I3: MaxBackoffMs <= 0 clamped every delay to zero — an operator-authorable
reconnect spin. Treated as unset, like a multiplier that cannot grow.

I4: the session published instanceId without NextSequence, so after a restart it
held the new agent's id beside the dead one's cursor, and a gap/OUT_OF_RANGE
re-baseline (id unchanged) never published the cursor at all. Both move in one
CAS, and the pump publishes its cursor as it exits.

I5: OnDataChange was a multicast Invoke — the first throwing handler aborted the
rest of the list, starving every later subscriber. Now walked by hand with the
catch inside the loop. The test that claimed to cover this registered the
recorder BEFORE the thrower, so it passed regardless; swapped, it was red.
Faults are latched to one Warning per stream generation (Debug thereafter) so a
consistently-throwing consumer cannot flood the log.

Also: the pump restarts after an unexpected fault (IsCompleted, not null); a
device-scope that matches nothing is a Warning, not a Debug tally; a device or
component with neither name nor id is skipped rather than emitting a blank path
segment; DiscoverAsync's remark no longer claims DriverInstanceActor retries
(it does not for a Once driver, and injection is dormant in v3); and two flake
seeds in the fake are gone (Dispose re-read its own counter; _disposeCts was
disposed under a racing SampleAsync).

439/439. Every changed behaviour falsified by mutation.
2026-07-27 13:57:46 -04:00
Joseph Doherty fa0e40796f feat(mtconnect): host-connectivity probe + instanceId rediscover (Task 13) 2026-07-27 13:57:46 -04:00
Joseph Doherty 13dd9fcd70 feat(mtconnect): ITagDiscovery streams tree + SupportsOnlineDiscovery opt-in (Task 12) 2026-07-27 13:57:46 -04:00
Joseph Doherty 9a67ed918a feat(mtconnect): ISubscribable /sample pump + ring-buffer re-baseline (Task 11)
One shared /sample long poll behind every subscription handle. Subscribe fires
initial data per reference from the primed /current (Part 4 convention) and
returns without waiting on the stream; the pump runs on its own task under its
own token — never the caller's Subscribe token, which is disposed the moment
that call returns.

The pump owns every way the sequence can break:

- sequence gap, measured against the RUNNING cursor (the previous chunk's
  nextSequence). Comparing against the opening `from` reports a gap on every
  chunk once the buffer rolls — a /current re-baseline storm.
- OUT_OF_RANGE under HTTP 200 (InvalidDataException out of SampleAsync), which
  IsSequenceGap cannot see at all — without it a driver recovers from falling a
  little behind but not from falling a lot.
- agent restart, checked BEFORE the gap because a restart resets sequences and
  so usually trips the gap check too; it clears the index (the held values
  describe a device model that no longer exists) and tells the subscribers.
- every chunk advances the cursor, heartbeats included.

Re-baseline calls CurrentAsync directly on the client the pump already holds:
routing it through ReinitializeAsync would take the non-reentrant lifecycle
semaphore from inside a pump a lifecycle method may already be awaiting, and
hang the driver with no exception and no log. StopSampleStreamAsync is filled
in (same call sites) and InitializeAsync now stops the stream before teardown
too, so a re-Initialize on a live instance cannot leave a pump enumerating a
disposed client.

MTConnectStreamEnded/TimeoutException reconnect under a geometric backoff with
a 100 ms growth floor (MinBackoffMs defaults to 0, and 0 x multiplier is still
0). MTConnectStreamNotSupportedException does NOT retry — it latches, reports
loudly, and is cleared only by a re-initialize.

Health precedence is explicit: /current and /sample fail independently, so each
path owns a flag, clears only its own, and Healthy requires both clear. Neither
can paper over the other's failure.

29 new tests, all deterministic — no sleeps, no wall-clock assertions. The fake
grew per-generation sample streams (so a reconnect has somewhere to land),
scripted /current answers, scripted sample failures, and a chunk-consumed
signal that also fires when the consumer abandons the stream. 375/375 green.
2026-07-27 13:57:46 -04:00
Joseph Doherty a127954cab fix(mtconnect): review remediation for the driver lifecycle (Task 9)
Important 1 - the probe-model triple was read outside _lifecycle while every
write happened under it. Packed (Client, ProbeModel, ProbeModelDataItemCount)
into one immutable AgentSession swapped by a single Volatile.Write, matching
what _index already does, rather than locking the readers: both await the
network, so taking _lifecycle would stall a shutdown for a whole
RequestTimeoutMs and would widen the non-reentrancy hazard. GetOrFetch and
Flush re-publish via CompareExchange so a late re-cache cannot undo a
concurrent teardown/re-init, and a client disposed mid-fetch now surfaces as
the same "not connected" InvalidOperationException browse already handles
instead of a raw ObjectDisposedException.

Important 2 - HasConfigBody decided emptiness by matching the literals "{}" /
"[]", so "{ }", "{\n}" and every pretty-printed empty document fell through to
ParseOptions, failed the required-AgentUri check, and turned a semantically
empty config into a cold-start fault. Now parsed: an object/array with no
elements (or a bare null) is empty; malformed text is deliberately NOT empty so
ParseOptions produces the real quoted error rather than silently starting on
stale options.

Important 3 - documented the _lifecycle non-reentrancy hazard in the code, on
both the semaphore and StopSampleStreamAsync, naming Task 11's re-baseline as
the specific path that would deadlock and giving the CurrentAsync-directly
pattern that avoids it.

Important 4 - removed the Initialize/Reinitialize asymmetry rather than
documenting it. Both now share one rule via ResolveIncomingOptions: an
unreadable config document never destroys working state, and faults only a
driver that had none. InitializeAsync previously tore down before parsing, so a
bad document on a live instance destroyed a healthy client - the exact outcome
ReinitializeAsync was written to avoid.

Minors: SafeDispose traces the swallowed disposal fault at Debug (a client that
cannot be released is how a handle leak starts); the RequirePositive test is a
Theory over all four timing knobs, not just RequestTimeoutMs (arch-review
01/S-6); the footprint test no longer overclaims "is zero before initialize".

CannedAgentClient gains a one-shot ProbeGate so a lifecycle change landing
mid-request is deterministic - no timers. 346/346 (329 + 17). Six mutations
verified: fetch-CAS, disposed-translation, flush retired-session guard,
literal HasConfigBody, teardown-before-parse, and two dropped RequirePositive
calls each fail only their own tests.
2026-07-27 13:57:46 -04:00
Joseph Doherty 1bfc1722b3 feat(mtconnect): IReadable via /current, ordered per-ref snapshots (Task 10)
One deadline-bounded /current per batch, answered out of that one whole-device
document: N references cost one round-trip, one snapshot each, in request order,
duplicates included.

DEVIATION (deliberate, from the plan's implied shared-index write): the response
is indexed into a PER-CALL snapshot index, never into the driver's shared
observation index. That index is written by Task 11's /sample pump and is
last-write-wins with no timestamp arbitration; a read folding its /current into
it would be a second writer racing the first, and /current is frequently the
OLDER document (the Agent composes it while the pump's newer delta is in
flight). Losing that race rolls a subscribed value backwards and leaves it there
until the data item next changes. Both paths coerce through the same
MTConnectObservationIndex logic, so a read and a subscription report a given
observation identically -- the read path is simply a pure function of
(document, authored tags).

Failure posture is the inverse of InitializeAsync: nothing but genuine caller
cancellation crosses the capability boundary. Unreachable Agent / blown deadline
/ unusable answer => every ref codes BadCommunicationError and the driver
degrades (never downgrading Faulted); no client at all (pre-initialize, faulted,
post-shutdown) => BadNotConnected with health untouched. The index's
BadWaitingForInitialData / BadNodeIdUnknown / BadNoCommunication / BadNotSupported
distinctions pass through unmodified.

21 tests pinning arity, order, duplicates, empty batch, one-round-trip, each Bad
code, read-time freshness, the anti-clobber invariant, the deadline, and
cancellation-propagates-but-agent-failure-does-not. CannedAgentClient gains a
cancellation-observing CurrentGate (still no timers or sleeps of its own).
Test csproj takes the AbCip/S7 OTOPCUA0001 NoWarn -- this suite calls the
capability directly by design.

329/329 green.
2026-07-27 13:57:46 -04:00
Joseph Doherty 4fe0af3693 feat(mtconnect): driver shell IDriver lifecycle over agent-client seam (Task 9)
MTConnectDriver implements IDriver only: connection-free ctor, initialize
(deadline-bounded /probe + /current, caching the device model, the agent
instanceId, and a primed observation index), reinitialize, shutdown, health,
footprint, and cache flush.

Decisions worth knowing:

- The driverConfigJson argument is HONOURED, not ignored. DriverInstanceActor
  delivers a config change only via ReinitializeAsync(json) on the live
  instance, so a driver reading only its ctor options would turn every MTConnect
  config edit into a silent no-op that still sealed the deployment green.
  ParseOptions lives on the driver; Task 15's factory must delegate to it rather
  than deserialize a second DTO.
- /probe ok but priming /current failing is Faulted, not Healthy-with-no-data:
  the index IS the read surface and the /sample cursor comes from /current.
- ReinitializeAsync rebuilds the client whenever ANY option the client reads at
  construction changed - AgentUri, DeviceName, RequestTimeoutMs, HeartbeatMs,
  SampleIntervalMs, SampleCount - not just the first two. The timing knobs are
  baked into the deadlines, the watchdog window, and the /sample query string.
- An unparseable config faults a cold start but leaves a RUNNING driver serving
  its previous valid configuration (the deployment fails instead).
- IMTConnectAgentClient now extends IDisposable rather than the driver doing
  `as IDisposable`, which would silently no-op against a non-disposable fake and
  leak an HttpClient + socket handler per re-deploy. Sync, not async: every
  resource behind the seam is sync-disposable, and the async unwind belongs to
  the SampleAsync enumerator the pump owns.
- Flush drops the probe/browse cache and keeps the index; every model consumer
  goes through GetOrFetchProbeModelAsync so a flush cannot wedge browse.

Tests: CannedAgentClient (shared, deterministic - channel-backed scripted
/sample with no timers, call counts, disposal tracking, failure injection) +
27 lifecycle tests. 308/308 in the project.
2026-07-27 13:57:46 -04:00