Compare commits

..

119 Commits

Author SHA1 Message Date
Joseph Doherty 813e445936 test(mtconnect): close live-gate leg 5 — deploy -> OPC UA read/subscribe PASSED
v2-ci / build (pull_request) Successful in 6m48s
v2-ci / unit-tests (pull_request) Failing after 18m6s
Re-ran the gate on the rebased branch against a freshly rebuilt
otopcua-host:mtconnect image, so it exercises the merged code.

The 2026-07-24 blocker (antiforgery cookie collision between the running
otopcua-dev rig on :9200 and the isolated stack on :9220) is sidestepped
rather than worked around: the headless deploy API
(POST /api/deployments + X-Api-Key) is AllowAnonymous().DisableAntiforgery()
by design, so cookie scoping is structurally irrelevant to it. No browser
needed — the recommended recipe for any future isolated-stack gate.

Both deploys sealed (DeploymentStatus.Sealed, FailureReason NULL) with both
MAIN nodes acking. Read + subscribe verified over opc.tcp://localhost:4860
across all three type paths: Float64 SAMPLE (sinusoid), Int64 EVENT
(monotonic), String EVENT (READY -> INTERRUPTED). This closes the last
unproven link — driver seam was already covered by the live integration
suite; leg 5 adds DeploymentArtifact -> address space -> OPC UA publish.

Two findings recorded in the plan:

- A bad authored dataType degrades to exactly ONE skipped tag, loudly
  (per-tag warning naming the expected keys, BadNodeIdUnknown, "every other
  tag is unaffected") while the rest keep streaming — the intended
  fail-isolated behaviour, demonstrated rather than argued. Correcting it
  took effect with no restart, showing MTConnect is not subject to the
  separately-tracked config-edits-discarded defect.

- fixture_asset_changed reading 0x80000000 is NOT an MTConnect defect. The
  driver maps the agent's UNAVAILABLE to BadNoCommunication correctly; the
  shared publish path collapses every status to a 3-state OpcUaQuality enum
  (DriverInstanceActor.QualityFromStatus, statusCode >> 30) and re-expands
  Bad to StatusCodes.Bad (OtOpcUaNodeManager.StatusFromQuality). Deliberate
  per the enum's own doc comment, but the client-visible consequence appears
  undocumented: no driver's Bad/Uncertain sub-code ever reaches a client.
  Tracked separately; out of scope here.

Also records the rebase onto master 123ddc3f and confirms StatusCodeParityTests
now covers 9 MTConnect constants, all passing.
2026-07-27 13:59:26 -04:00
Joseph Doherty 9d430bee62 docs(mtconnect): correct the MaintenanceMode note; record status-code parity pre-verification 2026-07-27 13:59:26 -04:00
Joseph Doherty 1d0989d3b4 test(mtconnect): live gate on an isolated rig — 4/5 legs pass, deploy leg blocked on a cookie collision (Task 21) 2026-07-27 13:59:26 -04:00
Joseph Doherty 71ccccef7c docs(mtconnect): P1 driver guide + tracking-doc update, defer write-back (Task 22)
- docs/drivers/MTConnect.md: getting-started guide for the P1 Agent MVP —
  config keys read off MTConnectDriverOptions/ConfigDto, capability surface,
  RawPath->dataItemId coercion precedence, UNAVAILABLE->BadNoCommunication +
  the rest of the quality-code table, CONDITION-as-String, browse-via-the-
  universal-browser, the fixture recipe (links the existing IntegrationTests
  Docker README instead of duplicating it), and a Known limitations section
  covering the two pre-existing fleet-wide gaps this build surfaced
  (IRediscoverable/IHostConnectivityProbe have no server consumer; no driver
  form blocks Save on validation). Records write-back (MTConnect Interfaces),
  P1.5 (native alarms, TIME_SERIES arrays, EVENT->enum), and P2 (SHDR ingest)
  as deferred per design doc SS3.6/SS9. Documents where the build diverged
  from the design: TrakHound MTConnect.NET was dropped entirely (no XML
  formatter in the pinned packages, no injectable HTTP handler) in favor of
  a hand-rolled System.Xml.Linq client, namespace-agnostic on LocalName.
- docs/drivers/README.md: adds MTConnect to the ground-truth driver table,
  the per-driver doc list, and the fixture coverage-map list.
- docs/plans/2026-07-24-driver-expansion-tracking.md: marks the MTConnect
  Wave-2 row Done (22/23 tasks; Task 21 live gate tracked separately in the
  plan file, not touched here), corrects mtconnect/cppagent -> mtconnect/agent
  in the fixture note, and links the new driver guide.

Deferred-cleanup decision (Task 1 review follow-up): removed
GenerateDocumentationFile/NoWarn CS1591 from Driver.MTConnect.Contracts's
csproj to match all eight sibling .Contracts projects, none of which carry
it. Verified both the Contracts project alone and the full solution still
build clean (0 warnings, 0 errors) with the flags gone — the existing XML
doc comments don't depend on GenerateDocumentationFile to compile.
2026-07-27 13:59:26 -04:00
Joseph Doherty cf9ce91e51 test(mtconnect): live-Agent docker fixture + integration suite (Tasks 19+20)
The first and only thing in the MTConnect workstream that exercises the driver
against a real Agent; everything else runs on canned XML.

Fixture (Docker/): two services, both stock images with this folder bind-mounted.
- `agent`   mtconnect/agent:2.7.0.12 on :5000. NOTE the repo name: the source
            project is mtconnect/cppagent but the PUBLISHED image is
            mtconnect/agent; mtconnect/cppagent does not exist on Docker Hub.
- `adapter` a stdlib SHDR feeder. Required, not optional: the agent image ships
            only the binary + schemas/styles, so a lone Agent answers /probe and
            then reports every observation UNAVAILABLE forever, which can prove
            nothing about reads or streaming.

Devices.xml seeds one DataItem per inference branch (SAMPLE+units, TIME_SERIES,
PART_COUNT/LINE_NUMBER, controlled vocab, DATA_SET, CONDITION), gives every named
item `name != id` (the inverse of the hand-authored unit fixtures), and leaves one
EVENT permanently unfed so UNAVAILABLE -> BadNoCommunication has a live subject.

Suite (12 tests) asserts STRUCTURALLY against the Agent's own /probe response, never
by literal id, so it survives an edit to Devices.xml and can be pointed at a real
machine tool via MTCONNECT_AGENT_ENDPOINT. It skips cleanly (12/12) when no Agent
answers, and each test carries a hard [Fact(Timeout)] so a half-up fixture cannot
wedge a build.

Three findings from bringing the fixture up, each now pinned in a comment:
- agent.cfg must be pure ASCII; one non-ASCII byte in a COMMENT makes the config
  parser reject the whole file with a bare "Failed / Stopped at line: N" and exit.
- `sampleCount` is an attribute of a TIME_SERIES observation, not of a DataItem
  declaration. A real Agent meeting one on a declaration DROPS THE ENTIRE DATA ITEM
  from the device model, so the inference only ever sees null and a live TIME_SERIES
  tag is always variable-length. Asserted, not ignored.
- The Agent's own <Agent> self-model publishes update-rate SAMPLEs that tick whether
  or not any adapter is attached. Excluding them is load-bearing: with them included
  the "live stream delivered a changed value" test passed with the fixture's data
  source deliberately stopped.

MTConnectError-under-HTTP-200 is NOT covered live and cannot be: cppagent 2.7 answers
an unknown device 404 and an out-of-range sequence 400, each with a well-formed error
body. That shape stays canned-only, and the suite says so.
2026-07-27 13:58:54 -04:00
Joseph Doherty 9227d03ca8 feat(mtconnect): make an MTConnect driver creatable + configurable in /raw (Task 18 Part B)
An MTConnect driver instance could not be authored at all: RawDriverTypeDialog
(the only DriverInstance creation surface) offered a hardcoded 9-entry list
without it, and DriverConfigModal's default: branch renders a warning with no
raw-JSON fallback — so even a hand-inserted row could not be configured and
DriverConfig stayed "{}", on which ParseOptions throws "missing required
AgentUri".

Adds the typed MTConnectDriverForm (the consistent sibling pattern) covering
agentUri / deviceName / the four polling knobs / probe / reconnect, wires the
dispatch case, and offers the type in the dialog.

- All decisions live in the static MTConnectDriverForm.BuildConfigJson +
  MTConnectFormModel (plain C#), because this project has no bUnit.
- The wire shape mirrors the driver's private MTConnectDriverConfigDto, NOT
  MTConnectDriverOptions: the latter's TimeSpan probe knobs would serialise as
  "00:00:05" and never bind to the driver's integer intervalMs.
- arch-review 01/S-6: a non-positive timing knob is REMOVED from the document
  rather than written, so ParseOptions substitutes the driver's own positive
  default instead of the driver faulting at Initialize on RequirePositive.
  Removal (not just skipping) also stops a stale positive value surviving and
  disagreeing with what the form shows.
- Unknown top-level keys survive a load/save, so the driver's hand-authored
  tags[] surface is not silently deleted by opening the form.
- _jsonOpts carries JsonStringEnumConverter, satisfying the
  DriverPageJsonConverterTests fleet guard properly rather than by exemption.

New RawDriverTypeDialogCoverageTests resolves the offered set through
DriverTypeNames.All, so the next registered driver fails here until the dialog
offers it.
2026-07-27 13:58:32 -04:00
Joseph Doherty 115a43dd47 feat(mtconnect): typed AdminUI tag-config editor + browse round-trip (Task 18 Part A)
Registers the MTConnect typed tag editor in TagConfigEditorMap /
TagConfigValidator (keyed off DriverTypeNames.MTConnect, never a literal) and
adds the razor shell over Task 17's MTConnectTagConfigModel.

The shell stays trivial: the /probe-sourced mtCategory/mtType/mtSubType/units
row is rendered disabled + readonly with NO binding of any kind, and no
inferred-type hint is shown — MTConnectDataTypeInference.Infer() also keys on
the DataItem `representation`, which the tag-config model does not carry, so
calling it here would agree with the browse picker for ordinary items and
silently disagree for exactly the DATA_SET/TABLE/TIME_SERIES ones.

Also fixes the picker -> editor round trip, which was broken in both
directions:

- RawBrowseCommitMapper.BuildTagConfig had no MTConnect branch, so a committed
  tag landed as the generic {"address": "<dataItemId>"} while the model read
  only `fullName`. The data plane kept working (the driver accepts all three id
  spellings), so the only symptom was an empty DataItem-id field in the editor
  — and saving from that state stamped fullName:"" over the tag.
- MTConnectTagConfigModel.FromJson now accepts dataItemId/address as fallbacks
  (fixing tags committed before this change) and ToJson normalises onto
  `fullName`, dropping the aliases so the two spellings cannot drift.
2026-07-27 13:58:19 -04:00
Joseph Doherty 51c8c7f30b fix(mtconnect): bind RawTags and key the data plane by RawPath
The driver could not receive data after a real deploy. The runtime hands every
driver its authored tags as a RawTags array (RawPath identity + TagConfig blob,
injected by DriverDeviceConfigMerger) and then subscribes, reads, and routes
published values by RawPath. MTConnectDriverConfigDto declared no RawTags
property, so System.Text.Json silently dropped the array, and every reference
the server passed missed an index keyed by DataItem id: every value came back
BadNodeIdUnknown, with the whole suite green because every test handed the
driver its own Tags array directly.

- MTConnectDriverOptions gains RawTags; the config DTO binds it (verbatim — a
  bad blob is skipped per-tag at binding time, never fails the deployment).
- A new TagBinding, derived from the options and swapped with them in one write,
  carries RawPath -> dataItemId, its one-to-many reverse, and the definition set
  the observation index coerces against. Subscribe/Read resolve through it; the
  fan-out expands each touched dataItemId into every RawPath bound to it and
  publishes under THAT reference — the line the blocker turned on.
- Coercion type precedence: the blob's driverDataType/dataType (both spellings —
  the AdminUI's MTConnectTagConfigModel writes the latter), else a Tags entry for
  the same DataItem, else the Agent's own /probe declaration via
  MTConnectDataTypeInference (what makes a browse-committed {"address": id} blob
  publish as a number), else String. A present-but-invalid type rejects that tag.
- An unmapped reference falls through unchanged, so the dataItemId-keyed CLI
  surface and every existing test keep working, and an unclaimed RawPath reports
  Bad rather than throwing.

15 new tests build their config through the real DriverDeviceConfigMerger; all
four mutations (publish the dataItemId, drop the RawTags binding, drop the
derived coercion type, drop the reverse-map expansion) are caught.
2026-07-27 13:58:06 -04:00
Joseph Doherty b0046d6959 feat(mtconnect): typed AdminUI tag-config model (Task 17)
MTConnectTagConfigModel: pure FromJson/ToJson/Validate for the MTConnect
tag TagConfig JSON blob, mirroring ModbusTagConfigModel/OpcUaClientTagConfigModel.
Fields: fullName (DataItem@id, required), dataType (DriverDataType override,
written as an enum name string), mtCategory/mtType/mtSubType/units (probe
metadata), mtDevice/mtComponent (author context). Preserves unknown JSON
keys across load->save (isHistorized/historianTagname and anything else).

Guards against the enum-serialization trap (numeric dataType would fault the
driver at deploy) both on write (always TagConfigJson.Set via enum name) and
on read (Validate() rejects a dataType that was stored as a bare digit
string, since Enum.TryParse silently accepts numeric text).

AdminUI.csproj gains a ProjectReference to Driver.MTConnect.Contracts,
matching the existing POCO-only driver-Contracts pattern used by the other
six typed tag editors.

Scope: pure model only (Task 17). The razor editor shell + TagConfigEditorMap/
TagConfigValidator registration is Task 18.
2026-07-27 13:58:06 -04:00
Joseph Doherty c232a3ce67 fix(mtconnect): cap CancelAfter overflow + stop leaking ex.Message in probe fallback (review follow-up)
Two review findings on Task 14's MTConnectDriverProbe, both red-first:

- CancelAfter(effectiveTimeout) sat before the try block with only a low-end
  floor on a non-positive timeout, never a high-end cap. A pathological
  timeout (e.g. TimeSpan.FromDays(1000)) threw an uncaught
  ArgumentOutOfRangeException straight out of CancelAfter's ~49.7-day legal
  range, violating the probe's own "Never throws" contract. Not reachable
  through today's only caller (AdminOperationsActor clamps to [1,60]s and
  wraps the call), but a landmine for any future caller without that clamp.
  Fixed by clamping effectiveTimeout to a new MaxTimeout (10 minutes).

- The final `catch (Exception ex) { return new(false, ex.Message, null); }`
  forwarded ex.Message verbatim for any exception type not enumerated above
  it, breaking the redaction discipline every other catch was audited for
  (AgentUri may carry userinfo credentials). Not exploitable by any
  currently-reachable exception type, but an open door for a future one.
  Fixed to build a message from the already-redacted safeUri + the
  exception's type name instead of its message.

Also: a genuine non-2xx stub-handler test (prior HttpRequestException
coverage was connection-refused only), a credentialed-agentUri +
successful-probe redaction test, and a one-line doc remark on the accepted
reachable-vs-deployable gap.
2026-07-27 13:58:06 -04:00
Joseph Doherty a107a1a8b9 feat(mtconnect): host registration + DriverTypeNames const + guard parity (Task 16)
Register the MTConnect factory + probe with the Host, add the
DriverTypeNames.MTConnect constant, and keep DriverTypeNamesGuardTests green.

- DriverTypeNames: new MTConnect const + appended to All.
- DriverFactoryBootstrap: MTConnectDriverFactoryExtensions.Register in
  Register(...) at the default Tier A (managed HttpClient/XML, in-process;
  Tier C is the only tier arming process-recycle, wrong here), and
  TryAddEnumerable(IDriverProbe -> MTConnectDriverProbe) in
  AddOtOpcUaDriverProbes so the probe reaches admin-only nodes (the
  admin-pinned Test-Connect singleton) without double-registering on a
  fused admin,driver node.
- Host.csproj: ProjectReference to Driver.MTConnect (required to compile
  the Register call).
- Core.Abstractions.Tests.csproj: ProjectReference to Driver.MTConnect so
  the guard's bin scan discovers the factory. This and the constant MUST
  land together -- verified RED (2/4 facts) with the constant alone.
- DriverProbeRegistrationTests: MTConnect added to AdminUiDriverTypeKeys;
  its idempotency fact asserts distinct-probe-count and goes RED without
  this (and RED if TryAddEnumerable is downgraded to AddSingleton).

Reconciles all four "MTConnect" literals onto the one constant: the
factory's DriverTypeName, MTConnectDriver.DriverType, and
MTConnectDriverProbe.DriverType now all alias DriverTypeNames.MTConnect
(no circular reference -- Core.Abstractions is a leaf both driver
projects already reference). This is the repo's TwinCat/Focas anti-drift
convention.
2026-07-27 13:58:06 -04:00
Joseph Doherty ba89d6d0ff feat(mtconnect): driver factory delegating to the single config parser (Task 15)
The factory carries NO config DTO of its own: CreateInstance delegates to
MTConnectDriver.ParseOptions, the single authority for the config document.
A second DTO would drift from the parse the runtime's ReinitializeAsync path
uses, and the drift would only surface at deploy time.

Construction is connection-free (agentClientFactory: null) because the Wave-0
universal browser builds a throwaway instance per browse probe, and the
returned instance is never narrowed — its five capability interfaces ARE the
runtime's dispatch surface. Pinned by tests asserting all five plus the
deliberate absence of IWritable.

DriverTypeName is the literal "MTConnect"; Task 16 adds DriverTypeNames.MTConnect
and the host registration that must equal it.
2026-07-27 13:57:46 -04:00
Joseph Doherty 23cd47d54b feat(mtconnect): IDriverProbe one-shot /probe reachability (Task 14) 2026-07-27 13:57:46 -04:00
Joseph Doherty dfbcb0e7f0 fix(mtconnect): close the silent-refusal + teardown-wedge defects in the pump (Task 11 review)
C1 (critical): a latched "this endpoint cannot stream" verdict could end with a
GREEN driver holding a permanently silent subscription. DriverInstanceActor
handles a tag-set change as Unsubscribe-then-Subscribe with no re-initialize;
the unsubscribe cleared the stream-degraded flag, the resubscribe was refused by
the latch without a word, and one successful read then reported Healthy.
StartSampleStreamCore now re-asserts the degradation and logs a Warning naming
the remedy, and teardown no longer clears the flag while the latch is set.

NOT promoted to Faulted: DriverHealthReport maps any Faulted driver to a /readyz
503, so that would de-ready the whole node over one misconfigured Agent whose
reads are perfectly healthy. Faulted stays for the case that earns it.

I1: teardown waits are bounded (5 s) and honour the caller's token. They were
unbounded and ignored it, so DriverInstanceActor's 5 s budget was inert and
PostStop — which blocks a dispatcher thread on ShutdownAsync while holding the
lifecycle semaphore — could be wedged forever by one blocking subscriber. The
cancellation source is disposed only when the loop is provably gone. Applied to
StopProbeLoopAsync too: Task 13 reproduced the identical shape, and closing the
class in one method while leaving the other as the pattern to copy is worse than
not closing it.

I2: the reconnect ladder was monotonic for the process lifetime, so ~9 unrelated
drops pinned the driver at MaxBackoffMs forever. A stream that delivered before
dropping now starts a fresh ladder (at attempt 1, not 0, so MinBackoffMs is
still honoured).

I3: MaxBackoffMs <= 0 clamped every delay to zero — an operator-authorable
reconnect spin. Treated as unset, like a multiplier that cannot grow.

I4: the session published instanceId without NextSequence, so after a restart it
held the new agent's id beside the dead one's cursor, and a gap/OUT_OF_RANGE
re-baseline (id unchanged) never published the cursor at all. Both move in one
CAS, and the pump publishes its cursor as it exits.

I5: OnDataChange was a multicast Invoke — the first throwing handler aborted the
rest of the list, starving every later subscriber. Now walked by hand with the
catch inside the loop. The test that claimed to cover this registered the
recorder BEFORE the thrower, so it passed regardless; swapped, it was red.
Faults are latched to one Warning per stream generation (Debug thereafter) so a
consistently-throwing consumer cannot flood the log.

Also: the pump restarts after an unexpected fault (IsCompleted, not null); a
device-scope that matches nothing is a Warning, not a Debug tally; a device or
component with neither name nor id is skipped rather than emitting a blank path
segment; DiscoverAsync's remark no longer claims DriverInstanceActor retries
(it does not for a Once driver, and injection is dormant in v3); and two flake
seeds in the fake are gone (Dispose re-read its own counter; _disposeCts was
disposed under a racing SampleAsync).

439/439. Every changed behaviour falsified by mutation.
2026-07-27 13:57:46 -04:00
Joseph Doherty fa0e40796f feat(mtconnect): host-connectivity probe + instanceId rediscover (Task 13) 2026-07-27 13:57:46 -04:00
Joseph Doherty 13dd9fcd70 feat(mtconnect): ITagDiscovery streams tree + SupportsOnlineDiscovery opt-in (Task 12) 2026-07-27 13:57:46 -04:00
Joseph Doherty 9a67ed918a feat(mtconnect): ISubscribable /sample pump + ring-buffer re-baseline (Task 11)
One shared /sample long poll behind every subscription handle. Subscribe fires
initial data per reference from the primed /current (Part 4 convention) and
returns without waiting on the stream; the pump runs on its own task under its
own token — never the caller's Subscribe token, which is disposed the moment
that call returns.

The pump owns every way the sequence can break:

- sequence gap, measured against the RUNNING cursor (the previous chunk's
  nextSequence). Comparing against the opening `from` reports a gap on every
  chunk once the buffer rolls — a /current re-baseline storm.
- OUT_OF_RANGE under HTTP 200 (InvalidDataException out of SampleAsync), which
  IsSequenceGap cannot see at all — without it a driver recovers from falling a
  little behind but not from falling a lot.
- agent restart, checked BEFORE the gap because a restart resets sequences and
  so usually trips the gap check too; it clears the index (the held values
  describe a device model that no longer exists) and tells the subscribers.
- every chunk advances the cursor, heartbeats included.

Re-baseline calls CurrentAsync directly on the client the pump already holds:
routing it through ReinitializeAsync would take the non-reentrant lifecycle
semaphore from inside a pump a lifecycle method may already be awaiting, and
hang the driver with no exception and no log. StopSampleStreamAsync is filled
in (same call sites) and InitializeAsync now stops the stream before teardown
too, so a re-Initialize on a live instance cannot leave a pump enumerating a
disposed client.

MTConnectStreamEnded/TimeoutException reconnect under a geometric backoff with
a 100 ms growth floor (MinBackoffMs defaults to 0, and 0 x multiplier is still
0). MTConnectStreamNotSupportedException does NOT retry — it latches, reports
loudly, and is cleared only by a re-initialize.

Health precedence is explicit: /current and /sample fail independently, so each
path owns a flag, clears only its own, and Healthy requires both clear. Neither
can paper over the other's failure.

29 new tests, all deterministic — no sleeps, no wall-clock assertions. The fake
grew per-generation sample streams (so a reconnect has somewhere to land),
scripted /current answers, scripted sample failures, and a chunk-consumed
signal that also fires when the consumer abandons the stream. 375/375 green.
2026-07-27 13:57:46 -04:00
Joseph Doherty a127954cab fix(mtconnect): review remediation for the driver lifecycle (Task 9)
Important 1 - the probe-model triple was read outside _lifecycle while every
write happened under it. Packed (Client, ProbeModel, ProbeModelDataItemCount)
into one immutable AgentSession swapped by a single Volatile.Write, matching
what _index already does, rather than locking the readers: both await the
network, so taking _lifecycle would stall a shutdown for a whole
RequestTimeoutMs and would widen the non-reentrancy hazard. GetOrFetch and
Flush re-publish via CompareExchange so a late re-cache cannot undo a
concurrent teardown/re-init, and a client disposed mid-fetch now surfaces as
the same "not connected" InvalidOperationException browse already handles
instead of a raw ObjectDisposedException.

Important 2 - HasConfigBody decided emptiness by matching the literals "{}" /
"[]", so "{ }", "{\n}" and every pretty-printed empty document fell through to
ParseOptions, failed the required-AgentUri check, and turned a semantically
empty config into a cold-start fault. Now parsed: an object/array with no
elements (or a bare null) is empty; malformed text is deliberately NOT empty so
ParseOptions produces the real quoted error rather than silently starting on
stale options.

Important 3 - documented the _lifecycle non-reentrancy hazard in the code, on
both the semaphore and StopSampleStreamAsync, naming Task 11's re-baseline as
the specific path that would deadlock and giving the CurrentAsync-directly
pattern that avoids it.

Important 4 - removed the Initialize/Reinitialize asymmetry rather than
documenting it. Both now share one rule via ResolveIncomingOptions: an
unreadable config document never destroys working state, and faults only a
driver that had none. InitializeAsync previously tore down before parsing, so a
bad document on a live instance destroyed a healthy client - the exact outcome
ReinitializeAsync was written to avoid.

Minors: SafeDispose traces the swallowed disposal fault at Debug (a client that
cannot be released is how a handle leak starts); the RequirePositive test is a
Theory over all four timing knobs, not just RequestTimeoutMs (arch-review
01/S-6); the footprint test no longer overclaims "is zero before initialize".

CannedAgentClient gains a one-shot ProbeGate so a lifecycle change landing
mid-request is deterministic - no timers. 346/346 (329 + 17). Six mutations
verified: fetch-CAS, disposed-translation, flush retired-session guard,
literal HasConfigBody, teardown-before-parse, and two dropped RequirePositive
calls each fail only their own tests.
2026-07-27 13:57:46 -04:00
Joseph Doherty 1bfc1722b3 feat(mtconnect): IReadable via /current, ordered per-ref snapshots (Task 10)
One deadline-bounded /current per batch, answered out of that one whole-device
document: N references cost one round-trip, one snapshot each, in request order,
duplicates included.

DEVIATION (deliberate, from the plan's implied shared-index write): the response
is indexed into a PER-CALL snapshot index, never into the driver's shared
observation index. That index is written by Task 11's /sample pump and is
last-write-wins with no timestamp arbitration; a read folding its /current into
it would be a second writer racing the first, and /current is frequently the
OLDER document (the Agent composes it while the pump's newer delta is in
flight). Losing that race rolls a subscribed value backwards and leaves it there
until the data item next changes. Both paths coerce through the same
MTConnectObservationIndex logic, so a read and a subscription report a given
observation identically -- the read path is simply a pure function of
(document, authored tags).

Failure posture is the inverse of InitializeAsync: nothing but genuine caller
cancellation crosses the capability boundary. Unreachable Agent / blown deadline
/ unusable answer => every ref codes BadCommunicationError and the driver
degrades (never downgrading Faulted); no client at all (pre-initialize, faulted,
post-shutdown) => BadNotConnected with health untouched. The index's
BadWaitingForInitialData / BadNodeIdUnknown / BadNoCommunication / BadNotSupported
distinctions pass through unmodified.

21 tests pinning arity, order, duplicates, empty batch, one-round-trip, each Bad
code, read-time freshness, the anti-clobber invariant, the deadline, and
cancellation-propagates-but-agent-failure-does-not. CannedAgentClient gains a
cancellation-observing CurrentGate (still no timers or sleeps of its own).
Test csproj takes the AbCip/S7 OTOPCUA0001 NoWarn -- this suite calls the
capability directly by design.

329/329 green.
2026-07-27 13:57:46 -04:00
Joseph Doherty 4fe0af3693 feat(mtconnect): driver shell IDriver lifecycle over agent-client seam (Task 9)
MTConnectDriver implements IDriver only: connection-free ctor, initialize
(deadline-bounded /probe + /current, caching the device model, the agent
instanceId, and a primed observation index), reinitialize, shutdown, health,
footprint, and cache flush.

Decisions worth knowing:

- The driverConfigJson argument is HONOURED, not ignored. DriverInstanceActor
  delivers a config change only via ReinitializeAsync(json) on the live
  instance, so a driver reading only its ctor options would turn every MTConnect
  config edit into a silent no-op that still sealed the deployment green.
  ParseOptions lives on the driver; Task 15's factory must delegate to it rather
  than deserialize a second DTO.
- /probe ok but priming /current failing is Faulted, not Healthy-with-no-data:
  the index IS the read surface and the /sample cursor comes from /current.
- ReinitializeAsync rebuilds the client whenever ANY option the client reads at
  construction changed - AgentUri, DeviceName, RequestTimeoutMs, HeartbeatMs,
  SampleIntervalMs, SampleCount - not just the first two. The timing knobs are
  baked into the deadlines, the watchdog window, and the /sample query string.
- An unparseable config faults a cold start but leaves a RUNNING driver serving
  its previous valid configuration (the deployment fails instead).
- IMTConnectAgentClient now extends IDisposable rather than the driver doing
  `as IDisposable`, which would silently no-op against a non-disposable fake and
  leak an HttpClient + socket handler per re-deploy. Sync, not async: every
  resource behind the seam is sync-disposable, and the async unwind belongs to
  the SampleAsync enumerator the pump owns.
- Flush drops the probe/browse cache and keeps the index; every model consumer
  goes through GetOrFetchProbeModelAsync so a flush cannot wedge browse.

Tests: CannedAgentClient (shared, deterministic - channel-backed scripted
/sample with no timers, call counts, disposal tracking, failure injection) +
27 lifecycle tests. 308/308 in the project.
2026-07-27 13:57:46 -04:00
Joseph Doherty 908bd7ba4a fix(mtconnect): deterministic concurrency test + DATA_SET -> BadNotSupported
The concurrency test asserted a post-Apply post-condition from inside the
race window: dev1_pos is authored, so a reader beating the first Apply
correctly observes BadWaitingForInitialData, which the test's two-member
"legal" set omitted. Consistently red (10/10) rather than intermittent.

Rewritten to assert only properties that hold at every instant — the whole
(status, value, source timestamp) triple against the three legal snapshots,
which is the real torn-snapshot check — plus a read counter guarding a
vacuous pass and one deterministic closing Apply for an exact end state.
The index is unchanged here; BadWaitingForInitialData is a designed state.

Also codes structured (DATA_SET / TABLE) observations BadNotSupported now
that the parser supplies MTConnectObservation.IsStructured, so a two-entry
data set no longer surfaces as the Good string "12".
2026-07-27 13:57:46 -04:00
Joseph Doherty 5ba9f1be6f fix(mtconnect): make a /sample stream end loud, and correct the gap contract (Task 7 review)
Code review of c46540ae approved the framing algorithm but moved the risk
onto the contract around it.

C1 (critical). SampleAsync returned normally on three different events:
EOF, the closing --boundary--, and a response that was not multipart at
all (where it yielded one parsed snapshot and finished). The seam
documents the stream as yielding until ct is cancelled, so a Task 11 pump
written to that contract would silently drop its subscription or hot-loop
reconnecting; the non-multipart leg reported a configuration error as a
healthy finished stream, which is the #485 quiet-successful-termination
shape aimed at the very next task. Every non-cancellation end now throws:
MTConnectStreamEndedException (ConnectionClosed / ClosingBoundary -
transient) or MTConnectStreamNotSupportedException (configuration -
reconnecting reproduces it forever), sharing one catchable base. The
non-multipart body is still parsed first so an Agent's own MTConnectError
text wins, but the document is NOT yielded.

I1. IsSequenceGap moved to IMTConnectAgentClient (static) and its first
parameter is now expectedFrom. The old name and doc were wrong for every
chunk after the first - the correct comparand is the PREVIOUS chunk's
NextSequence, and comparing against the opening "from" argument reports a
gap on every chunk once the ring buffer rolls, i.e. an endless /current
re-baseline storm. Both doc blocks also now record that cppagent answers
from < firstSequence with an OUT_OF_RANGE MTConnectError under HTTP 200,
which surfaces as InvalidDataException and NOT as a gap - so Task 11 must
treat that parse failure as a re-baseline trigger too.

I2. Every multipart test served the whole body from one buffer, so every
framing test completed in a single ReadAsync - the split-boundary path,
the split-header path and the multi-fill loop were correct by inspection
only. A ChunkedStream double (N bytes per read, N = 1/3/7) now drives the
fixtures through a split transport, plus a >16 KB document that forces a
buffer resize. Falsifiability: removing either scan overlap, or
collapsing the fill loop, fails ONLY the new chunked tests.

I3. Disposal mid-enumeration surfaces as OperationCanceledException via
an internal dispose-linked token, not a raw ObjectDisposedException that
Task 9's re-init would inflict on an enumerating pump.

I5. HeartbeatMs, SampleIntervalMs and SampleCount join RequestTimeoutMs
in construction-time positive-value validation. All three reach the query
string and an Agent answers heartbeat=0 with an HTTP-200 MTConnectError
that names no config key.

I4/M2. The no-Content-length framing fallback logs a one-shot warning
(the client takes an optional ILogger); a multipart response with no
boundary parameter fails fast instead of degrading into a watchdog
timeout.

M5/M6. Shared element/attribute reading rules extracted to MTConnectXml;
the probe parser documents why it alone does not trim BOM/whitespace. The
literal U+FEFF in source became a '' escape.

Also, from Task 8: MTConnectObservation gained IsStructured (default
false), set from element.HasElements. A DATA_SET/TABLE observation's
<Entry> children concatenate through element.Value to nonsense ("12" for
two entries) that the index would publish as Good, and only the parser
can tell that apart from a legitimate space-bearing Message. False for
CONDITION observations - their value comes from the element name, so
child content cannot corrupt it.

275/275 green in this suite's own files.
2026-07-27 13:57:46 -04:00
Joseph Doherty ac0a284055 feat(mtconnect): observation index + UNAVAILABLE->BadNoCommunication (Task 8) 2026-07-27 13:57:46 -04:00
Joseph Doherty cf43f8d0ab feat(mtconnect): current/sample parse + sequence-gap detection (Task 7)
Adds the /current and /sample legs to MTConnectAgentClient:

- MTConnectStreamsParser: hand-rolled System.Xml.Linq parse of an
  MTConnectStreams document into MTConnectStreamsResult. Observations are
  every element child of <Samples>/<Events>/<Condition>, keyed by
  dataItemId (the element NAME is the data item's type in PascalCase, so
  a fixed name set would drop observations). A CONDITION's value is its
  element name, not its text - <Normal/> is empty - and <Unavailable/>
  normalizes onto the same UNAVAILABLE sentinel a Sample/Event carries as
  text, so Task 8's no-comms mapping sees one token. Timestamps are
  normalized to DateTimeKind.Utc explicitly.
- IsSequenceGap(requestedFrom, chunk) => chunk.FirstSequence >
  requestedFrom, so Task 11's pump can re-baseline after a ring-buffer
  overflow. Equality is NOT a gap.
- SampleAsync: ResponseHeadersRead + a hand-rolled
  multipart/x-mixed-replace frame reader (Content-length fast path,
  boundary-scan fallback), bounded by a heartbeat watchdog derived from
  HeartbeatMs so a silent Agent faults instead of wedging the pump - the
  S7 frozen-peer shape (arch-review R2-01). The long poll gets its own
  HttpClient with an infinite Timeout so the unary /probe + /current
  deadline stays bounded.

Task 0 package decision, option (c): both TrakHound PackageReferences and
their PackageVersions are dropped. MTConnectHttpClientStream takes a URL
and dials its own socket (no handler/stream seam, so it cannot serve a
socket-free unit suite), and it emits already-parsed documents through the
formatter lookup Task 6 proved is missing - framing-only reuse was never
actually on offer. The driver now depends on nothing but the BCL. Recorded
in the plan's Task 0 CORRECTION block.

Also folds in two Task 6 review carry-overs: the stale
IMTConnectAgentClient doc comment claiming a TrakHound-backed
implementation, and a probe test constructing a mixed-namespace vendor
extension (<x:Pump> with default-namespace <DataItems> children) that
backs the parser's ignore-the-namespace claim with a test.

187/187 green.
2026-07-27 13:57:46 -04:00
Joseph Doherty d3aaa4a411 docs(mtconnect): correct the Task 0 decision — TrakHound cannot parse XML as pinned 2026-07-27 13:57:28 -04:00
Joseph Doherty 1990d21733 feat(mtconnect): agent client /probe parse to device model (Task 6) 2026-07-27 13:57:28 -04:00
Joseph Doherty b13cf5c6e8 fix(mtconnect): options carry the authored tag definitions (Task 2 follow-up) 2026-07-27 13:57:28 -04:00
Joseph Doherty 79e30d1ead fix(mtconnect): infer numeric EVENT types from the probe's UPPER_SNAKE spelling (Task 3 follow-up) 2026-07-27 13:57:28 -04:00
Joseph Doherty 3ca8ddf6f7 test(mtconnect): canned probe/current/sample XML fixtures incl. forced-gap (Task 4) 2026-07-27 13:57:28 -04:00
Joseph Doherty c0729ae5c0 feat(mtconnect): pure data-type inference table + golden test (Task 3) 2026-07-27 13:57:28 -04:00
Joseph Doherty 992ffe5620 feat(mtconnect): IMTConnectAgentClient seam + neutral parse DTOs (Task 5) 2026-07-27 13:57:28 -04:00
Joseph Doherty 18742ea92b feat(mtconnect): driver options + tag definition records (Task 2) 2026-07-27 13:57:28 -04:00
Joseph Doherty e36b920113 build(mtconnect): scaffold Contracts+Driver+Tests projects, add to slnx (Task 1) 2026-07-27 13:57:28 -04:00
Joseph Doherty 3b92b7cc83 docs(mtconnect): record TrakHound-vs-hand-rolled client decision (Task 0) 2026-07-27 13:57:11 -04:00
Joseph Doherty a0b5095b3b docs(drivers): un-stale the MQTT README rows + track every known gap as an issue
v2-ci / build (push) Successful in 4m41s
v2-ci / unit-tests (push) Failing after 16m11s
The cross-driver README still described MQTT as pre-P2 — "Sparkplug B is P2
(`mode: SparkplugB` is a placeholder)" and "Sparkplug B has no fixture yet (P2)".
Both were false as of the P2 merge (c3a2b0f7): SparkplugB is a fully implemented
mode and the live suite is 15 tests (7 plain + 8 Sparkplug) against a
project-owned C# edge-node simulator. Mqtt.md and the scadaproj index were both
already correct; only this file missed the P2 sweep.

Filed the nine Known-gaps rows as tracking issues and linked them, so the table
stops being the only record of them:

  #507  IRediscoverable is inert in v3 (platform gap — also Galaxy/TwinCAT/AbCip)
  #508  write-through (IWritable) — NCMD/DCMD + plain publish
  #509  a metric name containing '/' cannot be browse-committed
  #510  _canRebirth captured at browse-open — reaped session still draws button
  #511  Sparkplug primary-host STATE publishing (ActAsPrimaryHost)
  #512  .Browser references .Driver — extract the client-options builder
  #513  editor writes a meaningless payloadFormat into a Sparkplug blob
  #514  Request-rebirth unarmable on a plant that birthed before the window
  #515  DataSet / Template / PropertySet metric support

#507 is the one worth reading: it is a v3 platform gap, not an MQTT bug —
DriverHostActor.HandleDiscoveredNodes hard-returns, so every IRediscoverable
driver fires into a void and the edge gaining a point needs a redeploy to show up.

No code change; docs + issue links only.
2026-07-27 13:52:08 -04:00
Joseph Doherty 394558e6c2 docs: the umbrella-index edit belongs on scadaproj's main, and must be pushed
v2-ci / build (push) Successful in 5m39s
v2-ci / unit-tests (push) Failing after 3h12m23s
The cross-repo propagation rule said WHAT to update but not WHERE. scadaproj is
a separate repo with its own checkout, so committing there inherits whatever
branch it happens to be on — usually an unrelated feature branch, sometimes one
with no upstream — and the index then disagrees with reality for as long as that
branch stays unmerged.

Observed 2026-07-27: the Sql poll driver's index entry had sat unpushed on
feat/overview-dashboard since 2026-07-24 despite this rule, and the Mqtt driver's
entry was about to join it. Both recovered by cherry-picking onto main.

Records the explicit procedure and says to cherry-pick stranded index commits
rather than merge the branch they are stuck on. The canonical statement now also
lives in scadaproj's own CLAUDE.md (main @ 75f267c).

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-27 13:36:31 -04:00
Joseph Doherty c3a2b0f773 Merge branch 'feat/mqtt-sparkplug-driver'
v2-ci / build (push) Successful in 5m36s
v2-ci / unit-tests (push) Failing after 22m40s
# Conflicts:
#	src/Core/ZB.MOM.WW.OtOpcUa.Core.Abstractions/DriverTypeNames.cs
#	src/Server/ZB.MOM.WW.OtOpcUa.AdminUI/Components/Shared/Drivers/DriverConfigModal.razor
#	src/Server/ZB.MOM.WW.OtOpcUa.AdminUI/EndpointRouteBuilderExtensions.cs
#	src/Server/ZB.MOM.WW.OtOpcUa.AdminUI/Uns/TagEditors/TagConfigEditorMap.cs
#	src/Server/ZB.MOM.WW.OtOpcUa.AdminUI/Uns/TagEditors/TagConfigValidator.cs
#	src/Server/ZB.MOM.WW.OtOpcUa.AdminUI/ZB.MOM.WW.OtOpcUa.AdminUI.csproj
#	src/Server/ZB.MOM.WW.OtOpcUa.Host/Drivers/DriverFactoryBootstrap.cs
2026-07-27 13:31:39 -04:00
Joseph Doherty 123ddc3f40 fix(scripteditor): complete + hover absolute RawPaths (#490)
v2-ci / build (push) Successful in 4m25s
v2-ci / unit-tests (push) Failing after 14m58s
The issue reports a missing feature. The cause is a stale projection, and the
consequence was worse than the gap described.

`ScriptTagCatalog` still emitted `Tag.TagConfig.FullName` — the pre-v3 driver wire
address. Since v3 the mux key is the **RawPath**: `AddressSpaceComposer` harvests the
`ctx.GetTag("…")` literals into `DependencyRefs`, those become `VirtualTagActor`
dependency refs registered with `DependencyMuxActor`, whose `_byRef` is keyed Ordinal
by `AttributeValuePublished.FullReference` — and a driver's wire ref IS the RawPath
(`DriverHostActor._nodeIdByDriverRef` is keyed `(DriverInstanceId, RawPath)`). The
`{{equip}}/<RefName>` form already substitutes to RawPaths at both compose seams.

So the catalog was wrong in BOTH directions, not merely incomplete: completion
offered paths that could never resolve at runtime, while a CORRECT absolute RawPath
got no completion and hovered as "not a known configured tag path". The file's own
comment admitted the raw/UNS catalog was left for "Batch 2/3" and it never came.

Now projected through the shared `RawPathResolver` — the same byte-parity authority
`DraftValidator` and `AddressSpaceComposer` use — so a suggested path is identical to
the one the deploy gate and the runtime compute. A tag with a broken ancestry chain
resolves to null and is omitted: the runtime could not route it either, so offering
it would be a lie. Hover additionally gains the owning `DriverInstanceId`, which the
topology now makes available (it was always null before).

Editor-accepts <=> publish-accepts is preserved: no diagnostic keys off this catalog
(`OTSCRIPT_EQUIPREF` only marks `{{equip}}` paths), so this changes completion and
hover only — it cannot newly reject a script.

The existing tests pinned the v2 contract and were re-authored: a RawPath only exists
if the RawFolder -> DriverInstance -> Device -> TagGroup -> Tag chain does, so the
seed now builds one. Added coverage for group-nested paths, a driver at the cluster
root (no folder segment), a broken chain being omitted rather than guessed, and the
pre-v3 FullName no longer resolving.

Live-gated on docker-dev (rig rebuilt, both central nodes):
- completion, empty literal -> every RawPath incl. group-nested
  `opcua1/plc/OpcPlc/Telemetry/Basic/AlternatingBoolean`
- completion, prefix `pymod` -> `pymodbus/plc/HR200X`, `pymodbus/plc/ImportedTag`
- hover `pymodbus/plc/HR200X` -> "Tag · Type UInt16 · Driver MAIN-modbus"
- hover an unknown path -> "Not a known configured tag path"

AdminUI 742/742; script-analysis 55/55; solution builds 0 errors.
2026-07-26 11:08:11 -04:00
Joseph Doherty 0f38f48679 merge: Modbus RTU-over-TCP transport (#495)
Wave 1 of the driver-expansion program. The branch was complete and live-gated but
had never been merged; it sat 15 commits ahead and 50 behind master.

Purely additive behind the existing IModbusTransport seam — the application protocol
(function codes, codecs, planner, coalescing, write path, probe, materialisation,
HistoryRead) is identical across TCP and RTU; only the wire framing differs
([addr][PDU][CRC-16], no MBAP, no TxId). Zero changes to codecs/planner/health.

Selected by a new Transport config field (ModbusTransportMode: Tcp | RtuOverTcp),
routed through ModbusTransportFactory (throws on unknown mode), with the string-enum
guard on options/DTO/factory and a Transport selector on the AdminUI Modbus form.
ModbusSocketLifecycle was extracted from ModbusTcpTransport (behaviour-preserving)
and is composed by both transports.

Descoped by design: direct-serial (System.IO.Ports) and Modbus ASCII — RTU buses are
reached exclusively via a serial->Ethernet gateway. No per-tag config changes; the
existing per-tag UnitId already makes a multi-drop bus work.

Re-validated against current master at merge time (the branch's own gate predates 50
commits of master): no overlapping files, clean merge, solution builds 0 errors, and
Modbus 344/344, Modbus.Addressing 161/161, AdminUI 740/740, Runtime 507/507,
Configuration 131/131 all pass. Confirmed the Transport selector is still reachable
through master's current authoring path (DriverConfigModal -> ModbusDriverForm), so
the branch's live /run gate result still applies.
2026-07-26 10:58:51 -04:00
Joseph Doherty 0e93587b1c merge: test-correctness, config-secret and AdminUI fixes (#503 #501 #493 #499 #488 #505)
v2-ci / build (push) Successful in 4m14s
v2-ci / unit-tests (push) Failing after 14m48s
Six issues closed, one new one found, filed and fixed.

- #503 Galaxy EventPump flake — root cause was cross-test measurement leakage on a
  STATIC Meter, not a timing budget. 15/15 clean under the issue's own contention
  recipe. Three further defects in the same test fixed on the way, including a
  "held" dispatch loop that was never held (async void on EventHandler<T>) and an
  assertion built from a derived quantity, which could only fail spuriously.
- #501 Three absence assertions given positive controls (a fourth site the issue
  did not list). Proven falsifiable by breaking the production code: both breaks
  passed under the old AwaitAssert form.
- #493 docker-dev seed backfills ClusterNode.GrpcPort; upgrade step documented.
  Deliberately no EF migration — the right value is per-node config the DB cannot
  derive, and a wrong dial target is worse than a NULL that is already skipped.
- #499 Sql credentials refused in ClusterNode.DriverConfigOverridesJson, gated at
  the SAVE rather than the deploy: node overrides never enter the artifact, so a
  deploy gate would not stop the credential being stored. Live-verified.
- #488 Periodic desired-vs-actual subscription reconcile (option A), gated on
  ISubscribable so a non-subscribable driver is not Subscribe-spammed forever.
- #505 The AdminUI cluster-node editor had never worked at all — HTTP 500 on load
  (InputNumber<byte>) and Save a silent no-op (NodeId regex forbids the ':' every
  NodeId contains), with no ValidationSummary to reveal it. Pre-existing on master,
  invisible to 739 green AdminUI tests, found only by going to drive the page.

Also closed #492 (already fixed by 2964361a and drill-verified; no code change).

Verified: solution builds 0 errors; 2406 tests pass across Runtime (507), AdminUI
(739), Configuration (131), Galaxy (317), Cluster (147), Core (260),
ScriptedAlarms (82) and Sql (223). #499 and #505 live-gated on docker-dev with both
central nodes rebuilt; rig restored afterwards.
2026-07-26 10:51:43 -04:00
Joseph Doherty 4cc0215bb4 fix(adminui): the cluster-node editor was dead — 500 on load, silent no-op on save
Both defects are PRE-EXISTING (present on master, unrelated to #499) and were found
live-verifying the #499 guard, which lives on this page and was unreachable without
them fixed.

1. **HTTP 500 on load.** `FormModel.ServiceLevelBase` was `byte` to match the
   entity, and `InputNumber<T>`'s static constructor throws outright for
   `System.Byte` — "The type 'System.Byte' is not a supported numeric type". That
   escapes `BuildRenderTree` unhandled, which in Blazor takes the WHOLE page down,
   not one field. `/clusters/{id}/nodes/{nodeId}` and `.../nodes/new` have never
   rendered. Now `int` with the same `[Range(0, 255)]` and a cast on the way to the
   entity, so the stored type is unchanged. (Same failure mode as #504: an
   unhandled throw inside the render tree is a page-kill.)

2. **Save did nothing, silently.** `NodeId` was validated against
   `^[A-Za-z0-9_-]+$`, which forbids the colon — but every NodeId in this system is
   `host:port` (`ClusterRoleInfo` and `ConfigPublishCoordinator` derive it as
   `member.Address.Host:Port`, the seed writes `central-1:4053`, and
   `NodeDeploymentState.NodeId` is FK-bound to it). So every existing node failed
   validation, `OnValidSubmit` never ran, and the button was inert. Pattern now
   allows `:` and `.` (hostnames may be dotted).

3. **Added the missing `<ValidationSummary/>`.** The form had a
   `DataAnnotationsValidator` but nothing rendering its output, so a validation
   failure produced no feedback at all — which is precisely how (2) stayed
   invisible. This is the change that stops the next such defect being silent.

Live-verified on docker-dev (both central nodes rebuilt, since :9200 round-robins
the pair): page 500 → 200 on both routes, and a real edit now persists.
2026-07-26 10:42:04 -04:00
Joseph Doherty 4b9bcddbcc feat(drivers): periodic reconcile of desired vs actual subscriptions (#488)
Option A from the issue — the cheap consistency check, no interface change.

The desired ref set was re-applied on a Connected TRANSITION (`ResubscribeDesired`)
and on a deploy's `SetDesiredSubscriptions`, and nowhere else. A subscription lost
while the driver STAYED Connected — a subscribe that failed and was never retried,
a handle dropped down an error path — left the actor subscribed to nothing while
still reporting Healthy, indefinitely.

`ReconcileSubscription` now rides the existing 30 s `health-poll` tick: while
Connected, a driver holding a desired set but no live handle is re-subscribed.

Gated on `ISubscribable`. A driver that cannot subscribe at all can still be handed
a desired set, and without the guard it would be told `Subscribe` every tick and
fail every tick, forever.

This is a state-machine consistency check, NOT a probe — it can only see that this
actor holds no handle, never that a handle is present but dead server-side, because
`ISubscriptionHandle` is opaque. That is option B, and it stays deferred until a
real occurrence shows the handle outliving the subscription; it would need a
liveness member implemented across all eight drivers with different subscription
mechanics.

A redundant `Subscribe` is possible (a tick queued ahead of an already-queued one
observes the same no-handle state) and is harmless: `HandleSubscribeAsync` is a
`ReceiveAsync`, so no tick is processed mid-subscribe, and it drops any prior
subscription before establishing the new one. Suppressing it would need a
pending-subscribe flag threaded through every self-tell site — more state than the
race is worth.

`healthPollInterval` becomes an optional Props parameter, mirroring the existing
`rediscoverInterval` seam, so a test can drive the reconcile without waiting 30 s.

Three tests. The reconcile itself is proven by a stub whose FIRST subscribe throws:
nothing else would ever retry it, so a second attempt can only come from the
reconcile — disabling `ReconcileSubscription` turns it red. The two guard tests are
absence assertions and are bounded by a positive control rather than a sleep: a
counting health publisher makes ticks PROCESSED observable, so "still one subscribe"
means the reconcile ran and declined, not that the timer had not fired.

Explicitly not deploy-time drift (#486) — re-asserting an already-correct desired
set would not have helped there.
2026-07-26 10:30:19 -04:00
Joseph Doherty 2193d9f2e4 fix(config): refuse to store a Sql credential in a node config override (#499)
`ClusterNode.DriverConfigOverridesJson` is the third surface a Sql driver's config
is persisted on, and the only one nothing gated. It is a map keyed by
`DriverInstanceId` merged onto the cluster-level `DriverConfig`, so a literal
`connectionString` pasted there leaks exactly as one pasted into the driver — the
leak #498 closed for `DriverInstance.DriverConfig` and `Device.DeviceConfig`.

Gated at the SAVE, not at the deploy — a deliberate departure from both options the
issue offered:

- `DraftSnapshot` carries no `ClusterNode` rows, and adding them would widen the
  snapshot and every builder of one for a single rule about a value that never
  enters the artifact.
- More to the point, a deploy gate is the wrong instrument here. Node overrides are
  not in the artifact, so blocking a deploy would not stop the credential being
  stored — it is already in the database and replicated by then. Refusing the save
  is the only point where "refuse to store it" is literally true, which is #498's
  own framing: discarding a secret on read is not the same as refusing to store it.

The check moves into a shared `SqlCredentialGuard` so the two enforcement points
cannot drift; `DraftValidator` now calls it instead of its own private helper.

Kept deliberately narrow — the `Sql` driver's `connectionString` only, and only for
instance ids that ARE Sql drivers. The issue asked whether to generalise to
credential-shaped keys across every driver type; not doing that, for the same
reason #498 did not: a broader sweep would start refusing configs that are
legitimate today for drivers which never made an indirect-credential guarantee,
which is a regression rather than defence in depth. Widen per driver, as each gains
its own contract.

The error names the offending driver instance(s) and NEVER the value — it reaches
the AdminUI and the audit trail.

14 new cases cover both halves, including the case variants (System.Text.Json binds
`ConnectionString` to `connectionString`, so a case variant is the same key, not a
bypass), non-Sql ids being ignored, and every blank/malformed/non-object shape.
2026-07-26 10:30:04 -04:00
Joseph Doherty 8b7ea3213d fix(rig,docs): backfill ClusterNode.GrpcPort and document the upgrade step (#493)
`GrpcPort` is a nullable column added by Phase 5. Every insert in the docker-dev
seed is INSERT-if-not-exists — the right shape for a re-runnable seed, but it means
a new nullable column never reaches rows an earlier version of the file created. On
a persisted volume all six rows kept NULL, central found no dial targets, and
`TelemetryDial:Mode=Grpc` connected to nothing. (`AkkaPort` escaped this only
because it is NOT NULL with a default of 4053.)

Seed now backfills `GrpcPort IS NULL AND CreatedBy = 'docker-dev-seed'` to 4056.
NULL only — it must not overwrite a port an operator set deliberately on a
long-lived dev volume, and NULL is the unambiguous "predates the column" marker.

Deliberately NOT doing the suggested EF data migration for real deployments. The
correct value is each node's own `Telemetry:GrpcListenPort`, which is
per-deployment configuration the database cannot derive; a migration could only
invent one, and a WRONG dial target is worse than the NULL the dialer already skips
explicitly with a throttled warning. The rig can be backfilled because there the
port is fixed by docker-compose.yml and therefore actually knowable.

`docs/Telemetry.md` gains an upgrade section stating the prerequisite, the symptoms
(the throttled skip log, LISTEN with no ESTABLISHED, the pill never going live) and
the UPDATE to run before flipping any node to Grpc — plus why there is no migration.
2026-07-26 10:29:51 -04:00
Joseph Doherty 240d7aa8aa fix(test): give the three absence assertions a positive control (#501)
`AwaitAssert` returns on its FIRST success, so an assertion that is already true at
t=0 passes instantly and spends none of its duration. All three sites asserted an
absence against a collection that starts empty, so none of them ever gave the
handler the time their comments claimed.

Replaced with mailbox-ordering positive controls rather than sleeps:

- `RouteNativeAlarmAck_unknown_node_is_dropped` — follow the unmapped ack with a
  MAPPED one. The host forwards via `Tell` and the child handles `RouteAlarmAck`
  with `ReceiveAsync` (which suspends its own mailbox), so once the control reaches
  the driver the unmapped one has provably been handled.
- `RouteNativeAlarmAck_on_non_primary_is_dropped` — the role itself is the control:
  return the node to Primary and re-send. Comments (`gated` / `control`) are the
  discriminator, so a leaked ack is visible rather than merged into a count.
- `PeerProbeSupervisorTests.Single_node_snapshot_spawns_no_children` — a FOURTH site
  the issue did not list. It says the other `ChildCount.ShouldBe(0)` waits are each
  preceded by a `ShouldBe(1)`; that is true of three of the four, but not this one,
  which asserts the initial state. Now followed by a two-node snapshot: the count
  reaching 1 proves the single-node snapshot was processed, and a local-node child
  would have made it 2.

The issue warned these may turn red, which would be a real defect surfacing. They
did not — the routing is correct. Proven by breaking it deliberately instead:
disabling the Primary gate reddens the Secondary test, and routing unmapped ids to
an arbitrary mapping reddens the unknown-node test. Under the old form both breaks
passed.
2026-07-26 10:29:41 -04:00
Joseph Doherty 95295210b0 fix(test): scope the EventPump counter capture to its own pump (#503)
Root cause was cross-test measurement leakage, not a timing budget.

EventPump's `Meter` is STATIC — one instance per process — and a `MeterListener`
can only select by *meter* in `InstrumentPublished`. `StartMeterCapture` therefore
received measurements from every `EventPump` alive anywhere in the assembly. xUnit
runs test classes in parallel and several build pumps (`EventPumpStreamFaultTests`,
the `Driver-X` test in this same file), so a sibling's events landed in these
counters. That is why the flake needed a busy box: it needed two tests to overlap.
Measured directly during the gate — `Received = 11` in a run that emitted 10.

The old `Received == Dispatched + Dropped + InFlight` assertion is what turned the
leak into a failure. `InFlight` is DERIVED as `Received - Dispatched - Dropped`, so
the identity is a tautology that cannot fail on a real defect — only on a foreign
increment landing between its four separate `Interlocked.Read` calls. Hence the
reported subject, `counters.Received`. It is deleted, along with the `InFlight`
member, and replaced by direct assertions on the three MEASURED counters.

Three further defects in the same test, all found on the way:

- **The "held" dispatch loop was never held.** `OnDataChange` is `EventHandler<T>`
  — void-returning — so `async (_, _) => await gate` is async VOID: it returned to
  the dispatch loop at the first await and blocked nothing. Measured `Dispatched =
  2..3`, never 0. Drops happened only when the producer happened to outrun the
  consumer, so the arrangement the doc comment describes was fiction and the
  assertions rode on scheduling luck. Now a synchronous `ManualResetEventSlim.Wait()`
  on the loop's own thread, pinned by `Dispatched.ShouldBe(0)`.
- **Fixed `Task.Delay(150)`** replaced by a presence poll on the drop count (#500's
  rule). Dropped starts at 0, so it cannot be satisfied by the initial state.
- **Emission is staged** — one event, wait for the dispatcher to enter the handler,
  then the other nine — so the arithmetic is exact (`10 - 1 held - 2 buffered = 7`)
  rather than scheduler-dependent.

The gate release moved into `finally`, ahead of disposal: `DisposeAsync` awaits the
dispatch loop, which is now genuinely parked on that gate, so a failed assertion
would otherwise hang instead of failing.

Verified: 15/15 clean under the exact contention the issue measured (Runtime +
AdminUI + Core concurrently) — where the half-fixed version still failed on run 4.
Falsifiability: restoring the async-void handler turns the new assertions red.
2026-07-26 10:29:27 -04:00
Joseph Doherty e3155fbbf5 merge: never slice a DB-sourced string with the range operator (#504)
v2-ci / build (push) Successful in 4m11s
v2-ci / unit-tests (push) Failing after 14m22s
`@s.SourceHash[..12]` on an unvalidated nvarchar(64) threw out of
BuildRenderTree and took the whole /scripts page to an HTTP 500. Replaced with a
shared, null/short-safe `DisplayText.Abbreviate`, applied to all five unguarded
range-operator slices over DB-sourced strings in the AdminUI.

Live-verified on docker-dev: /scripts 500 -> 200 with the short hashes rendered
in full; long hashes still truncate with an ellipsis on the other pages.
2026-07-26 09:43:55 -04:00
Joseph Doherty 47d148daf9 fix(adminui): never slice a DB-sourced string with the range operator (#504)
`Scripts.razor` rendered a hash prefix as `@s.SourceHash[..12]…`. `SourceHash`
is a plain nvarchar(64) with no length floor, so any value shorter than 12
characters threw ArgumentOutOfRangeException out of BuildRenderTree — unhandled
in Blazor, which takes the WHOLE page to an HTTP 500, not just the offending
row. Live-reproduced on docker-dev, where two Script rows carry `h1` and
`h-abs-hr200`: /scripts was a hard 500.

New `Components/Shared/DisplayText.Abbreviate(string?, int)`:
- null/blank renders the Admin UI's em-dash placeholder rather than throwing;
- a value already within budget renders verbatim with NO ellipsis, so the
  display never implies text that isn't there (the old markup appended "…"
  unconditionally, outside the slice — `h1` would have shown as `h1…`);
- a maxLength < 1 is clamped, so a display bug cannot become a page crash.

Applied to every unguarded range-operator slice over a DB-sourced string found
by sweeping the AdminUI for `[..N]`:

  Scripts.razor                    SourceHash    (the reported crash)
  Certificates.razor               Thumbprint    (read off the on-disk store)
  Deployments.razor  x2            RevisionHash
  Clusters/ClusterOverview.razor   RevisionHash

The other `[..N]` hits are safe and untouched: Guid.ToString("N") slices
(always 32 chars) and the `Length > 60 ? [..60]` ternaries that self-guard.

NOTE — the issue text blamed the docker-dev seed; that was wrong and is
corrected on #504. `seed-clusters.sql` inserts only ServerCluster, ClusterNode
and LdapGroupRoleMapping. The short-hash rows came from earlier live-gate
authoring. A freshly seeded rig is fine; the trigger is any Script row that did
not go through ScriptEdit / UnsTreeService.HashSource — REST-API authored,
hand-inserted, or migrated in.

Tests: 14 new in DisplayTextTests, covering the short/null/blank/clamped cases
and a never-throws sweep over every value x budget combination. AdminUI suite
739/739.

Live-verified on docker-dev (AdminUI has no bUnit): /scripts 500 -> 200 across
both central nodes, rendering `hash=h1` and `hash=h-abs-hr200` in full;
/deployments still truncates 64-char hashes at 12 with the ellipsis;
/certificates and /clusters/MAIN unchanged.
2026-07-26 09:43:48 -04:00
Joseph Doherty 755ae1aa3f merge: re-assert non-normal scripted-alarm conditions after a (re)load (#487)
v2-ci / build (push) Successful in 4m21s
v2-ci / unit-tests (push) Failing after 3h14m40s
Closes the redeploy-reset gap on Part 9 scripted-alarm condition nodes — the
alarm analogue of the VirtualTag ReassertValue fix.

A full address-space rebuild re-materialises every condition node fresh, while
the engine's reload produces EmissionKind.None for an alarm whose state did not
change — so nothing wrote the node. Live-reproduced on docker-dev: a cleared-
but-still-UNACKNOWLEDGED alarm came back reporting acknowledged, dropping Retain
to false, so it vanished from ConditionRefresh entirely. The operator's
outstanding alarm silently disappeared from the alarm list on deploy.

ScriptedAlarmHostActor.ReassertConditionNodes now writes one AlarmStateUpdate
per alarm not in the Part 9 no-event position, sourced from the new
ScriptedAlarmEngine.GetProjections(). Node-only by construction — the projection
type carries no EmissionKind, so it cannot become an alerts row or a duplicate
historian record.

Live gate PASSED (pre-fix image vs. rebuilt, central-2): condition absent
pre-fix, present + Unacknowledged + Retain=True post-fix, stamped with its own
LastTransitionUtc; /alerts empty across two re-asserts with the panel proven
live by a real ACTIVATED transition.
2026-07-26 09:18:23 -04:00
Joseph Doherty 549c656489 fix(alarms): re-assert non-normal scripted-alarm conditions after a (re)load (#487)
A deploy that triggers a FULL address-space rebuild clears `_alarmConditions`
and re-materialises every condition node fresh — inactive, acked, confirmed,
unshelved. The engine reloads its persisted state and re-derives Active from the
predicate, so it ends up correct; but there is no transition to report, so
`LoadAsync` yields `EmissionKind.None` and `OnEngineEmission` drops it. Nothing
writes the node.

Live-reproduced on the docker-dev rig against the pre-fix image: an alarm that
had fired and cleared but was still UNACKNOWLEDGED came back from a rebuild
reporting Acknowledged with Retain=false — so it disappeared from
ConditionRefresh entirely, while the engine still held Unacknowledged. An
operator's outstanding alarm silently vanishes from the alarm list on deploy.

`ScriptedAlarmHostActor.ReassertConditionNodes` closes it. After each load it
reads the new `ScriptedAlarmEngine.GetProjections()` and sends one
`AlarmStateUpdate` per alarm NOT in the Part 9 no-event position.

Load-bearing properties:

- Node-only. It sends `AlarmStateUpdate` and nothing else — never the `alerts`
  topic, never the telemetry hub. `alerts` is the historization path, so a row
  per alarm per deploy would append a duplicate historian/AVEVA record every
  time anyone deploys. This is why it reads a projection rather than re-emitting:
  `ScriptedAlarmProjection` carries no `EmissionKind`, so there is nothing for
  the emission machinery — or a future refactor — to mistake for a transition.
  The issue's sketch said `GetState(alarmId)` + `LoadedAlarmIds`; that alone
  would have clobbered the node's severity and message, so the projection also
  carries the definition's severity, the resolved message template, and the
  last-observed worst-input quality.
- Only non-normal alarms. One in the no-event position already matches what the
  materialise built, so a never-fired alarm stays byte-identical to a deploy
  without this pass (including its Message, which the materialise seeds with the
  display name), and an all-normal deploy writes nothing at all.
- Timestamp is the condition's own `LastTransitionUtc`, not the deploy instant.
- No duplicate Part 9 events: `WriteAlarmCondition`'s delta-gate compares against
  the node's CURRENT state, so a re-assert onto a freshly materialised node is a
  genuine delta (fires once — correct, the rebuild reset what clients see) while
  one onto a surgically-preserved node is suppressed.

Tests: 6 new host tests + 6 new engine tests. Falsifiability checked by
disabling the call — 4 of the 6 host tests go red; the other 2 are absence
guards against an over-broad fix and pass either way by design.

Live gate (docker-dev, central-2, pre-fix image vs. rebuilt image):
- pre-fix: condition absent from ConditionRefresh after a rebuild;
- post-fix: present, Unacknowledged, Retain=True, stamped with its own
  LastTransitionUtc rather than the restart instant;
- /alerts stayed empty across two re-asserts, with the panel proven live by a
  real ACTIVATED transition immediately afterwards — no duplicate history.
2026-07-26 09:06:39 -04:00
Joseph Doherty 28c2866710 merge: Sql driver follow-ups #496/#497/#498 + the Runtime.Tests flake #500 (fix/sql-driver-followups)
v2-ci / build (push) Successful in 4m15s
v2-ci / unit-tests (push) Failing after 15m14s
Four commits closing the three follow-ups the Sql poll driver left behind, plus the
intermittent Runtime.Tests failure found while verifying them.

- #497 corrects 16 wrong OPC UA status-code constants across 6 drivers (the issue named
  3) and adds StatusCodeParityTests, a reflection guard checking every driver's
  hard-coded uint against the pinned SDK. That guard is what found the other 13.
- #498 fails the deploy when a Sql driver's persisted config carries a literal
  connectionString — the write-side half of a guarantee only enforced on read.
- #496 implements the design 8.1 catalog gate: authored table/column names are resolved
  against the live catalog at Initialize and replaced with the catalog's own spelling, so
  quoting becomes the backstop it was documented to be rather than the sole defence.
- #500 fixes the Runtime.Tests flake: four ordering/logic defects plus a class of tight
  presence budgets, now routed through a documented RuntimeActorTestBase.PresenceBudget.

Verified: full solution builds, all 41 unit-test projects green, Sql integration suite
21/21 against the real SQL Server on 10.100.0.35, and 30 consecutive Runtime.Tests runs
clean against a measured 13% baseline.

Still open by design: #499 (ClusterNode.DriverConfigOverridesJson is a third config
persistence surface DraftValidator cannot see) and #501 (two alarm-ack tests use
AwaitAssert for an absence assertion, so they prove nothing; the rewrite may
legitimately turn them red).
2026-07-25 22:27:26 -04:00
Joseph Doherty 154171f48c test(runtime): fix the intermittent Runtime.Tests failure — 4 ordering/logic defects + the presence-budget class (#500)
The reported symptom was "Runtime.Tests occasionally reports Failed: 1" with no test
name. Instrumenting 30 full-assembly runs reproduced it at 13% (4 runs) and showed it
was never ONE flaky test: five failures across three distinct tests in that batch, and
five more distinct tests over the verification rounds that followed.

Two hypotheses were tested and DISPROVED by measurement before anything was changed:

- Cluster-formation timeout. Every test in this assembly forms a real single-node Akka
  cluster over TCP in RuntimeActorTestBase's constructor and waits 5 s for MemberStatus.Up
  — 219 formations per run. Instrumented: idle max 960 ms, under-load max 1470 ms, zero
  timeouts. A 3.4x margin; not the cause.
- Ephemeral-port exhaustion from that bind churn. TIME_WAIT measured at ~0 throughout.
  Not the cause.

FOUR GENUINE ORDERING/LOGIC DEFECTS (none is a timeout):

1. VirtualTagHostActorTests.ApplyVirtualTags_respawns_child_when_plan_changes_in_place
   The test loops to accept RegisterInterest/UnregisterInterest in EITHER order — and
   then asserts on mux.LastSender, which DEPENDS on that order. Under the [Register,
   Unregister] interleaving LastSender is the DYING child and the assertion fails. Now
   captures the sender of the Register message as it arrives. Fully deterministic; no
   timing involved. (Swept the other four LastSender sites: all drain the Unregister
   first, so their ordering is fixed. Only this one was unsafe.)

2. DriverHostActorWriteRoutingTests.Primary_routes_write_via_uns_nodeid_to_driver_by_rawpath
   ApplyAck marks the end of the APPLY, not the point a write can succeed: the child is
   still in Connecting, which deliberately fast-fails writes ("driver not connected"),
   and the NodeId->driver reverse map has not been pushed. Now retries until accepted.
   This does not weaken the assertion — every rejection branch replies WITHOUT reaching
   the driver, so Writes.Count.ShouldBe(1) still means exactly what it did.

3. HistorianAdapterActorTests.Redundancy_snapshot_without_local_node_...
   One 500 ms constant served both presence budgets (AwaitAssert) and absence windows
   (ExpectNoMsg). Different quantities that happened to share a number, so the presence
   budget could not be raised without slowing every absence check. Split into
   AssertTimeout (5 s) and Settle (500 ms).

4. ContinuousHistorizationRecorderTests.Retry_after_writer_failure_eventually_acks
   Waited on `outbox == 0`, which is ALSO true before the first append — so the poll
   could win the race against the recorder and return having observed the INITIAL state,
   after which the CallCount>=2 check outside the block failed against a recorder that
   had not run. Both conditions are now polled together with the retry count as the
   discriminator. (The preceding test in the same file documents this exact trap.)

THE PRESENCE-BUDGET CLASS:

After those four, three consecutive 30-run rounds each still failed once — on a
DIFFERENT test, in a DIFFERENT file, every time (OpcUaPublishActorApplyFailureTests 2 s,
OpcUaPublishActorRebuildTests 2 s, OpcUaPublishActorTests 500 ms). That is a
distribution, not a defect, and fixing it one test at a time was the wrong method.

RuntimeActorTestBase now carries a documented PresenceBudget (15 s) and 57 sub-3 s
AwaitAssert/AwaitCondition durations across four files route through it.

Why this is safe rather than sloppy: a presence budget is an upper bound before giving
up, not a wait — these helpers poll and return the instant the condition holds. Raising
one costs nothing on the happy path and CANNOT make a genuinely failing assertion pass;
it only changes how quickly a real breakage reports. Absence windows are the opposite
(the elapsed time IS the assertion) and were deliberately left untouched with their own
short, individually-calibrated literals.

ScriptedAlarmHostActorTests is a special case, diagnosed rather than merely raised: its
8 s budget was sized for message passing but actually waits on a REAL Roslyn compile
(ApplyScriptedAlarms -> ScriptedAlarmEngine.LoadAsync compiles every predicate). Cold
script compiles are the slowest and most variable thing in the assembly, so that class
gets 30 s with the reason recorded.

Verification: 30 consecutive full-assembly runs, 30 passed / 0 failed, against a 13%
baseline. Full solution builds; all 41 unit-test projects green.

NOT fixed here, filed as #501: DriverHostActorNativeAlarmAckRoutingTests:90 and :116 use
AwaitAssert for an ABSENCE assertion, so they return instantly and prove nothing. They
are excluded from the budget change on purpose — a bigger number does not fix them, and
rewriting them to settle-then-assert may legitimately turn them red.
2026-07-25 20:20:06 -04:00
Joseph Doherty 4dc2ad8084 feat(sql): implement the design §8.1 catalog identifier-validation gate (#496)
Until now ISqlDialect.QuoteIdentifier was the SOLE defence for authored
table/column names: an identifier went from the TagConfig blob straight into a
command text, bracket-quoted. The residual risk was bounded — a hostile name is
quoted into one nonexistent object, the query fails after the connection opens,
and the tag Bad-codes — but "bounded" is not "filtered", and the design promised
a filter. Both ISqlDialect and SqlServerDialect carried a doc paragraph saying so.

SqlCatalogGate + SqlCatalogLoader now resolve every authored identifier against
the live catalog at Initialize, and REPLACE it with the catalog's own spelling.
The identifier text in an emitted poll query is therefore a string this driver
read back out of ListSchemas/ListTables/ListColumns — not operator input. Quoting
becomes the backstop it was documented to be.

Decisions worth knowing before touching this:

- Substitution, not just validation. Matching is exact-ordinal first, then a
  UNIQUE case-insensitive hit: SQL Server's default collation is CI so case
  variants have always worked, and rejecting them would break valid deployments.
  An ambiguous CI match under a case-sensitive collation is refused rather than
  guessed — picking one would publish another column's data under the operator's
  node, which is worse than rejecting the tag.

- Charset check BEFORE catalog lookup. Each identifier goes through
  QuoteIdentifier for its rejection rules first, so a name carrying a control or
  Unicode format character is rejected WITHOUT its value being echoed into a log
  line (Trojan-Source). A name that passes is safe to render, which is why
  catalog-miss messages do name it — an operator hunting a typo has to see what
  they wrote. Both halves are pinned by tests.

- A rejected tag keeps its node. It is dropped from the POLLED table but stays in
  the AUTHORED table, so it still materializes and reads BadNodeIdUnknown —
  §8.1's specified outcome. My first wiring dropped it from both, which deleted
  the node instead; an existing test (ReinitializeAsync_recoversFromFaulted)
  caught it. A status code can only be published by a node that exists, and a
  missing address-space entry is far harder to diagnose than a Bad quality.

- An unreadable catalog FAULTS Initialize; it does not reject every tag. That is
  the absence of evidence about the tags, not evidence against them — rejecting
  all of them would serve a confidently-empty address space and send the operator
  hunting typos that do not exist. Zero visible schemas is treated the same way,
  because that is exactly what a missing GRANT looks like. Faulting lands
  DriverInstanceActor in Reconnecting with its retry timer running.

- Bounded load: one schema list, one default-schema scalar, then one ListTables
  per distinct authored schema and one ListColumns per distinct authored relation
  — never a full catalog enumeration, and nothing after Initialize. Every
  authored name reaches the catalog queries as a bound @schema/@table parameter,
  so building the allow-list cannot itself be an injection vector.

- ISqlDialect gains DefaultSchemaSql ("SELECT SCHEMA_NAME()"). An unqualified
  `TagValues` must resolve in whatever schema the SERVER reports; hardcoding dbo
  would be a silent lie on any estate that maps service accounts to their own
  default schema. It is a query, not a constant, because the answer is
  per-connection.

- Accepted v1 limitation: a 3-part db.schema.table (or linked-server name)
  addresses a catalog this connection cannot enumerate, so it cannot be
  allow-listed and is rejected with a message pointing at the fix — expose the
  data through a view in the connected database.

VerifyLivenessAsync's wall-clock pattern is extracted to RunBoundedAsync and
reused, so the new I/O gets the same R2-01 / STAB-14 protection rather than a
second hand-rolled copy. (It also fixes a latent bug in the extracted code: the
OperationCanceledException arm detached the wrong task.)

Tests: 24 pure gate tests + 9 end-to-end driver tests against the real SQLite
catalog + 3 new live tests against the real SQL Server on 10.100.0.35 (21 in that
suite now pass, exercising the real SELECT SCHEMA_NAME() + INFORMATION_SCHEMA
path). Verified falsifiable: bypassing the gate turns exactly the three rejection
assertions red and leaves the rest green.

SqlInjectionRegressionTests is deliberately NOT rewritten to expect
BadNodeIdUnknown. It drives SqlPollReader directly, below the gate, and pins the
quoting backstop on its own — defence in depth is only worth the name if each
layer holds independently. Rewriting those assertions would delete the backstop's
only coverage and leave the gate a single point of failure. Its scope note, which
said the gate does not exist, is updated to say why it stays where it is.
2026-07-25 16:07:53 -04:00
Joseph Doherty 1919a8e538 feat(config): reject a persisted Sql connectionString at the deploy gate (#498)
The Sql driver's "a pasted literal connection string is never persisted" guarantee was
enforced only on the READ path: SqlDriverConfigDto has no connectionString property and
UnmappedMemberHandling.Skip drops the key on deserialization. Nothing stopped it being
WRITTEN. Config blobs are schemaless JSON columns and the fallback authoring pattern for
a driver without a typed form is a raw-JSON textarea, so an operator pasting
{"connectionString":"Server=...;Password=..."} would put a live credential in the
ConfigDb — persisted, replicated to every node's artifact cache, readable by anyone with
config access — while the runtime silently ignored it and the driver failed to connect.
Discarding a secret on read is not the same as refusing to store it.

DraftValidator.ValidateSqlConnectionStringNotPersisted fails the deploy
(SqlConnectionStringPersisted) when a Sql-typed driver instance's DriverConfig, or the
DeviceConfig of any device beneath it, carries a top-level connectionString key.

- Both surfaces are checked because DeviceConfig is merged onto DriverConfig before the
  DTO sees it — a credential pasted into the device blob is the identical leak.
- The key is matched case-insensitively: System.Text.Json binds ConnectionString to a
  connectionString property by default, so an ordinal match would leave the obvious
  bypass open.
- Only the top level is scanned; the DTO is flat, so a nested occurrence cannot bind and
  is not the mistake this rule exists to catch.
- The message never echoes the value — validation errors reach the AdminUI, the deploy
  log and the audit trail, and repeating the string there would leak the credential the
  rule is refusing to store. Pinned by a test.
- Malformed/blank/non-object JSON never throws: a parse failure would take down every
  other rule in the same pass.

16 tests. Verified falsifiable — with the rule disabled the 6 positive assertions fail
and the 10 negatives stay green, so none of them passes vacuously.

Design §8.2's "if an inline connectionString is ever permitted (dev convenience only),
the AdminUI redacts it" carve-out is withdrawn in the same change: redaction only hides a
credential that has already been written.

Not covered, and filed as a follow-up: ClusterNode.DriverConfigOverridesJson is a third
persistence surface for a driver config blob, but it is not part of DraftSnapshot, so
this validator cannot see it.
2026-07-25 15:37:48 -04:00
Joseph Doherty b95b99ba8e fix(drivers): correct 16 wrong OPC UA status-code constants + guard them by reflection (#497)
Every value verified against Opc.Ua.StatusCodes in the pinned SDK (1.5.378.106) by
reflection, not by copying siblings. All are client-visible: OPC UA clients branch on
status.

The three named in #497:
  - Historian.Gateway SampleMapper "BadNoData"  0x800E0000 -> 0x809B0000
    (0x800E0000 is BadServerHalted)
  - Galaxy "BadTimeout"                         0x800B0000 -> 0x800A0000
    (0x800B0000 is BadServiceUnsupported)
  - TwinCAT BadTypeMismatch                     0x80730000 -> 0x80740000
    (0x80730000 is BadWriteNotSupported)

BadTypeMismatch was wrong in three MORE drivers the issue did not name — FOCAS,
AbLegacy and AbCip carried the same 0x80730000. Every call site is a
FormatException/InvalidCastException conversion failure, so the name was right and
the value was wrong in all four.

Adding the reflection guard then surfaced nine further defects offline tests had
never checked:
  - Galaxy GoodLocalOverride  0x00D80000 -> 0x00960000  (not a UA code at all)
  - Galaxy UncertainLastUsableValue        0x40A40000 -> 0x40900000
  - Galaxy UncertainSensorNotAccurate      0x408D0000 -> 0x40930000
  - Galaxy UncertainEngineeringUnitsExceeded 0x408E0000 -> 0x40940000
  - Galaxy UncertainSubNormal              0x408F0000 -> 0x40950000
  - TwinCAT BadOutOfService   0x80BE0000 -> 0x808D0000 (was BadProtocolVersionUnsupported)
  - TwinCAT BadInvalidState   0x80350000 -> 0x80AF0000 (was BadAttributeIdInvalid)
  - AbCip/AbLegacy GoodMoreData 0x00A70000 -> 0x00A60000 (was GoodCommunicationEvent)
  - Historian.Gateway GatewayQualityMapper Good_LocalOverride 0x00D80000 -> 0x00960000

Galaxy's whole Uncertain block read as though the OPC DA quality byte could be
shifted into the UA substatus position; it cannot — the substatuses are an unrelated
enumeration. GatewayQualityMapper already held the correct table, which is what the
Galaxy values are now reconciled against.

Guard: StatusCodeParityTests (Core.Abstractions.Tests, which already project-
references every driver) reflects over every `const uint` in the deployed
ZB.MOM.WW.OtOpcUa.Driver.*.dll whose name reads as a status code, and asserts it
equals Opc.Ua.StatusCodes.<name>. Discovery is by convention, so a new driver
assembly referenced by that project is covered with no edit here. It went red on all
nine defects above before the fixes and now checks 109 constants green.

Because reflection can only see a NAMED constant, two inline literals were hoisted so
the guard can reach them: Galaxy's BadTimeout/Bad into StatusCodeMap, and
GatewayQualityMapper's 15-entry table into a new HistorianStatusCodes. An inline
`0x800B0000u, // BadTimeout` at a call site is exactly how the Galaxy defect survived.

The drivers stay SDK-free by design; only the test project references Opc.Ua.Core,
as the oracle.

Also corrected: tests that pinned the wrong values (Galaxy StatusCodeMapTests,
SampleMapperTests, GatewayQualityMapperTests — all written from the same bad source),
a mislabelled comment in DriverInstanceActorTests, and stale constants in two design
docs, one of them the unexecuted MTConnect driver design that would have propagated
BadNoData's wrong value into a new driver.

The CLI's SnapshotFormatter value->name table was checked against the SDK and is
correct; no change needed.
2026-07-25 15:34:31 -04:00
Joseph Doherty ee12568cab refactor(modbus-rtu): distinct CrcMismatch/UnitMismatch desync reasons
v2-ci / build (pull_request) Successful in 4m1s
v2-ci / unit-tests (pull_request) Failing after 13m33s
The RTU framer overloaded DesyncReason.TruncatedFrame for three distinct
failures — a genuine short read, a CRC-16 validation failure, and a
CRC-valid reply from the wrong RS-485 slave. Split the latter two into
dedicated CrcMismatch / UnitMismatch enum members so a future .Reason
consumer (metrics, operator diagnostics) can tell a wiring/EMI CRC fault
apart from a mis-addressed multi-drop reply. Behaviour is unchanged — all
three still throw ModbusTransportDesyncException and tear the socket down.

Framing tests now assert the specific Reason on each path.

Follow-up #12 from the Modbus RTU-over-TCP plan.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-25 15:13:48 -04:00
Joseph Doherty d5bd4226ee docs(mqtt): Sparkplug P2 live-verified; MQTT driver complete
Task 26 — the P2 milestone gate. Ran the live /run verification on an isolated
docker-dev rig (project otopcua-mqtt, ports 9210/4850-4851/14350, image built
from this branch) against the real Mosquitto TLS+auth broker and the C#
Sparkplug edge-node simulator on 10.100.0.35, and recorded the result.

The gate found three defects. Two are the same defect class the P1 gate found
twice — a hand-maintained AdminUI surface left behind by a driver-side feature —
and the third is a pre-existing cross-driver bug that only a Sparkplug flow
could surface.

  1. CRITICAL, FIXED — MqttDriverForm still shipped its P1 Sparkplug PLACEHOLDER.
     Switching Mode to SparkplugB rendered a "not available yet" notice and NO
     Group ID field, so with Sparkplug ingest fully shipped there was still no
     way to author a Sparkplug driver from the AdminUI at all. Sparkplug.GroupId
     is the driver's entire subscription filter (spBv1.0/{GroupId}/#): blank ⇒
     connected, Healthy, ingesting nothing. Now authors all five Sparkplug keys,
     MERGES over the existing sub-object rather than replacing it (so a key a
     newer driver adds inside it survives an older AdminUI), leaves it untouched
     in Plain mode, and validates the group id as the topic segment it is.
     Pinned by 8 MqttDriverFormModelTests cases, verified falsifiable — 8 RED
     with the fix stubbed out.

  2. PRE-EXISTING + CROSS-DRIVER, FIXED — every node label in the shared
     DriverBrowseTree was an <a href="#">. In a Blazor Web App, blazor.web.js's
     enhanced-navigation click interceptor resolves a bare "#" against
     <base href="/">, so clicking ANY browse-tree node label navigated the whole
     AdminUI to "/", tore down the circuit, and destroyed the hosting modal —
     losing the browse session AND the tag selection. @onclick:preventDefault
     does not help: it suppresses the browser's default action, not Blazor's own
     interceptor. Labels are now <button type="button">. Dates to the component's
     introduction (2026-05-28) and affects EVERY driver's picker; it surfaced
     only now because choosing a rebirth scope is the first flow that requires
     clicking a label rather than the ▶ toggle (always a button, always fine).

  3. FIXED (mitigation) — the browse tree rendered exactly once, at open, and
     never again, while OpenAsync returns as soon as the SUBSCRIBE is granted.
     On Sparkplug that means it is almost always empty forever, because births
     are never retained. Added a Refresh button that re-reads the root against
     the SAME session (nothing reconnects, nothing observed is lost, the tag
     selection is kept, only the armed rebirth scope is cleared).

Live gate result (all against the real broker + simulator, group OtOpcUaSim):

  - driver authored end-to-end through the fixed MqttDriverForm; blob is
    Mode:SparkplugB + Sparkplug.GroupId:OtOpcUaSim with enums as names
  - browse: BROWSER OPEN chip, Sparkplug-only rebirth panel, tree renders
    Group → EdgeNode → [Device] → Metric with the device folder Filler1 and
    node-level metric leaves SIDE BY SIDE, and "Node Control/Rebirth" as ONE
    leaf (the metric-name-is-not-a-topic-segment rule holding)
  - rebirth: group scope names the group + the 32-node cap, Cancel published
    NOTHING (simulator log unchanged); edge-node scope → "1 command published"
    and the simulator logged "rebirth NCMD #1 received; republishing metadata";
    group scope → "2 commands published", both edge nodes re-birthed; a device
    and a metric both resolve UP to "edge node EdgeA"; the panel does NOT render
    for a Plain MQTT device
  - browse-commit wrote bindable tuples with no topic/address key —
    {"groupId","edgeNodeId","metricName","deviceId"} for the device metric and
    deviceId ABSENT (not blank) for the node-level one, datatypes inherited from
    the birth (Int64 / Float)
  - deployed (Sealed), and both nodes served live changing values through OPC UA
    at the 2 s cadence — Good 0x00000000, no BadNodeIdUnknown
  - death→STALE: stopping the simulator drove both nodes to 0x80000000 with an
    empty value ~2 s later, and restarting it recovered them to Good on the new
    birth (FillCount back to its birth-declared 1000, then counting)
  - rediscovery fired exactly TWICE (once per authored scope: OtOpcUaSim/EdgeA
    7 metrics, OtOpcUaSim/EdgeA/Filler1 3 metrics) and was then gated across
    many subsequent births/rebirths — the anti-storm change gate working, and
    EdgeB correctly filtered out as an unauthored scope
  - tag editor: opens in SparkplugB (mode inference on reopen) with all four
    descriptor fields populated; blank group → "A Sparkplug group ID is
    required."; "Plant/1" rejected; "Node Control/Rebirth" ACCEPTED and saved
    with the slash whole; dataType override persists and its key is ABSENT again
    after selecting "(from birth certificate)"

Two things are NOT verified, and are recorded rather than claimed:

  - Rediscovery is INERT in v3 and this is a platform gap, not an MQTT one:
    nothing subscribes to IRediscoverable.OnRediscoveryNeeded and
    DriverHostActor.HandleDiscoveredNodes hard-returns. A DBIRTH introducing a
    metric does NOT change the OPC UA tree; redeploy. Verified from the driver's
    log line only, exactly as far as it can be verified.
  - A metric name containing "/" cannot be BROWSE-COMMITTED: the derived raw tag
    Name is a RawPath segment and may not contain "/", so the spec-mandatory
    "Node Control/Rebirth" is refused ("Row 3: Name must not contain '/'"). The
    refusal is loud and all-or-nothing, never a silently mis-bound tag, and
    Manual entry is the working path. Fixing it needs a name-sanitisation policy
    with a collision answer, so it was recorded rather than guessed at.

Also left open + documented: an empty browse tree cannot arm a rebirth (the
scope comes from a tree click, and the session side already accepts a bare
{group}/{edgeNode} — only a UI path is missing); _canRebirth is captured at
browse-open; the tag editor injects the Plain-only payloadFormat key into a
Sparkplug blob (cosmetic — the factory routes on the tuple).

Suites: offline 1528 passed / 0 failed (581 Driver.Mqtt + 809 AdminUI + 138
Core.Abstractions); live 15/15 against the broker + simulator.

Docs: docs/drivers/Mqtt.md brought fully up to date for P2 (Sparkplug config
sub-object, both tag shapes, the ingest state machine, rediscovery, browse +
rebirth + Refresh, a Known-gaps table); Mqtt-Test-Fixture.md gains the simulator
and drops its "Sparkplug has no fixture" claims; infra/README.md §3 gains the
simulator row; CLAUDE.md gains the simulator endpoint facts; the tracking doc
marks MQTT/Sparkplug COMPLETE. docker-dev/docker-compose.mqtt.yml is the
isolated-rig overlay the gate ran on.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-25 01:11:56 -04:00
Joseph Doherty 8d9155682d fix(mqtt): browse-commit writes a bindable address + a Request-rebirth affordance
Two gaps the Task 23/24 review found, both blocking Task 26's live gate.

Gap 1 — browse-commit produced a silently dead MQTT tag. RawBrowseCommitMapper
had no `Mqtt` case, so a committed leaf fell through to the generic
`{"address": …}` key. Neither MqttTagDefinitionFactory entry point reads
`address`: the tag deployed clean and reported BadNodeIdUnknown forever, with
no signal at commit time. Predates Task 23 (Plain was affected too), but Task 23
built the Sparkplug metric tree precisely so an operator could browse and commit
a binding.

The address is a DESCRIPTOR, not a reference string, and it cannot be recovered
from the browse node id: `{group}/{node}[/{device}]::{metric}` where a metric
name legitimately contains `/` (`Node Control/Rebirth`) — the ambiguity
MetricSeparator's remarks already name. So the session STATES it, via a new
`AttributeInfo.AddressFields` seam, and the mapper reads it. Which keys are
emitted is also how the mapper learns Plain vs Sparkplug — the driver type
reaching it is just `Mqtt`, and the mode lives on the driver config. Key names
are single-sourced in the new `MqttTagConfigKeys` (producer, mapper, factory),
with the literals pinned by a test so a symmetric rename cannot silently unbind
already-persisted blobs. A leaf with no stated address is refused at commit in
words rather than committed dead.

Gap 2 — RequestRebirthAsync had no UI. Task 23 shipped it backend-only; Task 26
step 1 assumes the button, and Task 25's runbook notes the picker tree stays
empty until a birth lands, so it is the only way to fill it on demand. Added to
the /raw browse modal: scope = the clicked tree node (device/metric resolve up to
their edge node, group fans out and is refused whole past 32), an explicit
two-click confirm naming the resolved scope and its consequence, the outcome or
the refusal shown rather than swallowed. Offered only for a Sparkplug session
(new `IRebirthCapableBrowseSession.RebirthAvailable` — one session class serves
both modes, so a type test alone would draw a button that can only throw) and
only to a DriverOperator; the server-side gate remains the boundary.

Tests: MQTT 545 → 581, AdminUI 781 → 800. The round-trip tests feed the emitted
blob to the REAL factory and were falsified by mutating the emitted key name.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 23:43:07 -04:00
Joseph Doherty 767e7031b3 test(mqtt): C# Sparkplug edge-node simulator + §3.6 live matrix
Adds a project-owned, controllable Sparkplug B edge node (SparkplugEdgeNode) that
encodes with the SAME generated Tahu schema the driver decodes with — so the live
gate checks encode/decode symmetry rather than the driver against itself — and
drives the §3.6 matrix end-to-end over the real TLS+auth Mosquitto fixture.

The simulator frames its topics independently of SparkplugTopic.Format: a
simulator borrowing the parser's formatter could not detect a bug in it, because
both sides of the comparison would be wrong together. It registers a real NDEATH
Last Will at CONNECT, restarts seq at 0 on NBIRTH, publishes DATA metrics as
alias-only, carries bdSeq per session, and encodes signed ints as two's
complement in the unsigned proto field.

SparkplugLiveTests (Category=LiveIntegration, env-gated, skip-clean) covers:
birth -> alias-only data -> OnDataChange under the RawPath; late-join rebirth
NCMD honoured by an independent decoder; alias reuse across a rebirth routing by
metric NAME; a stale-bdSeq NDEATH ignored while the current one stales; the real
broker-published Will staling and the next birth restoring; seq wrap 255->0
requesting no rebirth while a deliberate gap does; and the browser's passive
window asserted ON THE WIRE (a third client watching spBv1.0/{group}/NCMD/#),
plus RequestRebirthAsync node-vs-group enumeration.

Two things make the suite falsifiable rather than decorative: the seq-wrap test
builds its ingestor with rebirthDebounce: TimeSpan.Zero (at the shipped 10s
default a wrap-triggered NCMD would be swallowed and the test would pass for the
wrong reason) and pairs the silence with a deliberate gap that must produce one;
and the passivity assertion observes the broker rather than the session's own
publish counter, which can only see the seam it guards.

Docker/docker-compose.yml gains a profile-gated `sparkplug-sim` service running
the same engine standalone against the same broker (TLS, CA-pinned) — a manual
driving aid for the AdminUI picker, deliberately NOT a test dependency, since the
matrix needs an edge node it can command mid-test. Its app directory is publish
output (publish-simulator.sh), gitignored like secrets/.

Offline: 15 skipped, 0 failed with no env set. Live: 15/15 passed against
10.100.0.35:8883. Falsifiability spot-check: mutating AliasTable.RebuildFromBirth
to merge instead of replace, and neutering SparkplugCodec.ReinterpretSigned for
Int32, turned exactly the two corresponding tests red and nothing else; both
mutations reverted.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 23:18:08 -04:00
Joseph Doherty 64ec7aa43a feat(mqtt): Sparkplug birth-driven browser tree + scoped Request-rebirth action
Unseals the Sparkplug branch of the MQTT address-picker browser that Task 10
deliberately sealed shut, and adds the first (and only) sanctioned publish on a
browse session.

- MqttBrowseSession serves Group -> EdgeNode -> [Device] -> Metric in Sparkplug
  mode, decoded from observed NBIRTH/DBIRTH certificates (the only Sparkplug
  messages that name their metrics). Node-level metrics hang directly under
  their edge node -- no synthesised device folder, because the authored address
  tuple genuinely has deviceId = null for them. Metric node ids use a '::'
  separator: a metric name may contain '/', so slash-joining would make a
  node-level metric indistinguishable from a device folder.
- AttributesAsync is purely birth-derived in Sparkplug mode. The plain path's
  UTF-8 inference / payload-snippet machinery does not run and no payload bytes
  are retained: a Sparkplug body is protobuf and would be reported as
  "(N bytes, binary)" for every metric in the plant. A datatype ToDriverDataType
  cannot map is reported as the Sparkplug type name, never coerced.
- RequestRebirthAsync(scope) is the one publishing member, reached only by an
  explicit operator action. It routes through the counted PublishAsync
  chokepoint via a private RebirthTransport adapter, so PublishCountForTest
  still proves that OpenAsync/RootAsync/ExpandAsync/AttributesAsync/Observe/
  DisposeAsync publish nothing -- now in both modes. A device or metric scope
  resolves up to its owning edge node (NCMD is node-addressed); group scope
  enumerates observed edge nodes and is refused whole past
  MaxGroupRebirthNodes = 32.
- BrowserSessionService.RequestRebirthAsync gates on the same DriverOperator
  policy that gates the picker, evaluated server-side where the action happens
  rather than only at render time, fails closed, and Info-logs the scope and the
  published count. Exposed through the new IRebirthCapableBrowseSession contract
  so "does this session write?" stays a type question.
- ToBrowseOptions' Last-Will landmine re-verified: MqttDriverOptions still
  carries nothing will-shaped (an NDEATH is a broker-published will and would be
  invisible to PublishCountForTest), now pinned by a reflection guard.

Falsifiability: injecting a publish into RootAsync reddens the Sparkplug and
plain passivity tests; double-publishing per target reddens all four
one-NCMD-per-node assertions. 572/572 MQTT tests, 24/24 AdminUI browsing tests.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 22:51:16 -04:00
Joseph Doherty 6d7a458c4d fix(mqtt): reset Sparkplug ingest state at every session/authoring seam
Task 21 review — two Criticals sharing one root cause, plus four follow-ups.

C1/C2 root cause: `OnReconnectedAsync` opened with `Births.Clear(); _trackers.Clear();`
— "not housekeeping, the correctness step" — and that reset was missing from the
other two seams where the view underneath a surviving cache changes.

- C1: the ingestor OUTLIVES a session rebuild. `SessionIdentity` is a strict
  superset of `IngestIdentity`, so changing only Host/Port/ClientId/credentials/
  TLS/a timeout gives `!SameSession && SameIngest` — `ReinitializeAsync` tears the
  session down and re-establishes on the same instance, and `TeardownAsync` touches
  no ingest state (the resilience layer re-running `InitializeAsync` is the second
  path). A node that rebound alias 5 during the changeover would publish Pressure's
  value under the Temperature RawPath at Good. Hoisted into `ResetSessionState`,
  called from `AttachTo` (earliest seam — a persistent session can push between
  CONNACK and our SUBSCRIBE), `EstablishAsync` and `OnReconnectedAsync`.
- C2: `Register` pruned trackers but not births. While a node is unauthored its
  traffic is dropped at `Dispatch`, so its cached birth FREEZES rather than
  refreshing; drop-then-re-add across a week mis-routes at Good, and `Births` grew
  monotonically. Now evicted on edge-node departure.
  Deliberately NARROWER than "absent from ByScope": a device with no authored tags
  under a still-authored node keeps receiving traffic, so its birth is live, not
  frozen — evicting it would discard correct state. Pinned both ways.

- I1: a birth whose seq is unreadable produced a permanent debounce-paced rebirth
  loop. "A birth arrived" is now tracked separately from "the birth established a
  baseline". Writing the test showed the cap belonged wider than the review scoped
  it: data-before-birth and unknown-alias loop identically (20 messages → 23 NCMDs
  against a cap of 3), so all three missing-metadata reasons share one per-node
  consecutive cap, re-armed by any applied birth. A seq gap stays uncapped — it
  self-limits.
- I2: the oversize `WarnOnce` key was the raw topic, off a group-wide `#`
  subscription. `HandleMessage` now parses and filters to authored nodes BEFORE the
  size check (also skipping the protobuf parse for discarded traffic), and `_warned`
  carries a hard ceiling so a future bad derivation degrades to silence.
- I3: array metrics slipped the unsupported-type gate — `ToDriverDataType` returns
  the ELEMENT type, so a String-typed tag published a packed `byte[]` as base64 at
  Good. Gated on `IsSparkplugArray`.
- I4: driver-level Sparkplug coverage for the composition-root lines T22 left
  untested — the `OnDataChange` re-raise, `ReadAsync` off the shared cache,
  subscribe/unsubscribe dispatch, and both directions of the mode gate.

Minors: M1 the oversize test fed unparseable bytes and passed with the size check
deleted (now a real NBIRTH + an at-the-ceiling control); M2 answered in-place (an
absent seq is refused WITHOUT becoming the baseline, so the NCMD assertion is the
complete check); M3 `HostStateObserved` raise wrapped; M5 the legacy `STATE/{host}`
form is now subscribed, so the parser's tolerance is reachable in production; M8
`ResolveBinding` prefers the NAME over a disagreeing alias; M9 a narrowing float
conversion refuses instead of publishing ±Infinity at Good (a publisher's own
infinity still passes through). M4 was already resolved by T22.

Also fixed a race in the new tests: `publish.Topics.Clear()` after a call that
dispatches rebirths off-thread must drain first — it made the C1 sequence-baseline
pin pass vacuously in the first falsifiability run.

Falsifiability: dropping the reset from AttachTo+EstablishAsync reddens 4 (C1);
dropping the Register eviction reddens 2 (C2); widening the eviction predicate to
the reviewer's literal phrasing reddens the keep-live-births pin (C2b). 545/545 MQTT
unit tests green (was 524); whole-solution `--no-incremental` build clean under TWAE.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 22:44:43 -04:00
Joseph Doherty cd0157a3b8 feat(mqtt): typed tag editor Sparkplug mode + validation
Task 24 (P2): replaces the P1 Sparkplug Validate() stub (always null) and the
"not available yet" editor placeholder with real groupId/edgeNodeId/deviceId/
metricName authoring, mirroring MqttTagDefinitionFactory.FromSparkplugTagConfig
in both directions -- required ids + optional deviceId, dataType/qos read
strictly but only when present (dataType is optional: the birth certificate
declares it). One rule is deliberately stricter than the factory: group/edge-
node/device ids reject '/', '+', '#' (topic-segment characters a decoded
incoming id can never carry, so an authored one that does can never bind);
metricName is exempt since Sparkplug's own canonical names use '/' (e.g.
"Node Control/Rebirth"). No FullName key is written -- TagConfig carries no
identity under v3, and the plan's snippet asking for one was corrected per the
brief. A SparkplugB tag now survives save/reopen because its descriptor keys
re-infer Mode on load; there was never a 'mode' key to persist.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 22:43:44 -04:00
Joseph Doherty 6917d8805d feat(mqtt): Sparkplug UntilStable discovery + rediscover-on-DBIRTH
RediscoverPolicy is now mode-dependent: UntilStable in Sparkplug B (a tag's
dataType is optional -- the birth certificate declares it, so the discovered
tree fills in asynchronously), Once in Plain (nothing about the authored set
arrives on the wire).

DiscoverAsync resolves each Sparkplug tag's datatype per pass from the live
BirthCache, with the ingest path's own precedence: authored dataType wins,
else the birth's, else the record default. An unsupported Sparkplug type
(DataSet/Template/PropertySet/Unknown) falls back rather than blanking the
tag. The authored tag SET is never conditional on a birth -- an authored tag
is part of the declared configuration, not something the plant grants by
publishing.

SparkplugIngestor.BirthObserved is wired to RaiseRediscoveryNeeded behind a
two-part change gate, which is the anti-storm mechanism rather than an
optimisation: fire only for a scope an authored tag binds, and only when the
birth's metric-name SET (ordered-distinct, ordinal) differs from the last one
seen for that scope. Edge nodes re-birth freely -- on their own timer, on
every reconnect via the late-join rebirth fan-out, and once per rebirth NCMD
the gap policy sends -- and firing on each of those would make a healthy plant
a permanent address-space-rebuild loop. No extra debounce: what survives the
gate is a real edit to the served tree, and delaying it would need its own
trailing-edge flush to avoid dropping the last change.

The signature map is NOT cleared on death/reconnect/ingest rebuild (the
ingestor drops its whole birth cache on reconnect, but "the driver forgot" is
not "the address space changed"); it is pruned when a redeploy stops
authoring a scope.

ScopeHint is the "Mqtt" discovery folder, deliberately not the Sparkplug scope
path the plan named: the discovered tree is flat, so an edge-node/device path
would name a subtree that does not exist and a consumer scoping a rebuild on
it would rebuild nothing. The scope is carried in Reason instead.

Falsifiability: always-fire reddens the identical-rebirth + reordered-rebirth
pins; always-authored-datatype reddens the birth-fill pin; dropping the
authored-scope filter reddens the unauthored-device pin. 524/524 MQTT unit
tests green (was 513).

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 22:28:08 -04:00
Joseph Doherty 8538fb798b feat(mqtt): Sparkplug ingest state machine (birth/alias/seq/rebirth/death→stale)
Task 21 — the design doc's §3.6 correctness core. `SparkplugIngestor` owns the
authored `(group, node, device?, metric) → RawPath` index, the live `BirthCache`,
one `SequenceTracker` per authored edge node, and the rebirth policy; it routes
NBIRTH/DBIRTH/NDATA/DDATA/NDEATH/DDEATH/STATE into `OnDataChange` +
`LastValueCache`, or into a bounded, per-node-debounced rebirth NCMD.

Invariants held (and each pinned by a falsifiability control):
- Published reference is the RawPath, never a composed Sparkplug reference —
  `DriverHostActor`'s dual-namespace fan-out is RawPath-keyed.
- Tags bind by stable metric NAME; the alias is only the per-birth lookup, so an
  alias reused across a rebirth cannot mis-route.
- Every value goes through `SparkplugMetricBinding.Reinterpret` before publish.
- NDEATH never reaches `SequenceTracker.Accept` (it carries no seq); it is tied to
  its birth by bdSeq instead.
- An invalid payload mutates nothing — no tracker, no alias table, no eviction.
- `HandleMessage` never throws on MQTTnet's dispatcher thread.

Supporting changes:
- `MqttTagDefinitionFactory.FromSparkplugTagConfig` — the driver had no way to
  parse the Sparkplug descriptor keys at all; the P1 stub fields on
  `MqttTagDefinition` were never populated. Mode selects the parser, never a
  blob heuristic. `dataType` becomes optional in the Sparkplug shape (the birth
  declares it), recorded via the new `DataTypeAuthored` flag.
- `MqttConnection` implements `IMqttPublishTransport` (bounded, refusal throws).
- `MqttDriver` dispatches every capability by mode, subscribes `spBv1.0/{group}/#`
  (+ the configured STATE topic) once at connect, and re-subscribes + requests a
  late-join rebirth on reconnect. `IngestIdentity` now names `MqttSparkplugOptions`,
  closing the silent same-ingest gap its own ⚠️ comment warned about.

Primary-host STATE *publishing* is explicitly out of scope and warned about at
construction: it needs the retained-OFFLINE Last Will set at CONNECT time, and
shipping the ONLINE half alone is worse than shipping neither.

58 new tests (MQTT suite 455 → 513, all green).

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 22:11:41 -04:00
Joseph Doherty 31a98a1bf5 feat(mqtt): RebirthRequester NCMD encode
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 21:40:43 -04:00
Joseph Doherty 3fdf24659f fix(mqtt): AcceptBirth is NBIRTH-only -- a DBIRTH routed there wipes bdSeq
Self-review of the commit before this one caught a defect in the contract it
published rather than in the code it ran. AcceptBirth's doc said it records
"a birth (NBIRTH/DBIRTH)". That is wrong on both halves of this type, and the
wrong version plants exactly the defect the bdSeq tie exists to prevent.

Only the NBIRTH restarts the Sparkplug sequence; a DBIRTH carries the next
number in the edge node's ongoing stream, so it belongs to Accept. And only
the NBIRTH and NDEATH carry bdSeq at all -- a DBIRTH has none to pass. So a
Task 21 that followed the old doc and routed a DBIRTH to AcceptBirth would do
two wrong things at once: rebase the sequence mid-stream, and wipe the node's
session token to null. A stale Last Will arriving afterwards would then find
nothing to compare against, fall through the deliberate fail-toward-stale
rule in IsDeathForCurrentSession, and kill the live node.

Renamed AcceptBirth -> AcceptNodeBirth so the DBIRTH call site reads wrong at
a glance rather than only in prose, documented the hazard on both methods,
and added DeviceBirth_ContinuesTheNodeSequence_AndLeavesTheSessionTokenAlone
to pin the shape at this type's own boundary -- the ingest state machine now
has something to fail against if it wires the call sites the other way round.

Falsifiability: mutating Accept to clear _birthBdSeq reddens exactly the new
test (1 RED), reverted. 29 tests, 438/438 for the MQTT suite.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 21:32:44 -04:00
Joseph Doherty e4e313cf1c feat(mqtt): SequenceTracker seq-gap + bdSeq death-tie
SequenceTracker (Driver.Mqtt/Sparkplug/) holds one edge node's Sparkplug
stream-continuity state -- design doc SS3.6 invariant #3. Two independent
halves: the wrapping 0-255 seq counter that reports whether a message was
missed, and the bdSeq session token that ties an NDEATH to the NBIRTH it
belongs to.

Next-expected is (last + 1) & 0xFF, so 255 -> 0 is the sequence continuing.
The boundary is the whole task because it fails differently in each
direction: read the wrap as a gap and every edge node is asked to rebirth
once per 256 messages, so a busy node spends its life republishing births
instead of data; miss a real gap and the driver serves values it knows are
behind the device, at Good quality, indefinitely. Both directions are
asserted, plus two full wrap cycles so an off-by-one that only misfires at
one value has nowhere to hide.

A gap is reported once and the tracker then resynchronizes onto the observed
value -- holding the baseline at the pre-gap number would make every
following message report a gap too, so one lost message would become a
permanent rebirth storm, the same pathology as a wrong wrap reached by a
different route.

An out-of-range seq is refused, never masked. The codec hands seq over as a
ulong precisely so a spec-violating publisher's value is not truncated into a
plausible-looking one, and masking would finish the job it declined to do:
with the baseline at 255, a seq of 256 masks to 0 -- exactly what the wrap
expects next -- and the bogus message would be accepted as contiguous.
Refused values do not become the baseline either, so the next genuine message
is still measured against the last credible one.

A birth restarts the sequence via AcceptBirth rather than Accept, so it is
never itself a gap; a gap after a birth is still detected. An in-range but
non-zero birth seq is adopted rather than refused -- refusing it would demand
a fresh rebirth for every message from an otherwise-followable publisher. The
first message on a virgin tracker cannot be a gap (nothing to measure it
against) but leaves IsBirthSynchronized false, which is how Task 21 tells a
mid-stream late join from a synchronized node.

bdSeq exists to stop a late Last Will from killing a live node: a node's
connection drops, it reconnects and publishes a fresh NBIRTH, and only then
does the broker notice the old connection died and deliver its will carrying
the PREVIOUS session's bdSeq. Without the compare that stale NDEATH drives
every tag of a freshly-born healthy node to Bad. Unknowns fail toward stale
deliberately -- only a positive mismatch of two known tokens discards a
death, because acting on a death wrongly self-corrects at the next birth
(invariant #4) while ignoring a real one leaves a dead node's values flowing
Good forever. bdSeq is not a top-level field, so TryReadBdSeq extracts it
from the metric named bdSeq and runs it through ReinterpretSigned first: read
raw, a negative Int64 bdSeq becomes an enormous ulong -- a plausible-looking
session token minted out of nonsense.

Falsifiability verified, each mutation reverted after observing RED:
(a) no-wrap expected -> 2 RED; (b) Accept always true -> 5 RED; (c) bdSeq
compare always matches -> 2 RED; (d) mask out-of-range seq into range ->
2 RED. 28 tests, 437/437 for the MQTT suite.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 21:29:08 -04:00
Joseph Doherty 2afef00eaa feat(mqtt): AliasTable/BirthCache — bind-by-name, rebuild-per-birth
Design §3.6 invariants #1/#2. A Sparkplug alias is scoped to ONE birth: after a
rebirth an edge node may point alias 5 at a different metric. So authored tags
bind by the stable metric NAME and the alias is a per-birth cache only —
RebuildFromBirth builds both indexes fresh and publishes them in one assignment,
never merging or upserting, including for an empty birth (which legitimately
clears the scope).

BirthCache keys scopes by the full (group, edgeNode, device?) triple and applies
the spec's node→device rules: an NBIRTH invalidates every DBIRTH under that node
(returning their metric sets for the STALE fan-out), a death evicts rather than
blanks so data-before-rebirth stays detectable. Reads are lock-free on MQTTnet's
dispatcher thread — an immutable snapshot swapped atomically per scope, the same
discipline MqttSubscriptionManager.AuthoredTable uses, with rebuilds landing in
place so a held table reference always observes the newest birth.

Running values are deliberately NOT stored here — LastValueCache owns those,
keyed by RawPath. SparkplugMetricBinding.Reinterpret exposes the birth datatype
to the codec's two's-complement fix, since a DATA metric carries no datatype.

Falsifiability: mutating RebuildFromBirth to merge reddened 5/19 tests (both
plan invariants); making the alias carry metric identity across a birth reddened
3/19, including the headline alias-reuse case.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 21:28:41 -04:00
Joseph Doherty b245f380b1 feat(mqtt): Sparkplug topic parse/format + datatype map
SparkplugTopic (Driver.Mqtt/Sparkplug/) parses/formats spBv1.0 topics for
every message type (NBIRTH/DBIRTH/NDATA/DDATA/NDEATH/DDEATH/NCMD/DCMD/STATE).
TryParse never throws -- it is fed every topic a live spBv1.0/{group}/#
subscription delivers -- and rejects device/node-scope mismatches and MQTT
wildcard chars in a segment. STATE is handled honestly rather than force-fit
into the {group}/{type}/{node} mould: it parses both the v3.0
spBv1.0/STATE/{hostId} form and the legacy no-prefix STATE/{hostId} form
(tolerated on receive only -- Format/FormatState always emit v3.0). Format
gives Task 20's NCMD/DCMD write path a builder instead of hand-concatenation.

SparkplugDataType is a `global using` alias for the vendored proto's
generated Org.Eclipse.Tahu.Protobuf.DataType, not a second hand-duplicated
enum -- Metric.Datatype is a raw wire uint32 (no enum-typed field forces a
second CLR type to exist), and SparkplugCodec (Task 16, landed concurrently)
already casts straight to the generated type. A duplicate enum would be the
same enum-drift hazard this repo already names systemic (CLAUDE.md's driver
enum-serialization bug) and would force every downstream task to cast
between two value-compatible-but-nominally-different enums. ToDriverDataType()
maps per design doc SS3.5: Int8/UInt8 widen to Int16/UInt16 (no 8-bit
DriverDataType member), Float/Double to Float32/Float64 (there is no
DriverDataType.Double), Text/UUID/Bytes/File to String, *Array variants to
their element type (IsSparkplugArray carries the array bit separately), and
DataSet/Template/PropertySet/PropertySetList/Unknown return null -- an
explicit "unsupported, caller must skip+warn" rather than a guessed String.

The completeness test enumerates the live generated DataType member set
(via the alias) and asserts every member is mapped or on the explicit
unsupported list, so a future Tahu proto change is caught automatically
instead of silently falling through a stale duplicate enum's default.
Falsifiability verified by hand for three defect shapes (each reverted after
observing RED): wrong Int8 widening, a dropped mapping falling through
undetected, and inverted node/device topic scoping.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 21:16:02 -04:00
Joseph Doherty c164ec3da1 fix(mqtt): SparkplugCodec's never-throw guard must span the projection too
The try/catch stopped at Payload.Parser.ParseFrom, leaving the metric-projection
loop outside it — so the "never throws for any input" contract had a hole in
exactly the part that walks an attacker-shaped object graph. Widened to cover
parse AND projection.

Self-review finding; no test reddened, because a projection escape needs a
generated-code shape the vectors do not produce. That is the point: the guard is
the contract, not the observed behaviour of today's inputs.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 21:10:20 -04:00
Joseph Doherty 2589774480 feat(mqtt): SparkplugCodec decode + golden payload vectors
Decodes Sparkplug-B wire bytes into a driver-side projection (SparkplugPayload /
SparkplugMetric / SparkplugValueKind) for the Task 21 ingest state machine.

- Never throws, for ANY input. It sits on MQTTnet's shared dispatcher thread, so
  an escaping exception would stall delivery for every subscription on the
  connection, not one tag. Garbage, truncation, zero-length, a valid protobuf of
  another schema and an over-nested Template all resolve to a verdict.
- Zero-length input is INVALID, not "a payload with no metrics" — protobuf would
  parse it as all-defaults, and that is how a truncated-to-nothing body gets
  mistaken for a well-formed one.
- Explicit presence throughout (Has{Seq,Name,Alias,Datatype,Timestamp}), never a
  zero-check: an NBIRTH legitimately carries seq = 0, and every DATA metric after
  a birth carries an alias with no name and no datatype.
- Values are projected RAW, boxed, with the value oneof reported as an explicit
  ValueKind — Absent / Null / Scalar / Unsupported all mean different things and
  three of them carry a null value. DataSet/Template/extension decode as
  Unsupported (v1 scope) rather than throwing or silently vanishing.
- ReinterpretSigned() undoes Sparkplug's two's-complement-in-an-unsigned-field
  encoding of Int8/16/32/64. Kept out of decode because a DATA metric carries no
  datatype — only the consumer, holding the birth's alias table, knows which
  applies. Skipping it publishes 4294967254 for a tag whose value is -42.
- Datatype is carried as the generated Org.Eclipse.Tahu.Protobuf.DataType; the
  SparkplugDataType map is Task 17's and is applied downstream (Task 21).

Golden vectors: nbirth.bin (seq=0, named/aliased catalog, a negative Int32, an
is_null metric, a DataSet metric) + ndata.bin (alias-only metrics, no name, no
datatype) are hand-built by SparkplugGoldenPayloads, committed, and pinned by a
drift guard that rebuilds and byte-compares them on every run — plus a test that
they actually reach the output directory, since a missing copy item is invisible
in source.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 21:09:14 -04:00
Joseph Doherty 4ad540376d merge: SQL poll driver (feat/sql-poll-driver)
v2-ci / build (push) Successful in 4m13s
v2-ci / unit-tests (push) Failing after 3h14m20s
Read-only Sql Equipment-kind driver: polls SQL Server tables/views on an
interval and publishes selected columns/rows as OPC UA variable nodes. Three
projects mirror the Modbus split (Contracts / runtime / Browser); ISqlDialect
seam (SqlServer only in v1); grouped one-query-per-source reads with a
three-layer client-side deadline; schema-walk address picker; typed AdminUI
driver-config + tag-config editors; env-gated central-SQL integration + a
dedicated-container blackhole gate.

Live /run gate PASSED end-to-end on docker-dev: authored a Sql driver +
device + KeyValue tag through the AdminUI, deployed, and read the live value
(dbo.TagValues.Line1.Speed = 42.5, Good) back over OPC UA.

Follow-ups filed: Gitea #496 (§8.1 catalog gate), #497 (cross-driver status
codes), #498 (connectionString persist guard).

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 21:01:26 -04:00
Joseph Doherty 57a5bb3c60 chore(docker-dev): wire Sql__ConnectionStrings__DevSql for the Sql driver live gate
Adds the named connection-string ref the Sql poll driver resolves
(connectionStringRef: "DevSql") to both central nodes' env, pointing at a
DevSql database on the rig's own sql container. Committed-dev-secret exception,
same posture as the ConfigDb line above it. This is the env wiring the Task 21
live /run gate used to author + deploy a Sql driver and read a live value
(dbo.TagValues.Line1.Speed = 42.5) back over OPC UA.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 20:59:27 -04:00
Joseph Doherty 163dd7ab5c fix(mqtt): unbreak the .Contracts csproj -- '--' is illegal in an XML comment
The build-platform note added alongside the proto wiring quoted docker-dev's
`FROM --platform=linux/amd64` verbatim inside an MSBuild comment. XML forbids
'--' in a comment, so MSBuild refused to load the project at all (MSB4025) --
a total build stop for every consumer of .Contracts, i.e. the whole MQTT
driver, the Host, and the test suite. Reworded to name the platform in prose.

Caught by rebuilding after the edit; the preceding commit was made from a
build that predated it.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 20:56:20 -04:00
Joseph Doherty 1ae26675ac feat(mqtt): vendor Tahu sparkplug_b.proto + Grpc.Tools codegen
Vendors the NORMATIVE (proto2) Eclipse Tahu sparkplug_b.proto and compiles it
in .Contracts with build-time Grpc.Tools, message-only (GrpcServices="None") --
Sparkplug rides MQTT, there is no gRPC service, so no Grpc.Core.Api is pulled in.

Provenance rides in the file header: repo, path, pinned commit
5736e404889d4b95910613040a99ba79589ffb13, permalink, git blob SHA-1
bf72ab5f..., content SHA-256 4432c5c4..., EPL-2.0, and a re-verify recipe. The
body below line 38 is byte-for-byte upstream; the blob SHA-1 matches the tree
entry at that commit, so the copy is provably genuine and not reconstructed.

Chose the proto2 file over upstream's sibling sparkplug_b_c_sharp.proto: the
sibling is a lossy proto3 restatement that drops explicit presence, and protoc
generates valid C# from proto2 into the identical namespace
(Org.Eclipse.Tahu.Protobuf, PascalCased from the package -- no
csharp_namespace option). Presence matters downstream: an NBIRTH legitimately
carries seq = 0 and a DATA metric legitimately omits `name`, so Has{Seq,Name,
Alias,IsNull,Datatype} is the difference between "absent" and "present, zero".

.Contracts stays transport-free: Google.Protobuf is a serialization dependency
with a framework-only graph, Grpc.Tools is PrivateAssets=all, and the resolved
graph is exactly {Google.Protobuf, Grpc.Tools, Core.Abstractions} -- no MQTTnet.

Codegen tests pin the round-trip, the namespace, the Sparkplug field numbers,
the DataType spec indices, and protoc's enum-name mangling (UInt64 -> Uint64,
UUID -> Uuid) that Task 17's datatype map has to spell correctly.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 20:54:23 -04:00
Joseph Doherty 92ba120964 feat(mqtt): MqttDriverForm + MqttDeviceForm — the driver/device config forms P1 omitted
The P1 live gate had to insert MQTT driver config directly via SQL: DriverConfigModal and
DeviceModal both fell through to "No typed config form for driver type Mqtt" with no raw-JSON
fallback, so broker host/port/TLS/credentials were unauthorable from the AdminUI. Task 12 built
the *tag* editor; these are the driver + device surfaces.

MqttDriverForm authors the whole MqttDriverOptions connection surface (host/port/clientId,
TLS + CA pin, credentials, protocol version/clean-session/keep-alive, connect timeout, reconnect
backoffs, mode, maxPayloadBytes, and the Plain sub-object). The Sparkplug sub-object stays P2 —
a marked placeholder mirroring MqttTagConfigEditor's stub, with any existing sparkplug keys
preserved untouched.

The connection is authored on the DRIVER, not the device — unlike Modbus/S7/OpcUaClient and
contrary to what docs/drivers/Mqtt.md said. DriverDeviceConfigMerger merges a device's keys up
only when the driver has exactly ONE device, so a device-authored broker connection vanishes
silently the moment a second device is added. MqttDeviceForm is therefore informational (the
GalaxyDeviceForm shape) and round-trips DeviceConfig verbatim; it flags legacy connection keys
left on a device by a pre-form deployment, because those still win via merge-up.

Serialization goes through the ONE shared MqttJson.Options instance (Task 9's decision), never a
fifth per-form copy — pinned by an assertion that an ordinal-only reader CANNOT bind the emitted
blob, and by extending the fleet-wide DriverPageJsonConverterTests guard to resolve a form's
serializer from an explicit external-instance registry when it has no _jsonOpts field. Dropping
the converter from MqttJson.Options reddens 3 tests and leaves the symmetric round-trip green —
the Task 9 lesson, reproduced.

RawTags is stripped from the emitted blob (the deploy artifact owns it) and unknown top-level
keys survive a load->save. Validation mirrors the driver's own [Range] bounds inline, and
ToOptions() additionally clamps, so an ignored error still cannot persist a
connectTimeoutSeconds:0 driver-brick. A non-blank clientId raises the P1 live-gate warning
inline (fixed ids make redundant pair nodes evict each other while both report Healthy).

AdminUI 761/761 green (was 730), TWAE-clean; Driver.Mqtt 266/266 green.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 20:41:48 -04:00
Joseph Doherty b295b24419 feat(adminui): author + configure the Sql driver in /raw
The read-only Sql driver's runtime factory/probe and tag-side editors were
wired, but the AdminUI could not author the DRIVER: the /raw "New driver" type
list, the DriverConfigModal switch, and the legacy identity dropdown all omitted
Sql, and no SqlDriverForm existed. Live-driving the AdminUI surfaced this.

- Add SqlDriverForm.razor over the SqlDriverConfigDto shape: Provider
  (SqlServer-only in v1), required connectionStringRef (an env-var NAME, not a
  connection string — label is explicit), optional poll/operation/command
  timeouts, maxConcurrentGroups, nullIsBad. Enums serialize by NAME via
  JsonStringEnumConverter; a connection-string literal can never be authored
  (no such field is collected); rawTags (composer-owned) + allowWrites (inert,
  read-only v1) are never emitted.
- Wire Sql into DriverConfigModal switch, RawDriverTypeDialog type list, and
  DriverIdentitySection legacy dropdown.
- Add SqlDriverFormContractTests: the exact JSON the form emits constructs a
  driver through the real factory, missing connectionStringRef is rejected, and
  provider round-trips as the "SqlServer" name (never an ordinal).

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 20:41:33 -04:00
Joseph Doherty 07a4a5ff5e docs(mqtt): P1 plain-MQTT milestone live-verified; endpoint recorded
Task 14 — the P1 milestone gate. Ran the live /run verification on a docker-dev
rig against the real Mosquitto TLS+auth fixture, and recorded the result.

The gate did its job: it found MQTT was unauthorable in production.

  RawDriverTypeDialog's driver-type list is a hand-maintained array that nobody
  added MQTT to, so the /raw "New driver" picker never offered it — with every
  other layer (contracts, driver, browser, factory, host registration, typed tag
  editor) complete and 266 green unit tests. Nothing tied that array to
  DriverTypeNames and AdminUI has no bUnit, so no offline test could see it.
  Fixed, plus RawDriverTypeDialogParityTests: a reflection parity guard that
  fails for ANY DriverTypeNames entry missing from the picker. Verified
  falsifiable — RED with "…cannot be authored from the /raw New-driver picker,
  so they are unreachable in production: Mqtt" before the one-line fix.

A second gap is left OPEN and documented, not fixed (no task ever covered it):
MqttDriverForm / MqttDeviceForm do not exist, and DriverConfigModal/DeviceModal
have no raw-JSON fallback — so an operator can create an MQTT driver but cannot
author its broker endpoint or credentials. P1 is not operator-complete until
those two razor forms land. The gate authored config via SQL to get past it.

Live gate result (isolated 2-node MAIN stack, image built from this branch,
fixture at 10.100.0.35:8883, AllowUntrustedServerCertificate rather than
pinning the CA into the rig):

  - typed MQTT tag editor renders and round-trips; Json shows the JSON-path
    field, Raw/Scalar hide it; wildcard topic a/+/c blocked at save; data-type
    list offers Float64 and no Double; qos/retainSeed absent on "(driver
    default)" and qos:2 written as a number; isHistorized/historianTagname
    survive a topic-only edit; SparkplugB renders the P2 placeholder and
    switching back keeps Plain values; jsonPath is guidance not a gate in all
    three shapes ($ seeded only for a brand-new blank tag, blank stays absent,
    clearing saves fine)
  - deployed, and both editor-written blobs resolved and served changing live
    values through OPC UA (22.8 / 1629, StatusCode Good) — no BadNodeIdUnknown
  - the bespoke #-observation browser drove a real broker session and rendered
    the fixture's actual topic tree (line1/{counter,state,temperature},
    noise/chatter, retained/seed); a Modbus device's picker still correctly
    reports "Browsing unavailable"

Also found live and documented (not a code change): both nodes of a redundant
pair run the driver, so a FIXED clientId makes them evict each other forever —
the broker logs "already connected, closing old connection" and both nodes
reconnect every ~2s while still reporting Healthy. Unset (the default) is
correct; confirmed 0 reconnects/60s on both nodes after removing it.

Suites: offline 1134 passed / 0 failed (266 Driver.Mqtt + 730 AdminUI incl. the
3 new guards + 138 Core.Abstractions); live 7/7 against the fixture.

Docs: new docs/drivers/Mqtt.md + Mqtt-Test-Fixture.md (the per-driver + fixture
convention every other driver follows), MQTT rows in docs/drivers/README.md,
TestConnectProbes.md, root README.md, and infra/README.md §3.

CLAUDE.md drift corrected while adding the MQTT endpoints:
  - the "every fixture carries project: lmxopcua" claim was false — only the
    MQTT fixture does; the filter returns one stack, not the fleet
  - lmxopcua-fix.ps1 is Windows-VM-only; on macOS drive the host over ssh/rsync
    (and the docker-dev rig itself runs locally under OrbStack)

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 18:44:43 -04:00
Joseph Doherty 2409d05a28 fix(mqtt): a broker-refused CONNECT no longer reports a Healthy driver
MQTTnet 5 does not throw on an unsuccessful CONNACK — it returns the reason in
MqttClientConnectResult.ResultCode and v5 removed v4's ThrowOnNonSuccessfulConnectResponse,
so ConnectCoreAsync's unconditional SetState(Connected) sealed a wrong-password deployment
green: DriverState.Healthy / HostState.Running on a session that never authenticated.
Caught by the Task-13 Mosquitto fixture; 240 offline tests missed it.

ConnectCoreAsync now inspects the CONNACK and raises MqttConnectRejectedException.
Credentials/identity/protocol/Last-Will refusals are unrecoverable and set Faulted, which
stops the reconnect supervisor; transient refusals (ServerUnavailable, ServerBusy,
QuotaExceeded, UseAnotherServer, ConnectionRateExceeded, and anything unrecognised) stay
Reconnecting under backoff. MqttDriverProbe now shares the rejection wording, so the
AdminUI Test-connect button and the running driver can no longer disagree.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 18:02:00 -04:00
Joseph Doherty 033b3700c4 test(mqtt): Mosquitto TLS+auth fixture + env-gated plain live suite
Adds the MQTT driver integration-test project: an eclipse-mosquitto fixture running
auth AND TLS (allow_anonymous false on both listeners), a mosquitto_pub-based JSON
publisher emitting retained messages, a cert/password generator script whose output is
gitignored, and a live suite gated on MQTT_FIXTURE_ENDPOINT that skips clean offline.

The fixture is never anonymous by design: an anonymous broker would let the driver's
TLS/CA-pin and auth paths go untested, and those fail silently. The CA-pin leg carries
its own falsifiability control (a foreign CA generated in-process must be rejected).

Live gate on 10.100.0.35: 7/7 passed. It found a driver defect, reported not fixed
here (src/ is out of this task's scope): MqttConnection.ConnectAsync RETURNS NORMALLY
when Mosquitto rejects a CONNECT as not-authorized -- MQTTnet 5 does not throw on an
unsuccessful CONNACK, so a wrong broker password yields DriverState.Healthy from
InitializeAsync. The auth-negative test therefore asserts the invariant that holds
either way (no session is ever established) and documents the finding.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 17:49:14 -04:00
Joseph Doherty ff62b62a55 refactor(mqtt): accept a blank jsonPath; guide with a seeded "$" instead
Review follow-up. Dropped the hard Validate rejection of a blank jsonPath under
a Json payload: the runtime defaults it to the document root, which is the real
"publisher puts a bare JSON scalar on the topic" case, so rejecting it made the
editor refuse what the driver accepts — inverting "authoring surface accepts <=>
publish accepts" — and blocked whole CSV-import batches, since this validator
also gates RawManualTagEntryModal's review grid.

Replaced with guidance: an UNAUTHORED tag (no topic AND no jsonPath) seeds the
field with "$", so a fresh tag emits an explicit "jsonPath":"$" while an
existing path-less tag keeps the key ABSENT and the driver defaults it. The
wildcard-topic rule stays strict — a wildcard has no legitimate single-Tag
interpretation.

Also: omit a blank "topic" rather than writing "topic":"", and record why the
_lastConfigJson guard diverges from the Modbus template — the landmine is a
DERIVED, NON-PERSISTED UI FIELD INFERRED FROM BLOB CONTENT (Mode), which Task
24's Sparkplug field group will be tempted to add more of. Named the host-side
dependency (RawTagModal only mutates through this callback) that keeps the
staleness trade benign today.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 17:37:13 -04:00
Joseph Doherty a13ae926a8 fix(mqtt): gate DisposeAsync against in-flight lifecycle ops (C1) + rebuild-branch tests (I1)
C1 (Critical) — DisposeAsync went straight to TeardownAsync with no lifecycle-gate
wait, reopening the orphaned-connection class. Interleaving: ReinitializeAsync's
rebuild branch holds the gate mid-InitializeCoreAsync, past its own teardown but
before `_connection = connection`; an ungated dispose sees a null connection, does
nothing, and sets `_disposed`; the rebuild then publishes a live connection plus a
host-probe loop that every later dispose short-circuits past. Not proven reachable
today (DriverInstanceActor.PostStop calls the gated ShutdownAsync), but MqttDriver
publicly implements IAsyncDisposable, which invites `await using`.

- DisposeAsync now takes the gate with the same bounded wait + fallback teardown
  ShutdownAsync uses, and sets `_disposed` BEFORE the wait.
- InitializeCoreAsync re-checks `_disposed` after the connect and disposes the
  connection it just built rather than publishing it — this closes the residual
  window on the bounded-timeout fallback path.
- The doc comment no longer claims parity with ShutdownAsync it did not have, and
  stops conflating "don't Dispose() the semaphore object" with "don't WaitAsync".

I1 (Important) — the SameSession == false rebuild branch had no test driving it.
Adds three: an endpoint-changing delta rebuilds and Faults on a refused connect;
an ingest-only delta (MaxPayloadBytes) ALSO rebuilds, pinning that IngestIdentity
is nested inside SessionIdentity; and DisposeAsync serializes behind a lifecycle
operation parked at a new internal BeforeConnectHookForTests seam (mirroring
MqttConnection's AfterConnectHookForTests, which exists for the same race class).

Minors: comments recording that IngestIdentity names no Sparkplug field and must
grow one in P2 (tasks 21/22), and why _options/_subscriptions/_authoredRawPaths
get looser memory discipline than _health/_hostState.

Falsifiability: reverting DisposeAsync to the ungated form reddens exactly the new
serialization test ("Shouldly.ShouldAssertException : raced"); forcing SameSession
to always-true reddens exactly the two new rebuild tests. 240/240 MQTT tests pass;
forced rebuild of the driver project is 0 warnings under TreatWarningsAsErrors.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 17:31:04 -04:00
Joseph Doherty f79d13e2d8 feat(mqtt): factory + DriverTypeNames.Mqtt + host factory/probe registration
Task 9. Adds MqttDriverFactoryExtensions (DriverTypeName = DriverTypeNames.Mqtt,
direct deserialization of MqttDriverOptions — no intermediate DTO, mirroring
OpcUaClientDriverFactoryExtensions), the DriverTypeNames.Mqtt constant, and both
host wiring sites: the factory in DriverFactoryBootstrap.Register and the probe in
AddOtOpcUaDriverProbes (the admin-node path Program.cs calls in its hasAdmin
block — a probe wired only on driver nodes makes Test-connect silently dead).

Carried-forward items:

- Converges every MQTT config-parsing seam onto ONE JsonSerializerOptions —
  MqttJson.Options, in .Contracts alongside MqttDriverOptions. It replaces the
  probe's internal JsonOpts and the browser's separate private copy; the factory,
  the probe, MqttDriver.ParseOptions and MqttDriverBrowser now all parse through
  it. .Contracts is the only assembly all four consumers reference, and the
  browser's reference to the runtime .Driver project is a documented layering
  exception scheduled for removal — anchoring the options there would resurrect
  the duplicate the day it goes away.
- Replaces the "Mqtt" literals in MqttDriverProbe, MqttDriverBrowser and
  MqttDriver with the constant (string value unchanged).
- Tightens MqttDriverProbeTests.ProbeAsync_EnumAsName: it asserted only
  Ok == false + non-empty message, which is also exactly what a JSON-parse
  failure produces — so it stayed green under the very regression it names.
  It now asserts the probe got past the parse and reached the network.

Falsifiability: deleting JsonStringEnumConverter from MqttJson.Options reddens 9
tests across 4 suites, including the tightened probe test (message becomes
"Config JSON is invalid: The JSON value could not be converted to
MqttProtocolVersion") — which the pre-fix assertions would have passed.

Also references the MQTT driver from Core.Abstractions.Tests so
DriverTypeNamesGuardTests' reflective bin scan discovers the new factory and the
constant/factory parity check stays honest.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 17:24:13 -04:00
Joseph Doherty 36abd871a1 feat(mqtt): typed tag editor + validator (plain mode)
Thin typed model over a preserved JsonObject key bag, mirroring the sibling
<Driver>TagConfigModel template. Key names and strictness rules mirror
MqttTagDefinitionFactory rather than inventing an editor-side schema; enums
round-trip as NAMES (the factory reads them strictly by name).

No FullName key: under v3 a tag is identified by its RawPath, which the factory
keys the definition's Name off — a composed identity key here would be dead
weight nothing reads.

Mode is a UI-only sub-shape selector inferred from the blob (any Sparkplug
descriptor key present) and never serialised — the driver takes Plain vs
SparkplugB from its DRIVER config. Only Plain Validate() is implemented; the
Sparkplug branch is a stub Task 24 fills, and its keys are preserved untouched.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 17:17:37 -04:00
Joseph Doherty 420692b6e5 feat(mqtt): MqttDriver shell — lifecycle + authored-only discovery (Once)
Composes MqttConnection + MqttSubscriptionManager + LastValueCache into the
IDriver / ITagDiscovery / ISubscribable / IReadable / IHostConnectivityProbe /
IRediscoverable capability set.

- Discovery replays ONLY the authored raw tags (RediscoverPolicy = Once,
  SupportsOnlineDiscovery = false) — a chatty broker can never auto-provision.
- Composition order is register -> AttachTo -> ConnectAsync; the manager's
  Reconnected handler is passed through unwrapped so its throw-on-total-failure
  still tears a deaf session down.
- ReinitializeAsync applies a tag-only delta in place and never faults on a bad
  one; a session-changing delta rebuilds.
- ReadAsync serves the last-value cache and degrades per reference
  (BadWaitingForInitialData), never throwing for the batch.
- FlushOptionalCachesAsync is a no-op: the last-value cache IS IReadable.
- Adds MqttDriverOptions.RawTags (the authored-tag delivery mechanism every
  other driver's options DTO already has) and promotes MaxPayloadBytes from a
  manager ctor knob to an operator-facing key.
- Converges on MqttDriverProbe.JsonOpts rather than a second, divergent
  JsonSerializerOptions.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 17:06:39 -04:00
Joseph Doherty 211c6ea6d6 feat(mqtt): register MqttDriverBrowser (bespoke-first for Mqtt)
Wires the bespoke MqttDriverBrowser into AdminUI's IDriverBrowser
registrations alongside OpcUaClient/Galaxy, overriding the universal
DiscoveryDriverBrowser fallback for driverType "Mqtt".

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 16:48:55 -04:00
Joseph Doherty a856d61510 feat(mqtt): MqttDriverProbe CONNECT handshake
Test-connect probe for the Mqtt driver type: parses MqttDriverOptions
(shared enum-as-name JsonSerializerOptions), opens a bounded MQTT CONNECT
via MqttConnection.BuildClientOptions with a distinct -probe-{guid8}
client-id suffix, and classifies the outcome. Live-probed against
MQTTnet 5.2.0.1603 to confirm the real contract: a broker-rejected
CONNACK is a MqttClientConnectResult.ResultCode, never a thrown
exception, and timeout classification checks the deadline token itself
rather than switching on exception type (which varies unpredictably
between MqttConnectingFailedException and bare OperationCanceledException
for the same frozen-peer shape).

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 16:43:36 -04:00
Joseph Doherty abf3116bd4 fix(mqtt): wildcard tags went dark on reconnect — index by filter, not topic
C1 (CRITICAL). AuthoredTable routed any wildcard-authored tag into `Wildcards`
and never into the concrete topic index. But the two paths that reconstruct
state from a SUBSCRIBED FILTER STRING — OnReconnectedAsync's re-subscribe set
and DegradeTopic — both looked up that concrete-only index, and
`_subscribedTopics` keys on the wildcard PATTERN ("f/+/temp"). So:

- reconnect resolved zero filters for a wildcard tag, hit a silent early
  return, and issued NO subscribe. The cache keeps its last Good value, so the
  tag never turns Bad — it just stops updating forever behind a connection
  reporting healthy. Exactly the silent-death mode MqttConnection's own remarks
  warn about, left unguarded for wildcards.
- DegradeTopic missed too, so a SUBACK rejection of a wildcard filter degraded
  nothing and left the tag at BadWaitingForInitialData.

Fix: AuthoredTable now carries `ByFilter` (EVERY def, keyed by its authored
topic filter, wildcards included) alongside `ByExactTopic` (concrete only) and
`Wildcards`. Delivery matches an incoming published topic — which never carries
a wildcard — so it keeps using ByExactTopic + the comparer scan. Every path
keyed on a filter string uses ByFilter. The distinction is documented on the
record. The silent early return is now a Warning naming the orphaned topics.

Also in this pass:

- I2: bounded decode on the dispatcher thread. A body over `MaxPayloadBytes`
  (default 1 MiB, ctor-settable) is refused BEFORE any GetString/JsonDocument
  parse and degrades its own tags. Unbounded decode on the shared dispatcher is
  paid by every subscription, not just the offending topic. NOTE: promoting this
  to an operator-facing MqttDriverOptions key (+ WithMaximumPacketSize) needs a
  .Contracts edit this task is scoped out of; flagged in the XML docs.
- I3: integral-valued reals now coerce to integer tags. JS/Python edge gateways
  serialize an integer as 5.0, and TryGetInt32/int.TryParse are syntactic — an
  Int32 tag fed by such a gateway silently never received data. Parsed via
  decimal so integrality and range are exact across Int64/UInt64. A genuinely
  fractional 5.5 is still refused, never rounded.
- I4: the JSON document is parsed ONCE per message and shared across the whole
  fan-out (the documented "one document, one JSONPath per signal" shape was
  re-parsing per tag). Zero-alloc via a ref struct when no Json tag matches.
- I5: pinned the documented JSON-string-holding-a-number coercion end-to-end.
- Minor: Register now prunes `_subscribedTopics` / `_handleByRawPath` entries for
  tags a redeploy dropped, so a deleted topic stops being re-subscribed forever.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 16:40:31 -04:00
Joseph Doherty 4ec01bdfc1 fix(mqtt): bound browse observation against a chatty or malicious broker
Three review findings, all "unbounded work/memory from untrusted broker input"
rather than correctness bugs - but this points at a live plant broker.

1. InferPayload decoded the WHOLE payload on MQTTnet's dispatcher thread before
   trimming to 64 chars. The node cap never engages here: a multi-MB blob on an
   already-known topic creates no node, it just updates one, so the cost repeats
   per message forever. Now at most PayloadInspectMaxBytes (512) are examined.
   A blind slice can cut a multi-byte sequence, which the strict decoder would
   report as "binary" - TrimToUtf8Boundary backs the window off to the last whole
   sequence so good UTF-8 is never mislabelled. A clipped payload is String by
   construction rather than inferred from a partial view.

2. The node cap bounded COUNT, not bytes. MQTT topics run to ~65KB, so 50k nodes
   near that ceiling is a multi-GB tree. Topics now bounded at 1024 chars and
   segments at 256, rejected whole (a truncated topic is a DIFFERENT topic the
   picker would commit as a tag address) and surfaced via an __oversized__ marker
   kept distinct from __truncated__ - the operator's remedy differs.

3. Control characters in a snippet would break the picker's single-line row;
   Trim() only strips the ends. Now scrubbed to spaces.

Each guard was falsified independently by neutering it and confirming exactly one
test reddens. That found a real gap: the oversized-topic test was passing via the
SEGMENT bound, leaving the whole-topic bound untested - the test now uses many
short segments so only the topic bound can reject it.

Also records the .Driver project reference as a known, deliberate exception to the
.Browser -> .Contracts pattern, with its cost (AdminUI gains Core + Polly + Serilog)
and the clean fix (a leaf transport project; NOT a move into .Contracts, which is
deliberately transport-free).

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 16:30:45 -04:00
Joseph Doherty d421487bcd feat(mqtt): plain subscribe→OnDataChange with retained seed + ref resolver
MqttSubscriptionManager is the P1 plain-MQTT ingest path: it holds the authored
RawPath → MqttTagDefinition table behind the shared EquipmentTagRefResolver,
establishes one SUBSCRIBE per distinct authored topic, and turns each inbound
message into a LastValueCache update plus an OnDataChange notification.

Load-bearing choices, each pinned by a falsified test:

- The published reference IS the RawPath (def.Name), never the topic and never
  the TagConfig blob — DriverHostActor's dual-namespace fan-out is RawPath-keyed,
  so any other key passes every unit test and delivers nothing in production.
- One message fans out to EVERY tag on its topic; several tags routinely share a
  topic with different jsonPaths, and stopping at the first match silently starves
  the rest.
- An unauthored topic raises nothing: MQTT ingest never auto-provisions.
- Wildcard-authored topics (accepted by the parser, warned at deploy) match via
  MQTTnet's own MqttTopicFilterComparer, so '+'/'#' behave as a broker would.
  Concrete topics stay a single lookup; the wildcard scan runs only when one exists.
- Retained seeds are honoured natively on MQTT 5.0 (retain handling) AND filtered
  client-side on the retain flag, because 3.1.1 brokers always replay.
- Nothing throws on MQTTnet's dispatcher thread: a malformed payload degrades that
  tag (BadDecodingError / BadTypeMismatch), and a throwing OnDataChange subscriber
  is contained without starving the tags behind it.
- The manager issues the FIRST subscribe itself — Reconnected deliberately does not
  fire after the initial ConnectAsync.
- Reconnect re-subscribe: total failure THROWS (MqttConnection then tears the
  session down and retries under backoff, rather than serving a connected-but-deaf
  driver); a partial SUBACK rejection degrades only the denied refs, because
  retrying forever over one ACL-denied topic would take every tag dark.

MqttConnection gains a MessageReceived observer (fired on the dispatcher, throwing
observers contained) and a bounded SubscribeAsync under its own linked-CTS deadline
— deliberately not MqttClientOptions.Timeout, which already governs every MQTTnet
op. A refused filter is an outcome, not an exception; only the SUBSCRIBE itself
failing throws. Classify() is parameterised by operation name (messages unchanged
for connect).

JSONPath is a documented subset ($, $.a.b, $['a'], $.a[0]); filters/slices/
recursive descent report a miss rather than pulling in a JSONPath package for an
expression form that selects a set where a tag needs one scalar.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 16:12:35 -04:00
Joseph Doherty e27eb77187 docs(modbus-rtu): record Wave-1 RTU-over-TCP live /run gate result (PASSED)
v2-ci / build (pull_request) Successful in 3m45s
v2-ci / unit-tests (pull_request) Failing after 13m54s
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 15:50:10 -04:00
Joseph Doherty f2a9ee5d93 docs(mqtt): flag the Last-Will landmine at the browse-options seam
A will is published by the BROKER, so it is invisible to PublishCountForTest.
MqttDriverOptions carries none today, but Sparkplug NDEATH is a will message -
once P2 adds it, an ungraceful browse disconnect would fire NDEATH under the
plant's own edge-node identity. ToBrowseOptions is where it must be cleared.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 15:46:23 -04:00
Joseph Doherty 7980b34692 feat(mqtt): bespoke passive #-observation browser (plain)
MQTT has no discovery protocol, so the AdminUI address picker learns the topic
space the only way there is: subscribe to the wildcard and accumulate what
arrives. MqttBrowseSession serves that observation as a segment tree (split on
'/'); MqttDriverBrowser opens it with a distinct "{clientId}-browse-{guid8}"
identity under a 5-30 s clamped budget.

Browsing publishes NOTHING - the load-bearing safety property, since the picker
runs against a live plant broker. Every outbound message must route through the
single PublishAsync seam that counts into PublishCountForTest; P2's operator-
triggered Sparkplug rebirth is the one intended exception. Sparkplug mode fails
fast rather than serving a raw spBv1.0 topic tree that looks like a metric tree.

Deviation from the plan's ref list: the project also references .Driver so the
browse CONNECT reuses MqttConnection.BuildClientOptions (TLS posture, CA-pin
chain validator, credentials). A second copy would be a security divergence.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 15:45:32 -04:00
Joseph Doherty 585f827ef3 fix(sql): close the residual credential leak in HandlePollError
The C1 credential-hygiene fix converted InitializeAsync and ReadAsync to
surface only ex.GetType().Name, but left HandlePollError building its
operator-facing Degrade() message from raw ex.Message, gated only by
'if (ex is DbException) return;'. That guard exists for the I1
double-classification concern, not for message safety, and it is incomplete
for the leak vector: a malformed connection string throws ArgumentException
from the keyword parser (the same unquoted-';'-in-password shape the sibling
SqlDriverBrowser.Sanitize special-cases), which is not a DbException and so
reaches Degrade(ex.Message) unredacted.

It is unreachable in today's control flow only because the connection string
is static and validated identically at Initialize first — an emergent
property, not an enforced invariant, and a live per-poll leak the moment a
future edit (refresh-on-reinit, a different provider) breaks that assumption.

Make the fallback type-only like the other two sites; the full exception still
reaches the log sink via the exception parameter. The I1 defer + control tests
are unaffected (they assert Degraded, not message text). Pinned by a new test
driving a credential-bearing ArgumentException through HandlePollError;
verified load-bearing (red against Degrade(ex.Message), green after).

Closes the C1 review finding on commit 37f444b9.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 15:39:14 -04:00
Joseph Doherty be1df2d1e5 feat(sql): SqlTagConfigEditor razor shell
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 15:35:25 -04:00
Joseph Doherty 9b5a96876e feat(sql): register Sql factory + probe in Host, add DriverTypeNames.Sql
Add DriverTypeNames.Sql (+ All) as the shared source of truth, and wire the driver
into the Host: register the factory in DriverFactoryBootstrap.Register via
Driver.Sql.SqlDriverFactoryExtensions.Register(registry, loggerFactory), add the
SqlProbe (= Driver.Sql.SqlDriverProbe) probe via TryAddEnumerable, and add the
Driver.Sql project reference to the Host. Default tier A (SQL client is managed +
cross-platform, so ShouldStub needs no change).

Repoint the interim SqlDriver.DriverTypeName / SqlDriverFactoryExtensions.DriverTypeName
literals at DriverTypeNames.Sql; both still resolve to "Sql", so the AdminUI
TagConfigEditorMap/TagConfigValidator keys and the DriverTypeName-parity test stay
green. The Core guard test discovers factories from its own bin, so it also gains a
Driver.Sql project reference — that is the "registered-factory side" that lets
DriverTypeNamesGuardTests' bidirectional-parity check pass (4/4 green).

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 15:35:09 -04:00
Joseph Doherty fcc7ed698e harden(modbus-rtu): factory throws on unknown transport + document RTU FC-shape assumption
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 15:33:49 -04:00
Joseph Doherty a6370a26f8 fix(mqtt): make connect idempotent, bound the reconnect callback, stop State lying
Closes the connect-vs-connect defect: MQTTnet throws (and raises no DisconnectedAsync)
on a connect-while-connected, so a caller's ConnectAsync and the reconnect supervisor
corrupted each other in both directions. Both paths now check first; when the supervisor
finds the session already restored it stands down WITHOUT firing Reconnected.

Reconnected becomes Func<CancellationToken, Task>, fed from the lifetime token and capped
at ConnectTimeoutSeconds, so a hung re-subscribe can no longer park the supervisor.
Connected is published only after the re-subscribe succeeds; a supervisor that dies now
reports Faulted instead of Reconnecting forever.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 15:32:41 -04:00
Joseph Doherty 768fd87774 feat(mqtt): hand-rolled reconnect loop with bounded backoff + resubscribe
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:58:37 -04:00
Joseph Doherty cacc30a60d feat(mqtt): last-value cache backing IReadable (per-ref no-data)
LastValueCache bridges MQTT's subscribe-first push model to the OPC UA
server's polled IReadable.ReadAsync: the subscription path calls Update()
per RawPath, the read path calls Read(). Read never throws; an unseen
RawPath returns BadWaitingForInitialData (0x80320000) rather than an
exception or null, so a batch covering many references degrades per-ref
instead of failing wholesale.

Deviates from the plan's GoodNoData snippet: GoodNoData is reserved
repo-wide for "the historian window held no samples" (NullHistorianDataSource,
OtOpcUaNodeManager HistoryRead paths); BadWaitingForInitialData is the
established convention for "no live value observed yet" (CalculationDriver,
VirtualTagEngine, FOCAS, AddressSpaceApplier). Keyed by RawPath per the v3
driver-reference identity, not a topic/JSON-path-derived key.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:40:15 -04:00
Joseph Doherty 5629f3698d fix(mqtt): pin the CA accept path, honour presented intermediates, classify the dispose race
Review follow-ups on MqttConnection (Task 3):

- Every certificate test asserted rejection, so the accept branch was
  unreachable-by-regression. Adds a leaf genuinely issued by the pinned CA and
  asserts acceptance.
- ValidateAgainstPinnedCa never seeded ChainPolicy.ExtraStore from the incoming
  chain, so a leaf behind an intermediate delivered during the handshake failed
  despite a legitimate path to the pinned root. Seeds from both the incoming
  chain's elements and its ExtraStore; CustomRootTrust still means only the
  pinned roots may terminate the chain.
- A DisposeAsync racing an in-flight connect escaped as an unclassified
  exception; it now folds into ObjectDisposedException.
- Promotes the single-caller concurrency invariant into the type remarks, with
  the accurate blast radius (a leaked live connection, not a benign throw).
  Serialising the lifecycle remains Task 4's job.
- X509Chain.Build can throw; an exception escaping a TLS validation callback is
  an opaque handshake crash, so it is caught and refused.
- Adds a connect-retry test (Task 4's reconnect loop reuses the instance) and a
  disposed-then-connect test.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:38:16 -04:00
Joseph Doherty 4a8c9badce docs(modbus-rtu): note the co-running :5021 rtu_over_tcp fixture in Docker README + Dockerfile
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:36:00 -04:00
Joseph Doherty ded4c41798 fix(mqtt): key tag definitions by RawPath, not the TagConfig blob
Task 2 review follow-up. The plan specified MqttEquipmentTagParser.TryParse(reference)
with def.Name = the TagConfig blob, and told us to mirror a type named
ModbusEquipmentTagParser. Both are plan defects:

- EquipmentTagRefResolver documents that a v3 driver reference "is now always a
  RawPath" and that the blob-parse fallback is retired. Keying Name by the blob
  would make OnDataChange publish under a reference that never matches the
  RawPath-keyed fan-out in DriverHostActor - silently dead in production with
  every unit test still green.
- ModbusEquipmentTagParser does not exist. The six sibling drivers all use
  <Driver>TagDefinitionFactory.FromTagConfig(tagConfig, rawPath, out def).

Changes:
- Rename MqttEquipmentTagParser -> MqttTagDefinitionFactory; TryParse(reference,
  out def) -> FromTagConfig(tagConfig, rawPath, out def) setting Name: rawPath,
  matching ModbusTagDefinitionFactory's structure, param docs and guard order.
- Pin the identity contract with a dedicated test so a regression to blob-keying
  goes red.
- Read qos with the same strictness as the enums: a present-but-invalid qos
  ("high" / 1.5 / 5 / null) now rejects the tag and is warned by Inspect,
  instead of being silently absorbed into the driver-level default and handing
  the operator a weaker delivery guarantee than the one they authored.
- No ToTagConfig inverse: the siblings carry one solely for their Driver.<X>.Cli
  project, the MQTT plan defines none, and the AdminUI editor template
  references no driver factory. Recorded as an explicit YAGNI call in the type
  doc rather than added speculatively.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:31:36 -04:00
Joseph Doherty 650e27075f test(modbus-rtu): add rtu_over_tcp pymodbus fixture + RTU-over-TCP integration tests
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:29:10 -04:00
Joseph Doherty 726d6d198e feat(mqtt): MqttConnection connect/TLS/auth under bounded deadline
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:24:04 -04:00
Joseph Doherty 132e7b63f6 feat(modbus-rtu): add Transport selector to the AdminUI Modbus driver form
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:19:15 -04:00
Joseph Doherty 13bd9820b0 feat(modbus-rtu): route driver + probe transports through ModbusTransportFactory
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:12:12 -04:00
Joseph Doherty 1d90bf3611 feat(modbus-rtu): add ModbusTransportFactory switch on Transport mode
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:08:10 -04:00
Joseph Doherty e3fea6e409 feat(mqtt): plain tag parser + strict-enum descriptor
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:07:11 -04:00
Joseph Doherty 20be5416b9 feat(modbus-rtu): bind Transport mode on options/DTO/factory (string-enum guard)
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:04:34 -04:00
Joseph Doherty a5a8710af5 fix(mqtt): pin security-critical defaults with tests + redact password in ToString
Code review of f22db5d8 (approved-with-findings): the existing round-trip
tests only exercised explicit JSON payloads, so nothing failed if UseTls /
AllowUntrustedServerCertificate or the §5.1 numeric defaults were flipped.
Adds a defaults-pinning test plus a partial-JSON test that omits the TLS
knobs entirely (the production case — an operator config that just doesn't
mention TLS must not silently land insecure).

Also overrides the record's PrintMembers so ToString() no longer prints
Password in plaintext (verified RED before the fix: the new test failed
with the raw password in the rendered string). OpcUaClientDriverOptions has
the same unredacted-ToString() shape but is out of scope for this task per
the coordinator's note; flagging as a possible follow-up.

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 14:01:01 -04:00
Joseph Doherty 47e9cd56ef feat(modbus-rtu): add ModbusRtuOverTcpTransport (RTU framing over socket lifecycle)
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 13:56:19 -04:00
Joseph Doherty f22db5d801 feat(mqtt): Contracts options DTO + enums (name-serialized)
Task 1 of the MQTT/Sparkplug B driver plan: MqttMode, MqttPayloadFormat,
MqttProtocolVersion enums + MqttDriverOptions (broker conn + mode +
nullable Sparkplug/Plain sub-options) per design doc §5.1. Removes the
Task 0 temporary MQTTnet PackageReference from .Contracts (transport-free;
references only Core.Abstractions). MQTTnet stays pinned in
Directory.Packages.props for the .Driver project (Task 3).

Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 13:53:06 -04:00
Joseph Doherty 2f38b5b285 refactor(modbus): extract ModbusSocketLifecycle from ModbusTcpTransport (behaviour-preserving)
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 13:47:18 -04:00
Joseph Doherty d677ac7ad3 feat(mqtt): validate MQTTnet-5/net10 restore under central pinning (spike, P1 gate)
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 13:45:21 -04:00
Joseph Doherty 20100c36ad feat(modbus-rtu): validate response unit + cover partial-read/truncation in ModbusRtuFraming
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 13:42:25 -04:00
Joseph Doherty eed6617784 feat(modbus-rtu): add ModbusRtuFraming (ADU build + FC-aware length-less deframe)
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 13:33:59 -04:00
Joseph Doherty 8d0f60ec51 feat(modbus-rtu): add ModbusCrc CRC-16/MODBUS helper + golden vectors
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 13:30:05 -04:00
Joseph Doherty e787aa572b feat(modbus-rtu): add ModbusTransportMode enum (Tcp | RtuOverTcp)
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
2026-07-24 13:28:30 -04:00
269 changed files with 52736 additions and 694 deletions
+51 -3
View File
@@ -55,6 +55,28 @@ about OtOpcUa changes here — remote/push status, the driver set, the Galaxy da
shared-lib consumption, or per-project commands — update the **OtOpcUa entry in
`../scadaproj/CLAUDE.md`** in the same change so the index never drifts from this repo.
**The index edit belongs on `scadaproj`'s `main`, and must be pushed.** `scadaproj` is a
separate repo with its own checkout, so committing there inherits whatever branch it
happens to be on — which is usually *not* `main` and is often an unrelated feature branch
with no upstream. That silently drifts the index for as long as that branch stays
unmerged, which is the exact failure the rule above exists to prevent. Do this explicitly:
```bash
cd ~/Desktop/scadaproj
git rev-parse --abbrev-ref HEAD # note it, to restore afterwards
git stash list # never commit onto a dirty unrelated branch
git checkout main && git pull
# …edit the OtOpcUa entry in CLAUDE.md…
git commit -am "docs(index): <what changed about OtOpcUa>"
git push origin main
git checkout - # restore the original branch
```
Observed 2026-07-27: the `Sql` poll driver's index entry had sat unpushed on an unrelated
feature branch since 2026-07-24 because it was committed onto the current checkout, so the
index disagreed with reality for three days despite this rule. If you find index commits
stranded on a feature branch, cherry-pick them onto `main` rather than merging that branch.
## Architecture Overview
### Data Flow
@@ -168,13 +190,29 @@ central SQL Server.
> recipes (S7 blackhole, GLAuth outage, HistorianGateway LiveIntegration), the integration-test
> harness, the docker-dev rig, and the deploy setup. Start there; the sections below + `docs/v2/dev-environment.md` carry the detail.
> **Migrated 2026-04-28**: Docker config + host moved off this dev VM (DESKTOP-6JL3KKO) onto the shared Linux Docker host (`DOCKER`, 10.100.0.35) so the dev VM could shed WSL2/Hyper-V and have its GPU re-attached via ESXi passthrough. Docker Desktop is no longer installed here. All checked-in `appsettings.json` defaults, fixture-class default endpoints, and `e2e-config.sample.json` were rewritten to target `10.100.0.35`. The driver fixture compose files under `tests/.../Docker/docker-compose.yml` now carry a `project: lmxopcua` label on every service. See `docs/v2/dev-environment.md` for the full rewrite (header dated 2026-04-28).
> **Migrated 2026-04-28**: Docker config + host moved off this dev VM (DESKTOP-6JL3KKO) onto the shared Linux Docker host (`DOCKER`, 10.100.0.35) so the dev VM could shed WSL2/Hyper-V and have its GPU re-attached via ESXi passthrough. Docker Desktop is no longer installed here. All checked-in `appsettings.json` defaults, fixture-class default endpoints, and `e2e-config.sample.json` were rewritten to target `10.100.0.35`. See `docs/v2/dev-environment.md` for the full rewrite (header dated 2026-04-28).
> ⚠️ **The `project: lmxopcua` label is aspirational, not universal.** This note used to claim every fixture compose file carries it. As of 2026-07-24 only the **MQTT** fixture actually does — `grep -rn "labels:" tests --include=*.yml` matches `Driver.Mqtt.IntegrationTests/Docker/docker-compose.yml` and nothing else. So `docker ps --filter label=project=lmxopcua` returns the MQTT fixture alone, not the fleet. Add the label when you touch an older fixture; until then enumerate by container name.
Docker workloads run on a shared Linux host at **`10.100.0.35`** — not on this VM. Stacks live at `/opt/otopcua-<driver>/` on the host and carry the `project=lmxopcua` label so they're discoverable via `docker ps --filter label=project=lmxopcua`.
**`docker -H ssh://...` does NOT work from this VM.** Windows OpenSSH ↔ docker.exe stdio bridging hangs (`docker system dial-stdio` runs server-side but no API data flows). Use the helper below — it SSHes into the docker host and runs `docker compose` server-side.
**`docker -H ssh://...` does NOT work from the Windows VM.** Windows OpenSSH ↔ docker.exe stdio bridging hangs (`docker system dial-stdio` runs server-side but no API data flows). Use the helper below — it SSHes into the docker host and runs `docker compose` server-side.
**Use `lmxopcua-fix.ps1` (in `~/bin`) to control fixtures from this VM:**
> **On macOS there is no `lmxopcua-fix` — drive the host over SSH directly.** The helper is a Windows-VM
> script; `~/bin` is empty on the Mac and the commands below will not resolve. Use passwordless SSH
> (and `rsync` in place of `sync`), which is what `infra/README.md` §1 documents as the Mac path:
>
> ```bash
> ssh dohertj2@10.100.0.35 'docker ps --format "{{.Names}}\t{{.Status}}"'
> ssh dohertj2@10.100.0.35 'cd /opt/otopcua-mqtt && docker compose up -d'
> rsync -av tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.IntegrationTests/Docker/ \
> dohertj2@10.100.0.35:/opt/otopcua-mqtt/ # the `sync` step, by hand
> ```
>
> A local container runtime *does* exist on the Mac (OrbStack), so the `docker-dev/` rig itself runs
> locally there — it is only the shared **fixtures** that live on `10.100.0.35`.
**Use `lmxopcua-fix.ps1` (in `~/bin`) to control fixtures from the Windows VM:**
```powershell
lmxopcua-fix ls # list all lmxopcua-tagged containers on the host
@@ -195,6 +233,16 @@ lmxopcua-fix sync modbus # rsync this repo's tests/.../Docker/
- AB CIP: `10.100.0.35:44818` (`AB_SERVER_ENDPOINT`)
- S7: `10.100.0.35:1102` (`S7_SIM_ENDPOINT`)
- OPC UA reference (opc-plc): `opc.tcp://10.100.0.35:50000` (`OPCUA_SIM_ENDPOINT`)
- MQTT (Mosquitto, **auth on both listeners — no anonymous fallback**): `10.100.0.35:8883` TLS+auth
(`MQTT_FIXTURE_ENDPOINT`) · `10.100.0.35:1883` plaintext-but-authenticated (`MQTT_FIXTURE_PLAIN_ENDPOINT`).
Also needs `MQTT_FIXTURE_USERNAME` / `MQTT_FIXTURE_PASSWORD`, and `MQTT_FIXTURE_CA_CERT` to pin the
fixture CA (`scp dohertj2@10.100.0.35:/opt/otopcua-mqtt/secrets/ca.crt`). A co-running publisher loops
JSON under `otopcua/fixture/…`.
- MQTT **Sparkplug B**: the same broker, plus the `sparkplug` compose profile's `otopcua-sparkplug-sim`
(project-owned C# edge-node simulator) — group **`OtOpcUaSim`**, edge nodes **`EdgeA`**/**`EdgeB`**,
`EdgeA` device **`Filler1`**, ~2 s cadence. It subscribes to `spBv1.0/OtOpcUaSim/NCMD/{node}` and
re-births on request. **Births are never retained**, so `docker restart otopcua-sparkplug-sim` is how
you force a fresh one while a browse window is open.
Override any endpoint via the env var to point at a real PLC. The local OtOpcUa server runs on this VM at `opc.tcp://localhost:4840`**that's not on the docker host**.
+18
View File
@@ -75,8 +75,26 @@
<PackageVersion Include="Microsoft.NET.Test.Sdk" Version="17.12.0" />
<PackageVersion Include="Microsoft.Playwright" Version="1.51.0" />
<PackageVersion Include="Moq" Version="4.20.72" />
<!--
MQTT client for the Mqtt / Sparkplug B driver. MQTTnet v5 is the first line shipping a
net10.0 TFM (lib/net10.0, and an EMPTY net10.0 dependency group — zero transitive
packages, so it cannot participate in a diamond under this repo's deliberately-OFF
CentralPackageTransitivePinningEnabled). Sparkplug B payloads are hand-rolled rather than
taken from SparkplugNet, whose newest release (1.3.10) still pins MQTTnet 4.3.6.1152 and
ships net6.0/net8.0 only.
-->
<PackageVersion Include="MQTTnet" Version="5.2.0.1603" />
<!-- Driver.MTConnect carries NO backend NuGet. The TrakHound MTConnect.NET-Common/-HTTP pins
Task 0 added were removed in Task 7 (2026-07-24): neither package can parse an MTConnect
document (the XML formatter ships in a third, unreferenced package) nor frame the /sample
multipart stream socket-free. The driver uses HttpClient + System.Xml.Linq only. See the
Task 0 CORRECTION block in docs/plans/2026-07-24-mtconnect-driver.md. -->
<PackageVersion Include="Novell.Directory.Ldap.NETStandard" Version="3.6.0" />
<PackageVersion Include="OPCFoundation.NetStandard.Opc.Ua.Client" Version="1.5.378.106" />
<!-- Core carries Opc.Ua.StatusCodes, the oracle StatusCodeParityTests checks the drivers'
hard-coded uint constants against. Referenced by that TEST project only — the drivers
themselves stay SDK-free by design and keep spelling status codes as bare uints. -->
<PackageVersion Include="OPCFoundation.NetStandard.Opc.Ua.Core" Version="1.5.378.106" />
<PackageVersion Include="OPCFoundation.NetStandard.Opc.Ua.Configuration" Version="1.5.378.106" />
<PackageVersion Include="OPCFoundation.NetStandard.Opc.Ua.Server" Version="1.5.378.106" />
<!-- OpenTelemetry.Api < 1.15.3 has GHSA-g94r-2vxg-569j (header-parsing memory DoS). The trio
+3 -3
View File
@@ -1,6 +1,6 @@
# OtOpcUa
OPC UA server (.NET 10 AnyCPU) that exposes a fleet of industrial drivers as a single OPC UA address space. Drivers ship in-process for AVEVA System Platform Galaxy (via the sibling `mxaccessgw` repo), Modbus TCP, Siemens S7, Allen-Bradley CIP (ControlLogix / CompactLogix), Allen-Bradley Legacy (SLC 500 / MicroLogix), Beckhoff TwinCAT (ADS), FANUC FOCAS, and OPC UA Client (gateway).
OPC UA server (.NET 10 AnyCPU) that exposes a fleet of industrial drivers as a single OPC UA address space. Drivers ship in-process for AVEVA System Platform Galaxy (via the sibling `mxaccessgw` repo), Modbus TCP, Siemens S7, Allen-Bradley CIP (ControlLogix / CompactLogix), Allen-Bradley Legacy (SLC 500 / MicroLogix), Beckhoff TwinCAT (ADS), FANUC FOCAS, OPC UA Client (gateway), and MQTT (broker subscribe; Sparkplug B in progress).
A cross-platform client stack (.NET 10) — shared library, CLI, and Avalonia desktop app — connects to any OPC UA server.
@@ -15,7 +15,7 @@ A cross-platform client stack (.NET 10) — shared library, CLI, and Avalonia de
| address space + capability fan-out|
+-------------------------------------+
| | | | | | | |
Galaxy Modbus S7 AbCip AbLeg TwinCAT FOCAS OpcUaClient
Galaxy Modbus S7 AbCip AbLeg TwinCAT FOCAS OpcUaClient MQTT
|
v
mxaccessgw (sibling repo, gRPC)
@@ -91,7 +91,7 @@ See [docs/Client.CLI.md](docs/Client.CLI.md) and [docs/Client.UI.md](docs/Client
|---|---|
| Driver specs (per-driver capability surface, config, addressing) | [docs/v2/driver-specs.md](docs/v2/driver-specs.md) |
| Galaxy driver | [docs/drivers/Galaxy.md](docs/drivers/Galaxy.md) |
| Modbus / S7 / AbCip / AbLegacy / TwinCAT / FOCAS / OpcUaClient | [docs/drivers/](docs/drivers/) |
| Modbus / S7 / AbCip / AbLegacy / TwinCAT / FOCAS / OpcUaClient / MQTT | [docs/drivers/](docs/drivers/) |
| Galaxy parity rig (mxaccessgw setup) | [docs/v2/Galaxy.ParityRig.md](docs/v2/Galaxy.ParityRig.md) |
| Galaxy performance + tracing | [docs/v2/Galaxy.Performance.md](docs/v2/Galaxy.Performance.md) |
+10
View File
@@ -32,6 +32,8 @@
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Sql/ZB.MOM.WW.OtOpcUa.Driver.Sql.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Sql.Browser/ZB.MOM.WW.OtOpcUa.Driver.Sql.Browser.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Sql.Contracts/ZB.MOM.WW.OtOpcUa.Driver.Sql.Contracts.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.Contracts/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.Contracts.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.S7/ZB.MOM.WW.OtOpcUa.Driver.S7.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.S7.Contracts/ZB.MOM.WW.OtOpcUa.Driver.S7.Contracts.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.AbCip/ZB.MOM.WW.OtOpcUa.Driver.AbCip.csproj" />
@@ -45,6 +47,9 @@
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient.Contracts/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient.Contracts.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient.Browser/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient.Browser.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Mqtt/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Contracts/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Contracts.csproj" />
<Project Path="src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Browser/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Browser.csproj" />
</Folder>
<Folder Name="/src/Drivers/Driver CLIs/">
<Project Path="src/Drivers/Cli/ZB.MOM.WW.OtOpcUa.Driver.Cli.Common/ZB.MOM.WW.OtOpcUa.Driver.Cli.Common.csproj" />
@@ -96,6 +101,8 @@
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Sql.Tests/ZB.MOM.WW.OtOpcUa.Driver.Sql.Tests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Sql.Browser.Tests/ZB.MOM.WW.OtOpcUa.Driver.Sql.Browser.Tests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Sql.IntegrationTests/ZB.MOM.WW.OtOpcUa.Driver.Sql.IntegrationTests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.Tests/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.Tests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.IntegrationTests/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.IntegrationTests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.S7.Tests/ZB.MOM.WW.OtOpcUa.Driver.S7.Tests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.S7.IntegrationTests/ZB.MOM.WW.OtOpcUa.Driver.S7.IntegrationTests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.AbCip.Tests/ZB.MOM.WW.OtOpcUa.Driver.AbCip.Tests.csproj" />
@@ -110,6 +117,9 @@
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient.IntegrationTests/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient.IntegrationTests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient.Browser.Tests/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient.Browser.Tests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient.Browser.IntegrationTests/ZB.MOM.WW.OtOpcUa.Driver.OpcUaClient.Browser.IntegrationTests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Tests/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Tests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.IntegrationTests/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.IntegrationTests.csproj" />
<Project Path="tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.IntegrationTests/SparkplugSimulator/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.SparkplugSimulator.csproj" />
</Folder>
<Folder Name="/tests/Drivers/Driver CLIs/">
<Project Path="tests/Drivers/Cli/ZB.MOM.WW.OtOpcUa.Driver.Cli.Common.Tests/ZB.MOM.WW.OtOpcUa.Driver.Cli.Common.Tests.csproj" />
+34
View File
@@ -0,0 +1,34 @@
# Isolated MQTT/Sparkplug live-gate overlay (Task 26).
#
# Runs the MAIN central pair ONLY, from this worktree's image, under its own
# compose project + offset host ports, so the shared `otopcua-dev` rig and the
# sibling driver worktrees that share `otopcua-host:dev` are untouched.
#
# docker build -f docker-dev/Dockerfile --target runtime -t otopcua-host:mqtt .
# docker compose -p otopcua-mqtt -f docker-dev/docker-compose.yml \
# -f docker-dev/docker-compose.mqtt.yml \
# up -d sql migrator cluster-seed central-1 central-2
#
# AdminUI -> http://localhost:9210 (login disabled; auto-admin)
# OPC UA -> opc.tcp://localhost:4850 (central-1) / :4851 (central-2)
# SQL -> localhost,14350
#
# `!override` on every `ports` list: compose MERGES list-valued keys by
# appending, so a plain re-declaration keeps the base file's "14330:1433" /
# "4840:4840" and collides with the running `otopcua-dev` rig.
services:
sql:
ports: !override
- "14350:1433"
central-1:
image: otopcua-host:mqtt
ports: !override
- "4850:4840"
- "9210:9000"
central-2:
image: otopcua-host:mqtt
ports: !override
- "4851:4840"
+45
View File
@@ -0,0 +1,45 @@
# Isolated MTConnect live-gate overlay (Task 21).
#
# Runs the MAIN central pair ONLY, from this worktree's image, under its own
# compose project + offset host ports, so the shared `otopcua-dev` rig and the
# three sibling driver worktrees that share `otopcua-host:dev` are untouched.
# Mirrors how the mqtt-sparkplug worktree isolated itself (`otopcua-host:mqtt`,
# project `otopcua-mqtt`, ports 9210/4850).
#
# docker build -f docker-dev/Dockerfile --target runtime -t otopcua-host:mtconnect .
# docker compose -p otopcua-mtc -f docker-dev/docker-compose.yml \
# -f docker-dev/docker-compose.mtconnect.yml \
# up -d sql migrator cluster-seed central-1 central-2
#
# AdminUI -> http://localhost:9220 (login disabled; auto-admin)
# OPC UA -> opc.tcp://localhost:4860 (central-1) / :4861 (central-2)
# SQL -> localhost,14340
#
# Both MAIN nodes run because `cluster-seed` declares MAIN as a 2-node cluster:
# ConfigPublishCoordinator sources its expected-ack set from enabled ClusterNode
# rows, so deploying with one node down fails at the apply deadline.
#
# Traefik is deliberately not started — central-1 publishes its AdminUI port
# directly, which removes the :9200 round-robin that otherwise makes it
# ambiguous which node served a request.
# NOTE: `!override` is required on every `ports` list. Compose MERGES list-valued
# keys by appending, so a plain re-declaration would keep the base file's
# "14330:1433" / "4840:4840" alongside the offsets and collide with the running
# `otopcua-dev` rig ("Bind for 0.0.0.0:14330 failed: port is already allocated").
services:
sql:
ports: !override
- "14340:1433"
central-1:
image: otopcua-host:mtconnect
ports: !override
- "4860:4840"
- "9220:9000"
central-2:
image: otopcua-host:mtconnect
ports: !override
- "4861:4840"
+12
View File
@@ -272,6 +272,12 @@ services:
OTOPCUA_ROLES: "admin,driver,cluster-MAIN"
ASPNETCORE_URLS: "http://+:9000"
ConnectionStrings__ConfigDb: "Server=sql,1433;Database=OtOpcUa;User Id=sa;Password=OtOpcUa!Dev123;TrustServerCertificate=True;"
# Sql poll driver — the named connection string a Sql driver instance references by
# `connectionStringRef: "DevSql"`. Resolved in-process by SqlConnectionStringResolver (driver
# node) and SqlDriverBrowser (AdminUI browse), env-only, never persisted to ConfigDb. Points at
# the rig's own `sql` container, a SEPARATE `DevSql` database seeded with dbo.TagValues (KeyValue)
# + dbo.LatestStatus (WideRow). Committed-dev-secret exception, same as ConfigDb above.
Sql__ConnectionStrings__DevSql: "Server=sql,1433;Database=DevSql;User Id=sa;Password=OtOpcUa!Dev123;TrustServerCertificate=True;"
Cluster__Hostname: "0.0.0.0"
Cluster__Port: "4053"
Cluster__PublicHostname: "central-1"
@@ -391,6 +397,12 @@ services:
OTOPCUA_ROLES: "admin,driver,cluster-MAIN"
ASPNETCORE_URLS: "http://+:9000"
ConnectionStrings__ConfigDb: "Server=sql,1433;Database=OtOpcUa;User Id=sa;Password=OtOpcUa!Dev123;TrustServerCertificate=True;"
# Sql poll driver — the named connection string a Sql driver instance references by
# `connectionStringRef: "DevSql"`. Resolved in-process by SqlConnectionStringResolver (driver
# node) and SqlDriverBrowser (AdminUI browse), env-only, never persisted to ConfigDb. Points at
# the rig's own `sql` container, a SEPARATE `DevSql` database seeded with dbo.TagValues (KeyValue)
# + dbo.LatestStatus (WideRow). Committed-dev-secret exception, same as ConfigDb above.
Sql__ConnectionStrings__DevSql: "Server=sql,1433;Database=DevSql;User Id=sa;Password=OtOpcUa!Dev123;TrustServerCertificate=True;"
Cluster__Hostname: "0.0.0.0"
Cluster__Port: "4053"
Cluster__PublicHostname: "central-2"
+22
View File
@@ -106,6 +106,28 @@ IF NOT EXISTS (SELECT 1 FROM dbo.ClusterNode WHERE NodeId = 'site-b-2:4053')
(NodeId, ClusterId, Host, OpcUaPort, DashboardPort, AkkaPort, GrpcPort, ApplicationUri, ServiceLevelBase, Enabled, CreatedBy)
VALUES ('site-b-2:4053', 'SITE-B', 'site-b-2', 4840, 8081, 4053, 4056, 'urn:OtOpcUa:site-b-2', 150, 1, 'docker-dev-seed');
------------------------------------------------------------------------------
-- Backfill: GrpcPort on rows that predate the column (#493)
--
-- Every insert above is INSERT-if-not-exists, which is the right shape for a re-runnable seed
-- but means a *new nullable* column never reaches rows created by an earlier version of this
-- file. On a persisted volume that is exactly what happened when Phase 5 added GrpcPort: the
-- six ClusterNode rows already existed, so they kept NULL, TelemetryNodeSource found no dial
-- targets, and TelemetryDial:Mode=Grpc silently connected to nothing (throttled Warning per
-- node — the code degrades gracefully, so nothing failed loudly).
--
-- Backfills NULL only. It must not overwrite a port an operator set deliberately on a
-- long-lived dev volume; a NULL is the unambiguous "this row predates the column" marker.
--
-- AkkaPort deliberately has no equivalent here: it is NOT NULL with a default of 4053, so
-- pre-existing rows already carry a usable value and there is nothing to backfill.
------------------------------------------------------------------------------
UPDATE dbo.ClusterNode
SET GrpcPort = 4056 -- matches Telemetry__GrpcListenPort on all six nodes in docker-compose.yml
WHERE GrpcPort IS NULL
AND CreatedBy = 'docker-dev-seed';
------------------------------------------------------------------------------
-- Galaxy MxAccess gateway — INTENTIONALLY NOT SEEDED (retired 2026-06-15)
--
+39
View File
@@ -102,6 +102,45 @@ Persisted scope per plan decision #14: `Enabled`, `Acked`, `Confirmed`, `Shelvin
Every mutation the state machine produces is immediately persisted inside the engine's `_evalGate` semaphore, so the store's view is always consistent with the in-memory state.
### Post-(re)load condition-node re-assert (#487)
A deploy that triggers a **full address-space rebuild** clears `_alarmConditions` and re-materialises every
condition node fresh — inactive, acked, confirmed, unshelved. The engine, reloading in parallel, restores each
alarm's persisted state and re-derives `Active` from the predicate, so it ends up *correct*; but a still-active
alarm has **no transition to report**, so `LoadAsync` yields `EmissionKind.None` and `OnEngineEmission` drops it.
Nothing then writes the node. Before this fix, an active alarm with static dependencies read **normal** to every
OPC UA client after such a deploy, until its next real transition — which for a static-dependency alarm may
never come.
`ScriptedAlarmHostActor.ReassertConditionNodes` closes it. After each load completes it reads
`ScriptedAlarmEngine.GetProjections()` — a plain read returning `ScriptedAlarmProjection` (condition state +
the definition's severity + the resolved message template + last-observed worst-input quality) — and sends one
`OpcUaPublishActor.AlarmStateUpdate` per alarm that is **not** in the Part 9 no-event position.
Four properties are load-bearing:
- **Node-only.** It sends `AlarmStateUpdate` and nothing else — never the `alerts` topic, never the telemetry
hub. `alerts` is the historization path, so an alerts row per alarm per deploy would append a duplicate
historian / AVEVA record every time anyone deploys. This is why the re-assert reads a projection instead of
re-emitting: `ScriptedAlarmProjection` carries no `EmissionKind`, so there is nothing for the emission
machinery — or a future refactor — to mistake for a transition.
- **Only non-normal alarms.** An alarm in the no-event position (enabled, inactive, acked, confirmed,
unshelved) already matches what the materialise built, so it is skipped. That keeps a never-fired alarm
byte-identical to a deploy without this pass — including its Message, which the materialise seeds with the
alarm's display name — and writes nothing at all on the common all-normal deploy.
- **Ordering.** `DriverHostActor` tells `RebuildAddressSpace` to the publish actor *before* it tells
`ApplyScriptedAlarms` to the host, and the re-assert runs later still (after an awaited `LoadAsync` pipes
back). Both are local `Tell`s, which enqueue synchronously, so the re-materialise is already queued ahead of
these writes.
- **No duplicate Part 9 events.** `WriteAlarmCondition`'s delta-gate compares the incoming snapshot against
the node's *current* state: a re-assert onto a freshly materialised (normal) node is a genuine delta and
fires exactly one condition event — correct, since the rebuild reset what clients see — while a re-assert
onto a surgically-preserved node that already holds the same state is suppressed. The write is ungated by
redundancy role, like every other node write here; the Secondary keeps its address space warm for failover.
The timestamp carried is the condition's own `LastTransitionUtc`, **not** the deploy instant, so a client
reading `Time` / `ReceiveTime` after a rebuild still sees when the alarm actually went active.
## Source integration
`ScriptedAlarmSource` (`ScriptedAlarmSource.cs`) adapts the engine to the driver-agnostic `IAlarmSource` interface. The existing `AlarmSurfaceInvoker` + `GenericDriverNodeManager` fan-out consumes it the same way it consumes Galaxy / AB CIP / FOCAS sources — there is no scripted-alarm-specific code path in the server plumbing. From that point on, the flow into `AlarmConditionState` nodes, the `AlarmAck` session check, and the Historian sink is shared — see [AlarmTracking.md](AlarmTracking.md).
+33
View File
@@ -128,6 +128,39 @@ transports. ScadaBridge itself has since closed that gap with this identical
interceptor pattern, so Phase 5 ships authenticated from the start rather than deferring auth to a
later hardening pass. See the design doc's superseded note for the full history.
## Upgrading an existing deployment: populate `ClusterNode.GrpcPort` first (#493)
`GrpcPort` is a **nullable column added by Phase 5**, and nothing backfills it onto rows that
already existed. On an upgraded deployment every `ClusterNode` row therefore carries `NULL`,
central finds **no dial targets**, and flipping `TelemetryDial:Mode` to `Grpc` connects to
nothing. Fresh installs are unaffected. `AkkaPort` did not have this problem only because it is
`NOT NULL` with a default of 4053.
Nothing fails loudly, which is what makes this worth stating up front — the code degrades
gracefully, once per node, throttled:
```
ClusterNode <id> has no GrpcPort; it exposes no telemetry stream and is skipped in this dial refresh
(further skips of this node are silent)
```
Other symptoms: no `ESTABLISHED` sockets on the telemetry port (a `LISTEN` but nothing else in
`docker exec <node> cat /proc/net/tcp6`), and the `/alerts` / `/script-log` pill never turning
live in `Grpc` mode.
**Before flipping any node to `Grpc`**, set each driver node's `GrpcPort` to that node's own
`Telemetry:GrpcListenPort` — via the AdminUI node editor, or directly:
```sql
UPDATE dbo.ClusterNode SET GrpcPort = <that node's Telemetry:GrpcListenPort> WHERE NodeId = '<node>';
```
There is deliberately **no EF data migration** backfilling this. The correct value is that node's
own listener port, which is per-deployment configuration the database cannot derive; a migration
could only invent one, and a *wrong* dial target is worse than the `NULL` the dialer already skips
explicitly. The docker-dev seed does backfill (`GrpcPort IS NULL AND CreatedBy = 'docker-dev-seed'`
→ 4056) because there the port is fixed by `docker-compose.yml` and therefore actually knowable.
## Reconnect and reliability
Central's `TelemetryDialSupervisor` (an admin-role actor) runs **one reconnecting dialer per
+344
View File
@@ -0,0 +1,344 @@
# MTConnect Driver
Getting-started guide for the MTConnect Agent driver (P1 Agent MVP). This is
the short path — for the full design rationale read
[`docs/plans/2026-07-15-mtconnect-driver-design.md`](../plans/2026-07-15-mtconnect-driver-design.md)
and the build-vs-plan record in
[`docs/plans/2026-07-24-mtconnect-driver.md`](../plans/2026-07-24-mtconnect-driver.md); for the
fixture recipe read
[`tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.IntegrationTests/Docker/README.md`](../../tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.IntegrationTests/Docker/README.md)
(not duplicated here).
> **Live gate:** see the LIVE-GATE RESULT note in the plan's Task 21
> (`docs/plans/2026-07-24-mtconnect-driver.md`) for the docker-dev `/run` verification outcome
> (browse picker / typed editor / deploy / read / subscribe against the real Agent fixture).
## What it talks to
A **MTConnect Agent** — the vendor-neutral read-only telemetry endpoint that
front-ends a machine tool (or fronts an adapter that itself talks to the
machine over SHDR). The driver speaks plain HTTP + XML against the Agent's
three standard REST paths:
- `/probe` — the static device model (Device → Component → DataItem tree)
- `/current` — a snapshot of every DataItem's latest observation
- `/sample` — a `multipart/x-mixed-replace` long-poll stream of observation
deltas, keyed by a monotonic sequence number
v1 is **agent-first and read-only** — Discover + Read + Subscribe, no Write —
because the mainstream Agent surface has no "set value" operation.
## Built vs. planned — read this before trusting the design doc's §2
The design doc ([`2026-07-15-mtconnect-driver-design.md`](../plans/2026-07-15-mtconnect-driver-design.md))
picked the **TrakHound `MTConnect.NET-Common` / `-HTTP`** NuGet packages (MIT,
`netstandard2.0`) as the primary path, with a hand-rolled fallback. **The
hand-rolled fallback is what shipped.** Task 6/7 of the implementation plan
found by reflection + live invocation that the pinned TrakHound packages
ship **no XML formatter** (`Document Formatter Not found for "xml"` — the
formatter lives in a separate `MTConnect.NET-XML` package the design didn't
pin) and that `MTConnectHttpClientStream` exposes no injectable `HttpClient`
handler, so it cannot be unit-tested behind the driver's `IMTConnectAgentClient`
seam. Both package references were dropped. There is **no TrakHound
dependency anywhere in the shipped driver** — probe/current/sample parsing is
hand-rolled `System.Xml.Linq` plus a multipart boundary reader, entirely
behind `IMTConnectAgentClient`. See the CORRECTION block under Task 0 of the
implementation plan for the full record.
**Namespace-agnostic parsing.** Every parser matches XML elements on
`LocalName` only, ignoring the document's XML namespace entirely, so
MTConnect 1.3 through 2.x documents all parse without a schema-version
branch. This is deliberate, not sloppy: a real Agent injects vendor-extension
`Component`s in a foreign namespace whose child `DataItems` elements inherit
the *default* namespace, and namespace-strict matching would silently drop
the whole vendor component (and every DataItem beneath it) rather than fail
loudly. Attributes are read **unqualified** for the same reason — so an
`xsi:type` attribute is never mistaken for a DataItem's own `type`.
## Project split
| Project | Target | Role |
|---------|--------|------|
| `src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect/` | net10.0 | In-process driver — `MTConnectDriver`, the hand-rolled `IMTConnectAgentClient` (probe/current/sample), the observation index, the factory |
| `src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.Contracts/` | net10.0 | `MTConnectDriverOptions`, `MTConnectTagDefinition`, and the pure `MTConnectDataTypeInference` table shared by the driver, browse-commit, and the AdminUI typed editor |
No `.Browser` project — browse comes free from the Wave-0 universal
discovery browser (see [Browse](#browse--free-via-the-universal-browser)
below).
## Capability surface
`MTConnectDriver : IDriver, IReadable, ISubscribable, ITagDiscovery, IHostConnectivityProbe, IRediscoverable`
(`src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect/MTConnectDriver.cs`).
Deliberately **not** `IWritable` — the Agent surface is read-only by design.
Write-back exists only via optional, rarely-deployed MTConnect *Interfaces*
(a request/response handshake, not a "set value" operation), and is out of
scope for this build — see [Deferred](#deferred-not-in-this-build).
| Capability | Path | Notes |
|------------|------|-------|
| `ITagDiscovery` | `DiscoverAsync` — streams `/probe`'s Device→Component→DataItem tree into the address-space builder | `SupportsOnlineDiscovery = true`, `RediscoverPolicy = Once` |
| `IReadable` | `ReadAsync` → one `/current` per call | **Not the production data path** — see below |
| `ISubscribable` | `SubscribeAsync`/`OnDataChange` — the shared `/sample` long-poll pump | **The production data path** |
| `IHostConnectivityProbe` | periodic `/probe` under `Probe.*` | No consumer wires it today — see [Known limitations](#known-limitations) |
| `IRediscoverable` | watches the Agent's `Header.instanceId` | No consumer wires it today — see [Known limitations](#known-limitations) |
**`IReadable.ReadAsync` is never called by the running server.** The
production data plane is entirely `ISubscribable``DriverInstanceActor`
subscribes once and lives off `OnDataChange`. `ReadAsync` exists for the
Client CLI (`... read -n ...`) and for the unit/integration suites; it is
correct and tested, just not on the hot path.
## Minimum deployment
```jsonc
"Drivers": {
"mtconnect-1": {
"Type": "MTConnect",
"Config": {
"AgentUri": "http://10.100.0.35:5000",
"DeviceName": null,
"RequestTimeoutMs": 5000,
"SampleIntervalMs": 1000,
"SampleCount": 1000,
"HeartbeatMs": 10000,
"Probe": { "Enabled": true, "IntervalMs": 5000, "TimeoutMs": 2000 },
"Reconnect": { "MinBackoffMs": 0, "MaxBackoffMs": 30000, "BackoffMultiplier": 2.0 },
"Tags": []
}
}
}
```
`RawTags[]` is not authored by hand — `DriverDeviceConfigMerger` injects it
at deploy time from the tags authored on the `/raw` tree.
### Config keys
Read off `MTConnectDriverOptions` / `MTConnectDriverConfigDto`
(`src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.Contracts/MTConnectDriverOptions.cs`,
`src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect/MTConnectDriver.cs`):
| Key | Default | Notes |
|---|---|---|
| `AgentUri` | — (required) | Agent base URI; the driver appends `/probe`, `/current`, `/sample` |
| `DeviceName` | `null` | Scopes requests to `{AgentUri}/{DeviceName}/...`; `null` = whole Agent |
| `RequestTimeoutMs` | `5000` | Per-call deadline for `/probe` and `/current` |
| `SampleIntervalMs` | `1000` | The Agent's `/sample?interval=` query param |
| `SampleCount` | `1000` | The Agent's `/sample?count=` query param |
| `HeartbeatMs` | `10000` | The Agent's `/sample?heartbeat=` query param — also the watchdog's liveness window |
| `Probe.Enabled` / `Probe.IntervalMs` / `Probe.TimeoutMs` | `true` / `5000` / `2000` | Background connectivity-probe knobs (mirrors `ModbusProbeOptions`) |
| `Reconnect.MinBackoffMs` / `MaxBackoffMs` / `BackoffMultiplier` | `0` / `30000` / `2.0` | Geometric backoff after a failed request or a dropped `/sample` stream |
| `Tags[]` | `[]` | Pre-v3 / CLI authoring surface — one `MTConnectTagDefinition` per DataItem `id` |
| `RawTags[]` | `[]` (deploy-injected) | The v3 data-plane binding — see below |
**All timing knobs are validated strictly positive** at `InitializeAsync`
(`RequirePositive`, mirroring the arch-review 01/S-6 lesson: a `0` timeout
does not mean "wait forever," it faults the driver). This applies to
`RequestTimeoutMs`, `HeartbeatMs`, `SampleIntervalMs`, `SampleCount`, and
(when the probe is enabled) `Probe.Interval` / `Probe.Timeout`.
**No save-time gate on a blank `AgentUri` in the AdminUI.** The driver form
renders an inline validation notice (`_form.Validate()`), but
`DriverConfigModal.SaveAsync` has no per-form validation seam to block the
Save button on it — no sibling driver form has one either. A blank
`AgentUri` saves cleanly and fails at deploy with the driver's own error
message, not at authoring time.
## Data plane
### FullName == DataItem `id`
`FullName` (both `MTConnectTagDefinition.FullName` and every `RawTags[]`
blob's identifier) is the MTConnect `DataItem@id` attribute — the value the
universal browser commits and the key the driver resolves reads/subscribes
against. `MTConnectTagConfigModel.FromJson` (the AdminUI typed editor) tries
three spellings in order — `fullName``dataItemId``address` — and takes
the first non-blank one, normalising onto `fullName` on save. `address` is
accepted because that's the field name `RawBrowseCommitMapper` writes for a
driver with no typed editor path, so a browse-committed tag stays readable
even before the editor round-trips it.
### Coercion-type precedence
The driver keys its data plane by **RawPath** (the v3 raw-tag identity), not
by DataItem id directly, resolving RawPath → dataItemId internally. Each raw
tag's coercion type (the `DriverDataType` its Agent observation is parsed
into) is resolved in this order, first non-null wins:
1. The `RawTags[]` blob's `driverDataType` or `dataType` field (both
spellings accepted — the typed editor writes `dataType`, the driver's own
`tags[]` shape uses `driverDataType`).
2. A matching `Tags[]` entry naming the same DataItem id.
3. The Agent's own `/probe` declaration, via `MTConnectDataTypeInference.Infer`
(category/type/units/representation → `DriverDataType`).
4. `DriverDataType.String` — the coercion that cannot fail.
### Quality mapping
`UNAVAILABLE` — MTConnect's one explicit "I have no value" sentinel — maps to
**`BadNoCommunication`** (`0x80310000`). This is deliberately distinct from
the fleet-standard `BadCommunicationError` (`0x80050000`, used elsewhere for
the driver's *own* transport failure to the Agent): `UNAVAILABLE` means the
Agent is reachable and answered, it just has no device-backed value for that
item. A `CONDITION` observation's value is taken from its **element name**
(`<Normal/>` → the string `"Normal"`, `<Fault/>``"Fault"`); a `<Unavailable/>`
condition element normalises onto the same `UNAVAILABLE` sentinel as every
other category so one comparison covers all three.
Other status codes an observation can surface:
| Code | Meaning |
|---|---|
| `BadTypeMismatch` (`0x80740000`) | The Agent's text isn't a value of the tag's coerced type at all |
| `BadOutOfRange` (`0x803C0000`) | The Agent reported a number the coerced type can't represent |
| `BadNotSupported` (`0x803D0000`) | A shape this build doesn't materialize — a `TIME_SERIES` vector, or a structured `DATA_SET`/`TABLE` observation (deferred to P1.5; see [Known limitations](#known-limitations)) |
| `BadWaitingForInitialData` | Tag is authored but the Agent hasn't reported it yet |
| `BadNodeIdUnknown` | DataItem id is neither authored nor ever observed |
### Subscribe — the `/sample` pump
One shared `/sample` long-poll stream per driver instance (the Agent streams
the whole device model regardless of which subset is subscribed, so
per-tag streams would be wasted round-trips), run under a heartbeat
watchdog — `HttpClient.Timeout` cannot bound a long-lived stream, so a
missing chunk **and** missing keep-alive heartbeat within
`HeartbeatMs × N` is what detects a frozen peer.
Ring-buffer overflow (the Agent's circular observation buffer wrapped past
what the driver's cursor expects) is detected **two ways**, both triggering a
`/current` re-baseline before the stream resumes:
1. `IMTConnectAgentClient.IsSequenceGap` — the next chunk's `firstSequence`
is newer than the driver's expected cursor.
2. An `MTConnectError` document reporting `OUT_OF_RANGE` — real Agents return
this both under HTTP 200 (a normal MTConnect error document) and as a bare
HTTP 400, and the client handles both.
**An Agent `instanceId` change means the Agent restarted**, and is checked
**before** the sequence-gap check (a restart also usually trips the gap, so
the two need disambiguating, not just OR-ing together). On a changed
`instanceId` the driver clears its cached probe model and raises
`OnRediscoveryNeeded` — see [Known limitations](#known-limitations) for why
that signal currently has no consumer.
## Browse — free via the universal browser
MTConnect ships **no bespoke browser project**. Setting
`ITagDiscovery.SupportsOnlineDiscovery => true` is the entire integration:
the Wave-0 `DiscoveryDriverBrowser` sees a driver whose `TryCreate` succeeds
and whose instance reports online discovery, renders the AdminUI **Browse**
button, constructs the driver, runs `InitializeAsync` (the `/probe` connect)
+ `DiscoverAsync` into a `CapturingAddressSpaceBuilder`, then tears the
throwaway instance down. Each captured leaf's NodeId is the DataItem `id`,
committed directly as `TagConfig.FullName` on pick. See
[`docs/plans/2026-07-15-universal-discovery-browser-design.md`](../plans/2026-07-15-universal-discovery-browser-design.md).
## CONDITION modelling (v1: plain String)
Each `CONDITION` DataItem materializes as a `String` variable whose value is
the current state word — `Normal` / `Warning` / `Fault` / `UNAVAILABLE`.
There is no native OPC UA Part 9 alarm plumbing in this build;
`DriverAttributeInfo.IsAlarm = true` is still stamped on a CONDITION leaf so
the browse side-panel flags it and a future upgrade to native alarms doesn't
need re-authoring. See [Deferred](#deferred-not-in-this-build).
## Known limitations
These are real, not placeholders — read them before relying on the driver
for anything beyond values-and-conditions.
1. **`IRediscoverable` and `IHostConnectivityProbe` have no consumer in the
server.** `DriverInstanceActor` wires only `ISubscribable.OnDataChange` and
`IAlarmSource.OnAlarmEvent`; nothing in the server subscribes to
`OnRediscoveryNeeded` or `OnHostStatusChanged` except `GalaxyDriver`
wiring its own internal `DeployWatcher` sub-component. So a restarted
Agent (a changed `instanceId`) leaves a stale address space behind an
otherwise-Healthy driver. **This is a pre-existing fleet-wide gap
affecting every driver that implements either interface** (eight drivers
besides Galaxy), not something specific to MTConnect — this build simply
surfaced it again.
2. **`RawTagEntry.DeviceName` is accepted on the wire shape but ignored for
routing.** One driver instance owns exactly one Agent client scoped by
`MTConnectDriverOptions.DeviceName` (the top-level config key); the
per-tag `RawTagEntry.DeviceName` field the v3 raw-tag identity carries is
never read by the driver. Two `/raw` Devices authored under one MTConnect
driver instance both resolve to the same one Agent connection. Real
per-device routing (a client per device) would be a design change, not a
bug fix.
3. **Nothing validates a raw tag's declared OPC UA `DataType` against the
blob's coercion type.** The `/probe`-sourced inference
(`MTConnectDataTypeInference`) narrows how often an author picks a wrong
type, but it doesn't close the gap — an authored `DataType` mismatched
against the actual Agent observation still surfaces as `BadTypeMismatch`
/`BadOutOfRange` at runtime rather than at authoring time.
4. **`sampleCount` is not a legal `DataItem` attribute.** It belongs on a
`TIME_SERIES` *observation*, not the device-model declaration. A real
Agent that meets `sampleCount` on a `DataItem` declaration logs `The
following keys were present and not expected: sampleCount` and **drops
the entire data item** from the served model. A live TIME_SERIES tag
therefore always resolves to `ArrayDim = null` (variable-length array) in
practice — the fixture's canned `probe.xml`/`Devices.xml`, which does
declare a `sampleCount`, is unrealistic on this point and exists only to
pin the parsing rule itself.
5. **CONDITION is a plain `String`, not a native Part 9 alarm** (see above).
6. **`TIME_SERIES` arrays and structured `DATA_SET`/`TABLE` observations
surface as `BadNotSupported`**, not an approximation — see the Quality
mapping table.
## Deferred (not in this build)
- **Write-back (MTConnect Interfaces).** Design
[§3.6](../plans/2026-07-15-mtconnect-driver-design.md#36-iwritable--not-implemented-v1):
the mainstream Agent surface is read-only by design; Interfaces is a rare,
optional request/response handshake that would mislead if modelled as
`IWritable`. Revisit only on a concrete deployment need.
- **P1.5 fast-follow** (design
[§9](../plans/2026-07-15-mtconnect-driver-design.md#9-phasing--effort)):
CONDITION → native OPC UA Part 9 alarms via `IAlarmSource` (the Galaxy
native-alarm pattern); `TIME_SERIES` SAMPLE arrays materialized as real
OPC UA arrays; EVENT controlled-vocabulary values → OPC UA enumerations.
- **P2 (on demand): SHDR adapter ingest.** A `SourceMode: "Agent" | "Shdr"`
switch that opens the raw pipe-delimited SHDR TCP socket directly. Loses
auto-discovery (no device model without an Agent in front, so
`SupportsOnlineDiscovery` would have to report `false` and tags would be
authored by hand, like Modbus). Niche; not built.
## Testing
- **Unit tests**`tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.Tests/`
(canned-XML fixtures, no network) — 491/491 at time of writing. Covers
discovery tree shape, `MTConnectDataTypeInference`, observation indexing,
`UNAVAILABLE``BadNoCommunication`, CONDITION state-word mapping,
multipart chunk framing, and ring-buffer re-baseline paging (both the
sequence-gap and the `OUT_OF_RANGE` legs).
- **Integration tests**
`tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.IntegrationTests/` — env-gated
on `MTCONNECT_AGENT_ENDPOINT` (default `http://10.100.0.35:5000`), skips
cleanly (12 Skipped) when the fixture is unreachable, 12/12 against a real
Agent. Read the fixture's own
[`Docker/README.md`](../../tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.IntegrationTests/Docker/README.md)
for the full recipe — summary only, here:
- Image is **`mtconnect/agent:2.7.0.12`** — not `mtconnect/cppagent`, which
does not exist on Docker Hub (the design's original assumption).
- The stack is **two services**: the Agent plus a stdlib-only SHDR adapter
(`Docker/adapter.py`, on `python:3.13-alpine`). An Agent with no adapter
reports every observation `UNAVAILABLE` forever, which proves nothing.
- Deployed at `/opt/otopcua-mtconnect/` on the shared docker host
(`10.100.0.35`, `project=lmxopcua` label). Endpoint
`http://10.100.0.35:5000`.
- `agent.cfg` must be **pure ASCII** — one non-ASCII byte anywhere,
including a comment, makes the config parser reject the whole file with
a bare `Failed / Stopped at line: N`.
- Bring-up: `lmxopcua-fix sync mtconnect` then `lmxopcua-fix up mtconnect`
(PowerShell helper, Windows-only) — or, directly on the docker host,
`rsync` the repo's `Docker/` dir to `/opt/otopcua-mtconnect/` and run
`docker compose up -d --wait`.
## Further reading
- [`docs/plans/2026-07-15-mtconnect-driver-design.md`](../plans/2026-07-15-mtconnect-driver-design.md) — full design: capability wiring, browse reconciliation, typed-editor spec, resilience/timeout rules, phasing
- [`docs/plans/2026-07-24-mtconnect-driver.md`](../plans/2026-07-24-mtconnect-driver.md) — the executable implementation plan and the task-by-task build record, including where it diverged from the design (TrakHound → hand-rolled)
- [`docs/plans/2026-07-24-driver-expansion-tracking.md`](../plans/2026-07-24-driver-expansion-tracking.md) — Wave-2 program tracking
- [Docker fixture README](../../tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.IntegrationTests/Docker/README.md) — the authoritative fixture recipe (seeded device model, endpoint, Mac AirPlay port-5000 gotcha)
+133
View File
@@ -0,0 +1,133 @@
# MQTT test fixture
Coverage map + gap inventory for the MQTT driver's harness.
**TL;DR:** the unit suite (581 tests) carries the mapping, indexing, Sparkplug state-machine and
failure-classification logic; the live suite (15 tests) proves the parts only a real broker can prove
— TLS, real authentication, CA pinning in both directions, retained-message seeding, that a rejected
CONNACK is actually surfaced, and the Sparkplug birth/alias/rebirth/death path against a real
edge-node simulator.
## What the fixture is
`eclipse-mosquitto:2.0.22` plus a second container of the same image running `publisher/publish.sh`
(the image ships `mosquitto_pub`, so no bespoke image is needed), plus — behind the `sparkplug`
compose profile — **`otopcua-sparkplug-sim`**, a project-owned C# Sparkplug-B edge-node simulator
(`tests/Drivers/….Driver.Mqtt.IntegrationTests/SparkplugSimulator/`). Compose lives at
`tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.IntegrationTests/Docker/`; the deployed stack is
`/opt/otopcua-mqtt` on the shared Docker host.
**Two authenticated listeners, no anonymous fallback** (`allow_anonymous false`) — deliberate: an
anonymous broker would let a broken auth config pass every test.
| Listener | Endpoint | Env var |
|---|---|---|
| TLS + auth | `10.100.0.35:8883` | `MQTT_FIXTURE_ENDPOINT` |
| plaintext + auth | `10.100.0.35:1883` | `MQTT_FIXTURE_PLAIN_ENDPOINT` |
`MqttFixture` gates the suite on `MQTT_FIXTURE_ENDPOINT` / `MQTT_FIXTURE_USERNAME` /
`MQTT_FIXTURE_PASSWORD` (+ `MQTT_FIXTURE_CA_CERT`, `MQTT_FIXTURE_TOPIC_PREFIX`) and **skips cleanly**
when they are absent, so the suite is safe to run offline.
**Unlike every other fixture it needs generated material first** — `./gen-fixture-material.sh` writes
the gitignored `secrets/` (password file, CA, server cert/key). The broker will not start without it.
Bring-up and the CA-fetch recipe are in [`infra/README.md`](../../infra/README.md) §3.
### Published topics
| Topic | Shape | Retain | Why it exists |
|---|---|---|---|
| `otopcua/fixture/retained/seed` | `{"value":42.5,…}` | ✅ **published exactly once** | The only way to prove retained *seeding*. On a republished topic a retained seed is indistinguishable from a fast live publish |
| `otopcua/fixture/line1/temperature` | `{"value":…,"unit":"C","seq":n}` | ✅ | JSON + JSONPath extraction |
| `otopcua/fixture/line1/counter` | `{"value":n}` | ✅ | Monotonic, for ordering |
| `otopcua/fixture/line1/state` | `RUNNING` (bare text) | ✅ | Non-JSON `Raw`/`Scalar` payloads |
| `otopcua/fixture/noise/chatter` | JSON | ❌ | **Never authored as a tag** — a broker carrying unauthored traffic is the normal case and the driver must stay silent about it |
## What it actually covers
The 7 live tests: TLS+auth connect; connect pinned to the fixture CA; connect pinned to a *foreign*
CA is **rejected** (the falsifiability control — a pinning test that only ever passes proves nothing);
wrong password is rejected **and surfaced** (the MQTTnet-5 silent-`Healthy` defect); plain
subscribe→value under a RawPath; retained message seeds a value on subscribe; and an unauthored topic
stays silent while an authored tag keeps flowing.
### The Sparkplug simulator
Group **`OtOpcUaSim`**; edge nodes **`EdgeA`** and **`EdgeB`**; `EdgeA` additionally publishes device
**`Filler1`**. Node metrics `Temperature` (Float), `Pressure` (Double), `Count` (Int32), `Running`
(Boolean), `Serial` (String) plus the mandatory `bdSeq` and **`Node Control/Rebirth`** — the last one
deliberately carries a `/`, so it exercises the "a metric name is not a topic segment" rule. Device
metrics: `Temperature` (Float), `FillCount` (Int64), `Jammed` (Boolean). Cadence ~2 s.
It **subscribes to `spBv1.0/OtOpcUaSim/NCMD/{node}`** and re-births on a rebirth request, so it
exercises the real NCMD round trip rather than simulating it. Births are not retained (per spec), so
`docker restart otopcua-sparkplug-sim` is the way to force a fresh birth on demand:
```bash
ssh dohertj2@10.100.0.35 'cd /opt/otopcua-mqtt && docker compose --profile sparkplug up -d'
ssh dohertj2@10.100.0.35 'docker restart otopcua-sparkplug-sim' # force a birth
ssh dohertj2@10.100.0.35 'docker logs -f otopcua-sparkplug-sim'
```
Live-gated end-to-end (2026-07-24, P1 milestone): deployed on the docker-dev rig against this fixture,
both tags resolved and served changing values through OPC UA, and the AdminUI `#`-observation picker
rendered the real topic tree. Re-gated for P2 (2026-07-25) against the Sparkplug simulator — see
[Mqtt.md](Mqtt.md).
## What it does NOT cover
### 1. Sparkplug primary-host STATE, and write-through
The simulator never observes a Host Application `STATE`, because the driver never publishes one
(`ActAsPrimaryHost` is unimplemented). Nor is there any DCMD/NCMD **command** coverage: rebirth is the
only NCMD the driver ever sends, and `IWritable` is deferred.
Sparkplug **seq-gap injection** is unit-proven only — the simulator always publishes a well-formed
monotonic `seq`, so no live test drives the gap → rebirth path end-to-end.
### 2. Broker-side failure injection
No test kills the broker mid-session, partitions the network, or forces a SUBACK failure against a
real broker. Reconnect, backoff and per-topic degradation are **unit-proven only**.
### 3. Wildcard tags against a live broker
The by-filter indexing fix that keeps wildcard tags alive across a reconnect is covered by unit tests;
no live test authors a `+`/`#` tag and bounces the connection.
### 4. QoS 2 delivery semantics
The fixture publishes at QoS 0/1. QoS 2's exactly-once handshake is never exercised end-to-end.
### 5. Scale
One publisher, five topics, a ~2 s cadence. Nothing probes the 50 000-node browse cap, the 1 MiB
payload cap, or thousands of authored tags.
### 6. Redundant-pair client-id collision
Found by hand during the P1 gate (a fixed `clientId` makes pair nodes evict each other), **not** by a
test. No automated coverage asserts that two concurrent drivers coexist.
## When to trust the MQTT fixture, when to reach for a real broker
| Question | Fixture answers it? |
|---|---|
| Does TLS + auth work, and does bad auth fail loudly? | ✅ yes |
| Is a retained message seeded on subscribe? | ✅ yes |
| Does JSONPath extraction + coercion behave? | ✅ yes (unit + live) |
| Does the driver ignore unauthored traffic? | ✅ yes |
| Does it survive a broker restart / flapping link? | ❌ unit-only |
| Does it hold up at plant tag counts? | ❌ no |
| Does a Sparkplug birth → alias → DATA → death cycle work? | ✅ yes (live, vs. the simulator) |
| Does an NCMD rebirth request actually get answered? | ✅ yes (the simulator subscribes + re-births) |
| Sparkplug seq-gap → rebirth | ❌ unit-only (the sim never skips a seq) |
| Sparkplug primary-host STATE / DCMD writes | ❌ not implemented |
## Follow-up candidates
1. Broker-restart / SUBACK-failure live tests (closes gap 2).
2. A live wildcard-across-reconnect test (closes gap 3).
3. A concurrent-two-driver test pinning the client-id hazard (closes gap 6).
4. A seq-gap injection switch on the simulator, closing the last unit-only Sparkplug leg.
## Key fixture / config files
- `tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.IntegrationTests/Docker/docker-compose.yml`
- `tests/Drivers/…/Docker/mosquitto.conf` · `Docker/gen-fixture-material.sh` · `Docker/publisher/publish.sh`
- `tests/Drivers/…/MqttFixture.cs` · `PlainMqttLiveTests.cs`
- [`infra/README.md`](../../infra/README.md) §3 — bring-up, endpoints, CA fetch
+397
View File
@@ -0,0 +1,397 @@
# MQTT Driver
In-process MQTT **broker-subscribe** driver. It holds one MQTTnet 5 client against one broker,
subscribes to the topics its authored tags name, and forwards each PUBLISH straight to
`ISubscribable.OnDataChange`. Ingest is **push, not poll**, and **one-way** — there is no `IWritable`.
**Two ingest shapes, selected by the driver's `Mode`:** `Plain` binds a tag to a concrete MQTT topic;
`SparkplugB` binds it to a `group / edge node / device? / metric` tuple and decodes Eclipse-Tahu
protobuf under one `spBv1.0/{GroupId}/#` subscription. Both run over the same single client.
See [Mqtt-Test-Fixture.md](Mqtt-Test-Fixture.md) for the harness and
[`../plans/2026-07-15-mqtt-sparkplug-driver-design.md`](../plans/2026-07-15-mqtt-sparkplug-driver-design.md)
for the design.
> **Status.** P1 (plain MQTT) and P2 (Sparkplug B ingest) are both shipped and live-gated against the
> Mosquitto + Sparkplug-simulator fixture. Write-through (`IWritable` — NCMD/DCMD/plain publish) is
> **deferred** ([#508](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/508)); every MQTT node
> materializes read-only. See [Known gaps](#known-gaps) — every entry there carries a tracking issue
> (#507#515).
## Project Layout
| Project | Holds |
|---|---|
| `Driver.Mqtt.Contracts` | Config records, enums, `MqttTagDefinition` + the pure `MqttTagDefinitionFactory`, and the one shared `MqttJson.Options`. **Deliberately transport-free** — no MQTTnet reference |
| `Driver.Mqtt` | Runtime: `MqttDriver`, `MqttConnection` (connect + reconnect supervisor), `MqttSubscriptionManager`, `LastValueCache`, factory registration, `MqttDriverProbe` |
| `Driver.Mqtt.Browser` | `MqttDriverBrowser` + `MqttBrowseSession` — the AdminUI address picker's `#`-observation session |
> `.Browser` references `.Driver` (not just `.Contracts`) — a **documented, deliberate exception**
> so browse can reuse `MqttConnection.BuildClientOptions` (TLS / CA-pin / credentials / protocol
> mapping). Its cost is pulling `Core` into the AdminUI graph. The csproj carries a boxed warning;
> do not copy the pattern into a new driver's browser. The clean fix — extract the pure
> options builder down into `.Contracts` — is tracked as
> [#512](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/512).
## Capability Surface
```csharp
public sealed class MqttDriver
: IDriver, ITagDiscovery, ISubscribable, IReadable, IHostConnectivityProbe, IRediscoverable, IAsyncDisposable, IDisposable
```
| Capability | Entry point | Notes |
|---|---|---|
| `IDriver` | `InitializeAsync` / `ReinitializeAsync` / `ShutdownAsync` / `GetHealth` | Order **register → attach → connect** is contractual. `ReinitializeAsync` applies a tag-only delta in place, and rebuilds the session when the broker identity changed |
| `ITagDiscovery` | `DiscoverAsync` | Replays **only the authored `rawTags`** into an `Mqtt` folder. `SupportsOnlineDiscovery => false`. `RediscoverPolicy` is `Once` in Plain and **`UntilStable` in Sparkplug** (half of what a Sparkplug tag needs — its datatype — is only knowable from a birth) |
| `ISubscribable` | `SubscribeAsync` / `OnDataChange` | One SUBSCRIBE per distinct authored topic. **`publishingInterval` is ignored** — MQTT is push-based; this driver neither polls nor throttles |
| `IReadable` | `ReadAsync` | Serves the last received/retained value from `LastValueCache`; never throws. A never-seen tag reads `BadWaitingForInitialData` |
| `IHostConnectivityProbe` | `GetHostStatuses` | 1-second poll of `MqttConnection.State`; `HostName` is `"{host}:{port}"` |
| `IRediscoverable` | `OnRediscoveryNeeded` | Nothing raises it in Plain mode. **Sparkplug raises it** when a birth for an *authored* scope carries a metric-name set different from the last one seen for that scope — both conditions are anti-storm, see [Rediscovery](#rediscovery-sparkplug) |
**Not implemented:** `IWritable`, `IPerCallHostResolver`, `IAlarmSource`, `IHistoryProvider`.
Materialized nodes are `SecurityClassification.ViewOnly`, `IsAlarm: false`, `IsHistorized: false`
*"MQTT ingest is one-way … ViewOnly rather than an Operate node whose writes would silently vanish."*
Driver-type string: **`Mqtt`** (`DriverTypeNames.Mqtt`).
## Configuration
Bound through `MqttJson.Options`**case-insensitive**, enums by **name**, unknown members skipped.
**The whole connection is authored on the DRIVER** (`DriverConfig`), via `MqttDriverForm` in the `/raw`
driver-config modal. Unlike Modbus / S7 / OPC UA Client, MQTT does **not** use the v3
endpoint→`DeviceConfig` split: an MQTT driver instance holds exactly one broker session, so
`MqttDeviceForm` is informational (the same shape as `GalaxyDeviceForm`) and devices under an MQTT
driver are pure organisational grouping in the RawPath.
> **Why not the device.** `DriverDeviceConfigMerger` merges a device's keys up to the top level *only
> when the driver has exactly one device*. A device-authored broker connection therefore disappears
> from the merged blob the moment a second device is added — a silent outage with nothing to see at
> deploy time. Driver-level authoring is correct for 0, 1 or N devices. A pre-form deployment that put
> the connection on a sole device still works (merge-up wins over the driver's keys) and
> `MqttDeviceForm` flags those keys so the mismatch is visible; move them to the driver.
`DriverDeviceConfigMerger` still merges driver + device and injects `rawTags`.
```jsonc
// DriverConfig — authored by MqttDriverForm. PascalCase keys, enums as names.
{
"Host": "10.100.0.35",
"Port": 8883,
"UseTls": true,
"AllowUntrustedServerCertificate": false,
"CaCertificatePath": null, // pins the chain; null = OS trust store
"Username": "otopcua",
"Password": "", // never commit
"ProtocolVersion": "V500",
"CleanSession": true,
"KeepAliveSeconds": 30,
"ConnectTimeoutSeconds": 15,
"ReconnectMinBackoffSeconds": 1,
"ReconnectMaxBackoffSeconds": 30,
"MaxPayloadBytes": 1048576,
"Mode": "Plain",
"Plain": { "TopicPrefix": "", "DefaultQos": 1 }
}
// DeviceConfig — empty; MQTT has no per-device endpoint.
{}
```
### `Sparkplug` sub-object (`Mode: "SparkplugB"`)
```jsonc
{
"Mode": "SparkplugB",
"Sparkplug": {
"GroupId": "OtOpcUaSim", // REQUIRED — the driver's whole subscription scope
"HostId": "", // Host Application identity; receive-only (see below)
"ActAsPrimaryHost": false, // NOT IMPLEMENTED — logs a warning when true
"RequestRebirthOnGap": true, // a seq gap asks the edge node to re-announce (NCMD)
"BirthObservationWindowSeconds": 15 // how long discovery collects births before calling it stable
}
}
```
| Key | Default | Notes |
|---|---|---|
| `GroupId` | `""` | **Required.** The driver subscribes exactly `spBv1.0/{GroupId}/#`. Blank ⇒ it connects, reports `Healthy`, and ingests **nothing**. It is a literal topic segment, so `/`, `+`, `#` are refused by the form |
| `HostId` | `""` | Sparkplug Host Application identity. **Receive-only** — the driver observes `STATE` and never publishes one |
| `ActAsPrimaryHost` | `false` | **Not implemented.** Setting it logs a startup **warning** rather than being silently inert — being a primary host is a three-part contract (retained `ONLINE`, a matching `OFFLINE` registered as the MQTT Will *at connect time*, an explicit `OFFLINE` on clean shutdown), and shipping the `ONLINE` half without the Will is strictly worse than shipping neither: a crashed node would leave a retained `ONLINE` asserting forever that a dead host is alive |
| `RequestRebirthOnGap` | `true` | On a detected sequence-number gap, publish a rebirth-request NCMD instead of running on stale metric state |
| `BirthObservationWindowSeconds` | `15` | Discovery's birth-collection window. Must be ≥ 1 |
`MqttDriverForm` authors all five in Sparkplug mode. The sub-object is **merged, never replaced** (a
key a newer driver adds inside it survives an older AdminUI), and in `Plain` mode it is not touched at
all — so flipping a driver to Plain to look at it and back does not lose the group id. The form never
writes `RawTags`; the deploy artifact owns it.
> **This was a live-gate finding.** Through the P2 implementation `MqttDriverForm` still carried its P1
> placeholder — switching Mode to `SparkplugB` rendered a "not available yet" notice and **no group-id
> field**, so once Sparkplug ingest shipped there was no way to author a Sparkplug driver from the
> AdminUI at all. Pinned by `MqttDriverFormModelTests`' Sparkplug cases.
Every key has a default, so nothing is strictly required by the binder. Notable ones:
`port` **8883**, `useTls` **true**, `protocolVersion` **V500**, `cleanSession` **true**,
`keepAliveSeconds` **30**, `connectTimeoutSeconds` **15**,
`reconnectMinBackoffSeconds` **1** / `reconnectMaxBackoffSeconds` **30**,
`maxPayloadBytes` **1 MiB**, `allowUntrustedServerCertificate` **false**.
> ⚠️ **Leave `clientId` unset on a redundant pair.** Both nodes of a pair run the driver. MQTT
> requires client ids to be unique per broker, so a *fixed* `clientId` makes the two nodes evict each
> other forever — the broker logs `Client <id> already connected, closing old connection` and both
> nodes reconnect every few seconds while still reporting `Healthy`. Unset (the default) lets MQTTnet
> generate a unique id per connection, which is correct. This was observed live during the P1 gate.
> (The probe and browser already self-uniquify with `-probe-` / `-browse-` suffixes.) `MqttDriverForm`
> defaults the field blank and raises this warning inline the moment a value is typed.
The form validates every knob against the driver's own `[Range]` bounds (port 165535, keep-alive /
connect-timeout / backoffs ≥ 1 s, max-backoff ≥ min-backoff, `maxPayloadBytes` ≥ 1, QoS 02) and
additionally **clamps** on serialize, so an operator who ignores the inline error still cannot persist
a `connectTimeoutSeconds: 0` driver-brick.
## Tag Configuration
A tag binds in **one of two shapes**, and which one is decided by *which keys are present*, not by a
`mode` key on the tag. `MqttTagDefinitionFactory` reads the Sparkplug tuple when any Sparkplug
descriptor key is set, and the topic otherwise; the AdminUI's "Tag shape" selector infers the same way
on reopen. The key names are single-sourced in `MqttTagConfigKeys` so the browse-commit mapper and the
factory cannot drift.
### Plain shape
| Key | Type | Default | Notes |
|---|---|---|---|
| `topic` | string | — | **Required.** Absent/blank rejects the tag ⇒ `BadNodeIdUnknown` |
| `payloadFormat` | `Json` \| `Raw` \| `Scalar` | `Json` | **Strict** — a typo'd value rejects the tag |
| `jsonPath` | string | `"$"` | Absent *or empty* ⇒ document root. `Json` only |
| `dataType` | `DriverDataType` | `String` | **Strict.** `Float64`, not `Double` — see below |
| `qos` | int 02 | driver's `defaultQos` | **Strict** — a malformed value rejects rather than silently weakening delivery |
| `retainSeed` | bool | `true` | Seed this tag from the broker's retained message on subscribe |
There is **no `mode` key** — the AdminUI's "Tag shape" selector is UI-only, inferred from the blob and
never serialized.
**`dataType` uses `DriverDataType`, which has `Float64` and has no `Double`.** Authoring `"Double"` is
a strict-enum rejection, not a fallback.
Payload handling: `Json` selects via JSONPath then coerces (miss ⇒ `BadDecodingError`, coerce fail ⇒
`BadTypeMismatch`); `Raw` returns the payload's UTF-8 text verbatim and **ignores `dataType`**;
`Scalar` decodes then coerces. The JSONPath subset is `$`, `$.a.b`, `$['a']`, `$.a[0]` — **no
filters, slices, wildcards or recursive descent**.
### Wildcards
The **runtime accepts** a `+`/`#` topic (it is ambiguous, not unparseable) and the deploy-time
`Inspect` pass **warns**. The **AdminUI editor rejects** it outright — the one place the editor is
deliberately stricter than the driver, because one Tag holds one value and a wildcard would feed it
from many source topics. Wildcard tags are indexed **by filter**, not by concrete topic, so they
survive a reconnect; a message fans out to *every* matching tag, and unauthored topics are dropped.
### Sparkplug shape
| Key | Type | Default | Notes |
|---|---|---|---|
| `groupId` | string | — | **Required.** Must match the driver's `Sparkplug.GroupId` to ever receive a message |
| `edgeNodeId` | string | — | **Required.** |
| `deviceId` | string | *absent* | **Omit** for a metric published by the edge node itself (NBIRTH/NDATA). Absent is meaningfully different from `""` |
| `metricName` | string | — | **Required.** The tag's stable binding key across rebirths. **May contain `/`**`Node Control/Rebirth` and `Properties/Serial Number` are canonical Sparkplug names, and the whole name is one value, never split |
| `dataType` | `DriverDataType` | *from the birth* | **Optional override.** Absent ⇒ take whatever type the birth certificate declared. The AdminUI's "(from birth certificate)" option removes the key entirely |
| `isHistorized` / `historianTagname` | — | — | Server-side historization, unchanged from Plain |
`payloadFormat`, `jsonPath`, `qos` and `retainSeed` are **Plain-only** and meaningless here: a
Sparkplug body is protobuf, and the subscription is one driver-wide `spBv1.0/{GroupId}/#`, not a
per-tag filter.
`groupId` / `edgeNodeId` / `deviceId` are literal MQTT topic segments, so the AdminUI editor refuses
`/`, `+` and `#` in them — a decoded incoming id can never contain one, so such a tag could never
bind. `metricName` is deliberately **exempt**: it is not a topic segment.
```jsonc
// A device metric.
{"groupId":"OtOpcUaSim","edgeNodeId":"EdgeA","deviceId":"Filler1","metricName":"FillCount"}
// A node-level metric — note deviceId is ABSENT, not blank.
{"groupId":"OtOpcUaSim","edgeNodeId":"EdgeA","metricName":"Temperature"}
```
## Sparkplug B Ingest
One subscription (`spBv1.0/{GroupId}/#`) feeds a per-edge-node state machine
(`SparkplugIngestor` + `SparkplugCodec` / `AliasTable` / `BirthCache` / `SequenceTracker`):
- **Births are the only self-describing message.** NBIRTH/DBIRTH name every metric and assign its
alias; NDATA/DDATA carry an **alias and no name at all**. So the alias table is rebuilt *per birth*
and tags bind **by name**, never by alias — aliases are per-birth and may be reused for a different
metric after a rebirth.
- **NBIRTH is node-scoped, DBIRTH is device-scoped**, and routing a DBIRTH through the node path
wipes `bdSeq`. They are handled separately on purpose.
- **Sequence tracking**: `seq` is a wrapping 0255 counter. A detected gap raises a rebirth request
when `RequestRebirthOnGap` is on, rather than silently continuing on stale metric state.
- **Death → STALE.** An NDEATH whose `bdSeq` matches the live session (the bdSeq tie is what stops a
stale death from killing a fresh session) marks every metric under that edge node `Bad` /
`BadNoCommunication`. A DDEATH does the same for one device. The next birth restores them.
- **Unsupported datatypes are skipped with a warning, never guessed**: `DataSet`, `Template`,
`PropertySet`, `PropertySetList`, `Unknown`. `Int8``Int16` and `UInt8``UInt16` widen; `*Array`
variants map to the element type with the array bit set.
- **Ingest state is reset at every session/authoring seam** — a reconnect or a `ReinitializeAsync`
must not carry a previous session's alias table into a new one.
### Rediscovery (Sparkplug)
`OnRediscoveryNeeded` fires only when a birth is **for an authored scope** *and* its metric-name set
**differs** from the last one seen for that scope (order-insensitive, ordinal). Both conditions are
anti-storm, not optimisations: rediscovery triggers an address-space rebuild, and an edge node
re-births freely — on its own schedule, on every reconnect, and once per rebirth NCMD.
> ⚠️ **Nothing consumes the signal today.** No runtime component subscribes to
> `IRediscoverable.OnRediscoveryNeeded`, and `DriverHostActor.HandleDiscoveredNodes` hard-returns —
> discovered-node injection is dormant in v3. The served address space comes from the deploy artifact,
> so **a DBIRTH introducing a new metric does not change the OPC UA tree**; you must redeploy. The
> driver's decision is observable only as an Information log line
> (`… requesting rediscovery.`). This is a v3 platform gap, not an MQTT one.
## Connection + Failure Semantics
`MqttConnectionState`: `Disconnected → Connected → Reconnecting → Faulted → Disposed`.
A broker that is merely down stays **`Reconnecting` indefinitely** — unreachable is never a fault.
Backoff is exponential between the two `reconnectBackoff` bounds.
**MQTTnet 5 does not throw on a rejected CONNACK** — it returns the code and leaves the client
disconnected. Left unhandled, a wrong password produced a `Healthy` driver that never received a
value. `MqttConnection` therefore raises `MqttConnectRejectedException`:
- **Unrecoverable ⇒ `Faulted`, supervisor stops** (retrying cannot help): `MalformedPacket`,
`ProtocolError`, `UnsupportedProtocolVersion`, `ClientIdentifierNotValid`, `BadUserNameOrPassword`,
`NotAuthorized`, `Banned`, `BadAuthenticationMethod`, `TopicNameInvalid`, `PacketTooLarge`,
`PayloadFormatInvalid`, `RetainNotSupported`, `QoSNotSupported`, `ServerMoved`.
- **Everything else is transient** and keeps retrying — including `UnspecifiedError` and any future
code. (`ServerMoved` is permanent, `UseAnotherServer` is temporary — per the MQTT 5 spec.)
A reconnect that completes without re-subscribing would be a healthy-looking connection that receives
nothing forever, so `OnReconnectedAsync` **throws** when zero filters are granted, tearing the session
down rather than swallowing it. A SUBACK failure degrades just the affected RawPaths to
`BadCommunicationError`.
## Browse
`MqttDriverBrowser` opens a **read-only observation session**: it subscribes at QoS 0 to `#` (or
`{topicPrefix}/#`) under a `-browse-`-suffixed client id and builds a `/`-segment tree from whatever
publishes. **Nothing is ever published.** It is *inherently incomplete* — a topic that stays silent
during the window is invisible — so manual entry remains the escape hatch.
Bounds: 50 000 nodes (then `⚠ Observation truncated`), 1024-char topics / 256-char segments (then
`⚠ Some topics were too long`), 512-byte payload inspection. Sessions are reaped after 2 minutes idle.
**Sparkplug B browse serves a different tree from the same passive window**: observed NBIRTH/DBIRTH
certificates are decoded into `Group → EdgeNode → [Device] → Metric` (a birth is the only Sparkplug
message that names its metrics). Same node cap, same bounds, same "publishes nothing" guarantee.
The bespoke browser wins over the universal `DiscoveryDriverBrowser`, which would decline MQTT anyway
(`SupportsOnlineDiscovery` is `false`).
### Browse-commit writes the address the factory reads
A committed leaf's address is a **descriptor**, not a single reference string, so
`RawBrowseCommitMapper` reads it from the `AttributeInfo.AddressFields` the session states — Plain
emits `topic`; Sparkplug emits `groupId` / `edgeNodeId` / `deviceId?` / `metricName`, the names
`MqttTagConfigKeys` single-sources for both the mapper and `MqttTagDefinitionFactory`.
**It is never parsed out of the browse node id.** A Sparkplug node id is
`{group}/{node}[/{device}]::{metric}` and a metric name legitimately contains `/`
(`Node Control/Rebirth`), so splitting the id cannot recover the tuple. Which keys the session emits
is also how the mapper learns the shape — the driver *type* reaching it is just `Mqtt`, and the mode
lives on the driver config. A leaf whose attribute lookup returned nothing carries no address and is
refused **at commit, in words**, rather than committed as a tag that deploys clean and then reports
`BadNodeIdUnknown` forever.
### Request rebirth (Sparkplug only)
Births are never retained, so a healthy but quiet plant shows an **empty** picker tree until one
lands. The `Request rebirth` affordance in the browse modal is the way to fill it on demand: it
publishes a Sparkplug rebirth-request NCMD (`Node Control/Rebirth`) to the selected scope — the **one
and only** browse action that publishes anything.
- **Scope** is the last node clicked in the tree. A device or metric resolves **up to its owning edge
node** (Sparkplug has no device-scoped rebirth); a group fans out to every observed edge node
beneath it and is refused **whole** past 32, publishing nothing.
- **Two clicks, never one**`Request rebirth…` arms a confirm panel naming the resolved scope and
its consequence; changing the selection disarms it.
- **Gated on `DriverOperator`**, enforced server-side in `BrowserSessionService.RequestRebirthAsync`
(fail-closed) before the session is touched; the UI gate is defence in depth. A refusal — policy,
over-wide group, unresolvable scope — is **shown**, not swallowed.
- Offered only for a Sparkplug session (`IRebirthCapableBrowseSession.RebirthAvailable`); a Plain
MQTT window publishes nothing, ever.
### Refresh
An observation window **accumulates**, but the tree renders **once**, when it is created — and
`OpenAsync` returns as soon as the SUBSCRIBE is granted, without waiting for the birth-observation
window. So the first render is normally empty on Plain (a topic that has not published yet is
invisible) and *almost always* empty on Sparkplug (births are never retained). **Refresh** re-reads
the root against the same session: nothing reconnects, nothing already observed is lost, and the tag
selection is kept. Only the armed rebirth scope is cleared, because the node it named may no longer be
in the rebuilt tree.
> **Also fixed by this gate, and not MQTT-specific:** every node label in the shared
> `DriverBrowseTree` was an `<a href="#">`. `blazor.web.js`'s enhanced-navigation click interceptor
> resolves a bare `#` against `<base href="/">`, so clicking a node **label** navigated the whole
> AdminUI to `/` and destroyed the hosting modal (`@onclick:preventDefault` suppresses the browser's
> default action, not Blazor's interceptor). The labels are now `<button type="button">`. This
> affected **every** driver's picker since 2026-05-28; it surfaced here because choosing a rebirth
> scope is the first flow that requires clicking a label rather than the ▶ toggle.
> ⚠️ **Known gap — a completely empty tree cannot arm a rebirth.** The rebirth scope is taken from the
> last node clicked *in the tree*, so on a plant that has not birthed since the window opened there is
> nothing to click and `Request rebirth…` stays disabled — the exact case the affordance exists for.
> `MqttBrowseSession.ResolveRebirthTargets` already accepts a bare `{group}/{edgeNode}` for a node the
> window has never seen (it calls that "the prime rebirth target"); what is missing is a UI path to
> enter one, or to default the scope to the driver's own configured `GroupId`. Until then: use
> **Refresh** after any birth, or author the tags by **Manual entry**, which is always available.
## Test Connect
`MqttDriverProbe` performs **one CONNECT** (no SUBSCRIBE, no publish) under a `-probe-` client id and
disconnects. Success message: `"MQTT CONNECT OK"`. A rejected CONNACK is described by the same
`DescribeConnackRejection` the running driver uses, so the button and the driver report a rejection in
identical words. See [TestConnectProbes.md](TestConnectProbes.md).
## Testing
- **Unit**`tests/Drivers/…Driver.Mqtt.Tests` (**581 tests**): tag-config mapping (both shapes),
JSONPath subset, strict-enum rejection, subscription indexing (incl. the wildcard-by-filter case),
reconnect/rebuild branches, CONNACK classification, and the Sparkplug half — golden-payload decode,
topic parse/format, datatype map, alias/birth-cache rebuild, sequence + bdSeq-death ties, the
rediscovery change gate, and the browse session's "publishes nothing" assertion.
- **Live (env-gated)**`tests/Drivers/…Driver.Mqtt.IntegrationTests` (**15 tests**) against the
Mosquitto broker + the C# Sparkplug edge-node simulator; skips cleanly when `MQTT_FIXTURE_ENDPOINT`
is unset. See [Mqtt-Test-Fixture.md](Mqtt-Test-Fixture.md) and `infra/README.md` §3.
- **AdminUI**`MqttDriverFormModelTests` / `MqttTagConfigModelTests` cover both editors'
round-trip, mode inference and validation. AdminUI has no bUnit, so the razor shells themselves are
only ever proven by a live `/run` gate.
## Operational Notes
- **Read-only.** No writes reach the broker; every node is ViewOnly. Sparkplug NCMD is used for
**rebirth requests only** — never to command a device.
- **No alarms, no driver-side history.** Historization is a server-side concern (`isHistorized`).
- **Both pair nodes subscribe independently** — MQTT fan-out is per-connection, and the Primary gate
applies downstream, not to ingest. This is also why `clientId` must stay unset (above).
- **Unauthored traffic is ignored.** A busy broker is normal; the driver never auto-provisions a tag.
## Known gaps
Every row is tracked. None is a silent failure — each either refuses in words or logs.
| Gap | Issue | Effect | Notes |
|---|:---:|---|---|
| **Write-through (`IWritable`)** | [#508](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/508) | Every MQTT node is read-only | Deferred to P3 — NCMD/DCMD + plain publish, optimistic-Good + self-correct on echo |
| **Rediscovery is inert** | [#507](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/507) | A DBIRTH introducing a metric does not change the OPC UA tree; redeploy to pick it up | v3 **platform** gap, not MQTT-specific — nothing subscribes to `OnRediscoveryNeeded` and `DriverHostActor.HandleDiscoveredNodes` hard-returns. Also affects Galaxy / TwinCAT / AbCip. Driver-side decision is log-observable |
| **Rebirth needs a non-empty tree** | [#514](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/514) | `Request rebirth…` cannot be armed on a plant that has not birthed since the window opened | See [Refresh](#refresh). Session side already supports a bare `{group}/{edgeNode}` scope |
| **`_canRebirth` is captured at browse-open** | [#510](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/510) | A session reaped for idleness (2 min) still draws the button, and using it errors | The error is shown, not swallowed; reopen the browser |
| **Primary-host STATE publishing** | [#511](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/511) | `ActAsPrimaryHost` does nothing | Logs a startup warning rather than being silently inert — see the `Sparkplug` sub-object table |
| **A metric whose name contains `/` cannot be browse-committed** | [#509](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/509) | The commit is refused **whole**, in words (`Row N: Name must not contain '/'`) | The derived **tag Name** is a RawPath segment and may not contain `/`, while `metricName` legitimately may. This blocks the spec-mandatory `Node Control/Rebirth`. **Workaround: Manual entry** — author the tag with a `/`-free name and type the slashed `metricName` into the Sparkplug editor, which accepts it. Nothing is silently mis-bound |
| **`DataSet` / `Template` / `PropertySet` metrics** | [#515](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/515) | Skipped with a warning | v1-unsupported; never coerced into a guessed type |
| **The tag editor writes `payloadFormat` into a Sparkplug blob** | [#513](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/513) | Cosmetic only | `payloadFormat` is Plain-only; the factory routes on the Sparkplug keys and ignores it. Noise in the blob, not a mis-binding |
| **`.Browser` references `.Driver`, not just `.Contracts`** | [#512](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/512) | A documented layering exception | Taken to reuse `MqttConnection.BuildClientOptions` rather than duplicate TLS/CA-pin logic — see the note at the top of this doc |
+11 -2
View File
@@ -29,6 +29,8 @@ Driver type metadata is registered at startup in `DriverTypeRegistry` (`src/Core
| [TwinCAT](TwinCAT.md) | `Driver.TwinCAT` | B | Beckhoff `TwinCAT.Ads` (`TcAdsClient`) | IDriver, ITagDiscovery, IReadable, IWritable, ISubscribable, IHostConnectivityProbe, IPerCallHostResolver, IRediscoverable | The only native-notification driver outside Galaxy — ADS delivers `ValueChangedCallback` events the driver forwards straight to `ISubscribable.OnDataChange` without polling. Symbol tree uploaded via `SymbolLoaderFactory` |
| [FOCAS](FOCAS.md) | `Driver.FOCAS` | A | Pure-managed `FocasWireClient` — FOCAS/2 Ethernet binary protocol on TCP:8193, inlined into the driver assembly | IDriver, ITagDiscovery, IReadable, IWritable, ISubscribable, IHostConnectivityProbe, IPerCallHostResolver, IAlarmSource | `IWritable` is implemented but read-only by design — `WriteAsync` returns `BadNotWritable` for every point. CNC-shaped data model (axes, spindle, PMC, macros, alarms) not a flat tag map. Previously Tier-C (Host + P/Invoke + shim DLL); retired in the 2026-04-24 migration when the managed wire client landed |
| [OPC UA Client](OpcUaClient.md) | `Driver.OpcUaClient` | B | OPCFoundation `Opc.Ua.Client` | IDriver, ITagDiscovery, IReadable, IWritable, ISubscribable, IAlarmSource, IHistoryProvider, IHostConnectivityProbe | Gateway/aggregation driver — the only driver implementing driver-side `IHistoryProvider` (forwards HistoryRead to the upstream server). Opens a single `Session` against a remote OPC UA server and re-exposes its address space. Owns its own `ApplicationConfiguration` (distinct from `Client.Shared`) because it's always-on with keep-alive + `TransferSubscriptions` across SDK reconnect, not an interactive CLI |
| [MQTT](Mqtt.md) | `Driver.Mqtt` (+ `.Browser`, `.Contracts`) | A | [MQTTnet 5](https://github.com/dotnet/MQTTnet) | IDriver, ITagDiscovery, IReadable, ISubscribable, IHostConnectivityProbe, IRediscoverable | **Push, not poll, and read-only** — the broker delivers PUBLISH frames the driver forwards straight to `ISubscribable.OnDataChange`; `IReadable` serves the last retained/received value from cache, and there is **no `IWritable`** (nothing publishes back). Subscriptions are indexed **by topic filter, not by topic**, so wildcard tags survive a reconnect. A rejected CONNACK throws `MqttConnectRejectedException`: unrecoverable codes → `Faulted` + supervisor stop, transient → retry. **Two modes:** `Plain` binds a tag to a concrete topic + JSON path; `SparkplugB` decodes vendored Eclipse-Tahu protobuf under one `spBv1.0/{groupId}/#` subscription and binds by the `group/edgeNode/device?/metric` tuple (birth/alias/seq-gap state machine, death→STALE, scoped rebirth NCMD). `IRediscoverable` fires on a DBIRTH but is **inert platform-side** — see [#507](https://gitea.dohertylan.com/dohertj2/lmxopcua/issues/507) |
| [MTConnect](MTConnect.md) | `Driver.MTConnect` (+ `.Contracts`) | — | Hand-rolled `System.Xml.Linq` HTTP/XML client against a vendor-neutral MTConnect Agent (`/probe`, `/current`, `/sample`) — no TrakHound dependency (dropped; see the driver doc) | IDriver, ITagDiscovery, IReadable, ISubscribable, IHostConnectivityProbe, IRediscoverable | Read-only by design (no `IWritable`). Production data plane is entirely `ISubscribable`/`OnDataChange``IReadable` is CLI/test-only. Browse is free via the Wave-0 universal discovery browser, no bespoke browser project. `IRediscoverable`/`IHostConnectivityProbe` have no server-side consumer yet — a pre-existing fleet-wide gap, not MTConnect-specific |
| [Historian.Gateway](../Historian.md) | `Driver.Historian.Gateway` | — | `ZB.MOM.WW.HistorianGateway.Client` gRPC (`historian_gateway.v1`) | IHistorianDataSource (server-side read backend) + alarm `SendEvent` writer + `WriteLiveValues` recorder + `IHistorianProvisioning` | Not a tag driver — the sole historian backend. Registers `GatewayHistorianDataSource : IHistorianDataSource` for HistoryRead and serves alarm-write + continuous historization through the gateway. No `IDriver`/`ITagDiscovery` surface. (The retired Wonderware sidecar backend it replaced is documented at [Historian.Wonderware.md](Historian.Wonderware.md).) |
## Per-driver documentation
@@ -40,13 +42,18 @@ Driver type metadata is registered at startup in `DriverTypeRegistry` (`src/Core
- **FOCAS** has a short getting-started doc because the backend-selection env var + alarm projection opt-in need explaining up front:
- [FOCAS.md](FOCAS.md) — deployment, config, capability surface, alarm projection, troubleshooting
- **Modbus TCP**, **AB CIP**, **AB Legacy**, **Siemens S7**, **TwinCAT**, and **OPC UA Client** each have a per-driver overview page:
- **Modbus TCP**, **AB CIP**, **AB Legacy**, **Siemens S7**, **TwinCAT**, **OPC UA Client**, and **MQTT** each have a per-driver overview page:
- [Modbus.md](Modbus.md) — in-process Modbus-TCP driver: address formats, polled subscription model, DL205 octal mapping
- [AbCip.md](AbCip.md) — AB CIP / EtherNet-IP driver (ControlLogix / CompactLogix / Micro800 / GuardLogix): tag discovery, UDT resolution, alarm source
- [AbLegacy.md](AbLegacy.md) — AB Legacy PCCC driver (SLC 500 / MicroLogix / PLC-5): file-based addressing, user-authored tag list
- [S7.md](S7.md) — Siemens S7 driver (S7-300/400/1200/1500 + S7-200): getting started, config, data-block addressing, serialized single-connection model
- [TwinCAT.md](TwinCAT.md) — Beckhoff TwinCAT (ADS) driver: getting started, native-notification subscription, symbol-tree upload
- [OpcUaClient.md](OpcUaClient.md) — OPC UA Client (gateway/aggregation) driver: remote-server session, driver-side HistoryRead forwarding, reconnect behaviour
- [Mqtt.md](Mqtt.md) — MQTT broker-subscribe driver: broker/TLS/auth config, topic + JSON-path tag binding, retained-value seeding, `#`-observation browser, CONNACK-rejection semantics, **Sparkplug B** (birth/alias/seq/rebirth state machine, birth-driven browse picker, death→STALE), and the **Known gaps** table
- **MTConnect** has a short getting-started doc because the hand-rolled-vs-TrakHound divergence, the
data-plane precedence rules, and a non-trivial known-limitations list need explaining up front:
- [MTConnect.md](MTConnect.md) — Agent config, `FullName`==DataItem `id`, `UNAVAILABLE``BadNoCommunication`, CONDITION-as-`String`, browse-via-universal-browser, fixture recipe, deferred write-back
- **Historian.Gateway** (server-side historian backend, not a tag driver) is documented in the main guide:
- [../Historian.md](../Historian.md) — HistorianGateway backend: read-path registration, HistoryRead dispatch, alarm store-and-forward (`SendEvent`), continuous historization (`WriteLiveValues`), `EnsureTags` provisioning, config keys, deployment prerequisites. (The retired Wonderware sidecar backend it replaced: [Historian.Wonderware.md](Historian.Wonderware.md).)
@@ -64,12 +71,14 @@ Each driver has a dedicated fixture doc that lays out what the integration / uni
- [TwinCAT](TwinCAT-Test-Fixture.md) — XAR-VM integration scaffolding (task #221); three smoke tests skip when VM unreachable. Unit via `FakeTwinCATClient` with native-notification harness
- [FOCAS](FOCAS-Test-Fixture.md) — no integration fixture, unit-only via `FakeFocasClient`; Tier C out-of-process isolation scoped but not shipped
- [OPC UA Client](OpcUaClient-Test-Fixture.md) — Dockerized `opc-plc` integration suite (task #215): real Secure Channel + Session, read + subscribe verified end-to-end; write not yet exercised in the integration suite; exhaustive capability matrix (reconnect, failover, cert-auth, history, alarms) via unit suite with mocked `Session`
- [MQTT](Mqtt-Test-Fixture.md) — Dockerized `eclipse-mosquitto` with **TLS + real auth on both listeners (no anonymous fallback)** + a JSON-publisher sidecar; env-gated live suite (**15 tests**) — 7 plain (TLS/auth connect, CA pinning both ways, subscribe→value, retained seeding, wrong-password rejection) + 8 Sparkplug against a **project-owned C# edge-node simulator** on the `sparkplug` compose profile (birth→alias bind, DDATA, NDEATH→STALE, rebirth NCMD round-trip). Births are never retained, so `docker restart otopcua-sparkplug-sim` is how you force a fresh one
- [Galaxy](../v1/drivers/Galaxy-Test-Fixture.md) — richest harness: gateway E2E + ZB SQL live-smoke + MXAccess opt-in (v1 archive)
- [MTConnect](../../tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.IntegrationTests/Docker/README.md) — Dockerized `mtconnect/agent` + stdlib SHDR adapter (two services, the Agent alone reports everything `UNAVAILABLE`); 12/12 against the real Agent, canned XML unit fixtures cover the bulk (491/491)
## Related cross-driver docs
- [HistoricalDataAccess.md](../v1/HistoricalDataAccess.md) — `IHistoryProvider` dispatch, aggregate mapping, continuation points. The OPC UA Client driver is the only driver that implements driver-side `IHistoryProvider` (it forwards HistoryRead to the upstream server); the AVEVA Historian path is served server-side by the HistorianGateway-backed `IHistorianDataSource` instead. Other drivers do not implement the interface and return `BadHistoryOperationUnsupported`.
- [AlarmTracking.md](../AlarmTracking.md) — `IAlarmSource` event model and filtering. Implemented by Galaxy (native MxAccess alarms, working end-to-end), OPC UA Client, AB CIP, and FOCAS; AB Legacy, Modbus, S7, and TwinCAT have no alarm source.
- [AlarmTracking.md](../AlarmTracking.md) — `IAlarmSource` event model and filtering. Implemented by Galaxy (native MxAccess alarms, working end-to-end), OPC UA Client, AB CIP, and FOCAS; AB Legacy, Modbus, S7, TwinCAT, and MQTT have no alarm source.
- [Subscriptions.md](../v1/Subscriptions.md) — how the Server multiplexes subscriptions onto `ISubscribable.OnDataChange`.
- [docs/v2/driver-stability.md](../v2/driver-stability.md) — tier system (A / B / C), shared `CapabilityPolicy` defaults per tier × capability, `MemoryTracking` hybrid formula, and process-level recycle rules.
- [docs/v2/plan.md](../v2/plan.md) — authoritative vision, architecture decisions, migration strategy.
+2
View File
@@ -49,6 +49,7 @@ with a human-readable explanation rather than a false-green TCP-open tick.
| **TwinCAT** | `AdsClient.Connect` + `ReadStateAsync`. See [degrade semantics](#twincat-degrade) below. | `"ADS state: {state}"` | Deferred — no ADS target |
| **FOCAS** | `cnc_allclibhndl3` via a direct `DllImport("fwlib32")` in the probe. See [degrade semantics](#focas-degrade) below. | `"FOCAS handle OK"` | Deferred — no CNC + FWLIB |
| **Galaxy** | gRPC unary call to `GalaxyRepository.TestConnection` on the configured mxaccessgw endpoint. See [auth-rejection rule](#galaxy-auth-rejection) below. | `"gateway gRPC OK"` | `http://10.100.0.48:5120` (mxaccessgw) |
| **MQTT** | One MQTT `CONNECT` (no SUBSCRIBE, no publish) under a `-probe-`-suffixed client id, then disconnect. **A rejected CONNACK is a returned result code, not an exception** — MQTTnet 5 does not throw — so the probe inspects `ResultCode` and reuses the driver's own `DescribeConnackRejection`, making the button and the running driver word a rejection identically. | `"MQTT CONNECT OK"` | `10.100.0.35:8883` (Mosquitto TLS+auth fixture) |
**Historian.Wonderware** had a TCP `Hello``HelloAck` handshake probe before Phase 5, but the
Wonderware historian backend (and its driver-type / probe) has since been **retired** — the historian
@@ -133,6 +134,7 @@ above. (Live verification on `10.100.0.48:5120` with no key returns
| S7 | Verified | python-snap7 `10.100.0.35:1102` |
| AbCip | Verified | CIP sim `10.100.0.35:44818` |
| Galaxy | Verified | mxaccessgw `10.100.0.48:5120`; `Unauthenticated` reply counts as Ok |
| MQTT | Verified | Mosquitto `10.100.0.35:8883`; live suite covers TLS+auth Ok, foreign-CA rejection, and wrong-password rejection surfacing rather than reporting healthy |
| AbLegacy | Deferred | No PLC5/SLC sim; unit-proven + code path identical to AbCip |
| TwinCAT | Deferred | No ADS target; unit-proven + degrade guard tested |
| FOCAS | Deferred | No CNC + FWLIB on dev host; degrade guard is the CI-observable path |
@@ -271,7 +271,7 @@ public sealed class GatewayQualityMapperTests
{
[Theory]
[InlineData(192, 0x00000000u)] // Good
[InlineData(216, 0x00D80000u)] // Good_LocalOverride
[InlineData(216, 0x00960000u)] // Good_LocalOverride
[InlineData(64, 0x40000000u)] // Uncertain
[InlineData(0, 0x80000000u)] // Bad
[InlineData(8, 0x808A0000u)] // Bad_NotConnected
@@ -307,7 +307,7 @@ public sealed class SampleMapperTests
{
var a = new HistorianAggregateSample { Tag = "T", /* Value unset */ EndTime = Ts(2026,1,1,0,0,0) };
var snap = SampleMapper.ToAggregateSnapshot(a);
Assert.Equal(0x800E0000u, snap.StatusCode); // BadNoData
Assert.Equal(0x809B0000u, snap.StatusCode); // BadNoData
Assert.Null(snap.Value);
}
// Ts(...) builds a Google.Protobuf.WellKnownTypes.Timestamp from UTC parts.
@@ -323,7 +323,7 @@ public sealed class SampleMapperTests
- `StatusCode`: `GatewayQualityMapper.Map((byte)s.OpcQuality)` (prefer `opc_quality`; if zero/unset and `quality` carries the OPC-DA byte, fall back to `quality` — match whatever the gateway populates; document the choice).
- `SourceTimestampUtc`: `s.Timestamp.ToDateTime()` (UTC kind).
- `ServerTimestampUtc`: `DateTime.UtcNow`.
- `SampleMapper.ToAggregateSnapshot(HistorianAggregateSample)` → null aggregate value ⇒ `StatusCode 0x800E0000` (BadNoData), non-null ⇒ Good (`0x00000000`) with the value; `SourceTimestampUtc` ← the bucket end/start timestamp (match the Wonderware `ToAggregateSnapshots` convention — it stamps the bucket timestamp). Provide `IReadOnlyList<>` batch helpers too.
- `SampleMapper.ToAggregateSnapshot(HistorianAggregateSample)` → null aggregate value ⇒ `StatusCode 0x809B0000` (BadNoData), non-null ⇒ Good (`0x00000000`) with the value; `SourceTimestampUtc` ← the bucket end/start timestamp (match the Wonderware `ToAggregateSnapshots` convention — it stamps the bucket timestamp). Provide `IReadOnlyList<>` batch helpers too.
**Step 4: run, expect PASS.**
@@ -98,7 +98,7 @@ Pure function in `.Contracts` (shared by driver, browser-commit, and the typed e
| `EVENT` free text (`Program`, `Block`, `Message`) | `String` |
| `CONDITION` | `String` (state word; §3.6) |
**Quality mapping (`DataValueSnapshot.StatusCode`):** an observation value of `UNAVAILABLE` (MTConnect's explicit no-data sentinel) → **`BadNoCommunication` (`0x80310000u`)** with a null value, not the literal string — this is *the* UNAVAILABLE code, used everywhere in this design (read, subscribe, and the §8 fixtures). Rationale + fleet fit: semantically, `UNAVAILABLE` means the Agent is reachable but has no device-backed value (the adapter↔device link is down or the item has never reported) — exactly OPC UA's `BadNoCommunication` ("communication to the data source has failed"). The fleet grep shows no competing convention for this case: every existing driver's `BadCommunicationError` (`0x80050000u`, Modbus/FOCAS/TwinCAT/OpcUaClient/S7/AbCip/AbLegacy/Galaxy) marks the **driver's own transport failure** to its peer — a different failure (and the code this driver also uses when the *Agent* HTTP call fails, per §7) — while `BadNoData` (`0x800E0000u`) is historian-domain only (`Historian.Gateway/Mapping/SampleMapper.cs`). `BadNoCommunication` is already in the CLI's `SnapshotFormatter` name table (`src/Drivers/Cli/ZB.MOM.WW.OtOpcUa.Driver.Cli.Common/SnapshotFormatter.cs`), so it renders by name, not hex. Declare it as a `private const uint` in `MTConnectDriver` (the Modbus `StatusBadCommunicationError` pattern). Missing dataItem / empty condition → `Bad`. `observation.timestamp``SourceTimestamp` (not a separate node).
**Quality mapping (`DataValueSnapshot.StatusCode`):** an observation value of `UNAVAILABLE` (MTConnect's explicit no-data sentinel) → **`BadNoCommunication` (`0x80310000u`)** with a null value, not the literal string — this is *the* UNAVAILABLE code, used everywhere in this design (read, subscribe, and the §8 fixtures). Rationale + fleet fit: semantically, `UNAVAILABLE` means the Agent is reachable but has no device-backed value (the adapter↔device link is down or the item has never reported) — exactly OPC UA's `BadNoCommunication` ("communication to the data source has failed"). The fleet grep shows no competing convention for this case: every existing driver's `BadCommunicationError` (`0x80050000u`, Modbus/FOCAS/TwinCAT/OpcUaClient/S7/AbCip/AbLegacy/Galaxy) marks the **driver's own transport failure** to its peer — a different failure (and the code this driver also uses when the *Agent* HTTP call fails, per §7) — while `BadNoData` (`0x809B0000u`) is historian-domain only (`Historian.Gateway/Mapping/SampleMapper.cs`). `BadNoCommunication` is already in the CLI's `SnapshotFormatter` name table (`src/Drivers/Cli/ZB.MOM.WW.OtOpcUa.Driver.Cli.Common/SnapshotFormatter.cs`), so it renders by name, not hex. Declare it as a `private const uint` in `MTConnectDriver` (the Modbus `StatusBadCommunicationError` pattern). Missing dataItem / empty condition → `Bad`. `observation.timestamp``SourceTimestamp` (not a separate node).
### 3.4 `IReadable``/current`
@@ -588,6 +588,40 @@ SQL Server client works cross-platform, so `Sql` runs on macOS dev too (unlike G
`dialect.QuoteIdentifier` (which escapes/rejects embedded quote characters). A table/column string
in a `TagConfig` is validated against the live catalog (or an allow-list) before it's ever quoted
into text; an unknown identifier rejects the tag (→ `BadNodeIdUnknown`) rather than executing.
**Implemented (Gitea #496).** `SqlCatalogGate` + `SqlCatalogLoader`, run from
`SqlDriver.InitializeAsync` after the liveness check and before the driver reports Healthy. Shape,
and the decisions worth knowing before touching it:
- **The catalog's spelling is substituted, not merely checked.** An accepted identifier is rewritten
to the string the catalog returned, so the text quoted into a poll query is one this driver read
back out of `ListSchemas`/`ListTables`/`ListColumns` — not operator input. Matching is
exact-ordinal first, then a *unique* case-insensitive hit (SQL Server's default collation is CI, so
case variants have always worked and rejecting them would break valid configs); an ambiguous CI
match under a case-sensitive collation is refused rather than guessed, because picking one would
publish another column's data under the operator's node.
- **Charset check first, catalog lookup second.** Each identifier goes through `QuoteIdentifier` for
its rejection rules *before* being looked up, so a name carrying a control or Unicode format
character is rejected **without its value being echoed** into a log line (Trojan-Source). A name
that passes is safe to render, which is why catalog-miss messages *do* name it — an operator
hunting a typo has to see what they wrote.
- **Rejection is per-tag and keeps the node.** The rejected tag is dropped from the *polled* table
but stays in the *authored* table, so it still materializes as a node and reads
`BadNodeIdUnknown` (§8.1's specified outcome) instead of vanishing from the address space. Every
drop is logged at Warning with the tag, the field and the reason.
- **A catalog that cannot be read faults Initialize; it does not reject every tag.** An unreadable
catalog is the *absence* of evidence about the tags, not evidence against them — rejecting all of
them would serve a confidently-empty address space and send the operator hunting typos that do not
exist. Zero visible schemas is treated the same way, because that is what a missing GRANT looks
like. Faulting lands `DriverInstanceActor` in Reconnecting with its retry timer running.
- **Bounded load:** one schema-list query, one default-schema scalar (`ISqlDialect.DefaultSchemaSql`,
added for this — an unqualified `TagValues` must resolve in whatever schema the *server* says, not
a guessed `dbo`), then one `ListTables` per distinct authored schema and one `ListColumns` per
distinct authored relation. It never enumerates the whole catalog, and the load runs under the same
wall-clock bound as the liveness check (R2-01 / STAB-14).
- **Accepted v1 limitation:** a 3-part `db.schema.table` (or a linked-server name) addresses a
catalog this connection cannot enumerate, so it cannot be allow-listed and is rejected with a
message pointing at the fix — expose the data through a view in the connected database.
- **Treat any code path that builds SQL by string-concatenating a tag field as a defect** — enforce
with a review checklist item + a unit test that feeds a malicious `keyValue`/`table`
(`'; DROP TABLE …`) and asserts it either binds harmlessly (value) or is rejected (identifier),
@@ -607,8 +641,15 @@ comment) and the env-overridable ConfigDb connection string:
file. Direct env read (not `IConfiguration`) is deliberate: the factory registry materializes
drivers via a static `(id, json)` closure (Modbus pattern) with no `IConfiguration` in reach, so
this keeps the factory shape unchanged.
- If an inline `connectionString` is ever permitted (dev convenience only), the AdminUI **redacts** it
in display/logging and flags it dev-only.
- **An inline `connectionString` is not permitted at all** — the earlier "dev convenience, redacted in
the UI" carve-out is withdrawn (Gitea #498). Redaction only hides a credential that has *already been
written* to the config DB and replicated to every node's artifact cache. `SqlDriverConfigDto` has no
such property and `UnmappedMemberHandling.Skip` drops the key on read, but that protects only the read
path; config blobs are schemaless JSON columns, so nothing stopped the key being **written**.
`DraftValidator.ValidateSqlConnectionStringNotPersisted` now fails the deploy
(`SqlConnectionStringPersisted`) for a `connectionString` key at the top level of a Sql driver's
`DriverConfig` **or** of the `DeviceConfig` of any device beneath it (the two are merged before the DTO
sees them). The key is matched case-insensitively, and the error text never echoes the value.
- **Never log the resolved connection string.** Log only provider + server host + database name.
- Prefer integrated/managed auth where the estate supports it (`Integrated Security=true` / Azure AD /
Kerberos) so no password transits config at all.
+8 -1
View File
@@ -94,7 +94,14 @@ dual-namespace + event-delivery), Runtime 377 (+ VT reassert regression), AdminU
## Documented follow-ups (non-blocking)
- **Scripted-alarm redeploy recovery (CONFIRMED pre-existing bug, deferred — deserves its own careful fix). Filed as issue #487.**
- **Scripted-alarm redeploy recovery — FIXED (issue #487), see `docs/ScriptedAlarms.md` §"Post-(re)load
condition-node re-assert".** Landed as the proposed shape below: `ScriptedAlarmHostActor.ReassertConditionNodes`
in `OnAlarmsLoaded`, node-only, never the `alerts` topic. Two deviations from the sketch, both deliberate:
it reads a new `ScriptedAlarmEngine.GetProjections()` rather than `GetState(alarmId)` + `LoadedAlarmIds`
(a projection also carries the definition's severity, the resolved message template, and the last-observed
worst-input quality — `GetState` alone would have clobbered severity/message on the node); and it skips
alarms sitting in the Part 9 no-event position, so a deploy where nothing is in alarm writes no condition
nodes at all. The original deferral text follows for context.
The scripted-alarm Part 9 condition node has the same redeploy-reset race the VT `ReassertValue` fix closed:
`RebuildAddressSpace` clears `_alarmConditions` and `MaterialiseScriptedAlarms` recreates each condition
fresh/normal, but `ScriptedAlarmEngine.LoadAsync` reloads from the persisted state and yields
@@ -30,10 +30,10 @@ it reaches 📝.
| Wave | Deliverable | Design | Impl. plan | Status | Effort | Fixture / hardware |
|---|---|---|---|---|---|---|
| **0** | Universal Discover-backed browser | [design](2026-07-15-universal-discovery-browser-design.md) | [plan](2026-07-15-universal-discovery-browser-implementation.md) · [tasks](2026-07-15-universal-discovery-browser-implementation.md.tasks.json) | 🟡 **Live gate open** | SM | none (retrofits shipped drivers) |
| **1** | Modbus RTU (extend Modbus) | [design](2026-07-15-modbus-rtu-driver-design.md) | [plan](2026-07-24-modbus-rtu-driver.md) · [tasks](2026-07-24-modbus-rtu-driver.md.tasks.json) | 📝 **Plan ready** (11 tasks) | **S (lowest)** | `pymodbus` `rtu_over_tcp` — CI-simulatable, no new infra |
| **1** | Modbus RTU (extend Modbus) | [design](2026-07-15-modbus-rtu-driver-design.md) | [plan](2026-07-24-modbus-rtu-driver.md) · [tasks](2026-07-24-modbus-rtu-driver.md.tasks.json) | 🟢 **Built — live gate PASSED** (11 tasks, branch `feat/modbus-rtu-driver`, pending merge) | **S (lowest)** | `pymodbus` `rtu_over_tcp` — CI-simulatable, no new infra |
| **1** | SQL poll | [design](2026-07-15-sql-poll-driver-design.md) | [plan](2026-07-24-sql-poll-driver.md) · [tasks](2026-07-24-sql-poll-driver.md.tasks.json) | 📝 **Plan ready** (22 tasks) | SM | **existing central SQL Server** `10.100.0.35,14330` + SQLite unit |
| **2** | MTConnect Agent | [design](2026-07-15-mtconnect-driver-design.md) | [plan](2026-07-24-mtconnect-driver.md) · [tasks](2026-07-24-mtconnect-driver.md.tasks.json) | 📝 **Plan ready** (23 tasks) | SM | Dockerized `mtconnect/cppagent` — CI-simulatable |
| **2** | MQTT / Sparkplug B | [design](2026-07-15-mqtt-sparkplug-driver-design.md) | [plan](2026-07-24-mqtt-sparkplug-driver.md) · [tasks](2026-07-24-mqtt-sparkplug-driver.md.tasks.json) | 📝 **Plan ready** (27 tasks, P1+P2) | ML | Mosquitto/EMQX + C# Sparkplug sim — CI-simulatable |
| **2** | MTConnect Agent | [design](2026-07-15-mtconnect-driver-design.md) | [plan](2026-07-24-mtconnect-driver.md) · [tasks](2026-07-24-mtconnect-driver.md.tasks.json) | **COMPLETE** — P1 Agent MVP, all 23 tasks + the 5-leg live gate PASSED (branch `feat/mtconnect-driver`, PR #506) | SM | Dockerized `mtconnect/agent:2.7.0.12` (not `cppagent`) — CI-simulatable, live at `10.100.0.35:5000` |
| **2** | MQTT / Sparkplug B | [design](2026-07-15-mqtt-sparkplug-driver-design.md) | [plan](2026-07-24-mqtt-sparkplug-driver.md) · [tasks](2026-07-24-mqtt-sparkplug-driver.md.tasks.json) | **COMPLETE** — P1 + P2 shipped, both live-gated (Tasks 026) | ML | Mosquitto TLS+auth **+ C# Sparkplug edge-node simulator**, live at `10.100.0.35:8883` |
> Waves 3+ (BACnet/IP, Omron) and the deferred MELSEC are tracked in the program design doc §6/§8,
> not here. Omron CIP is the program's sole real-hardware wire gate; BACnet's broadcast/BBMD leg is
@@ -92,16 +92,30 @@ marginal cost.
Both are **fully CI-simulatable with zero new hardware**; Modbus RTU needs no new infra at all, and
SQL poll reuses the always-on central SQL Server. This is the recommended next build.
### Modbus RTU — 📝 Plan ready
### Modbus RTU — 🟢 Built, live gate PASSED (branch `feat/modbus-rtu-driver`, pending merge)
- **Design:** [`2026-07-15-modbus-rtu-driver-design.md`](2026-07-15-modbus-rtu-driver-design.md)
- **Implementation plan:** [`2026-07-24-modbus-rtu-driver.md`](2026-07-24-modbus-rtu-driver.md) · [`.tasks.json`](2026-07-24-modbus-rtu-driver.md.tasks.json) — **11 tasks** (1 trivial / 5 small / 2 standard / 3 high-risk).
- **Implementation plan:** [`2026-07-24-modbus-rtu-driver.md`](2026-07-24-modbus-rtu-driver.md) · [`.tasks.json`](2026-07-24-modbus-rtu-driver.md.tasks.json) — **11 tasks** (1 trivial / 5 small / 2 standard / 3 high-risk), all complete via subagent-driven-development.
- **Scope:** extend the existing Modbus driver with a `Transport` selector (RTU CRC-16 + FC-aware
length, no TxId) + a `rtu_over_tcp` `pymodbus` docker profile + the AdminUI selector.
**Direct-serial transport is descoped** (user, 2026-07-15) — RTU-over-TCP via a serial→Ethernet
gateway is the only shipped mode, so there is no serial hardware gate.
- **Fixture:** add an `rtu_over_tcp` profile to the existing Modbus integration fixture. First P1
step is confirming the pymodbus simulator exposes the RTU framer on a TCP server.
- **Effort:** **S — the lowest on the roadmap.** Recommended first.
- **Fixture:** added the `rtu_over_tcp` profile (`framer=rtu`, host `:5021`, co-runs with `standard`)
to the Modbus integration fixture. **Unknown confirmed:** pymodbus 3.13's simulator DOES honour
`framer=rtu` on a TCP server (wire-probed a pure RTU FC03 response, no MBAP) — so the simple profile
path was taken; **no stdlib fallback server was needed.**
- **Verification (all green):** unit suite **344/344**; full solution build 0 errors; the 2 RTU
integration tests pass **live** against the real `framer=rtu` fixture (read HR5→5, write+readback
HR200←4242). Per-task review chain (spec + code reviewers) + a final whole-branch review, all
approved-to-merge.
- **Live `/run` gate — PASSED (2026-07-24, docker-dev rebuilt from this branch):**
- AdminUI `/raw` Modbus **driver** config modal renders the new **Transport** selector (with the
operator hint) — set to `RtuOverTcp`, saved.
- **Device Test Connect → `OK · 23 MS`** — the probe built a `ModbusRtuOverTcpTransport` via the
factory and spoke **FC03 over real RTU framing** to `10.100.0.35:5021`.
- Authored raw tags HR5 (r/o) + HR200 (r/w) → **deployment `06ec1632` Sealed** (all 6 nodes acked).
- Client.CLI on `opc.tcp://localhost:4840` (`multi-role`/`password`): **read** `ns=2;s=rtu-gate/Device1/HR5`
**5, Good**; **write** `ns=2;s=rtu-gate/Device1/HR200` = 4242 → readback **4242, Good**.
- **Effort:** **S — the lowest on the roadmap.** Delivered first, as recommended.
### SQL poll — 📝 Plan ready
- **Design:** [`2026-07-15-sql-poll-driver-design.md`](2026-07-15-sql-poll-driver-design.md)
@@ -120,37 +134,118 @@ SQL poll reuses the always-on central SQL Server. This is the recommended next b
Both CI-simulatable on the shared docker host, no hardware.
### MTConnect Agent — 📝 Plan ready
### MTConnect Agent — ✅ Done (P1 Agent MVP)
- **Design:** [`2026-07-15-mtconnect-driver-design.md`](2026-07-15-mtconnect-driver-design.md)
- **Implementation plan:** [`2026-07-24-mtconnect-driver.md`](2026-07-24-mtconnect-driver.md) · [`.tasks.json`](2026-07-24-mtconnect-driver.md.tasks.json) — **23 tasks**. Task 0 is the TrakHound-vs-hand-rolled client decision; browse-picker live-verify is gated on Wave-0 #468.
- **Scope:** P1 Agent MVP (`IDriver`+`ITagDiscovery`+`IReadable`+`ISubscribable`+probe+rediscover),
browse **free via the Wave-0 universal browser** (`SupportsOnlineDiscovery=true`, no browser code),
typed editor, `UNAVAILABLE→BadNoCommunication` mapping, ring-buffer re-baseline paging.
- **Fixture:** canned XML unit fixtures (bulk of coverage) + a dockerized `mtconnect/cppagent`
integration fixture (env-gated). Depends on Wave 0's browse seam being live-verified (#468).
- **Effort:** SM (≈11.5 wk with TrakHound, ≈2.53 wk hand-rolled).
- **Implementation plan:** [`2026-07-24-mtconnect-driver.md`](2026-07-24-mtconnect-driver.md) · [`.tasks.json`](2026-07-24-mtconnect-driver.md.tasks.json) — **23 tasks; 22 built, Task 21 (live `/run` gate) tracked separately in the plan file.**
- **Driver guide:** [`docs/drivers/MTConnect.md`](../drivers/MTConnect.md) — config keys, capability
surface, data-plane precedence rules, quality-code mapping, and (most load-bearing) the **Known
limitations** section.
- **Scope shipped:** `IDriver`+`ITagDiscovery`+`IReadable`+`ISubscribable`+`IHostConnectivityProbe`+`IRediscoverable`
(deliberately no `IWritable` — read-only Agent surface), browse **free via the Wave-0 universal
browser** (`SupportsOnlineDiscovery=true`, no browser project), typed AdminUI tag editor,
`UNAVAILABLE→BadNoCommunication` mapping, CONDITION-as-`String`, ring-buffer re-baseline paging
(both the sequence-gap and `OUT_OF_RANGE` legs).
- **Diverged from the design:** Task 0/6/7 dropped the planned TrakHound `MTConnect.NET-Common`/`-HTTP`
dependency entirely — the pinned packages ship no XML formatter and the HTTP stream client has no
injectable handler to test behind the driver's seam. Parsing is hand-rolled `System.Xml.Linq`,
matching on `LocalName` only (namespace-agnostic, so 1.32.x all parse). See the driver guide's
"Built vs. planned" section.
- **Test results:** MTConnect unit suite 491/491; integration suite 12/12 against a real Agent
(`mtconnect/agent:2.7.0.12`, **not** `mtconnect/cppagent` — that Docker Hub repo doesn't exist);
skips cleanly (12 Skipped) offline. AdminUI 749/749; full solution 0 build errors.
- **Deferred (documented, not built):** write-back via MTConnect Interfaces (design §3.6); P1.5
CONDITION→native Part-9 alarms, `TIME_SERIES` array materialization, EVENT→enum vocab (design §9);
P2 SHDR adapter ingest (design §9). Also two pre-existing fleet-wide gaps this build surfaced
rather than caused: `IRediscoverable`/`IHostConnectivityProbe` have no server-side consumer (eight
drivers besides Galaxy), and no driver form blocks Save on a required-field validation error.
- **Fixture:** canned XML unit fixtures (bulk of coverage) + a dockerized `mtconnect/agent` +
stdlib-SHDR-adapter integration fixture (env-gated), deployed at `/opt/otopcua-mtconnect/` on the
shared docker host. See the fixture's own
[`Docker/README.md`](../../tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.IntegrationTests/Docker/README.md).
- **Effort:** SM, hand-rolled (the TrakHound estimate in the design doc did not apply once that
path was ruled out).
### MQTT / Sparkplug B — 📝 Plan ready
### MQTT / Sparkplug B — **COMPLETE** (P1 + P2, both live-gated)
- **Design:** [`2026-07-15-mqtt-sparkplug-driver-design.md`](2026-07-15-mqtt-sparkplug-driver-design.md)
- **Implementation plan:** [`2026-07-24-mqtt-sparkplug-driver.md`](2026-07-24-mqtt-sparkplug-driver.md) · [`.tasks.json`](2026-07-24-mqtt-sparkplug-driver.md.tasks.json) — **27 tasks** (P1 plain MQTT = Tasks 014, a complete shippable milestone; P2 Sparkplug B = Tasks 1526). Task 0 is the mandatory MQTTnet-5/net10 + central-pinning validation spike.
- **Scope:** P1 plain MQTT (MQTTnet-5 connect/TLS/auth + hand-rolled reconnect, subscribe→
- **Implementation plan:** [`2026-07-24-mqtt-sparkplug-driver.md`](2026-07-24-mqtt-sparkplug-driver.md) · [`.tasks.json`](2026-07-24-mqtt-sparkplug-driver.md.tasks.json) — **27 tasks, all done** (P1 plain MQTT = Tasks 014; P2 Sparkplug B = Tasks 1526).
- **Scope delivered:** P1 plain MQTT (MQTTnet-5 connect/TLS/auth + hand-rolled reconnect, subscribe→
`OnDataChange`, retained last-value read, `#`-observation browser); P2 Sparkplug B ingest
(vendored Tahu proto, birth/alias/seq-gap/rebirth state machine, death→STALE).
- **Fixture:** Mosquitto/EMQX broker (TLS + real auth, not anonymous) + a **project-owned C#
Sparkplug edge-node simulator** for the rebirth/seq-gap/death test matrix + env-gated live suite
(`MQTT_FIXTURE_ENDPOINT`).
- **Effort:** ML. **Top risks:** Sparkplug state-machine correctness; MQTTnet-5/net10 + the repo's
fragile central pinning (mitigated by hand-rolling Tahu over one MQTTnet-5 client).
(vendored Tahu proto + `Grpc.Tools` codegen, birth/alias/seq-gap/rebirth state machine,
death→STALE, birth-driven browse tree, scoped Request-rebirth NCMD, typed tag editor).
**Write-through (`IWritable`) is deferred to P3** — every MQTT node materializes read-only.
- **Fixture:** Mosquitto TLS+auth broker + JSON publisher sidecar, **live** at `10.100.0.35:8883`
(`:1883` plaintext-but-authenticated); stack `/opt/otopcua-mqtt`, compose in
`tests/Drivers/…​.Driver.Mqtt.IntegrationTests/Docker/`. Env-gated live suite
(`MQTT_FIXTURE_ENDPOINT`, see `infra/README.md` §3). The **project-owned C# Sparkplug edge-node
simulator** (`--profile sparkplug`, group `OtOpcUaSim`, nodes `EdgeA`/`EdgeB`, `EdgeA` device
`Filler1`) shipped with P2 and answers rebirth NCMDs.
- **P2 status (2026-07-25):** Tasks 1526 complete. **Offline 1528 passed / 0 failed**
(581 driver · 809 AdminUI · 138 Core.Abstractions); **live 15/15** against the broker + simulator.
- **P1 live gate found two AdminUI authoring gaps** (the gate's whole point — both invisible to green
unit tests, because AdminUI has no bUnit and nothing tied these hand-maintained lists to
`DriverTypeNames`):
1. **FIXED —** `RawDriverTypeDialog`'s option array never got an MQTT row, so **no operator could
create an MQTT driver at all**. Fixed, plus a reflection parity guard
(`RawDriverTypeDialogParityTests`) that fails for *any* future `DriverTypeNames` entry
missing from the picker.
2. **FIXED —** `MqttDriverForm` / `MqttDeviceForm` did not exist, so the broker connection was
unauthorable from the AdminUI. Both shipped after the P1 gate.
- **P2 live gate found a third one, of the same family — FIXED:** `MqttDriverForm` still carried its
P1 Sparkplug **placeholder**. Switching Mode to `SparkplugB` rendered a *"not available yet"* notice
and **no Group ID field** — so once Sparkplug ingest shipped there was still **no way to author a
Sparkplug driver from the AdminUI**, and `Sparkplug.GroupId` is the driver's entire subscription
filter (`spBv1.0/{GroupId}/#`): blank ⇒ connected, `Healthy`, ingesting nothing. Fixed (all five
Sparkplug keys, merge-not-replace on save, group-id segment validation) + 8 falsifiable
`MqttDriverFormModelTests` cases. **Three gates, three instances of the same defect class: a
hand-maintained AdminUI surface left behind by a driver-side feature.**
- **P2 live gate — two open UI gaps, documented not fixed** (see `docs/drivers/Mqtt.md` §Known gaps):
- **A completely empty browse tree cannot arm a rebirth.** The scope comes from a tree click, so on
a plant that has not birthed since the window opened `Request rebirth…` stays disabled — the very
case the affordance exists for. `MqttBrowseSession.ResolveRebirthTargets` already accepts a bare
`{group}/{edgeNode}` ("the prime rebirth target"); only a UI path to enter one is missing.
Mitigation shipped: a **Refresh** button on the browse tree (it re-reads the same session, which
previously rendered exactly once at open and never again).
- **P2 live gate also found a PRE-EXISTING, cross-driver AdminUI defect — FIXED:** every node label in
the shared `DriverBrowseTree` was an `<a href="#">`. In a Blazor Web App, `blazor.web.js`'s
enhanced-navigation click interceptor resolves a bare `#` against `<base href="/">`, so **clicking
any browse-tree node label navigated the whole AdminUI to `/`**, tore down the circuit and destroyed
the hosting modal — losing the browse session and the tag selection. `@onclick:preventDefault` does
not help: it suppresses the browser's default action, not Blazor's own interceptor. The labels are
now `<button type="button">`. It dates to the component's introduction (2026-05-28) and affects
**every** driver's picker; it only surfaced now because Sparkplug's Request-rebirth is the first
feature that requires clicking a *label* rather than the ▶ toggle (always a `<button>`, always fine).
- **A metric name containing `/` cannot be browse-committed.** The derived raw **tag Name** is a
RawPath segment and may not contain `/`, while `metricName` legitimately may — so the
spec-mandatory `Node Control/Rebirth` is refused at commit (`Row N: Name must not contain '/'`).
The refusal is loud and all-or-nothing, never a silently mis-bound tag, and **Manual entry** is
the working path (the Sparkplug tag editor accepts a slashed metric name). Fixing it needs a
name-sanitisation policy with a collision answer, so it was recorded rather than guessed at.
- `_canRebirth` is captured at browse-open, so a reaped session still draws the button.
- The tag editor injects the Plain-only `payloadFormat` key into a Sparkplug blob — cosmetic; the
factory routes on the Sparkplug tuple and ignores it.
- **Rediscovery is inert in v3 (platform gap, not MQTT's).** The Sparkplug driver raises
`OnRediscoveryNeeded` on a changed birth metric-set, but **nothing subscribes to it** and
`DriverHostActor.HandleDiscoveredNodes` hard-returns. A DBIRTH introducing a metric does **not**
change the OPC UA tree — redeploy. Driver-side behaviour is log-observable only.
- **Redundant-pair hazard (documented, not a code change):** both nodes of a pair run the driver, so a
**fixed `clientId` makes them evict each other forever** — the broker logs
`already connected, closing old connection` and both nodes reconnect every ~2 s while still
reporting `Healthy`. Unset (the default) is correct. `MqttDriverForm` warns inline the moment a
value is typed.
- **Effort:** ML, delivered. The MQTTnet-5/net10 + central-pinning risk is **retired** (Task 0's
spike held through both phases); Sparkplug state-machine correctness was carried by golden-payload
vectors + the live simulator.
---
## Next actions
1. **Close the Wave-0 live gate** (Gitea #468) — unblocks browse verification for MTConnect/BACnet.
1. **Close the Wave-0 live gate** (Gitea #468) — unblocks browse verification for BACnet (MTConnect
shipped its own Task-21 live `/run` gate independently — see the plan file).
2. **Build Wave 1** — Modbus RTU first (lowest effort, no new infra), then SQL poll. Both plans are
📝 ready; run the command from the [Executing a plan](#executing-a-plan-git-worktree--subagent-per-task)
table (subagent-driven-development in a git worktree).
3. **Build Wave 2** (MTConnect, then MQTT/Sparkplug) — both plans 📝 ready. MTConnect's browse leg
wants #468 closed first; MQTT's Task 0 pinning spike gates everything after it.
3. **Wave 2 is done.** Both deliverables shipped and live-gated: **MQTT / Sparkplug B** (P1 + P2) and **MTConnect Agent** (P1 Agent MVP, PR #506). Remaining work on both is optional follow-on — MQTT P3 write-through (`IWritable`: NCMD/DCMD + plain publish) plus its two browse-UI
gaps, and MTConnect write-back (deliberately deferred; the Agent surface is read-only).
Update this file's Summary table and per-wave status whenever a deliverable changes state.
+317
View File
@@ -48,6 +48,102 @@ git add docs/plans/2026-07-24-mtconnect-driver.md
git commit -m "docs(mtconnect): record TrakHound-vs-hand-rolled client decision (Task 0)"
```
**DECISION (verified 2026-07-24): TrakHound path.** All three checks passed against real nuget.org.
(1) `dotnet package search MTConnect.NET-Common --exact-match` and `...-HTTP --exact-match` both list
versions `3.2.0` through `6.9.0.2` on `nuget.org`, confirming `6.9.0.2` is the current top-of-list
version for both package ids (no version mismatch between the two). (2) A throwaway `net10.0` console
project in the scratchpad (outside the repo, so unaffected by this repo's `packageSourceMapping`) ran
`dotnet add package MTConnect.NET-Common -v 6.9.0.2` + `...-HTTP -v 6.9.0.2` + `dotnet restore` — both
restored cleanly with no source-mapping or compatibility errors; `project.assets.json` resolved
`MTConnect.NET-Common`, `MTConnect.NET-HTTP`, and the transitive `MTConnect.NET-TLS` (all `6.9.0.2`)
against the `net10.0` target framework, and each package's `lib/` folder contains a `netstandard2.0`
build (alongside `net6.0``net9.0`, `net461``net472`, `net47`, `net48`), so net10.0 resolves via the
in-box compatibility mapping. (3) The pinned version carries **no embedded LICENSE file** in the
`.nupkg` (`unzip -l` shows no `LICENSE*` entry in either package) — instead both nuspecs declare the
modern SPDX license-expression form `<license type="expression">MIT</license>` +
`<licenseUrl>https://licenses.nuget.org/MIT</licenseUrl>`, which nuget.org validates against the SPDX
list at push time and surfaces on the package page (confirmed live: the nuget.org page for
`MTConnect.NET-Common` `6.9.0.2` displays "MIT license" linked to `licenses.nuget.org/MIT`) — a
structurally stronger guarantee than a free-text embedded file for this specific version, so the
absence of a physical `LICENSE` file is not a fail. Task 1 adds
`PackageReference MTConnect.NET-Common 6.9.0.2` + `PackageReference MTConnect.NET-HTTP 6.9.0.2` to
`Driver.MTConnect.csproj`; the hand-roll fallback is not needed.
**CORRECTION (Task 6, commit `bdfefadc`) — the decision above rested on incomplete evidence, and the
hand-roll fallback IS in use for parsing.** Task 0 verified the packages *exist, restore, and are MIT*;
it never verified they can actually parse an MTConnect document. As pinned, they cannot. Task 6 proved
this empirically before writing any parser code: reflection over both referenced assemblies
(`MTConnect.NET-Common`, 1297 types; `-HTTP`, 410 types) found **zero `IResponseDocumentFormatter`
implementations and zero public static `FromXml`/parse entry points**, and a live invocation of
TrakHound's own entry point (`ResponseDocumentFormatter.CreateDevicesResponseDocument("xml", stream)`)
against this repo's `Fixtures/probe.xml`, with both DLLs loaded, returned:
```
Success = False
ERROR: Document Formatter Not found for "xml"
```
TrakHound 6.x resolves formatters from loaded assemblies at runtime, and the XML formatter ships in a
**third package, `MTConnect.NET-XML`**, which Task 0 did not evaluate. Even with it added, TrakHound
exposes no socket-free `Parse(string)` entry point of the shape this plan mandates. So `/probe` parsing
is a hand-rolled `System.Xml.Linq` walk (~250 LoC, `MTConnectProbeParser`); the driver above the
`IMTConnectAgentClient` seam is unaffected, exactly as the decision rule anticipated.
**The two `PackageReference`s remain and are, so far, dead weight.** Task 6 deliberately did not remove
them: the hard remaining problem is Task 7's `multipart/x-mixed-replace` chunk framing for `/sample`,
and `MTConnect.Clients.MTConnectHttpClientStream` in `-HTTP` may be usable for framing alone. **Task 7
owns the final call:** (a) use `-HTTP` for framing only, (b) add `MTConnect.NET-XML` and reconsider
wholesale, or (c) hand-roll the boundary reader and **drop both package references** — in which case the
Task-0 rationale comment in `Directory.Packages.props` goes with them. Record the outcome here.
**OUTCOME (Task 7) — option (c): hand-rolled boundary reader, BOTH package references DROPPED.**
`MTConnect.NET-Common`, `MTConnect.NET-HTTP` and the transitive `MTConnect.NET-TLS` are gone from
`Directory.Packages.props` and from `Driver.MTConnect.csproj`; the Task-0 rationale comments went with
them (each replaced by a comment naming this block). **The driver now depends on nothing but the BCL**
`HttpClient` + `System.Xml.Linq`. Reflection over `MTConnect.NET-HTTP` 6.9.0.2 (net9.0 lib, with
`-Common` resolved alongside) settled the one open candidate:
```
TYPE MTConnect.Clients.MTConnectHttpClientStream
.ctor(String url, String documentFormat)
Void Start(CancellationToken), Void Stop(), Task Run(CancellationToken)
event DocumentReceived / ErrorReceived / FormatError / ConnectionError / …
```
Three independent disqualifiers, any one of which is fatal:
1. **It cannot be exercised without a socket.** Its only constructor takes a *URL* and it dials the
connection itself inside `Run` — there is no `HttpMessageHandler`, no `HttpClient`, and no `Stream`
injection point. The whole `IMTConnectAgentClient` seam exists so Tasks 613 are unit-testable
against canned XML with no listener; adopting this type would have forced a live agent (or a real
loopback HTTP server) into the unit suite.
2. **"Framing alone" was never actually on offer.** The type does not surface raw part payloads — it
raises `DocumentReceived` with an already-parsed `IDocument`, resolved through the same
`documentFormat` formatter lookup that Task 6 proved is missing. Using it for framing would still
have required adding `MTConnect.NET-XML`, i.e. option (b) in disguise.
3. **The shape is inverted and the payload is heavy.** An event-driven `Start`/`Run` object has to be
adapted back into the `IAsyncEnumerable<MTConnectStreamsResult>` the seam declares, and `-HTTP`
vendors an entire embedded web server (`Ceen.*` types) that this driver would never touch.
The replacement is `MultipartStreamReader`, a ~200-line private nested class in
`MTConnectAgentClient.cs` that frames `multipart/x-mixed-replace` off the response stream. Two paths:
`Content-length` present (every agent in the wild writes it) → read by length and emit the chunk the
instant it completes; absent → scan to the next boundary, which necessarily emits one chunk late and
exists as a correctness fallback only. Both paths are covered by tests. Every read is bounded by a
heartbeat watchdog (`2 × HeartbeatMs + RequestTimeoutMs`) so a silent agent faults instead of wedging
the pump — the R2-01 shape — and the buffer is capped at 32 MiB.
**Also settled here: the streaming leg gets its own `HttpClient`.** `HttpClient.Timeout` is a
per-instance whole-response deadline, so the long poll cannot share the unary client. Raising the one
timeout would have un-bounded `/probe` + `/current`, which R2-01 forbids; instead `_streamHttp` runs
`Timeout.InfiniteTimeSpan` and is bounded by the request deadline on its *header* phase plus the
watchdog on its body.
**Process lesson:** "the dependency restores" is not "the dependency does the job." A library
verification checklist needs one end-to-end call against a real input, not just a restore. Task 7 adds
a second: **check the constructor for a test seam before counting a library as usable** — a type that
can only be handed a URL cannot participate in a socket-free unit suite, however good its internals.
---
## Task 1: Scaffold the two driver projects + the test project
@@ -289,6 +385,57 @@ dotnet test tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.Tests --filter "Ful
git add -A && git commit -m "feat(mtconnect): current/sample parse + sequence-gap detection (Task 7)"
```
### Task 7 review remediation (post-`c46540ae`) — three contract changes Tasks 9/11/13 must build against
Code review of `c46540ae` approved the framing algorithm but moved the risk onto the **contract**. Three
changes here are breaking relative to the sketch above; the rest are hardening.
1. **`SampleAsync` never ends normally except by cancellation.** The seam documents the stream as
yielding "indefinitely, until `ct` is cancelled", but the implementation returned normally on EOF, on
the closing `--boundary--`, and — worst — on a response that was not multipart at all, where it
yielded one parsed snapshot and finished. A pump written to the documented contract has no reason to
handle normal completion, so it would silently drop its subscription or hot-loop reconnecting; the
non-multipart leg additionally reported a **configuration error as a healthy finished stream**, which
is the #485 quiet-successful-termination shape aimed straight at Task 11. Every non-cancellation end
now throws: `MTConnectStreamEndedException` (with `MTConnectStreamEndReason.ConnectionClosed` /
`ClosingBoundary` — transient, reconnect) or `MTConnectStreamNotSupportedException` (configuration —
reconnecting reproduces it forever), sharing the base `MTConnectStreamException`. The non-multipart
body is still parsed first, so an Agent's own `MTConnectError` text still wins, but the document is
**not** yielded.
2. **`IsSequenceGap` moved to `IMTConnectAgentClient` (static) and its first parameter is now
`expectedFrom`.** The sketch's `requestedFrom` name and doc were wrong for every chunk after the
first: the correct comparand is the **previous chunk's `NextSequence`**, and comparing against the
opening `from` reports a gap on every chunk once the ring buffer rolls — an endless `/current`
re-baseline storm against a healthy stream. Rehomed onto the seam so Task 11 need not reference the
concrete client. **Task 11 must also treat an `OUT_OF_RANGE` `InvalidDataException` as a re-baseline
trigger:** cppagent answers a `from` below `firstSequence` with an `MTConnectError` under HTTP 200,
so ring-buffer overflow arrives as a *parse failure*, never as a gap-bearing chunk. `IsSequenceGap`
alone does not cover it. Documented on the seam.
3. **`MTConnectObservation` gained `IsStructured`** (default `false`), set from `element.HasElements`.
A DATA_SET/TABLE observation's `<Entry key=…>` children concatenate through `element.Value` to
nonsense (`"12"` for two entries) which the index would otherwise publish as Good, and only the
parser can tell that apart from a legitimate space-bearing `Message` — no value-shape heuristic
works. Deliberately **false for CONDITION** observations: their value comes from the element *name*,
so child content cannot corrupt it. The index maps the flag to a status; the parser does not.
Hardening in the same pass: disposal mid-enumeration now surfaces as `OperationCanceledException`
(via an internal dispose-linked token) rather than a raw `ObjectDisposedException`/`IOException` that
Task 9's re-init would otherwise inflict on an enumerating pump; `HeartbeatMs`, `SampleIntervalMs` and
`SampleCount` join `RequestTimeoutMs` in construction-time positive-value validation (all three reach
the query string, and an Agent answers `heartbeat=0` with an HTTP-200 `MTConnectError` that would name
no config key); a `multipart/*` response with no `boundary` parameter fails fast instead of degrading
into a watchdog timeout; the no-`Content-length` framing fallback logs a one-shot warning (the client
now takes an optional `ILogger`); and the shared element/attribute reading rules moved into
`MTConnectXml`, used by both parsers.
**Coverage gap closed, and worth remembering as a pattern.** Every multipart test served the whole body
from one buffer, so **every framing test completed in a single `ReadAsync`** — the split-boundary path,
the split-header path and the multi-fill loop were correct by inspection only, on the task whose
headline risk is framing. A `ChunkedStream` double returning N bytes per read (N = 1, 3, 7) now drives
the fixtures through a split transport. Falsifiability confirms the gap was real: removing the
split-boundary overlap, removing the split-header overlap, and collapsing `EnsureAsync`'s fill loop each
fail **only** the new chunked tests — every pre-existing multipart test stays green under all three.
---
## Task 8: `MTConnectObservationIndex` + `UNAVAILABLE → BadNoCommunication`
@@ -747,6 +894,176 @@ git commit -m "test(mtconnect): live /run gate on docker-dev — browse/read/sub
---
### LIVE-GATE RESULT (2026-07-24, completed 2026-07-27) — ALL 5 LEGS PASSED
**Rig.** The shared `otopcua-dev` rig was NOT used: three sibling driver worktrees
(`modbus-rtu`, `mqtt-sparkplug`, `sql-poll-driver`) share its `otopcua-host:dev` image, and it had
been up for hours. Instead an isolated stack was built from this worktree —
`otopcua-host:mtconnect`, compose project `otopcua-mtc`, overlay
`docker-dev/docker-compose.mtconnect.yml`, MAIN pair only, ports AdminUI **9220** / OPC UA **4860**
/ SQL **14340**. Mirrors how the mqtt worktree isolated itself. The `otopcua-dev` rig and every
sibling worktree were left untouched.
**Fixture.** `mtconnect/agent:2.7.0.12` + the SHDR adapter deployed to
`/opt/otopcua-mtconnect/` on the shared docker host and brought up healthy; `http://10.100.0.35:5000`
serves the seeded `OtFixtureCnc` model with live-moving values. Note it answers in the MTConnect
**2.0** namespace while the canned fixtures are 1.3 — the namespace-agnostic `LocalName` parsing
(Task 6) is what makes both work, now demonstrated rather than argued.
| Leg | Result | Evidence |
|---|---|---|
| Integration suite vs. a REAL Agent | ✅ **12/12** | incl. `Discovery_types_an_upper_snake_numeric_event_as_Int64` (the `PART_COUNT` regression), `Read_and_subscribe_key_on_RawPath_when_the_artifact_supplies_RawTags` (the blocker fix), `Subscribe_delivers_a_changed_value_from_the_live_sample_stream`, `Discovery_binds_FullName_to_the_DataItem_id_when_the_name_differs` |
| 1. Driver creatable in `/raw` | ✅ | `MTConnect` present in `RawDriverTypeDialog`; `mtc-1` created. Proves the Host `Register` call is genuinely wired — **the one thing no test guards** (deleting it leaves the guard test green and substitutes a `StubbedDriver` while the deploy still seals green). |
| 2. Typed config modal renders | ✅ | "Configure driver · MTConnect" with all four sections (Agent / Polling / Connectivity probe / Reconnect), **not** the "no typed config form" banner. **This is the mutation that survived all 749 AdminUI tests** — the `switch` in `DriverConfigModal.razor` compiles to `BuildRenderTree` and is unreachable by reflection. |
| 3. Config persists + round-trips | ✅ | Stored `{"agentUri":"http://10.100.0.35:5000","requestTimeoutMs":5000,…,"probe":{"enabled":true,…},"reconnect":{…}}`; reopening the modal displays the saved URI (Razor binding verified — no bUnit in this repo). |
| 4. Browse picker → commit | ✅ | "Connect & browse" dialled the live Agent; tree rendered `Agent` + `OtFixtureCnc` with `Favail` (name ≠ id) beside id-named leaves, proving `browseName = Name ?? Id`. Committed TagConfig is **`{"fullName":"fixture_asset_changed","dataType":"String"}`** — the `fullName`-not-`address` fix, live. Wave-0 #468 did **not** block this. |
| 5. Deploy → OPC UA read/subscribe | ✅ **PASSED (2026-07-27, post-rebase)** | see "Leg 5" below |
---
#### Leg 5 — PASSED 2026-07-27, on the rebased branch
Re-run after the rebase onto master (`123ddc3f`), against a **freshly rebuilt** `otopcua-host:mtconnect`
image, so this gate exercises the merged code and not the pre-rebase tree.
**The cookie blocker was sidestepped, not worked around.** Rather than moving hosts or stopping the
`otopcua-dev` rig, the deploy went through the **headless deploy API**
`POST /api/deployments` with `X-Api-Key` (`DeployApiEndpoints`, `Security:DeployApiKey`). It is
`AllowAnonymous().DisableAntiforgery()` by design ("machine endpoint, not a browser form post"), so
the antiforgery cookie collision is structurally irrelevant to it. This is the better recipe for any
future isolated-stack gate — it needs no browser at all:
```bash
curl -s -X POST http://localhost:9220/api/deployments \
-H "X-Api-Key: docker-dev-deploy-key" -H "Content-Type: application/json" \
-d '{"CreatedBy":"mtconnect-live-gate"}'
# {"outcome":"Accepted","deploymentId":"…","revisionHash":"…"} HTTP 202
```
Both deployments sealed `Status = 2` (`DeploymentStatus.Sealed`) with `FailureReason = NULL` — i.e.
**both** MAIN nodes acked. No `MaintenanceMode` hatch was needed: the isolated stack's topology is
MAIN-only with both central nodes enabled and running.
**Tag set.** Leg 4's AdminUI-committed tag (`fixture_asset_changed`) is an `AssetChanged` EVENT that
never moves, so three live-changing tags were added **directly via SQL** (deliberately bypassing the
typed editor — see the finding below) to exercise all three type paths.
**Results — read (`Client.CLI read`, `opc.tcp://localhost:4860`):**
| NodeId (`ns=2`, Raw realm) | Value | Status |
|---|---|---|
| `s=mtc-1/vmc/fixture_partcount` | `48099` (Int64) | `0x00000000` Good |
| `s=mtc-1/vmc/fixture_x_pos` | `96.4416` (Float64) | `0x00000000` Good |
| `s=mtc-1/vmc/fixture_execution` | `ACTIVE` (String) | `0x00000000` Good |
| `s=mtc-1/vmc/fixture_asset_changed` | *(null)* | `0x80000000` — see finding 6 |
The RawPath is `mtc-1/vmc/<dataItemId>` (driver/device/tag; no folder prefix), confirmed by
`browse -r`, which rendered all four as `[Variable]` under the `vmc` device object.
**Results — subscribe (`Client.CLI subscribe -r` over the whole device, 500 ms):** all four monitored;
live updates flowed continuously from the `/sample` long-poll pump through to OPC UA — `fixture_x_pos`
tracing its sinusoid (`103.66 → 123.18 → 141.90 → 155.27 → 159.99 → 154.93 → 141.32 → 122.48 → 103.03
→ 87.75 → 80.35`), `fixture_partcount` incrementing monotonically (`48128 → 48129 → 48130`), and
`fixture_execution` transitioning `READY → INTERRUPTED`. **This closes the last unproven link**: the
driver seam was already covered by the live integration suite, and leg 5 adds `DeploymentArtifact`
address-space materialisation → OPC UA publish on top of it.
**Additional findings from leg 5:**
5. **A bad `dataType` degrades to exactly one skipped tag, loudly — verified by accident.** The
hand-written SQL used `"dataType":"Double"`; the vocabulary is `DriverDataType`, whose member is
**`Float64`**. The driver logged, per tag:
> `could not map the raw tag mtc-1/vmc/fixture_x_pos to an Agent DataItem; it names no readable id
> (expected 'fullName', 'dataItemId' or 'address') or declares an unrecognised driverDataType. The
> tag is skipped and reports BadNodeIdUnknown; every other tag is unaffected.`
…and the other three tags kept streaming. That is the intended fail-isolated behaviour,
demonstrated live rather than argued. Correcting the row to `Float64` and redeploying turned the
tag Good with **no container restart** — which incidentally shows **MTConnect is not subject to the
separately-tracked config-edits-silently-discarded defect** that affects Modbus/FOCAS/OpcUaClient.
Note also that the typed AdminUI editor would have prevented the mistake outright (it offers a
`DriverDataType` dropdown); the error was an artifact of authoring straight into SQL.
6. **`fixture_asset_changed` reading `0x80000000` is NOT an MTConnect defect — it is a pre-existing,
deliberate fidelity gap on the shared publish path.** The agent does return the tag in `/current`
as `<AssetChanged …>UNAVAILABLE</AssetChanged>`, and the driver maps it correctly to
`BadNoCommunication` (`0x80310000`, `MTConnectObservationIndex.cs:255`). The node nevertheless
reports generic `Bad` (`0x80000000`) — while carrying the agent's exact source timestamp, which
proves the publish landed and only the status differs. Cause, verified by reading the source:
`DriverInstanceActor.QualityFromStatus` (`DriverInstanceActor.cs:1032-1041`) projects the driver's
`uint` onto the 3-state `OpcUaQuality` enum using only the top two severity bits (`statusCode >> 30`),
and `OtOpcUaNodeManager.StatusFromQuality` (`OtOpcUaNodeManager.cs:3318-3322`) re-expands `Bad` to
`StatusCodes.Bad`. It is intentional — `OpcUaQuality`'s own doc comment says *"Real SDK has
finer-grained codes; the engine actors only need this 3-state classification."*
The consequence appears undocumented and is client-visible: **no driver's Bad/Uncertain sub-code
ever reaches an OPC UA client.** Issue #497's 16 corrected constants and the new
`StatusCodeParityTests` guard keep the constants internally right, but a client cannot tell
`BadNoCommunication` from `BadTypeMismatch` from `BadNotSupported`. (The one survivor is
`BadWaitingForInitialData`, written directly at the node at `OtOpcUaNodeManager.cs:1806/1901`,
bypassing the projection.) Tracked separately; out of scope for this plan.
---
**Historical — why leg 5 was blocked on 2026-07-24 (superseded by the API-key recipe above).**
Browser cookies are scoped to a
**host, not a port**, so the AdminUI antiforgery cookie issued by the still-running `otopcua-dev`
rig on `localhost:9200` is sent to the isolated stack on `localhost:9220`, whose data-protection
keys differ:
```
Microsoft.AspNetCore.Antiforgery.AntiforgeryValidationException: The antiforgery token could not be decrypted.
```
The `Deploy current configuration` POST is rejected before reaching any MTConnect code — no
`Deployment` row is created, and the UI surfaces nothing. Clearing JS-visible cookies does not help
(the antiforgery cookie is `HttpOnly`). This is a two-stacks-on-one-host collision, entirely outside
the driver.
The three browser-side remedies originally listed (distinct host / incognito profile / stop the
`otopcua-dev` rig) all remain valid, but none was needed — the headless deploy API above avoids the
browser entirely and is the recommended recipe.
**Incidental findings from the rig work:**
1. ~~`ClusterNode` has no `MaintenanceMode` column, so CLAUDE.md is ahead of the migration.~~
**CORRECTED — this was my error, not a repo inconsistency.** The migration
`20260722125506_AddClusterNodeMaintenanceMode` exists on master, along with `AkkaPort`/`GrpcPort`.
**This branch is 39 commits behind master** (base `963eec1b`, master `28c28667`) and simply
predates them, so the isolated stack's schema lacked the column and I reduced the seeded topology
by deleting the SITE-A/SITE-B rows instead. CLAUDE.md is accurate. **DONE 2026-07-27 — rebased**
onto master `123ddc3f` (67 commits by then, not 39). Conflicts were all of one shape — this driver
and the newly-merged SQL-poll driver each adding their entry to the same list — resolved by keeping
**both** in: `ZB.MOM.WW.OtOpcUa.slnx`, `DriverTypeNames.cs`, `DriverFactoryBootstrap.cs`,
`Core.Abstractions.Tests.csproj`, `TagConfigEditorMap.cs`, `TagConfigValidator.cs`,
`DriverConfigModal.razor`, `RawDriverTypeDialog.razor`. Post-rebase: solution build 0 errors,
MTConnect 491/491, Core.Abstractions 256/256, AdminUI 800/800. For a future partial-topology gate,
`MaintenanceMode = 1` is now available and is the correct hatch (it is what the SQL-poll driver's
gate used), in place of the row deletion I resorted to.
2. `sqlcmd` against this schema needs `-I` (`SET QUOTED_IDENTIFIER ON`); without it every DML
statement fails with `Msg 1934` because of the filtered/computed indexes.
3. **Status-code parity is pre-verified against master's new guard.** Master gained
`StatusCodeParityTests`, which reflects over every `ZB.MOM.WW.OtOpcUa.Driver.*.dll` in its bin for
status-shaped `const uint` fields — including `private const` on `internal` types — and checks each
against `Opc.Ua.StatusCodes`. Task 16 added the `Driver.MTConnect` ProjectReference to that same
test project, so this driver's constants are in scope. All eight were verified against the pinned
SDK assembly directly: `BadCommunicationError 0x80050000`, `BadNoCommunication 0x80310000`,
`BadNodeIdUnknown 0x80340000`, `BadNotConnected 0x808A0000`, `BadNotSupported 0x803D0000`,
`BadOutOfRange 0x803C0000`, `BadTypeMismatch 0x80740000`, `BadWaitingForInitialData 0x80320000`
**0 mismatches**. Note the guard only sees *named* constants; an inline status literal at a call
site is invisible to it, so keep hoisting them. **Confirmed after the rebase**: the guard reports
9 MTConnect constants in scope (the 8 above plus `Good`), all passing — the pre-verification held.
4. **The separately-tracked `BadTypeMismatch` defect in FOCAS/TwinCAT/AbLegacy/AbCip is already fixed
on master** (issue #497 grew it to 16 wrong constants across 6 drivers). This driver independently
chose the correct `0x80740000` and pinned it with a mutation test, so the two agree; no action.
**Teardown when done:**
```bash
docker compose -p otopcua-mtc -f docker-dev/docker-compose.yml -f docker-dev/docker-compose.mtconnect.yml down -v
ssh dohertj2@10.100.0.35 'cd /opt/otopcua-mtconnect && docker compose down'
```
---
## Task 22: Docs + deferred-writeback note; update the tracking doc
**Classification:** trivial
+21 -2
View File
@@ -37,7 +37,7 @@ ssh dohertj2@10.100.0.35 'docker ps --filter label=project=lmxopcua --format "{{
| Stack dir | Purpose |
|---|---|
| `/opt/otopcua-mssql` | central SQL (always-on) |
| `/opt/otopcua-modbus` · `/opt/otopcua-abcip` · `/opt/otopcua-s7` · `/opt/otopcua-opcuaclient` | driver fixtures |
| `/opt/otopcua-modbus` · `/opt/otopcua-abcip` · `/opt/otopcua-s7` · `/opt/otopcua-opcuaclient` · `/opt/otopcua-mqtt` · `/opt/otopcua-mtconnect` | driver fixtures |
| `~/otopcua-ablegacy` · `~/otopcua-focas` | driver fixtures (user-owned) |
| `~/otopcua-harness` | Host.IntegrationTests real-mode deps (SQL + GLAuth) — see §4 |
@@ -69,9 +69,28 @@ when it isn't running. Bring one up, `dotnet test`, tear down. Repo compose live
| S7 | `otopcua-python-snap7` | `10.100.0.35:1102` | `docker compose --profile s7_1500 up -d` |
| OpcUaClient | `mcr.microsoft.com/iotedge/opc-plc:2.14.10` | `opc.tcp://10.100.0.35:50000` | `docker compose up -d` |
| FOCAS | `otopcua-focas-sim` (vendored focas-mock) | `10.100.0.35:8193` | `docker compose up -d` (stack `~/otopcua-focas`) |
| MQTT | `eclipse-mosquitto:2.0.22` (+ a publisher sidecar) | `10.100.0.35:8883` TLS+auth · `:1883` plaintext+auth | `MQTT_FIXTURE_USERNAME=otopcua MQTT_FIXTURE_PASSWORD=<pw> docker compose up -d` (stack `/opt/otopcua-mqtt`) |
| MQTT — Sparkplug B | `otopcua-sparkplug-sim` (project-owned C# edge-node simulator) | same broker; group **`OtOpcUaSim`**, edge nodes **`EdgeA`** / **`EdgeB`**, `EdgeA` device **`Filler1`**, ~2 s cadence | `docker compose --profile sparkplug up -d` (same stack). Answers rebirth NCMDs; restart it to force a fresh birth |
| MTConnect | `mtconnect/agent:2.7.0.12` + `python:3.13-alpine` adapter | `http://10.100.0.35:5000` | `docker compose up -d --wait` |
**Endpoint overrides** point a suite at a real PLC instead of the sim:
`MODBUS_SIM_ENDPOINT` · `AB_SERVER_ENDPOINT` · `S7_SIM_ENDPOINT` · `OPCUA_SIM_ENDPOINT`.
`MODBUS_SIM_ENDPOINT` · `AB_SERVER_ENDPOINT` · `S7_SIM_ENDPOINT` · `OPCUA_SIM_ENDPOINT` ·
`MQTT_FIXTURE_ENDPOINT` / `MQTT_FIXTURE_PLAIN_ENDPOINT` · `MTCONNECT_AGENT_ENDPOINT`.
> **MQTT needs credentials + material, unlike the other fixtures.** The broker has **no anonymous
> fallback on either listener** (deliberate — an anonymous broker would let a bad auth config pass),
> so it will not even start until `./gen-fixture-material.sh` has written `secrets/` (password file +
> CA + server cert/key, gitignored) and `MQTT_FIXTURE_USERNAME`/`MQTT_FIXTURE_PASSWORD` are exported
> at `docker compose up`. The live suite additionally wants `MQTT_FIXTURE_CA_CERT` pointing at a local
> copy of the CA: `scp dohertj2@10.100.0.35:/opt/otopcua-mqtt/secrets/ca.crt /tmp/mqtt-fixture-ca.crt`.
> It is also the only fixture whose services carry the `project: lmxopcua` label (see CLAUDE.md).
> **MTConnect is a TWO-service stack.** The `mtconnect/agent` image ships no simulator, so a lone
> Agent answers `/probe` and reports every observation `UNAVAILABLE` forever. The `adapter` service
> (a stdlib SHDR feeder, `Docker/adapter.py`) is the data source; the Agent dials out to it. Its
> suite also probes over **HTTP**, not TCP — an Agent's port accepts connections before its device
> model is parsed. Full detail + the seeded device model:
> [`tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.IntegrationTests/Docker/README.md`](../tests/Drivers/ZB.MOM.WW.OtOpcUa.Driver.MTConnect.IntegrationTests/Docker/README.md).
> **Fixture-cycling gotcha:** profile-gated services can share a port; a plain `docker compose down`
> may leave a stale container bound. Force-remove before switching profiles:
@@ -29,9 +29,31 @@ public enum BrowseNodeKind
/// <c>AlarmExtension</c> primitive). The picker pre-fills a default native-alarm <c>alarm</c> object
/// into the TagConfig when an alarm attribute is selected. Defaults to false so non-alarm-aware
/// drivers (e.g. the OPC UA client browser) aren't forced to flow a flag they don't produce.</param>
/// <param name="AddressFields">
/// The <b>structured</b> address this leaf binds by, as <c>TagConfig</c> key → value in the
/// driver's own key vocabulary, for a driver whose address is a <i>tuple</i> rather than a single
/// reference string. Null (the default) for every driver whose address IS the
/// <see cref="BrowseNode.NodeId"/> — nothing changes for them.
/// <para>
/// It exists because a tuple cannot be recovered from an id string in general. MQTT/Sparkplug
/// is the case in point: a browse node id is
/// <c>{group}/{node}[/{device}]::{metric}</c>, and a metric name may itself contain <c>/</c>
/// (<c>Node Control/Rebirth</c>) — so the session keeps the decomposition it already has and
/// <b>states</b> it here, rather than the AdminUI's browse-commit mapper re-deriving it by
/// splitting the id and hoping. Same discipline as the v3 address space carrying
/// <c>AddressSpaceRealm</c> explicitly instead of parsing it out of a NodeId.
/// </para>
/// <para>
/// The producer states the binding <em>shape</em> too, by which keys it emits — so a consumer
/// never has to infer a driver's sub-mode from the tree's geometry. Consumers must read the
/// keys they know and ignore the rest; this is a description, not a config blob to splat into
/// a persisted TagConfig.
/// </para>
/// </param>
public sealed record AttributeInfo(
string Name,
string DriverDataType,
bool IsArray,
string SecurityClass,
bool IsAlarm = false);
bool IsAlarm = false,
IReadOnlyDictionary<string, string>? AddressFields = null);
@@ -35,3 +35,66 @@ public interface IBrowseSession : IAsyncDisposable
/// <returns>The attributes of the node.</returns>
Task<IReadOnlyList<AttributeInfo>> AttributesAsync(string nodeId, CancellationToken cancellationToken);
}
/// <summary>
/// An <see cref="IBrowseSession"/> whose protocol offers an explicit, operator-triggered
/// <b>re-announce</b> action: a request that the remote peer republish its self-description so the
/// observation window can fill without waiting for the next natural announcement. Implemented
/// today only by the MQTT/Sparkplug browser, whose tree is built from observed NBIRTH/DBIRTH
/// certificates.
/// </summary>
/// <remarks>
/// <para>
/// <b>This is the only interface in the browse contract that causes an outbound message.</b>
/// Every other browse member is strictly read-only, and for the MQTT browser that read-only
/// property is load-bearing: a picker opened by an operator runs against a live production
/// broker. It is a separate interface, rather than an optional member on
/// <see cref="IBrowseSession"/>, precisely so "does this session write?" stays a type
/// question a caller cannot forget to ask.
/// </para>
/// <para>
/// <b>Callers must authorize before invoking it.</b> The AdminUI routes it through
/// <c>BrowserSessionService.RequestRebirthAsync</c>, which enforces the same
/// <c>DriverOperator</c> policy that gates the picker's Browse affordance — an implementation
/// cannot check that itself, since it holds no user identity.
/// </para>
/// </remarks>
public interface IRebirthCapableBrowseSession : IBrowseSession
{
/// <summary>
/// Whether this <i>particular</i> session can actually re-announce — the runtime half of the
/// type question, and the one a UI must ask before offering the affordance.
/// </summary>
/// <remarks>
/// One session class may serve several protocol shapes: the MQTT browser opens the same session
/// type for Plain and for Sparkplug B, and a Plain MQTT window publishes <b>nothing, ever</b> —
/// there is no plain-MQTT re-announce to offer. Implementing the interface therefore means "this
/// session type may re-announce"; this property means "this instance will". A UI gating on the
/// type alone would draw a button that can only ever throw.
/// <para>
/// This is <b>not</b> an authorization signal — see the interface remarks. False here hides
/// the affordance; the caller still authorizes before invoking
/// <see cref="RequestRebirthAsync"/>, and the implementation still refuses a call it cannot
/// serve.
/// </para>
/// </remarks>
bool RebirthAvailable { get; }
/// <summary>
/// Asks the addressed remote peer(s) to re-announce themselves.
/// </summary>
/// <param name="scope">
/// The protocol-specific target. For Sparkplug: a browse <c>NodeId</c> from this session's own
/// tree (group, edge node, device or metric — resolved up to the owning edge node), or a bare
/// <c>{group}/{edgeNode}</c> pair for a node that has not been observed yet.
/// </param>
/// <param name="cancellationToken">Cancellation token for the operation.</param>
/// <returns>The number of request messages published.</returns>
/// <exception cref="ArgumentException"><paramref name="scope"/> is empty or unusable.</exception>
/// <exception cref="InvalidOperationException">
/// The scope resolves to no target, or to more targets than the implementation will fan out to
/// in one action. Nothing is published in either case.
/// </exception>
/// <exception cref="NotSupportedException">This session's mode has no re-announce action.</exception>
Task<int> RequestRebirthAsync(string scope, CancellationToken cancellationToken);
}
@@ -41,9 +41,77 @@ public static class DraftValidator
ValidateUnsEffectiveLeafUniqueness(draft, errors);
ValidateEquipReferenceResolution(draft, errors);
ValidateCalculationTags(draft, errors);
ValidateSqlConnectionStringNotPersisted(draft, errors);
return errors;
}
/// <summary>
/// The <c>Sql</c> driver's "a pasted literal connection string is never persisted" guarantee, enforced
/// at the deploy gate (Gitea #498).
/// </summary>
/// <remarks>
/// <para>A Sql driver names its credentials indirectly — <c>connectionStringRef</c> resolves from the
/// environment / secret store at Initialize — so a deployed artifact carries no database password.
/// Until now that guarantee rested entirely on the <b>read</b> path: <c>SqlDriverConfigDto</c> has no
/// <c>connectionString</c> property and <c>UnmappedMemberHandling.Skip</c> drops the key on
/// deserialization. Nothing stopped the key being <b>written</b>. Config blobs are schemaless JSON
/// columns, and the fallback authoring pattern for a driver without a typed form is a raw-JSON
/// textarea, so an operator pasting <c>{"connectionString":"Server=…;Password=…"}</c> would put a live
/// credential in the ConfigDb — persisted, replicated to every node's artifact cache, and readable by
/// anyone with config access — while the runtime silently ignored it and the driver failed to connect.
/// Discarding a secret on read is not the same as refusing to store it.</para>
/// <para><b>Two of the three surfaces are checked here.</b> The third — a node's
/// <see cref="ClusterNode.DriverConfigOverridesJson"/> — is not reachable from
/// <see cref="DraftSnapshot"/> and is gated at its save instead (#499); see
/// <see cref="SqlCredentialGuard"/> for why a deploy gate is the wrong instrument for a value that
/// never enters the artifact.</para>
/// <para>Checked on <b>both</b> config surfaces a Sql driver reads: the instance's
/// <see cref="DriverInstance.DriverConfig"/> and the <see cref="Device.DeviceConfig"/> of every device
/// beneath it, because the two are merged before the DTO sees them — a credential pasted into the
/// device blob lands in exactly the same place.</para>
/// <para>The key is matched <b>case-insensitively</b>: <c>System.Text.Json</c> binds
/// <c>ConnectionString</c> to a <c>connectionString</c> property by default, so a case variant is the
/// same key, not a different one. Only the top level is scanned — the DTO is flat, so a nested
/// occurrence cannot bind and is not the credential-shaped mistake this rule exists to catch.</para>
/// <para><b>The message never echoes the value.</b> Validation errors reach the AdminUI, the deploy
/// log and the audit trail; repeating the offending string there would leak the very credential the
/// rule is refusing to store.</para>
/// </remarks>
private static void ValidateSqlConnectionStringNotPersisted(DraftSnapshot draft, List<ValidationError> errors)
{
const string ForbiddenKey = SqlCredentialGuard.ForbiddenKey;
var sqlInstanceIds = draft.DriverInstances
.Where(d => string.Equals(d.DriverType, Core.Abstractions.DriverTypeNames.Sql, StringComparison.Ordinal))
.Select(d => d.DriverInstanceId)
.ToHashSet(StringComparer.Ordinal);
if (sqlInstanceIds.Count == 0) return;
foreach (var d in draft.DriverInstances)
{
if (!sqlInstanceIds.Contains(d.DriverInstanceId)) continue;
if (!SqlCredentialGuard.CarriesLiteralConnectionString(d.DriverConfig)) continue;
errors.Add(new("SqlConnectionStringPersisted",
$"Sql driver instance '{d.DriverInstanceId}' has a '{ForbiddenKey}' key in its DriverConfig. " +
"A Sql driver must name its credentials indirectly via 'connectionStringRef', which resolves " +
"from the environment / secret store at Initialize; a literal connection string here would be " +
"stored in the config database and replicated to every node, and the runtime ignores it anyway. " +
"Remove the key and set 'connectionStringRef'.",
d.DriverInstanceId));
}
foreach (var dev in draft.Devices)
{
if (!sqlInstanceIds.Contains(dev.DriverInstanceId)) continue;
if (!SqlCredentialGuard.CarriesLiteralConnectionString(dev.DeviceConfig)) continue;
errors.Add(new("SqlConnectionStringPersisted",
$"Device '{dev.DeviceId}' on Sql driver instance '{dev.DriverInstanceId}' has a " +
$"'{ForbiddenKey}' key in its DeviceConfig. DeviceConfig is merged onto DriverConfig before " +
"the driver reads it, so this is the same leak: use 'connectionStringRef' on the driver instead.",
dev.DeviceId));
}
}
/// <summary>WP7 Calculation-driver deploy gates. For every tag bound to a <c>Calculation</c> driver:
/// <list type="number">
/// <item><b>scriptId existence</b> — the tag's <c>TagConfig.scriptId</c> must be present and resolve to
@@ -0,0 +1,139 @@
using System.Text.Json;
namespace ZB.MOM.WW.OtOpcUa.Configuration.Validation;
/// <summary>
/// The one place that decides whether a persisted config blob carries a literal Sql
/// <c>connectionString</c> (Gitea #498 / #499).
/// </summary>
/// <remarks>
/// <para>
/// A <c>Sql</c> driver names its credentials indirectly — <c>connectionStringRef</c> resolves from
/// the environment / secret store at Initialize — so nothing persisted should ever contain a
/// database password. The typed <c>SqlDriverConfigDto</c> already drops a <c>connectionString</c>
/// key on <b>read</b>, but discarding a secret on read is not the same as refusing to store it:
/// config blobs are schemaless JSON columns authored through raw-JSON textareas, so the key can
/// be written even though the runtime ignores it.
/// </para>
/// <para>
/// There are <b>three</b> surfaces a Sql driver's config is persisted on, and they need different
/// enforcement points:
/// <list type="bullet">
/// <item><c>DriverInstance.DriverConfig</c> and <c>Device.DeviceConfig</c> — both in the
/// deployed artifact, both visible to <c>DraftSnapshot</c>, both gated by
/// <c>DraftValidator</c>.</item>
/// <item><c>ClusterNode.DriverConfigOverridesJson</c> — the per-node override map. It is
/// <b>not</b> in <c>DraftSnapshot</c>, and deliberately does not go there: adding a
/// <c>ClusterNode</c> collection would widen the snapshot (and every builder of one) for a
/// single rule about a value that is not part of the artifact at all. More to the point, a
/// deploy gate is the wrong instrument here — the artifact never carries node overrides, so
/// blocking a deploy would not stop the credential being stored. It is already in the
/// database by then. This surface is therefore gated at the <b>save</b>, which is the only
/// point where "refuse to store it" is literally true.</item>
/// </list>
/// </para>
/// <para>
/// <b>Deliberately narrow.</b> This checks the <c>Sql</c> driver's <c>connectionString</c> key
/// only — not a general "credential-shaped key" sweep over every driver type. A broader rule
/// would start refusing configs that are legitimate today for drivers that never made this
/// guarantee, turning defence-in-depth into a regression. Widen it per driver, as each driver
/// gains an indirect-credential contract of its own.
/// </para>
/// </remarks>
public static class SqlCredentialGuard
{
/// <summary>The config key a Sql driver must never persist. Matched case-insensitively.</summary>
public const string ForbiddenKey = "connectionString";
/// <summary>
/// Whether <paramref name="configJson"/> is a JSON object carrying <see cref="ForbiddenKey"/> at
/// its top level.
/// </summary>
/// <remarks>
/// Matched <b>case-insensitively</b>: <c>System.Text.Json</c> binds <c>ConnectionString</c> to a
/// <c>connectionString</c> property by default, so a case variant is the same key, not a
/// different one. Only the top level is scanned — the DTO is flat, so a nested occurrence cannot
/// bind and is not the credential-shaped mistake this exists to catch.
/// </remarks>
/// <param name="configJson">The config blob to inspect.</param>
/// <returns><see langword="true"/> when the key is present at the top level.</returns>
public static bool CarriesLiteralConnectionString(string? configJson)
{
if (string.IsNullOrWhiteSpace(configJson)) return false;
try
{
using var doc = JsonDocument.Parse(configJson);
return doc.RootElement.ValueKind == JsonValueKind.Object
&& HasTopLevelKey(doc.RootElement, ForbiddenKey);
}
catch (JsonException)
{
// A malformed blob simply has no keys. Shaping the config JSON is another rule's job, and
// throwing here would turn a formatting mistake into a credential-guard failure.
return false;
}
}
/// <summary>
/// The <c>DriverInstanceId</c>s in a <c>ClusterNode.DriverConfigOverridesJson</c> map whose
/// override object carries a literal connection string, restricted to instances that are actually
/// Sql drivers.
/// </summary>
/// <remarks>
/// The override map is shaped <c>{ "&lt;DriverInstanceId&gt;": { …driver config keys… } }</c> and is
/// merged onto the cluster-level <c>DriverConfig</c>, so a credential pasted into one lands in
/// exactly the place <see cref="CarriesLiteralConnectionString"/> already refuses on the driver
/// itself.
/// </remarks>
/// <param name="overridesJson">The node's override map. Null / blank / malformed yields no
/// violations.</param>
/// <param name="sqlDriverInstanceIds">The ids of driver instances whose <c>DriverType</c> is
/// <c>Sql</c>. Keys outside this set are ignored — a non-Sql driver never made the indirect-credential
/// guarantee, so flagging it would be a regression, not defence in depth.</param>
/// <returns>The offending driver-instance ids, in document order; empty when there are none.</returns>
public static IReadOnlyList<string> FindNodeOverrideViolations(
string? overridesJson, IReadOnlyCollection<string> sqlDriverInstanceIds)
{
if (string.IsNullOrWhiteSpace(overridesJson) || sqlDriverInstanceIds.Count == 0)
{
return [];
}
var sqlIds = sqlDriverInstanceIds as IReadOnlySet<string>
?? sqlDriverInstanceIds.ToHashSet(StringComparer.Ordinal);
try
{
using var doc = JsonDocument.Parse(overridesJson);
if (doc.RootElement.ValueKind != JsonValueKind.Object) return [];
List<string>? offenders = null;
foreach (var entry in doc.RootElement.EnumerateObject())
{
if (!sqlIds.Contains(entry.Name)) continue;
if (entry.Value.ValueKind != JsonValueKind.Object) continue;
if (!HasTopLevelKey(entry.Value, ForbiddenKey)) continue;
(offenders ??= []).Add(entry.Name);
}
return (IReadOnlyList<string>?)offenders ?? [];
}
catch (JsonException)
{
return [];
}
}
/// <summary>Whether <paramref name="obj"/> carries <paramref name="key"/> at its top level,
/// case-insensitively.</summary>
private static bool HasTopLevelKey(JsonElement obj, string key)
{
foreach (var property in obj.EnumerateObject())
{
if (string.Equals(property.Name, key, StringComparison.OrdinalIgnoreCase)) return true;
}
return false;
}
}
@@ -53,6 +53,14 @@ public static class DriverTypeNames
/// <summary>Calculation pseudo-driver — tags computed by C# scripts over other tags' live values.</summary>
public const string Calculation = "Calculation";
/// <summary>Read-only SQL Server table/view poller — publishes columns as OPC UA variables.</summary>
public const string Sql = "Sql";
/// <summary>MQTT / Sparkplug B broker-subscription driver.</summary>
public const string Mqtt = "Mqtt";
/// <summary>MTConnect Agent driver — read-only HTTP/XML over an Agent's probe/current/sample surface.</summary>
public const string MTConnect = "MTConnect";
/// <summary>
/// Every driver-type string declared above, for callers that need to enumerate
/// the full set (e.g. validation of an authored <c>DriverType</c>).
@@ -68,5 +76,8 @@ public static class DriverTypeNames
OpcUaClient,
Galaxy,
Calculation,
Sql,
Mqtt,
MTConnect,
];
}
@@ -334,6 +334,48 @@ public sealed class ScriptedAlarmEngine : IDisposable
public IReadOnlyCollection<AlarmConditionState> GetAllStates()
=> _alarms.Values.Select(a => a.Condition).ToArray();
/// <summary>
/// #487 — the currently-held condition of every loaded alarm, packaged with the definition's
/// severity, the rendered message template and the last-observed input quality, i.e. everything
/// a caller needs to <b>project</b> the alarm onto an OPC UA condition node without waiting for
/// the next transition.
/// </summary>
/// <remarks>
/// <para>
/// <b>This is a read, not an emission.</b> It deliberately returns a distinct type rather
/// than a <see cref="ScriptedAlarmEvent"/>, and it never touches <see cref="OnEvent"/>. A
/// still-active alarm produces <see cref="EmissionKind.None"/> at load (there is no
/// transition to report), so the transition stream cannot carry the current state — that is
/// exactly why the caller needs this. Routing these through the emission path instead would
/// append a duplicate historian / <c>alerts</c> row on every deploy, so the shape here is
/// intentionally one the emission machinery cannot consume.
/// </para>
/// <para>
/// <b>Synchronization.</b> Reads <c>_alarms</c> and <c>_valueCache</c> — both
/// <see cref="ConcurrentDictionary{TKey, TValue}"/> — without taking <c>_evalGate</c>, the
/// same posture as <see cref="GetState"/> / <see cref="GetAllStates"/>. Each projection is
/// individually coherent; the set as a whole is not an atomic snapshot across a concurrent
/// re-evaluation. That is sufficient for the node-projection use: an evaluation racing this
/// read publishes its own transition through the normal path.
/// </para>
/// </remarks>
/// <returns>One projection per loaded alarm; empty before the first <see cref="LoadAsync"/>.</returns>
public IReadOnlyList<ScriptedAlarmProjection> GetProjections()
{
var projections = new List<ScriptedAlarmProjection>(_alarms.Count);
foreach (var (alarmId, state) in _alarms)
{
projections.Add(new ScriptedAlarmProjection(
AlarmId: alarmId,
Severity: state.Definition.Severity,
Message: MessageTemplate.Resolve(state.Definition.MessageTemplate, TryLookup),
Condition: state.Condition,
WorstInputStatusCode: LastWorstStatus(alarmId)));
}
return projections;
}
/// <summary>Acknowledges the specified alarm on behalf of the given user.</summary>
/// <param name="alarmId">The alarm identifier.</param>
/// <param name="user">The user performing the acknowledgment.</param>
@@ -974,6 +1016,26 @@ public sealed record ScriptedAlarmEvent(
// the top-2 severity bits. Default 0u == Good keeps every existing constructor call unchanged.
uint WorstInputStatusCode = 0u);
/// <summary>
/// #487 — a loaded alarm's current condition, ready to be projected onto its OPC UA node.
/// Produced by <see cref="ScriptedAlarmEngine.GetProjections"/> on demand; it is <b>not</b> an
/// emission and never travels the <see cref="ScriptedAlarmEngine.OnEvent"/> path, so it can never
/// become an <c>alerts</c> row or a historian record. Deliberately a separate type from
/// <see cref="ScriptedAlarmEvent"/> — it carries no <see cref="EmissionKind"/>, so there is nothing
/// for a future refactor to mistake for a transition.
/// </summary>
/// <param name="AlarmId">The alarm identifier (also its OPC UA condition NodeId).</param>
/// <param name="Severity">The definition's severity bucket.</param>
/// <param name="Message">The message template resolved against the current value cache.</param>
/// <param name="Condition">The Part 9 condition state the engine currently holds.</param>
/// <param name="WorstInputStatusCode">The last-observed worst OPC UA StatusCode across the alarm's inputs.</param>
public sealed record ScriptedAlarmProjection(
string AlarmId,
AlarmSeverity Severity,
string Message,
AlarmConditionState Condition,
uint WorstInputStatusCode);
/// <summary>
/// Upstream source abstraction — intentionally identical shape to the virtual-tag
/// engine's so Stream G can compose them behind one driver bridge.
@@ -35,7 +35,7 @@ namespace ZB.MOM.WW.OtOpcUa.Driver.AbCip;
public static class AbCipStatusMapper
{
public const uint Good = 0u;
public const uint GoodMoreData = 0x00A70000u;
public const uint GoodMoreData = 0x00A60000u;
public const uint BadInternalError = 0x80020000u;
public const uint BadNodeIdUnknown = 0x80340000u;
public const uint BadNotWritable = 0x803B0000u;
@@ -44,7 +44,7 @@ public static class AbCipStatusMapper
public const uint BadDeviceFailure = 0x808B0000u;
public const uint BadCommunicationError = 0x80050000u;
public const uint BadTimeout = 0x800A0000u;
public const uint BadTypeMismatch = 0x80730000u;
public const uint BadTypeMismatch = 0x80740000u;
/// <summary>Map a CIP general-status byte to an OPC UA StatusCode.</summary>
/// <param name="status">The CIP general-status byte value.</param>
@@ -10,7 +10,7 @@ namespace ZB.MOM.WW.OtOpcUa.Driver.AbLegacy;
public static class AbLegacyStatusMapper
{
public const uint Good = 0u;
public const uint GoodMoreData = 0x00A70000u;
public const uint GoodMoreData = 0x00A60000u;
public const uint BadInternalError = 0x80020000u;
public const uint BadNodeIdUnknown = 0x80340000u;
public const uint BadNotWritable = 0x803B0000u;
@@ -19,7 +19,7 @@ public static class AbLegacyStatusMapper
public const uint BadDeviceFailure = 0x808B0000u;
public const uint BadCommunicationError = 0x80050000u;
public const uint BadTimeout = 0x800A0000u;
public const uint BadTypeMismatch = 0x80730000u;
public const uint BadTypeMismatch = 0x80740000u;
/// <summary>
/// Map a libplctag return/status code to an OPC UA StatusCode. The integer passed here
@@ -17,7 +17,7 @@ public static class FocasStatusMapper
public const uint BadDeviceFailure = 0x808B0000u;
public const uint BadCommunicationError = 0x80050000u;
public const uint BadTimeout = 0x800A0000u;
public const uint BadTypeMismatch = 0x80730000u;
public const uint BadTypeMismatch = 0x80740000u;
/// <summary>
/// Map common FWLIB <c>EW_*</c> return codes. The values below match Fanuc's published
@@ -856,7 +856,7 @@ public sealed class GalaxyDriver
{
rejectedTcs.TrySetResult(new DataValueSnapshot(
Value: null,
StatusCode: 0x80000000u, // Bad
StatusCode: StatusCodeMap.Bad,
SourceTimestampUtc: null,
ServerTimestampUtc: DateTime.UtcNow));
}
@@ -875,7 +875,7 @@ public sealed class GalaxyDriver
{
tcs.TrySetResult(new DataValueSnapshot(
Value: null,
StatusCode: 0x800B0000u, // BadTimeout
StatusCode: StatusCodeMap.BadTimeout,
SourceTimestampUtc: null,
ServerTimestampUtc: DateTime.UtcNow));
}
@@ -26,14 +26,23 @@ internal static class StatusCodeMap
{
// OPC UA Part 4 standard StatusCodes — top-byte categories are 0x00 (Good),
// 0x40 (Uncertain), 0x80 (Bad). Specific codes layer onto the category byte.
//
// The substatus nibbles are NOT derivable from the OPC DA quality byte this mapper consumes —
// they are an unrelated OPC UA enumeration and have to be looked up. Five constants here were
// originally written as though the DA byte could be shifted into the substatus position, which
// produced values naming entirely different UA codes (Gitea #497): UncertainLastUsableValue held
// UncertainDataSubNormal's value, UncertainSubNormal held UncertainNoCommunicationLastUsableValue's,
// and GoodLocalOverride / UncertainSensorNotAccurate / UncertainEngineeringUnitsExceeded held values
// that are not OPC UA status codes at all. StatusCodeParityTests now checks every constant below
// against the pinned SDK's Opc.Ua.StatusCodes.
public const uint Good = 0x00000000u;
public const uint GoodLocalOverride = 0x00D80000u;
public const uint GoodLocalOverride = 0x00960000u;
public const uint Uncertain = 0x40000000u;
public const uint UncertainLastUsableValue = 0x40A40000u;
public const uint UncertainSensorNotAccurate = 0x408D0000u;
public const uint UncertainEngineeringUnitsExceeded = 0x408E0000u;
public const uint UncertainSubNormal = 0x408F0000u;
public const uint UncertainLastUsableValue = 0x40900000u;
public const uint UncertainSensorNotAccurate = 0x40930000u;
public const uint UncertainEngineeringUnitsExceeded = 0x40940000u;
public const uint UncertainSubNormal = 0x40950000u;
public const uint Bad = 0x80000000u;
public const uint BadConfigurationError = 0x80890000u;
public const uint BadNotConnected = 0x808A0000u;
@@ -44,6 +53,14 @@ internal static class StatusCodeMap
public const uint BadWaitingForInitialData = 0x80320000u;
public const uint BadInternalError = 0x80020000u;
/// <summary>
/// Fills a still-pending read when the caller's token fires before the gateway answers. Named here
/// rather than written inline at the call site so <c>StatusCodeParityTests</c> can reflect over it —
/// an inline literal is invisible to that guard, which is exactly how this constant spent its life
/// as <c>0x800B0000</c> (<c>BadServiceUnsupported</c>) under a <c>// BadTimeout</c> comment.
/// </summary>
public const uint BadTimeout = 0x800A0000u;
/// <summary>
/// Map a raw OPC DA quality byte (the low byte of an OPC DA <c>OpcQuality</c> ushort,
/// which is what Wonderware Historian + MXAccess surface as <c>OPCITEMSTATE.qLong</c>'s
@@ -18,29 +18,29 @@ internal static class GatewayQualityMapper
public static uint Map(byte q) => q switch
{
// Good family (192+)
192 => 0x00000000u, // Good
216 => 0x00D80000u, // Good_LocalOverride
192 => HistorianStatusCodes.Good,
216 => HistorianStatusCodes.GoodLocalOverride,
// Uncertain family (64-191)
64 => 0x40000000u, // Uncertain
68 => 0x40900000u, // Uncertain_LastUsableValue
80 => 0x40930000u, // Uncertain_SensorNotAccurate
84 => 0x40940000u, // Uncertain_EngineeringUnitsExceeded
88 => 0x40950000u, // Uncertain_SubNormal
64 => HistorianStatusCodes.Uncertain,
68 => HistorianStatusCodes.UncertainLastUsableValue,
80 => HistorianStatusCodes.UncertainSensorNotAccurate,
84 => HistorianStatusCodes.UncertainEngineeringUnitsExceeded,
88 => HistorianStatusCodes.UncertainSubNormal,
// Bad family (0-63)
0 => 0x80000000u, // Bad
4 => 0x80890000u, // Bad_ConfigurationError
8 => 0x808A0000u, // Bad_NotConnected
12 => 0x808B0000u, // Bad_DeviceFailure
16 => 0x808C0000u, // Bad_SensorFailure
20 => 0x80050000u, // Bad_CommunicationError
24 => 0x808D0000u, // Bad_OutOfService
32 => 0x80320000u, // Bad_WaitingForInitialData
0 => HistorianStatusCodes.Bad,
4 => HistorianStatusCodes.BadConfigurationError,
8 => HistorianStatusCodes.BadNotConnected,
12 => HistorianStatusCodes.BadDeviceFailure,
16 => HistorianStatusCodes.BadSensorFailure,
20 => HistorianStatusCodes.BadCommunicationError,
24 => HistorianStatusCodes.BadOutOfService,
32 => HistorianStatusCodes.BadWaitingForInitialData,
// Unknown — fall back to category bucket so callers still get something usable.
_ when q >= 192 => 0x00000000u,
_ when q >= 64 => 0x40000000u,
_ => 0x80000000u,
_ when q >= 192 => HistorianStatusCodes.Good,
_ when q >= 64 => HistorianStatusCodes.Uncertain,
_ => HistorianStatusCodes.Bad,
};
}
@@ -0,0 +1,75 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Historian.Gateway.Mapping;
/// <summary>
/// The OPC UA status codes this driver publishes, as named constants.
/// </summary>
/// <remarks>
/// <para>The driver layer is deliberately free of an OPC UA SDK reference, so status codes are spelled
/// as bare <c>uint</c>s. The cost of that is a value nobody checks against the name it is written
/// under — the defect class Gitea #497 found in six places across four drivers, one of them the
/// <c>BadNoData</c> in <see cref="SampleMapper"/> that was really <c>BadServerHalted</c>.</para>
/// <para><b>Named, not inline, on purpose.</b> <c>StatusCodeParityTests</c> guards these by reflecting
/// over <c>const uint</c> fields and comparing each against <c>Opc.Ua.StatusCodes</c> in the pinned SDK.
/// A literal written at a call site is invisible to that guard, which is exactly how
/// <see cref="GatewayQualityMapper"/> kept an incorrect <c>Good_LocalOverride</c> through every test it
/// had. Add a constant here rather than a literal in a <c>switch</c> arm.</para>
/// </remarks>
internal static class HistorianStatusCodes
{
// ---- Good family ----
/// <summary>The value is good; no qualification.</summary>
public const uint Good = 0x00000000u;
/// <summary>The value has been overridden locally (OPC DA quality 216).</summary>
public const uint GoodLocalOverride = 0x00960000u;
// ---- Uncertain family ----
/// <summary>The value is uncertain; no specific reason.</summary>
public const uint Uncertain = 0x40000000u;
/// <summary>Communication has failed; the last known value is returned (OPC DA quality 68).</summary>
public const uint UncertainLastUsableValue = 0x40900000u;
/// <summary>The sensor is known not to be accurate (OPC DA quality 80).</summary>
public const uint UncertainSensorNotAccurate = 0x40930000u;
/// <summary>The value is outside the engineering-unit range for the sensor (OPC DA quality 84).</summary>
public const uint UncertainEngineeringUnitsExceeded = 0x40940000u;
/// <summary>The value is derived from fewer sources than required (OPC DA quality 88).</summary>
public const uint UncertainSubNormal = 0x40950000u;
// ---- Bad family ----
/// <summary>The value is bad; no specific reason.</summary>
public const uint Bad = 0x80000000u;
/// <summary>A configuration problem prevents the value being produced (OPC DA quality 4).</summary>
public const uint BadConfigurationError = 0x80890000u;
/// <summary>The source is not connected (OPC DA quality 8).</summary>
public const uint BadNotConnected = 0x808A0000u;
/// <summary>The device reported a failure (OPC DA quality 12).</summary>
public const uint BadDeviceFailure = 0x808B0000u;
/// <summary>The sensor reported a failure (OPC DA quality 16).</summary>
public const uint BadSensorFailure = 0x808C0000u;
/// <summary>Communication with the source failed (OPC DA quality 20).</summary>
public const uint BadCommunicationError = 0x80050000u;
/// <summary>The source is out of service (OPC DA quality 24).</summary>
public const uint BadOutOfService = 0x808D0000u;
/// <summary>No initial value has arrived from the source yet (OPC DA quality 32).</summary>
public const uint BadWaitingForInitialData = 0x80320000u;
/// <summary>
/// The historian returned no data for the requested tag/interval. Distinct from a transport
/// failure: the query succeeded and the answer was empty.
/// </summary>
public const uint BadNoData = 0x809B0000u;
}
@@ -10,8 +10,8 @@ namespace ZB.MOM.WW.OtOpcUa.Driver.Historian.Gateway.Mapping;
/// </summary>
internal static class SampleMapper
{
private const uint StatusGood = 0x00000000u;
private const uint StatusBadNoData = 0x800E0000u;
private const uint StatusGood = HistorianStatusCodes.Good;
private const uint StatusBadNoData = HistorianStatusCodes.BadNoData;
/// <summary>OPC DA "Good" family floor — a quality byte at/above this carries usable data.</summary>
private const byte GoodQualityFloor = 192;
@@ -0,0 +1,266 @@
using System.Collections.Frozen;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// The type shape inferred for a single MTConnect <c>DataItem</c>: the scalar
/// <see cref="DriverDataType" /> plus the array shape that
/// <see cref="DriverAttributeInfo.IsArray" /> / <see cref="DriverAttributeInfo.ArrayDim" />
/// expect. Returned by value (a readonly record struct) so the discovery loop and the
/// AdminUI editor can call the inference on a hot path without allocating.
/// </summary>
/// <param name="DataType">The scalar element type of the observation.</param>
/// <param name="IsArray">
/// <c>true</c> only for a <c>SAMPLE</c> declared <c>representation="TIME_SERIES"</c>, whose
/// observation carries a vector of samples rather than a single value.
/// </param>
/// <param name="ArrayDim">
/// The probe-declared array length when <see cref="IsArray" /> is <c>true</c> and the device
/// model declared a positive <c>sampleCount</c>; <c>null</c> for a variable-length series
/// (many agents omit <c>sampleCount</c>) and always <c>null</c> for a scalar.
/// </param>
public readonly record struct MTConnectInferredType(DriverDataType DataType, bool IsArray, uint? ArrayDim);
/// <summary>
/// Translates an MTConnect <c>DataItem</c>'s device-model metadata (category / type / units /
/// representation) into the server's <see cref="DriverDataType" />. Pure and deterministic —
/// no I/O, no logging, no mutable state.
/// </summary>
/// <remarks>
/// <para>
/// This lives in <c>.Contracts</c> because three call sites must agree byte-for-byte, or a
/// tag's OPC UA type stops matching its authored config: <c>ITagDiscovery.DiscoverAsync</c>
/// (stamps <see cref="DriverAttributeInfo.DriverDataType" /> on every browsed leaf), the
/// universal browser's commit path (writes the browsed leaf into <c>TagConfig</c>), and the
/// AdminUI typed tag editor (shows the inferred type and offers an override).
/// </para>
/// <para>
/// <b>MTConnect is weakly typed on the wire</b> — every observation is text, and the standard
/// defines hundreds of <c>type</c> values with more added each revision. This is therefore a
/// <i>rule</i>, not an enumeration, and the result is stored per tag and author-overridable.
/// The rule is asymmetric on purpose: a wrong <i>numeric</i> guess produces a value that
/// fails to parse and surfaces as Bad quality at runtime, whereas <see cref="DriverDataType.String" />
/// always round-trips and the author can retype it. So every default leans to
/// <c>String</c> except where the standard itself guarantees a number.
/// </para>
/// <para>
/// <b>Defaulting rule, per category</b> (design §3.3):
/// </para>
/// <list type="bullet">
/// <item>
/// <description>
/// <c>SAMPLE</c> ⇒ <see cref="DriverDataType.Float64" />, <i>including for an
/// unrecognised or missing <c>type</c></i>. The standard defines the whole SAMPLE
/// category as a continuously-varying measured quantity, so the category alone is
/// the guarantee — a missing <c>units</c> attribute is a gap in the device model,
/// not evidence that the observation is text. <c>Float64</c> rather than an integer
/// type because a double parse accepts both <c>3</c> and <c>3.5</c>.
/// </description>
/// </item>
/// <item>
/// <description>
/// <c>EVENT</c> ⇒ <see cref="DriverDataType.String" /> by default, with a small
/// exception list (<see cref="NumericEventTypes" />) of standard types that are
/// defined as integers. The overwhelming majority of EVENT types are controlled
/// vocabulary (<c>Execution</c>, <c>ControllerMode</c>, <c>Availability</c>) or free
/// text (<c>Program</c>, <c>Block</c>, <c>Message</c>), and the exception list is
/// deliberately short — each entry is a coercion risk, while every omission is
/// merely a string the author can retype.
/// </description>
/// </item>
/// <item>
/// <description>
/// <c>CONDITION</c> ⇒ <see cref="DriverDataType.String" /> always. The observation
/// is a state word (<c>Normal</c> / <c>Warning</c> / <c>Fault</c> /
/// <c>Unavailable</c>), never a number, regardless of the item's <c>type</c>.
/// </description>
/// </item>
/// <item>
/// <description>
/// An <b>unknown, blank, or missing category</b> ⇒ <see cref="DriverDataType.String" />.
/// Probe XML in the wild is inconsistent; with no category we have no evidence the
/// observation is numeric, so we pick the type that cannot fail to parse.
/// </description>
/// </item>
/// </list>
/// <para>
/// Representation overrides the category where the two disagree about shape:
/// <c>DATA_SET</c> / <c>TABLE</c> observations are key-value maps, so they demote to
/// <see cref="DriverDataType.String" /> even on a <c>SAMPLE</c>; <c>TIME_SERIES</c> is
/// defined only for <c>SAMPLE</c> and is ignored (not treated as an array) elsewhere;
/// <c>DISCRETE</c> and <c>VALUE</c> are scalar and change nothing.
/// </para>
/// <para>
/// <b>Matching is case-insensitive AND separator-insensitive</b> on every input —
/// <c>_</c>, <c>-</c> and whitespace are ignored wherever they appear, so
/// <c>PART_COUNT</c> ≡ <c>PartCount</c> ≡ <c>part-count</c> ≡ <c>partcount</c>. This is not
/// mere leniency: <b>MTConnect spells the same concept two different ways in two different
/// documents</b>. The Devices document (<c>/probe</c>) writes <c>DataItem@type</c> in
/// UPPER_SNAKE (<c>PART_COUNT</c>, <c>POSITION</c>, <c>PATH_FEEDRATE</c>), while the Streams
/// documents (<c>/current</c>, <c>/sample</c>) name the observation element in PascalCase
/// (<c>PartCount</c>, <c>Position</c>). Discovery reads the <i>probe</i> spelling, and plain
/// case-insensitivity does not bridge the underscore — so matching only the PascalCase
/// spelling types every real agent's <c>PART_COUNT</c> as <see cref="DriverDataType.String" />
/// while every unit test using the sketch's PascalCase spelling stays green.
/// </para>
/// </remarks>
public static class MTConnectDataTypeInference
{
private const string CategorySample = "SAMPLE";
private const string CategoryEvent = "EVENT";
private const string RepresentationTimeSeries = "TIME_SERIES";
private const string RepresentationDataSet = "DATA_SET";
private const string RepresentationTable = "TABLE";
/// <summary>
/// The <c>EVENT</c> types the standard defines as integers — the exception list to the
/// "EVENT is text" default. Kept deliberately short: adding a type here is a claim that
/// every agent reports it as a parseable integer, and being wrong costs Bad quality.
/// <c>Line</c> is the deprecated spelling that <c>LineNumber</c> superseded in MTConnect 1.4;
/// both are mapped so a current-version agent is not silently mistyped.
/// <para>
/// The entries are written in the Streams (PascalCase) spelling, but the set's comparer
/// is <see cref="SeparatorInsensitiveComparer" />, so the probe document's
/// <c>PART_COUNT</c> / <c>LINE_NUMBER</c> hit the same entries. Adding a spelling variant
/// here would be redundant, not additional coverage.
/// </para>
/// </summary>
private static readonly FrozenSet<string> NumericEventTypes =
new[] { "PartCount", "Line", "LineNumber" }.ToFrozenSet(SeparatorInsensitiveComparer.Instance);
/// <summary>
/// Infers the OPC UA-facing type shape of an MTConnect <c>DataItem</c> from its device-model
/// metadata. See the type-level remarks for the full rule and its rationale.
/// </summary>
/// <param name="category">
/// The <c>DataItem</c> category — <c>SAMPLE</c>, <c>EVENT</c>, or <c>CONDITION</c>.
/// Case-insensitive; an unknown, blank, or <c>null</c> value yields
/// <see cref="DriverDataType.String" />.
/// </param>
/// <param name="type">
/// The <c>DataItem</c> type, in <b>either</b> document's spelling — the probe's
/// <c>PART_COUNT</c> / <c>POSITION</c> or the streams' <c>PartCount</c> / <c>Position</c>.
/// Matched case- and separator-insensitively (see the type-level remarks); an unrecognised
/// or <c>null</c> value falls back to the category default.
/// </param>
/// <param name="units">
/// The declared engineering units, if any. Accepted for signature stability and because every
/// call site already holds it, but <b>not load-bearing in v1</b>: <c>SAMPLE</c> is numeric by
/// definition of the category, so units add no discriminating power there, and treating a
/// units-bearing <c>EVENT</c> as numeric would be exactly the risky coercion this rule avoids.
/// </param>
/// <param name="representation">
/// The <c>DataItem</c> representation — <c>VALUE</c> (default), <c>TIME_SERIES</c>,
/// <c>DATA_SET</c>, <c>TABLE</c>, or <c>DISCRETE</c>. Case-insensitive.
/// </param>
/// <param name="sampleCount">
/// The probe-declared <c>sampleCount</c> for a <c>TIME_SERIES</c> item. Flows into
/// <see cref="MTConnectInferredType.ArrayDim" /> when positive; <c>null</c> or a non-positive
/// value yields a variable-length array (<c>ArrayDim = null</c>). Ignored for scalars.
/// </param>
/// <returns>The inferred scalar type plus array shape.</returns>
public static MTConnectInferredType Infer(
string? category,
string? type = null,
string? units = null,
string? representation = null,
int? sampleCount = null)
{
_ = units; // See the parameter doc: intentionally not load-bearing in v1.
var trimmedCategory = category.AsSpan().Trim();
var trimmedRepresentation = representation.AsSpan().Trim();
// A key-value observation is never a number, whatever the category claims. Checked first so
// a DATA_SET SAMPLE cannot slip through as Float64 and go Bad on every parse.
if (Matches(trimmedRepresentation, RepresentationDataSet) || Matches(trimmedRepresentation, RepresentationTable))
return new MTConnectInferredType(DriverDataType.String, IsArray: false, ArrayDim: null);
if (Matches(trimmedCategory, CategorySample))
{
// TIME_SERIES is defined for SAMPLE only; the vector is a vector of the same scalar type.
if (Matches(trimmedRepresentation, RepresentationTimeSeries))
return new MTConnectInferredType(
DriverDataType.Float64,
IsArray: true,
ArrayDim: sampleCount is > 0 ? (uint)sampleCount.Value : null);
return new MTConnectInferredType(DriverDataType.Float64, IsArray: false, ArrayDim: null);
}
if (Matches(trimmedCategory, CategoryEvent))
{
// No Trim() here: the set's comparer already ignores whitespace and separators
// wherever they appear, so the raw attribute value is looked up allocation-free.
var dataType = type is { Length: > 0 } && NumericEventTypes.Contains(type)
? DriverDataType.Int64
: DriverDataType.String;
return new MTConnectInferredType(dataType, IsArray: false, ArrayDim: null);
}
// CONDITION (a state word) and every unknown / blank / missing category.
return new MTConnectInferredType(DriverDataType.String, IsArray: false, ArrayDim: null);
}
private static bool Matches(ReadOnlySpan<char> value, string expected)
=> value.Equals(expected, StringComparison.OrdinalIgnoreCase);
/// <summary>
/// Ordinal-ignore-case equality that additionally skips <c>_</c>, <c>-</c> and whitespace
/// wherever they occur, so the probe document's <c>PART_COUNT</c> and the streams
/// document's <c>PartCount</c> are one key. Backs <see cref="NumericEventTypes" />.
/// </summary>
/// <remarks>
/// Comparing through a comparer rather than normalising the input keeps the lookup
/// allocation-free on the discovery hot path: no <c>string.Replace</c>, no
/// <c>Trim()</c>, no temporary buffer — the raw attribute value from the parser is passed
/// straight to <see cref="FrozenSet{T}.Contains" />.
/// </remarks>
private sealed class SeparatorInsensitiveComparer : IEqualityComparer<string>
{
/// <summary>The single shared instance; the comparer is stateless.</summary>
internal static readonly SeparatorInsensitiveComparer Instance = new();
private SeparatorInsensitiveComparer() { }
/// <inheritdoc />
public bool Equals(string? x, string? y)
{
if (ReferenceEquals(x, y)) return true;
if (x is null || y is null) return false;
int i = 0, j = 0;
while (true)
{
while (i < x.Length && IsSeparator(x[i])) i++;
while (j < y.Length && IsSeparator(y[j])) j++;
if (i == x.Length || j == y.Length) return i == x.Length && j == y.Length;
if (char.ToUpperInvariant(x[i]) != char.ToUpperInvariant(y[j])) return false;
i++;
j++;
}
}
/// <inheritdoc />
public int GetHashCode(string obj)
{
// Must hash exactly the characters Equals compares, or two equal spellings land in
// different buckets and the set silently misses.
var hash = new HashCode();
foreach (var c in obj)
{
if (IsSeparator(c)) continue;
hash.Add(char.ToUpperInvariant(c));
}
return hash.ToHashCode();
}
private static bool IsSeparator(char c) => c is '_' or '-' || char.IsWhiteSpace(c);
}
}
@@ -0,0 +1,134 @@
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// MTConnect Agent driver configuration. Bound from the driver's <c>DriverConfig</c> JSON at
/// <c>DriverHost.RegisterAsync</c>. The driver polls a remote MTConnect Agent's REST endpoints
/// (<c>/probe</c>, <c>/current</c>, <c>/sample</c>) rooted at <see cref="AgentUri"/> and maps
/// each returned DataItem into an OPC UA variable per <see cref="MTConnectTagDefinition"/>.
/// </summary>
public sealed class MTConnectDriverOptions
{
/// <summary>
/// Gets the MTConnect Agent's base URI (e.g. <c>http://agent:5000</c>). Required — the
/// driver appends the standard Agent request paths (<c>/probe</c>, <c>/current</c>,
/// <c>/sample</c>) to this base.
/// </summary>
public required string AgentUri { get; init; }
/// <summary>
/// Optional device-name scope. When set, requests are narrowed to
/// <c>{AgentUri}/{DeviceName}/...</c> so a multi-device Agent only serves one device's
/// DataItems through this driver instance. Default <c>null</c> = agent-wide (all devices).
/// </summary>
public string? DeviceName { get; init; }
/// <summary>
/// Per-call HTTP deadline, in milliseconds, applied to every Agent request
/// (<c>/probe</c>, <c>/current</c>, <c>/sample</c>). Default <c>5000</c> — long enough for
/// a large device probe response over a LAN, short enough that a hung Agent surfaces as a
/// failed poll rather than wedging the driver.
/// </summary>
public int RequestTimeoutMs { get; init; } = 5000;
/// <summary>
/// The <c>interval</c> query parameter (milliseconds) passed to the Agent's <c>/sample</c>
/// streaming request — the minimum time the Agent waits between successive chunks of the
/// response. Default <c>1000</c>; must be non-zero so the sample pump cannot spin the
/// connection into a busy-loop.
/// </summary>
public int SampleIntervalMs { get; init; } = 1000;
/// <summary>
/// The <c>count</c> query parameter passed to the Agent's <c>/sample</c> request — the
/// maximum number of DataItem observations returned per chunk. Default <c>1000</c>, the
/// MTConnect Agent's own conventional default.
/// </summary>
public int SampleCount { get; init; } = 1000;
/// <summary>
/// The <c>heartbeat</c> query parameter (milliseconds) passed to the Agent's <c>/sample</c>
/// streaming request — the Agent sends an empty chunk after this much idle time so the
/// client can detect a silently-stalled connection. Default <c>10000</c>, matching the
/// MTConnect Agent's own conventional default; must be non-zero or a quiet connection can
/// never be distinguished from a dead one.
/// </summary>
public int HeartbeatMs { get; init; } = 10000;
/// <summary>
/// The authored tags this driver instance serves — one <see cref="MTConnectTagDefinition"/>
/// per tag, keyed by the DataItem <c>id</c> it binds
/// (<see cref="MTConnectTagDefinition.FullName"/>). The factory config DTO carries these and
/// builds them into the options; the driver indexes them to know each tag's target
/// <see cref="MTConnectTagDefinition.DriverDataType"/> when coercing an Agent observation
/// (whose wire form is always text) into a published value.
/// <para>
/// Defaults to <b>empty, never <c>null</c></b>: a driver instance may legitimately be
/// authored with no tags yet, and a null collection here would surface as a
/// NullReferenceException at deploy time from operator-authored config.
/// </para>
/// </summary>
public IReadOnlyList<MTConnectTagDefinition> Tags { get; init; } = [];
/// <summary>
/// <b>The v3 data-plane binding: the authored raw tags the deploy artifact delivers.</b> Each
/// <see cref="RawTagEntry"/> pairs the tag's <b>RawPath</b> — the driver's wire reference for
/// read / subscribe / publish, and the key <c>DriverHostActor</c> routes published values by —
/// with the driver-specific <c>TagConfig</c> blob naming the Agent DataItem <c>id</c> it binds.
/// Injected into every driver's merged config by <c>DriverDeviceConfigMerger</c>, so this is the
/// collection a deployed instance actually serves from.
/// <para>
/// <b>Relationship to <see cref="Tags"/>.</b> The driver builds ONE observation index and
/// ONE <c>RawPath → dataItemId</c> table from both collections:
/// <list type="bullet">
/// <item>A tag in <c>RawTags</c> is reachable by its RawPath (production). Its coercion
/// type comes from the blob's <c>driverDataType</c>; failing that, from a
/// <see cref="Tags"/> entry naming the same DataItem id; failing that, from the Agent's
/// own <c>/probe</c> declaration (<see cref="MTConnectDataTypeInference"/>); failing
/// that, <see cref="DriverDataType.String"/> — the coercion that cannot fail.</item>
/// <item>A tag in <see cref="Tags"/> only is reachable by its DataItem <c>id</c>
/// (the driver CLI's authoring surface, and every pre-v3 test) and still contributes
/// its coercion type to the index.</item>
/// </list>
/// Both default to empty, so neither collection is required and a driver instance
/// authored with no tags yet starts cleanly.
/// </para>
/// </summary>
public IReadOnlyList<RawTagEntry> RawTags { get; init; } = [];
/// <summary>
/// Background connectivity-probe settings. When <see cref="MTConnectProbeOptions.Enabled"/>
/// is true the driver periodically issues a cheap <c>/probe</c> request and raises
/// <c>OnHostStatusChanged</c> on Running ↔ Stopped transitions.
/// </summary>
public MTConnectProbeOptions Probe { get; init; } = new();
/// <summary>
/// Reconnect backoff settings used after a failed Agent request or a dropped
/// <c>/sample</c> streaming connection.
/// </summary>
public MTConnectReconnectOptions Reconnect { get; init; } = new();
}
/// <summary>Background connectivity-probe knobs, mirroring <c>ModbusProbeOptions</c>.</summary>
public sealed class MTConnectProbeOptions
{
/// <summary>Gets a value indicating whether probing is enabled.</summary>
public bool Enabled { get; init; } = true;
/// <summary>Gets the interval between probe requests.</summary>
public TimeSpan Interval { get; init; } = TimeSpan.FromSeconds(5);
/// <summary>Gets the probe request timeout.</summary>
public TimeSpan Timeout { get; init; } = TimeSpan.FromSeconds(2);
}
/// <summary>Geometric-backoff settings for the post-failure reconnect loop, mirroring <c>ModbusReconnectOptions</c>.</summary>
public sealed class MTConnectReconnectOptions
{
/// <summary>Delay before the first reconnect attempt, in milliseconds. Default <c>0</c> = immediate.</summary>
public int MinBackoffMs { get; init; } = 0;
/// <summary>Upper bound on the geometric backoff sequence, in milliseconds.</summary>
public int MaxBackoffMs { get; init; } = 30000;
/// <summary>Multiplier applied each retry. Default <c>2.0</c> doubles each step.</summary>
public double BackoffMultiplier { get; init; } = 2.0;
}
@@ -0,0 +1,49 @@
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// One MTConnect-backed OPC UA variable. Each authored tag binds a single Agent DataItem by
/// its <c>id</c> attribute (<see cref="FullName"/>); the driver looks up the observation's
/// current value from the Agent's <c>/current</c> snapshot / <c>/sample</c> stream by that id.
/// </summary>
/// <param name="FullName">
/// The MTConnect DataItem's <c>id</c> attribute (the Agent's <c>/probe</c> response). This is
/// the driver's reference for read/subscribe — MTConnect is a read-only source protocol, so
/// there is no corresponding write path.
/// </param>
/// <param name="DriverDataType">Logical data type the DataItem's value is coerced to.</param>
/// <param name="IsArray">
/// When <c>true</c>, the tag is exposed as an OPC UA array (e.g. an MTConnect
/// <c>PathPosition</c> / vector-valued sample). Default <c>false</c> = scalar.
/// </param>
/// <param name="ArrayDim">
/// Element count when <see cref="IsArray"/> is <c>true</c>; ignored for scalar tags. Default
/// <c>0</c>.
/// </param>
/// <param name="MtCategory">
/// The DataItem's MTConnect <c>category</c> attribute — <c>SAMPLE</c>, <c>EVENT</c>, or
/// <c>CONDITION</c>. Optional metadata carried through for browse / display; not consulted by
/// the value-mapping path. Default <c>null</c>.
/// </param>
/// <param name="MtType">
/// The DataItem's MTConnect <c>type</c> attribute (e.g. <c>POSITION</c>, <c>EXECUTION</c>).
/// Optional metadata. Default <c>null</c>.
/// </param>
/// <param name="MtSubType">
/// The DataItem's MTConnect <c>subType</c> attribute (e.g. <c>ACTUAL</c>, <c>COMMANDED</c>).
/// Optional metadata. Default <c>null</c>.
/// </param>
/// <param name="Units">
/// The DataItem's MTConnect <c>units</c> attribute (e.g. <c>MILLIMETER</c>,
/// <c>REVOLUTION/MINUTE</c>). Optional metadata. Default <c>null</c>.
/// </param>
public sealed record MTConnectTagDefinition(
string FullName,
DriverDataType DriverDataType,
bool IsArray = false,
int ArrayDim = 0,
string? MtCategory = null,
string? MtType = null,
string? MtSubType = null,
string? Units = null);
@@ -0,0 +1,17 @@
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup>
<TargetFramework>net10.0</TargetFramework>
<Nullable>enable</Nullable>
<ImplicitUsings>enable</ImplicitUsings>
<TreatWarningsAsErrors>true</TreatWarningsAsErrors>
</PropertyGroup>
<!-- Pure records + the MTConnect data-type inference table, shared by the driver, browse-commit,
and the AdminUI typed tag editor. No backend NuGet — only the zero-dependency Core.Abstractions
leaf (DriverDataType etc.), mirroring Driver.Modbus.Contracts / Driver.AbCip.Contracts. -->
<ItemGroup>
<ProjectReference Include="..\..\Core\ZB.MOM.WW.OtOpcUa.Core.Abstractions\ZB.MOM.WW.OtOpcUa.Core.Abstractions.csproj"/>
</ItemGroup>
</Project>
@@ -0,0 +1,135 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// The seam between the MTConnect driver and an MTConnect Agent's HTTP surface
/// (<c>/probe</c>, <c>/current</c>, <c>/sample</c>). This interface is transport-neutral by
/// design — every method returns/yields the plain DTOs in <c>MTConnectDtos.cs</c>, never a
/// TrakHound type, an <c>XElement</c>, or any other wire/library artifact. That is what lets
/// every driver behaviour above this seam (Tasks 613) be unit-tested against a fake serving
/// canned probe/current/sample XML, with no sockets involved.
/// </summary>
/// <remarks>
/// <para>
/// The production implementation, <see cref="MTConnectAgentClient"/>, carries no backend NuGet
/// at all: it is an <see cref="HttpClient"/> over the Agent's three request paths, feeding
/// hand-rolled <c>System.Xml.Linq</c> parsers (<c>MTConnectProbeParser</c> /
/// <c>MTConnectStreamsParser</c>) plus a <c>multipart/x-mixed-replace</c> frame reader for the
/// <c>/sample</c> long poll. (Task 0 planned to build on the TrakHound MTConnect.NET libraries;
/// Tasks 6 and 7 proved they can neither parse a document nor frame the stream socket-free, and
/// dropped the references — see the Task 0 CORRECTION block in the plan.) A fake for tests only
/// needs to implement these three instance members; <see cref="IsSequenceGap"/> is a static
/// protocol rule that lives here, rather than on the implementation, so a consumer written
/// against this seam never has to reference the concrete client.
/// </para>
/// <para>
/// <b>Why the seam is <see cref="IDisposable"/> rather than
/// <see cref="IAsyncDisposable"/>.</b> The driver obtains its client through a
/// <c>Func&lt;MTConnectDriverOptions, IMTConnectAgentClient&gt;</c> and must release it on
/// <c>ShutdownAsync</c> and on every endpoint-changing <c>ReinitializeAsync</c>. Declaring
/// disposal here — instead of the driver doing <c>as IDisposable</c> + a null-conditional
/// call — is deliberate: that idiom silently no-ops against any future non-disposable
/// implementation, and the thing it would leak is an <see cref="HttpClient"/> and its socket
/// handler, once per re-deploy, for the life of the process.
/// </para>
/// <para>
/// Synchronous disposal is the honest shape because <i>every</i> resource behind this seam is
/// synchronously disposable — two <see cref="HttpClient"/>s and a
/// <see cref="CancellationTokenSource"/>. An <see cref="IAsyncDisposable"/> here would be a
/// synchronous method wearing a <c>ValueTask</c>, and it would oblige every canned-XML fake
/// to grow one too. The <b>async</b> unwind that a streaming client does need is already
/// covered elsewhere and is not this method's job: the <see cref="SampleAsync"/> enumerator
/// is itself <see cref="IAsyncDisposable"/> and is owned by the pump that enumerates it,
/// while <see cref="IDisposable.Dispose"/> on the client converts any read still in flight
/// into an ordinary <see cref="OperationCanceledException"/>.
/// </para>
/// </remarks>
public interface IMTConnectAgentClient : IDisposable
{
/// <summary>
/// Issues an Agent <c>/probe</c> request and returns the parsed device hierarchy. Called
/// once at driver initialization; the result is cached and later walked to build the OPC UA
/// browse tree, where each <see cref="MTConnectDataItem.Id"/> becomes a tag's FullName.
/// </summary>
/// <param name="ct">Cancellation token.</param>
Task<MTConnectProbeModel> ProbeAsync(CancellationToken ct);
/// <summary>
/// Issues an Agent <c>/current</c> request and returns the latest observed value for every
/// data item. Used both to prime the observation index at startup and to re-baseline after
/// a detected sequence gap in <see cref="SampleAsync"/>.
/// </summary>
/// <param name="ct">Cancellation token.</param>
Task<MTConnectStreamsResult> CurrentAsync(CancellationToken ct);
/// <summary>
/// Opens an Agent <c>/sample</c> long-poll stream starting at sequence <paramref name="from"/>
/// and yields one <see cref="MTConnectStreamsResult"/> per multipart chunk the Agent sends.
/// </summary>
/// <remarks>
/// <para>
/// <b>The only way this enumeration ends without throwing is <paramref name="ct"/> being
/// cancelled.</b> Every other way a stream can stop — the connection dropping, the Agent
/// writing its closing boundary, or the endpoint turning out not to stream at all —
/// throws an <see cref="MTConnectStreamException"/> naming which. A consumer therefore
/// never has to treat "the enumeration finished" as an ambiguous signal, and cannot
/// silently lose its subscription to a data path that quietly stopped (#485). See
/// <see cref="MTConnectStreamEndedException"/> (transient — reconnect) and
/// <see cref="MTConnectStreamNotSupportedException"/> (configuration — reconnecting will
/// reproduce it forever).
/// </para>
/// <para>
/// <b>Advancing the cursor.</b> After each chunk, the caller's next expected sequence is
/// that chunk's <see cref="MTConnectStreamsResult.NextSequence"/> — not the
/// <paramref name="from"/> this call was opened with. Feed that running value to
/// <see cref="IsSequenceGap"/> and, on reconnect, to a fresh <c>SampleAsync</c>.
/// </para>
/// </remarks>
/// <param name="from">The sequence number to resume streaming from.</param>
/// <param name="ct">Cancellation token; cancelling is the contracted way to end the stream.</param>
IAsyncEnumerable<MTConnectStreamsResult> SampleAsync(long from, CancellationToken ct);
/// <summary>
/// Has the Agent's ring buffer overflowed past the sequence the caller expected next? True
/// when the chunk's oldest retained sequence is <i>strictly</i> newer than
/// <paramref name="expectedFrom"/>, meaning observations between the two were evicted before
/// the caller could read them and an incremental update would silently skip values.
/// </summary>
/// <remarks>
/// <para>
/// <b><paramref name="expectedFrom"/> is a running cursor, not the original
/// <c>from</c>.</b> It equals the <c>from</c> passed to <see cref="SampleAsync"/> only
/// for the <i>first</i> chunk; from the second chunk on it is the <b>previous</b>
/// chunk's <see cref="MTConnectStreamsResult.NextSequence"/>. Comparing every chunk
/// against the original <c>from</c> instead reports a gap on every chunk once the
/// Agent's buffer has rolled past it — an endless <c>/current</c> re-baseline storm
/// against a perfectly healthy stream.
/// </para>
/// <para>
/// Equality is not a gap: <c>firstSequence == expectedFrom</c> is the ordinary
/// contiguous case where the Agent simply had nothing older to send.
/// </para>
/// <para>
/// <b>This check alone does NOT cover every ring-buffer overflow.</b> It can only judge
/// a chunk the Agent was willing to send — and cppagent answers a <c>from</c> that has
/// already fallen out of its buffer with an <c>MTConnectError</c> / <c>OUT_OF_RANGE</c>
/// document served under <b>HTTP 200</b>, which this client surfaces as an
/// <see cref="InvalidDataException"/> out of <see cref="SampleAsync"/>, never as a chunk.
/// A pump that re-baselines only on <c>IsSequenceGap</c> will therefore fault instead of
/// recovering exactly when it has fallen furthest behind. Treat an <c>OUT_OF_RANGE</c>
/// parse failure as a re-baseline trigger (re-prime from <see cref="CurrentAsync"/>,
/// resume from that snapshot's <c>NextSequence</c>) just like a detected gap.
/// </para>
/// </remarks>
/// <param name="expectedFrom">The sequence the caller expected this chunk to start at.</param>
/// <param name="chunk">The chunk the Agent answered with.</param>
/// <returns>
/// <c>true</c> when observations were lost and the caller must re-baseline via
/// <see cref="CurrentAsync"/> rather than trusting an incremental update.
/// </returns>
static bool IsSequenceGap(long expectedFrom, MTConnectStreamsResult chunk)
{
ArgumentNullException.ThrowIfNull(chunk);
return chunk.FirstSequence > expectedFrom;
}
}
@@ -0,0 +1,726 @@
using System.Globalization;
using System.Runtime.CompilerServices;
using System.Text;
using Microsoft.Extensions.Logging;
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// The production <see cref="IMTConnectAgentClient"/> — a thin <see cref="HttpClient"/> wrapper
/// over an MTConnect Agent's REST surface that hands every response body to a pure parser and
/// returns only the neutral DTOs in <c>MTConnectDtos.cs</c>.
/// </summary>
/// <remarks>
/// <para>
/// <b>The constructor opens no connection.</b> It only composes the request URIs and builds
/// its <see cref="HttpClient"/>s (which themselves dial nothing until a request is issued),
/// so a throwaway instance is safe to construct — the universal discovery browser's
/// <c>CanBrowse</c> pattern depends on that. <see cref="SampleAsync"/> likewise dials
/// nothing until it is enumerated.
/// </para>
/// <para>
/// <b>Two <see cref="HttpClient"/>s, and the split is deliberate.</b>
/// <see cref="HttpClient.Timeout"/> is a per-instance, whole-response deadline, so a single
/// client cannot serve both legs: the unary <c>/probe</c> and <c>/current</c> calls must be
/// bounded by <see cref="MTConnectDriverOptions.RequestTimeoutMs"/>, while <c>/sample</c> is
/// a long poll that is <i>supposed</i> to outlive that bound indefinitely. Raising the one
/// timeout to suit the stream would un-bound the unary deadline — precisely what arch-review
/// R2-01 says must never happen — so the streaming leg gets its own client with
/// <see cref="Timeout.InfiniteTimeSpan"/> and is bounded differently instead (below). In
/// production each client owns its own handler and connection pool; in tests both are built
/// over one injected handler, which the caller owns.
/// </para>
/// <para>
/// <b>Every call is deadline-bounded, but not all by the same mechanism.</b> The unary legs
/// run under a <see cref="CancellationTokenSource"/> linked to the caller's token and
/// cancelled after <c>RequestTimeoutMs</c>. The streaming leg bounds its <i>header</i> phase
/// the same way (a peer that accepts the connection and never answers must not wedge the
/// pump), then hands the body to a <b>heartbeat watchdog</b>: MTConnect Agents emit a
/// keep-alive chunk every <see cref="MTConnectDriverOptions.HeartbeatMs"/> precisely so a
/// client can tell "quiet but alive" from "dead", so every body read is bounded by a
/// two-missed-heartbeat window and a silent Agent surfaces as a
/// <see cref="TimeoutException"/>. This is the exact shape of this repo's S7 frozen-peer
/// production defect (arch-review R2-01) — an async read that ignores the socket deadline —
/// and the watchdog is what closes it here.
/// </para>
/// <para>
/// <b>Every operator-authorable duration/count is rejected at construction if non-positive.</b>
/// A <c>0</c> must never come to mean "wait forever" (arch-review 01/S-6), and the three
/// <c>/sample</c> query knobs additionally reach the wire: an Agent answers
/// <c>heartbeat=0</c> with an <c>MTConnectError</c> under HTTP 200, which would otherwise
/// surface as a parse failure that never names the offending config key.
/// </para>
/// <para>
/// <b>Disposal is safe against an in-flight enumeration.</b> Task 9 re-initializes the
/// driver by disposing and rebuilding this client, which can happen while a pump is still
/// enumerating <see cref="SampleAsync"/>. Rather than letting the in-flight read fault with
/// an unclassified <see cref="ObjectDisposedException"/> or <see cref="IOException"/> that no
/// caller is written to expect, every request runs under a token linked to an internal
/// dispose token, so disposal surfaces as an ordinary
/// <see cref="OperationCanceledException"/>.
/// </para>
/// <para>
/// This class does not swallow failures into an empty model: an unreachable Agent, a non-2xx
/// status, an unparseable body, and <b>the end of a <c>/sample</c> stream</b> all throw. See
/// <see cref="MTConnectStreamException"/> for why the last of those is an exception rather
/// than a normal end of enumeration. <c>InitializeAsync</c> (Task 9) is what catches these
/// and moves the driver to Faulted.
/// </para>
/// </remarks>
public sealed class MTConnectAgentClient : IMTConnectAgentClient, IDisposable
{
private readonly HttpClient _http;
private readonly HttpClient _streamHttp;
private readonly CancellationTokenSource _disposeCts = new();
private readonly ILogger? _logger;
private readonly Uri _probeUri;
private readonly Uri _currentUri;
private readonly string _sampleUriBase;
private readonly string _sampleQuerySuffix;
private readonly TimeSpan _requestTimeout;
private readonly TimeSpan _streamIdleTimeout;
private bool _disposed;
/// <summary>Creates a client for the Agent described by <paramref name="options"/>.</summary>
/// <param name="options">The driver options carrying the Agent URI, device scope, and per-call deadline.</param>
/// <param name="logger">
/// Optional logger. Used only for conditions an operator needs to see but which are not
/// failures — today, the degraded <c>/sample</c> framing mode.
/// </param>
/// <exception cref="ArgumentException"><see cref="MTConnectDriverOptions.AgentUri"/> is missing or is not an absolute HTTP(S) URI.</exception>
/// <exception cref="ArgumentOutOfRangeException">
/// Any of <see cref="MTConnectDriverOptions.RequestTimeoutMs"/>,
/// <see cref="MTConnectDriverOptions.HeartbeatMs"/>,
/// <see cref="MTConnectDriverOptions.SampleIntervalMs"/> or
/// <see cref="MTConnectDriverOptions.SampleCount"/> is not positive.
/// </exception>
public MTConnectAgentClient(MTConnectDriverOptions options, ILogger? logger = null)
: this(options, handler: null, logger)
{
}
/// <summary>
/// Test seam: as <see cref="MTConnectAgentClient(MTConnectDriverOptions, ILogger)"/>, but
/// sends through the supplied handler so the request/response round trip can be exercised
/// with no socket. The caller retains ownership of <paramref name="handler"/>.
/// </summary>
internal MTConnectAgentClient(MTConnectDriverOptions options, HttpMessageHandler? handler, ILogger? logger = null)
{
ArgumentNullException.ThrowIfNull(options);
RequirePositive(
options.RequestTimeoutMs,
nameof(MTConnectDriverOptions.RequestTimeoutMs),
"a non-positive per-call deadline would let a frozen Agent wedge the driver indefinitely");
RequirePositive(
options.HeartbeatMs,
nameof(MTConnectDriverOptions.HeartbeatMs),
"it is sent to the Agent as 'heartbeat', which an Agent rejects with an MTConnectError under HTTP 200, and it is what bounds the /sample watchdog — a quiet connection would otherwise be indistinguishable from a dead one");
RequirePositive(
options.SampleIntervalMs,
nameof(MTConnectDriverOptions.SampleIntervalMs),
"it is sent to the Agent as 'interval'; zero would spin the streaming connection into a busy-loop");
RequirePositive(
options.SampleCount,
nameof(MTConnectDriverOptions.SampleCount),
"it is sent to the Agent as 'count'; a non-positive count asks the Agent for no observations");
_logger = logger;
_requestTimeout = TimeSpan.FromMilliseconds(options.RequestTimeoutMs);
// Two missed heartbeats plus one request timeout of network slack. Deriving the window from
// HeartbeatMs is what makes it a liveness check rather than a throughput assumption: a busy
// Agent and an idle one both emit at least one chunk per heartbeat.
_streamIdleTimeout = TimeSpan.FromMilliseconds((2L * options.HeartbeatMs) + options.RequestTimeoutMs);
_probeUri = BuildRequestUri(options, "probe");
_currentUri = BuildRequestUri(options, "current");
_sampleUriBase = BuildRequestUri(options, "sample").AbsoluteUri;
_sampleQuerySuffix = string.Create(
CultureInfo.InvariantCulture,
$"&interval={options.SampleIntervalMs}&count={options.SampleCount}&heartbeat={options.HeartbeatMs}");
_http = handler is null ? new HttpClient() : new HttpClient(handler, disposeHandler: false);
_http.Timeout = _requestTimeout;
_streamHttp = handler is null ? new HttpClient() : new HttpClient(handler, disposeHandler: false);
// See the class remarks: a long poll must outlive the unary deadline by design, and the
// heartbeat watchdog — not this property — is what keeps it bounded.
_streamHttp.Timeout = Timeout.InfiniteTimeSpan;
}
/// <inheritdoc/>
public async Task<MTConnectProbeModel> ProbeAsync(CancellationToken ct)
{
ObjectDisposedException.ThrowIf(_disposed, this);
// ResponseContentRead (the default) buffers the whole body *under* this awaited, token-bound
// call, so the parse below reads from memory — no network read can outlive the deadline.
using var response = await SendAsync(_probeUri, ct).ConfigureAwait(false);
await using var body = await response.Content.ReadAsStreamAsync(ct).ConfigureAwait(false);
return MTConnectProbeParser.Parse(body);
}
/// <inheritdoc/>
public async Task<MTConnectStreamsResult> CurrentAsync(CancellationToken ct)
{
ObjectDisposedException.ThrowIf(_disposed, this);
using var response = await SendAsync(_currentUri, ct).ConfigureAwait(false);
await using var body = await response.Content.ReadAsStreamAsync(ct).ConfigureAwait(false);
return MTConnectStreamsParser.Parse(body);
}
/// <inheritdoc/>
public async IAsyncEnumerable<MTConnectStreamsResult> SampleAsync(
long from,
[EnumeratorCancellation] CancellationToken ct)
{
ObjectDisposedException.ThrowIf(_disposed, this);
var uri = BuildSampleUri(from);
// Linked to the caller's token AND to disposal, so tearing this client down mid-enumeration
// surfaces as a cancellation rather than as a raw ObjectDisposedException from the socket.
using var lifetime = CancellationTokenSource.CreateLinkedTokenSource(ct, _disposeCts.Token);
var live = lifetime.Token;
using var request = new HttpRequestMessage(HttpMethod.Get, uri);
// ResponseHeadersRead, not the default ResponseContentRead: buffering a
// multipart/x-mixed-replace body that is designed never to end would hang here forever.
using var response = await SendStreamRequestAsync(request, uri, lifetime).ConfigureAwait(false);
await using var body = await response.Content.ReadAsStreamAsync(live).ConfigureAwait(false);
var boundary = ResolveMultipartBoundary(response, uri);
if (boundary is null)
{
// Not a multipart stream: the Agent answered with a single document, which means it
// ignored the streaming request entirely. Parse it first — an MTConnectError arrives in
// exactly this shape (text/xml under HTTP 200) and the Agent's own error text is far
// more useful than ours — then report the misconfiguration. The parsed snapshot is
// deliberately NOT yielded: one document followed by a healthy-looking finished stream
// is how a dead data path disguises itself as a working one (#485).
var single = await ReadWholeBodyAsync(body, uri, lifetime).ConfigureAwait(false);
_ = MTConnectStreamsParser.Parse(single);
throw new MTConnectStreamNotSupportedException(
$"MTConnect Agent answered the /sample request '{uri}' with a single '{response.Content.Headers.ContentType?.MediaType ?? "unknown"}' document instead of a multipart/x-mixed-replace stream. It parsed as a valid MTConnectStreams document but will never stream, so no subscription is possible; check that the URI addresses an MTConnect Agent directly and is not fronted by a proxy that buffers streaming responses.");
}
var reader = new MultipartStreamReader(body, boundary, _streamIdleTimeout, uri, _logger);
var delivered = 0L;
while (true)
{
var part = await reader.ReadNextPartAsync(live).ConfigureAwait(false);
if (part is null)
{
// The caller cancelling is the ONE contracted way out; anything else is reported.
live.ThrowIfCancellationRequested();
throw new MTConnectStreamEndedException(reader.EndReason, uri.AbsoluteUri, delivered);
}
// A frame carrying nothing but whitespace is transport filler, not a document. The
// Agent's real keep-alive is a well-formed observation-free MTConnectStreams document,
// and that one IS yielded — it advances the sequence, so swallowing it would strand the
// caller's `from` behind the Agent's buffer.
if (string.IsNullOrWhiteSpace(part))
{
continue;
}
delivered++;
yield return MTConnectStreamsParser.Parse(part);
}
}
/// <inheritdoc/>
public void Dispose()
{
if (_disposed)
{
return;
}
_disposed = true;
// Cancel BEFORE disposing the clients so an in-flight read observes cancellation rather
// than a disposed handler.
_disposeCts.Cancel();
_http.Dispose();
_streamHttp.Dispose();
_disposeCts.Dispose();
}
private static void RequirePositive(int value, string optionName, string because)
{
if (value <= 0)
{
throw new ArgumentOutOfRangeException(
nameof(MTConnectDriverOptions),
value,
$"{optionName} must be positive; {because}.");
}
}
private Uri BuildSampleUri(long from) =>
new(string.Create(CultureInfo.InvariantCulture, $"{_sampleUriBase}?from={from}{_sampleQuerySuffix}"),
UriKind.Absolute);
private async Task<HttpResponseMessage> SendAsync(Uri uri, CancellationToken ct)
{
using var lifetime = CancellationTokenSource.CreateLinkedTokenSource(ct, _disposeCts.Token);
using var deadline = CancellationTokenSource.CreateLinkedTokenSource(lifetime.Token);
deadline.CancelAfter(_requestTimeout);
HttpResponseMessage response;
try
{
response = await _http.GetAsync(uri, deadline.Token).ConfigureAwait(false);
}
catch (OperationCanceledException) when (!lifetime.IsCancellationRequested)
{
// Neither the caller nor disposal cancelled, so this is our own deadline (or
// HttpClient.Timeout) firing. Reported as a TimeoutException because a bare "A task was
// canceled" tells an operator nothing about which Agent stopped answering.
throw TimedOut(uri);
}
response.EnsureSuccessStatusCode();
return response;
}
/// <summary>
/// Issues the <c>/sample</c> request and returns as soon as the response headers land. Only
/// this phase is bounded by the unary deadline; the body is bounded by the watchdog.
/// </summary>
private async Task<HttpResponseMessage> SendStreamRequestAsync(
HttpRequestMessage request, Uri uri, CancellationTokenSource lifetime)
{
using var deadline = CancellationTokenSource.CreateLinkedTokenSource(lifetime.Token);
deadline.CancelAfter(_requestTimeout);
HttpResponseMessage response;
try
{
response = await _streamHttp
.SendAsync(request, HttpCompletionOption.ResponseHeadersRead, deadline.Token)
.ConfigureAwait(false);
}
catch (OperationCanceledException) when (!lifetime.IsCancellationRequested)
{
throw TimedOut(uri);
}
response.EnsureSuccessStatusCode();
return response;
}
/// <summary>Reads a non-multipart streaming body whole, under the unary deadline.</summary>
private async Task<string> ReadWholeBodyAsync(Stream body, Uri uri, CancellationTokenSource lifetime)
{
using var deadline = CancellationTokenSource.CreateLinkedTokenSource(lifetime.Token);
deadline.CancelAfter(_requestTimeout);
using var reader = new StreamReader(body, Encoding.UTF8);
try
{
return await reader.ReadToEndAsync(deadline.Token).ConfigureAwait(false);
}
catch (OperationCanceledException) when (!lifetime.IsCancellationRequested)
{
throw TimedOut(uri);
}
}
private TimeoutException TimedOut(Uri uri) =>
new($"MTConnect Agent request to '{uri}' did not complete within {_requestTimeout.TotalMilliseconds:F0} ms.");
/// <summary>
/// The <c>boundary</c> parameter of a <c>multipart/*</c> response, or <c>null</c> when the
/// Agent answered with a single (non-multipart) document.
/// </summary>
/// <exception cref="MTConnectStreamNotSupportedException">
/// The response claims to be <c>multipart/*</c> but carries no <c>boundary</c> parameter, so
/// there is nothing to frame it by. Failed fast and named, because the alternative — falling
/// through to whole-body reading of a body that never ends — surfaces much later as a bare
/// watchdog <see cref="TimeoutException"/> that points an operator at the network instead of
/// at the malformed header.
/// </exception>
private static string? ResolveMultipartBoundary(HttpResponseMessage response, Uri uri)
{
var contentType = response.Content.Headers.ContentType;
if (contentType?.MediaType is null ||
!contentType.MediaType.StartsWith("multipart/", StringComparison.OrdinalIgnoreCase))
{
return null;
}
var boundary = contentType.Parameters
.FirstOrDefault(p => string.Equals(p.Name, "boundary", StringComparison.OrdinalIgnoreCase))?.Value?
.Trim('"');
if (string.IsNullOrEmpty(boundary))
{
throw new MTConnectStreamNotSupportedException(
$"MTConnect Agent answered the /sample request '{uri}' with Content-Type '{contentType.MediaType}' but no 'boundary' parameter, so the multipart stream cannot be framed.");
}
return boundary;
}
private static Uri BuildRequestUri(MTConnectDriverOptions options, string request)
{
var agentUri = options.AgentUri?.Trim();
if (string.IsNullOrEmpty(agentUri))
{
throw new ArgumentException(
$"{nameof(MTConnectDriverOptions.AgentUri)} is required (e.g. 'http://agent:5000').", nameof(options));
}
if (!Uri.TryCreate(agentUri, UriKind.Absolute, out var parsed) ||
(parsed.Scheme != Uri.UriSchemeHttp && parsed.Scheme != Uri.UriSchemeHttps))
{
throw new ArgumentException(
$"{nameof(MTConnectDriverOptions.AgentUri)} '{agentUri}' is not an absolute http(s) URI.", nameof(options));
}
var root = parsed.GetLeftPart(UriPartial.Path).TrimEnd('/');
var deviceName = options.DeviceName?.Trim();
var deviceSegment = string.IsNullOrEmpty(deviceName) ? string.Empty : $"/{Uri.EscapeDataString(deviceName)}";
return new Uri($"{root}{deviceSegment}/{request}", UriKind.Absolute);
}
/// <summary>
/// Frames a <c>multipart/x-mixed-replace</c> body into individual part payloads, reading
/// incrementally and never buffering more than the part currently in flight.
/// </summary>
/// <remarks>
/// <para>
/// <b>Why hand-rolled.</b> ASP.NET Core's <c>MultipartReader</c> is not in this
/// dependency set, and TrakHound's <c>MTConnectHttpClientStream</c> — the one candidate
/// the plan left open — takes a URL and dials its own socket, so it can neither be fed an
/// existing response stream nor be exercised without a listener. See the Task 0
/// correction in the plan.
/// </para>
/// <para>
/// <b>Two framing paths, and the first one matters for latency.</b> When the part carries
/// a <c>Content-length</c> header — every Agent in the wild writes one — the body is read
/// by length and the chunk is emitted the instant it is complete. Without it there is no
/// way to know a part has ended except by seeing the <i>next</i> boundary, so that path
/// necessarily emits each chunk one chunk late: with the default 10 s heartbeat a value
/// can be up to a heartbeat stale on a driver built for low-latency subscribe. It is a
/// correctness fallback, not the expected path, and it logs a one-shot warning when it
/// engages so the degradation is visible rather than merely slow. (It cannot strand the
/// final part indefinitely — a clean end of stream returns the buffered tail — though a
/// watchdog timeout on a stalled Agent does discard it.)
/// </para>
/// <para>
/// <b>Every read is watchdog-bounded.</b> A silent peer faults the stream rather than
/// parking a task forever (arch-review R2-01), and the buffer is capped so a hostile or
/// malfunctioning Agent cannot grow it without limit.
/// </para>
/// </remarks>
private sealed class MultipartStreamReader(
Stream stream, string boundary, TimeSpan idleTimeout, Uri uri, ILogger? logger)
{
private const int MaxBufferBytes = 32 * 1024 * 1024;
private static readonly byte[] CrLfCrLf = "\r\n\r\n"u8.ToArray();
private static readonly byte[] LfLf = "\n\n"u8.ToArray();
private readonly byte[] _delimiter = Encoding.ASCII.GetBytes("--" + boundary);
private byte[] _buffer = new byte[16 * 1024];
private int _length;
private bool _finished;
private bool _warnedNoContentLength;
/// <summary>How the stream ended. Only meaningful once a read has returned <c>null</c>.</summary>
public MTConnectStreamEndReason EndReason { get; private set; } = MTConnectStreamEndReason.ConnectionClosed;
/// <summary>
/// Reads the next part's payload, or <c>null</c> once the closing delimiter is seen or
/// the Agent closes the connection — see <see cref="EndReason"/> for which.
/// </summary>
public async Task<string?> ReadNextPartAsync(CancellationToken ct)
{
if (_finished)
{
return null;
}
if (!await SkipThroughDelimiterAsync(ct).ConfigureAwait(false) ||
!await EnsureAsync(2, ct).ConfigureAwait(false))
{
return Finish(MTConnectStreamEndReason.ConnectionClosed);
}
// "--boundary--" closes the stream.
if (_buffer[0] == (byte)'-' && _buffer[1] == (byte)'-')
{
return Finish(MTConnectStreamEndReason.ClosingBoundary);
}
var headerEnd = await FindHeaderEndAsync(ct).ConfigureAwait(false);
if (headerEnd < 0)
{
return Finish(MTConnectStreamEndReason.ConnectionClosed);
}
var headers = Encoding.ASCII.GetString(_buffer, 0, headerEnd);
Consume(headerEnd);
byte[]? body;
if (ParseContentLength(headers) is { } declared)
{
body = await ReadByLengthAsync(declared, ct).ConfigureAwait(false);
}
else
{
WarnOnceAboutMissingContentLength();
body = await ReadToNextDelimiterAsync(ct).ConfigureAwait(false);
}
return body is null ? null : Encoding.UTF8.GetString(body);
}
private void WarnOnceAboutMissingContentLength()
{
if (_warnedNoContentLength)
{
return;
}
_warnedNoContentLength = true;
logger?.LogWarning(
"MTConnect /sample stream from '{AgentUri}' sends parts with no Content-length header. Falling back to boundary-scan framing, which cannot emit a chunk until the NEXT chunk begins, so every observation is delayed by up to one Agent heartbeat. Values will read stale; this is the Agent's framing, not a driver fault.",
uri);
}
private string? Finish(MTConnectStreamEndReason reason)
{
_finished = true;
EndReason = reason;
return null;
}
private async Task<byte[]?> ReadByLengthAsync(int declared, CancellationToken ct)
{
if (!await EnsureAsync(declared, ct).ConfigureAwait(false))
{
Finish(MTConnectStreamEndReason.ConnectionClosed);
return null;
}
var body = _buffer[..declared];
Consume(declared);
return body;
}
private async Task<byte[]?> ReadToNextDelimiterAsync(CancellationToken ct)
{
var next = await FindDelimiterAsync(ct).ConfigureAwait(false);
if (next < 0)
{
// EOF with no closing delimiter: whatever is buffered is the last part. Returned
// rather than dropped, so a clean end of stream cannot strand a chunk.
var tail = _buffer[.._length];
_length = 0;
Finish(MTConnectStreamEndReason.ConnectionClosed);
return tail;
}
var body = _buffer[..TrimTrailingLineBreak(next)];
// Leave the delimiter itself in place; the next call's skip consumes it.
Consume(next);
return body;
}
/// <summary>Strips the single line break that terminates a part body before its delimiter.</summary>
private int TrimTrailingLineBreak(int end)
{
if (end >= 2 && _buffer[end - 2] == (byte)'\r' && _buffer[end - 1] == (byte)'\n')
{
return end - 2;
}
return end >= 1 && _buffer[end - 1] == (byte)'\n' ? end - 1 : end;
}
/// <summary>Advances past the next boundary delimiter; <c>false</c> at end of stream.</summary>
private async Task<bool> SkipThroughDelimiterAsync(CancellationToken ct)
{
var index = await FindDelimiterAsync(ct).ConfigureAwait(false);
if (index < 0)
{
return false;
}
Consume(index + _delimiter.Length);
return true;
}
private async Task<int> FindDelimiterAsync(CancellationToken ct)
{
var searchedTo = 0;
while (true)
{
// Restart the scan far enough back that a delimiter split across two reads is still
// found — without the overlap, a boundary straddling a packet boundary is invisible.
var index = IndexOf(_delimiter, Math.Max(0, searchedTo - (_delimiter.Length - 1)));
if (index >= 0)
{
return index;
}
searchedTo = _length;
if (!await FillAsync(ct).ConfigureAwait(false))
{
return -1;
}
}
}
/// <summary>
/// Finds the end of the part's header block — the first blank line after the delimiter.
/// A part with no headers at all is legal, in which case the blank line sits at offset 0.
/// </summary>
private async Task<int> FindHeaderEndAsync(CancellationToken ct)
{
var searchedTo = 0;
while (true)
{
var crlf = IndexOf(CrLfCrLf, Math.Max(0, searchedTo - 3));
var lf = IndexOf(LfLf, Math.Max(0, searchedTo - 1));
if (crlf >= 0 && (lf < 0 || crlf <= lf))
{
return crlf + CrLfCrLf.Length;
}
if (lf >= 0)
{
return lf + LfLf.Length;
}
searchedTo = _length;
if (!await FillAsync(ct).ConfigureAwait(false))
{
return -1;
}
}
}
private static int? ParseContentLength(string headers)
{
foreach (var line in headers.Split('\n'))
{
var trimmed = line.Trim();
if (!trimmed.StartsWith("content-length:", StringComparison.OrdinalIgnoreCase))
{
continue;
}
var raw = trimmed["content-length:".Length..].Trim();
return int.TryParse(raw, NumberStyles.Integer, CultureInfo.InvariantCulture, out var value) &&
value >= 0
? value
: null;
}
return null;
}
private async Task<bool> EnsureAsync(int count, CancellationToken ct)
{
while (_length < count)
{
if (!await FillAsync(ct).ConfigureAwait(false))
{
return false;
}
}
return true;
}
/// <summary>
/// Reads at least one more byte, bounded by the heartbeat watchdog. Returns <c>false</c>
/// at end of stream and throws <see cref="TimeoutException"/> when the Agent goes silent.
/// </summary>
private async Task<bool> FillAsync(CancellationToken ct)
{
if (_length == _buffer.Length)
{
if (_buffer.Length >= MaxBufferBytes)
{
throw new InvalidDataException(
$"MTConnect /sample stream from '{uri}' sent a part larger than {MaxBufferBytes} bytes without a boundary; refusing to buffer further.");
}
Array.Resize(ref _buffer, Math.Min(_buffer.Length * 2, MaxBufferBytes));
}
using var watchdog = CancellationTokenSource.CreateLinkedTokenSource(ct);
watchdog.CancelAfter(idleTimeout);
int read;
try
{
read = await stream.ReadAsync(_buffer.AsMemory(_length), watchdog.Token).ConfigureAwait(false);
}
catch (OperationCanceledException) when (!ct.IsCancellationRequested)
{
throw new TimeoutException(
$"MTConnect /sample stream from '{uri}' sent no data (not even a heartbeat) within {idleTimeout.TotalMilliseconds:F0} ms; treating the Agent as unreachable rather than waiting indefinitely.");
}
if (read <= 0)
{
return false;
}
_length += read;
return true;
}
private int IndexOf(byte[] pattern, int start)
{
if (start >= _length)
{
return -1;
}
var index = _buffer.AsSpan(start, _length - start).IndexOf(pattern.AsSpan());
return index < 0 ? -1 : index + start;
}
private void Consume(int count)
{
Buffer.BlockCopy(_buffer, count, _buffer, 0, _length - count);
_length -= count;
}
}
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,116 @@
using Microsoft.Extensions.Logging;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
using ZB.MOM.WW.OtOpcUa.Core.Hosting;
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// Static factory registration helper for <see cref="MTConnectDriver"/>. The Host registers it
/// once at startup; the bootstrapper then materialises MTConnect <c>DriverInstance</c> rows from
/// the deployed configuration into live driver instances. Mirrors
/// <c>ModbusDriverFactoryExtensions</c> / <c>FocasDriverFactoryExtensions</c>.
/// </summary>
/// <remarks>
/// <para>
/// <b>There is exactly one parser for the config document.</b> This factory does not carry a
/// config DTO of its own — it delegates to <see cref="MTConnectDriver.ParseOptions"/>, which
/// the driver itself must own because the runtime delivers a config <i>change</i> to a live
/// instance (<c>DriverInstanceActor</c> assigns its driver once and calls
/// <c>ReinitializeAsync(newJson)</c> thereafter). A second DTO here would drift from that
/// one silently, and the drift would first show up at deploy time as an operator's edit
/// being applied differently by the two paths — or not at all.
/// </para>
/// <para>
/// <b>Construction opens nothing.</b> The Wave-0 universal discovery browser builds a
/// throwaway instance through this factory purely to read
/// <c>ITagDiscovery.SupportsOnlineDiscovery</c>, and never initializes it — so a factory
/// that dialled the Agent (or that pre-built an <see cref="IMTConnectAgentClient"/>) would
/// open a connection per AdminUI browse render against a possibly-unreachable Agent. The
/// agent-client factory is therefore left null: <see cref="MTConnectDriver"/> builds the
/// production client inside <c>InitializeAsync</c>.
/// </para>
/// <para>
/// <b>The instance is returned as its concrete type.</b> The runtime resolves every optional
/// capability (<see cref="IReadable"/>, <see cref="ISubscribable"/>,
/// <see cref="ITagDiscovery"/>, <see cref="IHostConnectivityProbe"/>,
/// <see cref="IRediscoverable"/>) by pattern-matching the object
/// <see cref="DriverFactoryRegistry"/> hands back, so nothing here may narrow it — a
/// narrowed return is how a capability ends up implemented but never dispatched. There is
/// deliberately no <c>IWritable</c>: the MTConnect Agent surface is read-only.
/// </para>
/// </remarks>
public static class MTConnectDriverFactoryExtensions
{
/// <summary>
/// The <c>DriverInstance.DriverType</c> value this factory answers to.
/// </summary>
/// <remarks>
/// <b>Aliased to <see cref="DriverTypeNames.MTConnect"/>, not a literal.</b> The registry key,
/// the instance's self-reported <see cref="MTConnectDriver.DriverType"/>, the probe's
/// <c>DriverType</c>, and the constant every dispatch map keys off are therefore the same
/// symbol — the repo-wide fix for the historical <c>TwinCat</c>/<c>Focas</c> drift, where a
/// driver's own spelling silently diverged from the constant and dispatch maps missed it.
/// Retained as a named const (rather than deleted in favour of the constant) so the existing
/// factory callers and the sibling <c>*DriverFactoryExtensions.DriverTypeName</c> convention
/// keep working; it is now an alias with no independent value.
/// </remarks>
public const string DriverTypeName = DriverTypeNames.MTConnect;
/// <summary>
/// Register the MTConnect factory with the driver registry. The optional
/// <paramref name="loggerFactory"/> is captured at registration time and used to construct
/// an <see cref="ILogger{TCategoryName}"/> per driver instance — without it the driver runs
/// with the null logger (tests and standalone callers stay unchanged).
/// </summary>
/// <param name="registry">The driver factory registry to register with.</param>
/// <param name="loggerFactory">Optional logger factory for creating loggers per driver instance.</param>
public static void Register(DriverFactoryRegistry registry, ILoggerFactory? loggerFactory = null)
{
ArgumentNullException.ThrowIfNull(registry);
registry.Register(DriverTypeName, (id, json) => CreateInstance(id, json, loggerFactory));
}
/// <summary>Public for the Server-side bootstrapper + test consumers.</summary>
/// <param name="driverInstanceId">The unique identifier for the driver instance.</param>
/// <param name="driverConfigJson">The driver's <c>DriverConfig</c> JSON.</param>
/// <returns>The constructed <see cref="MTConnectDriver"/> instance.</returns>
public static MTConnectDriver CreateInstance(string driverInstanceId, string driverConfigJson)
=> CreateInstance(driverInstanceId, driverConfigJson, loggerFactory: null);
/// <summary>Logger-aware overload — used by <see cref="Register"/>'s closure when wired through DI.</summary>
/// <param name="driverInstanceId">The unique identifier for the driver instance.</param>
/// <param name="driverConfigJson">The driver's <c>DriverConfig</c> JSON.</param>
/// <param name="loggerFactory">Optional logger factory for creating loggers per driver instance.</param>
/// <returns>The constructed <see cref="MTConnectDriver"/> instance.</returns>
/// <exception cref="ArgumentException">
/// <paramref name="driverInstanceId"/> or <paramref name="driverConfigJson"/> is null or blank.
/// </exception>
/// <exception cref="InvalidOperationException">
/// The config document is unparseable, deserialises to null, omits the required
/// <c>agentUri</c>, or carries an unauthorable enum/tag value. Throwing is correct at this
/// seam: a deployment carrying such a row must fail rather than register a driver pointed at
/// nothing, and the discovery browser's <c>CanBrowse</c> catches factory exceptions so a
/// half-authored config merely disables the address picker.
/// </exception>
public static MTConnectDriver CreateInstance(
string driverInstanceId, string driverConfigJson, ILoggerFactory? loggerFactory)
{
ArgumentException.ThrowIfNullOrWhiteSpace(driverInstanceId);
ArgumentException.ThrowIfNullOrWhiteSpace(driverConfigJson);
// The single config authority — see the class remarks for why this is not a second DTO.
// Note the asymmetry with the driver's own lifecycle path: an empty ("{}" / "[]") document
// there means "no config supplied, keep the constructor options", but this factory HAS no
// constructor options to keep, so an empty document is simply a config with no agentUri and
// must fault.
var options = MTConnectDriver.ParseOptions(driverInstanceId, driverConfigJson);
// Named arguments deliberately: the ctor's trailing parameters are all optional, so a
// future insertion could silently re-bind a positional argument to the wrong one.
return new MTConnectDriver(
options,
driverInstanceId,
agentClientFactory: null,
logger: loggerFactory?.CreateLogger<MTConnectDriver>());
}
}
@@ -0,0 +1,237 @@
using System.Diagnostics;
using System.Net.Http;
using System.Text.Json;
using System.Text.Json.Serialization;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// One-shot MTConnect Agent reachability probe for the AdminUI's Test Connect button. Issues a
/// single <c>GET {AgentUri}/probe</c> under the supplied deadline and reports success/failure
/// with latency — it does not construct or initialize an <see cref="MTConnectDriver"/> instance.
/// </summary>
/// <remarks>
/// <para>
/// <b>Deliberately does NOT delegate to <see cref="MTConnectDriver.ParseOptions"/>.</b> That
/// method builds the full driver options document, including every authored
/// <c>MTConnectTagDefinition</c> (<c>BuildTag</c> can throw on a malformed tag entry that has
/// nothing to do with reachability). A probe only needs <c>agentUri</c>, so this type parses
/// its own minimal DTO — a config that is invalid for the driver (bad tag, bad data type) can
/// still be reachability-tested, matching the AdminUI's "does the endpoint answer" intent for
/// Test Connect. It also avoids a compile-time dependency on the factory DTO shape
/// (<c>MTConnectDriverFactoryExtensions</c>), which is authored separately (Task 15).
/// </para>
/// <para>
/// <b>Accepted gap: reachable ≠ deployable.</b> Because only <c>agentUri</c> is read, a
/// config can show green on Test Connect and still fail to deploy — e.g. a non-positive
/// <c>requestTimeoutMs</c> or a malformed tag entry the factory's full parse would reject.
/// This is the direct consequence of the independent-parsing decision above, judged
/// acceptable because Test Connect's job is "is the endpoint there", not "will this exact
/// config deploy".
/// </para>
/// <para>
/// <b>Deliberately DOES reuse <see cref="MTConnectAgentClient"/> and
/// <see cref="MTConnectProbeParser"/> for the actual round trip.</b> Hand-rolling the HTTP
/// call would have to re-derive the exact behaviour the client already provides: per-call
/// <see cref="HttpClient.Timeout"/> bounding, a linked-CTS deadline, and — most importantly —
/// correct handling of an Agent's <c>MTConnectError</c> document served under HTTP 200 (so
/// <c>EnsureSuccessStatusCode</c> never fires). <see cref="MTConnectProbeParser"/> already
/// lifts the Agent's own error text into the thrown exception's message; re-parsing by hand
/// here would either duplicate that logic or under-validate (accepting any 200 response as
/// healthy, which is precisely the "empty-but-successful" defect class this codebase has hit
/// before, #485). The one-off cost of a full probe-document parse is negligible next to that
/// risk.
/// </para>
/// <para>
/// <b>Never throws</b> (the <see cref="IDriverProbe"/> contract): every failure — unparseable
/// JSON, missing/malformed <c>agentUri</c>, DNS/TCP failure, a non-2xx response, a 200 body
/// that is not a valid <c>MTConnectDevices</c> document, and timeout — becomes
/// <c>Ok = false</c> with a message. Genuine caller cancellation is likewise converted rather
/// than left to propagate, matching <c>ModbusDriverProbe</c>'s posture (its linked-CTS catch
/// does not distinguish the caller's token from its own deadline).
/// </para>
/// <para>
/// <b>Non-positive <c>timeout</c> (arch-review 01/S-6).</b> A <c>0</c> or negative
/// timeout does NOT mean "wait forever" — it is replaced with <see cref="FallbackTimeout"/>
/// (5s, matching <see cref="MTConnectDriverOptions.RequestTimeoutMs"/>'s own default) so an
/// operator-authorable field can never brick the probe UI.
/// </para>
/// <para>
/// <b>Secrets discipline.</b> <c>AgentUri</c> may carry userinfo credentials
/// (<c>http://user:pass@host/</c>). Every message this type constructs from the URI uses a
/// redacted display form (<see cref="RedactedDisplay"/>) built ONLY from
/// <see cref="Uri.Scheme"/>/<see cref="Uri.Host"/>/<see cref="Uri.Port"/>/<see cref="Uri.AbsolutePath"/>
/// — never <see cref="Uri.UserInfo"/> and never the raw config string — so a password can
/// never reach the AdminUI via the returned <see cref="DriverProbeResult.Message"/>.
/// </para>
/// </remarks>
public sealed class MTConnectDriverProbe : IDriverProbe
{
/// <summary>Replaces a non-positive caller <c>timeout</c> — see the type remarks.</summary>
private static readonly TimeSpan FallbackTimeout = TimeSpan.FromSeconds(5);
/// <summary>
/// Caps an excessive caller <c>timeout</c>. <c>CancellationTokenSource.CancelAfter(TimeSpan)</c>'s
/// legal range tops out near ~49.7 days (<see cref="uint.MaxValue"/> ms minus one); an
/// uncapped pathological value (e.g. a caller passing <c>TimeSpan.FromDays(1000)</c>) throws
/// <see cref="ArgumentOutOfRangeException"/> straight out of
/// <c>CancelAfter</c>, which would violate the "Never throws" contract just as surely as a
/// non-positive timeout meaning "wait forever" would (arch-review 01/S-6 — this is the same
/// defect class, the high end rather than the low end). 10 minutes is generous for a
/// reachability probe and leaves an enormous margin under the legal ceiling.
/// </summary>
private static readonly TimeSpan MaxTimeout = TimeSpan.FromMinutes(10);
private static readonly JsonSerializerOptions JsonOptions = new()
{
PropertyNameCaseInsensitive = true,
UnmappedMemberHandling = JsonUnmappedMemberHandling.Skip,
Converters = { new JsonStringEnumConverter() },
};
/// <summary>Test seam: routes the reused <see cref="MTConnectAgentClient"/> through a stub handler so response-shape cases need no socket.</summary>
private readonly HttpMessageHandler? _handler;
/// <summary>Creates the probe used in production (real sockets).</summary>
public MTConnectDriverProbe()
: this(handler: null)
{
}
/// <summary>Test seam constructor — see <see cref="_handler"/>.</summary>
internal MTConnectDriverProbe(HttpMessageHandler? handler)
{
_handler = handler;
}
/// <inheritdoc />
/// <remarks>
/// Sourced from the <see cref="DriverTypeNames"/> constant, not a literal — the probe is
/// looked up by <c>AdminOperationsActor</c> under this key, so a drift from the constant
/// silently disables the Test Connect button rather than failing anything at build time.
/// </remarks>
public string DriverType => DriverTypeNames.MTConnect;
/// <summary>Minimal probe-only config shape. Deliberately independent of the factory DTO — see the type remarks.</summary>
private sealed class MTConnectProbeConfigDto
{
public string? AgentUri { get; set; }
}
/// <inheritdoc />
public async Task<DriverProbeResult> ProbeAsync(string configJson, TimeSpan timeout, CancellationToken ct)
{
MTConnectProbeConfigDto? dto;
try
{
dto = JsonSerializer.Deserialize<MTConnectProbeConfigDto>(configJson, JsonOptions);
}
catch (Exception ex)
{
return new(false, $"Config JSON is invalid: {ex.Message}", null);
}
if (dto is null)
{
return new(false, "Config JSON deserialized to null.", null);
}
var agentUri = dto.AgentUri?.Trim();
if (string.IsNullOrWhiteSpace(agentUri))
{
return new(false, "Config has no agentUri to probe.", null);
}
if (!Uri.TryCreate(agentUri, UriKind.Absolute, out var parsed) ||
(parsed.Scheme != Uri.UriSchemeHttp && parsed.Scheme != Uri.UriSchemeHttps))
{
return new(false, "Config's agentUri is not an absolute http(s) URI.", null);
}
var safeUri = RedactedDisplay(parsed);
// arch-review 01/S-6: a non-positive timeout must never mean "wait forever" (low end), and a
// pathologically large one must never overrun CancelAfter's legal range (high end) — see
// MaxTimeout.
var effectiveTimeout = timeout > TimeSpan.Zero ? timeout : FallbackTimeout;
if (effectiveTimeout > MaxTimeout)
{
effectiveTimeout = MaxTimeout;
}
var requestTimeoutMs = Math.Max(1, (int)Math.Ceiling(effectiveTimeout.TotalMilliseconds));
var sw = Stopwatch.StartNew();
using var cts = CancellationTokenSource.CreateLinkedTokenSource(ct);
cts.CancelAfter(effectiveTimeout);
var options = new MTConnectDriverOptions
{
AgentUri = agentUri,
RequestTimeoutMs = requestTimeoutMs,
};
try
{
using var client = new MTConnectAgentClient(options, _handler);
_ = await client.ProbeAsync(cts.Token).ConfigureAwait(false);
sw.Stop();
return new(true, $"MTConnect /probe OK ({safeUri})", sw.Elapsed);
}
catch (OperationCanceledException)
{
// Covers both the caller's token and our own CancelAfter deadline firing — see the type
// remarks on why cancellation is never allowed to propagate.
return new(false, $"Probe timed out after {effectiveTimeout.TotalSeconds:F0}s.", null);
}
catch (TimeoutException)
{
// MTConnectAgentClient's own request-deadline signal.
return new(false, $"Probe timed out after {effectiveTimeout.TotalSeconds:F0}s.", null);
}
catch (HttpRequestException ex)
{
// Connection failure or non-2xx status. HttpRequestException messages describe the
// socket endpoint or status code, never the request URI/userinfo, so ex.Message is safe.
return new(false, $"Request to {safeUri} failed: {ex.Message}", null);
}
catch (InvalidDataException ex)
{
// MTConnectProbeParser: not well-formed XML, not an MTConnectDevices document (including
// an MTConnectError-under-200, whose own errorCode/text ex.Message already carries), or
// a document with no devices. Never includes the request URI.
return new(false, $"Agent at {safeUri} did not return a valid MTConnect device model: {ex.Message}", null);
}
catch (ArgumentException)
{
// Defensive: MTConnectAgentClient's own AgentUri validation, reached only if this
// method's own Uri.TryCreate above passed but the client's stricter check disagreed.
// Never forward the exception's message verbatim — it can echo the raw (possibly
// credentialed) config string.
return new(false, "Config's agentUri was rejected by the MTConnect client as invalid.", null);
}
catch (Exception ex)
{
// Last-resort catch for an exception type none of the above enumerate. Deliberately does
// NOT forward ex.Message: unlike the typed catches above (each individually audited for
// whether its message can carry the raw AgentUri/userinfo), an unenumerated type has no
// such audit, so ex.Message is an open door for a credential leak. safeUri + the
// exception's TYPE name is enough for an operator to act on without that risk.
return new(false, $"Probe against {safeUri} failed unexpectedly ({ex.GetType().Name}).", null);
}
}
/// <summary>
/// Builds a display form of <paramref name="uri"/> for use in returned messages, composed
/// ONLY from scheme/host/port/path — never <see cref="Uri.UserInfo"/> — so a credentialed
/// <c>AgentUri</c> can never leak a password into the AdminUI.
/// </summary>
private static string RedactedDisplay(Uri uri)
{
var portSegment = uri.IsDefaultPort ? string.Empty : $":{uri.Port}";
return $"{uri.Scheme}://{uri.Host}{portSegment}{uri.AbsolutePath}";
}
}
@@ -0,0 +1,193 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// Result of an Agent <c>/probe</c> request: the full device hierarchy the Agent exposes.
/// Produced by <see cref="IMTConnectAgentClient.ProbeAsync"/>. This is the neutral shape both
/// the TrakHound-backed implementation and a hand-rolled XML fallback must produce — nothing
/// downstream of <see cref="IMTConnectAgentClient"/> ever sees a TrakHound type or an
/// <c>XElement</c>.
/// </summary>
/// <param name="Devices">Every device the Agent's probe response declares.</param>
public sealed record MTConnectProbeModel(IReadOnlyList<MTConnectDevice> Devices);
/// <summary>
/// Common shape shared by <see cref="MTConnectDevice"/> and <see cref="MTConnectComponent"/>:
/// each may directly own data items and/or nest further components, to arbitrary depth. Backs
/// the <see cref="MTConnectComponentTreeExtensions.AllDataItems"/> recursive walk.
/// </summary>
public interface IMTConnectComponentContainer
{
/// <summary>Components nested directly under this device/component (may be empty).</summary>
IReadOnlyList<MTConnectComponent> Components { get; }
/// <summary>Data items declared directly on this device/component (may be empty).</summary>
IReadOnlyList<MTConnectDataItem> DataItems { get; }
}
/// <summary>
/// One MTConnect device from the Agent's probe response — the root of a nested
/// component/data-item tree.
/// </summary>
/// <param name="Id">The device's <c>id</c> attribute.</param>
/// <param name="Name">The device's <c>name</c> attribute, when the Agent supplies one.</param>
/// <param name="Components">Child components declared directly under this device.</param>
/// <param name="DataItems">Data items declared directly on this device (outside any component).</param>
public sealed record MTConnectDevice(
string Id,
string? Name,
IReadOnlyList<MTConnectComponent> Components,
IReadOnlyList<MTConnectDataItem> DataItems) : IMTConnectComponentContainer;
/// <summary>
/// One MTConnect component (e.g. <c>Axes</c>, <c>Controller</c>, <c>Path</c>) from the Agent's
/// probe response. Components nest to arbitrary depth — a component may itself contain further
/// child components as well as data items.
/// </summary>
/// <param name="Id">The component's <c>id</c> attribute.</param>
/// <param name="Name">The component's <c>name</c> attribute, when the Agent supplies one.</param>
/// <param name="Components">Child components nested directly under this component.</param>
/// <param name="DataItems">Data items declared directly on this component.</param>
public sealed record MTConnectComponent(
string Id,
string? Name,
IReadOnlyList<MTConnectComponent> Components,
IReadOnlyList<MTConnectDataItem> DataItems) : IMTConnectComponentContainer;
/// <summary>
/// One MTConnect DataItem declaration from the Agent's probe response. This is pure metadata —
/// it carries no observed value; values arrive separately via
/// <see cref="IMTConnectAgentClient.CurrentAsync"/> / <see cref="IMTConnectAgentClient.SampleAsync"/>
/// and are correlated back to a data item by <see cref="Id"/>.
/// </summary>
/// <param name="Id">
/// The DataItem's <c>id</c> attribute — globally unique within the Agent's probe response and
/// the correlation key used against <see cref="MTConnectObservation.DataItemId"/>.
/// </param>
/// <param name="Name">The DataItem's <c>name</c> attribute, when the Agent supplies one.</param>
/// <param name="Category">The DataItem's <c>category</c> attribute — <c>SAMPLE</c>, <c>EVENT</c>, or <c>CONDITION</c>.</param>
/// <param name="Type">The DataItem's <c>type</c> attribute (e.g. <c>POSITION</c>, <c>EXECUTION</c>).</param>
/// <param name="SubType">The DataItem's <c>subType</c> attribute (e.g. <c>ACTUAL</c>, <c>COMMANDED</c>), when present.</param>
/// <param name="Units">The DataItem's <c>units</c> attribute (e.g. <c>MILLIMETER</c>, <c>REVOLUTION/MINUTE</c>), when present.</param>
/// <param name="Representation">
/// The DataItem's <c>representation</c> attribute (e.g. <c>TIME_SERIES</c>, <c>DISCRETE</c>);
/// absent means the MTConnect default (<c>VALUE</c>).
/// </param>
/// <param name="SampleCount">
/// The DataItem's <c>sampleCount</c> attribute — element count for a <c>TIME_SERIES</c>
/// representation. Present only on data items that carry one.
/// </param>
public sealed record MTConnectDataItem(
string Id,
string? Name,
string Category,
string Type,
string? SubType,
string? Units,
string? Representation,
int? SampleCount);
/// <summary>
/// Result of an Agent <c>/current</c> request, or of a single multipart chunk from a
/// <c>/sample</c> stream. Produced by <see cref="IMTConnectAgentClient.CurrentAsync"/> and
/// <see cref="IMTConnectAgentClient.SampleAsync"/>.
/// </summary>
/// <param name="InstanceId">
/// The Agent's current instance id (its MTConnectStreams header <c>instanceId</c>). Changes
/// whenever the Agent restarts or its underlying device model changes; the driver watches this
/// for change to raise rediscovery rather than trusting a stale observation snapshot.
/// </param>
/// <param name="NextSequence">
/// The header <c>nextSequence</c> — the sequence number to resume a subsequent
/// <see cref="IMTConnectAgentClient.SampleAsync"/> call <c>from</c>.
/// </param>
/// <param name="FirstSequence">
/// The header <c>firstSequence</c> — the oldest sequence number still held in the Agent's
/// circular buffer. Load-bearing for sequence-gap detection: a caller that requested
/// <c>from</c> a sequence older than this chunk's <see cref="FirstSequence"/> has fallen out of
/// the Agent's buffer and must re-baseline via <see cref="IMTConnectAgentClient.CurrentAsync"/>.
/// </param>
/// <param name="Observations">
/// The observations carried in this snapshot/chunk, in Agent-supplied order. Never <c>null</c>;
/// empty when the chunk carries no new observations (e.g. a heartbeat).
/// </param>
public sealed record MTConnectStreamsResult(
long InstanceId,
long NextSequence,
long FirstSequence,
IReadOnlyList<MTConnectObservation> Observations);
/// <summary>
/// One observed value for a single DataItem, as reported in a <c>/current</c> snapshot or a
/// <c>/sample</c> stream chunk. Deliberately dumb: <see cref="Value"/> is carried as the raw
/// string the Agent reported (including the literal <c>UNAVAILABLE</c>) with no interpretation
/// applied here — coercing <c>UNAVAILABLE</c> to a bad-quality status and converting the string
/// to the tag's <c>DriverDataType</c> are the observation index's job, not this layer's.
/// </summary>
/// <param name="DataItemId">
/// The reporting DataItem's <c>id</c> — correlates back to <see cref="MTConnectDataItem.Id"/>
/// from the probe model.
/// </param>
/// <param name="Value">
/// The observed value exactly as the Agent reported it, or the literal string
/// <c>"UNAVAILABLE"</c> when the Agent has no current value for the data item. Never
/// pre-parsed into a nullable or an enum at this layer.
/// </param>
/// <param name="TimestampUtc">
/// The observation's Agent-reported timestamp, normalized to UTC (<see cref="DateTime.Kind"/>
/// is always <see cref="DateTimeKind.Utc"/>). Becomes the OPC UA variable's SourceTimestamp.
/// </param>
/// <param name="IsStructured">
/// <c>true</c> when the Agent carried this observation's real content in <b>child elements</b>
/// rather than as text — an MTConnect 2.0 <c>DATA_SET</c> / <c>TABLE</c> observation, whose
/// content is a list of <c>&lt;Entry key="…"&gt;</c> elements.
/// <para>
/// <b>Why the flag exists, and why only the parser can set it.</b> <see cref="Value"/> is
/// the element's concatenated descendant text, so a two-entry data set reads as the single
/// token <c>"12"</c> — keys discarded, values run together. The observation index cannot
/// detect that after the fact: neither this record nor the tag definition carries the
/// DataItem's <c>representation</c>, and no value-shape heuristic can work, because
/// concatenated entries are indistinguishable from a legitimate space-bearing EVENT such as
/// a <c>Message</c> reading <c>"Coolant level low"</c>. The one reliable discriminator —
/// that the observation element had element children — exists only while the XML is still
/// XML.
/// </para>
/// <para>
/// Deliberately a <b>neutral fact about the wire shape</b>, not a status: mapping it to a
/// quality code is the observation index's job (it codes these
/// <c>BadNotSupported</c> rather than publishing concatenated noise as Good). Defaults to
/// <c>false</c>, the shape of every ordinary scalar observation.
/// </para>
/// </param>
public sealed record MTConnectObservation(
string DataItemId,
string Value,
DateTime TimestampUtc,
bool IsStructured = false);
/// <summary>
/// Recursive helpers over the <see cref="IMTConnectComponentContainer"/> device/component tree.
/// </summary>
public static class MTConnectComponentTreeExtensions
{
/// <summary>
/// Enumerates every data item owned by this device or component, plus every data item owned
/// by any component nested beneath it, to arbitrary depth. Order is depth-first: a node's
/// own data items first, then each child component's data items in turn.
/// </summary>
/// <param name="container">The device or component to walk.</param>
public static IEnumerable<MTConnectDataItem> AllDataItems(this IMTConnectComponentContainer container)
{
foreach (var dataItem in container.DataItems)
{
yield return dataItem;
}
foreach (var child in container.Components)
{
foreach (var dataItem in child.AllDataItems())
{
yield return dataItem;
}
}
}
}
@@ -0,0 +1,415 @@
using System.Collections.Concurrent;
using System.Collections.Frozen;
using System.Globalization;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// The driver's live <c>dataItemId → <see cref="DataValueSnapshot"/></c> map: the single place
/// an Agent observation's weakly-typed text becomes an OPC UA-shaped value + quality +
/// timestamp. Written by the <c>/current</c> baseline read and by the <c>/sample</c> long-poll
/// pump; read by <c>IReadable.ReadAsync</c> and the subscription fan-out.
/// </summary>
/// <remarks>
/// <para>
/// <b>Quality mapping (design §3.3).</b> MTConnect has exactly one way to say "I have no
/// value": the literal token <c>UNAVAILABLE</c>. It arrives as a Sample/Event's text, and —
/// after <see cref="MTConnectStreamsParser"/> normalizes the <c>&lt;Unavailable/&gt;</c>
/// element name — as a CONDITION's value on that same exact spelling. It maps to
/// <c>BadNoCommunication</c> with a <c>null</c> value: the Agent is reachable, the machine's
/// data item is not. The match is <b>ordinal and case-sensitive</b> against the one spelling
/// the parser guarantees; a case-insensitive match would swallow a legitimate free-text
/// EVENT whose value happens to read "Unavailable".
/// </para>
/// <para>
/// <b>Failure posture.</b> This type sits under <c>IReadable</c> and must never throw across
/// that boundary on account of observation content. Every unusable observation degrades to a
/// Bad-coded snapshot that still carries the Agent's source timestamp, so an operator can
/// always see <i>when</i> the driver last heard about the tag even when it cannot show
/// <i>what</i>. The only exceptions raised are <see cref="ArgumentNullException"/> for a null
/// argument — a programming error, not data.
/// </para>
/// <para>
/// <b>Status codes</b> are the canonical <c>Opc.Ua.StatusCodes</c> numeric values, declared
/// once each below. Note that <see cref="BadTypeMismatch"/> is <c>0x80740000</c>: the
/// <c>0x80730000</c> that four sibling drivers declare under that name is really
/// <c>BadWriteNotSupported</c>, which would be actively misleading on a read path and
/// renders as bare "Bad" in the driver CLI's <c>SnapshotFormatter</c>.
/// </para>
/// <para>
/// <b>Thread safety.</b> Backed by a <see cref="ConcurrentDictionary{TKey,TValue}"/> and
/// guaranteed <b>per-key atomic</b> — nothing more, deliberately. <see cref="Apply"/> writes
/// each dataItemId exactly once per call with a single whole-snapshot assignment, so a
/// reader can never observe a torn snapshot (a Good code beside a stale value). It is
/// explicitly <b>not</b> a whole-batch atomic swap: OPC UA reads are per-tag and the driver
/// offers no cross-tag consistency contract, <c>/sample</c> chunks are deltas whose
/// interleaving is indistinguishable from two adjacent legal chunks, and a whole-index lock
/// would serialize the pump against every read for no semantic gain.
/// </para>
/// <para>
/// <b>Shapes this build does not represent</b> are coded <c>BadNotSupported</c> with a null
/// value rather than approximated: a <c>TIME_SERIES</c> vector (array materialization is
/// deferred to P1.5) and a <c>DATA_SET</c> / <c>TABLE</c> observation. The latter carries its
/// real content in <c>&lt;Entry key=…&gt;</c> child elements, which reach a text-valued
/// observation concatenated with the keys discarded — a two-entry set reads as the single
/// token <c>"12"</c> — and because the type inference correctly demotes those
/// representations to <see cref="DriverDataType.String"/>, coercion would otherwise
/// "succeed" into a Good-coded snapshot carrying noise. This index cannot detect that shape
/// itself (no value heuristic can separate concatenated entries from a legitimate
/// space-bearing EVENT such as a <c>Message</c> reading "Coolant level low"), so it relies on
/// <see cref="MTConnectObservation.IsStructured"/>, which
/// <see cref="MTConnectStreamsParser"/> sets while the XML is still XML.
/// </para>
/// </remarks>
internal sealed class MTConnectObservationIndex
{
/// <summary>
/// The one token that means "the Agent has no value for this data item". Matched ordinally:
/// the parser already normalizes a CONDITION's <c>&lt;Unavailable/&gt;</c> element name onto
/// this exact spelling, so every category lands on one comparison.
/// </summary>
private const string UnavailableSentinel = "UNAVAILABLE";
private const uint Good = 0x00000000u;
/// <summary>The Agent reported <c>UNAVAILABLE</c>, or supplied no text at all.</summary>
private const uint BadNoCommunication = 0x80310000u;
/// <summary>The tag is authored but the Agent has not yet reported it.</summary>
private const uint BadWaitingForInitialData = 0x80320000u;
/// <summary>The dataItemId is neither authored nor ever observed.</summary>
private const uint BadNodeIdUnknown = 0x80340000u;
/// <summary>The Agent reported a number the authored type cannot represent.</summary>
private const uint BadOutOfRange = 0x803C0000u;
/// <summary>The observation's shape is real data this build cannot represent (TIME_SERIES vectors).</summary>
private const uint BadNotSupported = 0x803D0000u;
/// <summary>The Agent reported text that is not a value of the authored type at all.</summary>
private const uint BadTypeMismatch = 0x80740000u;
/// <summary>
/// MTConnect's CONDITION state vocabulary, ranked by severity. Used only to reconcile
/// several states reported for one dataItemId <i>within a single document</i> — see
/// <see cref="Reconcile"/>. An active <c>Fault</c> outranks <c>UNAVAILABLE</c> because a
/// fault is information and an absent value is not. Case-insensitive so a vendor's casing
/// of a state element still ranks; the parser itself emits the canonical spellings.
/// </summary>
private static readonly FrozenDictionary<string, int> ConditionSeverity =
new Dictionary<string, int>(StringComparer.OrdinalIgnoreCase)
{
[UnavailableSentinel] = 0,
["Normal"] = 1,
["Warning"] = 2,
["Fault"] = 3,
}.ToFrozenDictionary(StringComparer.OrdinalIgnoreCase);
private readonly ConcurrentDictionary<string, DataValueSnapshot> _snapshots = new(StringComparer.Ordinal);
private readonly FrozenDictionary<string, MTConnectTagDefinition> _tags;
private readonly TimeProvider _timeProvider;
/// <summary>
/// Builds an index over the driver instance's authored tags.
/// </summary>
/// <param name="tags">
/// The authored tags, keyed by <see cref="MTConnectTagDefinition.FullName"/> (the DataItem
/// <c>id</c>) — normally <see cref="MTConnectDriverOptions.Tags"/>. <c>null</c> or empty is
/// legal and means "no authored types": every observation is then indexed as its raw text
/// under <see cref="DriverDataType.String"/>.
/// <para>
/// Nothing enforces <c>FullName</c> uniqueness at the options layer, so a duplicate is
/// operator-authorable. <b>The last definition wins</b> rather than throwing: refusing to
/// construct would turn an authoring slip into a driver that will not start, whereas
/// last-wins is deterministic and costs only that one tag's coercion type.
/// </para>
/// </param>
/// <param name="timeProvider">Clock behind <c>ServerTimestampUtc</c>; defaults to <see cref="TimeProvider.System"/>.</param>
public MTConnectObservationIndex(
IEnumerable<MTConnectTagDefinition>? tags = null,
TimeProvider? timeProvider = null)
{
_timeProvider = timeProvider ?? TimeProvider.System;
var map = new Dictionary<string, MTConnectTagDefinition>(StringComparer.Ordinal);
foreach (var tag in tags ?? [])
{
if (tag is null || string.IsNullOrWhiteSpace(tag.FullName)) continue;
map[tag.FullName] = tag;
}
_tags = map.ToFrozenDictionary(StringComparer.Ordinal);
}
/// <summary>The number of distinct dataItemIds currently indexed.</summary>
public int Count => _snapshots.Count;
/// <summary>
/// Folds a <c>/current</c> snapshot or one <c>/sample</c> chunk into the index. Observations
/// with no authored tag are indexed too — the Agent streams the whole device, so they are
/// normal, and dropping them would make a streaming-but-not-yet-authored data item
/// indistinguishable from one that does not exist.
/// </summary>
/// <param name="result">The parsed streams document. An observation-free result is a legal
/// keep-alive and leaves the index untouched.</param>
/// <exception cref="ArgumentNullException"><paramref name="result"/> is <c>null</c>.</exception>
public void Apply(MTConnectStreamsResult result)
{
ArgumentNullException.ThrowIfNull(result);
if (result.Observations is not { Count: > 0 }) return;
// Duplicates are reconciled on the RAW observations first, so each dataItemId is written
// exactly once below as a single atomic whole-snapshot assignment. Doing it the other way
// — writing every observation and merging snapshots in place — would need a read-modify-write
// per key and could expose a half-merged state to a concurrent reader.
var resolved = new Dictionary<string, MTConnectObservation>(StringComparer.Ordinal);
foreach (var observation in result.Observations)
{
if (observation is null || string.IsNullOrWhiteSpace(observation.DataItemId)) continue;
resolved[observation.DataItemId] =
resolved.TryGetValue(observation.DataItemId, out var incumbent)
? Reconcile(incumbent, observation)
: observation;
}
foreach (var (dataItemId, observation) in resolved)
{
_snapshots[dataItemId] = BuildSnapshot(observation);
}
}
/// <summary>
/// Returns the current snapshot for a DataItem. Never <c>null</c> and never throws: an
/// unknown or not-yet-reported id yields a Bad-coded snapshot rather than an error.
/// </summary>
/// <param name="dataItemId">The DataItem <c>id</c> the tag binds.</param>
/// <returns>
/// The indexed snapshot; else <c>BadWaitingForInitialData</c> when the id is authored but
/// the Agent has not reported it yet (the standard OPC UA "subscribed, no value yet"
/// state); else <c>BadNodeIdUnknown</c>. The two are kept distinct because collapsing them
/// would tell an operator their correctly-authored tag does not exist.
/// </returns>
public DataValueSnapshot Get(string dataItemId)
{
if (!string.IsNullOrWhiteSpace(dataItemId) &&
_snapshots.TryGetValue(dataItemId, out var snapshot))
{
return snapshot;
}
var status = !string.IsNullOrWhiteSpace(dataItemId) && _tags.ContainsKey(dataItemId)
? BadWaitingForInitialData
: BadNodeIdUnknown;
return new DataValueSnapshot(null, status, null, Now);
}
/// <summary>
/// Drops every indexed observation. The Agent-restart re-baseline: when the streams header's
/// <c>instanceId</c> changes, every value the index holds describes a device model that no
/// longer exists, and serving it would be worse than serving nothing.
/// </summary>
public void Clear() => _snapshots.Clear();
private DateTime Now => _timeProvider.GetUtcNow().UtcDateTime;
/// <summary>
/// Picks the winner when one dataItemId appears more than once in a single document. A
/// <c>&lt;Condition&gt;</c> container legitimately carries several <b>simultaneously active</b>
/// states (a Fault and a Warning at once), so for a pair of recognized condition state words
/// the <b>most severe</b> wins — plain last-wins would under-report an active Fault as a
/// Warning purely because of element order. Anything else is last-wins, matching the Agent's
/// ascending-sequence ordering.
/// <para>
/// Deliberately scoped to one document: a later <see cref="Apply"/> is the Agent's new
/// statement of the state, so a cleared fault must fall back to Normal. A worst-of that
/// spanned Applies would latch every fault forever.
/// </para>
/// </summary>
private static MTConnectObservation Reconcile(MTConnectObservation incumbent, MTConnectObservation challenger)
{
if (ConditionSeverity.TryGetValue(incumbent.Value ?? string.Empty, out var incumbentSeverity) &&
ConditionSeverity.TryGetValue(challenger.Value ?? string.Empty, out var challengerSeverity))
{
return challengerSeverity >= incumbentSeverity ? challenger : incumbent;
}
return challenger;
}
private DataValueSnapshot BuildSnapshot(MTConnectObservation observation)
{
var serverTimestamp = Now;
var sourceTimestamp = DateTime.SpecifyKind(observation.TimestampUtc, DateTimeKind.Utc);
// No text at all is the degenerate spelling of "no value": MTConnect defines UNAVAILABLE as
// the only way to report absence, so an empty element is not an empty string value.
if (string.IsNullOrWhiteSpace(observation.Value) ||
string.Equals(observation.Value, UnavailableSentinel, StringComparison.Ordinal))
{
return new DataValueSnapshot(null, BadNoCommunication, sourceTimestamp, serverTimestamp);
}
// DATA_SET / TABLE: the Agent carried the content in <Entry key="…"> child elements, so
// Value is their concatenated text with the keys discarded (a two-entry set reads "12").
// Publishing that as a Good string would be noise an operator would trust. Checked before
// the tag lookup because it is a fact about the wire shape, true whether or not the data
// item is authored — and after the sentinel, so an unavailable data set still reports the
// truthful no-comms.
if (observation.IsStructured)
{
return new DataValueSnapshot(null, BadNotSupported, sourceTimestamp, serverTimestamp);
}
_tags.TryGetValue(observation.DataItemId, out var tag);
// TIME_SERIES vectors arrive as one space-separated string. P1 does not materialize OPC UA
// arrays (deferred to P1.5), and the honest answer is "real data I cannot represent" — not a
// type mismatch, and emphatically not the first element parsed out and served as Good.
// Checked AFTER the sentinel so an unavailable array tag still reports the truthful no-comms.
if (tag is { IsArray: true })
{
return new DataValueSnapshot(null, BadNotSupported, sourceTimestamp, serverTimestamp);
}
// An unauthored observation has no declared target type; String is the only coercion that
// cannot fail, and it carries the Agent's text with no interpretation applied.
var status = Coerce(observation.Value, tag?.DriverDataType ?? DriverDataType.String, out var value);
return status == Good
? new DataValueSnapshot(value, Good, sourceTimestamp, serverTimestamp)
: new DataValueSnapshot(null, status, sourceTimestamp, serverTimestamp);
}
/// <summary>
/// Converts the Agent's text to the tag's authored type. Every parse is
/// <see cref="CultureInfo.InvariantCulture"/>: MTConnect is an invariant-format wire, and a
/// host running a comma-decimal culture would otherwise read <c>123.4567</c> as
/// <c>1234567</c> — a silently wrong value under Good quality, the worst possible outcome.
/// </summary>
/// <returns><see cref="Good"/> on success; otherwise a Bad code and a <c>null</c> value.</returns>
private static uint Coerce(string raw, DriverDataType type, out object? value)
{
value = null;
var text = raw.Trim();
switch (type)
{
// Reference is a Galaxy-style attribute reference encoded as an OPC UA String. It is
// never a sensible authoring for MTConnect, but degrading it to text is harmless where
// rejecting it would break a tag for no protection.
case DriverDataType.String:
case DriverDataType.Reference:
value = raw;
return Good;
case DriverDataType.Boolean:
if (bool.TryParse(text, out var flag)) { value = flag; return Good; }
if (text is "1") { value = true; return Good; }
if (text is "0") { value = false; return Good; }
return BadTypeMismatch;
case DriverDataType.DateTime:
if (DateTime.TryParse(
text,
CultureInfo.InvariantCulture,
DateTimeStyles.AdjustToUniversal | DateTimeStyles.AssumeUniversal,
out var moment))
{
value = DateTime.SpecifyKind(moment, DateTimeKind.Utc);
return Good;
}
return BadTypeMismatch;
case DriverDataType.Float32:
if (float.TryParse(text, NumberStyles.Float, CultureInfo.InvariantCulture, out var single))
{
value = single;
return Good;
}
return BadTypeMismatch;
case DriverDataType.Float64:
if (double.TryParse(text, NumberStyles.Float, CultureInfo.InvariantCulture, out var real))
{
value = real;
return Good;
}
return BadTypeMismatch;
case DriverDataType.Int16:
if (short.TryParse(text, NumberStyles.Integer, CultureInfo.InvariantCulture, out var i16))
{
value = i16;
return Good;
}
break;
case DriverDataType.Int32:
if (int.TryParse(text, NumberStyles.Integer, CultureInfo.InvariantCulture, out var i32))
{
value = i32;
return Good;
}
break;
case DriverDataType.Int64:
if (long.TryParse(text, NumberStyles.Integer, CultureInfo.InvariantCulture, out var i64))
{
value = i64;
return Good;
}
break;
case DriverDataType.UInt16:
if (ushort.TryParse(text, NumberStyles.Integer, CultureInfo.InvariantCulture, out var u16))
{
value = u16;
return Good;
}
break;
case DriverDataType.UInt32:
if (uint.TryParse(text, NumberStyles.Integer, CultureInfo.InvariantCulture, out var u32))
{
value = u32;
return Good;
}
break;
case DriverDataType.UInt64:
if (ulong.TryParse(text, NumberStyles.Integer, CultureInfo.InvariantCulture, out var u64))
{
value = u64;
return Good;
}
break;
default:
// A DriverDataType this driver has no rule for. Text always round-trips, so serving
// it beats failing a tag over an enum member added later.
value = raw;
return Good;
}
// Integer parse failed. Distinguishing the two causes is worth one extra parse: it tells an
// operator whether to retype the tag or widen it.
return double.TryParse(text, NumberStyles.Float, CultureInfo.InvariantCulture, out _)
? BadOutOfRange // a number the authored type cannot represent (overflow, or a fraction)
: BadTypeMismatch; // not a number at all
}
}
@@ -0,0 +1,189 @@
using System.Xml;
using System.Xml.Linq;
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// Turns an MTConnect Agent <c>/probe</c> response (an <c>MTConnectDevices</c> XML document)
/// into the neutral <see cref="MTConnectProbeModel"/>. Deliberately a pure, socket-free static
/// so the whole parse is unit-testable against canned fixtures.
/// </summary>
/// <remarks>
/// <para>
/// <b>Why hand-rolled rather than TrakHound.</b> The TrakHound packages Task 0 pinned
/// (<c>MTConnect.NET-Common</c> + <c>-HTTP</c> 6.9.0.2 — both dropped in Task 7) contain <b>no</b>
/// <c>IResponseDocumentFormatter</c> implementation — the XML formatter ships in the
/// separate, unreferenced <c>MTConnect.NET-XML</c> package — so
/// <c>ResponseDocumentFormatter.CreateDevicesResponseDocument("xml", stream)</c> answers
/// <c>Success = false</c> / "Document Formatter Not found for &quot;xml&quot;" (verified by
/// reflection + live invocation against this repo's own fixture, 2026-07-24). TrakHound also
/// offers no socket-free public parse entry point, which the plan requires.
/// </para>
/// <para>
/// <b>Three shapes of the wire format this parser is built around.</b>
/// (1) A device may declare data items <i>directly</i> on <c>&lt;Device&gt;&lt;DataItems&gt;</c>,
/// outside any component — a walker that only recurses <c>&lt;Components&gt;</c> silently
/// drops them.
/// (2) The children of <c>&lt;Components&gt;</c> are named for the component <i>type</i>
/// (<c>&lt;Controller&gt;</c>, <c>&lt;Path&gt;</c>, <c>&lt;Axes&gt;</c>, …) — there is no
/// literal <c>&lt;Component&gt;</c> element — so <i>every</i> element child of
/// <c>&lt;Components&gt;</c> is a component.
/// (3) The document is XML-namespaced and the namespace URI carries the MTConnect version
/// (<c>urn:mtconnect.org:MTConnectDevices:1.3</c> … <c>:2.0</c>). Element matching and
/// attribute reading therefore go through <see cref="MTConnectXml"/>, whose two rules
/// (match on local name, read attributes unqualified) are shared verbatim with
/// <see cref="MTConnectStreamsParser"/> — see that type for why each rule exists.
/// </para>
/// <para>
/// <b>One deliberate divergence from the streams parser: no BOM/whitespace trim here.</b>
/// A <c>/probe</c> body arrives whole from <c>HttpContent</c>, so it begins exactly where
/// the Agent's document begins. A <c>/sample</c> chunk, by contrast, is lifted out of a
/// multipart frame and can carry a byte-order mark or the framing's own line break ahead of
/// the XML declaration, which <c>XmlReader</c> rejects — hence the trim there and not here.
/// This asymmetry is intentional, not an oversight.
/// </para>
/// <para>
/// <b>Failure posture.</b> Anything that is not a well-formed MTConnect device model throws
/// <see cref="InvalidDataException"/>. It never degrades to an empty-but-successful model:
/// an empty device model that looks successful is precisely the defect class that has bitten
/// this codebase before (#485), and here it would tear the driver's browse tree down to
/// nothing while reporting healthy.
/// </para>
/// </remarks>
internal static class MTConnectProbeParser
{
/// <summary>
/// Bound on component nesting depth. Real device models nest a handful of levels; this only
/// exists so a pathological or hostile document cannot recurse the parser into a
/// process-fatal (uncatchable) stack overflow.
/// </summary>
private const int MaxComponentDepth = 64;
/// <summary>How <see cref="MTConnectXml"/>'s shared error messages name this document.</summary>
private const string Subject = "/probe response";
/// <summary>Parses a <c>/probe</c> response held as a string.</summary>
/// <param name="xml">The raw <c>MTConnectDevices</c> XML document.</param>
/// <exception cref="InvalidDataException">
/// The payload is empty, not well-formed XML, not an <c>MTConnectDevices</c> document (an
/// Agent <c>MTConnectError</c> document included — Agents answer a bad device path with one
/// under HTTP 200), or declares no usable device.
/// </exception>
public static MTConnectProbeModel Parse(string xml)
{
if (string.IsNullOrWhiteSpace(xml))
{
throw new InvalidDataException("MTConnect /probe response was empty; expected an MTConnectDevices XML document.");
}
XDocument document;
try
{
document = XDocument.Parse(xml);
}
catch (XmlException ex)
{
throw new InvalidDataException($"MTConnect /probe response is not well-formed XML: {ex.Message}", ex);
}
return Build(document);
}
/// <summary>Parses a <c>/probe</c> response read from a stream.</summary>
/// <param name="stream">A stream positioned at the start of the XML document.</param>
/// <exception cref="InvalidDataException">As for <see cref="Parse(string)"/>.</exception>
public static MTConnectProbeModel Parse(Stream stream)
{
ArgumentNullException.ThrowIfNull(stream);
XDocument document;
try
{
document = XDocument.Load(stream);
}
catch (XmlException ex)
{
throw new InvalidDataException($"MTConnect /probe response is not well-formed XML: {ex.Message}", ex);
}
return Build(document);
}
private static MTConnectProbeModel Build(XDocument document)
{
var root = document.Root
?? throw new InvalidDataException("MTConnect /probe response has no root element.");
if (!MTConnectXml.IsNamed(root, "MTConnectDevices"))
{
throw new InvalidDataException(MTConnectXml.DescribeUnexpectedRoot(root, "MTConnectDevices", Subject));
}
var devicesContainer = MTConnectXml.ChildrenNamed(root, "Devices").FirstOrDefault()
?? throw new InvalidDataException(
"MTConnect /probe response has no <Devices> element; it does not describe a device model.");
var devices = devicesContainer.Elements().Select(ReadDevice).ToList();
if (devices.Count == 0)
{
throw new InvalidDataException(
"MTConnect /probe response declares no devices; refusing to report an empty device model as a successful probe.");
}
return new MTConnectProbeModel(devices);
}
private static MTConnectDevice ReadDevice(XElement element) =>
new(
MTConnectXml.RequiredAttribute(element, "id", "Device", Subject),
MTConnectXml.OptionalAttribute(element, "name"),
ReadComponents(element, depth: 1),
ReadDataItems(element));
private static MTConnectComponent ReadComponent(XElement element, int depth) =>
new(
MTConnectXml.RequiredAttribute(element, "id", $"Component <{element.Name.LocalName}>", Subject),
MTConnectXml.OptionalAttribute(element, "name"),
ReadComponents(element, depth + 1),
ReadDataItems(element));
/// <summary>
/// Every element child of this container's <c>&lt;Components&gt;</c> element, whatever it is
/// named — the element name is the component type, not the literal string "Component".
/// </summary>
private static IReadOnlyList<MTConnectComponent> ReadComponents(XElement container, int depth)
{
if (depth > MaxComponentDepth)
{
throw new InvalidDataException(
$"MTConnect /probe response nests components more than {MaxComponentDepth} levels deep; refusing to recurse further.");
}
return MTConnectXml.ChildrenNamed(container, "Components")
.SelectMany(components => components.Elements())
.Select(element => ReadComponent(element, depth))
.ToList();
}
private static IReadOnlyList<MTConnectDataItem> ReadDataItems(XElement container) =>
MTConnectXml.ChildrenNamed(container, "DataItems")
.SelectMany(dataItems => MTConnectXml.ChildrenNamed(dataItems, "DataItem"))
.Select(ReadDataItem)
.ToList();
private static MTConnectDataItem ReadDataItem(XElement element)
{
var id = MTConnectXml.RequiredAttribute(element, "id", "DataItem", Subject);
return new MTConnectDataItem(
id,
MTConnectXml.OptionalAttribute(element, "name"),
MTConnectXml.RequiredAttribute(element, "category", $"DataItem '{id}'", Subject),
MTConnectXml.RequiredAttribute(element, "type", $"DataItem '{id}'", Subject),
MTConnectXml.OptionalAttribute(element, "subType"),
MTConnectXml.OptionalAttribute(element, "units"),
MTConnectXml.OptionalAttribute(element, "representation"),
MTConnectXml.OptionalIntAttribute(element, "sampleCount", $"DataItem '{id}'", Subject));
}
}
@@ -0,0 +1,29 @@
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// Driver-internal identity of one <c>/sample</c> subscription, matching the sibling drivers'
/// shape (Galaxy's <c>GalaxySubscriptionHandle</c>, the shared <c>PollGroupEngine</c>'s
/// <c>PollSubscriptionHandle</c>): a monotonic per-driver id plus a diagnostic string that
/// carries it into logs.
/// </summary>
/// <remarks>
/// <para>
/// <b>Every handle shares ONE Agent stream.</b> Unlike a polled driver, where each
/// subscription owns a loop, MTConnect's <c>/sample</c> long poll is per-Agent: the driver
/// opens exactly one and fans each chunk out to whichever handles subscribe the reporting
/// DataItem. The handle is therefore purely an identity — it owns no connection, no task,
/// and no cursor — and dropping one only stops the stream when it was the last.
/// </para>
/// <para>
/// A <see langword="record"/> for value equality on the id, so a handle that has round-tripped
/// through the caller still resolves. The id is never reused within a driver instance.
/// </para>
/// </remarks>
/// <param name="SubscriptionId">The monotonic per-driver subscription id.</param>
internal sealed record MTConnectSampleHandle(long SubscriptionId) : ISubscriptionHandle
{
/// <inheritdoc/>
public string DiagnosticId => $"mtconnect-sub-{SubscriptionId}";
}
@@ -0,0 +1,161 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// Base type for every way an Agent <c>/sample</c> long poll can stop delivering chunks for a
/// reason that is <b>not</b> the caller cancelling it.
/// </summary>
/// <remarks>
/// <para>
/// <b>Why these are exceptions rather than a normal end of enumeration.</b>
/// <see cref="IMTConnectAgentClient.SampleAsync"/> is contracted to yield indefinitely until
/// its token is cancelled, so a pump written to that contract has no reason to handle normal
/// completion — it would either drop the subscription silently or hot-loop reconnecting on
/// an enumerator that ends immediately. Returning normally would make "the data path
/// stopped" indistinguishable from "the data path is idle", which is the
/// quiet-successful-looking-termination shape that produced this repo's #485 defect. Every
/// non-cancellation end is therefore loud, typed, and carries the Agent URI.
/// </para>
/// <para>
/// Catch this base type to handle "the stream is over" uniformly; catch the two derived
/// types to distinguish a <i>transient</i> end (the Agent closed a connection — reconnect)
/// from a <i>configuration</i> error (the Agent will never stream to this request — retrying
/// is pointless until an operator changes something).
/// </para>
/// </remarks>
public abstract class MTConnectStreamException : Exception
{
/// <summary>Creates the exception with no message.</summary>
protected MTConnectStreamException()
{
}
/// <summary>Creates the exception with a message.</summary>
/// <param name="message">The operator-facing description.</param>
protected MTConnectStreamException(string message)
: base(message)
{
}
/// <summary>Creates the exception with a message and an inner cause.</summary>
/// <param name="message">The operator-facing description.</param>
/// <param name="innerException">The underlying cause.</param>
protected MTConnectStreamException(string message, Exception innerException)
: base(message, innerException)
{
}
}
/// <summary>How an Agent <c>/sample</c> stream ended.</summary>
public enum MTConnectStreamEndReason
{
/// <summary>
/// The transport closed mid-stream — the connection dropped, the Agent restarted, or a proxy
/// timed the connection out. Ordinary and transient: reconnect from the last
/// <see cref="MTConnectStreamsResult.NextSequence"/>.
/// </summary>
ConnectionClosed = 0,
/// <summary>
/// The Agent wrote the closing <c>--boundary--</c> delimiter, ending the multipart response
/// deliberately. Usually means the request was bounded (an Agent honouring a <c>count</c>
/// that exhausted, or an administrative stop) rather than a fault. Also reconnectable, but
/// worth distinguishing in logs: a stream that keeps closing cleanly points at the request's
/// parameters, not at the network.
/// </summary>
ClosingBoundary = 1
}
/// <summary>
/// The Agent's <c>/sample</c> stream ended without the caller cancelling it. See
/// <see cref="MTConnectStreamException"/> for why this is not a normal end of enumeration.
/// </summary>
public sealed class MTConnectStreamEndedException : MTConnectStreamException
{
/// <summary>Creates the exception for a stream that ended for <paramref name="reason"/>.</summary>
/// <param name="reason">How the stream ended.</param>
/// <param name="agentUri">The <c>/sample</c> request URI whose stream ended.</param>
/// <param name="chunksDelivered">How many chunks were yielded before the end.</param>
public MTConnectStreamEndedException(MTConnectStreamEndReason reason, string agentUri, long chunksDelivered)
: base($"MTConnect /sample stream from '{agentUri}' ended after {chunksDelivered} chunk(s) without the caller cancelling it ({Describe(reason)}). The stream is contracted to run until cancelled, so this is reported rather than returned; resume from the last chunk's nextSequence.")
{
Reason = reason;
AgentUri = agentUri;
ChunksDelivered = chunksDelivered;
}
/// <summary>Creates the exception with a message.</summary>
/// <param name="message">The operator-facing description.</param>
public MTConnectStreamEndedException(string message)
: base(message)
{
AgentUri = string.Empty;
}
/// <summary>Creates the exception with a message and an inner cause.</summary>
/// <param name="message">The operator-facing description.</param>
/// <param name="innerException">The underlying cause.</param>
public MTConnectStreamEndedException(string message, Exception innerException)
: base(message, innerException)
{
AgentUri = string.Empty;
}
/// <summary>Creates the exception with no message.</summary>
public MTConnectStreamEndedException()
{
AgentUri = string.Empty;
}
/// <summary>How the stream ended.</summary>
public MTConnectStreamEndReason Reason { get; }
/// <summary>The <c>/sample</c> request URI whose stream ended.</summary>
public string AgentUri { get; }
/// <summary>How many chunks were yielded before the stream ended.</summary>
public long ChunksDelivered { get; }
private static string Describe(MTConnectStreamEndReason reason) =>
reason switch
{
MTConnectStreamEndReason.ClosingBoundary => "the Agent wrote the closing multipart boundary",
_ => "the connection closed mid-stream"
};
}
/// <summary>
/// The Agent answered a <c>/sample</c> request with something that cannot be consumed as a
/// stream at all — a single non-multipart document, or a <c>multipart/*</c> response with no
/// <c>boundary</c> parameter to frame it by.
/// </summary>
/// <remarks>
/// This is a <b>configuration</b> error, not a transient one, and is deliberately typed apart
/// from <see cref="MTConnectStreamEndedException"/>: reconnecting will produce the identical
/// response forever. The usual cause is an endpoint that is not an MTConnect Agent (a reverse
/// proxy, a load balancer's health page, or an Agent fronted by something that buffers and
/// collapses <c>multipart/x-mixed-replace</c>). Reported instead of yielding the one document
/// the Agent did return, because a single snapshot followed by a healthy-looking finished
/// stream is exactly how a dead data path disguises itself as a working one.
/// </remarks>
public sealed class MTConnectStreamNotSupportedException : MTConnectStreamException
{
/// <summary>Creates the exception with a message.</summary>
/// <param name="message">The operator-facing description.</param>
public MTConnectStreamNotSupportedException(string message)
: base(message)
{
}
/// <summary>Creates the exception with a message and an inner cause.</summary>
/// <param name="message">The operator-facing description.</param>
/// <param name="innerException">The underlying cause.</param>
public MTConnectStreamNotSupportedException(string message, Exception innerException)
: base(message, innerException)
{
}
/// <summary>Creates the exception with no message.</summary>
public MTConnectStreamNotSupportedException()
{
}
}
@@ -0,0 +1,240 @@
using System.Globalization;
using System.Xml;
using System.Xml.Linq;
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// Turns an MTConnect Agent <c>MTConnectStreams</c> document — a <c>/current</c> snapshot or one
/// chunk of a <c>/sample</c> multipart stream — into the neutral
/// <see cref="MTConnectStreamsResult"/>. Deliberately a pure, socket-free static so the whole
/// parse is unit-testable against canned fixtures; chunk framing lives in
/// <see cref="MTConnectAgentClient"/> and hands this parser one chunk's XML at a time.
/// </summary>
/// <remarks>
/// <para>
/// <b>Why hand-rolled rather than TrakHound.</b> Same finding as
/// <see cref="MTConnectProbeParser"/>: the packages pinned by Task 0 ship no XML document
/// formatter, and TrakHound's streaming client offers no socket-free entry point. Both
/// package references were dropped in Task 7; the driver depends only on the BCL.
/// </para>
/// <para>
/// <b>The Streams document is shaped nothing like the Devices document</b>, and three
/// differences drive this parser's design.
/// </para>
/// <list type="number">
/// <item>
/// <description>
/// <b>The observation element's name is the data item's type in PascalCase</b>
/// (<c>Position</c>, <c>PartCount</c>, <c>Execution</c>) — not the probe's
/// UPPER_SNAKE <c>POSITION</c> / <c>PART_COUNT</c>. The standard defines hundreds of
/// types and vendors add more, so matching a fixed set of element names would drop
/// observations silently. <b>Every</b> element child of <c>&lt;Samples&gt;</c>,
/// <c>&lt;Events&gt;</c> and <c>&lt;Condition&gt;</c> is therefore an observation,
/// keyed by its <c>dataItemId</c> attribute — the only identity the probe model and
/// the streams document actually share.
/// </description>
/// </item>
/// <item>
/// <description>
/// <b>A CONDITION observation's value is its ELEMENT NAME, not its text.</b>
/// <c>&lt;Normal/&gt;</c> / <c>&lt;Warning/&gt;</c> / <c>&lt;Fault/&gt;</c> /
/// <c>&lt;Unavailable/&gt;</c> carry the state in the name and are typically empty;
/// where text is present it is the operator-facing message, not the state. Reading
/// the text would report every condition as an empty string.
/// <c>&lt;Unavailable/&gt;</c> is normalized to the literal
/// <see cref="UnavailableSentinel"/> so a no-comms condition lands on exactly the
/// same sentinel as a Sample/Event whose text is <c>UNAVAILABLE</c> — the
/// observation index maps that one token to bad quality.
/// </description>
/// </item>
/// <item>
/// <description>
/// <b>Timestamps are normalized to UTC explicitly.</b> MTConnect timestamps are
/// ISO-8601 Zulu, but a bare <c>DateTime.Parse</c> yields <c>Local</c> or
/// <c>Unspecified</c> depending on machine settings — and this value becomes the
/// OPC UA SourceTimestamp, so a wrong <see cref="DateTimeKind"/> shifts every
/// timestamp by the host's UTC offset with nothing to show for it.
/// </description>
/// </item>
/// </list>
/// <para>
/// <b>Failure posture.</b> A malformed document, a non-<c>MTConnectStreams</c> root (an Agent
/// <c>MTConnectError</c> served under HTTP 200 included), an unusable header, or an
/// observation missing its <c>dataItemId</c>/<c>timestamp</c> all throw
/// <see cref="InvalidDataException"/>. A document with a valid header and <i>zero</i>
/// observations, by contrast, is entirely legal — <c>/sample</c> chunks are deltas and an
/// idle Agent's keep-alive is an observation-free document — so it parses to an empty
/// observation list rather than throwing.
/// </para>
/// </remarks>
internal static class MTConnectStreamsParser
{
/// <summary>
/// The single token that means "the Agent has no value for this data item". Sample/Event
/// observations carry it as text; a CONDITION carries it as the element name
/// <c>&lt;Unavailable/&gt;</c> and is normalized onto this spelling here.
/// </summary>
private const string UnavailableSentinel = "UNAVAILABLE";
private const string ConditionContainer = "Condition";
/// <summary>How <see cref="MTConnectXml"/>'s shared error messages name this document.</summary>
private const string Subject = "streams response";
/// <summary>How those messages name the header element.</summary>
private const string HeaderOwner = "the <Header>";
/// <summary>
/// The three elements whose element children are observations. A ComponentStream may carry
/// any subset of them, including none.
/// </summary>
private static readonly string[] ObservationContainers = ["Samples", "Events", ConditionContainer];
/// <summary>Parses a <c>/current</c> response, or one <c>/sample</c> chunk, held as a string.</summary>
/// <param name="xml">The raw <c>MTConnectStreams</c> XML document.</param>
/// <exception cref="InvalidDataException">
/// The payload is empty, not well-formed XML, not an <c>MTConnectStreams</c> document (an
/// Agent <c>MTConnectError</c> included — Agents answer a bad request with one under HTTP
/// 200), carries no usable <c>&lt;Header&gt;</c>, or carries a malformed observation.
/// </exception>
public static MTConnectStreamsResult Parse(string xml)
{
if (string.IsNullOrWhiteSpace(xml))
{
throw new InvalidDataException(
"MTConnect streams response was empty; expected an MTConnectStreams XML document.");
}
XDocument document;
try
{
// A chunk lifted out of a multipart frame can carry a byte-order mark or the framing's
// trailing line break; XmlReader rejects either ahead of the XML declaration.
document = XDocument.Parse(xml.Trim('\uFEFF', ' ', '\t', '\r', '\n'));
}
catch (XmlException ex)
{
throw new InvalidDataException($"MTConnect streams response is not well-formed XML: {ex.Message}", ex);
}
return Build(document);
}
/// <summary>Parses a <c>/current</c> response read from a stream.</summary>
/// <param name="stream">A stream positioned at the start of the XML document.</param>
/// <exception cref="InvalidDataException">As for <see cref="Parse(string)"/>.</exception>
public static MTConnectStreamsResult Parse(Stream stream)
{
ArgumentNullException.ThrowIfNull(stream);
XDocument document;
try
{
document = XDocument.Load(stream);
}
catch (XmlException ex)
{
throw new InvalidDataException($"MTConnect streams response is not well-formed XML: {ex.Message}", ex);
}
return Build(document);
}
private static MTConnectStreamsResult Build(XDocument document)
{
var root = document.Root
?? throw new InvalidDataException("MTConnect streams response has no root element.");
if (!MTConnectXml.IsNamed(root, "MTConnectStreams"))
{
throw new InvalidDataException(MTConnectXml.DescribeUnexpectedRoot(root, "MTConnectStreams", Subject));
}
var header = MTConnectXml.ChildrenNamed(root, "Header").FirstOrDefault()
?? throw new InvalidDataException(
"MTConnect streams response has no <Header>; without it the sequence bookkeeping the sample pump runs on is unknowable.");
return new MTConnectStreamsResult(
MTConnectXml.RequiredLongAttribute(header, "instanceId", HeaderOwner, Subject),
MTConnectXml.RequiredLongAttribute(header, "nextSequence", HeaderOwner, Subject),
MTConnectXml.RequiredLongAttribute(header, "firstSequence", HeaderOwner, Subject),
ReadObservations(root));
}
/// <summary>
/// Collects every observation in the document, in Agent-supplied order. Containers are found
/// by descending the whole <c>&lt;Streams&gt;</c> subtree rather than by walking a fixed
/// DeviceStream → ComponentStream chain, so a vendor's extra nesting level cannot silently
/// hide a device's observations.
/// </summary>
private static IReadOnlyList<MTConnectObservation> ReadObservations(XElement root)
{
var observations = new List<MTConnectObservation>();
foreach (var container in root.Descendants().Where(IsObservationContainer))
{
var isCondition = MTConnectXml.IsNamed(container, ConditionContainer);
foreach (var element in container.Elements())
{
observations.Add(ReadObservation(element, isCondition));
}
}
return observations;
}
private static MTConnectObservation ReadObservation(XElement element, bool isCondition)
{
var dataItemId = MTConnectXml.RequiredAttribute(element, "dataItemId", $"Observation <{element.Name.LocalName}>", Subject);
return new MTConnectObservation(
dataItemId,
isCondition ? ConditionState(element) : element.Value.Trim(),
ReadTimestampUtc(element, dataItemId),
// A DATA_SET/TABLE observation keeps its content in <Entry key="…"> children, so the
// text read above concatenates them into nonsense ("12" for two entries). Flagged here
// because this is the only layer that can still see the distinction — see
// MTConnectObservation.IsStructured. Never true for a CONDITION: its value comes from
// the element NAME, so child content cannot corrupt it, and flagging one would throw
// away a perfectly well-determined state.
IsStructured: !isCondition && element.HasElements);
}
/// <summary>
/// A condition's state is its element name. <c>Unavailable</c> is normalized to the
/// <see cref="UnavailableSentinel"/> spelling so downstream quality mapping has exactly one
/// token to recognize; every other state is reported as-named (<c>Normal</c>, <c>Warning</c>,
/// <c>Fault</c>, and any vendor state).
/// </summary>
private static string ConditionState(XElement element)
{
var state = element.Name.LocalName;
return string.Equals(state, "Unavailable", StringComparison.OrdinalIgnoreCase) ? UnavailableSentinel : state;
}
private static DateTime ReadTimestampUtc(XElement element, string dataItemId)
{
var raw = MTConnectXml.RequiredAttribute(element, "timestamp", $"Observation '{dataItemId}'", Subject);
if (!DateTime.TryParse(
raw,
CultureInfo.InvariantCulture,
DateTimeStyles.AdjustToUniversal | DateTimeStyles.AssumeUniversal,
out var parsed))
{
throw new InvalidDataException(
$"MTConnect streams response is malformed: observation '{dataItemId}' has an unparseable 'timestamp' ('{raw}').");
}
// AdjustToUniversal already yields Utc; stated explicitly so a future styles change cannot
// quietly hand the OPC UA layer a Local or Unspecified SourceTimestamp.
return DateTime.SpecifyKind(parsed, DateTimeKind.Utc);
}
private static bool IsObservationContainer(XElement element) =>
ObservationContainers.Any(name => MTConnectXml.IsNamed(element, name));
}
@@ -0,0 +1,146 @@
using System.Globalization;
using System.Xml.Linq;
namespace ZB.MOM.WW.OtOpcUa.Driver.MTConnect;
/// <summary>
/// The element/attribute reading rules shared by <see cref="MTConnectProbeParser"/> and
/// <see cref="MTConnectStreamsParser"/>. Both documents are the same dialect of XML and must be
/// read by the same rules; keeping one copy is what stops the two parsers drifting apart on the
/// two rules below, either of which fails silently rather than loudly.
/// </summary>
/// <remarks>
/// <para>
/// <b>Rule 1: elements are matched on <see cref="XName.LocalName"/>, ignoring the
/// namespace.</b> The namespace URI carries the MTConnect version
/// (<c>urn:mtconnect.org:MTConnectDevices:1.3</c> … <c>:2.0</c>), so a fully-qualified match
/// would hard-code a version. It also silently drops vendor-extension elements declared in
/// their own namespace, whose standard children still inherit the default MTConnect
/// namespace — the browse tree simply comes back short, with no error anywhere.
/// </para>
/// <para>
/// <b>Rule 2: attributes are read <i>unqualified</i>.</b> A prefixed attribute
/// (<c>xsi:type</c> on a document root, a vendor's <c>x:dataItemId</c>) must never be
/// mistaken for the real <c>type</c> / <c>dataItemId</c>, which would invent a data item the
/// probe never declared.
/// </para>
/// </remarks>
internal static class MTConnectXml
{
/// <summary>Does <paramref name="element"/> have the given local name, whatever its namespace?</summary>
/// <param name="element">The element to test.</param>
/// <param name="localName">The unqualified name to match.</param>
public static bool IsNamed(XElement element, string localName) =>
string.Equals(element.Name.LocalName, localName, StringComparison.Ordinal);
/// <summary>Direct element children of <paramref name="parent"/> with the given local name.</summary>
/// <param name="parent">The element whose children to filter.</param>
/// <param name="localName">The unqualified name to match.</param>
public static IEnumerable<XElement> ChildrenNamed(XElement parent, string localName) =>
parent.Elements().Where(child => IsNamed(child, localName));
/// <summary>
/// Reads an unqualified attribute, normalizing a present-but-empty value to <c>null</c>.
/// See the type-level remarks for why only unqualified attributes are considered.
/// </summary>
/// <param name="element">The element carrying the attribute.</param>
/// <param name="name">The unqualified attribute name.</param>
public static string? OptionalAttribute(XElement element, string name)
{
var value = element.Attribute(name)?.Value;
return string.IsNullOrWhiteSpace(value) ? null : value;
}
/// <summary>Reads a required unqualified attribute, or throws naming the owner and the attribute.</summary>
/// <param name="element">The element carrying the attribute.</param>
/// <param name="name">The unqualified attribute name.</param>
/// <param name="owner">How to describe the element in the error message.</param>
/// <param name="subject">How to describe the document in the error message (e.g. "/probe response").</param>
/// <exception cref="InvalidDataException">The attribute is absent or blank.</exception>
public static string RequiredAttribute(XElement element, string name, string owner, string subject)
{
var value = OptionalAttribute(element, name);
return value ?? throw new InvalidDataException(
$"MTConnect {subject} is malformed: {owner} has no '{name}' attribute.");
}
/// <summary>Reads an optional unqualified attribute as an <see cref="int"/>.</summary>
/// <param name="element">The element carrying the attribute.</param>
/// <param name="name">The unqualified attribute name.</param>
/// <param name="owner">How to describe the element in the error message.</param>
/// <param name="subject">How to describe the document in the error message.</param>
/// <exception cref="InvalidDataException">The attribute is present but not an integer.</exception>
public static int? OptionalIntAttribute(XElement element, string name, string owner, string subject)
{
var raw = OptionalAttribute(element, name);
if (raw is null)
{
return null;
}
if (!int.TryParse(raw, NumberStyles.Integer, CultureInfo.InvariantCulture, out var value))
{
throw new InvalidDataException(
$"MTConnect {subject} is malformed: {owner} has a non-numeric '{name}' attribute ('{raw}').");
}
return value;
}
/// <summary>Reads a required unqualified attribute as a <see cref="long"/>.</summary>
/// <param name="element">The element carrying the attribute.</param>
/// <param name="name">The unqualified attribute name.</param>
/// <param name="owner">How to describe the element in the error message.</param>
/// <param name="subject">How to describe the document in the error message.</param>
/// <exception cref="InvalidDataException">The attribute is absent, blank, or not an integer.</exception>
public static long RequiredLongAttribute(XElement element, string name, string owner, string subject)
{
var raw = RequiredAttribute(element, name, owner, subject);
if (!long.TryParse(raw, NumberStyles.Integer, CultureInfo.InvariantCulture, out var value))
{
throw new InvalidDataException(
$"MTConnect {subject} is malformed: {owner} has a non-numeric '{name}' attribute ('{raw}').");
}
return value;
}
/// <summary>
/// Builds the error message for a document whose root is not what was expected, lifting the
/// Agent's own error text when the response is an <c>MTConnectError</c> document — which
/// Agents return under <b>HTTP 200</b>, so <c>EnsureSuccessStatusCode</c> never fires and
/// this is the only place the operator can learn what the Agent objected to.
/// </summary>
/// <param name="root">The document's actual root element.</param>
/// <param name="expectedRootName">The root element name the caller required.</param>
/// <param name="subject">How to describe the document in the error message.</param>
public static string DescribeUnexpectedRoot(XElement root, string expectedRootName, string subject)
{
var message =
$"MTConnect {subject} root element is <{root.Name.LocalName}>, expected <{expectedRootName}>.";
if (!IsNamed(root, "MTConnectError"))
{
return message;
}
var errors = root.Descendants()
.Where(e => IsNamed(e, "Error"))
.Select(e =>
{
var code = OptionalAttribute(e, "errorCode");
var text = e.Value.Trim();
return code is null ? text : $"{code}: {text}";
})
.Where(text => text.Length > 0)
.ToList();
return errors.Count == 0
? message
: $"{message} The Agent reported: {string.Join("; ", errors)}";
}
}
@@ -0,0 +1,31 @@
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup>
<TargetFramework>net10.0</TargetFramework>
<Nullable>enable</Nullable>
<ImplicitUsings>enable</ImplicitUsings>
<LangVersion>latest</LangVersion>
<TreatWarningsAsErrors>true</TreatWarningsAsErrors>
<GenerateDocumentationFile>true</GenerateDocumentationFile>
<NoWarn>$(NoWarn);CS1591</NoWarn>
<RootNamespace>ZB.MOM.WW.OtOpcUa.Driver.MTConnect</RootNamespace>
</PropertyGroup>
<ItemGroup>
<ProjectReference Include="..\..\Core\ZB.MOM.WW.OtOpcUa.Core.Abstractions\ZB.MOM.WW.OtOpcUa.Core.Abstractions.csproj"/>
<ProjectReference Include="..\..\Core\ZB.MOM.WW.OtOpcUa.Core\ZB.MOM.WW.OtOpcUa.Core.csproj"/>
<ProjectReference Include="..\ZB.MOM.WW.OtOpcUa.Driver.MTConnect.Contracts\ZB.MOM.WW.OtOpcUa.Driver.MTConnect.Contracts.csproj"/>
</ItemGroup>
<!-- No backend NuGet: the Agent surface is plain HTTP + XML, served entirely by the BCL
(HttpClient + System.Xml.Linq). The TrakHound MTConnect.NET-Common/-HTTP references Task 0
added were removed in Task 7 — they can neither parse an MTConnect document (no XML
formatter ships in either package) nor frame the /sample stream socket-free (their client
type dials its own URL). See the Task 0 CORRECTION block in
docs/plans/2026-07-24-mtconnect-driver.md. -->
<ItemGroup>
<InternalsVisibleTo Include="ZB.MOM.WW.OtOpcUa.Driver.MTConnect.Tests"/>
</ItemGroup>
</Project>
@@ -130,6 +130,14 @@ public sealed class ModbusDriverOptions
/// </summary>
public MelsecFamily MelsecSubFamily { get; init; } = MelsecFamily.Q_L_iQR;
/// <summary>
/// Wire transport selector. <see cref="ModbusTransportMode.Tcp"/> (default) is the
/// historical Modbus TCP/MBAP framing. <see cref="ModbusTransportMode.RtuOverTcp"/>
/// frames Modbus RTU PDUs (CRC16, no MBAP header) over the same TCP socket — for
/// serial-to-Ethernet gateways that bridge RS-485 without re-encapsulating to MBAP.
/// </summary>
public ModbusTransportMode Transport { get; init; } = ModbusTransportMode.Tcp;
/// <summary>
/// When <c>true</c>, the driver suppresses redundant writes: if the most recent
/// successful write to a tag carried value V and a new write of V arrives, the second
@@ -0,0 +1,17 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Modbus;
/// <summary>
/// Wire transport for a Modbus driver instance. <see cref="Tcp"/> = Modbus/TCP (MBAP + TxId);
/// <see cref="RtuOverTcp"/> = RTU framing (address + CRC-16, no MBAP) tunnelled over a socket to
/// a serial→Ethernet gateway. Default <see cref="Tcp"/> — existing configs omit the field.
/// A direct-serial <c>Rtu</c> member is intentionally NOT reserved here (descoped 2026-07-15);
/// do not number-squat it.
/// </summary>
public enum ModbusTransportMode
{
/// <summary>Modbus/TCP — 7-byte MBAP header + transaction id (the historical default).</summary>
Tcp,
/// <summary>RTU framing (<c>[addr][PDU][CRC-16]</c>) tunnelled over a socket to a serial→Ethernet gateway.</summary>
RtuOverTcp,
}
@@ -0,0 +1,48 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Modbus;
/// <summary>
/// CRC-16/MODBUS helper: reflected polynomial <c>0xA001</c>, initial value <c>0xFFFF</c>.
/// On the wire the CRC is appended low byte first, high byte second — <see cref="Append"/>
/// does that; <see cref="Compute"/> just returns the 16-bit value. Byte-stream-agnostic:
/// no socket or framing knowledge lives here.
/// </summary>
public static class ModbusCrc
{
/// <summary>Computes the CRC-16/MODBUS checksum over <paramref name="data"/>.</summary>
/// <param name="data">The bytes to checksum (e.g. the RTU frame minus the trailing CRC).</param>
/// <returns>The 16-bit CRC value.</returns>
public static ushort Compute(ReadOnlySpan<byte> data)
{
ushort crc = 0xFFFF;
foreach (var b in data)
{
crc ^= b;
for (var i = 0; i < 8; i++)
{
var carry = (crc & 0x0001) != 0;
crc >>= 1;
if (carry)
{
crc ^= 0xA001;
}
}
}
return crc;
}
/// <summary>
/// Returns <paramref name="body"/> with its CRC-16/MODBUS appended, low byte first.
/// </summary>
/// <param name="body">The frame body to checksum.</param>
/// <returns>A new array: <paramref name="body"/> followed by <c>[crc &amp; 0xFF, crc &gt;&gt; 8]</c>.</returns>
public static byte[] Append(ReadOnlySpan<byte> body)
{
var crc = Compute(body);
var result = new byte[body.Length + 2];
body.CopyTo(result);
result[^2] = (byte)(crc & 0xFF);
result[^1] = (byte)(crc >> 8);
return result;
}
}
@@ -94,7 +94,8 @@ public sealed class ModbusDriver
/// <summary>Initializes a new Modbus TCP driver with the specified options and transport factory.</summary>
/// <param name="options">Driver configuration options.</param>
/// <param name="driverInstanceId">Unique identifier for this driver instance.</param>
/// <param name="transportFactory">Factory to create the Modbus transport; defaults to ModbusTcpTransport.</param>
/// <param name="transportFactory">Factory to create the Modbus transport; defaults to
/// <see cref="ModbusTransportFactory.Create"/>, which honours <see cref="ModbusDriverOptions.Transport"/>.</param>
/// <param name="logger">Logger instance; defaults to null logger if not provided.</param>
public ModbusDriver(ModbusDriverOptions options, string driverInstanceId,
Func<ModbusDriverOptions, IModbusTransport>? transportFactory = null,
@@ -107,11 +108,7 @@ public sealed class ModbusDriver
_resolver = new EquipmentTagRefResolver<ModbusTagDefinition>(
r => _tagsByRawPath.TryGetValue(r, out var t) ? t : null);
_transportFactory = transportFactory
?? (o => new ModbusTcpTransport(
o.Host, o.Port, o.Timeout, o.AutoReconnect,
keepAlive: o.KeepAlive,
idleDisconnect: o.IdleDisconnectTimeout,
reconnect: o.Reconnect));
?? (o => ModbusTransportFactory.Create(o));
_poll = new PollGroupEngine(
reader: ReadAsync,
onChange: (handle, tagRef, snapshot) =>
@@ -74,6 +74,8 @@ public static class ModbusDriverFactoryExtensions
: ParseEnum<ModbusFamily>(dto.Family, "<driver-level>", driverInstanceId, "Family"),
MelsecSubFamily = dto.MelsecSubFamily is null ? MelsecFamily.Q_L_iQR
: ParseEnum<MelsecFamily>(dto.MelsecSubFamily, "<driver-level>", driverInstanceId, "MelsecSubFamily"),
Transport = dto.Transport is null ? ModbusTransportMode.Tcp
: ParseEnum<ModbusTransportMode>(dto.Transport, "<driver-level>", driverInstanceId, "Transport"),
AutoProhibitReprobeInterval = dto.AutoProhibitReprobeMs is { } reprobeMs ? TimeSpan.FromMilliseconds(reprobeMs) : null,
AutoReconnect = dto.AutoReconnect ?? true,
// v3: the deploy artifact (Wave C) populates rawTags with the cluster's authored raw Modbus
@@ -159,6 +161,8 @@ public static class ModbusDriverFactoryExtensions
public string? Family { get; init; }
/// <summary>Gets or sets the Melsec subfamily.</summary>
public string? MelsecSubFamily { get; init; }
/// <summary>Gets or sets the wire transport mode (Tcp / RtuOverTcp). Null defaults to Tcp.</summary>
public string? Transport { get; init; }
/// <summary>Gets or sets the automatic prohibition reprobei interval in milliseconds.</summary>
public int? AutoProhibitReprobeMs { get; init; }
/// <summary>Gets or sets a value indicating whether automatic reconnection is enabled.</summary>
@@ -29,15 +29,19 @@ public sealed class ModbusDriverProbe : IDriverProbe
/// <inheritdoc />
public string DriverType => "Modbus";
/// <summary>The parsed probe target — host/port/unitId + the effective connection timeout, all read
/// from the SAME factory DTO shape (<c>ModbusDriverConfigDto</c>) the driver factory parses, so
/// <c>timeoutMs</c> binds identically to the factory (R2-11, 05/CONV-2; the OpcUaClient parity rule).</summary>
internal readonly record struct ProbeTarget(string Host, int Port, byte UnitId, TimeSpan Timeout);
/// <summary>The parsed probe target — host/port/unitId + the effective connection timeout + the wire
/// transport mode, all read from the SAME factory DTO shape (<c>ModbusDriverConfigDto</c>) the driver
/// factory parses, so <c>timeoutMs</c> binds identically to the factory (R2-11, 05/CONV-2; the
/// OpcUaClient parity rule) and an <c>RtuOverTcp</c>-authored config probes over RTU framing —
/// matching how the driver will actually run.</summary>
internal readonly record struct ProbeTarget(string Host, int Port, byte UnitId, TimeSpan Timeout, ModbusTransportMode Transport);
/// <summary>Parse the driver config JSON into a <see cref="ProbeTarget"/> using the factory DTO shape.
/// Returns a null target + an error message when the JSON is invalid or has no host/port. The effective
/// timeout mirrors the factory (<c>TimeSpan.FromMilliseconds(dto.TimeoutMs ?? …)</c>); when the config
/// omits <c>timeoutMs</c> the caller's <paramref name="fallbackTimeout"/> is used.</summary>
/// omits <c>timeoutMs</c> the caller's <paramref name="fallbackTimeout"/> is used. <c>Transport</c>
/// mirrors the factory's null-defaults-to-Tcp convention (<see cref="ModbusDriverFactoryExtensions"/>);
/// an unrecognized value is reported as a probe-target error rather than silently falling back.</summary>
/// <param name="configJson">The driver config JSON (factory DTO shape).</param>
/// <param name="fallbackTimeout">The timeout used when the config omits <c>timeoutMs</c>.</param>
/// <returns>The parsed target, or a null target with an error string.</returns>
@@ -52,8 +56,15 @@ public sealed class ModbusDriverProbe : IDriverProbe
if (string.IsNullOrWhiteSpace(dto.Host) || port <= 0)
return (null, "Config has no host/port to probe.");
var transportMode = ModbusTransportMode.Tcp;
if (dto.Transport is not null && !Enum.TryParse(dto.Transport, ignoreCase: true, out transportMode))
{
return (null, $"Config has unknown Transport '{dto.Transport}'. " +
$"Expected one of {string.Join(", ", Enum.GetNames<ModbusTransportMode>())}");
}
var timeout = dto.TimeoutMs is { } ms && ms > 0 ? TimeSpan.FromMilliseconds(ms) : fallbackTimeout;
return (new ProbeTarget(dto.Host!, port, dto.UnitId ?? 1, timeout), null);
return (new ProbeTarget(dto.Host!, port, dto.UnitId ?? 1, timeout, transportMode), null);
}
/// <inheritdoc />
@@ -72,9 +83,17 @@ public sealed class ModbusDriverProbe : IDriverProbe
using var cts = CancellationTokenSource.CreateLinkedTokenSource(ct);
cts.CancelAfter(timeout);
// Phase 1 — TCP connect (using ModbusTcpTransport which handles IPv4 preference).
// Phase 1 — TCP connect, over whichever transport the config's Transport mode selects
// (ModbusTransportFactory is the single mapping point — same switch the live driver uses).
// autoReconnect=false: this is a one-shot probe, no retry loops.
var transport = new ModbusTcpTransport(host, port, timeout, autoReconnect: false);
var transport = ModbusTransportFactory.Create(new ModbusDriverOptions
{
Host = host,
Port = port,
Timeout = timeout,
Transport = t.Transport,
AutoReconnect = false,
});
await using (transport.ConfigureAwait(false))
{
try
@@ -0,0 +1,187 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Modbus;
/// <summary>
/// Modbus RTU framing: builds a request ADU (<c>[unitId] + PDU + CRC</c>) and deframes a
/// response back to its bare PDU. Byte-stream-agnostic — it operates on a <see cref="Stream"/>
/// and knows nothing about sockets or serial ports, so a future serial transport can reuse it.
/// </summary>
/// <remarks>
/// <para>
/// The single genuinely-new correctness risk of the RTU path: an RTU frame carries <b>no
/// length field</b> (unlike TCP's MBAP header), so the response size must be derived from
/// the function code. After reading <c>addr(1)</c> + <c>fc(1)</c>, the remaining byte count
/// is one of three shapes:
/// </para>
/// <list type="table">
/// <listheader><term>Shape</term><description>Trailing bytes after <c>addr,fc</c></description></listheader>
/// <item>
/// <term>Exception (<c>fc &amp; 0x80</c>)</term>
/// <description><c>excCode(1)</c> + <c>CRC(2)</c></description>
/// </item>
/// <item>
/// <term>Read (FC 01/02/03/04)</term>
/// <description><c>byteCount(1)</c> then <c>byteCount</c> data bytes + <c>CRC(2)</c></description>
/// </item>
/// <item>
/// <term>Write echo (FC 05/06/15/16)</term>
/// <description>fixed <c>4</c> bytes + <c>CRC(2)</c></description>
/// </item>
/// </list>
/// <para>
/// The trailing CRC is validated over <c>[addr, ...pdu]</c> (everything except the two CRC
/// bytes) <b>before</b> the frame is interpreted — so a corrupt exception frame surfaces as a
/// desync, never as a bogus <see cref="ModbusException"/>. A CRC mismatch or a short read
/// (truncation / stream closed mid-frame) throws <see cref="ModbusTransportDesyncException"/>;
/// a CRC-valid exception PDU throws <see cref="ModbusException"/> carrying the original
/// function code (<c>fc &amp; 0x7F</c>) and exception code.
/// </para>
/// <para>
/// Once the CRC passes, the response's address byte is validated against the addressed unit.
/// On an RS-485 multi-drop bus a CRC-valid reply from the <b>wrong</b> slave would otherwise be
/// silently accepted as the addressed unit's data; such a frame is a unit-mismatch
/// <see cref="ModbusTransportDesyncException"/>. This ordering is deliberate — a corrupt frame
/// stays a CRC/desync failure, and only a coherent-but-mis-addressed frame becomes a
/// unit-mismatch desync.
/// </para>
/// </remarks>
public static class ModbusRtuFraming
{
/// <summary>
/// Builds an RTU request ADU: the unit id, the PDU, and the trailing CRC-16/MODBUS
/// (appended low byte first).
/// </summary>
/// <param name="unitId">The RTU slave/unit address.</param>
/// <param name="pdu">The bare PDU (function code + data).</param>
/// <returns>A new array: <c>[unitId, ...pdu, crcLo, crcHi]</c>.</returns>
public static byte[] BuildAdu(byte unitId, ReadOnlySpan<byte> pdu)
{
var frame = new byte[1 + pdu.Length];
frame[0] = unitId;
pdu.CopyTo(frame.AsSpan(1));
return ModbusCrc.Append(frame);
}
/// <summary>
/// Reads a single RTU response frame from <paramref name="stream"/> and returns its bare PDU
/// (<c>[fc, ...data]</c>), with the leading unit-id and trailing CRC stripped.
/// </summary>
/// <param name="stream">The byte stream to read the response from.</param>
/// <param name="expectedUnit">
/// The unit id the request was addressed to. The response's address byte is validated against
/// it after the CRC passes; a mismatch is a unit-mismatch desync.
/// </param>
/// <param name="ct">A token to cancel the read.</param>
/// <returns>The bare response PDU (function code followed by its data).</returns>
/// <exception cref="ModbusTransportDesyncException">
/// The frame was truncated (stream closed mid-frame), its trailing CRC did not validate, or its
/// address byte did not match <paramref name="expectedUnit"/>.
/// </exception>
/// <exception cref="ModbusException">
/// The frame was a CRC-valid Modbus exception PDU (high bit set on the function code).
/// </exception>
public static async Task<byte[]> ReadResponsePduAsync(
Stream stream, byte expectedUnit, CancellationToken ct)
{
// Header: addr(1) + fc(1). The function code selects how many more bytes to read.
var header = new byte[2];
await ReadExactlyAsync(stream, header, ct).ConfigureAwait(false);
var fc = header[1];
int trailing; // bytes after addr+fc, up to and including the 2 CRC bytes.
if ((fc & 0x80) != 0)
{
// Exception frame: excCode(1) + CRC(2).
trailing = 1 + 2;
}
else if (fc is 0x01 or 0x02 or 0x03 or 0x04)
{
// Read response: byteCount(1) then byteCount data bytes + CRC(2).
var bc = new byte[1];
await ReadExactlyAsync(stream, bc, ct).ConfigureAwait(false);
var byteCount = bc[0];
var rest = new byte[byteCount + 2];
await ReadExactlyAsync(stream, rest, ct).ConfigureAwait(false);
return ValidateAndStrip(expectedUnit, header, bc, rest);
}
else
{
// Fixed 4-byte echo: correct for the write FCs this driver emits (05/06/0F/10). A future
// variable-length response FC outside 01-04 (e.g. FC23) would be mis-sized here and
// surface as a desync — revisit the FC-shape table if the driver starts emitting one.
trailing = 4 + 2;
}
var tail = new byte[trailing];
await ReadExactlyAsync(stream, tail, ct).ConfigureAwait(false);
return ValidateAndStrip(expectedUnit, header, ReadOnlySpan<byte>.Empty, tail);
}
/// <summary>
/// Assembles the full frame from its pieces, validates the trailing CRC over everything
/// except the CRC bytes, then validates the response address against
/// <paramref name="expectedUnit"/> and strips <c>addr</c> + <c>CRC</c> to return the bare PDU.
/// A CRC-valid exception PDU throws <see cref="ModbusException"/>.
/// </summary>
private static byte[] ValidateAndStrip(
byte expectedUnit, ReadOnlySpan<byte> header, ReadOnlySpan<byte> middle, ReadOnlySpan<byte> tail)
{
// Reassemble the whole frame: header(addr,fc) + middle(optional byteCount) + tail.
var frame = new byte[header.Length + middle.Length + tail.Length];
header.CopyTo(frame);
middle.CopyTo(frame.AsSpan(header.Length));
tail.CopyTo(frame.AsSpan(header.Length + middle.Length));
// CRC covers everything except the trailing 2 CRC bytes.
var body = frame.AsSpan(0, frame.Length - 2);
var expected = ModbusCrc.Compute(body);
var actual = (ushort)(frame[^2] | (frame[^1] << 8));
if (actual != expected)
throw new ModbusTransportDesyncException(
ModbusTransportDesyncException.DesyncReason.CrcMismatch,
$"Modbus RTU CRC mismatch: computed {expected:X4} got {actual:X4}");
// Unit-address validation runs only AFTER the CRC passes: a corrupt frame is a CRC/desync
// failure, and only a coherent-but-mis-addressed frame (a CRC-valid reply from the wrong
// RS-485 slave) is a unit-mismatch desync. Never silently accept another unit's data.
var addr = frame[0];
if (addr != expectedUnit)
throw new ModbusTransportDesyncException(
ModbusTransportDesyncException.DesyncReason.UnitMismatch,
$"Modbus RTU unit mismatch: expected {expectedUnit} got {addr}");
// Bare PDU = frame minus leading addr and trailing CRC.
var pdu = body[1..].ToArray();
// Exception PDU: high bit set on the function code. The CRC already validated, so this is a
// coherent protocol-level error — surface the ORIGINAL function code (fc & 0x7F).
if ((pdu[0] & 0x80) != 0)
{
var fc = (byte)(pdu[0] & 0x7F);
var exCode = pdu[1];
throw new ModbusException(fc, exCode, $"Modbus exception fc={fc} code={exCode}");
}
return pdu;
}
/// <summary>
/// Fills <paramref name="buf"/> completely from <paramref name="stream"/>, throwing a
/// truncation desync if the stream ends first. Mirrors
/// <c>ModbusTcpTransport.ReadExactlyAsync</c>, but normalises the end-of-stream to a
/// <see cref="ModbusTransportDesyncException"/> so a length-less RTU frame that arrives
/// short is classified as a desync rather than a bare <see cref="EndOfStreamException"/>.
/// </summary>
private static async Task ReadExactlyAsync(Stream stream, byte[] buf, CancellationToken ct)
{
var read = 0;
while (read < buf.Length)
{
var n = await stream.ReadAsync(buf.AsMemory(read), ct).ConfigureAwait(false);
if (n == 0)
throw new ModbusTransportDesyncException(
ModbusTransportDesyncException.DesyncReason.TruncatedFrame,
$"Modbus RTU frame truncated: expected {buf.Length} bytes, got {read}");
read += n;
}
}
}
@@ -0,0 +1,190 @@
using System.Net.Sockets;
namespace ZB.MOM.WW.OtOpcUa.Driver.Modbus;
/// <summary>
/// Concrete Modbus RTU-over-TCP transport. Composes <see cref="ModbusSocketLifecycle"/> for the
/// connect / reconnect / keep-alive / idle machinery and <see cref="ModbusRtuFraming"/> for the
/// wire layout — the same lifecycle <see cref="ModbusTcpTransport"/> uses, but with CRC framing
/// instead of MBAP.
/// </summary>
/// <remarks>
/// <para>
/// Delta from <see cref="ModbusTcpTransport"/>: an RTU ADU is <c>[unitId][PDU][CRC]</c> with
/// <b>no MBAP header and no transaction id</b>. Because there is no TxId to correlate
/// interleaved responses, at most one transaction may be on the bus at a time — the
/// <see cref="_gate"/> single-flight is therefore <b>mandatory</b>, not merely a
/// diagnostics convenience.
/// </para>
/// <para>
/// Survives mid-transaction socket drops the same way the TCP transport does: a socket-level
/// failure (<see cref="IOException"/> / <see cref="SocketException"/> /
/// <see cref="EndOfStreamException"/>, and the <see cref="ModbusTransportDesyncException"/>
/// that derives from <see cref="IOException"/>) triggers a single reconnect-then-retry; a
/// second failure bubbles up so the driver's health surface reflects the real state.
/// </para>
/// <para>
/// Every transaction runs under a per-op deadline (a linked
/// <see cref="CancellationTokenSource.CancelAfter(TimeSpan)"/>, design §7 / R2-01) so a
/// frozen gateway can never wedge a poll: the deadline firing is normalised to a
/// <see cref="ModbusTransportDesyncException"/> and tears the socket down so the next attempt
/// reconnects. A <b>caller</b> cancellation is distinguished from that timeout and propagates
/// as a plain <see cref="OperationCanceledException"/> with no teardown.
/// </para>
/// <para>
/// The constructor is connection-free; all socket I/O happens in <see cref="ConnectAsync"/> /
/// <see cref="SendAsync"/>. The <see cref="ForTest"/> seam injects a pre-connected
/// <see cref="Stream"/> (an in-memory duplex fake) so the send/receive orchestration,
/// single-flight, and deadline can be exercised with no socket — in that path
/// <see cref="_life"/> is <see langword="null"/> and the idle / reconnect machinery is
/// skipped (a raw injected stream cannot reconnect).
/// </para>
/// </remarks>
public sealed class ModbusRtuOverTcpTransport : IModbusTransport
{
private readonly ModbusSocketLifecycle? _life;
private readonly Stream? _testStream;
private readonly TimeSpan _timeout;
private readonly SemaphoreSlim _gate = new(1, 1);
private bool _disposed;
/// <summary>Initializes a new instance of the <see cref="ModbusRtuOverTcpTransport"/> class.</summary>
/// <param name="host">The host address or hostname of the Modbus gateway.</param>
/// <param name="port">The TCP port of the Modbus gateway.</param>
/// <param name="timeout">The timeout for socket operations and the per-op response deadline.</param>
/// <param name="autoReconnect">Whether to automatically reconnect on socket failures.</param>
/// <param name="keepAlive">Optional keep-alive configuration for the socket.</param>
/// <param name="idleDisconnect">Optional duration after which an idle socket is disconnected.</param>
/// <param name="reconnect">Optional reconnect backoff configuration.</param>
public ModbusRtuOverTcpTransport(
string host, int port, TimeSpan timeout, bool autoReconnect = true,
ModbusKeepAliveOptions? keepAlive = null,
TimeSpan? idleDisconnect = null,
ModbusReconnectOptions? reconnect = null)
{
_timeout = timeout;
_life = new ModbusSocketLifecycle(host, port, timeout, autoReconnect, keepAlive, idleDisconnect, reconnect);
}
private ModbusRtuOverTcpTransport(Stream testStream, TimeSpan timeout)
{
_timeout = timeout;
_testStream = testStream;
}
/// <summary>
/// Test seam: builds a transport over a pre-connected, in-memory duplex <paramref name="stream"/>,
/// bypassing <see cref="ConnectAsync"/> and the idle / reconnect machinery so the send/receive
/// orchestration, single-flight, and per-op deadline can be exercised with no real socket.
/// </summary>
/// <param name="stream">A pre-connected duplex stream standing in for the socket.</param>
/// <param name="timeout">The per-op response deadline.</param>
/// <returns>A transport driving <paramref name="stream"/> directly.</returns>
internal static ModbusRtuOverTcpTransport ForTest(Stream stream, TimeSpan timeout) => new(stream, timeout);
/// <inheritdoc />
public Task ConnectAsync(CancellationToken ct) =>
// The ForTest seam is handed a pre-connected stream — nothing to dial.
_life is null ? Task.CompletedTask : _life.ConnectAsync(ct);
/// <inheritdoc />
public async Task<byte[]> SendAsync(byte unitId, byte[] pdu, CancellationToken ct)
{
if (_disposed) throw new ObjectDisposedException(nameof(ModbusRtuOverTcpTransport));
if (CurrentStream is null) throw new InvalidOperationException("Transport not connected");
// Single-flight: RTU carries no transaction id, so at most one transaction may be on the bus
// at a time — serialise every send/receive under the gate.
await _gate.WaitAsync(ct).ConfigureAwait(false);
try
{
// Proactive idle-disconnect (real socket only): if the socket has been quiet longer than
// the configured threshold, tear it down + reconnect before this PDU lands. Defends
// against silent NAT / firewall reaping.
if (_life is not null && _life.ShouldReconnectForIdle())
{
await _life.TearDownAsync().ConfigureAwait(false);
await _life.ConnectWithBackoffAsync(ct).ConfigureAwait(false);
}
try
{
var result = await SendOnceAsync(unitId, pdu, ct).ConfigureAwait(false);
_life?.MarkSuccess();
return result;
}
catch (Exception ex) when (_life is not null && _life.AutoReconnect && ModbusSocketLifecycle.IsSocketLevelFailure(ex))
{
// Mid-transaction drop: tear down the dead socket, reconnect (with backoff if
// configured), resend. Single retry — a second failure propagates so health/status
// reflect reality. Never reached in the ForTest seam (a raw injected stream has no
// lifecycle to reconnect).
await _life.TearDownAsync().ConfigureAwait(false);
await _life.ConnectWithBackoffAsync(ct).ConfigureAwait(false);
var result = await SendOnceAsync(unitId, pdu, ct).ConfigureAwait(false);
_life.MarkSuccess();
return result;
}
}
finally
{
_gate.Release();
}
}
/// <summary>The live stream: the lifecycle's <see cref="NetworkStream"/>, or the injected test stream.</summary>
private Stream? CurrentStream => _life?.Stream ?? _testStream;
/// <summary>
/// Executes exactly one RTU transaction on the current stream: build the ADU, write + flush,
/// read back one deframed response PDU. See the class remarks for the desync / timeout /
/// caller-cancel handling.
/// </summary>
private async Task<byte[]> SendOnceAsync(byte unitId, byte[] pdu, CancellationToken ct)
{
var stream = CurrentStream;
if (stream is null) throw new InvalidOperationException("Transport not connected");
// RTU ADU: [unitId] + PDU + CRC — no MBAP header, no TxId.
var adu = ModbusRtuFraming.BuildAdu(unitId, pdu);
using var cts = CancellationTokenSource.CreateLinkedTokenSource(ct);
cts.CancelAfter(_timeout);
try
{
await stream.WriteAsync(adu.AsMemory(), cts.Token).ConfigureAwait(false);
await stream.FlushAsync(cts.Token).ConfigureAwait(false);
// Framing owns the length-less RTU deframe + CRC + unit-address validation. A CRC-valid
// exception PDU throws ModbusException (socket still coherent — propagates, no teardown);
// truncation / CRC / unit-mismatch throw ModbusTransportDesyncException.
return await ModbusRtuFraming.ReadResponsePduAsync(stream, unitId, cts.Token).ConfigureAwait(false);
}
catch (ModbusTransportDesyncException)
{
// Framing violation: the stream is desynchronized — never reuse it.
if (_life is not null) await _life.TearDownAsync().ConfigureAwait(false);
throw;
}
catch (OperationCanceledException) when (!ct.IsCancellationRequested)
{
// Per-op timeout fired (NOT the caller). The response is unknown / may still land later
// and corrupt the next read — tear the socket down and normalise as a desync so the
// reconnect retry / status mapping engages.
if (_life is not null) await _life.TearDownAsync().ConfigureAwait(false);
throw new ModbusTransportDesyncException(
ModbusTransportDesyncException.DesyncReason.Timeout,
$"Modbus per-op response timeout after {_timeout.TotalMilliseconds:0}ms");
}
}
/// <summary>Asynchronously disposes the transport and underlying socket resources.</summary>
/// <returns>A task that represents the asynchronous operation.</returns>
public async ValueTask DisposeAsync()
{
if (_disposed) return;
_disposed = true;
if (_life is not null) await _life.DisposeAsync().ConfigureAwait(false);
_gate.Dispose();
}
}
@@ -0,0 +1,221 @@
using System.Net.Sockets;
namespace ZB.MOM.WW.OtOpcUa.Driver.Modbus;
/// <summary>
/// Owns the raw TCP socket lifecycle shared by every Modbus-over-TCP transport (MBAP and
/// RTU-over-TCP alike): the IPv4-preference connect, OS-level keep-alive, geometric-backoff
/// reconnect, idle-disconnect tracking, and teardown. It exposes the current
/// <see cref="NetworkStream"/> so the composing transport can drive its own framing on top —
/// the lifecycle knows nothing about MBAP / RTU wire layout.
/// </summary>
/// <remarks>
/// <para>
/// Extracted verbatim from <see cref="ModbusTcpTransport"/> so the connect / reconnect /
/// keep-alive / idle behaviour is written once and composed by both transports. The
/// constructor is connection-free — all socket I/O happens in <see cref="ConnectAsync"/> /
/// <see cref="ConnectWithBackoffAsync"/>.
/// </para>
/// <para>
/// Why keep-alive matters for DL205/DL260: the AutomationDirect H2-ECOM100 does NOT send
/// TCP keepalives per <c>docs/v2/dl205.md</c> §behavioral-oddities, so any NAT/firewall
/// between the gateway and PLC can silently close an idle socket after 2-5 minutes.
/// Enabling OS-level <c>SO_KEEPALIVE</c> lets the driver's own side detect a stuck socket
/// in reasonable time even when the application is mostly idle.
/// </para>
/// </remarks>
public sealed class ModbusSocketLifecycle : IAsyncDisposable
{
private readonly string _host;
private readonly int _port;
private readonly TimeSpan _timeout;
private readonly bool _autoReconnect;
private readonly ModbusKeepAliveOptions _keepAlive;
private readonly TimeSpan? _idleDisconnect;
private readonly ModbusReconnectOptions _reconnect;
private TcpClient? _client;
private NetworkStream? _stream;
private DateTime _lastSuccessUtc = DateTime.UtcNow;
/// <summary>Initializes a new instance of the <see cref="ModbusSocketLifecycle"/> class.</summary>
/// <param name="host">The host address or hostname of the Modbus server.</param>
/// <param name="port">The TCP port of the Modbus server.</param>
/// <param name="timeout">The timeout for socket operations.</param>
/// <param name="autoReconnect">Whether to automatically reconnect on socket failures.</param>
/// <param name="keepAlive">Optional keep-alive configuration for the socket.</param>
/// <param name="idleDisconnect">Optional duration after which an idle socket is disconnected.</param>
/// <param name="reconnect">Optional reconnect backoff configuration.</param>
public ModbusSocketLifecycle(
string host, int port, TimeSpan timeout, bool autoReconnect = true,
ModbusKeepAliveOptions? keepAlive = null,
TimeSpan? idleDisconnect = null,
ModbusReconnectOptions? reconnect = null)
{
_host = host;
_port = port;
_timeout = timeout;
_autoReconnect = autoReconnect;
_keepAlive = keepAlive ?? new ModbusKeepAliveOptions();
_idleDisconnect = idleDisconnect;
_reconnect = reconnect ?? new ModbusReconnectOptions();
}
/// <summary>Gets the current live <see cref="NetworkStream"/>, or <see langword="null"/> when not connected.</summary>
public NetworkStream? Stream => _stream;
/// <summary>Gets a value indicating whether auto-reconnect on socket failures is enabled.</summary>
public bool AutoReconnect => _autoReconnect;
/// <summary>Establishes a connection to the Modbus server, preferring IPv4.</summary>
/// <param name="ct">A cancellation token to observe for cancellation.</param>
/// <returns>A task representing the asynchronous connection operation.</returns>
public async Task ConnectAsync(CancellationToken ct)
{
// Resolve the host explicitly + prefer IPv4. .NET's TcpClient default-constructor is
// dual-stack (IPv6 first, fallback to IPv4) — but most Modbus TCP devices (PLCs and
// simulators like pymodbus) bind 0.0.0.0 only, so the IPv6 attempt times out and we
// burn the entire ConnectAsync budget before even trying IPv4. Resolving first +
// dialing the IPv4 address directly sidesteps that.
var addresses = await System.Net.Dns.GetHostAddressesAsync(_host, ct).ConfigureAwait(false);
var ipv4 = System.Linq.Enumerable.FirstOrDefault(addresses,
a => a.AddressFamily == System.Net.Sockets.AddressFamily.InterNetwork);
var target = ipv4 ?? (addresses.Length > 0 ? addresses[0] : System.Net.IPAddress.Loopback);
_client = new TcpClient(target.AddressFamily);
EnableKeepAlive(_client, _keepAlive);
using var cts = CancellationTokenSource.CreateLinkedTokenSource(ct);
cts.CancelAfter(_timeout);
await _client.ConnectAsync(target, _port, cts.Token).ConfigureAwait(false);
_stream = _client.GetStream();
_lastSuccessUtc = DateTime.UtcNow;
}
/// <summary>
/// Connect attempt with the configured geometric backoff. The first attempt fires after
/// <see cref="ModbusReconnectOptions.InitialDelay"/> (default zero — immediate); each
/// subsequent attempt sleeps for the previous delay times <c>BackoffMultiplier</c>,
/// capped at <c>MaxDelay</c>. Caller's cancellation token aborts the loop.
/// </summary>
/// <param name="ct">A cancellation token to observe for cancellation.</param>
/// <returns>A task representing the asynchronous reconnect operation.</returns>
public async Task ConnectWithBackoffAsync(CancellationToken ct)
{
var delay = _reconnect.InitialDelay;
var attempt = 0;
while (true)
{
if (delay > TimeSpan.Zero)
await Task.Delay(delay, ct).ConfigureAwait(false);
try
{
await ConnectAsync(ct).ConfigureAwait(false);
return;
}
catch (Exception ex) when (IsSocketLevelFailure(ex) && _autoReconnect)
{
attempt++;
// Geometric growth, capped. Use Math.Min on ticks so we don't overflow with
// pathological multipliers / long deployments.
var nextTicks = (long)(Math.Max(delay.Ticks, TimeSpan.FromMilliseconds(100).Ticks) * _reconnect.BackoffMultiplier);
delay = TimeSpan.FromTicks(Math.Min(nextTicks, _reconnect.MaxDelay.Ticks));
if (attempt >= 10)
{
// Bail after 10 attempts to surface persistent failure to the caller. With
// the default backoff (1s base, 2.0x mult, 30s cap) this is roughly 4 minutes
// of attempts; with InitialDelay=0 it's immediate up to the same cap.
throw;
}
}
}
}
/// <summary>
/// Reports whether the socket has been idle longer than the configured idle-disconnect
/// threshold and should be proactively torn down + reconnected before the next transaction.
/// Always <see langword="false"/> when no idle-disconnect timeout is configured.
/// </summary>
/// <returns><see langword="true"/> when the idle threshold has been exceeded.</returns>
public bool ShouldReconnectForIdle() =>
_idleDisconnect.HasValue && DateTime.UtcNow - _lastSuccessUtc > _idleDisconnect.Value;
/// <summary>Records that a transaction just succeeded, resetting the idle-disconnect clock.</summary>
public void MarkSuccess() => _lastSuccessUtc = DateTime.UtcNow;
/// <summary>Tears down the current socket + stream, leaving the lifecycle ready to reconnect.</summary>
/// <returns>A task representing the asynchronous teardown operation.</returns>
public async Task TearDownAsync()
{
try { if (_stream is not null) await _stream.DisposeAsync().ConfigureAwait(false); }
catch { /* best-effort */ }
_stream = null;
try { _client?.Dispose(); } catch { }
_client = null;
}
/// <summary>
/// Enable SO_KEEPALIVE with aggressive probe timing. DL205/DL260 doesn't send keepalives
/// itself; having the OS probe the socket every ~30s lets the driver notice a dead PLC
/// or broken NAT path long before the default 2-hour Windows idle timeout fires.
/// Non-fatal if the underlying OS rejects the option (some older Linux / container
/// sandboxes don't expose the fine-grained timing levers — the driver still works,
/// application-level probe still detects problems).
/// </summary>
private static void EnableKeepAlive(TcpClient client, ModbusKeepAliveOptions opts)
{
if (!opts.Enabled) return;
try
{
client.Client.SetSocketOption(SocketOptionLevel.Socket, SocketOptionName.KeepAlive, true);
// A TimeSpan < 1s previously truncated to 0 via the int cast,
// which Windows / Linux interpret as "use the default" — silently defeating the
// configured keep-alive timing. Round up to at least 1 second so a sub-second
// configuration still produces a real keep-alive cadence. Negative values are
// also clamped to 1 to avoid surfacing as OS errors.
client.Client.SetSocketOption(SocketOptionLevel.Tcp, SocketOptionName.TcpKeepAliveTime,
ClampToWholeSeconds(opts.Time));
client.Client.SetSocketOption(SocketOptionLevel.Tcp, SocketOptionName.TcpKeepAliveInterval,
ClampToWholeSeconds(opts.Interval));
client.Client.SetSocketOption(SocketOptionLevel.Tcp, SocketOptionName.TcpKeepAliveRetryCount, opts.RetryCount);
}
catch { /* best-effort; older OSes may not expose the granular knobs */ }
}
/// <summary>
/// Cast a <see cref="TimeSpan"/> to a whole number of seconds with a
/// minimum of 1 — protects callers from the int-cast truncation that turned 500 ms
/// keep-alive timing into "use the default" on most OSes.
/// </summary>
/// <param name="ts">The timespan to clamp to whole seconds.</param>
/// <returns>The clamped duration expressed as a whole number of seconds, never less than 1.</returns>
internal static int ClampToWholeSeconds(TimeSpan ts)
{
var seconds = (int)Math.Ceiling(ts.TotalSeconds);
return seconds < 1 ? 1 : seconds;
}
/// <summary>
/// Distinguish socket-layer failures (eligible for reconnect-and-retry) from
/// protocol-layer failures (must propagate — retrying the same PDU won't help if the
/// PLC just returned exception 02 Illegal Data Address).
/// </summary>
/// <param name="ex">The exception to classify.</param>
/// <returns><see langword="true"/> when the exception is a socket-level failure.</returns>
internal static bool IsSocketLevelFailure(Exception ex) =>
ex is EndOfStreamException
|| ex is IOException
|| ex is SocketException
|| ex is ObjectDisposedException;
/// <summary>Asynchronously disposes the underlying socket + stream resources.</summary>
/// <returns>A task that represents the asynchronous operation.</returns>
public async ValueTask DisposeAsync()
{
try
{
if (_stream is not null) await _stream.DisposeAsync().ConfigureAwait(false);
}
catch { /* best-effort */ }
_client?.Dispose();
}
}
@@ -3,13 +3,19 @@ using System.Net.Sockets;
namespace ZB.MOM.WW.OtOpcUa.Driver.Modbus;
/// <summary>
/// Concrete Modbus TCP transport. Wraps a single <see cref="TcpClient"/> and serializes
/// requests so at most one transaction is in-flight at a time — Modbus servers typically
/// support concurrent transactions, but the single-flight model keeps the wire trace
/// Concrete Modbus TCP transport. Wraps a single socket (owned by <see cref="ModbusSocketLifecycle"/>)
/// and serializes requests so at most one transaction is in-flight at a time — Modbus servers
/// typically support concurrent transactions, but the single-flight model keeps the wire trace
/// easy to diagnose and avoids interleaved-response correlation bugs.
/// </summary>
/// <remarks>
/// <para>
/// Owns the MBAP framing (<see cref="SendOnceAsync"/>: 7-byte header + transaction-id
/// pairing) and composes <see cref="ModbusSocketLifecycle"/> for the connect / reconnect /
/// keep-alive / idle-disconnect machinery — the same lifecycle the RTU-over-TCP transport
/// reuses with different framing.
/// </para>
/// <para>
/// Survives mid-transaction socket drops: when a send/read fails with a socket-level
/// error (<see cref="IOException"/>, <see cref="SocketException"/>, <see cref="EndOfStreamException"/>)
/// the transport disposes the dead socket, reconnects, and retries the PDU exactly
@@ -26,19 +32,11 @@ namespace ZB.MOM.WW.OtOpcUa.Driver.Modbus;
/// </remarks>
public sealed class ModbusTcpTransport : IModbusTransport
{
private readonly string _host;
private readonly int _port;
private readonly ModbusSocketLifecycle _life;
private readonly TimeSpan _timeout;
private readonly bool _autoReconnect;
private readonly ModbusKeepAliveOptions _keepAlive;
private readonly TimeSpan? _idleDisconnect;
private readonly ModbusReconnectOptions _reconnect;
private readonly SemaphoreSlim _gate = new(1, 1);
private TcpClient? _client;
private NetworkStream? _stream;
private ushort _nextTx;
private bool _disposed;
private DateTime _lastSuccessUtc = DateTime.UtcNow;
/// <summary>Initializes a new instance of the <see cref="ModbusTcpTransport"/> class.</summary>
/// <param name="host">The host address or hostname of the Modbus server.</param>
@@ -54,84 +52,18 @@ public sealed class ModbusTcpTransport : IModbusTransport
TimeSpan? idleDisconnect = null,
ModbusReconnectOptions? reconnect = null)
{
_host = host;
_port = port;
_timeout = timeout;
_autoReconnect = autoReconnect;
_keepAlive = keepAlive ?? new ModbusKeepAliveOptions();
_idleDisconnect = idleDisconnect;
_reconnect = reconnect ?? new ModbusReconnectOptions();
_life = new ModbusSocketLifecycle(host, port, timeout, autoReconnect, keepAlive, idleDisconnect, reconnect);
}
/// <inheritdoc />
public async Task ConnectAsync(CancellationToken ct)
{
// Resolve the host explicitly + prefer IPv4. .NET's TcpClient default-constructor is
// dual-stack (IPv6 first, fallback to IPv4) — but most Modbus TCP devices (PLCs and
// simulators like pymodbus) bind 0.0.0.0 only, so the IPv6 attempt times out and we
// burn the entire ConnectAsync budget before even trying IPv4. Resolving first +
// dialing the IPv4 address directly sidesteps that.
var addresses = await System.Net.Dns.GetHostAddressesAsync(_host, ct).ConfigureAwait(false);
var ipv4 = System.Linq.Enumerable.FirstOrDefault(addresses,
a => a.AddressFamily == System.Net.Sockets.AddressFamily.InterNetwork);
var target = ipv4 ?? (addresses.Length > 0 ? addresses[0] : System.Net.IPAddress.Loopback);
_client = new TcpClient(target.AddressFamily);
EnableKeepAlive(_client, _keepAlive);
using var cts = CancellationTokenSource.CreateLinkedTokenSource(ct);
cts.CancelAfter(_timeout);
await _client.ConnectAsync(target, _port, cts.Token).ConfigureAwait(false);
_stream = _client.GetStream();
_lastSuccessUtc = DateTime.UtcNow;
}
/// <summary>
/// Enable SO_KEEPALIVE with aggressive probe timing. DL205/DL260 doesn't send keepalives
/// itself; having the OS probe the socket every ~30s lets the driver notice a dead PLC
/// or broken NAT path long before the default 2-hour Windows idle timeout fires.
/// Non-fatal if the underlying OS rejects the option (some older Linux / container
/// sandboxes don't expose the fine-grained timing levers — the driver still works,
/// application-level probe still detects problems).
/// </summary>
private static void EnableKeepAlive(TcpClient client, ModbusKeepAliveOptions opts)
{
if (!opts.Enabled) return;
try
{
client.Client.SetSocketOption(SocketOptionLevel.Socket, SocketOptionName.KeepAlive, true);
// A TimeSpan < 1s previously truncated to 0 via the int cast,
// which Windows / Linux interpret as "use the default" — silently defeating the
// configured keep-alive timing. Round up to at least 1 second so a sub-second
// configuration still produces a real keep-alive cadence. Negative values are
// also clamped to 1 to avoid surfacing as OS errors.
client.Client.SetSocketOption(SocketOptionLevel.Tcp, SocketOptionName.TcpKeepAliveTime,
ClampToWholeSeconds(opts.Time));
client.Client.SetSocketOption(SocketOptionLevel.Tcp, SocketOptionName.TcpKeepAliveInterval,
ClampToWholeSeconds(opts.Interval));
client.Client.SetSocketOption(SocketOptionLevel.Tcp, SocketOptionName.TcpKeepAliveRetryCount, opts.RetryCount);
}
catch { /* best-effort; older OSes may not expose the granular knobs */ }
}
/// <summary>
/// Cast a <see cref="TimeSpan"/> to a whole number of seconds with a
/// minimum of 1 — protects callers from the int-cast truncation that turned 500 ms
/// keep-alive timing into "use the default" on most OSes.
/// </summary>
/// <param name="ts">The timespan to clamp to whole seconds.</param>
/// <returns>The clamped duration expressed as a whole number of seconds, never less than 1.</returns>
internal static int ClampToWholeSeconds(TimeSpan ts)
{
var seconds = (int)Math.Ceiling(ts.TotalSeconds);
return seconds < 1 ? 1 : seconds;
}
public Task ConnectAsync(CancellationToken ct) => _life.ConnectAsync(ct);
/// <inheritdoc />
public async Task<byte[]> SendAsync(byte unitId, byte[] pdu, CancellationToken ct)
{
if (_disposed) throw new ObjectDisposedException(nameof(ModbusTcpTransport));
if (_stream is null) throw new InvalidOperationException("Transport not connected");
if (_life.Stream is null) throw new InvalidOperationException("Transport not connected");
await _gate.WaitAsync(ct).ConfigureAwait(false);
try
@@ -140,27 +72,27 @@ public sealed class ModbusTcpTransport : IModbusTransport
// threshold, tear it down + reconnect before this PDU lands. Defends against silent
// NAT / firewall reaping where the socket looks alive locally but the upstream side
// dropped it minutes ago.
if (_idleDisconnect.HasValue && DateTime.UtcNow - _lastSuccessUtc > _idleDisconnect.Value)
if (_life.ShouldReconnectForIdle())
{
await TearDownAsync().ConfigureAwait(false);
await ConnectWithBackoffAsync(ct).ConfigureAwait(false);
await _life.TearDownAsync().ConfigureAwait(false);
await _life.ConnectWithBackoffAsync(ct).ConfigureAwait(false);
}
try
{
var result = await SendOnceAsync(unitId, pdu, ct).ConfigureAwait(false);
_lastSuccessUtc = DateTime.UtcNow;
_life.MarkSuccess();
return result;
}
catch (Exception ex) when (_autoReconnect && IsSocketLevelFailure(ex))
catch (Exception ex) when (_life.AutoReconnect && ModbusSocketLifecycle.IsSocketLevelFailure(ex))
{
// Mid-transaction drop: tear down the dead socket, reconnect (with backoff if
// configured), resend. Single retry — if it fails again, let it propagate so
// health/status reflect reality.
await TearDownAsync().ConfigureAwait(false);
await ConnectWithBackoffAsync(ct).ConfigureAwait(false);
await _life.TearDownAsync().ConfigureAwait(false);
await _life.ConnectWithBackoffAsync(ct).ConfigureAwait(false);
var result = await SendOnceAsync(unitId, pdu, ct).ConfigureAwait(false);
_lastSuccessUtc = DateTime.UtcNow;
_life.MarkSuccess();
return result;
}
}
@@ -170,43 +102,6 @@ public sealed class ModbusTcpTransport : IModbusTransport
}
}
/// <summary>
/// Connect attempt with the configured geometric backoff. The first attempt fires after
/// <see cref="ModbusReconnectOptions.InitialDelay"/> (default zero — immediate); each
/// subsequent attempt sleeps for the previous delay times <c>BackoffMultiplier</c>,
/// capped at <c>MaxDelay</c>. Caller's cancellation token aborts the loop.
/// </summary>
private async Task ConnectWithBackoffAsync(CancellationToken ct)
{
var delay = _reconnect.InitialDelay;
var attempt = 0;
while (true)
{
if (delay > TimeSpan.Zero)
await Task.Delay(delay, ct).ConfigureAwait(false);
try
{
await ConnectAsync(ct).ConfigureAwait(false);
return;
}
catch (Exception ex) when (IsSocketLevelFailure(ex) && _autoReconnect)
{
attempt++;
// Geometric growth, capped. Use Math.Min on ticks so we don't overflow with
// pathological multipliers / long deployments.
var nextTicks = (long)(Math.Max(delay.Ticks, TimeSpan.FromMilliseconds(100).Ticks) * _reconnect.BackoffMultiplier);
delay = TimeSpan.FromTicks(Math.Min(nextTicks, _reconnect.MaxDelay.Ticks));
if (attempt >= 10)
{
// Bail after 10 attempts to surface persistent failure to the caller. With
// the default backoff (1s base, 2.0x mult, 30s cap) this is roughly 4 minutes
// of attempts; with InitialDelay=0 it's immediate up to the same cap.
throw;
}
}
}
}
/// <summary>
/// Executes exactly one Modbus transaction on the current socket.
/// </summary>
@@ -230,7 +125,8 @@ public sealed class ModbusTcpTransport : IModbusTransport
/// </remarks>
private async Task<byte[]> SendOnceAsync(byte unitId, byte[] pdu, CancellationToken ct)
{
if (_stream is null) throw new InvalidOperationException("Transport not connected");
var stream = _life.Stream;
if (stream is null) throw new InvalidOperationException("Transport not connected");
var txId = ++_nextTx;
// MBAP: [TxId(2)][Proto=0(2)][Length(2)][UnitId(1)] + PDU
@@ -248,11 +144,11 @@ public sealed class ModbusTcpTransport : IModbusTransport
cts.CancelAfter(_timeout);
try
{
await _stream.WriteAsync(adu.AsMemory(), cts.Token).ConfigureAwait(false);
await _stream.FlushAsync(cts.Token).ConfigureAwait(false);
await stream.WriteAsync(adu.AsMemory(), cts.Token).ConfigureAwait(false);
await stream.FlushAsync(cts.Token).ConfigureAwait(false);
var header = new byte[7];
await ReadExactlyAsync(_stream, header, cts.Token).ConfigureAwait(false);
await ReadExactlyAsync(stream, header, cts.Token).ConfigureAwait(false);
var respTxId = (ushort)((header[0] << 8) | header[1]);
if (respTxId != txId)
throw new ModbusTransportDesyncException(
@@ -264,7 +160,7 @@ public sealed class ModbusTcpTransport : IModbusTransport
ModbusTransportDesyncException.DesyncReason.TruncatedFrame,
$"Modbus response length too small: {respLen}");
var respPdu = new byte[respLen - 1];
await ReadExactlyAsync(_stream, respPdu, cts.Token).ConfigureAwait(false);
await ReadExactlyAsync(stream, respPdu, cts.Token).ConfigureAwait(false);
// Exception PDU: function code has high bit set. This is a well-formed protocol-level
// error — the socket is still coherent, so it MUST propagate without teardown.
@@ -280,7 +176,7 @@ public sealed class ModbusTcpTransport : IModbusTransport
catch (ModbusTransportDesyncException)
{
// Framing violation: the socket is desynchronized — never reuse it.
await TearDownAsync().ConfigureAwait(false);
await _life.TearDownAsync().ConfigureAwait(false);
throw;
}
catch (OperationCanceledException) when (!ct.IsCancellationRequested)
@@ -288,33 +184,13 @@ public sealed class ModbusTcpTransport : IModbusTransport
// Per-op timeout fired (NOT the caller). The response is unknown / may still land later
// and corrupt the next read — tear the socket down and normalize as a desync so the
// reconnect retry / status mapping engages.
await TearDownAsync().ConfigureAwait(false);
await _life.TearDownAsync().ConfigureAwait(false);
throw new ModbusTransportDesyncException(
ModbusTransportDesyncException.DesyncReason.Timeout,
$"Modbus per-op response timeout after {_timeout.TotalMilliseconds:0}ms");
}
}
/// <summary>
/// Distinguish socket-layer failures (eligible for reconnect-and-retry) from
/// protocol-layer failures (must propagate — retrying the same PDU won't help if the
/// PLC just returned exception 02 Illegal Data Address).
/// </summary>
private static bool IsSocketLevelFailure(Exception ex) =>
ex is EndOfStreamException
|| ex is IOException
|| ex is SocketException
|| ex is ObjectDisposedException;
private async Task TearDownAsync()
{
try { if (_stream is not null) await _stream.DisposeAsync().ConfigureAwait(false); }
catch { /* best-effort */ }
_stream = null;
try { _client?.Dispose(); } catch { }
_client = null;
}
private static async Task ReadExactlyAsync(Stream s, byte[] buf, CancellationToken ct)
{
var read = 0;
@@ -332,12 +208,7 @@ public sealed class ModbusTcpTransport : IModbusTransport
{
if (_disposed) return;
_disposed = true;
try
{
if (_stream is not null) await _stream.DisposeAsync().ConfigureAwait(false);
}
catch { /* best-effort */ }
_client?.Dispose();
await _life.DisposeAsync().ConfigureAwait(false);
_gate.Dispose();
}
}
@@ -1,9 +1,10 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Modbus;
/// <summary>
/// Raised when a Modbus TCP transaction leaves the single-flight socket in an unknown /
/// Raised when a Modbus transaction leaves the single-flight socket in an unknown /
/// desynchronized state — a per-op response timeout, or a framing violation (transaction-id
/// mismatch or truncated MBAP length). Unlike a Modbus <em>exception PDU</em> (a well-formed
/// mismatch, truncated frame, RTU CRC mismatch, or RTU unit-address mismatch). Unlike a Modbus
/// <em>exception PDU</em> (a well-formed
/// protocol-level error where the socket is still coherent and the request must simply
/// propagate), a desync means bytes may still be in flight or half-read, so the socket is
/// unusable and MUST be torn down before this is thrown.
@@ -27,8 +28,20 @@ public sealed class ModbusTransportDesyncException : IOException
/// <summary>The response MBAP transaction id did not match the request's.</summary>
TxIdMismatch,
/// <summary>The response MBAP length field was below the mandatory minimum (truncated frame).</summary>
/// <summary>
/// The frame ended before it was complete — the MBAP length field was below the mandatory
/// minimum (TCP), or the stream closed mid-frame while reading a length-less RTU frame.
/// </summary>
TruncatedFrame,
/// <summary>The RTU frame's trailing CRC-16 did not validate over the received bytes.</summary>
CrcMismatch,
/// <summary>
/// A CRC-valid RTU frame carried a slave/unit address other than the one addressed — on an
/// RS-485 multi-drop bus, a reply from the wrong slave.
/// </summary>
UnitMismatch,
}
/// <summary>Gets the reason the socket was classified desynchronized.</summary>
@@ -0,0 +1,43 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Modbus;
/// <summary>
/// Single mapping point from <see cref="ModbusDriverOptions"/> to the concrete
/// <see cref="IModbusTransport"/> its <see cref="ModbusDriverOptions.Transport"/> selects.
/// Consumed by both <see cref="ModbusDriver"/>'s default transport-factory closure and the
/// AdminUI Test-Connect probe, so the <see cref="ModbusTransportMode"/> switch lives in
/// exactly one place.
/// </summary>
public static class ModbusTransportFactory
{
/// <summary>
/// Builds the transport <paramref name="options"/>.<see cref="ModbusDriverOptions.Transport"/>
/// selects, threading every shared connection setting (Host/Port/Timeout/AutoReconnect/
/// KeepAlive/IdleDisconnectTimeout/Reconnect) through to whichever concrete transport is built.
/// </summary>
/// <param name="options">Driver configuration options.</param>
/// <returns>A <see cref="ModbusRtuOverTcpTransport"/> when <see cref="ModbusTransportMode.RtuOverTcp"/>
/// is selected; a <see cref="ModbusTcpTransport"/> when <see cref="ModbusTransportMode.Tcp"/> is selected.</returns>
/// <exception cref="ArgumentOutOfRangeException">
/// <paramref name="options"/>.<see cref="ModbusDriverOptions.Transport"/> is not a recognized
/// <see cref="ModbusTransportMode"/> member — fail loudly rather than silently falling back to TCP/MBAP.
/// </exception>
public static IModbusTransport Create(ModbusDriverOptions options)
{
ArgumentNullException.ThrowIfNull(options);
return options.Transport switch
{
ModbusTransportMode.RtuOverTcp => new ModbusRtuOverTcpTransport(
options.Host, options.Port, options.Timeout, options.AutoReconnect,
keepAlive: options.KeepAlive,
idleDisconnect: options.IdleDisconnectTimeout,
reconnect: options.Reconnect),
ModbusTransportMode.Tcp => new ModbusTcpTransport(
options.Host, options.Port, options.Timeout, options.AutoReconnect,
keepAlive: options.KeepAlive,
idleDisconnect: options.IdleDisconnectTimeout,
reconnect: options.Reconnect),
_ => throw new ArgumentOutOfRangeException(
nameof(options), options.Transport, "Unknown Modbus transport mode."),
};
}
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,192 @@
using System.Buffers;
using System.Text.Json;
using Microsoft.Extensions.Logging;
using Microsoft.Extensions.Logging.Abstractions;
using MQTTnet;
using MQTTnet.Protocol;
using ZB.MOM.WW.OtOpcUa.Commons.Browsing;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Browser;
/// <summary>
/// Bespoke MQTT address-picker browser: opens a short-lived, <b>strictly passive</b> observation
/// window against the broker named by the form's JSON.
/// <para>
/// MQTT has no discovery protocol, so there is nothing to enumerate — the browser subscribes
/// to a wildcard (<c>#</c> / <c>{topicPrefix}/#</c> in plain mode, <c>spBv1.0/{groupId}/#</c>
/// in Sparkplug mode) and lets <see cref="MqttBrowseSession"/> accumulate whatever arrives:
/// a topic segment tree in plain mode, a Group → EdgeNode → Device → Metric tree decoded from
/// observed birth certificates in Sparkplug mode. Nothing on any browse path publishes — in
/// either mode: an operator opening a picker must not be able to inject a message into a
/// running plant. The session's operator-triggered <c>RequestRebirthAsync</c> is the sole
/// sanctioned publish, and it is never reached by <c>OpenAsync</c> or by any browse call.
/// </para>
/// <para>
/// <b>Layering note.</b> Unlike every other bespoke browser here, this project references the
/// runtime <c>.Driver</c> project (to reuse <c>MqttConnection.BuildClientOptions</c> rather
/// than fork the TLS / CA-pin path). That is a known, deliberate exception with a documented
/// cost and a documented clean fix — see the comment on that <c>ProjectReference</c> in
/// <c>ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Browser.csproj</c> before copying this project as a
/// template.
/// </para>
/// </summary>
public sealed class MqttDriverBrowser : IDriverBrowser
{
/// <summary>Floor on the open budget — a 1 s <c>ConnectTimeoutSeconds</c> would make browse unusable.</summary>
internal const int MinOpenBudgetSeconds = 5;
/// <summary>Ceiling on the open budget — the picker must never hang on an unreachable broker.</summary>
internal const int MaxOpenBudgetSeconds = 30;
/// <summary>Marks the transient browse identity in the broker's client-id/session logs.</summary>
internal const string BrowseClientIdPrefix = "-browse-";
private readonly ILogger<MqttDriverBrowser> _logger;
/// <summary>
/// Creates a browser. <b>Connection-free by contract</b> — the universal browser's
/// <c>CanBrowse</c> constructs a throwaway instance purely to ask which driver type it
/// handles, so the constructor must never touch the network.
/// </summary>
/// <param name="logger">Optional logger; defaults to <see cref="NullLogger{T}"/>. Never receives credentials.</param>
public MqttDriverBrowser(ILogger<MqttDriverBrowser>? logger = null) =>
_logger = logger ?? NullLogger<MqttDriverBrowser>.Instance;
/// <inheritdoc />
public string DriverType => DriverTypeNames.Mqtt;
/// <inheritdoc />
/// <remarks>
/// Connects with a browse-only client id, subscribes the wildcard at QoS 0, and hands back a
/// session that serves whatever the subscription observes. The whole open is bounded by
/// <see cref="OpenBudget"/>; nothing here publishes.
/// </remarks>
public async Task<IBrowseSession> OpenAsync(string configJson, CancellationToken cancellationToken)
{
// MqttJson.Options — the ONE shared instance across factory / probe / driver / browser
// (see its remarks). Parsing the same DriverConfig blob through a second, divergent
// JsonSerializerOptions is this repo's documented systemic enum bug: the picker would accept
// a `mode` / `protocolVersion` spelling the deployed driver rejects, or vice versa.
var opts = JsonSerializer.Deserialize<MqttDriverOptions>(configJson, MqttJson.Options)
?? throw new InvalidOperationException("Mqtt options deserialized to null.");
var filter = BuildRootFilter(opts);
var suffix = BuildBrowseClientIdSuffix();
var clientOptions = MqttConnection.BuildClientOptions(ToBrowseOptions(opts), suffix, _logger);
using var openCts = CancellationTokenSource.CreateLinkedTokenSource(cancellationToken);
openCts.CancelAfter(OpenBudget(opts));
var client = new MqttClientFactory().CreateMqttClient();
var session = new MqttBrowseSession(opts.Mode, client, _logger);
try
{
// Runs on MQTTnet's dispatcher: record and return, nothing else.
client.ApplicationMessageReceivedAsync += args =>
{
var payload = args.ApplicationMessage.Payload;
ReadOnlySpan<byte> body = payload.IsSingleSegment ? payload.FirstSpan : payload.ToArray();
session.Observe(args.ApplicationMessage.Topic, body);
return Task.CompletedTask;
};
await client.ConnectAsync(clientOptions, openCts.Token).ConfigureAwait(false);
// QoS 0: a transient observation window has no delivery guarantee to offer and should
// cost the broker as little as possible.
var subscribeOptions = new MqttClientSubscribeOptionsBuilder()
.WithTopicFilter(f => f
.WithTopic(filter)
.WithQualityOfServiceLevel(MqttQualityOfServiceLevel.AtMostOnce))
.Build();
await client.SubscribeAsync(subscribeOptions, openCts.Token).ConfigureAwait(false);
_logger.LogInformation(
"AdminUI MQTT browse session opened against {Host}:{Port} in {Mode} mode observing "
+ "'{Filter}' (read-only).",
opts.Host, opts.Port, opts.Mode, filter);
return session;
}
catch
{
await session.DisposeAsync().ConfigureAwait(false); // owns the client — disconnects + disposes
throw;
}
}
/// <summary>
/// The wildcard the observation window subscribes to.
/// <para>
/// <b>Plain mode:</b> the whole broker (<c>#</c>) unless the config scopes it with a
/// topic prefix.
/// </para>
/// <para>
/// <b>Sparkplug mode:</b> the Sparkplug namespace only —
/// <c>spBv1.0/{groupId}/#</c>, matching the group the deployed driver itself subscribes
/// (design §3.1), or <c>spBv1.0/#</c> when no group has been authored yet, so the picker
/// can discover which groups exist. The Plain-mode topic prefix is deliberately ignored:
/// it is a different mode's key, and honouring it would silently narrow the window to a
/// prefix that cannot match a Sparkplug topic at all.
/// </para>
/// </summary>
/// <param name="options">The browse configuration.</param>
/// <returns>The MQTT topic filter to subscribe.</returns>
internal static string BuildRootFilter(MqttDriverOptions options)
{
if (options.Mode == MqttMode.SparkplugB)
{
var groupId = (options.Sparkplug?.GroupId ?? string.Empty).Trim();
return groupId.Length == 0 ? "spBv1.0/#" : $"spBv1.0/{groupId}/#";
}
var prefix = (options.Plain?.TopicPrefix ?? string.Empty).Trim();
if (prefix.Length == 0) return "#";
if (!prefix.EndsWith('/')) prefix += "/";
return prefix + "#";
}
/// <summary>
/// A unique per-session client-id suffix. A broker disconnects an existing client when a new
/// one CONNECTs with the <i>same</i> client id, so sharing the driver's id would knock the
/// live driver offline every time an operator opened the picker.
/// </summary>
/// <returns>The suffix appended to the configured client id.</returns>
internal static string BuildBrowseClientIdSuffix() =>
BrowseClientIdPrefix + Guid.NewGuid().ToString("N")[..8];
/// <summary>
/// Projects the authored options into the ones the browse CONNECT uses. Clean session is
/// forced: a transient browse identity must not leave queued-message state behind on the
/// broker after the picker closes.
/// <para>
/// ⚠ <b>THE landmine — a Last Will would break the read-only guarantee from outside it.</b>
/// <c>MqttDriverOptions</c> (and <c>MqttSparkplugOptions</c>) carry no will, <b>re-verified
/// when Task 23 unsealed Sparkplug browse</b>, so there is nothing to strip and the
/// projection stays a single <c>CleanSession</c> override. If a will is ever added — NDEATH
/// <i>is</i> a will message — it MUST be cleared here: a will is published by the
/// <i>broker</i>, not by us, so it is invisible to
/// <see cref="MqttBrowseSession.PublishCountForTest"/>, and any ungraceful end to a browse
/// session (crash, dropped link, the disconnect deadline expiring) would then fire an
/// NDEATH under the plant's own edge-node identity — killing a live Sparkplug node because
/// an operator opened an address picker. <c>MqttBrowseSessionTests</c>'
/// <c>MqttDriverOptions_CarryNothingWillShaped_…</c> is the reflection guard that turns that
/// "there is nothing to strip" into a checked fact rather than a claim.
/// </para>
/// </summary>
/// <param name="options">The authored configuration.</param>
/// <returns>The configuration the browse CONNECT is built from.</returns>
internal static MqttDriverOptions ToBrowseOptions(MqttDriverOptions options) =>
options with { CleanSession = true };
/// <summary>
/// Whole-open bound (connect + subscribe), clamped to
/// <see cref="MinOpenBudgetSeconds"/><see cref="MaxOpenBudgetSeconds"/> around the config's
/// own connect deadline.
/// </summary>
/// <param name="options">The browse configuration.</param>
/// <returns>The clamped open budget.</returns>
internal static TimeSpan OpenBudget(MqttDriverOptions options) =>
TimeSpan.FromSeconds(Math.Clamp(options.ConnectTimeoutSeconds, MinOpenBudgetSeconds, MaxOpenBudgetSeconds));
}
@@ -0,0 +1,56 @@
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup>
<TargetFramework>net10.0</TargetFramework>
<Nullable>enable</Nullable>
<ImplicitUsings>enable</ImplicitUsings>
<LangVersion>latest</LangVersion>
<TreatWarningsAsErrors>true</TreatWarningsAsErrors>
<RootNamespace>ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Browser</RootNamespace>
<AssemblyName>ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Browser</AssemblyName>
</PropertyGroup>
<ItemGroup>
<ProjectReference Include="..\..\Core\ZB.MOM.WW.OtOpcUa.Commons\ZB.MOM.WW.OtOpcUa.Commons.csproj"/>
<ProjectReference Include="..\ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Contracts\ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Contracts.csproj"/>
<!--
┌──────────────────────────────────────────────────────────────────────────────────┐
│ KNOWN, DELIBERATE EXCEPTION to the .Browser → .Contracts pattern. │
│ DO NOT COPY THIS LINE INTO A NEW DRIVER'S BROWSER WITHOUT READING THIS. │
└──────────────────────────────────────────────────────────────────────────────────┘
Every other bespoke browser in this repo (OpcUaClient, Galaxy) references only its own
.Contracts project plus its transport package. This one also references the runtime
.Driver project. That is a real widening, not a free one, and it was accepted knowingly:
WHY: the browse CONNECT reuses MqttConnection.BuildClientOptions, which owns the TLS
posture, the PEM CA-pin chain validator, the credential handling and the protocol
version mapping (~100 lines). A second copy here would be a silent security
divergence — browse quietly not honouring a pinned CA while the runtime driver does.
The method is pure (no I/O, no network), so the dependency costs nothing at run time.
WHAT IT COSTS: the AdminUI's project graph does NOT otherwise include Core.csproj
(only Host.csproj and the runtime Driver.* projects do). Referencing this browser
from the AdminUI therefore pulls in Core + Polly.Core + Serilog at build/deploy time.
THE CLEAN FIX (deferred, not blocked): a small leaf project — say
ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Transport (net10, refs .Contracts + MQTTnet) — holding
BuildClientOptions/ConfigureTls/the CA-pin helpers, referenced by BOTH .Driver and
.Browser. Note the obvious-looking alternative does NOT work: those members cannot
move into .Contracts, because .Contracts is deliberately transport-free (Task 1
removed its MQTTnet reference to keep it so) and BuildClientOptions returns
MqttClientOptions, an MQTTnet type.
-->
<ProjectReference Include="..\ZB.MOM.WW.OtOpcUa.Driver.Mqtt\ZB.MOM.WW.OtOpcUa.Driver.Mqtt.csproj"/>
</ItemGroup>
<ItemGroup>
<PackageReference Include="MQTTnet"/>
<PackageReference Include="Microsoft.Extensions.Logging.Abstractions"/>
</ItemGroup>
<ItemGroup>
<InternalsVisibleTo Include="ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Tests"/>
</ItemGroup>
</Project>
@@ -0,0 +1,225 @@
using System.ComponentModel.DataAnnotations;
using System.Text;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// MQTT / Sparkplug B driver configuration. Bound from <c>DriverConfig</c> JSON at
/// driver-host registration time. Models the settings documented in
/// <c>docs/plans/2026-07-15-mqtt-sparkplug-driver-design.md</c> §5.1.
/// </summary>
/// <remarks>
/// A record (not a plain class) so a future secret-resolution seam can produce a
/// credential-resolved copy with a <c>with</c> expression — mirrors
/// <c>OpcUaClientDriverOptions</c>. <see cref="Mode"/> selects the ingest shape;
/// <see cref="Sparkplug"/> and <see cref="Plain"/> are nullable sub-objects — only the one
/// matching <see cref="Mode"/> is populated.
/// </remarks>
public sealed record MqttDriverOptions
{
/// <summary>Broker hostname or IP address.</summary>
public string Host { get; init; } = "localhost";
/// <summary>Broker TCP port.</summary>
[Range(1, 65535)]
public int Port { get; init; } = 8883;
/// <summary>
/// MQTT client identifier sent at CONNECT. Leave unset to let the driver generate one
/// (a stable per-instance id is recommended so broker-side ACLs / session state persist
/// across reconnects).
/// </summary>
public string? ClientId { get; init; }
/// <summary>
/// When <c>true</c>, connect over TLS. Default <c>true</c> — this driver never ships an
/// anonymous/plaintext-by-default posture; a deployment must opt into <c>false</c> for an
/// on-prem/dev broker with no TLS listener.
/// </summary>
public bool UseTls { get; init; } = true;
/// <summary>
/// When <c>true</c>, accept any self-signed / untrusted broker certificate. Dev/on-prem
/// escape hatch only — mirrors the <c>ServerHistorian</c> TLS knobs. Must stay
/// <c>false</c> in production so MITM attacks against the broker connection fail closed.
/// </summary>
public bool AllowUntrustedServerCertificate { get; init; } = false;
/// <summary>
/// PEM CA file pinning the broker's TLS chain. <c>null</c>/empty uses the OS trust
/// store.
/// </summary>
public string? CaCertificatePath { get; init; }
/// <summary>Username for broker authentication. <c>null</c> connects without a username.</summary>
public string? Username { get; init; }
/// <summary>
/// Password for broker authentication. Blank in committed JSON — supply via env, never
/// commit or log.
/// </summary>
public string? Password { get; init; } = "";
/// <summary>MQTT protocol version to negotiate at CONNECT.</summary>
public MqttProtocolVersion ProtocolVersion { get; init; } = MqttProtocolVersion.V500;
/// <summary>Whether to request a clean session (v3.1.1) / clean start (v5.0) at CONNECT.</summary>
public bool CleanSession { get; init; } = true;
/// <summary>Keep-alive interval sent at CONNECT.</summary>
[Range(1, int.MaxValue)]
public int KeepAliveSeconds { get; init; } = 30;
/// <summary>Bounded connect deadline (see design §8) — the driver never hangs past this.</summary>
[Range(1, int.MaxValue)]
public int ConnectTimeoutSeconds { get; init; } = 15;
/// <summary>Initial reconnect backoff after a connection drop.</summary>
[Range(1, int.MaxValue)]
public int ReconnectMinBackoffSeconds { get; init; } = 1;
/// <summary>Cap on the exponential reconnect backoff.</summary>
[Range(1, int.MaxValue)]
public int ReconnectMaxBackoffSeconds { get; init; } = 30;
/// <summary>
/// Ceiling on an inbound message body, in bytes. A larger message is refused before any
/// decode or parse and degrades its own tags to <c>BadDecodingError</c>. Default 1 MiB.
/// </summary>
/// <remarks>
/// Decode and parse both run synchronously on MQTTnet's shared dispatcher thread, so their
/// cost is paid by <i>every</i> subscription, not just the offending topic — without a bound,
/// one publisher shipping a multi-megabyte body stalls the whole driver's delivery. "Unbounded"
/// is deliberately not offered; a non-positive value falls back to the driver's own 1 MiB
/// default. The literal is duplicated from <c>MqttSubscriptionManager.DefaultMaxPayloadBytes</c>
/// because this <c>.Contracts</c> assembly is <i>referenced by</i> the driver assembly and
/// cannot reference it back; the manager's constructor is the single enforcement point.
/// </remarks>
[Range(1, int.MaxValue)]
public int MaxPayloadBytes { get; init; } = 1024 * 1024;
/// <summary>Selects the ingest shape — Plain topics or Sparkplug B.</summary>
public MqttMode Mode { get; init; } = MqttMode.Plain;
/// <summary>
/// The cluster's authored raw MQTT tags, as delivered by the deploy artifact: each carries the
/// tag's <b>RawPath</b> (its v3 identity and driver wire reference) plus its <c>TagConfig</c>
/// blob. <see cref="MqttDriver"/> maps each through <see cref="MqttTagDefinitionFactory"/> at
/// initialize and re-registers the set <b>wholesale</b> on every reinitialize, so a redeploy
/// that drops a tag stops feeding it.
/// </summary>
/// <remarks>
/// This is also the <b>only</b> source of the driver's discoverable node set — plain MQTT has
/// no browsable address space, and a tag that is not authored here is never materialized no
/// matter how much traffic its topic carries. Mirrors <c>ModbusDriverOptions.RawTags</c> /
/// <c>FocasDriverOptions.RawTags</c>.
/// </remarks>
public IReadOnlyList<RawTagEntry> RawTags { get; init; } = [];
/// <summary>
/// Sparkplug-only settings. Populated when <see cref="Mode"/> is
/// <see cref="MqttMode.SparkplugB"/>; <c>null</c> in Plain mode.
/// </summary>
public MqttSparkplugOptions? Sparkplug { get; init; }
/// <summary>
/// Plain-mode-only settings. Populated when <see cref="Mode"/> is
/// <see cref="MqttMode.Plain"/>; <c>null</c> in Sparkplug B mode.
/// </summary>
public MqttPlainOptions? Plain { get; init; }
/// <summary>
/// Record-generated <c>ToString()</c> member printer, overridden so <see cref="Password"/>
/// never renders in plaintext — this DTO routinely lands in logs / exception messages /
/// debugger watches, and the plan's cross-cutting rule is explicit: never commit or log
/// creds. A set password prints as <c>***</c>; unset (<c>null</c>/empty) renders
/// distinguishably so the redaction can't be mistaken for a real secret.
/// </summary>
private bool PrintMembers(StringBuilder builder)
{
var passwordDisplay = Password switch
{
null => "<null>",
"" => "<empty>",
_ => "***",
};
builder.Append("Host = ").Append(Host);
builder.Append(", Port = ").Append(Port);
builder.Append(", ClientId = ").Append(ClientId);
builder.Append(", UseTls = ").Append(UseTls);
builder.Append(", AllowUntrustedServerCertificate = ").Append(AllowUntrustedServerCertificate);
builder.Append(", CaCertificatePath = ").Append(CaCertificatePath);
builder.Append(", Username = ").Append(Username);
builder.Append(", Password = ").Append(passwordDisplay);
builder.Append(", ProtocolVersion = ").Append(ProtocolVersion);
builder.Append(", CleanSession = ").Append(CleanSession);
builder.Append(", KeepAliveSeconds = ").Append(KeepAliveSeconds);
builder.Append(", ConnectTimeoutSeconds = ").Append(ConnectTimeoutSeconds);
builder.Append(", ReconnectMinBackoffSeconds = ").Append(ReconnectMinBackoffSeconds);
builder.Append(", ReconnectMaxBackoffSeconds = ").Append(ReconnectMaxBackoffSeconds);
builder.Append(", MaxPayloadBytes = ").Append(MaxPayloadBytes);
builder.Append(", Mode = ").Append(Mode);
builder.Append(", Sparkplug = ").Append(Sparkplug);
builder.Append(", Plain = ").Append(Plain);
// Count only: a deployment routinely authors thousands of raw tags and each carries a full
// TagConfig blob, so printing the list would turn any log of this DTO into a config dump.
builder.Append(", RawTags = ").Append(RawTags.Count).Append(" tag(s)");
return true;
}
}
/// <summary>
/// Sparkplug B settings for an <see cref="MqttDriverOptions"/> instance in
/// <see cref="MqttMode.SparkplugB"/> mode. See
/// <c>docs/plans/2026-07-15-mqtt-sparkplug-driver-design.md</c> §5.1.
/// </summary>
public sealed record MqttSparkplugOptions
{
/// <summary>Sparkplug group id — the driver subscribes <c>spBv1.0/{GroupId}/#</c>.</summary>
public string GroupId { get; init; } = "";
/// <summary>
/// Sparkplug Host Application identity — published as <c>spBv1.0/STATE/{HostId}</c>
/// (Sparkplug v3.0).
/// </summary>
public string HostId { get; init; } = "";
/// <summary>
/// When <c>true</c>, this driver instance acts as the Sparkplug Primary Host
/// Application. At most one primary host may exist per host-id per broker.
/// </summary>
public bool ActAsPrimaryHost { get; init; } = false;
/// <summary>
/// When <c>true</c>, a detected sequence-number gap triggers a Sparkplug rebirth
/// request (NCMD) rather than silently continuing with stale metric state.
/// </summary>
public bool RequestRebirthOnGap { get; init; } = true;
/// <summary>
/// How long, in seconds, browse/discovery collects NBIRTH/DBIRTH traffic before
/// considering the observed metric set stable.
/// </summary>
[Range(1, int.MaxValue)]
public int BirthObservationWindowSeconds { get; init; } = 15;
}
/// <summary>
/// Plain-MQTT-mode settings for an <see cref="MqttDriverOptions"/> instance in
/// <see cref="MqttMode.Plain"/> mode. See
/// <c>docs/plans/2026-07-15-mqtt-sparkplug-driver-design.md</c> §5.1.
/// </summary>
public sealed record MqttPlainOptions
{
/// <summary>
/// Optional prefix prepended when the driver needs to compose a topic (e.g. discovery
/// scoping); per-tag topics are authored explicitly and are unaffected.
/// </summary>
public string TopicPrefix { get; init; } = "";
/// <summary>Default QoS used when a tag's own <c>qos</c> is unset.</summary>
[Range(0, 2)]
public int DefaultQos { get; init; } = 1;
}
@@ -0,0 +1,53 @@
using System.Text.Json;
using System.Text.Json.Serialization;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// The <b>single</b> <see cref="JsonSerializerOptions"/> instance every MQTT driver-config seam
/// parses <see cref="MqttDriverOptions"/> through — the runtime factory
/// (<c>MqttDriverFactoryExtensions</c>), the Test-connect probe (<c>MqttDriverProbe</c>), the
/// driver's own <c>ReinitializeAsync</c> re-parse, and the address-picker browser
/// (<c>MqttDriverBrowser</c>).
/// </summary>
/// <remarks>
/// <para>
/// <b>Why one instance, and why here.</b> Divergent per-seam options are this repo's
/// documented systemic enum bug: an AdminUI-authored config with an enum-valued field is
/// accepted by the seam that carries a <see cref="JsonStringEnumConverter"/> and faults the
/// one that does not, so "Test connect" goes green and the deployed driver dies. Two other
/// drivers already carry a copy per seam. Rather than repeat that, this driver keeps exactly
/// one instance.
/// </para>
/// <para>
/// It lives in <c>.Contracts</c> — the assembly that owns <see cref="MqttDriverOptions"/> and
/// the three enums the converter exists for — rather than in the factory or the probe,
/// because <c>.Contracts</c> is the <i>only</i> assembly all four consumers already reference.
/// In particular the browser lives in its own assembly and reaches the runtime <c>.Driver</c>
/// project only through a deliberate, documented layering exception that is scheduled to be
/// removed (see the <c>ProjectReference</c> comment in the browser's csproj); anchoring the
/// options in <c>.Driver</c> would resurrect the duplicate the day that reference goes away.
/// </para>
/// <para>
/// <b>Not a general-purpose JSON policy.</b> <c>UnmappedMemberHandling.Skip</c> means an
/// unknown key is ignored rather than rejected — deliberate, so a config blob authored
/// against a newer driver still binds — and <c>PropertyNameCaseInsensitive</c> accepts both
/// the camelCase the AdminUI emits and the PascalCase a hand-edited blob may carry. A
/// <see cref="JsonSerializerOptions"/> becomes read-only on first use, so this instance is
/// safe to share across threads and must never be mutated after startup.
/// </para>
/// </remarks>
public static class MqttJson
{
/// <summary>
/// The shared options. Enum-valued knobs (<see cref="MqttMode"/>,
/// <see cref="MqttProtocolVersion"/>, <see cref="MqttPayloadFormat"/>) round-trip by
/// <b>name</b>; numeric ordinals still bind, so an older blob is not broken by this.
/// </summary>
public static readonly JsonSerializerOptions Options = new()
{
PropertyNameCaseInsensitive = true,
UnmappedMemberHandling = JsonUnmappedMemberHandling.Skip,
Converters = { new JsonStringEnumConverter() },
};
}
@@ -0,0 +1,14 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// Selects the ingest shape an <see cref="MqttDriverOptions"/> instance uses. See
/// <c>docs/plans/2026-07-15-mqtt-sparkplug-driver-design.md</c> §5.1.
/// </summary>
public enum MqttMode
{
/// <summary>Plain MQTT — the driver subscribes to authored topics directly.</summary>
Plain,
/// <summary>Sparkplug B — the driver decodes Tahu-encoded NBIRTH/DBIRTH/NDATA/DDATA payloads.</summary>
SparkplugB,
}
@@ -0,0 +1,17 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// How a Plain-mode MQTT tag's payload is decoded. See
/// <c>docs/plans/2026-07-15-mqtt-sparkplug-driver-design.md</c> §5.3.
/// </summary>
public enum MqttPayloadFormat
{
/// <summary>The payload is a JSON document; a JSONPath selects the value (<c>jsonPath</c>).</summary>
Json,
/// <summary>The payload bytes are used as-is (e.g. binary / opaque).</summary>
Raw,
/// <summary>The payload is a bare scalar string (e.g. <c>"23.5"</c>), parsed by <c>dataType</c>.</summary>
Scalar,
}
@@ -0,0 +1,14 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// MQTT wire protocol version negotiated at CONNECT. See
/// <c>docs/plans/2026-07-15-mqtt-sparkplug-driver-design.md</c> §5.1.
/// </summary>
public enum MqttProtocolVersion
{
/// <summary>MQTT 3.1.1.</summary>
V311,
/// <summary>MQTT 5.0.</summary>
V500,
}
@@ -0,0 +1,40 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// The <c>TagConfig</c> JSON key names that make up an MQTT tag's <b>address</b>, in the two shapes
/// the driver binds by: a single <see cref="Topic"/> (Plain) or the
/// <see cref="GroupId"/>/<see cref="EdgeNodeId"/>/<see cref="DeviceId"/>/<see cref="MetricName"/>
/// tuple (Sparkplug B).
/// </summary>
/// <remarks>
/// <para>
/// These exist because the address is produced in one place and consumed in another, and the
/// two used to agree only by coincidence: the AdminUI browse-commit mapper writes the blob,
/// <see cref="MqttTagDefinitionFactory"/> reads it, and a key-name drift between them is not a
/// compile error — it is a tag that deploys clean, resolves to nothing, and reports
/// <c>BadNodeIdUnknown</c> at runtime with no authoring-time signal at all. Single-sourcing
/// the names makes that drift impossible rather than merely unlikely.
/// </para>
/// <para>
/// <b>Address keys only.</b> The behavioural keys (<c>payloadFormat</c>, <c>jsonPath</c>,
/// <c>dataType</c>, <c>qos</c>, <c>retainSeed</c>) are deliberately not here — nothing produces
/// them from a browse, so they carry no cross-component drift risk.
/// </para>
/// </remarks>
public static class MqttTagConfigKeys
{
/// <summary>The concrete MQTT topic a Plain-mode tag subscribes to.</summary>
public const string Topic = "topic";
/// <summary>The Sparkplug group id — the first segment of the tag's binding tuple.</summary>
public const string GroupId = "groupId";
/// <summary>The Sparkplug edge-node id.</summary>
public const string EdgeNodeId = "edgeNodeId";
/// <summary>The Sparkplug device id; absent for a metric published by the edge node itself.</summary>
public const string DeviceId = "deviceId";
/// <summary>The Sparkplug metric's stable name — the tag's binding key across rebirths.</summary>
public const string MetricName = "metricName";
}
@@ -0,0 +1,85 @@
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// The driver's internal per-tag descriptor — the parsed form of an authored raw tag's
/// <c>TagConfig</c> JSON, produced by <see cref="MqttTagDefinitionFactory.FromTagConfig"/> and
/// looked up through the shared <see cref="EquipmentTagRefResolver{TDef}"/> exactly as Modbus does.
/// See <c>docs/plans/2026-07-15-mqtt-sparkplug-driver-design.md</c> §5.2 / §5.3.
/// <para>
/// The record carries <b>both</b> ingest shapes: the Plain-MQTT fields
/// (<see cref="Topic"/> … <see cref="RetainSeed"/>), populated by
/// <see cref="MqttTagDefinitionFactory.FromTagConfig"/>, and the Sparkplug B descriptor fields
/// (<see cref="GroupId"/>, <see cref="EdgeNodeId"/>, <see cref="DeviceId"/>,
/// <see cref="MetricName"/>), populated by
/// <see cref="MqttTagDefinitionFactory.FromSparkplugTagConfig"/> (Task 21). A Sparkplug tag is
/// resolved by the stable <c>(group, node, device, metricName)</c> tuple, <b>never</b> by the
/// per-birth metric alias — an alias may be reused across a rebirth for a different metric.
/// </para>
/// <para>
/// <b>Which factory runs is chosen by the driver's <see cref="MqttMode"/>, never sniffed from
/// the blob.</b> A blob carrying both shapes' keys is legal (the AdminUI editor preserves
/// unknown keys through a load→save, so a tag retyped from Plain to Sparkplug keeps its old
/// <c>topic</c>), and a heuristic would make the same authored tag mean two different things
/// depending on which key happened to survive.
/// </para>
/// </summary>
/// <param name="Name">
/// The definition's identity: the tag's <b>RawPath</b> — the cluster-scoped slash path that is the
/// v3 driver wire reference. This is the key <see cref="EquipmentTagRefResolver{TDef}"/> resolves
/// on, the key <c>OnDataChange</c> must publish under, and the key <c>DriverHostActor</c> fans out
/// to the raw + UNS NodeIds. It is emphatically <b>not</b> the authored <c>TagConfig</c> blob: the
/// pre-v3 blob-as-reference shape is retired, and publishing under a blob key would silently miss
/// the RawPath-keyed fan-out while every unit test still passed.
/// </param>
/// <param name="Topic">The concrete MQTT topic this tag subscribes to (Plain mode).</param>
/// <param name="PayloadFormat">How the received payload is decoded (Plain mode).</param>
/// <param name="JsonPath">
/// JSONPath selecting the value inside a <see cref="MqttPayloadFormat.Json"/> payload. Defaults
/// to the document root (<c>$</c>) when the blob omits it; meaningless for
/// <see cref="MqttPayloadFormat.Raw"/> / <see cref="MqttPayloadFormat.Scalar"/>.
/// </param>
/// <param name="DataType">
/// The tag's declared value type. Explicit authoring is strongly preferred — payload-shape
/// inference is the brittle fallback (design §4).
/// </param>
/// <param name="Qos">
/// The per-tag subscription QoS (02), or <see langword="null"/> to inherit the driver-level
/// <see cref="MqttPlainOptions.DefaultQos"/>.
/// </param>
/// <param name="RetainSeed">
/// Whether the broker's retained message for <see cref="Topic"/> seeds this tag's initial value
/// on subscribe (the OPC UA initial-data convention). Defaults to <see langword="true"/>.
/// </param>
/// <param name="DataTypeAuthored">
/// Whether <paramref name="DataType"/> came from an authored <c>dataType</c> key or is merely this
/// record's default. Load-bearing in <b>Sparkplug</b> mode only, where <c>dataType</c> is optional:
/// a Sparkplug metric's type is declared by its own birth certificate, so an unauthored tag must
/// take the type the NBIRTH/DBIRTH declared rather than silently coercing every value to the
/// <see cref="DriverDataType.String"/> default. Both factories set it from the key's presence so it
/// means the same thing in either mode; nothing on the Plain ingest path reads it.
/// </param>
/// <param name="GroupId">The Sparkplug group id; <see langword="null"/> for a Plain-mode tag.</param>
/// <param name="EdgeNodeId">The Sparkplug edge-node id; <see langword="null"/> for a Plain-mode tag.</param>
/// <param name="DeviceId">
/// The Sparkplug device id, or <see langword="null"/> for a metric published by the edge node
/// itself (and for every Plain-mode tag).
/// </param>
/// <param name="MetricName">
/// The Sparkplug metric name — the tag's stable identity across rebirths, and the key the ingest
/// state machine binds by. <see langword="null"/> for a Plain-mode tag.
/// </param>
public sealed record MqttTagDefinition(
string Name,
string Topic,
MqttPayloadFormat PayloadFormat,
string JsonPath,
DriverDataType DataType,
int? Qos,
bool RetainSeed,
bool DataTypeAuthored = true,
string? GroupId = null,
string? EdgeNodeId = null,
string? DeviceId = null,
string? MetricName = null);
@@ -0,0 +1,290 @@
using System.Text.Json;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// v3 pure mapper: turns an authored raw tag's <c>TagConfig</c> JSON (the shape produced by the
/// AdminUI <c>MqttTagConfigModel</c>) into an <see cref="MqttTagDefinition"/>. Under v3 a tag's
/// identity is its <b>RawPath</b> (a cluster-scoped slash path), not the address blob — so the
/// produced definition's <see cref="MqttTagDefinition.Name"/> is the RawPath the driver was handed,
/// which is exactly the wire reference the driver's <c>RawPath → def</c> resolver keys on. The driver
/// builds that table by mapping each <see cref="RawTagEntry"/> the deploy artifact delivers through
/// <see cref="FromTagConfig"/>.
/// <para>
/// Two entry points carry deliberately different strictness contracts.
/// <see cref="FromTagConfig"/> is the <b>runtime</b> path: it never throws, and anything
/// malformed returns <see langword="false"/> so the driver surfaces <c>BadNodeIdUnknown</c>
/// rather than a misleading <c>Good</c> off a wrong-typed default. <see cref="Inspect"/> is the
/// <b>deploy-time</b> path: it returns human-readable warnings so a bad tag config surfaces at
/// deploy instead of silently going dark at runtime.
/// </para>
/// <para>
/// <b>Two runtime entry points, one per ingest shape</b> — <see cref="FromTagConfig"/> (Plain)
/// and <see cref="FromSparkplugTagConfig"/> (Sparkplug B). The caller picks by the driver's
/// <see cref="MqttMode"/>; neither sniffs the blob. That is deliberate: the AdminUI editor
/// preserves unknown keys through a load→save, so a tag retyped from Plain to Sparkplug keeps
/// its old <c>topic</c> and a presence heuristic would make one authored blob mean two
/// different things depending on which key happened to survive. The <i>driver's</i> mode is
/// the single authority.
/// </para>
/// <para>
/// <b>No <c>ToTagConfig</c> inverse — a deliberate YAGNI call.</b> The six sibling factories
/// each carry one solely to serve their <c>Driver.&lt;X&gt;.Cli</c> project, which synthesises
/// <see cref="RawTagEntry"/> blobs from operator flags; the MQTT plan defines no
/// <c>Mqtt.Cli</c>. The AdminUI typed editor does not need one either — the
/// <c>&lt;Driver&gt;TagConfigModel</c> template serialises through its own preserved
/// <c>JsonObject</c> key bag and references no driver factory (verified: no
/// <c>TagDefinitionFactory</c> reference exists anywhere under the AdminUI project). Add the
/// inverse when a real caller appears, not before.
/// </para>
/// </summary>
public static class MqttTagDefinitionFactory
{
/// <summary>The JSONPath applied when a Json-format blob omits <c>jsonPath</c>: the document root.</summary>
private const string RootJsonPath = "$";
/// <summary>The MQTT wildcard characters — illegal in a <em>tag's</em> (concrete) subscription topic.</summary>
private static readonly char[] TopicWildcards = ['+', '#'];
/// <summary>
/// Maps an authored <c>TagConfig</c> object to a typed definition keyed by <paramref name="rawPath"/>.
/// The input is always an authored TagConfig JSON object (there is no name-vs-blob heuristic in v3).
/// Enum fields (<c>payloadFormat</c> / <c>dataType</c>) and <c>qos</c> are read STRICTLY — a
/// present-but-invalid (typo'd) value rejects the whole tag (returns <see langword="false"/> ⇒ the
/// driver surfaces <c>BadNodeIdUnknown</c>) rather than silently defaulting to a wrong-typed Good or
/// a downgraded delivery guarantee. Never throws.
/// <para>
/// A <b>wildcard</b> topic (<c>+</c> / <c>#</c>) is deliberately NOT rejected here: it is
/// ambiguous rather than unparseable, and <see cref="Inspect"/> is the operator-visible surface
/// for it. Rejecting it at runtime would turn an authoring mistake into a silent
/// <c>BadNodeIdUnknown</c> with no stated cause.
/// </para>
/// </summary>
/// <param name="tagConfig">The authored equipment-tag TagConfig JSON.</param>
/// <param name="rawPath">The tag's RawPath — becomes the definition's identity (<c>Name</c>).</param>
/// <param name="def">The mapped definition when this returns <see langword="true"/>.</param>
/// <returns><see langword="true"/> when <paramref name="tagConfig"/> is a valid MQTT tag-config object.</returns>
public static bool FromTagConfig(string tagConfig, string rawPath, out MqttTagDefinition def)
{
def = null!;
if (string.IsNullOrWhiteSpace(tagConfig) || tagConfig[0] != '{') return false;
try
{
using var doc = JsonDocument.Parse(tagConfig);
var root = doc.RootElement;
if (root.ValueKind != JsonValueKind.Object) return false;
// topic is the whole point of a Plain-mode tag — without one there is nothing to subscribe
// to, so an absent/blank value is a hard reject rather than a defaulted empty subscription.
var topic = ReadString(root, MqttTagConfigKeys.Topic);
if (string.IsNullOrWhiteSpace(topic)) return false;
// Strict enum reads: a present-but-invalid (typo'd) value rejects the whole tag
// (→ BadNodeIdUnknown) instead of silently defaulting to a wrong-typed Good.
if (!TagConfigJson.TryReadEnumStrict(root, "payloadFormat", MqttPayloadFormat.Json, out var payloadFormat))
return false;
var dataTypeAuthored = root.TryGetProperty("dataType", out _);
if (!TagConfigJson.TryReadEnumStrict(root, "dataType", DriverDataType.String, out var dataType))
return false;
// qos is read with the same strictness, for the same reason: silently absorbing a malformed
// "qos":"high" / "qos":1.5 / "qos":5 into the driver-level default would hand the operator a
// WEAKER delivery guarantee than the one they authored, with nothing to show for it.
if (!TryReadQosStrict(root, out var qos)) return false;
var jsonPath = ReadString(root, "jsonPath");
if (string.IsNullOrEmpty(jsonPath)) jsonPath = RootJsonPath;
def = new MqttTagDefinition(
Name: rawPath,
Topic: topic,
PayloadFormat: payloadFormat,
JsonPath: jsonPath,
DataType: dataType,
Qos: qos,
RetainSeed: ReadBoolOrDefault(root, "retainSeed", defaultValue: true),
DataTypeAuthored: dataTypeAuthored);
return true;
}
catch (JsonException) { return false; }
catch (FormatException) { return false; }
catch (InvalidOperationException) { return false; }
}
/// <summary>
/// Maps an authored <c>TagConfig</c> object to a <b>Sparkplug B</b> definition keyed by
/// <paramref name="rawPath"/>. The tag's binding identity is the stable
/// <c>(groupId, edgeNodeId, deviceId?, metricName)</c> tuple; <c>deviceId</c> is optional (absent
/// means the metric is published by the edge node itself), and there is <b>no</b> <c>topic</c> —
/// a Sparkplug driver subscribes one group-wide filter and routes by the decoded topic + birth
/// certificate, never by a per-tag topic. Never throws.
/// </summary>
/// <remarks>
/// <para>
/// <b><c>dataType</c> is OPTIONAL here, unlike in Plain mode.</b> A Sparkplug metric declares
/// its own datatype in its birth certificate, so an unauthored tag legitimately takes the
/// type the birth declares — that is the whole point of Task 22's <c>UntilStable</c>
/// discovery. It is still read <b>strictly</b> when present (a typo'd value rejects the tag,
/// exactly as in Plain mode) and the outcome is recorded on
/// <see cref="MqttTagDefinition.DataTypeAuthored"/> so the ingest path can tell "the operator
/// declared String" from "nobody declared anything and the record defaulted".
/// </para>
/// <para>
/// <b>Plain-shape keys are read but not required.</b> <c>topic</c>/<c>payloadFormat</c>/
/// <c>jsonPath</c> are meaningless to the Sparkplug ingest path and are deliberately NOT
/// validated — a blob retyped in the editor keeps them, and rejecting the tag for a stale
/// leftover key would take a correctly-authored Sparkplug tag dark with no stated cause.
/// </para>
/// </remarks>
/// <param name="tagConfig">The authored equipment-tag TagConfig JSON.</param>
/// <param name="rawPath">The tag's RawPath — becomes the definition's identity (<c>Name</c>).</param>
/// <param name="def">The mapped definition when this returns <see langword="true"/>.</param>
/// <returns><see langword="true"/> when the blob carries a usable Sparkplug descriptor.</returns>
public static bool FromSparkplugTagConfig(string tagConfig, string rawPath, out MqttTagDefinition def)
{
def = null!;
if (string.IsNullOrWhiteSpace(tagConfig) || tagConfig[0] != '{') return false;
try
{
using var doc = JsonDocument.Parse(tagConfig);
var root = doc.RootElement;
if (root.ValueKind != JsonValueKind.Object) return false;
// The binding tuple. Without all three of these there is nothing a birth certificate could
// ever bind the tag to, so an absent/blank value is a hard reject (→ BadNodeIdUnknown)
// rather than a definition that can never resolve and reports nothing about why.
var groupId = ReadString(root, MqttTagConfigKeys.GroupId);
var edgeNodeId = ReadString(root, MqttTagConfigKeys.EdgeNodeId);
var metricName = ReadString(root, MqttTagConfigKeys.MetricName);
if (string.IsNullOrWhiteSpace(groupId)
|| string.IsNullOrWhiteSpace(edgeNodeId)
|| string.IsNullOrWhiteSpace(metricName))
{
return false;
}
// Optional: a node-level metric has no device segment. Blank normalises to absent so
// "deviceId":"" and an omitted key describe the same scope rather than two.
var deviceId = ReadString(root, MqttTagConfigKeys.DeviceId);
if (string.IsNullOrWhiteSpace(deviceId)) deviceId = null;
// Strict, but only when present — see the remarks.
var dataTypeAuthored = root.TryGetProperty("dataType", out _);
if (!TagConfigJson.TryReadEnumStrict(root, "dataType", DriverDataType.String, out var dataType))
return false;
if (!TryReadQosStrict(root, out var qos)) return false;
def = new MqttTagDefinition(
Name: rawPath,
// Plain-shape fields carried through verbatim; the Sparkplug ingest path reads none of
// them, and a retyped blob's leftovers must not change what the tag means.
Topic: ReadString(root, MqttTagConfigKeys.Topic),
PayloadFormat: MqttPayloadFormat.Json,
JsonPath: RootJsonPath,
DataType: dataType,
Qos: qos,
RetainSeed: ReadBoolOrDefault(root, "retainSeed", defaultValue: true),
DataTypeAuthored: dataTypeAuthored,
GroupId: groupId,
EdgeNodeId: edgeNodeId,
DeviceId: deviceId,
MetricName: metricName);
return true;
}
catch (JsonException) { return false; }
catch (FormatException) { return false; }
catch (InvalidOperationException) { return false; }
}
/// <summary>
/// Deploy-time inspection of an equipment-tag <c>TagConfig</c> blob. Returns human-readable
/// warnings for a structurally unparseable blob (which the runtime turns into a silent
/// <c>BadNodeIdUnknown</c>), for present-but-invalid <c>payloadFormat</c> / <c>dataType</c> /
/// <c>qos</c> values, and for a <b>wildcard</b> tag topic (a tag bound to <c>+</c> / <c>#</c>
/// would receive values from many topics — ambiguous, and almost never what the operator meant).
/// Empty when the blob is clean or is not an equipment-tag TagConfig object. Never throws.
/// </summary>
/// <param name="reference">The equipment tag's TagConfig JSON.</param>
/// <returns>The warnings; empty when clean.</returns>
public static IReadOnlyList<string> Inspect(string reference)
{
var warnings = new List<string>();
if (string.IsNullOrWhiteSpace(reference) || reference[0] != '{') return warnings;
try
{
using var doc = JsonDocument.Parse(reference);
var root = doc.RootElement;
if (root.ValueKind != JsonValueKind.Object)
{
warnings.Add("Mqtt TagConfig root is not a JSON object — the tag will not resolve (BadNodeIdUnknown).");
return warnings;
}
foreach (var w in new[]
{
TagConfigJson.DescribeInvalidEnum<MqttPayloadFormat>(root, "payloadFormat"),
TagConfigJson.DescribeInvalidEnum<DriverDataType>(root, "dataType"),
DescribeInvalidQos(root),
})
{
if (w is not null) warnings.Add(w);
}
var topic = ReadString(root, "topic");
if (topic.IndexOfAny(TopicWildcards) >= 0)
{
warnings.Add(
$"value '{topic}' for 'topic' contains an MQTT wildcard (+ or #); a tag's topic must be " +
"concrete, or the tag will be fed by every matching topic.");
}
}
catch (JsonException)
{
warnings.Add("Mqtt TagConfig is not valid JSON — the tag will not resolve (BadNodeIdUnknown).");
}
return warnings;
}
/// <summary>
/// Strict <c>qos</c> read, mirroring <see cref="TagConfigJson.TryReadEnumStrict{TEnum}"/>'s
/// absent / valid / present-but-invalid split: absent ⇒ <see langword="null"/> (the driver-level
/// <see cref="MqttPlainOptions.DefaultQos"/> wins); a JSON integer in 02 ⇒ that value; anything
/// else present (non-number, non-integer, or out of range) ⇒ the read FAILS.
/// </summary>
/// <param name="o">The TagConfig root object.</param>
/// <param name="qos">The parsed QoS, or <see langword="null"/> when the field is absent.</param>
/// <returns><see langword="false"/> only when <c>qos</c> is present but not a legal MQTT QoS.</returns>
private static bool TryReadQosStrict(JsonElement o, out int? qos)
{
qos = null;
if (!o.TryGetProperty("qos", out var e)) return true; // absent
if (e.ValueKind != JsonValueKind.Number || !e.TryGetInt32(out var v)) return false;
if (v is < 0 or > 2) return false;
qos = v;
return true;
}
/// <summary>
/// A human-readable warning for a present-but-invalid <c>qos</c> field, or <see langword="null"/>
/// when it is absent or valid — the <c>qos</c> counterpart of
/// <see cref="TagConfigJson.DescribeInvalidEnum{TEnum}"/>.
/// </summary>
/// <param name="o">The TagConfig root object.</param>
/// <returns>The warning text, or <see langword="null"/>.</returns>
private static string? DescribeInvalidQos(JsonElement o)
{
if (!o.TryGetProperty("qos", out var e)) return null;
if (e.ValueKind == JsonValueKind.Number && e.TryGetInt32(out var v) && v is >= 0 and <= 2) return null;
return $"value '{e.GetRawText()}' for 'qos' is not a valid MQTT QoS; valid: 0, 1, 2";
}
private static string ReadString(JsonElement o, string name)
=> o.TryGetProperty(name, out var e) && e.ValueKind == JsonValueKind.String
? e.GetString() ?? ""
: "";
private static bool ReadBoolOrDefault(JsonElement o, string name, bool defaultValue)
=> o.TryGetProperty(name, out var e) && e.ValueKind is JsonValueKind.True or JsonValueKind.False
? e.GetBoolean()
: defaultValue;
}
@@ -0,0 +1,261 @@
// =====================================================================================================
// VENDORED THIRD-PARTY FILE DO NOT EDIT. Re-vendor from upstream instead (recipe below).
//
// Upstream project Eclipse Tahu the Sparkplug B reference implementation
// Upstream repo https://github.com/eclipse-tahu/tahu
// Upstream path sparkplug_b/sparkplug_b.proto
// Pinned commit 5736e404889d4b95910613040a99ba79589ffb13 (master, 2023-11-06)
// Permalink https://raw.githubusercontent.com/eclipse-tahu/tahu/5736e404889d4b95910613040a99ba79589ffb13/sparkplug_b/sparkplug_b.proto
// git blob SHA-1 bf72ab5f09a333afabcb40fd45362ffbb0c8c5bd (matches the tree entry at that commit)
// content SHA-256 4432c5c483b7fb9732d0594c98a2e97dca5e517e39c5374a8b918d837f0b4a19 (8330 bytes)
// last modified 46f25e79f34234e6145d11108660dfd9133ae50d (2022-05-16, template_ref comment fix)
// License Eclipse Public License 2.0 (EPL-2.0) https://www.eclipse.org/legal/epl-2.0/
// Copyright (c) Cirrus Link Solutions and others. The upstream copyright + SPDX
// header is preserved verbatim as the first lines of the copied body below.
//
// Everything from line 38 down ("// * Copyright (c) 2015, 2018 Cirrus Link Solutions…") is a
// BYTE-FOR-BYTE copy of the upstream file. Only these header lines were added. To re-verify:
//
// curl -sSL <permalink above> -o /tmp/upstream.proto
// tail -n +38 src/Drivers/ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Contracts/Protos/sparkplug_b.proto > /tmp/vendored.proto
// diff /tmp/upstream.proto /tmp/vendored.proto && shasum -a 256 /tmp/vendored.proto
//
// WHY VENDORED rather than a NuGet package: SparkplugNet is the only .NET Sparkplug library and it is
// stale (1.3.10, 2024-07-02), transitively pins MQTTnet 4.3.6.1152 against this repo's MQTTnet 5.2.0,
// and ships net6.0/net8.0 only no net10.0 TFM. So the schema is vendored and decoded directly with
// Google.Protobuf, over the same single MQTTnet-5 client the plain-MQTT path already uses.
//
// WHY THE proto2 FILE and not the sibling sparkplug_b_c_sharp.proto that upstream also ships: this is
// the NORMATIVE Sparkplug B schema (the C# sibling is a lossy proto3 restatement that drops explicit
// presence and rewrites the `extensions` ranges as `google.protobuf.Any`). protoc from Grpc.Tools
// generates valid C# from proto2, and the generated namespace is identical (Org.Eclipse.Tahu.Protobuf,
// PascalCased from `package org.eclipse.tahu.protobuf` there is no `option csharp_namespace`). Keeping
// proto2 buys explicit presence Has{Name,Alias,Seq,IsNull,Datatype} which the decoder needs to tell
// "field absent" from "field present and zero" (an NBIRTH legitimately carries seq = 0, and a DATA
// metric legitimately omits `name` and carries only `alias`).
// =====================================================================================================
// * Copyright (c) 2015, 2018 Cirrus Link Solutions and others
// *
// * This program and the accompanying materials are made available under the
// * terms of the Eclipse Public License 2.0 which is available at
// * http://www.eclipse.org/legal/epl-2.0.
// *
// * SPDX-License-Identifier: EPL-2.0
// *
// * Contributors:
// * Cirrus Link Solutions - initial implementation
//
// To compile:
// cd client_libraries/java
// protoc --proto_path=../../ --java_out=src/main/java ../../sparkplug_b.proto
//
syntax = "proto2";
package org.eclipse.tahu.protobuf;
option java_package = "org.eclipse.tahu.protobuf";
option java_outer_classname = "SparkplugBProto";
enum DataType {
// Indexes of Data Types
// Unknown placeholder for future expansion.
Unknown = 0;
// Basic Types
Int8 = 1;
Int16 = 2;
Int32 = 3;
Int64 = 4;
UInt8 = 5;
UInt16 = 6;
UInt32 = 7;
UInt64 = 8;
Float = 9;
Double = 10;
Boolean = 11;
String = 12;
DateTime = 13;
Text = 14;
// Additional Metric Types
UUID = 15;
DataSet = 16;
Bytes = 17;
File = 18;
Template = 19;
// Additional PropertyValue Types
PropertySet = 20;
PropertySetList = 21;
// Array Types
Int8Array = 22;
Int16Array = 23;
Int32Array = 24;
Int64Array = 25;
UInt8Array = 26;
UInt16Array = 27;
UInt32Array = 28;
UInt64Array = 29;
FloatArray = 30;
DoubleArray = 31;
BooleanArray = 32;
StringArray = 33;
DateTimeArray = 34;
}
message Payload {
message Template {
message Parameter {
optional string name = 1;
optional uint32 type = 2;
oneof value {
uint32 int_value = 3;
uint64 long_value = 4;
float float_value = 5;
double double_value = 6;
bool boolean_value = 7;
string string_value = 8;
ParameterValueExtension extension_value = 9;
}
message ParameterValueExtension {
extensions 1 to max;
}
}
optional string version = 1; // The version of the Template to prevent mismatches
repeated Metric metrics = 2; // Each metric includes a name, datatype, and optionally a value
repeated Parameter parameters = 3;
optional string template_ref = 4; // MUST be a reference to a template definition if this is an instance (i.e. the name of the template definition) - MUST be omitted for template definitions
optional bool is_definition = 5;
extensions 6 to max;
}
message DataSet {
message DataSetValue {
oneof value {
uint32 int_value = 1;
uint64 long_value = 2;
float float_value = 3;
double double_value = 4;
bool boolean_value = 5;
string string_value = 6;
DataSetValueExtension extension_value = 7;
}
message DataSetValueExtension {
extensions 1 to max;
}
}
message Row {
repeated DataSetValue elements = 1;
extensions 2 to max; // For third party extensions
}
optional uint64 num_of_columns = 1;
repeated string columns = 2;
repeated uint32 types = 3;
repeated Row rows = 4;
extensions 5 to max; // For third party extensions
}
message PropertyValue {
optional uint32 type = 1;
optional bool is_null = 2;
oneof value {
uint32 int_value = 3;
uint64 long_value = 4;
float float_value = 5;
double double_value = 6;
bool boolean_value = 7;
string string_value = 8;
PropertySet propertyset_value = 9;
PropertySetList propertysets_value = 10; // List of Property Values
PropertyValueExtension extension_value = 11;
}
message PropertyValueExtension {
extensions 1 to max;
}
}
message PropertySet {
repeated string keys = 1; // Names of the properties
repeated PropertyValue values = 2;
extensions 3 to max;
}
message PropertySetList {
repeated PropertySet propertyset = 1;
extensions 2 to max;
}
message MetaData {
// Bytes specific metadata
optional bool is_multi_part = 1;
// General metadata
optional string content_type = 2; // Content/Media type
optional uint64 size = 3; // File size, String size, Multi-part size, etc
optional uint64 seq = 4; // Sequence number for multi-part messages
// File metadata
optional string file_name = 5; // File name
optional string file_type = 6; // File type (i.e. xml, json, txt, cpp, etc)
optional string md5 = 7; // md5 of data
// Catchalls and future expansion
optional string description = 8; // Could be anything such as json or xml of custom properties
extensions 9 to max;
}
message Metric {
optional string name = 1; // Metric name - should only be included on birth
optional uint64 alias = 2; // Metric alias - tied to name on birth and included in all later DATA messages
optional uint64 timestamp = 3; // Timestamp associated with data acquisition time
optional uint32 datatype = 4; // DataType of the metric/tag value
optional bool is_historical = 5; // If this is historical data and should not update real time tag
optional bool is_transient = 6; // Tells consuming clients such as MQTT Engine to not store this as a tag
optional bool is_null = 7; // If this is null - explicitly say so rather than using -1, false, etc for some datatypes.
optional MetaData metadata = 8; // Metadata for the payload
optional PropertySet properties = 9;
oneof value {
uint32 int_value = 10;
uint64 long_value = 11;
float float_value = 12;
double double_value = 13;
bool boolean_value = 14;
string string_value = 15;
bytes bytes_value = 16; // Bytes, File
DataSet dataset_value = 17;
Template template_value = 18;
MetricValueExtension extension_value = 19;
}
message MetricValueExtension {
extensions 1 to max;
}
}
optional uint64 timestamp = 1; // Timestamp at message sending time
repeated Metric metrics = 2; // Repeated forever - no limit in Google Protobufs
optional uint64 seq = 3; // Sequence number
optional string uuid = 4; // UUID to track message type in terms of schema definitions
optional bytes body = 5; // To optionally bypass the whole definition above
extensions 6 to max; // For third party extensions
}
@@ -0,0 +1,139 @@
// `SparkplugDataType` is an ALIAS for the vendored proto's generated `Org.Eclipse.Tahu.Protobuf.DataType`
// enum, not a second, hand-maintained enum. See the remarks on `SparkplugDataTypeExtensions` below for
// the reasoning; the short version is that a duplicate enum is a drift hazard this repo has a documented
// systemic bug class around (see CLAUDE.md "Driver enum-serialization bug"), and `Payload.Types.Metric`'s
// wire-level `Datatype` field is a raw `uint32` anyway — nothing structurally forces a second CLR enum to
// exist, so the lowest-risk shape is for `SparkplugDataType` to be the exact same type as the generated
// one, not a value-compatible lookalike that some cast has to bridge.
global using SparkplugDataType = Org.Eclipse.Tahu.Protobuf.DataType;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// Maps a Sparkplug B metric <see cref="SparkplugDataType"/> (the vendored Eclipse Tahu proto's
/// generated <c>Org.Eclipse.Tahu.Protobuf.DataType</c> enum — see the <c>global using</c> alias
/// above) to a driver-agnostic <see cref="DriverDataType"/>, per design doc §3.5.
/// </summary>
/// <remarks>
/// <para>
/// <b>Why an alias, not a duplicate enum.</b> The obvious "clean seam" shape is a fresh
/// <c>Driver.Mqtt.Contracts</c>-owned enum with its own <c>ToDriverDataType()</c> — but that
/// creates a second definition of the same 35-member vocabulary that has to be hand-kept in
/// sync with whatever Eclipse Tahu's <c>sparkplug_b.proto</c> defines, which is exactly the
/// enum-drift shape this repo already has a name for (CLAUDE.md's "Driver enum-serialization
/// bug (AdminUI authoring)" — a systemic mismatch between two enums meant to describe the same
/// thing). It also does not match how the decoder actually produces values: Sparkplug's
/// <c>Metric.datatype</c> wire field is a raw <c>uint32</c> (see <c>sparkplug_b.proto</c> line
/// 230 — <c>optional uint32 datatype = 4</c>), not the <c>DataType</c> enum type itself, so
/// <c>SparkplugCodec</c> (Task 16) already casts the decoded value straight to the generated
/// <c>Org.Eclipse.Tahu.Protobuf.DataType</c> (locally aliased there as <c>TahuDataType</c>).
/// A second, hand-duplicated enum would force every downstream consumer (Tasks 18/19/20/21) to
/// cast between two value-compatible-but-nominally-different enums to call this extension
/// method — extra surface for exactly zero benefit, since both "enums" would need to enumerate
/// the identical 35 members in the identical order to stay castable.
/// </para>
/// <para>
/// Making <see cref="SparkplugDataType"/> a <c>global using</c> alias for the generated type
/// sidesteps all of that: there is no second definition to drift, by construction — the alias
/// and the generated enum are the exact same CLR type. <see cref="ToDriverDataType"/> below is
/// written as an extension on <see cref="SparkplugDataType"/> purely for call-site readability;
/// it is equally callable as <c>someDecodedDatatype.ToDriverDataType()</c> against a plain
/// <c>Org.Eclipse.Tahu.Protobuf.DataType</c> value with no cast, which is exactly the shape
/// <c>SparkplugCodec</c>'s decoded metrics hand back.
/// </para>
/// <para>
/// <b>Drift guard.</b> Because there is only one enum, <c>SparkplugDataTypeTests</c>'
/// completeness test (<c>ToDriverDataType_HandlesEveryGeneratedDataTypeMember_...</c>) can
/// enumerate <c>Enum.GetValues&lt;SparkplugDataType&gt;()</c> — which, being the alias, is
/// literally the live generated member set — and assert every member is either mapped or on
/// the explicit unsupported list below. If Eclipse Tahu's proto ever gains a member, that test
/// picks it up automatically and fails until this map makes an explicit decision about it,
/// rather than the new member silently falling through a duplicate-enum's stale default.
/// </para>
/// <para>
/// Per design §3.5: <c>Int8</c>/<c>UInt8</c> widen to <see cref="DriverDataType.Int16"/> /
/// <see cref="DriverDataType.UInt16"/> (no OPC UA signed-byte distinction is needed and
/// <see cref="DriverDataType"/> has no 8-bit members at all); <c>Float</c>/<c>Double</c> map to
/// <see cref="DriverDataType.Float32"/> / <see cref="DriverDataType.Float64"/> — note there is
/// no <c>DriverDataType.Double</c> member, only <c>Float64</c>; <c>Text</c>/<c>UUID</c>
/// (generated as <c>Uuid</c> — protoc mangles the wire spelling, see
/// <c>SparkplugProtoCodegenTests.DataTypeEnum_CSharpNamesAreProtocMangled_NotTheProtoSpelling</c>)
/// /<c>Bytes</c>/<c>File</c> all fall back to <see cref="DriverDataType.String"/> (v1 base64/raw
/// fallback for Bytes/File, per design); every <c>*Array</c> variant maps to its scalar element
/// type (<see cref="IsSparkplugArray"/> carries the "and it's an array" bit separately, since
/// <see cref="DriverDataType"/> itself has no array concept — that lives at the OPC UA
/// ValueRank/ArrayDimensions layer the address-space builder owns). <c>DataSet</c>/
/// <c>Template</c>/<c>PropertySet</c>/<c>PropertySetList</c>/<c>Unknown</c> are unsupported in
/// v1 and map to <see langword="null"/> — deliberately, not a guessed
/// <see cref="DriverDataType.String"/>, so a caller has to make an explicit skip-or-warn
/// decision (unlike Galaxy's <c>DataTypeMap</c>, which silently defaults unknown codes to
/// <c>String</c> for legacy wire-compatibility reasons that do not apply here).
/// </para>
/// </remarks>
public static class SparkplugDataTypeExtensions
{
/// <summary>
/// Maps a Sparkplug metric datatype to the equivalent <see cref="DriverDataType"/>, or
/// <see langword="null"/> if the type is unsupported in v1 (<c>DataSet</c>, <c>Template</c>,
/// <c>PropertySet</c>, <c>PropertySetList</c>, <c>Unknown</c>). Callers must treat
/// <see langword="null"/> as an explicit "skip and warn", never coerce it to a guessed type.
/// </summary>
public static DriverDataType? ToDriverDataType(this SparkplugDataType dataType) => dataType switch
{
SparkplugDataType.Int8 => DriverDataType.Int16,
SparkplugDataType.Int16 => DriverDataType.Int16,
SparkplugDataType.Int32 => DriverDataType.Int32,
SparkplugDataType.Int64 => DriverDataType.Int64,
SparkplugDataType.Uint8 => DriverDataType.UInt16,
SparkplugDataType.Uint16 => DriverDataType.UInt16,
SparkplugDataType.Uint32 => DriverDataType.UInt32,
SparkplugDataType.Uint64 => DriverDataType.UInt64,
SparkplugDataType.Float => DriverDataType.Float32,
SparkplugDataType.Double => DriverDataType.Float64,
SparkplugDataType.Boolean => DriverDataType.Boolean,
SparkplugDataType.String => DriverDataType.String,
SparkplugDataType.DateTime => DriverDataType.DateTime,
SparkplugDataType.Text => DriverDataType.String,
SparkplugDataType.Uuid => DriverDataType.String,
SparkplugDataType.Bytes => DriverDataType.String,
SparkplugDataType.File => DriverDataType.String,
SparkplugDataType.Int8Array => DriverDataType.Int16,
SparkplugDataType.Int16Array => DriverDataType.Int16,
SparkplugDataType.Int32Array => DriverDataType.Int32,
SparkplugDataType.Int64Array => DriverDataType.Int64,
SparkplugDataType.Uint8Array => DriverDataType.UInt16,
SparkplugDataType.Uint16Array => DriverDataType.UInt16,
SparkplugDataType.Uint32Array => DriverDataType.UInt32,
SparkplugDataType.Uint64Array => DriverDataType.UInt64,
SparkplugDataType.FloatArray => DriverDataType.Float32,
SparkplugDataType.DoubleArray => DriverDataType.Float64,
SparkplugDataType.BooleanArray => DriverDataType.Boolean,
SparkplugDataType.StringArray => DriverDataType.String,
SparkplugDataType.DateTimeArray => DriverDataType.DateTime,
// Unsupported v1 (design §3.5): DataSet/Template are deferred scope; PropertySet/
// PropertySetList are PropertyValue-only metadata types that never legitimately appear as a
// Metric's own datatype; Unknown is the proto's explicit "placeholder for future expansion"
// (index 0). All five fall through here deliberately rather than being listed with a fake
// mapping.
_ => null,
};
/// <summary>
/// Whether <paramref name="dataType"/> is one of the 13 <c>*Array</c> variants. Combine with
/// <see cref="ToDriverDataType"/> (which already returns the *element* type for an array
/// variant) to build an OPC UA ValueRank=1 array node per design §3.5.
/// </summary>
public static bool IsSparkplugArray(this SparkplugDataType dataType) => dataType switch
{
SparkplugDataType.Int8Array or SparkplugDataType.Int16Array or SparkplugDataType.Int32Array
or SparkplugDataType.Int64Array or SparkplugDataType.Uint8Array or SparkplugDataType.Uint16Array
or SparkplugDataType.Uint32Array or SparkplugDataType.Uint64Array or SparkplugDataType.FloatArray
or SparkplugDataType.DoubleArray or SparkplugDataType.BooleanArray or SparkplugDataType.StringArray
or SparkplugDataType.DateTimeArray => true,
_ => false,
};
}
@@ -0,0 +1,36 @@
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup>
<TargetFramework>net10.0</TargetFramework>
<Nullable>enable</Nullable>
<ImplicitUsings>enable</ImplicitUsings>
<TreatWarningsAsErrors>true</TreatWarningsAsErrors>
</PropertyGroup>
<ItemGroup>
<ProjectReference Include="..\..\Core\ZB.MOM.WW.OtOpcUa.Core.Abstractions\ZB.MOM.WW.OtOpcUa.Core.Abstractions.csproj"/>
</ItemGroup>
<!-- Sparkplug B (P2): the vendored Eclipse Tahu schema is compiled here, in the one assembly all
four MQTT projects reference. Message-only codegen (GrpcServices="None") — Sparkplug rides MQTT,
there is no gRPC service, so no Grpc.Core.Api runtime reference is needed (Commons, this repo's
other locally-compiled proto, takes one only because it generates service stubs).
Google.Protobuf is a SERIALIZATION dependency, not a transport one, so .Contracts stays
transport-free: it still has no MQTTnet reference (dropped in Task 1) and Google.Protobuf's own
graph is framework-only. Grpc.Tools is build-time (PrivateAssets=all) and flows to no consumer.
Build-platform note (inherited, not new): Grpc.Tools' bundled linux_arm64 protoc segfaults
(exit 139) under Apple-Silicon Docker, which is why docker-dev/Dockerfile pins its build stage
to the linux/amd64 platform. That pin already exists for the Commons proto; this project adds
a second .proto to the same already-pinned build, not a new constraint. -->
<ItemGroup>
<PackageReference Include="Google.Protobuf"/>
<PackageReference Include="Grpc.Tools">
<PrivateAssets>all</PrivateAssets>
<IncludeAssets>runtime; build; native; contentfiles; analyzers; buildtransitive</IncludeAssets>
</PackageReference>
</ItemGroup>
<ItemGroup>
<Protobuf Include="Protos\sparkplug_b.proto" GrpcServices="None"/>
</ItemGroup>
</Project>
@@ -0,0 +1,63 @@
using System.Collections.Concurrent;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// Thread-safe last-observed-value store backing <see cref="IReadable"/> for MQTT/Sparkplug B.
/// MQTT is a subscribe-first protocol — the driver holds one live broker connection and values
/// arrive pushed, not polled. But the OPC UA server still issues polled <c>IReadable.ReadAsync</c>
/// batches, so this cache is the bridge: the subscription path (message handler) calls
/// <see cref="Update"/> with the newest observed value per reference, and the read path serves
/// the most recent one via <see cref="Read"/>.
/// </summary>
/// <remarks>
/// Keyed by <c>RawPath</c> — the v3 driver-reference identity (see
/// <c>EquipmentTagRefResolver&lt;TDef&gt;</c> and <see cref="MqttTagDefinition.Name"/>,
/// which <b>is</b> the RawPath) — never a topic/JSON-path-derived key. The parameter is
/// therefore named <c>rawPath</c> throughout rather than "topic" or "key" to keep that identity
/// fact visible at every call site.
/// </remarks>
public sealed class LastValueCache
{
// OPC UA BadWaitingForInitialData (0x80320000) — the repo-wide convention for "no value has
// been observed yet" (mirrored by CalculationDriver before its first evaluation,
// VirtualTagEngine for a freshly-materialised node, and OtOpcUaNodeManager /
// AddressSpaceApplier for a just-deployed variable). MQTT is subscribe-first, so an unseen
// RawPath is in exactly that state until its first publish arrives. Deliberately NOT
// GoodNoData — that code is reserved (see NullHistorianDataSource, OtOpcUaNodeManager
// HistoryRead paths) for "the historian window held no samples", a different question from
// "has this live reference ever been observed".
private const uint BadWaitingForInitialData = 0x80320000u;
private readonly ConcurrentDictionary<string, DataValueSnapshot> _values = new();
/// <summary>
/// Record the newest observed value for <paramref name="rawPath"/>. Called from the MQTT
/// subscription/message-handling path. A <c>null</c> or empty <paramref name="rawPath"/>
/// is a no-op — never throws.
/// </summary>
/// <param name="rawPath">The RawPath identifying the tag.</param>
/// <param name="snapshot">The newest observed value/quality/timestamps for that tag.</param>
public void Update(string rawPath, DataValueSnapshot snapshot)
{
if (string.IsNullOrEmpty(rawPath)) return;
_values[rawPath] = snapshot;
}
/// <summary>
/// Read the last observed value for <paramref name="rawPath"/>. Never throws: a
/// <c>null</c>/empty <paramref name="rawPath"/> or a RawPath never observed both return a
/// <see cref="BadWaitingForInitialData"/> snapshot rather than an exception, so a batch read
/// covering many references degrades per-reference instead of failing the whole call.
/// </summary>
/// <param name="rawPath">The RawPath identifying the tag.</param>
/// <returns>The last observed snapshot, or a BadWaitingForInitialData snapshot if none has been observed.</returns>
public DataValueSnapshot Read(string rawPath)
{
if (!string.IsNullOrEmpty(rawPath) && _values.TryGetValue(rawPath, out var snapshot))
return snapshot;
return new DataValueSnapshot(null, BadWaitingForInitialData, null, DateTime.UtcNow);
}
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,84 @@
using System.Text.Json;
using Microsoft.Extensions.Logging;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
using ZB.MOM.WW.OtOpcUa.Core.Hosting;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// Registers the MQTT / Sparkplug B driver with the <see cref="DriverFactoryRegistry"/>. The
/// Host's <c>DriverFactoryBootstrap</c> calls <see cref="Register"/> once at startup; the
/// driver-instance bootstrapper then materialises <c>DriverInstance</c> rows of type
/// <see cref="DriverTypeNames.Mqtt"/> into live <see cref="MqttDriver"/> instances.
/// </summary>
/// <remarks>
/// <para>
/// <b>Direct deserialization, no intermediate DTO.</b> The <c>DriverConfig</c> blob binds
/// straight onto <see cref="MqttDriverOptions"/> — the same shape
/// <see cref="MqttDriverProbe"/>, <c>MqttDriverBrowser</c> and
/// <see cref="MqttDriver.ReinitializeAsync"/> already bind, and the shape the options record
/// was designed for (every knob carries its own default, and <c>rawTags</c> is a plain
/// <see cref="RawTagEntry"/> list needing no translation). Mirrors
/// <c>OpcUaClientDriverFactoryExtensions</c>. <c>ModbusDriverFactoryExtensions</c>' separate
/// DTO exists to service a legacy nullable-everything blob plus string→enum parsing this
/// driver does with a converter instead; copying it here would add a <i>fourth</i> parse
/// shape for one config, which is precisely the divergence
/// <see cref="MqttJson.Options"/> exists to prevent.
/// </para>
/// <para>
/// <b>Connection-free.</b> <see cref="CreateInstance"/> parses and constructs only — the
/// <see cref="MqttDriver"/> constructor touches no network, and the registry contract
/// forbids the factory from calling <c>InitializeAsync</c> itself (the driver host owns the
/// retry semantics).
/// </para>
/// </remarks>
public static class MqttDriverFactoryExtensions
{
/// <summary>
/// Driver type name — matches <c>DriverInstance.DriverType</c> values. Sourced from
/// <see cref="DriverTypeNames.Mqtt"/> so the registration key can never drift from the
/// constant the dispatch maps, the probe and the browser reference.
/// </summary>
public const string DriverTypeName = DriverTypeNames.Mqtt;
/// <summary>
/// Register the MQTT factory with the driver registry. The optional
/// <paramref name="loggerFactory"/> is captured at registration time and used to construct an
/// <see cref="ILogger{MqttDriver}"/> per driver instance — without it the driver runs with no
/// logger (standalone/test callers stay unchanged).
/// </summary>
/// <param name="registry">The driver factory registry to register with.</param>
/// <param name="loggerFactory">Optional logger factory used to create per-instance loggers.</param>
public static void Register(DriverFactoryRegistry registry, ILoggerFactory? loggerFactory = null)
{
ArgumentNullException.ThrowIfNull(registry);
registry.Register(DriverTypeName, (id, json) => CreateInstance(id, json, loggerFactory));
}
/// <summary>Public for the Server-side bootstrapper + test consumers.</summary>
/// <param name="driverInstanceId">Stable logical id of the driver instance.</param>
/// <param name="driverConfigJson">The <c>DriverConfig</c> JSON blob for the instance.</param>
/// <param name="loggerFactory">Optional logger factory for the per-instance logger.</param>
/// <returns>A configured, not-yet-connected <see cref="MqttDriver"/>.</returns>
/// <exception cref="ArgumentException">
/// <paramref name="driverInstanceId"/> or <paramref name="driverConfigJson"/> is blank.
/// </exception>
/// <exception cref="InvalidOperationException">The config blob deserialised to <c>null</c>.</exception>
/// <exception cref="JsonException">The config blob is not valid JSON for the options shape.</exception>
public static MqttDriver CreateInstance(
string driverInstanceId,
string driverConfigJson,
ILoggerFactory? loggerFactory = null)
{
ArgumentException.ThrowIfNullOrWhiteSpace(driverInstanceId);
ArgumentException.ThrowIfNullOrWhiteSpace(driverConfigJson);
// MqttJson.Options — the one shared instance (see its remarks). A local copy here is the
// repo's documented systemic enum bug in the making.
var options = JsonSerializer.Deserialize<MqttDriverOptions>(driverConfigJson, MqttJson.Options)
?? throw new InvalidOperationException(
$"MQTT driver config for '{driverInstanceId}' deserialised to null");
return new MqttDriver(options, driverInstanceId, loggerFactory?.CreateLogger<MqttDriver>());
}
}
@@ -0,0 +1,226 @@
using System.Diagnostics;
using System.Net.Sockets;
using System.Security.Authentication;
using System.Text.Json;
using Microsoft.Extensions.Logging.Abstractions;
using MQTTnet;
using ZB.MOM.WW.OtOpcUa.Core.Abstractions;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt;
/// <summary>
/// CONNECT-handshake probe for the <see cref="MqttDriverOptions"/>-shaped driver config.
/// Opens a bounded MQTT CONNECT against the configured broker; a CONNACK-accepted response
/// (<c>MqttClientConnectResultCode.Success</c>) is the "device is answering" proof (green +
/// latency). Refused/TLS/auth/timeout failures each surface a targeted message. Mirrors
/// <c>ModbusDriverProbe</c>'s shape.
/// </summary>
/// <remarks>
/// <para>
/// <b>Reuses <see cref="MqttConnection.BuildClientOptions"/></b> for the connect (TLS,
/// CA-pin, credentials, protocol-version mapping) rather than hand-rolling a second
/// connect path — a duplicate would be a security divergence. A distinct client-id suffix
/// (see <see cref="ProbeClientIdPrefix"/>) keeps a probe from colliding with the driver's
/// own session: an MQTT broker disconnects an existing client when a new one CONNECTs
/// with the same client id, so a probe that reused the driver's id would knock the
/// running driver offline every time an operator clicked "Test connect".
/// </para>
/// <para>
/// <b>A broker-rejected CONNACK is not a thrown exception.</b> Live-probed against
/// MQTTnet 5.2.0.1603: <c>IMqttClient.ConnectAsync</c> returns a
/// <c>MqttClientConnectResult</c> whose <c>ResultCode</c> carries the broker's CONNACK
/// reason (e.g. <c>NotAuthorized</c>, <c>BadUserNameOrPassword</c>) — it does
/// <b>not</b> throw for a rejected CONNACK. Only transport-level failures throw, and
/// their shape varies by where the failure occurs: a closed/refused port throws
/// <c>MqttCommunicationException</c> wrapping a <see cref="SocketException"/>; an
/// untrusted broker certificate throws <c>MqttCommunicationException</c> wrapping an
/// <see cref="AuthenticationException"/>; a broker that accepts the TCP handshake and
/// then never answers CONNACK throws either <c>MqttConnectingFailedException</c>
/// (wrapping "Connection closed") or a bare <see cref="OperationCanceledException"/>
/// ("MQTT connect canceled"), inconsistently, depending on whether MQTTnet's own
/// internal <c>Timeout</c> or the deadline token wins the race. Consequently this probe —
/// like <see cref="MqttConnection.Classify"/> — never switches on exception type to
/// detect a timeout; it checks the deadline token itself, which is authoritative
/// regardless of which exception shape MQTTnet happened to throw.
/// </para>
/// </remarks>
public sealed class MqttDriverProbe : IDriverProbe
{
/// <summary>
/// Marks the transient probe identity in the broker's client-id/session logs — mirrors the
/// browser's own <c>-browse-{guid8}</c> suffix so the two transient sessions are
/// distinguishable in broker-side logs, and so a probe can never collide with (and knock
/// offline) the driver's own live session.
/// </summary>
internal const string ProbeClientIdPrefix = "-probe-";
/// <inheritdoc />
public string DriverType => DriverTypeNames.Mqtt;
/// <inheritdoc />
public async Task<DriverProbeResult> ProbeAsync(string configJson, TimeSpan timeout, CancellationToken ct)
{
MqttDriverOptions? options;
try
{
// MqttJson.Options — the one shared instance across factory / probe / driver / browser
// (see its remarks): enums must round-trip by NAME or an AdminUI-authored config that
// Test-connect accepts would fault the deployed driver.
options = JsonSerializer.Deserialize<MqttDriverOptions>(configJson, MqttJson.Options);
}
catch (Exception ex)
{
return new DriverProbeResult(false, $"Config JSON is invalid: {ex.Message}", null);
}
if (options is null)
{
return new DriverProbeResult(false, "Config JSON deserialized to null.", null);
}
if (string.IsNullOrWhiteSpace(options.Host) || options.Port is <= 0 or > 65535)
{
return new DriverProbeResult(false, "Config has no host/port to probe.", null);
}
// Honour the config's own connectTimeoutSeconds (factory parity, mirrors
// ModbusDriverProbe) — the caller's timeout is the fallback when the config omits it.
var effectiveSeconds = options.ConnectTimeoutSeconds > 0
? options.ConnectTimeoutSeconds
: (int)Math.Ceiling(timeout.TotalSeconds);
if (effectiveSeconds <= 0)
{
effectiveSeconds = 1;
}
// Rebuild with the resolved timeout so BuildClientOptions' own WithTimeout(...) and this
// method's deadline agree — both must expire at the same instant for the deadline check
// below to be a reliable signal regardless of which of the two actually threw.
var probeOptions = options with { ConnectTimeoutSeconds = effectiveSeconds };
MqttClientOptions clientOptions;
try
{
clientOptions = MqttConnection.BuildClientOptions(
probeOptions,
ProbeClientIdPrefix + Guid.NewGuid().ToString("N")[..8],
NullLogger.Instance);
}
catch (Exception ex)
{
// Never let a configured Password reach the message — BuildClientOptions itself never
// logs it, and no exception path through it can format credential values.
return new DriverProbeResult(false, $"Config could not be used to build a connection: {ex.Message}", null);
}
var target = $"{probeOptions.Host}:{probeOptions.Port}";
var sw = Stopwatch.StartNew();
using var deadline = new CancellationTokenSource(TimeSpan.FromSeconds(effectiveSeconds));
using var linked = CancellationTokenSource.CreateLinkedTokenSource(ct, deadline.Token);
using var client = new MqttClientFactory().CreateMqttClient();
MqttClientConnectResult result;
try
{
result = await client.ConnectAsync(clientOptions, linked.Token).ConfigureAwait(false);
}
catch (Exception ex)
{
return new DriverProbeResult(false, Classify(target, ex, ct, deadline, effectiveSeconds), null);
}
try
{
if (result.ResultCode != MqttClientConnectResultCode.Success)
{
return new DriverProbeResult(false, ClassifyRejectedConnack(target, result.ResultCode), null);
}
sw.Stop();
return new DriverProbeResult(true, "MQTT CONNECT OK", sw.Elapsed);
}
finally
{
await DisconnectBestEffortAsync(client, effectiveSeconds).ConfigureAwait(false);
}
}
/// <summary>
/// Maps a broker-rejected CONNACK (<see cref="MqttClientConnectResult.ResultCode"/> not
/// <c>Success</c> — never an exception, see the type remarks) to a targeted, credential-free
/// message.
/// </summary>
/// <remarks>
/// Delegates to <see cref="MqttConnection.DescribeConnackRejection"/> so the "Test connect"
/// button and the running driver report an identical rejection in identical words — they now
/// both classify CONNACK, and disagreeing about it was the exact confusion the Task-13 live
/// gate surfaced.
/// </remarks>
private static string ClassifyRejectedConnack(string target, MqttClientConnectResultCode code)
=> MqttConnection.DescribeConnackRejection(target, code);
/// <summary>
/// Classifies a thrown connect failure. Precedence mirrors
/// <see cref="MqttConnection.Classify"/>: caller cancellation, then the deadline, then the
/// exception's own shape — the deadline check is authoritative regardless of which exception
/// type MQTTnet happened to throw (see the type remarks for why exception type alone is not
/// reliable here).
/// </summary>
private static string Classify(
string target,
Exception ex,
CancellationToken callerToken,
CancellationTokenSource deadline,
int effectiveSeconds)
{
if (callerToken.IsCancellationRequested)
{
return "Probe was cancelled.";
}
if (deadline.IsCancellationRequested)
{
return $"Probe timed out after {effectiveSeconds}s.";
}
// Walk to the root cause — MQTTnet wraps transport/TLS failures one or two levels deep.
var root = ex;
while (root.InnerException is not null)
{
root = root.InnerException;
}
return root switch
{
SocketException se => $"Connect to {target} failed: {se.SocketErrorCode}",
AuthenticationException => $"TLS handshake with {target} failed: {root.Message}",
_ => $"Connect to {target} failed: {root.Message}",
};
}
/// <summary>
/// Best-effort teardown of a session this probe established. Failure here is not the
/// probe's business — the CONNECT outcome already decided Ok/failure — and it must never be
/// allowed to hang past the connect budget.
/// </summary>
private static async Task DisconnectBestEffortAsync(IMqttClient client, int timeoutSeconds)
{
if (!client.IsConnected)
{
return;
}
try
{
using var disconnectDeadline = new CancellationTokenSource(TimeSpan.FromSeconds(timeoutSeconds));
await client.DisconnectAsync(new MqttClientDisconnectOptions(), disconnectDeadline.Token)
.ConfigureAwait(false);
}
catch
{
// Best-effort only.
}
}
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,263 @@
using System.Collections.Frozen;
using System.Diagnostics.CodeAnalysis;
using TahuDataType = Org.Eclipse.Tahu.Protobuf.DataType;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Sparkplug;
/// <summary>
/// One <c>(edgeNode, device)</c> scope's metric catalog, as declared by its most recent
/// NBIRTH/DBIRTH: the alias→metric cache <b>and</b> the stable name→metric index, rebuilt wholesale
/// on every birth. Design doc §3.6 invariants #1 and #2.
/// </summary>
/// <remarks>
/// <para>
/// <b>The alias is a per-birth cache; the NAME is the binding key.</b> Sparkplug's alias
/// mechanism exists to shrink DATA payloads — a BIRTH declares <c>name ↔ alias</c> and every
/// subsequent DATA metric carries only the alias — and the alias is scoped to <i>that one
/// birth</i>. After a rebirth an edge node is entirely free to point alias 5 at a different
/// metric. So an authored tag binds by <see cref="SparkplugMetricBinding.Name"/>; the alias is
/// only ever the lookup that turns an incoming DATA metric back into a name. Binding a tag to
/// an alias instead means a rebirth silently starts routing Pressure into a Temperature tag —
/// <b>good quality, plausible values, completely wrong</b>, and invisible unless something
/// specifically reuses an alias across a rebirth.
/// </para>
/// <para>
/// <b><see cref="RebuildFromBirth"/> REPLACES; it never merges or upserts.</b> The whole
/// snapshot — both indexes — is built fresh and published in one assignment. An alias the new
/// birth did not declare is gone, including when the new birth declares no metrics at all. This
/// is the single behaviour the type exists to guarantee: a merge (or an "optimisation" that
/// skips an empty birth) reintroduces exactly the stale-alias mis-route above.
/// </para>
/// <para>
/// <b>Both indexes are replaced together, from the same birth.</b> They are deliberately held
/// in one type rather than split across an "alias table" and a separate "birth cache": they are
/// two views of one birth, and any seam between them is somewhere the two can be replaced
/// independently and skew — an alias resolving through the new birth into a name catalog still
/// holding the old one.
/// </para>
/// <para>
/// <b>Thread-safety: an immutable snapshot swapped atomically</b>, the same discipline
/// <c>MqttSubscriptionManager.AuthoredTable</c> uses and for the same reason — this is read on
/// MQTTnet's shared dispatcher thread as messages arrive, and a reader must never see a
/// half-rebuilt map. No lock is taken on any path.
/// </para>
/// <para>
/// <b>Last values do not live here.</b> The running last-observed value per tag is
/// <see cref="LastValueCache"/>'s, keyed by RawPath — the identity every publish in this driver
/// is made under. A second copy keyed by metric name would be a duplicate source of truth with
/// a different key, and it would grow with every metric the plant births whether or not any tag
/// was ever authored for it. A birth's own values flow straight through the ingest path to that
/// cache; nothing is retained here.
/// </para>
/// </remarks>
public sealed class AliasTable
{
private volatile Snapshot _snapshot = Snapshot.Empty;
/// <summary>How many distinct metrics the most recent birth declared.</summary>
public int Count => _snapshot.Metrics.Length;
/// <summary>
/// Every metric the most recent birth declared. The consumer's STALE fan-out on a death, and
/// the discovery path's "which metrics does this scope have" question, both read this.
/// </summary>
public IReadOnlyList<SparkplugMetricBinding> Metrics => _snapshot.Metrics;
/// <summary>
/// Replaces this scope's whole catalog with <paramref name="birthMetrics"/> — the metrics of one
/// NBIRTH/DBIRTH payload, exactly as <see cref="SparkplugCodec"/> projected them.
/// </summary>
/// <param name="birthMetrics">The birth payload's metrics.</param>
/// <returns>
/// What was installed, and what was refused — the caller's only window onto a malformed birth,
/// since this type takes no logger.
/// </returns>
/// <remarks>
/// <para>
/// <b>Only an unusable NAME is a rejection.</b> A metric with a blank/absent name cannot be
/// bound to an authored tag by any route, and storing it alias-only would create a binding
/// that can never resolve (or, worse, one that matches a tag whose authored metric name
/// happens to be blank), so it is dropped and counted.
/// </para>
/// <para>
/// <b>A missing datatype is NOT a rejection</b> — it becomes
/// <see cref="TahuDataType.Unknown"/>. Dropping such a metric would turn every DATA message
/// for its alias into an unknown-alias event, which the ingest state machine answers with a
/// rebirth request; the edge node would answer with the same malformed birth, and the pair
/// would loop. Kept as Unknown, the typed layer refuses the value once
/// (<c>ToDriverDataType()</c> returns null for it) and nothing storms.
/// </para>
/// <para>
/// <b>An alias claimed by two metrics in one birth resolves to neither.</b> Any choice
/// between them is a coin flip between two different signals, which is the mis-route this
/// type exists to prevent; the alias is dropped (the consumer sees an unknown alias and can
/// request a rebirth) while both metrics stay bindable by name. A duplicate <i>name</i>, by
/// contrast, is last-wins: the name is the binding key, so evicting it would take the
/// authored tag dark entirely.
/// </para>
/// </remarks>
public BirthApplyResult RebuildFromBirth(IEnumerable<SparkplugMetric> birthMetrics)
{
ArgumentNullException.ThrowIfNull(birthMetrics);
var byName = new Dictionary<string, SparkplugMetricBinding>(StringComparer.Ordinal);
var byAlias = new Dictionary<ulong, SparkplugMetricBinding>();
HashSet<ulong>? collidedAliases = null;
var rejectedUnnamed = 0;
var duplicateNames = 0;
var aliasCollisions = 0;
foreach (var metric in birthMetrics)
{
if (string.IsNullOrWhiteSpace(metric.Name))
{
rejectedUnnamed++;
continue;
}
var binding = new SparkplugMetricBinding(metric.Name, metric.Alias, metric.DataType ?? TahuDataType.Unknown);
if (!byName.TryAdd(binding.Name, binding))
{
duplicateNames++;
byName[binding.Name] = binding;
}
if (binding.Alias is { } alias && !byAlias.TryAdd(alias, binding))
{
aliasCollisions++;
(collidedAliases ??= []).Add(alias);
}
}
if (collidedAliases is not null)
{
foreach (var alias in collidedAliases)
{
byAlias.Remove(alias);
}
}
// A duplicate name is last-wins in `byName`, so the alias index can still hold the losing
// binding — drop those too rather than let an alias resolve to a name that resolves elsewhere.
if (duplicateNames > 0)
{
foreach (var alias in byAlias.Where(kv => !ReferenceEquals(byName[kv.Value.Name], kv.Value))
.Select(kv => kv.Key).ToList())
{
byAlias.Remove(alias);
}
}
// THE replace. One assignment of a fully-built snapshot — no merge, no upsert, and no
// short-circuit for an empty birth (an empty birth legitimately clears the scope).
_snapshot = new Snapshot(
byAlias.ToFrozenDictionary(),
byName.ToFrozenDictionary(StringComparer.Ordinal),
[.. byName.Values]);
return new BirthApplyResult(byName.Count, rejectedUnnamed, aliasCollisions, duplicateNames);
}
/// <summary>
/// Resolves a DATA metric's alias to the metric the <b>most recent</b> birth bound it to, or
/// <see langword="null"/> when this birth declared no such alias.
/// </summary>
/// <param name="alias">The alias carried by an inbound DATA metric.</param>
/// <returns>The binding, or <see langword="null"/> — which the consumer treats as an unknown alias.</returns>
public SparkplugMetricBinding? Resolve(ulong alias) =>
_snapshot.ByAlias.TryGetValue(alias, out var binding) ? binding : null;
/// <summary>The <see cref="Resolve(ulong)"/> try-shape.</summary>
/// <param name="alias">The alias carried by an inbound DATA metric.</param>
/// <param name="binding">The binding when this returns <see langword="true"/>.</param>
/// <returns><see langword="true"/> when the most recent birth declared this alias.</returns>
public bool TryResolve(ulong alias, [NotNullWhen(true)] out SparkplugMetricBinding? binding)
{
binding = Resolve(alias);
return binding is not null;
}
/// <summary>
/// Resolves by the stable metric name — the binding key. Used for a DATA metric that carried its
/// name rather than an alias (aliases are optional in Sparkplug), and to answer "does this scope
/// publish the metric this tag was authored against".
/// </summary>
/// <param name="name">The stable metric name. Matched ordinally; Sparkplug names are case-sensitive.</param>
/// <returns>The binding, or <see langword="null"/>.</returns>
public SparkplugMetricBinding? ResolveByName(string name) =>
!string.IsNullOrEmpty(name) && _snapshot.ByName.TryGetValue(name, out var binding) ? binding : null;
/// <summary>The <see cref="ResolveByName"/> try-shape.</summary>
/// <param name="name">The stable metric name.</param>
/// <param name="binding">The binding when this returns <see langword="true"/>.</param>
/// <returns><see langword="true"/> when the most recent birth declared this metric.</returns>
public bool TryResolveByName(string name, [NotNullWhen(true)] out SparkplugMetricBinding? binding)
{
binding = ResolveByName(name);
return binding is not null;
}
/// <summary>
/// One birth's catalog, immutable and published in a single assignment so the dispatcher thread
/// always reads one whole birth's worth of state.
/// </summary>
/// <param name="ByAlias">Alias → binding, for the birth that declared it. Collided aliases are absent.</param>
/// <param name="ByName">Stable metric name → binding — the index authored tags bind through.</param>
/// <param name="Metrics">Every declared metric, for the STALE fan-out and discovery.</param>
private sealed record Snapshot(
FrozenDictionary<ulong, SparkplugMetricBinding> ByAlias,
FrozenDictionary<string, SparkplugMetricBinding> ByName,
SparkplugMetricBinding[] Metrics)
{
public static readonly Snapshot Empty = new(
FrozenDictionary<ulong, SparkplugMetricBinding>.Empty,
FrozenDictionary<string, SparkplugMetricBinding>.Empty,
[]);
}
}
/// <summary>One metric as its birth declared it: the stable name, the per-birth alias, the datatype.</summary>
/// <param name="Name">
/// The stable metric name — <b>the binding key</b>. An authored tag's <c>metricName</c> matches
/// against this, never against <paramref name="Alias"/>.
/// </param>
/// <param name="Alias">
/// The alias this birth assigned, or <see langword="null"/> when the birth declared none (legal —
/// such a publisher sends names in DATA too). <b>Valid only until the next birth for this scope.</b>
/// </param>
/// <param name="DataType">
/// The datatype the birth declared, or <see cref="TahuDataType.Unknown"/> when it declared none.
/// This is the <i>only</i> place a DATA metric's datatype can come from — DATA carries none.
/// </param>
public sealed record SparkplugMetricBinding(string Name, ulong? Alias, TahuDataType DataType)
{
/// <summary>
/// Applies <see cref="SparkplugCodec.ReinterpretSigned"/> to a raw DATA value using
/// <b>this birth's</b> datatype.
/// </summary>
/// <param name="rawValue">The raw <see cref="SparkplugMetric.Value"/> from a DATA metric.</param>
/// <returns>The signed value for a signed-integer datatype; <paramref name="rawValue"/> otherwise.</returns>
/// <remarks>
/// Exposed here because this is the one object that holds the datatype at the moment a DATA
/// value is being resolved, and forgetting the call is silent: Sparkplug carries Int8/16/32/64
/// as two's complement in an <i>unsigned</i> proto field, so a metric whose value is -42
/// publishes as 4294967254 — plausible, Good quality, and invisible until someone reads a gauge.
/// </remarks>
public object? Reinterpret(object? rawValue) => SparkplugCodec.ReinterpretSigned(rawValue, DataType);
}
/// <summary>What one <see cref="AliasTable.RebuildFromBirth"/> installed, and what it refused.</summary>
/// <param name="Accepted">Distinct named metrics installed.</param>
/// <param name="RejectedUnnamed">Metrics dropped for carrying no usable name.</param>
/// <param name="AliasCollisions">Aliases claimed by more than one metric, and therefore left unresolvable.</param>
/// <param name="DuplicateNames">Metrics whose name a later metric in the same birth overwrote.</param>
public readonly record struct BirthApplyResult(
int Accepted,
int RejectedUnnamed,
int AliasCollisions,
int DuplicateNames)
{
/// <summary>Whether the birth was malformed in any way worth logging.</summary>
public bool HasAnomaly => RejectedUnnamed > 0 || AliasCollisions > 0 || DuplicateNames > 0;
}
@@ -0,0 +1,234 @@
using System.Collections.Concurrent;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Sparkplug;
/// <summary>
/// The driver's live Sparkplug birth state: one <see cref="AliasTable"/> per
/// <see cref="SparkplugScope"/>, with the spec's node→device invalidation rules applied on birth
/// and on death. Design doc §3.6 invariants #2 and #4.
/// </summary>
/// <remarks>
/// <para>
/// <b>Scoping.</b> A scope is the full <c>(group, edgeNode, device?)</c> triple — never the
/// device id alone and never <c>node/device</c> without the group, or one plant's
/// <c>Filler1</c> would answer for another's. A node-scoped birth (NBIRTH) owns the metrics the
/// edge node publishes itself; each DBIRTH owns its device's. They are separate scopes because
/// Sparkplug gives them separate alias spaces: alias 5 on <c>EdgeA</c> and alias 5 on
/// <c>EdgeA/Filler1</c> are unrelated metrics.
/// </para>
/// <para>
/// <b>An NBIRTH invalidates every device under that node.</b> Per the Sparkplug spec an NBIRTH
/// means the edge node (re)started and must re-publish a DBIRTH for each of its devices, so
/// every previously published DBIRTH is void. Keeping a device's alias table across its node's
/// rebirth is the same stale-alias mis-route as
/// <see cref="AliasTable.RebuildFromBirth"/> guards against, one scope up — and it is easier to
/// miss, because nothing about the DBIRTH itself changed. The invalidated devices are returned
/// (<see cref="NodeBirthOutcome.InvalidatedDevices"/>) rather than merely dropped, so the
/// consumer can fan STALE out over their metrics without a lookup against a table that has
/// already been evicted.
/// </para>
/// <para>
/// <b>A death evicts, it does not blank.</b> After an NDEATH/DDEATH the scope is <i>absent</i>,
/// which is what makes a DATA message arriving before the next birth a data-before-birth event
/// (⇒ request a rebirth, §3.6 invariant #3) rather than a silent unknown-alias drop against a
/// table that still looks alive.
/// </para>
/// <para>
/// <b>Thread-safety.</b> A <see cref="ConcurrentDictionary{TKey,TValue}"/> of scopes, each
/// holding an <see cref="AliasTable"/> whose own state is an atomically swapped immutable
/// snapshot. A rebirth <i>rebuilds the existing table in place</i> rather than replacing the
/// table object, so a consumer holding a reference to a scope's table always observes the newest
/// birth instead of quietly reading a detached older one. No lock is taken on any path — this is
/// read (and written) on MQTTnet's shared dispatcher thread.
/// </para>
/// <para>
/// <b>Not a value store.</b> See the remarks on <see cref="AliasTable"/>: running values belong
/// to <see cref="LastValueCache"/>, keyed by RawPath.
/// </para>
/// </remarks>
public sealed class BirthCache
{
private readonly ConcurrentDictionary<SparkplugScope, AliasTable> _scopes = new();
/// <summary>How many scopes currently hold a live birth.</summary>
public int Count => _scopes.Count;
/// <summary>A snapshot of the scopes currently holding a live birth.</summary>
public IReadOnlyList<SparkplugScope> Scopes => [.. _scopes.Keys];
/// <summary>
/// Applies an NBIRTH: rebuilds the edge node's own catalog and <b>invalidates every device</b>
/// under it, per the Sparkplug spec.
/// </summary>
/// <param name="groupId">The Sparkplug group id.</param>
/// <param name="edgeNodeId">The Sparkplug edge-node id.</param>
/// <param name="birthMetrics">The NBIRTH payload's metrics.</param>
/// <returns>The node's apply result plus the devices this birth invalidated, with their metrics.</returns>
public NodeBirthOutcome ApplyNodeBirth(string groupId, string edgeNodeId, IEnumerable<SparkplugMetric> birthMetrics)
{
ArgumentException.ThrowIfNullOrEmpty(groupId);
ArgumentException.ThrowIfNullOrEmpty(edgeNodeId);
ArgumentNullException.ThrowIfNull(birthMetrics);
// Devices first: a device table must never outlive the node birth that voided it, whatever the
// node's own rebuild does.
var invalidated = Evict(scope => scope.DeviceId is not null && scope.IsUnder(groupId, edgeNodeId));
var result = TableFor(new SparkplugScope(groupId, edgeNodeId, null)).RebuildFromBirth(birthMetrics);
return new NodeBirthOutcome(result, invalidated);
}
/// <summary>Applies a DBIRTH: rebuilds exactly that device's catalog, wholesale.</summary>
/// <param name="groupId">The Sparkplug group id.</param>
/// <param name="edgeNodeId">The Sparkplug edge-node id.</param>
/// <param name="deviceId">The Sparkplug device id.</param>
/// <param name="birthMetrics">The DBIRTH payload's metrics.</param>
/// <returns>What the birth installed, and what it refused.</returns>
public BirthApplyResult ApplyDeviceBirth(
string groupId,
string edgeNodeId,
string deviceId,
IEnumerable<SparkplugMetric> birthMetrics)
{
ArgumentException.ThrowIfNullOrEmpty(groupId);
ArgumentException.ThrowIfNullOrEmpty(edgeNodeId);
ArgumentException.ThrowIfNullOrEmpty(deviceId);
ArgumentNullException.ThrowIfNull(birthMetrics);
return TableFor(new SparkplugScope(groupId, edgeNodeId, deviceId)).RebuildFromBirth(birthMetrics);
}
/// <summary>
/// The scope's catalog, or <see langword="null"/> when no live birth has been seen for it —
/// which is the consumer's data-before-birth signal.
/// </summary>
/// <param name="scope">The scope to look up.</param>
/// <returns>The catalog, or <see langword="null"/>.</returns>
public AliasTable? Find(SparkplugScope scope) => _scopes.TryGetValue(scope, out var table) ? table : null;
/// <summary>
/// Applies an NDEATH: forgets the edge node and every device under it, returning what was
/// forgotten so the consumer can emit STALE for those metrics.
/// </summary>
/// <param name="groupId">The Sparkplug group id.</param>
/// <param name="edgeNodeId">The Sparkplug edge-node id.</param>
/// <returns>The evicted scopes and their metrics; empty when nothing was live.</returns>
public IReadOnlyList<SparkplugScopeMetrics> ForgetNode(string groupId, string edgeNodeId)
{
ArgumentException.ThrowIfNullOrEmpty(groupId);
ArgumentException.ThrowIfNullOrEmpty(edgeNodeId);
return Evict(scope => scope.IsUnder(groupId, edgeNodeId));
}
/// <summary>
/// Applies a DDEATH: forgets exactly one device, returning its metrics for the STALE fan-out.
/// The node's own catalog is untouched.
/// </summary>
/// <param name="groupId">The Sparkplug group id.</param>
/// <param name="edgeNodeId">The Sparkplug edge-node id.</param>
/// <param name="deviceId">The Sparkplug device id.</param>
/// <param name="removed">The evicted scope and its metrics when this returns <see langword="true"/>.</param>
/// <returns><see langword="true"/> when the device had a live birth to forget.</returns>
public bool TryForgetDevice(
string groupId,
string edgeNodeId,
string deviceId,
out SparkplugScopeMetrics removed)
{
ArgumentException.ThrowIfNullOrEmpty(groupId);
ArgumentException.ThrowIfNullOrEmpty(edgeNodeId);
ArgumentException.ThrowIfNullOrEmpty(deviceId);
var scope = new SparkplugScope(groupId, edgeNodeId, deviceId);
if (_scopes.TryRemove(scope, out var table))
{
removed = new SparkplugScopeMetrics(scope, table.Metrics);
return true;
}
removed = default;
return false;
}
/// <summary>Forgets every scope — the disconnect/reinitialise reset.</summary>
public void Clear() => _scopes.Clear();
/// <summary>
/// The scope's catalog, created empty on first use. Rebuilds happen <b>in place</b> on the
/// returned instance, so a consumer holding it observes the newest birth rather than a detached
/// older one.
/// </summary>
/// <param name="scope">The scope.</param>
/// <returns>The scope's catalog.</returns>
private AliasTable TableFor(SparkplugScope scope) => _scopes.GetOrAdd(scope, static _ => new AliasTable());
/// <summary>Removes every scope matching <paramref name="predicate"/>, capturing their metrics first.</summary>
/// <param name="predicate">Which scopes to evict.</param>
/// <returns>The evicted scopes and the metrics they held.</returns>
private List<SparkplugScopeMetrics> Evict(Func<SparkplugScope, bool> predicate)
{
List<SparkplugScopeMetrics>? evicted = null;
// ConcurrentDictionary enumeration is weakly consistent, which is exactly right here: this runs
// on the dispatcher thread and only needs to see the scopes that existed when the birth/death
// arrived.
foreach (var scope in _scopes.Keys)
{
if (!predicate(scope) || !_scopes.TryRemove(scope, out var table))
{
continue;
}
(evicted ??= []).Add(new SparkplugScopeMetrics(scope, table.Metrics));
}
return evicted ?? [];
}
}
/// <summary>
/// A Sparkplug metric-namespace scope: an edge node, or one device under it. The identity every
/// birth, death and alias resolution is keyed by.
/// </summary>
/// <param name="GroupId">The Sparkplug group id.</param>
/// <param name="EdgeNodeId">The Sparkplug edge-node id.</param>
/// <param name="DeviceId">The Sparkplug device id, or <see langword="null"/> for the node's own scope.</param>
public readonly record struct SparkplugScope(string GroupId, string EdgeNodeId, string? DeviceId)
{
/// <summary>Whether this is a device scope (a DBIRTH's) rather than an edge node's own.</summary>
public bool IsDevice => DeviceId is not null;
/// <summary>The owning edge node's scope — this instance itself when it is already node-scoped.</summary>
public SparkplugScope NodeScope => DeviceId is null ? this : new SparkplugScope(GroupId, EdgeNodeId, null);
/// <summary>Whether this scope belongs to the given edge node (the node's own scope included).</summary>
/// <param name="groupId">The Sparkplug group id.</param>
/// <param name="edgeNodeId">The Sparkplug edge-node id.</param>
/// <returns><see langword="true"/> when both segments match ordinally.</returns>
public bool IsUnder(string groupId, string edgeNodeId) =>
string.Equals(GroupId, groupId, StringComparison.Ordinal)
&& string.Equals(EdgeNodeId, edgeNodeId, StringComparison.Ordinal);
/// <inheritdoc/>
public override string ToString() =>
DeviceId is null ? $"{GroupId}/{EdgeNodeId}" : $"{GroupId}/{EdgeNodeId}/{DeviceId}";
}
/// <summary>
/// A scope's metrics, captured at the moment it was evicted — the STALE fan-out's input, taken
/// before the eviction so it cannot race a re-birth of the same scope.
/// </summary>
/// <param name="Scope">The scope that was evicted.</param>
/// <param name="Metrics">The metrics its last birth declared.</param>
public readonly record struct SparkplugScopeMetrics(SparkplugScope Scope, IReadOnlyList<SparkplugMetricBinding> Metrics);
/// <summary>The result of an NBIRTH: the node's own apply, plus the devices it invalidated.</summary>
/// <param name="Result">What the node's catalog rebuild installed and refused.</param>
/// <param name="InvalidatedDevices">
/// The device scopes this NBIRTH voided, with the metrics they held — the STALE fan-out's input
/// until each device re-DBIRTHs.
/// </param>
public readonly record struct NodeBirthOutcome(
BirthApplyResult Result,
IReadOnlyList<SparkplugScopeMetrics> InvalidatedDevices);
@@ -0,0 +1,146 @@
using Google.Protobuf;
using Org.Eclipse.Tahu.Protobuf;
using TahuDataType = Org.Eclipse.Tahu.Protobuf.DataType;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Sparkplug;
/// <summary>
/// The one seam through which a Sparkplug NCMD is published — <see cref="MqttConnection"/> is the
/// production implementation once it grows a publish leg (Task 21); tests substitute a fake so the
/// encode/topic/QoS/retain contract is exercisable without a broker. Deliberately narrower than
/// "publish anything": a Sparkplug driver's only outbound traffic is NCMD/DCMD, so this is not a
/// general MQTT client abstraction in waiting.
/// </summary>
public interface IMqttPublishTransport
{
/// <summary>
/// Publishes one application message. Implementations must be bounded — a broker that accepts
/// PUBLISH and never completes it (QoS &gt; 0 awaiting PUBACK) must fail at a deadline, not
/// hang; <see cref="RebirthRequester.RequestAsync"/> supplies one, but an implementation used
/// directly by a future caller must not rely on that.
/// </summary>
/// <param name="topic">The concrete topic to publish on.</param>
/// <param name="payload">The message body.</param>
/// <param name="qos">Requested QoS, 02.</param>
/// <param name="retain">The MQTT <c>retain</c> flag.</param>
/// <param name="cancellationToken">Caller cancellation.</param>
Task PublishAsync(string topic, byte[] payload, int qos, bool retain, CancellationToken cancellationToken);
}
/// <summary>
/// Encodes and publishes a Sparkplug B rebirth request — an NCMD carrying the well-known
/// <c>Node Control/Rebirth</c> boolean metric, addressed to one edge node's command topic. See
/// Sparkplug B v3.0 spec §6/§7 and design doc §3.6.
/// </summary>
/// <remarks>
/// <para>
/// <b><see cref="Build"/> is pure.</b> It does no I/O and touches no transport, so it is
/// trivially unit-testable and safe to call from anywhere (the ingest state machine on a
/// detected gap, the browser on an operator click) without pulling in a live connection. The
/// bounded, transport-facing publish is the separate <see cref="RequestAsync"/> overload.
/// </para>
/// <para>
/// <b>No <c>seq</c>.</b> Sparkplug's sequence-number stream (§6.4.1) belongs to the messages an
/// edge node itself publishes (NBIRTH/NDATA/DBIRTH/DDATA/NDEATH/DDEATH) — it is how a host
/// detects a gap in <i>that node's</i> outbound stream. An NCMD flows the other direction, host
/// → node, and is not a member of the node's sequence at all: stamping one on here would not be
/// "wrong per spec" so much as meaningless, and a strict edge-node implementation that validated
/// it against its own counter would have grounds to reject the command outright.
/// </para>
/// <para>
/// <b>QoS 0, retain <see langword="false"/> — always.</b> NCMD is ordinary Sparkplug command
/// traffic, published at the same QoS 0 the spec uses for NDATA/DDATA/NBIRTH/DBIRTH (only
/// NDEATH — the broker-issued Will — and STATE use QoS 1). <see langword="retain"/> is the one
/// that matters operationally: a retained NCMD would be replayed by the broker to the edge node
/// on <i>every</i> future subscribe — including the node's own reconnect — driving it into a
/// permanent rebirth loop from a single stale command. See the carry-forward landmine on
/// <c>MqttDriverBrowser</c>/NDEATH-as-Will for the sibling failure shape this type must not
/// introduce on the publish side.
/// </para>
/// <para>
/// <b><c>Datatype</c> and both timestamps are set.</b> <c>Node Control/Rebirth</c> is Boolean;
/// setting <see cref="Payload.Types.Metric.Datatype"/> explicitly (rather than leaving it for the
/// node to infer) and stamping both the metric and payload <c>timestamp</c> matches what a
/// well-behaved command payload carries per spec §6.4.13 and what the Task-25 simulator (and any
/// real Cirrus-Link-derived edge node) expects to validate against. An edge node that treats a
/// missing/ambiguous field as malformed and silently drops the command would otherwise leave the
/// driver retrying forever with no visible error — the whole point of a rebirth request is to
/// recover from exactly that kind of silent desync, so the request itself must not risk being
/// the next cause of one.
/// </para>
/// </remarks>
public static class RebirthRequester
{
/// <summary>The well-known Sparkplug B node-control metric name that triggers a rebirth.</summary>
public const string RebirthMetricName = "Node Control/Rebirth";
/// <summary>
/// Default bounded deadline for <see cref="RequestAsync"/> when the caller does not supply one.
/// Deliberately this type's own value rather than <c>MqttDriverOptions.ConnectTimeoutSeconds</c>
/// — that knob already governs every MQTTnet connect/subscribe operation, and repurposing it
/// would couple the rebirth-publish budget to the connect budget in a way neither side could
/// tune independently (the same reasoning <c>MqttConnection.SubscribeAsync</c> documents for its
/// own deadline).
/// </summary>
public static readonly TimeSpan DefaultRequestTimeout = TimeSpan.FromSeconds(5);
/// <summary>
/// Builds the wire bytes and target topic for a rebirth-request NCMD, without publishing it.
/// </summary>
/// <param name="groupId">The Sparkplug group id.</param>
/// <param name="edgeNodeId">The Sparkplug edge-node id to address the command to.</param>
/// <returns>The NCMD topic (<c>spBv1.0/{groupId}/NCMD/{edgeNodeId}</c>) and its encoded payload.</returns>
/// <exception cref="ArgumentException"><paramref name="groupId"/>/<paramref name="edgeNodeId"/> is null or empty.</exception>
public static (string Topic, byte[] Bytes) Build(string groupId, string edgeNodeId)
{
// Validates both ids and does the topic assembly — see the "don't hand-concatenate" remark on
// SparkplugTopic.Format itself.
var topic = SparkplugTopic.Format(groupId, SparkplugMessageType.NCMD, edgeNodeId);
var timestampMs = (ulong)DateTimeOffset.UtcNow.ToUnixTimeMilliseconds();
var payload = new Payload { Timestamp = timestampMs };
payload.Metrics.Add(new Payload.Types.Metric
{
Name = RebirthMetricName,
Datatype = (uint)TahuDataType.Boolean,
Timestamp = timestampMs,
BooleanValue = true,
});
return (topic, payload.ToByteArray());
}
/// <summary>
/// Builds and publishes a rebirth-request NCMD under a bounded deadline. QoS 0, retain
/// <see langword="false"/> — see the type remarks for why both are fixed rather than
/// configurable.
/// </summary>
/// <param name="transport">The publish seam — the live connection, or a test fake.</param>
/// <param name="groupId">The Sparkplug group id.</param>
/// <param name="edgeNodeId">The Sparkplug edge-node id to address the command to.</param>
/// <param name="cancellationToken">Caller cancellation; linked with the publish deadline.</param>
/// <param name="timeout">
/// The publish deadline. Defaults to <see cref="DefaultRequestTimeout"/> when omitted or
/// non-positive.
/// </param>
/// <exception cref="ArgumentException"><paramref name="groupId"/>/<paramref name="edgeNodeId"/> is null or empty.</exception>
/// <exception cref="ArgumentNullException"><paramref name="transport"/> is null.</exception>
public static async Task RequestAsync(
IMqttPublishTransport transport,
string groupId,
string edgeNodeId,
CancellationToken cancellationToken,
TimeSpan? timeout = null)
{
ArgumentNullException.ThrowIfNull(transport);
var (topic, bytes) = Build(groupId, edgeNodeId);
var deadlineSpan = timeout is { } t && t > TimeSpan.Zero ? t : DefaultRequestTimeout;
using var deadline = new CancellationTokenSource(deadlineSpan);
using var linked = CancellationTokenSource.CreateLinkedTokenSource(cancellationToken, deadline.Token);
await transport.PublishAsync(topic, bytes, qos: 0, retain: false, linked.Token).ConfigureAwait(false);
}
}
@@ -0,0 +1,386 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Sparkplug;
/// <summary>
/// Per-edge-node Sparkplug B stream-continuity state: the wrapping <c>seq</c> counter that says
/// whether the driver has missed a message, and the <c>bdSeq</c> session token that ties an
/// NDEATH to the NBIRTH it belongs to. Design doc §3.6 invariant #3.
/// </summary>
/// <remarks>
/// <para>
/// <b>One instance per edge node — this object <i>is</i> the edge node's stream state.</b> It
/// holds no key and no map. Sparkplug numbers messages per edge node (an edge node's devices
/// share its counter, so a DDATA advances the same sequence an NDATA does), and the ingest
/// state machine already keeps per-scope state for the alias table and birth cache; giving
/// this type a second internal dictionary would mean two maps whose lifetimes must be kept in
/// step by hand, and a tracker that outlived a removed edge node would go on answering
/// questions about a stream nobody is reading.
/// </para>
/// <para>
/// <b>This type detects; it does not react.</b> It never publishes, never requests a rebirth
/// and never emits a quality change — the ingest state machine reads its verdicts and decides.
/// The split matters because "was a message lost" is a fact about the wire, while "request a
/// rebirth now" is a policy gated on <c>requestRebirthOnGap</c>.
/// </para>
/// <para>
/// <b>Thread-safe under a plain monitor.</b> Sparkplug messages arrive on MQTTnet's shared
/// dispatcher thread while the reconnect path can call <see cref="Reset"/> from another, and
/// every operation here is a read-modify-write of two coupled fields — the immutable-snapshot
/// discipline <c>MqttSubscriptionManager</c> uses fits a set that is swapped wholesale, not a
/// counter that is advanced. The critical sections are a handful of instructions each and
/// nothing inside them can block, so the simple lock this repo uses elsewhere
/// (<c>MqttDriver._hostLock</c>) is the right shape.
/// </para>
/// <para>
/// <b>Nothing here throws, for any input.</b> Same reasoning as
/// <see cref="SparkplugCodec"/>: this sits behind an unauthenticated firehose on the
/// dispatcher thread, so a spec-violating publisher gets a verdict, never an exception that
/// would stall delivery for every subscription on the connection.
/// </para>
/// </remarks>
public sealed class SequenceTracker
{
/// <summary>
/// The largest legal Sparkplug <c>seq</c>. The counter is a byte on the wire in spirit but a
/// <c>uint64</c> in the schema, so the range check lives here — see <see cref="Accept"/>.
/// </summary>
public const ulong MaxSeq = 255;
private readonly object _gate = new();
/// <summary>The last credible sequence number, meaningful only when <see cref="_hasLastSeq"/>.</summary>
private byte _lastSeq;
/// <summary>Whether any credible sequence number has been observed since the last reset.</summary>
private bool _hasLastSeq;
/// <summary>Whether the current baseline was established by a birth rather than adopted mid-stream.</summary>
private bool _birthSynchronized;
/// <summary>The current session's <c>bdSeq</c>, or null when no birth has supplied one.</summary>
private ulong? _birthBdSeq;
/// <summary>
/// Whether the current sequence baseline came from a birth. <see langword="false"/> also after
/// a first message that was <i>not</i> a birth — the driver joined mid-stream and holds no
/// alias table, which is a different situation from a gap even though both call for a rebirth.
/// </summary>
public bool IsBirthSynchronized
{
get { lock (_gate) return _birthSynchronized; }
}
/// <summary>
/// The last credible sequence number observed, or <see langword="null"/> when none has been —
/// before the first message, or after <see cref="Reset"/>.
/// </summary>
public byte? LastSeq
{
get { lock (_gate) return _hasLastSeq ? _lastSeq : null; }
}
/// <summary>
/// The <c>bdSeq</c> of the most recent birth, or <see langword="null"/> when no birth has
/// supplied one. This is the token <see cref="IsDeathForCurrentSession"/> matches against.
/// </summary>
public ulong? BirthBdSeq
{
get { lock (_gate) return _birthBdSeq; }
}
/// <summary>
/// Records an <b>NBIRTH</b>: restarts the sequence at the birth's <c>seq</c> and adopts its
/// <c>bdSeq</c> as the current session token.
/// </summary>
/// <param name="seq">
/// The birth payload's <c>seq</c> — <see cref="SparkplugPayload.Seq"/>, which is
/// <see langword="null"/> when the message carried none.
/// </param>
/// <param name="bdSeq">
/// The birth's session token, read from its <c>bdSeq</c> metric with
/// <see cref="TryReadBdSeq"/>; <see langword="null"/> when the birth carried none.
/// </param>
/// <returns>
/// <see langword="true"/> when the birth established a usable sequence baseline;
/// <see langword="false"/> when its <c>seq</c> was absent or out of range.
/// </returns>
/// <remarks>
/// <para>
/// <b>An NBIRTH is never a gap.</b> It restarts the sequence by definition, so it must not
/// be run through <see cref="Accept"/> — doing so would flag a gap on the very message
/// that just resynchronized everything, and answer a birth with a request for another one.
/// </para>
/// <para>
/// <b>A DBIRTH must NOT be offered here — it belongs to <see cref="Accept"/>.</b> Only the
/// NBIRTH restarts the sequence; a DBIRTH carries the next number in the edge node's
/// ongoing stream, and only the NBIRTH (with the NDEATH) carries <c>bdSeq</c> at all.
/// Routing a DBIRTH here would do two wrong things at once: silently rebase the sequence
/// mid-stream, and — because the DBIRTH has no <c>bdSeq</c> to pass — <b>wipe the node's
/// session token to null</b>. A stale Last Will arriving afterwards would then find
/// nothing to be compared against, fall through the fail-toward-stale rule in
/// <see cref="IsDeathForCurrentSession"/>, and kill the live node this type exists to
/// protect. The method is named for the node deliberately, so that call site reads wrong.
/// </para>
/// <para>
/// <b>An in-range <c>seq</c> other than the spec-required 0 is adopted, not refused.</b>
/// Refusing it would leave the tracker unsynchronized against a publisher whose stream is
/// otherwise perfectly followable, and every message that followed would demand a fresh
/// rebirth — the storm this type exists to prevent, arrived at by being strict.
/// </para>
/// <para>
/// <b>The <c>bdSeq</c> is recorded even when the <c>seq</c> was unusable.</b> Session
/// identity and position-in-stream are independent facts; discarding the tie because the
/// sequence number was garbled would leave the next NDEATH unmatchable, which is the one
/// thing that lets a stale will kill a live node.
/// </para>
/// </remarks>
public bool AcceptNodeBirth(ulong? seq, ulong? bdSeq = null)
{
lock (_gate)
{
_birthBdSeq = bdSeq;
if (seq is not { } value || value > MaxSeq)
{
_hasLastSeq = false;
_birthSynchronized = false;
return false;
}
_lastSeq = (byte)value;
_hasLastSeq = true;
_birthSynchronized = true;
return true;
}
}
/// <summary>
/// Offers the next sequenced message's <c>seq</c> — <b>DBIRTH</b>, NDATA, DDATA or DDEATH —
/// and reports whether the stream is still contiguous.
/// </summary>
/// <param name="seq">
/// The payload's <c>seq</c> — <see cref="SparkplugPayload.Seq"/>, which is
/// <see langword="null"/> when the message carried none.
/// </param>
/// <returns>
/// <see langword="true"/> when this <c>seq</c> is the one that follows the last;
/// <see langword="false"/> when a message was missed, when the value is out of range, or when
/// it is absent. A <see langword="false"/> is the caller's cue to request a rebirth.
/// </returns>
/// <remarks>
/// <para>
/// <b>The wrap is the whole point: next-expected is <c>(last + 1) &amp; 0xFF</c>, so
/// 255 → 0 is the sequence continuing, not a gap.</b> Getting the boundary wrong fails
/// loudly in one direction and silently in the other. Read the wrap as a gap and every
/// edge node is asked to rebirth once per 256 messages, so a busy node spends its life
/// republishing births instead of publishing data. Miss a real gap and the driver serves
/// values it knows are behind the device, at Good quality, indefinitely.
/// </para>
/// <para>
/// <b>A gap is reported once, then the tracker resynchronizes onto the observed value.</b>
/// Holding the baseline at the pre-gap number would make every following message report a
/// gap too, so a single lost message would become a permanent rebirth storm.
/// </para>
/// <para>
/// <b>An out-of-range <c>seq</c> is refused, never masked into range.</b> The codec hands
/// this over as a <see cref="ulong"/> precisely so a spec-violating publisher's value is
/// not silently truncated into a plausible-looking one, and masking would finish the job
/// it declined to do: with the baseline at 255, a <c>seq</c> of 256 masks to 0 — exactly
/// the value the wrap expects next — and the bogus message would be accepted as
/// contiguous. Refused values do not become the baseline either, so the next genuine
/// message is still measured against the last credible one.
/// </para>
/// <para>
/// <b>The first message is not a gap</b> — there is nothing to measure it against. It is
/// adopted as the baseline, and <see cref="IsBirthSynchronized"/> stays
/// <see langword="false"/> to record that the driver joined mid-stream.
/// </para>
/// <para>
/// <b>NDEATH must not be offered here.</b> An NDEATH is the broker-published Last Will and
/// carries no <c>seq</c> at all; feeding it in would report a gap on a message that never
/// had a place in the sequence. It is tied to its birth by
/// <see cref="IsDeathForCurrentSession"/> instead. A DDEATH is published by the edge node
/// and <i>is</i> sequenced, so it does belong here.
/// </para>
/// </remarks>
public bool Accept(ulong? seq)
{
lock (_gate)
{
// Refused without becoming the baseline: see the out-of-range remarks above.
if (seq is not { } value || value > MaxSeq)
{
return false;
}
var current = (byte)value;
if (!_hasLastSeq)
{
_lastSeq = current;
_hasLastSeq = true;
return true;
}
var expected = (byte)((_lastSeq + 1) & 0xFF);
// Resynchronize even on a gap — one lost message must not become a permanent storm.
_lastSeq = current;
return current == expected;
}
}
/// <summary>
/// Reports whether an NDEATH belongs to the session currently believed alive — the
/// <c>bdSeq</c> tie of §3.6 invariant #3.
/// </summary>
/// <param name="deathBdSeq">
/// The death payload's <c>bdSeq</c>, read with <see cref="TryReadBdSeq"/>;
/// <see langword="null"/> when it carried none.
/// </param>
/// <returns>
/// <see langword="true"/> when the death applies and the node's metrics should go stale;
/// <see langword="false"/> when it is a previous session's will and must be ignored.
/// </returns>
/// <remarks>
/// <para>
/// <b>What this prevents.</b> An edge node's connection drops, it reconnects and publishes
/// a fresh NBIRTH — and only <i>then</i> does the broker notice the old connection died
/// and deliver its retained Last Will. The NDEATH arrives after the NBIRTH, carrying the
/// previous session's <c>bdSeq</c>. Without this comparison that stale will marks a live,
/// freshly-born node dead and drives every one of its tags to Bad quality, with no further
/// message coming to correct it until the node happens to rebirth again.
/// </para>
/// <para>
/// <b>Unknowns fail toward stale, deliberately.</b> A death arriving before any birth was
/// seen, a death carrying no <c>bdSeq</c>, or a birth that carried none all resolve to
/// <see langword="true"/>. The two error directions are not symmetric: acting on a death
/// wrongly marks tags stale until the next birth restores them (§3.6 invariant #4, a
/// self-correcting error an operator can see), while ignoring a real death leaves a dead
/// node's last values flowing at Good quality forever. Only a <i>positive mismatch</i> —
/// two known, different tokens — is evidence enough to discard a death.
/// </para>
/// <para>
/// <b>Pure.</b> It answers a question and changes nothing, so the caller can ask it while
/// logging or deciding. A death that <i>is</i> acted on needs no reset here: the node is
/// offline, and the next birth calls <see cref="AcceptNodeBirth"/> which restarts the sequence
/// anyway.
/// </para>
/// </remarks>
public bool IsDeathForCurrentSession(ulong? deathBdSeq)
{
lock (_gate)
{
if (_birthBdSeq is not { } birth || deathBdSeq is not { } death)
{
return true;
}
return birth == death;
}
}
/// <summary>
/// Returns the tracker to its virgin state: no sequence baseline, not birth-synchronized, no
/// session token.
/// </summary>
/// <remarks>
/// The reconnect seam (§3.6 invariant #5 — late join). After a dropped connection the driver
/// has missed an unknown number of messages, so the baseline it was holding is no longer
/// evidence of anything: carrying it across would let the stream look contiguous purely
/// because the edge node happened to publish the right number of messages while the driver was
/// away.
/// </remarks>
public void Reset()
{
lock (_gate)
{
_lastSeq = 0;
_hasLastSeq = false;
_birthSynchronized = false;
_birthBdSeq = null;
}
}
/// <summary>
/// Reads the <c>bdSeq</c> session token out of a decoded NBIRTH or NDEATH payload.
/// </summary>
/// <param name="payload">The decoded payload; <see langword="null"/> and invalid are both handled.</param>
/// <param name="bdSeq">The token on success; <c>0</c> otherwise.</param>
/// <returns><see langword="true"/> when the payload carried a usable <c>bdSeq</c> metric.</returns>
/// <remarks>
/// <para>
/// <b><c>bdSeq</c> is not a top-level payload field</b> — unlike <c>seq</c> it rides as an
/// ordinary metric named <c>bdSeq</c>, which is why reading it is a helper rather than a
/// property. The name is matched exactly (ordinal): the spec fixes it, so a case variant
/// is far more likely to be a different metric than a typo of this one.
/// </para>
/// <para>
/// <b>Signed values are reinterpreted before they are judged.</b> Sparkplug carries
/// <c>Int64</c> as two's complement in an unsigned field, so a negative <c>bdSeq</c>
/// reaches the projection as an enormous <see cref="ulong"/>. Read raw it would become a
/// plausible-looking session token minted out of nonsense; run through
/// <see cref="SparkplugCodec.ReinterpretSigned"/> it is visibly negative and refused.
/// </para>
/// <para>
/// <b>The value is treated as an opaque token, not as a 0255 counter.</b> Its only job
/// here is equality against the birth's, and a range check could never make a match more
/// or less correct — it could only discard the one piece of evidence tying a death to its
/// birth. This is the opposite call from <c>seq</c>, whose range is load-bearing precisely
/// because its arithmetic wraps.
/// </para>
/// </remarks>
public static bool TryReadBdSeq(SparkplugPayload? payload, out ulong bdSeq)
{
bdSeq = 0;
if (payload is not { IsValid: true })
{
return false;
}
foreach (var metric in payload.Metrics)
{
if (!string.Equals(metric.Name, "bdSeq", StringComparison.Ordinal))
{
continue;
}
if (metric.ValueKind != SparkplugValueKind.Scalar)
{
return false;
}
var value = metric.DataType is { } datatype
? SparkplugCodec.ReinterpretSigned(metric.Value, datatype)
: metric.Value;
switch (value)
{
case ulong u:
bdSeq = u;
return true;
case uint u:
bdSeq = u;
return true;
case long l when l >= 0:
bdSeq = (ulong)l;
return true;
case int i when i >= 0:
bdSeq = (ulong)i;
return true;
case short s when s >= 0:
bdSeq = (ulong)s;
return true;
case sbyte sb when sb >= 0:
bdSeq = (ulong)sb;
return true;
default:
// Negative, or a wire type no integer token can come from.
return false;
}
}
return false;
}
}
@@ -0,0 +1,313 @@
using Google.Protobuf;
using Org.Eclipse.Tahu.Protobuf;
using TahuDataType = Org.Eclipse.Tahu.Protobuf.DataType;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Sparkplug;
/// <summary>
/// Decodes Sparkplug-B wire bytes into <see cref="SparkplugPayload"/> — the driver-side projection
/// the ingest state machine consumes. Decode only; the NCMD <i>encode</i> path lives in
/// <c>RebirthRequester</c>.
/// </summary>
/// <remarks>
/// <para>
/// <b>Nothing here throws, for any input.</b> This runs on MQTTnet's shared dispatcher thread,
/// behind an unauthenticated firehose the plant's edge nodes publish into. An escaping
/// exception would not degrade one tag — it would stall or kill delivery for every
/// subscription on the connection. Garbage bytes, a truncated body, a zero-length payload, a
/// valid protobuf of some other schema and a pathologically nested Template all resolve to a
/// verdict: <see cref="TryDecode"/> returns <see langword="false"/> and
/// <see cref="Decode"/> returns <see cref="SparkplugPayload.Invalid"/>.
/// </para>
/// <para>
/// <b>Zero-length input is invalid, not empty.</b> Protobuf would happily parse zero bytes as
/// "a Payload with every field defaulted", but in Sparkplug an empty MQTT body is never a
/// legitimate message — treating it as a well-formed payload carrying no metrics is how a
/// truncated-to-nothing body gets mistaken for a real one. It is reported as invalid.
/// </para>
/// <para>
/// <b>Explicit presence is preserved, everywhere.</b> The vendored schema is proto2 precisely
/// so absence and zero stay distinguishable, and every projection below reads
/// <c>Has{Seq,Name,Alias,Datatype,Timestamp,IsNull}</c> rather than testing a value against its
/// default. Two cases make this load-bearing rather than pedantic: a Sparkplug NBIRTH is
/// REQUIRED to carry <c>seq = 0</c>, and every DATA metric after a birth carries an alias with
/// <b>no</b> name and <b>no</b> datatype. A zero-check decoder reports "no sequence" for every
/// birth and "" for every DATA metric name.
/// </para>
/// <para>
/// <b>Values are projected raw — nothing is coerced or reinterpreted here.</b> A metric's value
/// arrives as the CLR type of the wire field that carried it (see
/// <see cref="SparkplugMetric.Value"/>); coercion against the authored tag's declared
/// <c>DriverDataType</c> is the consumer's job, and follows this driver's standing rule that a
/// value which does not fit is refused rather than silently converted. The one wire-level
/// translation this type does offer is <see cref="ReinterpretSigned"/>, because Sparkplug's
/// two's-complement-in-an-unsigned-field encoding of signed integers is a property of the
/// <i>wire</i>, not of the tag.
/// </para>
/// <para>
/// <b><c>DataSet</c> and <c>Template</c> metrics are out of scope for v1.</b> They decode to
/// <see cref="SparkplugValueKind.Unsupported"/> with a null value — deliberately visible, so a
/// consumer can warn about a metric it cannot serve instead of silently treating it as one
/// that was never published.
/// </para>
/// </remarks>
public static class SparkplugCodec
{
/// <summary>Decodes Sparkplug-B wire bytes, reporting success rather than throwing.</summary>
/// <param name="wire">The MQTT message body.</param>
/// <param name="payload">
/// The decoded payload on success; <see cref="SparkplugPayload.Invalid"/> otherwise. Never
/// <see langword="null"/>.
/// </param>
/// <returns>
/// <see langword="true"/> when the bytes parsed as a Sparkplug-B <c>Payload</c>. Protobuf is a
/// permissive, self-describing-only-by-convention format, so this means "parsed" rather than
/// "was genuinely produced by a Sparkplug edge node" — unknown fields are preserved by
/// Google.Protobuf and simply ignored here.
/// </returns>
public static bool TryDecode(ReadOnlySpan<byte> wire, out SparkplugPayload payload)
{
payload = SparkplugPayload.Invalid;
if (wire.IsEmpty)
{
return false;
}
try
{
var proto = Payload.Parser.ParseFrom(wire);
var metrics = proto.Metrics.Count == 0
? []
: new SparkplugMetric[proto.Metrics.Count];
for (var i = 0; i < proto.Metrics.Count; i++)
{
metrics[i] = ProjectMetric(proto.Metrics[i]);
}
payload = new SparkplugPayload(
IsValid: true,
Seq: proto.HasSeq ? proto.Seq : null,
TimestampMs: proto.HasTimestamp ? proto.Timestamp : null,
Metrics: metrics);
return true;
}
catch (Exception)
{
// Deliberately broad, and deliberately spanning the projection as well as the parse.
// InvalidProtocolBufferException covers malformed, truncated and over-nested input, but
// this is the dispatcher-thread boundary and the projection walks an attacker-shaped
// object graph: the contract is a verdict for EVERY input, so a guard that stopped at
// the parse would be a contract with a hole in it. Narrowing the catch would trade the
// guarantee for a taxonomy nobody downstream can act on.
payload = SparkplugPayload.Invalid;
return false;
}
}
/// <summary>
/// Decodes Sparkplug-B wire bytes, returning <see cref="SparkplugPayload.Invalid"/> when they
/// cannot be decoded.
/// </summary>
/// <param name="wire">The MQTT message body.</param>
/// <returns>The decoded payload — never <see langword="null"/>.</returns>
/// <remarks>
/// The convenience shape. <b>Check <see cref="SparkplugPayload.IsValid"/>:</b> an undecodable
/// body and a genuinely metric-less one both present as an empty
/// <see cref="SparkplugPayload.Metrics"/>, and only the flag tells them apart. Prefer
/// <see cref="TryDecode"/> where the verdict drives control flow.
/// </remarks>
public static SparkplugPayload Decode(ReadOnlySpan<byte> wire) =>
TryDecode(wire, out var payload) ? payload : SparkplugPayload.Invalid;
/// <summary>
/// Reinterprets a raw metric value as the signed integer Sparkplug encoded it as, when its
/// datatype is a signed integer type; returns the value unchanged for every other datatype.
/// </summary>
/// <param name="value">
/// A <see cref="SparkplugMetric.Value"/> — a <see cref="uint"/> for the 8/16/32-bit arms, a
/// <see cref="ulong"/> for the 64-bit arm.
/// </param>
/// <param name="datatype">The metric's Sparkplug datatype, from its birth or its own field.</param>
/// <returns>
/// <see cref="sbyte"/> / <see cref="short"/> / <see cref="int"/> / <see cref="long"/> for the
/// signed integer datatypes; <paramref name="value"/> untouched otherwise (including when it is
/// <see langword="null"/> or not the wire type the datatype implies).
/// </returns>
/// <remarks>
/// Sparkplug carries <c>Int8</c>/<c>Int16</c>/<c>Int32</c> in the <b>unsigned</b>
/// <c>int_value</c> field and <c>Int64</c> in the unsigned <c>long_value</c> field, as two's
/// complement. Handing the raw field to a consumer publishes <c>4294967254</c> for a tag whose
/// value is <c>-42</c> — a wrong value that looks entirely plausible, arrives with Good
/// quality, and is invisible until someone reads a gauge. This is the single call that undoes
/// it, and it is separate from the decode itself because a DATA metric carries no datatype of
/// its own: only the consumer, holding the alias table built from the birth, knows which
/// datatype applies.
/// </remarks>
public static object? ReinterpretSigned(object? value, TahuDataType datatype) => datatype switch
{
TahuDataType.Int8 when value is uint raw => unchecked((sbyte)raw),
TahuDataType.Int16 when value is uint raw => unchecked((short)raw),
TahuDataType.Int32 when value is uint raw => unchecked((int)raw),
TahuDataType.Int64 when value is ulong raw => unchecked((long)raw),
_ => value,
};
/// <summary>Projects one generated metric onto the driver-side shape, presence intact.</summary>
/// <param name="metric">The generated metric.</param>
/// <returns>The projection.</returns>
private static SparkplugMetric ProjectMetric(Payload.Types.Metric metric)
{
var (kind, value) = ProjectValue(metric);
return new SparkplugMetric(
Name: metric.HasName ? metric.Name : null,
Alias: metric.HasAlias ? metric.Alias : null,
// Cast, never filter: an index this build's enum does not define is a future Sparkplug
// revision or a misbehaving publisher, and dropping it would deny the mapping layer the
// only evidence it has.
DataType: metric.HasDatatype ? (TahuDataType)metric.Datatype : null,
TimestampMs: metric.HasTimestamp ? metric.Timestamp : null,
ValueKind: kind,
Value: value);
}
/// <summary>Resolves a metric's value <c>oneof</c> to a kind plus a raw CLR value.</summary>
/// <param name="metric">The generated metric.</param>
/// <returns>The value kind and the projected value.</returns>
/// <remarks>
/// <c>is_null</c> wins over the <c>oneof</c>: the Sparkplug spec's whole reason for the field is
/// that some datatypes have no spare sentinel, so a publisher that sets it means "null" even if
/// a value field is also populated.
/// </remarks>
private static (SparkplugValueKind Kind, object? Value) ProjectValue(Payload.Types.Metric metric)
{
if (metric.HasIsNull && metric.IsNull)
{
return (SparkplugValueKind.Null, null);
}
return metric.ValueCase switch
{
Payload.Types.Metric.ValueOneofCase.IntValue => (SparkplugValueKind.Scalar, metric.IntValue),
Payload.Types.Metric.ValueOneofCase.LongValue => (SparkplugValueKind.Scalar, metric.LongValue),
Payload.Types.Metric.ValueOneofCase.FloatValue => (SparkplugValueKind.Scalar, metric.FloatValue),
Payload.Types.Metric.ValueOneofCase.DoubleValue => (SparkplugValueKind.Scalar, metric.DoubleValue),
Payload.Types.Metric.ValueOneofCase.BooleanValue => (SparkplugValueKind.Scalar, metric.BooleanValue),
Payload.Types.Metric.ValueOneofCase.StringValue => (SparkplugValueKind.Scalar, metric.StringValue),
// Copied out of the ByteString: the projection must outlive the parsed message.
Payload.Types.Metric.ValueOneofCase.BytesValue =>
(SparkplugValueKind.Scalar, metric.BytesValue.ToByteArray()),
// v1 scope boundary — decoded as a visible refusal, not as an exception or a silent drop.
Payload.Types.Metric.ValueOneofCase.DatasetValue => (SparkplugValueKind.Unsupported, null),
Payload.Types.Metric.ValueOneofCase.TemplateValue => (SparkplugValueKind.Unsupported, null),
Payload.Types.Metric.ValueOneofCase.ExtensionValue => (SparkplugValueKind.Unsupported, null),
_ => (SparkplugValueKind.Absent, null),
};
}
}
/// <summary>
/// A decoded Sparkplug-B payload: the driver-side projection of the generated <c>Payload</c>,
/// carrying only what the ingest state machine consumes.
/// </summary>
/// <param name="IsValid">
/// <see langword="false"/> for a payload that could not be decoded. An invalid payload also has an
/// empty <paramref name="Metrics"/> list, so <b>this flag is the only thing distinguishing an
/// undecodable body from a legitimately metric-less one</b>.
/// </param>
/// <param name="Seq">
/// The payload sequence number, or <see langword="null"/> when the message carried none. Kept as
/// the wire's <see cref="ulong"/> rather than narrowed to a byte: the Sparkplug range is 0255, but
/// a publisher that violates it is reporting a fact the consumer should be able to see and reject,
/// not one this layer should silently truncate into a plausible-looking sequence number.
/// </param>
/// <param name="TimestampMs">The payload timestamp in Sparkplug epoch milliseconds, or null when absent.</param>
/// <param name="Metrics">The payload's metrics, in wire order. Never <see langword="null"/>.</param>
public sealed record SparkplugPayload(
bool IsValid,
ulong? Seq,
ulong? TimestampMs,
IReadOnlyList<SparkplugMetric> Metrics)
{
/// <summary>The shared "could not decode this" result.</summary>
public static readonly SparkplugPayload Invalid = new(false, null, null, []);
}
/// <summary>One decoded Sparkplug metric.</summary>
/// <param name="Name">
/// The metric's stable name, or <see langword="null"/> when the message omitted it — which every
/// real DATA metric after a birth does, carrying only <paramref name="Alias"/>. Distinguishing
/// this from an empty string is the reason the vendored schema is proto2.
/// </param>
/// <param name="Alias">The per-birth alias, or <see langword="null"/> when the metric carried none.</param>
/// <param name="DataType">
/// The metric's declared datatype, or <see langword="null"/> when absent (again, the normal case
/// for a DATA metric — its datatype comes from the birth). An index this build's enum does not
/// define is preserved as an undefined enum value rather than dropped.
/// </param>
/// <param name="TimestampMs">
/// The metric's own acquisition timestamp in Sparkplug epoch milliseconds, or
/// <see langword="null"/> when it carried none — in which case the payload's timestamp applies.
/// </param>
/// <param name="ValueKind">
/// What <paramref name="Value"/> means. Check this before reading the value: a null value is
/// ambiguous between "explicitly null", "absent" and "a kind v1 does not support".
/// </param>
/// <param name="Value">
/// The <b>raw</b> wire value, boxed: <see cref="uint"/> (<c>int_value</c>), <see cref="ulong"/>
/// (<c>long_value</c>), <see cref="float"/>, <see cref="double"/>, <see cref="bool"/>,
/// <see cref="string"/> or <see cref="byte"/><c>[]</c> — and <see langword="null"/> for every
/// <paramref name="ValueKind"/> other than <see cref="SparkplugValueKind.Scalar"/>.
/// <para>
/// <b>Signed integers arrive as their unsigned two's-complement wire value</b> — a
/// <c>DataType.Int32</c> metric holding -42 is a <see cref="uint"/> of 4294967254 here. Run it
/// through <see cref="SparkplugCodec.ReinterpretSigned"/> with the metric's datatype (from the
/// birth, for a DATA metric) before publishing it.
/// </para>
/// <para>
/// Boxing is deliberate: the consumer coerces against the authored tag's declared
/// <c>DriverDataType</c> and hands the result to a <c>DataValueSnapshot</c>, which boxes
/// anyway, so a discriminated-union shape would buy an unboxing hop and cost every consumer a
/// switch over a dozen arms.
/// </para>
/// </param>
public readonly record struct SparkplugMetric(
string? Name,
ulong? Alias,
TahuDataType? DataType,
ulong? TimestampMs,
SparkplugValueKind ValueKind,
object? Value);
/// <summary>What a decoded metric's <see cref="SparkplugMetric.Value"/> represents.</summary>
/// <remarks>
/// Three of the four members have a null <see cref="SparkplugMetric.Value"/> and mean entirely
/// different things, which is exactly why the distinction is carried explicitly rather than left
/// for a consumer to infer from a null.
/// </remarks>
public enum SparkplugValueKind
{
/// <summary>The metric's value <c>oneof</c> carried nothing at all.</summary>
Absent = 0,
/// <summary>The metric set <c>is_null</c>: it exists, and its value is explicitly null.</summary>
Null,
/// <summary><see cref="SparkplugMetric.Value"/> holds the raw wire value.</summary>
Scalar,
/// <summary>
/// A <c>DataSet</c>, <c>Template</c> or extension value — decoded and reported, but not
/// supported in v1. The consumer should warn and skip the metric rather than treat it as
/// missing.
/// </summary>
Unsupported,
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,291 @@
using System.Diagnostics.CodeAnalysis;
namespace ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Sparkplug;
/// <summary>
/// The Sparkplug B topic-namespace element that identifies a message's purpose — the third
/// segment of <c>spBv1.0/{group}/{type}/{node}[/{device}]</c>, or the second segment of the
/// differently-shaped <c>spBv1.0/STATE/{hostId}</c>. See design doc §3.1/§3.6 and Sparkplug B
/// v3.0 spec §6 (Topic Namespace Elements).
/// </summary>
public enum SparkplugMessageType
{
/// <summary>Node birth certificate — (re)publishes an edge node's full metric/alias set.</summary>
NBIRTH,
/// <summary>Device birth certificate — (re)publishes a device's full metric/alias set.</summary>
DBIRTH,
/// <summary>Node data — incremental metric updates owned by the edge node itself.</summary>
NDATA,
/// <summary>Device data — incremental metric updates for a device under the edge node.</summary>
DDATA,
/// <summary>Node death certificate (the edge node's MQTT Will) — the node has gone offline.</summary>
NDEATH,
/// <summary>Device death certificate — the device has gone offline (published by its edge node).</summary>
DDEATH,
/// <summary>Node command — a write/rebirth request addressed to the edge node.</summary>
NCMD,
/// <summary>Device command — a write request addressed to a device under the edge node.</summary>
DCMD,
/// <summary>
/// Primary-host online/offline state. Shaped differently from every other message type — see
/// the remarks on <see cref="SparkplugTopic"/> and <see cref="SparkplugTopic.HostId"/>.
/// </summary>
STATE,
}
/// <summary>
/// A parsed Sparkplug B MQTT topic. See design doc §3.1/§3.6 and Sparkplug B v3.0 spec §6.
/// </summary>
/// <param name="Type">The message's purpose.</param>
/// <param name="GroupId">
/// The Sparkplug group id. <see langword="null"/> for <see cref="SparkplugMessageType.STATE"/>,
/// always populated otherwise.
/// </param>
/// <param name="EdgeNodeId">
/// The Sparkplug edge-node id. <see langword="null"/> for <see cref="SparkplugMessageType.STATE"/>,
/// always populated otherwise.
/// </param>
/// <param name="DeviceId">
/// The Sparkplug device id, present only for the device-scoped message types
/// (<see cref="SparkplugMessageType.DBIRTH"/>/<see cref="SparkplugMessageType.DDATA"/>/
/// <see cref="SparkplugMessageType.DDEATH"/>/<see cref="SparkplugMessageType.DCMD"/>);
/// <see langword="null"/> for node-scoped messages and for <see cref="SparkplugMessageType.STATE"/>.
/// </param>
/// <param name="HostId">
/// The primary-host id, populated only for <see cref="SparkplugMessageType.STATE"/>; carries the
/// third topic segment of <c>spBv1.0/STATE/{hostId}</c> (or the second segment of the legacy
/// pre-3.0 <c>STATE/{hostId}</c> form). <see langword="null"/> for every other message type.
/// </param>
/// <remarks>
/// <para>
/// <b>Never throws on arbitrary input.</b> <see cref="TryParse"/> is the primitive everything
/// else is built on — it is fed every topic string the broker delivers under a
/// <c>spBv1.0/{groupId}/#</c> subscription, none of it validated ahead of time, so it returns
/// <see langword="false"/> for anything malformed rather than throwing. <see cref="Parse"/> is
/// a throwing convenience wrapper for call sites (tests, hand-built topics) that already know
/// the string is well-formed.
/// </para>
/// <para>
/// <b>STATE does not fit the <c>{group}/{type}/{node}</c> mould — handled honestly, not
/// force-fit.</b> The Sparkplug v3.0 spec shapes the primary-host state topic as
/// <c>spBv1.0/STATE/{hostId}</c>: the message-type segment sits where a group id would
/// otherwise be, and there is no edge-node or device segment at all. Pre-3.0 peers additionally
/// published a bare <c>STATE/{hostId}</c> with no <c>spBv1.0</c> namespace prefix. This parser
/// targets the v3.0 form and tolerates the legacy one on receive (design §3.1); <see cref="Format"/>
/// and <see cref="FormatState"/> only ever produce the v3.0 form. A parsed STATE topic carries
/// its host id in <see cref="HostId"/> and leaves <see cref="GroupId"/>/<see cref="EdgeNodeId"/>/
/// <see cref="DeviceId"/> <see langword="null"/> — every other message type is the mirror image
/// (<see cref="GroupId"/>/<see cref="EdgeNodeId"/> populated, <see cref="HostId"/> null).
/// </para>
/// <para>
/// Group/edge-node/device/host segments are treated as opaque identifiers: validated only for
/// non-emptiness and the absence of the MQTT wildcard characters <c>+</c>/<c>#</c> (a broker
/// never delivers a PUBLISH on a topic containing either, so a topic string that does is
/// malformed, not a legitimate id worth preserving). No further character-set restriction is
/// applied — Sparkplug does not constrain id charsets beyond "not a topic-level separator".
/// </para>
/// </remarks>
public sealed record SparkplugTopic(
SparkplugMessageType Type,
string? GroupId,
string? EdgeNodeId,
string? DeviceId,
string? HostId)
{
private const string Namespace = "spBv1.0";
/// <summary>
/// Attempts to parse <paramref name="topic"/> as a Sparkplug B topic. Never throws — returns
/// <see langword="false"/> (and a <see langword="null"/> <paramref name="result"/>) for
/// anything that is not a well-formed Sparkplug topic, including <see langword="null"/>/empty
/// input, wrong namespace, an unrecognised message-type segment, a device-scoped message
/// missing its device segment (or vice versa), or a segment containing an MQTT wildcard.
/// </summary>
public static bool TryParse(string? topic, [NotNullWhen(true)] out SparkplugTopic? result)
{
result = null;
if (string.IsNullOrEmpty(topic))
{
return false;
}
var segments = topic.Split('/');
// Legacy pre-3.0 STATE form: `STATE/{hostId}`, no `spBv1.0` namespace prefix.
if (segments.Length == 2 && segments[0] == "STATE")
{
return TryBuildState(segments[1], out result);
}
if (segments.Length < 2 || segments[0] != Namespace)
{
return false;
}
// v3.0 STATE form: `spBv1.0/STATE/{hostId}`.
if (segments[1] == "STATE")
{
return segments.Length == 3 && TryBuildState(segments[2], out result);
}
// Every other message type: `spBv1.0/{group}/{type}/{node}[/{device}]`.
if (segments.Length is not (4 or 5))
{
return false;
}
if (!TryParseMessageType(segments[2], out var type))
{
return false;
}
var groupId = segments[1];
var edgeNodeId = segments[3];
if (!IsValidSegment(groupId) || !IsValidSegment(edgeNodeId))
{
return false;
}
var expectsDevice = IsDeviceScoped(type);
var hasDeviceSegment = segments.Length == 5;
if (expectsDevice != hasDeviceSegment)
{
return false;
}
string? deviceId = null;
if (hasDeviceSegment)
{
deviceId = segments[4];
if (!IsValidSegment(deviceId))
{
return false;
}
}
result = new SparkplugTopic(type, groupId, edgeNodeId, deviceId, null);
return true;
}
/// <summary>
/// Parses <paramref name="topic"/>, throwing <see cref="FormatException"/> if it is not a
/// well-formed Sparkplug B topic. See <see cref="TryParse"/> for the non-throwing form.
/// </summary>
public static SparkplugTopic Parse(string? topic)
{
if (TryParse(topic, out var result))
{
return result;
}
throw new FormatException($"'{topic}' is not a valid Sparkplug B topic.");
}
/// <summary>
/// Builds a Sparkplug B topic string for a non-STATE message — e.g.
/// <c>Format("Plant1", SparkplugMessageType.NCMD, "EdgeA")</c> for the Task 20 write path, so
/// it does not have to hand-concatenate segments itself.
/// </summary>
/// <exception cref="ArgumentException">
/// <paramref name="groupId"/>/<paramref name="edgeNodeId"/> is null/empty, or
/// <paramref name="type"/> is <see cref="SparkplugMessageType.STATE"/> (use
/// <see cref="FormatState"/> instead — STATE has no group/edge-node/device segments).
/// </exception>
public static string Format(string groupId, SparkplugMessageType type, string edgeNodeId, string? deviceId = null)
{
ArgumentException.ThrowIfNullOrEmpty(groupId);
ArgumentException.ThrowIfNullOrEmpty(edgeNodeId);
if (type == SparkplugMessageType.STATE)
{
throw new ArgumentException("Use FormatState to build a STATE topic.", nameof(type));
}
return deviceId is null
? $"{Namespace}/{groupId}/{type}/{edgeNodeId}"
: $"{Namespace}/{groupId}/{type}/{edgeNodeId}/{deviceId}";
}
/// <summary>Builds the v3.0 STATE topic string <c>spBv1.0/STATE/{hostId}</c>.</summary>
public static string FormatState(string hostId)
{
ArgumentException.ThrowIfNullOrEmpty(hostId);
return $"{Namespace}/STATE/{hostId}";
}
/// <summary>
/// Formats this instance back to its topic string (always the v3.0 STATE form for STATE
/// topics, even if this instance was parsed from the legacy no-prefix form).
/// </summary>
/// <exception cref="InvalidOperationException">
/// A required field for this instance's <see cref="Type"/> is <see langword="null"/> — i.e.
/// this instance was not built by <see cref="TryParse"/>/<see cref="Parse"/> and violates the
/// invariants documented on the type.
/// </exception>
public string ToTopicString() => Type == SparkplugMessageType.STATE
? FormatState(HostId ?? throw new InvalidOperationException("STATE topic is missing HostId."))
: Format(
GroupId ?? throw new InvalidOperationException("Non-STATE topic is missing GroupId."),
Type,
EdgeNodeId ?? throw new InvalidOperationException("Non-STATE topic is missing EdgeNodeId."),
DeviceId);
private static bool TryBuildState(string hostId, out SparkplugTopic? result)
{
result = null;
if (!IsValidSegment(hostId))
{
return false;
}
result = new SparkplugTopic(SparkplugMessageType.STATE, null, null, null, hostId);
return true;
}
private static bool TryParseMessageType(string segment, out SparkplugMessageType type)
{
switch (segment)
{
case nameof(SparkplugMessageType.NBIRTH):
type = SparkplugMessageType.NBIRTH;
return true;
case nameof(SparkplugMessageType.DBIRTH):
type = SparkplugMessageType.DBIRTH;
return true;
case nameof(SparkplugMessageType.NDATA):
type = SparkplugMessageType.NDATA;
return true;
case nameof(SparkplugMessageType.DDATA):
type = SparkplugMessageType.DDATA;
return true;
case nameof(SparkplugMessageType.NDEATH):
type = SparkplugMessageType.NDEATH;
return true;
case nameof(SparkplugMessageType.DDEATH):
type = SparkplugMessageType.DDEATH;
return true;
case nameof(SparkplugMessageType.NCMD):
type = SparkplugMessageType.NCMD;
return true;
case nameof(SparkplugMessageType.DCMD):
type = SparkplugMessageType.DCMD;
return true;
default:
type = default;
return false;
}
}
private static bool IsDeviceScoped(SparkplugMessageType type) => type is
SparkplugMessageType.DBIRTH or SparkplugMessageType.DDATA or SparkplugMessageType.DDEATH or SparkplugMessageType.DCMD;
private static bool IsValidSegment(string segment) =>
segment.Length > 0 && segment.IndexOf('+') < 0 && segment.IndexOf('#') < 0;
}
@@ -0,0 +1,29 @@
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup>
<TargetFramework>net10.0</TargetFramework>
<Nullable>enable</Nullable>
<ImplicitUsings>enable</ImplicitUsings>
<LangVersion>latest</LangVersion>
<TreatWarningsAsErrors>true</TreatWarningsAsErrors>
<GenerateDocumentationFile>true</GenerateDocumentationFile>
<NoWarn>$(NoWarn);CS1591</NoWarn>
<RootNamespace>ZB.MOM.WW.OtOpcUa.Driver.Mqtt</RootNamespace>
<AssemblyName>ZB.MOM.WW.OtOpcUa.Driver.Mqtt</AssemblyName>
</PropertyGroup>
<ItemGroup>
<ProjectReference Include="..\ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Contracts\ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Contracts.csproj"/>
<ProjectReference Include="..\..\Core\ZB.MOM.WW.OtOpcUa.Core.Abstractions\ZB.MOM.WW.OtOpcUa.Core.Abstractions.csproj"/>
<ProjectReference Include="..\..\Core\ZB.MOM.WW.OtOpcUa.Core\ZB.MOM.WW.OtOpcUa.Core.csproj"/>
</ItemGroup>
<ItemGroup>
<PackageReference Include="MQTTnet"/>
</ItemGroup>
<ItemGroup>
<InternalsVisibleTo Include="ZB.MOM.WW.OtOpcUa.Driver.Mqtt.Tests"/>
</ItemGroup>
</Project>
@@ -14,11 +14,13 @@ namespace ZB.MOM.WW.OtOpcUa.Driver.Sql;
/// parameterized in SQL, so the few that must appear as text are emitted through
/// <see cref="QuoteIdentifier"/>. Any code path that builds SQL by concatenating an authored tag field
/// <em>without</em> going through <see cref="QuoteIdentifier"/> is a defect (design §8.1).</para>
/// <para><b>Not yet implemented:</b> design §8.1 also specifies that an authored table/column is
/// validated against the live catalog before it is ever quoted into text, so that an unknown identifier
/// rejects the tag. No such gate exists yet — an identifier reaches <see cref="QuoteIdentifier"/>
/// straight from the authored <c>TagConfig</c> blob, so <b>quoting is currently the only defence</b>,
/// not a backstop behind an upstream filter. Scrutinise it accordingly.</para>
/// <para><b>Quoting is the backstop, not the only defence.</b> Design §8.1's catalog gate is
/// implemented: <see cref="SqlCatalogGate"/> resolves every authored table/column against the live
/// catalog at Initialize and <b>replaces it with the catalog's own spelling</b>, so an identifier
/// reaching <see cref="QuoteIdentifier"/> on the poll path is a string this driver read back out of
/// <see cref="ListSchemasSql"/> / <see cref="ListTablesSql"/> / <see cref="ListColumnsSql"/> — not
/// operator input. An identifier that resolves to nothing rejects its tag (which then publishes
/// <c>BadNodeIdUnknown</c>) rather than being quoted into a query against a nonexistent object.</para>
/// <para>Public because <c>Driver.Sql.Browser</c> consumes it — the catalog SQL <em>is</em> the browse
/// engine, so it is shared rather than duplicated. Implementations own their provider package;
/// <b>no provider-specific type appears in this signature</b> (<see cref="Factory"/> is the abstract
@@ -89,6 +91,20 @@ public interface ISqlDialect
/// <summary>Catalog query listing schemas — the browser's <c>RootAsync</c> level. Takes no parameters.</summary>
string ListSchemasSql { get; }
/// <summary>
/// Scalar query returning the schema an <b>unqualified</b> object name resolves to for the connecting
/// principal — T-SQL's <c>SELECT SCHEMA_NAME()</c>. Takes no parameters.
/// </summary>
/// <remarks>
/// <para>Needed by <see cref="SqlCatalogGate"/>: a tag authored as <c>TagValues</c> rather than
/// <c>dbo.TagValues</c> has to be looked up in <em>some</em> schema, and guessing <c>dbo</c> would be a
/// silent lie on any estate that maps its service accounts to their own default schema. Asking the
/// server is the only answer that matches how the poll query itself will resolve the name.</para>
/// <para>It is a <em>query</em> rather than a name because the answer is per-connection, not per-dialect
/// — the same dialect resolves differently for two different logins.</para>
/// </remarks>
string DefaultSchemaSql { get; }
/// <summary>Catalog query listing tables + views in one schema — the browser's schema-expand level. Bind <c>@schema</c>.</summary>
string ListTablesSql { get; }
@@ -0,0 +1,137 @@
namespace ZB.MOM.WW.OtOpcUa.Driver.Sql;
/// <summary>
/// An immutable snapshot of the parts of a database's catalog the authored tags actually name — the
/// allow-list design §8.1 requires identifiers to come from.
/// <para><b>This type holds catalog strings only.</b> Every name in it was read back out of
/// <see cref="ISqlDialect.ListSchemasSql"/> / <see cref="ISqlDialect.ListTablesSql"/> /
/// <see cref="ISqlDialect.ListColumnsSql"/>; nothing an operator typed is ever stored here. That is what
/// lets <see cref="SqlCatalogGate"/> hand the planner catalog spellings rather than authored ones, so the
/// identifier text in an emitted query is a string the database gave us.</para>
/// <para><b>Deliberately partial.</b> Only the schemas/tables the authored tags name are loaded — a
/// driver polling three tables must not enumerate a warehouse's ten thousand. A name absent from this
/// snapshot therefore means "not found <em>when we looked</em>", which is exactly the question the gate
/// asks.</para>
/// </summary>
public sealed class SqlCatalog
{
private readonly IReadOnlyList<string> _schemas;
/// <summary>Canonical schema → canonical table names in it.</summary>
private readonly IReadOnlyDictionary<string, IReadOnlyList<string>> _tablesBySchema;
/// <summary>Canonical <c>schema.table</c> → that relation's canonical column names.</summary>
private readonly IReadOnlyDictionary<string, IReadOnlyList<string>> _columnsByTable;
/// <summary>Constructs a catalog snapshot.</summary>
/// <param name="defaultSchema">The schema an unqualified object name resolves to for this connection.</param>
/// <param name="schemas">Every schema the connection can see.</param>
/// <param name="tablesBySchema">Canonical schema → its tables/views.</param>
/// <param name="columnsByTable">Canonical <c>schema.table</c> → its columns.</param>
internal SqlCatalog(
string defaultSchema,
IReadOnlyList<string> schemas,
IReadOnlyDictionary<string, IReadOnlyList<string>> tablesBySchema,
IReadOnlyDictionary<string, IReadOnlyList<string>> columnsByTable)
{
DefaultSchema = defaultSchema;
_schemas = schemas;
_tablesBySchema = tablesBySchema;
_columnsByTable = columnsByTable;
}
/// <summary>The schema an unqualified authored object name is resolved in.</summary>
public string DefaultSchema { get; }
/// <summary>Resolves an authored schema name to the catalog's own spelling.</summary>
/// <param name="authored">The authored schema, or null/blank for the default schema.</param>
/// <param name="canonical">The catalog's spelling, when resolved.</param>
/// <returns><see langword="true"/> when the schema exists (or the default schema is used).</returns>
public bool TryResolveSchema(string? authored, out string canonical)
{
if (string.IsNullOrWhiteSpace(authored))
{
canonical = DefaultSchema;
return true;
}
return TryMatch(_schemas, authored, out canonical);
}
/// <summary>Resolves an authored table/view name within a canonical schema.</summary>
/// <param name="canonicalSchema">The schema, already resolved through <see cref="TryResolveSchema"/>.</param>
/// <param name="authored">The authored table or view name.</param>
/// <param name="canonical">The catalog's spelling, when resolved.</param>
/// <returns><see langword="true"/> when the relation exists in that schema.</returns>
public bool TryResolveTable(string canonicalSchema, string authored, out string canonical)
{
canonical = string.Empty;
return _tablesBySchema.TryGetValue(canonicalSchema, out var tables)
&& TryMatch(tables, authored, out canonical);
}
/// <summary>Resolves an authored column name within a canonical relation.</summary>
/// <param name="canonicalSchema">The resolved schema.</param>
/// <param name="canonicalTable">The resolved table or view.</param>
/// <param name="authored">The authored column name.</param>
/// <param name="canonical">The catalog's spelling, when resolved.</param>
/// <returns><see langword="true"/> when the column exists on that relation.</returns>
public bool TryResolveColumn(
string canonicalSchema, string canonicalTable, string authored, out string canonical)
{
canonical = string.Empty;
return _columnsByTable.TryGetValue(QualifiedKey(canonicalSchema, canonicalTable), out var columns)
&& TryMatch(columns, authored, out canonical);
}
/// <summary>The dictionary key for a canonical relation.</summary>
/// <param name="schema">The canonical schema.</param>
/// <param name="table">The canonical table.</param>
/// <returns>The composite key.</returns>
internal static string QualifiedKey(string schema, string table) => schema + "." + table;
/// <summary>
/// Matches an authored name against catalog spellings: an exact (ordinal) hit wins, otherwise a
/// <b>unique</b> case-insensitive hit is accepted and its catalog spelling returned.
/// </summary>
/// <remarks>
/// <para><b>Why case-insensitive at all.</b> SQL Server's default collation is case-insensitive, so
/// <c>num_value</c> and <c>NUM_VALUE</c> genuinely name the same column and both work today. Matching
/// ordinally would reject configurations that have always been valid — a gate that breaks working
/// deployments is not defence in depth.</para>
/// <para><b>Why exact-first, and why ambiguity fails.</b> On a case-<em>sensitive</em> collation a
/// relation may legitimately carry both <c>Value</c> and <c>value</c>. Preferring the exact match keeps
/// such a config resolving to what the operator wrote, and refusing to choose between two
/// case-insensitive candidates is the only safe answer — silently picking one would publish a different
/// column's data under the operator's node, which is worse than rejecting the tag.</para>
/// </remarks>
/// <param name="candidates">The catalog spellings to match against.</param>
/// <param name="authored">The authored name.</param>
/// <param name="canonical">The catalog's spelling, when matched.</param>
/// <returns><see langword="true"/> on an unambiguous match.</returns>
internal static bool TryMatch(IReadOnlyList<string> candidates, string authored, out string canonical)
{
foreach (var candidate in candidates)
{
if (!string.Equals(candidate, authored, StringComparison.Ordinal)) continue;
canonical = candidate;
return true;
}
canonical = string.Empty;
var matches = 0;
foreach (var candidate in candidates)
{
if (!string.Equals(candidate, authored, StringComparison.OrdinalIgnoreCase)) continue;
if (++matches > 1)
{
canonical = string.Empty;
return false;
}
canonical = candidate;
}
return matches == 1;
}
}
@@ -0,0 +1,241 @@
using ZB.MOM.WW.OtOpcUa.Driver.Sql.Contracts;
namespace ZB.MOM.WW.OtOpcUa.Driver.Sql;
/// <summary>
/// One tag the gate refused, and why. Carried rather than logged in place so the caller decides the log
/// shape and the caller's tests can assert on the reason.
/// </summary>
/// <param name="RawPath">The rejected tag's identity — its RawPath, which is trusted structure, not blob content.</param>
/// <param name="Field">The <see cref="SqlTagDefinition"/> field that failed (e.g. <c>ValueColumn</c>).</param>
/// <param name="Reason">An operator-actionable explanation, safe to log.</param>
public sealed record SqlCatalogRejection(string RawPath, string Field, string Reason);
/// <summary>The outcome of applying a <see cref="SqlCatalog"/> to a set of authored tags.</summary>
/// <param name="Accepted">
/// The surviving definitions, <b>rewritten to catalog spellings</b>. Safe to hand to
/// <see cref="SqlGroupPlanner"/>.
/// </param>
/// <param name="Rejected">Every tag the gate refused, in input order.</param>
public sealed record SqlCatalogGateResult(
IReadOnlyList<SqlTagDefinition> Accepted,
IReadOnlyList<SqlCatalogRejection> Rejected);
/// <summary>
/// Design §8.1's identifier gate: validates every authored table/column against the live catalog
/// <b>before</b> it can be quoted into SQL text, and replaces it with the catalog's own spelling.
/// </summary>
/// <remarks>
/// <para><b>What this closes.</b> Until this existed, <see cref="ISqlDialect.QuoteIdentifier"/> was the
/// sole defence: an identifier went from the authored <c>TagConfig</c> blob straight into a command text,
/// bracket-quoted. The residual risk was bounded — a hostile name is quoted into one nonexistent object,
/// the query fails after the connection opens, and the tag Bad-codes — but "bounded" is not "filtered",
/// and the design promised a filter. With the gate in place the identifier text in an emitted query is a
/// string this driver read back out of the catalog, so quoting is a backstop behind an allow-list rather
/// than the whole story.</para>
/// <para><b>Charset first, catalog second.</b> Each identifier is passed through
/// <see cref="ISqlDialect.QuoteIdentifier"/> before it is looked up, purely for its rejection rules
/// (control characters, Unicode format characters, over-length). That ordering is deliberate: a name that
/// fails the charset check is rejected <em>without its value being echoed</em>, so nothing carrying a bidi
/// override or a NUL can reach a log line through this type's rejection messages. A name that passes has
/// been proven safe to render, which is why the catalog-miss messages can afford to name it — and they
/// must, or an operator has no way to find the typo.</para>
/// <para><b>Rejection is per-tag, never driver-wide.</b> A tag naming a dropped column is dropped from the
/// driver's table, so its RawPath resolves to nothing and it publishes
/// <see cref="SqlStatusCodes.BadNodeIdUnknown"/> — design §8.1's specified outcome — while every other tag
/// on that database keeps polling. This mirrors the "one malformed blob must not take the whole driver
/// down" rule the tag-table build already follows.</para>
/// </remarks>
public static class SqlCatalogGate
{
/// <summary>
/// The most name parts this gate can validate. A three-part <c>db.schema.table</c> (or a
/// four-part linked-server name) addresses a catalog this connection's <c>INFORMATION_SCHEMA</c>
/// cannot see, so it cannot be allow-listed.
/// </summary>
private const int MaxNameParts = 2;
/// <summary>
/// Applies <paramref name="catalog"/> to <paramref name="definitions"/>, returning the accepted tags
/// with catalog-spelled identifiers and the rejected ones with reasons.
/// </summary>
/// <param name="definitions">The authored definitions, as parsed from their <c>TagConfig</c> blobs.</param>
/// <param name="catalog">The catalog snapshot to validate against.</param>
/// <param name="dialect">Supplies the identifier charset rules via <see cref="ISqlDialect.QuoteIdentifier"/>.</param>
/// <returns>The accepted and rejected sets.</returns>
/// <exception cref="ArgumentNullException">A required argument is null.</exception>
public static SqlCatalogGateResult Apply(
IEnumerable<SqlTagDefinition> definitions, SqlCatalog catalog, ISqlDialect dialect)
{
ArgumentNullException.ThrowIfNull(definitions);
ArgumentNullException.ThrowIfNull(catalog);
ArgumentNullException.ThrowIfNull(dialect);
var accepted = new List<SqlTagDefinition>();
var rejected = new List<SqlCatalogRejection>();
foreach (var definition in definitions)
{
ArgumentNullException.ThrowIfNull(definition);
if (TryCanonicalize(definition, catalog, dialect, out var canonical, out var rejection))
accepted.Add(canonical!);
else
rejected.Add(rejection!);
}
return new SqlCatalogGateResult(accepted, rejected);
}
/// <summary>Resolves one definition's identifiers, or explains the first failure.</summary>
private static bool TryCanonicalize(
SqlTagDefinition definition,
SqlCatalog catalog,
ISqlDialect dialect,
out SqlTagDefinition? canonical,
out SqlCatalogRejection? rejection)
{
canonical = null;
rejection = null;
if (!TryResolveRelation(definition, catalog, dialect, out var schema, out var table, out rejection))
return false;
// Every identifier-bearing field, paired with its name for the rejection message. A field the tag's
// model does not use is null and simply skipped — the planner enforces which are required.
var resolved = new Dictionary<string, string?>(StringComparer.Ordinal);
foreach (var (field, authored) in new (string Field, string? Authored)[]
{
(nameof(SqlTagDefinition.KeyColumn), definition.KeyColumn),
(nameof(SqlTagDefinition.ValueColumn), definition.ValueColumn),
(nameof(SqlTagDefinition.TimestampColumn), definition.TimestampColumn),
(nameof(SqlTagDefinition.ColumnName), definition.ColumnName),
(nameof(SqlTagDefinition.RowSelectorColumn), definition.RowSelectorColumn),
(nameof(SqlTagDefinition.RowSelectorTopByTimestamp), definition.RowSelectorTopByTimestamp),
})
{
if (string.IsNullOrWhiteSpace(authored))
{
resolved[field] = authored;
continue;
}
if (!IsRenderableIdentifier(authored, dialect))
{
rejection = new SqlCatalogRejection(definition.Name, field,
$"the authored {field} is not a usable identifier (it is over-long, or contains control " +
"or Unicode format characters); the value is withheld from this message because it " +
"cannot be safely rendered");
return false;
}
if (!catalog.TryResolveColumn(schema, table, authored, out var column))
{
rejection = new SqlCatalogRejection(definition.Name, field,
$"column '{authored}' does not exist on '{schema}.{table}' (or matches more than one " +
"column under a case-sensitive collation)");
return false;
}
resolved[field] = column;
}
canonical = definition with
{
Table = SqlCatalog.QualifiedKey(schema, table),
KeyColumn = resolved[nameof(SqlTagDefinition.KeyColumn)],
ValueColumn = resolved[nameof(SqlTagDefinition.ValueColumn)],
TimestampColumn = resolved[nameof(SqlTagDefinition.TimestampColumn)],
ColumnName = resolved[nameof(SqlTagDefinition.ColumnName)],
RowSelectorColumn = resolved[nameof(SqlTagDefinition.RowSelectorColumn)],
RowSelectorTopByTimestamp = resolved[nameof(SqlTagDefinition.RowSelectorTopByTimestamp)],
};
return true;
}
/// <summary>Splits and resolves the authored <c>table</c> to a canonical schema + relation.</summary>
private static bool TryResolveRelation(
SqlTagDefinition definition,
SqlCatalog catalog,
ISqlDialect dialect,
out string schema,
out string table,
out SqlCatalogRejection? rejection)
{
schema = string.Empty;
table = string.Empty;
rejection = null;
const string Field = nameof(SqlTagDefinition.Table);
if (string.IsNullOrWhiteSpace(definition.Table))
{
rejection = new SqlCatalogRejection(definition.Name, Field,
"the tag has no 'table'; author the table or view to read");
return false;
}
var parts = definition.Table.Split('.');
if (parts.Length > MaxNameParts)
{
rejection = new SqlCatalogRejection(definition.Name, Field,
$"'{definition.Table}' is a {parts.Length}-part name. Only 'table' and 'schema.table' can be " +
"validated against this connection's catalog, so a cross-database or linked-server name " +
"cannot be allow-listed; expose the data through a view in this database instead");
return false;
}
var authoredSchema = parts.Length == MaxNameParts ? parts[0] : null;
var authoredTable = parts[^1];
foreach (var part in parts)
{
if (IsRenderableIdentifier(part, dialect)) continue;
rejection = new SqlCatalogRejection(definition.Name, Field,
"the authored table name has a part that is not a usable identifier (empty, over-long, or " +
"containing control or Unicode format characters); the value is withheld from this message " +
"because it cannot be safely rendered");
return false;
}
if (!catalog.TryResolveSchema(authoredSchema, out schema))
{
rejection = new SqlCatalogRejection(definition.Name, Field,
$"schema '{authoredSchema}' does not exist (or matches more than one schema under a " +
"case-sensitive collation)");
return false;
}
if (!catalog.TryResolveTable(schema, authoredTable, out table))
{
rejection = new SqlCatalogRejection(definition.Name, Field,
$"table or view '{authoredTable}' does not exist in schema '{schema}'" +
(authoredSchema is null
? $" (the connection's default schema — qualify the name as 'schema.{authoredTable}' if it lives elsewhere)"
: string.Empty));
return false;
}
return true;
}
/// <summary>
/// True when the dialect's quoting rules accept <paramref name="identifier"/> — i.e. it is safe both to
/// embed in SQL and to render into a log line or an AdminUI label.
/// </summary>
/// <remarks>
/// Uses <see cref="ISqlDialect.QuoteIdentifier"/> as the charset authority rather than duplicating its
/// rules, so the two can never disagree about what a legal identifier is. The quoted result is
/// discarded: only whether it threw is interesting here.
/// </remarks>
private static bool IsRenderableIdentifier(string identifier, ISqlDialect dialect)
{
try
{
dialect.QuoteIdentifier(identifier);
return true;
}
catch (ArgumentException)
{
return false;
}
}
}
@@ -0,0 +1,189 @@
using System.Data;
using System.Data.Common;
namespace ZB.MOM.WW.OtOpcUa.Driver.Sql;
/// <summary>
/// Loads the <see cref="SqlCatalog"/> slice the authored tags need, over one connection, using only the
/// dialect's catalog SQL (design §8.1).
/// </summary>
/// <remarks>
/// <para><b>Bounded by what is authored, not by what exists.</b> One query for the schema list, one for
/// the default schema, then one <see cref="ISqlDialect.ListTablesSql"/> per distinct authored schema and
/// one <see cref="ISqlDialect.ListColumnsSql"/> per distinct authored relation. A driver polling three
/// tables issues a handful of round-trips at Initialize and none thereafter; it never enumerates a
/// warehouse.</para>
/// <para><b>Every parameter is bound.</b> Authored names reach the catalog queries as
/// <c>@schema</c>/<c>@table</c> parameters — the same discipline the schema browser follows — so loading
/// the allow-list cannot itself be an injection vector. No identifier is quoted into text on this path.</para>
/// <para><b>Failures throw; they do not degrade to an empty catalog.</b> An empty catalog would reject
/// every tag, which is indistinguishable at the OPC UA surface from a database whose objects were all
/// dropped. A caller that cannot read the catalog has not learned that the tags are invalid — it has
/// learned nothing — and must fail its Initialize so the driver retries rather than serving a
/// confidently-empty address space.</para>
/// </remarks>
internal static class SqlCatalogLoader
{
/// <summary>Column alias <see cref="ISqlDialect.ListSchemasSql"/> projects.</summary>
private const string SchemaColumn = "TABLE_SCHEMA";
/// <summary>Column alias <see cref="ISqlDialect.ListTablesSql"/> projects.</summary>
private const string TableNameColumn = "TABLE_NAME";
/// <summary>Column alias <see cref="ISqlDialect.ListColumnsSql"/> projects.</summary>
private const string ColumnNameColumn = "COLUMN_NAME";
/// <summary>
/// Reads the catalog slice covering <paramref name="authoredTables"/>.
/// </summary>
/// <param name="connection">An <b>already-open</b> connection. Not disposed here; the caller owns it.</param>
/// <param name="dialect">Supplies the catalog SQL.</param>
/// <param name="authoredTables">The distinct authored <c>table</c> strings, as written in the blobs.</param>
/// <param name="commandTimeout">Per-command server-side backstop.</param>
/// <param name="cancellationToken">Cancellation token for the operation.</param>
/// <returns>The loaded catalog.</returns>
/// <exception cref="InvalidOperationException">The connection can see no schemas at all.</exception>
internal static async Task<SqlCatalog> LoadAsync(
DbConnection connection,
ISqlDialect dialect,
IReadOnlyCollection<string> authoredTables,
TimeSpan commandTimeout,
CancellationToken cancellationToken)
{
var timeoutSeconds = Math.Max(1, (int)Math.Ceiling(commandTimeout.TotalSeconds));
var schemas = await QueryColumnAsync(
connection, dialect.ListSchemasSql, SchemaColumn,
static _ => { }, timeoutSeconds, cancellationToken).ConfigureAwait(false);
// Zero schemas is a permissions or visibility symptom, never a real database. Rejecting every tag on
// that basis would report an authoring fault for what is actually a grant problem, and would send the
// operator to the wrong system — the same misdiagnosis the driver's health classification avoids.
if (schemas.Count == 0)
throw new InvalidOperationException(
"the catalog returned no schemas at all; the connecting principal most likely cannot read " +
"the catalog views, which is a grant problem rather than a tag-authoring one");
var defaultSchema = await ReadDefaultSchemaAsync(
connection, dialect, timeoutSeconds, cancellationToken).ConfigureAwait(false);
var tablesBySchema = new Dictionary<string, IReadOnlyList<string>>(StringComparer.Ordinal);
var columnsByTable = new Dictionary<string, IReadOnlyList<string>>(StringComparer.Ordinal);
// Resolving here goes through SqlCatalog.TryMatch — the SAME matcher the gate will use — so the
// loader and the gate can never disagree about which relation an authored name resolved to. Loading
// one table and validating against another would be a silent wrong-column defect.
foreach (var authored in authoredTables)
{
if (string.IsNullOrWhiteSpace(authored)) continue;
var parts = authored.Split('.');
if (parts.Length > 2) continue; // the gate rejects these by name; nothing to load.
var authoredTable = parts[^1];
string schema;
if (parts.Length == 2)
{
if (!SqlCatalog.TryMatch(schemas, parts[0], out schema)) continue;
}
else
{
schema = defaultSchema;
}
if (!tablesBySchema.TryGetValue(schema, out var tables))
{
tables = await QueryColumnAsync(
connection, dialect.ListTablesSql, TableNameColumn,
command => Bind(command, "@schema", schema),
timeoutSeconds, cancellationToken).ConfigureAwait(false);
tablesBySchema[schema] = tables;
}
if (!SqlCatalog.TryMatch(tables, authoredTable, out var table)) continue;
var key = SqlCatalog.QualifiedKey(schema, table);
if (columnsByTable.ContainsKey(key)) continue;
columnsByTable[key] = await QueryColumnAsync(
connection, dialect.ListColumnsSql, ColumnNameColumn,
command =>
{
Bind(command, "@schema", schema);
Bind(command, "@table", table);
},
timeoutSeconds, cancellationToken).ConfigureAwait(false);
}
return new SqlCatalog(defaultSchema, schemas, tablesBySchema, columnsByTable);
}
/// <summary>
/// Reads the connection's default schema, falling back to the sole visible schema when the dialect's
/// query answers null/blank.
/// </summary>
/// <remarks>
/// A null answer is possible (a login with no default schema mapped). Falling back to the single
/// visible schema keeps the common one-schema database working; with several visible schemas there is
/// no defensible guess, so unqualified names simply will not resolve and the gate says so by name.
/// </remarks>
private static async Task<string> ReadDefaultSchemaAsync(
DbConnection connection, ISqlDialect dialect, int timeoutSeconds, CancellationToken cancellationToken)
{
var command = connection.CreateCommand();
await using (command.ConfigureAwait(false))
{
command.CommandText = dialect.DefaultSchemaSql;
command.CommandTimeout = timeoutSeconds;
var value = await command.ExecuteScalarAsync(cancellationToken).ConfigureAwait(false);
var schema = value is null or DBNull ? null : value.ToString();
return string.IsNullOrWhiteSpace(schema) ? string.Empty : schema;
}
}
/// <summary>Runs one catalog query and projects a single string column from every row.</summary>
private static async Task<IReadOnlyList<string>> QueryColumnAsync(
DbConnection connection,
string sql,
string columnName,
Action<DbCommand> bind,
int timeoutSeconds,
CancellationToken cancellationToken)
{
var results = new List<string>();
var command = connection.CreateCommand();
await using (command.ConfigureAwait(false))
{
command.CommandText = sql;
command.CommandTimeout = timeoutSeconds;
bind(command);
var reader = await command
.ExecuteReaderAsync(CommandBehavior.SingleResult, cancellationToken).ConfigureAwait(false);
await using (reader.ConfigureAwait(false))
{
var ordinal = -1;
while (await reader.ReadAsync(cancellationToken).ConfigureAwait(false))
{
if (ordinal < 0) ordinal = reader.GetOrdinal(columnName);
if (reader.IsDBNull(ordinal)) continue;
var value = reader.GetValue(ordinal)?.ToString();
if (!string.IsNullOrWhiteSpace(value)) results.Add(value);
}
}
}
return results;
}
/// <summary>Binds one catalog-query parameter — the only way an authored name reaches the database here.</summary>
private static void Bind(DbCommand command, string name, string value)
{
var parameter = command.CreateParameter();
parameter.ParameterName = name;
parameter.DbType = DbType.String;
parameter.Value = value;
command.Parameters.Add(parameter);
}
}
@@ -29,14 +29,12 @@ public sealed class SqlDriver
: IDriver, ITagDiscovery, IReadable, ISubscribable, IHostConnectivityProbe, IAsyncDisposable
{
/// <summary>
/// The driver-type string this driver registers under.
/// <para><b>Deliberately a local constant, not <c>DriverTypeNames.Sql</c>.</b>
/// <c>DriverTypeNamesGuardTests</c> asserts bidirectional parity between the constants and the
/// driver factories actually registered in the process, so the constant may only be added in the
/// same change that wires the factory (this task ships no factory — see the plan's Task 11, which
/// adds <c>DriverTypeNames.Sql</c> and repoints this constant at it).</para>
/// The driver-type string this driver registers under — <see cref="DriverTypeNames.Sql"/>, the single
/// source of truth the deploy pipeline stores in <c>DriverInstance.DriverType</c> and every dispatch
/// map keys by. Chained off the shared constant (Task 11 wired the factory + probe, so
/// <c>DriverTypeNamesGuardTests</c>' bidirectional-parity check is satisfied).
/// </summary>
public const string DriverTypeName = "Sql";
public const string DriverTypeName = DriverTypeNames.Sql;
/// <summary>Shown for a connection string that names no recognisable server.</summary>
private const string UnknownEndpoint = "(unknown sql endpoint)";
@@ -76,7 +74,23 @@ public sealed class SqlDriver
private IReadOnlyDictionary<string, SqlTagDefinition> _tagsByRawPath =
new Dictionary<string, SqlTagDefinition>(StringComparer.Ordinal);
/// <summary>Resolves a read/subscribe RawPath to its authored definition — a single table hit, or a miss.</summary>
/// <summary>
/// The subset of <see cref="_tagsByRawPath"/> that survived the design §8.1 catalog gate, with every
/// identifier rewritten to the catalog's own spelling. <b>This — not the authored table — is what a
/// read or a poll resolves against</b>, so an identifier that never appeared in the catalog can never
/// reach a query.
/// <para><b>Why the two tables are separate.</b> A tag the gate rejected must still <em>exist</em> as a
/// node: §8.1's specified outcome is that it publishes
/// <see cref="SqlStatusCodes.BadNodeIdUnknown"/>, and a status code can only be published by a node
/// that is there. Dropping rejected tags from the authored table too would delete the node instead,
/// turning a diagnosable Bad quality into a silently missing address-space entry — a strictly worse
/// failure for the operator trying to find their typo.</para>
/// <para>Swapped atomically, for the same reason the authored table is.</para>
/// </summary>
private IReadOnlyDictionary<string, SqlTagDefinition> _polledByRawPath =
new Dictionary<string, SqlTagDefinition>(StringComparer.Ordinal);
/// <summary>Resolves a read/subscribe RawPath to its <b>catalog-validated</b> definition, or a miss.</summary>
private readonly EquipmentTagRefResolver<SqlTagDefinition> _resolver;
/// <summary>
@@ -184,8 +198,18 @@ public sealed class SqlDriver
/// <summary>The current authored-tag snapshot. Read through a barrier — <see cref="BuildTagTable"/> swaps it.</summary>
private IReadOnlyDictionary<string, SqlTagDefinition> Tags => Volatile.Read(ref _tagsByRawPath);
/// <summary>The resolver's lookup: RawPath → authored definition, or null on a miss.</summary>
private SqlTagDefinition? Lookup(string rawPath) => Tags.GetValueOrDefault(rawPath);
/// <summary>
/// The catalog-validated snapshot. Read through a barrier — <see cref="ApplyCatalogGateAsync"/> swaps it.
/// </summary>
private IReadOnlyDictionary<string, SqlTagDefinition> PolledTags => Volatile.Read(ref _polledByRawPath);
/// <summary>
/// The resolver's lookup: RawPath → <b>catalog-validated</b> definition, or null on a miss.
/// <para>Deliberately reads <see cref="PolledTags"/> rather than <see cref="Tags"/>: a tag the gate
/// rejected must miss here so the reader publishes
/// <see cref="SqlStatusCodes.BadNodeIdUnknown"/>, exactly as design §8.1 specifies.</para>
/// </summary>
private SqlTagDefinition? Lookup(string rawPath) => PolledTags.GetValueOrDefault(rawPath);
// ---- IDriver lifecycle ----
@@ -198,11 +222,16 @@ public sealed class SqlDriver
/// I/O, and on failure this method records <see cref="DriverState.Faulted"/> <b>and rethrows</b> —
/// <c>DriverInstanceActor</c> reads a throw as InitializeFailed and lands in Reconnecting with its
/// retry timer running, which is exactly the recovery a database that is merely down needs.</para>
/// <para>The liveness check is then followed by the design §8.1 <b>catalog gate</b>
/// (<see cref="ApplyCatalogGateAsync"/>), which is the second and last piece of I/O. It runs after
/// liveness because it needs a working connection, and before the driver reports Healthy because the
/// tag table it publishes must be the validated one — a poll must never see an unvalidated
/// identifier.</para>
/// </summary>
/// <param name="driverConfigJson">The driver configuration JSON (unused; see remarks).</param>
/// <param name="cancellationToken">Cancellation token for the operation.</param>
/// <returns>A task that represents the asynchronous operation.</returns>
/// <exception cref="InvalidOperationException">The database could not be reached.</exception>
/// <exception cref="InvalidOperationException">The database could not be reached, or its catalog could not be read.</exception>
public async Task InitializeAsync(string driverConfigJson, CancellationToken cancellationToken)
{
WriteHealth(new DriverHealth(DriverState.Initializing, null, null));
@@ -236,6 +265,36 @@ public sealed class SqlDriver
throw new InvalidOperationException(message, ex);
}
// Separate try from liveness so the operator surface names the stage that actually failed: "could
// not reach the database" and "reached it but could not read its catalog" send an operator to
// different places, and the second is usually a GRANT rather than an outage.
try
{
await ApplyCatalogGateAsync(cancellationToken).ConfigureAwait(false);
}
catch (OperationCanceledException) when (cancellationToken.IsCancellationRequested)
{
WriteHealth(new DriverHealth(DriverState.Unknown, null, null));
throw;
}
catch (Exception ex)
{
// Same discipline as above: exception TYPE only on the operator surface, full exception to the
// log sink. A catalog read that fails is NOT evidence the tags are wrong — it is the absence of
// evidence — so this faults the driver (and retries) rather than rejecting every tag and serving
// a confidently-empty address space.
var message =
$"Sql driver '{_driverInstanceId}' reached {Endpoint} but could not read its catalog to " +
$"validate authored tables/columns: {ex.GetType().Name}. Check that the connecting principal " +
"can read the catalog views.";
_logger.LogError(ex,
"Sql driver {DriverInstanceId} catalog validation against {Endpoint} failed.",
_driverInstanceId, Endpoint);
WriteHealth(new DriverHealth(DriverState.Faulted, null, message));
TransitionHostTo(HostState.Stopped);
throw new InvalidOperationException(message, ex);
}
WriteHealth(new DriverHealth(DriverState.Healthy, DateTime.UtcNow, null));
TransitionHostTo(HostState.Running);
_logger.LogInformation(
@@ -492,15 +551,17 @@ public sealed class SqlDriver
// A subscribed-read outage surfaces as a DbException the reader threw, which ReadAsync has ALREADY
// classified — good, endpoint-bearing LastError + host Stopped — before rethrowing. The poll engine
// then routes that same exception here. Re-degrading would overwrite the good LastError with a worse,
// guidance-free one (and, for a provider whose text echoes credentials, re-open the C1 leak) and log
// the outage a second time. So skip the connection class ReadAsync owns; the only exceptions that
// reach past this guard are the engine's own contract-violation throws — an InvalidOperationException
// carrying no provider text — which nothing else classifies.
// guidance-free one and log the outage a second time, so skip the connection class ReadAsync owns.
if (ex is DbException) return;
// Whatever reaches here — an engine contract-violation, or a non-DbException the provider raised
// (e.g. an ArgumentException from a malformed connection string, which the parser echoes with
// credential fragments) — the operator surface carries only the exception TYPE, never ex.Message.
// The full exception, with its text, still reaches the log SINK via the exception parameter. This is
// the same unconditional rule InitializeAsync/ReadAsync follow; do not weaken it to a type check.
_logger.LogWarning(ex, "Sql poll failed. Driver={DriverInstanceId} Endpoint={Endpoint}",
_driverInstanceId, Endpoint);
Degrade(ex.Message);
Degrade($"Sql driver '{_driverInstanceId}' poll failed: {ex.GetType().Name}.");
}
/// <summary>
@@ -594,6 +655,94 @@ public sealed class SqlDriver
}
Volatile.Write(ref _tagsByRawPath, table);
// Nothing is pollable until the catalog gate has validated it. Clearing here (rather than leaving a
// previous generation's validated table in place) means a Reinitialize whose gate then fails cannot
// keep polling the OLD identifiers against a database whose schema may be exactly what changed.
Volatile.Write(ref _polledByRawPath, new Dictionary<string, SqlTagDefinition>(StringComparer.Ordinal));
}
// ---- catalog gate (design §8.1) ----
/// <summary>
/// Validates every authored identifier against the live catalog and republishes the tag table with
/// catalog spellings, dropping the tags that do not resolve (design §8.1).
/// </summary>
/// <remarks>
/// <para><b>Why this must run before the driver reports Healthy.</b> The table this swaps in is what
/// every poll resolves against. Running it later — or in the background — would leave a window in
/// which an unvalidated identifier could be quoted into a query, which is the entire hole the gate
/// exists to close.</para>
/// <para><b>A dropped tag is not a driver fault.</b> It disappears from the table, so its RawPath
/// resolves to nothing and it publishes <see cref="SqlStatusCodes.BadNodeIdUnknown"/> — one operator
/// typo must not stop the other tags on that database, the same rule
/// <see cref="BuildTagTable"/> follows for a malformed blob. Each drop is logged at Warning with the
/// tag, the field and the reason, because a node that silently goes Bad with nothing in the log is
/// unsupportable.</para>
/// <para><b>No tags ⇒ no I/O.</b> A driver with nothing authored has nothing to validate, and issuing
/// catalog queries to prove that would be a round-trip that can only fail.</para>
/// <para>Bounded by the same wall-clock discipline as the liveness check (R2-01 / STAB-14): the work
/// runs on the thread pool so a provider implementing the async path synchronously blocks a pool
/// thread rather than wedging Initialize forever with no retry.</para>
/// </remarks>
/// <param name="cancellationToken">The caller's token.</param>
/// <returns>A task that completes when the tag table has been validated and republished.</returns>
private async Task ApplyCatalogGateAsync(CancellationToken cancellationToken)
{
var authored = Tags;
if (authored.Count == 0) return;
var tables = authored.Values
.Select(definition => definition.Table)
.Where(table => !string.IsNullOrWhiteSpace(table))
.Distinct(StringComparer.OrdinalIgnoreCase)
.ToArray();
var catalog = await RunBoundedAsync(
token => LoadCatalogAsync(tables, token),
_options.OperationTimeout,
"the catalog queries did not return",
cancellationToken).ConfigureAwait(false);
var result = SqlCatalogGate.Apply(authored.Values, catalog, _dialect);
foreach (var rejection in result.Rejected)
{
_logger.LogWarning(
"Sql tag '{Tag}' rejected by the catalog gate: {Field} — {Reason}. It will publish "
+ "BadNodeIdUnknown until the tag or the database is corrected. Driver={DriverInstanceId}",
rejection.RawPath, rejection.Field, rejection.Reason, _driverInstanceId);
}
var validated = new Dictionary<string, SqlTagDefinition>(result.Accepted.Count, StringComparer.Ordinal);
foreach (var definition in result.Accepted) validated[definition.Name] = definition;
// Only the POLLED table is replaced. The authored table is left intact so every authored tag keeps
// its node and a rejected one publishes BadNodeIdUnknown rather than vanishing (see _polledByRawPath).
Volatile.Write(ref _polledByRawPath, validated);
_logger.LogInformation(
"Sql driver {DriverInstanceId} catalog gate accepted {Accepted} of {Total} authored tag(s) "
+ "across {Tables} table(s) on {Endpoint}.",
_driverInstanceId, result.Accepted.Count, authored.Count, tables.Length, Endpoint);
}
/// <summary>Opens one connection and reads the catalog slice the authored tables need.</summary>
private async Task<SqlCatalog> LoadCatalogAsync(
IReadOnlyCollection<string> authoredTables, CancellationToken cancellationToken)
{
var connection = _factory.CreateConnection()
?? throw new InvalidOperationException(
$"the {_dialect.Provider} provider factory returned no connection.");
await using (connection.ConfigureAwait(false))
{
connection.ConnectionString = _connectionString;
await connection.OpenAsync(cancellationToken).ConfigureAwait(false);
return await SqlCatalogLoader
.LoadAsync(connection, _dialect, authoredTables, _options.CommandTimeout, cancellationToken)
.ConfigureAwait(false);
}
}
// ---- liveness ----
@@ -611,18 +760,50 @@ public sealed class SqlDriver
/// </summary>
/// <param name="cancellationToken">The caller's token.</param>
/// <returns>A task that completes when the database has answered.</returns>
private async Task VerifyLivenessAsync(CancellationToken cancellationToken)
private Task VerifyLivenessAsync(CancellationToken cancellationToken) =>
RunBoundedAsync(
async token => { await PingAsync(token).ConfigureAwait(false); return true; },
_options.CommandTimeout,
"the liveness statement did not return",
cancellationToken);
/// <summary>
/// Runs <paramref name="work"/> under a hard wall-clock <paramref name="budget"/>, on the thread pool,
/// converting a breach into a <see cref="TimeoutException"/> worded from <paramref name="what"/>.
/// </summary>
/// <remarks>
/// <para>The R2-01 / STAB-14 shape, shared by the two pieces of I/O Initialize performs. It exists
/// because a token is not a deadline: some ADO.NET providers implement the async path synchronously,
/// and a wedged socket can hang <em>inside</em> the provider's own cancellation handshake. Without an
/// independent wall clock, <c>DriverInstanceActor</c>'s init task would never complete and the driver
/// would sit in Connecting forever with no retry — the exact frozen-peer wedge the S7 driver shipped.</para>
/// <para>Both the linked CTS deadline and an outer <see cref="TaskExtensions"/> wait are used, because
/// each covers a case the other does not: the CTS handles a provider that <em>does</em> honour
/// cancellation, and the outer wait handles one that does not. An abandoned attempt keeps ownership of
/// its own connection and disposes it; <see cref="Detach"/> observes its eventual fault.</para>
/// </remarks>
/// <typeparam name="T">The work's result type.</typeparam>
/// <param name="work">The operation to bound. Receives the deadline-linked token.</param>
/// <param name="budget">The wall-clock allowance.</param>
/// <param name="what">Sentence fragment naming the operation, e.g. "the catalog queries did not return".</param>
/// <param name="cancellationToken">The caller's token.</param>
/// <returns>The work's result.</returns>
/// <exception cref="TimeoutException">The budget elapsed first.</exception>
private static async Task<T> RunBoundedAsync<T>(
Func<CancellationToken, Task<T>> work,
TimeSpan budget,
string what,
CancellationToken cancellationToken)
{
var budget = _options.CommandTimeout;
var deadline = CancellationTokenSource.CreateLinkedTokenSource(cancellationToken);
Task work;
Task<T> running;
try
{
deadline.CancelAfter(budget);
work = Task.Run(
running = Task.Run(
async () =>
{
try { await PingAsync(deadline.Token).ConfigureAwait(false); }
try { return await work(deadline.Token).ConfigureAwait(false); }
finally { deadline.Dispose(); }
},
CancellationToken.None);
@@ -636,20 +817,18 @@ public sealed class SqlDriver
try
{
await work.WaitAsync(budget, cancellationToken).ConfigureAwait(false);
return await running.WaitAsync(budget, cancellationToken).ConfigureAwait(false);
}
catch (TimeoutException)
{
Detach(work);
throw new TimeoutException(
$"the liveness statement did not return within {(int)budget.TotalMilliseconds} ms.");
Detach(running);
throw new TimeoutException($"{what} within {(int)budget.TotalMilliseconds} ms.");
}
catch (OperationCanceledException) when (!cancellationToken.IsCancellationRequested)
{
// Our deadline fired and the provider DID honour the linked token.
Detach(work);
throw new TimeoutException(
$"the liveness statement did not return within {(int)budget.TotalMilliseconds} ms.");
Detach(running);
throw new TimeoutException($"{what} within {(int)budget.TotalMilliseconds} ms.");
}
}
@@ -752,6 +931,7 @@ public sealed class SqlDriver
{
await _poll.DisposeAsync().ConfigureAwait(false);
Volatile.Write(ref _tagsByRawPath, new Dictionary<string, SqlTagDefinition>(StringComparer.Ordinal));
Volatile.Write(ref _polledByRawPath, new Dictionary<string, SqlTagDefinition>(StringComparer.Ordinal));
_resolver.Clear(); // no-op in v3 (the resolver reads the live table by closure); kept for symmetry
}
}
@@ -29,12 +29,9 @@ namespace ZB.MOM.WW.OtOpcUa.Driver.Sql;
public static class SqlDriverFactoryExtensions
{
/// <summary>
/// The driver-type string this factory registers under.
/// <para>Chained off <see cref="SqlDriver.DriverTypeName"/> rather than <c>DriverTypeNames.Sql</c>:
/// <c>DriverTypeNamesGuardTests</c> asserts bidirectional parity between the shared constants and the
/// factories actually registered in the process, so the constant may only be added in the change that
/// also wires this factory into the Host (the plan's Task 11, which repoints both at
/// <c>DriverTypeNames.Sql</c>).</para>
/// The driver-type string this factory registers under — chained off
/// <see cref="SqlDriver.DriverTypeName"/>, which is
/// <see cref="ZB.MOM.WW.OtOpcUa.Core.Abstractions.DriverTypeNames.Sql"/>.
/// </summary>
public const string DriverTypeName = SqlDriver.DriverTypeName;
@@ -44,6 +44,12 @@ public sealed class SqlServerDialect : ISqlDialect
public string ListSchemasSql =>
"SELECT DISTINCT TABLE_SCHEMA FROM INFORMATION_SCHEMA.TABLES ORDER BY TABLE_SCHEMA";
/// <summary>
/// <c>SCHEMA_NAME()</c> with no argument returns the calling principal's default schema — the one an
/// unqualified object name resolves in, which is precisely the question the catalog gate asks.
/// </summary>
public string DefaultSchemaSql => "SELECT SCHEMA_NAME()";
/// <inheritdoc/>
public string ListTablesSql =>
"SELECT TABLE_NAME, TABLE_TYPE FROM INFORMATION_SCHEMA.TABLES " +
@@ -70,9 +76,11 @@ public sealed class SqlServerDialect : ISqlDialect
/// control-character rule.</para>
/// <para>The rejection messages deliberately do <b>not</b> echo the offending value — it is untrusted
/// input and this exception's text reaches the driver log.</para>
/// <para><b>This is currently the only defence, not a backstop.</b> Design §8.1 specifies that an
/// identifier is validated against the live catalog before reaching here; that gate is not implemented
/// yet, so an authored <c>TagConfig</c> table/column arrives unfiltered.</para>
/// <para><b>This is a backstop, not the only defence.</b> Design §8.1's catalog gate
/// (<see cref="SqlCatalogGate"/>) resolves an authored table/column against the live catalog at
/// Initialize and substitutes the catalog's own spelling, so on the poll path the value arriving here
/// is a string read back out of the catalog. The rules below still run — the gate uses them as its own
/// charset filter, and they must hold for any future caller that has no catalog to check against.</para>
/// </remarks>
/// <param name="ident">The bare, unquoted identifier part.</param>
/// <returns>The bracket-quoted identifier.</returns>

Some files were not shown because too many files have changed in this diff Show More