64ec7aa43a696d9a32bf9e5e9b57a37088476db3
2787 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
64ec7aa43a |
feat(mqtt): Sparkplug birth-driven browser tree + scoped Request-rebirth action
Unseals the Sparkplug branch of the MQTT address-picker browser that Task 10 deliberately sealed shut, and adds the first (and only) sanctioned publish on a browse session. - MqttBrowseSession serves Group -> EdgeNode -> [Device] -> Metric in Sparkplug mode, decoded from observed NBIRTH/DBIRTH certificates (the only Sparkplug messages that name their metrics). Node-level metrics hang directly under their edge node -- no synthesised device folder, because the authored address tuple genuinely has deviceId = null for them. Metric node ids use a '::' separator: a metric name may contain '/', so slash-joining would make a node-level metric indistinguishable from a device folder. - AttributesAsync is purely birth-derived in Sparkplug mode. The plain path's UTF-8 inference / payload-snippet machinery does not run and no payload bytes are retained: a Sparkplug body is protobuf and would be reported as "(N bytes, binary)" for every metric in the plant. A datatype ToDriverDataType cannot map is reported as the Sparkplug type name, never coerced. - RequestRebirthAsync(scope) is the one publishing member, reached only by an explicit operator action. It routes through the counted PublishAsync chokepoint via a private RebirthTransport adapter, so PublishCountForTest still proves that OpenAsync/RootAsync/ExpandAsync/AttributesAsync/Observe/ DisposeAsync publish nothing -- now in both modes. A device or metric scope resolves up to its owning edge node (NCMD is node-addressed); group scope enumerates observed edge nodes and is refused whole past MaxGroupRebirthNodes = 32. - BrowserSessionService.RequestRebirthAsync gates on the same DriverOperator policy that gates the picker, evaluated server-side where the action happens rather than only at render time, fails closed, and Info-logs the scope and the published count. Exposed through the new IRebirthCapableBrowseSession contract so "does this session write?" stays a type question. - ToBrowseOptions' Last-Will landmine re-verified: MqttDriverOptions still carries nothing will-shaped (an NDEATH is a broker-published will and would be invisible to PublishCountForTest), now pinned by a reflection guard. Falsifiability: injecting a publish into RootAsync reddens the Sparkplug and plain passivity tests; double-publishing per target reddens all four one-NCMD-per-node assertions. 572/572 MQTT tests, 24/24 AdminUI browsing tests. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
6d7a458c4d |
fix(mqtt): reset Sparkplug ingest state at every session/authoring seam
Task 21 review — two Criticals sharing one root cause, plus four follow-ups.
C1/C2 root cause: `OnReconnectedAsync` opened with `Births.Clear(); _trackers.Clear();`
— "not housekeeping, the correctness step" — and that reset was missing from the
other two seams where the view underneath a surviving cache changes.
- C1: the ingestor OUTLIVES a session rebuild. `SessionIdentity` is a strict
superset of `IngestIdentity`, so changing only Host/Port/ClientId/credentials/
TLS/a timeout gives `!SameSession && SameIngest` — `ReinitializeAsync` tears the
session down and re-establishes on the same instance, and `TeardownAsync` touches
no ingest state (the resilience layer re-running `InitializeAsync` is the second
path). A node that rebound alias 5 during the changeover would publish Pressure's
value under the Temperature RawPath at Good. Hoisted into `ResetSessionState`,
called from `AttachTo` (earliest seam — a persistent session can push between
CONNACK and our SUBSCRIBE), `EstablishAsync` and `OnReconnectedAsync`.
- C2: `Register` pruned trackers but not births. While a node is unauthored its
traffic is dropped at `Dispatch`, so its cached birth FREEZES rather than
refreshing; drop-then-re-add across a week mis-routes at Good, and `Births` grew
monotonically. Now evicted on edge-node departure.
Deliberately NARROWER than "absent from ByScope": a device with no authored tags
under a still-authored node keeps receiving traffic, so its birth is live, not
frozen — evicting it would discard correct state. Pinned both ways.
- I1: a birth whose seq is unreadable produced a permanent debounce-paced rebirth
loop. "A birth arrived" is now tracked separately from "the birth established a
baseline". Writing the test showed the cap belonged wider than the review scoped
it: data-before-birth and unknown-alias loop identically (20 messages → 23 NCMDs
against a cap of 3), so all three missing-metadata reasons share one per-node
consecutive cap, re-armed by any applied birth. A seq gap stays uncapped — it
self-limits.
- I2: the oversize `WarnOnce` key was the raw topic, off a group-wide `#`
subscription. `HandleMessage` now parses and filters to authored nodes BEFORE the
size check (also skipping the protobuf parse for discarded traffic), and `_warned`
carries a hard ceiling so a future bad derivation degrades to silence.
- I3: array metrics slipped the unsupported-type gate — `ToDriverDataType` returns
the ELEMENT type, so a String-typed tag published a packed `byte[]` as base64 at
Good. Gated on `IsSparkplugArray`.
- I4: driver-level Sparkplug coverage for the composition-root lines T22 left
untested — the `OnDataChange` re-raise, `ReadAsync` off the shared cache,
subscribe/unsubscribe dispatch, and both directions of the mode gate.
Minors: M1 the oversize test fed unparseable bytes and passed with the size check
deleted (now a real NBIRTH + an at-the-ceiling control); M2 answered in-place (an
absent seq is refused WITHOUT becoming the baseline, so the NCMD assertion is the
complete check); M3 `HostStateObserved` raise wrapped; M5 the legacy `STATE/{host}`
form is now subscribed, so the parser's tolerance is reachable in production; M8
`ResolveBinding` prefers the NAME over a disagreeing alias; M9 a narrowing float
conversion refuses instead of publishing ±Infinity at Good (a publisher's own
infinity still passes through). M4 was already resolved by T22.
Also fixed a race in the new tests: `publish.Topics.Clear()` after a call that
dispatches rebirths off-thread must drain first — it made the C1 sequence-baseline
pin pass vacuously in the first falsifiability run.
Falsifiability: dropping the reset from AttachTo+EstablishAsync reddens 4 (C1);
dropping the Register eviction reddens 2 (C2); widening the eviction predicate to
the reviewer's literal phrasing reddens the keep-live-births pin (C2b). 545/545 MQTT
unit tests green (was 524); whole-solution `--no-incremental` build clean under TWAE.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
cd0157a3b8 |
feat(mqtt): typed tag editor Sparkplug mode + validation
Task 24 (P2): replaces the P1 Sparkplug Validate() stub (always null) and the "not available yet" editor placeholder with real groupId/edgeNodeId/deviceId/ metricName authoring, mirroring MqttTagDefinitionFactory.FromSparkplugTagConfig in both directions -- required ids + optional deviceId, dataType/qos read strictly but only when present (dataType is optional: the birth certificate declares it). One rule is deliberately stricter than the factory: group/edge- node/device ids reject '/', '+', '#' (topic-segment characters a decoded incoming id can never carry, so an authored one that does can never bind); metricName is exempt since Sparkplug's own canonical names use '/' (e.g. "Node Control/Rebirth"). No FullName key is written -- TagConfig carries no identity under v3, and the plan's snippet asking for one was corrected per the brief. A SparkplugB tag now survives save/reopen because its descriptor keys re-infer Mode on load; there was never a 'mode' key to persist. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
6917d8805d |
feat(mqtt): Sparkplug UntilStable discovery + rediscover-on-DBIRTH
RediscoverPolicy is now mode-dependent: UntilStable in Sparkplug B (a tag's dataType is optional -- the birth certificate declares it, so the discovered tree fills in asynchronously), Once in Plain (nothing about the authored set arrives on the wire). DiscoverAsync resolves each Sparkplug tag's datatype per pass from the live BirthCache, with the ingest path's own precedence: authored dataType wins, else the birth's, else the record default. An unsupported Sparkplug type (DataSet/Template/PropertySet/Unknown) falls back rather than blanking the tag. The authored tag SET is never conditional on a birth -- an authored tag is part of the declared configuration, not something the plant grants by publishing. SparkplugIngestor.BirthObserved is wired to RaiseRediscoveryNeeded behind a two-part change gate, which is the anti-storm mechanism rather than an optimisation: fire only for a scope an authored tag binds, and only when the birth's metric-name SET (ordered-distinct, ordinal) differs from the last one seen for that scope. Edge nodes re-birth freely -- on their own timer, on every reconnect via the late-join rebirth fan-out, and once per rebirth NCMD the gap policy sends -- and firing on each of those would make a healthy plant a permanent address-space-rebuild loop. No extra debounce: what survives the gate is a real edit to the served tree, and delaying it would need its own trailing-edge flush to avoid dropping the last change. The signature map is NOT cleared on death/reconnect/ingest rebuild (the ingestor drops its whole birth cache on reconnect, but "the driver forgot" is not "the address space changed"); it is pruned when a redeploy stops authoring a scope. ScopeHint is the "Mqtt" discovery folder, deliberately not the Sparkplug scope path the plan named: the discovered tree is flat, so an edge-node/device path would name a subtree that does not exist and a consumer scoping a rebuild on it would rebuild nothing. The scope is carried in Reason instead. Falsifiability: always-fire reddens the identical-rebirth + reordered-rebirth pins; always-authored-datatype reddens the birth-fill pin; dropping the authored-scope filter reddens the unauthored-device pin. 524/524 MQTT unit tests green (was 513). Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
8538fb798b |
feat(mqtt): Sparkplug ingest state machine (birth/alias/seq/rebirth/death→stale)
Task 21 — the design doc's §3.6 correctness core. `SparkplugIngestor` owns the
authored `(group, node, device?, metric) → RawPath` index, the live `BirthCache`,
one `SequenceTracker` per authored edge node, and the rebirth policy; it routes
NBIRTH/DBIRTH/NDATA/DDATA/NDEATH/DDEATH/STATE into `OnDataChange` +
`LastValueCache`, or into a bounded, per-node-debounced rebirth NCMD.
Invariants held (and each pinned by a falsifiability control):
- Published reference is the RawPath, never a composed Sparkplug reference —
`DriverHostActor`'s dual-namespace fan-out is RawPath-keyed.
- Tags bind by stable metric NAME; the alias is only the per-birth lookup, so an
alias reused across a rebirth cannot mis-route.
- Every value goes through `SparkplugMetricBinding.Reinterpret` before publish.
- NDEATH never reaches `SequenceTracker.Accept` (it carries no seq); it is tied to
its birth by bdSeq instead.
- An invalid payload mutates nothing — no tracker, no alias table, no eviction.
- `HandleMessage` never throws on MQTTnet's dispatcher thread.
Supporting changes:
- `MqttTagDefinitionFactory.FromSparkplugTagConfig` — the driver had no way to
parse the Sparkplug descriptor keys at all; the P1 stub fields on
`MqttTagDefinition` were never populated. Mode selects the parser, never a
blob heuristic. `dataType` becomes optional in the Sparkplug shape (the birth
declares it), recorded via the new `DataTypeAuthored` flag.
- `MqttConnection` implements `IMqttPublishTransport` (bounded, refusal throws).
- `MqttDriver` dispatches every capability by mode, subscribes `spBv1.0/{group}/#`
(+ the configured STATE topic) once at connect, and re-subscribes + requests a
late-join rebirth on reconnect. `IngestIdentity` now names `MqttSparkplugOptions`,
closing the silent same-ingest gap its own ⚠️ comment warned about.
Primary-host STATE *publishing* is explicitly out of scope and warned about at
construction: it needs the retained-OFFLINE Last Will set at CONNECT time, and
shipping the ONLINE half alone is worse than shipping neither.
58 new tests (MQTT suite 455 → 513, all green).
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
31a98a1bf5 |
feat(mqtt): RebirthRequester NCMD encode
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
3fdf24659f |
fix(mqtt): AcceptBirth is NBIRTH-only -- a DBIRTH routed there wipes bdSeq
Self-review of the commit before this one caught a defect in the contract it published rather than in the code it ran. AcceptBirth's doc said it records "a birth (NBIRTH/DBIRTH)". That is wrong on both halves of this type, and the wrong version plants exactly the defect the bdSeq tie exists to prevent. Only the NBIRTH restarts the Sparkplug sequence; a DBIRTH carries the next number in the edge node's ongoing stream, so it belongs to Accept. And only the NBIRTH and NDEATH carry bdSeq at all -- a DBIRTH has none to pass. So a Task 21 that followed the old doc and routed a DBIRTH to AcceptBirth would do two wrong things at once: rebase the sequence mid-stream, and wipe the node's session token to null. A stale Last Will arriving afterwards would then find nothing to compare against, fall through the deliberate fail-toward-stale rule in IsDeathForCurrentSession, and kill the live node. Renamed AcceptBirth -> AcceptNodeBirth so the DBIRTH call site reads wrong at a glance rather than only in prose, documented the hazard on both methods, and added DeviceBirth_ContinuesTheNodeSequence_AndLeavesTheSessionTokenAlone to pin the shape at this type's own boundary -- the ingest state machine now has something to fail against if it wires the call sites the other way round. Falsifiability: mutating Accept to clear _birthBdSeq reddens exactly the new test (1 RED), reverted. 29 tests, 438/438 for the MQTT suite. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
e4e313cf1c |
feat(mqtt): SequenceTracker seq-gap + bdSeq death-tie
SequenceTracker (Driver.Mqtt/Sparkplug/) holds one edge node's Sparkplug stream-continuity state -- design doc SS3.6 invariant #3. Two independent halves: the wrapping 0-255 seq counter that reports whether a message was missed, and the bdSeq session token that ties an NDEATH to the NBIRTH it belongs to. Next-expected is (last + 1) & 0xFF, so 255 -> 0 is the sequence continuing. The boundary is the whole task because it fails differently in each direction: read the wrap as a gap and every edge node is asked to rebirth once per 256 messages, so a busy node spends its life republishing births instead of data; miss a real gap and the driver serves values it knows are behind the device, at Good quality, indefinitely. Both directions are asserted, plus two full wrap cycles so an off-by-one that only misfires at one value has nowhere to hide. A gap is reported once and the tracker then resynchronizes onto the observed value -- holding the baseline at the pre-gap number would make every following message report a gap too, so one lost message would become a permanent rebirth storm, the same pathology as a wrong wrap reached by a different route. An out-of-range seq is refused, never masked. The codec hands seq over as a ulong precisely so a spec-violating publisher's value is not truncated into a plausible-looking one, and masking would finish the job it declined to do: with the baseline at 255, a seq of 256 masks to 0 -- exactly what the wrap expects next -- and the bogus message would be accepted as contiguous. Refused values do not become the baseline either, so the next genuine message is still measured against the last credible one. A birth restarts the sequence via AcceptBirth rather than Accept, so it is never itself a gap; a gap after a birth is still detected. An in-range but non-zero birth seq is adopted rather than refused -- refusing it would demand a fresh rebirth for every message from an otherwise-followable publisher. The first message on a virgin tracker cannot be a gap (nothing to measure it against) but leaves IsBirthSynchronized false, which is how Task 21 tells a mid-stream late join from a synchronized node. bdSeq exists to stop a late Last Will from killing a live node: a node's connection drops, it reconnects and publishes a fresh NBIRTH, and only then does the broker notice the old connection died and deliver its will carrying the PREVIOUS session's bdSeq. Without the compare that stale NDEATH drives every tag of a freshly-born healthy node to Bad. Unknowns fail toward stale deliberately -- only a positive mismatch of two known tokens discards a death, because acting on a death wrongly self-corrects at the next birth (invariant #4) while ignoring a real one leaves a dead node's values flowing Good forever. bdSeq is not a top-level field, so TryReadBdSeq extracts it from the metric named bdSeq and runs it through ReinterpretSigned first: read raw, a negative Int64 bdSeq becomes an enormous ulong -- a plausible-looking session token minted out of nonsense. Falsifiability verified, each mutation reverted after observing RED: (a) no-wrap expected -> 2 RED; (b) Accept always true -> 5 RED; (c) bdSeq compare always matches -> 2 RED; (d) mask out-of-range seq into range -> 2 RED. 28 tests, 437/437 for the MQTT suite. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
2afef00eaa |
feat(mqtt): AliasTable/BirthCache — bind-by-name, rebuild-per-birth
Design §3.6 invariants #1/#2. A Sparkplug alias is scoped to ONE birth: after a rebirth an edge node may point alias 5 at a different metric. So authored tags bind by the stable metric NAME and the alias is a per-birth cache only — RebuildFromBirth builds both indexes fresh and publishes them in one assignment, never merging or upserting, including for an empty birth (which legitimately clears the scope). BirthCache keys scopes by the full (group, edgeNode, device?) triple and applies the spec's node→device rules: an NBIRTH invalidates every DBIRTH under that node (returning their metric sets for the STALE fan-out), a death evicts rather than blanks so data-before-rebirth stays detectable. Reads are lock-free on MQTTnet's dispatcher thread — an immutable snapshot swapped atomically per scope, the same discipline MqttSubscriptionManager.AuthoredTable uses, with rebuilds landing in place so a held table reference always observes the newest birth. Running values are deliberately NOT stored here — LastValueCache owns those, keyed by RawPath. SparkplugMetricBinding.Reinterpret exposes the birth datatype to the codec's two's-complement fix, since a DATA metric carries no datatype. Falsifiability: mutating RebuildFromBirth to merge reddened 5/19 tests (both plan invariants); making the alias carry metric identity across a birth reddened 3/19, including the headline alias-reuse case. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
b245f380b1 |
feat(mqtt): Sparkplug topic parse/format + datatype map
SparkplugTopic (Driver.Mqtt/Sparkplug/) parses/formats spBv1.0 topics for
every message type (NBIRTH/DBIRTH/NDATA/DDATA/NDEATH/DDEATH/NCMD/DCMD/STATE).
TryParse never throws -- it is fed every topic a live spBv1.0/{group}/#
subscription delivers -- and rejects device/node-scope mismatches and MQTT
wildcard chars in a segment. STATE is handled honestly rather than force-fit
into the {group}/{type}/{node} mould: it parses both the v3.0
spBv1.0/STATE/{hostId} form and the legacy no-prefix STATE/{hostId} form
(tolerated on receive only -- Format/FormatState always emit v3.0). Format
gives Task 20's NCMD/DCMD write path a builder instead of hand-concatenation.
SparkplugDataType is a `global using` alias for the vendored proto's
generated Org.Eclipse.Tahu.Protobuf.DataType, not a second hand-duplicated
enum -- Metric.Datatype is a raw wire uint32 (no enum-typed field forces a
second CLR type to exist), and SparkplugCodec (Task 16, landed concurrently)
already casts straight to the generated type. A duplicate enum would be the
same enum-drift hazard this repo already names systemic (CLAUDE.md's driver
enum-serialization bug) and would force every downstream task to cast
between two value-compatible-but-nominally-different enums. ToDriverDataType()
maps per design doc SS3.5: Int8/UInt8 widen to Int16/UInt16 (no 8-bit
DriverDataType member), Float/Double to Float32/Float64 (there is no
DriverDataType.Double), Text/UUID/Bytes/File to String, *Array variants to
their element type (IsSparkplugArray carries the array bit separately), and
DataSet/Template/PropertySet/PropertySetList/Unknown return null -- an
explicit "unsupported, caller must skip+warn" rather than a guessed String.
The completeness test enumerates the live generated DataType member set
(via the alias) and asserts every member is mapped or on the explicit
unsupported list, so a future Tahu proto change is caught automatically
instead of silently falling through a stale duplicate enum's default.
Falsifiability verified by hand for three defect shapes (each reverted after
observing RED): wrong Int8 widening, a dropped mapping falling through
undetected, and inverted node/device topic scoping.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
c164ec3da1 |
fix(mqtt): SparkplugCodec's never-throw guard must span the projection too
The try/catch stopped at Payload.Parser.ParseFrom, leaving the metric-projection loop outside it — so the "never throws for any input" contract had a hole in exactly the part that walks an attacker-shaped object graph. Widened to cover parse AND projection. Self-review finding; no test reddened, because a projection escape needs a generated-code shape the vectors do not produce. That is the point: the guard is the contract, not the observed behaviour of today's inputs. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
2589774480 |
feat(mqtt): SparkplugCodec decode + golden payload vectors
Decodes Sparkplug-B wire bytes into a driver-side projection (SparkplugPayload /
SparkplugMetric / SparkplugValueKind) for the Task 21 ingest state machine.
- Never throws, for ANY input. It sits on MQTTnet's shared dispatcher thread, so
an escaping exception would stall delivery for every subscription on the
connection, not one tag. Garbage, truncation, zero-length, a valid protobuf of
another schema and an over-nested Template all resolve to a verdict.
- Zero-length input is INVALID, not "a payload with no metrics" — protobuf would
parse it as all-defaults, and that is how a truncated-to-nothing body gets
mistaken for a well-formed one.
- Explicit presence throughout (Has{Seq,Name,Alias,Datatype,Timestamp}), never a
zero-check: an NBIRTH legitimately carries seq = 0, and every DATA metric after
a birth carries an alias with no name and no datatype.
- Values are projected RAW, boxed, with the value oneof reported as an explicit
ValueKind — Absent / Null / Scalar / Unsupported all mean different things and
three of them carry a null value. DataSet/Template/extension decode as
Unsupported (v1 scope) rather than throwing or silently vanishing.
- ReinterpretSigned() undoes Sparkplug's two's-complement-in-an-unsigned-field
encoding of Int8/16/32/64. Kept out of decode because a DATA metric carries no
datatype — only the consumer, holding the birth's alias table, knows which
applies. Skipping it publishes 4294967254 for a tag whose value is -42.
- Datatype is carried as the generated Org.Eclipse.Tahu.Protobuf.DataType; the
SparkplugDataType map is Task 17's and is applied downstream (Task 21).
Golden vectors: nbirth.bin (seq=0, named/aliased catalog, a negative Int32, an
is_null metric, a DataSet metric) + ndata.bin (alias-only metrics, no name, no
datatype) are hand-built by SparkplugGoldenPayloads, committed, and pinned by a
drift guard that rebuilds and byte-compares them on every run — plus a test that
they actually reach the output directory, since a missing copy item is invisible
in source.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
163dd7ab5c |
fix(mqtt): unbreak the .Contracts csproj -- '--' is illegal in an XML comment
The build-platform note added alongside the proto wiring quoted docker-dev's `FROM --platform=linux/amd64` verbatim inside an MSBuild comment. XML forbids '--' in a comment, so MSBuild refused to load the project at all (MSB4025) -- a total build stop for every consumer of .Contracts, i.e. the whole MQTT driver, the Host, and the test suite. Reworded to name the platform in prose. Caught by rebuilding after the edit; the preceding commit was made from a build that predated it. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
1ae26675ac |
feat(mqtt): vendor Tahu sparkplug_b.proto + Grpc.Tools codegen
Vendors the NORMATIVE (proto2) Eclipse Tahu sparkplug_b.proto and compiles it
in .Contracts with build-time Grpc.Tools, message-only (GrpcServices="None") --
Sparkplug rides MQTT, there is no gRPC service, so no Grpc.Core.Api is pulled in.
Provenance rides in the file header: repo, path, pinned commit
5736e404889d4b95910613040a99ba79589ffb13, permalink, git blob SHA-1
bf72ab5f..., content SHA-256 4432c5c4..., EPL-2.0, and a re-verify recipe. The
body below line 38 is byte-for-byte upstream; the blob SHA-1 matches the tree
entry at that commit, so the copy is provably genuine and not reconstructed.
Chose the proto2 file over upstream's sibling sparkplug_b_c_sharp.proto: the
sibling is a lossy proto3 restatement that drops explicit presence, and protoc
generates valid C# from proto2 into the identical namespace
(Org.Eclipse.Tahu.Protobuf, PascalCased from the package -- no
csharp_namespace option). Presence matters downstream: an NBIRTH legitimately
carries seq = 0 and a DATA metric legitimately omits `name`, so Has{Seq,Name,
Alias,IsNull,Datatype} is the difference between "absent" and "present, zero".
.Contracts stays transport-free: Google.Protobuf is a serialization dependency
with a framework-only graph, Grpc.Tools is PrivateAssets=all, and the resolved
graph is exactly {Google.Protobuf, Grpc.Tools, Core.Abstractions} -- no MQTTnet.
Codegen tests pin the round-trip, the namespace, the Sparkplug field numbers,
the DataType spec indices, and protoc's enum-name mangling (UInt64 -> Uint64,
UUID -> Uuid) that Task 17's datatype map has to spell correctly.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
92ba120964 |
feat(mqtt): MqttDriverForm + MqttDeviceForm — the driver/device config forms P1 omitted
The P1 live gate had to insert MQTT driver config directly via SQL: DriverConfigModal and DeviceModal both fell through to "No typed config form for driver type Mqtt" with no raw-JSON fallback, so broker host/port/TLS/credentials were unauthorable from the AdminUI. Task 12 built the *tag* editor; these are the driver + device surfaces. MqttDriverForm authors the whole MqttDriverOptions connection surface (host/port/clientId, TLS + CA pin, credentials, protocol version/clean-session/keep-alive, connect timeout, reconnect backoffs, mode, maxPayloadBytes, and the Plain sub-object). The Sparkplug sub-object stays P2 — a marked placeholder mirroring MqttTagConfigEditor's stub, with any existing sparkplug keys preserved untouched. The connection is authored on the DRIVER, not the device — unlike Modbus/S7/OpcUaClient and contrary to what docs/drivers/Mqtt.md said. DriverDeviceConfigMerger merges a device's keys up only when the driver has exactly ONE device, so a device-authored broker connection vanishes silently the moment a second device is added. MqttDeviceForm is therefore informational (the GalaxyDeviceForm shape) and round-trips DeviceConfig verbatim; it flags legacy connection keys left on a device by a pre-form deployment, because those still win via merge-up. Serialization goes through the ONE shared MqttJson.Options instance (Task 9's decision), never a fifth per-form copy — pinned by an assertion that an ordinal-only reader CANNOT bind the emitted blob, and by extending the fleet-wide DriverPageJsonConverterTests guard to resolve a form's serializer from an explicit external-instance registry when it has no _jsonOpts field. Dropping the converter from MqttJson.Options reddens 3 tests and leaves the symmetric round-trip green — the Task 9 lesson, reproduced. RawTags is stripped from the emitted blob (the deploy artifact owns it) and unknown top-level keys survive a load->save. Validation mirrors the driver's own [Range] bounds inline, and ToOptions() additionally clamps, so an ignored error still cannot persist a connectTimeoutSeconds:0 driver-brick. A non-blank clientId raises the P1 live-gate warning inline (fixed ids make redundant pair nodes evict each other while both report Healthy). AdminUI 761/761 green (was 730), TWAE-clean; Driver.Mqtt 266/266 green. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
07a4a5ff5e |
docs(mqtt): P1 plain-MQTT milestone live-verified; endpoint recorded
Task 14 — the P1 milestone gate. Ran the live /run verification on a docker-dev
rig against the real Mosquitto TLS+auth fixture, and recorded the result.
The gate did its job: it found MQTT was unauthorable in production.
RawDriverTypeDialog's driver-type list is a hand-maintained array that nobody
added MQTT to, so the /raw "New driver" picker never offered it — with every
other layer (contracts, driver, browser, factory, host registration, typed tag
editor) complete and 266 green unit tests. Nothing tied that array to
DriverTypeNames and AdminUI has no bUnit, so no offline test could see it.
Fixed, plus RawDriverTypeDialogParityTests: a reflection parity guard that
fails for ANY DriverTypeNames entry missing from the picker. Verified
falsifiable — RED with "…cannot be authored from the /raw New-driver picker,
so they are unreachable in production: Mqtt" before the one-line fix.
A second gap is left OPEN and documented, not fixed (no task ever covered it):
MqttDriverForm / MqttDeviceForm do not exist, and DriverConfigModal/DeviceModal
have no raw-JSON fallback — so an operator can create an MQTT driver but cannot
author its broker endpoint or credentials. P1 is not operator-complete until
those two razor forms land. The gate authored config via SQL to get past it.
Live gate result (isolated 2-node MAIN stack, image built from this branch,
fixture at 10.100.0.35:8883, AllowUntrustedServerCertificate rather than
pinning the CA into the rig):
- typed MQTT tag editor renders and round-trips; Json shows the JSON-path
field, Raw/Scalar hide it; wildcard topic a/+/c blocked at save; data-type
list offers Float64 and no Double; qos/retainSeed absent on "(driver
default)" and qos:2 written as a number; isHistorized/historianTagname
survive a topic-only edit; SparkplugB renders the P2 placeholder and
switching back keeps Plain values; jsonPath is guidance not a gate in all
three shapes ($ seeded only for a brand-new blank tag, blank stays absent,
clearing saves fine)
- deployed, and both editor-written blobs resolved and served changing live
values through OPC UA (22.8 / 1629, StatusCode Good) — no BadNodeIdUnknown
- the bespoke #-observation browser drove a real broker session and rendered
the fixture's actual topic tree (line1/{counter,state,temperature},
noise/chatter, retained/seed); a Modbus device's picker still correctly
reports "Browsing unavailable"
Also found live and documented (not a code change): both nodes of a redundant
pair run the driver, so a FIXED clientId makes them evict each other forever —
the broker logs "already connected, closing old connection" and both nodes
reconnect every ~2s while still reporting Healthy. Unset (the default) is
correct; confirmed 0 reconnects/60s on both nodes after removing it.
Suites: offline 1134 passed / 0 failed (266 Driver.Mqtt + 730 AdminUI incl. the
3 new guards + 138 Core.Abstractions); live 7/7 against the fixture.
Docs: new docs/drivers/Mqtt.md + Mqtt-Test-Fixture.md (the per-driver + fixture
convention every other driver follows), MQTT rows in docs/drivers/README.md,
TestConnectProbes.md, root README.md, and infra/README.md §3.
CLAUDE.md drift corrected while adding the MQTT endpoints:
- the "every fixture carries project: lmxopcua" claim was false — only the
MQTT fixture does; the filter returns one stack, not the fleet
- lmxopcua-fix.ps1 is Windows-VM-only; on macOS drive the host over ssh/rsync
(and the docker-dev rig itself runs locally under OrbStack)
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
2409d05a28 |
fix(mqtt): a broker-refused CONNECT no longer reports a Healthy driver
MQTTnet 5 does not throw on an unsuccessful CONNACK — it returns the reason in MqttClientConnectResult.ResultCode and v5 removed v4's ThrowOnNonSuccessfulConnectResponse, so ConnectCoreAsync's unconditional SetState(Connected) sealed a wrong-password deployment green: DriverState.Healthy / HostState.Running on a session that never authenticated. Caught by the Task-13 Mosquitto fixture; 240 offline tests missed it. ConnectCoreAsync now inspects the CONNACK and raises MqttConnectRejectedException. Credentials/identity/protocol/Last-Will refusals are unrecoverable and set Faulted, which stops the reconnect supervisor; transient refusals (ServerUnavailable, ServerBusy, QuotaExceeded, UseAnotherServer, ConnectionRateExceeded, and anything unrecognised) stay Reconnecting under backoff. MqttDriverProbe now shares the rejection wording, so the AdminUI Test-connect button and the running driver can no longer disagree. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
033b3700c4 |
test(mqtt): Mosquitto TLS+auth fixture + env-gated plain live suite
Adds the MQTT driver integration-test project: an eclipse-mosquitto fixture running auth AND TLS (allow_anonymous false on both listeners), a mosquitto_pub-based JSON publisher emitting retained messages, a cert/password generator script whose output is gitignored, and a live suite gated on MQTT_FIXTURE_ENDPOINT that skips clean offline. The fixture is never anonymous by design: an anonymous broker would let the driver's TLS/CA-pin and auth paths go untested, and those fail silently. The CA-pin leg carries its own falsifiability control (a foreign CA generated in-process must be rejected). Live gate on 10.100.0.35: 7/7 passed. It found a driver defect, reported not fixed here (src/ is out of this task's scope): MqttConnection.ConnectAsync RETURNS NORMALLY when Mosquitto rejects a CONNECT as not-authorized -- MQTTnet 5 does not throw on an unsuccessful CONNACK, so a wrong broker password yields DriverState.Healthy from InitializeAsync. The auth-negative test therefore asserts the invariant that holds either way (no session is ever established) and documents the finding. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
ff62b62a55 |
refactor(mqtt): accept a blank jsonPath; guide with a seeded "$" instead
Review follow-up. Dropped the hard Validate rejection of a blank jsonPath under a Json payload: the runtime defaults it to the document root, which is the real "publisher puts a bare JSON scalar on the topic" case, so rejecting it made the editor refuse what the driver accepts — inverting "authoring surface accepts <=> publish accepts" — and blocked whole CSV-import batches, since this validator also gates RawManualTagEntryModal's review grid. Replaced with guidance: an UNAUTHORED tag (no topic AND no jsonPath) seeds the field with "$", so a fresh tag emits an explicit "jsonPath":"$" while an existing path-less tag keeps the key ABSENT and the driver defaults it. The wildcard-topic rule stays strict — a wildcard has no legitimate single-Tag interpretation. Also: omit a blank "topic" rather than writing "topic":"", and record why the _lastConfigJson guard diverges from the Modbus template — the landmine is a DERIVED, NON-PERSISTED UI FIELD INFERRED FROM BLOB CONTENT (Mode), which Task 24's Sparkplug field group will be tempted to add more of. Named the host-side dependency (RawTagModal only mutates through this callback) that keeps the staleness trade benign today. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
a13ae926a8 |
fix(mqtt): gate DisposeAsync against in-flight lifecycle ops (C1) + rebuild-branch tests (I1)
C1 (Critical) — DisposeAsync went straight to TeardownAsync with no lifecycle-gate
wait, reopening the orphaned-connection class. Interleaving: ReinitializeAsync's
rebuild branch holds the gate mid-InitializeCoreAsync, past its own teardown but
before `_connection = connection`; an ungated dispose sees a null connection, does
nothing, and sets `_disposed`; the rebuild then publishes a live connection plus a
host-probe loop that every later dispose short-circuits past. Not proven reachable
today (DriverInstanceActor.PostStop calls the gated ShutdownAsync), but MqttDriver
publicly implements IAsyncDisposable, which invites `await using`.
- DisposeAsync now takes the gate with the same bounded wait + fallback teardown
ShutdownAsync uses, and sets `_disposed` BEFORE the wait.
- InitializeCoreAsync re-checks `_disposed` after the connect and disposes the
connection it just built rather than publishing it — this closes the residual
window on the bounded-timeout fallback path.
- The doc comment no longer claims parity with ShutdownAsync it did not have, and
stops conflating "don't Dispose() the semaphore object" with "don't WaitAsync".
I1 (Important) — the SameSession == false rebuild branch had no test driving it.
Adds three: an endpoint-changing delta rebuilds and Faults on a refused connect;
an ingest-only delta (MaxPayloadBytes) ALSO rebuilds, pinning that IngestIdentity
is nested inside SessionIdentity; and DisposeAsync serializes behind a lifecycle
operation parked at a new internal BeforeConnectHookForTests seam (mirroring
MqttConnection's AfterConnectHookForTests, which exists for the same race class).
Minors: comments recording that IngestIdentity names no Sparkplug field and must
grow one in P2 (tasks 21/22), and why _options/_subscriptions/_authoredRawPaths
get looser memory discipline than _health/_hostState.
Falsifiability: reverting DisposeAsync to the ungated form reddens exactly the new
serialization test ("Shouldly.ShouldAssertException : raced"); forcing SameSession
to always-true reddens exactly the two new rebuild tests. 240/240 MQTT tests pass;
forced rebuild of the driver project is 0 warnings under TreatWarningsAsErrors.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
f79d13e2d8 |
feat(mqtt): factory + DriverTypeNames.Mqtt + host factory/probe registration
Task 9. Adds MqttDriverFactoryExtensions (DriverTypeName = DriverTypeNames.Mqtt, direct deserialization of MqttDriverOptions — no intermediate DTO, mirroring OpcUaClientDriverFactoryExtensions), the DriverTypeNames.Mqtt constant, and both host wiring sites: the factory in DriverFactoryBootstrap.Register and the probe in AddOtOpcUaDriverProbes (the admin-node path Program.cs calls in its hasAdmin block — a probe wired only on driver nodes makes Test-connect silently dead). Carried-forward items: - Converges every MQTT config-parsing seam onto ONE JsonSerializerOptions — MqttJson.Options, in .Contracts alongside MqttDriverOptions. It replaces the probe's internal JsonOpts and the browser's separate private copy; the factory, the probe, MqttDriver.ParseOptions and MqttDriverBrowser now all parse through it. .Contracts is the only assembly all four consumers reference, and the browser's reference to the runtime .Driver project is a documented layering exception scheduled for removal — anchoring the options there would resurrect the duplicate the day it goes away. - Replaces the "Mqtt" literals in MqttDriverProbe, MqttDriverBrowser and MqttDriver with the constant (string value unchanged). - Tightens MqttDriverProbeTests.ProbeAsync_EnumAsName: it asserted only Ok == false + non-empty message, which is also exactly what a JSON-parse failure produces — so it stayed green under the very regression it names. It now asserts the probe got past the parse and reached the network. Falsifiability: deleting JsonStringEnumConverter from MqttJson.Options reddens 9 tests across 4 suites, including the tightened probe test (message becomes "Config JSON is invalid: The JSON value could not be converted to MqttProtocolVersion") — which the pre-fix assertions would have passed. Also references the MQTT driver from Core.Abstractions.Tests so DriverTypeNamesGuardTests' reflective bin scan discovers the new factory and the constant/factory parity check stays honest. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
36abd871a1 |
feat(mqtt): typed tag editor + validator (plain mode)
Thin typed model over a preserved JsonObject key bag, mirroring the sibling <Driver>TagConfigModel template. Key names and strictness rules mirror MqttTagDefinitionFactory rather than inventing an editor-side schema; enums round-trip as NAMES (the factory reads them strictly by name). No FullName key: under v3 a tag is identified by its RawPath, which the factory keys the definition's Name off — a composed identity key here would be dead weight nothing reads. Mode is a UI-only sub-shape selector inferred from the blob (any Sparkplug descriptor key present) and never serialised — the driver takes Plain vs SparkplugB from its DRIVER config. Only Plain Validate() is implemented; the Sparkplug branch is a stub Task 24 fills, and its keys are preserved untouched. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
420692b6e5 |
feat(mqtt): MqttDriver shell — lifecycle + authored-only discovery (Once)
Composes MqttConnection + MqttSubscriptionManager + LastValueCache into the IDriver / ITagDiscovery / ISubscribable / IReadable / IHostConnectivityProbe / IRediscoverable capability set. - Discovery replays ONLY the authored raw tags (RediscoverPolicy = Once, SupportsOnlineDiscovery = false) — a chatty broker can never auto-provision. - Composition order is register -> AttachTo -> ConnectAsync; the manager's Reconnected handler is passed through unwrapped so its throw-on-total-failure still tears a deaf session down. - ReinitializeAsync applies a tag-only delta in place and never faults on a bad one; a session-changing delta rebuilds. - ReadAsync serves the last-value cache and degrades per reference (BadWaitingForInitialData), never throwing for the batch. - FlushOptionalCachesAsync is a no-op: the last-value cache IS IReadable. - Adds MqttDriverOptions.RawTags (the authored-tag delivery mechanism every other driver's options DTO already has) and promotes MaxPayloadBytes from a manager ctor knob to an operator-facing key. - Converges on MqttDriverProbe.JsonOpts rather than a second, divergent JsonSerializerOptions. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
211c6ea6d6 |
feat(mqtt): register MqttDriverBrowser (bespoke-first for Mqtt)
Wires the bespoke MqttDriverBrowser into AdminUI's IDriverBrowser registrations alongside OpcUaClient/Galaxy, overriding the universal DiscoveryDriverBrowser fallback for driverType "Mqtt". Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
a856d61510 |
feat(mqtt): MqttDriverProbe CONNECT handshake
Test-connect probe for the Mqtt driver type: parses MqttDriverOptions
(shared enum-as-name JsonSerializerOptions), opens a bounded MQTT CONNECT
via MqttConnection.BuildClientOptions with a distinct -probe-{guid8}
client-id suffix, and classifies the outcome. Live-probed against
MQTTnet 5.2.0.1603 to confirm the real contract: a broker-rejected
CONNACK is a MqttClientConnectResult.ResultCode, never a thrown
exception, and timeout classification checks the deadline token itself
rather than switching on exception type (which varies unpredictably
between MqttConnectingFailedException and bare OperationCanceledException
for the same frozen-peer shape).
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
abf3116bd4 |
fix(mqtt): wildcard tags went dark on reconnect — index by filter, not topic
C1 (CRITICAL). AuthoredTable routed any wildcard-authored tag into `Wildcards`
and never into the concrete topic index. But the two paths that reconstruct
state from a SUBSCRIBED FILTER STRING — OnReconnectedAsync's re-subscribe set
and DegradeTopic — both looked up that concrete-only index, and
`_subscribedTopics` keys on the wildcard PATTERN ("f/+/temp"). So:
- reconnect resolved zero filters for a wildcard tag, hit a silent early
return, and issued NO subscribe. The cache keeps its last Good value, so the
tag never turns Bad — it just stops updating forever behind a connection
reporting healthy. Exactly the silent-death mode MqttConnection's own remarks
warn about, left unguarded for wildcards.
- DegradeTopic missed too, so a SUBACK rejection of a wildcard filter degraded
nothing and left the tag at BadWaitingForInitialData.
Fix: AuthoredTable now carries `ByFilter` (EVERY def, keyed by its authored
topic filter, wildcards included) alongside `ByExactTopic` (concrete only) and
`Wildcards`. Delivery matches an incoming published topic — which never carries
a wildcard — so it keeps using ByExactTopic + the comparer scan. Every path
keyed on a filter string uses ByFilter. The distinction is documented on the
record. The silent early return is now a Warning naming the orphaned topics.
Also in this pass:
- I2: bounded decode on the dispatcher thread. A body over `MaxPayloadBytes`
(default 1 MiB, ctor-settable) is refused BEFORE any GetString/JsonDocument
parse and degrades its own tags. Unbounded decode on the shared dispatcher is
paid by every subscription, not just the offending topic. NOTE: promoting this
to an operator-facing MqttDriverOptions key (+ WithMaximumPacketSize) needs a
.Contracts edit this task is scoped out of; flagged in the XML docs.
- I3: integral-valued reals now coerce to integer tags. JS/Python edge gateways
serialize an integer as 5.0, and TryGetInt32/int.TryParse are syntactic — an
Int32 tag fed by such a gateway silently never received data. Parsed via
decimal so integrality and range are exact across Int64/UInt64. A genuinely
fractional 5.5 is still refused, never rounded.
- I4: the JSON document is parsed ONCE per message and shared across the whole
fan-out (the documented "one document, one JSONPath per signal" shape was
re-parsing per tag). Zero-alloc via a ref struct when no Json tag matches.
- I5: pinned the documented JSON-string-holding-a-number coercion end-to-end.
- Minor: Register now prunes `_subscribedTopics` / `_handleByRawPath` entries for
tags a redeploy dropped, so a deleted topic stops being re-subscribed forever.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
4ec01bdfc1 |
fix(mqtt): bound browse observation against a chatty or malicious broker
Three review findings, all "unbounded work/memory from untrusted broker input" rather than correctness bugs - but this points at a live plant broker. 1. InferPayload decoded the WHOLE payload on MQTTnet's dispatcher thread before trimming to 64 chars. The node cap never engages here: a multi-MB blob on an already-known topic creates no node, it just updates one, so the cost repeats per message forever. Now at most PayloadInspectMaxBytes (512) are examined. A blind slice can cut a multi-byte sequence, which the strict decoder would report as "binary" - TrimToUtf8Boundary backs the window off to the last whole sequence so good UTF-8 is never mislabelled. A clipped payload is String by construction rather than inferred from a partial view. 2. The node cap bounded COUNT, not bytes. MQTT topics run to ~65KB, so 50k nodes near that ceiling is a multi-GB tree. Topics now bounded at 1024 chars and segments at 256, rejected whole (a truncated topic is a DIFFERENT topic the picker would commit as a tag address) and surfaced via an __oversized__ marker kept distinct from __truncated__ - the operator's remedy differs. 3. Control characters in a snippet would break the picker's single-line row; Trim() only strips the ends. Now scrubbed to spaces. Each guard was falsified independently by neutering it and confirming exactly one test reddens. That found a real gap: the oversized-topic test was passing via the SEGMENT bound, leaving the whole-topic bound untested - the test now uses many short segments so only the topic bound can reject it. Also records the .Driver project reference as a known, deliberate exception to the .Browser -> .Contracts pattern, with its cost (AdminUI gains Core + Polly + Serilog) and the clean fix (a leaf transport project; NOT a move into .Contracts, which is deliberately transport-free). Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
d421487bcd |
feat(mqtt): plain subscribe→OnDataChange with retained seed + ref resolver
MqttSubscriptionManager is the P1 plain-MQTT ingest path: it holds the authored RawPath → MqttTagDefinition table behind the shared EquipmentTagRefResolver, establishes one SUBSCRIBE per distinct authored topic, and turns each inbound message into a LastValueCache update plus an OnDataChange notification. Load-bearing choices, each pinned by a falsified test: - The published reference IS the RawPath (def.Name), never the topic and never the TagConfig blob — DriverHostActor's dual-namespace fan-out is RawPath-keyed, so any other key passes every unit test and delivers nothing in production. - One message fans out to EVERY tag on its topic; several tags routinely share a topic with different jsonPaths, and stopping at the first match silently starves the rest. - An unauthored topic raises nothing: MQTT ingest never auto-provisions. - Wildcard-authored topics (accepted by the parser, warned at deploy) match via MQTTnet's own MqttTopicFilterComparer, so '+'/'#' behave as a broker would. Concrete topics stay a single lookup; the wildcard scan runs only when one exists. - Retained seeds are honoured natively on MQTT 5.0 (retain handling) AND filtered client-side on the retain flag, because 3.1.1 brokers always replay. - Nothing throws on MQTTnet's dispatcher thread: a malformed payload degrades that tag (BadDecodingError / BadTypeMismatch), and a throwing OnDataChange subscriber is contained without starving the tags behind it. - The manager issues the FIRST subscribe itself — Reconnected deliberately does not fire after the initial ConnectAsync. - Reconnect re-subscribe: total failure THROWS (MqttConnection then tears the session down and retries under backoff, rather than serving a connected-but-deaf driver); a partial SUBACK rejection degrades only the denied refs, because retrying forever over one ACL-denied topic would take every tag dark. MqttConnection gains a MessageReceived observer (fired on the dispatcher, throwing observers contained) and a bounded SubscribeAsync under its own linked-CTS deadline — deliberately not MqttClientOptions.Timeout, which already governs every MQTTnet op. A refused filter is an outcome, not an exception; only the SUBSCRIBE itself failing throws. Classify() is parameterised by operation name (messages unchanged for connect). JSONPath is a documented subset ($, $.a.b, $['a'], $.a[0]); filters/slices/ recursive descent report a miss rather than pulling in a JSONPath package for an expression form that selects a set where a tag needs one scalar. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
f2a9ee5d93 |
docs(mqtt): flag the Last-Will landmine at the browse-options seam
A will is published by the BROKER, so it is invisible to PublishCountForTest. MqttDriverOptions carries none today, but Sparkplug NDEATH is a will message - once P2 adds it, an ungraceful browse disconnect would fire NDEATH under the plant's own edge-node identity. ToBrowseOptions is where it must be cleared. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
7980b34692 |
feat(mqtt): bespoke passive #-observation browser (plain)
MQTT has no discovery protocol, so the AdminUI address picker learns the topic
space the only way there is: subscribe to the wildcard and accumulate what
arrives. MqttBrowseSession serves that observation as a segment tree (split on
'/'); MqttDriverBrowser opens it with a distinct "{clientId}-browse-{guid8}"
identity under a 5-30 s clamped budget.
Browsing publishes NOTHING - the load-bearing safety property, since the picker
runs against a live plant broker. Every outbound message must route through the
single PublishAsync seam that counts into PublishCountForTest; P2's operator-
triggered Sparkplug rebirth is the one intended exception. Sparkplug mode fails
fast rather than serving a raw spBv1.0 topic tree that looks like a metric tree.
Deviation from the plan's ref list: the project also references .Driver so the
browse CONNECT reuses MqttConnection.BuildClientOptions (TLS posture, CA-pin
chain validator, credentials). A second copy would be a security divergence.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
a6370a26f8 |
fix(mqtt): make connect idempotent, bound the reconnect callback, stop State lying
Closes the connect-vs-connect defect: MQTTnet throws (and raises no DisconnectedAsync) on a connect-while-connected, so a caller's ConnectAsync and the reconnect supervisor corrupted each other in both directions. Both paths now check first; when the supervisor finds the session already restored it stands down WITHOUT firing Reconnected. Reconnected becomes Func<CancellationToken, Task>, fed from the lifetime token and capped at ConnectTimeoutSeconds, so a hung re-subscribe can no longer park the supervisor. Connected is published only after the re-subscribe succeeds; a supervisor that dies now reports Faulted instead of Reconnecting forever. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
768fd87774 |
feat(mqtt): hand-rolled reconnect loop with bounded backoff + resubscribe
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
cacc30a60d |
feat(mqtt): last-value cache backing IReadable (per-ref no-data)
LastValueCache bridges MQTT's subscribe-first push model to the OPC UA server's polled IReadable.ReadAsync: the subscription path calls Update() per RawPath, the read path calls Read(). Read never throws; an unseen RawPath returns BadWaitingForInitialData (0x80320000) rather than an exception or null, so a batch covering many references degrades per-ref instead of failing wholesale. Deviates from the plan's GoodNoData snippet: GoodNoData is reserved repo-wide for "the historian window held no samples" (NullHistorianDataSource, OtOpcUaNodeManager HistoryRead paths); BadWaitingForInitialData is the established convention for "no live value observed yet" (CalculationDriver, VirtualTagEngine, FOCAS, AddressSpaceApplier). Keyed by RawPath per the v3 driver-reference identity, not a topic/JSON-path-derived key. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
5629f3698d |
fix(mqtt): pin the CA accept path, honour presented intermediates, classify the dispose race
Review follow-ups on MqttConnection (Task 3): - Every certificate test asserted rejection, so the accept branch was unreachable-by-regression. Adds a leaf genuinely issued by the pinned CA and asserts acceptance. - ValidateAgainstPinnedCa never seeded ChainPolicy.ExtraStore from the incoming chain, so a leaf behind an intermediate delivered during the handshake failed despite a legitimate path to the pinned root. Seeds from both the incoming chain's elements and its ExtraStore; CustomRootTrust still means only the pinned roots may terminate the chain. - A DisposeAsync racing an in-flight connect escaped as an unclassified exception; it now folds into ObjectDisposedException. - Promotes the single-caller concurrency invariant into the type remarks, with the accurate blast radius (a leaked live connection, not a benign throw). Serialising the lifecycle remains Task 4's job. - X509Chain.Build can throw; an exception escaping a TLS validation callback is an opaque handshake crash, so it is caught and refused. - Adds a connect-retry test (Task 4's reconnect loop reuses the instance) and a disposed-then-connect test. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
ded4c41798 |
fix(mqtt): key tag definitions by RawPath, not the TagConfig blob
Task 2 review follow-up. The plan specified MqttEquipmentTagParser.TryParse(reference)
with def.Name = the TagConfig blob, and told us to mirror a type named
ModbusEquipmentTagParser. Both are plan defects:
- EquipmentTagRefResolver documents that a v3 driver reference "is now always a
RawPath" and that the blob-parse fallback is retired. Keying Name by the blob
would make OnDataChange publish under a reference that never matches the
RawPath-keyed fan-out in DriverHostActor - silently dead in production with
every unit test still green.
- ModbusEquipmentTagParser does not exist. The six sibling drivers all use
<Driver>TagDefinitionFactory.FromTagConfig(tagConfig, rawPath, out def).
Changes:
- Rename MqttEquipmentTagParser -> MqttTagDefinitionFactory; TryParse(reference,
out def) -> FromTagConfig(tagConfig, rawPath, out def) setting Name: rawPath,
matching ModbusTagDefinitionFactory's structure, param docs and guard order.
- Pin the identity contract with a dedicated test so a regression to blob-keying
goes red.
- Read qos with the same strictness as the enums: a present-but-invalid qos
("high" / 1.5 / 5 / null) now rejects the tag and is warned by Inspect,
instead of being silently absorbed into the driver-level default and handing
the operator a weaker delivery guarantee than the one they authored.
- No ToTagConfig inverse: the siblings carry one solely for their Driver.<X>.Cli
project, the MQTT plan defines none, and the AdminUI editor template
references no driver factory. Recorded as an explicit YAGNI call in the type
doc rather than added speculatively.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
|
||
|
|
726d6d198e |
feat(mqtt): MqttConnection connect/TLS/auth under bounded deadline
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
e3fea6e409 |
feat(mqtt): plain tag parser + strict-enum descriptor
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
a5a8710af5 |
fix(mqtt): pin security-critical defaults with tests + redact password in ToString
Code review of
|
||
|
|
f22db5d801 |
feat(mqtt): Contracts options DTO + enums (name-serialized)
Task 1 of the MQTT/Sparkplug B driver plan: MqttMode, MqttPayloadFormat, MqttProtocolVersion enums + MqttDriverOptions (broker conn + mode + nullable Sparkplug/Plain sub-options) per design doc §5.1. Removes the Task 0 temporary MQTTnet PackageReference from .Contracts (transport-free; references only Core.Abstractions). MQTTnet stays pinned in Directory.Packages.props for the .Driver project (Task 3). Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
d677ac7ad3 |
feat(mqtt): validate MQTTnet-5/net10 restore under central pinning (spike, P1 gate)
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
963eec1b32 |
chore: gitignore .worktrees/ for subagent-driven-development
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
7a9caf141c |
docs(drivers): add worktree + subagent execution commands to the tracking doc
Per-plan copy-paste command (subagent-driven-development in a git worktree), plus the slash-command and executing-plans/resume alternatives, and a note that the four independent plans can run in concurrent worktrees. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
88adec047b |
refactor(health): adopt the shared active-node check (Health 0.3.0)
ClusterPrimaryHealthCheck was written three days' worth of debugging ago because the shared ActiveNodeHealthCheck selected by RoleLeader and reported Healthy for any node lacking the role. Health 0.3.0 fixes the shared check — oldest Up member, role-preference scoping, Unhealthy when the node owns no active work — so the private copy is now duplication rather than a workaround, and it is deleted. RolePreference [admin, driver] reproduces exactly what it did: a fused central node answers for admin (where the singletons and the AdminUI are pinned), a driver-only site node answers for driver. SelectOldestUpMemberOfRole now delegates to the shared ClusterActiveNode instead of re-implementing the age ordering. This is the point of the change rather than a tidy-up: the redundancy snapshot drives the OPC UA ServiceLevel 250/240 split while the health tier drives Traefik's admin routing, so two copies of "oldest Up member of a role" would be two chances to advertise one node as authoritative while gating the data plane on another. ControlPlane takes a ZB.MOM.WW.Health.Akka reference for it; the layering trade-off is recorded at the PackageReference. Behaviour note: a node carrying neither admin nor driver now answers 503 on the active tier instead of 200. That case is reachable — RoleParser admits dev-only and cluster-role-only nodes — and 503 is correct, since such a node owns no active work and must not be in an active-tier pool. No docker-dev node is affected: every rig node is admin+driver or driver-only. Tests replaced in kind. The rule itself is now pinned in the library against a real two-node cluster; what stays OtOpcUa's own decision is the wiring, so ActiveTierRegistrationTests pins which check is on the active tag, that it is registered on driver-only nodes too, and that it is not on the ready tier. Verified: Host.Tests 25/25, ControlPlane redundancy 15/15. |
||
|
|
76e55fc112 |
docs(drivers): Wave 0-2 tracking doc + Wave 1/2 implementation plans
Add a living progress tracker (docs/plans/2026-07-24-driver-expansion-tracking.md)
for the driver-expansion program's Waves 0-2, linked from the program design doc's
document map. One row per deliverable pointing at design + implementation plan + status.
Add executable implementation plans (plan .md + co-located .tasks.json) for the four
Wave 1/2 drivers, each authored by writing-plans methodology off its design doc and
mirroring the relevant existing driver:
- Modbus RTU 11 tasks (extends Driver.Modbus; RTU-over-TCP, direct-serial descoped)
- SQL poll 22 tasks (ISqlDialect seam, SQL Server only in v1; SQLite + central-SQL fixtures)
- MTConnect Agent 23 tasks (browse free via Wave-0 seam; TrakHound-vs-hand-rolled = Task 0)
- MQTT/Sparkplug B 27 tasks (P1 plain MQTT 0-14 then P2 Sparkplug 15-26; MQTTnet-5 pinning spike = Task 0)
Wave 0 (universal browser) remains merged (
|
||
|
|
4550486144 |
fix(health): active tier reports the cluster Primary, not the admin role leader
Closes the remaining half of #494, and corrects the half fixed in |
||
|
|
8dd9da7d4d |
fix(health): bump to ZB.MOM.WW.Health 0.2.1 — Traefik leader-pinning now works
0.2.1 makes the shared ActiveNodeHealthCheck return Unhealthy (503), not
Degraded (200), for a node that carries the role but is not the role leader.
That is what the shared health spec always specified, what docs/ServiceHosting.md,
docs/v2/Architecture-v2.md and the 2026-05-26 alignment design all describe
("200 only on the Admin-role leader; 503 elsewhere"), and what
docker-dev/traefik-dynamic.yml + scripts/install/traefik-dynamic.yml assume when
they use /health/active as the load-balancer probe.
The implementation returned Degraded, which MapZbHealth maps to 200, so BOTH
central nodes always passed the probe and Traefik load-balanced the admin UI
across the pair instead of pinning it to the leader. The docs were right; the
code was wrong. Closes the Traefik half of #494.
No OtOpcUa code change — package pin only.
Live-verified on docker-dev, with the MAIN pair genuinely formed (memberCount 2,
both members naming one leader):
central-1 (leader) /health/active -> 200 traefik UP
central-2 (standby) /health/active -> 503 traefik DOWN
Failover drill (README step 2), which had never actually exercised anything:
stopping central-1 flipped Traefik to central-2 within 20 s and
http://localhost:9200/ answered 200 continuously across the swap — no outage.
Restarting central-1 reclaimed the role-leader and Traefik flipped back.
README corrected on two counts: the drill now states that exactly ONE backend is
UP by design, and step 3 no longer claims central-2 keeps leadership — RoleLeader
is the lowest-address member, so central-1 deterministically reclaims it. Added
the cold-boot caveat that starting both nodes at once can leave each self-forming
a 1-member cluster, in which case both are their own leader and both answer 200.
STILL OPEN in #494: driver-only nodes. The check is scoped to the admin role and
returns Healthy for any node that lacks it, so all four site nodes answer 200 and
the real per-Cluster redundancy Primary (oldest Up driver member, IsDriverPrimary)
is still not what this tier reports.
|
||
|
|
9ed2251f37 |
chore(health): bump ZB.MOM.WW.Health* to 0.2.0
Picks up the additive per-entry `data` object in the canonical health JSON and the cluster-view data the shared AkkaClusterHealthCheck now publishes (leader / selfAddress / selfRoles / memberCount / unreachableCount). No OtOpcUa code change is needed: the "akka" ready/active check is already the shared ZB.MOM.WW.Health.Akka.AkkaClusterHealthCheck (Health/HealthEndpoints.cs), so /health/ready starts carrying entries["akka"].data.leader on the bump alone. Payloads from every other check are byte-identical — data is emitted only when a check publishes some. Prereq for the family overview dashboard (scadaproj docs/plans/2026-07-22-overview-dashboard-impl-plan.md, Phase 2). Verified: Host 21/21, AdminUI 692/692, Host builds 0 warnings. Live rig check (admin node serves data.leader) deferred to the dashboard's acceptance run. |
||
|
|
d1dac87f6f |
docs: record the bootstrap guard (CLAUDE.md + Phase 7 note)
The simultaneous-cold-start split-brain, carried as a Phase 7 "candidate follow-up,
not shipped", is now closed by the opt-in Cluster:BootstrapGuard (
|
||
|
|
279d1d0fb1 |
feat(cluster): simultaneous-cold-start split-brain guard (opt-in)
Closes the residual Phase-7 gap: when BOTH nodes of a 2-node pair cold-start at the same instant, each self-first seed runs FirstSeedNodeProcess, times out waiting for the other, and forms its own 1-node cluster — two Primaries in one pair (the Phase 6 live gate reproduced this reliably on docker). The prior mitigation was operational (staggered start / compose depends_on), which does not exist on production hardware. The guard (dark switch Cluster:BootstrapGuard:Enabled, default OFF) prevents the split without giving up cold-start-alone: - The lower-address node is the preferred founder: self-first, forms immediately. - The higher node probes its partner's Akka port (TCP connect) up to PartnerProbeSeconds: reachable => peer-first (join the founder, never race it); unreachable => self-first (partner is dead, form alone). The order is decided BEFORE the single JoinSeedNodes, from an explicit reachability signal — never re-formed mid-handshake (the retired SelfFormAfter failure mode). When on, Akka gets no config seeds (BuildClusterOptions) and ClusterBootstrapCoordinator drives the join. Review-driven hardening: case-insensitive tie-break (a hostname-casing mismatch would reopen the split); fail-fast validation of the timing knobs; the residual "founder dies in the probe->join window" hang is documented and made operator-visible (warning + a restart recovers). Real-ActorSystem coordinator tests cover the load-bearing higher-node cold-start-alone case; 147/147 Cluster tests pass. Live-gated on docker-dev (site-a = enablement demo, serialization removed, guard on; site-b keeps depends_on-serialization + guard off as the A/B control): simultaneous start -> site-a-1 founds, site-a-2 probes-reachable-joins -> 250/240, NO split; higher-node-alone -> probes dead founder, self-forms -> 250; founder rejoin -> 240. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |
||
|
|
c50ebcf7dc |
docs(mesh): Phase 7 DONE — failover drills PASSED, program complete
All 6 failover drills pass on the docker-dev three-mesh rig: - D1 MAIN graceful manual-failover via the AdminUI "Trigger failover" button — Primary leaves, peer→250, restarts and rejoins with no split. - D2 SITE-A auto-down failover, both directions; D3 SITE-B crash-the-oldest. - D4 auto-down 1-vs-1 survival — every survivor stayed Up (closes the gate deferred since Phase 0a). - D5 self-first cold-start-alone — a node boots with its partner down and forms its own mesh (gate b). - D6 recovery — restarted nodes rejoin as Secondary (240). Ships a one-page co-located operator runbook (OtOpcUa + ScadaBridge on shared site VMs). Manual-failover button confirmed MAIN-only by construction; site pairs (driver-only, no UI) fail over via auto-down. Carried Phase-6 finding — simultaneous cold-start of both pair VMs can split — mitigated operationally in the runbook (staggered start); a product-level join guard is a candidate follow-up, not shipped. The per-cluster mesh program (Phases 0a–7) is COMPLETE. Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW |