Compare commits
19 Commits
9f91e84d83
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
| b3dc17a5fe | |||
| 1c9159a7ce | |||
| 5d4853a1f0 | |||
| c1216044a4 | |||
| 248676ed16 | |||
| 47850c0f53 | |||
| bb138d5254 | |||
| 9a12826172 | |||
| 86d129de6c | |||
| dbec3ee263 | |||
| a1abbff760 | |||
| 266f001a2e | |||
| e04c2617cc | |||
| 97afa84fcd | |||
| 0eb44314cb | |||
| a5256e9b12 | |||
| e0f105c3b3 | |||
| 8524a7f746 | |||
| c4dcd9bc02 |
@@ -220,6 +220,7 @@ Other peers in the `scadaproj` family (see `scadaproj/CLAUDE.md` for details): `
|
||||
- Automatic dual-node recovery from persistent storage.
|
||||
- **Active/standby is decided by `ActiveNodeEvaluator.SelfIsOldestUp`, never by cluster leadership** — see the Architecture note above. `/health/active` is **central-only** (site nodes map no `/health/*` at all) and backs both Traefik's active-node routing and `IActiveNodeGate`, so the proxy and the Inbound API always agree on which node is active. Central never needs to know which *site* node is active: ClusterClient contact rotation reaches either receptionist and the site-internal `ClusterSingletonProxy` lands the work on the active node for free. The **exception is gRPC**, which picks `GrpcNodeAAddress`/`GrpcNodeBAddress` explicitly and flips on error.
|
||||
- **Seed-node ordering: every node lists ITSELF first (decision 2026-07-22) — the boot-alone gap is CLOSED.** Only `seed-nodes[0]` may self-join to form a new cluster (Akka runs `FirstSeedNodeProcess` for it, `JoinSeedNodeProcess` — which can never form one — for everyone else). All 14 shipped node appsettings now lead with the node's own address, so any node can cold-start alone and become operational unattended (~5s, `seed-node-timeout`); `StartupValidator` fails the boot if the ordering is broken (compares host AND port; Akka does no DNS canonicalisation). Two nodes cold-starting together while mutually reachable converge on ONE cluster via the `InitJoin` handshake — they split only under a genuine boot-time partition, the same class auto-down accepts. **An external self-form timer (`Cluster.Join(SelfAddress)` after a window) was implemented and REJECTED:** it sits outside the join handshake, so on a routine standby restart — where the peer is alive but the join is stalled behind removal of the node's own stale incarnation — it fires mid-join and permanently splits the pair (measured: still split after 90s). Regression tests: `SelfFirstSeedBootstrapTests`. The keep-oldest active-crash total outage was separately closed by the auto-down decision. See `docs/requirements/Component-ClusterInfrastructure.md` → Seed Node Ordering.
|
||||
- **Simultaneous-cold-start split-brain guard (opt-in, Gitea #33) — the residual self-first cost, closed.** Self-first-on-both means a *truly simultaneous* cold start (shared power/hypervisor event) races `FirstSeedNodeProcess` on BOTH nodes → two 1-node clusters that never merge (in-process loopback tests converge and hide it; real parallel VM starts don't — OtOpcUa reproduced it live). Ported from OtOpcUa (`lmxopcua` `d1dac87f`): a **dark switch `ScadaBridge:Cluster:BootstrapGuard:Enabled` (default OFF, guard-off byte-identical)**. On: `BuildHocon` emits an EMPTY seed list (Akka does not auto-join) and `ClusterBootstrapCoordinator` (`IHostedService`, registered in BOTH the Central and Site composition roots) issues ONE reachability-gated `JoinSeedNodes` — the pure `ClusterBootstrapGuard` core makes the lower canonical `host:port` the **founder** (self-first, forms immediately), the higher node TCP-probes the founder up to `PartnerProbeSeconds` (25s) and joins **peer-first** if reachable else **self-first** (cold-start-alone preserved). Decided BEFORE a single join, **never re-forms mid-handshake** (that was the rejected self-form-timer's flaw); case-insensitive founder tie-break; probe timings validated `>0` when enabled. Residual accepted trade: founder dies in the probe→join window → higher node hangs unjoined, coordinator WARNs, a restart recovers it. Tests: `ClusterBootstrapGuardTests` (pure) + `ClusterBootstrapCoordinatorTests` (real-ActorSystem, incl. both-cold-start-together-form-one-cluster). See `docs/requirements/Component-ClusterInfrastructure.md` → Simultaneous cold start.
|
||||
|
||||
### UI & Monitoring
|
||||
- Central UI: Blazor Server (ASP.NET Core + SignalR) with Bootstrap CSS. No third-party component frameworks (no Blazorise, MudBlazor, Radzen, etc.). Build custom Blazor components for tables, grids, forms, etc.
|
||||
|
||||
@@ -90,9 +90,9 @@
|
||||
to mark tests as Skipped (not silently Passed) when MSSQL is unreachable.
|
||||
-->
|
||||
<PackageVersion Include="Xunit.SkippableFact" Version="1.5.61" />
|
||||
<PackageVersion Include="ZB.MOM.WW.Health" Version="0.1.0" />
|
||||
<PackageVersion Include="ZB.MOM.WW.Health.Akka" Version="0.1.0" />
|
||||
<PackageVersion Include="ZB.MOM.WW.Health.EntityFrameworkCore" Version="0.1.0" />
|
||||
<PackageVersion Include="ZB.MOM.WW.Health" Version="0.3.0" />
|
||||
<PackageVersion Include="ZB.MOM.WW.Health.Akka" Version="0.3.0" />
|
||||
<PackageVersion Include="ZB.MOM.WW.Health.EntityFrameworkCore" Version="0.3.0" />
|
||||
<PackageVersion Include="ZB.MOM.WW.Telemetry" Version="0.1.0" />
|
||||
<PackageVersion Include="ZB.MOM.WW.Telemetry.Serilog" Version="0.1.0" />
|
||||
<PackageVersion Include="ZB.MOM.WW.MxGateway.Client" Version="0.1.1" />
|
||||
|
||||
@@ -215,7 +215,7 @@ Central callers interact through `CommunicationService`, which wraps each comman
|
||||
| Instance deployment (notify-and-fetch) | `RefreshDeploymentAsync` | 120 s |
|
||||
| Instance lifecycle | `DisableInstanceAsync`, `EnableInstanceAsync`, `DeleteInstanceAsync` | 30 s |
|
||||
| Artifact deployment | `DeployArtifactsAsync` | 60 s |
|
||||
| Integration routing | `RouteIntegrationCallAsync` | 30 s |
|
||||
| Integration routing (Inbound API routed-site-script) | `RouteToCallAsync`, `RouteToGetAttributesAsync`, `RouteToSetAttributesAsync`, `RouteToWaitForAttributeAsync` | 30 s (`IntegrationTimeout`) |
|
||||
| Debug snapshot | `RequestDebugSnapshotAsync` | 30 s |
|
||||
| Remote queries | `QueryEventLogsAsync`, `QueryParkedMessagesAsync`, etc. | 30 s |
|
||||
| OPC UA tag browse | `BrowseNodeAsync` | 30 s |
|
||||
|
||||
@@ -1,9 +1,17 @@
|
||||
# Integration call routing (`IntegrationCallRequest`) is dead on both ends
|
||||
|
||||
**Date:** 2026-07-22 · **Status:** OPEN (decision needed: wire or delete) · **Tracked:** Gitea
|
||||
**Date:** 2026-07-22 · **Status:** RESOLVED (DELETED 2026-07-23) · **Tracked:** Gitea
|
||||
[#32](https://gitea.dohertylan.com/dohertj2/ScadaBridge/issues/32) (filed 2026-07-23) · **Severity:**
|
||||
Low (no runtime impact — the path cannot be reached) · **Area:** Central–Site Communication
|
||||
|
||||
> **Resolution (2026-07-23):** Decision = **delete** (option 1 below). Removed the
|
||||
> `IntegrationCallRequest`/`IntegrationCallResponse` messages, `CommunicationService.RouteIntegrationCallAsync`,
|
||||
> the `SiteCommunicationActor` receive block + `_integrationHandler` field + `LocalHandlerType.Integration`,
|
||||
> and the four tests that covered them. `IntegrationTimeout` was **kept** — it is the live timeout for the
|
||||
> Inbound API's `RouteTo*` verbs, which are the actual implementation of this "External → Central → Site →
|
||||
> Central" pattern (design §4). The narrative comments that documented the exclusion (proto/mapper/dispatcher)
|
||||
> were updated. Full solution build clean; Communication suite 634 green. This note is retained as the record.
|
||||
|
||||
## What
|
||||
|
||||
"Pattern 4: Integration Routing" — `CommunicationService.RouteIntegrationCallAsync` →
|
||||
|
||||
@@ -475,5 +475,68 @@ default (Phase 4 flipped the shipped appsettings to `CentralGrpcEndpoints` — u
|
||||
redeploy from `main` now yields all-gRPC, not all-Akka). No ClusterClient remains anywhere in the
|
||||
running system or the source tree. Deferred, unchanged from earlier phases: the instance-dependent
|
||||
central→site RPC matrix (`ExecuteOpcUa`/`ExecuteRoute`/standby parked retry) and live
|
||||
`TriggerFailover`, all unit-proven; PSK **rotation** on a live pair; `docker-env2` was updated with
|
||||
its own key but not gated here.
|
||||
`TriggerFailover`, all unit-proven; PSK **rotation** on a live pair. `docker-env2` is now gated —
|
||||
see below.
|
||||
|
||||
---
|
||||
|
||||
## Env2 gate — `docker-env2` on the gRPC PSK build — **PASS** (2026-07-23, Gitea #31)
|
||||
|
||||
The secondary Transport-testing topology (`docker-env2/`: 2 central + 1 site-x × 2 nodes, host
|
||||
ports 91XX, shared infra) carried its `GrpcPsk` in config since Phase 0 but had never been
|
||||
redeployed or gated against the migration build. Rebuilt `scadabridge:latest` from `main` @
|
||||
`8524a7f7` (includes the #28 Transport execution-strategy fix) and recreated only the env2
|
||||
containers (`bash docker-env2/deploy.sh`; the primary rig's running image was untouched). Site-x
|
||||
PSK `dev-grpc-psk-docker-env2-site-x` on both site nodes matches central's
|
||||
`ScadaBridge__Communication__SitePsks__site-x`.
|
||||
|
||||
| # | Check | Result |
|
||||
|---|-------|--------|
|
||||
| 1 | Both site nodes boot with the key (StartupValidator passes) | **PASS** |
|
||||
| 2 | Unauthenticated control-plane call ⇒ PermissionDenied; correct key ⇒ success | **PASS** |
|
||||
| 3 | LocalDb unaffected | **PASS** |
|
||||
|
||||
### Check 1 — both site nodes boot
|
||||
|
||||
`scadabridge-env2-site-x-a` and `-b` both reached `Application started` with `Now listening on:
|
||||
http://[::]:8083` (the gRPC control/data listener). `StartupValidator` is fail-closed on a missing
|
||||
`ScadaBridge:Communication:GrpcPsk` — a site node without it refuses to boot — so reaching
|
||||
`Application started` is itself the proof the key is present and valid. Each node also logged a clean
|
||||
`Reconcile pass for site site-x node node-{a,b} complete: 0 fetched, 0 failed, 0 orphan(s)`, i.e.
|
||||
the site→central gRPC path was already carrying traffic.
|
||||
|
||||
### Check 2 — PSK auth on the control plane
|
||||
|
||||
`grpcurl` against the site-hosted `SiteStreamService/PullAuditEvents` on both nodes
|
||||
(`localhost:9123` = site-x-a, `localhost:9124` = site-x-b):
|
||||
|
||||
- **No auth header** ⇒ `PermissionDenied: Control plane authentication failed.`
|
||||
- **Wrong bearer key** (`Bearer WRONG-KEY`, `x-scadabridge-site: site-x`) ⇒ `PermissionDenied`.
|
||||
- **Correct key** (`Bearer dev-grpc-psk-docker-env2-site-x`) ⇒ `{}` (empty audit set — the env has
|
||||
no S&F driver seeded — but **not** a permission error: auth accepted). Identical on both nodes.
|
||||
|
||||
(Same contract clarification as the primary rig: the site interceptor authenticates the node's single
|
||||
`GrpcPsk` and does not require `x-scadabridge-site`; per-site isolation is proven by the wrong-key
|
||||
rejection.)
|
||||
|
||||
### Check 3 — LocalDb unaffected
|
||||
|
||||
Env2 site-x runs LocalDb **local-only** (no `LocalDb:Replication:PeerAddress` configured — the rig's
|
||||
replication posture is proven on the primary `docker/` stack, not here), so "unaffected" means the
|
||||
consolidated site DB stays operational on the gRPC build. Site logs hold **0** LocalDb
|
||||
errors/exceptions and both nodes boot healthy with the LocalDb at `/app/data/site-localdb.db`; the
|
||||
clean reconcile passes above confirm the site DB + site→central path work end to end.
|
||||
|
||||
### Bonus / observations
|
||||
|
||||
- **No real ClusterClient.** Central logs hold **0** ClusterClient/receptionist references; the site
|
||||
logs' only match is the benign string label `client=ClusterClientSiteAuditClient` in a
|
||||
`SiteAuditTelemetryActor created` line (a legacy identifier, **not** an Akka `ClusterClient`) —
|
||||
identical to the primary rig's site nodes. No real receptionist/ClusterClient actor path exists.
|
||||
- **Central registers site-x over gRPC.** `Site site-x registered online via heartbeat` — the
|
||||
site→central gRPC command/control path is live.
|
||||
- **Seed-data note (not a migration issue).** `ScadaBridgeConfig2.dbo.Sites` has **0 rows**, so
|
||||
central logs `Reconcile request from unknown site 'site-x' … replying with empty gap` and evicts
|
||||
it from the health aggregator as "no longer configured." This is a first-time-setup seeding gap
|
||||
(`docker-env2/seed-sites.sh` not yet run on the fresh DB), orthogonal to the transport — which
|
||||
demonstrably carries the heartbeat and reconcile regardless.
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
# OtOpcUa v3 native-alarm B/C live-gate (Gitea #14) — PASS
|
||||
|
||||
**Date:** 2026-07-23 · **Scope:** items **B** (native-alarm dedup) + **C** (UNS-bound alarm routing) of #14,
|
||||
against a UNS structure + scripted Part-9 alarm authored on the live `otopcua-dev` v3 cluster (see
|
||||
`2026-07-23-otopcua-v3-raw-path-live-gate.md` for the connection + `SiteAOnly` substrate). Instrument:
|
||||
ScadaBridge itself (its DCL A&C client is the thing under test).
|
||||
|
||||
## Headline
|
||||
|
||||
**Both B and C PASS, and item C's original premise turned out to be already-solved code.** No ScadaBridge
|
||||
code change is needed for B/C. The scope-doc worry — "a UNS-bound source won't `StartsWith`-match a
|
||||
condition whose `SourceName` is the RawPath/ScriptedAlarmId" — is **obsolete**: ScadaBridge routes native
|
||||
alarms by the **subscribed node reference**, not the event `SourceName` (deliberate fix, Gitea #17). The
|
||||
only remaining #14 work is the item-A data re-author (greenfield, operational).
|
||||
|
||||
## Substrate (on OtOpcUa, SITE-A cluster)
|
||||
|
||||
- UNS: `area-1/line-1/pump01` → signal `SiteAOnly` (equipment folder node `ns=3;s=EQ-54f7711e160d`,
|
||||
which has **`EventNotifier=1`**; signal `ns=3;s=area-1/line-1/pump01/SiteAOnly` mirrors the raw value).
|
||||
- Scripted Part-9 condition `ns=3;s=SA-siteahigh` (a `HasComponent` child of the equipment folder),
|
||||
`AlarmConditionType`, severity 700, predicate driven by `SiteAOnly` (set always-active for the gate).
|
||||
- `Server` (i=2253) `EventNotifier=1` (the connection-wide aggregate).
|
||||
|
||||
## What ScadaBridge did (bindings on template 2148, instance 98, connection `OtOpcUa-v3-raw`)
|
||||
|
||||
Two `TemplateNativeAlarmSource` bindings, deliberately overlapping to stress dedup:
|
||||
- `ServerAlarms` → `--source-ref i=2253` (Server aggregate)
|
||||
- `UnsEquipAlarms` → `--source-ref nsu=https://zb.com/otopcua/uns;s=EQ-54f7711e160d` (UNS equipment node)
|
||||
|
||||
## Findings / checks
|
||||
|
||||
| # | Check | Result |
|
||||
|---|-------|--------|
|
||||
| C1 | A **UNS-node-scoped** binding receives + routes the native alarm | **PASS** — `UnsEquipAlarms` mirrored `SA-siteahigh.SiteA High` active=True, sev 700, `kind=NativeOpcUa` |
|
||||
| C2 | The condition's `SourceName` = **ScriptedAlarmId** (`SA-siteahigh`), NOT RawPath | **Confirmed** (captured live) — and it does **not** break routing (routing keys off the subscribed node, Gitea #17), only feeds the alarm's identity (`alarmName = SA-siteahigh.SiteA High`) |
|
||||
| C3 | Item C's original "SourceName mismatch drops UNS-bound transitions" premise | **Obsolete** — `OpcUaAlarmMapper.BuildIdentity` stamps `SourceObjectReference` = the subscription ref; `DataConnectionActor` routes on that, not `SourceName` |
|
||||
| B1 | Fanned condition (Server aggregate + equipment-folder notifier) mirrored **once**, not duplicated | **PASS** — active condition appears exactly once (on the more-specific equipment feed); the Server binding stays an inactive placeholder. Stable across repeated snapshots |
|
||||
| B2 | Mechanism | ScadaBridge opens **one** OPC UA `Subscription` per connection with one `MonitoredItem` per source-ref; `ConditionRefresh` delivers each retained condition **once per subscription**, to the most-specific monitored notifier (the equipment folder) — so overlapping bindings do not double-mirror the fan-out |
|
||||
| B3 | Residual double-mirror risk | Latent, **not triggered** here: it needs two bindings whose `SourceReference` strings are literal prefixes of one another AND a delivery matching both (`DataConnectionActor` has no cross-binding, per-ConditionId dedup — `InstanceActor._latestAlarmEvents` is last-writer-wins by `alarmName`). Disjoint node refs (as here) don't collide |
|
||||
|
||||
## Captured real payload (what §5 phase 3 wanted)
|
||||
|
||||
`kind=NativeOpcUa`, `alarmName='SA-siteahigh.SiteA High'` (`SourceName`=`SA-siteahigh` = ScriptedAlarmId,
|
||||
`.` + `ConditionName`=`SiteA High`), `alarmTypeName='AlarmConditionType'`, `severity=700`,
|
||||
`condition.active=True`. A&C state is delivered via events/`ConditionRefresh` (ScadaBridge triggers a
|
||||
refresh on every subscribe), not a plain attribute read.
|
||||
|
||||
## Bottom line for #14
|
||||
|
||||
- **A** — raw/UNS binding re-author: validated end-to-end (raw-path gate proved `nsu=` bind→read; this gate
|
||||
proved UNS bind→alarm). Remaining work is data re-authoring (greenfield), **no code**.
|
||||
- **B** — dedup: **PASS**, no double-mirror across the v3 fan-out (one subscription + `ConditionRefresh`).
|
||||
- **C** — UNS-bound routing: **PASS**; the original concern was already solved by Gitea #17.
|
||||
- **D/E/F** — shipped in phase 2 (`2026-07-23-otopcua-v3-nsu-hardening-and-browse-ux.md`).
|
||||
|
||||
**#14 needs no further ScadaBridge code for B/C.** The only shipped code from this whole effort was the
|
||||
phase-2 `nsu=` hardening + the `RealOpcUaClient` host-rewrite fix (surfaced by the raw gate). #14 can close
|
||||
once item A's re-author is done in the target environment(s).
|
||||
|
||||
## Rig artifacts (removable on next `docker/deploy.sh`)
|
||||
|
||||
OtOpcUa ConfigDb: `UnsArea SITEA-area1` / `UnsLine SITEA-line1` / `Equipment EQ-54f7711e160d` /
|
||||
`UnsTagReference UTR-siteapump01-siteaonly` / `Script SC-siteahigh` / `ScriptedAlarm SA-siteahigh`.
|
||||
ScadaBridge: template 2148 native-alarm-sources `ServerAlarms` + `UnsEquipAlarms`; instance 98; connection 3044.
|
||||
@@ -0,0 +1,140 @@
|
||||
# OtOpcUa v3.0 dual-namespace cutover — scoping (Gitea #14)
|
||||
|
||||
**Date:** 2026-07-23 · **Status:** SCOPING (decisions needed — see §6) · **Tracked:** Gitea
|
||||
[#14](https://gitea.dohertylan.com/dohertj2/ScadaBridge/issues/14) · **Area:** Data Connection
|
||||
Layer (OPC UA adapter) · **Upstream:** OtOpcUa v3.0 (merged `master` 2026-07-16, PR #472, merge
|
||||
`ec6598ce`); design `~/Desktop/OtOpcUa/docs/plans/2026-07-15-raw-uns-two-subtree-v3-design.md`.
|
||||
|
||||
## 1. Headline
|
||||
|
||||
**ScadaBridge is already structurally v3-safe.** A prior refactor made the OPC UA reference a
|
||||
**namespace-URI-durable** string resolved dynamically against the live server `NamespaceArray`, so
|
||||
the retirement of OtOpcUa's `EquipmentNodeIds` scheme and the split into two namespaces does **not**
|
||||
break any parsing or binding logic. There is **no hardcoded `otopcua` namespace URI anywhere in
|
||||
`src/`** (only test fixtures), and **nothing in ScadaBridge parses the `{EquipmentId}/…` NodeId
|
||||
shape** — the identifier is opaque.
|
||||
|
||||
The genuine cutover is therefore **narrow**: one **data** problem (a namespace-index collision on
|
||||
legacy bindings), one real **code gap** (native-alarm dedup across the raw+uns notifier fan-out),
|
||||
and some **UX / validation** polish. Most of it can only be *closed* against a re-seeded live v3
|
||||
rig, so this is as much a live-gate as a code change.
|
||||
|
||||
## 2. The v3 wire contract (what changed upstream)
|
||||
|
||||
OtOpcUa replaced its single custom namespace `https://zb.com/otopcua/ns` with **two**:
|
||||
|
||||
| Subtree | Namespace URI | NodeId form | Shape |
|
||||
|---|---|---|---|
|
||||
| **Raw** (device tree, source of truth) | `https://zb.com/otopcua/raw` | `ns=<raw>;s=<RawPath>` | `Folder/…/Driver/Device/TagGroup/…/Tag` |
|
||||
| **UNS** (equipment projection) | `https://zb.com/otopcua/uns` | `ns=<uns>;s=<Area>/<Line>/<Equipment>/<EffectiveName>` | Area→Line→Equipment→signal |
|
||||
|
||||
Semantics ScadaBridge can rely on (from the v3 design doc):
|
||||
- Every value has **one source** (the raw tag), fanned to both the raw NodeId and each referencing
|
||||
UNS NodeId with **identical value/quality/timestamp**. Each UNS variable `Organizes`-references
|
||||
its raw node (cross-tree link is browseable).
|
||||
- **Writes** route through **either** NodeId (same `WriteOperate` gating).
|
||||
- **HistoryRead** works via **both** NodeIds under one historian tagname.
|
||||
- **Native Part 9 alarms:** one condition instance at the raw tag, `ConditionId`/primary identity =
|
||||
**RawPath**, fanned by a **single `ReportEvent`** to the raw device folder **and** every
|
||||
referencing equipment folder. A Server-object subscriber gets **exactly one** copy (the SDK
|
||||
dedups on the shared `InstanceStateSnapshot`); folder-scoped subscribers in each namespace see the
|
||||
condition's events.
|
||||
- The old `EquipmentNodeIds` (`{equipmentId}/{folderPath}/{name}`) scheme is **retired**.
|
||||
|
||||
OtOpcUa's own "Cross-repo impact" note sized this for us: **every ScadaBridge binding re-binds**;
|
||||
it's a *data* migration, not a code rewrite; the fragile part is that stored refs hard-code the
|
||||
**namespace index** and ScadaBridge stores no URI; the recommended remediation is to **move to
|
||||
`nsu=`-qualified references** while re-binding.
|
||||
|
||||
## 3. Current-state assessment (what is already v3-safe)
|
||||
|
||||
Grounded in the DCL code (`src/ZB.MOM.WW.ScadaBridge.DataConnectionLayer/`):
|
||||
|
||||
- **`Adapters/OpcUaNodeReference.cs` — the single translation seam, already URI-durable.**
|
||||
`Resolve` maps `nsu=<uri>` to the live index via `ExpandedNodeId.ToNodeId(expanded,
|
||||
namespaceUris)` against the session's `NamespaceTable`, and **throws** when a URI is absent from
|
||||
the server's `NamespaceArray` (no silent stale-index binding); rejects `svr=` cross-server refs.
|
||||
`ToDurable` emits the `nsu=<uri>` form by reading `namespaceUris.GetString(NamespaceIndex)`.
|
||||
**No change needed — this seam is the reason the cutover is low-risk.**
|
||||
- **`Adapters/RealOpcUaClient.cs`** — every read/write/subscribe/browse routes NodeIds through the
|
||||
seam with the live `_session.NamespaceUris`. No hardcoded URI, no assumed index.
|
||||
- **`Adapters/OpcUaDataConnection.cs`** — passes the configured tag-path string straight to the
|
||||
client; no path parsing/joining. There is **no browse-path→NodeId translation** anywhere
|
||||
(bindings are absolute NodeIds), so there is no indirection layer to update.
|
||||
- **`Commons/Types/DataConnections/OpcUaEndpointConfig.cs`** — carries only `EndpointUrl` + timing/
|
||||
auth/heartbeat; **no namespace URI, index, or NodeId field**. Namespaces are discovered from the
|
||||
live server, not configured — inherently v3-tolerant. Nothing to change here.
|
||||
- **`Adapters/OpcUaAlarmMapper.cs` `BuildIdentity`** — the per-condition key is already
|
||||
`SourceName (= RawPath post-v3) + "." + ConditionName`, i.e. **already aligned** with v3 keying on
|
||||
RawPath; the routing identity is the binding string verbatim (namespace-form-independent).
|
||||
|
||||
## 4. The genuine cutover work
|
||||
|
||||
| # | Item | Kind | Primary files | Confidence |
|
||||
|---|------|------|---------------|------------|
|
||||
| A | **Legacy `ns=<index>` bindings collide with v3.** v2's sole custom namespace and v3's `raw` both land at `ns=2`, so a stored `ns=2;s=…` resolves **without error** but now means a raw-tree node. Every stored binding must be re-authored (or migrated) to the new address space, ideally in `nsu=` form. | **Data** | `TemplateAttribute.DataSourceReference`, `InstanceConnectionBinding.DataSourceReferenceOverride` (config DB); heartbeat `TagPath`; Transport bundles | High (inspection) |
|
||||
| B | **Native-alarm dedup across the raw+uns fan-out.** `RealOpcUaClient.HandleAlarmEvent` keys on RawPath+ConditionName with **no `ConditionId`-based dedup**; it assumes one condition arrives on one feed. If a v3 server delivers the same condition through two notifier paths into the same feed, two `EventFieldList`s each call `onTransition`. Must verify the Server-object aggregate feed collapses to one copy (the v3 doc says the SDK dedups Server-object subscribers, so this *should* hold — but it is the biggest genuine gap and needs a live check; add ConditionId-based dedup only if the live server double-delivers). | **Code + live validation** | `Adapters/RealOpcUaClient.cs` (HandleAlarmEvent), `Adapters/OpcUaAlarmMapper.cs` | Medium (needs live v3) |
|
||||
| C | **Alarm routing for UNS-bound sources (confirmed by D-1 = ingest both).** `DataConnectionActor` routes transitions by `SourceReference.StartsWith(bindingRef)`, but a v3 condition's `SourceName` is the **RawPath** regardless of subtree — so a **UNS-bound** alarm source won't `StartsWith`-match. Needs a routing change: resolve a UNS-bound source to its backing RawPath (via the browseable UNS→Raw `Organizes` reference) or map identities. Confirmed scope; validate the exact `SourceName`/`SourceNode` payload on the live rig before finalizing the mapping. | **Code + live validation** | `Actors/DataConnectionActor.cs` (alarm routing), `Adapters/RealOpcUaClient.cs` | Medium (needs live v3) |
|
||||
| D | **Browse / search UX now shows two subtrees.** `BrowseChildrenAsync` / `AddressSpaceSearch.SearchAsync` start at the standard Objects root and will enumerate **both** raw and uns as sibling roots. Functional, not breaking — but the picker shows two roots, the manual-entry placeholder still nudges `ns=2;s=…`, and the deeper raw hierarchy stresses `VisitedNodeCeiling`/`maxDepth`. | **UX** | `CentralUI/Components/Dialogs/NodeBrowserDialog.razor`, `IOpcUaClient` search caps | High (inspection) |
|
||||
| E | **`nsu=` hardening (optional, recommended).** The picker already emits `nsu=` via `ToDurable`, but the seam still **accepts** bare `ns=<index>`. Decide whether to warn/reject bare `ns=` on save (closes A permanently) or leave it permissive. | **Code** | `OpcUaNodeReference`, binding validation, picker | High (inspection) |
|
||||
| F | **Doc drift.** `ns=` examples in code comments + `Component-DataConnectionLayer.md`; update to `nsu=`. Update the scadaproj umbrella index (cutover step 4). | **Docs** | doc comments, `docs/requirements/Component-DataConnectionLayer.md`, `../scadaproj/CLAUDE.md` | High |
|
||||
|
||||
## 5. Proposed phasing
|
||||
|
||||
1. **Decisions** (§6) — resolve D-1..D-3 before code.
|
||||
2. **`nsu=` hardening + UX** (items D, E, F) — picker emits/enforces `nsu=`, placeholder + search
|
||||
caps updated, docs swept. Pure ScadaBridge-side, unit-testable, no live server required.
|
||||
3. **Re-seed a v3 rig** — bring up an OtOpcUa v3.0 server (docker) and re-author a representative
|
||||
set of bindings across **both** subtrees via the picker (D-1); confirm subscribe/read/write
|
||||
round-trip and HistoryRead via each subtree, and capture the exact native-alarm `SourceName`/
|
||||
`SourceNode` payload so item C's UNS→Raw routing mapping is built against real data.
|
||||
4. **Alarm live-gate** (items B, C) — against the v3 rig, drive a native condition that fans to both
|
||||
raw + equipment notifiers and confirm ScadaBridge sees **exactly one** transition per state
|
||||
change, correctly routed to the bound source. Add ConditionId-based dedup / routing fix **only if
|
||||
the live server double-delivers**.
|
||||
5. **Data migration decision** (item A) — either force re-authoring (greenfield, no automatic map)
|
||||
or write a one-shot migration if a deterministic old→new mapping exists for the rig's data.
|
||||
6. **Umbrella index + issue close.**
|
||||
|
||||
## 6. Decisions needed
|
||||
|
||||
- **D-1 — Which subtree does ScadaBridge ingest? → DECIDED (2026-07-23): BOTH.** The picker
|
||||
surfaces both the Raw (`…/raw`, device tree) and UNS (`…/uns`, Area→Line→Equipment) subtrees, and
|
||||
an operator binds each attribute / native-alarm source against whichever fits — UNS for
|
||||
equipment-modelled signals, Raw where device-level identity is wanted. For **values** this is
|
||||
free: a given ScadaBridge attribute binds to one NodeId, and v3 guarantees identical
|
||||
value/quality/timestamp on both trees, so no per-value dedup arises.
|
||||
**Consequence for alarms (promotes item C from "validate" to "fix"):** a native alarm condition is
|
||||
keyed by its **RawPath** and its `SourceName` is the RawPath *regardless of which subtree the
|
||||
source was bound on*. So a **UNS-bound** alarm source (prefix `nsu=…/uns;s=<Area>/<Line>/…`) will
|
||||
**not** `StartsWith`-match a condition whose `SourceName` is the raw path. Ingesting both therefore
|
||||
**requires** a routing change so a UNS-bound source resolves to the RawPath of its backing raw node
|
||||
(the UNS→Raw `Organizes` reference is browseable and is the link to follow), or an equivalent
|
||||
identity mapping. This is confirmed scope, gated on the live v3 rig (§5 phase 4).
|
||||
- **D-2 — Enforce `nsu=` on save?** *Recommendation: yes* — warn (or reject) bare `ns=<index>`
|
||||
bindings at authoring time so the index-collision class (item A) can never recur. Low effort, high
|
||||
durability.
|
||||
- **D-3 — Migrate legacy bindings, or re-author?** Upstream is greenfield (no reliable automatic
|
||||
old→new mapping), so *re-authoring via the picker is the default*. A migration is only worth
|
||||
writing if this environment has a deterministic mapping worth automating.
|
||||
|
||||
## 7. Risks / what needs a live v3 server
|
||||
|
||||
- Items **B** and **C** (alarm dedup + routing) **cannot be fully closed by code inspection** — they
|
||||
depend on how a live v3 server actually delivers a fanned condition to an aggregate vs
|
||||
folder-scoped subscription. The plan gates them behind a re-seeded v3 rig.
|
||||
- No production config data is preserved upstream (greenfield), so there is **no migration
|
||||
correctness risk** to legacy production rows — only the dev/test rigs need re-authoring.
|
||||
- The `OpcUaEndpointConfig` has no namespace field, so **no schema/EF change** is anticipated; this
|
||||
cutover is expected to add **no EF migration** (to be confirmed once D-2's validation surface is
|
||||
decided).
|
||||
|
||||
## 8. Bottom line
|
||||
|
||||
The prior namespace-durability refactor did the heavy lifting: ScadaBridge already resolves
|
||||
`nsu=`-qualified references against the live `NamespaceArray` and never assumes an index or a NodeId
|
||||
shape. What remains is (1) getting existing bindings onto the new address space (data / re-author),
|
||||
(2) optionally enforcing `nsu=` so the index-collision class is closed for good, (3) tidying the
|
||||
two-subtree browse UX, and (4) a **live alarm-fan-out validation** that is the only item carrying
|
||||
real code risk. Once §6 is decided, phase 2 (hardening + UX + docs) is straightforward ScadaBridge
|
||||
work; phases 3–4 need a re-seeded v3 rig.
|
||||
@@ -0,0 +1,263 @@
|
||||
# OtOpcUa v3 cutover — `nsu=` hardening + browse UX (phase 2) Implementation Plan
|
||||
|
||||
> **For Claude:** REQUIRED SUB-SKILL: Use superpowers-extended-cc:executing-plans to implement this plan task-by-task.
|
||||
|
||||
**Goal:** Ship the pure ScadaBridge-side slice of the OtOpcUa v3 dual-namespace cutover (Gitea [#14](https://gitea.dohertylan.com/dohertj2/ScadaBridge/issues/14)) — nudge operators toward durable `nsu=` bindings and tidy the two-subtree browse UX — without a live v3 server.
|
||||
|
||||
**Architecture:** Add one pure string helper in Commons (`OpcUaReferenceForm.IsDurable`) that flags bare `ns=<index>` references (which the v3 raw/uns split makes ambiguous — see scope doc §4 item A). Surface it as a non-blocking authoring warning in the node-browser picker and fix the manual-entry placeholder. Sweep the `ns=`-only doc examples to `nsu=`. This is the D-2-recommended **warn, don't reject** posture: existing `ns=` bindings keep resolving (the `OpcUaNodeReference.Resolve` seam is unchanged and still accepts both forms), so nothing breaks on the current v2 rig.
|
||||
|
||||
**Tech Stack:** C#/.NET 10, xunit, Blazor Server (Bootstrap), `TreatWarningsAsErrors`.
|
||||
|
||||
**Scope boundary (read first):** This plan is **phase 2** of the scope doc `docs/plans/2026-07-23-otopcua-v3-dual-namespace-cutover-scope.md` — items **D, E, F** only. Items **A** (re-author legacy bindings), **B** (native-alarm dedup), and **C** (UNS-bound alarm routing) are **deferred — they require a re-seeded live OtOpcUa v3.0 rig** and live alarm-fan-out validation (scope doc §5 phases 3–5, §7). Do NOT attempt them here. Decisions folded in: **D-1 = ingest both subtrees** (already recorded); **D-2 = warn on bare `ns=`** (this plan); **D-3 = re-author, no migration code** (nothing to build).
|
||||
|
||||
---
|
||||
|
||||
### Task 1: Commons — `OpcUaReferenceForm.IsDurable` durability helper
|
||||
|
||||
**Classification:** small
|
||||
**Estimated implement time:** ~4 min
|
||||
**Parallelizable with:** Task 3
|
||||
|
||||
**Files:**
|
||||
- Create: `src/ZB.MOM.WW.ScadaBridge.Commons/Types/DataConnections/OpcUaReferenceForm.cs`
|
||||
- Test: `tests/ZB.MOM.WW.ScadaBridge.Commons.Tests/Types/DataConnections/OpcUaReferenceFormTests.cs`
|
||||
|
||||
A **pure string** helper — no `Opc.Ua` dependency (Commons has none, and CentralUI references only Commons). It answers one question for the authoring UI: *is this reference index-durable, or a bare server-namespace index we should warn about?* A bare `ns=<n>` with n ≥ 1 is exactly the form the v3 raw/uns namespace split makes ambiguous (scope doc §4 item A). `nsu=`, spec-fixed `ns=0`, and short forms with no namespace prefix (`i=85`, `s=Foo`) are durable and must not warn. Whitespace/empty returns `true` (nothing typed → nothing to warn about).
|
||||
|
||||
**Step 1: Write the failing tests**
|
||||
|
||||
```csharp
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Types.DataConnections;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Commons.Tests.Types.DataConnections;
|
||||
|
||||
public class OpcUaReferenceFormTests
|
||||
{
|
||||
[Theory]
|
||||
// durable → true
|
||||
[InlineData("nsu=https://zb.com/otopcua/raw;s=Line1.Pump.Speed", true)]
|
||||
[InlineData("nsu=https://zb.com/otopcua/uns;s=AreaA/Line1/Pump/Speed", true)]
|
||||
[InlineData("NSU=https://x;s=y", true)] // case-insensitive prefix
|
||||
[InlineData("ns=0;i=85", true)] // spec-fixed namespace 0
|
||||
[InlineData("i=85", true)] // short form, implicit ns0
|
||||
[InlineData("s=Devices.Pump1", true)] // no namespace prefix
|
||||
[InlineData("", true)] // nothing typed
|
||||
[InlineData(" ", true)]
|
||||
[InlineData(null, true)]
|
||||
// bare server-namespace index → false (warn)
|
||||
[InlineData("ns=2;s=MyDevice.Temperature", false)]
|
||||
[InlineData("ns=1;i=1001", false)]
|
||||
[InlineData("ns=10;s=x", false)]
|
||||
[InlineData(" ns=3;s=x ", false)] // trimmed before inspection
|
||||
public void IsDurable_classifies_reference_forms(string? reference, bool expected)
|
||||
{
|
||||
Assert.Equal(expected, OpcUaReferenceForm.IsDurable(reference));
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Step 2: Run to verify it fails**
|
||||
|
||||
Run: `dotnet test tests/ZB.MOM.WW.ScadaBridge.Commons.Tests/ZB.MOM.WW.ScadaBridge.Commons.Tests.csproj --filter OpcUaReferenceFormTests`
|
||||
Expected: FAIL — `OpcUaReferenceForm` does not exist (compile error).
|
||||
|
||||
**Step 3: Write the minimal implementation**
|
||||
|
||||
```csharp
|
||||
using System.Globalization;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Commons.Types.DataConnections;
|
||||
|
||||
/// <summary>
|
||||
/// Pure, protocol-package-free classification of an OPC UA node-reference string's
|
||||
/// <i>form</i> — used by authoring UI to nudge operators toward index-durable bindings.
|
||||
/// </summary>
|
||||
/// <remarks>
|
||||
/// A bare <c>ns=<index>;…</c> reference hard-codes a namespace <i>index</i>, which is
|
||||
/// only meaningful relative to one server's namespace table. The OtOpcUa v3.0 raw/uns
|
||||
/// split makes this actively dangerous: v2's sole custom namespace and v3's <c>raw</c>
|
||||
/// tree both land at <c>ns=2</c>, so a stored <c>ns=2;s=…</c> resolves without error but
|
||||
/// silently now means a raw-tree node (scope doc §4 item A). The durable form is
|
||||
/// <c>nsu=<uri>;…</c>, resolved against the live namespace table at use time by
|
||||
/// <c>OpcUaNodeReference.Resolve</c>. This helper only classifies the string; it does not
|
||||
/// parse or resolve NodeIds (that stays in the DataConnectionLayer seam, which owns the
|
||||
/// <c>Opc.Ua</c> dependency).
|
||||
/// </remarks>
|
||||
public static class OpcUaReferenceForm
|
||||
{
|
||||
/// <summary>
|
||||
/// True when <paramref name="reference"/> is index-durable — the durable
|
||||
/// <c>nsu=<uri>;…</c> form, spec-fixed <c>ns=0</c>, a short form with no namespace
|
||||
/// prefix, or empty/whitespace (nothing to warn about). False only for a bare
|
||||
/// server-namespace index (<c>ns=<n>;…</c> with n ≥ 1).
|
||||
/// </summary>
|
||||
public static bool IsDurable(string? reference)
|
||||
{
|
||||
if (string.IsNullOrWhiteSpace(reference))
|
||||
return true;
|
||||
|
||||
var trimmed = reference.Trim();
|
||||
|
||||
// nsu=<uri>;… is the durable form (and its "ns" prefix must be checked before the
|
||||
// bare "ns=" branch below, since "nsu=" also starts with "ns").
|
||||
if (trimmed.StartsWith("nsu=", StringComparison.OrdinalIgnoreCase))
|
||||
return true;
|
||||
|
||||
// Not a bare namespace-index reference at all → short forms (i=85, s=Foo) that
|
||||
// imply namespace 0, which is spec-fixed and already durable.
|
||||
if (!trimmed.StartsWith("ns=", StringComparison.OrdinalIgnoreCase))
|
||||
return true;
|
||||
|
||||
// ns=<digits>;… — parse the index; ns=0 is spec-fixed (durable), n ≥ 1 is bare.
|
||||
var semicolon = trimmed.IndexOf(';');
|
||||
var digits = semicolon >= 0 ? trimmed[3..semicolon] : trimmed[3..];
|
||||
if (int.TryParse(digits, NumberStyles.None, CultureInfo.InvariantCulture, out var index))
|
||||
return index == 0;
|
||||
|
||||
// Malformed "ns=" with non-numeric index — not a recognisable bare index; leave
|
||||
// the real parse/reject to OpcUaNodeReference.Resolve. Don't warn on it here.
|
||||
return true;
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Step 4: Run to verify it passes**
|
||||
|
||||
Run: `dotnet test tests/ZB.MOM.WW.ScadaBridge.Commons.Tests/ZB.MOM.WW.ScadaBridge.Commons.Tests.csproj --filter OpcUaReferenceFormTests`
|
||||
Expected: PASS (all theory rows).
|
||||
|
||||
**Step 5: Commit**
|
||||
|
||||
```bash
|
||||
git add src/ZB.MOM.WW.ScadaBridge.Commons/Types/DataConnections/OpcUaReferenceForm.cs \
|
||||
tests/ZB.MOM.WW.ScadaBridge.Commons.Tests/Types/DataConnections/OpcUaReferenceFormTests.cs
|
||||
git commit -m "feat(dcl): add OpcUaReferenceForm.IsDurable — flags bare ns= bindings (v3 cutover #14)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 2: Central UI — durable-binding nudge in the node-browser picker (items D + E)
|
||||
|
||||
**Classification:** small
|
||||
**Estimated implement time:** ~4 min
|
||||
**Parallelizable with:** none (uses the Task 1 helper)
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/ZB.MOM.WW.ScadaBridge.CentralUI/Components/Dialogs/NodeBrowserDialog.razor`
|
||||
|
||||
Two UX changes, both non-blocking (D-2 = warn, not reject):
|
||||
1. Manual-entry **placeholder** `ns=2;s=...` → `nsu=<namespace-uri>;s=...` (stop nudging the index form).
|
||||
2. A muted inline **warning** below the manual-entry input when the typed reference is a bare `ns=` index (`!OpcUaReferenceForm.IsDurable(_manualNodeId)`), explaining that index bindings can silently re-point after a server namespace change and to prefer the browse tree / `nsu=` form. Safe for every protocol that uses this dialog: MxGateway node ids carry no `ns=` prefix, so `IsDurable` returns `true` and the warning never shows.
|
||||
|
||||
> Not in scope here: the picker already renders whatever root nodes the server exposes, so the two v3 subtrees (`raw`/`uns`) appear as sibling roots with **no code change**. Tuning the search `maxDepth: 6` for the deeper raw hierarchy is deferred — it needs a live v3 rig to calibrate (scope doc §5 phase 3).
|
||||
|
||||
**Step 1: Add the using and update the placeholder**
|
||||
|
||||
At the top of the file, add to the existing `@using` block:
|
||||
|
||||
```razor
|
||||
@using ZB.MOM.WW.ScadaBridge.Commons.Types.DataConnections
|
||||
```
|
||||
|
||||
Change the manual-node-id input (currently line ~92) placeholder:
|
||||
|
||||
```razor
|
||||
<input class="form-control" @bind="_manualNodeId" @bind:event="oninput"
|
||||
placeholder="nsu=<namespace-uri>;s=..." />
|
||||
```
|
||||
|
||||
(Note the added `@bind:event="oninput"` so the warning below updates as the operator types.)
|
||||
|
||||
**Step 2: Add the inline durability warning**
|
||||
|
||||
Immediately after the manual-entry `input-group` `</div>` (line ~94, before the `modal-footer`), add:
|
||||
|
||||
```razor
|
||||
@if (!OpcUaReferenceForm.IsDurable(_manualNodeId))
|
||||
{
|
||||
<small class="text-warning-emphasis d-block mt-1" data-test="ns-index-warning">
|
||||
This is a bare namespace-<em>index</em> binding. Indexes can silently re-point after a
|
||||
server namespace change — prefer selecting from the tree above (it stores the durable
|
||||
<code>nsu=<uri>;…</code> form).
|
||||
</small>
|
||||
}
|
||||
```
|
||||
|
||||
**Step 3: Build to verify it compiles**
|
||||
|
||||
Run: `dotnet build src/ZB.MOM.WW.ScadaBridge.CentralUI/ZB.MOM.WW.ScadaBridge.CentralUI.csproj`
|
||||
Expected: Build succeeded, 0 warnings (project is `TreatWarningsAsErrors`).
|
||||
|
||||
**Step 4: Commit**
|
||||
|
||||
```bash
|
||||
git add src/ZB.MOM.WW.ScadaBridge.CentralUI/Components/Dialogs/NodeBrowserDialog.razor
|
||||
git commit -m "feat(ui): warn on bare ns= node bindings in the browse picker (v3 cutover #14)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 3: Docs — sweep `ns=` examples to `nsu=` (item F)
|
||||
|
||||
**Classification:** trivial
|
||||
**Estimated implement time:** ~4 min
|
||||
**Parallelizable with:** Task 1
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/ZB.MOM.WW.ScadaBridge.DataConnectionLayer/Adapters/OpcUaDataConnection.cs:15`
|
||||
- Modify: `src/ZB.MOM.WW.ScadaBridge.Commons/Interfaces/Protocol/IBrowsableDataConnection.cs:36`
|
||||
- Modify: `docs/requirements/Component-DataConnectionLayer.md`
|
||||
- Modify: `../scadaproj/CLAUDE.md` (umbrella index — CLAUDE.md propagation rule)
|
||||
|
||||
Only the sole-example `ns=2;s=…` comments get updated. **Leave `OpcUaNodeReference.cs` alone** — it deliberately documents *both* forms and explains why. Do not touch any `ns=` string in code/tests that is an actual binding value or assertion.
|
||||
|
||||
**Step 1: Update the two code-comment examples**
|
||||
|
||||
`OpcUaDataConnection.cs:15` — change:
|
||||
```
|
||||
/// - TagPath → NodeId (e.g., "ns=2;s=MyDevice.Temperature")
|
||||
```
|
||||
to:
|
||||
```
|
||||
/// - TagPath → NodeId (durable form, e.g. "nsu=https://server/ns;s=MyDevice.Temperature";
|
||||
/// the legacy "ns=2;s=..." index form is still accepted — see OpcUaNodeReference)
|
||||
```
|
||||
|
||||
`IBrowsableDataConnection.cs:36` — change the `<param name="NodeId">` example from `"ns=2;s=Devices.Pump1.Speed"` to `"nsu=https://server/ns;s=Devices.Pump1.Speed"`.
|
||||
|
||||
**Step 2: Update the component doc**
|
||||
|
||||
In `docs/requirements/Component-DataConnectionLayer.md`: update any `ns=<index>` example bindings to the `nsu=<uri>` durable form, and add a short note (cross-referencing `docs/plans/2026-07-23-otopcua-v3-dual-namespace-cutover-scope.md`) that OtOpcUa v3.0 splits its address space into `raw` + `uns` namespaces, that ScadaBridge is namespace-URI-durable via `OpcUaNodeReference`, and that the picker nudges operators to the `nsu=` form. First read the file to place the note in the OPC UA / node-reference section.
|
||||
|
||||
**Step 3: Update the umbrella index**
|
||||
|
||||
In `../scadaproj/CLAUDE.md`, find the ScadaBridge entry's OtOpcUa relationship line and note that ScadaBridge is v3-`nsu=`-durable and the #14 dual-namespace cutover is in progress (phase 2 shipped: `nsu=` authoring nudge + browse UX; alarm-fan-out live-gate deferred to a v3 rig). Read the surrounding lines first to match the index's style; keep it to one or two lines.
|
||||
|
||||
**Step 4: Commit**
|
||||
|
||||
```bash
|
||||
git add src/ZB.MOM.WW.ScadaBridge.DataConnectionLayer/Adapters/OpcUaDataConnection.cs \
|
||||
src/ZB.MOM.WW.ScadaBridge.Commons/Interfaces/Protocol/IBrowsableDataConnection.cs \
|
||||
docs/requirements/Component-DataConnectionLayer.md
|
||||
git commit -m "docs(dcl): sweep ns= examples to durable nsu= form (v3 cutover #14)"
|
||||
# umbrella index is a separate repo — commit it there
|
||||
git -C ../scadaproj add CLAUDE.md && git -C ../scadaproj commit -m "docs: note ScadaBridge OtOpcUa v3 nsu= cutover phase 2 (#14)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Deferred — requires a live OtOpcUa v3.0 rig (NOT in this plan)
|
||||
|
||||
Tracked in the scope doc (§5 phases 3–5, §7). Do not attempt without a re-seeded v3 server:
|
||||
|
||||
- **Item A** — re-author existing dev/test bindings across both subtrees via the picker (data, greenfield; D-3 = no migration code).
|
||||
- **Item B** — native-alarm dedup across the raw+uns notifier fan-out; add `ConditionId` dedup **only if** the live server double-delivers.
|
||||
- **Item C** — UNS-bound alarm routing: a v3 condition's `SourceName` is always the RawPath, so a UNS-bound source won't `StartsWith`-match — resolve UNS→Raw via the browseable `Organizes` reference. Build against the real `SourceName`/`SourceNode` payload captured on the rig.
|
||||
|
||||
## Verification (whole-plan, after Task 3)
|
||||
|
||||
```bash
|
||||
dotnet build ZB.MOM.WW.ScadaBridge.slnx
|
||||
dotnet test tests/ZB.MOM.WW.ScadaBridge.Commons.Tests/ZB.MOM.WW.ScadaBridge.Commons.Tests.csproj
|
||||
```
|
||||
Expected: build 0/0; Commons tests green (incl. `OpcUaReferenceFormTests`). No EF migration is added (scope doc §7 — `OpcUaEndpointConfig` has no namespace field).
|
||||
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"planPath": "docs/plans/2026-07-23-otopcua-v3-nsu-hardening-and-browse-ux.md",
|
||||
"tasks": [
|
||||
{"id": 1, "subject": "Task 1: Commons — OpcUaReferenceForm.IsDurable durability helper + tests", "status": "completed", "classification": "small", "parallelizableWith": [3], "commit": "0eb44314"},
|
||||
{"id": 2, "subject": "Task 2: Central UI — durable-binding nudge in node-browser picker (D+E)", "status": "completed", "classification": "small", "blockedBy": [1], "commit": "e04c2617"},
|
||||
{"id": 3, "subject": "Task 3: Docs — sweep ns= examples to nsu= (F)", "status": "completed", "classification": "trivial", "parallelizableWith": [1], "commit": "97afa84f"}
|
||||
],
|
||||
"lastUpdated": "2026-07-23"
|
||||
}
|
||||
@@ -0,0 +1,79 @@
|
||||
# OtOpcUa v3 raw-path live-gate (Gitea #14) — PASS
|
||||
|
||||
**Date:** 2026-07-23 · **Scope:** the raw-subtree slice of #14 (decision: "raw gate now, flag B/C").
|
||||
**Rig:** ScadaBridge docker rig (`scadabridge:latest`, rebuilt with the host-rewrite fix) against the
|
||||
**live `otopcua-dev` v3 cluster** (`otopcua-host:dev`, built 2026-07-23, branch `feat/mesh-phase5`,
|
||||
past the v3 merge `ec6598ce`). Related: scope doc `2026-07-23-otopcua-v3-dual-namespace-cutover-scope.md`,
|
||||
phase-2 plan `2026-07-23-otopcua-v3-nsu-hardening-and-browse-ux.md`.
|
||||
|
||||
## Headline
|
||||
|
||||
ScadaBridge is confirmed **structurally v3-safe against a real dual-namespace server** — durable
|
||||
`nsu=` references are emitted by browse, stored on bindings, and resolved live against the server's
|
||||
`NamespaceArray` to Good-quality reads. The gate also **surfaced and fixed a real DCL robustness
|
||||
bug** (couldn't dial a server advertising `0.0.0.0`). Items **B/C remain deferred** — the dev
|
||||
cluster has no UNS nodes and no native alarms to gate them against.
|
||||
|
||||
## Environment findings
|
||||
|
||||
- v3 cluster publishes both namespaces: `ns=2 = https://zb.com/otopcua/raw`,
|
||||
`ns=3 = https://zb.com/otopcua/uns`. **Raw lands at ns=2 — the exact index collision scope §4
|
||||
item A predicted** (v2's sole namespace and v3 raw both at ns=2).
|
||||
- Raw tree is **sparsely populated**: one materialized live Variable,
|
||||
`nsu=https://zb.com/otopcua/raw;s=sa-modbus2/plc/SiteAOnly` (Int16, AccessLevel=1 → **read-only**).
|
||||
Other drivers (calc1, opcua1, pymodbus, gate-modbus, Area1) expose folders but no live leaf tags.
|
||||
- **UNS (ns=3) is empty** — `ns=3;s=OtOpcUa` → `BadNodeIdUnknown`; no Equipment authored.
|
||||
- **No native alarms** — zero `HasCondition`/`HasEventSource` refs anywhere in the raw tree.
|
||||
|
||||
## The fix this gate produced
|
||||
|
||||
`RealOpcUaClient` took the server-advertised `EndpointDescription.EndpointUrl` verbatim for the
|
||||
session channel. OtOpcUa advertises `opc.tcp://0.0.0.0:4840/OtOpcUa` (its
|
||||
`OpcUaApplicationHostOptions.PublicHostname` defaults to `0.0.0.0`), and dialling `0.0.0.0` from
|
||||
inside the client container hits its own loopback → connect fails. Added
|
||||
`RealOpcUaClient.RewriteEndpointHostForReachability` (swaps the advertised authority for the
|
||||
reachable discovery host/port, preserves scheme+path, no-op when already reachable), called in both
|
||||
`ConnectAsync` and `VerifyEndpointAsync`. Mirrors `CoreClientUtils.SelectEndpoint` and OtOpcUa's own
|
||||
client (`IEndpointDiscovery`). Commit `a1abbff7` on `fix/opcua-discovery-host-rewrite`; 9 unit tests
|
||||
(`RealOpcUaClientEndpointRewriteTests`).
|
||||
|
||||
## Wiring (how the gate was run)
|
||||
|
||||
1. Rebuilt `scadabridge:latest` with the fix (`bash docker/deploy.sh`).
|
||||
2. `docker network connect otopcua-dev_default scadabridge-site-a-a` (and `-b`) — bridged the two
|
||||
docker networks so site-a can resolve `otopcua-dev-site-a-1-1`.
|
||||
3. `data-connection create --site-id 1 --name OtOpcUa-v3-raw --protocol OpcUa` →
|
||||
`endpointUrl=opc.tcp://otopcua-dev-site-a-1-1:4840/OtOpcUa`, securityMode none (id **3044**).
|
||||
4. `deploy artifacts --site-id 1` — pushed the def to site-a's LocalDb + made it live in the DCL.
|
||||
5. Template **2148** (`OtOpcUaV3Probe`) + attribute `SiteA` (Int32,
|
||||
`--data-source nsu=https://zb.com/otopcua/raw;s=sa-modbus2/plc/SiteAOnly`); instance **98**
|
||||
(`otopcua-v3-probe-1`) on site-a; `set-bindings [["SiteA",3044]]`; `deploy instance`.
|
||||
|
||||
## Checks
|
||||
|
||||
| # | Check | Result |
|
||||
|---|-------|--------|
|
||||
| 1 | Connect via `ConnectAsync` (through the rewrite fix) | **PASS** — `data-connection browse` returned the address space |
|
||||
| 2 | Connect via `VerifyEndpointAsync` (second rewrite call site) | **PASS** — `verify-endpoint` → `{"success":true}` |
|
||||
| 3 | Browse emits durable `nsu=` throughout | **PASS** — root `nsu=…/raw;s=OtOpcUa`, drivers + `…;s=sa-modbus2/plc/SiteAOnly` |
|
||||
| 4 | Both namespaces visible / discoverable | **PASS (partial)** — raw browsable; uns registered but empty (nothing to browse) |
|
||||
| 5 | Durable `nsu=` binding resolves live | **PASS** — attribute stores `nsu=…`, subscription resolves it against the live `NamespaceArray` |
|
||||
| 6 | Read round-trip, Good quality, live | **PASS** — `SiteA` Good, values changing 59689→59806→59822→59834 across snapshots |
|
||||
| 7 | Write round-trip | **N/A** — the only materialized tag is read-only (AccessLevel=1); no writable v3 tag to gate |
|
||||
|
||||
## Deferred (unchanged — need OtOpcUa-side authoring on the dev cluster)
|
||||
|
||||
- **Item A** — re-author legacy `ns=2` bindings to `nsu=` (data; the collision is confirmed real here).
|
||||
- **Item B** — native-alarm dedup across raw+uns fan-out: **no native alarms exist** to drive it.
|
||||
- **Item C** — UNS-bound → RawPath alarm routing: **no UNS nodes and no alarms exist** to drive it.
|
||||
|
||||
To gate B/C, the OtOpcUa dev cluster needs UNS Equipment authored (populate ns=3, referencing raw
|
||||
tags) and at least one alarm-configured tag raising a Part-9 condition that fans to raw + equipment
|
||||
notifiers.
|
||||
|
||||
## Rig artifacts created (non-default rig state)
|
||||
|
||||
- ScadaBridge site-a-a / site-a-b attached to `otopcua-dev_default` (dropped on next `docker/deploy.sh`).
|
||||
- Site-a config: DataConnection **3044** `OtOpcUa-v3-raw`, template **2148**, instance **98** (deployed).
|
||||
Harmless; remove with `instance delete 98` / `template delete 2148` / `data-connection delete 3044`
|
||||
if a clean rig is wanted.
|
||||
@@ -137,6 +137,27 @@ Behavior, covered by `SelfFirstSeedBootstrapTests` (real in-process clusters bui
|
||||
|
||||
The docker failover drill (`docker/failover-drill.sh`) proves both directions: `standby` mode kills the younger node (active untouched, zero routing blips); `active` mode kills the active/oldest node and asserts the survivor **takes over while the victim is still down**.
|
||||
|
||||
### Simultaneous cold start — the bootstrap guard (opt-in, Gitea #33)
|
||||
|
||||
Self-first-on-both has a residual cost. When **both** nodes cold-start at the same instant, each is `seed-nodes[0]` for itself, so **each runs `FirstSeedNodeProcess`**, times out waiting for the other, and forms its **own** single-node cluster — a split brain (two oldest members, two singletons, dual-active) that does **not** auto-merge and persists until an operator restarts one side. In-process loopback tests usually converge (the `InitJoin` handshake resolves before either self-join deadline, so row 3 above passes), which is exactly why the split hides until a real deployment powers up both VMs truly in parallel — a shared power event or a hypervisor host reboot. The sister project **OtOpcUa** (same `ZB.MOM.WW.*` Akka topology) reproduced it reliably on its docker rig.
|
||||
|
||||
The **bootstrap guard** eliminates the split without giving up cold-start-alone. It is an **opt-in dark switch** — `ScadaBridge:Cluster:BootstrapGuard:Enabled`, **default off**, so a node with the guard disabled keeps Akka's config-driven self-first auto-join byte-identical. When enabled, the node starts with **no config seed nodes** (`AkkaHostedService.BuildHocon` emits an empty seed list, so Akka does not auto-join) and a coordinator (`ClusterBootstrapCoordinator`, an `IHostedService` registered in both the Central and Site composition roots) picks the join order after a reachability probe:
|
||||
|
||||
- The node with the lexicographically **lower** canonical `host:port` is the **preferred founder**: it joins self-first and forms immediately if no peer answers — no probe, no delay. The tie-break is **case-insensitive** (a hostname-casing mismatch between the two configs must never make both think they are the founder and re-open the split).
|
||||
- The **higher** node TCP-probes its partner's Akka port (a node binds its port well before it forms) up to `PartnerProbeSeconds`: **reachable ⇒ peer-first** (join the founder, never race it); **unreachable after the window ⇒ self-first** (the partner is genuinely down, so cold-start-alone is preserved for the higher node too).
|
||||
- The order is decided **before** a single `Cluster.JoinSeedNodes`, from an explicit reachability signal. The coordinator **never re-forms mid-handshake** — that was the failure mode of the rejected self-form timer above (a `Join(self)` on a bare timeout is not ignored mid-handshake; it wins and islands a node).
|
||||
|
||||
**Residual trade-off (accepted, operator-visible).** Once the higher node commits peer-first it cannot self-form (`JoinSeedNodeProcess` retries forever). If the founder dies in the small probe→join window, the higher node hangs unjoined; the coordinator logs a WARNING after a bounded grace, and a **restart** recovers it (the guard re-runs, finds the founder down, and forms alone). The guard deliberately does not re-decide mid-handshake.
|
||||
|
||||
| Key | Default | Meaning |
|
||||
|---|---|---|
|
||||
| `BootstrapGuard:Enabled` | `false` | Dark switch. Only meaningful on a node that is one of its own two pair seeds; inert elsewhere (single-node install, self absent, 3+ seeds). |
|
||||
| `BootstrapGuard:PartnerProbeSeconds` | `25` | Higher node's probe window. Must comfortably exceed the founder's process-start-to-Akka-bind time, or a slow-but-alive founder is mistaken for dead and the split re-opens. Validated `> 0` at boot when the guard is on. |
|
||||
| `BootstrapGuard:PartnerProbeIntervalMs` | `500` | Interval between probes. Validated `> 0` when enabled. |
|
||||
| `BootstrapGuard:ProbeConnectTimeoutMs` | `1000` | Per-probe TCP connect timeout. Validated `> 0` when enabled. |
|
||||
|
||||
Decision core (`ClusterBootstrapGuard`) is a pure, fully unit-tested function; the probe + `JoinSeedNodes` runtime (`ClusterBootstrapCoordinator`) is covered by real-ActorSystem tests including the load-bearing higher-node-cold-start-alone case and the headline both-cold-start-together-form-one-cluster case (`ClusterBootstrapCoordinatorTests`). The alternative is purely operational (stagger the two VMs' service-manager start / start the founder first); the docker rig's compose `depends_on` does that today, but that does not exist on the real co-located VMs — the guard is the production-faithful fix. Ported from OtOpcUa (`lmxopcua` commit `d1dac87f`); implementing it in both products keeps their failure/recovery model identical, since a shared power event hits both pairs at once.
|
||||
|
||||
### Manual Failover (admin-triggered)
|
||||
|
||||
An Administrator can swap the central pair's roles deliberately — for planned maintenance on the active node, or to move singletons off a node behaving badly without waiting for a crash. The control is the **Trigger failover** button on the central-cluster card of the Health dashboard (`/monitoring/health`).
|
||||
@@ -191,7 +212,7 @@ These values balance failover speed with stability — fast enough that data col
|
||||
|
||||
If both nodes in a cluster fail simultaneously (e.g., site power outage):
|
||||
|
||||
1. **No manual intervention required.** Since both nodes are configured as seed nodes, whichever node starts first forms a new cluster. The second node joins when it starts.
|
||||
1. **No manual intervention required.** Since both nodes are configured as seed nodes, whichever node starts first forms a new cluster (self-first ordering) and the second node joins when it starts. If both power up truly simultaneously, each self-first seed can run `FirstSeedNodeProcess` and form its own 1-node cluster — a split brain; enable the opt-in **bootstrap guard** (Gitea #33, see *Simultaneous cold start* above) on co-located VM pairs that share a power domain to make them converge on one cluster instead.
|
||||
2. **State recovery** (each node has its own local copy of all required data):
|
||||
- **Site clusters**: The Deployment Manager singleton reads deployed configurations from local SQLite and re-creates the full Instance Actor hierarchy. Store-and-forward buffers are already persisted locally. Alarm states re-evaluate from incoming data values.
|
||||
- **Central cluster**: All state is in MS SQL (configuration database). The active node resumes normal operation.
|
||||
|
||||
@@ -46,6 +46,7 @@ Both central and site clusters. Each side has communication actors that handle m
|
||||
- Central routes the request to the appropriate site.
|
||||
- Site reads values from the Instance Actor and responds.
|
||||
- Central returns the response to the external system.
|
||||
- **Implementation**: this brokered round-trip is served by the **Inbound API's routed-site-script path** — a central-side inbound API method script composes the typed `RouteTo*` verbs on `CommunicationService` (`RouteToCall`, `RouteToGetAttributes`, `RouteToSetAttributes`, `RouteToWaitForAttribute`), driven from `InboundAPI/CommunicationServiceInstanceRouter` and sharing `IntegrationTimeout`. An earlier generic `IntegrationCallRequest` primitive was scaffolded for this pattern but never wired to a producer or a site handler; it was removed as dead code (Gitea #32, 2026-07-23).
|
||||
|
||||
### 5. Recipe/Command Delivery (External System → Central → Site)
|
||||
- **Pattern**: Fire-and-forget with acknowledgment.
|
||||
|
||||
@@ -203,6 +203,7 @@ DCL is a clean data pipe on the hot path. Browse is an **opt-in capability** for
|
||||
- `Resolve(reference, session.NamespaceUris)` accepts **both** `ns=<index>;s=<id>` (existing bindings, unchanged) and the durable `nsu=<namespace-uri>;s=<id>`, mapping the URI to the live index at use time. It is used at every parse site — subscribe, read, write, alarm-subscribe, browse. A URI absent from the server's `NamespaceArray` throws with a message naming the URI, rather than binding to whatever sits at that index; a `svr=` reference to another server is rejected outright.
|
||||
- `ToDurable(expandedNodeId, session.NamespaceUris)` is what the browser emits, so **what the picker shows is what gets stored** and newly-authored bindings are index-proof from the start. Namespace 0 is exempt (spec-fixed) and keeps its short form. It also closes a round-trip gap: the browse previously emitted `ExpandedNodeId.ToString()`, which for a URI- or server-index-carrying reference produced a string `NodeId.Parse` could not read back.
|
||||
- Bindings stored before this change keep their `ns=` form and keep working; they are only as durable as the server's namespace order. Re-authoring against a picker (which now emits `nsu=`) is what makes them durable — see `docs/plans/` and Gitea #14 for the OtOpcUa v3.0 cutover, where v2's sole custom namespace and v3's `raw` namespace both sit at index 2, so v2-era bindings resolve against v3 without error while meaning something else.
|
||||
- **OtOpcUa v3.0 dual-namespace cutover**: v3.0 splits OtOpcUa's address space into two namespaces, `raw` (`https://zb.com/otopcua/raw`) and `uns` (`https://zb.com/otopcua/uns`), where v2 exposed one. ScadaBridge is already namespace-URI-durable against this split: `OpcUaNodeReference` resolves a stored `nsu=<uri>;s=...` binding against the live session `NamespaceArray` at use time, so it is index-proof regardless of which namespace a given tag lands in or how many namespaces the server publishes. The node-browser picker already nudges operators toward the durable form, since `ToDurable` is what it emits for newly-authored bindings. See `docs/plans/2026-07-23-otopcua-v3-dual-namespace-cutover-scope.md` for the full cutover scope.
|
||||
- Browse runs against the live session; no caching at DCL.
|
||||
- **Frame-size guard**: the reply crosses the site→central Akka frame (default 128 KB) on a temp Ask actor; an oversized reply is silently discarded by remoting, hanging the picker. The child handler caps each `BrowseNodeResult` to a byte budget (~100 KB) before replying, OR-ing the adapter's own truncation signal into `Truncated`. This is protocol-agnostic (every adapter's reply funnels through it). Per-protocol upstream caps narrow the window first: OPC UA requests at most 500 references per node (continuation point → `Truncated`); MxGateway relies on the gateway's `BrowseChildren` page cap.
|
||||
- **`BrowseNext` paging** (OPC UA): a `Truncated` level no longer forces manual node-id entry. When the OPC UA browse is truncated, the adapter returns the session's continuation point as the reply's opaque `ContinuationToken`; a follow-up `BrowseNodeCommand` carrying that token calls `Session.BrowseNext` to fetch the next page (the picker exposes this as **"Load more"**). Continuation points are session-bound and can expire — `BadContinuationPointInvalid` is caught and the browse restarts from a fresh first page rather than failing. Manual node-id entry remains as a fallback when the site or its session is offline.
|
||||
|
||||
@@ -0,0 +1,169 @@
|
||||
# Environment Variables
|
||||
|
||||
A deep-scan inventory of every environment variable ScadaBridge reads — the Host, the CLI, the
|
||||
DelmiaNotifier client, EF design-time tooling, and the test/CI harness — plus the supporting infra
|
||||
containers. Compiled by scanning `Environment.GetEnvironmentVariable` call sites, the Host's
|
||||
`AddEnvironmentVariables()` config source, `docker/` + `docker-env2/` compose files,
|
||||
`deploy/wonder-app-vd03/install.ps1`, and every `appsettings*.json`.
|
||||
|
||||
**Scan date:** 2026-07-24 · **Branch:** `main`
|
||||
|
||||
Two columns qualify each variable:
|
||||
|
||||
- **Scope** — where it is read: **Runtime** (the running product — Host / CLI / DelmiaNotifier),
|
||||
**Design-time** (`dotnet ef` tooling), **Test** (unit / integration / E2E / perf harness only, never
|
||||
the shipped product), or **Infra** (a sidecar container, not ScadaBridge code).
|
||||
- **In appsettings?** — whether a shipped `appsettings*.json` already carries the same config key
|
||||
(so the env var is an *override* or *secret substitute*), versus being env-only. `${secret:...}`
|
||||
means the key is present but its value is a secret token, not a literal.
|
||||
|
||||
## How configuration resolves (read this first)
|
||||
|
||||
The Host builds its configuration in `src/ZB.MOM.WW.ScadaBridge.Host/Program.cs` in this precedence
|
||||
order (later overrides earlier):
|
||||
|
||||
1. `appsettings.json`
|
||||
2. `appsettings.{SCADABRIDGE_CONFIG}.json` (role file — `Central` or `Site`)
|
||||
3. **`.AddEnvironmentVariables()`** ← every variable below in the "config-override" table
|
||||
4. `.AddCommandLine(args)`
|
||||
|
||||
Then a **pre-host secrets expander** resolves any `${secret:NAME}` tokens left in the merged config,
|
||||
reading the encrypted secrets store whose master key comes from **`ZB_SECRETS_MASTER_KEY`**. Because
|
||||
step 3 runs *before* expansion, a whole-key env override (e.g. `ScadaBridge__Database__ConfigurationDb`)
|
||||
wins over — and short-circuits — the corresponding `${secret:...}` token. That is the deliberate
|
||||
"rollback / dev" path the docker rigs use.
|
||||
|
||||
**Config-key → env-var mapping:** replace each `:` in a config path with `__` (double underscore).
|
||||
So `ScadaBridge:Communication:GrpcPsk` is set via `ScadaBridge__Communication__GrpcPsk`. Any config
|
||||
key in the app can be overridden this way; the table below lists only the ones with an established
|
||||
purpose (secrets kept off disk, or dev-rig wiring).
|
||||
|
||||
---
|
||||
|
||||
## 1. Role & host environment selection
|
||||
|
||||
| Variable | Scope | In appsettings? | Consumed by | Purpose | Potential values | Required |
|
||||
|---|---|---|---|---|---|---|
|
||||
| `SCADABRIDGE_CONFIG` | Runtime | No (it *selects* the `appsettings.{value}.json` file) | Host (`Program.cs:42`) | Selects the role-specific overlay. **Primary role switch.** | `Central`, `Site` | Effectively yes (falls back to `DOTNET_ENVIRONMENT`, then `Production`) |
|
||||
| `DOTNET_ENVIRONMENT` | Runtime | No | Host (`Program.cs:43`) | Fallback role selector when `SCADABRIDGE_CONFIG` is unset; generic .NET environment name. | `Central`, `Site`, `Development`, `Production` | No |
|
||||
| `ASPNETCORE_ENVIRONMENT` | Runtime | No | ASP.NET Core + `DisableLoginGuard` (`Program.cs:438`) | Standard ASP.NET environment. `DisableLogin` is refused unless this is `Development`. | `Development`, `Production` | No (defaults to `Production` behavior) |
|
||||
| `ASPNETCORE_URLS` | Runtime | No (Kestrel binding; not in these appsettings) | Kestrel | HTTP bind addresses. Docker sets `http://+:5000`; install.ps1 sets `http://+:{WebPort}` (8085). | e.g. `http://+:5000`, `http://+:8085` | No |
|
||||
|
||||
> `SCADABRIDGE_CONFIG` and `ASPNETCORE_ENVIRONMENT` are independent. The docker rig runs
|
||||
> `SCADABRIDGE_CONFIG=Central` **and** `ASPNETCORE_ENVIRONMENT=Development` simultaneously.
|
||||
|
||||
## 2. Secrets store master key (KEK)
|
||||
|
||||
| Variable | Scope | In appsettings? | Consumed by | Purpose | Potential values | Required |
|
||||
|---|---|---|---|---|---|---|
|
||||
| `ZB_SECRETS_MASTER_KEY` | Runtime | Name only (`Secrets:MasterKey:EnvVarName` in `appsettings.json`; the value is never committed) | Host secrets expander | The KEK that decrypts the `ZB.MOM.WW.Secrets` store, from which every `${secret:...}` token resolves. A missing reference fails closed (`SecretNotFoundException`) before any SQL/LDAP/cluster wiring. | an opaque base64/hex master key | Yes on any node whose appsettings uses `${secret:...}` (shipped Central + Site do). Docker rigs avoid it with whole-key overrides. |
|
||||
|
||||
The variable *name* is itself configurable via `Secrets:MasterKey:EnvVarName` (defaults to
|
||||
`ZB_SECRETS_MASTER_KEY`). See `docs/operations/2026-07-16-secrets-clustered-master-key.md`.
|
||||
|
||||
## 3. Config-override env vars (`ScadaBridge__*` whole-key)
|
||||
|
||||
Map onto config keys via the `:`→`__` rule. Keep secrets off disk (production, via a secret manager)
|
||||
or wire the dev rigs without a KEK. The `__site-x` / `__site-a` leaf is the literal `siteId`.
|
||||
All are **Runtime** scope (the Host reads them; Host integration-test fixtures set the same names, but
|
||||
that is the same variable, not a separate one).
|
||||
|
||||
| Variable | In appsettings? | Config key | Purpose | Potential values | Required |
|
||||
|---|---|---|---|---|---|
|
||||
| `ScadaBridge__Database__ConfigurationDb` | Yes — `${secret:sql/scadabridge/configdb-connection}` in `appsettings.Central.json` | `ScadaBridge:Database:ConfigurationDb` | Central config DB connection string (MS SQL). Central-only. | `Server=host,1433;Database=ScadaBridgeConfig;User Id=...;Password=...;TrustServerCertificate=true` | Yes on Central — via env **or** the `${secret}` |
|
||||
| `ScadaBridge__Database__MachineDataDb` | No in shipped src appsettings; present in the deploy overlay (`deploy/wonder-app-vd03/appsettings.Central.json`, as `${...}`). install.ps1 derives it as `<DbName>MachineData`. | `ScadaBridge:Database:MachineDataDb` | Central machine-data DB connection string. Central-only. | connection string | Yes on Central |
|
||||
| `ScadaBridge__Security__JwtSigningKey` | Yes — `${secret:security/scadabridge/jwt-signing-key}` in `appsettings.Central.json` | `ScadaBridge:Security:JwtSigningKey` | HMAC-SHA256 signing key for the cookie-embedded session JWT. | ≥32-char secret | Yes on Central |
|
||||
| `ScadaBridge__Security__Ldap__ServiceAccountPassword` | Yes — `${secret:ldap/scadabridge/service-account-password}` in `appsettings.Central.json` | `ScadaBridge:Security:Ldap:ServiceAccountPassword` | LDAP bind-account password. (Renamed from the old flat `...__LdapServiceAccountPassword`.) | the bind password | Yes on Central if LDAP bind is used |
|
||||
| `ScadaBridge__InboundApi__ApiKeyPepper` | No — env/secret only (not in any shipped appsettings) | `ScadaBridge:InboundApi:ApiKeyPepper` | Per-environment pepper for the inbound API-key peppered-HMAC verifier. Hard Central startup requirement since 2026-06-03. | ≥16-char secret, unique per environment | Yes on Central |
|
||||
| `ScadaBridge__Communication__GrpcPsk` | Yes — `${secret:SB-GRPC-PSK-site-1}` in `appsettings.Site.json` | `ScadaBridge:Communication:GrpcPsk` | **Site side** of the per-site gRPC control-plane PSK. StartupValidator refuses to boot a Site node without it; both pair nodes carry the same value. | per-site secret (e.g. `dev-grpc-psk-docker-site-a`); prod `${secret:SB-GRPC-PSK-<siteId>}` | Yes on Site |
|
||||
| `ScadaBridge__Communication__SitePsks__<siteId>` | No in shipped src appsettings (central normally resolves the `SB-GRPC-PSK-<siteId>` secret); present only in the docker-compose env blocks | `ScadaBridge:Communication:SitePsks:<siteId>` | **Central side** of the same key: the value central presents when dialling that site. Must equal the site's `GrpcPsk`. Override for hosts without a master secret store. | matches the site's `GrpcPsk` | On Central, either this **or** the `SB-GRPC-PSK-<siteId>` store secret |
|
||||
|
||||
Docker dev values (dev-only-insecure, committed in `docker/docker-compose.yml`):
|
||||
`ScadaBridge__Communication__SitePsks__site-a = dev-grpc-psk-docker-site-a` (also `site-b`, `site-c`);
|
||||
`docker-env2` uses `SitePsks__site-x = dev-grpc-psk-docker-env2-site-x`.
|
||||
|
||||
> Any other config key is overridable the same way (e.g. `ScadaBridge__Cluster__SplitBrainResolverStrategy`,
|
||||
> `ScadaBridge__Logging__MinimumLevel`) — these are the standard .NET config binder, not "special" env
|
||||
> vars. Prefer `appsettings` files for non-secret values.
|
||||
|
||||
### `${secret:...}` tokens (resolved from the store, NOT env vars)
|
||||
|
||||
Not environment variables — config values resolved by the secrets expander using `ZB_SECRETS_MASTER_KEY`.
|
||||
All are **Runtime** scope and, by definition, present **In appsettings** as tokens.
|
||||
|
||||
| Token (in `appsettings`) | Meaning |
|
||||
|---|---|
|
||||
| `${secret:sql/scadabridge/configdb-connection}` | Central config DB connection string |
|
||||
| `${secret:ldap/scadabridge/service-account-password}` | LDAP bind password |
|
||||
| `${secret:security/scadabridge/jwt-signing-key}` | JWT signing key |
|
||||
| `${secret:SB-GRPC-PSK-<siteId>}` | Per-site gRPC PSK (site side); central resolves `SB-GRPC-PSK-{siteId}` at channel-build time |
|
||||
|
||||
## 4. CLI (`scadabridge` / `ZB.MOM.WW.ScadaBridge.CLI`)
|
||||
|
||||
Resolution order per value: command-line flag → env var → `~/.scadabridge/config.json`. The CLI does
|
||||
**not** read `appsettings.json` — its config file is `~/.scadabridge/config.json` (a separate schema).
|
||||
|
||||
| Variable | Scope | In appsettings? | Purpose | Potential values | Required |
|
||||
|---|---|---|---|---|---|
|
||||
| `SCADABRIDGE_MANAGEMENT_URL` | Runtime (CLI) | No — CLI config file has `managementUrl` instead | Management API base URL. Overridden by `--url`; overrides the config file. Missing all three → exit `1`, `NO_URL`. | e.g. `http://localhost:9000` | One of flag/env/config |
|
||||
| `SCADABRIDGE_FORMAT` | Runtime (CLI) | No — CLI config file has `defaultFormat` instead | Default output format. Overridden by `--format`. | `json`, `table` | No (defaults to `json`) |
|
||||
| `SCADABRIDGE_USERNAME` | Runtime (CLI) | No | LDAP username fallback when `--username` absent. | e.g. `multi-role` | No (some commands need auth) |
|
||||
| `SCADABRIDGE_PASSWORD` | Runtime (CLI) | No | LDAP password fallback when `--password` absent. **Preferred over `--password`** — avoids leaking into process listings / shell history. | the user's password | No (some commands need auth) |
|
||||
|
||||
## 5. DelmiaNotifier client (`WWNotifier.exe`)
|
||||
|
||||
| Variable | Scope | In appsettings? | Purpose | Potential values | Required |
|
||||
|---|---|---|---|---|---|
|
||||
| `SCADABRIDGE_API_KEY` | Runtime (client) | No — the notifier's `appsettings.json` holds only `BaseUrls`/`TimeoutSeconds`/`LogPath`; the key is **env-only by design** | Inbound API key sent as `X-API-Key`. Read only from the environment, never a file. Missing/empty → prints `API key not configured (SCADABRIDGE_API_KEY)`, exit `-1`, no HTTP attempt. | e.g. `sbk_...` | Yes |
|
||||
|
||||
> This is also why `System.Environment` is wholly forbidden in the script trust model — it would let a
|
||||
> site/inbound script exfiltrate `SCADABRIDGE_API_KEY` and other host env secrets.
|
||||
|
||||
## 6. EF Core design-time tooling
|
||||
|
||||
| Variable | Scope | In appsettings? | Purpose | Potential values | Required |
|
||||
|---|---|---|---|---|---|
|
||||
| `SCADABRIDGE_DESIGNTIME_CONNECTIONSTRING` | Design-time | Fallback for `ScadaBridge:Database:ConfigurationDb` — the factory reads that appsettings key first, then this env var | Connection string for `dotnet ef` migrations tooling when the Host's appsettings isn't the source. No hardcoded fallback — missing → actionable failure. | a config DB connection string | Only for `dotnet ef` when appsettings is unavailable |
|
||||
|
||||
## 7. Test / CI-only
|
||||
|
||||
Not read by the running product — the test harness only. Most default gracefully (skip, or use a
|
||||
local docker value) when unset.
|
||||
|
||||
| Variable | Scope | In appsettings? | Consumed by | Purpose | Potential values |
|
||||
|---|---|---|---|---|---|
|
||||
| `SCADABRIDGE_MSSQL_TEST_CONN` | Test | No | `MsSqlMigrationFixture`, AuditLog migration tests | Admin (sa) connection to a running MS SQL; absent → local docker sa default; drives `Skip.If`. | sa connection string |
|
||||
| `SCADABRIDGE_PLAYWRIGHT_DB` | Test | No | `PlaywrightDbConnection` | DB connection for Playwright E2E seeding; absent → `localhost:1433`. | connection string |
|
||||
| `SCADABRIDGE_CLI_DLL` | Test | No | `CliRunner` | Absolute path to the built CLI dll for Playwright when auto-probing fails. | path to `...CLI.dll` |
|
||||
| `SCADABRIDGE_AUDIT_FILTER_4KB_P95_US` | Test | No | `HotPathLatencyTests` | Override P95 latency threshold (µs) for the 4 KB audit-filter perf test. | positive number (default 5) |
|
||||
| `SCADABRIDGE_AUDIT_FILTER_RAW_P95_US` | Test | No | `HotPathLatencyTests` | Override P95 threshold for the raw-capture perf test. | positive number |
|
||||
| `SCADABRIDGE_AUDIT_FILTER_NOOP_P95_US` | Test | No | `HotPathLatencyTests` | Override P95 threshold for the no-op fast-path perf test. | positive number (default 5) |
|
||||
| `HOME` / `USERPROFILE` | Test | No | `CliConfig` tests | Redirected to locate `~/.scadabridge/config.json` under test. | temp dir path |
|
||||
|
||||
The Host test fixtures (`CentralDbTestEnvironment`, `ScadaBridgeWebApplicationFactory`,
|
||||
`StartupValidationTests`) set the **section-3 config-override vars** (`ScadaBridge__Database__*`,
|
||||
`ScadaBridge__InboundApi__ApiKeyPepper`, etc.) and `DOTNET_ENVIRONMENT=Central` at runtime; those are
|
||||
the same variables already documented above, not new ones.
|
||||
|
||||
## 8. Supporting infrastructure containers (not the app)
|
||||
|
||||
Set on the sidecar/infra containers in `infra/docker-compose.yml`, consumed by those images, not by
|
||||
ScadaBridge code — included for completeness when standing up a rig.
|
||||
|
||||
| Variable | Scope | In appsettings? | Container | Purpose | Value used |
|
||||
|---|---|---|---|---|---|
|
||||
| `MSSQL_SA_PASSWORD` | Infra | No | `scadabridge-mssql` | SQL Server `sa` password. | `ScadaBridge_Dev1#` (dev-only) |
|
||||
| `MSSQL_PID` | Infra | No | `scadabridge-mssql` | SQL Server edition. | `Developer` |
|
||||
|
||||
---
|
||||
|
||||
## Quick reference: minimum sets to boot a real deploy
|
||||
|
||||
**Central node** (via env, or the `${secret:...}` equivalents + `ZB_SECRETS_MASTER_KEY`):
|
||||
`SCADABRIDGE_CONFIG=Central`, `ScadaBridge__Database__ConfigurationDb`,
|
||||
`ScadaBridge__Database__MachineDataDb`, `ScadaBridge__Security__JwtSigningKey`,
|
||||
`ScadaBridge__Security__Ldap__ServiceAccountPassword`, `ScadaBridge__InboundApi__ApiKeyPepper`, and
|
||||
for each managed site either the `SB-GRPC-PSK-<siteId>` secret or `ScadaBridge__Communication__SitePsks__<siteId>`.
|
||||
|
||||
**Site node:** `SCADABRIDGE_CONFIG=Site` and `ScadaBridge__Communication__GrpcPsk` (or its
|
||||
`${secret:SB-GRPC-PSK-<siteId>}` form). StartupValidator hard-fails without the PSK.
|
||||
@@ -1,6 +1,7 @@
|
||||
@using Microsoft.AspNetCore.Components.Web
|
||||
@using ZB.MOM.WW.ScadaBridge.Commons.Interfaces.Protocol
|
||||
@using ZB.MOM.WW.ScadaBridge.Commons.Messages.Management
|
||||
@using ZB.MOM.WW.ScadaBridge.Commons.Types.DataConnections
|
||||
@using ZB.MOM.WW.ScadaBridge.CentralUI.Services
|
||||
@inject IBrowseService BrowseService
|
||||
|
||||
@@ -89,9 +90,17 @@
|
||||
|
||||
<div class="input-group">
|
||||
<span class="input-group-text">Manual node id:</span>
|
||||
<input class="form-control" @bind="_manualNodeId" placeholder="ns=2;s=..." />
|
||||
<input class="form-control" @bind="_manualNodeId" @bind:event="oninput" placeholder="nsu=<namespace-uri>;s=..." />
|
||||
<button class="btn btn-outline-secondary" @onclick="UseManual" disabled="@string.IsNullOrWhiteSpace(_manualNodeId)">Use</button>
|
||||
</div>
|
||||
@if (!OpcUaReferenceForm.IsDurable(_manualNodeId))
|
||||
{
|
||||
<small class="text-warning-emphasis d-block mt-1" data-test="ns-index-warning">
|
||||
This is a bare namespace-<em>index</em> binding. Indexes can silently re-point after a
|
||||
server namespace change — prefer selecting from the tree above (it stores the durable
|
||||
<code>nsu=<uri>;…</code> form).
|
||||
</small>
|
||||
}
|
||||
</div>
|
||||
<div class="modal-footer">
|
||||
<span class="me-auto text-muted">Selected: <code>@(_selectedNodeId ?? "(none)")</code></span>
|
||||
|
||||
@@ -0,0 +1,155 @@
|
||||
using System.Text.RegularExpressions;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.ClusterInfrastructure;
|
||||
|
||||
/// <summary>
|
||||
/// Pure decision logic for the simultaneous-cold-start split-brain guard (Gitea #33).
|
||||
/// </summary>
|
||||
/// <remarks>
|
||||
/// <para>
|
||||
/// Both nodes of a 2-node pair are self-first seeds so <b>either</b> can cold-start alone when
|
||||
/// its partner is dead (see <see cref="ClusterOptions.SeedNodes"/>). The cost is that when
|
||||
/// BOTH cold-start at the same instant, each runs Akka's <c>FirstSeedNodeProcess</c>, times out
|
||||
/// waiting for the other, and forms its OWN single-node cluster — a split brain (two oldest
|
||||
/// nodes, two singletons, dual-active). Two independent clusters do not auto-merge, so the
|
||||
/// split persists until an operator restarts one side. It bites on any event that powers up
|
||||
/// both site VMs together (site power restoration, hypervisor host reboot).
|
||||
/// </para>
|
||||
/// <para>
|
||||
/// This guard breaks the symmetry deterministically without giving up cold-start-alone. The
|
||||
/// node with the lexicographically <b>lower</b> canonical address is the <i>preferred
|
||||
/// founder</i>: it always uses self-first order and forms immediately. The <b>higher</b> node
|
||||
/// first probes whether its partner's Akka endpoint is reachable (see
|
||||
/// <see cref="ClusterBootstrapCoordinator"/>): reachable ⇒ peer-first order so it JOINS the
|
||||
/// founder rather than racing it; unreachable after the probe window ⇒ the partner is genuinely
|
||||
/// down, so it falls back to self-first and forms alone. Exactly one node founds when both are
|
||||
/// present; the lower node always founds when it starts alone.
|
||||
/// </para>
|
||||
/// <para>
|
||||
/// <b>Residual trade-off (accepted).</b> Once the higher node observes the partner reachable it
|
||||
/// commits to peer-first — Akka's <c>JoinSeedNodeProcess</c> retries <c>InitJoin</c> forever and
|
||||
/// never self-forms. If the founder dies in the small window between the probe succeeding and
|
||||
/// the join completing, the higher node hangs unjoined until it is restarted (a restart re-runs
|
||||
/// the guard, finds the founder unreachable, and self-forms). We deliberately do NOT re-decide
|
||||
/// mid-handshake — that is exactly the retired <c>SelfFormAfter</c> failure mode. Instead
|
||||
/// <see cref="ClusterBootstrapCoordinator"/> makes the hang operator-visible with a warning if
|
||||
/// the node has not come Up within a bounded time after committing peer-first.
|
||||
/// </para>
|
||||
/// <para>
|
||||
/// This decides the join order BEFORE issuing a single <c>JoinSeedNodes</c>, from an explicit
|
||||
/// reachability signal — unlike a timer-based watchdog, which fires a <c>Join(self)</c>
|
||||
/// mid-handshake on a bare timeout and could not tell "no seed answered" from "a join is in
|
||||
/// flight", islanding a node a failover had just bounced (the rejected self-form-watchdog
|
||||
/// design; see <c>SelfFirstSeedBootstrapTests</c>).
|
||||
/// </para>
|
||||
/// </remarks>
|
||||
public static class ClusterBootstrapGuard
|
||||
{
|
||||
private static readonly Regex SeedPattern =
|
||||
new(@"^akka(?:\.[a-z0-9]+)?://[^@/]+@(?<host>[^:/]+):(?<port>\d+)", RegexOptions.IgnoreCase | RegexOptions.Compiled);
|
||||
|
||||
/// <summary>The analyzed bootstrap role of this node within its seed list.</summary>
|
||||
/// <param name="Applies">True only when the guard engages: exactly two seeds, one of them this node.</param>
|
||||
/// <param name="IsFounder">True when this node is the preferred founder (lower address) — form immediately, no probe.</param>
|
||||
/// <param name="PartnerHost">The partner's advertised host (to probe when this node is the higher one); null when not applicable.</param>
|
||||
/// <param name="PartnerPort">The partner's Akka port.</param>
|
||||
/// <param name="SelfFirstOrder"><c>[self, partner]</c> — self-first; forms a new cluster when no peer answers.</param>
|
||||
/// <param name="PeerFirstOrder"><c>[partner, self]</c> — peer-first; joins the partner's cluster, never self-forms while it is up.</param>
|
||||
public sealed record BootstrapRole(
|
||||
bool Applies,
|
||||
bool IsFounder,
|
||||
string? PartnerHost,
|
||||
int PartnerPort,
|
||||
string[] SelfFirstOrder,
|
||||
string[] PeerFirstOrder);
|
||||
|
||||
/// <summary>
|
||||
/// Analyzes the configured seed list to decide this node's bootstrap role. The guard engages only
|
||||
/// for the pair shape (exactly two seeds, one of them this node); any other shape
|
||||
/// (single-seed / single-node install, three+ seeds, self absent) returns
|
||||
/// <see cref="BootstrapRole.Applies"/> = false and the caller joins the configured seeds unchanged.
|
||||
/// </summary>
|
||||
/// <param name="seeds">The configured seed URIs (self-first by convention, but order is not relied on here).</param>
|
||||
/// <param name="selfHost">This node's advertised host (<c>NodeOptions.NodeHostname</c>).</param>
|
||||
/// <param name="selfPort">This node's Akka remoting port (<c>NodeOptions.RemotingPort</c>).</param>
|
||||
/// <returns>The analyzed role; never null.</returns>
|
||||
public static BootstrapRole Analyze(IReadOnlyList<string> seeds, string selfHost, int selfPort)
|
||||
{
|
||||
ArgumentNullException.ThrowIfNull(seeds);
|
||||
ArgumentException.ThrowIfNullOrWhiteSpace(selfHost);
|
||||
|
||||
var na = new BootstrapRole(false, false, null, 0, Array.Empty<string>(), Array.Empty<string>());
|
||||
if (seeds.Count != 2) return na;
|
||||
|
||||
var selfKey = Key(selfHost, selfPort);
|
||||
string? self = null, partner = null;
|
||||
string? partnerHost = null;
|
||||
var partnerPort = 0;
|
||||
|
||||
foreach (var seed in seeds)
|
||||
{
|
||||
if (!TryParse(seed, out var host, out var port)) return na;
|
||||
if (string.Equals(Key(host, port), selfKey, StringComparison.OrdinalIgnoreCase))
|
||||
self = seed;
|
||||
else
|
||||
{
|
||||
partner = seed;
|
||||
partnerHost = host;
|
||||
partnerPort = port;
|
||||
}
|
||||
}
|
||||
|
||||
// Self must be exactly one of the two, and the other must be a distinct partner.
|
||||
if (self is null || partner is null || partnerHost is null) return na;
|
||||
|
||||
var selfFirst = new[] { self, partner };
|
||||
var peerFirst = new[] { partner, self };
|
||||
|
||||
// Deterministic tie-break: the lower canonical address is the preferred founder. Both nodes
|
||||
// compute the same comparison (same two seeds), so exactly one is the founder. Case-INSENSITIVE
|
||||
// to match how self/partner are classified above (and StartupValidator.SeedNodeIsSelf):
|
||||
// container/DNS hostnames are conventionally case-insensitive, and a casing difference between
|
||||
// the two sides' seed config must NOT make both think they are the founder (which would reopen
|
||||
// the very split this guard closes).
|
||||
var isFounder = string.Compare(
|
||||
Key(selfHost, selfPort),
|
||||
Key(partnerHost, partnerPort),
|
||||
StringComparison.OrdinalIgnoreCase) < 0;
|
||||
|
||||
return new BootstrapRole(true, isFounder, partnerHost, partnerPort, selfFirst, peerFirst);
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// The order the higher (non-founder) node uses once it has learned whether its partner is
|
||||
/// reachable: peer-first when the founder is up (join it), self-first when the founder is absent
|
||||
/// (form alone — cold-start-alone preserved).
|
||||
/// </summary>
|
||||
/// <param name="role">The analyzed role (must be applicable and NOT the founder).</param>
|
||||
/// <param name="partnerReachable">Whether the partner's Akka endpoint became reachable within the probe window.</param>
|
||||
/// <returns>The seed order to pass to <c>Cluster.JoinSeedNodes</c>.</returns>
|
||||
public static string[] HigherNodeOrder(BootstrapRole role, bool partnerReachable)
|
||||
{
|
||||
ArgumentNullException.ThrowIfNull(role);
|
||||
return partnerReachable ? role.PeerFirstOrder : role.SelfFirstOrder;
|
||||
}
|
||||
|
||||
/// <summary>Parses <c>akka.tcp://system@host:port[/...]</c> into host + port.</summary>
|
||||
/// <param name="seed">The seed URI.</param>
|
||||
/// <param name="host">The parsed advertised host.</param>
|
||||
/// <param name="port">The parsed Akka port.</param>
|
||||
/// <returns>True when the seed matched the expected shape.</returns>
|
||||
public static bool TryParse(string? seed, out string host, out int port)
|
||||
{
|
||||
host = string.Empty;
|
||||
port = 0;
|
||||
if (string.IsNullOrWhiteSpace(seed)) return false;
|
||||
|
||||
var m = SeedPattern.Match(seed.Trim());
|
||||
if (!m.Success) return false;
|
||||
if (!int.TryParse(m.Groups["port"].Value, out port)) return false;
|
||||
host = m.Groups["host"].Value;
|
||||
return true;
|
||||
}
|
||||
|
||||
private static string Key(string host, int port) => $"{host}:{port}";
|
||||
}
|
||||
@@ -114,4 +114,50 @@ public class ClusterOptions
|
||||
/// dials forever — log noise and a defeated validation intent (review 01).
|
||||
/// </summary>
|
||||
public bool AllowSingleNodeCluster { get; set; }
|
||||
|
||||
/// <summary>
|
||||
/// The simultaneous-cold-start split-brain guard (<c>ScadaBridge:Cluster:BootstrapGuard</c>).
|
||||
/// Default OFF — a dark switch, so existing deployments and tests keep Akka's config-driven
|
||||
/// self-first auto-join byte-identical. See <see cref="ClusterBootstrapGuard"/> for the
|
||||
/// decision logic and <c>ClusterBootstrapCoordinator</c> for the runtime (Gitea #33).
|
||||
/// </summary>
|
||||
public ClusterBootstrapGuardOptions BootstrapGuard { get; set; } = new();
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Configuration for the simultaneous-cold-start split-brain guard. Both nodes of a pair are
|
||||
/// self-first seeds so <b>either</b> can cold-start alone when its partner is dead — but the cost
|
||||
/// is that when BOTH cold-start at the same instant each runs Akka's <c>FirstSeedNodeProcess</c>,
|
||||
/// times out waiting for the other, and forms its own single-node cluster (a split brain that does
|
||||
/// not auto-merge). When <see cref="Enabled"/>, the node does NOT auto-join from its config seeds
|
||||
/// (<c>AkkaHostedService.BuildHocon</c> emits an empty seed list); a coordinator picks the join
|
||||
/// order — founder self-first / joiner peer-first — after a reachability probe. See
|
||||
/// <see cref="ClusterBootstrapGuard"/>.
|
||||
/// </summary>
|
||||
public sealed class ClusterBootstrapGuardOptions
|
||||
{
|
||||
/// <summary>
|
||||
/// Whether the guard is active. Default <see langword="false"/> — Akka auto-joins from the
|
||||
/// config seed list exactly as before. Only meaningful on a node that is one of its own two
|
||||
/// pair seeds; inert everywhere else (single-seed site node, self absent, 3+ seeds).
|
||||
/// </summary>
|
||||
public bool Enabled { get; set; }
|
||||
|
||||
/// <summary>
|
||||
/// How long the higher-address node probes its partner's Akka endpoint before concluding the
|
||||
/// partner is dead and forming alone. Must comfortably exceed the partner's worst-case
|
||||
/// process-start-to-Akka-bind time so a slow-but-alive partner is never mistaken for a dead one
|
||||
/// (which would re-open the split). Default 25s. Validated <c>> 0</c> at boot when the guard is on.
|
||||
/// </summary>
|
||||
public int PartnerProbeSeconds { get; set; } = 25;
|
||||
|
||||
/// <summary>
|
||||
/// The interval between partner reachability probes, in milliseconds. Default 500ms.
|
||||
/// </summary>
|
||||
public int PartnerProbeIntervalMs { get; set; } = 500;
|
||||
|
||||
/// <summary>
|
||||
/// The per-probe TCP connect timeout, in milliseconds. Default 1000ms.
|
||||
/// </summary>
|
||||
public int ProbeConnectTimeoutMs { get; set; } = 1000;
|
||||
}
|
||||
|
||||
@@ -30,6 +30,8 @@ public sealed class ClusterOptionsValidator : OptionsValidatorBase<ClusterOption
|
||||
/// <inheritdoc />
|
||||
protected override void Validate(ValidationBuilder builder, ClusterOptions options)
|
||||
{
|
||||
ValidateBootstrapGuard(builder, options.BootstrapGuard);
|
||||
|
||||
// The design doc states "both nodes are seed nodes — each node lists
|
||||
// both itself and its partner" so a properly-configured deployment lists
|
||||
// two. Accepting a single-seed configuration silently defeats the
|
||||
@@ -78,4 +80,33 @@ public sealed class ClusterOptionsValidator : OptionsValidatorBase<ClusterOption
|
||||
+ "oldest node can run as an isolated single-node cluster during a partition while the "
|
||||
+ "younger node forms its own, producing two live clusters.");
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// When the bootstrap guard is enabled (Gitea #33), its timing knobs must be positive. A
|
||||
/// zero/negative <see cref="ClusterBootstrapGuardOptions.PartnerProbeSeconds"/> in particular
|
||||
/// silently degrades the guard to "never wait, always conclude the partner is dead" — the
|
||||
/// higher node would form alone immediately and re-open the very split the guard exists to
|
||||
/// close. Fail fast at boot rather than producing that silent degradation. Nothing is checked
|
||||
/// when the guard is off (the knobs are inert), so a disabled guard never blocks a boot.
|
||||
/// </summary>
|
||||
private static void ValidateBootstrapGuard(ValidationBuilder builder, ClusterBootstrapGuardOptions? guard)
|
||||
{
|
||||
if (guard is null || !guard.Enabled)
|
||||
{
|
||||
return;
|
||||
}
|
||||
|
||||
builder.RequireThat(guard.PartnerProbeSeconds > 0,
|
||||
$"ClusterOptions.BootstrapGuard.PartnerProbeSeconds must be > 0 when the guard is enabled "
|
||||
+ $"(was {guard.PartnerProbeSeconds}); a non-positive value makes the higher node conclude its "
|
||||
+ "partner is dead without waiting and form alone, re-opening the split the guard prevents.");
|
||||
|
||||
builder.RequireThat(guard.PartnerProbeIntervalMs > 0,
|
||||
$"ClusterOptions.BootstrapGuard.PartnerProbeIntervalMs must be > 0 when the guard is enabled "
|
||||
+ $"(was {guard.PartnerProbeIntervalMs}).");
|
||||
|
||||
builder.RequireThat(guard.ProbeConnectTimeoutMs > 0,
|
||||
$"ClusterOptions.BootstrapGuard.ProbeConnectTimeoutMs must be > 0 when the guard is enabled "
|
||||
+ $"(was {guard.ProbeConnectTimeoutMs}).");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -33,7 +33,7 @@ public record BrowseChildrenResult(
|
||||
bool Truncated,
|
||||
string? ContinuationToken = null);
|
||||
|
||||
/// <param name="NodeId">Server-issued node identifier (e.g. <c>"ns=2;s=Devices.Pump1.Speed"</c>).</param>
|
||||
/// <param name="NodeId">Server-issued node identifier (e.g. <c>"nsu=https://server/ns;s=Devices.Pump1.Speed"</c>).</param>
|
||||
/// <param name="DisplayName">Human-readable display name from the server's DisplayName attribute.</param>
|
||||
/// <param name="NodeClass">Classifies the node for UI purposes (Variable rows are selectable; Object rows are navigable).</param>
|
||||
/// <param name="HasChildren">Hint so the UI can render an expand chevron without a second roundtrip.</param>
|
||||
|
||||
@@ -1,14 +0,0 @@
|
||||
namespace ZB.MOM.WW.ScadaBridge.Commons.Messages.Integration;
|
||||
|
||||
/// <summary>
|
||||
/// Request routed from central to site to invoke an integration method
|
||||
/// (external system call or notification) on behalf of the central UI or API.
|
||||
/// </summary>
|
||||
public record IntegrationCallRequest(
|
||||
string CorrelationId,
|
||||
string SiteId,
|
||||
string InstanceUniqueName,
|
||||
string TargetSystemName,
|
||||
string MethodName,
|
||||
IReadOnlyDictionary<string, object?> Parameters,
|
||||
DateTimeOffset Timestamp);
|
||||
@@ -1,12 +0,0 @@
|
||||
namespace ZB.MOM.WW.ScadaBridge.Commons.Messages.Integration;
|
||||
|
||||
/// <summary>
|
||||
/// Response for an integration call routed through central-site communication.
|
||||
/// </summary>
|
||||
public record IntegrationCallResponse(
|
||||
string CorrelationId,
|
||||
string SiteId,
|
||||
bool Success,
|
||||
string? ResultJson,
|
||||
string? ErrorMessage,
|
||||
DateTimeOffset Timestamp);
|
||||
@@ -0,0 +1,55 @@
|
||||
using System.Globalization;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Commons.Types.DataConnections;
|
||||
|
||||
/// <summary>
|
||||
/// Pure, protocol-package-free classification of an OPC UA node-reference string's
|
||||
/// <i>form</i> — used by authoring UI to nudge operators toward index-durable bindings.
|
||||
/// </summary>
|
||||
/// <remarks>
|
||||
/// A bare <c>ns=<index>;…</c> reference hard-codes a namespace <i>index</i>, which is
|
||||
/// only meaningful relative to one server's namespace table. The OtOpcUa v3.0 raw/uns
|
||||
/// split makes this actively dangerous: v2's sole custom namespace and v3's <c>raw</c>
|
||||
/// tree both land at <c>ns=2</c>, so a stored <c>ns=2;s=…</c> resolves without error but
|
||||
/// silently now means a raw-tree node. The durable form is <c>nsu=<uri>;…</c>,
|
||||
/// resolved against the live namespace table at use time by
|
||||
/// <c>OpcUaNodeReference.Resolve</c>. This helper only classifies the string; it does not
|
||||
/// parse or resolve NodeIds (that stays in the DataConnectionLayer seam, which owns the
|
||||
/// <c>Opc.Ua</c> dependency).
|
||||
/// </remarks>
|
||||
public static class OpcUaReferenceForm
|
||||
{
|
||||
/// <summary>
|
||||
/// True when <paramref name="reference"/> is index-durable — the durable
|
||||
/// <c>nsu=<uri>;…</c> form, spec-fixed <c>ns=0</c>, a short form with no namespace
|
||||
/// prefix, or empty/whitespace (nothing to warn about). False only for a bare
|
||||
/// server-namespace index (<c>ns=<n>;…</c> with n ≥ 1).
|
||||
/// </summary>
|
||||
public static bool IsDurable(string? reference)
|
||||
{
|
||||
if (string.IsNullOrWhiteSpace(reference))
|
||||
return true;
|
||||
|
||||
var trimmed = reference.Trim();
|
||||
|
||||
// nsu=<uri>;… is the durable form (and its "ns" prefix must be checked before the
|
||||
// bare "ns=" branch below, since "nsu=" also starts with "ns").
|
||||
if (trimmed.StartsWith("nsu=", StringComparison.OrdinalIgnoreCase))
|
||||
return true;
|
||||
|
||||
// Not a bare namespace-index reference at all → short forms (i=85, s=Foo) that
|
||||
// imply namespace 0, which is spec-fixed and already durable.
|
||||
if (!trimmed.StartsWith("ns=", StringComparison.OrdinalIgnoreCase))
|
||||
return true;
|
||||
|
||||
// ns=<digits>;… — parse the index; ns=0 is spec-fixed (durable), n ≥ 1 is bare.
|
||||
var semicolon = trimmed.IndexOf(';');
|
||||
var digits = semicolon >= 0 ? trimmed[3..semicolon] : trimmed[3..];
|
||||
if (int.TryParse(digits, NumberStyles.None, CultureInfo.InvariantCulture, out var index))
|
||||
return index == 0;
|
||||
|
||||
// Malformed "ns=" with non-numeric index — not a recognisable bare index; leave
|
||||
// the real parse/reject to OpcUaNodeReference.Resolve. Don't warn on it here.
|
||||
return true;
|
||||
}
|
||||
}
|
||||
@@ -35,11 +35,6 @@ namespace ZB.MOM.WW.ScadaBridge.Communication.Actors;
|
||||
/// singleton proxy (design §7.3). This is the same target the actor used; the
|
||||
/// extraction does not "fix" it.
|
||||
/// </para>
|
||||
/// <para>
|
||||
/// <b>Excluded by design:</b> <c>IntegrationCallRequest</c> — the 29th command, dead at
|
||||
/// both ends. It never enters this dispatcher; the actor keeps its own vestigial handler
|
||||
/// for it (28 of 29 migrate).
|
||||
/// </para>
|
||||
/// </remarks>
|
||||
public sealed class SiteCommandDispatcher
|
||||
{
|
||||
|
||||
@@ -61,12 +61,6 @@ public class SiteCommunicationActor : ReceiveActor, IWithTimers
|
||||
/// <summary>The transport supplied by the Host (null falls back to <see cref="NoOpCentralTransport"/>).</summary>
|
||||
private readonly ICentralTransport? _injectedTransport;
|
||||
|
||||
/// <summary>
|
||||
/// Handler for the vestigial <see cref="IntegrationCallRequest"/> — the one command NOT
|
||||
/// migrated to the dispatcher (dead at both ends), so it still routes on the actor.
|
||||
/// </summary>
|
||||
private IActorRef? _integrationHandler;
|
||||
|
||||
/// <summary>Akka timer scheduler injected by the framework via <see cref="IWithTimers"/>.</summary>
|
||||
public ITimerScheduler Timers { get; set; } = null!;
|
||||
|
||||
@@ -158,20 +152,6 @@ public class SiteCommunicationActor : ReceiveActor, IWithTimers
|
||||
Receive<RetryParkedOperation>(cmd => DispatchCommand(cmd));
|
||||
Receive<DiscardParkedOperation>(cmd => DispatchCommand(cmd));
|
||||
|
||||
// Integration Routing — the 29th command, NOT migrated to the dispatcher (dead at
|
||||
// both ends; no production code registers the handler). Kept on the actor so the
|
||||
// dispatcher's command surface stays the 28 that actually migrate.
|
||||
Receive<IntegrationCallRequest>(msg =>
|
||||
{
|
||||
if (_integrationHandler != null)
|
||||
_integrationHandler.Forward(msg);
|
||||
else
|
||||
{
|
||||
Sender.Tell(new IntegrationCallResponse(
|
||||
msg.CorrelationId, _siteId, false, null, "Integration handler not available", DateTimeOffset.UtcNow));
|
||||
}
|
||||
});
|
||||
|
||||
// Central→site manual failover relay. Central and the site are separate clusters,
|
||||
// so central can only ask — this node performs the graceful Leave locally, scoped to
|
||||
// the SITE-SPECIFIC role, because that is what site singletons (the Deployment
|
||||
@@ -281,9 +261,6 @@ public class SiteCommunicationActor : ReceiveActor, IWithTimers
|
||||
case LocalHandlerType.ParkedMessages:
|
||||
_dispatcher.RegisterParkedMessageHandler(msg.Handler);
|
||||
break;
|
||||
case LocalHandlerType.Integration:
|
||||
_integrationHandler = msg.Handler;
|
||||
break;
|
||||
case LocalHandlerType.Artifacts:
|
||||
_dispatcher.RegisterArtifactHandler(msg.Handler);
|
||||
break;
|
||||
@@ -364,6 +341,5 @@ public enum LocalHandlerType
|
||||
{
|
||||
EventLog,
|
||||
ParkedMessages,
|
||||
Integration,
|
||||
Artifacts
|
||||
}
|
||||
|
||||
@@ -1,4 +1,5 @@
|
||||
using Akka.Cluster;
|
||||
using ZB.MOM.WW.Health.Akka;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Communication.ClusterState;
|
||||
|
||||
@@ -32,17 +33,13 @@ public static class ActiveNodeEvaluator
|
||||
/// <param name="cluster">The Akka cluster to evaluate.</param>
|
||||
/// <param name="role">Optional role scope; when set, only members with this role are considered.</param>
|
||||
/// <returns><c>true</c> when self is Up and the oldest Up member in the role scope.</returns>
|
||||
public static bool SelfIsOldestUp(Cluster cluster, string? role = null)
|
||||
{
|
||||
var self = cluster.SelfMember;
|
||||
if (self.Status != MemberStatus.Up)
|
||||
return false;
|
||||
if (role != null && !self.HasRole(role))
|
||||
return false;
|
||||
|
||||
return cluster.State.Members
|
||||
.Where(m => m.Status == MemberStatus.Up)
|
||||
.Where(m => role == null || m.HasRole(role))
|
||||
.All(m => m.UniqueAddress.Equals(self.UniqueAddress) || self.IsOlderThan(m));
|
||||
}
|
||||
/// <remarks>
|
||||
/// Delegates to the shared <see cref="ClusterActiveNode.SelfIsActive"/>
|
||||
/// (<c>ZB.MOM.WW.Health.Akka</c> 0.3.0), which is this evaluator promoted into the shared library
|
||||
/// after OtOpcUa independently needed the same rule and wrote its own third copy. The rule now has
|
||||
/// one implementation family-wide; this method survives as Communication's entry point into it, so
|
||||
/// the delivery gate and heartbeat keep their existing call sites and layering.
|
||||
/// </remarks>
|
||||
public static bool SelfIsOldestUp(Cluster cluster, string? role = null) =>
|
||||
ClusterActiveNode.SelfIsActive(cluster, role);
|
||||
}
|
||||
|
||||
@@ -224,21 +224,13 @@ public class CommunicationService
|
||||
}
|
||||
|
||||
// ── Pattern 4: Integration Routing ──
|
||||
|
||||
/// <summary>
|
||||
/// Routes an integration call to a site.
|
||||
/// </summary>
|
||||
/// <param name="siteId">The target site identifier.</param>
|
||||
/// <param name="request">The integration call request.</param>
|
||||
/// <param name="cancellationToken">Cancellation token.</param>
|
||||
/// <returns>The integration call response.</returns>
|
||||
public async Task<IntegrationCallResponse> RouteIntegrationCallAsync(
|
||||
string siteId, IntegrationCallRequest request, CancellationToken cancellationToken = default)
|
||||
{
|
||||
var envelope = new SiteEnvelope(siteId, request);
|
||||
return await GetActor().Ask<IntegrationCallResponse>(
|
||||
envelope, _options.IntegrationTimeout, cancellationToken);
|
||||
}
|
||||
// The brokered "External System → Central → Site → Central → External System"
|
||||
// round-trip (design §4) is served by the Inbound API's routed-site-script
|
||||
// path — the RouteTo* verbs below (RouteToCall / RouteToGetAttributes /
|
||||
// RouteToSetAttributes / RouteToWaitForAttribute), driven from
|
||||
// InboundAPI/CommunicationServiceInstanceRouter and sharing IntegrationTimeout.
|
||||
// The earlier generic IntegrationCallRequest primitive was never wired to a
|
||||
// producer or a site handler and was removed as dead code (Gitea #32).
|
||||
|
||||
// ── Pattern 5: Debug View ──
|
||||
|
||||
|
||||
@@ -34,13 +34,6 @@ namespace ZB.MOM.WW.ScadaBridge.Communication.Grpc;
|
||||
/// server (T1B.2) and the central transport (T1B.3) can share one implementation
|
||||
/// without a project-reference cycle.
|
||||
/// </para>
|
||||
/// <para>
|
||||
/// <b>Excluded by design:</b> <c>IntegrationCallRequest</c>. It is the 29th
|
||||
/// command on <c>SiteCommunicationActor</c>'s receive table but is dead at both
|
||||
/// ends — no production code registers the integration handler, so the site
|
||||
/// always answers "Integration handler not available". See
|
||||
/// <c>docs/known-issues/2026-07-22-integration-call-routing-is-dead-code.md</c>.
|
||||
/// </para>
|
||||
/// <para><b>Null conventions.</b> proto3 has no presence for scalars, so:</para>
|
||||
/// <list type="bullet">
|
||||
/// <item>
|
||||
|
||||
@@ -17,10 +17,6 @@ import "google/protobuf/wrappers.proto";
|
||||
// place on the client and one dispatch switch on the server. A `oneof` request
|
||||
// / reply envelope preserves per-command typing inside the group.
|
||||
//
|
||||
// IntegrationCallRequest (the 29th command) is deliberately NOT here — it is
|
||||
// dead code at both ends; see
|
||||
// docs/known-issues/2026-07-22-integration-call-routing-is-dead-code.md.
|
||||
//
|
||||
// EVOLUTION RULE (same as sitestream.proto): field numbers are never reused and
|
||||
// changes are additive only. New commands take the next free oneof tag; an
|
||||
// older peer simply sees an unset oneof and answers with a clean error.
|
||||
|
||||
@@ -28,6 +28,13 @@
|
||||
<PackageReference Include="Grpc.Net.Client" />
|
||||
<PackageReference Include="Grpc.Tools" PrivateAssets="All" />
|
||||
<PackageReference Include="ZB.MOM.WW.Configuration" />
|
||||
<!-- For ClusterActiveNode — the shared "oldest Up member" primitive that ActiveNodeEvaluator was
|
||||
promoted into. The store-and-forward delivery gate, the heartbeat IsActive stamp and the
|
||||
/health/active tier must all name the SAME active node; delegating to one implementation is
|
||||
what makes that structural. It lives in Health.Akka because that is the family's Akka-cluster
|
||||
utility package (it already hosts AkkaClusterStatusPolicy and an endpoint gate); a dedicated
|
||||
cluster package would be a better home if a third such primitive appears. -->
|
||||
<PackageReference Include="ZB.MOM.WW.Health.Akka" />
|
||||
</ItemGroup>
|
||||
|
||||
<ItemGroup>
|
||||
|
||||
@@ -12,7 +12,8 @@ namespace ZB.MOM.WW.ScadaBridge.DataConnectionLayer.Adapters;
|
||||
/// Maps IDataConnection methods to OPC UA concepts via IOpcUaClient abstraction.
|
||||
///
|
||||
/// OPC UA mapping:
|
||||
/// - TagPath → NodeId (e.g., "ns=2;s=MyDevice.Temperature")
|
||||
/// - TagPath → NodeId (durable form, e.g. "nsu=https://server/ns;s=MyDevice.Temperature";
|
||||
/// the legacy "ns=2;s=..." index form is still accepted — see OpcUaNodeReference)
|
||||
/// - Subscribe → MonitoredItem with DataChangeNotification
|
||||
/// - Read/Write → Read/Write service calls
|
||||
/// - Quality → OPC UA StatusCode mapping
|
||||
|
||||
@@ -123,6 +123,11 @@ public class RealOpcUaClient : IOpcUaClient
|
||||
endpoint = new EndpointDescription(endpointUrl);
|
||||
}
|
||||
|
||||
// The server advertises its own base address, which is frequently unroutable from here
|
||||
// (a 0.0.0.0 wildcard bind, or an internal container/NAT hostname). Swap it back to the
|
||||
// host the operator actually reached, or the session channel dials an unreachable address.
|
||||
RewriteEndpointHostForReachability(endpoint, endpointUrl);
|
||||
|
||||
var endpointConfig = EndpointConfiguration.Create(appConfig);
|
||||
var configuredEndpoint = new ConfiguredEndpoint(null, endpoint, endpointConfig);
|
||||
|
||||
@@ -157,6 +162,52 @@ public class RealOpcUaClient : IOpcUaClient
|
||||
await _subscription.CreateAsync(cancellationToken);
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Rewrites a discovered endpoint's host/port to match the reachable URL the operator
|
||||
/// configured, when the server advertised a different (unroutable) one.
|
||||
/// </summary>
|
||||
/// <remarks>
|
||||
/// An OPC UA server advertises its base address from its OWN configuration, not the route the
|
||||
/// client took to reach it — so <c>GetEndpoints</c> routinely returns a wildcard bind
|
||||
/// (<c>opc.tcp://0.0.0.0:4840/…</c>) or an internal container/NAT hostname the client cannot
|
||||
/// dial. The OPC Foundation session opens its channel to <see cref="EndpointDescription.EndpointUrl"/>
|
||||
/// verbatim, so a 0.0.0.0 advertisement resolves to the client's OWN loopback and the connect
|
||||
/// fails. This swaps only the authority (host + port) to the reachable one, preserving the
|
||||
/// advertised scheme and path (e.g. <c>/OtOpcUa</c>). Mirrors <c>CoreClientUtils.SelectEndpoint</c>.
|
||||
/// A no-op when the advertised host/port is already reachable (so an opc-plc server that
|
||||
/// advertises its real container name is left untouched).
|
||||
/// </remarks>
|
||||
internal static void RewriteEndpointHostForReachability(EndpointDescription? endpoint, string discoveryUrl)
|
||||
{
|
||||
if (endpoint is null || string.IsNullOrWhiteSpace(endpoint.EndpointUrl))
|
||||
return;
|
||||
if (!Uri.TryCreate(discoveryUrl, UriKind.Absolute, out var reachable))
|
||||
return;
|
||||
if (!Uri.TryCreate(endpoint.EndpointUrl, UriKind.Absolute, out var advertised))
|
||||
return;
|
||||
|
||||
// Already reachable — leave it exactly as advertised.
|
||||
if (string.Equals(advertised.Host, reachable.Host, StringComparison.OrdinalIgnoreCase)
|
||||
&& advertised.Port == reachable.Port)
|
||||
return;
|
||||
|
||||
// An IPv6 literal must keep its brackets, or the reconstructed authority (`::1:4840`) is
|
||||
// ambiguous and unparseable. Uri.Host may or may not include them depending on the scheme,
|
||||
// so bracket only when missing (never double-wrap).
|
||||
var host = reachable.Host;
|
||||
if (reachable.HostNameType == UriHostNameType.IPv6 && !host.StartsWith('['))
|
||||
host = $"[{host}]";
|
||||
// opc.tcp has no default port known to Uri, so Uri.Port is -1 when the text omits one.
|
||||
// Prefer the reachable URL's port, fall back to the advertised one, and emit no ":port"
|
||||
// at all when neither carries one (rather than a malformed ":-1").
|
||||
var port = reachable.Port >= 0 ? reachable.Port : advertised.Port;
|
||||
var authority = port >= 0 ? $"{host}:{port}" : host;
|
||||
// Authority-only advertisements parse to AbsolutePath "/" — drop it so we don't append a
|
||||
// spurious trailing slash the server never advertised.
|
||||
var path = advertised.AbsolutePath == "/" ? string.Empty : advertised.AbsolutePath;
|
||||
endpoint.EndpointUrl = $"{advertised.Scheme}://{authority}{path}";
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Probes an OPC UA endpoint configuration WITHOUT persisting it or creating a
|
||||
/// long-lived connection — connect, capture the server certificate if it is untrusted,
|
||||
@@ -258,6 +309,10 @@ public class RealOpcUaClient : IOpcUaClient
|
||||
endpoint = new EndpointDescription(endpointUrl);
|
||||
}
|
||||
|
||||
// Same reachability rewrite as ConnectAsync — a probe against a server advertising a
|
||||
// 0.0.0.0 / NAT base address must dial the reachable host, not the advertised one.
|
||||
RewriteEndpointHostForReachability(endpoint, endpointUrl);
|
||||
|
||||
var endpointConfig = EndpointConfiguration.Create(appConfig);
|
||||
var configuredEndpoint = new ConfiguredEndpoint(null, endpoint, endpointConfig);
|
||||
|
||||
|
||||
@@ -261,8 +261,16 @@ public class AkkaHostedService : IHostedService
|
||||
TimeSpan transportHeartbeat,
|
||||
TimeSpan transportFailure)
|
||||
{
|
||||
var seedNodesStr = string.Join(",",
|
||||
clusterOptions.SeedNodes.Select(QuoteHocon));
|
||||
// Simultaneous-cold-start split-brain guard (Gitea #33): when enabled the node must NOT
|
||||
// auto-join from its config seeds — Akka's FirstSeedNodeProcess is exactly what races the
|
||||
// peer and forms two 1-node clusters. Emit an EMPTY seed list so the ActorSystem comes up
|
||||
// unjoined, and ClusterBootstrapCoordinator issues the single reachability-gated
|
||||
// JoinSeedNodes (founder self-first / joiner peer-first). Default OFF, so guard-off configs
|
||||
// render byte-identical to before.
|
||||
var effectiveSeeds = clusterOptions.BootstrapGuard.Enabled
|
||||
? Enumerable.Empty<string>()
|
||||
: clusterOptions.SeedNodes;
|
||||
var seedNodesStr = string.Join(",", effectiveSeeds.Select(QuoteHocon));
|
||||
var rolesStr = string.Join(",", roles.Select(QuoteHocon));
|
||||
|
||||
// auto-down (default): AutoDowning provider — the leader among the reachable
|
||||
|
||||
@@ -0,0 +1,211 @@
|
||||
using System.Diagnostics;
|
||||
using System.Net.Sockets;
|
||||
using Akka.Actor;
|
||||
using Akka.Cluster;
|
||||
using Microsoft.Extensions.Options;
|
||||
using ZB.MOM.WW.ScadaBridge.ClusterInfrastructure;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Host.Actors;
|
||||
|
||||
/// <summary>
|
||||
/// Runs the simultaneous-cold-start split-brain guard (Gitea #33): when
|
||||
/// <c>ScadaBridge:Cluster:BootstrapGuard:Enabled</c>, the node starts with NO config seed nodes
|
||||
/// (<c>AkkaHostedService.BuildHocon</c> emits an empty seed list), and this coordinator picks the join
|
||||
/// order and issues the single <see cref="Akka.Cluster.Cluster.JoinSeedNodes"/> once the ActorSystem is
|
||||
/// up. See <see cref="ClusterBootstrapGuard"/> for the decision logic and the split-brain rationale.
|
||||
/// </summary>
|
||||
/// <remarks>
|
||||
/// <para>
|
||||
/// The founder (lower canonical address) joins its self-first order immediately. The higher
|
||||
/// node probes its partner's Akka endpoint (a plain TCP connect — the partner has bound its
|
||||
/// port well before it forms) up to <see cref="ClusterBootstrapGuardOptions.PartnerProbeSeconds"/>:
|
||||
/// reachable ⇒ peer-first (join the founder); unreachable ⇒ self-first (partner is dead, form
|
||||
/// alone). The join runs on a background task so it never blocks host startup, and it is issued
|
||||
/// exactly once — the coordinator never re-forms a node mid-handshake, the failure mode of the
|
||||
/// rejected self-form-watchdog design (see <c>SelfFirstSeedBootstrapTests</c>).
|
||||
/// </para>
|
||||
/// <para>
|
||||
/// Registered unconditionally in both the Central and Site composition roots but a no-op unless
|
||||
/// the guard is enabled, so guard-off deployments keep Akka's config-driven self-first
|
||||
/// auto-join byte-identical.
|
||||
/// </para>
|
||||
/// </remarks>
|
||||
public sealed class ClusterBootstrapCoordinator : IHostedService
|
||||
{
|
||||
private readonly Func<ActorSystem> _system;
|
||||
private readonly ClusterOptions _clusterOptions;
|
||||
private readonly NodeOptions _nodeOptions;
|
||||
private readonly ILogger<ClusterBootstrapCoordinator> _log;
|
||||
private readonly CancellationTokenSource _cts = new();
|
||||
private Task? _joinTask;
|
||||
|
||||
/// <summary>Creates the coordinator.</summary>
|
||||
/// <param name="system">Lazy accessor for the node's ActorSystem (resolved after the Akka hosted service creates it).</param>
|
||||
/// <param name="clusterOptions">The bound cluster options (seed list + guard settings).</param>
|
||||
/// <param name="nodeOptions">The bound node identity (advertised hostname + remoting port).</param>
|
||||
/// <param name="log">Logger for the bootstrap decision.</param>
|
||||
public ClusterBootstrapCoordinator(
|
||||
Func<ActorSystem> system,
|
||||
IOptions<ClusterOptions> clusterOptions,
|
||||
IOptions<NodeOptions> nodeOptions,
|
||||
ILogger<ClusterBootstrapCoordinator> log)
|
||||
{
|
||||
_system = system ?? throw new ArgumentNullException(nameof(system));
|
||||
_clusterOptions = clusterOptions?.Value ?? throw new ArgumentNullException(nameof(clusterOptions));
|
||||
_nodeOptions = nodeOptions?.Value ?? throw new ArgumentNullException(nameof(nodeOptions));
|
||||
_log = log ?? throw new ArgumentNullException(nameof(log));
|
||||
}
|
||||
|
||||
/// <inheritdoc/>
|
||||
public Task StartAsync(CancellationToken cancellationToken)
|
||||
{
|
||||
if (!_clusterOptions.BootstrapGuard.Enabled)
|
||||
return Task.CompletedTask; // dark switch off — Akka auto-joined from config seeds already.
|
||||
|
||||
// Fire-and-forget so a higher node with a dead partner (which waits the full probe window)
|
||||
// never blocks host startup. Exceptions are logged; a failed join leaves the node unjoined,
|
||||
// which is visible and recoverable, never a silent split.
|
||||
_joinTask = Task.Run(() => RunAsync(_cts.Token), _cts.Token);
|
||||
return Task.CompletedTask;
|
||||
}
|
||||
|
||||
/// <inheritdoc/>
|
||||
public async Task StopAsync(CancellationToken cancellationToken)
|
||||
{
|
||||
await _cts.CancelAsync().ConfigureAwait(false);
|
||||
if (_joinTask is not null)
|
||||
{
|
||||
try { await _joinTask.ConfigureAwait(false); }
|
||||
catch (OperationCanceledException) { /* expected on shutdown */ }
|
||||
}
|
||||
}
|
||||
|
||||
private async Task RunAsync(CancellationToken ct)
|
||||
{
|
||||
try
|
||||
{
|
||||
var seeds = _clusterOptions.SeedNodes ?? new List<string>();
|
||||
var role = ClusterBootstrapGuard.Analyze(seeds, _nodeOptions.NodeHostname, _nodeOptions.RemotingPort);
|
||||
|
||||
string[] order;
|
||||
var committedPeerFirst = false;
|
||||
if (!role.Applies)
|
||||
{
|
||||
// Not a pair seed (single-node install, self absent, malformed, or 3+ seeds): join the
|
||||
// configured seeds unchanged — the guard arbitrates only the 2-node pair race.
|
||||
order = seeds.ToArray();
|
||||
_log.LogInformation(
|
||||
"Bootstrap guard: not a pair seed ({SeedCount} seed(s)); joining configured seeds unchanged.",
|
||||
seeds.Count);
|
||||
}
|
||||
else if (role.IsFounder)
|
||||
{
|
||||
order = role.SelfFirstOrder;
|
||||
_log.LogInformation(
|
||||
"Bootstrap guard: this node is the preferred founder (lower address); joining self-first, forming immediately if no peer answers.");
|
||||
}
|
||||
else
|
||||
{
|
||||
var reachable = await ProbePartnerAsync(role.PartnerHost!, role.PartnerPort, ct).ConfigureAwait(false);
|
||||
order = ClusterBootstrapGuard.HigherNodeOrder(role, reachable);
|
||||
committedPeerFirst = reachable;
|
||||
_log.LogInformation(
|
||||
"Bootstrap guard: this node is the higher address; partner {PartnerHost}:{PartnerPort} was {Reachability} within the probe window; joining {Order}.",
|
||||
role.PartnerHost, role.PartnerPort, reachable ? "REACHABLE" : "UNREACHABLE",
|
||||
reachable ? "peer-first (join the founder)" : "self-first (partner down, forming alone)");
|
||||
}
|
||||
|
||||
if (order.Length == 0)
|
||||
{
|
||||
_log.LogWarning("Bootstrap guard: no seed nodes to join; node stays unjoined until a seed is configured.");
|
||||
return;
|
||||
}
|
||||
|
||||
var cluster = Akka.Cluster.Cluster.Get(_system());
|
||||
var addresses = order.Select(Address.Parse).ToArray();
|
||||
cluster.JoinSeedNodes(addresses);
|
||||
|
||||
// Peer-first is a one-way commitment: JoinSeedNodeProcess never self-forms. If the founder
|
||||
// died in the probe→join window, this node hangs unjoined forever. We do NOT re-form here
|
||||
// (the rejected mid-handshake failure mode) — we make the hang operator-visible so a
|
||||
// restart, which re-runs the guard against the now-dead founder and self-forms, recovers
|
||||
// it. Nothing is done on the founder / self-first paths, which always self-form on their own.
|
||||
if (committedPeerFirst)
|
||||
await WarnIfNotUpAsync(cluster, ct).ConfigureAwait(false);
|
||||
}
|
||||
catch (OperationCanceledException)
|
||||
{
|
||||
// Host shutting down before the join completed — nothing to do.
|
||||
}
|
||||
catch (Exception ex)
|
||||
{
|
||||
_log.LogError(ex, "Bootstrap guard: failed to issue JoinSeedNodes; node remains unjoined.");
|
||||
}
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// After committing peer-first, waits a bounded time for this node to reach <c>Up</c> and logs a
|
||||
/// clear warning if it does not — the founder must have died in the probe→join window, leaving this
|
||||
/// node hung in <c>JoinSeedNodeProcess</c>. Read-only: it never re-forms or re-joins (that is the
|
||||
/// rejected mid-handshake failure mode); it only surfaces the condition so an operator restarts the node.
|
||||
/// </summary>
|
||||
private async Task WarnIfNotUpAsync(Akka.Cluster.Cluster cluster, CancellationToken ct)
|
||||
{
|
||||
// Give the join a generous grace: the founder's own self-form (seed-node-timeout) plus this
|
||||
// node's join round-trip. Two probe windows is comfortably beyond both.
|
||||
var grace = TimeSpan.FromSeconds(Math.Max(10, _clusterOptions.BootstrapGuard.PartnerProbeSeconds * 2));
|
||||
var sw = Stopwatch.StartNew();
|
||||
while (!ct.IsCancellationRequested && sw.Elapsed < grace)
|
||||
{
|
||||
if (cluster.SelfMember.Status == MemberStatus.Up) return; // joined — all good
|
||||
try { await Task.Delay(TimeSpan.FromSeconds(1), ct).ConfigureAwait(false); }
|
||||
catch (OperationCanceledException) { return; }
|
||||
}
|
||||
|
||||
if (!ct.IsCancellationRequested && cluster.SelfMember.Status != MemberStatus.Up)
|
||||
_log.LogWarning(
|
||||
"Bootstrap guard: committed peer-first but this node is still {Status} after {Grace}s — its founder likely died in the probe→join window. This node will NOT self-form on its own (peer-first never does). RESTART it to recover: the guard will re-run, find the founder down, and form alone.",
|
||||
cluster.SelfMember.Status, (int)grace.TotalSeconds);
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Polls the partner's Akka endpoint with a plain TCP connect until it is reachable or the probe
|
||||
/// window elapses. Reachable-at-TCP is a sound proxy for "the partner is coming up": a node binds
|
||||
/// its Akka port near the start of startup, well before it forms or joins a cluster.
|
||||
/// </summary>
|
||||
private async Task<bool> ProbePartnerAsync(string host, int port, CancellationToken ct)
|
||||
{
|
||||
var window = TimeSpan.FromSeconds(Math.Max(0, _clusterOptions.BootstrapGuard.PartnerProbeSeconds));
|
||||
var interval = TimeSpan.FromMilliseconds(Math.Max(50, _clusterOptions.BootstrapGuard.PartnerProbeIntervalMs));
|
||||
var sw = Stopwatch.StartNew();
|
||||
|
||||
while (!ct.IsCancellationRequested)
|
||||
{
|
||||
if (await TryConnectAsync(host, port, ct).ConfigureAwait(false)) return true;
|
||||
if (sw.Elapsed >= window) return false;
|
||||
try { await Task.Delay(interval, ct).ConfigureAwait(false); }
|
||||
catch (OperationCanceledException) { return false; }
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
private async Task<bool> TryConnectAsync(string host, int port, CancellationToken ct)
|
||||
{
|
||||
try
|
||||
{
|
||||
using var client = new TcpClient();
|
||||
using var cts = CancellationTokenSource.CreateLinkedTokenSource(ct);
|
||||
cts.CancelAfter(Math.Max(100, _clusterOptions.BootstrapGuard.ProbeConnectTimeoutMs));
|
||||
await client.ConnectAsync(host, port, cts.Token).ConfigureAwait(false);
|
||||
return client.Connected;
|
||||
}
|
||||
catch (OperationCanceledException) when (ct.IsCancellationRequested)
|
||||
{
|
||||
throw; // host shutdown — propagate
|
||||
}
|
||||
catch
|
||||
{
|
||||
return false; // connect refused / timed out / DNS not resolvable yet — partner not up
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -14,7 +14,8 @@ namespace ZB.MOM.WW.ScadaBridge.Host.Health;
|
||||
/// <see cref="ClusterActivityEvaluator.SelfIsOldest(Cluster, string?)"/> — NOT
|
||||
/// the cluster leader, whose address-ordered definition diverges from singleton
|
||||
/// placement after a restart (review 01 [High]). This is the same evaluator
|
||||
/// backing <see cref="OldestNodeActiveHealthCheck"/>, so
|
||||
/// backing the shared <c>ActiveNodeHealthCheck</c> registered on the active tier
|
||||
/// — both route to <c>ClusterActiveNode</c> — so
|
||||
/// <see cref="InboundApiEndpointFilter"/> can return HTTP 503 on a standby.
|
||||
///
|
||||
/// Registered only in the Central-role branch of <c>Program.cs</c>. The gate
|
||||
|
||||
@@ -1,4 +1,5 @@
|
||||
using Akka.Cluster;
|
||||
using ZB.MOM.WW.Health.Akka;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Host.Health;
|
||||
|
||||
@@ -27,14 +28,10 @@ public static class ClusterActivityEvaluator
|
||||
/// <param name="cluster">The Akka cluster to evaluate.</param>
|
||||
/// <param name="role">Optional role scope; when set, only members with this role are considered.</param>
|
||||
/// <returns>The oldest Up member in the role scope, or null when none is Up.</returns>
|
||||
public static Member? OldestUpMember(Cluster cluster, string? role = null)
|
||||
{
|
||||
Member? oldest = null;
|
||||
foreach (var m in cluster.State.Members.Where(m => m.Status == MemberStatus.Up))
|
||||
{
|
||||
if (role != null && !m.HasRole(role)) continue;
|
||||
if (oldest == null || m.IsOlderThan(oldest)) oldest = m;
|
||||
}
|
||||
return oldest;
|
||||
}
|
||||
/// <remarks>
|
||||
/// Delegates to the shared <see cref="ClusterActiveNode.OldestUpMember(Cluster, string?)"/>, so
|
||||
/// the label a node shows and the answer <see cref="SelfIsOldest"/> gates on come from one rule.
|
||||
/// </remarks>
|
||||
public static Member? OldestUpMember(Cluster cluster, string? role = null) =>
|
||||
ClusterActiveNode.OldestUpMember(cluster, role);
|
||||
}
|
||||
|
||||
@@ -1,34 +0,0 @@
|
||||
using Akka.Actor;
|
||||
using Microsoft.Extensions.Diagnostics.HealthChecks;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Host.Health;
|
||||
|
||||
/// <summary>
|
||||
/// /health/active check: Healthy only on the oldest Up member — the node that
|
||||
/// hosts the cluster singletons. Replaces the shared package's leader-based
|
||||
/// ActiveNodeHealthCheck (review 01 [High]): leadership is address-ordered and
|
||||
/// diverges from singleton placement after a restart, and during a partition
|
||||
/// BOTH nodes compute themselves leader, so Traefik served both. Oldest-member
|
||||
/// semantics keep exactly one active node through a partition (the non-oldest
|
||||
/// side still sees the old oldest member until SBR resolves membership).
|
||||
/// </summary>
|
||||
public sealed class OldestNodeActiveHealthCheck : IHealthCheck
|
||||
{
|
||||
private readonly ActorSystem _system;
|
||||
|
||||
/// <summary>Initializes a new <see cref="OldestNodeActiveHealthCheck"/> bound to the running actor system.</summary>
|
||||
/// <param name="system">The live cluster actor system.</param>
|
||||
public OldestNodeActiveHealthCheck(ActorSystem system) => _system = system;
|
||||
|
||||
/// <summary>Reports Healthy only when this node is the oldest Up cluster member (the singleton host); otherwise reports Unhealthy.</summary>
|
||||
/// <param name="context">The health check context (unused; the result depends only on cluster membership).</param>
|
||||
/// <param name="ct">Cancellation token.</param>
|
||||
/// <returns>A task that resolves to the health check result.</returns>
|
||||
public Task<HealthCheckResult> CheckHealthAsync(HealthCheckContext context, CancellationToken ct = default)
|
||||
{
|
||||
var cluster = Akka.Cluster.Cluster.Get(_system);
|
||||
return Task.FromResult(ClusterActivityEvaluator.SelfIsOldest(cluster)
|
||||
? HealthCheckResult.Healthy("oldest Up member (singleton host)")
|
||||
: HealthCheckResult.Unhealthy("not the oldest Up member (standby)"));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,77 @@
|
||||
using Microsoft.Extensions.Diagnostics.HealthChecks;
|
||||
using ZB.MOM.WW.LocalDb;
|
||||
using ZB.MOM.WW.LocalDb.Replication;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Host.Health;
|
||||
|
||||
/// <summary>
|
||||
/// /health/ready check for the consolidated site database (LocalDb). Site nodes have no EF
|
||||
/// context — central's <c>DatabaseHealthCheck<ScadaBridgeDbContext></c> is central-only — so
|
||||
/// the site store is probed directly through the registered <see cref="ILocalDb"/> with a trivial
|
||||
/// <c>SELECT 1</c>. That is the honest readiness question for a site node: everything it persists
|
||||
/// (OperationTracking, site events, store-and-forward, the replicated config tables) lives here.
|
||||
/// </summary>
|
||||
/// <remarks>
|
||||
/// Replication state is reported as <c>data</c> ENRICHMENT ONLY and never affects the status.
|
||||
/// Replication is default-OFF by design, so a node with no peer reports <c>connected=false</c> with
|
||||
/// a real 0 backlog — the healthy steady state of an unreplicated node — and failing the check on
|
||||
/// it would mark every correctly-configured single node unready.
|
||||
/// </remarks>
|
||||
public sealed class SiteLocalDbHealthCheck : IHealthCheck
|
||||
{
|
||||
private readonly ILocalDb _localDb;
|
||||
private readonly ISyncStatus? _syncStatus;
|
||||
|
||||
/// <summary>Initializes a new <see cref="SiteLocalDbHealthCheck"/>.</summary>
|
||||
/// <param name="localDb">The consolidated site database registered by <c>AddZbLocalDb</c>.</param>
|
||||
/// <param name="syncStatus">
|
||||
/// The replication status view, when the replication engine is registered. Optional so the check
|
||||
/// still works in a composition that wires storage without replication.
|
||||
/// </param>
|
||||
public SiteLocalDbHealthCheck(ILocalDb localDb, ISyncStatus? syncStatus = null)
|
||||
{
|
||||
_localDb = localDb ?? throw new ArgumentNullException(nameof(localDb));
|
||||
_syncStatus = syncStatus;
|
||||
}
|
||||
|
||||
/// <summary>Probes the site database and reports its reachability.</summary>
|
||||
/// <param name="context">The health check context (unused; the verdict depends only on the probe).</param>
|
||||
/// <param name="cancellationToken">Cancellation token.</param>
|
||||
/// <returns>Healthy when the probe query succeeds; Unhealthy with the fault otherwise.</returns>
|
||||
public async Task<HealthCheckResult> CheckHealthAsync(
|
||||
HealthCheckContext context,
|
||||
CancellationToken cancellationToken = default)
|
||||
{
|
||||
try
|
||||
{
|
||||
await _localDb.QueryAsync("SELECT 1;", static r => r.GetInt32(0), ct: cancellationToken)
|
||||
.ConfigureAwait(false);
|
||||
}
|
||||
catch (Exception ex) when (ex is not OperationCanceledException)
|
||||
{
|
||||
return HealthCheckResult.Unhealthy("Site LocalDb is not reachable.", ex);
|
||||
}
|
||||
|
||||
return HealthCheckResult.Healthy("Site LocalDb reachable.", BuildData());
|
||||
}
|
||||
|
||||
private IReadOnlyDictionary<string, object> BuildData()
|
||||
{
|
||||
var data = new Dictionary<string, object>(StringComparer.Ordinal);
|
||||
if (_syncStatus is null)
|
||||
return data;
|
||||
|
||||
data["replicationConnected"] = _syncStatus.Connected;
|
||||
|
||||
// Nullable on purpose upstream: null means "backlog unknown" (no provider wired or the poll
|
||||
// failed), which must not be flattened to 0 — that would render a pair which cannot read its
|
||||
// own oplog as perfectly caught up. Omit the key instead.
|
||||
if (_syncStatus.OplogBacklog is { } backlog)
|
||||
data["replicationOplogBacklog"] = backlog;
|
||||
|
||||
if (_syncStatus.PeerNodeId is { } peer)
|
||||
data["replicationPeerNodeId"] = peer;
|
||||
|
||||
return data;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
using Microsoft.Extensions.Diagnostics.HealthChecks;
|
||||
using ZB.MOM.WW.ScadaBridge.HealthMonitoring;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Host.Health;
|
||||
|
||||
/// <summary>
|
||||
/// /health/active check for a SITE pair: Healthy only on the pair's primary — the oldest Up member
|
||||
/// carrying the <c>site-{SiteId}</c> role. Same active/standby semantics central publishes, so a
|
||||
/// consumer reads one contract fleet-wide (200 = Active, 503 = Standby).
|
||||
/// </summary>
|
||||
/// <remarks>
|
||||
/// Deliberately NOT central's unscoped shared <c>ActiveNodeHealthCheck</c>: that one calls
|
||||
/// <c>ClusterActivityEvaluator.SelfIsOldest(cluster)</c> with no role argument, so on a mesh that
|
||||
/// carries more than one site's members it computes "oldest" across the wrong member set and a site
|
||||
/// node could report Active because it happens to be the oldest member overall. Delegating to
|
||||
/// <see cref="IClusterNodeProvider.SelfIsPrimary"/> reuses the role-scoped evaluation the site
|
||||
/// already trusts for singleton placement and event-log purge — one primary definition per site,
|
||||
/// not two that can disagree.
|
||||
/// </remarks>
|
||||
public sealed class SitePairActiveNodeHealthCheck : IHealthCheck
|
||||
{
|
||||
private readonly IClusterNodeProvider _nodeProvider;
|
||||
|
||||
/// <summary>Initializes a new <see cref="SitePairActiveNodeHealthCheck"/>.</summary>
|
||||
/// <param name="nodeProvider">The role-scoped site cluster node provider.</param>
|
||||
public SitePairActiveNodeHealthCheck(IClusterNodeProvider nodeProvider) =>
|
||||
_nodeProvider = nodeProvider ?? throw new ArgumentNullException(nameof(nodeProvider));
|
||||
|
||||
/// <summary>Reports Healthy on the site pair's primary and Unhealthy on the standby.</summary>
|
||||
/// <param name="context">The health check context (unused; the verdict depends only on cluster membership).</param>
|
||||
/// <param name="cancellationToken">Cancellation token.</param>
|
||||
/// <returns>Healthy on the primary; Unhealthy (503 = Standby) otherwise.</returns>
|
||||
public Task<HealthCheckResult> CheckHealthAsync(
|
||||
HealthCheckContext context,
|
||||
CancellationToken cancellationToken = default) =>
|
||||
Task.FromResult(_nodeProvider.SelfIsPrimary
|
||||
? HealthCheckResult.Healthy("site pair primary (oldest Up member of the site role)")
|
||||
: HealthCheckResult.Unhealthy("not the site pair primary (standby)"));
|
||||
}
|
||||
@@ -344,9 +344,10 @@ try
|
||||
// Ready tier (/health/ready): database + akka-cluster + required-singletons.
|
||||
// Active tier (/health/active): active-node = oldest Up member (singleton host) AND
|
||||
// database reachable (review 01 [High]) — a DB-dead active node drops from Traefik
|
||||
// rotation, and the active-node check is the host-local OldestNodeActiveHealthCheck,
|
||||
// NOT the shared leader-based ActiveNodeHealthCheck (leadership diverges from singleton
|
||||
// placement after a restart and both sides claim leadership during a partition).
|
||||
// rotation. The active-node check is the shared ActiveNodeHealthCheck, which selects by age
|
||||
// as of Health 0.3.0; through 0.2.1 it selected by leadership, which diverges from singleton
|
||||
// placement after a restart and names both sides during a partition, so this host kept a
|
||||
// private copy until the shared one was fixed.
|
||||
// The Akka checks resolve ActorSystem from DI via the singleton bridge registered below;
|
||||
// the DatabaseHealthCheck<TContext> resolves a scoped ScadaBridgeDbContext (no factory).
|
||||
builder.Services.AddHealthChecks()
|
||||
@@ -371,9 +372,12 @@ try
|
||||
"required-singletons",
|
||||
failureStatus: null,
|
||||
tags: new[] { ZbHealthTags.Ready })
|
||||
// Host-local oldest-member check (review 01 [High]) — see OldestNodeActiveHealthCheck.
|
||||
// Resolves ActorSystem from DI, like the akka-cluster check above.
|
||||
.AddTypeActivatedCheck<OldestNodeActiveHealthCheck>(
|
||||
// Oldest-Up-member check (review 01 [High]), unscoped: every central member competes for
|
||||
// the one active slot. Since ZB.MOM.WW.Health 0.3.0 this IS the shared check — it selects
|
||||
// by age, so the host-local OldestNodeActiveHealthCheck that existed only to avoid the old
|
||||
// leader-based selection is gone. Resolves ActorSystem from DI, like the akka-cluster
|
||||
// check above.
|
||||
.AddTypeActivatedCheck<ActiveNodeHealthCheck>(
|
||||
"active-node",
|
||||
failureStatus: null,
|
||||
tags: new[] { ZbHealthTags.Active });
|
||||
@@ -394,14 +398,28 @@ try
|
||||
builder.Services.AddSingleton<Akka.Actor.ActorSystem>(sp =>
|
||||
sp.GetRequiredService<AkkaHostedService>().GetOrCreateActorSystem());
|
||||
|
||||
// Simultaneous-cold-start split-brain guard (Gitea #33, dark switch
|
||||
// ScadaBridge:Cluster:BootstrapGuard:Enabled, default off). Registered AFTER the
|
||||
// AkkaHostedService/ActorSystem bridge so its StartAsync resolves the system the main path
|
||||
// built; when the guard is off it no-ops (Akka auto-joined from config seeds), so registering
|
||||
// it unconditionally is safe. When on, the node started unjoined (empty HOCON seed list, see
|
||||
// AkkaHostedService.BuildHocon) and this coordinator issues the single reachability-gated
|
||||
// JoinSeedNodes. It resolves the ActorSystem lazily via GetOrCreateActorSystem so it never
|
||||
// races startup or caches a null.
|
||||
builder.Services.AddHostedService(sp => new ClusterBootstrapCoordinator(
|
||||
() => sp.GetRequiredService<AkkaHostedService>().GetOrCreateActorSystem(),
|
||||
sp.GetRequiredService<IOptions<ClusterOptions>>(),
|
||||
sp.GetRequiredService<IOptions<NodeOptions>>(),
|
||||
sp.GetRequiredService<ILogger<ClusterBootstrapCoordinator>>()));
|
||||
|
||||
// Register the production IActiveNodeGate implementation so
|
||||
// standby-node gating is actually enforced (the InboundApiEndpointFilter
|
||||
// consults IActiveNodeGate and defaults to "allow" when none is registered,
|
||||
// which leaves the design's "central cluster only (active node)" guarantee
|
||||
// unenforced in deployed binaries). The gate is backed by the same oldest-member
|
||||
// evaluator (ClusterActivityEvaluator) as OldestNodeActiveHealthCheck above, so the
|
||||
// inbound API and the /health/active endpoint Traefik routes against agree on
|
||||
// which node is active.
|
||||
// evaluator (ClusterActivityEvaluator, which delegates to the shared ClusterActiveNode) as
|
||||
// the active-node check above, so the inbound API and the /health/active endpoint Traefik
|
||||
// routes against agree on which node is active.
|
||||
builder.Services.AddSingleton<ZB.MOM.WW.ScadaBridge.InboundAPI.IActiveNodeGate, ActiveNodeGate>();
|
||||
|
||||
// Admin-triggered manual failover of the central pair (Health page control,
|
||||
@@ -634,6 +652,20 @@ try
|
||||
// must be present before the Map* calls on the site role.
|
||||
app.UseRouting();
|
||||
|
||||
// Map the canonical three-tier health endpoints on the site node:
|
||||
// /health/ready — akka-cluster (site-pair membership, and the cluster-view data the
|
||||
// family overview dashboard reads) + localdb (the consolidated site
|
||||
// store — the only database a site node has).
|
||||
// /health/active — active-node = this site pair's primary; 200 on the primary, 503 on
|
||||
// the standby, the same contract central publishes.
|
||||
// /healthz — bare process liveness.
|
||||
// Checks are registered in SiteServiceRegistration.Configure. All three endpoints are
|
||||
// anonymous and use the canonical ZbHealthWriter JSON; the site pipeline runs no
|
||||
// authentication middleware and has no FallbackPolicy, so nothing extra is needed to keep
|
||||
// them reachable. They are served on the HTTP/1.1 listener (metricsPort, default :8084)
|
||||
// alongside /metrics — the gRPC listener is HTTP/2-only and unusable by a plain HTTP client.
|
||||
app.MapZbHealth();
|
||||
|
||||
// Observability — mount the always-on Prometheus /metrics scrape endpoint.
|
||||
// AddZbTelemetry (in SiteServiceRegistration.Configure → BindSharedOptions)
|
||||
// wires the OTel Resource + standard instrumentation + Prometheus exporter;
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
using Microsoft.Extensions.DependencyInjection.Extensions;
|
||||
using Microsoft.Extensions.Options;
|
||||
using ZB.MOM.WW.Health;
|
||||
using ZB.MOM.WW.Health.Akka;
|
||||
using ZB.MOM.WW.ScadaBridge.AuditLog;
|
||||
using ZB.MOM.WW.ScadaBridge.ClusterInfrastructure;
|
||||
using ZB.MOM.WW.ScadaBridge.Communication;
|
||||
@@ -168,6 +170,19 @@ public static class SiteServiceRegistration
|
||||
services.AddSingleton<Akka.Actor.ActorSystem>(sp =>
|
||||
sp.GetRequiredService<AkkaHostedService>().GetOrCreateActorSystem());
|
||||
|
||||
// Simultaneous-cold-start split-brain guard (Gitea #33, dark switch
|
||||
// ScadaBridge:Cluster:BootstrapGuard:Enabled, default off). This is the case the guard was
|
||||
// built for — a shared power event that powers up both site VMs together. Registered AFTER
|
||||
// the AkkaHostedService/ActorSystem bridge (mirrors the Central composition root in
|
||||
// Program.cs); it no-ops when the guard is off, so registering it unconditionally is safe.
|
||||
// When on, the node started unjoined (empty HOCON seed list) and this coordinator issues the
|
||||
// single reachability-gated JoinSeedNodes, resolving the ActorSystem lazily.
|
||||
services.AddHostedService(sp => new ClusterBootstrapCoordinator(
|
||||
() => sp.GetRequiredService<AkkaHostedService>().GetOrCreateActorSystem(),
|
||||
sp.GetRequiredService<IOptions<ClusterOptions>>(),
|
||||
sp.GetRequiredService<IOptions<NodeOptions>>(),
|
||||
sp.GetRequiredService<ILogger<ClusterBootstrapCoordinator>>()));
|
||||
|
||||
// Cluster node status provider for health reports
|
||||
services.AddSingleton<IClusterNodeProvider>(sp =>
|
||||
{
|
||||
@@ -190,6 +205,39 @@ public static class SiteServiceRegistration
|
||||
return () => nodeProvider.SelfIsPrimary;
|
||||
});
|
||||
|
||||
// Health checks — the shared ZB.MOM.WW.Health probes, mapped by MapZbHealth on the site's
|
||||
// HTTP/1.1 listener (Program.cs). Site nodes served no health endpoints before this; the
|
||||
// family overview dashboard probes every node the same way, so a site node has to answer
|
||||
// the same two questions central does.
|
||||
//
|
||||
// Ready (/health/ready) akka-cluster — this node's site-pair membership. It also carries
|
||||
// the cluster-view data (leader/memberCount/...) that the
|
||||
// dashboard reads, free with ZB.MOM.WW.Health 0.2.0.
|
||||
// localdb — the consolidated site store, which is the only
|
||||
// database a site node has (central's EF DatabaseHealthCheck is
|
||||
// central-only and does not apply here).
|
||||
// Active (/health/active) active-node — site-pair primary vs standby.
|
||||
//
|
||||
// Registered HERE rather than in Program.cs, like the LocalDb registrations above, so the
|
||||
// composition-root tests that build this graph actually cover them. The Akka check resolves
|
||||
// ActorSystem from DI per probe via the singleton bridge registered above.
|
||||
services.AddHealthChecks()
|
||||
.AddTypeActivatedCheck<AkkaClusterHealthCheck>(
|
||||
"akka-cluster",
|
||||
failureStatus: null,
|
||||
tags: new[] { ZbHealthTags.Ready },
|
||||
args: AkkaClusterStatusPolicy.Default)
|
||||
.AddTypeActivatedCheck<SiteLocalDbHealthCheck>(
|
||||
"localdb",
|
||||
failureStatus: null,
|
||||
tags: new[] { ZbHealthTags.Ready })
|
||||
// NOT central's unscoped shared ActiveNodeHealthCheck — that one is role-less; see
|
||||
// SitePairActiveNodeHealthCheck for why a site pair needs the role-scoped answer.
|
||||
.AddTypeActivatedCheck<SitePairActiveNodeHealthCheck>(
|
||||
"active-node",
|
||||
failureStatus: null,
|
||||
tags: new[] { ZbHealthTags.Active });
|
||||
|
||||
// Options binding
|
||||
BindSharedOptions(services, config);
|
||||
|
||||
|
||||
@@ -17,7 +17,14 @@
|
||||
"StableAfter": "00:00:15",
|
||||
"HeartbeatInterval": "00:00:02",
|
||||
"FailureDetectionThreshold": "00:00:10",
|
||||
"MinNrOfMembers": 1
|
||||
"MinNrOfMembers": 1,
|
||||
"_bootstrapGuard": "Gitea #33 simultaneous-cold-start split-brain guard. Dark switch, default off (Akka's config-driven self-first auto-join, unchanged). Enable ONLY where both pair VMs can power up together (shared power/hypervisor event) with no start-order serialization: then the lower host:port founds self-first and the higher probes-then-joins, so the pair converges to one cluster instead of two 1-node clusters. Timings validated > 0 when enabled.",
|
||||
"BootstrapGuard": {
|
||||
"Enabled": false,
|
||||
"PartnerProbeSeconds": 25,
|
||||
"PartnerProbeIntervalMs": 500,
|
||||
"ProbeConnectTimeoutMs": 1000
|
||||
}
|
||||
},
|
||||
"_secrets": "Host-003: Secrets are NOT committed in this file. The ${secret:...} references below are resolved at startup by the pre-host secrets expander (before StartupValidator runs) from the encrypted secrets store (SQLite, seeded via the ZB.MOM.WW.Secrets store/CLI/UI). The KEK is supplied out-of-band via the ZB_SECRETS_MASTER_KEY environment variable and never committed; an unseeded reference fails closed (SecretNotFoundException) before any SQL/LDAP/cluster wiring. ROLLBACK: the whole-key environment override still wins — set ScadaBridge__Database__ConfigurationDb / ScadaBridge__Security__Ldap__ServiceAccountPassword / ScadaBridge__Security__JwtSigningKey and .AddEnvironmentVariables() overlays the concrete value before expansion runs, so the ${secret:...} token is never evaluated. NOTE (Task 1.4): the LDAP settings moved into the nested Security:Ldap sub-section (bound to the shared ZB.MOM.WW.Auth LdapOptions) — the service-account-password env var is ScadaBridge__Security__Ldap__ServiceAccountPassword (was ScadaBridge__Security__LdapServiceAccountPassword).",
|
||||
"Database": {
|
||||
|
||||
@@ -20,7 +20,14 @@
|
||||
"StableAfter": "00:00:15",
|
||||
"HeartbeatInterval": "00:00:02",
|
||||
"FailureDetectionThreshold": "00:00:10",
|
||||
"MinNrOfMembers": 1
|
||||
"MinNrOfMembers": 1,
|
||||
"_bootstrapGuard": "Gitea #33 simultaneous-cold-start split-brain guard. Dark switch, default off (Akka's config-driven self-first auto-join, unchanged). Enable ONLY where both pair VMs can power up together (shared power/hypervisor event) with no start-order serialization: then the lower host:port founds self-first and the higher probes-then-joins, so the pair converges to one cluster instead of two 1-node clusters. Timings validated > 0 when enabled.",
|
||||
"BootstrapGuard": {
|
||||
"Enabled": false,
|
||||
"PartnerProbeSeconds": 25,
|
||||
"PartnerProbeIntervalMs": 500,
|
||||
"ProbeConnectTimeoutMs": 1000
|
||||
}
|
||||
},
|
||||
"Database": {
|
||||
// Migration-only as of LocalDb Phase 2. The site config tables now live in the
|
||||
|
||||
@@ -1272,265 +1272,96 @@ public sealed class BundleImporter : IBundleImporter
|
||||
}
|
||||
|
||||
var bundleImportId = Guid.NewGuid();
|
||||
var resolutionMap = resolutions.ToDictionary(
|
||||
r => (r.EntityType, r.Name),
|
||||
r => r);
|
||||
var summary = new ImportSummary();
|
||||
|
||||
// Inbound API keys are not transported between environments.
|
||||
// A legacy bundle may still contain a keys section; we ignore those keys
|
||||
// entirely (never re-create them) but count them so the result can tell
|
||||
// the operator to re-issue keys on this environment.
|
||||
var apiKeysIgnored = content.ApiKeys?.Count ?? 0;
|
||||
|
||||
// Set the correlation BEFORE the transaction so any audit writes
|
||||
// triggered during the apply pick up the BundleImportId — AuditService
|
||||
// reads the scoped context at the moment LogAsync is called.
|
||||
_correlationContext.BundleImportId = bundleImportId;
|
||||
|
||||
// BeginTransactionAsync is a no-op on the in-memory EF provider (which
|
||||
// logs an InMemoryEventId.TransactionIgnoredWarning by default). To keep
|
||||
// rollback semantics testable on in-memory AND correct on relational
|
||||
// providers, we defer the SINGLE SaveChangesAsync call until just before
|
||||
// CommitAsync — every Add*Async + LogAsync call only stages on the
|
||||
// change tracker, so throwing before SaveChangesAsync naturally undoes
|
||||
// the entire apply on both providers.
|
||||
await using var tx = await _dbContext.Database.BeginTransactionAsync(ct).ConfigureAwait(false);
|
||||
// The central ConfigurationDb context is configured with
|
||||
// EnableRetryOnFailure (ServiceCollectionExtensions), and
|
||||
// SqlServerRetryingExecutionStrategy REJECTS a user-initiated
|
||||
// transaction opened outside the strategy ("The configured execution
|
||||
// strategy 'SqlServerRetryingExecutionStrategy' does not support
|
||||
// user-initiated transactions."), so opening BeginTransactionAsync
|
||||
// directly here throws on any real SQL Server. Run the whole
|
||||
// BeginTransaction -> apply -> Commit unit INSIDE the execution strategy
|
||||
// so the strategy owns it and re-executes it as one retriable unit on a
|
||||
// transient fault. On the in-memory provider CreateExecutionStrategy()
|
||||
// returns a non-retrying strategy, so this stays a single straight-
|
||||
// through call there (behaviour unchanged).
|
||||
//
|
||||
// Retry idempotency: the strategy re-runs the WHOLE delegate on a
|
||||
// transient fault, so the change tracker is cleared at the top of the
|
||||
// delegate (adds staged by the failed attempt, whose transaction has
|
||||
// already rolled back, must not leak into the retry) and ApplyMergeAsync
|
||||
// rebuilds its own summary/resolution accumulators per call — otherwise
|
||||
// a retry would double-apply. A rollback failure captured inside the
|
||||
// delegate is surfaced to the outer catch via rollbackFailureMessage so
|
||||
// the BundleImportFailed row still records both faults.
|
||||
var strategy = _dbContext.Database.CreateExecutionStrategy();
|
||||
string? rollbackFailureMessage = null;
|
||||
ImportResult result;
|
||||
try
|
||||
{
|
||||
// Run semantic validation FIRST — before any writes are staged.
|
||||
// This is purely a name-resolution scan over the in-memory DTOs +
|
||||
// pre-existing target reads, so it has no ordering dependency on
|
||||
// the Apply* helpers. Failing here means the change tracker is
|
||||
// still empty, which keeps the rollback contract simple on both
|
||||
// the in-memory and relational providers (intermediate
|
||||
// SaveChangesAsync between Apply* and the second-pass rewire
|
||||
// would otherwise prevent the in-memory provider from undoing
|
||||
// already-flushed template rows on validation failure — the
|
||||
// in-memory transaction is a no-op, so the only safe pattern is
|
||||
// ONE deferred SaveChangesAsync at the very end of try-block).
|
||||
//
|
||||
// Skip-resolved DTOs are excluded from the in-bundle name set so
|
||||
// a Skip on a dependency surfaces as a missing-reference error
|
||||
// rather than silently passing.
|
||||
var (validationErrors, validationWarnings) =
|
||||
await RunSemanticValidationAsync(content, resolutionMap, ct).ConfigureAwait(false);
|
||||
// Validate every site / connection / template reference the
|
||||
// site-instance payload depends on BEFORE any row is staged. Running
|
||||
// this in the validation phase (not as an apply-pass guard) preserves
|
||||
// the rollback contract: a structurally-unresolvable bundle fails with
|
||||
// an empty change tracker, so nothing is half-written on the in-memory
|
||||
// provider (where the intermediate site/connection flush can't be
|
||||
// undone by ChangeTracker.Clear). The apply passes keep defensive
|
||||
// guards, but in normal operation those never fire.
|
||||
validationErrors = validationErrors.Count > 0
|
||||
? validationErrors
|
||||
: await ValidateSiteInstanceReferencesAsync(content, resolutionMap, nameMap, ct).ConfigureAwait(false);
|
||||
if (validationErrors.Count > 0)
|
||||
result = await strategy.ExecuteAsync(async () =>
|
||||
{
|
||||
throw new SemanticValidationException(validationErrors);
|
||||
}
|
||||
// Reset per-attempt state so a retried delegate re-runs clean.
|
||||
_dbContext.ChangeTracker.Clear();
|
||||
rollbackFailureMessage = null;
|
||||
|
||||
// Advisory template-script reference findings never block the import;
|
||||
// log them and surface them on the ImportResult so the operator sees
|
||||
// the same advisory the preview showed. The deploy-time gate is the
|
||||
// authoritative re-validation.
|
||||
if (validationWarnings.Count > 0)
|
||||
{
|
||||
_logger?.LogWarning(
|
||||
"Bundle import {BundleImportId}: {Count} advisory import warning(s): {Warnings}",
|
||||
bundleImportId, validationWarnings.Count, string.Join("; ", validationWarnings));
|
||||
}
|
||||
|
||||
// ---- Site/instance-scoped apply: sites + connections FIRST ----
|
||||
// Sites are the FK target for both data connections (SiteId) and
|
||||
// instances (SiteId); data connections are the FK target for instance
|
||||
// connection bindings (DataConnectionId) and the rewrite target for
|
||||
// native-alarm-source ConnectionNameOverride. Resolve-or-create both
|
||||
// BEFORE the central-config apply so the maps the instance pass needs
|
||||
// are fully populated, then flush so newly-created site/connection ids
|
||||
// materialise on the relational provider. (The instance pass itself
|
||||
// runs AFTER the central-config flush so it can also resolve template
|
||||
// ids by name.)
|
||||
var siteBySourceIdentifier = await ApplySitesAsync(
|
||||
content, nameMap, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
var connectionMaps = await ApplyDataConnectionsAsync(
|
||||
content, nameMap, siteBySourceIdentifier, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
// Flush so site + connection surrogate ids are assigned (relational
|
||||
// provider) before the instance pass wires its FKs. In-memory assigns
|
||||
// ids on AddAsync, so this is mostly a no-op there, but it keeps the
|
||||
// ordering correct on a real DB. Rides the same outer transaction.
|
||||
await _dbContext.SaveChangesAsync(ct).ConfigureAwait(false);
|
||||
|
||||
await ApplyTemplateFoldersAsync(content.TemplateFolders, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplyTemplatesAsync(content.Templates, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplySharedScriptsAsync(content.SharedScripts, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplyExternalSystemsAsync(content.ExternalSystems, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplyDatabaseConnectionsAsync(content.DatabaseConnections, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplyNotificationListsAsync(content.NotificationLists, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplySmtpConfigsAsync(content.SmtpConfigs, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplySmsConfigsAsync(content.SmsConfigs, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
// Inbound API keys are NOT applied from a bundle — any keys
|
||||
// in a legacy bundle were counted above (apiKeysIgnored) and are skipped.
|
||||
await ApplyApiMethodsAsync(content.ApiMethods, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
|
||||
// Second-pass rewire of name-keyed
|
||||
// FKs that can only be resolved AFTER every template's scripts and
|
||||
// child rows have been staged. We flush here so Pass A can look up
|
||||
// the just-persisted scripts by name and Pass B can look up the
|
||||
// just-persisted templates by name; both passes stage further
|
||||
// mutations that ride the SAME outer transaction (committed below).
|
||||
//
|
||||
// EF tracking note: AddAsync on the in-memory provider assigns
|
||||
// synthetic ids eagerly, so this intermediate flush mostly
|
||||
// materialises identity values on a relational provider. The
|
||||
// rollback contract is preserved because semantic validation
|
||||
// already ran above — any throw from this point onward represents
|
||||
// a successful merge that the user wants to keep, and the only
|
||||
// remaining failure surface is the final BundleImported audit
|
||||
// write itself.
|
||||
await _dbContext.SaveChangesAsync(ct).ConfigureAwait(false);
|
||||
await ResolveAlarmScriptLinksAsync(content.Templates, resolutionMap, user, ct).ConfigureAwait(false);
|
||||
await ResolveCompositionEdgesAsync(content.Templates, resolutionMap, user, ct).ConfigureAwait(false);
|
||||
await ResolveInheritanceEdgesAsync(content.Templates, resolutionMap, user, ct).ConfigureAwait(false);
|
||||
// Folder parent edges are name-keyed too — resolve them AFTER the flush
|
||||
// above so every imported folder has a materialised id to point at
|
||||
// (the Add pass created the folders parent-less). Mirrors the template
|
||||
// inheritance/composition rewires.
|
||||
await ResolveFolderParentEdgesAsync(content.TemplateFolders, resolutionMap, user, ct).ConfigureAwait(false);
|
||||
// External-system methods have no navigation on the parent definition,
|
||||
// so an Add can't cascade them — persist them here, once the parent
|
||||
// system has a real id from the flush above. (Overwrite already syncs
|
||||
// its methods inline against the existing parent id.)
|
||||
await PersistAddedExternalSystemMethodsAsync(content.ExternalSystems, resolutionMap, user, ct).ConfigureAwait(false);
|
||||
|
||||
// ---- Import-time acyclicity guard over the merged template graph ----
|
||||
// The two rewire passes above have STAGED (not yet saved) the bundle's
|
||||
// inheritance (ParentTemplateId) and composition edges onto the tracked
|
||||
// target templates. A bundle can still close a loop through a pre-existing
|
||||
// target template — e.g. Overwrite a base so it now inherits from (or
|
||||
// composes) one of its own descendants. The design-time save-gate enforces
|
||||
// acyclicity on individual edits, but an import merges many edges at once,
|
||||
// so re-check the whole merged graph here. On a cycle we throw and the
|
||||
// catch below rolls the whole import back (ChangeTracker.Clear on the
|
||||
// in-memory provider) BEFORE any of it is persisted; deliberately no
|
||||
// SaveChanges precedes this check so nothing survives that rollback.
|
||||
EnsureNoTemplateGraphCycles();
|
||||
|
||||
// ---- Site/instance-scoped apply: instances LAST ----
|
||||
// The instance pass needs every FK target materialised first:
|
||||
// • template id — resolved by name against the just-flushed central
|
||||
// config (Templates were flushed above for the alarm/composition
|
||||
// rewire), plus pre-existing target templates;
|
||||
// • site id — from the site map built by ApplySitesAsync;
|
||||
// • connection id — from the connection map built by
|
||||
// ApplyDataConnectionsAsync.
|
||||
// It writes the instance row + its four child collections (attribute /
|
||||
// alarm / native-alarm-source overrides + connection bindings) and
|
||||
// rewires every name-keyed FK to the target environment's ids.
|
||||
await ApplyInstancesAsync(
|
||||
content, resolutionMap, siteBySourceIdentifier, connectionMaps,
|
||||
user, summary, ct).ConfigureAwait(false);
|
||||
|
||||
// ---- Pre-commit point: stale-instance computation ----
|
||||
// Everything the import will write is staged on the change tracker at
|
||||
// this point but NOT yet committed. This computes StaleInstanceIds here
|
||||
// — target instances whose template was overwritten by this import and
|
||||
// therefore now carry a stale flattened-config revision — and threads
|
||||
// the resulting list into the ImportResult below (replacing the
|
||||
// Array.Empty<int>() placeholder). The single deferred SaveChangesAsync
|
||||
// is the next statement, so a read against the change tracker here sees
|
||||
// the full post-apply graph before commit.
|
||||
var staleInstanceIds = await ComputeStaleInstanceIdsAsync(
|
||||
content, resolutionMap, ct).ConfigureAwait(false);
|
||||
|
||||
await _auditService.LogAsync(
|
||||
user: user,
|
||||
action: "BundleImported",
|
||||
entityType: "Bundle",
|
||||
entityId: bundleImportId.ToString(),
|
||||
entityName: session.Manifest.SourceEnvironment,
|
||||
afterState: new
|
||||
// BeginTransactionAsync is a no-op on the in-memory EF provider
|
||||
// (which logs an InMemoryEventId.TransactionIgnoredWarning by
|
||||
// default). To keep rollback semantics testable on in-memory AND
|
||||
// correct on relational providers, ApplyMergeAsync defers the
|
||||
// SINGLE final SaveChangesAsync until the very end — every
|
||||
// Add*Async + LogAsync call only stages on the change tracker,
|
||||
// so throwing before that naturally undoes the whole apply on
|
||||
// both providers.
|
||||
var tx = await _dbContext.Database.BeginTransactionAsync(ct).ConfigureAwait(false);
|
||||
try
|
||||
{
|
||||
BundleImportId = bundleImportId,
|
||||
session.Manifest.SourceEnvironment,
|
||||
session.Manifest.ContentHash,
|
||||
Summary = new
|
||||
// ApplyMergeAsync stages the whole merge and issues the
|
||||
// single deferred final SaveChangesAsync itself; here we only
|
||||
// commit the transaction it ran under.
|
||||
var applied = await ApplyMergeAsync(
|
||||
content, resolutions, nameMap, session, bundleImportId, user, ct).ConfigureAwait(false);
|
||||
await tx.CommitAsync(ct).ConfigureAwait(false);
|
||||
await tx.DisposeAsync().ConfigureAwait(false);
|
||||
return applied;
|
||||
}
|
||||
catch
|
||||
{
|
||||
// Roll back explicitly (rather than leaning on Dispose) so a
|
||||
// rollback that itself throws does not mask the ORIGINAL
|
||||
// exception — capture it and let the original propagate to
|
||||
// the strategy, which decides whether to retry.
|
||||
try
|
||||
{
|
||||
summary.Added,
|
||||
summary.Overwritten,
|
||||
summary.Skipped,
|
||||
summary.Renamed,
|
||||
},
|
||||
// Legacy inbound API keys present in the bundle that
|
||||
// were ignored (not transported / not re-created here).
|
||||
ApiKeysIgnored = apiKeysIgnored,
|
||||
},
|
||||
cancellationToken: ct).ConfigureAwait(false);
|
||||
await tx.RollbackAsync(ct).ConfigureAwait(false);
|
||||
}
|
||||
catch (Exception rbEx)
|
||||
{
|
||||
rollbackFailureMessage = rbEx.Message;
|
||||
}
|
||||
try { await tx.DisposeAsync().ConfigureAwait(false); }
|
||||
catch { /* dispose-after-throw must not mask the original cause */ }
|
||||
|
||||
await _dbContext.SaveChangesAsync(ct).ConfigureAwait(false);
|
||||
await tx.CommitAsync(ct).ConfigureAwait(false);
|
||||
|
||||
PublishScriptArtifactChanges(resolutions);
|
||||
|
||||
// T-007: zero out the decrypted plaintext BEFORE remove so any
|
||||
// caller-held reference (e.g. the Razor page that built the
|
||||
// ImportPreview) sees the cleared buffer too. Remove drops the
|
||||
// dictionary entry; together they release the secrets immediately
|
||||
// instead of leaving them in process memory for the full TTL.
|
||||
ZeroDecryptedContent(session);
|
||||
_sessionStore.Remove(sessionId);
|
||||
|
||||
return new ImportResult(
|
||||
BundleImportId: bundleImportId,
|
||||
Added: summary.Added,
|
||||
Overwritten: summary.Overwritten,
|
||||
Skipped: summary.Skipped,
|
||||
Renamed: summary.Renamed,
|
||||
StaleInstanceIds: staleInstanceIds,
|
||||
AuditEventCorrelation: bundleImportId.ToString(),
|
||||
ApiKeysIgnored: apiKeysIgnored,
|
||||
Warnings: validationWarnings);
|
||||
// Clear the change tracker — on the in-memory provider the
|
||||
// rollback is a no-op and the staged adds would otherwise
|
||||
// persist on the next SaveChangesAsync (a retry or the
|
||||
// failure-row write below).
|
||||
_dbContext.ChangeTracker.Clear();
|
||||
throw;
|
||||
}
|
||||
}).ConfigureAwait(false);
|
||||
}
|
||||
catch (Exception ex)
|
||||
{
|
||||
// Rollback can itself throw (connection drop mid-rollback, provider
|
||||
// bug, etc). If it does, we must STILL write the BundleImportFailed
|
||||
// audit row — otherwise a rollback-failure path silently swallows
|
||||
// the import's audit trail. Capture the rollback exception (if any)
|
||||
// and surface it on the failure row alongside the original cause.
|
||||
Exception? rollbackFailure = null;
|
||||
try
|
||||
{
|
||||
await tx.RollbackAsync(ct).ConfigureAwait(false);
|
||||
}
|
||||
catch (Exception rbEx)
|
||||
{
|
||||
rollbackFailure = rbEx;
|
||||
}
|
||||
|
||||
// If rollback threw the IDbContextTransaction is in an indeterminate
|
||||
// state and still associated with the DbContext — a subsequent
|
||||
// SaveChangesAsync would attempt to enlist in (or commit to) that
|
||||
// broken transaction, and the failure-row would itself be rolled
|
||||
// back when the transaction is finally disposed. Dispose it now so
|
||||
// the audit-row write below uses a fresh implicit transaction. On
|
||||
// the happy rollback path Dispose is a benign no-op (the using
|
||||
// would call it on scope exit anyway).
|
||||
if (rollbackFailure is not null)
|
||||
{
|
||||
try { await tx.DisposeAsync().ConfigureAwait(false); }
|
||||
catch { /* dispose-after-throw must not mask the original cause */ }
|
||||
}
|
||||
|
||||
// Clear the change tracker before writing the failure row — on the
|
||||
// in-memory provider the rollback is a no-op and the staged adds
|
||||
// would otherwise persist when the next SaveChangesAsync runs. This
|
||||
// also matters when rollback threw: the change tracker is in an
|
||||
// ambiguous state and we don't want the failure-write to sweep up
|
||||
// any of the staged apply mutations.
|
||||
_dbContext.ChangeTracker.Clear();
|
||||
|
||||
// The strategy exhausted its retries (or threw a non-transient
|
||||
// error). The failed attempt already rolled back and cleared the
|
||||
// change tracker inside the delegate, so all that remains here is
|
||||
// the top-level BundleImportFailed audit row.
|
||||
//
|
||||
// Clear correlation FIRST so the failure row doesn't carry the now-
|
||||
// rolled-back BundleImportId. The contract is: BundleImportFailed
|
||||
// exists at top level (no correlation) so audit consumers can see
|
||||
@@ -1549,7 +1380,7 @@ public sealed class BundleImporter : IBundleImporter
|
||||
BundleImportId = bundleImportId,
|
||||
Reason = ex.Message,
|
||||
ExceptionType = ex.GetType().FullName,
|
||||
RollbackException = rollbackFailure?.Message,
|
||||
RollbackException = rollbackFailureMessage,
|
||||
},
|
||||
cancellationToken: ct).ConfigureAwait(false);
|
||||
await _dbContext.SaveChangesAsync(ct).ConfigureAwait(false);
|
||||
@@ -1580,6 +1411,236 @@ public sealed class BundleImporter : IBundleImporter
|
||||
// not inherit the import id.
|
||||
_correlationContext.BundleImportId = null;
|
||||
}
|
||||
|
||||
// Post-commit side effects — run ONCE, only after a durable commit, so a
|
||||
// retried attempt neither publishes twice nor releases the session early.
|
||||
PublishScriptArtifactChanges(resolutions);
|
||||
|
||||
// T-007: zero out the decrypted plaintext BEFORE remove so any
|
||||
// caller-held reference (e.g. the Razor page that built the
|
||||
// ImportPreview) sees the cleared buffer too. Remove drops the
|
||||
// dictionary entry; together they release the secrets immediately
|
||||
// instead of leaving them in process memory for the full TTL.
|
||||
ZeroDecryptedContent(session);
|
||||
_sessionStore.Remove(sessionId);
|
||||
|
||||
return result;
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Applies one bundle's merge inside the caller's transaction: semantic
|
||||
/// validation, every Apply* pass, the name-keyed FK rewires, the stale-
|
||||
/// instance computation and the BundleImported audit row, ending with the
|
||||
/// deferred final SaveChangesAsync staged by the caller. Does NOT begin or
|
||||
/// commit the transaction — <see cref="ApplyAsync"/> owns it so the whole
|
||||
/// unit can run inside the context's execution strategy (EnableRetryOnFailure
|
||||
/// rejects user-initiated transactions opened outside the strategy). Safe to
|
||||
/// re-run: the caller clears the change tracker before each attempt, and
|
||||
/// every accumulator here (summary counts, resolution map) is rebuilt per
|
||||
/// call.
|
||||
/// </summary>
|
||||
private async Task<ImportResult> ApplyMergeAsync(
|
||||
BundleContentDto content,
|
||||
IReadOnlyList<ImportResolution> resolutions,
|
||||
BundleNameMap nameMap,
|
||||
BundleSession session,
|
||||
Guid bundleImportId,
|
||||
string user,
|
||||
CancellationToken ct)
|
||||
{
|
||||
var resolutionMap = resolutions.ToDictionary(
|
||||
r => (r.EntityType, r.Name),
|
||||
r => r);
|
||||
var summary = new ImportSummary();
|
||||
|
||||
// Inbound API keys are not transported between environments.
|
||||
// A legacy bundle may still contain a keys section; we ignore those keys
|
||||
// entirely (never re-create them) but count them so the result can tell
|
||||
// the operator to re-issue keys on this environment.
|
||||
var apiKeysIgnored = content.ApiKeys?.Count ?? 0;
|
||||
|
||||
// Run semantic validation FIRST — before any writes are staged.
|
||||
// This is purely a name-resolution scan over the in-memory DTOs +
|
||||
// pre-existing target reads, so it has no ordering dependency on
|
||||
// the Apply* helpers. Failing here means the change tracker is
|
||||
// still empty, which keeps the rollback contract simple on both
|
||||
// the in-memory and relational providers (intermediate
|
||||
// SaveChangesAsync between Apply* and the second-pass rewire
|
||||
// would otherwise prevent the in-memory provider from undoing
|
||||
// already-flushed template rows on validation failure — the
|
||||
// in-memory transaction is a no-op, so the only safe pattern is
|
||||
// ONE deferred SaveChangesAsync at the very end of try-block).
|
||||
//
|
||||
// Skip-resolved DTOs are excluded from the in-bundle name set so
|
||||
// a Skip on a dependency surfaces as a missing-reference error
|
||||
// rather than silently passing.
|
||||
var (validationErrors, validationWarnings) =
|
||||
await RunSemanticValidationAsync(content, resolutionMap, ct).ConfigureAwait(false);
|
||||
// Validate every site / connection / template reference the
|
||||
// site-instance payload depends on BEFORE any row is staged. Running
|
||||
// this in the validation phase (not as an apply-pass guard) preserves
|
||||
// the rollback contract: a structurally-unresolvable bundle fails with
|
||||
// an empty change tracker, so nothing is half-written on the in-memory
|
||||
// provider (where the intermediate site/connection flush can't be
|
||||
// undone by ChangeTracker.Clear). The apply passes keep defensive
|
||||
// guards, but in normal operation those never fire.
|
||||
validationErrors = validationErrors.Count > 0
|
||||
? validationErrors
|
||||
: await ValidateSiteInstanceReferencesAsync(content, resolutionMap, nameMap, ct).ConfigureAwait(false);
|
||||
if (validationErrors.Count > 0)
|
||||
{
|
||||
throw new SemanticValidationException(validationErrors);
|
||||
}
|
||||
|
||||
// Advisory template-script reference findings never block the import;
|
||||
// log them and surface them on the ImportResult so the operator sees
|
||||
// the same advisory the preview showed. The deploy-time gate is the
|
||||
// authoritative re-validation.
|
||||
if (validationWarnings.Count > 0)
|
||||
{
|
||||
_logger?.LogWarning(
|
||||
"Bundle import {BundleImportId}: {Count} advisory import warning(s): {Warnings}",
|
||||
bundleImportId, validationWarnings.Count, string.Join("; ", validationWarnings));
|
||||
}
|
||||
|
||||
// ---- Site/instance-scoped apply: sites + connections FIRST ----
|
||||
// Sites are the FK target for both data connections (SiteId) and
|
||||
// instances (SiteId); data connections are the FK target for instance
|
||||
// connection bindings (DataConnectionId) and the rewrite target for
|
||||
// native-alarm-source ConnectionNameOverride. Resolve-or-create both
|
||||
// BEFORE the central-config apply so the maps the instance pass needs
|
||||
// are fully populated, then flush so newly-created site/connection ids
|
||||
// materialise on the relational provider. (The instance pass itself
|
||||
// runs AFTER the central-config flush so it can also resolve template
|
||||
// ids by name.)
|
||||
var siteBySourceIdentifier = await ApplySitesAsync(
|
||||
content, nameMap, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
var connectionMaps = await ApplyDataConnectionsAsync(
|
||||
content, nameMap, siteBySourceIdentifier, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
// Flush so site + connection surrogate ids are assigned (relational
|
||||
// provider) before the instance pass wires its FKs. In-memory assigns
|
||||
// ids on AddAsync, so this is mostly a no-op there, but it keeps the
|
||||
// ordering correct on a real DB. Rides the same outer transaction.
|
||||
await _dbContext.SaveChangesAsync(ct).ConfigureAwait(false);
|
||||
|
||||
await ApplyTemplateFoldersAsync(content.TemplateFolders, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplyTemplatesAsync(content.Templates, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplySharedScriptsAsync(content.SharedScripts, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplyExternalSystemsAsync(content.ExternalSystems, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplyDatabaseConnectionsAsync(content.DatabaseConnections, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplyNotificationListsAsync(content.NotificationLists, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplySmtpConfigsAsync(content.SmtpConfigs, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
await ApplySmsConfigsAsync(content.SmsConfigs, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
// Inbound API keys are NOT applied from a bundle — any keys
|
||||
// in a legacy bundle were counted above (apiKeysIgnored) and are skipped.
|
||||
await ApplyApiMethodsAsync(content.ApiMethods, resolutionMap, user, summary, ct).ConfigureAwait(false);
|
||||
|
||||
// Second-pass rewire of name-keyed
|
||||
// FKs that can only be resolved AFTER every template's scripts and
|
||||
// child rows have been staged. We flush here so Pass A can look up
|
||||
// the just-persisted scripts by name and Pass B can look up the
|
||||
// just-persisted templates by name; both passes stage further
|
||||
// mutations that ride the SAME outer transaction (committed below).
|
||||
//
|
||||
// EF tracking note: AddAsync on the in-memory provider assigns
|
||||
// synthetic ids eagerly, so this intermediate flush mostly
|
||||
// materialises identity values on a relational provider. The
|
||||
// rollback contract is preserved because semantic validation
|
||||
// already ran above — any throw from this point onward represents
|
||||
// a successful merge that the user wants to keep, and the only
|
||||
// remaining failure surface is the final BundleImported audit
|
||||
// write itself.
|
||||
await _dbContext.SaveChangesAsync(ct).ConfigureAwait(false);
|
||||
await ResolveAlarmScriptLinksAsync(content.Templates, resolutionMap, user, ct).ConfigureAwait(false);
|
||||
await ResolveCompositionEdgesAsync(content.Templates, resolutionMap, user, ct).ConfigureAwait(false);
|
||||
await ResolveInheritanceEdgesAsync(content.Templates, resolutionMap, user, ct).ConfigureAwait(false);
|
||||
// Folder parent edges are name-keyed too — resolve them AFTER the flush
|
||||
// above so every imported folder has a materialised id to point at
|
||||
// (the Add pass created the folders parent-less). Mirrors the template
|
||||
// inheritance/composition rewires.
|
||||
await ResolveFolderParentEdgesAsync(content.TemplateFolders, resolutionMap, user, ct).ConfigureAwait(false);
|
||||
// External-system methods have no navigation on the parent definition,
|
||||
// so an Add can't cascade them — persist them here, once the parent
|
||||
// system has a real id from the flush above. (Overwrite already syncs
|
||||
// its methods inline against the existing parent id.)
|
||||
await PersistAddedExternalSystemMethodsAsync(content.ExternalSystems, resolutionMap, user, ct).ConfigureAwait(false);
|
||||
|
||||
// ---- Import-time acyclicity guard over the merged template graph ----
|
||||
// The two rewire passes above have STAGED (not yet saved) the bundle's
|
||||
// inheritance (ParentTemplateId) and composition edges onto the tracked
|
||||
// target templates. A bundle can still close a loop through a pre-existing
|
||||
// target template — e.g. Overwrite a base so it now inherits from (or
|
||||
// composes) one of its own descendants. The design-time save-gate enforces
|
||||
// acyclicity on individual edits, but an import merges many edges at once,
|
||||
// so re-check the whole merged graph here. On a cycle we throw and the
|
||||
// catch below rolls the whole import back (ChangeTracker.Clear on the
|
||||
// in-memory provider) BEFORE any of it is persisted; deliberately no
|
||||
// SaveChanges precedes this check so nothing survives that rollback.
|
||||
EnsureNoTemplateGraphCycles();
|
||||
|
||||
// ---- Site/instance-scoped apply: instances LAST ----
|
||||
// The instance pass needs every FK target materialised first:
|
||||
// • template id — resolved by name against the just-flushed central
|
||||
// config (Templates were flushed above for the alarm/composition
|
||||
// rewire), plus pre-existing target templates;
|
||||
// • site id — from the site map built by ApplySitesAsync;
|
||||
// • connection id — from the connection map built by
|
||||
// ApplyDataConnectionsAsync.
|
||||
// It writes the instance row + its four child collections (attribute /
|
||||
// alarm / native-alarm-source overrides + connection bindings) and
|
||||
// rewires every name-keyed FK to the target environment's ids.
|
||||
await ApplyInstancesAsync(
|
||||
content, resolutionMap, siteBySourceIdentifier, connectionMaps,
|
||||
user, summary, ct).ConfigureAwait(false);
|
||||
|
||||
// ---- Pre-commit point: stale-instance computation ----
|
||||
// Everything the import will write is staged on the change tracker at
|
||||
// this point but NOT yet committed. This computes StaleInstanceIds here
|
||||
// — target instances whose template was overwritten by this import and
|
||||
// therefore now carry a stale flattened-config revision — and threads
|
||||
// the resulting list into the ImportResult below (replacing the
|
||||
// Array.Empty<int>() placeholder). The single deferred SaveChangesAsync
|
||||
// is the next statement, so a read against the change tracker here sees
|
||||
// the full post-apply graph before commit.
|
||||
var staleInstanceIds = await ComputeStaleInstanceIdsAsync(
|
||||
content, resolutionMap, ct).ConfigureAwait(false);
|
||||
|
||||
await _auditService.LogAsync(
|
||||
user: user,
|
||||
action: "BundleImported",
|
||||
entityType: "Bundle",
|
||||
entityId: bundleImportId.ToString(),
|
||||
entityName: session.Manifest.SourceEnvironment,
|
||||
afterState: new
|
||||
{
|
||||
BundleImportId = bundleImportId,
|
||||
session.Manifest.SourceEnvironment,
|
||||
session.Manifest.ContentHash,
|
||||
Summary = new
|
||||
{
|
||||
summary.Added,
|
||||
summary.Overwritten,
|
||||
summary.Skipped,
|
||||
summary.Renamed,
|
||||
},
|
||||
// Legacy inbound API keys present in the bundle that
|
||||
// were ignored (not transported / not re-created here).
|
||||
ApiKeysIgnored = apiKeysIgnored,
|
||||
},
|
||||
cancellationToken: ct).ConfigureAwait(false);
|
||||
|
||||
await _dbContext.SaveChangesAsync(ct).ConfigureAwait(false);
|
||||
|
||||
return new ImportResult(
|
||||
BundleImportId: bundleImportId,
|
||||
Added: summary.Added,
|
||||
Overwritten: summary.Overwritten,
|
||||
Skipped: summary.Skipped,
|
||||
Renamed: summary.Renamed,
|
||||
StaleInstanceIds: staleInstanceIds,
|
||||
AuditEventCorrelation: bundleImportId.ToString(),
|
||||
ApiKeysIgnored: apiKeysIgnored,
|
||||
Warnings: validationWarnings);
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
|
||||
+4
-2
@@ -24,7 +24,7 @@ namespace ZB.MOM.WW.ScadaBridge.CentralUI.PlaywrightTests.Notifications;
|
||||
/// </list>
|
||||
///
|
||||
/// <para>
|
||||
/// Idempotency: the SMS-config fact uses a STABLE Account SID (<c>ACtest123</c>) and probes
|
||||
/// Idempotency: the SMS-config fact uses a STABLE Account SID (<c>AC000…0001</c>) and probes
|
||||
/// <c>notification sms list</c> first — if a prior run already created it, the fact SKIPS the
|
||||
/// create and verifies the existing row instead (the page exposes no delete verb — neither UI
|
||||
/// nor CLI — so we never create a second copy we could not clean up; the render-crash +
|
||||
@@ -41,7 +41,9 @@ public class SmsNotificationE2ETests
|
||||
// Stable test fixture values. The Account SID is stable (not random) so reruns find the
|
||||
// prior config via 'notification sms list' and skip re-creation — the SMS config page has
|
||||
// no delete verb, so a unique-per-run SID would leak a config on every run.
|
||||
private const string TestAccountSid = "ACtest123";
|
||||
// Must match the ManagementActor create guard ^AC[0-9a-fA-F]{32}$ (added 2026-07-10) or
|
||||
// the create is rejected and the config-card / secret-non-leak assertions never run.
|
||||
private const string TestAccountSid = "AC00000000000000000000000000000001";
|
||||
private const string TestFromNumber = "+15551230000";
|
||||
// The Auth Token VALUE that must NEVER be echoed back to the page (presence flag only).
|
||||
private const string TestAuthTokenValue = "e2e-secret-token-xyz";
|
||||
|
||||
+3
-1
@@ -42,7 +42,9 @@ public class MonitoringAuthorizationTests
|
||||
public void HealthDashboard_IsIntentionallyAllAuthenticatedRoles()
|
||||
{
|
||||
// Health Dashboard stays all-roles (no policy) per the design doc.
|
||||
var attr = AuthorizeOf<Health>();
|
||||
// Fully qualified: the Communication.Health namespace otherwise makes
|
||||
// the bare `Health` page type ambiguous.
|
||||
var attr = AuthorizeOf<CentralUI.Components.Pages.Monitoring.Health>();
|
||||
|
||||
Assert.NotNull(attr);
|
||||
Assert.Null(attr!.Policy);
|
||||
|
||||
+147
@@ -0,0 +1,147 @@
|
||||
namespace ZB.MOM.WW.ScadaBridge.ClusterInfrastructure.Tests;
|
||||
|
||||
/// <summary>
|
||||
/// Unit tests for <see cref="ClusterBootstrapGuard"/> — the pure decision core of the
|
||||
/// simultaneous-cold-start split-brain guard (Gitea #33). The runtime probe + JoinSeedNodes wiring
|
||||
/// lives in <c>ClusterBootstrapCoordinator</c> and is covered by the real-ActorSystem
|
||||
/// <c>ClusterBootstrapCoordinatorTests</c> in the integration suite.
|
||||
/// </summary>
|
||||
public class ClusterBootstrapGuardTests
|
||||
{
|
||||
private const string A = "akka.tcp://scadabridge@scadabridge-site-a-a:8082";
|
||||
private const string B = "akka.tcp://scadabridge@scadabridge-site-a-b:8082";
|
||||
|
||||
[Fact]
|
||||
public void Lower_address_node_is_the_founder_and_needs_no_probe()
|
||||
{
|
||||
// scadabridge-site-a-a < scadabridge-site-a-b, so on node-a the guard makes it the founder.
|
||||
var role = ClusterBootstrapGuard.Analyze(new[] { A, B }, "scadabridge-site-a-a", 8082);
|
||||
|
||||
Assert.True(role.Applies);
|
||||
Assert.True(role.IsFounder);
|
||||
Assert.Equal(new[] { A, B }, role.SelfFirstOrder);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Higher_address_node_is_not_the_founder_and_targets_the_partner_for_probing()
|
||||
{
|
||||
// On node-b the partner is node-a (the lower/founder).
|
||||
var role = ClusterBootstrapGuard.Analyze(new[] { B, A }, "scadabridge-site-a-b", 8082);
|
||||
|
||||
Assert.True(role.Applies);
|
||||
Assert.False(role.IsFounder);
|
||||
Assert.Equal("scadabridge-site-a-a", role.PartnerHost);
|
||||
Assert.Equal(8082, role.PartnerPort);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Tie_break_is_symmetric_regardless_of_seed_order()
|
||||
{
|
||||
// Both nodes must agree on who founds no matter how each lists its seeds.
|
||||
var onLower = ClusterBootstrapGuard.Analyze(new[] { A, B }, "scadabridge-site-a-a", 8082);
|
||||
var onHigher = ClusterBootstrapGuard.Analyze(new[] { B, A }, "scadabridge-site-a-b", 8082);
|
||||
|
||||
Assert.True(onLower.IsFounder);
|
||||
Assert.False(onHigher.IsFounder);
|
||||
// Exactly one founder.
|
||||
Assert.True(onLower.IsFounder ^ onHigher.IsFounder);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Tie_break_is_case_insensitive_so_a_casing_mismatch_cannot_make_both_founders()
|
||||
{
|
||||
// The two sides' seed config differ only in host casing (review note 1): if the tie-break
|
||||
// were case-sensitive, both could conclude they are the founder and re-open the split.
|
||||
var lowerSelf = "akka.tcp://scadabridge@SITE-A-A:8082";
|
||||
var higherSelf = "akka.tcp://scadabridge@site-a-b:8082";
|
||||
|
||||
var onLower = ClusterBootstrapGuard.Analyze(new[] { lowerSelf, higherSelf }, "SITE-A-A", 8082);
|
||||
var onHigher = ClusterBootstrapGuard.Analyze(new[] { higherSelf, lowerSelf }, "site-a-b", 8082);
|
||||
|
||||
Assert.True(onLower.IsFounder ^ onHigher.IsFounder);
|
||||
Assert.True(onLower.IsFounder); // "site-a-a" < "site-a-b" ordinal-ignore-case
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Higher_node_joins_the_founder_when_partner_is_reachable()
|
||||
{
|
||||
var role = ClusterBootstrapGuard.Analyze(new[] { B, A }, "scadabridge-site-a-b", 8082);
|
||||
|
||||
// Reachable ⇒ peer-first ⇒ JoinSeedNodeProcess ⇒ joins the founder, never self-forms.
|
||||
Assert.Equal(new[] { A, B }, ClusterBootstrapGuard.HigherNodeOrder(role, partnerReachable: true)); // partner (node-a) first
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Higher_node_forms_alone_when_partner_is_unreachable()
|
||||
{
|
||||
var role = ClusterBootstrapGuard.Analyze(new[] { B, A }, "scadabridge-site-a-b", 8082);
|
||||
|
||||
// Unreachable after the window ⇒ partner is dead ⇒ self-first ⇒ cold-start-alone preserved.
|
||||
Assert.Equal(new[] { B, A }, ClusterBootstrapGuard.HigherNodeOrder(role, partnerReachable: false)); // self (node-b) first
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Guard_does_not_apply_to_a_single_seed_node()
|
||||
{
|
||||
// A single-node install lists exactly one seed (itself) — not a pair, guard is inert.
|
||||
var role = ClusterBootstrapGuard.Analyze(new[] { "akka.tcp://scadabridge@lone:8081" }, "lone", 8081);
|
||||
|
||||
Assert.False(role.Applies);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Guard_does_not_apply_when_self_is_absent_from_the_two_seeds()
|
||||
{
|
||||
var role = ClusterBootstrapGuard.Analyze(new[] { A, B }, "scadabridge-central-a", 8081);
|
||||
|
||||
Assert.False(role.Applies);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Guard_does_not_apply_to_three_or_more_seeds()
|
||||
{
|
||||
var role = ClusterBootstrapGuard.Analyze(
|
||||
new[] { A, B, "akka.tcp://scadabridge@scadabridge-site-a-c:8082" }, "scadabridge-site-a-a", 8082);
|
||||
|
||||
Assert.False(role.Applies);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Guard_does_not_apply_when_a_seed_is_malformed()
|
||||
{
|
||||
var role = ClusterBootstrapGuard.Analyze(new[] { A, "not-a-seed" }, "scadabridge-site-a-a", 8082);
|
||||
|
||||
Assert.False(role.Applies);
|
||||
}
|
||||
|
||||
[Theory]
|
||||
[InlineData("akka.tcp://scadabridge@scadabridge-site-a-a:8082", "scadabridge-site-a-a", 8082)]
|
||||
[InlineData("akka.tcp://scadabridge@10.0.0.5:8082/", "10.0.0.5", 8082)]
|
||||
[InlineData(" akka.tcp://scadabridge@host.lan:1234 ", "host.lan", 1234)]
|
||||
public void TryParse_extracts_host_and_port(string seed, string expectedHost, int expectedPort)
|
||||
{
|
||||
Assert.True(ClusterBootstrapGuard.TryParse(seed, out var host, out var port));
|
||||
Assert.Equal(expectedHost, host);
|
||||
Assert.Equal(expectedPort, port);
|
||||
}
|
||||
|
||||
[Theory]
|
||||
[InlineData("")]
|
||||
[InlineData("garbage")]
|
||||
[InlineData("akka.tcp://scadabridge@host")] // no port
|
||||
public void TryParse_rejects_malformed_seeds(string seed)
|
||||
{
|
||||
Assert.False(ClusterBootstrapGuard.TryParse(seed, out _, out _));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Port_participates_in_the_tie_break_when_hosts_are_equal()
|
||||
{
|
||||
// Shared-host loopback pair (test rig): same host, different ports — port breaks the tie.
|
||||
const string low = "akka.tcp://scadabridge@127.0.0.1:8081";
|
||||
const string high = "akka.tcp://scadabridge@127.0.0.1:8082";
|
||||
|
||||
Assert.True(ClusterBootstrapGuard.Analyze(new[] { low, high }, "127.0.0.1", 8081).IsFounder);
|
||||
Assert.False(ClusterBootstrapGuard.Analyze(new[] { high, low }, "127.0.0.1", 8082).IsFounder);
|
||||
}
|
||||
}
|
||||
+64
@@ -215,4 +215,68 @@ public class ClusterOptionsValidatorTests
|
||||
Assert.Contains("StableAfter", result.FailureMessage);
|
||||
Assert.Contains("HeartbeatInterval", result.FailureMessage);
|
||||
}
|
||||
|
||||
// ---- BootstrapGuard timing validation (Gitea #33) ----
|
||||
|
||||
[Fact]
|
||||
public void BootstrapGuard_disabled_with_zero_timings_passes_validation()
|
||||
{
|
||||
// The knobs are inert when the guard is off, so a disabled guard must never block a boot —
|
||||
// even if the timing values are nonsense (0 here).
|
||||
var options = ValidOptions();
|
||||
options.BootstrapGuard = new ClusterBootstrapGuardOptions
|
||||
{
|
||||
Enabled = false,
|
||||
PartnerProbeSeconds = 0,
|
||||
PartnerProbeIntervalMs = 0,
|
||||
ProbeConnectTimeoutMs = 0,
|
||||
};
|
||||
|
||||
var result = new ClusterOptionsValidator().Validate(null, options);
|
||||
|
||||
Assert.True(result.Succeeded, result.FailureMessage);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void BootstrapGuard_enabled_with_default_timings_passes_validation()
|
||||
{
|
||||
var options = ValidOptions();
|
||||
options.BootstrapGuard = new ClusterBootstrapGuardOptions { Enabled = true };
|
||||
|
||||
var result = new ClusterOptionsValidator().Validate(null, options);
|
||||
|
||||
Assert.True(result.Succeeded, result.FailureMessage);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void BootstrapGuard_enabled_with_nonpositive_probe_seconds_fails_validation()
|
||||
{
|
||||
// A zero PartnerProbeSeconds silently degrades the guard to "never wait, always conclude the
|
||||
// partner is dead" — the higher node forms alone immediately and re-opens the split. Reject it.
|
||||
var options = ValidOptions();
|
||||
options.BootstrapGuard = new ClusterBootstrapGuardOptions { Enabled = true, PartnerProbeSeconds = 0 };
|
||||
|
||||
var result = new ClusterOptionsValidator().Validate(null, options);
|
||||
|
||||
Assert.True(result.Failed);
|
||||
Assert.Contains("PartnerProbeSeconds", result.FailureMessage);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void BootstrapGuard_enabled_with_nonpositive_interval_or_timeout_fails_validation()
|
||||
{
|
||||
var options = ValidOptions();
|
||||
options.BootstrapGuard = new ClusterBootstrapGuardOptions
|
||||
{
|
||||
Enabled = true,
|
||||
PartnerProbeIntervalMs = 0,
|
||||
ProbeConnectTimeoutMs = -1,
|
||||
};
|
||||
|
||||
var result = new ClusterOptionsValidator().Validate(null, options);
|
||||
|
||||
Assert.True(result.Failed);
|
||||
Assert.Contains("PartnerProbeIntervalMs", result.FailureMessage);
|
||||
Assert.Contains("ProbeConnectTimeoutMs", result.FailureMessage);
|
||||
}
|
||||
}
|
||||
|
||||
+27
@@ -0,0 +1,27 @@
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Types.DataConnections;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Commons.Tests.Types.DataConnections;
|
||||
|
||||
public class OpcUaReferenceFormTests
|
||||
{
|
||||
[Theory]
|
||||
// durable → true
|
||||
[InlineData("nsu=https://zb.com/otopcua/raw;s=Line1.Pump.Speed", true)]
|
||||
[InlineData("nsu=https://zb.com/otopcua/uns;s=AreaA/Line1/Pump/Speed", true)]
|
||||
[InlineData("NSU=https://x;s=y", true)] // case-insensitive prefix
|
||||
[InlineData("ns=0;i=85", true)] // spec-fixed namespace 0
|
||||
[InlineData("i=85", true)] // short form, implicit ns0
|
||||
[InlineData("s=Devices.Pump1", true)] // no namespace prefix
|
||||
[InlineData("", true)] // nothing typed
|
||||
[InlineData(" ", true)]
|
||||
[InlineData(null, true)]
|
||||
// bare server-namespace index → false (warn)
|
||||
[InlineData("ns=2;s=MyDevice.Temperature", false)]
|
||||
[InlineData("ns=1;i=1001", false)]
|
||||
[InlineData("ns=10;s=x", false)]
|
||||
[InlineData(" ns=3;s=x ", false)] // trimmed before inspection
|
||||
public void IsDurable_classifies_reference_forms(string? reference, bool expected)
|
||||
{
|
||||
Assert.Equal(expected, OpcUaReferenceForm.IsDurable(reference));
|
||||
}
|
||||
}
|
||||
@@ -1,4 +1,3 @@
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.Integration;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.RemoteQuery;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Communication.Tests;
|
||||
@@ -8,25 +7,6 @@ namespace ZB.MOM.WW.ScadaBridge.Communication.Tests;
|
||||
/// </summary>
|
||||
public class MessageContractTests
|
||||
{
|
||||
[Fact]
|
||||
public void IntegrationCallRequest_HasCorrelationId()
|
||||
{
|
||||
var msg = new IntegrationCallRequest(
|
||||
"corr-123", "site1", "inst1", "ExtSys1", "GetData",
|
||||
new Dictionary<string, object?>(), DateTimeOffset.UtcNow);
|
||||
|
||||
Assert.Equal("corr-123", msg.CorrelationId);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void IntegrationCallResponse_HasCorrelationId()
|
||||
{
|
||||
var msg = new IntegrationCallResponse(
|
||||
"corr-123", "site1", true, "{}", null, DateTimeOffset.UtcNow);
|
||||
|
||||
Assert.Equal("corr-123", msg.CorrelationId);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void EventLogQueryRequest_HasCorrelationId()
|
||||
{
|
||||
@@ -79,10 +59,6 @@ public class MessageContractTests
|
||||
Assert.True(typeof(Commons.Messages.Artifacts.DeployArtifactsCommand).IsValueType == false);
|
||||
Assert.True(typeof(Commons.Messages.Artifacts.ArtifactDeploymentResponse).IsValueType == false);
|
||||
|
||||
// Pattern 4: Integration
|
||||
Assert.True(typeof(IntegrationCallRequest).IsValueType == false);
|
||||
Assert.True(typeof(IntegrationCallResponse).IsValueType == false);
|
||||
|
||||
// Pattern 5: Debug View
|
||||
Assert.True(typeof(Commons.Messages.DebugView.SubscribeDebugViewRequest).IsValueType == false);
|
||||
Assert.True(typeof(Commons.Messages.DebugView.DebugViewSnapshot).IsValueType == false);
|
||||
|
||||
@@ -1,7 +1,6 @@
|
||||
using Akka.Actor;
|
||||
using Akka.TestKit.Xunit2;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.Artifacts;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.Integration;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.RemoteQuery;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Types;
|
||||
using ZB.MOM.WW.ScadaBridge.Communication.Actors;
|
||||
@@ -243,18 +242,6 @@ public class SiteCommandDispatcherTests : TestKit
|
||||
dispatcher.ResolveRoute(new TriggerSiteFailover("c", SiteId)));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void ResolveRoute_RejectsTheExcludedIntegrationCommand()
|
||||
{
|
||||
// IntegrationCallRequest is the 29th command, dead at both ends and deliberately excluded
|
||||
// (28 of 29 migrate). It never enters the dispatcher.
|
||||
var dispatcher = Build(CreateTestProbe().Ref);
|
||||
var command = new IntegrationCallRequest(
|
||||
"c", SiteId, "inst", "es", "m", new Dictionary<string, object?>(), DateTimeOffset.UtcNow);
|
||||
|
||||
Assert.Throws<ArgumentException>(() => dispatcher.ResolveRoute(command));
|
||||
}
|
||||
|
||||
// ── Failover (local path) ──
|
||||
|
||||
[Fact]
|
||||
|
||||
@@ -94,10 +94,8 @@ public class SiteCommandDtoMapperGoldenTests
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// The contract carries exactly 28 commands — 29 on
|
||||
/// <c>SiteCommunicationActor</c>'s receive table minus the dead
|
||||
/// <c>IntegrationCallRequest</c>. If the site gains a command, this count
|
||||
/// moves deliberately, not silently.
|
||||
/// The contract carries exactly 28 commands across six domain groups. If the
|
||||
/// site gains a command, this count moves deliberately, not silently.
|
||||
/// </summary>
|
||||
[Fact]
|
||||
public void CommandInventory_Is28_AcrossSixGroups()
|
||||
@@ -431,17 +429,6 @@ public class SiteCommandDtoMapperGoldenTests
|
||||
// Group classification + rejection
|
||||
// ─────────────────────────────────────────────────────────────────────
|
||||
|
||||
/// <summary><c>IntegrationCallRequest</c> is excluded by design and must be rejected, not silently dropped.</summary>
|
||||
[Fact]
|
||||
public void IntegrationCallRequest_IsRejected()
|
||||
{
|
||||
var dead = new Commons.Messages.Integration.IntegrationCallRequest(
|
||||
"corr", "site-a", "Instance", "System", "Method",
|
||||
new Dictionary<string, object?>(), DateTimeOffset.UtcNow);
|
||||
|
||||
Assert.Throws<ArgumentException>(() => SiteCommandDtoMapper.GroupOf(dead));
|
||||
}
|
||||
|
||||
/// <summary>Packing a command into the wrong group's envelope is a hard error, not a silent no-op.</summary>
|
||||
[Fact]
|
||||
public void PackingIntoTheWrongGroup_Throws() =>
|
||||
|
||||
@@ -5,7 +5,6 @@ using ZB.MOM.WW.ScadaBridge.Commons.Messages.DataConnection;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.Deployment;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.Health;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.Lifecycle;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.Integration;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.Management;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.Notification;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.RemoteQuery;
|
||||
@@ -73,42 +72,6 @@ public class SiteCommunicationActorTests : TestKit
|
||||
dmProbe.ExpectMsg<DeploymentStateQueryRequest>(msg => msg.CorrelationId == "corr-q");
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void IntegrationCall_WithoutHandler_ReturnsFailure()
|
||||
{
|
||||
var dmProbe = CreateTestProbe();
|
||||
var siteActor = Sys.ActorOf(Props.Create(() =>
|
||||
new SiteCommunicationActor("site1", _options, dmProbe.Ref)));
|
||||
|
||||
var request = new IntegrationCallRequest(
|
||||
"corr1", "site1", "inst1", "ExtSys1", "GetData",
|
||||
new Dictionary<string, object?>(), DateTimeOffset.UtcNow);
|
||||
|
||||
siteActor.Tell(request);
|
||||
|
||||
ExpectMsg<IntegrationCallResponse>(msg =>
|
||||
!msg.Success && msg.ErrorMessage == "Integration handler not available");
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void IntegrationCall_WithHandler_ForwardedToHandler()
|
||||
{
|
||||
var dmProbe = CreateTestProbe();
|
||||
var handlerProbe = CreateTestProbe();
|
||||
var siteActor = Sys.ActorOf(Props.Create(() =>
|
||||
new SiteCommunicationActor("site1", _options, dmProbe.Ref)));
|
||||
|
||||
// Register integration handler
|
||||
siteActor.Tell(new RegisterLocalHandler(LocalHandlerType.Integration, handlerProbe.Ref));
|
||||
|
||||
var request = new IntegrationCallRequest(
|
||||
"corr1", "site1", "inst1", "ExtSys1", "GetData",
|
||||
new Dictionary<string, object?>(), DateTimeOffset.UtcNow);
|
||||
|
||||
siteActor.Tell(request);
|
||||
handlerProbe.ExpectMsg<IntegrationCallRequest>(msg => msg.CorrelationId == "corr1");
|
||||
}
|
||||
|
||||
// The site→central send-forwarding tests that used to live here (NotificationSubmit /
|
||||
// NotificationStatusQuery, with and without a registered central ClusterClient) were removed
|
||||
// with the Akka transport in the ClusterClient→gRPC migration's Phase 4: the actor no longer
|
||||
|
||||
+102
@@ -0,0 +1,102 @@
|
||||
using Opc.Ua;
|
||||
using ZB.MOM.WW.ScadaBridge.DataConnectionLayer.Adapters;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.DataConnectionLayer.Tests.Adapters;
|
||||
|
||||
/// <summary>
|
||||
/// Covers <see cref="RealOpcUaClient.RewriteEndpointHostForReachability"/> — the fix for OPC UA
|
||||
/// servers (e.g. OtOpcUa v3) that advertise an unroutable base address (a 0.0.0.0 wildcard bind or
|
||||
/// an internal container/NAT hostname). The session must dial the host the operator reached, not
|
||||
/// the advertised one, while preserving the advertised scheme + path.
|
||||
/// </summary>
|
||||
public class RealOpcUaClientEndpointRewriteTests
|
||||
{
|
||||
private static string? Rewrite(string advertised, string discovery)
|
||||
{
|
||||
var ep = new EndpointDescription(advertised);
|
||||
RealOpcUaClient.RewriteEndpointHostForReachability(ep, discovery);
|
||||
return ep.EndpointUrl;
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Rewrites_0_0_0_0_advertisement_to_the_reachable_host_preserving_path()
|
||||
{
|
||||
Assert.Equal(
|
||||
"opc.tcp://otopcua-dev-site-a-1-1:4840/OtOpcUa",
|
||||
Rewrite("opc.tcp://0.0.0.0:4840/OtOpcUa", "opc.tcp://otopcua-dev-site-a-1-1:4840/OtOpcUa"));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Rewrites_internal_nat_hostname_to_the_reachable_host()
|
||||
{
|
||||
Assert.Equal(
|
||||
"opc.tcp://plc.example.com:4840/OtOpcUa",
|
||||
Rewrite("opc.tcp://be5667fab9da:4840/OtOpcUa", "opc.tcp://plc.example.com:4840/OtOpcUa"));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Rewrites_port_too_when_the_advertised_port_differs()
|
||||
{
|
||||
Assert.Equal(
|
||||
"opc.tcp://gateway:51210/OtOpcUa",
|
||||
Rewrite("opc.tcp://0.0.0.0:4840/OtOpcUa", "opc.tcp://gateway:51210/OtOpcUa"));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Leaves_an_already_reachable_advertisement_untouched()
|
||||
{
|
||||
// opc-plc advertises its real container name — must not be rewritten.
|
||||
Assert.Equal(
|
||||
"opc.tcp://scadabridge-opcua:50000",
|
||||
Rewrite("opc.tcp://scadabridge-opcua:50000", "opc.tcp://scadabridge-opcua:50000"));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Handles_authority_only_advertisement_without_appending_a_trailing_slash()
|
||||
{
|
||||
Assert.Equal(
|
||||
"opc.tcp://host-b:4840",
|
||||
Rewrite("opc.tcp://0.0.0.0:4840", "opc.tcp://host-b:4840"));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Re_wraps_an_ipv6_reachable_host_in_brackets()
|
||||
{
|
||||
Assert.Equal(
|
||||
"opc.tcp://[::1]:4840/OtOpcUa",
|
||||
Rewrite("opc.tcp://0.0.0.0:4840/OtOpcUa", "opc.tcp://[::1]:4840/OtOpcUa"));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Emits_no_port_when_neither_url_carries_one()
|
||||
{
|
||||
// opc.tcp has no default port, so Uri.Port is -1 on both — must not emit ":-1".
|
||||
Assert.Equal(
|
||||
"opc.tcp://myhost/UaServer",
|
||||
Rewrite("opc.tcp://0.0.0.0/UaServer", "opc.tcp://myhost/UaServer"));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Uses_the_reachable_port_when_the_advertisement_omits_one()
|
||||
{
|
||||
Assert.Equal(
|
||||
"opc.tcp://myhost:4840/UaServer",
|
||||
Rewrite("opc.tcp://0.0.0.0/UaServer", "opc.tcp://myhost:4840/UaServer"));
|
||||
}
|
||||
|
||||
[Theory]
|
||||
[InlineData("")]
|
||||
[InlineData(" ")]
|
||||
[InlineData("not-a-uri")]
|
||||
public void Leaves_endpoint_unchanged_when_discovery_url_is_unusable(string discovery)
|
||||
{
|
||||
Assert.Equal("opc.tcp://0.0.0.0:4840/OtOpcUa", Rewrite("opc.tcp://0.0.0.0:4840/OtOpcUa", discovery));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Null_endpoint_is_a_no_op()
|
||||
{
|
||||
// Must not throw.
|
||||
RealOpcUaClient.RewriteEndpointHostForReachability(null, "opc.tcp://host:4840");
|
||||
}
|
||||
}
|
||||
@@ -5,6 +5,7 @@ using Microsoft.Extensions.DependencyInjection;
|
||||
using Microsoft.Extensions.Diagnostics.HealthChecks;
|
||||
using Microsoft.Extensions.Options;
|
||||
using ZB.MOM.WW.Health;
|
||||
using ZB.MOM.WW.Health.Akka;
|
||||
using ZB.MOM.WW.ScadaBridge.Host.Health;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Host.Tests;
|
||||
@@ -207,11 +208,13 @@ public class HealthCheckTests : IDisposable
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void ActiveNodeCheck_IsHostLocalOldestNodeCheck()
|
||||
public void ActiveNodeCheck_IsTheSharedOldestUpMemberCheck()
|
||||
{
|
||||
// The shared package's ActiveNodeHealthCheck is LEADER-based (review 01 [High]):
|
||||
// during a partition both nodes think they lead and Traefik serves both. The
|
||||
// active-node check must be the host-local oldest-member check instead.
|
||||
// The property under test is that the active tier selects by AGE, not by leadership: during a
|
||||
// partition both sides think they lead and Traefik would serve both (review 01 [High]).
|
||||
// Through ZB.MOM.WW.Health 0.2.1 the shared check was leader-based, so this host carried a
|
||||
// private OldestNodeActiveHealthCheck; 0.3.0 made the shared check age-based, so the private
|
||||
// copy is gone and the shared type is now the correct answer here.
|
||||
var previousEnv = Environment.GetEnvironmentVariable("DOTNET_ENVIRONMENT");
|
||||
try
|
||||
{
|
||||
@@ -222,7 +225,7 @@ public class HealthCheckTests : IDisposable
|
||||
Assert.Contains(ZbHealthTags.Active, active.Tags);
|
||||
|
||||
var check = active.Factory(factory.Services);
|
||||
Assert.IsType<OldestNodeActiveHealthCheck>(check);
|
||||
Assert.IsType<ActiveNodeHealthCheck>(check);
|
||||
}
|
||||
finally
|
||||
{
|
||||
|
||||
@@ -0,0 +1,239 @@
|
||||
using Microsoft.AspNetCore.Builder;
|
||||
using Microsoft.Extensions.Configuration;
|
||||
using Microsoft.Extensions.DependencyInjection;
|
||||
using Microsoft.Extensions.Diagnostics.HealthChecks;
|
||||
using Microsoft.Extensions.Options;
|
||||
using ZB.MOM.WW.Health;
|
||||
using ZB.MOM.WW.Health.Akka;
|
||||
using ZB.MOM.WW.LocalDb;
|
||||
using ZB.MOM.WW.LocalDb.Registration;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Messages.Health;
|
||||
using ZB.MOM.WW.ScadaBridge.HealthMonitoring;
|
||||
using ZB.MOM.WW.ScadaBridge.Host.Health;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Host.Tests;
|
||||
|
||||
/// <summary>
|
||||
/// Site-node health endpoints. Site nodes previously served no health surface at all — only gRPC on
|
||||
/// the HTTP/2 listener and <c>/metrics</c> on the HTTP/1.1 one — so the family overview dashboard
|
||||
/// had no uniform way to probe them. These tests pin the three checks the site now registers and
|
||||
/// the tier each one lands in.
|
||||
/// <para>
|
||||
/// The registration assertions build the REAL container from
|
||||
/// <see cref="SiteServiceRegistration.Configure"/> rather than a hand-assembled
|
||||
/// <c>ServiceCollection</c>, for the reason spelled out in <see cref="SiteLocalDbWiringTests"/>:
|
||||
/// this family's wiring defects have consistently been ones that only a container-built graph
|
||||
/// catches. Each registration is additionally ACTIVATED through its factory, so a check that is
|
||||
/// registered but cannot resolve its dependencies fails here instead of at the first live probe.
|
||||
/// </para>
|
||||
/// </summary>
|
||||
public class SiteHealthCheckTests : IDisposable
|
||||
{
|
||||
private readonly WebApplication _host;
|
||||
private readonly string _tempDbPath;
|
||||
private readonly string _tempSiteDbPath;
|
||||
private readonly string _tempTrackingDbPath;
|
||||
|
||||
public SiteHealthCheckTests()
|
||||
{
|
||||
var stamp = Guid.NewGuid();
|
||||
_tempDbPath = Path.Combine(Path.GetTempPath(), $"scadabridge_health_localdb_{stamp}.db");
|
||||
_tempSiteDbPath = Path.Combine(Path.GetTempPath(), $"scadabridge_health_site_{stamp}.db");
|
||||
_tempTrackingDbPath = Path.Combine(Path.GetTempPath(), $"scadabridge_health_track_{stamp}.db");
|
||||
|
||||
var builder = WebApplication.CreateBuilder();
|
||||
builder.Configuration.Sources.Clear();
|
||||
builder.Configuration.AddInMemoryCollection(new Dictionary<string, string?>
|
||||
{
|
||||
["ScadaBridge:Node:Role"] = "Site",
|
||||
["ScadaBridge:Node:NodeName"] = "node-a",
|
||||
["ScadaBridge:Node:NodeHostname"] = "test-site",
|
||||
["ScadaBridge:Node:SiteId"] = "TestSite",
|
||||
["ScadaBridge:Node:RemotingPort"] = "0",
|
||||
["ScadaBridge:Node:GrpcPort"] = "0",
|
||||
["ScadaBridge:Database:SiteDbPath"] = _tempSiteDbPath,
|
||||
["ScadaBridge:OperationTracking:ConnectionString"] = $"Data Source={_tempTrackingDbPath}",
|
||||
["ScadaBridge:Cluster:SeedNodes:0"] = "akka.tcp://scadabridge@localhost:2551",
|
||||
["ScadaBridge:Cluster:SeedNodes:1"] = "akka.tcp://scadabridge@localhost:2552",
|
||||
["LocalDb:Path"] = _tempDbPath,
|
||||
});
|
||||
|
||||
builder.Services.AddGrpc();
|
||||
builder.Services.AddSingleton<ZB.MOM.WW.ScadaBridge.Communication.Grpc.SiteStreamGrpcServer>();
|
||||
|
||||
SiteServiceRegistration.Configure(builder.Services, builder.Configuration);
|
||||
|
||||
AkkaHostedServiceRemover.RemoveAkkaHostedServiceOnly(builder.Services);
|
||||
|
||||
_host = builder.Build();
|
||||
}
|
||||
|
||||
public void Dispose()
|
||||
{
|
||||
(_host as IDisposable)?.Dispose();
|
||||
foreach (var path in new[] { _tempDbPath, _tempSiteDbPath, _tempTrackingDbPath })
|
||||
{
|
||||
try { File.Delete(path); } catch { /* best effort */ }
|
||||
}
|
||||
|
||||
GC.SuppressFinalize(this);
|
||||
}
|
||||
|
||||
private IReadOnlyList<HealthCheckRegistration> Registrations =>
|
||||
_host.Services.GetRequiredService<IOptions<HealthCheckServiceOptions>>().Value.Registrations.ToList();
|
||||
|
||||
[Fact]
|
||||
public void Site_RegistersExactlyTheThreeSiteChecks()
|
||||
{
|
||||
// Exact-set, not "contains": an extra check silently changes what /health/ready means for
|
||||
// every consumer, and a site node inheriting a central-only probe is precisely the mistake
|
||||
// this asserts against.
|
||||
Assert.Equal(
|
||||
new[] { "active-node", "akka-cluster", "localdb" },
|
||||
Registrations.Select(r => r.Name).OrderBy(n => n, StringComparer.Ordinal).ToArray());
|
||||
}
|
||||
|
||||
[Theory]
|
||||
[InlineData("akka-cluster", ZbHealthTags.Ready)]
|
||||
[InlineData("localdb", ZbHealthTags.Ready)]
|
||||
[InlineData("active-node", ZbHealthTags.Active)]
|
||||
public void Site_ChecksCarryTheirIntendedTier(string name, string expectedTag)
|
||||
{
|
||||
var registration = Assert.Single(Registrations, r => r.Name == name);
|
||||
|
||||
// Exactly one tier tag each: MapZbHealth selects by tag, so a check tagged both Ready and
|
||||
// Active would put standby-ness into readiness and drop the standby out of readiness too.
|
||||
Assert.Equal(new[] { expectedTag }, registration.Tags.ToArray());
|
||||
}
|
||||
|
||||
[Theory]
|
||||
[InlineData("akka-cluster", typeof(AkkaClusterHealthCheck))]
|
||||
[InlineData("localdb", typeof(SiteLocalDbHealthCheck))]
|
||||
[InlineData("active-node", typeof(SitePairActiveNodeHealthCheck))]
|
||||
public void Site_ChecksActivateToTheIntendedType(string name, Type expected)
|
||||
{
|
||||
var registration = Assert.Single(Registrations, r => r.Name == name);
|
||||
|
||||
// Running the factory is the point: a type-activated check whose constructor dependencies
|
||||
// are not registered on site would throw here rather than on the first probe in production.
|
||||
var check = registration.Factory(_host.Services);
|
||||
|
||||
Assert.IsType(expected, check);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Site_ActiveNodeCheck_IsRoleScoped_NotTheUnscopedCentralCheck()
|
||||
{
|
||||
// Central registers the shared ActiveNodeHealthCheck UNSCOPED — oldest Up member of the whole
|
||||
// cluster. On a mesh carrying more than one site's members that computes "oldest" across the
|
||||
// wrong member set, so a site node could report Active purely for being the oldest member
|
||||
// overall. A site must answer for its own site-{SiteId} role, which SitePairActiveNodeHealthCheck
|
||||
// does by delegating to the site's IClusterNodeProvider (itself now backed by the shared
|
||||
// ClusterActiveNode, so the rule is shared even though the scoping is local).
|
||||
var check = Assert.Single(Registrations, r => r.Name == "active-node").Factory(_host.Services);
|
||||
|
||||
Assert.IsType<SitePairActiveNodeHealthCheck>(check);
|
||||
Assert.IsNotType<ActiveNodeHealthCheck>(check);
|
||||
}
|
||||
|
||||
[Theory]
|
||||
[InlineData("database")]
|
||||
[InlineData("required-singletons")]
|
||||
public void Site_DoesNotRegisterCentralOnlyChecks(string centralOnlyName)
|
||||
{
|
||||
// Positive control first: if the harness ever stopped registering health checks at all,
|
||||
// every "is absent" assertion below would pass vacuously and this test would be worthless.
|
||||
Assert.Contains(Registrations, r => r.Name == "akka-cluster");
|
||||
|
||||
Assert.DoesNotContain(Registrations, r => r.Name == centralOnlyName);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public async Task Site_LocalDbCheck_ReportsHealthyAgainstTheRealSiteDatabase()
|
||||
{
|
||||
// Behaviour, not wiring: the probe runs against the consolidated site database this
|
||||
// composition actually built, so a check that resolves but cannot query fails here.
|
||||
var check = (SiteLocalDbHealthCheck)Assert.Single(Registrations, r => r.Name == "localdb")
|
||||
.Factory(_host.Services);
|
||||
|
||||
var result = await check.CheckHealthAsync(NewContext("localdb", check));
|
||||
|
||||
Assert.Equal(HealthStatus.Healthy, result.Status);
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public async Task Site_LocalDbCheck_EnrichesWithReplicationStateWithoutFailingOnDefaultOff()
|
||||
{
|
||||
// Replication ships default-OFF, so the unreplicated steady state must stay Healthy while
|
||||
// still surfacing the state as data — enrichment must never become a verdict.
|
||||
var check = (SiteLocalDbHealthCheck)Assert.Single(Registrations, r => r.Name == "localdb")
|
||||
.Factory(_host.Services);
|
||||
|
||||
var result = await check.CheckHealthAsync(NewContext("localdb", check));
|
||||
|
||||
Assert.Equal(HealthStatus.Healthy, result.Status);
|
||||
Assert.False(Assert.IsType<bool>(result.Data["replicationConnected"]));
|
||||
Assert.Equal(0L, Assert.IsType<long>(result.Data["replicationOplogBacklog"]));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public async Task Site_LocalDbCheck_ReportsUnhealthyWhenTheStoreIsUnreachable()
|
||||
{
|
||||
var check = new SiteLocalDbHealthCheck(new ThrowingLocalDb());
|
||||
|
||||
var result = await check.CheckHealthAsync(NewContext("localdb", check));
|
||||
|
||||
Assert.Equal(HealthStatus.Unhealthy, result.Status);
|
||||
}
|
||||
|
||||
[Theory]
|
||||
[InlineData(true, HealthStatus.Healthy)]
|
||||
[InlineData(false, HealthStatus.Unhealthy)]
|
||||
public async Task Site_ActiveNodeCheck_SplitsPrimaryFromStandby(bool selfIsPrimary, HealthStatus expected)
|
||||
{
|
||||
// 200 on the primary / 503 on the standby is the same contract central publishes, so a
|
||||
// consumer reads one active-node meaning fleet-wide.
|
||||
var check = new SitePairActiveNodeHealthCheck(new StubNodeProvider(selfIsPrimary));
|
||||
|
||||
var result = await check.CheckHealthAsync(NewContext("active-node", check));
|
||||
|
||||
Assert.Equal(expected, result.Status);
|
||||
}
|
||||
|
||||
private static HealthCheckContext NewContext(string name, IHealthCheck check) => new()
|
||||
{
|
||||
Registration = new HealthCheckRegistration(name, check, HealthStatus.Unhealthy, tags: null),
|
||||
};
|
||||
|
||||
private sealed class StubNodeProvider(bool selfIsPrimary) : IClusterNodeProvider
|
||||
{
|
||||
public bool SelfIsPrimary { get; } = selfIsPrimary;
|
||||
|
||||
public IReadOnlyList<NodeStatus> GetClusterNodes() => [];
|
||||
}
|
||||
|
||||
private sealed class ThrowingLocalDb : ILocalDb
|
||||
{
|
||||
public Task<int> ExecuteAsync(string sql, object? parameters = null, CancellationToken ct = default) =>
|
||||
throw new InvalidOperationException("store unreachable");
|
||||
|
||||
public Task<IReadOnlyList<T>> QueryAsync<T>(
|
||||
string sql,
|
||||
Func<Microsoft.Data.Sqlite.SqliteDataReader, T> map,
|
||||
object? parameters = null,
|
||||
CancellationToken ct = default) =>
|
||||
throw new InvalidOperationException("store unreachable");
|
||||
|
||||
public Task<ILocalDbTransaction> BeginTransactionAsync(CancellationToken ct = default) =>
|
||||
throw new InvalidOperationException("store unreachable");
|
||||
|
||||
public Microsoft.Data.Sqlite.SqliteConnection CreateConnection() =>
|
||||
throw new InvalidOperationException("store unreachable");
|
||||
|
||||
public ReplicatedTable RegisterReplicated(string tableName) =>
|
||||
throw new InvalidOperationException("store unreachable");
|
||||
|
||||
public IReadOnlyDictionary<string, ReplicatedTable> ReplicatedTables =>
|
||||
throw new InvalidOperationException("store unreachable");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,156 @@
|
||||
using System.Net;
|
||||
using System.Text.Json;
|
||||
using Microsoft.AspNetCore.Hosting;
|
||||
using Microsoft.AspNetCore.Mvc.Testing;
|
||||
using Microsoft.AspNetCore.TestHost;
|
||||
using Microsoft.Extensions.DependencyInjection;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Host.Tests;
|
||||
|
||||
/// <summary>
|
||||
/// Proves the site pipeline actually MAPS the health endpoints — the half
|
||||
/// <see cref="SiteHealthCheckTests"/> cannot cover, since that fixture builds the site container
|
||||
/// from <see cref="SiteServiceRegistration"/> and never runs <c>Program</c>'s endpoint wiring.
|
||||
/// Without this, deleting <c>app.MapZbHealth()</c> from the site branch would leave every
|
||||
/// registration test green while site nodes served 404 to the overview dashboard.
|
||||
/// <para>
|
||||
/// Boots the real <c>Program</c> in the Site role. <c>Program</c> builds its own
|
||||
/// <c>ConfigurationBuilder</c> before the web host exists, so its role and settings come from
|
||||
/// PROCESS-WIDE environment variables (<c>SCADABRIDGE_CONFIG</c> selects appsettings.Site.json)
|
||||
/// rather than from <c>WithWebHostBuilder</c> — which is why this fixture joins
|
||||
/// <see cref="HostBootCollection"/> and restores every var it sets.
|
||||
/// </para>
|
||||
/// </summary>
|
||||
[Collection(HostBootCollection.Name)]
|
||||
public class SiteHealthEndpointTests : IDisposable
|
||||
{
|
||||
private readonly List<IDisposable> _disposables = new();
|
||||
private readonly Dictionary<string, string?> _previousEnv = new(StringComparer.Ordinal);
|
||||
private readonly string _tempDbPath;
|
||||
|
||||
public SiteHealthEndpointTests()
|
||||
{
|
||||
_tempDbPath = Path.Combine(Path.GetTempPath(), $"scadabridge_health_ep_{Guid.NewGuid()}.db");
|
||||
|
||||
// Whole-key env overrides, the sanctioned path: supplying GrpcPsk concretely makes the
|
||||
// pre-host secrets expander skip appsettings.Site.json's ${secret:SB-GRPC-PSK-site-1}
|
||||
// token, so this test needs no seeded secrets store.
|
||||
//
|
||||
// The remoting port is moved off the site default (8082) so this fixture cannot collide
|
||||
// with a locally running node, but it must be a REAL port and seed-nodes[0] must be this
|
||||
// node itself — StartupValidator enforces both before the host is built, and it runs
|
||||
// whether or not the Akka hosted service is later removed. Nothing binds it: the hosted
|
||||
// service is removed below.
|
||||
SetEnv("SCADABRIDGE_CONFIG", "Site");
|
||||
SetEnv("ScadaBridge__Node__Role", "Site");
|
||||
SetEnv("ScadaBridge__Node__SiteId", "TestSite");
|
||||
SetEnv("ScadaBridge__Node__NodeHostname", "localhost");
|
||||
SetEnv("ScadaBridge__Node__RemotingPort", "18082");
|
||||
SetEnv("ScadaBridge__Cluster__SeedNodes__0", "akka.tcp://scadabridge@localhost:18082");
|
||||
SetEnv("ScadaBridge__Cluster__SeedNodes__1", "akka.tcp://scadabridge@localhost:18085");
|
||||
SetEnv("ScadaBridge__Communication__GrpcPsk", "test-psk-0123456789");
|
||||
SetEnv("LocalDb__Path", _tempDbPath);
|
||||
}
|
||||
|
||||
private void SetEnv(string key, string? value)
|
||||
{
|
||||
_previousEnv[key] = Environment.GetEnvironmentVariable(key);
|
||||
Environment.SetEnvironmentVariable(key, value);
|
||||
}
|
||||
|
||||
public void Dispose()
|
||||
{
|
||||
foreach (var d in _disposables)
|
||||
{
|
||||
try { d.Dispose(); } catch { /* best effort */ }
|
||||
}
|
||||
|
||||
foreach (var (key, value) in _previousEnv)
|
||||
{
|
||||
Environment.SetEnvironmentVariable(key, value);
|
||||
}
|
||||
|
||||
try { File.Delete(_tempDbPath); } catch { /* best effort */ }
|
||||
|
||||
GC.SuppressFinalize(this);
|
||||
}
|
||||
|
||||
private HttpClient CreateSiteClient()
|
||||
{
|
||||
var factory = new WebApplicationFactory<Program>()
|
||||
.WithWebHostBuilder(builder =>
|
||||
{
|
||||
// WebApplicationFactory defaults the environment to Development, which turns on
|
||||
// ValidateOnBuild/ValidateScopes. The site graph carries scoped registrations whose
|
||||
// dependencies are central-only (ReconcileService → IDeploymentManagerRepository,
|
||||
// AuditLogKpiSampleSource → IAuditLogRepository) and are never resolved on a site
|
||||
// node, so eager validation fails on a composition that runs fine in production.
|
||||
// That is a pre-existing property of the site container, not of the health wiring
|
||||
// under test — run this fixture the way a site node actually runs.
|
||||
builder.UseEnvironment("Production");
|
||||
|
||||
// ConfigureTestServices runs after the app's own registrations, so this drops the
|
||||
// hosted service that would form a real cluster while leaving the AkkaHostedService
|
||||
// singleton resolvable (the site pipeline resolves it for shutdown ordering).
|
||||
builder.ConfigureTestServices(AkkaHostedServiceRemover.RemoveAkkaHostedServiceOnly);
|
||||
});
|
||||
_disposables.Add(factory);
|
||||
|
||||
var client = factory.CreateClient();
|
||||
_disposables.Add(client);
|
||||
return client;
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public async Task Site_MapsHealthReady_WithTheCanonicalJsonBody()
|
||||
{
|
||||
var response = await CreateSiteClient().GetAsync("/health/ready");
|
||||
|
||||
// Never 404. 200 vs 503 depends on cluster state this fixture does not form, so both are
|
||||
// accepted — what is asserted is that the tier is mapped and answers in canonical shape.
|
||||
Assert.NotEqual(HttpStatusCode.NotFound, response.StatusCode);
|
||||
Assert.True(
|
||||
response.StatusCode is HttpStatusCode.OK or HttpStatusCode.ServiceUnavailable,
|
||||
$"Expected 200 or 503, got {(int)response.StatusCode}");
|
||||
|
||||
using var doc = JsonDocument.Parse(await response.Content.ReadAsStringAsync());
|
||||
var entries = doc.RootElement.GetProperty("entries");
|
||||
|
||||
Assert.True(entries.TryGetProperty("akka-cluster", out _), "ready tier must carry akka-cluster");
|
||||
Assert.True(entries.TryGetProperty("localdb", out _), "ready tier must carry localdb");
|
||||
// Readiness must NOT depend on being the primary, or a healthy standby would read unready.
|
||||
Assert.False(entries.TryGetProperty("active-node", out _), "active-node belongs to the Active tier");
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public async Task Site_MapsHealthActive_CarryingOnlyTheActiveNodeCheck()
|
||||
{
|
||||
var response = await CreateSiteClient().GetAsync("/health/active");
|
||||
|
||||
Assert.NotEqual(HttpStatusCode.NotFound, response.StatusCode);
|
||||
|
||||
using var doc = JsonDocument.Parse(await response.Content.ReadAsStringAsync());
|
||||
var entries = doc.RootElement.GetProperty("entries");
|
||||
|
||||
Assert.Equal(
|
||||
new[] { "active-node" },
|
||||
entries.EnumerateObject().Select(p => p.Name).ToArray());
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public async Task Site_HealthEndpointsAreAnonymous()
|
||||
{
|
||||
// The dashboard authenticates nowhere. The site pipeline runs no authentication middleware
|
||||
// today, so this is a regression pin: adding one later must not silently close these.
|
||||
var client = CreateSiteClient();
|
||||
|
||||
foreach (var path in new[] { "/health/ready", "/health/active", "/healthz" })
|
||||
{
|
||||
var response = await client.GetAsync(path);
|
||||
|
||||
Assert.NotEqual(HttpStatusCode.Unauthorized, response.StatusCode);
|
||||
Assert.NotEqual(HttpStatusCode.Forbidden, response.StatusCode);
|
||||
Assert.NotEqual(HttpStatusCode.NotFound, response.StatusCode);
|
||||
}
|
||||
}
|
||||
}
|
||||
+204
@@ -0,0 +1,204 @@
|
||||
using Akka.Actor;
|
||||
using Akka.Cluster;
|
||||
using Akka.Configuration;
|
||||
using Microsoft.Extensions.Logging.Abstractions;
|
||||
using Microsoft.Extensions.Options;
|
||||
using ZB.MOM.WW.ScadaBridge.ClusterInfrastructure;
|
||||
using ZB.MOM.WW.ScadaBridge.Host;
|
||||
using ZB.MOM.WW.ScadaBridge.Host.Actors;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.IntegrationTests.Cluster;
|
||||
|
||||
/// <summary>
|
||||
/// Real-ActorSystem tests for the simultaneous-cold-start split-brain guard (Gitea #33):
|
||||
/// <see cref="ClusterBootstrapCoordinator"/> + <see cref="ClusterBootstrapGuard"/>, driven exactly as
|
||||
/// production drives them — the node's ActorSystem is built from the production <c>BuildHocon</c> output
|
||||
/// with the guard enabled (which emits an EMPTY seed list, so Akka does not auto-join), then the real
|
||||
/// coordinator issues the reachability-gated <c>JoinSeedNodes</c>. This subsystem has a history of bugs
|
||||
/// invisible to pure unit tests, so the load-bearing correctness properties are asserted against running
|
||||
/// nodes, mirroring <see cref="SelfFirstSeedBootstrapTests"/>.
|
||||
/// </summary>
|
||||
/// <remarks>
|
||||
/// Ephemeral loopback ports share host <c>127.0.0.1</c>, so the port breaks the guard's ordinal
|
||||
/// <c>host:port</c> tie — <c>Min(port)</c> is the founder, <c>Max(port)</c> is the higher (probing) node.
|
||||
/// </remarks>
|
||||
public sealed class ClusterBootstrapCoordinatorTests : IAsyncLifetime
|
||||
{
|
||||
private readonly List<ActorSystem> _systems = new();
|
||||
private readonly List<ClusterBootstrapCoordinator> _coordinators = new();
|
||||
|
||||
/// <summary>Starts a node with the bootstrap guard ENABLED, wired exactly as the composition roots
|
||||
/// do: the ActorSystem is built with an empty HOCON seed list (guard on) and the real coordinator
|
||||
/// issues the reachability-gated join. Seeds are listed self-first (as every shipped config is); the
|
||||
/// guard, not the order, decides founder vs. joiner.</summary>
|
||||
private async Task<ActorSystem> StartGuardedNodeAsync(int selfPort, int peerPort, int probeSeconds = 4)
|
||||
{
|
||||
var self = $"akka.tcp://scadabridge@127.0.0.1:{selfPort}";
|
||||
var peer = $"akka.tcp://scadabridge@127.0.0.1:{peerPort}";
|
||||
var nodeOptions = new NodeOptions { Role = "Central", NodeHostname = "127.0.0.1", RemotingPort = selfPort };
|
||||
var clusterOptions = new ClusterOptions
|
||||
{
|
||||
SeedNodes = new List<string> { self, peer },
|
||||
BootstrapGuard = new ClusterBootstrapGuardOptions
|
||||
{
|
||||
Enabled = true,
|
||||
PartnerProbeSeconds = probeSeconds,
|
||||
PartnerProbeIntervalMs = 250,
|
||||
ProbeConnectTimeoutMs = 500,
|
||||
},
|
||||
StableAfter = TimeSpan.FromSeconds(3),
|
||||
HeartbeatInterval = TimeSpan.FromMilliseconds(500),
|
||||
FailureDetectionThreshold = TimeSpan.FromSeconds(2),
|
||||
MinNrOfMembers = 1,
|
||||
};
|
||||
|
||||
var hocon = AkkaHostedService.BuildHocon(
|
||||
nodeOptions, clusterOptions, new[] { "Central" },
|
||||
TimeSpan.FromSeconds(1), TimeSpan.FromSeconds(3));
|
||||
var system = ActorSystem.Create("scadabridge", ConfigurationFactory.ParseString(hocon));
|
||||
_systems.Add(system);
|
||||
|
||||
var coordinator = new ClusterBootstrapCoordinator(
|
||||
() => system, Options.Create(clusterOptions), Options.Create(nodeOptions),
|
||||
NullLogger<ClusterBootstrapCoordinator>.Instance);
|
||||
_coordinators.Add(coordinator);
|
||||
await coordinator.StartAsync(CancellationToken.None);
|
||||
return system;
|
||||
}
|
||||
|
||||
private static async Task<bool> WaitForUpAsync(ActorSystem system, int expected, TimeSpan timeout)
|
||||
{
|
||||
var cluster = Akka.Cluster.Cluster.Get(system);
|
||||
var deadline = DateTime.UtcNow + timeout;
|
||||
while (DateTime.UtcNow < deadline)
|
||||
{
|
||||
if (cluster.State.Members.Count(m => m.Status == MemberStatus.Up) >= expected) return true;
|
||||
await Task.Delay(200);
|
||||
}
|
||||
return cluster.State.Members.Count(m => m.Status == MemberStatus.Up) >= expected;
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// The FOUNDER (lower address) with the guard on forms a cluster alone when its partner is dead —
|
||||
/// exactly as un-guarded self-first does. The guard must not regress the preferred founder.
|
||||
/// </summary>
|
||||
[Fact]
|
||||
public async Task Founder_forms_alone_when_partner_is_dead()
|
||||
{
|
||||
var low = TwoNodeClusterFixture.GetFreeTcpPort();
|
||||
var high = TwoNodeClusterFixture.GetFreeTcpPort();
|
||||
var founderPort = Math.Min(low, high);
|
||||
var deadPeerPort = Math.Max(low, high); // nothing listening
|
||||
|
||||
var system = await StartGuardedNodeAsync(founderPort, deadPeerPort);
|
||||
|
||||
Assert.True(await WaitForUpAsync(system, 1, TimeSpan.FromSeconds(30)),
|
||||
"the founder (lower address) must self-form immediately when its partner is down");
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// THE LOAD-BEARING CASE. The HIGHER node with the guard on must still cold-start ALONE when its
|
||||
/// partner is genuinely dead: it probes the (dead) founder, times out, and falls back to self-first.
|
||||
/// If the guard's peer-first logic were wrong, this node would hang in JoinSeedNodeProcess forever —
|
||||
/// the very regression a naive "always let the lower node found" rule would introduce.
|
||||
/// </summary>
|
||||
[Fact]
|
||||
public async Task Higher_node_forms_alone_when_partner_is_dead_after_probing()
|
||||
{
|
||||
var low = TwoNodeClusterFixture.GetFreeTcpPort();
|
||||
var high = TwoNodeClusterFixture.GetFreeTcpPort();
|
||||
var higherPort = Math.Max(low, high);
|
||||
var deadFounderPort = Math.Min(low, high); // the founder is dead
|
||||
|
||||
var system = await StartGuardedNodeAsync(higherPort, deadFounderPort, probeSeconds: 3);
|
||||
|
||||
// Must come Up AFTER the probe window (~3s) expires and it falls back to self-first.
|
||||
Assert.True(await WaitForUpAsync(system, 1, TimeSpan.FromSeconds(30)),
|
||||
"the higher node must fall back to self-first and form alone once the dead partner never answers the probe");
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// FALSIFIABILITY CONTROL for the test above: the higher node does NOT form alone during the probe
|
||||
/// window — it is genuinely waiting/probing, not self-forming immediately (which would mean the
|
||||
/// guard is inert). If this ever starts seeing a member before the window, the probe is not gating.
|
||||
/// </summary>
|
||||
[Fact]
|
||||
public async Task Higher_node_does_not_form_before_the_probe_window_expires()
|
||||
{
|
||||
var low = TwoNodeClusterFixture.GetFreeTcpPort();
|
||||
var high = TwoNodeClusterFixture.GetFreeTcpPort();
|
||||
var higherPort = Math.Max(low, high);
|
||||
var deadFounderPort = Math.Min(low, high);
|
||||
|
||||
var system = await StartGuardedNodeAsync(higherPort, deadFounderPort, probeSeconds: 10);
|
||||
|
||||
// Well within the 10s probe window: still no cluster, because it is probing the dead founder.
|
||||
Assert.False(await WaitForUpAsync(system, 1, TimeSpan.FromSeconds(3)),
|
||||
"the higher node must probe for the founder, not self-form immediately");
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// The higher node JOINS a live founder rather than forming a second cluster: start the founder,
|
||||
/// then the higher node — its probe finds the founder reachable, it commits peer-first, and the
|
||||
/// pair converges to ONE cluster of two.
|
||||
/// </summary>
|
||||
[Fact]
|
||||
public async Task Higher_node_joins_a_live_founder()
|
||||
{
|
||||
var low = TwoNodeClusterFixture.GetFreeTcpPort();
|
||||
var high = TwoNodeClusterFixture.GetFreeTcpPort();
|
||||
var founderPort = Math.Min(low, high);
|
||||
var higherPort = Math.Max(low, high);
|
||||
|
||||
var founder = await StartGuardedNodeAsync(founderPort, higherPort);
|
||||
Assert.True(await WaitForUpAsync(founder, 1, TimeSpan.FromSeconds(30)),
|
||||
"the founder must be Up before the higher node probes it");
|
||||
|
||||
var higher = await StartGuardedNodeAsync(higherPort, founderPort);
|
||||
|
||||
Assert.True(await WaitForUpAsync(higher, 2, TimeSpan.FromSeconds(60)),
|
||||
"the higher node must join the founder, forming one cluster of two");
|
||||
Assert.True(await WaitForUpAsync(founder, 2, TimeSpan.FromSeconds(60)),
|
||||
"the founder must see the higher node as a member of ITS cluster");
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// THE HEADLINE #33 CASE. Both nodes cold-start truly simultaneously with the guard on. Without the
|
||||
/// guard each would run FirstSeedNodeProcess, race, and form its own 1-node cluster (split brain).
|
||||
/// With it, the founder forms and the higher node probes-then-joins, so the pair converges to
|
||||
/// EXACTLY ONE cluster of two — no two oldest, no dual singletons.
|
||||
/// </summary>
|
||||
[Fact]
|
||||
public async Task Both_nodes_cold_starting_together_form_exactly_one_cluster()
|
||||
{
|
||||
var low = TwoNodeClusterFixture.GetFreeTcpPort();
|
||||
var high = TwoNodeClusterFixture.GetFreeTcpPort();
|
||||
var founderPort = Math.Min(low, high);
|
||||
var higherPort = Math.Max(low, high);
|
||||
|
||||
// Start both without serialization — as a shared power event does.
|
||||
var founderTask = StartGuardedNodeAsync(founderPort, higherPort);
|
||||
var higherTask = StartGuardedNodeAsync(higherPort, founderPort);
|
||||
var founder = await founderTask;
|
||||
var higher = await higherTask;
|
||||
|
||||
Assert.True(await WaitForUpAsync(founder, 2, TimeSpan.FromSeconds(60)),
|
||||
"the pair must converge to one 2-member cluster, not two 1-member clusters");
|
||||
Assert.True(await WaitForUpAsync(higher, 2, TimeSpan.FromSeconds(60)),
|
||||
"the higher node must be in the SAME cluster as the founder");
|
||||
}
|
||||
|
||||
public Task InitializeAsync() => Task.CompletedTask;
|
||||
|
||||
public async Task DisposeAsync()
|
||||
{
|
||||
foreach (var c in _coordinators)
|
||||
{
|
||||
try { await c.StopAsync(CancellationToken.None); } catch { /* teardown */ }
|
||||
}
|
||||
foreach (var s in _systems)
|
||||
{
|
||||
try { await s.Terminate().WaitAsync(TimeSpan.FromSeconds(10)); } catch { /* teardown */ }
|
||||
}
|
||||
}
|
||||
}
|
||||
+219
@@ -0,0 +1,219 @@
|
||||
using System.Data.Common;
|
||||
using Microsoft.AspNetCore.DataProtection;
|
||||
using Microsoft.EntityFrameworkCore;
|
||||
using Microsoft.EntityFrameworkCore.Diagnostics;
|
||||
using Microsoft.EntityFrameworkCore.Storage;
|
||||
using Microsoft.Extensions.Configuration;
|
||||
using Microsoft.Extensions.DependencyInjection;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Entities.Scripts;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Entities.Templates;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Interfaces.Repositories;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Interfaces.Services;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Interfaces.Transport;
|
||||
using ZB.MOM.WW.ScadaBridge.Commons.Types.Transport;
|
||||
using ZB.MOM.WW.ScadaBridge.ConfigurationDatabase;
|
||||
using ZB.MOM.WW.ScadaBridge.ConfigurationDatabase.Repositories;
|
||||
using ZB.MOM.WW.ScadaBridge.ConfigurationDatabase.Services;
|
||||
using ZB.MOM.WW.ScadaBridge.DeploymentManager;
|
||||
using ZB.MOM.WW.ScadaBridge.TemplateEngine;
|
||||
using ZB.MOM.WW.ScadaBridge.Transport;
|
||||
using ZB.MOM.WW.ScadaBridge.Transport.Import;
|
||||
|
||||
namespace ZB.MOM.WW.ScadaBridge.Transport.IntegrationTests.Import;
|
||||
|
||||
/// <summary>
|
||||
/// Regression guard for Gitea #28: <see cref="BundleImporter.ApplyAsync"/> must
|
||||
/// run its multi-<c>SaveChanges</c> apply through the context's execution
|
||||
/// strategy, because the production central <c>ScadaBridgeDbContext</c> is
|
||||
/// configured with <c>EnableRetryOnFailure</c> and
|
||||
/// <c>SqlServerRetryingExecutionStrategy</c> REJECTS a user-initiated
|
||||
/// transaction opened outside the strategy
|
||||
/// (<c>CoreStrings.ExecutionStrategyExistingTransaction</c>) — so before the
|
||||
/// fix every bundle import threw on any real SQL Server.
|
||||
/// <para>
|
||||
/// The rest of the Transport suite runs on the in-memory provider, which has NO
|
||||
/// retrying strategy and where <c>BeginTransactionAsync</c> is a no-op, so it
|
||||
/// could never surface this. Here we run on SQLite (real relational
|
||||
/// transactions) AND configure a retrying execution strategy whose only job is
|
||||
/// to have <c>RetriesOnFailure == true</c>, which arms the exact same guard as
|
||||
/// the production SQL Server strategy. If <c>ApplyAsync</c> opened the
|
||||
/// transaction directly, this test would throw before importing anything.
|
||||
/// </para>
|
||||
/// <para>
|
||||
/// Reuses <see cref="SqliteCompatibleScadaBridgeDbContext"/> defined alongside
|
||||
/// <see cref="BundleImporterRollbackFailureTests"/> in this assembly.
|
||||
/// </para>
|
||||
/// </summary>
|
||||
public sealed class BundleImporterRetryingStrategyTests : IDisposable
|
||||
{
|
||||
private readonly ServiceProvider _provider;
|
||||
private readonly DbConnection _sharedConnection;
|
||||
|
||||
public BundleImporterRetryingStrategyTests()
|
||||
{
|
||||
var services = new ServiceCollection();
|
||||
services.AddSingleton<IConfiguration>(
|
||||
new ConfigurationBuilder().AddInMemoryCollection().Build());
|
||||
|
||||
// SQLite :memory: is per-connection, so pin one open connection for the
|
||||
// whole fixture or the schema vanishes between DbContext instances.
|
||||
_sharedConnection = new Microsoft.Data.Sqlite.SqliteConnection("DataSource=:memory:");
|
||||
_sharedConnection.Open();
|
||||
|
||||
services.AddSingleton<IDataProtectionProvider>(new EphemeralDataProtectionProvider());
|
||||
|
||||
services.AddSingleton(sp =>
|
||||
{
|
||||
var builder = new DbContextOptionsBuilder<ScadaBridgeDbContext>();
|
||||
// Configure a RETRYING execution strategy — this is the crux of the
|
||||
// regression. RetriesOnFailure == true (MaxRetryCount > 0) makes EF
|
||||
// reject any user-initiated BeginTransaction opened outside
|
||||
// strategy.ExecuteAsync, exactly like the production SQL Server
|
||||
// EnableRetryOnFailure configuration.
|
||||
builder.UseSqlite(
|
||||
_sharedConnection,
|
||||
sqlite => sqlite.ExecutionStrategy(deps => new TestRetryingExecutionStrategy(deps)));
|
||||
builder.ConfigureWarnings(w => w.Ignore(RelationalEventId.PendingModelChangesWarning));
|
||||
return builder.Options;
|
||||
});
|
||||
services.AddScoped<ScadaBridgeDbContext>(sp => new SqliteCompatibleScadaBridgeDbContext(
|
||||
sp.GetRequiredService<DbContextOptions<ScadaBridgeDbContext>>(),
|
||||
sp.GetRequiredService<IDataProtectionProvider>()));
|
||||
|
||||
services.AddScoped<ITemplateEngineRepository, TemplateEngineRepository>();
|
||||
services.AddScoped<IExternalSystemRepository, ExternalSystemRepository>();
|
||||
services.AddScoped<INotificationRepository, NotificationRepository>();
|
||||
services.AddScoped<IInboundApiRepository, InboundApiRepository>();
|
||||
services.AddScoped<ISiteRepository, SiteRepository>();
|
||||
services.AddScoped<IDeploymentManagerRepository, DeploymentManagerRepository>();
|
||||
services.AddScoped<ISharedSchemaRepository, SharedSchemaRepository>();
|
||||
services.AddScoped<IAuditCorrelationContext, AuditCorrelationContext>();
|
||||
services.AddScoped<IAuditService, AuditService>();
|
||||
services.AddTransport();
|
||||
// IStaleInstanceProbe (DeploymentManager) + flatten/hash (TemplateEngine)
|
||||
// so ApplyAsync's StaleInstanceIds computation resolves a real probe.
|
||||
services.AddTemplateEngine();
|
||||
services.AddDeploymentManager();
|
||||
|
||||
_provider = services.BuildServiceProvider();
|
||||
|
||||
using var scope = _provider.CreateScope();
|
||||
var ctx = scope.ServiceProvider.GetRequiredService<ScadaBridgeDbContext>();
|
||||
ctx.Database.EnsureCreated();
|
||||
}
|
||||
|
||||
public void Dispose()
|
||||
{
|
||||
_provider.Dispose();
|
||||
_sharedConnection.Dispose();
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public async Task ApplyAsync_succeeds_under_a_retrying_execution_strategy()
|
||||
{
|
||||
// Arrange: seed → export → wipe → apply(Add). Same shape as
|
||||
// BundleImporterApplyTests.ApplyAsync_adds_new_artifacts_in_single_transaction,
|
||||
// but the fixture's context carries a retrying execution strategy.
|
||||
await using (var scope = _provider.CreateAsyncScope())
|
||||
{
|
||||
var ctx = scope.ServiceProvider.GetRequiredService<ScadaBridgeDbContext>();
|
||||
ctx.SharedScripts.Add(new SharedScript("HelperFn", "return 1;"));
|
||||
ctx.Templates.Add(new Template("Pump") { Description = "fresh" });
|
||||
await ctx.SaveChangesAsync();
|
||||
}
|
||||
|
||||
var sessionId = await ExportAndLoadAsync();
|
||||
await WipeContentAsync();
|
||||
|
||||
// Act: before the fix this threw CoreStrings.ExecutionStrategyExistingTransaction
|
||||
// ("... does not support user-initiated transactions.") from the direct
|
||||
// BeginTransactionAsync. With the fix the apply runs inside the strategy.
|
||||
ImportResult result;
|
||||
await using (var scope = _provider.CreateAsyncScope())
|
||||
{
|
||||
var importer = scope.ServiceProvider.GetRequiredService<IBundleImporter>();
|
||||
var resolutions = new List<ImportResolution>
|
||||
{
|
||||
new("Template", "Pump", ResolutionAction.Add, null),
|
||||
new("SharedScript", "HelperFn", ResolutionAction.Add, null),
|
||||
};
|
||||
result = await importer.ApplyAsync(sessionId, resolutions, user: "bob");
|
||||
}
|
||||
|
||||
// Assert: the whole transactional apply committed.
|
||||
Assert.Equal(2, result.Added);
|
||||
Assert.Equal(0, result.Overwritten);
|
||||
Assert.NotEqual(Guid.Empty, result.BundleImportId);
|
||||
await using (var scope = _provider.CreateAsyncScope())
|
||||
{
|
||||
var ctx = scope.ServiceProvider.GetRequiredService<ScadaBridgeDbContext>();
|
||||
Assert.Equal(1, await ctx.Templates.CountAsync(t => t.Name == "Pump"));
|
||||
Assert.Equal(1, await ctx.SharedScripts.CountAsync(s => s.Name == "HelperFn"));
|
||||
// The BundleImported audit row landed inside the committed transaction.
|
||||
Assert.Equal(1, await ctx.AuditLogEntries.CountAsync(a => a.Action == "BundleImported"));
|
||||
}
|
||||
}
|
||||
|
||||
// ---- helpers (mirror BundleImporterApplyTests) ----
|
||||
|
||||
private async Task<Guid> ExportAndLoadAsync()
|
||||
{
|
||||
Stream bundleStream;
|
||||
await using (var scope = _provider.CreateAsyncScope())
|
||||
{
|
||||
var exporter = scope.ServiceProvider.GetRequiredService<IBundleExporter>();
|
||||
var ctx = scope.ServiceProvider.GetRequiredService<ScadaBridgeDbContext>();
|
||||
var templateIds = await ctx.Templates.Select(t => t.Id).ToListAsync();
|
||||
var sharedScriptIds = await ctx.SharedScripts.Select(s => s.Id).ToListAsync();
|
||||
var selection = new ExportSelection(
|
||||
TemplateIds: templateIds,
|
||||
SharedScriptIds: sharedScriptIds,
|
||||
ExternalSystemIds: Array.Empty<int>(),
|
||||
DatabaseConnectionIds: Array.Empty<int>(),
|
||||
NotificationListIds: Array.Empty<int>(),
|
||||
SmtpConfigurationIds: Array.Empty<int>(),
|
||||
ApiMethodIds: Array.Empty<int>(),
|
||||
IncludeDependencies: false);
|
||||
bundleStream = await exporter.ExportAsync(selection, user: "alice", sourceEnvironment: "dev",
|
||||
passphrase: null, cancellationToken: CancellationToken.None);
|
||||
}
|
||||
|
||||
using var ms = new MemoryStream();
|
||||
await bundleStream.CopyToAsync(ms);
|
||||
ms.Position = 0;
|
||||
|
||||
await using var loadScope = _provider.CreateAsyncScope();
|
||||
var importer = loadScope.ServiceProvider.GetRequiredService<IBundleImporter>();
|
||||
var session = await importer.LoadAsync(ms, passphrase: null);
|
||||
return session.SessionId;
|
||||
}
|
||||
|
||||
private async Task WipeContentAsync()
|
||||
{
|
||||
await using var scope = _provider.CreateAsyncScope();
|
||||
var ctx = scope.ServiceProvider.GetRequiredService<ScadaBridgeDbContext>();
|
||||
ctx.Templates.RemoveRange(ctx.Templates);
|
||||
ctx.SharedScripts.RemoveRange(ctx.SharedScripts);
|
||||
ctx.TemplateFolders.RemoveRange(ctx.TemplateFolders);
|
||||
await ctx.SaveChangesAsync();
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// A minimal retrying execution strategy for the test. Its sole purpose is
|
||||
/// to report <c>RetriesOnFailure == true</c> (inherited from
|
||||
/// <see cref="ExecutionStrategy"/>, since <c>MaxRetryCount > 0</c>), which
|
||||
/// arms EF's user-initiated-transaction guard just like the production
|
||||
/// <c>SqlServerRetryingExecutionStrategy</c>. It never actually retries —
|
||||
/// <see cref="ShouldRetryOn"/> returns false — so a genuine fault surfaces
|
||||
/// immediately instead of looping.
|
||||
/// </summary>
|
||||
private sealed class TestRetryingExecutionStrategy : ExecutionStrategy
|
||||
{
|
||||
public TestRetryingExecutionStrategy(ExecutionStrategyDependencies dependencies)
|
||||
: base(dependencies, maxRetryCount: 3, maxRetryDelay: TimeSpan.FromMilliseconds(1))
|
||||
{
|
||||
}
|
||||
|
||||
protected override bool ShouldRetryOn(Exception exception) => false;
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user