Files
ScadaBridge/docs/plans/2026-07-08-deferred-work-register.md
T

59 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Deferred-Work Register (established 2026-07-08, from architecture review 08)
Single tracked list of consciously-deferred work. Rules: every deferral gets a row
(rationale + revisit trigger); fix-now items reference the archreview plan that owns
them and are removed from this table when that plan's task lands.
## Fix-now (owned by archreview plans)
All 7 fix-now items landed via PLAN-04/05/06/07/08 (verified in review 08 round 2, 2026-07-12) — rows removed per the rule above; see the round-2 report §1 for the per-item evidence.
## Deferred (with rationale + revisit trigger)
| # | Item | Where noted | Rationale for deferral | Revisit trigger |
|---|------|-------------|------------------------|-----------------|
| 8 | Hash-chain tamper evidence (T1); CLI verify-chain is a no-op stub | audit-log roadmap :12 | v1.x by locked decision; append-only DB roles are the control | Compliance requirement for cryptographic tamper evidence |
| 9 | Parquet audit archival (T2); endpoint returns 501 | AuditEndpoints.cs:204 | v1.x; 501 + CLI messaging are honest | AuditLog partition volume nears retention ceiling |
| 11 | Central-persisted OPC UA cert-trust audit | m7 follow-ups | Broadcast-to-both-nodes covers HA | Governance/audit requirement for trust decisions |
| 17 | Unified notifications+site-calls outbox page | stillpending :118 | Explicit M9 decision to keep two pages | Operator confusion reports |
| 19 | Bundle signing / cluster-to-cluster pull / differential bundles | transport-design :402 | v1 manifest hash + AES-GCM held sufficient | Non-repudiation requirement across orgs |
| 23 | Live LDAP group-membership re-query for an active session | `docs/requirements/Component-Security.md` :61-69 (+ :78-79) | Blocked on an external package. The mid-session refresh re-maps the **stored** groups against the central DB with **no LDAP call**, so a directory group-membership change lands only at next login. A live re-query needs a passwordless service-account group-search method on the shared `ZB.MOM.WW.Auth.Ldap` library — an external NuGet `PackageReference` (`src/ZB.MOM.WW.ScadaBridge.Security/…csproj:23`) exposing only `AuthenticateAsync(username, password, ct)`. Central role-mapping/scope changes still apply within ~15 min (`RoleRefreshThresholdMinutes`). | `ZB.MOM.WW.Auth.Ldap` gains a standalone group-search API, or a requirement that a directory-side group revocation take effect mid-session rather than at next login |
| 24 | M8 large-bundle performance hardening | `docs/plans/2026-06-15-stillpending-completion-design.md:106` — "Small follow-ups logged (not blocking): … large-bundle/perf hardening" | Logged as a non-blocking follow-up when M8 shipped and never given an artifact: **no plan, no task entry, no perf/load test exists** (`tests/…Transport.Tests/Import/BundleImporterLoadTests.cs` is a `LoadAsync` unit suite despite the name). No measured problem; the only sizing controls in place are the 5-minute CLI transport timeout, `LineDiffer`'s `MaxInputLines`=4000 summary-only cap, and `MaxConcurrentImportSessions`=8. | First real bundle that times out, exhausts memory, or makes the import wizard's diff step unusable |
| 25 | Phase-8 WP-4 target-scale load test (10 sites × 500 instances × 75 tags = 37,500 subscriptions/site, 375,000 total) | `docs/plans/phase-8-production-readiness.md:152-170` (WP-4) + `:314-320` (test protocol); status claimed in `docs/plans/phase-8-checklist.md` | **Claimed complete but unevidenced.** The whole WP-4 deliverable is a **107-byte** checklist stub asserting "Status: Complete / Tests: All passing / Build: 0 errors, 0 warnings" with no per-work-package results and no linked run. Nearest real coverage is arithmetic/aggregation only — `PerformanceTests/StaggeredStartupTests.cs` (`TagCapacity_75TagsPer500Machines_37500Total`, 500-instances-over-10-sites distribution) and `HealthAggregationTests` (10-site report aggregation) — plus a **single-subscriber** 100k-event `Streaming/SiteStreamThroughputTests.cs`. No sustained multi-site run exists anywhere in `tests/` or `docker/`. | Before any production go-live at target scale; or the first site approaching ~500 instances / ~37.5k subscriptions |
| 26 | Ipsen MES MoveIn tail: leak-test (`-LT`) receivers + routing, PLC-output-flag writes, `Z28062` BTDB data completeness | `docs/plans/2026-06-16-ipsen-mes-movein.md:409` ("Out of scope (future)"); design `2026-06-16-ipsen-mes-movein-design.md:58-60, :196-198` | Customer-site scope, not a platform gap. `-LT` routing needs an MES-receiver child + Galaxy reference that do not exist on the reactor template (any `-LT`/unknown suffix returns `WasSuccessful=false` with an "unsupported side/target" message by decision); `MoveInComplete`/`Successful`/`ErrorText` are **PLC-owned** by locked decision, so ScadaBridge deliberately does not write them; `Z28062` completeness is an operational data fix, not code. Note the separate alarm-status path already handles the suffix — `_LT` is stripped before side-scoping (`2026-06-30-mes-alarm-status-api.md:158`). | Ipsen creates the leak-test receiver + Galaxy reference, or asks ScadaBridge to own the PLC-output flags — otherwise a **candidate won't-do** (`[PERM]`) at the next Ipsen scope review |
| 27 | External-system per-system retry config (`MaxRetries`/`RetryDelay`) never reaches sites, and has no CLI/management surface | Found live 2026-08-01 (rig session, #11 gRPC live checks) | Two stacked gaps: (a) `ExternalSystemArtifact` (Commons) carries `TimeoutSeconds` but NOT `MaxRetries`/`RetryDelay`, and the site `external_systems` table has no such columns — so a centrally-configured retry policy is silently ignored on sites; every cached call buffers with the S&F default (`DefaultMaxRetries` 50 × `DefaultRetryInterval` 30s ≈ 25 min to park). (b) `Create/UpdateExternalSystemCommand` don't expose the fields either — the only way to set them today is a direct DB edit of `ExternalSystemDefinitions`. Transport bundles DO carry them (arch-review 05 "ES retry config"), which masks the gap in export/import round-trips. Fix is additive: extend the artifact + site schema + apply path, and add `--max-retries`/`--retry-delay` to the CLI. | First operator who tunes retry policy on an external system and expects site cached calls to honor it |
| 28 | Health-dashboard "Trigger failover" confirm dialog's confirm button is labeled **"Delete"** | Found live 2026-08-01 (rig session, #11 TriggerSiteFailover check) | The DialogService confirmation host's default destructive-action label leaks through — the dialog copy is correct but the red confirm button says "Delete" for a failover. One-line fix: pass an explicit confirm label ("Fail over") at the Health-dashboard call site (and audit other confirm-dialog call sites for the same default). | Next Central UI session |
## Resolved (verified against the code 2026-07-10)
Rows removed from the Deferred table above once confirmed shipped. Kept here for traceability.
| # | Item | Resolution |
|---|------|-----------|
| 10 | Aggregated live alarm stream for Alarm Summary | Shipped 2026-07-10 (`docs/plans/2026-07-10-aggregated-live-alarm-stream-plan.md`): a **transient, in-memory** per-site central live alarm cache (`ISiteAlarmLiveCache`/`SiteAlarmLiveCacheService` + per-site `SiteAlarmAggregatorActor`) fed by a new site-wide, alarm-only `SubscribeSite` gRPC stream (`SiteStreamManager.SubscribeSiteAlarms`), seed-then-stream with dedup + NodeA↔NodeB re-seed + periodic reconcile. Alarm Summary is now live-cache-driven (`AlarmSummaryService.BuildFromLiveAlarms`) with the 15s poll retained as fallback + `NotReporting` authority. Honors the `[PERM]` no-central-store rule — nothing persisted (no EF table/migration). Options on `CommunicationOptions` (eagerly validated) + two `ScadaBridgeTelemetry` signals. |
| 7 | SecuredWrite audit rows leave SourceNode NULL | Resolved (PLAN-07): `ManagementActor.EmitSecuredWriteAuditAsync` routes through `ICentralAuditWriter`, which stamps `SourceNode` (`central-a`/`central-b`) from `INodeIdentityProvider`. |
| 12 (CLI/API) | Native-alarm-source-override CSV import | Shipped 2026-07-10: shared `CsvLineSplitter`, `NativeAlarmSourceOverrideCsvParser`, bulk all-or-nothing `SetInstanceNativeAlarmSourceOverridesCommand` + ManagementActor handler (Deployer-gated), CLI `instance native-alarm-source import --file`, parser/CLI/handler tests. **UI upload affordance shipped 2026-08-01** — second `InputFile` on the `InstanceConfigure` Native Alarm Source Overrides card reusing the shared parser, mirroring the attribute importer's UX and the server's all-or-nothing merge semantics (`InstanceConfigureNativeAlarmCsvImportTests`); row 12 removed from the Deferred table. |
| 13 | WaitForAttribute quality-gated ("Good"-only) mode | Already implemented (Commons `WaitForAttribute.RequireGoodQuality`, enforced in `InstanceActor`, threaded through `ScriptRuntimeContext`, tested in `InstanceActorWaitForAttributeTests`). Stale "planned enhancement" doc line corrected 2026-07-10. |
| 14 | WaitForAttribute in Test-Run sandbox | Shipped 2026-07-10 (full fidelity): sandbox `Attributes.WaitAsync`/`WaitForAsync` (value-equality) route to the bound instance via `ISandboxInstanceGateway.WaitForAttributeAsync` → the existing `CommunicationService.RouteToWaitForAttributeAsync` cross-site route. Additive `RouteToWaitForAttributeRequest.RequireGoodQuality` (honored by the site handler) makes quality-gated waits route too. **Predicate-form waits stay unsupported** (an in-process lambda can't be routed) and throw a labelled `ScriptSandboxException`. Tests: sandbox accessor routing (CentralUI), site-handler quality-flag threading (SiteRuntime). |
| 15 | BrowseNext final-page signal not surfaced | Already surfaced (M7 browse work): `RealOpcUaClient` sets `Truncated=false`/`ContinuationToken=null` on the last page; `BrowseNodeResult` carries both; `TreeRow.razor` renders "Load more" only when a continuation token remains — no wasted BrowseNext. |
| 16 | StubOpcUaClient throws on browse | Already resolved: `StubOpcUaClient` supports browse + address-space search, covered by `StubOpcUaClientBrowseTests`/`StubOpcUaClientSearchTests`. |
| 18 | Folder drag-drop | Closed — **permanently deferred (`[PERM]`)** by the M9 decision (`docs/plans/2026-06-15-stillpending-completion-design.md:122`): menu-based reorder (T23) shipped instead, and the folder-hierarchy design fixed the reorganization UX as "right-click context menus only (no drag-drop)" (`2026-05-11-templates-folder-hierarchy-design.md:27`). Row removed from the Deferred table 2026-08-01 — nothing left to revisit. |
| 20 | Deployment EXPIRED-row purge | Already resolved (PLAN-04): `PendingDeploymentPurgeActor` central singleton (spawned in `AkkaHostedService`) ticks `IDeploymentManagerRepository.PurgeExpiredPendingDeploymentsAsync` every `CommunicationOptions.PendingDeploymentPurgeInterval` (default 1h), options-validated, tested. |
| 21 | SiteAuditBacklogReporter threshold consolidation | Shipped 2026-07-10: `SqliteAuditWriterOptions.BacklogPollIntervalSeconds` (default 30) now drives the reporter's poll cadence; explicit ctor override still wins (tests), non-positive falls back to the 30 s default. Cadence tests added; stale "hard-code / follow-up" class-doc corrected. |
| 22 | KPI history hourly rollups | Shipped 2026-07-10 (`docs/plans/2026-07-10-kpi-history-hourly-rollups-plan.md`, T1T8): new `KpiRollupHourly` table (migration `20260710153953`) folded by a third recorder tick (`kpi-rollup`, `RollupInterval` default 1h) over a re-folded `RollupLookbackHours` window via an idempotent, failover-self-healing upsert; per-metric gauge-vs-rate aggregation (`KpiMetricAggregationCatalog`); a one-shot backfill of the retention window on start; raw-vs-rollup query routing by `RollupThresholdHours` (default 168h); longer rollup retention (`RollupRetentionDays` default 365 ≥ `RetentionDays`, dual daily purge); and 30 d / 90 d trend windows added to the four surfaces. Options + validator, docs (`Component-KpiHistory.md`), and tests shipped. |
## New deferrals from review 08 (this plan)
| Item | Rationale | Revisit trigger |
|------|-----------|-----------------|
| Communication → HealthMonitoring layering (ICentralHealthAggregator consumed by CentralCommunicationActor.cs:351) | Moving the interface + SiteHealthState to Commons ripples across 5 projects for a cosmetic inversion | Next breaking change to ICentralHealthAggregator |
| docs/components reference docs for ScriptAnalysis, KpiHistory, DelmiaNotifier | Reference docs are substantial (StyleGuide-conformant); README claim scoped instead (PLAN-08 Task 10) | Next doc-writing session touching those components |
| Test-coverage backfill: SiteCallAudit.Tests (31 tests/1.6k LOC), DeploymentManager.Tests | No defect identified; coverage partly lives in ManagementService/Host/Integration suites | First regression escaping either component |
| ~~Failover-timing measurement (the "~25s total failover" envelope)~~ **RESOLVED 2026-08-01** — split out of the combined row and closed. `tests/ZB.MOM.WW.ScadaBridge.PerformanceTests/Failover/FailoverTimingTests.cs` is no longer a skipped placeholder: it runs as a live `[Fact]` (`Category=Performance`) on the real two-node in-process rig (`TwoNodeClusterFixture`, production `BuildHocon`) at production timings — 2s heartbeat / 10s failure-detection threshold / 15s stable-after — hard-killing the younger node and timing the survivor's member REMOVAL with singleton continuity asserted on the oldest. Delivered by **PLAN-R2-01 Task 4** (`archreview/plans/PLAN-R2-01-cluster-host-failover.md:226`). The oldest-crash direction is covered behaviorally by `SbrFailoverTests.AutoDown_HardCrashOfOldestNode_*` and by `docker/failover-drill.sh`. | The 2026-07-08 "PLAN-01 rig landing" trigger had fired unnoticed (NF2); PLAN-R2-01 T4 wired the placeholder to the fixture rig rather than recording a blocker. | Closed. |
| Broader perf envelope — **S&F drain rate + per-subscriber stream backpressure** (the still-open half of the former combined row) | Never measured, and no owner plan survives now that PLAN-R2-01 closed the failover half. `PerformanceTests` covers failover timing, staggered startup, health aggregation, audit hot-path latency and a **single-subscriber** 100k-event `Streaming/SiteStreamThroughputTests.cs` — nothing measures store-and-forward drain throughput, nor what a slow/stalled subscriber does to the per-subscriber buffering in `Communication/Actors/StreamRelayActor.cs` / `Grpc/SiteStreamGrpcServer.cs` with many subscribers attached. No defect observed; deferred as measurement-only work. | First field S&F backlog that fails to drain within an operator's patience, a slow gRPC subscriber degrading a site stream for others, or the WP-4 target-scale run (row 25) being scheduled — that run should absorb this |
## Deferred — operational risk (from the initiative tracker, folded in 2026-07-12)
Two live items previously tracked ONLY in `archreview/plans/00-MASTER-TRACKER.md`'s registry are folded in here (NF5) so this register is the single tracking place. The tracker's narrative subsections remain as the historical evidence.
| # | Item | Where noted | Rationale for deferral | Revisit trigger |
|---|------|-------------|------------------------|-----------------|
| SBR | ~~**SBR oldest-crash total-outage gap**~~ **RESOLVED 2026-07-21 (owner decision — availability over partition-safety).** All clusters switched from `keep-oldest` to the `auto-down` downing strategy (Akka `AutoDowning`, `auto-down-unreachable-after` = 15s): a hard crash of EITHER node — active/oldest included — now fails over to the survivor in ~25s with no operator action. Accepted trade: a real network partition produces dual-active until an operator restarts one side. Decision record + evidence (live keep-oldest `DownReachable … including myself` log, Akka.NET 1.5.62 `KeepOldest.OldestDecision` source, rejected alternatives incl. the static-quorum-1 `DownAll` trap): `docs/plans/2026-07-21-auto-down-availability-decision.md`. | `archreview/plans/00-MASTER-TRACKER.md:194` + auto-memory `sbr-keep-oldest-2node-active-crash-gap` (both now historical) | — | Closed. Residual: seed-node boot-alone constraint (unchanged, documented in `Component-ClusterInfrastructure.md`); dual-active recovery is operator-driven. |
| vd03 | **`deploy/wonder-app-vd03/` overlay edits unapplied** — `appsettings.Central.json` needs `AllowSingleNodeCluster: true` + phantom-seed removal + `NodeName: central-a`; `install.ps1` needs `sc.exe failure` recovery actions. The `deploy/wonder-app-vd03/` artifact directory is intentionally untracked (production config out of source control), so the repo cannot ship the fix. | `archreview/plans/00-MASTER-TRACKER.md:198` (PLAN-01 T16/T20/T23) | Needs on-host access; without `NodeName` that deployment's audit rows stamp NULL `SourceNode`**partially mitigated once PLAN-R2-08 Task 7 lands: the host now FAILS AT BOOT with a key-naming error instead of silently NULLing, so applying the overlay becomes mandatory at the next upgrade.** Owner: whoever maintains the host (user). | Next wonder-app-vd03 deployment/upgrade — **the Task 7 validator makes this row unskippable then.** |