Files
ScadaBridge/docs/plans/2026-07-08-deferred-work-register.md
T

15 KiB
Raw Blame History

Deferred-Work Register (established 2026-07-08, from architecture review 08)

Single tracked list of consciously-deferred work. Rules: every deferral gets a row (rationale + revisit trigger); fix-now items reference the archreview plan that owns them and are removed from this table when that plan's task lands.

Fix-now (owned by archreview plans)

All 7 fix-now items landed via PLAN-04/05/06/07/08 (verified in review 08 round 2, 2026-07-12) — rows removed per the rule above; see the round-2 report §1 for the per-item evidence.

Deferred (with rationale + revisit trigger)

# Item Where noted Rationale for deferral Revisit trigger
8 Hash-chain tamper evidence (T1); CLI verify-chain is a no-op stub audit-log roadmap :12 v1.x by locked decision; append-only DB roles are the control Compliance requirement for cryptographic tamper evidence
9 Parquet audit archival (T2); endpoint returns 501 AuditEndpoints.cs:204 v1.x; 501 + CLI messaging are honest AuditLog partition volume nears retention ceiling
11 Central-persisted OPC UA cert-trust audit m7 follow-ups Broadcast-to-both-nodes covers HA Governance/audit requirement for trust decisions
17 Unified notifications+site-calls outbox page stillpending :118 Explicit M9 decision to keep two pages Operator confusion reports
19 Bundle signing / cluster-to-cluster pull / differential bundles transport-design :402 v1 manifest hash + AES-GCM held sufficient Non-repudiation requirement across orgs
23 Live LDAP group-membership re-query for an active session docs/requirements/Component-Security.md :61-69 (+ :78-79) Blocked on an external package. The mid-session refresh re-maps the stored groups against the central DB with no LDAP call, so a directory group-membership change lands only at next login. A live re-query needs a passwordless service-account group-search method on the shared ZB.MOM.WW.Auth.Ldap library — an external NuGet PackageReference (src/ZB.MOM.WW.ScadaBridge.Security/…csproj:23) exposing only AuthenticateAsync(username, password, ct). Central role-mapping/scope changes still apply within ~15 min (RoleRefreshThresholdMinutes). ZB.MOM.WW.Auth.Ldap gains a standalone group-search API, or a requirement that a directory-side group revocation take effect mid-session rather than at next login
24 M8 large-bundle performance hardening docs/plans/2026-06-15-stillpending-completion-design.md:106 — "Small follow-ups logged (not blocking): … large-bundle/perf hardening" Logged as a non-blocking follow-up when M8 shipped and never given an artifact: no plan, no task entry, no perf/load test exists (tests/…Transport.Tests/Import/BundleImporterLoadTests.cs is a LoadAsync unit suite despite the name). No measured problem; the only sizing controls in place are the 5-minute CLI transport timeout, LineDiffer's MaxInputLines=4000 summary-only cap, and MaxConcurrentImportSessions=8. First real bundle that times out, exhausts memory, or makes the import wizard's diff step unusable
25 Phase-8 WP-4 target-scale load test (10 sites × 500 instances × 75 tags = 37,500 subscriptions/site, 375,000 total) docs/plans/phase-8-production-readiness.md:152-170 (WP-4) + :314-320 (test protocol); status claimed in docs/plans/phase-8-checklist.md Claimed complete but unevidenced. The whole WP-4 deliverable is a 107-byte checklist stub asserting "Status: Complete / Tests: All passing / Build: 0 errors, 0 warnings" with no per-work-package results and no linked run. Nearest real coverage is arithmetic/aggregation only — PerformanceTests/StaggeredStartupTests.cs (TagCapacity_75TagsPer500Machines_37500Total, 500-instances-over-10-sites distribution) and HealthAggregationTests (10-site report aggregation) — plus a single-subscriber 100k-event Streaming/SiteStreamThroughputTests.cs. No sustained multi-site run exists anywhere in tests/ or docker/. Before any production go-live at target scale; or the first site approaching ~500 instances / ~37.5k subscriptions
26 Ipsen MES MoveIn tail: leak-test (-LT) receivers + routing, PLC-output-flag writes, Z28062 BTDB data completeness docs/plans/2026-06-16-ipsen-mes-movein.md:409 ("Out of scope (future)"); design 2026-06-16-ipsen-mes-movein-design.md:58-60, :196-198 Customer-site scope, not a platform gap. -LT routing needs an MES-receiver child + Galaxy reference that do not exist on the reactor template (any -LT/unknown suffix returns WasSuccessful=false with an "unsupported side/target" message by decision); MoveInComplete/Successful/ErrorText are PLC-owned by locked decision, so ScadaBridge deliberately does not write them; Z28062 completeness is an operational data fix, not code. Note the separate alarm-status path already handles the suffix — _LT is stripped before side-scoping (2026-06-30-mes-alarm-status-api.md:158). Ipsen creates the leak-test receiver + Galaxy reference, or asks ScadaBridge to own the PLC-output flags — otherwise a candidate won't-do ([PERM]) at the next Ipsen scope review

Resolved (verified against the code 2026-07-10)

Rows removed from the Deferred table above once confirmed shipped. Kept here for traceability.

# Item Resolution
10 Aggregated live alarm stream for Alarm Summary Shipped 2026-07-10 (docs/plans/2026-07-10-aggregated-live-alarm-stream-plan.md): a transient, in-memory per-site central live alarm cache (ISiteAlarmLiveCache/SiteAlarmLiveCacheService + per-site SiteAlarmAggregatorActor) fed by a new site-wide, alarm-only SubscribeSite gRPC stream (SiteStreamManager.SubscribeSiteAlarms), seed-then-stream with dedup + NodeA↔NodeB re-seed + periodic reconcile. Alarm Summary is now live-cache-driven (AlarmSummaryService.BuildFromLiveAlarms) with the 15s poll retained as fallback + NotReporting authority. Honors the [PERM] no-central-store rule — nothing persisted (no EF table/migration). Options on CommunicationOptions (eagerly validated) + two ScadaBridgeTelemetry signals.
7 SecuredWrite audit rows leave SourceNode NULL Resolved (PLAN-07): ManagementActor.EmitSecuredWriteAuditAsync routes through ICentralAuditWriter, which stamps SourceNode (central-a/central-b) from INodeIdentityProvider.
12 (CLI/API) Native-alarm-source-override CSV import Shipped 2026-07-10: shared CsvLineSplitter, NativeAlarmSourceOverrideCsvParser, bulk all-or-nothing SetInstanceNativeAlarmSourceOverridesCommand + ManagementActor handler (Deployer-gated), CLI instance native-alarm-source import --file, parser/CLI/handler tests. UI upload affordance shipped 2026-08-01 — second InputFile on the InstanceConfigure Native Alarm Source Overrides card reusing the shared parser, mirroring the attribute importer's UX and the server's all-or-nothing merge semantics (InstanceConfigureNativeAlarmCsvImportTests); row 12 removed from the Deferred table.
13 WaitForAttribute quality-gated ("Good"-only) mode Already implemented (Commons WaitForAttribute.RequireGoodQuality, enforced in InstanceActor, threaded through ScriptRuntimeContext, tested in InstanceActorWaitForAttributeTests). Stale "planned enhancement" doc line corrected 2026-07-10.
14 WaitForAttribute in Test-Run sandbox Shipped 2026-07-10 (full fidelity): sandbox Attributes.WaitAsync/WaitForAsync (value-equality) route to the bound instance via ISandboxInstanceGateway.WaitForAttributeAsync → the existing CommunicationService.RouteToWaitForAttributeAsync cross-site route. Additive RouteToWaitForAttributeRequest.RequireGoodQuality (honored by the site handler) makes quality-gated waits route too. Predicate-form waits stay unsupported (an in-process lambda can't be routed) and throw a labelled ScriptSandboxException. Tests: sandbox accessor routing (CentralUI), site-handler quality-flag threading (SiteRuntime).
15 BrowseNext final-page signal not surfaced Already surfaced (M7 browse work): RealOpcUaClient sets Truncated=false/ContinuationToken=null on the last page; BrowseNodeResult carries both; TreeRow.razor renders "Load more" only when a continuation token remains — no wasted BrowseNext.
16 StubOpcUaClient throws on browse Already resolved: StubOpcUaClient supports browse + address-space search, covered by StubOpcUaClientBrowseTests/StubOpcUaClientSearchTests.
18 Folder drag-drop Closed — permanently deferred ([PERM]) by the M9 decision (docs/plans/2026-06-15-stillpending-completion-design.md:122): menu-based reorder (T23) shipped instead, and the folder-hierarchy design fixed the reorganization UX as "right-click context menus only (no drag-drop)" (2026-05-11-templates-folder-hierarchy-design.md:27). Row removed from the Deferred table 2026-08-01 — nothing left to revisit.
20 Deployment EXPIRED-row purge Already resolved (PLAN-04): PendingDeploymentPurgeActor central singleton (spawned in AkkaHostedService) ticks IDeploymentManagerRepository.PurgeExpiredPendingDeploymentsAsync every CommunicationOptions.PendingDeploymentPurgeInterval (default 1h), options-validated, tested.
21 SiteAuditBacklogReporter threshold consolidation Shipped 2026-07-10: SqliteAuditWriterOptions.BacklogPollIntervalSeconds (default 30) now drives the reporter's poll cadence; explicit ctor override still wins (tests), non-positive falls back to the 30 s default. Cadence tests added; stale "hard-code / follow-up" class-doc corrected.
22 KPI history hourly rollups Shipped 2026-07-10 (docs/plans/2026-07-10-kpi-history-hourly-rollups-plan.md, T1T8): new KpiRollupHourly table (migration 20260710153953) folded by a third recorder tick (kpi-rollup, RollupInterval default 1h) over a re-folded RollupLookbackHours window via an idempotent, failover-self-healing upsert; per-metric gauge-vs-rate aggregation (KpiMetricAggregationCatalog); a one-shot backfill of the retention window on start; raw-vs-rollup query routing by RollupThresholdHours (default 168h); longer rollup retention (RollupRetentionDays default 365 ≥ RetentionDays, dual daily purge); and 30 d / 90 d trend windows added to the four surfaces. Options + validator, docs (Component-KpiHistory.md), and tests shipped.

New deferrals from review 08 (this plan)

Item Rationale Revisit trigger
Communication → HealthMonitoring layering (ICentralHealthAggregator consumed by CentralCommunicationActor.cs:351) Moving the interface + SiteHealthState to Commons ripples across 5 projects for a cosmetic inversion Next breaking change to ICentralHealthAggregator
docs/components reference docs for ScriptAnalysis, KpiHistory, DelmiaNotifier Reference docs are substantial (StyleGuide-conformant); README claim scoped instead (PLAN-08 Task 10) Next doc-writing session touching those components
Test-coverage backfill: SiteCallAudit.Tests (31 tests/1.6k LOC), DeploymentManager.Tests No defect identified; coverage partly lives in ManagementService/Host/Integration suites First regression escaping either component
~Failover-timing measurement (the "25s total failover" envelope) RESOLVED 2026-08-01 — split out of the combined row and closed. tests/ZB.MOM.WW.ScadaBridge.PerformanceTests/Failover/FailoverTimingTests.cs is no longer a skipped placeholder: it runs as a live [Fact] (Category=Performance) on the real two-node in-process rig (TwoNodeClusterFixture, production BuildHocon) at production timings — 2s heartbeat / 10s failure-detection threshold / 15s stable-after — hard-killing the younger node and timing the survivor's member REMOVAL with singleton continuity asserted on the oldest. Delivered by PLAN-R2-01 Task 4 (archreview/plans/PLAN-R2-01-cluster-host-failover.md:226). The oldest-crash direction is covered behaviorally by SbrFailoverTests.AutoDown_HardCrashOfOldestNode_* and by docker/failover-drill.sh. The 2026-07-08 "PLAN-01 rig landing" trigger had fired unnoticed (NF2); PLAN-R2-01 T4 wired the placeholder to the fixture rig rather than recording a blocker. Closed.
Broader perf envelope — S&F drain rate + per-subscriber stream backpressure (the still-open half of the former combined row) Never measured, and no owner plan survives now that PLAN-R2-01 closed the failover half. PerformanceTests covers failover timing, staggered startup, health aggregation, audit hot-path latency and a single-subscriber 100k-event Streaming/SiteStreamThroughputTests.cs — nothing measures store-and-forward drain throughput, nor what a slow/stalled subscriber does to the per-subscriber buffering in Communication/Actors/StreamRelayActor.cs / Grpc/SiteStreamGrpcServer.cs with many subscribers attached. No defect observed; deferred as measurement-only work. First field S&F backlog that fails to drain within an operator's patience, a slow gRPC subscriber degrading a site stream for others, or the WP-4 target-scale run (row 25) being scheduled — that run should absorb this

Deferred — operational risk (from the initiative tracker, folded in 2026-07-12)

Two live items previously tracked ONLY in archreview/plans/00-MASTER-TRACKER.md's registry are folded in here (NF5) so this register is the single tracking place. The tracker's narrative subsections remain as the historical evidence.

# Item Where noted Rationale for deferral Revisit trigger
SBR SBR oldest-crash total-outage gap RESOLVED 2026-07-21 (owner decision — availability over partition-safety). All clusters switched from keep-oldest to the auto-down downing strategy (Akka AutoDowning, auto-down-unreachable-after = 15s): a hard crash of EITHER node — active/oldest included — now fails over to the survivor in ~25s with no operator action. Accepted trade: a real network partition produces dual-active until an operator restarts one side. Decision record + evidence (live keep-oldest DownReachable … including myself log, Akka.NET 1.5.62 KeepOldest.OldestDecision source, rejected alternatives incl. the static-quorum-1 DownAll trap): docs/plans/2026-07-21-auto-down-availability-decision.md. archreview/plans/00-MASTER-TRACKER.md:194 + auto-memory sbr-keep-oldest-2node-active-crash-gap (both now historical) Closed. Residual: seed-node boot-alone constraint (unchanged, documented in Component-ClusterInfrastructure.md); dual-active recovery is operator-driven.
vd03 deploy/wonder-app-vd03/ overlay edits unappliedappsettings.Central.json needs AllowSingleNodeCluster: true + phantom-seed removal + NodeName: central-a; install.ps1 needs sc.exe failure recovery actions. The deploy/wonder-app-vd03/ artifact directory is intentionally untracked (production config out of source control), so the repo cannot ship the fix. archreview/plans/00-MASTER-TRACKER.md:198 (PLAN-01 T16/T20/T23) Needs on-host access; without NodeName that deployment's audit rows stamp NULL SourceNodepartially mitigated once PLAN-R2-08 Task 7 lands: the host now FAILS AT BOOT with a key-naming error instead of silently NULLing, so applying the overlay becomes mandatory at the next upgrade. Owner: whoever maintains the host (user). Next wonder-app-vd03 deployment/upgrade — the Task 7 validator makes this row unskippable then.