fix(opcua): hold last-known-good when a deploy's artifact cannot be read (#485)
A transient ConfigDb error while loading a deployment's artifact emptied the served address space. LoadArtifact caught the exception, logged "rebuild becomes no-op" and returned zero bytes; those parsed to an EMPTY composition, which the planner diffed against the live one as a PureRemove and the applier faithfully executed — 16 nodes gone, while the log claimed nothing happened. An artifact we could not obtain is not an empty configuration. Nothing legitimate produces a zero-length blob (an operator who really deletes everything still deploys a JSON document with empty arrays), so "no bytes" can only mean "no answer". HandleRebuild now abandons the rebuild on one, leaving both the materialised nodes and _lastApplied intact so the retry diffs against what is actually being served. The loader's own log line is now true. The two cases are logged differently: warn when there is a live address space being protected, debug when nothing has been deployed to this node yet. Covered by a RED-first actor test (verified failing: the load error tore down eq-2's subtree), plus a positive control proving a READABLE artifact that genuinely drops an equipment still removes it — the guard suppresses teardown only when the read failed. Found on the LocalDb Phase 1 follow-up live gate, which induced it by flapping SQL; that gate's log is the production evidence for the seam. Closes #485 Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
This commit is contained in:
@@ -48,7 +48,7 @@ scenario was believed impossible to recover from — it was a *documented limita
|
||||
gate is what turned the limitation into a reproducible case, and the second defect only existed
|
||||
behind the first.
|
||||
|
||||
## Unrelated finding (NOT a follow-up of this work — pre-existing; filed as `lmxopcua` issue #485)
|
||||
## Unrelated finding (NOT a follow-up of this work — pre-existing; `lmxopcua` issue #485, since FIXED)
|
||||
|
||||
While SQL was being stopped and started to drive check 4, site-b-1 took a deploy whose artifact load
|
||||
hit a transient ConfigDb error, and **emptied its served address space**:
|
||||
@@ -66,3 +66,13 @@ the node recovered fully on restart (`containers=16`). But a transient ConfigDb
|
||||
served address space instead of holding last-known-good is the same failure *class* as the Phase 1
|
||||
gate's check-3 defect. Untouched by this work — site-b-1 has replication off entirely — and induced by
|
||||
this gate's own SQL flapping.
|
||||
|
||||
**Fixed** (issue #485): `OpcUaPublishActor.HandleRebuild` now treats an artifact it could not obtain as
|
||||
*no answer* rather than *an empty configuration*, and abandons the rebuild with the served address space
|
||||
(and `_lastApplied`) intact. Nothing legitimate produces a zero-length blob — an operator who really
|
||||
deletes everything still deploys a JSON document with empty arrays — so "no bytes" can only mean the read
|
||||
failed. The log excerpt above is now impossible: the `PureRemove` never runs, and the loader's own
|
||||
"the rebuild is abandoned" line is finally true. Covered by
|
||||
`OpcUaPublishActorRebuildTests.Rebuild_keeps_the_last_known_good_address_space_when_the_artifact_load_fails`,
|
||||
with a positive control proving a *readable* artifact that genuinely drops an equipment still tears its
|
||||
subtree down.
|
||||
|
||||
Reference in New Issue
Block a user