Phase 3 Task 4. GrpcDeploymentArtifactFetcher streams the artifact from the first
configured central endpoint that answers, reassembles the chunks, and verifies
SHA-256(bytes) lowercase-hex == the dispatched RevisionHash (byte-identical to
ConfigComposer's Convert.ToHexStringLower(SHA256.HashData(blob))). Every per-endpoint
problem — RpcException, zero-length stream, or hash mismatch — is logged and the next
endpoint tried; if none yield verified bytes it returns null (apply failure, keep
last-known-good — the #485 contract), never throwing into the actor loop. A
Func<endpoint, client> seam keeps the gRPC boundary out of the unit tests (that is
Task 8's job); the real path caches one h2c GrpcChannel per endpoint.
Files live under the .DeploymentCache namespace (folder Deployment/) to avoid shadowing
the Configuration.Entities.Deployment type in the test assembly — same convention as
LocalDbDeploymentArtifactCache.
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
1. Wiped-node back-fill (live gate check 4). Fixed upstream in
ZB.MOM.WW.LocalDb 0.1.2 and pinned here. The gate's diagnosis was slightly
off: the oplog cap was not the gate. Snapshot detection measured the peer's
gap against the oldest surviving oplog row and read an empty oplog as "no
gap possible" — and an empty oplog is the steady state of a converged pair,
since ack-pruning deletes everything the peer confirmed. The healthy state
was the one state that could not heal a wiped peer. Now measured against
last_acked_seq when the oplog is empty.
Pinned at this level too, not just the library's: the pair harness grows a
WipePassiveAsync (a NEW database — a rebuilt node comes back with a new node
id and a zero watermark), and the new scenario asserts the emptied oplog as
its precondition before wiping, then requires the deployment artifact to
come back byte-identical with no new deploy. Verified RED against the pinned
0.1.1 (times out waiting for back-fill) and green on 0.1.2.
2. Oplog growth on default-OFF nodes (live gate check 8), fixed at the source:
StoreAsync now skips entirely when the pointer already names this
deployment/revision, the SHA matches the bytes, and the expected chunk count
is present. Re-caching an artifact the node already holds — every restart's
boot-from-cache, every RestoreApplied — writes nothing and mints no oplog
rows, where before it cost a delete plus an insert per chunk plus a pointer
update, all identical to what was already there.
Identity is over the bytes and the chunks are counted, both deliberately: a
re-composed artifact can carry the same ids with different bytes, and a
pointer can name a deployment whose chunks are missing, where skipping would
make an unreadable cache permanent. Both guards verified RED against a naive
pointer-only skip.
The gate's stated bound was also wrong and is corrected in the doc: growth
was never headed for the 1M row cap. AddZbLocalDbReplication is registered
unconditionally, so MaintenanceBackgroundService runs on default-OFF nodes
and the 7-day MaxOplogAge cap prunes.
Runtime.Tests 412/0/31, Host.IntegrationTests LocalDb 45/45. Pre-existing and
unrelated: AbCip_Green_AgainstSim (needs the AB CIP docker fixture, red on a
clean master too) and a flaky Roslyn race test (green 2/2 in isolation).
Claude-Session: https://claude.ai/code/session_01GASWkNEi68FSCtvr6rLoEW
128 KiB raw chunks base64-encoded into deployment_artifacts, with a per-cluster
pointer carrying a SHA-256 over the raw bytes. Reassembly verifies chunk count
AND the SHA before returning, so a partially replicated artifact reads back as a
clean miss rather than a silently truncated address space.
Adds GetCurrentUnkeyedAsync beyond the plan's interface: at the boot-failure seam
the actor cannot know its ClusterId (it is derivable only from an artifact you
already have, or from the central DB that is unreachable). Recon D-1.
The prune's 'deployment_id <> @DeploymentId' clause makes 'the pointer's target is
always present' structural rather than a side effect of clock resolution. Three
stores inside one timestamp tick fall through to a deployment_id tiebreak, and
since real ids are GUIDs that ordering is random - the row just written could lose,
leaving the pointer naming chunks that no longer exist. Verified red without the
clause.
The Runtime.Deployment namespace shadowed the Deployment EF entity type for
every file under ZB.MOM.WW.OtOpcUa.Runtime.Tests.*: C# name lookup walks the
enclosing namespace chain and finds the namespace before a using-imported type,
so 19 pre-existing call sites failed with CS0118. Caught only by a full-solution
build - the Host and its own tests compiled fine.