# LocalDb pair replication — operations runbook > **Phase 1 scope.** Every driver-role OtOpcUa node keeps a consolidated > [`ZB.MOM.WW.LocalDb`](../../CLAUDE.md) SQLite database. Today it caches exactly one thing: the > **deployed-configuration artifact**, chunked, so a node can **boot from cache when central SQL > Server is unreachable**. That cache can *optionally* replicate to the node's redundant pair peer > over the library's gRPC sync. **Replication is default-OFF and fail-closed.** This runbook covers > enabling it, the operational rules that keep a pair converging, and the DB-inspection safety > rules. ## What replicates, and what it buys you - **Replicated tables:** `deployment_artifacts` (the artifact, split into base64 chunks) and `deployment_pointer` (one current-deployment pointer per cluster). Registered — in this order, after the DDL — by `LocalDbSetup.OnReady`. - **The payoff:** in a redundant driver pair, *either* node can be the one that applied the latest deploy and wrote the cache. With replication on, the peer holds a byte-identical copy. So a node that restarts into a central-SQL outage can boot the current config **even if it never applied that deploy itself** — its peer did, and the row replicated. - **Convergence model:** last-writer-wins per primary key, HLC-stamped. Both nodes also store locally after every apply, so the pair converges to identical content regardless of which node deployed. Retention (newest-2 deployments per cluster) prunes on one node and the deletes replicate as tombstones. ## Enabling replication on a pair Replication has two knobs, both under `LocalDb`: | Key | Role | Notes | |---|---|---| | `LocalDb:SyncListenPort` | binds the dedicated **h2c** sync listener | `0` (default) = no listener, replication off. Set the **same** non-zero port on **both** nodes of the pair. | | `LocalDb:Replication:PeerAddress` | the address this node **dials** | Set on **one** node (the initiator); leave empty on the other (passive). The stream is bidirectional — one dial carries both directions. | | `LocalDb:Replication:ApiKey` | bearer token for the sync stream | **Must be byte-identical on both nodes.** See fail-closed rule below. Supply via `${secret:...}` / env — never a cleartext literal in production config. | | `LocalDb:Replication:MaxBatchSize` | rows per replication batch | `16` for the artifact cache. Batching is **row-count-only** against gRPC's 4 MB message cap; artifact chunk rows are ≈171 KB, so 16 × 171 KB ≈ 2.7 MB stays under the cap. Raising it risks tripping the cap. | Example (site-a-1 initiator, site-a-2 passive): ``` site-a-1: LocalDb__SyncListenPort=9001 LocalDb__Replication__PeerAddress=http://site-a-2:9001 LocalDb__Replication__ApiKey=${secret:site-a-localdb-sync-key} LocalDb__Replication__MaxBatchSize=16 site-a-2: LocalDb__SyncListenPort=9001 LocalDb__Replication__ApiKey=${secret:site-a-localdb-sync-key} # no PeerAddress LocalDb__Replication__MaxBatchSize=16 ``` > **Kestrel note.** Setting `SyncListenPort` makes the Host add an explicit `Listen*` for the h2c > sync listener. An explicit `Listen*` makes Kestrel **ignore `ASPNETCORE_URLS` entirely**, so the > Host also re-binds the primary HTTP port in the same block. If you change the primary port, > verify both `:` and `:` appear in the startup "Now listening on" lines. ### The ApiKey rule (fail-closed) The replication library's passive endpoint verifies **no** authentication — the host `LocalDbSyncAuthInterceptor` is the only gate, and it is **fail-closed**: - **No key configured ⇒ every sync call is refused** (`PermissionDenied`). "No key" is never "no auth required". - **A key mismatch ⇒ the pair silently stops converging.** The initiator's stream is rejected at the peer; nothing errors loudly. A typo in one node's key looks exactly like "replication is broken". If a pair is not converging, **check the keys match first.** ## Operational rules - **Stop / start the pair together where you can.** Each node keeps working (and caching locally) while its peer is down; the outage is not a data-loss event — the surviving node accumulates writes and the peer catches up on rejoin. But a long-lived solo node drifts further from its peer, so avoid leaving a pair split for extended periods. - **Tombstone-retention resurrection window.** Retention prunes to the newest 2 deployments and replicates the prune as tombstones. Tombstones are themselves retained only for a bounded window. If a node is offline **longer than the tombstone-retention window**, a delete that happened during its outage may no longer be expressible as a tombstone on rejoin — a pruned deployment could briefly reappear until the next deploy re-prunes it. Keep pair outages well inside that window. - **A rebuilt node is back-filled by its peer.** If a node loses its LocalDb file — a wiped volume, a re-imaged host, a fresh container — it rejoins with an empty database and its peer snapshots the cached configuration back to it, with no new deploy and no central SQL. Two things are worth knowing: this needs **`ZB.MOM.WW.LocalDb` ≥ 0.1.2** (on `0.1.1` a converged pair has pruned every oplog row on ack, and the wiped node was never snapshotted — it stayed empty until the next deploy), and it heals the *cache*, so the node regains boot-from-cache for future outages. A node wiped **during** a central outage is still repopulated by its peer under 0.1.2; a node with **no** peer replication configured self-heals only from central on its next successful apply. - **What boot-from-cache does NOT cover.** The cache is a *fallback for central-SQL outages at boot*, not a replacement for central. A **new deployment still requires central SQL** — the cache is only read when the central fetch fails at startup. When central returns, the node resumes fresh-config behavior. Boot-from-cache logs a running-from-cache signal; treat a node that stays on it as a central-connectivity incident, not steady state. ## Inspecting the LocalDb file — safety rules The database runs in **WAL mode**. These rules exist because violating them corrupted a live DB in the 2026-07-20 ScadaBridge incident: - **Never run `sqlite3` on the live file** (host-side, against a bind-mounted or container path). Opening a live WAL DB from a second process across virtiofs poisons the WAL. **Always copy the triplet first** and query the copy: ```bash docker cp :/app/data/otopcua-localdb.db /tmp/localdb.db docker cp :/app/data/otopcua-localdb.db-wal /tmp/localdb.db-wal docker cp :/app/data/otopcua-localdb.db-shm /tmp/localdb.db-shm sqlite3 /tmp/localdb.db 'SELECT cluster_id, deployment_id FROM deployment_pointer;' ``` - **Metrics** come from the container, not a host curl (`aspnet:10.0` has no `curl`): ```bash docker run --rm --network container: curlimages/curl:latest -s localhost:/metrics | grep localdb_ ``` - **If an anomaly appears within seconds of your own measurement, suspect the measurement.** Restart both nodes and re-observe untouched before blaming the library. ## Useful queries (on a copied triplet) ```sql -- What each node thinks the current deployment is SELECT cluster_id, deployment_id, revision_hash, applied_at_utc FROM deployment_pointer; -- Convergence check: the pointer's origin stamp should match on both nodes SELECT pk_json, hlc, node_id, is_tombstone FROM __localdb_row_version WHERE table_name = 'deployment_pointer' ORDER BY pk_json; -- Replication backlog (should drain to 0 when a pair is caught up) SELECT COUNT(*) FROM __localdb_oplog; ``` ## See also - [`docs/Redundancy.md`](../Redundancy.md) — the pair-local config cache section. - [`docs/Configuration.md`](../Configuration.md) — the `LocalDb` appsettings section. - The two-node convergence harness: `tests/Server/ZB.MOM.WW.OtOpcUa.Host.IntegrationTests/LocalDb/`.