fc553bd9ba
Task 1 of the Phase 2 plan. The gate stops the plan, but not for the reason it anticipated: oplog sizing is fine and D6 is resolved. BLOCKER: the Phase 1 consolidated LocalDb (site-localdb.db) throws SQLite Error 10 'disk I/O error' on essentially every write on the ACTIVE node under sustained load — site_events, OperationTracking and the audit telemetry paths all fail, and site event logging silently drops events. Isolated to the load (not the node, not host-side observation) by failing over between nodes, and to LocalDb specifically (legacy store-and-forward.db / scadabridge.db in the same bind-mounted directory take zero errors under identical load). Phase 2 would register 8 more tables into that database — including the two highest-volume ones — while deleting the bespoke mechanisms that currently carry them. Must be root-caused first. Also recorded: D6's premise corrected (largest known production config_json is ~60-70 KB, not >128 KB — but MaxBatchSize 500 is still unsafe, use 16); D4's premise corrected (alarm writes are bounded by per-SourceReference coalescing at a 100 ms flush); sf_messages has a hard 50 rows/sec structural ceiling; and exceeding the oplog caps is a graceful snapshot-resync, not a failure. Claude-Session: https://claude.ai/code/session_01BL2Vu1ESDQ9SCN4gVKkdts