fix(central): review findings — no client-side audit truncation, insert-first upsert, QI-safe scripts, honest operator replies
Six adversarial-review findings in the central SQL/ingest layer. F1 (AuditLogRepository.InsertChunkAsync) — the set-based ingest declared each string parameter at its COLUMN width (Actor/Target 256, Action 64, Outcome 16, Category 32, SourceNode 64), so SqlClient truncated an over-long value at bind time and committed the mutilated row — silent, in an append-only store, with no PayloadTruncated flag — while the per-row and reconciliation paths sent the same value in full and let the server reject it with 2628. Bind at the value's own length instead; explicit SqlDbType is kept (it fixes the VALUES constructor's derived column types and datetime2 precision). Design: reject everywhere, truncate nowhere — matching today's per-row behaviour. F2 (SiteCallAuditRepository.UpsertAsync) — the single-statement upsert ran the monotonic UPDATE first and INSERTed only if nothing matched. Two writers racing the first packet of one TrackedOperationId (the cached dual-write and the reconciliation pull carry DIFFERENT lifecycle states) both matched nothing, and the loser then skipped its INSERT or swallowed a 2627 — dropping its Status/RetryCount/HttpStatus/TerminalAtUtc. Legs swapped to `IF NOT EXISTS … INSERT; UPDATE <monotonic>` — still one round trip, and the loser's UPDATE now lands on the winner's row. The duplicate-key catch re-runs the monotonic UPDATE for the same reason. Moved to raw SQL with explicitly-typed parameters so the intricate rank predicate exists in exactly one place (an untyped DateTime would bind as `datetime` and round the freshness tiebreaker). F3 (docs/plans/sql/*.sql) — filtered-index DDL failed with error 1934 under the documented `docker exec … sqlcmd` path, which defaults QUOTED_IDENTIFIER OFF; once IX_Notifications_Delivered exists, QI-OFF DML on Notifications fails too. All four scripts now open with `SET QUOTED_IDENTIFIER ON; SET ANSI_NULLS ON; GO` (own batch, so it is in force when the next batch parses), and the migration convention in Component-ConfigurationDatabase.md documents `sqlcmd -I`. Verified live: the pre-fix script fails 1934 without -I, the fixed one applies. F4 (SiteCallAuditActor) — the off-mailbox reconciliation/purge passes reuse the injected repository, so tests drove one DbContext from the pass and a mailbox handler concurrently. Serialized at the CALL via a private SerializedRepository wrapper applied only by the test constructors, rather than running the pass on-mailbox: production keeps its PipeTo shape untouched, and the existing "a blocked drain does not stall ingest/query/KPI" regression tests stay meaningful (they would have been invalidated by suspending the mailbox). F5 (AuditLogIngestActor) — when the batch failed because the 20 s IngestBudget expired, the per-row fallback reused the same expired token: N instant failures, N counter bumps, zero accepted. The fallback now gets a fresh 5 s budget (inside the 30 s outer Ask), and a blown budget bumps the failure counter ONCE for the batch instead of once per row. F6 (NotificationOutboxRepository.UpdateAsync) — ExecuteUpdate's row count was discarded, so an operator Retry/Discard of a notification the retention purge had already deleted reported success (the pre-ExecuteUpdate code threw DbUpdateConcurrencyException). UpdateAsync now returns whether a row matched; the operator one-shots answer "notification not found" and emit no audit row for the action that did not happen, while the dispatcher logs a warning (its delivery already happened; nothing to retry). GetByIdAsync switched to AsNoTracking since the write is out-of-band. Tests: 5 new SQL-backed regressions (over-long Target rejected on both paths + boundary round-trip; concurrent first-write and already-created-by-another-writer upserts; vanished-row UpdateAsync), a token-identity pin on the ingest fallback, a repository-concurrency detector for the SiteCallAudit passes, and vanished-row operator-path tests. The F1/F2/F4 regressions were each confirmed failing against the pre-fix code. Suites: ConfigurationDatabase 369, AuditLog 378, SiteCallAudit 66, NotificationOutbox 152 — all green, solution builds with 0 warnings.
This commit is contained in:
@@ -322,6 +322,47 @@ public class AuditLogIngestActorTests : TestKit, IClassFixture<MsSqlMigrationFix
|
||||
Assert.Equal(2, counter.Count);
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// The per-row fallback must run on its OWN cancellation token, never the
|
||||
/// batch's. Sharing it meant that when the batch failed BECAUSE the 20 s
|
||||
/// ingest budget expired, every fallback insert was handed an
|
||||
/// already-cancelled token: N instant failures, N counter bumps, zero rows
|
||||
/// accepted — the fallback's entire purpose (land the good rows) defeated at
|
||||
/// the exact moment it was needed.
|
||||
/// </summary>
|
||||
/// <remarks>
|
||||
/// Pinned by token IDENTITY rather than by expiring a real budget: the budget
|
||||
/// is a fixed 20 s and a test that waited for it would cost 20 s of wall clock
|
||||
/// to assert something the token comparison establishes outright. Two
|
||||
/// <see cref="CancellationToken"/>s are equal iff they come from the same
|
||||
/// source, so "not equal" IS "a fresh CTS", and a fresh CTS cannot be
|
||||
/// pre-cancelled by the batch.
|
||||
/// </remarks>
|
||||
[Fact]
|
||||
public async Task Receive_WhenBatchFails_PerRowFallbackUsesAFreshToken()
|
||||
{
|
||||
var repository = new TokenRecordingRepository();
|
||||
var actor = CreateActor(repository);
|
||||
|
||||
var events = Enumerable.Range(0, 3).Select(_ => NewEvent(NewSiteId())).ToList();
|
||||
actor.Tell(new IngestAuditEventsCommand(events), TestActor);
|
||||
|
||||
var reply = ExpectMsg<IngestAuditEventsReply>(TimeSpan.FromSeconds(10));
|
||||
|
||||
// The fallback landed every row.
|
||||
Assert.Equal(3, reply.AcceptedEventIds.Count);
|
||||
|
||||
Assert.NotNull(repository.BatchToken);
|
||||
Assert.Equal(3, repository.RowTokens.Count);
|
||||
Assert.All(repository.RowTokens, rowToken =>
|
||||
{
|
||||
Assert.NotEqual(repository.BatchToken!.Value, rowToken);
|
||||
Assert.False(rowToken.IsCancellationRequested);
|
||||
});
|
||||
|
||||
await Task.CompletedTask;
|
||||
}
|
||||
|
||||
/// <summary>Counts how many times the guard's catch surfaced a write failure.</summary>
|
||||
private sealed class CountingFailureCounter : ICentralAuditWriteFailureCounter
|
||||
{
|
||||
@@ -329,6 +370,62 @@ public class AuditLogIngestActorTests : TestKit, IClassFixture<MsSqlMigrationFix
|
||||
public void Increment() => Count++;
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Fails the set-based insert (as an expired budget would) and records the
|
||||
/// token handed to each write so the fallback's token can be compared with
|
||||
/// the batch's.
|
||||
/// </summary>
|
||||
private sealed class TokenRecordingRepository : IAuditLogRepository
|
||||
{
|
||||
public CancellationToken? BatchToken { get; private set; }
|
||||
|
||||
public List<CancellationToken> RowTokens { get; } = new();
|
||||
|
||||
public Task<int> InsertManyIfNotExistsAsync(
|
||||
IReadOnlyList<AuditEvent> events, TimeSpan? commandTimeout = null, CancellationToken ct = default)
|
||||
{
|
||||
BatchToken = ct;
|
||||
throw new OperationCanceledException("simulated ingest-budget expiry", ct);
|
||||
}
|
||||
|
||||
public Task InsertIfNotExistsAsync(AuditEvent evt, CancellationToken ct = default)
|
||||
{
|
||||
RowTokens.Add(ct);
|
||||
return Task.CompletedTask;
|
||||
}
|
||||
|
||||
public Task<IReadOnlyList<AuditEvent>> QueryAsync(
|
||||
AuditLogQueryFilter filter, AuditLogPaging paging, CancellationToken ct = default) =>
|
||||
throw new NotSupportedException();
|
||||
|
||||
public Task<long> SwitchOutPartitionAsync(
|
||||
DateTime monthBoundary, TimeSpan? commandTimeout = null, CancellationToken ct = default) =>
|
||||
throw new NotSupportedException();
|
||||
|
||||
public Task<long> PurgeChannelOlderThanAsync(
|
||||
string channel, DateTime threshold, int batchSize, TimeSpan? commandTimeout = null, CancellationToken ct = default) =>
|
||||
throw new NotSupportedException();
|
||||
|
||||
public Task<long> BackfillSourceNodeAsync(
|
||||
string sentinel, DateTime before, int batchSize, CancellationToken ct = default) =>
|
||||
throw new NotSupportedException();
|
||||
|
||||
public Task<IReadOnlyList<DateTime>> GetPartitionBoundariesOlderThanAsync(
|
||||
DateTime threshold, CancellationToken ct = default) =>
|
||||
throw new NotSupportedException();
|
||||
|
||||
public Task<ZB.MOM.WW.ScadaBridge.Commons.Types.AuditLogKpiSnapshot> GetKpiSnapshotAsync(
|
||||
TimeSpan window, DateTime? nowUtc = null, CancellationToken ct = default) =>
|
||||
throw new NotSupportedException();
|
||||
|
||||
public Task<IReadOnlyList<ExecutionTreeNode>> GetExecutionTreeAsync(
|
||||
Guid executionId, CancellationToken ct = default) =>
|
||||
throw new NotSupportedException();
|
||||
|
||||
public Task<IReadOnlyList<string>> GetDistinctSourceNodesAsync(CancellationToken ct = default) =>
|
||||
throw new NotSupportedException();
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Tiny test double that delegates to a real repository but throws on a
|
||||
/// specified EventId. Used to verify per-row failure isolation: one bad
|
||||
|
||||
Reference in New Issue
Block a user