fix(audit): pair every audit request entry with a response (WP35) - #1601
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 04ef524d41
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5de578f1b0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
registry-stack/crates/registry-breg/src/action_evidence_maintenance.rs
Lines 203 to 206 in 09106c9
If PostgreSQL commits the expired-Evidence deletion but the connection disappears before acknowledging commit, this maps the result to Unavailable, after which erase_expired records a durable failed response even though the protected material was erased. Re-read the retained Evidence state on a fresh connection, as the request-retention path does, before deciding whether to record erased, failed, or an uncertain outcome.
AGENTS.md reference: products/breg/AGENTS.md:L175-L182
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
This answers the Codex review-body finding on
Generated by Claude Code |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8ee314f7ac
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8f329cac50
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| tokio::spawn(async move { | ||
| let result = writer.write(&entry).await; |
There was a problem hiding this comment.
Settle claimed requests when the spawned task is canceled
When the Tokio runtime shuts down while this task is awaiting the write, the task is dropped before lines 665–671 clear claimed/in_flight; the owning AuditRequest then observes an in-flight response and skips its unfinished entry, leaving a durable request permanently unpaired. AuditRequest::respond has the same non-RAII bookkeeping pattern. Fresh evidence after the earlier cancellation fix is that the spawned task survives cancellation of its caller but not cancellation of the runtime itself, so the claim needs a drop guard that settles or queues the unfinished response.
AGENTS.md reference: products/platform/AGENTS.md:L29-L31
Useful? React with 👍 / 👎.
| let _attempt = self | ||
| .begin_action_boundary_audit(claims, input.route_id, input.correlation) | ||
| .await?; |
There was a problem hiding this comment.
Carry the Evidence action handle through finalization
For every Evidence-backed action that passes preflight without replaying a receipt, _attempt is dropped when this function returns PreparedEvidenceAction, before invoke_evidence_action performs Evidence evaluation and calls finalize_evidence_action. Its destructor therefore writes an unfinished response even though the operation is still running, and finalization later appends a second terminal response under the same correlation. Return the handle with the prepared state or hold it in the outer coordinator until evaluation and finalization finish.
AGENTS.md reference: products/breg/AGENTS.md:L10-L15
Useful? React with 👍 / 👎.
| if reset.is_err() | ||
| && self | ||
| .transition_committed( | ||
| &PendingAudit { | ||
| outcome: DeliveryAuditOutcome::ReplayCommitted, | ||
| ..replay.clone() | ||
| }, | ||
| None, | ||
| ) | ||
| .await | ||
| == Some(true) |
There was a problem hiding this comment.
Distinguish another replay from this reset
When this replay's reset rolls back or fails before commit, a concurrent replay waiting on the same row can subsequently commit the identical next_generation before transition_committed reads it; this branch then reports the first replay as committed too, even though the API normally rejects the second use of an expected generation. Fresh evidence beyond the generation-only reconciliation fix is that another replay—not only this request—can write that replacement generation, so reconciliation needs a request-specific durable marker rather than generation equality alone.
AGENTS.md reference: products/platform/AGENTS.md:L29-L31
Useful? React with 👍 / 👎.
3815eb4 to
795f5cc
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 795f5cc221
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| state.in_flight -= 1; | ||
| state.in_flight == 0 && state.dropped | ||
| }); | ||
| if owes_unfinished { |
There was a problem hiding this comment.
Wait for the actual write before emitting unfinished
When runtime shutdown cancels the async response task after its stream or file write has started, that underlying blocking write can still finish successfully, but this guard decrements in_flight and allows the request handle to emit unfinished. The journal then contains both the real terminal response and an unfinished response for one request. Fresh evidence after the new drop-guard fix is that bookkeeping is still tied to the cancelable async task rather than the cancellation-surviving write operation.
AGENTS.md reference: products/platform/AGENTS.md:L29-L31
Useful? React with 👍 / 👎.
| match self.expired_evidence_remains(cutoff).await { | ||
| Some(false) => Ok(erased), |
There was a problem hiding this comment.
Tie erasure recovery to this transaction
When this erasure's commit rolls back after returning an ambiguous error, a concurrent erase-expired invocation can delete the same rows before expired_evidence_remains runs; Some(false) then attributes the second invocation's deletion to the first and records the first request as erased with its speculative count. Fresh evidence beyond the prior read-back fix is that the recovery predicate is not transaction-specific; use the original transaction status or a request-specific durable marker instead.
AGENTS.md reference: products/breg/AGENTS.md:L175-L182
Useful? React with 👍 / 👎.
| outcome: DeliveryAuditOutcome::ReplayRequested, | ||
| disposition: DeliveryAuditDisposition::ReplayPending, | ||
| }; | ||
| self.seams.record_audit(replay.record()).await?; |
There was a problem hiding this comment.
Audit replay requests before protected lookup
When an operator replay names an absent, stale, forbidden, expired, or nonterminal delivery, replay_in has already opened a transaction and read protected delivery state but returns before this request entry is written, so there is neither a request nor the newly promised replay_refused response. This also makes the newly enforced BREG-V1-27 claim false for replay refusals; begin a minimized replay audit before the lookup and answer it on every refusal path.
AGENTS.md reference: products/platform/AGENTS.md:L29-L31
Useful? React with 👍 / 👎.
| `request-retention erase` deletes the external attachment objects before it records its response, | ||
| which states how many objects still wait for deletion; when that deletion pass itself fails, the | ||
| count is `null` and the command reports the failure. If the erasure committed but its response |
There was a problem hiding this comment.
Document the actual erasure audit order
The implementation now records the committed erasure response before calling retry_external_deletions, and that response contains no pending-object count; if the deletion pass fails, the response is already durable and the command returns the retry error. These lines state the opposite ordering and promise a nullable count, which can mislead operators investigating a destructive partial cleanup. Update this procedure to match RequestRetentionOperatorService::erase (and the new test that explicitly pins commit recording before retries).
AGENTS.md reference: docs/site/AGENTS.md:L14-L16
Useful? React with 👍 / 👎.
| if let Some(unfinished) = unfinished { | ||
| writer.append_detached(&AuditEntry::response(key.0, key.1, unfinished)); |
There was a problem hiding this comment.
Await unfinished responses before releasing failures
When an audited operation exits through an early ?, dropping its handle only calls append_detached; the caller future can therefore return its HTTP or CLI failure before the file response is fsynced or the stream response is written. If that detached write subsequently fails, the failure has already been released and only an error is logged, contradicting the new fail-closed, response-before-release guarantee for the paths that rely on this unfinished backstop. These paths need an awaited cancellation-surviving completion, or the product contract must not claim durable acceptance before release.
AGENTS.md reference: products/platform/AGENTS.md:L29-L31
Useful? React with 👍 / 👎.
| *lifecycle_attempt = Some( | ||
| begin_maintenance_request( | ||
| request.audit, | ||
| lifecycle_request_entry(request, &lifecycle_reference)?, | ||
| ) | ||
| .await?, | ||
| ); |
There was a problem hiding this comment.
Keep progress responses from closing the lifecycle handle
When this lifecycle scrubs any request snapshots, the later request_scrub_entry is a response with the same FIELD_ENCRYPTION_AUDIT_SCHEMA and lifecycle correlation, so AuditWriter::append immediately marks this new handle answered after the first scrub transaction. If a subsequent record erasure or rebaseline then fails or is canceled, dropping lifecycle_attempt writes no unfinished response, leaving only an intermediate phase: "request-scrub" entry rather than the terminal or unfinished lifecycle outcome promised by this change. Use a distinct correlation/schema for progress entries or keep them from settling the lifecycle request.
AGENTS.md reference: products/breg/AGENTS.md:L175-L182
Useful? React with 👍 / 👎.
Add AuditWriter::begin, which appends the request entry and returns an AuditRequest handle that owes its response. The handle responds in the request's own schema and correlation, and a response appended through AuditWriter::append under the same schema and correlation answers it too. A handle dropped unanswered writes the product's unfinished record as the response, so an early return, a panic, or a canceled future still pairs the request entry. For the file destination that line is queued into the group commit and, if the runtime shuts down first, written when the last reference to the file is dropped. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
Hold the shared AuditRequest for each caller-requested operation, so a
refusal, a failed commit, or a canceled request that returns after the
request entry still writes {event, outcome: "unfinished"} as its response
under the same correlation. Audit add_review_note, which wrote no entries:
it now records who added a note and its history event, never the note's
text or audience.
Refs #1576
Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
An operation that has no response record to append still withholds its result, and now writes its unfinished outcome as the response, so its request entry is no longer left unpaired. Refs #1576 Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A failed capacity transaction, records replaced under a commitment, and a reused or expired idempotency key wrote a request entry and no response. They now write a response with the outcome unfinished and a closed reason (commitment.failed, commitment.facts-stale, idempotency.key-reused, idempotency.expired), never an authorization verdict. The request is held through the shared AuditRequest handle, so a commitment that returns or is canceled before it answers writes commitment.unfinished. A refusal, from the ledger or the permission check, is now answered only once its response entry is accepted, and service.unavailable otherwise, instead of logging the write failure and answering the refusal anyway. SCHEDULING-SEC-14 and the runtime configuration reference say so. Refs #1582 Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
The audit correlation was the caller's Idempotency-Key, which two calls may share, so their request and response entries could not be told apart. Every call now draws its own correlation and the caller's key is recorded only as correlationId. The request entry is held through the shared AuditRequest handle, so a call dropped before its outcome, such as a disconnected caller, writes an unfinished response. The answer is built before the response entry is written, so a rendered entry is never recorded for a document that could not be sent. Closes #1575 Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
Every attempt is held through the shared AuditRequest handle until its terminal or refusal entry answers it; a request that ends first, through an error, a timeout, or a dropped caller, writes an unfinished response under the same correlation. Reads that fail after their rows were read write the Refused terminal and log a terminal the destination refuses. Plain record_pre_io_audit now accepts only refusals, so an attempt can only be written through the handle. - A reviewed apply records its attempt before the receipt preflight's reads and the review authority, and holds it across the action. - Migration reconciliation answers a transition that fails after its request entry with a failed response. - Request detail erasure answers a refused or failed erasure, deletes external attachment objects before its response and records how many remain, and reports an erasure that committed without its response as request_retention.erasure.unaudited. Attachment cleanup is held too. - Evidence retention erasure is audited under breg-evidence-retention-audit/v1 with its cutoff and erased count. - An ingestion call refused after its ingestion request entry is answered in the ingestion schema and not recorded again as a general refusal. Refs #1592 #1587 #1597 #1590 #1593 #1507 Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
The delivery worker appended terminal and expiry entries inside the transaction it was about to commit, so a failed commit left an entry for a transition that never happened. Terminal dispositions and payload expiries are now recorded after their transaction commits. An attempt's request stays ahead of its lease commit, so it is on record before egress; if that commit fails the worker answers it with worker_interrupted. An operator replay is a replay_requested request before the reset and a replay_committed or replay_refused response after it. The seam no longer receives the transaction. This changes the shared seam, so Scheduling's hook audit follows the same order. Refs #1592 #1588 Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A mutation that stops after its attempt now answers it with an unfinished response, so each fault leaves two audit entries. Refs #1593 Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
Every caller-requested operation now answers its audit request entry, so the row returns to enforced with the fault-injection tests that prove the paths its gap listed. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
History rebaseline, a standalone history erasure, and a field-encryption erase-and-rebaseline run appended their request entry and then returned through refusals and database errors without a response. They now hold the shared request handle through the run, so an early end writes an unfinished response under the same correlation. The attachment verification worker holds its attempt the same way, which also pairs a job its time budget cancels. Refs #1592 Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
Review findings on the AuditRequest handle: - A caller canceled while its request entry was written could leave the request accepted with no handle to answer it. The write and the handle's registration now run in one task that outlives the caller, and a caller that stopped waiting drops the finished handle, which answers it. - A response accepted after its caller was canceled was not marked as answering its request, so the drop wrote a second, false unfinished response. append and respond now record the answer inside the task that writes it; a drop while a response is in flight leaves that response to settle the request, and writes the unfinished record only if it is refused. - The unfinished entry is validated before the request is written, so a request is never accepted with a response it could not write. - A stream destination writes a dropped request's unfinished entry on a dedicated thread, joined when the writer is dropped, instead of blocking a runtime thread on a stalled stream. wait_for_detached_entries lets tests and shutdown paths read after a drop; the product test captures use it. Refs #1561 Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A PostgreSQL commit error does not prove the transaction rolled back. After a failed commit the worker now reads the delivery state on a fresh connection: a terminal disposition or replay reset that did commit is recorded, and one that rolled back is left to expiry recovery or recorded as refused. A lease commit whose fate cannot be read is answered as interrupted, since a second interrupted answer is harmless and a missing one is not. Refs #1592 #1588 Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
An operator replay now runs to its response in a task of its own, so a caller that times out or disconnects after the replay_requested entry is accepted still leaves that request answered. A reset whose commit acknowledgement was lost is recognized by its replacement generation, which only a replay writes, rather than by the pending state the worker may already have moved it out of. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A request-detail erasure whose commit returned an error now reads the detail back on a fresh connection and answers its request with the erasure, failed, or unfinished, instead of assuming a rollback. A replayed ingestion chunk builds its answer inside the release transaction, before its disclosure entry, so an accepted disclosure is never followed by a failure that returns no receipt. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A response appended through AuditWriter::append now claims the oldest open request under its schema and correlation before its write starts, so a handle dropped while that response is written leaves the request to it instead of also writing its unfinished record. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…rors An Evidence retention erasure whose commit returned an error now checks on a fresh connection whether expired material remains, and a reconciliation transition that returned an error reads the maintenance state back. Each answers its request with the durable outcome, failed, or unfinished when that state cannot be read, instead of assuming a rollback. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A request dropped outside a Tokio runtime while the file state is busy now waits for the lock to queue its unfinished response, instead of discarding that response after a bounded number of attempts. No holder keeps the lock across an await, and the file's drop or the next group commit writes the queued line. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A scheduled retry whose commit acknowledgement was lost is recognized by a later attempt of the same generation, and a dead letter by a later replay generation, rather than only by the state the worker may already have moved on from. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…down A response claimed between a handle's in-flight check and its close no longer lets the drop write unfinished as a second answer, and a response task dropped by a runtime shutdown releases its in-flight count so the handle still pairs the request. A poisoned open-request map and a detached line lost at shutdown are logged. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…rs it Preflight dropped the attempt handle when it returned, so an admitted Evidence action was answered unfinished while evaluation still ran and answered again by its terminal or refusal. The admission now carries the held attempt to finalize, and the tests assert every audit request has exactly one response. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
The response of a committed request-detail erasure waited on up to thirty seconds of external-deletion retries, so a slow backend or an interrupted process left the committed erasure answered unfinished. The response is now recorded once the commit is confirmed, and the operation result alone reports the external deletions still pending. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…fails A run creation, cancellation, blocking, or receipt release whose commit returned an error answered its request entry refused, although the transition may have committed. Those commit errors now answer the request unfinished. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…nown A replay whose reset changed no row is refused without a read-back, so a generation another replay wrote is never taken for its own. A claim whose lease commit failed records each durable recovered or expired delivery whatever the lease's fate, and a finalize whose terminal commit cannot be read back answers its attempt as interrupted. The read-backs wait a short bounded backoff between attempts, and a replay read-back accepts any later generation. Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
… caller is canceled Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…e recording it unfinished Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…concerns by its pseudonym Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
… lost before recording it Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
… a request ends early Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
795f5cc to
3dd1be2
Compare
There was a problem hiding this comment.
💡 Codex Review
When replay A's reset rolls back after an ambiguous commit and concurrent replay B commits the same or a later generation before A's read-back, this >= returns true for A, so A emits replay_committed and succeeds despite not performing the reset; both callers can therefore report success. Fresh evidence in the current version is that reconciliation now explicitly accepts a later replay's generation, so the recovery check needs a request-specific durable marker rather than generation monotonicity alone.
AGENTS.md reference: products/platform/AGENTS.md:L29-L31
When deliver_once is canceled while this audit append is waiting for fsync, AuditWriter::append can still durably accept the AttemptStarted request in its spawned task, while dropping claim rolls back the lease transaction. Because no lease then exists for expiry recovery and this direct request entry has no AuditRequest guard, no worker_interrupted response is ever emitted, leaving the request permanently unpaired; retain a cancellation-safe handle or run the request, commit, and response bookkeeping in a task that survives its caller.
AGENTS.md reference: products/platform/AGENTS.md:L29-L31
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Pull Request
Summary
Closes WP35 of #1561 (audit request/response pairing). Every caller-requested operation now writes one audit
requestentry and at least oneresponseentry under the same correlation (PR #1560 decision D2). BREG-V1-27 returns toenforced.Base:
jeremi/audit-simplification.Closes #1576
Closes #1582
Closes #1592
Closes #1585
Closes #1588
Closes #1589
Closes #1597
Closes #1587
Closes #1590
Closes #1593
Closes #1575
Closes #1507
Mechanism: one shared request handle in
registry-platform-auditAuditWriter::begin(schema, correlation, request, unfinished) -> AuditRequestappends therequestentry and returns a#[must_use]handle that owes theresponse:AuditRequest::respond/finishappend a response in the request's own schema and correlation, awaited and fail closed.responseappended throughAuditWriter::appendwith the same schema and correlation also answers the oldest open handle. This lets products whose terminal entry is built far from the attempt keep that code: they only hold the handle for the operation's lifetime.unfinishedrecord as the response. That covers an early?return, a panic, and a canceled future (a request timeout or a disconnected caller). Stream destinations hand it to a dedicated writer thread that the writer joins when it is dropped, so a stalled stream never blocks the dropping task. The file destination queues it into the group commit and spawns the flush. If the runtime shuts down first, as in a one-shot operator command,Drop for GroupCommitFilewrites the queued line synchronously, unless a blocking write is still in flight.beginwrites and registers the request inside a spawned task, andappend/respondsettle their bookkeeping there too. A caller canceled while its request or response entry is being written therefore still gets exactly one answer, and a response accepted after the caller left is the only one. Theunfinishedrecord is validated before the request is written, so a request that could never be answered is refused up front.No config knobs, metrics or CLIs were added.
Why this shape. A closure or scoped API would need every BReg coordinator restructured around its commit. It would still need a drop guard for cancellation, which is the #1507 / #1563 case. Panicking in
Dropis wrong: cancellation is legitimate, and a panic during unwind aborts. The drop write cannot be awaited, so it cannot fail closed. Known outcomes (refusals, failed transitions) are therefore written explicitly and awaited; the drop path is the backstop for exits no code path names. A stopped writer refuses the drop write too, which is the existing "writer refused; repair and restart" state that readiness reports.#1507, settled consistently with #1563. A timed-out or canceled request's accepted attempt is answered by exactly one
unfinishedresponse when its handle drops. It is never lost unless the writer has already stopped, and never duplicated, because an answered handle writes nothing on drop. This is independent of the read path's connection cancel guard that #1563 tracks; that change needs nothing from this mechanism.Call-site audit
"Before" is the base branch. "After" is this PR.
AuditWriter::append(request)begin/AuditRequestCaseworkAudit::beginat 27 sites:store.rs(10),review.rs(6),assignment.rs(4),clocks.rs(2),task_grants.rs(4),source_retention.rs(1)?or refusal exit beforecomplete, plus the internal failures ofcommit()(invalidations,ensure_terminal, the database commit), left the request unpaired.add_review_notewrote no entry at all.AuditOperationholdsAuditRequest, and the unfinished record is{event, outcome: "unfinished"}.add_review_noteis audited ascasework.review_note_added.audit_requestat 6 commitment sitesrecord_refusalswallowed a failed append (fail open). Cancellation left requests unpaired.unfinishedresponses with closed reasons (commitment.failed,commitment.facts-stale,idempotency.key-reused,idempotency.expired). Refusals fail closed (service.unavailable). The handle covers everything else (commitment.unfinished). SEC-14 is revised.mutation.rs(2),mutation/action.rs(4),mutation/request.rs(2),postgres/read.rs(2),revision_read.rs,history_read.rs?after the attempt (bindings, transactions, fault points, timeouts, the evidence preflight). Post-read processing failures (#1593). Discarded terminal errors.begin_pre_io_audit/begin_action_pre_io_auditreturn the handle. Reads write the Refused terminal and log refused terminals.record_pre_io_auditnow refusesAttempt.request_reviewed_apply)attempt_recorded = true)migration_reconcile(Completable / Revertible)failedresponserequest_retention::eraserefusedorfailedresponse; external deletions before the terminal, with pending and tombstone counts;ErasureUnaudited→request_retention.erasure.unauditedrequest_retention::cleanup_attachmentsfailedappend_run_request(create, cancel, chunk replay, chunk receipt)breg-audit/v2refusal (#1597)RunAttempt; the service answers inbreg-ingestion-audit/v1(refused) and returnsIngestionRefusal { answered }, so the handler skips the general refusal. A fresh chunk's refusal belongs to the batch mutation's own v2 pair and is no longer recorded twice.evidence-retention erase-expiredbreg-evidence-retention-audit/v1: a request naming the cutoff and a response with the erased count orfailed, through thebregctlcompanion destinationregistry-platform-hooks)worker_interruptedif that commit fails. Terminals and expiries after commit. Replay: areplay_requestedrequest, then areplay_committedorreplay_refusedresponse.history rebaseline, standalonehistory erase, field-encryption erase-and-rebaseline lifecycle, attachment verification workerCoverageCompleteorNoPendingPlaintextHistory), database errors, and the worker's 180-second budget returned without a responsebegin_maintenance_requestholds the request through the run, and the lifecycle's request is held througherase_field_encryption_history. An early end answersunfinished.render_routeIdempotency-Key(#1575). A dropped call was unpaired. Arenderedentry could precede a failed response build.correlationId); the handle writesunfinished; the response is built before therenderedentry?exits.Evidence
Environment: PostgreSQL 17.x with PostGIS in this container (SSL off, port 5433), one disposable database per suite. Every PostgreSQL suite below ran against that database; none skipped for a missing URL. Base
jeremi/audit-simplificationat196affa.cargo fmt --checkcargo check --locked --workspace --all-targetscargo clippy --workspace --all-targets -- -D warningscargo test --lockedfor every workspace member-p <pkg> --no-fail-fast, binaries deleted between packages) rather than as one--workspacebuild. The failures are sixregistry-platform-auditpermission tests,registry-caseworkctldev::tests::config_change_keeps_the_owner_after_created_container_cannot_be_saved, andregistry-evidencectlaccess::revoke_leaves_the_record_untouched_when_the_private_directory_cannot_be_removed. Each relies on a0o500/0o200path that root ignores. The platform-audit ones fail identically on the base, and the other two are in crates this PR does not change. The platform-audit and caseworkctl ones pass (91/91 and 88/88) when rerun as an unprivileged user.products/breg/scripts/test-postgres.sh --lane postgrestargetsevidencebinary, exceptpostgres_startup::prepared_server_wires_services_and_static_jwks_readiness_tracks_database. That test writes{}\nas a "completed" companion audit file, and the base's own current-format check (#1581) refuses it. The test and that check are unchanged here.--lane immediate-actions(postgres_immediate_actions)postgres_history_rebaseline,postgres_history_erasure,postgres_history_migration,postgres_field_encryption,postgres_change_requests,cargo test -p registry-bregctl)review_postgres::cancellation_and_final_decision_commit_exactly_one_terminal_result, which fails deterministically on the base head without these changespostgres_commitments,schedulingctlintents_postgresandrecords_apply_postgrescargo test -p registry-renderpython3 products/breg/scripts/validate_product.pyandtest_validate_product.pyproducts/breg/scripts/check-contracts.shproducts/casework/scripts/check-checkpoint.shproducts/scheduling/scripts/check-checkpoint.shandcheck-contracts.shnpm test;check:markdown,check:style,check:evidence-anchors,check:contentcheck:evidence-linksnpm run checkbuild was not run for lack of disk.bregctldiagnostic is failure-report content onlyFault-injection proof. These tests were run against the unfixed code first and failed: the five new Scheduling tests, the Casework review-note test, and the Render
servepairing test. The Render concurrency test and the webhook terminal ordering were also mutation-checked: the fix was reverted locally, the test failed, and the fix was restored. The existing tests that pinned the old unpaired behaviour ("retains only its attempt", "a refused rebaseline records its request and no committed response") failed once each fix landed and were updated to the paired expectation. The remaining new BReg fault tests (reconcile, retention, ingestion, Evidence retention, strong-ETag read) were written after their fixes. Each asserts a response entry that the base code, on reading, never writes, but they were not executed against the base.Notes
Shared-seam change
registry-platform-hooks::DeliverySeams::record_auditno longer receives the transaction, and the worker now calls it after commit for terminals, expiries and replay outcomes. This also changes Scheduling's hook seam (crates/registry-scheduling/src/hooks.rs). One behaviour changes: a terminal entry the audit destination refuses now leaves the committed disposition in place, and the writer stays stopped, instead of rolling it back and redelivering after lease expiry. That matches how Scheduling already treats a response refused after commit.Security review notes (audit integrity is security-sensitive)
AuditRequestdrop writesunfinished; explicit awaited responses for known outcomeswriter::tests::a_request_dropped_unanswered_writes_its_unfinished_response,a_canceled_operation_pairs_its_request_in_the_file,a_panicking_operation_pairs_its_request,a_command_that_exits_after_an_early_return_pairs_its_request,a_response_in_another_schema_does_not_answer_the_requestAuditOperationholds the handlepostgres_transactions::a_refusal_after_the_request_entry_pairs_it_with_an_unfinished_response,audit::tests::a_requested_operation_without_a_response_record_is_refusedadd_review_notebegin/commit; text and audience are never recordedreview_postgres::review_notes_are_audited_without_their_textrecord_refusalreturnsservice.unavailablea_refusal_whose_response_entry_is_refused_answers_service_unavailable,a_permission_refusal_the_destination_refuses_answers_service_unavailableunfinishedresponsesa_failed_capacity_transaction_pairs_its_request_entry,a_records_swap_under_a_commitment_pairs_its_request_entry,an_idempotency_key_refusal_pairs_its_request_entrypostgres_read::real_postgres_read_is_authorized_bounded_minimized_and_audit_gated(StrongEtagandBeforeTerminalAuditfaults)postgres_change_requests::cached_review_result_cannot_authorize_fresh_apply_but_committed_receipt_recovers_offlinefailedresponsepostgres_migration::real_postgres_reconciliation_completes_reverts_or_refuses_a_pinned_target(a heldregistry_staterow fails activation)ErasureUnauditedpostgres_request_read_retention::request_detail_erasure_pairs_its_request_entry_on_every_outcome,postgres_request_upgrade_retention::operator_retention_service_counts_pages_erases_under_forced_rls_and_audits,request_retention::tests::an_unaudited_erasure_stays_distinct_from_a_refused_onepostgres_action_evidence_retention::expired_request_evidence_erases_only_retained_uses,retention_refuses_misbound_database_with_identical_roles_and_catalog_driftrefusedresponse; handler skips the v2 refusalpostgres_ingestion_runs::a_refusal_after_the_ingestion_request_is_answered_in_the_ingestion_schemaunfinishedpostgres_history_rebaseline::rebaseline_refuses_while_maintenance_is_not_readyandrebaseline_restores_snapshot_coverage_from_current_state_after_an_erasure;postgres_history_erasure::field_encryption_erasure_uses_flip_provenance_for_structured_plaintextpostgres_webhook_delivery::real_postgres_webhook_delivery_retry_dead_letter_replay_is_package_bound_audited_and_confined(deferred-trigger commit failures)server::tests::concurrent_calls_sharing_an_idempotency_key_pair_their_own_entries,serve::a_render_writes_a_request_entry_then_a_response_entry_sharing_correlationWhich of these tests were run failing first is stated under Evidence.
Review fixes
The Codex review raised five findings, and all five are fixed:
beginorappend: a request is now paired even when its caller is canceled mid-write (a_request_canceled_while_its_entry_is_written_is_still_paired,a_response_accepted_after_its_caller_left_is_the_only_answer,a_response_appended_after_its_caller_left_still_answers_the_request).unfinishedrecord: now refused before the request is written (an_unfinished_record_too_large_to_write_refuses_the_request).dropping_a_request_never_waits_on_a_stalled_stream).postgres_webhook_delivery. The committed-but-unacknowledged branch cannot be reproduced against a local PostgreSQL, so it has no dedicated test.A second Codex review of
5de578fraised four more findings, and all four are fixed:postgres_webhook_deliveryfails with that task removed.a_replay_whose_reset_committed_holds_after_the_worker_moves_it).failed, orunfinished. The refused-commit case is inrequest_detail_erasure_pairs_its_request_entry_on_every_outcome.After these fixes,
postgres_webhook_delivery,postgres_request_read_retentionandpostgres_ingestion_runspass, as do theregistry-platform-hookstests (149), fmt, and clippy on the changed crates.A third Codex review of
09106c9raised three more findings, and all three are fixed:AuditWriter::appendnow claims the request before its write starts (a_request_dropped_while_an_appended_response_is_written_is_answered_once, which fails with the claim removed).failed, orunfinishedwhen that state cannot be read.expired_request_evidence_erases_only_retained_usesadds a refused commit.The other request paths here already record an unknown outcome as
unfinished. A BReg mutation's terminal entry is written before its commit and gates it, which is the base design.After these fixes, the BReg targets
postgres_action_evidence_retention,postgres_migration,postgres_readandpostgres_mutationpass, as do the Casework and Scheduling PostgreSQL suites (apart from the known Casework failure below), theregistry-platform-audittests apart from the root-only ones, and the Casework and Scheduling checkpoint and contracts scripts.A fourth Codex review of
8ee314fraised two more findings, and both are fixed:a_request_dropped_outside_a_runtime_while_the_file_is_busy_still_pairs, which fails without the fix).a_committed_retry_holds_after_its_next_attempt_is_claimed,a_committed_dead_letter_holds_after_it_is_replayed).After these fixes, the
registry-platform-audittests (apart from the root-only ones), theregistry-platform-hookstests (151),postgres_webhook_delivery, Schedulingpostgres_commitmentsand clippy pass.Follow-ups (out of scope, not fixed here)
review_postgres::cancellation_and_final_decision_commit_exactly_one_terminal_resultfails deterministically on the base head without these changes (both the cancel and the decision succeed).submit_ingestion_chunkperforms run and binding reads before its ingestion request entry (the same class as BReg reviewed-apply preflight runs database reads before the request audit entry #1587).list, dry runs, ingestion list and read) and webhook expiry terminal duplicates.?exits; Relay: cancellation.registry-platform-auditwriter tests andcaseworkctl dev::tests::config_change_keeps_the_owner_after_created_container_cannot_be_savedfail when run as root and pass as an unprivileged user.request_store::load_postgres_envstill reads a hard-coded/private/tmp/...path. Tests here setBREG_TEST_DATABASE_URLand do not rely on it.DCO
Signed-off-bytrailer.🤖 Generated with Claude Code
https://claude.ai/code/session_01YcLEc9VkbA2p5WSEZXugQz
Generated by Claude Code