Skip to content

refactor(audit)!: replace chained journals with shared JSONL streams - #1560

Merged
jeremi merged 92 commits into
mainfrom
jeremi/audit-simplification
Sep 27, 2026
Merged

jeremi merged 92 commits into
mainfrom
jeremi/audit-simplification

Conversation

@jeremi

@jeremi jeremi commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Summary

BReg serialized audit writes through its PostgreSQL chain head, and the products maintained separate chain implementations and outbox publishers. This change moves BReg, Evidence, Relay, Render, Casework, and Scheduling onto one per-process JSONL AuditWriter.

The envelope is {schema,eventId,time,phase,correlation,record}. File acceptance follows group-commit fsync; stdout is explicitly best effort. Files have exclusive writer locks, private permissions, bounded rotation/retention, and companion CLI paths. Product record minimization and keyed identifier hashes are preserved.

#1519 has merged, and this branch is rebased onto current main. The branch has linear, signed-off history. The detailed security review is retained privately. The response-contract decisions are resolved as described below; the PR stays in draft until the maintainer undrafts it.

Review fixes

A staff review of the previous head found three P1 gaps against the approved contract (a request entry accepted before protected I/O, and a response entry accepted before any result is released). Each fix below has PostgreSQL tests that failed before the change and pass after it.

  • BReg pre-effect request entries. Ingestion-run creation, cancellation, and blocked submission; history erasure and rebaseline; request-detail erasure; migration reconciliation; and the field-encryption lifecycle now accept a request entry before any change. A refused writer leaves the database unchanged. These are registered as BREG-SEC-102 to BREG-SEC-108, and BREG-SEC-45's refusal text now describes the post-commit gap instead of atomic audit.
  • Scheduling replays. Every replayed receipt, including one a concurrent identical request won, is released only after a response entry that records the decision the receipt carries. If the entry is refused, the answer is service.unavailable. SCHEDULING-SEC-14 now covers replays.
  • Casework replays and no-ops. A caller-requested operation that records no domain event now appends one minimized {event, outcome} response (replayed or unchanged) before its result is returned. An operation with neither a domain event nor an outcome is refused before commit. This is registered as CASEWORK-SEC-22.
  • A Casework template-retirement test that injected its failure through the removed casework_audit_outbox now injects it through review history in the same transaction.

Response-contract decisions, resolved:

  • Non-idempotent recovery. An operation with no idempotent replay recovers operationally. When ingestion-run creation's response entry is refused after commit, the caller gets service.unavailable without the run id, and finds the run by listing open runs for its inputDigest. Operator reruns of erasure, rebaseline, and reconciliation find the committed state instead of replaying it. This is documented in INGESTION-RUNS.md, DEFINITION-OF-DONE.md, and the operator pages. A scoped recovery contract is tracked in Make BReg ingestion-run creation replayable after a refused response audit entry #1570.
  • Response cardinality. A caller-requested operation writes one request entry and at least one response entry sharing its correlation (BREG-V1-27). Casework writes one response entry per recorded event. Background work no caller requested writes only response entries and starts only while the writer is ready. Evidence writes one access entry per source stage.
  • Contract text. Remaining text that implied atomic or chained audit is aligned (HISTORY.md erasure coverage, the BReg API and retention pages, and Casework's runtime configuration and retention pages).
  • In-flight reads. BREG-SEC-03 now states that a read linearizes at query time. A result materialized under the active package may still be released after a later activation or maintenance start, because the response entry is not a second fence.
  • Attachment verification (Codex P1). The worker now accepts its terminal entry before the verdict commits. A refused entry rolls the verdict back, and the job is retried once its lease expires. A PostgreSQL test failed before the change and passes after it.
  • Companion commands with destination: stdout (Codex P2). Companion processes (bregctl, caseworkctl, schedulingctl) now write audit to stderr, so a --format json report keeps stdout. Configuration cannot select stderr; for_process derives it from stdout. The public bregctl --format json test lifecycle test runs against a TLS PostgreSQL with a stdout destination, failed before the change with trailing characters after the JSON report, and passes after it.
  • Audit stdout destination blocks Tokio workers when the collector stops reading #1566 and Audit rotation can reuse a sealed segment name after retention deletes it #1567 do not block merge.

Codex review of the rebased head:

A later Codex finding, that Evidence audit show --last-operation ignores an earlier unmatched attempt, is pre-existing on main and tracked in #1569.

CI on the rebased head also failed in the BReg tutorial job, because dev_lifecycle.rs still called the removed bregctl audit verify, and in Docs checks, because of an unintroduced "BReg" on the API stability page. Both are fixed. The lifecycle test passes locally, and the full docs check passes locally.

CI on the rebased head then failed in Casework PostgreSQL transactions. The native-exchange ThunderID test still wrote the retired Evidence auditStorage block. The test now writes audit.path, and both ignored ThunderID exchange tests pass locally the way CI runs them.

Review of the audit-pairing follow-up commits

A five-reviewer pass over the follow-up commits found no blockers. The pairing contract is now stated as "every audited request entry gets at least one answer while the process runs": a second answer is tolerated, and a process that stops with a write in flight can leave a request unanswered (docs, CHANGELOGs, and hooks comments are scoped accordingly). Fixed here, each with a test that failed first unless noted:

  • The audit writer no longer lets an appended response claim a request whose own response is still in flight.
  • The writer refuses an audit file name too long to leave room for the rotation suffix, requires read, write, and search permission on the ancestor it creates directories in, and fsyncs the parent of each directory it creates (Codex P2 threads). Requiring read permission is new: the parent sync must open that ancestor.
  • A hooks finalize canceled after its disposition commits still records its terminal entry.
  • A BReg ingestion transition that reached its commit is answered unfinished, never refused, when its commit errors or its response entry fails.
  • Scheduling re-checks the task grant's expiry with no database round trip before the commit (restores the pre-change ordering; not observable through a pinned clock, so no new red test).
  • The Evidence retention test now checks pairing on the journal that records it, and the BReg pairing helper fails on an empty journal.

Tracked as follow-ups: commit-fate read-backs after a lost COMMIT acknowledgment (#1604), unfinished answers at runtime shutdown (#1605), and stronger pairing assertions across BReg PostgreSQL tests (#1606).

Evidence

Rebased head

These checks ran on the rebased head with the review fixes. Only the two Codex fixes above came after them, and those were rechecked with cargo fmt --check, clippy (platform-audit, and Casework with postgres-test), the platform-audit, caseworkctl, and evidencectl tests, the BReg startup and runtime-config targets, and the Casework native-exchange tests.

  • Workspace tests: 7,017 passed, plus the same eight known macOS installer fixture failures described below.
  • PostgreSQL:
    • Scheduling passed 164 tests.
    • Every Casework target passed except the five installer fixtures.
    • BReg test-postgres.sh --lane all, run with no-fail-fast across 70 binaries: 661 passed, 11 failed, 10 ignored. All 11 failures are HTTPS loopback delivery tests (postgres_webhook_delivery, postgres_hook_proposals, and one request-lifecycle webhook case). postgres_webhook_delivery fails the same way on current main on this Mac, and CI's BReg PostgreSQL contracts job passes on this branch. I treat this as a local environment fault, not a regression, but the cause is not yet identified.
  • Gates: fmt, clippy with the PostgreSQL features, and the BReg, Scheduling, Relay, and identifier contract checks all passed. On macOS, the Relay contract check needs relayctl launched through a wrapper that sets DYLD_FALLBACK_LIBRARY_PATH for the aws-lc FIPS library, because SIP strips DYLD_* through /usr/bin/env.

Previous head

Evidence below is for the previous head 9467174eedd039cba5b773f0d47092b69c5da927, before the rebase and review fixes. Checks below used the pinned toolchain, locked dependencies, and disabled incremental compilation. No third-party dependency versions changed.

  • Rust checks: cargo fmt --check and cargo clippy --locked --workspace --all-targets -- -D warnings passed. cargo test --locked -p registry-cli-docs -p registry-language-server passed, including doctests. cargo deny check passed all four policy classes.
  • Workspace tests: the final serial cargo test --locked --workspace --no-fail-fast run completed with 6,978 passed and eight failed. All failures are the existing BReg (three) and Casework (five) macOS installer fixtures, unchanged from 9a05d3688: their v9.8.7 fixtures supply raw binary assets while the v0.33+ installers require FIPS tar archives. They fail at asset lookup before the intended installer fault injection. An isolated rerun reproduced the BReg failures, and independent source review confirmed the shared cause. All remaining executed tests and doctests passed; earlier concurrent-run rustdoc artifact errors did not recur.
  • Writer regressions: all 65 platform-audit tests passed. Controlled tests failed before and passed after the accepted-write cancellation, queued-stdout refusal, and incomplete-tail startup fixes.
  • Database tests: products/breg/scripts/test-postgres.sh passed 648 tests across 68 test binaries. Casework's documented PostgreSQL targets passed 145 tests across 13 targets; Scheduling's three PostgreSQL targets passed 82. Disposable databases were provided; these were executed tests, not missing-database skips. BReg's ten explicitly opt-in cases remain listed as ignored; the separately requested S3 journey passed.
  • Product gates: BReg contracts/client-contract/source-neutrality; Evidence contracts/source-neutrality/verifier-portability/configuration parity/authoring schema/no-I/O; Relay contracts/authoring schema; Casework and Scheduling checkpoints; and Scheduling contracts passed. products/identifiers/scripts/check.sh passed 18 tests and canonical artifact reproduction; generate.py --check-references passed.
  • BReg live workflows: PostgreSQL TLS, S3 attachment, adopter, historical, and immediate-action scripts passed. The change-request example script failed at its first fixture because its unchanged runner omits the required casework review-authority binding. Source comparison with 9a05d3688 confirmed this baseline omission; this is a failed gate, not execution proof for that journey.
  • Bindings: BReg/Casework/Evidence Node tests passed 44/19/53, Python tests 38/30/51. Native builds, TypeScript checks, dependency installation, license comparisons, and sync-registry-client-node.py --check passed. macOS tests used the maintained Cargo runtime-library helper; exact Node test scripts were invoked directly to retain its FIPS library path. No binding source or rpath was changed.
  • Upgrade journeys: real v0.33-to-development Evidence, BReg, and Casework rehearsals passed, preserving records/review work and retired audit archives, then emitting current-schema entries. These used staged development binaries with separately recorded hashes, not a final-HEAD rebuild. The local override is labeled unverified by the generic tool; retained release archive checksums and previously authenticated provenance supplied the old-binary identity evidence. Scheduling state was not rehearsed.
  • Documentation: npm test passed 614 tests; npm run check passed, including 207,148 link/asset checks. Final prose corrections passed 604 evidence anchors, 1,884 paths, and 1,727 symbols; architecture checks passed all 12 models.
  • Tooling and final review: upgrade rehearsal tests passed 27 cases; load evidence tests passed 11; Casework harness tests passed 9 with its unavailable k6 integration explicitly skipped. Shell checks passed. gitleaks git --config .gitleaks.toml --log-opts=9a05d3688..HEAD --redact scanned all 19 commits with no leaks. All commits are signed off and history is linear.

The parent #1519 advanced to 575bc4db9 with an unrelated harness fix after this branch was based on 9a05d3688; a merge-tree check confirmed a clean combination.

The workspace reported 41 intentionally ignored tests: opt-in benchmarks, public demos, issuer/container and installed-binary lifecycles, fixture regeneration, and externally fetched schema checks. Their reasons remain in the test output. The requested real S3 and Evidence sustained-load tests were executed separately and passed; ignored entries are not counted as verification of the remaining journeys.

Directional measurement

The paired BReg runs use 100,000 seeded records, seed 20260902, the same workload, and interleave before-1/after-1/before-2/after-2. Before is 9a05d3688; after is 8821b3a67, including the writer cancellation fix. The later final-tail startup check does not change normal append behavior.

The before environment retained historical measurement mutations; after was freshly seeded. Nominal seed settings and identifier pools matched, but the accumulated database state was not reset identically. This is an additional limit on interpreting the measured ratios.

Counters are the primary evidence. Transaction counts include monitoring and background activity; the denominator is completed workload operations, including failed operations. Achieved operations/second is not successful-request throughput.

Database signal, four pairs Before After
Committed transactions / completed operation 3.154–3.287 1.712–1.834
Sampled audit-lock waiters, peak 3 0
Sampled audit-lock waiters, mean 0.066–0.127 0
Audit head / audit tables present yes / two no / none

The observed transaction ratio was 41.8–47.4% lower across pairs. Different failure and work-completion counts limit interpretation of the magnitude. All eight steady runs saturated and failed the unchanged harness SLOs. Completed rates were 44.933–45.220 operations/s, with 2,938–3,929 HTTP 504 responses per run and p99 near 10 seconds. This does not establish successful-throughput improvement.

The ignored release-mode Evidence sustained-load test passed at 5,461 requests/s, 10,923 audit appends/s, p50/p95/p99 20.07/44.82/64.75 ms, and zero failures. It used the actual temporary file-backed group-commit writer, not an injected sink. The source baseline was 130,722 requests/s, 23.9 times the gateway rate. Audit cardinality assertions passed.

Host: shared Apple M5 Max, 18 logical cores, macOS 26.4.1, Docker PostgreSQL. BReg start/end one-minute load ranged 5.59–36.05. Evidence load moved from 31.11/28.39/23.68 to 45.46/32.14/25.17 (1/5/15 minutes), with 17 and 8 unrelated rustc processes at its endpoints. Full tables and reproduction commands are in the BReg loadtest README and Evidence operator contract.

All timings are directional observations from a shared development machine, not capacity claims or absolute pass/fail gates. Host load and build overlap were recorded for every run. The optional sweep was omitted; the paired steady repeats are the comparison.

Discarded timing comparisons: before-1/after-1 and before-2/after-2 both had lower after rates while builds overlapped. Each pair was repeated once as before-1-retry/after-1-retry and before-2-retry/after-2-retry; both retries also overlapped unrelated compilation, so their timing evidence remains inconclusive and is excluded from improvement claims. All raw observations are retained. The BReg sampler counts build-command matches, including wrappers, rather than exact compiler processes; separate executable checks confirmed real compiler overlap. No runtime setting or threshold was tuned from these runs.

Notes

Ordering and recovery

  • BReg writes request audit before protected work, including ingestion-run transitions and every maintenance command. Mutation transactions commit before their response entries; read response audit gates disclosure. Maintenance emits a post-commit response entry. Webhook terminal and replay entries are written after the delivery state commits. Security review note: an audit refusal no longer rolls back a delivery transition, so terminal audit fails open, but the journal never names a disposition the database does not hold. Finalize runs in its own task, so a worker canceled after the commit still records the terminal entry while the process runs.
  • Evidence accepts an access entry before each source read and a release entry after signing, before returning assertion bytes. Product batch and refusal schemas remain explicit.
  • Relay accepts the attempt before source I/O and the terminal entry before releasing the held response bytes.
  • Render accepts the request before rendering and the rendered response entry before document disclosure.
  • Casework and Scheduling write runtime audit directly while domain history remains in PostgreSQL. Their capacity/domain mutations commit before response audit. A replayed or no-op result is released only after its own response entry is accepted. Background and maintenance events do not imply a universal request/response pair.

A crash or audit refusal after a mutation commits can leave an unmatched request entry. An audit refusal can return unavailable while the effect remains committed. Recovery must inspect the receipt/history and use the operation's supported idempotency key. BReg ingestion-run creation, cancellation, hook proposals, and maintenance commands are not uniformly covered by idempotent replay; blindly resubmitting is unsafe. A refused writer needs repair and process restart.

Security review note, commit resolution (review round of 2026-09-27):

  • An unproven commit is answered unfinished, never refused. Once a BReg mutation, batch, request action, immediate action, hook proposal, or evidence action reaches its commit, a commit error or a failed post-commit audit append drops the attempt guard, which records unfinished. Previously a lost COMMIT acknowledgement was journaled as a refusal for an effect that could be durable. This is deliberately conservative: a commit that clearly rolled back is also recorded unfinished. Callers still see unavailable.
  • The attachment verification verdict (BREG-SEC-57) commits before its terminal entry. A refused terminal entry now leaves the verdict committed with the attempt answered unfinished, instead of rolling the verdict back.
  • Field-encryption erase-history no longer writes a mid-lifecycle request-scrub response under the lifecycle correlation; scrub counts appear only in the terminal entry, so a failure after the scrub commits journals unfinished rather than a committed lifecycle.
  • Casework background work whose commit read-back is unknown now appends unfinished response entries for its events, with no preceding request entry.
  • The platform audit writer validates a directory level created concurrently by another process with the same safe-directory checks as a pre-existing one.

Security review note, writer contract (staff review and final review round, 2026-09-27):

  • The operator creates the audit directory. The writer refuses only a world-writable, non-sticky ancestor (a group-writable parent such as Ubuntu's root:syslog 0775 /var/log is accepted), and each refusal names the fix.
  • Sealed segments are <path>.<sequence>. <path>.seq records the next sequence before every rotation, so numbering stays continuous across restarts, including after a shipper removes every sealed segment. <path>.lock, <path>.seq, and <path>.seq.tmp are reserved; for_process refuses a role whose sibling would land on one (ProcessRoleOverlapsStream).
  • A shipper may copy and remove sealed segments while the writer runs. Retention deletes oldest-first by sequence, skips files with a .lock companion, and a retention failure is logged instead of stopping the writer (it costs disk, never an entry).
  • The writer fsyncs the file and its directory before acknowledging, including when it reopens an existing file.
  • The active-file identity check uses device, inode, length, and modification time; ctime is not compared, so chmod or ACL re-application does not trip it.
  • Startup refusals in BReg, Relay, Casework, Scheduling, and Evidence carry the writer's own recovery text. BReg adds StartupError::AuditDestination, reported by bregctl doctor as startup.audit.refused. Evidence's audit startup error kinds are Configuration, Secret, Storage, and File.
  • Hook delivery whose commit cannot be resolved is recorded with the unknown disposition (BReg and Scheduling), never as retry_pending for a row that may be terminal.

Follow-ups tracked instead of held on this PR: #1614 (torn last line recovery), #1615 (shared segment naming), #1616 (descriptor-held audit directory, including retention through a swapped ancestor), #1617 (handle-based answers), #1619 (durable prefix when a grouped write fails during rotation), #1620 (evidencectl doctor and relative audit.path).

The former Relay claim that terminal audit cryptographically bound exact response bytes was inaccurate: the old implementation ignored _exact_response_bytes. The corrected claim is that terminal acceptance gates release of the exact held bytes. This is a documentation/test-name correction, not removal of a prior binding behavior.

Compatibility and migration

This is a breaking audit-format and configuration change. In-product chain verification, key-epoch detection, and SQL audit querying are removed. Operators must ship logs to append-only storage with separate deletion authority when tamper evidence and retained completeness are required.

  • Archive old active and numbered files outside the new retention directory, then start a fresh audit path. Give each runtime process a separate file and collect operator companion files too.
  • Preserve BReg's retired registry_audit and registry_audit_head data before package application removes them. Rebuild, sign, and apply the successor package for the changed catalog fingerprint. History-erasure coverage and encryption-erasure progress remain database state independent of audit retention.
  • Drain the previous Casework and Scheduling outbox publishers before migrations 17 and 8. Both migrations refuse unpublished rows. The upgrade rehearsal checks archived audit-table row counts while retaining row-preservation checks for every other table.
  • Evidence separates bundle hash-key configuration from runtime destination/rotation/retention settings; its bundle digest changes. Relay packages must be resealed for the new artifact schema and audit.path contract. Render moves to audit/render.jsonl.
  • Removed commands: bregctl audit verify/export/prune, evidence verify-audit, and registry-render audit-verify. CLI publication metadata, product changelogs, operational guidance, and architecture models follow the new contract.

Minimization remains at parity. In particular, Casework's published allowlist already excluded raw task/request/result identifiers, decision, and target fields before this change. Known-answer tests preserve the production identifier-hashing derivation.

Platform inventory

  • Added: AuditWriter, AuditEntry, AuditPhase, AuditDestination, FileDestination, destination/unavailability errors, rotation/retention bounds, for_process, and from_line_sink test support.
  • Kept: AuditProfile, AuditKeyHasher, identifier-key derivation, redaction helpers, authorization events, require_audit_under, and assert_json_absent_strings/AuditJsonLeakError.
  • Deleted sinks and envelopes: AuditSink, JsonlFileSink, JsonlStdoutSink, SyslogSink, AuditEnvelope, DurableSegmentedJsonlSink, DurableSegmentedAuditLog, and SegmentedAuditSummary.
  • Deleted chain machinery: ChainState, AuditChainHasher, AuditChainProfile, chain-key derivation, AuditProfile::chain_hasher/bootstrap_or_start_empty, verify_chain, JSONL/segmented-chain verification and visiting helpers, quarantine/recovery, chain-break records, chain verification errors, rotation/syslog helpers, and assert_chain_integrity/ChainAssertionError.
  • Removed platform-audit dependencies: async-trait, registry-platform-canonical-json, subtle, ulid, and development tracing-subscriber; platform-testing also drops async-trait.

Cargo.lock changes only the workspace packages' dependency edges, with no third-party version updates. The writer uses the existing schemars and uuid packages; evidencectl adds the shared audit crate for destination checks. Relay and Render shed obsolete chain-related dependencies.

Issues

This change removes the hash chains, the segmented-JSONL sink, and the BReg audit head those issues are about:

Closes #1447
Closes #1504
Closes #1125
Closes #1133

DCO

  • Every commit includes a Signed-off-by trailer.
  • I reviewed the submitted changes and am responsible for the contribution.

@jeremi

jeremi commented Sep 25, 2026

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-27T12:34:33.369206Z db06d67 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9467174eed

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/registry-platform-audit/src/writer.rs Outdated
@jeremi
jeremi force-pushed the jeremi/load-testing branch from 4820bb3 to 109975d Compare September 25, 2026 08:26
Base automatically changed from jeremi/load-testing to main September 25, 2026 09:06
@jeremi
jeremi force-pushed the jeremi/audit-simplification branch from 9467174 to 159b6ce Compare September 25, 2026 11:42
@jeremi

jeremi commented Sep 25, 2026

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 159b6ceaae

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/registry-platform-audit/src/writer.rs
Comment thread crates/registry-platform-audit/src/writer.rs Outdated
Comment thread crates/registry-platform-audit/src/writer.rs Outdated
@jeremi

jeremi commented Sep 25, 2026

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

} else {
self.last_completed
.take()
.filter(|completed| completed.operation == last)
.ok_or(EvidenceAuditError::InvalidEvent)?

P2 Badge Reject earlier unmatched Evidence audit attempts

When operation A has an access-attempt entry but no terminal entry and a later concurrent operation B completes, last_operation points to B and this branch returns B's completed view while silently leaving A in self.pending. Consequently evidencectl audit show --last-operation succeeds even though the retained stream contains an unmatched earlier access attempt, contrary to its completeness check; after selecting the last operation, reject any remaining pending operations other than the intentionally displayed access-only last operation.

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/registry-platform-audit/src/writer.rs Outdated
Comment thread crates/registry-breg/src/attachment_verification_worker.rs
@jeremi

jeremi commented Sep 26, 2026

Copy link
Copy Markdown
Member Author

Codex review-body finding (Evidence finish() returns the last operation while an earlier access attempt stays unmatched): confirmed, but finish() is identical on main, so this PR did not introduce it. Tracked in #1569, because refusing whenever any earlier attempt is unmatched would break --last-operation after a single crashed request, and choosing between refusing and reporting needs design.

@jeremi

jeremi commented Sep 26, 2026

Copy link
Copy Markdown
Member Author

@codex review

@jeremi
jeremi marked this pull request as ready for review September 26, 2026 04:12

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a9ad47ca18

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/registry-platform-audit/src/writer.rs
@jeremi
jeremi force-pushed the jeremi/audit-simplification branch from a9ad47c to 503b8cf Compare September 26, 2026 04:26

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

if event.phase == AuditPhase::AccessAttempt {
if self.completed.contains(&operation)
|| self
.pending
.insert(operation, PendingLocalOperation { event, view })
.is_some()

P2 Badge Aggregate multiple access entries in local audit inspection

When a local Evidence operation uses a multi-stage acquisition such as search-then-fetch, the runtime appends one AccessAttempt for each physical source call under the same operation correlation. The collector treats the second legitimate access entry as reuse of a pending operation and rejects the entire retained stream, so evidencectl audit show --last-operation becomes unusable after such a request. Retain all access entries for an operation and validate them against its single terminal entry instead of requiring exactly one.

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/registry-casework/src/audit.rs Outdated
Comment thread crates/registry-evidencectl/src/doctor.rs Outdated
Comment thread crates/registry-platform-audit/src/writer.rs
Comment thread crates/registry-platform-audit/src/writer.rs Outdated
@jeremi

jeremi commented Sep 26, 2026

Copy link
Copy Markdown
Member Author

Codex review-body finding (multi-stage Evidence operations break audit show --last-operation): confirmed. The collector and renderer are the same on main, but fixed here rather than ticketed. The collector keeps every access entry of an operation, requires each stage to share the operation context, and pairs the terminal entry with the last stage (whose source it names). The evidencectl renderer prints one ACCESS AUTHORIZED line per source call. Tests: local_inspection_returns_every_source_stage_of_a_multi_stage_operation and multi_stage_view_prints_one_access_line_per_source_call.

@jeremi
jeremi enabled auto-merge September 26, 2026 04:47

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 58537313a9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/registry-platform-audit/src/writer.rs Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e464b9d80f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/registry-platform-audit/src/writer.rs Outdated
Comment thread crates/registry-platform-audit/src/writer.rs Outdated
@jeremi
jeremi force-pushed the jeremi/audit-simplification branch from e464b9d to 4a688ff Compare September 26, 2026 05:16

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f40f30e46e

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/registry-platform-audit/src/writer.rs

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e94b878564

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/registry-platform-audit/src/writer.rs
Comment thread crates/registry-breg/src/runtime_config.rs
Comment thread crates/registry-evidencectl/src/doctor.rs Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 58a897b92c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/registry-platform-audit/src/writer.rs
Comment thread crates/registry-platform-audit/src/writer.rs Outdated
…concerns by its pseudonym

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
… lost before recording it

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
Folds in the follow-ups:
- fix(audit): state the parent-segment refusal in the runtime schemas

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A replay whose commit acknowledgement was lost and whose reset cannot be
read back may have committed, so recording it as refused and
dead-lettered could contradict the database. It is now answered with
`replay_unfinished` under the pending disposition, and `replay_refused`
stays for a reset that read back as rolled back or changed no row.

Folds in the follow-ups:
- docs(hooks): state the audit pairing contract as at least one answer

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…nfinished

A records swap whose commit acknowledgment was lost may have taken
effect, so `records apply` recorded a refusal the database could
contradict. The store now reads the swap's commit back like a capacity
commit, and a swap whose outcome cannot be read is answered `unfinished`
with reason `records.replace-unacknowledged`.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A client disconnect can drop a Relay handler after its attempt entry was
accepted and before its terminal entry, leaving the request unpaired.
The attempt now begins a platform audit request whose guard the handler
holds until its terminal returns; a dropped guard writes a terminal
record with the added unfinished outcome.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
`Path` drops a trailing `/` or `.`, so `/a/b/` passed the preflight as
the file `b` in `/a` and then failed to open at startup. The file
destination, Relay's shape check, and the runtime schema patterns now
require the path to end in a file name.

Folds in the follow-ups:
- chore(identifiers): refresh the catalog digests of the regenerated runtime schemas

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A worker aborted after the attempt entry was accepted and before the
lease committed rolled the lease back, leaving that attempt with no
lease for expiry recovery to answer. The claim now runs to its commit
in a task of its own, as the replay already does.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…s commit

Reading the transaction id for a lost-acknowledgment read-back put a
database round trip between the grant expiry re-check and COMMIT, so a
grant could lapse in that gap. The re-check now runs inside
commit_capacity, after the id is read and immediately before COMMIT.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
The retention page, the platform changelog, and the Render README said
only a crash could leave a request entry unanswered. An exit or runtime
shutdown with a write in flight can too, and a delivery attempt whose
terminal commit fails waits for lease-expiry recovery. State that a
request may be answered more than once, that an unfinished outcome also
covers refused and no-op operator commands, and that an ingestion commit
error is answered unfinished.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…eled

A worker aborted after its terminal disposition committed and before the
terminal entry was accepted left a delivered or dead-lettered row with no
entry, and a terminal row is never reaped, so nothing answered that
attempt. Finalize now runs to its terminal entry in a task of its own, as
the claim and the replay already do.

Folds in the follow-ups:
- docs(hooks): scope the detached-task answer promise to a running process

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…esponse

An appended response could claim a request whose handle was already writing
its own response, so that request ended with two responses and another open
request under the same correlation was answered unfinished.

Folds in the follow-ups:
- docs(audit): describe the detached queue's lock wait without poisoning

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…tes in

Creating the missing audit directory needs write and search permission on
its nearest existing ancestor, so the preflight probes both instead of
write alone.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A name the filesystem accepts for the active file could leave no room for
the sealed segment's sequence suffix, so the first rotation failed and
stopped the writer for good. Startup refuses such a name, and a process
role's sibling name, instead.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…r creates

A missing audit directory was created without syncing its parent, so a
power loss could remove the directory together with entries already
reported durable. The preflight also requires read permission on the
existing ancestor, since the writer opens it to sync.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…e a shipper

- Validate an audit directory level won by a concurrent creator.
- Refuse to create a missing audit directory only below a world-writable,
  non-sticky ancestor, and say to create it owned by the service user, 0700.
- Record the next sealed sequence in .seq at open and before each rotation,
  so numbering continues after a shipper removes every sealed segment.
- Keep retention and process roles out of another stream's active file.
- Apply retention oldest first, stop at the first unexpired segment, and log
  a retention failure instead of refusing open or rotation.
- Carry on when the sealed segment is removed right after rotation.
- Sync the directory on every open, not only when a name was created.
- Drop ctime from the active-file identity check; length, mtime and inode
  still catch truncation, append, in-place rewrite and replacement.
- Add AuditError::operator_description and pass it through Evidence, Relay,
  Casework and Scheduling startup refusals and logs.
- Refuse a multi-stage Evidence operation that retention cut short.
- Record a delivery terminal whose commit cannot be read back with the
  Unknown disposition instead of a guessed retry state.
- Describe the writer contract in products/platform/AGENTS.md.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…at startup

- Answer a committed ingestion transition unfinished when its response entry
  fails.
- Record an attachment verdict only after its commit resolves.
- Keep the request scrub from answering the erase-history lifecycle.
- Answer an unproven mutation commit unfinished, never refused.
- Accept the unknown webhook audit disposition for an interrupted terminal.
- Report an audit destination refusal as StartupError::AuditDestination with
  the writer's recovery text, in the startup log and in bregctl doctor.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…d ones

- Leave a trace of background work whose commit is unresolved.
- Settle a commit whose caller is canceled after COMMIT is sent.
- Pass the audit writer's recovery text through the startup refusal.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
…come

- Describe Render's server-drawn audit correlation in the retention table.
- State how Scheduling's audit-outbox retirement is covered outside the
  rehearsal.
- Document ordered, non-fatal retention, the .seq record, the narrowed
  ancestor rule, and the recovery text operators see at startup.
- Replace the lease-expiry wording with worker_interrupted and unknown.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
A child left with umask 0777 writes its coverage profile at exit with
mode 000, and llvm-profdata then crashes merging it, failing platform
line coverage in the merge queue.

Signed-off-by: Jeremi Joslin <jeremi@joslin.fr>
@jeremi
jeremi removed this pull request from the merge queue due to a manual request Sep 27, 2026
@jeremi
jeremi force-pushed the jeremi/audit-simplification branch from 404b7f4 to db06d67 Compare September 27, 2026 12:25
@jeremi
jeremi merged commit f002fc6 into main Sep 27, 2026
54 checks passed
@jeremi
jeremi deleted the jeremi/audit-simplification branch September 27, 2026 12:27

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: db06d67131

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +84 to +86
let writer = AuditWriter::open(destination)
.await
.map_err(|_| RegistryAuditError::Unavailable)?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve the companion audit recovery reason

When another bregctl maintenance command already holds the companion lock, AuditWriter::open returns a role-specific SinkLocked recovery message telling the operator to wait, but this mapping discards it. The checked history-erasure, rebaseline, reconciliation, and field-encryption callers then report only a generic audit failure—often advising the operator to check the path and permissions—so a valid deployment cannot distinguish lock contention from a broken destination. Preserve operator_description() (or the typed error) through RegistryAuditError so these commands report the actionable recovery.

Useful? React with 👍 / 👎.

release does not read or verify the chain; it appends its own entries to whatever file sits at
the configured path, and at startup and on each rotation it deletes any file named with an
eight-digit suffix beside it that is older than `audit.retainDays`, which includes files
Casework 0.34.0 sealed.

```sh
install -d -m 0700 /var/lib/registry-casework/audit-archive
cd /var/lib/registry-casework/audit
mv casework.ndjson casework.ndjson.[0-9]* /var/lib/registry-casework/audit-archive/

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make the archive command work without sealed segments

When the old Casework audit stream has never rotated, casework.ndjson.[0-9]* matches nothing and an ordinary shell passes that literal pathname to mv. The mv --help multi-source form is SOURCE... DIRECTORY, so the command moves the active file, then reports that the unmatched second source does not exist and exits nonzero; this makes the mandatory upgrade procedure fail in a common small deployment, especially under set -e. Enumerate existing segments conditionally, use find, or enable a null-glob behavior explicitly.

AGENTS.md reference: docs/site/AGENTS.md:L21-L24

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant