Test a model's action policy. Replay the evidence.
A ticket router needs rules for when to act, abstain, escalate or deny. Actseal measures accepted-action errors and coverage under a frozen policy.
Recorded Jev audit: INCONCLUSIVE. 580 of 639 verification cases received ACT; 24 disagreed with benchmark labels. Fixed benchmark; unreleased producer. Audit and offline replay
Boundary: the application owns execution. Replay checks consistency; it does not authenticate responses or prove label truth.
Engineering: 20 ADRs · 11 JSON Schemas Mutation harness · 1.0.1 release receipt
Synthetic demo. macOS/Linux + uv; use a new ./actseal-demo directory.
uvx --python 3.12 actseal demo --out ./actseal-demo
uvx --offline --python 3.12 actseal replay ./actseal-demo/fixed/evidence
uvx --offline --python 3.12 actseal replay ./actseal-demo/bad/evidenceNo model or API key is needed; install uv first. The first command downloads the package if needed. Expected result from synthetic fixtures, showing selected output:
[bad] expected BLOCK, observed BLOCK, replay BLOCK (match)
[fixed] expected PASS, observed PASS, replay PASS (match)
result: success
Demo exit: 0. Fixed replay: PASS / 0. Bad replay: BLOCK / 1.
The third command intentionally exits 1. These fixtures demonstrate the workflow; they do not measure a live model.
Use it in an application · Read the limits · Quickstart
Before enabling automatic ticket routing
Run the committed action-gate example to see the application boundary. It verifies a frozen policy, replays the recorded evidence, and routes authored tickets. Only ACT permits the example application’s local queue write; ABSTAIN, ESCALATE and DENY take non-execution paths.
From a development checkout, run
uv run --frozen python examples/action_gate/run.py --check.This is a synthetic integration example. The application owns execution; the result is not evidence of deployment performance.
The action-gate example
shows where Actseal sits in an application's control flow and draws that flow
from the ticket to the four decisions. A small ticket-routing application calls
the evaluator for every incoming ticket and executes its one action, a local
queue write, only when the decision is ACT. Actseal verifies the frozen policy
offline and replays the evidence; it does not execute, intercept or enforce the
application's action. The example uses authored synthetic data
(evidence_scope=demo) and runs without a key, model download or network.
Frozen policy. Measured risk and coverage. Offline replay.
A model can choose the right label often and still act on the wrong cases. Actseal checks a frozen action policy against labelled cases, bounds errors among accepted actions and coverage across all scheduled cases, and saves evidence for offline replay.
Five stages: Freeze turns the frozen policy, labelled inputs and model identity into one lock. Run collects provider answers and six synthetic faults into decisions: ACT, ABSTAIN, ESCALATE or DENY. Verify applies the risk and coverage bounds and fault rules to one verdict with its exit code. Seal writes one bounded evidence bundle. Replay recomputes the verdict offline with no model call. The stages are drawn in the how-it-works figure (dark version).
An ACT is a per-case decision permitted by the frozen allowlist and the selected-label probability threshold; it is not a statement that the answer is correct. Coverage is ACT decisions over all scheduled cases and risk is wrong ACT decisions over ACT decisions. A whole-run verdict bounds those rates:
| Exit | Verdict | Meaning |
|---|---|---|
| 0 | PASS | Valid evidence satisfies the frozen risk and coverage limits and the fault checks. |
| 1 | BLOCK | Valid evidence establishes a limit violation or a fault-check failure. |
| 2 | INCONCLUSIVE | Valid evidence proves neither PASS nor BLOCK. |
| 3 | ERROR | Invalid or incomplete evidence, a usage or setup failure, or an invalidated native-worker experiment. |
lock exits 0 on success. demo exits 0 only for its expected BLOCK/PASS
pair with matching replays; it never relabels the bad run. Every command
accepts --json for one versioned receipt.
Guarantees
- Complete scheduled-case and required fault evidence is checked.
- PASS requires the frozen risk/coverage bounds and fault rules to pass.
- Supported evidence is recomputed offline without calling a model.
Limits
- Hashes and replay cannot authenticate coherently rewritten responses.
- They cannot prove inference occurred or that labels are true.
- Population claims require the stated sampling assumptions; Actseal does not enforce application execution.
ACT is a per-case policy decision, not a claim that its label is correct. The application owns execution. Actseal is not an OS sandbox or permission firewall. An externally trusted lock digest anchors policy identity; it does not authenticate responses.
Actseal supports one categorical question with 2–16 labels and a frozen allowlist and threshold. The runtime core uses only the Python standard library on Python 3.12 and 3.13, macOS and Linux. Windows is unsupported. Native Laya support is limited to the documented tested CPU configurations.
The packaged demo and action-gate example are synthetic (evidence_scope=demo).
Population interpretation requires independent cases and one prespecified attempt
under a fixed policy. Do not retry until PASS.
The Jev adapter is PROVISIONAL and requires explicit opt-in:
--provider jev --experimental-provider. It has no 1.x compatibility promise.
The released 1.0.0 adapter's accepted evidence uses mocked transports; the separate audit below used an unreleased benchmark producer.
A preregistered audit of Jev on a fixed 16-intent Banking77 subset returned INCONCLUSIVE: 580/639 verification cases received ACT, with 24 accepted errors. The complete evidence and offline verification instructions are published with the unreleased benchmark producer snapshot identified explicitly.
Its offline replay runs with that archived snapshot and exits 2 (INCONCLUSIVE); the published 1.0.0 package returns ERROR integrity.lock for this bundle because that producer is not in the compatibility registry.
Replay never imports a provider, and the packaged demonstration establishes no population or model-quality result. The statistical contract states the bounds and sampling assumptions, providers the provider support, and the threat model the authenticity boundary.
Seven groups: the CLI and typed API; contracts and locks; providers;
normalization and policy; assessment, statistics and faults; evidence; and
replay. Replay reads the evidence bundle and never reaches a provider. The
architecture figure
(dark version)
names the actseal modules in each group.
This is an unedited capture of the three commands above against the published 1.0.0 wheel, rendered from the raw cast with the pinned authoring toolchain at speed 1; it shows the expected exits 0, 0 and 1. It was recorded after the v1.0.0 publication and is absent from the immutable v1.0.0 tag and that version's PyPI description. The recording illustrates the demo; it is not authenticated model evidence.
| Read | What it covers |
|---|---|
| Concepts | Frozen policy, selected-option probability, per-case decisions versus whole-run verdicts, evidence scope, one attempt |
| CLI reference and Python guide | Exact commands, exit codes, JSON receipts, public functions and runnable examples |
| Stability manifest, versioning and migration | The 1.x compatibility promise, what may change when, and the 0.1.0 evidence path |
| Statistical contract and threat model | Bounds, verdict rules, sampling assumptions and the authenticity boundary |
| Providers and FAQ | Fixture and optional native Laya setup; the PROVISIONAL experimental Jev cloud adapter behind --provider jev --experimental-provider (bring your own key, mocked-transport evidence for released 1.0.0; separate benchmark audit above); answers to "why not PASS" |
| Publishing, CHANGELOG, release notes and SECURITY | Release pipeline and receipts, changes per version, the receipt-backed release notes, private vulnerability reporting and support |
Actseal is part of a set of tools for independent testing and evidence for agent controls. ARCI gates repeated-trial agent regressions, injects faults and reduces failing fault sets; Actseal checks a frozen categorical policy’s accepted-action risk and coverage, with application-owned actions and offline replay. Frontier Scout compiles policies and verifies PR scope, Dorian checks claim warrants, and Evalopt Graph evaluates acceptance policies against supplied evidence.
The mutation harness applies eight prescribed changes to temporary copies of the gate: a wider per-tail error allocation, the accepted count as the coverage denominator, a zero-width risk interval when nothing is accepted, a threshold tie that abstains, PASS read from the wrong risk bound or the wrong coverage bound, faults that never block, and integrity errors outranked by a fault. A mutant counts as killed only when every designated test fails with an assertion error; import, setup or timeout failures invalidate the mutant instead.
Maintained by Ajay Surya Senthilrajan, with AI pair-programming recorded in commit trailers. See the tests, design records and release evidence linked here.
Apache-2.0; see LICENSE, NOTICE and dependency notices. Contributions: CONTRIBUTING.