On-chain compliance screening built around the error that matters: the illicit flow you MISS. A LangGraph agent that traces fund flows across hops, screens against a synthetic sanctions list, cites the rule that fired, and abstains to a human when it isn't sure — because freezing the wrong wallet and missing the right one are not symmetric mistakes.
This is a defensive compliance tool — the screening work a crypto exchange's compliance team or a vendor like Chainalysis does. It detects and flags suspicious patterns. It does not help anyone evade detection. All data is synthetic: the transaction graph, the address labels, and the sanctions list are generated with a fixed seed and committed to the repo. No real on-chain data is used.
On-chain data is transparent — you can trace every flow. The hard part is volume and pseudonymity. The two screening errors are not symmetric:
- A false positive (freezing a clean wallet) costs an analyst a review and annoys a customer.
- A false negative (clearing a wallet with sanctioned exposure) is the OFAC enforcement action, the fine, the consent order.
The eval reports false-negative rate as the headline. The abstention threshold is
calibrated at alpha=0.05 — when the agent isn't confident, it routes to a human.
- Ingest — parse the address/transaction, attach the known on-graph label (deterministic).
- Trace — walk the transaction graph N hops via
networkx, collecting exposure to labeled entities (mixers, exchanges, sanctioned addresses). - Screen — run heuristics: sanctioned-address proximity, mixer interaction, peel-chain pattern, rapid pass-through. Each firing heuristic is a candidate rule citation.
- Determination — Sonnet 4.6 produces a risk tier (low / medium / high) and recommendation (clear / review / freeze), citing every rule that fired via the Citations API. An uncited review/freeze is rejected in code.
- Abstain — below the calibrated confidence threshold, route to human review.
- Approval gate — high-risk freeze recommendations go through a Slack approval gate. A human approves or overrides. Never auto-freeze.
- Audit — every step is written to a hash-chained, examiner-grade log.
flowchart LR
A[Address / tx] --> IN[ingest<br/>labels]
IN --> TR[trace<br/>networkx N-hop<br/>exposure]
TR --> SC[screen<br/>sanctions · mixer · peel · passthrough]
SC --> DET[determination<br/>Sonnet 4.6 + Citations<br/>risk tier + confidence]
DET --> AB{confidence ≥ τ?}
AB -->|no| HR[human review]
AB -->|yes| RISK{high + freeze?}
RISK -->|yes| AG[Slack approval gate]
RISK -->|no| CLOSE[close with citation]
AG --> CLOSE
CLOSE --> AUD[(hash-chained audit log)]
HR --> AUD
git clone https://github.com/SebAustin/defi-compliance-agent && cd defi-compliance-agent
uv sync --all-extras && cp .env.example .env
make gen-data # generate synthetic tx graph + sanctions list + 80 cases (fixed seed)
make verify-data # confirm the generation is deterministic
make test # full mocked test suite, 85% coverage gate — ZERO API spend
make index # build the rulebook index (in-memory / on-disk)
make calibrate # fit the abstention threshold (needs API keys)
make screen ADDR="0xSANC...→0xCASE... (direct sanctioned exposure)"
make eval # full eval on 80 cases (needs API keys; ~$1.60)The test suite and the mocked eval path run with no keys and spend nothing. The live
make calibrate/make evalpaths call Anthropic + Voyage and requireANTHROPIC_API_KEYandVOYAGE_API_KEY.
CI gates every PR on these thresholds; the headline numbers are produced by running
make eval with API keys. Synthetic ground truth is co-generated with the graph, so
the metrics are real, not hand-set.
| Metric | Target |
|---|---|
| False-negative rate (missed illicit) | ≤ 0.03 |
| Risk-tier accuracy | ≥ 0.85 |
| Citation coverage | = 1.00 |
| Sanctioned-exposure recall | ≥ 0.95 |
| Abstention rate | report |
| Conditional accuracy (auto-decided) | ≥ 0.92 |
Citation coverage is a hard invariant: any review/freeze with zero cited rules raises
CitationContractError in code, so a passing run is 100% cited by construction.
- Anthropic. Citations API documentation. docs.anthropic.com, 2025.
- Yadkori et al. Mitigating LLM Hallucinations via Conformal Abstention. arXiv 2405.01563, 2024.
- FATF. Public typologies for virtual-asset risk (referenced conceptually; no data used). 2021.
MIT — see LICENSE.