Skip to content

About

LangGraph on-chain compliance agent. Multi-hop fund-flow tracing, sanctions screening, mixer-proximity heuristics, citation-grounded risk determinations with conformal abstention to human review, hash-chained audit log. Eval reports false-negative rate. Fully synthetic data.

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

defi-compliance-agent

On-chain compliance screening built around the error that matters: the illicit flow you MISS. A LangGraph agent that traces fund flows across hops, screens against a synthetic sanctions list, cites the rule that fired, and abstains to a human when it isn't sure — because freezing the wrong wallet and missing the right one are not symmetric mistakes.

CI License: MIT

Scope

This is a defensive compliance tool — the screening work a crypto exchange's compliance team or a vendor like Chainalysis does. It detects and flags suspicious patterns. It does not help anyone evade detection. All data is synthetic: the transaction graph, the address labels, and the sanctions list are generated with a fixed seed and committed to the repo. No real on-chain data is used.

The error that matters

On-chain data is transparent — you can trace every flow. The hard part is volume and pseudonymity. The two screening errors are not symmetric:

  • A false positive (freezing a clean wallet) costs an analyst a review and annoys a customer.
  • A false negative (clearing a wallet with sanctioned exposure) is the OFAC enforcement action, the fine, the consent order.

The eval reports false-negative rate as the headline. The abstention threshold is calibrated at alpha=0.05 — when the agent isn't confident, it routes to a human.

How it works

  1. Ingest — parse the address/transaction, attach the known on-graph label (deterministic).
  2. Trace — walk the transaction graph N hops via networkx, collecting exposure to labeled entities (mixers, exchanges, sanctioned addresses).
  3. Screen — run heuristics: sanctioned-address proximity, mixer interaction, peel-chain pattern, rapid pass-through. Each firing heuristic is a candidate rule citation.
  4. Determination — Sonnet 4.6 produces a risk tier (low / medium / high) and recommendation (clear / review / freeze), citing every rule that fired via the Citations API. An uncited review/freeze is rejected in code.
  5. Abstain — below the calibrated confidence threshold, route to human review.
  6. Approval gate — high-risk freeze recommendations go through a Slack approval gate. A human approves or overrides. Never auto-freeze.
  7. Audit — every step is written to a hash-chained, examiner-grade log.
flowchart LR
    A[Address / tx] --> IN[ingest<br/>labels]
    IN --> TR[trace<br/>networkx N-hop<br/>exposure]
    TR --> SC[screen<br/>sanctions · mixer · peel · passthrough]
    SC --> DET[determination<br/>Sonnet 4.6 + Citations<br/>risk tier + confidence]
    DET --> AB{confidence ≥ τ?}
    AB -->|no| HR[human review]
    AB -->|yes| RISK{high + freeze?}
    RISK -->|yes| AG[Slack approval gate]
    RISK -->|no| CLOSE[close with citation]
    AG --> CLOSE
    CLOSE --> AUD[(hash-chained audit log)]
    HR --> AUD
Loading

Quickstart

git clone https://github.com/SebAustin/defi-compliance-agent && cd defi-compliance-agent
uv sync --all-extras && cp .env.example .env
make gen-data         # generate synthetic tx graph + sanctions list + 80 cases (fixed seed)
make verify-data      # confirm the generation is deterministic
make test             # full mocked test suite, 85% coverage gate — ZERO API spend
make index            # build the rulebook index (in-memory / on-disk)
make calibrate        # fit the abstention threshold (needs API keys)
make screen ADDR="0xSANC...→0xCASE... (direct sanctioned exposure)"
make eval             # full eval on 80 cases (needs API keys; ~$1.60)

The test suite and the mocked eval path run with no keys and spend nothing. The live make calibrate / make eval paths call Anthropic + Voyage and require ANTHROPIC_API_KEY and VOYAGE_API_KEY.

Eval targets (gated in CI)

CI gates every PR on these thresholds; the headline numbers are produced by running make eval with API keys. Synthetic ground truth is co-generated with the graph, so the metrics are real, not hand-set.

Metric Target
False-negative rate (missed illicit) ≤ 0.03
Risk-tier accuracy ≥ 0.85
Citation coverage = 1.00
Sanctioned-exposure recall ≥ 0.95
Abstention rate report
Conditional accuracy (auto-decided) ≥ 0.92

Citation coverage is a hard invariant: any review/freeze with zero cited rules raises CitationContractError in code, so a passing run is 100% cited by construction.

Sources

  1. Anthropic. Citations API documentation. docs.anthropic.com, 2025.
  2. Yadkori et al. Mitigating LLM Hallucinations via Conformal Abstention. arXiv 2405.01563, 2024.
  3. FATF. Public typologies for virtual-asset risk (referenced conceptually; no data used). 2021.

License

MIT — see LICENSE.

About

LangGraph on-chain compliance agent. Multi-hop fund-flow tracing, sanctions screening, mixer-proximity heuristics, citation-grounded risk determinations with conformal abstention to human review, hash-chained audit log. Eval reports false-negative rate. Fully synthetic data.

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages