Permission-aware RAG, judged by evidence.
Authorization is compiled into the retrieval SQL, and every retrieval change is measured on a reproducible benchmark, negative results included.
English · 简体中文 · Architecture · Evaluation · Benchmark · Milestones
- 🛡️ Authorization inside the query: tenant, clearance, department and project rules are compiled into the SQL of every retrieval path, so unauthorized rows never leave PostgreSQL
- 🔎 Three retrieval strategies on one database: PostgreSQL full-text search, exact pgvector search and reciprocal rank fusion, each response tagged with the hash of its retrieval plan
- 📊 Evaluation with confidence intervals: 122 hand-checked cases, bootstrap intervals, paired comparisons and a BM25 reference row; a difference counts only if its interval excludes zero
- 🚨 A security gate in CI: 31 cases try to reach forbidden documents, and every returned chunk is checked against hand-written visibility; one unauthorized result fails the build
- 🧪 Negative results are published: hybrid does not beat dense on this dataset, and one earlier claim was withdrawn when a larger dataset stopped supporting it
- ⚙️ Real ingestion: asynchronous jobs with retry and resume, versioned documents, label changes that apply to the next query without re-embedding
- 🚀 Runs on a laptop: one
docker compose up, a CPU embedding model baked into the image, no API key and no GPU
Nothing about the request changes except the token. This is real output from ./scripts/demo-queries:
alice-engineer (tenant northstar) · dense-only · policy abac/1
1. hr-volunteer-policy › Volunteer Time Off Policy > European Union
Employees based in the EU receive two paid volunteer days per calendar year.
2. hr-volunteer-policy › Volunteer Time Off Policy > United States
Employees based in the US receive one paid volunteer day per calendar year.
mallory-outsider (tenant external) · dense-only · policy abac/1
1. volunteer-handbook › Community Volunteering Handbook > Volunteer days
Orbit Labs employees receive three volunteer days per year, which can be taken as half days.
The outsider gets no "permission denied", no hit count and no Northstar document title. The tenant condition is part of the SQL that selects candidates, so the Northstar policy is never a row in their result set.
The same holds inside a tenant. Alice has internal clearance and Carol has confidential; both ask how much on-call allowance staff engineers receive:
alice-engineer (tenant northstar) · dense-only · policy abac/1
1. eng-oncall-handbook › On-call Handbook > Compensation
Engineers receive an on-call allowance of 250 EUR per week of primary on-call, ...
2. eng-oncall-handbook › On-call Handbook > Acknowledging pages
The on-call engineer must acknowledge a page within 5 minutes. ...
carol-manager (tenant northstar) · dense-only · policy abac/1
1. hr-compensation-bands › Compensation Bands > On-call pay
Engineers at staff level and above receive an on-call allowance of 400 EUR per week ...
2. eng-oncall-handbook › On-call Handbook > Compensation
Engineers receive an on-call allowance of 250 EUR per week of primary on-call, ...
Alice gets the general handbook, and nothing tells her that a confidential document exists. Carol gets the confidential answer first.
Requirements: Docker, uv. The first build downloads Maven dependencies and a ~70 MB embedding model, which usually takes a few minutes. No API key, no GPU.
git clone https://github.com/poppycoderr/grounded-access.git
cd grounded-access
docker compose up -d --build --wait # PostgreSQL + pgvector, control plane, CPU model service
./scripts/load-demo # ingest the fictional Northstar and Orbit Labs corpora
./scripts/demo-queries # the comparison above
./scripts/benchmark # evaluate every strategy and enforce the security gateSearch as any demo identity, or mint a token and call the API yourself:
uv run --project packages/evaluation ga-eval search alice-engineer "How many paid volunteer days do EU employees receive?"
TOKEN=$(uv run scripts/mint-token.py alice-engineer --scope "query debug")
curl -s localhost:8080/api/v1/retrieval/search -H "Authorization: Bearer $TOKEN" \
-H 'content-type: application/json' \
-d '{"query":"paid volunteer days","strategy":"dense-only","k":3}'Most RAG demos retrieve first and filter afterwards. That leaks rows into application memory — and from there into rerankers, prompts and logs — while silently costing recall, because the filter eats the top-k that the index already chose.
The predicate carries the tenant, clearance, department and project rules. The shape of the solution is the point: whatever the rules are, they belong in the query that selects candidates.
A principal is compiled into one predicate with bound parameters, never string concatenation, and every chunk query embeds the same object (policy abac/1).
A missing or unknown clearance counts as the lowest level, and a principal without a department or projects sees only documents that do not restrict that attribute. A change of labels applies to the next query and needs no re-embedding.
Tenant, clearance, department and project are authorization and count toward the security gate. Region, validity dates and document status are scope: they shape relevance rather than access. Keeping them apart means the security metrics only ever count real access violations. The full decision table is in docs/architecture/authorization.md.
Guards in place today:
- an architecture test — only
AuthorizedChunkQuerymay read the chunk table; - a property-based test — random principals and labels, with every query path compared against a separate reference evaluator;
- integration tests on real pgvector — each rule of the decision table on all three strategies, cross-tenant isolation, hostile claim values, invalid and under-scoped tokens;
- the evaluation security gate — every returned chunk is checked for document and version against hand-labelled visibility, never against the compiler itself, and each principal's full chunk listing must match its visible set.
The threat model lists every control with the check that verifies it, and the risks that are accepted: above all, that the demo identity setup lets anyone mint any token.
benchmarks/reports/m2-authorization/ holds the committed run: run.json (dataset version, commit, retrieval plans, policy and chunker versions, bootstrap seed, platform and CPU), cases.jsonl (per-case rankings) and the rendered report.md. Regenerate it with ./scripts/benchmark --out benchmarks/reports/<name>.
Dataset v3, test split, 65 answerable cases, 95% bootstrap intervals, produced by the CI runner (Linux x86_64):
| Strategy | Recall@10 | MRR@10 | nDCG@10 | Security violations | Scope failures |
|---|---|---|---|---|---|
sparse-only (PostgreSQL FTS) |
0.923 [0.85, 0.98] | 0.683 [0.59, 0.77] | 0.742 [0.66, 0.82] | 0 | 0 |
dense-only (pgvector, exact) |
0.969 [0.92, 1.00] | 0.873 [0.80, 0.94] | 0.898 [0.84, 0.95] | 0 | 0 |
hybrid-rrf (reciprocal rank fusion of the two) |
0.969 [0.92, 1.00] | 0.822 [0.74, 0.89] | 0.857 [0.79, 0.91] | 0 | 0 |
bm25-reference (offline, same authorized chunks) |
0.931 [0.87, 0.98] | 0.720 [0.63, 0.80] | 0.770 [0.69, 0.84] | 0 | 0 |
Security. Zero violations across 122 cases, 31 of which try to reach a document the principal may not see: in another tenant, above its clearance, in a project it is not on, or in another department. Every returned chunk is checked for document and version, and before any query runs each principal's full chunk listing is compared with its hand-labelled visible set.
Scope. 13 cases ask about a region or a date, or name documents that are readable but do not apply: the other region's holiday calendar, last year's travel policy, a benefit that has not started. No strategy returned one. Scope failures are counted apart from security violations and never added to them.
What the paired comparisons support, and what they do not:
- Dense ranks the right evidence higher than FTS: MRR@10 +0.19 [+0.11, +0.28]. Whether the evidence appears in the top 10 at all shows no detectable difference (Recall@10 +0.05 [−0.02, +0.11]).
- Hybrid does not beat dense. MRR@10 −0.05 [−0.11, +0.01] against dense: no detectable difference, with the point estimate in favour of dense. The analysis of the first hybrid run explains why: equal-weight fusion gives the weaker FTS channel the same vote.
- FTS against BM25: a finding that did not hold. On dataset v1, BM25 was measurably ahead of FTS (MRR@10 +0.10 [+0.02, +0.18]). On v2 and v3 the difference is no longer detectable (v3: +0.04 [−0.02, +0.10]). The earlier reports stay in the repository; the claim is withdrawn until a larger dataset supports it.
- Dense beats BM25 on MRR@10 (+0.15 [+0.07, +0.24]).
Nothing was tuned on the test split. The dataset is 32 fictional documents and 122 hand-checked cases, with 36 of the 87 answerable ones deliberately worded so they share almost no words with their evidence. It is a demo benchmark: it shows the method and the direction of the differences, not production quality. Reports on earlier dataset versions are not comparable with this one. Published numbers are reproducible to the reported precision; dense result lists can differ in the order of near-tied candidates between CPUs (see benchmarks/README.md). See the dataset card for what it covers and what it does not.
Evidence is labelled as a document version plus a quote, not a chunk id, so chunking strategies can be compared on the same labels. Read the method in docs/evaluation/strategy.md.
| Capability | Today | Planned |
|---|---|---|
| Authorization in the retrieval query | Tenant, clearance, department and project rules compiled once per request into the SQL of every channel; region and validity scope compiled separately; both property-tested against a reference evaluator | A uniform "no answer" on the answering path (M3) |
| Retrieval | sparse-only (PostgreSQL FTS), dense-only (exact pgvector) and hybrid-rrf (reciprocal rank fusion with overlap deduplication); every response carries a plan hash |
Cross-encoder reranking (M3) |
| Ingestion | Asynchronous jobs (202 + poll) with a SKIP LOCKED worker, bounded retry and resume; content-hash versioning; Markdown and plain-text chunking with sentence-level splitting and overlap; disable and delete apply to the next query, with background cleanup |
– |
| Evaluation | 122 cases over 32 labelled documents: paraphrases, hard negatives, 31 authorization negatives and 13 scope cases; a BM25 reference row, bootstrap intervals and paired comparisons; a security gate in CI that checks every returned chunk and every principal's full listing | Answer metrics: citation validity, abstention (M3) |
| Answers | Ranked evidence from /api/v1/retrieval/search |
/api/v1/query with citations and abstention (M3) |
| Operations | Docker Compose, CI on every PR; a trace id per request, audit events and execution records written synchronously and failing closed | OpenTelemetry traces, dashboards (M4) |
Design goals that are not yet verified end to end, and the milestone that will verify them: a uniform "no answer" whether content is hidden or missing, and answers that cite only what the model was shown (both M3).
| Component | Stack | Role |
|---|---|---|
apps/control-plane |
Java 21 (CI on 21 and 25), Spring Boot 4.1 | Token verification, policy compilation, ingestion, retrieval, API |
apps/model-service |
Python 3.12, FastAPI, ONNX Runtime | Embeddings today, reranking in M3; no identities, no database access |
packages/contracts |
OpenAPI | The contract between them, with a drift test on both sides |
packages/evaluation |
Python 3.12 | ga-eval: dataset validation, retrieval metrics, security gate, reports |
| Storage | PostgreSQL 17 + pgvector | Documents, versions, chunks, full-text and vector search |
io.groundedaccess
├── identity # verify the JWT → Principal (attributes only from the token)
├── authorization # PolicyCompiler → one parameterized SQL predicate
├── corpus # normalize · chunking · versions and access labels · cleanup
├── ingestion # job queue: SKIP LOCKED worker, leases, retry and resume
├── retrieval # AuthorizedChunkQuery: the only reader of the chunk table · RRF · plan hash
├── modelclient # model-service client: batching, timeouts, pinned model
└── api # REST controllers, scopes, problem responses
| Milestone | Scope | Status |
|---|---|---|
| M0 Walking skeleton | Demo identities, Markdown ingestion, sparse and dense retrieval, tenant isolation, eval CLI, CI security gate | ✅ Done |
| M1 Retrieval baseline | Dataset with hard negatives, BM25 reference, confidence intervals; async ingestion, disable and delete, chunker v1; RRF hybrid with a published verdict | ✅ Done · v0.1.0-alpha.1 |
| M2 Authorization | Full decision table, property-based tests, labelled dataset v2 and the stricter gate, scope filters with asOf, audit events, existence-safe document reads, threat model, dataset v3 with scope cases, published report |
✅ Done |
| M3 Reranking and answers | Cross-encoder reranking with fallback, context builder, structured citations, abstention | Planned |
| M4 Operations and release | Traces and dashboards, failure and load tests, v0.1 benchmark report | Planned |
Not in the first phase: knowledge graphs or GraphRAG, autonomous agents, extra vector databases, OCR and multimodal input, fine-tuning, Kubernetes and multi-cloud, a no-code builder. See docs/project/milestones.md.
- Architecture overview — trust boundaries, data model, ingestion and query flows, failure behaviour
- Authorization model — invariants, decision table, compiled SQL, what is verified today
- Threat model — assets, actors, abuse cases with their checks, residual risks
- Evaluation strategy — case schema, metrics, CI gates, reproducibility rules
- Architecture decisions — modular monolith, PostgreSQL FTS + pgvector, retrieval-time authorization, the Python model service, what
asOfmeans - Milestones and open questions
Issues and pull requests are welcome, and the most useful ones right now are evaluation cases that break the retrieval baseline, and arguments against the design decisions in the ADRs. Read CONTRIBUTING.md and AGENTS.md first — the rules in AGENTS.md apply to humans and AI agents alike.
Security reports go through GitHub security advisories. The demo identity setup is insecure by design; see SECURITY.md.
- domain-driven-kit — executable DDD for Spring Boot. Grounded Access reuses its engineering conventions (architecture tests, null safety, CI layout) without depending on it.
Apache-2.0. The fictional demo corpus in data/corpus is published under CC BY 4.0 and describes companies that do not exist.