Skip to content

docs: publish the first dataset v2 report and withdraw the fts-versus-bm25 claim - #22

Merged
poppycoderr merged 1 commit into
mainfrom
docs/v2-baseline-report
Oct 1, 2026
Merged

poppycoderr merged 1 commit into
mainfrom
docs/v2-baseline-report

Conversation

@poppycoderr

Copy link
Copy Markdown
Owner

Background

Dataset v2 added access labels, restricted documents and authorization cases. Reports on v1 are not comparable with it, so the README needs numbers from a committed v2 run. This PR commits that run and updates the claims that depend on it.

What is published

 benchmarks/reports/
 ├── m1a-baseline/            # dataset v1
 ├── m1b-hybrid/              # dataset v1
+└── m2-labelled-dataset/     # dataset v2: the evaluation-results artifact of the main CI run at 1044de4, unchanged

test split, 57 answerable cases, Linux x86_64:

Strategy Recall@10 MRR@10 nDCG@10 Violations
sparse-only 0.930 [0.86, 0.98] 0.640 [0.54, 0.74] 0.711 [0.63, 0.79] 0
dense-only 0.965 [0.91, 1.00] 0.856 [0.78, 0.93] 0.884 [0.81, 0.94] 0
hybrid-rrf 0.965 [0.91, 1.00] 0.797 [0.71, 0.88] 0.837 [0.77, 0.90] 0
bm25-reference 0.921 [0.84, 0.98] 0.689 [0.59, 0.78] 0.744 [0.66, 0.82] 0

What the README now says

  • Security. Zero violations across 108 cases, 30 of which are authorization negatives. The gate checks every returned chunk and every principal's full listing.
  • Unchanged findings. Dense beats full-text search on MRR@10 (+0.22 [+0.12, +0.31]) and beats BM25 (+0.17 [+0.07, +0.26]). Hybrid shows no detectable difference from dense (−0.06 [−0.13, +0.01]).
  • One finding is withdrawn. On v1, BM25 was measurably ahead of PostgreSQL full-text search (MRR@10 +0.10 [+0.02, +0.18]). On v2 the difference is +0.05 [−0.02, +0.12], so there is no detectable difference. The README, the benchmarks README and the hybrid analysis all say so, and the v1 reports stay in the repository.

背景

数据集 v2 加入了访问标签、受限文档和授权用例,v1 上的报告不能与它直接比较,所以 README 需要引用一份已提交的 v2 运行结果。本 PR 提交这份结果,并更新依赖它的结论。

发布了什么

目录变化和指标见英文部分。m2-labelled-dataset/ 是 main 在 1044de4 上那次 CI 运行的 evaluation-results 产物,原样提交。

README 现在的说法

  • 安全。 108 条用例的越权结果为 0,其中 30 条是授权负例。门禁会检查每个返回的 chunk,以及每个身份能列出的全部内容。
  • 没有变化的结论。 dense 在 MRR@10 上优于全文检索(+0.22 [+0.12, +0.31]),也优于 BM25(+0.17 [+0.07, +0.26])。hybrid 与 dense 之间没有可检测的差异(−0.06 [−0.13, +0.01])。
  • 撤回一个结论。 在 v1 上,BM25 明显领先 PostgreSQL 全文检索(MRR@10 +0.10 [+0.02, +0.18])。在 v2 上差值是 +0.05 [−0.02, +0.12],没有可检测的差异。README、benchmarks README 和 hybrid 分析文档都已如实说明,v1 的报告仍保留在仓库里。

@poppycoderr
poppycoderr merged commit 4f792fc into main Oct 1, 2026
6 checks passed
@poppycoderr
poppycoderr deleted the docs/v2-baseline-report branch October 1, 2026 16:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant