Repository navigation
fix(evaluation): record the cpu model in every run and qualify dense reproducibility - #25
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Background
The reproducibility note said that dense rankings can differ between arm64 and x86_64, and it treated the hosted x86_64 runner as one reference machine. A CI run on dataset v2 then differed from the published
m2-labelled-datasetreport in 2 of 432 ranked lists. In each of the two, a pair of adjacent dense candidates with near-equal scores swapped places between ranks 7 and 10. No metric changed, and every sparse, hybrid and BM25 list was identical.Hosted runners do not all have the same CPU, and
run.jsondid not say which one a run used.Change
run.json "platform": "Linux-6.17.0-…-x86_64-with-glibc2.39", + "cpu": "<model name from /proc/cpuinfo>",ga-eval runrecords the CPU model. It reads/proc/cpuinfowhere that exists and falls back to what Python reports.benchmarks/README.mdand the evaluation strategy now state the claim that the evidence supports: published numbers are reproducible to the reported precision, and dense result lists are reproducible up to the order of near-ties.The earlier wording was too strong, and this PR corrects it. No published number changes.
背景
可复现性说明里原来写的是:dense 排名在 arm64 和 x86_64 之间可能不同,并把托管的 x86_64 runner 当作同一台参考机器。后来有一次数据集 v2 上的 CI 运行,和已发布的
m2-labelled-dataset报告相比,432 个排序列表里有 2 个不同。这两个列表里,各有一对得分几乎相同的相邻 dense 候选在第 7 到第 10 名之间互换了位置。没有任何指标变化,sparse、hybrid 和 BM25 的列表也全部一致。托管 runner 的 CPU 并不都相同,而
run.json此前没有记录一次运行用的是哪种 CPU。改动
run.json的变化见英文部分的 diff。ga-eval run现在会记录 CPU 型号:有/proc/cpuinfo时从那里读取,否则用 Python 报告的值。benchmarks/README.md和评测策略文档改成了证据能支持的说法:已发布的数字在报告的精度内可复现;dense 的结果列表可复现,但得分接近的候选之间的顺序除外。之前的表述过强,本 PR 做了更正。已发布的数字没有任何变化。