Issue mode: long reports use the whole-question ranking (SWE-bench Lite dev hit@3 0.250 → 0.583) - #131
Open
ParsaVictor wants to merge 5 commits into
Open
Issue mode: long reports use the whole-question ranking (SWE-bench Lite dev hit@3 0.250 → 0.583)#131ParsaVictor wants to merge 5 commits into
ParsaVictor wants to merge 5 commits into
Conversation
…er the packet by it SWE-bench Lite, dev split (flask, requests, seaborn, xarray, pylint; 24 issues the parameters were chosen on): hit@1 0.125 -> 0.333, hit@3 0.250 -> 0.583, hit@5 0.417 -> 0.667, in-packet 0.583 -> 0.667 (plain BM25: 0.250 / 0.625 / 0.750). The other seven repositories are the test split, measured once afterwards. - a prompt of 60+ words with anchors also seeds the BM25F ranking's best file and runner-up within 70%; nothing is pruned - gold::packet_file_order: packet files best first; for long reports reciprocal-rank fusion of activation order and the file ranking (lexical weight 2, k = 60); CLI selected_paths uses it - packet --json: ranked_paths (BM25F top 10) for evaluation - nine gold sets unchanged Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ard/--only for parallel runs Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With an embedding model on disk and the fast engine, a question that escalated to L3 built the whole workspace's file-tier sidecar synchronously before answering: django (3.5k files) held one SWE-bench issue for more than 10 minutes. The build now starts once on a background thread and the question is answered without it; a later question uses the sidecar. Same issue: >600 s -> 171 s, of which the query is ~19 s and the rest the cold index under 4 parallel jobs. NM_TIMING=1 also times the escalation stages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Issue mode ranked the whole 2 MB hostile prompt of the stage-4 security gate: 57 s against the 30 s bound. Same cap as the seed pipeline. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ress, next steps) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Long prompts (60+ words, e.g. issue reports) also seed the BM25F file ranking and the packet is ordered by reciprocal-rank fusion of activation and that ranking. SWE-bench Lite dev split (24 issues): hit@1 0.125→0.333, hit@3 0.250→0.583, hit@5 0.417→0.667. Test split (7 other repos) measured after merge. Nine gold sets unchanged.
🤖 Generated with Claude Code