Skip to content

Issue mode: long reports use the whole-question ranking (SWE-bench Lite dev hit@3 0.250 → 0.583) - #131

Open
ParsaVictor wants to merge 5 commits into
mainfrom
phase-n/issue-mode
Open

ParsaVictor wants to merge 5 commits into
mainfrom
phase-n/issue-mode

Conversation

@ParsaVictor

Copy link
Copy Markdown
Owner

Long prompts (60+ words, e.g. issue reports) also seed the BM25F file ranking and the packet is ordered by reciprocal-rank fusion of activation and that ranking. SWE-bench Lite dev split (24 issues): hit@1 0.125→0.333, hit@3 0.250→0.583, hit@5 0.417→0.667. Test split (7 other repos) measured after merge. Nine gold sets unchanged.

🤖 Generated with Claude Code

ParsaVictor and others added 5 commits October 1, 2026 16:53
…er the packet by it

SWE-bench Lite, dev split (flask, requests, seaborn, xarray, pylint; 24
issues the parameters were chosen on): hit@1 0.125 -> 0.333, hit@3
0.250 -> 0.583, hit@5 0.417 -> 0.667, in-packet 0.583 -> 0.667
(plain BM25: 0.250 / 0.625 / 0.750). The other seven repositories are
the test split, measured once afterwards.

- a prompt of 60+ words with anchors also seeds the BM25F ranking's best
  file and runner-up within 70%; nothing is pruned
- gold::packet_file_order: packet files best first; for long reports
  reciprocal-rank fusion of activation order and the file ranking
  (lexical weight 2, k = 60); CLI selected_paths uses it
- packet --json: ranked_paths (BM25F top 10) for evaluation
- nine gold sets unchanged

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ard/--only for parallel runs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With an embedding model on disk and the fast engine, a question that
escalated to L3 built the whole workspace's file-tier sidecar
synchronously before answering: django (3.5k files) held one SWE-bench
issue for more than 10 minutes. The build now starts once on a
background thread and the question is answered without it; a later
question uses the sidecar. Same issue: >600 s -> 171 s, of which the
query is ~19 s and the rest the cold index under 4 parallel jobs.

NM_TIMING=1 also times the escalation stages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Issue mode ranked the whole 2 MB hostile prompt of the stage-4 security
gate: 57 s against the 30 s bound. Same cap as the seed pipeline.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ress, next steps)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant