Optional code-aware embeddings (jina-code v2): ripgrep plain-language 0.500 → 0.667 recall, 0.156 → 0.347 precision - #130
Merged
Merged
Conversation
…ranking for plain-language questions With the model loaded, fusion off -> on: ripgrep holdout 0.500/0.156 -> 0.667/0.347, click 0.958/0.342 -> 0.958/0.439, this repo 0.679/0.404 -> 0.786/0.429. Without the model nothing changes; MiniLM is never fused (it lowered every set: ripgrep 0.375, this repo 0.607). - EmbeddingModelId::JinaCodeV2 (768-d), loaded from plain files (fastembed's hf-hub cache needs symlinks on Windows) - `neuromesh install embed jina-code`: spec with a sub-path file (onnx/model_quantized.onnx), per-spec readiness, 3 attempts per file - NEUROMESH_EMBED_MODEL also sets that model's vector width - convex combination (0.6 lexical / 0.4 dense, Bruch et al. 2023) - docs: CHANGELOG, configuration, measured, reply to upstream, handoff Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e file-localisation harness; research log - selected_files was a sorted set of base names: no hit@k, ambiguous __init__.py - scripts/swebench_localize.py: per-instance checkout of base_commit, packet on the issue text, plain BM25 on the same checkout, hit@1/3/5, resumable JSONL - docs/research/contributions-log.md: contributions, evidence, negative results, the gap list to a top-venue paper - with jina-code loaded the eight old sets are unchanged (web recall 0.917 -> 0.922, precision 0.643 -> 0.639) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ParsaVictor
marked this pull request as ready for review
October 1, 2026 12:36
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fuses the lexical file ranking with Jina code v2 file vectors for plain-language questions when the model is installed (
neuromesh install embed jina-code,NEUROMESH_EMBED_MODEL=jina_code_v2). Fusion off→on, model loaded: ripgrep 0.500/0.156→0.667/0.347, click 0.958/0.342→0.958/0.439, this repo 0.679/0.404→0.786/0.429. MiniLM is never fused. Draft until the old sets are checked with the model loaded.🤖 Generated with Claude Code