Skip to content

fix(entities): match file-path entities case-insensitively - #29

Merged
7vignesh merged 2 commits into
7vignesh:mainfrom
anush-933:fix/entity-file-path-case-insensitive
Sep 21, 2026
Merged

7vignesh merged 2 commits into
7vignesh:mainfrom
anush-933:fix/entity-file-path-case-insensitive

Conversation

@anush-933

Copy link
Copy Markdown
Contributor

extractEntities keeps file paths in their original case, but every other entity kind gets normalized. Since findEntityMatches compares with me.entity IN (...) and SQLite's IN on TEXT is case-sensitive, a memory that mentions src/Auth.ts never matches a query that mentions src/auth.ts.

I went with lowercasing the path at extraction time. Both recordEntities and findEntityMatches call extractEntities, so the two can't drift apart, and the index on memory_entities.entity still gets used. I skipped COLLATE NOCASE because it would force a scan on the retrieval path and would apply to every kind rather than just files. Nothing ever displays the entity string, so there's no display case to preserve.

Closes #26

Testing

npm run ci passes (234 tests). Added three tests to tests/entities.test.ts: lowercase normalization, mixed-case dedupe, and the repro from the issue. All three fail before the change and pass after.

I also ran npm run bench before and after the change and the output is identical, so there's no retrieval drift. bench:check does report drift against the committed snapshot, but it reports the same thing on a clean main, so I left bench/results/repo-memory.json alone rather than fold an unrelated snapshot update into this PR.

One question

This only fixes entities going forward. Rows already in an existing memory.db keep their old casing until that memory is re-recorded. I kept the change inside entities.ts since the issue scoped it that way, but I'm happy to add a migration that lowercases existing kind = 'file' rows if you'd rather have it here.

extractEntities stored file paths in their original case while every
other entity kind was normalized. findEntityMatches compares with
`me.entity IN (...)`, which is case-sensitive on TEXT in SQLite, so a
memory mentioning src/Auth.ts never matched a query mentioning
src/auth.ts even though they are the same file.

Lowercase file paths at extraction time. Both the write path
(recordEntities) and the read path (findEntityMatches) go through
extractEntities, so the two stay consistent by construction, and the
BINARY-collation index on memory_entities.entity is still usable.
Entities are only ever a matching key and are never displayed, so no
display case is lost.

Closes 7vignesh#26
@vercel

vercel Bot commented Sep 20, 2026

Copy link
Copy Markdown

@anush-933 is attempting to deploy a commit to the vignesh's projects Team on Vercel.

A member of the Team first needs to authorize it.

@7vignesh 7vignesh left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified locally and green on Node 20/22 + CodeQL. The reasoning in your description is exactly right - lowercasing at extraction keeps the write and read paths in agreement without forcing a scan, and leaving the bench snapshot alone was the correct call since the drift pre-exists on main.

@7vignesh
7vignesh merged commit 211e196 into 7vignesh:main Sep 21, 2026
4 of 5 checks passed
@7vignesh

Copy link
Copy Markdown
Owner

On the migration question: no need. Entities re-record whenever a memory is updated, and the entity leg is a recall booster rather than a correctness guarantee, so existing rows healing lazily is fine - not worth the migration risk. Appreciate you flagging it rather than assuming. Thanks for the clean fix!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Entity matching misses file paths that differ only in case

2 participants