fix: preserve mixed-script token boundaries in BM25 search - #415
Open
emecii wants to merge 1 commit into
Open
Conversation
3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Fixes BM25 retrieval of mixed Chinese/Japanese/Korean and Latin text. A passage such as
Python数据库SQLcurrently fails searches forPythonandSQLbecause n-gram expansion joins those words to the edge n-grams. Mixed queries such asPython数据库also fail because Unicode\wconsumes both scripts.Add boundaries around indexed CJK n-grams and keep CJK runs separate when parsing query terms. Existing OR-between-terms and AND-between-CJK-bigrams semantics are preserved. Extend the FTS5 tests with mixed-script retrieval after database reopen and negative controls.
Previously built artifacts remain readable; affected indexes need rebuilding to regenerate missing tokens. No new dependency or public API change.
Related Issues
Follow-up correctness fix to merged #401. No matching open issue or competing open PR was found. Does not address the metadata-sharing/incremental work in open #398.
Validation
AI assistance: implemented and tested with Codex (GPT-6); no claim of human review is made here.
Checklist
uv run pytest); focused 47-test suite passed.CI follow-up
The lint job reports exactly the seven pre-existing
F811duplicateindex_*methods incli.pydocumented above. Formatting passes. None of the errors is in the tokenizer or regression tests, and the same diagnostics were verified on the base. This PR leaves those handlers unchanged.