feat: scale bounded query rewriting - #8
Draft
hyeonsangjeon wants to merge 1 commit into
Draft
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
rewrite_many()andmap_sentences()APIs, stronger Microsoft Foundry terminal-status handling, batch-aware evaluation, and a reproducible 10k-term benchmark.Why
Candidate retrieval scaled linearly through a Python dynamic-programming implementation, concurrent mapping updates could expose partial indexes, short canonical terms could hide valid typo candidates, and incomplete Foundry responses could be treated as valid decisions. Similar projects support keeping the existing bounded candidate-selection architecture while accelerating deterministic retrieval and adding controlled batch concurrency.
Impact
Large vocabularies get substantially faster candidate lookup, online batches preserve input order and per-sentence traces, and V1 callers keep their public mutation and serialization patterns. Provider outputs and full sentences are not cached; normalized lexical tokens remain in the bounded LRU unless
candidate_cache_size=0is used.Validation