Releases: tobi/qmd
Release list
v2.8.3
[2.8.3] - 2026-08-16
Security
-
qmd updateno longer runs a project-local.qmd/index.yml'supdate:
commands without approval (#886). That file arrives with agit cloneand is
adopted automatically for any command run inside the tree, so cloning a
repository and runningqmd updateexecuted shell commands chosen by whoever
wrote it. On a terminal QMD now lists the commands and asks; with nobody to
ask it skips them and keeps indexing. Approvals are recorded per config file
and per command set in<config dir>/trusted.json, so editing a command — or
agit pullthat rewrites one — asks again. Newqmd trust,
qmd trust listandqmd trust revokemanage approvals, and
QMD_TRUST_UPDATE_HOOKS=1opts unattended runs back in. Commands in your own
~/.config/qmd/*.yml, including anythingqmd collection update-cmdwrites,
are unaffected. -
The same project-local trust gate now covers collection
pathvalues that
resolve outside the project and non-defaultmodels.embed/models.rerank
/models.generateURIs (#889). In-project paths still index unattended;
out-of-project directories are skipped untilqmd trust, and custom model
URIs are not loaded or downloaded.QMD_TRUST_LOCAL_CONFIG=1opts unattended
runs back in (andQMD_TRUST_UPDATE_HOOKS=1still does).qmd collection addrecords trust as it writes, the same wayupdate-cmddoes. -
Indexing no longer follows file symlinks or glob
..// absolute patterns
out of the collection directory.fast-globalready skipped symlinked
directories, but a file symlink (or a mask like../**/*.md) still resolved
viarealpathand ingested the target.qmd://filesystem resolution uses
the same containment check, soqmd://collection/../../../etc/passwdno
longer produces a path outside the collection. -
qmd mcp --httpnow validates theOriginandHostheaders on every
request and answers403when they name anything but a loopback address
(#881). Binding to localhost is no defence against the user's own browser:
a page can re-point its hostname at127.0.0.1(DNS rebinding) and read
the indexed corpus throughPOST /queryorPOST /mcp. Requests with no
Origin(curl, MCP clients, editors) are unaffected. Extend the allowlists
withQMD_ALLOWED_ORIGINS/QMD_ALLOWED_HOSTS, or set
QMD_ALLOWED_ORIGINS=*behind your own authenticating proxy. A wildcard
bind (--host 0.0.0.0) skips the host check and warns at startup.
Changed
-
Dependencies:
node-llama-cpp3.18.1 → 3.20.0 (llama.cpp b8390 → b10361, 2026-08-11). Also safe patch/minors:picomatch4.0.4 → 4.0.5,web-tree-sitter0.26.8 → 0.26.12,tsx4.21.0 → 4.23.12,vitest3.2.4 → 3.2.7. No zod/vitest major;@modelcontextprotocol/serverstays 2.0.0 (no 2.x patch).flake.nixFOD hashes are not updated here. -
generateand query expansion now awaitLlamaContextSequence.dispose()before disposing the parent context. node-llama-cpp 3.20 made sequence dispose async; the library's context-onDispose path does not wait. -
MCP server now speaks protocol revision 2026-07-28 via the official
TypeScript SDK 2.x (@modelcontextprotocol/server). HTTP is sessionless
(noMcp-Session-Id, no initialize handshake, no idle-session TTL / #816
reaper). Clients send version and capabilities in_meta;server/discover
is implemented; Streamable HTTP POST requiresMcp-Method/Mcp-Name
(mismatch →-32020);tools/listis deterministic and carriesttlMs/
cacheScope. 2025-era stdio clients still work (serveStdiodual-speak);
2025-era HTTPinitializeis answered per-request without minting a
session. Existing tools (query/get/multi_get/status), stdio
EOF shutdown, and named-index daemon PIDs are unchanged. No release. -
qmd pull(and implicit model downloads inembed/query) no longer print
node-llama-cpp's download progress bar. The bar redraws every few kilobytes
and flooded agent transcripts with thousands of tokens (#776). Pass
qmd pull --progressto show it on an interactive terminal. -
--full-pathno longer degrades silently when a result cannot be resolved on
disk (#785). A fallback there means the file moved or was deleted since the
last index, sosearch,query,getandmulti-getnow print a notice to
stderr naming how many results fell back and suggestingqmd update; stdout
stays machine-readable. -
search/querynow decide per result whether to show the docid under
--full-path, matchingmulti-getandget: a result that resolved shows
its on-disk path and no docid, one that did not keeps itsqmd://URI and
its docid, so it is still addressable. Previously the docid was dropped for
every row whenever the flag was set, leaving unresolved rows with neither a
usable path nor an identifier. -
search --format csvalways emits thedocidcolumn, empty for rows that
resolved to an on-disk path. Under--full-paththe header previously
dropped the column entirely — which also disagreed with the empty-result
header, always printed withdocid. Column positions are now stable across
runs and formats.
Fixed
-
Concurrent first-open of a cold index no longer fails with
table documents_fts already existson Bun/macOS. FTS5
CREATE VIRTUAL TABLE IF NOT EXISTSis not atomic across WAL
connections: two processes can both see a missing table on their
schema snapshot and the loser throws. Table create and legacy-schema
repair now use the sameBEGIN IMMEDIATE+ double-check as the FTS
sync triggers, and treat a concurrent "already exists" as success
when the table is present. -
Nix flake
qmd-node-modulesFOD hashes updated for x86_64-linux and
aarch64-darwin after the MCP SDK 2.0 bump.nix build/ Nix GHA was
failing with a fixed-output hash mismatch. -
CJK FTS rebuild no longer skips leftover
fts5(name, body, content='documents')
tables whenfts_cjk_normalized_versionis already stamped, and schema
repair now checks live FTS columns (PRAGMA table_info) as well as
sqlite_master.sql. The MCP HTTP test helper still seeds that legacy
table;startMcpHttpServer/createStoreon it must not throw
no such column: T.name(#792 regression). -
qmd collection add --globis no longer silently ignored. parseArgs ran
withstrict: false, so OpenClaw's--glob memory.md(and any other
--glob) fell through, the default**/*.mdwas used, and a second
collection on the same path collided as a duplicate instead of indexing
the requested mask (#536).--globis now an alias for--mask. -
The
bin/qmdtrampoline now execsprocess.execPathinstead of
re-resolvingnodefrom PATH. Native addons (better-sqlite3) are
compiled for the Node that installed qmd; a version manager (nvm, fnm,
mise) selecting a different major in the working directory used to spawn
that other binary and fail withNODE_MODULE_VERSION/ERR_DLOPEN_FAILED
(#577 leftover; #319).bun bin/qmdstill resolvesnodefrom PATH so
Node-ABI addons are not loaded into bun. -
qmd update/qmd collection addno longer swallow unreadable files
silently (#460).readFileSyncfailures (ETIMEDOUT on APFS compressed
files, EAGAIN, EACCES, …) still skip the file so the rest of the collection
indexes, but the CLI now warns with the path and error code and reports
the skip count. The SDKupdate()result includesskipped. -
The architecture diagram no longer draws Vec expansions into BM25 search
(#680).lexexpansions are FTS-only;vecandhydeexpansions are
vector-only. The original query still goes to both backends. -
qmd benchno longer runs to a wall of 0.00 when the fixture collection is
missing or empty (#716). It errors up front with the same "Collection not
found" / index hint asqmd search -c, and if every backend still scores
zero it warns on stderr to checkqmd ls. -
qmd cleanupnow reclaims the content and FTS space left behind after a
wrong-directoryqmd update. Deactivating files (the next update in the
right directory) only tombstoned thedocumentsrows; cleanup deleted those
rows and vacuumed, but never dropped the unreferencedcontenthashes and
never ran FTS5optimize, sodocuments_fts_datakept the old bodies
(#550). Cleanup now deletes inactive docs, then orphaned content, then
compact FTS, then vacuum.--dry-runreports the content hashes too. -
Concurrent
querycalls withrerank: trueon a cold MCP server no longer
raceensureRerankContexts(). Embed already serialized context creation;
rerank did not, so two overlapping first queries both saw an empty pool,
both created ranking contexts, and the inactivity timer disposed the loser
(Object is disposed, #682). Callers now await the in-flight create. -
Embedding-context pool size no longer assumes every GGUF costs 150 MB of
VRAM (the nomic-embed figure). Larger models such as Qwen3-Embedding-0.6B
are ~1190 MB per 2048-token context; opening 8 of those exhausted an 8 GB
card soqmd queryfailed withFailed to create any rerank contexteven
though the reranker itself was fine. The pool is now sized from the weight
file, and 1 GB is reserved for the reranker (#799). Default
embeddinggemma/nomic throughput is unchanged.QMD_EMBED_PARALLELISMstill
overrides. -
Multi-collection
-c A -c B(and SDK/MCPcollections: [A, B]) no longer
searches globally then post-filters. A large unrelated collection could fill
the FTS/ANN top-k so the requested collections vanished, yielding false-empty
results even though each collection matched on its own.searchFTS/
searchVecnow search each requested collection, then merge by score
(#775). Single-collection exact-scan (#791, #803) is unchanged. -
Query expansion no longer consumes caller
intent, and a cached expansion
whose sub-quer...
v2.5.3
[2.5.3] - 2026-05-28
Features
qmd getnow accepts a:from:countsuffix on a path or docid (e.g.
qmd get "#abc123:120:40"reads 40 lines starting at line 120). Explicit
--from/-lflags still override the suffix. The MCPgettool accepts the
same suffix.qmd getandqmd multi-getare now line-numbered by default and print
the document's#docidandqmd://path in the output header. Disable line
numbers with--no-line-numbers. The MCPget/multi_gettools default
lineNumberstotrueto match.qmd multi-getnow includes the#docidin every output format
(--md,--json,--csv,--xml,--files, and the default CLI view),
consistent withqmd search.qmd getandqmd multi-getaccept--full-path, which replaces the
qmd://path +#docidwith the document's on-disk filesystem path (handy for
piping intoRead/Edit/an editor). Falls back to the canonicalqmd://+
docid header when the file no longer exists on disk.qmd search/qmd querynow show a clearer hit identifier: the default CLI
view (and the new**file:**line in--mdoutput) always prints the full
qmd://collection/pathURI so you can pipe it straight back intoqmd get.qmd search/qmd queryaccept--full-pathwith the same semantics as
qmd get: the result label becomes the file's on-disk path —./-prefixed
relative path when the file lives in a subfolder of$PWD, absolute realpath
otherwise — and the per-result#docidis dropped because the path is the
identifier. The leading./is intentional so the output is unambiguously a
filesystem path. Applies to all output formats.qmd getandqmd multi-getnow also use the./-prefixed convention when
--full-pathrenders a path under$PWD, matchingsearch/query.- New
--format <kind>flag selects the output format (cli|json|csv|
md|xml|files) forsearch,query, andmulti-get. The legacy
boolean aliases (--json/--csv/--md/--xml/--files) still work but are
no longer in--help; prefer--format.
Fixes
- Launcher: source-mode runner selection now prefers Node + tsx over Bun when
bothpackage-lock.jsonandbun.lockare present in the package root,
mirroring the dist-mode "npm priority" rule. Fixes pnpm-global installs that
copy the entire working tree (including.gitandbun.lock) into the
install dir and previously routed through Bun, causing ABI mismatches with
the Node-builtbetter-sqlite3/sqlite-vecnative modules. - Darwin Metal: llama-using commands (
query,vsearch,embed) no longer
dump a multi-kB GGML/Metal backtrace at process exit even when output
succeeded. The libggml-metal staticggml_metal_devicedestructor asserts
[rsets->data count] == 0during__cxa_finalize_ranges, but the
buffer-free path never calls the symmetricggml_metal_device_rsets_rm
to remove released rsets from the device collection (upstream
ggml-org/llama.cpp#22593, one-line fix open as PR #22595). The assertion
only fires whenprocess.exit()skips Node'sbeforeExithook, which is
what node-llama-cpp uses to auto-dispose Metal contexts. Primary fix:
finishSuccessfulCliCommandnow setsprocess.exitCode = 0and returns
instead of callingprocess.exit(0), sobeforeExitfires and the native
binding cleans up before libc's static destructor runs. Defense-in-depth:
the launcher (bin/qmd) and the npm test driver (scripts/test-all.mjs- the
test:bun/test:unitpackage.json scripts) also set
GGML_METAL_NO_RESIDENCY=1on darwin before spawning node/bun, covering
error paths and tests that still terminate viaprocess.exit(). The env
var must be set before node/bun start — libggml-metal reads it via libc
getenvat module-load time, and Bun does not propagateprocess.env
mutations to libcsetenv— so it lives in the launcher rather than in
test-preload. Residency sets give no measurable speedup for QMD's
short-lived CLI workflow (benchmarked on M3 Pro). Opt back in with
QMD_METAL_KEEP_RESIDENCY=1for long-lived qmd processes (e.g. the MCP
daemon may benefit on hot reload) or to triage the upstream fix.
qmd doctorreports the mitigation state. Minimal reproduction:
scripts/repro-metal-rsets-crash.mjs.
- the
Docs
- qmd skill: emphasize reading line ranges with
get's built-in
:from:countsuffix /--from/-lflags instead of piping through
sed/head/tail; cite the docid and line numbers now present in retrieval
output; and author structuredintent:/lex:/vec:/hyde:queries yourself
rather than relying on built-in query expansion.
[2.5.2] - 2026-05-22
Fixes
- Launcher: Rewrite
bin/qmdas a Node-based shebang polyglot to fix global npm installation execution failures on Windows (#668 / #452), while supporting seamless fallback to Bun in Node-less environments.
[2.5.1] - 2026-05-20
Changes
- Release: publish from GitHub Actions via npm Trusted Publishing/OIDC instead of a long-lived
NPM_TOKENsecret.
[2.5.0] - 2026-05-19
Changes
- Dependencies: update core SQLite/config/chunking packages (
better-sqlite3,yaml,web-tree-sitter,tree-sitter-go, andtree-sitter-python) while keeping incompatiblezod,tsx, andvitestmajors pinned. - Agent skills: add
qmd skills list|get|pathto serve version-matched runtime skill instructions from the installed CLI, and makeqmd skill installwrite a stable discovery stub so installed agent skills do not go stale after QMD upgrades. - CLI: add
qmd doctorfor index/runtime diagnostics, including SQLite/sqlite-vec versions, embedding fingerprint freshness, mixed-fingerprint detection, safe legacy fingerprint adoption, and content-hash sampling.
Fixes
-
Launcher: prefer runnable TypeScript source in git checkouts even when ignored
dist/artifacts exist, while packaged installs continue to rundist/. -
GPU: keep node-llama-cpp's documented
gpu: "auto"initialization as the primary path, then perform no-build packaged CUDA/Vulkan/Metal probes only if auto falls back to CPU. -
CLI: move GPU/CPU runtime diagnostics out of
qmd status; useqmd doctorfor device probing and related environment guidance. -
CLI: point unexpected command/setup failures toward
qmd doctorso diagnostics are the default next step when QMD behaves incorrectly. -
Doctor: explicitly warn when
content_vectorscontains multiple non-empty embedding fingerprint names, with the per-fingerprint document/chunk breakdown. -
Embed: make the TTY progress line label byte-based input progress explicitly, show embedded chunks as a count, and shorten the displayed model name.
-
Embed: retain per-chunk failure details, retry failed chunks after later successful embeds and again when no other chunks remain, clear recovered errors, and cap retries to avoid endless loops.
-
Tests: expand the container smoke harness to cover npm-global, npx-style, and Bun-global install scenarios, always checking auto and
QMD_FORCE_CPU=1doctor modes, with opt-in tinyqmd embedand GPU probe runs for supported container runtimes. -
Embedding: fingerprint vector metadata using the active embedding model and formatting/chunking parameters so stale vectors are treated as pending after search semantics change. Legacy
content_vectorscolumns are migrated lazily on first vector-health/write use to preserve fast QMD startup. -
Skill: expand the packaged QMD skill with retrieval-first workflows, structured query examples, wiki/source collection guidance, and safe fallbacks when model-backed search is unavailable.
-
Tests: make
bun run testexecute the local unit suite under both Node/Vitest and Bun (test:node+test:bun) so runtime-specific regressions are caught before CI. -
Model config: centralize embedding/rerank/generation model resolution so
qmd embed,status,query,vsearch,pull, SDK vector search, andbenchuse the same active.qmd/index.yamlmodel hints and environment fallbacks. -
GPU/status:
qmd statusnow uses the same embedding model identity asqmd embedwhen computing pending embeddings, so URI-backed embeddings are not incorrectly reported as pending under the legacyembeddinggemmaalias. -
GPU status:
qmd statusnow always shows GPU mode/configuration without unsafe native probing, and CPU-fallback warnings point toQMD_STATUS_DEVICE_PROBE=1 qmd statusfor an actual backend probe. The no-GPU warning is emitted once per process instead of once per LLM instance during benchmarks. -
GPU: add
QMD_FORCE_CPU=1/--no-gputo bypass CUDA/Vulkan/Metal probing entirely, and route native llama.cpp stdout noise to stderr so JSON output stays parseable during search/query commands. -
Snippet line numbers:
qmd_query(MCP), HTTP/query, andqmd query
(CLI JSON output and snippet headers) now return absolute source-file
line numbers instead of chunk-local ones, so thelinefield can be
passed back toqmd_getasfromLinewithout a separate lookup.
Snippet selection remains scoped to the best matching chunk
(preserves #149). -
CLI:
qmd query --fullnow emits the full document body in all output
formats (json, csv, md, xml), restoring the documented behavior of the
flag. Previously it returned only the best matching chunk (~3.6KB max
per result). Output payload for--fullqueries is now proportional
to total document size. -
macOS Metal:
qmd query --jsonnow flushes successful JSON output and uses a safe immediate-exit path on Darwin to avoid ggml Metal finalizer aborts; other commands still dispose LLM contexts/models before the llama runtime. #368 -
Embedding: require complete chunk coverage before treating a document as
embedded, remove partial vectors when chunk/session failures leave a
document incomplete, and keepqmd statuspending counts honest after
interrupted long embed runs. #6...
v2.5.2
[2.5.2] - 2026-05-22
Fixes
- Launcher: Rewrite
bin/qmdas a Node-based shebang polyglot to fix global npm installation execution failures on Windows (#668 / #452), while supporting seamless fallback to Bun in Node-less environments.
[2.5.1] - 2026-05-20
Changes
- Release: publish from GitHub Actions via npm Trusted Publishing/OIDC instead of a long-lived
NPM_TOKENsecret.
[2.5.0] - 2026-05-19
Changes
- Dependencies: update core SQLite/config/chunking packages (
better-sqlite3,yaml,web-tree-sitter,tree-sitter-go, andtree-sitter-python) while keeping incompatiblezod,tsx, andvitestmajors pinned. - Agent skills: add
qmd skills list|get|pathto serve version-matched runtime skill instructions from the installed CLI, and makeqmd skill installwrite a stable discovery stub so installed agent skills do not go stale after QMD upgrades. - CLI: add
qmd doctorfor index/runtime diagnostics, including SQLite/sqlite-vec versions, embedding fingerprint freshness, mixed-fingerprint detection, safe legacy fingerprint adoption, and content-hash sampling.
Fixes
-
Launcher: prefer runnable TypeScript source in git checkouts even when ignored
dist/artifacts exist, while packaged installs continue to rundist/. -
GPU: keep node-llama-cpp's documented
gpu: "auto"initialization as the primary path, then perform no-build packaged CUDA/Vulkan/Metal probes only if auto falls back to CPU. -
CLI: move GPU/CPU runtime diagnostics out of
qmd status; useqmd doctorfor device probing and related environment guidance. -
CLI: point unexpected command/setup failures toward
qmd doctorso diagnostics are the default next step when QMD behaves incorrectly. -
Doctor: explicitly warn when
content_vectorscontains multiple non-empty embedding fingerprint names, with the per-fingerprint document/chunk breakdown. -
Embed: make the TTY progress line label byte-based input progress explicitly, show embedded chunks as a count, and shorten the displayed model name.
-
Embed: retain per-chunk failure details, retry failed chunks after later successful embeds and again when no other chunks remain, clear recovered errors, and cap retries to avoid endless loops.
-
Tests: expand the container smoke harness to cover npm-global, npx-style, and Bun-global install scenarios, always checking auto and
QMD_FORCE_CPU=1doctor modes, with opt-in tinyqmd embedand GPU probe runs for supported container runtimes. -
Embedding: fingerprint vector metadata using the active embedding model and formatting/chunking parameters so stale vectors are treated as pending after search semantics change. Legacy
content_vectorscolumns are migrated lazily on first vector-health/write use to preserve fast QMD startup. -
Skill: expand the packaged QMD skill with retrieval-first workflows, structured query examples, wiki/source collection guidance, and safe fallbacks when model-backed search is unavailable.
-
Tests: make
bun run testexecute the local unit suite under both Node/Vitest and Bun (test:node+test:bun) so runtime-specific regressions are caught before CI. -
Model config: centralize embedding/rerank/generation model resolution so
qmd embed,status,query,vsearch,pull, SDK vector search, andbenchuse the same active.qmd/index.yamlmodel hints and environment fallbacks. -
GPU/status:
qmd statusnow uses the same embedding model identity asqmd embedwhen computing pending embeddings, so URI-backed embeddings are not incorrectly reported as pending under the legacyembeddinggemmaalias. -
GPU status:
qmd statusnow always shows GPU mode/configuration without unsafe native probing, and CPU-fallback warnings point toQMD_STATUS_DEVICE_PROBE=1 qmd statusfor an actual backend probe. The no-GPU warning is emitted once per process instead of once per LLM instance during benchmarks. -
GPU: add
QMD_FORCE_CPU=1/--no-gputo bypass CUDA/Vulkan/Metal probing entirely, and route native llama.cpp stdout noise to stderr so JSON output stays parseable during search/query commands. -
Snippet line numbers:
qmd_query(MCP), HTTP/query, andqmd query
(CLI JSON output and snippet headers) now return absolute source-file
line numbers instead of chunk-local ones, so thelinefield can be
passed back toqmd_getasfromLinewithout a separate lookup.
Snippet selection remains scoped to the best matching chunk
(preserves #149). -
CLI:
qmd query --fullnow emits the full document body in all output
formats (json, csv, md, xml), restoring the documented behavior of the
flag. Previously it returned only the best matching chunk (~3.6KB max
per result). Output payload for--fullqueries is now proportional
to total document size. -
macOS Metal:
qmd query --jsonnow flushes successful JSON output and uses a safe immediate-exit path on Darwin to avoid ggml Metal finalizer aborts; other commands still dispose LLM contexts/models before the llama runtime. #368 -
Embedding: require complete chunk coverage before treating a document as
embedded, remove partial vectors when chunk/session failures leave a
document incomplete, and keepqmd statuspending counts honest after
interrupted long embed runs. #637 #378 -
Embedding:
qmd embed -c <collection>now scopes pending-doc selection
to the requested collection instead of embedding global pending work.
Scoped--forceclears only collection-owned vectors, preserves shared
hashes referenced by sibling collections, and dropsvectors_veconly
when the scoped clear empties all vectors. -
Hybrid search: weight RRF lists by query type so original FTS and original vector evidence get the intended 2x boost, instead of accidentally boosting the first lexical expansion. #591
-
MCP: seed llama.cpp/GGML quiet env vars before launching
qmd mcpso native logs cannot pollute stdio JSON-RPC framing. #593 -
CLI: remove CommonJS
require()calls from ESM index path normalization soqmd --index <path>no longer crashes withERR_AMBIGUOUS_MODULE_SYNTAXon Node 22+. #634 -
Windows CUDA: serialize llama.cpp embedding/reranking contexts by default to avoid intermittent
ggml-cuda.cu:98crashes inqmd query; setQMD_EMBED_PARALLELISMto opt back into parallel contexts if your driver is stable. #519 -
MCP: make
qmd mcp --index <name>use the selected index for both foreground and daemon HTTP servers instead of falling back to the default store. #343 -
Embedding: respect
QMD_EMBED_MODELconsistently for vector indexing and vector-backed search, with default-model fallback when unset. -
Config: use one home-directory resolver for YAML config and the default SQLite cache path, avoiding Windows CLI/MCP split-brain when
HOMEis unset. -
GPU: respect explicit
QMD_LLAMA_GPU=metal|vulkan|cudabackend overrides instead of always using auto GPU selection. #529 -
Fix: preserve original filename case in
handelize(). The previous
.toLowerCase()call made indexed paths unreachable on case-sensitive
filesystems (Linux).qmd updateautomatically migrates legacy
lowercase paths without re-embedding. -
CLI: make
qmd statusskip nativenode-llama-cppdevice probing by
default so status stays safe on machines with broken or unsupported GPU
drivers. SetQMD_STATUS_DEVICE_PROBE=1to opt in. -
CLI: lazy-load
node-llama-cppso lightweight commands such as
qmd statusdo not import native ML dependencies or trigger llama.cpp
builds on ARM/no-GPU machines. #491 -
Store: keep content rows referenced by inactive documents during orphan
cleanup soqmd updatepreserves soft-deleted tombstones for removed
files. #585 -
Packaging: install AST grammar WASM packages as required dependencies so
Bun global installs include TypeScript/TSX/JavaScript grammars, and add a
smoke:package-grammarsverification command. #595 -
Launcher: add wrapper smoke coverage for scoped package, npm/npx,
Homebrew/Linuxbrew, Bun global symlink layouts, and$BUN_INSTALL
false-positive runtime selection regressions. #351 #353 #354 #356 #358 #359
v2.5.1
[2.5.1] - 2026-05-20
Changes
- Release: publish from GitHub Actions via npm Trusted Publishing/OIDC instead of a long-lived
NPM_TOKENsecret.
[2.5.0] - 2026-05-19
Changes
- Dependencies: update core SQLite/config/chunking packages (
better-sqlite3,yaml,web-tree-sitter,tree-sitter-go, andtree-sitter-python) while keeping incompatiblezod,tsx, andvitestmajors pinned. - Agent skills: add
qmd skills list|get|pathto serve version-matched runtime skill instructions from the installed CLI, and makeqmd skill installwrite a stable discovery stub so installed agent skills do not go stale after QMD upgrades. - CLI: add
qmd doctorfor index/runtime diagnostics, including SQLite/sqlite-vec versions, embedding fingerprint freshness, mixed-fingerprint detection, safe legacy fingerprint adoption, and content-hash sampling.
Fixes
-
Launcher: prefer runnable TypeScript source in git checkouts even when ignored
dist/artifacts exist, while packaged installs continue to rundist/. -
GPU: keep node-llama-cpp's documented
gpu: "auto"initialization as the primary path, then perform no-build packaged CUDA/Vulkan/Metal probes only if auto falls back to CPU. -
CLI: move GPU/CPU runtime diagnostics out of
qmd status; useqmd doctorfor device probing and related environment guidance. -
CLI: point unexpected command/setup failures toward
qmd doctorso diagnostics are the default next step when QMD behaves incorrectly. -
Doctor: explicitly warn when
content_vectorscontains multiple non-empty embedding fingerprint names, with the per-fingerprint document/chunk breakdown. -
Embed: make the TTY progress line label byte-based input progress explicitly, show embedded chunks as a count, and shorten the displayed model name.
-
Embed: retain per-chunk failure details, retry failed chunks after later successful embeds and again when no other chunks remain, clear recovered errors, and cap retries to avoid endless loops.
-
Tests: expand the container smoke harness to cover npm-global, npx-style, and Bun-global install scenarios, always checking auto and
QMD_FORCE_CPU=1doctor modes, with opt-in tinyqmd embedand GPU probe runs for supported container runtimes. -
Embedding: fingerprint vector metadata using the active embedding model and formatting/chunking parameters so stale vectors are treated as pending after search semantics change. Legacy
content_vectorscolumns are migrated lazily on first vector-health/write use to preserve fast QMD startup. -
Skill: expand the packaged QMD skill with retrieval-first workflows, structured query examples, wiki/source collection guidance, and safe fallbacks when model-backed search is unavailable.
-
Tests: make
bun run testexecute the local unit suite under both Node/Vitest and Bun (test:node+test:bun) so runtime-specific regressions are caught before CI. -
Model config: centralize embedding/rerank/generation model resolution so
qmd embed,status,query,vsearch,pull, SDK vector search, andbenchuse the same active.qmd/index.yamlmodel hints and environment fallbacks. -
GPU/status:
qmd statusnow uses the same embedding model identity asqmd embedwhen computing pending embeddings, so URI-backed embeddings are not incorrectly reported as pending under the legacyembeddinggemmaalias. -
GPU status:
qmd statusnow always shows GPU mode/configuration without unsafe native probing, and CPU-fallback warnings point toQMD_STATUS_DEVICE_PROBE=1 qmd statusfor an actual backend probe. The no-GPU warning is emitted once per process instead of once per LLM instance during benchmarks. -
GPU: add
QMD_FORCE_CPU=1/--no-gputo bypass CUDA/Vulkan/Metal probing entirely, and route native llama.cpp stdout noise to stderr so JSON output stays parseable during search/query commands. -
Snippet line numbers:
qmd_query(MCP), HTTP/query, andqmd query
(CLI JSON output and snippet headers) now return absolute source-file
line numbers instead of chunk-local ones, so thelinefield can be
passed back toqmd_getasfromLinewithout a separate lookup.
Snippet selection remains scoped to the best matching chunk
(preserves #149). -
CLI:
qmd query --fullnow emits the full document body in all output
formats (json, csv, md, xml), restoring the documented behavior of the
flag. Previously it returned only the best matching chunk (~3.6KB max
per result). Output payload for--fullqueries is now proportional
to total document size. -
macOS Metal:
qmd query --jsonnow flushes successful JSON output and uses a safe immediate-exit path on Darwin to avoid ggml Metal finalizer aborts; other commands still dispose LLM contexts/models before the llama runtime. #368 -
Embedding: require complete chunk coverage before treating a document as
embedded, remove partial vectors when chunk/session failures leave a
document incomplete, and keepqmd statuspending counts honest after
interrupted long embed runs. #637 #378 -
Embedding:
qmd embed -c <collection>now scopes pending-doc selection
to the requested collection instead of embedding global pending work.
Scoped--forceclears only collection-owned vectors, preserves shared
hashes referenced by sibling collections, and dropsvectors_veconly
when the scoped clear empties all vectors. -
Hybrid search: weight RRF lists by query type so original FTS and original vector evidence get the intended 2x boost, instead of accidentally boosting the first lexical expansion. #591
-
MCP: seed llama.cpp/GGML quiet env vars before launching
qmd mcpso native logs cannot pollute stdio JSON-RPC framing. #593 -
CLI: remove CommonJS
require()calls from ESM index path normalization soqmd --index <path>no longer crashes withERR_AMBIGUOUS_MODULE_SYNTAXon Node 22+. #634 -
Windows CUDA: serialize llama.cpp embedding/reranking contexts by default to avoid intermittent
ggml-cuda.cu:98crashes inqmd query; setQMD_EMBED_PARALLELISMto opt back into parallel contexts if your driver is stable. #519 -
MCP: make
qmd mcp --index <name>use the selected index for both foreground and daemon HTTP servers instead of falling back to the default store. #343 -
Embedding: respect
QMD_EMBED_MODELconsistently for vector indexing and vector-backed search, with default-model fallback when unset. -
Config: use one home-directory resolver for YAML config and the default SQLite cache path, avoiding Windows CLI/MCP split-brain when
HOMEis unset. -
GPU: respect explicit
QMD_LLAMA_GPU=metal|vulkan|cudabackend overrides instead of always using auto GPU selection. #529 -
Fix: preserve original filename case in
handelize(). The previous
.toLowerCase()call made indexed paths unreachable on case-sensitive
filesystems (Linux).qmd updateautomatically migrates legacy
lowercase paths without re-embedding. -
CLI: make
qmd statusskip nativenode-llama-cppdevice probing by
default so status stays safe on machines with broken or unsupported GPU
drivers. SetQMD_STATUS_DEVICE_PROBE=1to opt in. -
CLI: lazy-load
node-llama-cppso lightweight commands such as
qmd statusdo not import native ML dependencies or trigger llama.cpp
builds on ARM/no-GPU machines. #491 -
Store: keep content rows referenced by inactive documents during orphan
cleanup soqmd updatepreserves soft-deleted tombstones for removed
files. #585 -
Packaging: install AST grammar WASM packages as required dependencies so
Bun global installs include TypeScript/TSX/JavaScript grammars, and add a
smoke:package-grammarsverification command. #595 -
Launcher: add wrapper smoke coverage for scoped package, npm/npx,
Homebrew/Linuxbrew, Bun global symlink layouts, and$BUN_INSTALL
false-positive runtime selection regressions. #351 #353 #354 #356 #358 #359
v2.1.0
[2.1.0] - 2026-04-05
Code files now chunk at function and class boundaries via tree-sitter,
clickable editor links land you at the right line from search results,
and per-collection model configuration means you can point different
collections at different embedding models. 25+ community PRs fix
embedding stability, BM25 accuracy, and cross-platform launcher issues.
Changes
- AST-aware chunking for code files via
web-tree-sitter. Supported
languages: TypeScript/JavaScript, Python, Go, and Rust. Code files
are chunked at function, class, and import boundaries instead of
arbitrary text positions. Markdown and unknown file types are unchanged.
--chunk-strategy <auto|regex>flag onqmd embedandqmd query
(defaultregex). SDK:chunkStrategyoption onembed()and
search().qmd statusshows grammar availability. qmd bench <fixture.json>command for search quality benchmarks.
Measures precision@k, recall, MRR, and F1 across BM25, vector, hybrid,
and full pipeline backends. Ships with an example fixture against
the eval-docs test collection. #470 (thanks @jmilinovich)models:section inindex.ymllets you configureembed,rerank,
andgeneratemodel URIs per collection. Resolution order is
config > env var (QMD_EMBED_MODEL,QMD_RERANK_MODEL,
QMD_GENERATE_MODEL) > built-in default. #502
(thanks @JohnRichardEnders)- CLI search output now emits clickable OSC 8 terminal hyperlinks when
stdout is a TTY. Links resolveqmd://paths to absolute filesystem
paths and open in editors via URI templates (default:
vscode://file/{path}:{line}:{col}). Configure withQMD_EDITOR_URI
oreditor_uriin the YAML config. #508 (thanks @danmackinlay) --no-rerankflag skips the reranking step inqmd query— useful
when you want fast results or don't have a GPU. Also exposed as
rerank: falseon the MCPquerytool. #370 (thanks @mvanhorn),
#478 (thanks @zestyboy)- ONNX conversion script for deploying embedding models via
Transformers.js. #399 (thanks @shreyaskarnik) - GitHub Actions workflow to build the Nix flake on Linux and macOS.
Fixes
- Embedding: prevent
qmd embedfrom running indefinitely when the
embedding loop stalls. #458 (thanks @ccc-fff) - Embedding: truncate oversized text before embedding to prevent GGML
crash, and bound memory usage during batch embedding. #393
(thanks @lskun), #395 (thanks @ProgramCaiCai) - Embedding: set explicit embed context size (default 2048, configurable
viaQMD_EMBED_CONTEXT_SIZE) instead of using the model's full
window. #500 - Embedding: error on dimension mismatch instead of silently rebuilding
the vec0 table. #501 - Embedding: handle vec0
OR REPLACElimitation ininsertEmbedding.
#456 (thanks @antonio-mello-ai) - Embedding: fix model selection when multiple models are configured.
#494 - BM25: correct field weights to include all 3 FTS columns — title,
body, and path were not weighted correctly. #462 (thanks @goldsr09) - BM25: handle hyphenated tokens in FTS5 lex queries so terms like
"real-time" match correctly. #463 (thanks @goldsr09) - BM25: preserve underscores in search terms instead of stripping them.
#404 - BM25: use CTE in
searchFTSto prevent query planner regression with
collection filter. - Reranker: increase default context size 2048→4096 and make
configurable viaQMD_RERANK_CONTEXT_SIZE. Fix template overhead
underestimate 200→512. #453 (thanks @builderjarvis) - GPU: catch initialization failures and fall back to CPU instead of
crashing. - MCP: read version from
package.jsoninstead of hardcoding. #431 - MCP: include collection name in status output. #416
- Multi-get: support brace expansion patterns in glob matching. #424
- Launcher: prioritize
package-lock.jsonto prevent Bun false
positive. #385 (thanks @rymalia) - Launcher: remove
$BUN_INSTALLcheck that caused false Bun detection.
#362 (thanks @syedair) - Launcher: skip Git Bash path detection on WSL. #371
(thanks @oysteinkrog) - Model cache: respect
XDG_CACHE_HOMEfor model cache directory. #457
(thanks @antonio-mello-ai) - SQLite: add macOS Homebrew SQLite support for Bun and restore
actionable errors. #377 (thanks @serhii12) - Pin zod to exact 4.2.1 to fix
tscbuild failure. #382
(thanks @rymalia) - Preserve dots and original case in
handelize()— filenames like
MEMORY.mdno longer becomememory-md. #475 (thanks @alexei-led) - Include
linein--jsonsearch output so editor integrations can
jump directly tofile:line. #506 (thanks @danmackinlay) - Nix: fix paths in flake and make Bun dependency a fixed-output
derivation so sandboxed Linux builds work offline. #479
(thanks @surma-dump) - Sync stale
bun.lock(better-sqlite311.x → 12.x). CI and release
script now use--frozen-lockfileto prevent recurrence. #386
(thanks @Mic92) - Approve native build scripts in pnpm so
better-sqlite3and
tree-sitter modules compile correctly. Update vitest ^3.0.0 → ^3.2.4.
v2.0.1
[2.0.1] - 2026-03-10
Changes
qmd skill installcopies the packaged QMD skill into
~/.claude/commands/for one-command setup. #355 (thanks @nibzard)
Fixes
- Fix Qwen3-Embedding GGUF filename case — HuggingFace filenames are
case-sensitive, the lowercase variant returned 404. #349 (thanks @byheaven) - Resolve symlinked global launcher path so
qmdworks correctly when
installed vianpm i -g. #352 (thanks @nibzard)
[2.0.0] - 2026-03-10
QMD 2.0 declares a stable library API. The SDK is now the primary interface —
the MCP server is a clean consumer of it, and the source is organized into
src/cli/ and src/mcp/. Also: Node 25 support and a runtime-aware bin wrapper
for bun installs.
Changes
- Stable SDK API with
QMDStoreinterface — search, retrieval, collection/context
management, indexing, lifecycle - Unified
search(): passqueryfor auto-expansion orqueriesfor
pre-expanded lex/vec/hyde — replaces the old query/search/structuredSearch split - New
getDocumentBody(),getDefaultCollectionNames(),Maintenanceclass - MCP server rewritten as a clean SDK consumer — zero internal store access
- CLI and MCP organized into
src/cli/andsrc/mcp/subdirectories - Runtime-aware
bin/qmdwrapper detects bun vs node to avoid ABI mismatches.
Closes #319 better-sqlite3bumped to ^12.4.5 for Node 25 support. Closes #257- Utility exports:
extractSnippet,addLineNumbers,DEFAULT_MULTI_GET_MAX_BYTES
Fixes
- Remove unused
import { resolve }in store.ts that shadowed local export
v2.0.0
[2.0.0] - 2026-03-10
QMD 2.0 declares a stable library API. The SDK is now the primary interface —
the MCP server is a clean consumer of it, and the source is organized into
src/cli/ and src/mcp/. Also: Node 25 support and a runtime-aware bin wrapper
for bun installs.
Changes
- Stable SDK API with
QMDStoreinterface — search, retrieval, collection/context
management, indexing, lifecycle - Unified
search(): passqueryfor auto-expansion orqueriesfor
pre-expanded lex/vec/hyde — replaces the old query/search/structuredSearch split - New
getDocumentBody(),getDefaultCollectionNames(),Maintenanceclass - MCP server rewritten as a clean SDK consumer — zero internal store access
- CLI and MCP organized into
src/cli/andsrc/mcp/subdirectories - Runtime-aware
bin/qmdwrapper detects bun vs node to avoid ABI mismatches.
Closes #319 better-sqlite3bumped to ^12.4.5 for Node 25 support. Closes #257- Utility exports:
extractSnippet,addLineNumbers,DEFAULT_MULTI_GET_MAX_BYTES
Fixes
- Remove unused
import { resolve }in store.ts that shadowed local export
v1.1.6
[1.1.6] - 2026-03-09
QMD can now be used as a library. import { createStore } from '@tobilu/qmd'
gives you the full search and indexing API — hybrid query, BM25, structured
search, collection/context management — without shelling out to the CLI.
Changes
- SDK / library mode:
createStore({ dbPath, config })returns a
QMDStorewithquery(),search(),structuredSearch(),get(),
multiGet(), and collection/context management methods. Supports inline
config (no files needed) or a YAML config path. - Package exports:
package.jsonnow declaresmain,types, and
exportsso bundlers and TypeScript resolve@tobilu/qmdcorrectly.
[1.1.5] - 2026-03-07
Ambiguous queries like "performance" now produce dramatically better results
when the caller knows what they mean. The new intent parameter steers all
five pipeline stages — expansion, strong-signal bypass, chunk selection,
reranking, and snippet extraction — without searching on its own. Design and
original implementation by Ilya Grigorik (@vyalamar) in #180.
Changes
- Intent parameter: optional
intentstring disambiguates queries across
the entire search pipeline. Available via CLI (--intentflag orintent:
line in query documents), MCP (intentfield on the query tool), and
programmatic API. Adapted from PR #180 (thanks @vyalamar). - Query expansion: when intent is provided, the expansion LLM prompt
includesQuery intent: {intent}, matching the finetune training data
format for better-aligned expansions. - Reranking: intent is prepended to the rerank query so Qwen3-Reranker
scores with domain context. - Chunk selection: intent terms scored at 0.5× weight alongside query
terms (1.0×) when selecting the best chunk per document for reranking. - Snippet extraction: intent terms scored at 0.3× weight to nudge
snippets toward intent-relevant lines without overriding query anchoring. - Strong-signal bypass disabled with intent: when intent is provided, the
BM25 strong-signal shortcut is skipped — the obvious keyword match may not
be what the caller wants. - MCP instructions: callers are now guided to provide
intenton every
search call for disambiguation. - Query document syntax:
intent:recognized as a line type. At most one
per document, cannot appear alone. Grammar updated indocs/SYNTAX.md.
[1.1.2] - 2026-03-07
13 community PRs merged. GPU initialization replaced with node-llama-cpp's
built-in autoAttempt — deleting ~220 lines of manual fallback code and
fixing GPU issues reported across 10+ PRs in one shot. Reranking is faster
through chunk deduplication and a parallelism cap that prevents VRAM
exhaustion.
Changes
- GPU init: use node-llama-cpp's
build: "autoAttempt"instead of manual
GPU backend detection. Automatically tries Metal/CUDA/Vulkan and falls back
gracefully. #310 (thanks @giladgd — the node-llama-cpp author) - Query
--explain:qmd query --explainexposes retrieval score traces
— backend scores, per-list RRF contributions, top-rank bonus, reranker
score, and final blended score. Works in JSON and CLI output. #242
(thanks @vyalamar) - Collection ignore patterns:
ignore: ["Sessions/**", "*.tmp"]in
collection config to exclude files from indexing. #304 (thanks @sebkouba) - Multilingual embeddings:
QMD_EMBED_MODELenv var lets you swap in
models like Qwen3-Embedding for non-English collections. #273 (thanks
@daocoding) - Configurable expansion context:
QMD_EXPAND_CONTEXT_SIZEenv var
(default 2048) — previously used the model's full 40960-token window,
wasting VRAM. #313 (thanks @0xble) candidateLimitexposed:-C/--candidate-limitflag and MCP
parameter to tune how many candidates reach the reranker. #255 (thanks
@pandysp)- MCP multi-session: HTTP transport now supports multiple concurrent
client sessions, each with its own server instance. #286 (thanks @joelev)
Fixes
- Reranking performance: cap parallel rerank contexts at 4 to prevent
VRAM exhaustion on high-core machines. Deduplicate identical chunk texts
before reranking — same content from different files now shares a single
reranker call. Cache scores by content hash instead of file path. - Deactivate stale docs when all files are removed from a collection and
qmd updateis run. #312 (thanks @0xble) - Handle emoji-only filenames (
🐘.md→1f418.md) instead of crashing.
#308 (thanks @debugerman) - Skip unreadable files during indexing (e.g. iCloud-evicted files returning
EAGAIN) instead of crashing. #253 (thanks @jimmynail) - Suppress progress bar escape sequences when stderr is not a TTY. #230
(thanks @dgilperez) - Emit format-appropriate empty output (
[]for JSON, CSV header for CSV,
etc.) instead of plain text "No results." #228 (thanks @amsminn) - Correct Windows sqlite-vec package name (
sqlite-vec-windows-x64) and add
sqlite-vec-linux-arm64. #225 (thanks @ilepn) - Fix claude plugin setup CLI commands in README. #311 (thanks @gi11es)
[1.1.1] - 2026-03-06
Fixes
- Reranker: truncate documents exceeding the 2048-token context window
instead of silently producing garbage scores. Long chunks (e.g. from
PDF ingestion) now get a fair ranking. - Nix: add python3 and cctools to build dependencies. #214 (thanks
@pcasaretto)
[1.1.0] - 2026-02-20
QMD now speaks in query documents — structured multi-line queries where every line is typed (lex:, vec:, hyde:), combining keyword precision with semantic recall. A single plain query still works exactly as before (it's treated as an implicit expand: and auto-expanded by the LLM). Lex now supports quoted phrases and negation ("C++ performance" -sports -athlete), making intent-aware disambiguation practical. The formal query grammar is documented in docs/SYNTAX.md.
The npm package now uses the standard #!/usr/bin/env node bin convention, replacing the custom bash wrapper. This fixes native module ABI mismatches when installed via bun and works on any platform with node >= 22 on PATH.
Changes
- Query document format: multi-line queries with typed sub-queries (
lex:,vec:,hyde:). Plain queries remain the default (expand:implicit, but not written inside the document). First sub-query gets 2× fusion weight — put your strongest signal first. Formal grammar indocs/SYNTAX.md. - Lex syntax: full BM25 operator support.
"exact phrase"for verbatim matching;-termand-"phrase"for exclusions. Essential for disambiguation when a term is overloaded across domains (e.g.performance -sports -athlete). expand:shortcut: send a single plain query (or start the document withexpand:on its only line) to auto-expand via the local LLM. Query documents themselves are limited tolex,vec, andhydelines.- MCP
querytool (renamed fromstructured_search): rewrote the tool description to fully teach AI agents the query document format, lex syntax, and combination strategy. Includes worked examples with intent-aware lex. - HTTP
/queryendpoint (renamed from/search;/searchkept as silent alias). collectionsarray filter: filter by multiple collections in a single query (collections: ["notes", "brain"]). Removed the singlecollectionstring param — array only.- Collection
include/exclude:includeByDefault: falsehides a collection from all queries unless explicitly named viacollections. CLI:qmd collection exclude <name>/qmd collection include <name>. - Collection
update-cmd: attach a shell command that runs before everyqmd update(e.g.git stash && git pull --rebase --ff-only && git stash pop). CLI:qmd collection update-cmd <name> '<cmd>'. qmd statustips: shows actionable tips when collections lack context descriptions or update commands.qmd collectionsubcommands:show,update-cmd,include,exclude. Bareqmd collectionnow prints help.- Packaging: replaced custom bash wrapper with standard
#!/usr/bin/env nodeshebang ondist/qmd.js. Fixes native module ABI mismatches when installed via bun, and works on any platform where node >= 22 is on PATH. - Removed MCP tools
search,vector_search,deep_search— all superseded byquery. - Removed
qmd context checkcommand. - CLI timing: each LLM step (expand, embed, rerank) prints elapsed time inline (
Expanding query... (4.2s)).
Fixes
qmd collection listshows[excluded]tag for collections withincludeByDefault: false.- Default searches now respect
includeByDefault— excluded collections are skipped unless explicitly named. - Fix main module detection when installed globally via npm/bun (symlink resolution).
v1.1.5
[1.1.5] - 2026-03-07
Ambiguous queries like "performance" now produce dramatically better results
when the caller knows what they mean. The new intent parameter steers all
five pipeline stages — expansion, strong-signal bypass, chunk selection,
reranking, and snippet extraction — without searching on its own. Design and
original implementation by Ilya Grigorik (@vyalamar) in #180.
Changes
- Intent parameter: optional
intentstring disambiguates queries across
the entire search pipeline. Available via CLI (--intentflag orintent:
line in query documents), MCP (intentfield on the query tool), and
programmatic API. Adapted from PR #180 (thanks @vyalamar). - Query expansion: when intent is provided, the expansion LLM prompt
includesQuery intent: {intent}, matching the finetune training data
format for better-aligned expansions. - Reranking: intent is prepended to the rerank query so Qwen3-Reranker
scores with domain context. - Chunk selection: intent terms scored at 0.5× weight alongside query
terms (1.0×) when selecting the best chunk per document for reranking. - Snippet extraction: intent terms scored at 0.3× weight to nudge
snippets toward intent-relevant lines without overriding query anchoring. - Strong-signal bypass disabled with intent: when intent is provided, the
BM25 strong-signal shortcut is skipped — the obvious keyword match may not
be what the caller wants. - MCP instructions: callers are now guided to provide
intenton every
search call for disambiguation. - Query document syntax:
intent:recognized as a line type. At most one
per document, cannot appear alone. Grammar updated indocs/SYNTAX.md.
[1.1.2] - 2026-03-07
13 community PRs merged. GPU initialization replaced with node-llama-cpp's
built-in autoAttempt — deleting ~220 lines of manual fallback code and
fixing GPU issues reported across 10+ PRs in one shot. Reranking is faster
through chunk deduplication and a parallelism cap that prevents VRAM
exhaustion.
Changes
- GPU init: use node-llama-cpp's
build: "autoAttempt"instead of manual
GPU backend detection. Automatically tries Metal/CUDA/Vulkan and falls back
gracefully. #310 (thanks @giladgd — the node-llama-cpp author) - Query
--explain:qmd query --explainexposes retrieval score traces
— backend scores, per-list RRF contributions, top-rank bonus, reranker
score, and final blended score. Works in JSON and CLI output. #242
(thanks @vyalamar) - Collection ignore patterns:
ignore: ["Sessions/**", "*.tmp"]in
collection config to exclude files from indexing. #304 (thanks @sebkouba) - Multilingual embeddings:
QMD_EMBED_MODELenv var lets you swap in
models like Qwen3-Embedding for non-English collections. #273 (thanks
@daocoding) - Configurable expansion context:
QMD_EXPAND_CONTEXT_SIZEenv var
(default 2048) — previously used the model's full 40960-token window,
wasting VRAM. #313 (thanks @0xble) candidateLimitexposed:-C/--candidate-limitflag and MCP
parameter to tune how many candidates reach the reranker. #255 (thanks
@pandysp)- MCP multi-session: HTTP transport now supports multiple concurrent
client sessions, each with its own server instance. #286 (thanks @joelev)
Fixes
- Reranking performance: cap parallel rerank contexts at 4 to prevent
VRAM exhaustion on high-core machines. Deduplicate identical chunk texts
before reranking — same content from different files now shares a single
reranker call. Cache scores by content hash instead of file path. - Deactivate stale docs when all files are removed from a collection and
qmd updateis run. #312 (thanks @0xble) - Handle emoji-only filenames (
🐘.md→1f418.md) instead of crashing.
#308 (thanks @debugerman) - Skip unreadable files during indexing (e.g. iCloud-evicted files returning
EAGAIN) instead of crashing. #253 (thanks @jimmynail) - Suppress progress bar escape sequences when stderr is not a TTY. #230
(thanks @dgilperez) - Emit format-appropriate empty output (
[]for JSON, CSV header for CSV,
etc.) instead of plain text "No results." #228 (thanks @amsminn) - Correct Windows sqlite-vec package name (
sqlite-vec-windows-x64) and add
sqlite-vec-linux-arm64. #225 (thanks @ilepn) - Fix claude plugin setup CLI commands in README. #311 (thanks @gi11es)
[1.1.1] - 2026-03-06
Fixes
- Reranker: truncate documents exceeding the 2048-token context window
instead of silently producing garbage scores. Long chunks (e.g. from
PDF ingestion) now get a fair ranking. - Nix: add python3 and cctools to build dependencies. #214 (thanks
@pcasaretto)
[1.1.0] - 2026-02-20
QMD now speaks in query documents — structured multi-line queries where every line is typed (lex:, vec:, hyde:), combining keyword precision with semantic recall. A single plain query still works exactly as before (it's treated as an implicit expand: and auto-expanded by the LLM). Lex now supports quoted phrases and negation ("C++ performance" -sports -athlete), making intent-aware disambiguation practical. The formal query grammar is documented in docs/SYNTAX.md.
The npm package now uses the standard #!/usr/bin/env node bin convention, replacing the custom bash wrapper. This fixes native module ABI mismatches when installed via bun and works on any platform with node >= 22 on PATH.
Changes
- Query document format: multi-line queries with typed sub-queries (
lex:,vec:,hyde:). Plain queries remain the default (expand:implicit, but not written inside the document). First sub-query gets 2× fusion weight — put your strongest signal first. Formal grammar indocs/SYNTAX.md. - Lex syntax: full BM25 operator support.
"exact phrase"for verbatim matching;-termand-"phrase"for exclusions. Essential for disambiguation when a term is overloaded across domains (e.g.performance -sports -athlete). expand:shortcut: send a single plain query (or start the document withexpand:on its only line) to auto-expand via the local LLM. Query documents themselves are limited tolex,vec, andhydelines.- MCP
querytool (renamed fromstructured_search): rewrote the tool description to fully teach AI agents the query document format, lex syntax, and combination strategy. Includes worked examples with intent-aware lex. - HTTP
/queryendpoint (renamed from/search;/searchkept as silent alias). collectionsarray filter: filter by multiple collections in a single query (collections: ["notes", "brain"]). Removed the singlecollectionstring param — array only.- Collection
include/exclude:includeByDefault: falsehides a collection from all queries unless explicitly named viacollections. CLI:qmd collection exclude <name>/qmd collection include <name>. - Collection
update-cmd: attach a shell command that runs before everyqmd update(e.g.git stash && git pull --rebase --ff-only && git stash pop). CLI:qmd collection update-cmd <name> '<cmd>'. qmd statustips: shows actionable tips when collections lack context descriptions or update commands.qmd collectionsubcommands:show,update-cmd,include,exclude. Bareqmd collectionnow prints help.- Packaging: replaced custom bash wrapper with standard
#!/usr/bin/env nodeshebang ondist/qmd.js. Fixes native module ABI mismatches when installed via bun, and works on any platform where node >= 22 is on PATH. - Removed MCP tools
search,vector_search,deep_search— all superseded byquery. - Removed
qmd context checkcommand. - CLI timing: each LLM step (expand, embed, rerank) prints elapsed time inline (
Expanding query... (4.2s)).
Fixes
qmd collection listshows[excluded]tag for collections withincludeByDefault: false.- Default searches now respect
includeByDefault— excluded collections are skipped unless explicitly named. - Fix main module detection when installed globally via npm/bun (symlink resolution).
v1.1.2
[1.1.2] - 2026-03-07
13 community PRs merged. GPU initialization replaced with node-llama-cpp's
built-in autoAttempt — deleting ~220 lines of manual fallback code and
fixing GPU issues reported across 10+ PRs in one shot. Reranking is faster
through chunk deduplication and a parallelism cap that prevents VRAM
exhaustion.
Changes
- GPU init: use node-llama-cpp's
build: "autoAttempt"instead of manual
GPU backend detection. Automatically tries Metal/CUDA/Vulkan and falls back
gracefully. #310 (thanks @giladgd — the node-llama-cpp author) - Query
--explain:qmd query --explainexposes retrieval score traces
— backend scores, per-list RRF contributions, top-rank bonus, reranker
score, and final blended score. Works in JSON and CLI output. #242
(thanks @vyalamar) - Collection ignore patterns:
ignore: ["Sessions/**", "*.tmp"]in
collection config to exclude files from indexing. #304 (thanks @sebkouba) - Multilingual embeddings:
QMD_EMBED_MODELenv var lets you swap in
models like Qwen3-Embedding for non-English collections. #273 (thanks
@daocoding) - Configurable expansion context:
QMD_EXPAND_CONTEXT_SIZEenv var
(default 2048) — previously used the model's full 40960-token window,
wasting VRAM. #313 (thanks @0xble) candidateLimitexposed:-C/--candidate-limitflag and MCP
parameter to tune how many candidates reach the reranker. #255 (thanks
@pandysp)- MCP multi-session: HTTP transport now supports multiple concurrent
client sessions, each with its own server instance. #286 (thanks @joelev)
Fixes
- Reranking performance: cap parallel rerank contexts at 4 to prevent
VRAM exhaustion on high-core machines. Deduplicate identical chunk texts
before reranking — same content from different files now shares a single
reranker call. Cache scores by content hash instead of file path. - Deactivate stale docs when all files are removed from a collection and
qmd updateis run. #312 (thanks @0xble) - Handle emoji-only filenames (
🐘.md→1f418.md) instead of crashing.
#308 (thanks @debugerman) - Skip unreadable files during indexing (e.g. iCloud-evicted files returning
EAGAIN) instead of crashing. #253 (thanks @jimmynail) - Suppress progress bar escape sequences when stderr is not a TTY. #230
(thanks @dgilperez) - Emit format-appropriate empty output (
[]for JSON, CSV header for CSV,
etc.) instead of plain text "No results." #228 (thanks @amsminn) - Correct Windows sqlite-vec package name (
sqlite-vec-windows-x64) and add
sqlite-vec-linux-arm64. #225 (thanks @ilepn) - Fix claude plugin setup CLI commands in README. #311 (thanks @gi11es)
[1.1.1] - 2026-03-06
Fixes
- Reranker: truncate documents exceeding the 2048-token context window
instead of silently producing garbage scores. Long chunks (e.g. from
PDF ingestion) now get a fair ranking. - Nix: add python3 and cctools to build dependencies. #214 (thanks
@pcasaretto)
[1.1.0] - 2026-02-20
QMD now speaks in query documents — structured multi-line queries where every line is typed (lex:, vec:, hyde:), combining keyword precision with semantic recall. A single plain query still works exactly as before (it's treated as an implicit expand: and auto-expanded by the LLM). Lex now supports quoted phrases and negation ("C++ performance" -sports -athlete), making intent-aware disambiguation practical. The formal query grammar is documented in docs/SYNTAX.md.
The npm package now uses the standard #!/usr/bin/env node bin convention, replacing the custom bash wrapper. This fixes native module ABI mismatches when installed via bun and works on any platform with node >= 22 on PATH.
Changes
- Query document format: multi-line queries with typed sub-queries (
lex:,vec:,hyde:). Plain queries remain the default (expand:implicit, but not written inside the document). First sub-query gets 2× fusion weight — put your strongest signal first. Formal grammar indocs/SYNTAX.md. - Lex syntax: full BM25 operator support.
"exact phrase"for verbatim matching;-termand-"phrase"for exclusions. Essential for disambiguation when a term is overloaded across domains (e.g.performance -sports -athlete). expand:shortcut: send a single plain query (or start the document withexpand:on its only line) to auto-expand via the local LLM. Query documents themselves are limited tolex,vec, andhydelines.- MCP
querytool (renamed fromstructured_search): rewrote the tool description to fully teach AI agents the query document format, lex syntax, and combination strategy. Includes worked examples with intent-aware lex. - HTTP
/queryendpoint (renamed from/search;/searchkept as silent alias). collectionsarray filter: filter by multiple collections in a single query (collections: ["notes", "brain"]). Removed the singlecollectionstring param — array only.- Collection
include/exclude:includeByDefault: falsehides a collection from all queries unless explicitly named viacollections. CLI:qmd collection exclude <name>/qmd collection include <name>. - Collection
update-cmd: attach a shell command that runs before everyqmd update(e.g.git stash && git pull --rebase --ff-only && git stash pop). CLI:qmd collection update-cmd <name> '<cmd>'. qmd statustips: shows actionable tips when collections lack context descriptions or update commands.qmd collectionsubcommands:show,update-cmd,include,exclude. Bareqmd collectionnow prints help.- Packaging: replaced custom bash wrapper with standard
#!/usr/bin/env nodeshebang ondist/qmd.js. Fixes native module ABI mismatches when installed via bun, and works on any platform where node >= 22 is on PATH. - Removed MCP tools
search,vector_search,deep_search— all superseded byquery. - Removed
qmd context checkcommand. - CLI timing: each LLM step (expand, embed, rerank) prints elapsed time inline (
Expanding query... (4.2s)).
Fixes
qmd collection listshows[excluded]tag for collections withincludeByDefault: false.- Default searches now respect
includeByDefault— excluded collections are skipped unless explicitly named. - Fix main module detection when installed globally via npm/bun (symlink resolution).