Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
91 commits
Select commit Hold shift + click to select a range
8ead9a0
perf(mcp): cache server instructions instead of rebuilding per session
fxstein Aug 21, 2026
dc4732b
fix(store): scope FTS by materializing the match set, not a per-rowid…
fxstein Aug 22, 2026
d0550ed
fix(search): one search per leg, not one per leg per collection
shalom-t Sep 8, 2026
be03377
refactor(metadata): name the filter condition's target `field`
aaronccasanova Sep 15, 2026
21a0b28
feat(index): add mtime+size fast-path via file_sync_state
rikvanriel Sep 16, 2026
b27634a
fix(search): bound hydration at DB level via substr, not JS
rikvanriel Sep 16, 2026
a3791ce
feat(metadata): add text and type operators to the filter language
aaronccasanova Sep 13, 2026
866aec2
feat(metadata): add metadata discovery store query
aaronccasanova Sep 16, 2026
68f62bb
feat(metadata): add qmd collection metadata drill-down
aaronccasanova Sep 12, 2026
2a29f35
feat(metadata): surface metadata in collection list, show, and status
aaronccasanova Sep 15, 2026
3a583d0
feat(metadata): expose metadata discovery on SDK, MCP, and HTTP
aaronccasanova Sep 12, 2026
f32848b
docs(metadata): document metadata discovery
aaronccasanova Sep 12, 2026
3c2a2f8
fix(search,embed,index): stream large queries with iterate()
rikvanriel Sep 16, 2026
86966dd
refactor(metadata): trim discovery to what ships
aaronccasanova Sep 17, 2026
8d43a23
test: deactivate sampling fixtures through the store API
naveenspark Sep 23, 2026
2e8b8e6
fix: compare doctor vectors at their stored positions
naveenspark Sep 23, 2026
2e55927
feat(store): vector layout resolver with integer collection ids
brettdavies Sep 3, 2026
e1642b7
feat(store): migration step that copies legacy vectors into per-colle…
brettdavies Sep 3, 2026
a1f55ce
feat(store): partition the vector index by collection
brettdavies Sep 25, 2026
d44eeb9
feat(embed): copy vectors to collections that gain an embedded hash a…
brettdavies Sep 3, 2026
51eae40
feat(update): remove stale vector rows at the end of every qmd update
brettdavies Sep 3, 2026
23d5a6c
docs: describe the per-collection vector index layout
brettdavies Sep 3, 2026
530a6fd
refactor(store): share the partition row helpers across the embed, co…
brettdavies Sep 4, 2026
5c7367d
fix(review): apply the code-review findings on the partitioned vector…
brettdavies Sep 4, 2026
84b953e
fix(store): widen a vector search until it fills its limit
brettdavies Sep 25, 2026
5807d23
fix(update): copy vectors into gained collections before removing sta…
brettdavies Sep 25, 2026
cd4723d
fix(store): make collection rename and vector index drops atomic
brettdavies Sep 25, 2026
5303f50
test(migrations): cover a legacy-table repair that fails on a stamped…
brettdavies Sep 25, 2026
58300da
feat(cleanup): repack the vector table when its chunks are mostly holes
brettdavies Sep 3, 2026
42fb497
fix(query): drop duplicate query expansions before searching (#921)
ParkerRex Sep 26, 2026
9e1e8a5
test(mcp): cover the per-store server instructions cache
brettdavies Sep 29, 2026
58e8986
test(bin): pin argv and exit status through the in-process CLI import
brettdavies Sep 29, 2026
4efbe96
fix: bound doctor vector sampling memory
naveenspark Sep 23, 2026
8e7fb48
test(doctor): bound the whole vector sample check, not just its query
brettdavies Sep 29, 2026
6af0bc7
fix(store): load legacy adoption sample body by rowid instead of per-…
mjaverto Sep 26, 2026
f69e97e
test(store): bound legacy adoption by active paths alone
brettdavies Sep 29, 2026
9908c6f
test(store): cover #962's reindex fast path, size guard and body cap
brettdavies Sep 29, 2026
d4567fa
fix(store): trust a file_sync_state row only while its document stands
brettdavies Sep 29, 2026
92ac75a
fix(store): deactivate a document whose file grows past 10 MB
brettdavies Sep 29, 2026
c6bd2c0
test(store): cover #1000's expansion dedupe on fresh output and vsearch
brettdavies Sep 29, 2026
c20f733
test(search): cover #946's single search per leg over a collection scope
brettdavies Sep 29, 2026
d983290
fix(search): fuse one ranked list per search over all named collections
xidus90 Sep 28, 2026
176b9b9
test(search): guard vec-leg ranking against the order collections are…
brettdavies Sep 29, 2026
ba2a2db
fix(store): scope keyword retrieval at the index, not after it
fxstein Aug 21, 2026
150b868
test(store): cover #918's scoped keyword search across several collec…
brettdavies Sep 29, 2026
f4633f9
fix(store): apply BM25 scope to the full match set, not a top-N window
Mr-Beasley Sep 14, 2026
bcf8b18
test(search): cover #953's collection scope combined with a metadata …
brettdavies Sep 29, 2026
6589100
fix(mcp): rebuild cached instructions stamped ahead of the clock
brettdavies Sep 29, 2026
0d0c350
perf: avoid redundant CLI runtime spawn
ilepn Aug 24, 2026
87e7b3f
test(store): check that the SQLite heap limit binds in the budget wor…
brettdavies Sep 29, 2026
1f75eb7
test(store): pin the active and content clauses of the sync-row check
brettdavies Sep 29, 2026
03691f2
fix(store): re-extract stale metadata for files the fast path skips
brettdavies Sep 29, 2026
aaf1911
perf(store): read metadata currency in the sync-state query
brettdavies Sep 29, 2026
5660ad6
test(store): cover a touched file whose content is unchanged
brettdavies Sep 29, 2026
133bc4f
fix(cli): skip files over 10 MB when adding a collection
brettdavies Sep 29, 2026
e7f36e9
test(store): cover #1000's dedupe on older cache rows and keyword lines
brettdavies Sep 29, 2026
74fc668
fix(store): drop sync rows with their collection and move them on rename
brettdavies Sep 29, 2026
94840c8
fix(store): do not trust a sync row for a file modified during the pass
brettdavies Sep 29, 2026
e27f9f8
fix(search): break score ties by filepath when merging collections
brettdavies Sep 29, 2026
1c061b0
docs(changelog): credit #1000 by PR number
brettdavies Sep 29, 2026
e999c61
docs(changelog): record #815's cache and credit #923, #978 and #995
brettdavies Sep 30, 2026
f277cc9
docs(changelog): record #962's re-index fast path and body cap
brettdavies Sep 30, 2026
5c98af4
docs(changelog): record #946 and #1009 and credit #953 and #918
brettdavies Sep 30, 2026
05cd983
perf(search): scope keyword search over several collections in one query
brettdavies Sep 30, 2026
eb0b105
fix(store): keep the body cap in partitioned vector search
brettdavies Sep 30, 2026
abf570f
test(store): point #1000's cached-expansion test at the partitioned t…
brettdavies Sep 29, 2026
84d39e0
test(store): check #983's atomic rename keeps the fast path
brettdavies Sep 30, 2026
7c7997f
fix(search): restrict a filtered vector scan with a subquery, not a list
brettdavies Sep 29, 2026
0f89d9b
fix(search): break vector distance ties by filepath across partitions
brettdavies Sep 29, 2026
20fa36e
fix(update): leave the index unchanged when a collection root is missing
ParkerRex Sep 26, 2026
5ed47af
feat(cleanup): repack the partitioned vector table within each partition
brettdavies Sep 30, 2026
03bd6f1
test(cleanup): check a partitioned repack keeps vectors, rows and rol…
brettdavies Sep 30, 2026
2262ed5
test(update): check a missing collection root keeps its vectors
brettdavies Sep 30, 2026
09a18b9
docs(changelog): credit #983 and #990
brettdavies Sep 30, 2026
454668b
refactor(search): drop the unused per-collection merge helper
brettdavies Sep 30, 2026
e4db0ef
fix(cli): report a file over the size limit as too large, not unreadable
brettdavies Sep 30, 2026
f36855c
perf(store): write a collection scan in short transactions
brettdavies Sep 30, 2026
7ba4f84
Merge pull request #956 from aaronccasanova/feature/metadata-filter-f…
tobi Oct 2, 2026
26b703c
Merge pull request #951 from aaronccasanova/feature/metadata-discover…
tobi Oct 2, 2026
6acaa67
test(search): name the filter target `field` in the scoped keyword tests
brettdavies Oct 2, 2026
6298d7f
test(search): name the filter target `field` in the partitioned vecto…
brettdavies Oct 2, 2026
06605b2
perf(search): find filter-eligible collections per document, not per …
brettdavies Oct 2, 2026
2fe8781
fix(store): keep a short body with an embedded NUL whole under the bo…
brettdavies Oct 2, 2026
c1c646e
Merge pull request #1020 from brettdavies/stack/1-small-fixes
tobi Oct 6, 2026
30a63ca
Merge pull request #1021 from brettdavies/stack/2-reindex-memory
tobi Oct 6, 2026
c66ff13
Merge pull request #1022 from brettdavies/stack/3-expansion-dedupe
tobi Oct 6, 2026
94e3f45
Merge pull request #1023 from brettdavies/stack/4-keyword-scope
tobi Oct 6, 2026
573be91
Merge pull request #1024 from brettdavies/stack/5-vec-partition
tobi Oct 6, 2026
93d211f
Merge pull request #1025 from brettdavies/stack/6-vec-repack
tobi Oct 6, 2026
c18b22c
Merge remote-tracking branch 'upstream/main' into sync/upstream-2026-…
systemfsoftware-maker Oct 7, 2026
1d06ffe
chore: migrate upstream sync code to strict tsconfig
systemfsoftware-maker Oct 7, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@ texts/
*.md
!README.md
!CLAUDE.md
!DESIGN.md
!CHANGELOG.md
!skills/**/*.md
!finetune/*.md
Expand Down
115 changes: 114 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,17 +2,130 @@

## [Unreleased]

### Fixed

- `qmd doctor` selects vector sample identities before loading document bodies,
avoiding excessive SQLite memory use on large indexes with duplicate paths.
#978 (thanks @naveenspark)
- Vector diagnostics match passages by their saved character position, avoiding
false mismatches when earlier chunks change the sequence numbering.
#978 (thanks @naveenspark)

### Added

- Added Oxlint lint fence.
- Document metadata and metadata filtering. Markdown documents can opt into typed metadata through a namespaced frontmatter block (`qmd.metadata` with strings, numbers, booleans, or flat homogeneous arrays), and every search surface — CLI `search`/`vsearch`/`query` via `--filter <json>`, the SDK's `filter` option on `search()`/`searchLex()`/`searchVector()`, the MCP `query` tool, and HTTP `POST /query` and `/search` — accepts one shared recursive filter AST discriminated by `operator`: `and`/`or`/`not` logical groups, `eq`/`ne`/`gt`/`gte`/`lt`/`lte` comparisons, `in`/`nin`/`all` membership, and `exists` presence. Every returned result satisfies the filter (applied before RRF fusion and reranking); like collection filtering, highly selective filters remain best-effort for top-K completeness. Frontmatter stays ordinary searchable content — no chunking, embedding, snippet, or line-number changes — and documents without `qmd.metadata` behave exactly as before. JSON/SDK/MCP/HTTP results now include each document's indexed metadata, and `qmd status` reports how many documents still need metadata extraction (a normal `qmd update` backfills existing indexes).
- Document metadata and metadata filtering. Markdown documents can opt into typed metadata through a namespaced frontmatter block (`qmd.metadata` with strings, numbers, booleans, or flat homogeneous arrays), and every search surface — CLI `search`/`vsearch`/`query` via `--filter <json>`, the SDK's `filter` option on `search()`/`searchLex()`/`searchVector()`, the MCP `query` tool, and HTTP `POST /query` and `/search` — accepts one shared recursive filter AST discriminated by `operator`: `and`/`or`/`not` logical groups, `eq`/`ne`/`gt`/`gte`/`lt`/`lte` comparisons, `in`/`nin`/`all` membership, `contains`/`prefix`/`suffix` text matching, `type` for the stored type of a key's values, and `exists` presence, with an optional `caseInsensitive` flag on any condition whose value is a string (ASCII folding). Every returned result satisfies the filter (applied before RRF fusion and reranking); like collection filtering, highly selective filters remain best-effort for top-K completeness. Frontmatter stays ordinary searchable content — no chunking, embedding, snippet, or line-number changes — and documents without `qmd.metadata` behave exactly as before. JSON/SDK/MCP/HTTP results now include each document's indexed metadata, and `qmd status` reports how many documents still need metadata extraction (a normal `qmd update` backfills existing indexes).
- Metadata discovery. Filtering is only useful when the caller knows what to filter on, so every surface now reports the metadata keys, types, and value counts already in the index — read from the existing metadata tables, with no re-indexing. The CLI adds `qmd collection metadata [name...]` with `--filter <json>` to count only documents matching a filter (same AST as search) and `--match <json>` to select which metadata entries are reported (the same AST evaluated per entry, where a condition's `field` is the entry's `key` or `value`, so one grammar covers a single key, a family of keys, reverse lookup of a value, typed thresholds, and any `and`/`or`/`not` composition), plus `--key-limit`/`--key-offset`/`--all-keys` and `--value-limit`/`--value-offset`/`--all-values` to page both windows, and `--sort count|value` and `--min-count <n>` to order and trim values. Every windowed list reports its exact remainder, counts are documents rather than values, numbers report min/median/max, keys whose documents disagree on type are split per type with their own counts, and with `--filter` every coverage count is measured against the documents the filter admits. `qmd collection list` names each collection's top keys, `qmd collection show` details them with a value preview, and `qmd status` summarizes coverage. The SDK adds `listMetadata(options)` (the same options, spelled `keyLimit`, `valueLimit`, and so on, with `MetadataOptionError` for a value outside its domain and `MetadataBindingBudgetError` for a filter and match that together bind too many SQL parameters), the MCP server adds a `metadata` tool and lists each collection's most covered keys in `status`, and the HTTP server adds `POST /metadata`. `MetadataFilter` and `MetadataMatch` are both `MetadataPredicate<Condition>`, one recursive grammar over the conditions a document or a metadata entry admits.

### Fixed

- Filtered vector search binds candidate IDs as one JSON list, so a parser-valid metadata filter cannot exhaust Node's SQL variable limit during document lookup. Applies to both exact scans and the capped global fallback.
- Collection- and metadata-scoped BM25 search (`qmd search -c`, the lex leg of
`qmd query`, MCP, SDK `searchLex`) no longer returns false-empty or
incomplete results when stronger matches outside the scope fill the old
`limit * 10` candidate window (#922). The scope is now applied to the full
FTS5 match set, materialized once, so `search -c <collection>` is exact for
common terms; unscoped search keeps its early-terminating plan. A search over
several collections, such as the default ones, runs one keyword query instead
of one per collection, which made keyword search about 12x faster over 13
collections (thanks @brettdavies). Builds on the approach in #918 (thanks
@fxstein). #953 (thanks @Mr-Beasley)

- `qmd update` no longer deactivates every document in a collection whose root
folder is missing, for example an unmounted drive or an offline network share.
The folder globbed to an empty list, which read as "every file was deleted",
so search went empty and a `qmd cleanup` before the drive came back deleted
the rows for good. The collection is now reported as not found and its index
is left unchanged. Permission and I/O errors still fail with their original
cause (#989). #990 (thanks @ParkerRex)
- Embedding generation and legacy fingerprint adoption now tokenize documents
with the store-selected embedding model instead of the global default. This
keeps chunk boundaries aligned with the model that creates and verifies the
stored vectors without initializing an unrelated provider.
- `qmd doctor` no longer stalls or fills the temp directory on large indexes
when checking legacy (empty-fingerprint) embeddings. The adoption sample
query joined `content` and grouped by the document body, so SQLite
materialized the full body once per legacy chunk and per active path before
`LIMIT 1` discarded it, the same pattern as the doctor vector-sample check
(#978). It now picks the sample row through indexes and loads only that
row's body; the sampled chunk is unchanged (#994). #995 (thanks @mjaverto)

### Changed

- The MCP server caches its instructions per store for up to 60 seconds instead
of rebuilding them, with a full index-status scan, for every HTTP request.
Concurrent requests share one build and a failed build is not cached; index
changes from other processes show up in the instructions within a minute.
#815 (thanks @fxstein)
- Avoid a redundant runtime process on packaged CLI calls where the launcher is
already running under its selected Node or Bun runtime. #923 (thanks @ilepn)
- `qmd update` no longer reads files whose modification time and size match
the last pass (tracked in a new `file_sync_state` table), so re-indexing an
unchanged collection takes seconds. A cached entry counts only while its
document is still active at that path with that content, so a collection
removed and added back is re-indexed, while a renamed one keeps its entries;
files with missing or outdated metadata are still re-extracted. `qmd update`
and `qmd collection add` skip files over 10 MB with `FILE_TOO_LARGE`, and
`qmd update` deactivates a previously indexed file that becomes empty or is
over 10 MB, including one indexed by an earlier release. #962 (thanks
@rikvanriel)
- Search results carry at most the first 262,144 characters of each document
body (`qmd search --full`, MCP results, keyword and vector hits), which
bounds memory on indexes with very large documents. A match past that point
is still found, but its snippet, and the passage the reranker scores, come
from the start of the document. #962 (thanks @rikvanriel)
- `qmd update` and `qmd collection add` write a collection scan in short
transactions instead of committing every row on its own, so a first index
of 20,000 files takes about 5 s instead of about 3 minutes. #1021 (thanks
@brettdavies)

- `qmd query` and `qmd vsearch` now drop repeated query expansions before
searching. The expansion model can repeat a line, and the cache kept every
copy: one reported query ran 23 expansions where 9 were distinct, and each
copy ran its own search. Cached expansions are deduplicated when read, so
existing caches need no rebuild. Repeated lex lines used to count more than
once in the rank fusion, so result order can shift slightly. (#921)
#1000 (thanks @ParkerRex)
- Structured searches over several collections (`qmd query` with
`lex:`/`vec:`/`hyde:` lines, the MCP `query` tool, SDK `queries`) run one
search per line over the whole collection list, a keyword search for a
`lex:` line and a vector search for a `vec:` or `hyde:` line, and fuse one
ranked list per search. Rankings no longer depend on the order the
collections are named. #946 (thanks @shalom-t), #1009 (thanks @xidus90)

### Changed

- The vector index is partitioned by collection. Each embedded chunk is
stored once per collection that holds it, under an integer collection id,
so a collection-scoped `qmd query` or `qmd vsearch` scans only that
collection's vectors instead of pre-filtering the whole index, and a small
collection is never crowded out of the results by a large one. The first
command after upgrading converts the index in place and prints its
progress; the conversion resumes from where it stopped if interrupted, two
commands started at once share it, and a VACUUM at the end reclaims the
space of the old table. `qmd collection rename` no longer touches vectors,
`qmd update` removes, as it finishes, the vectors of content that no
indexed document holds any more (content that comes back later is embedded
again), and `qmd update` and `qmd embed` copy an already-embedded document
into a collection that gains it instead of embedding it again. A
metadata-filtered vector search returns the nearest documents the filter
admits however many chunks it admits, and a vector search fills its limit
even when one long document holds all of the nearest chunks. Downgrading
to an older qmd afterwards needs `qmd embed -f`, and a `qmd mcp` server
started before the upgrade must be restarted. #983 (thanks @brettdavies)

### Changed

- `qmd cleanup` repacks the vector table when fewer than 90% of its chunk
reads hold live rows. vec0 fills the first free slot of its newest chunk
and reclaims a chunk only once it is empty, so re-embeds and orphan removal
leave holes in older chunks that every brute-force scan still reads; a
794k-vector index at 36% occupancy scanned twice as slowly as a packed copy
of the same rows. The repack moves the live rows of each mostly-empty chunk
to the tail, one chunk per short transaction, so it never holds the write
lock for long and an interrupted run leaves a consistent table. The cleanup
output and `qmd cleanup --dry-run` report the chunk counts.
#937 (thanks @brettdavies)

## [2.8.3] - 2026-08-16

Expand Down
10 changes: 9 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@ qmd collection add . --name <n> # Create/index collection
qmd collection list # List all collections with details
qmd collection remove <name> # Remove a collection by name
qmd collection rename <old> <new> # Rename a collection
qmd collection metadata [name...] # Discover metadata keys, types, and value counts (--match, --filter, --key-limit, --value-limit)
qmd init # Create a project-local .qmd index
qmd ls [collection[/path]] # List collections or files in a collection
qmd context add [path] "text" # Add context for path (defaults to current dir)
Expand Down Expand Up @@ -47,9 +48,16 @@ qmd collection remove mynotes
# Rename a collection
qmd collection rename mynotes my-notes

# Show collection details
# Show collection details, including the top metadata keys
qmd collection show mynotes

# Discover metadata keys and values to filter on
qmd collection metadata mynotes
qmd collection metadata mynotes --match '{"field":"key","operator":"eq","value":"topics"}'
qmd collection metadata mynotes --match '{"field":"value","operator":"eq","value":"docs-team"}'
qmd collection metadata mynotes --match '{"field":"key","operator":"eq","value":"topics"}' --filter '{"field":"status","operator":"eq","value":"published"}'
qmd collection metadata mynotes --key-limit 50 --key-offset 50 # next page of keys

# Set or clear the pre-update hook (runs before re-indexing on `qmd update`)
qmd collection update-cmd mynotes 'git pull --ff-only'
qmd collection update-cmd mynotes # clear
Expand Down
12 changes: 12 additions & 0 deletions DESIGN.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
# QMD interface

QMD is a command-line search tool, library, and MCP server. This change has no
graphical surface. Existing command output and structured result formats are
the observed interface contracts; see README.md and test/cli.test.ts.

The doctor vector check samples stored chunks and compares freshly generated
embeddings with their saved vectors. Its sampling must preserve active-document,
model, and fingerprint filters without expanding full document bodies across
all candidate chunks. No visual or typography changes are part of this work.
The saved character position identifies the passage to re-embed; a stored
sequence number is the vector key and may differ from today's chunk ordering.
Loading
Loading