Skip to content

chore: sync with upstream tobi/qmd (04e4dbd..93d211f) - #3

Merged
ryanleecode merged 91 commits into
mainfrom
sync/upstream-2026-10-07
Oct 7, 2026
Merged

ryanleecode merged 91 commits into
mainfrom
sync/upstream-2026-10-07

Conversation

@systemfsoftware-maker

Copy link
Copy Markdown

Syncs the fork with upstream tobi/qmd. The head is a real merge commit, not a rebase or squash.

Upstream range

Conflicts

There was one, in src/store.ts → expandQuery, at the cache-format migration. I took upstream's behavior: both legacy-format branches now pass through uniqueExpansions(...). The code keeps the fork's strict-mode style (r["query"], r["type"], r["text"], rows[0]?.["query"]). No other files conflicted.

Strict-mode fixes (1d06ffe, separate commit after the merge)

These bring upstream code under the fork's @systemfsoftware/tsconfig base. Runtime behavior is unchanged.

  • noPropertyAccessFromIndexSignature: switched Record<string, unknown> reads to bracket access in parseCliMetadataOptions (src/cli/qmd.ts) and in the REST /metadata handler (src/mcp/server.ts).
  • exactOptionalPropertyTypes: added | undefined to every ListMetadataOptions field (src/metadata-store.ts) and to the reindexCollectionIn options (src/store.ts). The consumers already treat undefined as absent (??, === undefined, truthiness).
  • No flags were dropped or loosened. There is no @ts-ignore, @ts-expect-error or as any, and no tests were skipped.

Fork commits preserved

Local gates (Linux x64, node v24.19.0, bun 1.4.0)

Command Result
./node_modules/.bin/tsc -p tsconfig.build.json --noEmit (TypeScript 7.0.2) exit 0. Before 1d06ffe there were 36 errors (30 TS4111, 4 TS2379, 1 TS2345, 1 TS2375).
npx oxlint (1.78.0) Found 0 warnings and 0 errors. on 94 files, exit 0
CI=true npx vitest run --testTimeout 60000 test/ 57 files passed; 1532 passed, 82 skipped, 0 failed; Type Errors: none; exit 0
CI=true LD_LIBRARY_PATH=/usr/lib/x86_64-linux-gnu bun test --timeout 60000 --preload ./src/test-preload.ts test/ 1528 pass, 90 skip, 0 fail, 56 files; exit 0
npm run build exit 0
nix build . exit 0; result/bin/qmd --version → qmd 2.8.3

The first vitest run had 4 failures. All of them came from the local container, not from this merge. Neither failing test file is in the upstream diff.

  • test/release-context.test.ts (3 tests): jq wasn't on the subprocess PATH. I put jq from the repo's locked nixpkgs (nix build --inputs-from . nixpkgs#jq) on PATH.
  • test/llm.test.ts › canWriteLlamaDir … not writable: the suite ran as root with CAP_DAC_OVERRIDE. I re-ran under setpriv --bounding-set -dac_override,-dac_read_search. GitHub runners don't run tests as root.

The vitest and bun numbers in the table come from the full re-run with those two fixes in place.

Notes:

  • Locally, bun install --frozen-lockfile fails at the better-sqlite3 install script because node-gyp is missing on this host. I installed with npm install instead. bun.lock is unchanged from main.
  • qmd --version from the nix build also prints *** stack smashing detected *** on exit. A nix build of fork main (2c8b6d2) prints the same thing, so it was there before this PR.

fxstein and others added 30 commits August 22, 2026 00:30
createMcpServer() rebuilds the server instructions on every call, and
the Streamable HTTP transport creates one McpServer per MCP session -
so every initialize re-runs the full index-status aggregation behind
buildInstructions(). On a large index that recomputation dominates the
handshake, and agent reconnect bursts stack the status scans (observed
as multi-second to tens-of-seconds initializes on an 81k-row index).

Cache the built instructions per store with a 60s TTL: concurrent
initializes share one in-flight build (a reconnect storm now costs one
status scan instead of N), failed builds are dropped so the next
initialize retries, and index updates from other processes still
surface within a minute without restarting the server. The status tool
is untouched and always live.

Co-Authored-By: Claude <noreply@anthropic.com>
(cherry picked from commit dcb58d3)
… prefilter

The rowid IN (...) prefilter defeats FTS5 early termination: SQLite drives the
query from the collection's whole doclist and probes the FTS5 index once per
candidate rowid (EXPLAIN: VIRTUAL TABLE INDEX 0:=M3), about 55ms per in-scope
document. On the live index that is 0.02s unscoped versus 21s scoped to a large
collection, and 507s for one pass over 19 collections, which on a single-threaded
MCP server is a multi-minute outage for every client.

Keep the MATCH global and unbounded, MATERIALIZE the complete match set once
(corpus-bounded), and apply the collection filter, ORDER BY and LIMIT in the
outer query. Correct (the inner set is complete, so the outer filter cannot
truncate it) and fast (natural FTS5 operation, EXPLAIN VIRTUAL TABLE INDEX 0:M3).
MATERIALIZED is load-bearing: without it the planner may flatten the CTE and
refold the collection predicate into the MATCH. Unscoped path unchanged; result
rows identical to the prefilter.

(cherry picked from commit 0c7a0b8)
structuredSearch loops the collection list and pushes each collection's
results as its own RRF ranked list, so a scope of N collections builds N
lists per search leg. RRF ranks with no relevance floor, so every
collection's best-of-a-bad-lot enters fusion as a rank 1 beside the real
answer. The 2x weight given to rankedLists[0] compounds it: it is meant for a
caller who ordered `searches` by importance, but the fan-out makes list 0
whichever collection happened to sort first.

The visible symptom is that ranking depends on argument order. Over a
21-collection scope and a 14-question gold set, reversing the collections
array moved rank 1 for 13 of 14 queries.

searchFTS and searchVec already accept `string | readonly string[]`, so this
just passes the scope straight through instead of looping it. On that same
set: P@1 0.000 -> 0.500, recall@5 0.643 -> 0.786, MRR 0.279 -> 0.607, and
reversing the array no longer changes anything. Median query time 0.499s ->
0.060s, since it is now one FTS and one vector query rather than 21 of each
-- searchVec in particular was pulling the same global top-k every time and
only then filtering by collection.

Single-collection scopes are unchanged, as they must be: the loop ran exactly
once there already.

(cherry picked from commit aa67e26)
A filter condition is a predicate over one field of the record under
evaluation: `field` names it, `operator` says how to compare, `value`
is the operand. For a document, the fields are its metadata keys, and
that special case is what the property was named after.

Rename the property from `key` to `field` so the grammar describes
itself without reference to what it happens to be applied to. Nothing
else in the AST changes: `operator`, `value`, `operands`, and `operand`
keep their names, and the parser rejects the old property the same way
it rejects any unknown one, with the JSON path of the failing node.

The rename lands before metadata filtering (tobi#910) ships, so no released
version accepts `key`.

Updates the README filter table and semantics, the skill, the MCP
`query` tool description, the CLI help and error examples, the
changelog entry, and every test that builds a condition.

Assisted-by: Claude Fable 5.1 via Pi
When a collection has no changes, reindexing still
reads every file to hash it. With thousands of files
this costs minutes of IO and hashing even when nothing
changed.

Add file_sync_state table tracking mtime_ms, size,
and content hash per (collection, path). On reindex,
stat the file first and skip the read entirely when
mtime and size match the cache. If mtime changed but
hash is identical only the cached mtime is updated.

Skip files over 10MB and empty files, and remove the
cache entry on orphan deletion.

Second reindex with no changes performs no file reads,
taking update from minutes to seconds when the
collection is unchanged.

(cherry picked from commit 3fcf230)
Previously, searchFTS and searchVec loaded the whole
document body as content.doc as body. On collections
with large transcripts this caused 20 candidates times 6
expansions times 100KB = 12MB of string copies through
better-sqlite3 into V8, observed as 4.2GB heap.

Now the bodies are bounded at the SQL level using
substr(content.doc, 1, 262144) as body. This limits each
body to 256KB. With 40 winners the max is 10MB vs
unbounded before. Rerank input stays 900 tokens, about
3.6KB, per chunk. The long term fix fetches only the
winning chunk via substr(doc, pos, len) for the 40 RRF
winners, but this commit prevents the OOM first.

(cherry picked from commit a75b659)
Extend the filter AST with three text operators, a type test, and an
optional case-folding flag, so a filter can reach values by substring
and reach one side of a key whose documents disagree on type.

- `contains`, `prefix`, and `suffix` take a non-empty string operand and
  match string values only. `contains` compiles to `instr()`. `prefix`
  and `suffix` compare UTF-8 bytes with the operand's byte length bound
  from JavaScript, because SQLite's text `length()` and `substr()` stop
  at an embedded NUL and the stored string domain admits one.
- `type` takes `string`, `number`, or `boolean` and matches the stored
  type of a key's values.
- `caseInsensitive: true` is accepted on any condition whose operand is
  a string or an array of strings. Both sides fold ASCII letters
  (SQLite's `lower()` and the same range in JavaScript), so the two
  agree exactly. Non-ASCII letters compare exactly.

Each compiled condition now names how its correlated subquery reaches a
document's rows. The covering indexes are (key, value, document_id), so
an exact equality (`eq`, `in`, `all`) is one probe by value. Every other
condition (ordered comparisons, `ne` and `nin` presence, `exists`,
`type`, the text operators, and folded equalities, since `lower()` is
opaque to the index) seeks the document's own rows through the primary
key: a unary `+` on the key term hides it from the planner, which
otherwise walks the key's whole value range once per document when
`sqlite_stat1` is absent. On 4,000 documents that plan made one
`priority > 10` count 376 ms and a `prefix` count 946 ms; both are
under 5 ms with the seek, on Node and Bun, with or without statistics.
A test asserts the plan for every operator in both states.

Validation reports the failing JSON path as before. The README filter
table, the MCP `query` tool description, the skill, and the changelog
entry name the new operators.

Bind vector candidate IDs as one JSON list during document lookup. A valid
near-ceiling filter otherwise overflows Node's SQL variable limit when
combined with candidate IDs, in both exact scans and the global fallback.
A model-free regression crosses the 20,000-vector boundary and verifies
that nonmatching documents sharing a content hash remain excluded.

The plan guard accepts SQLite 3.43's USING INDEX label for the same point
seek, retaining the check that rejects broad per-document key scans.

Assisted-by: Claude Fable 5.1 via Pi
Report keys, typed values, document coverage, collection provenance, and
numeric ranges over the same extraction gate as filtered search. `filter`
selects documents and `match` selects entries. Both windows run in SQL with
exact totals and remainders, including past-end pages.

Discovery compiles filters to uncorrelated document sets. Search retains
candidate-local seeks. Entry prefix/suffix comparisons return false, not
NULL, for empty strings so negation partitions the accepted value domain.

Read the report in one deferred transaction. Materialize key ranking once
in the type-statistics statement and reuse its selected names thereafter.
Read contributing collections with those statistics. Group value counts
once for both the distinct total and the value window, retaining one total
carrier per type even past the end. Preserve BINARY key and value ordering.

Compute even medians without overflow or loss of adjacent subnormals.
Validate safe-integer options and cast the production window operands to
INTEGER before adding. Check the widest planned statement against the
30,000-parameter application budget before preparing discovery queries.

Provide a lean key overview and a batched per-collection overview. Ordinal
zero counts each document/key once regardless of array length, without
reading values or maintaining another index or cache.

Tests cover aggregation, gates, filters, matches, ordering, both windows,
independent boolean expectations, live-writer snapshots, lifecycle changes,
actual endpoint SQL, binding limits, and structural query-work guards with
and without statistics. The independent 14,276-case predicate oracle and
120 aggregate comparisons pass on both drivers after these changes.

Sort contributing collection names by UTF-8 bytes after reading their
aggregate. This preserves SQLite BINARY order without requiring SQLite
3.44's aggregate ORDER BY syntax. The real discovery query and the filter,
discovery, and vector suites also pass on SQLite 3.43.2.

Assisted-by: Claude Fable 5.1 via Pi
Adds the CLI surface for metadata discovery so an agent can learn what
to filter on before it queries:

- `qmd collection metadata [name...]` prints one block per key: a
  header with type, coverage, and distinct count, then a value window.
  Strings list vertically with right-aligned counts, numbers show
  min/median/max plus the values when they fit, booleans print
  true/false counts. Keys whose documents disagree on type split per
  type, with contributing collections in the multi-collection view.
- `--filter <json>` counts only matching documents. The output then
  opens with a `filter:` line stating how many documents pass, and
  every key's coverage is measured against that population, so the
  numbers an agent reads are the numbers a filtered search would see.
- `--match <json>` selects which metadata entries are reported, with
  the same AST as `--filter` evaluated against each entry (a
  condition's `field` is the entry's `key` or `value`). Both JSON flags
  share one parser and report malformed JSON and invalid nodes the same
  way.
- Two windows, named the same on every surface: `--key-limit <n>`
  (default 50) with `--key-offset <n>` and `--all-keys`, and
  `--value-limit <n>` (default 10) with `--value-offset <n>` and
  `--all-values`. Every windowed list ends with the exact remainder and
  the flags that reach it, and an offset past the last key says how
  many keys there are. `-n` and `--all` are the single-window search
  flags and exit with a pointer at the discovery ones. `--sort
  count|value` orders values, `--min-count` drops the tail. Omitting
  the name covers the default collections, as an unscoped search does.
- src/metadata-format.ts holds the renderer so the MCP tool can print
  the same shape. It takes the CLI's color palette or none. The empty
  states render there too, so the filter line survives a filter whose
  documents declare no metadata. A string value prints bare only when
  the bare form is unambiguous. One that is empty, padded, contains a
  quote, a backslash, a control character, a line separator, or a
  delimiter of the compact `value (count), ...` list a type split
  prints, or that reads as a JSON number, boolean, or null, or as the
  list's remainder tail, prints as a JSON string with every control
  character escaped. One rule for both layouts, so a value never prints
  two ways, and `"a (1), b" (1)` cannot be read as two values. Ordinary
  values are unchanged, so the common case costs nothing.
- A filter and match that together exceed the SQL binding budget exit
  with the store's message.
- The pending-extraction warning prints on stderr whenever documents
  are gated out, since discovery always applies the gate.

The renderer suite reads the compact list back the way a reader must,
honoring quotes, and checks every item round-trips.

Covers the keys view, both windows with paging and their footers, the
filter line and denominator (including over an empty result), reverse
lookup, a composed match across both fields, type selection, sort and
min-count, default scope, empty matches, and exit codes for unknown
collections and invalid flags, filters, and matches. The renderer has
its own suite for string identity (round trip through JSON.parse, no
raw control characters) and the empty states.

Numeric flags reject unsafe integers as usage errors before opening the
index, preserving the original input in the message. Unsupported format
flags fail explicitly with a pointer to the structured surfaces, rather
than silently returning text to a caller expecting JSON.

Assisted-by: Claude Fable 5.1 via Pi
Makes metadata visible from the commands a user or agent already runs
first, so the drill-down is discoverable without knowing it exists:

- `qmd collection list` adds a `Metadata:` line per collection naming
  the top five keys by coverage and counting the rest (`+N more`).
  Omitted when the collection has no metadata, like `Ignore:`.
- `qmd collection show` adds a `Metadata` section: a coverage line
  (keys, documents with metadata, pending extraction), then the top
  five keys as aligned rows with type, coverage, distinct count, and a
  short value preview (top three strings, numeric range and median,
  boolean counts, or a pointer when types disagree), then a pointer at
  `qmd collection metadata <name>` when keys were left out.
- `qmd status` adds one summary line under Documents, placed with the
  existing pending-extraction line so the two read together.
- Every view asks the store for exactly what it prints. `show` requests
  a five-key, three-value window and reads the total and remainder from
  the result, `list` reads the first five key names and counts the
  rest, and `status` counts keys without listing them, so none of them
  grows with the vocabulary.
- src/metadata-store.ts gains the light queries these views need:
  `countDocumentsWithMetadata()` and an optional collection scope on
  `countDocumentsPendingMetadata()`. The pending-extraction warning
  that `collection metadata` and the filtered search commands print
  takes that scope, so it counts the collections the command reads
  and not the whole index. `status` keeps the index-wide count.

Collection list batches all key overviews in one query, including the
vocabulary totals, instead of counting and ranking separately for each
collection. CLI status counts metadata-bearing documents with existence
probes and counts keys from ordinal-zero rows, avoiding array expansion.

Type-conflict drill-down hints escape apostrophes in the shell-quoted JSON
operand. A POSIX shell round-trip regression ensures metadata keys remain
literal data, including command-substitution syntax.

Assisted-by: Claude Fable 5.1 via Pi
Ships discovery on the remaining surfaces with the same options object
and result shape the CLI uses:

- SDK: `listMetadata(options?)` on QMDStore under Collection
  Management, next to listCollections(). `match` and `filter` are
  validated at the runtime boundary like the search methods, and the
  window options by the store, so a caller sees MetadataFilterError,
  MetadataOptionError, or MetadataBindingBudgetError with the same
  message every surface prints. Discovery types, MetadataMatch, the
  error classes, the default limits, the binding budget, and
  parseMetadataMatch() are exported from the package root, and
  MetadataValueType now lives in metadata.ts beside the scalar types
  it describes.
- MCP `metadata` tool: flat and read-only, taking `collections`,
  `match`, `filter` (both validated like the query tool's filter),
  `keyLimit`, `keyOffset`, `valueLimit`, `valueOffset`, `sort`, and
  `minCount`. The description teaches the one grammar over two record
  types (a document for `filter`, a metadata entry for `match`), a
  table of matches and the question each answers, how to read
  `totalKeys`, `remainingKeys`, `remainingValues`, and `range`, and how
  to page. Text content renders the CLI shape through the shared
  formatter, empty states included so the filter line survives them,
  naming the option that reaches each remainder; structuredContent is
  the ListMetadataResult. The schema's integers are safe integers, and
  a filter and match over the binding budget return a tool error.
- MCP `status`: CollectionInfo gains `metadataKeyCount` and a windowed
  `metadataKeys` (the ten most covered keys with types), mirrored in
  StatusResult and listed in the text summary with `+N more` and a
  pointer at the metadata tool, so the call agents make first reveals
  that metadata exists without growing with the vocabulary.
- HTTP `POST /metadata`: same body as the tool. 400 on a non-object
  body, a non-object or invalid match or filter, a bad sort, a
  non-array `collections`, a non-number window option, a number
  outside its domain, or a filter and match over the binding budget
  (the store's message, verbatim). Returns the ListMetadataResult as
  its own envelope. Logged like POST /query.

Updates the exact tool-list assertions in test/mcp.test.ts and covers
each surface in test/metadata-surfaces.test.ts, including paging, the
400 paths, invalid matches and options, unsafe integers, the binding
budget on all three surfaces, the filter line over an empty tool
result, and status structured content.

Full status batches all collection metadata overviews in one query. An
internal getStatusSummary path returns only the index facts initialization
uses, without pretending uncomputed metadata counts are zero. MCP server
creation uses that path, including each stateless HTTP request. The public
SDK getStatus result and MCP status tool keep their metadata summaries.

Regression tests exercise empty-string negation through SDK, HTTP, and MCP.
A real MCP tools/list request succeeds with the metadata value table hidden
in the isolated fixture, proving initialization does not read it. Restoring
the eager status call makes that test fail.

Assisted-by: Claude Fable 5.1 via Pi
- README: a "Metadata Discovery" subsection after "Metadata Filtering"
  with the mental model (same metadata, different unit: filter narrows
  documents, match narrows the metadata itself, one predicate grammar
  over both), a table of matches and the question each answers, a
  wide-to-narrow walkthrough ending in a validated query, exact CLI
  output for string, number, boolean, threshold, filtered, and
  type-conflict keys, the two windows, how they page, and what they do
  and do not bound, the reading rules (including when a string prints
  quoted and that a result is one snapshot), and the SDK method with
  its predicate, option, and error types, MCP tool, and HTTP route. Also the new subcommand in the collection
  commands block, the `metadata` tool parameters, and the
  `POST /metadata` endpoint.
- CHANGELOG: one entry under Unreleased / Added.
- skills/qmd/SKILL.md: a "Discover metadata before filtering" section
  written for an agent deciding what to filter on: the match shapes
  worth knowing, how to read the header and the filter line, the
  window footers and how to page past them, numeric ranges, and type
  splits.
- CLAUDE.md: `qmd collection metadata` in the command table and the
  Collection Management block.

Examples use the status, topics, priority, and reviewed keys the
filtering docs already establish, so the two sections read as one.

Qualify the windows as bounds on key/value lists, not query work or
collection provenance. Document the current-data cost of status summaries
and their exclusion from MCP initialization. Name the filter target field
consistently and explain explicit rejection of unsupported CLI formats.

Assisted-by: Claude Fable 5.1 via Pi
Large result sets previously used all() which
materialized full arrays in V8 heap. On an index with
13k files and 89k chunk vectors this means tens of
thousands of rows (active paths, sync-state entries,
pending-embedding docs, candidate chunk vectors) in a
single array, which caused OOM.

Now they stream via iterate(): sync-state reads,
pending-embedding scans, active-path listings,
glob/fuzzy scans, hash_seq eligibility (with early
exit past the cap, falling back to ANN), and
embedding-body fetches (substr-bound, with a
bytes-only path).

Small queries remain with all() as they are bounded.

(cherry picked from commit 3c37c68)
A holistic pass over the feature after the correctness and performance
audits, asking whether each piece is as small and intentional as it can
be. No behavior changes. The full matrix, the close-out packet's semantic
checks (adapted for the removed helpers), and the SQLite 3.43.2 path all
pass on Bun and Node.

Code:

- Remove `listMetadataKeys()` and `getMetadataOverview()` from
  metadata-store.ts. Neither had a production caller: every status view
  reads `listMetadataCollectionSummaries()`. Their doc comment described
  the pre-batching design. `queryMetadataOverviews()` loses the SQL
  fragment parameter that only they varied.
- Stop exporting `DEFAULT_METADATA_KEY_LIMIT`,
  `DEFAULT_METADATA_VALUE_LIMIT`, and `METADATA_SQL_BINDING_BUDGET` from
  the SDK. The defaults are in the option docs and the budget is an
  internal safety valve. The error classes stay exported.
- Un-export `formatMetadataKeySummary()`, which only its own module calls.
- Measure text-operator byte lengths with `Buffer.byteLength`, the idiom
  the rest of the file already uses, instead of a second `TextEncoder`.
- Merge the two stacked comments on `warnPendingMetadata()` into one.
- Empty-result messages on the CLI and MCP tool say "see which keys
  exist" rather than "see every key", since the default window is 50.
- Note why vitest's typecheck ignores source errors: it would otherwise
  report pre-existing errors under tools/, and tsc covers src/.

Tests:

- Fold the `listMetadataKeys` assertions into the batched-overview tests.
- In the "rank once, group once" test, keep the statement-count guard
  and the materialization guard (the ranking CTE plans as MATERIALIZE,
  never CO-ROUTINE, with and without statistics). One ranking statement
  does not mean one evaluation of the ranking, and the guard fails the
  moment MATERIALIZED is dropped. Remove the EXISTS and LIST SUBQUERY
  assertions, which pin the filter compiler's plan and belong to its
  own suite.

Docs, sized to sit beside their neighbors:

- README: list the `metadata` tool under "Tools exposed" (it was
  missing) and move its parameter rows out of the middle of
  `multi_get`'s. Lead the Discovery section with the one-sentence model,
  keep one example per idea, render the window flags in the README's own
  flag-list style, shorten the reading rules, and drop the internals
  (binding budget constant, status cost, MCP initialization) that belong
  in the PR.
- CHANGELOG: trim the discovery entry to user-visible behavior.
- MCP `metadata` tool description: cut to the mental model, four match
  shapes, and how to read and page the result. It is now shorter than
  the `query` tool's.
- skills/qmd/SKILL.md: restructure the two metadata sections in the
  skill's own style (wrapped prose, short paragraphs, bulleted reading
  rules, five examples instead of nine). Add one sentence of judgment in
  the register the skill already uses for `-c`: add a metadata filter
  when it expresses a constraint the user intends, and check coverage
  first, because a value condition only sees documents that declare its
  key. A matching "Do not filter blind" pitfall. Bump the skill version
  to 2.3.0.

Assisted-by: Claude Fable 5.1 via Pi
Add src/vec-layout.ts as the one place that knows which sqlite-vec table is live and what it is called. It resolves `none`, the legacy `vectors_vec` (hash_seq or hash keyed) or the partitioned `vectors_by_collection`, with the vec0 shadow-table names and the stored dimension, and it owns the two plain tables the partitioned layout needs: `vector_collection_ids` (a stable integer id per collection name, so a rename is one row) and `vector_rows` (vec0 rowid to hash, seq, collection_id, NOT NULL and UNIQUE per triple). Both are created at open next to `content_vectors`.

The resolver also carries the collection id allocate, resolve, rename and delete helpers and the JSON rowid list used to bind a scoped search's rowids through `json_each` in one parameter. Nothing reads the partitioned table yet; the legacy table stays live.

(cherry picked from commit 3cdb92d)
…ction partitions

Add src/store-migrations.ts with the PRAGMA user_version ladder: version 1 installs the FTS sync triggers as before, and version 2 copies every row of the legacy `vectors_vec` table into the partitioned layout, one vec0 row per (chunk, active collection), keyed by the integer collection id. Nothing calls the step at open yet; that lands with the readers and writers that use the new layout.

The copy reads the legacy table's chunk blobs directly (validity bitmap, rowid list, packed float32 vectors) at one chunk per IMMEDIATE transaction and keeps the last copied chunk_id in store_config. A process killed between chunks resumes from that cursor, and two processes opening the same database take turns, because each reads the cursor inside its own write transaction and stops as soon as the other stamps the version. A verification pass then copies by key any chunk that content_vectors and an active document claim but no partition holds (a pre-upgrade repack can move a row into a chunk the walk already passed), content_vectors rows with no vector anywhere are deleted so the pending detector embeds them again, and the version stamp and the DROP of the legacy table share one transaction. A legacy table keyed by hash alone is dropped without a copy, as before. Without sqlite-vec the step defers and the version stays at 1.

vec0 checks the bound type of its partition key and rowid rather than applying affinity, and better-sqlite3 binds every JS number as a double, so those parameters are bound as bigint through `vecInteger`. The anti-join that finds chunks missing from a partition lives in vec-layout.ts because the embed path will need the same query.

(cherry picked from commit 7d17354)
Vector search, embedding writes, cleanup and the collection operations now use the layout vec-layout.ts describes: one vec0 table partitioned by an integer collection id with one row per (chunk, active collection), a `vector_rows` mapping from vec0 rowid to (hash, seq, collection_id), and the version 2 migration wired into store open so every process copies the legacy `vectors_vec` table before it can read vectors.

A collection-scoped `searchVec` is one KNN per collection id with `k = limit * 3` (capped at sqlite-vec's 4096), merged by distance, so a small collection costs its own rows instead of a pre-filter over the whole index and is never crowded out by a large one (tobi#775, tobi#791, tobi#803). The unscoped search is one KNN over all partitions with the same k. Step two joins the returned rowids through `json_each` in a single bound parameter, because thirteen partitions at the k cap exceed SQLite's separate-parameter limit. A scope naming an unknown collection skips it; an all-unknown scope is empty. The exact-scan, capped-ANN and post-filter paths are gone, as is the per-collection recursion; `mergeSearchResultsByScore` stays for FTS.

A metadata filter restricts each KNN to the vector rows of documents it admits, passed as `rowid IN (SELECT value FROM json_each(?))`, which vec0 0.1.9 applies inside the scan alongside the partition key. The result is an exact top-k of the eligible rows at any eligible-set size, so the 20,000-row exact-scan cutoff and the global over-fetch above it go away, and with them the false empty a large eligible set got when more than the over-fetch of closer rows were ineligible. The filter is re-applied on the document join, since one hash can back a matching and a non-matching document in the same collection. On the live index the restriction costs nothing measurable: a stars KNN (794k rows) restricted to half its rows took 1.04 s against 1.16 s unrestricted.

`insertEmbedding` writes the content_vectors row and one partition row per active collection of the hash in one transaction, replacing an existing row under the same rowid since vec0 ignores OR REPLACE. `cleanupOrphanedVectors` removes partition rows whose (hash, collection) has no active document, so a stale row in one collection cannot steal k slots there while the hash stays live elsewhere. `clearAllEmbeddings(collection)` clears the collection's partition rows for hashes exclusive to it and keeps shared ones in every partition, `removeCollection` deletes the partition and its id, `renameCollection` updates the id's name, and `removeIncompleteEmbeddings` drops every partition row of the hash. Doctor's stored-vector check and the legacy-fingerprint adoption sample resolve the vector through `vector_rows`. `ensureVecTable` keeps the dimension-mismatch check on the partitioned table and runs the copy itself if it meets a legacy table, which covers an older build having recreated one on an already-migrated index.

The metadata search suite gains the over-20,000-eligible case, which returns 0 of 5 documents on the unpartitioned code, and a same-collection shared-hash case. The store-level tests move to `insertEmbedding` fixtures, the MCP fixture seeds the legacy table and runs the migration, and the collection-filter suite gains the scoped-search semantics tests (shared hashes named per collection, inactive rows, unknown and empty scopes, a limit above the k cap, one-row-per-file collapse, per-partition k, rename, remove and scoped clear). No reader needs a retry for the moment the legacy table is dropped: the migration runs inside `createStore`, so no process using this code reads vectors before its own open has finished the copy.

(cherry picked from commit c8cbc4f)
…nd write each batch in one transaction

`generateEmbeddings` starts with a copy pass: every active (hash, collection) pair whose chunks are in content_vectors but missing from that collection's partition gets its rows copied from whichever partition holds them, so a document that appears in a second collection is searchable there after the next embed run without another model call. A chunk no partition holds has its hash's rows removed instead, which puts the hash back in front of the pending detector; that closes the case where content_vectors claims a chunk the index lost. `EmbedResult.chunksCopied` reports the copies and `qmd embed` prints them.

Each 32-chunk batch is written inside one IMMEDIATE transaction, with the success and failure bookkeeping applied after the commit. A chunk writes several rows across three tables, and a kill or a write error mid-batch now rolls the whole batch back instead of leaving content_vectors ahead of the index; the existing per-chunk fallback then retries the batch's chunks one by one.

(cherry picked from commit 48e564f)
A changed file keeps its document row and gets a new hash, which strands the old hash's rows in the collection's vector partition. The partition filter cannot see `documents.active`, so on a churning collection those rows take k slots from every scoped search until something removes them; with enough of them a scoped query comes back short or empty. `qmd update` now runs the orphan-vector cleanup as its last step and reports the rows it removed, and the hint that used to suggest `qmd cleanup` for orphaned chunks goes away with the orphans.

The regression test re-indexes a file through four rewrites with a fake embedder that places the stale bodies nearer the query than the live documents, shows the scoped search returning nothing, and checks that the cleanup restores the full count. Two tests cover the legacy-fingerprint adoption sample, which resolves its nearest stored vector through the partition mapping.

(cherry picked from commit 17589bc)
CHANGELOG gets the Unreleased entry for the index format change: what the partitioned layout does for scoped queries, that the first command after upgrading converts the index in place with progress, resumption and shared work between concurrent commands, what rename, update and embed now do with vectors, and that downgrading needs `qmd embed -f` and a running `qmd mcp` server a restart. The README schema block lists `vector_collection_ids`, `vector_rows` and `vectors_by_collection` in place of `vectors_vec`.

(cherry picked from commit c9cfc20)
…py and clear paths

`activeCollectionsOfHash`, `deletePartitionRowsOfHash` and a prepared-statement `storedEmbeddingLookup` in vec-layout.ts replace three copies of the same SQL in store.ts, the migration reuses the exported `parseDimensions` instead of its own regex, and the unused `STORE_SCHEMA_VERSION` alias goes. The collection-scoped `clearAllEmbeddings` now runs inside one IMMEDIATE transaction, so a force re-embed of one collection is a single commit and a crash cannot leave the mapping and the vec0 table out of step.

(cherry picked from commit 459ef0c)
… layout

Findings from the ce-code-review run 20260904-161233-cdf84496, all applied:

- The migration's verification pass, the vectorless content_vectors delete, the cursor delete, the version stamp and the DROP of the legacy table now run in one IMMEDIATE transaction, so a writer on an older build cannot land a row in the legacy table between the delete and the drop and lose its only vector. The partitioned table is created with IF NOT EXISTS inside a double-checked transaction, so two first openers no longer race on CREATE VIRTUAL TABLE. `runStoreMigrations` also runs the copy when a legacy table exists on an already-stamped store, as the module's own comment claimed; when that repair fails on a dimension mismatch it warns and defers instead of blocking every open.
- `qmd update`, `qmd collection add` and the library's `update()` now run the copy pass, so a document that appears in a second collection is searchable there right after indexing instead of after the next embed run, and the library's `update()` also removes stale partition rows and reports both counts (`staleVectorsRemoved`, `vectorsCopied`).
- `removeCollection` runs as one IMMEDIATE transaction, so an interrupted removal cannot leave the vec0 rows, the mapping and the collection id out of step.
- `copyVectorsToNewCollections` commits in slices of 1,000 rows, so a collection that gains many already-embedded documents no longer holds the write lock for the whole copy; a hash queued for re-embedding in one slice is skipped by later ones.

Tests cover each change: the flip holds one write lock, a row written after the copy phase still lands in a partition, a second opener finishes a migration another opener started, a legacy table on a stamped store is migrated at open, `qmd update` and the SDK `update()` copy vectors into a collection that gained a hash, `removeCollection` rolls back as a unit, and a 2,500-row copy commits in batches with straddling hashes queued correctly.

(cherry picked from commit 1921727)
A vector search asked each partition for `limit * 3` chunks and collapsed them to one row per file. When one long document held the nearest chunks, those slots all went to it and the search returned fewer documents than it asked for: one 20-chunk document and one single-chunk document gave 1 of 2 results at limit 2 and at limit 5, scoped or unscoped. The same happened when stale rows of a changed file sat nearer the query than any live document. The case was reported with a synthetic reproducer in a review comment on tobi#936.

Each scan target now resolves its own matches to documents. While they collapse into fewer than `limit` documents and the target holds rows beyond them, the KNN runs again with k doubled, up to sqlite-vec's 4096 cap. Each target returns its own nearest `limit` documents (or all it holds), so merging the targets by distance gives the exact nearest `limit` across the scope. The first KNN keeps `limit * 3`: on the live index, stars' top 60 chunks covered 46 to 60 distinct documents across ten sampled queries, so the second round is a guard for skewed data rather than a cost on typical queries.

Step two is prepared once per search and leaves document bodies out: one long document can hold every match of a widened KNN, so only the final results load their body. The metadata predicate (current, error-free extraction plus the compiled filter) is built in one place for the eligible-row query and the document join.

The store suite gains the one-long-document case, scoped and unscoped, and the metadata suite gains the reporter's two filtered cases: a long document with an ineligible copy of its content starving a second eligible document, and the best chunk of a 20- and a 450-chunk document surviving deduplication. All four returned one document before this change. The stale-row test's pre-cleanup search now answers instead of coming back empty.

(cherry picked from commit d8ae76e)
…le rows

`qmd update` and the library's `update()` removed stale partition rows first and ran the copy pass second. When a file moved from one collection to another between two updates, the cleanup deleted the old collection's partition rows for the hash, and with them its content_vectors rows, before the copy could read them; the copy then found no source, queued the hash, and the moved document needed a fresh embed. The copy now runs first and reads the rows the hash is leaving, and the cleanup afterwards removes only the stale row.

The SDK suite gains a test that moves an embedded file into a second collection and checks that its content_vectors rows survive, its only partition row is in the new collection, and nothing needs embedding; it failed with 0 of 1 content_vectors rows before the change. The CHANGELOG entry now says what the cleanup removes: the vectors of content no indexed document holds any more, which a later return of that content embeds again.

(cherry picked from commit bd86d16)
`renameCollection` ran its store_collections, documents and vector-id updates as separate statements. `vector_collection_ids.name` is unique, and a collection removed while sqlite-vec is not loaded keeps its id, so renaming onto that name failed after the documents had already moved, leaving them under a name whose partition id still belonged to the old collection. The rename now runs in one IMMEDIATE transaction: it renames the store_collections row first (so a name that already exists fails before anything moves), drops a leftover partition and id under the target name, refuses with an explicit message when that id cannot be dropped without sqlite-vec, then moves the documents and the id.

The unscoped `clearAllEmbeddings`, the scoped clear that empties the index, and `ensureVecTable`'s dimension reset each emptied `vector_rows` and dropped the vec0 table in separate statements. `vector_rows` ids are reused once the map is empty, so a process stopped between the two left vec0 rows whose rowids new inserts then collided with. One helper now drops the table and empties the map in a single commit, and the unscoped clear deletes content_vectors in the same transaction.

`searchVec` reads each final result's body after resolving the documents, so another process's orphaned-content cleanup can delete that row in between; the search now drops that result instead of throwing.

Tests cover a rename onto a leftover id with and without a droppable partition, a rename onto an existing collection, a failure injected on the last rename step (all names unchanged), a vanished content row during a search, and a DROP forced to fail during a clear and during a dimension reset (nothing committed). Each failed before the change.

(cherry picked from commit 8db74fa)
… store

When an older build recreates the legacy vector table on an index already at version 2, the next open repairs it; if that repair fails (here a partitioned table with different dimensions), the store must still open with a warning and leave the legacy table for a later attempt. The new test forces that failure and checks the warning, the version, and the layout.

(cherry picked from commit a7d9e17)
`qmd cleanup` compacts `vectors_vec` in place when fewer than 90% of its chunk reads hold live rows. vec0 puts each insert in the first free slot of its newest chunk and reclaims a chunk only once it is completely empty, so a delete in any older chunk leaves a hole that every brute-force scan still reads: re-embeds (DELETE then INSERT per chunk) and orphan removal accumulate them. A 794k-vector index measured at 36% occupancy (2,166 chunks for 776 needed) scanned in 1.8s where a packed copy of the same rows took 0.9s, and VACUUM never touches vec0's fixed-size chunk blobs.

`vectorTableLayout` reads the chunk and row counts from vec0's shadow tables and reports the chunks a packed table would need. `repackVectors` moves the live rows of every chunk filled below 90% to the tail, deleting and re-inserting them one chunk per IMMEDIATE transaction, so the emptied chunks vanish, the write lock is never held for more than one chunk of rows, and an interrupted run leaves a consistent table that the next cleanup finishes. A legacy table without a `hash_seq` key is left for `ensureVecTable` to rebuild. The repack runs after orphan removal and before VACUUM; `previewCleanup` projects the layout after orphan removal so `--dry-run` predicts the same decision, and the cleanup output reports the chunk counts either way.

Tests cover the layout arithmetic, the repack keeping every live row and the nearest-neighbour answer, a packed table left alone, the preview not rebuilding, a failing chunk move rolling back that chunk, the preview matching the run when orphans are pending, a legacy table, and a database without a vector table.

(cherry picked from commit 5ddf8ed)
The expansion model can emit the same line more than once, and nothing
removed the repeats. store.expandQuery only dropped entries equal to the
original query, then cached the rest as-is, and cache reads returned rows
unchanged. hybridQuery and vectorSearchQuery run one search per entry, so
a cached row with 12 identical hyde lines embedded and scanned 12 times.

Add uniqueExpansions(), an exact (type, query) filter, and apply it to
fresh model output before caching and to both cache-read branches, so
rows already bloated by older versions are cleaned as they are read.
vec and hyde entries with the same text are kept; they route to
different searches.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 309843b)
brettdavies and others added 28 commits September 29, 2026 19:36
searchFTS ran one FTS query per named collection and merged the lists
by score (tobi#775), so a scoped search over the default collections
matched and ranked the whole index once per collection. That split
existed because the scoped query only looked at a window of the best
matches, where a large collection could crowd the others out. Since
tobi#918 and tobi#953 the scope sees every match, so one query over the whole
collection list returns the same results as the merge.

searchFTS now filters by the whole list in one query and breaks score
ties by filepath, as the merge did. On a copy of a large index, the 8
fixture queries over 13 default collections took 77.6 s the old way
and 6.1 s in one query, with the same top 20 for each.
tobi#983 moved searchVec's body read into its own lookup, which fetched the
whole document again and dropped tobi#962's cap. The lookup now returns
the first 262,144 characters, like searchFTS, and the three reads that
cap a body (searchFTS, this lookup and getHashesForEmbedding) share
BODY_CAP_CHARS. tobi#983's race test intercepts the lookup's SQL, so its
expected string follows. The tobi#962 body-cap test "searchVec returns at
most 256 KiB of a document body" passes again.
…ed table

tobi#1000's test stubs a legacy vectors_vec table to switch on the vector
branch of hybridQuery. With tobi#983, hasVectorIndex() only recognizes the
partitioned table, so the branch never ran and the test failed with
"expected spy to be called 1 times, but got 0 times". The stub now
creates the partitioned table, as tobi#983 already does for the tests next
to it. The assertions are unchanged.
Renaming a collection moves its tobi#962 sync rows, and tobi#983's atomic
rename carries that move inside its transaction. The test renames
a collection through tobi#983's `renameCollection`, checks the sync row is
under the new name at once, and makes the file unreadable to show the
first pass under the new name reads nothing. Its file is backdated past
the racy-mtime window so the first pass stores a trusted row.
tobi#983 restricts a filtered vector scan to the rows the metadata filter
admits by reading them with all() and binding them as one JSON list, so
a broad filter on a large index put every eligible row on the heap once
per vector search. tobi#962's iterate() change had bounded that read at
20,000 rows.

The scan's `rowid IN` restriction is now the eligibility query itself,
which vec0 applies inside the scan as it did the list. Results stay
exact however many rows the filter admits, and none of those rows
reaches the heap. searchVec still skips a scoped collection with no
eligible row, now from a DISTINCT query over collection ids.
tobi#983's searchVec runs one KNN per collection in scope and merges the
results by distance. The sort kept equal distances in the order the
collections were named, so the same content in two collections came
back from whichever was named first. Ties now go to the smaller
filepath, as they do in the keyword merge.
A collection whose root folder is absent at update time (unmounted drive,
offline share, renamed folder) globbed to an empty file list. The
deactivation pass then read that as "every file was deleted", search went
empty, and a `qmd cleanup` before the folder returned hard-deleted the rows.

Check that the root is a directory before globbing. If it is not, return
without touching the index and report the root as ROOT_MISSING so the CLI
can warn. A genuinely empty folder still deactivates everything.

Fixes tobi#989

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit f8416e7)
The repack reads the partitioned table's shadow tables through the layout resolver, skips the newest chunk of every partition rather than one global newest chunk (vec0 appends each partition's inserts to its own tail), and re-inserts each moved row under its existing rowid and partition so the `vector_rows` mapping stays valid. Each chunk is read again inside its transaction, so a row deleted since the plan, whose rowid another collection has taken over, is not moved into this partition.

The trigger counts occupancy per partition, since rows of different partitions never share a chunk: a packed table needs ceil(rows / chunk size) chunks in each partition, and a repack also needs at least one chunk it would move. The `qmd cleanup --dry-run` projection counts partition rows whose (hash, collection) has no active document and leaves out a chunk whose rows are all orphans, as vec0 drops it with its last row, so the preview predicts the same decision. The test that kept a hash-keyed legacy table untouched goes with the code: a legacy table no longer survives a store open. Tests cover rows staying in their partition, a packed multi-partition table not repacking again, a table below the trigger with no chunk to move, the orphan-only chunk in a dry run, and a rowid reused by another partition mid-repack. The changelog credits tobi#937.
…lback

Two partitions of 1,100 documents, of which cleanup removes all but
three per partition as orphans. After runCleanup reports a repack,
every survivor's stored vector bytes are unchanged, the row in each
partition's newest chunk keeps its chunk and offset, the old chunk's
rows sit in that newest chunk, and the doctor vector sample check (tobi#978)
reports ok. Two more cases: a failed move of a chunk's last row, which
vec0 drops inside the transaction, rolls back with its row; and a hash
embedded in two collections but active in only one is orphaned in the
other alone, so the dry run drops that row and cleanup keeps the active
copy.
Since tobi#983, `qmd update` deletes the vectors of retired documents as it
finishes, so a collection whose directory disappears would lose its
embeddings along with its documents. With tobi#990's root check, update warns
that the root was not found and leaves the collection's documents active
and its vector rows in place.
The partitioned vector index bullet had no credit, and tobi#990's bullet
credited its issue (tobi#989) rather than the PR. Each now ends with its PR
number and author, as qmd's release guidelines ask; tobi#989 stays inline.
mergeSearchResultsByScore merged per-collection result lists by score.
Keyword search runs one query over the whole collection list, and
tobi#983's searchVec merges its per-partition results itself, so nothing
calls it at this level. The two comments that pointed at it now state
the filepath tie-break directly.
Files over 10 MB are skipped with FILE_TOO_LARGE, and the CLI printed
them through the unreadable-file branch as "Skipped unreadable file:
<path> (FILE_TOO_LARGE)" and counted them as unreadable. They now read
"Skipped file over 10 MB: <path>" with their own summary count.
Every write in a scan committed on its own: the content row, the
document, its metadata and tobi#962's sync row, one WAL commit each. A
first index of a large collection paid that per file, about 176 s for
20,000 files, and a 267,000-file collection's first sync pass wrote
4.7 GB of WAL at about 285 files/s.

scanWriteBatch groups the writes of reindexCollection and indexFiles
into transactions of up to 500 files or 250 ms, so another writer waits
at most one batch. The helpers that open their own transaction nest as
savepoints inside it. A failed write rolls back the open batch and
rethrows, leaving no transaction behind; the next run indexes the rolled
back files again. The same 20,000 files now index in 5.0 s.
tobi#956 renamed a metadata filter condition's target from `key` to `field`.
tobi#953's test of a match ranked below the unfiltered candidate window and
the tests of a collection scope combined with a metadata filter still
built `{ key: ... }` filters. Such a filter admits no document, so all
three searches came back empty. They now name `field`.
…r tests

tobi#956 renamed a metadata filter condition's target from `key` to `field`.
tobi#983's searchVec filter tests still built `{ key: ... }` filters, which
admit no document, so six of them came back empty and failed. They now
name `field`.
…chunk

searchVec skips a scoped collection that holds no vector row a metadata
filter admits. The query that finds those collections joined each
document to all of its vector rows. With a simple filter SQLite applies
the filter per document and then walks every row of each eligible
document. With a wide filter it scans the vector rows first and applies
the filter once per chunk: on tobi#951's near-ceiling test (one document of
20,000 chunks, a filter of 31,490 bindings) the query ran 4.1 s under
Bun and 6.9 s under Node per search.

The filter now applies once per document, and EXISTS stops at the
document's first vector row in the collection, whatever plan the filter
leads to. The same query runs in 3 ms under Bun and 5 ms under Node, and
returns the same collections.
…dy cap

The 256 KiB body cap reads `substr(doc, 1, 262144)`, and SQLite's
substr() stops at an embedded NUL. searchFTS, searchVec and
getHashesForEmbedding therefore returned only the part of a body before
its first NUL, however short the document. Indexing keeps NUL
characters, and without the cap these reads return such bodies whole.

The three reads now share cappedBodySql: a body whose UTF-8 encoding
fits the cap comes back as stored, and only a longer one goes through
substr(). A body over 256 KiB with a NUL before the cut still ends at
the NUL. Reading 40 bodies of 5 MB takes as long as before.

A test per read, each with a short body holding a NUL, fails on the
code below, where the body ends at the NUL. Reported in review on tobi#1021.
perf: bound MCP, launcher and doctor costs on large indexes (1/6)
perf(index): skip unchanged files on re-index and cap result bodies (2/6)
fix(query): drop duplicate query expansions (3/6)
fix(search): scope keyword search to the full match set, one ranked list per search (4/6)
perf(store): partition the vector index by collection (5/6)
feat(cleanup): repack the partitioned vector table (6/6)
Bring code merged from tobi/qmd (upstream/main 93d211f) onto the fork's
@systemfsoftware/tsconfig base, the same way #2 migrated the rest of src:

- noPropertyAccessFromIndexSignature: bracket access for Record<string, unknown>
  reads in parseCliMetadataOptions (src/cli/qmd.ts) and the REST /metadata
  handler (src/mcp/server.ts).
- exactOptionalPropertyTypes: widen ListMetadataOptions fields
  (src/metadata-store.ts) and reindexCollectionIn options (src/store.ts) to
  accept explicit undefined. Consumers already treat undefined as absent
  (`??`, `=== undefined`, truthiness), so no runtime behaviour changes.
@ryanleecode
ryanleecode merged commit c93d6c8 into main Oct 7, 2026
6 of 7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.