Repository navigation
chore: sync with upstream tobi/qmd (04e4dbd..93d211f) - #3
Merged
Merged
Conversation
createMcpServer() rebuilds the server instructions on every call, and the Streamable HTTP transport creates one McpServer per MCP session - so every initialize re-runs the full index-status aggregation behind buildInstructions(). On a large index that recomputation dominates the handshake, and agent reconnect bursts stack the status scans (observed as multi-second to tens-of-seconds initializes on an 81k-row index). Cache the built instructions per store with a 60s TTL: concurrent initializes share one in-flight build (a reconnect storm now costs one status scan instead of N), failed builds are dropped so the next initialize retries, and index updates from other processes still surface within a minute without restarting the server. The status tool is untouched and always live. Co-Authored-By: Claude <noreply@anthropic.com> (cherry picked from commit dcb58d3)
… prefilter The rowid IN (...) prefilter defeats FTS5 early termination: SQLite drives the query from the collection's whole doclist and probes the FTS5 index once per candidate rowid (EXPLAIN: VIRTUAL TABLE INDEX 0:=M3), about 55ms per in-scope document. On the live index that is 0.02s unscoped versus 21s scoped to a large collection, and 507s for one pass over 19 collections, which on a single-threaded MCP server is a multi-minute outage for every client. Keep the MATCH global and unbounded, MATERIALIZE the complete match set once (corpus-bounded), and apply the collection filter, ORDER BY and LIMIT in the outer query. Correct (the inner set is complete, so the outer filter cannot truncate it) and fast (natural FTS5 operation, EXPLAIN VIRTUAL TABLE INDEX 0:M3). MATERIALIZED is load-bearing: without it the planner may flatten the CTE and refold the collection predicate into the MATCH. Unscoped path unchanged; result rows identical to the prefilter. (cherry picked from commit 0c7a0b8)
structuredSearch loops the collection list and pushes each collection's results as its own RRF ranked list, so a scope of N collections builds N lists per search leg. RRF ranks with no relevance floor, so every collection's best-of-a-bad-lot enters fusion as a rank 1 beside the real answer. The 2x weight given to rankedLists[0] compounds it: it is meant for a caller who ordered `searches` by importance, but the fan-out makes list 0 whichever collection happened to sort first. The visible symptom is that ranking depends on argument order. Over a 21-collection scope and a 14-question gold set, reversing the collections array moved rank 1 for 13 of 14 queries. searchFTS and searchVec already accept `string | readonly string[]`, so this just passes the scope straight through instead of looping it. On that same set: P@1 0.000 -> 0.500, recall@5 0.643 -> 0.786, MRR 0.279 -> 0.607, and reversing the array no longer changes anything. Median query time 0.499s -> 0.060s, since it is now one FTS and one vector query rather than 21 of each -- searchVec in particular was pulling the same global top-k every time and only then filtering by collection. Single-collection scopes are unchanged, as they must be: the loop ran exactly once there already. (cherry picked from commit aa67e26)
A filter condition is a predicate over one field of the record under evaluation: `field` names it, `operator` says how to compare, `value` is the operand. For a document, the fields are its metadata keys, and that special case is what the property was named after. Rename the property from `key` to `field` so the grammar describes itself without reference to what it happens to be applied to. Nothing else in the AST changes: `operator`, `value`, `operands`, and `operand` keep their names, and the parser rejects the old property the same way it rejects any unknown one, with the JSON path of the failing node. The rename lands before metadata filtering (tobi#910) ships, so no released version accepts `key`. Updates the README filter table and semantics, the skill, the MCP `query` tool description, the CLI help and error examples, the changelog entry, and every test that builds a condition. Assisted-by: Claude Fable 5.1 via Pi
When a collection has no changes, reindexing still reads every file to hash it. With thousands of files this costs minutes of IO and hashing even when nothing changed. Add file_sync_state table tracking mtime_ms, size, and content hash per (collection, path). On reindex, stat the file first and skip the read entirely when mtime and size match the cache. If mtime changed but hash is identical only the cached mtime is updated. Skip files over 10MB and empty files, and remove the cache entry on orphan deletion. Second reindex with no changes performs no file reads, taking update from minutes to seconds when the collection is unchanged. (cherry picked from commit 3fcf230)
Previously, searchFTS and searchVec loaded the whole document body as content.doc as body. On collections with large transcripts this caused 20 candidates times 6 expansions times 100KB = 12MB of string copies through better-sqlite3 into V8, observed as 4.2GB heap. Now the bodies are bounded at the SQL level using substr(content.doc, 1, 262144) as body. This limits each body to 256KB. With 40 winners the max is 10MB vs unbounded before. Rerank input stays 900 tokens, about 3.6KB, per chunk. The long term fix fetches only the winning chunk via substr(doc, pos, len) for the 40 RRF winners, but this commit prevents the OOM first. (cherry picked from commit a75b659)
Extend the filter AST with three text operators, a type test, and an optional case-folding flag, so a filter can reach values by substring and reach one side of a key whose documents disagree on type. - `contains`, `prefix`, and `suffix` take a non-empty string operand and match string values only. `contains` compiles to `instr()`. `prefix` and `suffix` compare UTF-8 bytes with the operand's byte length bound from JavaScript, because SQLite's text `length()` and `substr()` stop at an embedded NUL and the stored string domain admits one. - `type` takes `string`, `number`, or `boolean` and matches the stored type of a key's values. - `caseInsensitive: true` is accepted on any condition whose operand is a string or an array of strings. Both sides fold ASCII letters (SQLite's `lower()` and the same range in JavaScript), so the two agree exactly. Non-ASCII letters compare exactly. Each compiled condition now names how its correlated subquery reaches a document's rows. The covering indexes are (key, value, document_id), so an exact equality (`eq`, `in`, `all`) is one probe by value. Every other condition (ordered comparisons, `ne` and `nin` presence, `exists`, `type`, the text operators, and folded equalities, since `lower()` is opaque to the index) seeks the document's own rows through the primary key: a unary `+` on the key term hides it from the planner, which otherwise walks the key's whole value range once per document when `sqlite_stat1` is absent. On 4,000 documents that plan made one `priority > 10` count 376 ms and a `prefix` count 946 ms; both are under 5 ms with the seek, on Node and Bun, with or without statistics. A test asserts the plan for every operator in both states. Validation reports the failing JSON path as before. The README filter table, the MCP `query` tool description, the skill, and the changelog entry name the new operators. Bind vector candidate IDs as one JSON list during document lookup. A valid near-ceiling filter otherwise overflows Node's SQL variable limit when combined with candidate IDs, in both exact scans and the global fallback. A model-free regression crosses the 20,000-vector boundary and verifies that nonmatching documents sharing a content hash remain excluded. The plan guard accepts SQLite 3.43's USING INDEX label for the same point seek, retaining the check that rejects broad per-document key scans. Assisted-by: Claude Fable 5.1 via Pi
Report keys, typed values, document coverage, collection provenance, and numeric ranges over the same extraction gate as filtered search. `filter` selects documents and `match` selects entries. Both windows run in SQL with exact totals and remainders, including past-end pages. Discovery compiles filters to uncorrelated document sets. Search retains candidate-local seeks. Entry prefix/suffix comparisons return false, not NULL, for empty strings so negation partitions the accepted value domain. Read the report in one deferred transaction. Materialize key ranking once in the type-statistics statement and reuse its selected names thereafter. Read contributing collections with those statistics. Group value counts once for both the distinct total and the value window, retaining one total carrier per type even past the end. Preserve BINARY key and value ordering. Compute even medians without overflow or loss of adjacent subnormals. Validate safe-integer options and cast the production window operands to INTEGER before adding. Check the widest planned statement against the 30,000-parameter application budget before preparing discovery queries. Provide a lean key overview and a batched per-collection overview. Ordinal zero counts each document/key once regardless of array length, without reading values or maintaining another index or cache. Tests cover aggregation, gates, filters, matches, ordering, both windows, independent boolean expectations, live-writer snapshots, lifecycle changes, actual endpoint SQL, binding limits, and structural query-work guards with and without statistics. The independent 14,276-case predicate oracle and 120 aggregate comparisons pass on both drivers after these changes. Sort contributing collection names by UTF-8 bytes after reading their aggregate. This preserves SQLite BINARY order without requiring SQLite 3.44's aggregate ORDER BY syntax. The real discovery query and the filter, discovery, and vector suites also pass on SQLite 3.43.2. Assisted-by: Claude Fable 5.1 via Pi
Adds the CLI surface for metadata discovery so an agent can learn what to filter on before it queries: - `qmd collection metadata [name...]` prints one block per key: a header with type, coverage, and distinct count, then a value window. Strings list vertically with right-aligned counts, numbers show min/median/max plus the values when they fit, booleans print true/false counts. Keys whose documents disagree on type split per type, with contributing collections in the multi-collection view. - `--filter <json>` counts only matching documents. The output then opens with a `filter:` line stating how many documents pass, and every key's coverage is measured against that population, so the numbers an agent reads are the numbers a filtered search would see. - `--match <json>` selects which metadata entries are reported, with the same AST as `--filter` evaluated against each entry (a condition's `field` is the entry's `key` or `value`). Both JSON flags share one parser and report malformed JSON and invalid nodes the same way. - Two windows, named the same on every surface: `--key-limit <n>` (default 50) with `--key-offset <n>` and `--all-keys`, and `--value-limit <n>` (default 10) with `--value-offset <n>` and `--all-values`. Every windowed list ends with the exact remainder and the flags that reach it, and an offset past the last key says how many keys there are. `-n` and `--all` are the single-window search flags and exit with a pointer at the discovery ones. `--sort count|value` orders values, `--min-count` drops the tail. Omitting the name covers the default collections, as an unscoped search does. - src/metadata-format.ts holds the renderer so the MCP tool can print the same shape. It takes the CLI's color palette or none. The empty states render there too, so the filter line survives a filter whose documents declare no metadata. A string value prints bare only when the bare form is unambiguous. One that is empty, padded, contains a quote, a backslash, a control character, a line separator, or a delimiter of the compact `value (count), ...` list a type split prints, or that reads as a JSON number, boolean, or null, or as the list's remainder tail, prints as a JSON string with every control character escaped. One rule for both layouts, so a value never prints two ways, and `"a (1), b" (1)` cannot be read as two values. Ordinary values are unchanged, so the common case costs nothing. - A filter and match that together exceed the SQL binding budget exit with the store's message. - The pending-extraction warning prints on stderr whenever documents are gated out, since discovery always applies the gate. The renderer suite reads the compact list back the way a reader must, honoring quotes, and checks every item round-trips. Covers the keys view, both windows with paging and their footers, the filter line and denominator (including over an empty result), reverse lookup, a composed match across both fields, type selection, sort and min-count, default scope, empty matches, and exit codes for unknown collections and invalid flags, filters, and matches. The renderer has its own suite for string identity (round trip through JSON.parse, no raw control characters) and the empty states. Numeric flags reject unsafe integers as usage errors before opening the index, preserving the original input in the message. Unsupported format flags fail explicitly with a pointer to the structured surfaces, rather than silently returning text to a caller expecting JSON. Assisted-by: Claude Fable 5.1 via Pi
Makes metadata visible from the commands a user or agent already runs first, so the drill-down is discoverable without knowing it exists: - `qmd collection list` adds a `Metadata:` line per collection naming the top five keys by coverage and counting the rest (`+N more`). Omitted when the collection has no metadata, like `Ignore:`. - `qmd collection show` adds a `Metadata` section: a coverage line (keys, documents with metadata, pending extraction), then the top five keys as aligned rows with type, coverage, distinct count, and a short value preview (top three strings, numeric range and median, boolean counts, or a pointer when types disagree), then a pointer at `qmd collection metadata <name>` when keys were left out. - `qmd status` adds one summary line under Documents, placed with the existing pending-extraction line so the two read together. - Every view asks the store for exactly what it prints. `show` requests a five-key, three-value window and reads the total and remainder from the result, `list` reads the first five key names and counts the rest, and `status` counts keys without listing them, so none of them grows with the vocabulary. - src/metadata-store.ts gains the light queries these views need: `countDocumentsWithMetadata()` and an optional collection scope on `countDocumentsPendingMetadata()`. The pending-extraction warning that `collection metadata` and the filtered search commands print takes that scope, so it counts the collections the command reads and not the whole index. `status` keeps the index-wide count. Collection list batches all key overviews in one query, including the vocabulary totals, instead of counting and ranking separately for each collection. CLI status counts metadata-bearing documents with existence probes and counts keys from ordinal-zero rows, avoiding array expansion. Type-conflict drill-down hints escape apostrophes in the shell-quoted JSON operand. A POSIX shell round-trip regression ensures metadata keys remain literal data, including command-substitution syntax. Assisted-by: Claude Fable 5.1 via Pi
Ships discovery on the remaining surfaces with the same options object and result shape the CLI uses: - SDK: `listMetadata(options?)` on QMDStore under Collection Management, next to listCollections(). `match` and `filter` are validated at the runtime boundary like the search methods, and the window options by the store, so a caller sees MetadataFilterError, MetadataOptionError, or MetadataBindingBudgetError with the same message every surface prints. Discovery types, MetadataMatch, the error classes, the default limits, the binding budget, and parseMetadataMatch() are exported from the package root, and MetadataValueType now lives in metadata.ts beside the scalar types it describes. - MCP `metadata` tool: flat and read-only, taking `collections`, `match`, `filter` (both validated like the query tool's filter), `keyLimit`, `keyOffset`, `valueLimit`, `valueOffset`, `sort`, and `minCount`. The description teaches the one grammar over two record types (a document for `filter`, a metadata entry for `match`), a table of matches and the question each answers, how to read `totalKeys`, `remainingKeys`, `remainingValues`, and `range`, and how to page. Text content renders the CLI shape through the shared formatter, empty states included so the filter line survives them, naming the option that reaches each remainder; structuredContent is the ListMetadataResult. The schema's integers are safe integers, and a filter and match over the binding budget return a tool error. - MCP `status`: CollectionInfo gains `metadataKeyCount` and a windowed `metadataKeys` (the ten most covered keys with types), mirrored in StatusResult and listed in the text summary with `+N more` and a pointer at the metadata tool, so the call agents make first reveals that metadata exists without growing with the vocabulary. - HTTP `POST /metadata`: same body as the tool. 400 on a non-object body, a non-object or invalid match or filter, a bad sort, a non-array `collections`, a non-number window option, a number outside its domain, or a filter and match over the binding budget (the store's message, verbatim). Returns the ListMetadataResult as its own envelope. Logged like POST /query. Updates the exact tool-list assertions in test/mcp.test.ts and covers each surface in test/metadata-surfaces.test.ts, including paging, the 400 paths, invalid matches and options, unsafe integers, the binding budget on all three surfaces, the filter line over an empty tool result, and status structured content. Full status batches all collection metadata overviews in one query. An internal getStatusSummary path returns only the index facts initialization uses, without pretending uncomputed metadata counts are zero. MCP server creation uses that path, including each stateless HTTP request. The public SDK getStatus result and MCP status tool keep their metadata summaries. Regression tests exercise empty-string negation through SDK, HTTP, and MCP. A real MCP tools/list request succeeds with the metadata value table hidden in the isolated fixture, proving initialization does not read it. Restoring the eager status call makes that test fail. Assisted-by: Claude Fable 5.1 via Pi
- README: a "Metadata Discovery" subsection after "Metadata Filtering" with the mental model (same metadata, different unit: filter narrows documents, match narrows the metadata itself, one predicate grammar over both), a table of matches and the question each answers, a wide-to-narrow walkthrough ending in a validated query, exact CLI output for string, number, boolean, threshold, filtered, and type-conflict keys, the two windows, how they page, and what they do and do not bound, the reading rules (including when a string prints quoted and that a result is one snapshot), and the SDK method with its predicate, option, and error types, MCP tool, and HTTP route. Also the new subcommand in the collection commands block, the `metadata` tool parameters, and the `POST /metadata` endpoint. - CHANGELOG: one entry under Unreleased / Added. - skills/qmd/SKILL.md: a "Discover metadata before filtering" section written for an agent deciding what to filter on: the match shapes worth knowing, how to read the header and the filter line, the window footers and how to page past them, numeric ranges, and type splits. - CLAUDE.md: `qmd collection metadata` in the command table and the Collection Management block. Examples use the status, topics, priority, and reviewed keys the filtering docs already establish, so the two sections read as one. Qualify the windows as bounds on key/value lists, not query work or collection provenance. Document the current-data cost of status summaries and their exclusion from MCP initialization. Name the filter target field consistently and explain explicit rejection of unsupported CLI formats. Assisted-by: Claude Fable 5.1 via Pi
Large result sets previously used all() which materialized full arrays in V8 heap. On an index with 13k files and 89k chunk vectors this means tens of thousands of rows (active paths, sync-state entries, pending-embedding docs, candidate chunk vectors) in a single array, which caused OOM. Now they stream via iterate(): sync-state reads, pending-embedding scans, active-path listings, glob/fuzzy scans, hash_seq eligibility (with early exit past the cap, falling back to ANN), and embedding-body fetches (substr-bound, with a bytes-only path). Small queries remain with all() as they are bounded. (cherry picked from commit 3c37c68)
A holistic pass over the feature after the correctness and performance audits, asking whether each piece is as small and intentional as it can be. No behavior changes. The full matrix, the close-out packet's semantic checks (adapted for the removed helpers), and the SQLite 3.43.2 path all pass on Bun and Node. Code: - Remove `listMetadataKeys()` and `getMetadataOverview()` from metadata-store.ts. Neither had a production caller: every status view reads `listMetadataCollectionSummaries()`. Their doc comment described the pre-batching design. `queryMetadataOverviews()` loses the SQL fragment parameter that only they varied. - Stop exporting `DEFAULT_METADATA_KEY_LIMIT`, `DEFAULT_METADATA_VALUE_LIMIT`, and `METADATA_SQL_BINDING_BUDGET` from the SDK. The defaults are in the option docs and the budget is an internal safety valve. The error classes stay exported. - Un-export `formatMetadataKeySummary()`, which only its own module calls. - Measure text-operator byte lengths with `Buffer.byteLength`, the idiom the rest of the file already uses, instead of a second `TextEncoder`. - Merge the two stacked comments on `warnPendingMetadata()` into one. - Empty-result messages on the CLI and MCP tool say "see which keys exist" rather than "see every key", since the default window is 50. - Note why vitest's typecheck ignores source errors: it would otherwise report pre-existing errors under tools/, and tsc covers src/. Tests: - Fold the `listMetadataKeys` assertions into the batched-overview tests. - In the "rank once, group once" test, keep the statement-count guard and the materialization guard (the ranking CTE plans as MATERIALIZE, never CO-ROUTINE, with and without statistics). One ranking statement does not mean one evaluation of the ranking, and the guard fails the moment MATERIALIZED is dropped. Remove the EXISTS and LIST SUBQUERY assertions, which pin the filter compiler's plan and belong to its own suite. Docs, sized to sit beside their neighbors: - README: list the `metadata` tool under "Tools exposed" (it was missing) and move its parameter rows out of the middle of `multi_get`'s. Lead the Discovery section with the one-sentence model, keep one example per idea, render the window flags in the README's own flag-list style, shorten the reading rules, and drop the internals (binding budget constant, status cost, MCP initialization) that belong in the PR. - CHANGELOG: trim the discovery entry to user-visible behavior. - MCP `metadata` tool description: cut to the mental model, four match shapes, and how to read and page the result. It is now shorter than the `query` tool's. - skills/qmd/SKILL.md: restructure the two metadata sections in the skill's own style (wrapped prose, short paragraphs, bulleted reading rules, five examples instead of nine). Add one sentence of judgment in the register the skill already uses for `-c`: add a metadata filter when it expresses a constraint the user intends, and check coverage first, because a value condition only sees documents that declare its key. A matching "Do not filter blind" pitfall. Bump the skill version to 2.3.0. Assisted-by: Claude Fable 5.1 via Pi
(cherry picked from commit ace57bd)
(cherry picked from commit 0bd6c19)
Add src/vec-layout.ts as the one place that knows which sqlite-vec table is live and what it is called. It resolves `none`, the legacy `vectors_vec` (hash_seq or hash keyed) or the partitioned `vectors_by_collection`, with the vec0 shadow-table names and the stored dimension, and it owns the two plain tables the partitioned layout needs: `vector_collection_ids` (a stable integer id per collection name, so a rename is one row) and `vector_rows` (vec0 rowid to hash, seq, collection_id, NOT NULL and UNIQUE per triple). Both are created at open next to `content_vectors`. The resolver also carries the collection id allocate, resolve, rename and delete helpers and the JSON rowid list used to bind a scoped search's rowids through `json_each` in one parameter. Nothing reads the partitioned table yet; the legacy table stays live. (cherry picked from commit 3cdb92d)
…ction partitions Add src/store-migrations.ts with the PRAGMA user_version ladder: version 1 installs the FTS sync triggers as before, and version 2 copies every row of the legacy `vectors_vec` table into the partitioned layout, one vec0 row per (chunk, active collection), keyed by the integer collection id. Nothing calls the step at open yet; that lands with the readers and writers that use the new layout. The copy reads the legacy table's chunk blobs directly (validity bitmap, rowid list, packed float32 vectors) at one chunk per IMMEDIATE transaction and keeps the last copied chunk_id in store_config. A process killed between chunks resumes from that cursor, and two processes opening the same database take turns, because each reads the cursor inside its own write transaction and stops as soon as the other stamps the version. A verification pass then copies by key any chunk that content_vectors and an active document claim but no partition holds (a pre-upgrade repack can move a row into a chunk the walk already passed), content_vectors rows with no vector anywhere are deleted so the pending detector embeds them again, and the version stamp and the DROP of the legacy table share one transaction. A legacy table keyed by hash alone is dropped without a copy, as before. Without sqlite-vec the step defers and the version stays at 1. vec0 checks the bound type of its partition key and rowid rather than applying affinity, and better-sqlite3 binds every JS number as a double, so those parameters are bound as bigint through `vecInteger`. The anti-join that finds chunks missing from a partition lives in vec-layout.ts because the embed path will need the same query. (cherry picked from commit 7d17354)
Vector search, embedding writes, cleanup and the collection operations now use the layout vec-layout.ts describes: one vec0 table partitioned by an integer collection id with one row per (chunk, active collection), a `vector_rows` mapping from vec0 rowid to (hash, seq, collection_id), and the version 2 migration wired into store open so every process copies the legacy `vectors_vec` table before it can read vectors. A collection-scoped `searchVec` is one KNN per collection id with `k = limit * 3` (capped at sqlite-vec's 4096), merged by distance, so a small collection costs its own rows instead of a pre-filter over the whole index and is never crowded out by a large one (tobi#775, tobi#791, tobi#803). The unscoped search is one KNN over all partitions with the same k. Step two joins the returned rowids through `json_each` in a single bound parameter, because thirteen partitions at the k cap exceed SQLite's separate-parameter limit. A scope naming an unknown collection skips it; an all-unknown scope is empty. The exact-scan, capped-ANN and post-filter paths are gone, as is the per-collection recursion; `mergeSearchResultsByScore` stays for FTS. A metadata filter restricts each KNN to the vector rows of documents it admits, passed as `rowid IN (SELECT value FROM json_each(?))`, which vec0 0.1.9 applies inside the scan alongside the partition key. The result is an exact top-k of the eligible rows at any eligible-set size, so the 20,000-row exact-scan cutoff and the global over-fetch above it go away, and with them the false empty a large eligible set got when more than the over-fetch of closer rows were ineligible. The filter is re-applied on the document join, since one hash can back a matching and a non-matching document in the same collection. On the live index the restriction costs nothing measurable: a stars KNN (794k rows) restricted to half its rows took 1.04 s against 1.16 s unrestricted. `insertEmbedding` writes the content_vectors row and one partition row per active collection of the hash in one transaction, replacing an existing row under the same rowid since vec0 ignores OR REPLACE. `cleanupOrphanedVectors` removes partition rows whose (hash, collection) has no active document, so a stale row in one collection cannot steal k slots there while the hash stays live elsewhere. `clearAllEmbeddings(collection)` clears the collection's partition rows for hashes exclusive to it and keeps shared ones in every partition, `removeCollection` deletes the partition and its id, `renameCollection` updates the id's name, and `removeIncompleteEmbeddings` drops every partition row of the hash. Doctor's stored-vector check and the legacy-fingerprint adoption sample resolve the vector through `vector_rows`. `ensureVecTable` keeps the dimension-mismatch check on the partitioned table and runs the copy itself if it meets a legacy table, which covers an older build having recreated one on an already-migrated index. The metadata search suite gains the over-20,000-eligible case, which returns 0 of 5 documents on the unpartitioned code, and a same-collection shared-hash case. The store-level tests move to `insertEmbedding` fixtures, the MCP fixture seeds the legacy table and runs the migration, and the collection-filter suite gains the scoped-search semantics tests (shared hashes named per collection, inactive rows, unknown and empty scopes, a limit above the k cap, one-row-per-file collapse, per-partition k, rename, remove and scoped clear). No reader needs a retry for the moment the legacy table is dropped: the migration runs inside `createStore`, so no process using this code reads vectors before its own open has finished the copy. (cherry picked from commit c8cbc4f)
…nd write each batch in one transaction `generateEmbeddings` starts with a copy pass: every active (hash, collection) pair whose chunks are in content_vectors but missing from that collection's partition gets its rows copied from whichever partition holds them, so a document that appears in a second collection is searchable there after the next embed run without another model call. A chunk no partition holds has its hash's rows removed instead, which puts the hash back in front of the pending detector; that closes the case where content_vectors claims a chunk the index lost. `EmbedResult.chunksCopied` reports the copies and `qmd embed` prints them. Each 32-chunk batch is written inside one IMMEDIATE transaction, with the success and failure bookkeeping applied after the commit. A chunk writes several rows across three tables, and a kill or a write error mid-batch now rolls the whole batch back instead of leaving content_vectors ahead of the index; the existing per-chunk fallback then retries the batch's chunks one by one. (cherry picked from commit 48e564f)
A changed file keeps its document row and gets a new hash, which strands the old hash's rows in the collection's vector partition. The partition filter cannot see `documents.active`, so on a churning collection those rows take k slots from every scoped search until something removes them; with enough of them a scoped query comes back short or empty. `qmd update` now runs the orphan-vector cleanup as its last step and reports the rows it removed, and the hint that used to suggest `qmd cleanup` for orphaned chunks goes away with the orphans. The regression test re-indexes a file through four rewrites with a fake embedder that places the stale bodies nearer the query than the live documents, shows the scoped search returning nothing, and checks that the cleanup restores the full count. Two tests cover the legacy-fingerprint adoption sample, which resolves its nearest stored vector through the partition mapping. (cherry picked from commit 17589bc)
CHANGELOG gets the Unreleased entry for the index format change: what the partitioned layout does for scoped queries, that the first command after upgrading converts the index in place with progress, resumption and shared work between concurrent commands, what rename, update and embed now do with vectors, and that downgrading needs `qmd embed -f` and a running `qmd mcp` server a restart. The README schema block lists `vector_collection_ids`, `vector_rows` and `vectors_by_collection` in place of `vectors_vec`. (cherry picked from commit c9cfc20)
…py and clear paths `activeCollectionsOfHash`, `deletePartitionRowsOfHash` and a prepared-statement `storedEmbeddingLookup` in vec-layout.ts replace three copies of the same SQL in store.ts, the migration reuses the exported `parseDimensions` instead of its own regex, and the unused `STORE_SCHEMA_VERSION` alias goes. The collection-scoped `clearAllEmbeddings` now runs inside one IMMEDIATE transaction, so a force re-embed of one collection is a single commit and a crash cannot leave the mapping and the vec0 table out of step. (cherry picked from commit 459ef0c)
… layout Findings from the ce-code-review run 20260904-161233-cdf84496, all applied: - The migration's verification pass, the vectorless content_vectors delete, the cursor delete, the version stamp and the DROP of the legacy table now run in one IMMEDIATE transaction, so a writer on an older build cannot land a row in the legacy table between the delete and the drop and lose its only vector. The partitioned table is created with IF NOT EXISTS inside a double-checked transaction, so two first openers no longer race on CREATE VIRTUAL TABLE. `runStoreMigrations` also runs the copy when a legacy table exists on an already-stamped store, as the module's own comment claimed; when that repair fails on a dimension mismatch it warns and defers instead of blocking every open. - `qmd update`, `qmd collection add` and the library's `update()` now run the copy pass, so a document that appears in a second collection is searchable there right after indexing instead of after the next embed run, and the library's `update()` also removes stale partition rows and reports both counts (`staleVectorsRemoved`, `vectorsCopied`). - `removeCollection` runs as one IMMEDIATE transaction, so an interrupted removal cannot leave the vec0 rows, the mapping and the collection id out of step. - `copyVectorsToNewCollections` commits in slices of 1,000 rows, so a collection that gains many already-embedded documents no longer holds the write lock for the whole copy; a hash queued for re-embedding in one slice is skipped by later ones. Tests cover each change: the flip holds one write lock, a row written after the copy phase still lands in a partition, a second opener finishes a migration another opener started, a legacy table on a stamped store is migrated at open, `qmd update` and the SDK `update()` copy vectors into a collection that gained a hash, `removeCollection` rolls back as a unit, and a 2,500-row copy commits in batches with straddling hashes queued correctly. (cherry picked from commit 1921727)
A vector search asked each partition for `limit * 3` chunks and collapsed them to one row per file. When one long document held the nearest chunks, those slots all went to it and the search returned fewer documents than it asked for: one 20-chunk document and one single-chunk document gave 1 of 2 results at limit 2 and at limit 5, scoped or unscoped. The same happened when stale rows of a changed file sat nearer the query than any live document. The case was reported with a synthetic reproducer in a review comment on tobi#936. Each scan target now resolves its own matches to documents. While they collapse into fewer than `limit` documents and the target holds rows beyond them, the KNN runs again with k doubled, up to sqlite-vec's 4096 cap. Each target returns its own nearest `limit` documents (or all it holds), so merging the targets by distance gives the exact nearest `limit` across the scope. The first KNN keeps `limit * 3`: on the live index, stars' top 60 chunks covered 46 to 60 distinct documents across ten sampled queries, so the second round is a guard for skewed data rather than a cost on typical queries. Step two is prepared once per search and leaves document bodies out: one long document can hold every match of a widened KNN, so only the final results load their body. The metadata predicate (current, error-free extraction plus the compiled filter) is built in one place for the eligible-row query and the document join. The store suite gains the one-long-document case, scoped and unscoped, and the metadata suite gains the reporter's two filtered cases: a long document with an ineligible copy of its content starving a second eligible document, and the best chunk of a 20- and a 450-chunk document surviving deduplication. All four returned one document before this change. The stale-row test's pre-cleanup search now answers instead of coming back empty. (cherry picked from commit d8ae76e)
…le rows `qmd update` and the library's `update()` removed stale partition rows first and ran the copy pass second. When a file moved from one collection to another between two updates, the cleanup deleted the old collection's partition rows for the hash, and with them its content_vectors rows, before the copy could read them; the copy then found no source, queued the hash, and the moved document needed a fresh embed. The copy now runs first and reads the rows the hash is leaving, and the cleanup afterwards removes only the stale row. The SDK suite gains a test that moves an embedded file into a second collection and checks that its content_vectors rows survive, its only partition row is in the new collection, and nothing needs embedding; it failed with 0 of 1 content_vectors rows before the change. The CHANGELOG entry now says what the cleanup removes: the vectors of content no indexed document holds any more, which a later return of that content embeds again. (cherry picked from commit bd86d16)
`renameCollection` ran its store_collections, documents and vector-id updates as separate statements. `vector_collection_ids.name` is unique, and a collection removed while sqlite-vec is not loaded keeps its id, so renaming onto that name failed after the documents had already moved, leaving them under a name whose partition id still belonged to the old collection. The rename now runs in one IMMEDIATE transaction: it renames the store_collections row first (so a name that already exists fails before anything moves), drops a leftover partition and id under the target name, refuses with an explicit message when that id cannot be dropped without sqlite-vec, then moves the documents and the id. The unscoped `clearAllEmbeddings`, the scoped clear that empties the index, and `ensureVecTable`'s dimension reset each emptied `vector_rows` and dropped the vec0 table in separate statements. `vector_rows` ids are reused once the map is empty, so a process stopped between the two left vec0 rows whose rowids new inserts then collided with. One helper now drops the table and empties the map in a single commit, and the unscoped clear deletes content_vectors in the same transaction. `searchVec` reads each final result's body after resolving the documents, so another process's orphaned-content cleanup can delete that row in between; the search now drops that result instead of throwing. Tests cover a rename onto a leftover id with and without a droppable partition, a rename onto an existing collection, a failure injected on the last rename step (all names unchanged), a vanished content row during a search, and a DROP forced to fail during a clear and during a dimension reset (nothing committed). Each failed before the change. (cherry picked from commit 8db74fa)
… store When an older build recreates the legacy vector table on an index already at version 2, the next open repairs it; if that repair fails (here a partitioned table with different dimensions), the store must still open with a warning and leave the legacy table for a later attempt. The new test forces that failure and checks the warning, the version, and the layout. (cherry picked from commit a7d9e17)
`qmd cleanup` compacts `vectors_vec` in place when fewer than 90% of its chunk reads hold live rows. vec0 puts each insert in the first free slot of its newest chunk and reclaims a chunk only once it is completely empty, so a delete in any older chunk leaves a hole that every brute-force scan still reads: re-embeds (DELETE then INSERT per chunk) and orphan removal accumulate them. A 794k-vector index measured at 36% occupancy (2,166 chunks for 776 needed) scanned in 1.8s where a packed copy of the same rows took 0.9s, and VACUUM never touches vec0's fixed-size chunk blobs. `vectorTableLayout` reads the chunk and row counts from vec0's shadow tables and reports the chunks a packed table would need. `repackVectors` moves the live rows of every chunk filled below 90% to the tail, deleting and re-inserting them one chunk per IMMEDIATE transaction, so the emptied chunks vanish, the write lock is never held for more than one chunk of rows, and an interrupted run leaves a consistent table that the next cleanup finishes. A legacy table without a `hash_seq` key is left for `ensureVecTable` to rebuild. The repack runs after orphan removal and before VACUUM; `previewCleanup` projects the layout after orphan removal so `--dry-run` predicts the same decision, and the cleanup output reports the chunk counts either way. Tests cover the layout arithmetic, the repack keeping every live row and the nearest-neighbour answer, a packed table left alone, the preview not rebuilding, a failing chunk move rolling back that chunk, the preview matching the run when orphans are pending, a legacy table, and a database without a vector table. (cherry picked from commit 5ddf8ed)
The expansion model can emit the same line more than once, and nothing removed the repeats. store.expandQuery only dropped entries equal to the original query, then cached the rest as-is, and cache reads returned rows unchanged. hybridQuery and vectorSearchQuery run one search per entry, so a cached row with 12 identical hyde lines embedded and scanned 12 times. Add uniqueExpansions(), an exact (type, query) filter, and apply it to fresh model output before caching and to both cache-read branches, so rows already bloated by older versions are cleaned as they are read. vec and hyde entries with the same text are kept; they route to different searches. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 309843b)
searchFTS ran one FTS query per named collection and merged the lists by score (tobi#775), so a scoped search over the default collections matched and ranked the whole index once per collection. That split existed because the scoped query only looked at a window of the best matches, where a large collection could crowd the others out. Since tobi#918 and tobi#953 the scope sees every match, so one query over the whole collection list returns the same results as the merge. searchFTS now filters by the whole list in one query and breaks score ties by filepath, as the merge did. On a copy of a large index, the 8 fixture queries over 13 default collections took 77.6 s the old way and 6.1 s in one query, with the same top 20 for each.
tobi#983 moved searchVec's body read into its own lookup, which fetched the whole document again and dropped tobi#962's cap. The lookup now returns the first 262,144 characters, like searchFTS, and the three reads that cap a body (searchFTS, this lookup and getHashesForEmbedding) share BODY_CAP_CHARS. tobi#983's race test intercepts the lookup's SQL, so its expected string follows. The tobi#962 body-cap test "searchVec returns at most 256 KiB of a document body" passes again.
…ed table tobi#1000's test stubs a legacy vectors_vec table to switch on the vector branch of hybridQuery. With tobi#983, hasVectorIndex() only recognizes the partitioned table, so the branch never ran and the test failed with "expected spy to be called 1 times, but got 0 times". The stub now creates the partitioned table, as tobi#983 already does for the tests next to it. The assertions are unchanged.
Renaming a collection moves its tobi#962 sync rows, and tobi#983's atomic rename carries that move inside its transaction. The test renames a collection through tobi#983's `renameCollection`, checks the sync row is under the new name at once, and makes the file unreadable to show the first pass under the new name reads nothing. Its file is backdated past the racy-mtime window so the first pass stores a trusted row.
tobi#983 restricts a filtered vector scan to the rows the metadata filter admits by reading them with all() and binding them as one JSON list, so a broad filter on a large index put every eligible row on the heap once per vector search. tobi#962's iterate() change had bounded that read at 20,000 rows. The scan's `rowid IN` restriction is now the eligibility query itself, which vec0 applies inside the scan as it did the list. Results stay exact however many rows the filter admits, and none of those rows reaches the heap. searchVec still skips a scoped collection with no eligible row, now from a DISTINCT query over collection ids.
tobi#983's searchVec runs one KNN per collection in scope and merges the results by distance. The sort kept equal distances in the order the collections were named, so the same content in two collections came back from whichever was named first. Ties now go to the smaller filepath, as they do in the keyword merge.
A collection whose root folder is absent at update time (unmounted drive, offline share, renamed folder) globbed to an empty file list. The deactivation pass then read that as "every file was deleted", search went empty, and a `qmd cleanup` before the folder returned hard-deleted the rows. Check that the root is a directory before globbing. If it is not, return without touching the index and report the root as ROOT_MISSING so the CLI can warn. A genuinely empty folder still deactivates everything. Fixes tobi#989 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit f8416e7)
The repack reads the partitioned table's shadow tables through the layout resolver, skips the newest chunk of every partition rather than one global newest chunk (vec0 appends each partition's inserts to its own tail), and re-inserts each moved row under its existing rowid and partition so the `vector_rows` mapping stays valid. Each chunk is read again inside its transaction, so a row deleted since the plan, whose rowid another collection has taken over, is not moved into this partition. The trigger counts occupancy per partition, since rows of different partitions never share a chunk: a packed table needs ceil(rows / chunk size) chunks in each partition, and a repack also needs at least one chunk it would move. The `qmd cleanup --dry-run` projection counts partition rows whose (hash, collection) has no active document and leaves out a chunk whose rows are all orphans, as vec0 drops it with its last row, so the preview predicts the same decision. The test that kept a hash-keyed legacy table untouched goes with the code: a legacy table no longer survives a store open. Tests cover rows staying in their partition, a packed multi-partition table not repacking again, a table below the trigger with no chunk to move, the orphan-only chunk in a dry run, and a rowid reused by another partition mid-repack. The changelog credits tobi#937.
…lback Two partitions of 1,100 documents, of which cleanup removes all but three per partition as orphans. After runCleanup reports a repack, every survivor's stored vector bytes are unchanged, the row in each partition's newest chunk keeps its chunk and offset, the old chunk's rows sit in that newest chunk, and the doctor vector sample check (tobi#978) reports ok. Two more cases: a failed move of a chunk's last row, which vec0 drops inside the transaction, rolls back with its row; and a hash embedded in two collections but active in only one is orphaned in the other alone, so the dry run drops that row and cleanup keeps the active copy.
Since tobi#983, `qmd update` deletes the vectors of retired documents as it finishes, so a collection whose directory disappears would lose its embeddings along with its documents. With tobi#990's root check, update warns that the root was not found and leaves the collection's documents active and its vector rows in place.
mergeSearchResultsByScore merged per-collection result lists by score. Keyword search runs one query over the whole collection list, and tobi#983's searchVec merges its per-partition results itself, so nothing calls it at this level. The two comments that pointed at it now state the filepath tie-break directly.
Files over 10 MB are skipped with FILE_TOO_LARGE, and the CLI printed them through the unreadable-file branch as "Skipped unreadable file: <path> (FILE_TOO_LARGE)" and counted them as unreadable. They now read "Skipped file over 10 MB: <path>" with their own summary count.
Every write in a scan committed on its own: the content row, the document, its metadata and tobi#962's sync row, one WAL commit each. A first index of a large collection paid that per file, about 176 s for 20,000 files, and a 267,000-file collection's first sync pass wrote 4.7 GB of WAL at about 285 files/s. scanWriteBatch groups the writes of reindexCollection and indexFiles into transactions of up to 500 files or 250 ms, so another writer waits at most one batch. The helpers that open their own transaction nest as savepoints inside it. A failed write rolls back the open batch and rethrows, leaving no transaction behind; the next run indexes the rolled back files again. The same 20,000 files now index in 5.0 s.
tobi#956 renamed a metadata filter condition's target from `key` to `field`. tobi#953's test of a match ranked below the unfiltered candidate window and the tests of a collection scope combined with a metadata filter still built `{ key: ... }` filters. Such a filter admits no document, so all three searches came back empty. They now name `field`.
…chunk searchVec skips a scoped collection that holds no vector row a metadata filter admits. The query that finds those collections joined each document to all of its vector rows. With a simple filter SQLite applies the filter per document and then walks every row of each eligible document. With a wide filter it scans the vector rows first and applies the filter once per chunk: on tobi#951's near-ceiling test (one document of 20,000 chunks, a filter of 31,490 bindings) the query ran 4.1 s under Bun and 6.9 s under Node per search. The filter now applies once per document, and EXISTS stops at the document's first vector row in the collection, whatever plan the filter leads to. The same query runs in 3 ms under Bun and 5 ms under Node, and returns the same collections.
…dy cap The 256 KiB body cap reads `substr(doc, 1, 262144)`, and SQLite's substr() stops at an embedded NUL. searchFTS, searchVec and getHashesForEmbedding therefore returned only the part of a body before its first NUL, however short the document. Indexing keeps NUL characters, and without the cap these reads return such bodies whole. The three reads now share cappedBodySql: a body whose UTF-8 encoding fits the cap comes back as stored, and only a longer one goes through substr(). A body over 256 KiB with a NUL before the cut still ends at the NUL. Reading 40 bodies of 5 MB takes as long as before. A test per read, each with a short body holding a NUL, fails on the code below, where the body ends at the NUL. Reported in review on tobi#1021.
perf: bound MCP, launcher and doctor costs on large indexes (1/6)
perf(index): skip unchanged files on re-index and cap result bodies (2/6)
fix(query): drop duplicate query expansions (3/6)
fix(search): scope keyword search to the full match set, one ranked list per search (4/6)
perf(store): partition the vector index by collection (5/6)
feat(cleanup): repack the partitioned vector table (6/6)
…10-07 # Conflicts: # src/store.ts
Bring code merged from tobi/qmd (upstream/main 93d211f) onto the fork's @systemfsoftware/tsconfig base, the same way #2 migrated the rest of src: - noPropertyAccessFromIndexSignature: bracket access for Record<string, unknown> reads in parseCliMetadataOptions (src/cli/qmd.ts) and the REST /metadata handler (src/mcp/server.ts). - exactOptionalPropertyTypes: widen ListMetadataOptions fields (src/metadata-store.ts) and reindexCollectionIn options (src/store.ts) to accept explicit undefined. Consumers already treat undefined as absent (`??`, `=== undefined`, truthiness), so no runtime behaviour changes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Syncs the fork with upstream
tobi/qmd. The head is a real merge commit, not a rebase or squash.Upstream range
04e4dbd..93d211f(merge-base04e4dbd8, upstream/main93d211f9ef4a869a9aed0d075ca767dda552627f). That is 89 commits; the first-parent merges are:fieldtobi/qmd#956 metadata filter field, Metadata discovery tobi/qmd#951 metadata discoverymatchc18b22chas parents2c8b6d2(fork main) and93d211f(upstream/main).package.json,bun.lock,pnpm-lock.yaml,flake.nixor.github/.git diff origin/main HEADon those paths is empty, so the lockfiles and the flake node-modules hash didn't need regenerating.Conflicts
There was one, in
src/store.ts→expandQuery, at the cache-format migration. I took upstream's behavior: both legacy-format branches now pass throughuniqueExpansions(...). The code keeps the fork's strict-mode style (r["query"],r["type"],r["text"],rows[0]?.["query"]). No other files conflicted.Strict-mode fixes (
1d06ffe, separate commit after the merge)These bring upstream code under the fork's
@systemfsoftware/tsconfigbase. Runtime behavior is unchanged.noPropertyAccessFromIndexSignature: switchedRecord<string, unknown>reads to bracket access inparseCliMetadataOptions(src/cli/qmd.ts) and in the REST/metadatahandler (src/mcp/server.ts).exactOptionalPropertyTypes: added| undefinedto everyListMetadataOptionsfield (src/metadata-store.ts) and to thereindexCollectionInoptions (src/store.ts). The consumers already treatundefinedas absent (??,=== undefined, truthiness).@ts-ignore,@ts-expect-errororas any, and no tests were skipped.Fork commits preserved
.github/(checked with grep formacos|darwin, no matches).origin/main.flake.nixstill has theaarch64-darwin/x86_64-darwinhash entries and thepkgs.darwin.cctoolsline, the same as forkmainbefore this sync. This merge didn't add them: upstream never touchedflake.nixin this range, and chore: drop macOS from CI (hosted larger-runner quota) #1 and chore: adopt org tsgo stack (TS7 + @systemfsoftware/tsconfig base + @effect/language-service) #2 never removed them. Taking them out would be a fork-only change outside a sync, so it should go in its own PR if it's wanted.Local gates (Linux x64, node v24.19.0, bun 1.4.0)
./node_modules/.bin/tsc -p tsconfig.build.json --noEmit(TypeScript 7.0.2)1d06ffethere were 36 errors (30 TS4111, 4 TS2379, 1 TS2345, 1 TS2375).npx oxlint(1.78.0)Found 0 warnings and 0 errors.on 94 files, exit 0CI=true npx vitest run --testTimeout 60000 test/CI=true LD_LIBRARY_PATH=/usr/lib/x86_64-linux-gnu bun test --timeout 60000 --preload ./src/test-preload.ts test/npm run buildnix build .result/bin/qmd --version→qmd 2.8.3The first vitest run had 4 failures. All of them came from the local container, not from this merge. Neither failing test file is in the upstream diff.
test/release-context.test.ts(3 tests):jqwasn't on the subprocess PATH. I put jq from the repo's locked nixpkgs (nix build --inputs-from . nixpkgs#jq) on PATH.test/llm.test.ts›canWriteLlamaDir … not writable: the suite ran as root withCAP_DAC_OVERRIDE. I re-ran undersetpriv --bounding-set -dac_override,-dac_read_search. GitHub runners don't run tests as root.The vitest and bun numbers in the table come from the full re-run with those two fixes in place.
Notes:
bun install --frozen-lockfilefails at thebetter-sqlite3install script becausenode-gypis missing on this host. I installed withnpm installinstead.bun.lockis unchanged frommain.qmd --versionfrom the nix build also prints*** stack smashing detected ***on exit. A nix build of forkmain(2c8b6d2) prints the same thing, so it was there before this PR.