diff --git a/.gitignore b/.gitignore index 165336e8c..ff2c477e8 100644 --- a/.gitignore +++ b/.gitignore @@ -11,6 +11,7 @@ texts/ *.md !README.md !CLAUDE.md +!DESIGN.md !CHANGELOG.md !skills/**/*.md !finetune/*.md diff --git a/CHANGELOG.md b/CHANGELOG.md index 2cb32c5f2..a89c471d5 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,17 +2,130 @@ ## [Unreleased] +### Fixed + +- `qmd doctor` selects vector sample identities before loading document bodies, + avoiding excessive SQLite memory use on large indexes with duplicate paths. + #978 (thanks @naveenspark) +- Vector diagnostics match passages by their saved character position, avoiding + false mismatches when earlier chunks change the sequence numbering. + #978 (thanks @naveenspark) + ### Added - Added Oxlint lint fence. -- Document metadata and metadata filtering. Markdown documents can opt into typed metadata through a namespaced frontmatter block (`qmd.metadata` with strings, numbers, booleans, or flat homogeneous arrays), and every search surface — CLI `search`/`vsearch`/`query` via `--filter `, the SDK's `filter` option on `search()`/`searchLex()`/`searchVector()`, the MCP `query` tool, and HTTP `POST /query` and `/search` — accepts one shared recursive filter AST discriminated by `operator`: `and`/`or`/`not` logical groups, `eq`/`ne`/`gt`/`gte`/`lt`/`lte` comparisons, `in`/`nin`/`all` membership, and `exists` presence. Every returned result satisfies the filter (applied before RRF fusion and reranking); like collection filtering, highly selective filters remain best-effort for top-K completeness. Frontmatter stays ordinary searchable content — no chunking, embedding, snippet, or line-number changes — and documents without `qmd.metadata` behave exactly as before. JSON/SDK/MCP/HTTP results now include each document's indexed metadata, and `qmd status` reports how many documents still need metadata extraction (a normal `qmd update` backfills existing indexes). +- Document metadata and metadata filtering. Markdown documents can opt into typed metadata through a namespaced frontmatter block (`qmd.metadata` with strings, numbers, booleans, or flat homogeneous arrays), and every search surface — CLI `search`/`vsearch`/`query` via `--filter `, the SDK's `filter` option on `search()`/`searchLex()`/`searchVector()`, the MCP `query` tool, and HTTP `POST /query` and `/search` — accepts one shared recursive filter AST discriminated by `operator`: `and`/`or`/`not` logical groups, `eq`/`ne`/`gt`/`gte`/`lt`/`lte` comparisons, `in`/`nin`/`all` membership, `contains`/`prefix`/`suffix` text matching, `type` for the stored type of a key's values, and `exists` presence, with an optional `caseInsensitive` flag on any condition whose value is a string (ASCII folding). Every returned result satisfies the filter (applied before RRF fusion and reranking); like collection filtering, highly selective filters remain best-effort for top-K completeness. Frontmatter stays ordinary searchable content — no chunking, embedding, snippet, or line-number changes — and documents without `qmd.metadata` behave exactly as before. JSON/SDK/MCP/HTTP results now include each document's indexed metadata, and `qmd status` reports how many documents still need metadata extraction (a normal `qmd update` backfills existing indexes). +- Metadata discovery. Filtering is only useful when the caller knows what to filter on, so every surface now reports the metadata keys, types, and value counts already in the index — read from the existing metadata tables, with no re-indexing. The CLI adds `qmd collection metadata [name...]` with `--filter ` to count only documents matching a filter (same AST as search) and `--match ` to select which metadata entries are reported (the same AST evaluated per entry, where a condition's `field` is the entry's `key` or `value`, so one grammar covers a single key, a family of keys, reverse lookup of a value, typed thresholds, and any `and`/`or`/`not` composition), plus `--key-limit`/`--key-offset`/`--all-keys` and `--value-limit`/`--value-offset`/`--all-values` to page both windows, and `--sort count|value` and `--min-count ` to order and trim values. Every windowed list reports its exact remainder, counts are documents rather than values, numbers report min/median/max, keys whose documents disagree on type are split per type with their own counts, and with `--filter` every coverage count is measured against the documents the filter admits. `qmd collection list` names each collection's top keys, `qmd collection show` details them with a value preview, and `qmd status` summarizes coverage. The SDK adds `listMetadata(options)` (the same options, spelled `keyLimit`, `valueLimit`, and so on, with `MetadataOptionError` for a value outside its domain and `MetadataBindingBudgetError` for a filter and match that together bind too many SQL parameters), the MCP server adds a `metadata` tool and lists each collection's most covered keys in `status`, and the HTTP server adds `POST /metadata`. `MetadataFilter` and `MetadataMatch` are both `MetadataPredicate`, one recursive grammar over the conditions a document or a metadata entry admits. ### Fixed +- Filtered vector search binds candidate IDs as one JSON list, so a parser-valid metadata filter cannot exhaust Node's SQL variable limit during document lookup. Applies to both exact scans and the capped global fallback. +- Collection- and metadata-scoped BM25 search (`qmd search -c`, the lex leg of + `qmd query`, MCP, SDK `searchLex`) no longer returns false-empty or + incomplete results when stronger matches outside the scope fill the old + `limit * 10` candidate window (#922). The scope is now applied to the full + FTS5 match set, materialized once, so `search -c ` is exact for + common terms; unscoped search keeps its early-terminating plan. A search over + several collections, such as the default ones, runs one keyword query instead + of one per collection, which made keyword search about 12x faster over 13 + collections (thanks @brettdavies). Builds on the approach in #918 (thanks + @fxstein). #953 (thanks @Mr-Beasley) + +- `qmd update` no longer deactivates every document in a collection whose root + folder is missing, for example an unmounted drive or an offline network share. + The folder globbed to an empty list, which read as "every file was deleted", + so search went empty and a `qmd cleanup` before the drive came back deleted + the rows for good. The collection is now reported as not found and its index + is left unchanged. Permission and I/O errors still fail with their original + cause (#989). #990 (thanks @ParkerRex) - Embedding generation and legacy fingerprint adoption now tokenize documents with the store-selected embedding model instead of the global default. This keeps chunk boundaries aligned with the model that creates and verifies the stored vectors without initializing an unrelated provider. +- `qmd doctor` no longer stalls or fills the temp directory on large indexes + when checking legacy (empty-fingerprint) embeddings. The adoption sample + query joined `content` and grouped by the document body, so SQLite + materialized the full body once per legacy chunk and per active path before + `LIMIT 1` discarded it, the same pattern as the doctor vector-sample check + (#978). It now picks the sample row through indexes and loads only that + row's body; the sampled chunk is unchanged (#994). #995 (thanks @mjaverto) + +### Changed + +- The MCP server caches its instructions per store for up to 60 seconds instead + of rebuilding them, with a full index-status scan, for every HTTP request. + Concurrent requests share one build and a failed build is not cached; index + changes from other processes show up in the instructions within a minute. + #815 (thanks @fxstein) +- Avoid a redundant runtime process on packaged CLI calls where the launcher is + already running under its selected Node or Bun runtime. #923 (thanks @ilepn) +- `qmd update` no longer reads files whose modification time and size match + the last pass (tracked in a new `file_sync_state` table), so re-indexing an + unchanged collection takes seconds. A cached entry counts only while its + document is still active at that path with that content, so a collection + removed and added back is re-indexed, while a renamed one keeps its entries; + files with missing or outdated metadata are still re-extracted. `qmd update` + and `qmd collection add` skip files over 10 MB with `FILE_TOO_LARGE`, and + `qmd update` deactivates a previously indexed file that becomes empty or is + over 10 MB, including one indexed by an earlier release. #962 (thanks + @rikvanriel) +- Search results carry at most the first 262,144 characters of each document + body (`qmd search --full`, MCP results, keyword and vector hits), which + bounds memory on indexes with very large documents. A match past that point + is still found, but its snippet, and the passage the reranker scores, come + from the start of the document. #962 (thanks @rikvanriel) +- `qmd update` and `qmd collection add` write a collection scan in short + transactions instead of committing every row on its own, so a first index + of 20,000 files takes about 5 s instead of about 3 minutes. #1021 (thanks + @brettdavies) + +- `qmd query` and `qmd vsearch` now drop repeated query expansions before + searching. The expansion model can repeat a line, and the cache kept every + copy: one reported query ran 23 expansions where 9 were distinct, and each + copy ran its own search. Cached expansions are deduplicated when read, so + existing caches need no rebuild. Repeated lex lines used to count more than + once in the rank fusion, so result order can shift slightly. (#921) + #1000 (thanks @ParkerRex) +- Structured searches over several collections (`qmd query` with + `lex:`/`vec:`/`hyde:` lines, the MCP `query` tool, SDK `queries`) run one + search per line over the whole collection list, a keyword search for a + `lex:` line and a vector search for a `vec:` or `hyde:` line, and fuse one + ranked list per search. Rankings no longer depend on the order the + collections are named. #946 (thanks @shalom-t), #1009 (thanks @xidus90) + +### Changed + +- The vector index is partitioned by collection. Each embedded chunk is + stored once per collection that holds it, under an integer collection id, + so a collection-scoped `qmd query` or `qmd vsearch` scans only that + collection's vectors instead of pre-filtering the whole index, and a small + collection is never crowded out of the results by a large one. The first + command after upgrading converts the index in place and prints its + progress; the conversion resumes from where it stopped if interrupted, two + commands started at once share it, and a VACUUM at the end reclaims the + space of the old table. `qmd collection rename` no longer touches vectors, + `qmd update` removes, as it finishes, the vectors of content that no + indexed document holds any more (content that comes back later is embedded + again), and `qmd update` and `qmd embed` copy an already-embedded document + into a collection that gains it instead of embedding it again. A + metadata-filtered vector search returns the nearest documents the filter + admits however many chunks it admits, and a vector search fills its limit + even when one long document holds all of the nearest chunks. Downgrading + to an older qmd afterwards needs `qmd embed -f`, and a `qmd mcp` server + started before the upgrade must be restarted. #983 (thanks @brettdavies) + +### Changed + +- `qmd cleanup` repacks the vector table when fewer than 90% of its chunk + reads hold live rows. vec0 fills the first free slot of its newest chunk + and reclaims a chunk only once it is empty, so re-embeds and orphan removal + leave holes in older chunks that every brute-force scan still reads; a + 794k-vector index at 36% occupancy scanned twice as slowly as a packed copy + of the same rows. The repack moves the live rows of each mostly-empty chunk + to the tail, one chunk per short transaction, so it never holds the write + lock for long and an interrupted run leaves a consistent table. The cleanup + output and `qmd cleanup --dry-run` report the chunk counts. + #937 (thanks @brettdavies) ## [2.8.3] - 2026-08-16 diff --git a/CLAUDE.md b/CLAUDE.md index ee952e4a0..429fd56e4 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -9,6 +9,7 @@ qmd collection add . --name # Create/index collection qmd collection list # List all collections with details qmd collection remove # Remove a collection by name qmd collection rename # Rename a collection +qmd collection metadata [name...] # Discover metadata keys, types, and value counts (--match, --filter, --key-limit, --value-limit) qmd init # Create a project-local .qmd index qmd ls [collection[/path]] # List collections or files in a collection qmd context add [path] "text" # Add context for path (defaults to current dir) @@ -47,9 +48,16 @@ qmd collection remove mynotes # Rename a collection qmd collection rename mynotes my-notes -# Show collection details +# Show collection details, including the top metadata keys qmd collection show mynotes +# Discover metadata keys and values to filter on +qmd collection metadata mynotes +qmd collection metadata mynotes --match '{"field":"key","operator":"eq","value":"topics"}' +qmd collection metadata mynotes --match '{"field":"value","operator":"eq","value":"docs-team"}' +qmd collection metadata mynotes --match '{"field":"key","operator":"eq","value":"topics"}' --filter '{"field":"status","operator":"eq","value":"published"}' +qmd collection metadata mynotes --key-limit 50 --key-offset 50 # next page of keys + # Set or clear the pre-update hook (runs before re-indexing on `qmd update`) qmd collection update-cmd mynotes 'git pull --ff-only' qmd collection update-cmd mynotes # clear diff --git a/DESIGN.md b/DESIGN.md new file mode 100644 index 000000000..618822490 --- /dev/null +++ b/DESIGN.md @@ -0,0 +1,12 @@ +# QMD interface + +QMD is a command-line search tool, library, and MCP server. This change has no +graphical surface. Existing command output and structured result formats are +the observed interface contracts; see README.md and test/cli.test.ts. + +The doctor vector check samples stored chunks and compares freshly generated +embeddings with their saved vectors. Its sampling must preserve active-document, +model, and fingerprint filters without expanding full document bodies across +all candidate chunks. No visual or typography changes are part of this work. +The saved character position identifies the passage to re-embed; a stored +sequence number is the vector key and may differ from today's chunk ordering. diff --git a/README.md b/README.md index 76b822b20..b89af6960 100644 --- a/README.md +++ b/README.md @@ -94,7 +94,8 @@ Although the tool works perfectly fine when you just tell your agent to use it o - `query` — Search with typed sub-queries (`lex`/`vec`/`hyde`), combined via RRF + reranking - `get` — Retrieve a document by path or docid (with fuzzy matching suggestions) - `multi_get` — Batch retrieve by glob pattern, comma-separated list, or docids -- `status` — Index health and collection info +- `status` — Index health and collection info, including each collection's top metadata keys +- `metadata` — Discover metadata keys, types, and value counts to filter on **Claude Desktop configuration** (`~/Library/Application Support/Claude/claude_desktop_config.json`): @@ -152,6 +153,7 @@ runs in a container and a liveness probe connects from a non-loopback address. The HTTP server exposes two endpoints: - `POST /mcp` — MCP Streamable HTTP (JSON responses, stateless) - `POST /query` (alias `/search`) — structured search without the MCP protocol. Accepts the same optional `filter` object as the `query` tool (invalid filters return `400`); see [Metadata Filtering](#metadata-filtering) +- `POST /metadata` — metadata discovery without the MCP protocol. Same body as the `metadata` tool (invalid filters return `400`); see [Metadata Discovery](#metadata-discovery) - `GET /health` — liveness check with uptime @@ -201,6 +203,15 @@ Point any MCP client at `http://localhost:8181/mcp` to connect. | `multi_get` | `maxBytes` | number | Skip files larger than N (default 10240) | | `multi_get` | `maxLines` | number | Limit lines per file | | `multi_get` | `lineNumbers` | boolean | Prefix lines with numbers (default **true**) | +| `metadata` | `collections` | string[] | Restrict discovery to collection names (default: the collections `query` searches) | +| `metadata` | `match` | object | Report only metadata entries matching this condition (same AST as `filter`, with `field` naming the entry's `key` or `value`) | +| `metadata` | `filter` | object | Count only documents matching this filter (same AST as `query`) | +| `metadata` | `keyLimit` | number | Keys reported (default 50). `totalKeys` and `remainingKeys` describe the rest | +| `metadata` | `keyOffset` | number | Keys skipped before the window, in report order (default 0) | +| `metadata` | `valueLimit` | number | Values reported per key and type (default 10). `remainingValues` reports the rest | +| `metadata` | `valueOffset` | number | Values skipped per key and type before the window, in `sort` order (default 0) | +| `metadata` | `sort` | string | `count` (default) or `value` | +| `metadata` | `minCount` | number | Hide values held by fewer documents (default 1) | Unknown parameters are silently ignored (not rejected) — double-check names if results seem unscoped. The HTTP `/query` and `/search` endpoints return @@ -302,8 +313,8 @@ const published = await store.search({ filter: { operator: "and", operands: [ - { key: "topics", operator: "all", value: ["typescript"] }, - { key: "status", operator: "ne", value: "draft" }, + { field: "topics", operator: "all", value: ["typescript"] }, + { field: "status", operator: "ne", value: "draft" }, ], }, }) @@ -644,9 +655,14 @@ qmd collection rename myproject my-project qmd ls notes qmd ls notes/subfolder -# Show collection details (path, glob mask, include status, context count) +# Show collection details (path, glob mask, include status, context count, top metadata keys) qmd collection show notes +# Discover metadata keys, types, and value counts (see Metadata Discovery) +qmd collection metadata notes +qmd collection metadata notes --match '{"field":"key","operator":"eq","value":"topics"}' +qmd collection metadata notes --match '{"field":"value","operator":"eq","value":"docs-team"}' + # Include or exclude a collection from default (unscoped) queries qmd collection include notes qmd collection exclude notes @@ -940,24 +956,25 @@ qmd: Supported values are strings, numbers, booleans, and flat homogeneous arrays of one of those. Nested objects, nulls, empty arrays, and mixed-type arrays are rejected (the document still indexes; it is excluded from filtered search until corrected). Metadata keys are user-defined data — `tags`, `topics`, and `labels` are all ordinary keys with no special semantics. -Every search surface (CLI, SDK, MCP, HTTP) accepts the same recursive filter, a JSON AST discriminated by `operator`: +Every search surface (CLI, SDK, MCP, HTTP) accepts the same recursive filter, a JSON AST discriminated by `operator`. A condition is a predicate over one field of the document's metadata: `field` names the metadata key, `operator` says how to compare, and `value` is what to compare against: ```sh # One condition qmd search "authentication" \ - --filter '{"key":"status","operator":"eq","value":"published"}' + --filter '{"field":"status","operator":"eq","value":"published"}' # Composed conditions — works with search, vsearch, and query qmd query "dependency injection" --filter '{ "operator": "and", "operands": [ - { "key": "topics", "operator": "all", "value": ["typescript", "programming"] }, - { "key": "status", "operator": "nin", "value": ["draft", "archived"] }, + { "field": "topics", "operator": "all", "value": ["typescript", "programming"] }, + { "field": "status", "operator": "nin", "value": ["draft", "archived"] }, { "operator": "or", "operands": [ - { "key": "priority", "operator": "gte", "value": 3 }, - { "key": "reviewed", "operator": "eq", "value": true } + { "field": "priority", "operator": "gte", "value": 3 }, + { "field": "reviewed", "operator": "eq", "value": true } ] }, - { "operator": "not", "operand": { "key": "audience", "operator": "eq", "value": "internal" } } + { "field": "topics", "operator": "prefix", "value": "sql", "caseInsensitive": true }, + { "operator": "not", "operand": { "field": "audience", "operator": "eq", "value": "internal" } } ] }' ``` @@ -966,25 +983,210 @@ qmd query "dependency injection" --filter '{ |------|-------| | Logical group | `{ "operator": "and" \| "or", "operands": […] }` | | Negation | `{ "operator": "not", "operand": {…} }` | -| Comparison | `{ "key", "operator": "eq" \| "ne" \| "gt" \| "gte" \| "lt" \| "lte", "value" }` | -| Membership | `{ "key", "operator": "in" \| "nin" \| "all", "value": […] }` | -| Presence | `{ "key", "operator": "exists", "value": true \| false }` | +| Comparison | `{ "field", "operator": "eq" \| "ne" \| "gt" \| "gte" \| "lt" \| "lte", "value" }` | +| Membership | `{ "field", "operator": "in" \| "nin" \| "all", "value": […] }` | +| Text | `{ "field", "operator": "contains" \| "prefix" \| "suffix", "value": "…" }` | +| Type | `{ "field", "operator": "type", "value": "string" \| "number" \| "boolean" }` | +| Presence | `{ "field", "operator": "exists", "value": true \| false }` | + +Any condition whose value is a string or an array of strings may add `"caseInsensitive": true`. Semantics: -- Matching is typed and exact — no string/number/boolean coercion, and a type mismatch never matches (including `ne` and `nin`). -- Array-valued metadata is a set: a condition matches when any element satisfies it, `all` requires every filter value to be present. +- Matching is typed and exact — no string/number/boolean coercion, and a type mismatch never matches (including `ne` and `nin`). Text operators match string values only. `type` matches the stored type of a key's values, which is how a filter reaches one side of a key whose documents disagree on type. +- Array-valued metadata is a set: a condition matches when any element satisfies it, `all` requires every filter value to be present, and `ne`/`nin` require that no element equals the operand (with at least one element of the operand's type present). - Missing keys do not match `ne`/`nin`; combine with `{ "operator": "exists", "value": false }` in an `or` group to include them. +- Matching is case-sensitive unless a condition sets `caseInsensitive`, which folds ASCII letters on both sides. Non-ASCII letters compare exactly. - Multiple conditions require an explicit `and` group — there is no implicit AND, and no `$`-prefixed shorthand. Guarantees and limits: - Every returned result satisfies the filter, before RRF fusion and reranking. -- Like collection filtering, highly selective filters are best-effort for top-K completeness: backends over-fetch and post-filter, so a very selective filter can return fewer than `limit` results. +- Highly selective filters can return fewer than `limit` results. Lexical search filters a bounded over-fetch window; vector search scans the filter-eligible set exactly when it holds at most 20,000 chunk vectors, and above that over-fetches and post-filters. - Filtered search only considers documents whose metadata has been extracted (run `qmd update` after upgrading; `qmd status` shows the pending count). JSON output (`--format json`), the SDK, MCP structured results, and the HTTP endpoints include each result's indexed metadata. +### Metadata Discovery + +Filtering is only useful if you know what to filter on. Discovery reports the metadata keys, types, and value counts already in the index, so a filter can be written from what is indexed instead of guessed. It reads the same tables filtering reads: no re-indexing, and every value it reports is one an `eq` filter can match. + +The command is one sentence: show metadata matching X for documents filtered by Y. `--filter` narrows **documents** by their metadata (same AST as search) and decides which are counted. `--match` narrows **the metadata itself** and decides which entries are reported. It takes the same AST, evaluated against each metadata entry instead of each document: a condition's `field` is `"key"` (the entry's key name) or `"value"` (its value). Every operator applies, including `type`, the text operators, `caseInsensitive`, and `and`/`or`/`not`. Only `exists` and `all` are rejected, as they have no meaning for a single entry. + +| `--match` | Question answered | +|-----------|-------------------| +| | Which keys exist, with a window of values each | +| `{"field":"key","operator":"eq","value":"topics"}` | Everything about one key | +| `{"field":"key","operator":"prefix","value":"mem-"}` | A family of keys | +| `{"field":"key","operator":"in","value":["tags","topics","labels"]}` | Which of these key names exist | +| `{"field":"value","operator":"eq","value":"docs-team"}` | Which keys hold this value | +| `{"field":"value","operator":"prefix","value":"2025-"}` | Which keys hold values shaped like this | +| `{"field":"value","operator":"type","value":"boolean"}` | Which keys hold booleans | +| `and` of `key eq priority` and `value gte 3` | Values of one key above a threshold | +| `and` of `key eq priority` and `value type number` | The numeric side of a key whose documents disagree on type | + +Start wide and narrow: + +```sh +# Which keys does this collection use? (`qmd collection show notes` previews the top five) +qmd collection metadata notes + +# Everything about one key: coverage, distinct count, top values +qmd collection metadata notes --match '{"field":"key","operator":"eq","value":"topics"}' + +# Reverse lookup: which keys hold this value +qmd collection metadata notes --match '{"field":"value","operator":"eq","value":"docs-team"}' + +# Values of one key matching a pattern +qmd collection metadata notes --match '{ + "operator": "and", + "operands": [ + { "field": "key", "operator": "eq", "value": "topics" }, + { "field": "value", "operator": "prefix", "value": "sql" } + ] +}' + +# What remains after a filter, before committing to it in a query +qmd collection metadata notes \ + --match '{"field":"key","operator":"eq","value":"topics"}' \ + --filter '{"field":"status","operator":"eq","value":"published"}' + +# Then search with the filter you just validated +qmd query "dependency injection" -c notes --filter '{ + "operator": "and", + "operands": [ + { "field": "status", "operator": "eq", "value": "published" }, + { "field": "topics", "operator": "all", "value": ["typescript"] } + ] +}' +``` + +The drill-down prints one block per key, in coverage order: + +```sh +qmd collection metadata notes +``` + +``` +topics string[] 388 of 480 documents 1,204 distinct + typescript 140 + sqlite 92 + search 77 + architecture 61 + sqlite-vec 44 + mcp 39 + embeddings 35 + agents 31 + cli 28 + testing 26 +1,194 more values, use --value-limit , --value-offset , or --all-values + +priority number 205 of 480 documents 5 distinct + min 1 median 3 max 5 + 1 (12) 2 (40) 3 (88) 4 (50) 5 (15) + +reviewed boolean 480 of 480 documents + true 61 false 419 +``` + +With `--filter`, the output opens with a `filter:` line stating how many documents pass, and every coverage count is measured against that population rather than the whole collection — the numbers a filtered search would see: + +```sh +qmd collection metadata notes \ + --match '{"field":"key","operator":"eq","value":"topics"}' \ + --filter '{"field":"status","operator":"eq","value":"published"}' \ + --value-limit 5 +``` + +``` +filter: 312 of 480 documents + +topics string[] 260 of 312 documents 811 distinct + typescript 104 + sqlite 70 + search 58 + architecture 40 + mcp 31 +806 more values, use --value-limit , --value-offset , or --all-values +``` + +Two windows page the key and value lists. Each windowed list ends with the exact remainder and the flags that reach it: + +```sh +--key-limit # Keys reported, in coverage order (default 50) +--key-offset # Keys skipped before the window +--all-keys # Remove the key window +--value-limit # Values reported per key and type (default 10) +--value-offset # Values skipped per key and type, in --sort order +--all-values # Remove the value window +--sort count|value # Order values by document count (default) or by value +--min-count # Drop values held by fewer documents +``` + +Omitting the collection name covers the default collections, exactly as an unscoped search does. + +Reading the output: + +- **Counts are documents, not values.** A document with `topics: [a, b]` contributes one to each. Coverage is "documents declaring this key", out of the documents the filter admits when there is one. +- **A bare string is exactly the value.** A string whose bare form could be read as something else (`"42"`, `""`, `"a, b"`) prints as a JSON string, so paste it into a filter as the JSON string it is. +- **Numbers report min, median, and max**, plus the enumerated values when they fit, which is enough to write a `gt`/`lt` threshold in one call. +- **Discovery sees exactly what filtering sees.** Same extraction gate, same active-document rule, same collection scope, and every count in a result comes from one database snapshot. Documents still pending extraction are reported on stderr and excluded until `qmd update` runs. +- **Type conflicts are reported, not resolved.** Metadata is validated one document at a time, so `priority: 3` in one file and `priority: high` in another both index, within one collection or across several. Discovery splits such a key by type, each with its own document count and (across collections) its contributing collections: + +```sh +qmd collection metadata --match '{"field":"key","operator":"eq","value":"priority"}' +``` + +``` +priority number | string 1,427 of 2,100 documents + number 223 docs min 1 median 3 max 5 notes, work + string 1,204 docs high (700), medium (380), low (124) work +``` + +To report one side only, add a `type` condition on the value field. The filter that reaches exactly those documents is the same condition with the metadata key in `field`: + +```sh +qmd collection metadata --match '{ + "operator": "and", + "operands": [ + { "field": "key", "operator": "eq", "value": "priority" }, + { "field": "value", "operator": "type", "value": "number" } + ] +}' +qmd query "release checklist" --filter '{"field":"priority","operator":"type","value":"number"}' +``` + +`qmd collection list` names each collection's top keys, `qmd collection show ` details the top five with a value preview, and `qmd status` summarizes how many keys and files carry metadata. + +The CLI prints text and rejects unsupported format flags. Structured discovery is available through the SDK, MCP, and HTTP with the same options (`collection`, `match`, `filter`, `keyLimit`, `keyOffset`, `valueLimit`, `valueOffset`, `sort`, `minCount`) and the same result shape: + +```typescript +// SDK: one flat result, keys in coverage order, each split per type +const discovery = await store.listMetadata({ + collection: "notes", + match: { field: "key", operator: "eq", value: "topics" }, + valueLimit: 5, +}) +discovery.documents // active documents in scope +discovery.totalKeys // keys with a matching entry, before the key window +discovery.remainingKeys // keys after the window +discovery.keys[0].types[0] // { type, multiValued, documents, distinctValues, values, remainingValues, range?, collections } + +// Narrowed by a filter: how many documents pass, and what is left to filter on +const published = await store.listMetadata({ + collection: "notes", + filter: { field: "status", operator: "eq", value: "published" }, +}) +published.filteredDocuments // the denominator for every coverage count in this result + +// Page through a wide vocabulary +const nextPage = await store.listMetadata({ keyLimit: 50, keyOffset: 50 }) +``` + +`filter` is a `MetadataFilter` and `match` is a `MetadataMatch`. Both are `MetadataPredicate`, the one recursive grammar over the conditions each record admits, so a document-only condition (`exists`, `all`, or a metadata key as the `field`) is a type error in a match as well as a runtime one. An option outside its domain throws `MetadataOptionError`, and a `filter` and `match` that together bind more SQL parameters than one statement allows throw `MetadataBindingBudgetError`. + +The MCP `metadata` tool takes the same options with `collections` spelled as on `query`, returns the CLI text plus the result as `structuredContent`, and the MCP `status` tool lists each collection's most covered keys so an agent's first call reveals that metadata exists. `POST /metadata` accepts the same body as the tool and returns the same result (`400` on an invalid match, filter, or option). + ### Output Format Default output is colorized CLI format (respects `NO_COLOR` env). @@ -1133,7 +1335,7 @@ qmd multi-get "docs/*.md" --max-bytes 20480 # Output multi-get as JSON for agent processing qmd multi-get "docs/*.md" --json -# Clean up cache and orphaned data +# Drop caches and orphans; repack the vector table when it has holes qmd cleanup ``` @@ -1213,13 +1415,15 @@ Index stored in: `~/.cache/qmd/index.sqlite` ### Schema ```sql -collections -- Indexed directories with name and glob patterns -path_contexts -- Context descriptions by virtual path (qmd://...) -documents -- Markdown content with metadata and docid (6-char hash) -documents_fts -- FTS5 full-text index -content_vectors -- Embedding chunks (hash, seq, pos, 900 tokens each) -vectors_vec -- sqlite-vec vector index (hash_seq key) -llm_cache -- Cached LLM responses (query expansion, rerank scores) +collections -- Indexed directories with name and glob patterns +path_contexts -- Context descriptions by virtual path (qmd://...) +documents -- Markdown content with metadata and docid (6-char hash) +documents_fts -- FTS5 full-text index +content_vectors -- Embedding chunks (hash, seq, pos, 900 tokens each) +vector_collection_ids -- Integer id per collection name (the vector partition key) +vector_rows -- Vector rowid to (hash, seq, collection_id) +vectors_by_collection -- sqlite-vec vector index, one row per chunk and collection +llm_cache -- Cached LLM responses (query expansion, rerank scores) ``` ## Environment Variables diff --git a/bin/qmd b/bin/qmd index 3ae0309d6..981663ab6 100755 --- a/bin/qmd +++ b/bin/qmd @@ -11,7 +11,7 @@ import { spawn, spawnSync } from "node:child_process"; import { existsSync, realpathSync } from "node:fs"; import { dirname, resolve, sep } from "node:path"; -import { fileURLToPath } from "node:url"; +import { fileURLToPath, pathToFileURL } from "node:url"; // Resolve symlinks so global installs (npm link / npm install -g) can find // the actual package directory instead of the global bin directory. @@ -160,33 +160,47 @@ if (existsSync(resolve(pkgDir, "package-lock.json"))) { } const selected = useSourceMode ? sourceRunner : (runnerName === "node" ? "node" : "bun"); -// Pin Node to the binary that launched this trampoline. Native addons -// (better-sqlite3) are compiled for that install's ABI; spawning PATH -// `node` picks up nvm/fnm/mise shims in the project directory and fails -// with NODE_MODULE_VERSION / ERR_DLOPEN_FAILED (leftover from #577; #319). -// If the trampoline itself is running under bun (`bun bin/qmd`), keep -// PATH `node` so we don't load Node-ABI addons into bun. -const runner = selected === "node" && typeof process.versions.bun !== "string" - ? process.execPath - : selected; -const args = useSourceMode ? sourceArgs : [jsEntry, ...process.argv.slice(2)]; -const needsShell = (runner === "bun") && process.platform === "win32"; - -const child = spawn(runner, args, { - stdio: "inherit", - shell: needsShell, -}); - -child.on("exit", (code, signal) => { - if (signal) { - process.kill(process.pid, signal); - } else { - process.exit(code ?? 0); - } -}); - -child.on("error", (err) => { - const name = useSourceMode ? sourceRunner : runnerName; - console.error(`qmd: failed to launch ${name}: ${err.message}`); - process.exit(1); -}); +const launchedByBun = typeof process.versions.bun === "string"; +const runtimeAlreadySelected = !useSourceMode + && ((selected === "node" && !launchedByBun) || (selected === "bun" && launchedByBun)); + +// Published packages can reach this launcher under the runtime selected by +// their lockfile. Starting that same runtime again adds an avoidable process +// startup to matching-runtime CLI calls. Align argv with the compiled entry +// point and import it in place; the CLI's main-module guard then behaves +// exactly as in a direct run. +if (runtimeAlreadySelected) { + process.argv.splice(1, 1, jsEntry); + await import(pathToFileURL(jsEntry).href); +} else { + const args = useSourceMode ? sourceArgs : [jsEntry, ...process.argv.slice(2)]; + // Pin Node to the binary that launched this trampoline. Native addons + // (better-sqlite3) are compiled for that install's ABI; spawning PATH + // `node` picks up nvm/fnm/mise shims in the project directory and fails + // with NODE_MODULE_VERSION / ERR_DLOPEN_FAILED (leftover from #577; #319). + // If the trampoline itself is running under bun (`bun bin/qmd`), keep + // PATH `node` so we don't load Node-ABI addons into bun. + const runner = selected === "node" && !launchedByBun + ? process.execPath + : selected; + const needsShell = (runner === "bun") && process.platform === "win32"; + + const child = spawn(runner, args, { + stdio: "inherit", + shell: needsShell, + }); + + child.on("exit", (code, signal) => { + if (signal) { + process.kill(process.pid, signal); + } else { + process.exit(code ?? 0); + } + }); + + child.on("error", (err) => { + const name = useSourceMode ? sourceRunner : runnerName; + console.error(`qmd: failed to launch ${name}: ${err.message}`); + process.exit(1); + }); +} diff --git a/scripts/package-smoke.mjs b/scripts/package-smoke.mjs index f9622e7a4..2a7ddae8f 100644 --- a/scripts/package-smoke.mjs +++ b/scripts/package-smoke.mjs @@ -57,9 +57,17 @@ assertPath("dist/cli/qmd.js", "compiled CLI"); run("compiled CLI under Node", process.execPath, ["dist/cli/qmd.js", "--help"], { quiet: true }); run("package wrapper", "sh", ["bin/qmd", "--help"], { quiet: true }); +run("dist package wrapper under Node", process.execPath, ["bin/qmd", "--help"], { + quiet: true, + env: { ...process.env, QMD_SOURCE_MODE: "0" }, +}); if (process.env.QMD_SKIP_BUN_SMOKE === "1") { console.log("==> compiled CLI under Bun (skipped by QMD_SKIP_BUN_SMOKE=1)"); } else { run("compiled CLI under Bun", "bun", ["dist/cli/qmd.js", "--help"], { quiet: true }); + run("dist package wrapper under Bun", "bun", ["bin/qmd", "--help"], { + quiet: true, + env: { ...process.env, QMD_SOURCE_MODE: "0" }, + }); } diff --git a/skills/qmd/SKILL.md b/skills/qmd/SKILL.md index 7693d4273..e80a3dbbd 100644 --- a/skills/qmd/SKILL.md +++ b/skills/qmd/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: Requires qmd CLI or MCP server. Install via `npm install -g @tobilu/qmd`. metadata: author: tobi - version: "2.2.0" + version: "2.3.0" allowed-tools: Bash(qmd:*), mcp__qmd__* --- @@ -201,14 +201,70 @@ Omit `-c` to search everything. ## Filter by metadata -Documents can carry typed metadata in a `qmd.metadata` frontmatter block (strings, numbers, booleans, or flat arrays). `search`, `vsearch`, and `query` accept `--filter` with a recursive JSON AST; every returned result satisfies it: +Documents can carry typed metadata in a `qmd.metadata` frontmatter block +(strings, numbers, booleans, or flat arrays). `search`, `vsearch`, and `query` +accept `--filter` with a recursive JSON AST, and every returned result satisfies +it. Add a metadata filter when it expresses a constraint the user intends, not +by default, and check coverage first: a condition on a key's value only sees +documents that declare that key, and discovery (below) shows how many do. ```bash -qmd search "authentication" --filter '{"key":"status","operator":"eq","value":"published"}' -qmd query "dependency injection" --filter '{"operator":"and","operands":[{"key":"topics","operator":"all","value":["typescript"]},{"key":"status","operator":"nin","value":["draft","archived"]}]}' +qmd search "authentication" --filter '{"field":"status","operator":"eq","value":"published"}' +qmd query "dependency injection" --filter '{"operator":"and","operands":[{"field":"topics","operator":"all","value":["typescript"]},{"field":"status","operator":"nin","value":["draft","archived"]}]}' ``` -Nodes are discriminated by `operator`: groups `and`/`or` take `operands`, `not` takes one `operand`, and conditions take `key` + `value` with operators `eq`/`ne`/`gt`/`gte`/`lt`/`lte` (comparison), `in`/`nin`/`all` (membership), or `exists` (presence). Matching is typed and exact; missing keys do not match `ne`/`nin` (add an `exists: false` branch in an `or` group to include them). The MCP `query` tool accepts the same AST as a `filter` object. JSON output includes each result's `metadata`. +Nodes are discriminated by `operator`: groups `and`/`or` take `operands`, `not` +takes one `operand`, and conditions take `field` + `value` with operators +`eq`/`ne`/`gt`/`gte`/`lt`/`lte` (comparison), `in`/`nin`/`all` (membership), +`contains`/`prefix`/`suffix` (text), `type` (the value is a `string`, `number`, +or `boolean`), or `exists` (presence). Matching is typed and exact; missing keys +do not match `ne`/`nin` (add an `exists: false` branch in an `or` group to +include them). Conditions with a string value may add `"caseInsensitive": true`, +which folds ASCII letters. The MCP `query` tool accepts the same AST as a +`filter` object. JSON output includes each result's `metadata`. + +## Discover metadata before filtering + +Do not guess keys or values. `qmd collection show ` lists the top keys +with types and a value preview, and `qmd collection metadata` drills in. Same +metadata, different unit: `--filter` narrows documents (which are counted), +`--match` narrows the metadata itself (which entries are reported). Both take +the predicate AST above. In a match, a condition's `field` is `"key"` (the +entry's key name) or `"value"` (its value), and every operator applies except +`exists` and `all`: + +```bash +qmd collection metadata notes # top keys, top values +qmd collection metadata notes --match '{"field":"key","operator":"eq","value":"topics"}' # one key in depth +qmd collection metadata notes --match '{"field":"value","operator":"eq","value":"docs-team"}' # which keys hold this value +qmd collection metadata notes --match '{"operator":"and","operands":[{"field":"key","operator":"eq","value":"priority"},{"field":"value","operator":"gte","value":3}]}' +qmd collection metadata notes --match '{"field":"key","operator":"eq","value":"topics"}' --filter '{"field":"status","operator":"eq","value":"published"}' +``` + +The last form shows what a filter leaves behind before you commit to it in a +query. Read the output like this: + +- **The header is the decision.** `topics string[] 388 of 480 documents + 1,204 distinct` gives coverage and cardinality. With `--filter`, a + `filter: 312 of 480 documents` line opens the output and every count is out + of those 312. Counts are documents, not values. +- **Quotes mark a string that reads like something else.** `"42"` is a string, + `42` a number, `""` is empty. Paste a quoted value into a filter as the JSON + string it is. +- **Footers mean windowed.** `N more values` or `N more keys` means the list + was cut. Page with `--value-offset`/`--key-offset` or raise + `--value-limit`/`--key-limit` rather than assuming the rest. +- **Numbers print `min`, `median`, `max`**, enough to write a `gt`/`lt` + threshold in one call. +- **`number | string` means documents disagree on type.** Each type reports its + own document count. Filter by the type that covers the documents you want. + +Every value shown can be matched with `eq` under the same collection scope. + +Over MCP, call the `metadata` tool (same options, `collections` as an array) +and read `totalKeys`, `remainingKeys`, `remainingValues`, and `range` from the +structured result. Page keys with `keyOffset` and values with `valueOffset`. +The `status` tool lists each collection's most covered keys, so check it first. ## MCP Tool: `query` @@ -294,6 +350,9 @@ server configuration. have. You expand the query; the model just ranks. - **Do not overuse semantic search.** If you know exact titles or terms, BM25 is faster and often better. +- **Do not filter blind.** Run `qmd collection metadata` first and read the + coverage. A key declared by 388 of 480 documents leaves 92 that no value + condition on it can reach; only `exists: false` selects them. - **Do not mutate indexes casually.** `qmd collection add`, `qmd update`, and `qmd embed` change local state and can be expensive. - **Model-backed commands can be environment-sensitive.** If `qmd query`, diff --git a/src/cli/embed-lock.ts b/src/cli/embed-lock.ts index ac95f5568..5fdf9f2db 100644 --- a/src/cli/embed-lock.ts +++ b/src/cli/embed-lock.ts @@ -1,10 +1,11 @@ /** * Process-level exclusive lock for `qmd embed`. * - * Concurrent embed runs against the same index can race on vectors_vec - * (UNIQUE constraint on hash_seq). This lockfile keeps a second process from - * starting while another embed holds the lock. Stale files left by crashed - * processes are recovered via PID identity checks (same spirit as mcp-pid.ts). + * Concurrent embed runs against the same index can race on vector_rows + * (UNIQUE constraint on hash, seq, collection_id). This lockfile keeps a + * second process from starting while another embed holds the lock. Stale + * files left by crashed processes are recovered via PID identity checks + * (same spirit as mcp-pid.ts). */ import { existsSync, readFileSync, unlinkSync, writeFileSync } from "node:fs"; diff --git a/src/cli/qmd.ts b/src/cli/qmd.ts index b7b05a152..992de916b 100644 --- a/src/cli/qmd.ts +++ b/src/cli/qmd.ts @@ -28,6 +28,7 @@ import { resolveCommaListName, matchFilesByGlob, getHashesNeedingEmbedding, + getEmbeddingVectorSamples, clearAllEmbeddings, insertEmbedding, getStatus, @@ -56,9 +57,12 @@ import { deactivateDocument, getActiveDocumentPaths, cleanupOrphanedContent, + cleanupOrphanedVectors, + copyVectorsToNewCollections, countOrphanedVectors, previewCleanup, runCleanup, + type VectorTableLayout, getCollectionsWithoutContext, getTopLevelPathsWithoutContext, handelize, @@ -80,15 +84,29 @@ import { createStore, getDefaultDbPath, reindexCollection, + scanWriteBatch, + REINDEX_MAX_FILE_SIZE, generateEmbeddings, maybeAdoptLegacyEmbeddingFingerprint, syncConfigToDb, type ReindexResult, type ChunkStrategy, } from "../store.js"; -import { syncDocumentMetadata, countDocumentsPendingMetadata } from "../metadata-store.js"; +import { + syncDocumentMetadata, + countDocumentsPendingMetadata, + countDocumentsWithMetadata, + countMetadataKeys, + listMetadata, + listMetadataCollectionSummaries, + MetadataBindingBudgetError, + type ListMetadataOptions, + type ListMetadataResult, +} from "../metadata-store.js"; +import { formatMetadataKeySummaries, formatMetadataOverview } from "../metadata-format.js"; import type { DocumentMetadata } from "../metadata.js"; -import { parseMetadataFilter, type MetadataFilter } from "../metadata-filter.js"; +import { parseMetadataFilter, parseMetadataMatch, type MetadataFilter, type MetadataMatch } from "../metadata-filter.js"; +import { hasVectorIndex, storedEmbedding } from "../vec-layout.js"; import { disposeDefaultLlamaCpp, getDefaultLlamaCpp, setDefaultLlamaCpp, LlamaCpp, withLLMSession, pullModels, DEFAULT_MODEL_CACHE_DIR, resolveEmbedModel, resolveGenerateModel, resolveRerankModel, resolveModels, inspectGgufFile, isDarwinMetalMitigationActive } from "../llm.js"; import { formatSearchResults, @@ -511,14 +529,6 @@ function sanitizeDiagnosticMessage(message: string): string { .join("; "); } -/** Hint after `qmd update` when orphaned embedding chunks exceed this share of vectors (#768). */ -const ORPHAN_VECTOR_HINT_RATIO = 0.1; - -function formatOrphanedVectorHint(orphaned: number, total: number): string { - const pct = total > 0 ? Math.round((orphaned / total) * 100) : 0; - return `${orphaned} orphaned embedding chunks (${pct}% of vectors) — run 'qmd cleanup' to reclaim space`; -} - async function showStatus(): Promise { const dbPath = getDbPath(); const db = getDb(); @@ -574,6 +584,10 @@ async function showStatus(): Promise { if (needsEmbedding > 0) { console.log(` ${c.yellow}Pending: ${needsEmbedding} need embedding${c.reset} (run 'qmd embed')`); } + const metadataKeyCount = countMetadataKeys(db); + if (metadataKeyCount > 0) { + console.log(` Metadata: ${metadataKeyCount} keys across ${countDocumentsWithMetadata(db)} files (explore with 'qmd collection metadata')`); + } const pendingMetadata = countDocumentsPendingMetadata(db); if (pendingMetadata > 0) { console.log(` ${c.yellow}Metadata: ${pendingMetadata} need extraction${c.reset} (run 'qmd update'; excluded from --filter searches)`); @@ -1004,19 +1018,32 @@ async function updateCollections(): Promise { console.log(""); } + // The pending count below only sees content_vectors, so a hash that joined + // a collection while already embedded elsewhere would be neither counted + // nor searchable there; copying its rows closes that gap without a model. + // The copy runs before the cleanup, which deletes the partition rows the + // copy reads from, so a document moved between collections keeps its vectors. + const copiedVectors = copyVectorsToNewCollections(db).copied; + // A changed file rewrites its document's hash in place, which strands the + // old hash's rows in the collection's vector partition; the partition + // filter cannot see documents.active, so those rows would take k slots + // from a scoped search until they are removed. + const staleVectors = cleanupOrphanedVectors(db); + // Check if any documents need embedding (show once at end) const needsEmbedding = getHashesNeedingEmbedding(db); - const vectorTotal = (db.prepare(`SELECT COUNT(*) as count FROM content_vectors`).get() as { count: number }).count; - const orphanedVectors = countOrphanedVectors(db); closeDb(); console.log(`${c.green}✓ All collections updated.${c.reset}`); + if (staleVectors > 0) { + console.log(`Removed ${staleVectors} stale vector row(s)`); + } + if (copiedVectors > 0) { + console.log(`Copied ${copiedVectors} vector(s) into collections that gained already-embedded documents`); + } if (needsEmbedding > 0) { console.log(`\nRun 'qmd embed' to update embeddings (${needsEmbedding} unique hashes need vectors)`); } - if (vectorTotal > 0 && orphanedVectors / vectorTotal >= ORPHAN_VECTOR_HINT_RATIO) { - console.log(`\n${formatOrphanedVectorHint(orphanedVectors, vectorTotal)}`); - } } /** @@ -1794,6 +1821,7 @@ function collectionList(): void { } console.log(`${c.bold}Collections (${collections.length}):${c.reset}\n`); + const metadataOverviews = listMetadataCollectionSummaries(db, COLLECTION_LIST_METADATA_KEYS); for (const coll of collections) { const updatedAt = coll.last_modified ? new Date(coll.last_modified) : new Date(); @@ -1810,6 +1838,12 @@ function collectionList(): void { console.log(` ${c.dim}Ignore:${c.reset} ${yamlColl.ignore.join(', ')}`); } console.log(` ${c.dim}Files:${c.reset} ${coll.active_count}`); + const metadataOverview = metadataOverviews.get(coll.name); + if (metadataOverview) { + const shownKeys = metadataOverview.keys.map(overview => overview.key); + const hiddenKeys = metadataOverview.totalKeys - shownKeys.length; + console.log(` ${c.dim}Metadata:${c.reset} ${shownKeys.join(', ')}${hiddenKeys > 0 ? `, +${hiddenKeys} more` : ''}`); + } console.log(` ${c.dim}Updated:${c.reset} ${timeAgo}`); console.log(); } @@ -1817,6 +1851,26 @@ function collectionList(): void { closeDb(); } +/** Key names shown on `collection list`, and keys detailed on `collection show`. */ +const COLLECTION_LIST_METADATA_KEYS = 5; + +// The Metadata section of `collection show`: top keys by coverage with a +// value preview, and a pointer at the drill-down for the rest. +function collectionShowMetadata(name: string): void { + const db = getDb(); + const result = listMetadata(db, { collection: name, keyLimit: COLLECTION_LIST_METADATA_KEYS, valueLimit: 3 }); + const documentsWithMetadata = countDocumentsWithMetadata(db, [name]); + const pendingMetadata = countDocumentsPendingMetadata(db, [name]); + closeDb(); + + console.log(formatMetadataOverview(result, { + documentsWithMetadata, + pendingMetadata, + drillDownHint: `qmd collection metadata ${name}`, + colors: c, + })); +} + /** Canonical --mask, with --glob as the alias OpenClaw and others already pass (#536). */ function collectionGlobFromCli(values: { mask?: unknown; glob?: unknown }): string { const mask = typeof values.mask === "string" && values.mask.length > 0 ? values.mask : undefined; @@ -1915,6 +1969,128 @@ function collectionRename(oldName: string, newName: string): void { console.log(` Virtual paths updated: ${c.cyan}qmd://${oldName}/${c.reset} → ${c.cyan}qmd://${newName}/${c.reset}`); } +// Metadata discovery drill-down. Collection names are already validated; +// an empty list means the default collections resolved to nothing, which +// the store reads as "every collection" exactly as search does. +function collectionMetadata(collectionNames: string[], options: ListMetadataOptions): void { + const db = getDb(); + + if (listCollections(db).length === 0) { + console.log("No collections found. Run 'qmd collection add .' to create one."); + closeDb(); + return; + } + + // Discovery always applies the extraction gate, filter or not. + warnPendingMetadata(db, collectionNames); + + let result: ListMetadataResult; + try { + result = listMetadata(db, { ...options, collection: collectionSearchFilter(collectionNames) }); + } catch (error) { + if (!(error instanceof MetadataBindingBudgetError)) throw error; + closeDb(); + console.error(`${c.yellow}${error.message}${c.reset}`); + process.exit(1); + } + closeDb(); + + const selection = options.match || options.filter; + console.log(formatMetadataKeySummaries(result, { + showCollections: collectionNames.length !== 1, + valueWindowHint: "--value-limit , --value-offset , or --all-values", + keyWindowHint: "--key-limit , --key-offset , or --all-keys", + keyOffset: options.keyOffset ?? 0, + keyOffsetLabel: "--key-offset", + emptyMessage: selection + ? "No metadata matches. Run 'qmd collection metadata' without --match or --filter to see which keys exist." + : "No metadata found. Add qmd.metadata frontmatter and run 'qmd update'.", + colors: c, + })); +} + +// Parse the discovery-specific flags; exits with usage on a bad value. +function parseCliMetadataOptions(values: Record): ListMetadataOptions { + const options: ListMetadataOptions = { + match: parseCliMetadataMatch(values["match"]), + filter: parseCliMetadataFilter(values["filter"]), + }; + + // Discovery windows two dimensions, so the single-window search flags + // have no reading here. Point at the flags that do. + if (values["n"] !== undefined) { + console.error("-n is not an option of 'qmd collection metadata'"); + console.error("Use --value-limit for values per key, or --key-limit for keys"); + process.exit(1); + } + if (values["all"]) { + console.error("--all is not an option of 'qmd collection metadata'"); + console.error("Use --all-values, --all-keys, or both"); + process.exit(1); + } + + const formatAlias = ["json", "csv", "md", "xml", "files"].find(flag => values[flag]); + const format = typeof values["format"] === "string" ? values["format"].trim().toLowerCase() : undefined; + if (formatAlias || (format !== undefined && format !== "cli")) { + console.error(`${formatAlias ? `--${formatAlias}` : `--format ${String(values["format"])}`} is not supported by 'qmd collection metadata'`); + console.error("This command prints text. Use the SDK, MCP metadata tool, or POST /metadata for structured output"); + process.exit(1); + } + + if (values["all-keys"]) { + options.keyLimit = Infinity; + } else if (values["key-limit"] !== undefined) { + options.keyLimit = parsePositiveInteger(values["key-limit"], "--key-limit"); + } + if (values["key-offset"] !== undefined) { + options.keyOffset = parseNonNegativeInteger(values["key-offset"], "--key-offset"); + } + + if (values["all-values"]) { + options.valueLimit = Infinity; + } else if (values["value-limit"] !== undefined) { + options.valueLimit = parsePositiveInteger(values["value-limit"], "--value-limit"); + } + if (values["value-offset"] !== undefined) { + options.valueOffset = parseNonNegativeInteger(values["value-offset"], "--value-offset"); + } + + if (values["min-count"] !== undefined) { + options.minCount = parsePositiveInteger(values["min-count"], "--min-count"); + } + + if (values["sort"] !== undefined) { + if (values["sort"] !== "count" && values["sort"] !== "value") { + console.error(`Invalid --sort value: ${String(values["sort"])}`); + console.error("Valid: count, value"); + process.exit(1); + } + options.sort = values["sort"]; + } + + return options; +} + +function parsePositiveInteger(raw: unknown, flag: string): number { + const parsed = Number(raw); + if (!Number.isSafeInteger(parsed) || parsed < 1) { + console.error(`Invalid ${flag} value: ${String(raw)}`); + console.error(`${flag} must be a positive safe integer`); + process.exit(1); + } + return parsed; +} + +function parseNonNegativeInteger(raw: unknown, flag: string): number { + const parsed = Number(raw); + if (!Number.isSafeInteger(parsed) || parsed < 0) { + console.error(`Invalid ${flag} value: ${String(raw)}`); + console.error(`${flag} must be a non-negative safe integer`); + process.exit(1); + } + return parsed; +} + async function indexFiles(pwd?: string, globPattern: string = DEFAULT_GLOB, collectionName?: string, suppressEmbedNotice: boolean = false, ignorePatterns?: string[]): Promise { const db = getDb(); const resolvedPwd = pwd || getPwd(); @@ -1965,7 +2141,9 @@ async function indexFiles(pwd?: string, globPattern: string = DEFAULT_GLOB, coll const livePaths = new Set(files.map(f => f.replace(/\\/g, '/'))); const startTime = Date.now(); + const batch = scanWriteBatch(db); for (const relativeFile of files) { + batch.next(); const filepath = getRealPath(resolve(resolvedPwd, relativeFile)); // Store the literal relative path — handelize() is NOT applied at index time. const path = relativeFile.replace(/\\/g, '/'); @@ -1975,6 +2153,14 @@ async function indexFiles(pwd?: string, globPattern: string = DEFAULT_GLOB, coll progress.set((processed / total) * 100); continue; } + let tooLarge = false; + try { tooLarge = statSync(filepath).size > REINDEX_MAX_FILE_SIZE; } catch { /* the read below reports it */ } + if (tooLarge) { + processed++; + skippedFiles.push({ file: relativeFile, code: "FILE_TOO_LARGE" }); + progress.set((processed / total) * 100); + continue; + } seenPaths.add(path); let content: string; @@ -2051,13 +2237,17 @@ async function indexFiles(pwd?: string, globPattern: string = DEFAULT_GLOB, coll let removed = 0; for (const path of allActive) { if (!seenPaths.has(path)) { + batch.next(); deactivateDocument(db, collectionName, path); removed++; } } + batch.commit(); + // Clean up orphaned content hashes (content not referenced by any document) const orphanedContent = cleanupOrphanedContent(db); + const copiedVectors = copyVectorsToNewCollections(db, collectionName).copied; // Check if vector index needs updating const needsEmbedding = getHashesNeedingEmbedding(db); @@ -2069,6 +2259,9 @@ async function indexFiles(pwd?: string, globPattern: string = DEFAULT_GLOB, coll if (orphanedContent > 0) { console.log(`Cleaned up ${orphanedContent} orphaned content hash(es)`); } + if (copiedVectors > 0) { + console.log(`Copied ${copiedVectors} vector(s) into collections that gained already-embedded documents`); + } if (needsEmbedding > 0 && !suppressEmbedNotice) { console.log(`\nRun 'qmd embed' to update embeddings (${needsEmbedding} unique hashes need vectors)`); @@ -2092,16 +2285,24 @@ function reportMetadataErrors(metadataErrors: number): void { function reportSkippedReads(skippedFiles: { file: string; code: string }[]): void { if (skippedFiles.length === 0) return; + const sizeLimitMb = Math.round(REINDEX_MAX_FILE_SIZE / (1024 * 1024)); for (const skipped of skippedFiles) { - if (skipped.code === "OUTSIDE_COLLECTION") { + if (skipped.code === "ROOT_MISSING") { + console.warn(`⚠ Collection root not found, index left unchanged: ${skipped.file}`); + } else if (skipped.code === "OUTSIDE_COLLECTION") { console.warn(`⚠ Skipped file outside collection: ${skipped.file}`); + } else if (skipped.code === "FILE_TOO_LARGE") { + console.warn(`⚠ Skipped file over ${sizeLimitMb} MB: ${skipped.file}`); } else { console.warn(`⚠ Skipped unreadable file: ${skipped.file} (${skipped.code})`); } } const escaped = skippedFiles.filter(f => f.code === "OUTSIDE_COLLECTION").length; - const unreadable = skippedFiles.length - escaped; + const tooLarge = skippedFiles.filter(f => f.code === "FILE_TOO_LARGE").length; + const rootMissing = skippedFiles.filter(f => f.code === "ROOT_MISSING").length; + const unreadable = skippedFiles.length - escaped - tooLarge - rootMissing; if (escaped) console.warn(`Skipped ${escaped} file(s) outside the collection root`); + if (tooLarge) console.warn(`Skipped ${tooLarge} file(s) over ${sizeLimitMb} MB`); if (unreadable) console.warn(`Skipped ${unreadable} unreadable file(s)`); } @@ -2193,7 +2394,7 @@ async function vectorIndex( const storeInstance = getStore(); const db = storeInstance.db; - // Exclusive process lock — concurrent embeds race on vectors_vec (#825) + // Exclusive process lock — concurrent embeds race on vector_rows (#825) const embedLock = tryAcquireEmbedLock(embedLockPathForDb(getDbPath())); if (!embedLock) { console.log(EMBED_LOCK_BUSY_MESSAGE); @@ -2265,6 +2466,9 @@ async function vectorIndex( const totalTimeSec = result.durationMs / 1000; + if (result.chunksCopied > 0) { + console.log(`${c.green}✓${c.reset} Copied ${formatCount(result.chunksCopied)} vectors into collections that gained already-embedded documents`); + } if (result.chunksEmbedded === 0 && result.docsProcessed === 0) { console.log(`${c.green}✓ No non-empty documents to embed.${c.reset}`); } else { @@ -2702,6 +2906,10 @@ function outputResults(results: OutputRow[], query: string, opts: OutputOptions) warnUnresolvedFullPaths(unresolvedCount, filtered.length); } +function describeVectorLayout(layout: VectorTableLayout): string { + return `${layout.chunks} chunks for ${layout.neededChunks} needed, ${Math.round(layout.occupancy * 100)}% useful`; +} + // Resolve -c collection filter: supports single string, array, or undefined. // Returns validated collection names (exits on unknown collection). function resolveCollectionFilter(raw: string | string[] | undefined, useDefaults: boolean = false): string[] { @@ -2824,29 +3032,54 @@ function parseStructuredQuery(query: string): ParsedStructuredQuery | null { // Parse and validate a --filter JSON string; exits with an actionable // message on malformed JSON or an invalid filter AST. function parseCliMetadataFilter(rawFilter: unknown): MetadataFilter | undefined { - if (rawFilter === undefined) return undefined; + return parseCliPredicateFlag(rawFilter, { + flag: "--filter", + example: `{"field":"status","operator":"eq","value":"published"}`, + parse: parseMetadataFilter, + }); +} - let filterJson: unknown; +// Same grammar as --filter, evaluated against metadata entries for discovery. +function parseCliMetadataMatch(rawMatch: unknown): MetadataMatch | undefined { + return parseCliPredicateFlag(rawMatch, { + flag: "--match", + example: `{"field":"key","operator":"eq","value":"topics"}`, + parse: parseMetadataMatch, + }); +} + +interface CliPredicateFlag { + flag: string; + example: string; + parse: (input: unknown) => Predicate; +} + +// Parse a JSON predicate flag with the parser for its record type. +function parseCliPredicateFlag(raw: unknown, predicateFlag: CliPredicateFlag): Predicate | undefined { + if (raw === undefined) return undefined; + + let astJson: unknown; try { - filterJson = JSON.parse(String(rawFilter)); + astJson = JSON.parse(String(raw)); } catch (err) { - console.error(`Invalid --filter JSON: ${err instanceof Error ? err.message : String(err)}`); - console.error(`Example: --filter '{"key":"status","operator":"eq","value":"published"}'`); + console.error(`Invalid ${predicateFlag.flag} JSON: ${err instanceof Error ? err.message : String(err)}`); + console.error(`Example: ${predicateFlag.flag} '${predicateFlag.example}'`); process.exit(1); } try { - return parseMetadataFilter(filterJson); + return predicateFlag.parse(astJson); } catch (err) { console.error(err instanceof Error ? err.message : String(err)); process.exit(1); } } -// Filtered search excludes documents without current metadata extraction; -// tell the user when that makes results incomplete. -function warnPendingMetadata(db: Database): void { - const pendingMetadata = countDocumentsPendingMetadata(db); +// Filtered search and discovery exclude documents without current metadata +// extraction; tell the user when that makes results incomplete, scoped to +// the collections the command reads. +function warnPendingMetadata(db: Database, collectionNames: string[]): void { + const pendingMetadata = countDocumentsPendingMetadata(db, collectionNames.length > 0 ? collectionNames : undefined); if (pendingMetadata === 0) return; process.stderr.write(`${c.yellow}Warning: ${pendingMetadata} document(s) lack current metadata extraction and are excluded from filtered results. Run 'qmd update'.${c.reset}\n`); } @@ -2858,7 +3091,7 @@ function search(query: string, opts: OutputOptions): void { // Use default collections if none specified const collectionNames = resolveCollectionFilter(opts.collection, true); - if (opts.filter) warnPendingMetadata(db); + if (opts.filter) warnPendingMetadata(db, collectionNames); // Use large limit for --all, otherwise fetch more than needed and let outputResults filter const fetchLimit = opts.all ? 100000 : Math.max(50, opts.limit * 2); @@ -2909,7 +3142,7 @@ async function vectorSearch(query: string, opts: OutputOptions, _model: string = const collectionNames = resolveCollectionFilter(opts.collection, true); checkIndexHealth(store.db); - if (opts.filter) warnPendingMetadata(store.db); + if (opts.filter) warnPendingMetadata(store.db, collectionNames); await withLLMSession(async () => { let results = await vectorSearchQuery(store, query, { @@ -2954,7 +3187,7 @@ async function querySearch(query: string, opts: OutputOptions, _embedModel: stri const collectionNames = resolveCollectionFilter(opts.collection, true); checkIndexHealth(store.db); - if (opts.filter) warnPendingMetadata(store.db); + if (opts.filter) warnPendingMetadata(store.db, collectionNames); // Check for structured query syntax (lex:/vec:/hyde:/intent: prefixes) const parsed = parseStructuredQuery(query); @@ -3111,7 +3344,17 @@ function parseCLI() { json: { type: "boolean" }, explain: { type: "boolean" }, collection: { type: "string", short: "c", multiple: true }, // Filter by collection(s) - filter: { type: "string" }, // Metadata filter (JSON AST) for search/vsearch/query + filter: { type: "string" }, // Metadata filter (JSON AST) for search/vsearch/query/collection metadata + // Metadata discovery options (collection metadata) + match: { type: "string" }, // Metadata match (JSON AST) over the entries reported + "key-limit": { type: "string" }, // keys reported (default 50) + "key-offset": { type: "string" }, // keys skipped before the window + "all-keys": { type: "boolean" }, // remove the key window + "value-limit": { type: "string" }, // values reported per key (default 10) + "value-offset": { type: "string" }, // values skipped per key before the window + "all-values": { type: "boolean" }, // remove the value window + sort: { type: "string" }, // count (default) | value + "min-count": { type: "string" }, // drop values held by fewer documents // Collection options name: { type: "string" }, // collection name mask: { type: "string" }, // glob pattern @@ -3651,6 +3894,7 @@ function showHelp(): void { console.log(""); console.log("Collections & context:"); console.log(" qmd collection add/list/remove/rename/show - Manage indexed folders"); + console.log(" qmd collection metadata [name] [--match J] - Discover metadata keys and values to filter on"); console.log(" qmd context add/list/rm - Attach human-written summaries"); console.log(" qmd ls [collection[/path]] - Inspect indexed files"); console.log(""); @@ -3729,7 +3973,7 @@ function showHelp(): void { console.log(" --format - Output format: cli (default) | json | csv | md | xml | files"); console.log(" -c, --collection - Filter by one or more collections"); console.log(" --filter - Metadata filter (recursive JSON AST; search/vsearch/query)"); - console.log(" e.g. '{\"key\":\"status\",\"operator\":\"eq\",\"value\":\"published\"}'"); + console.log(" e.g. '{\"field\":\"status\",\"operator\":\"eq\",\"value\":\"published\"}'"); console.log(""); console.log("Embed/query options:"); console.log(" --chunk-strategy - Chunking mode (default: regex; auto uses AST for code files)"); @@ -3997,27 +4241,17 @@ function checkModelCache(activeModels: { embed: string; generate: string; rerank } } -async function checkEmbeddingVectorSamples(db: Database, model: string, fingerprint: string, sampleSize: number = 3): Promise { +export async function checkEmbeddingVectorSamples(db: Database, model: string, fingerprint: string, sampleSize: number = 3): Promise { const activeDocs = (db.prepare(`SELECT COUNT(*) AS count FROM documents WHERE active = 1`).get() as { count: number }).count; if (activeDocs === 0) { return { ok: true, details: "no active documents indexed" }; } - const vecTableExists = db.prepare(`SELECT 1 FROM sqlite_master WHERE type='table' AND name='vectors_vec'`).get(); - if (!vecTableExists) { + if (!hasVectorIndex(db)) { return { ok: false, details: "no vector table to test; please run qmd embed again" }; } - const samples = db.prepare(` - SELECT cv.hash, cv.seq, c.doc AS body, MIN(d.path) AS path - FROM content_vectors cv - JOIN documents d ON d.hash = cv.hash AND d.active = 1 - JOIN content c ON c.hash = cv.hash - WHERE cv.model = ? AND cv.embed_fingerprint = ? - GROUP BY cv.hash, cv.seq, c.doc - ORDER BY random() - LIMIT ? - `).all(model, fingerprint, sampleSize) as { hash: string; seq: number; body: string; path: string }[]; + const samples = getEmbeddingVectorSamples(db, model, fingerprint, sampleSize); if (samples.length === 0) { return { ok: false, details: "no current embedded chunks to test; please run qmd embed again" }; @@ -4030,7 +4264,9 @@ async function checkEmbeddingVectorSamples(db: Database, model: string, fingerpr for (const sample of samples) { const hashSeq = `${sample.hash}_${sample.seq}`; const chunks = await chunkDocumentByTokens(sample.body, undefined, undefined, undefined, sample.path, undefined, session.signal); - const chunk = chunks[sample.seq]; + // Sequence numbers identify stored vectors, but earlier chunks can split + // differently after a tokenizer/chunker change. Compare the saved passage. + const chunk = chunks.find(chunk => chunk.pos === sample.pos); if (!chunk) { mismatches.push(`${shortHashSeq(hashSeq)}: chunk no longer exists`); continue; @@ -4043,13 +4279,13 @@ async function checkEmbeddingVectorSamples(db: Database, model: string, fingerpr continue; } - const stored = db.prepare(`SELECT embedding FROM vectors_vec WHERE hash_seq = ?`).get(hashSeq) as { embedding: Uint8Array } | undefined; + const stored = storedEmbedding(db, sample.hash, sample.seq); if (!stored) { mismatches.push(`${shortHashSeq(hashSeq)}: stored vector missing`); continue; } - const distance = cosineDistance(result.embedding, decodeStoredEmbedding(stored.embedding)); + const distance = cosineDistance(result.embedding, decodeStoredEmbedding(stored)); if (distance > threshold) { mismatches.push(`${shortHashSeq(hashSeq)}: stored vector distance ${distance.toFixed(6)}`); } @@ -4643,6 +4879,16 @@ if (isMain) { const ctxCount = Object.keys(col.context).length; console.log(` Contexts: ${ctxCount}`); } + collectionShowMetadata(name); + break; + } + + case "metadata": { + // Positional names are optional; omitted means the default + // collections, as an unscoped search does. + const rawNames = cli.args.length > 1 ? cli.args.slice(1) : undefined; + const collectionNames = resolveCollectionFilter(rawNames, true); + collectionMetadata(collectionNames, parseCliMetadataOptions(cli.values)); break; } @@ -4656,6 +4902,13 @@ if (isMain) { console.log(" remove Remove a collection"); console.log(" rename Rename a collection"); console.log(" show Show collection details"); + console.log(" metadata [name...] Discover metadata keys, types, and value counts"); + console.log(" --match Report only metadata entries matching this condition"); + console.log(" (same AST as --filter; 'field' is the entry's key or value)"); + console.log(" --filter Count only documents matching a metadata filter"); + console.log(" --key-limit Keys reported (default 50), --key-offset pages, --all-keys removes the window"); + console.log(" --value-limit Values per key (default 10), --value-offset pages, --all-values removes the window"); + console.log(" --sort count|value Value order (default count), --min-count drops the tail"); console.log(" update-cmd [cmd] Set pre-update command (e.g., 'git pull')"); console.log(" include Include in default queries"); console.log(" exclude Exclude from default queries"); @@ -4665,6 +4918,9 @@ if (isMain) { console.log(" qmd collection add ~/notes --name notes --mask 'a.md,journals/*.md'"); console.log(" qmd collection update-cmd brain 'git pull'"); console.log(" qmd collection exclude archive"); + console.log(" qmd collection metadata notes --match '{\"field\":\"key\",\"operator\":\"eq\",\"value\":\"topics\"}'"); + console.log(" qmd collection metadata notes --match '{\"field\":\"value\",\"operator\":\"eq\",\"value\":\"docs-team\"}'"); + console.log(" qmd collection metadata notes --match '{\"field\":\"key\",\"operator\":\"eq\",\"value\":\"topics\"}' --filter '{\"field\":\"status\",\"operator\":\"eq\",\"value\":\"published\"}'"); process.exit(0); } @@ -4994,12 +5250,19 @@ if (isMain) { if (stats.orphanedContent > 0) { console.log(`Would remove ${stats.orphanedContent} orphaned content hashes`); } + if (stats.vectorLayout) { + console.log(stats.vectorsRepacked + ? `Would repack the vector table (${describeVectorLayout(stats.vectorLayout)})` + : `${c.dim}Vector table is packed (${describeVectorLayout(stats.vectorLayout)})${c.reset}`); + } console.log("Would compact FTS and vacuum the database"); closeDb(); break; } - const stats = runCleanup(db); + const stats = runCleanup(db, { + onVectorRepack: (layout) => console.log(`Repacking the vector table (${describeVectorLayout(layout)})...`), + }); console.log(`${c.green}✓${c.reset} Cleared ${stats.cacheCount} cached API responses`); if (stats.orphanedVectors > 0) { console.log(`${c.green}✓${c.reset} Removed ${stats.orphanedVectors} orphaned embedding chunks`); @@ -5012,6 +5275,11 @@ if (isMain) { if (stats.orphanedContent > 0) { console.log(`${c.green}✓${c.reset} Removed ${stats.orphanedContent} orphaned content hashes`); } + if (stats.vectorLayout) { + console.log(stats.vectorsRepacked + ? `${c.green}✓${c.reset} Repacked the vector table (${describeVectorLayout(stats.vectorLayout)})` + : `${c.dim}Vector table is packed (${describeVectorLayout(stats.vectorLayout)})${c.reset}`); + } console.log(`${c.green}✓${c.reset} FTS compacted, database vacuumed`); closeDb(); diff --git a/src/db.ts b/src/db.ts index 09367fcca..8bef6bf9b 100644 --- a/src/db.ts +++ b/src/db.ts @@ -127,6 +127,8 @@ export function openDatabase(path: string): Database { * Common subset of the Database interface used throughout QMD. */ export interface Database { + /** Both drivers expose it; true while a transaction or savepoint is open. */ + readonly inTransaction: boolean; exec(sql: string): void; prepare(sql: string): Statement; loadExtension(path: string): void; diff --git a/src/index.ts b/src/index.ts index 37fbd8269..caa4c68d0 100644 --- a/src/index.ts +++ b/src/index.ts @@ -41,6 +41,7 @@ import { vacuumDatabase, cleanupOrphanedContent, cleanupOrphanedVectors, + copyVectorsToNewCollections, deleteLLMCache, deleteInactiveDocuments, clearAllEmbeddings, @@ -73,15 +74,34 @@ import type { MetadataScalar, MetadataScalarArray, MetadataValue, + MetadataValueType, } from "./metadata.js"; import { parseMetadataFilter, + parseMetadataMatch, MetadataFilterError, type MetadataFilter, type MetadataFilterGroup, type MetadataFilterNegation, type MetadataCondition, + type MetadataPredicate, + type MetadataPredicateGroup, + type MetadataPredicateNegation, + type MetadataMatch, + type MetadataEntryCondition, + type MetadataEntryField, } from "./metadata-filter.js"; +import { + listMetadata as storeListMetadata, + MetadataBindingBudgetError, + MetadataOptionError, + type ListMetadataOptions, + type ListMetadataResult, + type MetadataKeySummary, + type MetadataKeyTypeSummary, + type MetadataValueCount, + type MetadataKeyOverview, +} from "./metadata-store.js"; import { setConfigSource, loadConfig, @@ -129,12 +149,30 @@ export type { MetadataScalar, MetadataScalarArray, MetadataValue, + MetadataValueType, MetadataFilter, MetadataFilterGroup, MetadataFilterNegation, MetadataCondition, + MetadataPredicate, + MetadataPredicateGroup, + MetadataPredicateNegation, + MetadataMatch, + MetadataEntryCondition, + MetadataEntryField, +}; +export { parseMetadataFilter, parseMetadataMatch, MetadataFilterError }; + +// Re-export metadata discovery types (listMetadata() and status metadata keys) +export type { + ListMetadataOptions, + ListMetadataResult, + MetadataKeySummary, + MetadataKeyTypeSummary, + MetadataValueCount, + MetadataKeyOverview, }; -export { parseMetadataFilter, MetadataFilterError }; +export { MetadataBindingBudgetError, MetadataOptionError }; // Re-export the internal Store type for advanced consumers export type { InternalStore }; @@ -169,6 +207,10 @@ export type UpdateResult = { unchanged: number; removed: number; skipped: number; + /** Vector rows removed because their (hash, collection) no longer has an active document. */ + staleVectorsRemoved: number; + /** Vector rows copied into the partition of a collection that gained an already-embedded hash. */ + vectorsCopied: number; needsEmbedding: number; }; @@ -303,6 +345,22 @@ export interface QMDStore { /** List all collections with document stats */ listCollections(): Promise<{ name: string; pwd: string; glob_pattern: string; doc_count: number; active_count: number; last_modified: string | null; includeByDefault: boolean }[]>; + /** + * Discover metadata keys, types, and value counts across the documents in + * scope. `filter` selects which documents are counted. `match` selects + * which of their metadata entries are reported, with the filter grammar + * evaluated against each entry (a condition's `field` is the entry's `key` or + * `value`). Keys and values are windowed by `keyLimit`/`keyOffset` and + * `valueLimit`/`valueOffset`, and the result carries the totals and + * remainders needed to page. Every value reported is one an `eq` filter can + * match under the same scope, and the whole result comes from one database + * snapshot. Throws MetadataFilterError for an invalid predicate, + * MetadataOptionError for an option outside its domain, and + * MetadataBindingBudgetError when `filter` and `match` together bind more + * SQL parameters than one statement allows. + */ + listMetadata(options?: ListMetadataOptions): Promise; + /** Get names of collections included by default in queries */ getDefaultCollectionNames(): Promise; @@ -510,6 +568,11 @@ export async function createStore(options: StoreOptions): Promise { return result; }, listCollections: async () => storeListCollections(db), + listMetadata: async (opts) => storeListMetadata(db, { + ...opts, + match: opts?.match === undefined ? undefined : parseMetadataMatch(opts.match), + filter: opts?.filter === undefined ? undefined : parseMetadataFilter(opts.filter), + }), getDefaultCollectionNames: async () => { const collections = storeListCollections(db); return collections.filter(c => c.includeByDefault).map(c => c.name); @@ -564,6 +627,15 @@ export async function createStore(options: StoreOptions): Promise { totalSkipped += result.skipped; } + // A changed file rewrites its document's hash in place and strands the + // old hash's partition rows, which take k slots from scoped searches; + // a hash that joined a collection while embedded elsewhere stays + // unsearchable there until its rows are copied. The copy runs first: + // the cleanup deletes the partition rows it copies from, so a document + // moved between collections would otherwise need a fresh embed. + const vectorsCopied = copyVectorsToNewCollections(db).copied; + const staleVectorsRemoved = cleanupOrphanedVectors(db); + return { collections: filtered.length, indexed: totalIndexed, @@ -571,6 +643,8 @@ export async function createStore(options: StoreOptions): Promise { unchanged: totalUnchanged, removed: totalRemoved, skipped: totalSkipped, + staleVectorsRemoved, + vectorsCopied, needsEmbedding: internal.getHashesNeedingEmbedding(), }; }, diff --git a/src/maintenance.ts b/src/maintenance.ts index 5393555b3..9af91b431 100644 --- a/src/maintenance.ts +++ b/src/maintenance.ts @@ -15,8 +15,11 @@ import { clearAllEmbeddings, optimizeDocumentsFts, previewCleanup, + repackVectors, runCleanup, + vectorTableLayout, type CleanupStats, + type VectorTableLayout, } from "./store.js"; export class Maintenance { @@ -61,14 +64,25 @@ export class Maintenance { optimizeDocumentsFts(this.store.db); } + /** Chunk layout of the vector table, or null without one. */ + vectorLayout(): VectorTableLayout | null { + return vectorTableLayout(this.store.db); + } + + /** Move the live rows of mostly-empty vector chunks to the tail; see {@link repackVectors}. */ + repackVectors(): VectorTableLayout | null { + return repackVectors(this.store.db); + } + /** Preview what {@link run} would remove, without writing. */ preview(): CleanupStats { return previewCleanup(this.store.db); } /** - * Full cleanup: cache, orphaned vectors, inactive docs, orphaned content, - * FTS optimize, vacuum. Same sequence as `qmd cleanup`. + * Full cleanup: cache, orphaned vectors, vector repack when the table is + * mostly holes, inactive docs, orphaned content, FTS optimize, vacuum. Same + * sequence as `qmd cleanup`. */ run(): CleanupStats { return runCleanup(this.store.db); diff --git a/src/mcp/server.ts b/src/mcp/server.ts index d66888014..c222170bc 100644 --- a/src/mcp/server.ts +++ b/src/mcp/server.ts @@ -23,13 +23,20 @@ import { getDefaultDbPath, DEFAULT_MULTI_GET_MAX_BYTES, parseMetadataFilter, + parseMetadataMatch, type QMDStore, type ExpandedQuery, type IndexStatus, type DocumentMetadata, type MetadataFilter, + type MetadataMatch, + type MetadataKeyOverview, + type ListMetadataResult, + MetadataBindingBudgetError, + MetadataOptionError, } from "../index.js"; import { getConfigPath } from "../collections.js"; +import { formatMetadataKeySummaries } from "../metadata-format.js"; import { enableProductionMode } from "../store.js"; import { checkRequestOrigin, resolveOriginGuard } from "./origin-guard.js"; @@ -61,6 +68,28 @@ function validateFilterArgument(filter: unknown): { filter?: MetadataFilter; err } } +/** Same grammar as the filter, validated against metadata entries for discovery. */ +function validateMatchArgument(match: unknown): { match?: MetadataMatch; error?: string } { + if (match === undefined) return {}; + try { + return { match: parseMetadataMatch(match) }; + } catch (err) { + return { error: err instanceof Error ? err.message : String(err) }; + } +} + +/** Pick the named fields of a JSON body that must be numbers when present. */ +function readNumberFields(params: Record, names: string[]): { values: Record; error?: string } { + const values: Record = {}; + for (const name of names) { + const value = params[name]; + if (value === undefined) continue; + if (typeof value !== "number") return { values, error: `Invalid field: ${name} (must be a number)` }; + values[name] = value; + } + return { values }; +} + type StatusResult = { totalDocuments: number; needsEmbedding: number; @@ -71,6 +100,8 @@ type StatusResult = { pattern: string | null; documents: number; lastUpdated: string; + metadataKeyCount: number; + metadataKeys: MetadataKeyOverview[]; }[]; }; @@ -122,7 +153,9 @@ function getPackageVersion(): string { * searchable without a tool call. */ async function buildInstructions(store: QMDStore): Promise { - const status = await store.getStatus(); + // Instructions use index facts, not metadata overviews. Do not rank keys on + // initialization or on each stateless HTTP request just to discard them. + const status = store.internal.getStatusSummary(); const globalCtx = await store.getGlobalContext(); const lines: string[] = []; @@ -180,6 +213,37 @@ async function buildInstructions(store: QMDStore): Promise { return lines.join("\n"); } +/** + * Cache built instructions per store. The HTTP transport builds a fresh + * McpServer per request (sessionless MCP 2026-07-28), and buildInstructions() + * runs a full index-status aggregation, so rebuilding it per request puts that + * scan in front of every call. Concurrent builds share one in-flight promise, + * and entries expire after a short TTL so index updates from other processes + * still surface without a server restart. + */ +const INSTRUCTIONS_CACHE_TTL_MS = 60_000; +const instructionsCache = new WeakMap; builtAt: number }>(); + +function getInstructions(store: QMDStore): Promise { + const cached = instructionsCache.get(store); + // A negative age means the clock stepped back since the build; that entry is + // not known to be fresh, so it is rebuilt rather than kept past the TTL. + const age = cached ? Date.now() - cached.builtAt : -1; + if (cached && age >= 0 && age < INSTRUCTIONS_CACHE_TTL_MS) { + return cached.promise; + } + const promise = buildInstructions(store); + instructionsCache.set(store, { promise, builtAt: Date.now() }); + // Drop failed builds so the next initialize retries instead of replaying + // the cached rejection for the rest of the TTL window. + promise.catch(() => { + if (instructionsCache.get(store)?.promise === promise) { + instructionsCache.delete(store); + } + }); + return promise; +} + /** * Create an MCP server with all QMD tools, resources, and prompts registered. * Shared by both stdio and HTTP transports. @@ -191,7 +255,7 @@ async function createMcpServer(store: QMDStore, inflight?: InflightGate): Promis const server = new McpServer( { name: "qmd", version: getPackageVersion() }, { - instructions: await buildInstructions(store), + instructions: await getInstructions(store), // tools/list is static for the process lifetime; resources/read stays // uncacheable because the index can change under us. cacheHints: { @@ -348,11 +412,13 @@ Intent-aware lex (C++ performance, not sports): filter: z.record(z.string(), z.unknown()).optional().describe( "Metadata filter (recursive JSON AST). Every returned result satisfies it. " + "Nodes are operator-discriminated: logical groups {operator:'and'|'or', operands:[...]}, " + - "negation {operator:'not', operand:{...}}, and conditions {key, operator, value} with " + - "operators eq/ne/gt/gte/lt/lte (comparison), in/nin/all (membership), exists (presence). " + + "negation {operator:'not', operand:{...}}, and conditions {field, operator, value} with " + + "operators eq/ne/gt/gte/lt/lte (comparison), in/nin/all (membership), contains/prefix/suffix (text), " + + "type (value is 'string'|'number'|'boolean'), exists (presence). " + "Values are typed exactly (no coercion); missing keys do not match ne/nin. " + - "Example: {\"operator\":\"and\",\"operands\":[{\"key\":\"topics\",\"operator\":\"all\",\"value\":[\"typescript\"]}," + - "{\"key\":\"status\",\"operator\":\"ne\",\"value\":\"draft\"}]}" + "Conditions with a string value may add caseInsensitive:true (ASCII folding). " + + "Example: {\"operator\":\"and\",\"operands\":[{\"field\":\"topics\",\"operator\":\"all\",\"value\":[\"typescript\"]}," + + "{\"field\":\"status\",\"operator\":\"ne\",\"value\":\"draft\"}]}" ), intent: z.string().optional().describe( "Background context to disambiguate the query. Example: query='performance', intent='web page load times and Core Web Vitals'. Does not search on its own." @@ -608,6 +674,15 @@ Intent-aware lex (C++ performance, not sports): for (const col of status.collections) { summary.push(` - ${col.name}: ${col.path} (${col.documents} docs)`); + if (col.metadataKeys.length > 0) { + const keyLabels = col.metadataKeys.map(overview => `${overview.key} (${overview.types.join(" | ")})`); + const hiddenKeys = col.metadataKeyCount - col.metadataKeys.length; + summary.push(` metadata keys: ${keyLabels.join(", ")}${hiddenKeys > 0 ? `, +${hiddenKeys} more` : ""}`); + } + } + + if (status.collections.some(col => col.metadataKeys.length > 0)) { + summary.push(` Metadata: call the 'metadata' tool to see values and counts before writing a 'filter'`); } return { @@ -617,6 +692,104 @@ Intent-aware lex (C++ performance, not sports): }) ); + // --------------------------------------------------------------------------- + // Tool: qmd_metadata (Metadata discovery) + // --------------------------------------------------------------------------- + + server.registerTool( + "metadata", + { + title: "Metadata Discovery", + description: `Discover which metadata keys exist, what types they hold, and how many documents share each value, so you can write a \`filter\` for the query tool from what is indexed instead of guessing. + +## Mental model + +\`filter\` selects WHICH documents are counted. \`match\` selects WHICH metadata entries of those documents are reported. Both take the same recursive AST as the query tool's \`filter\`. A condition tests one \`field\` of the record under evaluation: for \`filter\` the record is a document and \`field\` names one of its metadata keys, for \`match\` the record is a metadata entry and \`field\` is \`"key"\` or \`"value"\`. Every operator applies to a match except \`exists\` and \`all\`, which have no meaning for a single entry. + +| match | Question answered | +|---|---| +| (none) | Which keys exist, with a window of values each | +| \`{"field":"key","operator":"eq","value":"topics"}\` | Everything about one key | +| \`{"field":"key","operator":"prefix","value":"mem-"}\` | A family of keys | +| \`{"field":"value","operator":"eq","value":"docs-team"}\` | Which keys hold this value | +| \`{"operator":"and","operands":[{"field":"key","operator":"eq","value":"priority"},{"field":"value","operator":"gte","value":3}]}\` | Values of one key above a threshold | + +Add \`filter\` to any of these to see what remains after narrowing. The result then reports \`filteredDocuments\`, and every coverage count is measured against that population. + +## Reading the result + +Counts are documents, not values. \`keys\` is one window of \`totalKeys\` in coverage order, and \`remainingKeys\` is the exact count after it. Each key splits by type: metadata is validated per document, so a key can hold numbers in some files and strings in others, and each type reports its own \`documents\`, \`distinctValues\`, a window of \`values\`, \`remainingValues\`, and, for numbers, \`range\` (min, median, max). Page keys with \`keyOffset\` and values with \`valueOffset\`, which applies to every key in the result and so reads best after \`match\` narrows to one key. + +Every value reported here can be matched with \`{field: '', operator: 'eq', value}\` under the same collections.`, + annotations: { readOnlyHint: true, openWorldHint: false }, + inputSchema: z.object({ + collections: z.array(z.string()).optional().describe("Restrict to these collections (default: the same collections query searches)"), + match: z.record(z.string(), z.unknown()).optional().describe( + "Report only metadata entries matching this condition. Same recursive AST as 'filter', evaluated per entry: " + + "a condition's field is 'key' (the entry's key name) or 'value' (its value). Default: every entry. " + + "Example: {\"field\":\"key\",\"operator\":\"eq\",\"value\":\"topics\"}" + ), + filter: z.record(z.string(), z.unknown()).optional().describe( + "Count only documents matching this metadata filter. Same recursive AST as the query tool's 'filter'. " + + "Example: {\"field\":\"status\",\"operator\":\"eq\",\"value\":\"published\"}" + ), + keyLimit: z.number().int().positive().optional().default(50).describe("Keys reported (default: 50). 'remainingKeys' says how many follow the window"), + keyOffset: z.number().int().nonnegative().optional().default(0).describe("Keys skipped before the window, in report order (default: 0)"), + valueLimit: z.number().int().positive().optional().default(10).describe("Values reported per key and type (default: 10). 'remainingValues' says how many follow the window"), + valueOffset: z.number().int().nonnegative().optional().default(0).describe("Values skipped per key and type before the window, in 'sort' order (default: 0)"), + sort: z.enum(["count", "value"]).optional().default("count").describe("Order values by document count descending (default) or by value ascending"), + minCount: z.number().int().positive().optional().default(1).describe("Hide values held by fewer documents than this (default: 1)"), + }), + }, + track(async ({ collections, match, filter, keyLimit, keyOffset, valueLimit, valueOffset, sort, minCount }) => { + const matchValidation = validateMatchArgument(match); + const filterValidation = validateFilterArgument(filter); + const validationError = matchValidation.error ?? filterValidation.error; + if (validationError) { + return { + content: [{ type: "text" as const, text: `Error: ${validationError}` }], + isError: true, + }; + } + + const effectiveCollections = collections ?? defaultCollectionNames; + let result: ListMetadataResult; + try { + result = await store.listMetadata({ + collection: effectiveCollections.length > 0 ? effectiveCollections : undefined, + match: matchValidation.match, + filter: filterValidation.filter, + keyLimit, + keyOffset, + valueLimit, + valueOffset, + sort, + minCount, + }); + } catch (err) { + if (!(err instanceof MetadataBindingBudgetError)) throw err; + return { + content: [{ type: "text" as const, text: `Error: ${err.message}` }], + isError: true, + }; + } + + const text = formatMetadataKeySummaries(result, { + showCollections: effectiveCollections.length !== 1, + valueWindowHint: "a higher 'valueLimit' or a 'valueOffset'", + keyWindowHint: "a higher 'keyLimit' or a 'keyOffset'", + keyOffset, + keyOffsetLabel: "keyOffset", + emptyMessage: "No metadata matches. Call without match/filter to see which keys exist, or check the status tool for collections with metadata.", + }); + + return { + content: [{ type: "text", text }], + structuredContent: result, + }; + }) + ); + return server; } @@ -1103,6 +1276,103 @@ export async function startMcpHttpServer( return; } + // REST endpoint: POST /metadata — metadata discovery, same body as the metadata tool + if (pathname === "/metadata" && nodeReq.method === "POST") { + const rawBody = await collectBody(nodeReq); + let parsedParams: unknown; + try { + parsedParams = rawBody.trim() === "" ? {} : JSON.parse(rawBody); + } catch { + nodeRes.writeHead(400, { "Content-Type": "application/json" }); + nodeRes.end(JSON.stringify({ error: "Invalid JSON body" })); + return; + } + if (typeof parsedParams !== "object" || parsedParams === null || Array.isArray(parsedParams)) { + nodeRes.writeHead(400, { "Content-Type": "application/json" }); + nodeRes.end(JSON.stringify({ error: "JSON body must be an object" })); + return; + } + const params = parsedParams as Record; + + // Optional metadata filter — must be an object and a valid filter AST + let restFilter: MetadataFilter | undefined; + if (params["filter"] !== undefined) { + if (typeof params["filter"] !== "object" || params["filter"] === null || Array.isArray(params["filter"])) { + nodeRes.writeHead(400, { "Content-Type": "application/json" }); + nodeRes.end(JSON.stringify({ error: "Invalid field: filter (must be an object)" })); + return; + } + const filterValidation = validateFilterArgument(params["filter"]); + if (filterValidation.error) { + nodeRes.writeHead(400, { "Content-Type": "application/json" }); + nodeRes.end(JSON.stringify({ error: filterValidation.error })); + return; + } + restFilter = filterValidation.filter; + } + + // Optional metadata match, validated the same way against entries + let restMatch: MetadataMatch | undefined; + if (params["match"] !== undefined) { + if (typeof params["match"] !== "object" || params["match"] === null || Array.isArray(params["match"])) { + nodeRes.writeHead(400, { "Content-Type": "application/json" }); + nodeRes.end(JSON.stringify({ error: "Invalid field: match (must be an object)" })); + return; + } + const matchValidation = validateMatchArgument(params["match"]); + if (matchValidation.error) { + nodeRes.writeHead(400, { "Content-Type": "application/json" }); + nodeRes.end(JSON.stringify({ error: matchValidation.error })); + return; + } + restMatch = matchValidation.match; + } + + if (params["sort"] !== undefined && params["sort"] !== "count" && params["sort"] !== "value") { + nodeRes.writeHead(400, { "Content-Type": "application/json" }); + nodeRes.end(JSON.stringify({ error: "Invalid field: sort (must be 'count' or 'value')" })); + return; + } + + if (params["collections"] !== undefined && !Array.isArray(params["collections"])) { + nodeRes.writeHead(400, { "Content-Type": "application/json" }); + nodeRes.end(JSON.stringify({ error: "Invalid field: collections (must be an array)" })); + return; + } + + // The JSON type is checked here; the store checks each number's domain. + const numberFields = readNumberFields(params, ["keyLimit", "keyOffset", "valueLimit", "valueOffset", "minCount"]); + if (numberFields.error) { + nodeRes.writeHead(400, { "Content-Type": "application/json" }); + nodeRes.end(JSON.stringify({ error: numberFields.error })); + return; + } + + // Use default collections if none specified + const effectiveCollections = params["collections"] ? params["collections"].map(String) : defaultCollectionNames; + + let result: ListMetadataResult; + try { + result = await store.listMetadata({ + collection: effectiveCollections.length > 0 ? effectiveCollections : undefined, + match: restMatch, + filter: restFilter, + sort: params["sort"], + ...numberFields.values, + }); + } catch (err) { + if (!(err instanceof MetadataOptionError) && !(err instanceof MetadataBindingBudgetError)) throw err; + nodeRes.writeHead(400, { "Content-Type": "application/json" }); + nodeRes.end(JSON.stringify({ error: err.message })); + return; + } + + nodeRes.writeHead(200, { "Content-Type": "application/json" }); + nodeRes.end(JSON.stringify(result)); + log(`${ts()} POST /metadata ${result.keys.length} keys (${Date.now() - reqStart}ms)`); + return; + } + if (pathname === "/mcp") { const rawBody = nodeReq.method !== "GET" && nodeReq.method !== "HEAD" ? await collectBody(nodeReq) diff --git a/src/metadata-filter.ts b/src/metadata-filter.ts index aa16e7cd6..fba627583 100644 --- a/src/metadata-filter.ts +++ b/src/metadata-filter.ts @@ -1,60 +1,118 @@ /** - * QMD Metadata Filter - Recursive filter AST, strict runtime validation, and - * parameterized SQL compilation. + * QMD Metadata Filter - Recursive predicate AST, strict runtime validation, + * and parameterized SQL compilation. * - * The filter has one canonical, `operator`-discriminated recursive shape shared - * by every public search surface (CLI, SDK, MCP, HTTP): + * The predicate has one canonical, `operator`-discriminated recursive shape + * shared by every public search surface (CLI, SDK, MCP, HTTP): * * { "operator": "and", "operands": [ ... ] } * { "operator": "not", "operand": { ... } } - * { "key": "status", "operator": "eq", "value": "published" } + * { "field": "status", "operator": "eq", "value": "published" } + * { "field": "topics", "operator": "prefix", "value": "sql", "caseInsensitive": true } * - * Compilation emits correlated EXISTS/NOT EXISTS subqueries over - * `document_metadata_values` with every user value bound as a parameter — - * metadata keys and values are data, never SQL. + * A condition tests one field of the record under evaluation, named by `field`, + * against `value` using `operator`. The grammar evaluates two kinds of record: + * + * - A document, whose fields are its metadata keys. This is the filter + * (`MetadataFilter`) every search surface accepts, and it compiles to one + * subquery per condition over `document_metadata_values`, correlated for + * a few candidates or a document set for the corpus (FilterScope). + * - A metadata entry (one row of that table), whose fields are `key` and + * `value`. This is the match (`MetadataMatch`) metadata discovery accepts, + * and it compiles to a predicate over one row. + * + * Both are `MetadataPredicate`, the one recursive grammar, parameterized by + * the conditions the record admits. + * + * Comparison, membership, and text operators are typed by their operand: a + * string operand only ever compares against string values, a number against + * numbers, a boolean against booleans, and a mismatch never matches. Every + * user value binds as a parameter. Metadata keys and values are data, never + * SQL. */ -import type { MetadataScalar, MetadataScalarArray } from "./metadata.js"; +import type { MetadataScalar, MetadataScalarArray, MetadataValueType } from "./metadata.js"; import { METADATA_LIMITS } from "./metadata.js"; // ============================================================================= // Public types // ============================================================================= -export type MetadataFilter = - | MetadataFilterGroup - | MetadataFilterNegation - | MetadataCondition; +/** + * The recursive grammar: a condition, or `and`/`or`/`not` over predicates. + * `Condition` is the set of conditions the record under evaluation admits. + */ +export type MetadataPredicate = + | Condition + | MetadataPredicateGroup + | MetadataPredicateNegation; -export interface MetadataFilterGroup { +export interface MetadataPredicateGroup { operator: "and" | "or"; - operands: readonly MetadataFilter[]; + operands: readonly MetadataPredicate[]; } -export interface MetadataFilterNegation { +export interface MetadataPredicateNegation { operator: "not"; - operand: MetadataFilter; + operand: MetadataPredicate; } +/** A predicate over a document, whose fields are its metadata keys. */ +export type MetadataFilter = MetadataPredicate; +export type MetadataFilterGroup = MetadataPredicateGroup; +export type MetadataFilterNegation = MetadataPredicateNegation; + +/** A predicate over a metadata entry, whose fields are `key` and `value`. */ +export type MetadataMatch = MetadataPredicate; + +/** The fields of a metadata entry. */ +export type MetadataEntryField = "key" | "value"; + +/** + * The conditions every record admits, over the fields it has. + * `caseInsensitive` folds ASCII letters on both sides and is accepted only + * where the operand is a string or an array of strings. + */ +type SharedMetadataCondition = + | { field: Field; operator: "eq" | "ne"; value: MetadataScalar; caseInsensitive?: boolean } + | { field: Field; operator: "gt" | "gte" | "lt" | "lte"; value: string | number; caseInsensitive?: boolean } + | { field: Field; operator: "in" | "nin"; value: MetadataScalarArray; caseInsensitive?: boolean } + | { field: Field; operator: "contains" | "prefix" | "suffix"; value: string; caseInsensitive?: boolean } + | { field: Field; operator: "type"; value: MetadataValueType }; + +/** + * One condition over a document. `all` and `exists` speak about the set of + * values a document holds under a key, so only a document admits them. + */ export type MetadataCondition = - | { key: string; operator: "eq" | "ne"; value: MetadataScalar } - | { key: string; operator: "gt" | "gte" | "lt" | "lte"; value: string | number } - | { key: string; operator: "in" | "nin" | "all"; value: MetadataScalarArray } - | { key: string; operator: "exists"; value: boolean }; + | SharedMetadataCondition + | { field: string; operator: "all"; value: MetadataScalarArray; caseInsensitive?: boolean } + | { field: string; operator: "exists"; value: boolean }; + +/** One condition over a metadata entry: its `key` (a string) or its `value` (typed). */ +export type MetadataEntryCondition = SharedMetadataCondition; export interface CompiledMetadataFilter { sql: string; params: (string | number)[]; } -/** Raised by parseMetadataFilter with the JSON path of the failing node. */ +/** Which record a filter is evaluated against. */ +export type MetadataRecordType = "document" | "entry"; + +const ENTRY_FIELDS: ReadonlySet = new Set(["key", "value"]); + +/** Raised by parseMetadataFilter and parseMetadataMatch with the JSON path of the failing node. */ export class MetadataFilterError extends Error { readonly path: string; + /** The message without the `Invalid metadata ... at` prefix. */ + readonly detail: string; - constructor(path: string, message: string) { - super(`Invalid metadata filter at ${path}: ${message}`); + constructor(path: string, detail: string, recordType: MetadataRecordType = "document") { + super(`Invalid metadata ${recordType === "entry" ? "match" : "filter"} at ${path}: ${detail}`); this.name = "MetadataFilterError"; this.path = path; + this.detail = detail; } } @@ -76,14 +134,19 @@ const GROUP_OPERATORS = new Set(["and", "or"]); const COMPARISON_OPERATORS = new Set(["eq", "ne", "gt", "gte", "lt", "lte"]); const ORDERED_OPERATORS = new Set(["gt", "gte", "lt", "lte"]); const MEMBERSHIP_OPERATORS = new Set(["in", "nin", "all"]); -const CONDITION_OPERATORS = new Set([...COMPARISON_OPERATORS, ...MEMBERSHIP_OPERATORS, "exists"]); +const TEXT_OPERATORS = new Set(["contains", "prefix", "suffix"]); +const CONDITION_OPERATORS = new Set([...COMPARISON_OPERATORS, ...MEMBERSHIP_OPERATORS, ...TEXT_OPERATORS, "type", "exists"]); const ALL_OPERATORS = [...GROUP_OPERATORS, "not", ...CONDITION_OPERATORS]; +const VALUE_TYPES: readonly MetadataValueType[] = ["string", "number", "boolean"]; // ============================================================================= // Validation // ============================================================================= -type FilterParseState = { nodes: number }; +type FilterParseState = { nodes: number; recordType: MetadataRecordType }; + +/** Conditions whose operand can be a string, and so can fold case. */ +type CaseFoldableCondition = Exclude; /** * Strictly validate an untrusted value as a MetadataFilter. @@ -92,10 +155,28 @@ type FilterParseState = { nodes: number }; * arrays by de-duplicating while preserving order. */ export function parseMetadataFilter(input: unknown): MetadataFilter { - const state: FilterParseState = { nodes: 0 }; + const state: FilterParseState = { nodes: 0, recordType: "document" }; return parseFilterNode(input, "$", 1, state); } +/** + * Strictly validate an untrusted value as a match over metadata entries. Same + * grammar and limits as parseMetadataFilter. A condition's `field` names a field + * of the entry, `key` or `value`. `exists` and `all` have no meaning for a + * single entry and are rejected. + */ +export function parseMetadataMatch(input: unknown): MetadataMatch { + const state: FilterParseState = { nodes: 0, recordType: "entry" }; + try { + // The walker enforces the entry rules through `state.recordType`, so + // every condition it returns is a MetadataEntryCondition. + return parseFilterNode(input, "$", 1, state) as MetadataMatch; + } catch (error) { + if (error instanceof MetadataFilterError) throw new MetadataFilterError(error.path, error.detail, "entry"); + throw error; + } +} + function parseFilterNode(input: unknown, path: string, depth: number, state: FilterParseState): MetadataFilter { if (depth > METADATA_FILTER_LIMITS.maxDepth) { throw new MetadataFilterError(path, `exceeds maximum nesting depth of ${METADATA_FILTER_LIMITS.maxDepth}`); @@ -123,7 +204,7 @@ function parseFilterNode(input: unknown, path: string, depth: number, state: Fil return parseFilterNegation(node, path, depth, state); } if (CONDITION_OPERATORS.has(operator)) { - return parseFilterCondition(node, operator, path); + return parseFilterCondition(node, operator, path, state.recordType); } throw new MetadataFilterError(path, `unknown operator '${operator}' — expected one of: ${ALL_OPERATORS.join(", ")}`); @@ -174,15 +255,29 @@ function parseFilterNegation( }; } -function parseFilterCondition(node: Record, operator: string, path: string): MetadataCondition { - rejectUnknownProperties(node, ["key", "operator", "value"], path); +function parseFilterCondition( + node: Record, + operator: string, + path: string, + recordType: MetadataRecordType, +): MetadataCondition { + rejectUnknownProperties(node, ["field", "operator", "value", "caseInsensitive"], path); - const key = node["key"]; - if (typeof key !== "string" || key.length === 0) { - throw new MetadataFilterError(path, `'${operator}' requires a non-empty string 'key'`); + const field = node["field"]; + if (typeof field !== "string" || field.length === 0) { + throw new MetadataFilterError(path, `'${operator}' requires a non-empty string 'field'`); } - if (Buffer.byteLength(key, "utf-8") > METADATA_FILTER_LIMITS.maxKeyBytes) { - throw new MetadataFilterError(path, `'key' exceeds ${METADATA_FILTER_LIMITS.maxKeyBytes} bytes`); + if (Buffer.byteLength(field, "utf-8") > METADATA_FILTER_LIMITS.maxKeyBytes) { + throw new MetadataFilterError(path, `'field' exceeds ${METADATA_FILTER_LIMITS.maxKeyBytes} bytes`); + } + + if (recordType === "entry") { + if (!ENTRY_FIELDS.has(field)) { + throw new MetadataFilterError(path, `'${field}' is not a field of a metadata entry, expected 'key' or 'value'`); + } + if (operator === "exists" || operator === "all") { + throw new MetadataFilterError(path, `'${operator}' has no meaning for a single metadata entry`); + } } if (!("value" in node)) { @@ -190,19 +285,39 @@ function parseFilterCondition(node: Record, operator: string, p } const value = node["value"]; + const caseInsensitive = node["caseInsensitive"]; + if (caseInsensitive !== undefined && typeof caseInsensitive !== "boolean") { + throw new MetadataFilterError(`${path}.caseInsensitive`, "'caseInsensitive' must be a boolean"); + } + if (operator === "exists") { if (typeof value !== "boolean") { throw new MetadataFilterError(`${path}.value`, "'exists' requires a boolean value"); } - return { key, operator, value }; + rejectCaseInsensitive(caseInsensitive, operator, path); + return { field, operator, value }; + } + + if (operator === "type") { + const valueType = VALUE_TYPES.find(candidate => candidate === value); + if (!valueType) { + throw new MetadataFilterError(`${path}.value`, `'type' requires one of: ${VALUE_TYPES.join(", ")}`); + } + rejectCaseInsensitive(caseInsensitive, operator, path); + return { field, operator, value: valueType }; + } + + if (TEXT_OPERATORS.has(operator)) { + const text = parseScalarValue(value, `${path}.value`); + if (typeof text !== "string" || text.length === 0) { + throw new MetadataFilterError(`${path}.value`, `'${operator}' requires a non-empty string value`); + } + return withCaseFolding({ field, operator: operator as "contains" | "prefix" | "suffix", value: text }, caseInsensitive, path); } if (MEMBERSHIP_OPERATORS.has(operator)) { - return { - key, - operator: operator as "in" | "nin" | "all", - value: parseMembershipValues(value, operator, path), - }; + const values = parseMembershipValues(value, operator, path); + return withCaseFolding({ field, operator: operator as "in" | "nin" | "all", value: values }, caseInsensitive, path); } // Comparison operators: eq, ne, gt, gte, lt, lte. @@ -210,7 +325,26 @@ function parseFilterCondition(node: Record, operator: string, p if (ORDERED_OPERATORS.has(operator) && typeof scalar === "boolean") { throw new MetadataFilterError(`${path}.value`, `'${operator}' requires a string or number value`); } - return { key, operator, value: scalar } as MetadataCondition; + return withCaseFolding({ field, operator, value: scalar } as CaseFoldableCondition, caseInsensitive, path); +} + +/** Attach `caseInsensitive` when given, which only string operands accept. */ +function withCaseFolding(condition: CaseFoldableCondition, caseInsensitive: unknown, path: string): MetadataCondition { + if (caseInsensitive === undefined) return condition; + + const operand: unknown = condition.value; + const isText = typeof operand === "string"; + const isTextArray = Array.isArray(operand) && operand.every(element => typeof element === "string"); + if (!isText && !isTextArray) { + throw new MetadataFilterError(`${path}.caseInsensitive`, "'caseInsensitive' applies to string values only"); + } + + return { ...condition, caseInsensitive: caseInsensitive === true }; +} + +function rejectCaseInsensitive(caseInsensitive: unknown, operator: string, path: string): void { + if (caseInsensitive === undefined) return; + throw new MetadataFilterError(`${path}.caseInsensitive`, `'caseInsensitive' does not apply to '${operator}'`); } function parseMembershipValues(value: unknown, operator: string, path: string): MetadataScalarArray { @@ -270,33 +404,87 @@ function rejectUnknownProperties(node: Record, allowed: string[ // ============================================================================= /** - * Compile a validated filter into one parameterized SQL predicate correlated - * against a documents-table alias (e.g. `d`). All keys and values are bound - * parameters. The caller is responsible for restricting the surrounding query - * to active documents with current, error-free metadata extraction. + * The documents a compiled filter is applied to, which decides the SQL each + * condition takes. + * + * - `candidates`: a few documents, as when search checks its results. Each + * condition is a correlated `EXISTS` that seeks the document's own rows. + * - `corpus`: every document, as when discovery counts them. Each condition + * is an uncorrelated `IN (SELECT document_id ...)` that SQLite evaluates + * once per statement and probes per document, where a correlated predicate + * would be re-run for every metadata row the statement joins. + */ +export type FilterScope = "candidates" | "corpus"; + +/** + * Compile a validated filter into one parameterized SQL predicate over a + * documents-table alias (e.g. `d`). All keys and values are bound parameters. + * The caller is responsible for restricting the surrounding query to active + * documents with current, error-free metadata extraction. + */ +export function compileMetadataFilter(filter: MetadataFilter, documentsAlias: string, scope: FilterScope = "candidates"): CompiledMetadataFilter { + const params: (string | number)[] = []; + const sql = compileFilterNode(filter, { alias: documentsAlias, scope }, params); + return { sql, params }; +} + +/** The documents-table alias a filter tests and the scope it is applied to. */ +interface FilterTarget { + alias: string; + scope: FilterScope; +} + +/** + * Compile a validated match into one parameterized SQL predicate over a + * `document_metadata_values` row alias (e.g. `mv`). A condition's `field` names + * the field of the entry its operand is tested against: `key`, which is always + * a string, or `value`, which is typed by `value_type`. The caller is + * responsible for scoping the rows to eligible documents. */ -export function compileMetadataFilter(filter: MetadataFilter, documentsAlias: string): CompiledMetadataFilter { +export function compileMetadataMatch(match: MetadataMatch, valuesAlias: string): CompiledMetadataFilter { const params: (string | number)[] = []; - const sql = compileFilterNode(filter, documentsAlias, params); + const sql = compileMatchNode(match, valuesAlias, params); return { sql, params }; } -function compileFilterNode(filter: MetadataFilter, alias: string, params: (string | number)[]): string { +/** The typed columns a condition's operand is tested against. */ +interface ValueColumns { + type: string; + text: string; + number: string; + boolean: string; +} + +/** The values of a document, as seen from inside a correlated subquery. */ +const DOCUMENT_VALUE_COLUMNS: ValueColumns = { + type: "mv.value_type", + text: "mv.text_value", + number: "mv.number_value", + boolean: "mv.boolean_value", +}; + +function compileFilterNode(filter: MetadataFilter, target: FilterTarget, params: (string | number)[]): string { + const columns = DOCUMENT_VALUE_COLUMNS; + switch (filter.operator) { case "and": case "or": { const joiner = filter.operator === "and" ? " AND " : " OR "; - return `(${filter.operands.map(operand => compileFilterNode(operand, alias, params)).join(joiner)})`; + return `(${filter.operands.map(operand => compileFilterNode(operand, target, params)).join(joiner)})`; } case "not": - return `NOT ${compileFilterNode(filter.operand, alias, params)}`; + return `NOT ${compileFilterNode(filter.operand, target, params)}`; case "exists": - params.push(filter.key); + params.push(filter.field); return filter.value - ? buildValueExistsSql(alias, "mv.key = ?") - : `NOT ${buildValueExistsSql(alias, "mv.key = ?")}`; + ? buildConditionSql(target, "byDocument") + : `NOT ${buildConditionSql(target, "byDocument")}`; + + case "type": + params.push(filter.field, filter.value); + return buildConditionSql(target, "byDocument", `${columns.type} = ?`); case "eq": case "gt": @@ -304,10 +492,12 @@ function compileFilterNode(filter: MetadataFilter, alias: string, params: (strin case "lt": case "lte": { const sqlOperator = { eq: "=", gt: ">", gte: ">=", lt: "<", lte: "<=" }[filter.operator]; - params.push(filter.key, bindScalar(filter.value)); - return buildValueExistsSql( - alias, - `mv.key = ? AND mv.value_type = '${valueTypeOf(filter.value)}' AND mv.${valueColumnOf(filter.value)} ${sqlOperator} ?`, + const columnSql = buildOperandColumnSql(columns, filter.value, filter.caseInsensitive); + params.push(filter.field, bindOperand(filter.value, filter.caseInsensitive)); + return buildConditionSql( + target, + filter.operator === "eq" ? equalitySeekOf(filter.caseInsensitive) : "byDocument", + `${columns.type} = '${valueTypeOf(filter.value)}' AND ${columnSql} ${sqlOperator} ?`, ); } @@ -315,61 +505,198 @@ function compileFilterNode(filter: MetadataFilter, alias: string, params: (strin // Key must have at least one same-type value, and no same-type value // may equal the operand. Missing keys and type mismatches do not match. const valueType = valueTypeOf(filter.value); - params.push(filter.key); - const presentSql = buildValueExistsSql(alias, `mv.key = ? AND mv.value_type = '${valueType}'`); - params.push(filter.key, bindScalar(filter.value)); - const equalSql = buildValueExistsSql( - alias, - `mv.key = ? AND mv.value_type = '${valueType}' AND mv.${valueColumnOf(filter.value)} = ?`, - ); + const columnSql = buildOperandColumnSql(columns, filter.value, filter.caseInsensitive); + params.push(filter.field); + const presentSql = buildConditionSql(target, "byDocument", `${columns.type} = '${valueType}'`); + params.push(filter.field, bindOperand(filter.value, filter.caseInsensitive)); + const equalSql = buildConditionSql(target, equalitySeekOf(filter.caseInsensitive), `${columns.type} = '${valueType}' AND ${columnSql} = ?`); return `(${presentSql} AND NOT ${equalSql})`; } case "in": case "nin": { const valueType = valueTypeOf(filter.value[0]!); - const column = valueColumnOf(filter.value[0]!); + const columnSql = buildOperandColumnSql(columns, filter.value[0]!, filter.caseInsensitive); const placeholders = filter.value.map(() => "?").join(", "); + const operands = filter.value.map(element => bindOperand(element, filter.caseInsensitive)); if (filter.operator === "in") { - params.push(filter.key, ...filter.value.map(bindScalar)); - return buildValueExistsSql(alias, `mv.key = ? AND mv.value_type = '${valueType}' AND mv.${column} IN (${placeholders})`); + params.push(filter.field, ...operands); + return buildConditionSql(target, equalitySeekOf(filter.caseInsensitive), `${columns.type} = '${valueType}' AND ${columnSql} IN (${placeholders})`); } - params.push(filter.key); - const presentSql = buildValueExistsSql(alias, `mv.key = ? AND mv.value_type = '${valueType}'`); - params.push(filter.key, ...filter.value.map(bindScalar)); - const memberSql = buildValueExistsSql(alias, `mv.key = ? AND mv.value_type = '${valueType}' AND mv.${column} IN (${placeholders})`); + params.push(filter.field); + const presentSql = buildConditionSql(target, "byDocument", `${columns.type} = '${valueType}'`); + params.push(filter.field, ...operands); + const memberSql = buildConditionSql(target, equalitySeekOf(filter.caseInsensitive), `${columns.type} = '${valueType}' AND ${columnSql} IN (${placeholders})`); return `(${presentSql} AND NOT ${memberSql})`; } case "all": { const valueType = valueTypeOf(filter.value[0]!); - const column = valueColumnOf(filter.value[0]!); + const columnSql = buildOperandColumnSql(columns, filter.value[0]!, filter.caseInsensitive); const memberSqls = filter.value.map(element => { - params.push(filter.key, bindScalar(element)); - return buildValueExistsSql(alias, `mv.key = ? AND mv.value_type = '${valueType}' AND mv.${column} = ?`); + params.push(filter.field, bindOperand(element, filter.caseInsensitive)); + return buildConditionSql(target, equalitySeekOf(filter.caseInsensitive), `${columns.type} = '${valueType}' AND ${columnSql} = ?`); }); return `(${memberSqls.join(" AND ")})`; } + + case "contains": + case "prefix": + case "suffix": { + params.push(filter.field); + const columnSql = buildOperandColumnSql(columns, filter.value, filter.caseInsensitive); + const textSql = compileTextTestSql(filter.operator, filter.value, columnSql, filter.caseInsensitive, params); + return buildConditionSql(target, "byDocument", `${columns.type} = 'string' AND ${textSql}`); + } + } +} + +/** + * How a `candidates` subquery reaches one document's rows for a key. The + * covering indexes are (key, , document_id), so a condition that pins + * the value column with an equality is a single probe by value. Any other + * condition would walk the key's whole value range once per document (the + * stats-free planner assumes a range is narrow), so it seeks the document's + * own rows through the primary key instead. A `corpus` subquery runs once, + * so it always takes the covering index. + */ +type RowSeek = "byValue" | "byDocument"; + +/** An equality pins the value column unless it is folded, `lower()` is opaque to the index. */ +function equalitySeekOf(caseInsensitive: boolean | undefined): RowSeek { + return caseInsensitive ? "byDocument" : "byValue"; +} + +/** + * One condition over the document's rows for the key, whose parameter the + * caller binds before `conditionSql`'s. In the `candidates` scope the unary + * `+` on the key term hides it from the planner, which leaves `document_id` + * as the only indexable term and makes the seek independent of `sqlite_stat1`. + */ +function buildConditionSql(target: FilterTarget, seek: RowSeek, conditionSql?: string): string { + if (target.scope === "corpus") { + const whereSql = conditionSql ? `mv.key = ? AND ${conditionSql}` : "mv.key = ?"; + return `${target.alias}.id IN (SELECT mv.document_id FROM document_metadata_values mv WHERE ${whereSql})`; + } + + const keyTermSql = seek === "byValue" ? "mv.key = ?" : "+mv.key = ?"; + const whereSql = conditionSql ? `${keyTermSql} AND ${conditionSql}` : keyTermSql; + return `EXISTS (SELECT 1 FROM document_metadata_values mv WHERE mv.document_id = ${target.alias}.id AND ${whereSql})`; +} + +function compileMatchNode(match: MetadataMatch, alias: string, params: (string | number)[]): string { + switch (match.operator) { + case "and": + case "or": { + const joiner = match.operator === "and" ? " AND " : " OR "; + return `(${match.operands.map(operand => compileMatchNode(operand, alias, params)).join(joiner)})`; + } + + case "not": + return `NOT ${compileMatchNode(match.operand, alias, params)}`; + + default: + return compileMatchCondition(match, alias, params); + } +} + +/** + * One condition over one entry. The entry's `key` field is a string with no + * other typed columns, so a number or boolean operand fails its type guard and + * matches nothing, the same outcome a type mismatch has in a filter. + */ +function compileMatchCondition(condition: MetadataEntryCondition, alias: string, params: (string | number)[]): string { + const columns: ValueColumns = condition.field === "key" + ? { type: "'string'", text: `${alias}.key`, number: "NULL", boolean: "NULL" } + : { type: `${alias}.value_type`, text: `${alias}.text_value`, number: `${alias}.number_value`, boolean: `${alias}.boolean_value` }; + + switch (condition.operator) { + case "type": + params.push(condition.value); + return `${columns.type} = ?`; + + case "eq": + case "ne": + case "gt": + case "gte": + case "lt": + case "lte": { + const sqlOperator = { eq: "=", ne: "<>", gt: ">", gte: ">=", lt: "<", lte: "<=" }[condition.operator]; + const columnSql = buildOperandColumnSql(columns, condition.value, condition.caseInsensitive); + params.push(bindOperand(condition.value, condition.caseInsensitive)); + return `(${columns.type} = '${valueTypeOf(condition.value)}' AND ${columnSql} ${sqlOperator} ?)`; + } + + case "in": + case "nin": { + const valueType = valueTypeOf(condition.value[0]!); + const columnSql = buildOperandColumnSql(columns, condition.value[0]!, condition.caseInsensitive); + const placeholders = condition.value.map(() => "?").join(", "); + const membership = condition.operator === "in" ? "IN" : "NOT IN"; + params.push(...condition.value.map(element => bindOperand(element, condition.caseInsensitive))); + return `(${columns.type} = '${valueType}' AND ${columnSql} ${membership} (${placeholders}))`; + } + + case "contains": + case "prefix": + case "suffix": { + const columnSql = buildOperandColumnSql(columns, condition.value, condition.caseInsensitive); + const textSql = compileTextTestSql(condition.operator, condition.value, columnSql, condition.caseInsensitive, params); + return `(${columns.type} = 'string' AND ${textSql})`; + } } } -function buildValueExistsSql(alias: string, conditionSql: string): string { - return `EXISTS (SELECT 1 FROM document_metadata_values mv WHERE mv.document_id = ${alias}.id AND ${conditionSql})`; +/** + * Substring tests over a string column. Prefix and suffix compare UTF-8 bytes, + * with the operand's byte length bound from JavaScript: SQLite's text + * `length()` and `substr()` stop at an embedded NUL, and blobs do not. UTF-8 + * is self-synchronizing, so a byte-prefix (or byte-suffix) of a whole operand + * is exactly a character-prefix (or -suffix). + */ +function compileTextTestSql( + operator: "contains" | "prefix" | "suffix", + text: string, + columnSql: string, + caseInsensitive: boolean | undefined, + params: (string | number)[], +): string { + const operand = caseInsensitive ? foldAsciiCase(text) : text; + + if (operator === "contains") { + params.push(operand); + return `instr(${columnSql}, ?) > 0`; + } + + params.push(Buffer.byteLength(operand, "utf-8"), operand); + // SQLite returns NULL for a substring of an empty BLOB. Each primitive + // must return a boolean so entry matches and their negations partition rows. + return operator === "prefix" + ? `COALESCE(substr(CAST(${columnSql} AS BLOB), 1, ?) = CAST(? AS BLOB), 0)` + : `COALESCE(substr(CAST(${columnSql} AS BLOB), -?) = CAST(? AS BLOB), 0)`; } -function valueTypeOf(scalar: MetadataScalar): "string" | "number" | "boolean" { - return typeof scalar as "string" | "number" | "boolean"; +function valueTypeOf(scalar: MetadataScalar): MetadataValueType { + return typeof scalar as MetadataValueType; } -function valueColumnOf(scalar: MetadataScalar): "text_value" | "number_value" | "boolean_value" { - if (typeof scalar === "string") return "text_value"; - if (typeof scalar === "number") return "number_value"; - return "boolean_value"; +/** The typed column an operand compares against, folded when the condition ignores case. */ +function buildOperandColumnSql(columns: ValueColumns, scalar: MetadataScalar, caseInsensitive: boolean | undefined): string { + if (typeof scalar === "number") return columns.number; + if (typeof scalar === "boolean") return columns.boolean; + return caseInsensitive ? `lower(${columns.text})` : columns.text; } -function bindScalar(scalar: MetadataScalar): string | number { +/** Booleans bind as 0/1. Case-insensitive strings bind folded the same way SQLite's `lower()` folds the column. */ +function bindOperand(scalar: MetadataScalar, caseInsensitive: boolean | undefined): string | number { if (typeof scalar === "boolean") return scalar ? 1 : 0; + if (typeof scalar === "string" && caseInsensitive) return foldAsciiCase(scalar); return scalar; } + +/** SQLite's built-in `lower()` folds ASCII letters only, so the operand is folded over the same range. */ +function foldAsciiCase(text: string): string { + return text.replace(/[A-Z]/g, letter => letter.toLowerCase()); +} diff --git a/src/metadata-format.ts b/src/metadata-format.ts new file mode 100644 index 000000000..351ae0fca --- /dev/null +++ b/src/metadata-format.ts @@ -0,0 +1,311 @@ +/** + * QMD Metadata Format - Plain-text rendering of metadata discovery results. + * + * Shared by the CLI and the MCP `metadata` tool so both print one shape: a + * filter line when a filter narrowed the documents, a header per key, a body + * per type, and footers naming the remainder and the options that reach it + * whenever a value list or the key list is windowed. The empty states render + * here too, so the filter line survives them. + * + * String values print bare only when the bare form is unambiguous. A value + * that is empty, padded, contains a quote, a backslash, a control character, + * a line separator, or a delimiter of the compact `value (count), ...` list, + * or that reads as a JSON number, boolean, or null, or as the list's + * remainder tail, is printed as a JSON string, so what the reader sees is + * exactly the value. + */ + +import type { + ListMetadataResult, + MetadataKeySummary, + MetadataKeyTypeSummary, + MetadataValueCount, +} from "./metadata-store.js"; + +export interface MetadataFormatColors { + reset: string; + dim: string; + bold: string; + cyan: string; +} + +export interface FormatMetadataOptions { + /** Show which collections contribute each type on type-split keys. */ + showCollections?: boolean; + /** How the caller widens or pages the value window, e.g. `--value-limit , --value-offset , or --all-values`. */ + valueWindowHint: string; + /** How the caller widens or pages the key window, e.g. `--key-limit , --key-offset , or --all-keys`. */ + keyWindowHint: string; + /** The key offset the caller requested, and how the caller spells that option, e.g. `--key-offset`. */ + keyOffset: number; + keyOffsetLabel: string; + /** Printed in place of the key blocks when no metadata entry is in the result at all. */ + emptyMessage: string; + /** ANSI sequences; omit for plain text. */ + colors?: MetadataFormatColors; +} + +const NO_COLORS: MetadataFormatColors = { reset: "", dim: "", bold: "", cyan: "" }; + +/** + * Render the whole result: the filter line when a filter is in play, then + * one block per key separated by blank lines and the key footer when keys + * were left out of the window, or the one-line empty state when there are no + * blocks to show. + */ +export function formatMetadataKeySummaries(result: ListMetadataResult, options: FormatMetadataOptions): string { + const colors = options.colors ?? NO_COLORS; + const blocks: string[] = []; + + if (result.filteredDocuments !== undefined) { + blocks.push(`${colors.dim}filter:${colors.reset} ${formatCount(result.filteredDocuments)} of ${formatCount(result.documents)} documents`); + } + + if (result.totalKeys === 0) { + blocks.push(`${colors.dim}${options.emptyMessage}${colors.reset}`); + } else if (result.keys.length === 0) { + const keyLabel = result.totalKeys === 1 ? "key" : "keys"; + blocks.push(`${colors.dim}No keys at ${options.keyOffsetLabel} ${formatCount(options.keyOffset)}, ${formatCount(result.totalKeys)} ${keyLabel} in total.${colors.reset}`); + } + + blocks.push(...result.keys.map(summary => formatMetadataKeySummary(summary, result, options))); + + if (result.remainingKeys > 0) { + blocks.push(`${colors.dim}${formatCount(result.remainingKeys)} more ${result.remainingKeys === 1 ? "key" : "keys"}, use ${options.keyWindowHint}${colors.reset}`); + } + + return blocks.join("\n\n"); +} + +/** + * One key: a header of name, types, coverage, and distinct count, then the + * body per type. Coverage is measured against the documents the filter + * admitted when there is one, otherwise against every active document. + */ +function formatMetadataKeySummary(summary: MetadataKeySummary, result: ListMetadataResult, options: FormatMetadataOptions): string { + const colors = options.colors ?? NO_COLORS; + const typeLabel = summary.types.map(typeLabelOf).join(" | "); + const coverage = `${formatCount(summary.documents)} of ${formatCount(result.filteredDocuments ?? result.documents)} documents`; + const header = [`${colors.cyan}${colors.bold}${summary.key}${colors.reset}`, `${colors.dim}${typeLabel}${colors.reset}`, coverage]; + const lines: string[] = []; + + if (summary.types.length === 1) { + const typeSummary = summary.types[0]!; + if (typeSummary.type !== "boolean") header.push(`${formatCount(typeSummary.distinctValues)} distinct`); + lines.push(header.join(" "), ...formatTypeBody(typeSummary)); + } else { + lines.push(header.join(" "), ...formatTypeSplit(summary.types, options)); + } + + const remainingValues = summary.types.reduce((sum, typeSummary) => sum + typeSummary.remainingValues, 0); + if (remainingValues > 0) { + lines.push(`${colors.dim}${formatCount(remainingValues)} more ${remainingValues === 1 ? "value" : "values"}, use ${options.valueWindowHint}${colors.reset}`); + } + + return lines.join("\n"); +} + +export interface FormatMetadataOverviewOptions { + /** Active, extracted documents declaring at least one key. */ + documentsWithMetadata: number; + /** Active documents awaiting extraction, mentioned so the coverage reads honestly. */ + pendingMetadata: number; + /** Command that shows the rest, e.g. `qmd collection metadata notes`. */ + drillDownHint: string; + colors?: MetadataFormatColors; +} + +/** + * The `Metadata:` section of `collection show`: a coverage line, then the + * keys in the result's window as aligned rows with a short value preview, + * then a pointer at the drill-down when keys were left out. Indented to sit + * under the other `show` fields. + */ +export function formatMetadataOverview(result: ListMetadataResult, options: FormatMetadataOverviewOptions): string { + const colors = options.colors ?? NO_COLORS; + const pendingNote = options.pendingMetadata > 0 ? ` (${formatCount(options.pendingMetadata)} pending extraction)` : ""; + + if (result.totalKeys === 0) return ` Metadata: none${pendingNote}`; + + const keyLabel = result.totalKeys === 1 ? "key" : "keys"; + const lines = [` Metadata: ${formatCount(result.totalKeys)} ${keyLabel}, ${formatCount(options.documentsWithMetadata)} of ${formatCount(result.documents)} documents${pendingNote}`]; + + const shownKeys = result.keys; + const keyWidth = Math.max(...shownKeys.map(summary => summary.key.length)); + const typeWidth = Math.max(...shownKeys.map(summary => summary.types.map(typeLabelOf).join(" | ").length)); + const documentsWidth = Math.max(...shownKeys.map(summary => formatCount(summary.documents).length)); + const distinctWidth = Math.max(...shownKeys.map(summary => formatCount(distinctValuesOf(summary)).length)); + + for (const summary of shownKeys) { + const typeLabel = summary.types.map(typeLabelOf).join(" | "); + const columns = [ + `${colors.cyan}${summary.key.padEnd(keyWidth)}${colors.reset}`, + `${colors.dim}${typeLabel.padEnd(typeWidth)}${colors.reset}`, + `${formatCount(summary.documents).padStart(documentsWidth)} ${documentsLabelOf(summary.documents)}`, + `${formatCount(distinctValuesOf(summary)).padStart(distinctWidth)} distinct`, + ]; + const preview = formatValuePreview(summary, options.drillDownHint); + if (preview) columns.push(preview); + lines.push(` ${columns.join(" ")}`); + } + + if (result.remainingKeys > 0) { + lines.push(` ${colors.dim}${formatCount(result.remainingKeys)} more ${result.remainingKeys === 1 ? "key" : "keys"}, see '${options.drillDownHint}'${colors.reset}`); + } + + return lines.join("\n"); +} + +/** + * One-line value preview for the overview row. Strings list the window with + * a trailing ellipsis when truncated, and nothing at all when every value is + * unique (a value list would be noise). Numbers give the range, booleans the + * two counts, and a type conflict points at the drill-down. + */ +function formatValuePreview(summary: MetadataKeySummary, drillDownHint: string): string { + if (summary.types.length > 1) { + // JSON quoting does not protect an apostrophe inside a shell single quote. + const match = JSON.stringify({ field: "key", operator: "eq", value: summary.key }).replaceAll("'", "'\\''"); + return `types disagree, see: ${drillDownHint} --match '${match}'`; + } + + const typeSummary = summary.types[0]!; + if (typeSummary.type === "boolean") return formatBooleanCounts(typeSummary.values).replace(" ", ", "); + if (typeSummary.type === "number") { + const range = typeSummary.range!; + return `${formatValue(range.min)} to ${formatValue(range.max)}, median ${formatValue(range.median)}`; + } + if (typeSummary.distinctValues === typeSummary.documents) return ""; + + const preview = typeSummary.values.map(count => `${formatValue(count.value)} (${formatCount(count.documents)})`); + if (typeSummary.remainingValues > 0) preview.push("..."); + return preview.join(", "); +} + +function distinctValuesOf(summary: MetadataKeySummary): number { + return summary.types.reduce((sum, typeSummary) => sum + typeSummary.distinctValues, 0); +} + +/** Body for a key with one type: vertical values for strings, a range for numbers, one line for booleans. */ +function formatTypeBody(typeSummary: MetadataKeyTypeSummary): string[] { + if (typeSummary.type === "boolean") return [` ${formatBooleanCounts(typeSummary.values)}`]; + + if (typeSummary.type === "number") { + const lines = [` ${formatRange(typeSummary)}`]; + if (typeSummary.values.length > 0) lines.push(` ${formatInlineValues(typeSummary).join(" ")}`); + return lines; + } + + const valueWidth = Math.max(...typeSummary.values.map(count => formatValue(count.value).length)); + const countWidth = Math.max(...typeSummary.values.map(count => formatCount(count.documents).length)); + return typeSummary.values.map(count => ` ${formatValue(count.value).padEnd(valueWidth)} ${formatCount(count.documents).padStart(countWidth)}`); +} + +/** + * Body for a key whose documents disagree on type: one line per type with + * its own document count and a compact value summary, plus the contributing + * collections when the view spans more than one. + */ +function formatTypeSplit(typeSummaries: MetadataKeyTypeSummary[], options: FormatMetadataOptions): string[] { + const typeWidth = Math.max(...typeSummaries.map(typeSummary => typeSummary.type.length)); + const documentsWidth = Math.max(...typeSummaries.map(typeSummary => formatCount(typeSummary.documents).length)); + + const rows = typeSummaries.map(typeSummary => { + let valuesSummary: string; + if (typeSummary.type === "boolean") valuesSummary = formatBooleanCounts(typeSummary.values); + else if (typeSummary.type === "number") valuesSummary = formatRange(typeSummary); + else { + const inlineValues = typeSummary.values.map(count => `${formatValue(count.value)} (${formatCount(count.documents)})`); + if (typeSummary.remainingValues > 0) inlineValues.push(`${formatCount(typeSummary.remainingValues)} more`); + valuesSummary = inlineValues.join(", "); + } + const row = ` ${typeSummary.type.padEnd(typeWidth)} ${formatCount(typeSummary.documents).padStart(documentsWidth)} ${documentsLabelOf(typeSummary.documents)} ${valuesSummary}`; + return { row, collections: typeSummary.collections.join(", ") }; + }); + + if (!options.showCollections) return rows.map(({ row }) => row); + + const rowWidth = Math.max(...rows.map(({ row }) => row.length)); + return rows.map(({ row, collections }) => `${row.padEnd(rowWidth)} ${collections}`); +} + +function typeLabelOf(typeSummary: MetadataKeyTypeSummary): string { + return typeSummary.multiValued ? `${typeSummary.type}[]` : typeSummary.type; +} + +function formatRange(typeSummary: MetadataKeyTypeSummary): string { + const range = typeSummary.range!; + return `min ${formatValue(range.min)} median ${formatValue(range.median)} max ${formatValue(range.max)}`; +} + +/** + * Numbers enumerate inline as `value (documents)`. When the whole + * distribution fits it reads in value order; a truncated window keeps the + * requested order so the most common values stay visible. + */ +function formatInlineValues(typeSummary: MetadataKeyTypeSummary): string[] { + const values = typeSummary.remainingValues === 0 + ? [...typeSummary.values].sort((a, b) => Number(a.value) - Number(b.value)) + : typeSummary.values; + return values.map(count => `${formatValue(count.value)} (${formatCount(count.documents)})`); +} + +function formatBooleanCounts(values: MetadataValueCount[]): string { + const trueCount = values.find(count => count.value === true); + const falseCount = values.find(count => count.value === false); + const parts: string[] = []; + if (trueCount) parts.push(`true ${formatCount(trueCount.documents)}`); + if (falseCount) parts.push(`false ${formatCount(falseCount.documents)}`); + return parts.join(" "); +} + +/** Padded so `doc` and `docs` rows stay column-aligned. */ +function documentsLabelOf(documents: number): string { + return documents === 1 ? "doc " : "docs"; +} + +function formatValue(value: string | number | boolean): string { + if (typeof value !== "string") return String(value); + return isUnambiguousBare(value) ? value : formatQuoted(value); +} + +// Controls (C0, DEL, C1) and the line and paragraph separators, none of which +// print as themselves. +const CONTROL_CHARACTERS = /[\u0000-\u001f\u007f-\u009f\u2028\u2029]/u; +// Whitespace other than one interior ASCII space between other characters. +const AMBIGUOUS_WHITESPACE = /^\s|\s$|\s\s|[^\S ]/u; +// The compact list joins `value (count)` items with ", " and ends with a +// remainder tail, so a bare string can neither contain the delimiters nor +// read as a tail. One rule for both layouts, so a value never prints two ways. +const LIST_DELIMITERS = /[,()]/u; +const LIST_TAIL = /^(?:\.\.\.|\d+ more)$/u; + +/** True when printing the string as-is cannot be mistaken for another value or for layout. */ +function isUnambiguousBare(text: string): boolean { + if (text === "" || text.includes('"') || text.includes("\\")) return false; + if (CONTROL_CHARACTERS.test(text) || AMBIGUOUS_WHITESPACE.test(text)) return false; + if (LIST_DELIMITERS.test(text) || LIST_TAIL.test(text)) return false; + return !readsAsJsonLiteral(text); +} + +/** A string that JSON would parse as something other than a string, e.g. `42`, `true`, `null`, `[1]`. */ +function readsAsJsonLiteral(text: string): boolean { + try { + JSON.parse(text); + return true; + } catch { + return false; + } +} + +/** A JSON string, with every character in CONTROL_CHARACTERS escaped, not only the ones JSON.stringify escapes. */ +function formatQuoted(text: string): string { + return JSON.stringify(text).replace( + /[\u007f-\u009f\u2028\u2029]/gu, + character => `\\u${character.codePointAt(0)!.toString(16).padStart(4, "0")}`, + ); +} + +function formatCount(count: number): string { + return count.toLocaleString("en-US"); +} diff --git a/src/metadata-store.ts b/src/metadata-store.ts index 704b2268b..2d7257bb2 100644 --- a/src/metadata-store.ts +++ b/src/metadata-store.ts @@ -14,13 +14,25 @@ * for filtering. */ -import type { Database } from "./db.js"; +import * as buffer from "node:buffer"; + +import type { Database, SQLiteValue } from "./db.js"; import { extractDocumentMetadata, METADATA_EXTRACTION_VERSION, type DocumentMetadata, type MetadataExtractionResult, + type MetadataScalar, + type MetadataValueType, } from "./metadata.js"; +import { + compileMetadataFilter, + compileMetadataMatch, + parseMetadataMatch, + type CompiledMetadataFilter, + type MetadataFilter, + type MetadataMatch, +} from "./metadata-filter.js"; // ============================================================================= // Schema @@ -163,9 +175,11 @@ function isDocumentMetadataCurrent(db: Database, documentId: number): boolean { /** * Count active documents without a current, error-free metadata extraction. * These documents are excluded from filtered search until `qmd update` runs. + * Scoped to `collectionNames` when given, otherwise the whole index. */ -export function countDocumentsPendingMetadata(db: Database): number { - const row = db.prepare(` +export function countDocumentsPendingMetadata(db: Database, collectionNames?: string[]): number { + const params: SQLiteValue[] = [METADATA_EXTRACTION_VERSION]; + let sql = ` SELECT COUNT(*) as c FROM documents d WHERE d.active = 1 AND NOT EXISTS ( @@ -173,8 +187,12 @@ export function countDocumentsPendingMetadata(db: Database): number { WHERE dm.document_id = d.id AND dm.extraction_version = ? AND dm.extraction_error IS NULL - ) - `).get(METADATA_EXTRACTION_VERSION) as { c: number }; + )`; + if (collectionNames) { + sql += ` AND d.collection IN (SELECT value FROM json_each(?))`; + params.push(JSON.stringify(collectionNames)); + } + const row = db.prepare(sql).get(...params) as { c: number }; return row.c; } @@ -210,3 +228,681 @@ export function parseMetadataJson(metadataJson: string | null | undefined): Docu return {}; } } + +// ============================================================================= +// Discovery +// ============================================================================= + +export interface ListMetadataOptions { + /** Restrict to these collections. Undefined means every collection in the index. */ + collection?: string | string[] | undefined; + /** + * Report only the metadata entries matching this condition. Same grammar as + * `filter`, evaluated against each entry: a condition's `field` names the + * entry's `key` or `value`. Undefined reports every entry. + */ + match?: MetadataMatch | undefined; + /** Count only documents matching this filter. Same AST as search. */ + filter?: MetadataFilter | undefined; + /** Keys reported (default 50). `Infinity` removes the window. */ + keyLimit?: number | undefined; + /** Keys skipped before the window, in report order (default 0). */ + keyOffset?: number | undefined; + /** Values reported per key and type (default 10). `Infinity` removes the window. */ + valueLimit?: number | undefined; + /** Values skipped per key and type before the window, in `sort` order (default 0). */ + valueOffset?: number | undefined; + /** Order of values within a key (default "count"). */ + sort?: "count" | "value" | undefined; + /** Drop values held by fewer documents than this (default 1). */ + minCount?: number | undefined; +} + +export interface ListMetadataResult { + /** Active documents in scope. */ + documents: number; + /** + * Documents in scope that pass `filter`. Present only when a filter was + * given, and then the denominator for every coverage count. + */ + filteredDocuments?: number; + /** Keys with a matching entry, before the key window. */ + totalKeys: number; + /** The key window, by documents descending then key ascending. */ + keys: MetadataKeySummary[]; + /** Keys after the window: `totalKeys - keyOffset - keys.length`, floored at 0. */ + remainingKeys: number; +} + +export interface MetadataKeySummary { + key: string; + /** Distinct documents holding a matching entry under any type. */ + documents: number; + /** One entry per value_type present. Length > 1 is a type conflict. */ + types: MetadataKeyTypeSummary[]; +} + +export interface MetadataKeyTypeSummary { + type: MetadataValueType; + /** + * True when any document holds more than one value for this key. A + * one-element array is indistinguishable from a scalar in the index. + */ + multiValued: boolean; + /** Distinct documents holding a matching value of this type. */ + documents: number; + /** Distinct matching values that meet `minCount`. */ + distinctValues: number; + /** The value window, narrowed by `match` and `minCount`, ordered by `sort`. */ + values: MetadataValueCount[]; + /** Distinct values after the window: `distinctValues - valueOffset - values.length`, floored at 0. */ + remainingValues: number; + /** Numbers only. Computed over every matching value row, so array elements each count. */ + range?: { min: number; median: number; max: number }; + /** Collections contributing a matching value of this type, ascending. */ + collections: string[]; +} + +export interface MetadataValueCount { + value: MetadataScalar; + documents: number; +} + +/** The light form of a key summary for status views: name, coverage, and types, no values. */ +export interface MetadataKeyOverview { + key: string; + /** Distinct documents declaring the key with any type. */ + documents: number; + /** By documents descending then name. Length > 1 is a type conflict. */ + types: MetadataValueType[]; +} + +/** Exact vocabulary size and a bounded key overview, computed together. */ +export interface MetadataOverview { + totalKeys: number; + keys: MetadataKeyOverview[]; +} + +export const DEFAULT_METADATA_KEY_LIMIT = 50; +export const DEFAULT_METADATA_VALUE_LIMIT = 10; + +/** + * Application budget for the bound parameters of one discovery statement, under + * SQLite's default variable limit (32,766) with headroom. A filter and a match + * each pass the parser's own limits independently, and the match is bound + * twice (once to rank keys, once to aggregate), so the combination is checked + * here, before any statement is prepared. + */ +export const METADATA_SQL_BINDING_BUDGET = 30_000; + +/** Raised by listMetadata when an option is outside its domain. */ +export class MetadataOptionError extends Error { + readonly option: string; + + constructor(option: string, expected: string, received: unknown) { + super(`Invalid ${option}: expected ${expected}, received ${formatReceived(received)}`); + this.name = "MetadataOptionError"; + this.option = option; + } +} + +/** Raised by listMetadata when `filter` and `match` together bind more SQL parameters than the budget. */ +export class MetadataBindingBudgetError extends Error { + readonly bindings: number; + readonly budget: number; + + constructor(bindings: number, budget: number) { + super(`filter and match together bind ${bindings} SQL parameters, over the budget of ${budget}. Shorten their membership lists.`); + this.name = "MetadataBindingBudgetError"; + this.bindings = bindings; + this.budget = budget; + } +} + +/** The window and ordering options with defaults applied and domains checked. */ +interface ResolvedWindow { + keyLimit: number; + keyOffset: number; + valueLimit: number; + valueOffset: number; + sort: "count" | "value"; + minCount: number; +} + +/** + * The value rows discovery aggregates over. `withSql` defines the `eligible` + * CTE (and, once a key window is applied, `selected_keys`); `fromSql` joins + * `document_metadata_values mv` to them, narrowed by the match when one is + * set. Each query supplies its own SELECT list, WHERE, and GROUP BY around + * these two parts. + */ +interface Region { + withSql: string; + withParams: SQLiteValue[]; + fromSql: string; + fromParams: SQLiteValue[]; +} + +type ValueRow = { + key: string; + value_type: MetadataValueType; + text_value: string | null; + number_value: number | null; + boolean_value: number | null; +}; + +/** + * Summarize metadata keys, types, and value counts for the documents in + * scope. Discovery sees exactly what filtering sees: the same extraction gate, + * active-document rule, and collection scope, so every value reported here is + * a value an `eq` filter can match. + * + * `filter` selects which documents are counted. `match` selects which of their + * metadata entries are reported, with the same grammar evaluated per entry. + * Counts are documents, not values: a document with `topics: [a, b]` + * contributes one to each. + * + * The key window limits which keys receive detailed aggregation and the value + * window limits the values returned per key and type. Exact counts, medians, + * and ordering still process the relevant rows. These windows do not bound + * database work, temporary storage, or the contributing collection lists. + * + * The report is read inside one deferred transaction, so every count comes + * from the same database snapshot even while another connection writes. + */ +export function listMetadata(db: Database, options: ListMetadataOptions = {}): ListMetadataResult { + const window = resolveWindow(options); + // A match built in code gets the same checks as one parsed from JSON. + const match = options.match ? compileMetadataMatch(parseMetadataMatch(options.match), "mv") : undefined; + return db.transaction(() => readMetadataReport(db, options, match, window))(); +} + +function readMetadataReport( + db: Database, + options: ListMetadataOptions, + match: CompiledMetadataFilter | undefined, + window: ResolvedWindow, +): ListMetadataResult { + const collectionNames = options.collection === undefined ? undefined : [options.collection].flat(); + const eligible = buildEligibleCte(collectionNames, options.filter); + const matchedRegion = buildRegion(eligible, match); + const keyWindowRegion = buildKeyWindowRegion(matchedRegion, window.keyLimit, window.keyOffset); + assertBindingBudget(keyWindowRegion, matchedRegion); + + const result: ListMetadataResult = { + documents: countActiveDocuments(db, collectionNames), + totalKeys: 0, + keys: [], + remainingKeys: 0, + }; + if (options.filter) result.filteredDocuments = countEligibleDocuments(db, eligible); + + result.totalKeys = countKeys(db, matchedRegion); + if (result.totalKeys === 0) return result; + + const typeStats = queryTypeStats(db, keyWindowRegion); + if (typeStats.length === 0) return result; + + // Carry the selected names through the report. MATERIALIZED prevents repeat + // ranking within one statement, but cannot share it across statements. + const keyNames = [...new Set(typeStats.map(row => row.key))]; + const region = narrowRegionToKeys(matchedRegion, keyNames); + const medianByKey = queryNumberMedians(db, region); + const typeSummaryByKeyType = new Map(); + const typeSummariesByKey = new Map(); + + for (const row of typeStats) { + const typeSummary: MetadataKeyTypeSummary = { + type: row.value_type, + multiValued: false, + documents: row.documents, + distinctValues: 0, + values: [], + remainingValues: 0, + collections: (JSON.parse(row.collections) as string[]).sort(compareBinary), + }; + if (row.value_type === "number") { + typeSummary.range = { min: row.min_value!, median: medianByKey.get(row.key)!, max: row.max_value! }; + } + typeSummaryByKeyType.set(`${row.key}\0${row.value_type}`, typeSummary); + + const typeSummaries = typeSummariesByKey.get(row.key) ?? []; + typeSummaries.push(typeSummary); + typeSummariesByKey.set(row.key, typeSummaries); + } + + // Array-ness describes the key, not the matched entries, so it is read over + // every entry of the keys in the window. + const keyRegion = buildRegion(eligible, buildKeySetPredicate([...typeSummariesByKey.keys()])); + for (const row of queryTypeArrayness(db, keyRegion)) { + const typeSummary = typeSummaryByKeyType.get(`${row.key}\0${row.value_type}`); + if (typeSummary) typeSummary.multiValued = row.multi_valued === 1; + } + + for (const row of queryValueCounts(db, region, window)) { + const typeSummary = typeSummaryByKeyType.get(`${row.key}\0${row.value_type}`); + if (!typeSummary) continue; + typeSummary.distinctValues = row.distinct_values; + if (row.in_window) typeSummary.values.push({ value: scalarOf(row), documents: row.documents }); + } + + for (const typeSummary of typeSummaryByKeyType.values()) { + typeSummary.remainingValues = Math.max(0, typeSummary.distinctValues - window.valueOffset - typeSummary.values.length); + } + + // Keys arrive in window rank order and stay in it. A document holds a key + // under exactly one type (arrays are homogeneous), so per-type document + // counts partition the key's documents. + for (const [key, typeSummaries] of typeSummariesByKey) { + typeSummaries.sort(compareTypeSummaries); + result.keys.push({ + key, + documents: typeSummaries.reduce((sum, typeSummary) => sum + typeSummary.documents, 0), + types: typeSummaries, + }); + } + result.remainingKeys = Math.max(0, result.totalKeys - window.keyOffset - result.keys.length); + + return result; +} + +/** + * Every collection's key names, coverage, and types, in coverage order, + * windowed to the first `keyLimit` keys per collection. One GROUP BY, no + * values: what `collection list`, `status`, and the MCP status tool print so + * a first look reveals that metadata exists. + */ +export function listMetadataCollectionSummaries(db: Database, keyLimit: number = DEFAULT_METADATA_KEY_LIMIT): Map { + return queryMetadataOverviews(db, buildEligibleCte(undefined, undefined), keyLimit); +} + +interface MetadataKeyCoverage { + collection: string; + key: string; + documents: number; + string_documents: number; + number_documents: number; + boolean_documents: number; + total_keys: number; +} + +function queryMetadataOverviews(db: Database, eligible: Region, keyLimit: number): Map { + const region = buildRegion(eligible); + // Ordinal zero represents a document/key once. Extraction guarantees one + // homogeneous type per key, so arrays need neither value reads nor DISTINCT. + const coverages = db.prepare(` + ${region.withSql}, + key_coverages AS ( + SELECT e.collection AS collection, mv.key, COUNT(*) AS documents, + SUM(mv.value_type = 'string') AS string_documents, + SUM(mv.value_type = 'number') AS number_documents, + SUM(mv.value_type = 'boolean') AS boolean_documents, + COUNT(*) OVER (PARTITION BY e.collection) AS total_keys, + ROW_NUMBER() OVER (PARTITION BY e.collection ORDER BY COUNT(*) DESC, mv.key) AS rank + ${region.fromSql} + WHERE mv.ordinal = 0 + GROUP BY e.collection, mv.key + ) + SELECT * FROM key_coverages WHERE rank <= ? OR ? = -1 ORDER BY collection, rank + `).all(...region.withParams, ...region.fromParams, Number.isFinite(keyLimit) ? keyLimit : -1, Number.isFinite(keyLimit) ? keyLimit : -1) as MetadataKeyCoverage[]; + + const overviews = new Map(); + for (const coverage of coverages) { + const overview = overviews.get(coverage.collection) ?? { totalKeys: coverage.total_keys, keys: [] }; + const typeCounts: [MetadataValueType, number][] = [["string", coverage.string_documents], ["number", coverage.number_documents], ["boolean", coverage.boolean_documents]]; + typeCounts.sort((left, right) => right[1] - left[1] || compareBinary(left[0], right[0])); + overview.keys.push({ key: coverage.key, documents: coverage.documents, types: typeCounts.filter(([, count]) => count > 0).map(([type]) => type) }); + overviews.set(coverage.collection, overview); + } + return overviews; +} + +function compareTypeSummaries(a: MetadataKeyTypeSummary, b: MetadataKeyTypeSummary): number { + return b.documents - a.documents || compareBinary(a.type, b.type); +} + +/** SQLite BINARY order compares UTF-8 bytes, not JavaScript's UTF-16 code units. */ +function compareBinary(left: string, right: string): number { + return buffer.Buffer.compare(buffer.Buffer.from(left), buffer.Buffer.from(right)); +} + +/** Distinct metadata keys declared by active, extracted documents in scope. */ +export function countMetadataKeys(db: Database, collectionNames?: string[]): number { + const region = buildRegion(buildEligibleCte(collectionNames, undefined)); + const row = db.prepare(` + ${region.withSql} + SELECT COUNT(DISTINCT mv.key) AS c ${region.fromSql} WHERE mv.ordinal = 0 + `).get(...region.withParams, ...region.fromParams) as { c: number }; + return row.c; +} + +/** Active, extracted documents in scope that declare at least one metadata key. */ +export function countDocumentsWithMetadata(db: Database, collectionNames?: string[]): number { + const eligible = buildEligibleCte(collectionNames, undefined); + const row = db.prepare(` + ${eligible.withSql} + SELECT COUNT(*) AS c FROM eligible e + WHERE EXISTS (SELECT 1 FROM document_metadata_values mv WHERE mv.document_id = e.document_id) + `).get(...eligible.withParams) as { c: number }; + return row.c; +} + +// ----------------------------------------------------------------------------- +// Options +// ----------------------------------------------------------------------------- + +function resolveWindow(options: ListMetadataOptions): ResolvedWindow { + return { + keyLimit: readLimit(options.keyLimit, "keyLimit", DEFAULT_METADATA_KEY_LIMIT), + keyOffset: readOffset(options.keyOffset, "keyOffset"), + valueLimit: readLimit(options.valueLimit, "valueLimit", DEFAULT_METADATA_VALUE_LIMIT), + valueOffset: readOffset(options.valueOffset, "valueOffset"), + sort: readSort(options.sort), + minCount: readPositiveInteger(options.minCount, "minCount", 1), + }; +} + +// Safe integers are exactly representable in JavaScript. The drivers may bind +// them as REAL, so SQL arithmetic must cast explicitly when it needs INTEGER. +function readLimit(value: unknown, option: string, fallback: number): number { + if (value === undefined) return fallback; + if (value === Infinity) return Infinity; + if (isSafeInteger(value) && value >= 1) return value; + throw new MetadataOptionError(option, "a positive integer or Infinity", value); +} + +function readOffset(value: unknown, option: string): number { + if (value === undefined) return 0; + if (isSafeInteger(value) && value >= 0) return value; + throw new MetadataOptionError(option, "a non-negative integer", value); +} + +function readPositiveInteger(value: unknown, option: string, fallback: number): number { + if (value === undefined) return fallback; + if (isSafeInteger(value) && value >= 1) return value; + throw new MetadataOptionError(option, "a positive integer", value); +} + +function isSafeInteger(value: unknown): value is number { + return Number.isSafeInteger(value); +} + +function readSort(value: unknown): "count" | "value" { + if (value === undefined) return "count"; + if (value === "count" || value === "value") return value; + throw new MetadataOptionError("sort", "'count' or 'value'", value); +} + +function formatReceived(value: unknown): string { + if (typeof value === "string") return JSON.stringify(value); + if (typeof value === "number" || typeof value === "boolean" || value === null) return String(value); + return typeof value; +} + +// ----------------------------------------------------------------------------- +// Regions +// ----------------------------------------------------------------------------- + +/** + * Documents that filtered search can see: active, with a current, error-free + * extraction, in scope, and passing the filter. Lists bind as one JSON + * parameter so the SQLite variable limit never applies. The filter is + * compiled for the corpus: every statement that reads this CTE joins it to + * metadata rows, and a correlated predicate would be re-run for each of + * them, where a document set is built once per statement and probed. + */ +function buildEligibleCte(collectionNames: string[] | undefined, filter: MetadataFilter | undefined): Region { + const withParams: SQLiteValue[] = []; + let withSql = ` + WITH eligible AS ( + SELECT d.id AS document_id, d.collection + FROM documents d + JOIN document_metadata dm ON dm.document_id = d.id + WHERE d.active = 1 + AND dm.extraction_version = ${METADATA_EXTRACTION_VERSION} + AND dm.extraction_error IS NULL`; + + if (collectionNames) { + withSql += ` + AND d.collection IN (SELECT value FROM json_each(?))`; + withParams.push(JSON.stringify(collectionNames)); + } + + if (filter) { + const compiledFilter = compileMetadataFilter(filter, "d", "corpus"); + withSql += ` + AND ${compiledFilter.sql}`; + withParams.push(...compiledFilter.params); + } + + withSql += ` + )`; + + return { withSql, withParams, fromSql: "", fromParams: [] }; +} + +/** + * Join eligible documents' values, narrowed to the entries an optional + * predicate over `mv` admits. The predicate lives in the JOIN so queries can + * add their own WHERE. + */ +function buildRegion(eligible: Region, predicate?: CompiledMetadataFilter): Region { + const region: Region = { + withSql: eligible.withSql, + withParams: [...eligible.withParams], + fromSql: ` + FROM document_metadata_values mv + JOIN eligible e ON e.document_id = mv.document_id`, + fromParams: [], + }; + + if (predicate) { + region.fromSql += ` + AND ${predicate.sql}`; + region.fromParams.push(...predicate.params); + } + + return region; +} + +/** + * Narrow a region to one window of its keys, in report order (documents + * descending, then key in BINARY collation), each carrying its `rank` so the + * order survives the joins and GROUP BYs that read it. The window is a CTE + * every later aggregation joins. It is MATERIALIZED because a small LIMIT + * otherwise invites the planner to inline it as a coroutine and re-run the + * ranking once per value row. SQLite reads `LIMIT -1` as no limit. + */ +function buildKeyWindowRegion(region: Region, keyLimit: number, keyOffset: number): Region { + const withSql = `${region.withSql}, + selected_keys AS MATERIALIZED ( + SELECT mv.key, ROW_NUMBER() OVER (ORDER BY COUNT(DISTINCT mv.document_id) DESC, mv.key ASC) AS rank + ${region.fromSql} + GROUP BY mv.key + ORDER BY rank + LIMIT ? OFFSET ? + )`; + + return { + withSql, + withParams: [...region.withParams, ...region.fromParams, Number.isFinite(keyLimit) ? keyLimit : -1, keyOffset], + fromSql: `${region.fromSql} + JOIN selected_keys sk ON sk.key = mv.key`, + fromParams: [...region.fromParams], + }; +} + +/** Reuse the report's selected keys without repeating coverage and ranking. */ +function narrowRegionToKeys(region: Region, keyNames: string[]): Region { + const predicate = buildKeySetPredicate(keyNames); + return { + ...region, + fromSql: `${region.fromSql} AND ${predicate.sql}`, + fromParams: [...region.fromParams, ...predicate.params], + }; +} + +/** Every entry of these keys, bound as one JSON parameter. */ +function buildKeySetPredicate(keyNames: string[]): CompiledMetadataFilter { + return { sql: "mv.key IN (SELECT value FROM json_each(?))", params: [JSON.stringify(keyNames)] }; +} + +/** Selected names as JSON, minCount, offset, and the two endpoint operands. */ +const VALUE_QUERY_BINDINGS = 5; + +function assertBindingBudget(keyWindowRegion: Region, matchedRegion: Region): void { + const bindings = Math.max( + keyWindowRegion.withParams.length + keyWindowRegion.fromParams.length, + matchedRegion.withParams.length + matchedRegion.fromParams.length + VALUE_QUERY_BINDINGS, + ); + if (bindings > METADATA_SQL_BINDING_BUDGET) throw new MetadataBindingBudgetError(bindings, METADATA_SQL_BINDING_BUDGET); +} + +// ----------------------------------------------------------------------------- +// Queries +// ----------------------------------------------------------------------------- + +function countActiveDocuments(db: Database, collectionNames: string[] | undefined): number { + let sql = `SELECT COUNT(*) AS c FROM documents d WHERE d.active = 1`; + const params: SQLiteValue[] = []; + if (collectionNames) { + sql += ` AND d.collection IN (SELECT value FROM json_each(?))`; + params.push(JSON.stringify(collectionNames)); + } + const row = db.prepare(sql).get(...params) as { c: number }; + return row.c; +} + +function countEligibleDocuments(db: Database, eligible: Region): number { + const row = db.prepare(`${eligible.withSql} SELECT COUNT(*) AS c FROM eligible`).get(...eligible.withParams) as { c: number }; + return row.c; +} + +function countKeys(db: Database, region: Region): number { + const row = db.prepare(` + ${region.withSql} + SELECT COUNT(DISTINCT mv.key) AS c + ${region.fromSql} + `).get(...region.withParams, ...region.fromParams) as { c: number }; + return row.c; +} + +type TypeStatsRow = { key: string; value_type: MetadataValueType; documents: number; min_value: number | null; max_value: number | null; collections: string }; + +/** + * Per key and type over a key-window region, in the window's rank order. + * Collection names are sorted after reading: aggregate ORDER BY would + * require SQLite 3.44, newer than QMD's declared Bun minimum. + */ +function queryTypeStats(db: Database, region: Region): TypeStatsRow[] { + return db.prepare(` + ${region.withSql} + SELECT mv.key, mv.value_type, + COUNT(DISTINCT mv.document_id) AS documents, + MIN(mv.number_value) AS min_value, + MAX(mv.number_value) AS max_value, + json_group_array(DISTINCT e.collection) AS collections + ${region.fromSql} + GROUP BY mv.key, mv.value_type + ORDER BY sk.rank, mv.value_type + `).all(...region.withParams, ...region.fromParams) as TypeStatsRow[]; +} + +type TypeArraynessRow = { key: string; value_type: MetadataValueType; multi_valued: number }; + +function queryTypeArrayness(db: Database, keyRegion: Region): TypeArraynessRow[] { + return db.prepare(` + ${keyRegion.withSql} + SELECT mv.key, mv.value_type, MAX(mv.ordinal) > 0 AS multi_valued + ${keyRegion.fromSql} + GROUP BY mv.key, mv.value_type + `).all(...keyRegion.withParams, ...keyRegion.fromParams) as TypeArraynessRow[]; +} + +type MiddleValueRow = { key: string; number_value: number }; + +/** + * Median over value rows per key: the middle row for odd counts, the midpoint + * of the two middle rows for even. The midpoint is taken in JavaScript because + * SQL's AVG sums first and the sum of two finite doubles can overflow. + */ +function queryNumberMedians(db: Database, region: Region): Map { + const middleRows = db.prepare(` + ${region.withSql} + SELECT key, number_value + FROM ( + SELECT mv.key, mv.number_value, + ROW_NUMBER() OVER (PARTITION BY mv.key ORDER BY mv.number_value) AS position, + COUNT(*) OVER (PARTITION BY mv.key) AS total + ${region.fromSql} + WHERE mv.value_type = 'number' + ) + WHERE position IN ((total + 1) / 2, (total + 2) / 2) + ORDER BY key, number_value + `).all(...region.withParams, ...region.fromParams) as MiddleValueRow[]; + + const medianByKey = new Map(); + for (const row of middleRows) { + const lower = medianByKey.get(row.key); + medianByKey.set(row.key, lower === undefined ? row.number_value : midpointOf(lower, row.number_value)); + } + return medianByKey; +} + +/** + * The correctly rounded midpoint when the sum is representable, which also + * keeps adjacent subnormals exact. Halving each side first is reserved for the + * overflow region, where both operands are far from underflow. + */ +function midpointOf(lower: number, upper: number): number { + const sum = lower + upper; + if (Number.isFinite(sum)) return sum / 2; + return lower / 2 + upper / 2; +} + +type ValueCountRow = ValueRow & { documents: number; distinct_values: number; in_window: number }; + +/** + * Group distinct values once for both the total and the window. The first + * row of each partition carries its total even when the window is past the + * end. Only rows marked in_window become returned values. + * Only one typed column is non-null per partition, so ordering all three is stable. + */ +function queryValueCounts(db: Database, region: Region, window: ResolvedWindow): ValueCountRow[] { + const valueOrder = "mv.text_value, mv.number_value, mv.boolean_value"; + const rankOrder = window.sort === "count" ? `COUNT(DISTINCT mv.document_id) DESC, ${valueOrder}` : valueOrder; + const params = [...region.withParams, ...region.fromParams, window.minCount, window.valueOffset]; + let windowSql = "rank > ?"; + // Both drivers may bind REALs. Cast before adding so the endpoint remains + // exact even when it exceeds JavaScript's safe-integer range. + if (Number.isFinite(window.valueLimit)) { + windowSql += " AND rank <= CAST(? AS INTEGER) + CAST(? AS INTEGER)"; + params.push(window.valueOffset, window.valueLimit); + } + + return db.prepare(` + ${region.withSql}, + value_counts AS ( + SELECT mv.key, mv.value_type, mv.text_value, mv.number_value, mv.boolean_value, + COUNT(DISTINCT mv.document_id) AS documents, + COUNT(*) OVER (PARTITION BY mv.key, mv.value_type) AS distinct_values, + ROW_NUMBER() OVER (PARTITION BY mv.key, mv.value_type ORDER BY ${rankOrder}) AS rank + ${region.fromSql} + GROUP BY mv.key, mv.value_type, mv.text_value, mv.number_value, mv.boolean_value + HAVING COUNT(DISTINCT mv.document_id) >= ? + ), + value_window AS ( + SELECT *, (${windowSql}) AS in_window FROM value_counts + ) + SELECT key, value_type, text_value, number_value, boolean_value, documents, distinct_values, in_window + FROM value_window + WHERE rank = 1 OR in_window + ORDER BY key, value_type, rank + `).all(...params) as ValueCountRow[]; +} + +function scalarOf(row: ValueRow): MetadataScalar { + if (row.value_type === "string") return row.text_value!; + if (row.value_type === "number") return row.number_value!; + return row.boolean_value === 1; +} diff --git a/src/metadata.ts b/src/metadata.ts index 602bebfd1..5ddc9082e 100644 --- a/src/metadata.ts +++ b/src/metadata.ts @@ -28,6 +28,9 @@ import YAML from "yaml"; export type MetadataScalar = string | number | boolean; +/** The stored type of one metadata value. Arrays are homogeneous, so a key holds one type per document. */ +export type MetadataValueType = "string" | "number" | "boolean"; + export type MetadataScalarArray = | readonly string[] | readonly number[] diff --git a/src/store-migrations.ts b/src/store-migrations.ts new file mode 100644 index 000000000..cb938ccf0 --- /dev/null +++ b/src/store-migrations.ts @@ -0,0 +1,294 @@ +/** + * store-migrations.ts - PRAGMA user_version steps applied when a store opens. + * + * Version 1 installs the FTS sync triggers. Version 2 moves every vector out + * of the legacy `vectors_vec` table (one row per chunk, keyed by hash_seq) + * into the layout `vec-layout.ts` describes: one row per (chunk, active + * collection) in a vec0 table partitioned by collection id. + * + * The copy walks the legacy table's chunk blobs in chunk_id order, one chunk + * per IMMEDIATE transaction, and keeps the last copied chunk_id in + * store_config. A killed process resumes from that cursor, and two processes + * that open the same database take turns on chunks because each reads the + * cursor inside its own write transaction. A verification pass that copies + * rows the walk missed, the version stamp and the drop of the legacy table + * share one IMMEDIATE transaction, so that drop is the flip. + */ + +import type { Database } from "./db.js"; +import { + LEGACY_VEC_TABLE, + PartitionWriter, + VEC_ROWS_TABLE, + VEC_TABLE, + createPartitionedVecTable, + missingPartitionRows, + parseDimensions, + vecLayout, + vecTableReadable, + type ReadableVecLayout, +} from "./vec-layout.js"; + +export const FTS_SYNC_TRIGGERS_VERSION = 1; +export const VECTOR_PARTITION_VERSION = 2; + +const CURSOR_KEY = "vector_partition_cursor"; + +export function getUserVersion(db: Database): number { + const row = db.prepare(`PRAGMA user_version`).get() as Record | undefined; + const value = row ? Object.values(row)[0] : 0; + return typeof value === "number" ? value : Number(value) || 0; +} + +/** + * Run `step` and stamp `version` in one IMMEDIATE transaction. The + * double-checked read makes concurrent first opens of one database apply the + * step once: busy_timeout serializes the transactions, and the loser sees the + * version the winner stamped. + */ +export function applyVersionedStep(db: Database, version: number, step: () => void): void { + if (getUserVersion(db) >= version) return; + db.exec(`BEGIN IMMEDIATE`); + try { + if (getUserVersion(db) < version) { + step(); + db.exec(`PRAGMA user_version = ${version}`); + } + db.exec(`COMMIT`); + } catch (err) { + db.exec(`ROLLBACK`); + throw err; + } +} + +export type VectorMigrationPhase = "copy" | "verify" | "flip" | "vacuum" | "done"; + +export type VectorMigrationProgress = { + phase: VectorMigrationPhase; + /** Legacy rows processed so far. */ + copied: number; + /** Legacy rows at the start of this process's run. */ + total: number; +}; + +export type VectorMigrationOptions = { + sqliteVecAvailable: boolean; + onProgress?: (progress: VectorMigrationProgress) => void; +}; + +export type VectorMigrationResult = "applied" | "deferred"; + +interface LegacyChunk { + chunkId: number; + size: number; + validity: Uint8Array; + rowids: Uint8Array; + vectors: Uint8Array; +} + +function liveSlots(chunk: { size: number; validity: Uint8Array }): number[] { + const slots: number[] = []; + for (let i = 0; i < chunk.size; i++) { + if ((chunk.validity[i >> 3]! >> (i & 7)) & 1) slots.push(i); + } + return slots; +} + +function splitHashSeq(key: string): { hash: string; seq: number } | null { + const at = key.lastIndexOf("_"); + if (at <= 0) return null; + const seq = Number(key.slice(at + 1)); + return Number.isInteger(seq) ? { hash: key.slice(0, at), seq } : null; +} + +function readCursor(db: Database): number { + const row = db.prepare(`SELECT value FROM store_config WHERE key = ?`).get(CURSOR_KEY) as { value: string } | undefined; + const value = row ? Number(row.value) : Number.NaN; + return Number.isFinite(value) ? value : -1; +} + +function writeCursor(db: Database, chunkId: number): void { + db.prepare(`INSERT INTO store_config (key, value) VALUES (?, ?) ON CONFLICT(key) DO UPDATE SET value = excluded.value`) + .run(CURSOR_KEY, String(chunkId)); +} + +function copyLegacyChunks( + db: Database, + legacy: ReadableVecLayout, + dimensions: number, + onProgress: (copied: number) => void, +): void { + const writer = new PartitionWriter(db); + const nextChunk = db.prepare(`SELECT chunk_id AS chunkId, size, validity, rowids FROM ${legacy.shadow.chunks} WHERE chunk_id > ? ORDER BY chunk_id LIMIT 1`); + const vectorsOf = db.prepare(`SELECT vectors FROM ${legacy.shadow.vectorChunks} WHERE rowid = ?`); + const keyOf = db.prepare(`SELECT id FROM ${legacy.shadow.rowids} WHERE rowid = ?`); + const hasChunkRow = db.prepare(`SELECT 1 FROM content_vectors WHERE hash = ? AND seq = ?`); + const bytesPerVector = dimensions * 4; + let copied = 0; + + const copyOne = db.transaction((): boolean => { + if (vecLayout(db).kind !== "legacy") return false; + const chunk = nextChunk.get(readCursor(db)) as Omit | undefined; + if (!chunk) return false; + const blob = (vectorsOf.get(chunk.chunkId) as { vectors: Uint8Array } | undefined)?.vectors; + const rowids = new DataView(chunk.rowids.buffer, chunk.rowids.byteOffset, chunk.rowids.byteLength); + for (const slot of liveSlots(chunk)) { + copied++; + const key = (keyOf.get(rowids.getBigInt64(slot * 8, true)) as { id: string } | undefined)?.id; + const parsed = key === undefined ? null : splitHashSeq(key); + if (!parsed || !blob || blob.byteLength < (slot + 1) * bytesPerVector) continue; + if (!hasChunkRow.get(parsed.hash, parsed.seq)) continue; + const embedding = blob.subarray(slot * bytesPerVector, (slot + 1) * bytesPerVector); + writer.writeToActiveCollections(parsed.hash, parsed.seq, embedding); + } + writeCursor(db, chunk.chunkId); + return true; + }); + + while (copyOne.immediate()) { + onProgress(copied); + } +} + +/** + * Rows the chunk walk can miss: a pre-upgrade `qmd cleanup` repacking the + * legacy table moves live rows into its newest chunk while the walk is past + * it, and a process on the old layout keeps writing into chunks the walk has + * left behind. They are copied by key from the legacy table. Runs inside the + * caller's write transaction. + */ +function copyStragglers(db: Database, legacy: ReadableVecLayout): void { + const legacyVector = db.prepare(`SELECT embedding FROM ${legacy.table} WHERE hash_seq = ?`); + const writer = new PartitionWriter(db); + for (const missing of missingPartitionRows(db)) { + const row = legacyVector.get(`${missing.hash}_${missing.seq}`) as { embedding: Uint8Array } | undefined; + if (row) writer.writeToCollection(missing.hash, missing.seq, missing.collection, row.embedding); + } +} + +/** + * A content_vectors row with no vector in any partition is a chunk to embed + * again; without the row the pending detector picks the hash up. + */ +export function deleteVectorlessContentVectors(db: Database): number { + return db.prepare(` + DELETE FROM content_vectors + WHERE NOT EXISTS (SELECT 1 FROM ${VEC_ROWS_TABLE} vr WHERE vr.hash = content_vectors.hash AND vr.seq = content_vectors.seq) + `).run().changes; +} + +/** + * The straggler copy, the vectorless cleanup and the drop hold one write + * lock: a legacy row written between them by a process on the old layout + * would otherwise leave a content_vectors row whose only vector is in the + * dropped table, which the pending detector never re-embeds. `onFlip` runs + * with that lock held, after the verification and before the drop. + */ +function flipToPartitioned(db: Database, legacy: ReadableVecLayout, onFlip: () => void): void { + db.transaction(() => { + if (vecLayout(db).kind !== "legacy") return; + copyStragglers(db, legacy); + deleteVectorlessContentVectors(db); + onFlip(); + db.prepare(`DELETE FROM store_config WHERE key = ?`).run(CURSOR_KEY); + db.exec(`PRAGMA user_version = ${Math.max(getUserVersion(db), VECTOR_PARTITION_VERSION)}`); + db.exec(`DROP TABLE IF EXISTS ${LEGACY_VEC_TABLE}`); + }).immediate(); +} + +/** Dimensions of the partitioned table; undefined when there is no table. */ +function partitionedTableDimensions(db: Database): number | null | undefined { + const row = db.prepare(`SELECT sql FROM sqlite_master WHERE type = 'table' AND name = ?`).get(VEC_TABLE) as { sql: string } | undefined; + return row ? parseDimensions(row.sql) : undefined; +} + +/** + * Check and create share one IMMEDIATE transaction with a double-checked + * read, matching applyVersionedStep: concurrent first openers create the + * partitioned table once and the losers reuse it. + */ +function ensurePartitionedTable(db: Database, dimensions: number): void { + const exists = (): boolean => { + const existing = partitionedTableDimensions(db); + if (existing === undefined) return false; + if (existing !== dimensions) { + throw new Error( + `Vector migration found a ${VEC_TABLE} table with ${existing ?? "unknown"} dimensions while the legacy ${LEGACY_VEC_TABLE} table holds ${dimensions}d vectors` + ); + } + return true; + }; + if (exists()) return; + db.transaction(() => { + if (!exists()) createPartitionedVecTable(db, dimensions); + }).immediate(); +} + +/** + * The version 2 step. Returns "deferred" when the legacy table cannot be read + * because sqlite-vec is not loaded: the version stays behind and the next open + * with the extension retries, while FTS keeps working. The step keys on the + * legacy table rather than the version so that a legacy table an older build + * created on an already-stamped database is copied too. + */ +export function migrateVectorLayout(db: Database, options: VectorMigrationOptions): VectorMigrationResult { + const layout = vecLayout(db); + if (layout.kind !== "legacy") { + applyVersionedStep(db, VECTOR_PARTITION_VERSION, () => {}); + return "applied"; + } + if (!options.sqliteVecAvailable || !vecTableReadable(db, layout)) return "deferred"; + if (!layout.keyedByHashSeq || layout.dimensions === null) { + applyVersionedStep(db, VECTOR_PARTITION_VERSION, () => { + db.exec(`DROP TABLE IF EXISTS ${LEGACY_VEC_TABLE}`); + }); + return "applied"; + } + + const dimensions = layout.dimensions; + const total = (db.prepare(`SELECT COUNT(*) AS c FROM ${layout.shadow.rowids}`).get() as { c: number }).c; + const report = (phase: VectorMigrationPhase, copied: number) => options.onProgress?.({ phase, copied, total }); + + ensurePartitionedTable(db, dimensions); + copyLegacyChunks(db, layout, dimensions, (copied) => report("copy", copied)); + if (vecLayout(db).kind === "legacy") { + report("verify", total); + flipToPartitioned(db, layout, () => report("flip", total)); + report("vacuum", total); + try { + db.exec(`VACUUM`); + } catch (err) { + console.warn(`VACUUM after the vector migration did not run (${err instanceof Error ? err.message : String(err)}); run 'qmd cleanup' to reclaim the space.`); + } + } + report("done", total); + return "applied"; +} + +export type StoreMigrationDeps = VectorMigrationOptions & { + installFtsSyncTriggers: (db: Database) => void; +}; + +/** + * The vector step also runs on a stamped database that holds a legacy table: + * an older build recreates `vectors_vec` on any database it embeds into, and + * the resolver reports that table as the layout until it is gone. + */ +export function runStoreMigrations(db: Database, deps: StoreMigrationDeps): void { + applyVersionedStep(db, FTS_SYNC_TRIGGERS_VERSION, () => deps.installFtsSyncTriggers(db)); + if (getUserVersion(db) < VECTOR_PARTITION_VERSION) { + migrateVectorLayout(db, deps); + return; + } + // A legacy table on a stamped store came from an older build; copying it + // here is a repair, so a failure (a dimension mismatch with the partitioned + // table) degrades vector search instead of blocking every open. The embed + // path raises the actionable error when it next runs. + if (vecLayout(db).kind === "legacy") { + try { + migrateVectorLayout(db, deps); + } catch (err) { + console.warn(`Legacy vector table left in place: ${err instanceof Error ? err.message : String(err)}`); + } + } +} diff --git a/src/store.ts b/src/store.ts index db63d2a9e..4a46ed863 100644 --- a/src/store.ts +++ b/src/store.ts @@ -12,7 +12,41 @@ */ import { openDatabase, loadSqliteVec } from "./db.js"; -import type { Database } from "./db.js"; +import { + PartitionWriter, + VEC_COLLECTION_IDS_TABLE, + VEC_ROWS_TABLE, + VEC_TABLE, + activeCollectionsOfHash, + allocateCollectionId, + createPartitionedVecTable, + createVectorMetadataTables, + deleteCollectionId, + deletePartitionRows, + deletePartitionRowsOfHash, + hasVectorIndex, + missingPartitionRows, + partitionRowKey, + renameCollectionId, + resolveCollectionId, + resolveCollectionIds, + rowidList, + storedEmbeddingLookup, + upsertPartitionVector, + vecInteger, + vecLayout, + vecTableReadable, + type MissingPartitionRow, + type VecLayout, +} from "./vec-layout.js"; +import { + FTS_SYNC_TRIGGERS_VERSION, + getUserVersion, + migrateVectorLayout, + runStoreMigrations, + type VectorMigrationProgress, +} from "./store-migrations.js"; +import type { Database, SQLiteValue } from "./db.js"; import picomatch from "picomatch"; import { createHash } from "crypto"; import { readFileSync, realpathSync, statSync, mkdirSync } from "node:fs"; @@ -44,7 +78,9 @@ import { syncDocumentMetadata, countDocumentsPendingMetadata, getMetadataByFilepath, + listMetadataCollectionSummaries, parseMetadataJson, + type MetadataKeyOverview, } from "./metadata-store.js"; // ============================================================================= @@ -101,6 +137,22 @@ export function splitGlobMask(mask: string): string[] { } export const DEFAULT_MULTI_GET_MAX_BYTES = 64 * 1024; // 64KB + +/** + * Characters of a document body that search results, the vector body lookup + * and getHashesForEmbedding return (SQLite's substr counts characters), so one + * very large document cannot put its whole text on the heap per result. + */ +const BODY_CAP_CHARS = 262_144; + +/** + * SQL for the first BODY_CAP_CHARS characters of a body column. substr() + * stops at an embedded NUL, so a body whose UTF-8 encoding fits the cap is + * returned whole and only a longer one goes through substr(). + */ +function cappedBodySql(column: string): string { + return `CASE WHEN length(CAST(${column} AS BLOB)) <= ${BODY_CAP_CHARS} THEN ${column} ELSE substr(${column}, 1, ${BODY_CAP_CHARS}) END`; +} export const DEFAULT_EMBED_MAX_DOCS_PER_BATCH = 64; export const DEFAULT_EMBED_MAX_BATCH_BYTES = 64 * 1024 * 1024; // 64MB export const DEFAULT_EMBED_MAX_DURATION_MS = 30 * 60 * 1000; // 30 minutes; see EmbedOptions.maxDurationMs @@ -871,10 +923,6 @@ const CJK_CHAR_PATTERN = /[\p{Script=Han}\p{Script=Hiragana}\p{Script=Katakana}\ const CJK_RUN_PATTERN = /[\p{Script=Han}\p{Script=Hiragana}\p{Script=Katakana}\p{Script=Hangul}]+/gu; const FTS_CJK_NORMALIZED_VERSION = "1"; -// Bump when any FTS sync trigger body in applyFtsSyncTriggers changes, so the -// new definition is reapplied to existing databases on next open. -const STORE_SCHEMA_VERSION = 1; - /** * FTS5's unicode61 tokenizer does not segment CJK text into searchable words. * Normalize CJK runs by spacing every character so exact CJK queries can be @@ -901,12 +949,6 @@ function sanitizeFTS5Phrase(phrase: string): string { .join(' '); } -function getUserVersion(db: Database): number { - const row = db.prepare(`PRAGMA user_version`).get() as Record | undefined; - const value = row ? Object.values(row)[0] : 0; - return typeof value === "number" ? value : Number(value) || 0; -} - // FTS sync triggers keep documents_fts current for callers that write directly // to documents (production indexing rebuilds FTS in TypeScript to normalize CJK // first). The bodies use DROP+CREATE rather than CREATE IF NOT EXISTS so a @@ -914,9 +956,11 @@ function getUserVersion(db: Database): number { // autocommit statements, so concurrent opens of one database interleave across // connections (A drops, B drops, A creates, B creates -> "trigger already // exists"); busy_timeout serializes individual statements but not the pair. -// Gate the work behind PRAGMA user_version and apply it inside one IMMEDIATE -// transaction: the DROP+CREATE pair is atomic across connections, and a -// double-checked read skips it once any process has stamped the version. +// runStoreMigrations gates the work behind PRAGMA user_version and applies it +// inside one IMMEDIATE transaction: the DROP+CREATE pair is atomic across +// connections, and a double-checked read skips it once any process has +// stamped the version. Bump FTS_SYNC_TRIGGERS_VERSION when a trigger body +// changes so existing databases reinstall it on the next open. function installFtsSyncTriggers(db: Database): void { db.exec(`DROP TRIGGER IF EXISTS documents_ai`); db.exec(` @@ -959,19 +1003,25 @@ function installFtsSyncTriggers(db: Database): void { `); } -function applyFtsSyncTriggers(db: Database): void { - if (getUserVersion(db) >= STORE_SCHEMA_VERSION) return; - db.exec(`BEGIN IMMEDIATE`); - try { - if (getUserVersion(db) < STORE_SCHEMA_VERSION) { - installFtsSyncTriggers(db); - db.exec(`PRAGMA user_version = ${STORE_SCHEMA_VERSION}`); +/** Progress of the one-time vector layout upgrade, on stderr. */ +function vectorMigrationReporter(): (progress: VectorMigrationProgress) => void { + let startedAt: number | null = null; + const tty = Boolean(process.stderr.isTTY); + return ({ phase, copied, total }) => { + if (startedAt === null) { + startedAt = Date.now(); + process.stderr.write(`Moving ${total} vectors to the per-collection index layout (one-time upgrade)...\n`); } - db.exec(`COMMIT`); - } catch (err) { - db.exec(`ROLLBACK`); - throw err; - } + if (phase === "copy") { + if (tty) process.stderr.write(`\r copied ${copied}/${total}`); + return; + } + if (tty) process.stderr.write("\r"); + if (phase === "verify") process.stderr.write(` copied ${copied}/${total}; verifying...\n`); + else if (phase === "flip") process.stderr.write(` dropping the old vector table...\n`); + else if (phase === "vacuum") process.stderr.write(` vacuuming the index (minutes on a large index)...\n`); + else process.stderr.write(` vector index upgraded in ${Math.round((Date.now() - startedAt) / 1000)}s\n`); + }; } /** @@ -1048,7 +1098,7 @@ function recreateDocumentsFts(db: Database): void { } // Missing-table create and legacy-schema repair share one IMMEDIATE -// transaction with a double-checked read, matching applyFtsSyncTriggers: +// transaction with a double-checked read, matching applyVersionedStep: // the DROP+CREATE (or first CREATE) is atomic across connections, and // losers skip once any process has published the current table. function ensureDocumentsFtsSchema(db: Database): void { @@ -1058,10 +1108,10 @@ function ensureDocumentsFtsSchema(db: Database): void { if (!documentsFtsSchemaIsCurrent(db)) { if (documentsFtsExists(db)) { recreateDocumentsFts(db); - // recreateDocumentsFts dropped the sync triggers. applyFtsSyncTriggers + // recreateDocumentsFts dropped the sync triggers. runStoreMigrations // only reinstalls them when user_version is stale, so a DB that already // has the current user_version would otherwise be left untriggered. - if (getUserVersion(db) >= STORE_SCHEMA_VERSION) { + if (getUserVersion(db) >= FTS_SYNC_TRIGGERS_VERSION) { installFtsSyncTriggers(db); } } else { @@ -1257,6 +1307,7 @@ function initializeDatabase(db: Database): void { `); ensureContentVectorsStatusIndex(db); + createVectorMetadataTables(db); // Document metadata — extraction state plus normalized value rows for // metadata filtering. Keyed by document identity, not content hash. @@ -1283,11 +1334,30 @@ function initializeDatabase(db: Database): void { ) `); + // File sync state — mtime+size fast-path. + // Inspired by qmd-py incremental sync. + db.exec(` + CREATE TABLE IF NOT EXISTS file_sync_state ( + collection TEXT NOT NULL, + relative_path TEXT NOT NULL, + mtime_ms INTEGER NOT NULL, + size INTEGER NOT NULL, + content_hash TEXT NOT NULL, + document_id INTEGER NOT NULL, + PRIMARY KEY (collection, relative_path) + ) + `); + db.exec(`CREATE INDEX IF NOT EXISTS idx_file_sync_state_collection ON file_sync_state(collection)`); + // FTS - index filepath (collection/path), title, and content. // Do not CREATE VIRTUAL TABLE here as an autocommit statement: FTS5 // IF NOT EXISTS races under WAL (see createDocumentsFtsTable). ensureDocumentsFtsSchema(db); - applyFtsSyncTriggers(db); + runStoreMigrations(db, { + installFtsSyncTriggers, + sqliteVecAvailable: _sqliteVecAvailable === true, + onProgress: vectorMigrationReporter(), + }); rebuildFTSForCjkNormalization(db); } @@ -1471,28 +1541,41 @@ export function isSqliteVecAvailable(): boolean { return _sqliteVecAvailable === true; } +/** + * Drop the vec0 table and empty its rowid map in one commit. vector_rows ids + * are reused once the map is empty, so a map emptied without the drop hands + * new rows rowids that surviving vec0 rows still hold, and every later insert + * fails on the vec0 primary key. + */ +function dropVectorIndex(db: Database): void { + db.transaction(() => { + db.exec(`DROP TABLE IF EXISTS ${VEC_TABLE}`); + db.exec(`DELETE FROM ${VEC_ROWS_TABLE}`); + }).immediate(); +} + function ensureVecTableInternal(db: Database, dimensions: number): void { if (!_sqliteVecAvailable) { throw createSqliteVecUnavailableError( _sqliteVecUnavailableReason ?? "vector operations require a SQLite build with extension loading support" ); } - const tableInfo = db.prepare(`SELECT sql FROM sqlite_master WHERE type='table' AND name='vectors_vec'`).get() as { sql: string } | null; - if (tableInfo) { - const match = tableInfo.sql.match(/float\[(\d+)\]/); - const hasHashSeq = tableInfo.sql.includes('hash_seq'); - const hasCosine = tableInfo.sql.includes('distance_metric=cosine'); - const existingDims = match?.[1] ? parseInt(match[1], 10) : null; - if (existingDims === dimensions && hasHashSeq && hasCosine) return; - if (existingDims !== null && existingDims !== dimensions) { + let layout = vecLayout(db); + if (layout.kind === "legacy") { + migrateVectorLayout(db, { sqliteVecAvailable: true, onProgress: vectorMigrationReporter() }); + layout = vecLayout(db); + } + if (layout.kind === "partitioned") { + if (layout.dimensions === dimensions) return; + if (layout.dimensions !== null) { throw new Error( - `Embedding dimension mismatch: existing vectors are ${existingDims}d but the current model produces ${dimensions}d. ` + + `Embedding dimension mismatch: existing vectors are ${layout.dimensions}d but the current model produces ${dimensions}d. ` + `Run 'qmd embed -f' to re-embed with the new model.` ); } - db.exec("DROP TABLE IF EXISTS vectors_vec"); + dropVectorIndex(db); } - db.exec(`CREATE VIRTUAL TABLE vectors_vec USING vec0(hash_seq TEXT PRIMARY KEY, embedding float[${dimensions}] distance_metric=cosine)`); + createPartitionedVecTable(db, dimensions); } // ============================================================================= @@ -1511,6 +1594,7 @@ export type Store = { getHashesNeedingEmbedding: (model?: string) => number; getIndexHealth: (model?: string) => IndexHealthInfo; getStatus: (model?: string) => IndexStatus; + getStatusSummary: (model?: string) => IndexStatusSummary; // Caching getCacheKey: typeof getCacheKey; @@ -1610,9 +1694,152 @@ export type ReindexResult = { metadataErrors: number; }; +/** + * File sync state row — mtime+size fast-path cache. + * Inspired by qmd-py incremental sync. + */ +type FileSyncStateRow = { + relative_path: string; + mtime_ms: number; + size: number; + content_hash: string; + document_id: number; + /** 1 when the document has a metadata extraction at the current version. */ + metadata_current: number; +}; + +function getFileSyncStateMap(db: Database, collectionName: string): Map { + try { + // A row is trusted only while its document is still the active row for + // this collection, path and content. Removing a collection, or renaming it + // away, leaves rows behind; trusting those would count a file that has no + // active document as unchanged. A distrusted row is rewritten on reindex. + const stmt = db.prepare(` + SELECT s.relative_path, s.mtime_ms, s.size, s.content_hash, s.document_id, + COALESCE(dm.extraction_version = ?, 0) AS metadata_current + FROM file_sync_state s + JOIN documents d ON d.id = s.document_id + LEFT JOIN document_metadata dm ON dm.document_id = s.document_id + WHERE s.collection = ? + AND d.active = 1 + AND d.collection = s.collection + AND d.path = s.relative_path + AND d.hash = s.content_hash + `); + const map = new Map(); + // Large-result query: use iterate() to stream rows instead of .all() materializing at once + for (const r of stmt.iterate(METADATA_EXTRACTION_VERSION, collectionName) as IterableIterator) { + map.set(r.relative_path, r); + } + return map; + } catch { + // Table may not exist yet on legacy DBs — initializeDatabase creates it on next open, + // but guard here for safety. + return new Map(); + } +} + +function upsertFileSyncState(db: Database, collectionName: string, relPath: string, mtimeMs: number, size: number, contentHash: string, documentId: number): void { + try { + db.prepare(` + INSERT INTO file_sync_state(collection, relative_path, mtime_ms, size, content_hash, document_id) + VALUES (?, ?, ?, ?, ?, ?) + ON CONFLICT(collection, relative_path) DO UPDATE SET + mtime_ms = excluded.mtime_ms, + size = excluded.size, + content_hash = excluded.content_hash, + document_id = excluded.document_id + `).run(collectionName, relPath, Math.floor(mtimeMs), size, contentHash, documentId); + } catch { + // Legacy DB without table — will be created on next open; skip caching this run + } +} + +function deleteFileSyncStateForCollection(db: Database, collectionName: string, relPath: string): void { + try { + db.prepare(`DELETE FROM file_sync_state WHERE collection = ? AND relative_path = ?`).run(collectionName, relPath); + } catch {} +} + +/** + * A file modified within this long of the stat that read it gets its sync row + * stored with UNTRUSTED_SYNC_MTIME, which the fast path never matches. On a + * filesystem whose mtime granule is 1-2 s (HFS+, FAT, some network mounts), a + * same-size rewrite in that window would keep both mtime and size, and the + * fast path would skip it for good. The next update reads such a file once + * more and trusts it once its mtime is older: git's "racily clean" rule. + */ +const RACY_SYNC_WINDOW_MS = 2000; +const UNTRUSTED_SYNC_MTIME = -1; + +/** + * Maximum file size to index — prevents OOM on accidental binary inclusion. + */ +export const REINDEX_MAX_FILE_SIZE = 10 * 1024 * 1024; // 10MB + +/** + * Deactivate a previously indexed file that can no longer be indexed (it + * became empty or grew past REINDEX_MAX_FILE_SIZE) and drop its sync row, so + * its old content stops being searchable. + */ +function retireIndexedFile( + db: Database, + collectionName: string, + path: string, + livePaths: ReadonlySet, + syncStateMap: Map, +): void { + if (!findOrMigrateLegacyDocument(db, collectionName, path, livePaths)) return; + deactivateDocument(db, collectionName, path); + deleteFileSyncStateForCollection(db, collectionName, path); + syncStateMap.delete(path); +} + +/** + * Groups the per-file writes of a collection scan into short transactions. + * Committed one by one, a first index of a large collection writes a content + * row, a document, its metadata and a sync row per file, each its own WAL + * commit: 20,000 files took about 176 s that way. A batch closes after + * `maxFiles` files or `maxMs`, so other writers never wait longer than that. + */ +export function scanWriteBatch(db: Database, maxFiles: number = 500, maxMs: number = 250) { + let open = false; + let files = 0; + let startedAt = 0; + const commit = (): void => { + if (!open) return; + open = false; + db.exec("COMMIT"); + }; + return { + /** Call before a file's writes; closes the batch when it is full or old. */ + next(): void { + if (open && (files >= maxFiles || Date.now() - startedAt >= maxMs)) commit(); + if (!open) { + db.exec("BEGIN"); + open = true; + files = 0; + startedAt = Date.now(); + } + files++; + }, + commit, + rollback(): void { + if (!open) return; + open = false; + db.exec("ROLLBACK"); + }, + }; +} + /** * Re-index a single collection by scanning the filesystem and updating the database. + * Uses mtime+size fast-path (file_sync_state) to avoid re-reading unchanged files. * Pure function — no console output, no db lifecycle management. + * + * Fast-path: stat mtime_ms+size against cached row to skip file read. + * If mtime changed but content hash identical, only mtime cache is updated. + * Skips >10MB and empty files, cleans sync table entry on orphan removal. */ export async function reindexCollection( store: Store, @@ -1623,6 +1850,28 @@ export async function reindexCollection( ignorePatterns?: string[] | undefined; onProgress?: ((info: ReindexProgress) => void) | undefined; } +): Promise { + const batch = scanWriteBatch(store.db); + try { + const result = await reindexCollectionIn(batch, store, collectionPath, globPattern, collectionName, options); + batch.commit(); + return result; + } catch (err) { + batch.rollback(); + throw err; + } +} + +async function reindexCollectionIn( + batch: ReturnType, + store: Store, + collectionPath: string, + globPattern: string, + collectionName: string, + options?: { + ignorePatterns?: string[] | undefined; + onProgress?: ((info: ReindexProgress) => void) | undefined; + } ): Promise { const db = store.db; const now = new Date().toISOString(); @@ -1632,6 +1881,26 @@ export async function reindexCollection( ...excludeDirs.map(d => `**/${d}/**`), ...(options?.ignorePatterns || []), ]; + + // A missing root (unmounted drive, offline share, deleted folder) globs to + // an empty list, which the deactivation pass below would read as "every + // file was deleted". Leave the index alone and report the root instead. + let rootIsDirectory = false; + try { + rootIsDirectory = statSync(collectionPath).isDirectory(); + } catch (error) { + const code = fsErrorCode(error); + if (code !== "ENOENT" && code !== "ENOTDIR") throw error; + } + if (!rootIsDirectory) { + return { + indexed: 0, updated: 0, unchanged: 0, removed: 0, orphanedCleaned: 0, + skipped: 1, + skippedFiles: [{ file: collectionPath, code: "ROOT_MISSING" }], + metadataErrors: 0, + }; + } + const allFiles: string[] = await fastGlob(splitGlobMask(globPattern), { cwd: collectionPath, onlyFiles: true, @@ -1649,19 +1918,15 @@ export async function reindexCollection( let indexed = 0, updated = 0, unchanged = 0, processed = 0, metadataErrors = 0; const skippedFiles: ReindexSkippedFile[] = []; const seenPaths = new Set(); - // Literal paths of every file in this scan. Passed to the legacy-path - // migration so it never adopts a row that still belongs to a live file. const livePaths = new Set(files.map(f => normalizePathSeparators(f))); + // Load file_sync_state for this collection (mtime+size fast-path) + const syncStateMap = getFileSyncStateMap(db, collectionName); + for (const relativeFile of files) { - const filepath = getRealPath(resolve(collectionPath, relativeFile)); - // Store the literal relative path so the filesystem path can always be - // reconstructed as: resolve(collection.path, storedPath). - // handelize() is NOT applied at index time — it is display-only. + batch.next(); const path = normalizePathSeparators(relativeFile); - // Glob `../` segments, absolute patterns, and file symlinks can resolve - // outside the collection root. Do not ingest those files, and do not mark - // them seen so a previous escaped row is deactivated on this pass. + const filepath = getRealPath(resolve(collectionPath, relativeFile)); if (!isPathInsideDir(collectionPath, filepath)) { processed++; skippedFiles.push({ file: relativeFile, code: "OUTSIDE_COLLECTION" }); @@ -1670,13 +1935,57 @@ export async function reindexCollection( } seenPaths.add(path); + // Stat first — mtime+size fast-path (no read) + let stat: ReturnType | null = null; + const statTimeMs = Date.now(); + try { + stat = statSync(filepath); + } catch (err) { + processed++; + skippedFiles.push({ file: relativeFile, code: fsErrorCode(err) }); + options?.onProgress?.({ file: relativeFile, current: processed, total }); + continue; + } + + if (!stat) { + processed++; + options?.onProgress?.({ file: relativeFile, current: processed, total }); + continue; + } + + const mtimeMs = stat.mtimeMs; + const size = stat.size; + const syncMtimeMs = statTimeMs - mtimeMs < RACY_SYNC_WINDOW_MS ? UNTRUSTED_SYNC_MTIME : mtimeMs; + + // Skip large files (>10MB) — prevents OOM + if (size > REINDEX_MAX_FILE_SIZE) { + retireIndexedFile(db, collectionName, path, livePaths, syncStateMap); + processed++; + skippedFiles.push({ file: relativeFile, code: "FILE_TOO_LARGE" }); + options?.onProgress?.({ file: relativeFile, current: processed, total }); + continue; + } + + // Fast-path: stat matches cached sync state — skip read entirely + const cached = syncStateMap.get(path); + // Missing or stale metadata (an index from before the metadata schema, or + // an extraction-version bump) needs the content, so such a file is read and + // re-extracted through the hash-match branch below. + if (cached && cached.mtime_ms === Math.floor(mtimeMs) && cached.size === size && cached.metadata_current) { + unchanged++; + processed++; + options?.onProgress?.({ file: relativeFile, current: processed, total }); + // Still need to ensure document exists (might have been deactivated externally) + // But we count as unchanged and avoid expensive read+hash+metadata sync. + // Note: metadata sync for unchanged is skipped in fast-path; if needed, disable fast-path or force re-read. + continue; + } + + // Need to read file let content: string; try { content = readFileSync(filepath, "utf-8"); } catch (err) { - // Skip files that can't be read (ETIMEDOUT on APFS compressed files, - // EAGAIN on iCloud evicted files, EACCES, etc.) instead of aborting - // the rest of the collection (#460). processed++; skippedFiles.push({ file: relativeFile, code: fsErrorCode(err) }); options?.onProgress?.({ file: relativeFile, current: processed, total }); @@ -1684,13 +1993,34 @@ export async function reindexCollection( } if (!content.trim()) { + // Empty file — if previously indexed, deactivate it (treat as removed) + retireIndexedFile(db, collectionName, path, livePaths, syncStateMap); processed++; + options?.onProgress?.({ file: relativeFile, current: processed, total }); continue; } const hash = await hashContent(content); - const title = extractTitle(content, relativeFile); + // Hash matches cached sync state but mtime differed (clock skew, backup restore) — only update mtime cache + if (cached && cached.content_hash === hash) { + // Update sync state mtime/size only + upsertFileSyncState(db, collectionName, path, syncMtimeMs, size, hash, cached.document_id); + unchanged++; + processed++; + // Keep content in memory for metadata sync if needed? For speed, skip metadata sync on hash-match fast-path. + // Existing behavior for hash-same was to still do metadata backfill; we preserve it by loading documentId from cache. + // However we already have content here, so do metadata backfill for hash-match case. + const existingForMeta = findOrMigrateLegacyDocument(db, collectionName, path, livePaths); + if (existingForMeta) { + const extraction = syncDocumentMetadata(db, existingForMeta.id, content, path, { onlyIfStale: true }); + if (extraction?.error) metadataErrors++; + } + options?.onProgress?.({ file: relativeFile, current: processed, total }); + continue; + } + + const title = extractTitle(content, relativeFile); const existing = findOrMigrateLegacyDocument(db, collectionName, path, livePaths); let documentId: number; @@ -1708,21 +2038,21 @@ export async function reindexCollection( } } else { insertContent(db, hash, content, now); - const stat = statSync(filepath); - updateDocument(db, existing.id, title, hash, - stat ? new Date(stat.mtime).toISOString() : now); + updateDocument(db, existing.id, title, hash, new Date(stat.mtime).toISOString()); updated++; } } else { indexed++; insertContent(db, hash, content, now); - const stat = statSync(filepath); documentId = insertDocument(db, collectionName, path, title, hash, stat ? new Date(stat.birthtime).toISOString() : now, - stat ? new Date(stat.mtime).toISOString() : now); + new Date(stat.mtime).toISOString()); } - // Unchanged content still backfills missing or stale extraction state. + // Upsert sync state after successful indexing + upsertFileSyncState(db, collectionName, path, syncMtimeMs, size, hash, documentId); + + // Metadata extraction const extraction = syncDocumentMetadata(db, documentId, content, path, contentChanged ? undefined : { onlyIfStale: true }); if (extraction?.error) metadataErrors++; @@ -1731,16 +2061,20 @@ export async function reindexCollection( options?.onProgress?.({ file: relativeFile, current: processed, total }); } - // Deactivate documents that no longer exist + // Deactivate documents that no longer exist + cleanup sync_state const allActive = getActiveDocumentPaths(db, collectionName); let removed = 0; for (const path of allActive) { if (!seenPaths.has(path)) { + batch.next(); deactivateDocument(db, collectionName, path); + deleteFileSyncStateForCollection(db, collectionName, path); removed++; } } + batch.commit(); + const orphanedCleaned = cleanupOrphanedContent(db); return { indexed, updated, unchanged, removed, orphanedCleaned, skipped: skippedFiles.length, skippedFiles, metadataErrors }; @@ -1767,6 +2101,8 @@ export type EmbedProgress = { export type EmbedResult = { docsProcessed: number; chunksEmbedded: number; + /** Chunks copied into the partition of a collection that gained an already-embedded hash. */ + chunksCopied: number; /** Active failed chunks that did not recover after retries. */ errors: number; failures?: EmbedFailure[]; @@ -1924,7 +2260,15 @@ function getPendingEmbeddingDocs(db: Database, collection?: string, model: strin GROUP BY d.hash ORDER BY MIN(d.path) `); - return (collection ? stmt.all(model, fingerprint, collection) : stmt.all(model, fingerprint)) as PendingEmbeddingDoc[]; + // Large-result query (up to 9k docs): stream via iterate() instead of .all() to bound V8 heap + const results: PendingEmbeddingDoc[] = []; + const iter = collection + ? stmt.iterate(model, fingerprint, collection) + : stmt.iterate(model, fingerprint); + for (const row of iter as IterableIterator) { + results.push(row); + } + return results; }); } @@ -1963,12 +2307,17 @@ function getEmbeddingDocsForBatch(db: Database, batch: PendingEmbeddingDoc[]): E if (batch.length === 0) return []; const placeholders = batch.map(() => "?").join(","); - const rows = db.prepare(` + // Bounded: batch size max 64 (maxDocsPerBatch), so IN list max 64 hashes. + // Use iterate() to stream rows instead of materializing all at once, nicer for large batches. + const stmt = db.prepare(` SELECT hash, doc as body FROM content WHERE hash IN (${placeholders}) - `).all(...batch.map(doc => doc.hash)) as { hash: string; body: string }[]; - const bodyByHash = new Map(rows.map(row => [row.hash, row.body])); + `); + const bodyByHash = new Map(); + for (const row of stmt.iterate(...batch.map(doc => doc.hash)) as IterableIterator<{ hash: string; body: string }>) { + bodyByHash.set(row.hash, row.body); + } return batch.map((doc) => ({ ...doc, @@ -1997,10 +2346,11 @@ export async function generateEmbeddings( clearAllEmbeddings(db, options?.collection); } + const chunksCopied = copyVectorsToNewCollections(db, options?.collection).copied; const docsToEmbed = getPendingEmbeddingDocs(db, options?.collection, model); if (docsToEmbed.length === 0) { - return { docsProcessed: 0, chunksEmbedded: 0, errors: 0, durationMs: 0 }; + return { docsProcessed: 0, chunksEmbedded: 0, chunksCopied, errors: 0, durationMs: 0 }; } const totalBytes = docsToEmbed.reduce((sum, doc) => sum + Math.max(0, doc.bytes), 0); const totalDocs = docsToEmbed.length; @@ -2176,19 +2526,30 @@ export async function generateEmbeddings( try { const embeddings = await session.embedBatch(texts, { model }); - for (let i = 0; i < chunkBatch.length; i++) { - const chunk = chunkBatch[i]!; - const embedding = embeddings[i]; - if (embedding) { - insertEmbedding(db, chunk.hash, chunk.seq, chunk.pos, new Float32Array(embedding.embedding), model, now, chunk.expectedTotalChunks, fingerprint); - chunksEmbedded++; - successesSinceRetry++; - clearFailure(chunk); - } else { - recordFailure(chunk, "batch embedding returned no vector"); + // One IMMEDIATE transaction per batch: every chunk writes several + // rows across three tables, and a kill mid-batch must not leave + // them out of step. Failure bookkeeping runs after the commit. + const stored: ChunkItem[] = []; + const unembedded: ChunkItem[] = []; + db.transaction(() => { + for (let i = 0; i < chunkBatch.length; i++) { + const chunk = chunkBatch[i]!; + const embedding = embeddings[i]; + if (embedding) { + insertEmbedding(db, chunk.hash, chunk.seq, chunk.pos, new Float32Array(embedding.embedding), model, now, chunk.expectedTotalChunks, fingerprint); + stored.push(chunk); + } else { + unembedded.push(chunk); + } } - batchChunkBytesProcessed += chunk.bytes; + }).immediate(); + for (const chunk of stored) { + chunksEmbedded++; + successesSinceRetry++; + clearFailure(chunk); } + for (const chunk of unembedded) recordFailure(chunk, "batch embedding returned no vector"); + batchChunkBytesProcessed += chunkBatch.reduce((sum, chunk) => sum + chunk.bytes, 0); await retryFailedChunks(); } catch (error) { // Batch failed — try individual embeddings as fallback. If an @@ -2237,6 +2598,7 @@ export async function generateEmbeddings( return { docsProcessed: totalDocs, chunksEmbedded: result.chunksEmbedded, + chunksCopied, errors: result.errors, failures: result.failures, durationMs: Date.now() - startTime, @@ -2265,6 +2627,7 @@ export function createStore(dbPath?: string): Store { getHashesNeedingEmbedding: (model?: string) => getHashesNeedingEmbedding(db, undefined, model ?? store.llm?.embedModelName ?? DEFAULT_EMBED_MODEL), getIndexHealth: (model?: string) => getIndexHealth(db, model ?? store.llm?.embedModelName ?? DEFAULT_EMBED_MODEL), getStatus: (model?: string) => getStatus(db, model ?? store.llm?.embedModelName ?? DEFAULT_EMBED_MODEL), + getStatusSummary: (model?: string) => getStatusSummary(db, model ?? store.llm?.embedModelName ?? DEFAULT_EMBED_MODEL), // Caching getCacheKey, @@ -2528,8 +2891,19 @@ export type CollectionInfo = { pattern: string | null; documents: number; lastUpdated: string; + /** Distinct metadata keys declared in this collection. */ + metadataKeyCount: number; + /** + * The most covered metadata keys in this collection, with coverage and + * types, by coverage. Windowed to STATUS_METADATA_KEY_LIMIT; `listMetadata` + * pages through the rest. + */ + metadataKeys: MetadataKeyOverview[]; }; +/** Keys a status view names per collection before pointing at discovery. */ +export const STATUS_METADATA_KEY_LIMIT = 10; + export type IndexStatus = { totalDocuments: number; needsEmbedding: number; @@ -2539,10 +2913,46 @@ export type IndexStatus = { collections: CollectionInfo[]; }; +/** Index facts for initialization, without computing unused metadata overviews. */ +export type IndexStatusSummary = Omit & { + collections: Omit[]; +}; + // ============================================================================= // Index health // ============================================================================= +export type EmbeddingVectorSample = { + hash: string; + seq: number; + pos: number; + body: string; + path: string; +}; + +export function getEmbeddingVectorSamples(db: Database, model: string, fingerprint: string, sampleSize: number = 3): EmbeddingVectorSample[] { + // Limit chunk identities before reading bodies. Joining bodies to every chunk + // and duplicate path can make a three-row sample sort gigabytes of text. + return db.prepare(` + WITH sampled AS MATERIALIZED ( + SELECT cv.hash, cv.seq, cv.pos + FROM content_vectors cv + JOIN content c ON c.hash = cv.hash + WHERE cv.model = ? AND cv.embed_fingerprint = ? + AND EXISTS ( + SELECT 1 FROM documents d WHERE d.hash = cv.hash AND d.active = 1 + ) + ORDER BY random() + LIMIT ? + ) + SELECT sampled.hash, sampled.seq, sampled.pos, c.doc AS body, + (SELECT MIN(d.path) FROM documents d + WHERE d.hash = sampled.hash AND d.active = 1) AS path + FROM sampled + JOIN content c ON c.hash = sampled.hash + `).all(model, fingerprint, sampleSize); +} + export function getHashesNeedingEmbedding(db: Database, collection?: string, model: string = DEFAULT_EMBED_MODEL): number { const collectionFilter = collection ? `AND d.collection = ?` : ``; const fingerprint = getEmbeddingFingerprint(model); @@ -2588,24 +2998,46 @@ export async function maybeAdoptLegacyEmbeddingFingerprint(store: Store, model: return { checked: false, adopted: 0, reason: "no legacy empty-fingerprint embeddings" }; } - const sample = withLazyContentVectorMigration(db, () => db.prepare(` - SELECT cv.hash, cv.seq, cv.pos, cv.total_chunks, c.doc AS body, MIN(d.path) AS path - FROM content_vectors cv - JOIN documents d ON d.hash = cv.hash AND d.active = 1 - JOIN content c ON c.hash = cv.hash - WHERE cv.model = ? AND cv.embed_fingerprint = '' - GROUP BY cv.hash, cv.seq, cv.pos, cv.total_chunks, c.doc - ORDER BY cv.hash, cv.seq - LIMIT 1 - `).get(model) as { hash: string; seq: number; pos: number; total_chunks: number; body: string; path: string } | undefined); + // Pick the sample through indexes, then read one body by rowid; a join that + // carries c.doc costs one body copy per legacy chunk × active path. Both + // EXISTS filters depend only on hash, so this is still the lowest (hash, seq). + // One read transaction keeps the pick and the load on the same snapshot. + const sample = withLazyContentVectorMigration(db, () => { + const pickSample = db.prepare(` + SELECT cv.rowid AS rid + FROM content_vectors cv + WHERE cv.model = ? AND cv.embed_fingerprint = '' + AND cv.hash = ( + SELECT l.hash + FROM content_vectors l + WHERE l.model = ? AND l.embed_fingerprint = '' + AND EXISTS (SELECT 1 FROM documents d WHERE d.hash = l.hash AND d.active = 1) + AND EXISTS (SELECT 1 FROM content c WHERE c.hash = l.hash) + ORDER BY l.hash + LIMIT 1 + ) + ORDER BY cv.seq + LIMIT 1 + `); + const loadSample = db.prepare(` + SELECT cv.hash, cv.seq, cv.pos, cv.total_chunks, c.doc AS body, + (SELECT MIN(d.path) FROM documents d WHERE d.hash = cv.hash AND d.active = 1) AS path + FROM content_vectors cv + JOIN content c ON c.hash = cv.hash + WHERE cv.rowid = ? + `); + return db.transaction(() => { + const pick = pickSample.get(model, model) as { rid: number } | null | undefined; + return pick ? loadSample.get(pick.rid) : undefined; + })() as { hash: string; seq: number; pos: number; total_chunks: number; body: string; path: string } | null | undefined; + }); if (!sample) { return { checked: false, adopted: 0, reason: `${legacyCount} legacy docs have no active sample` }; } - const tableExists = db.prepare(`SELECT name FROM sqlite_master WHERE type='table' AND name='vectors_vec'`).get(); - if (!tableExists) { - return { checked: false, adopted: 0, reason: "vectors_vec table is missing" }; + if (!hasVectorIndex(db)) { + return { checked: false, adopted: 0, reason: "vector index is missing" }; } const expectedHashSeq = `${sample.hash}_${sample.seq}`; @@ -2634,18 +3066,20 @@ export async function maybeAdoptLegacyEmbeddingFingerprint(store: Store, model: } const nearest = db.prepare(` - SELECT hash_seq, distance - FROM vectors_vec + SELECT rowid, distance + FROM ${VEC_TABLE} WHERE embedding MATCH ? AND k = 1 - `).get(new Float32Array(result.embedding)) as { hash_seq: string; distance: number } | undefined; + `).get(new Float32Array(result.embedding)) as { rowid: number; distance: number } | undefined; + const nearestKey = nearest ? partitionRowKey(db, nearest.rowid) : undefined; - if (!nearest) { + if (!nearest || !nearestKey) { return { checked: true, adopted: 0, reason: "legacy sample vector not found" }; } const threshold = 0.0001; - if (nearest.hash_seq !== expectedHashSeq || nearest.distance > threshold) { - return { checked: true, adopted: 0, reason: `legacy sample differs from current fingerprint (nearest ${nearest.hash_seq}, distance ${nearest.distance.toFixed(6)})` }; + const nearestHashSeq = `${nearestKey.hash}_${nearestKey.seq}`; + if (nearestHashSeq !== expectedHashSeq || nearest.distance > threshold) { + return { checked: true, adopted: 0, reason: `legacy sample differs from current fingerprint (nearest ${nearestHashSeq}, distance ${nearest.distance.toFixed(6)})` }; } const update = withLazyContentVectorMigration(db, () => db.prepare(`UPDATE content_vectors SET embed_fingerprint = ? WHERE model = ? AND embed_fingerprint = ''`).run(fingerprint, model)); @@ -2775,74 +3209,205 @@ export function countOrphanedVectors(db: Database): number { } /** - * Remove orphaned vector embeddings that are not referenced by any active document. - * Returns the number of orphaned embedding chunks deleted. + * Remove vector rows whose (hash, collection) no active document references, + * and the content_vectors rows of hashes no active document references at + * all. Returns the number of orphaned chunks deleted. */ export function cleanupOrphanedVectors(db: Database): number { // sqlite-vec may not be loaded (e.g. Bun's bun:sqlite lacks loadExtension). - // The vectors_vec virtual table can appear in sqlite_master from a prior - // session, but querying it without the vec0 module loaded will crash (#380). + // The vec0 table can appear in sqlite_master from a prior session, but + // querying it without the module loaded would crash (#380). if (!isSqliteVecAvailable()) { return 0; } - - // The schema entry can exist even when sqlite-vec itself is unavailable - // (for example when reopening a DB without vec0 loaded). In that case, - // touching the virtual table throws "no such module: vec0" and cleanup - // should degrade gracefully like the rest of the vector features. - try { - db.prepare(`SELECT 1 FROM vectors_vec LIMIT 0`).get(); - } catch { + const layout = vecLayout(db); + if (layout.kind !== "partitioned" || !vecTableReadable(db, layout)) { return 0; } return withLazyContentVectorMigration(db, () => { - // Count and both DELETEs share one transaction. An interruption between the - // two DELETEs (crash, SQLITE_BUSY) desyncs the tables: vectors_vec loses - // the rows while content_vectors still records the chunks as embedded. - // These rows are orphaned (no active document), so live vector search — - // which post-filters on documents.active = 1 — is unaffected right away. - // The failure is latent: if that content hash is later reactivated (qmd is - // content-addressable, so the same content returning revives the hash), the - // stale content_vectors rows make getHashesNeedingEmbedding treat it as - // already embedded, so qmd embed skips it and the document is silently - // unsearchable by vector with no orphan left to clean up. Keeping the count - // inside the same transaction also makes the returned number match the rows - // the DELETEs actually remove if another connection mutates documents - // concurrently. Run it BEGIN IMMEDIATE: the count reads before the DELETEs - // write, and upgrading a deferred read snapshot under a concurrent WAL - // writer fails with SQLITE_BUSY_SNAPSHOT instead of honoring the busy + // The count and both DELETEs share one transaction: an interruption + // between them (crash, SQLITE_BUSY) would leave content_vectors claiming + // chunks the index no longer holds, and a revived hash would then be + // skipped by getHashesNeedingEmbedding and stay unsearchable. Run it + // BEGIN IMMEDIATE: upgrading a deferred read snapshot under a concurrent + // WAL writer fails with SQLITE_BUSY_SNAPSHOT instead of honoring the busy // timeout. Nested callers still get a savepoint. const cleanup = db.transaction(() => { - const orphaned = (db.prepare(ORPHANED_VECTOR_COUNT_SQL).get() as { c: number }).c; - if (orphaned === 0) { + const orphanRows = db.prepare(` + SELECT vr.id FROM ${VEC_ROWS_TABLE} vr + JOIN ${VEC_COLLECTION_IDS_TABLE} ci ON ci.id = vr.collection_id + WHERE NOT EXISTS ( + SELECT 1 FROM documents d WHERE d.hash = vr.hash AND d.collection = ci.name AND d.active = 1 + ) + `).all() as { id: number }[]; + const orphanedChunks = (db.prepare(ORPHANED_VECTOR_COUNT_SQL).get() as { c: number }).c; + if (orphanRows.length === 0 && orphanedChunks === 0) { return 0; } - // Delete from vectors_vec first - db.exec(` - DELETE FROM vectors_vec WHERE hash_seq IN ( - SELECT cv.hash || '_' || cv.seq FROM content_vectors cv - WHERE NOT EXISTS ( - SELECT 1 FROM documents d WHERE d.hash = cv.hash AND d.active = 1 - ) - ) - `); - - // Delete from content_vectors + deletePartitionRows(db, orphanRows.map((row) => row.id)); db.exec(` DELETE FROM content_vectors WHERE hash NOT IN ( SELECT hash FROM documents WHERE active = 1 ) `); - return orphaned; + return Math.max(orphanRows.length, orphanedChunks); }); return cleanup.immediate(); }); } +/** + * Repack the vector table when fewer than this share of its chunk reads hold + * live rows. vec0 places an insert in the first free slot of its newest chunk + * and reclaims a chunk only once it is empty, so a delete in any older chunk + * leaves a hole that every brute-force scan still reads; at 36% occupancy a + * scan took twice as long as on a packed copy of the same rows. + */ +const VEC_REPACK_BELOW_OCCUPANCY = 0.9; + +/** Chunks filled below this share have their live rows moved to the tail. */ +const VEC_REPACK_CHUNK_FILL = 0.9; + +export type VectorTableLayout = { + rows: number; + chunks: number; + /** Chunks a freshly packed table needs for the same rows. */ + neededChunks: number; + /** neededChunks over chunks: the share of a scan's chunk reads that hold live rows. */ + occupancy: number; +}; + +interface VecChunkRow { + chunkId: number; + size: number; + partition: number; + validity: Uint8Array; + rowids: Uint8Array; +} + +/** A plan for repackVectors: the table's chunk layout and the chunks it would move. */ +type VecRepackPlan = { layout: VectorTableLayout; sparse: number[] }; + +/** The partitioned vec0 table when sqlite-vec is loaded and the table answers; null otherwise. */ +function readableVectorLayout(db: Database): Extract | null { + if (!isSqliteVecAvailable()) return null; + const layout = vecLayout(db); + return layout.kind === "partitioned" && vecTableReadable(db, layout) ? layout : null; +} + +/** + * Chunk layout of the vector table from vec0's shadow tables, and the chunks + * a repack would move: those filled below VEC_REPACK_CHUNK_FILL, except each + * partition's newest. Rows of different partitions never share a chunk, so a + * packed table needs ceil(rows / chunk size) chunks in each partition. With + * `dropOrphans` the plan is for the table as cleanupOrphanedVectors will + * leave it; vec0 drops a chunk with its last row, so a chunk holding only + * orphans is gone from that layout. + */ +function vectorRepackPlan(db: Database, dropOrphans: boolean = false): VecRepackPlan | null { + const layout = readableVectorLayout(db); + if (!layout) return null; + const orphansByChunk = new Map(); + if (dropOrphans) { + const rows = db.prepare(` + SELECT r.chunk_id AS chunkId, COUNT(*) AS n + FROM ${VEC_ROWS_TABLE} vr + JOIN ${VEC_COLLECTION_IDS_TABLE} ci ON ci.id = vr.collection_id + JOIN ${layout.shadow.rowids} r ON r.rowid = vr.id + WHERE NOT EXISTS (SELECT 1 FROM documents d WHERE d.hash = vr.hash AND d.collection = ci.name AND d.active = 1) + GROUP BY r.chunk_id + `).all() as { chunkId: number; n: number }[]; + for (const row of rows) orphansByChunk.set(row.chunkId, row.n); + } + const chunks = (db.prepare(`SELECT chunk_id AS chunkId, size, partition00 AS partition, validity FROM ${layout.shadow.chunks} ORDER BY chunk_id`).all() as Omit[]) + .map((chunk) => ({ ...chunk, live: liveSlots(chunk).length - (orphansByChunk.get(chunk.chunkId) ?? 0) })) + .filter((chunk) => chunk.live > 0); + const rowsByPartition = new Map(); + const newestByPartition = new Map(); + for (const chunk of chunks) { + rowsByPartition.set(chunk.partition, (rowsByPartition.get(chunk.partition) ?? 0) + chunk.live); + newestByPartition.set(chunk.partition, chunk.chunkId); + } + let rows = 0; + let neededChunks = 0; + for (const chunk of chunks) { + if (newestByPartition.get(chunk.partition) !== chunk.chunkId) continue; + const partitionRows = rowsByPartition.get(chunk.partition)!; + rows += partitionRows; + neededChunks += Math.ceil(partitionRows / chunk.size); + } + const occupancy = chunks.length > 0 ? neededChunks / chunks.length : 1; + const sparse = chunks + .filter((chunk) => newestByPartition.get(chunk.partition) !== chunk.chunkId && chunk.live < chunk.size * VEC_REPACK_CHUNK_FILL) + .map((chunk) => chunk.chunkId); + return { layout: { rows, chunks: chunks.length, neededChunks, occupancy }, sparse }; +} + +/** Whether a plan is worth running: the table's scans read mostly holes and some chunk would move. */ +function repackWanted(plan: VecRepackPlan | null): plan is VecRepackPlan { + return plan !== null && plan.layout.occupancy < VEC_REPACK_BELOW_OCCUPANCY && plan.sparse.length > 0; +} + +/** Chunk layout of the vector table from vec0's shadow tables, or null without a readable table. */ +export function vectorTableLayout(db: Database): VectorTableLayout | null { + return vectorRepackPlan(db)?.layout ?? null; +} + +function liveSlots(chunk: { size: number; validity: Uint8Array }): number[] { + const slots: number[] = []; + for (let i = 0; i < chunk.size; i++) { + if ((chunk.validity[i >> 3]! >> (i & 7)) & 1) slots.push(i); + } + return slots; +} + +/** + * Compact the vector table in place: the live rows of every chunk that is + * mostly holes are deleted and re-inserted under their rowid and partition, + * one chunk per IMMEDIATE transaction. vec0 puts each insert in the first + * free slot of its partition's newest chunk and drops a chunk once it is + * empty, so the moved rows fill that partition's tail and the emptied chunks + * vanish; each partition's newest chunk is left alone because that is where + * its rows land. A transaction holds the write lock for at most one chunk of + * rows, so concurrent embed and update runs interleave with it, and an + * interrupted run leaves a consistent table that the next cleanup finishes. + */ +export function repackVectors(db: Database, onChunk?: (moved: number, total: number) => void): VectorTableLayout | null { + const layout = readableVectorLayout(db); + const plan = vectorRepackPlan(db); + if (!layout || !plan) return null; + const chunkOf = db.prepare(`SELECT chunk_id AS chunkId, size, partition00 AS partition, validity, rowids FROM ${layout.shadow.chunks} WHERE chunk_id = ?`); + const newestOf = db.prepare(`SELECT MAX(chunk_id) AS chunkId FROM ${layout.shadow.chunks} WHERE partition00 = ?`); + const vectorOf = db.prepare(`SELECT embedding FROM ${layout.table} WHERE rowid = ?`); + const remove = db.prepare(`DELETE FROM ${layout.table} WHERE rowid = ?`); + const insert = db.prepare(`INSERT INTO ${layout.table} (rowid, collection_id, embedding) VALUES (?, ?, ?)`); + plan.sparse.forEach((chunkId, i) => { + db.transaction(() => { + // Read the chunk again under the write lock: since the plan was made, a + // concurrent embed or update may have emptied it, made it its + // partition's newest, or deleted a row whose rowid now belongs to + // another partition. + const chunk = chunkOf.get(chunkId) as VecChunkRow | undefined; + if (!chunk) return; + if ((newestOf.get(chunk.partition) as { chunkId: number }).chunkId === chunkId) return; + const rowids = new DataView(chunk.rowids.buffer, chunk.rowids.byteOffset, chunk.rowids.byteLength); + for (const slot of liveSlots(chunk)) { + const rowid = rowids.getBigInt64(slot * 8, true); + const row = vectorOf.get(rowid) as { embedding: Uint8Array } | undefined; + if (row === undefined) continue; + remove.run(rowid); + insert.run(rowid, vecInteger(chunk.partition), row.embedding); + } + }).immediate(); + onChunk?.(i + 1, plan.sparse.length); + }); + return vectorTableLayout(db); +} + /** * Run VACUUM to reclaim unused space in the database. * This operation rebuilds the database file to eliminate fragmentation. @@ -2870,6 +3435,10 @@ export type CleanupStats = { orphanedVectors: number; inactiveDocs: number; orphanedContent: number; + /** Chunk layout of the vector table before any repack, null without a vector table. */ + vectorLayout: VectorTableLayout | null; + /** True when the vector table was repacked (or, in a preview, would be). */ + vectorsRepacked: boolean; }; /** Counts what `runCleanup` would remove, including content only held by inactive docs. */ @@ -2878,21 +3447,34 @@ export function previewCleanup(db: Database): CleanupStats { const orphanedVectors = countOrphanedVectors(db); const inactiveDocs = (db.prepare(`SELECT COUNT(*) as c FROM documents WHERE active = 0`).get() as { c: number }).c; const orphanedContent = countOrphanedContent(db); - return { cacheCount, orphanedVectors, inactiveDocs, orphanedContent }; + const plan = vectorRepackPlan(db, true); + return { cacheCount, orphanedVectors, inactiveDocs, orphanedContent, vectorLayout: plan?.layout ?? null, vectorsRepacked: repackWanted(plan) }; } +export type CleanupHooks = { + /** Called with the layout right before a vector repack starts. */ + onVectorRepack?: (layout: VectorTableLayout) => void; +}; + /** - * Full `qmd cleanup` sequence: drop cache, orphaned vectors, inactive document - * rows, then the content those rows were pinning, compact FTS5, vacuum. + * Full `qmd cleanup` sequence: drop cache, orphaned vectors, repack the vector + * table when its chunks are mostly holes, inactive document rows, then the + * content those rows were pinning, compact FTS5, vacuum. */ -export function runCleanup(db: Database): CleanupStats { +export function runCleanup(db: Database, hooks: CleanupHooks = {}): CleanupStats { const cacheCount = deleteLLMCache(db); const orphanedVectors = cleanupOrphanedVectors(db); + const plan = vectorRepackPlan(db); + const vectorsRepacked = repackWanted(plan); + if (vectorsRepacked) { + hooks.onVectorRepack?.(plan.layout); + repackVectors(db); + } const inactiveDocs = deleteInactiveDocuments(db); const orphanedContent = cleanupOrphanedContent(db); optimizeDocumentsFts(db); vacuumDatabase(db); - return { cacheCount, orphanedVectors, inactiveDocs, orphanedContent }; + return { cacheCount, orphanedVectors, inactiveDocs, orphanedContent, vectorLayout: plan?.layout ?? null, vectorsRepacked }; } // ============================================================================= @@ -3137,10 +3719,15 @@ export function deactivateDocument(db: Database, collectionName: string, path: s * Get all active document paths for a collection. */ export function getActiveDocumentPaths(db: Database, collectionName: string): string[] { - const rows = db.prepare(` + const stmt = db.prepare(` SELECT path FROM documents WHERE collection = ? AND active = 1 - `).all(collectionName) as { path: string }[]; - return rows.map(r => r.path); + `); + // Large-result query (up to 5k per collection): use iterate() to bound heap + const paths: string[] = []; + for (const r of stmt.iterate(collectionName) as IterableIterator<{ path: string }>) { + paths.push(r.path); + } + return paths; } export { formatQueryForEmbedding, formatDocForEmbedding }; @@ -3450,22 +4037,26 @@ export function findDocumentByDocid(db: Database, docid: string): { filepath: st } export function findSimilarFiles(db: Database, query: string, maxDistance: number = 3, limit: number = 5): string[] { - const allFiles = db.prepare(` + const stmt = db.prepare(` SELECT d.path FROM documents d WHERE d.active = 1 - `).all() as { path: string }[]; + `); + // Large-result query (all active docs, up to 9k): iterate to bound heap const queryLower = query.toLowerCase(); - const scored = allFiles - .map(f => ({ path: f.path, dist: levenshtein(f.path.toLowerCase(), queryLower) })) - .filter(f => f.dist <= maxDistance) + const scored: { path: string; dist: number }[] = []; + for (const f of stmt.iterate() as IterableIterator<{ path: string }>) { + const dist = levenshtein(f.path.toLowerCase(), queryLower); + if (dist <= maxDistance) scored.push({ path: f.path, dist }); + } + return scored .sort((a, b) => a.dist - b.dist) - .slice(0, limit); - return scored.map(f => f.path); + .slice(0, limit) + .map(f => f.path); } export function matchFilesByGlob(db: Database, pattern: string): { filepath: string; displayPath: string; bodyLength: number }[] { - const allFiles = db.prepare(` + const stmt = db.prepare(` SELECT 'qmd://' || d.collection || '/' || d.path as virtual_path, LENGTH(content.doc) as body_length, @@ -3474,16 +4065,20 @@ export function matchFilesByGlob(db: Database, pattern: string): { filepath: str FROM documents d JOIN content ON content.hash = d.hash WHERE d.active = 1 - `).all() as { virtual_path: string; body_length: number; path: string; collection: string }[]; - + `); + // Large-result query: iterate to bound heap (all active docs) const isMatch = picomatch(pattern); - return allFiles - .filter(f => isMatch(f.virtual_path) || isMatch(f.path) || isMatch(f.collection + '/' + f.path)) - .map(f => ({ - filepath: f.virtual_path, // Virtual path for precise lookup - displayPath: f.path, // Relative path for display - bodyLength: f.body_length - })); + const results: { filepath: string; displayPath: string; bodyLength: number }[] = []; + for (const f of stmt.iterate() as IterableIterator<{ virtual_path: string; body_length: number; path: string; collection: string }>) { + if (isMatch(f.virtual_path) || isMatch(f.path) || isMatch(f.collection + '/' + f.path)) { + results.push({ + filepath: f.virtual_path, + displayPath: f.path, + bodyLength: f.body_length, + }); + } + } + return results; } // ============================================================================= @@ -3674,39 +4269,82 @@ export function listCollections(db: Database): { name: string; pwd: string; glob } /** - * Remove a collection and clean up its documents. - * Uses collections.ts to remove from YAML config and cleans up database. + * Drop a collection's vector partition and its id. With the vec0 table + * present but sqlite-vec not loaded, the rows stay for cleanupOrphanedVectors + * to remove once the extension loads, so the mapping never runs ahead of the + * index. + */ +function deleteVectorPartition(db: Database, collectionName: string): void { + const collectionId = resolveCollectionId(db, collectionName); + if (collectionId === undefined) return; + const layout = vecLayout(db); + if (layout.kind === "legacy") return; + if (layout.kind === "partitioned") { + if (!isSqliteVecAvailable()) return; + db.prepare(`DELETE FROM ${VEC_TABLE} WHERE collection_id = ?`).run(vecInteger(collectionId)); + } + db.prepare(`DELETE FROM ${VEC_ROWS_TABLE} WHERE collection_id = ?`).run(collectionId); + deleteCollectionId(db, collectionName); +} + +/** + * Remove a collection: its vector partition, its documents, the content no + * active document references any more, and its store_collections row. */ export function removeCollection(db: Database, collectionName: string): { deletedDocs: number; cleanedHashes: number } { - // Delete documents from database - const docResult = db.prepare(`DELETE FROM documents WHERE collection = ?`).run(collectionName); + // One commit: a partition delete that lands without the documents delete + // leaves the mapping and id row out of step with the vec0 table, and the + // update-time orphan cleanup skips rows whose documents are still active. + return db.transaction(() => { + deleteVectorPartition(db, collectionName); - // Clean up orphaned content hashes - const cleanupResult = db.prepare(` - DELETE FROM content - WHERE hash NOT IN (SELECT DISTINCT hash FROM documents WHERE active = 1) - `).run(); + const docResult = db.prepare(`DELETE FROM documents WHERE collection = ?`).run(collectionName); + db.prepare(`DELETE FROM file_sync_state WHERE collection = ?`).run(collectionName); - // Remove from store_collections - deleteStoreCollection(db, collectionName); + const cleanupResult = db.prepare(` + DELETE FROM content + WHERE hash NOT IN (SELECT DISTINCT hash FROM documents WHERE active = 1) + `).run(); - return { - deletedDocs: docResult.changes, - cleanedHashes: cleanupResult.changes - }; + deleteStoreCollection(db, collectionName); + + return { + deletedDocs: docResult.changes, + cleanedHashes: cleanupResult.changes, + }; + }).immediate(); } /** - * Rename a collection. - * Updates both YAML config and database documents table. + * Rename a collection: its documents, its vector partition id, and its + * store_collections row. */ export function renameCollection(db: Database, oldName: string, newName: string): void { - // Update all documents with the new collection name in database - db.prepare(`UPDATE documents SET collection = ? WHERE collection = ?`) - .run(newName, oldName); + // One commit: the vector id name is UNIQUE, so a rename that fails after + // the documents moved would leave them under a name whose partition id + // still belongs to the old one, and vector search would miss them. + db.transaction(() => { + renameStoreCollection(db, oldName, newName); + + // A removed collection keeps its id while sqlite-vec is not loaded to + // drop its partition; the renamed collection's id cannot take that name. + if (resolveCollectionId(db, newName) !== undefined) { + deleteVectorPartition(db, newName); + if (resolveCollectionId(db, newName) !== undefined) { + throw new Error( + `Cannot rename to '${newName}': a removed collection of that name still has a vector partition, which only qmd with sqlite-vec loaded can drop. ` + + `Open qmd with sqlite-vec loaded and retry the rename.` + ); + } + } - // Rename in store_collections - renameStoreCollection(db, oldName, newName); + db.prepare(`UPDATE documents SET collection = ? WHERE collection = ?`).run(newName, oldName); + // The documents keep their ids and paths, so their sync rows stay valid + // under the new name. Rows already under it belong to no collection. + db.prepare(`DELETE FROM file_sync_state WHERE collection = ?`).run(newName); + db.prepare(`UPDATE file_sync_state SET collection = ? WHERE collection = ?`).run(newName, oldName); + renameCollectionId(db, oldName, newName); + }).immediate(); } // ============================================================================= @@ -4065,27 +4703,12 @@ function scopedCollectionNames(scope: CollectionScope): string[] | undefined { return names.length > 0 ? names : undefined; } -function mergeSearchResultsByScore(lists: SearchResult[][], limit: number): SearchResult[] { - const best = new Map(); - for (const list of lists) { - for (const r of list) { - const prev = best.get(r.filepath); - if (!prev || r.score > prev.score) best.set(r.filepath, r); - } - } - return Array.from(best.values()) - .sort((a, b) => b.score - a.score) - .slice(0, limit); +function compareFilepaths(a: { filepath: string }, b: { filepath: string }): number { + return a.filepath < b.filepath ? -1 : a.filepath > b.filepath ? 1 : 0; } export function searchFTS(db: Database, query: string, limit: number = 20, collectionName?: string | readonly string[], filter?: MetadataFilter): SearchResult[] { const names = scopedCollectionNames(collectionName); - // Search each requested collection before merging/truncating so a large - // unrelated collection cannot occupy global top-k and starve the rest (#775). - if (names && names.length > 1) { - return mergeSearchResultsByScore(names.map(name => searchFTS(db, query, limit, name, filter)), limit); - } - const collectionFilter = names?.[0]; const ftsQuery = buildFTS5Query(query); if (!ftsQuery) return []; @@ -4097,26 +4720,30 @@ export function searchFTS(db: Database, query: string, limit: number = 20, colle // query into a 17-second query on large collections. const params: (string | number)[] = [ftsQuery]; - // When filtering by collection or metadata, fetch extra candidates from the - // FTS index since some will be filtered out. Without a filter we can fetch - // exactly the requested limit. Selective filters remain best-effort: an - // eligible document outside this candidate window is missed (same - // completeness contract as collection filtering). - const ftsLimit = (collectionFilter || filter) ? limit * 10 : limit; + // Unscoped, the MATCH is the whole answer: LIMIT inside the CTE lets FTS5 + // stop at the requested count. Scoped by collection or metadata, the filter + // must see the COMPLETE match set — any inner LIMIT (the old `limit * 10`) + // returns false-empty results whenever stronger out-of-scope matches fill + // the window (#922). MATERIALIZED keeps the planner from flattening the CTE + // and folding the filter back into the MATCH. The set is corpus-bounded: + // at most one row per matching document. + // Because the scope sees every match, one query over the whole collection + // list returns what a query per collection merged by score would (#775): + // a large collection can no longer crowd the others out of a window. + const scoped = Boolean(names || filter); let sql = ` - WITH fts_matches AS ( + WITH fts_matches AS ${scoped ? "MATERIALIZED " : ""}( SELECT rowid, bm25(documents_fts, 1.5, 4.0, 1.0) as bm25_score FROM documents_fts WHERE documents_fts MATCH ? - ORDER BY bm25_score ASC - LIMIT ${ftsLimit} + ${scoped ? "" : `ORDER BY bm25_score ASC LIMIT ${limit}`} ) SELECT 'qmd://' || d.collection || '/' || d.path as filepath, d.collection || '/' || d.path as display_path, d.title, - content.doc as body, + ${cappedBodySql("content.doc")} as body, d.hash, fm.bm25_score, dm.metadata_json @@ -4127,9 +4754,9 @@ export function searchFTS(db: Database, query: string, limit: number = 20, colle WHERE d.active = 1 `; - if (collectionFilter) { - sql += ` AND d.collection = ?`; - params.push(String(collectionFilter)); + if (names) { + sql += ` AND d.collection IN (SELECT value FROM json_each(?))`; + params.push(JSON.stringify(names)); } if (filter) { @@ -4140,8 +4767,9 @@ export function searchFTS(db: Database, query: string, limit: number = 20, colle params.push(...compiledFilter.params); } - // bm25 lower is better; sort ascending. - sql += ` ORDER BY fm.bm25_score ASC LIMIT ?`; + // bm25 lower is better; sort ascending. Ties go to the smaller filepath, so + // the order the collections were named never decides which equal hit survives. + sql += ` ORDER BY fm.bm25_score ASC, filepath ASC LIMIT ?`; params.push(limit); const rows = db.prepare(sql).all(...params) as { filepath: string; display_path: string; title: string; body: string; hash: string; bm25_score: number; metadata_json: string | null }[]; @@ -4177,194 +4805,223 @@ export function searchFTS(db: Database, query: string, limit: number = 20, colle /** sqlite-vec rejects k above this in MATCH queries (v0.1.9). */ const SQLITE_VEC_MAX_K = 4096; -/** - * Max filter-eligible vectors for an exact cosine scan. Above this we fall - * back to global ANN with a capped over-fetch. Exact scan avoids the - * post-filter starvation of small eligible sets — originally small - * collections (#791, #803), now also selective metadata filters; ANN remains - * for very large eligible sets where a full scan would be expensive. - */ -const FILTERED_VEC_EXACT_SCAN_MAX = 20_000; - -const VEC_HASH_SEQ_IN_CHUNK = 400; - -/** - * Exact cosine-distance scan over a known set of hash_seq keys. - * Uses vec_distance_cosine with chunked IN lists (no JOIN with vectors_vec). - */ -function exactVecScanByHashSeq( - db: Database, - embedding: number[], - hashSeqs: string[], - limit: number, -): { hash_seq: string; distance: number }[] { - if (hashSeqs.length === 0 || limit <= 0) return []; - - const queryVec = new Float32Array(embedding); - // Over-fetch a bit so multi-chunk docs can still yield `limit` unique files. - const fetchLimit = Math.max(limit * 3, limit); - const scored: { hash_seq: string; distance: number }[] = []; - - for (let i = 0; i < hashSeqs.length; i += VEC_HASH_SEQ_IN_CHUNK) { - const chunk = hashSeqs.slice(i, i + VEC_HASH_SEQ_IN_CHUNK); - const placeholders = chunk.map(() => "?").join(","); - const rows = db.prepare(` - SELECT hash_seq, vec_distance_cosine(embedding, ?) AS distance - FROM vectors_vec - WHERE hash_seq IN (${placeholders}) - `).all(queryVec, ...chunk) as { hash_seq: string; distance: number }[]; - scored.push(...rows); - } - - scored.sort((a, b) => a.distance - b.distance); - return scored.slice(0, fetchLimit); +interface VecMatch { + rowid: number; + distance: number; } -function annVecScan( - db: Database, - embedding: number[], - k: number, -): { hash_seq: string; distance: number }[] { - const vecK = Math.max(1, Math.min(SQLITE_VEC_MAX_K, k)); - return db.prepare(` - SELECT hash_seq, distance - FROM vectors_vec - WHERE embedding MATCH ? AND k = ? - `).all(new Float32Array(embedding), vecK) as { hash_seq: string; distance: number }[]; +/** One KNN scan target: a collection's partition, or the whole table when no scope is given. */ +interface VecScanTarget { + collectionId?: number; } -export async function searchVec(db: Database, query: string, model: string, limit: number = 20, collectionName?: string | readonly string[], session?: ILLMSession, precomputedEmbedding?: number[], llm?: LlamaCpp, filter?: MetadataFilter): Promise { - const tableExists = db.prepare(`SELECT name FROM sqlite_master WHERE type='table' AND name='vectors_vec'`).get(); - if (!tableExists) return []; - - const embedding = precomputedEmbedding ?? await getEmbedding(query, model, true, session, llm); - if (!embedding) return []; - - const names = scopedCollectionNames(collectionName); - if (names && names.length > 1) { - const lists = await Promise.all( - names.map(name => searchVec(db, query, model, limit, name, session, embedding, llm, filter)), - ); - return mergeSearchResultsByScore(lists, limit); - } - const collectionFilter = names?.[0]; - - // IMPORTANT: We use a two-step query approach here because sqlite-vec virtual tables - // hang indefinitely when combined with JOINs in the same query. Do NOT try to - // "optimize" this by combining into a single query with JOINs - it will break. - // See: https://github.com/tobi/qmd/pull/23 - - // Step 1: Get vector matches from sqlite-vec (no JOINs allowed). - // - // Collection and metadata filters cannot be pushed into MATCH (sqlite-vec - // has no join-safe predicate here). Global ANN + post-filter starves small - // eligible sets: they never enter the top-k (#791, #803). Multiplier - // over-fetch alone is not enough either — sqlite-vec caps k at 4096. For a - // filter we therefore exact-scan the eligible vectors when the set is small - // enough, and only then fall back to capped ANN + post-filter. - let vecResults: { hash_seq: string; distance: number }[]; - - if (collectionFilter || filter) { - let eligibleSql = ` - SELECT DISTINCT cv.hash || '_' || cv.seq AS hash_seq - FROM content_vectors cv - JOIN documents d ON d.hash = cv.hash AND d.active = 1 - `; - const eligibleConditions: string[] = []; - const eligibleParams: (string | number)[] = []; - - if (collectionFilter) { - eligibleConditions.push(`d.collection = ?`); - eligibleParams.push(collectionFilter); - } - - if (filter) { - const compiledFilter = compileMetadataFilter(filter, "d"); - eligibleSql += ` JOIN document_metadata dm ON dm.document_id = d.id`; - eligibleConditions.push(`dm.extraction_version = ${METADATA_EXTRACTION_VERSION}`); - eligibleConditions.push(`dm.extraction_error IS NULL`); - eligibleConditions.push(compiledFilter.sql); - eligibleParams.push(...compiledFilter.params); - } - - eligibleSql += ` WHERE ${eligibleConditions.join(" AND ")}`; - - const eligibleHashSeqs = withLazyContentVectorMigration(db, () => - db.prepare(eligibleSql).all(...eligibleParams) as { hash_seq: string }[], - ).map((r) => r.hash_seq); +/** The document behind a vector match, at its nearest chunk. */ +interface VecDocumentMatch { + rowid: number; + hash: string; + pos: number; + filepath: string; + display_path: string; + title: string; + metadata_json: string | null; + distance: number; +} - if (eligibleHashSeqs.length === 0) return []; +/** + * A metadata filter as SQL over a `document_metadata dm` join, admitting only + * documents whose metadata extraction is current and error-free. + */ +function compileCurrentMetadataFilter(filter: MetadataFilter): { sql: string; params: SQLiteValue[] } { + const compiled = compileMetadataFilter(filter, "d"); + return { + sql: `dm.extraction_version = ${METADATA_EXTRACTION_VERSION} AND dm.extraction_error IS NULL AND ${compiled.sql}`, + params: compiled.params, + }; +} - if (eligibleHashSeqs.length <= FILTERED_VEC_EXACT_SCAN_MAX) { - vecResults = exactVecScanByHashSeq(db, embedding, eligibleHashSeqs, limit); - } else { - // Large eligible set: ANN with over-fetch, hard-capped at sqlite-vec's max k. - vecResults = annVecScan(db, embedding, Math.max(limit * 30, limit * 3)); +/** + * Prepares an exact top-k cosine scan of the vector table, run once per + * target with k in 1..SQLITE_VEC_MAX_K. vec0 evaluates the partition equality + * inside the scan, so a small collection costs its own rows rather than the + * whole index and is never crowded out by a larger one (#775, #791, #803). A + * `rowid IN` restriction is applied inside the scan the same way, so a + * selective metadata filter gets an exact top-k of its own rows instead of + * whatever survives a post-filter of a larger top-k. The restriction is the + * eligibility subquery itself rather than a bound list of rowids, so however + * many rows a filter admits, none of them is read onto the heap. + */ +function knnVecScanner(db: Database, partitioned: boolean, filter?: MetadataFilter): (embedding: Float32Array, k: number, target: VecScanTarget) => VecMatch[] { + const conditions = ["embedding MATCH ?", "k = ?"]; + if (partitioned) conditions.push("collection_id = ?"); + const eligible = filter ? metadataEligibleRowsSql(filter) : undefined; + if (eligible) conditions.push(`rowid IN (SELECT vr.id ${eligible.sql}${partitioned ? " AND vr.collection_id = ?" : ""})`); + const statement = db.prepare(` + SELECT rowid, distance + FROM ${VEC_TABLE} + WHERE ${conditions.join(" AND ")} + `); + return (embedding, k, target) => { + const params: SQLiteValue[] = [embedding, k]; + const partition = vecInteger(target.collectionId ?? 0); + if (partitioned) params.push(partition); + if (eligible) { + params.push(...eligible.params); + if (partitioned) params.push(partition); } - } else { - vecResults = annVecScan(db, embedding, limit * 3); - } + return statement.all(...params) as VecMatch[]; + }; +} - if (vecResults.length === 0) return []; +/** + * FROM and WHERE clauses selecting, as `vr`, the vector rows of active + * documents a metadata filter admits. A row is eligible when any document of + * its collection that holds the row's content passes the filter. + */ +function metadataEligibleRowsSql(filter: MetadataFilter): { sql: string; params: SQLiteValue[] } { + const current = compileCurrentMetadataFilter(filter); + return { + sql: ` + FROM documents d + JOIN document_metadata dm ON dm.document_id = d.id + JOIN ${VEC_COLLECTION_IDS_TABLE} ci ON ci.name = d.collection + JOIN ${VEC_ROWS_TABLE} vr ON vr.hash = d.hash AND vr.collection_id = ci.id + WHERE d.active = 1 AND ${current.sql}`, + params: current.params, + }; +} - // Step 2: Get chunk info and document data - const hashSeqs = vecResults.map(r => r.hash_seq); - const distanceMap = new Map(vecResults.map(r => [r.hash_seq, r.distance])); +/** + * Ids of the collections holding at least one vector row a metadata filter + * admits. The filter runs once per document and EXISTS probes for a row, so + * the cost follows the documents in scope rather than their chunk count. + */ +function metadataEligibleCollections(db: Database, filter: MetadataFilter): Set { + const current = compileCurrentMetadataFilter(filter); + const rows = db.prepare(` + SELECT DISTINCT ci.id AS collectionId + FROM documents d + JOIN document_metadata dm ON dm.document_id = d.id + JOIN ${VEC_COLLECTION_IDS_TABLE} ci ON ci.name = d.collection + WHERE d.active = 1 AND ${current.sql} + AND EXISTS (SELECT 1 FROM ${VEC_ROWS_TABLE} vr WHERE vr.hash = d.hash AND vr.collection_id = ci.id) + `).all(...current.params) as { collectionId: number }[]; + return new Set(rows.map(row => row.collectionId)); +} - // Build query for document lookup - const placeholders = hashSeqs.map(() => '?').join(','); - let docSql = ` +/** + * Prepares step 2 of a vector search: the documents behind a set of vector + * rows, one per file at its nearest chunk, nearest first. The rowids are bound + * as one JSON parameter, so the statement text stays fixed and a match set up + * to the 4096 k cap never meets SQLite's bound-parameter limit. Bodies are + * left out: one long document can hold every match, and only the final results + * load theirs. + */ +function vecDocumentResolver(db: Database, filter?: MetadataFilter): (matches: readonly VecMatch[]) => VecDocumentMatch[] { + // Re-apply the filter on the document join: vectors are content-scoped, + // so one hash can belong to both matching and non-matching documents. + const current = filter ? compileCurrentMetadataFilter(filter) : undefined; + const statement = withLazyContentVectorMigration(db, () => db.prepare(` SELECT - cv.hash || '_' || cv.seq as hash_seq, + vr.id AS rowid, cv.hash, cv.pos, 'qmd://' || d.collection || '/' || d.path as filepath, d.collection || '/' || d.path as display_path, d.title, - content.doc as body, dm.metadata_json - FROM content_vectors cv - JOIN documents d ON d.hash = cv.hash AND d.active = 1 + FROM json_each(?) j + JOIN ${VEC_ROWS_TABLE} vr ON vr.id = j.value + JOIN ${VEC_COLLECTION_IDS_TABLE} ci ON ci.id = vr.collection_id + JOIN content_vectors cv ON cv.hash = vr.hash AND cv.seq = vr.seq + JOIN documents d ON d.hash = vr.hash AND d.collection = ci.name JOIN content ON content.hash = d.hash LEFT JOIN document_metadata dm ON dm.document_id = d.id - WHERE cv.hash || '_' || cv.seq IN (${placeholders}) - `; - const params: (string | number)[] = [...hashSeqs]; + WHERE d.active = 1${current ? ` AND ${current.sql}` : ""} + `)); + return (matches) => { + if (matches.length === 0) return []; + const distanceByRowid = new Map(matches.map(r => [r.rowid, r.distance])); + const rows = statement.all(rowidList(matches.map(r => r.rowid)), ...(current?.params ?? [])) as Omit[]; + const best = new Map(); + for (const row of rows) { + const distance = distanceByRowid.get(row.rowid) ?? 1; + const existing = best.get(row.filepath); + if (!existing || distance < existing.distance) best.set(row.filepath, { ...row, distance }); + } + return Array.from(best.values()).sort((a, b) => a.distance - b.distance); + }; +} - if (collectionFilter) { - docSql += ` AND d.collection = ?`; - params.push(collectionFilter); +/** + * Nearest documents of one scan target. The first KNN asks for three chunks + * per requested document. Chunks of one long document can still fill every + * slot, so while the matches collapse into fewer than `limit` documents and + * the target holds rows beyond them, k doubles, up to sqlite-vec's cap. + */ +function nearestVecDocuments( + scan: ReturnType, + resolve: ReturnType, + queryVec: Float32Array, + limit: number, + target: VecScanTarget, +): VecDocumentMatch[] { + for (let k = limit * 3; ; k *= 2) { + const vecK = Math.max(1, Math.min(SQLITE_VEC_MAX_K, k)); + const matches = scan(queryVec, vecK, target); + const documents = resolve(matches); + if (documents.length >= limit || matches.length < vecK || vecK === SQLITE_VEC_MAX_K) return documents; } +} - if (filter) { - // Re-apply the filter on the document join: vectors are content-scoped, - // so one hash can belong to both matching and non-matching documents. - const compiledFilter = compileMetadataFilter(filter, "d"); - docSql += ` AND dm.extraction_version = ${METADATA_EXTRACTION_VERSION} AND dm.extraction_error IS NULL AND ${compiledFilter.sql}`; - params.push(...compiledFilter.params); - } +export async function searchVec(db: Database, query: string, model: string, limit: number = 20, collectionName?: string | readonly string[], session?: ILLMSession, precomputedEmbedding?: number[], llm?: LlamaCpp, filter?: MetadataFilter): Promise { + if (!hasVectorIndex(db)) return []; - const docRows = withLazyContentVectorMigration(db, () => db.prepare(docSql).all(...params) as { - hash_seq: string; hash: string; pos: number; filepath: string; - display_path: string; title: string; body: string; metadata_json: string | null; - }[]); + const embedding = precomputedEmbedding ?? await getEmbedding(query, model, true, session, llm); + if (!embedding) return []; - // Combine with distances and dedupe by filepath - const seen = new Map(); - for (const row of docRows) { - const distance = distanceMap.get(row.hash_seq) ?? 1; - const existing = seen.get(row.filepath); - if (!existing || distance < existing.bestDist) { - seen.set(row.filepath, { row, bestDist: distance }); - } + const names = scopedCollectionNames(collectionName); + let collectionIds: number[] | undefined; + if (names) { + collectionIds = Array.from(resolveCollectionIds(db, names).values()); + if (collectionIds.length === 0) return []; } + const eligible = filter ? metadataEligibleCollections(db, filter) : undefined; - return Array.from(seen.values()) - .sort((a, b) => a.bestDist - b.bestDist) + // IMPORTANT: We use a two-step query approach here because sqlite-vec virtual tables + // hang indefinitely when combined with JOINs in the same query. Do NOT try to + // "optimize" this by combining into a single query with JOINs - it will break. + // See: https://github.com/tobi/qmd/pull/23 + + // Step 1 gets vector matches from sqlite-vec (no JOINs allowed), one KNN per + // collection in scope; step 2 resolves them to documents. One statement per + // member rather than `collection_id IN (...)`: the IN form yields k rows per + // value only because SQLite runs vec0's filter once per value, which is a + // planner detail rather than a vec0 contract. + const scanned = collectionIds && eligible ? collectionIds.filter(id => eligible.has(id)) : collectionIds; + const scanTargets: VecScanTarget[] = scanned ? scanned.map(collectionId => ({ collectionId })) : eligible?.size === 0 ? [] : [{}]; + if (scanTargets.length === 0) return []; + const scan = knnVecScanner(db, collectionIds !== undefined, filter); + const resolve = vecDocumentResolver(db, filter); + const queryVec = new Float32Array(embedding); + // Bodies are capped at BODY_CAP_CHARS, as in searchFTS, so a large document cannot + // put its whole text on the heap for each result. + const bodyOf = db.prepare(`SELECT ${cappedBodySql("doc")} AS doc FROM content WHERE hash = ?`); + + // Each target yields its own nearest `limit` documents (or all it holds), so + // merging them by distance gives the scope's exact nearest `limit`. Ties go + // to the smaller filepath, as in searchFTS. + return scanTargets + .flatMap(target => nearestVecDocuments(scan, resolve, queryVec, limit, target)) + .sort((a, b) => a.distance - b.distance || compareFilepaths(a, b)) .slice(0, limit) - .map(({ row, bestDist }) => { + .flatMap((row): SearchResult[] => { + // The body is read after resolution, outside its snapshot: another + // process's orphaned-content cleanup can delete the row in between. + const content = bodyOf.get(row.hash) as { doc: string } | null | undefined; + if (content == null) return []; + const body = content.doc; const collectionName = row.filepath.split('//')[1]?.split('/')[0] || ""; - return { + return [{ filepath: row.filepath, displayPath: row.display_path, title: row.title, @@ -4372,14 +5029,14 @@ export async function searchVec(db: Database, query: string, model: string, limi docid: getDocid(row.hash), collectionName, modifiedAt: "", // Not available in vec query - bodyLength: row.body.length, - body: row.body, + bodyLength: body.length, + body, context: getContextForFile(db, row.filepath), metadata: parseMetadataJson(row.metadata_json), - score: 1 - bestDist, // Cosine similarity = 1 - cosine distance + score: 1 - row.distance, // Cosine similarity = 1 - cosine distance source: "vec" as const, chunkPos: row.pos, - }; + }]; }); } @@ -4402,8 +5059,9 @@ async function getEmbedding(text: string, model: string, isQuery: boolean, sessi */ export function getHashesForEmbedding(db: Database, model: string = DEFAULT_EMBED_MODEL): { hash: string; body: string; path: string }[] { const fingerprint = getEmbeddingFingerprint(model); - return withLazyContentVectorMigration(db, () => db.prepare(` - SELECT d.hash, c.doc as body, MIN(d.path) as path + return withLazyContentVectorMigration(db, () => { + const stmt = db.prepare(` + SELECT d.hash, ${cappedBodySql("c.doc")} as body, MIN(d.path) as path FROM documents d JOIN content c ON d.hash = c.hash LEFT JOIN ( @@ -4415,29 +5073,38 @@ export function getHashesForEmbedding(db: Database, model: string = DEFAULT_EMBE WHERE d.active = 1 AND (v.hash IS NULL OR v.chunk_count < v.expected_chunks) GROUP BY d.hash - `).all(model, fingerprint) as { hash: string; body: string; path: string }[]); + `); + // Large-result query (up to 9k): use iterate() to stream, bound heap + const results: { hash: string; body: string; path: string }[] = []; + for (const row of stmt.iterate(model, fingerprint) as IterableIterator<{ hash: string; body: string; path: string }>) { + results.push(row); + } + return results; + }); } /** * Clear embeddings for the whole index, or just for one collection. * - * When `collection` is omitted the entire content_vectors table is emptied and - * the vectors_vec virtual table is dropped (it is recreated with the right + * When `collection` is omitted the content_vectors and vector_rows tables are + * emptied and the vec0 table is dropped (it is recreated with the right * dimensions on the next embed run). * - * When `collection` is provided, only vectors whose hash is referenced - * exclusively by active documents in that collection are removed. Hashes - * shared with active documents in other collections are left in place so - * vector search keeps working there (content_vectors is keyed globally by - * content hash; identical document bodies across collections share a row). - * vectors_vec is preserved so other collections keep working unless the scoped - * clear empties content_vectors entirely, in which case it is dropped so the - * next embed can recreate the table with the current dimensions. + * When `collection` is provided, only chunks whose hash is referenced + * exclusively by active documents in that collection are removed, from that + * collection's partition and from content_vectors. Hashes shared with active + * documents in other collections are left in place so vector search keeps + * working there (content_vectors is keyed globally by content hash; identical + * document bodies across collections share a row). The vec0 table is dropped + * only when the scoped clear empties content_vectors entirely, so the next + * embed can recreate it with the current dimensions. */ export function clearAllEmbeddings(db: Database, collection?: string): void { if (!collection) { - db.exec(`DELETE FROM content_vectors`); - db.exec(`DROP TABLE IF EXISTS vectors_vec`); + db.transaction(() => { + db.exec(`DELETE FROM content_vectors`); + dropVectorIndex(db); + }).immediate(); return; } @@ -4452,23 +5119,15 @@ export function clearAllEmbeddings(db: Database, collection?: string): void { AND d2.collection != d.collection ) `; - - const vecTableExists = db - .prepare(`SELECT 1 FROM sqlite_master WHERE type='table' AND name='vectors_vec'`) - .get(); - - withLazyContentVectorMigration(db, () => { - if (vecTableExists) { - const hashSeqRows = db.prepare(` - SELECT cv.hash, cv.seq - FROM content_vectors cv - WHERE cv.hash IN (${exclusiveHashesQuery}) - `).all(collection) as { hash: string; seq: number }[]; - - const delVec = db.prepare(`DELETE FROM vectors_vec WHERE hash_seq = ?`); - for (const row of hashSeqRows) { - delVec.run(`${row.hash}_${row.seq}`); - } + const collectionId = resolveCollectionId(db, collection); + + withLazyContentVectorMigration(db, () => db.transaction(() => { + if (collectionId !== undefined) { + const rows = db.prepare(` + SELECT id FROM ${VEC_ROWS_TABLE} + WHERE collection_id = ? AND hash IN (${exclusiveHashesQuery}) + `).all(collectionId, collection) as { id: number }[]; + deletePartitionRows(db, rows.map((row) => row.id)); } db.prepare(` @@ -4479,21 +5138,68 @@ export function clearAllEmbeddings(db: Database, collection?: string): void { const remaining = db .prepare(`SELECT COUNT(*) AS n FROM content_vectors`) .get() as { n: number }; - if (remaining.n === 0) { - db.exec(`DROP TABLE IF EXISTS vectors_vec`); + if (remaining.n === 0) dropVectorIndex(db); + }).immediate()); +} + +/** Partition rows written per commit by copyVectorsToNewCollections. */ +export const VECTOR_COPY_BATCH_ROWS = 1000; + +/** + * Give every active (hash, collection) pair whose chunks are embedded a row + * in that collection's partition, copied from any partition that holds the + * chunk, so a hash that gains a collection is searchable there without + * another model call. A chunk no partition holds is queued for embedding by + * deleting its hash's rows, which the pending detector then picks up. + * + * Rows commit in batches of VECTOR_COPY_BATCH_ROWS so a collection that gains + * thousands of embedded hashes does not hold the write lock for the whole + * copy; an interrupted run resumes from whatever missingPartitionRows still + * reports. + */ +export function copyVectorsToNewCollections(db: Database, collection?: string): { copied: number; queued: number } { + const layout = vecLayout(db); + if (layout.kind === "legacy" || !isSqliteVecAvailable()) return { copied: 0, queued: 0 }; + return withLazyContentVectorMigration(db, () => { + const missing = missingPartitionRows(db, collection); + if (missing.length === 0) return { copied: 0, queued: 0 }; + const writer = layout.kind === "partitioned" ? new PartitionWriter(db) : null; + const storedVector = writer ? storedEmbeddingLookup(db) : null; + const deleteChunks = db.prepare(`DELETE FROM content_vectors WHERE hash = ?`); + const queued = new Set(); + let copied = 0; + const copyBatch = (rows: MissingPartitionRow[]) => { + const queuedNow: string[] = []; + for (const row of rows) { + if (queued.has(row.hash)) continue; + const vector = storedVector?.(row.hash, row.seq); + if (!vector || !writer) { + queued.add(row.hash); + queuedNow.push(row.hash); + continue; + } + if (writer.writeToCollection(row.hash, row.seq, row.collection, vector)) copied++; + } + // Deleting by hash also removes rows an earlier batch already committed + // for it; the run-level set keeps a later batch from copying it again. + for (const hash of queuedNow) { + deletePartitionRowsOfHash(db, hash); + deleteChunks.run(hash); + } + }; + for (let start = 0; start < missing.length; start += VECTOR_COPY_BATCH_ROWS) { + const batch = missing.slice(start, start + VECTOR_COPY_BATCH_ROWS); + db.transaction(() => copyBatch(batch)).immediate(); } + return { copied, queued: queued.size }; }); } /** - * Insert a single embedding into both content_vectors and vectors_vec tables. - * The hash_seq key is formatted as "hash_seq" for the vectors_vec table. - * - * content_vectors is inserted first so that getHashesForEmbedding (which checks - * only content_vectors) won't re-select the hash on a crash between the two inserts. - * - * vectors_vec uses DELETE + INSERT instead of INSERT OR REPLACE because sqlite-vec's - * vec0 virtual tables silently ignore the OR REPLACE conflict clause. + * Insert a single embedding: the content_vectors row plus one row in the + * partition of every active collection that holds the hash, in one + * transaction so a crash cannot leave the tables out of step. A hash with no + * active document gets its content_vectors row only. */ export function insertEmbedding( db: Database, @@ -4506,18 +5212,14 @@ export function insertEmbedding( totalChunks: number = 1, fingerprint: string = getEmbeddingFingerprint(model) ): void { - const hashSeq = `${hash}_${seq}`; - withLazyContentVectorMigration(db, () => { - // Insert content_vectors first — crash-safe ordering (see getHashesForEmbedding) - const insertContentVectorStmt = db.prepare(`INSERT OR REPLACE INTO content_vectors (hash, seq, pos, model, embed_fingerprint, total_chunks, embedded_at) VALUES (?, ?, ?, ?, ?, ?, ?)`); - insertContentVectorStmt.run(hash, seq, pos, model, fingerprint, totalChunks, embeddedAt); - - // vec0 virtual tables don't support OR REPLACE — use DELETE + INSERT - const deleteVecStmt = db.prepare(`DELETE FROM vectors_vec WHERE hash_seq = ?`); - const insertVecStmt = db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`); - deleteVecStmt.run(hashSeq); - insertVecStmt.run(hashSeq, embedding); + db.transaction(() => { + db.prepare(`INSERT OR REPLACE INTO content_vectors (hash, seq, pos, model, embed_fingerprint, total_chunks, embedded_at) VALUES (?, ?, ?, ?, ?, ?, ?)`) + .run(hash, seq, pos, model, fingerprint, totalChunks, embeddedAt); + for (const collection of activeCollectionsOfHash(db, hash)) { + upsertPartitionVector(db, hash, seq, allocateCollectionId(db, collection), embedding); + } + })(); }); } @@ -4526,15 +5228,12 @@ function removeIncompleteEmbeddings(db: Database, expectedChunksByHash: Map(); + return expansions.filter((e) => { + const key = `${e.type}\n${e.query}`; + if (seen.has(key)) return false; + seen.add(key); + return true; + }); +} + export async function expandQuery(query: string, model: string = DEFAULT_QUERY_MODEL, db: Database, llmOverride?: LlamaCpp): Promise { // Check cache first — stored as JSON preserving types. Intent is // deliberately absent from both the cache key and the generation call: @@ -4562,9 +5276,9 @@ export async function expandQuery(query: string, model: string = DEFAULT_QUERY_M const rows = parsed as Array>; // Migrate old cache format: { type, text } → { type, query } if (rows.length > 0 && typeof rows[0]?.["query"] === "string") { - return rows.map((r) => ({ type: r["type"] as ExpandedQuery["type"], query: String(r["query"]) })); + return uniqueExpansions(rows.map((r) => ({ type: r["type"] as ExpandedQuery["type"], query: String(r["query"]) }))); } else if (rows.length > 0 && typeof rows[0]?.["text"] === "string") { - return rows.map((r) => ({ type: r["type"] as ExpandedQuery["type"], query: String(r["text"]) })); + return uniqueExpansions(rows.map((r) => ({ type: r["type"] as ExpandedQuery["type"], query: String(r["text"]) }))); } } catch { // Old cache format (pre-typed, newline-separated text) — re-expand @@ -4577,9 +5291,9 @@ export async function expandQuery(query: string, model: string = DEFAULT_QUERY_M // Map Queryable[] → ExpandedQuery[] (same shape, decoupled from llm.ts internals). // Filter out entries that duplicate the original query text. - const expanded: ExpandedQuery[] = results + const expanded = uniqueExpansions(results .filter(r => r.text !== query) - .map(r => ({ type: r.type, query: r.text })); + .map(r => ({ type: r.type, query: r.text }))); if (expanded.length > 0) { setCachedResult(db, cacheKey, JSON.stringify(expanded)); @@ -5237,6 +5951,18 @@ export function findDocuments( // ============================================================================= export function getStatus(db: Database, model: string = DEFAULT_EMBED_MODEL): IndexStatus { + const statusSummary = getStatusSummary(db, model); + const metadataOverviews = listMetadataCollectionSummaries(db, STATUS_METADATA_KEY_LIMIT); + return { + ...statusSummary, + collections: statusSummary.collections.map(collection => { + const metadataOverview = metadataOverviews.get(collection.name); + return { ...collection, metadataKeyCount: metadataOverview?.totalKeys ?? 0, metadataKeys: metadataOverview?.keys ?? [] }; + }), + }; +} + +export function getStatusSummary(db: Database, model: string = DEFAULT_EMBED_MODEL): IndexStatusSummary { // DB is source of truth for collections — config provides supplementary metadata const dbCollections = db.prepare(` SELECT @@ -5252,7 +5978,7 @@ export function getStatus(db: Database, model: string = DEFAULT_EMBED_MODEL): In const storeCollections = getStoreCollections(db); const configLookup = new Map(storeCollections.map(c => [c.name, { path: c.path, pattern: c.pattern }])); - const collections: CollectionInfo[] = dbCollections.map(row => { + const collections: IndexStatusSummary["collections"] = dbCollections.map(row => { const config = configLookup.get(row.name); return { name: row.name, @@ -5272,7 +5998,7 @@ export function getStatus(db: Database, model: string = DEFAULT_EMBED_MODEL): In const totalDocs = (db.prepare(`SELECT COUNT(*) as c FROM documents WHERE active = 1`).get() as { c: number }).c; const needsEmbedding = getHashesNeedingEmbedding(db, undefined, model); - const hasVectors = !!db.prepare(`SELECT name FROM sqlite_master WHERE type='table' AND name='vectors_vec'`).get(); + const hasVectors = hasVectorIndex(db); return { totalDocuments: totalDocs, @@ -5553,9 +6279,7 @@ export async function hybridQuery( const rankedLists: RankedResult[][] = []; const rankedListMeta: RankedListMeta[] = []; const docidMap = new Map(); // filepath -> docid - const hasVectors = !!store.db.prepare( - `SELECT name FROM sqlite_master WHERE type='table' AND name='vectors_vec'` - ).get(); + const hasVectors = hasVectorIndex(store.db); // Step 1: BM25 probe — strong signal skips expensive LLM expansion // When intent is provided, disable strong-signal bypass — the obvious BM25 @@ -5876,9 +6600,7 @@ export async function vectorSearchQuery( const filter = options?.filter; const intent = options?.intent; - const hasVectors = !!store.db.prepare( - `SELECT name FROM sqlite_master WHERE type='table' AND name='vectors_vec'` - ).get(); + const hasVectors = hasVectorIndex(store.db); if (!hasVectors) return []; // Expand query — filter to vec/hyde only (lex queries target FTS, not vector) @@ -5997,30 +6719,27 @@ export async function structuredSearch( const rankedLists: RankedResult[][] = []; const rankedListMeta: RankedListMeta[] = []; const docidMap = new Map(); // filepath -> docid - const hasVectors = !!store.db.prepare( - `SELECT name FROM sqlite_master WHERE type='table' AND name='vectors_vec'` - ).get(); + const hasVectors = hasVectorIndex(store.db); - // Helper to run search across collections (or all if undefined) - const collectionList = collections ?? [undefined]; // undefined = all collections + // Each search yields ONE ranked list over the union of the named collections + // (undefined = all). searchFTS/searchVec merge per-collection results by score, + // so the RRF weight below boosts the first search, not the first collection. // Step 1: Run FTS for all lex searches (sync, instant) for (const search of searches) { if (search.type === 'lex') { - for (const coll of collectionList) { - const ftsResults = store.searchFTS(search.query, 20, coll, filter); - if (ftsResults.length > 0) { - for (const r of ftsResults) docidMap.set(r.filepath, r.docid); - rankedLists.push(ftsResults.map(r => ({ - file: r.filepath, displayPath: r.displayPath, - title: r.title, body: r.body || "", score: r.score, - }))); - rankedListMeta.push({ - source: "fts", - queryType: "lex", - query: search.query, - }); - } + const ftsResults = store.searchFTS(search.query, 20, collections, filter); + if (ftsResults.length > 0) { + for (const r of ftsResults) docidMap.set(r.filepath, r.docid); + rankedLists.push(ftsResults.map(r => ({ + file: r.filepath, displayPath: r.displayPath, + title: r.title, body: r.body || "", score: r.score, + }))); + rankedListMeta.push({ + source: "fts", + queryType: "lex", + query: search.query, + }); } } } @@ -6044,23 +6763,21 @@ export async function structuredSearch( const embedding = embeddings[i]?.embedding; if (!embedding) continue; - for (const coll of collectionList) { - const vecResults = await store.searchVec( - vecSearches[i]!.query, embedModel, 20, coll, - undefined, embedding, filter - ); - if (vecResults.length > 0) { - for (const r of vecResults) docidMap.set(r.filepath, r.docid); - rankedLists.push(vecResults.map(r => ({ - file: r.filepath, displayPath: r.displayPath, - title: r.title, body: r.body || "", score: r.score, - }))); - rankedListMeta.push({ - source: "vec", - queryType: vecSearches[i]!.type, - query: vecSearches[i]!.query, - }); - } + const vecResults = await store.searchVec( + vecSearches[i]!.query, embedModel, 20, collections, + undefined, embedding, filter + ); + if (vecResults.length > 0) { + for (const r of vecResults) docidMap.set(r.filepath, r.docid); + rankedLists.push(vecResults.map(r => ({ + file: r.filepath, displayPath: r.displayPath, + title: r.title, body: r.body || "", score: r.score, + }))); + rankedListMeta.push({ + source: "vec", + queryType: vecSearches[i]!.type, + query: vecSearches[i]!.query, + }); } } } diff --git a/src/vec-layout.ts b/src/vec-layout.ts new file mode 100644 index 000000000..3242a58b6 --- /dev/null +++ b/src/vec-layout.ts @@ -0,0 +1,303 @@ +/** + * vec-layout.ts - Names and resolution of the sqlite-vec vector layout. + * + * Vectors live in one vec0 table partitioned by an integer collection id, one + * row per (chunk, active collection). `vector_rows` maps each vec0 rowid back + * to its (hash, seq, collection_id), and `vector_collection_ids` maps names to + * ids. The integer key is what makes `qmd collection rename` a one-row update: + * vec0 rejects UPDATE on a partition key, so a name-keyed partition would have + * to rewrite every row of the collection. + * + * A layout change needs a new permanent table name: `ALTER TABLE RENAME` on a + * vec0 table leaves its shadow tables under the old name (sqlite-vec 0.1.9), + * and SQLite refuses to rename shadow tables by hand. `vecLayout` is the one + * place that knows which table is live, so no other module spells the names. + */ + +import type { Database } from "./db.js"; + +export const VEC_TABLE = "vectors_by_collection"; +export const VEC_ROWS_TABLE = "vector_rows"; +export const VEC_COLLECTION_IDS_TABLE = "vector_collection_ids"; +export const LEGACY_VEC_TABLE = "vectors_vec"; + +/** vec0's shadow tables for a virtual table name. */ +export type VecShadowTables = { + chunks: string; + rowids: string; + vectorChunks: string; + info: string; +}; + +export type VecLayout = + | { kind: "none" } + | { kind: "legacy"; table: typeof LEGACY_VEC_TABLE; dimensions: number | null; keyedByHashSeq: boolean; shadow: VecShadowTables } + | { kind: "partitioned"; table: typeof VEC_TABLE; dimensions: number | null; shadow: VecShadowTables }; + +export type ReadableVecLayout = Exclude; + +function shadowTables(table: string): VecShadowTables { + return { + chunks: `${table}_chunks`, + rowids: `${table}_rowids`, + vectorChunks: `${table}_vector_chunks00`, + info: `${table}_info`, + }; +} + +export function parseDimensions(sql: string | null): number | null { + const match = sql?.match(/float\[(\d+)\]/); + return match?.[1] ? parseInt(match[1], 10) : null; +} + +/** + * Which vector layout the database holds. A legacy table wins while it + * exists: the migration drops it as its last step, so that drop is the flip. + */ +export function vecLayout(db: Database): VecLayout { + const rows = db.prepare( + `SELECT name, sql FROM sqlite_master WHERE type = 'table' AND name IN (?, ?)` + ).all(LEGACY_VEC_TABLE, VEC_TABLE) as { name: string; sql: string | null }[]; + + const legacy = rows.find((row) => row.name === LEGACY_VEC_TABLE); + if (legacy) { + return { + kind: "legacy", + table: LEGACY_VEC_TABLE, + dimensions: parseDimensions(legacy.sql), + keyedByHashSeq: (legacy.sql ?? "").includes("hash_seq"), + shadow: shadowTables(LEGACY_VEC_TABLE), + }; + } + const partitioned = rows.find((row) => row.name === VEC_TABLE); + if (partitioned) { + return { + kind: "partitioned", + table: VEC_TABLE, + dimensions: parseDimensions(partitioned.sql), + shadow: shadowTables(VEC_TABLE), + }; + } + return { kind: "none" }; +} + +/** True when the partitioned vector index exists; vector search reads nothing else. */ +export function hasVectorIndex(db: Database): boolean { + return vecLayout(db).kind === "partitioned"; +} + +/** + * A schema entry can outlive the vec0 module (reopening without sqlite-vec + * loaded), in which case touching the table throws "no such module". + */ +export function vecTableReadable(db: Database, layout: ReadableVecLayout): boolean { + try { + db.prepare(`SELECT 1 FROM ${layout.table} LIMIT 0`).get(); + return true; + } catch { + return false; + } +} + +export function createVectorMetadataTables(db: Database): void { + db.exec(` + CREATE TABLE IF NOT EXISTS ${VEC_COLLECTION_IDS_TABLE} ( + id INTEGER PRIMARY KEY, + name TEXT NOT NULL UNIQUE + ) + `); + db.exec(` + CREATE TABLE IF NOT EXISTS ${VEC_ROWS_TABLE} ( + id INTEGER PRIMARY KEY, + hash TEXT NOT NULL, + seq INTEGER NOT NULL, + collection_id INTEGER NOT NULL, + UNIQUE(hash, seq, collection_id) + ) + `); + db.exec(`CREATE INDEX IF NOT EXISTS idx_vector_rows_collection ON ${VEC_ROWS_TABLE}(collection_id)`); +} + +export function createPartitionedVecTable(db: Database, dimensions: number): void { + db.exec( + `CREATE VIRTUAL TABLE IF NOT EXISTS ${VEC_TABLE} USING vec0(collection_id INTEGER PARTITION KEY, embedding float[${dimensions}] distance_metric=cosine)` + ); +} + +export function resolveCollectionId(db: Database, name: string): number | undefined { + const row = db.prepare(`SELECT id FROM ${VEC_COLLECTION_IDS_TABLE} WHERE name = ?`).get(name) as { id: number } | undefined; + return row?.id; +} + +/** Ids of the named collections that have one; names without an id are absent. */ +export function resolveCollectionIds(db: Database, names: readonly string[]): Map { + const ids = new Map(); + const stmt = db.prepare(`SELECT id FROM ${VEC_COLLECTION_IDS_TABLE} WHERE name = ?`); + for (const name of names) { + const row = stmt.get(name) as { id: number } | undefined; + if (row) ids.set(name, row.id); + } + return ids; +} + +export function allocateCollectionId(db: Database, name: string): number { + db.prepare(`INSERT OR IGNORE INTO ${VEC_COLLECTION_IDS_TABLE} (name) VALUES (?)`).run(name); + const id = resolveCollectionId(db, name); + if (id === undefined || id === null) { + throw new Error(`Could not allocate a vector collection id for '${name}'`); + } + return id; +} + +export function renameCollectionId(db: Database, oldName: string, newName: string): void { + db.prepare(`UPDATE ${VEC_COLLECTION_IDS_TABLE} SET name = ? WHERE name = ?`).run(newName, oldName); +} + +export function deleteCollectionId(db: Database, name: string): void { + db.prepare(`DELETE FROM ${VEC_COLLECTION_IDS_TABLE} WHERE name = ?`).run(name); +} + +/** + * Integer parameter for the vec0 table. better-sqlite3 binds every JS number + * as a double, and vec0 checks the bound type of its partition key and rowid + * instead of applying column affinity, so those parameters go in as bigint. + */ +export function vecInteger(value: number | bigint): bigint { + return BigInt(value); +} + +/** + * One bound parameter for `json_each(?)`: a scoped search over many + * partitions can return more rowids than SQLite allows as separate + * parameters (32,766). + */ +export function rowidList(rowids: readonly number[]): string { + return JSON.stringify(rowids); +} + +/** Deletes vector rows by rowid from the vec0 table (when it exists) and the mapping. */ +export function deletePartitionRows(db: Database, rowids: readonly number[]): void { + if (rowids.length === 0) return; + const deleteVec = hasVectorIndex(db) ? db.prepare(`DELETE FROM ${VEC_TABLE} WHERE rowid = ?`) : null; + const deleteRow = db.prepare(`DELETE FROM ${VEC_ROWS_TABLE} WHERE id = ?`); + for (const id of rowids) { + deleteVec?.run(vecInteger(id)); + deleteRow.run(id); + } +} + +/** Deletes every partition row of a hash, in every collection. */ +export function deletePartitionRowsOfHash(db: Database, hash: string): void { + const rows = db.prepare(`SELECT id FROM ${VEC_ROWS_TABLE} WHERE hash = ?`).all(hash) as { id: number }[]; + deletePartitionRows(db, rows.map((row) => row.id)); +} + +const ACTIVE_COLLECTIONS_OF_HASH_SQL = `SELECT DISTINCT collection FROM documents WHERE hash = ? AND active = 1`; + +/** Names of the collections with an active document for the hash: the partitions that must hold its vectors. */ +export function activeCollectionsOfHash(db: Database, hash: string): string[] { + return (db.prepare(ACTIVE_COLLECTIONS_OF_HASH_SQL).all(hash) as { collection: string }[]).map((row) => row.collection); +} + +/** + * Replace or insert the vector of (hash, seq) in one collection's partition. + * vec0 ignores OR REPLACE, so an existing row is deleted and its rowid reused. + */ +export function upsertPartitionVector(db: Database, hash: string, seq: number, collectionId: number, embedding: Float32Array | Uint8Array): void { + const existing = db.prepare(`SELECT id FROM ${VEC_ROWS_TABLE} WHERE hash = ? AND seq = ? AND collection_id = ?`) + .get(hash, seq, collectionId) as { id: number } | undefined; + let rowid: number; + if (existing) { + db.prepare(`DELETE FROM ${VEC_TABLE} WHERE rowid = ?`).run(vecInteger(existing.id)); + rowid = existing.id; + } else { + rowid = Number(db.prepare(`INSERT INTO ${VEC_ROWS_TABLE} (hash, seq, collection_id) VALUES (?, ?, ?)`).run(hash, seq, collectionId).lastInsertRowid); + } + db.prepare(`INSERT INTO ${VEC_TABLE} (rowid, collection_id, embedding) VALUES (?, ?, ?)`).run(vecInteger(rowid), vecInteger(collectionId), embedding); +} + +/** Bulk writer that inserts a vector under a collection unless that partition row exists. */ +export class PartitionWriter { + private readonly ids = new Map(); + private readonly collectionsOf; + private readonly insertRow; + private readonly insertVec; + + constructor(private readonly db: Database) { + this.collectionsOf = db.prepare(ACTIVE_COLLECTIONS_OF_HASH_SQL); + this.insertRow = db.prepare(`INSERT OR IGNORE INTO ${VEC_ROWS_TABLE} (hash, seq, collection_id) VALUES (?, ?, ?)`); + this.insertVec = db.prepare(`INSERT INTO ${VEC_TABLE} (rowid, collection_id, embedding) VALUES (?, ?, ?)`); + } + + private idOf(collection: string): number { + let id = this.ids.get(collection); + if (id === undefined) { + id = allocateCollectionId(this.db, collection); + this.ids.set(collection, id); + } + return id; + } + + /** True when a row was written; false when the partition already held (hash, seq). */ + writeToCollection(hash: string, seq: number, collection: string, embedding: Uint8Array | Float32Array): boolean { + const id = this.idOf(collection); + const inserted = this.insertRow.run(hash, seq, id); + if (inserted.changes === 0) return false; + this.insertVec.run(vecInteger(inserted.lastInsertRowid), vecInteger(id), embedding); + return true; + } + + writeToActiveCollections(hash: string, seq: number, embedding: Uint8Array | Float32Array): number { + let written = 0; + for (const row of this.collectionsOf.all(hash) as { collection: string }[]) { + if (this.writeToCollection(hash, seq, row.collection, embedding)) written++; + } + return written; + } +} + +/** + * Lookup of stored vector bytes by (hash, seq) from whichever partition holds + * the chunk, with its two statements prepared once for callers that loop. + */ +export function storedEmbeddingLookup(db: Database): (hash: string, seq: number) => Uint8Array | undefined { + const rowOf = db.prepare(`SELECT id FROM ${VEC_ROWS_TABLE} WHERE hash = ? AND seq = ? LIMIT 1`); + const vectorOf = db.prepare(`SELECT embedding FROM ${VEC_TABLE} WHERE rowid = ?`); + return (hash, seq) => { + const row = rowOf.get(hash, seq) as { id: number } | undefined; + if (!row) return undefined; + const stored = vectorOf.get(vecInteger(row.id)) as { embedding: Uint8Array } | undefined; + return stored?.embedding; + }; +} + +/** Stored vector bytes of (hash, seq) from whichever partition holds it. */ +export function storedEmbedding(db: Database, hash: string, seq: number): Uint8Array | undefined { + return storedEmbeddingLookup(db)(hash, seq); +} + +/** (hash, seq) behind a vec0 rowid. */ +export function partitionRowKey(db: Database, rowid: number): { hash: string; seq: number } | undefined { + return db.prepare(`SELECT hash, seq FROM ${VEC_ROWS_TABLE} WHERE id = ?`).get(rowid) as { hash: string; seq: number } | undefined; +} + +export type MissingPartitionRow = { hash: string; seq: number; collection: string }; + +/** + * Chunks recorded in content_vectors for an active document that have no row + * in that document's collection partition. Each one is a vector to copy from + * another partition (a hash that gained a collection) or, with no partition + * holding it, a chunk to embed again. + */ +export function missingPartitionRows(db: Database, collection?: string): MissingPartitionRow[] { + const filter = collection ? `AND collection = ?` : ``; + return db.prepare(` + SELECT cv.hash, cv.seq, d.collection + FROM content_vectors cv + JOIN (SELECT DISTINCT hash, collection FROM documents WHERE active = 1 ${filter}) d ON d.hash = cv.hash + LEFT JOIN ${VEC_COLLECTION_IDS_TABLE} ci ON ci.name = d.collection + LEFT JOIN ${VEC_ROWS_TABLE} vr ON vr.hash = cv.hash AND vr.seq = cv.seq AND vr.collection_id = ci.id + WHERE vr.id IS NULL + ORDER BY cv.hash, cv.seq, d.collection + `).all(...(collection ? [collection] : [])) as MissingPartitionRow[]; +} diff --git a/test/_helpers/doctor-vector-check-worker.ts b/test/_helpers/doctor-vector-check-worker.ts new file mode 100644 index 000000000..d0d23af05 --- /dev/null +++ b/test/_helpers/doctor-vector-check-worker.ts @@ -0,0 +1,37 @@ +import { createStore } from "../../src/store.js"; +import { checkEmbeddingVectorSamples } from "../../src/cli/qmd.js"; +import { LlamaCpp, setDefaultLlamaCpp } from "../../src/llm.js"; +import { heapLimitBinds } from "./heap-limit.js"; + +// Every stored vector is [1, 0] and every re-embedded chunk returns [1, 0], so +// the check passes whenever it can sample, read and chunk the bodies. +class ConstantLlm extends LlamaCpp { + async tokenize(text: string) { return new Array(Math.ceil(text.length / 16)).fill(1); } + async embed() { return { embedding: [1, 0], model: "model" }; } +} + +const store = createStore(":memory:"); +try { + setDefaultLlamaCpp(new ConstantLlm()); + const body = "word ".repeat(13_107); // About 64 KiB per document. + store.ensureVecTable(2); + store.db.transaction(() => { + for (let doc = 0; doc < 16; doc++) { + const hash = `document-${doc}`; + store.insertContent(hash, body, "2026-01-01"); + for (let path = 0; path < 8; path++) { + store.insertDocument("test", `${hash}-${path}.md`, hash, hash, "2026-01-01", "2026-01-01"); + } + for (let seq = 0; seq < 32; seq++) { + store.insertEmbedding(hash, seq, 0, new Float32Array([1, 0]), "model", "2026-01-01", 32, "current"); + } + } + })(); + store.db.exec("PRAGMA temp_store = MEMORY"); + store.db.exec("PRAGMA hard_heap_limit = 33554432"); + const result = await checkEmbeddingVectorSamples(store.db, "model", "current"); + console.log(JSON.stringify({ ...result, limitBinds: heapLimitBinds(store.db, 33554432) })); +} finally { + setDefaultLlamaCpp(null); + store.close(); +} diff --git a/test/_helpers/doctor-vector-sample-worker.ts b/test/_helpers/doctor-vector-sample-worker.ts new file mode 100644 index 000000000..3c2135949 --- /dev/null +++ b/test/_helpers/doctor-vector-sample-worker.ts @@ -0,0 +1,30 @@ +import { createStore, getEmbeddingVectorSamples } from "../../src/store.js"; + +const store = createStore(":memory:"); +try { + const body = "word ".repeat(13_107); // About 64 KiB per document. + const insertChunk = store.db.prepare(` + INSERT INTO content_vectors (hash, seq, model, embed_fingerprint, embedded_at) + VALUES (?, ?, 'model', 'current', '2026-01-01') + `); + store.db.transaction(() => { + for (let doc = 0; doc < 16; doc++) { + const hash = `document-${doc}`; + store.insertContent(hash, body, "2026-01-01"); + for (let path = 0; path < 8; path++) { + store.insertDocument("test", `${hash}-${path}.md`, hash, hash, "2026-01-01", "2026-01-01"); + } + for (let seq = 0; seq < 32; seq++) insertChunk.run(hash, seq); + } + })(); + store.db.exec("PRAGMA temp_store = MEMORY"); + store.db.exec("PRAGMA hard_heap_limit = 33554432"); + const samples = getEmbeddingVectorSamples(store.db, "model", "current"); + console.log(JSON.stringify({ + samples: samples.length, + distinctChunks: new Set(samples.map(sample => `${sample.hash}:${sample.seq}`)).size, + bodiesComplete: samples.every(sample => sample.body === body), + })); +} finally { + store.close(); +} diff --git a/test/_helpers/heap-limit.ts b/test/_helpers/heap-limit.ts new file mode 100644 index 000000000..0a01a7ccc --- /dev/null +++ b/test/_helpers/heap-limit.ts @@ -0,0 +1,16 @@ +import type { Database } from "../../src/db.js"; + +/** + * Whether SQLite enforces the hard heap limit just set on `db`: an allocation + * past it fails with SQLITE_NOMEM only when SQLite tracks memory. Bun's bundled + * SQLite on Linux does; better-sqlite3 (DEFAULT_MEMSTATUS=0) does not, so a + * heap-budget test proves its budget only where this returns true. + */ +export function heapLimitBinds(db: Database, limitBytes: number): boolean { + try { + db.prepare(`SELECT length(randomblob(?)) AS n`).get(limitBytes + 8 * 1024 * 1024); + return false; + } catch { + return true; + } +} diff --git a/test/_helpers/legacy-adoption-sample-worker.ts b/test/_helpers/legacy-adoption-sample-worker.ts new file mode 100644 index 000000000..768da69dd --- /dev/null +++ b/test/_helpers/legacy-adoption-sample-worker.ts @@ -0,0 +1,46 @@ +/** + * legacy-adoption-sample-worker - runs legacy fingerprint adoption under a + * 128 MiB SQLite heap limit. + * + * Spawned by store.test.ts: PRAGMA hard_heap_limit is process-wide and cannot + * be raised again, so it must not leak into the test runner. The fixture has + * large bodies with many legacy chunks and duplicate paths; materializing one + * body per chunk (per path) needs far more than the budget. + */ +import { createStore, maybeAdoptLegacyEmbeddingFingerprint } from "../../src/store.ts"; + +const model = "model"; +const store = createStore(":memory:"); +try { + const body = "word ".repeat(13_107); // About 64 KiB per document. + const insertChunk = store.db.prepare(` + INSERT INTO content_vectors (hash, seq, model, embed_fingerprint, embedded_at) + VALUES (?, ?, ?, '', '2026-01-01') + `); + store.db.transaction(() => { + for (let doc = 0; doc < 16; doc++) { + const hash = `document-${doc}`; + store.insertContent(hash, body, "2026-01-01"); + for (let path = 0; path < 8; path++) { + store.insertDocument("test", `${hash}-${path}.md`, hash, hash, "2026-01-01", "2026-01-01"); + } + for (let seq = 0; seq < 32; seq++) insertChunk.run(hash, seq, model); + } + })(); + // Store the sample chunk's vector: the one the stub embedder returns. + store.ensureVecTable(3); + store.insertEmbedding("document-0", 0, 0, new Float32Array([0.1, 0.2, 0.3]), model, "2026-01-01", 32, ""); + // Adoption only tokenizes, detokenizes and embeds; a stub covers that surface. + store.llm = { + async tokenize(text: string) { return new Array(Math.max(1, Math.ceil(text.length / 16))).fill(1); }, + async detokenize(tokens: readonly number[]) { return "x".repeat(tokens.length * 16); }, + async embed() { return { embedding: [0.1, 0.2, 0.3], model }; }, + } as any; + + store.db.exec("PRAGMA temp_store = MEMORY"); + store.db.exec("PRAGMA hard_heap_limit = 134217728"); + const result = await maybeAdoptLegacyEmbeddingFingerprint(store, model); + console.log(JSON.stringify(result)); +} finally { + store.close(); +} diff --git a/test/_helpers/legacy-adoption-shared-body-worker.ts b/test/_helpers/legacy-adoption-shared-body-worker.ts new file mode 100644 index 000000000..02a0a0664 --- /dev/null +++ b/test/_helpers/legacy-adoption-shared-body-worker.ts @@ -0,0 +1,44 @@ +/** + * legacy-adoption-shared-body-worker - runs legacy fingerprint adoption for + * one legacy chunk whose 1 MiB body sits behind 200 active paths, under a + * 16 MiB SQLite heap limit. + * + * A sample query that carries c.doc through its join copies the body once per + * active path, far past the budget, while the chunk count stays at one. The + * index lives in a temp file: the FTS table stores every path's copy of the + * body, which an in-memory database would count against the heap limit. + */ +import { mkdtempSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { createStore, maybeAdoptLegacyEmbeddingFingerprint } from "../../src/store.ts"; +import { heapLimitBinds } from "./heap-limit.ts"; + +const model = "model"; +const dir = mkdtempSync(join(tmpdir(), "qmd-legacy-shared-body-")); +const store = createStore(join(dir, "index.sqlite")); +try { + const body = "word ".repeat(209_716); // About 1 MiB. + store.insertContent("shared", body, "2026-01-01"); + store.db.transaction(() => { + for (let path = 0; path < 200; path++) { + store.insertDocument("test", `shared-${path}.md`, "shared", "shared", "2026-01-01", "2026-01-01"); + } + })(); + store.ensureVecTable(3); + store.insertEmbedding("shared", 0, 0, new Float32Array([0.1, 0.2, 0.3]), model, "2026-01-01", 1, ""); + // Adoption only tokenizes, detokenizes and embeds; a stub covers that surface. + store.llm = { + async tokenize(text: string) { return new Array(Math.max(1, Math.ceil(text.length / 16))).fill(1); }, + async detokenize(tokens: readonly number[]) { return "x".repeat(tokens.length * 16); }, + async embed() { return { embedding: [0.1, 0.2, 0.3], model }; }, + } as any; + + store.db.exec("PRAGMA temp_store = MEMORY"); + store.db.exec("PRAGMA hard_heap_limit = 16777216"); + const result = await maybeAdoptLegacyEmbeddingFingerprint(store, model); + console.log(JSON.stringify({ checked: result.checked, adopted: result.adopted, limitBinds: heapLimitBinds(store.db, 16777216) })); +} finally { + store.close(); + rmSync(dir, { recursive: true, force: true }); +} diff --git a/test/bin-wrapper.test.ts b/test/bin-wrapper.test.ts index ae31bb19a..687118ab9 100644 --- a/test/bin-wrapper.test.ts +++ b/test/bin-wrapper.test.ts @@ -1,7 +1,7 @@ import { afterEach, describe, expect, test } from "vitest"; import { chmodSync, copyFileSync, mkdtempSync, mkdirSync, readFileSync, realpathSync, rmSync, symlinkSync, writeFileSync } from "node:fs"; import { tmpdir } from "node:os"; -import { dirname, join, relative } from "node:path"; +import { delimiter, dirname, join, relative } from "node:path"; import { execFileSync, spawnSync } from "node:child_process"; import { fileURLToPath } from "node:url"; @@ -58,24 +58,35 @@ fi return { root, capturePath, runtimeBin }; } +function guardedEsmCli(body: string): string { + return [ + 'import { realpathSync, writeFileSync } from "node:fs";', + 'import { fileURLToPath } from "node:url";', + 'const __filename = fileURLToPath(import.meta.url);', + 'const argv1 = process.argv[1];', + 'const isMain = argv1 === __filename', + ' || argv1?.endsWith("/qmd.ts")', + ' || argv1?.endsWith("/qmd.js")', + ' || (argv1 != null && realpathSync(argv1) === __filename);', + 'if (isMain) {', + ` ${body}`, + '}', + '', + ].join("\n"); +} + function makePackage(root: string, packagePath: string, lockfiles: string[] = [], options: { dist?: boolean; source?: boolean; tsx?: boolean; git?: boolean } = {}) { const packageRoot = join(root, packagePath); const includeDist = options.dist ?? true; mkdirSync(join(packageRoot, "bin"), { recursive: true }); + writeFileSync(join(packageRoot, "package.json"), '{"type":"module"}\n'); copyFileSync(join(repoRoot, "bin", "qmd"), join(packageRoot, "bin", "qmd")); chmodSync(join(packageRoot, "bin", "qmd"), 0o755); if (includeDist) { mkdirSync(join(packageRoot, "dist", "cli"), { recursive: true }); writeFileSync( join(packageRoot, "dist", "cli", "qmd.js"), - [ - 'const { writeFileSync } = require("node:fs");', - 'const capture = process.env.QMD_WRAPPER_CAPTURE;', - 'if (capture) {', - ' writeFileSync(capture, ["node", process.argv[1], ...process.argv.slice(2)].join("\\n") + "\\n");', - '}', - '', - ].join("\n"), + guardedEsmCli('writeFileSync(process.env.QMD_WRAPPER_CAPTURE, ["node", process.argv[1], ...process.argv.slice(2)].join("\\n") + "\\n");'), ); } if (options.source) { @@ -299,15 +310,104 @@ describe("bin/qmd package wrapper", () => { expect(result.args).toEqual([realpathSync(join(packageRoot, "src", "cli", "qmd.ts")), "--version"]); }); - test("node child uses process.execPath, not a different PATH node (#577 leftover)", () => { + test("source-mode Node child uses process.execPath, not a different PATH node (#577 leftover)", () => { const { root, runtimeBin, capturePath } = makeTempFixture(); + const packageRoot = makePackage(root, "qmd", [], { source: true, tsx: true, git: true }); + + const child = spawnSync(REAL_NODE, [join(packageRoot, "bin", "qmd"), "--version"], { + env: { + ...process.env, + PATH: `${runtimeBin}${delimiter}${process.env.PATH ?? ""}`, + QMD_WRAPPER_CAPTURE: capturePath, + }, + encoding: "utf8", + }); + const [runtime, scriptPath, ...args] = readFileSync(capturePath, "utf8").trimEnd().split("\n"); + + expect(child.error).toBeUndefined(); + expect(child.status).toBe(0); + expect(runtime).toBe("node"); + expect(scriptPath).toBe(realpathSync(join(packageRoot, "node_modules", "tsx", "dist", "cli.mjs"))); + expect(args).toEqual([realpathSync(join(packageRoot, "src", "cli", "qmd.ts")), "--version"]); + }); + + test("imports dist in the current Node process when Node is already selected", () => { + const { root, capturePath } = makeTempFixture(); const packageRoot = makePackage(root, "node_modules/@tobilu/qmd"); + writeFileSync( + join(packageRoot, "dist", "cli", "qmd.js"), + guardedEsmCli('writeFileSync(process.env.QMD_WRAPPER_CAPTURE, String(process.pid));'), + ); - const result = runWrapper(join(packageRoot, "bin", "qmd"), runtimeBin, capturePath); + const result = spawnSync(REAL_NODE, [join(packageRoot, "bin", "qmd"), "--version"], { + env: { ...process.env, QMD_WRAPPER_CAPTURE: capturePath }, + encoding: "utf8", + }); - expect(result.runtime).toBe("node"); - expect(result.scriptPath).toBe(realpathSync(join(packageRoot, "dist", "cli", "qmd.js"))); - expect(result.args).toEqual(["--version"]); + expect(result.error).toBeUndefined(); + expect(result.status).toBe(0); + expect(Number(readFileSync(capturePath, "utf8"))).toBe(result.pid); + }); + + test.skipIf(typeof process.versions.bun !== "string")("imports dist in the current Bun process when Bun is already selected", () => { + const { root, capturePath } = makeTempFixture(); + const packageRoot = makePackage(root, "node_modules/@tobilu/qmd", ["bun.lock"]); + writeFileSync( + join(packageRoot, "dist", "cli", "qmd.js"), + guardedEsmCli('writeFileSync(process.env.QMD_WRAPPER_CAPTURE, String(process.pid));'), + ); + + const result = spawnSync(process.execPath, [join(packageRoot, "bin", "qmd"), "--version"], { + env: { ...process.env, QMD_WRAPPER_CAPTURE: capturePath }, + encoding: "utf8", + }); + + expect(result.error).toBeUndefined(); + expect(result.status).toBe(0); + expect(Number(readFileSync(capturePath, "utf8"))).toBe(result.pid); + }); + + test("an in-process dist import keeps the CLI's argv and exit status", () => { + const { root, capturePath } = makeTempFixture(); + const packageRoot = makePackage(root, "node_modules/@tobilu/qmd"); + const distEntry = join(packageRoot, "dist", "cli", "qmd.js"); + writeFileSync( + distEntry, + guardedEsmCli('writeFileSync(process.env.QMD_WRAPPER_CAPTURE, JSON.stringify({ pid: process.pid, argv: process.argv.slice(1) })); process.exit(3);'), + ); + + const result = spawnSync(REAL_NODE, [join(packageRoot, "bin", "qmd"), "search", "x"], { + env: { ...process.env, QMD_WRAPPER_CAPTURE: capturePath }, + encoding: "utf8", + }); + + expect(result.error).toBeUndefined(); + expect(result.status).toBe(3); + const captured = JSON.parse(readFileSync(capturePath, "utf8")); + expect(captured.pid).toBe(result.pid); + expect(realpathSync(captured.argv[0])).toBe(realpathSync(distEntry)); + expect(captured.argv.slice(1)).toEqual(["search", "x"]); + }); + + test("a Node launch of a Bun-locked package still hands off to a bun child", () => { + const { root, capturePath, runtimeBin } = makeTempFixture(); + const packageRoot = makePackage(root, "node_modules/@tobilu/qmd", ["bun.lock"]); + writeFileSync( + join(packageRoot, "dist", "cli", "qmd.js"), + guardedEsmCli('writeFileSync(process.env.QMD_WRAPPER_CAPTURE, "in-process\\n");'), + ); + + const result = spawnSync(REAL_NODE, [join(packageRoot, "bin", "qmd"), "search", "x"], { + env: { ...process.env, PATH: `${runtimeBin}${delimiter}${process.env.PATH ?? ""}`, QMD_WRAPPER_CAPTURE: capturePath }, + encoding: "utf8", + }); + + expect(result.error).toBeUndefined(); + expect(result.status).toBe(0); + const [runtime, scriptPath, ...args] = readFileSync(capturePath, "utf8").trimEnd().split("\n"); + expect(runtime).toBe("bun"); + expect(realpathSync(scriptPath)).toBe(realpathSync(join(packageRoot, "dist", "cli", "qmd.js"))); + expect(args).toEqual(["search", "x"]); }); test("explains how to build when dist is missing and source cannot run", () => { diff --git a/test/cleanup-repack-partitions.test.ts b/test/cleanup-repack-partitions.test.ts new file mode 100644 index 000000000..ca0cd263f --- /dev/null +++ b/test/cleanup-repack-partitions.test.ts @@ -0,0 +1,111 @@ +/** + * qmd cleanup repacks a partitioned vector table in place: every surviving + * vector keeps its bytes, each partition's newest chunk keeps its rows where + * they are, and the doctor's stored-vector check still passes afterwards. + */ +import { describe, test, expect, afterEach } from "vitest"; +import { mkdtemp, rm } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { createStore, getEmbeddingFingerprint, runCleanup, type Store } from "../src/store.js"; +import { VEC_ROWS_TABLE, VEC_TABLE, resolveCollectionId, storedEmbedding, vecInteger } from "../src/vec-layout.js"; +import { checkEmbeddingVectorSamples } from "../src/cli/qmd.js"; +import { LlamaCpp, setDefaultLlamaCpp } from "../src/llm.js"; + +const MODEL = "repack-model"; +// Two vec0 chunks of 1024 slots per partition: two live rows stay in the old +// chunk and one in the newest, after cleanup removes the rest as orphans. +const TOTAL = 1100; +const KEEP = [0, 1, 1099]; +const COLLECTIONS = ["alpha", "beta"]; + +function vectorFor(i: number): number[] { + return [1, i / TOTAL, 0]; +} + +// Re-embeds a chunk as the vector its document was stored with, so the doctor +// check compares like with like. +class StoredVectorLlm extends LlamaCpp { + async tokenize(text: string) { return new Array(Math.max(1, Math.ceil(text.length / 16))).fill(1); } + async embed(text: string) { + const index = Number(/\d{5}/.exec(text)?.[0]); + return { embedding: vectorFor(index), model: MODEL }; + } +} + +let dir: string | undefined; +let store: Store | undefined; + +afterEach(async () => { + setDefaultLlamaCpp(null); + store?.close(); + store = undefined; + if (dir) await rm(dir, { recursive: true, force: true }); + dir = undefined; +}); + +function hashOf(collection: string, i: number): string { + return `${collection}${String(i).padStart(5, "0")}`; +} + +function seedPartition(s: Store, collection: string): void { + const now = new Date().toISOString(); + const kept = new Set(KEEP); + s.db.transaction(() => { + for (let i = 0; i < TOTAL; i++) { + const hash = hashOf(collection, i); + s.insertContent(hash, `# ${hash}\n\nbody ${hash}`, now); + s.insertDocument(collection, `${hash}.md`, hash, hash, now, now); + s.insertEmbedding(hash, 0, 0, new Float32Array(vectorFor(i)), MODEL, now, 1); + if (!kept.has(i)) s.deactivateDocument(collection, `${hash}.md`); + } + })(); +} + +/** "chunk_id:chunk_offset" of a hash's vector row in the vec0 table. */ +function rowPosition(s: Store, hash: string): string { + const row = s.db.prepare(` + SELECT r.chunk_id AS chunk, r.chunk_offset AS offset + FROM ${VEC_ROWS_TABLE} vr JOIN ${VEC_TABLE}_rowids r ON r.rowid = vr.id + WHERE vr.hash = ? + `).get(hash) as { chunk: number; offset: number }; + return `${row.chunk}:${row.offset}`; +} + +function newestChunk(s: Store, collection: string): number { + const partition = vecInteger(resolveCollectionId(s.db, collection)!); + const row = s.db.prepare(`SELECT MAX(chunk_id) AS chunk FROM ${VEC_TABLE}_chunks WHERE partition00 = ?`) + .get(partition) as { chunk: number }; + return row.chunk; +} + +describe("qmd cleanup repack on a partitioned vector table", () => { + test("keeps every surviving vector's bytes and each partition's newest chunk, and the doctor check still passes", async () => { + dir = await mkdtemp(join(tmpdir(), "qmd-repack-partitions-")); + const s = createStore(join(dir, "index.sqlite")); + store = s; + s.ensureVecTable(3); + for (const collection of COLLECTIONS) seedPartition(s, collection); + const survivors = COLLECTIONS.flatMap(collection => KEEP.map(i => hashOf(collection, i))); + const bytesBefore = new Map(survivors.map(hash => [hash, Array.from(storedEmbedding(s.db, hash, 0)!)])); + const newestBefore = new Map(COLLECTIONS.map(collection => [collection, rowPosition(s, hashOf(collection, 1099))])); + for (const collection of COLLECTIONS) { + expect(rowPosition(s, hashOf(collection, 1099)).split(":")[0]).toBe(String(newestChunk(s, collection))); + } + + const stats = runCleanup(s.db); + + expect(stats.vectorsRepacked).toBe(true); + for (const hash of survivors) { + expect(Array.from(storedEmbedding(s.db, hash, 0)!)).toEqual(bytesBefore.get(hash)); + } + for (const collection of COLLECTIONS) { + expect(rowPosition(s, hashOf(collection, 1099))).toBe(newestBefore.get(collection)); + // The old chunk's two rows moved behind the newest one, which emptied it. + expect(rowPosition(s, hashOf(collection, 0)).split(":")[0]).toBe(String(newestChunk(s, collection))); + } + setDefaultLlamaCpp(new StoredVectorLlm()); + expect(await checkEmbeddingVectorSamples(s.db, MODEL, getEmbeddingFingerprint(MODEL), survivors.length)) + .toMatchObject({ ok: true }); + }); +}); diff --git a/test/cleanup.test.ts b/test/cleanup.test.ts index c9ec9c02f..cc1c214b3 100644 --- a/test/cleanup.test.ts +++ b/test/cleanup.test.ts @@ -16,9 +16,14 @@ import { cleanupOrphanedContent, countOrphanedContent, previewCleanup, + repackVectors, runCleanup, + searchVec, + vectorTableLayout, type Store, } from "../src/store.js"; +import type { Database } from "../src/db.js"; +import { VEC_ROWS_TABLE, VEC_TABLE, deletePartitionRows, partitionRowKey, resolveCollectionId, vecInteger } from "../src/vec-layout.js"; let store: Store | null = null; @@ -107,3 +112,241 @@ describe("qmd cleanup reclaim (#550)", () => { expect(s.db.prepare(`SELECT count(*) as c FROM documents`).get()).toEqual({ c: 1 }); }); }); + +describe("qmd cleanup vector repack", () => { + const DIMS = 3; + + const hashOf = (prefix: string, i: number) => `${prefix}${String(i).padStart(5, "0")}`; + + /** + * Inserts `total` embedded documents into `collection`, then deletes every + * vector row except the indexes in `keep`. vec0 drops a chunk only when it + * is completely empty, so keeping one row in each chunk leaves the holes. + */ + async function seedVectors(s: Store, total: number, keep: readonly number[], collection = "docs", prefix = "vec"): Promise { + const now = new Date().toISOString(); + s.ensureVecTable(DIMS); + s.db.transaction(() => { + for (let i = 0; i < total; i++) { + const hash = hashOf(prefix, i); + insertContent(s.db, hash, `body ${hash}`, now); + insertDocument(s.db, collection, `${hash}.md`, hash, hash, now, now); + s.insertEmbedding(hash, 0, 0, new Float32Array([1, i * 0.001, 0]), "test-model", now, 1); + } + const kept = new Set(keep.map((i) => hashOf(prefix, i))); + const rows = s.db.prepare(`SELECT vr.id, vr.hash FROM ${VEC_ROWS_TABLE} vr JOIN vector_collection_ids ci ON ci.id = vr.collection_id WHERE ci.name = ?`).all(collection) as { id: number; hash: string }[]; + deletePartitionRows(s.db, rows.filter((row) => !kept.has(row.hash)).map((row) => row.id)); + })(); + } + + /** Two chunks of 1024 slots holding three live rows: two in the first chunk, one in the second. */ + const HOLEY = { total: 1100, keep: [0, 1, 1099] } as const; + + function nearest(s: Store): string { + const row = s.db.prepare(`SELECT rowid FROM ${VEC_TABLE} WHERE embedding MATCH ? AND k = 1`).get(new Float32Array([1, 0, 0])) as { rowid: number }; + return partitionRowKey(s.db, row.rowid)!.hash; + } + + /** Hashes whose vector rows are present in the vec0 table, in hash order. */ + function liveHashes(s: Store): string[] { + return (s.db.prepare(`SELECT hash FROM ${VEC_ROWS_TABLE} WHERE id IN (SELECT rowid FROM ${VEC_TABLE}) ORDER BY hash`).all() as { hash: string }[]).map((row) => row.hash); + } + + function partitionCount(s: Store, collection: string): number { + return (s.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_TABLE} WHERE collection_id = ?`).get(vecInteger(resolveCollectionId(s.db, collection)!)) as { c: number }).c; + } + + test("layout counts the chunks a packed table would need against the chunks in use", async () => { + const s = await openStore(); + await seedVectors(s, HOLEY.total, HOLEY.keep); + + expect(vectorTableLayout(s.db)).toEqual({ rows: 3, chunks: 2, neededChunks: 1, occupancy: 0.5 }); + }); + + test("layout is null without a vector table", async () => { + const s = await openStore(); + expect(vectorTableLayout(s.db)).toBeNull(); + expect(previewCleanup(s.db)).toMatchObject({ vectorLayout: null, vectorsRepacked: false }); + }); + + test("runCleanup repacks a table that is mostly holes and keeps every live row", async () => { + const s = await openStore(); + await seedVectors(s, HOLEY.total, HOLEY.keep); + + const stats = runCleanup(s.db); + expect(stats.vectorsRepacked).toBe(true); + expect(stats.vectorLayout).toMatchObject({ chunks: 2, neededChunks: 1 }); + expect(vectorTableLayout(s.db)).toEqual({ rows: 3, chunks: 1, neededChunks: 1, occupancy: 1 }); + expect(liveHashes(s)).toEqual(["vec00000", "vec00001", "vec01099"]); + expect(nearest(s)).toBe("vec00000"); + }); + + test("runCleanup leaves a packed table alone", async () => { + const s = await openStore(); + await seedVectors(s, 3, [0, 1, 2]); + + const stats = runCleanup(s.db); + expect(stats.vectorsRepacked).toBe(false); + expect(stats.vectorLayout).toEqual({ rows: 3, chunks: 1, neededChunks: 1, occupancy: 1 }); + }); + + test("previewCleanup reports the repack without rebuilding", async () => { + const s = await openStore(); + await seedVectors(s, HOLEY.total, HOLEY.keep); + + expect(previewCleanup(s.db)).toMatchObject({ vectorsRepacked: true, vectorLayout: { chunks: 2, neededChunks: 1 } }); + expect(vectorTableLayout(s.db)).toMatchObject({ chunks: 2 }); + }); + + test("a chunk move that fails midway leaves that chunk's rows in place", async () => { + const s = await openStore(); + await seedVectors(s, HOLEY.total, HOLEY.keep); + const failing: Database = { + prepare: (sql: string) => { + const real = s.db.prepare(sql); + if (!sql.startsWith(`INSERT INTO ${VEC_TABLE} (`)) return real; + return { ...real, get: real.get.bind(real), all: real.all.bind(real), iterate: real.iterate.bind(real), run: () => { throw new Error("injected failure during re-insert"); } }; + }, + transaction: (fn) => s.db.transaction(fn), + exec: (sql: string) => s.db.exec(sql), + loadExtension: (path: string) => s.db.loadExtension(path), + close: () => s.db.close(), + }; + + expect(() => repackVectors(failing)).toThrow("injected failure during re-insert"); + expect(vectorTableLayout(s.db)).toEqual({ rows: 3, chunks: 2, neededChunks: 1, occupancy: 0.5 }); + expect(nearest(s)).toBe("vec00000"); + expect(liveHashes(s)).toEqual(["vec00000", "vec00001", "vec01099"]); + }); + + test("previewCleanup projects the layout after orphan removal, matching what runCleanup does", async () => { + const s = await openStore(); + const now = new Date().toISOString(); + s.ensureVecTable(DIMS); + s.db.transaction(() => { + for (let i = 0; i < 2048; i++) { + const hash = `orphan${String(i).padStart(5, "0")}`; + insertContent(s.db, hash, `body ${hash}`, now); + insertDocument(s.db, "docs", `${hash}.md`, hash, hash, now, now); + s.insertEmbedding(hash, 0, 0, new Float32Array([1, i * 0.001, 0]), "test-model", now, 1); + } + })(); + for (let i = 0; i < 2048; i += 2) deactivateDocument(s.db, "docs", `orphan${String(i).padStart(5, "0")}.md`); + + expect(vectorTableLayout(s.db)).toEqual({ rows: 2048, chunks: 2, neededChunks: 2, occupancy: 1 }); + expect(previewCleanup(s.db)).toMatchObject({ orphanedVectors: 1024, vectorsRepacked: true, vectorLayout: { rows: 1024, chunks: 2, neededChunks: 1 } }); + + const stats = runCleanup(s.db); + expect(stats.vectorsRepacked).toBe(true); + expect(vectorTableLayout(s.db)).toEqual({ rows: 1024, chunks: 1, neededChunks: 1, occupancy: 1 }); + }); + + test("repack moves rows within their own partition and leaves each partition's newest chunk alone", async () => { + const s = await openStore(); + await seedVectors(s, HOLEY.total, HOLEY.keep, "alpha", "alpha"); + await seedVectors(s, HOLEY.total, HOLEY.keep, "beta", "beta"); + // Rows of different partitions never share a chunk: a packed copy needs one chunk each. + expect(vectorTableLayout(s.db)).toEqual({ rows: 6, chunks: 4, neededChunks: 2, occupancy: 0.5 }); + + const stats = runCleanup(s.db); + + expect(stats.vectorsRepacked).toBe(true); + expect(vectorTableLayout(s.db)).toEqual({ rows: 6, chunks: 2, neededChunks: 2, occupancy: 1 }); + expect(runCleanup(s.db).vectorsRepacked).toBe(false); + expect(partitionCount(s, "alpha")).toBe(3); + expect(partitionCount(s, "beta")).toBe(3); + const alpha = await searchVec(s.db, "ignored", "test-model", 5, "alpha", undefined, [1, 0, 0]); + expect(alpha.map((r) => r.hash).sort()).toEqual(["alpha00000", "alpha00001", "alpha01099"]); + const beta = await searchVec(s.db, "ignored", "test-model", 5, "beta", undefined, [1, 0, 0]); + expect(beta.map((r) => r.hash).sort()).toEqual(["beta00000", "beta00001", "beta01099"]); + }); + + test("a table below the occupancy trigger with no chunk to move is not repacked", async () => { + const s = await openStore(); + // Two chunks at 973 of 1024 slots (above the per-chunk fill) and a newest chunk of 51. + const keep = Array.from({ length: 2099 }, (_, i) => i).filter((i) => !(i <= 50 || (i >= 1024 && i <= 1074))); + await seedVectors(s, 2099, keep); + expect(vectorTableLayout(s.db)).toEqual({ rows: 1997, chunks: 3, neededChunks: 2, occupancy: 2 / 3 }); + + expect(previewCleanup(s.db).vectorsRepacked).toBe(false); + expect(runCleanup(s.db).vectorsRepacked).toBe(false); + }); + + test("previewCleanup leaves out a chunk that orphan removal empties, as runCleanup finds it", async () => { + const s = await openStore(); + await seedVectors(s, 1100, Array.from({ length: 1100 }, (_, i) => i)); + s.db.transaction(() => { + for (let i = 0; i < 1024; i++) deactivateDocument(s.db, "docs", `${hashOf("vec", i)}.md`); + })(); + + const preview = previewCleanup(s.db); + expect(preview).toMatchObject({ orphanedVectors: 1024, vectorsRepacked: false, vectorLayout: { rows: 76, chunks: 1, neededChunks: 1, occupancy: 1 } }); + + const stats = runCleanup(s.db); + expect(stats.vectorsRepacked).toBe(false); + expect(stats.vectorLayout).toEqual(preview.vectorLayout); + }); + + test("a repack reads each chunk again, so a rowid reused by another partition stays there", async () => { + const s = await openStore(); + // Alpha holds two sparse chunks and a newest one; beta is small and packed. + await seedVectors(s, 2100, [0, 1, 1024, 1025, 2099], "alpha", "alpha"); + await seedVectors(s, 3, [0, 1, 2], "beta", "beta"); + const betaId = resolveCollectionId(s.db, "beta")!; + const reused = (s.db.prepare(`SELECT id FROM ${VEC_ROWS_TABLE} WHERE hash = ?`).get(hashOf("alpha", 1024)) as { id: number }).id; + + repackVectors(s.db, (done) => { + if (done !== 1) return; + // Between two chunk transactions, another writer drops a row of the second + // sparse chunk and a beta embed takes over its rowid. + deletePartitionRows(s.db, [reused]); + s.db.prepare(`INSERT INTO ${VEC_ROWS_TABLE} (id, hash, seq, collection_id) VALUES (?, ?, 0, ?)`).run(reused, "betanew", betaId); + s.db.prepare(`INSERT INTO ${VEC_TABLE} (rowid, collection_id, embedding) VALUES (?, ?, ?)`).run(vecInteger(reused), vecInteger(betaId), new Float32Array([0, 1, 0])); + }); + + const row = s.db.prepare(`SELECT collection_id AS collectionId FROM ${VEC_TABLE} WHERE rowid = ?`).get(vecInteger(reused)) as { collectionId: number }; + expect(row.collectionId).toBe(betaId); + expect(partitionCount(s, "alpha")).toBe(4); + expect(partitionCount(s, "beta")).toBe(4); + }); + + test("a failed move of a chunk's last row rolls back the emptied chunk", async () => { + const s = await openStore(); + // The first chunk keeps one row, so moving it empties the chunk inside the transaction. + await seedVectors(s, 1100, [0, 1099]); + const failing: Database = { + prepare: (sql: string) => { + const real = s.db.prepare(sql); + if (!sql.startsWith(`INSERT INTO ${VEC_TABLE} (`)) return real; + return { ...real, get: real.get.bind(real), all: real.all.bind(real), iterate: real.iterate.bind(real), run: () => { throw new Error("injected failure during re-insert"); } }; + }, + transaction: (fn) => s.db.transaction(fn), + exec: (sql: string) => s.db.exec(sql), + loadExtension: (path: string) => s.db.loadExtension(path), + close: () => s.db.close(), + }; + + expect(() => repackVectors(failing)).toThrow("injected failure during re-insert"); + expect(vectorTableLayout(s.db)).toEqual({ rows: 2, chunks: 2, neededChunks: 1, occupancy: 0.5 }); + expect(liveHashes(s)).toEqual(["vec00000", "vec01099"]); + expect(nearest(s)).toBe("vec00000"); + }); + + test("previewCleanup counts a shared hash's row as orphaned only in the collection that dropped it", async () => { + const s = await openStore(); + const now = new Date().toISOString(); + s.ensureVecTable(DIMS); + insertContent(s.db, "sharedhash", "body sharedhash", now); + insertDocument(s.db, "alpha", "shared.md", "shared", "sharedhash", now, now); + insertDocument(s.db, "beta", "shared.md", "shared", "sharedhash", now, now); + s.insertEmbedding("sharedhash", 0, 0, new Float32Array([1, 0, 0]), "test-model", now, 1); + expect(partitionCount(s, "alpha")).toBe(1); + expect(partitionCount(s, "beta")).toBe(1); + deactivateDocument(s.db, "beta", "shared.md"); + + expect(previewCleanup(s.db).vectorLayout).toMatchObject({ rows: 1, chunks: 1 }); + runCleanup(s.db); + expect(partitionCount(s, "alpha")).toBe(1); + expect(partitionCount(s, "beta")).toBe(0); + }); +}); diff --git a/test/cli.test.ts b/test/cli.test.ts index 8fdbcf50c..4c00979a0 100644 --- a/test/cli.test.ts +++ b/test/cli.test.ts @@ -6,7 +6,7 @@ */ import { describe, test, expect, beforeAll, afterAll, beforeEach } from "vitest"; -import { chmod, copyFile, mkdtemp, rm, writeFile, mkdir } from "fs/promises"; +import { chmod, copyFile, mkdtemp, rm, writeFile, mkdir, rename } from "fs/promises"; import { existsSync, lstatSync, readFileSync, symlinkSync, writeFileSync, unlinkSync } from "fs"; import { tmpdir } from "os"; import { join, dirname } from "path"; @@ -15,6 +15,8 @@ import { spawn } from "child_process"; import { setTimeout as sleep } from "timers/promises"; import { buildEditorUri, termLink, resolveEmbedModelForCli } from "../src/cli/qmd.ts"; import { openDatabase } from "../src/db.ts"; +import { createStore, insertContent, insertDocument } from "../src/store.ts"; +import { VEC_COLLECTION_IDS_TABLE, VEC_ROWS_TABLE, VEC_TABLE } from "../src/vec-layout.ts"; import { DEFAULT_EMBED_MODEL_URI, DEFAULT_GENERATE_MODEL_URI, DEFAULT_RERANK_MODEL_URI } from "../src/llm.ts"; import { setConfigSource } from "../src/collections.ts"; @@ -687,6 +689,24 @@ describe("CLI Add Command", () => { } }); + test("collection add skips files over 10 MB", async () => { + const env = await createIsolatedTestEnv("skip-too-large"); + const collectionDir = join(testDir, `skip-too-large-${testCounter}`); + await mkdir(collectionDir, { recursive: true }); + await writeFile(join(collectionDir, "good.md"), "alpha\n"); + await writeFile(join(collectionDir, "big.md"), "a".repeat(10 * 1024 * 1024 + 1)); + + const { stdout, stderr, exitCode } = await runQmd( + ["collection", "add", collectionDir, "--name", "skip-too-large"], + { dbPath: env.dbPath, configDir: env.configDir, cwd: collectionDir }, + ); + expect(exitCode).toBe(0); + expect(stdout).toContain("Indexed: 1 new"); + expect(stderr).toContain("Skipped file over 10 MB: big.md"); + expect(stderr).toContain("Skipped 1 file(s) over 10 MB"); + expect(stderr).not.toContain("unreadable"); + }); + test("can recreate collection with remove and add", async () => { // First add await runQmd(["collection", "add", "."]); @@ -1247,6 +1267,27 @@ describe("CLI Multi-Get Command", () => { }); describe("CLI Update Command", () => { + test.skipIf(process.platform === "win32" || process.getuid?.() === 0)("reports collection root permission failures and exits non-zero", async () => { + const env = await createIsolatedTestEnv("update-root-permissions"); + const parent = join(testDir, "restricted-root"); + const collectionPath = join(parent, "docs"); + await mkdir(collectionPath, { recursive: true }); + await writeFile(join(collectionPath, "readme.md"), "# Readme\n\nPreserve this indexed document.\n"); + const added = await runQmd(["collection", "add", collectionPath, "--name", "docs"], env); + expect(added.exitCode).toBe(0); + await chmod(parent, 0o000); + try { + const result = await runQmd(["update"], env); + expect(result.exitCode).toBe(1); + expect(result.stderr).toContain("EACCES"); + const search = await runQmd(["search", "Preserve", "--json"], env); + expect(search.exitCode).toBe(0); + expect(search.stdout).toContain("readme.md"); + } finally { + await chmod(parent, 0o700); + } + }, 60000); + let localDbPath: string; beforeEach(async () => { @@ -1351,6 +1392,115 @@ describe("CLI Cleanup Command", () => { }); }); +describe("qmd update stale vector rows", () => { + test("a document whose file disappeared loses its vector rows at the end of update", async () => { + const env = await createIsolatedTestEnv("stale-vector-rows"); + const add = await runQmd(["collection", "add", fixturesDir, "--name", "fixtures"], { dbPath: env.dbPath, configDir: env.configDir }); + expect(add.exitCode).toBe(0); + + const store = createStore(env.dbPath); + try { + const now = new Date().toISOString(); + insertContent(store.db, "ghosthash", "# Ghost\n\ngone", now); + insertDocument(store.db, "fixtures", "ghost.md", "Ghost", "ghosthash", now, now); + store.ensureVecTable(3); + store.insertEmbedding("ghosthash", 0, 0, new Float32Array([1, 2, 3]), "test", now); + expect(store.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_TABLE}`).get()).toEqual({ c: 1 }); + } finally { + store.close(); + } + + const update = await runQmd(["update"], { dbPath: env.dbPath, configDir: env.configDir }); + expect(update.exitCode).toBe(0); + expect(update.stdout).toContain("Removed 1 stale vector row(s)"); + + const db = openDatabase(env.dbPath); + try { + expect(db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_ROWS_TABLE}`).get()).toEqual({ c: 0 }); + expect(db.prepare(`SELECT COUNT(*) AS c FROM content_vectors WHERE hash = 'ghosthash'`).get()).toEqual({ c: 0 }); + expect(db.prepare(`SELECT active FROM documents WHERE path = 'ghost.md'`).get()).toEqual({ active: 0 }); + } finally { + db.close(); + } + }); + + test("update leaves a collection whose directory is missing with its documents and vectors", async () => { + const env = await createIsolatedTestEnv("missing-root"); + const root = join(testDir, `missing-root-${testCounter}`); + await mkdir(root, { recursive: true }); + await writeFile(join(root, "kept.md"), "# Kept\n\nstill here\n"); + const add = await runQmd(["collection", "add", root, "--name", "mounted"], { dbPath: env.dbPath, configDir: env.configDir }); + expect(add.exitCode).toBe(0); + + const store = createStore(env.dbPath); + try { + const { hash } = store.db.prepare(`SELECT hash FROM documents WHERE path = 'kept.md'`).get() as { hash: string }; + store.ensureVecTable(3); + store.insertEmbedding(hash, 0, 0, new Float32Array([1, 2, 3]), "test", new Date().toISOString()); + } finally { + store.close(); + } + await rename(root, `${root}-unmounted`); + + const update = await runQmd(["update"], { dbPath: env.dbPath, configDir: env.configDir }); + expect(update.exitCode).toBe(0); + + const db = openDatabase(env.dbPath); + try { + expect(db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_ROWS_TABLE}`).get()).toEqual({ c: 1 }); + expect(db.prepare(`SELECT active FROM documents WHERE path = 'kept.md'`).get()).toEqual({ active: 1 }); + } finally { + db.close(); + } + expect(update.stderr).toContain("Collection root not found"); + }); + + test("a collection that gains an already-embedded document receives its vector rows at the end of update", async () => { + const env = await createIsolatedTestEnv("copied-vector-rows"); + const firstDir = join(testDir, `copied-vector-rows-first-${testCounter}`); + const secondDir = join(testDir, `copied-vector-rows-second-${testCounter}`); + await mkdir(firstDir, { recursive: true }); + await mkdir(secondDir, { recursive: true }); + await writeFile(join(firstDir, "shared.md"), "# Shared\n\nembedded once, then joins a second collection\n"); + await writeFile(join(secondDir, "other.md"), "# Other\n\nonly in the second collection\n"); + + const addFirst = await runQmd(["collection", "add", firstDir, "--name", "first"], { dbPath: env.dbPath, configDir: env.configDir }); + expect(addFirst.exitCode).toBe(0); + const addSecond = await runQmd(["collection", "add", secondDir, "--name", "second"], { dbPath: env.dbPath, configDir: env.configDir }); + expect(addSecond.exitCode).toBe(0); + + const store = createStore(env.dbPath); + let sharedHash: string; + try { + const now = new Date().toISOString(); + sharedHash = (store.db.prepare(`SELECT hash FROM documents WHERE collection = 'first' AND path = 'shared.md' AND active = 1`).get() as { hash: string }).hash; + store.ensureVecTable(3); + store.insertEmbedding(sharedHash, 0, 0, new Float32Array([1, 2, 3]), "test", now); + } finally { + store.close(); + } + + await copyFile(join(firstDir, "shared.md"), join(secondDir, "shared.md")); + + const update = await runQmd(["update"], { dbPath: env.dbPath, configDir: env.configDir }); + expect(update.exitCode).toBe(0); + expect(update.stdout).toContain("Copied 1 vector(s) into collections that gained already-embedded documents"); + + const db = openDatabase(env.dbPath); + try { + const partitions = db.prepare(` + SELECT ci.name AS collection FROM ${VEC_ROWS_TABLE} vr + JOIN ${VEC_COLLECTION_IDS_TABLE} ci ON ci.id = vr.collection_id + WHERE vr.hash = ? + ORDER BY ci.name + `).all(sharedHash); + expect(partitions).toEqual([{ collection: "first" }, { collection: "second" }]); + } finally { + db.close(); + } + }); +}); + describe("orphaned embedding vectors (#768)", () => { let localDbPath: string; let localConfigDir: string; @@ -1411,11 +1561,21 @@ describe("orphaned embedding vectors (#768)", () => { expect(stdout).toContain("qmd cleanup"); }); - test("update hints when orphan ratio exceeds 10%", async () => { + test("update leaves orphaned chunks for cleanup when there is no vector table", async () => { const { stdout, exitCode } = await runQmd(["update"], { dbPath: localDbPath, configDir: localConfigDir }); expect(exitCode).toBe(0); - expect(stdout).toContain("3 orphaned embedding chunks (75% of vectors)"); - expect(stdout).toContain("run 'qmd cleanup' to reclaim space"); + expect(stdout).not.toContain("orphaned embedding chunks"); + expect(stdout).not.toContain("stale vector row"); + const status = await runQmd(["status"], { dbPath: localDbPath, configDir: localConfigDir }); + expect(status.stdout).toMatch(/Orphaned:\s+3 embedding chunks/); + + const db = openDatabase(localDbPath); + const vectorlessLiveChunks = (db.prepare(` + SELECT COUNT(*) as c FROM content_vectors cv + WHERE EXISTS (SELECT 1 FROM documents d WHERE d.hash = cv.hash AND d.active = 1) + `).get() as { c: number }).c; + db.close(); + expect(vectorlessLiveChunks).toBe(0); }); test("cleanup --dry-run reports what would be removed without deleting", async () => { @@ -1436,7 +1596,7 @@ describe("orphaned embedding vectors (#768)", () => { const cache = (db.prepare(`SELECT COUNT(*) as c FROM llm_cache`).get() as { c: number }).c; const inactive = (db.prepare(`SELECT COUNT(*) as c FROM documents WHERE active = 0`).get() as { c: number }).c; db.close(); - expect(vectors).toBe(4); + expect(vectors).toBe(3); expect(cache).toBe(2); expect(inactive).toBe(1); }); diff --git a/test/doctor-vector-sampling.test.ts b/test/doctor-vector-sampling.test.ts new file mode 100644 index 000000000..f912f350d --- /dev/null +++ b/test/doctor-vector-sampling.test.ts @@ -0,0 +1,126 @@ +import { afterEach, beforeEach, describe, expect, test } from "vitest"; +import { spawnSync } from "node:child_process"; +import { dirname, join } from "node:path"; +import { fileURLToPath } from "node:url"; +import { isBun } from "../src/db.js"; +import { chunkDocumentByTokens, createStore, extractTitle, formatDocForEmbedding, getEmbeddingVectorSamples, type Store } from "../src/store.js"; +import { checkEmbeddingVectorSamples } from "../src/cli/qmd.js"; +import { setDefaultLlamaCpp, LlamaCpp } from "../src/llm.js"; + +const projectRoot = join(dirname(fileURLToPath(import.meta.url)), ".."); +let store: Store; + +beforeEach(() => { store = createStore(":memory:"); }); +afterEach(() => { setDefaultLlamaCpp(null); store.close(); }); + +function addDocument(hash: string, path: string, active: boolean = true): void { + store.insertContent(hash, `Body for ${hash}`, "2026-01-01"); + store.insertDocument("test", path, hash, hash, "2026-01-01", "2026-01-01"); + if (!active) store.deactivateDocument("test", path); +} + +function addChunk(hash: string, seq: number, model: string = "model", fingerprint: string = "current", pos: number = 0): void { + store.db.prepare(` + INSERT INTO content_vectors (hash, seq, model, embed_fingerprint, pos, embedded_at) + VALUES (?, ?, ?, ?, ?, '2026-01-01') + `).run(hash, seq, model, fingerprint, pos); +} + +describe("doctor vector sampling", () => { + async function addStoredPassage(seq: number, positionExists: boolean = true, vectorMatches: boolean = true): Promise { + let expectedText = ""; + class PassageLlm extends LlamaCpp { + async tokenize(text: string) { return new Array(Math.ceil(text.length / 16)).fill(1); } + async embed(text: string) { + return { embedding: text === expectedText ? [1, 0] : [0, 1], model: "model" }; + } + } + setDefaultLlamaCpp(new PassageLlm()); + const body = "First passage. ".repeat(300) + "\n\nSecond passage. ".repeat(300); + const chunks = await chunkDocumentByTokens(body); + const passage = chunks[1]!; + expect(passage.pos).toBeGreaterThan(0); + expectedText = formatDocForEmbedding(passage.text, extractTitle(body, "sample.md"), "model"); + store.insertContent("sample", body, "2026-01-01"); + store.insertDocument("test", "sample.md", "Sample", "sample", "2026-01-01", "2026-01-01"); + store.ensureVecTable(2); + store.insertEmbedding("sample", seq, positionExists ? passage.pos : 1, + new Float32Array(vectorMatches ? [1, 0] : [0, 1]), "model", "2026-01-01", 1, "current"); + } + + test("checks the stored passage when earlier chunks shift its sequence number", async () => { + await addStoredPassage(0); + expect((await checkEmbeddingVectorSamples(store.db, "model", "current")).ok).toBe(true); + }); + + test("checks the stored passage when its old sequence is beyond today's chunk count", async () => { + await addStoredPassage(100); + expect((await checkEmbeddingVectorSamples(store.db, "model", "current")).ok).toBe(true); + }); + + test("still rejects a changed vector at the matching position", async () => { + await addStoredPassage(1, true, false); + const result = await checkEmbeddingVectorSamples(store.db, "model", "current"); + expect(result.ok).toBe(false); + expect(result.details).toContain("stored vector distance"); + }); + + test("does not substitute the sequence number when the stored position disappeared", async () => { + await addStoredPassage(1, false); + const result = await checkEmbeddingVectorSamples(store.db, "model", "current"); + expect(result.ok).toBe(false); + expect(result.details).toContain("chunk no longer exists"); + }); + + test("samples each eligible chunk once and uses the first active path", () => { + addDocument("shared", "z.md"); + addDocument("shared", "b.md"); + addDocument("shared", "a.md", false); + addChunk("shared", 0); + addChunk("shared", 1, "model", "current", 1234); + addDocument("other", "other.md"); + addChunk("other", 0); + addDocument("inactive", "inactive.md", false); + addChunk("inactive", 0); + addDocument("stale", "stale.md"); + addChunk("stale", 0, "model", "old"); + addDocument("different-model", "different.md"); + addChunk("different-model", 0, "other-model"); + addChunk("orphan", 0); + + const samples = getEmbeddingVectorSamples(store.db, "model", "current", 20); + expect(samples.sort((a, b) => a.hash.localeCompare(b.hash) || a.seq - b.seq)).toEqual([ + { hash: "other", seq: 0, pos: 0, body: "Body for other", path: "other.md" }, + { hash: "shared", seq: 0, pos: 0, body: "Body for shared", path: "b.md" }, + { hash: "shared", seq: 1, pos: 1234, body: "Body for shared", path: "b.md" }, + ]); + expect(getEmbeddingVectorSamples(store.db, "model", "current", 2)).toHaveLength(2); + expect(getEmbeddingVectorSamples(store.db, "model", "current", 0)).toEqual([]); + expect(getEmbeddingVectorSamples(store.db, "missing", "current")).toEqual([]); + }); + + test("samples large documents with duplicate paths within a 32 MiB SQLite budget", () => { + // SQLite's hard heap limit is process-wide and cannot be raised again. + const worker = join(projectRoot, "test", "_helpers", "doctor-vector-sample-worker.ts"); + const args = isBun ? [worker] : [join(projectRoot, "node_modules", "tsx", "dist", "cli.mjs"), worker]; + const result = spawnSync(process.execPath, args, { encoding: "utf8", timeout: 20_000 }); + expect(result.error).toBeUndefined(); + expect(result.stderr).toBe(""); + expect(result.status).toBe(0); + expect(JSON.parse(result.stdout)).toEqual({ samples: 3, distinctChunks: 3, bodiesComplete: true }); + }); + + test("the doctor check verifies large documents with duplicate paths within a 32 MiB SQLite budget", () => { + // SQLite's hard heap limit is process-wide and cannot be raised again. + const worker = join(projectRoot, "test", "_helpers", "doctor-vector-check-worker.ts"); + const args = isBun ? [worker] : [join(projectRoot, "node_modules", "tsx", "dist", "cli.mjs"), worker]; + const result = spawnSync(process.execPath, args, { encoding: "utf8", timeout: 60_000 }); + expect(result.error).toBeUndefined(); + expect(result.stderr).toBe(""); + expect(result.status).toBe(0); + const out = JSON.parse(result.stdout); + expect(out).toMatchObject({ ok: true, details: "3 sampled chunks reproduce stored vectors" }); + // The budget only binds where SQLite tracks memory; on Bun under Linux it must. + if (isBun && process.platform === "linux") expect(out.limitBinds).toBe(true); + }); +}); diff --git a/test/eval.test.ts b/test/eval.test.ts index d575ff847..1e2dfa4d8 100644 --- a/test/eval.test.ts +++ b/test/eval.test.ts @@ -15,6 +15,7 @@ import { mkdtempSync, rmSync, readFileSync, readdirSync } from "fs"; import { join } from "path"; import { tmpdir } from "os"; import { openDatabase } from "../src/db.js"; +import { VEC_TABLE, hasVectorIndex } from "../src/vec-layout.js"; import type { Database } from "../src/db.js"; import { createHash } from "crypto"; import { fileURLToPath } from "url"; @@ -166,12 +167,8 @@ describe.skipIf(!!process.env.CI)("Vector Search", () => { db = store.db; // Check if embeddings already exist (from previous test run) - const vecTable = db.prepare( - `SELECT name FROM sqlite_master WHERE type='table' AND name='vectors_vec'` - ).get(); - - if (vecTable) { - const count = db.prepare(`SELECT COUNT(*) as cnt FROM vectors_vec`).get() as { cnt: number }; + if (hasVectorIndex(db)) { + const count = db.prepare(`SELECT COUNT(*) as cnt FROM ${VEC_TABLE}`).get() as { cnt: number }; if (count.cnt > 0) { hasEmbeddings = true; return; @@ -277,11 +274,8 @@ describe.skipIf(!!process.env.CI)("Hybrid Search (RRF)", () => { store = createStore(); db = store.db; // Check if vectors exist - const vecTable = db.prepare( - `SELECT name FROM sqlite_master WHERE type='table' AND name='vectors_vec'` - ).get(); - if (vecTable) { - const count = db.prepare(`SELECT COUNT(*) as cnt FROM vectors_vec`).get() as { cnt: number }; + if (hasVectorIndex(db)) { + const count = db.prepare(`SELECT COUNT(*) as cnt FROM ${VEC_TABLE}`).get() as { cnt: number }; hasVectors = count.cnt > 0; } }); diff --git a/test/mcp.test.ts b/test/mcp.test.ts index dc0e4c523..4a3b9bac5 100644 --- a/test/mcp.test.ts +++ b/test/mcp.test.ts @@ -5,7 +5,7 @@ * Uses mocked Ollama responses and a test database. */ -import { describe, test, expect, beforeAll, afterAll, beforeEach, afterEach } from "vitest"; +import { describe, test, expect, vi, beforeAll, afterAll, beforeEach, afterEach } from "vitest"; import { openDatabase, loadSqliteVec } from "../src/db.js"; import type { Database } from "../src/db.js"; import { getDefaultLlamaCpp, disposeDefaultLlamaCpp } from "../src/llm"; @@ -18,6 +18,8 @@ import type { CollectionConfig } from "../src/collections"; import { setConfigIndexName } from "../src/collections"; import { syncConfigToDb } from "../src/store"; import { initializeMetadataSchema } from "../src/metadata-store"; +import { VEC_TABLE, createVectorMetadataTables, vecLayout } from "../src/vec-layout"; +import { migrateVectorLayout } from "../src/store-migrations"; // ============================================================================= // Test Database Setup @@ -128,8 +130,10 @@ function initTestDatabase(db: Database): void { // Document metadata tables — searchFTS/searchVec join them for result metadata initializeMetadataSchema(db); + createVectorMetadataTables(db); } +/** Seeds the legacy vector table, then runs the layout migration on it. */ function seedTestData(db: Database): void { const now = new Date().toISOString(); @@ -192,6 +196,8 @@ function seedTestData(db: Database): void { db.prepare(`INSERT INTO content_vectors (hash, seq, pos, model, embed_fingerprint, embedded_at) VALUES (?, 0, 0, ?, ?, ?)`).run(doc.hash, DEFAULT_EMBED_MODEL, getEmbeddingFingerprint(DEFAULT_EMBED_MODEL), now); db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${doc.hash}_0`, embedding); } + expect(migrateVectorLayout(db, { sqliteVecAvailable: true })).toBe("applied"); + expect(vecLayout(db)).toMatchObject({ kind: "partitioned", dimensions: 768 }); } // ============================================================================= @@ -342,6 +348,7 @@ describe("MCP Server", () => { const emptyDb = openDatabase(":memory:"); initTestDatabase(emptyDb); emptyDb.exec("DROP TABLE IF EXISTS vectors_vec"); + emptyDb.exec(`DROP TABLE IF EXISTS ${VEC_TABLE}`); const results = await searchVec(emptyDb, "test", DEFAULT_EMBED_MODEL, 10); expect(results.length).toBe(0); @@ -1124,6 +1131,55 @@ describe.skipIf(!!process.env.CI)("MCP HTTP Transport", () => { expect(json.result.serverInfo.name).toBe("qmd"); }); + test("POST /mcp initialize serves cached instructions across requests", async () => { + // The legacy initialize path answers either JSON or a single SSE frame + // depending on how the transport negotiates; read both the same way. + const legacyInitialize = async (id: number) => { + const res = await fetch(`${baseUrl}/mcp`, { + method: "POST", + headers: { + "Content-Type": "application/json", + "Accept": "application/json, text/event-stream", + }, + body: JSON.stringify({ + jsonrpc: "2.0", + id, + method: "initialize", + params: { + protocolVersion: "2025-03-26", + capabilities: {}, + clientInfo: { name: "test-client", version: "1.0.0" }, + }, + }), + }); + const text = await res.text(); + const data = text.includes("data:") + ? text.split("\n").find(line => line.startsWith("data:"))!.slice(5).trim() + : text; + return JSON.parse(data) as any; + }; + + const first = await legacyInitialize(10); + const firstInstructions = first.result.instructions; + expect(typeof firstInstructions).toBe("string"); + expect(firstInstructions.length).toBeGreaterThan(0); + + // Change the index under the server. HTTP builds a fresh McpServer per + // request, so without the cache the next initialize rebuilds the + // instructions and reports the new document count; with it, the same + // string is served until the TTL lapses. + const db = openDatabase(httpTestDbPath); + const now = new Date().toISOString(); + db.prepare(`INSERT OR IGNORE INTO content (hash, doc, created_at) VALUES (?, ?, ?)`) + .run("hash-instr-cache", "# Cache probe\nbody", now); + db.prepare(`INSERT INTO documents (collection, path, title, hash, created_at, modified_at, active) VALUES ('docs', ?, ?, ?, ?, ?, 1)`) + .run("instructions-cache-probe.md", "Cache Probe", "hash-instr-cache", now, now); + db.close(); + + const second = await legacyInitialize(11); + expect(second.result.instructions).toBe(firstInstructions); + }); + test("POST /mcp tools/list returns registered tools without a session", async () => { const { status, json, contentType, headers } = await mcpRequest("tools/list"); expect(status).toBe(200); @@ -1131,7 +1187,7 @@ describe.skipIf(!!process.env.CI)("MCP HTTP Transport", () => { expect(headers.get("mcp-session-id")).toBeNull(); const toolNames = json.result.tools.map((t: any) => t.name); - expect(toolNames).toEqual(["query", "get", "multi_get", "status"]); + expect(toolNames).toEqual(["query", "get", "multi_get", "status", "metadata"]); }); test("POST /mcp tools/call query returns results", async () => { @@ -1309,7 +1365,7 @@ describe("MCP HTTP Transport — 2026-07-28 protocol", () => { expect(json.result.ttlMs).toBe(60_000); expect(json.result.cacheScope).toBe("private"); const toolNames = json.result.tools.map((t: { name: string }) => t.name); - expect(toolNames).toEqual(["query", "get", "multi_get", "status"]); + expect(toolNames).toEqual(["query", "get", "multi_get", "status", "metadata"]); const serverInfo = json.result._meta?.["io.modelcontextprotocol/serverInfo"]; expect(serverInfo?.name).toBe("qmd"); }); @@ -1380,6 +1436,106 @@ describe("MCP HTTP Transport — 2026-07-28 protocol", () => { expect(res.status).toBe(405); expect(res.headers.get("mcp-session-id")).toBeNull(); }); + + // Each request gets a fresh McpServer, so these observe the per-store + // instructions cache through server/discover. + async function discoverInstructions(id: number): Promise<{ status: number; instructions?: string }> { + const { status, json } = await postMcp( + { jsonrpc: "2.0", id, method: "server/discover", params: { _meta: mcp2026Meta } }, + { "MCP-Protocol-Version": MCP_2026, "Mcp-Method": "server/discover" }, + ); + return { status, instructions: json.result?.instructions }; + } + + function documentCount(instructions: string | undefined): number { + const match = instructions?.match(/over (\d+) markdown documents/); + if (!match) throw new Error(`no document count in instructions: ${instructions}`); + return Number(match[1]); + } + + function activeDocuments(): number { + const db = openDatabase(dbPath); + try { + const row = db.prepare(`SELECT COUNT(*) AS n FROM documents WHERE active = 1`).get() as { n: number }; + return row.n; + } finally { + db.close(); + } + } + + function addDocument(slug: string): void { + const db = openDatabase(dbPath); + const now = new Date().toISOString(); + db.prepare(`INSERT INTO content (hash, doc, created_at) VALUES (?, ?, ?)`) + .run(`hash-${slug}`, `# ${slug}\nbody`, now); + db.prepare(`INSERT INTO documents (collection, path, title, hash, created_at, modified_at, active) VALUES ('docs', ?, ?, ?, ?, ?, 1)`) + .run(`${slug}.md`, slug, `hash-${slug}`, now, now); + db.close(); + } + + test("server/discover serves cached instructions after the index changes", async () => { + const first = await discoverInstructions(20); + expect(first.status).toBe(200); + addDocument("instructions-cache-hit"); + const second = await discoverInstructions(21); + expect(second.status).toBe(200); + expect(second.instructions).toBe(first.instructions); + }); + + test("server/discover rebuilds instructions once the 60 s cache entry expires", async () => { + const realNow = Date.now.bind(Date); + const first = await discoverInstructions(22); + addDocument("instructions-cache-expiry"); + const cached = await discoverInstructions(23); + expect(cached.instructions).toBe(first.instructions); + const clock = vi.spyOn(Date, "now").mockImplementation(() => realNow() + 61_000); + try { + const rebuilt = await discoverInstructions(24); + expect(rebuilt.status).toBe(200); + expect(documentCount(rebuilt.instructions)).toBe(activeDocuments()); + expect(documentCount(rebuilt.instructions)).toBeGreaterThan(documentCount(first.instructions)); + } finally { + clock.mockRestore(); + } + }); + + test("server/discover retries a failed instructions build instead of caching the failure", async () => { + const realNow = Date.now.bind(Date); + // Past every earlier entry's TTL, so the next discover has to build. + const clock = vi.spyOn(Date, "now").mockImplementation(() => realNow() + 10 * 60_000); + const db = openDatabase(dbPath); + try { + db.exec(`ALTER TABLE store_collections RENAME TO store_collections_offline`); + const failed = await discoverInstructions(25); + expect(failed.instructions).toBeUndefined(); + db.exec(`ALTER TABLE store_collections_offline RENAME TO store_collections`); + + const retried = await discoverInstructions(26); + expect(retried.status).toBe(200); + addDocument("instructions-cache-after-failure"); + const cached = await discoverInstructions(27); + expect(cached.instructions).toBe(retried.instructions); + } finally { + const offline = db.prepare(`SELECT name FROM sqlite_master WHERE name = 'store_collections_offline'`).get(); + if (offline) db.exec(`ALTER TABLE store_collections_offline RENAME TO store_collections`); + db.close(); + clock.mockRestore(); + } + }); + + test("server/discover rebuilds an entry stamped ahead of the clock once the clock steps back", async () => { + const realNow = Date.now.bind(Date); + // Built while the clock read ten minutes ahead, as after a backward step. + const clock = vi.spyOn(Date, "now").mockImplementation(() => realNow() + 20 * 60_000); + const ahead = await discoverInstructions(28); + clock.mockRestore(); + addDocument("instructions-cache-clock-step"); + + const rebuilt = await discoverInstructions(29); + expect(rebuilt.status).toBe(200); + expect(rebuilt.instructions).not.toBe(ahead.instructions); + expect(documentCount(rebuilt.instructions)).toBe(activeDocuments()); + }); }); diff --git a/test/metadata-cli.test.ts b/test/metadata-cli.test.ts index 213f0eede..e99475f35 100644 --- a/test/metadata-cli.test.ts +++ b/test/metadata-cli.test.ts @@ -9,6 +9,7 @@ import { tmpdir } from "node:os"; import { join, dirname } from "node:path"; import { fileURLToPath } from "node:url"; import { spawn } from "node:child_process"; +import { openDatabase } from "../src/db.js"; const thisDir = dirname(fileURLToPath(import.meta.url)); const projectRoot = join(thisDir, ".."); @@ -93,7 +94,7 @@ describe("qmd search --filter", () => { const { stdout, exitCode } = await runQmd([ "search", "cli filter keyword", "--format", "json", - "--filter", '{"key":"status","operator":"eq","value":"published"}', + "--filter", '{"field":"status","operator":"eq","value":"published"}', ]); expect(exitCode).toBe(0); @@ -107,8 +108,8 @@ describe("qmd search --filter", () => { const filter = JSON.stringify({ operator: "and", operands: [ - { key: "topics", operator: "all", value: ["typescript", "programming"] }, - { operator: "not", operand: { key: "status", operator: "eq", value: "draft" } }, + { field: "topics", operator: "all", value: ["typescript", "programming"] }, + { operator: "not", operand: { field: "status", operator: "eq", value: "draft" } }, ], }); const { stdout, exitCode } = await runQmd([ @@ -122,7 +123,7 @@ describe("qmd search --filter", () => { const { stdout, exitCode } = await runQmd([ "search", "cli filter keyword", "--format", "json", - "--filter", '{"key":"status","operator":"eq","value":"missing"}', + "--filter", '{"field":"status","operator":"eq","value":"missing"}', ]); expect(exitCode).toBe(0); expect(JSON.parse(stdout)).toEqual([]); @@ -152,10 +153,397 @@ describe("qmd search --filter", () => { test("rejects valid JSON with an invalid filter AST", async () => { const { stderr, exitCode } = await runQmd([ - "search", "cli filter keyword", "--filter", '{"key":"status","operator":"equal","value":"x"}', + "search", "cli filter keyword", "--filter", '{"field":"status","operator":"equal","value":"x"}', ]); expect(exitCode).toBe(1); expect(stderr).toMatch(/Invalid metadata filter at \$/); expect(stderr).toMatch(/unknown operator 'equal'/); }, 30000); }); + +describe("qmd collection metadata", () => { + beforeAll(async () => { + await writeFile(join(fixturesDir, "discovery.md"), [ + "---", + "qmd:", + " metadata:", + " status: published", + " topics: [typescript, sqlite, search, architecture, embeddings, mcp, cli, testing, agents, indexing, chunking, reranking]", + " priority: 5", + " owner: docs-team", + "---", + "", + "# Discovery doc", + "", + "cli filter keyword body", + "", + ].join("\n")); + await writeFile(join(fixturesDir, "priority.md"), [ + "---", + "qmd:", + " metadata:", + " status: draft", + " priority: 1", + " reviewed: false", + " reviewers: [docs-team, security-team]", + "---", + "", + "# Priority doc", + "", + "cli filter keyword body", + "", + ].join("\n")); + // Nothing enforces a type across documents, so one file in the same + // collection may spell priority as a label where the others use numbers. + await writeFile(join(fixturesDir, "conflict.md"), [ + "---", + "qmd:", + " metadata:", + " status: archived", + " priority: high", + "---", + "", + "# Conflict doc", + "", + "cli filter keyword body", + "", + ].join("\n")); + const updateResult = await runQmd(["update"]); + expect(updateResult.exitCode).toBe(0); + }, 60000); + + test("lists every key with coverage, type, and a value window", async () => { + const { stdout, exitCode } = await runQmd(["collection", "metadata", "notes"]); + expect(exitCode).toBe(0); + + // Keys arrive in coverage order; the denominator is every active document. + expect(stdout).toMatch(/^status {2}string {2}5 of 6 documents {2}3 distinct\n {2}draft {6}2\n {2}published {2}2\n {2}archived {3}1\n/); + expect(stdout).toContain("topics string[] 2 of 6 documents 13 distinct"); + expect(stdout).toContain("reviewed boolean 1 of 6 documents\n false 1"); + expect(stdout).toContain("reviewers string[] 1 of 6 documents 2 distinct\n docs-team 1\n security-team 1"); + }, 30000); + + test("splits a key whose documents disagree on type, even within one collection", async () => { + const { stdout, exitCode } = await runQmd(["collection", "metadata", "notes", "--match", '{"field":"key","operator":"eq","value":"priority"}']); + expect(exitCode).toBe(0); + + // One row per type with its own document count. The collection column + // is omitted since the view covers a single collection. + expect(stdout).toBe([ + "priority number | string 3 of 6 documents", + " number 2 docs min 1 median 3 max 5", + " string 1 doc high (1)", + "", + ].join("\n")); + }, 30000); + + test("truncates value lists with an explicit remainder and escape hatch", async () => { + const { stdout, exitCode } = await runQmd(["collection", "metadata", "notes", "--match", '{"field":"key","operator":"eq","value":"topics"}']); + expect(exitCode).toBe(0); + + const valueLines = stdout.split("\n").filter(line => /^ {2}\S/.test(line)); + expect(valueLines).toHaveLength(10); + expect(valueLines[0]).toBe(" typescript 2"); + expect(stdout).toContain("3 more values, use --value-limit , --value-offset , or --all-values"); + expect(stdout).not.toContain("status"); + }, 30000); + + test("--value-limit, --value-offset, and --all-values size and page the value window", async () => { + const topics = ["--match", '{"field":"key","operator":"eq","value":"topics"}']; + + const limited = await runQmd(["collection", "metadata", "notes", ...topics, "--value-limit", "2"]); + expect(limited.stdout).toContain("11 more values, use --value-limit , --value-offset , or --all-values"); + + const paged = await runQmd(["collection", "metadata", "notes", ...topics, "--value-limit", "2", "--value-offset", "11"]); + expect(paged.stdout.split("\n").filter(line => /^ {2}\S/.test(line))).toHaveLength(2); + expect(paged.stdout).not.toContain("more values"); + + const all = await runQmd(["collection", "metadata", "notes", ...topics, "--all-values"]); + expect(all.stdout).not.toContain("more values"); + expect(all.stdout.split("\n").filter(line => /^ {2}\S/.test(line))).toHaveLength(13); + }, 30000); + + test("--key-limit, --key-offset, and --all-keys size and page the key window", async () => { + const keyHeaders = (stdout: string) => stdout.split("\n").filter(line => /^\S.* of 6 documents/.test(line)).map(line => line.split(" ")[0]); + + const everything = await runQmd(["collection", "metadata", "notes"]); + expect(keyHeaders(everything.stdout)).toHaveLength(6); + expect(everything.stdout).not.toContain("more keys"); + + const limited = await runQmd(["collection", "metadata", "notes", "--key-limit", "2"]); + expect(keyHeaders(limited.stdout)).toEqual(["status", "priority"]); + expect(limited.stdout).toContain("4 more keys, use --key-limit , --key-offset , or --all-keys"); + + const paged = await runQmd(["collection", "metadata", "notes", "--key-limit", "2", "--key-offset", "2"]); + expect(keyHeaders(paged.stdout)).toEqual(["topics", "owner"]); + expect(paged.stdout).toContain("2 more keys, use --key-limit , --key-offset , or --all-keys"); + + const lastPage = await runQmd(["collection", "metadata", "notes", "--key-limit", "2", "--key-offset", "4"]); + expect(keyHeaders(lastPage.stdout)).toEqual(["reviewed", "reviewers"]); + expect(lastPage.stdout).not.toContain("more keys"); + + const pastTheEnd = await runQmd(["collection", "metadata", "notes", "--key-offset", "6"]); + expect(pastTheEnd.exitCode).toBe(0); + expect(pastTheEnd.stdout).toContain("No keys at --key-offset 6, 6 keys in total."); + + const all = await runQmd(["collection", "metadata", "notes", "--all-keys", "--key-limit", "1"]); + expect(keyHeaders(all.stdout)).toHaveLength(6); + }, 60000); + + test("a match on the value field is a reverse lookup across keys", async () => { + const { stdout, exitCode } = await runQmd(["collection", "metadata", "notes", "--match", '{"field":"value","operator":"eq","value":"docs-team"}']); + expect(exitCode).toBe(0); + + expect(stdout).toBe([ + "owner string 1 of 6 documents 1 distinct", + " docs-team 1", + "", + "reviewers string[] 1 of 6 documents 1 distinct", + " docs-team 1", + "", + ].join("\n")); + }, 30000); + + test("--filter counts only matching documents and says so in the header", async () => { + const { stdout, exitCode } = await runQmd([ + "collection", "metadata", "notes", "--match", '{"field":"key","operator":"eq","value":"priority"}', + "--filter", '{"field":"status","operator":"eq","value":"published"}', + ]); + expect(exitCode).toBe(0); + // The filter line reports the population, and every key's coverage is + // measured against it. Two documents are published and one of them has + // a priority, so the type split above collapses to a flat number view. + expect(stdout).toBe([ + "filter: 2 of 6 documents", + "", + "priority number 1 of 2 documents 1 distinct", + " min 5 median 5 max 5", + " 5 (1)", + "", + ].join("\n")); + }, 30000); + + test("--sort value and --min-count reshape the window", async () => { + const sorted = await runQmd(["collection", "metadata", "notes", "--match", '{"field":"key","operator":"eq","value":"topics"}', "--sort", "value", "--value-limit", "2"]); + expect(sorted.stdout).toContain(" agents 1\n architecture 1\n"); + + const common = await runQmd(["collection", "metadata", "notes", "--match", '{"field":"key","operator":"eq","value":"topics"}', "--min-count", "2"]); + expect(common.stdout).toContain("topics string[] 2 of 6 documents 1 distinct\n typescript 2\n"); + expect(common.stdout).not.toContain("more values"); + }, 30000); + + test("omitting the collection covers the default collections", async () => { + const { stdout, exitCode } = await runQmd(["collection", "metadata"]); + expect(exitCode).toBe(0); + expect(stdout).toContain("status string 5 of 6 documents 3 distinct"); + }, 30000); + + test("a composed match selects entries across the key and value fields", async () => { + const { stdout, exitCode } = await runQmd([ + "collection", "metadata", "notes", + "--match", JSON.stringify({ + operator: "and", + operands: [ + { field: "key", operator: "eq", value: "topics" }, + { operator: "or", operands: [ + { field: "value", operator: "prefix", value: "sql" }, + { field: "value", operator: "contains", value: "ARCH", caseInsensitive: true }, + ] }, + ], + }), + ]); + expect(exitCode).toBe(0); + expect(stdout).toBe([ + "topics string[] 1 of 6 documents 3 distinct", + " architecture 1", + " search 1", + " sqlite 1", + "", + ].join("\n")); + }, 30000); + + test("a type match reports one side of a split key", async () => { + const { stdout, exitCode } = await runQmd([ + "collection", "metadata", "notes", + "--match", '{"field":"value","operator":"type","value":"number"}', + ]); + expect(exitCode).toBe(0); + expect(stdout).toBe([ + "priority number 2 of 6 documents 2 distinct", + " min 1 median 3 max 5", + " 1 (1) 5 (1)", + "", + ].join("\n")); + }, 30000); + + test("reports when nothing matches", async () => { + const { stdout, exitCode } = await runQmd([ + "collection", "metadata", "notes", "--match", '{"field":"key","operator":"prefix","value":"missing-"}', + ]); + expect(exitCode).toBe(0); + expect(stdout).toContain("No metadata matches"); + }, 30000); + + test("keeps the filter line when the filtered documents declare no metadata", async () => { + // plain.md is the one document without a status, and it has no metadata at all. + const { stdout, exitCode } = await runQmd([ + "collection", "metadata", "notes", "--filter", '{"field":"status","operator":"exists","value":false}', + ]); + expect(exitCode).toBe(0); + expect(stdout).toBe([ + "filter: 1 of 6 documents", + "", + "No metadata matches. Run 'qmd collection metadata' without --match or --filter to see which keys exist.", + "", + ].join("\n")); + }, 30000); + + test("exits on an unknown collection", async () => { + const { stderr, exitCode } = await runQmd(["collection", "metadata", "missing"]); + expect(exitCode).toBe(1); + expect(stderr).toContain("Collection not found: missing"); + }, 30000); + + test("exits on an invalid filter", async () => { + const { stderr, exitCode } = await runQmd([ + "collection", "metadata", "notes", "--filter", '{"field":"status","operator":"equal","value":"x"}', + ]); + expect(exitCode).toBe(1); + expect(stderr).toMatch(/Invalid metadata filter at \$/); + }, 30000); + + test("exits on an invalid match, naming the match", async () => { + const badJson = await runQmd(["collection", "metadata", "notes", "--match", "{nope"]); + expect(badJson.exitCode).toBe(1); + expect(badJson.stderr).toMatch(/Invalid --match JSON/); + expect(badJson.stderr).toContain(`Example: --match '{"field":"key","operator":"eq","value":"topics"}'`); + + const badField = await runQmd([ + "collection", "metadata", "notes", "--match", '{"field":"topics","operator":"eq","value":"x"}', + ]); + expect(badField.exitCode).toBe(1); + expect(badField.stderr).toMatch(/Invalid metadata match at \$: 'topics' is not a field of a metadata entry, expected 'key' or 'value'/); + + const badOperator = await runQmd([ + "collection", "metadata", "notes", "--match", '{"field":"value","operator":"exists","value":true}', + ]); + expect(badOperator.exitCode).toBe(1); + expect(badOperator.stderr).toMatch(/'exists' has no meaning for a single metadata entry/); + }, 30000); + + test("exits on invalid window, --min-count, and --sort values", async () => { + const badLimit = await runQmd(["collection", "metadata", "notes", "--value-limit", "0"]); + expect(badLimit.exitCode).toBe(1); + expect(badLimit.stderr).toContain("Invalid --value-limit value: 0"); + + const badKeyLimit = await runQmd(["collection", "metadata", "notes", "--key-limit", "x"]); + expect(badKeyLimit.exitCode).toBe(1); + expect(badKeyLimit.stderr).toContain("Invalid --key-limit value: x"); + + const badOffset = await runQmd(["collection", "metadata", "notes", "--key-offset=-1"]); + expect(badOffset.exitCode).toBe(1); + expect(badOffset.stderr).toContain("Invalid --key-offset value: -1"); + expect(badOffset.stderr).toContain("--key-offset must be a non-negative safe integer"); + + const badMinCount = await runQmd(["collection", "metadata", "notes", "--min-count", "x"]); + expect(badMinCount.exitCode).toBe(1); + expect(badMinCount.stderr).toContain("Invalid --min-count value: x"); + + const badSort = await runQmd(["collection", "metadata", "notes", "--sort", "size"]); + expect(badSort.exitCode).toBe(1); + expect(badSort.stderr).toContain("Invalid --sort value: size"); + }, 30000); + + test("rejects unsafe integers with usage, preserving the input rather than its rounded value", async () => { + for (const flag of ["--key-limit", "--key-offset", "--value-limit", "--value-offset", "--min-count"]) { + const result = await runQmd(["collection", "metadata", "notes", flag, "9007199254740993"]); + expect(result.exitCode).toBe(1); + expect(result.stdout).toBe(""); + expect(result.stderr).toContain(`Invalid ${flag} value: 9007199254740993`); + expect(result.stderr).toContain("safe integer"); + expect(result.stderr).not.toContain("MetadataOptionError"); + expect(result.stderr).not.toContain("at collectionMetadata"); + } + }, 60000); + + test("rejects unsupported formats rather than silently returning text", async () => { + for (const flags of [["--json"], ["--format", "json"], ["--csv"], ["--md"], ["--xml"], ["--files"]]) { + const result = await runQmd(["collection", "metadata", "notes", ...flags]); + expect(result.exitCode).toBe(1); + expect(result.stdout).toBe(""); + expect(result.stderr).toContain("is not supported by 'qmd collection metadata'"); + expect(result.stderr).toContain("Use the SDK, MCP metadata tool, or POST /metadata for structured output"); + } + const result = await runQmd(["collection", "metadata", "notes", "--format", "cli"]); + expect(result.exitCode).toBe(0); + expect(result.stdout).toContain("topics string[]"); + }, 60000); + + test("points the search window flags at the discovery ones", async () => { + const n = await runQmd(["collection", "metadata", "notes", "-n", "3"]); + expect(n.exitCode).toBe(1); + expect(n.stderr).toContain("-n is not an option of 'qmd collection metadata'"); + expect(n.stderr).toContain("Use --value-limit for values per key, or --key-limit for keys"); + + const all = await runQmd(["collection", "metadata", "notes", "--all"]); + expect(all.exitCode).toBe(1); + expect(all.stderr).toContain("--all is not an option of 'qmd collection metadata'"); + expect(all.stderr).toContain("Use --all-values, --all-keys, or both"); + }, 30000); +}); + +describe("metadata in collection list, show, and status", () => { + test("collection list names the top keys and counts the rest", async () => { + const { stdout, exitCode } = await runQmd(["collection", "list"]); + expect(exitCode).toBe(0); + expect(stdout).toContain(" Metadata: status, priority, topics, owner, reviewed, +1 more\n"); + }, 30000); + + test("collection show details the top keys and points at the drill-down", async () => { + const { stdout, exitCode } = await runQmd(["collection", "show", "notes"]); + expect(exitCode).toBe(0); + + const metadataSection = stdout.slice(stdout.indexOf(" Metadata:")); + expect(metadataSection).toBe([ + " Metadata: 6 keys, 5 of 6 documents", + " status string 5 docs 3 distinct draft (2), published (2), archived (1)", + ` priority number | string 3 docs 3 distinct types disagree, see: qmd collection metadata notes --match '{"field":"key","operator":"eq","value":"priority"}'`, + " topics string[] 2 docs 13 distinct typescript (2), agents (1), architecture (1), ...", + " owner string 1 doc 1 distinct", + " reviewed boolean 1 doc 1 distinct false 1", + " 1 more key, see 'qmd collection metadata notes'", + "", + ].join("\n")); + }, 30000); + + test("status summarizes metadata and points at the drill-down", async () => { + const { stdout, exitCode } = await runQmd(["status"]); + expect(exitCode).toBe(0); + expect(stdout).toContain(" Metadata: 6 keys across 5 files (explore with 'qmd collection metadata')\n"); + }, 30000); + + test("warns about pending extraction only within the collections the command reads", async () => { + const otherDir = join(testDir, "other"); + await mkdir(otherDir, { recursive: true }); + await writeFile(join(otherDir, "aged.md"), "---\nqmd:\n metadata:\n status: draft\n---\n\n# Aged\n"); + expect((await runQmd(["collection", "add", otherDir, "--name", "other"])).exitCode).toBe(0); + + const ageOther = "UPDATE document_metadata SET extraction_version = 0 WHERE document_id IN (SELECT id FROM documents WHERE collection = 'other')"; + const db = openDatabase(dbPath); + db.prepare(ageOther).run(); + db.close(); + + try { + const notes = await runQmd(["collection", "metadata", "notes"]); + expect(notes.exitCode).toBe(0); + expect(notes.stderr).not.toContain("lack current metadata extraction"); + + const other = await runQmd(["collection", "metadata", "other"]); + expect(other.exitCode).toBe(0); + expect(other.stderr).toContain("Warning: 1 document(s) lack current metadata extraction"); + } finally { + // Leave the index as the other tests expect it. + expect((await runQmd(["collection", "remove", "other"])).exitCode).toBe(0); + } + }, 60000); +}); diff --git a/test/metadata-discovery.test.ts b/test/metadata-discovery.test.ts new file mode 100644 index 000000000..14234222e --- /dev/null +++ b/test/metadata-discovery.test.ts @@ -0,0 +1,991 @@ +/** + * metadata-discovery.test.ts - Store-level metadata discovery: key summaries, + * per-type value windows, the match over metadata entries, and scope and gate + * agreement with filtered search. + */ + +import { describe, test, expect, beforeAll, afterAll, beforeEach, afterEach } from "vitest"; +import { mkdtemp, rm } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { + createStore, + insertContent, + insertDocument, + hashContent, + searchFTS, + getStatus, + getStatusSummary, + type Store, +} from "../src/store.js"; +import { + countDocumentsPendingMetadata, + countDocumentsWithMetadata, + countMetadataKeys, + listMetadata, + listMetadataCollectionSummaries, + METADATA_SQL_BINDING_BUDGET, + MetadataBindingBudgetError, + MetadataOptionError, + replaceDocumentMetadata, + type ListMetadataOptions, + type MetadataKeySummary, +} from "../src/metadata-store.js"; +import type { Database, SQLiteValue } from "../src/db.js"; +import { compileMetadataFilter, compileMetadataMatch, MetadataFilterError, type MetadataFilter, type MetadataMatch } from "../src/metadata-filter.js"; +import { METADATA_EXTRACTION_VERSION, type DocumentMetadata } from "../src/metadata.js"; + +let testDir: string; +let store: Store; + +beforeAll(async () => { + testDir = await mkdtemp(join(tmpdir(), "qmd-metadata-discovery-")); +}); + +afterAll(async () => { + await rm(testDir, { recursive: true, force: true }); +}); + +beforeEach(() => { + const dbPath = join(testDir, `test-${Date.now()}-${Math.random().toString(36).slice(2)}.sqlite`); + store = createStore(dbPath); +}); + +afterEach(() => { + store.close(); +}); + +let documentCounter = 0; + +/** Insert an active document with extracted metadata. Every body contains "doc" so FTS can reach it. */ +async function insertMetadataDoc(collection: string, metadata: DocumentMetadata): Promise { + documentCounter += 1; + const path = `doc-${documentCounter}.md`; + const content = `# doc ${documentCounter}\n\nbody of doc ${documentCounter}\n`; + const now = new Date().toISOString(); + const hash = await hashContent(content); + insertContent(store.db, hash, content, now); + const documentId = insertDocument(store.db, collection, path, path, hash, now, now); + replaceDocumentMetadata(store.db, documentId, { metadata, extractionVersion: METADATA_EXTRACTION_VERSION }); + return documentId; +} + +/** The store's database with every prepared statement's SQL reported before it runs. */ +function observeStatements(db: Database, onPrepare: (sql: string) => void, onAll?: (sql: string, params: SQLiteValue[]) => void): Database { + return new Proxy(db, { + get(target, property) { + if (property === "prepare") { + return (sql: string) => { + onPrepare(sql); + const statement = target.prepare(sql); + if (!onAll) return statement; + + return new Proxy(statement, { + get(prepared, member) { + if (member === "all") return (...params: SQLiteValue[]) => { + onAll(sql, params); + return prepared.all(...params); + }; + const value = prepared[member as keyof typeof prepared]; + return typeof value === "function" ? value.bind(prepared) : value; + }, + }); + }; + } + // Bun's Database keeps private fields, so its methods must run against the real instance. + const member = target[property as keyof Database]; + return typeof member === "function" ? member.bind(target) : member; + }, + }); +} + +function planOf(sql: string): string[] { + const placeholders = Array.from(sql.matchAll(/\?/g), () => null); + const rows = store.db.prepare(`EXPLAIN QUERY PLAN ${sql}`).all(...placeholders) as { detail: string }[]; + return rows.map(row => row.detail); +} + +function summaryOf(keys: MetadataKeySummary[], key: string): MetadataKeySummary { + const summary = keys.find(candidate => candidate.key === key); + if (!summary) throw new Error(`key ${key} missing from ${keys.map(candidate => candidate.key).join(", ")}`); + return summary; +} + +function valuesOf(keys: MetadataKeySummary[], key: string): [string | number | boolean, number][] { + return summaryOf(keys, key).types[0]!.values.map(count => [count.value, count.documents]); +} + +describe("listMetadata counting", () => { + test("counts documents, not values, and reports coverage per key", async () => { + await insertMetadataDoc("notes", { topics: ["a", "b"], status: "draft" }); + await insertMetadataDoc("notes", { topics: ["a"], status: "published" }); + await insertMetadataDoc("notes", { status: "published" }); + + const result = listMetadata(store.db); + + expect(result.documents).toBe(3); + expect(result.filteredDocuments).toBeUndefined(); + expect(result.keys.map(summary => summary.key)).toEqual(["status", "topics"]); + + const topics = summaryOf(result.keys, "topics"); + expect(topics.documents).toBe(2); + expect(topics.types).toHaveLength(1); + expect(topics.types[0]!.multiValued).toBe(true); + expect(topics.types[0]!.distinctValues).toBe(2); + expect(valuesOf(result.keys, "topics")).toEqual([["a", 2], ["b", 1]]); + + const status = summaryOf(result.keys, "status"); + expect(status.documents).toBe(3); + expect(status.types[0]!.multiValued).toBe(false); + expect(valuesOf(result.keys, "status")).toEqual([["published", 2], ["draft", 1]]); + }); + + test("returns an empty key list when nothing has metadata", async () => { + await insertMetadataDoc("notes", {}); + + const result = listMetadata(store.db); + + expect(result).toEqual({ documents: 1, totalKeys: 0, keys: [], remainingKeys: 0 }); + }); + + test("applies the extraction gate filtered search applies", async () => { + await insertMetadataDoc("notes", { status: "visible" }); + const pendingId = await insertMetadataDoc("notes", { status: "pending" }); + const erroredId = await insertMetadataDoc("notes", { status: "errored" }); + const staleId = await insertMetadataDoc("notes", { status: "stale" }); + const inactiveId = await insertMetadataDoc("notes", { status: "inactive" }); + + store.db.prepare(`DELETE FROM document_metadata WHERE document_id = ?`).run(pendingId); + store.db.prepare(`UPDATE document_metadata SET extraction_error = 'boom' WHERE document_id = ?`).run(erroredId); + store.db.prepare(`UPDATE document_metadata SET extraction_version = ? WHERE document_id = ?`).run(METADATA_EXTRACTION_VERSION - 1, staleId); + store.db.prepare(`UPDATE documents SET active = 0 WHERE id = ?`).run(inactiveId); + + const result = listMetadata(store.db); + + // The denominator counts active documents whether or not they are extracted. + expect(result.documents).toBe(4); + expect(valuesOf(result.keys, "status")).toEqual([["visible", 1]]); + }); +}); + +describe("listMetadata scope and filter", () => { + beforeEach(async () => { + await insertMetadataDoc("notes", { status: "published", priority: 3 }); + await insertMetadataDoc("notes", { status: "draft", priority: 1 }); + await insertMetadataDoc("work", { status: "published", priority: 5 }); + await insertMetadataDoc("work", { status: "archived" }); + }); + + test("undefined collection means every collection", () => { + const result = listMetadata(store.db); + + expect(result.documents).toBe(4); + expect(summaryOf(result.keys, "status").documents).toBe(4); + expect(summaryOf(result.keys, "status").types[0]!.collections).toEqual(["notes", "work"]); + }); + + test("a single collection scopes counts and the denominator", () => { + const result = listMetadata(store.db, { collection: "notes" }); + + expect(result.documents).toBe(2); + expect(valuesOf(result.keys, "status")).toEqual([["draft", 1], ["published", 1]]); + expect(summaryOf(result.keys, "status").types[0]!.collections).toEqual(["notes"]); + }); + + test("a collection list scopes to exactly those collections", async () => { + await insertMetadataDoc("other", { status: "elsewhere" }); + + const result = listMetadata(store.db, { collection: ["notes", "work"] }); + + expect(result.documents).toBe(4); + expect(valuesOf(result.keys, "status").map(([value]) => value)).not.toContain("elsewhere"); + }); + + test("an unknown collection yields an empty scope", () => { + const result = listMetadata(store.db, { collection: "missing" }); + + expect(result).toEqual({ documents: 0, totalKeys: 0, keys: [], remainingKeys: 0 }); + }); + + test("filter narrows which documents are counted and reports how many pass", () => { + const result = listMetadata(store.db, { + filter: { field: "status", operator: "eq", value: "published" }, + }); + + expect(result.documents).toBe(4); + expect(result.filteredDocuments).toBe(2); + expect(summaryOf(result.keys, "priority").documents).toBe(2); + expect(valuesOf(result.keys, "priority")).toEqual([[3, 1], [5, 1]]); + expect(valuesOf(result.keys, "status")).toEqual([["published", 2]]); + }); + + test("filter composes with the same AST search accepts", () => { + const result = listMetadata(store.db, { + filter: { + operator: "and", + operands: [ + { field: "status", operator: "eq", value: "published" }, + { field: "priority", operator: "gte", value: 4 }, + ], + }, + }); + + expect(result.filteredDocuments).toBe(1); + expect(summaryOf(result.keys, "status").types[0]!.collections).toEqual(["work"]); + }); +}); + +describe("listMetadata match", () => { + beforeEach(async () => { + await insertMetadataDoc("notes", { topics: ["typescript", "sqlite"], owner: "docs-team", "mem-kind": "fact", priority: 3, reviewed: true }); + await insertMetadataDoc("notes", { topics: ["typescript", "PostgreSQL"], owner: "search-team", reviewers: ["docs-team", "security-team"], "mem-scope": "user", priority: 10, reviewed: false }); + }); + + function keysOf(match: MetadataFilter): string[] { + return listMetadata(store.db, { match }).keys.map(summary => summary.key); + } + + test("eq on the key field reports only that key", () => { + const result = listMetadata(store.db, { match: { field: "key", operator: "eq", value: "topics" } }); + + expect(result.keys.map(summary => summary.key)).toEqual(["topics"]); + expect(valuesOf(result.keys, "topics")).toEqual([["typescript", 2], ["PostgreSQL", 1], ["sqlite", 1]]); + }); + + test("text operators on the key field select a family of keys", () => { + expect(keysOf({ field: "key", operator: "prefix", value: "mem-" })).toEqual(["mem-kind", "mem-scope"]); + expect(keysOf({ field: "key", operator: "suffix", value: "-scope" })).toEqual(["mem-scope"]); + expect(keysOf({ field: "key", operator: "contains", value: "review" })).toEqual(["reviewed", "reviewers"]); + }); + + test("membership on the key field answers which of several names exist", () => { + expect(keysOf({ field: "key", operator: "in", value: ["tags", "topics", "labels"] })).toEqual(["topics"]); + expect(keysOf({ field: "key", operator: "nin", value: ["topics", "owner", "priority", "reviewed"] })).toEqual(["mem-kind", "mem-scope", "reviewers"]); + }); + + test("a match nothing satisfies yields an empty key list and keeps the denominator", () => { + const result = listMetadata(store.db, { match: { field: "key", operator: "prefix", value: "missing-" } }); + + expect(result.keys).toEqual([]); + expect(result.documents).toBe(2); + }); + + test("eq on the value field is a reverse lookup across keys", () => { + const result = listMetadata(store.db, { match: { field: "value", operator: "eq", value: "docs-team" } }); + + expect(result.keys.map(summary => summary.key)).toEqual(["owner", "reviewers"]); + expect(summaryOf(result.keys, "owner").documents).toBe(1); + expect(summaryOf(result.keys, "owner").types[0]!.distinctValues).toBe(1); + expect(valuesOf(result.keys, "reviewers")).toEqual([["docs-team", 1]]); + }); + + test("a key and a value condition compose with and, keeping the key's array-ness", () => { + const result = listMetadata(store.db, { + match: { + operator: "and", + operands: [ + { field: "key", operator: "eq", value: "topics" }, + { field: "value", operator: "prefix", value: "type" }, + ], + }, + }); + + expect(valuesOf(result.keys, "topics")).toEqual([["typescript", 2]]); + expect(summaryOf(result.keys, "topics").types[0]!.remainingValues).toBe(0); + // Array-ness describes the key, not the matched entries. + expect(summaryOf(result.keys, "topics").types[0]!.multiValued).toBe(true); + }); + + test("values match by their own type, never through a text form", () => { + expect(keysOf({ field: "value", operator: "eq", value: 10 })).toEqual(["priority"]); + expect(keysOf({ field: "value", operator: "eq", value: "10" })).toEqual([]); + expect(keysOf({ field: "value", operator: "eq", value: true })).toEqual(["reviewed"]); + expect(keysOf({ field: "value", operator: "contains", value: "true" })).toEqual([]); + expect(keysOf({ field: "value", operator: "prefix", value: "1" })).toEqual([]); + }); + + test("ordered comparisons on the value field select values and narrow the range", () => { + const result = listMetadata(store.db, { + match: { + operator: "and", + operands: [ + { field: "key", operator: "eq", value: "priority" }, + { field: "value", operator: "gte", value: 5 }, + ], + }, + }); + + expect(valuesOf(result.keys, "priority")).toEqual([[10, 1]]); + expect(summaryOf(result.keys, "priority").types[0]!.range).toEqual({ min: 10, median: 10, max: 10 }); + expect(summaryOf(result.keys, "priority").documents).toBe(1); + }); + + test("type on the value field selects keys by the type they hold", () => { + expect(keysOf({ field: "value", operator: "type", value: "boolean" })).toEqual(["reviewed"]); + expect(keysOf({ field: "value", operator: "type", value: "number" })).toEqual(["priority"]); + expect(keysOf({ operator: "not", operand: { field: "value", operator: "type", value: "string" } })).toEqual(["priority", "reviewed"]); + }); + + test("type on the value field reports one side of a key whose documents disagree", async () => { + await insertMetadataDoc("work", { priority: "high" }); + + const result = listMetadata(store.db, { + match: { + operator: "and", + operands: [ + { field: "key", operator: "eq", value: "priority" }, + { field: "value", operator: "type", value: "number" }, + ], + }, + }); + + const priority = summaryOf(result.keys, "priority"); + expect(priority.types.map(typeSummary => typeSummary.type)).toEqual(["number"]); + expect(priority.documents).toBe(2); + expect(priority.types[0]!.collections).toEqual(["notes"]); + }); + + test("membership and negation on the value field exclude known values", () => { + const result = listMetadata(store.db, { + match: { + operator: "and", + operands: [ + { field: "key", operator: "eq", value: "topics" }, + { operator: "not", operand: { field: "value", operator: "in", value: ["typescript"] } }, + ], + }, + }); + + expect(valuesOf(result.keys, "topics")).toEqual([["PostgreSQL", 1], ["sqlite", 1]]); + expect(keysOf({ field: "value", operator: "nin", value: ["docs-team", "search-team", "security-team", "typescript", "sqlite", "PostgreSQL", "fact", "user"] })).toEqual([]); + }); + + test("or composes across the key and value fields", () => { + const result = listMetadata(store.db, { + match: { + operator: "or", + operands: [ + { field: "key", operator: "eq", value: "owner" }, + { field: "value", operator: "eq", value: "docs-team" }, + ], + }, + }); + + expect(result.keys.map(summary => summary.key)).toEqual(["owner", "reviewers"]); + // Every owner entry satisfies the first branch, only one reviewers entry the second. + expect(valuesOf(result.keys, "owner")).toEqual([["docs-team", 1], ["search-team", 1]]); + expect(valuesOf(result.keys, "reviewers")).toEqual([["docs-team", 1]]); + }); + + test("caseInsensitive folds ASCII on either field", () => { + expect(keysOf({ field: "value", operator: "contains", value: "sql" })).toEqual(["topics"]); + expect(valuesOf(listMetadata(store.db, { match: { field: "value", operator: "contains", value: "sql" } }).keys, "topics")).toEqual([["sqlite", 1]]); + expect(valuesOf(listMetadata(store.db, { match: { field: "value", operator: "contains", value: "sql", caseInsensitive: true } }).keys, "topics")).toEqual([["PostgreSQL", 1], ["sqlite", 1]]); + expect(keysOf({ field: "key", operator: "eq", value: "TOPICS", caseInsensitive: true })).toEqual(["topics"]); + }); + + test("a number or boolean operand against the key field matches nothing", () => { + expect(keysOf({ field: "key", operator: "eq", value: 3 })).toEqual([]); + expect(keysOf({ field: "key", operator: "gte", value: 0 })).toEqual([]); + expect(keysOf({ field: "key", operator: "in", value: [true] })).toEqual([]); + // The key field is always a string, so `type string` on it is every key. + const everyKey = listMetadata(store.db).keys.map(summary => summary.key); + expect(keysOf({ field: "key", operator: "type", value: "string" })).toEqual(everyKey); + expect(keysOf({ field: "key", operator: "type", value: "number" })).toEqual([]); + }); + + test("match composes with filter: filter counts documents, match selects entries", () => { + const result = listMetadata(store.db, { + match: { field: "key", operator: "eq", value: "topics" }, + filter: { field: "reviewed", operator: "eq", value: true }, + }); + + expect(result.filteredDocuments).toBe(1); + expect(valuesOf(result.keys, "topics")).toEqual([["sqlite", 1], ["typescript", 1]]); + }); + + test("rejects conditions that have no meaning for a metadata entry", () => { + const cases: [unknown, RegExp][] = [ + [{ field: "topics", operator: "eq", value: "x" }, /^Invalid metadata match at \$: 'topics' is not a field of a metadata entry, expected 'key' or 'value'/], + [{ field: "value", operator: "exists", value: true }, /'exists' has no meaning for a single metadata entry/], + [{ field: "value", operator: "all", value: ["a"] }, /'all' has no meaning for a single metadata entry/], + [{ operator: "and", operands: [{ field: "key", operator: "eq", value: "a" }, { field: "nope", operator: "eq", value: 1 }] }, /at \$\.operands\[1\]:/], + [{ field: "value", operator: "eq", value: 3, caseInsensitive: true }, /^Invalid metadata match at \$\.caseInsensitive:/], + ]; + + for (const [match, expected] of cases) { + expect(() => listMetadata(store.db, { match: match as MetadataMatch })).toThrow(expected); + expect(() => listMetadata(store.db, { match: match as MetadataMatch })).toThrow(MetadataFilterError); + } + }); +}); + +describe("listMetadata key window", () => { + // Coverage: e 5, d 4, c 3, b 2, a 1. Report order is e, d, c, b, a. + beforeEach(async () => { + const keyNames = ["a", "b", "c", "d", "e"]; + for (const [index, key] of keyNames.entries()) { + for (let count = 0; count <= index; count += 1) await insertMetadataDoc("notes", { [key]: "x" }); + } + }); + + function keysOf(options: ListMetadataOptions = {}): string[] { + return listMetadata(store.db, options).keys.map(summary => summary.key); + } + + test("reports every key with no remainder when the vocabulary fits the window", () => { + const result = listMetadata(store.db); + + expect(result.keys.map(summary => summary.key)).toEqual(["e", "d", "c", "b", "a"]); + expect(result.totalKeys).toBe(5); + expect(result.remainingKeys).toBe(0); + }); + + test("keyLimit windows keys in report order and reports the exact remainder", () => { + const result = listMetadata(store.db, { keyLimit: 2 }); + + expect(result.keys.map(summary => summary.key)).toEqual(["e", "d"]); + expect(result.totalKeys).toBe(5); + expect(result.remainingKeys).toBe(3); + }); + + test("keyOffset pages through the vocabulary without gaps or repeats", () => { + const pages = [0, 2, 4].map(keyOffset => listMetadata(store.db, { keyLimit: 2, keyOffset })); + + expect(pages.map(page => page.keys.map(summary => summary.key))).toEqual([["e", "d"], ["c", "b"], ["a"]]); + expect(pages.map(page => page.remainingKeys)).toEqual([3, 1, 0]); + expect(pages.every(page => page.totalKeys === 5)).toBe(true); + }); + + test("an offset past the end reports no keys and keeps the total", () => { + const result = listMetadata(store.db, { keyOffset: 10 }); + + expect(result.keys).toEqual([]); + expect(result.totalKeys).toBe(5); + expect(result.remainingKeys).toBe(0); + }); + + test("an infinite keyLimit removes the window", () => { + expect(keysOf({ keyLimit: Infinity })).toEqual(["e", "d", "c", "b", "a"]); + }); + + test("the window applies to the keys the match admits", () => { + const result = listMetadata(store.db, { match: { field: "key", operator: "in", value: ["a", "c", "e"] }, keyLimit: 2 }); + + expect(result.keys.map(summary => summary.key)).toEqual(["e", "c"]); + expect(result.totalKeys).toBe(3); + expect(result.remainingKeys).toBe(1); + }); + + test("a key summary inside the window is complete", async () => { + await insertMetadataDoc("work", { e: ["y", "z"], d: 1 }); + + const result = listMetadata(store.db, { keyLimit: 1 }); + const e = summaryOf(result.keys, "e"); + + expect(e.documents).toBe(6); + expect(e.types[0]!.multiValued).toBe(true); + expect(e.types[0]!.distinctValues).toBe(3); + expect(e.types[0]!.collections).toEqual(["notes", "work"]); + expect(result.keys).toHaveLength(1); + }); + + test("collection provenance uses UTF-8 ordering without SQLite 3.44 aggregate syntax", async () => { + for (const collection of ["𐀀", "", "notes", "B"]) { + await insertMetadataDoc(collection, { provenance: "shared" }); + } + const observed = observeStatements(store.db, sql => { + expect(sql).not.toMatch(/json_group_array\([^)]*\bORDER\s+BY/iu); + }); + const report = listMetadata(observed, { match: { field: "key", operator: "eq", value: "provenance" } }); + expect(summaryOf(report.keys, "provenance").types[0]!.collections).toEqual(["B", "notes", "", "𐀀"]); + }); + + test("pages are prefixes of the full result under one collation, on both readers", async () => { + // Equal coverage, so order falls to the key name. BINARY collation puts + // uppercase before "_", "_" before lowercase, and lowercase before "é". + await insertMetadataDoc("work", { _id: 1, a: 1, B: 1, "é": 1, Z: 1 }); + const names = ["B", "Z", "_id", "a", "é"]; + + const fullResult = listMetadata(store.db, { collection: "work" }).keys.map(summary => summary.key); + const pages = [0, 2, 4].flatMap(keyOffset => + listMetadata(store.db, { collection: "work", keyLimit: 2, keyOffset }).keys.map(summary => summary.key)); + + expect(fullResult).toEqual(names); + expect(pages).toEqual(names); + expect(listMetadataCollectionSummaries(store.db).get("work")!.keys.map(overview => overview.key)).toEqual(names); + expect(listMetadataCollectionSummaries(store.db, 3).get("work")!.keys.map(overview => overview.key)).toEqual(names.slice(0, 3)); + }); + + test("ranks keys once per report and groups values once for both totals and windows", () => { + const statements: string[] = []; + const observed = observeStatements(store.db, sql => statements.push(sql)); + const options: ListMetadataOptions[] = [ + { keyLimit: 5, valueLimit: 3 }, + { collection: "notes", keyLimit: 5, valueLimit: 3 }, + { filter: { field: "e", operator: "prefix", value: "x" }, keyLimit: 5 }, + {}, + ]; + // One statement ranks the keys, but the planner may still evaluate its + // CTE once per joined row unless it is materialized. Check both, with + // and without table statistics, which change the planner's choice. + for (const statistics of [false, true]) { + if (statistics) store.db.exec("ANALYZE"); + for (const option of options) { + statements.length = 0; + listMetadata(observed, option); + const rankingStatements = statements.filter(sql => sql.includes("selected_keys AS")); + expect(rankingStatements).toHaveLength(1); + expect(statements.filter(sql => sql.includes("GROUP BY mv.key, mv.value_type, mv.text_value"))).toHaveLength(1); + if (option.filter) expect(statements.some(sql => sql.includes("d.id IN (SELECT mv.document_id"))).toBe(true); + const plan = planOf(rankingStatements[0]!); + expect(plan.some(step => step.includes("MATERIALIZE selected_keys")), rankingStatements[0]).toBe(true); + expect(plan.some(step => step.includes("CO-ROUTINE selected_keys")), rankingStatements[0]).toBe(false); + } + } + + statements.length = 0; + listMetadataCollectionSummaries(observed, 10); + expect(statements).toHaveLength(1); + expect(statements[0]).toContain("mv.ordinal = 0"); + expect(statements[0]).not.toContain("number_value"); + }); +}); + +describe("metadata overviews", () => { + test("batched collection overviews agree with full reports through lifecycle changes", async () => { + await insertMetadataDoc("notes", { tags: ["a", "b", "c"], priority: 1 }); + const changed = await insertMetadataDoc("notes", { tags: 7, priority: "high" }); + await insertMetadataDoc("work", { tags: true, "é": "x", "😀": "x" }); + const stale = await insertMetadataDoc("work", { old: "hidden" }); + const inactive = await insertMetadataDoc("work", { deleted: "hidden" }); + store.db.prepare("UPDATE document_metadata SET extraction_version = 0 WHERE document_id = ?").run(stale); + store.db.prepare("UPDATE documents SET active = 0 WHERE id = ?").run(inactive); + + for (const phase of ["initial", "replaced", "renamed", "deleted"] as const) { + if (phase === "replaced") replaceDocumentMetadata(store.db, changed, { metadata: { tags: [true, false] }, extractionVersion: METADATA_EXTRACTION_VERSION }); + if (phase === "renamed") store.db.prepare("UPDATE documents SET collection = 'renamed' WHERE collection = 'notes'").run(); + if (phase === "deleted") store.db.prepare("DELETE FROM documents WHERE id = ?").run(changed); + for (const keyLimit of [1, 10, Infinity]) { + const statements: string[] = []; + const observed = observeStatements(store.db, sql => statements.push(sql)); + const overviews = listMetadataCollectionSummaries(observed, keyLimit); + expect(statements).toHaveLength(1); + for (const collection of ["notes", "renamed", "work", "missing"]) { + const report = listMetadata(store.db, { collection, keyLimit }); + const expected = { totalKeys: report.totalKeys, keys: report.keys.map(key => ({ key: key.key, documents: key.documents, types: key.types.map(type => type.type) })) }; + expect(overviews.get(collection) ?? { totalKeys: 0, keys: [] }).toEqual(expected); + } + } + } + }); +}); + +describe("entry text predicates", () => { + test("positive, negated, and compound matches partition empty and nonempty strings", async () => { + const values = ["", "draft", "PUBLISHED", "abc\u0000XYZ", "\u0000", "😀XYZ", 7, true]; + for (const value of values) await insertMetadataDoc("notes", { label: value }); + + for (const operator of ["contains", "prefix", "suffix"] as const) { + for (const operand of ["ft", "PUB", "XYZ", "\u0000", "😀", "nomatch"]) { + for (const caseInsensitive of [false, true]) { + const fold = (text: string) => caseInsensitive ? text.replace(/[A-Z]/gu, letter => letter.toLowerCase()) : text; + const expected = values.map(value => { + if (typeof value !== "string") return false; + const text = fold(value); + const needle = fold(operand); + return operator === "contains" ? text.includes(needle) : operator === "prefix" ? text.startsWith(needle) : text.endsWith(needle); + }); + const condition = { field: "value", operator, value: operand, caseInsensitive } as const; + const predicates: MetadataMatch[] = [ + condition, + { operator: "not", operand: condition }, + { operator: "and", operands: [{ field: "key", operator: "eq", value: "label" }, { operator: "not", operand: condition }] }, + { operator: "or", operands: [condition, { operator: "not", operand: condition }] }, + ]; + const expectations = [expected, expected.map(value => !value), expected.map(value => !value), values.map(() => true)]; + for (const [index, predicate] of predicates.entries()) { + const compiled = compileMetadataMatch(predicate, "mv"); + const rows = store.db.prepare(`SELECT ${compiled.sql} AS matched FROM document_metadata_values mv ORDER BY document_id`).all(...compiled.params) as { matched: number }[]; + expect(rows.map(row => row.matched)).toEqual(expectations[index]!.map(Number)); + const report = listMetadata(store.db, { match: predicate, valueLimit: Infinity }); + expect(report.keys[0]?.documents ?? 0).toBe(expectations[index]!.filter(Boolean).length); + } + + const filter: MetadataFilter = { operator: "not", operand: { ...condition, field: "label" } }; + for (const scope of ["candidates", "corpus"] as const) { + const compiled = compileMetadataFilter(filter, "d", scope); + const rows = store.db.prepare(`SELECT ${compiled.sql} AS matched FROM documents d ORDER BY id`).all(...compiled.params) as { matched: number }[]; + expect(rows.map(row => row.matched)).toEqual(expected.map(value => Number(!value))); + } + } + } + } + }); +}); + +describe("listMetadata consistency", () => { + test("reads the whole report from one snapshot while another connection writes", async () => { + const documentId = await insertMetadataDoc("notes", { priority: 1 }); + const writer = createStore(store.db.prepare("PRAGMA database_list").all().map(row => (row as { file: string }).file)[0]!); + + try { + let rewritten = false; + const observed = observeStatements(store.db, sql => { + // Commit a change from the second connection after the type statistics + // (min and max) have been read and before the medians and values are, + // the interleaving that would report a range no document ever had. + if (rewritten || !sql.includes("PARTITION BY mv.key ORDER BY mv.number_value")) return; + rewritten = true; + replaceDocumentMetadata(writer.db, documentId, { metadata: { priority: 100 }, extractionVersion: METADATA_EXTRACTION_VERSION }); + }); + + const priority = summaryOf(listMetadata(observed).keys, "priority").types[0]!; + + expect(rewritten).toBe(true); + expect(priority.range).toEqual({ min: 1, median: 1, max: 1 }); + expect(priority.values).toEqual([{ value: 1, documents: 1 }]); + expect(summaryOf(listMetadata(store.db).keys, "priority").types[0]!.values).toEqual([{ value: 100, documents: 1 }]); + } finally { + writer.close(); + } + }); +}); + +describe("listMetadata options", () => { + test("rejects a window or ordering option outside its domain", () => { + const cases: [ListMetadataOptions, RegExp][] = [ + [{ keyLimit: 0 }, /^Invalid keyLimit: expected a positive integer or Infinity, received 0$/], + [{ keyLimit: 1.5 }, /^Invalid keyLimit: expected a positive integer or Infinity, received 1.5$/], + [{ keyOffset: -1 }, /^Invalid keyOffset: expected a non-negative integer, received -1$/], + [{ valueLimit: -Infinity }, /^Invalid valueLimit: expected a positive integer or Infinity, received -Infinity$/], + [{ valueOffset: 0.5 }, /^Invalid valueOffset: expected a non-negative integer, received 0.5$/], + [{ minCount: 0 }, /^Invalid minCount: expected a positive integer, received 0$/], + [{ sort: "size" as "count" }, /^Invalid sort: expected 'count' or 'value', received "size"$/], + ]; + + for (const [options, expected] of cases) { + expect(() => listMetadata(store.db, options)).toThrow(expected); + expect(() => listMetadata(store.db, options)).toThrow(MetadataOptionError); + } + }); + + test("names the offending option on the error", () => { + try { + listMetadata(store.db, { keyOffset: -1 }); + throw new Error("expected listMetadata to throw"); + } catch (error) { + expect(error).toBeInstanceOf(MetadataOptionError); + expect((error as MetadataOptionError).option).toBe("keyOffset"); + } + }); + + test("rejects integers SQLite cannot bind as INTEGER, and binds the largest it can", async () => { + await insertMetadataDoc("notes", { status: "x" }); + const largest = Number.MAX_SAFE_INTEGER; + + expect(() => listMetadata(store.db, { keyLimit: 1e30 })).toThrow(MetadataOptionError); + expect(() => listMetadata(store.db, { keyOffset: 2 ** 53 })).toThrow(/^Invalid keyOffset: expected a non-negative integer, received 9007199254740992$/); + expect(() => listMetadata(store.db, { valueOffset: 1e30 })).toThrow(MetadataOptionError); + expect(() => listMetadata(store.db, { minCount: 1e30 })).toThrow(MetadataOptionError); + + expect(listMetadata(store.db, { keyLimit: largest, keyOffset: largest }).keys).toEqual([]); + expect(summaryOf(listMetadata(store.db, { valueLimit: largest, valueOffset: largest }).keys, "status").types[0]!.values).toEqual([]); + expect(summaryOf(listMetadata(store.db, { valueLimit: largest }).keys, "status").types[0]!.values).toHaveLength(1); + }); + + test("the production value window includes its exact integer endpoint past 2^53", async () => { + await insertMetadataDoc("notes", { label: "x" }); + let windowChecked = false; + const observed = observeStatements(store.db, () => {}, (sql, params) => { + const windowSql = /\((rank > \? AND rank <= .*)\) AS in_window/u.exec(sql)?.[1]; + if (!windowSql) return; + windowChecked = true; + + // Execute the actual emitted predicate and bindings on a tiny exact-rank + // fixture. Reading the rank itself as a JS number would round it again. + const rows = store.db.prepare(` + WITH ranks(rank) AS (VALUES (9007199254740991), (9007199254740992), (9007199254740993), (9007199254740994)) + SELECT CAST(rank AS TEXT) AS rank FROM ranks WHERE ${windowSql} ORDER BY rank + `).all(...params.slice(-3)); + expect(rows).toEqual([{ rank: "9007199254740992" }, { rank: "9007199254740993" }]); + }); + + listMetadata(observed, { valueOffset: Number.MAX_SAFE_INTEGER, valueLimit: 2 }); + expect(windowChecked).toBe(true); + }); + + test("rejects a filter and match whose bindings together exceed the statement budget", async () => { + await insertMetadataDoc("notes", { status: "x" }); + const members = Array.from({ length: 64 }, (_, index) => `v${index}`); + // 7 groups of 31 conditions: 225 nodes, inside every parser limit. + const wideFilter: MetadataFilter = { + operator: "and", + operands: Array.from({ length: 7 }, () => ({ + operator: "and" as const, + operands: Array.from({ length: 31 }, () => ({ field: "status", operator: "all" as const, value: members })), + })), + }; + const wideMatch: MetadataMatch = { + operator: "or", + operands: Array.from({ length: 7 }, () => ({ + operator: "or" as const, + operands: Array.from({ length: 31 }, () => ({ field: "value" as const, operator: "in" as const, value: members })), + })), + }; + + // 217 conditions binding field and value per member: 27,776, under budget alone. + expect(listMetadata(store.db, { filter: wideFilter }).totalKeys).toBe(0); + expect(listMetadata(store.db, { match: wideMatch }).totalKeys).toBe(0); + + try { + listMetadata(store.db, { filter: wideFilter, match: wideMatch }); + throw new Error("expected listMetadata to throw"); + } catch (error) { + expect(error).toBeInstanceOf(MetadataBindingBudgetError); + const budgetError = error as MetadataBindingBudgetError; + expect(budgetError.budget).toBe(METADATA_SQL_BINDING_BUDGET); + expect(budgetError.bindings).toBeGreaterThan(METADATA_SQL_BINDING_BUDGET); + expect(budgetError.message).toMatch(/^filter and match together bind \d+ SQL parameters, over the budget of 30000\. Shorten their membership lists\.$/); + } + }); +}); + +describe("listMetadata value window", () => { + beforeEach(async () => { + const tags = ["a", "a", "a", "b", "b", "c", "d", "d", "d", "d"]; + for (const tag of tags) await insertMetadataDoc("notes", { tag }); + }); + + test("defaults to count order with a stable tiebreak", () => { + expect(valuesOf(listMetadata(store.db).keys, "tag")).toEqual([["d", 4], ["a", 3], ["b", 2], ["c", 1]]); + }); + + test("sort by value orders ascending", () => { + expect(valuesOf(listMetadata(store.db, { sort: "value" }).keys, "tag")).toEqual([["a", 3], ["b", 2], ["c", 1], ["d", 4]]); + }); + + test("valueLimit windows values and reports the exact remainder", () => { + const tag = summaryOf(listMetadata(store.db, { valueLimit: 2 }).keys, "tag").types[0]!; + + expect(tag.values.map(count => count.value)).toEqual(["d", "a"]); + expect(tag.distinctValues).toBe(4); + expect(tag.remainingValues).toBe(2); + }); + + test("an infinite valueLimit removes the window", () => { + const tag = summaryOf(listMetadata(store.db, { valueLimit: Infinity }).keys, "tag").types[0]!; + + expect(tag.values).toHaveLength(4); + expect(tag.remainingValues).toBe(0); + }); + + test("valueOffset pages through the values without gaps or repeats", () => { + const pages = [0, 2, 4].map(valueOffset => + summaryOf(listMetadata(store.db, { valueLimit: 2, valueOffset }).keys, "tag").types[0]!); + + expect(pages.map(tag => tag.values.map(count => count.value))).toEqual([["d", "a"], ["b", "c"], []]); + expect(pages.map(tag => tag.remainingValues)).toEqual([2, 0, 0]); + expect(pages.every(tag => tag.distinctValues === 4)).toBe(true); + }); + + test("valueOffset follows the requested sort", () => { + const tag = summaryOf(listMetadata(store.db, { sort: "value", valueLimit: 2, valueOffset: 1 }).keys, "tag").types[0]!; + + expect(tag.values.map(count => count.value)).toEqual(["b", "c"]); + }); + + test("minCount drops the tail from the values, distinct count, and remainder", () => { + const tag = summaryOf(listMetadata(store.db, { minCount: 3, valueLimit: 1 }).keys, "tag").types[0]!; + + expect(tag.values).toEqual([{ value: "d", documents: 4 }]); + expect(tag.distinctValues).toBe(2); + expect(tag.remainingValues).toBe(1); + // Coverage still counts every document holding the key. + expect(tag.documents).toBe(10); + }); + + test("a minCount nothing meets leaves the key with an empty window", () => { + const tag = summaryOf(listMetadata(store.db, { minCount: 99 }).keys, "tag").types[0]!; + + expect(tag.values).toEqual([]); + expect(tag.distinctValues).toBe(0); + expect(tag.remainingValues).toBe(0); + }); +}); + +describe("listMetadata numbers", () => { + test("reports min, median, and max for an odd count", async () => { + for (const priority of [5, 1, 3]) await insertMetadataDoc("notes", { priority }); + + const priority = summaryOf(listMetadata(store.db).keys, "priority").types[0]!; + + expect(priority.range).toEqual({ min: 1, median: 3, max: 5 }); + expect(priority.values.map(count => count.value)).toEqual([1, 3, 5]); + }); + + test("averages the middle values for an even count", async () => { + for (const priority of [1, 2, 3, 10]) await insertMetadataDoc("notes", { priority }); + + expect(summaryOf(listMetadata(store.db).keys, "priority").types[0]!.range).toEqual({ min: 1, median: 2.5, max: 10 }); + }); + + test("the even-count median is finite and correctly rounded across the double range", async () => { + const cases: [number, number, number][] = [ + [1e308, 1.5e308, 1.25e308], + [-1.5e308, -1e308, -1.25e308], + [-1e308, 1e308, 0], + [Number.MIN_VALUE, 2 * Number.MIN_VALUE, 2 * Number.MIN_VALUE], + [-2 * Number.MIN_VALUE, -Number.MIN_VALUE, -2 * Number.MIN_VALUE], + [Number.MAX_VALUE, Number.MAX_VALUE, Number.MAX_VALUE], + ]; + + for (const [index, [lower, upper]] of cases.entries()) { + await insertMetadataDoc(`c${index}`, { n: lower }); + await insertMetadataDoc(`c${index}`, { n: upper }); + } + + for (const [index, [lower, upper, median]] of cases.entries()) { + const n = summaryOf(listMetadata(store.db, { collection: `c${index}` }).keys, "n").types[0]!; + expect(n.range, `${lower}, ${upper}`).toEqual({ min: lower, median, max: upper }); + } + }); + + test("median counts every value row, including array elements", async () => { + await insertMetadataDoc("notes", { scores: [1, 1, 1] }); + await insertMetadataDoc("notes", { scores: [9] }); + + const scores = summaryOf(listMetadata(store.db).keys, "scores").types[0]!; + + expect(scores.range).toEqual({ min: 1, median: 1, max: 9 }); + expect(scores.values).toEqual([{ value: 1, documents: 1 }, { value: 9, documents: 1 }]); + }); + + test("strings and booleans carry no range", async () => { + await insertMetadataDoc("notes", { status: "x", reviewed: true }); + + const result = listMetadata(store.db); + + expect(summaryOf(result.keys, "status").types[0]!.range).toBeUndefined(); + expect(summaryOf(result.keys, "reviewed").types[0]!.range).toBeUndefined(); + }); +}); + +describe("listMetadata type conflicts and attribution", () => { + test("splits a key by type and partitions its documents", async () => { + await insertMetadataDoc("notes", { priority: 3 }); + await insertMetadataDoc("notes", { priority: 1 }); + await insertMetadataDoc("work", { priority: "high" }); + + const priority = summaryOf(listMetadata(store.db).keys, "priority"); + + expect(priority.documents).toBe(3); + expect(priority.types.map(typeSummary => [typeSummary.type, typeSummary.documents])).toEqual([["number", 2], ["string", 1]]); + expect(priority.types[0]!.range).toEqual({ min: 1, median: 2, max: 3 }); + expect(priority.types[0]!.collections).toEqual(["notes"]); + expect(priority.types[1]!.collections).toEqual(["work"]); + expect(priority.types[1]!.values).toEqual([{ value: "high", documents: 1 }]); + }); + + test("orders equally covered types by name", async () => { + await insertMetadataDoc("notes", { flag: true }); + await insertMetadataDoc("notes", { flag: "yes" }); + + expect(summaryOf(listMetadata(store.db).keys, "flag").types.map(typeSummary => typeSummary.type)).toEqual(["boolean", "string"]); + }); + + test("orders keys by coverage descending, then name", async () => { + await insertMetadataDoc("notes", { zeta: 1, alpha: 1, mid: 1 }); + await insertMetadataDoc("notes", { zeta: 1, alpha: 1 }); + await insertMetadataDoc("notes", { zeta: 1 }); + + expect(listMetadata(store.db).keys.map(summary => summary.key)).toEqual(["zeta", "alpha", "mid"]); + }); +}); + +describe("listMetadata agrees with filtered search", () => { + test("every reported value is reachable through an eq filter under the same scope", async () => { + await insertMetadataDoc("notes", { topics: ["a", "b"], priority: 3, reviewed: true, owner: "docs-team" }); + await insertMetadataDoc("notes", { topics: ["b"], priority: 1.5, reviewed: false }); + await insertMetadataDoc("work", { topics: ["c"], priority: 3 }); + const pendingId = await insertMetadataDoc("notes", { topics: ["ghost"] }); + store.db.prepare(`DELETE FROM document_metadata WHERE document_id = ?`).run(pendingId); + + const scopes: ListMetadataOptions[] = [{}, { collection: "notes" }, { collection: ["notes", "work"] }]; + for (const scope of scopes) { + const result = listMetadata(store.db, { ...scope, valueLimit: Infinity }); + expect(result.keys.length).toBeGreaterThan(0); + + for (const summary of result.keys) { + for (const typeSummary of summary.types) { + for (const count of typeSummary.values) { + const hits = searchFTS(store.db, "doc", 100, scope.collection, { field: summary.key, operator: "eq", value: count.value }); + expect(hits, `${summary.key} = ${String(count.value)}`).toHaveLength(count.documents); + } + } + } + } + }); +}); + +describe("status view helpers", () => { + beforeEach(async () => { + await insertMetadataDoc("notes", { status: "published", priority: 3 }); + await insertMetadataDoc("notes", { status: "draft" }); + await insertMetadataDoc("notes", {}); + await insertMetadataDoc("work", { priority: "high", source: "jira" }); + const pendingId = await insertMetadataDoc("work", { status: "ghost" }); + store.db.prepare(`DELETE FROM document_metadata WHERE document_id = ?`).run(pendingId); + }); + + test("listMetadataCollectionSummaries reports names, coverage, and types per collection in coverage order", () => { + const overviews = listMetadataCollectionSummaries(store.db); + expect(overviews.get("notes")).toEqual({ totalKeys: 2, keys: [ + { key: "status", documents: 2, types: ["string"] }, + { key: "priority", documents: 1, types: ["number"] }, + ] }); + expect(overviews.get("work")).toEqual({ totalKeys: 2, keys: [ + { key: "priority", documents: 1, types: ["string"] }, + { key: "source", documents: 1, types: ["string"] }, + ] }); + expect(overviews.has("missing")).toBe(false); + + const windowed = listMetadataCollectionSummaries(store.db, 1); + expect(windowed.get("notes")).toEqual({ totalKeys: 2, keys: [{ key: "status", documents: 2, types: ["string"] }] }); + expect(listMetadataCollectionSummaries(store.db, Infinity).get("notes")!.keys).toHaveLength(2); + }); + + test("countMetadataKeys reports the vocabulary size in scope", () => { + expect(countMetadataKeys(store.db)).toBe(3); + expect(countMetadataKeys(store.db, ["work"])).toBe(2); + expect(countMetadataKeys(store.db, ["missing"])).toBe(0); + }); + + test("countDocumentsWithMetadata counts extracted documents declaring a key", () => { + expect(countDocumentsWithMetadata(store.db)).toBe(3); + expect(countDocumentsWithMetadata(store.db, ["notes"])).toBe(2); + }); + + test("status batches metadata overviews and initialization reads no metadata values", () => { + const statements: string[] = []; + const observed = observeStatements(store.db, sql => statements.push(sql)); + const status = getStatus(observed); + expect(status.collections).toHaveLength(2); + expect(status.collections.find(collection => collection.name === "notes")?.metadataKeyCount).toBe(2); + expect(statements.filter(sql => sql.includes("document_metadata_values"))).toHaveLength(1); + + statements.length = 0; + const summary = getStatusSummary(observed); + expect(summary.totalDocuments).toBe(status.totalDocuments); + expect(summary.pendingMetadata).toBe(status.pendingMetadata); + expect(summary.collections.map(collection => collection.name)).toEqual(status.collections.map(collection => collection.name)); + expect(summary.collections[0]).not.toHaveProperty("metadataKeys"); + expect(statements.some(sql => sql.includes("document_metadata_values"))).toBe(false); + }); + + test("countDocumentsPendingMetadata accepts a collection scope", () => { + expect(countDocumentsPendingMetadata(store.db)).toBe(1); + expect(countDocumentsPendingMetadata(store.db, ["notes"])).toBe(0); + expect(countDocumentsPendingMetadata(store.db, ["work"])).toBe(1); + }); +}); diff --git a/test/metadata-filter.test-d.ts b/test/metadata-filter.test-d.ts new file mode 100644 index 000000000..9f6cb4e76 --- /dev/null +++ b/test/metadata-filter.test-d.ts @@ -0,0 +1,74 @@ +/** + * metadata-filter.test-d.ts - Compile-time shape of the predicate grammar. + * MetadataFilter and MetadataMatch share one recursive grammar and differ in + * the conditions they admit. Checked by vitest's typecheck pass. + */ + +import { describe, test, expectTypeOf } from "vitest"; +import { + compileMetadataMatch, + parseMetadataFilter, + parseMetadataMatch, + type MetadataCondition, + type MetadataEntryCondition, + type MetadataFilter, + type MetadataMatch, + type MetadataPredicate, +} from "../src/metadata-filter.js"; +import type { ListMetadataOptions } from "../src/metadata-store.js"; + +describe("MetadataPredicate", () => { + test("filter and match are the one grammar over different conditions", () => { + expectTypeOf().toEqualTypeOf>(); + expectTypeOf().toEqualTypeOf>(); + expectTypeOf(parseMetadataFilter).returns.toEqualTypeOf(); + expectTypeOf(parseMetadataMatch).returns.toEqualTypeOf(); + expectTypeOf(compileMetadataMatch).parameter(0).toEqualTypeOf(); + expectTypeOf().toEqualTypeOf(); + }); + + test("a match composes with and, or, and not like a filter", () => { + const match: MetadataMatch = { + operator: "and", + operands: [ + { field: "key", operator: "prefix", value: "mem-" }, + { operator: "not", operand: { field: "value", operator: "type", value: "boolean" } }, + { operator: "or", operands: [ + { field: "value", operator: "gte", value: 3 }, + { field: "value", operator: "in", value: ["a", "b"], caseInsensitive: true }, + ] }, + ], + }; + expectTypeOf(match).toMatchTypeOf(); + }); + + test("a match names only the fields of a metadata entry", () => { + // @ts-expect-error a document's metadata key is not a field of an entry + const byMetadataKey: MetadataMatch = { field: "topics", operator: "eq", value: "typescript" }; + // @ts-expect-error nested in a group as well + const nested: MetadataMatch = { operator: "and", operands: [{ field: "topics", operator: "eq", value: "x" }] }; + expectTypeOf(byMetadataKey).toEqualTypeOf(); + expectTypeOf(nested).toEqualTypeOf(); + }); + + test("a match admits no condition about the set of values a document holds", () => { + // @ts-expect-error `exists` has no meaning for a single entry + const exists: MetadataMatch = { field: "value", operator: "exists", value: true }; + // @ts-expect-error `all` has no meaning for a single entry + const all: MetadataMatch = { field: "value", operator: "all", value: ["a"] }; + expectTypeOf(exists).toEqualTypeOf(); + expectTypeOf(all).toEqualTypeOf(); + }); + + test("a filter still admits every document condition", () => { + const filter: MetadataFilter = { + operator: "and", + operands: [ + { field: "topics", operator: "all", value: ["typescript", "sqlite"] }, + { field: "reviewed", operator: "exists", value: true }, + { field: "key", operator: "eq", value: "a metadata key may be named key" }, + ], + }; + expectTypeOf(filter).toMatchTypeOf(); + }); +}); diff --git a/test/metadata-filter.test.ts b/test/metadata-filter.test.ts index 6e36e921e..7df69e7dc 100644 --- a/test/metadata-filter.test.ts +++ b/test/metadata-filter.test.ts @@ -8,7 +8,10 @@ import { openDatabase } from "../src/db.js"; import type { Database } from "../src/db.js"; import { parseMetadataFilter, + parseMetadataMatch, compileMetadataFilter, + type FilterScope, + compileMetadataMatch, MetadataFilterError, METADATA_FILTER_LIMITS, type MetadataFilter, @@ -23,32 +26,63 @@ import { METADATA_EXTRACTION_VERSION, type DocumentMetadata } from "../src/metad describe("parseMetadataFilter", () => { test("accepts every condition operator shape", () => { const conditions: unknown[] = [ - { key: "status", operator: "eq", value: "published" }, - { key: "status", operator: "ne", value: "draft" }, - { key: "priority", operator: "gt", value: 3 }, - { key: "priority", operator: "gte", value: 3 }, - { key: "priority", operator: "lt", value: 10 }, - { key: "name", operator: "lte", value: "m" }, - { key: "topics", operator: "in", value: ["a", "b"] }, - { key: "topics", operator: "nin", value: [1, 2] }, - { key: "flags", operator: "all", value: [true, false] }, - { key: "status", operator: "exists", value: false }, + { field: "status", operator: "eq", value: "published" }, + { field: "status", operator: "ne", value: "draft" }, + { field: "priority", operator: "gt", value: 3 }, + { field: "priority", operator: "gte", value: 3 }, + { field: "priority", operator: "lt", value: 10 }, + { field: "name", operator: "lte", value: "m" }, + { field: "topics", operator: "in", value: ["a", "b"] }, + { field: "topics", operator: "nin", value: [1, 2] }, + { field: "flags", operator: "all", value: [true, false] }, + { field: "status", operator: "exists", value: false }, + { field: "topics", operator: "contains", value: "vec" }, + { field: "topics", operator: "prefix", value: "sql" }, + { field: "owner", operator: "suffix", value: "-team" }, + { field: "priority", operator: "type", value: "number" }, ]; for (const condition of conditions) { expect(parseMetadataFilter(condition)).toEqual(condition); } }); + test("accepts caseInsensitive on string operands only", () => { + const accepted: unknown[] = [ + { field: "status", operator: "eq", value: "Published", caseInsensitive: true }, + { field: "status", operator: "ne", value: "Draft", caseInsensitive: false }, + { field: "name", operator: "gt", value: "M", caseInsensitive: true }, + { field: "topics", operator: "in", value: ["SQL", "TypeScript"], caseInsensitive: true }, + { field: "topics", operator: "all", value: ["SQL"], caseInsensitive: true }, + { field: "topics", operator: "contains", value: "SQL", caseInsensitive: true }, + { field: "topics", operator: "prefix", value: "SQL", caseInsensitive: true }, + { field: "topics", operator: "suffix", value: "SQL", caseInsensitive: true }, + ]; + for (const condition of accepted) { + expect(parseMetadataFilter(condition)).toEqual(condition); + } + + const rejected: [unknown, RegExp][] = [ + [{ field: "priority", operator: "eq", value: 3, caseInsensitive: true }, /at \$\.caseInsensitive:.*string values only/], + [{ field: "reviewed", operator: "in", value: [true], caseInsensitive: true }, /string values only/], + [{ field: "reviewed", operator: "exists", value: true, caseInsensitive: true }, /does not apply to 'exists'/], + [{ field: "priority", operator: "type", value: "number", caseInsensitive: true }, /does not apply to 'type'/], + [{ field: "status", operator: "eq", value: "x", caseInsensitive: "yes" }, /must be a boolean/], + ]; + for (const [input, expected] of rejected) { + expect(() => parseMetadataFilter(input)).toThrow(expected); + } + }); + test("accepts nested groups and negation", () => { const filter = { operator: "and", operands: [ - { key: "topics", operator: "all", value: ["typescript", "programming"] }, + { field: "topics", operator: "all", value: ["typescript", "programming"] }, { operator: "or", operands: [ - { key: "status", operator: "eq", value: "published" }, - { operator: "not", operand: { key: "audience", operator: "eq", value: "internal" } }, + { field: "status", operator: "eq", value: "published" }, + { operator: "not", operand: { field: "audience", operator: "eq", value: "internal" } }, ], }, ], @@ -57,31 +91,37 @@ describe("parseMetadataFilter", () => { }); test("canonicalizes membership arrays by de-duplicating", () => { - const parsed = parseMetadataFilter({ key: "topics", operator: "in", value: ["a", "b", "a"] }); - expect(parsed).toEqual({ key: "topics", operator: "in", value: ["a", "b"] }); + const parsed = parseMetadataFilter({ field: "topics", operator: "in", value: ["a", "b", "a"] }); + expect(parsed).toEqual({ field: "topics", operator: "in", value: ["a", "b"] }); }); test("rejects invalid node shapes with the failing JSON path", () => { const cases: [unknown, RegExp][] = [ ["not-an-object", /at \$:.*must be an object/], - [{ key: "a", value: 1 }, /at \$:.*missing 'operator'/], - [{ operator: "equal", key: "a", value: 1 }, /unknown operator 'equal'/], + [{ field: "a", value: 1 }, /at \$:.*missing 'operator'/], + [{ operator: "equal", field: "a", value: 1 }, /unknown operator 'equal'/], [{ operator: "and", operands: [] }, /non-empty 'operands'/], [{ operator: "and", operands: "nope" }, /'operands' array/], - [{ operator: "not", operands: [{ key: "a", operator: "eq", value: 1 }] }, /unknown property 'operands'/], + [{ operator: "not", operands: [{ field: "a", operator: "eq", value: 1 }] }, /unknown property 'operands'/], [{ operator: "not" }, /exactly one 'operand'/], [{ operator: "and", operands: [{ operator: "eq" }] }, /at \$\.operands\[0\]/], - [{ operator: "eq", value: 1 }, /non-empty string 'key'/], - [{ operator: "eq", key: "a" }, /requires a 'value'/], - [{ operator: "eq", key: "a", value: 1, extra: true }, /unknown property 'extra'/], - [{ operator: "and", operands: [{ key: "a", operator: "eq", value: 1 }], key: "a" }, /unknown property 'key'/], - [{ operator: "gt", key: "a", value: true }, /string or number value/], - [{ operator: "eq", key: "a", value: NaN }, /finite/], - [{ operator: "eq", key: "a", value: { nested: 1 } }, /string, number, or boolean/], - [{ operator: "in", key: "a", value: "x" }, /array value/], - [{ operator: "in", key: "a", value: [] }, /non-empty array/], - [{ operator: "in", key: "a", value: [1, "two"] }, /homogeneous array/], - [{ operator: "exists", key: "a", value: "yes" }, /boolean value/], + [{ operator: "eq", value: 1 }, /non-empty string 'field'/], + [{ operator: "eq", field: "a" }, /requires a 'value'/], + [{ operator: "eq", field: "a", value: 1, extra: true }, /unknown property 'extra'/], + [{ operator: "and", operands: [{ field: "a", operator: "eq", value: 1 }], field: "a" }, /unknown property 'field'/], + [{ operator: "eq", key: "a", value: 1 }, /unknown property 'key'/], + [{ operator: "gt", field: "a", value: true }, /string or number value/], + [{ operator: "eq", field: "a", value: NaN }, /finite/], + [{ operator: "eq", field: "a", value: { nested: 1 } }, /string, number, or boolean/], + [{ operator: "in", field: "a", value: "x" }, /array value/], + [{ operator: "in", field: "a", value: [] }, /non-empty array/], + [{ operator: "in", field: "a", value: [1, "two"] }, /homogeneous array/], + [{ operator: "exists", field: "a", value: "yes" }, /boolean value/], + [{ operator: "contains", field: "a", value: "" }, /at \$\.value:.*'contains' requires a non-empty string/], + [{ operator: "prefix", field: "a", value: 3 }, /'prefix' requires a non-empty string/], + [{ operator: "suffix", field: "a", value: ["x"] }, /string, number, or boolean/], + [{ operator: "type", field: "a", value: "integer" }, /at \$\.value:.*'type' requires one of: string, number, boolean/], + [{ operator: "type", field: "a", value: 1 }, /'type' requires one of/], ]; for (const [input, expected] of cases) { @@ -91,13 +131,13 @@ describe("parseMetadataFilter", () => { }); test("rejects excessive depth, node count, and operand count", () => { - let deepFilter: unknown = { key: "a", operator: "eq", value: 1 }; + let deepFilter: unknown = { field: "a", operator: "eq", value: 1 }; for (let i = 0; i <= METADATA_FILTER_LIMITS.maxDepth; i++) { deepFilter = { operator: "not", operand: deepFilter }; } expect(() => parseMetadataFilter(deepFilter)).toThrow(/nesting depth/); - const condition = { key: "a", operator: "eq", value: 1 }; + const condition = { field: "a", operator: "eq", value: 1 }; const wideGroup = { operator: "or", operands: Array.from({ length: METADATA_FILTER_LIMITS.maxGroupOperands + 1 }, () => condition), @@ -114,7 +154,7 @@ describe("parseMetadataFilter", () => { expect(() => parseMetadataFilter(manyNodes)).toThrow(/nodes/); const manyValues = { - key: "a", + field: "a", operator: "in", value: Array.from({ length: METADATA_FILTER_LIMITS.maxMembershipValues + 1 }, (_, i) => i), }; @@ -122,6 +162,60 @@ describe("parseMetadataFilter", () => { }); }); +describe("parseMetadataMatch", () => { + test("accepts the filter grammar with 'key' and 'value' as the condition fields", () => { + const match = { + operator: "and", + operands: [ + { field: "key", operator: "prefix", value: "mem-" }, + { operator: "or", operands: [ + { field: "value", operator: "gte", value: 3 }, + { field: "value", operator: "type", value: "boolean" }, + { operator: "not", operand: { field: "value", operator: "in", value: ["Draft"], caseInsensitive: true } }, + ] }, + ], + }; + expect(parseMetadataMatch(match)).toEqual(match); + }); + + test("rejects other fields and the two set operators, naming the match", () => { + const cases: [unknown, RegExp][] = [ + [{ field: "topics", operator: "eq", value: "x" }, /^Invalid metadata match at \$: 'topics' is not a field of a metadata entry, expected 'key' or 'value'$/], + [{ field: "value", operator: "exists", value: true }, /^Invalid metadata match at \$: 'exists' has no meaning for a single metadata entry$/], + [{ field: "key", operator: "all", value: ["a"] }, /'all' has no meaning for a single metadata entry/], + [{ operator: "not", operand: { field: "value", operator: "eq" } }, /^Invalid metadata match at \$\.operand: 'eq' requires a 'value'$/], + ["nope", /^Invalid metadata match at \$: each filter node must be an object$/], + ]; + for (const [input, expected] of cases) { + expect(() => parseMetadataMatch(input)).toThrow(expected); + expect(() => parseMetadataMatch(input)).toThrow(MetadataFilterError); + } + // The filter keeps its own name. + expect(() => parseMetadataFilter("nope")).toThrow(/^Invalid metadata filter at \$:/); + }); +}); + +describe("compileMetadataMatch", () => { + test("compiles to a predicate over the row alias with every operand bound", () => { + const compiled = compileMetadataMatch(parseMetadataMatch({ + operator: "and", + operands: [ + { field: "key", operator: "eq", value: "k'; --" }, + { field: "value", operator: "suffix", value: "V'; --", caseInsensitive: true }, + { field: "value", operator: "nin", value: [1, 2] }, + ], + }), "row"); + + expect(compiled.sql).not.toContain("'; --"); + expect(compiled.sql).not.toContain("EXISTS"); + expect(compiled.sql).toContain("row.key = ?"); + expect(compiled.sql).toContain("lower(row.text_value)"); + expect(compiled.sql).toContain("row.number_value NOT IN (?, ?)"); + // The suffix binds its operand's UTF-8 byte length ahead of the operand. + expect(compiled.params).toEqual(["k'; --", 6, "v'; --", 1, 2]); + }); +}); + // ============================================================================= // SQL compilation and semantics // ============================================================================= @@ -152,8 +246,8 @@ describe("compileMetadataFilter semantics", () => { return documentId; } - function matchPaths(filter: MetadataFilter): string[] { - const compiled = compileMetadataFilter(parseMetadataFilter(filter), "d"); + function matchPathsIn(filter: MetadataFilter, scope: FilterScope): string[] { + const compiled = compileMetadataFilter(parseMetadataFilter(filter), "d", scope); const rows = db.prepare(` SELECT d.path FROM documents d JOIN document_metadata dm ON dm.document_id = d.id @@ -165,15 +259,22 @@ describe("compileMetadataFilter semantics", () => { return rows.map(row => row.path); } + /** Both scopes are one semantics in two SQL forms, so every case checks they agree. */ + function matchPaths(filter: MetadataFilter): string[] { + const candidates = matchPathsIn(filter, "candidates"); + expect(matchPathsIn(filter, "corpus"), `corpus scope for ${JSON.stringify(filter)}`).toEqual(candidates); + return candidates; + } + test("eq matches each scalar type exactly, without coercion", () => { insertDoc("str.md", { status: "published" }); insertDoc("num.md", { status: 1 }); insertDoc("bool.md", { status: true }); - expect(matchPaths({ key: "status", operator: "eq", value: "published" })).toEqual(["str.md"]); - expect(matchPaths({ key: "status", operator: "eq", value: 1 })).toEqual(["num.md"]); - expect(matchPaths({ key: "status", operator: "eq", value: true })).toEqual(["bool.md"]); - expect(matchPaths({ key: "status", operator: "eq", value: "1" })).toEqual([]); + expect(matchPaths({ field: "status", operator: "eq", value: "published" })).toEqual(["str.md"]); + expect(matchPaths({ field: "status", operator: "eq", value: 1 })).toEqual(["num.md"]); + expect(matchPaths({ field: "status", operator: "eq", value: true })).toEqual(["bool.md"]); + expect(matchPaths({ field: "status", operator: "eq", value: "1" })).toEqual([]); }); test("ne requires key presence and matching type", () => { @@ -182,7 +283,7 @@ describe("compileMetadataFilter semantics", () => { insertDoc("missing.md", { other: "x" }); insertDoc("typed.md", { status: 1 }); - expect(matchPaths({ key: "status", operator: "ne", value: "draft" })).toEqual(["published.md"]); + expect(matchPaths({ field: "status", operator: "ne", value: "draft" })).toEqual(["published.md"]); }); test("ordered comparisons match numbers and binary-ordered strings", () => { @@ -192,11 +293,11 @@ describe("compileMetadataFilter semantics", () => { insertDoc("alpha.md", { name: "alpha" }); insertDoc("zulu.md", { name: "zulu" }); - expect(matchPaths({ key: "priority", operator: "gte", value: 3 })).toEqual(["high.md", "mid.md"]); - expect(matchPaths({ key: "priority", operator: "lt", value: 3 })).toEqual(["low.md"]); - expect(matchPaths({ key: "name", operator: "gt", value: "alpha" })).toEqual(["zulu.md"]); + expect(matchPaths({ field: "priority", operator: "gte", value: 3 })).toEqual(["high.md", "mid.md"]); + expect(matchPaths({ field: "priority", operator: "lt", value: 3 })).toEqual(["low.md"]); + expect(matchPaths({ field: "name", operator: "gt", value: "alpha" })).toEqual(["zulu.md"]); // Type mismatch: no string 'priority' values exist. - expect(matchPaths({ key: "priority", operator: "gte", value: "3" })).toEqual([]); + expect(matchPaths({ field: "priority", operator: "gte", value: "3" })).toEqual([]); }); test("in, nin, and all evaluate membership over value sets", () => { @@ -204,18 +305,18 @@ describe("compileMetadataFilter semantics", () => { insertDoc("go.md", { topics: ["go"] }); insertDoc("none.md", { other: "x" }); - expect(matchPaths({ key: "topics", operator: "in", value: ["typescript", "rust"] })).toEqual(["ts.md"]); - expect(matchPaths({ key: "topics", operator: "nin", value: ["typescript", "rust"] })).toEqual(["go.md"]); - expect(matchPaths({ key: "topics", operator: "all", value: ["typescript", "programming"] })).toEqual(["ts.md"]); - expect(matchPaths({ key: "topics", operator: "all", value: ["typescript", "rust"] })).toEqual([]); + expect(matchPaths({ field: "topics", operator: "in", value: ["typescript", "rust"] })).toEqual(["ts.md"]); + expect(matchPaths({ field: "topics", operator: "nin", value: ["typescript", "rust"] })).toEqual(["go.md"]); + expect(matchPaths({ field: "topics", operator: "all", value: ["typescript", "programming"] })).toEqual(["ts.md"]); + expect(matchPaths({ field: "topics", operator: "all", value: ["typescript", "rust"] })).toEqual([]); }); test("exists matches presence and absence", () => { insertDoc("has.md", { status: "ok" }); insertDoc("hasnt.md", { other: "x" }); - expect(matchPaths({ key: "status", operator: "exists", value: true })).toEqual(["has.md"]); - expect(matchPaths({ key: "status", operator: "exists", value: false })).toEqual(["hasnt.md"]); + expect(matchPaths({ field: "status", operator: "exists", value: true })).toEqual(["has.md"]); + expect(matchPaths({ field: "status", operator: "exists", value: false })).toEqual(["hasnt.md"]); }); test("and, or, and not compose recursively", () => { @@ -226,22 +327,22 @@ describe("compileMetadataFilter semantics", () => { expect(matchPaths({ operator: "and", operands: [ - { key: "status", operator: "eq", value: "published" }, - { key: "priority", operator: "gte", value: 3 }, + { field: "status", operator: "eq", value: "published" }, + { field: "priority", operator: "gte", value: 3 }, ], })).toEqual(["a.md"]); expect(matchPaths({ operator: "or", operands: [ - { key: "priority", operator: "gte", value: 9 }, - { key: "priority", operator: "lte", value: 1 }, + { field: "priority", operator: "gte", value: 9 }, + { field: "priority", operator: "lte", value: 1 }, ], })).toEqual(["b.md", "c.md"]); expect(matchPaths({ operator: "not", - operand: { key: "status", operator: "eq", value: "draft" }, + operand: { field: "status", operator: "eq", value: "draft" }, })).toEqual(["a.md", "b.md"]); }); @@ -252,8 +353,8 @@ describe("compileMetadataFilter semantics", () => { const rangeFilter: MetadataFilter = { operator: "and", operands: [ - { key: "priority", operator: "gte", value: 3 }, - { key: "priority", operator: "lt", value: 10 }, + { field: "priority", operator: "gte", value: 3 }, + { field: "priority", operator: "lt", value: 10 }, ], }; // Document-level semantics: [1, 20] satisfies both conditions via @@ -261,19 +362,106 @@ describe("compileMetadataFilter semantics", () => { expect(matchPaths(rangeFilter)).toEqual(["narrow.md", "wide.md"]); }); + test("text operators match any string element and never a number or boolean", () => { + insertDoc("sqlite.md", { topics: ["sqlite", "search"] }); + insertDoc("vec.md", { topics: ["sqlite-vec", "embeddings"] }); + insertDoc("pg.md", { topics: "postgres" }); + insertDoc("num.md", { topics: 42 }); + insertDoc("bool.md", { topics: true }); + + expect(matchPaths({ field: "topics", operator: "prefix", value: "sql" })).toEqual(["sqlite.md", "vec.md"]); + expect(matchPaths({ field: "topics", operator: "suffix", value: "vec" })).toEqual(["vec.md"]); + expect(matchPaths({ field: "topics", operator: "contains", value: "arch" })).toEqual(["sqlite.md"]); + expect(matchPaths({ field: "topics", operator: "contains", value: "4" })).toEqual([]); + expect(matchPaths({ field: "topics", operator: "prefix", value: "tru" })).toEqual([]); + }); + + test("prefix and suffix compare whole characters and never overrun the value", () => { + insertDoc("exact.md", { code: "abc" }); + insertDoc("emoji.md", { code: "\u{1F600}abc" }); + + // The operand equal to the value is both its prefix and its suffix. + expect(matchPaths({ field: "code", operator: "prefix", value: "abc" })).toEqual(["exact.md"]); + expect(matchPaths({ field: "code", operator: "suffix", value: "abc" })).toEqual(["emoji.md", "exact.md"]); + // An operand longer than the value cannot match either end. + expect(matchPaths({ field: "code", operator: "prefix", value: "abcd" })).toEqual([]); + expect(matchPaths({ field: "code", operator: "suffix", value: "zabc" })).toEqual([]); + // Astral characters count as one character on both sides. + expect(matchPaths({ field: "code", operator: "prefix", value: "\u{1F600}a" })).toEqual(["emoji.md"]); + expect(matchPaths({ field: "code", operator: "suffix", value: "\u{1F600}abc" })).toEqual(["emoji.md"]); + }); + + test("text operators see the whole value past an embedded NUL", () => { + insertDoc("nul.md", { code: "abc\u0000XYZ" }); + insertDoc("plain.md", { code: "abc" }); + + expect(matchPaths({ field: "code", operator: "suffix", value: "XYZ" })).toEqual(["nul.md"]); + expect(matchPaths({ field: "code", operator: "suffix", value: "abc" })).toEqual(["plain.md"]); + expect(matchPaths({ field: "code", operator: "prefix", value: "abc\u0000XYZ" })).toEqual(["nul.md"]); + expect(matchPaths({ field: "code", operator: "suffix", value: "abc\u0000XYZ" })).toEqual(["nul.md"]); + expect(matchPaths({ field: "code", operator: "prefix", value: "abc\u0000" })).toEqual(["nul.md"]); + expect(matchPaths({ field: "code", operator: "contains", value: "\u0000X" })).toEqual(["nul.md"]); + expect(matchPaths({ field: "code", operator: "suffix", value: "xyz", caseInsensitive: true })).toEqual(["nul.md"]); + expect(matchPaths({ field: "code", operator: "eq", value: "abc\u0000XYZ" })).toEqual(["nul.md"]); + }); + + test("type matches the stored type of a key's values", () => { + insertDoc("num.md", { priority: 3 }); + insertDoc("nums.md", { priority: [1, 2] }); + insertDoc("str.md", { priority: "high" }); + insertDoc("bool.md", { priority: true }); + insertDoc("none.md", { other: 1 }); + + expect(matchPaths({ field: "priority", operator: "type", value: "number" })).toEqual(["num.md", "nums.md"]); + expect(matchPaths({ field: "priority", operator: "type", value: "string" })).toEqual(["str.md"]); + expect(matchPaths({ field: "priority", operator: "type", value: "boolean" })).toEqual(["bool.md"]); + expect(matchPaths({ + operator: "not", + operand: { field: "priority", operator: "type", value: "string" }, + })).toEqual(["bool.md", "none.md", "num.md", "nums.md"]); + }); + + test("caseInsensitive folds ASCII letters on both sides for every string operator", () => { + insertDoc("upper.md", { status: "PUBLISHED", topics: ["PostgreSQL", "MySQL"] }); + insertDoc("lower.md", { status: "published", topics: ["sqlite"] }); + insertDoc("draft.md", { status: "Draft", topics: ["Search"] }); + + expect(matchPaths({ field: "status", operator: "eq", value: "Published" })).toEqual([]); + expect(matchPaths({ field: "status", operator: "eq", value: "Published", caseInsensitive: true })).toEqual(["lower.md", "upper.md"]); + expect(matchPaths({ field: "status", operator: "ne", value: "published", caseInsensitive: true })).toEqual(["draft.md"]); + expect(matchPaths({ field: "status", operator: "in", value: ["DRAFT"], caseInsensitive: true })).toEqual(["draft.md"]); + expect(matchPaths({ field: "status", operator: "nin", value: ["DRAFT"], caseInsensitive: true })).toEqual(["lower.md", "upper.md"]); + expect(matchPaths({ field: "topics", operator: "all", value: ["postgresql", "mysql"], caseInsensitive: true })).toEqual(["upper.md"]); + expect(matchPaths({ field: "topics", operator: "contains", value: "sql", caseInsensitive: true })).toEqual(["lower.md", "upper.md"]); + expect(matchPaths({ field: "topics", operator: "prefix", value: "postgres", caseInsensitive: true })).toEqual(["upper.md"]); + expect(matchPaths({ field: "topics", operator: "suffix", value: "SQL", caseInsensitive: true })).toEqual(["upper.md"]); + // Lexical comparison folds too: "Draft" sorts before "published" once lowered. + expect(matchPaths({ field: "status", operator: "lt", value: "M", caseInsensitive: true })).toEqual(["draft.md"]); + }); + + test("caseInsensitive leaves non-ASCII letters exact", () => { + insertDoc("upper.md", { city: "\u00C9VORA" }); + insertDoc("lower.md", { city: "\u00E9vora" }); + + // The ASCII part folds and the accented initial does not, so each + // spelling matches itself only. + expect(matchPaths({ field: "city", operator: "eq", value: "\u00C9vora", caseInsensitive: true })).toEqual(["upper.md"]); + expect(matchPaths({ field: "city", operator: "eq", value: "\u00E9VORA", caseInsensitive: true })).toEqual(["lower.md"]); + }); + test("boolean values round-trip through membership operators", () => { insertDoc("flagged.md", { reviewed: true }); insertDoc("unflagged.md", { reviewed: false }); - expect(matchPaths({ key: "reviewed", operator: "in", value: [true] })).toEqual(["flagged.md"]); - expect(matchPaths({ key: "reviewed", operator: "nin", value: [true] })).toEqual(["unflagged.md"]); + expect(matchPaths({ field: "reviewed", operator: "in", value: [true] })).toEqual(["flagged.md"]); + expect(matchPaths({ field: "reviewed", operator: "nin", value: [true] })).toEqual(["unflagged.md"]); }); test("SQL injection payloads in keys and values stay data", () => { insertDoc("safe.md", { "key'; DROP TABLE documents; --": "v'; DROP TABLE documents; --" }); expect(matchPaths({ - key: "key'; DROP TABLE documents; --", + field: "key'; DROP TABLE documents; --", operator: "eq", value: "v'; DROP TABLE documents; --", })).toEqual(["safe.md"]); @@ -282,12 +470,89 @@ describe("compileMetadataFilter semantics", () => { expect((db.prepare(`SELECT COUNT(*) as c FROM documents`).get() as { c: number }).c).toBe(1); }); + test("every condition seeks by value or by document, with and without statistics", () => { + // The covering indexes are (key, , document_id). An equality on the + // value column is one probe by value. Everything else must seek the + // document's rows through the primary key, or the planner walks the key's + // whole value range once per document. `ne` and `nin` pair one of each. + // ANALYZE must not change any of it (below about 32 rows, statistics + // make a full scan of the tiny index the cheaper plan, correctly). + for (let index = 0; index < 64; index++) { + insertDoc(`d${index}.md`, { status: index % 2 ? "published" : "draft", priority: index, topics: [`t${index}`] }); + } + + // SQLite 3.43 labels the same point probe USING INDEX, without COVERING. + const byValue = "INDEX idx_metadata_"; + const byDocument = "INDEX sqlite_autoindex_document_metadata_values_1 (document_id=?)"; + const expectations: [MetadataFilter, string][] = [ + [{ field: "status", operator: "eq", value: "published" }, byValue], + [{ field: "topics", operator: "in", value: ["t1", "t2"] }, byValue], + [{ field: "topics", operator: "all", value: ["t1"] }, byValue], + [{ field: "status", operator: "eq", value: "PUBLISHED", caseInsensitive: true }, byDocument], + [{ field: "priority", operator: "gt", value: 3 }, byDocument], + [{ field: "priority", operator: "lte", value: 3 }, byDocument], + [{ field: "status", operator: "ne", value: "draft" }, byDocument], + [{ field: "topics", operator: "nin", value: ["t1"] }, byDocument], + [{ field: "status", operator: "exists", value: true }, byDocument], + [{ field: "status", operator: "type", value: "string" }, byDocument], + [{ field: "topics", operator: "prefix", value: "t" }, byDocument], + [{ field: "topics", operator: "contains", value: "1" }, byDocument], + [{ field: "topics", operator: "suffix", value: "1", caseInsensitive: true }, byDocument], + ]; + + for (const statistics of [false, true]) { + if (statistics) db.exec("ANALYZE"); + for (const [filter, seek] of expectations) { + const compiled = compileMetadataFilter(parseMetadataFilter(filter), "d"); + const plan = (db.prepare(`EXPLAIN QUERY PLAN SELECT COUNT(*) FROM documents d WHERE ${compiled.sql}`) + .all(...compiled.params) as { detail: string }[]) + .map(row => row.detail) + .filter(detail => detail.includes(" mv ")); + const label = `${filter.operator}${"caseInsensitive" in filter ? " folded" : ""}${statistics ? " with statistics" : ""}`; + expect(plan.some(detail => detail.includes(seek)), `${label}: ${plan.join(" | ")}`).toBe(true); + expect(plan.some(detail => detail.includes("key=?") && !detail.includes("document_id=?")), `${label} walks a key range: ${plan.join(" | ")}`).toBe(false); + } + } + }); + + test("the corpus scope compiles each condition to one uncorrelated document set", () => { + for (let index = 0; index < 64; index++) { + insertDoc(`d${index}.md`, { status: index % 2 ? "published" : "draft", priority: index }); + } + const filter = parseMetadataFilter({ + operator: "and", + operands: [ + { field: "priority", operator: "gt", value: 3 }, + { operator: "not", operand: { field: "status", operator: "exists", value: false } }, + ], + }); + const compiled = compileMetadataFilter(filter, "d", "corpus"); + expect(compiled.sql).toBe( + "(d.id IN (SELECT mv.document_id FROM document_metadata_values mv WHERE mv.key = ? AND mv.value_type = 'number' AND mv.number_value > ?)" + + " AND NOT NOT d.id IN (SELECT mv.document_id FROM document_metadata_values mv WHERE mv.key = ?))", + ); + expect(compiled.params).toEqual(["priority", 3, "status"]); + + for (const statistics of [false, true]) { + if (statistics) db.exec("ANALYZE"); + const plan = (db.prepare(`EXPLAIN QUERY PLAN SELECT COUNT(*) FROM documents d WHERE ${compiled.sql}`) + .all(...compiled.params) as { detail: string }[]) + .map(row => row.detail); + // The document sets are built once (LIST SUBQUERY), by covering index, and probed per document. + expect(plan.filter(step => step.includes("LIST SUBQUERY"))).toHaveLength(2); + expect(plan.some(step => step.includes(" EXISTS ")), plan.join(" | ")).toBe(false); + expect(plan.some(step => step.includes("idx_metadata_number_lookup (key=? AND number_value>?)")), plan.join(" | ")).toBe(true); + } + }); + test("compiled SQL never interpolates user keys or values", () => { const filter = parseMetadataFilter({ operator: "and", operands: [ - { key: "key'; --", operator: "eq", value: "value'; --" }, - { key: "topics", operator: "in", value: ["a'; --"] }, + { field: "key'; --", operator: "eq", value: "value'; --" }, + { field: "topics", operator: "in", value: ["a'; --"] }, + { field: "topics", operator: "contains", value: "c'; --" }, + { field: "topics", operator: "suffix", value: "S'; --", caseInsensitive: true }, ], }); const compiled = compileMetadataFilter(filter, "d"); @@ -295,5 +560,7 @@ describe("compileMetadataFilter semantics", () => { expect(compiled.params).toContain("key'; --"); expect(compiled.params).toContain("value'; --"); expect(compiled.params).toContain("a'; --"); + expect(compiled.params).toContain("c'; --"); + expect(compiled.params).toContain("s'; --"); }); }); diff --git a/test/metadata-format.test.ts b/test/metadata-format.test.ts new file mode 100644 index 000000000..87250ecca --- /dev/null +++ b/test/metadata-format.test.ts @@ -0,0 +1,156 @@ +/** + * Metadata discovery rendering: the plain-text shape the CLI and the MCP + * `metadata` tool share, exercised directly where the CLI fixture cannot + * reach (string identity, control characters, empty states under a filter). + */ + +import * as childProcess from "node:child_process"; + +import { describe, test, expect } from "vitest"; +import { formatMetadataKeySummaries, formatMetadataOverview, type FormatMetadataOptions } from "../src/metadata-format.js"; +import type { ListMetadataResult, MetadataKeyTypeSummary, MetadataScalar } from "../src/metadata-store.js"; + +const OPTIONS: FormatMetadataOptions = { + valueWindowHint: "--value-limit ", + keyWindowHint: "--key-limit ", + keyOffset: 0, + keyOffsetLabel: "--key-offset", + emptyMessage: "No metadata matches.", +}; + +function stringTypeSummary(values: string[], remainingValues = 0): MetadataKeyTypeSummary { + return { + type: "string", + multiValued: false, + documents: values.length, + distinctValues: values.length + remainingValues, + values: values.map(value => ({ value, documents: 1 })), + remainingValues, + collections: ["notes"], + }; +} + +function stringResult(values: string[]): ListMetadataResult { + return { documents: values.length, totalKeys: 1, keys: [{ key: "label", documents: values.length, types: [stringTypeSummary(values)] }], remainingKeys: 0 }; +} + +/** A key held as a string by some documents and a number by one, which prints the compact `value (count)` list. */ +function splitResult(values: string[], remainingValues = 0): ListMetadataResult { + const numberSummary: MetadataKeyTypeSummary = { + type: "number", multiValued: false, documents: 1, distinctValues: 1, + values: [{ value: 7, documents: 1 }], remainingValues: 0, range: { min: 7, median: 7, max: 7 }, collections: ["notes"], + }; + const documents = values.length + 1; + return { documents, totalKeys: 1, keys: [{ key: "label", documents, types: [stringTypeSummary(values, remainingValues), numberSummary] }], remainingKeys: 0 }; +} + +/** + * The items of the compact list on the string row of a type split, read back + * the way a reader must: a quoted item is a JSON string (so it may contain the + * delimiters), a bare item runs to its count, and the tail closes the list. + */ +function compactItems(output: string): { value: MetadataScalar; documents: number }[] { + const row = output.split("\n").find(line => line.startsWith(" string "))!; + const list = row.replace(/^ {2}string {2}\d+ docs? {2}/u, ""); + const items: { value: MetadataScalar; documents: number }[] = []; + let rest = list; + while (rest !== "") { + const tail = /^(\d+) more$/u.exec(rest); + if (tail) { + items.push({ value: `${tail[1]} more`, documents: -1 }); + break; + } + const item = rest.startsWith('"') + ? /^("(?:[^"\\]|\\.)*") \((\d+)\)(?:, |$)/u.exec(rest)! + : /^([^,()]*) \((\d+)\)(?:, |$)/u.exec(rest)!; + const printed = item[1]!; + items.push({ value: printed.startsWith('"') ? JSON.parse(printed) as MetadataScalar : printed, documents: Number(item[2]) }); + rest = rest.slice(item[0].length); + } + return items; +} + +/** The printed value column of each row: everything before the two-space gap and the count. */ +function printedValues(output: string): string[] { + return output.split("\n").slice(1).filter(line => line !== "").map(line => line.slice(2).replace(/ {2,}\d+$/, "")); +} + +test.skipIf(process.platform === "win32")("a type-conflict hint preserves apostrophes and shell syntax in the metadata key", () => { + const result = splitResult(["a", "b"]); + const key = "owner'$(printf injected)'s"; + result.keys[0]!.key = key; + const output = formatMetadataOverview(result, { documentsWithMetadata: result.documents, pendingMetadata: 0, drillDownHint: "qmd collection metadata notes" }); + const command = output.split("types disagree, see: ")[1]!; + + // Stub qmd to capture its arguments. The fixture's substitution is harmless + // if quoting regresses, but it must remain literal data in the JSON operand. + const argumentsText = childProcess.execFileSync("sh", ["-c", `qmd() { printf '%s\\n' "$@"; }\n${command}`], { encoding: "utf8" }); + const args = argumentsText.trimEnd().split("\n"); + expect(args.slice(0, 4)).toEqual(["collection", "metadata", "notes", "--match"]); + expect(JSON.parse(args[4]!)).toEqual({ field: "key", operator: "eq", value: key }); +}); + +describe("formatMetadataKeySummaries values", () => { + test("prints a plain string bare", () => { + expect(printedValues(formatMetadataKeySummaries(stringResult(["published", "docs team", "v1.2-rc", "é"]), OPTIONS))) + .toEqual(["published", "docs team", "v1.2-rc", "é"]); + }); + + test("quotes every string whose bare form could be read as another value or as layout", () => { + const ambiguous = ["", '""', " padded", "padded ", "two spaces", "tab\tin", "line\nbreak", "42", "-1.5", "true", "null", "[1]", "back\\slash"]; + + const printed = printedValues(formatMetadataKeySummaries(stringResult(ambiguous), OPTIONS)); + + expect(printed.every(value => value.startsWith('"') && value.endsWith('"'))).toBe(true); + expect(printed.map(value => JSON.parse(value) as MetadataScalar)).toEqual(ambiguous); + }); + + test("escapes every control character, including the ones JSON.stringify leaves literal", () => { + const controls = ["bell\u0007", "del\u007f", "csi\u009b", "line\u2028sep", "para\u2029sep", "nul\u0000"]; + + const output = formatMetadataKeySummaries(stringResult(controls), OPTIONS); + + expect(output).not.toMatch(/[\u0000-\u0008\u000b-\u001f\u007f-\u009f\u2028\u2029]/u); + expect(printedValues(output).map(value => JSON.parse(value) as MetadataScalar)).toEqual(controls); + }); + + test("quotes a string that the compact list's delimiters or tails would otherwise absorb", () => { + // "a (1), b" and the three items a, b, c would print identically without quotes. + const values = ["a (1), b", "c", "3 more", "...", "(x)", "plain value"]; + + const output = formatMetadataKeySummaries(splitResult(values, 2), OPTIONS); + + expect(output.split("\n")[1]).toBe(' string 6 docs "a (1), b" (1), c (1), "3 more" (1), "..." (1), "(x)" (1), plain value (1), 2 more'); + expect(compactItems(output)).toEqual([ + ...values.map(value => ({ value, documents: 1 })), + { value: "2 more", documents: -1 }, + ]); + // The same value prints the same way in the vertical layout. + expect(printedValues(formatMetadataKeySummaries(stringResult(values), OPTIONS))).toEqual(['"a (1), b"', "c", '"3 more"', '"..."', '"(x)"', "plain value"]); + }); + + test("prints numbers and booleans bare", () => { + const typeSummary: MetadataKeyTypeSummary = { + type: "number", multiValued: false, documents: 2, distinctValues: 2, + values: [{ value: 1.5, documents: 1 }, { value: -2, documents: 1 }], + remainingValues: 0, range: { min: -2, median: -0.25, max: 1.5 }, collections: ["notes"], + }; + const result: ListMetadataResult = { documents: 2, totalKeys: 1, keys: [{ key: "n", documents: 2, types: [typeSummary] }], remainingKeys: 0 }; + + expect(formatMetadataKeySummaries(result, OPTIONS)).toBe("n number 2 of 2 documents 2 distinct\n min -2 median -0.25 max 1.5\n -2 (1) 1.5 (1)"); + }); +}); + +describe("formatMetadataKeySummaries empty states", () => { + test("keeps the filter line when no entry matched", () => { + const result: ListMetadataResult = { documents: 3, filteredDocuments: 1, totalKeys: 0, keys: [], remainingKeys: 0 }; + + expect(formatMetadataKeySummaries(result, OPTIONS)).toBe("filter: 1 of 3 documents\n\nNo metadata matches."); + }); + + test("keeps the filter line when the key window falls past the end", () => { + const result: ListMetadataResult = { documents: 3, filteredDocuments: 2, totalKeys: 4, keys: [], remainingKeys: 0 }; + + expect(formatMetadataKeySummaries(result, { ...OPTIONS, keyOffset: 10 })).toBe("filter: 2 of 3 documents\n\nNo keys at --key-offset 10, 4 keys in total."); + }); +}); diff --git a/test/metadata-search.test.ts b/test/metadata-search.test.ts index 6e4b1741b..a97876e72 100644 --- a/test/metadata-search.test.ts +++ b/test/metadata-search.test.ts @@ -22,7 +22,8 @@ import { } from "../src/store.js"; import { replaceDocumentMetadata, syncDocumentMetadata } from "../src/metadata-store.js"; import { METADATA_EXTRACTION_VERSION, type DocumentMetadata } from "../src/metadata.js"; -import type { MetadataFilter } from "../src/metadata-filter.js"; +import { parseMetadataFilter, type MetadataFilter } from "../src/metadata-filter.js"; +import type { Database, SQLiteValue } from "../src/db.js"; let testDir: string; let store: Store; @@ -71,12 +72,30 @@ describe("searchFTS with metadata filter", () => { expect(unfiltered.length).toBe(3); const filtered = searchFTS(store.db, "authentication", 10, undefined, { - key: "status", operator: "eq", value: "published", + field: "status", operator: "eq", value: "published", }); expect(filtered.map(r => r.displayPath)).toEqual(["notes/published.md"]); expect(filtered[0]!.metadata).toEqual({ status: "published" }); }); + test("finds a matching document ranked below the unfiltered candidate window", async () => { + // limit=5 made the old filtered window 50; every noise doc outranks the target. + for (let i = 0; i < 50; i++) { + await insertDoc("notes", `noise-${i}.md`, `# N${i}\n\nalpha alpha alpha`, { status: "draft" }); + } + await insertDoc( + "notes", + "target.md", + `# Target\n\n${"Unrelated prose. ".repeat(40)}One weaker mention of alpha.`, + { status: "published" }, + ); + + const filtered = searchFTS(store.db, "alpha", 5, undefined, { + field: "status", operator: "eq", value: "published", + }); + expect(filtered.map(r => r.displayPath)).toEqual(["notes/target.md"]); + }); + test("excludes pending, stale, and errored documents from filtered search", async () => { const { documentId: erroredId } = await insertDoc("notes", "errored.md", "# One\n\ncommon term"); await insertDoc("notes", "pending.md", "# Two\n\ncommon term"); @@ -95,7 +114,7 @@ describe("searchFTS with metadata filter", () => { // An unprocessed document must not satisfy `exists: false`. const filtered = searchFTS(store.db, "common", 10, undefined, { - key: "status", operator: "exists", value: false, + field: "status", operator: "exists", value: false, }); expect(filtered.map(r => r.displayPath)).toEqual(["notes/extracted.md"]); @@ -110,7 +129,7 @@ describe("searchFTS with metadata filter", () => { await insertDoc("other", "d.md", "# D\n\nshared topic", { status: "published" }); const filtered = searchFTS(store.db, "shared", 10, ["notes", "docs"], { - key: "status", operator: "eq", value: "published", + field: "status", operator: "eq", value: "published", }); expect(filtered.map(r => r.displayPath).sort()).toEqual(["docs/b.md", "notes/a.md"]); }); @@ -129,12 +148,12 @@ describe("searchFTS with metadata filter", () => { const filter: MetadataFilter = { operator: "and", operands: [ - { key: "topics", operator: "all", value: ["typescript", "programming"] }, + { field: "topics", operator: "all", value: ["typescript", "programming"] }, { operator: "or", operands: [ - { key: "status", operator: "eq", value: "published" }, - { key: "priority", operator: "gte", value: 3 }, + { field: "status", operator: "eq", value: "published" }, + { field: "priority", operator: "gte", value: 3 }, ], }, ], @@ -151,7 +170,7 @@ describe("searchFTS with metadata filter", () => { } const filtered = searchFTS(store.db, "repeated keyword", 5, undefined, { - key: "status", operator: "eq", value: "published", + field: "status", operator: "eq", value: "published", }); expect(filtered.map(r => r.displayPath)).toEqual(["notes/doc-0.md"]); }); @@ -182,19 +201,50 @@ describe("searchVec with metadata filter", () => { const filtered = await searchVec( store.db, "q", model, 10, undefined, undefined, queryEmbedding, undefined, - { key: "status", operator: "eq", value: "published" }, + { field: "status", operator: "eq", value: "published" }, ); expect(filtered.map(r => r.displayPath)).toEqual(["notes/far-published.md"]); expect(filtered[0]!.metadata).toEqual({ status: "published" }); }); + test("a near-ceiling filter fits both vector lookup paths and still excludes shared nonmatching copies", async () => { + store.ensureVecTable(3); + const body = "# Book\n\nDeterministic vector fixture"; + const { hash } = await insertDoc("notes", "published.md", body, { eligible: true }); + await insertDoc("notes", "draft-copy.md", body, { eligible: false }); + const timestamp = new Date().toISOString(); + + store.db.transaction(() => { + for (let sequence = 0; sequence < 20_000; sequence++) { + insertEmbedding(store.db, hash, sequence, 0, new Float32Array([1, sequence / 20_001, 0]), model, timestamp, 20_001); + } + })(); + + // 256 nodes, 64 distinct members per wide leaf: a valid filter with + // 31,490 bindings. Candidate IDs must not exhaust the remaining budget. + const members = Array.from({ length: 64 }, (_, index) => `value-${index}`); + const leaves = Array.from({ length: 246 }, () => ({ field: "absent", operator: "all", value: members })); + const groups = Array.from({ length: 8 }, (_, index) => ({ operator: "or", operands: leaves.slice(index * 32, (index + 1) * 32) })); + const metadataFilter = parseMetadataFilter({ operator: "or", operands: [{ field: "eligible", operator: "eq", value: true }, ...groups] }); + + // At 20,000 eligible chunks the exact path returns up to limit * 3 IDs. + const exactResults = await searchVec(store.db, "q", model, 500, "notes", undefined, queryEmbedding, undefined, metadataFilter); + expect(exactResults.map(result => result.displayPath)).toEqual(["notes/published.md"]); + + // The extra chunk crosses into capped global lookup. Its 4,096 candidate + // IDs used to push the final document lookup past Node's variable limit. + insertEmbedding(store.db, hash, 20_000, 0, new Float32Array([1, 1, 0]), model, timestamp, 20_001); + const fallbackResults = await searchVec(store.db, "q", model, 137, "notes", undefined, queryEmbedding, undefined, metadataFilter); + expect(fallbackResults.map(result => result.displayPath)).toEqual(["notes/published.md"]); + }); + test("returns empty when no documents are eligible", async () => { store.ensureVecTable(3); await insertEmbeddedDoc("notes", "doc.md", "# Doc", [1, 0, 0], { status: "draft" }); const filtered = await searchVec( store.db, "q", model, 10, undefined, undefined, queryEmbedding, undefined, - { key: "status", operator: "eq", value: "published" }, + { field: "status", operator: "eq", value: "published" }, ); expect(filtered).toEqual([]); }); @@ -218,7 +268,7 @@ describe("searchVec with metadata filter", () => { const filtered = await searchVec( store.db, "q", model, 10, undefined, undefined, queryEmbedding, undefined, - { key: "status", operator: "eq", value: "published" }, + { field: "status", operator: "eq", value: "published" }, ); expect(filtered.map(r => r.displayPath)).toEqual(["notes/published-copy.md"]); }); @@ -230,10 +280,116 @@ describe("searchVec with metadata filter", () => { const filtered = await searchVec( store.db, "q", model, 10, "docs", undefined, queryEmbedding, undefined, - { key: "status", operator: "eq", value: "published" }, + { field: "status", operator: "eq", value: "published" }, ); expect(filtered.map(r => r.displayPath)).toEqual(["docs/b.md"]); }); + + const eligibleOnly: MetadataFilter = { field: "eligible", operator: "eq", value: true }; + + /** An eligible document and an ineligible copy of its content in one collection, one chunk per vector. */ + async function insertLongDocumentWithExcludedCopy(vectors: number[][]): Promise { + const body = "# Many chunks"; + const { hash } = await insertDoc("book", "many-chunks.md", body, { eligible: true }); + await insertDoc("book", "excluded-copy.md", body, { eligible: false }); + const now = new Date().toISOString(); + vectors.forEach((vector, seq) => insertEmbedding(store.db, hash, seq, seq * 100, new Float32Array(vector), model, now, vectors.length)); + } + + test("a document with many close chunks does not starve another eligible document", async () => { + store.ensureVecTable(3); + await insertLongDocumentWithExcludedCopy(Array.from({ length: 20 }, () => [1, 0, 0])); + await insertEmbeddedDoc("book", "second-document.md", "# Second document", [0, 1, 0], { eligible: true }); + + const results = await searchVec(store.db, "q", model, 2, undefined, undefined, queryEmbedding, undefined, eligibleOnly); + expect(results.map(r => r.displayPath)).toEqual(["book/many-chunks.md", "book/second-document.md"]); + }); + + test.each([20, 450])("the best chunk of a %i-chunk document survives deduplication", async (chunks) => { + store.ensureVecTable(3); + await insertLongDocumentWithExcludedCopy( + Array.from({ length: chunks }, (_, seq) => (seq === chunks - 1 ? [1, 0, 0] : [0, 0, 1])), + ); + await insertEmbeddedDoc("book", "second-document.md", "# Second document", [-1, 0.1, 0], { eligible: true }); + + const results = await searchVec(store.db, "q", model, 2, undefined, undefined, queryEmbedding, undefined, eligibleOnly); + expect(results.map(r => r.displayPath)).toEqual(["book/many-chunks.md", "book/second-document.md"]); + expect(results[0]!.chunkPos).toBe((chunks - 1) * 100); + }); + + test("a filter admitting more than 20,000 chunks still returns its nearest eligible documents", async () => { + store.ensureVecTable(3); + store.db.exec("BEGIN"); + for (let i = 0; i < 20_001; i++) { + await insertEmbeddedDoc("book", `eligible-${i}.md`, `# Eligible ${i}`, [0, 1, 0], { eligible: true }); + } + for (let i = 0; i < 200; i++) { + await insertEmbeddedDoc("book", `closer-${i}.md`, `# Closer ${i}`, [1, 0, 0], { eligible: false }); + } + store.db.exec("COMMIT"); + + const filtered = await searchVec( + store.db, "q", model, 5, "book", undefined, queryEmbedding, undefined, + { field: "eligible", operator: "eq", value: true }, + ); + expect(filtered).toHaveLength(5); + expect(filtered.every(r => r.metadata.eligible === true)).toBe(true); + }, 120_000); + + test("a filtered vector scan binds no list of eligible rows, so none is held on the heap", async () => { + store.ensureVecTable(3); + // One eligible document of 2,000 chunks, and closer ineligible documents. + const { hash } = await insertDoc("book", "long-eligible.md", "# Long eligible", { eligible: true }); + const now = new Date().toISOString(); + store.db.transaction(() => { + for (let seq = 0; seq < 2_000; seq++) insertEmbedding(store.db, hash, seq, seq, new Float32Array([0, 1, 0]), model, now, 2_000); + })(); + for (let i = 0; i < 50; i++) { + await insertEmbeddedDoc("book", `closer-${i}.md`, `# Closer ${i}`, [1, 0, 0], { eligible: false }); + } + // Every string bound to a vector scan: a JSON list of eligible rowids would show here. + const bound: string[] = []; + const recording: Database = { + prepare: (sql: string) => { + const real = store.db.prepare(sql); + if (!sql.includes("MATCH")) return real; + return { + ...real, + run: real.run.bind(real), + get: real.get.bind(real), + iterate: real.iterate.bind(real), + all: (...params: SQLiteValue[]) => { + for (const param of params) if (typeof param === "string" && param.startsWith("[")) bound.push(param); + return real.all(...params); + }, + }; + }, + transaction: (fn) => store.db.transaction(fn), + exec: (sql: string) => store.db.exec(sql), + loadExtension: (path: string) => store.db.loadExtension(path), + close: () => store.db.close(), + }; + + for (const scope of ["book", undefined]) { + const results = await searchVec(recording, "q", model, 5, scope, undefined, queryEmbedding, undefined, eligibleOnly); + expect(results.map(r => r.displayPath)).toEqual(["book/long-eligible.md"]); + } + expect(bound).toEqual([]); + }); + + test("shared content hash within one collection returns only the matching document path", async () => { + store.ensureVecTable(3); + const body = "# Shared body"; + const { hash } = await insertDoc("notes", "published-copy.md", body, { status: "published" }); + await insertDoc("notes", "draft-copy.md", body, { status: "draft" }); + insertEmbedding(store.db, hash, 0, 0, new Float32Array([1, 0, 0]), model, new Date().toISOString(), 1); + + const filtered = await searchVec( + store.db, "q", model, 10, "notes", undefined, queryEmbedding, undefined, + { field: "status", operator: "eq", value: "published" }, + ); + expect(filtered.map(r => r.displayPath)).toEqual(["notes/published-copy.md"]); + }); }); describe("structuredSearch with metadata filter", () => { @@ -248,7 +404,7 @@ describe("structuredSearch with metadata filter", () => { syncDocumentMetadata(store.db, draftId, draftBody, "draft.md"); const results = await structuredSearch(store, [{ type: "lex", query: "structured keyword" }], { - filter: { key: "status", operator: "eq", value: "published" }, + filter: { field: "status", operator: "eq", value: "published" }, skipRerank: true, }); @@ -270,3 +426,33 @@ describe("structuredSearch with metadata filter", () => { expect(byPath.get("notes/b.md")).toEqual({}); }); }); + +describe("searchFTS with a collection scope and a metadata filter together", () => { + const published: MetadataFilter = { field: "status", operator: "eq", value: "published" }; + + test("returns the in-scope document the filter admits, past both the window and a stronger draft", async () => { + // Every noise document outranks both small-collection documents globally, + // and the draft outranks the target inside the small collection. + for (let i = 0; i < 50; i++) { + await insertDoc("large", `noise-${i}.md`, `# N${i}\n\nalpha alpha alpha`, { status: "published" }); + } + await insertDoc("small", "draft.md", "# Draft\n\nalpha alpha", { status: "draft" }); + await insertDoc( + "small", + "target.md", + `# Target\n\n${"Unrelated prose. ".repeat(40)}One weaker mention of alpha.`, + { status: "published" }, + ); + + const results = searchFTS(store.db, "alpha", 1, "small", published); + expect(results.map(r => r.displayPath)).toEqual(["small/target.md"]); + }); + + test("keeps the 256 KiB body cap on the scoped path", async () => { + await insertDoc("small", "long.md", `# Long\n\nalpha ${"z".repeat(300 * 1024)}`, { status: "published" }); + + const results = searchFTS(store.db, "alpha", 5, "small", published); + expect(results).toHaveLength(1); + expect(results[0]!.body!.length).toBe(262_144); + }); +}); diff --git a/test/metadata-store.test.ts b/test/metadata-store.test.ts index a7e11becb..9a0c8301b 100644 --- a/test/metadata-store.test.ts +++ b/test/metadata-store.test.ts @@ -228,6 +228,21 @@ describe("reindexCollection metadata synchronization", () => { expect(getMetadataForPath("doc.md")).toEqual({ status: "ok" }); }); + test("re-extracts an unchanged file whose metadata came from an older extraction version", async () => { + await writeFile(join(collectionDir, "doc.md"), buildDoc("qmd:\n metadata:\n status: ok\n", "# Doc\n")); + await reindex(); + + store.db.prepare(`UPDATE document_metadata SET extraction_version = 0`).run(); + expect(countDocumentsPendingMetadata(store.db)).toBe(1); + + const result = await reindex(); + expect(result.unchanged).toBe(1); + const row = store.db.prepare(`SELECT extraction_version FROM document_metadata`).get() as { extraction_version: number }; + expect(row.extraction_version).toBe(METADATA_EXTRACTION_VERSION); + expect(countDocumentsPendingMetadata(store.db)).toBe(0); + expect(getMetadataForPath("doc.md")).toEqual({ status: "ok" }); + }); + test("counts extraction errors without aborting the collection", async () => { await writeFile(join(collectionDir, "bad.md"), buildDoc("qmd:\n metadata:\n mixed: [1, two]\n", "# Bad\n")); await writeFile(join(collectionDir, "good.md"), buildDoc("qmd:\n metadata:\n status: ok\n", "# Good\n")); diff --git a/test/metadata-surfaces.test.ts b/test/metadata-surfaces.test.ts index 3b5eec3cc..77d16a027 100644 --- a/test/metadata-surfaces.test.ts +++ b/test/metadata-surfaces.test.ts @@ -10,7 +10,14 @@ import { unlinkSync } from "node:fs"; import { tmpdir } from "node:os"; import { join } from "node:path"; import YAML from "yaml"; -import { createStore, type QMDStore } from "../src/index.js"; +import { + createStore, + MetadataBindingBudgetError, + MetadataOptionError, + type MetadataFilter, + type MetadataMatch, + type QMDStore, +} from "../src/index.js"; import { createStore as createInternalStore, insertContent, @@ -38,6 +45,16 @@ function buildDoc(status: string, body: string): string { return `---\nqmd:\n metadata:\n status: ${status}\n topics: [typescript]\n---\n\n${body}`; } +/** A filter and a match that each pass the parser's limits and together exceed the SQL binding budget. */ +function buildWidePredicates(): { filter: MetadataFilter; match: MetadataMatch } { + const members = Array.from({ length: 64 }, (_, index) => `v${index}`); + const groups = (operand: () => Operand) => Array.from({ length: 7 }, () => Array.from({ length: 31 }, operand)); + return { + filter: { operator: "and", operands: groups(() => ({ field: "status", operator: "all" as const, value: members })).map(operands => ({ operator: "and" as const, operands })) }, + match: { operator: "or", operands: groups(() => ({ field: "value" as const, operator: "in" as const, value: members })).map(operands => ({ operator: "or" as const, operands })) }, + }; +} + // ============================================================================= // SDK // ============================================================================= @@ -68,7 +85,7 @@ describe("SDK metadata filter", () => { expect(unfiltered.length).toBe(2); const filtered = await store.searchLex("sdk keyword", { - filter: { key: "status", operator: "eq", value: "published" }, + filter: { field: "status", operator: "eq", value: "published" }, }); expect(filtered.map(r => r.displayPath)).toEqual(["docs/published.md"]); expect(filtered[0]!.metadata).toEqual({ status: "published", topics: ["typescript"] }); @@ -77,7 +94,7 @@ describe("SDK metadata filter", () => { test("search with pre-expanded queries applies the filter", async () => { const results = await store.search({ queries: [{ type: "lex", query: "sdk keyword" }], - filter: { key: "status", operator: "ne", value: "draft" }, + filter: { field: "status", operator: "ne", value: "draft" }, rerank: false, }); expect(results.map(r => r.displayPath)).toEqual(["docs/published.md"]); @@ -94,6 +111,80 @@ describe("SDK metadata filter", () => { filter: { operator: "and", operands: [] }, })).rejects.toThrow(/non-empty 'operands'/); }); + + test("listMetadata summarizes keys, types, and value counts", async () => { + const result = await store.listMetadata(); + expect(result.documents).toBe(2); + expect(result.filteredDocuments).toBeUndefined(); + expect(result.keys.map(summary => summary.key)).toEqual(["status", "topics"]); + + const topics = result.keys[1]!.types[0]!; + expect(topics).toMatchObject({ type: "string", documents: 2, distinctValues: 1, remainingValues: 0, collections: ["docs"] }); + expect(result).toMatchObject({ totalKeys: 2, remainingKeys: 0 }); + expect(topics.values).toEqual([{ value: "typescript", documents: 2 }]); + }); + + test("listMetadata scopes, narrows by filter, and selects entries by match", async () => { + const filtered = await store.listMetadata({ + collection: "docs", + match: { field: "key", operator: "eq", value: "status" }, + filter: { field: "status", operator: "eq", value: "published" }, + }); + expect(filtered.filteredDocuments).toBe(1); + expect(filtered.keys[0]!.types[0]!.values).toEqual([{ value: "published", documents: 1 }]); + + const reverse = await store.listMetadata({ match: { field: "value", operator: "prefix", value: "dra" } }); + expect(reverse.keys.map(summary => summary.key)).toEqual(["status"]); + + const scoped = await store.listMetadata({ collection: "missing" }); + expect(scoped).toEqual({ documents: 0, totalKeys: 0, keys: [], remainingKeys: 0 }); + }); + + test("listMetadata windows keys and values and rejects options outside their domain", async () => { + const page = await store.listMetadata({ keyLimit: 1, keyOffset: 1 }); + expect(page.keys.map(summary => summary.key)).toEqual(["topics"]); + expect(page).toMatchObject({ totalKeys: 2, remainingKeys: 0 }); + + await expect(store.listMetadata({ keyLimit: 0 })).rejects.toThrow(/^Invalid keyLimit: expected a positive integer or Infinity, received 0$/); + await expect(store.listMetadata({ valueOffset: -1 })).rejects.toThrow(MetadataOptionError); + await expect(store.listMetadata({ keyLimit: 1e30 })).rejects.toThrow(MetadataOptionError); + await expect(store.listMetadata(buildWidePredicates())).rejects.toThrow(MetadataBindingBudgetError); + }); + + test("listMetadata validates filters and matches at the SDK runtime boundary", async () => { + // Plain-JS callers bypass the declarations, so the SDK must reject a bad AST at runtime. + const untrustedFilter: import("../src/index.js").ListMetadataOptions = JSON.parse('{"filter":{"field":"status","operator":"equal","value":"x"}}'); + await expect(store.listMetadata(untrustedFilter)).rejects.toThrow(/unknown operator 'equal'/); + + const untrustedMatch: import("../src/index.js").ListMetadataOptions = JSON.parse('{"match":{"field":"status","operator":"eq","value":"x"}}'); + await expect(store.listMetadata(untrustedMatch)).rejects.toThrow(/Invalid metadata match at \$: 'status' is not a field of a metadata entry/); + }); + + test("negated text matches preserve extracted empty strings through the SDK", async () => { + const document = store.internal.db.prepare("SELECT id FROM documents WHERE path = 'published.md'").get() as { id: number }; + replaceDocumentMetadata(store.internal.db, document.id, { metadata: { status: "", topics: ["typescript"] }, extractionVersion: METADATA_EXTRACTION_VERSION }); + try { + for (const operator of ["prefix", "suffix"] as const) { + const result = await store.listMetadata({ match: { operator: "and", operands: [ + { field: "key", operator: "eq", value: "status" }, + { operator: "not", operand: { field: "value", operator, value: "zzz" } }, + ] } }); + expect(result.keys[0]!.documents).toBe(2); + expect(result.keys[0]!.types[0]!.values).toEqual([{ value: "", documents: 1 }, { value: "draft", documents: 1 }]); + } + } finally { + replaceDocumentMetadata(store.internal.db, document.id, { metadata: { status: "published", topics: ["typescript"] }, extractionVersion: METADATA_EXTRACTION_VERSION }); + } + }); + + test("getStatus lists metadata keys per collection", async () => { + const status = await store.getStatus(); + expect(status.collections[0]!.metadataKeyCount).toBe(2); + expect(status.collections[0]!.metadataKeys).toEqual([ + { key: "status", documents: 2, types: ["string"] }, + { key: "topics", documents: 2, types: ["string"] }, + ]); + }); }); // ============================================================================= @@ -121,6 +212,7 @@ describe("MCP and HTTP metadata filter", () => { const internal = createInternalStore(dbPath); await seedDoc(internal.db, "published.md", "# Pub\n\nhttp keyword body", { status: "published" }); await seedDoc(internal.db, "draft.md", "# Draft\n\nhttp keyword body", { status: "draft" }); + await seedDoc(internal.db, "tagged.md", "# Tagged\n\nhttp keyword body", { status: "archived", topics: ["sqlite", "search"], priority: 3 }); const testConfig: CollectionConfig = { collections: { docs: { path: "/test/docs", pattern: "**/*.md" } }, @@ -160,6 +252,10 @@ describe("MCP and HTTP metadata filter", () => { } async function callQueryTool(args: Record): Promise<{ status: number; json: any }> { + return callTool("query", args); + } + + async function callTool(name: string, args: Record): Promise<{ status: number; json: any }> { const res = await fetch(`${baseUrl}/mcp`, { method: "POST", headers: { @@ -167,14 +263,14 @@ describe("MCP and HTTP metadata filter", () => { "Accept": "application/json, text/event-stream", "MCP-Protocol-Version": "2026-07-28", "Mcp-Method": "tools/call", - "Mcp-Name": "query", + "Mcp-Name": name, }, body: JSON.stringify({ jsonrpc: "2.0", id: 1, method: "tools/call", params: { - name: "query", + name, arguments: args, _meta: { "io.modelcontextprotocol/protocolVersion": "2026-07-28", @@ -187,10 +283,59 @@ describe("MCP and HTTP metadata filter", () => { return { status: res.status, json: await res.json() }; } + test("MCP initialization lists tools without reading metadata value rows", async () => { + const internal = createInternalStore(dbPath); + // Make value-table reads fail deterministically instead of testing a noisy + // timing threshold. The running server already has its database connection. + internal.db.exec("ALTER TABLE document_metadata_values RENAME TO hidden_metadata_values"); + try { + const response = await fetch(`${baseUrl}/mcp`, { + method: "POST", + headers: { "Content-Type": "application/json", "Accept": "application/json, text/event-stream", "MCP-Protocol-Version": "2026-07-28", "Mcp-Method": "tools/list" }, + body: JSON.stringify({ jsonrpc: "2.0", id: 1, method: "tools/list", params: { _meta: { + "io.modelcontextprotocol/protocolVersion": "2026-07-28", + "io.modelcontextprotocol/clientInfo": { name: "metadata-test", version: "1.0.0" }, + "io.modelcontextprotocol/clientCapabilities": {}, + } } }), + }); + expect(response.status).toBe(200); + const body = await response.json() as { result: { tools: { name: string }[] } }; + expect(body.result.tools.map(tool => tool.name)).toContain("metadata"); + } finally { + internal.db.exec("ALTER TABLE hidden_metadata_values RENAME TO document_metadata_values"); + internal.close(); + } + }); + + test("HTTP and MCP negated text matches include empty strings", async () => { + const internal = createInternalStore(dbPath); + const document = internal.db.prepare("SELECT id FROM documents WHERE path = 'draft.md'").get() as { id: number }; + replaceDocumentMetadata(internal.db, document.id, { metadata: { status: "" }, extractionVersion: METADATA_EXTRACTION_VERSION }); + try { + for (const operator of ["prefix", "suffix"] as const) { + const args = { match: { operator: "and", operands: [ + { field: "key", operator: "eq", value: "status" }, + { operator: "not", operand: { field: "value", operator, value: "zzz" } }, + ] } }; + const http = await postJson("/metadata", args); + const mcp = await callTool("metadata", args); + expect(http.status).toBe(200); + expect(mcp.status).toBe(200); + expect(mcp.json.result.structuredContent).toEqual(http.json); + expect(http.json.keys[0].documents).toBe(3); + expect(http.json.keys[0].types[0].values).toContainEqual({ value: "", documents: 1 }); + expect(mcp.json.result.content[0].text).toContain('""'); + } + } finally { + replaceDocumentMetadata(internal.db, document.id, { metadata: { status: "draft" }, extractionVersion: METADATA_EXTRACTION_VERSION }); + internal.close(); + } + }); + test("POST /query applies the filter and includes metadata", async () => { const { status, json } = await postJson("/query", { searches: [{ type: "lex", query: "http keyword" }], - filter: { key: "status", operator: "eq", value: "published" }, + filter: { field: "status", operator: "eq", value: "published" }, rerank: false, }); expect(status).toBe(200); @@ -202,7 +347,7 @@ describe("MCP and HTTP metadata filter", () => { test("POST /search alias accepts the same filter", async () => { const { status, json } = await postJson("/search", { searches: [{ type: "lex", query: "http keyword" }], - filter: { key: "status", operator: "eq", value: "draft" }, + filter: { field: "status", operator: "eq", value: "draft" }, rerank: false, }); expect(status).toBe(200); @@ -220,7 +365,7 @@ describe("MCP and HTTP metadata filter", () => { const invalidAst = await postJson("/query", { searches: [{ type: "lex", query: "http keyword" }], - filter: { key: "status", operator: "equal", value: "published" }, + filter: { field: "status", operator: "equal", value: "published" }, }); expect(invalidAst.status).toBe(400); expect(invalidAst.json.error).toMatch(/unknown operator 'equal'/); @@ -232,8 +377,8 @@ describe("MCP and HTTP metadata filter", () => { filter: { operator: "and", operands: [ - { key: "status", operator: "eq", value: "published" }, - { operator: "not", operand: { key: "status", operator: "eq", value: "draft" } }, + { field: "status", operator: "eq", value: "published" }, + { operator: "not", operand: { field: "status", operator: "eq", value: "draft" } }, ], }, rerank: false, @@ -256,4 +401,160 @@ describe("MCP and HTTP metadata filter", () => { expect(json.result.isError).toBe(true); expect(json.result.content[0].text).toMatch(/non-empty 'operands'/); }); + + test("MCP metadata tool returns key summaries as structured content and CLI-shaped text", async () => { + const { status, json } = await callTool("metadata", { match: { field: "key", operator: "eq", value: "status" } }); + expect(status).toBe(200); + expect(json.result.isError).toBeFalsy(); + + const result = json.result.structuredContent; + expect(result.documents).toBe(3); + expect(result.keys).toHaveLength(1); + expect(result.keys[0].types[0]).toMatchObject({ type: "string", documents: 3, distinctValues: 3, remainingValues: 0 }); + expect(result.keys[0].types[0].values).toEqual([ + { value: "archived", documents: 1 }, + { value: "draft", documents: 1 }, + { value: "published", documents: 1 }, + ]); + expect(json.result.content[0].text).toBe("status string 3 of 3 documents 3 distinct\n archived 1\n draft 1\n published 1"); + }); + + test("MCP metadata tool narrows by filter, windows values, and reports the remainder", async () => { + const { json } = await callTool("metadata", { + match: { field: "key", operator: "eq", value: "topics" }, + valueLimit: 1, + filter: { field: "status", operator: "eq", value: "archived" }, + }); + const result = json.result.structuredContent; + expect(result.filteredDocuments).toBe(1); + const topics = result.keys[0].types[0]; + expect(topics.multiValued).toBe(true); + expect(topics.values).toHaveLength(1); + expect(topics.remainingValues).toBe(1); + expect(json.result.content[0].text).toContain("filter: 1 of 3 documents\n\ntopics string[] 1 of 1 documents 2 distinct"); + expect(json.result.content[0].text).toContain("1 more value, use a higher 'valueLimit' or a 'valueOffset'"); + }); + + test("MCP metadata tool windows and pages keys", async () => { + const firstPage = await callTool("metadata", { keyLimit: 2 }); + const first = firstPage.json.result.structuredContent; + expect(first.keys.map((summary: { key: string }) => summary.key)).toEqual(["status", "priority"]); + expect(first).toMatchObject({ totalKeys: 3, remainingKeys: 1 }); + expect(firstPage.json.result.content[0].text).toContain("1 more key, use a higher 'keyLimit' or a 'keyOffset'"); + + const secondPage = await callTool("metadata", { keyLimit: 2, keyOffset: 2 }); + expect(secondPage.json.result.structuredContent.keys.map((summary: { key: string }) => summary.key)).toEqual(["topics"]); + expect(secondPage.json.result.structuredContent.remainingKeys).toBe(0); + + const pastTheEnd = await callTool("metadata", { keyOffset: 3 }); + expect(pastTheEnd.json.result.isError).toBeFalsy(); + expect(pastTheEnd.json.result.content[0].text).toBe("No keys at keyOffset 3, 3 keys in total."); + + const badOffset = await callTool("metadata", { keyOffset: -1 }); + expect(badOffset.json.error ?? badOffset.json.result?.isError).toBeTruthy(); + + const unsafeLimit = await callTool("metadata", { keyLimit: 1e30 }); + expect(unsafeLimit.json.error ?? unsafeLimit.json.result?.isError).toBeTruthy(); + }); + + test("MCP metadata tool keeps the filter line over an empty result and reports the binding budget", async () => { + // Every seeded document declares a status, so this filter admits none. + const empty = await callTool("metadata", { filter: { field: "status", operator: "exists", value: false } }); + expect(empty.json.result.isError).toBeFalsy(); + expect(empty.json.result.structuredContent).toMatchObject({ documents: 3, filteredDocuments: 0, totalKeys: 0 }); + expect(empty.json.result.content[0].text).toBe( + "filter: 0 of 3 documents\n\nNo metadata matches. Call without match/filter to see which keys exist, or check the status tool for collections with metadata.", + ); + + const overBudget = await callTool("metadata", buildWidePredicates()); + expect(overBudget.json.result.isError).toBe(true); + expect(overBudget.json.result.content[0].text).toMatch(/^Error: filter and match together bind \d+ SQL parameters, over the budget of 30000\./); + }); + + test("MCP metadata tool rejects invalid filters and matches", async () => { + const { json } = await callTool("metadata", { filter: { field: "status", operator: "equal", value: "x" } }); + expect(json.result.isError).toBe(true); + expect(json.result.content[0].text).toMatch(/unknown operator 'equal'/); + + const badMatch = await callTool("metadata", { match: { field: "value", operator: "all", value: ["x"] } }); + expect(badMatch.json.result.isError).toBe(true); + expect(badMatch.json.result.content[0].text).toMatch(/Invalid metadata match at \$: 'all' has no meaning for a single metadata entry/); + }); + + test("MCP status tool lists metadata keys per collection", async () => { + const { json } = await callTool("status", {}); + expect(json.result.isError).toBeFalsy(); + expect(json.result.structuredContent.collections[0].metadataKeyCount).toBe(3); + expect(json.result.structuredContent.collections[0].metadataKeys).toEqual([ + { key: "status", documents: 3, types: ["string"] }, + { key: "priority", documents: 1, types: ["number"] }, + { key: "topics", documents: 1, types: ["string"] }, + ]); + expect(json.result.content[0].text).toContain("metadata keys: status (string), priority (number), topics (string)"); + expect(json.result.content[0].text).toContain("call the 'metadata' tool"); + }); + + test("POST /metadata returns the same result as the tool", async () => { + const { status, json } = await postJson("/metadata", { match: { field: "value", operator: "type", value: "number" } }); + expect(status).toBe(200); + expect(json.documents).toBe(3); + expect(json.keys.map((summary: { key: string }) => summary.key)).toEqual(["priority"]); + expect(json.keys[0].types[0]).toMatchObject({ type: "number", range: { min: 3, median: 3, max: 3 } }); + + const reverse = await postJson("/metadata", { match: { field: "value", operator: "eq", value: "search" }, valueLimit: 5 }); + expect(reverse.json.keys.map((summary: { key: string }) => summary.key)).toEqual(["topics"]); + + const page = await postJson("/metadata", { keyLimit: 1, keyOffset: 1 }); + expect(page.json.keys.map((summary: { key: string }) => summary.key)).toEqual(["priority"]); + expect(page.json).toMatchObject({ totalKeys: 3, remainingKeys: 1 }); + }); + + test("POST /metadata rejects invalid bodies, filters, matches, and sort with 400", async () => { + const stringFilter = await postJson("/metadata", { filter: "status = published" }); + expect(stringFilter.status).toBe(400); + expect(stringFilter.json.error).toMatch(/must be an object/); + + const stringMatch = await postJson("/metadata", { match: "key = topics" }); + expect(stringMatch.status).toBe(400); + expect(stringMatch.json.error).toMatch(/Invalid field: match \(must be an object\)/); + + const invalidMatch = await postJson("/metadata", { match: { field: "topics", operator: "eq", value: "x" } }); + expect(invalidMatch.status).toBe(400); + expect(invalidMatch.json.error).toMatch(/'topics' is not a field of a metadata entry/); + + const invalidAst = await postJson("/metadata", { filter: { field: "status", operator: "equal", value: "x" } }); + expect(invalidAst.status).toBe(400); + expect(invalidAst.json.error).toMatch(/unknown operator 'equal'/); + + const badSort = await postJson("/metadata", { sort: "size" }); + expect(badSort.status).toBe(400); + expect(badSort.json.error).toMatch(/sort/); + + const stringLimit = await postJson("/metadata", { keyLimit: "5" }); + expect(stringLimit.status).toBe(400); + expect(stringLimit.json.error).toBe("Invalid field: keyLimit (must be a number)"); + + const zeroLimit = await postJson("/metadata", { valueLimit: 0 }); + expect(zeroLimit.status).toBe(400); + expect(zeroLimit.json.error).toBe("Invalid valueLimit: expected a positive integer or Infinity, received 0"); + + const fractionalOffset = await postJson("/metadata", { keyOffset: 1.5 }); + expect(fractionalOffset.status).toBe(400); + expect(fractionalOffset.json.error).toBe("Invalid keyOffset: expected a non-negative integer, received 1.5"); + + const stringCollections = await postJson("/metadata", { collections: "docs" }); + expect(stringCollections.status).toBe(400); + expect(stringCollections.json.error).toBe("Invalid field: collections (must be an array)"); + + const unsafeLimit = await postJson("/metadata", { keyLimit: 1e30 }); + expect(unsafeLimit.status).toBe(400); + expect(unsafeLimit.json.error).toBe("Invalid keyLimit: expected a positive integer or Infinity, received 1e+30"); + + const overBudget = await postJson("/metadata", buildWidePredicates()); + expect(overBudget.status).toBe(400); + expect(overBudget.json.error).toMatch(/^filter and match together bind \d+ SQL parameters, over the budget of 30000\./); + + const arrayBody = await fetch(`${baseUrl}/metadata`, { method: "POST", headers: { "Content-Type": "application/json" }, body: "[]" }); + expect(arrayBody.status).toBe(400); + }); }); diff --git a/test/multi-collection-filter.test.ts b/test/multi-collection-filter.test.ts index 63aae7d5b..ac779178e 100644 --- a/test/multi-collection-filter.test.ts +++ b/test/multi-collection-filter.test.ts @@ -1,9 +1,10 @@ /** * Unit tests for multi-collection CLI filter input (#191, #775). * - * Collection scoping is applied in searchFTS/searchVec (search each collection, - * then merge). These tests cover CLI parseArgs / input normalization only. - * Starvation regressions live in store.test.ts. + * Collection scoping is applied in searchFTS (search each collection, then + * merge) and searchVec (one KNN per collection partition, merged by distance). + * These tests cover CLI parseArgs / input normalization only. Starvation + * regressions live in store.test.ts. */ import { describe, test, expect } from "vitest"; diff --git a/test/sdk.test.ts b/test/sdk.test.ts index 5ae0d0e01..de9af83a1 100644 --- a/test/sdk.test.ts +++ b/test/sdk.test.ts @@ -6,7 +6,7 @@ */ import { describe, test, expect, beforeAll, afterAll, beforeEach, afterEach } from "vitest"; -import { mkdtemp, writeFile, mkdir, rm } from "node:fs/promises"; +import { mkdtemp, writeFile, mkdir, rm, rename } from "node:fs/promises"; import { tmpdir } from "node:os"; import { join } from "node:path"; import { existsSync, writeFileSync, mkdirSync, readFileSync } from "node:fs"; @@ -23,6 +23,7 @@ import { type ExpandQueryOptions, } from "../src/index.js"; import { setDefaultLlamaCpp } from "../src/llm.js"; +import { VEC_COLLECTION_IDS_TABLE, VEC_ROWS_TABLE } from "../src/vec-layout.js"; // ============================================================================= // Test Helpers @@ -850,6 +851,8 @@ describe("update", () => { expect(result.unchanged).toBe(0); expect(result.removed).toBe(0); expect(result.skipped).toBe(0); + expect(result.staleVectorsRemoved).toBe(0); + expect(result.vectorsCopied).toBe(0); expect(typeof result.needsEmbedding).toBe("number"); await store.close(); @@ -1101,6 +1104,116 @@ describe("embed", () => { } }); + test("store.update drops stale vector rows and copies rows into a collection that gained an embedded hash", async () => { + const leftDir = join(testDir, `vector-rows-left-${Date.now()}`); + const rightDir = join(testDir, `vector-rows-right-${Date.now()}`); + await mkdir(leftDir, { recursive: true }); + await mkdir(rightDir, { recursive: true }); + const shared = "# Shared\n\nEmbedded once, then joins a second collection.\n"; + await writeFile(join(leftDir, "shared.md"), shared); + await writeFile(join(rightDir, "gone.md"), "# Gone\n\nDisappears before the second update.\n"); + + const store = await createStore({ + dbPath: freshDbPath(), + config: { + collections: { + left: { path: leftDir, pattern: "**/*.md" }, + right: { path: rightDir, pattern: "**/*.md" }, + }, + }, + }); + setDefaultLlamaCpp(createFakeTokenizer() as any); + store.internal.llm = createFakeEmbedLlm() as any; + + const partitionsOf = (path: string): string[] => { + const doc = store.internal.db.prepare(`SELECT hash FROM documents WHERE path = ? LIMIT 1`).get(path) as { hash: string }; + const rows = store.internal.db.prepare(` + SELECT ci.name AS collection FROM ${VEC_ROWS_TABLE} vr + JOIN ${VEC_COLLECTION_IDS_TABLE} ci ON ci.id = vr.collection_id + WHERE vr.hash = ? + ORDER BY ci.name + `).all(doc.hash) as Array<{ collection: string }>; + return rows.map((row) => row.collection); + }; + + try { + await store.update(); + await store.embed(); + expect(partitionsOf("shared.md")).toEqual(["left"]); + expect(partitionsOf("gone.md")).toEqual(["right"]); + + await writeFile(join(rightDir, "shared.md"), shared); + await rm(join(rightDir, "gone.md")); + const result = await store.update(); + + expect(result.removed).toBe(1); + expect(result.staleVectorsRemoved).toBe(1); + expect(result.vectorsCopied).toBe(1); + expect(result.needsEmbedding).toBe(0); + expect(partitionsOf("shared.md")).toEqual(["left", "right"]); + expect(partitionsOf("gone.md")).toEqual([]); + } finally { + setDefaultLlamaCpp(null); + await store.close(); + } + }); + + test("store.update keeps the vectors of a document moved to another collection", async () => { + const fromDir = join(testDir, `vector-move-from-${Date.now()}`); + const toDir = join(testDir, `vector-move-to-${Date.now()}`); + await mkdir(fromDir, { recursive: true }); + await mkdir(toDir, { recursive: true }); + await writeFile(join(fromDir, "moved.md"), "# Moved\n\nEmbedded in one collection, then moved to another.\n"); + + const store = await createStore({ + dbPath: freshDbPath(), + config: { + collections: { + from: { path: fromDir, pattern: "**/*.md" }, + to: { path: toDir, pattern: "**/*.md" }, + }, + }, + }); + setDefaultLlamaCpp(createFakeTokenizer() as any); + store.internal.llm = createFakeEmbedLlm() as any; + + const db = store.internal.db; + const hashOf = () => + (db.prepare(`SELECT hash FROM documents WHERE path = 'moved.md' LIMIT 1`).get() as { hash: string }).hash; + const chunkCount = (hash: string) => + (db.prepare(`SELECT COUNT(*) AS n FROM content_vectors WHERE hash = ?`).get(hash) as { n: number }).n; + const partitionsOf = (hash: string): string[] => + (db.prepare(` + SELECT ci.name AS collection FROM ${VEC_ROWS_TABLE} vr + JOIN ${VEC_COLLECTION_IDS_TABLE} ci ON ci.id = vr.collection_id + WHERE vr.hash = ? + ORDER BY ci.name + `).all(hash) as Array<{ collection: string }>).map((row) => row.collection); + + try { + await store.update(); + await store.embed(); + const hash = hashOf(); + const embeddedChunks = chunkCount(hash); + expect(embeddedChunks).toBeGreaterThan(0); + expect(partitionsOf(hash)).toEqual(["from"]); + + await rename(join(fromDir, "moved.md"), join(toDir, "moved.md")); + const result = await store.update(); + + expect(result.removed).toBe(1); + expect(result.indexed).toBe(1); + expect(chunkCount(hash)).toBe(embeddedChunks); + expect(partitionsOf(hash)).toEqual(["to"]); + expect(result.vectorsCopied).toBe(1); + expect(result.staleVectorsRemoved).toBe(1); + expect(result.needsEmbedding).toBe(0); + } finally { + setDefaultLlamaCpp(null); + await store.close(); + } + }); + test("store.embed rejects invalid batch limits", async () => { const store = await createStore({ dbPath: freshDbPath(), diff --git a/test/store-migrations.test.ts b/test/store-migrations.test.ts new file mode 100644 index 000000000..8bc567000 --- /dev/null +++ b/test/store-migrations.test.ts @@ -0,0 +1,461 @@ +/** + * The version 2 step copies the legacy `vectors_vec` rows into the + * per-collection partitioned layout, chunk by chunk, resumably, and drops the + * legacy table in the transaction that stamps the version. + */ +import { describe, test, expect, afterEach, vi } from "vitest"; +import { mkdtemp, rm } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { openDatabase, loadSqliteVec, type Database } from "../src/db.js"; +import { createStore, insertContent, insertDocument, type Store } from "../src/store.js"; +import { + VECTOR_PARTITION_VERSION, + getUserVersion, + migrateVectorLayout, + runStoreMigrations, + type VectorMigrationPhase, + type VectorMigrationProgress, +} from "../src/store-migrations.js"; +import { + LEGACY_VEC_TABLE, + VEC_ROWS_TABLE, + VEC_TABLE, + createPartitionedVecTable, + resolveCollectionId, + vecInteger, + vecLayout, +} from "../src/vec-layout.js"; + +let store: Store | null = null; +let extra: Database[] = []; +let dir: string | null = null; + +async function openStore(): Promise { + dir = await mkdtemp(join(tmpdir(), "qmd-migration-")); + store = createStore(join(dir, "index.sqlite")); + return store; +} + +function openSecondConnection(s: Store): Database { + const db = openDatabase(s.dbPath); + loadSqliteVec(db); + extra.push(db); + return db; +} + +afterEach(async () => { + for (const db of extra) db.close(); + extra = []; + store?.close(); + store = null; + if (dir) await rm(dir, { recursive: true, force: true }); + dir = null; +}); + +const DIMS = 3; + +function vec(x: number, y: number, z: number): Float32Array { + return new Float32Array([x, y, z]); +} + +function storedVector(db: Database, hash: string, seq: number, collection: string): number[] | undefined { + const id = resolveCollectionId(db, collection); + if (id === undefined) return undefined; + const row = db.prepare(`SELECT id FROM ${VEC_ROWS_TABLE} WHERE hash = ? AND seq = ? AND collection_id = ?`).get(hash, seq, id) as { id: number } | undefined; + if (!row) return undefined; + const stored = db.prepare(`SELECT embedding FROM ${VEC_TABLE} WHERE rowid = ?`).get(vecInteger(row.id)) as { embedding: Uint8Array } | undefined; + if (!stored) return undefined; + return Array.from(new Float32Array(stored.embedding.buffer, stored.embedding.byteOffset, DIMS)); +} + +function partitionCount(db: Database, collection: string): number { + const id = resolveCollectionId(db, collection); + if (id === undefined) return 0; + return (db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_TABLE} WHERE collection_id = ?`).get(vecInteger(id)) as { c: number }).c; +} + +function tableNames(db: Database, prefix: string): string[] { + return (db.prepare(`SELECT name FROM sqlite_master WHERE name LIKE ? ORDER BY name`).all(`${prefix}%`) as { name: string }[]).map((r) => r.name); +} + +/** A store still at version 1 with a legacy vec0 table of small chunks. */ +class LegacyFixture { + readonly vectors = new Map(); + + constructor(readonly db: Database, chunkSize = 8) { + db.exec(`PRAGMA user_version = 1`); + db.exec(`CREATE VIRTUAL TABLE ${LEGACY_VEC_TABLE} USING vec0(hash_seq TEXT PRIMARY KEY, embedding float[${DIMS}] distance_metric=cosine, chunk_size=${chunkSize})`); + } + + doc(collection: string, hash: string, active = 1): void { + const now = new Date().toISOString(); + insertContent(this.db, hash, `Document ${hash}`, now); + insertDocument(this.db, collection, `${hash}.md`, hash, hash, now, now); + if (!active) this.db.prepare(`UPDATE documents SET active = 0 WHERE collection = ? AND hash = ?`).run(collection, hash); + } + + chunkRow(hash: string, seq: number, total = 1): void { + this.db.prepare(`INSERT OR IGNORE INTO content_vectors (hash, seq, pos, model, embedded_at, total_chunks) VALUES (?, ?, ?, 'test', ?, ?)`) + .run(hash, seq, seq * 100, new Date().toISOString(), total); + } + + legacyVector(hash: string, seq: number, embedding: Float32Array): void { + this.db.prepare(`INSERT INTO ${LEGACY_VEC_TABLE} (hash_seq, embedding) VALUES (?, ?)`).run(`${hash}_${seq}`, embedding); + this.vectors.set(`${hash}_${seq}`, embedding); + } + + /** Document with content_vectors rows and legacy vectors for every chunk. */ + embedded(collection: string, hash: string, chunks: readonly Float32Array[]): void { + this.doc(collection, hash); + chunks.forEach((embedding, seq) => { + this.chunkRow(hash, seq, chunks.length); + this.legacyVector(hash, seq, embedding); + }); + } + + legacyChunkCount(): number { + return (this.db.prepare(`SELECT COUNT(*) AS c FROM ${LEGACY_VEC_TABLE}_chunks`).get() as { c: number }).c; + } +} + +/** + * Two collections sharing one hash, a two-chunk document, an inactive-only + * hash, a legacy row without a content_vectors row, a content_vectors row + * without a vector, and enough filler to span several legacy chunks. + */ +function seedStandardFixture(db: Database): LegacyFixture { + const f = new LegacyFixture(db); + f.embedded("a", "h1", [vec(1, 0, 0), vec(0.9, 0.1, 0)]); + f.embedded("a", "h2", [vec(0, 1, 0)]); + f.embedded("a", "hs", [vec(0, 0, 1)]); + f.doc("b", "hs"); + f.embedded("b", "h3", [vec(0.5, 0.5, 0)]); + f.doc("a", "h4", 0); + f.chunkRow("h4", 0); + f.legacyVector("h4", 0, vec(0.1, 0.2, 0.3)); + f.legacyVector("ghost", 0, vec(0.3, 0.2, 0.1)); + f.doc("a", "h5"); + f.chunkRow("h5", 0); + for (let i = 0; i < 12; i++) { + f.embedded("a", `fill${String(i).padStart(2, "0")}`, [vec(0.2, 0.2, i * 0.05)]); + } + return f; +} + +const STANDARD_A_ROWS = 2 + 1 + 1 + 12; +const STANDARD_B_ROWS = 1 + 1; +/** h1 x2, h2, hs, h3, the inactive h4, the ghost row, and 12 fillers. */ +const STANDARD_LEGACY_ROWS = 2 + 1 + 1 + 1 + 1 + 1 + 12; + +describe("migrateVectorLayout", () => { + test("a fresh store has nothing to copy and reaches the partition version at open", async () => { + const s = await openStore(); + expect(getUserVersion(s.db)).toBeGreaterThanOrEqual(1); + expect(migrateVectorLayout(s.db, { sqliteVecAvailable: true })).toBe("applied"); + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + expect(vecLayout(s.db).kind).toBe("none"); + }); + + test("copies a hash_seq legacy table into per-collection partitions", async () => { + const s = await openStore(); + const fixture = seedStandardFixture(s.db); + expect(fixture.legacyChunkCount()).toBeGreaterThan(1); + const phases: VectorMigrationProgress[] = []; + + expect(migrateVectorLayout(s.db, { sqliteVecAvailable: true, onProgress: (p) => phases.push(p) })).toBe("applied"); + + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + expect(vecLayout(s.db)).toMatchObject({ kind: "partitioned", dimensions: DIMS }); + expect(tableNames(s.db, LEGACY_VEC_TABLE)).toEqual([]); + expect(s.db.prepare(`SELECT value FROM store_config WHERE key = 'vector_partition_cursor'`).get()).toBeFalsy(); + + expect(partitionCount(s.db, "a")).toBe(STANDARD_A_ROWS); + expect(partitionCount(s.db, "b")).toBe(STANDARD_B_ROWS); + expect((s.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_ROWS_TABLE}`).get() as { c: number }).c).toBe(STANDARD_A_ROWS + STANDARD_B_ROWS); + + expect(storedVector(s.db, "h1", 0, "a")).toEqual([1, 0, 0]); + expect(storedVector(s.db, "h1", 1, "a")).toEqual(Array.from(vec(0.9, 0.1, 0))); + expect(storedVector(s.db, "hs", 0, "a")).toEqual([0, 0, 1]); + expect(storedVector(s.db, "hs", 0, "b")).toEqual([0, 0, 1]); + expect(storedVector(s.db, "h3", 0, "b")).toEqual([0.5, 0.5, 0]); + expect(storedVector(s.db, "h3", 0, "a")).toBeUndefined(); + expect(storedVector(s.db, "h4", 0, "a")).toBeUndefined(); + + const chunkRows = (s.db.prepare(`SELECT hash FROM content_vectors ORDER BY hash`).all() as { hash: string }[]).map((r) => r.hash); + expect(chunkRows).not.toContain("h4"); + expect(chunkRows).not.toContain("h5"); + expect(chunkRows).toContain("h1"); + expect(chunkRows).toContain("hs"); + + const nearest = s.db.prepare(`SELECT rowid, distance FROM ${VEC_TABLE} WHERE embedding MATCH ? AND k = 1 AND collection_id = ?`) + .get(vec(0.5, 0.5, 0), vecInteger(resolveCollectionId(s.db, "b")!)) as { rowid: number }; + const nearestRow = s.db.prepare(`SELECT hash FROM ${VEC_ROWS_TABLE} WHERE id = ?`).get(nearest.rowid) as { hash: string }; + expect(nearestRow.hash).toBe("h3"); + + expect(phases.map((p) => p.phase)).toEqual([...phases.filter((p) => p.phase === "copy").map(() => "copy"), "verify", "flip", "vacuum", "done"]); + const lastCopy = phases.filter((p) => p.phase === "copy").at(-1)!; + expect(lastCopy.copied).toBe(lastCopy.total); + expect(lastCopy.total).toBe(STANDARD_LEGACY_ROWS); + }); + + test("drops a legacy table keyed by hash alone and stamps the version", async () => { + const s = await openStore(); + s.db.exec(`PRAGMA user_version = 1`); + s.db.exec(`CREATE VIRTUAL TABLE ${LEGACY_VEC_TABLE} USING vec0(hash TEXT PRIMARY KEY, embedding float[${DIMS}] distance_metric=cosine)`); + s.db.prepare(`INSERT INTO ${LEGACY_VEC_TABLE} (hash, embedding) VALUES ('h1', ?)`).run(vec(1, 0, 0)); + + expect(migrateVectorLayout(s.db, { sqliteVecAvailable: true })).toBe("applied"); + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + expect(vecLayout(s.db).kind).toBe("none"); + expect(tableNames(s.db, LEGACY_VEC_TABLE)).toEqual([]); + }); + + test("defers without sqlite-vec and leaves the legacy table and version alone", async () => { + const s = await openStore(); + const fixture = seedStandardFixture(s.db); + + expect(migrateVectorLayout(s.db, { sqliteVecAvailable: false })).toBe("deferred"); + expect(getUserVersion(s.db)).toBe(1); + expect(vecLayout(s.db).kind).toBe("legacy"); + expect((s.db.prepare(`SELECT COUNT(*) AS c FROM ${LEGACY_VEC_TABLE}`).get() as { c: number }).c).toBe(fixture.vectors.size); + expect((s.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_ROWS_TABLE}`).get() as { c: number }).c).toBe(0); + }); + + test("keeps the legacy dimension", async () => { + const s = await openStore(); + s.db.exec(`PRAGMA user_version = 1`); + s.db.exec(`CREATE VIRTUAL TABLE ${LEGACY_VEC_TABLE} USING vec0(hash_seq TEXT PRIMARY KEY, embedding float[5] distance_metric=cosine)`); + const now = new Date().toISOString(); + insertContent(s.db, "h1", "Document h1", now); + insertDocument(s.db, "a", "h1.md", "h1", "h1", now, now); + s.db.prepare(`INSERT INTO content_vectors (hash, seq, pos, model, embedded_at) VALUES ('h1', 0, 0, 'test', ?)`).run(now); + s.db.prepare(`INSERT INTO ${LEGACY_VEC_TABLE} (hash_seq, embedding) VALUES ('h1_0', ?)`).run(new Float32Array([1, 2, 3, 4, 5])); + + expect(migrateVectorLayout(s.db, { sqliteVecAvailable: true })).toBe("applied"); + expect(vecLayout(s.db)).toMatchObject({ kind: "partitioned", dimensions: 5 }); + expect(partitionCount(s.db, "a")).toBe(1); + }); + + test("resumes from the cursor after a failure between chunks", async () => { + const s = await openStore(); + seedStandardFixture(s.db); + let copies = 0; + + expect(() => migrateVectorLayout(s.db, { + sqliteVecAvailable: true, + onProgress: (p) => { + if (p.phase === "copy" && ++copies === 1) throw new Error("killed after the first chunk"); + }, + })).toThrow("killed after the first chunk"); + + expect(getUserVersion(s.db)).toBe(1); + expect(vecLayout(s.db).kind).toBe("legacy"); + const cursor = s.db.prepare(`SELECT value FROM store_config WHERE key = 'vector_partition_cursor'`).get() as { value: string }; + expect(Number(cursor.value)).toBeGreaterThanOrEqual(1); + const partial = (s.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_ROWS_TABLE}`).get() as { c: number }).c; + expect(partial).toBeGreaterThan(0); + expect(partial).toBeLessThan(STANDARD_A_ROWS + STANDARD_B_ROWS); + + const resumed: number[] = []; + expect(migrateVectorLayout(s.db, { sqliteVecAvailable: true, onProgress: (p) => { if (p.phase === "copy") resumed.push(p.copied); } })).toBe("applied"); + expect(resumed.length).toBeLessThan(copies + 10); + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + expect(partitionCount(s.db, "a")).toBe(STANDARD_A_ROWS); + expect(partitionCount(s.db, "b")).toBe(STANDARD_B_ROWS); + expect((s.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_ROWS_TABLE}`).get() as { c: number }).c).toBe(STANDARD_A_ROWS + STANDARD_B_ROWS); + }); + + test("two openers take turns on chunks and both finish with one copy of every row", async () => { + const s = await openStore(); + seedStandardFixture(s.db); + const second = openSecondConnection(s); + let secondResult: string | null = null; + let secondCopies = 0; + + const firstResult = migrateVectorLayout(s.db, { + sqliteVecAvailable: true, + onProgress: (p) => { + if (p.phase === "copy" && secondResult === null) { + secondResult = migrateVectorLayout(second, { + sqliteVecAvailable: true, + onProgress: (q) => { if (q.phase === "copy") secondCopies++; }, + }); + } + }, + }); + + expect(firstResult).toBe("applied"); + expect(secondResult).toBe("applied"); + expect(secondCopies).toBeGreaterThan(0); + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + expect(vecLayout(s.db).kind).toBe("partitioned"); + expect(partitionCount(s.db, "a")).toBe(STANDARD_A_ROWS); + expect(partitionCount(s.db, "b")).toBe(STANDARD_B_ROWS); + expect((s.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_ROWS_TABLE}`).get() as { c: number }).c).toBe(STANDARD_A_ROWS + STANDARD_B_ROWS); + }); + + test("the verification pass copies a row that landed in an already-walked chunk", async () => { + const s = await openStore(); + const fixture = seedStandardFixture(s.db); + let injected = false; + + expect(migrateVectorLayout(s.db, { + sqliteVecAvailable: true, + onProgress: (p) => { + if (p.phase === "copy" && p.copied === p.total && !injected) { + injected = true; + fixture.embedded("b", "late", [vec(0.7, 0.7, 0.1)]); + } + }, + })).toBe("applied"); + + expect(injected).toBe(true); + expect(storedVector(s.db, "late", 0, "b")).toEqual(Array.from(vec(0.7, 0.7, 0.1))); + expect(partitionCount(s.db, "b")).toBe(STANDARD_B_ROWS + 1); + }); + + test("a legacy row written after the copy phase and before the flip ends up in the partition", async () => { + const s = await openStore(); + const fixture = seedStandardFixture(s.db); + const phases: VectorMigrationPhase[] = []; + let injected = false; + + expect(migrateVectorLayout(s.db, { + sqliteVecAvailable: true, + onProgress: (p) => { + phases.push(p.phase); + if (p.phase === "verify" && !injected) { + injected = true; + fixture.embedded("b", "late", [vec(0.7, 0.7, 0.1)]); + } + }, + })).toBe("applied"); + + expect(injected).toBe(true); + expect(phases.slice(phases.indexOf("verify"))).toEqual(["verify", "flip", "vacuum", "done"]); + expect(storedVector(s.db, "late", 0, "b")).toEqual(Array.from(vec(0.7, 0.7, 0.1))); + expect(partitionCount(s.db, "b")).toBe(STANDARD_B_ROWS + 1); + expect(s.db.prepare(`SELECT 1 FROM content_vectors WHERE hash = 'late' AND seq = 0`).get()).toBeTruthy(); + expect(tableNames(s.db, LEGACY_VEC_TABLE)).toEqual([]); + }); + + test("the verification pass and the drop hold one write lock", async () => { + const s = await openStore(); + seedStandardFixture(s.db); + const other = openSecondConnection(s); + other.exec(`PRAGMA busy_timeout = 0`); + let lockedAtFlip: boolean | null = null; + + expect(migrateVectorLayout(s.db, { + sqliteVecAvailable: true, + onProgress: (p) => { + if (p.phase !== "flip") return; + try { + other.exec(`BEGIN IMMEDIATE`); + other.exec(`ROLLBACK`); + lockedAtFlip = false; + } catch { + lockedAtFlip = true; + } + }, + })).toBe("applied"); + + expect(lockedAtFlip).toBe(true); + expect(vecLayout(s.db).kind).toBe("partitioned"); + expect(tableNames(s.db, LEGACY_VEC_TABLE)).toEqual([]); + }); + + test("creating the partitioned table a second time is a no-op", async () => { + const s = await openStore(); + createPartitionedVecTable(s.db, DIMS); + expect(() => createPartitionedVecTable(s.db, DIMS)).not.toThrow(); + expect(vecLayout(s.db)).toMatchObject({ kind: "partitioned", dimensions: DIMS }); + expect(tableNames(s.db, VEC_TABLE)).toContain(`${VEC_TABLE}_chunks`); + }); + + test("a second opener finishes the migration with the partitioned table another opener created", async () => { + const s = await openStore(); + seedStandardFixture(s.db); + createPartitionedVecTable(s.db, DIMS); + expect(vecLayout(s.db).kind).toBe("legacy"); + const second = openSecondConnection(s); + + expect(migrateVectorLayout(second, { sqliteVecAvailable: true })).toBe("applied"); + + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + expect(vecLayout(s.db)).toMatchObject({ kind: "partitioned", dimensions: DIMS }); + expect(partitionCount(s.db, "a")).toBe(STANDARD_A_ROWS); + expect(partitionCount(s.db, "b")).toBe(STANDARD_B_ROWS); + expect(migrateVectorLayout(s.db, { sqliteVecAvailable: true })).toBe("applied"); + }); + + test("refuses a partitioned table whose dimensions differ from the legacy table's", async () => { + const s = await openStore(); + seedStandardFixture(s.db); + createPartitionedVecTable(s.db, DIMS + 1); + + expect(() => migrateVectorLayout(s.db, { sqliteVecAvailable: true })).toThrow(/dimensions/); + expect(getUserVersion(s.db)).toBe(1); + expect(vecLayout(s.db).kind).toBe("legacy"); + }); + + test("the flip drops the legacy table and its shadow tables with the version stamp", async () => { + const s = await openStore(); + seedStandardFixture(s.db); + expect(tableNames(s.db, LEGACY_VEC_TABLE).length).toBeGreaterThan(1); + + migrateVectorLayout(s.db, { sqliteVecAvailable: true }); + + expect(tableNames(s.db, LEGACY_VEC_TABLE)).toEqual([]); + expect(tableNames(s.db, VEC_TABLE)).toContain(`${VEC_TABLE}_chunks`); + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + expect(migrateVectorLayout(s.db, { sqliteVecAvailable: true })).toBe("applied"); + expect(partitionCount(s.db, "a")).toBe(STANDARD_A_ROWS); + }); +}); + +describe("runStoreMigrations", () => { + test("migrates a legacy table an older build created on a stamped store", async () => { + const s = await openStore(); + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + seedStandardFixture(s.db); + s.db.exec(`PRAGMA user_version = ${VECTOR_PARTITION_VERSION}`); + expect(vecLayout(s.db).kind).toBe("legacy"); + + runStoreMigrations(s.db, { installFtsSyncTriggers: () => {}, sqliteVecAvailable: true }); + + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + expect(vecLayout(s.db)).toMatchObject({ kind: "partitioned", dimensions: DIMS }); + expect(tableNames(s.db, LEGACY_VEC_TABLE)).toEqual([]); + expect(partitionCount(s.db, "a")).toBe(STANDARD_A_ROWS); + expect(partitionCount(s.db, "b")).toBe(STANDARD_B_ROWS); + }); + + test("keeps the legacy table in place when the repair fails on a stamped store", async () => { + const s = await openStore(); + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + seedStandardFixture(s.db); + s.db.exec(`PRAGMA user_version = ${VECTOR_PARTITION_VERSION}`); + createPartitionedVecTable(s.db, DIMS + 1); + expect(vecLayout(s.db).kind).toBe("legacy"); + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + + expect(() => runStoreMigrations(s.db, { installFtsSyncTriggers: () => {}, sqliteVecAvailable: true })).not.toThrow(); + + expect(warn).toHaveBeenCalledWith(expect.stringContaining("Legacy vector table left in place")); + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + expect(vecLayout(s.db).kind).toBe("legacy"); + warn.mockRestore(); + }); + + test("leaves a stamped store without a legacy table alone", async () => { + const s = await openStore(); + createPartitionedVecTable(s.db, DIMS); + + runStoreMigrations(s.db, { installFtsSyncTriggers: () => { throw new Error("must not reinstall"); }, sqliteVecAvailable: true }); + + expect(getUserVersion(s.db)).toBe(VECTOR_PARTITION_VERSION); + expect(vecLayout(s.db)).toMatchObject({ kind: "partitioned", dimensions: DIMS }); + }); +}); diff --git a/test/store.helpers.unit.test.ts b/test/store.helpers.unit.test.ts index aaaf3a0ee..0a058e2cd 100644 --- a/test/store.helpers.unit.test.ts +++ b/test/store.helpers.unit.test.ts @@ -6,6 +6,7 @@ import { describe, test, expect } from "vitest"; import { mkdtempSync, mkdirSync, writeFileSync, symlinkSync, rmSync, chmodSync } from "node:fs"; import { join } from "node:path"; import { tmpdir } from "node:os"; +import { VEC_TABLE } from "../src/vec-layout.js"; import { homedir, resolve, @@ -169,10 +170,11 @@ describe("countOrphanedVectors", () => { describe("cleanupOrphanedVectors", () => { test("returns 0 when vec table exists in schema but sqlite-vec is unavailable", () => { const prepare = (sql: string) => { - if (sql.includes("sqlite_master") && sql.includes("vectors_vec")) { - return { get: () => ({ name: "vectors_vec" }) }; + if (sql.includes("sqlite_master")) { + const row = { name: VEC_TABLE, sql: `CREATE VIRTUAL TABLE ${VEC_TABLE} USING vec0(collection_id INTEGER PARTITION KEY, embedding float[3] distance_metric=cosine)` }; + return { get: () => row, all: () => [row] }; } - if (sql.includes("SELECT 1 FROM vectors_vec LIMIT 0")) { + if (sql.includes(`SELECT 1 FROM ${VEC_TABLE} LIMIT 0`)) { return { get: () => { throw new Error("no such module: vec0"); } }; } throw new Error(`Unexpected SQL in test: ${sql}`); diff --git a/test/store.test.ts b/test/store.test.ts index 51f913a3d..f2a160f72 100644 --- a/test/store.test.ts +++ b/test/store.test.ts @@ -7,16 +7,19 @@ */ import { describe, test, expect, beforeAll, afterAll, beforeEach, afterEach, vi } from "vitest"; -import { openDatabase, loadSqliteVec } from "../src/db.js"; -import type { Database } from "../src/db.js"; -import { unlink, mkdtemp, rmdir, writeFile, rm, mkdir, rename, chmod, readFile, symlink } from "node:fs/promises"; +import { openDatabase, loadSqliteVec, isBun } from "../src/db.js"; +import type { Database, SQLiteValue } from "../src/db.js"; +import { unlink, mkdtemp, rmdir, writeFile, rm, mkdir, rename, chmod, readFile, symlink, utimes } from "node:fs/promises"; import { tmpdir } from "node:os"; import { join } from "node:path"; +import { spawnSync } from "node:child_process"; +import { fileURLToPath } from "node:url"; import YAML from "yaml"; import * as llmModule from "../src/llm.js"; import { disposeDefaultLlamaCpp, setDefaultLlamaCpp } from "../src/llm.js"; import { createStore, + DEFAULT_EMBED_MODEL, DEFAULT_QUERY_MODEL, DEFAULT_RERANK_MODEL, verifySqliteVecLoaded, @@ -56,12 +59,18 @@ import { insertContent, insertDocument, cleanupOrphanedVectors, - generateEmbeddings, + clearAllEmbeddings, + copyVectorsToNewCollections, + VECTOR_COPY_BATCH_ROWS, maybeAdoptLegacyEmbeddingFingerprint, + removeCollection, + renameCollection, + generateEmbeddings, getHybridRrfWeights, _resetProductionModeForTesting, hybridQuery, structuredSearch, + searchVec, vectorSearchQuery, type Store, type DocumentResult, @@ -70,6 +79,7 @@ import { type RankedListMeta, } from "../src/store.js"; import type { CollectionConfig } from "../src/collections.js"; +import { LEGACY_VEC_TABLE, VEC_COLLECTION_IDS_TABLE, VEC_ROWS_TABLE, VEC_TABLE, deletePartitionRows, resolveCollectionId, vecInteger } from "../src/vec-layout.js"; // ============================================================================= // LlamaCpp Setup @@ -185,6 +195,35 @@ async function insertTestDocument( return row?.id ?? 0; } +/** + * Same connection, but any statement whose SQL contains `fragment` throws + * `message` instead of running, whether prepared or exec'd. + */ +function failingDb(db: Database, fragment: string, message: string): Database { + const guard = (sql: string) => { + if (sql.includes(fragment)) throw new Error(message); + }; + return { + prepare: (sql: string) => { + guard(sql); + return db.prepare(sql); + }, + transaction: (fn) => db.transaction(fn), + exec: (sql: string) => { + guard(sql); + db.exec(sql); + }, + loadExtension: (path: string) => db.loadExtension(path), + close: () => db.close(), + }; +} + +function vectorRowCount(store: Store, collection: string): number { + const id = resolveCollectionId(store.db, collection); + if (id === undefined) return 0; + return (store.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_ROWS_TABLE} WHERE collection_id = ?`).get(id) as { c: number }).c; +} + /** Sync YAML config file to SQLite store_collections in the current test store */ async function syncTestConfig(): Promise { if (!currentTestStore) return; @@ -1370,6 +1409,128 @@ describe("Query expansion cache (#818)", () => { await cleanupTestDb(store); } }); + + test("hybridQuery embeds each distinct cached expansion once (#921)", async () => { + const store = await createTestStore(); + const embedModel = "hf:ggml-org/embeddinggemma-300M-GGUF/embeddinggemma-300M-Q8_0.gguf"; + const embedBatchSpy = vi.fn(async (texts: string[]) => texts.map(() => ({ embedding: [1, 2, 3], model: embedModel }))); + store.db.exec(`CREATE TABLE ${VEC_TABLE} (collection_id INTEGER, embedding BLOB)`); + store.llm = { embedModelName: embedModel, embedBatch: embedBatchSpy } as any; + store.searchVec = vi.fn(async () => [] as SearchResult[]) as any; + try { + // The row from #921: one hyde string cached 12 times, one vec string twice. + const cached = [ + ...Array.from({ length: 12 }, () => ({ type: "hyde", query: "musubi reconstruction guide" })), + { type: "vec", query: "methods for musubi" }, + { type: "vec", query: "methods for musubi" }, + { type: "lex", query: "musubi reconstruction" }, + ]; + store.setCachedResult(getCacheKey("expandQuery", { query: "musubi", model: DEFAULT_QUERY_MODEL }), JSON.stringify(cached)); + + await hybridQuery(store, "musubi", { limit: 5, minScore: 0, skipRerank: true, intent: "x" }); + + expect(embedBatchSpy).toHaveBeenCalledTimes(1); + expect(embedBatchSpy.mock.calls[0]![0]).toHaveLength(3); + } finally { + await cleanupTestDb(store); + } + }); + + test("expandQuery drops repeated lines from fresh model output before caching them", async () => { + const store = await createTestStore(); + const generateModelName = "dedupe-generate-model"; + store.llm = { + generateModelName, + expandQuery: async () => [ + { type: "hyde", text: "g" }, { type: "hyde", text: "g" }, { type: "hyde", text: "g" }, + { type: "vec", text: "v" }, { type: "vec", text: "v" }, + ], + } as any; + try { + const expanded = await store.expandQuery("q"); + expect(expanded).toEqual([{ type: "hyde", query: "g" }, { type: "vec", query: "v" }]); + const cached = store.getCachedResult(getCacheKey("expandQuery", { query: "q", model: generateModelName })); + expect(JSON.parse(cached!)).toHaveLength(2); + } finally { + await cleanupTestDb(store); + } + }); + + test("vectorSearchQuery searches each distinct cached expansion once", async () => { + const store = await createTestStore(); + const embedModel = "dedupe-embed-model"; + store.ensureVecTable(3); + store.llm = { embedModelName: embedModel } as any; + const searchVecSpy = vi.fn(async () => [] as SearchResult[]); + store.searchVec = searchVecSpy as any; + try { + const cached = [ + ...Array.from({ length: 12 }, () => ({ type: "hyde", query: "g" })), + { type: "vec", query: "v" }, + { type: "vec", query: "v" }, + ]; + store.setCachedResult(getCacheKey("expandQuery", { query: "q", model: DEFAULT_QUERY_MODEL }), JSON.stringify(cached)); + + await vectorSearchQuery(store, "q", { limit: 5 }); + + // The original query plus one search per distinct expansion. + expect(searchVecSpy).toHaveBeenCalledTimes(3); + expect(searchVecSpy.mock.calls.map(call => call[0])).toEqual(["q", "g", "v"]); + } finally { + await cleanupTestDb(store); + } + }); + + test("vectorSearchQuery searches each distinct expansion once from a cached row in the older text shape", async () => { + const store = await createTestStore(); + const embedModel = "dedupe-embed-model"; + store.ensureVecTable(3); + store.llm = { embedModelName: embedModel } as any; + const searchVecSpy = vi.fn(async () => [] as SearchResult[]); + store.searchVec = searchVecSpy as any; + try { + // Rows written before the cache stored `query` carry the line as `text`. + const cached = [ + ...Array.from({ length: 12 }, () => ({ type: "hyde", text: "g" })), + { type: "vec", text: "v" }, + { type: "vec", text: "v" }, + ]; + store.setCachedResult(getCacheKey("expandQuery", { query: "q", model: DEFAULT_QUERY_MODEL }), JSON.stringify(cached)); + + await vectorSearchQuery(store, "q", { limit: 5 }); + + expect(searchVecSpy.mock.calls.map(call => call[0])).toEqual(["q", "g", "v"]); + } finally { + await cleanupTestDb(store); + } + }); + + test("hybridQuery runs one keyword search per distinct cached lex line", async () => { + const store = await createTestStore(); + const embedModel = "dedupe-embed-model"; + store.ensureVecTable(3); + store.llm = { + embedModelName: embedModel, + embedBatch: async (texts: string[]) => texts.map(() => ({ embedding: [1, 2, 3], model: embedModel })), + } as any; + store.searchVec = vi.fn(async () => [] as SearchResult[]) as any; + const searchFTSSpy = vi.fn(() => [] as SearchResult[]); + store.searchFTS = searchFTSSpy as any; + try { + const cached = [ + { type: "lex", query: "k" }, { type: "lex", query: "k" }, { type: "lex", query: "k" }, + { type: "vec", query: "v" }, + ]; + store.setCachedResult(getCacheKey("expandQuery", { query: "q", model: DEFAULT_QUERY_MODEL }), JSON.stringify(cached)); + + await hybridQuery(store, "q", { limit: 5, minScore: 0, skipRerank: true, intent: "x" }); + + // The original query's probe, then the distinct lex line once. + expect(searchFTSSpy.mock.calls.map(call => call[0])).toEqual(["q", "k"]); + } finally { + await cleanupTestDb(store); + } + }); }); @@ -1628,6 +1789,46 @@ describe("FTS Search", () => { await cleanupTestDb(store); }); + test("searchFTS scope is exact when an out-of-scope collection fills the candidate window (#922)", async () => { + const store = await createTestStore(); + const noise = await createTestCollection({ name: "noise", pwd: "/test/noise" }); + const target = await createTestCollection({ name: "target", pwd: "/test/target" }); + const other = await createTestCollection({ name: "other", pwd: "/test/other" }); + + // limit=5 made the old scoped window limit*10 = 50: exactly as many + // stronger out-of-scope hits as fit in it. + for (let i = 0; i < 50; i++) { + await insertTestDocument(store.db, noise, { + name: `noise-${i}`, + title: "alpha alpha", + body: `Noise ${i}: alpha alpha alpha.`, + displayPath: `noise-${i}.md`, + }); + } + await insertTestDocument(store.db, target, { + name: "t", + title: "Target", + body: `${"Unrelated prose. ".repeat(40)}One weaker mention of alpha.`, + displayPath: "t.md", + }); + await insertTestDocument(store.db, other, { + name: "o", + title: "Other", + body: "Nothing relevant here.", + displayPath: "o.md", + }); + + expect(store.searchFTS("alpha", 5).every(r => r.collectionName === noise)).toBe(true); + + const single = store.searchFTS("alpha", 5, target); + expect(single.map(r => r.displayPath)).toEqual([`${target}/t.md`]); + + const multi = store.searchFTS("alpha", 5, [target, other]); + expect(multi.map(r => r.displayPath)).toEqual([`${target}/t.md`]); + + await cleanupTestDb(store); + }); + test("searchFTS finds CJK documents by exact and mixed queries", async () => { const store = await createTestStore(); const collectionName = await createTestCollection(); @@ -2826,6 +3027,42 @@ describe("Reindex Collection", () => { expect(paths.map(r => r.path)).toEqual(["X - b.md", "a.md"]); }); + test("leaves the index unchanged when the collection root is missing", async () => { + const store = await createTestStore(); + const collectionName = "unmounted"; + const parent = join(testDir, `unmounted-${Date.now()}-${Math.random().toString(36).slice(2)}`); + const collectionPath = join(parent, "notes"); + await mkdir(collectionPath, { recursive: true }); + await writeFile(join(collectionPath, "a.md"), "# A\n\nalpha\n"); + await writeFile(join(collectionPath, "b.md"), "# B\n\nbravo\n"); + + try { + const initial = await reindexCollection(store, collectionPath, "**/*.md", collectionName); + expect(initial.indexed).toBe(2); + + // An unmounted drive or offline share looks exactly like an empty folder to + // the glob. That must not read as "every file was deleted". + await rename(collectionPath, join(parent, "notes.unmounted")); + const whileMissing = await reindexCollection(store, collectionPath, "**/*.md", collectionName); + expect(whileMissing.removed).toBe(0); + expect(whileMissing.skipped).toBe(1); + expect(whileMissing.skippedFiles[0]!.code).toBe("ROOT_MISSING"); + + const active = store.db.prepare(` + SELECT COUNT(*) AS count FROM documents WHERE collection = ? AND active = 1 + `).get(collectionName) as { count: number }; + expect(active.count).toBe(2); + + // A genuinely empty folder still deactivates everything. + await mkdir(collectionPath); + const whileEmpty = await reindexCollection(store, collectionPath, "**/*.md", collectionName); + expect(whileEmpty.removed).toBe(2); + } finally { + await rm(parent, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); + test("does not index a file symlink whose target is outside the collection", async () => { const store = await createTestStore(); const parent = join(testDir, `escape-sym-${Date.now()}-${Math.random().toString(36).slice(2)}`); @@ -3018,156 +3255,722 @@ describe("Reindex Collection", () => { }); }); -// ============================================================================= -// Index Status Tests -// ============================================================================= +describe("Reindex Collection file sync state (#962)", () => { + const BODY_CAP = 262_144; -describe("Index Status", () => { - test("getStatus returns correct structure", async () => { - const store = await createTestStore(); - const status = store.getStatus(); - expect(status).toHaveProperty("totalDocuments"); - expect(status).toHaveProperty("needsEmbedding"); - expect(status).toHaveProperty("hasVectorIndex"); - expect(status).toHaveProperty("collections"); - expect(Array.isArray(status.collections)).toBe(true); + async function collectionDir(prefix: string): Promise { + const dir = join(testDir, `${prefix}-${Date.now()}-${Math.random().toString(36).slice(2)}`); + await mkdir(dir, { recursive: true }); + return dir; + } - await cleanupTestDb(store); - }); + function activeBody(store: Store, collection: string, path: string): string | undefined { + const row = store.db.prepare(` + SELECT content.doc AS body FROM documents d JOIN content ON content.hash = d.hash + WHERE d.collection = ? AND d.path = ? AND d.active = 1 + `).get(collection, path) as { body: string } | undefined; + return row?.body; + } - test("getStatus counts documents correctly", async () => { - const store = await createTestStore(); - const collectionName = await createTestCollection(); + function syncRowCount(store: Store, collection: string, path: string): number { + const row = store.db.prepare(` + SELECT COUNT(*) AS n FROM file_sync_state WHERE collection = ? AND relative_path = ? + `).get(collection, path) as { n: number }; + return row.n; + } - await insertTestDocument(store.db, collectionName, { name: "doc1", active: 1 }); - await insertTestDocument(store.db, collectionName, { name: "doc2", active: 1 }); - await insertTestDocument(store.db, collectionName, { name: "doc3", active: 0 }); // inactive + test("reindex trusts an unchanged mtime and size without reading the file", async () => { + const store = await createTestStore(); + const dir = await collectionDir("sync-fast-path"); + const file = join(dir, "doc.md"); + try { + // A whole-second mtime survives utimes exactly under both runtimes; a + // millisecond Date can come back a fraction lower through Node's seconds. + const mtime = new Date("2026-01-02T03:04:05Z"); + await writeFile(file, "# A\n\nalpha\n"); + await utimes(file, mtime, mtime); + await reindexCollection(store, dir, "**/*.md", "notes"); + // Different bytes of the same size, with the mtime put back. + await writeFile(file, "# A\n\nbravo\n"); + await utimes(file, mtime, mtime); + + const result = await reindexCollection(store, dir, "**/*.md", "notes"); + expect(result).toMatchObject({ indexed: 0, updated: 0, unchanged: 1 }); + expect(activeBody(store, "notes", "doc.md")).toContain("alpha"); + } finally { + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); - const status = store.getStatus(); - expect(status.totalDocuments).toBe(2); // Only active docs + test("reindex does not read an unchanged file", async () => { + const store = await createTestStore(); + const dir = await collectionDir("sync-no-read"); + const file = join(dir, "doc.md"); + try { + await writeFile(file, "# A\n\nalpha\n"); + // Older than the racy window, so the first pass stores a trusted row. + await utimes(file, new Date("2026-01-02T03:04:05Z"), new Date("2026-01-02T03:04:05Z")); + await reindexCollection(store, dir, "**/*.md", "notes"); + // Stat still works on an unreadable file; a read would fail with EACCES. + await chmod(file, 0o000); + + const result = await reindexCollection(store, dir, "**/*.md", "notes"); + expect(result.skippedFiles).toEqual([]); + expect(result).toMatchObject({ indexed: 0, updated: 0, unchanged: 1, removed: 0 }); + expect(activeBody(store, "notes", "doc.md")).toContain("alpha"); + } finally { + await chmod(file, 0o644); + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); - await cleanupTestDb(store); + test("a touched file whose content is unchanged refreshes its sync row without re-indexing", async () => { + const store = await createTestStore(); + const dir = await collectionDir("sync-touch"); + const file = join(dir, "doc.md"); + try { + await writeFile(file, "# A\n\nalpha\n"); + await utimes(file, new Date("2026-01-02T03:04:05Z"), new Date("2026-01-02T03:04:05Z")); + await reindexCollection(store, dir, "**/*.md", "notes"); + const touched = new Date("2026-01-02T03:05:05Z"); + await utimes(file, touched, touched); + + const first = await reindexCollection(store, dir, "**/*.md", "notes"); + expect(first).toMatchObject({ indexed: 0, updated: 0, unchanged: 1, removed: 0 }); + const row = store.db.prepare(`SELECT mtime_ms FROM file_sync_state WHERE collection = ? AND relative_path = ?`) + .get("notes", "doc.md") as { mtime_ms: number }; + expect(row.mtime_ms).toBe(touched.getTime()); + + // The refreshed row puts the file back on the fast path: a read would now report EACCES. + await chmod(file, 0o000); + const second = await reindexCollection(store, dir, "**/*.md", "notes"); + expect(second.skippedFiles).toEqual([]); + expect(second).toMatchObject({ indexed: 0, updated: 0, unchanged: 1, removed: 0 }); + expect(activeBody(store, "notes", "doc.md")).toContain("alpha"); + } finally { + await chmod(file, 0o644); + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } }); - test("getStatus reports collection info", async () => { + test("a same-size rewrite that keeps a recent mtime is still read", async () => { const store = await createTestStore(); - const collectionName = await createTestCollection({ pwd: "/test/path", glob: "**/*.md" }); - await insertTestDocument(store.db, collectionName, { name: "doc1" }); + const dir = await collectionDir("sync-racy"); + const file = join(dir, "doc.md"); + try { + // A whole second ahead of the clock: inside the racy window however slow the run. + const mtime = new Date((Math.floor(Date.now() / 1000) + 1) * 1000); + await writeFile(file, "# A\n\nalpha\n"); + await utimes(file, mtime, mtime); + await reindexCollection(store, dir, "**/*.md", "notes"); + // Rewritten within the same mtime granule: same size, same mtime. + await writeFile(file, "# A\n\nbravo\n"); + await utimes(file, mtime, mtime); + + const result = await reindexCollection(store, dir, "**/*.md", "notes"); + expect(result).toMatchObject({ indexed: 0, updated: 1, unchanged: 0 }); + expect(activeBody(store, "notes", "doc.md")).toContain("bravo"); + } finally { + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); - const status = store.getStatus(); - expect(status.collections.length).toBeGreaterThanOrEqual(1); - const col = status.collections.find(c => c.name === collectionName); - expect(col).toBeDefined(); - expect(col?.path).toBe("/test/path"); - expect(col?.pattern).toBe("**/*.md"); - expect(col?.documents).toBe(1); + test("reindexCollection groups its writes into transactions", async () => { + const store = await createTestStore(); + const collectionPath = join(testDir, `batched-${Date.now()}-${Math.random().toString(36).slice(2)}`); + await mkdir(collectionPath, { recursive: true }); + for (let i = 0; i < 5; i++) await writeFile(join(collectionPath, `d${i}.md`), `# D${i}\n\nbody ${i}\n`); + const seen: boolean[] = []; - await cleanupTestDb(store); + try { + const result = await reindexCollection(store, collectionPath, "**/*.md", "batched", { + onProgress: () => seen.push(store.db.inTransaction), + }); + expect(result.indexed).toBe(5); + expect(seen).toContain(true); + expect(store.db.inTransaction).toBe(false); + const rows = store.db.prepare(`SELECT COUNT(*) AS c FROM file_sync_state WHERE collection = ?`).get("batched") as { c: number }; + expect(rows.c).toBe(5); + } finally { + await rm(collectionPath, { recursive: true, force: true }); + await cleanupTestDb(store); + } }); - test("getHashesNeedingEmbedding counts correctly", async () => { + test("a failed write rolls back the open batch and leaves no transaction behind", async () => { const store = await createTestStore(); - const collectionName = await createTestCollection(); + const collectionPath = join(testDir, `batch-fail-${Date.now()}-${Math.random().toString(36).slice(2)}`); + await mkdir(collectionPath, { recursive: true }); + for (const name of ["a.md", "b.md", "c.md"]) await writeFile(join(collectionPath, name), `# ${name}\n\nbody\n`); + store.db.exec(`CREATE TRIGGER fail_b BEFORE INSERT ON documents WHEN NEW.path = 'b.md' BEGIN SELECT RAISE(ABORT, 'injected'); END`); - // Add documents with different hashes - await insertTestDocument(store.db, collectionName, { name: "doc1", hash: "hash1" }); - await insertTestDocument(store.db, collectionName, { name: "doc2", hash: "hash2" }); - await insertTestDocument(store.db, collectionName, { name: "doc3", hash: "hash1" }); // same hash as doc1 + try { + await expect(reindexCollection(store, collectionPath, "**/*.md", "batch-fail")).rejects.toThrow("injected"); + expect(store.db.inTransaction).toBe(false); + + store.db.exec(`DROP TRIGGER fail_b`); + const retry = await reindexCollection(store, collectionPath, "**/*.md", "batch-fail"); + expect(retry.indexed + retry.unchanged + retry.updated).toBe(3); + const rows = store.db.prepare(`SELECT COUNT(*) AS c FROM file_sync_state WHERE collection = ?`).get("batch-fail") as { c: number }; + expect(rows.c).toBe(3); + } finally { + await rm(collectionPath, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); - const needsEmbedding = store.getHashesNeedingEmbedding(); - expect(needsEmbedding).toBe(2); // hash1 and hash2 + test("files over 10 MB are skipped with FILE_TOO_LARGE and not indexed", async () => { + const store = await createTestStore(); + const dir = await collectionDir("sync-too-large"); + try { + await writeFile(join(dir, "big.md"), "a".repeat(10 * 1024 * 1024 + 1)); + await writeFile(join(dir, "small.md"), "# Small\n\nfits\n"); - await cleanupTestDb(store); + const result = await reindexCollection(store, dir, "**/*.md", "notes"); + expect(result.indexed).toBe(1); + expect(result.skippedFiles).toEqual([{ file: "big.md", code: "FILE_TOO_LARGE" }]); + const rows = store.db.prepare(`SELECT COUNT(*) AS n FROM documents WHERE collection = ? AND path = ?`) + .get("notes", "big.md") as { n: number }; + expect(rows.n).toBe(0); + } finally { + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } }); - test("embedding health is scoped to the active embed model", async () => { + test("a previously indexed file that grows past 10 MB is reported and deactivated", async () => { const store = await createTestStore(); - const collectionName = await createTestCollection(); - const activeModel = "hf:active/embed-model.gguf"; - const staleModel = "hf:stale/embed-model.gguf"; - const now = new Date().toISOString(); + const dir = await collectionDir("sync-grown"); + const file = join(dir, "doc.md"); + try { + await writeFile(file, "# A\n\nalpha\n"); + await reindexCollection(store, dir, "**/*.md", "notes"); + await writeFile(file, "# A\n\n" + "a".repeat(10 * 1024 * 1024)); + + const result = await reindexCollection(store, dir, "**/*.md", "notes"); + expect(result.skippedFiles).toEqual([{ file: "doc.md", code: "FILE_TOO_LARGE" }]); + expect(activeBody(store, "notes", "doc.md")).toBeUndefined(); + expect(syncRowCount(store, "notes", "doc.md")).toBe(0); + } finally { + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); - store.llm = { embedModelName: activeModel } as any; - store.ensureVecTable(3); - await insertTestDocument(store.db, collectionName, { name: "doc1", hash: "hash1" }); - store.insertEmbedding("hash1", 0, 0, new Float32Array([1, 2, 3]), staleModel, now, 1); + test("a previously indexed file that becomes empty is deactivated", async () => { + const store = await createTestStore(); + const dir = await collectionDir("sync-emptied"); + const file = join(dir, "doc.md"); + try { + await writeFile(file, "# A\n\nalpha\n"); + await reindexCollection(store, dir, "**/*.md", "notes"); + await writeFile(file, ""); - expect(store.getHashesNeedingEmbedding()).toBe(1); - expect(store.getStatus().needsEmbedding).toBe(1); - expect(store.getIndexHealth().needsEmbedding).toBe(1); - expect(store.getHashesNeedingEmbedding(staleModel)).toBe(0); + await reindexCollection(store, dir, "**/*.md", "notes"); + expect(activeBody(store, "notes", "doc.md")).toBeUndefined(); + } finally { + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); - await cleanupTestDb(store); + test("deleting an indexed file removes its sync row on the next reindex", async () => { + const store = await createTestStore(); + const dir = await collectionDir("sync-deleted"); + try { + await writeFile(join(dir, "doc.md"), "# A\n\nalpha\n"); + await writeFile(join(dir, "keep.md"), "# B\n\nbravo\n"); + await reindexCollection(store, dir, "**/*.md", "notes"); + expect(syncRowCount(store, "notes", "doc.md")).toBe(1); + await rm(join(dir, "doc.md")); + + const result = await reindexCollection(store, dir, "**/*.md", "notes"); + expect(result.removed).toBe(1); + expect(syncRowCount(store, "notes", "doc.md")).toBe(0); + expect(syncRowCount(store, "notes", "keep.md")).toBe(1); + } finally { + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } }); - test("embedding health treats stale fingerprints as needing re-embedding", async () => { + test("removing a collection and adding it back re-indexes its files", async () => { const store = await createTestStore(); - const collectionName = await createTestCollection(); - const model = "hf:test/embed-model.gguf"; - const now = new Date().toISOString(); + const dir = await collectionDir("sync-readd"); + try { + await writeFile(join(dir, "doc.md"), "# A\n\nalpha\n"); + await reindexCollection(store, dir, "**/*.md", "notes"); + removeCollection(store.db, "notes"); - store.llm = { embedModelName: model } as any; - store.ensureVecTable(3); - await insertTestDocument(store.db, collectionName, { name: "doc1", hash: "hash1" }); - store.insertEmbedding("hash1", 0, 0, new Float32Array([1, 2, 3]), model, now, 1, "stale1"); + const result = await reindexCollection(store, dir, "**/*.md", "notes"); + expect(result.indexed).toBe(1); + expect(activeBody(store, "notes", "doc.md")).toContain("alpha"); + } finally { + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); - expect(getEmbeddingFingerprint(model)).toMatch(/^[a-f0-9]{6}$/); - expect(store.getHashesNeedingEmbedding()).toBe(1); + test("removing and re-adding a collection re-indexes a file whose mtime moved but content did not", async () => { + const store = await createTestStore(); + const dir = await collectionDir("sync-readd-touched"); + const file = join(dir, "doc.md"); + try { + await writeFile(file, "# A\n\nalpha\n"); + await utimes(file, new Date("2026-01-02T03:04:05Z"), new Date("2026-01-02T03:04:05Z")); + await reindexCollection(store, dir, "**/*.md", "notes"); + removeCollection(store.db, "notes"); + await utimes(file, new Date("2026-01-02T03:05:05Z"), new Date("2026-01-02T03:05:05Z")); - await cleanupTestDb(store); + const result = await reindexCollection(store, dir, "**/*.md", "notes"); + expect(result.indexed).toBe(1); + expect(activeBody(store, "notes", "doc.md")).toContain("alpha"); + } finally { + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } }); - test("getIndexHealth returns health info", async () => { + test("removing a collection deletes its sync rows", async () => { const store = await createTestStore(); - const collectionName = await createTestCollection(); - await insertTestDocument(store.db, collectionName, { name: "doc1" }); + const dir = await collectionDir("sync-remove-rows"); + try { + await writeFile(join(dir, "doc.md"), "# A\n\nalpha\n"); + await reindexCollection(store, dir, "**/*.md", "notes"); + expect(syncRowCount(store, "notes", "doc.md")).toBe(1); - const health = store.getIndexHealth(); - expect(health).toHaveProperty("needsEmbedding"); - expect(health).toHaveProperty("totalDocs"); - expect(health).toHaveProperty("daysStale"); - expect(health.totalDocs).toBe(1); + removeCollection(store.db, "notes"); + expect(syncRowCount(store, "notes", "doc.md")).toBe(0); + } finally { + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); - await cleanupTestDb(store); + test("renaming a collection moves its sync rows, so the new name keeps the fast path", async () => { + const store = await createTestStore(); + const dir = await collectionDir("sync-rename-rows"); + const file = join(dir, "doc.md"); + try { + await writeFile(file, "# A\n\nalpha\n"); + await utimes(file, new Date("2026-01-02T03:04:05Z"), new Date("2026-01-02T03:04:05Z")); + await reindexCollection(store, dir, "**/*.md", "notes"); + renameCollection(store.db, "notes", "archive"); + expect(syncRowCount(store, "notes", "doc.md")).toBe(0); + expect(syncRowCount(store, "archive", "doc.md")).toBe(1); + + // A read would now report EACCES. + await chmod(file, 0o000); + const result = await reindexCollection(store, dir, "**/*.md", "archive"); + expect(result.skippedFiles).toEqual([]); + expect(result).toMatchObject({ indexed: 0, updated: 0, unchanged: 1, removed: 0 }); + } finally { + await chmod(file, 0o644); + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } }); -}); + + test("renaming a collection and adding the old name at another path indexes both", async () => { + const store = await createTestStore(); + const first = await collectionDir("sync-rename-first"); + const second = await collectionDir("sync-rename-second"); + const mtime = new Date("2026-01-02T03:04:05Z"); + try { + // Same relative path, size and mtime in both directories. + await writeFile(join(first, "doc.md"), "# A\n\nalpha\n"); + await writeFile(join(second, "doc.md"), "# A\n\nbravo\n"); + await utimes(join(first, "doc.md"), mtime, mtime); + await utimes(join(second, "doc.md"), mtime, mtime); + await reindexCollection(store, first, "**/*.md", "notes"); + renameCollection(store.db, "notes", "archive"); + + await reindexCollection(store, first, "**/*.md", "archive"); + const result = await reindexCollection(store, second, "**/*.md", "notes"); + expect(result.indexed).toBe(1); + expect(activeBody(store, "archive", "doc.md")).toContain("alpha"); + expect(activeBody(store, "notes", "doc.md")).toContain("bravo"); + } finally { + await rm(first, { recursive: true, force: true }); + await rm(second, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); + + test("a file whose document was deactivated elsewhere is re-indexed", async () => { + const store = await createTestStore(); + const dir = await collectionDir("sync-deactivated"); + try { + await writeFile(join(dir, "doc.md"), "# A\n\nalpha\n"); + await reindexCollection(store, dir, "**/*.md", "notes"); + // Unchanged file and sync row; only the document went inactive. + store.deactivateDocument("notes", "doc.md"); + + await reindexCollection(store, dir, "**/*.md", "notes"); + expect(activeBody(store, "notes", "doc.md")).toContain("alpha"); + } finally { + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); + + test("a file whose document now points at other content is re-read", async () => { + const store = await createTestStore(); + const dir = await collectionDir("sync-rehashed"); + try { + await writeFile(join(dir, "doc.md"), "# A\n\nalpha\n"); + await reindexCollection(store, dir, "**/*.md", "notes"); + // Unchanged file and sync row; only the document's content changed. + const now = new Date().toISOString(); + store.insertContent("elsewhere-hash", "# A\n\nwritten elsewhere\n", now); + const doc = store.db.prepare(`SELECT id FROM documents WHERE collection = ? AND path = ?`) + .get("notes", "doc.md") as { id: number }; + store.updateDocument(doc.id, "A", "elsewhere-hash", now); + + await reindexCollection(store, dir, "**/*.md", "notes"); + expect(activeBody(store, "notes", "doc.md")).toContain("alpha"); + } finally { + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); + + test("after #983's atomic collection rename the new name takes the fast path at once", async () => { + const store = await createTestStore(); + const dir = await collectionDir("sync-atomic-rename"); + const file = join(dir, "doc.md"); + try { + await writeFile(file, "# A\n\nalpha\n"); + await utimes(file, new Date("2026-01-02T03:04:05Z"), new Date("2026-01-02T03:04:05Z")); + await reindexCollection(store, dir, "**/*.md", "notes"); + renameCollection(store.db, "notes", "archive"); + expect(syncRowCount(store, "archive", "doc.md")).toBe(1); + + // The first pass under the new name would report EACCES if it read the file. + await chmod(file, 0o000); + const first = await reindexCollection(store, dir, "**/*.md", "archive"); + expect(first.skippedFiles).toEqual([]); + expect(first).toMatchObject({ indexed: 0, updated: 0, unchanged: 1, removed: 0 }); + expect(activeBody(store, "archive", "doc.md")).toContain("alpha"); + } finally { + await chmod(file, 0o644); + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); + + test("searchFTS returns at most 256 KiB of a document body", async () => { + const store = await createTestStore(); + try { + const body = "# Long\n\nzebracap " + "x".repeat(300 * 1024); + await insertTestDocument(store.db, "docs", { name: "long", body, displayPath: "long.md" }); + + const results = store.searchFTS("zebracap", 5); + expect(results).toHaveLength(1); + expect(results[0]!.body!.length).toBe(BODY_CAP); + expect(results[0]!.body).toBe(body.slice(0, BODY_CAP)); + } finally { + await cleanupTestDb(store); + } + }); + + test("searchVec returns at most 256 KiB of a document body", async () => { + const store = await createTestStore(); + try { + const body = "# Long\n\nvector cap " + "y".repeat(300 * 1024); + const hash = await hashContent(body); + await insertTestDocument(store.db, "docs", { name: "long", body, hash, displayPath: "long.md" }); + store.ensureVecTable(3); + store.insertEmbedding(hash, 0, 0, new Float32Array([1, 0, 0]), "cap-model", new Date().toISOString()); + + const results = await store.searchVec("q", "cap-model", 5, undefined, undefined, [1, 0, 0]); + expect(results).toHaveLength(1); + expect(results[0]!.body!.length).toBe(BODY_CAP); + expect(results[0]!.body).toBe(body.slice(0, BODY_CAP)); + } finally { + await cleanupTestDb(store); + } + }); + + test("searchFTS returns a short body with an embedded NUL in full", async () => { + const store = await createTestStore(); + try { + const body = "# Nul\n\nzebranul before\u0000after \u00e4\u{1F600}"; + await insertTestDocument(store.db, "docs", { name: "nul", body, displayPath: "nul.md" }); + + const results = store.searchFTS("zebranul", 5); + expect(results).toHaveLength(1); + expect(results[0]!.body).toBe(body); + } finally { + await cleanupTestDb(store); + } + }); + + test("searchVec returns a short body with an embedded NUL in full", async () => { + const store = await createTestStore(); + try { + const body = "# Nul\n\nvector before\u0000after \u00e4\u{1F600}"; + const hash = await hashContent(body); + await insertTestDocument(store.db, "docs", { name: "nul", body, hash, displayPath: "nul.md" }); + store.ensureVecTable(3); + store.insertEmbedding(hash, 0, 0, new Float32Array([1, 0, 0]), "cap-model", new Date().toISOString()); + + const results = await store.searchVec("q", "cap-model", 5, undefined, undefined, [1, 0, 0]); + expect(results).toHaveLength(1); + expect(results[0]!.body).toBe(body); + } finally { + await cleanupTestDb(store); + } + }); + + test("getHashesForEmbedding returns a short body with an embedded NUL in full", async () => { + const store = await createTestStore(); + try { + const body = "# Nul\n\nembed before\u0000after \u00e4\u{1F600}"; + await insertTestDocument(store.db, "docs", { name: "nul", body, displayPath: "nul.md" }); + + expect(store.getHashesForEmbedding().map(row => row.body)).toEqual([body]); + } finally { + await cleanupTestDb(store); + } + }); +}); + +describe("Collection-scoped keyword search", () => { + // Short noise documents with the term in their titles outrank, globally, one + // long target in each of two small collections. The old scoped search took a + // global top (limit * 10) and filtered it, so the noise must fill that window. + async function crowdedCollections(store: Store, noiseCount: number): Promise<{ smallA: string; smallB: string }> { + const large = await createTestCollection({ name: "crowd-large", pwd: "/test/crowd-large" }); + const smallA = await createTestCollection({ name: "crowd-small-a", pwd: "/test/crowd-small-a" }); + const smallB = await createTestCollection({ name: "crowd-small-b", pwd: "/test/crowd-small-b" }); + for (let i = 0; i < noiseCount; i++) { + await insertTestDocument(store.db, large, { + name: `noise-${i}`, + title: `zebra zebra ${i}`, + body: `# Noise ${i}\n\nzebra zebra zebra, noise document ${i}.`, + displayPath: `noise-${i}.md`, + }); + } + for (const collection of [smallA, smallB]) { + await insertTestDocument(store.db, collection, { + name: `target-${collection}`, + title: "Target", + body: `# Target\n\n${"filler prose without the search term. ".repeat(40)}zebra.`, + displayPath: "target.md", + }); + } + return { smallA, smallB }; + } + + test("searchFTS over two crowded-out collections returns a match from each", async () => { + const store = await createTestStore(); + try { + const { smallA, smallB } = await crowdedCollections(store, 40); + const results = store.searchFTS("zebra", 2, [smallA, smallB]); + expect(results.map(r => r.displayPath).sort()).toEqual([`${smallA}/target.md`, `${smallB}/target.md`]); + } finally { + await cleanupTestDb(store); + } + }); + + test("searchFTS over several collections runs one keyword query and returns the same results", async () => { + const store = await createTestStore(); + try { + const { smallA, smallB } = await crowdedCollections(store, 40); + const scope = ["crowd-large", smallA, smallB]; + const perCollection = scope.flatMap(name => store.searchFTS("zebra", 5, name)) + .sort((a, b) => b.score - a.score || (a.filepath < b.filepath ? -1 : a.filepath > b.filepath ? 1 : 0)) + .slice(0, 5); + // Counts the keyword queries the scoped search runs. + let ftsQueries = 0; + const counting: Database = { + prepare: (sql: string) => { + const real = store.db.prepare(sql); + if (!sql.includes("documents_fts MATCH")) return real; + return { + ...real, + run: real.run.bind(real), + get: real.get.bind(real), + iterate: real.iterate.bind(real), + all: (...params: Parameters) => { ftsQueries++; return real.all(...params); }, + }; + }, + transaction: (fn) => store.db.transaction(fn), + exec: (sql: string) => store.db.exec(sql), + loadExtension: (path: string) => store.db.loadExtension(path), + close: () => store.db.close(), + }; + + const { searchFTS } = await import("../src/store.js"); + const results = searchFTS(counting, "zebra", 5, scope); + expect(ftsQueries).toBe(1); + expect(results.map(r => r.filepath)).toEqual(perCollection.map(r => r.filepath)); + } finally { + await cleanupTestDb(store); + } + }); + + test("structuredSearch over two crowded-out collections returns a match from each", async () => { + const store = await createTestStore(); + try { + // structuredSearch asks searchFTS for 20 results, a window of 200. + const { smallA, smallB } = await crowdedCollections(store, 250); + const results = await structuredSearch(store, [{ type: "lex", query: "zebra" }], { + collections: [smallA, smallB], skipRerank: true, + }); + expect(results.map(r => r.displayPath).sort()).toEqual([`${smallA}/target.md`, `${smallB}/target.md`]); + } finally { + await cleanupTestDb(store); + } + }); +}); + +// ============================================================================= +// Index Status Tests +// ============================================================================= + +describe("Index Status", () => { + test("getStatus returns correct structure", async () => { + const store = await createTestStore(); + const status = store.getStatus(); + expect(status).toHaveProperty("totalDocuments"); + expect(status).toHaveProperty("needsEmbedding"); + expect(status).toHaveProperty("hasVectorIndex"); + expect(status).toHaveProperty("collections"); + expect(Array.isArray(status.collections)).toBe(true); + + await cleanupTestDb(store); + }); + + test("getStatus counts documents correctly", async () => { + const store = await createTestStore(); + const collectionName = await createTestCollection(); + + await insertTestDocument(store.db, collectionName, { name: "doc1", active: 1 }); + await insertTestDocument(store.db, collectionName, { name: "doc2", active: 1 }); + await insertTestDocument(store.db, collectionName, { name: "doc3", active: 0 }); // inactive + + const status = store.getStatus(); + expect(status.totalDocuments).toBe(2); // Only active docs + + await cleanupTestDb(store); + }); + + test("getStatus reports collection info", async () => { + const store = await createTestStore(); + const collectionName = await createTestCollection({ pwd: "/test/path", glob: "**/*.md" }); + await insertTestDocument(store.db, collectionName, { name: "doc1" }); + + const status = store.getStatus(); + expect(status.collections.length).toBeGreaterThanOrEqual(1); + const col = status.collections.find(c => c.name === collectionName); + expect(col).toBeDefined(); + expect(col?.path).toBe("/test/path"); + expect(col?.pattern).toBe("**/*.md"); + expect(col?.documents).toBe(1); + + await cleanupTestDb(store); + }); + + test("getHashesNeedingEmbedding counts correctly", async () => { + const store = await createTestStore(); + const collectionName = await createTestCollection(); + + // Add documents with different hashes + await insertTestDocument(store.db, collectionName, { name: "doc1", hash: "hash1" }); + await insertTestDocument(store.db, collectionName, { name: "doc2", hash: "hash2" }); + await insertTestDocument(store.db, collectionName, { name: "doc3", hash: "hash1" }); // same hash as doc1 + + const needsEmbedding = store.getHashesNeedingEmbedding(); + expect(needsEmbedding).toBe(2); // hash1 and hash2 + + await cleanupTestDb(store); + }); + + test("embedding health is scoped to the active embed model", async () => { + const store = await createTestStore(); + const collectionName = await createTestCollection(); + const activeModel = "hf:active/embed-model.gguf"; + const staleModel = "hf:stale/embed-model.gguf"; + const now = new Date().toISOString(); + + store.llm = { embedModelName: activeModel } as any; + store.ensureVecTable(3); + await insertTestDocument(store.db, collectionName, { name: "doc1", hash: "hash1" }); + store.insertEmbedding("hash1", 0, 0, new Float32Array([1, 2, 3]), staleModel, now, 1); + + expect(store.getHashesNeedingEmbedding()).toBe(1); + expect(store.getStatus().needsEmbedding).toBe(1); + expect(store.getIndexHealth().needsEmbedding).toBe(1); + expect(store.getHashesNeedingEmbedding(staleModel)).toBe(0); + + await cleanupTestDb(store); + }); + + test("embedding health treats stale fingerprints as needing re-embedding", async () => { + const store = await createTestStore(); + const collectionName = await createTestCollection(); + const model = "hf:test/embed-model.gguf"; + const now = new Date().toISOString(); + + store.llm = { embedModelName: model } as any; + store.ensureVecTable(3); + await insertTestDocument(store.db, collectionName, { name: "doc1", hash: "hash1" }); + store.insertEmbedding("hash1", 0, 0, new Float32Array([1, 2, 3]), model, now, 1, "stale1"); + + expect(getEmbeddingFingerprint(model)).toMatch(/^[a-f0-9]{6}$/); + expect(store.getHashesNeedingEmbedding()).toBe(1); + + await cleanupTestDb(store); + }); + + test("getIndexHealth returns health info", async () => { + const store = await createTestStore(); + const collectionName = await createTestCollection(); + await insertTestDocument(store.db, collectionName, { name: "doc1" }); + + const health = store.getIndexHealth(); + expect(health).toHaveProperty("needsEmbedding"); + expect(health).toHaveProperty("totalDocs"); + expect(health).toHaveProperty("daysStale"); + expect(health.totalDocs).toBe(1); + + await cleanupTestDb(store); + }); +}); describe("cleanupOrphanedVectors atomicity", () => { - // Seeds one active document (1 chunk) and one inactive document (2 chunks), - // so cleanup should remove exactly the 2 orphaned chunks from both tables. + // Seeds one active document (1 chunk) and one document (2 chunks) that goes + // inactive after embedding, so cleanup should remove exactly the 2 orphaned + // chunks from both tables. async function seedOrphanFixture(store: Store): Promise { const collectionName = await createTestCollection(); const now = new Date().toISOString(); store.ensureVecTable(3); await insertTestDocument(store.db, collectionName, { name: "kept-doc", hash: "keephash" }); - await insertTestDocument(store.db, collectionName, { name: "orphaned-doc", hash: "orphanhash", active: 0 }); + await insertTestDocument(store.db, collectionName, { name: "orphaned-doc", hash: "orphanhash" }); store.insertEmbedding("keephash", 0, 0, new Float32Array([1, 2, 3]), "test-model", now, 1); store.insertEmbedding("orphanhash", 0, 0, new Float32Array([4, 5, 6]), "test-model", now, 2); store.insertEmbedding("orphanhash", 1, 10, new Float32Array([7, 8, 9]), "test-model", now, 2); + store.db.prepare(`UPDATE documents SET active = 0 WHERE hash = 'orphanhash'`).run(); } function vecCounts(db: Database): { vec: number; meta: number } { - const vec = (db.prepare(`SELECT COUNT(*) AS c FROM vectors_vec`).get() as { c: number }).c; + const vec = (db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_TABLE}`).get() as { c: number }).c; const meta = (db.prepare(`SELECT COUNT(*) AS c FROM content_vectors`).get() as { c: number }).c; return { vec, meta }; } // Fault injection: same connection, but the content_vectors DELETE throws — - // after the vectors_vec DELETE already executed inside the transaction. + // after the vector DELETEs already executed inside the transaction. function makeFailingDb(db: Database): Database { - return { - prepare: (sql: string) => db.prepare(sql), - transaction: (fn) => db.transaction(fn), - exec: (sql: string) => { - if (sql.includes("DELETE FROM content_vectors")) { - throw new Error("injected failure between deletes"); - } - return db.exec(sql); - }, - loadExtension: (path: string) => db.loadExtension(path), - close: () => db.close(), - }; + return failingDb(db, "DELETE FROM content_vectors", "injected failure between deletes"); } test("removes orphaned chunks from both tables and returns the count", async () => { @@ -3181,25 +3984,25 @@ describe("cleanupOrphanedVectors atomicity", () => { expect(vecCounts(store.db)).toEqual({ vec: 1, meta: 1 }); const survivor = store.db.prepare(`SELECT hash FROM content_vectors`).get() as { hash: string }; expect(survivor.hash).toBe("keephash"); - const survivorVec = store.db.prepare(`SELECT hash_seq FROM vectors_vec`).get() as { hash_seq: string }; - expect(survivorVec.hash_seq).toBe("keephash_0"); + const survivorVec = store.db.prepare(`SELECT hash, seq FROM ${VEC_ROWS_TABLE}`).get() as { hash: string; seq: number }; + expect(survivorVec).toEqual({ hash: "keephash", seq: 0 }); } finally { await cleanupTestDb(store); } }); - test("rolls back the vectors_vec DELETE when the content_vectors DELETE fails", async () => { + test("rolls back the vector DELETEs when the content_vectors DELETE fails", async () => { const store = await createTestStore(); try { await seedOrphanFixture(store); const db = store.db; - // Without the transaction wrap this used to leave vectors_vec already + // Without the transaction wrap this used to leave the vector table already // purged while content_vectors still claimed the chunks were embedded // (silent desync). expect(() => cleanupOrphanedVectors(makeFailingDb(db))).toThrow("injected failure between deletes"); - // Both tables must be untouched — the vectors_vec DELETE was rolled back. + // Both tables must be untouched — the vector DELETEs were rolled back. expect(vecCounts(db)).toEqual({ vec: 3, meta: 3 }); // The connection is left in a clean state: a plain retry succeeds. @@ -3249,7 +4052,7 @@ describe("cleanupOrphanedVectors atomicity", () => { } } // If the cleanup ran inline instead of inside its own savepoint, the - // vectors_vec DELETE would survive the caught failure and commit with + // vector DELETEs would survive the caught failure and commit with // the outer transaction below. db.prepare(`INSERT INTO content (hash, doc, created_at) VALUES (?, ?, ?)`) .run("outer-survivor", "outer doc", new Date().toISOString()); @@ -3387,7 +4190,7 @@ describe("Vector Table", () => { // Initially no vector table let exists = store.db.prepare(` - SELECT name FROM sqlite_master WHERE type='table' AND name='vectors_vec' + SELECT name FROM sqlite_master WHERE type='table' AND name='${VEC_TABLE}' `).get(); expect(exists).toBeFalsy(); // null or undefined @@ -3395,7 +4198,7 @@ describe("Vector Table", () => { store.ensureVecTable(768); exists = store.db.prepare(` - SELECT name FROM sqlite_master WHERE type='table' AND name='vectors_vec' + SELECT name FROM sqlite_master WHERE type='table' AND name='${VEC_TABLE}' `).get(); expect(exists).toBeTruthy(); @@ -3410,7 +4213,7 @@ describe("Vector Table", () => { // Check dimensions const tableInfo = store.db.prepare(` - SELECT sql FROM sqlite_master WHERE type='table' AND name='vectors_vec' + SELECT sql FROM sqlite_master WHERE type='table' AND name='${VEC_TABLE}' `).get() as { sql: string }; expect(tableInfo.sql).toContain("float[768]"); @@ -3419,39 +4222,65 @@ describe("Vector Table", () => { // Original table should still exist untouched const tableInfoAfter = store.db.prepare(` - SELECT sql FROM sqlite_master WHERE type='table' AND name='vectors_vec' + SELECT sql FROM sqlite_master WHERE type='table' AND name='${VEC_TABLE}' `).get() as { sql: string }; expect(tableInfoAfter.sql).toContain("float[768]"); await cleanupTestDb(store); }); - test("insertEmbedding is idempotent for an existing vec0 hash_seq (#598)", async () => { + test("insertEmbedding replaces the stored vector of an existing chunk (#598)", async () => { const store = await createTestStore(); + const collection = await createTestCollection(); store.ensureVecTable(2); const hash = "existinghashseq"; const first = new Float32Array([0.1, 0.2]); const second = new Float32Array([0.3, 0.4]); const now = new Date().toISOString(); + await insertTestDocument(store.db, collection, { name: "doc", hash }); + store.insertEmbedding(hash, 0, 0, first, "test-model", now); + const rowid = (store.db.prepare(`SELECT id FROM ${VEC_ROWS_TABLE} WHERE hash = ? AND seq = 0`).get(hash) as { id: number }).id; - store.db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${hash}_0`, first); - - // Reproduces sqlite-vec's broken conflict handling: vec0 does not honor OR REPLACE. - expect(() => { - store.db.prepare(`INSERT OR REPLACE INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${hash}_0`, second); - }).toThrow(/UNIQUE constraint failed/i); - - // QMD must therefore use DELETE + INSERT when upserting the vector row. + // vec0 does not honor OR REPLACE: the partition row is deleted and + // re-inserted under the same rowid. expect(() => store.insertEmbedding(hash, 0, 0, second, "test-model", now)).not.toThrow(); - const vectorCount = store.db.prepare(`SELECT COUNT(*) AS count FROM vectors_vec WHERE hash_seq = ?`).get(`${hash}_0`) as { count: number }; + expect(store.db.prepare(`SELECT id FROM ${VEC_ROWS_TABLE} WHERE hash = ? AND seq = 0`).all(hash)).toEqual([{ id: rowid }]); + const vectorCount = store.db.prepare(`SELECT COUNT(*) AS count FROM ${VEC_TABLE}`).get() as { count: number }; const metadataCount = store.db.prepare(`SELECT COUNT(*) AS count FROM content_vectors WHERE hash = ? AND seq = 0`).get(hash) as { count: number }; expect(vectorCount.count).toBe(1); expect(metadataCount.count).toBe(1); + const stored = store.db.prepare(`SELECT embedding FROM ${VEC_TABLE} WHERE rowid = ?`).get(BigInt(rowid)) as { embedding: Uint8Array }; + expect(Array.from(new Float32Array(stored.embedding.buffer, stored.embedding.byteOffset, 2))).toEqual(Array.from(second)); await cleanupTestDb(store); }); + + test("ensureVecTable keeps a dimensionless vector table and its row map when clearing the rows fails", async () => { + const store = await createTestStore(); + try { + // No float[N] in the declaration, so ensureVecTable drops and recreates it. + const dimensionless = `CREATE TABLE ${VEC_TABLE} (collection_id INTEGER, embedding BLOB)`; + store.db.exec(dimensionless); + store.db.prepare(`INSERT INTO ${VEC_ROWS_TABLE} (hash, seq, collection_id) VALUES ('h1', 0, 1)`).run(); + store.db.exec(`CREATE TRIGGER block_vector_rows_delete BEFORE DELETE ON ${VEC_ROWS_TABLE} BEGIN SELECT RAISE(ABORT, 'injected failure clearing vector rows'); END`); + const vecTableSql = () => (store.db.prepare(`SELECT sql FROM sqlite_master WHERE name = ?`).get(VEC_TABLE) as { sql: string } | null | undefined)?.sql; + const rowCount = () => (store.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_ROWS_TABLE}`).get() as { c: number }).c; + + expect(() => store.ensureVecTable(8)).toThrow("injected failure clearing vector rows"); + + expect(vecTableSql()).toBe(dimensionless); + expect(rowCount()).toBe(1); + + store.db.exec(`DROP TRIGGER block_vector_rows_delete`); + store.ensureVecTable(8); + expect(vecTableSql()).toContain("float[8]"); + expect(rowCount()).toBe(0); + } finally { + await cleanupTestDb(store); + } + }); }); // ============================================================================= @@ -3598,6 +4427,23 @@ describe("Integration", () => { // ============================================================================= describe("Vector Search collection filter", () => { + test("an exact vector tie between collections goes the same way whichever is named first", async () => { + const store = await createTestStore(); + const alpha = await createTestCollection({ name: "alpha", pwd: "/test/alpha" }); + const beta = await createTestCollection({ name: "beta", pwd: "/test/beta" }); + store.ensureVecTable(3); + for (const collection of [alpha, beta]) { + await insertTestDocument(store.db, collection, { name: "same", hash: "samehash", body: "Same body", displayPath: "same.md" }); + } + store.insertEmbedding("samehash", 0, 0, new Float32Array([1, 0, 0]), "test", new Date().toISOString()); + + for (const collections of [["alpha", "beta"], ["beta", "alpha"]]) { + const results = await store.searchVec("ignored", "test-model", 1, collections, undefined, [1, 0, 0]); + expect(results.map(r => r.filepath)).toEqual(["qmd://alpha/same.md"]); + } + await cleanupTestDb(store); + }); + test("searchVec finds docs in a small collection crowded by a large one (#791, #803)", async () => { const store = await createTestStore(); const large = await createTestCollection({ name: "large", pwd: "/test/large" }); @@ -3610,10 +4456,10 @@ describe("Vector Search collection filter", () => { queryEmbedding[0] = 1; // 250 nearer neighbours in the large collection. With limit=3: - // - old global k=limit*3=9 never sees `small` - // - a plain multiplier (limit*30=90) still misses it - // - sqlite-vec also caps k at 4096, so multipliers cannot fix tiny - // collections in huge indexes. Collection-scoped exact scan does. + // - an unscoped top-k (k=limit*3=9) never sees `small` + // - sqlite-vec caps k at 4096, so a wider over-fetch cannot fix tiny + // collections in huge indexes; a KNN inside the collection's + // partition does. for (let i = 0; i < 250; i++) { const hash = `largehash${String(i).padStart(3, "0")}`; await insertTestDocument(store.db, large, { @@ -3624,8 +4470,7 @@ describe("Vector Search collection filter", () => { }); const embedding = new Float32Array(dims); embedding[0] = 1; - store.db.prepare(`INSERT INTO content_vectors (hash, seq, pos, model, embedded_at) VALUES (?, 0, 0, 'test', ?)`).run(hash, now); - store.db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${hash}_0`, embedding); + store.insertEmbedding(hash, 0, 0, embedding, 'test', new Date().toISOString()); } const targetHash = "smallhash001"; @@ -3635,12 +4480,11 @@ describe("Vector Search collection filter", () => { body: "Target document in the small collection", displayPath: "target.md", }); - // Farther than the noise vectors — only found via collection-scoped scan. + // Farther than the noise vectors; only a scan of the small partition reaches it. const targetEmbedding = new Float32Array(dims); targetEmbedding[0] = 0.6; targetEmbedding[1] = 0.8; - store.db.prepare(`INSERT INTO content_vectors (hash, seq, pos, model, embedded_at) VALUES (?, 0, 0, 'test', ?)`).run(targetHash, now); - store.db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${targetHash}_0`, targetEmbedding); + store.insertEmbedding(targetHash, 0, 0, targetEmbedding, 'test', new Date().toISOString()); const filtered = await store.searchVec( "ignored — embedding precomputed", @@ -3691,8 +4535,7 @@ describe("Vector Search collection filter", () => { }); const embedding = new Float32Array(dims); embedding[0] = 1; - store.db.prepare(`INSERT INTO content_vectors (hash, seq, pos, model, embedded_at) VALUES (?, 0, 0, 'test', ?)`).run(hash, now); - store.db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${hash}_0`, embedding); + store.insertEmbedding(hash, 0, 0, embedding, 'test', new Date().toISOString()); } const kbHash = "kbhash001"; @@ -3705,8 +4548,7 @@ describe("Vector Search collection filter", () => { const kbEmbedding = new Float32Array(dims); kbEmbedding[0] = 0.55; kbEmbedding[1] = 0.84; - store.db.prepare(`INSERT INTO content_vectors (hash, seq, pos, model, embedded_at) VALUES (?, 0, 0, 'test', ?)`).run(kbHash, now); - store.db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${kbHash}_0`, kbEmbedding); + store.insertEmbedding(kbHash, 0, 0, kbEmbedding, 'test', new Date().toISOString()); const notesHash = "noteshash001"; await insertTestDocument(store.db, notes, { @@ -3718,30 +4560,674 @@ describe("Vector Search collection filter", () => { const notesEmbedding = new Float32Array(dims); notesEmbedding[0] = 0.5; notesEmbedding[1] = 0.87; - store.db.prepare(`INSERT INTO content_vectors (hash, seq, pos, model, embedded_at) VALUES (?, 0, 0, 'test', ?)`).run(notesHash, now); - store.db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${notesHash}_0`, notesEmbedding); + store.insertEmbedding(notesHash, 0, 0, notesEmbedding, 'test', new Date().toISOString()); + + const global = await store.searchVec( + "ignored — embedding precomputed", + "test-model", + 3, + undefined, + undefined, + queryEmbedding, + ); + expect(global).toHaveLength(3); + expect(global.every((r) => r.collectionName === large)).toBe(true); + + const union = await store.searchVec( + "ignored — embedding precomputed", + "test-model", + 3, + [knowledge, notes], + undefined, + queryEmbedding, + ); + expect(union.map(r => r.collectionName).sort()).toEqual([knowledge, notes].sort()); + expect(union.every((r) => r.collectionName !== large)).toBe(true); + + await cleanupTestDb(store); + }); + + test("searchFTS finds a scoped doc the global candidate window does not contain", async () => { + const store = await createTestStore(); + const large = await createTestCollection({ name: "large-fts", pwd: "/test/large-fts" }); + const small = await createTestCollection({ name: "small-fts", pwd: "/test/small-fts" }); + + // 40 short noise docs carrying the term in the title (bm25 title weight + // 4.0), so every one of them outranks the target globally. + for (let i = 0; i < 40; i++) { + await insertTestDocument(store.db, large, { + name: `noise-${i}`, + title: `zebra zebra ${i}`, + body: `# Noise ${i}\n\nzebra zebra zebra, noise document ${i}.`, + displayPath: `noise-${i}.md`, + }); + } + + // One long target doc that mentions the term once, dead last globally. + await insertTestDocument(store.db, small, { + name: "target", + title: "Target", + body: `# Target\n\n${"filler prose without the search term. ".repeat(40)}zebra.`, + displayPath: "target.md", + }); + + // Unscoped, the target is nowhere near the top of the ranking. + const global = store.searchFTS("zebra", 2); + expect(global).toHaveLength(2); + expect(global.every(r => r.collectionName === large)).toBe(true); + + // Scoped, it is the only answer. Taking a global top-(limit * 10) and + // filtering afterwards returned nothing here: all 20 candidates were + // large-fts documents. + const scoped = store.searchFTS("zebra", 2, small); + expect(scoped.map(r => r.displayPath)).toEqual([`${small}/target.md`]); + + await cleanupTestDb(store); + }); + + const DIMS = 8; + const query: number[] = Array(DIMS).fill(0); + query[0] = 1; + + function vector(x: number, y: number): Float32Array { + const embedding = new Float32Array(DIMS); + embedding[0] = x; + embedding[1] = y; + return embedding; + } + + /** Chunk `seq` of `hash` sits at pos seq * 100, with one partition row per active collection of the hash. */ + function insertChunkVectors(store: Store, hash: string, chunks: readonly Float32Array[]): void { + const now = new Date().toISOString(); + chunks.forEach((embedding, seq) => store.insertEmbedding(hash, seq, seq * 100, embedding, "test", now, chunks.length)); + } + + async function insertVecDoc(store: Store, collection: string, hash: string, chunks: readonly Float32Array[]): Promise { + await insertTestDocument(store.db, collection, { name: hash, hash, body: `Document ${hash}`, displayPath: `${hash}.md` }); + insertChunkVectors(store, hash, chunks); + } + + /** Documents orthogonal to the query. */ + async function insertFillerDocs(store: Store, collection: string, prefix: string, count: number): Promise { + for (let i = 0; i < count; i++) { + await insertVecDoc(store, collection, `${prefix}${String(i).padStart(2, "0")}`, [vector(0, 1)]); + } + } + + test("searchVec scoped to two of three collections returns exactly those two", async () => { + const store = await createTestStore(); + const first = await createTestCollection({ name: "first", pwd: "/test/first" }); + const second = await createTestCollection({ name: "second", pwd: "/test/second" }); + const third = await createTestCollection({ name: "third", pwd: "/test/third" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, first, "firsthash", [vector(1, 0)]); + await insertVecDoc(store, second, "secondhash", [vector(0.9, 0.44)]); + await insertVecDoc(store, third, "thirdhash", [vector(0.8, 0.6)]); + await insertFillerDocs(store, second, "secondfill", 4); + + const results = await store.searchVec("ignored", "test-model", 3, [first, third], undefined, query); + expect(results.map((r) => r.collectionName).sort()).toEqual([first, third].sort()); + + await cleanupTestDb(store); + }); + + test("searchVec runs one KNN statement per collection in scope and an unpartitioned one otherwise", async () => { + const store = await createTestStore(); + const first = await createTestCollection({ name: "first", pwd: "/test/first" }); + const second = await createTestCollection({ name: "second", pwd: "/test/second" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, first, "firsthash", [vector(1, 0)]); + await insertVecDoc(store, second, "secondhash", [vector(0.6, 0.8)]); + const prepare = vi.spyOn(store.db, "prepare"); + const knnStatements = () => prepare.mock.calls.map((call) => String(call[0])).filter((sql) => sql.includes("MATCH")); + try { + const scoped = await store.searchVec("ignored", "test-model", 3, [first, second], undefined, query); + expect(scoped.map((r) => r.hash)).toEqual(["firsthash", "secondhash"]); + expect(knnStatements()).toHaveLength(1); + expect(knnStatements()[0]).toContain("collection_id = ?"); + expect(knnStatements()[0]).not.toContain(" IN "); + + prepare.mockClear(); + await store.searchVec("ignored", "test-model", 3, undefined, undefined, query); + expect(knnStatements()).toHaveLength(1); + expect(knnStatements()[0]).not.toContain("collection_id"); + } finally { + prepare.mockRestore(); + } + + await cleanupTestDb(store); + }); + + test("searchVec returns a hash shared with an out-of-scope collection once, named for the scoped collection", async () => { + const store = await createTestStore(); + const included = await createTestCollection({ name: "included", pwd: "/test/included" }); + const excluded = await createTestCollection({ name: "excluded", pwd: "/test/excluded" }); + store.ensureVecTable(DIMS); + await insertTestDocument(store.db, excluded, { name: "sharedhash", hash: "sharedhash", body: "Document sharedhash", displayPath: "sharedhash.md" }); + await insertTestDocument(store.db, included, { name: "copy", hash: "sharedhash", body: "Document sharedhash", displayPath: "copy.md" }); + insertChunkVectors(store, "sharedhash", [vector(1, 0)]); + await insertFillerDocs(store, excluded, "excludedfill", 3); + + const results = await store.searchVec("ignored", "test-model", 5, included, undefined, query); + expect(results).toHaveLength(1); + expect(results[0]!.collectionName).toBe(included); + expect(results[0]!.displayPath).toBe(`${included}/copy.md`); + + await cleanupTestDb(store); + }); + + test("searchVec returns a hash shared by two in-scope collections once per collection", async () => { + const store = await createTestStore(); + const alpha = await createTestCollection({ name: "alpha", pwd: "/test/alpha" }); + const beta = await createTestCollection({ name: "beta", pwd: "/test/beta" }); + store.ensureVecTable(DIMS); + await insertTestDocument(store.db, alpha, { name: "shared", hash: "sharedhash", body: "Document sharedhash", displayPath: "shared.md" }); + await insertTestDocument(store.db, beta, { name: "shared", hash: "sharedhash", body: "Document sharedhash", displayPath: "shared.md" }); + insertChunkVectors(store, "sharedhash", [vector(1, 0)]); + + const results = await store.searchVec("ignored", "test-model", 5, [alpha, beta], undefined, query); + expect(results.map((r) => r.collectionName).sort()).toEqual([alpha, beta].sort()); + expect(new Set(results.map((r) => r.filepath)).size).toBe(2); + + await cleanupTestDb(store); + }); + + test("searchVec ignores a hash whose only in-scope row is inactive", async () => { + const store = await createTestStore(); + const included = await createTestCollection({ name: "included", pwd: "/test/included" }); + const excluded = await createTestCollection({ name: "excluded", pwd: "/test/excluded" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, excluded, "sharedhash", [vector(1, 0)]); + await insertTestDocument(store.db, included, { name: "stale", hash: "sharedhash", body: "Document sharedhash", displayPath: "stale.md", active: 0 }); + + const results = await store.searchVec("ignored", "test-model", 5, included, undefined, query); + expect(results).toEqual([]); + + await cleanupTestDb(store); + }); + + test("searchVec returns a document once when its hash has an active and an inactive row in scope", async () => { + const store = await createTestStore(); + const included = await createTestCollection({ name: "included", pwd: "/test/included" }); + const outside = await createTestCollection({ name: "outside", pwd: "/test/outside" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, included, "duphash", [vector(1, 0)]); + await insertFillerDocs(store, outside, "outsidefill", 3); + await insertTestDocument(store.db, included, { name: "stale", hash: "duphash", body: "Document duphash", displayPath: "stale.md", active: 0 }); + + const results = await store.searchVec("ignored", "test-model", 5, included, undefined, query); + expect(results).toHaveLength(1); + expect(results[0]!.displayPath).toBe(`${included}/duphash.md`); + + await cleanupTestDb(store); + }); + + test("searchVec returns empty for a scoped collection that has documents but no vectors", async () => { + const store = await createTestStore(); + const embedded = await createTestCollection({ name: "embedded", pwd: "/test/embedded" }); + const bare = await createTestCollection({ name: "bare", pwd: "/test/bare" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, embedded, "embeddedhash", [vector(1, 0)]); + await insertFillerDocs(store, embedded, "embeddedfill", 2); + await insertTestDocument(store.db, bare, { name: "plain", body: "Not embedded", displayPath: "plain.md" }); + + const results = await store.searchVec("ignored", "test-model", 3, bare, undefined, query); + expect(results).toEqual([]); + + await cleanupTestDb(store); + }); + + test("searchVec returns empty for a collection name that no document carries", async () => { + const store = await createTestStore(); + const embedded = await createTestCollection({ name: "embedded", pwd: "/test/embedded" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, embedded, "embeddedhash", [vector(1, 0)]); + + expect(await store.searchVec("ignored", "test-model", 3, "never-indexed", undefined, query)).toEqual([]); + const mixed = await store.searchVec("ignored", "test-model", 3, ["never-indexed", embedded], undefined, query); + expect(mixed.map((r) => r.hash)).toEqual(["embeddedhash"]); + + await cleanupTestDb(store); + }); + + test("searchVec accepts a scoped limit above sqlite-vec's k cap", async () => { + const store = await createTestStore(); + const embedded = await createTestCollection({ name: "embedded", pwd: "/test/embedded" }); + const outside = await createTestCollection({ name: "outside", pwd: "/test/outside" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, embedded, "nearhash", [vector(1, 0)]); + await insertVecDoc(store, embedded, "farhash", [vector(0.6, 0.8)]); + await insertFillerDocs(store, outside, "outsidefill", 4); + + const scoped = await store.searchVec("ignored", "test-model", 2000, embedded, undefined, query); + expect(scoped.map((r) => r.hash)).toEqual(["nearhash", "farhash"]); + const unscoped = await store.searchVec("ignored", "test-model", 2000, undefined, undefined, query); + expect(unscoped).toHaveLength(6); + + await cleanupTestDb(store); + }); + + test("searchVec treats an empty collection list like no scope", async () => { + const store = await createTestStore(); + const first = await createTestCollection({ name: "first", pwd: "/test/first" }); + const second = await createTestCollection({ name: "second", pwd: "/test/second" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, first, "firsthash", [vector(1, 0)]); + await insertVecDoc(store, second, "secondhash", [vector(0.6, 0.8)]); + + const unscoped = await store.searchVec("ignored", "test-model", 3, undefined, undefined, query); + const emptyScope = await store.searchVec("ignored", "test-model", 3, [], undefined, query); + expect(unscoped.map((r) => r.hash)).toEqual(["firsthash", "secondhash"]); + expect(emptyScope).toEqual(unscoped); + + await cleanupTestDb(store); + }); + + test("searchVec collapses a scoped over-fetch to one row per file at its best chunk", async () => { + const store = await createTestStore(); + const alpha = await createTestCollection({ name: "alpha", pwd: "/test/alpha" }); + const beta = await createTestCollection({ name: "beta", pwd: "/test/beta" }); + const outside = await createTestCollection({ name: "outside", pwd: "/test/outside" }); + store.ensureVecTable(DIMS); + await insertFillerDocs(store, outside, "outsidefill", 30); + const near = vector(1, 0); + const far = vector(0.6, 0.8); + for (let i = 0; i < 3; i++) { + await insertVecDoc(store, alpha, `long${i}`, Array.from({ length: 10 }, () => near)); + } + for (let i = 0; i < 20; i++) { + await insertVecDoc(store, i % 2 === 0 ? alpha : beta, `short${String(i).padStart(2, "0")}`, [far]); + } + + const results = await store.searchVec("ignored", "test-model", 10, [alpha, beta], undefined, query); + expect(results).toHaveLength(10); + expect(results.slice(0, 3).map((r) => r.hash).sort()).toEqual(["long0", "long1", "long2"]); + expect(new Set(results.map((r) => r.filepath)).size).toBe(results.length); + + await cleanupTestDb(store); + }); + + test("searchVec fills its limit when one document's chunks take every candidate slot", async () => { + const store = await createTestStore(); + const book = await createTestCollection({ name: "book", pwd: "/test/book" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, book, "manychunks", Array.from({ length: 20 }, () => vector(1, 0))); + await insertVecDoc(store, book, "seconddoc", [vector(0.2, 1)]); + + for (const scope of [book, undefined]) { + for (const limit of [2, 5]) { + const results = await store.searchVec("ignored", "test-model", limit, scope, undefined, query); + expect(results.map((r) => r.hash)).toEqual(["manychunks", "seconddoc"]); + } + } + + await cleanupTestDb(store); + }); + + test("searchVec rows carry the vec source, a cosine score, and the best chunk position", async () => { + const store = await createTestStore(); + const embedded = await createTestCollection({ name: "embedded", pwd: "/test/embedded" }); + const outside = await createTestCollection({ name: "outside", pwd: "/test/outside" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, embedded, "twochunks", [vector(0, 1), vector(0.6, 0.8)]); + await insertFillerDocs(store, outside, "outsidefill", 3); + + const results = await store.searchVec("ignored", "test-model", 3, embedded, undefined, query); + expect(results).toHaveLength(1); + expect(results[0]!.source).toBe("vec"); + expect(results[0]!.chunkPos).toBe(100); + expect(results[0]!.score).toBeCloseTo(0.6, 5); + + await cleanupTestDb(store); + }); + + test("searchVec union keeps a sibling collection reachable behind a long in-scope document", async () => { + const store = await createTestStore(); + const small = await createTestCollection({ name: "small", pwd: "/test/small" }); + const large = await createTestCollection({ name: "large", pwd: "/test/large" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, large, "longhash", Array.from({ length: 12 }, () => vector(1, 0))); + await insertVecDoc(store, small, "smallhash", [vector(0.6, 0.8)]); + const outside = await createTestCollection({ name: "outside", pwd: "/test/outside" }); + await insertFillerDocs(store, outside, "outsidefill", 4); + + const results = await store.searchVec("ignored", "test-model", 3, [small, large], undefined, query); + expect(results.map((r) => r.collectionName).sort()).toEqual([large, small].sort()); + expect(results.some((r) => r.hash === "smallhash")).toBe(true); + + await cleanupTestDb(store); + }); + + test("searchVec scoped to a collection still answers when one chunk row has no vector", async () => { + const store = await createTestStore(); + const partial = await createTestCollection({ name: "partial", pwd: "/test/partial" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, partial, "wholehash0", [vector(1, 0)]); + await insertVecDoc(store, partial, "wholehash1", [vector(0.9, 0.44)]); + await insertVecDoc(store, partial, "ghosthash", [vector(0.8, 0.6)]); + const other = await createTestCollection({ name: "other", pwd: "/test/other" }); + await insertFillerDocs(store, other, "otherfill", 4); + const ghost = store.db.prepare(`SELECT id FROM ${VEC_ROWS_TABLE} WHERE hash = 'ghosthash' AND seq = 0`).get() as { id: number }; + deletePartitionRows(store.db, [ghost.id]); + + const scoped = await store.searchVec("ignored", "test-model", 5, partial, undefined, query); + const unscoped = await store.searchVec("ignored", "test-model", 5, undefined, undefined, query); + expect(scoped.map((r) => r.hash)).toEqual(["wholehash0", "wholehash1"]); + expect(unscoped.filter((r) => r.collectionName === partial).map((r) => r.hash)).toEqual(["wholehash0", "wholehash1"]); + + await cleanupTestDb(store); + }); + + test("searchVec scoped to a collection that holds most of the index returns its nearest documents", async () => { + const store = await createTestStore(); + const big = await createTestCollection({ name: "big", pwd: "/test/big" }); + const other = await createTestCollection({ name: "other", pwd: "/test/other" }); + store.ensureVecTable(DIMS); + for (let i = 0; i < 30; i++) { + await insertVecDoc(store, big, `bighash${String(i).padStart(2, "0")}`, [vector(1, 0.02 + i * 0.01)]); + } + for (let i = 0; i < 12; i++) { + await insertVecDoc(store, other, `otherhash${String(i).padStart(2, "0")}`, [vector(1, 0)]); + } + + const results = await store.searchVec("ignored", "test-model", 3, big, undefined, query); + expect(results.map((r) => r.hash)).toEqual(["bighash00", "bighash01", "bighash02"]); + + await cleanupTestDb(store); + }); + + test("searchVec scoped to a majority collection is not starved by nearer vectors outside it", async () => { + const store = await createTestStore(); + const big = await createTestCollection({ name: "big", pwd: "/test/big" }); + const other = await createTestCollection({ name: "other", pwd: "/test/other" }); + store.ensureVecTable(DIMS); + for (let i = 0; i < 100; i++) { + await insertVecDoc(store, big, `bighash${String(i).padStart(3, "0")}`, [vector(0.6, 0.8)]); + } + for (let i = 0; i < 80; i++) { + await insertVecDoc(store, other, `otherhash${String(i).padStart(3, "0")}`, [vector(1, 0)]); + } + + const results = await store.searchVec("ignored", "test-model", 3, big, undefined, query); + expect(results).toHaveLength(3); + expect(results.every((r) => r.collectionName === big)).toBe(true); + + await cleanupTestDb(store); + }); + + test("searchVec scoped to a majority collection larger than the over-fetch returns its nearest documents", async () => { + const store = await createTestStore(); + const big = await createTestCollection({ name: "big", pwd: "/test/big" }); + const other = await createTestCollection({ name: "other", pwd: "/test/other" }); + store.ensureVecTable(DIMS); + for (let i = 0; i < 300; i++) { + await insertVecDoc(store, big, `bighash${String(i).padStart(3, "0")}`, [vector(1, 0.02 + i * 0.001)]); + } + for (let i = 0; i < 20; i++) { + await insertVecDoc(store, other, `otherhash${String(i).padStart(2, "0")}`, [vector(1, 0)]); + } + + const results = await store.searchVec("ignored", "test-model", 3, big, undefined, query); + expect(results.map((r) => r.hash)).toEqual(["bighash000", "bighash001", "bighash002"]); + + await cleanupTestDb(store); + }); + + test("renameCollection keeps every vector reachable under the new name", async () => { + const store = await createTestStore(); + const before = await createTestCollection({ name: "before", pwd: "/test/before" }); + const other = await createTestCollection({ name: "other", pwd: "/test/other" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, before, "renamedhash", [vector(1, 0)]); + await insertFillerDocs(store, other, "otherfill", 3); + + renameCollection(store.db, before, "after"); + + const renamed = await store.searchVec("ignored", "test-model", 3, "after", undefined, query); + expect(renamed.map((r) => r.hash)).toEqual(["renamedhash"]); + expect(renamed[0]!.collectionName).toBe("after"); + expect(await store.searchVec("ignored", "test-model", 3, before, undefined, query)).toEqual([]); + expect(vectorRowCount(store, "after")).toBe(1); + + await cleanupTestDb(store); + }); + + test("searchVec drops a result whose content row vanishes after its document resolved", async () => { + const store = await createTestStore(); + const collection = await createTestCollection({ name: "racing", pwd: "/test/racing" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, collection, "nearhash", [vector(1, 0)]); + await insertVecDoc(store, collection, "vanishhash", [vector(0.9, 0.44)]); + await insertVecDoc(store, collection, "farhash", [vector(0.8, 0.6)]); + + // Replays another process's orphaned-content cleanup landing between + // document resolution and the body read. + const bodySql = "SELECT CASE WHEN length(CAST(doc AS BLOB)) <= 262144 THEN doc ELSE substr(doc, 1, 262144) END AS doc FROM content WHERE hash = ?"; + const racing: Database = { + prepare: (sql: string) => { + const statement = store.db.prepare(sql); + if (!sql.includes(bodySql)) return statement; + return { + run: (...params: SQLiteValue[]) => statement.run(...params), + all: (...params: SQLiteValue[]) => statement.all(...params), + iterate: (...params: SQLiteValue[]) => statement.iterate(...params), + get: (...params: SQLiteValue[]) => { + if (params[0] === "vanishhash") store.db.prepare(`DELETE FROM content WHERE hash = ?`).run("vanishhash"); + return statement.get(...params); + }, + }; + }, + transaction: (fn) => store.db.transaction(fn), + exec: (sql: string) => store.db.exec(sql), + loadExtension: (path: string) => store.db.loadExtension(path), + close: () => store.db.close(), + }; + + const results = await searchVec(racing, "ignored", "test-model", 3, undefined, undefined, query); + + expect(results.map((r) => r.hash)).toEqual(["nearhash", "farhash"]); + expect(results.map((r) => r.body)).toEqual(["Document nearhash", "Document farhash"]); + + await cleanupTestDb(store); + }); + + function documentsIn(store: Store, collection: string): number { + return (store.db.prepare(`SELECT COUNT(*) AS c FROM documents WHERE collection = ?`).get(collection) as { c: number }).c; + } + + function storeCollectionNames(store: Store, ...names: string[]): string[] { + return (store.db.prepare(`SELECT name FROM store_collections WHERE name IN (${names.map(() => "?").join(", ")}) ORDER BY name`) + .all(...names) as { name: string }[]).map((row) => row.name); + } + + test("renameCollection onto a removed collection's leftover partition drops the leftover and keeps the renamed vectors", async () => { + const store = await createTestStore(); + const before = await createTestCollection({ name: "before", pwd: "/test/before" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, before, "renamedhash", [vector(1, 0)]); + const beforeId = resolveCollectionId(store.db, before)!; + // A removed collection whose partition and id outlived its documents. + await insertVecDoc(store, "after", "leftoverhash", [vector(0.9, 0.44), vector(0.8, 0.6)]); + store.db.prepare(`DELETE FROM documents WHERE collection = 'after'`).run(); + expect(vectorRowCount(store, "after")).toBe(2); + + renameCollection(store.db, before, "after"); + + const renamed = await store.searchVec("ignored", "test-model", 3, "after", undefined, query); + expect(renamed.map((r) => r.hash)).toEqual(["renamedhash"]); + expect(renamed[0]!.collectionName).toBe("after"); + expect(resolveCollectionId(store.db, "after")).toBe(beforeId); + expect(resolveCollectionId(store.db, before)).toBeUndefined(); + expect(vectorRowCount(store, "after")).toBe(1); + expect((store.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_TABLE}`).get() as { c: number }).c).toBe(1); + expect(documentsIn(store, "after")).toBe(1); + expect(storeCollectionNames(store, before, "after")).toEqual(["after"]); + + await cleanupTestDb(store); + }); + + test("renameCollection rolls back every table when a later step fails", async () => { + const store = await createTestStore(); + const before = await createTestCollection({ name: "before", pwd: "/test/before" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, before, "renamedhash", [vector(1, 0)]); + const beforeId = resolveCollectionId(store.db, before)!; + + const failing = failingDb(store.db, `UPDATE ${VEC_COLLECTION_IDS_TABLE} SET name = ?`, "injected failure renaming the vector id"); + expect(() => renameCollection(failing, before, "after")).toThrow("injected failure renaming the vector id"); + + expect(documentsIn(store, before)).toBe(1); + expect(documentsIn(store, "after")).toBe(0); + expect(resolveCollectionId(store.db, before)).toBe(beforeId); + expect(resolveCollectionId(store.db, "after")).toBeUndefined(); + expect(storeCollectionNames(store, before, "after")).toEqual([before]); + expect((await store.searchVec("ignored", "test-model", 3, before, undefined, query)).map((r) => r.hash)).toEqual(["renamedhash"]); + + // The connection is left clean: a plain retry renames everything. + renameCollection(store.db, before, "after"); + expect(documentsIn(store, "after")).toBe(1); + expect(resolveCollectionId(store.db, "after")).toBe(beforeId); + expect(storeCollectionNames(store, before, "after")).toEqual(["after"]); + + await cleanupTestDb(store); + }); + + test("renameCollection onto an existing collection moves nothing and keeps the target's vectors", async () => { + const store = await createTestStore(); + const before = await createTestCollection({ name: "before", pwd: "/test/before" }); + const taken = await createTestCollection({ name: "taken", pwd: "/test/taken" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, before, "renamedhash", [vector(1, 0)]); + await insertVecDoc(store, taken, "takenhash", [vector(0.9, 0.44)]); + const takenId = resolveCollectionId(store.db, taken)!; + + expect(() => renameCollection(store.db, before, taken)).toThrow(`Collection '${taken}' already exists`); + + expect(documentsIn(store, before)).toBe(1); + expect(documentsIn(store, taken)).toBe(1); + expect(resolveCollectionId(store.db, taken)).toBe(takenId); + expect(vectorRowCount(store, taken)).toBe(1); + expect((await store.searchVec("ignored", "test-model", 3, taken, undefined, query)).map((r) => r.hash)).toEqual(["takenhash"]); + + await cleanupTestDb(store); + }); + + test("renameCollection refuses a target whose leftover id cannot be dropped yet and changes nothing", async () => { + const store = await createTestStore(); + const before = await createTestCollection({ name: "before", pwd: "/test/before" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, before, "renamedhash", [vector(1, 0)]); + const beforeId = resolveCollectionId(store.db, before)!; + store.db.prepare(`INSERT INTO ${VEC_COLLECTION_IDS_TABLE} (name) VALUES ('after')`).run(); + const leftoverId = resolveCollectionId(store.db, "after")!; + // A legacy table awaiting migration: partitions are not dropped until it runs. + store.db.exec(`CREATE VIRTUAL TABLE ${LEGACY_VEC_TABLE} USING vec0(hash_seq TEXT PRIMARY KEY, embedding float[${DIMS}] distance_metric=cosine)`); + + expect(() => renameCollection(store.db, before, "after")).toThrow(/'after'.*sqlite-vec/s); + + expect(documentsIn(store, before)).toBe(1); + expect(documentsIn(store, "after")).toBe(0); + expect(resolveCollectionId(store.db, before)).toBe(beforeId); + expect(resolveCollectionId(store.db, "after")).toBe(leftoverId); + expect(storeCollectionNames(store, before, "after")).toEqual([before]); + + await cleanupTestDb(store); + }); + + test("removeCollection drops the collection's partition and leaves the others", async () => { + const store = await createTestStore(); + const gone = await createTestCollection({ name: "gone", pwd: "/test/gone" }); + const kept = await createTestCollection({ name: "kept", pwd: "/test/kept" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, gone, "gonehash", [vector(1, 0), vector(0.9, 0.1)]); + await insertVecDoc(store, kept, "kepthash", [vector(0.6, 0.8)]); + + removeCollection(store.db, gone); + + expect(vectorRowCount(store, gone)).toBe(0); + expect(resolveCollectionId(store.db, gone)).toBeUndefined(); + expect((store.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_TABLE}`).get() as { c: number }).c).toBe(1); + const results = await store.searchVec("ignored", "test-model", 5, undefined, undefined, query); + expect(results.map((r) => r.hash)).toEqual(["kepthash"]); + + await cleanupTestDb(store); + }); + + test("removeCollection rolls back the partition delete when the documents delete fails", async () => { + const store = await createTestStore(); + const gone = await createTestCollection({ name: "gone", pwd: "/test/gone" }); + const kept = await createTestCollection({ name: "kept", pwd: "/test/kept" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, gone, "gonehash", [vector(1, 0), vector(0.9, 0.1)]); + await insertVecDoc(store, kept, "kepthash", [vector(0.6, 0.8)]); + const goneId = resolveCollectionId(store.db, gone)!; + const partitionVectors = (id: number) => + (store.db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_TABLE} WHERE collection_id = ?`).get(vecInteger(id)) as { c: number }).c; + const documentsIn = (collection: string) => + (store.db.prepare(`SELECT COUNT(*) AS c FROM documents WHERE collection = ?`).get(collection) as { c: number }).c; + + const failing = failingDb(store.db, "DELETE FROM documents WHERE collection = ?", "injected failure after the partition delete"); + expect(() => removeCollection(failing, gone)).toThrow("injected failure after the partition delete"); + + // The partition delete ran first in the same transaction; none of it survives the rollback. + expect(resolveCollectionId(store.db, gone)).toBe(goneId); + expect(vectorRowCount(store, gone)).toBe(2); + expect(partitionVectors(goneId)).toBe(2); + expect(documentsIn(gone)).toBe(1); + expect((await store.searchVec("ignored", "test-model", 5, gone, undefined, query)).map((r) => r.hash)).toEqual(["gonehash"]); + + // The connection is left clean: a plain retry removes everything. + removeCollection(store.db, gone); + expect(resolveCollectionId(store.db, gone)).toBeUndefined(); + expect(vectorRowCount(store, gone)).toBe(0); + expect(partitionVectors(goneId)).toBe(0); + expect(documentsIn(gone)).toBe(0); + expect(vectorRowCount(store, kept)).toBe(1); + + await cleanupTestDb(store); + }); - const global = await store.searchVec( - "ignored — embedding precomputed", - "test-model", - 3, - undefined, - undefined, - queryEmbedding, - ); - expect(global).toHaveLength(3); - expect(global.every((r) => r.collectionName === large)).toBe(true); + test("clearAllEmbeddings for one collection keeps a shared hash in every collection", async () => { + const store = await createTestStore(); + const cleared = await createTestCollection({ name: "cleared", pwd: "/test/cleared" }); + const other = await createTestCollection({ name: "other", pwd: "/test/other" }); + store.ensureVecTable(DIMS); + await insertTestDocument(store.db, cleared, { name: "shared", hash: "sharedhash", body: "Document sharedhash", displayPath: "shared.md" }); + await insertTestDocument(store.db, other, { name: "shared", hash: "sharedhash", body: "Document sharedhash", displayPath: "shared.md" }); + insertChunkVectors(store, "sharedhash", [vector(1, 0)]); + await insertVecDoc(store, cleared, "onlyhash", [vector(0.9, 0.44)]); - const union = await store.searchVec( - "ignored — embedding precomputed", - "test-model", - 3, - [knowledge, notes], - undefined, - queryEmbedding, - ); - expect(union.map(r => r.collectionName).sort()).toEqual([knowledge, notes].sort()); - expect(union.every((r) => r.collectionName !== large)).toBe(true); + clearAllEmbeddings(store.db, cleared); + + expect((await store.searchVec("ignored", "test-model", 5, other, undefined, query)).map((r) => r.hash)).toEqual(["sharedhash"]); + expect((await store.searchVec("ignored", "test-model", 5, cleared, undefined, query)).map((r) => r.hash)).toEqual(["sharedhash"]); + expect(store.db.prepare(`SELECT COUNT(*) AS c FROM content_vectors WHERE hash = 'onlyhash'`).get()).toEqual({ c: 0 }); + expect(vectorRowCount(store, cleared)).toBe(1); + + await cleanupTestDb(store); + }); + + test("clearAllEmbeddings for the whole index rolls back every table when the vec0 drop fails", async () => { + const store = await createTestStore(); + const collection = await createTestCollection({ name: "cleared", pwd: "/test/cleared" }); + store.ensureVecTable(DIMS); + await insertVecDoc(store, collection, "firsthash", [vector(1, 0), vector(0.9, 0.44)]); + await insertVecDoc(store, collection, "secondhash", [vector(0.8, 0.6)]); + const count = (table: string) => (store.db.prepare(`SELECT COUNT(*) AS c FROM ${table}`).get() as { c: number }).c; + const counts = () => ({ contentVectors: count("content_vectors"), rows: count(VEC_ROWS_TABLE), vectors: count(VEC_TABLE) }); + expect(counts()).toEqual({ contentVectors: 3, rows: 3, vectors: 3 }); + + const failing = failingDb(store.db, `DROP TABLE IF EXISTS ${VEC_TABLE}`, "injected failure dropping the vec0 table"); + expect(() => clearAllEmbeddings(failing)).toThrow("injected failure dropping the vec0 table"); + + expect(counts()).toEqual({ contentVectors: 3, rows: 3, vectors: 3 }); + expect((await store.searchVec("ignored", "test-model", 5, collection, undefined, query)).map((r) => r.hash)).toEqual(["firsthash", "secondhash"]); + + // The connection is left clean: a plain retry clears everything. + clearAllEmbeddings(store.db); + expect(count("content_vectors")).toBe(0); + expect(count(VEC_ROWS_TABLE)).toBe(0); + expect(store.db.prepare(`SELECT name FROM sqlite_master WHERE name = ?`).get(VEC_TABLE)).toBeFalsy(); await cleanupTestDb(store); }); @@ -3760,7 +5246,7 @@ describe.skipIf(!!process.env.CI)("LlamaCpp Integration", () => { body: "Some content", }); - // No vectors_vec table exists, should return empty + // No vector table exists, should return empty const results = await store.searchVec("query", "embeddinggemma", 10); expect(results).toHaveLength(0); @@ -3783,8 +5269,7 @@ describe.skipIf(!!process.env.CI)("LlamaCpp Integration", () => { // Create vector table and insert a vector store.ensureVecTable(768); const embedding = Array(768).fill(0).map(() => Math.random()); - store.db.prepare(`INSERT INTO content_vectors (hash, seq, pos, model, embedded_at) VALUES (?, 0, 0, 'test', ?)`).run(hash, new Date().toISOString()); - store.db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${hash}_0`, new Float32Array(embedding)); + store.insertEmbedding(hash, 0, 0, new Float32Array(embedding), 'test', new Date().toISOString()); const results = await store.searchVec("test query", "embeddinggemma", 10); expect(results).toHaveLength(1); @@ -3815,14 +5300,12 @@ describe.skipIf(!!process.env.CI)("LlamaCpp Integration", () => { body: "Content in collection two", }); - // Create vectors_vec table with correct dimensions (768 for embeddinggemma) + // Create the vector table with correct dimensions (768 for embeddinggemma) store.ensureVecTable(768); const embedding1 = Array(768).fill(0).map(() => Math.random()); const embedding2 = Array(768).fill(0).map(() => Math.random()); - store.db.prepare(`INSERT INTO content_vectors (hash, seq, pos, model, embedded_at) VALUES (?, 0, 0, 'test', ?)`).run(hash1, new Date().toISOString()); - store.db.prepare(`INSERT INTO content_vectors (hash, seq, pos, model, embedded_at) VALUES (?, 0, 0, 'test', ?)`).run(hash2, new Date().toISOString()); - store.db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${hash1}_0`, new Float32Array(embedding1)); - store.db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${hash2}_0`, new Float32Array(embedding2)); + store.insertEmbedding(hash1, 0, 0, new Float32Array(embedding1), 'test', new Date().toISOString()); + store.insertEmbedding(hash2, 0, 0, new Float32Array(embedding2), 'test', new Date().toISOString()); // Search without filter - should return both const allResults = await store.searchVec("content", "embeddinggemma", 10); @@ -3855,8 +5338,7 @@ describe.skipIf(!!process.env.CI)("LlamaCpp Integration", () => { // Create vector table and insert a test vector store.ensureVecTable(768); const embedding = Array(768).fill(0).map(() => Math.random()); - store.db.prepare(`INSERT INTO content_vectors (hash, seq, pos, model, embedded_at) VALUES (?, 0, 0, 'test', ?)`).run(hash, new Date().toISOString()); - store.db.prepare(`INSERT INTO vectors_vec (hash_seq, embedding) VALUES (?, ?)`).run(`${hash}_0`, new Float32Array(embedding)); + store.insertEmbedding(hash, 0, 0, new Float32Array(embedding), 'test', new Date().toISOString()); // This should complete quickly (not hang) due to the two-step fix // The old code with JOINs in the sqlite-vec query would hang indefinitely @@ -4191,6 +5673,169 @@ describe("Embedding batching", () => { } }); + const legacyModel = "hf:test/legacy-sample.gguf"; + // createFakeEmbedLlm().embed returns this vector for every text. + const matchingVector = new Float32Array([0.1, 0.2, 0.3]); + const otherVector = new Float32Array([0.3, -0.2, 0.1]); + + function addLegacyChunk( + store: Store, + hash: string, + seq: number, + opts: { model?: string; fingerprint?: string; vector?: Float32Array } = {}, + ): void { + const model = opts.model ?? legacyModel; + const fingerprint = opts.fingerprint ?? ""; + const embeddedAt = new Date(0).toISOString(); + if (opts.vector) { + store.insertEmbedding(hash, seq, 0, opts.vector, model, embeddedAt, 3, fingerprint); + } else { + store.db.prepare(` + INSERT INTO content_vectors (hash, seq, pos, model, embed_fingerprint, total_chunks, embedded_at) + VALUES (?, ?, 0, ?, ?, 3, ?) + `).run(hash, seq, model, fingerprint, embeddedAt); + } + } + + test("legacy fingerprint adoption samples the lowest hash and seq when active documents share content", async () => { + const store = await createTestStore(); + const db = store.db; + const fakeLlm = createFakeEmbedLlm(); + store.llm = fakeLlm as any; + + try { + store.ensureVecTable(3); + await insertTestDocument(db, "docs", { hash: "hash-b", body: "Other legacy body.", displayPath: "other.md" }); + addLegacyChunk(store, "hash-b", 0, { vector: otherVector }); + // One body behind two active paths and an inactive one, with seqs inserted + // out of order so rowid order disagrees with ORDER BY hash, seq. + const body = "Shared legacy body without a heading."; + await insertTestDocument(db, "docs", { hash: "hash-a", body, displayPath: "z.md" }); + await insertTestDocument(db, "docs", { hash: "hash-a", body, displayPath: "b.md" }); + await insertTestDocument(db, "docs", { hash: "hash-a", body, displayPath: "a.md", active: 0 }); + addLegacyChunk(store, "hash-a", 2); + addLegacyChunk(store, "hash-a", 1); + addLegacyChunk(store, "hash-a", 0, { vector: matchingVector }); + + const result = await maybeAdoptLegacyEmbeddingFingerprint(store, legacyModel); + + expect(result).toMatchObject({ checked: true, adopted: 4 }); + expect(result.reason).toMatch(/^sample hash-a_0 matched/); + // The title comes from the first active path (b.md), not the inactive a.md. + expect(fakeLlm.embedCalls.map(call => call.text)).toEqual([formatDocForEmbedding(body, "b", legacyModel)]); + } finally { + await cleanupTestDb(store); + } + }); + + test("legacy fingerprint adoption skips rows without an active document or content", async () => { + const store = await createTestStore(); + const db = store.db; + const fakeLlm = createFakeEmbedLlm(); + store.llm = fakeLlm as any; + + try { + store.ensureVecTable(3); + await insertTestDocument(db, "docs", { hash: "hash-a", body: "Inactive body.", displayPath: "inactive.md", active: 0 }); + addLegacyChunk(store, "hash-a", 0, { vector: otherVector }); + await insertTestDocument(db, "docs", { hash: "hash-b", body: "Missing body.", displayPath: "missing.md" }); + addLegacyChunk(store, "hash-b", 0, { vector: otherVector }); + // With foreign keys on, deleting content cascades to the document; turn + // them off so an active document is left pointing at missing content. + db.exec("PRAGMA foreign_keys = OFF"); + db.prepare(`DELETE FROM content WHERE hash = ?`).run("hash-b"); + db.exec("PRAGMA foreign_keys = ON"); + + const skipped = await maybeAdoptLegacyEmbeddingFingerprint(store, legacyModel); + + expect(skipped).toEqual({ checked: false, adopted: 0, reason: "2 legacy docs have no active sample" }); + expect(fakeLlm.embedCalls).toHaveLength(0); + + await insertTestDocument(db, "docs", { hash: "hash-c", body: "Active body.", displayPath: "active.md" }); + addLegacyChunk(store, "hash-c", 0, { vector: matchingVector }); + + const result = await maybeAdoptLegacyEmbeddingFingerprint(store, legacyModel); + + expect(result).toMatchObject({ checked: true, adopted: 3 }); + expect(result.reason).toMatch(/^sample hash-c_0 matched/); + } finally { + await cleanupTestDb(store); + } + }); + + test("legacy fingerprint adoption ignores other models and fingerprinted rows", async () => { + const store = await createTestStore(); + const db = store.db; + const fakeLlm = createFakeEmbedLlm(); + store.llm = fakeLlm as any; + + try { + store.ensureVecTable(3); + await insertTestDocument(db, "docs", { hash: "hash-a", body: "Other model body.", displayPath: "other-model.md" }); + addLegacyChunk(store, "hash-a", 0, { model: "hf:test/other-model.gguf", vector: otherVector }); + await insertTestDocument(db, "docs", { hash: "hash-b", body: "Fingerprinted body.", displayPath: "fingerprinted.md" }); + addLegacyChunk(store, "hash-b", 0, { fingerprint: "abc123", vector: otherVector }); + await insertTestDocument(db, "docs", { hash: "hash-c", body: "Legacy body.", displayPath: "legacy.md" }); + addLegacyChunk(store, "hash-c", 0, { vector: matchingVector }); + + const result = await maybeAdoptLegacyEmbeddingFingerprint(store, legacyModel); + + expect(result).toMatchObject({ checked: true, adopted: 1 }); + expect(result.reason).toMatch(/^sample hash-c_0 matched/); + expect(db.prepare(`SELECT hash, embed_fingerprint FROM content_vectors ORDER BY hash`).all()).toEqual([ + { hash: "hash-a", embed_fingerprint: "" }, + { hash: "hash-b", embed_fingerprint: "abc123" }, + { hash: "hash-c", embed_fingerprint: getEmbeddingFingerprint(legacyModel) }, + ]); + + // No legacy rows remain for this model: a second pass is a no-op. + const again = await maybeAdoptLegacyEmbeddingFingerprint(store, legacyModel); + + expect(again).toEqual({ checked: false, adopted: 0, reason: "no legacy empty-fingerprint embeddings" }); + expect(fakeLlm.embedCalls).toHaveLength(1); + } finally { + await cleanupTestDb(store); + } + }); + + test("legacy fingerprint adoption loads one sample body within a 128 MiB SQLite budget", () => { + // SQLite's hard heap limit is process-wide and cannot be raised again, so + // the fixture runs in a worker. The limit binds only under Bun with an + // SQLite that tracks memory (Bun's bundled SQLite on Linux, Homebrew SQLite + // on macOS), where the old per-chunk query fails with SQLITE_NOMEM. Apple's + // system libsqlite3 and better-sqlite3 are built with + // SQLITE_DEFAULT_MEMSTATUS=0, so there this only checks the adoption result. + const projectRoot = fileURLToPath(new URL("..", import.meta.url)); + const worker = join(projectRoot, "test", "_helpers", "legacy-adoption-sample-worker.ts"); + const args = isBun ? [worker] : [join(projectRoot, "node_modules", "tsx", "dist", "cli.mjs"), worker]; + const result = spawnSync(process.execPath, args, { encoding: "utf8", timeout: 20_000 }); + + expect(result.error).toBeUndefined(); + expect(result.stderr).toBe(""); + expect(result.status).toBe(0); + const adoption = JSON.parse(result.stdout) as { checked: boolean; adopted: number; reason: string }; + // 16 documents x 32 legacy chunks, all adopted after one sample matched. + expect(adoption).toMatchObject({ checked: true, adopted: 512 }); + expect(adoption.reason).toMatch(/^sample document-0_0 matched/); + }); + + test("legacy fingerprint adoption reads one copy of a body shared by 200 active paths within a 16 MiB SQLite budget", () => { + // One legacy chunk, so only the active-path count multiplies the body. The + // limit binds under Bun only, as in the 128 MiB test above. + const projectRoot = fileURLToPath(new URL("..", import.meta.url)); + const worker = join(projectRoot, "test", "_helpers", "legacy-adoption-shared-body-worker.ts"); + const args = isBun ? [worker] : [join(projectRoot, "node_modules", "tsx", "dist", "cli.mjs"), worker]; + const result = spawnSync(process.execPath, args, { encoding: "utf8", timeout: 60_000 }); + + expect(result.error).toBeUndefined(); + expect(result.stderr).toBe(""); + expect(result.status).toBe(0); + const out = JSON.parse(result.stdout); + expect(out).toMatchObject({ checked: true, adopted: 1 }); + // The budget only binds where SQLite tracks memory; on Bun under Linux it must. + if (isBun && process.platform === "linux") expect(out.limitBinds).toBe(true); + }); + test("generateEmbeddings flushes batches when maxDocsPerBatch is reached", async () => { const store = await createTestStore(); const db = store.db; @@ -4376,7 +6021,7 @@ describe("Embedding batching", () => { expect(result.errors).toBeGreaterThan(0); expect(result.failures?.[0]?.attempts).toBe(3); expect(db.prepare(`SELECT COUNT(*) as count FROM content_vectors`).get()).toEqual({ count: 0 }); - expect(db.prepare(`SELECT COUNT(*) as count FROM vectors_vec`).get()).toEqual({ count: 0 }); + expect(db.prepare(`SELECT COUNT(*) as count FROM ${VEC_TABLE}`).get()).toEqual({ count: 0 }); expect(store.getHashesNeedingEmbedding()).toBe(1); expect(store.getStatus().needsEmbedding).toBe(1); } finally { @@ -4452,7 +6097,7 @@ describe("Embedding batching", () => { const db = store.db; // Store is pinned to a 3-dim embed model. Docs AND the query must use it, - // so vectors_vec is created as float[3]. + // so the vector table is created as float[3]. const storeModel = "hf:store/embeddinggemma-300M.gguf"; const storeLlm = { ...createFakeTokenizer(), @@ -4506,7 +6151,7 @@ describe("Embedding batching", () => { const model = "hf:Qwen/Qwen3-Embedding-0.6B-GGUF/Qwen3-Embedding-0.6B-Q8_0.gguf"; const searchVecSpy = vi.fn(async () => [] as SearchResult[]) as any; - store.db.exec(`CREATE TABLE vectors_vec (hash_seq TEXT PRIMARY KEY, embedding BLOB)`); + store.db.exec(`CREATE TABLE ${VEC_TABLE} (collection_id INTEGER, embedding BLOB)`); store.llm = { embedModelName: model } as any; store.searchVec = searchVecSpy as any; store.expandQuery = vi.fn(async () => []) as any; @@ -4532,7 +6177,7 @@ describe("Embedding batching", () => { }))); const searchVecSpy = vi.fn(async () => [] as SearchResult[]) as any; - store.db.exec(`CREATE TABLE vectors_vec (hash_seq TEXT PRIMARY KEY, embedding BLOB)`); + store.db.exec(`CREATE TABLE ${VEC_TABLE} (collection_id INTEGER, embedding BLOB)`); store.llm = { embedModelName: model, embedBatch: embedBatchSpy, @@ -4564,7 +6209,7 @@ describe("Embedding batching", () => { const searchVecSpy = vi.fn(async () => [] as SearchResult[]) as any; const searchFtsSpy = vi.fn(() => [] as SearchResult[]) as any; - store.db.exec(`CREATE TABLE vectors_vec (hash_seq TEXT PRIMARY KEY, embedding BLOB)`); + store.db.exec(`CREATE TABLE ${VEC_TABLE} (collection_id INTEGER, embedding BLOB)`); store.llm = { embedModelName: model, embedBatch: embedBatchSpy, @@ -4603,7 +6248,7 @@ describe("Embedding batching", () => { }))); const searchVecSpy = vi.fn(async () => [] as SearchResult[]) as any; - store.db.exec(`CREATE TABLE vectors_vec (hash_seq TEXT PRIMARY KEY, embedding BLOB)`); + store.db.exec(`CREATE TABLE ${VEC_TABLE} (collection_id INTEGER, embedding BLOB)`); store.llm = { embedModelName: model, embedBatch: embedBatchSpy, @@ -4627,6 +6272,388 @@ describe("Embedding batching", () => { } }); + test("generateEmbeddings copies vectors to a collection that gained an embedded hash instead of embedding again", async () => { + const store = await createTestStore(); + const db = store.db; + const fakeLlm = createFakeEmbedLlm(); + + setDefaultLlamaCpp(createFakeTokenizer() as any); + store.llm = fakeLlm as any; + + try { + const first = await createTestCollection({ name: "first", pwd: "/test/first" }); + const second = await createTestCollection({ name: "second", pwd: "/test/second" }); + await insertTestDocument(db, first, { name: "shared", hash: "sharedhash", body: "# Shared\n\nShared body", displayPath: "shared.md" }); + const embedded = await generateEmbeddings(store); + expect(embedded.chunksEmbedded).toBe(1); + expect(embedded.chunksCopied).toBe(0); + + await insertTestDocument(db, second, { name: "copy", hash: "sharedhash", body: "# Shared\n\nShared body", displayPath: "copy.md" }); + expect(await store.searchVec("ignored", "test-model", 5, second, undefined, [1, 2, 3])).toEqual([]); + expect(store.getHashesNeedingEmbedding()).toBe(0); + + const result = await generateEmbeddings(store); + + expect(fakeLlm.embedBatchCalls).toHaveLength(1); + expect(result.chunksCopied).toBe(1); + expect(result.chunksEmbedded).toBe(0); + const found = await store.searchVec("ignored", "test-model", 5, second, undefined, [1, 2, 3]); + expect(found.map((r) => r.displayPath)).toEqual([`${second}/copy.md`]); + expect((await store.searchVec("ignored", "test-model", 5, first, undefined, [1, 2, 3])).map((r) => r.displayPath)).toEqual([`${first}/shared.md`]); + } finally { + setDefaultLlamaCpp(null); + await cleanupTestDb(store); + } + }); + + test("generateEmbeddings embeds a hash again when no partition holds its vector", async () => { + const store = await createTestStore(); + const db = store.db; + const fakeLlm = createFakeEmbedLlm(); + + setDefaultLlamaCpp(createFakeTokenizer() as any); + store.llm = fakeLlm as any; + + try { + const docs = await createTestCollection({ name: "docs", pwd: "/test/docs" }); + await insertTestDocument(db, docs, { name: "lost", hash: "losthash", body: "# Lost\n\nLost body", displayPath: "lost.md" }); + await generateEmbeddings(store); + const rows = db.prepare(`SELECT id FROM ${VEC_ROWS_TABLE}`).all() as { id: number }[]; + expect(rows).toHaveLength(1); + deletePartitionRows(db, rows.map((row) => row.id)); + expect(store.getHashesNeedingEmbedding()).toBe(0); + + const result = await generateEmbeddings(store); + + expect(fakeLlm.embedBatchCalls).toHaveLength(2); + expect(result.chunksEmbedded).toBe(1); + expect(result.chunksCopied).toBe(0); + expect((await store.searchVec("ignored", "test-model", 5, docs, undefined, [1, 2, 3])).map((r) => r.hash)).toEqual(["losthash"]); + } finally { + setDefaultLlamaCpp(null); + await cleanupTestDb(store); + } + }); + + /** Active single-chunk documents for `hashes` in `collection`, in one transaction. */ + function seedDocs(db: Database, collection: string, hashes: readonly string[]): void { + const now = new Date().toISOString(); + db.transaction(() => { + for (const hash of hashes) { + insertContent(db, hash, `# ${hash}\n\nBody of ${hash}`, now); + insertDocument(db, collection, `${hash}.md`, hash, hash, now, now); + } + })(); + } + + /** `chunks` embedded chunks per hash under the default model, in every collection active for the hash, in one transaction. */ + function seedEmbeddings(store: Store, hashes: readonly string[], chunks = 1): void { + const now = new Date().toISOString(); + store.db.transaction(() => { + for (const hash of hashes) { + for (let seq = 0; seq < chunks; seq++) { + store.insertEmbedding(hash, seq, seq * 100, new Float32Array([1, 2, 3]), DEFAULT_EMBED_MODEL, now, chunks); + } + } + })(); + } + + function dropPartitionRows(db: Database, hash: string, seq: number): void { + const rows = db.prepare(`SELECT id FROM ${VEC_ROWS_TABLE} WHERE hash = ? AND seq = ?`).all(hash, seq) as { id: number }[]; + deletePartitionRows(db, rows.map((row) => row.id)); + } + + function rowsOfHash(db: Database, table: string, hash: string): number { + return (db.prepare(`SELECT COUNT(*) AS c FROM ${table} WHERE hash = ?`).get(hash) as { c: number }).c; + } + + /** Same connection, counting the IMMEDIATE transactions that run through it. */ + function commitCountingDb(db: Database, commits: { count: number }): Database { + return { + prepare: (sql: string) => db.prepare(sql), + exec: (sql: string) => db.exec(sql), + transaction: (fn) => { + const wrapped = db.transaction(fn); + const immediate = ((...args: Parameters) => { + const result = wrapped.immediate(...args); + commits.count++; + return result; + }) as typeof fn; + return Object.assign(((...args: Parameters) => wrapped(...args)) as typeof fn, { immediate }); + }, + loadExtension: (path: string) => db.loadExtension(path), + close: () => db.close(), + }; + } + + const paddedHash = (i: number) => `copy${String(i).padStart(4, "0")}`; + + test("generateEmbeddings copies more than one batch of vectors to a collection that gained embedded hashes", async () => { + const store = await createTestStore(); + const db = store.db; + const fakeLlm = createFakeEmbedLlm(); + + setDefaultLlamaCpp(createFakeTokenizer() as any); + store.llm = fakeLlm as any; + + try { + const first = await createTestCollection({ name: "first", pwd: "/test/first" }); + const second = await createTestCollection({ name: "second", pwd: "/test/second" }); + const hashes = Array.from({ length: VECTOR_COPY_BATCH_ROWS * 2.5 }, (_, i) => paddedHash(i)); + store.ensureVecTable(3); + seedDocs(db, first, hashes); + seedEmbeddings(store, hashes); + seedDocs(db, second, hashes); + expect(store.getHashesNeedingEmbedding()).toBe(0); + + const result = await generateEmbeddings(store); + + expect(fakeLlm.embedBatchCalls).toHaveLength(0); + expect(result.chunksCopied).toBe(hashes.length); + expect(result.chunksEmbedded).toBe(0); + expect(vectorRowCount(store, second)).toBe(hashes.length); + const found = await store.searchVec("ignored", "test-model", 5, second, undefined, [1, 2, 3]); + expect(found).toHaveLength(5); + expect(found.every((r) => r.collectionName === second)).toBe(true); + expect(copyVectorsToNewCollections(db)).toEqual({ copied: 0, queued: 0 }); + } finally { + setDefaultLlamaCpp(null); + await cleanupTestDb(store); + } + }); + + test("copyVectorsToNewCollections queues a hash whose chunks straddle a batch boundary and keeps none of its copies", async () => { + const store = await createTestStore(); + const db = store.db; + try { + const first = await createTestCollection({ name: "first", pwd: "/test/first" }); + const second = await createTestCollection({ name: "second", pwd: "/test/second" }); + store.ensureVecTable(3); + + // missingPartitionRows orders by hash, seq, collection, so single-chunk + // hashes place each two-chunk hash's rows on both sides of a batch + // boundary. Suffix "a" sorts a hash right after the single it follows. + const batch = VECTOR_COPY_BATCH_ROWS; + const group1 = Array.from({ length: batch - 1 }, (_, i) => paddedHash(i)); + // Rows: (seq 0, second) closes batch 1; (seq 1, first) and (seq 1, second) open batch 2. + const copiedThenQueued = `${paddedHash(batch - 2)}a`; + const group2 = Array.from({ length: batch - 4 }, (_, i) => paddedHash(batch - 1 + i)); + // Rows: (seq 0, first) and (seq 0, second) close batch 2; (seq 1, second) opens batch 3. + const queuedThenSkipped = `${paddedHash(2 * batch - 6)}a`; + const singles = [...group1, ...group2]; + const straddlers = [copiedThenQueued, queuedThenSkipped]; + + seedDocs(db, first, [...singles, ...straddlers]); + seedEmbeddings(store, singles); + seedEmbeddings(store, straddlers, 2); + dropPartitionRows(db, copiedThenQueued, 1); + dropPartitionRows(db, queuedThenSkipped, 0); + seedDocs(db, second, [...singles, ...straddlers]); + const missingRows = singles.length + 3 + 3; + const commits = { count: 0 }; + + const result = copyVectorsToNewCollections(commitCountingDb(db, commits)); + + expect(commits.count).toBe(Math.ceil(missingRows / batch)); + expect(result.queued).toBe(2); + // Every single, plus copiedThenQueued's seq 0 written by batch 1 before batch 2 found its seq 1 missing. + expect(result.copied).toBe(singles.length + 1); + for (const hash of straddlers) { + expect(rowsOfHash(db, VEC_ROWS_TABLE, hash)).toBe(0); + expect(rowsOfHash(db, "content_vectors", hash)).toBe(0); + } + expect(vectorRowCount(store, first)).toBe(singles.length); + expect(vectorRowCount(store, second)).toBe(singles.length); + expect(store.getHashesNeedingEmbedding()).toBe(2); + expect(copyVectorsToNewCollections(db)).toEqual({ copied: 0, queued: 0 }); + } finally { + await cleanupTestDb(store); + } + }); + + test("generateEmbeddings writes a batch in one transaction and rolls it back when a chunk write fails", async () => { + const store = await createTestStore(); + const real = store.db; + const fakeLlm = createFakeEmbedLlm(); + const doomed = "doomedhash"; + const observed = { inserted: [] as string[], insideTransaction: [] as boolean[], firstHashRowsAfterRollback: -1 }; + const rowsOf = (hash: string) => (real.prepare(`SELECT COUNT(*) AS c FROM content_vectors WHERE hash = ?`).get(hash) as { c: number }).c; + const guard = unknown>(run: T): T => ((...args: unknown[]) => { + try { + return run(...args); + } catch (error) { + // The failing write unwinds an inner savepoint first; sample once the + // outermost transaction has rolled back. + if (!(real as any).inTransaction && observed.firstHashRowsAfterRollback < 0 && observed.inserted.length > 0) { + observed.firstHashRowsAfterRollback = rowsOf(observed.inserted[0]!); + } + throw error; + } + }) as T; + const failingDb = { + prepare: (sql: string) => { + const stmt = real.prepare(sql); + if (!sql.includes("INSERT OR REPLACE INTO content_vectors")) return stmt; + return { + run: (...params: any[]) => { + if (params[0] === doomed) throw new Error("injected write failure"); + observed.inserted.push(params[0]); + observed.insideTransaction.push((real as any).inTransaction); + return stmt.run(...params); + }, + get: (...params: any[]) => stmt.get(...params), + all: (...params: any[]) => stmt.all(...params), + iterate: (...params: any[]) => stmt.iterate(...params), + }; + }, + transaction: (fn: any) => { + const tx = real.transaction(fn); + const wrapped = guard(tx) as any; + wrapped.immediate = guard(tx.immediate); + return wrapped; + }, + exec: (sql: string) => real.exec(sql), + loadExtension: (path: string) => real.loadExtension(path), + close: () => real.close(), + }; + + setDefaultLlamaCpp(createFakeTokenizer() as any); + store.llm = fakeLlm as any; + store.db = failingDb as any; + + try { + await insertTestDocument(real, "docs", { name: "one", hash: "onehash", body: "# One\n\nAlpha" }); + await insertTestDocument(real, "docs", { name: "two", hash: doomed, body: "# Two\n\nBeta" }); + await insertTestDocument(real, "docs", { name: "three", hash: "threehash", body: "# Three\n\nGamma" }); + + const result = await generateEmbeddings(store); + + expect(fakeLlm.embedBatchCalls).toHaveLength(1); + expect(observed.inserted[0]).toBe("onehash"); + expect(observed.insideTransaction.every(Boolean)).toBe(true); + expect(observed.firstHashRowsAfterRollback).toBe(0); + expect(result.errors).toBe(1); + expect(result.failures?.[0]?.hash).toBe(doomed); + expect(rowsOf(doomed)).toBe(0); + expect(rowsOf("onehash")).toBe(1); + expect(rowsOf("threehash")).toBe(1); + expect(store.getHashesNeedingEmbedding()).toBe(1); + } finally { + store.db = real; + setDefaultLlamaCpp(null); + await cleanupTestDb(store); + } + }); + + test("cleanupOrphanedVectors after re-indexing a changed file removes the stale rows that starve a scoped search", async () => { + const store = await createTestStore(); + const db = store.db; + const near = [1, 0, 0]; + const far = [0, 1, 0]; + const embeddingFor = (text: string) => ({ embedding: text.includes("stale") ? near : far, model: "fake-embed" }); + const fakeLlm = { + ...createFakeTokenizer(), + async embed(text: string, _options?: { model?: string }) { return embeddingFor(text); }, + async embedBatch(texts: string[], _options?: { model?: string }) { return texts.map(embeddingFor); }, + }; + + setDefaultLlamaCpp(createFakeTokenizer() as any); + store.llm = fakeLlm as any; + const dir = await mkdtemp(join(tmpdir(), "qmd-stale-vectors-")); + + try { + await writeFile(join(dir, "live.md"), "# Live\n\nlive body\n"); + const versions = ["stale 0", "stale 1", "stale 2", "stale 3", "final body"]; + for (const body of versions) { + await writeFile(join(dir, "churn.md"), `# Churn\n\n${body}\n`); + await reindexCollection(store, dir, "**/*.md", "docs"); + await generateEmbeddings(store); + } + expect(db.prepare(`SELECT COUNT(*) AS c FROM documents WHERE active = 1`).get()).toEqual({ c: 2 }); + expect(db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_TABLE}`).get()).toEqual({ c: 6 }); + + // Four stale rows sit nearer the query than both live documents and take + // every one of the first scoped KNN's three slots at limit 1; the search + // only answers after widening its KNN past them. + expect(await store.searchVec("ignored", "test-model", 1, "docs", undefined, near)).toHaveLength(1); + + expect(cleanupOrphanedVectors(db)).toBe(4); + + expect(db.prepare(`SELECT COUNT(*) AS c FROM ${VEC_TABLE}`).get()).toEqual({ c: 2 }); + const after = await store.searchVec("ignored", "test-model", 1, "docs", undefined, near); + expect(after).toHaveLength(1); + expect(await store.searchVec("ignored", "test-model", 5, "docs", undefined, near)).toHaveLength(2); + } finally { + setDefaultLlamaCpp(null); + await rm(dir, { recursive: true, force: true }); + await cleanupTestDb(store); + } + }); + + test("maybeAdoptLegacyEmbeddingFingerprint finds the sample's stored vector through the partition mapping", async () => { + const store = await createTestStore(); + const db = store.db; + const model = "hf:test/embed-model.gguf"; + const stored = [0.1, 0.2, 0.3]; + const fakeLlm = { + ...createFakeTokenizer(), + embedModelName: model, + async embed(_text: string, _options?: { model?: string }) { return { embedding: stored, model }; }, + async embedBatch(texts: string[], _options?: { model?: string }) { return texts.map(() => ({ embedding: stored, model })); }, + }; + + setDefaultLlamaCpp(createFakeTokenizer() as any); + store.llm = fakeLlm as any; + + try { + const docs = await createTestCollection({ name: "docs", pwd: "/test/docs" }); + await insertTestDocument(db, docs, { name: "legacy", hash: "legacyhash", body: "# Legacy\n\nLegacy body", displayPath: "legacy.md" }); + store.ensureVecTable(3); + store.insertEmbedding("legacyhash", 0, 0, new Float32Array(stored), model, new Date().toISOString(), 1, ""); + expect(store.getHashesNeedingEmbedding(model)).toBe(1); + + const result = await maybeAdoptLegacyEmbeddingFingerprint(store, model); + + expect(result).toMatchObject({ checked: true, adopted: 1 }); + expect(result.reason).toContain("legacyhash_0"); + expect(store.getHashesNeedingEmbedding(model)).toBe(0); + } finally { + setDefaultLlamaCpp(null); + await cleanupTestDb(store); + } + }); + + test("maybeAdoptLegacyEmbeddingFingerprint keeps a legacy fingerprint whose sample no longer matches", async () => { + const store = await createTestStore(); + const db = store.db; + const model = "hf:test/embed-model.gguf"; + const fakeLlm = { + ...createFakeTokenizer(), + embedModelName: model, + async embed(_text: string, _options?: { model?: string }) { return { embedding: [0.3, 0.2, 0.1], model }; }, + async embedBatch(texts: string[], _options?: { model?: string }) { return texts.map(() => ({ embedding: [0.3, 0.2, 0.1], model })); }, + }; + + setDefaultLlamaCpp(createFakeTokenizer() as any); + store.llm = fakeLlm as any; + + try { + const docs = await createTestCollection({ name: "docs", pwd: "/test/docs" }); + await insertTestDocument(db, docs, { name: "legacy", hash: "legacyhash", body: "# Legacy\n\nLegacy body", displayPath: "legacy.md" }); + store.ensureVecTable(3); + store.insertEmbedding("legacyhash", 0, 0, new Float32Array([0.1, 0.2, 0.3]), model, new Date().toISOString(), 1, ""); + + const result = await maybeAdoptLegacyEmbeddingFingerprint(store, model); + + expect(result).toMatchObject({ checked: true, adopted: 0 }); + expect(result.reason).toContain("nearest legacyhash_0"); + expect(store.getHashesNeedingEmbedding(model)).toBe(1); + } finally { + setDefaultLlamaCpp(null); + await cleanupTestDb(store); + } + }); + test("generateEmbeddings rejects invalid batch limits", async () => { const store = await createTestStore(); diff --git a/test/structured-search.test.ts b/test/structured-search.test.ts index 70da7fd1e..8f2ab08d6 100644 --- a/test/structured-search.test.ts +++ b/test/structured-search.test.ts @@ -9,12 +9,16 @@ * Run with: bun test structured-search.test.ts */ -import { describe, test, expect, beforeAll, afterAll } from "vitest"; +import { describe, test, expect, vi, beforeAll, afterAll, beforeEach, afterEach } from "vitest"; import { mkdtemp, rm } from "node:fs/promises"; import { tmpdir } from "node:os"; import { join } from "node:path"; import { createStore, + hashContent, + insertContent, + insertDocument, + searchFTS, structuredSearch, validateSemanticQuery, validateLexQuery, @@ -345,6 +349,40 @@ describe("structuredSearch", () => { { type: "lex", query: "\"unfinished phrase", line: 2 } ])).rejects.toThrow(/unmatched double quote/); }); + + test("ranks over the union of named collections, not by the order they are named", async () => { + const now = new Date().toISOString(); + const add = async (collection: string, path: string, body: string) => { + const hash = await hashContent(body); + insertContent(store.db, hash, body, now); + insertDocument(store.db, collection, path, path, hash, now, now); + }; + const filler = "lorem ipsum dolor sit amet consectetur adipiscing elit ".repeat(40); + await add("alpha", "alpha/weak.md", `${filler} zanzibarquartz ${filler}`); + await add("beta", "beta/strong.md", "zanzibarquartz zanzibarquartz zanzibarquartz notes"); + + const searches: ExpandedQuery[] = [{ type: "lex", query: "zanzibarquartz" }]; + const unscoped = await structuredSearch(store, searches, { skipRerank: true }); + expect(unscoped[0]?.displayPath).toContain("strong.md"); + + for (const collections of [["alpha", "beta"], ["beta", "alpha"]]) { + const results = await structuredSearch(store, searches, { collections, skipRerank: true }); + expect(results.map(r => r.displayPath)).toEqual(unscoped.map(r => r.displayPath)); + } + }); + + test("an exact tie between collections goes the same way whichever is named first", async () => { + const now = new Date().toISOString(); + const body = "quillmarrow notes"; + const hash = await hashContent(body); + insertContent(store.db, hash, body, now); + insertDocument(store.db, "alpha", "same.md", "same.md", hash, now, now); + insertDocument(store.db, "beta", "same.md", "same.md", hash, now, now); + + for (const collections of [["alpha", "beta"], ["beta", "alpha"]]) { + expect(searchFTS(store.db, "quillmarrow", 1, collections).map(r => r.filepath)).toEqual(["qmd://alpha/same.md"]); + } + }); }); // ============================================================================= @@ -592,3 +630,83 @@ describe("buildFTS5Query (lex parser)", () => { expect(buildFTS5Query("performance -sports")).toBe('"performance"* NOT "sports"*'); }); }); + +// ============================================================================= +// structuredSearch collection scope +// ============================================================================= + +describe("structuredSearch collection scope", () => { + let testDir: string; + let store: Store; + const origConfigDir = process.env.QMD_CONFIG_DIR; + + beforeEach(async () => { + testDir = await mkdtemp(join(tmpdir(), "qmd-structured-scope-")); + process.env.QMD_CONFIG_DIR = await mkdtemp(join(testDir, "config-")); + store = createStore(join(testDir, "index.sqlite")); + }); + + afterEach(async () => { + store.close(); + if (origConfigDir === undefined) delete process.env.QMD_CONFIG_DIR; + else process.env.QMD_CONFIG_DIR = origConfigDir; + await rm(testDir, { recursive: true, force: true }); + }); + + function addDocument(collection: string, path: string, body: string): void { + const now = new Date().toISOString(); + const hash = `hash-${collection}-${path}`; + store.insertContent(hash, body, now); + store.insertDocument(collection, path, path, hash, now, now); + } + + test("a vec leg over several collections runs one vector search over the whole list", async () => { + store.ensureVecTable(3); + store.llm = { + embedModelName: "scope-embed-model", + embedBatch: async (texts: string[]) => texts.map(() => ({ embedding: [1, 0, 0], model: "scope-embed-model" })), + } as any; + const searchVec = vi.fn(async () => []); + store.searchVec = searchVec as any; + + await structuredSearch(store, [{ type: "vec", query: "x" }], { collections: ["alpha", "beta"], skipRerank: true }); + + expect(searchVec).toHaveBeenCalledTimes(1); + expect(searchVec.mock.calls[0]![3]).toEqual(["alpha", "beta"]); + }); + + test("a vec leg ranks the same whichever order the collections are named in", async () => { + const embedModel = "scope-embed-model"; + store.ensureVecTable(3); + store.llm = { + embedModelName: embedModel, + embedBatch: async (texts: string[]) => texts.map(() => ({ embedding: [1, 0, 0], model: embedModel })), + } as any; + const now = new Date().toISOString(); + const vectors: [string, string, number[]][] = [ + ["alpha", "far.md", [0, 1, 0]], + ["beta", "near.md", [1, 0, 0]], + ["alpha", "mid.md", [0.8, 0.6, 0]], + ]; + for (const [collection, path, vector] of vectors) { + addDocument(collection, path, `# ${path}\n\nbody`); + store.insertEmbedding(`hash-${collection}-${path}`, 0, 0, new Float32Array(vector), embedModel, now); + } + + const searches: ExpandedQuery[] = [{ type: "vec", query: "x" }]; + const unscoped = await structuredSearch(store, searches, { skipRerank: true }); + expect(unscoped.map(r => r.displayPath)).toEqual(["beta/near.md", "alpha/mid.md", "alpha/far.md"]); + for (const collections of [["alpha", "beta"], ["beta", "alpha"]]) { + const scoped = await structuredSearch(store, searches, { collections, skipRerank: true }); + expect(scoped.map(r => r.displayPath)).toEqual(unscoped.map(r => r.displayPath)); + } + }); + + test("an empty collection list searches every collection", async () => { + addDocument("alpha", "a.md", "# A\n\nquokkascope in alpha"); + addDocument("beta", "b.md", "# B\n\nquokkascope in beta"); + + const results = await structuredSearch(store, [{ type: "lex", query: "quokkascope" }], { collections: [], skipRerank: true }); + expect(results.map(r => r.displayPath).sort()).toEqual(["alpha/a.md", "beta/b.md"]); + }); +}); diff --git a/test/vec-layout.test.ts b/test/vec-layout.test.ts new file mode 100644 index 000000000..110d9c132 --- /dev/null +++ b/test/vec-layout.test.ts @@ -0,0 +1,157 @@ +/** + * The vector layout resolver owns the names of the sqlite-vec tables and the + * integer collection ids that partition the vec0 table. + */ +import { describe, test, expect, afterEach } from "vitest"; +import { mkdtemp, rm } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { createStore, type Store } from "../src/store.js"; +import { + LEGACY_VEC_TABLE, + VEC_COLLECTION_IDS_TABLE, + VEC_ROWS_TABLE, + VEC_TABLE, + allocateCollectionId, + createPartitionedVecTable, + hasVectorIndex, + renameCollectionId, + resolveCollectionId, + resolveCollectionIds, + rowidList, + vecLayout, + vecTableReadable, +} from "../src/vec-layout.js"; + +let store: Store | null = null; +let dir: string | null = null; + +async function openStore(): Promise { + dir = await mkdtemp(join(tmpdir(), "qmd-vec-layout-")); + store = createStore(join(dir, "index.sqlite")); + return store; +} + +afterEach(async () => { + store?.close(); + store = null; + if (dir) await rm(dir, { recursive: true, force: true }); + dir = null; +}); + +function tableExists(s: Store, name: string): boolean { + return Boolean(s.db.prepare(`SELECT name FROM sqlite_master WHERE type = 'table' AND name = ?`).get(name)); +} + +describe("vecLayout", () => { + test("resolves none on a store without a vector table", async () => { + const s = await openStore(); + expect(vecLayout(s.db)).toEqual({ kind: "none" }); + expect(hasVectorIndex(s.db)).toBe(false); + }); + + test("creates the collection id and row mapping tables at open", async () => { + const s = await openStore(); + expect(tableExists(s, VEC_COLLECTION_IDS_TABLE)).toBe(true); + expect(tableExists(s, VEC_ROWS_TABLE)).toBe(true); + }); + + test("resolves the partitioned table with its dimensions and shadow tables", async () => { + const s = await openStore(); + createPartitionedVecTable(s.db, 3); + + const layout = vecLayout(s.db); + expect(layout.kind).toBe("partitioned"); + if (layout.kind !== "partitioned") return; + expect(layout.table).toBe(VEC_TABLE); + expect(layout.dimensions).toBe(3); + expect(layout.shadow).toEqual({ + chunks: `${VEC_TABLE}_chunks`, + rowids: `${VEC_TABLE}_rowids`, + vectorChunks: `${VEC_TABLE}_vector_chunks00`, + info: `${VEC_TABLE}_info`, + }); + expect(tableExists(s, layout.shadow.chunks)).toBe(true); + expect(tableExists(s, layout.shadow.rowids)).toBe(true); + expect(hasVectorIndex(s.db)).toBe(true); + expect(vecTableReadable(s.db, layout)).toBe(true); + }); + + test("resolves a legacy hash_seq table, and keeps resolving it while both tables exist", async () => { + const s = await openStore(); + s.db.exec(`CREATE VIRTUAL TABLE ${LEGACY_VEC_TABLE} USING vec0(hash_seq TEXT PRIMARY KEY, embedding float[4] distance_metric=cosine)`); + + const legacy = vecLayout(s.db); + expect(legacy).toMatchObject({ kind: "legacy", table: LEGACY_VEC_TABLE, dimensions: 4, keyedByHashSeq: true }); + expect(hasVectorIndex(s.db)).toBe(false); + + createPartitionedVecTable(s.db, 4); + expect(vecLayout(s.db).kind).toBe("legacy"); + }); + + test("marks a legacy table keyed by hash alone", async () => { + const s = await openStore(); + s.db.exec(`CREATE VIRTUAL TABLE ${LEGACY_VEC_TABLE} USING vec0(hash TEXT PRIMARY KEY, embedding float[2] distance_metric=cosine)`); + expect(vecLayout(s.db)).toMatchObject({ kind: "legacy", dimensions: 2, keyedByHashSeq: false }); + }); + + test("reports a schema entry whose module is not loaded as unreadable", async () => { + const s = await openStore(); + s.db.exec(`CREATE TABLE ${VEC_TABLE} (collection_id INTEGER, embedding BLOB)`); + const layout = vecLayout(s.db); + expect(layout.kind).toBe("partitioned"); + if (layout.kind !== "partitioned") return; + expect(layout.dimensions).toBeNull(); + expect(vecTableReadable(s.db, layout)).toBe(true); + s.db.exec(`DROP TABLE ${VEC_TABLE}`); + expect(vecTableReadable(s.db, layout)).toBe(false); + }); +}); + +describe("collection ids", () => { + test("allocates a stable integer id per name and resolves it back", async () => { + const s = await openStore(); + const docs = allocateCollectionId(s.db, "docs"); + const notes = allocateCollectionId(s.db, "notes"); + expect(docs).not.toBe(notes); + expect(allocateCollectionId(s.db, "docs")).toBe(docs); + expect(resolveCollectionId(s.db, "docs")).toBe(docs); + expect(resolveCollectionId(s.db, "unknown")).toBeUndefined(); + expect(resolveCollectionIds(s.db, ["notes", "unknown", "docs"])).toEqual(new Map([["docs", docs], ["notes", notes]])); + }); + + test("rename keeps the id so partition rows stay reachable", async () => { + const s = await openStore(); + const id = allocateCollectionId(s.db, "old"); + renameCollectionId(s.db, "old", "new"); + expect(resolveCollectionId(s.db, "old")).toBeUndefined(); + expect(resolveCollectionId(s.db, "new")).toBe(id); + }); + + test("rename of a name without an id is a no-op", async () => { + const s = await openStore(); + expect(() => renameCollectionId(s.db, "missing", "other")).not.toThrow(); + expect(resolveCollectionId(s.db, "other")).toBeUndefined(); + }); + + test("the row mapping refuses a null collection id", async () => { + const s = await openStore(); + expect(() => s.db.prepare(`INSERT INTO ${VEC_ROWS_TABLE} (hash, seq, collection_id) VALUES ('h', 0, NULL)`).run()).toThrow(/NOT NULL/); + }); + + test("the row mapping is unique per (hash, seq, collection)", async () => { + const s = await openStore(); + const id = allocateCollectionId(s.db, "docs"); + s.db.prepare(`INSERT INTO ${VEC_ROWS_TABLE} (hash, seq, collection_id) VALUES ('h', 0, ?)`).run(id); + expect(() => s.db.prepare(`INSERT INTO ${VEC_ROWS_TABLE} (hash, seq, collection_id) VALUES ('h', 0, ?)`).run(id)).toThrow(/UNIQUE/); + }); +}); + +describe("rowid lists", () => { + test("json_each expands a rowid list on the loaded SQLite build", async () => { + const s = await openStore(); + const ids = Array.from({ length: 60_000 }, (_, i) => i + 1); + const row = s.db.prepare(`SELECT COUNT(*) AS n, MAX(value) AS top FROM json_each(?)`).get(rowidList(ids)) as { n: number; top: number }; + expect(row).toEqual({ n: 60_000, top: 60_000 }); + }); +}); diff --git a/vitest.config.ts b/vitest.config.ts index 463e72385..3c95c2419 100644 --- a/vitest.config.ts +++ b/vitest.config.ts @@ -5,5 +5,11 @@ export default defineConfig({ testTimeout: 30000, fileParallelism: false, include: ["test/**/*.test.ts"], + typecheck: { + enabled: true, + include: ["test/**/*.test-d.ts"], + // Judge only the .test-d.ts assertions; tsc covers src/ separately. + ignoreSourceErrors: true, + }, }, });