Repository navigation
[diskann-quantization] PQ Infrastructure for diskann-inmem - #1458
Mark Hildebrand (hildebrandmw) wants to merge 22 commits into
Conversation
diskann-inmem
Benchmarking PerformanceUsing the following input Benchmark JSON{
"search_directories": [
"..."
],
"output_directory": null,
"jobs": [
{
"type": "exhaustive-product-quantization",
"content": {
"compression_threads": 8,
"data": "wikipedia/wikipedia_base_100K.bin",
"data_type": "float32",
"distance": "inner_product",
"num_pq_centers": 256,
"num_pq_chunks": 192,
"search": {
"groundtruth": "wikipedia/wikipedia-100K",
"num_threads": 8,
"queries": "wikipedia/wikipedia_query.bin",
"recalls": {
"recall_k": [
10,
20,
30,
40
],
"recall_n": [
10,
20,
30,
40
]
}
},
"seed": 7831252621480178695,
"table_style": "fixed-chunk"
}
},
{
"type": "exhaustive-product-quantization",
"content": {
"compression_threads": 8,
"data": "wikipedia/wikipedia_base_100K.bin",
"data_type": "float32",
"distance": "inner_product",
"num_pq_centers": 256,
"num_pq_chunks": 192,
"search": {
"groundtruth": "wikipedia/wikipedia-100K",
"num_threads": 8,
"queries": "wikipedia/wikipedia_query.bin",
"recalls": {
"recall_k": [
10,
20,
30,
40
],
"recall_n": [
10,
20,
30,
40
]
}
},
"seed": 7831252621480178695,
"table_style": "transposed"
}
},
{
"type": "exhaustive-product-quantization",
"content": {
"compression_threads": 8,
"data": "openai-v3/base-100k.bin",
"data_type": "float32",
"distance": "squared_l2",
"num_pq_centers": 256,
"num_pq_chunks": 384,
"search": {
"groundtruth": "openai-v3/gt-100k.bin",
"num_threads": 8,
"queries": "openai-v3/query.bin",
"recalls": {
"recall_k": [
10,
20,
30,
40
],
"recall_n": [
10,
20,
30,
40
]
}
},
"seed": 7831252621480178695,
"table_style": "fixed-chunk"
}
},
{
"type": "exhaustive-product-quantization",
"content": {
"compression_threads": 8,
"data": "openai-v3/base-100k.bin",
"data_type": "float32",
"distance": "squared_l2",
"num_pq_centers": 256,
"num_pq_chunks": 384,
"search": {
"groundtruth": "openai-v3/gt-100k.bin",
"num_threads": 8,
"queries": "openai-v3/query.bin",
"recalls": {
"recall_k": [
10,
20,
30,
40
],
"recall_n": [
10,
20,
30,
40
]
}
},
"seed": 7831252621480178695,
"table_style": "transposed"
}
},
{
"type": "exhaustive-product-quantization",
"content": {
"compression_threads": 8,
"data": "openai-v3/base-100k.bin",
"data_type": "float32",
"distance": "cosine",
"num_pq_centers": 256,
"num_pq_chunks": 384,
"search": {
"groundtruth": "openai-v3/gt-100k.bin",
"num_threads": 8,
"queries": "openai-v3/query.bin",
"recalls": {
"recall_k": [
10,
20,
30,
40
],
"recall_n": [
10,
20,
30,
40
]
}
},
"seed": 7831252621480178695,
"table_style": "fixed-chunk"
}
},
{
"type": "exhaustive-product-quantization",
"content": {
"compression_threads": 8,
"data": "openai-v3/base-100k.bin",
"data_type": "float32",
"distance": "cosine",
"num_pq_centers": 256,
"num_pq_chunks": 384,
"search": {
"groundtruth": "openai-v3/gt-100k.bin",
"num_threads": 8,
"queries": "openai-v3/query.bin",
"recalls": {
"recall_k": [
10,
20,
30,
40
],
"recall_n": [
10,
20,
30,
40
]
}
},
"seed": 7831252621480178695,
"table_style": "padded"
}
},
{
"type": "exhaustive-product-quantization",
"content": {
"compression_threads": 8,
"data": "openai-v3/base-100k.bin",
"data_type": "float32",
"distance": "cosine",
"num_pq_centers": 256,
"num_pq_chunks": 384,
"search": {
"groundtruth": "openai-v3/gt-100k.bin",
"num_threads": 8,
"queries": "openai-v3/query.bin",
"recalls": {
"recall_k": [
10,
20,
30,
40
],
"recall_n": [
10,
20,
30,
40
]
}
},
"seed": 7831252621480178695,
"table_style": "transposed"
}
}
]
}
For Wikipedia, the transposed table is signficantly faster at preprocessing and compression (reflected in compression time and preprocess time). In addition, even though both the OpenAI with L2 is a similar story. OpenAI with Cosine gives some insight into the Recall in all cases is identical. |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The newly required table-style field makes existing product-exhaustive benchmark configurations fail deserialization.
Review effort: Balanced
Findings: 1
What changed in this PR
Adds PQ lookup, cosine-distance, and padded-table infrastructure for faster quantized distance computation and benchmarking.
Changes:
- Adds precomputed lookup tables with cosine support.
- Introduces SIMD-aware padded tables for full-quant and quant-quant distances.
- Extends benchmarks and shared distance tests across table implementations.
| File | Description |
|---|---|
diskann-quantization/src/views.rs |
Computes maximum chunk dimensions. |
diskann-quantization/src/test_util.rs |
Adds exact-value test checks. |
diskann-quantization/src/product/tables/transposed/table.rs |
Supports typed lookup outputs and cosine preprocessing. |
diskann-quantization/src/product/tables/transposed/pivots.rs |
Generates cosine dot-and-norm values. |
diskann-quantization/src/product/tables/transposed/mod.rs |
Adds end-to-end distance tests. |
diskann-quantization/src/product/tables/test.rs |
Adds shared PQ distance test infrastructure. |
diskann-quantization/src/product/tables/padded.rs |
Implements padded SIMD distance tables. |
diskann-quantization/src/product/tables/mod.rs |
Exposes new table APIs. |
diskann-quantization/src/product/tables/lookup.rs |
Implements precomputed distance lookup. |
diskann-quantization/src/product/tables/basic.rs |
Adds borrowed table views. |
diskann-quantization/src/product/mod.rs |
Makes table modules public. |
diskann-quantization/src/distances.rs |
Adds the cosine operation marker. |
diskann-disk/src/storage/quant/generator.rs |
Corrects a comment. |
diskann-disk/src/search/pq/quantizer_preprocess.rs |
Adapts to generic lookup outputs. |
diskann-benchmark/src/inputs/exhaustive.rs |
Adds PQ table-style configuration. |
diskann-benchmark/src/exhaustive/product.rs |
Benchmarks all PQ table implementations. |
diskann-benchmark/example/product-exhaustive.json |
Selects the transposed benchmark style. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| pub(crate) seed: u64, | ||
| pub(crate) num_pq_chunks: NonZeroUsize, | ||
| pub(crate) num_pq_centers: NonZeroUsize, | ||
| pub(crate) table_style: PQTableStyle, |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #1458 +/- ##
==========================================
- Coverage 91.90% 91.89% -0.01%
==========================================
Files 582 585 +3
Lines 114947 116271 +1324
==========================================
+ Hits 105640 106852 +1212
- Misses 9307 9419 +112
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|

Add infrastructure for:
Distance Lookups (
diskann_quantization/src/product/tables/lookup.rs)PQ distance computations for a fixed query can be made faster by pre-computing distances between a query and all PQ pivots into a lookup table and indexing into that table by PQ codes. This PR introduces
lookup_single, which unrolls the lookup table computation.In addition, it supports cosine computations via the
DotAndNormtype. This works by computing both the inner product between a query chunk and pivot chunk as the "dot" field, and storing the squared pivot-chunk norm as well. After lookup, the finalDotAndNormcontains both the inner product and square norm between the query vector and the entire reconstruct PQ vector, and can be combined with the query norm for the final cosine similarity.Since distance lookup tables are non-trivial allocations, this API (and the preparation stage in the
TransposedTable) need a little bit of work to use for full distance computations (see the changes indiskann-benchmark). This is to provide full control of how allocations happen to the caller. The current lookup tables indiskann-providersrequire pooled objects, which is too opinionated fordiskann-quantization.Cosine Distance Preparation
Infrastructure is added to
diskann-quantization/src/product/tables/transposed/{pivots.rs, table.rs}to do fast(er) preparation ofDotAndNormbased lookup tables. This is functionality that is lacking in the current PQ infrastructure, requiring either use of L2 or a slow fallback.Supporting fast(er) quant-quant distances
To support graph pruning, we need a way of doing relatively fast quant-quant distances. The methods in
FixedChunkPQTableare quite branchy, and a dedicated quant-quant distance lookup table requires on the order of 32KB to 131KB of storage per-chunk which is prohibitively large.This PR introduces
diskann-quantization/src/product/tables/padded.rs. The idea here is to ensure that the pivots for each chunk are contiguous in memory, and then padding all pivots to the same length (a multiple of an underlying SIMD width). This organization and padding makes quant-quant distance computations much more regular. Full-quant distances are still a little messy, unfortunately.Testing
diskann-providershas decent infrastructure for testing PQ based distances. This PR ports this infrastructure todiskann-quantization/src/product/tables/test.rsand uses it to test end-to-end distances via both thePaddedTableand the look-up table basedTransposedTable.Suggested Reviewing Order
diskann-quantization/src/product/tables/test.rs: Shared distance test infrastructure. Extends theChecktype indiskann-quantization/src/test_util.rs.diskann-quantization/src/product/tables/lookup.rs: Implementation of distance table lookup.diskann-quantization/src/product/tables/transposed/pivots.rs: Support populating pre-processed cosine changes intoDotAndNorm.diskann-quantization/src/product/tables/transposed/table.rs: Exposing pre-processing API forDotAndNormbased cosine distances.diskann-quantization/src/product/tables/transposed/mod.rs: End-to-end distance tests.diskann-quantization/src/product/tables/padded.rs: New padded table for quant-quant distances.diskann-benchmark/src/exhaustive/product.rs: Extending the exhaustive search benchmark to use the outgoingFixedChunkPQTableas well as thePaddedTableandTransposedTable.Benchmark Performance
See PR comment to keep commit description smaller.