Skip to content

i8 x i8 MaxSim Kernels - #1429

Open
Suryansh Gupta (suri-kumkaran) wants to merge 10 commits into
mainfrom
users/suryangupta/i8_x_i8_maxsim_kernel
Open

Suryansh Gupta (suri-kumkaran) wants to merge 10 commits into
mainfrom
users/suryangupta/i8_x_i8_maxsim_kernel

Conversation

@suri-kumkaran

@suri-kumkaran Suryansh Gupta (suri-kumkaran) commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Integer MaxSim matrix kernels for i8 x i8, alongside the existing f32 and f16 paths.
Scores stay in i32 end to end instead of routing through f32, so results are exact.

Status

Scalar, V3, V4 and Neon each have their own kernel. Known gaps:

  • Scalar is still slower than the reference. It uses the emulated i16 pair dot, which widens each lane to i32 before multiplying. PACK = 1 versions were slower still.
  • V4 gains less with 8 or 16 queries, where half or more of its 32 row query panel is padding. A 16 row panel is left for a follow up.

What is here

The packed View and Panel in matrix_kernels::blocks::packed gain a PACK parameter (default 1, as in BlockTransposed), so contraction elements can be interleaved to match the width of the dot product instruction.

ISA MR NR PACK Query packed as Instruction
Scalar 8 2 2 i8 emulated
V3 16 6 2 i16 vpmaddwd
V4 32 6 4 u8, as x + 128 vpdpbusd
Neon 8 6 4 i8 sdot

PACK differs by target because the instruction does. vpmaddwd is a 2 way dot per 32 bit lane, vpdpbusd and sdot are 4 way.

A PrepareB seam in the driver prepares each L1 sized block of B (the documents) before the micro-kernels use it:

  • V3 widens B to i16. The query is widened when the kernel is built, so the inner loop does no widening.
  • V4 sums each column of B. vpdpbusd is unsigned by signed, so the query is stored as x + 128 in u8, and each accumulator starts at -128 times its column sum to cancel the shift.
  • Scalar and Neon pass B through untouched.

MaxSimElement is implemented for i8 and gains a Score type (i32 for i8, f32 for f32 and f16) and a NO_MATCH score for empty documents. MaxSimKernel<T> and Erase<T> now take T: MaxSimElement instead of T: Copy, and compute_max_sim writes to &mut [T::Score]. f32 and f16 callers still pass &mut [f32]. Generic code needs the new bound.

build_max_sim now returns BuildMaxSimError instead of NotSupported. It wraps NotSupported and adds DimTooLarge, which the i8 build returns for queries with more than 131_071 dimensions. Products are bounded by 128 * 128, so that is the largest dimension where i32 cannot overflow.

Also adds:

  • An i8 factory test that checks Auto and Reference against MaxSim, including the full i8 range, and tests on either side of the dimension limit.
  • max_sim_kernel_i8 in fallback.rs for the i8 reference ISA.
  • int8 as an element type for multi-vector-op, with example and perf test jobs.
  • SIMDReinterpret<i8x16> for u32x4 on aarch64 in diskann-wide, for the Neon broadcast, with a test.

Performance

V3 and V4 on a Xeon Gold 6426Y, best of 30 interleaved runs over two job orders, reference held at 1.000 as control:

  • V3: 3.68x geomean over reference, 3.7x to 4.6x with 32 or more queries, 1.87x and 3.39x with 8 and 16.
  • V4: 14.88x geomean over reference, 16.0x to 20.7x with 32 or more queries, 4.15x and 9.48x with 8 and 16.

Neon on an Apple M5 Pro, best of 15 interleaved runs: 1.75x geomean over reference, 1.66x to 1.88x across shapes. Measured on an earlier revision and not re-run since.

@codecov-commenter

Codecov Comments Bot (codecov-commenter) commented Sep 21, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.50898% with 45 lines in your changes missing coverage. Please review.
✅ Project coverage is 91.95%. Comparing base (bceaf45) to head (c91a7fa).

Files with missing lines Patch % Lines
...-quantization/src/multi_vector/distance/factory.rs 79.80% 42 Missing ⚠️
...c/matrix_kernels/maxsim/packed_i8_x_unpacked_i8.rs 99.53% 3 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1429      +/-   ##
==========================================
+ Coverage   91.92%   91.95%   +0.02%     
==========================================
  Files         580      581       +1     
  Lines      114938   115864     +926     
==========================================
+ Hits       105657   106542     +885     
- Misses       9281     9322      +41     
Flag Coverage Δ
miri 91.95% <95.50%> (+0.02%) ⬆️
unittests 91.89% <95.33%> (+0.02%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
diskann-benchmark/src/multi_vector/driver.rs 88.52% <100.00%> (ø)
diskann-benchmark/src/multi_vector/kernels.rs 91.89% <100.00%> (+0.11%) ⬆️
...n-quantization/src/matrix_kernels/blocks/packed.rs 99.70% <100.00%> (+0.05%) ⬆️
...quantization/src/matrix_kernels/blocks/unpacked.rs 100.00% <100.00%> (ø)
...matrix_kernels/maxsim/packed_f32_x_unpacked_f16.rs 100.00% <100.00%> (ø)
...matrix_kernels/maxsim/packed_f32_x_unpacked_f32.rs 99.54% <100.00%> (+<0.01%) ⬆️
...ann-quantization/src/matrix_kernels/maxsim/test.rs 100.00% <100.00%> (ø)
...skann-quantization/src/matrix_kernels/test_util.rs 78.26% <100.00%> (+3.26%) ⬆️
diskann-quantization/src/matrix_kernels/util.rs 100.00% <100.00%> (ø)
...quantization/src/multi_vector/distance/fallback.rs 97.97% <100.00%> (+0.21%) ⬆️
... and 3 more

... and 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Comment thread diskann-quantization/src/matrix_kernels/test_util.rs Outdated
Comment thread diskann-quantization/src/matrix_kernels/blocks/packed.rs
Comment thread diskann-quantization/src/matrix_kernels/maxsim/packed_i8_x_unpacked_i8.rs Outdated
@suri-kumkaran
Suryansh Gupta (suri-kumkaran) marked this pull request as ready for review October 1, 2026 18:10
@suri-kumkaran
Suryansh Gupta (suri-kumkaran) requested review from a team and a balanced review from Copilot October 1, 2026 18:10

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Dimensions above 131,071 can overflow the new i32 accumulation paths and produce panics or incorrect scores.

Review effort: Balanced
Findings: 1 High severity

Open (1)
What changed in this PR

Adds exact i8 × i8 MaxSim support with architecture-specific kernels and i32 scores.

Changes:

  • Adds Scalar, V3, V4, and Neon integer kernels.
  • Extends packed layouts and MaxSim APIs for typed scores.
  • Adds correctness tests and benchmark configurations.
File Description
diskann-wide/​src/​arch/​aarch64/​u32x4_.rs Adds Neon reinterpretation support.
diskann-quantization/​src/​multi_vector/​distance/​kernel.rs Generalizes kernel score types.
diskann-quantization/​src/​multi_vector/​distance/​fallback.rs Adds reference i8 scoring.
diskann-quantization/​src/​multi_vector/​distance/​factory.rs Wires i8 kernels and dispatch.
diskann-quantization/​src/​matrix_kernels/​util.rs Adds integer conversion and load/store helpers.
diskann-quantization/​src/​matrix_kernels/​test_util.rs Adds random i8 test data.
diskann-quantization/​src/​matrix_kernels/​maxsim/​test.rs Adds integer test generation.
diskann-quantization/​src/​matrix_kernels/​maxsim/​packed_i8_x_unpacked_i8.rs Implements integer MaxSim kernels.
diskann-quantization/​src/​matrix_kernels/​maxsim/​packed_f32_x_unpacked_f32.rs Updates renamed test helper usage.
diskann-quantization/​src/​matrix_kernels/​maxsim/​packed_f32_x_unpacked_f16.rs Updates renamed test helper usage.
diskann-quantization/​src/​matrix_kernels/​maxsim/​mod.rs Exposes the integer kernel module.
diskann-quantization/​src/​matrix_kernels/​blocks/​unpacked.rs Exposes remainder start offsets.
diskann-quantization/​src/​matrix_kernels/​blocks/​packed.rs Adds packed contraction grouping.
diskann-benchmark/​src/​multi_vector/​kernels.rs Registers i8 benchmarks.
diskann-benchmark/​src/​multi_vector/​driver.rs Supports element-specific score buffers.
diskann-benchmark/​perf_test_inputs/​multi-vector.json Adds i8 performance jobs.
diskann-benchmark/​example/​multi-vector.json Adds example i8 jobs.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread diskann-quantization/src/multi_vector/distance/factory.rs

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mainly looked at the stuff in matrix_kernels. I think this is mostly okay to merge. It is starting to show some cracks in our abstractions (I made a few notes in the inline comments), and this gets worse with #1394.

The general problems that are emerging are

  1. The fixed tiling iteration order. I'm still not sure if it's the best nor how k-blocking will affect things.
  2. The long-range implicit contracts between panel preparation and the micro-kernels. I would feel much better if there was a single source of truth for how these things interacted. However, I don't have an actual concrete suggestion yet.

Comment thread diskann-quantization/src/matrix_kernels/util.rs Outdated
Comment thread diskann-quantization/src/matrix_kernels/maxsim/packed_i8_x_unpacked_i8.rs Outdated
Comment thread diskann-quantization/src/multi_vector/distance/factory.rs
Comment thread diskann-wide/src/arch/aarch64/u32x4_.rs
Comment thread diskann-quantization/src/matrix_kernels/blocks/packed.rs
@suri-kumkaran

Copy link
Copy Markdown
Contributor Author

I mainly looked at the stuff in matrix_kernels. I think this is mostly okay to merge. It is starting to show some cracks in our abstractions (I made a few notes in the inline comments), and this gets worse with #1394.

The general problems that are emerging are

  1. The fixed tiling iteration order. I'm still not sure if it's the best nor how k-blocking will affect things.
  2. The long-range implicit contracts between panel preparation and the micro-kernels. I would feel much better if there was a single source of truth for how these things interacted. However, I don't have an actual concrete suggestion yet.

Thanks for the review, Mark. Agreed on both points. I felt the second one most with V4: the query is packed as  x + 128  in the factory, while the  -128 * column sum  that cancels it is computed in  PrepareB , and only comments and tests tie the two together. For the tiling order I reused the f32 one as is, so I don't have a better idea there yet. Happy to help work out a cleaner design along with #1394.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants