Repository navigation
perf(ioloader): decode NPY payloads unboxed, and per leading-axis block concurrently - #29
NicolasRouquette wants to merge 1 commit into
Conversation
07c5540 to
6d824f7
Compare
|
Thanks, Nicolas. We still want to merge this performance work, but the data-loader cleanup in Could you rebase onto the current After the rebase, please rerun |
6d824f7 to
a9d1541
Compare
|
Hey Nicolas, can you rebase this onto the latest main? |
|
Thanks! The NPY decoding improvements still look useful. Could you rebase this onto current main and update the tests for the current loader, preserving its input validation? |
…ck concurrently The `.npy` element loop allocated on every step: an `Option` per byte, an `Except` per element, and a heap cell per element to box the `Float`. It also re-tested the dtype once per element. `decodeNpyPayload` now selects a range decoder once per file and fills a `FloatArray`, so `NpyData.values` is unboxed. `NN.API.Data.Sources` wraps that buffer into a tensor through `Tensor.from` without a copy or a boxed intermediate. Header validation, bounded handle reads, and the truncation checks are unchanged; the in-memory and file-backed readers share one decoder. `parseNpyLeadingAxisBlocks` and `readNpyLeadingAxisBlocks` return one `FloatArray` per index along the leading axis of a C-order array and decode those blocks concurrently. `Parallelism` and `resolveTaskCount` bound the task count by the request, the block count, and the payload size. `NN/Tests/Data/IO/Npy.lean` checks bit-for-bit round trips for both dtypes, that every task count returns the same blocks, determinism across repeated concurrent decodes, the task-count bounds, Fortran reordering, rejections, and that the file-backed readers agree with the in-memory ones.
a9d1541 to
181edd4
Compare
|
Rebased onto current What changed in the port:
Re-measured on this tree, same binary for all rows (details in the description):
Checks on The PR description is updated to match. |
Brings in the upstream PR branch (lean-dojo#29) at 181edd4, rebased onto upstream main b062b9a. combined is rebuilt from b062b9a, the first upstream tip on the LibTorch backend (csrc/cuda is gone). Its previous tip d162950 (Lean 4.33, custom CUDA kernels, arena, cache cap, TexTable) is kept as bkp/pre-libtorch/combined and is not merged: nothing on it applies to the LibTorch tree.
|
Hey Nicolas, thanks! This looks like a useful improvement. The loader, its tests, and the updated data API compile locally. My isolated interpreter run could not load the native parseNpy implementation, so I have not confirmed the runtime suite locally yet; that is a scratch-run setup limitation, not evidence of a bug in this PR. Could you share the native runtime test command and results? It would also be good to cover signed zero, infinities, and NaNs explicitly in the decode tests. Keeping this open until the runtime check is confirmed. |
Body
What
The
.npyelement loop allocates on every step: anOptionper byte insidereadLittleEndian,an
Exceptper element fromreadNpyElement, and a heap cell per element again to store theFloatin anArray Float. It also re-tests the dtype string once per element.This selects the decoder once per file inside
decodeNpyPayload, reads the bytes without anintermediate
Option, and decodes into aFloatArray.NpyData.valuesbecomes aFloatArray:the payload is bulk numeric data, so it costs one machine word per element instead of a word
plus a heap cell, and
Tensor.fromwraps that buffer into a tensor without copying it. Headervalidation, the bounded handle reads, prefix validation, and the truncation checks are unchanged;
the in-memory and file-backed readers share the one decoder.
It also adds
parseNpyLeadingAxisBlocksandreadNpyLeadingAxisBlocks, which return oneFloatArrayper index along the leading axis. Those blocks are physically contiguous in aC-order payload, so they are independent and can be decoded concurrently. Callers that want the
blocks apart anyway (one channel, one band, one batch element) get the concurrency for free.
Callers that want a single flat payload keep using
parseNpy/readNpy, which stay sequential:producing one contiguous result concurrently would mean filling several arrays and then copying
them all into one, and that copy walks every element a second time for several times the cost of
the concurrency it was meant to buy.
ParallelismandresolveTaskCountmake the task count explicit rather than tying it to thefile's shape. The count follows the machine and the block count follows the file, and those are
independent, so blocks are grouped across tasks rather than handed out one per task: a
(50000, 1, 28, 28)dataset does not spawn 50 000 tasks.Measured
One 764 MB
<f8file of shape(8, 3456, 3456)(95 551 488 elements), already in memory, Xeonw7-3465X (28 cores, 56 threads). One binary holds all three decoders; the first row is the
element loop
mainships atb062b9a3, copied verbatim. Five trials per row, each in its ownprocess so peak RSS is per path; the rate is file bytes over the median; every trial passes a
distinct tag so the pure decode cannot be shared across trials. Peak RSS includes the 764 MB
input held in memory.
mainparseNpy, unboxed sequential.sequential.tasks 8.auto.autoresolves to 8 tasks here: the leading dimension is 8, andresolveTaskCountneversplits a block, so the file's leading axis is the ceiling on this decode's parallelism whatever
the core count.
Tests
NN/Tests/Data/IO/Npy.leansynthesises.npyfiles in memory, so the expected values areknown exactly, and writes the same bytes to
.lake/build/tmpfor the file-backed readers.It checks:
<f8and<f4, scalars included.==onFloatwould be the wrong test: it identifies0.0with-0.0and makes everyNaNunequal to itself, so it would both miss sign-bit transcription errors and reject acorrectly decoded
NaN;parseNpyreturns, for 1-D, 2-D,and 3-D shapes and for zero-sized axes on either side;
above the block count and above the core count, plus
sequentialandauto. This is whatmakes the task count a free tuning knob rather than part of the contract;
resolveTaskCountrespects each of its three bounds;dtype, Fortran order and scalar shape for the block path;
readNpyandreadNpyLeadingAxisBlocksagree bit for bit with their in-memory counterparts,and the file-backed block reader rejects a Fortran-order or scalar header without reading a
payload (the fixture is a bare header, and the error names the storage order, not truncation).
The existing checks in
NN/Tests/API/Data.leanare unchanged except where they compare thepayload, which now reads
.values.data.Note on
Task.spawnThe block decode uses
Task.spawn, which takes the work as aUnit → αclosure. That isload-bearing and there is a comment saying so: handing a pure expression to an
IO-level spawninstead lets the compiler float it out of the closure onto the spawning thread, and the fan-out
then runs sequentially at exactly the single-threaded rate while still looking parallel.
Compatibility
NpyData.valueschanges type fromArray FloattoFloatArray. In-tree consumers:NN/API/Data/Sources.lean(nowTensor.fromon the buffer) and the comparisons inNN/Tests/API/Data.lean.NN/Examples/Data/Loaders/Npy.leanreads onlydtypeandshape.Checks run