Repository navigation
feat(libtorch): report device identity per index - #35
Open
NicolasRouquette wants to merge 1 commit into
Open
NicolasRouquette wants to merge 1 commit into
NicolasRouquette wants to merge 1 commit into
Conversation
`LibTorch.deviceInfo index` reads one device's name, compute capability, SM count, clocks, bus width, total memory, and the driver and runtime versions through the SDK's per-device cached properties; `currentDeviceInfo` reads the device `getDevice` selects. Querying by explicit index keeps the answer independent of the selected device, so switching devices cannot report a stale record. Without LibTorch, the queries fail through `IO` like the other controls. `NN/Tests/Runtime/Cuda/DeviceInfo.lean` checks every visible device under its own index, the current-device agreement, index rejection, and, in a fresh process selected by `TORCHLEAN_LIBTORCH_DEVICE_PROBE=switch`, that switching devices changes the reading. Verified in the default build only: the `torchlean.cpp` half has not been compiled against a LibTorch SDK yet.
NicolasRouquette
added a commit
to NicolasRouquette/TorchLean
that referenced
this pull request
Oct 3, 2026
Brings in the upstream PR branch (lean-dojo#35) at 65fd9d9, one commit on upstream main b062b9a: Runtime.Autograd.LibTorch.deviceInfo and currentDeviceInfo read at::cuda::getDeviceProperties(index) and cudaDeviceGetAttribute per index, with DeviceInfo.format for benchmark headers. Successor of the closed lean-dojo#30. The compiled-architecture fields are gone with the custom kernels.
Member
|
Hey Nicolas, thanks! Per-device identity is useful. Could you check this together with #36? The CUDA property queries need CPU-only SDK handling, and the tests should not assume nativeAvailable means a CUDA device exists. Also, the statement in DeviceInfo.lean that TorchLean compiles no device code of its own will need updating when our pending custom-kernel/NVRTC work lands. Please keep the SDK-packaged architecture list separate from runtime-compiled kernels in that explanation. Keeping this open for those integration checks. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Runtime.Autograd.LibTorch.deviceInfo (index : UInt32) : IO DeviceInfoandcurrentDeviceInfo, the LibTorch follow-up to #30: name, compute capability (major * 10 + minor, so86reads assm_86), SM count, core and memory clocks, memory bus width, totalmemory, and the driver and runtime versions, plus
DeviceInfo.peakBandwidthGBsand a one-lineDeviceInfo.formatfor a benchmark header.Buffer.allocatorStatsanswers how much devicememory there is; this answers which device.
The review finding from #30
#30 cached one
cudaDevicePropunder a process-widepthread_once, so after switching devicesit kept reporting device 0 and could combine fields from two cards. Here every query takes an
explicit index and reads
at::cuda::getDeviceProperties(index), the SDK's own per-device cache,so the answer never depends on which device was selected earlier;
currentDeviceInfoisdeviceInfo (← getDevice). The test switches through every visible device in a fresh processand checks that
currentDeviceInforeports each index in turn.Clocks and bus width come from
cudaDeviceGetAttribute, since CUDA 13 removedclockRateandmemoryClockRatefromcudaDevicePropand the attribute enums exist in 12.x and 13.x. An indexat or above
device_countis aTORCH_CHECKfailure that surfaces as anIOerror, like everyother runtime control; the build without LibTorch fails the same way through
unavailable.crather than returning a record of zeroes.
What is not carried over
#30 also reported the architectures the binary held native code for (
__CUDA_ARCH_LIST__) andwarned when the device was being served through PTX JIT. TorchLean no longer compiles device
code, and ATen's C++ API does not expose the SDK's packaged architecture list (only the Python
package reads it back), so those fields are gone. The module docstring says so.
Tests
NN/Tests/Runtime/Cuda/DeviceInfo.lean, called fromNN/Tests/Suite.leanunder every runtimestatus:
count, and memory;
currentDeviceInfoagrees withgetDevice; the index at the visible countis rejected; and a child process (
TORCHLEAN_LIBTORCH_DEVICE_PROBE=switch, the pattern of thememory probes) walks
setDevicethrough every device before any buffer exists, becausesetDevicerefuses while wrappers are live;deviceInfo 0fails.Checks run
RTX A4500 (sm_86, driver 580.126), pip torch 2.11.0+cu128 as the SDK, CUDA 12.8 toolkit for
SDK discovery, Lean v4.34.0:
GPU suite output of the new section:
Not run: the sanitizer harness and the elementwise C++ harness. This host has one GPU, so the
switching probe exercised one index; the multi-device case is covered by construction (per-index
reads) rather than by an observation.