Repository navigation
feat(libtorch): run the bridge on the host when the SDK is CPU-only - #36
NicolasRouquette wants to merge 2 commits into
Conversation
The LibTorch bridge pinned `c10::kCUDA`, so a build without a CUDA device had no backend at all: `unavailable.c` fails every buffer operation, and the device buffer suite could only run on a GPU host. The bridge now selects its device once, at its first call, and runs on the host through ATen's CPU kernels when the SDK has no CUDA support or when `TORCHLEAN_LIBTORCH_DEVICE=cpu` asks for it. Nothing is reimplemented: every operation already goes through `torchlean::options()`. Selection: `TORCHLEAN_LIBTORCH_DEVICE=cpu` selects the host and `=cuda` insists on a CUDA device. Unset, a visible CUDA device is selected when the SDK has CUDA support; otherwise the host is selected only when the SDK has no CUDA support at all, so a CUDA-enabled SDK without a visible device stays `RuntimeStatus.nativeUnavailable` rather than silently running on the host. CMake no longer requires the `torch_cuda` target. A CPU-only SDK compiles the backend with `TORCHLEAN_LIBTORCH_CUDA=0`, without CUDA headers or a toolkit; the CPU-only pip wheel is a complete SDK for it. `sdk.txt` records the choice. Lean: `Runtime.Autograd.LibTorch.DeviceKind` and `deviceKind` report the selection. `RuntimeStatus` keeps its three states, with `.nativeAvailable` meaning that a device is selected, CUDA or host, so `requireNativeRuntime` and every existing match are unchanged. On the host the native allocator counters and driver memory read zero, `synchronize` and `emptyCache` return at once, `setMemoryFraction` and `setDevice` are rejected, and the suite skips its native memory probes. The suite prints the device and checks that the selection agrees with the visible device count and with the environment variable. Verified with pip torch 2.11.0+cu128 on an RTX A4500 (device cuda) and with the same binary under `TORCHLEAN_LIBTORCH_DEVICE=cpu` (device host), and with pip torch 2.11.0+cpu and no toolkit (`TORCHLEAN_LIBTORCH_CUDA=0`, device host): all curated tests passed in each run, with only the three memory probes skipped on the host. The default build and its CPU suite are unchanged.
…nd CUDA Programs made of the IEEE basic operations, gathers and lookups give bit-identical results on the two devices; transcendental functions come from different libraries on the host and on CUDA and differ at the ulp level, and reductions differ in order.
|
Hey Nicolas, thanks! CPU-only LibTorch support is worth having. One issue before merging: NN/Tests/Suite.lean uses requireNativeRuntime for TORCHLEAN_REQUIRE_CUDA=1, but nativeAvailable now includes host execution. That lets a CUDA-required run pass on CPU. Could you make that strict check require an actual CUDA device, while retaining a separate native-runtime check for CPU sessions? Please also check the combination with #35, whose device-information tests currently assume CUDA. Our pending NVRTC/custom-kernel work will need conditional CUDA compilation and a clear unsupported error on the host; that is an integration requirement for our local changes, not a claim that this PR fails to build independently. Holding off until the strict CUDA check is fixed. |
What
The LibTorch bridge pins
c10::kCUDA, so the backend has one device and the default build hasnone:
unavailable.cfails every buffer operation. Before the move to LibTorch, the default buildlinked the
_stub.chost implementations of the same ABI, which is what kept the device buffersuite runnable on hosted CI ("CPU hosted CI does not validate GPU execution" in
csrc/libtorch/README.md) and what let a CPU-only deployment run code written againstRuntime.Autograd.LibTorch.Buffer.This PR lets the same bridge run on the host through ATen's CPU kernels. Nothing is reimplemented:
the device is chosen once, and every operation already goes through
torchlean::options().TORCHLEAN_LIBTORCH_DEVICE=cpuselects the host,=cudainsists on a CUDAdevice. Unset, a visible CUDA device is selected when the SDK has CUDA support; otherwise the
host is selected only when the SDK has no CUDA support at all. A CUDA-enabled SDK without a
visible device stays
RuntimeStatus.nativeUnavailable, as today, rather than silently runningon the host.
torch_cudatarget is no longer an error; it compiles the backendwith
TORCHLEAN_LIBTORCH_CUDA=0, with no CUDA headers and no toolkit. The CPU-only pip wheel(
pip install --index-url https://download.pytorch.org/whl/cpu torch==2.11.0, 746 MB, notoolkit) is a complete SDK for it.
Runtime.Autograd.LibTorch.DeviceKind(cuda | host) anddeviceKind.RuntimeStatuskeeps its three states;.nativeAvailablenow means "a device is selected",CUDA or host, so
requireNativeRuntimeand every existingmatchare unchanged. The suiteprints the device and checks that the selection agrees with the visible count and with the
environment variable.
synchronizeandemptyCachereturn at once;setMemoryFractionandsetDeviceare rejected; the suite'snative memory probes (accounting, saved attention buffers, OOM recovery) are skipped, and
nothing else is.
-K cuda=truekeeps its meaning, "link the LibTorch backend"; the SDK decides the device. Theextern names (
torchlean_cuda_*) and theConfig.device := .cudasession name also keep theirmeaning of "the bridge"; if you would rather have a
libtorch-cpudevice name for sessions, thatis a small follow-up and I am happy to do it, but I kept the user-facing vocabulary unchanged here.
Verification
cudaTORCHLEAN_LIBTORCH_DEVICE=cpuhostTORCHLEAN_LIBTORCH_CUDA=0,lddshowslibtorch_cpuand no CUDA library)host-Kcuda=false), CPU suiteThe host run covers the same sections as the CUDA run, including the CUDA kernel coverage suite,
the float32 contract checks, determinism, attention, convolution and pooling, and the stress
tests; the only difference between the two logs is the skipped memory probes. Results are
bit-identical across the two devices only for programs made of the IEEE basic operations, gathers
and lookups (a downstream table-lookup retrieval was checked bit for bit, host against GPU);
ATen's transcendental functions come from different libraries on the host and on CUDA and differ
at the ulp level (a downstream Newton inverse differed by at most 1e-6 in float32), and reductions
differ in order. The suite's tolerances were met on both.
Repo lint:
python3 scripts/checks/repo_lint.py --fail-on-warnOK. Lean lint (lake lint): OK.Rationale
The curated device suite can run on a hosted runner with the CPU wheel and no toolkit, so the real
buffer ABI is exercised on every push instead of only on a GPU box (the "CI job" below). It also
gives downstream code written against the buffer ABI a CPU deployment target again, which the old
_stub.cfiles provided andunavailable.cdoes not.This PR enables executing the same binary on GPU or on CPU.
Files
csrc/libtorch/torchlean.cpp,torchlean_libtorch.h,CMakeLists.txt,unavailable.c,README.md;scripts/libtorch_build.py;NN/Runtime/Autograd/Engine/LibTorch/{Controls,Buffer, Trusted}.lean;NN/Tests/Suite.lean,NN/Tests/Runtime/Cuda/Stress.lean;lakefile.lean,README.md.