Skip to content

feat(libtorch): run the bridge on the host when the SDK is CPU-only - #36

Open
NicolasRouquette wants to merge 2 commits into
lean-dojo:mainfrom
NicolasRouquette:libtorch-cpu-device
Open

NicolasRouquette wants to merge 2 commits into
lean-dojo:mainfrom
NicolasRouquette:libtorch-cpu-device

Conversation

@NicolasRouquette

Copy link
Copy Markdown
Contributor

What

The LibTorch bridge pins c10::kCUDA, so the backend has one device and the default build has
none: unavailable.c fails every buffer operation. Before the move to LibTorch, the default build
linked the _stub.c host implementations of the same ABI, which is what kept the device buffer
suite runnable on hosted CI ("CPU hosted CI does not validate GPU execution" in
csrc/libtorch/README.md) and what let a CPU-only deployment run code written against
Runtime.Autograd.LibTorch.Buffer.

This PR lets the same bridge run on the host through ATen's CPU kernels. Nothing is reimplemented:
the device is chosen once, and every operation already goes through torchlean::options().

  • Selection. TORCHLEAN_LIBTORCH_DEVICE=cpu selects the host, =cuda insists on a CUDA
    device. Unset, a visible CUDA device is selected when the SDK has CUDA support; otherwise the
    host is selected only when the SDK has no CUDA support at all. A CUDA-enabled SDK without a
    visible device stays RuntimeStatus.nativeUnavailable, as today, rather than silently running
    on the host.
  • CMake. An SDK without the torch_cuda target is no longer an error; it compiles the backend
    with TORCHLEAN_LIBTORCH_CUDA=0, with no CUDA headers and no toolkit. The CPU-only pip wheel
    (pip install --index-url https://download.pytorch.org/whl/cpu torch==2.11.0, 746 MB, no
    toolkit) is a complete SDK for it.
  • Lean. Runtime.Autograd.LibTorch.DeviceKind (cuda | host) and deviceKind.
    RuntimeStatus keeps its three states; .nativeAvailable now means "a device is selected",
    CUDA or host, so requireNativeRuntime and every existing match are unchanged. The suite
    prints the device and checks that the selection agrees with the visible count and with the
    environment variable.
  • On the host: the native allocator counters and driver memory read zero; synchronize and
    emptyCache return at once; setMemoryFraction and setDevice are rejected; the suite's
    native memory probes (accounting, saved attention buffers, OOM recovery) are skipped, and
    nothing else is.

-K cuda=true keeps its meaning, "link the LibTorch backend"; the SDK decides the device. The
extern names (torchlean_cuda_*) and the Config.device := .cuda session name also keep their
meaning of "the bridge"; if you would rather have a libtorch-cpu device name for sessions, that
is a small follow-up and I am happy to do it, but I kept the user-facing vocabulary unchanged here.

Verification

run SDK device result
curated suite, A4500 pip torch 2.11.0+cu128 cuda all curated tests passed
same binary, TORCHLEAN_LIBTORCH_DEVICE=cpu pip torch 2.11.0+cu128 host all curated tests passed; the three memory probes skipped
curated suite, no toolkit (TORCHLEAN_LIBTORCH_CUDA=0, ldd shows libtorch_cpu and no CUDA library) pip torch 2.11.0+cpu host all curated tests passed; the three memory probes skipped
default build (-Kcuda=false), CPU suite none — 11107 jobs; all curated tests passed, CUDA kernels skipped as before

The host run covers the same sections as the CUDA run, including the CUDA kernel coverage suite,
the float32 contract checks, determinism, attention, convolution and pooling, and the stress
tests; the only difference between the two logs is the skipped memory probes. Results are
bit-identical across the two devices only for programs made of the IEEE basic operations, gathers
and lookups (a downstream table-lookup retrieval was checked bit for bit, host against GPU);
ATen's transcendental functions come from different libraries on the host and on CUDA and differ
at the ulp level (a downstream Newton inverse differed by at most 1e-6 in float32), and reductions
differ in order. The suite's tolerances were met on both.

Repo lint: python3 scripts/checks/repo_lint.py --fail-on-warn OK. Lean lint (lake lint): OK.

Rationale

The curated device suite can run on a hosted runner with the CPU wheel and no toolkit, so the real
buffer ABI is exercised on every push instead of only on a GPU box (the "CI job" below). It also
gives downstream code written against the buffer ABI a CPU deployment target again, which the old
_stub.c files provided and unavailable.c does not.

This PR enables executing the same binary on GPU or on CPU.

Files

csrc/libtorch/torchlean.cpp, torchlean_libtorch.h, CMakeLists.txt, unavailable.c,
README.md; scripts/libtorch_build.py; NN/Runtime/Autograd/Engine/LibTorch/{Controls,Buffer, Trusted}.lean; NN/Tests/Suite.lean, NN/Tests/Runtime/Cuda/Stress.lean; lakefile.lean,
README.md.

The LibTorch bridge pinned `c10::kCUDA`, so a build without a CUDA device had no backend at all:
`unavailable.c` fails every buffer operation, and the device buffer suite could only run on a
GPU host. The bridge now selects its device once, at its first call, and runs on the host through
ATen's CPU kernels when the SDK has no CUDA support or when `TORCHLEAN_LIBTORCH_DEVICE=cpu` asks
for it. Nothing is reimplemented: every operation already goes through `torchlean::options()`.

Selection: `TORCHLEAN_LIBTORCH_DEVICE=cpu` selects the host and `=cuda` insists on a CUDA device.
Unset, a visible CUDA device is selected when the SDK has CUDA support; otherwise the host is
selected only when the SDK has no CUDA support at all, so a CUDA-enabled SDK without a visible
device stays `RuntimeStatus.nativeUnavailable` rather than silently running on the host.

CMake no longer requires the `torch_cuda` target. A CPU-only SDK compiles the backend with
`TORCHLEAN_LIBTORCH_CUDA=0`, without CUDA headers or a toolkit; the CPU-only pip wheel is a
complete SDK for it. `sdk.txt` records the choice.

Lean: `Runtime.Autograd.LibTorch.DeviceKind` and `deviceKind` report the selection.
`RuntimeStatus` keeps its three states, with `.nativeAvailable` meaning that a device is
selected, CUDA or host, so `requireNativeRuntime` and every existing match are unchanged. On the
host the native allocator counters and driver memory read zero, `synchronize` and `emptyCache`
return at once, `setMemoryFraction` and `setDevice` are rejected, and the suite skips its native
memory probes. The suite prints the device and checks that the selection agrees with the visible
device count and with the environment variable.

Verified with pip torch 2.11.0+cu128 on an RTX A4500 (device cuda) and with the same binary under
`TORCHLEAN_LIBTORCH_DEVICE=cpu` (device host), and with pip torch 2.11.0+cpu and no toolkit
(`TORCHLEAN_LIBTORCH_CUDA=0`, device host): all curated tests passed in each run, with only the
three memory probes skipped on the host. The default build and its CPU suite are unchanged.
…nd CUDA

Programs made of the IEEE basic operations, gathers and lookups give bit-identical results on
the two devices; transcendental functions come from different libraries on the host and on
CUDA and differ at the ulp level, and reductions differ in order.
@Robertboy18

Copy link
Copy Markdown
Member

Hey Nicolas, thanks! CPU-only LibTorch support is worth having. One issue before merging: NN/Tests/Suite.lean uses requireNativeRuntime for TORCHLEAN_REQUIRE_CUDA=1, but nativeAvailable now includes host execution. That lets a CUDA-required run pass on CPU. Could you make that strict check require an actual CUDA device, while retaining a separate native-runtime check for CPU sessions? Please also check the combination with #35, whose device-information tests currently assume CUDA. Our pending NVRTC/custom-kernel work will need conditional CUDA compilation and a clear unsupported error on the host; that is an integration requirement for our local changes, not a claim that this PR fails to build independently. Holding off until the strict CUDA check is fixed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants