Repository navigation
Conversation
The EPG, pooled and PERK kernels are C++ over a small tile runtime (_tile.hpp), compiled by nvcc into blochsim._gpu, one file per kernel with 256- and 1024-thread bounded variants, linking the CUDA runtime statically. CMake builds it wherever it finds nvcc (BLOCHSIM_CUDA overrides), and the x86-64 manylinux wheel carries it. The same source compiled for the host (blochsim._gpu_host) replaces Triton's interpreter in the suite and is held to the C++ kernels. Triton, its modules, the interpreted marker and the cache-pruning script are gone. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ND2A7uRuhWiay6B4rWF1H
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ND2A7uRuhWiay6B4rWF1H
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ND2A7uRuhWiay6B4rWF1H
This was referenced Oct 7, 2026
mcencini
added a commit
that referenced
this pull request
Oct 8, 2026
…hem as blochsim[cu12]/[cu13] (#33) The EPG, many-pool and PERK kernels are CUDA compiled ahead of time and written for their layouts: a layout is what is compiled, every other switch steers whole blocks at run time, and each loop is written once over a number type, so the JVPs and the second-order adjoints are the same source at a dual. The adjoints keep checkpoints instead of a recorded trajectory. Every case measured runs at or under Triton's device time; device memory is equal or lower except the three-pool second-order adjoint, which leaves 84 MiB of local memory against Triton's 52. The card's module ships as blochsim-cuda12 and blochsim-cuda13, installed with blochsim[cu12] or blochsim[cu13] beside a torch of the same CUDA major version: it links torch's CUDA runtime by rpath, carries code for 7.5, 8.0 and 9.0 and PTX for 9.0, and is refused by name for another major or release. The package's own structural derivatives are taken in reverse mode, so a simulation and its gradients make PyTorch compile nothing through TorchScript. Supersedes #31 and #32. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requested by Matteo · project thread
What this changes
Before: the GPU kernels were Triton and were JIT-compiled on the first call. The suite checked them through Triton's CPU interpreter, behind the
interpretedmarker.After: the EPG, pooled and PERK kernels are C++, written over a small tile runtime (
_tile.hpp).nvcccompiles them ahead of time intoblochsim._gpu, which links the CUDA runtime statically. Nothing compiles at the first call, and there is no Triton fallback. The public API is unchanged and the dispatch is the same, so consumers need no update.How:
_epg_kernels.hpp,_pools_kernels.hppand_perk_kernels.hpp._kernels.hppis the launch table (parameters, block axes)._gpu_launch.Kernel(name)[grid](...)keeps Triton's call shape. CUDA tensors go to_gpuon torch's current stream; CPU tensors go to_gpu_host, the same source compiled for the host, one program at a time._gpuwherever it findsnvcc;BLOCHSIM_CUDA=ON/OFFoverrides that._gpu_kernel.cu.in, so the build runs in parallel.__launch_bounds__(256)and(1024). The launcher picks the 256 variant whenever the block fits.75-real;80-real;86-real;89-real;90, where90carries PTX for newer cards.before-allinstallscuda-nvccandcuda-cudart-devel). Wheels leave out_gpu_host._epg_triton.py,_pools_triton.py,_perk_triton.py, theinterpretedmarker, the Triton CI step andprune_triton_cache.py. The docs, skills and CLAUDE.md are updated to match.Known limit: a tile is one element per thread, so a launch is at most 1024 threads. An EPG run with more than 1024 state orders on a card, or a pooled run with orders × pools above 1024, is refused with a clear error.
How it was checked
There is no GPU in the container, so nothing here has run on a card.
tests/sequence/test_host_kernels.py).test_many_pools_host.py).test_perk_kernel.py).pytest tests/ -n 6gave 1269 passed, 315 skipped, 13 failed and 12 errors. Every failure and error comes from this container's brokentorchvisioninstall (operator torchvision::nms does not exist, the deepinv import). The one exception istest_an_image_quality_design_fits_in_its_budget, a timing budget that passes when run alone.nvcc12.x for sm_80: every kernel compiles. The heaviest,_epg_vjp_jvp_kernel, takes about 4.5 min per architecture and uses 255 registers with some spills at the 256 bound. At the 1024 bound it spills heavily, and that variant only runs for blocks wider than 256.pre-commit run --all-filesis clean.Needs a card (Matteo):
pip install -e . --config-settings=cmake.define.BLOCHSIM_CUDA=ON, optionally adding--config-settings=cmake.define.CMAKE_CUDA_ARCHITECTURES=native.pytest tests/ -n auto:test_cuda_parity.py,test_both_pools.py,test_perk_kernel.py,test_dynamic_transmit.pyandtest_subspace_streams.pyexercise the device path.Checklist
pytest tests/passes locally (apart from the environment failures above), and new tests cover the host build.pre-commit run --all-filesis clean.🤖 Generated with Claude Code
https://claude.ai/code/session_014ND2A7uRuhWiay6B4rWF1H
Generated by Claude Code