Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
16 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/ISSUE_TEMPLATE/bug_report.yml
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,7 @@ body:
multiple: true
options:
- CPU
- CUDA (Triton kernels)
- CUDA
- Both
validations:
required: true
2 changes: 1 addition & 1 deletion .github/SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ a signal model are all Python that runs in your process, so a malicious
In scope is anything that turns *data* into execution or into memory
corruption -- a sequence description, a pulse waveform, a phantom or a
dictionary read from a file, or values passed to the simulator, reaching the
C++ and Triton kernels. Those kernels index raw pointers, so an out-of-bounds
C++ and CUDA kernels. Those kernels index raw pointers, so an out-of-bounds
read or write reachable from ordinary arguments is a vulnerability and not
merely a bug.

Expand Down
2 changes: 1 addition & 1 deletion .github/skills/add-a-sequence/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,7 @@ evaluates. Adding an effect is one line in a model's property declaration, plus
the term itself in **both** kernels.

The shared parameter ABI is `src/blochsim/sequence/_parameters.py`, read by the
Python dispatch, the C++ extension and the Triton kernels. A parameter added
Python dispatch, the C++ extension and the GPU kernels. A parameter added
there is added in all three or in none.

## What the change has to arrive with
Expand Down
38 changes: 11 additions & 27 deletions .github/skills/build-and-test/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: build-and-test
description: Compile the C++ kernels and run the BlochSim suite, including the Triton paths on a machine with no GPU. Use when asked to build, install, test, or reproduce a failure in blochsim.
description: Compile the C++ kernels and run the BlochSim suite, including the GPU kernels on a machine with no GPU. Use when asked to build, install, test, or reproduce a failure in blochsim.
---

# Build and test BlochSim
Expand All @@ -16,7 +16,8 @@ run alone, `pip install -e ".[test]"` is enough and much smaller.

An editable install puts the Python sources on the path, so an edit under
`src/blochsim` takes effect on the next import. The two C++ extensions are
compiled artifacts and do not: re-run the install after editing any `.cpp`.
compiled artifacts and do not: re-run the install after editing any `.cpp`,
`.hpp` or `.cu`.

**Read the exit status, not the output.** A failed compile leaves the
previously built `.so` importable, so the suite runs green against a kernel
Expand All @@ -41,32 +42,15 @@ against closed forms — the shift, the RF rotation, relaxation, diffusion, flow
spoiling, the two-pool and three-pool longitudinal steps — while `sequence/`,
`model/`, `estimators/`, `recon/` and `optim/` cover the layers above.

## The Triton paths, without a GPU
## The GPU kernels, without a GPU

The tests carrying the `interpreted` marker are deselected by `addopts`. They
run a Triton kernel through Triton's CPU interpreter — no card, no compile,
about a minute each — and are how the GPU plumbing is verified anywhere:

```sh
pytest tests/ -m interpreted
TRITON_INTERPRET=1 python your_script.py # the same trick, by hand
```

Triton has no wheel for macOS or Windows. On those platforms the default
deselection is not an optimisation, it is the only thing that runs.

## Before you time anything

Kernel compiles dominate a cold GPU run, not the arithmetic. A suite that takes
minutes on a card is mostly Triton compiling one specialization per feature
combination it meets; the second run of the same suite is a different number
entirely. Run the whole suite at natural boundaries rather than after every
edit.

The cache keys on the *source* of `_epg_triton.py`, not on what it means, so a
formatting pass or a deleted dead local costs the same full recompile as a new
`tl.constexpr`. If the suite suddenly takes an hour where it took minutes,
check whether that file changed before looking for a performance regression.
On Linux the install compiles the GPU kernels for the host as
`blochsim._gpu_host`, and
for the card as `blochsim._gpu` wherever CMake finds `nvcc`
(`--config-settings=cmake.define.BLOCHSIM_CUDA=ON` insists on it).
`tests/sequence/test_host_kernels.py` and `test_many_pools_host.py` run the
host build against the C++ kernels, which is how the GPU kernels are verified
on any machine; the CUDA tests skip themselves without a card.

## Style

Expand Down
17 changes: 4 additions & 13 deletions .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -48,14 +48,6 @@ jobs:
if: runner.os == 'Linux'
run: python -m pip install torch --index-url https://download.pytorch.org/whl/cpu

# Triton reaches a runner as a dependency of the CUDA build of PyTorch,
# which the CPU index above deliberately leaves out, and it exists for
# Linux alone. Installing it is what lets the kernel-gate tests run; on
# the other two platforms they skip themselves.
- name: Install Triton (Linux only)
if: runner.os == 'Linux'
run: python -m pip install triton

# Not editable and not from a wheel: the C++ kernels are compiled here,
# on this runner, with this compiler, which is the half of the package a
# pure-Python job would never exercise. CMake and Ninja arrive as
Expand All @@ -69,13 +61,12 @@ jobs:
import importlib.util, pathlib
spec = importlib.util.find_spec("blochsim")
root = pathlib.Path(next(iter(spec.submodule_search_locations)))
for kernel in sorted(root.glob("_*_cpu*")):
for kernel in sorted(root.glob("_*_cpu*")) + sorted(root.glob("_gpu*")):
print(kernel.name, kernel.stat().st_size, "bytes")

# What a runner with no card runs: the C++ kernels and every layer above
# them. The GPU tests skip themselves on ``torch.cuda.is_available()``,
# and the ones that reach for Triton's CPU interpreter carry the
# ``interpreted`` marker, which ``addopts`` deselects.
# What a runner with no card runs: the C++ kernels, the GPU kernels
# compiled for the host, and every layer above them. The GPU tests skip
# themselves on ``torch.cuda.is_available()``.
- name: Run the tests
run: pytest tests/ -n auto

Expand Down
160 changes: 159 additions & 1 deletion .github/workflows/wheels.yml
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,9 @@ on:
- setup.py
- "src/blochsim/*.cpp"
- "src/blochsim/*.hpp"
- "src/blochsim/*.cu"
- "src/blochsim/*.cu.in"
- "src/cuda/**"
- scripts/check_wheel.py
- .github/workflows/wheels.yml
workflow_dispatch:
Expand Down Expand Up @@ -88,9 +91,144 @@ jobs:
name: dist-${{ matrix.os }}
path: wheelhouse/*.whl

# The GPU kernels are a package of their own per CUDA major version,
# blochsim-cuda12 and blochsim-cuda13 (src/cuda/), which blochsim loads when
# torch is built for the same major version: `pip install blochsim[cu12]`.
# Each carries the card's module alone and links the CUDA runtime that
# torch's CUDA builds bring, and is compiled with the oldest toolkit of its
# major that torch ships, so it runs against every torch of it.
wheels-cuda:
name: Build blochsim-cuda${{ matrix.cuda }} (linux x86_64)
runs-on: ubuntu-latest
timeout-minutes: 150
container:
image: quay.io/pypa/manylinux_2_28_x86_64
strategy:
fail-fast: false
matrix:
include:
# CUDA 12.6's nvcc takes GCC 13 at the newest as its host compiler.
- cuda: "12"
toolkit: "12-6"
home: /usr/local/cuda-12.6
host: gcc-toolset-13-gcc-c++
hostcxx: /opt/rh/gcc-toolset-13/root/usr/bin/g++
- cuda: "13"
toolkit: "13-0"
home: /usr/local/cuda-13.0
host: ""
hostcxx: ""
env:
CUDA_HOME: ${{ matrix.home }}
steps:
- uses: actions/checkout@v7
with:
# setuptools-scm reads the version from the tag history, and the
# CUDA build pins the blochsim of the same version.
fetch-depth: 0
- name: Install the CUDA toolkit
run: |
dnf install -y dnf-plugins-core
dnf config-manager --add-repo \
https://developer.download.nvidia.com/compute/cuda/repos/rhel8/x86_64/cuda-rhel8.repo
dnf install -y --nogpgcheck cuda-nvcc-${{ matrix.toolkit }} cuda-cudart-devel-${{ matrix.toolkit }} ${{ matrix.host }}
echo $CUDA_HOME/bin >> $GITHUB_PATH
echo /opt/python/cp312-cp312/bin >> $GITHUB_PATH
- name: Build the wheel
env:
CUDACXX: ${{ matrix.home }}/bin/nvcc
run: |
if [ -n "${{ matrix.hostcxx }}" ]; then export CUDAHOSTCXX=${{ matrix.hostcxx }}; fi
python -m pip install build
python -m build --wheel --outdir dist/ src/cuda/${{ matrix.cuda }}
# The CUDA runtime is excluded rather than vendored: torch already brings
# it, from the nvidia wheel the module's rpath points into.
- name: Repair to a manylinux tag
run: |
python -m pip install auditwheel
LD_LIBRARY_PATH=$CUDA_HOME/lib64 python -m auditwheel repair dist/*.whl \
-w wheelhouse/ --plat manylinux_2_28_x86_64 --exclude libcudart.so.${{ matrix.cuda }}
# The card's module and the package around it, and nothing vendored.
# PyPI takes a file of 100 MB at most unless a project is granted more.
- name: Inspect the CUDA wheel
run: |
python -m auditwheel show wheelhouse/*.whl
python -m zipfile -l wheelhouse/*.whl | tee contents.txt
! grep -Ei 'libcudart|\.libs/' contents.txt
grep -q 'blochsim_cuda${{ matrix.cuda }}/_gpu' contents.txt
! grep -E '^ *blochsim/' contents.txt
size=$(stat -c %s wheelhouse/*.whl)
echo "wheel: $size bytes"
test "$size" -lt 100000000
- uses: actions/upload-artifact@v7
with:
name: cuda-wheel-${{ matrix.cuda }}
path: wheelhouse/*.whl

# What a user installs: the blochsim wheel with its extra, beside torch's
# build for the same CUDA major version, on the oldest and the newest Python
# the package supports. The runners have no card, so this proves that the
# install resolves -- the CUDA build's pins against torch's -- that blochsim
# loads the CUDA build, and that the runtime it links is the one torch's
# nvidia wheel installed, found from the module's own directory.
test-cuda-wheels:
name: Test blochsim[cu${{ matrix.cuda }}] on Python ${{ matrix.python }}
needs: [wheels, wheels-cuda]
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
cuda: ["12", "13"]
python: ["3.10", "3.14"]
include:
- cuda: "12"
torch: cu126
- cuda: "13"
torch: cu130
steps:
- uses: actions/checkout@v7
- uses: actions/setup-python@v7
with:
python-version: ${{ matrix.python }}
- uses: actions/download-artifact@v8
with:
name: dist-ubuntu-latest
path: base
- uses: actions/download-artifact@v8
with:
name: cuda-wheel-${{ matrix.cuda }}
path: cuda
# Both wheels are named as files, so pip takes them whatever their
# version, and the extra's requirement is the CUDA wheel given beside it.
- name: Install blochsim[cu${{ matrix.cuda }}] with torch's ${{ matrix.torch }} build
run: |
python -m pip install --upgrade pip
pip install "$(ls base/*manylinux*x86_64*.whl)[cu${{ matrix.cuda }}]" cuda/*.whl \
torch --index-url https://download.pytorch.org/whl/${{ matrix.torch }} \
--extra-index-url https://pypi.org/simple
pip list | grep -Ei '^(torch|blochsim|nvidia-cuda-runtime)'
- name: The card's module loads bare, against torch's runtime
run: python scripts/check_wheel.py
- name: blochsim loads the CUDA build for torch's major version
run: |
python -c "
import torch
from blochsim import _gpu_launch
module = _gpu_launch._module('cuda')
print(torch.__version__, module.__name__)
assert torch.version.cuda.split('.')[0] == '${{ matrix.cuda }}', torch.version.cuda
assert module.__name__ == 'blochsim_cuda${{ matrix.cuda }}._gpu', module.__name__
assert _gpu_launch.available()
"

# PyPI hands a trusted publisher a token for one project, the one whose
# publisher matches the job's repository, workflow and environment, so each
# project is published from an environment of its own: pypi for blochsim,
# pypi-cuda12 and pypi-cuda13 for the CUDA builds. The CUDA builds pin the
# blochsim they were built with, so they follow it.
publish:
name: Publish to PyPI
needs: [sdist, wheels]
needs: [sdist, wheels, wheels-cuda, test-cuda-wheels]
if: startsWith(github.ref, 'refs/tags/v')
runs-on: ubuntu-latest
environment:
Expand All @@ -107,3 +245,23 @@ jobs:
path: dist

- uses: pypa/gh-action-pypi-publish@release/v1

publish-cuda:
name: Publish blochsim-cuda${{ matrix.cuda }} to PyPI
needs: [publish]
if: startsWith(github.ref, 'refs/tags/v')
runs-on: ubuntu-latest
strategy:
matrix:
cuda: ["12", "13"]
environment:
name: pypi-cuda${{ matrix.cuda }}
url: https://pypi.org/p/blochsim-cuda${{ matrix.cuda }}
permissions:
id-token: write
steps:
- uses: actions/download-artifact@v8
with:
name: cuda-wheel-${{ matrix.cuda }}
path: dist
- uses: pypa/gh-action-pypi-publish@release/v1
2 changes: 2 additions & 0 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,8 @@ repos:
hooks:
- id: check-added-large-files
args: [--maxkb=512]
# The EPG kernels for the card, in every mode.
exclude: ^src/blochsim/_epg_kernels\.hpp$
- id: check-case-conflict
- id: check-merge-conflict
- id: check-symlinks
Expand Down
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,18 @@
- **The package is `blochsim`.** The distribution, the import name and the
repository are `blochsim`; `import torchsim` becomes `import blochsim`, with
the same modules and names beneath it.
- **The GPU kernels are CUDA compiled ahead of time, and Triton is not used.**
The EPG, pooled and PERK kernels are C++ (`_epg_kernels.hpp`,
`_pools_kernels.hpp`, `_perk_kernels.hpp`, and the layouts of `_layout.hpp`)
compiled by `nvcc` into a module of their own per CUDA major version:
`pip install blochsim[cu12]` or `blochsim[cu13]` installs `blochsim-cuda12`
or `blochsim-cuda13` beside a torch of the same major version, which links
the CUDA runtime that torch brings and carries code for 7.5, 8.0 and 9.0
cards. A source build compiles it beside the package wherever CMake finds
`nvcc`. No kernel is compiled at the first call. The same kernels are
compiled for the host as `blochsim._gpu_host`, which the suite holds to the
C++ kernels; the `interpreted` marker is gone. A launch is at most 1024
threads, so an EPG run of more than 1024 state orders on a card is refused.

### Added

Expand Down
Loading
Loading