Repository navigation
Move BART to Codeberg master with the rebased downstream stack - #104
Merged
Merged
Conversation
The submodule moves from 694ab5fe on bartorch-v1.0.00 to 7fab96d7 on downstream: Codeberg master 3beb5572 plus the 22-commit portability stack. What the library needs from that BART: - C23, which BART now compiles as (bool is a keyword, stdbool.h is gone). - USE_GPU beside USE_CUDA, the macro BART's device code is now under, and simu/gpu_bloch.cu in the CUDA sources. - nufft_create2 and sense_nc_init take a field map and a time map. Both are only served by BART's explicit transform (--nufft-conf dft), which is handed straight to BART; a field map without it is refused by name. sense.c includes modelnc.h so the compiler holds the substitute to the header. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ok1jumatvU8YSzRCEaj8Q
BART no longer has pocsense, sqpics or rtnlinv, so their derived tools go, with the CLI's pocsense adapter and the tests that ran them. apps.pocsense stays, built as before from BART's pocs iteration, and is held to a fully sampled phantom being its fixed point and to its answer lying in the range of the coils. New commands are private: pulseq reads and writes files, sort is an array operation torch has, bet and extractdc are not wrapped yet. unwrap takes a set of axes rather than one; mobafit now requires its output and raga refuses to run without one, so the optional-output exception moves from mobafit to raga. The catalogue reader learns OPTL_FLVEC7 and a help string declared as a pointer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ok1jumatvU8YSzRCEaj8Q
mobafit's -P help writes |M|, which reStructuredText reads as a substitution, and the docs build treats that warning as an error. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ok1jumatvU8YSzRCEaj8Q
…UFFT test Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ok1jumatvU8YSzRCEaj8Q
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
The new misc/dllspec.h gives BARTLIB_API default visibility outside Windows, which put seq_*, num_rand_init and debug_printf in the library's dynamic table beside the bartorch_* ABI. Defining the header's guard with BARTLIB_API empty leaves them hidden, as Windows already has them through BARTLIB_STATIC, and a test holds the table's C names to the ABI. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ok1jumatvU8YSzRCEaj8Q
nvcc's host code took no visibility preset, so the CUDA build exported BART's cuda_* wrappers, eigenmapscu and wl3_cuda_* beside the bartorch_* ABI, which the dynamic-table test now catches. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ok1jumatvU8YSzRCEaj8Q
The Linux wheels carried a libgomp of their own, so a process that imported torch had two OpenMP runtimes. torch carries libgomp.so.1 in torch/lib from 2.7.1 on, which is the name the library's NEEDED entry asks for: the library is loaded after torch on Linux as it already is on macOS and Windows, the wheels are repaired with --exclude libgomp.so.1, and $ORIGIN/../torch/lib is on the runpath. Linux requires torch 2.7.1, the first release whose copy has that name, and the OpenMP test now holds every platform to one runtime, torch's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ok1jumatvU8YSzRCEaj8Q
mcencini
marked this pull request as ready for review
September 29, 2026 09:55
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requested by Matteo · project thread
What changed
external/bartgoes from694ab5fe(onbartorch-v1.0.00, now also taggeddownstream-20260929) to7fab96d7(ondownstream). The new commit is Codebergmaster3beb5572plus 22 downstream commits..gitmodulesis unchanged.12cbf14fis superseded upstream;694ab5feis folded in).-std=gnu2x. It definesUSE_GPUbesideUSE_CUDA, since BART's device code now sits underUSE_GPU.simu/gpu_bloch.cuis added to the CUDA sources.misc/dllspec.hnow givesBARTLIB_APIdefault visibility outside Windows. That putseq_*,num_rand_initanddebug_printfin the dynamic table beside thebartorch_*ABI; the dry-run Linux wheel had 35 of them.BARTLIB_STATICalready does on Windows.cuda_*,eigenmapscuandwl3_cuda_*wrappers; that leak predates this PR, and the new test caught it.tests/test_abi.pynow holds the table's C names to the ABI. It fails on that wheel and passes on this build.auditwheel repair --exclude libgomp.so.1, and$ORIGIN/../torch/libis on the library's runpath._lib._loadimports torch before loading the library on every platform, so the copy torch loaded is the one bound.torch/lib/libgomp.so.1from 2.7.1 on (earlier releases use a hashed name), so Linux now requirestorch>=2.7.1. macOS stays at 2.3.tests/test_openmp.pynow requires exactly one runtime, torch's, on every platform.AGENTS.md,cmake/openmp.cmake, the FINUFFT design page, the installation and prerequisites pages and the license page are updated to match.nufft_create2andsense_nc_initnow take a field map and a time map.--nufft-conf dft(BART's explicit transform, the only one that uses a field map) is handed straight to BART.dftis refused with a message that names the problem.sense.cnow includessense/modelnc.h, so the compiler checks the substitute against BART's declaration. The previous copy compiled against the old signature without complaint._catalogue.pyis regenerated. The reader now handlesOPTL_FLVEC7and a help string declared as a pointer (pulseq). Option help is escaped for|(mobafit's|M|broke the docs build).pulseqreads and writes files.sortis covered by torch.betandextractdcare not wrapped yet.unwrapnow takes a set of axes; the wrapper passes one bit.mobafitnow requires its output, andragarefuses to run without one. The optional-output exception therefore moves frommobafittoraga.test_a_library_without_its_modules_transforms_at_the_baselinecopied the library into an empty folder, which loses anything the wheel vendors beside it. It now keeps the wheel's layout.Why
This moves the fork workflow onto Codeberg history:
masterequals Codeberg, anddownstreamismasterplus the patch stack.Validation
CI on this PR (Tests, Docs, Lint, Typos): green on
58b06cc0.Publish dry run:
workflow_dispatchon58b06cc0, run 36549695519, all green.Publish to PyPIandAttach the CUDA wheel to the releasewere skipped (they run only on tags), so nothing was published.Wheel inspection (downloaded artifacts,
auditwheel,readelf/nm/objdump,macholib,pefile). Linux and CUDA are from run 36549695519; macOS and Windows are from run 36541419983 on87e42582, and nothing since changes what they link.manylinux_2_28_x86_64. The highest GLIBC version any of its libraries references is 2.28.auditwheel showon the runner reports 2_38. That comes from the runner's own libgomp, which it now resolves because the wheel no longer carries one (inferred; the wheel's own objects stop at 2.28).NEEDEDis libc, libm, libdl, libpthread, ld-linux andlibgomp.so.1. There is nobartorch.libs, libstdc++ or libgcc_s. The runpath is$ORIGIN/../torch/lib._mklmodules link no MKL and import no FFTW symbols, and export onlyfinufft*andbartorch_fftw_bind.torch/lib/libgomp.so.1.test_openmp.pyandtest_finufft.pypass against the extracted wheel (77 passed, 15 skipped).NEEDEDislibcudart.so.12,libcufft.so.11,libcublas(Lt).so.12andlibgomp.so.1, with no vendored libgomp.nvidia/*/liband then$ORIGIN/../torch/lib./usr/lib/libSystem.B.dylib,@rpath/libomp.dyliband/usr/lib/libc++.1.dylib.@loader_path/../torch/lib, so the libomp is torch's. No Homebrew,/usr/localor/opt/localpath appears anywhere in the binary.KERNEL32.dll, the UCRTapi-ms-win-crt-*set andlibiomp5md.dll(torch's Intel OpenMP).Local, Linux x86-64, GCC 14:
pytest tests/gave 1806 passed, 97 skipped and 20 failed on40866e73. All 20 failures are packages missing from the container: torchsim, mrinufft, a torch/torchvision mismatch that breaks deepinv, and the MKL FINUFFT module, which needsmkl-include. Lint is clean. Every renamed BART symbol exists in both itsbart_and its substituted form.Needs real hardware (not run; no results claimed)
NVIDIA card:
python scripts/check_device.pywith the CUDA wheel. It covers:picswith and without Toeplitz-gbeing redundantAlso run
tests/test_cuda.py. CI only compiles the device code and loads it without a card; this bump moved it underUSE_GPUand addedsimu/gpu_bloch.cu.Intel and AMD x86-64 (Linux and Windows):
_finufft.simd()picks the newest level the CPU runs: v4 only where AVX-512 exists; AMD Zen 4 and later is v4, and earlier Zen is v3.pip install mkl, the_mklmodule is preferred and agrees with DUCC0.BARTORCH_FINUFFT_SIMDoverrides the choice.Clean Windows machine without MSYS2 on PATH:
pip installthe wheel and torch, then import and run a tool and a NUFFT.libiomp5md.dllis loaded.Apple Silicon outside CI: the wheel with a torch from PyPI, running a tool, a NUFFT and threaded torch in one process, with one libomp mapped.
Public API and documentation
bartorch.tools.pocsense,tools.sqpicsandtools.rtnlinv, andbartorch pocsenseon the command line. Upstream deleted these commands.bartorch.apps.pocsenseis still built from BART'spocsiteration. It is now tested against a fully sampled phantom being a fixed point and against its answer lying in the range of the coils, instead of against the command's bits.torch>=2.7.1on Linux, for the libgomp above.docs/api/tools.md,docs/api/apps.md,docs/explanation/non-cartesian.md, the OpenMP pages listed above, andAGENTS.md.🤖 Generated with Claude Code
https://claude.ai/code/session_014ok1jumatvU8YSzRCEaj8Q