Skip to content

Support numba 0.66-0.67 with LLVM 22 - #78

Merged
ggeorgakoudis merged 3 commits into
mainfrom
support-numba-0.67
Sep 29, 2026
Merged

ggeorgakoudis merged 3 commits into
mainfrom
support-numba-0.67

Conversation

@ggeorgakoudis

Copy link
Copy Markdown
Contributor

This PR moves PyOMP to LLVM 22.1.8 and makes numba 0.66–0.67 the supported range, for the 0.6.x release line.

Supersedes #75 and gives credit to selected commits.

Fixes #71 (the LLVM 22 libomptarget CUDA plugin includes llvm/llvm-project#159354).
Fixes #68 (llvmlite 0.48 knows sm_110).

Changes

Pass library

  • Port to the LLVM 22 OpenMPIRBuilder API:
    • __kmpc_parallel_51 becomes __kmpc_parallel_60, with nt_strict = 0.
    • TargetKernelArgs gets the new dyn-groupprivate fallback field.
    • ReductionInfo gets the new DataPtrPtrGen field, set to nullptr.
    • PassPlugin.h is included from its new location.
  • Set OMPBuilder.Config. In LLVM 22, createReductions reads Config.isGPU(), and an unset config is undefined behaviour in release builds of LLVM.
  • The pass now supports LLVM 22 only, so remove every LLVM_VERSION_MAJOR guard.

OpenMP runtimes

  • Keep three patches for 22.1.8: link LLVM statically, skip building liboffload, and add __tgt_get_device_info. Delete the patch sets for LLVM 14, 15, 16 and 20.
  • Build the nvptx and amdgcn device runtime bitcode (libomptarget-{nvptx,amdgpu}.bc) with a separate CMake build per target from openmp/device. LLVM 22 no longer builds these files as part of offload.
  • Clear CMAKE_C_FLAGS and CMAKE_CXX_FLAGS for the device runtime builds to avoid conda host flags breaking GPU targets.

Packaging and CI

  • Linux wheels: manylinux has no LLVM 22 packages, so build with the conda-forge clang/llvmdev 22.1.8 toolchain. Pass --gcc-toolchain=/opt/rh/gcc-toolset-14/root/usr so the libraries link against the image's libstdc++ and stay manylinux_2_28 compatible.
  • The pyomp_loader wrapper is now compiled with CFLAGS, so it uses the same toolchain as the other libraries.
  • Require numba >=0.66,<0.68 in both pyproject and the conda recipe. The test matrices run numba 0.66.0 and 0.67.0.
  • lld 22.1.8 is a new build dependency; it links the AMDGPU device runtime.
  • Remove the conda is_freethreading variant. No numba pin uses it any more.
  • Conda test script: numba 0.66 removed the "NumPy AVX512_SKX detected" sysinfo entry, so read the "NumPy Supported SIMD features" list instead.
  • Docs: add the 0.6.x row to the compatibility tables in the README and installation guide.

ggeorgakoudis and others added 3 commits September 28, 2026 22:51
numba 0.66 and 0.67 use llvmlite 0.48 (LLVM 22), so build the pass library
and the OpenMP runtimes against LLVM 22.1.8.

- Port the pass library to the LLVM 22 OpenMPIRBuilder API
  (__kmpc_parallel_60, TargetKernelArgs, ReductionInfo, PassPlugin.h), and
  set the OpenMPIRBuilder config, which createReductions now reads.
- Build Linux wheels with the conda-forge LLVM 22 toolchain, compiling
  against the manylinux gcc-toolset libstdc++.
- Build the nvptx and amdgcn device runtime bitcode from openmp/device,
  ignoring host CFLAGS and CXXFLAGS.
- Keep three runtime patches for 22.1.8 (static LLVM, skip liboffload,
  __tgt_get_device_info) and drop the patch sets for older LLVM versions.
- Require numba >=0.66,<0.68 and test numba 0.66.0 and 0.67.0.
- Drop the conda is_freethreading variant, which no numba pin uses now.

The libomptarget CUDA plugin in LLVM 22 includes llvm/llvm-project#159354
(fixes #71), and llvmlite 0.48 knows sm_110 (fixes #68).

The OpenMPIRBuilder config setting, the device runtime host-flag reset, and
the __tgt_get_device_info patch port are from #75.

Co-authored-by: Aleksander Wennersteen <aleksander.wennersteen@pasqal.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
numba 0.66 removed the "NumPy AVX512_SKX detected" sysinfo entry, so the
conda test script failed with a KeyError before running the tests. Check the
"NumPy Supported SIMD features" list instead. The fix is from #75.

Co-authored-by: Aleksander Wennersteen <aleksander.wennersteen@pasqal.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@awennersteen

Copy link
Copy Markdown
Contributor

Thank you for taking this over!
Looks good to me, compiles and runs on my machine.

@DrTodd13 DrTodd13 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@ggeorgakoudis
ggeorgakoudis merged commit bc2e2c6 into main Sep 29, 2026
62 checks passed
@ggeorgakoudis
ggeorgakoudis deleted the support-numba-0.67 branch September 29, 2026 15:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Offload fails for Blackwell gpu arch and CUDA 13 Offload problems on Nvidia Jetson AGX Thor (sm_110)

3 participants