Skip to content

Support LLVM 22 with Numba 0.66 - #75

Closed
awennersteen wants to merge 11 commits into
Python-for-HPC:mainfrom
awennersteen:aw/support-llvm-22
Closed

awennersteen wants to merge 11 commits into
Python-for-HPC:mainfrom
awennersteen:aw/support-llvm-22

Conversation

@awennersteen

@awennersteen awennersteen commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

README First

This PR is largely developed with Codex on GPT6-Astra medium. It was originally done for me to be able to test some experimental code of my own.
I provide it for the maintainers, since I've seen that in #71 you note that you were going to make the update.
I'm happy to make the effort needed to make this merged, but I'm also happy for the maintainers to decide that its easier to do it yourself.

Description

Build PyOMP's OpenMP pass, host runtime, offload runtime, and GPU device bitcode with LLVM 22.1.8, targeting Numba 0.66.x and llvmlite 0.48.x.

The Numba dependency changes from >=0.62,<0.64 to >=0.66,<0.67 to match llvmlite 0.48's LLVM 22 stack, following the compiler-stack updates in upstream PR #37 and PR #45. Broader cross-version bitcode compatibility has not been validated. Python 3.10–3.14 and Linux x86_64, Linux ARM64, and macOS ARM64 remain in CI.

The implementation adapts LLVM APIs and libomptarget patches and builds NVPTX/AMDGPU device bitcode separately, excluding host CPU flags. Package builds disable LLVM's offload tests and unit tests. Linux and macOS wheels use LLVM 22.1.8 from conda-forge through the existing Miniforge setup. Linux wheels link the compiler runtimes statically for manylinux compatibility.

Validation

CI results for current commit 66690ff:

  • Wheel CI: all three platform wheel builds, all 15 OS/Python installation-test jobs, and the source-distribution build passed. The Modal GPU job stopped before testing because this fork lacks MODAL_TOKEN_ID and MODAL_TOKEN_SECRET.
  • Conda CI: the macOS/Python 3.10 build reached the test suite but failed test_omp_get_wtime (0.2597 seconds versus 0.25 expected). The remaining matrix jobs were cancelled by fail-fast, so the Conda matrix is not fully validated on this commit.
  • Four basic mandatory NVIDIA offload tests previously passed locally on an RTX 3080 with CUDA 12.8. The full GPU suite and Blackwell with CUDA 13 remain unverified; issue Offload fails for Blackwell gpu arch and CUDA 13 #71 is not yet confirmed fixed.

Numba 0.66 selects llvmlite 0.48 and LLVM 22. This dependency change must land with the following LLVM 22 compiler and runtime update; it is not validated as a standalone release.

Assisted-by: Codex
Adapt OpenMPIRBuilder configuration, reduction and target arguments, plugin headers, and the GPU parallel runtime ABI. Port the runtime patches and build the relocated GPU device bitcode separately. Update LLVM build pins and document the tested platform limits.

Validated together with Numba 0.66: 120 host tests, 68 mandatory host-offload tests, and four mandatory RTX 3080 offload tests passed. The full GPU suite and other platforms remain unverified.

Assisted-by: Codex
@ggeorgakoudis

Copy link
Copy Markdown
Contributor

Thank you @awennersteen! I opened #78, which supersedes this PR and also adds numba 0.67. It includes three of your changes, with you credited as co-author: the OpenMPIRBuilder config setting, the host-flag reset for the device runtime build, and the __tgt_get_device_info patch port. I'm closing this PR now that #78 is underway. Reviews on #78 are welcome.

ggeorgakoudis added a commit that referenced this pull request Sep 29, 2026
* Port to LLVM 22 and numba 0.66-0.67

numba 0.66 and 0.67 use llvmlite 0.48 (LLVM 22), so build the pass library
and the OpenMP runtimes against LLVM 22.1.8.

- Port the pass library to the LLVM 22 OpenMPIRBuilder API
  (__kmpc_parallel_60, TargetKernelArgs, ReductionInfo, PassPlugin.h), and
  set the OpenMPIRBuilder config, which createReductions now reads.
- Build Linux wheels with the conda-forge LLVM 22 toolchain, compiling
  against the manylinux gcc-toolset libstdc++.
- Build the nvptx and amdgcn device runtime bitcode from openmp/device,
  ignoring host CFLAGS and CXXFLAGS.
- Keep three runtime patches for 22.1.8 (static LLVM, skip liboffload,
  __tgt_get_device_info) and drop the patch sets for older LLVM versions.
- Require numba >=0.66,<0.68 and test numba 0.66.0 and 0.67.0.
- Drop the conda is_freethreading variant, which no numba pin uses now.

The libomptarget CUDA plugin in LLVM 22 includes llvm/llvm-project#159354
(fixes #71), and llvmlite 0.48 knows sm_110 (fixes #68).

The OpenMPIRBuilder config setting, the device runtime host-flag reset, and
the __tgt_get_device_info patch port are from #75.

* Read AVX512_SKX from the NumPy SIMD features in conda tests

numba 0.66 removed the "NumPy AVX512_SKX detected" sysinfo entry, so the
conda test script failed with a KeyError before running the tests. Check the
"NumPy Supported SIMD features" list instead. The fix is from #75.

* Document PyOMP 0.6 compatibility

---------

Co-authored-by: Aleksander Wennersteen <aleksander.wennersteen@pasqal.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants