Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 13 additions & 1 deletion cmake/MFCTargets.cmake
Original file line number Diff line number Diff line change
Expand Up @@ -203,7 +203,19 @@ exit 0
-fopenmp-assume-threads-oversubscription
-fopenmp-assume-teams-oversubscription
-fopenmp-assume-no-nested-parallelism)
target_link_options(${a_target} PRIVATE -fopenmp --offload-arch=gfx90a -flto-partitions=${MFC_BUILD_JOBS})
# attributor-max-pi-accesses: amdflang generates device code for the WHOLE
# image at link time, and once the image carries enough target regions the
# device link's Attributor exceeds its AAPointerInfo access cap on a
# heavily-shared object. Pointer information then goes pessimistic and
# OpenMPOpt's __kmpc_parallel cleanup fails module-wide: UNTOUCHED kernels
# regenerate with 2.4-4.5x worse ISA (register spills, +512 B LDS in every
# kernel) whenever ANY kernel is added or removed anywhere in the code.
# Raising the cap restores full pointer precision for the whole image and
# makes kernel quality independent of unrelated edits, at the price of a
# longer device link. See docs/documentation/gpuParallelization.md
# ("AMD flang known issues") for the failure signature.
target_link_options(${a_target} PRIVATE -fopenmp --offload-arch=gfx90a -flto-partitions=${MFC_BUILD_JOBS}
"SHELL:-Xoffload-linker -mllvm -Xoffload-linker -attributor-max-pi-accesses=16384")
endif()
endif()

Expand Down
36 changes: 36 additions & 0 deletions docs/documentation/gpuParallelization.md
Original file line number Diff line number Diff line change
Expand Up @@ -829,6 +829,42 @@ LIBOMPTARGET_JIT_SKIP_OPT=1
- If set, the image will only be passed through the backend.
- The backend is invoked with the `LIBOMPTARGET_JIT_OPT_LEVEL` flag.

## AMD flang (amdflang) Known Issues

### Whole-image device codegen instability (worked around in the build)

amdflang generates device code for the whole image at link time. Once the image carries
enough OpenMP target regions, the device link's `Attributor` pass exceeds its
`AAPointerInfo` access cap on a heavily shared object; pointer information goes
pessimistic and `OpenMPOpt`'s `__kmpc_parallel` cleanup then fails for the whole module.
The visible effect: adding (or removing) ANY kernel anywhere silently regenerates
UNTOUCHED kernels with far worse ISA — measured 2.4-4.5x slower, with register spills
and an extra 512 B of LDS in every kernel. A wall-time A/B between two commits that
differ in target-region count is confounded by this whole-image effect.

MFC's build raises the cap (`-attributor-max-pi-accesses=16384`, passed to the offload
linker in `cmake/MFCTargets.cmake`), which restores full pointer precision for the whole
image and makes kernel quality independent of unrelated edits. The cost is a longer
device link. If a build's device link is unexpectedly slow, this flag is why — do not
remove it; kernel performance becomes nondeterministic across commits without it.

The failure signature without the flag: after adding a kernel, unrelated kernels'
resource usage shifts image-wide (uniform LDS increase, scratch/spill jumps visible in
`rocprofv3` dispatch records) and previously fast kernels slow several-fold.

### Target regions inside Fortran BLOCK constructs are silently dropped

A `GPU_PARALLEL_LOOP` (OpenMP target region) written inside a Fortran `block ...
end block` construct compiles cleanly, but amdflang omits it from the device image
while the host still registers it. The first launch aborts with

hsa_executable_get_symbol_by_name(__omp_offloading_..._l<line>.kd):
HSA_STATUS_ERROR_INVALID_SYMBOL_NAME
omptarget error: Failed to load kernel ...

followed by a segmentation fault. Never place a GPU kernel inside a `block` construct;
hoist it into its own (module) subroutine with the locals passed as arguments.

Comment on lines +865 to +867
## Compiler Documentation

- [Cray & OpenMP Docs](https://cpe.ext.hpe.com/docs/24.11/cce/man7/intro_openmp.7.html#environment-variables)
Expand Down
2 changes: 1 addition & 1 deletion src/simulation/m_rhs.fpp
Original file line number Diff line number Diff line change
Expand Up @@ -275,7 +275,7 @@ contains
if (adv_src_mode == adv_src_mode_vel_iface) then
! u-interface: flux_src(adv%beg) holds one shared face-normal velocity. Pointer-alias adv%beg+1:adv%end to
! the same memory so loops over adv%beg:adv%end can keep fluid indexing while still reading one value. This
! saves (num_fluids - 1) 3D field allocations.
! saves (num_fluids - 1) 3D field allocations
do l = eqn_idx%adv%beg + 1, eqn_idx%adv%end
flux_src_n(i)%vf(l)%sf => flux_src_n(i)%vf(eqn_idx%adv%beg)%sf
$:GPU_ENTER_DATA(attach='[flux_src_n(i)%vf(l)%sf]')
Expand Down
Loading