Load-balance, active-box windowing, block-structured AMR - #1628
Load-balance, active-box windowing, block-structured AMR#1628sbryngelson wants to merge 595 commits into
Conversation
There was a problem hiding this comment.
Pull request overview
This PR introduces an opt-in (“default-off”) family of performance/diagnostic features (load-weight and SFC partition diagnostics, weighted init-time decomposition, rank timing), plus major simulation capabilities (active-box RHS windowing and block-structured AMR) and corresponding post-processing support and documentation/validation updates.
Changes:
- Adds new runtime parameters and toolchain metadata/validation hooks for the experimental performance/AMR feature family.
- Extends the simulation code with new modules for active-box restriction, load-weight diagnostics, SFC partition reporting, rank timing, and AMR integration points (including restart/output plumbing).
- Updates post_process to read/write AMR fine-block overlays and adds/updates golden metadata plus documentation/indexing.
Reviewed changes
Copilot reviewed 82 out of 94 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
| toolchain/mfc/params/descriptions.py | Adds user-facing descriptions for new experimental/performance parameters. |
| toolchain/mfc/params/definitions.py | Registers new parameters (AMR, hybrid sensors, load-balance diagnostics) and target applicability. |
| toolchain/mfc/lint_docs.py | Treats new validator checks as non-physics doc checks. |
| tests/F980C769/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/ECABA006/golden-metadata.txt | Adds golden metadata for active-box test coverage. |
| tests/DD4CD8F3/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/CC4213FD/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/BD21A5C0/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/BCBA6E74/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/ACE05393/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/987D9025/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/852CCB81/golden-metadata.txt | Adds golden metadata for AMR-related golden tests. |
| tests/65C375B4/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/4DADE04B/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/454C565F/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/3A474BEE/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/2FC423D3/golden-metadata.txt | Adds golden metadata for a new/updated test. |
| tests/13945217/golden-metadata.txt | Adds golden metadata for moving-IB under AMR test coverage. |
| src/simulation/m_viscous.fpp | Clamps FD coefficient indexing to avoid ghost-region coefficient OOB in IB drag gradient evaluation. |
| src/simulation/m_time_steppers.fpp | Integrates active-box bounds into RK update loops and interleaves AMR fine-stage/subcycle operations. |
| src/simulation/m_start_up.fpp | Wires up new modules (rank timing, active-box, load-weight, SFC partition, AMR) into init/timestep/finalize and restart I/O. |
| src/simulation/m_sfc_partition.fpp | Adds analysis-only SFC tiling + weighted partition prediction and reporting. |
| src/simulation/m_rank_timing.fpp | Adds per-rank wall-time imbalance measurement helpers and reporting. |
| src/simulation/m_load_weight.fpp | Adds per-cell load-weight field construction and rank-level imbalance reporting. |
| src/simulation/m_hypoelastic.fpp | Refactors FD coefficient setup into a callable update routine (supporting AMR grid swaps). |
| src/simulation/m_global_parameters.fpp | Adds AMR working-state mirrors and slot selection helper plus defaults for new parameters. |
| src/simulation/m_data_output.fpp | Adds output/report hooks for load-weight, SFC partition, and rank-time diagnostics. |
| src/simulation/m_checker.fpp | Adds input validation/prohibits for active-box, hybrid sensors, load-balance, and AMR configurations. |
| src/simulation/m_active_box.fpp | Adds active-box initialization/growth and debug envelope checking. |
| src/simulation/m_acoustic_src.fpp | Adds AMR-aware handling of acoustic source support (bounding boxes and overlap abort). |
| src/post_process/m_start_up.fpp | Calls AMR fine-data reader and AMR overlay writer when amr is enabled. |
| src/post_process/m_global_parameters.fpp | Adds default-off amr flag for post_process overlay behavior. |
| src/post_process/m_data_output.fpp | Implements AMR fine-block overlay mesh/variables output (Silo/binary) and multimesh registration. |
| src/common/m_phase_change.fpp | Exposes per-cell Newton iteration count and threads it through relaxation to support load-weighting. |
| src/common/m_global_parameters_common.fpp | Adjusts start_idx lifecycle/allocation and makes load_weight_wrt visible to GPU macros. |
| src/common/m_derived_types.fpp | Introduces a simple t_box type used by new partitioning infrastructure. |
| src/common/m_box.fpp | Adds box/partition arithmetic helpers (equal/weighted splits, box-from-splits). |
| src/common/m_boundary_common.fpp | Skips BC buffer population during AMR fine advance to rely on coarse-driven ghost fill. |
| docs/module_categories.json | Registers new modules under documentation categories. |
| docs/documentation/readme.md | Adds AMR section link to the documentation index. |
| .typos.toml | Adds project-specific abbreviations to the spelling allowlist. |
| D = ((gs_min(lp) - 1.0_wp)*cvs(lp))/((gs_min(vp) - 1.0_wp)*cvs(vp)) | ||
|
|
||
| #ifdef MFC_SIMULATION | ||
| if (relax .and. load_weight_wrt) then |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## master #1628 +/- ##
==========================================
+ Coverage 60.77% 61.98% +1.21%
==========================================
Files 83 93 +10
Lines 20872 25637 +4765
Branches 3101 4206 +1105
==========================================
+ Hits 12685 15892 +3207
- Misses 6121 6993 +872
- Partials 2066 2752 +686 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
Upstream latent gap found during the MHD+AMR investigation (independent of this PR): Also for the record: MHD+AMR was attempted and re-gated on measured evidence rather than assumption — the coarse/fine seam is a continuous O(1) div(B) source that cleaning spreads but cannot remove (details in the amr.md support matrix row and commit ac203b1). |
…sking the full block volume" This reverts commit 61b71e4.
Maintainability review follow-up (comment-only, no behavior change): (1) Correct three comments that asserted falsehoods - amr_isect_lo/hi frame is level-dependent (L0 for level-1, parent-fine for level>=2, not a hard-coded 2x); the 'multi-level not yet landed / slots always 1' phrases (multi-level and multi-block ship in this PR); and the L0-reflux 'remaining piece' note (level>=2 blocks reflux to their PARENT via s_amr_reflux_to_parent, implemented). (2) Document the swap contract at the sw_* declarations and in common-pitfalls.md - grid-derived module state a kernel reads on the fine grid must be swapped or refreshed, device copy too (the ab_int regression class); flag amr_rvw as the next candidate; add an AMR note at s_compute_rhs. (3) Complete the TWIN lockstep convention on the q<->pb/mv axis (12 pairs), the coarse RK stage combination (s_tvd_rk <-> the two fine RK updaters), and the freg<->creg capture policy (the 'energy only when not viscous' rule now names all four sites).
The np=1 branches of s_amr_gather_coarse_patch and its pb/mv twin inlined a device kernel byte-identical to s_amr_gather_own_box_device / s_amr_gather_own_box_pbmv_device, which the np>1 path already calls. Replace the inline copies with the shared calls (-30 LOC), removing a drift surface. No behavior change - the surviving kernel is the one already exercised (verified: 21C71558/476AA3A4/1CBACEB5/BCBA6E74/DDD79C8B pass).
s_amr_igr_swap_sigma seeds the fine block's sigma Dirichlet data from the OWNER's LOCAL coarse jac, clamped to the owner's buffer bounds. Unlike q_cons that seed is not P2P-gathered, so a block whose footprint or ghost shell crosses a rank boundary reads clamped edge values instead of the neighbour rank's sigma - a silent wrong answer with zero coverage. Gate to num_procs = 1 (runtime-only, no case_validator mirror) until the sigma seed is distributed. IGR AMR goldens run ppn=1 and are unaffected (6C20B752/660FFBFE pass).
First user-facing example for block-structured AMR: a dense droplet in uniform flow, its interface tracked by a dynamically-regridded, subcycled 2:1 block while smooth regions stay coarse. Exercises the four required AMR settings (amr + initial block, amr_regrid_int > 0, amr_subcycle) plus amr_cluster_eff for the closed-interface case. Validated and smoke-run 150 steps / 15 regrids, clean (no cap, no NaN).
Group the refinement-ratio parameter with the amr_ family so it surfaces under ./mfc.sh params amr and cannot become a breaking rename after merge (reviewer's last fix-before-merge item). Pure mechanical rename of the namelist parameter and its Fortran variable across 8 source files, the param registry/validator/tests, and the AMR docs; the test trace label and golden field values are unchanged (no golden regeneration). Rebuilt all 3 targets; AMR/IGR goldens pass incl. DB9BF199 (the ref_ratio-4 case) and multi-level restart.
Privatize four symbols with no external caller (s_amr_swap_to_fine, s_amr_restore_coarse, s_amr_fill_fine_ghosts, amr_dt_fine) so the audited swap call-sites become a compiler guarantee; narrow m_start_up's bare 'use m_amr' to the five symbols it actually references. t_level stays public (it backs the externally-consumed amr_slots). Build + AMR goldens pass.
amr_block_beg/end note they are 0-based inclusive and required when amr = T; amr_tag_eps states the actual tagging criterion (max axial |rho(i+1)-rho(i-1)|/(2 rho_i)); add amr.md to the lint_docs freshness set (it passes clean).
75AD6885 (multi-level static block) had byte-identical case mods to 4AF96C49 (multi-level restart); restart_check runs the same straight-run golden compare before the restart roundtrip (test.py:547 then :578), so 4AF96C49 fully subsumes it. Removed the stanza + golden dir and folded its description into the surviving test.
Flip restart_check on the existing 3D static-block golden (476AA3A4): the restart reader validates the fine extent per axis, so a z-extent slip would pass every non-restart 3D golden. Reuses the straight-run golden (no new golden); the midpoint roundtrip passes, confirming 3D multi-axis AMR restart round-trips correctly.
Terse pass over the comments in the PR's new AMR/load-balance modules (m_amr, m_amr_regrid, m_amr_restart, m_amr_registers, m_active_box, m_load_balance, m_load_weight, m_sfc_partition, m_rank_timing, m_box): cut filler, redundant restatement, and obvious-code narration, and re-flow the mangled nest_children comments. Comment-only (~37% fewer words, net -124 lines) - every invariant, index/frame convention, compiler-quirk note, TWIN lockstep marker, and trap warning is preserved in substance. Build + 12 AMR/IGR/MHD/load-balance/active-box goldens bit-identical.
Terse pass over the PR-ADDED comments in m_ibm, m_checker, m_time_steppers, and m_rhs (pre-existing files - only the PR's own AMR / fine-IB / swap / gate comments were edited; verified no pre-existing comment or PROHIBIT message string was modified). Cut filler and run-ons while preserving every gate rationale, compiler-quirk note (CCE lib-4425 module separation, declare-target/present-table, ab_int device-copy, AMD defaultmap/ROCm), TWIN marker, and trap warning. Comment-only; build + 8 AMR/IB/IGR/subcycle/QBMM goldens bit-identical.
The Example suite caps grids at 25x25 (modify_example_case), but an AMR block must sit >= buff_size cells inside the domain and span at most half of it - there is no room for a valid block at 25x25, and the example's block indices (sized to its own 128-cell grid) fall outside the capped grid, so the CI Example run failed case validation ('amr_block_end must be <= global cell max per axis'). Reproduced locally. The example stays valid/runnable at native resolution and AMR is covered directly by amr_golden_tests; this only removes it from the downsized Example smoke suite (same reason as the existing 2D_reacting_mixing_layer skip).
|
…forced cross-rank remap (l0_migrate_step)
… bit-identical (l0_rebalance_interval)
… (gather at output), all paths bit-identical
stepfill 71.2% dead / rb-L1 67.9% dead at np=8 - both clear the pre-registered 50% bar; next code increment is the step-fill ring clip. Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
Advection_AmrCore's IC replaced with S0's periodic blob field; per-rank work exactly flat (76.8M cells/rank/step) and matching MFC's S0 (77.1M/rank). MFC's 1.598x/2.63x doublings now have their target. Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
… and amended Two reviewers (consumer lens + state/lifetime lens) both proved the shell claim exact (reach 1 on every branch, zero margin) and both caught the same spec defect: the pull_host gate cannot express lockstep-only, so the scope is now ALL runtime gathers with the subcycle proof recorded. The validation plan is rebuilt around their blindness findings: whole-patch poison (not core - shell holes), MPI_GET_COUNT transport asserts (not tautological recomputation), shadow word accounting (the cov counter is circular with the slab arithmetic), and a boundary-block golden (zero existing coverage). Bindings F1-F11 recorded. Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
The attempted boundary-block golden aborts in the existing runtime checker (blocks must sit buff_size inside the domain), and mar <= buff_size for buff_size >= 2, so no valid patch crosses the boundary. The implementation asserts the premise instead of adding a test for an invalid configuration. Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
Every runtime consumer of amr_cg (fill_cons/gsta/gstb) reads only floor(f/rr)+-1, so of the gathered patch only the shell - the patch minus the open core [region_lo+1, region_hi-1] - is live; [amr-cov] measured the difference at 71% of stepfill words at np=8. Runtime gathers (level-1 pull_host, level>=2 to_host=.false., subcycle included) now ship/copy <= 6 disjoint shell slabs through fused single-launch kernels; contributors whose box lies in the core send zero-word messages so the owner's request set is unchanged. Rebuild/init gathers, the pbmv twin, and the amr_cg coherence walls are untouched. Validation per docs/documentation/amr_stepfill_ring_clip.md: slab disjointness/ coverage asserts, MPI_GET_COUNT transport asserts on every clipped recv, consumer-frame asserts, mar<=buff_size premise assert, and an MFC_DEBUG whole-patch NaN poison before every clipped write. Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
The five shell kernels took their slab bounds as per-launch mapped array dummies (8 arrays each). At the standing LIBOMPTARGET_MEMORY_MANAGER_THRESHOLD=0 default every map is a real hipMalloc/hipFree, and the same-day A/B (rcab-0822: HEAD 407.3 s vs clip 555.8 s at np=4) showed the churn slowing EVERY subsequent kernel launch - rhs +68%/call, rk 3x, restrict +83% - not just the clipped ones. The metadata now lives in a GPU_DECLARE'd 6x8 buffer staged by one 48-int GPU_UPDATE per clipped gather; the kernels read it as a present device array with no per-launch maps. Data path unchanged: same cells, same order, same wire layout. Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
Exact-width slices mapped the same pool column with a different extent every call; reverting to the pre-clip full-column map shape removes the last structural mapping difference vs the unclipped path. (The same-day 5-step probe later showed the wall regression is a constant-factor kernel tax present from step one, so this is hygiene, not the fix - the hunt continues with the 3-minute probe as discriminator.) Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
Reverts a797074, bd85c79, dc6d412. The clip itself is proven correct (output bit-identity at np=4/np=8, zero transport-assert trips, wire words -64 to -72%, gather -33% at np=8), but adding its 7 target regions makes amdflang's whole-image device link deterministically regenerate UNTOUCHED kernels with 2.4-4.5x worse ISA (weno scratch 28->140 B, riemann VGPR 128->40 with AccVGPR 8->136, LDS 2048->2560 image-wide), slowing every compute kernel: rhs 22.1->41 ms/call, np=4 wall 407->553 s. Reproduced deterministically in a second tree; independent of -flto-partitions; same AGPR-blowup family as the weno-agpr reproducer. Parked, not abandoned: the re-landing trigger and full evidence live in amr_action_plan.md '2026-08-22 (final)' and the status header of amr_stepfill_ring_clip.md. Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
Counts level-1 tags at each regrid that fall outside the pre-regrid level-1 coverage - a feature that evolved unrefined because amr_buf did not cover its drift over amr_regrid_int steps. The first regrid (hierarchy population from the seed block) is skipped; the totals print once at finalize alongside [amr-cov]. The case validator gains a non-fatal advisory when amr_buf < amr_regrid_int (a hard rule would wrongly reject the suite's golden int=5/buf=2-3 low-CFL cases; the runtime count is the per-run truth). Verified on S0 at int=2/buf=4: 2.3M steady-state tags, 0 escaped. This is the validity half of the regrid-cadence ledger item: moving the benchmark cadence toward production intervals now has its evidence instrument. Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
spack/rpack allocated (largest-block words) x old_np columns per regrid - GBs at production block counts with nearly every column unused - and rq was O(old_np x num_procs), one of the endstate's forbidden scale terms. Both are now sized to the blocks actually sent/received via dense column maps from a pre-pass that applies the send loop's own destination criterion, and the exact request count. The message set, sizes, tags, and posting order are unchanged: gated on exact [amr-xa] F4 equality (39 msgs, 417682080 words both directions on the S0 np=4 probe), identical [amr-cad] counts, and the 75-case AMR golden subset (75/75). Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
…th (T1/I4b-a) mg:slot measured 57.9 s mean / 90.2 s max at np=8 (5.5% of wall): every 8-16 incremental replica-slot allocs re-staged the whole store through the host. One s_amr_prereserve_stash call per wave grows the store at most once, and s_amr_st_reserve now zeroes only the NEW columns (one full-store host pass saved per growth event). No new device code; store contents byte-identical. Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
…ehind cadence Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
…ttach stash-alloc doc Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
XA_NH header words ([site, blk, bl, bh]) ride ahead of every F1/F2/F3 payload under MFC_DEBUG and are verified before unpacking; zero-width in production. Device kernels untouched (offset via argument slices). Gates: 75/75 AMR goldens (production); debug np=8 probe clean with headers live on all 858 F1 + 1646 F2 messages; a seeded consume-order bug aborted with the full expected-vs-got diagnostic. Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
s_amr_stage_fill_wave: F1+F3 per-stage gathers as one aggregated message per (peer, family) per RK stage - all recvs posted, device packs into pool slices via the existing kernels, one WAITALL, box-major consume through the single amr_cg. Replaces the per-box owner-WAITALL / contributor-flush / F3-blocking-SEND chain. Level>=2 keeps the per-box F2 path (I3); subcycle keeps its sites (I8). Gates: [amr-xa] F1 payload words exact vs baseline (msgs 858->381), F2/F4-F7 byte-identical; live identity headers on every wave transfer (F1 np=8, F3 np=2 with real traffic) + per-message length asserts; seeded offset-shift arm aborts at the header check; adversarial review (grow-helper data-loss bug fixed pre-gate). Outstanding at commit time: full-suite goldens job 383666 (baseline worktree). Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
s_amr_parent_fill_wave(lev): the per-step level>=2 parent gather - previously a pooled ISEND + blocking MPI_RECV per box per stage on the majority of boxes - as one aggregated message per (parent-owner, child-owner) pair per level, levels ascending. Pair lists derive on all ranks from replicated metadata only; reuses the I2a wave's scratch; zero new device kernels. s_amr_fine_stage_fill lost its last caller and is deleted. Regrid chunked, subcycle, and init/static F2 paths unchanged. Also fixes restart leaving amr_num_levels at 1 until the first regrid (the per-level driver needs it truthful; the old per-box loop was immune). Gates: [amr-xa] F2 payload words exact vs baseline, msgs 1646->524, F1 and F4-F7 byte-identical; live identity headers + per-message length asserts on the F2 wave; seeded offset-shift arm aborts at the header check; adversarial review clean (10/10 invariants); local AMR goldens 75/75 incl. multi-level restart np=2 and ppn=4 dynamic regrid. Claude-Session: https://claude.ai/code/session_01N8xV1fowU5LmyfxCivNLDH
The cross-rank branch of s_amr_fine_fine_halo (one blocking MPI_SENDRECV per pair through two shared seam buffers) becomes one aggregated message per (peer, direction) per call: both owners walk the same replicated pair list ascending, so per-peer offsets agree with no metadata exchange. Same-rank pairs keep the batched device kernel. The shared seam buffers and their tile-grow reconciliation are deleted. Gates: [amr-xa] F6 payload words exact vs baseline (380,849,184), msgs 2646->378; all other families byte-identical; debug identity headers + length asserts clean; seeded header-shift arm aborts; local AMR-75 goldens 75/75.
The last per-box rendezvous chains: the level-1 reflux-face exchange (s_amr_p2p_reflux_faces per box) becomes s_amr_reflux_faces_wave — receives post zero-copy into the freg register host mirrors, owner D2H + multicast ISENDs, one WAITALL, receivers push H2D — and the level>=2 split-ownership freg handoff becomes s_amr_freg_wave. s_amr_reflux_to_parent gains do_xchg so the subcycle path keeps its per-box exchange. Message count is unchanged by design (zero-copy into per-box register slots); the wave removes the O(boxes) rendezvous chain. s_amr_reg_reserve hoists ahead of both waves since the apply can reallocate the registers. Gates: [amr-xa] F5 payload words exact (659,423,232), msgs 6708 unchanged; all other families byte-identical; debug-only companion identity headers + length asserts clean; seeded blk-shift arm aborts at the companion check; local AMR-75 goldens 75/75.
…1 blocks The per-box s_amr_apply_reflux launched up to 3 tiny face kernels per block per step; the launch overhead, not the arithmetic, dominated the reflux-apply phase (10.4 pct of the np8 step post-wave). A host precompute now walks the level-1 slots with the same select_slot + face-flags logic and fills per-slot descriptor arrays; one kernel per face direction corrects the coarse rhs for every block, mirroring the capture-side batching. The per-box form is deleted (single call site); the subcycle path is untouched. Block corrections are disjoint (merge invariant) and a block's x/y/z outside layers are distinct cells, so per-(face, eq, cell) arithmetic and child-sum order are identical. Gates: step-5 full-state output byte-identical to the pre-batch binary on the np8 probe; amr-xa byte-identical; reflux calls/rank 759 to 15; local AMR-75 goldens 75/75.
…npack, overlap carry-forward) The migration half of the regrid budget was host-staged end to end: full-slot device-host round trips per owned old block at stash creation, serial host cast loops for the MPI pack and unpack, a full-slot push per received replica, and a host overlap carry-forward. All four now run where the store is authoritative: a device cons-to-stor stash kernel, device pack/unpack kernels whose copyin/copyout stage exactly the packed interior (wire layout byte-identical, message set and amr-xa F4 totals unchanged), and a device overlap kernel with the per-box prolong push hoisted ahead of it (same final device state). The stash no longer touches the host mirror, retiring the two grow-hazard push sites. Gates: step-5 output byte-identical to the previous binary on the np8 probe; amr-xa byte-identical; probe rg:mig -41 pct, unpack 74 to 11 ms/call; AMR-75 75/75. Also documents a new compiler trap: amdflang silently drops target regions nested in Fortran block constructs from the device image (INVALID_SYMBOL_NAME at first launch) - kernels must live in ordinary subroutines.
…shes s_prolong_one_var and the alphas/species closure prolongs were host loops writing the cons host mirror, forcing a full-slot H2D push per built box at three call sites (rebuild, startup populate, persistent-L2 build). All three are now GPU kernels: the gathered patch's device mirror is pushed once per dispatch (patch-sized, 8x smaller than the slot pushes it replaces), the slot is built in place, and every full-slot push is deleted. minmod was already GPU_ROUTINE-decorated; the shared alpha limiter switch is inlined into its kernel and the helper deleted. With the device-side migration this completes the rebuild's device-resident data path. CPU builds compile the kernels to the identical plain loops (CPU results unchanged); GPU prolong arithmetic moves device-side and gates on golden tolerance. Gates: np8 probe rc=0 with amr-xa byte-exact; AMR-75 75/75.
s_amr_st_reserve preserved live blocks through a growth by pulling the entire store device to host, reallocating, and pushing it back - for each of up to four arrays per event. The round trip's one remaining purpose (carrying the migration stash's host writes across a growth) died with the device-side migration, so growth now stages through a device-mapped temporary: two on-device copies, zero PCIe. The host mirror comes out of a growth undefined, within the existing contract (every host reader pulls its slot first; compaction already leaves the host stale by design). The 5-step np8 probe fell 100.8 to 43.4 s (-57 pct) with byte-identical output - the growth round trips were a large untimed cost in every short run; the operating-point yield is smaller (high-water reached early at int=20). Gates: step-5 output byte-identical; amr-xa unchanged; AMR-75 75/75. Follow-up: the registers REG_GROW macro keeps the same pattern.
Probe verdict recorded in the ledger: the plan walks cost 0.01 ms/call, refuting the plan-caching increment (I6) before it was built; the wave's residue is pack copyouts and the WAITALL. Next comm target re-ranked to ring-clip-on-waves.
…t, restart pad zeroing Adversarial review of the day's four increments returned three findings, all fixed here. (1) The batched reflux apply relies on block corrections being disjoint, but the IB path re-broke the invariant: s_amr_expand_box_over_bodies runs after clustering and the follow-up merge fused only overlapping pairs, so two boxes left with a 1-cell gap had coincident outside coarse cells - an unsynchronized read-modify-write inside one kernel. The IB merge now fuses pairs closer than a 2-cell gap. (2) The device-native store grow transiently held old plus staging = 2x the array on device at exactly the memory high-water mark, a measured OOM class; device staging now applies only up to 32 old columns (where the short-run win concentrates) and larger grows keep the OOM-safe host round trip. (3) Restart pushes full padded columns after writing only interiors, and the host pad bytes are undefined since the device-native grow; the host column is now zeroed (owner-only) before each restart read. Gates: bit-identity vs the pre-fix binary on the np8 probe; AMR-75 75/75 with the IB and restart-roundtrip cases green.
The reverted clip (proven correct; killed by the amdflang whole-image codegen bug, since root-caused with a verified link-flag workaround) is reimplemented on the stage-fill wave: after each pair's box intersection, the slab is clipped against the patch's hollow shell (the open core is provably dead - amr_stepfill_ring_clip.md), yielding up to six sub-slab transfers derived identically on both sides from replicated metadata. The primitives (shell slab decomposition, clip, the debug NaN-poison arm, the shell-only own-box copy) are lifted verbatim from the reverted implementation; pack, unpack, and consume were already generic over slab bounds, so sub-slabs are just more transfers. The pbmv gather keeps its full-box wire contract (qbmm plus non-polytropic runs stay unclipped, as the original deliberately did). Messages stay at the per-peer count; only payload drops. Gates: F1 payload words 1,071,084,168 to 416,141,172 (-61 pct) with message count unchanged and every other family byte-identical; step-5 output byte-identical on the np8 probe; the debug NaN-poison arm runs clean (any consumer read of an unshipped cell aborts); seeded header arm aborts on a shell sub-slab; AMR-75 75/75; wall not in the slow codegen class.
Lines of Code
|
Summary
An opt-in, default-off family of performance features and the measurement infrastructure they rest on. With all flags at their defaults the only touched production path is
s_mpi_decompose_computational_domain, refactored through the newm_boxmodule (byte-identical; covered by the existing suite).m_box(partition arithmetic),m_load_weight/load_weight_wrt(per-cell load-weight field + imbalance metric),m_sfc_partition/sfc_partition_wrt(Morton-SFC predicted-imbalance diagnostic),m_load_balance/load_balance(weighted static decomposition at init; AMR-fine-work-aware),m_rank_timing/rank_time_wrt(per-rank compute-time diagnostic).m_active_box/active_box: restricts reconstruction/Riemann/RK windows to a light-cone-grown box around non-ambient flow; strict-subset golden-tested.hybrid_wenoandhybrid_riemann(+hybrid_weno_eps,hybrid_smooth_flux): linear-optimal weights / central-or-Rusanov flux in smooth cells, full WENO/HLLC at flagged discontinuities (Jameson sensor, stencil-dilated, per-level under AMR).m_amr+m_amr_registers: two-level 2:1 refined block hierarchy; conservative restriction and conservative-linear prolongation with physics-specific closures; per-stage flux registers with Berger–Colella refluxing; Berger–Rigoutsos multi-block dynamic regrid; optional dt/2 subcycling; multi-rank (single-owner blocks assigned by Morton-SFC work balancing at each regrid, with migration; blocks may span rank seams via P2P coarse↔fine gather/scatter; same-level seam halo; distributed registers); restart (both IO modes, regridded-layout persistence); AMR-aware post-processing (fine blocks visualizable as Silo overlay domains); GPU-resident fine level on both OpenACC and OpenMP offload.Full algorithm and user documentation:
docs/documentation/amr.md(support matrix enforced at runtime by the checker — unsupported combinations abort with named messages, never silently).AMR physics support matrix (abridged; authoritative table in amr.md)
Supported and golden-tested: single- and multi-fluid (5-eq,
mpp_lim) · 6-eq with per-block pressure relaxation · viscous (refluxed) · phase change (relax) · chemistry incl. species diffusion · Euler–Euler bubbles (polytropic/non-polytropic, mono/polydisperse, QBMM incl. non-polytropic with per-blockpb/mvside-state; dynamic regrid + subcycle) · acoustic sources (coarse-grid support with regrid exclusion) · immersed boundaries (multi-body, static or prescribed-motion, incl. dynamic regrid with body-containment expansion and per-substage guards) · 2D axisymmetric (per-block WENO-coefficient recompute) · stretched grids (exact parent-bisection ghost coordinates + per-swap coefficient recompute) · hybrid WENO/Riemann sensors (per-level) · Lagrangian bubbles (cloud excluded from blocks; two-way coupling on the coarse grid; regrid clips around the moving cloud) ·active_box(blocks contained in the growing window; agrees with plain AMR to ~1e-14) · IGR (restriction-only coupling: fine sigma solve seeded/Dirichlet-bounded by the coarse solve; documented truncation-order seam, exact free-stream) · 1D MHD/RMHD (div(B)=0 by construction in 1D; HLL and HLLD, incl. relativistic).Gated with named aborts (documented rationale): surface tension (seam force imbalance is structural — three fixes attempted and diagnosed in amr.md) · 2D/3D MHD (attempted and measured: the c/f seam is a continuous O(1) div(B) source GLM cleaning cannot remove — needs constrained-transport-class B prolongation/reflux) · hyperelasticity · 3D cylindrical (global azimuthal filter) · force-driven IB (
moving_ibm=2) · STL bodies · Riemann-extrapolation BCs (bc=-4) ·amr_subcycleunder IGR · stretched grids with Lagrangian/IB-regrid (uniform-spacing index conversions).Validation evidence
Known issues (all non-gating or in progress)
continue-on-error): an intermittent post-detected NaN on the two Lagrangian+AMR goldens. Exhaustively unreproducible off GitHub's runners — the exact failing stack (NVHPC 24.3 SDK,-tp=px -Kieee, HPC-X MPI, and the CI docker image itself under apptainer) passes elsewhere, as do native/zen2 builds; 24.5+ green. Documented at the golden definitions.Review guide
The commit history is arc-ordered (active-box → load-weight → SFC → weighted decomposition → rank timing → hybrid → m_box → AMR rungs → physics envelope → CI/GPU hardening); reviewing by arc is much easier than by file. The AMR arc builds stepwise: static hierarchy → restriction/prolongation → fine advance → refluxing → regrid → subcycling → multi-rank → GPU → each physics rung with its own validation. Commit messages carry the validation evidence for their change (measured defects, golden UUIDs, repro details for CI fixes).
All parameters ship default-off with
case_validatorentries, runtime checker gates, andcase.md/amr.mddocumentation.