Skip to content

Let each event of the fused EPG kernel pay only for what it does - #29

Merged
mcencini merged 4 commits into
mainfrom
epg-kernel-event-branches
Oct 6, 2026
Merged

mcencini merged 4 commits into
mainfrom
epg-kernel-event-branches

Conversation

@mcencini

@mcencini mcencini commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

The fused forward EPG kernel (_epg_kernel) does less work per event, with unchanged results. Stacked on #28: the base is cuda-off-axis-buffers, so the diff shows only the three kernel commits.

What changes

  • Each event does only its own work. Every program reads the same event, so a branch on the event's kind and action is taken by all programs alike and costs no divergence. A pulse that turns rotates the states, a recorded readout reads them, and a shift or spoil moves or clears them. No event computes another kind's work only to discard it with tl.where.
  • Phase trig is taken once, at launch. RF-spoiled phases grow without bound, and every program used to take the accurate cosine and sine at every pulse and readout. simulate_into now takes them once per event, in double precision. The kernel turns them by the transmit field's phase and gets the double angle from their products.
  • The flip angle uses one reduction. The flip angle times B1 is per atom, so it can't be tabulated at launch. One Cody–Waite reduction by a quarter turn plus the single-precision Cephes polynomials gives its sine and cosine to about an ulp. Previously two library calls each made their own reduction.

Measured

Kernel time alone, best of repeats, on an RTX 4060 laptop GPU. The stream is a 7,132-event spoiled stream captured from a 3D GRE design:

trains × atoms before branches + phase tables + flip sincos
690 22.4 ms — 15.4 ms 15.5 ms
6,210 193 ms 138 ms 116 ms 99 ms

Above about 200 trains × atoms the kernel is throughput-bound, at about 4.4 ns per event per problem before this change. Below that, the per-event latency of one problem's chain sets a floor of about 0.86 µs per event.

Two variants made no difference in this regime, so they are not in this PR:

  • more problems per program (_TILE_ELEMENTS 128/256/512: 30, 39 and 203 ms at 690 problems, against 22 ms);
  • a software-pipelined event loop (tl.range(..., num_stages=2..4)).

Checks

  • Agreement with the old kernel on the captured stream:
    • the branches alone: 1.3e-8 relative;
    • with the phase tables: 4.8e-8;
    • with the flip-angle sincos: 3.7e-6, from the polynomial against libdevice's accurate routines.
  • Test suite: 498 tests pass in tests/sequence, covering CUDA parity, pools, exchange, simulation, SSFP readouts, train batching, features, diffusion, flow, washout, dynamic transmit, exact profile, lineshape, transmit array, steady state, settling and fixed point.
  • Not run: test_against_epgpy.py is skipped here because epgpy isn't installed.

🤖 Generated with Claude Code

mcencini and others added 4 commits October 5, 2026 23:00
The CUDA kernels carry the two through one launch flag and read both per
voxel when either is declared, while the tissue buffers narrowed each to a
single value by its own feature. A tissue giving off-resonance beside a
transmit map, without a transmit phase, then read the phase past the end of
its buffer: a wrong echo phase that changed with the batch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every program reads the same event, so a branch on its kind and action is
taken uniformly: a pulse that turns rotates the states, a recorded readout
reads them, a shift shifts them, and no event computes the others' work to
throw it away. On a 7132-event spoiled stream over 6210 trains and atoms the
kernel takes 138 ms where it took 193 ms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Under RF spoiling a phase grows without bound, and every program took the
accurate cosine and sine of it at every pulse and every readout. The launch
takes them once per event in double precision; a pulse turns them by the
transmit field's phase and reads the double angle off their products. The
7132-event stream over 6210 trains and atoms takes 116 ms where it took 138.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tion

The flip angle times the transmit field is an atom's own, so it cannot be
taken at the launch; one Cody-Waite reduction and the single-precision
Cephes polynomials give both to about an ulp where two library calls each
reduced it. The 6210-train stream takes 99 ms where it took 116.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mcencini
mcencini changed the base branch from cuda-off-axis-buffers to main October 6, 2026 13:25
@mcencini
mcencini merged commit 51f5a40 into main Oct 6, 2026
18 checks passed
@mcencini
mcencini deleted the epg-kernel-event-branches branch October 6, 2026 13:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant