Repository navigation
Let each event of the fused EPG kernel pay only for what it does - #29
Merged
Merged
Conversation
The CUDA kernels carry the two through one launch flag and read both per voxel when either is declared, while the tissue buffers narrowed each to a single value by its own feature. A tissue giving off-resonance beside a transmit map, without a transmit phase, then read the phase past the end of its buffer: a wrong echo phase that changed with the batch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every program reads the same event, so a branch on its kind and action is taken uniformly: a pulse that turns rotates the states, a recorded readout reads them, a shift shifts them, and no event computes the others' work to throw it away. On a 7132-event spoiled stream over 6210 trains and atoms the kernel takes 138 ms where it took 193 ms. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Under RF spoiling a phase grows without bound, and every program took the accurate cosine and sine of it at every pulse and every readout. The launch takes them once per event in double precision; a pulse turns them by the transmit field's phase and reads the double angle off their products. The 7132-event stream over 6210 trains and atoms takes 116 ms where it took 138. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tion The flip angle times the transmit field is an atom's own, so it cannot be taken at the launch; one Cody-Waite reduction and the single-precision Cephes polynomials give both to about an ulp where two library calls each reduced it. The 6210-train stream takes 99 ms where it took 116. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The fused forward EPG kernel (
_epg_kernel) does less work per event, with unchanged results. Stacked on #28: the base iscuda-off-axis-buffers, so the diff shows only the three kernel commits.What changes
tl.where.simulate_intonow takes them once per event, in double precision. The kernel turns them by the transmit field's phase and gets the double angle from their products.Measured
Kernel time alone, best of repeats, on an RTX 4060 laptop GPU. The stream is a 7,132-event spoiled stream captured from a 3D GRE design:
Above about 200 trains × atoms the kernel is throughput-bound, at about 4.4 ns per event per problem before this change. Below that, the per-event latency of one problem's chain sets a floor of about 0.86 µs per event.
Two variants made no difference in this regime, so they are not in this PR:
_TILE_ELEMENTS128/256/512: 30, 39 and 203 ms at 690 problems, against 22 ms);tl.range(..., num_stages=2..4)).Checks
tests/sequence, covering CUDA parity, pools, exchange, simulation, SSFP readouts, train batching, features, diffusion, flow, washout, dynamic transmit, exact profile, lineshape, transmit array, steady state, settling and fixed point.test_against_epgpy.pyis skipped here because epgpy isn't installed.🤖 Generated with Claude Code