Repository navigation
Round a pulse the same whichever terms the kernel was compiled for - #30
Merged
Merged
Conversation
The pulse's sine and cosine, its phase turned by the transmit field's, its double angle and the rotation of the states are written as fused multiply-adds in a fixed order. Left to the compiler, each kernel variant contracted them its own way, and a tissue declaring fewer terms got an answer one ulp away from the full kernel's on the same input. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
On a card,
tests/sequence/test_feature_gates.py::test_the_forward_answer_does_not_depend_on_the_gatefails onmainfor all six cases since #29: the forward kernel compiled for the terms a tissue declares answers about one ulp (1.3e-7) away from the kernel compiled with every term, on the same input. The CI runners have no GPU, so #29's checks did not reach it.Cause
#29 replaced library calls with arithmetic: a Cody–Waite reduction and Cephes polynomials for the flip angle, the event phase turned by the transmit phase through products, and the double angle from those products. Those sums of products feed
_rotate_flip_phase. The compiler contractsa*b + c*dinto a fused multiply-add, and which product it fuses depends on the surrounding code. A variant whereoff_axisortransmitis off foldsb1_cos = 1,b1_sin = 0andatom_b1 = 1into constants, and so contracts differently. WithTRITON_DEFAULT_FP_FUSION=0, every gate gives the same bits.Change
The reduction, both polynomials, the transmit-phase turn, the double angle and every term of
_rotate_flip_phaseare explicittl.fmain a fixed order. Nothing is left for the compiler to contract, so every variant rounds the same way.Checks (RTX 4060 Laptop, Triton 3.7.1)
pytest tests/sequence: 1013 passed, 1 skipped (onmain: 6 failed, all intest_feature_gates).main, 2.00–2.03 s here. Turning fusion off for the whole kernel instead gives 2.27 s.🤖 Generated with Claude Code