Conversation
The parallel linear algebra publishes new pivot rows with __sync_bool_compare_and_swap, a GCC builtin that is not available with MSVC. Moreover, the other threads read these pivot arrays with plain loads, which is a data race under the C11 memory model. The pivot arrays that threads add pivots to concurrently are now arrays of _Atomic pointers. A new pivot is published with a compare and swap using release ordering, the reducers load entries with acquire ordering, so a published row and its coefficients are fully visible to them. Sequential code after the parallel regions uses relaxed accesses. Arrays only read during parallel regions stay plain, and the compare and swap in the sequential kernel computation of the saturation is replaced by a plain check and store. The prototypes of the internal 32 bit reducers taking atomic arguments move from the installed data.h to data.c, so that the public headers stay usable from C++. configure now checks for C11 atomics. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The parallel linear algebra in neogb publishes new pivot rows with
__sync_bool_compare_and_swap. That GCC builtin isn't available with MSVC. Also, other threads read the same pivot arrays with plain loads, which is a data race under the C11 memory model. This PR switches to standard<stdatomic.h>.Changes
_Atomic(T *)arrays:pivsin the sparse echelon forms andnpsin the dense ones. So do the sequential interreduction steps that pass them to the same reducers. This coversla_ff_8.c,la_ff_16.c,la_ff_32.c,la_qq.cand theio.cwrappers.allocate_atomic_pivots()helper indata.c.sparse_AB_CD_*, the known pivots inprobabilistic_sparse_dense_echelon_form_*, SBA). The compare and swap in the sequential kernel computation ofexact_sparse_reduced_echelon_form_sat_ff_32becomes a plain check and store.data.his installed and reached frommsolve.h, which hasextern "C"guards for C++ users. GCC and MSVC reject_Atomicin C++ before C++23, so the four internal 32-bit reducer prototypes whose signatures changed now live only indata.c, which is compiled as part ofgb.c. The installed headers don't use atomics.Testing
-Watomic-implicit-seq-cst, so every access to the atomic arrays is explicit.make checkpasses (68/68).-t 8produced reduced Gröbner bases identical to single-threaded runs. That covers-l 1/2/42/44over 8-, 16- and 31-bit primes (eco11 with characteristic 251, 65521 and 1073741827), plus-l 2/42/44over QQ (kat7).-l 2and-l 44.Notes
./msolve -g 2 -l 1 -f input_files/kat7-qq.mssegfaults even single-threaded. It does so on master as well, so it's unrelated to this PR.flagandbad_primein some echelon form functions are still plainints written by several threads. This PR doesn't touch them.This PR was prepared with the assistance of generative AI (Claude Code). I reviewed the changes and the testing.
🤖 Generated with Claude Code