Skip to content

Replace __sync builtins with C11 atomics in linear algebra - #377

Draft
wegank wants to merge 1 commit into
algebraic-solving:masterfrom
wegank:c11-atomics
Draft

wegank wants to merge 1 commit into
algebraic-solving:masterfrom
wegank:c11-atomics

Conversation

@wegank

@wegank wegank commented Oct 1, 2026

Copy link
Copy Markdown
Member

The parallel linear algebra in neogb publishes new pivot rows with __sync_bool_compare_and_swap. That GCC builtin isn't available with MSVC. Also, other threads read the same pivot arrays with plain loads, which is a data race under the C11 memory model. This PR switches to standard <stdatomic.h>.

Changes

  • Atomic pivot arrays. The arrays that threads publish into concurrently become _Atomic(T *) arrays: pivs in the sparse echelon forms and nps in the dense ones. So do the sequential interreduction steps that pass them to the same reducers. This covers la_ff_8.c, la_ff_16.c, la_ff_32.c, la_qq.c and the io.c wrappers.
  • Memory ordering.
    • New pivots are published by a compare and swap with release ordering, after the row is fully reduced and normalized.
    • The reducers load entries with acquire ordering, so a published row and its coefficients are fully visible to them. Each entry is now loaded once instead of twice.
    • Sequential code after the parallel regions uses relaxed accesses.
    • The rule is documented once, above the new allocate_atomic_pivots() helper in data.c.
  • Plain arrays. Arrays that are only read during parallel regions stay plain (sparse_AB_CD_*, the known pivots in probabilistic_sparse_dense_echelon_form_*, SBA). The compare and swap in the sequential kernel computation of exact_sparse_reduced_echelon_form_sat_ff_32 becomes a plain check and store.
  • Public headers. data.h is installed and reached from msolve.h, which has extern "C" guards for C++ users. GCC and MSVC reject _Atomic in C++ before C++23, so the four internal 32-bit reducer prototypes whose signatures changed now live only in data.c, which is compiled as part of gb.c. The installed headers don't use atomics.
  • configure. It now checks for C11 atomics and fails with a clear message if they're missing.
  • Small cleanup. The unrolled pivot-counting loops in the dense linear algebra become a plain loop, which avoids five atomic loads per iteration.

Testing

  • neogb compiles without warnings with clang's -Watomic-implicit-seq-cst, so every access to the atomic arrays is explicit.
  • make check passes (68/68).
  • Stress test: 120 runs with -t 8 produced reduced Gröbner bases identical to single-threaded runs. That covers -l 1/2/42/44 over 8-, 16- and 31-bit primes (eco11 with characteristic 251, 65521 and 1073741827), plus -l 2/42/44 over QQ (kat7).
  • Runtime on arm64 (Apple Silicon) is within ±2% of master for eco11 with 1 and 8 threads, -l 2 and -l 44.

Notes

  • ./msolve -g 2 -l 1 -f input_files/kat7-qq.ms segfaults even single-threaded. It does so on master as well, so it's unrelated to this PR.
  • flag and bad_prime in some echelon form functions are still plain ints written by several threads. This PR doesn't touch them.
  • Full MSVC support needs more work beyond this PR (C11 atomics flags, OpenMP 2.0 loop variables, other GCC-isms).

This PR was prepared with the assistance of generative AI (Claude Code). I reviewed the changes and the testing.

🤖 Generated with Claude Code

The parallel linear algebra publishes new pivot rows with
__sync_bool_compare_and_swap, a GCC builtin that is not available with
MSVC. Moreover, the other threads read these pivot arrays with plain
loads, which is a data race under the C11 memory model.

The pivot arrays that threads add pivots to concurrently are now arrays
of _Atomic pointers. A new pivot is published with a compare and swap
using release ordering, the reducers load entries with acquire
ordering, so a published row and its coefficients are fully visible to
them. Sequential code after the parallel regions uses relaxed accesses.
Arrays only read during parallel regions stay plain, and the compare and
swap in the sequential kernel computation of the saturation is replaced
by a plain check and store.

The prototypes of the internal 32 bit reducers taking atomic arguments
move from the installed data.h to data.c, so that the public headers
stay usable from C++. configure now checks for C11 atomics.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant