audit: the #76 claims, measured - 23/28 killed, zero new - #115
Conversation
The five documented claims of #76 were measured rather than assumed. 28 mutants over every behaviour of LibHashNoAlloc: the suite as it stands kills 28/28, and the seven tests the issue proposes, written in their strongest library-calling form and probed alone, kill 23 of those 28 and nothing the suite misses. One of the seven kills nothing because its claim is about the compiler's memory layout and it never calls the library. So no test is added and the scan record is the whole change. The mutants and both probe transcripts are on hash-76-measurement, which is evidence, not a proposal. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V8ViHcKLVk2YoS2joH4HdN
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. WalkthroughThe audit data adds a mutation-test scan for ChangesMutation Scan
Priority: ⬇️ Low Estimated code review effort: 1 (Trivial) | ~2 minutes Change: Other Merge Risk: ⚪ Minimal · up to This PR records mutation-scan results without changing production or test behavior, so it is mergeable. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Linked Issues checkExplanation Issue
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
#114 landed on main after the record was written. It changes only /// lines in LibHashNoAlloc, so no mutated line moved, but "src/ and test/ are byte-identical to the scanned commit" is no longer literally true and the note now says what is. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V8ViHcKLVk2YoS2joH4HdN
Closes #76.
#76 lists five documented claims that no test names: the packed-collision
motivating example, "a
Foois ALWAYS 4 words" with populated members,deterministic boundary lengths, pointer non-determinism, and
bytes1[]as aword list.
#111 filed all five as tests and was closed unmerged: its own measurement showed
the seven tests killed no mutant the suite did not already kill, and three of
them never called the library. #76 was then reopened deliberately, undecided -
that was a verdict on the PR, not on the question of whether this repo wants a
pinning test per documented claim.
This decides it by measurement. #111 measured six mutants against tests that in
three cases could not have killed any. Here the mutant set is every behaviour
src/lib/LibHashNoAlloc.solhas, and the candidates were first rewritten tocall the library wherever the claim permits it, so the result is not an artifact
of how they were written.
A test earns its place here only by killing a mutant the suite does not already
kill, so both sides were probed with
mutation-probeover the same 28 mutants:the read offset and byte count of
hashBytesand of bothhashWordsoverloads,the two scratch writes and the hashed range of
combineHashes, theHASH_NILliteral, and the no-allocation promise of all four.
The pre-existing 77-test suite kills 28/28. The seven tests #76 proposes,
written in their strongest library-calling form and probed alone, kill 23/28
with zero new kills. So no test is added; the diff is the scan record, and
this is the evidence.
Claim by claim
"abc"+"def"and"ab"+"cdef"pack identically and the composition separates every pair (LibHashNoAllocNatSpec)testAbcDefAndAbCdefPackAlikeAndComposeApart,testCompositionSeparatesSplitsOfOneStringtestLeafNodeCollisionpinscombineHashes(a, b) == hashBytes(abi.encodePacked(a, b))over fuzzed bytes,testHashBytespinshashBytes == LibHashSlowat every fuzzed length,testFoldEqualsPackedConcatFoldpins the fold against the packed-concat oracle (crossType / LibHashNoAlloc / HashPatternFold)Foois ALWAYS 4 words", with populated members (README)testFooIsFourWordsWithPopulatedMemberstestFooIsFourWordspinsptr == fmpBeforeand a0x80region;testDynamicMembersArePointerWordsandtestOneWordReferenceMembersArePointerWordspin that each member is one word whatever it holds;testFooListIsWordListmeasuresFoos in a list (MemoryLayout / HashPattern)hashWords([x])being the bare word's hashtestHashBytesAtWordBoundaryLengths,testSingletonWordListHashesAsItsBareWordtestEmptyCollision(length 0 through every entry point),testWordsOneTwoKnownAnswer(2 words / 64 bytes through every entry point against a pinned constant),testKnownAnswersAreBuiltinKeccak,testHashWordsGasOneWordandtestHashWordsGasTwoWords,testHashBytesCostsTheSameAsBuiltinLongat 4096 bytes,testBytesTrueLengthfor 1 vs 2 bytes, thetestHashBytes/testHashWords/testHashWordsUint256fuzzers againstLibHashSlow, andtestFoldSingletonIsNotItemfor the contrast with the seeded foldtestEqualFoosDifferByRegionAndAgreeByCompositiontestStructWithPointersHashesAsNestedNodesruns the README's steps A-E over a fuzzedFooagainstLibFooOracle.hashFoo;testPointerOnlyStructEqualsFoldOfListandtestPointerOnlyStructEqualsFoldOfListAnyLengthwalk the pointers of a fuzzed struct and equate the walk with the fold;testDynamicMembersArePointerWords(HashPattern / crossType / MemoryLayout)bytes1[]named as a word list for hashing (README)testBytes1ArrayHashesAsItsWordstestBytes1ArrayIsWordListpins the layout - length prefix then one left-aligned word per element;testHashWordListpins the word-list Yul againstkeccak256(abi.encodePacked(...));testFixedBytesIsLeftAligned;testHashWordsagainstLibHashSlow(MemoryLayout / HashPattern / LibHashNoAlloc)Union of the proposals: 23 of 28. Every one of those 23 already dies to the
suite, by the tests the last column names and others; the full kill list per
mutant is in the transcript linked below.
The five the proposals never reach
HB10, HW07, HU04 and CH06 replace each function's in-place hash with a copy
through a fresh allocation - the no-allocation promise, which is the whole
reason this library exists. HN01 moves
HASH_NILby one nibble. None of theseven proposals asserts either, because none of the five documented claims is
about allocation or about the nil constant. The suite kills all five:
testHashBytesNoAlloc,testHashWordsNoAlloc,testHashWordsUint256NoAlloc,testCombineHashesNoAlloc, the four...TouchesNoMemory/testCombineHashesTouchesOnlyScratchmemory tests, andtestHashNilwithtestEmptyCollision.The proposals were measured at their strongest, not as filed
Three of the fix texts in #76 route through a private copy of the README's
assembly or through builtins only, which is the objection that closed #111 and
which would guarantee a zero result for an uninteresting reason. Each was
rewritten to call
LibHashNoAllocwherever the claim permits it, and named infull words rather than the single letters #111 was also faulted for, before
being probed. What is measured is the best case for adding them.
Claim 2 is the one that cannot be routed: "a
Foois ALWAYS 4 words" is astatement about the Solidity compiler's memory layout, not about this library.
Its test calls nothing in
src/, so there is no library line whose mutation itcould catch, and it kills nothing. That is a property of the claim, not a
weakness in how the test was written.
The record
The appended
audit/mutation-test-scans.jsonentry is the scan itself: 28behaviours, 28 killed before, 28 after, 0 candidates, 0 confirmed, 77 tests
before and after, nothing filed. It is recorded against
f2a9f5f, the commitscanned;
mainhas since moved to092293eby a README-only commit (#113) anda NatSpec-only one (#114).
test/is byte-identical to the scanned tree andsrc/differs from it only in///lines, so every mutated line is unchangedand the 28/28 stands for
mainas it is now. The suite figures in this PR's QAwere re-run here, not carried over.
The 28 mutant definitions and both probe transcripts are pushed on the
evidence-only branch
hash-76-measurementata3c78d7, on top ofd8fe182which carries the seven candidate tests as they were measured. That branch is
artifact, not a proposal, and is not for merge.
QA
the diff is a JSON scan record that no code reads. The discrimination question
is instead the subject of this PR and was answered by measurement: the seven
candidates were built and probed in isolation (
mutation-probebaseline green,7 passed) and kill 23 of the 28 mutants, none of which the suite misses.
mutation-probeoversrc/lib/LibHashNoAlloc.sol, 28mutants, run twice against the same mutant file. Against the suite as it
stands (baseline green, 77 passed): 28/28 KILLED, 0 survived, 0 no-run, 0
harness errors. Against the seven candidates alone (baseline green, 7 passed):
23/28 KILLED, 5 survived (HB10, HW07, HU04, CH06, HN01), 0 no-run, 0 harness
errors. Both transcripts, with the per-mutant kill lists, are on
hash-76-measurement.LibHashNoAllocNatSpec and theREADME themselves rather than from [RLH-36] [INFO] README/NatSpec claims with no pinning test (residue): the packed-collision motivating example, 'a Foo is ALWAYS 4 words' with populated members, deterministic boundary lengths, pointer non-determinism, bytes1[] as a word list #76's summary of them, and the existing
coverage out of the test files rather than from the audit finding's evidence
lines. The kill/survive verdicts are
mutation-probe's, from the suite's ownpass/fail on a mutated
src/, not from reasoning about what a test ought tocatch. The suite's own oracles throughout are
LibHashSlowand thekeccak256/abi.encode/abi.encodePackedbuiltins, never the libraryunder test.
closed, offers a sixth optional README bullet, and was reopened undecided on
exactly that fork. All five are measured above, each with the test [RLH-36] [INFO] README/NatSpec claims with no pinning test (residue): the packed-collision motivating example, 'a Foo is ALWAYS 4 words' with populated members, deterministic boundary lengths, pointer non-determinism, bytes1[] as a word list #76 named
for it, written in its strongest form and probed against every behaviour the
library has rather than test: pin the five documented claims that had no witness #111's six mutants; the answer is that none of the five
adds coverage, so none is taken and the fork is decided by what the mutants
did. The optional README bullet is not taken either - it would restate the
n= 1 case of an existing "Across types" bullet. What is taken is the scanrecord, so the next audit pass finds this measured rather than re-deriving it.
The branch of test: pin the five documented claims that had no witness #111 is untouched, as its closing comment asked.
Suite: 77 passed / 0 failed / 0 skipped across 9 suites.
forge lint -D warningsclean,forge fmt --checkclean,pre-commit run --all-filespassesand rewrites nothing.
🤖 Generated with Claude Code
https://claude.ai/code/session_01V8ViHcKLVk2YoS2joH4HdN