fix(stellar): make the WASM, ABI and binding artifact checks blocking - #222
Open
Elthelthetallgirl wants to merge 5 commits into
Open
Elthelthetallgirl wants to merge 5 commits into
Elthelthetallgirl wants to merge 5 commits into
Conversation
A plain `cargo build --target wasm32-unknown-unknown --release` in a virtual workspace selects every member. bench, bench-crossover and integration-tests enable soroban-sdk/testutils, which the SDK rejects for wasm with a compile_error, so the whole wasm build failed inside the SDK's transitive deps instead of in the crates that were actually wrong. default-members now selects only the ten deployable members, so the WASM, optimize and size-gate steps in CI can become blocking again. Host-only members remain reachable with --workspace or -p.
wraith-names converted the owner Address into an xdr::ScAddress, but soroban-sdk gates that conversion behind cfg(not(target_family = "wasm")), so the crate only built on the host and every wasm32 build of it failed. The previous documentation blamed soroban-sdk 22.0.11 and told CI to keep the old 9,755-byte baseline forever; the contract simply could not be deployed as written. owner_public_key now decodes the strkey on chain: unpadded base32, version byte 6 << 3, CRC16-XMODEM over the first 33 bytes compared little-endian. Contract, muxed and corrupt addresses still map to NamesError::InvalidSigner, exactly as the ScAddress match did. Also gate `extern crate alloc` behind cfg(test): linking alloc into the cdylib requires a global allocator the contract does not have, which was the second reason the wasm artifact never built.
The artifact steps have been non-blocking for so long that the checked-in output no longer described the contracts. Regenerating with stellar-cli 26.1.0, the new CI pin, shows what was invisible: - 47 of the 72 ABI-exported functions had no client method. wraith-names exposed 5 of 39, stealth-sender 3 of 15 and stealth-registry 2 of 3, so those three clients could not call most of their own contract. The regenerated clients are 711, 339 and 145 lines instead of 147, 123 and 106. - stealth-batch-sender's client had no package.json, tsconfig.json, README.md or .gitignore, so it was not installable outside this repository even though its methods were complete. - The ABI snapshots did not cover the current wraith-names and stealth-sender surfaces. Also drop the timestamp from the generated top-level index.ts. The bindings step now fails the PR on any diff, and a generation date changes on every run, which would make every CI run look like drift. Both generators are idempotent on the committed tree: re-running `pnpm bindings:stellar` and `stellar/abi/update.sh` produces no diff.
Every artifact step in the stellar job had continue-on-error, and the drift checks were guarded by `if: steps.x.outcome == 'success'`, so the job reported success while the WASM build, the size gate, the TypeScript bindings and the ABI snapshots were all failing. The comments blamed a soroban-sdk 22.0.11 and rustc incompatibility; the build failure was actually the workspace selecting testutils-enabled members, and the toolchain bump that was being waited on would not have fixed anything. - Pin stellar-cli 26.1.0. Its bundled wasm-opt validates the four contracts that emit memory.copy/memory.fill, which 22.0.1's rejected outright, so the size gate could never have passed for them. The pinned hash is now the digest upstream publishes for the release asset instead of a hash taken from the archive on first use. - Fix the optimized filename the size gate reads. `contract optimize` writes <name>.optimized.wasm; the step stat'ed <name>_optimized.wasm, and the "no artifacts, skipping" branch turned an empty build into a pass. - Fail when the build produces no artifacts instead of skipping. - Scope both drift checks to stellar/bindings and abi/. A bare `git diff --exit-code` also covered the soroban test_snapshots that `cargo test --workspace` rewrites earlier in the same job, so it would have failed PRs for unrelated snapshot churn.
…grade path SUPPLY_CHAIN.md gains a Stellar artifact checks section: the three things that decide the pins, namely testutils staying out of the default wasm target set, contract code not being allowed to use host-only SDK APIs, and the optimizer having to accept bulk memory. Plus the upgrade path: a CLI bump is an artifact-regeneration PR because the template drives every client's package.json, 28.x renames type_ to type in the ABI output and deprecates the TypeScript bindings generator, and the release container still builds with stellar-cli 22.0.0, which needs a planned redeploy rather than a silent bump. SIZE.md corrected: it claimed wraith_names cannot be compiled for wasm32 with soroban-sdk 22.0.11 and told CI to keep a stale 9,755-byte baseline forever. Neither the claim nor the workaround is right; the contract now builds, and the fresh measurement of all nine payloads is recorded (largest is wraith_names at 57,575 bytes, 51.11% of the 112,640-byte gate). The documented glob <name>_optimized.wasm was also wrong, and the "-p only" workaround is replaced by the default-members explanation. DEPLOYMENT.md moves the CLI requirement to 26.1.0 with the reason. stellar/README.md and scripts/keeper/README.md now say --workspace, because a plain `cargo test` at the workspace root no longer includes the host-only members.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What was actually wrong
The four artifact steps in the
stellarjob werecontinue-on-error: true, andthe two drift checks behind them were gated on
if: steps.*.outcome == 'success',so contract artifacts could drift without failing a PR. The comments in
ci.ymlblamed a "soroban-sdk 22.0.11 vs current rustc" incompatibility that would be
"fixed in 23.x/27.x". That diagnosis is wrong, and the SDK needed no migration.
Five separate defects were hiding behind it:
The wasm build selected host-only crates. A plain
cargo buildin avirtual workspace builds every member.
bench,bench-crossoverandintegration-testsenablesoroban-sdk/testutils, and the SDK guards thatfeature with
compile_error!("'testutils' feature is not supported on 'wasm' target"). All 162 errors were insidesoroban-sdk/src/testutils*, which iswhy they looked like an SDK/compiler problem. They reproduce identically on
the release container's pinned Rust 1.88.0, so no toolchain upgrade fixes
them. The scheduled reproducible-build run on today's
developproves it:job 111776651234
fails at "Generate Attestation" with
'testutils' feature is not supported on 'wasm' targetfollowed bycould not compile soroban-sdk (lib) due to 162 previous errors— the same 162 the workflow comment quotes, from thecontainer's Rust 1.88.0 — while the run still reports
successbecause thatjob is
continue-on-errortoo.default-membersinstellar/Cargo.tomlnowlimits the default target set to the deployable members.
wraith-namescould not build for wasm at all. It converted the ownerAddressinto anxdr::ScAddress, an impl the SDK gates behindcfg(not(target_family = "wasm")), and it linkedallocunconditionally,which leaves the cdylib without a global allocator. The owner's ed25519 key is
now recovered on chain by decoding the account strkey (base32 +
CRC16-XMODEM), and
extern crate allocis#[cfg(test)].The size gate measured a file that never exists.
stellar contract optimizewrites<name>.optimized.wasm; the step stat'ed<name>_optimized.wasm, so the loop aborted on the firststatunder therunner's
bash -e, andcontinue-on-errorhid it. The step alsoexit 0'dwhen there were no artifacts at all.
The pinned optimizer cannot validate this workspace. rustc emits
bulk-memory instructions (
memory.copy/memory.fill) for four of the ninecontracts, and the wasm-opt bundled with stellar-cli 22.0.1 rejects them
during validation. Reproduced on both Rust 1.88.0 and Rust 1.98.1, so
contract optimizecould never have passed forgovernance,stealth_batch_sender,stealth_vaultandwraith_nameson the old pin.stellar-cli 26.1.0 validates all nine, producing output identical to the
optimizer in 28.1.0.
Bindings generation could not run, and could not be diffed. CLI 22.0.1
makes
--contract-ida required argument forcontract bindings typescript,so the wasm-mode invocation in
stellar/scripts/generate-bindings.tsfailedbefore writing anything. Separately, the generator stamped
// Generated on <ISO timestamp>intobindings/typescript/index.tson everyrun, so
git diff --exit-codewould have reported drift on every push evenwith byte-identical inputs.
The version story matters here: the committed clients pin
@stellar/stellar-sdk^14.5.0, carry nonetworksblock, and the staleindex.tstimestamp is2026-05-28T02:32:58.942Z, about 55 minutes afterstellar-cli 26.1.0 was published (
2026-05-28T01:38:04Z). The artifacts in thisrepository were generated with 26.1.0 while CI pinned 22.0.1. Pinning 26.1.0 in
CI is what makes the committed artifacts reproducible by the reviewer's own
toolchain, and it is the version whose release assets publish a sha256 digest.
How completely the old gating hid all of this, from the last
developpush inwhich the
stellarjob ran (run 37058976359, commit7c65132): every stepreported
success, including the failing wasm build, optimize and bindingssteps, while both drift checks read
skipped, becausecontinue-on-errormakesa step's outcome
failureand theif:guards keyed on that outcome. The jobwas green and checked nothing.
Changes
stellar/Cargo.toml: adddefault-membersfor the ten deployable members,with a comment explaining why host-only members must stay out of the default
set.
stellar/wraith-names/src/lib.rs: wasm-compatible owner key recovery(
base32_quintet,crc16_xmodem,account_public_key),#[cfg(test)]extern crate alloc, plus four tests: an SDK round-trip, a known all-zeroaccount vector, a corrupt-checksum rejection and a contract-address
rejection. The checksum is compared little-endian, which is how stellar-strkey
writes it.
stellar/scripts/generate-bindings.ts: deterministic top-levelindex.ts..github/workflows/ci.yml: pinSTELLAR_CLI_VERSION: 26.1.0with theupstream-published Linux x86_64 digest; remove every
continue-on-errorandboth
if: steps.*.outcomeguards from the artifact steps; fix the optimizedfile name; fail instead of skipping when no artifacts exist; use
stellar contract optimize(the step calledsoroban, which only works viathe symlink); scope the two drift checks to their own paths.
stellar/abi/*.json(5 files, +3064/-50) andstellar/bindings/typescript/**(10 files, +1353/-126, including thepackage.json,tsconfig.json,README.mdand.gitignorethatstealth-batch-sendernever had).SUPPLY_CHAIN.mdcovering thethree constraints that decide the pins and the upgrade path, the inventory and
known-gaps rows, corrected false claims in
stellar/SIZE.mdplus a currentper-contract size table,
stellar/DEPLOYMENT.md's CLI guidance, and--workspaceinstellar/README.mdandstellar/scripts/keeper/README.mdbecause a plain
cargo testat the workspace root no longer includes thehost-only members.
The drift checks use
git status --porcelain --untracked-files=allrather thangit diff --exit-codebecause the old unscopedgit diff --exit-codecomparedthe whole repository:
cargo test --workspaceruns earlier in the same job andrewrites soroban
test_snapshots, so the bindings check was destined to fail onunrelated churn while still missing newly generated, uncommitted files.
Measured drift this PR uncovers
The checked-in ABI snapshots were missing 47 of the 72 exported functions:
stealth_announcerstealth_registryremove_keysstealth_senderwithdraw_manystealth_batch_senderWraithMetricEventUDTwraith_namesThe committed TypeScript clients had the same holes, since they are generated
from those snapshots:
wraith-namesbound 5 of 39 functions in 146 lines(710 now),
stealth-sender3 of 15 in 122 (338 now),stealth-registry2 of 3in 105 (144 now).
stealth-batch-senderhad all 14 methods but nopackage.json, so it was not installable outside the repository at all. That isexactly the invisible artifact drift the issue describes, and from now on it
fails the PR.
Verification
Run locally with CI's toolchain (Rust 1.98.1 =
env.RUST_TOOLCHAIN,stellar-cli 26.1.0, pnpm 10.28.2):
cargo build --target wasm32-unknown-unknown --releaseproduces the ninecontract artifacts. On the old tree this command failed with 162 SDK errors.
stellar contract optimizesucceeds for all nine (it failed for four on22.0.1). Optimized payloads:
wraith_names57,575 (51.11% of the 112,640budget, the largest),
stealth_sender24,986,stealth_batch_sender21,902,governance18,506,stealth_splitter16,899,stealth_vault16,585,stealth_announcer6,575,stealth_registry5,973,wraith_asset_policy4,559.
pnpm bindings:stellarandstellar/abi/update.shproduces no diff, so thetwo new gates pass as committed.
committed state: editing a tracked binding fails, deleting a tracked binding
fails, adding an untracked file under
stellar/bindingsfails (this is thecase
--untracked-files=allexists for), and the same three cases againstabi/fail. A clean tree passes both.cargo test --workspace --no-fail-fast: 393 passed, 0 failed, 12 ignored,including
wraith_names(30 passed, 5 ignored) with the four new strkeytests.
cargo fmt --all --checkclean.pnpm test:stellar-bindingspasses against the regenerated client. Oneenvironment note: if
stellar/node_modulesexists locally it shadows the rootinstall and resolves
@stellar/stellar-sdk13.3.0 instead of the root's15.1.0, and 13.3.0's
RpcServerrejects the test'shttp://localhost:8000URL. CI only installs the root project, so it resolves 15.1.0 and passes; the
failure is a local-layout artifact, reproducible on the old bindings too.
STELLAR_CLI_SHA256is the digest upstream publishes forstellar-cli-26.1.0-x86_64-unknown-linux-gnu.tar.gz(
e18d5a76…5c0c); it was re-computed from the downloaded archive, and thearchive still contains a single top-level
stellarbinary, so the existingtar xzf/sudo mvlines are unchanged.This PR comes from a fork, so its workflows will sit at
action_requireduntila maintainer approves a run; every gate above was therefore measured locally
rather than asserted from a green checkmark. Expect the
stellarjob to go from"green, checking nothing" to "green, checking these five things" on the first
approved run.
Out of scope, on purpose
stellar/build/build.shoptimizes every contract, so with the container'sstellar-cli 22.0.0 it aborts on the same four bulk-memory artifacts. Bumping
that pin changes optimized bytes, therefore the hashes
stellar/build/verify.jscompares against deployed contracts, therefore it is a planned redeploy under
the re-audit rules in
SUPPLY_CHAIN.md. Recorded as a known gap instead.old comments waited for was never the fix.
stellar-verification.ymlstays non-blocking. Itsbuild-and-verifyjob isthe Docker reproducible build plus
verify.js, and its job-levelcontinue-on-erroris a documented decision, not one of the artifact steps inissue [Wave 9] Re-enable blocking Stellar WASM, ABI, and binding checks #187.
default-membersclears the "pinned combination stops compiling"half of that job's failure (the testutils errors above), but a 22.x optimizer
then rejects the same four bulk-memory artifacts — reproduced locally with the
22.0.1 release binary against artifacts built by the container's own Rust
1.88.0. It goes from red to red, so flipping it stays the redeploy's job rather
than this PR's.
type_totypeincontract info interface --output json-formatted, which would rewrite everyABI snapshot in this PR, and marks
contract bindings typescriptdeprecated infavour of the JavaScript SDK generator. Its optimizer output is byte-identical
to 26.1.0, so nothing is lost by holding the pin at 26.1.0.
.github/is owned by@truthixify, so the workflow pin bump in this PR needsthat review per
SUPPLY_CHAIN.md.closes #187