Skip to content

fix(stellar): make the WASM, ABI and binding artifact checks blocking - #222

Open
Elthelthetallgirl wants to merge 5 commits into
wraith-protocol:developfrom
Elthelthetallgirl:fix/stellar-blocking-artifact-checks
Open

Elthelthetallgirl wants to merge 5 commits into
wraith-protocol:developfrom
Elthelthetallgirl:fix/stellar-blocking-artifact-checks

Conversation

@Elthelthetallgirl

Copy link
Copy Markdown

What was actually wrong

The four artifact steps in the stellar job were continue-on-error: true, and
the two drift checks behind them were gated on if: steps.*.outcome == 'success',
so contract artifacts could drift without failing a PR. The comments in ci.yml
blamed a "soroban-sdk 22.0.11 vs current rustc" incompatibility that would be
"fixed in 23.x/27.x". That diagnosis is wrong, and the SDK needed no migration.
Five separate defects were hiding behind it:

  1. The wasm build selected host-only crates. A plain cargo build in a
    virtual workspace builds every member. bench, bench-crossover and
    integration-tests enable soroban-sdk/testutils, and the SDK guards that
    feature with compile_error!("'testutils' feature is not supported on 'wasm' target"). All 162 errors were inside soroban-sdk/src/testutils*, which is
    why they looked like an SDK/compiler problem. They reproduce identically on
    the release container's pinned Rust 1.88.0, so no toolchain upgrade fixes
    them. The scheduled reproducible-build run on today's develop proves it:
    job 111776651234
    fails at "Generate Attestation" with 'testutils' feature is not supported on 'wasm' target followed by could not compile soroban-sdk (lib) due to 162 previous errors — the same 162 the workflow comment quotes, from the
    container's Rust 1.88.0 — while the run still reports success because that
    job is continue-on-error too. default-members in stellar/Cargo.toml now
    limits the default target set to the deployable members.

  2. wraith-names could not build for wasm at all. It converted the owner
    Address into an xdr::ScAddress, an impl the SDK gates behind
    cfg(not(target_family = "wasm")), and it linked alloc unconditionally,
    which leaves the cdylib without a global allocator. The owner's ed25519 key is
    now recovered on chain by decoding the account strkey (base32 +
    CRC16-XMODEM), and extern crate alloc is #[cfg(test)].

  3. The size gate measured a file that never exists. stellar contract optimize writes <name>.optimized.wasm; the step stat'ed
    <name>_optimized.wasm, so the loop aborted on the first stat under the
    runner's bash -e, and continue-on-error hid it. The step also exit 0'd
    when there were no artifacts at all.

  4. The pinned optimizer cannot validate this workspace. rustc emits
    bulk-memory instructions (memory.copy/memory.fill) for four of the nine
    contracts, and the wasm-opt bundled with stellar-cli 22.0.1 rejects them
    during validation. Reproduced on both Rust 1.88.0 and Rust 1.98.1, so
    contract optimize could never have passed for governance,
    stealth_batch_sender, stealth_vault and wraith_names on the old pin.
    stellar-cli 26.1.0 validates all nine, producing output identical to the
    optimizer in 28.1.0.

  5. Bindings generation could not run, and could not be diffed. CLI 22.0.1
    makes --contract-id a required argument for contract bindings typescript,
    so the wasm-mode invocation in stellar/scripts/generate-bindings.ts failed
    before writing anything. Separately, the generator stamped
    // Generated on <ISO timestamp> into bindings/typescript/index.ts on every
    run, so git diff --exit-code would have reported drift on every push even
    with byte-identical inputs.

The version story matters here: the committed clients pin
@stellar/stellar-sdk ^14.5.0, carry no networks block, and the stale
index.ts timestamp is 2026-05-28T02:32:58.942Z, about 55 minutes after
stellar-cli 26.1.0 was published (2026-05-28T01:38:04Z). The artifacts in this
repository were generated with 26.1.0 while CI pinned 22.0.1. Pinning 26.1.0 in
CI is what makes the committed artifacts reproducible by the reviewer's own
toolchain, and it is the version whose release assets publish a sha256 digest.

How completely the old gating hid all of this, from the last develop push in
which the stellar job ran (run 37058976359, commit 7c65132): every step
reported success, including the failing wasm build, optimize and bindings
steps, while both drift checks read skipped, because continue-on-error makes
a step's outcome failure and the if: guards keyed on that outcome. The job
was green and checked nothing.

Changes

  • stellar/Cargo.toml: add default-members for the ten deployable members,
    with a comment explaining why host-only members must stay out of the default
    set.
  • stellar/wraith-names/src/lib.rs: wasm-compatible owner key recovery
    (base32_quintet, crc16_xmodem, account_public_key), #[cfg(test)]
    extern crate alloc, plus four tests: an SDK round-trip, a known all-zero
    account vector, a corrupt-checksum rejection and a contract-address
    rejection. The checksum is compared little-endian, which is how stellar-strkey
    writes it.
  • stellar/scripts/generate-bindings.ts: deterministic top-level index.ts.
  • .github/workflows/ci.yml: pin STELLAR_CLI_VERSION: 26.1.0 with the
    upstream-published Linux x86_64 digest; remove every continue-on-error and
    both if: steps.*.outcome guards from the artifact steps; fix the optimized
    file name; fail instead of skipping when no artifacts exist; use
    stellar contract optimize (the step called soroban, which only works via
    the symlink); scope the two drift checks to their own paths.
  • Regenerated artifacts: stellar/abi/*.json (5 files, +3064/-50) and
    stellar/bindings/typescript/** (10 files, +1353/-126, including the
    package.json, tsconfig.json, README.md and .gitignore that
    stealth-batch-sender never had).
  • Docs: a new "Stellar artifact checks" section in SUPPLY_CHAIN.md covering the
    three constraints that decide the pins and the upgrade path, the inventory and
    known-gaps rows, corrected false claims in stellar/SIZE.md plus a current
    per-contract size table, stellar/DEPLOYMENT.md's CLI guidance, and
    --workspace in stellar/README.md and stellar/scripts/keeper/README.md
    because a plain cargo test at the workspace root no longer includes the
    host-only members.

The drift checks use git status --porcelain --untracked-files=all rather than
git diff --exit-code because the old unscoped git diff --exit-code compared
the whole repository: cargo test --workspace runs earlier in the same job and
rewrites soroban test_snapshots, so the bindings check was destined to fail on
unrelated churn while still missing newly generated, uncommitted files.

Measured drift this PR uncovers

The checked-in ABI snapshots were missing 47 of the 72 exported functions:

Snapshot Committed exports Actual exports Previously missing
stealth_announcer 1 1 doc text only
stealth_registry 2 3 remove_keys
stealth_sender 3 15 pause/signers/multisig/rotation/withdraw_many
stealth_batch_sender 14 14 WraithMetricEvent UDT
wraith_names 5 39 auction, bulk, multisig and rotation entry points

The committed TypeScript clients had the same holes, since they are generated
from those snapshots: wraith-names bound 5 of 39 functions in 146 lines
(710 now), stealth-sender 3 of 15 in 122 (338 now), stealth-registry 2 of 3
in 105 (144 now). stealth-batch-sender had all 14 methods but no
package.json, so it was not installable outside the repository at all. That is
exactly the invisible artifact drift the issue describes, and from now on it
fails the PR.

Verification

Run locally with CI's toolchain (Rust 1.98.1 = env.RUST_TOOLCHAIN,
stellar-cli 26.1.0, pnpm 10.28.2):

  • cargo build --target wasm32-unknown-unknown --release produces the nine
    contract artifacts. On the old tree this command failed with 162 SDK errors.
  • stellar contract optimize succeeds for all nine (it failed for four on
    22.0.1). Optimized payloads: wraith_names 57,575 (51.11% of the 112,640
    budget, the largest), stealth_sender 24,986, stealth_batch_sender 21,902,
    governance 18,506, stealth_splitter 16,899, stealth_vault 16,585,
    stealth_announcer 6,575, stealth_registry 5,973, wraith_asset_policy
    4,559.
  • Both generators are idempotent on the committed tree: re-running
    pnpm bindings:stellar and stellar/abi/update.sh produces no diff, so the
    two new gates pass as committed.
  • Both gates were then negative-tested with the exact CI command, on the
    committed state: editing a tracked binding fails, deleting a tracked binding
    fails, adding an untracked file under stellar/bindings fails (this is the
    case --untracked-files=all exists for), and the same three cases against
    abi/ fail. A clean tree passes both.
  • cargo test --workspace --no-fail-fast: 393 passed, 0 failed, 12 ignored,
    including wraith_names (30 passed, 5 ignored) with the four new strkey
    tests. cargo fmt --all --check clean.
  • pnpm test:stellar-bindings passes against the regenerated client. One
    environment note: if stellar/node_modules exists locally it shadows the root
    install and resolves @stellar/stellar-sdk 13.3.0 instead of the root's
    15.1.0, and 13.3.0's RpcServer rejects the test's http://localhost:8000
    URL. CI only installs the root project, so it resolves 15.1.0 and passes; the
    failure is a local-layout artifact, reproducible on the old bindings too.
  • STELLAR_CLI_SHA256 is the digest upstream publishes for
    stellar-cli-26.1.0-x86_64-unknown-linux-gnu.tar.gz
    (e18d5a76…5c0c); it was re-computed from the downloaded archive, and the
    archive still contains a single top-level stellar binary, so the existing
    tar xzf/sudo mv lines are unchanged.

This PR comes from a fork, so its workflows will sit at action_required until
a maintainer approves a run; every gate above was therefore measured locally
rather than asserted from a green checkmark. Expect the stellar job to go from
"green, checking nothing" to "green, checking these five things" on the first
approved run.

Out of scope, on purpose

  • The release container is still broken and stays that way here.
    stellar/build/build.sh optimizes every contract, so with the container's
    stellar-cli 22.0.0 it aborts on the same four bulk-memory artifacts. Bumping
    that pin changes optimized bytes, therefore the hashes stellar/build/verify.js
    compares against deployed contracts, therefore it is a planned redeploy under
    the re-audit rules in SUPPLY_CHAIN.md. Recorded as a known gap instead.
  • No soroban-sdk migration. 22.0.11 builds and tests fine; the "SDK bump" the
    old comments waited for was never the fix.
  • stellar-verification.yml stays non-blocking. Its build-and-verify job is
    the Docker reproducible build plus verify.js, and its job-level
    continue-on-error is a documented decision, not one of the artifact steps in
    issue [Wave 9] Re-enable blocking Stellar WASM, ABI, and binding checks #187. default-members clears the "pinned combination stops compiling"
    half of that job's failure (the testutils errors above), but a 22.x optimizer
    then rejects the same four bulk-memory artifacts — reproduced locally with the
    22.0.1 release binary against artifacts built by the container's own Rust
    1.88.0. It goes from red to red, so flipping it stays the redeploy's job rather
    than this PR's.
  • Not 28.x. 28.1.0 renames type_ to type in
    contract info interface --output json-formatted, which would rewrite every
    ABI snapshot in this PR, and marks contract bindings typescript deprecated in
    favour of the JavaScript SDK generator. Its optimizer output is byte-identical
    to 26.1.0, so nothing is lost by holding the pin at 26.1.0.
  • .github/ is owned by @truthixify, so the workflow pin bump in this PR needs
    that review per SUPPLY_CHAIN.md.

closes #187

A plain `cargo build --target wasm32-unknown-unknown --release` in a virtual
workspace selects every member. bench, bench-crossover and integration-tests
enable soroban-sdk/testutils, which the SDK rejects for wasm with a
compile_error, so the whole wasm build failed inside the SDK's transitive
deps instead of in the crates that were actually wrong.

default-members now selects only the ten deployable members, so the WASM,
optimize and size-gate steps in CI can become blocking again. Host-only
members remain reachable with --workspace or -p.
wraith-names converted the owner Address into an xdr::ScAddress, but soroban-sdk
gates that conversion behind cfg(not(target_family = "wasm")), so the crate only
built on the host and every wasm32 build of it failed. The previous
documentation blamed soroban-sdk 22.0.11 and told CI to keep the old 9,755-byte
baseline forever; the contract simply could not be deployed as written.

owner_public_key now decodes the strkey on chain: unpadded base32, version byte
6 << 3, CRC16-XMODEM over the first 33 bytes compared little-endian. Contract,
muxed and corrupt addresses still map to NamesError::InvalidSigner, exactly as
the ScAddress match did.

Also gate `extern crate alloc` behind cfg(test): linking alloc into the cdylib
requires a global allocator the contract does not have, which was the second
reason the wasm artifact never built.
The artifact steps have been non-blocking for so long that the checked-in
output no longer described the contracts. Regenerating with stellar-cli
26.1.0, the new CI pin, shows what was invisible:

- 47 of the 72 ABI-exported functions had no client method. wraith-names
  exposed 5 of 39, stealth-sender 3 of 15 and stealth-registry 2 of 3, so
  those three clients could not call most of their own contract. The
  regenerated clients are 711, 339 and 145 lines instead of 147, 123 and 106.
- stealth-batch-sender's client had no package.json, tsconfig.json,
  README.md or .gitignore, so it was not installable outside this repository
  even though its methods were complete.
- The ABI snapshots did not cover the current wraith-names and stealth-sender
  surfaces.

Also drop the timestamp from the generated top-level index.ts. The bindings
step now fails the PR on any diff, and a generation date changes on every
run, which would make every CI run look like drift.

Both generators are idempotent on the committed tree: re-running
`pnpm bindings:stellar` and `stellar/abi/update.sh` produces no diff.
Every artifact step in the stellar job had continue-on-error, and the drift
checks were guarded by `if: steps.x.outcome == 'success'`, so the job reported
success while the WASM build, the size gate, the TypeScript bindings and the
ABI snapshots were all failing. The comments blamed a soroban-sdk 22.0.11 and
rustc incompatibility; the build failure was actually the workspace selecting
testutils-enabled members, and the toolchain bump that was being waited on
would not have fixed anything.

- Pin stellar-cli 26.1.0. Its bundled wasm-opt validates the four contracts
  that emit memory.copy/memory.fill, which 22.0.1's rejected outright, so the
  size gate could never have passed for them. The pinned hash is now the
  digest upstream publishes for the release asset instead of a hash taken from
  the archive on first use.
- Fix the optimized filename the size gate reads. `contract optimize` writes
  <name>.optimized.wasm; the step stat'ed <name>_optimized.wasm, and the
  "no artifacts, skipping" branch turned an empty build into a pass.
- Fail when the build produces no artifacts instead of skipping.
- Scope both drift checks to stellar/bindings and abi/. A bare
  `git diff --exit-code` also covered the soroban test_snapshots that
  `cargo test --workspace` rewrites earlier in the same job, so it would have
  failed PRs for unrelated snapshot churn.
…grade path

SUPPLY_CHAIN.md gains a Stellar artifact checks section: the three things that
decide the pins, namely testutils staying out of the default wasm target set,
contract code not being allowed to use host-only SDK APIs, and the optimizer
having to accept bulk memory. Plus the upgrade path: a CLI bump is an
artifact-regeneration PR because the template drives every client's
package.json, 28.x renames type_ to type in the ABI output and deprecates the
TypeScript bindings generator, and the release container still builds with
stellar-cli 22.0.0, which needs a planned redeploy rather than a silent bump.

SIZE.md corrected: it claimed wraith_names cannot be compiled for wasm32 with
soroban-sdk 22.0.11 and told CI to keep a stale 9,755-byte baseline forever.
Neither the claim nor the workaround is right; the contract now builds, and the
fresh measurement of all nine payloads is recorded (largest is wraith_names at
57,575 bytes, 51.11% of the 112,640-byte gate). The documented glob
<name>_optimized.wasm was also wrong, and the "-p only" workaround is replaced
by the default-members explanation.

DEPLOYMENT.md moves the CLI requirement to 26.1.0 with the reason.
stellar/README.md and scripts/keeper/README.md now say --workspace, because a
plain `cargo test` at the workspace root no longer includes the host-only
members.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Wave 9] Re-enable blocking Stellar WASM, ABI, and binding checks

1 participant