Skip to content

feat: checkpoint keeper watch progress so a restart cannot resettle (#380) - #447

Merged
karagozemin merged 3 commits into
Sub-Rosa-Issue:mainfrom
adeboladee:feat/issue-380-keeper-watch-checkpoint
Oct 1, 2026
Merged

karagozemin merged 3 commits into
Sub-Rosa-Issue:mainfrom
adeboladee:feat/issue-380-keeper-watch-checkpoint

Conversation

@adeboladee

Copy link
Copy Markdown
Contributor

Summary

The keeper's per-round watch progress is now durable. The in-memory settlement guard dies with the process, so a keeper that crashed after broadcasting clear/settle had no record of the in-flight transaction and the next process rebroadcasted it. This adds a local watch checkpoint and makes the watch loop, the one-shot keeper, and the dry-run planner share it.

  • services/keeper/src/checkpoint.ts — schema, pure planners (planCheckpointStep / planCheckpointRollback), file store (KeeperCheckpointStore, KEEPER_CHECKPOINT_PATH, default .keeper-checkpoint.json), and a dryRun mode that tracks progress in memory and never writes.
  • services/keeper/src/keeper.ts — KeeperDeps.checkpoint: skip the steps the cursor records as complete and record each step on success (including when the contract confirms it idempotently, e.g. AlreadySettled).
  • services/keeper/src/watch-loop.ts — resumeCheckpoint() startup gate, wired into runWatchLoop, plus checkpoint / verifyTransaction params.
  • services/keeper/src/dry-run.ts + run.ts — the dry-run summary now carries the checkpoint a live pass would write, and warns when a live keeper would refuse to start.
  • watch.ts / serve.ts — construct the checkpoint; a binding mismatch aborts the process.

Checkpoint schema

{
  "version": 1,
  "network": "Test SDF Network ; September 2015",
  "contractId": "C...",
  "rounds": {
    "1": {
      "roundId": "1",
      "completedSteps": ["open-reveal", "reveal", "clear", "settle"],
      "lastCompletedStep": "settle",
      "lastTransactionHash": "0x…",
      "stepHashes": { "settle": "0x…" },
      "updatedAt": "2026-09-30T00:00:00.000Z"
    }
  }
}

Steps tracked: open-reveal, reveal, clear, settle, void. The file is bound to one network + contractId pair. The preview the dry-run prints is produced by the same planner the store uses, so it is byte-for-byte what a live pass would persist (asserted by a test).

Validation checks on startup

  1. Binding — the constructor throws KeeperCheckpointMismatchError when the stored network or contractId differs from the process config (also true for an unsupported version); a corrupt/unparseable file is backed up to *.corrupted.<ts> and the keeper starts from a fresh cursor instead of guessing.
  2. Hash verification — every recorded stepHashes entry is re-checked. failed/missing rolls the step back so it is retried; confirmed is trusted even when the RPC replica still reports the pre-step status. A failing hash lookup keeps the cursor (an unreachable RPC must not discard durable progress).
  3. Chain reconciliation — entries with no hash to verify are checked against the on-chain status (STEP_SATISFIED_BY_STATUS). If the chain cannot prove a step happened, the entry is dropped rather than stranding the round.

open-reveal is recorded for auditability but is not skip-eligible (CHECKPOINT_SKIP_STEPS): whether the reveal window is open is already authoritative on-chain, so trusting the cursor there could strand an Open round. The reveal step is only advanced when every bidder ended up revealed, keeping undecryptable seals retryable.

Dry-run behavior

KEEPER_DRY_RUN=true npm run start prints the checkpoint it would write — path, network, contractId, proposedStep, mismatch, and the full proposedFile — inside the summary. It submits nothing (transactionsSubmitted: 0), writes nothing (checkpoint.filesWritten: 0), and KeeperCheckpointStore({ dryRun: true }) never touches the filesystem. A checkpoint that would block a live keeper is reported as checkpoint.mismatch ("network" or "contractId").

Acceptance criteria → tests

Acceptance criterion Test
Restart after a confirmed settle does not submit settle again checkpoint-restart.test.ts → a confirmed settle is not submitted again by a lagging replica (also asserted end-to-end through runWatchLoop in a resumed watch loop does not settle a round the cursor already settled)
A checkpoint for a different contract id refuses to start checkpoint.test.ts → refuses to start when the checkpoint contract id does not match (and the network variant)
Dry-run writes nothing and submits nothing dry-run.test.ts → writes nothing to the checkpoint path and submits nothing, plus checkpoint.test.ts → tracks progress in memory and never writes to disk
Tests cover crash-before-checkpoint and crash-after-checkpoint with a fake clock checkpoint-restart.test.ts describes restart after the checkpoint was written (4 tests) and restart before the checkpoint was written (4 tests), all on createFakeTime with an in-memory chain; the crash-before flow walks settle → idempotent skip → recorded → never rebroadcast
Store round id, last completed step, transaction hash, and network checkpoint.test.ts → persists round id, last completed step, transaction hash, and network
On startup, skip steps marked complete after confirming hash status keeps steps whose transaction hash is confirmed / rolls back a step whose transaction failed / a lost cursor is rebuilt from the chain, not trusted blindly

Test execution evidence

$ cd services/keeper && npm test
ℹ tests 116
ℹ suites 27
ℹ pass 116
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 6745.900642

(73 pre-existing tests + 43 new: 27 in checkpoint.test.ts, 12 in checkpoint-restart.test.ts, 4 added to dry-run.test.ts.)

Also green:

$ pnpm keeper:typecheck          # tsc --noEmit, no errors
$ pnpm coverage:test             # services/keeper 86.08% lines (2548/2960); gate 70% → passed
$ pnpm logging:check             # No direct console calls
$ pnpm errors:normalize:check    # No ad hoc caught-value stringification
$ pnpm errors:check              # types.rs and ERRORS.md agree
$ pnpm snapshot:check            # 51/51 categories present
$ pnpm docs:check / docs:check-links / threat-model:check

pnpm time:guard still fails, but only on pre-existing violations in apps/web/src/hooks/useDashboardData.test.ts and packages/sdk/src/status-client.ts — verified identical on main (stashed) and untouched by this PR.

All tests are offline: no RPC, no Drand, no signing material. The live testnet scripts (keeper:e2e, lifecycle:e2e) were not run — they require a funded testnet account and a live round.

Closes #380

…ence

The in-memory settlement guard dies with the process: a keeper that crashes
after broadcasting clear/settle has no record of the in-flight transaction and
the next process rebroadcasts it. Persist the per-round watch cursor locally so
a restart cannot resettle.

- checkpoint.ts: durable cursor (round id, completed steps, last completed step,
  transaction hash, network, contract id) with pure planners, a file store, and
  dry-run (in-memory, never writes) mode.
- Startup validation: refuse to start when the checkpoint network or contract id
  differs from the process config; re-verify recorded transaction hashes and
  reconcile hashless entries against the on-chain status.
- keeper.ts: skip steps the cursor records as complete (reveal, clear, settle,
  void) and record each completed step, including when the contract confirms the
  step idempotently. open-reveal stays on-chain-driven so a round cannot strand.
- dry-run.ts/run.ts: print the checkpoint a live pass would write, plus any
  binding conflict, without submitting a transaction or touching disk.
- watch/serve: construct the checkpoint (mismatch aborts the process).

Closes Sub-Rosa-Issue#380
@drips-wave

drips-wave Bot commented Sep 30, 2026

Copy link
Copy Markdown

@adeboladee Great news! 🎉 Based on an automated assessment of this PR, the linked Wave issue(s) no longer count against your application limits.

You can now already apply to more issues while waiting for a review of this PR. Keep up the great work! 🚀

Learn more about application limits

@karagozemin
karagozemin merged commit 8de1d66 into Sub-Rosa-Issue:main Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Checkpoint keeper watch progress so a restart cannot resettle

2 participants