Skip to content

Restore sbatch --wait in Cromwell's submit-docker - #195

Merged
kennedydane merged 1 commit into
masterfrom
dane/update/cromwell_again
Oct 8, 2026
Merged

kennedydane merged 1 commit into
masterfrom
dane/update/cromwell_again

Conversation

@kennedydane

@kennedydane kennedydane commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Put --wait back on the sbatch call in Cromwell's submit-docker, matching the upstream docs example. It was dropped in 7870b01 ("Submit docker tasks asynchronously"); a user's docker tasks have been failing since and they believe this restores working behaviour.
  • With --wait, sbatch returns the job's exit status, so a job SLURM kills for time or memory fails the task immediately instead of Cromwell waiting forever for an rc file (the config sets no exit-code-timeout-seconds, so check-alive is only run after a server restart).
  • The trade-off is recorded in a NOTE: at the site: Cromwell learns no job id until the job ends, so abort cannot scancel it and each running docker task holds a backend thread. exit-code-timeout-seconds is the non-blocking alternative to evaluate once the failure is understood.

Test plan

  • Rerun the Cromwell install so the config is re-rendered, e.g. ansible-playbook site.yaml -t cromwell,cromwell92
  • Have the affected user rerun their workflow and confirm the docker tasks complete
  • If they still fail, inspect stderr.submit in the failing task's execution directory

🤖 Generated with Claude Code

https://claude.ai/code/session_01L5um94i6XqNc8EfAwJxEeF

A user's docker tasks are failing since the flag was dropped, and they
believe blocking submission fixes it. Put `--wait` back, as in the
upstream docs example, so the task can be rerun against the rendered
config and the cause narrowed down.

With `--wait`, sbatch returns the job's exit status, so a job SLURM
kills for time or memory fails the task at once rather than leaving
Cromwell waiting forever for an rc file. The cost, recorded in a NOTE at
the site, is that Cromwell learns no job id until the job ends: abort
cannot scancel it and each running docker task holds a backend thread.
`exit-code-timeout-seconds` is the non-blocking alternative to evaluate
once the user's failure is understood.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5um94i6XqNc8EfAwJxEeF
Copilot AI balanced review requested due to automatic review settings October 8, 2026 13:35

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@kennedydane
kennedydane requested a review from MikeCTZA October 8, 2026 13:49
@kennedydane
kennedydane merged commit a2ebec8 into master Oct 8, 2026
1 check passed
@kennedydane
kennedydane deleted the dane/update/cromwell_again branch October 8, 2026 13:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants