Skip to content

slurm: make run-parts.sh exclusive detection work with custom prefix and recent Slurm - #1391

Open
100milliongold wants to merge 1 commit into
NVIDIA:masterfrom
xiilab:fix/run-parts-exclusive-detection
Open

slurm: make run-parts.sh exclusive detection work with custom prefix and recent Slurm#1391
100milliongold wants to merge 1 commit into
NVIDIA:masterfrom
xiilab:fix/run-parts-exclusive-detection

Conversation

@100milliongold

Copy link
Copy Markdown
Contributor

Problem

The exclusive-job detection in roles/slurm/templates/etc/slurm/shared/bin/run-parts.sh fails in two independent ways:

  1. PATH. It calls scontrol/squeue through PATH. slurmd's environment does not include a custom slurm_install_prefix, so both commands return nothing, numcpus_sys and numcpus_job are both empty, "" == "" is true, and every job runs the *-exclusive-* scripts. On a shared node that resets power limits and application clocks on all GPUs and drops page caches for each job (prolog took 6–8 s; srun: Prolog hung on node).

  2. Parsing. grep -Eio "TRES=cpu=[0-9]+" matches both ReqTRES= and AllocTRES= lines on recent Slurm, so numcpus_job becomes a multi-line value ('1\n56' in a bash -x trace) that never compares equal. With scontrol on PATH, exclusive jobs are therefore never detected.

Reproduced on DGX OS 7.5.0, Slurm 26.05.1, slurm_install_prefix: /raid/slurm/usr/local.

Fix

  • Call {{ slurm_install_prefix }}/bin/squeue by absolute path (the script is deployed with the template module, so the variable is available).
  • Read allocated CPUs and node count with squeue -o %C / -o %D instead of parsing scontrol show job.
  • Guard against an empty result so a lookup failure means "not exclusive".

Verification

On the system above, before the fix every srun --gres=gpu:1 job logged Running .../50-exclusive-gpu; after symlinking the binaries onto PATH (which exercises failure 2) a bash -x run showed numcpus_job='1\n56'. The patched logic yields numcpus_job=56, numcpus_sys=256, exclusive=0 for that job.

…and recent Slurm

The exclusive-job check in run-parts.sh had two independent failures:

1. It called scontrol/squeue through PATH. slurmd's environment does not
   include a custom slurm_install_prefix, so both commands produced empty
   output, numcpus_sys and numcpus_job were both "", the comparison was
   true, and every job ran the *-exclusive-* prolog/epilog scripts. On a
   shared node this reset power limits and clocks on all GPUs and dropped
   page caches for every job.

2. It parsed "scontrol show job" with grep -Eio "TRES=cpu=[0-9]+". On
   recent Slurm the output has both ReqTRES= and AllocTRES= lines, so the
   pattern matched twice and numcpus_job became a multi-line value that
   never compared equal. With scontrol on PATH, exclusive jobs were
   therefore never detected.

Use {{ slurm_install_prefix }}/bin/squeue by absolute path (the file is
already deployed via the template module) and read allocated CPUs and node
count with -o %C / -o %D instead of parsing scontrol. Guard against an
empty result so a lookup failure means "not exclusive" rather than
"exclusive".

Observed on DGX OS 7.5.0, Slurm 26.05.1, slurm_install_prefix=/raid/slurm/usr/local:
- before the binaries were symlinked into /usr/local/bin, every srun
  --gres=gpu:1 job logged "Running .../50-exclusive-gpu" and prolog took
  6-8 s (srun: Prolog hung on node);
- after symlinking, a bash -x run of the script showed numcpus_job='1<nl>56'.

Signed-off-by: Jea-Eok-Kim <je.kim@xiilab.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants