Skip to content

Run the PR test on OSDC runners - #2711

Closed
huydhn wants to merge 3 commits into
mainfrom
osdc/pr-test-arc
Closed

huydhn wants to merge 3 commits into
mainfrom
osdc/pr-test-arc

Conversation

@huydhn

@huydhn huydhn commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Migrates pr-test.yml off EC2.

  • Drops the docker run / docker exec wrapper. OSDC runners have no host docker daemon and every job already runs in a container, so ghcr.io/pytorch/torchbench:latest moves into container: and the two .ci/torchbench scripts run directly. The image puts its own venv on PATH, so python3 still resolves there.
  • linux.aws.a100 → mt-l-x86iavx512-11-125-a100, linux.24xlarge → mt-l-x86iavx512-94-192, per pytorch/pytorch .github/arc.yaml. --gpus all comes from the matrix so only the cuda leg gets it.
  • HF models download fresh into RUNNER_TEMP rather than reading the shared /mnt/hf_cache mount, same as Build the Linux wheels on OSDC runners TensorRT#4745.
  • Adds unittest-xml-reporting to requirements.txt — see below.

Not using linux_job_v3 because HUGGING_FACE_HUB_TOKEN is scoped to the docker-s3-upload environment, and a job calling a reusable workflow cannot declare an environment.

Two docker flags are dropped, neither settable per job on a pod: --shm-size=32g (OSDC gives every pod a 2Gi /dev/shm, what pytorch's own docker CI uses) and --cap-add=SYS_PTRACE --security-opt seccomp=unconfined.

Why xmlrunner

The cuda leg runs torchbenchmark._components.test.test_subprocess, which calls torch's run_tests(). That imports xmlrunner when TEST_SAVE_XML is set — CI sets it — so the job died before any test ran. The module comes from unittest-xml-reporting; pin matches pytorch's .ci/docker/requirements-ci.txt.

It has never been installed, so the cuda leg has never been able to pass. Unnoticed because Test cuda does not run on main — linux.aws.a100 queues it until something cancels it, in all 8 runs on record. Moving to an available runner is what surfaced it. Takes effect after the next nightly image rebuild.

Known failure, out of scope

Test cpu fails with FAILED (errors=8, skipped=136) / WeightsUnpickler error. This is pre-existing on main — identical markers in all 8 recent runs, including on EC2. Not a runner issue and not fixed here.

Authored with Claude Code.

@meta-cla meta-cla Bot added the cla signed label Sep 30, 2026
huydhn added a commit to pytorch/test-infra that referenced this pull request Sep 30, 2026
cache_dir rejected a bare name outright, so seeding anything torchbench needs
was impossible:

    ValueError: cannot parse repo id 'albert-base-v2'

gpt2, t5-base, t5-large, albert-base-v2, distilbert-base-uncased and
xlm-roberta-base are all real repo ids in that form, and the hub caches them as
models--<name> with no org segment. That is not the same directory as the
namespaced alias -- models--gpt2 and models--openai-community--gpt2 are
distinct -- so a cache holding one still misses a job asking for the other,
which is exactly how pytorch/benchmark#2711 failed:

    OSError: [Errno 30] Read-only file system:
        '/mnt/hf_cache/hub/models--albert-base-v2'

with models--albert--albert-base-v2 sitting right there in the same bucket.

--from-hub had to learn the same shape: it unpacked the directory into exactly
three parts, which a bare id does not have.

A bare *dataset* id is still rejected. There is no namespace to copy from, so
that one really is a typo.

The test asserting bare names are rejected was asserting my own bug, so it is
replaced rather than kept.

Authored with Claude Code.
huydhn added 2 commits October 1, 2026 10:47
Drops the docker wrapper. On EC2 the runner is a bare host, so the job had to
pull ghcr.io/pytorch/torchbench:latest, `docker run` it with the workspace
bind-mounted at /tmp/workspace, and `docker exec` the two .ci scripts into it.
An OSDC runner has no host docker daemon and every job already runs inside a
container, so the same thing is expressed by naming that image in `container:`
and running the scripts directly. The image sets PATH to its own
/workspace/benchmark/.venv, so `python3` still resolves to the venv.

Runners follow pytorch/pytorch's .github/arc.yaml mapping: linux.aws.a100 ->
mt-l-x86iavx512-11-125-a100 (1x A100, 11 vCPU / 125 GiB), and linux.24xlarge
(c5.24xlarge, 96/192) -> mt-l-x86iavx512-94-192. `--gpus all` comes from the
matrix so only the cuda leg gets it.

The HF step is the standard read-only-mount dance, same shape as
test-infra's linux_job_v3: models come from /mnt/hf_cache, everything HF
writes anyway goes to RUNNER_TEMP.

Two docker flags are dropped because a pod cannot set them per job:
--shm-size=32g (OSDC gives every pod a 2Gi /dev/shm, which is what pytorch's
own docker CI runs with) and --cap-add=SYS_PTRACE --security-opt
seccomp=unconfined.

`environment: docker-s3-upload` stays -- HUGGING_FACE_HUB_TOKEN is scoped to
that environment, not the repo, which is also why this cannot use
linux_job_v3: a job calling a reusable workflow cannot declare an environment.

Authored with Claude Code.
Same shape as pytorch/TensorRT#4745: HF_HOME, HF_HUB_CACHE and
HF_DATASETS_CACHE all point at RUNNER_TEMP, so nothing touches the read-only
/mnt/hf_cache mount.

Reading from the mount is a poor fit here. install.py pulls the whole hf_*
model set, and most of those are the legacy un-namespaced ids -- gpt2,
t5-base, t5-large, albert-base-v2, distilbert-base-uncased, xlm-roberta-base.
Those cache as models--<name>, a different directory from the namespaced alias
models--<org>--<name>, so a bucket holding one still misses a job asking for
the other. Keeping twelve repos in that exact form in step across four regions
is a standing obligation for one workflow, and the failure mode is the bad one:
green in whichever region the PR happened to land in, broken on main from the
region it did not.

Authored with Claude Code.
The cuda leg of .ci/torchbench/test.sh runs

    python3 -m torchbenchmark._components.test.test_subprocess

which calls torch's run_tests(). That imports xmlrunner when TEST_SAVE_XML is
set, and CI sets it, so the job dies before any test runs:

    File ".../torch/testing/_internal/common_utils.py", line 1366
      import xmlrunner
    ModuleNotFoundError: No module named 'xmlrunner'

The module comes from unittest-xml-reporting, not from a package named
xmlrunner; the pin matches pytorch's .ci/docker/requirements-ci.txt.

This has never been installed, so the cuda leg has never been able to pass.
It went unnoticed because Test cuda does not run on main -- linux.aws.a100
queues it until something cancels it, in every run on record. It only showed
up once the job moved to a runner that was actually available.

Authored with Claude Code.
@huydhn
huydhn requested a review from xuzhao9 October 1, 2026 21:44
@huydhn
huydhn marked this pull request as ready for review October 1, 2026 21:44
@meta-codesync

meta-codesync Bot commented Oct 1, 2026

Copy link
Copy Markdown

@huydhn has imported this pull request. If you are a Meta employee, you can view this in D122886793.

@meta-codesync meta-codesync Bot closed this in 284984a Oct 2, 2026
@meta-codesync meta-codesync Bot added the Merged label Oct 2, 2026
@meta-codesync

meta-codesync Bot commented Oct 2, 2026

Copy link
Copy Markdown

@huydhn merged this pull request in 284984a.

This branch had an error being deployed

1 failed deployment
docker-s3-upload — 5be83824 Deployed Oct 1, 2026 by huydhn via Test cuda #1699
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant