Conversation
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 05:34 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 05:34 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 05:34 — with
GitHub Actions
Failure
huydhn
added a commit
to pytorch/test-infra
that referenced
this pull request
Sep 30, 2026
cache_dir rejected a bare name outright, so seeding anything torchbench needs
was impossible:
ValueError: cannot parse repo id 'albert-base-v2'
gpt2, t5-base, t5-large, albert-base-v2, distilbert-base-uncased and
xlm-roberta-base are all real repo ids in that form, and the hub caches them as
models--<name> with no org segment. That is not the same directory as the
namespaced alias -- models--gpt2 and models--openai-community--gpt2 are
distinct -- so a cache holding one still misses a job asking for the other,
which is exactly how pytorch/benchmark#2711 failed:
OSError: [Errno 30] Read-only file system:
'/mnt/hf_cache/hub/models--albert-base-v2'
with models--albert--albert-base-v2 sitting right there in the same bucket.
--from-hub had to learn the same shape: it unpacked the directory into exactly
three parts, which a bare id does not have.
A bare *dataset* id is still rejected. There is no namespace to copy from, so
that one really is a typo.
The test asserting bare names are rejected was asserting my own bug, so it is
replaced rather than kept.
Authored with Claude Code.
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 06:20 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 06:20 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 06:20 — with
GitHub Actions
Failure
Drops the docker wrapper. On EC2 the runner is a bare host, so the job had to pull ghcr.io/pytorch/torchbench:latest, `docker run` it with the workspace bind-mounted at /tmp/workspace, and `docker exec` the two .ci scripts into it. An OSDC runner has no host docker daemon and every job already runs inside a container, so the same thing is expressed by naming that image in `container:` and running the scripts directly. The image sets PATH to its own /workspace/benchmark/.venv, so `python3` still resolves to the venv. Runners follow pytorch/pytorch's .github/arc.yaml mapping: linux.aws.a100 -> mt-l-x86iavx512-11-125-a100 (1x A100, 11 vCPU / 125 GiB), and linux.24xlarge (c5.24xlarge, 96/192) -> mt-l-x86iavx512-94-192. `--gpus all` comes from the matrix so only the cuda leg gets it. The HF step is the standard read-only-mount dance, same shape as test-infra's linux_job_v3: models come from /mnt/hf_cache, everything HF writes anyway goes to RUNNER_TEMP. Two docker flags are dropped because a pod cannot set them per job: --shm-size=32g (OSDC gives every pod a 2Gi /dev/shm, which is what pytorch's own docker CI runs with) and --cap-add=SYS_PTRACE --security-opt seccomp=unconfined. `environment: docker-s3-upload` stays -- HUGGING_FACE_HUB_TOKEN is scoped to that environment, not the repo, which is also why this cannot use linux_job_v3: a job calling a reusable workflow cannot declare an environment. Authored with Claude Code.
Same shape as pytorch/TensorRT#4745: HF_HOME, HF_HUB_CACHE and HF_DATASETS_CACHE all point at RUNNER_TEMP, so nothing touches the read-only /mnt/hf_cache mount. Reading from the mount is a poor fit here. install.py pulls the whole hf_* model set, and most of those are the legacy un-namespaced ids -- gpt2, t5-base, t5-large, albert-base-v2, distilbert-base-uncased, xlm-roberta-base. Those cache as models--<name>, a different directory from the namespaced alias models--<org>--<name>, so a bucket holding one still misses a job asking for the other. Keeping twelve repos in that exact form in step across four regions is a standing obligation for one workflow, and the failure mode is the bad one: green in whichever region the PR happened to land in, broken on main from the region it did not. Authored with Claude Code.
huydhn
force-pushed
the
osdc/pr-test-arc
branch
from
October 1, 2026 17:51
92a0846 to
0818a02
Compare
huydhn
had a problem deploying
to
docker-s3-upload
October 1, 2026 17:52 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
October 1, 2026 17:52 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
October 1, 2026 19:14 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
October 1, 2026 19:14 — with
GitHub Actions
Failure
The cuda leg of .ci/torchbench/test.sh runs
python3 -m torchbenchmark._components.test.test_subprocess
which calls torch's run_tests(). That imports xmlrunner when TEST_SAVE_XML is
set, and CI sets it, so the job dies before any test runs:
File ".../torch/testing/_internal/common_utils.py", line 1366
import xmlrunner
ModuleNotFoundError: No module named 'xmlrunner'
The module comes from unittest-xml-reporting, not from a package named
xmlrunner; the pin matches pytorch's .ci/docker/requirements-ci.txt.
This has never been installed, so the cuda leg has never been able to pass.
It went unnoticed because Test cuda does not run on main -- linux.aws.a100
queues it until something cancels it, in every run on record. It only showed
up once the job moved to a runner that was actually available.
Authored with Claude Code.
huydhn
had a problem deploying
to
docker-s3-upload
October 1, 2026 21:41 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
October 1, 2026 21:41 — with
GitHub Actions
Failure
huydhn
marked this pull request as ready for review
October 1, 2026 21:44
|
@huydhn has imported this pull request. If you are a Meta employee, you can view this in D122886793. |
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Migrates
pr-test.ymloff EC2.docker run/docker execwrapper. OSDC runners have no host docker daemon and every job already runs in a container, soghcr.io/pytorch/torchbench:latestmoves intocontainer:and the two.ci/torchbenchscripts run directly. The image puts its own venv onPATH, sopython3still resolves there.linux.aws.a100→mt-l-x86iavx512-11-125-a100,linux.24xlarge→mt-l-x86iavx512-94-192, per pytorch/pytorch.github/arc.yaml.--gpus allcomes from the matrix so only the cuda leg gets it.RUNNER_TEMPrather than reading the shared/mnt/hf_cachemount, same as Build the Linux wheels on OSDC runners TensorRT#4745.unittest-xml-reportingtorequirements.txt— see below.Not using
linux_job_v3becauseHUGGING_FACE_HUB_TOKENis scoped to thedocker-s3-uploadenvironment, and a job calling a reusable workflow cannot declare an environment.Two docker flags are dropped, neither settable per job on a pod:
--shm-size=32g(OSDC gives every pod a 2Gi/dev/shm, what pytorch's own docker CI uses) and--cap-add=SYS_PTRACE --security-opt seccomp=unconfined.Why xmlrunner
The cuda leg runs
torchbenchmark._components.test.test_subprocess, which calls torch'srun_tests(). That importsxmlrunnerwhenTEST_SAVE_XMLis set — CI sets it — so the job died before any test ran. The module comes fromunittest-xml-reporting; pin matches pytorch's.ci/docker/requirements-ci.txt.It has never been installed, so the cuda leg has never been able to pass. Unnoticed because
Test cudadoes not run on main —linux.aws.a100queues it until something cancels it, in all 8 runs on record. Moving to an available runner is what surfaced it. Takes effect after the next nightly image rebuild.Known failure, out of scope
Test cpufails withFAILED (errors=8, skipped=136)/WeightsUnpickler error. This is pre-existing on main — identical markers in all 8 recent runs, including on EC2. Not a runner issue and not fixed here.Authored with Claude Code.