This directory contains configuration and utilities for offloading complex evaluation benchmarks to a separate environment using the eval_delegate mechanism. It is designed to integrate with Nemo Skills for running benchmarks like AIME25, Arena-Hard, and HLE, which may require specific environments distinct from the main training setup.
The setup allows miles to delegate evaluation tasks to a dedicated "Skills" server. This creates a clear separation of concerns:
- miles Container: Runs the main training loop and hosts the model using SGLang.
- Skills Container: Hosts the
nemo_skillsenvironment, runs the evaluation logic, and queries the model running in the miles container.
- A writable host directory for cached data (e.g.,
/data/.cache). - Docker installed with NVIDIA GPU support.
Create a Docker network to allow communication between the miles and Skills containers.
docker network create skills-netStart the main container where miles and the model will run. Replace <miles container name> with your desired name (e.g., miles_main).
docker run \
-itd \
--shm-size 32g \
--gpus all \
-v /data/.cache:/root/.cache \
-v /dev/shm:/shm \
--ipc=host \
--privileged \
--network skills-net \
--name <miles container name> \
radixark/miles:latest \
/bin/bashStart the container that will run the evaluation benchmarks. Replace <env container name> with your desired name (e.g., skills_env).
docker run \
-itd \
--shm-size 32g \
--gpus all \
-v /data/.cache:/root/.cache \
-v /dev/shm:/shm \
--ipc=host \
--privileged \
--network skills-net \
--name <env container name> \
--network-alias skills_server \
guapisolo/nemoskills:0.7.1 \
/bin/bashEnter the Skills container and set up the environment.
a) Install Dependencies
# Clone repositories
git clone -b miles_skills https://github.com/guapisolo/miles.git /opt/miles
git clone -b miles https://github.com/guapisolo/Skills.git /opt/Skills
# Install Skills package
cd /opt/Skills
pip install -e . --no-depsb) Prepare Datasets
Download and prepare the datasets you intend to use.
cd /opt/Skills/nemo_skills/dataset
python3 aime25/prepare.py
python3 hle/prepare.py
python3 arena-hard/prepare.pyc) Start the Evaluation Server
Start the server that listens for evaluation requests from miles.
cd /opt/miles
python examples/eval/nemo_skills/skills_server.py \
--host 0.0.0.0 \
--port 9050 \
--output-root /opt/skills-eval \
--config-dir examples/eval/nemo_skills/config \
--cluster local_cluster \
--max-concurrent-requests 512 \
--openai-model-name miles-openai-model*Note: You can now connect to the server at skills_server:9050 from within the skills-net Docker network. The server always proxies evaluation traffic to an OpenAI-compatible sglang router (miles starts and manage the router), so adjust --openai-model-name and --max-concurrent-requests as needed for your deployment.
The example scripts are located in examples/eval/scripts. Here is an example workflow for training Qwen3-4B with delegated evaluation.
Enter the miles container and install the package.
cd /root/miles
git pull
pip install -e . --no-deps# Download model weights (Qwen3-4B)
hf download Qwen/Qwen3-4B --local-dir /root/Qwen3-4B
# Download training dataset (dapo-math-17k)
hf download --repo-type dataset zhuzilin/dapo-math-17k \
--local-dir /root/dapo-math-17kYou need to convert the HF model to the format required by Megatron-LM. Ensure you load the correct model arguments first.
# Source model arguments
source scripts/models/qwen3-4B.sh
# Convert model
PYTHONPATH=/root/Megatron-LM python tools/convert_hf_to_torch_dist.py \
${MODEL_ARGS[@]} \
--hf-checkpoint /root/Qwen3-4B \
--save /root/Qwen3-4B_torch_distRun the training script.
bash examples/eval/scripts/run-qwen3-4B.shThe evaluation configuration is defined in examples/eval/scripts/multi_tasks.yaml. It specifies:
delegate: Configurations for the external skills server (URL, timeouts).datasets: List of datasets to evaluate on (e.g.,aime25,arena-hard).