Skip to content

Repository files navigation

Spindle

Spindle is a Tinker SDK-compatible backend run on Modal. Trainers run forward_backward and optim_step calls, then publish updated weights to autoscaling sampling replicas managed by the Stitch protocol (hence the name!). Currently, Spindle supports single-tenant full-parameter training as well as multi-tenant LoRA training.

Getting Started

Full-parameter training runs

For a dedicated full-parameter fine-tuning (FFT) run, use Python 3.12 and configure your Modal credentials and spindle-proxy secret as described below. Then:

import spindle
import tinker
from spindle.engines import qwen3_5_4b_full_64k

engine = qwen3_5_4b_full_64k()
with spindle.run(engine=engine) as (url, api_key):
    service = tinker.ServiceClient(base_url=url, api_key=api_key)
    training = spindle.create_full_training_client(service, engine.model)
    # Train and sample through the Tinker SDK here.

Our FFT path is not Tinker compatible, but roughly obeys the same abstractions.

See scoped runs for recovery and custom engines, and the Codeforces example for a complete training loop with sandbox judging and checkpoints.

LoRA training runs

The traditional Tinker path uses LoRA training, which is implemented via a multi-tenant Miles/Megatron backend in our system. Our LoRA path is Tinker-compatible out of the box on any of our supported models:

import os
import tinker

service = tinker.ServiceClient(
    base_url=os.environ["TINKER_BASE_URL"],
    api_key=os.environ["TINKER_API_KEY"],
)
training = service.create_lora_training_client(
    base_model="Qwen/Qwen3.5-9B-Base",
    rank=16,
)
# Train and sample through the Tinker SDK here.

Shared deployment quick start

Shared deployments use Python recipes inheriting from BaseConfig. See Python deployment configs. Keep the active Python config list in scripts/deploy_models.sh; run it to deploy the complete list.

Install Spindle into your own Python project, deploy it once to Modal, then call its API from your training scripts. The commands below work in Bash or Zsh.

If someone has already deployed Spindle for you, install the package in step 1, then skip to step 4 with the server URL and Spindle API key they provide. API clients do not need Modal deployment credentials or sampler proxy tokens.

1. Install into your project

With uv installed:

uv init --python 3.12 my-spindle-project
cd my-spindle-project
uv add 'modal-spindle @ git+https://github.com/modal-projects/spindle.git'

2. Configure Modal and secrets once

Use a Modal account with permission to deploy apps and create secrets in your chosen environment. Authenticate if you have not already configured credentials for that workspace:

uv run modal token new
export MODAL_ENVIRONMENT=your-environment
uv run modal environment list

Set your-environment to an existing environment before creating secrets so the secrets and deployment use the same environment. For automation, existing MODAL_TOKEN_ID / MODAL_TOKEN_SECRET credentials can be supplied instead of the interactive login.

There are three separate credentials:

Credential Purpose Who needs it
Modal API token / local profile Manage Modal resources Deployer
TINKER_API_KEY in the spindle-api secret Authenticate calls to the Spindle API Deployer and API clients
Proxy token in the spindle-proxy secret Let Spindle reach protected sampler pools Deployed control plane and trainers

For a new deployment, generate a Spindle API key and store it as a Modal secret in the existing environment:

export TINKER_API_KEY="your api key here"
uv run modal secret create spindle-api \
  TINKER_API_KEY="$TINKER_API_KEY"

Sampler pools use Modal proxy authentication. Create a proxy token and allow it in the deployment environment. If you deploy with service-user credentials or lack permission to create workspace proxy tokens, have a workspace owner or manager provision an allowed token first; set MODAL_PROXY_TOKEN_ID and MODAL_PROXY_TOKEN_SECRET to that pair and skip the token-creation block below.

read -r MODAL_PROXY_TOKEN_ID MODAL_PROXY_TOKEN_SECRET < <(uv run python -c '
import os
import modal
tokens = modal.Workspace.from_context().proxy_tokens
token = tokens.create()
tokens.allow(token.token_id, os.environ["MODAL_ENVIRONMENT"])
print(token.token_id, token.token_secret)
')

Store the token in the same environment under the secret spindle-proxy:

uv run modal secret create spindle-proxy \
  MODAL_PROXY_TOKEN_ID="$MODAL_PROXY_TOKEN_ID" \
  MODAL_PROXY_TOKEN_SECRET="$MODAL_PROXY_TOKEN_SECRET"

3. Deploy the installed package

Create a configuration from a preset, validate it, and deploy it with Python 3.12:

uv run spindle config init --preset qwen35-9b-lora-16k > deployment.py
uv run spindle config validate deployment.py
uv run spindle deploy deployment.py

This deploys the shared app and prints its server URL. Add more Python config files to the same command to serve more recipes. Always supply the complete current set. The Miles commit is pinned in miles_image.py; see Python deployment configs.

From a repository checkout, maintain the list in scripts/deploy_models.sh and run that script. spindle deploy supplies the current configs and frontend platform settings to Modal.

Deploying the server doesn't allocate any GPUs; rather, this allocation for both the training and sampling sides are done on demand. See cold starts and capacity configuration before running a larger workload.

4. Clean up

After the script exits, session heartbeats stop and Spindle's periodic cleaner reclaims idle training models and their latest sampler pools. Check that cleanup has finished in the Modal dashboard or list apps with:

uv run modal app list

To tear down the deployment, stop its spindle-fft-... sampler apps, then the frontend selected with --app (spindle by default), using uv run modal app stop <app-id>. Stopping the frontend does not stop sampler apps.

Next steps

Refer to the docs for design and for more advanced features when working with either the full-parameter or LoRA paths:

Read Working with Full Fine-Tunes for full training, or Working with Multi-LoRA for shared Miles adapters, batch submission, scheduling, and sampling.

and the raw Tinker RL example for sampling and a toy policy update. Copy examples you want to run into your project; repository scripts/ are not installed with the package.

The W&B RL example extends it to a multi-step loop that logs reward, response length, and Spindle's training metrics to Weights & Biases from the client side; tinker-cookbook users can instead set wandb_project/wandb_name on the cookbook Config.

See Design for the control-plane, training-engine, and sampling architecture.

See Profiling for how to enable the torch.profiler trace of a training step and read it in Perfetto.

See Observability for OTLP export to Datadog or a custom destination, experiment labels, and the complete span/metric inventory.

Validation

See FFT validation and LoRA validation for end-to-end training runs we've done with both parameterizations. The Codeforces codegolf example provides a larger-scale e2e code-RL training run, which trains Qwen3.5-9B with GRPO or TailRL advantages for correctness and short solutions. It includes a sandboxed judge, checkpoint recovery, and commands to continue a checkpoint with a different reward or advantage estimator, as well as pass@k and best-of-k evaluation.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages