Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
71 changes: 71 additions & 0 deletions .github/workflows/package.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
name: Package
on:
pull_request:
push:
branches: [main]
workflow_call:
inputs:
release:
type: boolean
default: false
permissions:
contents: read
env:
UV_DEFAULT_INDEX: https://pypi.org/simple
jobs:
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: python -m pip install uv==0.9.26
# uv builds the wheel from the sdist, checking that source releases rebuild.
- run: uv build
- run: uvx --from twine twine check --strict dist/*
- run: uv run --no-project --with packaging python scripts/check_distribution.py
- name: Check PyPI release prerequisites
if: inputs.release
env:
RELEASE_TAG: ${{ github.event.release.tag_name }}
run: uv run --no-project --with packaging python scripts/check_distribution.py --for-pypi --tag "$RELEASE_TAG"
- uses: actions/upload-artifact@v4
with:
name: distributions
path: dist/*
if-no-files-found: error
install:
needs: build
runs-on: ubuntu-latest
strategy:
matrix:
python: ['3.11', '3.12']
distribution: [wheel, sdist]
steps:
- uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python }}
- uses: actions/download-artifact@v4
with:
name: distributions
path: dist
- name: Install in a clean environment
env:
DISTRIBUTION: ${{ matrix.distribution }}
run: |
python -m venv "$RUNNER_TEMP/spindle-install"
echo "$RUNNER_TEMP/spindle-install/bin" >> "$GITHUB_PATH"
if [ "$DISTRIBUTION" = wheel ]; then
"$RUNNER_TEMP/spindle-install/bin/python" -m pip install --index-url https://pypi.org/simple dist/*.whl
else
"$RUNNER_TEMP/spindle-install/bin/python" -m pip install --index-url https://pypi.org/simple dist/*.tar.gz
fi
- name: Check installed API, CLI, and packaged presets
working-directory: ${{ runner.temp }}
run: |
python -m pip check
python -c "import spindle; from importlib.metadata import version; assert spindle.__version__ == version('modal-spindle'); from spindle.engines import qwen3_5_4b_full_64k; qwen3_5_4b_full_64k()"
spindle --help
spindle config init --preset qwen35-9b-lora-16k > deployment.py
spindle config validate deployment.py
30 changes: 30 additions & 0 deletions .github/workflows/publish.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
name: Publish to PyPI
on:
release:
types: [published]
permissions:
contents: read
concurrency:
group: pypi-${{ github.event.release.tag_name }}
cancel-in-progress: false
jobs:
tests:
uses: ./.github/workflows/tests.yml
package:
uses: ./.github/workflows/package.yml
with:
release: true
publish:
needs: [tests, package]
runs-on: ubuntu-latest
environment:
name: pypi
url: https://pypi.org/p/modal-spindle
permissions:
id-token: write
steps:
- uses: actions/download-artifact@v4
with:
name: distributions
path: dist
- uses: pypa/gh-action-pypi-publish@release/v1
4 changes: 3 additions & 1 deletion .github/workflows/tests.yml
Original file line number Diff line number Diff line change
@@ -1,10 +1,13 @@
name: Core CPU tests
on:
workflow_call:
pull_request:
push:
branches: [main]
permissions:
contents: read
env:
UV_DEFAULT_INDEX: https://pypi.org/simple
jobs:
test:
runs-on: ubuntu-latest
Expand All @@ -19,4 +22,3 @@ jobs:
# so the loss, gradient, and FP32-head regressions also run in CI.
- run: uv pip install --python .venv/bin/python torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu
- run: uv run --no-sync pytest -q
- run: uv build
39 changes: 24 additions & 15 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,11 @@

Spindle is a Tinker SDK-compatible backend run on Modal. Trainers run `forward_backward` and `optim_step` calls, then publish updated weights to autoscaling sampling replicas managed by the [Stitch](https://github.com/modal-projects/stitch) protocol (hence the name!). Currently, Spindle supports single-tenant full-parameter training as well as multi-tenant LoRA training.

# Getting Started
Spindle supports Python 3.11 and 3.12; use Python 3.12 for Modal deployments.
The package is being prepared for PyPI as `modal-spindle`. Until the first release,
install from Git using the instructions below.

# Getting Started

## Full-parameter training runs

Expand All @@ -23,8 +27,8 @@ with spindle.run(engine=engine) as (url, api_key):

Our FFT path is *not* Tinker compatible, but roughly obeys the same abstractions.

See [scoped runs](docs/scoped-runs.md) for recovery and custom engines,
and the [Codeforces example](examples/codeforces-codegolf/README.md) for a complete
See [scoped runs](https://github.com/modal-projects/spindle/blob/main/docs/scoped-runs.md) for recovery and custom engines,
and the [Codeforces example](https://github.com/modal-projects/spindle/blob/main/examples/codeforces-codegolf/README.md) for a complete
training loop with sandbox judging and checkpoints.

## LoRA training runs
Expand All @@ -48,7 +52,7 @@ training = service.create_lora_training_client(

## Shared deployment quick start

Shared deployments use Python recipes inheriting from `BaseConfig`. See [Python deployment configs](docs/deployment-configs.md). Keep the active Python config list in [scripts/deploy_models.sh](scripts/deploy_models.sh); run it to deploy the complete list.
Shared deployments use Python recipes inheriting from `BaseConfig`. See [Python deployment configs](https://github.com/modal-projects/spindle/blob/main/docs/deployment-configs.md). Keep the active Python config list in [scripts/deploy_models.sh](https://github.com/modal-projects/spindle/blob/main/scripts/deploy_models.sh); run it to deploy the complete list.

Install Spindle into your own Python project, deploy it once to Modal, then call
its API from your training scripts. The commands below work in Bash or Zsh.
Expand Down Expand Up @@ -137,11 +141,11 @@ uv run spindle config validate deployment.py
uv run spindle deploy deployment.py
```

This deploys the shared app and prints its `server` URL. Add more Python config files to the same command to serve more recipes. Always supply the complete current set. The Miles commit is pinned in `miles_image.py`; see [Python deployment configs](docs/deployment-configs.md).
This deploys the shared app and prints its `server` URL. Add more Python config files to the same command to serve more recipes. Always supply the complete current set. The Miles commit is pinned in `miles_image.py`; see [Python deployment configs](https://github.com/modal-projects/spindle/blob/main/docs/deployment-configs.md).

From a repository checkout, maintain the list in `scripts/deploy_models.sh` and run that script. `spindle deploy` supplies the current configs and frontend platform settings to Modal.

Deploying the server doesn't allocate any GPUs; rather, this allocation for both the training and sampling sides are done on demand. See [cold starts and capacity configuration](docs/full-fine-tunes.md#performance-and-behavior-considerations)
Deploying the server doesn't allocate any GPUs; rather, this allocation for both the training and sampling sides are done on demand. See [cold starts and capacity configuration](https://github.com/modal-projects/spindle/blob/main/docs/full-fine-tunes.md#performance-and-behavior-considerations)
before running a larger workload.

### 4. Clean up
Expand All @@ -161,32 +165,37 @@ using `uv run modal app stop <app-id>`. Stopping the frontend does not stop samp

Refer to the docs for design and for more advanced features when working with either the full-parameter or LoRA paths:

Read [Working with Full Fine-Tunes](docs/full-fine-tunes.md) for full training,
or [Working with Multi-LoRA](docs/multi-lora.md) for shared Miles adapters, batch
Read [Working with Full Fine-Tunes](https://github.com/modal-projects/spindle/blob/main/docs/full-fine-tunes.md) for full training,
or [Working with Multi-LoRA](https://github.com/modal-projects/spindle/blob/main/docs/multi-lora.md) for shared Miles adapters, batch
submission, scheduling, and sampling.

and the [raw Tinker RL example](scripts/rl_example.py) for sampling and a toy
and the [raw Tinker RL example](https://github.com/modal-projects/spindle/blob/main/scripts/rl_example.py) for sampling and a toy
policy update. Copy examples you want to run into your project; repository
`scripts/` are not installed with the package.

The [W&B RL example](scripts/wandb_rl_example.py) extends it to a multi-step
The [W&B RL example](https://github.com/modal-projects/spindle/blob/main/scripts/wandb_rl_example.py) extends it to a multi-step
loop that logs reward, response length, and Spindle's training metrics to Weights
& Biases from the client side; tinker-cookbook users can instead set
`wandb_project`/`wandb_name` on the cookbook `Config`.

See [Design](docs/design.md) for the control-plane, training-engine, and sampling
See [Design](https://github.com/modal-projects/spindle/blob/main/docs/design.md) for the control-plane, training-engine, and sampling
architecture.

See [Profiling](docs/profiling.md) for how to enable the `torch.profiler` trace of
See [Profiling](https://github.com/modal-projects/spindle/blob/main/docs/profiling.md) for how to enable the `torch.profiler` trace of
a training step and read it in Perfetto.

See [Observability](docs/observability.md) for OTLP export to Datadog or a custom
See [Observability](https://github.com/modal-projects/spindle/blob/main/docs/observability.md) for OTLP export to Datadog or a custom
destination, experiment labels, and the complete span/metric inventory.

## Validation

See [FFT validation](docs/validation.md) and [LoRA validation](docs/lora_validation.md)
for end-to-end training runs we've done with both parameterizations. The [Codeforces codegolf](examples/codeforces-codegolf/README.md) example provides a larger-scale e2e code-RL training run, which trains Qwen3.5-9B
See [FFT validation](https://github.com/modal-projects/spindle/blob/main/docs/validation.md) and [LoRA validation](https://github.com/modal-projects/spindle/blob/main/docs/lora_validation.md)
for end-to-end training runs we've done with both parameterizations. The [Codeforces codegolf](https://github.com/modal-projects/spindle/blob/main/examples/codeforces-codegolf/README.md) example provides a larger-scale e2e code-RL training run, which trains Qwen3.5-9B
with GRPO or TailRL advantages for correctness and short solutions. It includes
a sandboxed judge, checkpoint recovery, and commands to continue a checkpoint
with a different reward or advantage estimator, as well as pass@k and best-of-k evaluation.

## Development and releases

See [Publishing](https://github.com/modal-projects/spindle/blob/main/docs/publishing.md)
for first-release prerequisites, package validation, and the PyPI release process.
18 changes: 17 additions & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
@@ -1,11 +1,20 @@
[build-system]
requires = ["setuptools>=68", "wheel"]
requires = ["setuptools>=77.0.3"]
build-backend = "setuptools.build_meta"

[project]
name = "modal-spindle"
version = "0.1.0"
description = "Training and disaggregated sampling on Modal with a Tinker-compatible API"
readme = "README.md"
requires-python = ">=3.11,<3.13"
classifiers = [
"Development Status :: 3 - Alpha",
"Programming Language :: Python :: 3",
"Programming Language :: Python :: 3.11",
"Programming Language :: Python :: 3.12",
"Topic :: Scientific/Engineering :: Artificial Intelligence",
]
dependencies = [
"opentelemetry-exporter-otlp-proto-http>=1.39,<2",
"opentelemetry-sdk>=1.39,<2",
Expand All @@ -22,8 +31,15 @@ dependencies = [
"zstandard>=0.25.0",
]

[project.urls]
Homepage = "https://github.com/modal-projects/spindle"
Documentation = "https://github.com/modal-projects/spindle/blob/main/README.md"
Repository = "https://github.com/modal-projects/spindle"
Issues = "https://github.com/modal-projects/spindle/issues"

[tool.setuptools.packages.find]
where = ["src"]
include = ["spindle*"]

[tool.pytest.ini_options]
testpaths = ["tests"]
Expand Down
91 changes: 91 additions & 0 deletions scripts/check_distribution.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,91 @@
"""Check built metadata; --for-pypi also enforces release prerequisites."""

import argparse
import ast
from email.parser import BytesParser
from pathlib import Path
import tarfile
import tomllib
import zipfile

from packaging.requirements import Requirement
from packaging.specifiers import SpecifierSet


def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--for-pypi", action="store_true")
parser.add_argument("--tag")
args = parser.parse_args()
root = Path(__file__).resolve().parents[1]
project = tomllib.loads((root / "pyproject.toml").read_text())["project"]
errors = []

version = None
for statement in ast.parse((root / "src/spindle/__init__.py").read_text()).body:
if isinstance(statement, ast.Assign) and any(
isinstance(target, ast.Name) and target.id == "__version__"
for target in statement.targets
):
version = ast.literal_eval(statement.value)
if version != project["version"]:
errors.append("spindle.__version__ must match project.version")
if args.tag is not None and args.tag != f"v{project['version']}":
errors.append(f"Release tag must be v{project['version']}, got {args.tag!r}")

wheels = list((root / "dist").glob("*.whl"))
sdists = list((root / "dist").glob("*.tar.gz"))
if len(wheels) != 1 or len(sdists) != 1:
raise SystemExit(
"Expected one wheel and one sdist in dist/; use a clean build directory"
)
with zipfile.ZipFile(wheels[0]) as archive:
(metadata_path,) = (
name for name in archive.namelist() if name.endswith(".dist-info/METADATA")
)
wheel_metadata = BytesParser().parsebytes(archive.read(metadata_path))
for source in (root / "src/spindle").rglob("*.py"):
if source.relative_to(root / "src").as_posix() not in archive.namelist():
errors.append(f"Wheel is missing {source.relative_to(root)}")
with tarfile.open(sdists[0]) as archive:
(metadata_path,) = (
member
for member in archive.getmembers()
if member.name.count("/") == 1 and member.name.endswith("/PKG-INFO")
)
sdist_metadata = BytesParser().parsebytes(
archive.extractfile(metadata_path).read()
)

for kind, metadata in [("wheel", wheel_metadata), ("sdist", sdist_metadata)]:
for field, expected in [
("Name", "modal-spindle"),
("Version", project["version"]),
]:
if metadata[field] != expected:
errors.append(f"{kind}: {field} must be {expected!r}")
if SpecifierSet(metadata["Requires-Python"]) != SpecifierSet(
project["requires-python"]
):
errors.append(f"{kind}: Requires-Python must match pyproject.toml")
if not metadata["Summary"] or not metadata.get_payload().strip():
errors.append(f"{kind}: missing description or README")
if args.for_pypi:
if not metadata["License-Expression"] or not metadata.get_all(
"License-File"
):
errors.append(
f"{kind}: choose a license and include its file before release"
)
for dependency in metadata.get_all("Requires-Dist", []):
if Requirement(dependency).url:
errors.append(
f"{kind}: PyPI does not accept direct URL dependency: {dependency}"
)
if errors:
raise SystemExit("\n".join(errors))
print(f"Validated modal-spindle {project['version']} wheel and sdist")


if __name__ == "__main__":
main()
3 changes: 2 additions & 1 deletion scripts/deploy_models.sh
Original file line number Diff line number Diff line change
Expand Up @@ -9,9 +9,10 @@ cd "$(dirname "$0")/.."
deployment_files=(
src/spindle/configs/qwen35_9b_lora_16k.py
src/spindle/configs/qwen35_9b_lora_64k.py
src/spindle/configs/qwen35_9b_lora_128k.py
src/spindle/configs/qwen35_4b_fft_64k.py
src/spindle/configs/gpt_oss_20b_lora_64k.py
src/spindle/configs/qwen36_35b_a3b_lora_32k.py
)

spindle deploy "${deployment_files[@]}" "$@"
uv run --python 3.12 spindle deploy "${deployment_files[@]}" "$@"
35 changes: 0 additions & 35 deletions src/spindle/control_plane/deployments.py

This file was deleted.

Loading
Loading