Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,9 @@ Entry: `cmd/modelmove`. License: MIT.
Stay on **synthetic fixtures** (a few MiB) in this lab. Do not download
multi-gigabyte checkpoints onto the 50 GiB lab disk.

Hub models are downloaded locally with the `hf` CLI and then moved as
plain directories — see `docs/hub-workflow.md`.

## Build and test

```sh
Expand Down
8 changes: 7 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -263,7 +263,13 @@ make check # everything CI runs
The local and SSH transports are complete and covered end to end. SSH
planning-phase manifests are gzip-compressed when both sides advertise
the `manifest-gzip` feature (protocol version stays 1, so older helpers
still work). Planned:
still work).

For models that start life on the Hugging Face Hub, see
[docs/hub-workflow.md](docs/hub-workflow.md): download locally with the
`hf` CLI, then transfer with `copy` / `sync` / `verify` / `diff`.

Planned:

- QUIC transport for high-latency links
- Direct Hugging Face Hub and Ollama registry sources
Expand Down
72 changes: 72 additions & 0 deletions docs/hub-workflow.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
# Hugging Face Hub workflow

`modelmove` does not talk to the Hugging Face Hub. It moves a model
**directory** between machines. The Hub step is a plain local download with
the `huggingface_hub` CLI; everything after that is `modelmove`.

## 1. Download locally

```bash
pip install -U "huggingface_hub[cli]"

# new CLI
hf download meta-llama/Llama-3-8B --local-dir ./llama-3-8b

# or the legacy entry point
huggingface-cli download meta-llama/Llama-3-8B --local-dir ./llama-3-8b
```

The result is an ordinary directory tree (`.safetensors` shards, config,
tokenizer). That is all `modelmove` needs.

## 2. Transfer

```bash
# First transfer to a new machine
modelmove copy ./llama-3-8b gpu-box:/srv/models/llama-3-8b

# Ship a fine-tune; only the changed tensors move
modelmove sync ./llama-3-8b-ft gpu-box:/srv/models/llama-3-8b-ft

# See what it would cost first
modelmove sync ./llama-3-8b-ft gpu-box:/srv/models/llama-3-8b-ft --dry-run
```

## 3. Verify

`copy` and `sync` record the manifest they applied in
`<dst>/.modelmove/`, so verification needs no arguments:

```bash
modelmove verify /srv/models/llama-3-8b
```

A mismatch exits with status 2 and names the exact byte ranges that no longer
match. Re-running `sync` repairs exactly those chunks.

## 4. Optional: price a migration with diff

If you already have manifests of two revisions (e.g. a base checkpoint and a
fine-tune), you can compute what a sync would move without touching the
network:

```bash
modelmove manifest ./llama-3-8b > base.manifest
modelmove manifest ./llama-3-8b-ft > finetune.manifest
modelmove diff base.manifest finetune.manifest
```

`diff` runs the same calculation a transfer would, so the numbers are what a
`sync` would actually send.

## Non-goals

- **No `hf://` source.** `modelmove` accepts local paths and SSH destinations
only. There is no Hub URL scheme and no registry client.
- **No tokens.** Authentication for gated models is handled by the `hf`
CLI (its own `HF_TOKEN` / login state), not by `modelmove`.
- **No LFS streaming.** Large files are downloaded by the Hub CLI into a
local directory first; `modelmove` never streams through the Hub.

The lab fixtures follow the same shape: a synthetic tree downloaded (or
created) locally, then moved with `modelmove`.
9 changes: 6 additions & 3 deletions internal/receiver/receiver.go
Original file line number Diff line number Diff line change
Expand Up @@ -299,6 +299,9 @@ func (r *Receiver) Chunk(d chunk.Digest, data []byte) error {
return nil
}

// statusFailed marks a per-file result whose transfer or verification broke.
const statusFailed = "failed"

// EndFile fills any chunks that were not sent from local sources, verifies the
// result and, under per-file atomicity, replaces the original.
func (r *Receiver) EndFile() (*FileResult, error) {
Expand All @@ -314,7 +317,7 @@ func (r *Receiver) EndFile() (*FileResult, error) {
a.f.Close()
a.f = nil
}
res.Status = "failed"
res.Status = statusFailed
res.Error = err.Error()
r.results = append(r.results, *res)
return res, err
Expand All @@ -338,7 +341,7 @@ func (r *Receiver) EndFile() (*FileResult, error) {
}
if err := a.f.Close(); err != nil {
a.f = nil
res.Status = "failed"
res.Status = statusFailed
res.Error = err.Error()
r.results = append(r.results, *res)
return res, err
Expand Down Expand Up @@ -366,7 +369,7 @@ func (r *Receiver) EndFile() (*FileResult, error) {
r.pending = append(r.pending, sf)
} else {
if err := r.commit(sf); err != nil {
res.Status = "failed"
res.Status = statusFailed
res.Error = err.Error()
r.results = append(r.results, *res)
return res, err
Expand Down