diff --git a/AGENTS.md b/AGENTS.md index 6657439..815058c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -13,6 +13,9 @@ Entry: `cmd/modelmove`. License: MIT. Stay on **synthetic fixtures** (a few MiB) in this lab. Do not download multi-gigabyte checkpoints onto the 50 GiB lab disk. +Hub models are downloaded locally with the `hf` CLI and then moved as +plain directories — see `docs/hub-workflow.md`. + ## Build and test ```sh diff --git a/README.md b/README.md index fc991ef..847eda8 100644 --- a/README.md +++ b/README.md @@ -263,7 +263,13 @@ make check # everything CI runs The local and SSH transports are complete and covered end to end. SSH planning-phase manifests are gzip-compressed when both sides advertise the `manifest-gzip` feature (protocol version stays 1, so older helpers -still work). Planned: +still work). + +For models that start life on the Hugging Face Hub, see +[docs/hub-workflow.md](docs/hub-workflow.md): download locally with the +`hf` CLI, then transfer with `copy` / `sync` / `verify` / `diff`. + +Planned: - QUIC transport for high-latency links - Direct Hugging Face Hub and Ollama registry sources diff --git a/docs/hub-workflow.md b/docs/hub-workflow.md new file mode 100644 index 0000000..ccc3325 --- /dev/null +++ b/docs/hub-workflow.md @@ -0,0 +1,72 @@ +# Hugging Face Hub workflow + +`modelmove` does not talk to the Hugging Face Hub. It moves a model +**directory** between machines. The Hub step is a plain local download with +the `huggingface_hub` CLI; everything after that is `modelmove`. + +## 1. Download locally + +```bash +pip install -U "huggingface_hub[cli]" + +# new CLI +hf download meta-llama/Llama-3-8B --local-dir ./llama-3-8b + +# or the legacy entry point +huggingface-cli download meta-llama/Llama-3-8B --local-dir ./llama-3-8b +``` + +The result is an ordinary directory tree (`.safetensors` shards, config, +tokenizer). That is all `modelmove` needs. + +## 2. Transfer + +```bash +# First transfer to a new machine +modelmove copy ./llama-3-8b gpu-box:/srv/models/llama-3-8b + +# Ship a fine-tune; only the changed tensors move +modelmove sync ./llama-3-8b-ft gpu-box:/srv/models/llama-3-8b-ft + +# See what it would cost first +modelmove sync ./llama-3-8b-ft gpu-box:/srv/models/llama-3-8b-ft --dry-run +``` + +## 3. Verify + +`copy` and `sync` record the manifest they applied in +`/.modelmove/`, so verification needs no arguments: + +```bash +modelmove verify /srv/models/llama-3-8b +``` + +A mismatch exits with status 2 and names the exact byte ranges that no longer +match. Re-running `sync` repairs exactly those chunks. + +## 4. Optional: price a migration with diff + +If you already have manifests of two revisions (e.g. a base checkpoint and a +fine-tune), you can compute what a sync would move without touching the +network: + +```bash +modelmove manifest ./llama-3-8b > base.manifest +modelmove manifest ./llama-3-8b-ft > finetune.manifest +modelmove diff base.manifest finetune.manifest +``` + +`diff` runs the same calculation a transfer would, so the numbers are what a +`sync` would actually send. + +## Non-goals + +- **No `hf://` source.** `modelmove` accepts local paths and SSH destinations + only. There is no Hub URL scheme and no registry client. +- **No tokens.** Authentication for gated models is handled by the `hf` + CLI (its own `HF_TOKEN` / login state), not by `modelmove`. +- **No LFS streaming.** Large files are downloaded by the Hub CLI into a + local directory first; `modelmove` never streams through the Hub. + +The lab fixtures follow the same shape: a synthetic tree downloaded (or +created) locally, then moved with `modelmove`. diff --git a/internal/receiver/receiver.go b/internal/receiver/receiver.go index 8e56f8e..c020946 100644 --- a/internal/receiver/receiver.go +++ b/internal/receiver/receiver.go @@ -299,6 +299,9 @@ func (r *Receiver) Chunk(d chunk.Digest, data []byte) error { return nil } +// statusFailed marks a per-file result whose transfer or verification broke. +const statusFailed = "failed" + // EndFile fills any chunks that were not sent from local sources, verifies the // result and, under per-file atomicity, replaces the original. func (r *Receiver) EndFile() (*FileResult, error) { @@ -314,7 +317,7 @@ func (r *Receiver) EndFile() (*FileResult, error) { a.f.Close() a.f = nil } - res.Status = "failed" + res.Status = statusFailed res.Error = err.Error() r.results = append(r.results, *res) return res, err @@ -338,7 +341,7 @@ func (r *Receiver) EndFile() (*FileResult, error) { } if err := a.f.Close(); err != nil { a.f = nil - res.Status = "failed" + res.Status = statusFailed res.Error = err.Error() r.results = append(r.results, *res) return res, err @@ -366,7 +369,7 @@ func (r *Receiver) EndFile() (*FileResult, error) { r.pending = append(r.pending, sf) } else { if err := r.commit(sf); err != nil { - res.Status = "failed" + res.Status = statusFailed res.Error = err.Error() r.results = append(r.results, *res) return res, err