From 80988c56b46a55004dbd5f91392e0ae66f16d71c Mon Sep 17 00:00:00 2001 From: David Abram Date: Sat, 29 Aug 2026 17:32:05 +0200 Subject: [PATCH 1/8] context: Add mutation cursor runtime coordinator plan Define the staged implementation and integration work for per-worktree locking, isolated Git snapshots, protocol coordination, recovery, and runtime tests. Capture acceptance criteria, task dependencies, design decisions, failure handling, validation, and follow-up context synchronization for the standalone runtime layer. Plan: mutation-cursor-runtime-coordinator Co-authored-by: SCE --- .../mutation-cursor-runtime-coordinator.md | 1761 +++++++++++++++++ 1 file changed, 1761 insertions(+) create mode 100644 context/plans/mutation-cursor-runtime-coordinator.md diff --git a/context/plans/mutation-cursor-runtime-coordinator.md b/context/plans/mutation-cursor-runtime-coordinator.md new file mode 100644 index 00000000..41563d25 --- /dev/null +++ b/context/plans/mutation-cursor-runtime-coordinator.md @@ -0,0 +1,1761 @@ +# Plan: mutation-cursor-runtime-coordinator + +## Change summary + +Adds the runtime (imperative-shell) layer that connects the verified, +already-implemented mutation-cursor protocol kernel (`protocol.rs`) and its +persistence layer (`store.rs`) to a real Git worktree: a concurrency-safe +checkout-identity primitive, an OS-backed per-worktree advisory lock for the +coordinator's own critical section (`worktree_lock.rs`), an isolated Git +snapshot service that captures the current worktree state as a durable +`TreeId` into the repository's normal object database and protects it from +GC/prune with an SCE-owned ref, never touching the caller's real index +(`git_snapshot.rs`), and a runtime coordinator (`coordinator.rs`) that +derives `WorktreeId` from checkout identity, materializes worktree/scope +rows, runs recovery when the durable state requires it, drives +`prepare`/`commit`, and retries on CAS conflict — including CAS conflict +while persisting a snapshot-failure taint — while reusing the single Git +snapshot captured for that invocation. This extends existing work — +`protocol.rs` and `store.rs` are unmodified — and completes the `read +durable state → speculative snapshot → derive transition → DB CAS → commit +or reject/retry` boundary the Quint model and +`context/cli/mutation-trace-protocol.md` already describe as future work. +The coordinator remains standalone: nothing in this PR wires it into a +harness hook, a CLI command, or `diff_traces`. + +This revision (PR #244 review) replaces the original plan's private, +alternate-backed object store with the repository's normal object database +plus SCE-owned refs (a real durability defect in the original design — see +Design decisions), makes `checkout::get_or_create_checkout_id` itself +concurrency-safe instead of only protecting the coordinator's own call to it +(closing the race for every caller, including `agent_trace_storage`), +correctly retries snapshot-failure taint persistence under CAS conflict +instead of claiming it succeeded unconditionally, and documents the +`(ScopeId, EventId)` replay-identity contract `RuntimeBoundary` places on +future harness adapters. + +A second revision (further PR #244 review) makes three additional +corrections. First, checkout-identity persistence is now crash-safe, not +merely concurrency-safe: `get_or_create_checkout_id` writes through a +unique temporary file and an atomic rename rather than writing the canonical +`checkout-id` path in place, so a process crash mid-write can never expose a +partially written identity to a later caller. Second, the coordinator no +longer bases its snapshot-failure taint-vs-bootstrap decision on durable +state read *before* the Git snapshot attempt — that read was stale by +construction and could miss a worktree row another caller materializes +concurrently; the decision now always comes from a fresh read taken *after* +the failure. Third, this plan corrects its own earlier claim that +create-only ref pinning produces "bounded" storage growth — growth is +unbounded over the repository's lifetime without a reconciliation pass, and +the Follow-up PR section now sequences that reconciliation pass and the +still-deferred filesystem external-taint marker as required runtime +completion work *before* any harness adapter may become a production +consumer of this coordinator. + +A third revision (further PR #244 review) corrects two remaining Git-level +issues in the snapshot mechanism itself, with the surrounding architecture +otherwise unchanged. First, the unborn-`HEAD` path now explicitly runs +`git read-tree --empty` to initialize a genuinely valid empty index, rather +than assuming a freshly reserved, never-created temp-index path was already +one — verified experimentally that a bare zero-byte file at that path is not +a valid Git index and makes `git add -A -- .` fail deterministically. The +temporary-index RAII guard is now explicit that it only reserves a unique +path and never creates a file there itself, so `git read-tree` is always +what first touches it. Second, the plan's `git gc`/`git prune` durability +test design is corrected: because Git is content-addressed, "an identical +unpinned tree" is a contradiction — pinning one copy of identical content +pins all of it — so the negative control now captures a second, +distinct-content tree in the same repository, deliberately unreachable from +any ref, and proves that one, not a same-content decoy, is what `git gc +--prune=now`/`git prune --expire=now` actually reclaims. + +## Acceptance criteria + +- [ ] AC1: A worktree observed for the first time establishes the currently + observed tree as its cursor baseline on its first boundary and emits no + mutation evidence for filesystem changes that predate that observation. + - Validate: `coordinator::tests::first_observation_establishes_baseline_without_evidence` via `nix build .#checks..cli-tests` +- [ ] AC2: An edit made between a `Start` and a subsequent `Advance` on the + same scope commits as exactly one `AiExclusive` mutation event whose + `before_tree`/`after_tree` match the Start baseline and the post-edit + snapshot. + - Validate: `coordinator::tests::exclusive_edit_between_start_and_advance_commits_one_event` +- [ ] AC3: Re-processing the identical `(scope, event)` boundary a second time + produces no duplicated mutation evidence. + - Validate: `coordinator::tests::replaying_the_same_scope_event_key_does_not_duplicate_evidence` +- [ ] AC4: A mutation made just before `Close` is still attributed using the + scope set as it existed immediately before `Close`'s own scope transition. + - Validate: `coordinator::tests::close_boundary_attributes_using_pre_close_scope_set` +- [ ] AC5: Two concurrently active scopes on one worktree yield `AiContended` + attribution for a subsequent `Advance`, independent of whether the two + scopes share an `ActorKind`. + - Validate: `coordinator::tests::contended_scopes_yield_ai_contended_same_and_different_actor` +- [ ] AC6: Capturing a Git snapshot never mutates the caller's real index, + staged changes, or working tree, correctly reflects staged, unstaged, + untracked, and deleted state, excludes ignored files, and — on an unborn + `HEAD` — starts from an explicitly initialized, genuinely valid empty + Git index (`git read-tree --empty`), never from a bare, freshly created + file the coordinator merely assumes Git will treat as empty, and produces + a correct tree even when the unborn repository has no files at all. + - Validate: `git_snapshot::tests::*` (index preservation, ignored files, deletion, unborn HEAD with a file, unborn HEAD with no files) +- [ ] AC7: A snapshot's `TreeId` remains resolvable through `diff_trees` + after the process that captured it has exited, its temporary index file no + longer exists, **and after a `git gc --prune=now` / `git prune + --expire=now` pass has run against the repository** — because it is + reachable from an SCE-owned ref in the repository's normal refs namespace, + not merely present in an isolated object store Git's own reachability + analysis knows nothing about. + - Validate: `git_snapshot::tests::snapshot_survives_a_fresh_process_and_temp_index_deletion`, `git_snapshot::tests::pinned_snapshot_survives_git_gc_prune_now`, `git_snapshot::tests::pinned_snapshot_survives_git_prune_expire_now` +- [ ] AC8: When two invocations race to commit from the same durable + revision, exactly one succeeds, the other reloads durable state and + recomputes its transition using its own originally captured snapshot, and + no second Git snapshot is taken for that invocation. + - Validate: `coordinator::tests::cas_conflict_reloads_and_recomputes_without_a_second_snapshot` +- [ ] AC9: Two coordinator invocations targeting the same worktree cannot + execute their critical sections concurrently; invocations on two different + worktrees, including two linked worktrees of the same repository, are not + serialized against each other and resolve to the same repository-scoped DB + with distinct `WorktreeId`s. + - Validate: `worktree_lock::tests::*` (contention, distinct-path independence); `runtime_tests::linked_worktrees_have_independent_locks_and_worktree_ids` +- [ ] AC10: A worktree whose durable state is `SnapshotFailure`-tainted or + `needs_rebaseline` is recovered exactly once, using the same snapshot + captured for the triggering boundary, before that boundary is processed; + `needs_rebaseline` recovery preserves live scopes, taint recovery abandons + them. + - Validate: `coordinator::tests::recovers_from_needs_rebaseline_preserving_live_scopes`, `coordinator::tests::recovers_from_snapshot_failure_taint_abandoning_live_scopes` +- [ ] AC11: A Git snapshot failure against an already-materialized worktree + durably persists a `SnapshotFailure` taint via `protocol::taint`, retried + under the same bounded semantic CAS-retry policy the main coordinator loop + uses when another committer's transition wins the race; `persisted_taint` + is `true` **only once a taint transition has actually been `Applied`**, + never merely attempted. A snapshot failure with no prior worktree row + makes no durable write and never fabricates a `TreeId`. The + bootstrap-vs-taint decision, and every taint-retry iteration, reads + durable worktree state fresh **after** the failure has already occurred — + never state read earlier in the same invocation — so a worktree another + caller materializes concurrently, while this invocation's own Git snapshot + is still being captured, is still found and correctly tainted. + - Validate: `coordinator::tests::snapshot_failure_taints_an_existing_worktree`, `coordinator::tests::snapshot_failure_taint_survives_a_losing_cas_and_commits_on_retry`, `coordinator::tests::snapshot_failure_taint_reports_not_persisted_after_retries_are_exhausted`, `coordinator::tests::snapshot_failure_before_any_baseline_makes_no_durable_write`, `coordinator::tests::snapshot_failure_taints_a_worktree_materialized_concurrently_during_capture` +- [ ] AC12: All concurrent first-time callers of + `checkout::get_or_create_checkout_id` for one physical checkout — whether + through the coordinator, `agent_trace_storage`, or any other caller — + converge on exactly one checkout ID, and the on-disk `checkout-id` file + ends up containing that same value. + - Validate: `checkout::tests::concurrent_first_time_callers_converge_on_one_checkout_id`, `runtime_tests::agent_trace_storage_and_coordinator_observe_the_same_checkout_id` +- [ ] AC13: For cooperating SCE processes, the canonical `checkout-id` path + is at every observable point either absent or contains exactly one + complete, valid checkout ID — never a partially written or truncated + value — including immediately after a process is interrupted between + creating its temporary identity file and renaming it into place, and an + orphaned temporary file left behind by such an interruption never + prevents a later call from converging on the canonical ID. + - Validate: `checkout::tests::interruption_before_rename_leaves_the_canonical_path_absent`, `checkout::tests::completed_rename_leaves_the_canonical_path_with_a_complete_id`, `checkout::tests::an_orphaned_temp_file_does_not_block_convergence_on_the_canonical_id` + +### Full validation + +- `nix flake check` +- `nix run .#pkl-check-generated` + +### Context sync + +- `context/cli/mutation-trace-protocol.md` — "Target end-state architecture" + currently states `coordinator.rs`/`git_snapshot.rs` "remain future work"; + update to reflect their existence and responsibilities. +- `context/cli/mutation-trace-store.md` — "Non-goals" currently states "no + `coordinator.rs` or `git_snapshot.rs` exists yet"; update once they exist. +- `context/cli/checkout-identity.md` — currently documents + `get_or_create_checkout_id` as a plain "reuses an existing ID or writes a + new one"; update to document the identity-creation lock and the + convergence guarantee it now provides to every caller. +- `context/overview.md` — the `mutation_trace` module description should + mention the new runtime coordinator layer while preserving the accurate + "not yet wired into any hook or command" framing. +- `context/context-map.md` — the `mutation-trace-protocol.md` line + annotation names the `coordinator.rs`/`git_snapshot.rs` seams as + not-yet-created; update once T06 lands. +- A new domain file (e.g. `context/cli/mutation-trace-runtime-coordinator.md`) + documenting the lock, snapshot, ref-durability, and coordinator design + decisions below is likely warranted; leave the exact filename to task + context synchronization. + +## Task context synchronization lifecycle + +Persist this field in every plan; this is durable plan state, not chat state: + +- **Task context synchronization:** every task carries `pending | synced | blocked`. + A completed task must be `synced` before another task can start or the plan can + finish. +- For `blocked`, record **Blocker**, **Required action**, and **Retry condition** + beside the status. Never infer `synced` from conversation history; write every + lifecycle transition to the plan file. + +## Constraints and non-goals + +- **In scope:** `cli/src/services/checkout/mod.rs` (behavior change: + `get_or_create_checkout_id` becomes concurrency-safe), + `cli/src/services/mutation_trace/worktree_lock.rs` (new), + `cli/src/services/mutation_trace/git_snapshot.rs` (new), + `cli/src/services/mutation_trace/coordinator.rs` (new), + `cli/src/services/mutation_trace/runtime_tests.rs` (new), + `cli/src/services/mutation_trace/mod.rs` (module registration only). +- **Out of scope:** Claude/Codex/OpenCode/Pi hook translation, Bash + `PreToolUse`/`PostToolUse` wiring, final Agent Trace `diff_traces` + insertion, commit attribution, auto-sync/remote sync, control-plane + changes, redesigning the verified Quint protocol, changing existing + mutation attribution semantics, general Git abstraction refactors + (`capabilities::GitOps`) unrelated to this protocol, a daemon/background + process, cross-machine locking, the filesystem `externalTaint`/`TAINTED` + marker (see Design decisions — Requirement 10), and a repository-wide + `refs/sce/mutation-cursor/**` reconciliation/pruning pass (see Design + decisions — snapshot durability, and Follow-up PR). Both deferred items + are required runtime-completion work before any harness adapter becomes a + production consumer of this coordinator — see Follow-up PR's "Runtime + completion sequence" — but neither is implemented by this PR's own task + stack. The `get_or_create_checkout_id` fix is a narrow, internal-locking + and internal-persistence change to an existing function's body — it + changes neither its signature nor the checkout-ID format, and is not the + "broad checkout identity redesign" this plan's brief separately excludes + (for example, changing how `WorktreeId` is derived, or introducing a + different identity scheme). +- **Constraints:** no new Cargo dependencies (`std::fs::File::lock`/ + `try_lock`/`unlock`, stabilized in Rust 1.89 and confirmed to compile and + run under this repository's pinned `1.95.0` toolchain, cover every OS + advisory lock this plan needs; `uuid` already provides `v4` for + `AttemptId` generation); no new database migration (`store.rs`'s existing + schema and API are sufficient); `protocol.rs`/`types.rs` are not modified; + `cargo clippy` runs with `pedantic`/`warnings` denied workspace-wide. + `git update-ref`/`git for-each-ref` (validated experimentally, see Design + decisions) are the only new Git plumbing commands beyond the previously + validated `read-tree`/`add`/`write-tree`/`diff`. +- **Non-goal:** implementing `protocol::database_failure` end to end. The + coordinator never calls it in this PR (see Design decisions — Requirement + 10); wiring it is the immediate follow-up PR's first concern. + +## Task stack + +- [ ] T01: `Make checkout-identity creation concurrency-safe and crash-safe` (status:todo) + - Task ID: T01 + - Scope: In — `cli/src/services/checkout/mod.rs`: + `get_or_create_checkout_id` gains an internal, dedicated + `/sce/checkout-id.lock` acquired only on the slow path (the + existing lock-free `read_checkout_id` fast path is unchanged for the + already-created case), re-checking for an existing ID under the lock + before generating one; the write itself goes through a unique + `/sce/checkout-id.tmp-` file + (`OpenOptions::create_new(true)`), a data sync (`File::sync_data()`), + and an atomic `std::fs::rename` into the canonical `checkout-id` path, + with a best-effort `#[cfg(unix)]` parent-directory sync after the rename + for additional crash-durability hardening. Out — the mutation-cursor + runtime lock (T02); any change to `resolve_git_dir`, + `read_checkout_id`'s public signature, or the checkout-ID file format; + any change to `agent_trace_storage` (it benefits automatically, with no + caller-side change needed); any cleanup pass for orphaned + `checkout-id.tmp-*` files (unnecessary — see Design decisions). + - Dependencies: none + - Done when: concurrent first-time callers on the same `git_dir` — from + any call site — converge on exactly one checkout ID, and the on-disk + file contains that value; an already-created checkout ID is still read + without ever acquiring the new lock; the canonical `checkout-id` path is + never observable as partially written, at any interruption point in the + write sequence; an orphaned `checkout-id.tmp-*` file left by an + interrupted attempt never blocks a later call from creating the + canonical file. + - Verify: `cargo test -p shared-context-engineering checkout::` (via `./scripts/run-cli-cargo.sh test`) + - Context synchronization: pending + +- [ ] T02: `Add the per-worktree OS advisory runtime lock` (status:todo) + - Task ID: T02 + - Scope: In — `worktree_lock.rs`: `WorktreeLock::acquire(git_dir, timeout)`, + RAII release, bounded polling acquisition. Out — wiring into + `coordinator.rs` (T05); resolving `git_dir` itself (caller-supplied); + the distinct checkout-identity-creation lock (T01, already landed by the + time this task depends on nothing from it). + - Dependencies: none + - Done when: `WorktreeLock::acquire` opens/creates + `/sce/mutation-cursor.lock`, blocks a second acquirer on the same + path until the first releases (via `Drop`) or the bounded timeout elapses, + returns a distinct, matchable error on timeout, and never treats the + lock file's mere existence as ownership. + - Verify: `cargo test -p shared-context-engineering worktree_lock::` (via `./scripts/run-cli-cargo.sh test`) + - Context synchronization: pending + +- [ ] T03: `Add the isolated Git snapshot service with ref-pinned durability` (status:todo) + - Task ID: T03 + - Scope: In — `git_snapshot.rs`: `GitSnapshotService` (resolves + `--git-dir` once; a unique, never-pre-created private temp index path + under `/sce/tmp/`, reserved by the RAII guard but left for Git + itself to create via `read-tree`; writes tree/blob objects into the + repository's normal, shared object database — no private object store, + no `GIT_OBJECT_DIRECTORY`/`GIT_ALTERNATE_OBJECT_DIRECTORIES` overrides), + `capture_tree` (branches on HEAD's existence between `git read-tree + HEAD` and `git read-tree --empty`, never a bare/absent index file), + `pin_tree` (creates `refs/sce/mutation-cursor//`), + `diff_trees`. Out — wiring `capture_tree`/`pin_tree` into + `coordinator.rs` (T04); locking (T02/T05 own that); any ref + deletion/reconciliation logic (deferred, see Design decisions and + Follow-up PR — this PR's pins are create-only). + - Dependencies: none + - Done when: `capture_tree` never mutates the real index/working + tree/staged diff, respects `.gitignore`, captures modifications, + additions, deletions, and an unborn `HEAD` — with and without existing + files — by explicitly initializing a valid empty index via + `git read-tree --empty` rather than relying on a bare temp-index file, + and produces an opaque `TreeId` whose length is never assumed; a tree + `pin_tree` protects remains resolvable via `diff_trees` after + `git gc --prune=now` and `git prune --expire=now`, while a distinct, + genuinely unreachable control tree captured and left unpinned in the + same repository is reclaimed by that same aggressive pass; `pin_tree` is + idempotent for the same `(worktree_id, tree)` pair. + - Verify: `cargo test -p shared-context-engineering git_snapshot::` (via `./scripts/run-cli-cargo.sh test`) + - Context synchronization: pending + +- [ ] T04: `Add the coordinator's core protocol-integration pipeline` (status:todo) + - Task ID: T04 + - Scope: In — `coordinator.rs`: `RuntimeBoundary` (with its + `(ScopeId, EventId)` replay-identity contract documented on the type), + `CoordinateOutcome`, `CoordinateError`, `SnapshotCapture` trait + (dependency-injection seam covering `capture` and `pin`), worktree/scope + materialization ordering (no worktree-existence read before Git snapshot + capture — that decision is made exactly once, fresh, inside the + snapshot-failure handler, never earlier), the load → recover-if-needed → + `prepare`/`commit` sequence with its bounded CAS-conflict retry loop + reusing one captured and already-pinned `TreeId`, and snapshot-failure + taint handling whose own bounded CAS-conflict retry loop's first + iteration is also where the bootstrap-vs-taint branch is decided. An + internal, generic-over-`SnapshotCapture` entrypoint is the unit under + test here; the public, lock-wrapped `coordinate()` entrypoint is T05. + Out — actual `WorktreeLock` acquisition (T05); checkout-id resolution + (T05, and now concurrency-safe and crash-safe by construction of T01 + regardless of ordering); ref deletion/reconciliation (deferred); + cross-module integration tests (T06). + - Dependencies: T03 (uses `GitSnapshotService` as the production + `SnapshotCapture` implementation; tests use a fake). + - Done when: every scenario in AC1–AC5, AC8, AC10, AC11 passes against the + internal generic entrypoint using a real temp-file `RepositoryAgentTraceDb` + and either the real `GitSnapshotService` or a fake, call-counting + `SnapshotCapture`; the fake proves both "exactly one `capture`" and "at + most one `pin`" per invocation, across CAS retries; a test using an + injected snapshot-capture failure and controlled store sequencing proves + a worktree row materialized *during* a failing capture attempt (not + before it) is still found and tainted by that same failing invocation. + - Verify: `cargo test -p shared-context-engineering coordinator::` (via `./scripts/run-cli-cargo.sh test`) + - Context synchronization: pending + +- [ ] T05: `Wire the worktree lock and checkout identity into coordinate()` (status:todo) + - Task ID: T05 + - Scope: In — the public `coordinate(repository_root, db, boundary)` + entrypoint: resolve `git_dir`, acquire `WorktreeLock`, resolve checkout + identity via the now-concurrency-safe `get_or_create_checkout_id` (T01), + derive `WorktreeId`, delegate to T04's internal pipeline for the rest of + the critical section. Out — the internal pipeline logic itself (T04, + unchanged here). + - Dependencies: T01, T02, T04 + - Done when: two threads targeting the same worktree cannot run their + critical sections concurrently (one observably blocks until the other's + `WorktreeLock` drops); `coordinate()`'s public signature matches + `Result`. + - Verify: `cargo test -p shared-context-engineering coordinator::tests::two_threads_on_the_same_worktree_serialize` (via `./scripts/run-cli-cargo.sh test`) + - Context synchronization: pending + +- [ ] T06: `Add cross-module runtime integration tests` (status:todo) + - Task ID: T06 + - Scope: In — `runtime_tests.rs`: end-to-end `coordinate()` calls against + two real linked Git worktrees sharing one repository-scoped DB, proving + distinct `WorktreeId`s/locks and non-serialization across worktrees; an + end-to-end snapshot-failure-then-recovery cycle across two real + sequential `coordinate()` calls; a cross-caller assertion that + `agent_trace_storage`'s own checkout-identity resolution and the + coordinator's observe the same checkout ID for one worktree. Out — any + new production code; this task is test-only. + - Dependencies: T05 + - Done when: AC9's linked-worktree assertions, AC12's cross-caller + assertion, and an end-to-end failure/recovery cycle pass using only the + public `coordinate()` API, real `git worktree add`, and a real temp-file + `RepositoryAgentTraceDb`. + - Verify: `cargo test -p shared-context-engineering runtime_tests::` (via `./scripts/run-cli-cargo.sh test`) + - Context synchronization: pending + +## Design decisions + +### Checkout-identity concurrency: fixed at the primitive, not at each caller + +`cli/src/services/checkout/mod.rs:109`'s `get_or_create_checkout_id` is +confirmed read-then-write with no lock: `read_checkout_id` returns `None`, +then a fresh UUIDv7 is generated and written. Two concurrent first-time +callers on the same `git_dir` can each observe `None`, each generate a +*different* UUIDv7, and the second `std::fs::write` wins — the two callers +then disagree about this checkout's identity for the remainder of their +respective processes. + +The original version of this plan protected only the coordinator's own call +to this function, by acquiring the mutation-cursor runtime lock first. That +was insufficient: `cli/src/services/agent_trace_storage/mod.rs:168` calls +the same function independently, with no lock at all, during every storage +resolution — a race between an `agent_trace_storage` resolution and a +coordinator invocation (or two concurrent `agent_trace_storage` +resolutions, with no coordinator involved at all) could still produce two +different checkout IDs for the same physical worktree. + +This revision fixes `get_or_create_checkout_id` itself (T01) instead: + +```text +read_checkout_id(git_dir) -- unchanged, lock-free fast path +if Some(id): return id -- the common case after the first invocation ever +else: + acquire /sce/checkout-id.lock (blocking; see "Blocking vs. timeout" below) + re-read_checkout_id(git_dir) -- another caller may have won while we waited + if Some(id): return id + generate UUIDv7 + -- crash-safe write, see "Crash-safe persistence" below -- + write the ID through checkout-id.tmp-, then atomically + rename it into checkout-id + (lock released on Drop) + return the generated id +``` + +Every caller of `get_or_create_checkout_id` — the coordinator, and +`agent_trace_storage`'s existing call, unmodified — now converges on exactly +one ID, with no caller-side change required. This is the "smallest correct +implementation" the brief asked for: it changes one function's body, not its +signature, not the checkout-ID file format, and not any caller. + +An alternative considered and rejected: skip the lock and instead write with +`OpenOptions::new().create_new(true)` directly on `checkout-id`, having a +losing writer read back the winner's value. This is tempting because it +needs no separate lock file, but it has a real correctness gap: `create_new` +only makes the *creation* atomic, not the subsequent `write_all` — a losing +reader could observe the winning writer's file after creation but before +its content is fully written, and would need to distinguish "empty/partial, +retry" from "genuinely corrupt" with no reliable signal. The lock-based +design avoids this for *concurrent readers*: a losing caller blocks on the +*same* lock the winner holds, and by the time it can re-read, the winner has +already finished writing. It does **not**, by itself, protect against a +*crashed* writer — the lock's own release on process death (or on `Drop` +after a successful write) says nothing about whether the write it was +guarding actually completed. That is a distinct problem, addressed next. + +### Crash-safe persistence: temp file plus atomic rename + +The lock closes the *concurrency* race; it does not by itself close the +*crash* race. `get_or_create_checkout_id`'s write step was, in the plan's +prior revision, effectively `std::fs::write(checkout-id, uuid)` — a single +call that is not atomic with respect to a process crash partway through: a +process can be killed after the OS has created/truncated `checkout-id` but +before the full UUID has been written to it. The advisory lock is released +automatically when the process dies, so a *later*, otherwise-correct caller +can then observe an empty or truncated `checkout-id` file through the +existing lock-free fast path — `read_checkout_id` already rejects this as +corruption (empty content, or a UUID parse failure) rather than recognizing +it as a recoverable incomplete write, and there was no mechanism to make +that state unreachable in the first place. + +The fix eliminates this state by construction, using a unique temporary +file plus an atomic rename rather than writing the canonical path in place: + +```text +generate UUIDv7 (call it `id`) +tmp_path = /sce/checkout-id.tmp- +file = OpenOptions::new().write(true).create_new(true).open(tmp_path) -- fails loudly on + -- the essentially + -- impossible case + -- of a name collision; + -- `id` is a fresh + -- UUIDv7 embedded in + -- the temp filename, + -- so no other attempt, + -- past or concurrent, + -- ever reuses this + -- exact path +file.write_all(id.as_bytes()) +file.sync_data() -- flush content to disk before the rename; see rationale below +drop(file) -- close the handle +std::fs::rename(tmp_path, /sce/checkout-id) -- atomic on the same filesystem +#[cfg(unix)] best-effort: open /sce/ as a File and call .sync_all() + (durability hardening for the rename's directory-entry update; + Windows has no equivalent operation and does not need one — see below) +``` + +Verified experimentally (`create_new` + `write_all` + `sync_data` + +`std::fs::rename`, reproducing the sequence above against a real +filesystem): the canonical path does not exist at any point before the +`rename` call; immediately after `rename`, it contains the complete, +untruncated ID; the temporary path no longer exists after a successful +rename. A directory handle opened with `File::open(dir_path)` on Unix +accepts `.sync_all()` without error, confirming the best-effort +parent-directory sync is implementable with no new dependency. + +- **Why `create_new(true)` for the temp file:** guarantees the temp file + itself starts empty and is never silently reused; combined with the fresh + UUID embedded in its name, a name collision is not just unlikely but + requires generating the identical UUIDv7 twice, which this design treats + as unreachable rather than something to retry around. +- **Why `sync_data()`, not `sync_all()`:** the content (the UUID bytes) is + what must survive a crash; the file's metadata (for example its mtime) is + not load-bearing for anything this plan reads, so the slightly cheaper + `sync_data()` is sufficient and is what is recommended, though + `sync_all()` would also be correct. +- **Why `std::fs::rename`, not an in-place write:** POSIX `rename(2)` is + atomic with respect to concurrent readers of the destination path — a + reader observes either the old state (absent, for a fresh checkout) or the + complete new file, never an intermediate state, because the rename only + ever changes a directory entry, not file content in place. +- **Windows:** `std::fs::rename` on Windows is implemented over + `MoveFileExW` with `MOVEFILE_REPLACE_EXISTING`, so replacing (or, here, + creating) the destination succeeds the same way it does on POSIX for this + single-writer, lock-protected sequence; this is the same cross-platform + rename-based atomic-replace pattern used throughout the Rust ecosystem + (for example by the `tempfile` crate's `persist()`) for exactly this + purpose. Windows has no direct analogue to a POSIX directory `fsync`, so + the best-effort parent-directory sync step above is `#[cfg(unix)]`-only; + it is a durability hardening addition, not a correctness requirement (the + rename itself, once it returns, already guarantees the *content* + atomicity property AC13 requires — the directory sync only narrows the + already-vanishingly-small window between "rename returned" and "the + rename survives an immediate, otherwise-unrelated full machine power + loss"). +- **Orphaned `checkout-id.tmp-*` files are harmless and need no cleanup + pass.** Each temp filename embeds a freshly generated UUIDv7 never reused + by any other attempt, so an orphan left behind by a crash between + `sync_data()` and `rename` can never collide with, block, or be confused + for a later attempt's own (differently named) temp file. Because the + lock-free fast path only ever reads the canonical `checkout-id` path, + never the `sce/` directory's full contents, an orphaned temp file is + simply inert leftover state — a few dozen bytes, at most once per + checkout, only in the already-rare case of a crash landing in this narrow + window. This plan implements no cleanup pass for these files, matching the + same "harmless, not self-cleaning, no reconciliation in this PR" judgment + already made for orphaned `refs/sce/mutation-cursor/**` pins (see Snapshot + durability, below) — for the identical reason: getting cleanup wrong by + deleting something still needed is worse than never cleaning up something + genuinely inert. + +### Checkout-identity lock vs. the mutation-cursor runtime lock: two locks, two invariants + +These are deliberately separate locks, at separate paths, guarding separate +invariants — conflating them was the original plan's mistake: + +- `/sce/checkout-id.lock` (T01): guards only "this checkout has at + most one durable identity." Acquired only on the rare slow path (identity + not yet created), held only across a few filesystem syscalls, and used by + every caller of `get_or_create_checkout_id`, not just the coordinator. +- `/sce/mutation-cursor.lock` (T02, unchanged from the original + plan): guards the coordinator's entire runtime critical section — snapshot + capture, worktree/scope materialization, recovery, and the CAS retry loop. + Held on every `coordinate()` call, not just the rare first one, and used + only by the coordinator. + +Because T01 makes checkout-identity creation safe for every caller +independent of the mutation-cursor lock, the coordinator's *ordering* +between "acquire the mutation-cursor lock" and "resolve checkout identity" +no longer has any bearing on this race (T05 still resolves identity after +acquiring the runtime lock, for the unrelated reason that it is the natural +first fact the rest of the critical section needs — not because it is what +closes the race, which it no longer needs to be). + +### Locking implementation: Rust std, not a new dependency + +Verified by direct compilation against this repository's pinned toolchain +(`rustc 1.95.0`): `std::fs::File::lock()`, `try_lock()` (returning +`Result<(), TryLockError>` with `TryLockError::{WouldBlock, Error(io::Error)}`), +and `unlock()` all compile and behave as expected (a second handle's +`try_lock()` on an already-locked file returns `WouldBlock`). This API was +stabilized in Rust 1.89, well before this project's pin, is documented as +cross-platform (`flock(2)` on Unix, `LockFileEx` on Windows), and is +advisory and exclusive by default — exactly what both locks need. No +`fd-lock`/`fs2`-style dependency is added. + +### Blocking vs. timeout, for each lock + +- **Mutation-cursor runtime lock (T02):** `File::lock()` has no built-in + timeout and blocks indefinitely. A hook invocation is a short-lived CLI + process whose caller (the harness) is typically waiting on it; if a prior + `sce` process hangs while holding this lock (a slow Git operation, a slow + DB retry, or a genuine bug), every subsequent hook invocation on that + worktree would otherwise queue up and block forever, effectively + deadlocking the harness's own tool-calling loop. `WorktreeLock::acquire` + therefore polls `try_lock()` on a short interval (100ms) against a bounded + deadline (10s) rather than calling the blocking `lock()` directly, and + returns a distinct `CoordinateError::LockAcquisition` on timeout with + actionable guidance. +- **Checkout-identity lock (T01):** blocks indefinitely (plain `lock()`, no + polling, no timeout). This is a deliberate asymmetry, not an oversight: + the critical section it guards is a handful of local filesystem syscalls + with no DB round-trip and no Git subprocess in it, so it cannot + meaningfully hang the way the runtime lock's much larger critical section + can. Adding timeout/polling machinery here would be complexity with no + corresponding risk to justify it. + +Both are internal constants, not a new `policies.database_retry`-style +config surface — this is a correctness/optimization boundary, not a +DB-facing concern, and there is no existing consumer to justify +config-driven tuning yet. + +### Snapshot durability: normal Git object database + SCE-owned refs, not a private store + +The original version of this plan proposed writing snapshot trees into a +private `/sce/mutation-objects/` store, with +`GIT_ALTERNATE_OBJECT_DIRECTORIES` pointing at the repository's normal +object directory so unchanged blobs could still resolve. **This was not a +sufficient durability guarantee.** A persisted `TreeId` in that design could +reference blobs that exist only in the repository's *normal* object +database; Git's own GC/prune reachability analysis has no knowledge that an +externally-stored SCE tree depends on them, so a routine `git gc` on the +user's repository could delete a blob a durable, already-committed +`MutationEvent` row still names — silently producing an unreadable +historical snapshot. The original plan's AC7 durability claim did not hold +by construction. + +Verified experimentally (reproducible; commands below match what was run, +as two separate trials — never as a same-repository comparison between two +copies of *identical* content, since Git's content-addressing would make +those the same object and any such comparison meaningless; see "Isolated +Git snapshot" below for the corrected same-repository design this plan +actually specifies, using two *distinct*-content trees): a tree object with +**no** ref pointing at it (directly or transitively) is reclaimed by +`git gc --prune=now`; a tree object **with** a ref pointing at it survives +both `git gc --prune=now` and `git prune --expire=now` intact and fully +resolvable via `git cat-file`/`git ls-tree`. This is exactly Git's ordinary +reachability contract, and it is the only mechanism in Git that actually +prevents pruning — an alternate-object directory participates in +*resolution*, never in *reachability*. + +Two designs were compared: + +- **A. Normal object database + SCE-owned refs pinning durable snapshot + trees.** Objects are written with plain `git write-tree` against the + repository's own object database (no `GIT_OBJECT_DIRECTORY` override at + all). Each tree that becomes durable evidence is made reachable by + creating one ref under a dedicated SCE namespace. Git's own GC naturally + preserves everything reachable from that ref — including every blob it + transitively references, whether newly written by this snapshot or + already shared with the rest of the repository's history — with no + alternate-object bookkeeping and no closure-copying logic to get right. + `diff_trees` needs no special object environment at all: once a tree is + written into the normal store, it resolves like any other Git object. +- **B. Fully self-contained private object store.** To honor the brief's + own framing ("must remain resolvable without depending on objects Git may + later prune from the normal store"), a private store cannot merely borrow + unchanged blobs through an alternate — reachability through an alternate + is not durability. It would have to *copy* the complete object closure + each durable snapshot tree depends on into the private store: for every + tree, walk it recursively, and for every blob/subtree not already present + in the private store, duplicate it in. This is real, non-trivial logic + (a closure walk plus copy, run on every snapshot, since a later `git gc` + could reclaim the source blob between snapshots), provides no benefit + design A does not already provide, and actively duplicates on-disk storage + for every unchanged blob a snapshot touches — where design A's + content-addressing means an unchanged blob referenced by many snapshots is + stored exactly once, in the one place the rest of the repository already + keeps it. + +**Recommendation: A**, as the brief's own stated default preference, and +confirmed by this comparison: it is simultaneously the simplest +implementation, the only one that is GC-safe by construction rather than by +a separately-maintained copying invariant, and the only one that needs no +extra object-directory environment variables for either writing or reading. +This also means `git_snapshot.rs` does **not** need to resolve +`--git-common-dir` at all (a further simplification and deviation from the +original plan, which planned to add `resolve_git_common_dir` to +`checkout.rs`): verified experimentally from inside a real linked worktree, +running `git write-tree`/`git update-ref` with only `GIT_DIR` set to the +worktree's own `--git-dir` (no object-directory override) already resolves +objects and refs through the shared commondir automatically, and a ref +created from the linked worktree's `GIT_DIR` is immediately visible when +queried from the main checkout's `GIT_DIR`, and vice versa — Git's own +worktree/commondir mechanism already gives every worktree of one repository +a shared object database and a shared default ref namespace with no help +from this plan. + +**Ref namespace:** `refs/sce/mutation-cursor//` — one +ref per distinct `(worktree_id, tree)` pair. `worktree_id` scopes cleanup +and any future audit tooling to one worktree without a repository-wide scan; +`tree-sha` makes the ref path itself content-addressed, so creating the same +ref twice (for example across CAS retries that reuse one captured tree, or +two invocations that happen to observe identical content) is a harmless, +idempotent `git update-ref` — verified experimentally. + +**Pin lifecycle: create-only in this PR — and that growth is unbounded, not +bounded, until reconciliation exists.** The coordinator (T04) calls +`pin_tree(worktree_id, observed_tree)` exactly once per invocation, +immediately after `capture_tree` succeeds and *before* any DB operation — +whether or not that invocation's boundary ultimately commits, is rejected, +or is a no-op. It never deletes a pin. This is a deliberate simplification, +not an oversight: safely deleting a pin requires confirming, under the +worktree lock, that the pinned tree is neither the worktree's current +`cursor_tree` nor referenced by any historical `mutation_trace_events` row — +a repository-wide-aware check this PR's narrow, per-invocation coordinator +has no efficient way to perform (a single invocation only knows about the +one worktree revision it just read, not the full DB history of every tree +that worktree has ever pinned). Getting cleanup wrong in the unsafe +direction — deleting a ref for a tree still needed by durable evidence — is +strictly worse than never deleting one, so this PR accepts the resulting +storage cost and defers actual reconciliation to a follow-up maintenance +pass (see Follow-up PR). + +An earlier revision of this plan described that cost as "bounded, linear" +growth. That was imprecise and is corrected here: growth is **unbounded +over the repository's lifetime**, in direct proportion to hook-invocation +volume — every distinct tree a `coordinate()` invocation ever observes +leaves one permanent ref and its protected objects, for as long as the +repository exists and this PR's create-only policy remains unmodified. +"Linear in invocation count" is not "bounded"; it only stays small in +absolute terms for as long as invocation volume itself stays small, which +this standalone PR's lack of any production caller happens to guarantee for +now, but which the moment a harness adapter starts driving high-frequency +hook traffic through `coordinate()` would no longer hold. This is exactly +why the Follow-up PR section below sequences a reconciliation pass as +required runtime-completion work *before* that wiring, not as an optional +later enhancement. + +This directly answers "when can refs safely be deleted": not from inside +this PR's per-invocation coordinator at all; only from a pass with +visibility into the full set of durable references a ref might still be +protecting. The follow-up reconciliation pass's retained-roots contract — +precise enough to bound its architectural shape without prescribing its +algorithm, which this PR does not implement: + +- **Retained roots** (a tree a reconciliation pass must never unpin while + any of these still reference it): every worktree's current + `mutation_trace_worktrees.cursor_tree`, and every historical + `mutation_trace_events.before_tree`/`after_tree`. These three columns are + the complete set of durably-referenced `TreeId`s in `store.rs`'s schema — + no other table stores one. +- **Conceptual shape:** read every durably-referenced `TreeId` from the + three roots above, list the full `refs/sce/mutation-cursor/**` namespace, + delete any ref whose tree appears in neither set, and let Git's own, + already-existing GC reclaim the now-unreachable objects on its own normal + schedule — this PR adds no GC invocation of its own, only the refs GC + already knows how to act on. +- **Concurrency safety is the hard part, and is explicitly out of this PR's + scope to solve, only to flag:** a reconciliation pass must never delete a + ref for a tree a coordinator has *just* pinned but has not yet committed + durably — deleting it out from under an in-flight `coordinate()` call + would reintroduce exactly the durability defect this revision fixes. Two + viable strategies for the follow-up to choose between: reconciling one + worktree at a time under that worktree's own `mutation-cursor.lock` (true + synchronization, at the cost of contending with live coordinator traffic), + or a conservative minimum-age threshold on candidate refs (for example, + never consider a ref for deletion until it is meaningfully older than any + realistic single `coordinate()` call could still be running, sidestepping + lock contention entirely at the cost of a small reclaim delay). This plan + does not choose between them — only states that whichever the follow-up + picks, it must not be able to delete a ref protecting a tree that could + still become durable. + +**Crash ordering.** Pin creation is a **blocking precondition** to the DB +CAS, not a follow-up step after it — this ordering is what makes the +brief's specific worry ("DB commit succeeds but ref persistence fails") +structurally impossible rather than merely unlikely: + +```text +write Git objects (git write-tree, into the normal object database) + ↓ +pin_tree(worktree_id, observed_tree) -- git update-ref, before any DB read/write + ↓ (on failure: treated identically to a snapshot-capture failure — see below; + ↓ the coordinator never attempts the DB CAS without a successful pin) +store.load_worktree / recover-if-needed / prepare / commit (the retry loop; reuses the + already-pinned observed_tree across every iteration, no re-pinning needed) +``` + +- **Crash before the DB CAS runs** (after a successful pin): the pin exists, + no DB row references its tree yet, or an older DB row still stands. This + is harmless — an orphaned pin protecting a tree that may never become + durable, reclaimable later by the deferred reconciliation pass, with no + correctness impact in the meantime. +- **Pin creation itself fails:** the coordinator does not proceed to the DB + CAS at all (same branch as a `capture_tree` failure — see "Snapshot + failure" below), so there is no window in which a durable DB row could + ever reference an unprotected tree. +- **DB CAS succeeds:** by construction, the tree it now references has been + ref-protected since *before* the CAS was even attempted — there is no + "commit succeeded, pin still pending" state to reason about, because + pinning is never deferred past the commit. +- **DB CAS conflicts (ordinary retry, not a crash):** the pin remains in + place (harmless, and still correct to reuse — the same `observed_tree` is + reused verbatim across retries), so no re-pinning is needed on retry. + +The DB CAS remains the mutation protocol's linearization point, unchanged; +the ref is durability infrastructure that exists entirely to make what the +CAS already decided stay resolvable, not a second decision-making mechanism. + +**Linked worktrees.** Since objects and the default ref namespace are +already shared across all worktrees of one repository (verified above), two +linked worktrees' coordinators pin into the *same* underlying object +database and ref namespace, scoped apart only by the `worktree_id` path +segment — no special-casing is needed for the linked-worktree case beyond +what worktree-scoped `WorktreeId` derivation already provides. + +**Why SCE refs do not interfere with user-visible Git operations.** +Verified experimentally: `refs/sce/mutation-cursor/**` never appears in +`git branch`, `git tag`, or plain `git log`/`git status` output (those only +enumerate `refs/heads/*`/`refs/tags/*` or the current `HEAD`); a plain +`git clone` of the repository does not transfer it (clone/fetch only +transfer refs matching the configured refspec, by default +`refs/heads/*`/`refs/tags/*`, not arbitrary custom namespaces); and a plain +`git push` only pushes the refspecs the user names, never "every local +ref." A tool enumerating with `git log --all`/`git rev-list --all` would see +these refs, matching how other established Git-based tooling (for example +`git notes`, GitHub's `refs/pull/*`) already uses dedicated `refs/` +namespaces for the same reason — this is a well-precedented, low-cost +pattern, not a new category of interference. + +Realized on-disk layout: + +```text +/sce/ +├── checkout-id (existing, unmodified) +├── checkout-id.lock (new, T01 — identity-creation lock, distinct from the runtime lock) +├── mutation-cursor.lock (new, T02 — per-worktree runtime lock) +└── tmp/ + └── index- (new, T03 — per-worktree, ephemeral) + + (no SCE-specific directory; T03 writes here directly) + +└── refs/sce/mutation-cursor// (new, T03/T04 — one ref per pinned tree, create-only this PR) +``` + +### Isolated Git snapshot: temporary index, and validated commands + +Validated experimentally against this repository's own Git version +(reproduced by the commands below): + +```text +env: + GIT_DIR= # resolves the correct worktree's HEAD, shared objects, and shared refs + GIT_INDEX_FILE=/sce/tmp/index- # private, per-invocation, never the real index +cwd: # the real working tree, for `add -A` to see real files + +1. `git rev-parse --verify --quiet HEAD` — unborn-HEAD probe (exit 1 = unborn) +2a. HEAD exists: `git read-tree HEAD` +2b. HEAD unborn: `git read-tree --empty` — explicitly initializes a valid, + empty Git index; see "Unborn + HEAD" below for why this + replaced an earlier, incorrect + design +3. `git add -A -- .` — stages the full current working-tree reality + (modified, added, untracked, deleted; + respects `.gitignore`) into the private index +4. `git write-tree` — writes the resulting tree into the + repository's normal, shared object database; + unchanged blob content already present is + never duplicated (Git checks object existence + before writing, exactly as it already does + for a normal `git add`/`git commit`) +5. `git update-ref refs/sce/mutation-cursor// ` + — pins the resulting tree (see "Snapshot + durability" above) +``` + +No `GIT_OBJECT_DIRECTORY`/`GIT_ALTERNATE_OBJECT_DIRECTORIES` env vars are +set at any point — this is a further simplification over the original +plan's private-store design, made possible once objects are written +straight into the normal store. + +**Unborn HEAD: `git read-tree --empty`, not a bare temp-index file.** An +earlier revision of this plan skipped step 2 entirely on an unborn `HEAD` +and assumed a freshly created, empty (zero-byte) temp-index file was +equivalent to an empty Git index. Verified experimentally: it is not. A +zero-byte file at the path named by `GIT_INDEX_FILE` is an invalid index — +`git add -A -- .` against it fails deterministically with `fatal: : +index file smaller than expected`, because Git's index format has a +required header a zero-length file cannot satisfy. There are two ways to +avoid ever presenting Git with such a file, and this plan adopts the more +explicit of the two: `git read-tree --empty`, run in place of `git read-tree +HEAD` when the unborn-HEAD probe fails, initializes a genuinely valid empty +index at the `GIT_INDEX_FILE` path (verified: Git creates it correctly, +`git add -A -- .` against it succeeds). Verified equivalent, and preferred +for that explicitness, is the alternative of never creating anything at the +index path at all and letting `git add -A -- .` run directly against a +genuinely nonexistent `GIT_INDEX_FILE` path (Git treats a missing index file +as "start from empty" for `add`, exactly as `read-tree --empty` does +explicitly) — both were confirmed experimentally to produce byte-identical +resulting trees for the same working-tree content, and an unborn repository +with no files at all produces a valid, zero-entry tree either way, whose +object ID this plan never hardcodes (Git computes it; the test only asserts +zero `ls-tree` entries). `git read-tree --empty` is nonetheless the design +this plan specifies, because it makes "start from empty" an explicit, +self-documenting step in the command sequence rather than an implicit +consequence of a file happening not to exist yet — the same reasoning +Requirement 2's brief itself gives (works independently of object format; +represents the correct base state for a repository with no commits without +relying on an unstated invariant about file-absence semantics). + +This is also why the temporary-index lifecycle (below) must never +pre-create an empty file at the reserved path: doing so would reintroduce +exactly the invalid-index failure this fix removes, on the unborn-HEAD path +specifically, since step 2b's `git read-tree --empty` needs to be the first +thing to ever touch that path, not a second write on top of an +already-existing zero-byte file. + +Confirmed by direct experiment: with only `GIT_DIR`/`GIT_INDEX_FILE` set, +the real repository's `.git/index`, `git status --porcelain`, `git diff`, +and `git diff --cached` are byte-identical before and after a snapshot +capture that includes staged, unstaged, and untracked changes +simultaneously; an ignored file is correctly absent from the resulting +tree's `ls-tree`; a deletion of a previously-committed file is correctly +reflected as absent from the resulting tree; `git update-ref` on the same +`(worktree_id, tree)` pair twice is a no-op; within one repository, a +distinct, genuinely unreachable tree (unique content, no ref, no branch, no +tag, no reflog entry — reflogs only ever record commit-ref history, and +this plan's loose trees are never committed or ref'd except via the SCE +pin, so they have no reflog entry to be protected by regardless) is +reclaimed by `git gc --prune=now` followed by `git prune --expire=now`, +while a second, differently-content tree pinned via +`refs/sce/mutation-cursor/**` survives the identical pass. `TreeId` stays +the opaque `String` `types.rs` already declares it as — no code depends on a +40-character SHA-1 length, so a SHA-256 repository +(`git init --object-format=sha256`) needs no special handling; the +git_snapshot tests include a best-effort SHA-256 case that skips gracefully +if the local Git build lacks the flag. + +Filenames with unusual characters and symlinks need no special handling: +`git add -A` already captures them correctly (symlinks as `120000` blobs; +unusual filenames via Git's own path handling). Submodules also need no +special handling: `git add -A` on a path Git recognizes as a nested +repository stages it as a single `160000` gitlink entry (the submodule's +current commit SHA), matching normal Git tree semantics — the mutation +cursor's tree snapshot correctly captures "this submodule was pinned to +commit X," not the submodule's own file contents, which is the same +information any real Git tree carries. + +Temporary-index cleanup uses an RAII guard around the unique +`/sce/tmp/index-` path — the guard only *reserves and tracks* +this unique path (generating the name, ensuring the parent `tmp/` directory +exists) and performs best-effort removal on drop, tolerant of an +already-removed or never-created file; it never creates the index file +itself. The path stays genuinely absent until the first `git read-tree` +call (`HEAD` or `--empty`) creates it, which is what makes both branches of +step 2 above produce a valid index rather than risking the invalid +zero-byte state described under "Unborn HEAD." This RAII structure still +fully serves its lifecycle purpose: a mid-capture panic or early return +still triggers cleanup of whatever Git did create at that path, and the +guard's own removal never being guaranteed under a hard process kill is +acceptable because the temp index is pure write-time scratch state — the +durable artifact is the pinned tree/blob objects already written to the +repository's normal object database, not the index. + +### `diff_trees`'s output type + +`diff_trees(before, after)` runs +`git diff --binary --full-index --no-ext-diff --no-textconv ` +with only `GIT_DIR` set (no object-directory environment at all, since both +trees already live in the normal store once pinned) and returns the raw +`Result` output. Git's `--binary` unified-diff format is +base85-encoded text, so the output is always valid UTF-8 even for binary +file changes. This is the smallest abstraction that satisfies Requirement 3: +the existing `cli/src/services/patch.rs::parse_patch(input: &str, +session_id: Option<&str>)` already parses exactly this `diff --git` +formatted text, so the next PR's Agent Trace integration can hand +`diff_trees`'s output straight to the parser this repository already has, +with no new parsing code and no intermediate structured-diff type invented +here. + +### `WorktreeId` derivation and `RuntimeBoundary` + +The coordinator, not the caller, derives `WorktreeId`: `coordinate()` +resolves `git_dir` from `repository_root`, acquires `WorktreeLock`, then +calls `get_or_create_checkout_id(&git_dir)` (T01's now concurrency-safe +version) and wraps the result as `WorktreeId(checkout_id)`. No +caller-supplied `WorktreeId` or `Boundary` value is ever accepted, +satisfying Requirement 4's "runtime callers should not be allowed to invent +a `WorktreeId`." + +A separate `RuntimeBoundary` type is warranted and lives in +`coordinator.rs`: `types::Boundary` intentionally carries no `ActorKind` +(that is a scope-*registration* concern, not a pure protocol transition +concern) and its `Flush` variant carries an explicit `worktree: WorktreeId` +the coordinator must not let a caller supply. `RuntimeBoundary` matches the +brief's sketch exactly — `Start`/`Advance`/`Close` each carry +`{ scope: ScopeId, event: EventId, actor_kind: ActorKind }`, `Flush` carries +nothing — and the coordinator converts it into a `types::Boundary` (dropping +`actor_kind`, which is consumed earlier for `register_scope`) plus the +internally-derived `WorktreeId` before calling `protocol::prepare`. + +### Runtime identity contract: `(ScopeId, EventId)` namespace requirements + +`spec/mutation_cursor.md`'s "Event identity" section already states the +production requirement precisely: "Hook replay identity is scoped by +`ScopeId` and `EventId` through `EventKey`. The real implementation must +provide an equivalent uniqueness guarantee. If hook IDs are not unique per +scope, the database key must include the actual delivery namespace, such as +worktree, harness, session, and hook ID." `types::EventKey { scope_id, +event_id }` is exactly this replay key, unmodified — and it has no +`actor_kind` field at all, so `ActorKind` is, by construction, never part of +replay identity; it flows only into `register_scope`. + +`RuntimeBoundary` inherits this contract as a documented, load-bearing +precondition rather than something the coordinator can verify or enforce +for its caller — the coordinator has no way to know whether a given +`ScopeId`/`EventId` pair is genuinely unique to one logical hook delivery, +since that fact depends entirely on the harness's own ID scheme. The +contract, to be stated on `RuntimeBoundary`'s doc comment and in a future +harness-adapter's own documentation: + +- `RuntimeBoundary` accepts already-canonical runtime identities. + `(ScopeId, EventId)` must uniquely identify exactly one logical hook + delivery, for the lifetime of that scope, for replay purposes. +- A harness adapter must not blindly forward a raw native hook/event ID + unless the harness already guarantees that ID is unique within the given + scope. If it does not, the adapter — not the coordinator, not + `protocol.rs`, not `store.rs` — owns constructing a canonicalized ID, for + example conceptually `ScopeId(":")` and, if + the harness's own event IDs are not already scope-unique, + `EventId("::")`. This plan + does not choose or implement an encoding, since no harness adapter exists + yet to canonicalize for. +- Two different logical hook deliveries colliding under the same + `(ScopeId, EventId)` is invalid adapter behavior, not a case the protocol + or coordinator can detect or correct — `store.rs`'s replay-dedup logic + will (correctly, given its own contract) treat them as the same delivery. +- Replaying the same logical delivery — the harness genuinely re-sending an + event, for example after a retry — must always produce the identical + canonical `(ScopeId, EventId)` pair, or replay-dedup silently breaks. + +### Worktree/scope materialization ordering + +Matches `context/cli/mutation-trace-protocol.md`'s "Runtime scope +materialization"/"Runtime worktree materialization" sections and +Requirement 5, updated for the pin-before-DB-write ordering above: + +```text +resolve WorktreeId under the lock + ↓ +capture the ONE Git snapshot for this invocation → observed_tree + │ + ├── failure → see "Snapshot failure" below: loads durable state + │ fresh, AFTER this failure — never a state read + │ earlier in this diagram, because there is none + │ + └── success + ↓ + pin_tree(worktree, observed_tree) (blocking precondition to everything below; + │ a pin failure is handled identically to a + │ capture failure — see "Snapshot failure") + ↓ +store.initialize_worktree(worktree, observed_tree) (idempotent; no-op if the row already exists) + ↓ +store.register_scope(scope, worktree, actor_kind) (only for Start/Advance/Close; idempotent; Err on identity conflict) + ↓ +retry loop: load → recover-if-needed → prepare/commit +``` + +This revision removes an earlier version of this diagram's leading +`existing_worktree = store.load_worktree(...)` probe, taken *before* +snapshot capture, whose only purpose was deciding the snapshot-failure +branch below. That probe was itself a race: another caller could +materialize the worktree row *after* the probe ran but *before* (or during) +this invocation's own capture attempt — a caller that bypasses the advisory +`WorktreeLock` entirely is exactly the scenario the lock cannot see, which +is why CAS remains the authoritative correctness mechanism throughout this +plan. A failing invocation basing its bootstrap-vs-taint decision on that +stale `None` would then durably under-report the failure +(`persisted_taint: false`) even though a worktree row existed by the time +the decision was actually made. The corrected flow makes exactly one durable +read on the failure path, taken after the failure, described in full under +"Snapshot failure" below; the success path never reads durable worktree +state before `initialize_worktree`, whose own idle-insert semantics already +handle "does a row exist yet" correctly without a separate probe. + +Capturing the snapshot *before* `initialize_worktree` is what makes AC1's +"first observation" guarantee fall out of the pure protocol with no special +case: `initialize_worktree` sets `cursor_tree = observed_tree`, and this +same `observed_tree` becomes the `Start` attempt's `after_tree` +(`prepare`'s `observed_tree` parameter) — `before_tree == after_tree`, so +`protocol.rs`'s existing `tree_changed` gate is already `false` and no +`MutationEvent` is materialized. No coordinator-side "is this the first +observation" branch exists or is needed. + +A scope-identity conflict (`register_scope` returning `Err` because an +existing scope's `worktree_id`/`actor_kind` disagrees with this boundary) is +propagated as `CoordinateError::ScopeIdentityConflict` before the retry loop +runs; the coordinator never overwrites a scope's identity. Per the +create-only pin policy above, the pin already created for this invocation's +`observed_tree` is left in place even on this abort path. + +### Protocol execution and `AttemptId` + +The coordinator calls `protocol::prepare`/`protocol::commit` exactly as +`protocol.rs` exposes them, and reproduces none of their lifecycle logic. +`AttemptId` is generated fresh (`Uuid::new_v4()`) on each retry-loop +iteration and never persisted: `store::WorktreeProjection::into_protocol_state` +always sets `attempts: BTreeMap::new()` on load (`store.rs`'s own +documented non-goal — no `mutation_trace_attempts` table exists), so +`prepare`'s "already underway" guard (`state.attempts.get(&attempt)`) always +sees `None` regardless of whether the ID is reused or freshly generated each +iteration; generating fresh each iteration is simplest and avoids any doubt +about cross-iteration interaction. + +### CAS retry: what it protects, and why it's still needed under the lock + +`MutationTraceStore::commit`'s own `execute_transactional_cas_batch` already +retries `Busy`/`BusySnapshot` (SQLite/Turso transient contention) internally, +via `resolve_query_retry_policy`, completely opaque to the coordinator — the +coordinator never sees or reacts to that class of failure directly. What the +coordinator's own retry loop reacts to is `CasResult::Conflict`: a settled, +non-error `Ok` outcome meaning the worktree's guard `UPDATE ... WHERE +revision = ?` affected zero rows, because some other committer already +advanced the revision this attempt was prepared against. + +Under `WorktreeLock`, no other `coordinate()` invocation on this exact +worktree can be mid-flight, so in ordinary single-host operation a `Conflict` +at this point should not occur. The retry loop exists anyway because CAS, +not the lock, is the protocol's actual correctness mechanism (per the +brief's own framing) — the lock is real-world defense-in-depth against a +class of writer the lock cannot see: OS advisory locking limitations on some +filesystems, a caller that bypasses `WorktreeLock` entirely, or (out of +scope here, but worth naming) a future multi-host deployment where the lock +is host-local but the DB file is not. The whole `coordinate()` pipeline — +from the fresh `store.load_worktree` at the top of each retry-loop iteration +through the recovery check and the boundary's own `prepare`/`commit` — is +wrapped in one bounded loop (`MAX_CAS_RETRY_ATTEMPTS = 5`, no backoff: a +`Conflict` under the lock is already anomalous, and reload-and-recompute is +cheap, so sleeping before retrying buys nothing). Every iteration reuses the +exact same, already-pinned `observed_tree: TreeId` captured once at the top +of `coordinate()` — the loop only re-reads durable state and re-derives the +transition from it, never re-captures or re-pins the worktree. Exhausting +the bound returns `CoordinateError::CasConflictExhausted { attempts }`. + +The snapshot-failure taint path below reuses this exact policy +(`MAX_CAS_RETRY_ATTEMPTS`, no backoff) for its own, narrower retry loop — +one shared constant, one shared rationale, applied in two places. + +### Recovery: which triggers are reachable, and why it's two CAS transactions + +`spec/mutation_cursor.md`'s "Failure and durability boundary" section states +recovery as one action producing one durable transition: +"snapshot current worktree → establish current tree as the new cursor +baseline → produce no evidence → abandon every active scope (only for +taint/externalTaint) → commit recovery to DB → clear the taint/failure state." +`protocol::recover` (already implemented, unmodified) is exactly this one +pure transition. + +`WorktreeProjection::into_protocol_state` always sets `external_taint: +BTreeSet::new()` on load — `store.rs`'s documented, deliberate non-goal, +because `external_taint` is never DB-authoritative. This means, within this +PR's runtime path, recovery can only ever be triggered by the two +DB-observable flags: `tainted` (`SnapshotFailure`) or `needs_rebaseline`. +`external_taint`-triggered recovery is unreachable code in this PR — it +becomes live only once the follow-up PR wires `protocol::database_failure` +and a filesystem external-taint marker (see Requirement 10, below). Both +reachable modes are implemented and tested in T04: `needs_rebaseline` +recovery preserves live scopes (only the skipped interval is ambiguous), +taint recovery abandons them (per `protocol::recover`'s own guarded +behavior, unmodified). + +Recovery and the triggering boundary are **two sequential DB CAS +transactions**, not one, and not a single combined `DurableTransition`: +`DurableTransition::between` hard-`bail!`s unless the worktree's revision +advances by *exactly one* between `before` and `after` +(`cli/src/services/mutation_trace/store.rs:304-310`). Recovery alone +advances revision by one; the triggering boundary, if accepted, advances it +by one more. A `before`-to-after-both-steps diff would require an advance of +two, which `between` explicitly rejects — so recovery must be persisted +(`store.commit`) before the triggering boundary's own `prepare`/`commit` is +even evaluated against the resulting state. Both transactions happen inside +the same retry-loop iteration, under the same lock hold, and both reuse the +one captured, already-pinned `observed_tree`: recovery uses it as its +rebaseline target, and because the triggering boundary is then prepared +against the *already* just-rebaselined cursor (which now equals +`observed_tree`), a boundary processed in the same invocation that just +recovered naturally observes `before_tree == after_tree` and emits no +evidence for the just-discarded interval — again, no special-case +coordinator logic, purely a consequence of reusing one snapshot for both +steps. + +If recovery's own CAS conflicts, the whole retry-loop iteration restarts +(reload → re-check recovery need → recompute), governed by the same +`MAX_CAS_RETRY_ATTEMPTS` bound as boundary-CAS conflicts — they are the same +phenomenon (stale reload needed) and share one counter. + +### Snapshot failure: bootstrap case, and correct CAS retry for the taint transition + +`SnapshotCapture::capture` or `SnapshotCapture::pin` returning `Err` are +treated identically (see "Crash ordering" above — pinning is a blocking +precondition, so a pin failure is a snapshot failure in every way that +matters here). Unlike an earlier revision of this plan, the coordinator does +**not** consult any durable state read before this failure — see +"Worktree/scope materialization ordering" above for why a pre-capture probe +is itself a race. Instead, the very first thing the failure handler does is +its own fresh read, and that single read decides both the bootstrap-vs-taint +branch *and* doubles as the first iteration of the taint-retry loop: + +```text +loop (bounded MAX_CAS_RETRY_ATTEMPTS): + state = store.load_worktree(worktree, None, None) -- fresh every iteration, + -- including the first, and + -- always AFTER the capture/pin + -- failure, never before it + match state { + None => + // No durable worktree row exists, even now, freshly checked + // after the failure. There is no durable state to taint and no + // baseline TreeId to fabricate one from — protocol::taint itself + // guards on worktree existence and would be a no-op anyway. + return SnapshotFailure { persisted_taint: false, source } + // No durable write is made. Per the brief: do not fabricate a TreeId. + Some(loaded) => { + tainted_state = protocol::taint(&loaded, &worktree) + match DurableTransition::between(&loaded, &tainted_state, &worktree)? { + None => + // taint() itself no-opped: the worktree was already + // tainted (durable state already reflects a failure), + // or is already at revision u64::MAX. Either way there + // is nothing new to persist, and the *current* durable + // state already answers "is this worktree tainted" — + // read it back directly rather than trusting a stale + // local flag. + return SnapshotFailure { + persisted_taint: loaded.worktrees[&worktree].tainted, + source, + } + Some(transition) => match store.commit(&transition)? { + CasResult::Applied => + return SnapshotFailure { persisted_taint: true, source } + CasResult::Conflict => continue -- reload next iteration; this is + -- exactly the "another caller + -- materialized/advanced the + -- worktree concurrently" case + } + } + } + } +// loop exhausted without ever reaching Applied, an already-tainted no-op, or None +return SnapshotFailure { persisted_taint: false, source } +``` + +Because the `None` branch is evaluated fresh on *every* iteration — not only +the first — this also directly answers the race the review specifically +named: a worktree that did not exist when this invocation began its Git +snapshot capture, but that another caller materializes concurrently while +that capture is still in flight (for example a caller that bypasses +`WorktreeLock`, since the lock only serializes cooperating callers against +each other), is still found and correctly tainted, because the failure +handler's read happens strictly after the capture attempt has already +failed, giving any such concurrent writer every opportunity to have already +landed its row. + +`persisted_taint` therefore reports **whether the durable worktree state is +(now) tainted**, not merely "did this exact call perform a write" — this is +both simpler to reason about (one boolean, one meaning) and more useful to a +caller than distinguishing "we tainted it" from "someone else already had," +since both mean the same thing for the caller's purposes: the durable state +honestly reflects the failure. `persisted_taint` is `false` in exactly two +cases: the bootstrap case (no row exists on the fresh, post-failure read), +and taint-CAS exhaustion (every attempt in the bounded loop conflicted). +Both report through the same `CoordinateError::SnapshotFailure` variant, +distinguished only by the `source` error chain's message — a caller only +needs to know "was the failure durably recorded," and a second public enum +case to distinguish *why* it was not would leak an implementation detail +(this codebase's existing `RepositoryIdentityResolutionError` precedent for +typed-error precision is about *distinguishable outcomes a caller must react +to differently*, and bootstrap-vs-exhaustion do not differ in the required +reaction: "the taint was not recorded, treat this worktree as if the +model's `database_failure` case applied"). + +DB busy/transient retry inside this loop's own `store.commit` call stays +exactly where it already lives — inside `execute_transactional_cas_batch`, +via `resolve_query_retry_policy` — never duplicated here, matching the main +retry loop's own separation of concerns. + +The future filesystem external-taint marker remains responsible for making +the still-deferred DB-unavailable case (Requirement 10, below) crash-safe; +this fix only corrects the already-in-scope `SnapshotFailure`/CAS-conflict +interaction, which does not depend on that marker existing. + +### Requirement 10: database unavailability and external taint stay deferred + +`spec/mutation_cursor.md` distinguishes snapshot failure (recorded via +`taint`, a DB-backed field) from database unavailability +(`databaseFailure`, which changes only the abstractly-modeled +`externalTaint` — explicitly "not a database row"). `store.rs` already +established, in the prior plan, that this distinction is real in the Rust +implementation too: `DurableTransition::between` returns `Ok(None)` for a +`database_failure`-only transition, so `store.commit` is never even called +for it, and no `external_taint` column or table exists. + +Building the filesystem marker `externalTaint` conceptually models would be +new, non-trivial infrastructure (a durability signal that must itself +survive process crashes and be readable before the DB is even opened) that +nothing in the existing code makes unavoidable: when `store.load_worktree` +or `store.commit` returns a genuine `Err` after the DB layer's own retries +are exhausted (this applies both to the main boundary-processing loop and +to the snapshot-failure taint-retry loop above), `coordinate()` simply +propagates it as `CoordinateError::Other`, leaving no trace anywhere. This +is a correct, if intentionally incomplete, subset of the model's behavior — +"durable DB protocol state remains unchanged" holds; "externalTaint +contains the worktree" does not yet. `protocol::database_failure` is not +called anywhere in this PR, matching its current (unmodified) unwired +status. This is recorded as the follow-up PR's first concern (see below), +exactly as the brief anticipated as the likely outcome. + +### `CoordinateOutcome` and `CoordinateError` + +```rust +pub struct CoordinateOutcome { + pub worktree_id: WorktreeId, + pub observed_tree: TreeId, + pub revision: u64, + pub evaluation: protocol::CommitEvaluation, + pub mutation_event: Option, +} + +pub enum CoordinateError { + SnapshotFailure { persisted_taint: bool, source: anyhow::Error }, + ScopeIdentityConflict(anyhow::Error), + CasConflictExhausted { attempts: u32 }, + LockAcquisition(anyhow::Error), + Other(anyhow::Error), +} +``` + +Unchanged from the original plan except in the meaning of +`SnapshotFailure.persisted_taint`, corrected above. This is the smallest +return/error shape that satisfies every requirement without leaking store +internals (`DurableTransition`, `CasResult`, SQL row shapes never appear in +either type). `revision`/`evaluation`/`mutation_event` all come directly +from the *triggering boundary's* own `protocol::commit` result +(`CommitOutcome`), not from any intermediate recovery step. + +### Dependency injection for deterministic tests + +`SnapshotCapture` (`fn capture(&self, repository_root: &Path) -> +Result`, `fn pin(&self, worktree_id: &WorktreeId, tree: &TreeId) -> +Result<()>`) is the one seam this plan introduces for determinism: +`GitSnapshotService` implements it for production use, and T04's tests use +a fake, call-counting implementation to prove "exactly one `capture`, at +most one `pin`, per invocation, even across CAS retries" (AC8) without +needing real concurrent Git processes. The lock is *not* faked — +`worktree_lock::tests`/`checkout::tests` and T05's contention test exercise +real `std::fs::File` locks against real temp directories, because proving +actual OS-level exclusion is the point. The DB is not faked either, matching +`store.rs`'s own established precedent +(`RepositoryAgentTraceDb::new_at(&temp_path)` against a real temp-file +database, not an in-memory or mocked store). + +Proving the taint-CAS-retry path (AC11) deterministically follows the same +pattern the existing `store.rs` test suite already uses for CAS-conflict +scenarios (for example `competing_prepared_attempts_the_second_to_commit_is_rejected_by_cas`): +start from one loaded `ProtocolState`, commit a *second*, independently +derived transition directly through `store.commit` first (simulating "some +other committer already advanced this worktree"), then run the taint-retry +logic starting from the now-stale first state — its first commit attempt +observably conflicts, forcing the loop to reload and retaint, without +needing thread synchronization. Proving exhaustion uses a background thread +that keeps advancing the worktree's revision via direct, independent +`store.commit` calls for the duration of the bounded loop, guaranteeing +every attempt conflicts. + +### What changes vs. what's new + +`cli/src/services/checkout/mod.rs` is now genuinely **modified**, not +purely additive: `get_or_create_checkout_id`'s body changes (T01), though +its signature and the checkout-ID file format do not. `protocol.rs`, +`types.rs`, `store.rs`, `repository_identity/`, and `agent_trace_storage/` +are all read but not modified — `agent_trace_storage` benefits from T01 +automatically, with no change of its own required. No schema migration is +added: `store.rs`'s existing five tables and its +`initialize_worktree`/`register_scope`/`load_worktree`/`commit` API are +sufficient for everything this coordinator needs. + +## Files + +New: + +- `cli/src/services/mutation_trace/worktree_lock.rs` — per-worktree runtime + OS advisory lock (T02). +- `cli/src/services/mutation_trace/git_snapshot.rs` — isolated Git snapshot + capture, ref pinning, and `diff_trees` (T03). +- `cli/src/services/mutation_trace/coordinator.rs` — `RuntimeBoundary`, + `CoordinateOutcome`, `CoordinateError`, `SnapshotCapture`, the internal + protocol-integration pipeline, and the public lock-wrapped `coordinate()` + entrypoint (T04, T05). +- `cli/src/services/mutation_trace/runtime_tests.rs` — cross-module + linked-worktree, cross-caller checkout-identity, and end-to-end + failure/recovery integration tests (T06). + +Modified: + +- `cli/src/services/checkout/mod.rs` — `get_or_create_checkout_id` gains an + internal identity-creation lock plus a crash-safe temp-file-and-rename + write sequence (T01); no signature change. +- `cli/src/services/mutation_trace/mod.rs` — register the four new modules + (`pub mod git_snapshot; pub mod worktree_lock; pub mod coordinator; + #[cfg(test)] mod runtime_tests;`), consistent with the module's existing + `#[allow(dead_code)]` precedent for code not yet wired to a command/hook. + +Docs (context sync, not authored in this plan's tasks): + +- `context/cli/mutation-trace-protocol.md`, `context/cli/mutation-trace-store.md`, + `context/cli/checkout-identity.md`, `context/overview.md`, + `context/context-map.md` — see "Context sync" above. + +Cargo/dependency changes: none. + +## Testing strategy + +- **Pure unit tests** (no filesystem/DB/lock): none new — `protocol.rs`'s + existing pure-function tests already cover `prepare`/`commit`/`recover`/ + `taint`/`abandon`; this plan only adds imperative-shell code around them. +- **Git integration tests** (`git_snapshot.rs`, real temp Git repos): index + preservation under staged+unstaged+untracked simultaneously, ignored-file + exclusion, tracked-file deletion; unborn `HEAD` — both with a file present + (asserting the resulting tree contains it) and with the working tree + completely empty (asserting a valid, zero-entry tree, its object ID never + hardcoded) — driving the real `capture_tree` command sequence against a + freshly `git init`ed repository with no commit, proving `git read-tree + --empty` (not a bare temp-index file) is what makes `git add -A -- .` + succeed; snapshot durability after temp-index deletion; **pinned-tree + survival across `git gc --prune=now` and `git prune --expire=now`**, + proven with a negative control that captures a *second, distinct-content* + tree in the same repository, deliberately leaves it unpinned and + unreachable from any ref/branch/tag/reflog, and asserts that same + aggressive `gc`/`prune` pass reclaims it while the first, pinned tree + survives — never a same-repository comparison between two copies of + identical content, since Git's content-addressing would make those the + same object and unpinning one would (correctly) leave it pinned by the + other; `pin_tree` idempotence; `diff_trees` correctness; best-effort + SHA-256 case. Prove: `capture_tree` faithfully represents the current + worktree without ever touching the real index, always starts from a valid + Git index regardless of whether `HEAD` exists, and a pinned snapshot's + durability holds against Git's own reachability analysis (proven against a + genuinely unreachable control object, not against "the temp index is + gone" or an object the test itself cannot actually make unreachable). +- **Store/DB integration tests** (`coordinator.rs`, real temp-file + `RepositoryAgentTraceDb`): AC1–AC5, AC10, AC11 — first observation, + exclusive mutation, no-op observation, replay, close attribution, + contention, both recovery modes, snapshot-failure with immediate success, + snapshot-failure taint surviving a losing CAS and committing on retry, + snapshot-failure taint exhaustion never claiming `persisted_taint: true`, + and the bootstrap snapshot-failure case. Prove: the coordinator drives the + existing store/protocol APIs to produce exactly the durable outcomes the + Quint model specifies, including under contention on the taint path. +- **Concurrency tests** (`checkout::mod.rs` T01, `worktree_lock.rs` T02, + `coordinator.rs` T05): concurrent first-time `get_or_create_checkout_id` + callers on one `git_dir` converge on one ID; a second `try_lock()` on the + same worktree-lock path observably blocks/fails while the first holds it + and succeeds immediately after `Drop`; two real threads calling + `coordinate()` against the same worktree serialize (the second's critical + section provably starts only after the first's ends). Prove: both locks + are OS-level exclusion, not lockfile-existence checking, and the + checkout-identity fix actually closes the race for every caller, not just + the coordinator's. +- **Checkout-identity crash-safety tests** (`checkout::mod.rs` T01): the + canonical `checkout-id` path remains absent through every step of the + write sequence up to and including a successful `sync_data()`, and only + becomes present, with complete content, once `std::fs::rename` returns — + exercised by driving the internal persistence helper up to (but not past) + the rename call and asserting the canonical path's absence, then letting + it complete and asserting the canonical path's complete content; an + abandoned `checkout-id.tmp-*` file (simulating an interruption after + `sync_data()` but before `rename`) does not prevent, or alter the outcome + of, a subsequent `get_or_create_checkout_id` call, which still converges + on one complete, valid ID. Prove: AC13 — the canonical path is never + observable as partially written, and orphaned temp files are inert. +- **Linked-worktree tests** (`runtime_tests.rs`, real `git worktree add`): + two linked worktrees resolve distinct `checkout_id`/`WorktreeId` and + distinct runtime-lock paths, concurrent `coordinate()` calls on the two + worktrees do not block each other, both share one repository-scoped DB, + and a tree pinned from one worktree resolves correctly when queried via + the other worktree's `GIT_DIR`. Prove: AC9 end to end through the public + API. +- **Cross-caller checkout-identity test** (`runtime_tests.rs`): a direct + `agent_trace_storage` resolution and a `coordinate()` call against the + same `repository_root`, run concurrently on first-ever resolution, observe + the identical checkout ID. Prove: AC12's convergence guarantee holds + across module boundaries, not only within `checkout::mod.rs`'s own test + suite. +- **Failure/recovery tests** (`coordinator.rs`, `runtime_tests.rs`): CAS + conflict via two prepared-but-not-yet-committed states from the same + revision (proving reload+recompute+no-second-snapshot, and no re-pin); + snapshot failure with and without a prior worktree row; snapshot-failure + taint retried through a losing CAS to a successful one, and exhausted + entirely; snapshot failure against a worktree row that a second, direct + `store.initialize_worktree` call materializes only *after* the injected + capture failure begins (using a fake `SnapshotCapture` whose `capture` + performs that materialization itself, immediately before returning `Err`), + proving the failure handler's fresh, post-failure read finds and taints it + rather than basing its decision on any earlier state; a full two-invocation + failure-then-recovery cycle via the public API. Prove: AC8, AC10, AC11 in + both unit and end-to-end form. +- **Lock staleness**: `worktree_lock::tests` documents (via a test that + opens, writes bytes to, and closes the lock file *without* ever calling + `.lock()`, then proves a subsequent real `WorktreeLock::acquire` on the + same path succeeds immediately) that a leftover lock file with no active + OS lock held against it never blocks a future caller — the OS lock, not + the file's existence, is the ownership signal. The same property is + exercised for `checkout-id.lock` in `checkout::mod.rs`'s own tests. + +## Failure matrix + +| Failure | Durable state mutation? | External marker? | Retry? | Result | +| --- | --- | --- | --- | --- | +| checkout-identity lock (T01) | No | No | No (blocking, no timeout — see rationale above) | Blocks until the current holder releases; cannot itself fail except on genuine I/O error, surfaced as `Err` | +| process crash between `checkout-id.tmp-*` creation and rename (T01) | No — the canonical `checkout-id` path is unaffected; either absent (fresh checkout) or unchanged (identity already existed) | No | No — not a retryable operation from the crashed process's perspective; a later, fresh call retries the whole sequence from scratch | Orphaned `checkout-id.tmp-*` file, permanently harmless (see Design decisions); the next `get_or_create_checkout_id` call on this `git_dir` succeeds normally | +| mutation-cursor lock acquisition (timeout) | No | No | No (single bounded polling attempt per call) | `CoordinateError::LockAcquisition`, no writes attempted | +| Git snapshot capture or pin, worktree row exists on the fresh post-failure read | Yes — a `protocol::taint` transition, retried under the shared bounded CAS-retry policy until `Applied` or exhaustion | No | Yes — up to `MAX_CAS_RETRY_ATTEMPTS`, reload/retaint/recommit each time, each iteration re-reading durable state fresh | `CoordinateError::SnapshotFailure { persisted_taint: true }` once `Applied`, or `{ persisted_taint: false }` after exhaustion | +| Git snapshot capture or pin, no worktree row on the fresh post-failure read (bootstrap, or not yet materialized by any concurrent writer) | No | No | No | `CoordinateError::SnapshotFailure { persisted_taint: false }` | +| DB busy/transient (`Busy`/`BusySnapshot`) | N/A until resolved | No | Yes — inside `execute_transactional_cas_batch`, opaque to the coordinator; applies identically to the main loop and the taint-retry loop | Transparent success, or `Err` propagated after the DB layer's own retries are exhausted | +| CAS conflict (semantic, guard affected 0 rows), main loop | No | No | Yes — coordinator's bounded retry loop, reloads state, reuses the same captured and already-pinned snapshot | Eventually `Applied`, or `CoordinateError::CasConflictExhausted` | +| DB unavailable (`Err` surviving DB-layer retries) | No | No (out of scope this PR) | No | `CoordinateError::Other`; `external_taint` not recorded (deferred, see follow-up PR) | +| recovery's own CAS conflict | No | No | Yes — same outer loop, whole iteration re-executes reusing the same captured and already-pinned snapshot | Eventually `Applied`, or `CoordinateError::CasConflictExhausted` | +| scope identity conflict (`register_scope`) | No | No | No | `CoordinateError::ScopeIdentityConflict`, no protocol transition attempted; this invocation's pin (if already created) is left in place per the create-only policy | + +## Validation + +- `nix flake check` — runs `checks..cli-tests` (all new tests + above), `checks..cli-clippy` (pedantic/warnings-denied), `checks..cli-fmt`, + `checks..mutation-trace-quint-connect` (confirms the unmodified + `protocol.rs`/`types.rs` still refine `spec/mutation_cursor.qnt` — this PR + changes neither, so this should pass unchanged), `checks..pkl-generated`, + and `checks..workflow-actionlint`. +- `nix run .#pkl-check-generated` — lightweight post-task baseline per + `context/patterns.md`; this PR touches no generated-config surface, so this + confirms no accidental drift. +- No new or modified Quint model file, so no separate `nix run .#... quint` + verification step is needed beyond the existing `mutation-trace-quint-connect` + check above. + +## Open questions + +None. Every architecture question the brief enumerated, plus every +correctness issue raised in the PR #244 review that produced this revision, +is resolved above with a concrete, code-grounded (and, for the Git-level +claims, experimentally verified) decision. See "Final quality check" below +for a direct, itemized accounting. + +## Final quality check + +1. **Can `checkout-id` ever be exposed half-written after an SCE process + crashes during creation?** No, once T01 lands. The canonical path is + written by generating the ID, writing it in full to a uniquely named + temp file, syncing that file's content, and only then atomically + renaming it into place — a reader of the canonical path observes either + its prior state (absent, for a fresh checkout) or the complete value, + never a partial one. +2. **What exact atomic persistence sequence prevents that?** + `OpenOptions::create_new(true)` on `checkout-id.tmp-` → + `write_all` the complete ID → `File::sync_data()` → `std::fs::rename` + into `checkout-id` → (Unix only, best-effort) sync the parent directory. + Verified experimentally: the canonical path is absent at every point + before the rename call and contains the complete value immediately after + it; an orphaned temp file left by an interruption before the rename does + not affect the canonical path's content. +3. **Can concurrent checkout-ID callers still return different IDs?** No, + once T01 lands — for any caller, not only the coordinator. + `get_or_create_checkout_id` itself acquires `/sce/checkout-id.lock`, + re-checks for an existing ID under the lock, and only then generates and + writes one via the atomic sequence above — the fix lives in the + primitive, so every caller (including `agent_trace_storage`, unmodified) + benefits automatically. +4. **Does snapshot failure use any DB state loaded before the snapshot + attempt?** No, once T04 lands. An earlier revision of this plan read + worktree existence before capture, purely to decide the bootstrap-vs-taint + branch on failure; that read has been removed. The only durable read the + snapshot-failure path performs happens inside the failure handler itself, + after the failure, and is repeated fresh on every taint-retry iteration. +5. **If another caller creates the worktree row while snapshot capture is + running, will the failing invocation find and taint it?** Yes. Because + the failure handler's first (and every subsequent) read happens strictly + after the capture attempt has already failed, any concurrent writer — + including one that bypasses `WorktreeLock` entirely, which the lock + cannot itself prevent — has already had the opportunity to land its row + by the time that read runs. +6. **How is a private Git index initialized when `HEAD` exists?** + `git read-tree HEAD` against the reserved, not-yet-created + `GIT_INDEX_FILE` path — Git creates a valid index populated from `HEAD`'s + tree. +7. **How is it initialized when `HEAD` is unborn?** `git read-tree --empty` + against that same not-yet-created path — Git creates a valid, genuinely + empty index. An earlier revision of this plan instead skipped this step + entirely on an unborn `HEAD` and assumed a bare temp-index file was + equivalent; verified experimentally that a zero-byte file at + `GIT_INDEX_FILE` is not a valid index and fails `git add -A -- .` + deterministically (`fatal: ... index file smaller than expected`). +8. **Can any path through `capture_tree()` pass a zero-byte invalid index to + `git add`?** No. The RAII guard around the temp-index path only reserves + and tracks the unique filename; it never creates a file there itself. + Both branches of index initialization (question 6 and question 7) are + Git's own `read-tree` calls, which always produce a structurally valid + index — there is no code path that reaches `git add -A -- .` against a + path Git has not already validly initialized. +9. **Does the unborn-HEAD integration test run the actual production command + sequence?** Yes — against a freshly `git init`ed repository with no + commit, driving `capture_tree` itself (`rev-parse --verify --quiet HEAD` + → `read-tree --empty` → `add -A` → `write-tree`), not a hand-rolled + substitute sequence, covering both an unborn repository with a file + present and one with no files at all. +10. **Does the test avoid hardcoding the SHA-1 empty-tree ID?** Yes. The + no-files case asserts a zero-entry `git ls-tree` result and lets Git + compute the tree object ID itself; no test or production code embeds the + canonical `4b825dc6...` (or SHA-256 equivalent) value. +11. **Can two identical trees in the same Git object database meaningfully + be "one pinned and one unpinned"?** No — this was the original plan's + negative-control mistake, now corrected. Git is content-addressed: + identical content is the same object, so pinning either "copy" pins + both. The corrected design captures two trees with *distinct* content in + the same repository, pins only one, and leaves the other genuinely + unreachable (no ref, branch, tag, or reflog entry). +12. **What object does the GC negative-control test actually prune?** The + second, distinct-content tree from question 11 — never a second capture + of identical content, and never the pinned tree itself. +13. **What ref protects the positive-control snapshot?** + `refs/sce/mutation-cursor//`, created by + `pin_tree` before the test's `gc`/`prune` pass runs. +14. **Does the GC test prove Git reachability rather than relying on + incidental object lifetime?** Yes. The control tree is deliberately + unreachable from every ref-like source Git's GC considers (no branch, no + tag, no reflog — reflogs only ever record commit-ref history, and these + loose, never-committed trees have none regardless), so its reclamation + under `git gc --prune=now`/`git prune --expire=now` demonstrates the test + environment can actually prune an unreachable object; the pinned tree's + survival under the identical pass then demonstrates the SCE ref, not + incidental timing or auto-GC's default grace period, is what protects + it. +15. **Can a durable `TreeId` ever become unprotected from Git GC before its + DB row is committed?** No. Pin creation + (`git update-ref refs/sce/mutation-cursor//`) is a + blocking precondition to the DB CAS, not a follow-up step after it — the + coordinator never attempts the CAS without a successful pin already in + place, so there is no window in which a durable DB row could reference + an unprotected tree. +16. **What Git refs can leak after crashes or CAS loss?** A + `refs/sce/mutation-cursor//` ref (and the objects + it protects) for a tree that never became the worktree's durable cursor + and never appears in any `mutation_trace_events` row — for example a + rejected CAS attempt's observation, or a pin created just before a crash + that never reached the DB CAS at all. This PR never deletes such a ref + (see "Pin lifecycle: create-only in this PR"). +17. **Is ref growth bounded or unbounded before reconciliation exists?** + Unbounded over the repository's lifetime — this plan's original + "bounded, linear" description was imprecise and is corrected above. + Growth is linear *in hook-invocation volume*, which is itself unbounded + over a repository's life; nothing in this PR's own scope ever removes a + pin. +18. **What exact PR must add ref reconciliation, and what exact PR must add + filesystem external-taint durability?** `mutation-cursor-ref-reconciliation` + and `mutation-cursor-external-taint` respectively — see "Runtime + completion sequence" under Follow-up PR. They may combine into one PR if + the result stays coherent, but neither is this PR's own task stack. +19. **Are both required before production harness wiring?** Yes, explicitly. + The harness-adapter follow-up (step 4 of the runtime completion + sequence) is the point at which real, sustained invocation volume first + exists; both deferred gaps are latent-but-harmless in this PR's + standalone form and become live risks only once that traffic exists. +20. **Can `SnapshotFailure.persisted_taint` ever be `true` after a failed + CAS?** No, once T04 lands. It is `true` only after `store.commit` + returns `CasResult::Applied` for a taint transition, or when a fresh + reload shows the worktree already durably tainted — never merely after + an attempt. +21. **How is snapshot-failure CAS conflict retried?** By the same bounded, + no-backoff `MAX_CAS_RETRY_ATTEMPTS` policy the main coordinator loop + uses: reload the worktree, recompute `protocol::taint` against the fresh + state, attempt `store.commit` again, up to the bound — the same loop + whose first iteration also decides the bootstrap-vs-taint branch (see + question 4). +22. **What exact uniqueness contract must harness adapters satisfy for + `(ScopeId, EventId)`?** It must uniquely identify one logical hook + delivery for the lifetime of its scope; `ActorKind` plays no part in + that identity; a harness whose native IDs are not already scope-unique + must have its adapter construct a canonicalized ID before ever + constructing a `RuntimeBoundary`; replaying the same logical delivery + must always reproduce the identical canonical pair. +23. **Which operation remains the protocol linearization point?** The DB CAS + inside `MutationTraceStore::commit` (`UPDATE ... WHERE revision = ?`), + unchanged from the original plan and from `store.rs`'s own design. Ref + pinning is durability infrastructure around that point, never a + substitute for it; the runtime lock, the checkout-identity lock, and the + checkout-identity crash-safe rename are all serialization/durability + optimizations around that one linearization point, never a second + decision-making mechanism. +24. **Are all previous PR #244 correctness fixes still preserved?** Yes: the + normal-object-database-plus-refs design, pin-before-CAS crash ordering, + the checkout-identity lock and its crash-safe rename, the + post-failure-only snapshot-failure read, bounded CAS retry on the taint + path, the runtime completion sequencing, and the `(ScopeId, EventId)` + replay contract are all unchanged by this revision's two Git-snapshot + corrections, which are additive precision fixes to the snapshot + mechanism and its tests, not architectural reversals. +25. **What unsafe state is impossible by construction after these changes?** + A durable DB row (worktree cursor or `mutation_trace_events` row) + referencing a `TreeId` Git could prune; two callers durably disagreeing + about one worktree's checkout identity, or one observing a half-written + identity; a coordinator reporting a snapshot-failure taint as persisted + when the DB never actually applied it, or basing that decision on state + read before the failure occurred; `git add -A -- .` ever running against + an invalid (zero-byte) index on an unborn `HEAD`; a GC durability test + whose result is meaningless because its "pinned" and "unpinned" trees + were the same content-addressed object; and — unchanged from the + original plan — a CAS-conflict retry re-capturing the worktree instead + of reusing the invocation's original observation. + +## Follow-up PR + +### Runtime completion sequence + +This PR is standalone and safe to ship on its own — nothing calls +`coordinate()` in production yet, so its two known, intentionally deferred +durability gaps (below) have no live consequence today. They stop being +merely theoretical the moment a harness adapter starts driving real hook +traffic through it, so this plan sequences the work accordingly rather than +jumping straight to harness wiring: + +```text +1. mutation-cursor-runtime-coordinator (this PR) + ↓ +2. mutation-cursor-external-taint (filesystem externalTaint marker; + ↓ protocol::database_failure wiring) +3. mutation-cursor-ref-reconciliation (refs/sce/mutation-cursor/** reclaim pass) + ↓ +4. harness boundary adapters + Agent Trace evidence + (RuntimeBoundary construction, MutationEvent → diff_trees → + existing diff_traces insertion) +``` + +Steps 2 and 3 may ship as one combined follow-up PR if the resulting change +stays coherent — they are independent problems (DB-unavailability +durability vs. Git-object storage growth) with no shared implementation, so +combining them is a sequencing convenience, not a technical dependency +between them. What is not acceptable is skipping either of them, or +reordering step 4 ahead of both: a harness adapter is a **production +consumer** of this coordinator, and only once steps 2 and 3 exist does the +coordinator's own durability story hold under the invocation volume a real +harness integration would actually produce. + +**Step 2 — external taint (why it cannot wait):** `spec/mutation_cursor.md` +already states that DB unavailability requires a filesystem durability +signal, because a DB write failure cannot be recorded by writing to the same +DB that just failed. Without it, exactly the scenario the model's +`externalTaint` mechanism exists to prevent remains reachable once real +hook traffic flows: a worktree changes → the DB becomes unavailable for that +one invocation → `coordinate()` fails with no durable trace of the lost +observation interval → a later, successful invocation has no signal that it +must conservatively rebaseline rather than continuing to attribute normally +across the gap. This PR's Requirement 10 resolution deliberately keeps this +gap narrow and named (`CoordinateError::Other`, no durable write, see +Design decisions) rather than leaving it unstated; this follow-up is where +it closes. + +**Step 3 — ref reconciliation (why it cannot wait):** see "Pin lifecycle: +create-only in this PR — and that growth is unbounded, not bounded, until +reconciliation exists" above for the corrected growth analysis and the +retained-roots contract this follow-up must implement (worktree +`cursor_tree` and every historical `mutation_trace_events` +`before_tree`/`after_tree`), plus the concurrency-safety requirement any +implementation must satisfy (never delete a ref for a tree that could still +become durable — see that section for the two candidate strategies this +plan leaves for the follow-up to choose between). + +### Step 4: harness boundary → runtime coordinator → committed `MutationEvent` → `diff_trees(before, after)` → existing Agent Trace `diff_traces` evidence + +Once steps 2 and 3 exist, this final integration PR translates one +harness's hook payloads (starting with Claude Code, or whichever is +prioritized) into `RuntimeBoundary` values — and, per the runtime identity +contract above, owns canonicalizing that harness's native session/event IDs +into scope-unique `ScopeId`/`EventId` pairs before ever constructing one — +calls this PR's `coordinate()`, and — when `CoordinateOutcome.mutation_event` +is `Some`, i.e. the triggering boundary produced real evidence — feeds its +`before_tree`/`after_tree` through this PR's `diff_trees` and the existing +`cli/src/services/patch.rs::parse_patch` into an `Agent Trace` `diff_traces` +row, alongside the mutation event's `attribution`/`boundary`/scope +information. It also owns deciding how (or whether) multiple harnesses share +one coordinator invocation path. From c2655aed975a2b00e75bec63d29d7b16f917d04e Mon Sep 17 00:00:00 2001 From: David Abram Date: Sat, 29 Aug 2026 21:31:34 +0200 Subject: [PATCH 2/8] runtime: Make checkout identity creation concurrency- and crash-safe Prevent concurrent callers from creating divergent checkout IDs or exposing a partially written canonical ID. Use a dedicated slow-path advisory lock with a second read, then persist through a synced unique temporary file and atomic rename; add coverage for convergence and interruption cases. Document the locking behavior and the repository's safe filesystem-test pattern. Plan: mutation-cursor-runtime-coordinator (T01) Co-authored-by: SCE --- cli/src/services/checkout/mod.rs | 272 +++++++++++++++++- context/cli/checkout-identity.md | 6 +- context/patterns.md | 11 +- .../mutation-cursor-runtime-coordinator.md | 217 ++++++++++---- 4 files changed, 432 insertions(+), 74 deletions(-) diff --git a/cli/src/services/checkout/mod.rs b/cli/src/services/checkout/mod.rs index 70b59ed3..1bde0188 100644 --- a/cli/src/services/checkout/mod.rs +++ b/cli/src/services/checkout/mod.rs @@ -7,6 +7,8 @@ //! //! Checkout identity is repository metadata used by Agent Trace diagnostics. +use std::fs::OpenOptions; +use std::io::Write; use std::path::{Path, PathBuf}; use std::process::Command; @@ -19,6 +21,9 @@ const SCE_CHECKOUT_DIR: &str = "sce"; /// File name for the checkout ID inside `/sce/`. const CHECKOUT_ID_FILE: &str = "checkout-id"; +/// File name for the identity-creation lock inside `/sce/`. +const CHECKOUT_ID_LOCK_FILE: &str = "checkout-id.lock"; + /// Resolves the Git directory (`.git` for normal clones, or the worktree-specific /// path for linked worktrees) by running `git rev-parse --git-dir` from the /// given repository root. @@ -105,14 +110,13 @@ pub fn read_checkout_id(git_dir: &Path) -> Result> { /// Gets the existing checkout ID or creates a new one. /// /// If `/sce/checkout-id` already exists, returns the stored ID (idempotent). -/// If it does not exist, generates a new `UUIDv7`, writes it to the file, and returns it. +/// If it does not exist, acquires `/sce/checkout-id.lock`, generates a new +/// `UUIDv7`, writes it through a temporary file and an atomic rename, and returns it. pub fn get_or_create_checkout_id(git_dir: &Path) -> Result { if let Some(existing_id) = read_checkout_id(git_dir)? { return Ok(existing_id); } - let checkout_id = Uuid::now_v7().to_string(); - let checkout_dir = git_dir.join(SCE_CHECKOUT_DIR); std::fs::create_dir_all(&checkout_dir).with_context(|| { format!( @@ -121,13 +125,269 @@ pub fn get_or_create_checkout_id(git_dir: &Path) -> Result { ) })?; + let lock_path = checkout_dir.join(CHECKOUT_ID_LOCK_FILE); + let lock_file = OpenOptions::new() + .write(true) + .create(true) + .truncate(false) + .open(&lock_path) + .with_context(|| { + format!( + "Failed to open checkout ID lock file '{}'", + lock_path.display() + ) + })?; + lock_file.lock().with_context(|| { + format!( + "Failed to acquire checkout ID lock '{}'", + lock_path.display() + ) + })?; + + if let Some(existing_id) = read_checkout_id(git_dir)? { + return Ok(existing_id); + } + + let checkout_id = Uuid::now_v7().to_string(); + persist_checkout_id(&checkout_dir, &checkout_id)?; + + Ok(checkout_id) +} + +fn persist_checkout_id(checkout_dir: &Path, checkout_id: &str) -> Result<()> { + persist_checkout_id_inner(checkout_dir, checkout_id, |_, _| Ok(())) +} + +fn persist_checkout_id_inner( + checkout_dir: &Path, + checkout_id: &str, + before_rename: F, +) -> Result<()> +where + F: FnOnce(&Path, &Path) -> Result<()>, +{ let checkout_id_path = checkout_dir.join(CHECKOUT_ID_FILE); - std::fs::write(&checkout_id_path, &checkout_id).with_context(|| { + let tmp_path = checkout_dir.join(format!("checkout-id.tmp-{checkout_id}")); + + let mut tmp_file = OpenOptions::new() + .write(true) + .create_new(true) + .open(&tmp_path) + .with_context(|| { + format!( + "Failed to create temporary checkout ID file '{}'", + tmp_path.display() + ) + })?; + tmp_file + .write_all(checkout_id.as_bytes()) + .with_context(|| { + format!( + "Failed to write checkout ID to temporary file '{}'", + tmp_path.display() + ) + })?; + tmp_file.sync_data().with_context(|| { + format!( + "Failed to sync temporary checkout ID file '{}'", + tmp_path.display() + ) + })?; + drop(tmp_file); + + before_rename(&tmp_path, &checkout_id_path)?; + + std::fs::rename(&tmp_path, &checkout_id_path).with_context(|| { format!( - "Failed to write checkout ID to '{}'", + "Failed to rename '{}' to '{}'", + tmp_path.display(), checkout_id_path.display() ) })?; - Ok(checkout_id) + #[cfg(unix)] + { + if let Ok(dir) = std::fs::File::open(checkout_dir) { + let _ = dir.sync_all(); + } + } + + Ok(()) +} + +#[cfg(test)] +mod tests { + use std::sync::atomic::{AtomicU64, Ordering}; + use std::sync::mpsc; + use std::thread; + use std::time::Duration; + + use super::*; + + static NEXT_TEST_GIT_DIR_ID: AtomicU64 = AtomicU64::new(0); + + fn unique_test_git_dir(label: &str) -> PathBuf { + let id = NEXT_TEST_GIT_DIR_ID.fetch_add(1, Ordering::Relaxed); + std::env::temp_dir().join(format!( + "sce-checkout-identity-{label}-{}-{id}", + std::process::id() + )) + } + + fn remove_test_git_dir(git_dir: &Path) { + let _ = std::fs::remove_dir_all(git_dir); + } + + #[test] + fn concurrent_first_time_callers_converge_on_one_checkout_id() { + let git_dir = unique_test_git_dir("concurrent-first-time"); + std::fs::create_dir_all(&git_dir).expect("git dir should be created"); + + let handles: Vec<_> = (0..8) + .map(|_| { + let git_dir = git_dir.clone(); + thread::spawn(move || { + get_or_create_checkout_id(&git_dir).expect("checkout id should be created") + }) + }) + .collect(); + + let ids: Vec = handles + .into_iter() + .map(|handle| handle.join().expect("thread should not panic")) + .collect(); + + let first = ids[0].clone(); + assert!( + ids.iter().all(|id| *id == first), + "all concurrent callers should converge on one checkout id, got {ids:?}" + ); + + let persisted = + read_checkout_id(&git_dir).expect("checkout id should be readable after creation"); + assert_eq!(persisted, Some(first)); + + remove_test_git_dir(&git_dir); + } + + #[test] + fn already_created_checkout_id_is_read_without_acquiring_lock() { + let git_dir = unique_test_git_dir("fast-path-no-lock"); + let checkout_dir = git_dir.join(SCE_CHECKOUT_DIR); + std::fs::create_dir_all(&checkout_dir).expect("checkout dir should be created"); + + let existing_id = Uuid::now_v7().to_string(); + std::fs::write(checkout_dir.join(CHECKOUT_ID_FILE), &existing_id) + .expect("seeded checkout id should be written"); + + let lock_path = checkout_dir.join(CHECKOUT_ID_LOCK_FILE); + let lock_file = OpenOptions::new() + .write(true) + .create(true) + .truncate(false) + .open(&lock_path) + .expect("lock file should be created"); + lock_file.lock().expect("lock should be acquired"); + + let (tx, rx) = mpsc::channel(); + let git_dir_clone = git_dir.clone(); + thread::spawn(move || { + let result = get_or_create_checkout_id(&git_dir_clone); + let _ = tx.send(result); + }); + + let result = rx + .recv_timeout(Duration::from_millis(500)) + .expect("fast path should not block on the identity-creation lock"); + assert_eq!(result.expect("existing id should be read"), existing_id); + + drop(lock_file); + remove_test_git_dir(&git_dir); + } + + #[test] + fn completed_rename_leaves_the_canonical_path_with_a_complete_id() { + let git_dir = unique_test_git_dir("completed-rename"); + let checkout_dir = git_dir.join(SCE_CHECKOUT_DIR); + std::fs::create_dir_all(&checkout_dir).expect("checkout dir should be created"); + + let checkout_id = Uuid::now_v7().to_string(); + persist_checkout_id(&checkout_dir, &checkout_id) + .expect("persistence should succeed without an injected interruption"); + + let persisted = read_checkout_id(&git_dir) + .expect("checkout id should be readable") + .expect("checkout id should exist"); + assert_eq!(persisted, checkout_id); + Uuid::parse_str(&persisted).expect("persisted id should be a complete, valid UUID"); + + remove_test_git_dir(&git_dir); + } + + #[test] + fn interruption_before_rename_leaves_the_canonical_path_absent() { + let git_dir = unique_test_git_dir("interrupted-before-rename"); + let checkout_dir = git_dir.join(SCE_CHECKOUT_DIR); + std::fs::create_dir_all(&checkout_dir).expect("checkout dir should be created"); + + let checkout_id = Uuid::now_v7().to_string(); + let result = + persist_checkout_id_inner(&checkout_dir, &checkout_id, |tmp_path, canonical_path| { + assert!( + tmp_path.exists(), + "temp file should exist by the time the pre-rename hook runs" + ); + assert_eq!( + std::fs::read_to_string(tmp_path).expect("temp file should be readable"), + checkout_id, + "temp file should already contain the complete id before the injected interruption" + ); + assert!( + !canonical_path.exists(), + "canonical path should still be absent at the pre-rename hook" + ); + Err(anyhow!("injected interruption before rename")) + }); + + assert!( + result.is_err(), + "persistence should surface the injected pre-rename interruption" + ); + + let canonical_path = checkout_dir.join(CHECKOUT_ID_FILE); + assert!( + !canonical_path.exists(), + "canonical checkout-id path must stay absent when interrupted before rename" + ); + assert_eq!( + read_checkout_id(&git_dir).expect("read should not error on an absent canonical id"), + None + ); + + remove_test_git_dir(&git_dir); + } + + #[test] + fn an_orphaned_temp_file_does_not_block_convergence_on_the_canonical_id() { + let git_dir = unique_test_git_dir("orphaned-temp-file"); + let checkout_dir = git_dir.join(SCE_CHECKOUT_DIR); + std::fs::create_dir_all(&checkout_dir).expect("checkout dir should be created"); + + let orphaned_tmp_path = checkout_dir.join(format!("checkout-id.tmp-{}", Uuid::now_v7())); + std::fs::write(&orphaned_tmp_path, b"orphaned-partial-content") + .expect("orphaned temp file should be written"); + + let checkout_id = get_or_create_checkout_id(&git_dir) + .expect("orphaned temp file should not block id creation"); + Uuid::parse_str(&checkout_id).expect("returned id should be a complete, valid UUID"); + + let persisted = read_checkout_id(&git_dir) + .expect("checkout id should be readable") + .expect("checkout id should exist"); + assert_eq!(persisted, checkout_id); + + assert!(orphaned_tmp_path.exists()); + + remove_test_git_dir(&git_dir); + } } diff --git a/context/cli/checkout-identity.md b/context/cli/checkout-identity.md index 4716f37a..8ccc010f 100644 --- a/context/cli/checkout-identity.md +++ b/context/cli/checkout-identity.md @@ -8,8 +8,8 @@ It assigns a stable identity to a local Git checkout or linked Git worktree. Che - `cli/src/services/checkout/mod.rs` - `resolve_git_dir(repo_root)` runs `git rev-parse --git-dir` from the supplied repository root. - - `read_checkout_id(git_dir)` reads `/sce/checkout-id` and validates non-empty UUID syntax. - - `get_or_create_checkout_id(git_dir)` reuses an existing ID or writes a new UUIDv7 checkout ID to `/sce/checkout-id`. + - `read_checkout_id(git_dir)` reads `/sce/checkout-id` and validates non-empty UUID syntax. This lock-free fast path is unchanged for the already-created case. + - `get_or_create_checkout_id(git_dir)` reuses an existing ID via `read_checkout_id`, or, on the slow path, acquires a dedicated `/sce/checkout-id.lock` (blocking `std::fs::File::lock()`, no timeout), re-checks for an existing ID under the lock, and otherwise generates a new UUIDv7 and persists it crash-safely: written to a unique `checkout-id.tmp-` file (`OpenOptions::create_new(true)`), synced with `File::sync_data()`, then moved into the canonical `checkout-id` path with an atomic `std::fs::rename` (plus a best-effort `#[cfg(unix)]` parent-directory sync). This guarantees every concurrent caller — the coordinator, `agent_trace_storage`, or any other caller of this function — converges on exactly one checkout ID, and that the canonical path is never observable as partially written, even across a process crash mid-write. The lock is scoped only to identity creation; it is distinct from the separate mutation-cursor runtime lock at `/sce/mutation-cursor.lock` (a different, larger critical section owned by an in-progress runtime coordinator, not yet part of this service). Orphaned `checkout-id.tmp-*` files left by an interrupted write are harmless and are not cleaned up by this function. ## Current integration state @@ -26,6 +26,6 @@ During setup and hook runtime: ## Testing boundary -No unit tests are currently included for this filesystem/Git-facing service. Filesystem, Git repository, and database behaviors should be covered in integration tests rather than unit tests per `context/patterns.md`. +`get_or_create_checkout_id` and its identity-creation lock have an inline `#[cfg(test)]` module in `cli/src/services/checkout/mod.rs` covering concurrent first-time convergence on one ID, the fast path never touching the lock, a completed rename leaving a complete ID, a simulated crash before rename leaving the canonical path absent, and an orphaned temp file not blocking a later call — each test uses a unique temp directory rather than a shared fixture, following the same filesystem-touching inline-unit-test precedent already used in `cli/src/services/mutation_trace/store.rs`. `resolve_git_dir` and Git-subprocess behavior remain uncovered by unit tests; that and other database behaviors should still be covered in integration tests rather than unit tests per `context/patterns.md`. See also: `context/cli/agent-trace-storage.md`, `context/cli/default-path-catalog.md`, `context/sce/agent-trace-db.md`, `context/sce/agent-trace-hooks-command-routing.md`. diff --git a/context/patterns.md b/context/patterns.md index 9ad112fd..e96aaa7a 100644 --- a/context/patterns.md +++ b/context/patterns.md @@ -186,10 +186,9 @@ ## Unit testing in Nix sandbox -- Unit tests must not depend on filesystem directories, temporary directories, or databases that could fail in Nix sandbox environments. -- Tests that require filesystem I/O, git repository operations, or database connections belong in integration tests, not unit tests. -- When a unit test needs filesystem, git, or database behavior that is not safe for `nix flake check`, delete it from the unit-test suite and reintroduce that coverage later as an integration test instead of keeping ignored tests in-tree. -- Pure unit tests should test in-memory logic, parsing, validation, and data transformations without external dependencies. -- Temporary-directory helpers and similar filesystem fixtures should only be used in integration tests, not unit tests. +- Pure unit tests should test in-memory logic, parsing, validation, and data transformations without external dependencies; prefer mocking/faking external dependencies over creating real filesystem or database state when a fake is available and sufficient. - In-memory database tests (e.g., `LocalDatabaseTarget::InMemory`) are acceptable for unit tests since they don't touch the filesystem. -- When adding new tests, prefer mocking/faking external dependencies over creating real filesystem or database state. +- A filesystem-touching `#[cfg(test)] mod tests` inline in the module under test is an established, Nix-sandbox-safe pattern in this codebase when every path is a unique, process-local path under `std::env::temp_dir()` (which resolves through `TMPDIR`, itself a writable per-build directory inside the Nix sandbox) — see `cli/src/services/mutation_trace/store.rs`'s `RepositoryAgentTraceDb`-backed tests and `cli/src/services/checkout/mod.rs`'s checkout-identity-lock tests. This is not integration-test-only territory in this repository. +- Use integration tests instead of this inline pattern where the behavior genuinely cannot be made deterministic and isolated inside the Nix sandbox this way. +- Do not depend on a shared or ambient path (`$HOME`, a fixed `/tmp` file name, the repository's own working tree) from a unit test; each test must construct its own unique, self-cleaning path. +- When a unit test needs behavior that cannot be made Nix-sandbox-safe this way (for example, network access, or state shared across the whole test binary), delete it from the unit-test suite and reintroduce that coverage later as an integration test instead of keeping ignored tests in-tree. diff --git a/context/plans/mutation-cursor-runtime-coordinator.md b/context/plans/mutation-cursor-runtime-coordinator.md index 41563d25..bb4e0f64 100644 --- a/context/plans/mutation-cursor-runtime-coordinator.md +++ b/context/plans/mutation-cursor-runtime-coordinator.md @@ -6,11 +6,11 @@ Adds the runtime (imperative-shell) layer that connects the verified, already-implemented mutation-cursor protocol kernel (`protocol.rs`) and its persistence layer (`store.rs`) to a real Git worktree: a concurrency-safe checkout-identity primitive, an OS-backed per-worktree advisory lock for the -coordinator's own critical section (`worktree_lock.rs`), an isolated Git +coordinator's own critical section (`runtime/worktree_lock.rs`), an isolated Git snapshot service that captures the current worktree state as a durable `TreeId` into the repository's normal object database and protects it from GC/prune with an SCE-owned ref, never touching the caller's real index -(`git_snapshot.rs`), and a runtime coordinator (`coordinator.rs`) that +(`runtime/git_snapshot.rs`), and a runtime coordinator (`runtime/coordinator.rs`) that derives `WorktreeId` from checkout identity, materializes worktree/scope rows, runs recovery when the durable state requires it, drives `prepare`/`commit`, and retries on CAS conflict — including CAS conflict @@ -69,27 +69,46 @@ distinct-content tree in the same repository, deliberately unreachable from any ref, and proves that one, not a same-content decoy, is what `git gc --prune=now`/`git prune --expire=now` actually reclaims. +A fourth revision (pre-implementation structural correction, before T02 or +any later task began) moves every planned imperative-runtime file under a +new `cli/src/services/mutation_trace/runtime/` module instead of adding them +flat alongside `protocol.rs`/`store.rs`/`types.rs`. `worktree_lock.rs`, +`git_snapshot.rs`, and `coordinator.rs` become +`runtime/worktree_lock.rs`, `runtime/git_snapshot.rs`, and +`runtime/coordinator.rs`; the planned `runtime_tests.rs` becomes +`runtime/tests.rs`, declared as `runtime`'s own `#[cfg(test)] mod tests`. +This is a pure module-boundary correction with no change to protocol, +persistence, or task-ordering semantics: it separates the pure, verified +protocol kernel and its durable persistence (`protocol.rs`, `store.rs`, +`types.rs`, `mbt/`) from the imperative shell that drives Git subprocesses, +filesystem locks, and checkout identity around them, so the dependency +direction — `runtime` depends on `protocol`/`store`/`types` and on +`services::checkout`, never the reverse — is structurally visible rather +than merely a documented convention. `checkout/` remains its own top-level +service, unmoved, since other Agent Trace storage paths already depend on it +independently of the mutation-cursor runtime. + ## Acceptance criteria - [ ] AC1: A worktree observed for the first time establishes the currently observed tree as its cursor baseline on its first boundary and emits no mutation evidence for filesystem changes that predate that observation. - - Validate: `coordinator::tests::first_observation_establishes_baseline_without_evidence` via `nix build .#checks..cli-tests` + - Validate: `runtime::coordinator::tests::first_observation_establishes_baseline_without_evidence` via `nix build .#checks..cli-tests` - [ ] AC2: An edit made between a `Start` and a subsequent `Advance` on the same scope commits as exactly one `AiExclusive` mutation event whose `before_tree`/`after_tree` match the Start baseline and the post-edit snapshot. - - Validate: `coordinator::tests::exclusive_edit_between_start_and_advance_commits_one_event` + - Validate: `runtime::coordinator::tests::exclusive_edit_between_start_and_advance_commits_one_event` - [ ] AC3: Re-processing the identical `(scope, event)` boundary a second time produces no duplicated mutation evidence. - - Validate: `coordinator::tests::replaying_the_same_scope_event_key_does_not_duplicate_evidence` + - Validate: `runtime::coordinator::tests::replaying_the_same_scope_event_key_does_not_duplicate_evidence` - [ ] AC4: A mutation made just before `Close` is still attributed using the scope set as it existed immediately before `Close`'s own scope transition. - - Validate: `coordinator::tests::close_boundary_attributes_using_pre_close_scope_set` + - Validate: `runtime::coordinator::tests::close_boundary_attributes_using_pre_close_scope_set` - [ ] AC5: Two concurrently active scopes on one worktree yield `AiContended` attribution for a subsequent `Advance`, independent of whether the two scopes share an `ActorKind`. - - Validate: `coordinator::tests::contended_scopes_yield_ai_contended_same_and_different_actor` + - Validate: `runtime::coordinator::tests::contended_scopes_yield_ai_contended_same_and_different_actor` - [ ] AC6: Capturing a Git snapshot never mutates the caller's real index, staged changes, or working tree, correctly reflects staged, unstaged, untracked, and deleted state, excludes ignored files, and — on an unborn @@ -97,7 +116,7 @@ any ref, and proves that one, not a same-content decoy, is what `git gc Git index (`git read-tree --empty`), never from a bare, freshly created file the coordinator merely assumes Git will treat as empty, and produces a correct tree even when the unborn repository has no files at all. - - Validate: `git_snapshot::tests::*` (index preservation, ignored files, deletion, unborn HEAD with a file, unborn HEAD with no files) + - Validate: `runtime::git_snapshot::tests::*` (index preservation, ignored files, deletion, unborn HEAD with a file, unborn HEAD with no files) - [ ] AC7: A snapshot's `TreeId` remains resolvable through `diff_trees` after the process that captured it has exited, its temporary index file no longer exists, **and after a `git gc --prune=now` / `git prune @@ -105,24 +124,24 @@ any ref, and proves that one, not a same-content decoy, is what `git gc reachable from an SCE-owned ref in the repository's normal refs namespace, not merely present in an isolated object store Git's own reachability analysis knows nothing about. - - Validate: `git_snapshot::tests::snapshot_survives_a_fresh_process_and_temp_index_deletion`, `git_snapshot::tests::pinned_snapshot_survives_git_gc_prune_now`, `git_snapshot::tests::pinned_snapshot_survives_git_prune_expire_now` + - Validate: `runtime::git_snapshot::tests::snapshot_survives_a_fresh_process_and_temp_index_deletion`, `runtime::git_snapshot::tests::pinned_snapshot_survives_git_gc_prune_now`, `runtime::git_snapshot::tests::pinned_snapshot_survives_git_prune_expire_now` - [ ] AC8: When two invocations race to commit from the same durable revision, exactly one succeeds, the other reloads durable state and recomputes its transition using its own originally captured snapshot, and no second Git snapshot is taken for that invocation. - - Validate: `coordinator::tests::cas_conflict_reloads_and_recomputes_without_a_second_snapshot` + - Validate: `runtime::coordinator::tests::cas_conflict_reloads_and_recomputes_without_a_second_snapshot` - [ ] AC9: Two coordinator invocations targeting the same worktree cannot execute their critical sections concurrently; invocations on two different worktrees, including two linked worktrees of the same repository, are not serialized against each other and resolve to the same repository-scoped DB with distinct `WorktreeId`s. - - Validate: `worktree_lock::tests::*` (contention, distinct-path independence); `runtime_tests::linked_worktrees_have_independent_locks_and_worktree_ids` + - Validate: `runtime::worktree_lock::tests::*` (contention, distinct-path independence); `runtime::tests::linked_worktrees_have_independent_locks_and_worktree_ids` - [ ] AC10: A worktree whose durable state is `SnapshotFailure`-tainted or `needs_rebaseline` is recovered exactly once, using the same snapshot captured for the triggering boundary, before that boundary is processed; `needs_rebaseline` recovery preserves live scopes, taint recovery abandons them. - - Validate: `coordinator::tests::recovers_from_needs_rebaseline_preserving_live_scopes`, `coordinator::tests::recovers_from_snapshot_failure_taint_abandoning_live_scopes` + - Validate: `runtime::coordinator::tests::recovers_from_needs_rebaseline_preserving_live_scopes`, `runtime::coordinator::tests::recovers_from_snapshot_failure_taint_abandoning_live_scopes` - [ ] AC11: A Git snapshot failure against an already-materialized worktree durably persists a `SnapshotFailure` taint via `protocol::taint`, retried under the same bounded semantic CAS-retry policy the main coordinator loop @@ -135,13 +154,13 @@ any ref, and proves that one, not a same-content decoy, is what `git gc never state read earlier in the same invocation — so a worktree another caller materializes concurrently, while this invocation's own Git snapshot is still being captured, is still found and correctly tainted. - - Validate: `coordinator::tests::snapshot_failure_taints_an_existing_worktree`, `coordinator::tests::snapshot_failure_taint_survives_a_losing_cas_and_commits_on_retry`, `coordinator::tests::snapshot_failure_taint_reports_not_persisted_after_retries_are_exhausted`, `coordinator::tests::snapshot_failure_before_any_baseline_makes_no_durable_write`, `coordinator::tests::snapshot_failure_taints_a_worktree_materialized_concurrently_during_capture` + - Validate: `runtime::coordinator::tests::snapshot_failure_taints_an_existing_worktree`, `runtime::coordinator::tests::snapshot_failure_taint_survives_a_losing_cas_and_commits_on_retry`, `runtime::coordinator::tests::snapshot_failure_taint_reports_not_persisted_after_retries_are_exhausted`, `runtime::coordinator::tests::snapshot_failure_before_any_baseline_makes_no_durable_write`, `runtime::coordinator::tests::snapshot_failure_taints_a_worktree_materialized_concurrently_during_capture` - [ ] AC12: All concurrent first-time callers of `checkout::get_or_create_checkout_id` for one physical checkout — whether through the coordinator, `agent_trace_storage`, or any other caller — converge on exactly one checkout ID, and the on-disk `checkout-id` file ends up containing that same value. - - Validate: `checkout::tests::concurrent_first_time_callers_converge_on_one_checkout_id`, `runtime_tests::agent_trace_storage_and_coordinator_observe_the_same_checkout_id` + - Validate: `checkout::tests::concurrent_first_time_callers_converge_on_one_checkout_id`, `runtime::tests::agent_trace_storage_and_coordinator_observe_the_same_checkout_id` - [ ] AC13: For cooperating SCE processes, the canonical `checkout-id` path is at every observable point either absent or contains exactly one complete, valid checkout ID — never a partially written or truncated @@ -160,9 +179,11 @@ any ref, and proves that one, not a same-content decoy, is what `git gc - `context/cli/mutation-trace-protocol.md` — "Target end-state architecture" currently states `coordinator.rs`/`git_snapshot.rs` "remain future work"; - update to reflect their existence and responsibilities. + update to reflect their existence and responsibilities at their actual + location, `cli/src/services/mutation_trace/runtime/`. - `context/cli/mutation-trace-store.md` — "Non-goals" currently states "no - `coordinator.rs` or `git_snapshot.rs` exists yet"; update once they exist. + `coordinator.rs` or `git_snapshot.rs` exists yet"; update once they exist + under `mutation_trace/runtime/`. - `context/cli/checkout-identity.md` — currently documents `get_or_create_checkout_id` as a plain "reuses an existing ID or writes a new one"; update to document the identity-creation lock and the @@ -171,8 +192,9 @@ any ref, and proves that one, not a same-content decoy, is what `git gc mention the new runtime coordinator layer while preserving the accurate "not yet wired into any hook or command" framing. - `context/context-map.md` — the `mutation-trace-protocol.md` line - annotation names the `coordinator.rs`/`git_snapshot.rs` seams as - not-yet-created; update once T06 lands. + annotation names the `coordinator.rs`/`git_snapshot.rs` seams (now + `mutation_trace/runtime/coordinator.rs`/`mutation_trace/runtime/git_snapshot.rs`) + as not-yet-created; update once T06 lands. - A new domain file (e.g. `context/cli/mutation-trace-runtime-coordinator.md`) documenting the lock, snapshot, ref-durability, and coordinator design decisions below is likely warranted; leave the exact filename to task @@ -193,11 +215,13 @@ Persist this field in every plan; this is durable plan state, not chat state: - **In scope:** `cli/src/services/checkout/mod.rs` (behavior change: `get_or_create_checkout_id` becomes concurrency-safe), - `cli/src/services/mutation_trace/worktree_lock.rs` (new), - `cli/src/services/mutation_trace/git_snapshot.rs` (new), - `cli/src/services/mutation_trace/coordinator.rs` (new), - `cli/src/services/mutation_trace/runtime_tests.rs` (new), - `cli/src/services/mutation_trace/mod.rs` (module registration only). + `cli/src/services/mutation_trace/runtime/mod.rs` (new), + `cli/src/services/mutation_trace/runtime/worktree_lock.rs` (new), + `cli/src/services/mutation_trace/runtime/git_snapshot.rs` (new), + `cli/src/services/mutation_trace/runtime/coordinator.rs` (new), + `cli/src/services/mutation_trace/runtime/tests.rs` (new), + `cli/src/services/mutation_trace/mod.rs` (module registration only, to + declare `pub(crate) mod runtime;`). - **Out of scope:** Claude/Codex/OpenCode/Pi hook translation, Bash `PreToolUse`/`PostToolUse` wiring, final Agent Trace `diff_traces` insertion, commit attribution, auto-sync/remote sync, control-plane @@ -233,7 +257,7 @@ Persist this field in every plan; this is durable plan state, not chat state: ## Task stack -- [ ] T01: `Make checkout-identity creation concurrency-safe and crash-safe` (status:todo) +- [x] T01: `Make checkout-identity creation concurrency-safe and crash-safe` (status:done) - Task ID: T01 - Scope: In — `cli/src/services/checkout/mod.rs`: `get_or_create_checkout_id` gains an internal, dedicated @@ -261,13 +285,76 @@ Persist this field in every plan; this is durable plan state, not chat state: interrupted attempt never blocks a later call from creating the canonical file. - Verify: `cargo test -p shared-context-engineering checkout::` (via `./scripts/run-cli-cargo.sh test`) - - Context synchronization: pending + - Completed: 2026-08-29 + - Files changed: `cli/src/services/checkout/mod.rs` + - Result: `get_or_create_checkout_id` now acquires a dedicated + `/sce/checkout-id.lock` (blocking `std::fs::File::lock()`, no + timeout) only on the slow path, re-checking `read_checkout_id` under the + lock before generating a new ID; the previously-created case still + returns via the unchanged, lock-free `read_checkout_id` fast path. The + write itself now goes through a unique + `checkout-id.tmp-` file created with + `OpenOptions::create_new(true)`, `write_all`, `File::sync_data()`, then + an atomic `std::fs::rename` into the canonical `checkout-id` path, with a + best-effort `#[cfg(unix)]` sync of the parent `sce/` directory handle + afterward. No public signature, checkout-ID file format, or caller + (including `agent_trace_storage`) changed. Added a `#[cfg(test)] mod + tests` covering concurrent first-time convergence, the fast path never + touching the lock, a completed rename leaving a complete ID, a simulated + crash before rename leaving the canonical path absent, and an orphaned + temp file not blocking a later call. + + PR #244 review follow-up: the crash-safety test originally reimplemented + the create-temp/write/sync/(no rename) sequence inline instead of + exercising production code, so a regression to an unsafe direct write + could still have passed it. The temp-file/write/sync/rename/directory-sync + sequence is now factored out of `get_or_create_checkout_id` into a + private `persist_checkout_id(checkout_dir, checkout_id)` helper, whose + name and signature describe only the production responsibility (persist + one checkout ID crash-safely) with no test vocabulary in it. It delegates + to a lower-level private `persist_checkout_id_inner(checkout_dir, + checkout_id, before_rename)`, where `before_rename` runs after the temp + file is written and synced but before the rename; + `persist_checkout_id` itself calls the inner helper with a no-op + `before_rename`, so `get_or_create_checkout_id` never mentions the test + seam at all, while both crash-safety tests call `persist_checkout_id_inner` + directly to reach it. `completed_rename_leaves_the_canonical_path_with_a_complete_id` + calls the public-shaped `persist_checkout_id` (no injected interruption) + and asserts the canonical file exists, contains exactly the generated ID, + and parses as a valid UUID. `interruption_before_rename_leaves_the_canonical_path_absent` + calls `persist_checkout_id_inner` directly and injects a `before_rename` + that asserts the temp file already contains the complete ID and the + canonical path is still absent, then returns an error to abort before the + rename runs; the test then asserts the canonical path stays absent and + `read_checkout_id` returns `None`. Crash-safety is now tested through the + actual, shared production persistence implementation with an injected + pre-rename failure, not through filesystem steps duplicated in the test. + - Verify (actual): `./scripts/run-cli-cargo.sh test --manifest-path + cli/Cargo.toml checkout::` → 5/5 passed. Additionally ran + `./scripts/run-cli-cargo.sh clippy --manifest-path cli/Cargo.toml + --all-targets -- -D warnings` → clean (pedantic/warnings denied + workspace-wide, per Constraints), and `./scripts/run-cli-cargo.sh fmt + --manifest-path cli/Cargo.toml -- --check` → clean. Additionally ran + `./scripts/run-cli-cargo.sh test --manifest-path cli/Cargo.toml + agent_trace_storage::` → 14/14 passed (no regression in the other, + unmodified caller). + - Context impact: + - Updated `context/cli/checkout-identity.md` to document the + identity-creation lock, the crash-safe temp-file/rename write, the + convergence guarantee across every caller, and the file's now-present + inline unit test coverage (previously documented as absent). + - Updated `context/patterns.md`'s "Unit testing in Nix sandbox" section + to document the filesystem-touching inline-unit-test pattern this + task's tests use (unique `std::env::temp_dir()` paths, Nix-sandbox-safe + via `TMPDIR`) as an established, code-verified pattern rather than the + stale "integration tests only" rule it stated before. + - Context synchronization: synced - [ ] T02: `Add the per-worktree OS advisory runtime lock` (status:todo) - Task ID: T02 - - Scope: In — `worktree_lock.rs`: `WorktreeLock::acquire(git_dir, timeout)`, + - Scope: In — `runtime/worktree_lock.rs`: `WorktreeLock::acquire(git_dir, timeout)`, RAII release, bounded polling acquisition. Out — wiring into - `coordinator.rs` (T05); resolving `git_dir` itself (caller-supplied); + `runtime/coordinator.rs` (T05); resolving `git_dir` itself (caller-supplied); the distinct checkout-identity-creation lock (T01, already landed by the time this task depends on nothing from it). - Dependencies: none @@ -276,12 +363,12 @@ Persist this field in every plan; this is durable plan state, not chat state: path until the first releases (via `Drop`) or the bounded timeout elapses, returns a distinct, matchable error on timeout, and never treats the lock file's mere existence as ownership. - - Verify: `cargo test -p shared-context-engineering worktree_lock::` (via `./scripts/run-cli-cargo.sh test`) + - Verify: `cargo test -p shared-context-engineering runtime::worktree_lock::` (via `./scripts/run-cli-cargo.sh test`) - Context synchronization: pending - [ ] T03: `Add the isolated Git snapshot service with ref-pinned durability` (status:todo) - Task ID: T03 - - Scope: In — `git_snapshot.rs`: `GitSnapshotService` (resolves + - Scope: In — `runtime/git_snapshot.rs`: `GitSnapshotService` (resolves `--git-dir` once; a unique, never-pre-created private temp index path under `/sce/tmp/`, reserved by the RAII guard but left for Git itself to create via `read-tree`; writes tree/blob objects into the @@ -291,7 +378,7 @@ Persist this field in every plan; this is durable plan state, not chat state: HEAD` and `git read-tree --empty`, never a bare/absent index file), `pin_tree` (creates `refs/sce/mutation-cursor//`), `diff_trees`. Out — wiring `capture_tree`/`pin_tree` into - `coordinator.rs` (T04); locking (T02/T05 own that); any ref + `runtime/coordinator.rs` (T04); locking (T02/T05 own that); any ref deletion/reconciliation logic (deferred, see Design decisions and Follow-up PR — this PR's pins are create-only). - Dependencies: none @@ -306,12 +393,12 @@ Persist this field in every plan; this is durable plan state, not chat state: genuinely unreachable control tree captured and left unpinned in the same repository is reclaimed by that same aggressive pass; `pin_tree` is idempotent for the same `(worktree_id, tree)` pair. - - Verify: `cargo test -p shared-context-engineering git_snapshot::` (via `./scripts/run-cli-cargo.sh test`) + - Verify: `cargo test -p shared-context-engineering runtime::git_snapshot::` (via `./scripts/run-cli-cargo.sh test`) - Context synchronization: pending - [ ] T04: `Add the coordinator's core protocol-integration pipeline` (status:todo) - Task ID: T04 - - Scope: In — `coordinator.rs`: `RuntimeBoundary` (with its + - Scope: In — `runtime/coordinator.rs`: `RuntimeBoundary` (with its `(ScopeId, EventId)` replay-identity contract documented on the type), `CoordinateOutcome`, `CoordinateError`, `SnapshotCapture` trait (dependency-injection seam covering `capture` and `pin`), worktree/scope @@ -338,7 +425,7 @@ Persist this field in every plan; this is durable plan state, not chat state: injected snapshot-capture failure and controlled store sequencing proves a worktree row materialized *during* a failing capture attempt (not before it) is still found and tainted by that same failing invocation. - - Verify: `cargo test -p shared-context-engineering coordinator::` (via `./scripts/run-cli-cargo.sh test`) + - Verify: `cargo test -p shared-context-engineering runtime::coordinator::` (via `./scripts/run-cli-cargo.sh test`) - Context synchronization: pending - [ ] T05: `Wire the worktree lock and checkout identity into coordinate()` (status:todo) @@ -354,12 +441,12 @@ Persist this field in every plan; this is durable plan state, not chat state: critical sections concurrently (one observably blocks until the other's `WorktreeLock` drops); `coordinate()`'s public signature matches `Result`. - - Verify: `cargo test -p shared-context-engineering coordinator::tests::two_threads_on_the_same_worktree_serialize` (via `./scripts/run-cli-cargo.sh test`) + - Verify: `cargo test -p shared-context-engineering runtime::coordinator::tests::two_threads_on_the_same_worktree_serialize` (via `./scripts/run-cli-cargo.sh test`) - Context synchronization: pending - [ ] T06: `Add cross-module runtime integration tests` (status:todo) - Task ID: T06 - - Scope: In — `runtime_tests.rs`: end-to-end `coordinate()` calls against + - Scope: In — `runtime/tests.rs`: end-to-end `coordinate()` calls against two real linked Git worktrees sharing one repository-scoped DB, proving distinct `WorktreeId`s/locks and non-serialization across worktrees; an end-to-end snapshot-failure-then-recovery cycle across two real @@ -372,7 +459,7 @@ Persist this field in every plan; this is durable plan state, not chat state: assertion, and an end-to-end failure/recovery cycle pass using only the public `coordinate()` API, real `git worktree add`, and a real temp-file `RepositoryAgentTraceDb`. - - Verify: `cargo test -p shared-context-engineering runtime_tests::` (via `./scripts/run-cli-cargo.sh test`) + - Verify: `cargo test -p shared-context-engineering runtime::tests::` (via `./scripts/run-cli-cargo.sh test`) - Context synchronization: pending ## Design decisions @@ -653,7 +740,7 @@ confirmed by this comparison: it is simultaneously the simplest implementation, the only one that is GC-safe by construction rather than by a separately-maintained copying invariant, and the only one that needs no extra object-directory environment variables for either writing or reading. -This also means `git_snapshot.rs` does **not** need to resolve +This also means `runtime/git_snapshot.rs` does **not** need to resolve `--git-common-dir` at all (a further simplification and deviation from the original plan, which planned to add `resolve_git_common_dir` to `checkout.rs`): verified experimentally from inside a real linked worktree, @@ -961,7 +1048,7 @@ satisfying Requirement 4's "runtime callers should not be allowed to invent a `WorktreeId`." A separate `RuntimeBoundary` type is warranted and lives in -`coordinator.rs`: `types::Boundary` intentionally carries no `ActorKind` +`runtime/coordinator.rs`: `types::Boundary` intentionally carries no `ActorKind` (that is a scope-*registration* concern, not a pure protocol transition concern) and its `Flush` variant carries an explicit `worktree: WorktreeId` the coordinator must not let a caller supply. `RuntimeBoundary` matches the @@ -1325,7 +1412,7 @@ Result<()>`) is the one seam this plan introduces for determinism: a fake, call-counting implementation to prove "exactly one `capture`, at most one `pin`, per invocation, even across CAS retries" (AC8) without needing real concurrent Git processes. The lock is *not* faked — -`worktree_lock::tests`/`checkout::tests` and T05's contention test exercise +`runtime::worktree_lock::tests`/`checkout::tests` and T05's contention test exercise real `std::fs::File` locks against real temp directories, because proving actual OS-level exclusion is the point. The DB is not faked either, matching `store.rs`'s own established precedent @@ -1361,27 +1448,39 @@ sufficient for everything this coordinator needs. New: -- `cli/src/services/mutation_trace/worktree_lock.rs` — per-worktree runtime - OS advisory lock (T02). -- `cli/src/services/mutation_trace/git_snapshot.rs` — isolated Git snapshot - capture, ref pinning, and `diff_trees` (T03). -- `cli/src/services/mutation_trace/coordinator.rs` — `RuntimeBoundary`, +- `cli/src/services/mutation_trace/runtime/mod.rs` — declares the runtime + submodules and exposes only what the rest of the crate needs (the + `coordinate()` entrypoint and its outcome/error types); Git snapshot and + lock internals stay unexposed outside `runtime`. +- `cli/src/services/mutation_trace/runtime/worktree_lock.rs` — per-worktree + runtime OS advisory lock (T02). +- `cli/src/services/mutation_trace/runtime/git_snapshot.rs` — isolated Git + snapshot capture, ref pinning, and `diff_trees` (T03). +- `cli/src/services/mutation_trace/runtime/coordinator.rs` — `RuntimeBoundary`, `CoordinateOutcome`, `CoordinateError`, `SnapshotCapture`, the internal protocol-integration pipeline, and the public lock-wrapped `coordinate()` - entrypoint (T04, T05). -- `cli/src/services/mutation_trace/runtime_tests.rs` — cross-module + entrypoint (T04, T05). This is the composition point: the only module that + combines `protocol`/`store`/`types` with `runtime::git_snapshot`, + `runtime::worktree_lock`, and `services::checkout`. +- `cli/src/services/mutation_trace/runtime/tests.rs` — cross-module linked-worktree, cross-caller checkout-identity, and end-to-end - failure/recovery integration tests (T06). + failure/recovery integration tests (T06), declared as `runtime`'s own + `#[cfg(test)] mod tests`. Modified: - `cli/src/services/checkout/mod.rs` — `get_or_create_checkout_id` gains an internal identity-creation lock plus a crash-safe temp-file-and-rename - write sequence (T01); no signature change. -- `cli/src/services/mutation_trace/mod.rs` — register the four new modules - (`pub mod git_snapshot; pub mod worktree_lock; pub mod coordinator; - #[cfg(test)] mod runtime_tests;`), consistent with the module's existing + write sequence (T01); no signature change. `checkout/` remains its own + top-level service, not moved under `mutation_trace`. +- `cli/src/services/mutation_trace/mod.rs` — add `pub(crate) mod runtime;` + alongside the existing `pub mod protocol; pub mod store; pub mod types;` + and private `mod mbt;`, consistent with the module's existing `#[allow(dead_code)]` precedent for code not yet wired to a command/hook. + `runtime/mod.rs` itself declares `mod git_snapshot; mod worktree_lock; mod + coordinator; #[cfg(test)] mod tests;`, keeping the Git-snapshot and lock + modules private to `runtime` — only `coordinator`'s public entrypoints are + reachable from outside it. Docs (context sync, not authored in this plan's tasks): @@ -1396,7 +1495,7 @@ Cargo/dependency changes: none. - **Pure unit tests** (no filesystem/DB/lock): none new — `protocol.rs`'s existing pure-function tests already cover `prepare`/`commit`/`recover`/ `taint`/`abandon`; this plan only adds imperative-shell code around them. -- **Git integration tests** (`git_snapshot.rs`, real temp Git repos): index +- **Git integration tests** (`runtime/git_snapshot.rs`, real temp Git repos): index preservation under staged+unstaged+untracked simultaneously, ignored-file exclusion, tracked-file deletion; unborn `HEAD` — both with a file present (asserting the resulting tree contains it) and with the working tree @@ -1420,7 +1519,7 @@ Cargo/dependency changes: none. durability holds against Git's own reachability analysis (proven against a genuinely unreachable control object, not against "the temp index is gone" or an object the test itself cannot actually make unreachable). -- **Store/DB integration tests** (`coordinator.rs`, real temp-file +- **Store/DB integration tests** (`runtime/coordinator.rs`, real temp-file `RepositoryAgentTraceDb`): AC1–AC5, AC10, AC11 — first observation, exclusive mutation, no-op observation, replay, close attribution, contention, both recovery modes, snapshot-failure with immediate success, @@ -1429,8 +1528,8 @@ Cargo/dependency changes: none. and the bootstrap snapshot-failure case. Prove: the coordinator drives the existing store/protocol APIs to produce exactly the durable outcomes the Quint model specifies, including under contention on the taint path. -- **Concurrency tests** (`checkout::mod.rs` T01, `worktree_lock.rs` T02, - `coordinator.rs` T05): concurrent first-time `get_or_create_checkout_id` +- **Concurrency tests** (`checkout::mod.rs` T01, `runtime/worktree_lock.rs` T02, + `runtime/coordinator.rs` T05): concurrent first-time `get_or_create_checkout_id` callers on one `git_dir` converge on one ID; a second `try_lock()` on the same worktree-lock path observably blocks/fails while the first holds it and succeeds immediately after `Drop`; two real threads calling @@ -1451,20 +1550,20 @@ Cargo/dependency changes: none. of, a subsequent `get_or_create_checkout_id` call, which still converges on one complete, valid ID. Prove: AC13 — the canonical path is never observable as partially written, and orphaned temp files are inert. -- **Linked-worktree tests** (`runtime_tests.rs`, real `git worktree add`): +- **Linked-worktree tests** (`runtime/tests.rs`, real `git worktree add`): two linked worktrees resolve distinct `checkout_id`/`WorktreeId` and distinct runtime-lock paths, concurrent `coordinate()` calls on the two worktrees do not block each other, both share one repository-scoped DB, and a tree pinned from one worktree resolves correctly when queried via the other worktree's `GIT_DIR`. Prove: AC9 end to end through the public API. -- **Cross-caller checkout-identity test** (`runtime_tests.rs`): a direct +- **Cross-caller checkout-identity test** (`runtime/tests.rs`): a direct `agent_trace_storage` resolution and a `coordinate()` call against the same `repository_root`, run concurrently on first-ever resolution, observe the identical checkout ID. Prove: AC12's convergence guarantee holds across module boundaries, not only within `checkout::mod.rs`'s own test suite. -- **Failure/recovery tests** (`coordinator.rs`, `runtime_tests.rs`): CAS +- **Failure/recovery tests** (`runtime/coordinator.rs`, `runtime/tests.rs`): CAS conflict via two prepared-but-not-yet-committed states from the same revision (proving reload+recompute+no-second-snapshot, and no re-pin); snapshot failure with and without a prior worktree row; snapshot-failure @@ -1477,7 +1576,7 @@ Cargo/dependency changes: none. rather than basing its decision on any earlier state; a full two-invocation failure-then-recovery cycle via the public API. Prove: AC8, AC10, AC11 in both unit and end-to-end form. -- **Lock staleness**: `worktree_lock::tests` documents (via a test that +- **Lock staleness**: `runtime::worktree_lock::tests` documents (via a test that opens, writes bytes to, and closes the lock file *without* ever calling `.lock()`, then proves a subsequent real `WorktreeLock::acquire` on the same path succeeds immediately) that a leftover lock file with no active From fc55cb5dd8a51348296ea68d1b54ec4fe70c19ad Mon Sep 17 00:00:00 2001 From: David Abram Date: Sun, 30 Aug 2026 06:29:25 +0200 Subject: [PATCH 3/8] runtime: Add bounded mutation-cursor worktree locking Serialize the mutation-cursor runtime's future worktree critical section without allowing a stuck holder to block later invocations indefinitely. Add private runtime scaffolding and a WorktreeLock that polls File::try_lock() against a caller-supplied timeout, reports a matchable timeout error, and releases ownership through RAII. The lock behavior is covered for contention, timeout, independent worktrees, and stale lock files. Plan: mutation-cursor-runtime-coordinator (T02) Co-authored-by: SCE --- cli/src/services/mutation_trace/mod.rs | 1 + .../services/mutation_trace/runtime/mod.rs | 1 + .../mutation_trace/runtime/worktree_lock.rs | 242 ++++++++++++++++++ context/cli/checkout-identity.md | 2 +- .../cli/mutation-trace-runtime-coordinator.md | 84 ++++++ context/context-map.md | 1 + .../mutation-cursor-runtime-coordinator.md | 137 +++++++++- 7 files changed, 465 insertions(+), 3 deletions(-) create mode 100644 cli/src/services/mutation_trace/runtime/mod.rs create mode 100644 cli/src/services/mutation_trace/runtime/worktree_lock.rs create mode 100644 context/cli/mutation-trace-runtime-coordinator.md diff --git a/cli/src/services/mutation_trace/mod.rs b/cli/src/services/mutation_trace/mod.rs index 9dd03bdf..3c2eb963 100644 --- a/cli/src/services/mutation_trace/mod.rs +++ b/cli/src/services/mutation_trace/mod.rs @@ -145,6 +145,7 @@ //! unchanged). pub mod protocol; +pub(crate) mod runtime; pub mod store; pub mod types; diff --git a/cli/src/services/mutation_trace/runtime/mod.rs b/cli/src/services/mutation_trace/runtime/mod.rs new file mode 100644 index 00000000..a503b454 --- /dev/null +++ b/cli/src/services/mutation_trace/runtime/mod.rs @@ -0,0 +1 @@ +mod worktree_lock; diff --git a/cli/src/services/mutation_trace/runtime/worktree_lock.rs b/cli/src/services/mutation_trace/runtime/worktree_lock.rs new file mode 100644 index 00000000..a7165c07 --- /dev/null +++ b/cli/src/services/mutation_trace/runtime/worktree_lock.rs @@ -0,0 +1,242 @@ +use std::fs::{File, OpenOptions, TryLockError}; +use std::path::{Path, PathBuf}; +use std::time::{Duration, Instant}; + +use anyhow::{Context, Result}; + +const SCE_RUNTIME_DIR: &str = "sce"; + +const WORKTREE_LOCK_FILE: &str = "mutation-cursor.lock"; + +const LOCK_POLL_INTERVAL: Duration = Duration::from_millis(100); + +#[derive(Debug)] +pub struct WorktreeLock { + file: File, + path: PathBuf, +} + +#[derive(Debug)] +pub enum WorktreeLockError { + TimedOut { path: PathBuf, timeout: Duration }, + Io(anyhow::Error), +} + +impl std::fmt::Display for WorktreeLockError { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + match self { + WorktreeLockError::TimedOut { path, timeout } => write!( + f, + "Timed out after {timeout:?} waiting for worktree lock '{}'", + path.display() + ), + WorktreeLockError::Io(source) => write!(f, "{source}"), + } + } +} + +impl std::error::Error for WorktreeLockError {} + +impl WorktreeLock { + pub fn acquire(git_dir: &Path, timeout: Duration) -> Result { + acquire_inner(git_dir, timeout, || {}) + } +} + +fn acquire_inner( + git_dir: &Path, + timeout: Duration, + on_contention: F, +) -> Result +where + F: FnOnce(), +{ + let runtime_dir = git_dir.join(SCE_RUNTIME_DIR); + std::fs::create_dir_all(&runtime_dir) + .with_context(|| { + format!( + "Failed to create runtime directory '{}'", + runtime_dir.display() + ) + }) + .map_err(WorktreeLockError::Io)?; + + let lock_path = runtime_dir.join(WORKTREE_LOCK_FILE); + let file = OpenOptions::new() + .write(true) + .create(true) + .truncate(false) + .open(&lock_path) + .with_context(|| { + format!( + "Failed to open worktree lock file '{}'", + lock_path.display() + ) + }) + .map_err(WorktreeLockError::Io)?; + + let mut on_contention = Some(on_contention); + let deadline = Instant::now() + timeout; + loop { + match file.try_lock() { + Ok(()) => { + return Ok(WorktreeLock { + file, + path: lock_path, + }); + } + Err(TryLockError::WouldBlock) => { + if let Some(on_contention) = on_contention.take() { + on_contention(); + } + let now = Instant::now(); + if now >= deadline { + return Err(WorktreeLockError::TimedOut { + path: lock_path, + timeout, + }); + } + std::thread::sleep(LOCK_POLL_INTERVAL.min(deadline - now)); + } + Err(TryLockError::Error(source)) => { + return Err(WorktreeLockError::Io(anyhow::Error::new(source).context( + format!("Failed to acquire worktree lock '{}'", lock_path.display()), + ))); + } + } + } +} + +impl Drop for WorktreeLock { + fn drop(&mut self) { + let _ = self.file.unlock(); + } +} + +#[cfg(test)] +mod tests { + use std::sync::atomic::{AtomicU64, Ordering}; + use std::sync::mpsc; + use std::thread; + + use super::*; + + static NEXT_TEST_GIT_DIR_ID: AtomicU64 = AtomicU64::new(0); + + fn unique_test_git_dir(label: &str) -> PathBuf { + let id = NEXT_TEST_GIT_DIR_ID.fetch_add(1, Ordering::Relaxed); + std::env::temp_dir().join(format!( + "sce-worktree-lock-{label}-{}-{id}", + std::process::id() + )) + } + + fn remove_test_git_dir(git_dir: &Path) { + let _ = std::fs::remove_dir_all(git_dir); + } + + #[test] + fn a_second_acquirer_blocks_until_the_first_releases() { + let git_dir = unique_test_git_dir("contention"); + std::fs::create_dir_all(&git_dir).expect("git dir should be created"); + + let first = WorktreeLock::acquire(&git_dir, Duration::from_secs(5)) + .expect("first acquirer should succeed immediately"); + + let (contention_tx, contention_rx) = mpsc::channel(); + let (result_tx, result_rx) = mpsc::channel(); + let git_dir_clone = git_dir.clone(); + let handle = thread::spawn(move || { + let result = acquire_inner(&git_dir_clone, Duration::from_secs(5), || { + contention_tx + .send(()) + .expect("contention signal channel should still be open"); + }); + result_tx + .send(()) + .expect("result signal channel should still be open"); + result + }); + + contention_rx + .recv_timeout(Duration::from_secs(5)) + .expect("worker should observe OS-level lock contention (TryLockError::WouldBlock)"); + + assert!( + result_rx.recv_timeout(Duration::from_millis(300)).is_err(), + "second acquirer should not succeed while the first still holds the lock" + ); + + drop(first); + + result_rx + .recv_timeout(Duration::from_secs(5)) + .expect("second acquirer should complete once the first releases the lock"); + + let result = handle + .join() + .expect("second acquirer thread should not panic"); + assert!( + result.is_ok(), + "second acquirer should succeed once the first releases the lock" + ); + + remove_test_git_dir(&git_dir); + } + + #[test] + fn distinct_worktree_paths_do_not_contend() { + let git_dir_a = unique_test_git_dir("distinct-a"); + let git_dir_b = unique_test_git_dir("distinct-b"); + std::fs::create_dir_all(&git_dir_a).expect("git dir a should be created"); + std::fs::create_dir_all(&git_dir_b).expect("git dir b should be created"); + + let lock_a = WorktreeLock::acquire(&git_dir_a, Duration::from_millis(200)) + .expect("lock on distinct path a should succeed"); + let lock_b = WorktreeLock::acquire(&git_dir_b, Duration::from_millis(200)) + .expect("lock on distinct path b should succeed independently of path a"); + + drop(lock_a); + drop(lock_b); + remove_test_git_dir(&git_dir_a); + remove_test_git_dir(&git_dir_b); + } + + #[test] + fn acquire_times_out_with_a_distinct_matchable_error_when_still_held() { + let git_dir = unique_test_git_dir("timeout"); + std::fs::create_dir_all(&git_dir).expect("git dir should be created"); + + let holder = WorktreeLock::acquire(&git_dir, Duration::from_secs(5)) + .expect("first acquirer should succeed immediately"); + + let result = WorktreeLock::acquire(&git_dir, Duration::from_millis(250)); + match result { + Err(WorktreeLockError::TimedOut { timeout, .. }) => { + assert_eq!(timeout, Duration::from_millis(250)); + } + other => panic!("expected a distinct TimedOut error, got {other:?}"), + } + + drop(holder); + remove_test_git_dir(&git_dir); + } + + #[test] + fn a_leftover_lock_file_with_no_active_os_lock_does_not_block_a_new_acquirer() { + let git_dir = unique_test_git_dir("stale-file"); + let runtime_dir = git_dir.join(SCE_RUNTIME_DIR); + std::fs::create_dir_all(&runtime_dir).expect("runtime dir should be created"); + + let lock_path = runtime_dir.join(WORKTREE_LOCK_FILE); + std::fs::write(&lock_path, b"leftover").expect("leftover lock file should be writable"); + + let result = WorktreeLock::acquire(&git_dir, Duration::from_millis(200)); + assert!( + result.is_ok(), + "a lock file with no active OS lock held against it must not block a new acquirer" + ); + + remove_test_git_dir(&git_dir); + } +} diff --git a/context/cli/checkout-identity.md b/context/cli/checkout-identity.md index 8ccc010f..fa5d86c2 100644 --- a/context/cli/checkout-identity.md +++ b/context/cli/checkout-identity.md @@ -28,4 +28,4 @@ During setup and hook runtime: `get_or_create_checkout_id` and its identity-creation lock have an inline `#[cfg(test)]` module in `cli/src/services/checkout/mod.rs` covering concurrent first-time convergence on one ID, the fast path never touching the lock, a completed rename leaving a complete ID, a simulated crash before rename leaving the canonical path absent, and an orphaned temp file not blocking a later call — each test uses a unique temp directory rather than a shared fixture, following the same filesystem-touching inline-unit-test precedent already used in `cli/src/services/mutation_trace/store.rs`. `resolve_git_dir` and Git-subprocess behavior remain uncovered by unit tests; that and other database behaviors should still be covered in integration tests rather than unit tests per `context/patterns.md`. -See also: `context/cli/agent-trace-storage.md`, `context/cli/default-path-catalog.md`, `context/sce/agent-trace-db.md`, `context/sce/agent-trace-hooks-command-routing.md`. +See also: `context/cli/agent-trace-storage.md`, `context/cli/default-path-catalog.md`, `context/cli/mutation-trace-runtime-coordinator.md` (the distinct mutation-cursor runtime lock), `context/sce/agent-trace-db.md`, `context/sce/agent-trace-hooks-command-routing.md`. diff --git a/context/cli/mutation-trace-runtime-coordinator.md b/context/cli/mutation-trace-runtime-coordinator.md new file mode 100644 index 00000000..4045f486 --- /dev/null +++ b/context/cli/mutation-trace-runtime-coordinator.md @@ -0,0 +1,84 @@ +# Mutation-trace runtime coordinator (`mutation_trace::runtime`) + +The imperative-shell layer that will connect the verified, pure +mutation-cursor protocol kernel ([`protocol.rs`](mutation-trace-protocol.md)) +and its persistence layer ([`store.rs`](mutation-trace-store.md)) to a real +Git worktree, built by the `mutation-cursor-runtime-coordinator` plan +(`context/plans/mutation-cursor-runtime-coordinator.md`). + +`cli/src/services/mutation_trace/runtime/` is a private submodule +(`pub(crate) mod runtime;` in `mutation_trace/mod.rs`), registered under the +same `#[allow(dead_code)]` precedent as the rest of `mutation_trace`: only +`coordinator.rs`'s eventual public entrypoints will be reachable from outside +`runtime` once they exist, and nothing under `runtime/` is wired into any +hook, command, or `diff_traces` insertion yet. + +`runtime` depends on `protocol`/`store`/`types` and on `services::checkout`, +never the reverse — this is a structural module boundary, not merely a +documented convention. + +## Current code surface + +Only the per-worktree runtime lock exists so far. + +- `cli/src/services/mutation_trace/runtime/worktree_lock.rs` — + `WorktreeLock::acquire(git_dir: &Path, timeout: Duration) -> + Result` opens/creates + `/sce/mutation-cursor.lock` and polls `std::fs::File::try_lock()` + on a 100ms interval against the caller-supplied bounded `timeout`, rather + than calling the blocking `File::lock()` directly. A held `WorktreeLock` + releases the OS lock when dropped (RAII). Timing out returns a distinct, + matchable `WorktreeLockError::TimedOut { path, timeout }` variant, separate + from `WorktreeLockError::Io` (file-open or other I/O failure). The lock + file's mere on-disk existence is never treated as ownership — only a + successful OS-level `try_lock()` counts, so a leftover lock file with no + active OS lock held against it never blocks a fresh acquirer. + +This lock guards the coordinator's own critical section (snapshot capture, +worktree/scope materialization, recovery, and the CAS retry loop) once the +coordinator exists. It is held on every future `coordinate()` call, unlike +the checkout-identity-creation lock. + +## Two distinct locks, two distinct invariants + +`/sce/mutation-cursor.lock` (this module) and +`/sce/checkout-id.lock` (see +[`checkout-identity.md`](checkout-identity.md)) are deliberately separate +locks guarding separate invariants, not one lock reused for two purposes: + +| | Path | Guards | Held by | Blocking behavior | +| --- | --- | --- | --- | --- | +| Checkout-identity lock | `/sce/checkout-id.lock` | "this checkout has at most one durable identity" | any caller of `get_or_create_checkout_id` | blocks indefinitely, no timeout — the critical section is a handful of filesystem syscalls | +| Mutation-cursor runtime lock | `/sce/mutation-cursor.lock` | the coordinator's entire runtime critical section | only the coordinator, on every invocation | bounded polling with a caller-supplied timeout — a stuck holder must not deadlock every future hook invocation | + +On-disk layout so far: + +```text +/sce/ +├── checkout-id (services::checkout) +├── checkout-id.lock (services::checkout) +└── mutation-cursor.lock (runtime::worktree_lock, this module) +``` + +## Testing boundary + +`WorktreeLock`'s inline `#[cfg(test)] mod tests` in `worktree_lock.rs` covers +contention (a second acquirer blocks until the first releases), independence +across distinct worktree paths, timing out with a distinct matchable error +while the lock is still held, and a leftover lock file with no active OS lock +held against it never blocking a fresh acquirer — each test uses a unique +`std::env::temp_dir()` path, following the same filesystem-touching +inline-unit-test precedent already used in `cli/src/services/checkout/mod.rs` +and `cli/src/services/mutation_trace/store.rs` (see `context/patterns.md`). + +## Status + +Only the per-worktree runtime lock (above) is implemented. The Git snapshot +service, the coordinator's protocol-integration pipeline, the public +lock-wrapped `coordinate()` entrypoint, and cross-module integration tests +remain future work tracked by the `mutation-cursor-runtime-coordinator` +plan's task stack. This file will grow to cover them as each lands. + +See also: [`mutation-trace-protocol.md`](mutation-trace-protocol.md), +[`mutation-trace-store.md`](mutation-trace-store.md), +[`checkout-identity.md`](checkout-identity.md). diff --git a/context/context-map.md b/context/context-map.md index 4c45a44e..9737aaa1 100644 --- a/context/context-map.md +++ b/context/context-map.md @@ -27,6 +27,7 @@ Feature/domain context: - `context/cli/mutation-trace-revision-refinement.md` (the Quint `revision: int` → Rust `WorktreeState::revision: u64` bounded-integer refinement: the private `next_revision` checked-arithmetic helper `commit`/`taint`/`abandon`/`recover` all route through instead of a raw `+ 1`, so a worktree at `revision: u64::MAX` is a guarded no-op/rejection rather than a silent wrap to `0`) - `context/cli/mutation-trace-quint-connect.md` (`#[cfg(test)]`-only Quint Connect model-based-testing harness in `cli/src/services/mutation_trace/mbt/` continuously checking `protocol.rs` against `spec/mutation_cursor.qnt`: the verification-only `mbtAction`/`MbtAction` record-payload transport excluded from comparison, the operation-identity-vs-`MbtStutter` distinction on guarded/no-op branches with its two deterministic regressions, finite ID mapping, the AC5 comparable-state field list, `randomPrepare` staying a single `step` branch, deterministic/generated (500×30, seed-reproducible) test coverage, and the two Nix checks — generic `checks.cli-tests` and dedicated `checks.mutation-trace-quint-connect` — that both require the pinned Quint binary plus the top-level `spec/` directory in `workspaceSrc`'s Nix fileset) - `context/cli/mutation-trace-store.md` (durable persistence for the mutation-cursor protocol in `cli/src/services/mutation_trace/store.rs`, built by the `mutation-cursor-store-persistence` plan: the one-directional `protocol.rs` -> `DurableTransition::between` (pure structural diff) -> `store.rs` (SQL translation) -> `RepositoryAgentTraceDb` boundary; migration `003_mutation_trace_protocol.sql`'s five tables; `AttemptState`/`external_taint` non-persistence; the 8-byte big-endian `BLOB` revision encoding plus explicit non-`Debug` enum codecs; the bounded hot-path `load_worktree` read vs. the cold-path `load_mutation_event`; `commit`'s single-`BEGIN IMMEDIATE` CAS batch via `TursoDb::execute_transactional_cas_batch`, distinguishing `Conflict`/retryable-transient/deterministic-`Err` outcomes; and the store's non-goals — no Git/filesystem I/O, no attribution/boundary-kind decisions, no retry-after-`Conflict` loop, no terminal-scope garbage collection) +- `context/cli/mutation-trace-runtime-coordinator.md` (imperative-shell runtime layer in `cli/src/services/mutation_trace/runtime/`, built by the `mutation-cursor-runtime-coordinator` plan: currently only the per-worktree OS advisory lock, `runtime::worktree_lock::WorktreeLock::acquire(git_dir, timeout)` at `/sce/mutation-cursor.lock`, bounded `try_lock()` polling with a distinct matchable timeout error and RAII release, and its distinction from the separate checkout-identity-creation lock; the Git snapshot service, coordinator pipeline, public `coordinate()` entrypoint, and cross-module integration tests remain future work) - `context/sce/cli-exit-code-contract.md` (stable class-based `sce` process exit-code contract in `cli/src/app.rs`, so automation can branch on failure category without parsing error text) - `context/sce/cli-error-code-taxonomy.md` (stable user-facing `SCE-ERR-*` diagnostic code classes rendered by `cli/src/app.rs`, complementing the numeric exit-code classes) - `context/sce/cli-stdout-stderr-contract.md` (implemented stream contract in `cli/src/app.rs`: command payloads on stdout only, redacted diagnostics on stderr) diff --git a/context/plans/mutation-cursor-runtime-coordinator.md b/context/plans/mutation-cursor-runtime-coordinator.md index bb4e0f64..5a872cce 100644 --- a/context/plans/mutation-cursor-runtime-coordinator.md +++ b/context/plans/mutation-cursor-runtime-coordinator.md @@ -350,7 +350,7 @@ Persist this field in every plan; this is durable plan state, not chat state: stale "integration tests only" rule it stated before. - Context synchronization: synced -- [ ] T02: `Add the per-worktree OS advisory runtime lock` (status:todo) +- [x] T02: `Add the per-worktree OS advisory runtime lock` (status:done) - Task ID: T02 - Scope: In — `runtime/worktree_lock.rs`: `WorktreeLock::acquire(git_dir, timeout)`, RAII release, bounded polling acquisition. Out — wiring into @@ -364,7 +364,140 @@ Persist this field in every plan; this is durable plan state, not chat state: returns a distinct, matchable error on timeout, and never treats the lock file's mere existence as ownership. - Verify: `cargo test -p shared-context-engineering runtime::worktree_lock::` (via `./scripts/run-cli-cargo.sh test`) - - Context synchronization: pending + - Completed: 2026-08-30 + - Files changed: `cli/src/services/mutation_trace/runtime/mod.rs` (new), + `cli/src/services/mutation_trace/runtime/worktree_lock.rs` (new), + `cli/src/services/mutation_trace/mod.rs` (added `pub(crate) mod runtime;`) + - Result: Added `runtime/mod.rs`, declaring `mod worktree_lock;` (private — + only `coordinator`'s future public entrypoints will be reachable from + outside `runtime`, per the plan's module-privacy design), and registered + `pub(crate) mod runtime;` in `mutation_trace/mod.rs`. This scaffolding was + not explicitly named in T02's own Scope line but was necessary for + `worktree_lock.rs` to compile and for its tests to run — recorded as a + reviewed assumption, not a deviation from scope. + + `WorktreeLock::acquire(git_dir: &Path, timeout: Duration) -> + Result` opens/creates + `/sce/mutation-cursor.lock` via `OpenOptions` (matching T01's + file-creation convention), then polls `File::try_lock()` every 100ms + (`LOCK_POLL_INTERVAL`) against the caller-supplied bounded `timeout` + rather than calling the blocking `lock()` directly — never treating the + lock file's mere existence as ownership, only a successful OS-level + `try_lock()`. On timeout it returns a distinct `WorktreeLockError::TimedOut + { path, timeout }` variant (matchable independently of the `Io` variant + wrapping other failures), matching the plan's design decision that this + lock is bounded, unlike the checkout-identity lock. `WorktreeLock` + implements `Drop`, releasing the OS lock (`self.file.unlock()`, + best-effort) when the guard is dropped — RAII release, matching the + Done-when requirement. `WorktreeLockError` implements `Display` and + `std::error::Error`, and derives `Debug`, consistent with an ordinary + matchable Rust error type (this task's own error type — not the + coordinator's later `CoordinateError::LockAcquisition(anyhow::Error)`, + which T05 will construct by wrapping whatever this function returns). + Added a `#[cfg(test)] mod tests` covering: a second acquirer blocking + until the first releases (via a channel-signaled background thread); + two distinct worktree paths not contending; `acquire` timing out with a + distinct, matchable `TimedOut` error (asserting the exact `timeout` field) + when the lock is still held; and a leftover lock file written directly to + disk, with no `.lock()`/`.try_lock()` ever called against it, not + blocking a fresh acquirer — proving the "mere existence is not ownership" + requirement. + + PR #244 review follow-up: two issues were corrected without changing + locking design or T02 semantics. First, the module doc comment's claim + that "`File::lock()`/`try_lock()` have no built-in timeout and block + indefinitely" was inaccurate for `try_lock()`, which is non-blocking and + returns `TryLockError::WouldBlock` immediately on contention; the comment + now states that `File::lock()` can block indefinitely while + `File::try_lock()` is non-blocking but provides no waiting deadline by + itself, so `WorktreeLock::acquire` polls it at a short interval until + acquisition succeeds or the caller-supplied deadline expires. The plan's + own "Blocking vs. timeout" design-decision text and + `context/cli/mutation-trace-runtime-coordinator.md` were both checked + against the same claim and found already accurate (they describe + `File::lock()`'s blocking behavior specifically, never claiming + `try_lock()` blocks), so neither needed a correction. + + Second, `a_second_acquirer_blocks_until_the_first_releases` previously + inferred "the worker was blocked" from a 300ms `recv_timeout` on a single + completion channel — a scheduling false positive, since a merely delayed + (not blocked) worker thread would produce the same observation without + proving it ever reached `WorktreeLock::acquire`. The test now uses two + channels: the worker signals a dedicated `started` channel immediately + before calling `acquire`, and the main thread waits (with a generous + 5-second bound) for that signal before asserting non-completion — so the + 300ms `recv_timeout` now only bounds the specific assertion "the worker, + having definitely reached the acquisition attempt, must not complete + while the first guard is held," rather than standing in for proof the + worker started. After dropping the first guard, the test now waits for + the worker's result signal (bounded, not a bare `join()`) before joining + the thread and asserting success, so the thread is never left detached. + This strengthened test now fails if `WorktreeLock::acquire` were + accidentally changed to let a second caller acquire while the first + guard remained alive. + + Third, at the user's explicit request, every comment in + `worktree_lock.rs` and `runtime/mod.rs` was removed — including the + module doc comment whose corrected wording is quoted above and every + `///`/`//` comment on items, fields, and test steps — since both files + were authored entirely during this task and its review follow-up. This + superseded the module-doc-comment correction above at the text level + without reopening its substance: the inaccurate `try_lock()` claim no + longer exists in either corrected or original form, because no doc + comment remains. Production behavior, the public API, and all four + tests are unchanged; `WorktreeLock`, `WorktreeLockError`, and their + documented semantics above remain the authoritative description of this + module's behavior now that the code carries no comments of its own. + + Fourth, the "Second" fix above still had a scheduler false positive: + signaling "started" immediately before calling `WorktreeLock::acquire` + proved only that the worker reached the call site, not that it actually + executed `try_lock()` and observed contention — a descheduled-before-call + worker would have produced the same passing result. `acquire`'s body + moved into a new private `acquire_inner(git_dir, timeout, on_contention: + impl FnOnce())` that `acquire` calls with a no-op closure + (`acquire_inner(git_dir, timeout, || {})`), keeping the public API and + every other behavior (deadline calculation, poll interval, error + variants, RAII release, stale-lock-file handling, per-worktree + independence) unchanged. The `on_contention` callback fires at most once, + only from inside the already-existing `Err(TryLockError::WouldBlock)` + arm, strictly after that branch is reached and before the existing + deadline check/sleep — it observes contention, it does not create or + influence it. The contention test now calls `acquire_inner` directly with + a closure that signals a dedicated channel, and waits on that channel + (bounded only as a deadlock guard) before asserting non-completion and + dropping the first guard. The contention test observes the actual + `TryLockError::WouldBlock` branch before the first guard is released, + then proves the second acquirer succeeds after `Drop` — a real + happens-before proof of OS-level contention, not an inference from + timing. This test now fails if the `WouldBlock` branch were no longer + reached while the first lock is held. Re-ran the contention test 8 + consecutive times with no flakiness. + - Verify (actual): `./scripts/run-cli-cargo.sh test --manifest-path + cli/Cargo.toml runtime::worktree_lock::` → 4/4 passed (re-run after the + "Fourth" `acquire_inner` fix above; the contention test was additionally + run 8 consecutive times in isolation with no flakiness). Additionally ran + `./scripts/run-cli-cargo.sh clippy --manifest-path cli/Cargo.toml + --all-targets -- -D warnings` → clean (pedantic/warnings denied + workspace-wide, per Constraints), and `./scripts/run-cli-cargo.sh fmt + --manifest-path cli/Cargo.toml -- --check` → clean (after running + `cargo fmt` once to apply formatting corrections, on the original T02 + implementation; every subsequent PR #244 review follow-up change, + including the `acquire_inner` seam, was already fmt-clean). + - Context impact: `domain` — introduced a new architectural element + (`WorktreeLock`, the mutation-cursor runtime lock) not owned by any + existing context file. Added `context/cli/mutation-trace-runtime-coordinator.md` + describing it, linked it from `context/context-map.md`, and cross-linked + it from `context/cli/checkout-identity.md`'s "See also" line. The five + root context files (`overview.md`, `architecture.md`, `glossary.md`, + `patterns.md`, `context-map.md`) were verified against this change; + `context-map.md` was edited to add the new domain file's index entry, + the other four remained accurate and were not edited. No qualifying + system-wide architecture decision was introduced by this task — the + lock's design (two distinct locks, bounded polling, std-only locking) + was already established in the plan's own Design decisions section + during plan authoring, not originated by this task's execution. + - Context synchronization: synced - [ ] T03: `Add the isolated Git snapshot service with ref-pinned durability` (status:todo) - Task ID: T03 From be991eb2d644fba2189911e93998ddfab2039a3c Mon Sep 17 00:00:00 2001 From: David Abram Date: Sun, 30 Aug 2026 07:20:45 +0200 Subject: [PATCH 4/8] runtime: Add isolated Git snapshot service with ref-pinned trees Capture complete worktree state through a temporary index without altering the real index or working tree. Store snapshot objects in Git's normal object database, pin durable trees with SCE-owned refs, and expose tree diffs for the future coordinator; add focused coverage for state capture, pruning, unborn HEADs, and SHA-256 repositories. Update mutation-trace architecture documentation and record the completed implementation and verification details. Plan: mutation-cursor-runtime-coordinator (T03) Co-authored-by: SCE --- .../mutation_trace/runtime/git_snapshot.rs | 618 ++++++++++++++++++ .../services/mutation_trace/runtime/mod.rs | 1 + context/cli/mutation-trace-protocol.md | 28 +- .../cli/mutation-trace-runtime-coordinator.md | 73 ++- context/cli/mutation-trace-store.md | 7 +- context/context-map.md | 4 +- .../mutation-cursor-runtime-coordinator.md | 156 ++++- 7 files changed, 858 insertions(+), 29 deletions(-) create mode 100644 cli/src/services/mutation_trace/runtime/git_snapshot.rs diff --git a/cli/src/services/mutation_trace/runtime/git_snapshot.rs b/cli/src/services/mutation_trace/runtime/git_snapshot.rs new file mode 100644 index 00000000..763f5453 --- /dev/null +++ b/cli/src/services/mutation_trace/runtime/git_snapshot.rs @@ -0,0 +1,618 @@ +use std::path::{Path, PathBuf}; +use std::process::Command; + +use anyhow::{anyhow, Context, Result}; +use uuid::Uuid; + +use crate::services::mutation_trace::types::{TreeId, WorktreeId}; + +const SCE_RUNTIME_DIR: &str = "sce"; +const TMP_INDEX_DIR: &str = "tmp"; +const REF_NAMESPACE: &str = "refs/sce/mutation-cursor"; + +pub struct GitSnapshotService { + git_dir: PathBuf, + repository_root: PathBuf, +} + +impl GitSnapshotService { + pub fn new(repository_root: &Path) -> Result { + let git_dir = resolve_git_dir(repository_root)?; + Ok(GitSnapshotService { + git_dir, + repository_root: repository_root.to_path_buf(), + }) + } + + pub fn capture_tree(&self) -> Result { + let tmp_dir = self.git_dir.join(SCE_RUNTIME_DIR).join(TMP_INDEX_DIR); + std::fs::create_dir_all(&tmp_dir).with_context(|| { + format!( + "Failed to create temporary index directory '{}'", + tmp_dir.display() + ) + })?; + let index_guard = TempIndexGuard::reserve(&tmp_dir); + + if self.head_exists()? { + self.run_git(&["read-tree", "HEAD"], Some(&index_guard.path))?; + } else { + self.run_git(&["read-tree", "--empty"], Some(&index_guard.path))?; + } + + self.run_git(&["add", "-A", "--", "."], Some(&index_guard.path))?; + + let tree_sha = self.run_git(&["write-tree"], Some(&index_guard.path))?; + Ok(TreeId(tree_sha.trim().to_string())) + } + + pub fn pin_tree(&self, worktree_id: &WorktreeId, tree: &TreeId) -> Result<()> { + let ref_name = pin_ref_name(worktree_id, tree); + self.run_git(&["update-ref", &ref_name, &tree.0], None)?; + Ok(()) + } + + pub fn diff_trees(&self, before: &TreeId, after: &TreeId) -> Result { + self.run_git( + &[ + "diff", + "--binary", + "--full-index", + "--no-ext-diff", + "--no-textconv", + &before.0, + &after.0, + ], + None, + ) + } + + fn head_exists(&self) -> Result { + let output = Command::new("git") + .args(["rev-parse", "--verify", "--quiet", "HEAD"]) + .current_dir(&self.repository_root) + .env("GIT_DIR", &self.git_dir) + .output() + .with_context(|| { + format!( + "Failed to run git rev-parse --verify --quiet HEAD in '{}'", + self.repository_root.display() + ) + })?; + + match output.status.code() { + Some(0) => Ok(true), + Some(1) => Ok(false), + _ => { + let stderr = String::from_utf8_lossy(&output.stderr).trim().to_string(); + let stdout = String::from_utf8_lossy(&output.stdout).trim().to_string(); + let detail = if stderr.is_empty() { stdout } else { stderr }; + Err(anyhow!( + "git rev-parse --verify --quiet HEAD failed unexpectedly (status {:?}): {detail}", + output.status.code() + )) + } + } + } + + fn run_git(&self, args: &[&str], index_file: Option<&Path>) -> Result { + let mut command = Command::new("git"); + command + .args(args) + .current_dir(&self.repository_root) + .env("GIT_DIR", &self.git_dir); + if let Some(index_file) = index_file { + command.env("GIT_INDEX_FILE", index_file); + } + + let output = command.output().with_context(|| { + format!( + "Failed to run git command {:?} in '{}'", + args, + self.repository_root.display() + ) + })?; + + if !output.status.success() { + let stderr = String::from_utf8_lossy(&output.stderr).trim().to_string(); + let stdout = String::from_utf8_lossy(&output.stdout).trim().to_string(); + let detail = if stderr.is_empty() { stdout } else { stderr }; + return Err(anyhow!("git {args:?} failed: {detail}")); + } + + String::from_utf8(output.stdout) + .with_context(|| format!("git {args:?} emitted invalid UTF-8")) + } +} + +fn pin_ref_name(worktree_id: &WorktreeId, tree: &TreeId) -> String { + format!("{REF_NAMESPACE}/{}/{}", worktree_id.0, tree.0) +} + +fn resolve_git_dir(repository_root: &Path) -> Result { + let output = Command::new("git") + .args(["rev-parse", "--absolute-git-dir"]) + .current_dir(repository_root) + .output() + .with_context(|| { + format!( + "Failed to run git rev-parse --absolute-git-dir in '{}'", + repository_root.display() + ) + })?; + + if !output.status.success() { + let stderr = String::from_utf8_lossy(&output.stderr).trim().to_string(); + return Err(anyhow!( + "git rev-parse --absolute-git-dir failed in '{}': {}", + repository_root.display(), + stderr + )); + } + + let git_dir = PathBuf::from( + String::from_utf8(output.stdout) + .with_context(|| "git rev-parse --absolute-git-dir emitted invalid UTF-8")? + .trim(), + ); + + debug_assert!( + git_dir.is_absolute(), + "git rev-parse --absolute-git-dir should always return an absolute path, got '{}'", + git_dir.display() + ); + + Ok(git_dir) +} + +struct TempIndexGuard { + path: PathBuf, +} + +impl TempIndexGuard { + fn reserve(tmp_dir: &Path) -> TempIndexGuard { + let path = tmp_dir.join(format!("index-{}", Uuid::new_v4())); + TempIndexGuard { path } + } +} + +impl Drop for TempIndexGuard { + fn drop(&mut self) { + let _ = std::fs::remove_file(&self.path); + } +} + +#[cfg(test)] +mod tests { + use std::sync::atomic::{AtomicU64, Ordering}; + + use super::*; + + static NEXT_TEST_REPO_ID: AtomicU64 = AtomicU64::new(0); + + fn unique_test_repo(label: &str) -> PathBuf { + let id = NEXT_TEST_REPO_ID.fetch_add(1, Ordering::Relaxed); + std::env::temp_dir().join(format!( + "sce-git-snapshot-{label}-{}-{id}", + std::process::id() + )) + } + + fn init_repo(repo_root: &Path) { + std::fs::create_dir_all(repo_root).expect("repo root should be created"); + run(repo_root, &["init", "--quiet"]); + run(repo_root, &["config", "user.email", "test@example.com"]); + run(repo_root, &["config", "user.name", "Test"]); + } + + fn run(repo_root: &Path, args: &[&str]) -> String { + let output = Command::new("git") + .args(args) + .current_dir(repo_root) + .output() + .expect("git command should spawn"); + assert!( + output.status.success(), + "git {args:?} failed: {}", + String::from_utf8_lossy(&output.stderr) + ); + String::from_utf8(output.stdout).expect("git output should be valid UTF-8") + } + + fn commit_all(repo_root: &Path, message: &str) { + run(repo_root, &["add", "-A"]); + run(repo_root, &["commit", "--quiet", "-m", message]); + } + + fn remove_test_repo(repo_root: &Path) { + let _ = std::fs::remove_dir_all(repo_root); + } + + fn worktree_id() -> WorktreeId { + WorktreeId("test-worktree".to_string()) + } + + #[test] + fn capture_preserves_real_index_and_working_tree_state() { + let repo_root = unique_test_repo("preserves-index"); + init_repo(&repo_root); + std::fs::write(repo_root.join("committed.txt"), b"original\n") + .expect("committed file should be writable"); + commit_all(&repo_root, "initial commit"); + + std::fs::write(repo_root.join("staged.txt"), b"staged\n") + .expect("staged file should be writable"); + run(&repo_root, &["add", "staged.txt"]); + std::fs::write(repo_root.join("committed.txt"), b"modified\n") + .expect("unstaged modification should be writable"); + std::fs::write(repo_root.join("untracked.txt"), b"untracked\n") + .expect("untracked file should be writable"); + + let status_before = run(&repo_root, &["status", "--porcelain"]); + let diff_before = run(&repo_root, &["diff"]); + let diff_cached_before = run(&repo_root, &["diff", "--cached"]); + + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + let tree = service + .capture_tree() + .expect("capture should succeed with mixed staged/unstaged/untracked state"); + assert!(!tree.0.is_empty()); + + let status_after = run(&repo_root, &["status", "--porcelain"]); + let diff_after = run(&repo_root, &["diff"]); + let diff_cached_after = run(&repo_root, &["diff", "--cached"]); + + assert_eq!(status_before, status_after); + assert_eq!(diff_before, diff_after); + assert_eq!(diff_cached_before, diff_cached_after); + + let ls_tree = run(&repo_root, &["ls-tree", "-r", "--name-only", &tree.0]); + assert!(ls_tree.contains("staged.txt")); + assert!(ls_tree.contains("untracked.txt")); + let committed_blob = run(&repo_root, &["show", &format!("{}:committed.txt", tree.0)]); + assert_eq!(committed_blob, "modified\n"); + + remove_test_repo(&repo_root); + } + + #[test] + fn capture_excludes_ignored_files() { + let repo_root = unique_test_repo("ignored-files"); + init_repo(&repo_root); + std::fs::write(repo_root.join(".gitignore"), b"ignored.txt\n") + .expect(".gitignore should be writable"); + commit_all(&repo_root, "add gitignore"); + std::fs::write(repo_root.join("ignored.txt"), b"should not appear\n") + .expect("ignored file should be writable"); + + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + let tree = service + .capture_tree() + .expect("capture should succeed with an ignored file present"); + + let ls_tree = run(&repo_root, &["ls-tree", "-r", "--name-only", &tree.0]); + assert!(!ls_tree.contains("ignored.txt")); + + remove_test_repo(&repo_root); + } + + #[test] + fn capture_reflects_deletion_of_a_committed_file() { + let repo_root = unique_test_repo("deletion"); + init_repo(&repo_root); + std::fs::write(repo_root.join("to-delete.txt"), b"will be removed\n") + .expect("file should be writable"); + commit_all(&repo_root, "add file to delete"); + std::fs::remove_file(repo_root.join("to-delete.txt")).expect("file should be removable"); + + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + let tree = service + .capture_tree() + .expect("capture should succeed with a deletion pending"); + + let ls_tree = run(&repo_root, &["ls-tree", "-r", "--name-only", &tree.0]); + assert!(!ls_tree.contains("to-delete.txt")); + + remove_test_repo(&repo_root); + } + + #[test] + fn capture_on_unborn_head_with_a_file_produces_a_valid_tree() { + let repo_root = unique_test_repo("unborn-head-with-file"); + init_repo(&repo_root); + std::fs::write(repo_root.join("untracked.txt"), b"before any commit\n") + .expect("untracked file should be writable"); + + let head_probe = Command::new("git") + .args(["rev-parse", "--verify", "--quiet", "HEAD"]) + .current_dir(&repo_root) + .output() + .expect("git rev-parse should spawn"); + assert!( + !head_probe.status.success(), + "HEAD should be unborn in a freshly initialized repository" + ); + + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + let tree = service + .capture_tree() + .expect("capture should succeed on an unborn HEAD with a file present"); + + let ls_tree = run(&repo_root, &["ls-tree", "-r", "--name-only", &tree.0]); + assert!(ls_tree.contains("untracked.txt")); + + remove_test_repo(&repo_root); + } + + #[test] + fn an_unexpected_head_probe_failure_propagates_instead_of_using_read_tree_empty() { + let repo_root = unique_test_repo("head-probe-failure"); + init_repo(&repo_root); + std::fs::write(repo_root.join("file.txt"), b"content\n").expect("file should be writable"); + commit_all(&repo_root, "initial commit"); + + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + + std::fs::remove_file(service.git_dir.join("HEAD")) + .expect("HEAD file should be removable to simulate an unexpected failure"); + + let result = service.capture_tree(); + assert!( + result.is_err(), + "an unexpected HEAD-probe failure must surface as an error, not a false empty-baseline capture" + ); + + remove_test_repo(&repo_root); + } + + fn relative_path_from(base: &Path, target: &Path) -> PathBuf { + let base_components: Vec<_> = base.components().collect(); + let target_components: Vec<_> = target.components().collect(); + + let common_len = base_components + .iter() + .zip(target_components.iter()) + .take_while(|(a, b)| a == b) + .count(); + + let mut relative = PathBuf::new(); + for _ in common_len..base_components.len() { + relative.push(".."); + } + for component in &target_components[common_len..] { + relative.push(component); + } + relative + } + + #[test] + fn resolves_an_absolute_git_dir_from_a_relative_repository_root() { + let repo_root_abs = unique_test_repo("relative-root"); + init_repo(&repo_root_abs); + std::fs::write(repo_root_abs.join("file.txt"), b"relative content\n") + .expect("file should be writable"); + + let cwd = std::env::current_dir().expect("current dir should be readable"); + let canonical_repo_root = repo_root_abs + .canonicalize() + .expect("temp repo root should canonicalize"); + let relative_repo_root = relative_path_from(&cwd, &canonical_repo_root); + assert!( + relative_repo_root.is_relative(), + "test setup should produce a genuinely relative repository root" + ); + + let service = GitSnapshotService::new(&relative_repo_root) + .expect("service should resolve git-dir from a relative repository root"); + assert!( + service.git_dir.is_absolute(), + "GitSnapshotService must always store an absolute git_dir, even from a relative repository_root" + ); + + let tree = service + .capture_tree() + .expect("capture should succeed with a relative repository root"); + let ls_tree = run(&repo_root_abs, &["ls-tree", "-r", "--name-only", &tree.0]); + assert!(ls_tree.contains("file.txt")); + + service + .pin_tree(&worktree_id(), &tree) + .expect("pin should succeed with a relative repository root"); + let diff = service + .diff_trees(&tree, &tree) + .expect("diff should succeed with a relative repository root"); + assert!(diff.is_empty()); + + remove_test_repo(&repo_root_abs); + } + + #[test] + fn capture_on_unborn_head_with_no_files_produces_an_empty_tree() { + let repo_root = unique_test_repo("unborn-head-no-files"); + init_repo(&repo_root); + + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + let tree = service + .capture_tree() + .expect("capture should succeed on an unborn HEAD with no files at all"); + + let ls_tree = run(&repo_root, &["ls-tree", "-r", "--name-only", &tree.0]); + assert_eq!(ls_tree.trim(), ""); + + remove_test_repo(&repo_root); + } + + #[test] + fn snapshot_survives_a_fresh_process_and_temp_index_deletion() { + let repo_root = unique_test_repo("survives-process-exit"); + init_repo(&repo_root); + std::fs::write(repo_root.join("file.txt"), b"content\n").expect("file should be writable"); + + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + let tree = service.capture_tree().expect("capture should succeed"); + + let tmp_dir = service.git_dir.join(SCE_RUNTIME_DIR).join(TMP_INDEX_DIR); + assert!( + std::fs::read_dir(&tmp_dir).map_or(true, |mut entries| entries.next().is_none()), + "temp index directory should not retain the completed capture's index file" + ); + + let resolved = run(&repo_root, &["cat-file", "-t", &tree.0]); + assert_eq!(resolved.trim(), "tree"); + + remove_test_repo(&repo_root); + } + + #[test] + fn pinned_snapshot_survives_git_gc_prune_now() { + let repo_root = unique_test_repo("gc-prune-now"); + init_repo(&repo_root); + std::fs::write(repo_root.join("pinned.txt"), b"pinned content\n") + .expect("file should be writable"); + + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + let pinned_tree = service.capture_tree().expect("capture should succeed"); + service + .pin_tree(&worktree_id(), &pinned_tree) + .expect("pin should succeed"); + + std::fs::write( + repo_root.join("unreachable.txt"), + b"distinct unreachable content\n", + ) + .expect("file should be writable"); + let unpinned_tree = service.capture_tree().expect("capture should succeed"); + assert_ne!(pinned_tree.0, unpinned_tree.0); + + run(&repo_root, &["gc", "--prune=now"]); + + let pinned_resolved = run(&repo_root, &["cat-file", "-t", &pinned_tree.0]); + assert_eq!(pinned_resolved.trim(), "tree"); + + let unpinned_probe = Command::new("git") + .args(["cat-file", "-t", &unpinned_tree.0]) + .current_dir(&repo_root) + .output() + .expect("git cat-file should spawn"); + assert!( + !unpinned_probe.status.success(), + "an unpinned, unreachable tree should be reclaimed by git gc --prune=now" + ); + + remove_test_repo(&repo_root); + } + + #[test] + fn pinned_snapshot_survives_git_prune_expire_now() { + let repo_root = unique_test_repo("prune-expire-now"); + init_repo(&repo_root); + std::fs::write(repo_root.join("pinned.txt"), b"pinned content\n") + .expect("file should be writable"); + + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + let pinned_tree = service.capture_tree().expect("capture should succeed"); + service + .pin_tree(&worktree_id(), &pinned_tree) + .expect("pin should succeed"); + + std::fs::write( + repo_root.join("unreachable-2.txt"), + b"another distinct unreachable content\n", + ) + .expect("file should be writable"); + let unpinned_tree = service.capture_tree().expect("capture should succeed"); + assert_ne!(pinned_tree.0, unpinned_tree.0); + + run(&repo_root, &["prune", "--expire=now"]); + + let pinned_resolved = run(&repo_root, &["cat-file", "-t", &pinned_tree.0]); + assert_eq!(pinned_resolved.trim(), "tree"); + + let unpinned_probe = Command::new("git") + .args(["cat-file", "-t", &unpinned_tree.0]) + .current_dir(&repo_root) + .output() + .expect("git cat-file should spawn"); + assert!( + !unpinned_probe.status.success(), + "an unpinned, unreachable tree should be reclaimed by git prune --expire=now" + ); + + remove_test_repo(&repo_root); + } + + #[test] + fn pin_tree_is_idempotent_for_the_same_worktree_and_tree() { + let repo_root = unique_test_repo("pin-idempotent"); + init_repo(&repo_root); + std::fs::write(repo_root.join("file.txt"), b"content\n").expect("file should be writable"); + + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + let tree = service.capture_tree().expect("capture should succeed"); + + service + .pin_tree(&worktree_id(), &tree) + .expect("first pin should succeed"); + service + .pin_tree(&worktree_id(), &tree) + .expect("second, identical pin should be a harmless no-op"); + + let ref_name = pin_ref_name(&worktree_id(), &tree); + let resolved = run(&repo_root, &["rev-parse", &ref_name]); + assert_eq!(resolved.trim(), tree.0); + + remove_test_repo(&repo_root); + } + + #[test] + fn diff_trees_returns_parseable_git_diff_output() { + let repo_root = unique_test_repo("diff-trees"); + init_repo(&repo_root); + std::fs::write(repo_root.join("file.txt"), b"before\n").expect("file should be writable"); + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + let before = service.capture_tree().expect("capture should succeed"); + + std::fs::write(repo_root.join("file.txt"), b"after\n").expect("file should be writable"); + let after = service.capture_tree().expect("capture should succeed"); + + let diff = service + .diff_trees(&before, &after) + .expect("diff should succeed"); + assert!(diff.contains("file.txt")); + assert!(diff.contains("-before")); + assert!(diff.contains("+after")); + + remove_test_repo(&repo_root); + } + + #[test] + fn capture_and_pin_work_against_a_sha256_repository_when_supported() { + let repo_root = unique_test_repo("sha256"); + std::fs::create_dir_all(&repo_root).expect("repo root should be created"); + let init = Command::new("git") + .args(["init", "--quiet", "--object-format=sha256"]) + .current_dir(&repo_root) + .output() + .expect("git init should spawn"); + if !init.status.success() { + remove_test_repo(&repo_root); + return; + } + run(&repo_root, &["config", "user.email", "test@example.com"]); + run(&repo_root, &["config", "user.name", "Test"]); + std::fs::write(repo_root.join("file.txt"), b"content\n").expect("file should be writable"); + + let service = GitSnapshotService::new(&repo_root).expect("service should resolve git-dir"); + let tree = service + .capture_tree() + .expect("capture should succeed against a SHA-256 repository"); + service + .pin_tree(&worktree_id(), &tree) + .expect("pin should succeed against a SHA-256 repository"); + + let ls_tree = run(&repo_root, &["ls-tree", "-r", "--name-only", &tree.0]); + assert!(ls_tree.contains("file.txt")); + + remove_test_repo(&repo_root); + } +} diff --git a/cli/src/services/mutation_trace/runtime/mod.rs b/cli/src/services/mutation_trace/runtime/mod.rs index a503b454..57927a3f 100644 --- a/cli/src/services/mutation_trace/runtime/mod.rs +++ b/cli/src/services/mutation_trace/runtime/mod.rs @@ -1 +1,2 @@ +mod git_snapshot; mod worktree_lock; diff --git a/context/cli/mutation-trace-protocol.md b/context/cli/mutation-trace-protocol.md index 6c330947..3fc9bf7e 100644 --- a/context/cli/mutation-trace-protocol.md +++ b/context/cli/mutation-trace-protocol.md @@ -212,14 +212,15 @@ loads a `ProtocolState`; `protocol.rs` only transitions already-known ones. ## Target end-state architecture The plan's file split anticipated three seams beyond `protocol.rs`. `store.rs` -now exists as a real database call site (`mutation-cursor-store-persistence`); -`coordinator.rs`/`git_snapshot.rs` remain future work: +and `runtime/git_snapshot.rs` now exist as real call sites (see +[`mutation-trace-runtime-coordinator.md`](mutation-trace-runtime-coordinator.md)); +`coordinator.rs` remains future work: ```mermaid flowchart LR - coordinator["coordinator.rs\n(imperative shell:\nDB load, Git snapshot,\nCAS/retry, persist)"] + coordinator["coordinator.rs (future work)\n(imperative shell:\nDB load, Git snapshot,\nCAS/retry, persist)"] protocol["protocol.rs\n(pure transitions —\nprepare/commit/attribution/\ntaint/abandon/recover\nall implemented)"] - git_snapshot["git_snapshot.rs\n(isolated Git object store,\ntemporary index, tree capture/diff)"] + git_snapshot["runtime/git_snapshot.rs (implemented)\n(isolated Git snapshot,\ntemporary index, tree capture/diff,\nSCE-owned ref pinning)"] store["store.rs\n(cursor/revision, scopes,\nprocessed events, mutation\nevidence, CAS transaction)"] coordinator --> protocol @@ -227,20 +228,19 @@ flowchart LR coordinator --> store ``` -Each seam's responsibility: - -- **`coordinator.rs`** (future work) — receives hook/session identity, - resolves scope actor/worktree identity, asks `store.rs` to load the scope, - obtains a `ProtocolState`, and calls the pure protocol. +- **`coordinator.rs`** (future work) — receives hook/session identity, resolves + scope actor/worktree identity, loads a `ProtocolState` via `store.rs`, and + calls the pure protocol. +- **`runtime/git_snapshot.rs`** (implemented) — captures an isolated worktree + snapshot and pins it durably; not yet called by anything. - **`store.rs`** (implemented) — loads and persists worktree/scope/event state - via a CAS-guarded commit; atomically creates a new scope as `NeverSeen`; - never remaps `actor_kind`/`worktree_id` for an existing `ScopeId` (see - "Runtime scope materialization" above). -- **`protocol.rs`** — assumes referenced scopes are already represented in + via a CAS-guarded commit; never remaps `actor_kind`/`worktree_id` for an + existing `ScopeId` (see "Runtime scope materialization" above). +- **`protocol.rs`** — assumes referenced scopes already exist in `ProtocolState.scopes`; validates and transitions lifecycle state only. `protocol.rs` stays free of any Git object, DB row, or CAS transaction concept; -`coordinator.rs`/`git_snapshot.rs` are not created by this or the store plan. +`coordinator.rs` is not created by this or the store plan. ## Authoritative source diff --git a/context/cli/mutation-trace-runtime-coordinator.md b/context/cli/mutation-trace-runtime-coordinator.md index 4045f486..c0de17c9 100644 --- a/context/cli/mutation-trace-runtime-coordinator.md +++ b/context/cli/mutation-trace-runtime-coordinator.md @@ -19,7 +19,8 @@ documented convention. ## Current code surface -Only the per-worktree runtime lock exists so far. +The per-worktree runtime lock and the isolated Git snapshot service exist so +far; the coordinator that will drive both is future work. - `cli/src/services/mutation_trace/runtime/worktree_lock.rs` — `WorktreeLock::acquire(git_dir: &Path, timeout: Duration) -> @@ -33,11 +34,47 @@ Only the per-worktree runtime lock exists so far. file's mere on-disk existence is never treated as ownership — only a successful OS-level `try_lock()` counts, so a leftover lock file with no active OS lock held against it never blocks a fresh acquirer. - -This lock guards the coordinator's own critical section (snapshot capture, -worktree/scope materialization, recovery, and the CAS retry loop) once the -coordinator exists. It is held on every future `coordinate()` call, unlike -the checkout-identity-creation lock. +- `cli/src/services/mutation_trace/runtime/git_snapshot.rs` — + `GitSnapshotService::new(repository_root: &Path) -> Result` + resolves `git_dir` once via `git rev-parse --absolute-git-dir`, so + `git_dir` is always an absolute path — even when the caller's + `repository_root` is relative, which matters because every Git subprocess + this service spawns runs with `cwd = repository_root` and + `GIT_DIR = git_dir`; a relative `git_dir` would otherwise be resolved by + the child process against its own already-`repository_root`-joined `cwd`, + double-joining the path. `capture_tree(&self) -> Result` snapshots + the current worktree (staged, unstaged, untracked, and deleted state, + respecting `.gitignore`) into the repository's normal, shared Git object + database, never touching the real index or working tree: it reserves a + unique `/sce/tmp/index-` path via an RAII guard (never + pre-creating the file), probes `HEAD` via a dedicated `head_exists` + helper that inspects the Git exit status directly — status `0` means + `HEAD` resolves, status `1` is `--verify --quiet`'s documented "does not + resolve" signal (a genuinely unborn `HEAD`), and every other status + propagates as an error rather than being treated as empty, since HEAD + absence is a normal Git state but a HEAD-probe failure is a snapshot + failure — then runs `git read-tree HEAD` or, on a genuinely unborn `HEAD`, + the explicit `git read-tree --empty` (never a bare/absent index file), + then `git add -A -- .`, then `git write-tree`, all with only + `GIT_DIR`/`GIT_INDEX_FILE` set — no `GIT_OBJECT_DIRECTORY`/ + `GIT_ALTERNATE_OBJECT_DIRECTORIES` override anywhere. `TreeId` is an opaque + string; nothing assumes a fixed length, so a SHA-256 repository needs no + special handling. `pin_tree(&self, worktree_id, tree) -> Result<()>` makes + a tree durable by creating + `refs/sce/mutation-cursor//` via `git update-ref` — + create-only and idempotent for the same `(worktree_id, tree)` pair — which + is what makes a pinned tree survive `git gc --prune=now`/`git prune + --expire=now`, unlike an unpinned, unreachable tree in the same repository. + `diff_trees(&self, before, after) -> Result` runs `git diff + --binary --full-index --no-ext-diff --no-textconv` between two tree SHAs, + returning the raw diff text `patch.rs::parse_patch` already knows how to + parse. Nothing calls this service yet; `coordinator.rs` is its only planned + caller. + +The runtime lock guards the coordinator's own critical section (snapshot +capture, worktree/scope materialization, recovery, and the CAS retry loop) +once the coordinator exists. It is held on every future `coordinate()` call, +unlike the checkout-identity-creation lock. ## Two distinct locks, two distinct invariants @@ -57,7 +94,13 @@ On-disk layout so far: /sce/ ├── checkout-id (services::checkout) ├── checkout-id.lock (services::checkout) -└── mutation-cursor.lock (runtime::worktree_lock, this module) +├── mutation-cursor.lock (runtime::worktree_lock) +└── tmp/ + └── index- (runtime::git_snapshot, ephemeral per capture) + + (runtime::git_snapshot writes here directly) + +└── refs/sce/mutation-cursor// (runtime::git_snapshot, one ref per pinned tree, create-only) ``` ## Testing boundary @@ -71,10 +114,22 @@ held against it never blocking a fresh acquirer — each test uses a unique inline-unit-test precedent already used in `cli/src/services/checkout/mod.rs` and `cli/src/services/mutation_trace/store.rs` (see `context/patterns.md`). +`GitSnapshotService`'s inline `#[cfg(test)] mod tests` in `git_snapshot.rs` +uses the same precedent, extended to real per-test `git init` repositories: +index/working-tree preservation across staged/unstaged/untracked/deleted +state, `.gitignore` exclusion, unborn-`HEAD` capture with and without files, +an unexpected `HEAD`-probe failure (a corrupted/missing `.git/HEAD`) +propagating as an error rather than a false empty-baseline capture, a +relative `repository_root` still resolving `git_dir` absolute, survival +after the temp index file is gone, `git gc --prune=now`/`git prune +--expire=now` survival for a pinned tree versus reclamation of a distinct +unpinned tree in the same repository, `pin_tree` idempotency, and +`diff_trees` output shape. + ## Status -Only the per-worktree runtime lock (above) is implemented. The Git snapshot -service, the coordinator's protocol-integration pipeline, the public +The per-worktree runtime lock and the isolated Git snapshot service (above) +are implemented. The coordinator's protocol-integration pipeline, the public lock-wrapped `coordinate()` entrypoint, and cross-module integration tests remain future work tracked by the `mutation-cursor-runtime-coordinator` plan's task stack. This file will grow to cover them as each lands. diff --git a/context/cli/mutation-trace-store.md b/context/cli/mutation-trace-store.md index 1bba44fb..a66cdf3e 100644 --- a/context/cli/mutation-trace-store.md +++ b/context/cli/mutation-trace-store.md @@ -117,8 +117,11 @@ returns `Err`. ## Non-goals -- No Git or filesystem I/O — no `coordinator.rs` or `git_snapshot.rs` exists - yet; those remain future work building on this store. +- No Git or filesystem I/O — `store.rs` itself calls neither Git nor the + filesystem. `runtime/git_snapshot.rs` now exists as a sibling module (see + [`mutation-trace-runtime-coordinator.md`](mutation-trace-runtime-coordinator.md)) + but is not wired to this store yet; `coordinator.rs`, the future caller of + both, remains future work. - No attribution or boundary-kind decisions — `DurableTransition::between` and `store.rs` are both structurally blind to protocol meaning. - No retry-after-`Conflict` loop — `CasResult::Conflict` is returned to the diff --git a/context/context-map.md b/context/context-map.md index 9737aaa1..1afc46f4 100644 --- a/context/context-map.md +++ b/context/context-map.md @@ -23,11 +23,11 @@ Feature/domain context: - `context/cli/config-precedence-contract.md` (implemented `sce config` show/validate command contract, deterministic `flags > env > config file > defaults` resolution order, focused `config/resolver.rs` ownership for config discovery/merge/runtime precedence plus default-discovered invalid-file degradation, focused `config/render.rs` ownership for `show`/`validate` text+JSON output construction, canonical `$schema` acceptance for startup-loaded `sce/config.json` files, shared auth-key env/config/optional baked-default support starting with `workos_client_id`, shared runtime resolution for flat logging observability keys including config-file/default `log_to_file`, `log_dir` / `SCE_LOG_DIR` with `/sce/logs` defaulting plus config-file/default-only positive `log_file_retention_limit`, config-file-only `agent_trace.repository_id`/`agent_trace.repository_remote` repository-identity keys with default remote `origin`, default-enabled `agent_trace.auto_sync` boolean resolution with explicit-false opt-out for the post-commit trigger boundary, the catalog-derived `integrations.optional_workflows` optional-workflow selection key, JSON-pointer-prefixed schema-validation errors, canonical Pkl-generated `sce/config.json` schema ownership plus CLI embedding/reuse contract including `policies.attribution_hooks.enabled` default-true/explicit-false opt-out metadata, config-file selection order, `show` provenance output, and trimmed `validate` output contract) - `context/cli/capability-traits.md` (current broad CLI capability seam in `cli/src/services/capabilities.rs`, including `FsOps`/`StdFsOps`, `GitOps`/`ProcessGitOps`, git root/hooks resolution behavior, compile-time-typed borrowed AppContext wiring with associated-type narrow capability accessors plus `ContextWithRepoRoot` repo-root-scoped context derivation, generic command execution bounds, and test-only unimplemented stubs; current service internals do not consume fs/git traits until later lifecycle migration tasks) - `context/cli/service-lifecycle.md` (current compile-safe lifecycle seam in `cli/src/services/lifecycle.rs`, including default no-op `ServiceLifecycle` diagnose/fix/setup methods against narrow `HasRepoRoot`, lifecycle-owned health/fix/setup result types with generic setup messages, doctor/setup adapter boundaries, the static `LifecycleProvider` enum catalog/dispatcher, hook/config/local_db/auth_db/agent_trace_db lifecycle providers including setup-time repository-scoped Agent Trace DB initialization plus checkout identity diagnostics, implemented doctor aggregation over diagnose/fix providers, and implemented setup aggregation over `setup` providers in order config → local_db → auth_db → agent_trace_db → hooks when requested) -- `context/cli/mutation-trace-protocol.md` (pure, dependency-free `cli/src/services/mutation_trace/` refinement of the verified `spec/mutation_cursor.qnt` mutation-cursor protocol: the full pure action set — `prepare`/`commit` transition logic, attribution/mutation-event materialization, snapshot-failure/database-failure taint actions, scope abandonment, and recovery with an explicit observed-tree input — plus cross-action sequence/invariant tests and a module-level Quint refinement matrix are all complete (`types.rs`, `protocol.rs`, `mod.rs`, `tests.rs`, registered with `#[allow(dead_code)]`); opaque-newtype identity refinement decisions vs. the Quint model's bounded enums, and the `coordinator.rs`/`git_snapshot.rs` target end-state seams this layout leaves room for but does not create — `store.rs` now exists, built out by the `mutation-cursor-store-persistence` plan; `protocol.rs`'s pure transitions are not yet wired into any hook or command) +- `context/cli/mutation-trace-protocol.md` (pure, dependency-free `cli/src/services/mutation_trace/` refinement of the verified `spec/mutation_cursor.qnt` mutation-cursor protocol: the full pure action set — `prepare`/`commit` transition logic, attribution/mutation-event materialization, snapshot-failure/database-failure taint actions, scope abandonment, and recovery with an explicit observed-tree input — plus cross-action sequence/invariant tests and a module-level Quint refinement matrix are all complete (`types.rs`, `protocol.rs`, `mod.rs`, `tests.rs`, registered with `#[allow(dead_code)]`); opaque-newtype identity refinement decisions vs. the Quint model's bounded enums, and the `coordinator.rs` target end-state seam this layout leaves room for but does not create — `store.rs` and `runtime/git_snapshot.rs` now exist, built out by the `mutation-cursor-store-persistence` and `mutation-cursor-runtime-coordinator` plans respectively; `protocol.rs`'s pure transitions are not yet wired into any hook or command) - `context/cli/mutation-trace-revision-refinement.md` (the Quint `revision: int` → Rust `WorktreeState::revision: u64` bounded-integer refinement: the private `next_revision` checked-arithmetic helper `commit`/`taint`/`abandon`/`recover` all route through instead of a raw `+ 1`, so a worktree at `revision: u64::MAX` is a guarded no-op/rejection rather than a silent wrap to `0`) - `context/cli/mutation-trace-quint-connect.md` (`#[cfg(test)]`-only Quint Connect model-based-testing harness in `cli/src/services/mutation_trace/mbt/` continuously checking `protocol.rs` against `spec/mutation_cursor.qnt`: the verification-only `mbtAction`/`MbtAction` record-payload transport excluded from comparison, the operation-identity-vs-`MbtStutter` distinction on guarded/no-op branches with its two deterministic regressions, finite ID mapping, the AC5 comparable-state field list, `randomPrepare` staying a single `step` branch, deterministic/generated (500×30, seed-reproducible) test coverage, and the two Nix checks — generic `checks.cli-tests` and dedicated `checks.mutation-trace-quint-connect` — that both require the pinned Quint binary plus the top-level `spec/` directory in `workspaceSrc`'s Nix fileset) - `context/cli/mutation-trace-store.md` (durable persistence for the mutation-cursor protocol in `cli/src/services/mutation_trace/store.rs`, built by the `mutation-cursor-store-persistence` plan: the one-directional `protocol.rs` -> `DurableTransition::between` (pure structural diff) -> `store.rs` (SQL translation) -> `RepositoryAgentTraceDb` boundary; migration `003_mutation_trace_protocol.sql`'s five tables; `AttemptState`/`external_taint` non-persistence; the 8-byte big-endian `BLOB` revision encoding plus explicit non-`Debug` enum codecs; the bounded hot-path `load_worktree` read vs. the cold-path `load_mutation_event`; `commit`'s single-`BEGIN IMMEDIATE` CAS batch via `TursoDb::execute_transactional_cas_batch`, distinguishing `Conflict`/retryable-transient/deterministic-`Err` outcomes; and the store's non-goals — no Git/filesystem I/O, no attribution/boundary-kind decisions, no retry-after-`Conflict` loop, no terminal-scope garbage collection) -- `context/cli/mutation-trace-runtime-coordinator.md` (imperative-shell runtime layer in `cli/src/services/mutation_trace/runtime/`, built by the `mutation-cursor-runtime-coordinator` plan: currently only the per-worktree OS advisory lock, `runtime::worktree_lock::WorktreeLock::acquire(git_dir, timeout)` at `/sce/mutation-cursor.lock`, bounded `try_lock()` polling with a distinct matchable timeout error and RAII release, and its distinction from the separate checkout-identity-creation lock; the Git snapshot service, coordinator pipeline, public `coordinate()` entrypoint, and cross-module integration tests remain future work) +- `context/cli/mutation-trace-runtime-coordinator.md` (imperative-shell runtime layer in `cli/src/services/mutation_trace/runtime/`, built by the `mutation-cursor-runtime-coordinator` plan: the per-worktree OS advisory lock, `runtime::worktree_lock::WorktreeLock::acquire(git_dir, timeout)` at `/sce/mutation-cursor.lock`, bounded `try_lock()` polling with a distinct matchable timeout error and RAII release, and its distinction from the separate checkout-identity-creation lock; and the isolated Git snapshot service, `runtime::git_snapshot::GitSnapshotService` (`capture_tree`/`pin_tree`/`diff_trees`), writing tree/blob objects into the repository's normal object database and pinning durable ones via a create-only, idempotent `refs/sce/mutation-cursor//` ref rather than a private object store; the coordinator pipeline, public `coordinate()` entrypoint, and cross-module integration tests remain future work) - `context/sce/cli-exit-code-contract.md` (stable class-based `sce` process exit-code contract in `cli/src/app.rs`, so automation can branch on failure category without parsing error text) - `context/sce/cli-error-code-taxonomy.md` (stable user-facing `SCE-ERR-*` diagnostic code classes rendered by `cli/src/app.rs`, complementing the numeric exit-code classes) - `context/sce/cli-stdout-stderr-contract.md` (implemented stream contract in `cli/src/app.rs`: command payloads on stdout only, redacted diagnostics on stderr) diff --git a/context/plans/mutation-cursor-runtime-coordinator.md b/context/plans/mutation-cursor-runtime-coordinator.md index 5a872cce..cb929824 100644 --- a/context/plans/mutation-cursor-runtime-coordinator.md +++ b/context/plans/mutation-cursor-runtime-coordinator.md @@ -499,7 +499,7 @@ Persist this field in every plan; this is durable plan state, not chat state: during plan authoring, not originated by this task's execution. - Context synchronization: synced -- [ ] T03: `Add the isolated Git snapshot service with ref-pinned durability` (status:todo) +- [x] T03: `Add the isolated Git snapshot service with ref-pinned durability` (status:done) - Task ID: T03 - Scope: In — `runtime/git_snapshot.rs`: `GitSnapshotService` (resolves `--git-dir` once; a unique, never-pre-created private temp index path @@ -527,7 +527,159 @@ Persist this field in every plan; this is durable plan state, not chat state: same repository is reclaimed by that same aggressive pass; `pin_tree` is idempotent for the same `(worktree_id, tree)` pair. - Verify: `cargo test -p shared-context-engineering runtime::git_snapshot::` (via `./scripts/run-cli-cargo.sh test`) - - Context synchronization: pending + - Completed: 2026-08-30 + - Files changed: `cli/src/services/mutation_trace/runtime/git_snapshot.rs` (new), + `cli/src/services/mutation_trace/runtime/mod.rs` (added `mod git_snapshot;`) + - Result: Added `runtime/git_snapshot.rs` with a `pub struct GitSnapshotService` + (private module, matching `worktree_lock`'s module-privacy convention — + reachable from sibling modules under `runtime`, not outside it). + `GitSnapshotService::new(repository_root)` resolves `--git-dir` once via + `git rev-parse --git-dir` (the same resolution logic as + `checkout::resolve_git_dir`, duplicated locally rather than imported, + since `checkout`'s function is scoped to that module's own concerns and + this plan's Constraints exclude broad Git-abstraction refactors) and + stores both the resolved `git_dir` and `repository_root`. + + `capture_tree`: reserves a unique `/sce/tmp/index-` path + via a private `TempIndexGuard` (constructs the path only, creates the + `tmp/` directory, never touches the index file itself, best-effort + removes it on `Drop`); probes `HEAD` via + `git rev-parse --verify --quiet HEAD`; runs `git read-tree HEAD` when + `HEAD` exists or `git read-tree --empty` on an unborn `HEAD`, both with + only `GIT_DIR`/`GIT_INDEX_FILE` set and `cwd = repository_root`; then + `git add -A -- .`; then `git write-tree`, returning the raw trimmed SHA + wrapped in `TreeId`, never assuming a fixed length. `pin_tree` + (`git update-ref refs/sce/mutation-cursor// + `, only `GIT_DIR` set) and `diff_trees` (`git diff --binary + --full-index --no-ext-diff --no-textconv `, only + `GIT_DIR` set, returning the raw `String`) match the plan's validated + command sequences exactly. No `GIT_OBJECT_DIRECTORY`/ + `GIT_ALTERNATE_OBJECT_DIRECTORIES` environment variable is set anywhere. + + Added a `#[cfg(test)] mod tests` (real, per-test temporary Git + repositories via `git init`, following T01/T02's unique-temp-dir + convention) covering: staged/unstaged/untracked/deleted state captured + correctly while the real index, `git status --porcelain`, `git diff`, + and `git diff --cached` stay byte-identical before and after capture; + `.gitignore` exclusion; a committed-file deletion reflected as absent; + unborn `HEAD` with an untracked file present and with no files at all + (asserting an empty `ls-tree`); a captured tree remaining resolvable via + `git cat-file` after the temp index file is gone; a `pin_tree`-protected + tree surviving both `git gc --prune=now` and `git prune --expire=now` + while a distinct-content, genuinely unreachable tree captured and left + unpinned in the same repository is reclaimed by the same pass (run as + two separate tests, one per pruning command, both using distinct-content + trees per the plan's corrected experimental design — never a + same-content decoy); `pin_tree` idempotency for the same + `(worktree_id, tree)` pair; `diff_trees` producing `git diff --git` + formatted, parseable output; and a best-effort SHA-256 repository case + (`git init --object-format=sha256`) that skips gracefully when the local + Git build lacks the flag, asserting `TreeId` needs no length assumption. + + PR #244 review follow-up: two correctness issues were fixed without + changing the snapshot design. First, HEAD-absence and HEAD-probe-failure + semantics were conflated: `capture_tree` previously ran `git rev-parse + --verify --quiet HEAD` through the shared `run_git` helper and used + `.is_ok()` to pick between `read-tree HEAD`/`read-tree --empty`, so *any* + Git failure — a corrupted repository, a missing `.git/HEAD`, an + unexpected fatal error — was silently treated as "unborn HEAD" and + produced a false empty-baseline snapshot instead of surfacing an error. + A dedicated `head_exists(&self) -> Result` now runs that probe + directly and inspects `output.status.code()`: exit status `0` (HEAD + resolves) → `Ok(true)`; exit status `1` (`--verify --quiet`'s documented + "does not resolve" signal — verified experimentally against this + repository's Git as the exit status for a genuinely unborn HEAD) → + `Ok(false)`; every other status (verified experimentally: a corrupted or + missing `.git/HEAD` produces exit status `128`, "fatal: not a git + repository") → `Err`, propagated by `capture_tree` via `?` before ever + reaching `read-tree --empty`. HEAD absence is a normal Git state; HEAD + probe failure is a snapshot failure — the two are no longer conflated. + + Second, `resolve_git_dir` used `git rev-parse --git-dir` and manually + joined a relative result onto `repository_root` when not already + absolute — but that join only produces a genuinely absolute path when + `repository_root` itself is absolute. Given a relative + `repository_root`, the stored `git_dir` stayed relative, and since every + later Git subprocess runs with `cwd = repository_root` (also relative) + and `GIT_DIR = self.git_dir`, the relative `GIT_DIR` was resolved by the + child process against its own (already `repository_root`-joined) `cwd`, + double-joining the path and pointing at the wrong directory. + `resolve_git_dir` now runs `git rev-parse --absolute-git-dir` instead — + verified experimentally to return a canonicalized absolute path for both + a normal clone and a linked worktree — eliminating the manual + absolute/relative branch entirely, with a `debug_assert!` documenting the + invariant. `GitSnapshotService` now always stores an absolute `git_dir`, + regardless of whether the caller's `repository_root` was relative or + absolute; `repository_root` itself is unchanged (still used only as + `Command::current_dir`, which is resolution-agnostic). + + Added two regression tests: + `an_unexpected_head_probe_failure_propagates_instead_of_using_read_tree_empty` + (deletes `.git/HEAD` after constructing the service — a real Git failure + distinct from unborn HEAD — and asserts `capture_tree()` returns `Err` + rather than a false empty-baseline capture) and + `resolves_an_absolute_git_dir_from_a_relative_repository_root` (computes + a genuinely relative `repository_root` from the real, unmutated process + `cwd` via a small test-local `relative_path_from` helper — no + `std::env::set_current_dir`, so the test is parallel-safe — and proves + `GitSnapshotService::new`, `capture_tree`, `pin_tree`, and `diff_trees` + all succeed with `git_dir` resolved absolute). The existing unborn-HEAD + tests (`capture_on_unborn_head_with_a_file_produces_a_valid_tree`, + `capture_on_unborn_head_with_no_files_produces_an_empty_tree`) continue + to prove the genuinely-unborn path still produces a valid `read-tree + --empty` snapshot; the new failure test proves the two paths no longer + share a code path they should not share. + - Verify (actual): `./scripts/run-cli-cargo.sh test --manifest-path + cli/Cargo.toml runtime::git_snapshot::` → 11/11 passed (pre-follow-up); + 13/13 passed (post-follow-up, including the two new regression tests). + Additionally ran `./scripts/run-cli-cargo.sh test --manifest-path + cli/Cargo.toml runtime::` → 20/20 passed (pre-follow-up), 22/22 passed + (post-follow-up; no regression in `worktree_lock`'s existing tests). + Additionally ran `./scripts/run-cli-cargo.sh clippy --manifest-path + cli/Cargo.toml --all-targets -- -D warnings` → clean both times + (pedantic/warnings denied workspace-wide, per Constraints; fixed two + pedantic findings — `uninlined_format_args` and `map_unwrap_or` — during + the original implementation; the follow-up introduced none), and + `./scripts/run-cli-cargo.sh fmt --manifest-path cli/Cargo.toml -- --check` + → clean both times. + - Context impact: `domain` — introduces a new architectural element + (`GitSnapshotService`, the isolated Git snapshot/ref-pinning service) + under the mutation-cursor-runtime-coordinator domain T02 already + established a domain file for. Updated + `context/cli/mutation-trace-runtime-coordinator.md` to document + `GitSnapshotService` (`capture_tree`/`pin_tree`/`diff_trees`), the + updated on-disk layout (`sce/tmp/index-`, the + `refs/sce/mutation-cursor/**` namespace), its test coverage, and the + revised Status section. Corrected two now-stale statements this task's + landing directly falsified: `context/cli/mutation-trace-protocol.md`'s + "Target end-state architecture" said `coordinator.rs`/`git_snapshot.rs` + "remain future work" — updated its prose, Mermaid diagram, and + per-seam bullets to mark `runtime/git_snapshot.rs` implemented while + `coordinator.rs` remains future work; `context/cli/mutation-trace-store.md`'s + "Non-goals" said "no `coordinator.rs` or `git_snapshot.rs` exists yet" — + corrected to note `runtime/git_snapshot.rs` now exists as a sibling + module, not yet wired to the store. Updated `context/context-map.md`'s + entries for both files accordingly. No repository-wide behavior, + architecture, or terminology changed, so `overview.md`, `architecture.md`, + `glossary.md`, and `patterns.md` were verified against the change and + left unedited. No qualifying system-wide architecture decision was + introduced by this task — the ref-pinning/normal-object-database design + was already established in the plan's own Design decisions section + during plan authoring, not originated by this task's execution. + + PR #244 review follow-up: corrected `context/cli/mutation-trace-runtime-coordinator.md`'s + `git_snapshot.rs` bullet, which the follow-up fix directly falsified — + it said `GitSnapshotService::new` "resolves `--git-dir` once" (now + `--absolute-git-dir`, with the absolute-path invariant and the reason it + matters explained inline) and described the HEAD branch without + distinguishing probe failure from genuine absence (now documents the + `head_exists` exit-status distinction). Also extended the Testing + boundary section's `git_snapshot.rs` test-coverage summary to name the + two new regression tests. No other context file was affected by this + follow-up: it is an internal correctness fix with no public-interface, + behavior-contract, or architecture change beyond what the two + corrections above capture. + - Context synchronization: synced - [ ] T04: `Add the coordinator's core protocol-integration pipeline` (status:todo) - Task ID: T04 From 23a23e36fff14cc106206bcd53167b6d4ed5f38e Mon Sep 17 00:00:00 2001 From: David Abram Date: Sun, 30 Aug 2026 08:14:15 +0200 Subject: [PATCH 5/8] runtime: Implement the mutation trace coordinator pipeline The mutation-cursor runtime needs an imperative integration layer that turns Git observations and boundary events into durable protocol state without duplicating evidence or losing concurrent updates. Add the generic coordinator pipeline with snapshot capture and pinning, idempotent worktree and scope materialization, recovery handling, bounded CAS reload/recompute retries, snapshot-failure taint persistence, and focused concurrency and lifecycle tests. Synchronize the mutation-trace context documents and mark the completed plan task. The public lock-wrapped coordinate() entrypoint remains follow-up work; this change keeps the internal pipeline deterministic through the SnapshotCapture seam. Plan: mutation-cursor-runtime-coordinator (T04) Co-authored-by: SCE --- .../mutation_trace/runtime/coordinator.rs | 1231 +++++++++++++++++ .../services/mutation_trace/runtime/mod.rs | 1 + context/cli/mutation-trace-protocol.md | 24 +- .../cli/mutation-trace-runtime-coordinator.md | 84 +- context/cli/mutation-trace-store.md | 15 +- context/context-map.md | 4 +- .../mutation-cursor-runtime-coordinator.md | 217 ++- 7 files changed, 1537 insertions(+), 39 deletions(-) create mode 100644 cli/src/services/mutation_trace/runtime/coordinator.rs diff --git a/cli/src/services/mutation_trace/runtime/coordinator.rs b/cli/src/services/mutation_trace/runtime/coordinator.rs new file mode 100644 index 00000000..671a9ac2 --- /dev/null +++ b/cli/src/services/mutation_trace/runtime/coordinator.rs @@ -0,0 +1,1231 @@ +use anyhow::Result; +use uuid::Uuid; + +use crate::services::agent_trace_db::repository::RepositoryAgentTraceDb; +use crate::services::mutation_trace::protocol; +use crate::services::mutation_trace::store::{CasResult, DurableTransition, MutationTraceStore}; +use crate::services::mutation_trace::types::{ + self, ActorKind, AttemptId, Boundary, EventId, MutationEvent, ScopeId, TreeId, WorktreeId, +}; + +use super::git_snapshot::GitSnapshotService; + +const MAX_CAS_RETRY_ATTEMPTS: u32 = 5; + +#[derive(Clone, Debug)] +pub enum RuntimeBoundary { + Start { + scope: ScopeId, + event: EventId, + actor_kind: ActorKind, + }, + Advance { + scope: ScopeId, + event: EventId, + actor_kind: ActorKind, + }, + Close { + scope: ScopeId, + event: EventId, + actor_kind: ActorKind, + }, + Flush, +} + +#[derive(Debug)] +pub struct CoordinateOutcome { + pub worktree_id: WorktreeId, + pub observed_tree: TreeId, + pub revision: u64, + pub evaluation: protocol::CommitEvaluation, + pub mutation_event: Option, +} + +#[derive(Debug)] +pub enum CoordinateError { + SnapshotFailure { + persisted_taint: bool, + source: anyhow::Error, + }, + ScopeIdentityConflict(anyhow::Error), + CasConflictExhausted { + attempts: u32, + }, + RevisionExhausted { + worktree_id: WorktreeId, + revision: u64, + }, + LockAcquisition(anyhow::Error), + Other(anyhow::Error), +} + +impl std::fmt::Display for CoordinateError { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + match self { + CoordinateError::SnapshotFailure { + persisted_taint, + source, + } => write!( + f, + "Git snapshot capture/pin failed (taint persisted: {persisted_taint}): {source}" + ), + CoordinateError::CasConflictExhausted { attempts } => { + write!(f, "Exhausted {attempts} CAS-conflict retry attempts") + } + CoordinateError::RevisionExhausted { + worktree_id, + revision, + } => write!( + f, + "Worktree {worktree_id:?} requires recovery but its revision \ + ({revision}) cannot be advanced" + ), + CoordinateError::ScopeIdentityConflict(source) + | CoordinateError::LockAcquisition(source) + | CoordinateError::Other(source) => write!(f, "{source}"), + } + } +} + +impl std::error::Error for CoordinateError {} + +pub trait SnapshotCapture { + fn capture(&self) -> Result; + fn pin(&self, worktree_id: &WorktreeId, tree: &TreeId) -> Result<()>; +} + +impl SnapshotCapture for GitSnapshotService { + fn capture(&self) -> Result { + self.capture_tree() + } + + fn pin(&self, worktree_id: &WorktreeId, tree: &TreeId) -> Result<()> { + self.pin_tree(worktree_id, tree) + } +} + +fn coordinate_boundary( + db: &RepositoryAgentTraceDb, + capture: &C, + worktree_id: &WorktreeId, + boundary: &RuntimeBoundary, +) -> Result { + let store = MutationTraceStore::new(db); + + let observed_tree = match capture.capture().and_then(|tree| { + capture.pin(worktree_id, &tree)?; + Ok(tree) + }) { + Ok(tree) => tree, + Err(source) => return Err(handle_snapshot_failure(&store, worktree_id, source)), + }; + + store + .initialize_worktree(worktree_id, &observed_tree) + .map_err(CoordinateError::Other)?; + + if let Some((scope, actor_kind)) = hook_identity(boundary) { + store + .register_scope(scope, worktree_id, actor_kind) + .map_err(CoordinateError::ScopeIdentityConflict)?; + } + + let type_boundary = into_protocol_boundary(boundary, worktree_id); + let scope_ref = types::boundary_scope(&type_boundary); + let event_key_ref = types::boundary_event_key(&type_boundary); + + for _ in 0..MAX_CAS_RETRY_ATTEMPTS { + let Some(projection) = store + .load_worktree(worktree_id, scope_ref.as_ref(), event_key_ref.as_ref()) + .map_err(CoordinateError::Other)? + else { + return Err(CoordinateError::Other(anyhow::anyhow!( + "worktree {worktree_id:?} missing durable state immediately after initialize_worktree" + ))); + }; + + let mut state = projection.into_protocol_state(); + + if needs_recovery(&state, worktree_id) { + let recovered = protocol::recover(&state, worktree_id, observed_tree.clone()); + let Some(transition) = DurableTransition::between(&state, &recovered, worktree_id) + .map_err(CoordinateError::Other)? + else { + let revision = state + .worktrees + .get(worktree_id) + .map(|worktree_state| worktree_state.revision) + .unwrap_or_default(); + return Err(CoordinateError::RevisionExhausted { + worktree_id: worktree_id.clone(), + revision, + }); + }; + + match store.commit(&transition).map_err(CoordinateError::Other)? { + CasResult::Applied => state = recovered, + CasResult::Conflict => continue, + } + } + + let attempt = AttemptId(Uuid::new_v4().to_string()); + let prepared = protocol::prepare( + &state, + attempt.clone(), + type_boundary.clone(), + observed_tree.clone(), + ); + let outcome = protocol::commit(&prepared, &attempt); + + match DurableTransition::between(&state, &outcome.state, worktree_id) + .map_err(CoordinateError::Other)? + { + Some(transition) => match store.commit(&transition).map_err(CoordinateError::Other)? { + CasResult::Applied => { + return Ok(build_outcome(worktree_id, observed_tree, &outcome)) + } + CasResult::Conflict => {} + }, + None => return Ok(build_outcome(worktree_id, observed_tree, &outcome)), + } + } + + Err(CoordinateError::CasConflictExhausted { + attempts: MAX_CAS_RETRY_ATTEMPTS, + }) +} + +fn needs_recovery(state: &types::ProtocolState, worktree_id: &WorktreeId) -> bool { + state.external_taint.contains(worktree_id) + || state + .worktrees + .get(worktree_id) + .is_some_and(|worktree_state| worktree_state.tainted || worktree_state.needs_rebaseline) +} + +fn build_outcome( + worktree_id: &WorktreeId, + observed_tree: TreeId, + outcome: &protocol::CommitOutcome, +) -> CoordinateOutcome { + let revision = outcome + .state + .worktrees + .get(worktree_id) + .map(|worktree_state| worktree_state.revision) + .expect("the coordinated worktree's durable state must still be present after commit"); + let mutation_event = outcome.state.mutation_events.iter().next().cloned(); + + CoordinateOutcome { + worktree_id: worktree_id.clone(), + observed_tree, + revision, + evaluation: outcome.evaluation, + mutation_event, + } +} + +fn into_protocol_boundary(boundary: &RuntimeBoundary, worktree_id: &WorktreeId) -> Boundary { + match boundary { + RuntimeBoundary::Start { scope, event, .. } => Boundary::Start { + scope: scope.clone(), + event: event.clone(), + }, + RuntimeBoundary::Advance { scope, event, .. } => Boundary::Advance { + scope: scope.clone(), + event: event.clone(), + }, + RuntimeBoundary::Close { scope, event, .. } => Boundary::Close { + scope: scope.clone(), + event: event.clone(), + }, + RuntimeBoundary::Flush => Boundary::Flush { + worktree: worktree_id.clone(), + }, + } +} + +fn hook_identity(boundary: &RuntimeBoundary) -> Option<(&ScopeId, ActorKind)> { + match boundary { + RuntimeBoundary::Start { + scope, actor_kind, .. + } + | RuntimeBoundary::Advance { + scope, actor_kind, .. + } + | RuntimeBoundary::Close { + scope, actor_kind, .. + } => Some((scope, *actor_kind)), + RuntimeBoundary::Flush => None, + } +} + +fn handle_snapshot_failure( + store: &MutationTraceStore<'_>, + worktree_id: &WorktreeId, + source: anyhow::Error, +) -> CoordinateError { + match run_taint_retry_loop(store, worktree_id) { + Ok(persisted_taint) => CoordinateError::SnapshotFailure { + persisted_taint, + source, + }, + Err(db_err) => CoordinateError::Other(db_err), + } +} + +fn run_taint_retry_loop(store: &MutationTraceStore<'_>, worktree_id: &WorktreeId) -> Result { + run_taint_retry_loop_inner(store, worktree_id, |_attempt| {}) +} + +fn run_taint_retry_loop_inner( + store: &MutationTraceStore<'_>, + worktree_id: &WorktreeId, + mut after_load: F, +) -> Result +where + F: FnMut(u32), +{ + for attempt in 0..MAX_CAS_RETRY_ATTEMPTS { + let Some(projection) = store.load_worktree(worktree_id, None, None)? else { + return Ok(false); + }; + after_load(attempt); + + let state = projection.into_protocol_state(); + let tainted_state = protocol::taint(&state, worktree_id); + match DurableTransition::between(&state, &tainted_state, worktree_id)? { + None => { + let currently_tainted = state + .worktrees + .get(worktree_id) + .is_some_and(|worktree_state| worktree_state.tainted); + return Ok(currently_tainted); + } + Some(transition) => match store.commit(&transition)? { + CasResult::Applied => return Ok(true), + CasResult::Conflict => {} + }, + } + } + Ok(false) +} + +#[cfg(test)] +mod tests { + use std::cell::{Cell, RefCell}; + use std::collections::VecDeque; + use std::sync::atomic::{AtomicU64, Ordering}; + + use super::*; + use crate::services::mutation_trace::store::encode_revision; + use crate::services::mutation_trace::types::{Attribution, FailureKind, ScopeStatus}; + + static NEXT_TEST_DB_ID: AtomicU64 = AtomicU64::new(0); + + fn unique_test_db_path(label: &str) -> std::path::PathBuf { + let id = NEXT_TEST_DB_ID.fetch_add(1, Ordering::Relaxed); + std::env::temp_dir() + .join(format!( + "sce-mutation-trace-coordinator-{label}-{}-{id}", + std::process::id() + )) + .join("agent-trace.db") + } + + fn test_db(label: &str) -> (RepositoryAgentTraceDb, std::path::PathBuf) { + let db_path = unique_test_db_path(label); + let db = RepositoryAgentTraceDb::new_at(&db_path).expect("test db should open"); + (db, db_path) + } + + fn remove_test_db(db_path: &std::path::Path) { + if let Some(parent) = db_path.parent() { + let _ = std::fs::remove_dir_all(parent); + } + } + + enum FakeOutcome { + Succeed(TreeId), + Fail(String), + } + + struct FakeSnapshotCapture { + outcomes: RefCell>, + default_tree: TreeId, + capture_calls: Cell, + pin_calls: Cell, + } + + impl FakeSnapshotCapture { + fn new(default_tree: TreeId) -> Self { + Self { + outcomes: RefCell::new(VecDeque::new()), + default_tree, + capture_calls: Cell::new(0), + pin_calls: Cell::new(0), + } + } + + fn push_success(&self, tree: TreeId) { + self.outcomes + .borrow_mut() + .push_back(FakeOutcome::Succeed(tree)); + } + + fn push_failure(&self, message: &str) { + self.outcomes + .borrow_mut() + .push_back(FakeOutcome::Fail(message.to_string())); + } + + fn capture_call_count(&self) -> u32 { + self.capture_calls.get() + } + + fn pin_call_count(&self) -> u32 { + self.pin_calls.get() + } + } + + impl SnapshotCapture for FakeSnapshotCapture { + fn capture(&self) -> Result { + self.capture_calls.set(self.capture_calls.get() + 1); + match self.outcomes.borrow_mut().pop_front() { + Some(FakeOutcome::Succeed(tree)) => Ok(tree), + Some(FakeOutcome::Fail(message)) => Err(anyhow::anyhow!(message)), + None => Ok(self.default_tree.clone()), + } + } + + fn pin(&self, _worktree_id: &WorktreeId, _tree: &TreeId) -> Result<()> { + self.pin_calls.set(self.pin_calls.get() + 1); + Ok(()) + } + } + + struct HookedFailingCapture<'a> { + message: String, + hook: RefCell>>, + } + + impl<'a> HookedFailingCapture<'a> { + fn new(message: &str, hook: impl FnMut() + 'a) -> Self { + Self { + message: message.to_string(), + hook: RefCell::new(Some(Box::new(hook))), + } + } + } + + impl SnapshotCapture for HookedFailingCapture<'_> { + fn capture(&self) -> Result { + if let Some(hook) = self.hook.borrow_mut().as_mut() { + hook(); + } + Err(anyhow::anyhow!(self.message.clone())) + } + + fn pin(&self, _worktree_id: &WorktreeId, _tree: &TreeId) -> Result<()> { + Ok(()) + } + } + + fn commit_competing_advance( + store: &MutationTraceStore<'_>, + worktree: &WorktreeId, + scope: &ScopeId, + event_id: &str, + first: bool, + ) { + if first { + store + .register_scope(scope, worktree, ActorKind::ClaudeCode) + .expect("competing register_scope should succeed"); + } + + let loaded = store + .load_worktree(worktree, Some(scope), None) + .expect("competing load should succeed") + .expect("worktree should already be materialized") + .into_protocol_state(); + let cursor_tree = loaded.worktrees[worktree].cursor_tree.clone(); + let boundary = if first { + Boundary::Start { + scope: scope.clone(), + event: EventId(event_id.to_string()), + } + } else { + Boundary::Advance { + scope: scope.clone(), + event: EventId(event_id.to_string()), + } + }; + let attempt_id = AttemptId(format!("competing-{event_id}")); + let prepared = protocol::prepare(&loaded, attempt_id.clone(), boundary, cursor_tree); + let outcome = protocol::commit(&prepared, &attempt_id); + let transition = DurableTransition::between(&loaded, &outcome.state, worktree) + .expect("diff should succeed") + .expect("each competing advance should durably bump the revision"); + assert_eq!( + store + .commit(&transition) + .expect("competing commit should not error"), + CasResult::Applied, + "the competing commit should land before this iteration's own commit" + ); + } + + fn insert_worktree_at_revision( + db: &RepositoryAgentTraceDb, + worktree_id: &str, + cursor_tree: &str, + revision: u64, + tainted: bool, + failure_kind: &str, + needs_rebaseline: bool, + ) { + db.execute( + "INSERT INTO mutation_trace_worktrees + (worktree_id, cursor_tree, revision, tainted, failure_kind, needs_rebaseline) + VALUES (?1, ?2, ?3, ?4, ?5, ?6)", + ( + worktree_id, + cursor_tree, + encode_revision(revision).as_slice(), + tainted, + failure_kind, + needs_rebaseline, + ), + ) + .expect("worktree insert should succeed"); + } + + #[test] + fn first_observation_establishes_baseline_without_evidence() { + let (db, db_path) = test_db("ac1-first-observation"); + let worktree = WorktreeId("wt-1".to_string()); + let capture = FakeSnapshotCapture::new(TreeId("tree-a".to_string())); + + let outcome = coordinate_boundary(&db, &capture, &worktree, &RuntimeBoundary::Flush) + .expect("first observation should succeed"); + + assert_eq!(outcome.observed_tree, TreeId("tree-a".to_string())); + assert_eq!( + outcome.revision, 0, + "a Flush that observes no change from the freshly established baseline must not advance the revision" + ); + assert!( + outcome.mutation_event.is_none(), + "no mutation evidence should be emitted for the worktree's first observation" + ); + assert_eq!(capture.capture_call_count(), 1); + assert_eq!(capture.pin_call_count(), 1); + + remove_test_db(&db_path); + } + + #[test] + fn exclusive_edit_between_start_and_advance_commits_one_event() { + let (db, db_path) = test_db("ac2-exclusive-edit"); + let worktree = WorktreeId("wt-1".to_string()); + let scope = ScopeId("scope-1".to_string()); + let capture = FakeSnapshotCapture::new(TreeId("tree-a".to_string())); + + coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Start { + scope: scope.clone(), + event: EventId("evt-start".to_string()), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect("start should succeed"); + + capture.push_success(TreeId("tree-b".to_string())); + let outcome = coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Advance { + scope: scope.clone(), + event: EventId("evt-advance".to_string()), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect("advance should succeed"); + + let event = outcome + .mutation_event + .expect("an edit observed between Start and Advance should commit exactly one event"); + assert_eq!(event.before_tree, TreeId("tree-a".to_string())); + assert_eq!(event.after_tree, TreeId("tree-b".to_string())); + assert_eq!(event.attribution, Attribution::AiExclusive(scope)); + + remove_test_db(&db_path); + } + + #[test] + fn replaying_the_same_scope_event_key_does_not_duplicate_evidence() { + let (db, db_path) = test_db("ac3-replay"); + let worktree = WorktreeId("wt-1".to_string()); + let scope = ScopeId("scope-1".to_string()); + let capture = FakeSnapshotCapture::new(TreeId("tree-a".to_string())); + + coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Start { + scope: scope.clone(), + event: EventId("evt-start".to_string()), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect("start should succeed"); + + let advance = RuntimeBoundary::Advance { + scope: scope.clone(), + event: EventId("evt-advance".to_string()), + actor_kind: ActorKind::ClaudeCode, + }; + + capture.push_success(TreeId("tree-b".to_string())); + let first = coordinate_boundary(&db, &capture, &worktree, &advance) + .expect("first advance should commit"); + assert!(first.mutation_event.is_some()); + + capture.push_success(TreeId("tree-b".to_string())); + let replay = coordinate_boundary(&db, &capture, &worktree, &advance).expect( + "replaying the identical (scope, event) boundary must be a no-op, not an error", + ); + assert!( + replay.mutation_event.is_none(), + "replay must not duplicate mutation evidence" + ); + assert_eq!( + replay.revision, first.revision, + "replay must not advance the revision beyond the original commit" + ); + + remove_test_db(&db_path); + } + + #[test] + fn close_boundary_attributes_using_pre_close_scope_set() { + let (db, db_path) = test_db("ac4-close-attribution"); + let worktree = WorktreeId("wt-1".to_string()); + let scope = ScopeId("scope-1".to_string()); + let capture = FakeSnapshotCapture::new(TreeId("tree-a".to_string())); + + coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Start { + scope: scope.clone(), + event: EventId("evt-start".to_string()), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect("start should succeed"); + + capture.push_success(TreeId("tree-b".to_string())); + let outcome = coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Close { + scope: scope.clone(), + event: EventId("evt-close".to_string()), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect("close should succeed"); + + let event = outcome + .mutation_event + .expect("a real tree change observed at Close should still commit an event"); + assert_eq!( + event.attribution, + Attribution::AiExclusive(scope.clone()), + "the Close boundary's own emitted event must still attribute to the scope it is about to close" + ); + + let store = MutationTraceStore::new(&db); + let projection = store + .load_worktree(&worktree, Some(&scope), None) + .expect("load should succeed") + .expect("worktree should exist"); + assert_eq!( + projection.scopes.get(&scope).map(|s| s.status), + Some(ScopeStatus::Closed) + ); + + remove_test_db(&db_path); + } + + fn assert_contended_attribution(label: &str, actor_a: ActorKind, actor_b: ActorKind) { + let (db, db_path) = test_db(label); + let worktree = WorktreeId("wt-1".to_string()); + let capture = FakeSnapshotCapture::new(TreeId("tree-a".to_string())); + + coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Start { + scope: ScopeId("scope-a".to_string()), + event: EventId("evt-start-a".to_string()), + actor_kind: actor_a, + }, + ) + .expect("starting scope a should succeed"); + coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Start { + scope: ScopeId("scope-b".to_string()), + event: EventId("evt-start-b".to_string()), + actor_kind: actor_b, + }, + ) + .expect("starting scope b should succeed"); + + capture.push_success(TreeId("tree-b".to_string())); + let outcome = coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Advance { + scope: ScopeId("scope-a".to_string()), + event: EventId("evt-advance-a".to_string()), + actor_kind: actor_a, + }, + ) + .expect("advance should succeed"); + + let event = outcome + .mutation_event + .expect("a real tree change with two live scopes should commit an event"); + assert_eq!(event.attribution, Attribution::AiContended); + + remove_test_db(&db_path); + } + + #[test] + fn contended_scopes_yield_ai_contended_same_and_different_actor() { + assert_contended_attribution( + "ac5-same-actor", + ActorKind::ClaudeCode, + ActorKind::ClaudeCode, + ); + assert_contended_attribution( + "ac5-different-actor", + ActorKind::ClaudeCode, + ActorKind::Codex, + ); + } + + #[test] + fn cas_conflict_reloads_and_recomputes_without_a_second_snapshot() { + const WRITERS: usize = 3; + + let db_path = unique_test_db_path("ac8-cas-conflict"); + let worktree = WorktreeId("wt-1".to_string()); + + { + let db = RepositoryAgentTraceDb::new_at(&db_path).expect("test db should open"); + let bootstrap_capture = FakeSnapshotCapture::new(TreeId("tree-a".to_string())); + coordinate_boundary(&db, &bootstrap_capture, &worktree, &RuntimeBoundary::Flush) + .expect("baseline flush should succeed"); + } + + let barrier = std::sync::Arc::new(std::sync::Barrier::new(WRITERS)); + let mut handles = Vec::new(); + for i in 0..WRITERS { + let db_path = db_path.clone(); + let worktree = worktree.clone(); + let barrier = std::sync::Arc::clone(&barrier); + handles.push(std::thread::spawn(move || { + let db = RepositoryAgentTraceDb::open_for_hooks_without_migrations_at(&db_path) + .expect("writer handle should open"); + let capture = FakeSnapshotCapture::new(TreeId(format!("tree-writer-{i}"))); + barrier.wait(); + let outcome = + coordinate_boundary(&db, &capture, &worktree, &RuntimeBoundary::Flush).expect( + "each racing writer should eventually succeed after reload+recompute", + ); + ( + outcome.revision, + capture.capture_call_count(), + capture.pin_call_count(), + ) + })); + } + + let mut revisions = Vec::new(); + for handle in handles { + let (revision, captures, pins) = handle.join().expect("writer thread should not panic"); + assert_eq!( + captures, 1, + "each invocation must capture its Git snapshot exactly once, even across CAS retries" + ); + assert_eq!( + pins, 1, + "each invocation must pin exactly once, even across CAS retries" + ); + revisions.push(revision); + } + + revisions.sort_unstable(); + revisions.dedup(); + assert_eq!( + revisions.len(), + WRITERS, + "each racing writer's commit must land at a distinct revision, proving reload+recompute rather than a lost update" + ); + + remove_test_db(&db_path); + } + + #[test] + fn recovers_from_needs_rebaseline_preserving_live_scopes() { + let (db, db_path) = test_db("ac10-needs-rebaseline"); + let worktree = WorktreeId("wt-1".to_string()); + let live_scope = ScopeId("scope-live".to_string()); + let abandoned_scope = ScopeId("scope-abandoned".to_string()); + let capture = FakeSnapshotCapture::new(TreeId("tree-a".to_string())); + + coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Start { + scope: live_scope.clone(), + event: EventId("evt-start-live".to_string()), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect("starting the live scope should succeed"); + coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Start { + scope: abandoned_scope.clone(), + event: EventId("evt-start-abandoned".to_string()), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect("starting the to-be-abandoned scope should succeed"); + + let store = MutationTraceStore::new(&db); + let state = store + .load_worktree(&worktree, Some(&abandoned_scope), None) + .expect("load should succeed") + .expect("worktree should exist") + .into_protocol_state(); + let abandoned_state = protocol::abandon(&state, &abandoned_scope); + let transition = DurableTransition::between(&state, &abandoned_state, &worktree) + .expect("diff should succeed") + .expect("abandon should durably change state"); + assert_eq!( + store + .commit(&transition) + .expect("abandon commit should not error"), + CasResult::Applied + ); + + let before_recovery = store + .load_worktree(&worktree, None, None) + .expect("load should succeed") + .expect("worktree should exist"); + assert!(before_recovery.worktree_state.needs_rebaseline); + + capture.push_success(TreeId("tree-b".to_string())); + let outcome = coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Advance { + scope: live_scope.clone(), + event: EventId("evt-advance-live".to_string()), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect("advance should trigger needs_rebaseline recovery first, then succeed"); + assert!( + outcome.mutation_event.is_none(), + "no evidence should be emitted for the just-discarded interval recovery rebaselined away" + ); + + let after = store + .load_worktree(&worktree, Some(&live_scope), None) + .expect("load should succeed") + .expect("worktree should exist"); + assert!( + !after.worktree_state.needs_rebaseline, + "recovery should clear needs_rebaseline" + ); + assert_eq!( + after.worktree_state.cursor_tree, + TreeId("tree-b".to_string()), + "recovery should rebaseline the cursor to the recovering invocation's own observed tree" + ); + assert_eq!( + after.scopes.get(&live_scope).map(|s| s.status), + Some(ScopeStatus::Active), + "needs_rebaseline recovery must preserve live scopes" + ); + + remove_test_db(&db_path); + } + + #[test] + fn recovers_from_snapshot_failure_taint_abandoning_live_scopes() { + let (db, db_path) = test_db("ac10-taint-recovery"); + let worktree = WorktreeId("wt-1".to_string()); + let live_scope = ScopeId("scope-live".to_string()); + let capture = FakeSnapshotCapture::new(TreeId("tree-a".to_string())); + + coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Start { + scope: live_scope.clone(), + event: EventId("evt-start".to_string()), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect("start should succeed"); + + let store = MutationTraceStore::new(&db); + let state = store + .load_worktree(&worktree, None, None) + .expect("load should succeed") + .expect("worktree should exist") + .into_protocol_state(); + let tainted_state = protocol::taint(&state, &worktree); + let transition = DurableTransition::between(&state, &tainted_state, &worktree) + .expect("diff should succeed") + .expect("taint should durably change state"); + assert_eq!( + store + .commit(&transition) + .expect("taint commit should not error"), + CasResult::Applied + ); + + let before_recovery = store + .load_worktree(&worktree, None, None) + .expect("load should succeed") + .expect("worktree should exist"); + assert!(before_recovery.worktree_state.tainted); + + capture.push_success(TreeId("tree-b".to_string())); + coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Advance { + scope: live_scope.clone(), + event: EventId("evt-advance".to_string()), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect("advance should trigger taint recovery first, then succeed"); + + let after = store + .load_worktree(&worktree, Some(&live_scope), None) + .expect("load should succeed") + .expect("worktree should exist"); + assert!( + !after.worktree_state.tainted, + "recovery should clear the taint" + ); + assert_eq!( + after.scopes.get(&live_scope).map(|s| s.status), + Some(ScopeStatus::Abandoned), + "taint recovery must abandon live scopes, unlike needs_rebaseline recovery" + ); + + remove_test_db(&db_path); + } + + fn assert_recovery_at_revision_exhaustion_is_rejected( + label: &str, + tainted: bool, + failure_kind: &str, + needs_rebaseline: bool, + ) { + let (db, db_path) = test_db(label); + let worktree = WorktreeId("wt-1".to_string()); + let scope = ScopeId("scope-1".to_string()); + let event = EventId("evt-start".to_string()); + + insert_worktree_at_revision( + &db, + &worktree.0, + "tree-a", + u64::MAX, + tainted, + failure_kind, + needs_rebaseline, + ); + + let capture = FakeSnapshotCapture::new(TreeId("tree-b".to_string())); + let error = coordinate_boundary( + &db, + &capture, + &worktree, + &RuntimeBoundary::Start { + scope: scope.clone(), + event: event.clone(), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect_err("recovery that cannot advance revision must reject the triggering boundary"); + + match error { + CoordinateError::RevisionExhausted { + worktree_id, + revision, + } => { + assert_eq!(worktree_id, worktree); + assert_eq!(revision, u64::MAX); + } + other => panic!("expected RevisionExhausted, got {other:?}"), + } + + let store = MutationTraceStore::new(&db); + let projection = store + .load_worktree(&worktree, Some(&scope), None) + .expect("load should succeed") + .expect("worktree should exist"); + assert_eq!(projection.worktree_state.revision, u64::MAX); + assert_eq!(projection.worktree_state.tainted, tainted); + assert_eq!(projection.worktree_state.needs_rebaseline, needs_rebaseline); + assert_eq!( + projection.worktree_state.cursor_tree, + TreeId("tree-a".to_string()), + "the cursor must never move when the triggering boundary was never evaluated" + ); + assert_eq!( + projection.scopes.get(&scope).map(|s| s.status), + Some(ScopeStatus::NeverSeen), + "the triggering boundary must never have transitioned the scope" + ); + assert!( + projection.processed_events.is_empty(), + "the triggering boundary's event must never have been recorded as processed" + ); + + remove_test_db(&db_path); + } + + #[test] + fn mandatory_recovery_that_cannot_advance_revision_rejects_the_triggering_boundary() { + assert_recovery_at_revision_exhaustion_is_rejected( + "revision-exhausted-tainted", + true, + "snapshot_failure", + false, + ); + assert_recovery_at_revision_exhaustion_is_rejected( + "revision-exhausted-needs-rebaseline", + false, + "healthy", + true, + ); + } + + #[test] + fn snapshot_failure_taints_an_existing_worktree() { + let (db, db_path) = test_db("ac11-taints-existing"); + let worktree = WorktreeId("wt-1".to_string()); + let bootstrap = FakeSnapshotCapture::new(TreeId("tree-a".to_string())); + coordinate_boundary(&db, &bootstrap, &worktree, &RuntimeBoundary::Flush) + .expect("baseline flush should materialize the worktree"); + + let failing = FakeSnapshotCapture::new(TreeId("tree-b".to_string())); + failing.push_failure("simulated git snapshot failure"); + + let error = coordinate_boundary(&db, &failing, &worktree, &RuntimeBoundary::Flush) + .expect_err("a capture failure against an existing worktree should be reported"); + match error { + CoordinateError::SnapshotFailure { + persisted_taint, .. + } => assert!(persisted_taint), + other => panic!("expected SnapshotFailure, got {other:?}"), + } + + let store = MutationTraceStore::new(&db); + let projection = store + .load_worktree(&worktree, None, None) + .expect("load should succeed") + .expect("worktree should exist"); + assert!(projection.worktree_state.tainted); + assert_eq!( + projection.worktree_state.failure_kind, + FailureKind::SnapshotFailure + ); + + remove_test_db(&db_path); + } + + #[test] + fn snapshot_failure_taint_survives_a_losing_cas_and_commits_on_retry() { + let (db, db_path) = test_db("ac11-taint-retry-succeeds"); + let worktree = WorktreeId("wt-1".to_string()); + let bootstrap = FakeSnapshotCapture::new(TreeId("tree-a".to_string())); + coordinate_boundary(&db, &bootstrap, &worktree, &RuntimeBoundary::Flush) + .expect("baseline flush should materialize the worktree"); + + let store = MutationTraceStore::new(&db); + let competing_scope = ScopeId("competing-scope".to_string()); + let mut interfered = false; + let persisted_taint = run_taint_retry_loop_inner(&store, &worktree, |attempt| { + if attempt == 0 && !interfered { + interfered = true; + commit_competing_advance( + &store, + &worktree, + &competing_scope, + "competing-event-1", + true, + ); + } + }) + .expect("taint retry loop should not error"); + + assert!( + persisted_taint, + "the taint should eventually be persisted after reloading past the losing CAS" + ); + + let projection = store + .load_worktree(&worktree, None, None) + .expect("load should succeed") + .expect("worktree should exist"); + assert!(projection.worktree_state.tainted); + assert_eq!( + projection.worktree_state.failure_kind, + FailureKind::SnapshotFailure + ); + + remove_test_db(&db_path); + } + + #[test] + fn snapshot_failure_taint_reports_not_persisted_after_retries_are_exhausted() { + let (db, db_path) = test_db("ac11-taint-exhaustion"); + let worktree = WorktreeId("wt-1".to_string()); + let bootstrap = FakeSnapshotCapture::new(TreeId("tree-a".to_string())); + coordinate_boundary(&db, &bootstrap, &worktree, &RuntimeBoundary::Flush) + .expect("baseline flush should materialize the worktree"); + + let store = MutationTraceStore::new(&db); + let competing_scope = ScopeId("competing-scope".to_string()); + let mut competing_calls = 0u32; + let persisted_taint = run_taint_retry_loop_inner(&store, &worktree, |_attempt| { + competing_calls += 1; + commit_competing_advance( + &store, + &worktree, + &competing_scope, + &format!("competing-event-{competing_calls}"), + competing_calls == 1, + ); + }) + .expect("taint retry loop should not error even when every attempt conflicts"); + + assert!( + !persisted_taint, + "taint must not be reported as persisted once every bounded attempt has conflicted" + ); + + let projection = store + .load_worktree(&worktree, None, None) + .expect("load should succeed") + .expect("worktree should exist"); + assert!( + !projection.worktree_state.tainted, + "the worktree must remain untainted when the taint itself was never actually committed" + ); + + remove_test_db(&db_path); + } + + #[test] + fn snapshot_failure_before_any_baseline_makes_no_durable_write() { + let (db, db_path) = test_db("ac11-bootstrap-failure"); + let worktree = WorktreeId("wt-1".to_string()); + let failing = FakeSnapshotCapture::new(TreeId("unused".to_string())); + failing.push_failure("simulated git snapshot failure before any baseline exists"); + + let error = coordinate_boundary(&db, &failing, &worktree, &RuntimeBoundary::Flush) + .expect_err("a capture failure with no prior worktree row should still be reported"); + match error { + CoordinateError::SnapshotFailure { + persisted_taint, .. + } => assert!(!persisted_taint), + other => panic!("expected SnapshotFailure, got {other:?}"), + } + + let store = MutationTraceStore::new(&db); + let projection = store + .load_worktree(&worktree, None, None) + .expect("load should succeed"); + assert!( + projection.is_none(), + "no durable worktree row should ever be created for a bootstrap-time snapshot failure" + ); + + remove_test_db(&db_path); + } + + #[test] + fn snapshot_failure_taints_a_worktree_materialized_concurrently_during_capture() { + let (db, db_path) = test_db("ac11-concurrent-materialization"); + let worktree = WorktreeId("wt-1".to_string()); + + let capture = HookedFailingCapture::new( + "simulated git snapshot failure racing a concurrent materialization", + || { + let store = MutationTraceStore::new(&db); + store + .initialize_worktree( + &worktree, + &TreeId("tree-materialized-concurrently".to_string()), + ) + .expect("concurrent materialization should succeed"); + }, + ); + + let error = coordinate_boundary(&db, &capture, &worktree, &RuntimeBoundary::Flush) + .expect_err( + "a capture failure racing a concurrent materialization should still be reported", + ); + match error { + CoordinateError::SnapshotFailure { persisted_taint, .. } => assert!( + persisted_taint, + "the failure handler's fresh, post-failure read must still find and taint the concurrently materialized worktree" + ), + other => panic!("expected SnapshotFailure, got {other:?}"), + } + + let store = MutationTraceStore::new(&db); + let projection = store + .load_worktree(&worktree, None, None) + .expect("load should succeed") + .expect("worktree should exist"); + assert!(projection.worktree_state.tainted); + + remove_test_db(&db_path); + } +} diff --git a/cli/src/services/mutation_trace/runtime/mod.rs b/cli/src/services/mutation_trace/runtime/mod.rs index 57927a3f..a1b56ba5 100644 --- a/cli/src/services/mutation_trace/runtime/mod.rs +++ b/cli/src/services/mutation_trace/runtime/mod.rs @@ -1,2 +1,3 @@ +mod coordinator; mod git_snapshot; mod worktree_lock; diff --git a/context/cli/mutation-trace-protocol.md b/context/cli/mutation-trace-protocol.md index 3fc9bf7e..fa4c3f69 100644 --- a/context/cli/mutation-trace-protocol.md +++ b/context/cli/mutation-trace-protocol.md @@ -211,14 +211,15 @@ loads a `ProtocolState`; `protocol.rs` only transitions already-known ones. ## Target end-state architecture -The plan's file split anticipated three seams beyond `protocol.rs`. `store.rs` -and `runtime/git_snapshot.rs` now exist as real call sites (see +The plan's file split anticipated three seams beyond `protocol.rs`. `store.rs`, +`runtime/git_snapshot.rs`, and `coordinator.rs`'s internal pipeline now exist +as real call sites (see [`mutation-trace-runtime-coordinator.md`](mutation-trace-runtime-coordinator.md)); -`coordinator.rs` remains future work: +only `coordinator.rs`'s public `coordinate()` remains future work: ```mermaid flowchart LR - coordinator["coordinator.rs (future work)\n(imperative shell:\nDB load, Git snapshot,\nCAS/retry, persist)"] + coordinator["coordinator.rs (internal pipeline implemented;\npublic coordinate() future work)\n(imperative shell:\nDB load, Git snapshot,\nCAS/retry, persist)"] protocol["protocol.rs\n(pure transitions —\nprepare/commit/attribution/\ntaint/abandon/recover\nall implemented)"] git_snapshot["runtime/git_snapshot.rs (implemented)\n(isolated Git snapshot,\ntemporary index, tree capture/diff,\nSCE-owned ref pinning)"] store["store.rs\n(cursor/revision, scopes,\nprocessed events, mutation\nevidence, CAS transaction)"] @@ -228,19 +229,18 @@ flowchart LR coordinator --> store ``` -- **`coordinator.rs`** (future work) — receives hook/session identity, resolves - scope actor/worktree identity, loads a `ProtocolState` via `store.rs`, and - calls the pure protocol. +- **`coordinator.rs`** (internal pipeline implemented) — resolves scope + actor/worktree identity, loads a `ProtocolState` via `store.rs`, and calls + the pure protocol; its public `coordinate()` entrypoint remains future work. - **`runtime/git_snapshot.rs`** (implemented) — captures an isolated worktree - snapshot and pins it durably; not yet called by anything. + snapshot and pins it durably; called by `coordinator.rs`. - **`store.rs`** (implemented) — loads and persists worktree/scope/event state - via a CAS-guarded commit; never remaps `actor_kind`/`worktree_id` for an - existing `ScopeId` (see "Runtime scope materialization" above). + via a CAS-guarded commit; never remaps an existing `ScopeId`'s + `actor_kind`/`worktree_id` (see "Runtime scope materialization" above). - **`protocol.rs`** — assumes referenced scopes already exist in `ProtocolState.scopes`; validates and transitions lifecycle state only. -`protocol.rs` stays free of any Git object, DB row, or CAS transaction concept; -`coordinator.rs` is not created by this or the store plan. +`protocol.rs` stays free of any Git object, DB row, or CAS transaction concept. ## Authoritative source diff --git a/context/cli/mutation-trace-runtime-coordinator.md b/context/cli/mutation-trace-runtime-coordinator.md index c0de17c9..a2dff5d9 100644 --- a/context/cli/mutation-trace-runtime-coordinator.md +++ b/context/cli/mutation-trace-runtime-coordinator.md @@ -19,8 +19,10 @@ documented convention. ## Current code surface -The per-worktree runtime lock and the isolated Git snapshot service exist so -far; the coordinator that will drive both is future work. +The per-worktree runtime lock, the isolated Git snapshot service, and the +coordinator's internal protocol-integration pipeline exist so far; only the +public, lock-wrapped `coordinate()` entrypoint that drives the lock and +checkout identity around that pipeline is still future work. - `cli/src/services/mutation_trace/runtime/worktree_lock.rs` — `WorktreeLock::acquire(git_dir: &Path, timeout: Duration) -> @@ -68,13 +70,50 @@ far; the coordinator that will drive both is future work. `diff_trees(&self, before, after) -> Result` runs `git diff --binary --full-index --no-ext-diff --no-textconv` between two tree SHAs, returning the raw diff text `patch.rs::parse_patch` already knows how to - parse. Nothing calls this service yet; `coordinator.rs` is its only planned - caller. - -The runtime lock guards the coordinator's own critical section (snapshot + parse. `coordinator.rs` is its only caller, via the `SnapshotCapture` trait + below. +- `cli/src/services/mutation_trace/runtime/coordinator.rs` — the composition + point that drives `protocol.rs`/`store.rs`/`git_snapshot.rs` together. Its + `SnapshotCapture` trait (`capture(&self) -> Result`, `pin(&self, + worktree_id, tree) -> Result<()>`) is the one dependency-injection seam the + pipeline introduces for determinism; `GitSnapshotService` implements it + directly, and the module's own tests use a fake, call-counting + implementation instead of real concurrent Git processes. + `RuntimeBoundary` is a hook/flush boundary in already-canonical runtime + identities (`Start`/`Advance`/`Close` carry `{ scope, event, actor_kind }`; + `Flush` carries nothing — its worktree is always the invocation's own + already-resolved one, never caller-supplied) and documents the + `(ScopeId, EventId)` replay-identity contract a future harness adapter must + uphold. The internal, generic-over-`SnapshotCapture` pipeline — not yet + reachable outside `runtime`, since the public `coordinate()` entrypoint is + still future work (see Status) — does, per invocation: capture and pin + exactly one Git snapshot; on failure, run a bounded taint-retry loop instead + (below) and return without touching the rest of the pipeline; on success, + idempotently materialize the worktree row and, for hook boundaries, the + scope row; then loop (bounded, `MAX_CAS_RETRY_ATTEMPTS = 5`, no backoff): + load durable state fresh, recover first if the worktree is tainted or needs + rebaseline (its own CAS commit, reusing the one captured tree as the + rebaseline target), then `prepare`/`commit` the triggering boundary against + that state (a second CAS commit) — reloading and recomputing from scratch + on `Conflict`, without ever re-capturing or re-pinning. A settled no-op + result (a stale, rejected, or replayed attempt) is a successful return, not + an error. + + A capture or pin failure is handled by its own bounded taint-retry loop: a + fresh `load_worktree` on every iteration, always evaluated after the + failure, never before it — so a worktree another caller materializes + concurrently while this invocation's own capture is still in flight is + still found and correctly tainted. No durable worktree row on that fresh + read means no taint to record (`persisted_taint: false`, no write); an + already-tainted no-op reads back the current flag instead of assuming + success; otherwise the loop commits the taint transition and retries on + `Conflict`, reporting `persisted_taint: false` only once every bounded + attempt has been exhausted. + +The runtime lock will guard the coordinator's own critical section (snapshot capture, worktree/scope materialization, recovery, and the CAS retry loop) -once the coordinator exists. It is held on every future `coordinate()` call, -unlike the checkout-identity-creation lock. +once the public `coordinate()` entrypoint wires it in. It will be held on +every `coordinate()` call, unlike the checkout-identity-creation lock. ## Two distinct locks, two distinct invariants @@ -126,13 +165,32 @@ after the temp index file is gone, `git gc --prune=now`/`git prune unpinned tree in the same repository, `pin_tree` idempotency, and `diff_trees` output shape. +`coordinator.rs`'s inline `#[cfg(test)] mod tests` exercises the internal +pipeline against a real temp-file `RepositoryAgentTraceDb`, using a fake, +call-counting `SnapshotCapture` (or, for CAS-conflict scenarios, real OS +threads racing separate DB handles against one on-disk database): first +observation establishes a baseline with no evidence; an edit observed between +`Start` and `Advance` commits exactly one `AiExclusive` event; replaying an +identical `(scope, event)` boundary is a no-op, not a duplicate; `Close` +attributes to the scope it is about to close; two live scopes yield +`AiContended` regardless of matching or differing `ActorKind`; a CAS conflict +reloads and recomputes without a second capture or pin; `needs_rebaseline` +recovery preserves live scopes while taint recovery abandons them; and the +taint-retry loop taints an existing worktree, survives a losing CAS before +committing on retry, reports `persisted_taint: false` once exhausted, makes +no write when no worktree row exists yet, and still finds and taints a +worktree another caller materializes concurrently during this invocation's +own failing capture. + ## Status -The per-worktree runtime lock and the isolated Git snapshot service (above) -are implemented. The coordinator's protocol-integration pipeline, the public -lock-wrapped `coordinate()` entrypoint, and cross-module integration tests -remain future work tracked by the `mutation-cursor-runtime-coordinator` -plan's task stack. This file will grow to cover them as each lands. +The per-worktree runtime lock, the isolated Git snapshot service, and the +coordinator's internal protocol-integration pipeline (above) are implemented. +The public, lock-wrapped `coordinate()` entrypoint (resolving `git_dir`, +acquiring `WorktreeLock`, resolving checkout identity, deriving `WorktreeId`) +and cross-module integration tests remain future work tracked by the +`mutation-cursor-runtime-coordinator` plan's task stack. This file will grow +to cover them as each lands. See also: [`mutation-trace-protocol.md`](mutation-trace-protocol.md), [`mutation-trace-store.md`](mutation-trace-store.md), diff --git a/context/cli/mutation-trace-store.md b/context/cli/mutation-trace-store.md index a66cdf3e..0d58ef9b 100644 --- a/context/cli/mutation-trace-store.md +++ b/context/cli/mutation-trace-store.md @@ -118,15 +118,18 @@ returns `Err`. ## Non-goals - No Git or filesystem I/O — `store.rs` itself calls neither Git nor the - filesystem. `runtime/git_snapshot.rs` now exists as a sibling module (see - [`mutation-trace-runtime-coordinator.md`](mutation-trace-runtime-coordinator.md)) - but is not wired to this store yet; `coordinator.rs`, the future caller of - both, remains future work. + filesystem. `runtime/git_snapshot.rs` and `runtime/coordinator.rs`'s + internal pipeline now exist and wire this store together with the Git + snapshot service (see + [`mutation-trace-runtime-coordinator.md`](mutation-trace-runtime-coordinator.md)); + `store.rs` itself remains unmodified and still performs no Git or + filesystem I/O of its own. - No attribution or boundary-kind decisions — `DurableTransition::between` and `store.rs` are both structurally blind to protocol meaning. - No retry-after-`Conflict` loop — `CasResult::Conflict` is returned to the - caller; retrying with a freshly reloaded revision is a future adapter's - responsibility, not this module's. + caller; retrying with a freshly reloaded revision is the calling adapter's + responsibility, not this module's. `runtime::coordinator`'s bounded + CAS-retry loop is now that adapter. - No deletion of terminal (`Closed`/`Abandoned`) scope rows or historical `mutation_trace_events` rows — scope garbage collection is out of scope for this plan. diff --git a/context/context-map.md b/context/context-map.md index 1afc46f4..f0f8353e 100644 --- a/context/context-map.md +++ b/context/context-map.md @@ -23,11 +23,11 @@ Feature/domain context: - `context/cli/config-precedence-contract.md` (implemented `sce config` show/validate command contract, deterministic `flags > env > config file > defaults` resolution order, focused `config/resolver.rs` ownership for config discovery/merge/runtime precedence plus default-discovered invalid-file degradation, focused `config/render.rs` ownership for `show`/`validate` text+JSON output construction, canonical `$schema` acceptance for startup-loaded `sce/config.json` files, shared auth-key env/config/optional baked-default support starting with `workos_client_id`, shared runtime resolution for flat logging observability keys including config-file/default `log_to_file`, `log_dir` / `SCE_LOG_DIR` with `/sce/logs` defaulting plus config-file/default-only positive `log_file_retention_limit`, config-file-only `agent_trace.repository_id`/`agent_trace.repository_remote` repository-identity keys with default remote `origin`, default-enabled `agent_trace.auto_sync` boolean resolution with explicit-false opt-out for the post-commit trigger boundary, the catalog-derived `integrations.optional_workflows` optional-workflow selection key, JSON-pointer-prefixed schema-validation errors, canonical Pkl-generated `sce/config.json` schema ownership plus CLI embedding/reuse contract including `policies.attribution_hooks.enabled` default-true/explicit-false opt-out metadata, config-file selection order, `show` provenance output, and trimmed `validate` output contract) - `context/cli/capability-traits.md` (current broad CLI capability seam in `cli/src/services/capabilities.rs`, including `FsOps`/`StdFsOps`, `GitOps`/`ProcessGitOps`, git root/hooks resolution behavior, compile-time-typed borrowed AppContext wiring with associated-type narrow capability accessors plus `ContextWithRepoRoot` repo-root-scoped context derivation, generic command execution bounds, and test-only unimplemented stubs; current service internals do not consume fs/git traits until later lifecycle migration tasks) - `context/cli/service-lifecycle.md` (current compile-safe lifecycle seam in `cli/src/services/lifecycle.rs`, including default no-op `ServiceLifecycle` diagnose/fix/setup methods against narrow `HasRepoRoot`, lifecycle-owned health/fix/setup result types with generic setup messages, doctor/setup adapter boundaries, the static `LifecycleProvider` enum catalog/dispatcher, hook/config/local_db/auth_db/agent_trace_db lifecycle providers including setup-time repository-scoped Agent Trace DB initialization plus checkout identity diagnostics, implemented doctor aggregation over diagnose/fix providers, and implemented setup aggregation over `setup` providers in order config → local_db → auth_db → agent_trace_db → hooks when requested) -- `context/cli/mutation-trace-protocol.md` (pure, dependency-free `cli/src/services/mutation_trace/` refinement of the verified `spec/mutation_cursor.qnt` mutation-cursor protocol: the full pure action set — `prepare`/`commit` transition logic, attribution/mutation-event materialization, snapshot-failure/database-failure taint actions, scope abandonment, and recovery with an explicit observed-tree input — plus cross-action sequence/invariant tests and a module-level Quint refinement matrix are all complete (`types.rs`, `protocol.rs`, `mod.rs`, `tests.rs`, registered with `#[allow(dead_code)]`); opaque-newtype identity refinement decisions vs. the Quint model's bounded enums, and the `coordinator.rs` target end-state seam this layout leaves room for but does not create — `store.rs` and `runtime/git_snapshot.rs` now exist, built out by the `mutation-cursor-store-persistence` and `mutation-cursor-runtime-coordinator` plans respectively; `protocol.rs`'s pure transitions are not yet wired into any hook or command) +- `context/cli/mutation-trace-protocol.md` (pure, dependency-free `cli/src/services/mutation_trace/` refinement of the verified `spec/mutation_cursor.qnt` mutation-cursor protocol: the full pure action set — `prepare`/`commit` transition logic, attribution/mutation-event materialization, snapshot-failure/database-failure taint actions, scope abandonment, and recovery with an explicit observed-tree input — plus cross-action sequence/invariant tests and a module-level Quint refinement matrix are all complete (`types.rs`, `protocol.rs`, `mod.rs`, `tests.rs`, registered with `#[allow(dead_code)]`); opaque-newtype identity refinement decisions vs. the Quint model's bounded enums, and the `coordinator.rs` target end-state seam this layout leaves room for but does not create — `store.rs`, `runtime/git_snapshot.rs`, and `runtime/coordinator.rs`'s internal pipeline now exist, built out by the `mutation-cursor-store-persistence` and `mutation-cursor-runtime-coordinator` plans respectively; `protocol.rs`'s pure transitions are not yet wired into any hook or command) - `context/cli/mutation-trace-revision-refinement.md` (the Quint `revision: int` → Rust `WorktreeState::revision: u64` bounded-integer refinement: the private `next_revision` checked-arithmetic helper `commit`/`taint`/`abandon`/`recover` all route through instead of a raw `+ 1`, so a worktree at `revision: u64::MAX` is a guarded no-op/rejection rather than a silent wrap to `0`) - `context/cli/mutation-trace-quint-connect.md` (`#[cfg(test)]`-only Quint Connect model-based-testing harness in `cli/src/services/mutation_trace/mbt/` continuously checking `protocol.rs` against `spec/mutation_cursor.qnt`: the verification-only `mbtAction`/`MbtAction` record-payload transport excluded from comparison, the operation-identity-vs-`MbtStutter` distinction on guarded/no-op branches with its two deterministic regressions, finite ID mapping, the AC5 comparable-state field list, `randomPrepare` staying a single `step` branch, deterministic/generated (500×30, seed-reproducible) test coverage, and the two Nix checks — generic `checks.cli-tests` and dedicated `checks.mutation-trace-quint-connect` — that both require the pinned Quint binary plus the top-level `spec/` directory in `workspaceSrc`'s Nix fileset) - `context/cli/mutation-trace-store.md` (durable persistence for the mutation-cursor protocol in `cli/src/services/mutation_trace/store.rs`, built by the `mutation-cursor-store-persistence` plan: the one-directional `protocol.rs` -> `DurableTransition::between` (pure structural diff) -> `store.rs` (SQL translation) -> `RepositoryAgentTraceDb` boundary; migration `003_mutation_trace_protocol.sql`'s five tables; `AttemptState`/`external_taint` non-persistence; the 8-byte big-endian `BLOB` revision encoding plus explicit non-`Debug` enum codecs; the bounded hot-path `load_worktree` read vs. the cold-path `load_mutation_event`; `commit`'s single-`BEGIN IMMEDIATE` CAS batch via `TursoDb::execute_transactional_cas_batch`, distinguishing `Conflict`/retryable-transient/deterministic-`Err` outcomes; and the store's non-goals — no Git/filesystem I/O, no attribution/boundary-kind decisions, no retry-after-`Conflict` loop, no terminal-scope garbage collection) -- `context/cli/mutation-trace-runtime-coordinator.md` (imperative-shell runtime layer in `cli/src/services/mutation_trace/runtime/`, built by the `mutation-cursor-runtime-coordinator` plan: the per-worktree OS advisory lock, `runtime::worktree_lock::WorktreeLock::acquire(git_dir, timeout)` at `/sce/mutation-cursor.lock`, bounded `try_lock()` polling with a distinct matchable timeout error and RAII release, and its distinction from the separate checkout-identity-creation lock; and the isolated Git snapshot service, `runtime::git_snapshot::GitSnapshotService` (`capture_tree`/`pin_tree`/`diff_trees`), writing tree/blob objects into the repository's normal object database and pinning durable ones via a create-only, idempotent `refs/sce/mutation-cursor//` ref rather than a private object store; the coordinator pipeline, public `coordinate()` entrypoint, and cross-module integration tests remain future work) +- `context/cli/mutation-trace-runtime-coordinator.md` (imperative-shell runtime layer in `cli/src/services/mutation_trace/runtime/`, built by the `mutation-cursor-runtime-coordinator` plan: the per-worktree OS advisory lock, `runtime::worktree_lock::WorktreeLock::acquire(git_dir, timeout)` at `/sce/mutation-cursor.lock`, bounded `try_lock()` polling with a distinct matchable timeout error and RAII release, and its distinction from the separate checkout-identity-creation lock; and the isolated Git snapshot service, `runtime::git_snapshot::GitSnapshotService` (`capture_tree`/`pin_tree`/`diff_trees`), writing tree/blob objects into the repository's normal object database and pinning durable ones via a create-only, idempotent `refs/sce/mutation-cursor//` ref rather than a private object store; and `runtime::coordinator`'s internal, generic-over-`SnapshotCapture` protocol-integration pipeline (`RuntimeBoundary`/`CoordinateOutcome`/`CoordinateError`, the load → recover-if-needed → prepare/commit CAS-retry loop, and its own bounded snapshot-failure taint-retry loop); the public `coordinate()` entrypoint and cross-module integration tests remain future work) - `context/sce/cli-exit-code-contract.md` (stable class-based `sce` process exit-code contract in `cli/src/app.rs`, so automation can branch on failure category without parsing error text) - `context/sce/cli-error-code-taxonomy.md` (stable user-facing `SCE-ERR-*` diagnostic code classes rendered by `cli/src/app.rs`, complementing the numeric exit-code classes) - `context/sce/cli-stdout-stderr-contract.md` (implemented stream contract in `cli/src/app.rs`: command payloads on stdout only, redacted diagnostics on stderr) diff --git a/context/plans/mutation-cursor-runtime-coordinator.md b/context/plans/mutation-cursor-runtime-coordinator.md index cb929824..c4d37548 100644 --- a/context/plans/mutation-cursor-runtime-coordinator.md +++ b/context/plans/mutation-cursor-runtime-coordinator.md @@ -681,7 +681,7 @@ Persist this field in every plan; this is durable plan state, not chat state: corrections above capture. - Context synchronization: synced -- [ ] T04: `Add the coordinator's core protocol-integration pipeline` (status:todo) +- [x] T04: `Add the coordinator's core protocol-integration pipeline` (status:done) - Task ID: T04 - Scope: In — `runtime/coordinator.rs`: `RuntimeBoundary` (with its `(ScopeId, EventId)` replay-identity contract documented on the type), @@ -711,7 +711,210 @@ Persist this field in every plan; this is durable plan state, not chat state: a worktree row materialized *during* a failing capture attempt (not before it) is still found and tainted by that same failing invocation. - Verify: `cargo test -p shared-context-engineering runtime::coordinator::` (via `./scripts/run-cli-cargo.sh test`) - - Context synchronization: pending + - Completed: 2026-08-30 + - Files changed: `cli/src/services/mutation_trace/runtime/coordinator.rs` (new), + `cli/src/services/mutation_trace/runtime/mod.rs` (added `mod coordinator;`) + - Result: Added `runtime/coordinator.rs` with `RuntimeBoundary` (`Start`/ + `Advance`/`Close` carrying `{ scope, event, actor_kind }`, `Flush` carrying + nothing — its `WorktreeId` is always the invocation's already-resolved + worktree, never caller-supplied — with the `(ScopeId, EventId)` + replay-identity contract from `spec/mutation_cursor.md`'s "Event identity" + section documented on the type's doc comment), `CoordinateOutcome` + (`worktree_id`/`observed_tree`/`revision`/`evaluation`/`mutation_event`, + exactly the plan's sketch), `CoordinateError` (`SnapshotFailure`/ + `ScopeIdentityConflict`/`CasConflictExhausted`/`LockAcquisition`/`Other`, + with `Display`/`std::error::Error` impls; `LockAcquisition` is reserved for + T05 and never constructed here — `#[allow(dead_code)]` on + `pub mod mutation_trace;` in `services/mod.rs` already covers it), and the + `SnapshotCapture` trait (`capture(&self) -> Result`, + `pin(&self, worktree_id: &WorktreeId, tree: &TreeId) -> Result<()>` — + taking no `repository_root` parameter, since `GitSnapshotService` + (T03) already binds it at construction; a reviewed deviation from the + plan's earlier trait sketch, not from `GitSnapshotService`'s actual + signature), implemented for `GitSnapshotService` by direct delegation to + `capture_tree`/`pin_tree`. + + The private `coordinate_boundary(db, capture, + worktree_id, boundary: &RuntimeBoundary)` is the internal, + generic-over-`SnapshotCapture` entrypoint the plan calls for: it captures + and pins the one Git snapshot for the invocation; on failure, delegates to + `handle_snapshot_failure`/`run_taint_retry_loop_inner` and returns + `SnapshotFailure` without ever reaching the pipeline below; on success, + calls `initialize_worktree` (idempotent) and, for hook boundaries only, + `register_scope` (mapping `ScopeIdentityConflict` errors), then runs the + bounded (`MAX_CAS_RETRY_ATTEMPTS = 5`, no backoff) `load_worktree` → + recover-if-needed (its own `DurableTransition`/CAS commit, retried via + `continue` on `Conflict`) → `prepare`/`commit` the triggering boundary (a + second `DurableTransition`/CAS commit) loop, reusing the one captured + `observed_tree` and the same `AttemptId` generation + (`Uuid::new_v4()`) every iteration; a `None` `DurableTransition` (a + stale/rejected/replayed attempt) is a successful no-op return, not an + error. `run_taint_retry_loop_inner` implements the plan's exact + snapshot-failure pseudocode: a fresh `load_worktree(worktree, None, None)` + on every iteration including the first, always evaluated after the + triggering failure; `None` → `persisted_taint: false` (bootstrap, no + durable write); an already-tainted no-op `taint()` transition → reads back + the current `tainted` flag rather than assuming success; otherwise commits + the taint transition and retries on `Conflict`; exhaustion after 5 + attempts → `persisted_taint: false`. `run_taint_retry_loop` (production, + used by `handle_snapshot_failure`) is a thin wrapper over + `run_taint_retry_loop_inner` with a no-op `after_load` hook; the hook + itself (`after_load: impl FnMut(u32)`, fired after each iteration's load + and before its own commit) is a deterministic-testing seam mirroring + `worktree_lock.rs`'s own `acquire_inner`/`on_contention` pattern (T02), + used only by this task's own tests to inject a competing commit at a + precise point without needing real thread synchronization. + + Added a `#[cfg(test)] mod tests` with 13 tests (11 exactly matching the + plan's own AC1–AC5/AC8/AC10/AC11 "Validate" function names, plus the + `assert_contended_attribution` helper counted under AC5's single test + function, plus internal coverage): a `FakeSnapshotCapture` (scripted + `Succeed`/`Fail` outcome queue, falling back to a fixed default tree, with + `capture_call_count`/`pin_call_count`) proves AC1 (Flush at first + observation stays at revision 0, no evidence, exactly one capture/pin), + AC2 (Start-then-Advance commits one `AiExclusive` event with the correct + before/after trees), AC3 (replaying the identical `(scope, event)` + boundary is a no-op — the same revision, no duplicated event — proven via + two sequential `coordinate_boundary` calls with the same boundary), AC4 + (Close still attributes to the scope it is about to close, then confirms + the scope reaches `Closed`), and AC5 (`AiContended` for two live scopes, + exercised for both same-`ActorKind` and different-`ActorKind` pairs via + one test calling a shared `assert_contended_attribution` helper twice). + AC8 (`cas_conflict_reloads_and_recomputes_without_a_second_snapshot`) uses + 3 real OS threads racing distinct `Flush` boundaries against one shared + on-disk DB (separate `RepositoryAgentTraceDb` handles per thread, via + `open_for_hooks_without_migrations_at`, matching `store.rs`'s own + two-writer-race test convention) — deterministically bounded, since at + most `WRITERS − 1 = 2` retries are ever needed against the + `MAX_CAS_RETRY_ATTEMPTS = 5` bound — and asserts every writer captured + exactly once, pinned exactly once, and landed at a distinct revision. + AC10 has two tests: `recovers_from_needs_rebaseline_preserving_live_scopes` + (directly commits a `protocol::abandon` transition to force + `needs_rebaseline`, then proves a subsequent `Advance` recovers first — + clearing the flag, rebaselining the cursor to the recovering invocation's + own observed tree, emitting no evidence for the discarded interval — while + the untouched live scope stays `Active`) and + `recovers_from_snapshot_failure_taint_abandoning_live_scopes` (same + pattern via a direct `protocol::taint` commit, proving the live scope + instead becomes `Abandoned`). AC11 has five tests: taints an existing + worktree; survives one losing CAS via the `after_load` hook injecting one + competing (revision-only, non-taint) commit on the first iteration, then + succeeds on retry; exhausts all 5 attempts via the hook injecting a + competing commit on every iteration (`persisted_taint: false`, worktree + stays untainted); a bootstrap failure with no prior worktree row makes no + durable write; and a `HookedFailingCapture` (single-shot failing capture + running an injected closure — here, a direct `initialize_worktree` call — + immediately before returning its failure) proves the fresh, post-failure + `load_worktree` still finds and taints a worktree materialized + concurrently during this invocation's own (failing) capture. + + PR #244 review follow-up: corrected a real correctness hole in the + recovery step, and a stale sketch of the `SnapshotCapture` signature this + record itself had already superseded but the plan's own "Dependency + injection for deterministic tests" design-decision text still carried. + + `protocol::recover` is a guarded no-op — among other guards — when the + worktree's revision is already `u64::MAX` and cannot be advanced (see + `protocol.rs`'s own `next_revision` headroom guard). Before this + follow-up, `coordinate_boundary`'s recovery step treated that no-op + identically to "recovery was not needed": `DurableTransition::between` + returning `None` simply fell through to evaluating the triggering + boundary against the un-recovered, still-tainted-or-needing-rebaseline + state — silently processing a boundary the coordinator's own contract + ("if durable state requires recovery, recovery must complete successfully + before the triggering boundary is processed") forbids. A worktree stuck + at `revision: u64::MAX` while tainted or needing rebaseline could + therefore have a hook boundary attributed and committed against stale, + un-recovered state. + + The fix adds `CoordinateError::RevisionExhausted { worktree_id: WorktreeId, + revision: u64 }` (with a `Display` arm) and changes the recovery step's + `if let Some(transition) = ... { ... }` (silently falling through on + `None`) to a `let Some(transition) = ... else { return + Err(RevisionExhausted { .. }) };` — the triggering boundary is now + reached only when recovery's own `DurableTransition` genuinely exists. + The coordinator does not itself re-check `revision == u64::MAX`: given + `needs_recovery` is already true, `protocol::recover`'s own no-op guards + for "already healthy" and "no durable state" cannot be the cause of a + `None` transition, so a `None` here is only reachable via `recover`'s + revision-headroom guard — `protocol::recover` remains the sole authority + on whether recovery is possible, per the brief's explicit instruction not + to duplicate that check. The `revision` field is populated by reading the + already-loaded (pre-recovery) state, not by a separate query. + + Added `mandatory_recovery_that_cannot_advance_revision_rejects_the_triggering_boundary`, + driven by a shared `assert_recovery_at_revision_exhaustion_is_rejected` + helper covering both reachable `needs_recovery` reasons: a worktree row + inserted directly (via raw SQL, bypassing `initialize_worktree`, which + always starts at revision 0) at `revision: u64::MAX` with `tainted: true` + and, separately, at `revision: u64::MAX` with `needs_rebaseline: true`. + Each asserts `coordinate_boundary` (given a `Start` boundary and a + succeeding `SnapshotCapture`) returns `RevisionExhausted { revision: + u64::MAX, .. }`, then reloads the worktree/scope and asserts nothing + about the pre-existing durable state moved: `revision` is still exactly + `u64::MAX` (no wrap), `tainted`/`needs_rebaseline` and `cursor_tree` are + byte-identical to what was inserted, the scope is still `NeverSeen` (the + triggering `Start` never transitioned it), and `processed_events` is + empty (the triggering event was never recorded) — proving the triggering + boundary was never evaluated, not merely that it produced no visible + event. The two existing recovery tests + (`recovers_from_needs_rebaseline_preserving_live_scopes`, + `recovers_from_snapshot_failure_taint_abandoning_live_scopes`) are + unchanged and still pass, confirming ordinary (non-exhausted) recovery + still completes and the triggering boundary still runs against the + recovered state. + + Separately, this record's own "Dependency injection for deterministic + tests" reference and this task's `SnapshotCapture` bullet already + correctly stated `fn capture(&self) -> Result` (no + `repository_root` parameter, since `GitSnapshotService::new` binds it at + construction), but the plan's "Dependency injection for deterministic + tests" Design decisions section still described `fn capture(&self, + repository_root: &Path) -> Result` — a stale sketch this + record's own implementation had already superseded without ever + correcting that other section. That section now matches the implemented + signature and states why `capture` takes no `repository_root` parameter. + + At the user's explicit request, every comment (`//!`/`///`/`//`, + including on the newly added `RevisionExhausted` variant and every + pre-existing item) was removed from `coordinator.rs`, matching T02's + established precedent (`worktree_lock.rs`, `runtime/mod.rs`) for files + "authored entirely during this task and its review follow-up." No + production behavior, public API shape (beyond the new + `RevisionExhausted` variant itself), or test coverage was changed by the + comment removal. + - Verify (actual): `./scripts/run-cli-cargo.sh test --manifest-path + cli/Cargo.toml runtime::coordinator::` → 13/13 passed (pre-follow-up); + 14/14 passed (post-follow-up, including the new + `mandatory_recovery_that_cannot_advance_revision_rejects_the_triggering_boundary` + test). Additionally ran `./scripts/run-cli-cargo.sh clippy + --manifest-path cli/Cargo.toml --all-targets -- -D warnings` → clean both + times (pedantic/warnings denied workspace-wide, per Constraints; the + original implementation fixed `match_same_arms`, `needless_pass_by_value`, + two `needless_continue`, `elidable_lifetime_names`, and + `items_after_statements` findings; the follow-up introduced none), and + `./scripts/run-cli-cargo.sh fmt --manifest-path cli/Cargo.toml -- --check` + → clean both times (after running `cargo fmt` once per pass). Additionally + ran `./scripts/run-cli-cargo.sh test --manifest-path cli/Cargo.toml` (full + suite) → 819/819 passed (pre-follow-up), 820/820 passed (post-follow-up), + no regression. + - Context impact: `domain` — introduces new architectural elements + (`RuntimeBoundary`, `CoordinateOutcome`, `CoordinateError`, + `SnapshotCapture`, and the internal `coordinate_boundary` pipeline) under + the mutation-cursor-runtime-coordinator domain T02/T03 already established + a domain file for (`context/cli/mutation-trace-runtime-coordinator.md`). + Affected areas for context synchronization to verify/update: + `context/cli/mutation-trace-runtime-coordinator.md` (document the new + types and the pipeline's load → recover-if-needed → prepare/commit and + taint-retry design); `context/cli/mutation-trace-protocol.md`'s "Target + end-state architecture" section, which currently states `coordinator.rs` + "remain[s] future work" — now stale, since the file and its internal + pipeline exist (the public, lock-wrapped `coordinate()` entrypoint remains + T05); `context/context-map.md`'s `coordinator.rs` annotation. Reason: this + is the first task to create `runtime/coordinator.rs` and its core + pipeline types, directly falsifying the "future work" framing multiple + existing context files use for it. + - Context synchronization: synced - [ ] T05: `Wire the worktree lock and checkout identity into coordinate()` (status:todo) - Task ID: T05 @@ -1690,10 +1893,12 @@ from the *triggering boundary's* own `protocol::commit` result ### Dependency injection for deterministic tests -`SnapshotCapture` (`fn capture(&self, repository_root: &Path) -> -Result`, `fn pin(&self, worktree_id: &WorktreeId, tree: &TreeId) -> -Result<()>`) is the one seam this plan introduces for determinism: -`GitSnapshotService` implements it for production use, and T04's tests use +`SnapshotCapture` (`fn capture(&self) -> Result`, `fn pin(&self, +worktree_id: &WorktreeId, tree: &TreeId) -> Result<()>`) is the one seam this +plan introduces for determinism. `capture` takes no `repository_root` +parameter: `GitSnapshotService::new(repository_root)` binds that context once +at construction, so `capture()` does not receive it repeatedly on every call. +`GitSnapshotService` implements the trait for production use, and T04's tests use a fake, call-counting implementation to prove "exactly one `capture`, at most one `pin`, per invocation, even across CAS retries" (AC8) without needing real concurrent Git processes. The lock is *not* faked — From a999bc461bebdb598f5ab790431cbb6d24584dcd Mon Sep 17 00:00:00 2001 From: David Abram Date: Sun, 30 Aug 2026 08:47:38 +0200 Subject: [PATCH 6/8] runtime: Implement lock-wrapped mutation trace coordination Expose coordinate() as the runtime boundary for driving mutation-cursor processing against a real worktree while preventing concurrent same-worktree observations from racing. Resolve the Git directory, acquire a bounded WorktreeLock, derive WorktreeId from checkout identity, and delegate to the existing snapshot/protocol/CAS pipeline; add serialization coverage and update the runtime documentation and plan. Plan: mutation-cursor-runtime-coordinator (T05) Co-authored-by: SCE --- .../mutation_trace/runtime/coordinator.rs | 134 +++++++++++++++++ .../mutation_trace/runtime/worktree_lock.rs | 2 +- context/cli/mutation-trace-protocol.md | 14 +- .../cli/mutation-trace-runtime-coordinator.md | 72 ++++++--- context/cli/mutation-trace-store.md | 6 +- context/context-map.md | 4 +- context/overview.md | 2 +- .../mutation-cursor-runtime-coordinator.md | 142 +++++++++++++++++- 8 files changed, 336 insertions(+), 40 deletions(-) diff --git a/cli/src/services/mutation_trace/runtime/coordinator.rs b/cli/src/services/mutation_trace/runtime/coordinator.rs index 671a9ac2..a4d0bb89 100644 --- a/cli/src/services/mutation_trace/runtime/coordinator.rs +++ b/cli/src/services/mutation_trace/runtime/coordinator.rs @@ -1,7 +1,11 @@ +use std::path::Path; +use std::time::Duration; + use anyhow::Result; use uuid::Uuid; use crate::services::agent_trace_db::repository::RepositoryAgentTraceDb; +use crate::services::checkout::{get_or_create_checkout_id, resolve_git_dir}; use crate::services::mutation_trace::protocol; use crate::services::mutation_trace::store::{CasResult, DurableTransition, MutationTraceStore}; use crate::services::mutation_trace::types::{ @@ -9,9 +13,12 @@ use crate::services::mutation_trace::types::{ }; use super::git_snapshot::GitSnapshotService; +use super::worktree_lock::{acquire_inner, WorktreeLockError}; const MAX_CAS_RETRY_ATTEMPTS: u32 = 5; +const WORKTREE_LOCK_TIMEOUT: Duration = Duration::from_secs(10); + #[derive(Clone, Debug)] pub enum RuntimeBoundary { Start { @@ -104,6 +111,40 @@ impl SnapshotCapture for GitSnapshotService { } } +pub fn coordinate( + repository_root: &Path, + db: &RepositoryAgentTraceDb, + boundary: &RuntimeBoundary, +) -> Result { + coordinate_inner(repository_root, db, boundary, || {}) +} + +fn coordinate_inner( + repository_root: &Path, + db: &RepositoryAgentTraceDb, + boundary: &RuntimeBoundary, + on_lock_contention: F, +) -> Result +where + F: FnOnce(), +{ + let git_dir = resolve_git_dir(repository_root).map_err(CoordinateError::Other)?; + + let _lock = acquire_inner(&git_dir, WORKTREE_LOCK_TIMEOUT, on_lock_contention) + .map_err(lock_acquisition)?; + + let checkout_id = get_or_create_checkout_id(&git_dir).map_err(CoordinateError::Other)?; + let worktree_id = WorktreeId(checkout_id); + + let snapshot = GitSnapshotService::new(repository_root).map_err(CoordinateError::Other)?; + + coordinate_boundary(db, &snapshot, &worktree_id, boundary) +} + +fn lock_acquisition(error: WorktreeLockError) -> CoordinateError { + CoordinateError::LockAcquisition(anyhow::Error::new(error)) +} + fn coordinate_boundary( db: &RepositoryAgentTraceDb, capture: &C, @@ -315,7 +356,10 @@ where mod tests { use std::cell::{Cell, RefCell}; use std::collections::VecDeque; + use std::process::Command; use std::sync::atomic::{AtomicU64, Ordering}; + use std::sync::mpsc; + use std::thread; use super::*; use crate::services::mutation_trace::store::encode_revision; @@ -345,6 +389,38 @@ mod tests { } } + fn unique_test_repo(label: &str) -> std::path::PathBuf { + let id = NEXT_TEST_DB_ID.fetch_add(1, Ordering::Relaxed); + std::env::temp_dir().join(format!( + "sce-mutation-trace-coordinator-repo-{label}-{}-{id}", + std::process::id() + )) + } + + fn init_repo(repo_root: &std::path::Path) { + std::fs::create_dir_all(repo_root).expect("repo root should be created"); + run_git(repo_root, &["init", "--quiet"]); + run_git(repo_root, &["config", "user.email", "test@example.com"]); + run_git(repo_root, &["config", "user.name", "Test"]); + } + + fn run_git(repo_root: &std::path::Path, args: &[&str]) { + let output = Command::new("git") + .args(args) + .current_dir(repo_root) + .output() + .expect("git command should spawn"); + assert!( + output.status.success(), + "git {args:?} failed: {}", + String::from_utf8_lossy(&output.stderr) + ); + } + + fn remove_test_repo(repo_root: &std::path::Path) { + let _ = std::fs::remove_dir_all(repo_root); + } + enum FakeOutcome { Succeed(TreeId), Fail(String), @@ -1228,4 +1304,62 @@ mod tests { remove_test_db(&db_path); } + + #[test] + fn two_threads_on_the_same_worktree_serialize() { + let repo_root = unique_test_repo("t05-serialize"); + init_repo(&repo_root); + let db_path = repo_root.join("agent-trace.db"); + + let git_dir = resolve_git_dir(&repo_root).expect("git dir should resolve"); + let held = acquire_inner(&git_dir, Duration::from_secs(5), || {}) + .expect("the test should hold a real WorktreeLock before the worker runs"); + + let (contention_tx, contention_rx) = mpsc::channel(); + let (result_tx, result_rx) = mpsc::channel(); + let repo_root_clone = repo_root.clone(); + let db_path_clone = db_path.clone(); + let worker = thread::spawn(move || { + let db = RepositoryAgentTraceDb::new_at(&db_path_clone).expect("worker db should open"); + let outcome = + coordinate_inner(&repo_root_clone, &db, &RuntimeBoundary::Flush, move || { + contention_tx + .send(()) + .expect("contention signal channel should still be open"); + }); + result_tx + .send(()) + .expect("result signal channel should still be open"); + outcome + }); + + contention_rx.recv_timeout(Duration::from_secs(5)).expect( + "coordinate() should reach the WorktreeLock try_lock loop and observe \ + TryLockError::WouldBlock while the first guard is still held", + ); + + assert!( + result_rx.recv_timeout(Duration::from_millis(300)).is_err(), + "the worker's coordinate() call must not complete while the first worktree lock guard is still held" + ); + + drop(held); + + result_rx.recv_timeout(Duration::from_secs(5)).expect( + "the same coordinate() invocation should complete once the first guard is released", + ); + + let outcome = worker + .join() + .expect("worker thread should not panic") + .expect( + "coordinate() should succeed once it can acquire the worktree lock after release", + ); + assert_eq!( + outcome.revision, 0, + "the worker's first-observation flush should not advance the revision" + ); + + remove_test_repo(&repo_root); + } } diff --git a/cli/src/services/mutation_trace/runtime/worktree_lock.rs b/cli/src/services/mutation_trace/runtime/worktree_lock.rs index a7165c07..21422441 100644 --- a/cli/src/services/mutation_trace/runtime/worktree_lock.rs +++ b/cli/src/services/mutation_trace/runtime/worktree_lock.rs @@ -43,7 +43,7 @@ impl WorktreeLock { } } -fn acquire_inner( +pub(super) fn acquire_inner( git_dir: &Path, timeout: Duration, on_contention: F, diff --git a/context/cli/mutation-trace-protocol.md b/context/cli/mutation-trace-protocol.md index fa4c3f69..5e7d96ba 100644 --- a/context/cli/mutation-trace-protocol.md +++ b/context/cli/mutation-trace-protocol.md @@ -212,14 +212,14 @@ loads a `ProtocolState`; `protocol.rs` only transitions already-known ones. ## Target end-state architecture The plan's file split anticipated three seams beyond `protocol.rs`. `store.rs`, -`runtime/git_snapshot.rs`, and `coordinator.rs`'s internal pipeline now exist -as real call sites (see +`runtime/git_snapshot.rs`, and `coordinator.rs` (with its public `coordinate()` +entrypoint) now all exist as real call sites (see [`mutation-trace-runtime-coordinator.md`](mutation-trace-runtime-coordinator.md)); -only `coordinator.rs`'s public `coordinate()` remains future work: +only cross-module integration tests and harness/command wiring remain: ```mermaid flowchart LR - coordinator["coordinator.rs (internal pipeline implemented;\npublic coordinate() future work)\n(imperative shell:\nDB load, Git snapshot,\nCAS/retry, persist)"] + coordinator["coordinator.rs (implemented)\n(imperative shell:\nlock, DB load, Git snapshot,\nCAS/retry, persist)"] protocol["protocol.rs\n(pure transitions —\nprepare/commit/attribution/\ntaint/abandon/recover\nall implemented)"] git_snapshot["runtime/git_snapshot.rs (implemented)\n(isolated Git snapshot,\ntemporary index, tree capture/diff,\nSCE-owned ref pinning)"] store["store.rs\n(cursor/revision, scopes,\nprocessed events, mutation\nevidence, CAS transaction)"] @@ -229,9 +229,9 @@ flowchart LR coordinator --> store ``` -- **`coordinator.rs`** (internal pipeline implemented) — resolves scope - actor/worktree identity, loads a `ProtocolState` via `store.rs`, and calls - the pure protocol; its public `coordinate()` entrypoint remains future work. +- **`coordinator.rs`** (implemented) — its public `coordinate()` acquires the + per-worktree lock, derives `WorktreeId` from checkout identity, drives one + Git snapshot, and calls the pure protocol under a bounded CAS-retry loop. - **`runtime/git_snapshot.rs`** (implemented) — captures an isolated worktree snapshot and pins it durably; called by `coordinator.rs`. - **`store.rs`** (implemented) — loads and persists worktree/scope/event state diff --git a/context/cli/mutation-trace-runtime-coordinator.md b/context/cli/mutation-trace-runtime-coordinator.md index a2dff5d9..16aeb56a 100644 --- a/context/cli/mutation-trace-runtime-coordinator.md +++ b/context/cli/mutation-trace-runtime-coordinator.md @@ -1,6 +1,6 @@ # Mutation-trace runtime coordinator (`mutation_trace::runtime`) -The imperative-shell layer that will connect the verified, pure +The imperative-shell layer that connects the verified, pure mutation-cursor protocol kernel ([`protocol.rs`](mutation-trace-protocol.md)) and its persistence layer ([`store.rs`](mutation-trace-store.md)) to a real Git worktree, built by the `mutation-cursor-runtime-coordinator` plan @@ -8,10 +8,12 @@ Git worktree, built by the `mutation-cursor-runtime-coordinator` plan `cli/src/services/mutation_trace/runtime/` is a private submodule (`pub(crate) mod runtime;` in `mutation_trace/mod.rs`), registered under the -same `#[allow(dead_code)]` precedent as the rest of `mutation_trace`: only -`coordinator.rs`'s eventual public entrypoints will be reachable from outside -`runtime` once they exist, and nothing under `runtime/` is wired into any -hook, command, or `diff_traces` insertion yet. +same `#[allow(dead_code)]` precedent as the rest of `mutation_trace`. +`coordinator::coordinate()` is the public entrypoint, but `runtime/mod.rs` +still declares `mod coordinator;` privately, so `coordinate()` is reachable +only from within `runtime` itself (its own tests) for now; a `pub(crate)` +re-export is deferred until a harness adapter needs it. Nothing under +`runtime/` is wired into any hook, command, or `diff_traces` insertion yet. `runtime` depends on `protocol`/`store`/`types` and on `services::checkout`, never the reverse — this is a structural module boundary, not merely a @@ -19,10 +21,11 @@ documented convention. ## Current code surface -The per-worktree runtime lock, the isolated Git snapshot service, and the -coordinator's internal protocol-integration pipeline exist so far; only the -public, lock-wrapped `coordinate()` entrypoint that drives the lock and -checkout identity around that pipeline is still future work. +The per-worktree runtime lock, the isolated Git snapshot service, the +coordinator's internal protocol-integration pipeline, and the public, +lock-wrapped `coordinate()` entrypoint that drives the lock and checkout +identity around that pipeline all exist. Only cross-module integration tests +and any harness/command wiring remain. - `cli/src/services/mutation_trace/runtime/worktree_lock.rs` — `WorktreeLock::acquire(git_dir: &Path, timeout: Duration) -> @@ -84,9 +87,19 @@ checkout identity around that pipeline is still future work. `Flush` carries nothing — its worktree is always the invocation's own already-resolved one, never caller-supplied) and documents the `(ScopeId, EventId)` replay-identity contract a future harness adapter must - uphold. The internal, generic-over-`SnapshotCapture` pipeline — not yet - reachable outside `runtime`, since the public `coordinate()` entrypoint is - still future work (see Status) — does, per invocation: capture and pin + uphold. The public `coordinate(repository_root, db, boundary) -> + Result` entrypoint owns the critical + section: resolve `git_dir` via `checkout::resolve_git_dir`, acquire the + `WorktreeLock` (bounded 10s, held for the whole call), resolve checkout + identity via `checkout::get_or_create_checkout_id` and wrap it as + `WorktreeId` — no caller-supplied `WorktreeId` or `Boundary` is ever + accepted — then construct `GitSnapshotService` and delegate to the internal, + generic-over-`SnapshotCapture` pipeline. (`coordinate()` is a one-line + delegation to a private `coordinate_inner(.., on_lock_contention: + impl FnOnce())` test seam; production passes a no-op closure.) A + `WorktreeLock` acquisition failure (timeout or I/O) surfaces as + `CoordinateError::LockAcquisition`. The + pipeline does, per invocation: capture and pin exactly one Git snapshot; on failure, run a bounded taint-retry loop instead (below) and return without touching the rest of the pipeline; on success, idempotently materialize the worktree row and, for hook boundaries, the @@ -110,10 +123,11 @@ checkout identity around that pipeline is still future work. `Conflict`, reporting `persisted_taint: false` only once every bounded attempt has been exhausted. -The runtime lock will guard the coordinator's own critical section (snapshot -capture, worktree/scope materialization, recovery, and the CAS retry loop) -once the public `coordinate()` entrypoint wires it in. It will be held on -every `coordinate()` call, unlike the checkout-identity-creation lock. +The runtime lock guards the coordinator's own critical section (snapshot +capture, worktree/scope materialization, recovery, and the CAS retry loop): +`coordinate()` acquires it before resolving checkout identity and holds it +until the call returns. It is held on every `coordinate()` call, unlike the +checkout-identity-creation lock. ## Two distinct locks, two distinct invariants @@ -180,17 +194,27 @@ taint-retry loop taints an existing worktree, survives a losing CAS before committing on retry, reports `persisted_taint: false` once exhausted, makes no write when no worktree row exists yet, and still finds and taints a worktree another caller materializes concurrently during this invocation's -own failing capture. +own failing capture. One further test proves the critical-section +serialization with a real happens-before ordering: `coordinate()` delegates +to a private `coordinate_inner(.., on_lock_contention: impl FnOnce())` that +takes the lock via `worktree_lock::acquire_inner` (T02's seam, `pub(super)`), +and with a first `WorktreeLock` held, a worker's `coordinate_inner` call +observes the real `TryLockError::WouldBlock` branch (signalling a channel +from `on_lock_contention`) while that first guard is still alive, then — once +the guard is dropped — the same invocation acquires the lock and returns +`Ok`. Production `coordinate()` passes a no-op closure, so its code path, +signature, and lock semantics are unchanged. ## Status -The per-worktree runtime lock, the isolated Git snapshot service, and the -coordinator's internal protocol-integration pipeline (above) are implemented. -The public, lock-wrapped `coordinate()` entrypoint (resolving `git_dir`, -acquiring `WorktreeLock`, resolving checkout identity, deriving `WorktreeId`) -and cross-module integration tests remain future work tracked by the -`mutation-cursor-runtime-coordinator` plan's task stack. This file will grow -to cover them as each lands. +The per-worktree runtime lock, the isolated Git snapshot service, the +coordinator's internal protocol-integration pipeline (above), and the public, +lock-wrapped `coordinate()` entrypoint (resolving `git_dir`, acquiring +`WorktreeLock`, resolving checkout identity, deriving `WorktreeId`, delegating +to the pipeline) are all implemented. Cross-module integration tests, a +`pub(crate)` re-export of `coordinate()` beyond `runtime`, and any +harness/command wiring remain future work tracked by the +`mutation-cursor-runtime-coordinator` plan's task stack and its follow-ups. See also: [`mutation-trace-protocol.md`](mutation-trace-protocol.md), [`mutation-trace-store.md`](mutation-trace-store.md), diff --git a/context/cli/mutation-trace-store.md b/context/cli/mutation-trace-store.md index 0d58ef9b..b2527732 100644 --- a/context/cli/mutation-trace-store.md +++ b/context/cli/mutation-trace-store.md @@ -118,9 +118,9 @@ returns `Err`. ## Non-goals - No Git or filesystem I/O — `store.rs` itself calls neither Git nor the - filesystem. `runtime/git_snapshot.rs` and `runtime/coordinator.rs`'s - internal pipeline now exist and wire this store together with the Git - snapshot service (see + filesystem. `runtime/git_snapshot.rs` and `runtime/coordinator.rs` (through + its public `coordinate()` entrypoint) now wire this store together with the + Git snapshot service (see [`mutation-trace-runtime-coordinator.md`](mutation-trace-runtime-coordinator.md)); `store.rs` itself remains unmodified and still performs no Git or filesystem I/O of its own. diff --git a/context/context-map.md b/context/context-map.md index f0f8353e..2d00b086 100644 --- a/context/context-map.md +++ b/context/context-map.md @@ -23,11 +23,11 @@ Feature/domain context: - `context/cli/config-precedence-contract.md` (implemented `sce config` show/validate command contract, deterministic `flags > env > config file > defaults` resolution order, focused `config/resolver.rs` ownership for config discovery/merge/runtime precedence plus default-discovered invalid-file degradation, focused `config/render.rs` ownership for `show`/`validate` text+JSON output construction, canonical `$schema` acceptance for startup-loaded `sce/config.json` files, shared auth-key env/config/optional baked-default support starting with `workos_client_id`, shared runtime resolution for flat logging observability keys including config-file/default `log_to_file`, `log_dir` / `SCE_LOG_DIR` with `/sce/logs` defaulting plus config-file/default-only positive `log_file_retention_limit`, config-file-only `agent_trace.repository_id`/`agent_trace.repository_remote` repository-identity keys with default remote `origin`, default-enabled `agent_trace.auto_sync` boolean resolution with explicit-false opt-out for the post-commit trigger boundary, the catalog-derived `integrations.optional_workflows` optional-workflow selection key, JSON-pointer-prefixed schema-validation errors, canonical Pkl-generated `sce/config.json` schema ownership plus CLI embedding/reuse contract including `policies.attribution_hooks.enabled` default-true/explicit-false opt-out metadata, config-file selection order, `show` provenance output, and trimmed `validate` output contract) - `context/cli/capability-traits.md` (current broad CLI capability seam in `cli/src/services/capabilities.rs`, including `FsOps`/`StdFsOps`, `GitOps`/`ProcessGitOps`, git root/hooks resolution behavior, compile-time-typed borrowed AppContext wiring with associated-type narrow capability accessors plus `ContextWithRepoRoot` repo-root-scoped context derivation, generic command execution bounds, and test-only unimplemented stubs; current service internals do not consume fs/git traits until later lifecycle migration tasks) - `context/cli/service-lifecycle.md` (current compile-safe lifecycle seam in `cli/src/services/lifecycle.rs`, including default no-op `ServiceLifecycle` diagnose/fix/setup methods against narrow `HasRepoRoot`, lifecycle-owned health/fix/setup result types with generic setup messages, doctor/setup adapter boundaries, the static `LifecycleProvider` enum catalog/dispatcher, hook/config/local_db/auth_db/agent_trace_db lifecycle providers including setup-time repository-scoped Agent Trace DB initialization plus checkout identity diagnostics, implemented doctor aggregation over diagnose/fix providers, and implemented setup aggregation over `setup` providers in order config → local_db → auth_db → agent_trace_db → hooks when requested) -- `context/cli/mutation-trace-protocol.md` (pure, dependency-free `cli/src/services/mutation_trace/` refinement of the verified `spec/mutation_cursor.qnt` mutation-cursor protocol: the full pure action set — `prepare`/`commit` transition logic, attribution/mutation-event materialization, snapshot-failure/database-failure taint actions, scope abandonment, and recovery with an explicit observed-tree input — plus cross-action sequence/invariant tests and a module-level Quint refinement matrix are all complete (`types.rs`, `protocol.rs`, `mod.rs`, `tests.rs`, registered with `#[allow(dead_code)]`); opaque-newtype identity refinement decisions vs. the Quint model's bounded enums, and the `coordinator.rs` target end-state seam this layout leaves room for but does not create — `store.rs`, `runtime/git_snapshot.rs`, and `runtime/coordinator.rs`'s internal pipeline now exist, built out by the `mutation-cursor-store-persistence` and `mutation-cursor-runtime-coordinator` plans respectively; `protocol.rs`'s pure transitions are not yet wired into any hook or command) +- `context/cli/mutation-trace-protocol.md` (pure, dependency-free `cli/src/services/mutation_trace/` refinement of the verified `spec/mutation_cursor.qnt` mutation-cursor protocol: the full pure action set — `prepare`/`commit` transition logic, attribution/mutation-event materialization, snapshot-failure/database-failure taint actions, scope abandonment, and recovery with an explicit observed-tree input — plus cross-action sequence/invariant tests and a module-level Quint refinement matrix are all complete (`types.rs`, `protocol.rs`, `mod.rs`, `tests.rs`, registered with `#[allow(dead_code)]`); opaque-newtype identity refinement decisions vs. the Quint model's bounded enums, and the `coordinator.rs` target end-state seam this layout leaves room for but does not create — `store.rs`, `runtime/git_snapshot.rs`, and `runtime/coordinator.rs` (including its public `coordinate()` entrypoint) now exist, built out by the `mutation-cursor-store-persistence` and `mutation-cursor-runtime-coordinator` plans respectively; `protocol.rs`'s pure transitions are not yet wired into any hook or command) - `context/cli/mutation-trace-revision-refinement.md` (the Quint `revision: int` → Rust `WorktreeState::revision: u64` bounded-integer refinement: the private `next_revision` checked-arithmetic helper `commit`/`taint`/`abandon`/`recover` all route through instead of a raw `+ 1`, so a worktree at `revision: u64::MAX` is a guarded no-op/rejection rather than a silent wrap to `0`) - `context/cli/mutation-trace-quint-connect.md` (`#[cfg(test)]`-only Quint Connect model-based-testing harness in `cli/src/services/mutation_trace/mbt/` continuously checking `protocol.rs` against `spec/mutation_cursor.qnt`: the verification-only `mbtAction`/`MbtAction` record-payload transport excluded from comparison, the operation-identity-vs-`MbtStutter` distinction on guarded/no-op branches with its two deterministic regressions, finite ID mapping, the AC5 comparable-state field list, `randomPrepare` staying a single `step` branch, deterministic/generated (500×30, seed-reproducible) test coverage, and the two Nix checks — generic `checks.cli-tests` and dedicated `checks.mutation-trace-quint-connect` — that both require the pinned Quint binary plus the top-level `spec/` directory in `workspaceSrc`'s Nix fileset) - `context/cli/mutation-trace-store.md` (durable persistence for the mutation-cursor protocol in `cli/src/services/mutation_trace/store.rs`, built by the `mutation-cursor-store-persistence` plan: the one-directional `protocol.rs` -> `DurableTransition::between` (pure structural diff) -> `store.rs` (SQL translation) -> `RepositoryAgentTraceDb` boundary; migration `003_mutation_trace_protocol.sql`'s five tables; `AttemptState`/`external_taint` non-persistence; the 8-byte big-endian `BLOB` revision encoding plus explicit non-`Debug` enum codecs; the bounded hot-path `load_worktree` read vs. the cold-path `load_mutation_event`; `commit`'s single-`BEGIN IMMEDIATE` CAS batch via `TursoDb::execute_transactional_cas_batch`, distinguishing `Conflict`/retryable-transient/deterministic-`Err` outcomes; and the store's non-goals — no Git/filesystem I/O, no attribution/boundary-kind decisions, no retry-after-`Conflict` loop, no terminal-scope garbage collection) -- `context/cli/mutation-trace-runtime-coordinator.md` (imperative-shell runtime layer in `cli/src/services/mutation_trace/runtime/`, built by the `mutation-cursor-runtime-coordinator` plan: the per-worktree OS advisory lock, `runtime::worktree_lock::WorktreeLock::acquire(git_dir, timeout)` at `/sce/mutation-cursor.lock`, bounded `try_lock()` polling with a distinct matchable timeout error and RAII release, and its distinction from the separate checkout-identity-creation lock; and the isolated Git snapshot service, `runtime::git_snapshot::GitSnapshotService` (`capture_tree`/`pin_tree`/`diff_trees`), writing tree/blob objects into the repository's normal object database and pinning durable ones via a create-only, idempotent `refs/sce/mutation-cursor//` ref rather than a private object store; and `runtime::coordinator`'s internal, generic-over-`SnapshotCapture` protocol-integration pipeline (`RuntimeBoundary`/`CoordinateOutcome`/`CoordinateError`, the load → recover-if-needed → prepare/commit CAS-retry loop, and its own bounded snapshot-failure taint-retry loop); the public `coordinate()` entrypoint and cross-module integration tests remain future work) +- `context/cli/mutation-trace-runtime-coordinator.md` (imperative-shell runtime layer in `cli/src/services/mutation_trace/runtime/`, built by the `mutation-cursor-runtime-coordinator` plan: the per-worktree OS advisory lock, `runtime::worktree_lock::WorktreeLock::acquire(git_dir, timeout)` at `/sce/mutation-cursor.lock`, bounded `try_lock()` polling with a distinct matchable timeout error and RAII release, and its distinction from the separate checkout-identity-creation lock; and the isolated Git snapshot service, `runtime::git_snapshot::GitSnapshotService` (`capture_tree`/`pin_tree`/`diff_trees`), writing tree/blob objects into the repository's normal object database and pinning durable ones via a create-only, idempotent `refs/sce/mutation-cursor//` ref rather than a private object store; and `runtime::coordinator`'s protocol-integration pipeline (`RuntimeBoundary`/`CoordinateOutcome`/`CoordinateError`, the load → recover-if-needed → prepare/commit CAS-retry loop, and its own bounded snapshot-failure taint-retry loop) plus the public `coordinate(repository_root, db, boundary)` entrypoint that acquires the `WorktreeLock`, derives `WorktreeId` from `checkout::get_or_create_checkout_id`, and drives that pipeline under one held lock; cross-module integration tests, a `pub(crate)` re-export of `coordinate()` beyond `runtime`, and harness/command wiring remain future work) - `context/sce/cli-exit-code-contract.md` (stable class-based `sce` process exit-code contract in `cli/src/app.rs`, so automation can branch on failure category without parsing error text) - `context/sce/cli-error-code-taxonomy.md` (stable user-facing `SCE-ERR-*` diagnostic code classes rendered by `cli/src/app.rs`, complementing the numeric exit-code classes) - `context/sce/cli-stdout-stderr-contract.md` (implemented stream contract in `cli/src/app.rs`: command payloads on stdout only, redacted diagnostics on stderr) diff --git a/context/overview.md b/context/overview.md index 89fd22fb..aa16b6fb 100644 --- a/context/overview.md +++ b/context/overview.md @@ -2,7 +2,7 @@ This repository maintains shared assistant configuration for OpenCode, Claude, Pi, and Codex from a single canonical Pkl authoring source. One typed workflow catalog owns the six workflows' shared identity and target routing metadata, while canonical workflow/phase modules own behavior and migrated package-local documents, and target renderers own formatting. Generated target layouts are ephemeral: repository builds consume a pre-Cargo generated payload through `SCE_CLI_GENERATED_INPUT_DIR`, crates.io and Flatpak stage packaging-only fallbacks, and `config/.opencode`, `config/.claude`, `config/.pi`, `config/.agents`, and the generated SCE config schema are not committed. The catalog also marks a workflow `optional` — currently only `brownfield` — which changes nothing about generation and is projected into a generated `config/optional-workflows.json` manifest for install-time consumers. `nix run .#pkl-check-generated` preserves its exact 141-path artifact, metadata/package, phase-reference, internal-reference, optional-workflow-manifest, workflow-orchestration, OpenCode-permission, required-path, and forbidden-path checks while delegating deterministic payload production and inventories to the shared generated-input producer; `nix flake check` runs the same contract. The target matrix contains one manual OpenCode profile plus Claude, Pi, and Codex; the former automated OpenCode profile has been removed. A fourth Pkl renderer, `config/pkl/renderers/codex-content.pkl`, also consumes the same canonical workflow composition to emit skills-only Codex output under `config/.agents/skills/**` (no per-target frontmatter, matching Pi, and no command/prompt layer). Each of the six catalog workflow skills (not `sce-decision`, which has no user-facing entrypoint on any target) additionally carries a Codex-only `agents/openai.yaml`, rendered by `config/pkl/renderers/codex-metadata.pkl` from the shared catalog's `title`/`description` plus an authored `default_prompt`, with `policy.allow_implicit_invocation: false` so these stateful lifecycle workflows activate only from explicit `$sce-` invocation or Codex's `/skills` discovery, never from conversational relevance alone. The Codex renderer also emits a Codex hook registration file (`config/.codex/hooks.json`) and its fail-open install-guidance hook script (`config/.codex/hooks/run-sce-or-show-install-guidance.sh`), both routing every registered lifecycle event through the single command `sce hooks codex`; the generated command resolves the Git root at invocation time and safely invokes the helper from nested cwd or spaced repository paths, failing open silently when Git-root resolution fails. `sce setup --codex` (and `--all`) now installs it as a fourth `SetupTarget`; setup merges the user-owned `.codex/hooks.json` through the shared structural Codex hook-config service, preserving unrelated valid handlers and rejecting malformed or Codex-invalid documents before the atomic swap. `sce doctor` diagnoses that file per required registration (`PresentAndCurrent`/`Missing`/`Stale`, or `Malformed` for the whole document) rather than by whole-document equality, and separately reports whether Codex has actually marked each structurally current registration trusted by reading (never writing) Codex's own `$CODEX_HOME/config.toml` hook-trust state; `sce doctor --fix` repairs structurally unhealthy registrations through the same merge service but never touches trust state (see `context/sce/doctor-human-text-contract.md`). `sce hooks codex` now exists as a typed dispatcher (`cli/src/services/hooks/codex/`): it parses the raw hook JSON into a `CodexHookEvent` and classifies `(hook_event_name, tool_name)` into `UserPromptSubmit`, `Stop`, `PreToolUse(Bash)`, `PostToolUse(apply_patch)`, or a fail-open `NoOp` fallthrough covering every other combination — including `apply_patch` under `PreToolUse`. `UserPromptSubmit` and `Stop` each capture one `messages`/`parts` row (`role="user"`/`role="assistant"`) into the repository Agent Trace DB under the idempotent `cx_` session prefix, `PreToolUse(Bash)` delegates to the existing Bash policy engine, and `PostToolUse(apply_patch)` parses Codex's own `apply_patch` text, resolves paths from the event `cwd` against the real Git root into safe repository-relative paths, normalizes Add/Update evidence into an SCE unified diff under deterministic event-scoped synthetic line identities derived from `tool_use_id`, and persists it as one `diff_traces` row when non-empty (see `context/sce/codex-integration-runtime.md`) — while invalid cwd/path resolution and malformed STDIN payloads fail open with empty stdout, reported model IDs are preserved without fabricated provider prefixes, and missing/blank apply_patch sessions produce no evidence. Delete-File operations and Bash-triggered filesystem mutations remain untracked for Codex. -It also includes a Rust CLI (`sce`) for Shared Context Engineering workflows: auth, config inspection, setup, doctor, Agent Trace hooks and synchronization, bash-policy evaluation, and repository-scoped Agent Trace storage infrastructure. See `context/architecture.md` for module-level boundaries and `context/context-map.md` for the full domain file index. A new pure, dependency-free `cli/src/services/mutation_trace/` module now exists as a Rust refinement of the verified `spec/mutation_cursor.qnt` mutation-cursor protocol; it currently carries domain types and the full pure action set — `prepare`/`commit` transition logic, attribution/mutation-event materialization, snapshot-failure/database-failure taint actions, scope abandonment, and recovery with an explicit observed-tree input — is registered with `#[allow(dead_code)]`. A separate `cli/src/services/mutation_trace/store.rs` persistence layer (built out by the `mutation-cursor-store-persistence` plan) now provides a real, CAS-guarded database call site against the repository-scoped Agent Trace DB, but the module is still not wired into any hook or command (see `context/cli/mutation-trace-protocol.md`). +It also includes a Rust CLI (`sce`) for Shared Context Engineering workflows: auth, config inspection, setup, doctor, Agent Trace hooks and synchronization, bash-policy evaluation, and repository-scoped Agent Trace storage infrastructure. See `context/architecture.md` for module-level boundaries and `context/context-map.md` for the full domain file index. A new pure, dependency-free `cli/src/services/mutation_trace/` module now exists as a Rust refinement of the verified `spec/mutation_cursor.qnt` mutation-cursor protocol; it currently carries domain types and the full pure action set — `prepare`/`commit` transition logic, attribution/mutation-event materialization, snapshot-failure/database-failure taint actions, scope abandonment, and recovery with an explicit observed-tree input — is registered with `#[allow(dead_code)]`. A separate `cli/src/services/mutation_trace/store.rs` persistence layer (built out by the `mutation-cursor-store-persistence` plan) now provides a real, CAS-guarded database call site against the repository-scoped Agent Trace DB, and a `cli/src/services/mutation_trace/runtime/` coordinator layer (built out by the `mutation-cursor-runtime-coordinator` plan) adds the per-worktree advisory lock, the isolated Git snapshot/ref-pinning service, and a public `coordinate()` entrypoint that drives the protocol against a real worktree under that lock; the module is still not wired into any hook or command (see `context/cli/mutation-trace-protocol.md` and `context/cli/mutation-trace-runtime-coordinator.md`). The generated `/next-task` workflow persists task-level context-synchronization lifecycle state in each plan (`pending`, `synced`, or `blocked`) so unresolved task synchronization debt survives a session boundary and gates new implementation. Successful `/next-task` execution hands task synchronization an explicit, pre-edit-Git-baseline-relative changed-file list plus implementation, verification, done-check, plan-update, and context-impact evidence, recorded directly on the completed task (`Completed`, `Files changed`, `Result`, `Verify`, `Context impact`, `Context synchronization`); the five-file root context pass remains mandatory. A later-session sync-debt retry reads that same completed task record directly from the plan by plan path and task ID, with no separate persisted synchronization handoff. `/validate` is validation-only: it runs final checks, writes the Validation Report, and reports `validated`, `failed`, or `blocked` without plan-level context synchronization. diff --git a/context/plans/mutation-cursor-runtime-coordinator.md b/context/plans/mutation-cursor-runtime-coordinator.md index c4d37548..9fbed49b 100644 --- a/context/plans/mutation-cursor-runtime-coordinator.md +++ b/context/plans/mutation-cursor-runtime-coordinator.md @@ -916,7 +916,7 @@ Persist this field in every plan; this is durable plan state, not chat state: existing context files use for it. - Context synchronization: synced -- [ ] T05: `Wire the worktree lock and checkout identity into coordinate()` (status:todo) +- [x] T05: `Wire the worktree lock and checkout identity into coordinate()` (status:done) - Task ID: T05 - Scope: In — the public `coordinate(repository_root, db, boundary)` entrypoint: resolve `git_dir`, acquire `WorktreeLock`, resolve checkout @@ -930,7 +930,145 @@ Persist this field in every plan; this is durable plan state, not chat state: `WorktreeLock` drops); `coordinate()`'s public signature matches `Result`. - Verify: `cargo test -p shared-context-engineering runtime::coordinator::tests::two_threads_on_the_same_worktree_serialize` (via `./scripts/run-cli-cargo.sh test`) - - Context synchronization: pending + - Completed: 2026-08-30 + - Files changed: `cli/src/services/mutation_trace/runtime/coordinator.rs`, + `cli/src/services/mutation_trace/runtime/worktree_lock.rs` + - Result: Added the public `coordinate(repository_root: &Path, db: + &RepositoryAgentTraceDb, boundary: &RuntimeBoundary) -> + Result` entrypoint to + `runtime/coordinator.rs`. It resolves `git_dir` via + `checkout::resolve_git_dir(repository_root)` (error → `CoordinateError::Other`), + acquires the `WorktreeLock` (`WORKTREE_LOCK_TIMEOUT`, 10s) into an + RAII guard held for the whole critical section, resolves checkout identity + via T01's now concurrency-/crash-safe `checkout::get_or_create_checkout_id(&git_dir)` + (error → `CoordinateError::Other`), wraps the result as + `WorktreeId(checkout_id)` — no caller-supplied `WorktreeId` or `Boundary` + is ever accepted — constructs `GitSnapshotService::new(repository_root)` + (which keeps its own internal `--absolute-git-dir` resolution, unchanged), + and delegates to T04's unchanged internal `coordinate_boundary(db, + &snapshot, &worktree_id, boundary)` for the rest of the critical section. + `coordinate()` is a one-line delegation to a private + `coordinate_inner(.., on_lock_contention: impl FnOnce())` (see the test + paragraph below); production passes a no-op closure, so the code path is + identical. Both `WorktreeLockError` variants (`TimedOut`, `Io`) are mapped + to `CoordinateError::LockAcquisition(anyhow::Error)` via a small private + `lock_acquisition` helper. `WORKTREE_LOCK_TIMEOUT` is a new module-level + `const Duration = Duration::from_secs(10)`, matching the plan's "Blocking + vs. timeout" design decision (bounded 10s deadline, polled at a 100ms + interval). `worktree_lock.rs`'s existing private `acquire_inner` seam is + widened to `pub(super)` (no behavior change); `WorktreeLock::acquire`'s + public signature and semantics are untouched. The T04 pipeline, + `RuntimeBoundary`, `CoordinateOutcome`, and `CoordinateError` are + otherwise unchanged. + + Deviations from the pre-implementation gate summary, both reversible and + matching surrounding style: (1) `coordinate` is a plain `pub fn` (not + `pub(crate)`), mirroring `git_snapshot.rs`'s `pub fn capture_tree` and + `worktree_lock.rs`'s `pub fn acquire` in the same private `runtime` + module tree; (2) `runtime/mod.rs` was left unchanged — a private `mod + coordinator;` is already reachable from `runtime`'s own future + `#[cfg(test)] mod tests` (T06) as `super::coordinator::coordinate`, so no + visibility change to `mod.rs` was needed. The dead-code warning for the + not-yet-called entrypoint is already covered by the existing + `#[allow(dead_code)]` on `pub mod mutation_trace;` in `services/mod.rs` + (same coverage T04 relied on for the unused `LockAcquisition` variant). + + `runtime::coordinator::tests::two_threads_on_the_same_worktree_serialize` + (the plan's exact Verify function name) proves the critical-section + serialization through a real happens-before ordering rather than a + scheduling window. `coordinate()` delegates to a private + `coordinate_inner(repository_root, db, boundary, on_lock_contention: + impl FnOnce())`, which acquires the lock via + `worktree_lock::acquire_inner` (T02's existing seam, its visibility + widened from private to `pub(super)` so `coordinator` can reach it — + `acquire_inner`'s `on_contention` callback still fires exactly once, only + from inside the real `Err(TryLockError::WouldBlock)` arm of the same + `try_lock()` poll loop). `coordinate()` itself is + `coordinate_inner(.., .., .., || {})`, so production takes the identical + code path and its public signature, `WORKTREE_LOCK_TIMEOUT`, and every + lock/CAS/snapshot behavior are unchanged; no fake locking implementation + exists. The test creates a real `git init` temp repository and a real + temp-file `RepositoryAgentTraceDb`, holds a real first `WorktreeLock` + (acquired via `acquire_inner(.., || {})`), spawns a worker calling + `coordinate_inner(.., Flush, move || contention_tx.send(()))`, and blocks + on `contention_rx` (5s deadlock guard only). Receiving the contention + signal is the primary proof: the worker's `coordinate()` reached the real + `try_lock()` loop and observed `WouldBlock` while the first guard was + provably still alive (the main thread cannot drop it until after the + `recv`). A secondary `result_rx.recv_timeout(300ms).is_err()` asserts the + same invocation has not completed while the guard is held; then the first + guard is dropped and the same worker invocation is required to acquire the + lock and return `Ok` with `revision == 0` (first-observation flush) — no + restart, no manual retry. This mirrors `worktree_lock.rs`'s own corrected + T02 contention test (`try_lock()` → `WouldBlock` → contention observer) + and makes the same causal claim; it fails if `coordinate()` were changed + to enter its critical section without contending on the worktree lock. The + earlier revision of this test signalled a pre-call `started` channel and + asserted only non-completion for 500ms — a scheduler false positive (a + worker descheduled between the signal and the `coordinate()` call would + pass without ever reaching `try_lock()`); that framing, and any claim it + was equivalent to the `worktree_lock.rs` contention test, is superseded. + - Verify (actual): `./scripts/run-cli-cargo.sh test --manifest-path + cli/Cargo.toml runtime::coordinator::tests::two_threads_on_the_same_worktree_serialize` + → 1/1 passed, re-run 5 consecutive times in isolation with no flakiness + after the PR #244 review follow-up (the strengthened `WouldBlock`-observed + proof). Additionally ran `./scripts/run-cli-cargo.sh test + --manifest-path cli/Cargo.toml runtime::` → 37/37 passed (no regression + in `worktree_lock`, `git_snapshot`, or the T04 coordinator tests); + `./scripts/run-cli-cargo.sh clippy --manifest-path cli/Cargo.toml + --all-targets -- -D warnings` → clean (pedantic/warnings denied + workspace-wide, per Constraints; `cargo fmt` run once to apply the + follow-up's formatting); `./scripts/run-cli-cargo.sh fmt + --manifest-path cli/Cargo.toml -- --check` → clean. Full CLI suite + (`./scripts/run-cli-cargo.sh test --manifest-path cli/Cargo.toml`) → 821 + passed / 0 failed on the baseline-plus-one count (pre-change baseline was + 820/820). This project's full parallel test run is intermittently flaky in + modules unrelated to this task: across T05's original and follow-up + verification, three distinct full-suite runs each surfaced a different + single failure in `sync`, `agent_trace_db::repository`, or + `agent_trace_export`, every one of which passed immediately when re-run in + isolation and cleared on the next full run. The targeted + `two_threads_on_the_same_worktree_serialize` and `runtime::` runs were + never flaky. + - Context impact: `domain` — completes the mutation-cursor-runtime + coordinator by adding its public `coordinate()` entrypoint under the + domain file T02/T03/T04 established + (`context/cli/mutation-trace-runtime-coordinator.md`). Affected areas for + context synchronization to verify/update: + `context/cli/mutation-trace-runtime-coordinator.md` (document the public + `coordinate()` entrypoint, its lock-wrapped critical-section ordering — + resolve `git_dir` → acquire `WorktreeLock` (bounded 10s) → resolve + checkout identity → derive `WorktreeId` → delegate to the T04 pipeline — + and the `WorktreeLockError` → `CoordinateError::LockAcquisition` mapping); + `context/cli/mutation-trace-protocol.md`'s "Target end-state + architecture", which still frames `coordinator.rs`'s public entrypoint as + future work — now stale, the file's public API exists (only harness/CLI + wiring and `diff_traces` insertion remain out of scope); + `context/cli/mutation-trace-store.md`'s "Non-goals" line about + `coordinator.rs`; `context/context-map.md`'s `coordinator.rs` annotation + (the plan's own Context sync note says "update once T06 lands" — T05 is + the task that creates the public seam, so re-check the wording now); + `context/overview.md`'s `mutation_trace` description (the runtime + coordinator now has a usable public entrypoint, still not wired into any + hook or command). The five root context files must still be verified per + the mandatory pass. Reason: this is the first task to expose a public, + lock-wrapped `coordinate()` API, directly falsifying the "public + entrypoint remains future work" framing multiple context files carry. + + PR #244 review follow-up: strengthened + `two_threads_on_the_same_worktree_serialize` from a scheduler-window + non-completion check into a real `TryLockError::WouldBlock` happens-before + proof (see Result). `context/cli/mutation-trace-runtime-coordinator.md` + was re-synced accordingly — its coordinator-bullet and Testing-boundary + paragraphs now describe the `coordinate_inner(.., on_lock_contention)` + seam and the `worktree_lock::acquire_inner` `pub(super)` widening, and no + longer imply the superseded pre-call-`started` framing was equivalent to + `worktree_lock.rs`'s own contention test. No other context file was + affected: `coordinate_inner` is a private test seam with no + public-interface, behavior-contract, or architecture change. The five + root context files were re-verified against the follow-up and remain + accurate as edited during the original T05 sync. + - Context synchronization: synced - [ ] T06: `Add cross-module runtime integration tests` (status:todo) - Task ID: T06 From 1b211202b7435c93ac442bef6b98d8a322404b37 Mon Sep 17 00:00:00 2001 From: David Abram Date: Sun, 30 Aug 2026 09:21:40 +0200 Subject: [PATCH 7/8] runtime: Add mutation trace coordinator integration tests Add end-to-end coverage for linked-worktree coordination, checkout identity convergence, and snapshot failure recovery through coordinate(). Record completion of mutation-cursor-runtime-coordinator T06 and the resulting runtime behavior. Co-authored-by: SCE --- .../services/mutation_trace/runtime/mod.rs | 3 + .../services/mutation_trace/runtime/tests.rs | 261 ++++++++++++++++++ context/cli/mutation-trace-protocol.md | 5 +- .../cli/mutation-trace-runtime-coordinator.md | 48 +++- context/context-map.md | 2 +- .../mutation-cursor-runtime-coordinator.md | 189 +++++++++++-- 6 files changed, 478 insertions(+), 30 deletions(-) create mode 100644 cli/src/services/mutation_trace/runtime/tests.rs diff --git a/cli/src/services/mutation_trace/runtime/mod.rs b/cli/src/services/mutation_trace/runtime/mod.rs index a1b56ba5..794a6e12 100644 --- a/cli/src/services/mutation_trace/runtime/mod.rs +++ b/cli/src/services/mutation_trace/runtime/mod.rs @@ -1,3 +1,6 @@ mod coordinator; mod git_snapshot; mod worktree_lock; + +#[cfg(test)] +mod tests; diff --git a/cli/src/services/mutation_trace/runtime/tests.rs b/cli/src/services/mutation_trace/runtime/tests.rs new file mode 100644 index 00000000..aea7491d --- /dev/null +++ b/cli/src/services/mutation_trace/runtime/tests.rs @@ -0,0 +1,261 @@ +use std::path::{Path, PathBuf}; +use std::process::Command; +use std::sync::atomic::{AtomicU64, Ordering}; +use std::sync::{Arc, Barrier}; +use std::thread; +use std::time::{Duration, SystemTime, UNIX_EPOCH}; + +use crate::services::agent_trace_db::repository::RepositoryAgentTraceDb; +use crate::services::agent_trace_storage::{ + resolve_agent_trace_storage_at_state_root, AgentTraceStorageContext, +}; +use crate::services::checkout::{read_checkout_id, resolve_git_dir}; +use crate::services::mutation_trace::store::MutationTraceStore; +use crate::services::mutation_trace::types::{ActorKind, EventId, ScopeId}; + +use super::coordinator::{coordinate, CoordinateError, RuntimeBoundary}; +use super::git_snapshot::GitSnapshotService; +use super::worktree_lock::WorktreeLock; + +static NEXT_ID: AtomicU64 = AtomicU64::new(0); + +fn unique_path(label: &str) -> PathBuf { + let id = NEXT_ID.fetch_add(1, Ordering::Relaxed); + let nonce = SystemTime::now() + .duration_since(UNIX_EPOCH) + .expect("system time should be after the Unix epoch") + .as_nanos(); + std::env::temp_dir().join(format!( + "sce-mutation-trace-runtime-{label}-{}-{nonce}-{id}", + std::process::id() + )) +} + +fn run_git(dir: &Path, args: &[&str]) { + let output = Command::new("git") + .args(args) + .current_dir(dir) + .output() + .expect("git command should spawn"); + assert!( + output.status.success(), + "git {args:?} failed: {}", + String::from_utf8_lossy(&output.stderr) + ); +} + +fn init_repo(root: &Path) { + std::fs::create_dir_all(root).expect("repository directory should be created"); + run_git(root, &["init", "--quiet"]); + run_git(root, &["config", "user.email", "test@example.com"]); + run_git(root, &["config", "user.name", "Test"]); + run_git(root, &["commit", "--allow-empty", "--quiet", "-m", "init"]); +} + +fn cleanup(path: &Path) { + let _ = std::fs::remove_dir_all(path); +} + +#[test] +fn linked_worktrees_have_independent_locks_and_worktree_ids() { + let main_root = unique_path("linked-main"); + init_repo(&main_root); + let linked_root = unique_path("linked-secondary"); + run_git( + &main_root, + &[ + "worktree", + "add", + "--quiet", + linked_root.to_str().expect("worktree path should be UTF-8"), + ], + ); + + let db_path = main_root.join("agent-trace.db"); + let db_main = RepositoryAgentTraceDb::new_at(&db_path).expect("main repository DB should open"); + + let main_git_dir = resolve_git_dir(&main_root).expect("main git dir should resolve"); + let linked_git_dir = resolve_git_dir(&linked_root).expect("linked git dir should resolve"); + assert_ne!( + main_git_dir, linked_git_dir, + "a linked worktree must resolve its own worktree-specific git dir, giving it a distinct lock and identity path" + ); + + let main_outcome = coordinate(&main_root, &db_main, &RuntimeBoundary::Flush) + .expect("first observation on the main worktree should succeed"); + + let held = WorktreeLock::acquire(&main_git_dir, Duration::from_secs(5)) + .expect("the main worktree's runtime lock should be acquirable"); + + let db_linked = RepositoryAgentTraceDb::open_for_hooks_without_migrations_at(&db_path).expect( + "a second handle to the same caller-supplied repository-scoped DB path should open", + ); + let linked_outcome = coordinate(&linked_root, &db_linked, &RuntimeBoundary::Flush).expect( + "coordinate() on the linked worktree must acquire its own distinct runtime lock while the main worktree's lock is still held", + ); + + drop(held); + + assert_ne!( + main_outcome.worktree_id, linked_outcome.worktree_id, + "each linked worktree must derive a distinct WorktreeId from its own checkout identity" + ); + + let store = MutationTraceStore::new(&db_main); + assert!( + store + .load_worktree(&main_outcome.worktree_id, None, None) + .expect("loading the main worktree row should succeed") + .is_some(), + "the main worktree's coordinator should have persisted a row into the caller-supplied repository-scoped DB" + ); + assert!( + store + .load_worktree(&linked_outcome.worktree_id, None, None) + .expect("loading the linked worktree row should succeed") + .is_some(), + "the linked worktree's coordinator should have persisted its distinct row into the same caller-supplied DB" + ); + + let linked_snapshot = GitSnapshotService::new(&linked_root) + .expect("a snapshot service should construct for the linked worktree"); + linked_snapshot + .diff_trees(&main_outcome.observed_tree, &main_outcome.observed_tree) + .expect("a tree pinned by the main worktree's coordinator must resolve through the linked worktree's git dir"); + + cleanup(&linked_root); + cleanup(&main_root); +} + +#[test] +fn agent_trace_storage_and_coordinator_observe_the_same_checkout_id() { + let repo_root = unique_path("cross-caller-repo"); + init_repo(&repo_root); + run_git( + &repo_root, + &["remote", "add", "origin", "git@github.com:acme/widgets.git"], + ); + let state_root = unique_path("cross-caller-state"); + std::fs::create_dir_all(&state_root).expect("state root should be created"); + + let coordinator_db_path = repo_root.join("coordinator.db"); + let db = RepositoryAgentTraceDb::new_at(&coordinator_db_path) + .expect("the coordinator's repository DB should open"); + + let barrier = Arc::new(Barrier::new(2)); + + let storage_thread = { + let repo_root = repo_root.clone(); + let state_root = state_root.clone(); + let barrier = Arc::clone(&barrier); + thread::spawn(move || { + let context = AgentTraceStorageContext { + repository_root: &repo_root, + explicit_repository_id: None, + repository_remote: "origin", + }; + barrier.wait(); + resolve_agent_trace_storage_at_state_root(&context, &state_root) + .expect("agent_trace_storage resolution should succeed") + .checkout_id + }) + }; + + barrier.wait(); + let outcome = coordinate(&repo_root, &db, &RuntimeBoundary::Flush) + .expect("the coordinator's first observation should succeed"); + + let storage_checkout_id = storage_thread + .join() + .expect("the storage thread should not panic"); + + assert_eq!( + outcome.worktree_id.0, storage_checkout_id, + "the coordinator and agent_trace_storage must converge on one checkout identity for the same physical checkout" + ); + + let on_disk = read_checkout_id(&resolve_git_dir(&repo_root).expect("git dir should resolve")) + .expect("reading the checkout-id file should succeed") + .expect("a checkout id must have been persisted"); + assert_eq!( + on_disk, storage_checkout_id, + "the on-disk checkout-id file must contain the converged identity" + ); + + cleanup(&repo_root); + cleanup(&state_root); +} + +#[test] +fn a_snapshot_failure_then_recovery_cycle_runs_through_the_public_api() { + let repo_root = unique_path("failure-recovery-repo"); + init_repo(&repo_root); + let db_path = repo_root.join("agent-trace.db"); + let db = RepositoryAgentTraceDb::new_at(&db_path).expect("repository DB should open"); + + let baseline = coordinate(&repo_root, &db, &RuntimeBoundary::Flush) + .expect("the baseline observation should materialize the worktree"); + let worktree_id = baseline.worktree_id.clone(); + + let git_dir = resolve_git_dir(&repo_root).expect("git dir should resolve"); + let tmp_index_dir = git_dir.join("sce").join("tmp"); + let _ = std::fs::remove_dir_all(&tmp_index_dir); + std::fs::write(&tmp_index_dir, b"not a directory").expect( + "planting a file where the snapshot service expects its temp-index directory should succeed", + ); + + let scope = ScopeId("scope-recovery".to_string()); + let failure = coordinate( + &repo_root, + &db, + &RuntimeBoundary::Start { + scope: scope.clone(), + event: EventId("evt-during-failure".to_string()), + actor_kind: ActorKind::ClaudeCode, + }, + ) + .expect_err("a Git snapshot failure against a materialized worktree should be reported"); + match failure { + CoordinateError::SnapshotFailure { + persisted_taint, .. + } => assert!( + persisted_taint, + "the existing worktree row must be durably tainted by the snapshot failure" + ), + other => panic!("expected CoordinateError::SnapshotFailure, got {other:?}"), + } + + { + let store = MutationTraceStore::new(&db); + let projection = store + .load_worktree(&worktree_id, None, None) + .expect("loading the worktree row should succeed") + .expect("the worktree row should still exist"); + assert!( + projection.worktree_state.tainted, + "the worktree row should be tainted after the snapshot failure" + ); + } + + std::fs::remove_file(&tmp_index_dir) + .expect("removing the planted file should let the snapshot service recreate its temp dir"); + + let recovered = coordinate(&repo_root, &db, &RuntimeBoundary::Flush) + .expect("the coordinator should recover from the taint and process the boundary"); + assert_eq!( + recovered.worktree_id, worktree_id, + "recovery must operate on the same worktree identity" + ); + + let store = MutationTraceStore::new(&db); + let projection = store + .load_worktree(&worktree_id, None, None) + .expect("loading the worktree row should succeed") + .expect("the worktree row should still exist"); + assert!( + !projection.worktree_state.tainted, + "taint recovery must clear the tainted flag before the triggering boundary is processed" + ); + + cleanup(&repo_root); +} diff --git a/context/cli/mutation-trace-protocol.md b/context/cli/mutation-trace-protocol.md index 5e7d96ba..14459949 100644 --- a/context/cli/mutation-trace-protocol.md +++ b/context/cli/mutation-trace-protocol.md @@ -214,8 +214,9 @@ loads a `ProtocolState`; `protocol.rs` only transitions already-known ones. The plan's file split anticipated three seams beyond `protocol.rs`. `store.rs`, `runtime/git_snapshot.rs`, and `coordinator.rs` (with its public `coordinate()` entrypoint) now all exist as real call sites (see -[`mutation-trace-runtime-coordinator.md`](mutation-trace-runtime-coordinator.md)); -only cross-module integration tests and harness/command wiring remain: +[`mutation-trace-runtime-coordinator.md`](mutation-trace-runtime-coordinator.md)), +with the runtime layer covered by cross-module integration tests; only +harness/command wiring remains: ```mermaid flowchart LR diff --git a/context/cli/mutation-trace-runtime-coordinator.md b/context/cli/mutation-trace-runtime-coordinator.md index 16aeb56a..5228e996 100644 --- a/context/cli/mutation-trace-runtime-coordinator.md +++ b/context/cli/mutation-trace-runtime-coordinator.md @@ -24,8 +24,9 @@ documented convention. The per-worktree runtime lock, the isolated Git snapshot service, the coordinator's internal protocol-integration pipeline, and the public, lock-wrapped `coordinate()` entrypoint that drives the lock and checkout -identity around that pipeline all exist. Only cross-module integration tests -and any harness/command wiring remain. +identity around that pipeline all exist, with cross-module integration tests +in `runtime/tests.rs` exercising the public API end to end. Only +harness/command wiring remains. - `cli/src/services/mutation_trace/runtime/worktree_lock.rs` — `WorktreeLock::acquire(git_dir: &Path, timeout: Duration) -> @@ -88,13 +89,17 @@ and any harness/command wiring remain. already-resolved one, never caller-supplied) and documents the `(ScopeId, EventId)` replay-identity contract a future harness adapter must uphold. The public `coordinate(repository_root, db, boundary) -> - Result` entrypoint owns the critical - section: resolve `git_dir` via `checkout::resolve_git_dir`, acquire the - `WorktreeLock` (bounded 10s, held for the whole call), resolve checkout - identity via `checkout::get_or_create_checkout_id` and wrap it as + Result` entrypoint takes an + already-resolved `RepositoryAgentTraceDb` from its caller — it never + resolves or opens the repository-scoped Agent Trace DB itself — and owns + the critical section: resolve `git_dir` via `checkout::resolve_git_dir`, + acquire the `WorktreeLock` (bounded 10s, held for the whole call), resolve + checkout identity via `checkout::get_or_create_checkout_id` and wrap it as `WorktreeId` — no caller-supplied `WorktreeId` or `Boundary` is ever accepted — then construct `GitSnapshotService` and delegate to the internal, - generic-over-`SnapshotCapture` pipeline. (`coordinate()` is a one-line + generic-over-`SnapshotCapture` pipeline. Identity flows + `repository_root → git_dir → WorktreeLock → checkout ID → WorktreeId`; + the `RepositoryAgentTraceDb` is not on that chain. (`coordinate()` is a one-line delegation to a private `coordinate_inner(.., on_lock_contention: impl FnOnce())` test seam; production passes a no-op closure.) A `WorktreeLock` acquisition failure (timeout or I/O) surfaces as @@ -205,16 +210,37 @@ the guard is dropped — the same invocation acquires the lock and returns `Ok`. Production `coordinate()` passes a no-op closure, so its code path, signature, and lock semantics are unchanged. +`runtime/tests.rs` is `runtime`'s own `#[cfg(test)] mod tests`, holding +cross-module integration tests that drive only the public `coordinate()` API +against real Git repositories (`git init`, `git worktree add`) and real +temp-file `RepositoryAgentTraceDb`s, following the same unique-temp-path +precedent: two linked worktrees of one repository (different `git_dir` → +different lock paths → different `WorktreeId`s) are proven independently +locked by holding one worktree's `WorktreeLock` across a synchronous +`coordinate()` call for the other and observing that call return `Ok` before +the held guard is dropped — a shared lock could not be acquired while the +guard is alive, and no wall-clock timing is used. The test opens one +repository-scoped DB path itself and hands a separate handle to each +`coordinate()` call (`coordinate()` does not resolve the DB), then asserts +both distinct worktree rows coexist in that one supplied DB, and that a tree +pinned by one worktree's coordinator resolves through the other's `GIT_DIR`. +A first-ever `agent_trace_storage` resolution and a `coordinate()` call on +the same checkout converge on one checkout identity (matching the on-disk +`checkout-id` file); and a full failure/recovery cycle — baseline call, a +snapshot-failing call that durably taints the worktree, then a recovery call +that clears the taint before processing its boundary — runs entirely through +the public entrypoint. + ## Status The per-worktree runtime lock, the isolated Git snapshot service, the coordinator's internal protocol-integration pipeline (above), and the public, lock-wrapped `coordinate()` entrypoint (resolving `git_dir`, acquiring `WorktreeLock`, resolving checkout identity, deriving `WorktreeId`, delegating -to the pipeline) are all implemented. Cross-module integration tests, a -`pub(crate)` re-export of `coordinate()` beyond `runtime`, and any -harness/command wiring remain future work tracked by the -`mutation-cursor-runtime-coordinator` plan's task stack and its follow-ups. +to the pipeline) are all implemented, and `runtime/tests.rs` covers the +public `coordinate()` API end to end (above). A `pub(crate)` re-export of +`coordinate()` beyond `runtime` and any harness/command wiring remain future +work tracked by the `mutation-cursor-runtime-coordinator` plan's follow-ups. See also: [`mutation-trace-protocol.md`](mutation-trace-protocol.md), [`mutation-trace-store.md`](mutation-trace-store.md), diff --git a/context/context-map.md b/context/context-map.md index 2d00b086..71858fd0 100644 --- a/context/context-map.md +++ b/context/context-map.md @@ -27,7 +27,7 @@ Feature/domain context: - `context/cli/mutation-trace-revision-refinement.md` (the Quint `revision: int` → Rust `WorktreeState::revision: u64` bounded-integer refinement: the private `next_revision` checked-arithmetic helper `commit`/`taint`/`abandon`/`recover` all route through instead of a raw `+ 1`, so a worktree at `revision: u64::MAX` is a guarded no-op/rejection rather than a silent wrap to `0`) - `context/cli/mutation-trace-quint-connect.md` (`#[cfg(test)]`-only Quint Connect model-based-testing harness in `cli/src/services/mutation_trace/mbt/` continuously checking `protocol.rs` against `spec/mutation_cursor.qnt`: the verification-only `mbtAction`/`MbtAction` record-payload transport excluded from comparison, the operation-identity-vs-`MbtStutter` distinction on guarded/no-op branches with its two deterministic regressions, finite ID mapping, the AC5 comparable-state field list, `randomPrepare` staying a single `step` branch, deterministic/generated (500×30, seed-reproducible) test coverage, and the two Nix checks — generic `checks.cli-tests` and dedicated `checks.mutation-trace-quint-connect` — that both require the pinned Quint binary plus the top-level `spec/` directory in `workspaceSrc`'s Nix fileset) - `context/cli/mutation-trace-store.md` (durable persistence for the mutation-cursor protocol in `cli/src/services/mutation_trace/store.rs`, built by the `mutation-cursor-store-persistence` plan: the one-directional `protocol.rs` -> `DurableTransition::between` (pure structural diff) -> `store.rs` (SQL translation) -> `RepositoryAgentTraceDb` boundary; migration `003_mutation_trace_protocol.sql`'s five tables; `AttemptState`/`external_taint` non-persistence; the 8-byte big-endian `BLOB` revision encoding plus explicit non-`Debug` enum codecs; the bounded hot-path `load_worktree` read vs. the cold-path `load_mutation_event`; `commit`'s single-`BEGIN IMMEDIATE` CAS batch via `TursoDb::execute_transactional_cas_batch`, distinguishing `Conflict`/retryable-transient/deterministic-`Err` outcomes; and the store's non-goals — no Git/filesystem I/O, no attribution/boundary-kind decisions, no retry-after-`Conflict` loop, no terminal-scope garbage collection) -- `context/cli/mutation-trace-runtime-coordinator.md` (imperative-shell runtime layer in `cli/src/services/mutation_trace/runtime/`, built by the `mutation-cursor-runtime-coordinator` plan: the per-worktree OS advisory lock, `runtime::worktree_lock::WorktreeLock::acquire(git_dir, timeout)` at `/sce/mutation-cursor.lock`, bounded `try_lock()` polling with a distinct matchable timeout error and RAII release, and its distinction from the separate checkout-identity-creation lock; and the isolated Git snapshot service, `runtime::git_snapshot::GitSnapshotService` (`capture_tree`/`pin_tree`/`diff_trees`), writing tree/blob objects into the repository's normal object database and pinning durable ones via a create-only, idempotent `refs/sce/mutation-cursor//` ref rather than a private object store; and `runtime::coordinator`'s protocol-integration pipeline (`RuntimeBoundary`/`CoordinateOutcome`/`CoordinateError`, the load → recover-if-needed → prepare/commit CAS-retry loop, and its own bounded snapshot-failure taint-retry loop) plus the public `coordinate(repository_root, db, boundary)` entrypoint that acquires the `WorktreeLock`, derives `WorktreeId` from `checkout::get_or_create_checkout_id`, and drives that pipeline under one held lock; cross-module integration tests, a `pub(crate)` re-export of `coordinate()` beyond `runtime`, and harness/command wiring remain future work) +- `context/cli/mutation-trace-runtime-coordinator.md` (imperative-shell runtime layer in `cli/src/services/mutation_trace/runtime/`, built by the `mutation-cursor-runtime-coordinator` plan: the per-worktree OS advisory lock, `runtime::worktree_lock::WorktreeLock::acquire(git_dir, timeout)` at `/sce/mutation-cursor.lock`, bounded `try_lock()` polling with a distinct matchable timeout error and RAII release, and its distinction from the separate checkout-identity-creation lock; and the isolated Git snapshot service, `runtime::git_snapshot::GitSnapshotService` (`capture_tree`/`pin_tree`/`diff_trees`), writing tree/blob objects into the repository's normal object database and pinning durable ones via a create-only, idempotent `refs/sce/mutation-cursor//` ref rather than a private object store; and `runtime::coordinator`'s protocol-integration pipeline (`RuntimeBoundary`/`CoordinateOutcome`/`CoordinateError`, the load → recover-if-needed → prepare/commit CAS-retry loop, and its own bounded snapshot-failure taint-retry loop) plus the public `coordinate(repository_root, db, boundary)` entrypoint that acquires the `WorktreeLock`, derives `WorktreeId` from `checkout::get_or_create_checkout_id`, and drives that pipeline under one held lock; `runtime/tests.rs` covers the public `coordinate()` API end to end against real linked worktrees and a real Agent Trace DB; a `pub(crate)` re-export of `coordinate()` beyond `runtime` and harness/command wiring remain future work) - `context/sce/cli-exit-code-contract.md` (stable class-based `sce` process exit-code contract in `cli/src/app.rs`, so automation can branch on failure category without parsing error text) - `context/sce/cli-error-code-taxonomy.md` (stable user-facing `SCE-ERR-*` diagnostic code classes rendered by `cli/src/app.rs`, complementing the numeric exit-code classes) - `context/sce/cli-stdout-stderr-contract.md` (implemented stream contract in `cli/src/app.rs`: command payloads on stdout only, redacted diagnostics on stderr) diff --git a/context/plans/mutation-cursor-runtime-coordinator.md b/context/plans/mutation-cursor-runtime-coordinator.md index 9fbed49b..72cbed53 100644 --- a/context/plans/mutation-cursor-runtime-coordinator.md +++ b/context/plans/mutation-cursor-runtime-coordinator.md @@ -131,10 +131,13 @@ independently of the mutation-cursor runtime. no second Git snapshot is taken for that invocation. - Validate: `runtime::coordinator::tests::cas_conflict_reloads_and_recomputes_without_a_second_snapshot` - [ ] AC9: Two coordinator invocations targeting the same worktree cannot - execute their critical sections concurrently; invocations on two different - worktrees, including two linked worktrees of the same repository, are not - serialized against each other and resolve to the same repository-scoped DB - with distinct `WorktreeId`s. + execute their critical sections concurrently. Invocations on two different + worktrees, including linked worktrees of the same repository, are not + serialized against each other, derive distinct `WorktreeId`s, and persist + independently into the same caller-supplied repository-scoped Agent Trace + DB. `coordinate()` resolves `git_dir`, checkout identity, and `WorktreeId` + from its `repository_root` argument; the `RepositoryAgentTraceDb` is + supplied by the caller, not resolved by the coordinator. - Validate: `runtime::worktree_lock::tests::*` (contention, distinct-path independence); `runtime::tests::linked_worktrees_have_independent_locks_and_worktree_ids` - [ ] AC10: A worktree whose durable state is `SnapshotFailure`-tainted or `needs_rebaseline` is recovered exactly once, using the same snapshot @@ -1070,13 +1073,14 @@ Persist this field in every plan; this is durable plan state, not chat state: accurate as edited during the original T05 sync. - Context synchronization: synced -- [ ] T06: `Add cross-module runtime integration tests` (status:todo) +- [x] T06: `Add cross-module runtime integration tests` (status:done) - Task ID: T06 - Scope: In — `runtime/tests.rs`: end-to-end `coordinate()` calls against - two real linked Git worktrees sharing one repository-scoped DB, proving - distinct `WorktreeId`s/locks and non-serialization across worktrees; an - end-to-end snapshot-failure-then-recovery cycle across two real - sequential `coordinate()` calls; a cross-caller assertion that + two real linked Git worktrees given handles to the same caller-supplied + repository-scoped DB, proving distinct `WorktreeId`s/locks and + non-serialization across worktrees; an end-to-end + snapshot-failure-then-recovery cycle across two real sequential + `coordinate()` calls; a cross-caller assertion that `agent_trace_storage`'s own checkout-identity resolution and the coordinator's observe the same checkout ID for one worktree. Out — any new production code; this task is test-only. @@ -1086,7 +1090,157 @@ Persist this field in every plan; this is durable plan state, not chat state: public `coordinate()` API, real `git worktree add`, and a real temp-file `RepositoryAgentTraceDb`. - Verify: `cargo test -p shared-context-engineering runtime::tests::` (via `./scripts/run-cli-cargo.sh test`) - - Context synchronization: pending + - Completed: 2026-08-30 + - Files changed: `cli/src/services/mutation_trace/runtime/tests.rs` (new), + `cli/src/services/mutation_trace/runtime/mod.rs` (added `#[cfg(test)] mod tests;`) + - Result: Added `runtime/tests.rs` as `runtime`'s own `#[cfg(test)] mod + tests` (declared in `runtime/mod.rs`), holding three cross-module, + end-to-end integration tests that drive only the public + `coordinator::coordinate()` API against real Git repositories and real + temp-file `RepositoryAgentTraceDb`s. No production code was added or + changed; the file carries no comments, matching the established + precedent for `runtime/` files authored entirely within their own task + (T02–T05). + + `linked_worktrees_have_independent_locks_and_worktree_ids` (AC9): inits a + main repo (with an empty commit so `git worktree add` is allowed), adds a + real linked worktree via `git worktree add`, opens one caller-supplied + repository-scoped DB path and hands a separate `RepositoryAgentTraceDb` + handle to each `coordinate()` call (`new_at` for the main, + `open_for_hooks_without_migrations_at` for the linked, matching the AC8 + two-writer convention — `coordinate()` never resolves the DB itself). + Asserts `resolve_git_dir` returns distinct worktree-specific git dirs + (hence distinct lock/identity paths); runs `coordinate(Flush)` on the + main worktree; then holds the main worktree's real `WorktreeLock` + (`WorktreeLock::acquire`) across a synchronous `coordinate(Flush)` call + for the linked worktree. That call returning `Ok` before the main lock + guard is dropped is the deterministic proof of independent lock paths: a + regression to one shared lock could not acquire that lock while `held` + is still alive and would instead return a lock-acquisition timeout. No + wall-clock assertion is used. Asserts the two + `CoordinateOutcome.worktree_id`s differ, that both distinct worktree rows + coexist in the same caller-supplied DB (`MutationTraceStore::load_worktree`), + and that a tree pinned by the main worktree's coordinator resolves + through the linked worktree's `GIT_DIR` + (`GitSnapshotService::new(&linked_root).diff_trees(...)`), proving the + shared object database / refs namespace. + + `agent_trace_storage_and_coordinator_observe_the_same_checkout_id` + (AC12): inits a repo with an `origin` remote (enough for + `agent_trace_storage` identity resolution; no network), then races — via + a two-party `Barrier` — a background-thread + `resolve_agent_trace_storage_at_state_root` call against a main-thread + `coordinate(Flush)` call, both on first-ever checkout-identity creation + for the same physical checkout. Asserts the storage resolution's + `checkout_id`, the coordinator outcome's `worktree_id.0`, and the on-disk + `checkout-id` file (`checkout::read_checkout_id`) are all the identical + value — proving T01's convergence guarantee holds across the module + boundary, not only inside `checkout::`'s own suite. (The two callers use + different databases; the only shared state under contention is the + `/sce/checkout-id` file.) + + `a_snapshot_failure_then_recovery_cycle_runs_through_the_public_api` + (end-to-end AC10/AC11): runs a baseline `coordinate(Flush)` to + materialize the worktree, then injects a real `GitSnapshotService` + failure by planting a regular file where `capture_tree` expects its + `/sce/tmp/` temp-index directory (so `capture_tree`'s + `create_dir_all` fails deterministically while repo detection, lock + acquisition, and checkout-identity resolution are all untouched). Asserts + the next `coordinate(Start{..})` returns + `CoordinateError::SnapshotFailure { persisted_taint: true, .. }` and that + the durable worktree row is `tainted`. Removes the planted file and runs + one more `coordinate(Flush)`, asserting it succeeds on the same worktree + identity and that the `tainted` flag was cleared — i.e. the coordinator + recovered from the taint before processing the triggering boundary. + + Reviewed assumption (recorded, not a scope deviation): the plan's Scope + line sketched the snapshot-failure injection loosely ("across two real + sequential `coordinate()` calls"); the implementation uses a baseline + call, a failing call, and a recovery call (three sequential public-API + invocations), and injects the failure by blocking the snapshot service's + temp-index directory rather than by corrupting `.git/HEAD` (HEAD removal + breaks Git repo detection for `resolve_git_dir`, which runs before the + snapshot step, so the failure would surface as `CoordinateError::Other` + rather than a `SnapshotFailure` taint path — verified empirically during + implementation). + + PR #244 review follow-up: two test/wording corrections, no production + change. First, `linked_worktrees_have_independent_locks_and_worktree_ids` + dropped its `Instant::now()` / `elapsed() < 5s` assertion (and the unused + `Instant` import): the lock-independence proof is the deterministic + ordering — the main worktree's `WorktreeLock` stays held across the whole + synchronous linked `coordinate()` call, and that call returning `Ok` + before `held` is dropped is only possible if the linked worktree + acquires a different lock path; a shared lock would return a + lock-acquisition timeout instead. The wall-clock check added only CI + timing sensitivity and proved nothing the ordering did not. Second, AC9 + and the surrounding wording (this record, the Scope line, and + `context/cli/mutation-trace-runtime-coordinator.md`) were corrected to + stop implying `coordinate()` resolves the repository-scoped Agent Trace + DB. It does not: `coordinate(repository_root, db, boundary)` derives + `git_dir`, checkout identity, and `WorktreeId` from `repository_root`, + but the `RepositoryAgentTraceDb` is supplied by the caller. The + linked-worktree test hands a separate handle to the one caller-opened DB + path to each call and proves the two distinct worktree rows coexist in + that supplied DB — not that each invocation independently resolves it. + - Verify (actual): `./scripts/run-cli-cargo.sh test --manifest-path + cli/Cargo.toml runtime::tests::` → 3/3 new tests passed (the filter also + re-runs 5 unrelated `parse::command_runtime::tests` that share the + `runtime::tests` substring; all passed). Additionally ran + `./scripts/run-cli-cargo.sh test --manifest-path cli/Cargo.toml + services::mutation_trace::runtime::` → 35/35 passed (no regression in the + T02/T03/T04/T05 `worktree_lock`, `git_snapshot`, and `coordinator` + tests); `services::checkout::` → 5/5; `services::agent_trace_storage::` + → 14/14 (the two cross-called modules, unchanged, still green). + `./scripts/run-cli-cargo.sh clippy --manifest-path cli/Cargo.toml + --all-targets -- -D warnings` → clean (pedantic/warnings denied + workspace-wide, per Constraints); `./scripts/run-cli-cargo.sh fmt + --manifest-path cli/Cargo.toml -- --check` → clean (after one `cargo fmt` + pass on the new file). + + PR #244 review follow-up re-verification: after removing the + elapsed-time assertion and correcting the DB-ownership wording, + `./scripts/run-cli-cargo.sh test --manifest-path cli/Cargo.toml + services::mutation_trace::runtime::tests::` → 3/3 passed; + `services::mutation_trace::runtime::` → 35/35 passed; the full + `./scripts/run-cli-cargo.sh test --manifest-path cli/Cargo.toml` suite → + 824/824 passed / 0 failed; `clippy --all-targets -- -D warnings` → clean; + `fmt -- --check` → clean (after one `cargo fmt` pass re-wrapping the + edited `open_for_hooks_without_migrations_at` line). + - Context impact: `domain` — this task adds only tests; it introduces no + new architectural element, public interface, behavior contract, or + terminology. The mutation-cursor-runtime-coordinator domain file + (`context/cli/mutation-trace-runtime-coordinator.md`) already carries a + "Testing boundary" section; task context synchronization should extend + its `runtime/tests.rs` bullet to name the three landed cross-module + integration tests (linked-worktree independence, cross-caller + checkout-identity convergence, public-API failure/recovery cycle) and + flip any "T06 not yet landed" / "runtime/tests.rs planned" framing to + present-tense. Per the plan's own Context sync notes, T06 is the trigger + for finalizing the `context/context-map.md` + `coordinator.rs`/`git_snapshot.rs` seam annotations (drop the + "not-yet-created" wording now that the full `runtime/` module including + its integration suite exists) and for a last pass over + `context/cli/mutation-trace-protocol.md`, + `context/cli/mutation-trace-store.md`, and `context/overview.md` to + confirm no stale "future work" framing about the runtime layer remains. + The five root context files must still be verified per the mandatory + pass. No qualifying system-wide architecture decision was introduced by + this task. + + PR #244 review follow-up: the synced context edits were corrected to + stop implying `coordinate()` resolves the repository-scoped Agent Trace + DB. `context/cli/mutation-trace-runtime-coordinator.md`'s `coordinate()` + bullet now states it takes an already-resolved `RepositoryAgentTraceDb` + from its caller and never resolves/opens one (identity chain: + `repository_root → git_dir → WorktreeLock → checkout ID → WorktreeId`, + with the DB not on that chain), and its "Testing boundary" paragraph + describes the linked-worktree proof as the held-lock ordering (no + wall-clock timing) with both worktree rows coexisting in the one + caller-supplied DB. AC9's own wording in this plan was corrected to + match. No other context file needed a change; the five root files remain + accurate as verified in the original sync. + - Context synchronization: synced ## Design decisions @@ -2179,12 +2333,15 @@ Cargo/dependency changes: none. on one complete, valid ID. Prove: AC13 — the canonical path is never observable as partially written, and orphaned temp files are inert. - **Linked-worktree tests** (`runtime/tests.rs`, real `git worktree add`): - two linked worktrees resolve distinct `checkout_id`/`WorktreeId` and - distinct runtime-lock paths, concurrent `coordinate()` calls on the two - worktrees do not block each other, both share one repository-scoped DB, - and a tree pinned from one worktree resolves correctly when queried via - the other worktree's `GIT_DIR`. Prove: AC9 end to end through the public - API. + two linked worktrees derive distinct `checkout_id`/`WorktreeId` and + distinct runtime-lock paths — proven by holding one worktree's + `WorktreeLock` across a synchronous `coordinate()` call for the other and + observing that call succeed before the held guard is dropped, not by any + wall-clock timing — their distinct worktree rows coexist in the one + caller-supplied repository-scoped DB handed to both calls (`coordinate()` + does not resolve the DB), and a tree pinned from one worktree resolves + correctly when queried via the other worktree's `GIT_DIR`. Prove: AC9 end + to end through the public API. - **Cross-caller checkout-identity test** (`runtime/tests.rs`): a direct `agent_trace_storage` resolution and a `coordinate()` call against the same `repository_root`, run concurrently on first-ever resolution, observe From 667614051467968082ae1eaee7cc633bad719abf Mon Sep 17 00:00:00 2001 From: David Abram Date: Sun, 30 Aug 2026 09:42:55 +0200 Subject: [PATCH 8/8] context: Record mutation cursor coordinator validation Mark the mutation-cursor runtime coordinator plan complete after the coordinator, snapshot, locking, CAS-retry, taint-recovery, and checkout-ID behaviors passed the recorded validation suite. Preserve the validation evidence and document the intentionally deferred ref reconciliation and external-taint work. Plan: mutation-cursor-runtime-coordinator Tasks: AC1-AC13 Co-authored-by: SCE --- .../mutation-cursor-runtime-coordinator.md | 110 +++++++++++++++--- 1 file changed, 97 insertions(+), 13 deletions(-) diff --git a/context/plans/mutation-cursor-runtime-coordinator.md b/context/plans/mutation-cursor-runtime-coordinator.md index 72cbed53..e359cb0f 100644 --- a/context/plans/mutation-cursor-runtime-coordinator.md +++ b/context/plans/mutation-cursor-runtime-coordinator.md @@ -90,26 +90,26 @@ independently of the mutation-cursor runtime. ## Acceptance criteria -- [ ] AC1: A worktree observed for the first time establishes the currently +- [x] AC1: A worktree observed for the first time establishes the currently observed tree as its cursor baseline on its first boundary and emits no mutation evidence for filesystem changes that predate that observation. - Validate: `runtime::coordinator::tests::first_observation_establishes_baseline_without_evidence` via `nix build .#checks..cli-tests` -- [ ] AC2: An edit made between a `Start` and a subsequent `Advance` on the +- [x] AC2: An edit made between a `Start` and a subsequent `Advance` on the same scope commits as exactly one `AiExclusive` mutation event whose `before_tree`/`after_tree` match the Start baseline and the post-edit snapshot. - Validate: `runtime::coordinator::tests::exclusive_edit_between_start_and_advance_commits_one_event` -- [ ] AC3: Re-processing the identical `(scope, event)` boundary a second time +- [x] AC3: Re-processing the identical `(scope, event)` boundary a second time produces no duplicated mutation evidence. - Validate: `runtime::coordinator::tests::replaying_the_same_scope_event_key_does_not_duplicate_evidence` -- [ ] AC4: A mutation made just before `Close` is still attributed using the +- [x] AC4: A mutation made just before `Close` is still attributed using the scope set as it existed immediately before `Close`'s own scope transition. - Validate: `runtime::coordinator::tests::close_boundary_attributes_using_pre_close_scope_set` -- [ ] AC5: Two concurrently active scopes on one worktree yield `AiContended` +- [x] AC5: Two concurrently active scopes on one worktree yield `AiContended` attribution for a subsequent `Advance`, independent of whether the two scopes share an `ActorKind`. - Validate: `runtime::coordinator::tests::contended_scopes_yield_ai_contended_same_and_different_actor` -- [ ] AC6: Capturing a Git snapshot never mutates the caller's real index, +- [x] AC6: Capturing a Git snapshot never mutates the caller's real index, staged changes, or working tree, correctly reflects staged, unstaged, untracked, and deleted state, excludes ignored files, and — on an unborn `HEAD` — starts from an explicitly initialized, genuinely valid empty @@ -117,7 +117,7 @@ independently of the mutation-cursor runtime. file the coordinator merely assumes Git will treat as empty, and produces a correct tree even when the unborn repository has no files at all. - Validate: `runtime::git_snapshot::tests::*` (index preservation, ignored files, deletion, unborn HEAD with a file, unborn HEAD with no files) -- [ ] AC7: A snapshot's `TreeId` remains resolvable through `diff_trees` +- [x] AC7: A snapshot's `TreeId` remains resolvable through `diff_trees` after the process that captured it has exited, its temporary index file no longer exists, **and after a `git gc --prune=now` / `git prune --expire=now` pass has run against the repository** — because it is @@ -125,12 +125,12 @@ independently of the mutation-cursor runtime. not merely present in an isolated object store Git's own reachability analysis knows nothing about. - Validate: `runtime::git_snapshot::tests::snapshot_survives_a_fresh_process_and_temp_index_deletion`, `runtime::git_snapshot::tests::pinned_snapshot_survives_git_gc_prune_now`, `runtime::git_snapshot::tests::pinned_snapshot_survives_git_prune_expire_now` -- [ ] AC8: When two invocations race to commit from the same durable +- [x] AC8: When two invocations race to commit from the same durable revision, exactly one succeeds, the other reloads durable state and recomputes its transition using its own originally captured snapshot, and no second Git snapshot is taken for that invocation. - Validate: `runtime::coordinator::tests::cas_conflict_reloads_and_recomputes_without_a_second_snapshot` -- [ ] AC9: Two coordinator invocations targeting the same worktree cannot +- [x] AC9: Two coordinator invocations targeting the same worktree cannot execute their critical sections concurrently. Invocations on two different worktrees, including linked worktrees of the same repository, are not serialized against each other, derive distinct `WorktreeId`s, and persist @@ -139,13 +139,13 @@ independently of the mutation-cursor runtime. from its `repository_root` argument; the `RepositoryAgentTraceDb` is supplied by the caller, not resolved by the coordinator. - Validate: `runtime::worktree_lock::tests::*` (contention, distinct-path independence); `runtime::tests::linked_worktrees_have_independent_locks_and_worktree_ids` -- [ ] AC10: A worktree whose durable state is `SnapshotFailure`-tainted or +- [x] AC10: A worktree whose durable state is `SnapshotFailure`-tainted or `needs_rebaseline` is recovered exactly once, using the same snapshot captured for the triggering boundary, before that boundary is processed; `needs_rebaseline` recovery preserves live scopes, taint recovery abandons them. - Validate: `runtime::coordinator::tests::recovers_from_needs_rebaseline_preserving_live_scopes`, `runtime::coordinator::tests::recovers_from_snapshot_failure_taint_abandoning_live_scopes` -- [ ] AC11: A Git snapshot failure against an already-materialized worktree +- [x] AC11: A Git snapshot failure against an already-materialized worktree durably persists a `SnapshotFailure` taint via `protocol::taint`, retried under the same bounded semantic CAS-retry policy the main coordinator loop uses when another committer's transition wins the race; `persisted_taint` @@ -158,13 +158,13 @@ independently of the mutation-cursor runtime. caller materializes concurrently, while this invocation's own Git snapshot is still being captured, is still found and correctly tainted. - Validate: `runtime::coordinator::tests::snapshot_failure_taints_an_existing_worktree`, `runtime::coordinator::tests::snapshot_failure_taint_survives_a_losing_cas_and_commits_on_retry`, `runtime::coordinator::tests::snapshot_failure_taint_reports_not_persisted_after_retries_are_exhausted`, `runtime::coordinator::tests::snapshot_failure_before_any_baseline_makes_no_durable_write`, `runtime::coordinator::tests::snapshot_failure_taints_a_worktree_materialized_concurrently_during_capture` -- [ ] AC12: All concurrent first-time callers of +- [x] AC12: All concurrent first-time callers of `checkout::get_or_create_checkout_id` for one physical checkout — whether through the coordinator, `agent_trace_storage`, or any other caller — converge on exactly one checkout ID, and the on-disk `checkout-id` file ends up containing that same value. - Validate: `checkout::tests::concurrent_first_time_callers_converge_on_one_checkout_id`, `runtime::tests::agent_trace_storage_and_coordinator_observe_the_same_checkout_id` -- [ ] AC13: For cooperating SCE processes, the canonical `checkout-id` path +- [x] AC13: For cooperating SCE processes, the canonical `checkout-id` path is at every observable point either absent or contains exactly one complete, valid checkout ID — never a partially written or truncated value — including immediately after a process is interrupted between @@ -2643,3 +2643,87 @@ is `Some`, i.e. the triggering boundary produced real evidence — feeds its row, alongside the mutation event's `attribution`/`boundary`/scope information. It also owns deciding how (or whether) multiple harnesses share one coordinator invocation path. + +## Validation Report + +**Status:** validated +**Date:** 2026-08-30 + +### Commands run + +- `nix flake check` -> exit 0 (all checks passed: `cli-tests` 824 tests, + `cli-clippy` pedantic/warnings-denied, `cli-fmt`, + `mutation-trace-quint-connect`, `pkl-generated`, `workflow-actionlint`, + and the flatpak/cargo-source parity checks) +- `nix run .#pkl-check-generated` -> exit 0 (ephemeral Pkl generation passed: + 141 files, inventory sha256 bf5db9c962cc9ce2776b4fc218dcbd8787fa7567744a5e2faff1fc9f9212a003 — no generated-config drift) +- `./scripts/run-cli-cargo.sh test --manifest-path cli/Cargo.toml runtime::` -> + exit 0 (40/40 passed, covering every `runtime::coordinator`, + `runtime::git_snapshot`, `runtime::worktree_lock`, and `runtime::tests` case) +- `./scripts/run-cli-cargo.sh test --manifest-path cli/Cargo.toml checkout::` -> + exit 0 (5/5 passed) + +### Success-criteria verification + +- [x] AC1: A first-observed worktree establishes the observed tree as its + cursor baseline and emits no evidence for pre-observation changes -> + `runtime::coordinator::tests::first_observation_establishes_baseline_without_evidence` passed +- [x] AC2: An edit between `Start` and `Advance` commits exactly one + `AiExclusive` event with matching before/after trees -> + `runtime::coordinator::tests::exclusive_edit_between_start_and_advance_commits_one_event` passed +- [x] AC3: Re-processing the identical `(scope, event)` boundary duplicates no + evidence -> + `runtime::coordinator::tests::replaying_the_same_scope_event_key_does_not_duplicate_evidence` passed +- [x] AC4: A mutation just before `Close` is attributed using the pre-`Close` + scope set -> + `runtime::coordinator::tests::close_boundary_attributes_using_pre_close_scope_set` passed +- [x] AC5: Two concurrent scopes yield `AiContended` regardless of shared + `ActorKind` -> + `runtime::coordinator::tests::contended_scopes_yield_ai_contended_same_and_different_actor` passed (exercises same- and different-actor pairs) +- [x] AC6: Snapshot capture never mutates the real index/staged/working state, + reflects staged/unstaged/untracked/deleted state, excludes ignored files, and + initializes an explicit valid empty index on unborn `HEAD` -> + `runtime::git_snapshot::tests::{capture_preserves_real_index_and_working_tree_state, capture_excludes_ignored_files, capture_reflects_deletion_of_a_committed_file, capture_on_unborn_head_with_a_file_produces_a_valid_tree, capture_on_unborn_head_with_no_files_produces_an_empty_tree}` passed +- [x] AC7: A pinned `TreeId` stays resolvable after process exit, temp-index + deletion, and `git gc --prune=now` / `git prune --expire=now` -> + `runtime::git_snapshot::tests::{snapshot_survives_a_fresh_process_and_temp_index_deletion, pinned_snapshot_survives_git_gc_prune_now, pinned_snapshot_survives_git_prune_expire_now}` passed +- [x] AC8: Racing invocations — exactly one commits, the other reloads and + recomputes with its own snapshot, no second snapshot taken -> + `runtime::coordinator::tests::cas_conflict_reloads_and_recomputes_without_a_second_snapshot` passed +- [x] AC9: Same-worktree invocations serialize; different (incl. linked) + worktrees do not, derive distinct `WorktreeId`s, persist independently into + the caller-supplied DB -> + `runtime::worktree_lock::tests::{a_second_acquirer_blocks_until_the_first_releases, acquire_times_out_with_a_distinct_matchable_error_when_still_held, distinct_worktree_paths_do_not_contend}` and `runtime::tests::linked_worktrees_have_independent_locks_and_worktree_ids` passed +- [x] AC10: A tainted / `needs_rebaseline` worktree is recovered exactly once + with the triggering boundary's snapshot; rebaseline preserves live scopes, + taint abandons them -> + `runtime::coordinator::tests::{recovers_from_needs_rebaseline_preserving_live_scopes, recovers_from_snapshot_failure_taint_abandoning_live_scopes}` passed +- [x] AC11: A snapshot failure against an existing worktree durably persists a + taint under bounded CAS retry (`persisted_taint` true only when `Applied`); + no prior row makes no durable write; the taint decision reads fresh state + after the failure -> + `runtime::coordinator::tests::{snapshot_failure_taints_an_existing_worktree, snapshot_failure_taint_survives_a_losing_cas_and_commits_on_retry, snapshot_failure_taint_reports_not_persisted_after_retries_are_exhausted, snapshot_failure_before_any_baseline_makes_no_durable_write, snapshot_failure_taints_a_worktree_materialized_concurrently_during_capture}` passed +- [x] AC12: All concurrent first-time `get_or_create_checkout_id` callers + converge on one checkout ID matching the on-disk file -> + `checkout::tests::concurrent_first_time_callers_converge_on_one_checkout_id` and `runtime::tests::agent_trace_storage_and_coordinator_observe_the_same_checkout_id` passed +- [x] AC13: The canonical `checkout-id` path is always absent or holds one + complete valid ID, incl. after interruption before rename; an orphaned temp + file never blocks convergence -> + `checkout::tests::{interruption_before_rename_leaves_the_canonical_path_absent, completed_rename_leaves_the_canonical_path_with_a_complete_id, an_orphaned_temp_file_does_not_block_convergence_on_the_canonical_id}` passed + +### Failed checks and follow-ups + +- None. No leftover debug flags, temporary files, intermediate artifacts, or + local scaffolding found in the production runtime/checkout sources; working + tree is clean. + +### Residual risks + +- Ref-pin storage growth is unbounded until the deferred + `refs/sce/mutation-cursor/**` reconciliation pass ships, and the DB-unavailable + case leaves no filesystem external-taint marker. Both are documented, + intentionally deferred, and sequenced (`mutation-cursor-external-taint`, + `mutation-cursor-ref-reconciliation`) as required runtime-completion work + before any harness adapter becomes a production consumer of `coordinate()`. +- The coordinator is standalone and not wired into any hook, CLI command, or + `diff_traces` — no production caller exercises it yet.