Repository navigation
fix: harness uplift follow-ups (provider_started, graph lifecycle events, completion recording) - #351
Conversation
Add tests asserting that a summarizer which reached its provider before failing without usage still marks provider_started, while one rejected before dispatch leaves it false. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…out usage Co-authored-by: Medulla <medulla@tinyhumans.ai>
The graph driver now announces turn and message lifecycle events like the direct loop, so the parity filter that excluded them is gone and the expected kinds are matched in full. The test pinning the old direct-loop-only gap was removed, with the exact event shape covered by the new graph_lifecycle_events test. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The agent loop's compilation and runtime logic now live in dedicated modules, with lifecycle and phase handling moved into the harness crate. This separates graph construction from execution so each can evolve independently. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The graph driver now closes the current turn and flushes any messages appended by the after-agent middleware, matching the direct loop so transcript events are announced on every exit path. The settle node receives the run context so it can emit those events. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…iver Co-authored-by: Medulla <medulla@tinyhumans.ai>
…h the completion router Co-authored-by: Medulla <medulla@tinyhumans.ai>
The subagent README now spells out which terminal outcomes are recorded by the completion router, covering cancellations, executor errors, and the pause/resume case, and notes that a re-run under the same task id does not add a second record. The harness terminal-outcome doc records that the graph loop driver now publishes the turn and message lifecycle events with the same semantics as the direct loop, and explains how `provider_started` is set for summarizer calls. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…atch tracking The graph agent loop now documents how it announces TurnStarted, TurnCompleted and MessageAppended at the same points as the direct loop, including that the seed transcript is never announced and nested tool calls never reach the transcript. The summarization README gains a section on dispatch tracking, explaining why a summarizer that fails without usage after dispatching still reports provider_started while a pre-dispatch rejection does not. Remaining edits are rustfmt reflows and README wording cleanups. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…nd completion recording Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add integration tests covering graph lifecycle events, verifying that the expected events are emitted as a graph runs through its lifecycle. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Interrupted nodes discard their state and re-run from entry on resume, so the appends they had already announced are now retracted with MessageRetracted, keeping event mirrors consistent with the transcript that is actually kept. The built-in model-backed summarizers also mark dispatch before calling their model, so provider_started is reported correctly for host-defined summarizers that cannot call mark_dispatched themselves. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Tiny Sweeper review
Last completed reportTiny Sweeper reviewTiny Sweeper reviewed this change across 6 lane(s) and found 6 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below. State: Ready for maintainer review Review snapshot
Completeness: Complete What changedThe review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below. FeaturesNone identified with supported citations. TestsNo supported feature-to-test mapping was produced. Test execution is not inferred. Findings
Resolved this pass
Before mergeNone. How this fits togetherflowchart LR
n0["RunContext"]:::impacted
n1["LoopState"]:::impacted
n2["compile_loop"]:::impacted
n3["tools_node"]:::impacted
n4["model_node"]:::impacted
n5["PLAN"]:::impacted
n2 -->|uses| n1
n2 -->|calls| n3
n2 -->|calls| n4
n2 -->|uses| n5
n3 -->|uses| n0
n3 -->|uses| n1
n3 -->|uses| n5
n4 -->|uses| n0
n4 -->|uses| n1
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Warning Review limit reached
This review includes 26 billable files and costs up to $6.50. Or wait 22 minutes for your next included review. View limit detailsLimit details: You’ve used all 2 included reviews currently available. Review configuration: ⚙️ Run configuration
📒 Files selected for processing (26)
Comment |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 21b9dcb7b1
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
tinysweeper found nothing blocking. Approving.
$0.0653 · 1,144,532 in / 56,232 out · 179,848 cached (16%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0398 · 606,282 in / 34,421 out · 113,697 cached (19%) · gpt-5.6-luna, glm-5.3-flash
security: $0.0246 · 431,501 in / 17,296 out · 64,039 cached (15%) · gpt-5.6-luna
tests: $0.0004 · 54,678 in / 1,463 out · 1,920 cached (4%) · glm-5.3-flash
description: $0.0002 · 25,189 in / 133 out · 64 cached (0%) · glm-5.3-flash
Add integration tests covering graph lifecycle events, verifying that the expected events are emitted as a graph runs through its lifecycle. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The agent loop runtime now emits graph lifecycle events as a run progresses, so subscribers can observe start, step, and completion transitions. Integration tests cover the new event stream. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
… interrupt check, close interrupted model turn Co-authored-by: Medulla <medulla@tinyhumans.ai>
There was a problem hiding this comment.
tinysweeper found nothing blocking. Approving.
$0.0201 · 387,726 in / 40,439 out · 37,109 cached (10%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0111 · 183,882 in / 21,016 out · 17,931 cached (10%) · gpt-5.6-luna, glm-5.3-flash
security: $0.0082 · 118,723 in / 13,138 out · 16,106 cached (14%) · gpt-5.6-luna
tests: $0.0003 · 27,929 in / 2,084 out · 1,536 cached (5%) · glm-5.3-flash
description: $0.0002 · 27,662 in / 1,486 out · 1,408 cached (5%) · glm-5.3-flash
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ef0aaa6187
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…mpletion test double Co-authored-by: Medulla <medulla@tinyhumans.ai>
The graph agent loop now reports lifecycle events through the harness so subagent completion and phase transitions are observable. This lets orchestration track subagent progress and integration tests assert on the event stream. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Moved the agent loop runtime helpers out of the main module into their own file to keep the loop logic easier to follow. No behaviour change. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
There was a problem hiding this comment.
tinysweeper found nothing blocking. Approving.
$0.0105 · 239,537 in / 21,391 out · 12,243 cached (5%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0050 · 87,412 in / 7,294 out · 6,679 cached (8%) · gpt-5.6-luna, glm-5.3-flash
security: $0.0048 · 66,210 in / 7,340 out · 5,372 cached (8%) · gpt-5.6-luna
tests: $0.0003 · 28,211 in / 2,250 out · 64 cached (0%) · glm-5.3-flash
description: $0.0003 · 27,944 in / 2,021 out · 64 cached (0%) · glm-5.3-flash
…result for cancelled completions Co-authored-by: Medulla <medulla@tinyhumans.ai>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 74b554511e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…ails Co-authored-by: Medulla <medulla@tinyhumans.ai>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 26e518d96b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| pub(crate) fn adopt(&mut self, completed_turns: u32, open_from: Option<usize>) { | ||
| self.turn = self.turn.max(completed_turns); |
There was a problem hiding this comment.
Carry interrupted attempts into resumed turn numbering
When a compiled graph resumes in a fresh LoopRuntime after after_model requests an interrupt, the interrupted activation has already emitted and closed turn 1, but its local loop_state.turn += 1 is discarded, leaving the checkpoint at turn 0. This max therefore restores 0, and the resumed model emits another TurnStarted { turn: 1 } / TurnCompleted { turn: 1 }, causing consumers keyed by turn number to conflate separate attempts. Fresh evidence beyond the earlier interrupted-turn thread is that adopt now derives its counter solely from that stale checkpoint value, while the new resume test checks message indices but not turn-number uniqueness; persist the lifecycle counter across the interrupt and add a uniqueness assertion.
AGENTS.md reference: AGENTS.md:L70-L78
Useful? React with 👍 / 👎.
Summary
Three follow-ups from the harness uplift (#329, #347 D10), each with a failing test first.
provider_startedon summarizer failures with no usage. Summarizer calls bypass the run context's dispatch marker, so a compaction that failed without usage leftTerminalOutcome::provider_startedfalse.ContextCompressionMiddlewarenow scopes each summarization with a task-local dispatch tracker (summarization::dispatch);ModelSummarizerandTaskStateSummarizermark it right before calling their model, and the middleware sets the run-wide flag. A pre-dispatch rejection (validation) leaves it false. Only the run-wide flag is set, so timeout-phase classification is unchanged.GraphLoopDriver,LoopIterandcompile_loopnow emitTurnStarted,TurnCompletedandMessageAppendedat the same points as the direct loop, via newphases::lifecycle_*wrappers over the harnessTurnTracker. Input is never announced (the tracker is seeded on first node entry, so checkpoint resume in a fresh runtime does not re-announce history), nested tool calls never reach the transcript so they emit nothing, and an interrupting node retracts its announced appends because it re-runs on resume. The test that pinned the gap inloop_as_graph.rsis flipped into parity tests (graph_lifecycle_events.rs), and the parity helper no longer filters lifecycle kinds.SubagentDriver). With a router and a notify mode: cancelled children are recorded asCancelled(also a cancel after planning but before launch), executor errors (Execution, orTransientafter retries) asFailedwhilerunstill returns the error. Host seam faults (task id mismatch, persistence, missing capability) are not recorded. Paused children are not recorded (not terminal); the resume that finishes the task records once, and an error while resuming a paused task is not recorded because its pause is still durable. Behaviour without a router or notify mode is unchanged.Notes
LoopIter/compile_loopdo not close an open turn on node error.docs/modules/harness/terminal-outcome.md,subagent/README.md,summarization/README.md, graphagent_loopmodule docs.Verification
cargo test --workspace(only the knownvalidate_repo_root_rejects_non_repofails), clippy-D warnings,cargo fmt --checkclean;cargo docwarning set unchanged in the touched items. An opus code review found no Critical issues; its Important findings (interrupt duplicates, seeding in non-plan nodes,TaskStateSummarizermarking) are fixed with tests.Co-authored-by: Medulla medulla@tinyhumans.ai