Skip to content

fix: harness uplift follow-ups (provider_started, graph lifecycle events, completion recording) - #351

Merged
senamakel merged 21 commits into
mainfrom
harness-uplift/uplift-followups
Oct 9, 2026
Merged

senamakel merged 21 commits into
mainfrom
harness-uplift/uplift-followups

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

Three follow-ups from the harness uplift (#329, #347 D10), each with a failing test first.

  1. provider_started on summarizer failures with no usage. Summarizer calls bypass the run context's dispatch marker, so a compaction that failed without usage left TerminalOutcome::provider_started false. ContextCompressionMiddleware now scopes each summarization with a task-local dispatch tracker (summarization::dispatch); ModelSummarizer and TaskStateSummarizer mark it right before calling their model, and the middleware sets the run-wide flag. A pre-dispatch rejection (validation) leaves it false. Only the run-wide flag is set, so timeout-phase classification is unchanged.
  2. Graph-driver lifecycle events. GraphLoopDriver, LoopIter and compile_loop now emit TurnStarted, TurnCompleted and MessageAppended at the same points as the direct loop, via new phases::lifecycle_* wrappers over the harness TurnTracker. Input is never announced (the tracker is seeded on first node entry, so checkpoint resume in a fresh runtime does not re-announce history), nested tool calls never reach the transcript so they emit nothing, and an interrupting node retracts its announced appends because it re-runs on resume. The test that pinned the gap in loop_as_graph.rs is flipped into parity tests (graph_lifecycle_events.rs), and the parity helper no longer filters lifecycle kinds.
  3. Completion recording (SubagentDriver). With a router and a notify mode: cancelled children are recorded as Cancelled (also a cancel after planning but before launch), executor errors (Execution, or Transient after retries) as Failed while run still returns the error. Host seam faults (task id mismatch, persistence, missing capability) are not recorded. Paused children are not recorded (not terminal); the resume that finishes the task records once, and an error while resuming a paused task is not recorded because its pause is still durable. Behaviour without a router or notify mode is unchanged.

Notes

  • The router keeps the first record per task id, so a task re-run under the same id after a recorded failure or cancellation adds no second completion. Documented in the subagent README.
  • Known graph/direct differences (documented): output-retry prompt ordering, and LoopIter/compile_loop do not close an open turn on node error.
  • Docs updated: docs/modules/harness/terminal-outcome.md, subagent/README.md, summarization/README.md, graph agent_loop module docs.

Verification

cargo test --workspace (only the known validate_repo_root_rejects_non_repo fails), clippy -D warnings, cargo fmt --check clean; cargo doc warning set unchanged in the touched items. An opus code review found no Critical issues; its Important findings (interrupt duplicates, seeding in non-plan nodes, TaskStateSummarizer marking) are fixed with tests.

Co-authored-by: Medulla medulla@tinyhumans.ai

senamakel and others added 13 commits October 9, 2026 08:42
Add tests asserting that a summarizer which reached its provider before
failing without usage still marks provider_started, while one rejected
before dispatch leaves it false.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…out usage

Co-authored-by: Medulla <medulla@tinyhumans.ai>
The graph driver now announces turn and message lifecycle events like the
direct loop, so the parity filter that excluded them is gone and the
expected kinds are matched in full. The test pinning the old direct-loop-only
gap was removed, with the exact event shape covered by the new
graph_lifecycle_events test.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The agent loop's compilation and runtime logic now live in dedicated
modules, with lifecycle and phase handling moved into the harness crate.
This separates graph construction from execution so each can evolve
independently.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The graph driver now closes the current turn and flushes any messages
appended by the after-agent middleware, matching the direct loop so
transcript events are announced on every exit path. The settle node
receives the run context so it can emit those events.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…iver

Co-authored-by: Medulla <medulla@tinyhumans.ai>
…h the completion router

Co-authored-by: Medulla <medulla@tinyhumans.ai>
The subagent README now spells out which terminal outcomes are recorded by the completion router, covering cancellations, executor errors, and the pause/resume case, and notes that a re-run under the same task id does not add a second record. The harness terminal-outcome doc records that the graph loop driver now publishes the turn and message lifecycle events with the same semantics as the direct loop, and explains how `provider_started` is set for summarizer calls.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…atch tracking

The graph agent loop now documents how it announces TurnStarted, TurnCompleted and MessageAppended at the same points as the direct loop, including that the seed transcript is never announced and nested tool calls never reach the transcript. The summarization README gains a section on dispatch tracking, explaining why a summarizer that fails without usage after dispatching still reports provider_started while a pre-dispatch rejection does not. Remaining edits are rustfmt reflows and README wording cleanups.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…nd completion recording

Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add integration tests covering graph lifecycle events, verifying that
the expected events are emitted as a graph runs through its lifecycle.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Interrupted nodes discard their state and re-run from entry on resume, so the
appends they had already announced are now retracted with MessageRetracted,
keeping event mirrors consistent with the transcript that is actually kept.
The built-in model-backed summarizers also mark dispatch before calling their
model, so provider_started is reported correctly for host-defined summarizers
that cannot call mark_dispatched themselves.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@tinysweeper

tinysweeper Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

⚠️ Review failed for 74b554511ebe. forge: GitHub: API rate limit exceeded for installation ID 152184043. If you reach out to GitHub Support for help, please include the request ID E668:4DF5F:1005CE4:34FEB67:6AC88857 and timestamp 2026-10-09 06:23:20 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs\.github\.com/en/site\-policy/github\-terms/github\-terms\-of\-service\)
Documentation URL: https://docs\.github\.com/en/rest/using\-the\-rest\-api/getting\-started\-with\-the\-rest\-api\#rate\-limiting

Last completed report

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lane(s) and found 6 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below.

State: Ready for maintainer review
Priority: medium
Reviewed head: 8b8e7b925202
Updated: 1791526634 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 19 Active findings 11
Tests 4 Noted findings 0
Documentation 3 Resolved findings 25
Configuration 0 Pending checks/questions 0

Completeness: Complete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

The review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

  • medium · critique · Honor terminal and pause conflicts when saving pauses — `save_pause` unconditionally overwrites the pause map and reports `Inserted`. If a terminal outcome already exists for this key, this leaves a stale resumable pause behind instead (crates/tinyagents\-orchestration/src/subagent/driver\_completion\_tests\.rs:95)
  • medium · critique · Verify pause ownership before recording a terminal — This removes whatever pause currently exists and records the terminal without checking the `resume` token supplied by the completing invocation. If invocation A loaded pause A, inv (crates/tinyagents\-orchestration/src/subagent/driver\_completion\_tests\.rs:108)
  • medium · security · Preserve existing pauses when saving duplicate keys — This unconditionally replaces an existing pause and always returns `Inserted`. The persistence contract distinguishes `Replaced` and `Existing`, so a second lifecycle can overwrite (crates/tinyagents\-orchestration/src/subagent/driver\_completion\_tests\.rs:95)
  • medium · security · Reject pauses after terminal persistence — Saving a pause never checks `terminals`, so a pause can be inserted after another invocation has already closed the task. The persistence contract provides `TerminalExisting` for t (crates/tinyagents\-orchestration/src/subagent/driver\_completion\_tests\.rs:95)
  • medium · security · Do not consume a pause from another invocation — Terminal recording removes any pause for the key without examining the supplied resume token. A stale invocation can therefore consume a newer or unconsumed pause and insert a term (crates/tinyagents\-orchestration/src/subagent/driver\_completion\_tests\.rs:108)
  • medium · tests · Mark summarizer dispatch only after the model accepts the request — `mark_dispatched` fires before `model.invoke`, so a request the model layer rejects locally (validation, middleware rejection before any provider dispatch) still sets the flag, and (crates/tinyagents\-harness/src/summarization/model\_summarizer\.rs:201)
  • medium · tests · Mark task-state dispatch only after the model accepts the request — Same issue as the model summarizer: the dispatch flag is set before `model.invoke`, so a locally rejected request counts as a dispatched provider call and inflates `provider_starte (crates/tinyagents\-harness/src/summarization/task\_state/mod\.rs:263)
  • medium · tests · Sort message indices before deduplicating — `dedup` only collapses *adjacent* equal elements. Message indices are formatted into strings (`append:10:...` sorts before `append:2:...`) and a duplicate announcement of the same (crates/tinyagents\-integration\-tests/tests/graph\_lifecycle\_events\.rs:137)
  • medium · description · Sort message indices before deduplicating — `dedup` only removes *consecutive* equal elements. The lifecycle events are emitted in transcript order, but on a resume-with-retraction path the same index could be announced, ret (\(pull request description\))
  • medium · description · Mark dispatch only after the model accepts the request — `mark_dispatched()` runs before `model.invoke`, so a request the model layer rejects before any provider dispatch (local validation, request construction failure) still sets the fl (\(pull request description\))
  • medium · description · Mark task-state dispatch only after the model accepts the request — Same pattern as `ModelSummarizer`: the flag is set before `model.invoke` runs, so a pre-dispatch rejection inside the model layer is indistinguishable from a dispatched provider ca (\(pull request description\))

Resolved this pass

  • Record cancellations returned as lifecycle errors
  • Honor terminal persistence conflicts in the test memory
  • Mark dispatch only after the model accepts the request
  • Close the tool turn only after handling interrupts
  • Record cancellations returned as lifecycle errors
  • Mark dispatch only after model validation succeeds
  • Require retraction of the discarded tool message
  • Model terminal persistence conflicts in the test memory
  • Preserve existing pauses when saving duplicate keys
  • Preserve partial tool results on batch failure
  • Mark dispatch only after model validation succeeds in the state summarizer
  • Mark task-state dispatch only after the model accepts the request
  • Model terminal persistence conflicts in the test memory
  • Honor terminal persistence conflicts in the test memory
  • Model terminal persistence conflicts in the test memory
  • Close the tool turn only after handling interrupts
  • Record cancellations returned as lifecycle errors
  • Preserve partial tool results on batch failure
  • Require retraction of the discarded tool message
  • Record cancellations returned as lifecycle errors
  • Close the tool turn only after handling interrupts
  • Require retraction of the discarded tool message
  • Preserve partial tool results on batch failure
  • Model terminal persistence conflicts in the test memory
  • Honor terminal persistence conflicts in the test memory

Before merge

None.

How this fits together

flowchart LR
  n0["RunContext"]:::impacted
  n1["LoopState"]:::impacted
  n2["compile_loop"]:::impacted
  n3["tools_node"]:::impacted
  n4["model_node"]:::impacted
  n5["PLAN"]:::impacted
  n2 -->|uses| n1
  n2 -->|calls| n3
  n2 -->|calls| n4
  n2 -->|uses| n5
  n3 -->|uses| n0
  n3 -->|uses| n1
  n3 -->|uses| n5
  n4 -->|uses| n0
  n4 -->|uses| n1
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 2 files; 3 findings. (1 already reported on an earlier push) (13 earlier finding(s) still open) (2 observation(s) grouped into shared inline comments) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinyagents\-orchestration/src/subagent/driver\_completion\_tests\.rs — Honor terminal and pause conflicts when saving pauses
  • Evidence: crates/tinyagents\-orchestration/src/subagent/driver\_completion\_tests\.rs — Verify pause ownership before recording a terminal

security

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 2 files; 4 findings. (1 already reported on an earlier push) (1 observation(s) grouped into shared inline comments) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinyagents\-orchestration/src/subagent/driver\_completion\_tests\.rs — Preserve existing pauses when saving duplicate keys
  • Evidence: crates/tinyagents\-orchestration/src/subagent/driver\_completion\_tests\.rs — Reject pauses after terminal persistence
  • Evidence: crates/tinyagents\-orchestration/src/subagent/driver\_completion\_tests\.rs — Do not consume a pause from another invocation

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The graph driver now emits lifecycle events with retraction on interrupt, partial tool results survive batch failure, and subagent completions record cancellations and executor failures; most prior findings are fixed by the new tests. Three concerns still stand: the dispatch flag is set before the model call in both built-in summarizers, and the new deduplication assertion only removes adjacent duplicates. (5 earlier finding(s) still open) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinyagents\-harness/src/summarization/model\_summarizer\.rs — Mark summarizer dispatch only after the model accepts the request
  • Evidence: crates/tinyagents\-harness/src/summarization/task\_state/mod\.rs — Mark task-state dispatch only after the model accepts the request
  • Evidence: crates/tinyagents\-integration\-tests/tests/graph\_lifecycle\_events\.rs — Sort message indices before deduplicating

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The completion-recording work now behaves as described, with the test memory modeling terminal conflicts and duplicate-pause concerns partly addressed; the graph lifecycle tests cover interrupts, partial batches, resumes and nested tools, and the retract-on-interrupt and partial-batch fixes from earlier cycles are in place. Remaining: the dispatch trackers still mark before the model call, one test still deduplicates unsorted indices, and the test memory still overwrites an existing pause on duplicate save. (3 earlier finding(s) still open) (1 observation(s) grouped into shared inline comments) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: \(pull request description\) — Sort message indices before deduplicating
  • Evidence: \(pull request description\) — Mark dispatch only after the model accepts the request
  • Evidence: \(pull request description\) — Mark task-state dispatch only after the model accepts the request

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: gpt-5.6-luna, glm-5.3-flash
  • Spend: $0.010541
  • Tokens: 239537 input · 21391 output · 12243 cached · 0 embedding
Head State Pass summary
21b9dcb7b11c ready for maintainer review 6 active finding(s), 0 resolved finding(s) (at 1791525903)
ef0aaa618715 ready for maintainer review 12 active finding(s), 14 resolved finding(s) (at 1791526354)
8b8e7b925202 ready for maintainer review 11 active finding(s), 25 resolved finding(s) (at 1791526634)

tinysweeper 0.1.0

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-09T06:30:24.418119Z 26e518d New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

coderabbitai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

This review includes 26 billable files and costs up to $6.50.

Or wait 22 minutes for your next included review.

Check out review usage here.

View limit details

Limit details: You’ve used all 2 included reviews currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: ee9bb644-dffd-481b-a912-826d99076259
📥 Commits

Reviewing files that changed from the base of the PR and between d54a853 and 26e518d.

📒 Files selected for processing (26)
  • crates/tinyagents-graph/src/agent_loop/compile.rs
  • crates/tinyagents-graph/src/agent_loop/driver.rs
  • crates/tinyagents-graph/src/agent_loop/iter.rs
  • crates/tinyagents-graph/src/agent_loop/mod.rs
  • crates/tinyagents-graph/src/agent_loop/runtime.rs
  • crates/tinyagents-harness/src/agent_loop/entry.rs
  • crates/tinyagents-harness/src/agent_loop/lifecycle.rs
  • crates/tinyagents-harness/src/agent_loop/phases.rs
  • crates/tinyagents-harness/src/agent_loop/terminal_outcome_tests.rs
  • crates/tinyagents-harness/src/context/mod.rs
  • crates/tinyagents-harness/src/middleware/library/context.rs
  • crates/tinyagents-harness/src/middleware/library/context/overflow.rs
  • crates/tinyagents-harness/src/middleware/library/context/summary.rs
  • crates/tinyagents-harness/src/summarization/README.md
  • crates/tinyagents-harness/src/summarization/dispatch.rs
  • crates/tinyagents-harness/src/summarization/mod.rs
  • crates/tinyagents-harness/src/summarization/model_summarizer.rs
  • crates/tinyagents-harness/src/summarization/task_state/mod.rs
  • crates/tinyagents-integration-tests/tests/graph_lifecycle_events.rs
  • crates/tinyagents-integration-tests/tests/loop_as_graph.rs
  • crates/tinyagents-orchestration/src/subagent/README.md
  • crates/tinyagents-orchestration/src/subagent/completion.rs
  • crates/tinyagents-orchestration/src/subagent/driver.rs
  • crates/tinyagents-orchestration/src/subagent/driver_completion_tests.rs
  • crates/tinyagents-orchestration/src/subagent/outcome_status_map.rs
  • docs/modules/harness/terminal-outcome.md
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 21b9dcb7b1

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/tinyagents-graph/src/agent_loop/driver.rs
Comment thread crates/tinyagents-graph/src/agent_loop/runtime.rs Outdated

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0653 · 1,144,532 in / 56,232 out · 179,848 cached (16%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0398 · 606,282 in   / 34,421 out · 113,697 cached (19%) · gpt-5.6-luna, glm-5.3-flash
security:    $0.0246 · 431,501 in   / 17,296 out · 64,039 cached (15%)  · gpt-5.6-luna
tests:       $0.0004 · 54,678 in    / 1,463 out  · 1,920 cached (4%)    · glm-5.3-flash
description: $0.0002 · 25,189 in    / 133 out    · 64 cached (0%)       · glm-5.3-flash

Comment thread crates/tinyagents-integration-tests/tests/graph_lifecycle_events.rs
Comment thread crates/tinyagents-harness/src/summarization/task_state/mod.rs
Comment thread crates/tinyagents-graph/src/agent_loop/runtime.rs Outdated
Comment thread crates/tinyagents-orchestration/src/subagent/completion.rs
Comment thread crates/tinyagents-harness/src/summarization/model_summarizer.rs
@tinysweeper tinysweeper Bot added the priority: p2 Soon. Real but survivable — a rough edge, a gap, a thing that will bite later. label Oct 9, 2026
senamakel and others added 3 commits October 9, 2026 09:06
Add integration tests covering graph lifecycle events, verifying that
the expected events are emitted as a graph runs through its lifecycle.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The agent loop runtime now emits graph lifecycle events as a run
progresses, so subscribers can observe start, step, and completion
transitions. Integration tests cover the new event stream.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
… interrupt check, close interrupted model turn

Co-authored-by: Medulla <medulla@tinyhumans.ai>

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0201 · 387,726 in / 40,439 out · 37,109 cached (10%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0111 · 183,882 in / 21,016 out · 17,931 cached (10%) · gpt-5.6-luna, glm-5.3-flash
security:    $0.0082 · 118,723 in / 13,138 out · 16,106 cached (14%) · gpt-5.6-luna
tests:       $0.0003 · 27,929 in  / 2,084 out  · 1,536 cached (5%)   · glm-5.3-flash
description: $0.0002 · 27,662 in  / 1,486 out  · 1,408 cached (5%)   · glm-5.3-flash

Comment thread crates/tinyagents-integration-tests/tests/graph_lifecycle_events.rs Outdated
Comment thread crates/tinyagents-orchestration/src/subagent/driver_completion_tests.rs Outdated
Comment thread crates/tinyagents-graph/src/agent_loop/runtime.rs
Comment thread crates/tinyagents-graph/src/agent_loop/runtime.rs
Comment thread crates/tinyagents-integration-tests/tests/graph_lifecycle_events.rs
Comment thread crates/tinyagents-graph/src/agent_loop/runtime.rs
Comment thread crates/tinyagents-harness/src/summarization/model_summarizer.rs
Comment thread crates/tinyagents-harness/src/summarization/task_state/mod.rs

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ef0aaa6187

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/tinyagents-graph/src/agent_loop/runtime.rs Outdated
Comment thread crates/tinyagents-orchestration/src/subagent/completion.rs
senamakel and others added 3 commits October 9, 2026 09:14
…mpletion test double

Co-authored-by: Medulla <medulla@tinyhumans.ai>
The graph agent loop now reports lifecycle events through the harness so
subagent completion and phase transitions are observable. This lets
orchestration track subagent progress and integration tests assert on the
event stream.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Moved the agent loop runtime helpers out of the main module into their own
file to keep the loop logic easier to follow. No behaviour change.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0105 · 239,537 in / 21,391 out · 12,243 cached (5%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0050 · 87,412 in  / 7,294 out  · 6,679 cached (8%)  · gpt-5.6-luna, glm-5.3-flash
security:    $0.0048 · 66,210 in  / 7,340 out  · 5,372 cached (8%)  · gpt-5.6-luna
tests:       $0.0003 · 28,211 in  / 2,250 out  · 64 cached (0%)     · glm-5.3-flash
description: $0.0003 · 27,944 in  / 2,021 out  · 64 cached (0%)     · glm-5.3-flash

Comment thread crates/tinyagents-harness/src/summarization/model_summarizer.rs
Comment thread crates/tinyagents-harness/src/summarization/task_state/mod.rs
Comment thread crates/tinyagents-integration-tests/tests/graph_lifecycle_events.rs
…result for cancelled completions

Co-authored-by: Medulla <medulla@tinyhumans.ai>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 74b554511e

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/tinyagents-graph/src/agent_loop/runtime.rs
Comment thread crates/tinyagents-graph/src/agent_loop/runtime.rs
…ails

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@senamakel
senamakel merged commit c2661cc into main Oct 9, 2026
11 checks passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 26e518d96b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +65 to +66
pub(crate) fn adopt(&mut self, completed_turns: u32, open_from: Option<usize>) {
self.turn = self.turn.max(completed_turns);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Carry interrupted attempts into resumed turn numbering

When a compiled graph resumes in a fresh LoopRuntime after after_model requests an interrupt, the interrupted activation has already emitted and closed turn 1, but its local loop_state.turn += 1 is discarded, leaving the checkpoint at turn 0. This max therefore restores 0, and the resumed model emits another TurnStarted { turn: 1 } / TurnCompleted { turn: 1 }, causing consumers keyed by turn number to conflate separate attempts. Fresh evidence beyond the earlier interrupted-turn thread is that adopt now derives its counter solely from that stale checkpoint value, while the new resume test checks message indices but not turn-number uniqueness; persist the lifecycle counter across the interrupt and add a uniqueness assertion.

AGENTS.md reference: AGENTS.md:L70-L78

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p2 Soon. Real but survivable — a rough edge, a gap, a thing that will bite later.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant