Repository navigation
chore(vendor): bump tinyagents to main (typed TerminalOutcome) - #7125
Conversation
Update the vendored tinyagents submodule to a newer commit. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add tests for the session host driver covering progress tracing and journal projection cost rollup, and adjust the driver so these paths are exercised correctly. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Extend the stalled model stream test to check that the typed terminal outcome carries a ProviderFailed reason classified from the generation stall, so the failure path is verified rather than only the partial narration. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Tiny Sweeper reviewTiny Sweeper reviewed this change across 6 lane(s) and found 1 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below. State: Reviewing pending checks Review snapshot
Completeness: Complete What changedThe review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below. FeaturesNone identified with supported citations. TestsNo supported feature-to-test mapping was produced. Test execution is not inferred. Findings
Resolved this pass
Pending checks: Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS) Before merge
How this fits togetherflowchart LR
n0["..._usage_does_not_roll_into_the_parent_turn<br/>changed"]:::changed
n1["turn_with_usage<br/>changed"]:::changed
n2["...ide_a_subagent_projects_a_child_tool_span<br/>changed"]:::changed
n3["...own_tool_call_projects_a_failed_tool_span<br/>changed"]:::changed
n4["...cts_failed_subagent_from_child_run_failed<br/>changed"]:::changed
n5["projects_turn_content_from_root_model_io<br/>changed"]:::changed
n6["spans_from_observations"]:::impacted
n7["with_capture_content"]:::impacted
n8["vec"]:::impacted
n9["turn_span"]:::impacted
n10["ctx"]:::impacted
n0 -->|calls| n6
n0 -->|tests| n6
n0 -->|calls| n8
n0 -->|tests| n8
n0 -->|calls| n9
n0 -->|tests| n9
n1 -->|calls| n8
n2 -->|calls| n6
n2 -->|tests| n6
n2 -->|calls| n8
n2 -->|tests| n8
n3 -->|calls| n6
n3 -->|tests| n6
n3 -->|calls| n8
n3 -->|tests| n8
n4 -->|calls| n6
n4 -->|tests| n6
n4 -->|calls| n8
n4 -->|tests| n8
n4 -->|calls| n10
n4 -->|tests| n10
n5 -->|calls| n6
n5 -->|tests| n6
n5 -->|calls| n7
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 8 remain after this review. 📝 WalkthroughWalkthroughThe driver now classifies chained typed errors and carries terminal outcomes on failures. Successful driver outcomes set the outcome to ChangesTerminal outcomes
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~15 minutes Change: Other Suggested reviewers: Merge Risk: ⚪ Minimal · up to The change adds typed terminal outcomes to driver failures and keeps successful runs unchanged. No actionable merge-blocking risk remains. Architecture SummaryArchitecture risk: 🔵 Low · up to The change affects 1 system. Changed systems: Architecture concerns Review detailsSystems and components
Before / after behavior
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
A rabbit checks the typed error trail, Comment |
There was a problem hiding this comment.
tinysweeper found nothing blocking. Approving.
$0.0064 · 90,328 in / 6,716 out · 12,086 cached (13%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0039 · 43,172 in / 2,129 out · 7,883 cached (18%) · gpt-5.6-luna
security: $0.0022 · 22,245 in / 655 out · 4,011 cached (18%) · gpt-5.6-luna
tests: $0.0001 · 6,323 in / 743 out · 64 cached (1%) · glm-5.3-flash
description: $0.0001 · 6,270 in / 191 out · 64 cached (1%) · glm-5.3-flash
e2e: $0.0001 · 7,191 in / 698 out · 64 cached (1%) · glm-5.3-flash
…ilure Add a driver test asserting that a typed error still yields a terminal outcome when the transcript snapshot is empty, while an untyped error leaves the outcome unset. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
There was a problem hiding this comment.
tinysweeper found nothing blocking. Approving.
$0.0049 · 81,760 in / 4,607 out · 7,644 cached (9%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0025 · 27,616 in / 1,088 out · 4,066 cached (15%) · gpt-5.6-luna
security: $0.0022 · 26,393 in / 523 out · 3,578 cached (14%) · gpt-5.6-luna
tests: $0.0001 · 7,019 in / 308 out · 0 cached (0%) · glm-5.3-flash
description: $0.0001 · 7,055 in / 104 out · 0 cached (0%) · glm-5.3-flash
e2e: $0.0001 · 7,887 in / 856 out · 0 cached (0%) · glm-5.3-flash
Summary
vendor/tinyagentsto currentmain(c8b5c1f2, includes tinyagents#329).DriverOutcome/DriverFailurecarry an optional typedTerminalOutcome;AgentEvent::RunCompleted/RunFailedcarryoutcome: Option<TerminalOutcome>.TinyAgentsErrorin the error chain (stalled generation check, and aTerminalOutcomeattached to theDriverFailure) instead of matching rendered error text.Problem
tinyagents#329 changed public struct shapes (
DriverFailure,DriverOutcome,AgentEvent::RunFailed/RunCompleted), so the host no longer compiled against the new pin. The host also detected a stalled generation by substring-matching the error string.Solution
TerminalOutcome(reason, class, timeout phase,provider_started, message) onRunFailed/RunCompletedandAgentRun::terminal;DriverFailure.outcome/DriverOutcome.outcome; new lifecycle eventsTurnStarted,TurnCompletedandMessageAppended; a graph-driven agent loop driver; summarizer failures classified from their inner error; timeout phase backfill.outcome: Noneon successfulDriverOutcomes (runtime derives it frominterrupted), a typed outcome viaTerminalOutcome::from_erroron failures;journal_projectiontest literals updated. The new lifecycle events are not projected into spans by this PR.Submission Checklist
TerminalOutcome; existing literals updated)agent::session_host::driver_testsCloses #NNN: N/A, dependency bumpImpact
cargo check --tests(root andopenhuman-cliwith product features),openhuman-appcheck, clippy-D warnings,cargo fmt --check,pnpm rust:layout, feature-forwarding and submodule-monotonic checks,agent::lib tests (1792 passed,RUST_MIN_STACK=16777216).Related
Co-authored-by: Medulla medulla@tinyhumans.ai
Summary by CodeRabbit