Skip to content

chore(vendor): bump tinyagents to main (typed TerminalOutcome) - #7125

Merged
senamakel merged 4 commits into
tinyhumansai:mainfrom
senamakel:tinyagents-uplift-bump-3
Oct 8, 2026
Merged

senamakel merged 4 commits into
tinyhumansai:mainfrom
senamakel:tinyagents-uplift-bump-3

Conversation

@senamakel

@senamakel senamakel commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Bump vendor/tinyagents to current main (c8b5c1f2, includes tinyagents#329).
  • Adapt the session host to the new driver contract: DriverOutcome / DriverFailure carry an optional typed TerminalOutcome; AgentEvent::RunCompleted / RunFailed carry outcome: Option<TerminalOutcome>.
  • The partial-failure path now classifies from the typed TinyAgentsError in the error chain (stalled generation check, and a TerminalOutcome attached to the DriverFailure) instead of matching rendered error text.

Problem

tinyagents#329 changed public struct shapes (DriverFailure, DriverOutcome, AgentEvent::RunFailed / RunCompleted), so the host no longer compiled against the new pin. The host also detected a stalled generation by substring-matching the error string.

Solution

  • tinyagents changes brought in: typed TerminalOutcome (reason, class, timeout phase, provider_started, message) on RunFailed / RunCompleted and AgentRun::terminal; DriverFailure.outcome / DriverOutcome.outcome; new lifecycle events TurnStarted, TurnCompleted and MessageAppended; a graph-driven agent loop driver; summarizer failures classified from their inner error; timeout phase backfill.
  • Host: outcome: None on successful DriverOutcomes (runtime derives it from interrupted), a typed outcome via TerminalOutcome::from_error on failures; journal_projection test literals updated. The new lifecycle events are not projected into spans by this PR.
  • Test: the stalled-generation test now asserts the typed outcome.

Submission Checklist

  • Tests added or updated (stalled-generation driver test asserts the typed TerminalOutcome; existing literals updated)
  • Diff coverage: changed host lines are exercised by agent::session_host::driver_tests
  • Coverage matrix updated: N/A, no feature rows change
  • Affected feature IDs listed under Related: N/A
  • No new external network dependencies introduced
  • Manual smoke checklist updated: N/A, no release-cut surface touched
  • Linked issue closed via Closes #NNN: N/A, dependency bump

Impact

  • Desktop, CLI and TUI hosts share the agent loop; behavior is unchanged except that failures now carry a typed outcome.
  • Validation: cargo check --tests (root and openhuman-cli with product features), openhuman-app check, clippy -D warnings, cargo fmt --check, pnpm rust:layout, feature-forwarding and submodule-monotonic checks, agent:: lib tests (1792 passed, RUST_MIN_STACK=16777216).

Related

Co-authored-by: Medulla medulla@tinyhumans.ai

Summary by CodeRabbit

  • Bug Fixes
    • Stalled responses are now recognized as provider failures, with the corresponding failover reason recorded.
    • Failure classifications are preserved when a run ends without messages or returns partial results, improving the accuracy of reported outcomes.
    • Runs without a typed failure classification no longer receive an inferred terminal outcome, while typed cancellation and other classified failures retain their reason.

senamakel and others added 3 commits October 8, 2026 11:23
Update the vendored tinyagents submodule to a newer commit.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add tests for the session host driver covering progress tracing and journal projection cost rollup, and adjust the driver so these paths are exercised correctly.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Extend the stalled model stream test to check that the typed terminal
outcome carries a ProviderFailed reason classified from the generation
stall, so the failure path is verified rather than only the partial
narration.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@tinysweeper

tinysweeper Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lane(s) and found 1 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below.

State: Reviewing pending checks
Priority: medium
Reviewed head: a2c492629863
Updated: 1791449676 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 1 Active findings 1
Tests 4 Noted findings 0
Documentation 0 Resolved findings 4
Configuration 0 Pending checks/questions 4

Completeness: Complete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

The review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

  • medium · e2e · Cover the empty-messages branch that now attaches the terminal outcome — The new `outcome` field propagates a typed terminal outcome (`TerminalReason::Cancelled`, `ProviderFailed`) onto `DriverFailure` and through `AgentEvent::RunCompleted`/`RunFailed`. (crates/openhuman\-core/src/agent/session\_host/driver\.rs:417)

Resolved this pass

  • Cover the empty-messages branch that now attaches the terminal outcome
  • Cover the empty-messages branch that now attaches the terminal outcome
  • Cover the empty-messages branch that now attaches the terminal outcome
  • Cover the empty-messages branch that now attaches the terminal outcome

Pending checks: Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS)

Before merge

  • Wait for Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS).

How this fits together

flowchart LR
  n0["..._usage_does_not_roll_into_the_parent_turn<br/>changed"]:::changed
  n1["turn_with_usage<br/>changed"]:::changed
  n2["...ide_a_subagent_projects_a_child_tool_span<br/>changed"]:::changed
  n3["...own_tool_call_projects_a_failed_tool_span<br/>changed"]:::changed
  n4["...cts_failed_subagent_from_child_run_failed<br/>changed"]:::changed
  n5["projects_turn_content_from_root_model_io<br/>changed"]:::changed
  n6["spans_from_observations"]:::impacted
  n7["with_capture_content"]:::impacted
  n8["vec"]:::impacted
  n9["turn_span"]:::impacted
  n10["ctx"]:::impacted
  n0 -->|calls| n6
  n0 -->|tests| n6
  n0 -->|calls| n8
  n0 -->|tests| n8
  n0 -->|calls| n9
  n0 -->|tests| n9
  n1 -->|calls| n8
  n2 -->|calls| n6
  n2 -->|tests| n6
  n2 -->|calls| n8
  n2 -->|tests| n8
  n3 -->|calls| n6
  n3 -->|tests| n6
  n3 -->|calls| n8
  n3 -->|tests| n8
  n4 -->|calls| n6
  n4 -->|tests| n6
  n4 -->|calls| n8
  n4 -->|tests| n8
  n4 -->|calls| n10
  n4 -->|tests| n10
  n5 -->|calls| n6
  n5 -->|tests| n6
  n5 -->|calls| n7
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The updated tests correctly adapt error inputs to the `anyhow::Error` interface and add coverage for typed terminal outcomes on an empty snapshot. The prior empty-messages coverage finding is fixed, and this change looks safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

security

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The change adds coverage for typed terminal outcomes, including the empty-snapshot path, and looks safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The revision addresses the prior gap: `empty_snapshot_failure_still_carries_the_typed_terminal_outcome` now pins the empty-messages branch, covering both the typed (Cancelled → TerminalReason::Cancelled) and untyped (outcome stays None) paths, and the stalled test asserts the derived terminal reason. The new tests fail if the classification or attachment regresses, so the change looks sound and safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The vendor bump adapts the host to the typed TerminalOutcome contract; the previously missing coverage for the empty-messages branch is now supplied by empty_snapshot_failure_still_carries_the_typed_terminal_outcome, and the stalled test asserts the typed outcome. Looks sound to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: This change adds a typed `terminal::TerminalOutcome` to driver failures and agent run events, classified from the harness error chain. The empty-messages branch is now unit-tested in driver_tests.rs, but no end-to-end test drives the new `outcome` field; the prior coverage finding stands. Waiting on end-to-end jobs: `Rust E2E (mock backend)`, `Build Playwright E2E Artifact`, `E2E (Playwright / web lane)`, `Desktop E2E (full suite, 3 OS)`.
  • Unresolved questions/checks: Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS)
  • Evidence: crates/openhuman\-core/src/agent/session\_host/driver\.rs — Cover the empty-messages branch that now attaches the terminal outcome
Evidence and run details
  • Models: gpt-5.6-luna, glm-5.3-flash
  • Spend: $0.004931
  • Tokens: 81760 input · 4607 output · 7644 cached · 0 embedding
Head State Pass summary
b92f2e99ff54 pending 1 active finding(s), 0 resolved finding(s) (at 1791448749)
a2c492629863 pending 1 active finding(s), 4 resolved finding(s) (at 1791449676)

tinysweeper 0.1.0

@coderabbitai

coderabbitai Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: fe45d9b6-c941-47f0-8bc6-94cad853b66e
📥 Commits

Reviewing files that changed from the base of the PR and between b92f2e9 and a2c4926.

📒 Files selected for processing (1)
  • crates/openhuman-core/src/agent/session_host/driver_tests.rs

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 8 remain after this review.


📝 Walkthrough

Walkthrough

The driver now classifies chained typed errors and carries terminal outcomes on failures. Successful driver outcomes set the outcome to None. Journal projection test fixtures set event outcomes to None.

Changes

Terminal outcomes

Layer / File(s) Summary
Driver outcome classification
vendor/tinyagents, crates/openhuman-core/src/agent/session_host/driver.rs, crates/openhuman-core/src/agent/session_host/driver_tests.rs, crates/openhuman-core/src/agent/session_host/runtime_adapter_tests.rs
The driver accepts anyhow::Error, classifies chained TinyAgentsError values, and attaches classified outcomes to failures. Tests cover GenerationStalled classification and empty-snapshot outcomes. The test driver sets its outcome to None. The TinyAgents submodule pointer changes.
Journal projection fixture outcomes
crates/openhuman-core/src/agent/progress_tracing/journal_projection_tests.rs, crates/openhuman-core/src/agent/progress_tracing/journal_projection_cost_rollup_tests.rs
Fixtures now explicitly set RunCompleted and RunFailed outcomes to None.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~15 minutes

Change: Other

Suggested reviewers: m3ga-mind

Merge Risk: ⚪ Minimal · up to a2c49

The change adds typed terminal outcomes to driver failures and keeps successful runs unchanged. No actionable merge-blocking risk remains.

Architecture Summary

Architecture risk: 🔵 Low · up to a2c49

The change affects 1 system.

Changed systems: crates

Architecture concerns
No architecture-level concerns identified.

Review details

Systems and components

  • observed — crates (service) was modified; 5 changed files map to changed impact.

Before / after behavior

  • observed — Modified behavior in crates/openhuman-core/src/agent/progress_tracing/journal_projection_cost_rollup_tests.rs: The turn_with_usage fixture now explicitly sets the run-completion outcome to None.
  • observed — Modified behavior in crates/openhuman-core/src/agent/progress_tracing/journal_projection_cost_rollup_tests.rs: The subagent usage fixture now explicitly sets the run-completion outcome to None.
  • observed — Modified behavior in crates/openhuman-core/src/agent/progress_tracing/journal_projection_cost_rollup_tests.rs: The recovered unknown-tool-call fixture now explicitly sets the run-completion outcome to None.
  • observed — Modified behavior in crates/openhuman-core/src/agent/progress_tracing/journal_projection_cost_rollup_tests.rs: The subagent unknown-tool-call fixture now explicitly sets the run-completion outcome to None.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 66.67% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 18 functions across 5 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately identifies the vendor submodule update and the typed TerminalOutcome integration. It is concise and specific to the main change.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit checks the typed error trail,
And marks the outcomes, clear and pale.
The empty snapshot keeps its clue,
While fixtures state what they do.
The journal rests, its fields set right,
Then hops away into the night.

Comment @coderabbitai help to get the list of available commands.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0064 · 90,328 in / 6,716 out · 12,086 cached (13%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0039 · 43,172 in / 2,129 out · 7,883 cached (18%)  · gpt-5.6-luna
security:    $0.0022 · 22,245 in / 655 out   · 4,011 cached (18%)  · gpt-5.6-luna
tests:       $0.0001 · 6,323 in  / 743 out   · 64 cached (1%)      · glm-5.3-flash
description: $0.0001 · 6,270 in  / 191 out   · 64 cached (1%)      · glm-5.3-flash
e2e:         $0.0001 · 7,191 in  / 698 out   · 64 cached (1%)      · glm-5.3-flash

Comment thread crates/openhuman-core/src/agent/session_host/driver_tests.rs
@tinysweeper tinysweeper Bot added the priority: p2 Soon. Real but survivable — a rough edge, a gap, a thing that will bite later. label Oct 8, 2026
coderabbitai[bot]
coderabbitai Bot previously approved these changes Oct 8, 2026
…ilure

Add a driver test asserting that a typed error still yields a terminal
outcome when the transcript snapshot is empty, while an untyped error
leaves the outcome unset.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0049 · 81,760 in / 4,607 out · 7,644 cached (9%)  · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0025 · 27,616 in / 1,088 out · 4,066 cached (15%) · gpt-5.6-luna
security:    $0.0022 · 26,393 in / 523 out   · 3,578 cached (14%) · gpt-5.6-luna
tests:       $0.0001 · 7,019 in  / 308 out   · 0 cached (0%)      · glm-5.3-flash
description: $0.0001 · 7,055 in  / 104 out   · 0 cached (0%)      · glm-5.3-flash
e2e:         $0.0001 · 7,887 in  / 856 out   · 0 cached (0%)      · glm-5.3-flash

Comment thread crates/openhuman-core/src/agent/session_host/driver.rs
@senamakel
senamakel merged commit ceacd54 into tinyhumansai:main Oct 8, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p2 Soon. Real but survivable — a rough edge, a gap, a thing that will bite later.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant