Skip to content

chore(vendor): bump tinyagents to 5d322814 (dead-call recovery, tinytools) - #7185

Merged
senamakel merged 6 commits into
tinyhumansai:mainfrom
sanil-23:pr/bump-tinyagents-dead-call
Oct 9, 2026
Merged

senamakel merged 6 commits into
tinyhumansai:mainfrom
sanil-23:pr/bump-tinyagents-dead-call

Conversation

@sanil-23

@sanil-23 sanil-23 commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Moves vendor/tinyagents from 1df6ad01 to 5d322814, the merge of tinyhumansai/tinyagents#343. That range is exactly two upstream PRs:

Lockfile. tinytools e2bf1be moves tinytools-std to sha2 0.11; Cargo.lock records that one line. No crate versions move and no release is needed.

Restores two calls #7145 had to drop. With RunContext::request_reasoning now vendored, the half-time and late deliverable notes ask for reasoning on the next call again. The clock-band test asserts both.

Not included on purpose. tinyagents main has since taken five harness-uplift merges and a new nested tinystoragedrivers submodule; that is a separate bump for whoever owns those changes.

Validation: cargo test -p openhuman --lib -- unmet_deliverable verify_before_finish (24 pass), cargo clippy -p openhuman --all-targets -- -D warnings clean, against a vendor tree at exactly these gitlinks.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Improvements
    • After a late-stage progress reminder, the agent now considers its next steps. At the halfway point, it does so when requested candidates are still missing, helping it continue working toward outstanding deliverables.

… for reasoning again

Moves vendor/tinyagents from 1df6ad01 to 5d322814, the merge of
tinyagents tinyhumansai#343, which carries exactly tinyagents tinyhumansai#342 (dead-call
recovery: reasoning-off fallback after a dead call, the streaming
reasoning watchdog, the dead call's reasoning carried into the retry,
reasoning handed back on a repeat note and for the finish check) and
tinyhumansai#343 (vendored tinytools to e2bf1be: tinytools tinyhumansai#52 to tinyhumansai#56).

tinytools e2bf1be moves tinytools-std to sha2 0.11, so Cargo.lock
records that; no other dependency changes.

With RunContext::request_reasoning now in the vendored harness, the
half-time and late notes ask for reasoning on the next call again, as
they did before tinyhumansai#7145 had to drop the calls. The clock-band test
asserts both requests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-09T13:06:10.923925Z 4a813ff New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

coderabbitai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Important

Review skipped

Review was skipped as selected files did not have any reviewable changes.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 90137065-0765-4fa6-af54-a102e8a92504

📥 Commits

Reviewing files that changed from the base of the PR and between bf5f006 and 4a813ff.


You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 02a6ceed-0f6a-4bfb-a03d-d8334638daf4

📥 Commits

Reviewing files that changed from the base of the PR and between b9bd7cd and bf5f006.


⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock

📒 Files selected for processing (3)
  • crates/openhuman-core/src/agent/tinyagents/middleware/unmet_deliverable.rs
  • crates/openhuman-core/src/agent/tinyagents/middleware/unmet_deliverable_tests.rs
  • vendor/tinyagents

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review.



📝 Walkthrough

Walkthrough

Clock-triggered notes now request reasoning on the next call. Tests verify that both note paths set the run context’s repeat-noted flag. The vendored tinyagents reference points to a new commit.

Changes

Clock-Triggered Notes

Layer / File(s) Summary
Clock-note emission and repeat-flag checks
crates/openhuman-core/src/agent/tinyagents/middleware/unmet_deliverable.rs, crates/openhuman-core/src/agent/tinyagents/middleware/unmet_deliverable_tests.rs, vendor/tinyagents
The late note and the half-time note for missing candidates request reasoning. Tests verify that consuming the repeat-noted flag returns true for both notes. The vendored tinyagents reference changes to commit 5d3228142f4a202ccb13d77c4a23b97bbc85510.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Suggested reviewers: senamakel


Merge Risk

Merge Risk: ⚪ Minimal · up to bf5f0

The available source confirms the clock-note paths issue the repeat-noted signal and tests check it. The exact vendored consumer is unavailable, but no concrete failure is established to block merging.

Security Architecture Review

Security architecture risk: 🔵 Low · up to bf5f0

The visible local change does not expand access or privileges. However, the accompanying recovery and tool updates could affect execution behavior, and their exact implementations were unavailable for review. Recovery and cancellation guarantees therefore remain unconfirmed.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The demonstrated local effect is within existing agent runs and their model/tool interaction. The bundled tool changes also reach HTTP handling and file-patch behavior according to the PR description, but their effective authorization, filesystem, network, and tenant exposure cannot be bounded without the recorded vendor source.

Trust Boundaries and Controls

  • inferred — The local additions do not introduce a new caller, replace run identity, change candidate access checks, or grant tool authority. Existing clock conditions select when reasoning is requested. This counterevidence applies to the local patch, not to the unverified vendor update.

Resilience and Maintainability Implications

  • observed — The changed test checks the repeat-noted signal after both clock notes. It directly invokes after_tool and does not establish production next-call consumption, recovery ordering, or flag cleanup after interruption.



🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check Passed The title clearly identifies the vendor bump and names the main included changes: dead-call recovery and tinytools updates.
Docstring Coverage Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 2 functions across 2 files. (1 skipped: 1 …
Linked Issues check Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check Passed Check skipped because no linked issues were found for this pull request.


✨ Finishing Touches 💡 2
⚔️ Resolve merge conflicts 💡
  • Resolve merge conflict in branch pr/bump-tinyagents-dead-call


🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR


  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit taps the clock at night,
A note appears, then prompts take flight.
The half-time flag is checked with care,
The late note joins it in the air.
New tinyagents hop into view.

Comment @coderabbitai help to get the list of available commands.

@tinysweeper

tinysweeper Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lane(s) and found 0 active actionable finding(s). The diff adds ctx.request_reasoning() calls at the half-time and late unmet-deliverable clock rungs, with unit tests asserting both signals. Reviewers judged it safe to merge; end-to-end jobs are still pending. Detailed lane evidence and any incomplete work are listed below.

State: Reviewing pending checks
Priority: none
Reviewed head: 4a813ffd9d14
Updated: 1791550833 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 1 Active findings 0
Tests 1 Noted findings 0
Documentation 0 Resolved findings 0
Configuration 0 Pending checks/questions 4

Completeness: Complete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

No supported behavioral explanation was produced.

Features

  • Modified — Unmet-deliverable clock-rung reasoning requests: At the half-time and late clock rungs, the middleware now asks the agent loop for reasoning on the next call in addition to appending its guidance note to the tool result. This shapes how the tinyagents run loop drives subsequent model calls at the moments where reasoning is most valuable; no route, wire format, or UI surface changes. (crates/openhuman-core/src/agent/tinyagents/middleware/unmet_deliverable.rs#impl<C: Send + Sync> Middleware<(), C> for UnmetDeliverableMiddleware {)

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

No active actionable findings.

Pending checks: Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS)

Before merge

  • Wait for Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS).
Agent review details

critique

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The change requests reasoning after the half-time and late unmet-deliverable notes, and the accompanying tests verify both signals. It looks safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

security

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The change requests model reasoning when the middleware emits its half-time or late guidance, and the accompanying tests verify both paths. It looks safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Positive: The updated test asserts the reasoning-request signal for both the half-time and late paths via ctx.take_repeat_noted(), directly covering the behaviour change.
  • Lane summary: The change adds `ctx.request_reasoning()` at the half-time and late note points, and the existing sibling test file is updated to assert via `ctx.take_repeat_noted()` that each path actually requests reasoning — so removing either call would fail the test. The assertions target the real behaviour change, not just the return value. No new problems found; safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Positive: Nothing sensitive found in what this pull request commits.
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Positive: The diff matches the PR description: ctx.request_reasoning() is restored at the half-time and late unmet-deliverable notes, with tests asserting both paths.
  • Lane summary: The PR restores ctx.request_reasoning() calls at the half-time and late unmet-deliverable notes, exactly as the description claims, with tests asserting both via take_repeat_noted. The diff matches the description and looks safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: The change makes the unmet-deliverable middleware request reasoning from the agent loop at the half-time and late clock rungs. Its effect is internal to the tinyagents run loop — it shapes how the loop drives subsequent model calls and is directly asserted by the sibling unit tests via take_repeat_noted(); no route, wire format, or UI surface changes, so no end-to-end test is needed and none could meaningfully assert it. The change looks sound. Waiting on end-to-end jobs: `Rust E2E (mock backend)`, `Build Playwright E2E Artifact`, `E2E (Playwright / web lane)`, `Desktop E2E (full suite, 3 OS)`.
  • Unresolved questions/checks: Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS)
Evidence and run details
  • Models: gpt-5.6-luna, glm-5.3-flash
  • Spend: $0.000430
  • Tokens: 44277 input · 2219 output · 9136 cached · 0 embedding
Head State Pass summary
bf5f006a64a1 pending 0 active finding(s), 0 resolved finding(s) (at 1791545964)
4a813ffd9d14 pending 0 active finding(s), 0 resolved finding(s) (at 1791550833)

tinysweeper 0.1.0

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0004 · 41,221 in / 3,051 out · 4,272 cached (10%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0001 · 12,767 in / 611 out   · 2,130 cached (17%) · gpt-5.6-luna
security:    $0.0001 · 11,123 in / 339 out   · 1,886 cached (17%) · gpt-5.6-luna
tests:       $0.0000 · 4,574 in  / 296 out   · 64 cached (1%)     · glm-5.3-flash
description: $0.0000 · 4,363 in  / 139 out   · 64 cached (1%)     · glm-5.3-flash
e2e:         $0.0001 · 5,441 in  / 715 out   · 64 cached (1%)     · glm-5.3-flash

@tinysweeper tinysweeper Bot added the priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect. label Oct 9, 2026
coderabbitai[bot]
coderabbitai Bot previously approved these changes Oct 9, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bf5f006a64

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread Cargo.lock
"rustix",
"serde_json",
"sha2 0.10.9",
"sha2 0.11.0",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Refresh the desktop lockfile with tinytools

crates/openhuman-app/Cargo.lock is a separate Cargo world, but its tinytools-std dependency still says sha2 0.10.9 at line 8279 while this updated path crate requires 0.11. The checked cargo tree --help describes --locked as “Assert that Cargo.lock will remain unchanged”, and scripts/check-linux-tls-dependencies.sh:11,90 runs that exact mode against the Tauri manifest, so every Rust linux-tls-policy lane for this vendor bump exits before the dependency policy check; refresh the app lockfile along with the root one.

AGENTS.md reference: AGENTS.md:L16-L16

Useful? React with 👍 / 👎.

senamakel and others added 5 commits October 9, 2026 15:58
Update the vendored submodule pointers for tinyagents, tinymcp, tinysearch and tinyskills to their latest upstream commits.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…e dead-call bump)

Co-authored-by: Medulla <medulla@tinyhumans.ai>
Update the vendored submodule pointers for tinyagents, tinymcp, tinysearch and tinyskills to their latest revisions.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@senamakel
senamakel merged commit 8dc4656 into tinyhumansai:main Oct 9, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants