Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 33 additions & 0 deletions crates/tinyagents-harness/src/agent_loop/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -172,6 +172,39 @@ allowed call ends on the blank reply instead of failing with
`0` to restore exact-replay behavior. The recovery lives in the shared
`run_loop`, so it applies identically to the unary and streaming paths.

Neither a smaller cap nor a lower effort label reliably stops a hosted model
that deliberates past its cap; `reasoning.effort = none` does. So after a dead
call the retry or nudged call goes out with reasoning switched off
(`RunPolicy::truncated_empty_reasoning_fallback`, default `true`), at the same
cap (the cap was for the deliberation), and the nudge tells the model to do its
working-out in the workspace. The hold-off backs off per death (1, 2, 4, 8, 16
live calls without reasoning, no ceiling) and the configured effort always
returns once it is spent: kept off for good, a model spends the rest of a run
writing probe programs instead of the deliverable;
the switch is announced as `AgentEvent::ControlApplied` (`reasoning_fallback`).
The state is run-wide (`agent_loop/reasoning_fallback.rs`): it outlives the
turn that set it.

A request's `reasoning.budget_tokens` is a promise the provider may not keep,
and every streamed call measured that reasoned past its budget with nothing
visible went on to die at the output cap. The reasoning watchdog
(`RunPolicy::reasoning_watchdog`, default `RequestBudget`) ends such a call at
the budget instead, dropping the stream, and hands the loop the same
`finish_reason = length`, no-content response the cap would have produced, so
the recovery above runs after a fraction of the wait
(`AgentEvent::ControlApplied`, `reasoning_watchdog`). Visible text or a
tool-call fragment disarms it; unary calls are not bounded.

The reasoning a call dies in is usually real work (a correct derivation, cut
off), and every retry used to begin it again from nothing. The tail of that
reasoning (`RunPolicy::truncated_empty_carry_reasoning_chars`, default 8,000
characters, `0` carries nothing) rides into the transcript as a user message
ahead of the retry or nudged call, framed as the model's own interrupted notes
with the instruction to continue from there in code rather than re-derive
(`AgentEvent::ControlApplied`, `truncated_empty_reasoning_carried`). The
watchdog's synthetic dead response keeps the reasoning that streamed, so a call
it ends is carried the same way as one the cap ended.

A provider can also end a stream normally after emitting only reasoning, with
no visible text or tool call. Hosts may set
`RunPolicy::empty_response_retries` to retry that non-truncated blank result;
Expand Down
1 change: 1 addition & 0 deletions crates/tinyagents-harness/src/agent_loop/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -130,6 +130,7 @@ mod model_switch;
mod model_turn;
mod nested;
pub mod phases;
mod reasoning_fallback;
mod response_recovery;
mod run_loop;
pub(crate) mod stream;
Expand Down
Loading
Loading