fix(agent): heal threads left with an unanswered tool call - #174
Open
ciaransweet wants to merge 2 commits into
Open
ciaransweet wants to merge 2 commits into
ciaransweet wants to merge 2 commits into
Conversation
A run cancelled while a tool is running (the client closing the stream, a pod rolled mid-turn) checkpoints the model's tool call with no result. Providers then reject every later turn on that thread; Mistral answers 400 "Not the same number of function calls and responses". with_session_state now always wires a before_agent middleware that closes each unanswered call with an error ToolMessage placed straight after it, written back to the checkpoint, so the thread heals on its next turn. A thread paused on interrupt() is untouched: a message on one is refused before the graph runs, and a resume skips before_agent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The plain agent build_agent returns under MCP_AGENT_STATE=0 bypasses with_session_state but keeps conversations the same way, so a cancelled run broke its threads just the same. It now carries the repair too. CONSUMING.md says both bundled agents carry it, and tells a host assembling its own checkpointed agent to add it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
j08lue
approved these changes
Oct 5, 2026
| raise ValueError(MISTRAL_3230) | ||
| if not calls: | ||
| reply = AIMessage( | ||
| "", tool_calls=[{"name": "search", "args": {}, "id": "SH2lrEzCG"}] |
Member
There was a problem hiding this comment.
Better parameterise this ID. But nvm.
| ) | ||
| from mcp_agent.main import with_session_state | ||
|
|
||
| MISTRAL_3230 = "Not the same number of function calls and responses" |
Member
There was a problem hiding this comment.
Interesting static detail. So this expectation is provider-dependent?
This will likely break when we use another provider.
Contributor
Author
There was a problem hiding this comment.
In that they report it differently, but the high level expectation is the same.
Contributor
Author
There was a problem hiding this comment.
Agree that it will likely break if we switch, but probably an 'expected' break.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
A run cancelled while a tool is running leaves the model's tool call in the checkpoint with no result. This happens when the client closes the stream or a pod is rolled mid-turn. Providers then reject every later turn on that thread, and nothing the user sends can fix it.
Seen in production on the ECMWF deployment (ecmwf/dss-agentic-ai-services#231):
/runs8 seconds into asearch_knowledgecall, and uvicorn cancelled the run (Cancelled via cancel scope … by RequestResponseCycle.run_asgi()).400 Not the same number of function calls and responses(code 3230).Fix
New
mcp_agent.interrupted_tool_calls: abefore_agentmiddleware that closes each unanswered tool call with an errorToolMessageplaced straight after it. It writes that back to the checkpoint, so the thread heals once, on its next turn, including threads that are already broken.with_session_statealways wires it in, afterStateCaptureMiddlewareand before the host's own middleware.build_agentwith session state off (MCP_AGENT_STATE=0) skipswith_session_statebut still keeps conversations, so it carries the repair as well.docs/CONSUMING.mdsays both bundled agents carry it, and tells a host assembling its own checkpointed agent (section 4b) to add it tomiddleware.interrupt()is never touched. A message on a paused thread is refused before the graph runs, and a resume re-enters the paused node without going throughbefore_agent.Proof
tests/mcp_agent/test_interrupted_tool_calls.pyreally cancels a run mid-tool against a model that validates history the way Mistral does.create_agent, the next turn raises the 3230 error.with_session_state, the next turn answers, and the repaired checkpoint has the placeholder right after the call.mistral-small-latestreturn the 400 as stored. Repaired, they succeed, and the model reissues the interruptedsearch_knowledgecall on its own.🤖 Generated with Claude Code