Skip to content

Tool-path rate limits: retry at the bridge, then fail the run fast - #492

Open
rejojer wants to merge 3 commits into
mainfrom
fix/tool-429-fail-fast
Open

Tool-path rate limits: retry at the bridge, then fail the run fast#492
rejojer wants to merge 3 commits into
mainfrom
fix/tool-429-fail-fast

Conversation

@rejojer

@rejojer rejojer commented Sep 8, 2026

Copy link
Copy Markdown
Member

A PageIndex cloud 429 (or 5xx) on a tool call used to reach the model as an INTERNAL_ERROR envelope saying "try again": the model re-called once with no wait, then wrote the failure into its answer, and chat() returned normally with no status anywhere. The same 429 before the loop (the doc_id targeting lookup) already propagated raw.

Changes

  • McpBridge mounts a urllib3 Retry: 429, 500/502/503/504 and connection failures, three attempts, 0/2/4 s apart. Retry-After is ignored: a long one is a quota, not a blip, and parsing it was the one way a 429 could turn into "could not reach the server" (urllib3 raises on a non-integer value). Read timeouts are never replayed (240 s each, and the server may have acted). Exhausted, the last response falls through to the existing >= 400 branch, so status_code survives.
  • _bridge_invoker re-raises 429/5xx and an unreachable server (a transport failure that outlived the connection retries) alongside 401/403. The frameworks turn a raised tool exception back into model-visible text, so each chat() door gets its own escape:
    • openai-agents: the in-process MCPServer passes a failure_error_function that lets a PageIndex-caused failure propagate, and _translate_run_error unwraps it from the framework's wrapper (this also un-flattens the mid-session 401 case).
    • Messages lane: each turn's tools run through the runner's public generate_tool_call_response(), and a recorded failure raises before the next model call.
  • _model_backend_error keeps the provider's status_code.
  • The handshake error blames the API key only on 401/403; a rate-limited handshake no longer tells the user to rotate a working key.

Claude Agent SDK tools cannot fail fast: the SDK MCP server converts handler exceptions into JSON-RPC errors for Claude Code by design.

Behaviour change on the public tool surfaces

as_openai_tools() users running their own Runner.run now get an exception for a post-retry 429/5xx or an unreachable server (the 401/403 re-raise always intended this; the framework absorbed it). as_anthropic_tools() users' runners absorb it into an is_error result and log a traceback. The plain functions from agent_tools() raise PageIndexAPIError for the same set, where their docstring used to promise only 401/403; the docstrings now state the real raise set.

Verification

  • New tests against a local HTTP stub: the retry schedule (fixed backoff with a Retry-After header present and ignored), exhausted-retry status for 429 and 504, read-timeout non-replay, an unreachable server re-raised through the invoker, the handshake wording per status; plus the invoker re-raise matrix, openai-agents escape, run-error unwrap, provider status_code, end-to-end fail-fast on the chat and Messages doors, and a regression guard for model-side slips.
  • Full suite green; the without-frameworks leg green; both agent test modules green on the openai-agents 0.18.1 floor; anthropic 0.108.0 already has generate_tool_call_response.
  • Live against the cloud with real models: a forced post-retry 429 raises PageIndexAPIError(status_code=429) on chat() default, streamed, and protocol="messages", after one tool call.

https://claude.ai/code/session_014S88dcSz7jykegAWyWZk8E
https://claude.ai/code/session_013xk3xt9KgHNTjsYmFLKxbu

A PageIndex cloud 429 (or 5xx) on a tool call used to reach the model as
an INTERNAL_ERROR envelope saying "try again": the model re-called once
with no wait, then wrote the failure into its answer, and chat() returned
normally with no status anywhere. The same 429 before the loop (the
doc_id targeting lookup) already propagated raw.

- McpBridge mounts a urllib3 Retry: 429/502/503 and connection failures,
  three attempts, 0/2/4 s apart or as Retry-After says; read timeouts
  are never replayed (240 s each, and the server may have acted); a
  Retry-After past a minute is a quota, not a blip, so the backoff runs
  instead of sleeping it out. Exhausted, the last response falls through
  to the existing >= 400 branch, so the status_code survives.
- _bridge_invoker re-raises 429/5xx alongside 401/403. The frameworks
  turn a raised tool exception back into model-visible text, so each
  chat() door gets its own escape: the in-process MCPServer's
  failure_error_function lets a PageIndex-caused failure propagate and
  _translate_run_error unwraps it from the framework's wrapper (which
  also un-flattens the 401 case); the Messages lane runs each turn's
  tools through the runner's public generate_tool_call_response() and
  raises before the next model call.
- _model_backend_error keeps the provider's status_code.

Claude Agent SDK tools cannot fail fast: the SDK MCP server converts
handler exceptions into JSON-RPC errors for Claude Code by design.

Claude-Session: https://claude.ai/code/session_014S88dcSz7jykegAWyWZk8E
Comment thread tests/test_agent_tools.py
def log_message(self, *args):
pass

def do_POST(self):
Comment thread tests/test_local_chat.py Fixed
Comment thread tests/test_agent_tools.py
def test_bridge_read_timeout_is_not_retried(mcp_stub, monkeypatch):
"""A read timeout is a full wait the server may have acted on:
surfaced once, never replayed."""
import pageindex.mcp_bridge as mcp_bridge
Comment thread tests/test_agent_tools.py
escapes the run (a model-side slip staying model-visible is covered end
to end in test_local_chat)."""
pytest.importorskip("agents")
import pageindex.mcp_bridge as mcp_bridge
The bridge retry is now a plain urllib3 Retry: 429 and every 5xx retried
three times at the fixed 0/2/4 s backoff, Retry-After ignored. That drops
the _Retry subclass, whose get_retry_after raised InvalidHeader on a
non-integer header (turning a 429 into "could not reach the server"),
honoured a 60 s Retry-After three times over, and let a 413 carrying
Retry-After replay. 500 and 504 join the forcelist so the invoker's "what
survived the bridge's retries" holds for every status it re-raises. Retry
is imported from requests.adapters, the declared dependency.

The invoker re-raises transport failures too: once the bridge's own
connection retries fail, the model cannot reach the server either, and
the envelope only sent it round the retry loop.

The handshake error blames the API key only on 401/403: a rate-limited
handshake is now a run-terminating error and was telling users to rotate
a working key.

Docstrings on agent_tools()/build_agent_tools and the Anthropic adapter
state the real raise set: 401/403, post-retry 429/5xx, unreachable server.

Claude-Session: https://claude.ai/code/session_013xk3xt9KgHNTjsYmFLKxbu
Comment thread tests/test_agent_tools.py

def test_handshake_failure_blames_the_key_only_on_auth_statuses(monkeypatch):
"""A rate-limited or failing handshake is not a key problem."""
import types
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant