feat: user-specified retries for transient API failures - #155
Merged
Merged
Conversation
timzsu
force-pushed
the
zsu/api-executor-retry
branch
from
September 23, 2026 14:54
a8cc9c6 to
04639c2
Compare
timzsu
marked this pull request as ready for review
September 23, 2026 14:57
timzsu
added this pull request to stack #156
September 24, 2026 04:55
The API executor marks transient failures (5xx, 408, 429, connection errors) as retryable but never retried them. Add a spec.api.retries field (default 0, preserving current behavior) that controls how many times a transient failure is re-issued before the task fails, with a fixed 1s backoff between attempts. Non-retryable 4xx statuses and cancelled tasks are never retried. Co-Authored-By: Claude Code <noreply@anthropic.com> Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu>
timzsu
force-pushed
the
zsu/api-executor-retry
branch
from
September 27, 2026 01:13
04639c2 to
22b222e
Compare
…during backoff A cancellation left over from a previous task no longer leaks into the next task on a reused warm executor: cancel() records the task it targets, and run() clears a cancellation addressed to a different task while one aimed at the task now starting still stands, all under a single lock so a racing cancel is never lost. The retry backoff now waits on the cancel event instead of sleeping, so a cancelled task raises its cancellation as soon as it is signalled rather than sitting out the full backoff. Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu> Co-Authored-By: Claude Code <noreply@anthropic.com>
Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu> Co-Authored-By: Claude Code <noreply@anthropic.com>
The interrupt monitor checks the runner's current task id and then calls executor.cancel(task_id) without a lock spanning both, so a late cancellation for a finished task can reach a warm executor that has since started another task. The executor now records the active task id under its cancel lock, sets the cancel event only when no run is in flight or the cancellation matches the active task, and clears the active id when the run returns or raises. Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu> Co-Authored-By: Claude Code <noreply@anthropic.com>
Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu> Co-Authored-By: Claude Code <noreply@anthropic.com>
A single recorded cancellation id let a late cancel for a prior task overwrite a recorded cancel for the task about to start, so that task ran despite its own cancellation. Track pending cancelled ids in a set guarded by the cancel lock; run() consumes only its own id and clears the rest as stale. Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu> Co-Authored-By: Claude Code <noreply@anthropic.com>
kaiitunnz
requested changes
Sep 27, 2026
kaiitunnz
left a comment
Collaborator
There was a problem hiding this comment.
A few comments. PTAL.
Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu>
kaiitunnz
requested changes
Sep 27, 2026
kaiitunnz
left a comment
Collaborator
There was a problem hiding this comment.
One more minor issue.
Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu>
4 of 8 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
The API executor marks transient failures (5xx, 408, 429, connection errors) as
retryable=Truebut never retried them — a request failed immediately on the first transient error. In production the LLM serving endpoint intermittently returns 504 Gateway Timeout on large requests; the same request succeeds when retried. A user-specified retry count lets the executor re-issue and succeed.Changes
spec.api.retries(new field, default0preserving current behavior) — how many times a transient failure is re-issued before the task fails._request_with_retries— wraps the request in a retry loop with a fixed 1s backoff between attempts. A retryable failure (connection error, 5xx, 408, 429) retries up toretriestimes; non-retryable 4xx and cancelled tasks stop immediately.Design
The retry is per-request, preserving the executor's concurrency model. The final attempt's failure propagates to the caller — a request that exhausts retries still fails loudly, never silently swallowed.
Test Plan
The repo's own pre-commit gate over every changed file, and the full suite on the repo's interpreter.
Test Result
E2E against a live serving endpoint. A FlowMesh worker running this change serves Lumilake's paper-ingestion runs against
qwen3.8-27bwithspec.api.retries: 3. Its per-call log shows both outcomes:Pre-submission Checklist
pre-commit run --all-filesand fixed any issues.uv run pytest tests/passes locally.uv sync --all-packages --group ci --frozen).[BREAKING]and described migration steps above.