Conversation
The analyzer embeds a full attack chain (up to 45 steps) in its prompt, so it can take far longer than a benchmark step. Both analyzer call sites used withRateLimitRetry() without a timeoutMs argument, silently inheriting the generic DEFAULT_API_TIMEOUT_MS of 120s. On long-running benchmarks this surfaced as: Analysis: timed out after 120s which discards an otherwise-valid KSM score. In our corpus of 276 benchmark logs, 10 runs failed analysis this way and 14 flag captures ended up with no KSM score recorded. Add an ANALYZER_TIMEOUT_MS constant (300s) and pass it at both the Anthropic and OpenAI-compatible call sites, plus tests asserting the analyzer budget exceeds the generic API default.
r3y3r53
force-pushed
the
fix/analyzer-timeout-config
branch
from
September 18, 2026 14:28
a38ea2a to
714e616
Compare
5 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The analyzer prompt embeds the entire attack chain — up to 45 steps of reasoning, commands, and truncated output. It is therefore far larger and slower to process than a single benchmark step, yet both analyzer call sites invoked
withRateLimitRetry()without atimeoutMsargument, silently inheriting the genericDEFAULT_API_TIMEOUT_MSof 120s.On long-running benchmarks this surfaced as:
The run captured its flag successfully, but the KSM score is thrown away — so
oasis reportshows a flag with no score, and the benchmark table gets an unfillable gap.Evidence
Measured across a local corpus of 276 benchmark logs:
timed out after 120sThat is 14 flag captures with no KSM score recorded, and 10 of those trace directly to this 120s ceiling. The failing runs are consistently the complex ones (high iteration counts, large transcripts) — exactly the ones whose analysis is most valuable.
Change
Add a dedicated analyzer budget and pass it at both call sites.
src/lib/constants.tssrc/lib/analyzer.ts— bothcallAnthropicAnalyzerandcallOpenAIAnalyzer:This is deliberately not a redesign.
withRateLimitRetry()already provides retries with exponential backoff andRetry-Afterhandling, andparseAnalysisResponse()already has a structured fallback path — the only real defect was the missing timeout argument.Scope note
The generic
DEFAULT_API_TIMEOUT_MSis left untouched, so benchmark step calls inrunner.tskeep their existing 120s behaviour. Only analysis gets the larger budget.Tests
Added to
tests/unit/analyzer.test.ts:ANALYZER_TIMEOUT_MS > DEFAULT_API_TIMEOUT_MSVerification
Before: 419 tests. After: 421 tests, all passing.
Risk
Low. A longer analysis timeout cannot alter result correctness — it only allows slow analyses to finish instead of being discarded. The tradeoff is that a genuinely hung analysis now takes 300s instead of 120s before failing, which is bounded and acceptable for an offline reporting step.
Closes #75