Skip to content

fix(analyzer): give analysis calls a dedicated 300s timeout - #78

Open
r3y3r53 wants to merge 1 commit into
mainfrom
fix/analyzer-timeout-config
Open

r3y3r53 wants to merge 1 commit into
mainfrom
fix/analyzer-timeout-config

Conversation

@r3y3r53

@r3y3r53 r3y3r53 commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator

Problem

The analyzer prompt embeds the entire attack chain — up to 45 steps of reasoning, commands, and truncated output. It is therefore far larger and slower to process than a single benchmark step, yet both analyzer call sites invoked withRateLimitRetry() without a timeoutMs argument, silently inheriting the generic DEFAULT_API_TIMEOUT_MS of 120s.

On long-running benchmarks this surfaced as:

→ ✖ Analysis failed
  Analysis: timed out after 120s
  Retry later with: oasis analyze 5512af4f

The run captured its flag successfully, but the KSM score is thrown away — so oasis report shows a flag with no score, and the benchmark table gets an unfillable gap.

Evidence

Measured across a local corpus of 276 benchmark logs:

Metric Count
Log files 276
Logs that captured a flag 71
Logs that failed analysis with timed out after 120s 10
Logs with both a flag and a KSM score 57

That is 14 flag captures with no KSM score recorded, and 10 of those trace directly to this 120s ceiling. The failing runs are consistently the complex ones (high iteration counts, large transcripts) — exactly the ones whose analysis is most valuable.

Change

Add a dedicated analyzer budget and pass it at both call sites.

src/lib/constants.ts

// Analyzer LLM call timeout. The analyzer prompt embeds a full attack chain
// (up to 45 steps), which can be far larger than a benchmark step, so it needs
// more headroom than the generic 120s API default.
export const ANALYZER_TIMEOUT_MS = 300_000;   // 5 minutes

src/lib/analyzer.ts — both callAnthropicAnalyzer and callOpenAIAnalyzer:

const response = await withRateLimitRetry(
  () => client.messages.create({ /* ... */ }),
  'Analysis',
  false,
  ANALYZER_TIMEOUT_MS,   // ← was omitted, defaulting to 120s
);

This is deliberately not a redesign. withRateLimitRetry() already provides retries with exponential backoff and Retry-After handling, and parseAnalysisResponse() already has a structured fallback path — the only real defect was the missing timeout argument.

Scope note

The generic DEFAULT_API_TIMEOUT_MS is left untouched, so benchmark step calls in runner.ts keep their existing 120s behaviour. Only analysis gets the larger budget.

Tests

Added to tests/unit/analyzer.test.ts:

  • asserts ANALYZER_TIMEOUT_MS > DEFAULT_API_TIMEOUT_MS
  • asserts it is a positive finite number of ms

Verification

npm run typecheck   # clean
npm run test        # 421 passed (16 files)
npm run build       # clean

Before: 419 tests. After: 421 tests, all passing.

Risk

Low. A longer analysis timeout cannot alter result correctness — it only allows slow analyses to finish instead of being discarded. The tradeoff is that a genuinely hung analysis now takes 300s instead of 120s before failing, which is bounded and acceptable for an offline reporting step.

Closes #75

The analyzer embeds a full attack chain (up to 45 steps) in its prompt, so
it can take far longer than a benchmark step. Both analyzer call sites used
withRateLimitRetry() without a timeoutMs argument, silently inheriting the
generic DEFAULT_API_TIMEOUT_MS of 120s.

On long-running benchmarks this surfaced as:

  Analysis: timed out after 120s

which discards an otherwise-valid KSM score. In our corpus of 276 benchmark
logs, 10 runs failed analysis this way and 14 flag captures ended up with no
KSM score recorded.

Add an ANALYZER_TIMEOUT_MS constant (300s) and pass it at both the Anthropic
and OpenAI-compatible call sites, plus tests asserting the analyzer budget
exceeds the generic API default.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Implement progressive analyzer timeout based on run complexity

1 participant