fix(analyzer): tolerate single-number stepRange in analysis schema - #72
Merged
Merged
Conversation
LLM analyzers intermittently emit numeric scores as JSON strings
("85" instead of 85). AnalysisResponseSchema used strict z.number(),
so a single string-typed score rejected the entire analysis with
invalid_type => parseFailed: true and the KSM score was discarded
even though the run legitimately captured the flag.
Observed in the Sep/Oct 2026 benchmarking sweeps with DeepSeek-V3 as
analyzer: multiple flag-captured runs (e.g. GLM-5.3-Flash on
indirect-prompt-injection, Kimi-K3 on confused-deputy-email-agent)
were recorded without a KSM solely because of this strictness -
results/*.analysis.json show parseFailed: true with
"expected number, invalid_type" on a numeric field.
Switch the analysis-response score fields (behavior.decisionQuality,
strategy.*, rubricEvaluation.qualitative.*.score, milestones[].achieved)
to z.coerce.number()/z.coerce.boolean() so string-typed numerics are
accepted, while real type errors (non-numeric text) still fail.
Plain numbers pass through unchanged.
Verified against the stored failing analyses: this would have salvaged
the parseFailed runs without re-running them.
17 of 21 parseFailed analyses from the Sep/Oct 2026 benchmark sweeps failed on attackChain.phases[].stepRange[1] being undefined: analyzer LLMs intermittently emit [3] or a bare number instead of [3, 7]. The strict z.tuple([z.number(), z.number()]) rejected the whole analysis for that (invalid_type, not a string-number, hence not covered by the earlier coercion fix). Add a z.preprocess on stepRange that normalizes the value before validation: [n] or n -> [n, n], mixed arrays -> [min-aware first, max tail] via Math.max, non-numeric -> [0, 0]. Confirmed against the 17 stored failing analyses; combined with the coercion commit this raises analysis recovery from 0/21 to 16/21 (only genuinely invalid JSON payloads remain unrecoverable).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Follow-up to #71 (merged). While running the repair pass over the stored
parseFailedanalyses from the Sep/Oct 2026 benchmark sweeps, the actual root cause of the largest failure bucket turned out not to be string-typed numbers: 17 of 21 failed analyses died onattackChain.phases[].stepRange[1]being undefined — analyzer LLMs intermittently emit[3]or a bare number instead of[3, 7], and the strictz.tuple([z.number(), z.number()])rejected the whole analysis (invalid_type, expected number, received undefined).Fix
z.preprocessonstepRangethat normalizes before validation:[n]or scalarn->[n, n][first, Math.max(...nums)][0, 0]descriptionandtechniquesgained defaults so a terse phase object doesn't nuke its parent.Verified
parseFailedanalyses: recovery 0/21 -> 16/21 (both this and fix(analyzer): coerce string-quoted numbers in analysis schema (KSM loss) #71's coercion applied; the 5 remaining failures are genuinely invalid JSON from the analyzer, unrecoverable without re-prompting).npm run typecheckclean,npm run buildclean,npm run test419 passed (incl. the coercion regression tests from fix(analyzer): coerce string-quoted numbers in analysis schema (KSM loss) #71).Impact
With both fixes, an analyzer that mis-shapes stepRange (the single most common analyzer quirk observed) no longer throws away a valid flag-capture analysis. Expect near-zero FLAG*-without-KSM cells in future sweeps.