Skip to content

fix(analyzer): tolerate single-number stepRange in analysis schema - #72

Merged
Treelovah merged 2 commits into
mainfrom
fix/analyzer-steprange-preprocess
Sep 17, 2026
Merged

Treelovah merged 2 commits into
mainfrom
fix/analyzer-steprange-preprocess

Conversation

@r3y3r53

@r3y3r53 r3y3r53 commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator

Problem

Follow-up to #71 (merged). While running the repair pass over the stored parseFailed analyses from the Sep/Oct 2026 benchmark sweeps, the actual root cause of the largest failure bucket turned out not to be string-typed numbers: 17 of 21 failed analyses died on attackChain.phases[].stepRange[1] being undefined — analyzer LLMs intermittently emit [3] or a bare number instead of [3, 7], and the strict z.tuple([z.number(), z.number()]) rejected the whole analysis (invalid_type, expected number, received undefined).

Fix

z.preprocess on stepRange that normalizes before validation:

  • [n] or scalar n -> [n, n]
  • mixed arrays -> [first, Math.max(...nums)]
  • non-numeric / absent -> [0, 0]

description and techniques gained defaults so a terse phase object doesn't nuke its parent.

Verified

Impact

With both fixes, an analyzer that mis-shapes stepRange (the single most common analyzer quirk observed) no longer throws away a valid flag-capture analysis. Expect near-zero FLAG*-without-KSM cells in future sweeps.

LLM analyzers intermittently emit numeric scores as JSON strings
("85" instead of 85). AnalysisResponseSchema used strict z.number(),
so a single string-typed score rejected the entire analysis with
invalid_type => parseFailed: true and the KSM score was discarded
even though the run legitimately captured the flag.

Observed in the Sep/Oct 2026 benchmarking sweeps with DeepSeek-V3 as
analyzer: multiple flag-captured runs (e.g. GLM-5.3-Flash on
indirect-prompt-injection, Kimi-K3 on confused-deputy-email-agent)
were recorded without a KSM solely because of this strictness -
results/*.analysis.json show parseFailed: true with
"expected number, invalid_type" on a numeric field.

Switch the analysis-response score fields (behavior.decisionQuality,
strategy.*, rubricEvaluation.qualitative.*.score, milestones[].achieved)
to z.coerce.number()/z.coerce.boolean() so string-typed numerics are
accepted, while real type errors (non-numeric text) still fail.
Plain numbers pass through unchanged.

Verified against the stored failing analyses: this would have salvaged
the parseFailed runs without re-running them.
17 of 21 parseFailed analyses from the Sep/Oct 2026 benchmark sweeps
failed on attackChain.phases[].stepRange[1] being undefined: analyzer
LLMs intermittently emit [3] or a bare number instead of [3, 7].
The strict z.tuple([z.number(), z.number()]) rejected the whole
analysis for that (invalid_type, not a string-number, hence not
covered by the earlier coercion fix).

Add a z.preprocess on stepRange that normalizes the value before
validation: [n] or n -> [n, n], mixed arrays -> [min-aware first,
max tail] via Math.max, non-numeric -> [0, 0]. Confirmed against the
17 stored failing analyses; combined with the coercion commit this
raises analysis recovery from 0/21 to 16/21 (only genuinely invalid
JSON payloads remain unrecoverable).

@Treelovah Treelovah left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Product review: stepRange preprocess is the real win (17/21 parseFailed). Includes #71 coercion so both land together if #71 not yet on main. Defaults on description/techniques are safe. Approve.

@Treelovah
Treelovah merged commit af399b9 into main Sep 17, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants