Skip to content

feat: eval baseline arm — measure LCD vs bare Claude Code - #12

Merged
stid merged 5 commits into
mainfrom
feat/eval-baseline-arm
Aug 23, 2026
Merged

stid merged 5 commits into
mainfrom
feat/eval-baseline-arm

Conversation

@stid

@stid stid commented Aug 23, 2026

Copy link
Copy Markdown
Owner

T1 of the 2026-08 product review (dev-harness PROPOSALS-2026-08.md, work-item t1-baseline-eval, decision D-025).

Every results row so far compares LCD variants against each other; the headline claim — context economy pays against a plain agent — has never been measured. This PR adds the missing arm.

What

  • evals/run-eval.sh --baseline: bare-agent arm — no plugin, workspace stripped of the fixture's LCD onboarding (strip verified fail-closed), driven by a new frozen prompt whose feature paragraph is byte-identical to the LCD prompt's (test-enforced), EVAL-DONE finish line. --baseline + --plugin-dir refused as a contradiction.
  • evals/grade.sh --baseline: identical row shape with audit: n/a; fail-closed both directions; golden-locked criteria untouched (golden test passes unedited).
  • Symmetric config isolation for BOTH arms (recon vs current CLI docs): user-level plugins and ~/.claude/CLAUDE.md leak into claude -p by default, and --plugin-dir adds rather than replaces — every model call now runs with a fresh CLAUDE_CONFIG_DIR + --settings '{"enabledPlugins":[]}'. Arms differ by exactly one flag.
  • tests/test-eval-baseline.sh: red-first fixture tests for AC-1..AC-4 (own file, so the old eval-harness work-item's identical AC literals can't green this audit).

Notes

  • Version bump 0.16.0 → 0.17.0 lockstep in the branch's first commit.
  • Campaign rows + conclusion land as a follow-up PR once ≥3 runs per arm complete (evidence-in-PR policy — no eval in CI).

🤖 Generated with Claude Code

https://claude.ai/code/session_01VRajFkyg9aBwbUFMhVaxNe

stid and others added 5 commits August 22, 2026 17:25
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRajFkyg9aBwbUFMhVaxNe
Red-first per the Deep-lane pipeline (work-item t1-baseline-eval in the dev harness).
Covers: --baseline dry-run + symmetric isolation contract (AC-1), workspace strip /
fail-closed residue + flag contradiction (AC-2), grader baseline row shape + two-way
fail-closed (AC-3), cross-arm row comparability + frozen-prompt feature-paragraph
parity (AC-4). AC-5 is a harness-closeout criterion, verified at audit/closeout, not
unit-testable here.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRajFkyg9aBwbUFMhVaxNe
…olation in run-eval.sh

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRajFkyg9aBwbUFMhVaxNe
…aseline, two-way fail-closed

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRajFkyg9aBwbUFMhVaxNe
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRajFkyg9aBwbUFMhVaxNe
@stid
stid merged commit ca41b35 into main Aug 23, 2026
3 checks passed
@stid
stid deleted the feat/eval-baseline-arm branch August 23, 2026 00:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant