Skip to content

docs(evals): T1 campaign results — LCD vs bare agent (single-shot) - #13

Merged
stid merged 3 commits into
mainfrom
docs/t1-campaign-results
Aug 29, 2026
Merged

stid merged 3 commits into
mainfrom
docs/t1-campaign-results

Conversation

@stid

@stid stid commented Aug 23, 2026

Copy link
Copy Markdown
Owner

Campaign rows + conclusion for the first baseline-arm run (v0.17.0 harness). 3 runs/arm, claude-fable-5, settings-only isolation (config-dir isolation broke auth on the host — noted in the conclusion).

Headline (single-shot, this task class): bare agent ~4.5× cheaper / ~5.8× fewer tokens, all 6 runs green, rubric quality near-parity (lcd 69/75, none 72/75). LCD's multi-session claims remain unmeasured — follow-up longitudinal benchmark (T1b) proposed in the dev harness.

Published per the evidence-in-PR policy; unfavorable direction included by design (the harness is an instrument, not a marketing device).

🤖 Generated with Claude Code

https://claude.ai/code/session_01VRajFkyg9aBwbUFMhVaxNe

…per single-shot, quality near-parity

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRajFkyg9aBwbUFMhVaxNe
@stid
stid marked this pull request as draft August 23, 2026 01:41
@stid

stid commented Aug 23, 2026

Copy link
Copy Markdown
Owner Author

HOLD — do not merge. The independent closeout evaluator refuted the campaign's isolation claim: --settings '{"enabledPlugins":[]}' is shape-invalid (correct schema is an object map with explicit false per plugin) and the fallback wrapper had unset CLAUDE_CONFIG_DIR, so all three none rows ran with the lcd plugin loaded (transcript evidence: lcd skills advertised in the baseline sessions). The conclusion's isolation statement is wrong as written. Fix + verified re-run to follow; this PR will be updated or superseded.

…ows + rewritten conclusion

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRajFkyg9aBwbUFMhVaxNe
@stid
stid marked this pull request as ready for review August 29, 2026 01:11
….5x) — evaluator round-2 finding

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRajFkyg9aBwbUFMhVaxNe
@stid
stid merged commit 05d9215 into main Aug 29, 2026
3 checks passed
@stid
stid deleted the docs/t1-campaign-results branch August 29, 2026 01:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant