Repository navigation
Configure OpenClaw Forge full-brief evaluation - #101
Merged
sanafayyaz315 merged 7 commits intoOct 6, 2026
Merged
Conversation
sanafayyaz315
force-pushed
the
sana/morning-brief-eval
branch
from
September 29, 2026 09:43
5ef0f35 to
947be68
Compare
sanafayyaz315
marked this pull request as ready for review
September 29, 2026 10:43
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
openclaw-forgesubmission's GLM-5-3-Flash agent withmaxTokens: 128000and set its separate preflight request to 512 tokens.brief.json, reject OpenClaw's no-final-summary fallback, and require a brief with anevidenceIdandscope: fullfor the morning-briefing case.llm-base-urloverrides in the existing OpenShell PipelineRun example; note where to change its sandbox image digest.Why
The Flash output budget was too small for a larger briefing run. The evaluation also needs to distinguish a completed full brief from an attention brief or a fallback response. These settings are confined to the
openclaw-forgesubmission; the shared Pipeline and Evaluate Task are unchanged.Verification
scope: fulland false forscope: attentionor missing scope.Scope
The branch excludes the test-only inbox-cap parameter, morning-only case copy, DEBUG/raw-stream settings, personal change log, and
saw-mpkGateway PipelineRun template. The template remains available locally for testing. Thepublished_briefcheck verifies full scope and an evidence ID; it does not score the brief's content. Content-quality evaluation is the next step.Companion harness PR: GuyZivRH/agent-eval-harness#9