Skip to content

Configure OpenClaw Forge full-brief evaluation - #101

Merged
sanafayyaz315 merged 7 commits into
RHEcosystemAppEng:mainfrom
sanafayyaz315:sana/morning-brief-eval
Oct 6, 2026
Merged

sanafayyaz315 merged 7 commits into
RHEcosystemAppEng:mainfrom
sanafayyaz315:sana/morning-brief-eval

Conversation

@sanafayyaz315

@sanafayyaz315 sanafayyaz315 commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Configure the openclaw-forge submission's GLM-5-3-Flash agent with maxTokens: 128000 and set its separate preflight request to 512 tokens.
  • Collect the published brief.json, reject OpenClaw's no-final-summary fallback, and require a brief with an evidenceId and scope: full for the morning-briefing case.
  • Show commented, optional submission-repo, harness-repo/revision, and prepare-stage llm-base-url overrides in the existing OpenShell PipelineRun example; note where to change its sandbox image digest.

Why

The Flash output budget was too small for a larger briefing run. The evaluation also needs to distinguish a completed full brief from an attention brief or a fallback response. These settings are confined to the openclaw-forge submission; the shared Pipeline and Evaluate Task are unchanged.

Verification

  • The changed YAML files parse successfully, and the full-brief check returns true for scope: full and false for scope: attention or missing scope.
  • Pre-commit passed on the final commit.
  • Earlier Gateway evaluations with test-specific run configuration produced full briefs at caps 30, 100, and 400; the cap-400 run found 104 available emails. These runs predate this cleaned PR commit and do not establish content quality or a live rerun of this exact diff.

Scope

The branch excludes the test-only inbox-cap parameter, morning-only case copy, DEBUG/raw-stream settings, personal change log, and saw-mpk Gateway PipelineRun template. The template remains available locally for testing. The published_brief check verifies full scope and an evidence ID; it does not score the brief's content. Content-quality evaluation is the next step.

Companion harness PR: GuyZivRH/agent-eval-harness#9

@sanafayyaz315
sanafayyaz315 force-pushed the sana/morning-brief-eval branch from 5ef0f35 to 947be68 Compare September 29, 2026 09:43
@sanafayyaz315 sanafayyaz315 changed the title Add morning briefing Gateway evaluation configuration Configure OpenClaw Forge full-brief evaluation Sep 29, 2026
@sanafayyaz315
sanafayyaz315 marked this pull request as ready for review September 29, 2026 10:43

@GuyZivRH GuyZivRH left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@sanafayyaz315
sanafayyaz315 merged commit 58e14a9 into RHEcosystemAppEng:main Oct 6, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants