Repository navigation
Add controlled Forge eval suites and verified pipeline reporting - #99
Open
tarun-etikala wants to merge 3 commits into
Open
tarun-etikala wants to merge 3 commits into
tarun-etikala wants to merge 3 commits into
Conversation
GuyZivRH
requested changes
Sep 29, 2026
GuyZivRH
left a comment
Collaborator
There was a problem hiding this comment.
Blocking
scripts/trigger_forge_controlled.py:248-258
The generated inline gate TaskSpec references:
$(params.enable-mlflow)
$(params.mlflow-tracking-uri)
but the TaskSpec declares no params, and the generated Pipeline task does not pass any. These are Pipeline-level parameters, not automatically available as TaskSpec parameters. The gate may fail Tekton validation or receive unresolved substitutions.
Fix by declaring these as TaskSpec parameters and passing them from the generated Pipeline task:
taskSpec:
params:
- name: enable-mlflow
- name: mlflow-tracking-uri
and add corresponding task params using
Validation
- 62 relevant Forge/publish tests passed.
- The controlled-evaluation design, immutable overlays, paired comparison, durable-storage checks, and fixture coverage are otherwise strong.
- The removed personal deployment manifest is appropriate.
tarun-etikala
force-pushed
the
codex/strengthen-forge-evals
branch
from
September 29, 2026 16:05
d0bdbfa to
685c6be
Compare
tarun-etikala
marked this pull request as ready for review
September 29, 2026 16:20
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem and result
The Forge evaluation setup needed a repeatable way to distinguish infrastructure failures, invalid evaluations, and actual skill failures. Missing artifacts or skipped storage could appear successful, and comparisons could mix different inputs or omit child usage. This adds a concrete controlled-evaluation workflow with preserved evidence and correctness-before-cost reporting. It does not change Forge skills or upgrade the live application.
Requires companion harness PR: GuyZivRH/agent-eval-harness#8 (targets the existing
feat/aeh-openshell-openclawintegration branch).Review scope: working integration, phase 1
This draft captures the pipeline path validated in
saw-nommen, including the cases and tooling used to prove it. Review failure propagation, artifact durability, source reproducibility, and the result gates. Moving Forge-owned evaluation content intorh-forge/forge-evalis the next phase, not part of this draft. Namespace configuration stays outside the repository. The source-overlay trigger remains development tooling, not the proposed final CI interface.For navigating the diff: 85 of the 114 changed files are case-fixture inputs. Start with
scripts/publish.py,abevalflow/artifact_storage.py,scripts/forge_gate.py, andDocs/forge_controlled_evals.md; then review the trigger and suite expectations.Changes
AEH_REQUIRE_ARTIFACTS=1/--require-aeh-artifacts); preserve optional best-effort storage for other AEH callers. Fail strict storage on missing configuration or artifacts; upload and SHA-256 read-back verify every object. Run the final gate after storage and require its attestation plus FINISHED MLflow reporting. Reclaim resources only after durable evidence is verified.Validation
Full local suite: 1,166 passed;
ruff check abevalflow/ scripts/ tests/passed.saw-nommen: final smokeforge-controlled-ts9xdpassed all six tasks; canonical briefingforge-controlled-vlztland mock M365forge-controlled-wmgdlpassed end to end.Verified storage read-back, retained evidence on failed runs, and actual PVC reclamation after explicitly archiving a completed smoke run's logs.
Fault cases exposed Git download failures, missing/incomplete model output, invalid fixtures, and confounded comparisons; these are preserved as failures rather than converted into successful benchmarks.
Cluster revalidation on 2026-09-25 using published flow
ae6c2aa/ harness5b340bd:forge-controlled-62wmlpassed all five corrected briefing cases and all six pipeline tasks. Complete usage was recorded for each case; storage attestation verified 141 objects, MLflow reporting passed, and the final gate reported no issues.Successful full briefing run
Negative run
forge-controlled-csqljrejected the unsupported evidence reason before model preflight or agent execution (invalid_eval, no trajectory files). Its analysis/storage succeeded with 13 verified objects and the final gate failed as intended. A stalled source fetch recovered on bounded retry 1/3.Cleanup validation
forge-controlled-zwh4sexercised the explicit strict-storage opt-in with overlayf1b0805a0aba: malformed evidence was rejected before model execution, storage verified 13 objects, and the final gate failed as expected. No trajectory files were produced. The full five-case success above validates the preceding integration; this follow-up specifically validates the cleanup path.Draft review limits
forge-controlled-2b7t6is preserved as failed due to its malformed fixture; the corrected full-suite and negative runs above close the final evidence-schema cluster-validation gap.improvementfor compatibility; choose and freeze--policy regressionfor a regression check. Invalid, incomplete, or inconclusive comparisons still fail the command.Follow-up: Forge Eval ownership and CI
rh-forge/forge-eval; retain reusable scheduling/storage reliability fixes in Flow and execution support in AEH.No repository migration is included in this phase. The cluster proof is preserved so the next phase can demonstrate equivalent behavior rather than rebuild the pipeline without a reference.
Please review task integration, artifact/failure handling, fixture expectations, and comparison gates. Raw run artifacts and credentials are excluded.