Skip to content

Add controlled Forge eval suites and verified pipeline reporting - #99

Open
tarun-etikala wants to merge 3 commits into
RHEcosystemAppEng:mainfrom
tarun-etikala:codex/strengthen-forge-evals
Open

tarun-etikala wants to merge 3 commits into
RHEcosystemAppEng:mainfrom
tarun-etikala:codex/strengthen-forge-evals

Conversation

@tarun-etikala

@tarun-etikala tarun-etikala commented Sep 24, 2026 •

Copy link
Copy Markdown

Problem and result

The Forge evaluation setup needed a repeatable way to distinguish infrastructure failures, invalid evaluations, and actual skill failures. Missing artifacts or skipped storage could appear successful, and comparisons could mix different inputs or omit child usage. This adds a concrete controlled-evaluation workflow with preserved evidence and correctness-before-cost reporting. It does not change Forge skills or upgrade the live application.

Requires companion harness PR: GuyZivRH/agent-eval-harness#8 (targets the existing feat/aeh-openshell-openclaw integration branch).

Review scope: working integration, phase 1

This draft captures the pipeline path validated in saw-nommen, including the cases and tooling used to prove it. Review failure propagation, artifact durability, source reproducibility, and the result gates. Moving Forge-owned evaluation content into rh-forge/forge-eval is the next phase, not part of this draft. Namespace configuration stays outside the repository. The source-overlay trigger remains development tooling, not the proposed final CI interface.

For navigating the diff: 85 of the 114 changed files are case-fixture inputs. Start with scripts/publish.py, abevalflow/artifact_storage.py, scripts/forge_gate.py, and Docs/forge_controlled_evals.md; then review the trigger and suite expectations.

Changes

  • Add reusable trigger, report-fetch, and explicitly scoped cleanup commands plus a namespace-neutral runbook. Keep the configured base PipelineRun local instead of committing personal endpoints and Secret references. Clone installed Tasks under immutable content-addressed names; preserve the existing pipeline. Pin source revisions, verify overlay hashes, and bound pre-agent Git fetch retries.
  • Add four separate suites: runtime smoke, five sealed briefing cases, eleven drafting lifecycle cases, and governed mock M365 collection. Sealed cases use fixed time and synthetic identity; expected answers remain outside the agent workspace. Fixture generation validates the image-owned evidence schema.
  • Require canonical output and complete usage. Distinguish infrastructure/validity/quality failures; preserve artifacts even when quality fails.
  • For controlled Forge runs, explicitly require durable artifacts (AEH_REQUIRE_ARTIFACTS=1 / --require-aeh-artifacts); preserve optional best-effort storage for other AEH callers. Fail strict storage on missing configuration or artifacts; upload and SHA-256 read-back verify every object. Run the final gate after storage and require its attestation plus FINISHED MLflow reporting. Reclaim resources only after durable evidence is verified.
  • Add predeclared paired experiments with immutable skill inputs, fresh state, alternating arm order, and provenance checks. Report input/output/reasoning/cache tokens, requests, tool calls, and time. Enforce correctness before the no-output-token-increase gate; missing/confounded runs cannot count as improvement. Separate regression acceptance (unchanged passing results permitted) from improvement benchmarking (measured quality gain required). Both retain critical correctness and no-output-token-increase gates. Propagate the comparison exit code through the experiment runner. No automatic skill promotion or outcome-dependent reruns.

Validation

  • Full local suite: 1,166 passed; ruff check abevalflow/ scripts/ tests/ passed.

  • saw-nommen: final smoke forge-controlled-ts9xd passed all six tasks; canonical briefing forge-controlled-vlztl and mock M365 forge-controlled-wmgdl passed end to end.

  • Verified storage read-back, retained evidence on failed runs, and actual PVC reclamation after explicitly archiving a completed smoke run's logs.

  • Fault cases exposed Git download failures, missing/incomplete model output, invalid fixtures, and confounded comparisons; these are preserved as failures rather than converted into successful benchmarks.

  • Cluster revalidation on 2026-09-25 using published flow ae6c2aa / harness 5b340bd: forge-controlled-62wml passed all five corrected briefing cases and all six pipeline tasks. Complete usage was recorded for each case; storage attestation verified 141 objects, MLflow reporting passed, and the final gate reported no issues.
    Successful full briefing run

  • Negative run forge-controlled-csqlj rejected the unsupported evidence reason before model preflight or agent execution (invalid_eval, no trajectory files). Its analysis/storage succeeded with 13 verified objects and the final gate failed as intended. A stalled source fetch recovered on bounded retry 1/3.

  • Cleanup validation forge-controlled-zwh4s exercised the explicit strict-storage opt-in with overlay f1b0805a0aba: malformed evidence was rejected before model execution, storage verified 13 objects, and the final gate failed as expected. No trajectory files were produced. The full five-case success above validates the preceding integration; this follow-up specifically validates the cleanup path.

Draft review limits

  • The earlier breadth run forge-controlled-2b7t6 is preserved as failed due to its malformed fixture; the corrected full-suite and negative runs above close the final evidence-schema cluster-validation gap.
  • The paired skill benchmark stopped on a pre-agent Git fetch failure after one valid pair. There is no skill-improvement or token-reduction claim here; skill outcomes are sample pipeline reports.
  • Mock Slack integration remains to be configured. Deterministic judges cover this delivery; nuanced tone/synthesis scoring is deferred.
  • This extends an installed AEH/OpenShell pipeline. A new namespace still needs its own egress, current gateway mTLS, Forge CA, synthetic identity, storage, and private-image access. Its configured base PipelineRun stays outside the repository.
  • The trigger packages local source in a ConfigMap and rewrites copies of installed Task scripts. It provides reproducible development runs through recorded hashes, but this mechanism should be replaced for routine CI.
  • The default comparison policy remains improvement for compatibility; choose and freeze --policy regression for a regression check. Invalid, incomplete, or inconclusive comparisons still fail the command.

Follow-up: Forge Eval ownership and CI

  1. Move Forge suites, fixtures, validators, and comparison policy into rh-forge/forge-eval; retain reusable scheduling/storage reliability fixes in Flow and execution support in AEH.
  2. Replace source overlays/Task rewrites with pinned published dependencies and supported pipeline parameters, backed by validated environment configuration.
  3. Run from a fresh CI checkout without local overlays, then enable trusted manual dispatch, followed by trusted PR checks and scheduled broad suites. Keep expensive repeated comparisons and optional LLM judges explicitly budgeted.

No repository migration is included in this phase. The cluster proof is preserved so the next phase can demonstrate equivalent behavior rather than rebuild the pipeline without a reference.

Please review task integration, artifact/failure handling, fixture expectations, and comparison gates. Raw run artifacts and credentials are excluded.

@GuyZivRH GuyZivRH left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking
​scripts/trigger_forge_controlled.py:248-258
The generated inline gate TaskSpec references:
$(params.enable-mlflow)
$(params.mlflow-tracking-uri)
but the TaskSpec declares no params, and the generated Pipeline task does not pass any. These are Pipeline-level parameters, not automatically available as TaskSpec parameters. The gate may fail Tekton validation or receive unresolved substitutions.
Fix by declaring these as TaskSpec parameters and passing them from the generated Pipeline task:
taskSpec:
params:
- name: enable-mlflow
- name: mlflow-tracking-uri
and add corresponding task params using $(params.enable-mlflow) and $(params.mlflow-tracking-uri).
Validation

  • 62 relevant Forge/publish tests passed.
  • The controlled-evaluation design, immutable overlays, paired comparison, durable-storage checks, and fixture coverage are otherwise strong.
  • The removed personal deployment manifest is appropriate.

@tarun-etikala
tarun-etikala force-pushed the codex/strengthen-forge-evals branch from d0bdbfa to 685c6be Compare September 29, 2026 16:05
@tarun-etikala
tarun-etikala marked this pull request as ready for review September 29, 2026 16:20

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants