diff --git a/docs/decisions/0002-use-promptfoo-native-agent-execution.md b/docs/decisions/0002-use-promptfoo-native-agent-execution.md new file mode 100644 index 00000000..5d5d3584 --- /dev/null +++ b/docs/decisions/0002-use-promptfoo-native-agent-execution.md @@ -0,0 +1,277 @@ +# ADR 0002: Use Promptfoo native agent providers with managed disposable workspaces + +- Status: Accepted +- Date: 2026-09-21 +- Updated: 2026-09-28 + +## Decision + +Coding-agent evaluations will run directly through Promptfoo's built-in agent +providers inside one disposable local or CI job. V1 will not introduce an +AllAgents execution gateway, a custom Promptfoo provider, or a separate runner +service. + +Both supported providers receive the same fixed job-private working directory: + +```yaml +working_dir: ./.eval/workspace +``` + +A Promptfoo lifecycle extension owns the mutable directory around every +evaluation row: + +1. job bootstrap resolves and materializes exact source inputs into immutable, + job-private seeds; +2. `beforeEach` removes any previous workspace and creates a private copy of the + selected seed at `.eval/workspace`; +3. Promptfoo invokes the selected built-in agent provider in that directory; +4. deterministic JavaScript assertions inspect the final filesystem and run + declared checks before teardown; +5. `afterEach` records bounded diagnostic metadata and removes the workspace; +6. `afterAll` removes remaining job-private evaluation state. + +Setup fails closed in `beforeEach`: any reset or copy error throws before the +provider runs. Promptfoo currently catches and logs `afterEach` failures, so an +`afterEach` exception alone cannot fail the evaluation command. The extension +records a failure sentinel, and the surrounding job wrapper checks that sentinel +and workspace absence before accepting the run. `beforeEach` always deletes the +fixed workspace before copying a seed, even when the previous cleanup appeared +successful. + +Promptfoo runs these rows serially and without its response cache: + +```yaml +evaluateOptions: + maxConcurrency: 1 + cache: false +``` + +Serial execution makes one fixed working directory unambiguous. Parallelism, if +needed later, is job-level: each disposable job receives its own `.eval` root. + +Promptfoo remains the evaluation system of record. It owns prompts, provider and +model matrices, repetitions, assertions, scores, pass/fail decisions, traces, +and reports. The workspace extension owns only setup, reset, bounded diagnostic +collection, and cleanup. + +## Why + +Promptfoo already implements the agent-facing behavior the earlier gateway +design planned to recreate: + +- the built-in Claude Agent SDK and Codex SDK providers both accept + `working_dir`, resolved relative to the configuration file; +- Promptfoo documents extension hooks for `beforeAll`, `beforeEach`, + `afterEach`, and `afterAll`; +- Promptfoo's own write-capable Claude example combines an extension-managed + workspace with `maxConcurrency: 1`; +- external JavaScript assertions can inspect the final workspace and return + structured grading results; +- built-in providers return final output, usage, provider metadata, and session + identifiers where supported; and +- Promptfoo emits and ingests OpenTelemetry traces and projects recognized tool + spans into `trajectory:*` assertions. + +The proposed gateway added an HTTP API, custom provider, queue, database, +artifact service, source cache, worker state machine, custom agent adapters, +ATIF conversion, sandbox implementation, cancellation protocol, and recovery +semantics. None of those components is necessary for a trusted, single-tenant +evaluation that already runs inside a disposable job. + +Supporting evidence and provider-specific limits are recorded in +[Promptfoo native agent workspaces](../research/promptfoo-native-agent-workspaces.md). +The implementation sequence is in the +[Promptfoo coding-agent evaluation plan](../plans/2026-09-18-0837-feat-promptfoo-coding-agent-evals-plan.md). + +## Execution boundary + +The disposable job is the outer lifecycle and isolation boundary. It may be a +CI job, rootless container, or VM. The job: + +- starts without mutable state from another evaluation job; +- receives only the source and model credentials required for that job; +- materializes exact source revisions before starting Promptfoo; +- runs Promptfoo and all assertions; +- exports the requested Promptfoo results, traces, and bounded diagnostics; and +- destroys the complete job filesystem and process tree when finished. + +Promptfoo provider sandboxes constrain agent operations but are not a substitute +for a hostile multi-tenant execution boundary. Write-capable or adversarial +evaluations must run in a disposable container or VM rather than directly on a +developer workstation. + +Source acquisition occurs before agent execution. Acquisition credentials must +not be copied into `.eval/seeds`, `.eval/workspace`, result metadata, or trace +attributes. Job bootstrap removes them from the environment before Promptfoo +starts whenever the source transport permits that separation. + +## Promptfoo configuration + +The initial provider matrix uses Promptfoo's providers directly: + +```yaml +providers: + - id: anthropic:claude-agent-sdk + config: + working_dir: ./.eval/workspace + append_allowed_tools: [Write, Edit, MultiEdit, Bash] + permission_mode: acceptEdits + persist_session: false + sandbox: + enabled: true + failIfUnavailable: true + + - id: openai:codex-sdk + config: + working_dir: ./.eval/workspace + sandbox_mode: workspace-write + approval_policy: never + enable_streaming: true + persist_threads: false + +extensions: + - file://extensions/workspace.cjs:workspaceLifecycle + +evaluateOptions: + maxConcurrency: 1 + cache: false + +tracing: + enabled: true + otlp: + http: {} +``` + +Provider-specific permissions remain explicit. A common working directory does +not imply identical tool or sandbox behavior. + +Codex `enable_streaming` is enabled because Promptfoo uses its SDK events to +emit provider-level command, file, search, MCP, and turn spans. Deep native +tracing is optional, not the default: it can expose additional payloads and, for +Codex, disables thread persistence. + +## Workspace contract + +`.eval/seeds` contains resolved inputs for the current job, published through a +read-only mount or owned by a bootstrap identity that the unprivileged +Promptfoo/agent user cannot modify. `.eval/workspace` is always disposable and +writable. `.eval/artifacts` may contain explicitly selected bounded diagnostics. + +The source manifest records the requested identity and resolved immutable +identity for each seed. A mutable Git ref may be an input to resolution but is +never the recorded resolved identity. OCI input, if used, records the verified +manifest digest. + +The workspace copy must not use hardlinks or any writable alias back to a seed. +A reflink or another copy-on-write primitive is acceptable only when later +writes cannot mutate the seed. A full recursive copy is the portable fallback. + +Every row starts from the same selected seed state. Workspace mutations, +installed dependencies, generated files, and provider session state must not +cross row boundaries. Cross-row provider thread persistence is disabled in V1. +Agent-started background services are unsupported because the extension cannot +guarantee process-tree cleanup between rows; disposable job teardown is the +process cleanup boundary. + +## Assertions and evidence + +Behavioral success is decided by Promptfoo assertions, not lifecycle hooks. +Rows that inspect the live workspace use deterministic assertions only. +Promptfoo may defer a row's complete assertion set when model-graded assertions +are present; a later `beforeEach` could then replace the shared workspace before +the earlier filesystem assertion executes. + +For a deterministic-only row, the filesystem assertion runs after the provider +returns and before `afterEach` removes the workspace. It may: + +- verify required files and contents; +- execute bounded commands without shell interpolation; +- check exit status, timeout, and selected output; +- inspect Promptfoo provider metadata; and +- inspect OpenTelemetry trace data or use built-in `trajectory:*` assertions. + +If model grading is required, the deterministic phase first serializes all +needed facts to a unique row artifact and a separate evaluation grades that +immutable evidence. It must not read the shared live workspace later. + +`afterEach` may add diagnostic metadata or named scores that Promptfoo permits +hooks to mutate, but it cannot override `success`, `score`, or +`response.output`. It therefore must not contain the authoritative grader. + +V1 stores Promptfoo's native result and OpenTelemetry trace exports. It does not +convert provider events to ATIF. Promptfoo's normalized trajectory view is +sufficient for tool-use, argument, sequence, step-count, and goal assertions. +An ATIF adapter may be added later at an explicit interoperability boundary; it +must not fabricate reasoning, messages, or tool results absent from provider +telemetry. + +Hidden checks are not promised by this design. Keeping a verifier outside +`working_dir` does not prove that a shell-capable agent in the same job cannot +read it. A requirement for secret verifier bytes needs a separate isolation +decision. + +## Scope + +V1 includes: + +- direct Promptfoo execution through its Claude Agent SDK and Codex SDK + providers; +- one fixed extension-managed workspace; +- serial, uncached evaluation rows; +- exact source staging and private per-row copies; +- deterministic filesystem and command assertions; +- Promptfoo-native output, metadata, usage, and OpenTelemetry traces; and +- disposable job-level cleanup. + +V1 excludes: + +- a network execution API or shared remote service; +- a custom Promptfoo provider; +- a durable run database, queue, or artifact service; +- shared mutable workspaces or cross-row provider sessions; +- gateway-owned agent adapters or grading; +- ATIF normalization; +- hostile multi-tenant isolation; +- secret post-run verifier injection; and +- resumable runs, checkpoints, or workspace recovery. + +## Consequences + +Benefits: + +- the implementation is a small Promptfoo configuration, lifecycle extension, + source-staging helper, and deterministic assertion module; +- Claude and Codex provider behavior stays aligned with Promptfoo releases; +- Promptfoo's result, trace, assertion, repetition, and report machinery remains + authoritative; +- source seeds can be reused within a job without sharing mutable workspaces; +- the fixed working directory keeps provider configuration static; and +- deleting the gateway removes a second protocol and telemetry model. + +Costs and limits: + +- V1 is a trusted single-tenant job design, not a hosted execution service; +- rows run serially inside a job; +- live-workspace assertions cannot be mixed with deferred model grading; +- provider tool, transcript, and trace coverage differs; +- `afterEach` failures need a wrapper-visible sentinel because Promptfoo logs + them rather than converting a passing row to an error; +- a process crash can bypass extension cleanup, so disposable job teardown is + required; +- Promptfoo's trace is observability data, not a lossless replayable transcript; + and +- strong credential brokering, hidden verifiers, and tenant isolation remain + unsolved because they are outside the selected scope. + +## Reconsider when + +Introduce a separate execution service only when a concrete requirement needs +one of the boundaries the native design does not provide: mutually untrusted +tenants, remote API callers, centrally enforced network policy, source/model +credential brokering, secret verifier injection, durable cancellation and +recovery, retention beyond the disposable job, or shared scheduling across +machines. + +Need for more providers, matrices, repetitions, deterministic assertions, +workspace copies, or Promptfoo trajectory checks does not by itself justify a +gateway. diff --git a/docs/plans/2026-09-18-0837-feat-promptfoo-coding-agent-evals-plan.md b/docs/plans/2026-09-18-0837-feat-promptfoo-coding-agent-evals-plan.md new file mode 100644 index 00000000..c87e6ff9 --- /dev/null +++ b/docs/plans/2026-09-18-0837-feat-promptfoo-coding-agent-evals-plan.md @@ -0,0 +1,549 @@ +--- +title: "Promptfoo Coding-Agent Evaluations - Implementation Plan" +date: 2026-09-18 +updated: 2026-09-28 +type: feat +artifact_contract: ce-unified-plan/v1 +artifact_readiness: implementation-ready +execution: code +--- + +# Promptfoo Coding-Agent Evaluations - Implementation Plan + +## Goal capsule + +- **Objective:** Evaluate write-capable Claude and Codex agents against fresh, + reproducible workspaces and grade their final filesystem state. +- **Means:** Run Promptfoo's built-in agent providers in one disposable job, + using a lifecycle extension to copy an immutable seed into + `.eval/workspace` before every row and remove it afterward. +- **Evaluation owner:** Promptfoo owns prompts, provider matrices, repetitions, + assertions, scores, pass/fail decisions, OpenTelemetry traces, and reports. +- **Stop conditions:** Do not build an execution gateway, custom Promptfoo + provider, agent adapter, ATIF converter, durable run service, or reusable + session layer. + +The authoritative decision is +[ADR 0002](../decisions/0002-use-promptfoo-native-agent-execution.md). +Primary-source findings are in +[Promptfoo native agent workspaces](../research/promptfoo-native-agent-workspaces.md). +The earlier +[one-shot gateway research](../research/one-shot-coding-agent-gateway-boundary.md) +is retained as analysis of the rejected hosted-service alternative. + +## Product contract + +One evaluation job owns one private `.eval` root. Job bootstrap materializes +exact source inputs once as read-only seeds. Promptfoo then runs every expanded +test/provider/repetition row serially: + +```mermaid +flowchart LR + B[Job bootstrap] --> S[Resolve exact source seeds] + S --> H[Promptfoo beforeEach] + H --> C[Private copy to .eval/workspace] + C --> P[Built-in Claude or Codex provider] + P --> A[Promptfoo assertions inspect final workspace and trace] + A --> T[Promptfoo afterEach captures bounded diagnostics] + T --> D[Delete workspace] + D --> N{More rows?} + N -->|yes| H + N -->|no| J[Export results and destroy job] +``` + +The fixed provider working directory is: + +```text +/.eval/workspace +``` + +The extension is lifecycle glue, not a provider. It never calls a model, +interprets agent output, assigns a reward, or replaces Promptfoo's provider +metadata and tracing. + +## Scope + +### Included + +- Promptfoo configuration for `anthropic:claude-agent-sdk` and + `openai:codex-sdk`. +- One source-staging command that creates immutable job-private seeds from + declared local, Git, or OCI inputs. +- One JavaScript lifecycle extension implementing `beforeAll`, `beforeEach`, + `afterEach`, and `afterAll`. +- One fixed `.eval/workspace` configured on every write-capable provider. +- Serial execution with Promptfoo response caching disabled. +- Deterministic JavaScript assertions over files, commands, provider metadata, + and trace data. +- Promptfoo's built-in OpenTelemetry receiver and trajectory assertions. +- Result, trace, source-provenance, and bounded diagnostic artifacts retained by + the surrounding job. +- Local and CI/VPS execution inside the same disposable container or VM shape. + +### Excluded + +- `allagentsdev/allagents-gateway` or another network execution service. +- A custom Promptfoo provider or local/remote provider mode. +- Queues, databases, idempotency APIs, artifact download APIs, leases, or + cancellation protocols. +- Custom Claude, Codex, or OMP adapters. +- ATIF normalization or a second public trajectory format. +- Cross-row session or thread persistence. +- Concurrent rows sharing one workspace. +- Hidden verifier bytes that must be inaccessible to a shell-capable agent. +- Hostile multi-tenant isolation or caller-specific authorization policy. +- Checkpoints, resumable runs, workspace recovery, or generic patch export. + +## Repository layout + +Add the evaluation implementation to this repository: + +```text +evals/coding-agent/ + promptfooconfig.yaml + sources.yaml + extensions/ + workspace.cjs + assertions/ + workspace.cjs + scripts/ + stage-sources.ts + fixtures/ + ... + .eval/ # ignored, job-private runtime state + seeds/ + workspace/ + artifacts/ +``` + +The configuration, extension, assertions, and source catalog are checked in. +`.eval` is always generated and ignored. + +## Fixed Promptfoo configuration + +Start from this shape: + +```yaml +description: AllAgents coding-agent evaluations + +providers: + - id: anthropic:claude-agent-sdk + config: + working_dir: ./.eval/workspace + append_allowed_tools: [Write, Edit, MultiEdit, Bash] + permission_mode: acceptEdits + persist_session: false + sandbox: + enabled: true + failIfUnavailable: true + + - id: openai:codex-sdk + config: + working_dir: ./.eval/workspace + sandbox_mode: workspace-write + approval_policy: never + enable_streaming: true + persist_threads: false + +extensions: + - file://extensions/workspace.cjs:workspaceLifecycle + +evaluateOptions: + maxConcurrency: 1 + cache: false + +tracing: + enabled: true + otlp: + http: {} + +defaultTest: + assert: + - type: javascript + value: file://assertions/workspace.cjs +``` + +Provider settings remain provider-specific: + +- Claude receives only the explicit tools required by a case. Use + `acceptEdits` for unattended edits, set `persist_session: false`, and require + its sandbox with `failIfUnavailable: true`; do not use permission bypass by + default. +- Codex uses `workspace-write`, `approval_policy: never`, explicit + `persist_threads: false`, and its minimal process environment. + `enable_streaming` supplies provider-level operation and turn spans. +- Do not enable provider thread or session persistence. Every row is a new + attempt. +- Do not enable deep tracing globally. Add it to a focused configuration only + when native SDK spans answer a specific question and their additional data + exposure is acceptable. + +The normal evaluation command uses `--no-cache` as defense in depth even though +the checked-in configuration sets `evaluateOptions.cache: false`. + +## Case contract + +Every test row supplies: + +```yaml +vars: + case_id: safe-stable-id + source_id: staged-source-id + task: rendered agent instruction + check: + command: [bun, test] + timeout_ms: 120000 + expected_exit: 0 +``` + +`case_id` and `source_id` are identifiers, not paths. Restrict them to a short +ASCII identifier grammar such as `^[a-z0-9][a-z0-9._-]*$`. The extension maps +`source_id` beneath its own resolved `.eval/seeds` directory and rejects +unknown identifiers, symlinks escaping the seed, and any resolved path outside +the evaluation root. + +The extension assigns each expanded provider/test/repetition row a monotonic +`workspace_row_id` such as `000001-safe-stable-id` during `beforeEach` and +returns it in the test variables. Authors do not supply this value. It is the +artifact and receipt key, so repeated cases and provider matrices cannot +overwrite one another. + +Checks are closed data consumed by the checked-in assertion module. A command +is an executable plus literal argument vector; it is never a shell string. +Cases cannot supply an environment map, arbitrary assertion module, working +directory, or cleanup command. + +Promptfoo `metadata` is descriptive report data, not the execution control +plane. Repository URLs, refs, workdir paths, Git-cache settings, skill-copy +commands, and verifier commands belong in the checked-in source/workspace +catalog. A suite may select a catalog entry with `defaultTest.vars.source_id`; +the extension consumes that validated identifier. Keep suite metadata for +source links, experiment tags, and other annotations. + +## Source staging + +`scripts/stage-sources.ts` runs before Promptfoo. It reads `sources.yaml` and +materializes only source IDs used by the selected evaluation: + +- local sources are copied from an explicitly allowed repository-relative + path; +- Git sources resolve a requested ref to a full commit and check out that exact + commit; +- OCI sources, when present, resolve and verify an exact manifest digest before + extracting the declared filesystem content. + +Each staged seed contains a provenance record with: + +- source ID and kind; +- requested source and ref, when applicable; +- resolved Git commit or OCI manifest digest; +- materializer version; and +- a deterministic digest of the staged tree or source descriptor. + +Staging uses private temporary directories and publishes a seed only after +materialization and verification succeed. Before Promptfoo starts, bootstrap +exposes published seeds through a read-only mount or transfers them to an +identity the unprivileged Promptfoo/agent user cannot modify. Mode bits owned by +that same user are not an immutability boundary. No credentials, Git +credential-helper responses, registry tokens, or temporary acquisition files +enter the seed. + +Seeds are reusable only inside the current disposable job. V1 does not define a +cross-job cache, eviction protocol, or shared authorization boundary. + +For a composed workspace, the source catalog defines non-overlapping +destinations and staging produces one complete seed tree. Composition happens +once during bootstrap rather than during every Promptfoo row. + +## Workspace lifecycle extension + +Export one extension function: + +```javascript +module.exports.workspaceLifecycle = async function workspaceLifecycle(hookName, context) { + // beforeAll | beforeEach | afterEach | afterAll +}; +``` + +Resolve all paths from CommonJS `__dirname`, not `process.cwd()`. + +### `beforeAll` + +- Refuse to run when `.eval` or the selected seeds resolve outside the + evaluation directory. +- Remove a stale `.eval/workspace` from an interrupted local run. +- Verify that every referenced source ID has a complete staged seed and + provenance record. +- Create a fresh bounded artifacts directory for this evaluation. + +### `beforeEach` + +- Validate `context.test.vars.case_id` and `source_id`. +- Allocate the next unique `workspace_row_id` and add it to the returned test + variables. +- Remove `.eval/workspace` unconditionally. +- Copy the selected read-only seed to a private temporary sibling. +- Never hardlink files or create another writable alias to seed content. +- Use a reflink/copy-on-write copy only when writes cannot reach the seed; use a + full recursive copy otherwise. +- Make the private copy writable, then atomically rename it to + `.eval/workspace`. +- Return the modified context so `workspace_row_id` reaches assertions and + `afterEach`. + +Any setup failure throws and prevents the provider call. + +### `afterEach` + +For deterministic-only rows, Promptfoo's normal path invokes `afterEach` after +provider execution and assertions. The hook: + +- records the row ID, case ID, source provenance digest, and bounded workspace + status in `context.result.metadata`; +- optionally copies explicitly allowlisted diagnostic files to + `.eval/artifacts//`; +- removes `.eval/workspace` in `finally`; +- writes the row's success receipt only after diagnostics and cleanup finish; + and +- writes `.eval/hook-failure.json` before throwing if diagnostic collection or + cleanup fails. + +Promptfoo currently catches and logs `afterEach` exceptions. Throwing is still +useful for logs, but it does not make the CLI exit nonzero. The job wrapper must +reject any run with a hook-failure sentinel, a surviving workspace, duplicate +row IDs, or a missing receipt for any Promptfoo result row. + +The hook must not attempt to rewrite `success`, `score`, or `response.output`; +Promptfoo does not persist such overrides from `afterEach`. + +### `afterAll` + +- Verify that `.eval/workspace` is absent. +- Write a compact source/artifact manifest for the surrounding job. +- Remove remaining temporary directories. +- Leave only explicitly retained seeds or artifacts required by job export. + +The job wrapper performs the authoritative post-Promptfoo sentinel, workspace, +receipt, and manifest checks. Disposable job teardown remains the final cleanup +boundary if Promptfoo or the extension process crashes before hooks complete. + +## Deterministic workspace assertion + +`assertions/workspace.cjs` receives `output` and Promptfoo's assertion context. +It resolves the same fixed workspace path independently of test-controlled +strings. + +For each declared check it: + +1. validates the closed check object; +2. resolves the executable from the disposable job's fixed `PATH`; +3. starts the executable directly with a literal argument vector and + `cwd = .eval/workspace`; +4. enforces the per-check timeout and terminates the spawned process tree; +5. captures bounded stdout and stderr; +6. compares the actual exit status with `expected_exit`; and +7. returns a `GradingResult` with a factual reason and named scores. + +Additional file assertions read only declared workspace-relative paths, reject +absolute paths and `..`, reject symlink escape, and bound bytes read. + +The assertion is authoritative for filesystem behavior because a +deterministic-only row executes it before `afterEach` removes the workspace. The +hook may retain its summary for diagnostics but does not re-grade it. + +Do not combine a live-workspace assertion with a model-graded assertion in the +same evaluation row. Promptfoo can defer the complete assertion set for grouped +model grading, allowing a later `beforeEach` to replace the shared workspace +first. When semantic grading is required, the deterministic evaluation must +serialize all required facts to a unique row artifact; a separate +evaluation grades that artifact without reading `.eval/workspace`. + +Tests for the assertion cover: + +- passing and failing exit statuses; +- timeout and process termination; +- missing executable and launch failure; +- stdout/stderr truncation; +- missing, non-regular, oversized, and symlink-escaping files; and +- an assertion observing an agent mutation before teardown. + +## Traces and transcripts + +Promptfoo OpenTelemetry is the only V1 trace model. + +- Claude's built-in provider emits an `invoke_agent` span, per-turn markers, + and completed tool spans; detailed tool calls also remain in + `response.metadata.toolCalls`. +- Codex uses `enable_streaming: true` to emit provider-level turn, command, + file, search, MCP, response, and reasoning-item spans where the SDK exposes + them. +- Built-in `trajectory:*` assertions check tool use, arguments, sequence, step + count, and goal success. +- JavaScript assertions may inspect `context.trace` when a built-in assertion + is insufficient. + +Retain the Promptfoo result export and trace JSON as job artifacts. These are +observability and evaluation records, not guaranteed lossless transcripts. +Provider coverage differs, subagent text may be summarized or omitted, and +native reasoning may be unavailable. + +Do not add ATIF in V1. Add a converter only when a named downstream consumer +requires ATIF, and preserve explicit missing fields rather than inventing +content absent from Promptfoo/provider telemetry. + +Trace retention and redaction need explicit job settings. Promptfoo's OTLP +receiver redaction does not filter all spans emitted by built-in providers +before local storage. The disposable trace store must contain no test variables +or custom attributes with credentials, and the job must delete local trace +storage after exporting approved artifacts. + +## Isolation and security boundary + +Run write-capable evaluations inside a disposable rootless container or VM with: + +- source bootstrap separated from the unprivileged Promptfoo/agent user; +- `.eval/seeds` mounted read-only or owned by a non-agent identity; +- no host workspace mounted writable; +- no container runtime socket or host device access; +- a private writable evaluation workspace and artifacts directory; +- explicit CPU, memory, process, disk, and wall-time limits; +- network disabled unless the provider call requires an allowlisted endpoint; + and +- only scoped credentials required by source staging and model access. + +Source acquisition should finish before Promptfoo starts so acquisition +credentials can be removed. Model credentials remain provider concerns and must +not be copied into the workspace or test variables. + +This is a trusted single-tenant evaluation design. It does not safely expose a +remote endpoint to arbitrary callers. It also does not guarantee that verifier +files elsewhere in the same job are hidden from an agent with shell access. + +Agent-started background processes may outlive one tool call depending on the +provider runtime. V1 cases must not depend on persistent background services, +and disposable job teardown is the guaranteed process cleanup boundary. A +requirement for process-perfect row isolation changes the design to one +disposable sandbox per row. + +## Failure semantics + +- Source staging failure stops the job before Promptfoo. +- `beforeAll` or `beforeEach` failure prevents affected provider execution. +- Provider launch, timeout, or SDK failure remains a Promptfoo error row. +- A deterministic check returning the wrong exit status is a behavioral + assertion failure, not an infrastructure error. +- Check launch failure, timeout, invalid configuration, or unsafe path is an + assertion error with its exact cause. +- `afterEach` collection or cleanup failure writes a failure sentinel; the job + wrapper fails the run even though Promptfoo itself only logs the hook error. +- A surviving workspace, missing source/artifact manifest, job export failure, + or outer cleanup failure fails the job. + +No automatic agent retry exists. Promptfoo repetitions are intentional new +attempts, each with a new workspace. Provider-internal transport retries remain +provider behavior and must be visible through its output or trace where +supported. + +## Delivery phases and proof + +### Phase 0 — Pin the native execution surface + +Add Promptfoo and the two optional SDK dependencies at tested versions. Add the +evaluation directory, `.eval` ignore rule, configuration, and one read-only +fixture case. + +**Proof:** Promptfoo validates the configuration and both providers resolve the +same absolute `.eval/workspace` from the config directory. + +### Phase 1 — Stage exact reusable seeds + +Implement the source catalog reader and staging command with local and exact +Git sources first. Add OCI materialization only when an initial case requires +it. Record source provenance and reject path escape, mutable published seeds, +credentials in output, and partial staging directories. + +**Proof:** stage the same exact source twice in one job, observe one immutable +seed identity, and verify a failed staging attempt publishes nothing. + +### Phase 2 — Reset one private workspace per row + +Implement the lifecycle extension. Exercise two serial rows against one seed: +the first mutates and adds files; the second must observe only seed content. +Force an `afterEach` failure and verify the next `beforeEach` still deletes the +stale workspace before copying. + +**Proof:** both rows start from the same seed digest, receive different writable +workspace instances, cannot mutate the seed, and leave no workspace after the +suite. + +### Phase 3 — Grade the final filesystem + +Implement closed command and file checks. Keep check programs outside the +workspace and treat them as trusted job code, not secret material. Add focused +tests for the behavioral and safety branches listed above. + +**Proof:** a fixture agent mutation is visible to the assertion, a correct +change passes, an incorrect change fails with exact command/file evidence, and +cleanup runs after either outcome. + +### Phase 4 — Capture native traces and metadata + +Enable Promptfoo tracing, Codex streaming spans, Claude tool metadata, and +trajectory assertions. Configure approved result and trace exports plus local +trace-store deletion. + +**Proof:** one Claude and one Codex run each show final output, usage when +reported, at least one provider/tool span for a tool-using case, deterministic +workspace evidence, and no ATIF artifact. + +### Phase 5 — Run in the disposable job + +Package the exact local and CI/VPS invocation in a rootless container or VM. +Apply resource, filesystem, environment, and network limits. Export results only +after Promptfoo and the extension finish, then destroy the job. + +**Proof:** two consecutive jobs cannot see one another's workspace, seeds, +provider sessions, processes, or local trace database; approved result +artifacts remain available to CI. + +## Focused release E2E + +Run the built evaluation path, not an isolated test helper: + +1. Build the disposable job image. +2. Stage an exact fixture repository revision. +3. Run Promptfoo with `maxConcurrency: 1`, caching disabled, and both built-in + providers. +4. Give each provider a task that must edit a file and run a command. +5. Let the external assertion run the repository's deterministic check. +6. Repeat the case and verify the second row starts pristine. +7. Inspect Promptfoo output, provider metadata, and the trace timeline. +8. Verify the seed is unchanged and `.eval/workspace` is absent. +9. Export approved results/traces and destroy the container. +10. Start another container and verify no mutable state or provider session is + present. + +Record the exact source revision, image digest, Promptfoo/provider versions, +commands, and observed results in the eventual pull request. + +## Completion checklist + +- [ ] Promptfoo invokes Claude and Codex through built-in providers only. +- [ ] Every provider uses `working_dir: ./.eval/workspace`. +- [ ] `maxConcurrency: 1` and response-cache disablement are checked in. +- [ ] Source staging records exact immutable identities and leaks no credential + material. +- [ ] Every row receives a fresh private copy and cannot mutate its seed. +- [ ] Setup failures stop provider execution; cleanup failures create a + wrapper-checked failure sentinel. +- [ ] Live-workspace assertions are deterministic-only and execute before + teardown; model grading, when needed, consumes separately persisted + evidence. +- [ ] Assertions execute literal argument vectors with bounded output and time. +- [ ] Promptfoo results and OpenTelemetry traces are retained as the native + evidence formats. +- [ ] No custom provider, gateway protocol, runner service, or ATIF conversion + remains. +- [ ] Write-capable E2E runs inside a disposable rootless container or VM. +- [ ] The focused release E2E proves row and job isolation, grading, tracing, + export, and cleanup. diff --git a/docs/research/harbor-repository-materialization.md b/docs/research/harbor-repository-materialization.md new file mode 100644 index 00000000..7929c584 --- /dev/null +++ b/docs/research/harbor-repository-materialization.md @@ -0,0 +1,186 @@ +# Harbor repository materialization lessons + +## Status + +This note remains future-adapter research. Harbor, Terminal-Bench, and +SWE-bench are not part of the native Promptfoo V1. + +Harbor is still useful evidence because it packages an instruction, +environment, workdir, verifier, and reward artifact around one disposable task. +The selected design does not put Harbor between Promptfoo and its built-in +Claude or Codex providers, and does not treat Harbor as a source kind. + +Any future Harbor integration needs its own decision covering task provenance, +workspace ownership, verifier visibility, result mapping, and which system owns +the sandbox. It must not silently reintroduce the rejected gateway contracts or +move behavioral judgment out of the selected evaluation owner. + +The current boundary is defined by +[ADR 0002](../decisions/0002-use-promptfoo-native-agent-execution.md) and +[Promptfoo native agent workspaces](./promptfoo-native-agent-workspaces.md). + +## What Harbor fetches + +### Task packages from Git + +Harbor's `GitRepoRegistryClient` resolves a dataset registry ref, inspects the selected +commit's tree without checking out blobs, and identifies task directories containing +`task.toml`. When task content is requested, `TaskClient` groups requested task paths by +Git URL and performs one shallow, no-checkout clone per URL. It uses a blobless partial +clone where supported, configures sparse checkout for only the selected task paths, +fetches each requested commit at depth one, checks it out, and records the resolved +commit. + +This is efficient for a large repository containing many independent Harbor tasks. It +is not a mechanism for assembling several application repositories into one agent +workspace. + +Harbor also accepts an omitted commit or a mutable ref and resolves it to a +commit. The native source catalog may declare a branch, tag, or commit for +developer convenience, but job bootstrap resolves and records the full commit +before publishing the seed. Reproducibility-sensitive catalog entries use a full +commit; OCI entries resolve and record the verified manifest digest before seed +publication. + +### Task packages from the package registry + +Package-registry tasks are downloaded as tar archives into a cache keyed by the task's +content hash. A direct `sha256:` reference can hit that cache without registry +resolution. Dataset manifests likewise refer to task packages by SHA-256 digest. This +is the closest Harbor analogue to an OCI workspace snapshot: a content-addressed, +reusable input bundle. + +Before publishing a Git task directory, Harbor stages it in a temporary directory, +rejects source paths containing symlinks, materializes only relative symlinks that stay +inside the task root, rejects cycles and special entries, then replaces the target. +Those containment and publish-after-validation properties are useful for any cached +workspace artifact. + +### The repository edited by the agent + +Once the task package is present, Harbor asks the selected environment provider to +start the task's `environment/` definition. For Docker this can be: + +- `[environment].docker_image`; +- `environment/Dockerfile`; or +- `environment/docker-compose.yaml`. + +The task format deliberately leaves the environment flexible. Harbor builds the +Dockerfile/Compose definition or uses the prebuilt image, then runs the agent in that +environment. There is no core repository-source schema carrying URL, exact commit, +destination, and per-repository provenance. + +The Multi-SWE-bench adapter makes the distinction concrete. Each generated task uses +an upstream `mswebench/...:pr-...` base image that already contains the repository at +`/home/{repo_name}`. Its Dockerfile creates `/workspace/{repo_name}` as a symlink and +sets that as `WORKDIR`; Harbor itself never clones that application repository. + +## Superseded materializer lessons (historical) + +The sections below preserve conclusions from the rejected builder/snapshot +architecture. They are evidence about Harbor's implementation, not current +recommendations for the evaluation-only runner. + +### Adopt + +1. **Separate descriptor acquisition from execution.** Resolve and validate immutable + inputs before starting the coding-agent runtime. +2. **Use content-addressed snapshot caches.** Key reusable OCI workspace + snapshots by their immutable OCI and workspace-manifest digests. Direct Git + mode resolves revisions independently and records the resulting commits. + Reauthorize every remote acquisition. +3. **Avoid downloading irrelevant content.** For Git-backed descriptor catalogs, + Harbor's tree-only discovery and sparse checkout are sound optimizations. For an + application repository, use partial/shallow acquisition only when it preserves the + required commit and evidence semantics. +4. **Stage, validate, then publish.** Materialize into a temporary location, enforce + path/link/type/size limits, verify every requested identity, and atomically expose + the completed workspace to the worker. +5. **Support prebuilt immutable artifacts.** A digest-pinned OCI workspace snapshot is + the scalable path for very large repositories and expensive setup. + +### Adapt + +Keep a first-class workspace manifest instead of hiding source inside an +environment image. Each materialized repository or snapshot should retain at +least: + +- canonical source URL or configured snapshot identity; +- requested revision and resolved commit, or OCI manifest digest; +- destination path and optional source subdirectory; +- acquisition implementation identity; +- resulting tree/content identity; and +- completeness and verification-versus-attestation facts. + +Use exactly two initial source modes: + +1. **Direct declared Git repositories** for the normal case. A request selects + configured repository names and may override only their revisions. The + materializer resolves and records full commits and enforces collision-free + destinations. +2. **Named OCI workspace snapshots** for large, preassembled workspaces. The + project workspace declares the repository; the request supplies immutable + OCI and workspace-manifest digests. + +Both modes produce the same standard workspace manifest. Neither mode falls +through to the other after admission. + +### Do not copy + +- Unresolved mutable Git refs as terminal execution identities. Branch and tag + overrides are valid only when the materializer resolves and records a full + commit before agent execution. +- Mutable OCI tags or package `latest` as accepted snapshot identities. +- Harbor's broad Git transport set (`http`, `ssh`, and `git` as well as HTTPS) at + the materializer boundary. AllAgents permits only canonical credential-free + HTTPS with configured hosts, disabled redirects/helpers/filters/hooks/ + submodules, and full-commit verification. +- A non-fatal Git LFS miss. If declared workspace content cannot be materialized, + preparation fails before agent execution. +- Hashing a prebuilt image reference string as environment identity. Resolve and + pin the OCI manifest digest. +- Arbitrary task-authored Dockerfiles, Compose files, or public-network setup as + caller input. Harbor runs benchmark definitions trusted by the evaluator; + AllAgents accepts authenticated service requests with a different trust + boundary. +- Treating a container image alone as sufficient provenance. An image can carry the + correct files while obscuring which repositories, commits, generator, and setup + produced them. + +## Future-adapter questions + +Before implementing any Harbor, Terminal-Bench, or SWE-bench adapter: + +1. Write a dedicated ADR and adapter mapping. Do not add a Harbor source kind to + the native workspace catalog implicitly. +2. Decide and specify who owns the sandbox, agent invocation, checks, + cancellation, cleanup, and result publication. +3. Define authoritative registry, task, version, environment, and agent + provenance. +4. Preserve bounded raw agent and check evidence without translating Harbor + scores into an unrelated pass/fail shape. +5. Define artifact ownership, retention, integrity, and consumer access when + evidence must outlive the disposable job. +6. Prove cleanup and define behavior for Harbor outages and cleanup failure. +7. Keep direct Promptfoo runs independent of Harbor and let Promptfoo assertions + or graders make every behavioral judgment. + +The practical conclusion is narrow: Harbor is useful primary-source evidence +for content-addressed task packages, isolated tasks, and colocated checks. None +of Harbor, Terminal-Bench, or SWE-bench is an approved current adapter, control +plane, provenance variant, or result contract. They do not justify reusable +sessions. + +## Primary sources + +Inspected Harbor commit +[`b83e7686999a18ba90a8603794d7d18d42cab010`](https://github.com/harbor-framework/harbor/tree/b83e7686999a18ba90a8603794d7d18d42cab010): + +- [`src/harbor/registry/client/git_repo.py`](https://github.com/harbor-framework/harbor/blob/b83e7686999a18ba90a8603794d7d18d42cab010/src/harbor/registry/client/git_repo.py) — ref resolution, tree-only discovery, and sparse registry checkout. +- [`src/harbor/tasks/client.py`](https://github.com/harbor-framework/harbor/blob/b83e7686999a18ba90a8603794d7d18d42cab010/src/harbor/tasks/client.py) — Git/local/package task acquisition, content-hash cache, LFS behavior, safe staging, and resolved commits. +- [`src/harbor/models/task/id.py`](https://github.com/harbor-framework/harbor/blob/b83e7686999a18ba90a8603794d7d18d42cab010/src/harbor/models/task/id.py) — Git, local, and package task identities. +- [`src/harbor/models/dataset/manifest.py`](https://github.com/harbor-framework/harbor/blob/b83e7686999a18ba90a8603794d7d18d42cab010/src/harbor/models/dataset/manifest.py) — digest-addressed dataset task references. +- [`docs/content/docs/tasks/index.mdx`](https://github.com/harbor-framework/harbor/blob/b83e7686999a18ba90a8603794d7d18d42cab010/docs/content/docs/tasks/index.mdx) — task structure and Docker image/Dockerfile/Compose environment contract. +- [`src/harbor/environments/definition.py`](https://github.com/harbor-framework/harbor/blob/b83e7686999a18ba90a8603794d7d18d42cab010/src/harbor/environments/definition.py) — environment selection and content identity. +- [`src/harbor/environments/docker/docker.py`](https://github.com/harbor-framework/harbor/blob/b83e7686999a18ba90a8603794d7d18d42cab010/src/harbor/environments/docker/docker.py) — prebuilt-image versus build behavior and container startup. +- [`adapters/multi-swe-bench/README.md`](https://github.com/harbor-framework/harbor/blob/b83e7686999a18ba90a8603794d7d18d42cab010/adapters/multi-swe-bench/README.md) and its [`environment/Dockerfile`](https://github.com/harbor-framework/harbor/blob/b83e7686999a18ba90a8603794d7d18d42cab010/adapters/multi-swe-bench/src/multi_swe_bench_adapter/task-template/environment/Dockerfile) — application repository supplied by an upstream prebuilt image rather than cloned by Harbor core. diff --git a/docs/research/one-shot-coding-agent-gateway-boundary.md b/docs/research/one-shot-coding-agent-gateway-boundary.md new file mode 100644 index 00000000..aa60d1f0 --- /dev/null +++ b/docs/research/one-shot-coding-agent-gateway-boundary.md @@ -0,0 +1,302 @@ +# One-shot coding-agent gateway boundary + +## Status + +This note records the hosted one-shot gateway alternative that was considered +and rejected for the initial trusted local/CI evaluation scope. Its source, +sandbox, credential, and artifact analysis remains useful if AllAgents later +needs a remote multi-tenant execution service. + +The current decision uses Promptfoo's built-in Claude and Codex providers with +an extension-managed disposable workspace. See +[ADR 0002](../decisions/0002-use-promptfoo-native-agent-execution.md) and +[Promptfoo native agent workspaces](./promptfoo-native-agent-workspaces.md). + +## Historical decision + +The prior proposal placed a general one-shot `AgentRun` API between Promptfoo +and coding agents. It assigned source materialization, immutable caching, +runtime profiles, agent adapters, post-run evidence, artifacts, and cleanup to +`allagentsdev/allagents-gateway`. + +That boundary is not part of V1. Promptfoo now invokes its built-in agent +providers directly inside one disposable job. A lifecycle extension resets +`.eval/workspace` for every serial row, deterministic assertions inspect the +final filesystem, and Promptfoo retains its native result and OpenTelemetry +trace formats. + +The gateway alternative remains relevant only if later requirements demand +remote callers, hostile tenant isolation, credential brokering, hidden +verifiers, durable cancellation/recovery, or retention independent of a +disposable job. + +## Why Promptfoo is the control plane + +Promptfoo already owns the evaluation-shaped abstractions: + +- its configuration expands prompts, providers, and test variables into a matrix and applies + per-test assertions ([configuration guide](https://www.promptfoo.dev/docs/configuration/guide/)); +- its assertion layer supports deterministic JavaScript/Python checks, structured-output checks, + weighted scores, cost/latency limits, and model-graded rubrics + ([assertions and metrics](https://www.promptfoo.dev/docs/configuration/expected-outputs/)); +- its coding-agent guidance recommends repeated runs for stochastic agents, disposable + workspaces for write-capable tests, and checking files after the run when final text is not + sufficient evidence + ([Evaluate Coding Agents](https://www.promptfoo.dev/docs/guides/evaluate-coding-agents/)); and +- a custom JavaScript or TypeScript provider needs only `id()` and `callApi()`, and may return + structured `output`, `error`, usage, cost, and arbitrary metadata + ([custom JavaScript provider](https://www.promptfoo.dev/docs/providers/custom-api/)). + +Promptfoo's stock Codex and Claude Agent SDK providers accept an explicit +working directory and expose provider-native output, usage, metadata, and +tracing. They do not materialize pristine source trees themselves. The selected +design supplies that missing lifecycle with job bootstrap plus Promptfoo +`beforeEach` and `afterEach` hooks; it does not require a custom provider or +gateway. + +## Multi-turn and sandboxed-code boundaries + +Promptfoo's simulated-user provider has two different transport modes. Its default +resends the complete transcript on each turn. With `stateful: true`, Promptfoo +sends only the newest user message after the target returns a session ID and +expects that target to retain its own history +([simulated-user provider](https://www.promptfoo.dev/docs/providers/simulated-user/)). +The native V1 deliberately disables cross-row session persistence. Provider +turns that occur inside one Promptfoo row share that row's private workspace, +but the next provider/test/repetition row starts from a new seed copy. A later +multi-turn evaluation that intentionally preserves filesystem state needs an +explicit case-level lifecycle rather than accidental thread pooling. + +Promptfoo's +[sandboxed-code guide](https://www.promptfoo.dev/docs/guides/sandboxed-code-evals/) +does not put Promptfoo or its provider inside a sandbox. Its `type: python` +assertion runs trusted user code, and that assertion explicitly calls Epicbox +to launch generated code in a one-time Docker container. The current design +therefore uses a disposable outer container or VM for write-capable agent +execution and uses Promptfoo JavaScript assertions only for trusted +deterministic checks against the retained final workspace. + +## Historical `AgentRun` contract + +The rejected gateway proposal defined closed `AgentRunRequest`, +`PostRunSpec`, and `AgentRunResult` wire contracts. They are not current +implementation contracts and the superseded implementation plan has been +replaced by the +[Promptfoo coding-agent evaluation plan](../plans/2026-09-18-0837-feat-promptfoo-coding-agent-evals-plan.md). +The details below are retained only to document what a future hosted execution +service would need to decide. + +The request contract is closed. At a logical level it carries request identity, +the instruction, ordered workspace sources and working directory, an agent +selection, a required `runtime_profile_id`, an agent network-policy name, +`post_run` as either `null` or the closed post-run specification, and all +required limits including `max_agent_output_bytes`. Admission resolves and +persists one authorized immutable runtime-profile revision and its canonical +`profile_digest`; dispatch uses only that revision and re-verifies the profile, +runtime image, tool/service implementation, and containerized-service image +digests. Callers do not submit an image, environment map, tool path, service +command, implementation digest, or mutable runtime configuration. + +Git sources carry a canonical HTTPS URL, ref, history policy, destination, and +access mode. OCI sources carry a canonical repository fetch location, direct +descriptor, destination, and access mode; a digest alone has no fetch location. +Uploaded bundles carry an authorized `BundleReference`. For Git and OCI, the +authenticated caller plus canonical source must match exactly one +operator-configured policy/credential route; zero or ambiguous matches reject +before network access. Caller JSON carries no credential or route selector. +Source acquisition denies redirects. Any OCI bearer-token realm is pinned by +the matched operator route rather than trusted from an arbitrary challenge. + +`PostRunSpec v1` uses an optional authorized bundle reference, one named +post-run network policy, ordered structured executable variants with literal +arguments, and ordered output-file requests. It permits no shell parsing, +caller environment, command-specific working directory, or command-specific +network policy. Runtime-only commands and output-only collection need no bundle. + +`AgentRunResult v1` has status `completed`, `cancelled`, or +`infrastructure_error`; direct workspace provenance when available; immutable +runtime provenance; bounded agent output, usage, timing, trajectory, raw +post-run evidence, and typed error data. `RuntimeProvenance` records +`runtime_profile_id`, `profile_digest`, `image_digest`, sandbox-policy version, +each tool's name/version/`implementation_digest`, and each service's +name/version/`implementation_digest` plus container image digest when applicable. + +`AgentResult.final_output` is a bounded `CapturedText` union. Its inline form +records UTF-8 text, digest, byte size, and truncation; its artifact form records +an `ArtifactReference` and truncation. `PostRunEvidence` keeps complete +request-order command and output-file observation vectors. Command observations +distinguish `completed`, `timed_out`, `not_run`, and `unavailable`; file +observations distinguish collected, missing, limit-exceeded, non-regular, +not-run, and unavailable states. Each completed observation is sealed durably, +so cancellation and infrastructure errors can return raw partial evidence +instead of discarding it. Promptfoo alone interprets that evidence as behavioral +pass/fail or reward. + +An `ArtifactReference` includes `artifact_id`, media type, digest, byte size, and +expiry. +Status, cancellation, terminal-result, and artifact-dereference operations +authenticate the caller and enforce tenant/run ownership. The provider obtains +referenced bytes through the authenticated artifact endpoint before expiry and +verifies streamed size and digest; an opaque artifact ID is never treated as +self-authenticating evidence. + +Post-run commands execute only after the agent reaches a terminal state and its +process/cgroup and network namespace are torn down, descendant absence is +verified, and source/model credentials are removed. Hidden-bundle bytes do not +exist in the run filesystem before that boundary. The worker then materializes +any authorized bundle. Commands run in the retained runtime and final workspace +with the exact dependencies, services, filesystem state, and declared source +modes produced by setup and the agent. + +Network access is phase-separated and default-drop: private acquisition, +task-service, agent, and post-run namespaces expose only the destinations +authorized for that phase. Post-run external egress is denied by default; its +named policy may expose only declared localhost or sidecars. A nonzero or timed +out command remains a raw observation, while failure to create/materialize the +workspace, enforce isolation, or launch a command is infrastructure failure. + +Terminal visibility is cleanup-gated. The gateway first assembles and seals all +available evidence, then cleans the workspace/runtime and releases cache and +artifact leases, and only afterward publishes `AgentRunResult v1`. Cleanup or +lease-release failure returns `infrastructure_error` with the partial evidence +collected so far; it never publishes a misleading completed result. + +Keep the API, disposable trial worker, agent adapters, source materializers, +post-run command executor, artifact service, and first Promptfoo provider +together in `allagentsdev/allagents-gateway`. This is a general one-shot +coding-agent boundary, not an evaluation-specific service or session platform. + +## Policy-bound immutable source cache + +The gateway owns a shared cache of verified, immutable source generations so +concurrent trials do not refetch or recopy large inputs. A cache lookup never +bypasses policy: canonicalize the source and repeat the authenticated-caller +route match on every request, including hits. Credentials are used only to +populate a missing generation and are not stored in cached content. + +Cache identity is an exact Git commit/object generation or an OCI repository +plus direct manifest descriptor and the materializer/cache format version. The +descriptor supplies content identity; the repository supplies fetch and policy +context. Populate misses in private staging, verify identity and limits, remove +acquisition-only state, then publish the generation atomically and read-only. A +mutable Git ref or OCI tag may be an input to resolution but never a cache +identity. + +Materialize each declared source according to access: + +- **read-only:** mount the policy-admitted cached generation directly into every + concurrent trial; +- **writable:** create a private reflink/copy-on-write clone; if the filesystem + cannot clone, make a full private copy. + +No trial may write the cache or another trial's view. Agent files, hidden check +bundles, post-run mutations, credentials, processes, and service state remain +trial-private. Cleanup removes those trial views but not the immutable cache +generation. + +Linux reflinks provide the local copy-on-write primitive, with same-filesystem +constraints ([`FICLONE`](https://man7.org/linux/man-pages/man2/ioctl_ficlonerange.2.html)). + +## Authenticated private precedent + +The authenticated WiseTechGlobal example +[`exercises/coding-agent-harness`](https://github.com/WiseTechGlobal/ai-evals-examples/tree/main/exercises/coding-agent-harness) +is direct evidence for this boundary (the repository is private and the links require access): + +- [`lib/coding-agent.mjs`](https://github.com/WiseTechGlobal/ai-evals-examples/blob/main/exercises/coding-agent-harness/lib/coding-agent.mjs) + creates a unique temporary root with `mkdtempSync`, recursively copies the fixture into a fresh + workspace, launches the SDK agent in a constrained Docker container, and removes the complete + temporary root in `finally`; +- after the agent exits, checks run with the final workspace mounted into the + check environment and networking disabled; +- [`lib/coding-agent.mjs`](https://github.com/WiseTechGlobal/ai-evals-examples/blob/main/exercises/coding-agent-harness/lib/coding-agent.mjs) + collects fixture-test exits and hidden-check JSON from the workspace, proving + that filesystem checks must execute where the final files and task runtime are + available; +- the example uses separate check containers. The gateway should instead keep + checks in the same trial runtime and final workspace so installed dependencies + and declared services remain available without changing declared source modes; + hidden checks are injected only after the agent and its descendants stop and + credentials are removed; +- the example currently converts observations into pass/fail/reward inside its + provider. The gateway boundary should stop one step earlier: assemble raw + exits, output, and requested artifacts, complete cleanup, then return the + terminal result for Promptfoo code or LLM graders to judge; +- the trial path contains no Git initialization, diff, commit, or produced-file + cursor—the authoritative object is the final filesystem state; and +- [`promptfooconfig.yaml`](https://github.com/WiseTechGlobal/ai-evals-examples/blob/main/exercises/coding-agent-harness/promptfooconfig.yaml) + leaves repetition, prompt variants, task rows, assertions, tracing, timeout, + and concurrency to Promptfoo. + +The example currently runs a Copilot SDK agent in-process with Promptfoo rather +than calling a general gateway. Its reusable evidence is the disposable trial +lifecycle and the requirement that checks execute beside the final workspace. +Keep all judgment in Promptfoo. + +## Harbor and Terminal-Bench are future adapter research + +Harbor packages an instruction, environment, and test script as a self-contained +task. A trial starts the environment, runs the agent, then runs the test script +in that environment; the script writes a numeric or structured reward under +`/logs/verifier/` +([task overview](https://docs.harborframework.com/core-concepts/tasks/overview), +[task tutorial](https://docs.harborframework.com/tutorials/create-a-task)). + +Harbor and Terminal-Bench are outside the native Promptfoo V1. A future +integration should define its own adapter boundary and provenance rather than +revive the rejected gateway implicitly. Promptfoo remains the owner of +behavioral pass/fail and reward. + +## SWE-bench is future adapter research + +SWE-bench's official evaluator consumes a prediction record containing +`instance_id`, model identity, and `model_patch`; it creates a Docker +environment, applies that patch, runs tests, and writes per-instance reports and +logs +([evaluation guide](https://www.swebench.com/SWE-bench/guides/evaluation/), +[harness reference](https://www.swebench.com/SWE-bench/reference/harness/)). +That is useful evidence that any future SWE-bench adapter owns its exact patch +transport. + +SWE-bench is outside the native Promptfoo V1. A future adapter must define its +base-commit, filename, diff-format, size, and evaluator compatibility rules. +The current workspace extension does not expose a generic diff, patch, +modified-workspace, or change-artifact API. + + +## Historical recommendation + +Implement the smallest complete loop: + +```mermaid +flowchart LR + P[Promptfoo JSON matrix] --> Q[Promptfoo provider] + Q --> G[AgentRun API] + G --> K[Map caller + canonical source to one policy route] + K --> M[RO mount or private CoW / full copy] + M --> W[Disposable trial + authorized runtime profile] + W --> A[Codex or OMP] + A --> T[Teardown agent process/netns + remove credentials] + T --> C{Post-run commands?} + C -->|yes| H[Only now materialize authorized bundle] + H --> E[Same runtime + declared source modes] + E --> R[Raw exits + output + artifacts] + C -->|no| O[Agent output + evidence] + R --> Z[Assemble terminal evidence] + O --> Z + Z --> D[Cleanup + release leases] + D --> V[Return terminal AgentRunResult to provider] + V --> X[Dereference authorized artifacts + verify digest] + X --> J[Promptfoo code / LLM graders] +``` + +The diagram above summarizes the rejected hosted-service recommendation. It +would be appropriate only if AllAgents needed a remote execution product with +tenant authorization, strong sandboxing, policy-bound source and model +credentials, cleanup-gated results, and durable artifacts. + +For the selected trusted local/CI scope, those controls would duplicate the +disposable job and Promptfoo's built-in providers while adding a second API, +provider, adapter, trace, and persistence stack. The current design therefore +keeps the reusable insight—fresh private workspaces and post-agent filesystem +checks—but implements it with Promptfoo lifecycle hooks and assertions. diff --git a/docs/research/promptfoo-native-agent-workspaces.md b/docs/research/promptfoo-native-agent-workspaces.md new file mode 100644 index 00000000..80e09392 --- /dev/null +++ b/docs/research/promptfoo-native-agent-workspaces.md @@ -0,0 +1,496 @@ +# Promptfoo-native coding-agent workspaces + +## Conclusion + +For a **trusted, single-host local or CI evaluation**, Promptfoo's built-in Claude +Agent SDK or OpenAI Codex SDK provider can replace the proposed execution gateway's +basic evaluation loop: + +1. a Promptfoo `beforeEach` extension deletes and re-copies a fixture into the fixed + `./.eval/workspace` path; +2. `evaluateOptions.maxConcurrency: 1` prevents two cases from sharing that path; +3. the provider receives `config.working_dir: ./.eval/workspace` and runs one agent; +4. a deterministic `javascript` assertion reads the final filesystem; +5. `afterEach` records bounded observations if needed and removes the workspace; and +6. `afterAll` retries cleanup of remaining job-private state. + +That is enough when the source fixture is already local, the CI runner is the +security boundary, the checks are trusted, and raw Promptfoo results plus +OpenTelemetry traces are sufficient. It is composition, not a built-in disposable +workspace feature: the reset/materialization and filesystem verifier are glue code. +Promptfoo explicitly recommends serial execution, extension hooks, wrapper scripts, +Git, or containers for side-effecting Claude runs, and its official advanced example +uses `maxConcurrency: 1` plus an extension reset +([side-effect guidance](https://www.promptfoo.dev/docs/providers/claude-agent-sdk/#managing-side-effects), +[pinned example config](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/examples/claude-agent-sdk/advanced/promptfooconfig.yaml#L8-L29), +[pinned reset hook](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/examples/claude-agent-sdk/advanced/hooks.js#L58-L101)). + +It does **not** replace the gateway's stronger contracts: authenticated and +provenance-preserving Git/OCI/bundle composition, isolation from an untrusted agent, +phase-separated credentials and networking, resource/process-tree enforcement, +hidden post-run checks, bounded and durable evidence, normalized ATIF trajectories, +idempotency/cancellation, or an OMP adapter. Those requirements can justify an +external runner even when there is no remote multi-tenant service. + +## Evidence labels and source snapshot + +- **Documented** means Promptfoo's public documentation promises the behavior. +- **Source-observed** means the current implementation does it, but the public docs do + not define it as a stable contract. +- **Proposed glue** means code this evaluation repository must own; it is not supplied + by Promptfoo or either agent SDK provider. + +Source was inspected at Promptfoo commit +[`712a506de6412ca6879fe8dba1319569ea820cbf`](https://github.com/promptfoo/promptfoo/tree/712a506de6412ca6879fe8dba1319569ea820cbf). +All source links below are pinned to that revision. Public documentation links are +unversioned and describe the current docs as inspected on 2026-09-28. + +## Direct answer: can the fixed-workspace loop work? + +Yes, with the following boundaries. + +| Step | Status | Exact behavior and boundary | +|---|---|---| +| Serialize cases | **Documented** | `evaluateOptions.maxConcurrency: 1` sets the maximum concurrent requests to one; the CLI equivalent is `--max-concurrency 1`. `tests[].options.runSerially: true` is also supported, but global concurrency one is simpler when every case shares one path ([configuration reference](https://www.promptfoo.dev/docs/configuration/reference/#config), [test-case reference](https://www.promptfoo.dev/docs/configuration/reference/#test-case)). | +| Reset before a case | **Documented hook, proposed reset** | Root `extensions` supports `beforeEach`, whose context is `{ test }`, before each individual evaluation. Promptfoo provides the callback point, not copy/clone/reset logic ([extension hooks](https://www.promptfoo.dev/docs/configuration/reference/#extension-hooks)). | +| Point the agent at the path | **Documented** | Both providers accept `config.working_dir`; relative values resolve from the config file's directory ([Claude working directory](https://www.promptfoo.dev/docs/providers/claude-agent-sdk/#with-working-directory), [Codex working directory](https://www.promptfoo.dev/docs/providers/openai-codex-sdk/#with-working-directory)). The shared resolver in source confirms config-relative resolution ([`resolveAgenticWorkingDir`](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/providers/agentic-utils.ts#L74-L90)). | +| Run a write-capable agent | **Documented** | Codex uses `sandbox_mode: workspace-write` by default. Claude requires explicit write/edit tools and a permission mode such as `acceptEdits`; a configured directory is read-only by default ([Codex sandbox modes](https://www.promptfoo.dev/docs/providers/openai-codex-sdk/#sandbox-modes), [Claude tools and permissions](https://www.promptfoo.dev/docs/providers/claude-agent-sdk/#tools-and-permissions)). | +| Inspect the final filesystem | **Documented assertion API; source-observed ordering** | An external `javascript` assertion is trusted Node code and receives `output` plus `context`, so it can read files or invoke a fixed verifier ([JavaScript assertions](https://www.promptfoo.dev/docs/configuration/expected-outputs/javascript/#external-script)). Current source invokes `beforeEach`, then `runEvalInternal`, and only afterward invokes `afterEach` and persists the row ([provider/evaluation order](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluator.ts#L3567-L3593), [post-row hook order](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluator.ts#L3595-L3653)). Thus a normal deterministic assertion sees the provider's final filesystem. | +| Reset for the next case | **Proposed glue** | `afterEach` may archive bounded evidence and then remove the workspace; the next `beforeEach` removes it again before copying the seed. Cleanup failures are only logged by the current evaluator and do not automatically turn a passing row into an error ([caught `afterEach` failure](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluator.ts#L3614-L3650)). Therefore `beforeEach` must throw on a failed reset, `afterAll` must retry cleanup, and the outer job should verify teardown when cleanup is a correctness requirement. | +| Prevent a cache hit from skipping the run | **Documented and source-observed** | Set `evaluateOptions.cache: false` or use `--no-cache`. This disables the scoped cache used by agentic providers; the Claude provider otherwise fingerprints the working directory and can return a prior response ([caching configuration](https://www.promptfoo.dev/docs/configuration/caching/), [evaluation cache scope](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluate.ts#L361-L364), [agentic cache initialization](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/providers/agentic-utils.ts#L203-L273)). A cached response cannot recreate cached filesystem mutations. | + +### Important grading-order caveat + +The fixed shared path is safe for the deterministic filesystem assertion shown below. +Do not assume it remains safe if the same test also contains a model-graded assertion. +At concurrency one, the current evaluator can group model-graded assertions by provider: +it performs several target calls and defers each row's complete assertion set before +flushing grading ([grouping predicate](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluator.ts#L519-L537), +[grouped execution](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluator.ts#L3912-L4007), +[deferred assertion call](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluator.ts#L1449-L1469)). +A later `beforeEach` can therefore replace the workspace before the earlier row's +filesystem assertion executes. + +Use one of these clean boundaries: + +- keep filesystem-observing rows deterministic-only; +- have the deterministic assertion capture all needed facts and grade later from those + persisted facts in a separate evaluation; or +- use a wrapper/custom provider that returns a per-run evidence snapshot with the agent + response. + +A nonzero `evaluateOptions.timeoutMs` currently disables grouped grading, but that is +an implementation detail rather than a documented workspace-lifecycle guarantee; it +should not be the foundation of evidence correctness +([grouping condition](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluator.ts#L4939-L4960)). + +## Minimal verified composition + +The following shape is verified against the configuration reference, both provider +implementations, and Promptfoo's official Claude reset example at the pinned revision. +It intentionally uses one prompt, one provider, deterministic-only assertions, global +serialization, and disabled response caching. + +Expected repository layout: + +```text +promptfooconfig.yaml +fixtures/base/ # immutable or otherwise protected seed +.eval/workspace/ # disposable; ignored by version control +promptfoo/workspace-hooks.cjs +promptfoo/assert-workspace.cjs +``` + +### `promptfooconfig.yaml` + +```yaml +# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json + +description: Native fixed-workspace coding-agent evaluation + +prompts: + - | + Implement the requested change. Write the exact value "{{ expected }}" to result.txt. + +providers: + - id: openai:codex-sdk + config: + working_dir: ./.eval/workspace + skip_git_repo_check: true # remove when fixtures contain a Git repository + sandbox_mode: workspace-write + approval_policy: never + network_access_enabled: false + web_search_mode: disabled + enable_streaming: true + persist_threads: false + +extensions: + - file://./promptfoo/workspace-hooks.cjs:extensionHook + +evaluateOptions: + maxConcurrency: 1 + cache: false + +tracing: + enabled: true + otlp: + http: {} + +outputPath: ./.eval/results.json + +tests: + - description: writes the requested result + vars: + expected: READY + assert: + - type: javascript + value: file://./promptfoo/assert-workspace.cjs +``` + +`skip_git_repo_check: true` is necessary only for a non-Git fixture; Codex otherwise +requires the working directory or a parent to be a Git repository. `approval_policy: +never` is the documented unattended-CI recommendation, and network, search, and full +process-environment inheritance are separate settings from the filesystem sandbox +([Codex parameters and caveats](https://www.promptfoo.dev/docs/providers/openai-codex-sdk/#supported-parameters), +[Codex sandbox modes](https://www.promptfoo.dev/docs/providers/openai-codex-sdk/#sandbox-modes)). + +For Claude, replace only the provider block: + +```yaml +providers: + - id: anthropic:claude-agent-sdk + config: + working_dir: ./.eval/workspace + append_allowed_tools: [Write, Edit, MultiEdit, Bash] + permission_mode: acceptEdits + persist_session: false + sandbox: + enabled: true + failIfUnavailable: true +``` + +Claude's `sandbox.enabled` is not implied by `working_dir` or `permission_mode`. +`failIfUnavailable` defaults to true when the sandbox is enabled, but spelling it out +makes the CI requirement visible. Network domains, local binding, Unix sockets, and +credential masking have their own `sandbox.network.*` and +`sandbox.credentials.envVars` keys +([Claude sandbox configuration](https://www.promptfoo.dev/docs/providers/claude-agent-sdk/#sandbox-configuration)). + +`deep_tracing` is intentionally omitted from the baseline. Root tracing plus Claude's +provider spans, or Codex `enable_streaming`, supplies the normal evaluation trajectory +with less payload exposure. Opt into deep tracing only when SDK/CLI-internal spans are +required: Codex deep tracing disables `persist_threads`, `thread_id`, and +`thread_pool_size`, and native CLI spans can contain payloads outside Promptfoo's +stream-event sanitizer +([Codex deep tracing](https://www.promptfoo.dev/docs/providers/openai-codex-sdk/#deep-tracing)). + +### `promptfoo/workspace-hooks.cjs` — proposed glue + +```js +const fs = require('node:fs/promises'); +const path = require('node:path'); + +const root = path.resolve(__dirname, '..'); +const seed = path.join(root, 'fixtures', 'base'); +const workspace = path.join(root, '.eval', 'workspace'); + +async function materializeFreshWorkspace() { + await fs.rm(workspace, { recursive: true, force: true }); + await fs.mkdir(path.dirname(workspace), { recursive: true }); + await fs.cp(seed, workspace, { recursive: true, errorOnExist: true }); +} + +module.exports = async function extensionHook(hookName, context) { + if (hookName === 'beforeEach') { + // Do not catch this error: a failed reset must prevent the agent call. + await materializeFreshWorkspace(); + } else if (hookName === 'afterEach' || hookName === 'afterAll') { + // Assertions have completed before afterEach on the normal deterministic path. + await fs.rm(workspace, { recursive: true, force: true }); + } + return context; +}; +``` + +This materializes a local directory tree; it does not resolve a Git ref, pull OCI +content, validate a bundle digest, enforce read-only subtrees, or record provenance. +A CI checkout step or wrapper must prepare `fixtures/base`. The seed must not be +agent-writable. If checks need the final tree after the entire evaluation, copy selected +evidence to a case-specific artifact directory in `afterEach` before deleting the +workspace. + +### `promptfoo/assert-workspace.cjs` — proposed deterministic grader + +```js +const fs = require('node:fs/promises'); +const path = require('node:path'); + +const workspace = path.resolve(__dirname, '..', '.eval', 'workspace'); + +module.exports = async function assertWorkspace(_output, context) { + try { + const actual = await fs.readFile(path.join(workspace, 'result.txt'), 'utf8'); + const expected = String(context.vars.expected); + const pass = actual.trim() === expected; + return { + pass, + score: pass ? 1 : 0, + reason: pass + ? 'result.txt has the expected content' + : `result.txt was ${JSON.stringify(actual.trim())}`, + }; + } catch (error) { + return { pass: false, score: 0, reason: `result.txt unavailable: ${error.message}` }; + } +}; +``` + +For executable checks, this module can use `execFile`/`spawn` with a fixed executable +and literal argument array. That remains trusted assertion code running as the +Promptfoo process. It is not the gateway's separately sandboxed, credential-stripped, +structured `post_run` phase. + +## Lifecycle hooks: exact names and timing + +Two unrelated hook systems must not be conflated. + +### Promptfoo evaluation extensions + +Root `extensions` accepts JavaScript or Python functions. The exact lifecycle names +are `beforeAll`, `beforeEach`, `afterEach`, and `afterAll` +([configuration reference](https://www.promptfoo.dev/docs/configuration/reference/#available-hooks)). + +| Hook | Documented context | Relevant timing | Mutation contract | +|---|---|---|---| +| `beforeAll` | `{ suite }` | Once before the evaluation | May return selected mutated suite fields. | +| `beforeEach` | `{ test }` | Before one expanded evaluation step invokes the provider | Returning `{ test }` replaces the test context; a thrown reset error propagates. Current source calls it immediately before `runEvalInternal` ([source](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluator.ts#L3567-L3593)). | +| `afterEach` | `{ test, result }` | After the row is graded, before persistence in the normal non-deferred path | Only `result.namedScores`, `result.metadata`, and `result.response.metadata` are persisted; it cannot override `success`, `score`, or `response.output` ([mutation reference](https://www.promptfoo.dev/docs/configuration/reference/#aftereach)). | +| `afterAll` | `{ results, prompts, suite, evalId, config }` | Once after all rows | Side effects only; its return value is not persisted. | + +A path naming a custom function, such as +`file://./workspace-hooks.cjs:extensionHook`, receives every event as +`(hookName, context)`. A path whose function name is exactly one lifecycle name runs +only for that event and is called as `(context, { hookName })` +([implementation rules](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluatorHelpers.ts#L774-L816)). +Returned mutable fields are shallow-merged, not deep-merged +([extension mutation docs](https://www.promptfoo.dev/docs/configuration/reference/#extension-hooks)). + +`beforeEach` is invoked for each expanded `RunEvalOptions`, not merely once for the +original YAML object (**source-observed**). With multiple prompts, providers, or +repeats, the same logical YAML test can therefore be reset several times. The scheduler +runs `options.runSerially` steps first and other steps through a concurrency-limited +loop; global concurrency one makes both phases single-file +([scheduler source](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluator.ts#L4048-L4111), +[matrix partition](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/evaluator.ts#L4924-L4951)). + +### Claude Agent SDK provider hooks + +`providers[].config.hooks` is a different, SDK-native interception system for events +such as `PreToolUse` and `PostToolUse`. Promptfoo preserves SDK input/return shapes; +these callbacks are programmatic-only and must be defined in a JS/TS provider file, +not expressed as YAML functions. For example, `PostToolUse` can return +`updatedToolOutput` before the model sees a tool result +([Claude provider hooks](https://www.promptfoo.dev/docs/providers/claude-agent-sdk/#hooks)). +These hooks are useful inside a run, but they do not reset the shared workspace between +Promptfoo cases. + +## Provider comparison + +| Capability | Claude Agent SDK (`anthropic:claude-agent-sdk`) | OpenAI Codex SDK (`openai:codex-sdk`) | +|---|---|---| +| Working directory | `working_dir`; omitted means a provider-created temporary directory that is removed after the call. An explicit directory persists and is read-only by default ([docs](https://www.promptfoo.dev/docs/providers/claude-agent-sdk/#quick-start), [cleanup source](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/providers/claude-agent-sdk.ts#L1801-L1824)). | `working_dir`; omitted means the current directory. It must be in a Git repo unless `skip_git_repo_check: true` ([docs](https://www.promptfoo.dev/docs/providers/openai-codex-sdk/#quick-start)). | +| Write enablement | Add `Write`, `Edit`, `MultiEdit`, and optionally `Bash` via `append_allowed_tools`/`custom_allowed_tools`/`tools`; set `permission_mode`. | `sandbox_mode: workspace-write` is default; set `approval_policy: never` for unattended evaluation. | +| Filesystem isolation | Claude SDK `sandbox.enabled`; `failIfUnavailable`, network/socket policy, exclusions, and credential masking are explicit. | `sandbox_mode` controls filesystem access only; network/search, environment inheritance, and approvals are separate. | +| Extra paths | `additional_directories` | `additional_directories` | +| Fresh session default | Auto-generated SDK session; `persist_session` defaults true on disk, while `continue`/`resume` are opt-in. A new Promptfoo call is not by itself a fresh filesystem. | New ephemeral thread for each case by default. `persist_threads`, `thread_id`, and `thread_pool_size` opt into reuse; `deep_tracing` ignores all three and makes a fresh SDK client/thread ([thread docs](https://www.promptfoo.dev/docs/providers/openai-codex-sdk/#thread-management), [deep-trace caveat](https://www.promptfoo.dev/docs/providers/openai-codex-sdk/#deep-tracing)). | +| Provider response metadata | `metadata.toolCalls`, `skillCalls`, `numTurns`, `durationMs`, `durationApiMs`, `modelUsage`, permission denials, terminal reason, structured output, and surfaced assistant errors when present ([response construction](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/providers/claude-agent-sdk.ts#L2267-L2339)). | General Codex items are not copied into stable `metadata`; current `metadata` is skill-detection data when present. `output` is final text, `sessionId` is the thread, token usage/cost are separate, and `raw` serializes the SDK turn ([response construction](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/providers/openai/codex-sdk.ts#L2419-L2449)). | +| Tool/trajectory visibility | Completed tool calls are available directly in `metadata.toolCalls`; provider tracing also emits tool and turn spans. | Set `enable_streaming: true` for command, file-change, MCP, search, reasoning, message, and turn spans; set `deep_tracing: true` for CLI-native spans. | +| Full transcript | No stable full-transcript result. `raw` is the terminal SDK result; `metadata.toolCalls` carries tool I/O. Subagent text/thinking is omitted by default and requires `forward_subagent_text: true`, with documented redaction behavior ([tool tracking](https://www.promptfoo.dev/docs/providers/claude-agent-sdk/#tool-call-tracking)). | No stable normalized transcript result. Final text is `output`; streaming-mode `raw` includes provider-specific `items`, sanitized `reasoningTexts`, and `conversationMessages`, while traces carry operation events ([stream-result source](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/providers/openai/codex-sdk.ts#L1544-L1561)). | + +Neither provider materializes an application repository into an explicit directory. +Claude's automatic temporary directory is empty and deleted, and Codex defaults to the +current directory. An explicit `working_dir` means "operate here," not "make this path +fresh." The provider validates/accesses the directory after the evaluation extension +has prepared it +([Claude validation and `cwd`](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/providers/claude-agent-sdk.ts#L1801-L1844), +[Codex resolution and validation](https://github.com/promptfoo/promptfoo/blob/712a506de6412ca6879fe8dba1319569ea820cbf/src/providers/openai/codex-sdk.ts#L2222-L2273)). + +## Deterministic filesystem grading and result metadata + +A JavaScript assertion can return a boolean, number, or `{ pass, score, reason, +namedScores?, componentResults? }`; its `context` includes the prompt, vars, test, +provider, complete `providerResponse`, response metadata shortcut, and trace data when +enabled +([JavaScript assertion context](https://www.promptfoo.dev/docs/configuration/expected-outputs/javascript/#using-test-context)). +This supports deterministic checks of: + +- exact file content, mode, or absence; +- a parsed manifest or compiler output; +- a fixed executable's exit code/stdout/stderr; and +- Claude's `context.providerResponse.metadata.toolCalls`. + +Promptfoo can export complete evaluation data to JSON or one row per line to JSONL +([output formats](https://www.promptfoo.dev/docs/configuration/outputs/)). An +`afterEach` hook can add structured observations to `result.metadata`, add numeric +metrics to `result.namedScores`, or add provider-level details to +`result.response.metadata`; it cannot retroactively change the grade +([afterEach contract](https://www.promptfoo.dev/docs/configuration/reference/#aftereach)). + +Limitations of this native pattern: + +1. Assertions run with the Promptfoo process's authority, not a separate restricted + checker identity. +2. There is no built-in hidden-check-bundle timing boundary. Keeping the assertion + outside `working_dir` is not a guarantee that a shell-capable agent cannot read it. +3. There is no built-in structured post-run command/result schema, byte limit, file + promotion, digest, or artifact retention contract. +4. A fixed workspace has only one live final state. A later `beforeEach` overwrites it; + evidence that must survive must be copied or serialized per case. +5. `afterEach` cleanup is best-effort in the current evaluator because its exception + is caught. The next `beforeEach` and final `afterAll` should retry and throw, and the + surrounding job should independently verify cleanup when it is a release condition. +6. Provider and evaluation timeouts stop the call, but native composition does not + establish the gateway's process-tree/cgroup cleanup guarantee for arbitrary daemons + the agent launched. + +## OpenTelemetry, trajectory assertions, and transcript limits + +Root `tracing.enabled: true` creates a distinct trace per test-case execution and puts +`traceId` plus `evaluationId` on each result row. Promptfoo's built-in OTLP receiver is +configured under `tracing.otlp.http`; traces can be inspected in the UI, fetched through +`GET /api/traces/:traceId` or `GET /api/traces/evaluation/:evaluationId`, or exported as +JSON +([tracing overview](https://www.promptfoo.dev/docs/tracing/#built-in-provider-instrumentation), +[result-row linkage and API](https://www.promptfoo.dev/docs/tracing/#trace-linkage-on-result-rows), +[JSON export](https://www.promptfoo.dev/docs/tracing/#exporting-traces)). +JavaScript assertions receive trace spans as `context.trace`, and built-in +`trajectory:tool-used`, `trajectory:tool-args-match`, `trajectory:tool-sequence`, +`trajectory:step-count`, and `trajectory:goal-success` assertions consume normalized +span information +([traced-workflow assertions](https://www.promptfoo.dev/docs/tracing/#4-assert-on-traced-workflows)). + +Provider differences matter: + +- Claude emits an `invoke_agent` span, `gen_ai.turn N` spans, and a child span for each + completed tool call. `deep_tracing: true` asks the SDK subprocess to export native + model/tool/subagent spans to the receiver + ([Claude tracing](https://www.promptfoo.dev/docs/providers/claude-agent-sdk/#tracing)). +- Codex needs `enable_streaming: true` for Promptfoo to turn SDK events into command, + file-change, MCP, search, reasoning, message, and turn spans. `deep_tracing: true` + additionally injects OTEL context into the Codex CLI + ([Codex tracing](https://www.promptfoo.dev/docs/providers/openai-codex-sdk/#tracing-and-observability)). +- Codex streaming still returns only after the turn completes; it is event aggregation, + not live partial-token delivery to assertions + ([streaming behavior](https://www.promptfoo.dev/docs/providers/openai-codex-sdk/#streaming)). + +These traces are useful trajectory evidence, but they are not the proposed gateway's +bounded, normalized ATIF v1 artifact. No ATIF exporter or stable cross-provider full +transcript contract was found in the inspected Promptfoo docs/source (**source-review +finding, not a documented guarantee of absence**). The supported export is Promptfoo's +OpenTelemetry-shaped trace JSON. Provider `raw` payloads remain SDK-specific, and +neither provider documents them as a complete, size-bounded transcript schema. + +Treat trace redaction as defense in depth. Promptfoo warns that Codex CLI-native spans +created by `deep_tracing` are outside Promptfoo's stream-event sanitizer, and the OTLP +receiver's `redactAttributes` does not filter in-process built-in provider spans before +local storage +([Codex deep-trace warning](https://www.promptfoo.dev/docs/providers/openai-codex-sdk/#deep-tracing), +[trace redaction scope](https://www.promptfoo.dev/docs/tracing/#configuration-reference)). + +## Native composition versus the proposed gateway + +The accepted ADR now selects this native composition and explicitly rejects a gateway +for V1. It makes the disposable job the outer isolation/lifecycle boundary and assigns +exact source staging plus bounded diagnostics to job-owned glue +([native-execution decision](../decisions/0002-use-promptfoo-native-agent-execution.md#decision), +[evaluation plan contract](../plans/2026-09-18-0837-feat-promptfoo-coding-agent-evals-plan.md#product-contract)). +The table below compares that decision with the stronger responsibilities of the +earlier proposed gateway so the point at which a future external runner becomes +justified remains explicit. + +| Gateway responsibility | Native Promptfoo + fixed local path | Assessment for trusted local/CI | +|---|---|---| +| Prompt/test/provider matrices, repeats, grading, reports | Native Promptfoo strength | **Replace gateway portion; Promptfoo already owns this.** | +| Run Claude or Codex in an existing checkout | Built-in SDK providers with `working_dir` | **Sufficient.** | +| Fresh workspace per case | Extension/wrapper deletes and copies a fixture | **Sufficient as repo-owned glue**, provided reset failure aborts and the seed is protected. | +| Deterministic final-filesystem checks | Trusted JS assertion reads files or runs a fixed verifier | **Sufficient** for deterministic-only rows and trusted checks. | +| Basic tool/step evidence | Claude metadata plus OTLP traces; Codex streamed and optionally deep OTLP traces | **Usually sufficient**, if Promptfoo JSON/trace JSON is the accepted evidence format. | +| Exact Git/OCI/bundle acquisition and provenance | Not provided by these providers or extensions | Use CI checkout/container tooling for a simple local case; retain an external materializer when exact multi-source provenance is a requirement. | +| Read-only and writable source views in one composed tree | No first-class source/access-mode model | External runner required when access modes are security properties rather than fixture convention. | +| Agent isolation, resource ceilings, process-tree reaping | Partial provider-specific sandbox controls; no gateway-equivalent lifecycle contract | External sandbox/runner required for untrusted code or strict CPU/memory/PID/IO cleanup. | +| Separate agent/check credentials and network namespaces | Assertions share the host process/runtime; no hidden late bundle | External runner required for secret tests, phase separation, or policy-enforced egress. | +| Bounded output, requested files, digests, immutable artifacts, retention | Promptfoo outputs/traces, plus arbitrary glue | External evidence/artifact service required when these are contractual. | +| Stable normalized ATIF trajectory | OTLP spans and provider-specific raw/metadata | External normalization required if ATIF is mandatory. | +| Idempotency, cancellation, queue recovery, durable terminal states | Local process semantics only | External control plane required if ambiguous/retried infrastructure execution matters. | +| OMP execution and cross-agent parity | No built-in OMP provider among the two evaluated here | Custom provider or external adapter required. | + +## Application to the PR 679 Promptfoo experiment + +The authenticated +[`framework-parity/promptfoo/pr-679`](https://github.com/EntityProcess/wtg-ai-prompts-experiment/tree/main/framework-parity/promptfoo/pr-679) +experiment currently uses top-level Promptfoo `metadata` as an execution +configuration channel: + +- the + [`with-agentrules` suite](https://github.com/EntityProcess/wtg-ai-prompts-experiment/blob/main/framework-parity/promptfoo/pr-679/with-agentrules.suite.yaml) + puts repository URLs, revisions, a workdir, Git-cache configuration, and a + skills-config path under `metadata`; +- `setup_environment_extension.ts` reads `suite.metadata.environment`, creates a + timestamped workspace, and publishes its path through process environment + variables; and +- `skills_extension.ts` reads `suite.metadata.skills`, while the PI provider + consumes the resulting workspace and manifest through `workdirEnv` and + `environmentManifestEnv`. + +That works, but Promptfoo documents top-level `metadata` as arbitrary data stored +with the eval config, not as a workspace lifecycle schema. The fixed-workspace +composition provides a cleaner replacement: + +1. Move repository and exact-revision declarations to the checked-in source + catalog owned by the evaluation harness. +2. Let the job launcher stage the CargoWise seed before Promptfoo starts. +3. Put `source_id` or a workspace-profile ID in suite `defaultTest.vars`; the + `beforeEach` extension copies that seed to `./.eval/workspace`. +4. Configure a built-in agent with + `working_dir: ./.eval/workspace`. If the PI provider remains, give it the same + ordinary `working_dir` config field instead of discovering a path through + suite metadata and environment-variable indirection. +5. Treat the with-skill and without-skill variants as distinct staged workspace + profiles. Skill materialization belongs in source/profile setup, or in the + built-in provider's documented skill configuration when that provider owns + skill loading. +6. Keep Promptfoo `metadata` descriptive only: source PR, source eval, experiment + tags, and other report annotations. + +The shared PR 679 cases currently model-grade the agent's textual review and do +not inspect a mutable final filesystem, so Promptfoo's deferred-grading behavior +does not invalidate them. If those cases later add live-workspace assertions, +the deterministic phase must persist row evidence before a separate model- +grading evaluation, as described above. + +## Narrow recommendation + +Adopt the native composition first for evaluations that meet **all** of these +conditions: + +- one trusted local/CI runner owns the workspace and credentials; +- inputs are already checked out or can be copied from a protected local fixture; +- one shared `./.eval/workspace` is acceptable with `maxConcurrency: 1`; +- tests can use deterministic filesystem assertions without same-row deferred + model grading; +- provider-specific Claude/Codex metadata plus Promptfoo OTLP trace JSON is adequate; +- caching is disabled so every row actually executes; and +- the provider sandbox plus the CI/container boundary is an acceptable risk boundary. + +Under those conditions, a gateway adds little evaluation value. Keep the composition +small: one fail-closed `beforeEach` materializer, one deterministic assertion module, +one built-in provider, and optional tracing. Do not recreate an API, job queue, upload +protocol, or artifact store around a local run. + +Retain or introduce an external gateway/runner only when at least one concrete +requirement crosses that boundary: untrusted agent execution; multi-source Git/OCI +composition with verified provenance; read-only mount enforcement; hidden checks; +credential/network phase separation; strict cgroup/process cleanup; durable bounded +artifacts; idempotent cancellation/recovery; mandatory ATIF; OMP support; or execution +on a machine other than the Promptfoo process. Remote multi-tenancy is one reason for +those contracts, not a prerequisite for them. diff --git a/docs/research/source-credential-broker-precedents.md b/docs/research/source-credential-broker-precedents.md new file mode 100644 index 00000000..88f65533 --- /dev/null +++ b/docs/research/source-credential-broker-precedents.md @@ -0,0 +1,220 @@ +# Source credential broker precedents + +## Status + +This note records credential and policy precedents for the rejected hosted +gateway design. Its primary-source findings remain relevant if AllAgents later +accepts remote callers or must broker source and model credentials across a +tenant boundary. + +The current trusted local/CI design has no gateway credential broker. Disposable +job bootstrap resolves exact sources and materializes job-private seeds before +Promptfoo starts. Acquisition credentials must not enter the seed, mutable +workspace, test variables, result metadata, or trace attributes, and should be +removed from the environment before agent execution whenever the source +transport permits it. + +Promptfoo then invokes its built-in Claude or Codex provider in +`.eval/workspace`. The outer disposable container or VM is the selected +isolation boundary; this is not equivalent to the operator-authorized, +phase-separated credential and network controls described below. + +See [ADR 0002](../decisions/0002-use-promptfoo-native-agent-execution.md) and +[Promptfoo native agent workspaces](./promptfoo-native-agent-workspaces.md) for +the current decision. The detailed broker design below is historical +future-service research, not a V1 implementation contract. + +## Precedents + +### Git credential helpers and Git Credential Manager + +**Trust boundary.** Git credential helpers are external programs. Git invokes a +configured helper through the shell, supplies an operation and credential context, +and stops consulting helpers after it has a username and a non-expired password +([Git `gitcredentials`](https://git-scm.com/docs/gitcredentials#Documentation/gitcredentials.txt-helper)). +The scriptable `git credential fill` interface sends the repository context on +standard input and returns the resolved username and password on standard output +([Git `git-credential`](https://git-scm.com/docs/git-credential#_typical_use_of_git_credential)). +Consequently, the Git/acquisition process receives the resulting bearer secret; +the helper is not a membrane that makes an untrusted caller safe. + +GCM is an implementation of this local contract, not a required remote service. +Its executable is a console application; on every invocation it reads Git's +request from standard input, retrieves or generates a credential, serializes the +credential to standard output, and terminates +([GCM architecture, “Command execution”](https://github.com/git-ecosystem/git-credential-manager/blob/main/docs/architecture.md#command-execution)). +Git calls it implicitly, and later Git commands reuse stored credentials or tokens +while they remain valid +([GCM README, “How to use”](https://github.com/git-ecosystem/git-credential-manager#how-to-use)). +GCM can put credentials in OS-controlled stores such as Windows Credential +Manager or macOS Keychain, use Secret Service or GPG-backed storage, use Git's +ephemeral cache, or disable its store entirely +([GCM credential stores](https://github.com/git-ecosystem/git-credential-manager/blob/main/docs/credstores.md)). + +**Lifetime.** Helper-process lifetime and credential lifetime are separate. GCM +exits after each request, while the selected store controls token persistence. +Git's built-in cache is an optional local daemon reachable over a Unix-domain +socket restricted to the current user; it forgets credentials after 900 seconds +by default or sooner if the daemon dies +([Git `git-credential-cache`](https://git-scm.com/docs/git-credential-cache#_description), +[options](https://git-scm.com/docs/git-credential-cache#_options)). This is a +local process/socket boundary, not a remotely reachable credential service. + +**Relevance.** Git helpers and GCM prove that a local credential provider can be +an on-demand process rather than a network service. The AllAgents materializer +does not inherit or invoke the host's configured helper chain. It creates a +closed helper for the selected deployment credential, invokes Git with an +isolated home and system/global configuration disabled, and removes the helper +before returning. The coding-agent runtime inherits neither the helper +configuration nor its credential. + +### SSH agent forwarding + +**Trust boundary.** `ssh-agent` holds private keys and exposes operations through +a Unix-domain socket. With forwarding, private keys and passphrases do not cross +the network; the SSH connection carries requests to the local agent and returns +the results +([OpenSSH `ssh-agent`](https://man.openbsd.org/ssh-agent#DESCRIPTION)). The +forwarded socket is nevertheless an authentication capability. OpenSSH warns +that anyone able to bypass the remote socket's permissions can use loaded +identities to authenticate even though they cannot extract the key material +([OpenSSH `ForwardAgent`](https://man.openbsd.org/ssh_config#ForwardAgent)). +GitHub gives the same operational warning: a trusted server can use the keys as +the user while the connection is established, so forwarding should be enabled +only for specifically trusted hosts +([GitHub, “Using SSH agent forwarding”](https://docs.github.com/en/authentication/connecting-to-github-with-ssh/using-ssh-agent-forwarding#setting-up-ssh-agent-forwarding)). + +**Lifetime.** The remote forwarding capability lasts for the SSH connection. +The underlying identity may live longer: `ssh-agent` has no default maximum +identity lifetime unless configured, while `ssh-add -t` can impose one and +`ssh-add -c` can require confirmation for each use +([OpenSSH `ssh-agent -t`](https://man.openbsd.org/ssh-agent#t), +[OpenSSH `ssh-add`](https://man.openbsd.org/ssh-add#c)). + +**Relevance.** Agent forwarding is precedent for reusing a local identity +without copying the long-lived private key, but it is not selected for +AllAgents direct Git acquisition, which accepts canonical HTTPS repository URLs +only. A remote process with the forwarded socket could authenticate as the +user, so the socket must never reach a remote worker, setup code, or the +coding-agent runtime. + +### GitHub App installation tokens and Actions checkout + +**Trust boundary.** A GitHub App uses an RS256 JWT, created with the App private +key, to request an installation access token +([GitHub, “Generating a JSON Web Token”](https://docs.github.com/en/apps/creating-github-apps/authenticating-with-a-github-app/generating-a-json-web-token-jwt-for-a-github-app)). +The mint request can narrow the token to selected repositories and permissions, +and GitHub will not grant repositories or permissions beyond those already +granted to the installation +([GitHub, “Generating an installation access token”](https://docs.github.com/en/apps/creating-github-apps/authenticating-with-a-github-app/generating-an-installation-access-token-for-a-github-app#generating-an-installation-access-token)). +This separates high-value issuer material from the disposable credential handed +to a worker. + +GitHub Actions applies that model per job. GitHub creates a unique +`GITHUB_TOKEN` before each job; it is a GitHub App installation token limited to +the workflow repository, with permissions reducible through workflow policy +([GitHub Actions `GITHUB_TOKEN`](https://docs.github.com/en/actions/concepts/security/github_token#about-the-github_token)). +`actions/checkout` uses the token for Git commands, stores persisted credentials +in a separate file under `RUNNER_TEMP`, references that file from Git config, and +removes the references and file during post-job cleanup +([checkout README, v6 credential storage](https://github.com/actions/checkout#checkout-v6), +[checkout credential setup](https://github.com/actions/checkout/blob/main/src/git-auth-helper.ts#L329-L436), +[checkout credential cleanup](https://github.com/actions/checkout/blob/main/src/git-auth-helper.ts#L475-L510)). + +**Lifetime.** A normal GitHub App installation token expires after one hour +([GitHub installation token documentation](https://docs.github.com/en/apps/creating-github-apps/authenticating-with-a-github-app/generating-an-installation-access-token-for-a-github-app#generating-an-installation-access-token)). +The Actions token expires when its job finishes or at its effective maximum +lifetime; GitHub documents a six-hour maximum on GitHub-hosted runners and at +most 24 hours of refresh for longer self-hosted jobs +([GitHub Actions `GITHUB_TOKEN`](https://docs.github.com/en/actions/concepts/security/github_token#about-the-github_token)). +Checkout's credential file is a convenience capability inside that job, not a +long-term credential store, and its post-job deletion is defense in depth rather +than the token's revocation mechanism. + +**Relevance.** GitHub App installation tokens are useful deployment inputs +because repository scope, read-only contents permission, and expiry are enforced +by GitHub. Token minting remains outside the materializer contract. If an +operator supplies such a token through the configured environment reference, +the materializer still treats it as a phase-scoped acquisition secret and binds +the resulting source identity to the session's effective descriptor digest. + +### BuildKit secret and SSH mounts + +**Trust boundary.** BuildKit distinguishes secret delivery from ordinary build +arguments and environment variables, which can persist in an image. A secret +mount makes a client-provided secret temporarily available only to a particular +build instruction; an SSH mount supplies an agent socket or key and is intended +for cases such as fetching private Git repositories +([Docker build secrets](https://docs.docker.com/build/building/secrets/#types-of-build-secrets)). +`RUN --mount=type=secret` makes the value available without baking it into the +image, while `RUN --mount=type=ssh` exposes SSH-agent access through a mounted +socket +([Dockerfile secret mount](https://docs.docker.com/reference/dockerfile/#run---mounttypesecret), +[Dockerfile SSH mount](https://docs.docker.com/reference/dockerfile/#run---mounttypessh)). + +The isolation guarantee is intentionally narrow. BuildKit states that secret +values must not be written to disk or included in cache checksums and that an +untrusted frontend cannot access forwarded SSH private keys; it also states that +a container explicitly run with a secret mount can read that secret +([BuildKit security boundary](https://github.com/moby/buildkit/blob/master/PROJECT.md#security-boundary)). +A mount therefore limits *where and when* a capability appears; it does not make +code within the mounted step trustworthy. + +**Lifetime.** The secret mount is available for the duration of its build +instruction, rather than becoming part of the resulting image +([Docker build secrets](https://docs.docker.com/build/building/secrets/#secret-mounts)). +When an agent socket is supplied, SSH access is available for the mounted +instruction without adding the private key to the image +([Dockerfile SSH mount](https://docs.docker.com/reference/dockerfile/#run---mounttypessh)). + +**Relevance.** AllAgents uses the same phase-scoping pattern: inject a token only +into the trusted source-acquisition operation, then remove the +mount/socket/environment before agent execution. Like BuildKit, this delivery +mechanism does not mint credentials and does not make code with access to the +secret trustworthy. + +## Historical gateway recommendation + +If a future hosted execution service needs the stronger boundary studied here, +the prior recommendation was: + +1. Use the normative `AgentRunRequest v1`, `PostRunSpec v1`, and + `AgentRunResult v1` contracts rather than a second credential-specific shape. +2. Configure source matchers, redirect policy, OCI auth realms, credential + mapping, runtime profiles, network policies, and secret handles out-of-band. +3. Combine authenticated caller identity with canonical Git URL or OCI + repository. Require exactly one source route; reject zero or ambiguous + matches before any network access. +4. Deny source redirects and pin any OCI bearer-token realm through the matched + operator route. Repeat source authorization on every immutable-cache lookup. +5. Give the matched source credential only to bounded private acquisition on a + miss; remove it before publishing the verified immutable generation. Never + store credentials or mutable trial state in the shared cache. +6. Resolve and persist an authorized immutable runtime-profile revision and + `profile_digest` at admission. Dispatch uses only that revision, re-verifies + profile/image, tool/service implementation, and applicable service-image + digests, and returns them in `RuntimeProvenance`. Give the agent model access + only through its run-scoped local proxy; never expose model credentials in + the process, environment, workspace, or artifacts. +7. Keep acquisition, task-service, agent, and post-run networks phase-separated + and default-drop. Apply only the named policy authorized for each phase. +8. Tear down the agent process/cgroup and network namespace, verify descendants + are absent, and remove credentials before any hidden-bundle bytes are + materialized. Only then authorize/materialize the optional `BundleReference` + and run structured commands with declared source modes intact. +9. Seal bounded complete/partial raw evidence as observations finish. Promptfoo + alone owns pass/fail and reward. +10. Destroy the workspace, agent home, containers, temporary credential + material, and network namespaces and release leases before publishing + `AgentRunResult v1`. Cleanup failure is `infrastructure_error` with sealed + partial evidence. +11. Tenant/run-authorize status, cancellation, result, and result-artifact + operations; tenant-authorize bundle upload. Dereference expiring result + artifacts through the authenticated endpoint and verify streamed size and + digest. + +A central token minter, delivery lease, or generic credential-broker protocol +remains out of scope until remote or otherwise untrusted workers create a +concrete need. The future-service rule is: **credentials exist only in the +phase that consumes them and never enter raw post-run evidence or durable trial +state.** diff --git a/docs/research/workspace-contract-incumbents.md b/docs/research/workspace-contract-incumbents.md new file mode 100644 index 00000000..e27e5774 --- /dev/null +++ b/docs/research/workspace-contract-incumbents.md @@ -0,0 +1,175 @@ +# Workspace contract incumbents + +## Current conclusion + +No examined incumbent is needed between Promptfoo and the coding agents for the +initial trusted local/CI scope. Promptfoo's built-in Claude Agent SDK and Codex +SDK providers already accept a prepared `working_dir`, while Promptfoo lifecycle +hooks can reset that directory around each serial evaluation row. + +The selected boundary is one disposable job: + +- job bootstrap resolves exact sources and creates read-only seeds; +- `beforeEach` privately copies the selected seed to `.eval/workspace`; +- the built-in provider runs there; +- Promptfoo assertions inspect the final filesystem and native trace; and +- `afterEach` removes the workspace before the next row. + +Promptfoo owns matrices, repetitions, assertions, pass/fail, rewards, metrics, +OpenTelemetry traces, and result presentation. The extension owns only +workspace setup, bounded diagnostics, reset, and cleanup. A container or VM owns +the outer process and filesystem isolation boundary. + +The gateway, source-builder, snapshot, and long-lived session contracts +evaluated below are rejected for V1. Their incumbent comparisons remain useful +if a future remote multi-tenant execution service needs stronger authorization, +credential, artifact, cancellation, or recovery boundaries. + +The current decision is +[ADR 0002](../decisions/0002-use-promptfoo-native-agent-execution.md), supported +by [Promptfoo native agent workspaces](./promptfoo-native-agent-workspaces.md). + +## Historical contract evaluated + +The comparisons below originally evaluated a product execution stack with a +separate source builder, an immutable OCI handoff, and a long-lived gateway. +That architecture is rejected for the chosen evaluation-only scope. Statements +below that prescribe source descriptors, snapshot manifests, builder ownership, +or gateway behavior are retained as historical comparison, not current +recommendations. Their primary-source descriptions of incumbents remain useful. + +## Incumbent comparison + +### Harbor: benchmark materialization incumbent, not a direct invocation schema + +Harbor absolutely materializes runnable workspaces. Its `--repo` input clones a benchmark repository from GitHub, GitLab, or Hugging Face, optionally pinned to a branch, tag, or commit. Each selected task then supplies an instruction, verifier, and environment. The environment can be built from a Dockerfile or Compose file or pulled as a prebuilt image through `environment.docker_image`; `environment.workdir` selects the command working directory. A `BaseEnvironment` provider starts that filesystem and exposes execution and file-transfer operations to the agent and verifier. + +The important distinction is between two repositories that coding benchmarks often collapse: + +1. the **benchmark/task repository**, selected by Harbor `--repo`, which contains `task.toml`, instructions, environment definitions, and tests; and +2. the **target application repository**, which is normally baked into the task image or acquired by task-authored environment setup. + +Harbor has a strong, reusable contract for the first item and for the resulting runnable environment. Its Git dataset identifier does not independently describe an arbitrary set of target repositories, their checkout destinations, source credentials, or the resolved provenance returned to a caller. The prebuilt image field likewise selects the whole task environment rather than a workspace source artifact with a separately verified manifest. An AllAgents source artifact may preserve normalized offline Git history, but it still excludes tools, services, verifier assumptions, and runtime configuration. + +AllAgents should therefore follow Harbor's **architecture** for benchmark interoperability: immutable task packages, prebuilt OCI environments, explicit workdir/resources/network policy, isolated trials, and verifier separation. A Harbor adapter can compile a selected task and environment into the AllAgents execution inputs. The Harbor task schema should not replace the smaller direct-execution descriptor used when a caller supplies target repositories or a workspace snapshot without a benchmark package. + +Harbor's newer [Agent Sandbox Protocol (ASP)](https://docs.harborframework.com/core-concepts/sandboxes/asp) is a separate layer. Its draft v0 `.asp.json` describes an already provisioned sandbox reachable over SSH and supplies an absolute sandbox workspace path. It standardizes harness-to-sandbox execute/read/write behavior, explicitly leaving provisioning to the orchestrator. ASP may become a useful southbound runtime adapter, but it does not specify how Git or OCI content becomes that workspace. + +Primary evidence: [Git datasets](https://docs.harborframework.com/core-concepts/datasets/git-repos), [task packages](https://docs.harborframework.com/core-concepts/tasks/overview), [environment materialization](https://docs.harborframework.com/core-concepts/tasks/environment), [task configuration](https://docs.harborframework.com/core-concepts/tasks/configuration), pinned [task config](https://github.com/harbor-framework/harbor/blob/cdb76bae6dc88d5bca1c8f0754bbba300d6574b4/src/harbor/models/task/config.py), [job config](https://github.com/harbor-framework/harbor/blob/cdb76bae6dc88d5bca1c8f0754bbba300d6574b4/src/harbor/models/job/config.py), and [Git acquisition](https://github.com/harbor-framework/harbor/blob/cdb76bae6dc88d5bca1c8f0754bbba300d6574b4/src/harbor/registry/client/git_repo.py). + +### Hugging Face and SWE-bench: registry plus specialized materializer + +Hugging Face also participates in real workspace materialization, but the responsibility is split. The Hub stores every dataset as a Git repository. A SWE-bench dataset row then identifies the target GitHub repository with `repo`, pins its state with `base_commit`, and can provide `environment_setup_commit`, patches, tests, and issue text. The SWE-bench harness converts that benchmark record into layered Docker artifacts—base, repository environment, and per-instance images—then starts the instance image, applies the model patch, runs tests, and grades the result. + +That is a concrete and widely used Git-to-OCI workspace pipeline. Hugging Face itself supplies registry and dataset-repository contracts; SWE-bench supplies the coding-task schema and execution harness. Neither exposes one general multi-repository invocation schema. SWE-bench is intentionally specialized to one repository/base commit and its test-transition metadata. + +For compatibility, an AllAgents benchmark adapter should map a SWE-bench `repo` plus `base_commit` to repository mode. A prepared instance image must instead pass through a future task/environment boundary, or an importer must extract its checkout into a `workspaceSnapshot` with a separately verified workspace manifest. The importer may preserve normalized Git history for offline evaluation, but it must remove remotes, credentials, and unsafe administrative state. The adapter must never register the runnable image itself as a workspace snapshot. This is an ingestion mapping, not a reason to replace the direct descriptor: AllAgents still needs multiple destinations, Git-or-OCI selection, strict credential and egress policy, continuation binding, and resolved attachment provenance. + +Primary evidence: the [SWE-bench dataset schema on Hugging Face](https://huggingface.co/datasets/princeton-nlp/SWE-bench), [Hugging Face dataset repository model](https://huggingface.co/docs/hub/datasets-overview), and [SWE-bench evaluation harness](https://www.swebench.com/SWE-bench/reference/harness). + +### Devfile 2.3: closest portable source-layout precedent + +Devfile is the strongest portable comparison. It is a CNCF Sandbox project with an open governance process and documented implementations including Eclipse Che and `odo`. Its normative schema defines `projects[]` with: + +- a required project `name`; +- `git.remotes`, mapping remote names to URLs; +- optional `checkoutFrom.remote` and `checkoutFrom.revision`; +- optional relative `clonePath`, defaulting to the project name; and +- ZIP sources as an alternative to Git. + +This is a clear semantic match for source URL, revision, and destination. Devfile also defines a projects root and projects are mapped into runtime components through `sourceMapping`. + +It is not a safe wholesale replacement: + +- The schema permits multiple named remotes where AllAgents deliberately accepts one canonical source URL per repository. +- Devfile's revision description permits the default branch when the requested revision is absent or not found; AllAgents requires a supplied ref to fail closed. +- `clonePath` is optional and defaults from the project name; AllAgents requires an explicit collision-checked destination. +- ZIP sources have no required content digest. There is no OCI workspace-source variant. +- Devfile has no standard resolved-commit result, canonical source-visible manifest, generation identity, or attachment-commit acknowledgement. +- Runtime implementations own credential behavior. The DevWorkspace Operator, for example, may expose configured Git credentials to workspace containers; that is weaker than acquisition-only credentials. +- Devfile lifecycle events and component `sourceMapping` configure a development environment. They do not define the AllAgents timing rule that source is attached before the agent starts and metadata appears only after attachment completes. + +AllAgents should cite and follow Devfile's vocabulary where it fits, while preserving stricter semantics: + +| AllAgents | Devfile precedent | Decision | +| --- | --- | --- | +| `url` | `projects[].git.remotes.` | Keep one canonical public HTTPS URL; do not import named-remote ambiguity. | +| `ref` | `checkoutFrom.revision` | Keep `ref`, strict resolution, and requested/resolved identity. Do not adopt fallback-on-miss behavior. | +| `destination` | `clonePath` | Treat this as the closest direct precedent, but keep it required and collision checked. | +| `workspaceRoot` | projects root / `$PROJECTS_ROOT` | Same conceptual root; keep the typed logical value rather than exposing a container path. | +| `workspacePath` | component `sourceMapping` is adjacent, not equivalent | Keep manifest-verified relative path semantics. | +| resolved Git/OCI provenance | no equivalent | Retain the AllAgents result model. | + +Primary evidence: the pinned [Devfile 2.3 JSON Schema](https://github.com/devfile/api/blob/v2.3.0/schemas/latest/devfile.json), [schema reference](https://devfile.io/docs/2.3.0/devfile-schema), [project authoring guide](https://devfile.io/docs/2.3.0/adding-projects), [governance](https://github.com/devfile/api/blob/v2.3.0/GOVERNANCE.md), [CNCF project record](https://www.cncf.io/projects/devfile/), [documented users](https://devfile.io/docs/2.3.0/users-of-devfile), and the DevWorkspace Operator's pinned [Git-credential behavior](https://github.com/devfile/devworkspace-operator/blob/9df10c1baba8e7d88948a21077e3a06fd2cca639/docs/additional-configuration.adoc). + +### Development Containers: environment standard, not source standard + +The Development Container Specification assumes a project workspace/source folder already exists. It standardizes how that folder is mounted or opened in an image-, Dockerfile-, or Compose-based development container. Relevant fields include `workspaceMount`, `workspaceFolder`, image/build/Compose selection, Features, and ordered lifecycle commands. + +`workspaceFolder` is a container/editor path, not a source descriptor. `workspaceMount` is a runtime mount expression, not a portable source identity. Lifecycle hooks run after implementations have made source available, and the specification does not define repository URL/ref/destination, Git resolution, OCI workspace snapshots, or a resolved provenance response. Feature lockfiles add integrity for Dev Container Features, not for the application workspace. + +This is the most credible portable standard for a possible future **development-environment layer**, with official support listed for VS Code, Visual Studio, IntelliJ, the reference CLI, GitHub Codespaces, CodeSandbox, DevPod, and Ona. It should not be stretched into the acquisition layer. + +Primary evidence: pinned [normative specification](https://github.com/devcontainers/spec/blob/c95ffeed1d059abfe9ffbe79762dc2fa4e7c2421/docs/specs/devcontainer-reference.md), [JSON Schema](https://github.com/devcontainers/spec/blob/c95ffeed1d059abfe9ffbe79762dc2fa4e7c2421/schemas/devContainer.base.schema.json), [field and lifecycle reference](https://github.com/devcontainers/spec/blob/c95ffeed1d059abfe9ffbe79762dc2fa4e7c2421/docs/specs/devcontainerjson-reference.md), [supporting tools](https://github.com/devcontainers/spec/blob/c95ffeed1d059abfe9ffbe79762dc2fa4e7c2421/docs/specs/supporting-tools.md), and [contribution process](https://github.com/devcontainers/spec/blob/c95ffeed1d059abfe9ffbe79762dc2fa4e7c2421/CONTRIBUTING.md). + +### Daytona: runtime provider with clone operations + +Daytona is the closest field-level operational match: its Git clone operation accepts `url`, `path`, optional branch or commit, credentials, depth, and an insecure-TLS option. But this is an imperative operation against an already-created Daytona sandbox. It does not standardize multi-source declaration, strict canonicalization, immutable result provenance, or committed attachment timing. Its per-operation credentials and optional TLS bypass also conflict with the AllAgents trust boundary. Older Daytona workspace models coupled repository metadata, devcontainer/build configuration, and provider workspace state, illustrating the portability cost of adopting a vendor workspace object. + +Primary evidence: Daytona [Git operations](https://www.daytona.io/docs/en/git-operations), and pinned Daytona [workspace](https://github.com/daytonaio/daytona/blob/dfb50e8a31e9a93b31181113d7b44b657cf27168/pkg/models/workspace.go), [repository](https://github.com/daytonaio/daytona/blob/dfb50e8a31e9a93b31181113d7b44b657cf27168/pkg/apiclient/model_git_repository.go), and [workspace-creation](https://github.com/daytonaio/daytona/blob/dfb50e8a31e9a93b31181113d7b44b657cf27168/pkg/apiclient/model_create_workspace_from_git_repository.go) models. + +### GitHub Codespaces and Gitpod Classic: lifecycle precedents, not portable standards + +GitHub Codespaces creates a managed environment in the context of a GitHub repository. Its repository-scoped REST endpoint accepts `ref`, machine/location choices, `devcontainer_path`, `working_directory`, idle timeout, and retention. The repository is implied by the endpoint and the service delegates environment setup to Dev Containers. This is strong evidence for resolving source before environment startup and for treating working directory separately from source identity. It is not suitable as the AllAgents descriptor because it is GitHub-specific, single-repository, and does not expose the same resolved Git/OCI provenance or attachment transaction. + +Gitpod Classic/Ona similarly combines a context URL with workspace initialization and supports `additionalRepositories` plus checkout locations in `.gitpod.yml`. It is a useful product precedent for multiple checkouts, but its API and configuration are service-specific and mix source, image/build, and task lifecycle concerns. + +Primary evidence: GitHub's [create-codespace REST operation](https://docs.github.com/en/rest/codespaces/codespaces?apiVersion=2022-11-28#create-a-codespace-in-a-repository), [Codespaces CLI](https://cli.github.com/manual/gh_codespace_create), [Dev Container introduction](https://docs.github.com/en/codespaces/setting-up-your-project-for-codespaces/adding-a-dev-container-configuration/introduction-to-dev-containers), Gitpod Classic's [public API](https://ona.com/docs/classic/user/references/gitpod-public-api), and [`.gitpod.yml` reference](https://ona.com/docs/classic/user/references/gitpod-yml). + +## Normative artifact standards + +The source descriptor should remain AllAgents-owned, but its immutable artifact identities should not be invented locally. + +For OCI snapshots, the [OCI Image Specification 1.1.1 descriptor](https://github.com/opencontainers/image-spec/blob/v1.1.1/descriptor.md) defines content identity with media type, digest, and size, including verification against the digest. The [image manifest](https://github.com/opencontainers/image-spec/blob/v1.1.1/manifest.md) defines the config descriptor and ordered layer descriptors. Those are the correct normative identities for the snapshot artifact. OCI does not define the source-visible workspace tree, destination layout, logical working directory, or the semantic state of an embedded Git repository. The AllAgents workspace manifest therefore declares each repository root as tree-only or history-bearing. A history-bearing root records its resolved commit and canonical object-set digest, while the runner verifies the offline `.git` state and absence of remotes. + +For Git, a full commit object ID is the resolved source identity. The request still needs the original ref because a branch/tag name and its resolved commit answer different audit questions. OCI snapshot provenance instead retains the verified `snapshotName` and `imageManifestDigest`; each history-bearing root adds only its destination, resolved commit, and object-set digest, with no repository URL or requested ref. Neither Git nor OCI defines when a runner has successfully attached that content, so `effectiveDescriptorDigest`, `generationId`, `sourceIdentity`, `workingDirectory`, and `workspaceManifestDigest` must remain AllAgents result fields. + +## Current recommendation + +Adopt the native Promptfoo boundary for initial evaluation work: + +- **Evaluation owner:** Promptfoo owns prompt/provider matrices, repeats, + assertions, grading, rewards, metrics, traces, and result presentation. +- **Agent invocation:** use Promptfoo's built-in Claude Agent SDK and Codex SDK + providers rather than a custom provider or AllAgents adapter. +- **Working directory:** configure every write-capable provider with + `working_dir: ./.eval/workspace`. +- **Workspace lifecycle:** stage exact read-only seeds once per disposable job; + `beforeEach` creates a private writable copy and `afterEach` removes it. +- **Concurrency:** set `evaluateOptions.maxConcurrency: 1`; parallelize only by + running independent disposable jobs with separate `.eval` roots. +- **Caching:** disable Promptfoo response caching so a cached response cannot + bypass the agent and leave filesystem assertions grading an unrelated + workspace. +- **Judgment:** run deterministic command and file assertions after the provider + returns and before workspace teardown. +- **Trace:** retain Promptfoo OpenTelemetry data and use its built-in + `trajectory:*` assertions. Do not add ATIF without a named interoperability + consumer. +- **Isolation:** use a disposable rootless container or VM for write-capable + evals. Provider sandboxes are not a hostile tenant boundary. +- **Future service:** reconsider a gateway only for remote callers, mutually + untrusted tenants, centrally enforced network policy, credential brokering, + secret verifier injection, durable cancellation/recovery, or cross-machine + scheduling and retention. + +## Existing research status + +- [Promptfoo native agent workspaces](./promptfoo-native-agent-workspaces.md) is + the current execution research. +- [One-shot coding-agent gateway boundary](./one-shot-coding-agent-gateway-boundary.md) + records the rejected hosted-service alternative. +- [Harbor repository materialization](./harbor-repository-materialization.md) + remains useful evidence for task packages and separate verifiers. +- The incumbent descriptions above remain primary-source evidence; their + historical gateway prescriptions are not current implementation requirements.