diff --git a/.fullsend/.gitignore b/.fullsend/.gitignore new file mode 100644 index 000000000..302039c0d --- /dev/null +++ b/.fullsend/.gitignore @@ -0,0 +1,6 @@ +# fullsend-managed local resource cache — regenerated from the pinned URLs in +# lock.yaml on every run/lock; never committed. +.fullsend-cache/ + +# Local-only run scaffolding (env files with live secrets, launchers). +.local-run/ diff --git a/.fullsend/config.base.yaml b/.fullsend/config.base.yaml new file mode 100644 index 000000000..6416a39f5 --- /dev/null +++ b/.fullsend/config.base.yaml @@ -0,0 +1,26 @@ +# fullsend per-repo configuration +# https://github.com/fullsend-ai/fullsend +# +# This file configures fullsend for per-repo installation mode. +# See ADR 0033 for details. +# +# The "runtime" key selects which agent runtime runs the agents, claude +# (default when unset) or pi. For one run, the 'fullsend run --runtime' +# flag wins, then FULLSEND_RUNTIME, then this file. See docs/runtimes.md. +version: "1" +kill_switch: false +# The registered source is the LOCAL composing child harness (base + runner-local +# host_files); its `base:` pins the repo content by raw URL. See +# .fullsend/harness/verify-pr.yaml. The base URL's host prefix must remain listed +# in allowed_remote_resources below (base URLs are not inherited — they are +# validated against this allowlist). +agents: + - name: verify-pr + source: harness/verify-pr.yaml + - name: triage-security + source: harness/triage-security.yaml +allowed_remote_resources: + # Exactly one prefix: this harness's base URL is self-hosted on + # RHEcosystemAppEng and its children resolve relative to that base (no + # fullsend-ai remote overlay in the native v0.37.0 pinned-URL model). + - https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/ diff --git a/.fullsend/config.yaml b/.fullsend/config.yaml new file mode 100644 index 000000000..a0e77f668 --- /dev/null +++ b/.fullsend/config.yaml @@ -0,0 +1,20 @@ +# fullsend per-repo configuration (overlay) +# https://github.com/fullsend-ai/fullsend +# +# This file is the per-repo overlay for fullsend configuration. +# Base settings are provided by config.base.yaml (vendor preset). +# Values set here override the base layer. Omitted fields inherit +# from config.base.yaml, then from compiled-in code defaults. +# +# See ADR 0069 for the layered configuration model. +# +# The verify-pr agent is registered in config.base.yaml (the base layer), so +# `fullsend dispatch` / `fullsend run` resolve it while the vendor shim — which +# greps only THIS file for agents — does not dispatch it. CI-gated dispatch +# lives in .github/workflows/fullsend-verify-pr.yml. + +# Repo-specific inference override (hosted Vertex AI via Workload Identity +# Federation). The base preset supplies everything else. +inference: + project: it-gcp-tpa + wif_provider: projects/442181572212/locations/global/workloadIdentityPools/fullsend-inference/providers/gh-rhecosystemappeng-sdlc-plugin diff --git a/.fullsend/harness/triage-security.yaml b/.fullsend/harness/triage-security.yaml new file mode 100644 index 000000000..342851282 --- /dev/null +++ b/.fullsend/harness/triage-security.yaml @@ -0,0 +1,12 @@ +# Local composing child for the triage-security harness. +# +# The URL-pinned base supplies the tokenless triage contract. This local child +# supplies the runner-local pre-script output because Fullsend rejects an +# absolute host-file source inherited from a URL-sourced harness. The child +# mapping wins for this destination during composition. +base: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/harness/triage-security.yaml#sha256=110fb33d880caf1add20e73c8017451378966eeef52ef34aa988e41c35fc1b02 + +host_files: + - src: ${FULLSEND_RUN_DIR}/pre/triage-security-input.json + dest: /sandbox/workspace/.pre-script/triage-security-input.json + optional: true diff --git a/.fullsend/harness/verify-pr.yaml b/.fullsend/harness/verify-pr.yaml new file mode 100644 index 000000000..027f82d89 --- /dev/null +++ b/.fullsend/harness/verify-pr.yaml @@ -0,0 +1,68 @@ +# Local composing child for the verify-pr harness. +# +# fullsend v0.37.0 rejects an absolute host_files.src that is inherited from a +# URL-sourced harness, so the runner-local credential and pre_script-output +# mounts cannot live in the URL-pinned base (harness/verify-pr.yaml). This local +# child pins that base by raw URL (so the pinned bytes are byte-identical across +# local and CI) and adds those absolute mounts directly. host_files are +# concatenated base+child, deduplicated by dest with the child winning (ADR-0045; +# harness-fields.md). Absolute host_files declared directly in this local child +# are NOT URL-sourced, so they are accepted. +# +# ONE child serves BOTH environments — the mounts are env-expanded and the OIDC +# token is optional: +# GOOGLE_APPLICATION_CREDENTIALS local: service-account key JSON +# CI: WIF external_account config +# GCP_OIDC_TOKEN_FILE local: unset → optional mount skipped +# CI: runner-refreshed OIDC token file +# Only the runtime environment differs — fullsend's sanctioned "same harness, +# different runtime env" model (ADR-0055; running-agents-locally.md). +# +# PIN: the base cannot be a cross-boundary local path — fullsend rejects a +# `base:` that escapes the .fullsend workspace root, and a URL base is required +# anyway so the base's relative children resolve as pinned raw URLs. The base is +# pinned to a commit SHA (not a branch) plus the base file's sha256; both local +# and CI resolve these exact bytes. lock.yaml freezes every child SHA256. +# The base URL host prefix must stay listed in config.yaml allowed_remote_resources +# (base URLs are validated against that allowlist, not inherited). +# Re-pin after any base edit: push harness/verify-pr.yaml, set the SHA to the new +# commit + the sha256 to `shasum -a 256 harness/verify-pr.yaml`, `fullsend lock`. +base: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/harness/verify-pr.yaml#sha256=3c9dc221e2a37e7aa7ff294381b6b5b1580771ba70432c42924fc3a85f078e63 + +# CEL trigger (ADR 0061) — MUST be declared here, on the composing child, not +# only on the base. `fullsend dispatch` registers this child (config.base.yaml +# source: harness/verify-pr.yaml resolves under .fullsend/) and composes the +# base, but mergeBaseIntoChild (internal/harness/compose.go) carries `role` and +# other scalars from base->child while intentionally NOT carrying `trigger`. A +# trigger set only on the base is therefore inert; ListTriggeredHarnesses reads +# the composed child's Trigger. Kept byte-identical to the base's trigger (which +# documents intent and covers direct-base consumption). Fires on a PR +# change_proposal that is opened / synchronized / reopened — the dispatch-path +# equivalent of the CI poller's `on: pull_request` types in +# .github/workflows/fullsend-verify-pr.yml. +trigger: | + event.entity.kind == "change_proposal" && + event.transition.kind in ["synchronized", "opened", "reopened"] + +host_files: + - src: ${GOOGLE_APPLICATION_CREDENTIALS} + dest: /tmp/.gcp-credentials.json + - src: ${GCP_OIDC_TOKEN_FILE} + dest: /sandbox/workspace/.gcp-oidc-token + optional: true + # pre_script prefetch output — pre-verify-pr.sh writes /tmp/fullsend-pre-output + # on the runner; mounted read-only into the sandbox (optional — absent until + # the pre_script runs, and skipped on an ADR-0072 skip). + - src: /tmp/fullsend-pre-output/verify-pr-input.json + dest: /sandbox/workspace/.pre-script/verify-pr-input.json + optional: true + # Concatenated failed-check logs — pre-verify-pr.sh always writes this file + # (empty when no check failed) next to verify-pr-input.json. correctness.md + # Check 1b reads it on a FAIL via the bundle's github.check_run_logs_path (dest + # below); the large log text stays off the input bundle and off the agent's + # context until then. host_files mounts single files only, so per-check logs + # are concatenated into this one file. optional: absent until the pre_script + # runs, and skipped on an ADR-0072 skip. + - src: /tmp/fullsend-pre-output/check-run-logs.txt + dest: /sandbox/workspace/.pre-script/check-run-logs.txt + optional: true diff --git a/.fullsend/lock.yaml b/.fullsend/lock.yaml new file mode 100644 index 000000000..9e276e08e --- /dev/null +++ b/.fullsend/lock.yaml @@ -0,0 +1,540 @@ +# Generated by fullsend lock — DO NOT EDIT +version: 1 +generated_at: 2026-09-04T13:39:01.038924Z +harnesses: + triage-security: + source: harness/triage-security.yaml + sha256: eb1329bb309c00c339ca34e7da54bafd5faf2e5ef0a3f11bf331faaa8dcd45ca + resolved_at: 2026-09-30T11:35:13.025497Z + dependencies: + - field: base + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/harness/triage-security.yaml + sha256: 110fb33d880caf1add20e73c8017451378966eeef52ef34aa988e41c35fc1b02 + type: file + fetched_at: 2026-09-30T11:35:04.92833Z + - field: pre_script + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/plugins/sdlc-workflow/scripts/pre-triage-security.sh + sha256: fa779388a5bece507d3d670704e7f56c9099e646ecea570b471fdc6c8b9181a6 + type: directory + files: + - path: execute-actions.py + sha256: 595220e3cac4426acaecfda17e9c7b73532cd9bf59a453c4093dc8e9a57a0936 + - path: execute-triage-security-actions.py + sha256: 43263de469c4e997d20c64e89b99e49ec0cb0a386d3af72d42584cbb62bb827a + - path: jira-client.py + sha256: d2ed3d8188f3f0ceda5f8ed13deadaba19e0f4a3ec9a4041d7f9bef4c101e5e7 + - path: post-triage-security.sh + sha256: 2e5527eddd13f3beb9a87afec9e9e443e6f72bccde32a3620144584af1b9b140 + - path: post-verify-pr.sh + sha256: 0cdd85d68e1c94546964747badcb1837602aae0e47a168903f39e3760d1be687 + - path: pre-triage-security.sh + sha256: 3b82b146ec0c91ec93e93c7206d357121a256a8496be75b92a98ad6fe034921a + - path: pre-verify-pr.sh + sha256: 89e37062e78de2f2064863b0395ccc46c89489d1e4ee510b1b20386c20fe30ee + - path: pre_triage_security.py + sha256: 46dbf5fd2061c774774af35ba6da6bd652bba0c9f3144da20573f1ca25aab0c6 + - path: pre_verify_pr.py + sha256: c2bdade236aa8a2fe374a08a2aab0f78d287125a8ae38b0d3a938b98252fae85 + - path: strip_extra_properties.py + sha256: 9099200014a0a4a9da8795a3346f963efc34c53f4603d340551c46b8d135d122 + - path: test_execute_actions.py + sha256: 4170eea78768a516097d212a48b4b459ffb47370c0cc224bc157113d581ab0d4 + - path: test_execute_triage_security_actions.py + sha256: eb296eea89b4cbf588b3d0c7a62e520ae6b20fb417fbcada20529d9af876f3f4 + - path: test_jira_client.py + sha256: 1b9a6cc3aa36942e373df3ed847c550b2e396bc6df6cff5f03361253f7bbc8f1 + - path: test_jira_client_cli.py + sha256: d36b790c1a310111a0d2f3e81415f9576e80acc5f0bc7d0ff8084ffad001285f + - path: test_pre_triage_security.py + sha256: 31ed6b80c80a934d6f15d9959e1df0d089dea7b8aa6f0ecdf58e36a5e575a860 + - path: test_pre_verify_pr.py + sha256: df06c53591349693228e51e56f8636a4482bd0b748096671a0198301be1c51bd + - path: test_triage_security_fullsend.py + sha256: 6c5e923ec652d44459b7db12cccb4b72eaea511aacbe6ea04c931a26ed5b1333 + - path: validate-output-schema.sh + sha256: 56b964145da0b62f438dbc1885780d86e2d589c5340adb93c00e3c8777979cbe + fetched_at: 2026-09-30T11:35:08.106377Z + - field: post_script + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/plugins/sdlc-workflow/scripts/post-triage-security.sh + sha256: fa779388a5bece507d3d670704e7f56c9099e646ecea570b471fdc6c8b9181a6 + type: directory + files: + - path: execute-actions.py + sha256: 595220e3cac4426acaecfda17e9c7b73532cd9bf59a453c4093dc8e9a57a0936 + - path: execute-triage-security-actions.py + sha256: 43263de469c4e997d20c64e89b99e49ec0cb0a386d3af72d42584cbb62bb827a + - path: jira-client.py + sha256: d2ed3d8188f3f0ceda5f8ed13deadaba19e0f4a3ec9a4041d7f9bef4c101e5e7 + - path: post-triage-security.sh + sha256: 2e5527eddd13f3beb9a87afec9e9e443e6f72bccde32a3620144584af1b9b140 + - path: post-verify-pr.sh + sha256: 0cdd85d68e1c94546964747badcb1837602aae0e47a168903f39e3760d1be687 + - path: pre-triage-security.sh + sha256: 3b82b146ec0c91ec93e93c7206d357121a256a8496be75b92a98ad6fe034921a + - path: pre-verify-pr.sh + sha256: 89e37062e78de2f2064863b0395ccc46c89489d1e4ee510b1b20386c20fe30ee + - path: pre_triage_security.py + sha256: 46dbf5fd2061c774774af35ba6da6bd652bba0c9f3144da20573f1ca25aab0c6 + - path: pre_verify_pr.py + sha256: c2bdade236aa8a2fe374a08a2aab0f78d287125a8ae38b0d3a938b98252fae85 + - path: strip_extra_properties.py + sha256: 9099200014a0a4a9da8795a3346f963efc34c53f4603d340551c46b8d135d122 + - path: test_execute_actions.py + sha256: 4170eea78768a516097d212a48b4b459ffb47370c0cc224bc157113d581ab0d4 + - path: test_execute_triage_security_actions.py + sha256: eb296eea89b4cbf588b3d0c7a62e520ae6b20fb417fbcada20529d9af876f3f4 + - path: test_jira_client.py + sha256: 1b9a6cc3aa36942e373df3ed847c550b2e396bc6df6cff5f03361253f7bbc8f1 + - path: test_jira_client_cli.py + sha256: d36b790c1a310111a0d2f3e81415f9576e80acc5f0bc7d0ff8084ffad001285f + - path: test_pre_triage_security.py + sha256: 31ed6b80c80a934d6f15d9959e1df0d089dea7b8aa6f0ecdf58e36a5e575a860 + - path: test_pre_verify_pr.py + sha256: df06c53591349693228e51e56f8636a4482bd0b748096671a0198301be1c51bd + - path: test_triage_security_fullsend.py + sha256: 6c5e923ec652d44459b7db12cccb4b72eaea511aacbe6ea04c931a26ed5b1333 + - path: validate-output-schema.sh + sha256: 56b964145da0b62f438dbc1885780d86e2d589c5340adb93c00e3c8777979cbe + fetched_at: 2026-09-30T11:35:08.09954Z + - field: validation_loop.script + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/plugins/sdlc-workflow/scripts/validate-output-schema.sh + sha256: fa779388a5bece507d3d670704e7f56c9099e646ecea570b471fdc6c8b9181a6 + type: directory + files: + - path: execute-actions.py + sha256: 595220e3cac4426acaecfda17e9c7b73532cd9bf59a453c4093dc8e9a57a0936 + - path: execute-triage-security-actions.py + sha256: 43263de469c4e997d20c64e89b99e49ec0cb0a386d3af72d42584cbb62bb827a + - path: jira-client.py + sha256: d2ed3d8188f3f0ceda5f8ed13deadaba19e0f4a3ec9a4041d7f9bef4c101e5e7 + - path: post-triage-security.sh + sha256: 2e5527eddd13f3beb9a87afec9e9e443e6f72bccde32a3620144584af1b9b140 + - path: post-verify-pr.sh + sha256: 0cdd85d68e1c94546964747badcb1837602aae0e47a168903f39e3760d1be687 + - path: pre-triage-security.sh + sha256: 3b82b146ec0c91ec93e93c7206d357121a256a8496be75b92a98ad6fe034921a + - path: pre-verify-pr.sh + sha256: 89e37062e78de2f2064863b0395ccc46c89489d1e4ee510b1b20386c20fe30ee + - path: pre_triage_security.py + sha256: 46dbf5fd2061c774774af35ba6da6bd652bba0c9f3144da20573f1ca25aab0c6 + - path: pre_verify_pr.py + sha256: c2bdade236aa8a2fe374a08a2aab0f78d287125a8ae38b0d3a938b98252fae85 + - path: strip_extra_properties.py + sha256: 9099200014a0a4a9da8795a3346f963efc34c53f4603d340551c46b8d135d122 + - path: test_execute_actions.py + sha256: 4170eea78768a516097d212a48b4b459ffb47370c0cc224bc157113d581ab0d4 + - path: test_execute_triage_security_actions.py + sha256: eb296eea89b4cbf588b3d0c7a62e520ae6b20fb417fbcada20529d9af876f3f4 + - path: test_jira_client.py + sha256: 1b9a6cc3aa36942e373df3ed847c550b2e396bc6df6cff5f03361253f7bbc8f1 + - path: test_jira_client_cli.py + sha256: d36b790c1a310111a0d2f3e81415f9576e80acc5f0bc7d0ff8084ffad001285f + - path: test_pre_triage_security.py + sha256: 31ed6b80c80a934d6f15d9959e1df0d089dea7b8aa6f0ecdf58e36a5e575a860 + - path: test_pre_verify_pr.py + sha256: df06c53591349693228e51e56f8636a4482bd0b748096671a0198301be1c51bd + - path: test_triage_security_fullsend.py + sha256: 6c5e923ec652d44459b7db12cccb4b72eaea511aacbe6ea04c931a26ed5b1333 + - path: validate-output-schema.sh + sha256: 56b964145da0b62f438dbc1885780d86e2d589c5340adb93c00e3c8777979cbe + fetched_at: 2026-09-30T11:35:08.09954Z + - field: validation_loop.schema + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/plugins/sdlc-workflow/schemas/triage-security-result.schema.json + sha256: f1100b5624433fae2163bad82c0ccd641212d3a84e7b0562448919d015b5d1d2 + type: resource + fetched_at: 2026-09-30T11:35:08.380721Z + - field: agent + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/plugins/sdlc-workflow/agents/triage-security.md + sha256: 3e0680274ff99326fa219b45ccbea6ea1ed1a7eba3f78056acf9e47b29c7c9fd + type: resource + fetched_at: 2026-09-30T11:35:08.639209Z + - field: policy + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/plugins/sdlc-workflow/policies/triage-security.yaml + sha256: 312f947a7c0eceab078111406c8288853c760eb0304b080262da322066046b91 + type: resource + fetched_at: 2026-09-30T11:35:08.910263Z + - field: host_files[0].src + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/plugins/sdlc-workflow/env/gcp-vertex.env + sha256: 10b2ba695b1d4e65e0964a233b75d6a06853f1eaefc260c56b9eedd989ae7d41 + type: resource + fetched_at: 2026-09-30T11:35:09.211078Z + - field: openshell.profiles[0] + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/plugins/sdlc-workflow/profiles/fullsend-vertex-ai.yaml + sha256: 76535a148387b1281be2bfb23b6724e4133327189b38c8ce1de2c6dddb8d8341 + type: resource + fetched_at: 2026-09-30T11:35:09.473046Z + - field: providers[0] + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/plugins/sdlc-workflow/providers/vertex-ai.yaml + sha256: ae5ebe527e5d7b0fa6994346fde3f6ba11b632a3ba78ece96591b5ed85083ec1 + type: resource + fetched_at: 2026-09-30T11:35:09.762266Z + - field: plugins[0] + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/59f9ff4b4f49646c64ed3412271ce7c017e781d4/plugins/sdlc-workflow/ + sha256: 802a1316374582a6b8ffd6720f30b76c45b78f00f86269aa7758e6f548e0e27c + type: directory + files: + - path: .claude-plugin/plugin.json + sha256: a10c3191ab7957bd496b0ebbbb212844c48d3a95c1556097d2466d5770cf1c20 + - path: agents/triage-security.md + sha256: 3e0680274ff99326fa219b45ccbea6ea1ed1a7eba3f78056acf9e47b29c7c9fd + - path: agents/verify-pr.md + sha256: 229fef270b4b195f91cc46ef3ba50cbf7701f49ee546291882cff893d53844b3 + - path: env/gcp-vertex.env + sha256: 10b2ba695b1d4e65e0964a233b75d6a06853f1eaefc260c56b9eedd989ae7d41 + - path: plugin.json + sha256: 727186c0e7e8d0d6b241ff61cb57a979a2d9c7c0c684a4d1c12940534e9f2b18 + - path: policies/triage-security.yaml + sha256: 312f947a7c0eceab078111406c8288853c760eb0304b080262da322066046b91 + - path: policies/verify-pr.yaml + sha256: 62d1374be253b0b225921bd8e22e679781076ecdb3a9ea22d5e6a4805753c5a7 + - path: profiles/fullsend-vertex-ai.yaml + sha256: 76535a148387b1281be2bfb23b6724e4133327189b38c8ce1de2c6dddb8d8341 + - path: providers/vertex-ai.yaml + sha256: ae5ebe527e5d7b0fa6994346fde3f6ba11b632a3ba78ece96591b5ed85083ec1 + - path: schemas/triage-security-input.schema.json + sha256: c72a624e821eed0c12ac7dd8b60140f97c58160b449e55974afa644ee5ecd3af + - path: schemas/triage-security-result.schema.json + sha256: f1100b5624433fae2163bad82c0ccd641212d3a84e7b0562448919d015b5d1d2 + - path: schemas/verify-pr-input.schema.json + sha256: 3403de2f42c88bdc4213a10ac7e5241888e08642c495d78456d0cd44611a9e6c + - path: schemas/verify-pr-result.schema.json + sha256: 62c3e5ce0c47e73823b5cab9a577be7b721d8fb31389bc8602dc82ee05526ad6 + - path: scripts/execute-actions.py + sha256: 595220e3cac4426acaecfda17e9c7b73532cd9bf59a453c4093dc8e9a57a0936 + - path: scripts/execute-triage-security-actions.py + sha256: 43263de469c4e997d20c64e89b99e49ec0cb0a386d3af72d42584cbb62bb827a + - path: scripts/jira-client.py + sha256: d2ed3d8188f3f0ceda5f8ed13deadaba19e0f4a3ec9a4041d7f9bef4c101e5e7 + - path: scripts/post-triage-security.sh + sha256: 2e5527eddd13f3beb9a87afec9e9e443e6f72bccde32a3620144584af1b9b140 + - path: scripts/post-verify-pr.sh + sha256: 0cdd85d68e1c94546964747badcb1837602aae0e47a168903f39e3760d1be687 + - path: scripts/pre-triage-security.sh + sha256: 3b82b146ec0c91ec93e93c7206d357121a256a8496be75b92a98ad6fe034921a + - path: scripts/pre-verify-pr.sh + sha256: 89e37062e78de2f2064863b0395ccc46c89489d1e4ee510b1b20386c20fe30ee + - path: scripts/pre_triage_security.py + sha256: 46dbf5fd2061c774774af35ba6da6bd652bba0c9f3144da20573f1ca25aab0c6 + - path: scripts/pre_verify_pr.py + sha256: c2bdade236aa8a2fe374a08a2aab0f78d287125a8ae38b0d3a938b98252fae85 + - path: scripts/strip_extra_properties.py + sha256: 9099200014a0a4a9da8795a3346f963efc34c53f4603d340551c46b8d135d122 + - path: scripts/test_execute_actions.py + sha256: 4170eea78768a516097d212a48b4b459ffb47370c0cc224bc157113d581ab0d4 + - path: scripts/test_execute_triage_security_actions.py + sha256: eb296eea89b4cbf588b3d0c7a62e520ae6b20fb417fbcada20529d9af876f3f4 + - path: scripts/test_jira_client.py + sha256: 1b9a6cc3aa36942e373df3ed847c550b2e396bc6df6cff5f03361253f7bbc8f1 + - path: scripts/test_jira_client_cli.py + sha256: d36b790c1a310111a0d2f3e81415f9576e80acc5f0bc7d0ff8084ffad001285f + - path: scripts/test_pre_triage_security.py + sha256: 31ed6b80c80a934d6f15d9959e1df0d089dea7b8aa6f0ecdf58e36a5e575a860 + - path: scripts/test_pre_verify_pr.py + sha256: df06c53591349693228e51e56f8636a4482bd0b748096671a0198301be1c51bd + - path: scripts/test_triage_security_fullsend.py + sha256: 6c5e923ec652d44459b7db12cccb4b72eaea511aacbe6ea04c931a26ed5b1333 + - path: scripts/validate-output-schema.sh + sha256: 56b964145da0b62f438dbc1885780d86e2d589c5340adb93c00e3c8777979cbe + - path: shared/comment-footnote.md + sha256: bd78a3a34db53aaec53149c2d5557d1c55d10c6f4cd1622d906652cf9b72ca5c + - path: shared/convention-applicability-rules.md + sha256: 1ccabd6491148fdfc9911fcd2f233c9092d4dd54bc82cd67f52512e3e35d1823 + - path: shared/description-digest-protocol.md + sha256: 06cf26d8677d2e9866f7c72d8fb8037cca13df8a7ffcdbf3cd333775283fb7ea + - path: shared/eval-coverage-propagation.md + sha256: d04efdb625e6b1382b2a4b8682bf41accfd99f286c40856238e0735c1f4a00f0 + - path: shared/jira-access-strategy.md + sha256: c00aebec96792156cdb2a80a5eb1431cbf9b1d35ba7b54d5fbca37227a2b6303 + - path: shared/jira-api-token-guide.md + sha256: f8fefd4e12ffca9127e692de99b18aa8e3b007cb5f5e9fb2c7b16336fd84518a + - path: shared/jira-rest-fallback.md + sha256: fe39a0eab5febfb77c7c5112c0ec1c3f1ab5ec581d9f3c40b4cd968228d23cb5 + - path: shared/task-description-template.md + sha256: 8f0d8fbddb8b8f662db3c1c5c47e9224aa5c0d504fd615e935a960ead6660a7b + - path: skills/define-feature/SKILL.md + sha256: c314f3f549c3ef835e03c38ed8ce2212422c77a876d9b786dd4e08806f5d5e4e + - path: skills/implement-task/SKILL.md + sha256: bd4ea474c2504cbdc0500bdfb33949af53584b05dc2696bcf9f12dc5d898a16a + - path: skills/plan-feature/SKILL.md + sha256: a8603de28ac9ff7d163b3ef5e11c44ab3f190d735e7ec6fd57604b6b9766a9ce + - path: skills/report-bug/SKILL.md + sha256: cc841f523aa9b2620f01426c32a84ee29710c732157356743a87e8ff60823cc4 + - path: skills/run-evals/SKILL.md + sha256: 221c45e692d4becfd839bc99ce6d2c4a3fabc9a68dd5479c40de28b2870ec5c9 + - path: skills/run-evals/scripts/aggregate_benchmark.py + sha256: b89b6471d6abad6684863afcdc8634c2097986afb3ff64f80f179cd71c3dd325 + - path: skills/run-evals/scripts/render_summary.py + sha256: 32da827a4db496a20d80148583b1f4906f595bcc5b77393a5e39dfdab30db1a5 + - path: skills/run-evals/scripts/test_render_summary.py + sha256: 34a2333e5bee8b0f07c23e64a9e7ab096cf16e4a027cacf91b86c4b39e584cf7 + - path: skills/setup/SKILL.md + sha256: a9d4a55fc35123607c445ea4ad2609e8306443f2032d360fa55b54827871bde6 + - path: skills/setup/constraints.template.md + sha256: 6a0c75f5e390eea5c42daffd8517b19ac58d41d6600d46a067bc31d3056b52b0 + - path: skills/setup/conventions.template.md + sha256: 264e53e4b2c2cad10d69293e009ca3fd7be3437de0422c81f6a9e4430f018361 + - path: skills/setup/project-config.template.md + sha256: d2b607d8d226bbe303b00c9e6e875b115f8f138e8e4afa29c38f37fce7b03dbd + - path: skills/setup/security-config.template.md + sha256: 2422a28db45539c16cd6bfc45ab9d390dc41c816784065ad84c82cc8b5aea120 + - path: skills/triage-bug/SKILL.md + sha256: e3c5f28d1c7bafdcf706d8be89110cef28217ef878a754fb89c571c5731242d4 + - path: skills/triage-security/SKILL.md + sha256: 23e198c71b8b33af584cfb679d149d416c50f3444b540677d4f0fea8e3cf3102 + - path: skills/triage-security/jira-triage-operations.md + sha256: 2421a2e0fb2aa748b5b3b34a523b75f015f6c862fcd82854d1efd3886595aa80 + - path: skills/triage-security/remediation-templates.md + sha256: 05572e97310f23bc9834a4be92fb50af86f2dec16fdb613cb02ca8e936937eb9 + - path: skills/triage-security/version-impact-analysis.md + sha256: 79c9bec47f07f76da2479997e0e345f0bc961485a5c4b962563cbbacdc16faa5 + - path: skills/verify-pr/SKILL.md + sha256: 9fb1e3bbe01d8c57cff6489415b06ecfd14fa63795cabd122580f0bb7408790d + - path: skills/verify-pr/correctness.md + sha256: 09579d68e36cf6256915ff1e839804a2a1394457ec820063369dd5c260491426 + - path: skills/verify-pr/dispatch-template.md + sha256: 8246497d73dce88f18f19830a10356f04abbd9862c8e06db3b6e326a658a4617 + - path: skills/verify-pr/finding-template.md + sha256: b3a7c52d60d1ffc30f5d80219b2fc9db6c28eabeca83bcc0ffa1e424b3a4d0c2 + - path: skills/verify-pr/intent-alignment.md + sha256: 8a1ffc88f4072df77fc50a12bb6443be6e6884ac5c51b4501dac0e0e18c37f01 + - path: skills/verify-pr/security.md + sha256: 128877cf7e156666dbb403805a0112059cc8cc3255e148c6794da29d48b8f5cf + - path: skills/verify-pr/style-conventions.md + sha256: 5ab920c596615c0d23354cb0e90fe5ebf5f08635ffa3602e8e2d2abd74028a19 + fetched_at: 2026-09-30T11:35:13.009869Z + verify-pr: + source: harness/verify-pr.yaml + sha256: cefd31d27a3bb7f61c790725821a808643846afb8bf8f3b1167ea8220a4de884 + resolved_at: 2026-09-17T16:59:14.885757Z + dependencies: + - field: base + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/harness/verify-pr.yaml + sha256: 3c9dc221e2a37e7aa7ff294381b6b5b1580771ba70432c42924fc3a85f078e63 + type: file + fetched_at: 2026-09-16T08:49:11.789227Z + - field: pre_script + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/plugins/sdlc-workflow/scripts/pre-verify-pr.sh + sha256: 656f5b7024f1fc9b9b9b43f0c02e6aacbecc8af281405c6ca86f5b3dc74c972e + type: directory + files: + - path: execute-actions.py + sha256: 595220e3cac4426acaecfda17e9c7b73532cd9bf59a453c4093dc8e9a57a0936 + - path: jira-client.py + sha256: 67eb537abe8223d53b2905085d5e98c9013a3bb847df2ccfc18ecd92c8204c27 + - path: post-verify-pr.sh + sha256: 0cdd85d68e1c94546964747badcb1837602aae0e47a168903f39e3760d1be687 + - path: pre-verify-pr.sh + sha256: 89e37062e78de2f2064863b0395ccc46c89489d1e4ee510b1b20386c20fe30ee + - path: pre_verify_pr.py + sha256: c2bdade236aa8a2fe374a08a2aab0f78d287125a8ae38b0d3a938b98252fae85 + - path: strip_extra_properties.py + sha256: 9099200014a0a4a9da8795a3346f963efc34c53f4603d340551c46b8d135d122 + - path: test_execute_actions.py + sha256: 4170eea78768a516097d212a48b4b459ffb47370c0cc224bc157113d581ab0d4 + - path: test_jira_client.py + sha256: 1b9a6cc3aa36942e373df3ed847c550b2e396bc6df6cff5f03361253f7bbc8f1 + - path: test_jira_client_cli.py + sha256: d36b790c1a310111a0d2f3e81415f9576e80acc5f0bc7d0ff8084ffad001285f + - path: test_pre_verify_pr.py + sha256: df06c53591349693228e51e56f8636a4482bd0b748096671a0198301be1c51bd + - path: validate-output-schema.sh + sha256: 56b964145da0b62f438dbc1885780d86e2d589c5340adb93c00e3c8777979cbe + fetched_at: 2026-09-17T16:59:04.386171Z + - field: post_script + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/plugins/sdlc-workflow/scripts/post-verify-pr.sh + sha256: 656f5b7024f1fc9b9b9b43f0c02e6aacbecc8af281405c6ca86f5b3dc74c972e + type: directory + files: + - path: execute-actions.py + sha256: 595220e3cac4426acaecfda17e9c7b73532cd9bf59a453c4093dc8e9a57a0936 + - path: jira-client.py + sha256: 67eb537abe8223d53b2905085d5e98c9013a3bb847df2ccfc18ecd92c8204c27 + - path: post-verify-pr.sh + sha256: 0cdd85d68e1c94546964747badcb1837602aae0e47a168903f39e3760d1be687 + - path: pre-verify-pr.sh + sha256: 89e37062e78de2f2064863b0395ccc46c89489d1e4ee510b1b20386c20fe30ee + - path: pre_verify_pr.py + sha256: c2bdade236aa8a2fe374a08a2aab0f78d287125a8ae38b0d3a938b98252fae85 + - path: strip_extra_properties.py + sha256: 9099200014a0a4a9da8795a3346f963efc34c53f4603d340551c46b8d135d122 + - path: test_execute_actions.py + sha256: 4170eea78768a516097d212a48b4b459ffb47370c0cc224bc157113d581ab0d4 + - path: test_jira_client.py + sha256: 1b9a6cc3aa36942e373df3ed847c550b2e396bc6df6cff5f03361253f7bbc8f1 + - path: test_jira_client_cli.py + sha256: d36b790c1a310111a0d2f3e81415f9576e80acc5f0bc7d0ff8084ffad001285f + - path: test_pre_verify_pr.py + sha256: df06c53591349693228e51e56f8636a4482bd0b748096671a0198301be1c51bd + - path: validate-output-schema.sh + sha256: 56b964145da0b62f438dbc1885780d86e2d589c5340adb93c00e3c8777979cbe + fetched_at: 2026-09-17T16:59:04.386171Z + - field: validation_loop.script + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/plugins/sdlc-workflow/scripts/validate-output-schema.sh + sha256: 656f5b7024f1fc9b9b9b43f0c02e6aacbecc8af281405c6ca86f5b3dc74c972e + type: directory + files: + - path: execute-actions.py + sha256: 595220e3cac4426acaecfda17e9c7b73532cd9bf59a453c4093dc8e9a57a0936 + - path: jira-client.py + sha256: 67eb537abe8223d53b2905085d5e98c9013a3bb847df2ccfc18ecd92c8204c27 + - path: post-verify-pr.sh + sha256: 0cdd85d68e1c94546964747badcb1837602aae0e47a168903f39e3760d1be687 + - path: pre-verify-pr.sh + sha256: 89e37062e78de2f2064863b0395ccc46c89489d1e4ee510b1b20386c20fe30ee + - path: pre_verify_pr.py + sha256: c2bdade236aa8a2fe374a08a2aab0f78d287125a8ae38b0d3a938b98252fae85 + - path: strip_extra_properties.py + sha256: 9099200014a0a4a9da8795a3346f963efc34c53f4603d340551c46b8d135d122 + - path: test_execute_actions.py + sha256: 4170eea78768a516097d212a48b4b459ffb47370c0cc224bc157113d581ab0d4 + - path: test_jira_client.py + sha256: 1b9a6cc3aa36942e373df3ed847c550b2e396bc6df6cff5f03361253f7bbc8f1 + - path: test_jira_client_cli.py + sha256: d36b790c1a310111a0d2f3e81415f9576e80acc5f0bc7d0ff8084ffad001285f + - path: test_pre_verify_pr.py + sha256: df06c53591349693228e51e56f8636a4482bd0b748096671a0198301be1c51bd + - path: validate-output-schema.sh + sha256: 56b964145da0b62f438dbc1885780d86e2d589c5340adb93c00e3c8777979cbe + fetched_at: 2026-09-17T16:59:04.386171Z + - field: validation_loop.schema + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/plugins/sdlc-workflow/schemas/verify-pr-result.schema.json + sha256: 62c3e5ce0c47e73823b5cab9a577be7b721d8fb31389bc8602dc82ee05526ad6 + type: resource + fetched_at: 2026-09-17T16:59:10.434287Z + - field: agent + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/plugins/sdlc-workflow/agents/verify-pr.md + sha256: 229fef270b4b195f91cc46ef3ba50cbf7701f49ee546291882cff893d53844b3 + type: resource + fetched_at: 2026-09-17T16:59:10.709192Z + - field: policy + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/plugins/sdlc-workflow/policies/verify-pr.yaml + sha256: 62d1374be253b0b225921bd8e22e679781076ecdb3a9ea22d5e6a4805753c5a7 + type: resource + fetched_at: 2026-09-17T16:59:10.978282Z + - field: host_files[0].src + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/plugins/sdlc-workflow/env/gcp-vertex.env + sha256: 10b2ba695b1d4e65e0964a233b75d6a06853f1eaefc260c56b9eedd989ae7d41 + type: resource + fetched_at: 2026-09-17T16:59:11.282366Z + - field: openshell.profiles[0] + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/plugins/sdlc-workflow/profiles/fullsend-vertex-ai.yaml + sha256: 76535a148387b1281be2bfb23b6724e4133327189b38c8ce1de2c6dddb8d8341 + type: resource + fetched_at: 2026-09-17T16:59:11.553974Z + - field: providers[0] + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/plugins/sdlc-workflow/providers/vertex-ai.yaml + sha256: ae5ebe527e5d7b0fa6994346fde3f6ba11b632a3ba78ece96591b5ed85083ec1 + type: resource + fetched_at: 2026-09-17T16:59:11.834269Z + - field: plugins[0] + url: https://raw.githubusercontent.com/RHEcosystemAppEng/sdlc-plugins/da053ba1649ef9368436d69a639e121138b97801/plugins/sdlc-workflow/plugin.json + sha256: a410b7e9f53dbe44d0d14b3f96e37b163499eae29d290aeddb520a34c8c35df0 + type: directory + files: + - path: .claude-plugin/plugin.json + sha256: a10c3191ab7957bd496b0ebbbb212844c48d3a95c1556097d2466d5770cf1c20 + - path: agents/verify-pr.md + sha256: 229fef270b4b195f91cc46ef3ba50cbf7701f49ee546291882cff893d53844b3 + - path: env/gcp-vertex.env + sha256: 10b2ba695b1d4e65e0964a233b75d6a06853f1eaefc260c56b9eedd989ae7d41 + - path: plugin.json + sha256: 727186c0e7e8d0d6b241ff61cb57a979a2d9c7c0c684a4d1c12940534e9f2b18 + - path: policies/verify-pr.yaml + sha256: 62d1374be253b0b225921bd8e22e679781076ecdb3a9ea22d5e6a4805753c5a7 + - path: profiles/fullsend-vertex-ai.yaml + sha256: 76535a148387b1281be2bfb23b6724e4133327189b38c8ce1de2c6dddb8d8341 + - path: providers/vertex-ai.yaml + sha256: ae5ebe527e5d7b0fa6994346fde3f6ba11b632a3ba78ece96591b5ed85083ec1 + - path: schemas/verify-pr-input.schema.json + sha256: 3403de2f42c88bdc4213a10ac7e5241888e08642c495d78456d0cd44611a9e6c + - path: schemas/verify-pr-result.schema.json + sha256: 62c3e5ce0c47e73823b5cab9a577be7b721d8fb31389bc8602dc82ee05526ad6 + - path: scripts/execute-actions.py + sha256: 595220e3cac4426acaecfda17e9c7b73532cd9bf59a453c4093dc8e9a57a0936 + - path: scripts/jira-client.py + sha256: 67eb537abe8223d53b2905085d5e98c9013a3bb847df2ccfc18ecd92c8204c27 + - path: scripts/post-verify-pr.sh + sha256: 0cdd85d68e1c94546964747badcb1837602aae0e47a168903f39e3760d1be687 + - path: scripts/pre-verify-pr.sh + sha256: 89e37062e78de2f2064863b0395ccc46c89489d1e4ee510b1b20386c20fe30ee + - path: scripts/pre_verify_pr.py + sha256: c2bdade236aa8a2fe374a08a2aab0f78d287125a8ae38b0d3a938b98252fae85 + - path: scripts/strip_extra_properties.py + sha256: 9099200014a0a4a9da8795a3346f963efc34c53f4603d340551c46b8d135d122 + - path: scripts/test_execute_actions.py + sha256: 4170eea78768a516097d212a48b4b459ffb47370c0cc224bc157113d581ab0d4 + - path: scripts/test_jira_client.py + sha256: 1b9a6cc3aa36942e373df3ed847c550b2e396bc6df6cff5f03361253f7bbc8f1 + - path: scripts/test_jira_client_cli.py + sha256: d36b790c1a310111a0d2f3e81415f9576e80acc5f0bc7d0ff8084ffad001285f + - path: scripts/test_pre_verify_pr.py + sha256: df06c53591349693228e51e56f8636a4482bd0b748096671a0198301be1c51bd + - path: scripts/validate-output-schema.sh + sha256: 56b964145da0b62f438dbc1885780d86e2d589c5340adb93c00e3c8777979cbe + - path: shared/comment-footnote.md + sha256: bd78a3a34db53aaec53149c2d5557d1c55d10c6f4cd1622d906652cf9b72ca5c + - path: shared/convention-applicability-rules.md + sha256: 1ccabd6491148fdfc9911fcd2f233c9092d4dd54bc82cd67f52512e3e35d1823 + - path: shared/description-digest-protocol.md + sha256: 06cf26d8677d2e9866f7c72d8fb8037cca13df8a7ffcdbf3cd333775283fb7ea + - path: shared/eval-coverage-propagation.md + sha256: d04efdb625e6b1382b2a4b8682bf41accfd99f286c40856238e0735c1f4a00f0 + - path: shared/jira-access-strategy.md + sha256: c00aebec96792156cdb2a80a5eb1431cbf9b1d35ba7b54d5fbca37227a2b6303 + - path: shared/jira-api-token-guide.md + sha256: f8fefd4e12ffca9127e692de99b18aa8e3b007cb5f5e9fb2c7b16336fd84518a + - path: shared/jira-rest-fallback.md + sha256: fe39a0eab5febfb77c7c5112c0ec1c3f1ab5ec581d9f3c40b4cd968228d23cb5 + - path: shared/task-description-template.md + sha256: 8f0d8fbddb8b8f662db3c1c5c47e9224aa5c0d504fd615e935a960ead6660a7b + - path: skills/define-feature/SKILL.md + sha256: c314f3f549c3ef835e03c38ed8ce2212422c77a876d9b786dd4e08806f5d5e4e + - path: skills/implement-task/SKILL.md + sha256: 0276a52292d68908e56ab213746c93a931bede189d3b6e7a914f6a2024f6f4c1 + - path: skills/plan-feature/SKILL.md + sha256: a8603de28ac9ff7d163b3ef5e11c44ab3f190d735e7ec6fd57604b6b9766a9ce + - path: skills/report-bug/SKILL.md + sha256: cc841f523aa9b2620f01426c32a84ee29710c732157356743a87e8ff60823cc4 + - path: skills/run-evals/SKILL.md + sha256: 221c45e692d4becfd839bc99ce6d2c4a3fabc9a68dd5479c40de28b2870ec5c9 + - path: skills/run-evals/scripts/aggregate_benchmark.py + sha256: b89b6471d6abad6684863afcdc8634c2097986afb3ff64f80f179cd71c3dd325 + - path: skills/run-evals/scripts/render_summary.py + sha256: 32da827a4db496a20d80148583b1f4906f595bcc5b77393a5e39dfdab30db1a5 + - path: skills/run-evals/scripts/test_render_summary.py + sha256: 34a2333e5bee8b0f07c23e64a9e7ab096cf16e4a027cacf91b86c4b39e584cf7 + - path: skills/setup/SKILL.md + sha256: a9d4a55fc35123607c445ea4ad2609e8306443f2032d360fa55b54827871bde6 + - path: skills/setup/constraints.template.md + sha256: 6a0c75f5e390eea5c42daffd8517b19ac58d41d6600d46a067bc31d3056b52b0 + - path: skills/setup/conventions.template.md + sha256: 264e53e4b2c2cad10d69293e009ca3fd7be3437de0422c81f6a9e4430f018361 + - path: skills/setup/project-config.template.md + sha256: d2b607d8d226bbe303b00c9e6e875b115f8f138e8e4afa29c38f37fce7b03dbd + - path: skills/setup/security-config.template.md + sha256: 2422a28db45539c16cd6bfc45ab9d390dc41c816784065ad84c82cc8b5aea120 + - path: skills/triage-bug/SKILL.md + sha256: e3c5f28d1c7bafdcf706d8be89110cef28217ef878a754fb89c571c5731242d4 + - path: skills/triage-security/SKILL.md + sha256: 040ba49ffcedb2c0301389ff4ca1572f45c9b9c0d33f7a8f463b0b1ad472d459 + - path: skills/triage-security/jira-triage-operations.md + sha256: 665aef8f2529ee44be4fdee57783e6d589867f71793e9b7d280ce8f3973a9dcb + - path: skills/triage-security/remediation-templates.md + sha256: 0b45136eb63e18a0bb81eebcb1dd033a27a69134117949eedf9f8e43d0d76010 + - path: skills/triage-security/version-impact-analysis.md + sha256: a18376094360d435e96436c9fd1bd6a04831123d903e44cf6cf57e600285e3c9 + - path: skills/verify-pr/SKILL.md + sha256: 9fb1e3bbe01d8c57cff6489415b06ecfd14fa63795cabd122580f0bb7408790d + - path: skills/verify-pr/correctness.md + sha256: 09579d68e36cf6256915ff1e839804a2a1394457ec820063369dd5c260491426 + - path: skills/verify-pr/dispatch-template.md + sha256: 8246497d73dce88f18f19830a10356f04abbd9862c8e06db3b6e326a658a4617 + - path: skills/verify-pr/finding-template.md + sha256: b3a7c52d60d1ffc30f5d80219b2fc9db6c28eabeca83bcc0ffa1e424b3a4d0c2 + - path: skills/verify-pr/intent-alignment.md + sha256: 8a1ffc88f4072df77fc50a12bb6443be6e6884ac5c51b4501dac0e0e18c37f01 + - path: skills/verify-pr/security.md + sha256: 128877cf7e156666dbb403805a0112059cc8cc3255e148c6794da29d48b8f5cf + - path: skills/verify-pr/style-conventions.md + sha256: 5ab920c596615c0d23354cb0e90fe5ebf5f08635ffa3602e8e2d2abd74028a19 + fetched_at: 2026-09-17T16:59:14.868165Z diff --git a/.github/workflows/eval-pr-run.yml b/.github/workflows/eval-pr-run.yml index ca3110c84..82f2f2dd2 100644 --- a/.github/workflows/eval-pr-run.yml +++ b/.github/workflows/eval-pr-run.yml @@ -37,8 +37,8 @@ permissions: statuses: write env: - # Reviewed PR299 suite; normal activation switches this to trusted github.sha. - NATIVE_EVAL_SOURCE_SHA: 91698d4dca24bc199e763dc6462dc54523e473c0 + # After integration merge, suite execution always comes from trusted main. + NATIVE_EVAL_SOURCE_SHA: ${{ github.sha }} jobs: discover: @@ -229,9 +229,6 @@ jobs: } const skills = confirmed.join(','); - // TC-6726 bootstrap restriction: remove only this identity condition - // during the separately reviewed PR299 activation delivery. - const bootstrap = prNumber === 299 && process.env.SOURCE_BRANCH === 'verify-pr-fullsend'; const nativeRelevant = files.some(f => /^evals\/fullsend\//.test(f.filename) || /^evals\/triage-security\//.test(f.filename) || @@ -239,7 +236,7 @@ jobs: /^plugins\/sdlc-workflow\/(policies\/triage-security\.yaml|providers\/vertex-ai\.yaml|profiles\/fullsend-vertex-ai\.yaml|env\/gcp-vertex\.env|schemas\/triage-security-(input|result)\.schema\.json|scripts\/(validate-output-schema\.sh|strip_extra_properties\.py|test_(fullsend_gate_eval|native_fullsend_eval_ci)\.py))$/.test(f.filename) || /^\.github\/(workflows\/eval-pr(-run)?\.yml|scripts\/run-native-fullsend-evals\.sh)$/.test(f.filename) ); - core.setOutput('native', String(bootstrap && nativeRelevant)); + core.setOutput('native', String(nativeRelevant)); core.setOutput('skills', skills); console.log(`Discovered changed skills with evals: ${skills || 'none'}`); @@ -295,6 +292,7 @@ jobs: permissions: contents: read pull-requests: write + issues: write actions: read id-token: write steps: @@ -410,9 +408,37 @@ jobs: exit 1 fi + - name: Render eval results + env: + SKILLS_CSV: ${{ needs.discover.outputs.skills }} + run: | + set -euo pipefail + IFS=',' read -ra SKILLS <<< "$SKILLS_CSV" + version=$(jq -r '.version' plugins/sdlc-workflow/.claude-plugin/plugin.json) + + for skill in "${SKILLS[@]}"; do + workspace="/tmp/${skill}-eval-pr" + python3 plugins/sdlc-workflow/skills/run-evals/scripts/aggregate_benchmark.py \ + --results "$workspace" + + render_args=( + --results "$workspace" + --skill "$skill" + --version "$version" + ) + baseline="evals/${skill}/baselines/latest" + if [[ -d "$baseline" ]]; then + render_args+=(--baseline "$baseline") + fi + + python3 plugins/sdlc-workflow/skills/run-evals/scripts/render_summary.py \ + "${render_args[@]}" + test -s "$workspace/summary.md" + done + - *publication-guard - - name: Post eval results review + - name: Post eval results comment if: steps.publication.outputs.latest == 'true' env: SKILLS_CSV: ${{ needs.discover.outputs.skills }} @@ -426,7 +452,8 @@ jobs: const skills = process.env.SKILLS_CSV.split(',').filter(Boolean); const prNumber = parseInt(process.env.PR_NUMBER); - let body = `## Eval Results\n\nSource head: ${process.env.HEAD_SHA}\nMerge: ${process.env.MERGE_SHA}\n\n`; + const marker = ''; + let body = `${marker}\n## Eval Results\n\nSource head: ${process.env.HEAD_SHA}\nMerge: ${process.env.MERGE_SHA}\n\n`; for (const skill of skills) { const summaryPath = `/tmp/${skill}-eval-pr/summary.md`; if (fs.existsSync(summaryPath)) { @@ -436,31 +463,38 @@ jobs: } } - const reviews = await github.paginate(github.rest.pulls.listReviews, { + const comments = await github.paginate(github.rest.issues.listComments, { owner: context.repo.owner, repo: context.repo.repo, - pull_number: prNumber + issue_number: prNumber, + per_page: 100 }); - const marker = '## Eval Results'; - const existing = reviews.find(r => - r.user?.login === 'github-actions[bot]' && r.commit_id === process.env.HEAD_SHA && r.body?.startsWith(marker) + const existing = comments.find(comment => + comment.user?.login === 'github-actions[bot]' && comment.body?.startsWith(marker) ); + // The sticky comment is shared across heads; a same-head run guard + // alone cannot stop an older head from replacing the current report. + const { data: current } = await github.rest.pulls.get({ + ...context.repo, pull_number: prNumber + }); + if (current.state !== 'open' || current.head?.sha !== process.env.HEAD_SHA) { + core.info('PR head changed; skipping stale sticky report'); + return; + } + if (existing) { - await github.rest.pulls.updateReview({ + await github.rest.issues.updateComment({ owner: context.repo.owner, repo: context.repo.repo, - pull_number: prNumber, - review_id: existing.id, + comment_id: existing.id, body }); } else { - await github.rest.pulls.createReview({ + await github.rest.issues.createComment({ owner: context.repo.owner, repo: context.repo.repo, - pull_number: prNumber, - event: 'COMMENT', - commit_id: process.env.HEAD_SHA, + issue_number: prNumber, body }); } @@ -533,8 +567,7 @@ jobs: const { data: pr } = await github.rest.pulls.get({ ...context.repo, pull_number: Number(process.env.PR_NUMBER) }); - if (pr.state !== 'open' || pr.number !== 299 || pr.base?.ref !== 'main' || - pr.head?.ref !== 'verify-pr-fullsend' || + if (pr.state !== 'open' || pr.base?.ref !== 'main' || pr.head?.sha !== process.env.HEAD_SHA || pr.base?.sha !== process.env.BASE_SHA || pr.merge_commit_sha !== process.env.MERGE_SHA || pr.head?.repo?.full_name !== process.env.SOURCE_REPO || diff --git a/.github/workflows/fullsend-verify-pr.yml b/.github/workflows/fullsend-verify-pr.yml index 9b6f2ed5e..bca2442a8 100644 --- a/.github/workflows/fullsend-verify-pr.yml +++ b/.github/workflows/fullsend-verify-pr.yml @@ -64,12 +64,11 @@ on: # and fired NOTHING for any PR (fork or same-repo). Landing it on main # registers the event; the base-branch copy still runs (trusted base context). # - # `branches:` scopes the rollout to PRs targeting the verify-pr-fullsend feature - # branch only, so verify-pr does not yet review PRs into main. Widen/remove this - # filter when the feature graduates to main. + # `branches:` covers PRs targeting main and the verify-pr-fullsend feature + # branch, so the feature can verify its own deployment before graduation. pull_request_target: types: [opened, synchronize, reopened, labeled] - branches: [verify-pr-fullsend] + branches: [main, verify-pr-fullsend] # One in-flight verify-pr per PR; a new push cancels the superseded run (both # the CI wait and any dispatch already handed off). diff --git a/.github/workflows/fullsend.yaml b/.github/workflows/fullsend.yaml new file mode 100644 index 000000000..62b3965af --- /dev/null +++ b/.github/workflows/fullsend.yaml @@ -0,0 +1,121 @@ +# This file is managed by fullsend. Do not edit it directly. +# Upstream: https://github.com/fullsend-ai/fullsend/blob/main/internal/scaffold/fullsend-repo/.github/workflows/fullsend.yaml +--- +# fullsend shim workflow (per-repo installation mode) +# Routes events to agent workflows via reusable-dispatch.yml. +# All agent execution happens in this repo's context — no external +# config repo is needed. +# +# Security: pull_request_target runs the BASE branch version of this workflow, +# preventing PRs from modifying it to exfiltrate credentials. +# This shim never checks out PR code, so it is not vulnerable to "pwn request" +# attacks. +# +# Routing: this shim forwards the raw event context to reusable-dispatch.yml, +# which determines the stage and runs the agent inline (ADR 62). +# Adding a new stage requires only a job in reusable-dispatch.yml — zero changes to this repo. +# +# Concurrency: per-role cancel-in-progress groups live in reusable-dispatch.yml +# stage jobs with -agent- suffix. Roles operate independently (#2452). +name: fullsend + +on: + issues: + types: [opened, edited, labeled] + issue_comment: + types: [created] + pull_request_target: + types: [opened, synchronize, ready_for_review, closed, labeled, unlabeled] + pull_request_review: + types: [submitted] + +permissions: {} + +jobs: + dispatch: + if: >- + (github.event_name != 'pull_request_target' && github.event_name != 'pull_request_review' + || github.event.pull_request.head.ref != 'fullsend/scaffold-install') + && (github.event_name != 'issue_comment' + || github.event.comment.user.type != 'Bot') + permissions: + actions: write + id-token: write + contents: write + issues: write + packages: read + pull-requests: write + uses: fullsend-ai/fullsend/.github/workflows/reusable-dispatch.yml@v0 + with: + event_action: ${{ github.event.action }} + install_mode: per-repo + mint_url: ${{ vars.FULLSEND_MINT_URL }} + gcp_region: ${{ vars.FULLSEND_GCP_REGION }} + project_number: ${{ vars.FULLSEND_PROJECT_NUMBER }} + runner_image: ubuntu-24.04 + secrets: + FULLSEND_GCP_WIF_PROVIDER: ${{ secrets.FULLSEND_GCP_WIF_PROVIDER }} + FULLSEND_GCP_PROJECT_ID: ${{ secrets.FULLSEND_GCP_PROJECT_ID }} + OTEL_EXPORTER_OTLP_TRACES_HEADERS: ${{ secrets.OTEL_EXPORTER_OTLP_TRACES_HEADERS }} + OTEL_EXPORTER_OTLP_HEADERS: ${{ secrets.OTEL_EXPORTER_OTLP_HEADERS }} + + stop-fix: + # Job-level if: is intentionally coarse — it only screens for the + # /fs-fix-stop command on a PR from a non-bot. The authoritative + # authorization decision (collaborator permission API + PR-author escape + # hatch) is made in the step below, so a maintainer whose author_association + # is not MEMBER (e.g. private org membership) is not filtered out (ADR 0054). + if: >- + github.event_name == 'issue_comment' + && github.event.issue.pull_request + && github.event.comment.user.type != 'Bot' + && github.event.comment.body == '/fs-fix-stop' + runs-on: ubuntu-24.04 + permissions: + contents: read + issues: write + pull-requests: write + steps: + - name: Add fullsend-no-fix label and notify + env: + GH_TOKEN: ${{ github.token }} + PR_NUMBER: ${{ github.event.issue.number }} + REPO: ${{ github.repository }} + COMMENT_USER_LOGIN: ${{ github.event.comment.user.login }} + ISSUE_USER_LOGIN: ${{ github.event.issue.user.login }} + run: | + set -euo pipefail + # ADR 0054: authorize via the collaborator permission API + # (admin|maintain|write), not author_association — the latter grants + # contributor status to anyone with a single merged PR (issue #5421). + # Mirrors has_repo_permission() in dispatch.yml; keep the two in sync. + # The PR author may always stop the fix agent on their own PR. + authorized=false + if [[ -n "$COMMENT_USER_LOGIN" && "$COMMENT_USER_LOGIN" == "$ISSUE_USER_LOGIN" ]]; then + authorized=true + else + if api_err=$(mktemp); then + if role=$(gh api "repos/$REPO/collaborators/$COMMENT_USER_LOGIN/permission" \ + --jq '.role_name' 2>"$api_err"); then + case "$role" in + admin|maintain|write) authorized=true ;; + esac + else + echo "::warning::Permission API call failed for $COMMENT_USER_LOGIN: $(cat "$api_err")" + fi + rm -f "$api_err" + else + echo "::warning::Failed to create temp file for permission check of $COMMENT_USER_LOGIN" + fi + fi + if [[ "$authorized" != "true" ]]; then + echo "::notice::User $COMMENT_USER_LOGIN is not authorized to stop the fix agent (requires write access or PR authorship)" + exit 0 + fi + gh label create "fullsend-no-fix" --repo "$REPO" \ + --description "Skip bot-triggered fix agent runs" --color "FBCA04" \ + --force 2>/dev/null || true + gh pr edit "$PR_NUMBER" --repo "$REPO" \ + --add-label "fullsend-no-fix" + gh pr comment "$PR_NUMBER" --repo "$REPO" \ + --body "Fix agent disabled for this PR. Remove the \`fullsend-no-fix\` label or use \`/fs-fix\` to re-engage." diff --git a/.github/workflows/python-tests.yml b/.github/workflows/python-tests.yml new file mode 100644 index 000000000..1854cc33d --- /dev/null +++ b/.github/workflows/python-tests.yml @@ -0,0 +1,39 @@ +name: Python Tests + +on: + push: + # TEMPORARY: the verify-pr-fullsend entry runs the suite on the feature + # branch before it merges to main. Remove it (leave only `main`) once the + # feature branch is merged — tracked by TC-5816. + branches: [main, verify-pr-fullsend] + pull_request: + # TEMPORARY: the verify-pr-fullsend entry makes this workflow run as a + # required check on PRs targeting the feature branch (e.g. PR #298). + # Remove it (leave only `main`) once the feature branch merges to main — TC-5816. + branches: [main, verify-pr-fullsend] + +jobs: + pytest: + name: Script Unit Tests + runs-on: ubuntu-latest + strategy: + # Run every version so one failure doesn't hide others. + fail-fast: false + matrix: + # No declared minimum Python; test the range contributors are likely + # to run locally, up to the latest stable (3.14, Oct 2025). + python-version: ['3.11', '3.12', '3.13', '3.14'] + steps: + - name: Checkout repository + uses: actions/checkout@v7 + + - name: Set up Python + uses: actions/setup-python@v5 + with: + python-version: ${{ matrix.python-version }} + + - name: Install test dependencies + run: pip install pytest 'jsonschema[format]' pyyaml + + - name: Run pytest + run: python3 -m pytest plugins/sdlc-workflow/scripts/ -q diff --git a/.github/workflows/skillsaw.yml b/.github/workflows/skillsaw.yml index fb3718c9f..72e4a073a 100644 --- a/.github/workflows/skillsaw.yml +++ b/.github/workflows/skillsaw.yml @@ -2,9 +2,15 @@ name: Skillsaw on: push: - branches: [main] + # TEMPORARY: the verify-pr-fullsend entry runs the skill lint on the + # feature branch before it merges to main. Remove it (leave only `main`) + # once the feature branch is merged — tracked by TC-5816. + branches: [main, verify-pr-fullsend] pull_request: - branches: [main] + # TEMPORARY: the verify-pr-fullsend entry makes this workflow run as a + # check on PRs targeting the feature branch (e.g. PR #298). Remove it + # (leave only `main`) once the feature branch merges to main — TC-5816. + branches: [main, verify-pr-fullsend] jobs: lint: diff --git a/.github/workflows/validate-plugins.yml b/.github/workflows/validate-plugins.yml index 15e664fd4..ae952caf3 100644 --- a/.github/workflows/validate-plugins.yml +++ b/.github/workflows/validate-plugins.yml @@ -14,9 +14,15 @@ name: Validate Plugins on: push: - branches: [main] + # TEMPORARY: the verify-pr-fullsend entry runs plugin validation on the + # feature branch before it merges to main. Remove it (leave only `main`) + # once the feature branch is merged — tracked by TC-5816. + branches: [main, verify-pr-fullsend] pull_request: - branches: [main] + # TEMPORARY: the verify-pr-fullsend entry makes this workflow run as a + # check on PRs targeting the feature branch (e.g. PR #298). Remove it + # (leave only `main`) once the feature branch merges to main — TC-5816. + branches: [main, verify-pr-fullsend] jobs: validate: diff --git a/CHANGELOG.md b/CHANGELOG.md index cdc7c76ec..464b5cc1d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,13 @@ All notable changes to the sdlc-workflow plugin are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [Unreleased] + +### Added + +- Fullsend dual-mode execution for `triage-security`, with validated trusted input, + tokenless sandbox analysis, and structured authorized or report-only Jira output + ## [0.13.10] - 2026-10-07 ### Added @@ -153,7 +160,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - Epic creation and grouping strategies with configurable hierarchy preferences in `plan-feature` (TC-4869) - Parent issue linking step in `plan-feature` (TC-4870) -- Early assignment and Assigned transition at Step 0.7 in `triage-security` (TC-5008) +- Early assignment and Assigned transition at Step 0.8 in `triage-security` (TC-5008) - Cross-CVE traceability links and comments at Step 4.3 in `triage-security` (TC-5009) ### Fixed diff --git a/CONVENTIONS.md b/CONVENTIONS.md index 36c88b393..c9d9a3270 100644 --- a/CONVENTIONS.md +++ b/CONVENTIONS.md @@ -9,7 +9,7 @@ - **Primary format**: Markdown documentation - **Configuration**: YAML (`.serena/project.yml`) and JSON (plugin manifests) - **Plugin system**: Claude Code plugin format -- **No source code**: This is a documentation-heavy repository — skills are defined in Markdown (`SKILL.md` files) rather than traditional programming languages +- **Documentation-first, with Python tooling**: Skills are defined in Markdown (`SKILL.md` files), but the repository also ships executable Python and shell helpers under `plugins/sdlc-workflow/scripts/` that back the workflow. Changes to those scripts require running their automated test suite (see **Testing Conventions**) ## Code Style @@ -34,7 +34,7 @@ - **`plugins/sdlc-workflow/`** — main plugin directory - **`skills//`** — individual skill directories, each containing a `SKILL.md` file - **`shared/`** — shared resources like `task-description-template.md` - - **`scripts/`** — utility scripts (if any) + - **`scripts/`** — executable Python and shell helpers (e.g., `execute-actions.py`, `jira-client.py`, `pre-verify-pr.sh`); the core Python modules have a `pytest` unit-test suite (`test_*.py`) alongside them, while `strip_extra_properties.py` and the shell helpers are not yet unit-tested - **`.claude-plugin/`** — plugin manifest (`plugin.json`) - **`.claude-plugin/`** — marketplace manifest at root level (`marketplace.json`) - **`.serena/`** — Serena configuration files @@ -46,7 +46,11 @@ ## Error Handling -Not applicable — this is a documentation repository with no runtime code. +The Markdown skills have no runtime error handling, but the Python scripts under +`plugins/sdlc-workflow/scripts/` do. They must fail fast and loud: validate inputs, +exit non-zero (`sys.exit(1)`) with a message on `stderr` on error, and never swallow +an exception into a silent fallback. Cover both the success and the failure path with +tests (see **Testing Conventions**). ## Testing Conventions @@ -56,7 +60,20 @@ Not applicable — this is a documentation repository with no runtime code. 3. Run `/agents` to verify no plugin agents are missing 4. Edit a `SKILL.md`, then `/reload-plugins` to verify changes are picked up - **CI validation**: Uses `claude plugin validate` on all plugin directories under `plugins/` -- **No automated tests**: Skills are validated through CI and manual testing; no unit test framework +- **Automated unit tests (mandatory)**: The core Python modules under + `plugins/sdlc-workflow/scripts/` (`execute-actions.py`, `jira-client.py`, + `pre_verify_pr.py`) have a `pytest` suite in sibling `test_*.py` files + (`test_execute_actions.py`, `test_jira_client.py`, `test_jira_client_cli.py`, + `test_pre_verify_pr.py`). `strip_extra_properties.py` and the shell helpers + (`pre-verify-pr.sh`, `post-verify-pr.sh`, `validate-output-schema.sh`) are not yet + unit-tested. Any change to a script under `scripts/` **must** add or update the + matching test and keep the whole suite green. Run it before every commit and before opening or updating + a PR: + ```bash + python3 -m pytest plugins/sdlc-workflow/scripts/ -q + ``` + A passing run is a precondition for merge, not an optional step — do not commit a + script change without running it. - **Fixture documentation**: Eval and test fixture files must include a leading comment header in the file's native comment syntax (e.g., `` for Markdown/HTML, `// ...` for JSON with comments, `# ...` for YAML) explaining that the content is deliberate test material. Use the canonical prefixes below so tooling (linters, scanners, grep filters) can reliably identify annotated fixtures. Two categories require annotation: - **Adversarial fixtures** — files containing intentionally adversarial, malicious-looking, or unusual content (e.g., injection vectors, malformed input, security-sensitive patterns). Use the prefix `ADVERSARIAL TEST FIXTURE — ` (e.g., ``). - **Synthetic data fixtures** — files representing synthetic or mock entities (e.g., fake repository structures, mock Jira issues, fabricated API responses). Use the prefix `SYNTHETIC TEST DATA — ` (e.g., ``). @@ -70,10 +87,17 @@ Not applicable — this is a documentation repository with no runtime code. ## CI Checks -All CI checks must pass before merging. Run locally before pushing: +All checks must pass before merging. Run locally before pushing: ```bash +# 1. Skill instruction lint uvx skillsaw + +# 2. Plugin manifest validation +claude plugin validate plugins/sdlc-workflow + +# 3. Python unit tests — required whenever anything under scripts/ changes +python3 -m pytest plugins/sdlc-workflow/scripts/ -q ``` ### Skill Lint (Skillsaw) @@ -90,6 +114,23 @@ claude plugin validate plugins/sdlc-workflow Validates plugin manifests under `plugins/`. CI workflow: `.github/workflows/validate-plugins.yml`. +### Python Unit Tests + +The core Python modules under `plugins/sdlc-workflow/scripts/` (`execute-actions.py`, +`jira-client.py`, `pre_verify_pr.py`) are covered by a `pytest` suite in sibling +`test_*.py` files; `strip_extra_properties.py` and the shell helpers +(`pre-verify-pr.sh`, `post-verify-pr.sh`, `validate-output-schema.sh`) are not yet +covered. Run the full suite and keep it green whenever you touch anything under +`scripts/`: + +```bash +python3 -m pytest plugins/sdlc-workflow/scripts/ -q +``` + +This suite is **required before every commit that changes a script and before merge**, +and is enforced in CI by `.github/workflows/python-tests.yml` (runs on pushes and pull +requests targeting `main`). Run it locally before pushing so failures surface before CI. + ## Commit Messages - **Format**: Conventional Commits — `type(scope): description` @@ -143,7 +184,7 @@ Validates plugin manifests under `plugins/`. CI workflow: `.github/workflows/val ## Dependencies -- **No external dependencies** — this repository contains only documentation and configuration files +- **No external dependencies for the plugins themselves** — the Claude Code plugins are Markdown, YAML, and JSON. The repository also ships executable Python and shell helpers under `plugins/sdlc-workflow/scripts/`: these require Python 3 (`validate-output-schema.sh` additionally requires the `jsonschema` package at runtime), and the test suite requires `pytest` (plus `jsonschema`) - **Runtime dependency**: Claude Code CLI (users must have Claude Code installed to use the plugins) - **Plugin system**: Uses Claude Code's plugin marketplace and validation system (`claude plugin validate`) - **Version synchronization**: The plugin version must be kept in sync between: diff --git a/docs/constraints.md b/docs/constraints.md index b5bf1cb99..efa454956 100644 --- a/docs/constraints.md +++ b/docs/constraints.md @@ -101,6 +101,14 @@ existing instruction in a SKILL.md or CLAUDE.md file. | 1.89 | `triage-bug` MUST present the matched version to the user for confirmation before setting the affectsVersions field. | `triage-bug/SKILL.md` — Step 4.5 | | 1.90 | `triage-bug` MUST NOT set affectsVersions without explicit user confirmation. | `triage-bug/SKILL.md` — Step 4.5 | | 1.91 | `triage-bug` MUST post a gap-flagging comment when version info cannot be determined from the bug description. | `triage-bug/SKILL.md` — Step 4.5 | +| 1.92 | `triage-security` MUST preserve the interactive workflow only when `FULLSEND_OUTPUT_DIR` is absent; a present but empty value MUST fail closed, and Fullsend mode MUST NOT fall back to interactive behavior. | `triage-security/SKILL.md` — Step 0.6 | +| 1.93 | In Fullsend mode, `triage-security` MUST consume only schema-validated trusted input containing issue, remote-link, configuration, external, matrix/source, Jira metadata, idempotency, and authorization evidence before analysis or mutation planning. The shipped prefetch collector MUST set `authorization.mutation_authorized` to `false`. | `triage-security/SKILL.md` — Step 0.7, Trusted-evidence map; `triage-security-input.schema.json`; `pre_triage_security.py` | +| 1.94 | The Fullsend sandbox MUST NOT call Jira, GitHub, web, Git, or credentialed source services, and MUST NOT write a security matrix. | `triage-security/SKILL.md` — Step 0.6, Guardrails | +| 1.95 | A Fullsend `report-only` result MUST contain only report-only actions; mutation-authorized actions MUST be schema-defined and require trusted `authorization.mutation_authorized: true`. | `triage-security/SKILL.md` — Fullsend action serialization; `triage-security-result.schema.json` | +| 1.96 | Only the trusted post-execution executor MAY perform Fullsend Jira mutations using the prevalidated trusted input after independently validating the result and its authorization flag. | `post-triage-security.sh`; `execute-triage-security-actions.py` | +| 1.97 | The trusted executor MUST resolve generated references before dependent Jira calls and fail unresolved placeholders rather than inventing Jira keys. | `triage-security/SKILL.md` — Fullsend action serialization; `execute-triage-security-actions.py` | +| 1.98 | Fullsend retries MUST use stable `triage-security:` markers and trusted retry state to avoid duplicate comments, labels, links, transitions, remediation tasks, and remediation description digests. | `triage-security/SKILL.md` — Fullsend action serialization; `execute-triage-security-actions.py` | +| 1.99 | Missing, malformed, or contradictory required Fullsend evidence MUST fail closed before mutation. The sandbox MUST treat incomplete or inconsistent supplied evidence as a blocked report-only outcome, and MUST NOT refresh or repair stale matrix data. | `triage-security/SKILL.md` — Steps 0.7, 0.3; `agents/triage-security.md` — Constraints | ### Prior Art — Cross-phase integrity (§1.33–1.35) @@ -195,7 +203,12 @@ Each constraint above references its source. The full source files are: - `plugins/sdlc-workflow/skills/define-feature/SKILL.md` — Guardrails (§1.7–1.8), Important Rules (§1.9), Step 3.5 Select Priority and Fix Version (§1.72), Step 6.2 Create Feature Issue (§1.73) - `plugins/sdlc-workflow/skills/report-bug/SKILL.md` — Guardrails (§1.50–1.51), Step 4 Preview and Confirm (§1.52) - `plugins/sdlc-workflow/skills/triage-bug/SKILL.md` — Guardrails (§1.53–1.54), Step 1 Parse bug description (§1.88), Step 4 Post root cause comment (§1.56), Step 4.5 Affects Version Resolution (§1.89, §1.90, §1.91), Step 5 Front-load reproducer test (§1.55), Step 6 Decomposition Guard (§1.57) -- `plugins/sdlc-workflow/skills/triage-security/SKILL.md` — Guardrails (§1.37, §1.38, §1.47), Step 0 (§1.49, §1.68), Step 0 Deployment Context (§1.77, §1.78), Step 1 Ecosystem detection (§1.48), Step 1 Data Extraction (§1.49), Step 1.5 External CVE Data Enrichment (§1.62), Step 1.7 Embargo Check (§1.70, §1.71), Discovery mode (§1.69), Step 2.1 (§1.47), Step 2.2 (§1.42), Step 2.3 (§1.48), Step 2.4 (§1.44), Step 8 (§1.43), Step 8 Case B (§1.61), Post-Triage Summary (§1.67), Important Rules (§1.38–§1.43, §1.45, §1.46), Remediation Task Creation (§1.46, §1.66) +- `plugins/sdlc-workflow/skills/triage-security/SKILL.md` — Guardrails (§1.37, §1.38, §1.47, §1.94), Step 0 (§1.49, §1.68), Step 0.3 Matrix Staleness (§1.99), Steps 0.6–0.7 Fullsend mode and trusted input (§1.92–§1.95, §1.99), Fullsend action serialization (§1.95, §1.97, §1.98), Step 0 Deployment Context (§1.77, §1.78), Step 1 Ecosystem detection (§1.48), Step 1 Data Extraction (§1.49), Step 1.5 External CVE Data Enrichment (§1.62), Step 1.7 Embargo Check (§1.70, §1.71), Discovery mode (§1.69), Step 2.1 (§1.47), Step 2.2 (§1.42), Step 2.3 (§1.48), Step 2.4 (§1.44), Step 8 (§1.43), Step 8 Case B (§1.61), Post-Triage Summary (§1.67), Important Rules (§1.38–§1.43, §1.45, §1.46), Remediation Task Creation (§1.46, §1.66) +- `plugins/sdlc-workflow/agents/triage-security.md` — Sandbox constraints (§1.99) +- `plugins/sdlc-workflow/schemas/triage-security-input.schema.json` — Trusted input contract (§1.93) +- `plugins/sdlc-workflow/schemas/triage-security-result.schema.json` — Result modes and actions (§1.95) +- `plugins/sdlc-workflow/scripts/post-triage-security.sh` — Trusted result handoff (§1.96) +- `plugins/sdlc-workflow/scripts/execute-triage-security-actions.py` — Authorization, reference resolution, and idempotent execution (§1.96–§1.98) - `plugins/sdlc-workflow/skills/triage-security/version-impact-analysis.md` — Step 2.3 enriched fix threshold (§1.62) - `plugins/sdlc-workflow/skills/triage-security/remediation-templates.md` — Jira Issue Creation digest comment (§1.66) - `plugins/sdlc-workflow/skills/triage-security/jira-triage-operations.md` — Step 3 ProdSec @mention (§1.68), Step 4.2 Idempotent sibling linking (§1.58), Step 4.3 Cross-CVE overlap detection (§1.59), Step 4.3 ProdSec @mention (§1.68), Step 4.4 Proactive reconciliation (§1.61) diff --git a/docs/plans/notes/task1-findings.md b/docs/plans/notes/task1-findings.md new file mode 100644 index 000000000..a4aed67e5 --- /dev/null +++ b/docs/plans/notes/task1-findings.md @@ -0,0 +1,112 @@ +# Task 1 findings — Environment + delivery/resolution confirmation + +**Jira:** [TC-5805](https://redhat.atlassian.net/browse/TC-5805) (Epic A: Delivery foundation, parent TC-5800) +**Date:** 2026-08-27 +**fullsend CLI:** v0.37.0 (built from `github.com/fullsend-ai/fullsend` tag `v0.37.0`) + +## IMAGE digest + +``` +IMAGE=ghcr.io/fullsend-ai/fullsend-code@sha256:9743bc7b6e451e0bcea25ae4a67e0c040c296f1fee04c08988ae80c53fafcfe6 +``` + +Source: `image:` in `/Users/mrizzi/git/cloned/agents/harness/review.yaml`. + +## 1. Toolchain — fullsend v0.37.0 + +Built with the version ldflag (plain `go build` stamps `version dev`): + +```bash +go build -C /Users/mrizzi/git/cloned/fullsend \ + -ldflags "-X github.com/fullsend-ai/fullsend/internal/cli.version=v0.37.0" \ + -o /Users/mrizzi/.local/bin/fullsend ./cmd/fullsend +fullsend --version # -> fullsend version v0.37.0 +``` + +**Confirmed:** `fullsend --version` reports `v0.37.0`. + +## 2. Stock image binaries + +`podman run --rm --entrypoint "" "$IMAGE" bash -lc 'command -v ...'`: + +| binary | path | +|---------|-----------------------------| +| claude | /usr/local/bin/claude | +| node | /usr/bin/node | +| python3 | /sandbox/.venv/bin/python3 | +| gh | /usr/bin/gh | +| git | /usr/bin/git | +| jq | /usr/bin/jq | +| curl | /usr/bin/curl | + +**Confirmed:** all required binaries present, none MISSING. `claude`/`node`/`python3`/`gh`/`git` +all resolve (python3 = `/sandbox/.venv/bin/python3`, required for ADR-0090 python sandbox hooks). + +➡️ **TC-5806 (sandbox image extension) is a NO-OP** — no required binary is missing. + +## 3. Root-level harness resolution + +Validated against a throwaway probe (`role: verify-pr`, `plugins: [plugins/sdlc-workflow]`, +`image: $IMAGE`) run from the repo root; probe never committed (working tree clean afterwards). + +**Confirmed:** a root-level harness bases its relative children at the repo root, so +`plugins/sdlc-workflow` resolves in place and the `sdlc-workflow:verify-pr` skill is present +with its sibling `shared/` intact: + +- `plugin.json` name = `sdlc-workflow`; `skills/` and `shared/` are siblings at the plugin root. +- `plugins/sdlc-workflow/skills/verify-pr/SKILL.md` exists; the skill's `../../shared/*.md` + links resolve to `plugins/sdlc-workflow/shared/` (comment-footnote.md, jira-rest-fallback.md, …). +- Skill id therefore = `sdlc-workflow:verify-pr`. + +Evidence chain (resolution-only, no dispatch): + +- `fullsend agent add harness/probe-verify-pr.yaml --name probe-verify-pr --fullsend-dir .` + → `✓ Added agent "probe-verify-pr"` (the root-level harness path resolves at the repo root). +- `fullsend lock probe-verify-pr --fullsend-dir . --offline --max-depth 0` + → loads + validates the harness and runs `ResolveRelativeTo(absFullsendDir)` with **no** + "resolves outside fullsend directory" error → `plugins/sdlc-workflow` resolves cleanly + inside the repo root; "no remote dependencies" confirms it is treated as a local, in-place path. +- Source confirmation (fullsend v0.37.0): + - `internal/cli/run.go:464` and `internal/cli/lock.go:248` call + `Harness.ResolveRelativeTo(absFullsendDir)` — relative local paths resolve against the + **fullsend-dir**, which must be the repo root. + - `internal/cli/run.go:670` calls `Harness.ValidateFilesExist()`, which `os.Stat`s every + plugin path (`internal/harness/harness.go:798`). + +### ⚠️ Deviations from the task's literal commands (important for TC-5807) + +These are real constraints the delivery model must account for; the task's example commands +do **not** work verbatim on fullsend v0.37.0: + +1. **`--fullsend-dir` must be the repo root**, not a scratch `/tmp` dir. fullsend resolves the + harness source path *and* the harness's relative children (`plugins/…`) against + `absFullsendDir` = the `--fullsend-dir` value (`ResolveRelativeTo`, `validateLocalPath`). + With `--fullsend-dir /tmp/probe-fs` the probe fails + (`local path does not exist: /tmp/probe-fs/harness/probe-verify-pr.yaml`) and + `plugins/sdlc-workflow` would resolve under `/tmp`, not the repo. The probe here used + `--fullsend-dir .` at the repo root and a scratch `config.yaml`/`harness/` (both removed). + +2. **`verify-pr` is NOT a valid `config.yaml` role.** `config.yaml` `roles:` are validated + against a fixed enum — `fullsend, triage, coder, review, fix, retro, prioritize, e2e` + (`internal/config/config.go:196`); `verify-pr` is rejected + (`invalid role "verify-pr": must be one of …`). The **harness** `role:` field is only + regex-validated (`internal/harness/harness.go:474`), so `role: verify-pr` is accepted in the + harness YAML. The scratch `config.yaml` therefore declared a valid enum role (`review`) while + the harness kept `role: verify-pr`, and `agent add` succeeded. TC-5807 must reconcile this: + the installation `config.yaml` cannot list `verify-pr` as a role in v0.37.0. + +3. **`fullsend agent add` does not stage plugins.** It only validates the harness source path + exists and records the source string in `config.yaml` (`internal/cli/agent.go:174-205`). + Plugin resolution/staging and existence checks happen at `fullsend run` time. The + `fullsend run … --offline --max-depth 0` "secondary gate" was **not** executed here because + `fullsend run` sets up the sandbox and dispatches a real agent (not resolution-only); + resolution was confirmed via `agent add` + `lock` + source review instead. + +## Acceptance + +- [x] `fullsend --version` reports v0.37.0; IMAGE digest recorded. +- [x] Required binaries present (none MISSING) → TC-5806 no-op. +- [x] Throwaway probe: `agent add` resolves the in-place plugin at the repo root; + `sdlc-workflow:verify-pr` present with `skills/verify-pr/SKILL.md` + sibling `shared/`; + probe removed (working tree clean). +- [x] Findings committed (trailers required). diff --git a/docs/superpowers/specs/2026-09-22-triage-security-fullsend-design.md b/docs/superpowers/specs/2026-09-22-triage-security-fullsend-design.md new file mode 100644 index 000000000..c2a9f6756 --- /dev/null +++ b/docs/superpowers/specs/2026-09-22-triage-security-fullsend-design.md @@ -0,0 +1,103 @@ +# Triage-Security Fullsend Dual-Mode Design + +## Goal + +Allow `triage-security` to perform the existing CVE decision tree in a tokenless +Fullsend sandbox while preserving the interactive Jira, web, and Git behavior when +`FULLSEND_OUTPUT_DIR` is genuinely absent. + +## Scope + +The implementation changes only the triage-security skill instructions, its three +companion procedure documents, and its eval contract. The existing Fullsend harness, +trusted input schema, result schema, prefetch script, action executor, and output +validator are dependencies supplied by TC-6207 through TC-6209 and are not changed. + +## Mode Selection and Input + +The skill detects Fullsend from the *presence* of `FULLSEND_OUTPUT_DIR`. An exported +empty value is a configuration error: it must fail closed, rather than selecting the +credentialed interactive path. When the variable is absent, every current interactive +step and confirmation gate remains in force. + +When the variable is non-empty, the skill reads only +`/sandbox/workspace/.pre-script/triage-security-input.json` and validates it against +`${CLAUDE_PLUGIN_ROOT}/schemas/triage-security-input.schema.json`. Missing, +unparseable, or schema-invalid data produces a failure result and stops the workflow; +it never triggers Jira, GitHub, web, CVE, or credentialed repository fallback. Mounted +repository evidence is read-only. The sandbox never writes `security-matrix.md`. + +The failure result is deliberately `{ "error": "..." }`, without the result schema's +required `schema_version`, `mode`, `report`, and `actions` members. The runner's +output validation must reject it visibly, matching verify-pr's fail-closed behavior; +the skill stops immediately after writing it. + +## Analysis Flow + +Fullsend preserves the interactive order: configuration and staleness evaluation; +issue extraction, enrichment, and embargo assessment; version-impact analysis; Jira +operations; remediation; and the post-triage summary. Each stage consumes the matching +validated bundle section instead of performing an external read: + +| Interactive source | Fullsend source | +|---|---| +| CLAUDE.md/Jira configuration | `configuration` | +| Jira issue and remote links | `issue`, `remote_links`, `jira_metadata` | +| MITRE, OSV, lifecycle pages | `external_evidence` | +| matrices and release/source lookups | `matrix`, `source_evidence` | +| existing actions/tasks/comments/links | `idempotency` | + +Evidence requirements do not weaken in Fullsend: all supported versions, released +pinned commits, development heads, retags, lock-file and dependency-chain evidence, +external fix-threshold precedence, duplicate checks, and reporter/ProdSec data retain +their current semantics. + +## Result and Mutations + +The sandbox accumulates a result conforming to +`triage-security-result.schema.json`: `schema_version`, `mode`, an evidence-backed +`report`, and ordered `actions`. Its stable action markers make an action-plan rerun +idempotent. + +When `authorization.mutation_authorized` is false, `mode` is `report-only` and the +only action is `report-only`. The report names every withheld mutation and explains +that trusted-runner authorization is required. + +When it is true, each interactive mutation becomes one existing schema action: + +| Interactive operation | Fullsend action | +|---|---| +| Assignment, Affects Versions, VEX, labels | `field-edit` | +| Assigned, In Progress, Closed transitions | `status-transition` | +| Triage, digest, reconciliation, and summary comments | `comment` | +| Related, Depend, and Blocks relationships | `link` | +| Remediation task creation | `remediation-task`, followed by references and links | + +Action order preserves the original protocol: task creation, description digest +comment, reference resolution, links, and later comments. Interactive confirmation +prompts become deterministic authorization in Fullsend; interactive confirmations are +unchanged. + +## Documentation Changes + +`SKILL.md` defines the top-level dual-mode contract and applies it to every numbered +step. `version-impact-analysis.md` substitutes trusted matrix/source/lifecycle evidence +for every external read and blocks matrix repair in the sandbox. +`jira-triage-operations.md` turns every write surface into an action or report-only +recommendation. `remediation-templates.md` specifies serialized task descriptions, +labels, links, digest comments, action-reference ordering, and report-only behavior. + +## Evaluation + +The triage-security evals retain current standard, already-fixed, duplicate, +split-stream, RPM, enrichment, overlap, reconciliation, and rerun scenarios. New +assertions cover gate detection, valid-bundle-only execution, no sandbox-side external +calls or writes, report-only output, authorized structured actions, preserved ordering, +and idempotent retry markers. They also assert that an absent Fullsend gate continues +to use the established interactive path. + +## Non-Goals + +This work does not change the JSON schemas, harness, pre/post scripts, action executor, +security-matrix format, existing interactive decision rules, or Jira state outside the +task’s normal implementation lifecycle. diff --git a/docs/testing/fullsend-gate-evals.md b/docs/testing/fullsend-gate-evals.md new file mode 100644 index 000000000..43b041ea6 --- /dev/null +++ b/docs/testing/fullsend-gate-evals.md @@ -0,0 +1,313 @@ +# Native Fullsend gate evals + +TC-6677 moves cases 033–036 into a separate native Fullsend suite. Ordinary +`sdlc-workflow:run-evals` retains the original triage32/164 and verify6/68. +This suite invokes the actual `sdlc-workflow:triage-security` Skill for synthetic +TC-8101 through a test agent. It covers real Skill execution in Fullsend with +synthetic bundles; full production pre/post integration and native verify-pr coverage remain separate. +TC-6213 stays closed for its approved scope; hosted rollout evidence belongs to TC-6726. + +**Local native execution is proven:** fresh run +`tc6677-8d4aab96a2684e47ab5f1fdf65700a8b` passed all 21/21 on 2026-10-05, +with real Skill/native tool evidence, malformed rejection and report-only final +validation. **Hosted WIF execution remains unproven until a real GitHub run.** +Static tests and preflight do not establish hosted credentials, infrastructure or grading. + +## Versions and prerequisites + +Use Python3.12, Git and curl on macOS or Linux (amd64/arm64). `setup` installs +only into the chosen cache. It verifies Fullsendv0.43.0 release archives against +the SHA256 values in [dependencies.json](../../evals/fullsend/dependencies.json), +and fetches canonical agent-eval-harness1.22.0 at immutable commit +`4b540c652f5ed325e18abf6b4bd0eb4414a4bb3c`. Python runtime/Vertex/build +dependencies are exactly versioned and hash-locked in +[requirements.lock](../../evals/fullsend/requirements.lock). The harness wheel +is built from that verified source with locked build tools; its generated wheel +bytes are not claimed reproducible. OpenShell, Podman, the host OS, service +configuration and the delivered model runtime inside the production image are +outside the Python lock. + +The operator must provide installed **OpenShell CLI and gateway0.0.116**, Podman, +an approved running gateway and container-driver configuration, and an accessible +production sandbox image. The pin is Fullsendv0.43.0's +[OpenShell pin file](https://github.com/fullsend-ai/fullsend/blob/d5f36921ac754705619f38c637ef692873809fbc/.github/scripts/openshell-version.sh). +The agents `LOCAL.md` copy mentions older0.0.83; do not use that version. +The inspected0.0.116 CLI has `gateway add/select/list`, **no `gateway start`**; +provision/run `openshell-gateway` using your existing approved authenticated +configuration. The Python entrypoint does not install or start host services, create +credentials, change gateway/TLS configuration, or relax sandbox policy. + +Platform differences: + +- **macOS:** Python3.12 and curl may need separate installation; Podman needs + an initialized/running machine. Paths are resolved to physical paths (including + `/private/tmp`) for delivery. Setup downloads the Darwin host binary and the + matching Linux binary for the sandbox. The VM/image CPU architecture must match + the selected host architecture; cross-architecture execution is not prepared here. +- **Linux/CI:** provide Python3.12 with venv support, Git, curl and CA certificates. + Rootless Podman needs valid subordinate UID/GID mappings and an active API socket, + normally `${XDG_RUNTIME_DIR}/podman/podman.sock`. Provision these through the + runner's approved host setup, along with OpenShell's authenticated gateway and + required supervisor image. The existing Fullsend + [functional CI source](https://github.com/fullsend-ai/fullsend/blob/d5f36921ac754705619f38c637ef692873809fbc/.github/workflows/functional-tests.yml) + documents this host dependency. The trusted CI wrapper reuses the pinned upstream + installers and rootless Podman setup; Python `setup` remains dependency-only. + +For further platform context, consult the pinned +[Fullsend local guide](https://github.com/fullsend-ai/fullsend/blob/d5f36921ac754705619f38c637ef692873809fbc/docs/guides/user/running-agents-locally.md). +Use the actual installed OpenShell CLI/help and approved gateway configuration +where older guide commands differ. + +## Common local and CI commands + +Run from the reviewed repository checkout. The following two commands perform +**dependency/CLI checks only** and need no inference credentials: + +```bash +python3.12 evals/fullsend/run.py setup --cache /tmp/tc-6677-eval-deps +python3.12 evals/fullsend/run.py preflight --cache /tmp/tc-6677-eval-deps +``` + +`setup` fetches pinned source/packages/releases; it performs no global install. +It overrides host pip user-install defaults and keeps pip's cache in the selected +directory. curl retains normal host TLS verification; SHA256 verification is +mandatory before installing a binary. A corrupt cached archive fails visibly; +remove that specific archive and repeat setup after investigating the mismatch. +`preflight` checks exact package versions, clean pinned framework source, +upstream workspace/execute/collect/score CLI imports, suite configuration, +Fullsend CLI flags, and OpenShell/Podman versions. It does not prove gateway +reachability, authentication, image availability, environment propagation or tools. + +For actual execution, the operator supplies existing Vertex inference credentials +and host judge credentials through these environment variables: + +| Variable | Required input | +|---|---| +| `GOOGLE_APPLICATION_CREDENTIALS` | Absolute path to the operator-provided GCP credential file, accessible to native Fullsend and the host judge | +| `ANTHROPIC_VERTEX_PROJECT_ID` | Vertex project with access to the selected Claude models | +| `GOOGLE_CLOUD_PROJECT` | GCP project used by the production Vertex environment mount | +| `CLOUD_ML_REGION` | Vertex region supporting both selected models | + +The production credential provider/profile/environment mount is reused unchanged. +The test harness uses the existing native host-file pattern to upload the +operator-provided `GOOGLE_APPLICATION_CREDENTIALS` file directly to +`/tmp/.gcp-credentials.json`, the path referenced by that environment template. +This mount is required and is not expanded as text. The adapter never reads or +stages credential contents in `native-config`, the checkout or output. Native +Fullsend controls the upload into the sandbox; the host judge keeps using the +original operator-provided path. +Do not print environment values, place credential files in the checkout/output, +or redirect `GOOGLE_APPLICATION_CREDENTIALS` to its sandbox path for host scoring. +`FULLSEND_MINT_URL` must be unset for this synthetic suite: the adapter refuses it +because native Fullsend would otherwise attempt live forge-token minting. No Jira +or GitHub fixture/token/issue URL is needed. No live prefetch, post-script or status +notification is configured. + +After host services and those inputs are ready, the operator can run: + +```bash +python3.12 evals/fullsend/run.py run \ + --cache /tmp/tc-6677-eval-deps \ + --output /tmp/tc-6677-native-evals \ + --model claude-opus-4-8 \ + --judge-model claude-opus-4-6 \ + --effort high +``` + +This command **does perform paid inference**: four serial native Fullsend runs +and21 upstream Boolean LLM judgments. Select model IDs supported by your Vertex +project/region. Model availability is not checked by preflight. The native +agent timeout is30minutes, the opaque CLI case timeout40minutes, and the test +validation loop has one iteration. Framework budget hints are advisory, not a +spend cap. A zero/unknown framework cost is not evidence of zero actual cost; +the unchanged native metrics use `total_cost_usd`, while CliRunner looks for +`cost_usd`. Use the retained native metrics. + +## Automatic CI, WIF and rollout + +PR323 adds only four trusted CI files to main: the two eval workflows, the CI +wrapper and workflow contract tests. This native suite, dependencies, cases, +fixtures, judges, companions and documentation are delivered by PR299. + +The bootstrap initially runs native evals only for PR299 targeting main from +`verify-pr-fullsend`. It separately checks out the explicitly reviewed immutable +PR299 suite commit and the exact approved PR merge used as sandbox plugin input. +The wrapper verifies the reviewed suite SHA before installing or executing its +code. Workflow and suite provenance are both included in the safe result. +PR-controlled head scripts, dependency locks or validation commands never run on +the credentialed host. Plugin root, ancestor and child symlinks are rejected. + +Approval reuses the existing write/admin collaborator check and `eval-protected` +environment for external authors. Discovery pins the API-associated head/base/ +merge and verifies merge parents; native execution rechecks after approval and +before WIF. Changed revisions require a new run. Checkouts use immutable SHAs +and `persist-credentials: false`. + +The readonly native job reuses FULLSEND_GCP_WIF_PROVIDER/PROJECT_ID and region. +Pinned upstream Fullsend source supplies OpenShell/Podman setup and sandbox WIF +preparation. Original host ADC stays with Vertex judging; separate prepared +sandbox ADC plus native OIDC token/refresh inputs go to Fullsend. No long-lived +key, GitHub App, Jira token, forge mint or production pre/post hook is added. +Trusted suite resources supply the pre-script, policy/profile/provider and host +schema validator. Only the selected PR plugin is sandbox test content. + +All four cases and21 Boolean judgments retain their original texts. Missing, +null, skipped applicable judgments, errors or failed scoring fail reporting and +the combined ordinary/native check. Expected negative native exits remain intact. +Only source pins, Boolean outcomes, counts and exit/completeness status are +published as `native-result.json`, retained14days. Raw logs, arbitrary rationale, +transcripts, configs and credentials stay private in runner temporary storage. +Reporting writes use a separate job without inference credentials. + +Merge sequence: + +1. Review the native suite source on PR299 and its explicit immutable pin in the + minimal PR323 bootstrap. Human merges PR323 first. +2. The restricted main workflow evaluates PR299's exact approved plugin revision + with that reviewed suite source. Require actual WIF-backed21/21 alongside + successful ordinary evals; local21/21 and preflight do not prove hosted success. +3. Human merges PR299 after successful hosted validation. PR299's normal + activation removes only the initial identity restriction and selects trusted + `github.sha` as suite source. Future runs then use the trusted suite on main. + +No arbitrary PR suite script is promoted at runtime. Hosted WIF audience/policy, +gateway/image/model availability and refresh remain unproven until the real run. +Native verify-pr and production pre/post mutation coverage remain separate. + +## Cases and raw evidence + +| Case | Input/gate | Strict assertions | +|---|---|---:| +| 033-absent | Test fragment unsets native gate; target CLAUDE.md lacks Security Configuration | 4 | +| 034-empty | Test fragment exports an empty gate; no target CLAUDE.md/input | 5 | +| 035-malformed | Native nonempty gate; exact retained malformed bytes mounted | 5 | +| 036-valid | Native nonempty gate; retained trusted report-only bundle mounted | 7 | + +The input mount is `/sandbox/workspace/.pre-script/triage-security-input.json`. +The native output directory is `/sandbox/workspace/output`, **not `/sandbox/output`**. +Negative gate injection uses a mounted test-only `.env.d` fragment sourced before +model launch, as supported by Fullsendv0.43.0 `bootstrapEnv` and Claude runtime +`buildRunCommand`. It is deliberate negative configuration, not a normal Fullsend +configuration. Source support does not establish runtime propagation. + +The test agent bypasses only the production agent's input-before-Skill startup +guard by invoking the actual Skill first. It does not duplicate gate/validator +logic or run nested model CLIs. The absent case must reach the existing interactive +missing-configuration guard before credentials. Empty must fail at the precise +gate instruction. Malformed must execute the actual input validator and error-only +abort. Valid must perform real analysis, write the completed result, then execute +the actual inline final JSON/schema validator. Every assertion requires genuine +Skill/tool records and intended plugin binding; narrated outcomes fail. + +Each local case copies the unchanged delivered plugin (including script companions), +test agent/pre-script and the three retained synthetic fixtures into its isolated +`native-config` directory. Resource paths and fixture/schema references point +inside that directory, which Fullsend uses as the resolver workspace root through +`--fullsend-dir`. The separate synthetic target is a fresh local `git init` +repository, with no remote or commit. Its only project file is the absent case's +`CLAUDE.md`; the other cases have no project configuration or input content. +Pinned Fullsend0.43.0 `UploadDir` includes `.git`, and read-only setup requires +that metadata directory; neither operation requires a commit. The fresh local 21/21 run exercised this source +contract; hosted setup still needs its own real WIF run. The target never points at the +real repository. No containment check is disabled and external resource symlinks +are not used. + +Generated host mounts resolve to `native-config/pre/`. The test pre-script writes +the required exact gate fragment there for every case, and the exact retained +input for malformed/valid. It uses the known configuration root; Fullsend0.43.0 +does not provide `FULLSEND_RUN_DIR` to pre-scripts or host-file bootstrap (only +host validation commands receive it). Both generated mounts are optional during +early environment/file validation because preparation has not run yet. A failed +or stale preparation returns nonzero and Fullsend aborts before sandbox creation; +optional mounting does not turn that failure into success. The unchanged strict +runtime assertions still require the real mounted gate state and Skill execution. + +The run prints its fresh destination: +`/triage-security-gate//`. Keep that entire directory and, if needed +for diagnosing framework failures, the printed upstream temporary workspace. +Upstream `workspace.py`, `execute.py`, `collect.py`, and `score.py` own case iteration, +artifact collection and grading. The adapter forwards native stdout/stderr and +actual exit unchanged; collection copies native bytes without rewriting records. + +Expected retained evidence includes: + +- `cases//run_result.json`, `stdout.log`, `stderr.log`: upstream actual process + exit and native console output, distinct from individual Bash tool exits. +- `cases//output/native/agent-*/iteration-1/transcripts/`: actual native runtime + records showing Skill invocation/body, subsequent tool calls/results and failures. +- `cases//output/native/agent-*/iteration-1/output/`: retained native output inventory after host validation. + Malformed may produce no file or sole `agent-result.json` containing `{}` after host stripping; + valid produces sole `agent-result.json`, while absent/empty produce none. +- Native metrics/logs/traces and the unchanged root metrics copy when uniquely found. +- Upstream judge results and summary, with all21 individual Boolean results/rationales. + +Negative cases may cause a nonzero native CLI exit because the production host +schema intentionally rejects absent/no result or the malformed abort result. +For malformed input, actual delivered Skill/tool records must prove invalid JSON +rejection with parser detail and tool exit1, native nonzero and no successful +analysis, fallback or actions. Either no result file with host rejection of the +absent result, or sole collected `agent-result.json` containing `{}` after intentional +stripping, is the expected negative outcome. An absent output directory or failed +attempted abort write **after proven real Skill invalid JSON rejection** is accepted +on the no-result path, not a disqualifying bootstrap/inference failure. No recovery +write is required when the abort creates no file. If the error-only file was written, raw tools must prove +the prescribed object was successfully written **before host validation**, followed +by host `strip_extra_properties.py` removing `error` and rejecting the success schema. +Empty output or nonzero alone cannot pass; genuine rejection and host records for +the applicable path are mandatory. Nonempty unexpected output or a success report +fails. No extra sandbox evidence file is required. The host validation loop is +separate from the Skill inline validator. +For malformed, infrastructure/inference failure is disqualifying when it prevents +actual Skill input validation. Without genuine invalid JSON proof, the case fails. +Failures preventing the required Skill execution remain failures for every case. The +entrypoint continues collection/scoring after upstream execution failure while +retaining raw exits; it refuses missing case results and delegates verdicts to +upstream judges. + +At the pinned framework revision, partial judge exceptions yield `value:null` +and are omitted from aggregate values. After successful upstream scoring, the +entrypoint checks the unchanged `summary.yaml`: exactly four expected cases and +all21 applicable Boolean outcomes (4/5/5/7) are mandatory. Missing, null, +non-Boolean, error or skipped applicable outcomes fail the command. Invalid YAML, +duplicate keys, wrong run/case identities and unexpected assertion names also +fail. The scorer emits seven named records per case; only the configured +nonapplicable assertions may have its precise conditional-skip record. + +This is a result-completeness gate, not grading: `False` remains a complete +outcome, upstream thresholds decide pass/fail, and nonzero upstream scoring exits +are preserved. No summary bytes, scores, rationales or native exits are changed. +Operators must still reconcile judgments/rationales against genuine raw execution +records; a complete summary alone does not prove correct Skill execution. +No task/bug closure follows from static tests, host schema validation or a +successful dependency preflight. + +## Deterministic development checks + +```bash +python3 -m pytest plugins/sdlc-workflow/scripts/ -q +python3 -m pytest plugins/sdlc-workflow/skills/run-evals/scripts/ -q +git diff --check +uvx skillsaw +claude plugin validate plugins/sdlc-workflow +``` + +The fixture/CLI tests use explicitly synthetic process doubles and never execute +an agent. Production source-contract tests establish instruction contracts only. +The fresh local run proves the four paths locally; hosted WIF acceptance remains +pending until the bootstrap is merged and PR299 runs successfully. + +An optional no-inference resolver regression uses actual Fullsend APIs from a +temporary snapshot of pinned source `d5f36921ac754705619f38c637ef692873809fbc`. +It needs Go1.26.5 or newer and already cached module dependencies (downloads and +automatic toolchain installation are disabled). It proves resource resolution and +early environment validation without a generated run-directory variable, confirms +generated source paths and exact prepared bytes, and rejects external profiles +and symlink escapes; it does not launch +Fullsend or a model. Consumer setup/run commands do not require this source or Go. + +```bash +TC6677_FULLSEND_SOURCE=/path/to/read-only/fullsend-clone \ +TC6677_GO_CACHE=/tmp/tc-6677-go-build-cache \ +python3 -m pytest plugins/sdlc-workflow/scripts/test_fullsend_gate_eval.py \ + -q -k actual_pinned_fullsend_resolver +``` diff --git a/docs/tools.md b/docs/tools.md index 448b97a61..bf2d72031 100644 --- a/docs/tools.md +++ b/docs/tools.md @@ -130,3 +130,45 @@ If no Serena instance is available for a repository, skills fall back to Read, G ### Limitations Check the **Code Intelligence** > **Limitations** section in your project's CLAUDE.md for per-instance limitations (e.g., language server features that are not supported). + +--- + +## Fullsend triage-security runner artifacts + +`triage-security` has a file-based Fullsend contract for non-interactive execution. +The trusted runner and sandbox are deliberately separate: the sandbox plans work from +prefetched evidence, and only the trusted runner can execute Jira mutations. + +| Artifact | Responsibility | +|---|---| +| `.fullsend/harness/triage-security.yaml` (repo root) | Defines the branch-independent Fullsend harness and its trusted pre/post phases. | +| `scripts/pre-triage-security.sh` / `scripts/pre_triage_security.py` | Fetch, normalize, and validate trusted Jira, remote-link, configuration, external, matrix/source, metadata, and idempotency evidence into `triage-security-input.json`; the shipped collector sets `authorization.mutation_authorized` to `false`. | +| `schemas/triage-security-input.schema.json` | Requires issue data, remote links, configuration, external evidence, matrix/source evidence, Jira metadata, idempotency context, and authorization. | +| `agents/triage-security.md` | Constrains the sandbox to mounted evidence and directs it to write `agent-result.json` to `FULLSEND_OUTPUT_DIR`. | +| `schemas/triage-security-result.schema.json` | Defines `report-only` and `mutation-authorized` results and the supported action schema. | +| `policies/triage-security.yaml` | Supplies the runner policy used by the harness. | +| `scripts/post-triage-security.sh` | Locates a sandbox result under `FULLSEND_RUN_DIR`, verifies its path and JSON, and invokes the executor. | +| `scripts/execute-triage-security-actions.py` | Independently validates authorization, resolves references, deduplicates retry actions, and performs trusted Jira writes. | + +The sandbox uses the mounted `triage-security-input.json` plus the delivered +`triage-security` skill and schema artifacts. It must not call Jira, GitHub, web, Git, +or other credentialed evidence sources, and it cannot write a matrix. It returns a +schema-valid `agent-result.json` after successful validated analysis, with one of these +modes: + +- `report-only` — exactly `report-only` actions; it reports withheld or blocked work + without mutation. +- `mutation-authorized` — permits only schema-defined actions when a trusted prefetch + implementation sets `authorization.mutation_authorized` to `true`. The shipped + collector does not provision this authorization. + +The result schema supports these action types: `report-only`, `field-edit`, +`status-transition`, `comment`, `link`, `resolve-reference`, and `remediation-task`. +`resolve-reference` and remediation-task references are resolved by the trusted +executor before dependent Jira calls; unresolved placeholders are execution errors. +The executor uses `triage-security:` markers and the prefetched Jira snapshot to make +comments, labels, links, transitions, and remediation tasks idempotent across retries. + +When evidence is missing, malformed, or contradictory, the prefetch/validation path +fails closed before mutation. Operators should repair the trusted bundle or its matrix +and lock-file evidence, not grant the sandbox new access or use an interactive fallback. diff --git a/docs/workflow.md b/docs/workflow.md index 6c27a768f..35e489727 100644 --- a/docs/workflow.md +++ b/docs/workflow.md @@ -337,6 +337,62 @@ Triages a Jira Vulnerability issue (CVE-based, auto-created by PSIRT) with full - Every Jira mutation requires explicit engineer confirmation - No fabricated data — all evidence from actual lock file output or Jira API responses +#### Fullsend non-interactive mode (TC-6201) + +`triage-security` also supports non-interactive execution through the +`.fullsend/harness/triage-security.yaml` Fullsend harness. This mode is selected by the +presence of `FULLSEND_OUTPUT_DIR`; an unset variable preserves the interactive +Jira/web/Git workflow above, while an exported-but-empty value is a fail-closed +configuration error. A sandbox run never falls back to the interactive workflow. + +The branch-independent execution flow is split across trust boundaries: + +1. The trusted pre-script (`scripts/pre-triage-security.sh` and + `scripts/pre_triage_security.py`) gathers Jira, remote-link, configuration, + external, matrix, source, metadata, idempotency, and authorization evidence. + It validates and mounts `triage-security-input.json` for the sandbox. The shipped + collector sets `authorization.mutation_authorized` to `false`, so its default + harness output is report-only. +2. The sandbox reads only that bundle and the delivered skill/schema artifacts. It + has no Jira, GitHub, web, credentialed source, or direct mutation access; it + writes `agent-result.json` to `FULLSEND_OUTPUT_DIR`. +3. The trusted post-script (`scripts/post-triage-security.sh`) selects the result, + validates its JSON/path, and passes it with the trusted input to + `scripts/execute-triage-security-actions.py`. Only this trusted executor can + perform authorized Jira actions. + +The input bundle must satisfy `schemas/triage-security-input.schema.json`. Missing, +malformed, or contradictory required evidence fails before mutation. The result must +satisfy `schemas/triage-security-result.schema.json` and use either `report-only` or +`mutation-authorized` mode. A report-only result contains only a `report-only` +action. Mutation-authorized output additionally requires trusted +`authorization.mutation_authorized: true` before the executor can act. Enabling +mutation-authorized execution requires a trusted prefetch implementation that supplies +that authorization; the shipped collector does not do so. + +Supported result actions are `report-only`, `field-edit`, `status-transition`, +`comment`, `link`, `resolve-reference`, and `remediation-task`. The executor resolves +generated references before dependent Jira calls; an unresolved placeholder fails +execution. Stable `triage-security:` markers and the prefetched Jira snapshot prevent +duplicate comments, labels, links, transitions, and remediation tasks on retries. A +remediation task receives its description digest before its dependent links or +follow-up comments are executed. + +Fullsend does not refresh, repair, or write a security matrix in the sandbox. The +trusted runner must supply matrix and lock-file evidence; absent or malformed required +evidence fails closed, and incomplete or inconsistent supplied evidence is reported as +a blocked report-only outcome. The existing interactive path retains its matrix-refresh +and engineer-confirmation behavior. + +**Fullsend troubleshooting:** + +| Symptom | Operator response | +|---|---| +| Invalid or incomplete `triage-security-input.json` | Fix the trusted prefetch evidence and rerun; do not attempt an interactive fallback from the sandbox. | +| Stale matrix or missing lock-file evidence | Refresh or correct evidence on the trusted runner, then rerun. The sandbox neither repairs nor writes matrix data. | +| Mutation plan is withheld | Inspect the report-only result and trusted authorization; `mutation-authorized` output requires `authorization.mutation_authorized: true`. | +| Post-execution failure | Correct the trusted-runner failure and retry. Stable markers and the prefetched snapshot suppress duplicate writes and re-register existing remediation references safely. | + --- ### Report Phase diff --git a/evals/fullsend/dependencies.json b/evals/fullsend/dependencies.json new file mode 100644 index 000000000..f25b327be --- /dev/null +++ b/evals/fullsend/dependencies.json @@ -0,0 +1,19 @@ +{ + "python": "3.12", + "harness": { + "repository": "https://github.com/opendatahub-io/agent-eval-harness.git", + "commit": "4b540c652f5ed325e18abf6b4bd0eb4414a4bb3c", + "version": "1.22.0" + }, + "fullsend": { + "version": "0.43.0", + "source_commit": "d5f36921ac754705619f38c637ef692873809fbc", + "archives": { + "darwin-amd64": "87ecbec25518ec04273baca648c52433fefa6e3012b83b488008541c23fddf5b", + "darwin-arm64": "71e9d07c45c5d20e30c9da3b7c85c0dba387be55afe50bde1fbb0a8c58b62e66", + "linux-amd64": "e56be72bb2af7210418307e784d65a8ead4ff736540c077f75d258865e595f10", + "linux-arm64": "30b7a2a62556196c7f579dd695d0c30242aaaffabd6c5ba464b41a6fd97add78" + } + }, + "openshell": "0.0.116" +} diff --git a/evals/fullsend/requirements.in b/evals/fullsend/requirements.in new file mode 100644 index 000000000..102433227 --- /dev/null +++ b/evals/fullsend/requirements.in @@ -0,0 +1,10 @@ +# Agent-eval-harness1.22.0 runtime/Vertex extra and isolated build tools. +# Source itself is verified by immutable commit in dependencies.json. +pyyaml>=6.0 +jinja2>=3.1 +jsonschema>=4 +truststore>=0.9,<1.0 +anthropic[vertex]>=0.40 +setuptools>=68.0 +wheel +pip diff --git a/evals/fullsend/requirements.lock b/evals/fullsend/requirements.lock new file mode 100644 index 000000000..28714259c --- /dev/null +++ b/evals/fullsend/requirements.lock @@ -0,0 +1,1082 @@ +# This file was autogenerated by uv via the following command: +# uv pip compile evals/fullsend/requirements.in --constraint evals/fullsend/requirements.lock --generate-hashes --python-version 3.12 --output-file evals/fullsend/requirements.lock --cache-dir /private/tmp/tc6726-uv-cache +annotated-types==0.8.0 \ + --hash=sha256:13b2beaad985e05e2d6407ee4c4f35590b11f8d693a258a561055cac8f64cab7 \ + --hash=sha256:f072f4d804ea359e4eaf198b1af7a8b0943881a87f31bb764f8bf219bb9419e0 + # via + # -c evals/fullsend/requirements.lock + # pydantic +anthropic==1.11.0 \ + --hash=sha256:3906fabac7ad7b5b46c6186040398fc7826885c77ce34e4dd7849de16fc8d0f8 \ + --hash=sha256:52f97b2c485cca7ac66058374f5073a3febb7e6849b16989c145d602a3efee21 + # via + # -c evals/fullsend/requirements.lock + # -r evals/fullsend/requirements.in +anyio==4.15.1 \ + --hash=sha256:6152fdbbf9a77fdec97731721bebf7c4c44f7c29b424b0065826173efc7ed101 \ + --hash=sha256:9f28306018cbd6d329e64a36d58256edff76dd996fe423bc957326e578b82a94 + # via + # -c evals/fullsend/requirements.lock + # anthropic + # httpx2 +attrs==26.1.0 \ + --hash=sha256:c647aa4a12dfbad9333ca4e71fe62ddc36f4e63b2d260a37a8b83d2f043ac309 \ + --hash=sha256:d03ceb89cb322a8fd706d4fb91940737b6642aa36998fe130a9bc96c985eff32 + # via + # -c evals/fullsend/requirements.lock + # jsonschema + # referencing +certifi==2026.7.22 \ + --hash=sha256:62f22742b58a1a33014a2b6b706588a8d7e2a88ae7bd1a6ebe8c992928483775 \ + --hash=sha256:741e2c3b351ddf169a738da9f2c048608ff7f2c5cc02f1ebc6b118bb090d5d55 + # via + # -c evals/fullsend/requirements.lock + # requests +cffi==2.1.1 \ + --hash=sha256:046bfc24911b37851ee1b51aab8bffe713d89c68c6a057b09484ce9fd5f69b4e \ + --hash=sha256:06c72bb76605a4b0cd0aad6930b69d4baf7dd5d806cfc409b824191099700e66 \ + --hash=sha256:0beceaabe56af686895136a2de78db54ecd8e4046b236b8fd6d6cb61389e9bf2 \ + --hash=sha256:154852545011f779917b11c78db2358d095da62a9a172b78ad0a583ee5adc0d0 \ + --hash=sha256:194cffa889098ced9976c3fc6340305e43f6303657d298da55366907c05c22d6 \ + --hash=sha256:19ee6127ee34de7d83ce3d371ebc5ed91addbdcc39f9ab15ce4eb35a4e534971 \ + --hash=sha256:1a18a57b58cfb21fc28d72e876acf10eaed67a1ed96226f92af4df681d571c4c \ + --hash=sha256:1aa5645c30469b09530c4ebca77ebf8f17618293c58f8549cb1a543a50236e7d \ + --hash=sha256:1dea0e4d7d4f11f619fe8c1d76caf49e24405b4b5743c0e3be16a500ecd930c9 \ + --hash=sha256:208f941bb9d18e768138677f0a6d2ce01f590df56043dda1df1535ac57c88517 \ + --hash=sha256:210019b6c7cf07f081b4c54635c8cf744377001350e29cc0f81c4377b4797735 \ + --hash=sha256:246fa40ce8645a614ff682e0b70f37134e460eaf93a775e0cbe3cca585a67a80 \ + --hash=sha256:25792eac27877609e7bb06d42ff88278a6624fff2ba9bbb523c09616b117e80f \ + --hash=sha256:27350daa11d4f10c540e6e89dada4c54feb7256ad03e9a4dc075ebad7ba360d1 \ + --hash=sha256:28907ab9bfb6aa13184cfc17c6b8e1023c5ab6fd7076d8c20a35e59fe04f8f29 \ + --hash=sha256:2ae64be792b8966f2c69538199728b290e34726562896df1e5dc8ffd8d8188e8 \ + --hash=sha256:31348097ff5bbe827ccc41795d4dd099d9f0625e7def00ee653c137a490c2a6c \ + --hash=sha256:3143d81e29e1e20a9ce10901ec369012947876596f75a222235965f2b7ae832e \ + --hash=sha256:3222ba5d678f80a030e6afbcc33dc1ae5cb45facabb61cee2c7016b8432fde48 \ + --hash=sha256:3311ed60d36f83378794e1009ac6258bafbf81f7888b4caa7b35a521e3f95813 \ + --hash=sha256:334644fbac4eff73d985a17a91226df55d0f394160c4cfb880e084c8f7161cac \ + --hash=sha256:34e261f78cb6ceaaa36f42f2613f4380d94d9c759a9c73c769ee6e0247364632 \ + --hash=sha256:363e05fa78e15116c3c32c210ee36884fd6b9afa6d440e47112c3bd511d64cb6 \ + --hash=sha256:398aff33cee2767e3e781d2554c54bd0dff386bb437581e0d8011fde1a942ec1 \ + --hash=sha256:3d22a20b1fb1632cc72c22f95f7b0d2961c3e1c235f245ba4c606c4771035659 \ + --hash=sha256:42a494cee34437f05546455144f2b5d9ac09b1face62bcfce597d2e521066688 \ + --hash=sha256:42e2f76b9455f5a9a844f770bf3e200ed3da0e15f5df3db9c31fe80b04b3d004 \ + --hash=sha256:42f6930c31dc7f50732c9ae793c2786c7b6b044195967bbdde40bb9be81c4cc0 \ + --hash=sha256:456a61fa52d579ebf9df2e9552ead5129855dbaff6c1e5a9b1bc408809bdc062 \ + --hash=sha256:471cee653ae88de62096552e6d24ccb4a5adb8c8c9f10b5054d0122c15bf2779 \ + --hash=sha256:49cbc70e6542d4ccccb936558d1064a8012541e78f821f955cff24e357776c94 \ + --hash=sha256:4a7c934f7360e8cd64fe9efadcbd10c7c6364f531e432b9a4bf5ccbc9e0e8b50 \ + --hash=sha256:4be96343e422f2dfcd12ab5c9f5aebe03f82f737c6bffeca6830b3875cb44aab \ + --hash=sha256:4f42141fc14250de6dde5ee7ea4432be017252d91f19c5ad043c084cea629cac \ + --hash=sha256:507a24c282e0f42f8ed737cf048572cbf580468da5555764a8331735e9c736b6 \ + --hash=sha256:51b31d1c98274844cfd7838ce00bfc27c7423a4dc00fc0772fc3331c2cc90676 \ + --hash=sha256:58acb8ab8e295e6c5ea12f888cbb13cf21511ef2a3303a23f4325c29d17fe5c1 \ + --hash=sha256:5a59cc1c4442bc3d5c703bf720b51138d0bfc173618807c9ee2490a7541dd3d9 \ + --hash=sha256:5bb4e7ea95dcd6a014a6fef62e62467d67d8e582326443f3d68e71d6320a9fcf \ + --hash=sha256:5c58fe613dc5e5336357eff555824a314d8e43282600435c8d1cb6a7a2fedd13 \ + --hash=sha256:5e7cecbaadb83884793e05828cee59b210b24583b9c7425d0ba6a754fe22eb4e \ + --hash=sha256:616f097f2fe415bc92a247f02e11f634e1f9e9a83d327e3c915c15089c87869e \ + --hash=sha256:63bbfd5ded17c4840ac07cd8f1c21ba9d9708141f840b324f422f41b207e3973 \ + --hash=sha256:64faea20f4e2613363a1a9b9c7dd73058f3ecd00133a511e72ad7c511658f527 \ + --hash=sha256:661c298b4821edebead0c91edd2b00374d67ad7c5a1f7a91d4442633b79d6a72 \ + --hash=sha256:68e62fe11f30d5ca8289242866f0a5291402d8529ca2178ab8afc5c9694ae890 \ + --hash=sha256:6a8dddef476fab96d066d578fc88526767b836ab5ab21754e1d5bf3879c31c7c \ + --hash=sha256:6e192623c49c94421616a5778fba35cf0d5a8d000650c1967ef4448ee5cdd990 \ + --hash=sha256:7225e4514edb64eb6740324353e0da0711954fd8d7da4576755b1c6e09b697cd \ + --hash=sha256:75f80557d1389eddbd0de2681f6a390a0c5338c31ddaa821381c203fc3fd50d9 \ + --hash=sha256:770de9db11e84213beec501cfcaa013b019820ca881e03344dea5844f7876d94 \ + --hash=sha256:7750c6449dff7864bb9bb27ddfb0267756189201a3afc911d82b3caacd70dfc3 \ + --hash=sha256:7bde5e4cc5c10140859842b9d383af292b22639a4dffb725314baf45968cef80 \ + --hash=sha256:7ce713ace7c0e4520535b42b77eaa742c16dab813978064913e5a3cf82973b41 \ + --hash=sha256:7da0c5eff80f0197f3b3d1232ec5a682a9325f4ae9016a78f5f5ca35f9ced1f5 \ + --hash=sha256:7dbb61fe3a7699468030f71bbe5f8a0e326a151daa91beb11a6fc1f980c55e1c \ + --hash=sha256:811bd1e21d32de12efca32393a0ab3f5133b54fce9bd44b8bd77ab07da14bf6a \ + --hash=sha256:8ef53b2de9bcb9197d31854256575d59dbac0cba72ac627bb291ef5eceb74be4 \ + --hash=sha256:937c0052c05a31ca1daf18de3158eed4dbfcb9cc107adbea227728d647be701e \ + --hash=sha256:9d2055050ea716bd38b7f7f1579c275386646b4894c155a3e2f3cd62ed41b7c6 \ + --hash=sha256:9f8d177621de5cb38ee3e731eda45d421db093ec0739f46a5594babda7987a98 \ + --hash=sha256:a2d7755bef5a12ed488f4ef1f1b69ee9191d7396083b755a5d2295f6edb4768b \ + --hash=sha256:a48d62ab9d6f4f98c983223a547af44be6ca3691074c31cecced6facd3ba2dc1 \ + --hash=sha256:a4f00aa42f75d6e4595e8866e748cc1705adc0cddfeb2ca86d0d03993d63ba03 \ + --hash=sha256:a6e721d4b0e45d5b65e87534470e67b18dcd092c83f68fba09f152b9cbc061af \ + --hash=sha256:a730a083190634c65cca36ba5f489531576ebd79bcd5c8e172130f6453127231 \ + --hash=sha256:a931079504ecc49efed7744c476a5c343a92fabf66dec2db95edb1b2fdc770e2 \ + --hash=sha256:aa9511c62d14da7aacc9b4bf51f3f697a621e83b2d6919008243c3aad168eea3 \ + --hash=sha256:ab36d55f9ed2d067327667c2fea18dda018eb628dd6347aa01dda6cf1f5d3836 \ + --hash=sha256:ad2c86c495b899d862ea0f4b42891b8713a3bd45dd4105c7fd51c2a72f39f3a5 \ + --hash=sha256:aeae0e330c9f6acd681f647d46cefd30c29f93e3392882e792e82080c9691399 \ + --hash=sha256:b0431303acaea1089ad4b3e9ce4e6518193def1118d4073ca848635ee4ea2e96 \ + --hash=sha256:b5bdfd1c873d4e093aabc0ca84c4ca6dbc4f752afb5c86f146d9742580c9da2e \ + --hash=sha256:baed1e86cc735622097354b9d1281406caf42ff42a886d29faa8e8d1630333be \ + --hash=sha256:c1453022f490d2459a11819d83ad1d586e9ff65a12ac3e705ffebd46d3685dcf \ + --hash=sha256:c26608d2222fb1e94487e4a387d85f13eb55d5ed725cb25a0c589ac4ee60e7bc \ + --hash=sha256:c7659f22557c5a0bc4855cd635f55edec690cc008a40768527762cb9fb263455 \ + --hash=sha256:c8c69575568085ba0b1b10c0249d779a214aea6f6522e949a0fc9fb0fcb449d0 \ + --hash=sha256:c8d2c9fd1f2d16f780d15127abb050d13d1a76c03a4bd87d7e4980e45e511e12 \ + --hash=sha256:ca82be1a1d406ecfe1d25dc16cb33488e5a16bf4438c9fb590484ea29d92478b \ + --hash=sha256:cc572dace3f60ef98d7b12ff411d20f5362feb31a0439eab0085bbfd349982d7 \ + --hash=sha256:d18e5ac0f2f03f4f518d3e23db0f0cad7faa1da8620e9c09461d443bbf6e6692 \ + --hash=sha256:d28630f5854ab07ab1fd4aba756de52326c82e6be15d414b12793f1975048b54 \ + --hash=sha256:d9c275eaacd24aa73f94ffd6de08fc3f932424d8b6c376f4bed7cde376fe7bc3 \ + --hash=sha256:da0e573f9f97159390c89d9f1a9e41908b66d408cc5b58d08cf3847d844c531b \ + --hash=sha256:dd31f52ea1086513bb9df30f8fcee9b8918323ae067a3d5b78bc826a000712be \ + --hash=sha256:dddad92b554513a31f272570678ba307fb9f618f05e3d4a5eacafff9eae03e1d \ + --hash=sha256:df423d40ee8654634421812bc3b196da3f9bd7d32929da813f8394c4348a5358 \ + --hash=sha256:df913725b79db7bcf03448f36b7bf8815363417d5b58deecf9305e3e30f0f21a \ + --hash=sha256:e0bcb7e0f677f543555d2adff3bf19c05f66cdb4796e5ff602442ab2fe3c4ef7 \ + --hash=sha256:e2d65b31f36619cda3999b78b2aa9632e76b78448e7a56fc4240824200e7c4fc \ + --hash=sha256:e6e8cff14d6fb0be70a09c0bdc58096f501952d04624ebf867e0e56da2df8960 \ + --hash=sha256:f16c709686a78c727bbbf059f92b0bf41c6fc60deec706d2dc19f529175a6125 \ + --hash=sha256:f24fb43132a4c6b4cb4eb029492919b2db645be6808d738f244fd146c03c32cb \ + --hash=sha256:f53e442b08449d42821fa4a4fba000095af9f62742a500f978a9f557ec44339a \ + --hash=sha256:f5cfbc5fe74540d335175b656c725d74d90e3730c626d92575eea35029d9afaa \ + --hash=sha256:f81b3b8f3d4e343550fa4baa0e479bba9f2d29ce9c2e9b51d1ce1718d7442fcf \ + --hash=sha256:f8ec5e643a9a937f64e1999eb9f75d072263751912dc5cd06d3c85f8f44be7c3 \ + --hash=sha256:fb92203a88b3d3053034db775110081c49d28be6551923805e039924093761e4 \ + --hash=sha256:fcd22650c908d7b7da162bbfaab594a1227a15d1643a98c68b122ac642fa2264 + # via + # -c evals/fullsend/requirements.lock + # cryptography +charset-normalizer==3.5.2 \ + --hash=sha256:01077390b03f7988f11d700a2194e69b119741a86b1a638b1db88891e3eced8e \ + --hash=sha256:01b0c0d2262a9e28e8484a278c7e1b5d650e3ac8cf2683d2967e25899f208bdf \ + --hash=sha256:04851f73ae72b8413dddadb16a49dfee95263553741fd42d546f7d66907e6be5 \ + --hash=sha256:0521c5665880b33d603717defa76c094048900010897909952397feb3039da56 \ + --hash=sha256:0774bf9bf620249fee3e0b8b9fd3065de213be30f3aa94ce2494b3b638949e26 \ + --hash=sha256:0891b9d3903c5571c03771ca669a4b0ec5618ca722a5c957d3d29cd4e5062848 \ + --hash=sha256:0c951d5e6dd9c2ff60609476752bee49da4206adde960ebc247766937f72e718 \ + --hash=sha256:0fed1d06615f022ee3b13caf5e8b180cfea32bb2c5aded8a9d44277afc040f93 \ + --hash=sha256:114e4d0c92d618409ed82a99e22b5c5e768fe995f2973f78265f4524f49d4640 \ + --hash=sha256:11912e4bb14baae7c5d8791aa55ba0a3a03ec6729073307b0f57270abaa713d3 \ + --hash=sha256:11a4d68a6ecda3292cb1e50239e111543ba5d709bb62a6b4ea1afcfa729d8875 \ + --hash=sha256:124fbf1a8ff966d87ae05bb8bd45a71f966055ed8bba320d0c7cf450bc5f4d0e \ + --hash=sha256:1461ac396c4fdb983a675f20aa555624f0ee18ac83d832b9244ffff3d8055275 \ + --hash=sha256:1503bccbeb36d5527790c3930327704c39af22de3112f1b1666a9f3ce15ee204 \ + --hash=sha256:15bb4005af6320d259dc7593ca84a38d7fe06a421dbcf7b910ae23979101e787 \ + --hash=sha256:15c44f7edfd477b06f517a5cc317fc1707edb9de2c865f43d4b6513907473234 \ + --hash=sha256:16fa0eccf81304b79c5cd87f9271c3b85dd9dd99245e4422ae9c0dd45e0f99d3 \ + --hash=sha256:183b88127acdb4fabe59d951ab424faf1af7b63cdbb5f776186c1ea2ffcaed98 \ + --hash=sha256:195c26fb65950f8fce54e26349852b7bdd7c5f120aeefbcc440b8a20faaed4a3 \ + --hash=sha256:1afb975bd5d68d5ce9f6b6d44fdf2f7e34b895a35e95708a7a91b20a3b51d187 \ + --hash=sha256:1b4cbc7c3491ccb4aa17fcd8165649d01cf39f76de1696da8631b5f71b85401d \ + --hash=sha256:1bc0baf5ef96b6ede57d47f4b8fe4d9d84019c3bfcbeb20a41edc6a6ee341f1f \ + --hash=sha256:1c50fe28bbc2ced33386f298650d91218076c05420e6cbd790b913adc41659e7 \ + --hash=sha256:1db38f4c5496827c1a501846d64d14c3b80c7e6714e406cd7dc36a9899fa1011 \ + --hash=sha256:211d5a3eb6af8f513b8d4ca19a8c1b7accab1b5f0d3175f9826b03c1a920dc1f \ + --hash=sha256:23851fb4e1b85ed3f6c2a27b777cdfe2e19fb5b38429a8faf38c7542b7665869 \ + --hash=sha256:254eb48b9fa5ee9898a3c445825a1f340fe53712a098904b39b0bddba8ea3cb1 \ + --hash=sha256:2625388c6c754520c37abaf3b41eb34d1cc4a373f457898f08606c8e362b891d \ + --hash=sha256:281cb91036248400f4cc957495cccd44c275c2e0c5854f7e45ac5cf7dc193847 \ + --hash=sha256:28a15fdad492a99b6eccfaaed66ef3f74050680545ea61ec8b2f4c538f1f1320 \ + --hash=sha256:28b4f0d66fb834ff90f28209ac7bce77868c45d8c93e26f906709d9b7c2e1af9 \ + --hash=sha256:2a925889534b3748302dae5dead07cc13480de1dac3aea80a941b729b471ef93 \ + --hash=sha256:2b7b3bbfb4fe8ef40600792d762fbaa9057559f9d3fad209525b7a22b99e91fd \ + --hash=sha256:2c9ad19a6cfcd5ea5c0d41161d22f9df1dcc277e9bef2751391334546a314c00 \ + --hash=sha256:2cc961b171b3f3440f410489ab3573e86aea8736134ebbb40ea1338b7f0831bc \ + --hash=sha256:2ce45c6627b22c47e390bc91a41c3d13032192e699fa0bea96e9671b373d69b0 \ + --hash=sha256:2e06a3a98f916dd41d27f3105e02e7a40181c98c94b9158733d03a6f80506c09 \ + --hash=sha256:304d5463e65a35d7bb0850550e0780395395f6fcf452f04db7d5ca7cecc425ac \ + --hash=sha256:304d8e4d493af723536393eee0c689eb7813f4a474c8b479dee63f1fdd98f621 \ + --hash=sha256:30fcd120b732aa79317f08dee04d7de0847822e4cf7ee0e9f445bb958832252c \ + --hash=sha256:31f3930700408d211f13378ccbe1c40845d8da54bd0681fac3a9b5aae81c7aa8 \ + --hash=sha256:34276fd796040bf0993ab33a369aa572e6979c7aab225a88893667ad8eac8f7a \ + --hash=sha256:355ad8011081dec5412240c087a9a0c9d4d5039f3ed11a3f13e18c2b29b56c51 \ + --hash=sha256:38a873987f3be698494da8b2e3085e29da02da7b633dce73e79c699a113d7bf0 \ + --hash=sha256:39de2a259fc954455c57274dc94c79d5842774e1247a016aff30bc0efed0f4ef \ + --hash=sha256:3d14b50de6bf4d0edf857a9386836846f982b8f524e188e2e68b96d702bcf4aa \ + --hash=sha256:3d21b8b13c7592db2ac5e544a6d83187b995257472b0c9e8351b6d507ae37ed6 \ + --hash=sha256:3d31298449090ab8d47b7b1b2a555ff73cac7ed438a08b7ac160980c7ebed649 \ + --hash=sha256:3ddacd27458c45bdacd6bd6db644bfb730efbf9e830310186e3045c9c5be8fb2 \ + --hash=sha256:3df041de8887954562c9b261cba85ca0e9ded74048daf125f45edcfaa4832229 \ + --hash=sha256:40ab6bffa02ae10a0581e6c198be7d2d8ca5c2a0c64e4ed3465d766df457573e \ + --hash=sha256:4275811936e2f06feff5e598fb42a1b7ae852da8e39605211892b56b81a34efd \ + --hash=sha256:443eae2bf318abeaf6f15d785138f71fd6de770e99a92158b8b814265e079115 \ + --hash=sha256:447441e76ec720b15e64418d32e092297340387053047c7c694f579efb0ee1d9 \ + --hash=sha256:4495c5002a7b28557e7e222e77e0b661183e432b7d6d2e788101e3f240e05b8c \ + --hash=sha256:44bd4fbb29dfbeba60e7d2bd000c59e4b21ddb3cc53912b14048d37092706d7c \ + --hash=sha256:4685902cf26edf013ed7a3da0f426ebba7a00ebb9541386d835afbf002c11cab \ + --hash=sha256:498dc3188ca05a68231ac3fdbfc7f57eb67e1343c30e0fea17f8218c1599b253 \ + --hash=sha256:4c2b5031f63e331e3839b40aed2dd6f191e9c07edbde303e7876846ea1946995 \ + --hash=sha256:4d48f2d08b9de5864e2c8744d4461b862fb149a18274abc8b698c45975573438 \ + --hash=sha256:4f87960d57feabfb618e4e0af6e7371645fa26a277860739d6e5d6e0012c92f0 \ + --hash=sha256:50e3adfb96fc189eb27b1cf62d3b598b89b4bb0420d93a3d3e42e137409011be \ + --hash=sha256:51cf45226a9b588d0d2b4880c62d686934b63ab0bd79ca23ab0e9762eb27441b \ + --hash=sha256:52aa6992700996af31f375de0c6bacd402b0097fe40b53c426b9f51a90ebabc7 \ + --hash=sha256:55ea99acb17b9325618de155a0cd6a2e8f5d10be008113e1d433bbb58db543b2 \ + --hash=sha256:56bc200a365efb37383b7852e4cc5898d3b2da5987289b543956cf8cad71018a \ + --hash=sha256:588461c2e8384d309bd63e5826019b6977bc66d629b99ac8737bb795d7b2cb5a \ + --hash=sha256:58ca3755ee7ff7f59b57789ec9833c9de9ea275405cdd240eda1f193112e398a \ + --hash=sha256:58f361dcbab699cf8f42db3f47c8e7fd1036f138c23a5d08de9fde5f425a730c \ + --hash=sha256:598a11a2c7ebaa5334bf698bf29568c9c390abac6a154d8170fedecd1cea38c5 \ + --hash=sha256:59f63901b0031c3136cf64704dcb21de0bbae62ce2c9529bc39d27665463de37 \ + --hash=sha256:5cde776b7cc66e4f6c99612cea4aa7269aa65863f7a15841b2c264f103822f4e \ + --hash=sha256:5e2b6b57e9733d39f0c9fd3185efa6b8e29652c4cd8fe94180272cf6ed9a78c4 \ + --hash=sha256:5fb29fb8cd1a46c27a1bf9613ad5ec2599310d46b4025d9556404a6b6a292800 \ + --hash=sha256:6045373d5a89a5ec71afde535db987ca28e76dfa276c2d4c818265b375d4b055 \ + --hash=sha256:619799369eeef6366ed3e8755a5670f4f2f0fb6b30a0fd7264dc0fdc2357058e \ + --hash=sha256:62588a277bfb59def052abd940703fa35107152bf479781a878617d60faf8fb5 \ + --hash=sha256:62603db9a7caa0802eaa28c1c46fecd7b3a263a774069c24c3c28c302448721c \ + --hash=sha256:65cd72beeeca9d3aaea1201e5923859f308f952f9c71de93f06063c79f0f7a3b \ + --hash=sha256:68eb192d85ab8e5f6ec69c2bc6ac0179fbf04a5ac1569d12fbef74883fe102d0 \ + --hash=sha256:6bd128f206a7752ae1f2ab6c61bf8a24ba28913a10df8b14c2637b973ff97a80 \ + --hash=sha256:6be488a102b8cf28d0391d8c4ba7748938ae28b78ad901f8585520fca33ead1a \ + --hash=sha256:7218e8f32b0956cfcd048fd42d9d5779809745ca1d86113ca56f66e7ae1549c4 \ + --hash=sha256:7441d755b7ab94f8d4eb3e43ec05482d760842fd263d003a99102d742cd835e2 \ + --hash=sha256:749e97e1b32313717a565abbe321bc2190bc8b35f1a67e4cdbc7c56c8d8ffe58 \ + --hash=sha256:75a3ceed0724d625d64b86ca20aba182e4df462e04c2414fc941c0f523f06aac \ + --hash=sha256:780fbe7cab297b81dad9fb8dc5eb003c0468ffb0d9e5f65068c53a34661a96bc \ + --hash=sha256:78456a747de8dc58360ffa581f30a002baf5aa28cb262536545e91f113ed7639 \ + --hash=sha256:7967d08cf06dee78443b874f98c98036f624f3a4e73e11f9f64f5be4d25393cf \ + --hash=sha256:7a881931aa470808df94a8c380eed2bbbc76cd9dc622310f99665658c821eb6d \ + --hash=sha256:7dcd882da75ef9adf94903b1e3b9419e8aa8fb4c7396822b834b9ef7fb96954f \ + --hash=sha256:7e841fb9010836c992c9f12fcbd43a831de93a5f726fc1ccd8ca1d0268c5014c \ + --hash=sha256:7fdde2c9fd9e3eca40631e024664cf2584272cc8f96308cbe5fdfc930f51d8bc \ + --hash=sha256:8024d00c3faf3fc0c16e07a69f4405e8eac7cc0ab15f65fe6cf43827c4cf72b4 \ + --hash=sha256:80d02b6f04e92601a081dd97b23d3128033098bff5d35d392ddcc0476ea11253 \ + --hash=sha256:838dcc90063569a0448120554591a1d6c4a4ffe11babf048908793154ab86ade \ + --hash=sha256:849df64e889b2e17230d58410a03dba311a65b163508fd33679b2b737d4b7858 \ + --hash=sha256:87475fabc8d9996fd9c27debb395e642e8c838d78a00b6e932227a0e06b81e26 \ + --hash=sha256:87e50a3e7cb90af586b6c5faf23e302a970415ac73bd7bd90a515a04b427ef96 \ + --hash=sha256:89b53f3cda69831909888e0494f4fa0bcd3537e3e138dabeb620bd6ad946bae8 \ + --hash=sha256:8a893cc101149f80a653f82062ebc95b34525a2614382e1da5458fe7c6997249 \ + --hash=sha256:8b2bfab86aa71ae13aa41a6a26aab338e0db2b8bc75434b05aea89e011ff35a4 \ + --hash=sha256:8d86d6fc60743dc916eb79e2eb1ec4818e21e427731543af40a3021851174a13 \ + --hash=sha256:915563965d418f986e7e145accc592eae9e1a1be3566ff98a05d7a9ec42a76e1 \ + --hash=sha256:92888bb3187c5ba50500b00b3b310c9f2c651709d28036077680cb5255450a03 \ + --hash=sha256:93223adc95033dd47133a46ccfc316a0139176fd79085762e27202ec56018f03 \ + --hash=sha256:9373ad13ef0d2c0fb761e04e55bfdee5a08b52cef2c882c8fbe9935b1517152e \ + --hash=sha256:9409a8bf35cf78353942504b24a57de3d75b708997a1e4bd8db71ac8633ce364 \ + --hash=sha256:9b7f416ff0978e2f2249330527f0ad6fa02f4932e6199692d3b52da2048c19e4 \ + --hash=sha256:9bde855991b7e362c146535e3136a50bfaffc0487d38b33ca7e5edefc6e23849 \ + --hash=sha256:9cae88599c7219005d879f98e5ed53341e9a122af585e1091200358a3003d2a0 \ + --hash=sha256:9cf9b1a857e25c4baceeb3624e92a56df3668f398c4acba74e174d81fb4d1d3a \ + --hash=sha256:9f56f72050826f63dcee7a7f55b0a77168cb3bfc553fd405e7f8f9ece75a4036 \ + --hash=sha256:a090bb2c68df85450502e3e20d665e3a5af9c65a84d6508ed477badd49166fd3 \ + --hash=sha256:a192e2c40070d92c3ccf777e3a5c4ff515573cd2bb7ed0c537fdadbbec5bbf21 \ + --hash=sha256:a19a731138fc27d5682277d3b9df22855cea1239bce7fcec5f78f42ef2d1f3c3 \ + --hash=sha256:a66c3bc5ab1f0ff2164fc9965ddd611ff0802173f4b9d24554c563f6ab7e1d6e \ + --hash=sha256:a815775b6c38d4e0ff7bcffbeba67feded90202bb6a226b8dd35f1c855217413 \ + --hash=sha256:a89012d6d5476ee112d20d998570ed58df2260a852afb1758809cd6900411d21 \ + --hash=sha256:ae4f5fea5b8b8ccff88238cc8569303e5ee95efae67fa62922a311397a71f346 \ + --hash=sha256:b6856554c4f44d79fc2307d5768854310a8f0096e501c75637542c82292b0429 \ + --hash=sha256:b6b751274acb69d77b3323d6b7dbaa3c7fdfc1eb829b7eb61d262f32e1af9685 \ + --hash=sha256:b736353c0a625bbd5fcec108576e2385db3496f4f771f785ff32e108d3c3bc45 \ + --hash=sha256:b7fd005a73d9e657273b7a10dc71a9e03c8fb9ee6999798d6918ce095b81ac7f \ + --hash=sha256:b91363207bd9dc966a691e959bb47f64b30f7ac4b072be9968b366982f7db77c \ + --hash=sha256:ba0b1d2620edf869789c3879223f52bf2afc5d31b3cb47cc57b3a12c05e2aa9d \ + --hash=sha256:bbbfc8e28816f19d7c0f1816664980c0a9875d01b27cdf8eedddb639d9e108ad \ + --hash=sha256:bd16aabe4a02a297c23417aa17ac6299dbd8c49f673bcd645b4929b11f5a4400 \ + --hash=sha256:c0afc6800ba57ccc350374c5bd6150419915d95ce93cdbab2d783d75eaf30ecb \ + --hash=sha256:c6708715abcf3c73b99508253e961a9967f02fe536532834149574eda6de0d1c \ + --hash=sha256:c7c9ab723cde841fefb34efbad91e87f00a674b1fe1cd0784fde742bf2c154dc \ + --hash=sha256:c8f3d67aeaf55f017982b73683f0e7342ba2f6635a78f69ce89ebb26aa411e5c \ + --hash=sha256:c9790464842f85f437dbbb54417eda1e0e6bfc52dd8d22d6fd1c994b73b2dc74 \ + --hash=sha256:ca403d7e4798f525fdfc78e258820419cbbd0f0ecbab9de7840e3c017cf6b8cf \ + --hash=sha256:d008d90a7f2471519aef0c90dfbe73b3e6e4d5e66ac48e19154c17e89e98b604 \ + --hash=sha256:d19fbd981a488e22cd04883659ca6b08f50b5974f9fd7c95655ef6a043e5893f \ + --hash=sha256:d1befeed746d247c81127bb14de9dc3d30edb6e5976d34f83f86ed262b1d9105 \ + --hash=sha256:d2374b62878abb00cd8309b32af6c0b715cd02dec0ca74ef12e5069bdc64144a \ + --hash=sha256:d376bbd28b3a8999db1a103b3b388aee6f1ddeb3e51bc2172993efdcd86e064d \ + --hash=sha256:d4a7319f304a774bed22115bc891618e45f85065ab44ea6acd07d274e750519a \ + --hash=sha256:d6734d2ef8a50fbf8445c139477da401f50d62a0606bf00e20ec6d87773fefb1 \ + --hash=sha256:d760fe2a4d7c3b226cb9026d6a842868d52a7901bd98420e1baf14e80da85cf5 \ + --hash=sha256:d913de495d90407cd859d263bee2e5d1a4ed3eb6573c04e70d9ec619a7cbed7f \ + --hash=sha256:db19d07e2e0129e974a0e65d0064fc222a446cd5122c2fd4184d2af9fc734a9e \ + --hash=sha256:dca9ab98072a5a54ebacebdc45f53e645336b320c667410b061be1ca588ae709 \ + --hash=sha256:ddc7dacc8ece3a182e7f15cb862d1fd616b46d076cb1ae9dd232b2c38b655874 \ + --hash=sha256:ddf19c062bea7a0cc80f519243d2c01dd091be0cf952a0750d4ad576709559f5 \ + --hash=sha256:def79fa35ef0cef8d2accec024f4fdc7ead3012ff02f5215c783f39f03ef8cfc \ + --hash=sha256:df29a0a7107f7011e77f4eebdddec4c7331e24d787a0b21a46d63bdf7445da95 \ + --hash=sha256:e09a3942ecbdee5cce73ea9d42da82b81b72ac1bf031ce069b93b5adf4eac8cd \ + --hash=sha256:e242bb1c5e76e97dfa9e7f209a71e93a01d7f19ffdd5cfbb2e2d55b4f08f8ab0 \ + --hash=sha256:e243bd13217235fc7290c621941c3f5cc8b66e4872495be821d7436ba2fb838d \ + --hash=sha256:e2af3aad578aa6bd1384bcf4750fc285e5a9de53f40b7d41e5a0bf748edeb2b3 \ + --hash=sha256:e4e81e09c1578b8df602e3db08b0b3ea0a6947ad612f52bf8dc5ea8d47691f0c \ + --hash=sha256:e54da4baf05720032d527874d40b65fa4d7e5c6c6a43d0c3adbeffcaf275a2b3 \ + --hash=sha256:e80e6c2f55656b4824d72065abb4ddd6a525c74bd78a0aab5d9fc2cf4fb5af50 \ + --hash=sha256:ed2a239c0ea213acc1908150a3037257083c7c083128f1a4cec2ec4b97dca491 \ + --hash=sha256:ed905975ab14056a2e5eb1c376cb2e1ebc5396baf84163939c518556fccde9f5 \ + --hash=sha256:ee21e28f0430bd6dc9086c6e525d5e818a44a5ad19720c8a0ef766792f3eb5e5 \ + --hash=sha256:ee43c17b173d46a3212baa6ead3ae258eeabdae48c263a01ccf0218c366dd655 \ + --hash=sha256:ef4fcbf3327382cd4c9f540babd61248208af7b93eec4de397b4d5f58a09e288 \ + --hash=sha256:eff0ac9dbe711a4aee69bf04a83896aa9b85f19641264053a9f6d48573abb7dd \ + --hash=sha256:f0aa869112ef88429ae17820d99c3dd9504c9e9c671d3c246f3d7442cb051084 \ + --hash=sha256:f3c96f633825733f735c5a9cf21d21a257d8e1edf0b1cee0a064b9c424ca0f7d \ + --hash=sha256:f5833ad231be5eb6553de524a70f48d71b2c8563101750531e0b80184e175cd4 \ + --hash=sha256:f5ec61164adcec446f8969a3358ec3f9b26bbda3b9213e5586d219afa8df2915 \ + --hash=sha256:f7d486c83842422badd511868fd8a9a20e9407ace71564b6af47ce7e60a336c1 \ + --hash=sha256:fb9e68df06293761f9fe66ade60a9bc6d0f5e42b8acf2939a9158af86ab0e5bd \ + --hash=sha256:fc14a032f813bf5fe624d991960ea83e9715adc27e4c1830a2361eb1d02ac341 \ + --hash=sha256:fcff63213e8e6e47770541a4607175404f47cbb3ebea7b6058cc82d524a0e424 \ + --hash=sha256:fd1fbe0f116b6e55da77aca2c6ddcddcfac2186cbf78bdebf40fc156efca389d \ + --hash=sha256:fe9753dfee015c570d73df76f899f18444d41388bffcde097deba51c4fadbb9f + # via + # -c evals/fullsend/requirements.lock + # requests +cryptography==50.0.2 \ + --hash=sha256:0ddc924c04591c2811ca024d62ecad4f7f6f08af8939c211438f48a16bd23602 \ + --hash=sha256:0ec5f09541743261e66e291b4a0cbf0fb2997aeaab6d9e9c740b9dba1b58d1c2 \ + --hash=sha256:0ecbc5652bdb6fc9eaf89a7d196e20941adfe812f43bc4ca05d9150496821047 \ + --hash=sha256:1981f1db4630889b9ef7803fadef12b056f428cb6b85c27ba57b774793b6093c \ + --hash=sha256:1ba34f04897fcdaa73f74145c25f3ec146fbd56593853e88adc2e811303c5f42 \ + --hash=sha256:241449bf940a5d27309bd317e6f9a2af6932113818bb2b8f5c59ddc7ef16da18 \ + --hash=sha256:25784ce8b9621c90c643efb9e1e2162ab3b0224cae446ad5e70e7fcb1ce18b51 \ + --hash=sha256:3dc4fd8058cea1644971207d530e1a03a184a805ffc8ebdddf0599d78a331b81 \ + --hash=sha256:4061c0079120205fb760c58acab6443e217307dcf05e3702cf970e0689972856 \ + --hash=sha256:4a20ce1e5cb4284a86692fdcba7cb8754185c6b2e5c56fcef3751cf451d3cdc2 \ + --hash=sha256:4e81d95e5bafc2d6e34e4bed780e53e4d5b9a2f928573428aa4d35fbec1eb0de \ + --hash=sha256:58a0c478eeca76fe5e07993c5a0703def34a6dc6a0cda4f5564639b33112ffe7 \ + --hash=sha256:58ddb5a8e3179d12f19e4ea34d2d32e9d63a4baa142c875c1eb59f41b7243acd \ + --hash=sha256:630ebfea3bf689d075f82316324ff7433dc447fe6bc1bfc76524b74b4a9567d2 \ + --hash=sha256:6f8700550aa1474a91e5dc07049c46f98b423b5b1ddd0483e0b51362eeeaf5be \ + --hash=sha256:78198641e5be9521beea5aa782bb551a58068d10e6eb04c9c680c1b69f2e7d45 \ + --hash=sha256:79def8d059362e7831389ed3be0ecdf58a89386e1271e35dd9f5af84e81bffd0 \ + --hash=sha256:7a8701d6b584d76e909e3d305b7d126b41439876a5aaf76cddc67fc230eafa2e \ + --hash=sha256:7afa5a6602a9f29af1f3a2965f831bae7c9d5d597b7cbb716d41ab3b7d89879c \ + --hash=sha256:7b46165bb56eb4704e2eaaf86f3c940d19154535d9b0ca7d6d590b04060e00d5 \ + --hash=sha256:7b75de3c8b3be1cdb1052747c929440c3eea46c1bc2cb8a6e3a48388e9b7b452 \ + --hash=sha256:7c6d0330c472d96f6a6afe24d80dfdf15176c33096f0a4397ae4c60f3dd3be48 \ + --hash=sha256:828d49b0ff5a0e3975865571c5d91dbbdd0d38d8289b249a163e9425413a5e05 \ + --hash=sha256:84f964e537f916e2cc85199e5a88742e964939b575ac8598b3f9d6cc416cdaf1 \ + --hash=sha256:85d0d9a31b9098e98534226d5686b47264b95e62ce459dc2e62fdfc809f9fe93 \ + --hash=sha256:87e9ce85beb6b328ba370cc6e6aea483c92617b4c95b1d33a49297eb662bfb04 \ + --hash=sha256:8c71ba2cd31fc93748c38e1b613200ff1c2665cbfd5341fe3a61cfde35a1430e \ + --hash=sha256:92e665960f25fcdc73725b9cec7a3824f279ba97a98653afe9ffac2e43668f67 \ + --hash=sha256:94e5e9f108ee10471288214d3d233fbfbb492840a8457eb85178d643ddeb32c7 \ + --hash=sha256:9c8402a82ea0dc4ceeab793db05f0fafa8ca139ca34fcde5df0f596103c74107 \ + --hash=sha256:9dab55f57c74c3cad24c323bacbbd04be4705ba6eb0d92e920b1fc4837ed5079 \ + --hash=sha256:a582ab2ae1d34f67112cadc86702774c9ea4374df6bca6afe672817203c99134 \ + --hash=sha256:a6557e5f38e065ca9fbdaf7cfc7435ecb1d113aa81a022d1b51921ee7432e227 \ + --hash=sha256:a9f7355e6fab51f6c369b86fb7571cffa05edee2c2121e0380a37fb9ac1cd5c1 \ + --hash=sha256:ab50ee449bf968271e820086f10a33d101dd060370abc10bcd22279be2656539 \ + --hash=sha256:ac9ed99d81760c62fe89d5f0815cdfa1ba9a35141cf30f1c2d044f04b4803d2e \ + --hash=sha256:b13478603dcd0a2479ff8e87e2c19a7d525734686fe3c49542472293a204212d \ + --hash=sha256:c423ab384a46c4dff7217b2ea5ba2e11cffdeab6441acd04cf65a369caf0366c \ + --hash=sha256:c5e67125c7dca78d199ec4e116aa93dbb83494808ecbb8211a2cb09b1bf41dbd \ + --hash=sha256:c71be1cbfa5cd9a41ee452acf1eccd82b2c05950358b106ec8ceb83411d1a020 \ + --hash=sha256:cbc8738fd8526d80f35cb3a40d41f41a2e7030bb3b18b09a6778ef63d291c2fd \ + --hash=sha256:ce47f66801c20ec6c6632453bb5960fe38939e9306970b48b3a5a26de7745d94 \ + --hash=sha256:d370b8d1dfcdf7130178137f6fbee6140774a1acc6cacefc4b42643ec11d0a3a \ + --hash=sha256:d38cdff612d06fa6a32840d5e1b1f7a27cee4a349aa9085d94a67789d6bfd408 \ + --hash=sha256:d8947001be83df1394050758ce0e745dd74fb134eef0a4b5124208dfc3a68c37 \ + --hash=sha256:deb9fde5c60e437ee4821bc9bc39ff31b42135c27e1dc61ef0a629389c1de62e \ + --hash=sha256:dfe9763530994147d9af1def057a5b9658b00e8f8fe8743d144d1e0911c2e454 \ + --hash=sha256:e105ab60406787da31fccc883fc0f733af1efd78f0136a4599692c4083a73d0c \ + --hash=sha256:e275096ea1e60cc595cda2836fd4a6c725d1125108b868be17f53684d164e2cc \ + --hash=sha256:edc3342adf8f697fc5f59c887a304356f147b397809440ed64e2fa6af2f50f37 \ + --hash=sha256:ee247f5c245c9a2fe7c8e2214e295918838e44e00a45a6718451e4004219e767 \ + --hash=sha256:eef4c2f3423810b3070ab391f85436d2f8bbfcb286ac15cbc73190b3563b1f1a \ + --hash=sha256:f21e8a22c8605750c7af886bab299a363721264061b4ac0a30efb73cfd58efc5 \ + --hash=sha256:f265528741e048bce55c3463ed721fb0aa45a5888d8add8cfeccb3035451bbdc \ + --hash=sha256:f2f9bd7f90c64fe89253f0a2c05e3c4856072660429ce8831b4235bf29403a67 \ + --hash=sha256:f785f6161f202ab04d8ca194158968798e480ca058943907972da5f12e2881e8 \ + --hash=sha256:f9f6143a8c75945eb960d9eb98905a441394abfa24afaae239d514ffb2586480 \ + --hash=sha256:fa8f5efb344d6908a1ce62f4a24e2e5780f825d6f53f5f50ec5ffacac72936cb \ + --hash=sha256:fdd28f912fccfec1846a94e2e1e8f9b0012f557f0c46fe4f3eb0d7a87afcf90b + # via + # -c evals/fullsend/requirements.lock + # google-auth +docstring-parser==0.18.0 \ + --hash=sha256:292510982205c12b1248696f44959db3cdd1740237a968ea1e2e7a900eeb2015 \ + --hash=sha256:b3fcbed555c47d8479be0796ef7e19c2670d428d72e96da63f3a40122860374b + # via + # -c evals/fullsend/requirements.lock + # anthropic +google-auth==2.59.1 \ + --hash=sha256:89c3f931683a482ac97e61df7eb9da5e08a91703f3c752b2377d72cfb7d69e6a \ + --hash=sha256:ce50fc533ac02f489a2b183a0c156672c376ecb2091b1127bc7efba2975fff27 + # via + # -c evals/fullsend/requirements.lock + # anthropic +h11==0.16.0 \ + --hash=sha256:4e35b956cf45792e4caa5885e69fba00bdbc6ffafbfa020300e549b208ee5ff1 \ + --hash=sha256:63cf8bbe7522de3bf65932fda1d9c2772064ffb3dae62d55932da54b31cb6c86 + # via + # -c evals/fullsend/requirements.lock + # httpcore2 +httpcore2==2.13.1 \ + --hash=sha256:e0aa977abe17e69a3b820a24542a6fa88702676d83880b8d194dcd18408e5103 \ + --hash=sha256:e1e05d4f25f7d7d496bfb96748f6f4b67657b03da069b3a68c36069f3db73d0a + # via + # -c evals/fullsend/requirements.lock + # httpx2 +httpx2==2.13.1 \ + --hash=sha256:6dff50fabc270ee5fd25d845d0b078ed20564579744d6d962850975996d2f9a4 \ + --hash=sha256:e48744a19e3af5ee48313d0ce5fe941d5422fae5705ea922a4aabf94d7800dfa + # via + # -c evals/fullsend/requirements.lock + # anthropic +idna==3.20 \ + --hash=sha256:a7db850025b95ded1eae8a46181a1a6c56c92c96f0e2b005d9ff8dc0210cab44 \ + --hash=sha256:ab7ae7122974553370f0bdb919e1a960b2cd1bc1ef0276416d896db81c14582c + # via + # -c evals/fullsend/requirements.lock + # anyio + # httpx2 + # requests +jinja2==3.1.6 \ + --hash=sha256:0137fb05990d35f1275a587e9aee6d56da821fc83491a0fb838183be43f66d6d \ + --hash=sha256:85ece4451f492d0c13c5dd7c13a64681a86afae63a5f347908daf103ce6d2f67 + # via + # -c evals/fullsend/requirements.lock + # -r evals/fullsend/requirements.in +jiter==0.17.0 \ + --hash=sha256:00b5a98df3e3a3e8cf7b619f4ac2f8bf975bbf3d95d02c5d17b8dbfe5c8b8245 \ + --hash=sha256:00d783a779c5664e16dbad5e3a3c3a75e128b07dd5f4765159658d9210a50ca5 \ + --hash=sha256:0239520085cac678e77a606fd7e3f1c60c371d719790c5e3807388d3da4354c2 \ + --hash=sha256:02a360707033d8cef53f7f3480817a1489177a259ec6ec01e98c37e0b922ddca \ + --hash=sha256:02adebb7ce6413c44d40af9ad59d1c1cd79630ccdcb6f7bdd2d461e48c03d8f9 \ + --hash=sha256:03e432f226a453851079fb84cd17c6da9991eab723e28d716f14ae3d906e0c12 \ + --hash=sha256:0619d806e260ecf0c2a64521942c94af5d547c9ec99b55ae4f51b538b5576a76 \ + --hash=sha256:073dc68c1a700c8fc480e877864a6b6ffc887533e261f4380c08c16bf09d057a \ + --hash=sha256:0b52d52035b3907c5b1f6277857b29c1cbfc965e24e0f27330dbed83edb591ec \ + --hash=sha256:10c5349312e5cb02b7a21e123a57665afa895953f05bf252a9dd4c13a572b7ab \ + --hash=sha256:10cd64a5720ad7f809ac5466ff1705813f1b6b510f195a73acafba0ac0e1f675 \ + --hash=sha256:10f5558eed511b830488003449d942bd75829ad6257dc58cb9a03e596a7777b1 \ + --hash=sha256:11902505d401691720f5785c15b02204248526edee11b635cd6c40cd52b81599 \ + --hash=sha256:155be7355bdb7ca76ab0961be8982c225f964a5c073a83984183f22391cc29fc \ + --hash=sha256:16dd0c1baf098ae70b8f3616574eb3fedf34e26670b89e16a7e67561f737ed2d \ + --hash=sha256:1b18434638228c0c184281609bf3d9459026a0f1ea48fb76c205e3ef72069caa \ + --hash=sha256:29f49b325e0234e4ad9ecca5b861ffbd09b95ccac9bd46fa55841b6e56eea5fe \ + --hash=sha256:2c45ad7c973ef33fe5114a953377b35a95240f4542c0724d9f781e47dc24bac7 \ + --hash=sha256:300ce01ab0215e3dea4d00090143c909aedc65c0f809b3c07983e1d038f291b9 \ + --hash=sha256:30793a24a31e968969757c9e08d830cbb15a2cd3c4959b4498b38f4b1c2258eb \ + --hash=sha256:30c692d567ba206c7cca38c9d1d0ccc70c9786290173c184d871ca12e9981ed7 \ + --hash=sha256:32aaaa764604496610a3ad2d98503ae88ccb2fbe769e892ff4533e778e85f708 \ + --hash=sha256:362bb47423886d45a9f705d2d9d4008c6eedd4e41eb1bab4e96fb6daa06b33fd \ + --hash=sha256:36ee6e69027396664e59995b9a635a947a5304ee9837279584a0bb8145c8f6b8 \ + --hash=sha256:370d8fe5bf201dc6925e8a84c81ac7291f74d9fd1778234fc79d517064a5c76b \ + --hash=sha256:37150a9e02e869475854fa20b7d0d5e26d18d0f8bc17293999973ff27e99ae7a \ + --hash=sha256:37f33d327900bf2879613b3363fd48df97b4232d0c41f54bcf2e790c2fc40a71 \ + --hash=sha256:3ad556afc289f15d2b181b941982d01f06190863c07440185b9f354e1bd2def3 \ + --hash=sha256:3bf4dc2b84a464117fb097d15a25c58d100d2692888e3b0d92df5b48ed16b7c0 \ + --hash=sha256:3c1a5336c04a41b1f1cf9572e294aec27cc569767ff73de7bf87a91f0bea7cb9 \ + --hash=sha256:3e05f5adbf68c4bd11e1610f394034d984152988e84be6f8314235ce6f2139e5 \ + --hash=sha256:40d2c240f8f80b5b0f201b29f0ae129c81448c60c772227a41747b5e0026f6a2 \ + --hash=sha256:42b0260445251b1bc520a63baa94a32d88e0f931fba234f1764db7feb7c72174 \ + --hash=sha256:454c4997d73cc466c71fd565d91e603b0274e48ea0c6b0b7a7aee6967e4ceb7c \ + --hash=sha256:455e4ab35cb2a4a91a8404e08fd3c621bae433922e59bf1c494fe20a426b013b \ + --hash=sha256:4607ec7d93355fbc25b8dc5189153cf21d66063b9f9cd04dd2774e6e783f9b6a \ + --hash=sha256:470e1b1e4c42f1ead2189166a299691871a2df5056c976e7fb96feafaf5f9d44 \ + --hash=sha256:492f37230bbf9581ab2c17bcda862c249afb9ae2e3ab2dd6db59943bc4cc3153 \ + --hash=sha256:4dfbfe5a6e1e80a7082af559f66386405025ec278833e0c649f69cbc6e1004cc \ + --hash=sha256:4e3f052c671d5f425cca5ea5901cf11a831369fba4a55a3862cab93c323b4c3b \ + --hash=sha256:5078ab00664307fab2019b522a93aeb191122789f085daf5fd9e362154021d4a \ + --hash=sha256:51e1519d676a9f14dad9c2a411170d43b022ddb7989562df4e849b261ce127b2 \ + --hash=sha256:523c499235fb65add25d4bb01b1c4709ce695efdc7deb6c0a7bc515b5c44e0fb \ + --hash=sha256:545c36a0f3b2238c242cc9785439d3242a871b7bc39fe3f441bcaa07bf3aa83e \ + --hash=sha256:55d0e0e613a3f9ad600cf436e0e2b8057d1b52bcf1d91b2d36ac53451231e6a8 \ + --hash=sha256:5888fe5abc1ca2fa834a3e1b4c7ef0dcece286a7d7e95a609ef0934b777b9fc9 \ + --hash=sha256:58df29268a95e910f17db7ec9178eb7f15aa8619aaca3575275c4e6b3f4fe4c5 \ + --hash=sha256:59bddbe6f9ffecc68d641e1e2d619ce64cf8a9e9eeb74e5c518f74fc87abf1b0 \ + --hash=sha256:5a52a430d04225ffde633e6840bf2381d34c019ff98526b5929755b9052fb199 \ + --hash=sha256:5bf350452a43173e69e1fc74847c57a60e3d7515807287f29849baa2a85d8718 \ + --hash=sha256:5c23849235d2142ce444b2b8c6eceee9f82f4cc0bd5c9081602e4155c6197807 \ + --hash=sha256:61aed66ee042b3b49ef85fdf75714234d055d89d8496ac1c6e47f89e7a30d5e4 \ + --hash=sha256:6219adaf59711ba7063a52496e8ec6d3fa3e209d7827d83eee3b2abc780a1744 \ + --hash=sha256:64846211a2debe7c071d2146d2283d2b0c1c93dc8fd5fb7794faac2ca6061b5c \ + --hash=sha256:686c93d86f2b426c803024b805bd161a6cd10e9627c23e901640eab646c0ad8a \ + --hash=sha256:6871973bfbd4408f7f1c632b30bbb5bbd9671c1bc8650af6823e24b7be13709b \ + --hash=sha256:6af5b74073bd25bae695e6d00919f6a9be7ed5a9f8836d981eb1ffe84139e6fb \ + --hash=sha256:6b303d88e6a0bda789ec4b7801c7bad68e27230ba1fe4baffc756d1fbd32dc9d \ + --hash=sha256:6cb41cd1432f1dc19a231cf70b54d42b2c9f05085155859263fce06fa4d41388 \ + --hash=sha256:6cf564d43c4388149ca58ee571d0f5ccf875e20d1fd4662fd94cc0d1ea3b10ef \ + --hash=sha256:6eb6aedeb7352b8f3b6af9cbd67983840165c00428e63f1b420a85885128ea31 \ + --hash=sha256:70f19a2ca8429f91e82eeffb2f51cb87bc2d6e953b009b91a92d29c3a16ccb03 \ + --hash=sha256:71dbd74314c5df52a1bccf7b8bca46d14e943af7a2012e73b23f49977ef194c8 \ + --hash=sha256:73b64e69c4150748e020356d958af94bec33c70a0a93d665cfa8f6d580fe1a63 \ + --hash=sha256:746243a080b4ca790b8499af3d7cf9825d5f5987933950cd818e767ee353d826 \ + --hash=sha256:755079792868ce5d4938e83b91a0939b34fb858a1ca65a104f2d771bea57faa1 \ + --hash=sha256:7573e80232c5bcf80c24c038cf7e53a463f5c3b1dd1dd4109d66304f4dccc233 \ + --hash=sha256:76eb4a5c20e86f9f848286f167024890f2862258a965d254774deb7fc1545ca1 \ + --hash=sha256:77f6aac0137309b31448c1bdcda4c6c77077664a6d018ece8d94019c68a5a5b9 \ + --hash=sha256:785a216bbaf8f15fc974e964ced7322cd3d774bb0e86949edd78c6bffd6ba35b \ + --hash=sha256:7b68d3495d95da120651a5628c7ebadee84ed001a1b76e6afc325c42482f15b5 \ + --hash=sha256:8079849db9a1371bfd90bad088458a8fb836261879df2233cc9632464ecf64e1 \ + --hash=sha256:81c83c0abe614446a283d994d2c07c4f58632dea2cdf66ba9e2921bb8ccd593e \ + --hash=sha256:826871c42cebaae22f0a2b5673a4a1a75c851bb2d13b3c17764a630a6b298984 \ + --hash=sha256:84963d3f395ef5e9a32ce47155e08a7962fa292c159a10cb98b931cef1416925 \ + --hash=sha256:84ac78df457e1ee3f7e733bd114823302ae8c5ad5542d7e6647d92ffaa090a04 \ + --hash=sha256:86d703d9faa1ffc8ae4e9de0fa007712ed2171b5c0d93811a8e2e105ac729b0d \ + --hash=sha256:86f3f9343a288eb85a81ef20a752b2f84564296636db54a9fff0b5c8deaf1df2 \ + --hash=sha256:8adca2e793288e5f1bb29279bb439d0d3cfbb50eddca7e7e6ffd42ff4f482406 \ + --hash=sha256:8c21265b251d99bbb40080d178a8953e35601d3a1564e05c4de4c0d2ca616797 \ + --hash=sha256:8c286860abfe8b100cac1c02e225e5776eb9216edd71ba17cdb237da4af32bc9 \ + --hash=sha256:8f770b0c77e5fac482e1ba03ca1a7e18286bfb213d749932a00a7e4cd5de5e06 \ + --hash=sha256:93946d89fa04d5ba64dd323a8dd8d901676cb8a3c81d99ae4f6c051a9b4c3f2f \ + --hash=sha256:96b8b0c6dc5d78682f54a450785e075aa929cde768304cad363cd4efba5a82ac \ + --hash=sha256:9bd3caac219df476dd0cc3fe01d2f1581ed588906feac767abd9614c1c12f8b3 \ + --hash=sha256:a277f97eba7d66b1ee27eb5dab5b774ff46a10c78d89a1d3dcce04ce1357c8ca \ + --hash=sha256:a3cebb1fe4a1abb00465f3f8a17e09112603e8b7c59e5c3adbcd9f7815a64acd \ + --hash=sha256:ac3c6ee3264d6f5c44c617f90bc7e8b9e1587e7d6708c9d8f811cb65582ee312 \ + --hash=sha256:af2f7501580f274b63c4b2283bc425f5df7edf06ae5b171e5f87d912ff359a20 \ + --hash=sha256:b550585523339b71cb852b811aae49d08d7601ad8ffe9f5dc1562f4c3d22fd87 \ + --hash=sha256:b75f85660108965a94be77911a25a253429307294d9415b3c597118977a614de \ + --hash=sha256:b847b18d066c46b3b7ae49d6c94a7634c5e4a8983146ee25562a092000f5e3ad \ + --hash=sha256:bcc064f99183a9cbe7f26ed648c352031a74145cd61ed75d34632c73eb46a5a8 \ + --hash=sha256:c19b9357309b8cc6de8a48fca8e44a8c9c2feaaa2f5896d037fa505d48fcab80 \ + --hash=sha256:c4289293e5278d9314b00f15c37f2120fa51d3d68565292e715524c750e775a9 \ + --hash=sha256:cfafd7be8b16ceadd298db542cead37cddc211c4c49e04ad2596924df18625b1 \ + --hash=sha256:d0ce4feb52493e3513335b2accdcd75605652e4632772d3c8c2f7b86954d7f39 \ + --hash=sha256:d2c0bf24c72fd0491405dce5d40194f2070e9021ce648c1a1d46234b93d848ff \ + --hash=sha256:d47687806f9c54c84ea38733507081337922beca90ce819c7d852dd485bc0f23 \ + --hash=sha256:d85c558c9f8532bba287a990ac63767c7daf756f0d8c030219f62499b1fa228a \ + --hash=sha256:da139721f4b7cafdbff580a4f511ea24cb91f4909330c6b926a1ca53836c0a59 \ + --hash=sha256:dbbfe4e3c21c8166980cddc5bee1a315df082454f007947dfb6fb73800768165 \ + --hash=sha256:dc0288ce39190ee33fe6e4ec73161eed34e7e2da509b525546ca061778d62b64 \ + --hash=sha256:e088612ff90ebc9247e1a43074b72835804261c47e6a6c01cb3ddcb55360d688 \ + --hash=sha256:e654b6b04e39c9cb19cb8b04c6ddf1f2db07751fa14156413969fd78bad0e5cb \ + --hash=sha256:eaba834b72d573547b9d966465b3394b749d5e14208cc70acb63aca37619ab33 \ + --hash=sha256:eae86b1f027031e39db2e0e9c4842221edb7b8cd474d23f87a79b3bd4b651768 \ + --hash=sha256:eb2295da7c3769f6719b227a237aa6a5cfa6550e478bc838001b592c57e16575 \ + --hash=sha256:ebf918dfd6a74adc1b9ad71f63c4ab00902fcd3b7fd39f2e24d871db8d713b91 \ + --hash=sha256:ec89771f4272b989487a6364e519db6bbaba323e8bbf949ac89a45ea9c18b7a3 \ + --hash=sha256:ed1a24005daac667d577402d75a2922f9775a165b146b883ff1ad3602d8be689 \ + --hash=sha256:efe9f61bb30174d2f5c8396445c360c96c44e78164d0815dfe627ccf57849574 \ + --hash=sha256:f0bc7f684b65bcda9c20434267577db71bf9905ceddd32b60d1d93278d8c8d3a \ + --hash=sha256:f3d7f7b34114f7ddc6d72a8e882d49de636b35d9fd12b4d420d3c5729f6c9812 \ + --hash=sha256:f753eb70b1474a29e635e7542ff7312e6d6b951e0b25e8a2e8c34eeb1ddcd478 \ + --hash=sha256:fa13acf1046f95df808c64b1310705e143fab87aee73ae00cc42d640867fd2c1 \ + --hash=sha256:fd7790aa79c8b518e512ebcdfce9f11d8ef5f30efd43720c8a19a548b39fa489 \ + --hash=sha256:fe15ddf316f1f1f643347d3a474e74ce61880c79a11ec5dca53df20c071bd3e8 \ + --hash=sha256:ffa0380ad091de7d3fc33e17a97ff479851ee18a0a2a3ee56ff3215cdc886656 + # via + # -c evals/fullsend/requirements.lock + # anthropic +jsonschema==4.26.0 \ + --hash=sha256:0c26707e2efad8aa1bfc5b7ce170f3fccc2e4918ff85989ba9ffa9facb2be326 \ + --hash=sha256:d489f15263b8d200f8387e64b4c3a75f06629559fb73deb8fdfb525f2dab50ce + # via + # -c evals/fullsend/requirements.lock + # -r evals/fullsend/requirements.in +jsonschema-specifications==2025.9.1 \ + --hash=sha256:98802fee3a11ee76ecaca44429fda8a41bff98b00a0f2838151b113f210cc6fe \ + --hash=sha256:b540987f239e745613c7a9176f3edb72b832a4ac465cf02712288397832b5e8d + # via + # -c evals/fullsend/requirements.lock + # jsonschema +markupsafe==3.0.3 \ + --hash=sha256:0303439a41979d9e74d18ff5e2dd8c43ed6c6001fd40e5bf2e43f7bd9bbc523f \ + --hash=sha256:068f375c472b3e7acbe2d5318dea141359e6900156b5b2ba06a30b169086b91a \ + --hash=sha256:0bf2a864d67e76e5c9a34dc26ec616a66b9888e25e7b9460e1c76d3293bd9dbf \ + --hash=sha256:0db14f5dafddbb6d9208827849fad01f1a2609380add406671a26386cdf15a19 \ + --hash=sha256:0eb9ff8191e8498cca014656ae6b8d61f39da5f95b488805da4bb029cccbfbaf \ + --hash=sha256:0f4b68347f8c5eab4a13419215bdfd7f8c9b19f2b25520968adfad23eb0ce60c \ + --hash=sha256:1085e7fbddd3be5f89cc898938f42c0b3c711fdcb37d75221de2666af647c175 \ + --hash=sha256:116bb52f642a37c115f517494ea5feb03889e04df47eeff5b130b1808ce7c219 \ + --hash=sha256:12c63dfb4a98206f045aa9563db46507995f7ef6d83b2f68eda65c307c6829eb \ + --hash=sha256:133a43e73a802c5562be9bbcd03d090aa5a1fe899db609c29e8c8d815c5f6de6 \ + --hash=sha256:1353ef0c1b138e1907ae78e2f6c63ff67501122006b0f9abad68fda5f4ffc6ab \ + --hash=sha256:15d939a21d546304880945ca1ecb8a039db6b4dc49b2c5a400387cdae6a62e26 \ + --hash=sha256:177b5253b2834fe3678cb4a5f0059808258584c559193998be2601324fdeafb1 \ + --hash=sha256:1872df69a4de6aead3491198eaf13810b565bdbeec3ae2dc8780f14458ec73ce \ + --hash=sha256:1b4b79e8ebf6b55351f0d91fe80f893b4743f104bff22e90697db1590e47a218 \ + --hash=sha256:1b52b4fb9df4eb9ae465f8d0c228a00624de2334f216f178a995ccdcf82c4634 \ + --hash=sha256:1ba88449deb3de88bd40044603fafffb7bc2b055d626a330323a9ed736661695 \ + --hash=sha256:1cc7ea17a6824959616c525620e387f6dd30fec8cb44f649e31712db02123dad \ + --hash=sha256:218551f6df4868a8d527e3062d0fb968682fe92054e89978594c28e642c43a73 \ + --hash=sha256:26a5784ded40c9e318cfc2bdb30fe164bdb8665ded9cd64d500a34fb42067b1c \ + --hash=sha256:2713baf880df847f2bece4230d4d094280f4e67b1e813eec43b4c0e144a34ffe \ + --hash=sha256:2a15a08b17dd94c53a1da0438822d70ebcd13f8c3a95abe3a9ef9f11a94830aa \ + --hash=sha256:2f981d352f04553a7171b8e44369f2af4055f888dfb147d55e42d29e29e74559 \ + --hash=sha256:32001d6a8fc98c8cb5c947787c5d08b0a50663d139f1305bac5885d98d9b40fa \ + --hash=sha256:3524b778fe5cfb3452a09d31e7b5adefeea8c5be1d43c4f810ba09f2ceb29d37 \ + --hash=sha256:3537e01efc9d4dccdf77221fb1cb3b8e1a38d5428920e0657ce299b20324d758 \ + --hash=sha256:35add3b638a5d900e807944a078b51922212fb3dedb01633a8defc4b01a3c85f \ + --hash=sha256:38664109c14ffc9e7437e86b4dceb442b0096dfe3541d7864d9cbe1da4cf36c8 \ + --hash=sha256:3a7e8ae81ae39e62a41ec302f972ba6ae23a5c5396c8e60113e9066ef893da0d \ + --hash=sha256:3b562dd9e9ea93f13d53989d23a7e775fdfd1066c33494ff43f5418bc8c58a5c \ + --hash=sha256:457a69a9577064c05a97c41f4e65148652db078a3a509039e64d3467b9e7ef97 \ + --hash=sha256:4bd4cd07944443f5a265608cc6aab442e4f74dff8088b0dfc8238647b8f6ae9a \ + --hash=sha256:4e885a3d1efa2eadc93c894a21770e4bc67899e3543680313b09f139e149ab19 \ + --hash=sha256:4faffd047e07c38848ce017e8725090413cd80cbc23d86e55c587bf979e579c9 \ + --hash=sha256:509fa21c6deb7a7a273d629cf5ec029bc209d1a51178615ddf718f5918992ab9 \ + --hash=sha256:5678211cb9333a6468fb8d8be0305520aa073f50d17f089b5b4b477ea6e67fdc \ + --hash=sha256:591ae9f2a647529ca990bc681daebdd52c8791ff06c2bfa05b65163e28102ef2 \ + --hash=sha256:5a7d5dc5140555cf21a6fefbdbf8723f06fcd2f63ef108f2854de715e4422cb4 \ + --hash=sha256:69c0b73548bc525c8cb9a251cddf1931d1db4d2258e9599c28c07ef3580ef354 \ + --hash=sha256:6b5420a1d9450023228968e7e6a9ce57f65d148ab56d2313fcd589eee96a7a50 \ + --hash=sha256:722695808f4b6457b320fdc131280796bdceb04ab50fe1795cd540799ebe1698 \ + --hash=sha256:729586769a26dbceff69f7a7dbbf59ab6572b99d94576a5592625d5b411576b9 \ + --hash=sha256:77f0643abe7495da77fb436f50f8dab76dbc6e5fd25d39589a0f1fe6548bfa2b \ + --hash=sha256:795e7751525cae078558e679d646ae45574b47ed6e7771863fcc079a6171a0fc \ + --hash=sha256:7be7b61bb172e1ed687f1754f8e7484f1c8019780f6f6b0786e76bb01c2ae115 \ + --hash=sha256:7c3fb7d25180895632e5d3148dbdc29ea38ccb7fd210aa27acbd1201a1902c6e \ + --hash=sha256:7e68f88e5b8799aa49c85cd116c932a1ac15caaa3f5db09087854d218359e485 \ + --hash=sha256:83891d0e9fb81a825d9a6d61e3f07550ca70a076484292a70fde82c4b807286f \ + --hash=sha256:8485f406a96febb5140bfeca44a73e3ce5116b2501ac54fe953e488fb1d03b12 \ + --hash=sha256:8709b08f4a89aa7586de0aadc8da56180242ee0ada3999749b183aa23df95025 \ + --hash=sha256:8f71bc33915be5186016f675cd83a1e08523649b0e33efdb898db577ef5bb009 \ + --hash=sha256:915c04ba3851909ce68ccc2b8e2cd691618c4dc4c4232fb7982bca3f41fd8c3d \ + --hash=sha256:949b8d66bc381ee8b007cd945914c721d9aba8e27f71959d750a46f7c282b20b \ + --hash=sha256:94c6f0bb423f739146aec64595853541634bde58b2135f27f61c1ffd1cd4d16a \ + --hash=sha256:9a1abfdc021a164803f4d485104931fb8f8c1efd55bc6b748d2f5774e78b62c5 \ + --hash=sha256:9b79b7a16f7fedff2495d684f2b59b0457c3b493778c9eed31111be64d58279f \ + --hash=sha256:a320721ab5a1aba0a233739394eb907f8c8da5c98c9181d1161e77a0c8e36f2d \ + --hash=sha256:a4afe79fb3de0b7097d81da19090f4df4f8d3a2b3adaa8764138aac2e44f3af1 \ + --hash=sha256:ad2cf8aa28b8c020ab2fc8287b0f823d0a7d8630784c31e9ee5edea20f406287 \ + --hash=sha256:b8512a91625c9b3da6f127803b166b629725e68af71f8184ae7e7d54686a56d6 \ + --hash=sha256:bc51efed119bc9cfdf792cdeaa4d67e8f6fcccab66ed4bfdd6bde3e59bfcbb2f \ + --hash=sha256:bdc919ead48f234740ad807933cdf545180bfbe9342c2bb451556db2ed958581 \ + --hash=sha256:bdd37121970bfd8be76c5fb069c7751683bdf373db1ed6c010162b2a130248ed \ + --hash=sha256:be8813b57049a7dc738189df53d69395eba14fb99345e0a5994914a3864c8a4b \ + --hash=sha256:c0c0b3ade1c0b13b936d7970b1d37a57acde9199dc2aecc4c336773e1d86049c \ + --hash=sha256:c47a551199eb8eb2121d4f0f15ae0f923d31350ab9280078d1e5f12b249e0026 \ + --hash=sha256:c4ffb7ebf07cfe8931028e3e4c85f0357459a3f9f9490886198848f4fa002ec8 \ + --hash=sha256:ccfcd093f13f0f0b7fdd0f198b90053bf7b2f02a3927a30e63f3ccc9df56b676 \ + --hash=sha256:d2ee202e79d8ed691ceebae8e0486bd9a2cd4794cec4824e1c99b6f5009502f6 \ + --hash=sha256:d53197da72cc091b024dd97249dfc7794d6a56530370992a5e1a08983ad9230e \ + --hash=sha256:d6dd0be5b5b189d31db7cda48b91d7e0a9795f31430b7f271219ab30f1d3ac9d \ + --hash=sha256:d88b440e37a16e651bda4c7c2b930eb586fd15ca7406cb39e211fcff3bf3017d \ + --hash=sha256:de8a88e63464af587c950061a5e6a67d3632e36df62b986892331d4620a35c01 \ + --hash=sha256:df2449253ef108a379b8b5d6b43f4b1a8e81a061d6537becd5582fba5f9196d7 \ + --hash=sha256:e1c1493fb6e50ab01d20a22826e57520f1284df32f2d8601fdd90b6304601419 \ + --hash=sha256:e1cf1972137e83c5d4c136c43ced9ac51d0e124706ee1c8aa8532c1287fa8795 \ + --hash=sha256:e2103a929dfa2fcaf9bb4e7c091983a49c9ac3b19c9061b6d5427dd7d14d81a1 \ + --hash=sha256:e56b7d45a839a697b5eb268c82a71bd8c7f6c94d6fd50c3d577fa39a9f1409f5 \ + --hash=sha256:e8afc3f2ccfa24215f8cb28dcf43f0113ac3c37c2f0f0806d8c70e4228c5cf4d \ + --hash=sha256:e8fc20152abba6b83724d7ff268c249fa196d8259ff481f3b1476383f8f24e42 \ + --hash=sha256:eaa9599de571d72e2daf60164784109f19978b327a3910d3e9de8c97b5b70cfe \ + --hash=sha256:ec15a59cf5af7be74194f7ab02d0f59a62bdcf1a537677ce67a2537c9b87fcda \ + --hash=sha256:f190daf01f13c72eac4efd5c430a8de82489d9cff23c364c3ea822545032993e \ + --hash=sha256:f34c41761022dd093b4b6896d4810782ffbabe30f2d443ff5f083e0cbbb8c737 \ + --hash=sha256:f3e98bb3798ead92273dc0e5fd0f31ade220f59a266ffd8a4f6065e0a3ce0523 \ + --hash=sha256:f42d0984e947b8adf7dd6dde396e720934d12c506ce84eea8476409563607591 \ + --hash=sha256:f71a396b3bf33ecaa1626c255855702aca4d3d9fea5e051b41ac59a9c1c41edc \ + --hash=sha256:f9e130248f4462aaa8e2552d547f36ddadbeaa573879158d721bbd33dfe4743a \ + --hash=sha256:fed51ac40f757d41b7c48425901843666a6677e3e8eb0abcff09e4ba6e664f50 + # via + # -c evals/fullsend/requirements.lock + # jinja2 +packaging==26.3 \ + --hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \ + --hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c + # via + # -c evals/fullsend/requirements.lock + # wheel +pip==26.2.1 \ + --hash=sha256:71138adf1f4ca900cdb7d289c21b7494329f2332b6d85f0e1c42108c0384ed3e \ + --hash=sha256:f6ad667e89a1fe78046c8f13232b247200f5258d7828f3f7883d660878e0813f + # via + # -c evals/fullsend/requirements.lock + # -r evals/fullsend/requirements.in +pyasn1==0.6.4 \ + --hash=sha256:9c447d8431c947fe4c8febc4ed9e760bc29011a5b01e5c74b67025bd9fb8ce81 \ + --hash=sha256:deda9277cfd454080ec40b207fb6df82206a3a2688735233cdcd8d3d565f088b + # via + # -c evals/fullsend/requirements.lock + # pyasn1-modules +pyasn1-modules==0.4.2 \ + --hash=sha256:29253a9207ce32b64c3ac6600edc75368f98473906e8fd1043bd6b5b1de2c14a \ + --hash=sha256:677091de870a80aae844b1ca6134f54652fa2c8c5a52aa396440ac3106e941e6 + # via + # -c evals/fullsend/requirements.lock + # google-auth +pycparser==3.0 \ + --hash=sha256:600f49d217304a5902ac3c37e1281c9fe94e4d0489de643a9504c5cdfdfc6b29 \ + --hash=sha256:b727414169a36b7d524c1c3e31839a521725078d7b2ff038656844266160a992 + # via + # -c evals/fullsend/requirements.lock + # cffi +pydantic==2.13.5 \ + --hash=sha256:346a034f080da3755d8e9cb5e00e8b07de1d39e4f6e2c87d8ab7cafa0b269a73 \ + --hash=sha256:51a9c5f7b2f8e636f04c6cada605d9b6a3bf1348fdf945a3d8869b19bba0ee08 + # via + # -c evals/fullsend/requirements.lock + # anthropic +pydantic-core==2.46.5 \ + --hash=sha256:013d6f3483d81e02e7c328831808f336c8596ee33b4bd4026b9ffb1e960b8942 \ + --hash=sha256:03b9666e41e35d8909852ba191a0607520f81b74eaf12ccf8737005dbb313821 \ + --hash=sha256:045ab3b6d308439e32b81cc173bba5b9018bc6ed896afd0c65b3b009b1699af5 \ + --hash=sha256:0bddb4020d8f04175865ccd17eff3040874fc11fb593f424edb452653b4b947c \ + --hash=sha256:0cdbada856a1c69a7624a64d3d9aefe79300bd6ef827b43a4f265010b9b55184 \ + --hash=sha256:0fc5be0abd4a407e200d844b404e33639a554e7bd0d448e7b9ae181be4789ac2 \ + --hash=sha256:10416c15b8839ecc4ef4d0885da76da6fd0f67333a0eb8aff6d93c4b8f2910fc \ + --hash=sha256:15f4a94963c95accac15b7b657bb177d3ad82bb90b0d0526d9a9b85079925db5 \ + --hash=sha256:18a09e1e1011b462f2e32774f25859ef1223d5c2b0546a633cf56654710721e0 \ + --hash=sha256:193375f3548919d3f0b60936ca113ada3e38f264f91b9b8e0508efaad57be931 \ + --hash=sha256:1a353f84de772f423b5ffb11d7ae352fbbef0f446f3c0b0af0f8236d7233606e \ + --hash=sha256:1e449def1945a462c464331254e5a44fca7c3b4f9aedf59ec2f50f8066dd8e25 \ + --hash=sha256:1e5aad1220a1192c42341c8fd4a8686657e73ab2a920c970bdc4de334fe3193d \ + --hash=sha256:200aa3dc9f8d54f0754f43247c0bad0999fdcfbfd2488384dd44f37279271fe6 \ + --hash=sha256:2471fd51c61c610e1dcf7de44d7299283661654d11264ab4802b303368d69c47 \ + --hash=sha256:24922243639cbdac66c75fcb6fd6495a9cb52b213d62f9a0d16f0310b1ff8038 \ + --hash=sha256:28a6a556cd3b6066bea827857f9d9cce027c96f776e512f544a581f9e42161f8 \ + --hash=sha256:2bc9419666990c06d7397831f2126a1ecc3594aaa3ff7de5bf2d066802f4e07b \ + --hash=sha256:2cbd9a5eff05e51c447c34dfa4632145b26b09120cf04bd0c871e44c1a5e1c9a \ + --hash=sha256:2d330aaba8621b1edcec8ae2c4050f63b84ccf6d98723a8f212e9684713abf0e \ + --hash=sha256:2d5d76654becf5efd62c9e51c3756c67b49498b0c9a40884934c40807adbd074 \ + --hash=sha256:337639ba62a11acde6ef3aeb08c8ea755f8ef1fe5e513356c0f36a2b0d7568b0 \ + --hash=sha256:347ec774390c87326a2e4929d58d3f7e8763a104d5d35f4cd595a4c952366433 \ + --hash=sha256:356c8368cbc321050b169595683a2e1d63413b1e0e2868b330af9fc14c616d3f \ + --hash=sha256:37ae34309d7bd8c0d61ab839668058f2a7962ea1fc51d105d2db228fe0618034 \ + --hash=sha256:37ea7b83c935e5b0d68c9449b82651accf78a10828b2c02b2f2d9e9496446c21 \ + --hash=sha256:3a3e26b6a8274211bddee2d0e4d0d42778f17a34510f49d2ec44b58abfc41736 \ + --hash=sha256:3aa166e99c4f2985407fb8714aebede877ecb5455cf321b606adca926d30d5a0 \ + --hash=sha256:3d2652072b2d774947ba5cf78a9e59644ac62ee572daf6dd2e1dfe905e15b2b7 \ + --hash=sha256:40375c2d05acec10323e45dfe2077ac44bc74659008614af5069034e2cfc781c \ + --hash=sha256:413a717a410d0c817ef5b786a059415550b3794e1d0c2abffd9efb93a3d9f7b4 \ + --hash=sha256:46c25dda9d092a06c08db76ffe0a197107904d0dfac653f7d5306bbcd6d6119c \ + --hash=sha256:49776eab08766a08dfff7012f8b422dcd7e25e43b316eedf0477c24fcfa84b7c \ + --hash=sha256:4d44cf99ddebf875f9b68cc267aa684c99b7b44fe63ee1cac4ec163807290069 \ + --hash=sha256:4dedce55295becb61921e386b99d4f2706045306e7fa52249a33004c837379fb \ + --hash=sha256:4f8507560a9284e1370bb048ed4282012fbef4e8d109875b95e884d228552061 \ + --hash=sha256:4fdc8b93a41521988916eeaa271173fcca7fa0803d62f87675aac8dcec1c8e29 \ + --hash=sha256:5086029a57366b8cf81b130a43908738095c270c21a8d7f0e8bdfdb89718e2f3 \ + --hash=sha256:52e24eacdb536cade636aa90fb851835222becff8484b7001fdc78cb0290f2aa \ + --hash=sha256:53feb344243bb9510a9dec7bf3cf1b64d88a98af5dc7872a5160465f8b198c8e \ + --hash=sha256:545f26c504b27c3758439a5e6d9349931f0a04f855668d5fe323c89e82300a38 \ + --hash=sha256:54d510bac3ee52247af28ed4bb18a1e799f040ac60fd2bf5ccd4c92f1fbe786f \ + --hash=sha256:5cb482e9e84c851f4e623fe4acc1ced89168cf1fe18f7089db4548c8f5bbb65b \ + --hash=sha256:5e81740c09e310f5aa5cbd3e434a01c154d4bef93241c7877b39f211d2b78ba8 \ + --hash=sha256:5ee239d575f80b08eca11f6e20f90c4c695de7825c67eefe6091fbf20dda648e \ + --hash=sha256:5f194189415698233dd1114a093a9b56e61e2c57e11b469be3b0506f46f0771c \ + --hash=sha256:5f93c5fe914d75fbec9a49209b00da5f08e9e467d69da2b1510c81940cfd10be \ + --hash=sha256:657b40d6240c0a7b6a64b30f22d1e3aa631c7e846c621b0c0f6d1d75e2e15ea6 \ + --hash=sha256:6d30e1a4f138b8951063e9a394752a9179b51da288ffa507b1e659222f4c1793 \ + --hash=sha256:6f7b393a8b3da82f5c1fc0751e6d01ac6c55b93c18226a60bdfba4a724efafd1 \ + --hash=sha256:701b2e04b560eeb4bddf7a25ab8ca476176e34fdbd9a0e18196f0d12d4685f0b \ + --hash=sha256:771cf63ae0b1b50dd22e5f3e3549fab5f3f4ff1635d352a9e1a97fe01c7b2e64 \ + --hash=sha256:79bdfa52f843137045b2d081cc05c120ba6665d29b7559c2c47690906f39279f \ + --hash=sha256:7ac031912d54f3d83ef3b3eb98dfabc1608802e2202263d25957eeed40b94761 \ + --hash=sha256:7b0fc826b16c55e561e5d2a0c5c77b051ba1d92808118c4e4b5390f5e0cf191d \ + --hash=sha256:7c6be839a5a8312626b32029a415644a0846b420bc8b52b95b28cd92da162168 \ + --hash=sha256:816ff0a6550ffc06c098ccd2e0698600f9aa7da192a79eaa6f9af504a35db869 \ + --hash=sha256:82a36973cf8a2ef5406f4fe2edbf8ed0c99629535d959e0b100c76a32535a111 \ + --hash=sha256:837b396ca3d7b74091ca623f6cbd8351bd42d670a79c2683e79fb089f06a2de5 \ + --hash=sha256:850a08d167dde16db8702c274f320c7be9d7da6f6dff2b58b18f9e815bd94f5b \ + --hash=sha256:8816f3d218beb4b787de5c9759c259b8fa61f9dec42dc7811f320a33771778b7 \ + --hash=sha256:892a881d5f68c2b9ea304b7a6c2c60d9343df578a311b0f86b94bc8f1ffe8129 \ + --hash=sha256:895395f8918627b04efb1ad2a4cf605387143300ba03304cd1dfa6d03f5e095e \ + --hash=sha256:8b10e3e8fd7ddc2bd915848a2768e44c15b22936f1cc54c462ad1164deb02655 \ + --hash=sha256:8e24d8f05fa2d28513d94e877e9c75ad66175376209b3977f916e240e623193c \ + --hash=sha256:8feeac04b5794e513e710af2f9c87d49f31a6dc47967bb264a1fed61a8989bec \ + --hash=sha256:9432f3598db432cb51c5b37fdbf29a60fcccc79e30d37a05022776a6bc4ab689 \ + --hash=sha256:976e1128455aa595ea04c79ccfedff1aaeab96ee013fcc916bed120c4f0ad94f \ + --hash=sha256:978e7b97d4824b5be09c69fb70507cbde3b0323fc147332ca40a94d9a6a0ebbf \ + --hash=sha256:97bf8de4d541598c94a59344eeb988a94c08ff76b5723c41f6567ec18c7892ea \ + --hash=sha256:97cf3eb53a8cccacf9d46686a0926186c9bfb5574f2ed66d3639d5fe117cd3a9 \ + --hash=sha256:9b68938dd5b0c783d88ff8e2dcc69451b5eb936fe212d516b21b9d5567f6d464 \ + --hash=sha256:9c4b71f10dd532fb7a5cbc8f58707779e64f03a258c2bf8bfbaecfcd9970b519 \ + --hash=sha256:9f47b8a949e60f027f0aa0a6f6c7b7e9c55cbf4380d10b344e282fa4e7ab1e1b \ + --hash=sha256:a1dee1b804ff4d11c663636cf15d2ea47e9f79cd56c033fb1cbf08924842a48f \ + --hash=sha256:a2468d93d181667a7abd66e1b64bb9f76f361b0fef8faddf687456453576f5ee \ + --hash=sha256:a2a5e1d0ff29adddc9f6d6821a66302e4493f8ca898b715b6b1182c2c201ea0a \ + --hash=sha256:a39ac25a9a2fa4072efdb429833c4a4c8009a51ff9eea3eeae131713cd27991e \ + --hash=sha256:a445486499897b88a7d6c310c88ed64dd37b1b59bfd7ae9107490bbb362f47d6 \ + --hash=sha256:a91c17edf6eea2402cb5457b4c89e99bc5ed1004aa34c4adf1d4258c1a5c22c2 \ + --hash=sha256:ab4b66edffb32d9e951efb3814bd104b8367a7501b81b955cacb5726d897389f \ + --hash=sha256:aca6c767f552b21b10f774aeac128e828eafb796adfa1b666a18bf6321453c3a \ + --hash=sha256:acf8a67ba51f4ca9ddbd0e6b3000a65ac51ab734661778b3e7ba64d99a710f2f \ + --hash=sha256:b10ec717381bdbfafef34607824db4c91de69ff085e4fca3b2af91b4fa17e68a \ + --hash=sha256:b49924c73a235e969511bf2aabdff3beebf9820931f646c80274d5d780010c47 \ + --hash=sha256:b6acfb46a814762367fb7ba0828b0a17d441b92ce249a0e007474c9072662dda \ + --hash=sha256:b7ca9034437b6022f941f4857459562ee00a560b97e7cce8a0ec5a74fc6766e0 \ + --hash=sha256:b98134087d9de723658d17a42c7d0da8d6e2ef08015dee7dc93889047315f5e4 \ + --hash=sha256:b9fe6fb92520e3fd61f2e49000b6911b188824f089b75973ea06d6267f0b476d \ + --hash=sha256:bce57638e08ac148e5778cce7feb968307a727d66f8e2274a543d0cf0c9ad6a3 \ + --hash=sha256:c14ad3bdc85ee7f318742c457ca3968a92126d144b15721c759033bfb06296c2 \ + --hash=sha256:c1c43ad4339643d70ebb8124e1305a7dab423001eff58bb41a0f731adbc98355 \ + --hash=sha256:c3471e5c4a949c26ec00a77f01df59096aa9495877de76fd60a980f8ee6be461 \ + --hash=sha256:c583b927a8838dab890706a6fa7573fbb8b70e24000ef9f7238e2d6f6435a5ed \ + --hash=sha256:c76fe65e607be28c7fd4d56fc3c42b1583aa058ce3408b7ad0fd540171d31f9f \ + --hash=sha256:c7ea57fc63aa7da93a1bd2d644e6577befae10c52c4e36377635eea1056a74f5 \ + --hash=sha256:cd5214352ae68f3b5e9af7768bdc5253695ee069675db3480518420b3be881f2 \ + --hash=sha256:cdbb78909f52b981d3b2d56b97328d71eb0b974c36bd77c920123a7ebb192829 \ + --hash=sha256:cdc8b74ecc48c0cb1e9607a05ec4e9e88db60a19ffcc9a1d5f9088ede40c8dc0 \ + --hash=sha256:d0a24b40877af2de4950252be9d21eaf7fb07660f3c2cae1f56c6b599ada5266 \ + --hash=sha256:d22a945598fb91236b4dd793a6e42e4f3dd7740bb5aace5ebd7d4c08d13bb575 \ + --hash=sha256:d2f9fc07a8042a8f95925b35c4f04f469707c981fc33245b6ca187cf5d2dd290 \ + --hash=sha256:d625a186a65201c23a9e3b8ed9c47e90a026e03256608cc91851c6709096844f \ + --hash=sha256:d925f3d9afd05a8c0fb3a1031463a8d59ebe5e2afad297e29c78be19e13b4e62 \ + --hash=sha256:e64e88d5585bea9ce95861079de72006c7fa6d3df4e3a3b65ba31eb979c15c9f \ + --hash=sha256:e652ab17569c94bff5475520f907b7148b8c24036a8ebbe5cf7cf7493d28579a \ + --hash=sha256:e7b891faeedeafba41b2983e5001a81b6a915b69544c7e7570d1989ce1c36ac7 \ + --hash=sha256:e80675d75ae2cd14372cb65cad5400d9347a3d3f6c13000183f22dfd027283ed \ + --hash=sha256:e9c134bb666dd54b778b9fc0d2b50cbb7f979b9e3716f26a88c9ab3b6fc1dd0f \ + --hash=sha256:eb7d8d0e5886a89a55d2eef490e272fa965a9d57c6b29a5b5088a7997ec2cad1 \ + --hash=sha256:ecb42011e12ee19cafbc312887cbf3546959fe02fbad44f272d4be5baa997615 \ + --hash=sha256:ef3fbbf161dc9351a2fe0422e51b129f9e97e42385bd0320b309c15f7d287dd8 \ + --hash=sha256:efd62a42486f1bda5d24cb4f63d15a3c7768375fe83d36f9417b4ad7a2fb20b3 \ + --hash=sha256:f077d0b97ab11fa7dcc633fca53515f290bca8a8a633e966d5b6d1879d9ed01a \ + --hash=sha256:f332f0e72a5a0400141f830744e141bf9f97917878dbe968669e8a7fefea78ff \ + --hash=sha256:f7b0ec93a2893de856652154d73b7ba622f26fa97726487dcac373de5f4c6084 \ + --hash=sha256:fa10ef4112775900e7a0661068635eb67b2ab824fbde764de6e0e21982a93db0 \ + --hash=sha256:fc5d783bd4a2387e97b8a2d5ec781cfb92b3d893bf82370548e99db5915935d3 \ + --hash=sha256:fc8515076c11f3cfdf4fb142dcca0fe384b1230a3b5415458ac84f3e0903ec13 \ + --hash=sha256:ff218293c9c806138dca139765e3b067621be52bcd93cdc14c7711be7ddc90a9 + # via + # -c evals/fullsend/requirements.lock + # pydantic +pyyaml==6.0.3 \ + --hash=sha256:00c4bdeba853cc34e7dd471f16b4114f4162dc03e6b7afcc2128711f0eca823c \ + --hash=sha256:0150219816b6a1fa26fb4699fb7daa9caf09eb1999f3b70fb6e786805e80375a \ + --hash=sha256:02893d100e99e03eda1c8fd5c441d8c60103fd175728e23e431db1b589cf5ab3 \ + --hash=sha256:02ea2dfa234451bbb8772601d7b8e426c2bfa197136796224e50e35a78777956 \ + --hash=sha256:0f29edc409a6392443abf94b9cf89ce99889a1dd5376d94316ae5145dfedd5d6 \ + --hash=sha256:10892704fc220243f5305762e276552a0395f7beb4dbf9b14ec8fd43b57f126c \ + --hash=sha256:16249ee61e95f858e83976573de0f5b2893b3677ba71c9dd36b9cf8be9ac6d65 \ + --hash=sha256:1d37d57ad971609cf3c53ba6a7e365e40660e3be0e5175fa9f2365a379d6095a \ + --hash=sha256:1ebe39cb5fc479422b83de611d14e2c0d3bb2a18bbcb01f229ab3cfbd8fee7a0 \ + --hash=sha256:214ed4befebe12df36bcc8bc2b64b396ca31be9304b8f59e25c11cf94a4c033b \ + --hash=sha256:2283a07e2c21a2aa78d9c4442724ec1eb15f5e42a723b99cb3d822d48f5f7ad1 \ + --hash=sha256:22ba7cfcad58ef3ecddc7ed1db3409af68d023b7f940da23c6c2a1890976eda6 \ + --hash=sha256:27c0abcb4a5dac13684a37f76e701e054692a9b2d3064b70f5e4eb54810553d7 \ + --hash=sha256:28c8d926f98f432f88adc23edf2e6d4921ac26fb084b028c733d01868d19007e \ + --hash=sha256:2e71d11abed7344e42a8849600193d15b6def118602c4c176f748e4583246007 \ + --hash=sha256:34d5fcd24b8445fadc33f9cf348c1047101756fd760b4dacb5c3e99755703310 \ + --hash=sha256:37503bfbfc9d2c40b344d06b2199cf0e96e97957ab1c1b546fd4f87e53e5d3e4 \ + --hash=sha256:3c5677e12444c15717b902a5798264fa7909e41153cdf9ef7ad571b704a63dd9 \ + --hash=sha256:3ff07ec89bae51176c0549bc4c63aa6202991da2d9a6129d7aef7f1407d3f295 \ + --hash=sha256:41715c910c881bc081f1e8872880d3c650acf13dfa8214bad49ed4cede7c34ea \ + --hash=sha256:418cf3f2111bc80e0933b2cd8cd04f286338bb88bdc7bc8e6dd775ebde60b5e0 \ + --hash=sha256:44edc647873928551a01e7a563d7452ccdebee747728c1080d881d68af7b997e \ + --hash=sha256:4a2e8cebe2ff6ab7d1050ecd59c25d4c8bd7e6f400f5f82b96557ac0abafd0ac \ + --hash=sha256:4ad1906908f2f5ae4e5a8ddfce73c320c2a1429ec52eafd27138b7f1cbe341c9 \ + --hash=sha256:501a031947e3a9025ed4405a168e6ef5ae3126c59f90ce0cd6f2bfc477be31b7 \ + --hash=sha256:5190d403f121660ce8d1d2c1bb2ef1bd05b5f68533fc5c2ea899bd15f4399b35 \ + --hash=sha256:5498cd1645aa724a7c71c8f378eb29ebe23da2fc0d7a08071d89469bf1d2defb \ + --hash=sha256:5cf4e27da7e3fbed4d6c3d8e797387aaad68102272f8f9752883bc32d61cb87b \ + --hash=sha256:5e0b74767e5f8c593e8c9b5912019159ed0533c70051e9cce3e8b6aa699fcd69 \ + --hash=sha256:5ed875a24292240029e4483f9d4a4b8a1ae08843b9c54f43fcc11e404532a8a5 \ + --hash=sha256:5fcd34e47f6e0b794d17de1b4ff496c00986e1c83f7ab2fb8fcfe9616ff7477b \ + --hash=sha256:5fdec68f91a0c6739b380c83b951e2c72ac0197ace422360e6d5a959d8d97b2c \ + --hash=sha256:6344df0d5755a2c9a276d4473ae6b90647e216ab4757f8426893b5dd2ac3f369 \ + --hash=sha256:64386e5e707d03a7e172c0701abfb7e10f0fb753ee1d773128192742712a98fd \ + --hash=sha256:652cb6edd41e718550aad172851962662ff2681490a8a711af6a4d288dd96824 \ + --hash=sha256:66291b10affd76d76f54fad28e22e51719ef9ba22b29e1d7d03d6777a9174198 \ + --hash=sha256:66e1674c3ef6f541c35191caae2d429b967b99e02040f5ba928632d9a7f0f065 \ + --hash=sha256:6adc77889b628398debc7b65c073bcb99c4a0237b248cacaf3fe8a557563ef6c \ + --hash=sha256:79005a0d97d5ddabfeeea4cf676af11e647e41d81c9a7722a193022accdb6b7c \ + --hash=sha256:7c6610def4f163542a622a73fb39f534f8c101d690126992300bf3207eab9764 \ + --hash=sha256:7f047e29dcae44602496db43be01ad42fc6f1cc0d8cd6c83d342306c32270196 \ + --hash=sha256:8098f252adfa6c80ab48096053f512f2321f0b998f98150cea9bd23d83e1467b \ + --hash=sha256:850774a7879607d3a6f50d36d04f00ee69e7fc816450e5f7e58d7f17f1ae5c00 \ + --hash=sha256:8d1fab6bb153a416f9aeb4b8763bc0f22a5586065f86f7664fc23339fc1c1fac \ + --hash=sha256:8da9669d359f02c0b91ccc01cac4a67f16afec0dac22c2ad09f46bee0697eba8 \ + --hash=sha256:8dc52c23056b9ddd46818a57b78404882310fb473d63f17b07d5c40421e47f8e \ + --hash=sha256:9149cad251584d5fb4981be1ecde53a1ca46c891a79788c0df828d2f166bda28 \ + --hash=sha256:93dda82c9c22deb0a405ea4dc5f2d0cda384168e466364dec6255b293923b2f3 \ + --hash=sha256:96b533f0e99f6579b3d4d4995707cf36df9100d67e0c8303a0c55b27b5f99bc5 \ + --hash=sha256:9c57bb8c96f6d1808c030b1687b9b5fb476abaa47f0db9c0101f5e9f394e97f4 \ + --hash=sha256:9c7708761fccb9397fe64bbc0395abcae8c4bf7b0eac081e12b809bf47700d0b \ + --hash=sha256:9f3bfb4965eb874431221a3ff3fdcddc7e74e3b07799e0e84ca4a0f867d449bf \ + --hash=sha256:a33284e20b78bd4a18c8c2282d549d10bc8408a2a7ff57653c0cf0b9be0afce5 \ + --hash=sha256:a80cb027f6b349846a3bf6d73b5e95e782175e52f22108cfa17876aaeff93702 \ + --hash=sha256:b30236e45cf30d2b8e7b3e85881719e98507abed1011bf463a8fa23e9c3e98a8 \ + --hash=sha256:b3bc83488de33889877a0f2543ade9f70c67d66d9ebb4ac959502e12de895788 \ + --hash=sha256:b865addae83924361678b652338317d1bd7e79b1f4596f96b96c77a5a34b34da \ + --hash=sha256:b8bb0864c5a28024fac8a632c443c87c5aa6f215c0b126c449ae1a150412f31d \ + --hash=sha256:ba1cc08a7ccde2d2ec775841541641e4548226580ab850948cbfda66a1befcdc \ + --hash=sha256:bdb2c67c6c1390b63c6ff89f210c8fd09d9a1217a465701eac7316313c915e4c \ + --hash=sha256:c1ff362665ae507275af2853520967820d9124984e0f7466736aea23d8611fba \ + --hash=sha256:c2514fceb77bc5e7a2f7adfaa1feb2fb311607c9cb518dbc378688ec73d8292f \ + --hash=sha256:c3355370a2c156cffb25e876646f149d5d68f5e0a3ce86a5084dd0b64a994917 \ + --hash=sha256:c458b6d084f9b935061bc36216e8a69a7e293a2f1e68bf956dcd9e6cbcd143f5 \ + --hash=sha256:d0eae10f8159e8fdad514efdc92d74fd8d682c933a6dd088030f3834bc8e6b26 \ + --hash=sha256:d76623373421df22fb4cf8817020cbb7ef15c725b9d5e45f17e189bfc384190f \ + --hash=sha256:ebc55a14a21cb14062aa4162f906cd962b28e2e9ea38f9b4391244cd8de4ae0b \ + --hash=sha256:eda16858a3cab07b80edaf74336ece1f986ba330fdb8ee0d6c0d68fe82bc96be \ + --hash=sha256:ee2922902c45ae8ccada2c5b501ab86c36525b883eff4255313a253a3160861c \ + --hash=sha256:efd7b85f94a6f21e4932043973a7ba2613b059c4a000551892ac9f1d11f5baf3 \ + --hash=sha256:f7057c9a337546edc7973c0d3ba84ddcdf0daa14533c2065749c9075001090e6 \ + --hash=sha256:fa160448684b4e94d80416c0fa4aac48967a969efe22931448d853ada8baf926 \ + --hash=sha256:fc09d0aa354569bc501d4e787133afc08552722d3ab34836a80547331bb5d4a0 + # via + # -c evals/fullsend/requirements.lock + # -r evals/fullsend/requirements.in +referencing==0.37.0 \ + --hash=sha256:381329a9f99628c9069361716891d34ad94af76e461dcb0335825aecc7692231 \ + --hash=sha256:44aefc3142c5b842538163acb373e24cce6632bd54bdb01b21ad5863489f50d8 + # via + # -c evals/fullsend/requirements.lock + # jsonschema + # jsonschema-specifications +requests==2.34.2 \ + --hash=sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0 \ + --hash=sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed + # via + # -c evals/fullsend/requirements.lock + # google-auth +rpds-py==2026.9.1 \ + --hash=sha256:00ba2d8c7dd4ee537978ddf4b3fbd712bef2d8751603f7f3146b3f4287768e25 \ + --hash=sha256:01445c8d194aa032a08e944f16567672da1c62dbdbefd8b6d0693032e290cf68 \ + --hash=sha256:028ad274ea951dac64491b5d1e65712a4aeabfdbdb9fccf797b57bd899b0c495 \ + --hash=sha256:0483515261947e4e8b8e1375bf7463e7eb6ccfb3d86e7b554d90cd5285f20f32 \ + --hash=sha256:068c37bba854ec2fe42f7365c640af11dd9895890ccbf2df5070d0c059bd7f96 \ + --hash=sha256:074a4d198bc34d9a8ea425114fc3ded6d11ec01f6a314a8db67454a5152d8834 \ + --hash=sha256:07deecbfce94c78473018bc7d10b337cc651d12df87a1eb2cb3e4024bc9c33d0 \ + --hash=sha256:08dae4a4095150a7c4545a1fb40b98e1ab1744fbc2770d92c977b9dadaa49ab6 \ + --hash=sha256:0da298fb372dc192610a4b9ecbc68a0cd8b675bbbd1fc519d01b41cfd658333e \ + --hash=sha256:0f045bb053c9057720d72c56dffe30dffdc05997b2897a827b9325f0ab6623fa \ + --hash=sha256:10e208f2425d973938afcd56e28a7c4be32e27b6a60b5d381f49fb9d8acf9759 \ + --hash=sha256:136a1c3fe4402b7008bc81cb62ee538481795b61a7e83df88dff3b3f02b726ff \ + --hash=sha256:159a7aab5c5e8b112c8830f54717ce56da1252ebdbb526f5be2df2309280b9e7 \ + --hash=sha256:172e47169583f46ce118cbec68e6795d0da0f4606b488b6434f8276bca0a058c \ + --hash=sha256:1c2d1f6da5128eabf34e963d7163a818846075a52568250d006c4c953b40f903 \ + --hash=sha256:1d55198263bb51f557550c6ed2e6d1cb6a6fed6eb5c9120b741c5926bef8a45d \ + --hash=sha256:1d77b649e6f7cdf12ca5c2a98dad0ad37f9ea9b6f960408a92f0cb12bb3d04d9 \ + --hash=sha256:1e8d4d79d828299bf44a55db22a9388ab967b49d17132c88eab0f4360b48da8e \ + --hash=sha256:22ffd29a63d71fb1b81552c21f2c2b734949b7ac751a9be70675a939a900839b \ + --hash=sha256:2693b2728bbcc48d09a981a356954b0c47c53ff25b545856f28a889ea619f69a \ + --hash=sha256:270bdcdaac5d5b6f73c5e22e7e135c7f2a50e789f71d9e241d5be8d90026e19a \ + --hash=sha256:2711d29b653b3bce48a63d18b9c6b53274669e6d6c4094dddeb4d9a0e45128b2 \ + --hash=sha256:2c16ab111bc27c646ba8aa005d0527754edc538ebb636f0b1bf8e244b48d1945 \ + --hash=sha256:306ee1850d8105b5baf977e78d45fcadd12c1a54678d614c9baf217708446e91 \ + --hash=sha256:3231c4c0e521dafa5be0c9f114ee2c2ad46650836f2d72caa86801950c3e7044 \ + --hash=sha256:3890a6aa36e6baa53d5258a2a25d3ef8b37ad165a6ab27a892d7c3e3a432cd69 \ + --hash=sha256:3a72c11530d71abfb66c8d7696a2f86c43e63fca8b948f1a784ac490f4ec688e \ + --hash=sha256:3b5a6f40f0a1486b4b36c888123afc67acdbd9f33235927acf5ff295429a0ba3 \ + --hash=sha256:3c91c210ae7645626c608400e3519b4a642f837cce09ca830db3beb2e9f274d4 \ + --hash=sha256:3cd182d7291d29b92c521a0069d9c01ba6193628a9a105531d11b40a6d731a33 \ + --hash=sha256:3e524c7874ac72884d28e16dd5b8d839fd09e0fe76b020d3fbca23212a7b8c52 \ + --hash=sha256:3e93b2cd69a9830be33e03945cd7cda940a0a8bfcfbff41d6144f0cb0d3d8bd9 \ + --hash=sha256:3edae8c5ddfdb6985d49ae9d150516e5076888879022f91a26c2de9276ce0bdb \ + --hash=sha256:3f0e9ac28fc067d4d34b88ae43c48e9489455c97fee9633d851f7eeed5a05d35 \ + --hash=sha256:42e75466f83cd43f6026c81eab74246efb2bdadafb307b85700632d06c68f299 \ + --hash=sha256:44b32a7c4f0da3d28af31c259e38ddcff096f855e205ed0671d02fcf44f1ea1c \ + --hash=sha256:457866b85daf5034296666168b84a69e0b2e89dc4f1af102b46f6448a60b9063 \ + --hash=sha256:45bc6bccf78b20fd834237d18db64965d7ee68ba7f60440a26c7ab71e7b8d51a \ + --hash=sha256:46d80bc76b51a6c24f9944368c28d38b8bcbcea1da4f2f8d3ebc31a67e8c6ec6 \ + --hash=sha256:4793ef7f78268b124b73fa933440f01d258bbae01de9fa53e9080c9ab0425a12 \ + --hash=sha256:492e5e428cbe126221611f47e068f01660352feec4ad18bc0f5ea9b2ae88fb14 \ + --hash=sha256:4b26b03d9d2658ee2fa234f8f4f19f38a09773fe5261028025032e26d4d35af0 \ + --hash=sha256:4c0d2cb595a420b34d5086db0add011e26e2c09d6a024afbac4228bf8f863a30 \ + --hash=sha256:4cfaf02209061880210819934de2f4f6aa83dc04dafe6770276acc240a56da31 \ + --hash=sha256:501909f2e4a1e2dee528ef766fe3c469060ebc17e54a8383d404ba07a81a6f02 \ + --hash=sha256:50906f5aea24b5a865cbd0a589698288631d9f3a54c3a937c83aefa95a0d14af \ + --hash=sha256:54ac2158a6f96cfbabff0b2eedaf94b90c5ec7ca8317fcadc61e1c2b2e0ff6ef \ + --hash=sha256:56c6952a9b15047466d0c2347c446a761d4527f89976156341e68f0ce5cc08b0 \ + --hash=sha256:56cd8b3f77d7b6812f533b662186a1f28316931166ddc00fb893b1b0db7e9888 \ + --hash=sha256:57492a550a1d88d29d003247e5f78dd8cf04a701fac0e4c8db8745a6d2504e0a \ + --hash=sha256:5943980471829f6de242a20b109de3111ba6b77e3af0ffc587028ac854b05e6c \ + --hash=sha256:5c6ee90dee3e85e055ddfd502d611643d9b0fd94c818220bda84ec3dacd9b27b \ + --hash=sha256:5c90e7fa02e8f5de0d10c17595c568ada48c5302e749462c0ea1a4c362111a86 \ + --hash=sha256:5ce8943f79c2210f7abcc28e86367b03b28d95027fd01c46d2472373ae70c86f \ + --hash=sha256:617f59cde379b4f648a09797b7f683d04b90a46344cddab85639da5aff0f5531 \ + --hash=sha256:6307a0da524939decb8ca4a3933b8ab62525794411d6984fca6726e732804af6 \ + --hash=sha256:684fd492fff4fead00587544e059be2bbcb6f93454f21fa2a91b66fc7508be82 \ + --hash=sha256:6b5b393eda5ea42cca1c1a6665f2a4882b4fd5d1777e41ce0545a107fb008c9d \ + --hash=sha256:6b723eb406dec5bc9ec516c73ab9c3239a3284e017f7eb89ee2b3258bd504fb7 \ + --hash=sha256:6b9bf3135b4ad5981df9a73d71a35272d650a2985ae9c2746357b24d59de2448 \ + --hash=sha256:6beb738155fe8ab8091afdfa5a3226b21c2b1593f1e50ebb90eb25b44dbc0391 \ + --hash=sha256:6c0dbbcc19735fe5f8b0a54c07659d154a9e69f47e15d0a6ab7299215daf62cb \ + --hash=sha256:6cdc537c8633d7fd92a82e2e0d2ab74320a3f63d5e59fb9cf08711e08fe151c4 \ + --hash=sha256:6eae33003518fd4cb4f83a218d5371469dd3001aa3b87128c005b07762f7fe5e \ + --hash=sha256:740d0a99cf9de0b17a3943388e9294a59becf75e7c43421f387bd3c7a9901f7c \ + --hash=sha256:75c38c50ab9aca840225d9a9a3810bf11d04bd5c1f186cabbb8aee56db3e9b15 \ + --hash=sha256:761fdae6728ceb99ab182fad2f0cc1e262f610834dc891aea1d1a2a2e634776f \ + --hash=sha256:7664419f27db41d4f1c43a78dccda7dd6e8ef2428df3ee01d0c2a07a6b071297 \ + --hash=sha256:76d3af9732d2dab69f28179b40ba2d87e2f1d5824b4a694780aa787d685e8f36 \ + --hash=sha256:78326f4cb4427a56ba4996c0762b63be45f06b85f086526420d2b3a66e40f84d \ + --hash=sha256:7868b85224291c6cb6759f9b5adb9745f486d226f62b16a614dd5a2a5ab2b35b \ + --hash=sha256:815d26356930846a40c7bc1366e7b1b0320ab8a063e66c11298a208bed0fd237 \ + --hash=sha256:8171b44a054e5c67fd748ada04187f1250bf35b95f85e52ab64bcf3331a923bb \ + --hash=sha256:821b2755db9194409254012f429c56643416fb96ef9be090be82ec8826b7f477 \ + --hash=sha256:837c6b305e26fe0f75b15c92cf3b2ba29e0ae19dc40b1c557b026cb426347d0c \ + --hash=sha256:839dde845559254f34885267c6878f60d61d5205180226d976fe488d45fa128e \ + --hash=sha256:84a6ecc0c940169190d2c23bd969debd48c94dbc855acd60188a68d71d421608 \ + --hash=sha256:8601470267d938bcb7f3ab1a336100af51a4fd5b6ed030ef52461bb3ef5e7e07 \ + --hash=sha256:88b5268892fde430d5531f95bc560b6efbbd67c929662c586afd729a96e7461c \ + --hash=sha256:8aa5dda18d39b6143eb24809d158f9252c88f402749b6f1b62a506cc7d96cc35 \ + --hash=sha256:926bdd3e3b5998ddf70cc64bc8cf57209571f9044542913afb673799fec77dd0 \ + --hash=sha256:96beca19ec79de272e8668585380ff9092c47077c1d7a1e098e00bbd921f4785 \ + --hash=sha256:9a0460d43603d1fd9ef59c30278531e15d78581721ddb538fa560aa7817ea4ad \ + --hash=sha256:a03d57b86d2a51d0a66c92177e2be154ad015f357791d306e714569999cdb4cc \ + --hash=sha256:a36b70596407634ca82d4b989a3729074a008537a0522e4c8046a67c729103e9 \ + --hash=sha256:a3a52a3ba86436ab3aef510fbe21512abc2ddd1993005dfe50514bd2284ef025 \ + --hash=sha256:a3dbc5ed9514908d5046107d7b1346bde71eea61de6e0e4919c19354f97e769f \ + --hash=sha256:a431156bb41865fc14cd5d79bb9d7bbed83110b0159e34e62ae30951f96c0009 \ + --hash=sha256:a575404ebc9cf2e91edd32eaf570ec1430eb900d4f56724ba7dd4bc1fc9c176d \ + --hash=sha256:a5cf77eb04f20b720be95265a3e00eb2a14814074255cc27069c551b2db53118 \ + --hash=sha256:a8763f20692da7df39b0afdd1ba3042b004c50a45994f76c2d9a25641f7673db \ + --hash=sha256:ab4b2fda7c2b542f7f9d886cc6a838c5079d2b76f72e6081411faba11adde2c9 \ + --hash=sha256:addeda51556dac7c1a2f14cda62db8b621cd12afba3091d03a96c72932387eab \ + --hash=sha256:b242c27c8f836305a4a72df9cdd564386ac57b807bd252a063223331c9316b37 \ + --hash=sha256:b4f062343e7ad3fa94f2c66e5ae667dee47ee74dd41a9057c4fbe163236a123d \ + --hash=sha256:b5b8b0753718d258fd454283fbd57e14545d3b40583fa672e27cb4f987626bcc \ + --hash=sha256:be3e47e2d91aa3942ff9bf4077a505226005abfc39b6f7554a91c1b9393986b9 \ + --hash=sha256:befc2d6a953e563f8a7bfd87a42c22ebf8a3e980dcb7b6a4d17b70b0e914e8a3 \ + --hash=sha256:bf35d0568abda97233239ce32896d3ad53fccc537832c104e30c94aa5fb93569 \ + --hash=sha256:c933c6678c6f116ff8af47a4c6db0868b8ace74af0343016c0ef00f00272ea69 \ + --hash=sha256:c9d1aca01f49170fdcf5c92761b1fafe97f554b721ca4570c5949fff778f0d4b \ + --hash=sha256:cdeaa99ce822dca76cfb1b993e9120c5ea212f2eb66d48950ad63c349668a018 \ + --hash=sha256:ce4d4f52e2a4324396caddbd45a97d8d7be5f42edd25d2355282a9c34f9b2f7f \ + --hash=sha256:d1028417bb44037eb3069c1009bd7b7277212876cda22fbe565b0bca9fab6d2c \ + --hash=sha256:d151e148117294133bf8af7eeace085e7e87432db15ab6adf640330298a47f6f \ + --hash=sha256:d7841166b7fa64c9c56404617ae4341448847482d45933b13135d26c130519e5 \ + --hash=sha256:d7fca4eb6df565e2a928f1c7dad92d27db8f9df0f449e76423ed5d7e713ed445 \ + --hash=sha256:d95a354e02393eada6d7351184671aced9d4cce109dabf927cb7aa99624352a1 \ + --hash=sha256:d9edf30457d74eebfd76b045535e36f1cd89062566a128a0db2145ca042d787e \ + --hash=sha256:dbc2673f9223d420c91145599b3ba45a8a50c207d1976908e5fb5ddb0c9b9429 \ + --hash=sha256:e01b3c878c8641913e688edd1b3f08658c6783d29cf6b826bd3c0d1ae7a1ffaa \ + --hash=sha256:e21c1429e205828ea886a2293a4a2c8e01f4c25d9893ca330e97a6cf73f52e7b \ + --hash=sha256:e43d4a1f673e8a1cbd8533e809e02b4bf9d4f2280269bb640436556312121250 \ + --hash=sha256:e6d198bad4e49dd6732fbd636e2fc5c082f45c8cad0b4acb756b00c82c76072e \ + --hash=sha256:e6ea1cda8d8c688278430e4268a42f5e5da3bdd74578dfadc0820c3f1766ce83 \ + --hash=sha256:ea394a937f17a54c51239348bdbe2e3518124c8d4a8951ba04a311d3095bd18f \ + --hash=sha256:eac2f5dbafd585dfe31f86a23ebf0d3ba480a9d49ebc87947267b5608d4ea0cd \ + --hash=sha256:eb61be926bb81567c1f48bdc8aa22b9855048dc2efd53871f9f7e6e9a5632346 \ + --hash=sha256:eba5d173f7d5708b22a93815017a4611873ed54db9f268077c0dd1ed99cfc858 \ + --hash=sha256:ec450527cbf485e13c8d3602a54f428ab0432fdade0ede75efd74b735421c871 \ + --hash=sha256:eef6a03b0b6d08d0835ccfa8ec8d1bc70525e3801387567137b50c557695e6da \ + --hash=sha256:ef0d8c843e2827d6c120ab4687e9423fb1d893db1df27b7c1506615bcb9734a0 \ + --hash=sha256:ef6b65b03247c54692ad4fd9ee97cb772781927db72e3cb05e70b3db6d1ff14f \ + --hash=sha256:f3d6ed6a98cfd19155996605474982cc470d7601746a6439078f1a5a3fa8b050 \ + --hash=sha256:fce4b85234a0cbad67bf8e6e1201ee815d172c9aebad75f25645bc4d834f8e31 \ + --hash=sha256:fda1d96e542c37b6c804547dbf489c129fe7c97183a76a5ec275909ba1a063df \ + --hash=sha256:fdcd198979b4ecffcc1beba366a7fbcf4eb41243691a82fe52ceb0b902f09c12 \ + --hash=sha256:fe5ad0664ec772b02c45859041aa17655709cced7a31005817fbbbd988c25567 + # via + # -c evals/fullsend/requirements.lock + # jsonschema + # referencing +setuptools==84.0.0 \ + --hash=sha256:51a52592b3b99e102b609654876bd65f19f999935166d1352678931132b0c670 \ + --hash=sha256:f4695c21257f0d9b537ec2692c941d02ee143b7cc1276941349a546573b2ef73 + # via + # -c evals/fullsend/requirements.lock + # -r evals/fullsend/requirements.in +sniffio==1.3.1 \ + --hash=sha256:2f6da418d1f1e0fddd844478f41680e794e6051915791a034ff65e5f100525a2 \ + --hash=sha256:f4324edc670a0f49750a81b895f35c3adb843cca46f0530f79fc1babb23789dc + # via + # -c evals/fullsend/requirements.lock + # anthropic +truststore==0.10.4 \ + --hash=sha256:9d91bd436463ad5e4ee4aba766628dd6cd7010cf3e2461756b3303710eebc301 \ + --hash=sha256:adaeaecf1cbb5f4de3b1959b42d41f6fab57b2b1666adb59e89cb0b53361d981 + # via + # -c evals/fullsend/requirements.lock + # -r evals/fullsend/requirements.in + # httpcore2 + # httpx2 +typing-extensions==4.16.0 \ + --hash=sha256:481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8 \ + --hash=sha256:dc983d19a509c94dba722ee6abd33940f7c05a89e243c47e907eb4db6f1a43e5 + # via + # -c evals/fullsend/requirements.lock + # anthropic + # anyio + # httpx2 + # pydantic + # pydantic-core + # referencing + # typing-inspection +typing-inspection==0.4.4 \ + --hash=sha256:547274fa6b0a561ccf549cc9524b999a578e737d015d8709d021f9d0d13bea47 \ + --hash=sha256:65b8397ba37ccbce054456aaccddfc91e6e3083c92824df348d96ca832f3f147 + # via + # -c evals/fullsend/requirements.lock + # pydantic +urllib3==2.8.0 \ + --hash=sha256:0cf3cae568d36aa9576b28dfb35f11328f1cb974ca7647d9475ebb86c75ac6e3 \ + --hash=sha256:63bf2ead4c879426ebf22ef2a781eeb4aa3b4ae798a0435506f8687fd5bb9b63 + # via + # -c evals/fullsend/requirements.lock + # requests +wheel==0.48.0 \ + --hash=sha256:3217dcc807155e45db462d7ef2431f5ddda0d7273b700d05a67b271ceb1287ab \ + --hash=sha256:94800765601e9171bf5d58d066e640662842bcedcbab982b2c90787a2c987322 + # via + # -c evals/fullsend/requirements.lock + # -r evals/fullsend/requirements.in diff --git a/evals/fullsend/run.py b/evals/fullsend/run.py new file mode 100644 index 000000000..544a3b5ef --- /dev/null +++ b/evals/fullsend/run.py @@ -0,0 +1,351 @@ +#!/usr/bin/env python3 +"""Common local/CI entrypoint for the native Fullsend gate suite.""" + +import argparse +import hashlib +import json +import os +from pathlib import Path +import platform +import re +import shutil +import subprocess +import sys +import tarfile +import tempfile +import uuid +import venv + + +ROOT = Path(__file__).resolve().parents[2] +HERE = Path(__file__).resolve().parent +CASES = ["033-absent", "034-empty", "035-malformed", "036-valid"] +ASSERTION_COUNTS = dict(zip(CASES, [4, 5, 5, 7])) + + +def pins(): + """Read the repository-owned immutable dependency manifest.""" + return json.loads((HERE / "dependencies.json").read_text()) + + +def binaries(cache): + """Select host and Linux sandbox binaries; macOS cannot run in the sandbox.""" + arch = {"x86_64": "amd64", "arm64": "arm64", "aarch64": "arm64"}.get(platform.machine()) + system = {"Darwin": "darwin", "Linux": "linux"}.get(platform.system()) + if not arch or not system: + raise ValueError("Supported hosts are macOS/Linux amd64/arm64") + return cache / "bin" / f"fullsend-{system}-{arch}", cache / "bin" / f"fullsend-linux-{arch}" + + +def install_binary(cache, key, dependency): + """Verify an archive's immutable digest before installing only its CLI file.""" + destination = cache / "bin" / f"fullsend-{key}" + archive = cache / f"fullsend-{key}.tar.gz" + if not archive.exists(): + url = f'https://github.com/fullsend-ai/fullsend/releases/download/v{dependency["version"]}/fullsend_{dependency["version"]}_{key.replace("-", "_")}.tar.gz' + subprocess.run(["curl", "--fail", "--location", "--silent", "--show-error", + "--output", str(archive), url], check=True) + if hashlib.sha256(archive.read_bytes()).hexdigest() != dependency["archives"][key]: + raise ValueError(f"Fullsend archive digest mismatch for {key}") + with tarfile.open(archive, "r:gz") as package: + candidates = [m for m in package.getmembers() if m.isfile() and Path(m.name).name == "fullsend"] + if len(candidates) != 1: + raise ValueError("Fullsend archive must contain exactly one regular CLI binary") + destination.parent.mkdir(parents=True, exist_ok=True) + with package.extractfile(candidates[0]) as source, destination.open("wb") as target: + shutil.copyfileobj(source, target) + destination.chmod(0o755) + + +def verify_source(source, dependency): + """Reject changed tracked framework code or a different source revision.""" + actual = subprocess.check_output(["git", "-C", str(source), "rev-parse", "HEAD"], text=True).strip() + dirty = subprocess.check_output(["git", "-C", str(source), "status", "--porcelain", "--untracked-files=no"], text=True) + if actual != dependency["commit"] or dirty: + raise ValueError("Eval-harness source is not the clean approved immutable revision") + + +def setup(cache): + """Install isolated locked dependencies and CLI binaries; never inference/services.""" + if sys.version_info[:2] != (3, 12): + raise ValueError("Use Python3.12 for the locked dependency setup") + dependency = pins() + cache.mkdir(parents=True, exist_ok=True) + source = cache / "agent-eval-harness" + if not source.exists(): + subprocess.run(["git", "init", "-q", str(source)], check=True) + subprocess.run(["git", "-C", str(source), "fetch", "--depth", "1", dependency["harness"]["repository"], dependency["harness"]["commit"]], check=True) + subprocess.run(["git", "-C", str(source), "checkout", "--detach", "FETCH_HEAD"], check=True) + verify_source(source, dependency["harness"]) + environment = cache / "venv" + if not (environment / "bin/python").is_file(): + venv.EnvBuilder(with_pip=True).create(environment) + python = environment / "bin/python" + pip = [str(python), "-m", "pip", "--cache-dir", str(cache / "pip-cache"), "install", "--no-user"] + subprocess.run([*pip, "--require-hashes", "-r", str(HERE / "requirements.lock")], check=True) + subprocess.run([*pip, "--no-deps", "--no-build-isolation", str(source)], check=True) + for binary in set(binaries(cache)): + install_binary(cache, binary.name.removeprefix("fullsend-"), dependency["fullsend"]) + print("Isolated dependency setup complete; no inference or host-service changes") + + +def resolved_config(python, model, judge_model, effort, plugin_root=None): + """Resolve repository paths before upstream workspace/execute path handling.""" + import yaml + config = yaml.safe_load((HERE / "triage-security/eval.yaml").read_text()) + config["dataset"]["path"] = str(HERE / "triage-security/cases") + config["runner"]["command"][0:2] = [str(python), str(HERE / "triage-security/run-fullsend.py")] + if plugin_root is not None: + config["runner"]["command"].extend(["--plugin-root", str(plugin_root)]) + config["runner"]["effort"] = effort + config["models"] = {"skill": model, "judge": judge_model} + return config + + +def verify_locked_dependencies(lock): + """Reject installed version drift from the checked-in transitive hash lock.""" + import importlib.metadata + requirements = re.findall(r"^([A-Za-z0-9_.-]+)==([^\s;\\]+)", lock.read_text(), re.MULTILINE) + if not requirements: + raise ValueError("Locked dependency manifest is empty") + for package, expected in requirements: + try: + actual = importlib.metadata.version(package) + except importlib.metadata.PackageNotFoundError: + actual = None + if actual != expected: + raise ValueError(f"Locked dependency mismatch: {package}, expected {expected}, installed {actual}") + + +def preflight(cache, model, judge_model, effort): + """Validate dependencies/config/CLI contracts only; no sandbox or model launch.""" + import importlib.metadata + import jsonschema # Required by the trusted host output validator. + import yaml + from agent_eval.config import EvalConfig + dependency = pins() + if sys.version_info[:2] != (3, 12): + raise ValueError("Preflight requires the Python3.12 isolated environment") + verify_source(cache / "agent-eval-harness", dependency["harness"]) + if importlib.metadata.version("agent-eval-harness") != dependency["harness"]["version"]: + raise ValueError("Unexpected installed eval-harness version") + verify_locked_dependencies(HERE / "requirements.lock") + subprocess.run([sys.executable, "-m", "pip", "--cache-dir", str(cache / "pip-cache"), "check"], check=True) + for phase in ["workspace", "execute", "collect", "score"]: + subprocess.run([sys.executable, str(cache / f"agent-eval-harness/skills/eval-run/scripts/{phase}.py"), "--help"], + cwd=ROOT, stdout=subprocess.DEVNULL, check=True) + host, sandbox = binaries(cache) + if not sandbox.is_file(): + raise ValueError("Missing Linux sandbox Fullsend binary; run setup") + for executable, version in [(host, dependency["fullsend"]["version"]), ("openshell", dependency["openshell"]), + ("openshell-gateway", dependency["openshell"])]: + result = subprocess.check_output([str(executable), "--version"], text=True) + if not re.search(r"(? {index - 1}' + if result["value"] is not None or rationale != f"Skipped: condition '{condition}' is false": + raise ValueError(f"Invalid upstream summary: {case}/{name} requires its nonapplicable skip") + + +def phase_failure_code(stderr): + """Map observed stderr markers to fixed advisory categories, never raw messages.""" + for code, pattern in [ + ("authentication-error", r"unauthenticated|unauthorized|invalid_grant|\b401\b"), + ("permission-error", r"permission.?denied|forbidden|\b403\b"), + ("quota-error", r"resource_exhausted|rate.limit|\b429\b"), + ("timeout", r"timed? ?out|deadline exceeded"), + ("connection-error", r"connection refused|connection reset|name resolution"), + ]: + if re.search(pattern, stderr, re.IGNORECASE): + return code + return "phase-exit" + + +def pipeline(python, source, config, workspace, run_dir, run_id, environment, diagnostics=None): + """Delegate case execution/collection/grading entirely to the upstream framework.""" + diagnostics = diagnostics if diagnostics is not None else {} + diagnostics["phase_exits"] = {} + scripts = source / "skills/eval-run/scripts" + phases = [ + ("workspace", ["--config", str(config), "--run-id", run_id, "--symlinks", "none"]), + ("execute", ["--workspace", str(workspace), "--config", str(config), "--output", str(run_dir), "--run-id", run_id]), + ("collect", ["--config", str(config), "--workspace", str(workspace), "--output", str(run_dir)]), + ("score", ["judges", "--config", str(config), "--run-id", run_id]), + ] + for phase, arguments in phases: + diagnostics.update(phase=phase, code="runtime-error") + with tempfile.TemporaryFile() as stderr: + process = subprocess.run([str(python), str(scripts / f"{phase}.py"), *arguments], + cwd=ROOT, env=environment, stderr=stderr, check=False) + stderr.seek(0) + text = stderr.read(65536).decode("utf-8", errors="replace") + sys.stderr.write(text) # The trusted wrapper keeps this stream private. + while chunk := stderr.read(65536): + sys.stderr.write(chunk.decode("utf-8", errors="replace")) + diagnostics["phase_exits"][phase] = process.returncode + diagnostics["code"] = phase_failure_code(text) if process.returncode else "none" + if phase == "execute": + if not all((run_dir / "cases" / case / "run_result.json").is_file() for case in CASES): + diagnostics["code"] = "missing-case-results" + raise ValueError("Missing native case results; infrastructure failure, refusing zero-case grading") + if process.returncode: + print(f"Upstream execution exit {process.returncode}; retaining actual exits for native evidence judges", file=sys.stderr) + elif process.returncode: + return process.returncode + diagnostics.update(phase="summary", code="invalid-summary") + validate_summary(run_dir, run_id) + diagnostics.update(phase="complete", code="none") + return 0 + + +def publish_report(run_dir, destination, source, exit_code, diagnostics=None): + """Export only source pins and Boolean outcomes; raw evidence stays private.""" + import yaml + sha_keys = {"head_sha", "merge_sha", "base_sha", "trusted_sha", "eval_source_sha"} + if (set(source) != sha_keys | {"pr_number"} or type(source["pr_number"]) is not int + or source["pr_number"] <= 0 + or any(not re.fullmatch(r"[0-9a-f]{40}", source[key]) for key in sha_keys)): + raise ValueError("Invalid CI source provenance") + outcomes = {} + complete = False + try: + validate_summary(run_dir, run_dir.name) + summary = yaml.safe_load((run_dir / "summary.yaml").read_text()) + outcomes = {case: {f"assertion_{i}": summary["per_case"][case][f"assertion_{i}"]["value"] + for i in range(1, count + 1)} for case, count in ASSERTION_COUNTS.items()} + complete = True + except ValueError: + exit_code = exit_code or 1 + report = {"source": source, "exit_code": exit_code, "complete": complete, "outcomes": outcomes, + "passed": sum(value is True for case in outcomes.values() for value in case.values()), "total": 21} + if diagnostics is not None: + phases = {"preflight", "configuration", "workspace", "execute", "collect", "score", "summary", "complete"} + codes = {"none", "runtime-error", "phase-exit", "missing-case-results", "invalid-summary", + "authentication-error", "permission-error", "quota-error", "timeout", "connection-error"} + if (set(diagnostics) != {"phase", "code", "phase_exits"} + or diagnostics["phase"] not in phases or diagnostics["code"] not in codes + or not isinstance(diagnostics["phase_exits"], dict) + or any(key not in {"workspace", "execute", "collect", "score"} or type(value) is not int + for key, value in diagnostics["phase_exits"].items())): + raise ValueError("Invalid safe diagnostic values") + report["diagnostics"] = diagnostics + destination.mkdir(parents=True, exist_ok=True) + (destination / "native-result.json").write_text(json.dumps(report, indent=2) + "\n") + + +def main(): + """Use the same setup/preflight/run command in local shells and trusted CI.""" + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("command", choices=["setup", "preflight", "run"]) + parser.add_argument("--cache", type=Path, default=Path(tempfile.gettempdir()) / "tc-6677-eval-deps") + parser.add_argument("--output", type=Path, default=Path(tempfile.gettempdir()) / "tc-6677-native-evals") + parser.add_argument("--model", default="claude-opus-4-8") + parser.add_argument("--judge-model", default="claude-opus-4-6") + parser.add_argument("--effort", choices=["low", "medium", "high", "max"], default="high") + parser.add_argument("--plugin-root", type=Path, help="Sandbox-tested plugin; tooling always comes from this checkout") + parser.add_argument("--report-dir", type=Path, help="Allowlisted CI result directory; never raw credentials/transcripts") + args = parser.parse_args() + cache = args.cache.resolve() + run_dir = args.output.resolve() / "missing-run" + exit_code = 1 + diagnostics = {"phase": "preflight", "code": "runtime-error", "phase_exits": {}} + try: + if args.command == "setup": + setup(cache) + return 0 + python = cache / "venv/bin/python" + if not python.is_file(): + raise ValueError("Missing isolated dependencies; run setup first") + if Path(sys.prefix).resolve() != (cache / "venv").resolve(): + os.execv(str(python), [str(python), str(Path(__file__).resolve()), *sys.argv[1:]]) + preflight(cache, args.model, args.judge_model, args.effort) + if args.command == "preflight": + return 0 + diagnostics["phase"] = "configuration" + import yaml + run_id = "tc6677-" + uuid.uuid4().hex + workspace = Path(tempfile.gettempdir()) / "agent-eval" / run_id + run_dir = args.output.resolve() / "triage-security-gate" / run_id + if workspace.exists() or run_dir.exists(): + raise ValueError("Refusing existing run identifier") + run_dir.mkdir(parents=True) + config = run_dir / "eval.yaml" + config.write_text(yaml.safe_dump(resolved_config(python, args.model, args.judge_model, args.effort, + args.plugin_root.absolute() if args.plugin_root else None), sort_keys=False)) + host, sandbox = binaries(cache) + environment = dict(os.environ, TC6677_FULLSEND_BIN=str(host), TC6677_SANDBOX_FULLSEND_BIN=str(sandbox), + AGENT_EVAL_RUNS_DIR=str(args.output.resolve())) + print(f"Native evidence destination: {run_dir}", flush=True) + exit_code = pipeline(python, cache / "agent-eval-harness", config, workspace, run_dir, run_id, environment, diagnostics) + return exit_code + except (OSError, ValueError, subprocess.SubprocessError) as exc: + print(f"Fullsend eval setup/runner failed: {exc}", file=sys.stderr) + return 1 + finally: + if args.command == "run" and args.report_dir: + source = {key: os.environ.get("TC6726_" + key.upper(), "") + for key in ["head_sha", "merge_sha", "base_sha", "trusted_sha", "eval_source_sha"]} + source["pr_number"] = int(os.environ.get("TC6726_PR_NUMBER", "0")) + publish_report(run_dir, args.report_dir.resolve(), source, exit_code, diagnostics) + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/evals/fullsend/triage-security/agent.md b/evals/fullsend/triage-security/agent.md new file mode 100644 index 000000000..179014203 --- /dev/null +++ b/evals/fullsend/triage-security/agent.md @@ -0,0 +1,33 @@ +--- +name: triage-security-gate +description: Synthetic native Fullsend execution of the actual triage-security Skill +--- + + + +Evaluation clock: 2026-09-21T12:00:00Z. Invoke **sdlc-workflow:triage-security** +through the actual **Skill tool** for synthetic issue **TC-8101** as your first +tool call. Do not read or parse trusted input before invoking it. If it cannot +load, expose that failure and stop; do not manually read a Skill file or replace +its execution with expected behavior. + +Follow every applicable instruction in the delivered Skill and its companions. +Before its gate instruction, observe only FULLSEND_OUTPUT_DIR presence and value +through a real Bash tool call (never the entire environment or credentials). +Execute each delivered gate/input/final-validation Bash instruction separately, +preserving substituted plugin paths and actual failures. Do not extract commands +from a source file or duplicate them from this fixture. Stop immediately when the +Skill prescribes stopping. Do not mask exits or create a narrated fallback. + +For interactive routing, use the synthetic target CLAUDE.md and its existing +missing-Security-Configuration guard. For sandbox routing, use only the actual +mounted trusted input and complete the applicable trusted-evidence analysis. +After successful analysis, write the completed result to the Skill's native +output location and execute its actual inline final JSON/schema validator. +The separate native host validation loop does not substitute for that instruction. + +Preserve the Skill's single sandbox result-file contract. Never write local +analysis, logs, receipts, matrices or supplementary sandbox artifacts. Do not +call Jira/GitHub/CVE/WebFetch/lifecycle services, inspect credentials, mutate +anything externally or launch another model CLI. Fullsend retains native runtime +artifacts outside the Skill's output directory; do not fabricate those artifacts. diff --git a/evals/fullsend/triage-security/cases/033-absent/annotations.yaml b/evals/fullsend/triage-security/cases/033-absent/annotations.yaml new file mode 100644 index 000000000..9792ad646 --- /dev/null +++ b/evals/fullsend/triage-security/cases/033-absent/annotations.yaml @@ -0,0 +1,12 @@ +# SYNTHETIC TEST DATA — strict native execution assertions, never expected-behavior receipts +assertion_count: 4 +assertions: + - Native runtime records show actual sdlc-workflow:triage-security Skill invocation for TC-8101, delivered PR + plugin instructions and genuine subsequent Bash calls/results. Fixture/bootstrap/runtime errors, missing or + wrong-version Skill binding, narrated fallbacks, duplicated commands and schema-only sample output do not count. + - An actual Bash probe proves FULLSEND_OUTPUT_DIR is absent, not exported empty. The delivered Step 0.6 actually + prints interactive mode and exits0. + - After that gate, genuine tools read the synthetic target CLAUDE.md and reach the existing missing-Security-Configuration + guard directing /setup. No trusted-input validation or result generation occurs. + - The Skill stops before Jira initialization, credential inspection, external lookups, mutations or later triage. + Native Skill output inventory is empty. diff --git a/evals/fullsend/triage-security/cases/033-absent/input.yaml b/evals/fullsend/triage-security/cases/033-absent/input.yaml new file mode 100644 index 000000000..633eca313 --- /dev/null +++ b/evals/fullsend/triage-security/cases/033-absent/input.yaml @@ -0,0 +1,3 @@ +# SYNTHETIC TEST DATA — independent native Fullsend gate case +scenario: absent +issue: TC-8101 diff --git a/evals/fullsend/triage-security/cases/034-empty/annotations.yaml b/evals/fullsend/triage-security/cases/034-empty/annotations.yaml new file mode 100644 index 000000000..d41f9bb9b --- /dev/null +++ b/evals/fullsend/triage-security/cases/034-empty/annotations.yaml @@ -0,0 +1,13 @@ +# SYNTHETIC TEST DATA — strict native execution assertions, never expected-behavior receipts +assertion_count: 5 +assertions: + - Native runtime records show actual sdlc-workflow:triage-security Skill invocation for TC-8101, delivered PR + plugin instructions and genuine subsequent Bash calls/results. Fixture/bootstrap/runtime errors, missing or + wrong-version Skill binding, narrated fallbacks, duplicated commands and schema-only sample output do not count. + - An actual Bash probe proves FULLSEND_OUTPUT_DIR is present and empty, distinct from absent. + - 'Actual delivered Step0.6 Bash fails with exit1 and exact stderr ERROR: FULLSEND_OUTPUT_DIR is set but empty + followed by newline. Native CLI or host-validator exit is not the gate exit.' + - Raw tools show no trusted-input validation, target CLAUDE.md read, credential inspection, external tools, interactive + fallback or later triage after the gate failure. + - Actual native Skill output inventory is empty; no agent-result.json or synthetic expected-behavior output is + created. diff --git a/evals/fullsend/triage-security/cases/034-empty/input.yaml b/evals/fullsend/triage-security/cases/034-empty/input.yaml new file mode 100644 index 000000000..996d792e3 --- /dev/null +++ b/evals/fullsend/triage-security/cases/034-empty/input.yaml @@ -0,0 +1,3 @@ +# SYNTHETIC TEST DATA — independent native Fullsend gate case +scenario: empty +issue: TC-8101 diff --git a/evals/fullsend/triage-security/cases/035-malformed/annotations.yaml b/evals/fullsend/triage-security/cases/035-malformed/annotations.yaml new file mode 100644 index 000000000..3ceb86f94 --- /dev/null +++ b/evals/fullsend/triage-security/cases/035-malformed/annotations.yaml @@ -0,0 +1,31 @@ +# SYNTHETIC TEST DATA — strict native execution assertions, never expected-behavior receipts +assertion_count: 5 +assertions: + - Native runtime records show actual sdlc-workflow:triage-security Skill invocation for TC-8101, delivered PR + plugin instructions and genuine subsequent Bash calls/results. Fixture/bootstrap/runtime errors, missing or + wrong-version Skill binding, narrated fallbacks, duplicated commands and schema-only sample output do not count. + - 'Actual Bash probe and delivered Step0.6 prove native nonempty FULLSEND_OUTPUT_DIR=/sandbox/workspace/output + and sandbox mode: /sandbox/workspace/output with exit0, before the separate input failure.' + - 'Delivered Step0.7 really reads /sandbox/workspace/.pre-script/triage-security-input.json and delivered plugin + input schema, emits ERROR: trusted triage-security input is invalid JSON: with parser detail, and fails with + exit1.' + - >- + Genuine raw tools must prove actual sdlc-workflow:triage-security invocation and delivered invalid JSON parser + error with tool exit1, native Fullsend CLI exit nonzero, and no successful analysis, fallback or actions. + Output directory absence or a failed attempted abort write AFTER proven real Skill invalid JSON rejection + is accepted on the no-result path, not a disqualifying bootstrap/inference failure. Infrastructure/inference + failure is disqualifying only when it prevents actual Skill input validation. Without genuine invalid JSON proof, FAIL. + Accept either no result file and an empty output inventory with host validate-output-schema.sh rejecting the + absent result (no recovery write is required), OR collected agent-result.json exactly {} after intentional + host stripping. If an error-only file was written, raw tools must prove its successful write BEFORE host + validation, exactly {"error":"triage-security aborted: trusted input is missing, invalid JSON, or fails + triage-security-input.schema.json; no interactive fallback is available in the sandbox."}, with no + schema_version/mode/report/actions or extra diagnostics; subsequent host strip_extra_properties.py records + must show stripped: ['error'] and rejection of the success schema. Require the complete ordered evidence + chain for the applicable output path. Mere empty output or nonzero alone, infrastructure/inference failure + preventing actual Skill input validation, + nonempty unexpected output or success report, and narrated outcomes are FAIL. + - Genuine raw tools show no analysis, successful final validation, Jira/GitHub/WebFetch/CVE/lifecycle/Git/credential + inspection, executed actions or interactive fallback. Native inventory contains no output files OR sole + agent-result.json containing {}; the expected host rejection of an absent or schema-invalid result must be + distinct from fixture/runtime failure. diff --git a/evals/fullsend/triage-security/cases/035-malformed/input.yaml b/evals/fullsend/triage-security/cases/035-malformed/input.yaml new file mode 100644 index 000000000..baa9797df --- /dev/null +++ b/evals/fullsend/triage-security/cases/035-malformed/input.yaml @@ -0,0 +1,3 @@ +# SYNTHETIC TEST DATA — independent native Fullsend gate case +scenario: malformed +issue: TC-8101 diff --git a/evals/fullsend/triage-security/cases/036-valid/annotations.yaml b/evals/fullsend/triage-security/cases/036-valid/annotations.yaml new file mode 100644 index 000000000..31a24f9c9 --- /dev/null +++ b/evals/fullsend/triage-security/cases/036-valid/annotations.yaml @@ -0,0 +1,24 @@ +# SYNTHETIC TEST DATA — strict native execution assertions, never expected-behavior receipts +assertion_count: 7 +assertions: + - Native runtime records show actual sdlc-workflow:triage-security Skill invocation for TC-8101, delivered PR + plugin instructions and genuine subsequent Bash calls/results. Fixture/bootstrap/runtime errors, missing or + wrong-version Skill binding, narrated fallbacks, duplicated commands and schema-only sample output do not count. + - Actual probe and delivered Step0.6 show native nonempty FULLSEND_OUTPUT_DIR=/sandbox/workspace/output and sandbox + routing with exit0. Delivered Step0.7 actually prints Trusted triage-security input available and succeeds without + writing the abort object. + - Completed TC-8101 report grounds analysis in mounted release openssl-libs:3.0.7-1 at abcdef0 and development + openssl-libs:3.0.8-2 at main, names unavailable evidence honestly, and describes withheld assignment, field + edits, transition, comment, link and remediation task. Genuine analysis precedes output; the illustrative initial + result is insufficient. + - 'Actual completed agent-result.json satisfies delivered result schema: schema_version="1", mode="report-only", + evidence-backed report and exactly one report-only action. No mutation action is executed or serialized. Output/schema + validity alone is insufficient.' + - 'After writing that completed result, actual Bash executes the delivered inline final validator: json.load on + the native agent-result.json and jsonschema.validate against delivered plugin result schema, with exit0 and + Final Fullsend result validated. Narration, input validation, host validate-output-schema.sh or independent + sample validation cannot substitute.' + - Raw tools show no Jira/GitHub/WebFetch/CVE/lifecycle/Git/credential inspection, executed actions or interactive + fallback; analysis uses mounted trusted evidence only. + - Actual native output inventory is exactly [agent-result.json]. No sandbox logs/receipts/matrices/Markdown or + supplementary files are written; genuine Fullsend runtime artifacts are retained outside Skill output. diff --git a/evals/fullsend/triage-security/cases/036-valid/input.yaml b/evals/fullsend/triage-security/cases/036-valid/input.yaml new file mode 100644 index 000000000..fa355b6c2 --- /dev/null +++ b/evals/fullsend/triage-security/cases/036-valid/input.yaml @@ -0,0 +1,3 @@ +# SYNTHETIC TEST DATA — independent native Fullsend gate case +scenario: valid +issue: TC-8101 diff --git a/evals/fullsend/triage-security/eval.yaml b/evals/fullsend/triage-security/eval.yaml new file mode 100644 index 000000000..1fa29dead --- /dev/null +++ b/evals/fullsend/triage-security/eval.yaml @@ -0,0 +1,95 @@ +# SYNTHETIC TEST DATA — separate opaque CLI suite, not sdlc run-evals +name: fullsend-triage-security-gates +execution: + skill: triage-security-gate + mode: case + timeout: 2400 + parallelism: 1 +runner: + type: cli + effort: high + command: + - '{python}' + - '{runner}' + - --agent + - '{agent}' + - --workspace + - '{workspace}' + - --output-dir + - '{output_dir}' + - --scenario + - '{scenario}' + - --model + - '{model}' + - --effort + - '{effort}' +models: + skill: claude-opus-4-8 + judge: claude-opus-4-6 +dataset: + path: cases +outputs: + - path: output +traces: + stdout: true + stderr: true + events: false + metrics: true +judges: + - name: assertion_1 + if: annotations.get("assertion_count", 0) > 0 + prompt_file: evals/fullsend/triage-security/judge.md + arguments: + assertion_index: 0 + feedback_type: bool + - name: assertion_2 + if: annotations.get("assertion_count", 0) > 1 + prompt_file: evals/fullsend/triage-security/judge.md + arguments: + assertion_index: 1 + feedback_type: bool + - name: assertion_3 + if: annotations.get("assertion_count", 0) > 2 + prompt_file: evals/fullsend/triage-security/judge.md + arguments: + assertion_index: 2 + feedback_type: bool + - name: assertion_4 + if: annotations.get("assertion_count", 0) > 3 + prompt_file: evals/fullsend/triage-security/judge.md + arguments: + assertion_index: 3 + feedback_type: bool + - name: assertion_5 + if: annotations.get("assertion_count", 0) > 4 + prompt_file: evals/fullsend/triage-security/judge.md + arguments: + assertion_index: 4 + feedback_type: bool + - name: assertion_6 + if: annotations.get("assertion_count", 0) > 5 + prompt_file: evals/fullsend/triage-security/judge.md + arguments: + assertion_index: 5 + feedback_type: bool + - name: assertion_7 + if: annotations.get("assertion_count", 0) > 6 + prompt_file: evals/fullsend/triage-security/judge.md + arguments: + assertion_index: 6 + feedback_type: bool +thresholds: + assertion_1: + min_pass_rate: 1.0 + assertion_2: + min_pass_rate: 1.0 + assertion_3: + min_pass_rate: 1.0 + assertion_4: + min_pass_rate: 1.0 + assertion_5: + min_pass_rate: 1.0 + assertion_6: + min_pass_rate: 1.0 + assertion_7: + min_pass_rate: 1.0 diff --git a/evals/fullsend/triage-security/harness.yaml b/evals/fullsend/triage-security/harness.yaml new file mode 100644 index 000000000..51080cc01 --- /dev/null +++ b/evals/fullsend/triage-security/harness.yaml @@ -0,0 +1,44 @@ +# SYNTHETIC TEST DATA — native test agent; production policy/provider/profile/schema unchanged +role: triage +agent: evals/fullsend/triage-security/agent.md +model: claude-opus-4-8 +effort: high +image: ghcr.io/fullsend-ai/fullsend-code@sha256:9743bc7b6e451e0bcea25ae4a67e0c040c296f1fee04c08988ae80c53fafcfe6 +readonly_repo: true +plugins: + - plugins/sdlc-workflow +policy: plugins/sdlc-workflow/policies/triage-security.yaml +providers: + - plugins/sdlc-workflow/providers/vertex-ai.yaml +openshell: + profiles: + - plugins/sdlc-workflow/profiles/fullsend-vertex-ai.yaml +host_files: + - src: plugins/sdlc-workflow/env/gcp-vertex.env + dest: /sandbox/workspace/.env.d/gcp-vertex.env + expand: true + # Generated by the test pre-script after early environment/file validation. + # Fullsend resolves these against native-config; no host run-dir env is needed. + - src: pre/triage-security-input.json + dest: /sandbox/workspace/.pre-script/triage-security-input.json + optional: true + - src: pre/tc-6677-gate.env + dest: /sandbox/workspace/.env.d/zz-tc-6677-gate.env + optional: true + # Operator-provided inference credential; native upload only, never staged. + - src: ${GOOGLE_APPLICATION_CREDENTIALS} + dest: /tmp/.gcp-credentials.json +pre_script: evals/fullsend/triage-security/prepare-fixture.py +validation_loop: + script: plugins/sdlc-workflow/scripts/validate-output-schema.sh + schema: plugins/sdlc-workflow/schemas/triage-security-result.schema.json + max_iterations: 1 +env: + runner: + PYTHONDONTWRITEBYTECODE: '1' + TC6677_SCENARIO: ${TC6677_SCENARIO} + TC6677_REPO_ROOT: ${TC6677_REPO_ROOT} + FULLSEND_OUTPUT_SCHEMA: ${TC6677_REPO_ROOT}/plugins/sdlc-workflow/schemas/triage-security-result.schema.json + FULLSEND_OUTPUT_FILE: agent-result.json +timeout_minutes: 30 +version: 1 diff --git a/evals/fullsend/triage-security/judge.md b/evals/fullsend/triage-security/judge.md new file mode 100644 index 000000000..39bc5c8dd --- /dev/null +++ b/evals/fullsend/triage-security/judge.md @@ -0,0 +1,39 @@ + + +Judge only this assertion: +{{ annotations.assertions[arguments.assertion_index] }} + +Case input: {{ inputs }} +Raw case record, actual stdout/stderr/exit and collected native files: {{ outputs }} + +Return a boolean using the upstream judge response contract. PASS requires +genuine native runtime evidence, never an expected-behavior story or a manually +constructed JSON sample. For EVERY assertion, require this case's actual Skill +tool invocation of sdlc-workflow:triage-security, delivered instructions containing +the evaluated Fullsend presence gate and inline final validator, actual plugin +binding to the supplied repository's plugin, and real subsequent tool calls/results +matching those instructions. Missing/wrong-version/manual-read Skill execution +fails even if an output matches an expected value. Locate genuine native Claude +transcripts/logs under output/native; framework stdout alone may be insufficient. +Do not infer tool success from literal source or an agent's claims. Fixture +preparation, sandbox/bootstrap/inference failures, missing/truncated decisive +tool results or timeouts must not masquerade as expected Skill rejection. + +Judge subsequent calls/results up to the case's prescribed stop, not completion +of the entire triage workflow. In 033-absent, the real presence-gate Bash result +is interactive mode; reading CLAUDE.md and stopping for missing Security +Configuration is expected. In 034-empty, the real presence-gate Bash result +exits 1 with ERROR: FULLSEND_OUTPUT_DIR is set but empty; stopping immediately +is expected. These intentionally supplied environment states are case inputs, +not fixture/bootstrap/inference failures. Neither stop requires later triage, +input validation, a result write or execution of the final validator. Both still +require the actual Skill invocation, supplied plugin binding, delivered gate +and validator instructions, and genuine gate tool calls/results described above. + +Distinguish native fullsend CLI exit, the Skill's Bash tool exit, and the host +validation loop. Expected negative-case CLI/schema failure is not itself proof +of correct gate/input behavior. The valid case requires the actual Skill inline +validator after the completed result write, not merely host validate-output-schema.sh +or standalone validation of a sample. Inspect actual output file inventory and +raw tools to establish that no prohibited sandbox files or external calls occurred. +Opaque CLI metrics may lack cost_usd; that is not a runtime success signal. diff --git a/evals/fullsend/triage-security/prepare-fixture.py b/evals/fullsend/triage-security/prepare-fixture.py new file mode 100755 index 000000000..68ad803cc --- /dev/null +++ b/evals/fullsend/triage-security/prepare-fixture.py @@ -0,0 +1,45 @@ +#!/usr/bin/env python3 +"""SYNTHETIC TEST DATA — native pre-script; no live prefetch or triage logic.""" + +import os +from pathlib import Path +import sys + + +def prepare(root, scenario): + """Prepare only host mounts; Fullsend sources the fragment before inference.""" + if scenario not in {"absent", "empty", "malformed", "valid"}: + raise ValueError("Unknown synthetic gate scenario") + fixtures = root / "evals/triage-security/files" + pre = root / "pre" + # This is the case-private native-config root, not Fullsend's output run dir. + # Native resolution binds the generated mounts here; never reuse old mounts. + if pre.exists() and any(pre.iterdir()): + raise ValueError("Refusing stale pre-script fixture files") + pre.mkdir(parents=True, exist_ok=True) + fragment = "# SYNTHETIC TEST DATA — deliberate native gate condition injection\n" + if scenario == "absent": + fragment += "unset FULLSEND_OUTPUT_DIR\n" + elif scenario == "empty": + fragment += "export FULLSEND_OUTPUT_DIR=''\n" + if scenario in {"malformed", "valid"}: + name = "fullsend-invalid-trusted-input.md" if scenario == "malformed" else "fullsend-report-only-trusted-input.json" + data = (fixtures / name).read_bytes() + if scenario == "malformed": + data = data.split(b"```json\n", 1)[1].split(b"\n```", 1)[0] + (pre / "triage-security-input.json").write_bytes(data) + (pre / "tc-6677-gate.env").write_text(fragment) + + +def main(): + """Consume the native trusted pre-script environment without exposing it.""" + try: + prepare(Path(os.environ["TC6677_REPO_ROOT"]), os.environ["TC6677_SCENARIO"]) + except (KeyError, OSError, ValueError, IndexError) as exc: + print(f"Synthetic gate fixture preparation failed: {exc}", file=sys.stderr) + return 1 + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/evals/fullsend/triage-security/run-fullsend.py b/evals/fullsend/triage-security/run-fullsend.py new file mode 100644 index 000000000..fa50ae1ca --- /dev/null +++ b/evals/fullsend/triage-security/run-fullsend.py @@ -0,0 +1,135 @@ +#!/usr/bin/env python3 +"""Opaque CLI adapter: delegate one synthetic case to native Fullsend unchanged.""" + +import argparse +import os +from pathlib import Path +import shutil +import signal +import subprocess +import sys + +import yaml + + +def run_case(root, workspace, output, scenario, model, effort, host_binary, sandbox_binary, plugin_root=None): + """Stage test-only resources, call native CLI once, preserve its actual exit.""" + if scenario not in {"absent", "empty", "malformed", "valid"}: + raise ValueError("Unknown synthetic gate scenario") + if os.environ.get("FULLSEND_MINT_URL"): + raise ValueError("Synthetic suite requires FULLSEND_MINT_URL unset; live forge minting is forbidden") + setup = workspace / "native-config" + target = workspace / "synthetic-target" + native = output / "native" + if any(p.exists() for p in [setup, target, native, output / "metrics.json"]): + raise ValueError("Refusing stale native configuration or output") + suite = root / "evals/fullsend/triage-security" + h = yaml.safe_load((suite / "harness.yaml").read_text()) + # PR content is only uploaded as a sandbox plugin. Trusted host executables, + # policy, providers and schema remain independent even if the PR replaces them. + plugin_root = plugin_root or root / "plugins/sdlc-workflow" + plugin_root = plugin_root.absolute() + if (any(p.is_symlink() for p in [plugin_root, *plugin_root.parents]) + or any(p.is_symlink() for p in plugin_root.rglob("*"))): + raise ValueError("Tested plugin must contain regular files, not symlinks") + plugin_destination = setup / ("tested-plugin" if plugin_root != root / "plugins/sdlc-workflow" else "plugins/sdlc-workflow") + shutil.copytree(plugin_root, plugin_destination, + ignore=shutil.ignore_patterns("__pycache__", "*.pyc")) + if plugin_destination == setup / "tested-plugin": + for name in ["policies/triage-security.yaml", "providers/vertex-ai.yaml", "profiles/fullsend-vertex-ai.yaml", + "env/gcp-vertex.env", "schemas/triage-security-result.schema.json", + "scripts/validate-output-schema.sh", "scripts/strip_extra_properties.py"]: + destination = setup / "plugins/sdlc-workflow" / name + destination.parent.mkdir(parents=True, exist_ok=True) + shutil.copy2(root / "plugins/sdlc-workflow" / name, destination) + staged_suite = setup / "evals/fullsend/triage-security" + staged_suite.mkdir(parents=True) + for name in ["agent.md", "prepare-fixture.py"]: + shutil.copy2(suite / name, staged_suite / name) + staged_fixtures = setup / "evals/triage-security/files" + staged_fixtures.mkdir(parents=True) + for name in ["fullsend-gate-interactive-config.md", "fullsend-invalid-trusted-input.md", + "fullsend-report-only-trusted-input.json"]: + shutil.copy2(root / "evals/triage-security/files" / name, staged_fixtures / name) + for field in ["agent", "policy", "pre_script"]: + h[field] = str(setup / h[field]) + h["plugins"] = [str(plugin_destination)] + for field in ["providers"]: + h[field] = [str(setup / p) for p in h[field]] + h["openshell"]["profiles"] = [str(setup / p) for p in h["openshell"]["profiles"]] + for field in ["script", "schema"]: + h["validation_loop"][field] = str(setup / h["validation_loop"][field]) + h["host_files"][0]["src"] = str(setup / h["host_files"][0]["src"]) + if os.environ.get("GCP_OIDC_TOKEN_FILE"): + token = Path(os.environ["GCP_OIDC_TOKEN_FILE"]) + if not token.is_file(): + raise ValueError("Missing prepared sandbox OIDC token") + h["host_files"].append({"src": str(token), "dest": "/sandbox/workspace/.gcp-oidc-token"}) + h["env"]["runner"]["TC6677_SCENARIO"] = scenario + h["env"]["runner"]["TC6677_REPO_ROOT"] = str(setup) + h["env"]["runner"]["FULLSEND_OUTPUT_SCHEMA"] = str(setup / "plugins/sdlc-workflow/schemas/triage-security-result.schema.json") + (setup / "harness").mkdir(parents=True) + (setup / "harness/triage-security-gate.yaml").write_text(yaml.safe_dump(h, sort_keys=False)) + (setup / "config.yaml").write_text(yaml.safe_dump({ + "version": "1", "runtime": "claude", + "agents": [{"source": "harness/triage-security-gate.yaml"}], + }, sort_keys=False)) + target.mkdir() + # Native UploadDir retains .git; read-only setup requires it, not a commit. + subprocess.run(["git", "init", "--quiet", str(target)], check=True) + if scenario == "absent": + shutil.copy2(root / "evals/triage-security/files/fullsend-gate-interactive-config.md", target / "CLAUDE.md") + output.mkdir(parents=True, exist_ok=True) + native.mkdir() + command = [str(host_binary), "run", "triage-security-gate", "--fullsend-dir", str(setup), + "--target-repo", str(target), "--output-dir", str(native), + "--fullsend-binary", str(sandbox_binary), "--runtime", "claude", + "--model", model, "--effort", effort, "--no-post-script"] + # Stdout/stderr go directly to CliRunner; no nested inference or env rewriting. + environment = dict(os.environ) + if os.environ.get("TC6726_SANDBOX_CREDENTIALS"): + credential = Path(os.environ["TC6726_SANDBOX_CREDENTIALS"]) + if not credential.is_file(): + raise ValueError("Missing prepared sandbox credential") + environment["GOOGLE_APPLICATION_CREDENTIALS"] = str(credential) + # Judge processes keep original host ADC; only native Fullsend gets prepared + # sandbox ADC. Its reserved OIDC refresh variables remain host-only upstream. + process = subprocess.run(command, cwd=workspace, env=environment, check=False) + metrics = list(native.glob("agent-*/metrics.json")) + if len(metrics) == 1: + shutil.copy2(metrics[0], output / "metrics.json") + elif len(metrics) > 1: + print("Ambiguous native metrics; originals retained, no root copy selected", file=sys.stderr) + return process.returncode + + +def main(): + """Accept the upstream opaque CLI placeholders, not fixture-generated code.""" + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--agent", choices=["triage-security-gate"], required=True) + parser.add_argument("--workspace", type=Path, required=True) + parser.add_argument("--output-dir", type=Path, required=True) + parser.add_argument("--scenario", choices=["absent", "empty", "malformed", "valid"], required=True) + parser.add_argument("--model", required=True) + parser.add_argument("--effort", required=True) + parser.add_argument("--plugin-root", type=Path) + args = parser.parse_args() + try: + return run_case(Path(__file__).resolve().parents[3], args.workspace.resolve(), + args.output_dir.resolve(), args.scenario, args.model, args.effort, + Path(os.environ["TC6677_FULLSEND_BIN"]), + Path(os.environ["TC6677_SANDBOX_FULLSEND_BIN"]), + args.plugin_root.absolute() if args.plugin_root else None) + except (KeyError, OSError, ValueError) as exc: + print(f"Native gate eval fixture/CLI failure: {exc}", file=sys.stderr) + return 1 + + +if __name__ == "__main__": + status = main() + if status < 0: + # Python sys.exit(-signal) wraps it modulo256; preserve the native signal. + if -status not in {signal.SIGKILL, signal.SIGSTOP}: + signal.signal(-status, signal.SIG_DFL) + os.kill(os.getpid(), -status) + sys.exit(status) diff --git a/evals/triage-security/evals.json b/evals/triage-security/evals.json index c6b79fe05..be04e0d17 100644 --- a/evals/triage-security/evals.json +++ b/evals/triage-security/evals.json @@ -320,8 +320,8 @@ }, { "id": 19, - "prompt": "Triage Vulnerability issue TC-8001. The issue details are in vuln-issue-standard.md, the security matrix is in security-matrix-stale-mock.md (which has a Last-Updated timestamp of 2026-05-01T10:00:00Z — more than 14 days old), and the project CLAUDE.md is in claude-md-security-config.md. Do NOT actually call Jira MCP, git show, or any external tools. Instead, write your triage analysis to the workspace outputs/ directory: write outputs/staleness-check.md with the Step 0.3 staleness detection output including the warning message and user options.", - "expected_output": "A staleness check that detects the security-matrix.md timestamp (2026-05-01T10:00:00Z) is older than the 14-day default threshold and presents a warning with three options: refresh now, proceed anyway, or stop.", + "prompt": "Evaluation clock: 2026-07-12T10:00:00Z (fixed UTC scenario time for execution and grading). Use this as the current date for Step 0.3, regardless of the runner or grader wall clock. Triage Vulnerability issue TC-8001. The issue details are in vuln-issue-standard.md, the security matrix is in security-matrix-stale-mock.md (which has a Last-Updated timestamp of 2026-05-01T10:00:00Z — 72 days old at the evaluation clock), and the project CLAUDE.md is in claude-md-security-config.md. Do NOT actually call Jira MCP, git show, or any external tools. Instead, write your triage analysis to the workspace outputs/ directory: write outputs/staleness-check.md with the Step 0.3 staleness detection output including the warning message and user options. Record the evaluation clock and computed matrix age in outputs/staleness-check.md as fixture evidence; this does not introduce a user-facing warning for a fresh matrix. Record the source Last-Updated HTML comment, the parsed ISO 8601 timestamp, and the comparison clock in outputs/staleness-check.md as extraction evidence.", + "expected_output": "Evaluation clock: 2026-07-12T10:00:00Z (fixed UTC scenario time for execution and grading). Use this as the current date for Step 0.3, regardless of the runner or grader wall clock. A staleness check that detects the security-matrix.md timestamp (2026-05-01T10:00:00Z) is older than the 14-day default threshold and presents a warning with three options: refresh now, proceed anyway, or stop. The matrix is 72 days old at the evaluation clock.", "files": [ "files/vuln-issue-standard.md", "files/security-matrix-stale-mock.md", @@ -337,8 +337,8 @@ }, { "id": 20, - "prompt": "Triage Vulnerability issue TC-8001. The issue details are in vuln-issue-standard.md, the security matrix is in security-matrix-mock.md (which has a recent Last-Updated timestamp within the 14-day threshold), and the project CLAUDE.md is in claude-md-security-config.md. For all time-based calculations in this evaluation, treat today's date as 2026-06-29T10:00:00Z. Do NOT actually call Jira MCP, git show, or any external tools. Instead, write your triage analysis to the workspace outputs/ directory: write outputs/staleness-check.md documenting the Step 0.3 result, and write outputs/data-extraction.md with the parsed CVE data table from Step 1.", - "expected_output": "Step 0.3 reads the Last-Updated timestamp, finds it within the 14-day threshold, and proceeds silently to the next step without any warning or user prompt.", + "prompt": "Evaluation clock: 2026-07-12T10:00:00Z (fixed UTC scenario time for execution and grading). Use this as the current date for Step 0.3, regardless of the runner or grader wall clock. Triage Vulnerability issue TC-8001. The issue details are in vuln-issue-standard.md, the security matrix is in security-matrix-mock.md (which has a Last-Updated timestamp of 2026-06-28T10:00:00Z — exactly 14 days old at the evaluation clock, within the threshold), and the project CLAUDE.md is in claude-md-security-config.md. Do NOT actually call Jira MCP, git show, or any external tools. Instead, write your triage analysis to the workspace outputs/ directory: write outputs/staleness-check.md documenting the Step 0.3 result, and write outputs/data-extraction.md with the parsed CVE data table from Step 1. Record the evaluation clock and computed matrix age in outputs/staleness-check.md as fixture evidence; this does not introduce a user-facing warning for a fresh matrix.", + "expected_output": "Evaluation clock: 2026-07-12T10:00:00Z (fixed UTC scenario time for execution and grading). Use this as the current date for Step 0.3, regardless of the runner or grader wall clock. Step 0.3 reads the Last-Updated timestamp, finds it within the 14-day threshold, and proceeds silently to the next step without any warning or user prompt. The matrix is exactly 14 days old at the evaluation clock; the strict older-than-14-day rule must not warn at equality.", "files": [ "files/vuln-issue-standard.md", "files/security-matrix-mock.md", diff --git a/evals/triage-security/files/fullsend-authorized-trusted-bundle.md b/evals/triage-security/files/fullsend-authorized-trusted-bundle.md new file mode 100644 index 000000000..9f5ee1aec --- /dev/null +++ b/evals/triage-security/files/fullsend-authorized-trusted-bundle.md @@ -0,0 +1,19 @@ + + +# Trusted bundle facts + +The mounted `triage-security-input.json` is already schema-valid. It contains +only the following evidence needed for this case: + +- Issue `TC-8100`; authorization `mutation_authorized: true`. +- Trusted assignee account ID `557058:fullsend-engineer`. +- Affects Versions `2.2.0`, VEX Justification `Affected`, and new label + `security-triaged`. +- Transition `In Progress` and one triage comment represented as valid ADF. +- A remediation task for `TC-8100` with stable ref `remediation-8100`; its + existing Jira task has no description digest yet. +- A Depend link from `TC-8100` to `{{remediation-8100.key}}`. +- Existing idempotency marker `triage-security:tc-8100:field-edit:label:security-triaged`. + +No other files, credentials, or external evidence are available. All supplied +values must be used as-is; no account, version, or reference lookup is needed. diff --git a/evals/triage-security/files/fullsend-authorized-trusted-input.json b/evals/triage-security/files/fullsend-authorized-trusted-input.json new file mode 100644 index 000000000..a55fed10f --- /dev/null +++ b/evals/triage-security/files/fullsend-authorized-trusted-input.json @@ -0,0 +1,12 @@ +{ + "schema_version": "1", + "issue": {"key": "TC-8100", "summary": "CVE-2026-8100", "description": {}, "status": "New", "labels": [], "versions": [], "reporter": {"account_id": "reporter-1", "display_name": "Reporter"}, "comments": [], "fields": {"fullsend_actions": {"assignee_account_id": "557058:fullsend-engineer", "affects_versions": ["2.2.0"], "vex_justification": "Affected", "label": "security-triaged", "transition": "In Progress", "remediation_ref": "remediation-8100"}}}, + "remote_links": [{"url": "https://example.com/evidence", "title": "Evidence"}], + "configuration": {"project_key": "TC", "jira_version_prefix": "RHTPA", "vulnerability_issue_type_id": "10016", "component_label_pattern": "pscomponent:", "version_streams": [{"name": "2.2.x", "matrix_path": "security-matrix.md", "release_repository": "release-repo"}], "source_repositories": [{"name": "release-repo", "url": "https://github.com/example/release-repo", "deployment_context": "internal"}]}, + "external_evidence": {"mitre": {"source_url": "https://example.com/mitre", "retrieved_at": "2026-09-21T12:00:00Z", "status": 200, "body": {}}, "osv": {"source_url": "https://example.com/osv", "retrieved_at": "2026-09-21T12:00:00Z", "status": 200, "body": {}}, "lifecycle": {"source_url": "https://example.com/lifecycle", "retrieved_at": "2026-09-21T12:00:00Z", "status": 200, "body": {}}}, + "matrix": {"streams": [{"name": "2.2.x", "matrix_source": "trusted", "rows": [{"version": "2.2.0", "source_commits": {"release-repo": "abcdef0"}, "retag_of": null}]}]}, + "source_evidence": {"lock_files": [{"repository": "release-repo", "ref": "abcdef0", "path": "rpms.lock.yaml", "command": "git show abcdef0:rpms.lock.yaml", "content": "openssl-libs: 3.0.7-1"}], "development_streams": [{"repository": "release-repo", "ref": "main", "path": "rpms.lock.yaml", "command": "git show main:rpms.lock.yaml", "content": "openssl-libs: 3.0.8-2"}]}, + "jira_metadata": {"versions": [], "sibling_searches": [], "related_issues": []}, + "idempotency": {"action_markers": ["triage-security:tc-8100:field-edit:label:security-triaged"], "existing_remediation": [{"key": "TC-9100", "summary": "Remediation task for TC-8100", "status": "In Progress", "labels": ["ai-generated-jira"], "description": {}, "comments": [], "links": []}]}, + "authorization": {"mutation_authorized": true} +} diff --git a/evals/triage-security/files/fullsend-eval-1-trusted-input.json b/evals/triage-security/files/fullsend-eval-1-trusted-input.json new file mode 100644 index 000000000..bbdd40824 --- /dev/null +++ b/evals/triage-security/files/fullsend-eval-1-trusted-input.json @@ -0,0 +1,336 @@ +{ + "schema_version": "1", + "issue": { + "key": "TC-8001", + "summary": "CVE-2026-31812 quinn-proto - Panic on large stream counts [rhtpa-2.2]", + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "A vulnerability was found in quinn-proto. The quinn-proto crate before version 0.11.14 allows a remote attacker to cause a panic by sending a QUIC transport frame that creates an excessive number of streams. This vulnerability is classified as a denial of service (DoS).\n\n**Affected package**: quinn-proto\n**Affected versions**: versions before 0.11.14\n**Fixed version**: 0.11.14\n**CVSS**: 7.5 (High)\n\nThe vulnerability exists because quinn-proto does not properly validate the number of streams requested in a STREAMS frame. An attacker can send a specially crafted frame that causes the server to allocate an unbounded number of stream state objects, leading to a panic when the allocation exceeds internal limits.\n\n### References\n\n- https://github.com/advisories/GHSA-2026-qp73-x4mq\n- https://rustsec.org/advisories/RUSTSEC-2026-0042.html" + } + ] + } + ] + }, + "status": "New", + "labels": [ + "CVE-2026-31812", + "pscomponent:org/rhtpa-server" + ], + "versions": [], + "reporter": { + "account_id": "synthetic-reporter", + "display_name": "Synthetic Reporter" + }, + "comments": [], + "fields": { + "fixture_purpose": "SYNTHETIC TEST DATA — independent Fullsend counterpart of interactive eval 1; all identities, URLs, CVEs, commit pins and evidence are deliberate public test material, not live advisory evidence.", + "current_user": { + "accountId": "synthetic-engineer", + "displayName": "Synthetic Engineer" + }, + "ecosystem": "Cargo", + "affected_package": "quinn-proto", + "fixed_version": "0.11.14", + "affected_range": "< 0.11.14", + "affectsVersions": [ + { + "id": "version-stale", + "name": "RHTPA 2.0.0" + } + ], + "assignee": null, + "issuelinks": [], + "customfield_10632": "quinn-proto" + } + }, + "remote_links": [ + { + "url": "https://example.com/advisories/CVE-2026-31812", + "title": "Synthetic CVE-2026-31812 advisory" + } + ], + "configuration": { + "project_key": "TC", + "jira_version_prefix": "RHTPA", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "version_streams": [ + { + "name": "2.1.x", + "matrix_path": "2.1.x/security-matrix.md", + "release_repository": "release-repo" + }, + { + "name": "2.2.x", + "matrix_path": "2.2.x/security-matrix.md", + "release_repository": "release-repo" + } + ], + "source_repositories": [ + { + "name": "backend", + "url": "https://github.com/example/backend", + "deployment_context": "customer-shipped" + } + ], + "vex_justification_field": "customfield_10665", + "upstream_affected_component_field": "customfield_10632", + "stream_field": "customfield_10832" + }, + "external_evidence": { + "mitre": { + "source_url": "https://example.com/mitre/CVE-2026-31812", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "cveMetadata": { + "cveId": "CVE-2026-31812" + }, + "containers": { + "cna": { + "affected": [ + { + "product": "quinn-proto", + "versions": [ + { + "status": "affected", + "lessThan": "0.11.14", + "versionType": "semver" + } + ] + } + ] + } + } + } + }, + "osv": { + "source_url": "https://example.com/osv/CVE-2026-31812", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "id": "CVE-2026-31812", + "affected": [ + { + "package": { + "ecosystem": "Cargo", + "name": "quinn-proto" + }, + "ranges": [ + { + "type": "SEMVER", + "events": [ + { + "introduced": "0" + }, + { + "fixed": "0.11.14" + } + ] + } + ] + } + ] + } + }, + "lifecycle": { + "source_url": "https://example.com/lifecycle", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "supported_streams": [ + "2.1.x", + "2.2.x" + ], + "eol_streams": [] + } + } + }, + "matrix": { + "streams": [ + { + "name": "2.1.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.1.x matrix snapshot", + "rows": [ + { + "version": "2.1.0", + "source_commits": { + "backend": "a003008" + }, + "retag_of": null + }, + { + "version": "2.1.1", + "source_commits": { + "backend": "a003012" + }, + "retag_of": null + } + ] + }, + { + "name": "2.2.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.2.x matrix snapshot", + "rows": [ + { + "version": "2.2.0", + "source_commits": { + "backend": "a004005" + }, + "retag_of": null + }, + { + "version": "2.2.1", + "source_commits": { + "backend": "a004008" + }, + "retag_of": null + }, + { + "version": "2.2.2", + "source_commits": { + "backend": "a004008" + }, + "retag_of": "2.2.1" + }, + { + "version": "2.2.3", + "source_commits": { + "backend": "a004011" + }, + "retag_of": null + }, + { + "version": "2.2.4", + "source_commits": { + "backend": "a004012" + }, + "retag_of": null + } + ] + } + ] + }, + "source_evidence": { + "lock_files": [ + { + "repository": "backend", + "ref": "a003008", + "path": "Cargo.lock", + "command": "git show a003008:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + }, + { + "repository": "backend", + "ref": "a003012", + "path": "Cargo.lock", + "command": "git show a003012:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + }, + { + "repository": "backend", + "ref": "a004005", + "path": "Cargo.lock", + "command": "git show a004005:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.9\"\n" + }, + { + "repository": "backend", + "ref": "a004008", + "path": "Cargo.lock", + "command": "git show a004008:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.12\"\n" + }, + { + "repository": "backend", + "ref": "a004011", + "path": "Cargo.lock", + "command": "git show a004011:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + }, + { + "repository": "backend", + "ref": "a004012", + "path": "Cargo.lock", + "command": "git show a004012:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + } + ], + "development_streams": [ + { + "repository": "backend", + "ref": "release/0.3.z", + "path": "Cargo.lock", + "command": "git show release/0.3.z:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + }, + { + "repository": "backend", + "ref": "release/0.4.z", + "path": "Cargo.lock", + "command": "git show release/0.4.z:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.12\"\n" + } + ] + }, + "jira_metadata": { + "versions": [ + { + "id": "version-2-1-0", + "name": "RHTPA 2.1.0", + "released": true + }, + { + "id": "version-2-1-1", + "name": "RHTPA 2.1.1", + "released": true + }, + { + "id": "version-2-2-0", + "name": "RHTPA 2.2.0", + "released": true + }, + { + "id": "version-2-2-1", + "name": "RHTPA 2.2.1", + "released": true + }, + { + "id": "version-2-2-2", + "name": "RHTPA 2.2.2", + "released": true + }, + { + "id": "version-2-2-3", + "name": "RHTPA 2.2.3", + "released": true + }, + { + "id": "version-2-2-4", + "name": "RHTPA 2.2.4", + "released": true + }, + { + "id": "version-dev", + "name": "RHTPA 2.2.5", + "released": false + } + ], + "sibling_searches": [], + "related_issues": [] + }, + "idempotency": { + "action_markers": [], + "existing_remediation": [] + }, + "authorization": { + "mutation_authorized": true + } +} diff --git a/evals/triage-security/files/fullsend-eval-11-trusted-input.json b/evals/triage-security/files/fullsend-eval-11-trusted-input.json new file mode 100644 index 000000000..45842db53 --- /dev/null +++ b/evals/triage-security/files/fullsend-eval-11-trusted-input.json @@ -0,0 +1,379 @@ +{ + "schema_version": "1", + "issue": { + "key": "TC-8021", + "summary": "CVE-2026-55123 tokio - Use-after-free in task abort [rhtpa-2.1]", + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "A vulnerability was found in the tokio crate. Versions of tokio before 1.42.0 are vulnerable to a use-after-free when a spawned task is aborted while holding a borrowed reference. This can lead to memory corruption and potential code execution.\n\n**Affected package**: tokio\n**Affected versions**: versions before 1.42.0\n**Fixed version**: 1.42.0\n**CVSS**: 8.1 (High)\n\nThis issue is scoped to stream [rhtpa-2.1]. A proactive remediation task (TC-8022) already exists for this stream, created by a prior cross-stream triage of TC-8020 (stream [rhtpa-2.2]).\n\n### Existing preemptive task (provided by Step 4.4 JQL search)\n\nA JQL search for `labels = 'security-preemptive' AND labels = 'CVE-2026-55123'` returns:\n\n- **TC-8022** — Remediate CVE-2026-55123: bump tokio to 1.42.0 (rhtpa-2.1)\n - **Status**: Open\n - **Labels**: ai-generated-jira, Security, CVE-2026-55123, security-preemptive\n - **Issue Links**:\n - **Related**: TC-8020 (originating CVE Jira, stream [rhtpa-2.2])\n\n### References\n\n- https://github.com/advisories/GHSA-2026-tk91-v5pp\n- https://rustsec.org/advisories/RUSTSEC-2026-0088.html" + } + ] + } + ] + }, + "status": "New", + "labels": [ + "CVE-2026-55123", + "pscomponent:org/rhtpa-server" + ], + "versions": [], + "reporter": { + "account_id": "synthetic-reporter", + "display_name": "Synthetic Reporter" + }, + "comments": [], + "fields": { + "fixture_purpose": "SYNTHETIC TEST DATA — independent Fullsend counterpart of interactive eval 11; all identities, URLs, CVEs, commit pins and evidence are deliberate public test material, not live advisory evidence.", + "current_user": { + "accountId": "synthetic-engineer", + "displayName": "Synthetic Engineer" + }, + "ecosystem": "Cargo", + "affected_package": "tokio", + "fixed_version": "1.42.0", + "affected_range": "< 1.42.0", + "affectsVersions": [ + { + "id": "version-stale", + "name": "RHTPA 2.0.0" + } + ], + "assignee": null, + "issuelinks": [], + "customfield_10632": "tokio" + } + }, + "remote_links": [ + { + "url": "https://example.com/advisories/CVE-2026-55123", + "title": "Synthetic CVE-2026-55123 advisory" + } + ], + "configuration": { + "project_key": "TC", + "jira_version_prefix": "RHTPA", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "version_streams": [ + { + "name": "2.1.x", + "matrix_path": "2.1.x/security-matrix.md", + "release_repository": "release-repo" + }, + { + "name": "2.2.x", + "matrix_path": "2.2.x/security-matrix.md", + "release_repository": "release-repo" + } + ], + "source_repositories": [ + { + "name": "backend", + "url": "https://github.com/example/backend", + "deployment_context": "customer-shipped" + } + ], + "vex_justification_field": "customfield_10665", + "upstream_affected_component_field": "customfield_10632", + "stream_field": "customfield_10832" + }, + "external_evidence": { + "mitre": { + "source_url": "https://example.com/mitre/CVE-2026-55123", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "cveMetadata": { + "cveId": "CVE-2026-55123" + }, + "containers": { + "cna": { + "affected": [ + { + "product": "tokio", + "versions": [ + { + "status": "affected", + "lessThan": "1.42.0", + "versionType": "semver" + } + ] + } + ] + } + } + } + }, + "osv": { + "source_url": "https://example.com/osv/CVE-2026-55123", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "id": "CVE-2026-55123", + "affected": [ + { + "package": { + "ecosystem": "Cargo", + "name": "tokio" + }, + "ranges": [ + { + "type": "SEMVER", + "events": [ + { + "introduced": "0" + }, + { + "fixed": "1.42.0" + } + ] + } + ] + } + ] + } + }, + "lifecycle": { + "source_url": "https://example.com/lifecycle", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "supported_streams": [ + "2.1.x", + "2.2.x" + ], + "eol_streams": [] + } + } + }, + "matrix": { + "streams": [ + { + "name": "2.1.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.1.x matrix snapshot", + "rows": [ + { + "version": "2.1.0", + "source_commits": { + "backend": "a003008" + }, + "retag_of": null + }, + { + "version": "2.1.1", + "source_commits": { + "backend": "a003012" + }, + "retag_of": null + } + ] + }, + { + "name": "2.2.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.2.x matrix snapshot", + "rows": [ + { + "version": "2.2.0", + "source_commits": { + "backend": "a004005" + }, + "retag_of": null + }, + { + "version": "2.2.1", + "source_commits": { + "backend": "a004008" + }, + "retag_of": null + }, + { + "version": "2.2.2", + "source_commits": { + "backend": "a004008" + }, + "retag_of": "2.2.1" + }, + { + "version": "2.2.3", + "source_commits": { + "backend": "a004011" + }, + "retag_of": null + }, + { + "version": "2.2.4", + "source_commits": { + "backend": "a004012" + }, + "retag_of": null + } + ] + } + ] + }, + "source_evidence": { + "lock_files": [ + { + "repository": "backend", + "ref": "a003008", + "path": "Cargo.lock", + "command": "git show a003008:Cargo.lock", + "content": "[[package]]\nname = \"tokio\"\nversion = \"1.40.0\"\n" + }, + { + "repository": "backend", + "ref": "a003012", + "path": "Cargo.lock", + "command": "git show a003012:Cargo.lock", + "content": "[[package]]\nname = \"tokio\"\nversion = \"1.40.0\"\n" + }, + { + "repository": "backend", + "ref": "a004005", + "path": "Cargo.lock", + "command": "git show a004005:Cargo.lock", + "content": "[[package]]\nname = \"tokio\"\nversion = \"1.42.0\"\n" + }, + { + "repository": "backend", + "ref": "a004008", + "path": "Cargo.lock", + "command": "git show a004008:Cargo.lock", + "content": "[[package]]\nname = \"tokio\"\nversion = \"1.42.0\"\n" + }, + { + "repository": "backend", + "ref": "a004011", + "path": "Cargo.lock", + "command": "git show a004011:Cargo.lock", + "content": "[[package]]\nname = \"tokio\"\nversion = \"1.42.0\"\n" + }, + { + "repository": "backend", + "ref": "a004012", + "path": "Cargo.lock", + "command": "git show a004012:Cargo.lock", + "content": "[[package]]\nname = \"tokio\"\nversion = \"1.42.0\"\n" + } + ], + "development_streams": [ + { + "repository": "backend", + "ref": "release/0.3.z", + "path": "Cargo.lock", + "command": "git show release/0.3.z:Cargo.lock", + "content": "[[package]]\nname = \"tokio\"\nversion = \"1.42.0\"\n" + }, + { + "repository": "backend", + "ref": "release/0.4.z", + "path": "Cargo.lock", + "command": "git show release/0.4.z:Cargo.lock", + "content": "[[package]]\nname = \"tokio\"\nversion = \"1.42.0\"\n" + } + ] + }, + "jira_metadata": { + "versions": [ + { + "id": "version-2-1-0", + "name": "RHTPA 2.1.0", + "released": true + }, + { + "id": "version-2-1-1", + "name": "RHTPA 2.1.1", + "released": true + }, + { + "id": "version-2-2-0", + "name": "RHTPA 2.2.0", + "released": true + }, + { + "id": "version-2-2-1", + "name": "RHTPA 2.2.1", + "released": true + }, + { + "id": "version-2-2-2", + "name": "RHTPA 2.2.2", + "released": true + }, + { + "id": "version-2-2-3", + "name": "RHTPA 2.2.3", + "released": true + }, + { + "id": "version-2-2-4", + "name": "RHTPA 2.2.4", + "released": true + }, + { + "id": "version-dev", + "name": "RHTPA 2.2.5", + "released": false + } + ], + "sibling_searches": [ + { + "purpose": "preemptive-remediation", + "jql": "project = TC AND labels = \"security-preemptive\" AND labels = \"CVE-2026-55123\"", + "issues": [ + { + "key": "TC-8022", + "summary": "Remediate CVE-2026-55123: bump tokio to 1.42.0 (rhtpa-2.1)", + "status": "Open", + "labels": [ + "ai-generated-jira", + "Security", + "CVE-2026-55123", + "security-preemptive" + ], + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Prior cross-stream remediation from TC-8020." + } + ] + } + ] + }, + "comments": [], + "links": [ + { + "type": "Related", + "outward": "TC-8020" + } + ] + } + ] + } + ], + "related_issues": [] + }, + "idempotency": { + "action_markers": [ + "triage-security:tc-8021:assignment", + "triage-security:tc-8021:status:assigned" + ], + "existing_remediation": [] + }, + "authorization": { + "mutation_authorized": true + } +} diff --git a/evals/triage-security/files/fullsend-eval-12-trusted-input.json b/evals/triage-security/files/fullsend-eval-12-trusted-input.json new file mode 100644 index 000000000..7a74072dd --- /dev/null +++ b/evals/triage-security/files/fullsend-eval-12-trusted-input.json @@ -0,0 +1,313 @@ +{ + "schema_version": "1", + "issue": { + "key": "TC-8030", + "summary": "CVE-2026-48901 h2 - HTTP/2 CONTINUATION flood [rhtpa-2.2]", + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "A vulnerability was found in h2. The h2 crate is affected by an HTTP/2 CONTINUATION flood vulnerability. An attacker can send a large number of CONTINUATION frames that causes excessive memory allocation and CPU usage.\n\n**Affected package**: h2\n**Affected versions**: versions prior to the fix\n**Fixed version**: see advisory\n**CVSS**: 7.5 (High)\n\nThe vulnerability exists because h2 does not limit the number of CONTINUATION frames that can follow a HEADERS frame. An attacker can exploit this to cause a denial of service.\n\n### References\n\n- https://github.com/advisories/GHSA-2026-r7f2-kk9p\n- https://rustsec.org/advisories/RUSTSEC-2026-0089.html" + } + ] + } + ] + }, + "status": "New", + "labels": [ + "CVE-2026-48901", + "pscomponent:org/rhtpa-server" + ], + "versions": [], + "reporter": { + "account_id": "synthetic-reporter", + "display_name": "Synthetic Reporter" + }, + "comments": [], + "fields": { + "fixture_purpose": "SYNTHETIC TEST DATA — independent Fullsend counterpart of interactive eval 12; all identities, URLs, CVEs, commit pins and evidence are deliberate public test material, not live advisory evidence.", + "current_user": { + "accountId": "synthetic-engineer", + "displayName": "Synthetic Engineer" + }, + "ecosystem": "Cargo", + "affected_package": "h2", + "fixed_version": "0.4.8", + "affected_range": "< 0.4.8", + "affectsVersions": [ + { + "id": "version-stale", + "name": "RHTPA 2.0.0" + } + ], + "assignee": null, + "issuelinks": [], + "customfield_10632": "h2" + } + }, + "remote_links": [ + { + "url": "https://example.com/advisories/CVE-2026-48901", + "title": "Synthetic CVE-2026-48901 advisory" + } + ], + "configuration": { + "project_key": "TC", + "jira_version_prefix": "RHTPA", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "version_streams": [ + { + "name": "2.1.x", + "matrix_path": "2.1.x/security-matrix.md", + "release_repository": "release-repo" + }, + { + "name": "2.2.x", + "matrix_path": "2.2.x/security-matrix.md", + "release_repository": "release-repo" + } + ], + "source_repositories": [ + { + "name": "backend", + "url": "https://github.com/example/backend", + "deployment_context": "customer-shipped" + } + ], + "vex_justification_field": "customfield_10665", + "upstream_affected_component_field": "customfield_10632", + "stream_field": "customfield_10832" + }, + "external_evidence": { + "mitre": { + "source_url": "https://example.com/mitre/CVE-2026-48901", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "cveMetadata": { + "cveId": "CVE-2026-48901" + }, + "containers": { + "cna": { + "affected": [ + { + "product": "h2", + "versions": [ + { + "status": "affected", + "lessThan": "0.4.8", + "versionType": "semver" + } + ] + } + ] + } + } + } + }, + "osv": { + "source_url": "https://example.com/osv/CVE-2026-48901", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 503, + "body": {} + }, + "lifecycle": { + "source_url": "https://example.com/lifecycle", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "supported_streams": [ + "2.1.x", + "2.2.x" + ], + "eol_streams": [] + } + } + }, + "matrix": { + "streams": [ + { + "name": "2.1.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.1.x matrix snapshot", + "rows": [ + { + "version": "2.1.0", + "source_commits": { + "backend": "a003008" + }, + "retag_of": null + }, + { + "version": "2.1.1", + "source_commits": { + "backend": "a003012" + }, + "retag_of": null + } + ] + }, + { + "name": "2.2.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.2.x matrix snapshot", + "rows": [ + { + "version": "2.2.0", + "source_commits": { + "backend": "a004005" + }, + "retag_of": null + }, + { + "version": "2.2.1", + "source_commits": { + "backend": "a004008" + }, + "retag_of": null + }, + { + "version": "2.2.2", + "source_commits": { + "backend": "a004008" + }, + "retag_of": "2.2.1" + }, + { + "version": "2.2.3", + "source_commits": { + "backend": "a004011" + }, + "retag_of": null + }, + { + "version": "2.2.4", + "source_commits": { + "backend": "a004012" + }, + "retag_of": null + } + ] + } + ] + }, + "source_evidence": { + "lock_files": [ + { + "repository": "backend", + "ref": "a003008", + "path": "Cargo.lock", + "command": "git show a003008:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.5\"\n" + }, + { + "repository": "backend", + "ref": "a003012", + "path": "Cargo.lock", + "command": "git show a003012:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.5\"\n" + }, + { + "repository": "backend", + "ref": "a004005", + "path": "Cargo.lock", + "command": "git show a004005:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.8\"\n" + }, + { + "repository": "backend", + "ref": "a004008", + "path": "Cargo.lock", + "command": "git show a004008:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.8\"\n" + }, + { + "repository": "backend", + "ref": "a004011", + "path": "Cargo.lock", + "command": "git show a004011:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.9\"\n" + }, + { + "repository": "backend", + "ref": "a004012", + "path": "Cargo.lock", + "command": "git show a004012:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.9\"\n" + } + ], + "development_streams": [ + { + "repository": "backend", + "ref": "release/0.3.z", + "path": "Cargo.lock", + "command": "git show release/0.3.z:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.8\"\n" + }, + { + "repository": "backend", + "ref": "release/0.4.z", + "path": "Cargo.lock", + "command": "git show release/0.4.z:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.8\"\n" + } + ] + }, + "jira_metadata": { + "versions": [ + { + "id": "version-2-1-0", + "name": "RHTPA 2.1.0", + "released": true + }, + { + "id": "version-2-1-1", + "name": "RHTPA 2.1.1", + "released": true + }, + { + "id": "version-2-2-0", + "name": "RHTPA 2.2.0", + "released": true + }, + { + "id": "version-2-2-1", + "name": "RHTPA 2.2.1", + "released": true + }, + { + "id": "version-2-2-2", + "name": "RHTPA 2.2.2", + "released": true + }, + { + "id": "version-2-2-3", + "name": "RHTPA 2.2.3", + "released": true + }, + { + "id": "version-2-2-4", + "name": "RHTPA 2.2.4", + "released": true + }, + { + "id": "version-dev", + "name": "RHTPA 2.2.5", + "released": false + } + ], + "sibling_searches": [], + "related_issues": [] + }, + "idempotency": { + "action_markers": [], + "existing_remediation": [] + }, + "authorization": { + "mutation_authorized": false + } +} diff --git a/evals/triage-security/files/fullsend-eval-18-trusted-input.json b/evals/triage-security/files/fullsend-eval-18-trusted-input.json new file mode 100644 index 000000000..d033b9fef --- /dev/null +++ b/evals/triage-security/files/fullsend-eval-18-trusted-input.json @@ -0,0 +1,461 @@ +{ + "schema_version": "1", + "issue": { + "key": "TC-8001", + "summary": "CVE-2026-31812 quinn-proto - Panic on large stream counts [rhtpa-2.2]", + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "A vulnerability was found in quinn-proto. The quinn-proto crate before version 0.11.14 allows a remote attacker to cause a panic by sending a QUIC transport frame that creates an excessive number of streams. This vulnerability is classified as a denial of service (DoS).\n\n**Affected package**: quinn-proto\n**Affected versions**: versions before 0.11.14\n**Fixed version**: 0.11.14\n**CVSS**: 7.5 (High)\n\nThe vulnerability exists because quinn-proto does not properly validate the number of streams requested in a STREAMS frame. An attacker can send a specially crafted frame that causes the server to allocate an unbounded number of stream state objects, leading to a panic when the allocation exceeds internal limits.\n\n### References\n\n- https://github.com/advisories/GHSA-2026-qp73-x4mq\n- https://rustsec.org/advisories/RUSTSEC-2026-0042.html" + } + ] + } + ] + }, + "status": "In Progress", + "labels": [ + "CVE-2026-31812", + "pscomponent:org/rhtpa-server", + "ai-cve-triaged" + ], + "versions": [], + "reporter": { + "account_id": "synthetic-reporter", + "display_name": "Synthetic Reporter" + }, + "comments": [ + { + "body": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Triage summary already posted." + } + ] + } + ] + } + } + ], + "fields": { + "fixture_purpose": "SYNTHETIC TEST DATA — independent Fullsend counterpart of interactive eval 18; all identities, URLs, CVEs, commit pins and evidence are deliberate public test material, not live advisory evidence.", + "current_user": { + "accountId": "synthetic-engineer", + "displayName": "Synthetic Engineer" + }, + "ecosystem": "Cargo", + "affected_package": "quinn-proto", + "fixed_version": "0.11.14", + "affected_range": "< 0.11.14", + "affectsVersions": [ + { + "id": "version-2-2-0", + "name": "RHTPA 2.2.0", + "released": true + }, + { + "id": "version-2-2-1", + "name": "RHTPA 2.2.1", + "released": true + }, + { + "id": "version-2-2-2", + "name": "RHTPA 2.2.2", + "released": true + } + ], + "assignee": { + "accountId": "synthetic-engineer" + }, + "issuelinks": [ + { + "type": { + "name": "Depend" + }, + "outwardIssue": { + "key": "TC-8100" + } + } + ], + "customfield_10632": "quinn-proto" + } + }, + "remote_links": [ + { + "url": "https://example.com/advisories/CVE-2026-31812", + "title": "Synthetic CVE-2026-31812 advisory" + } + ], + "configuration": { + "project_key": "TC", + "jira_version_prefix": "RHTPA", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "version_streams": [ + { + "name": "2.1.x", + "matrix_path": "2.1.x/security-matrix.md", + "release_repository": "release-repo" + }, + { + "name": "2.2.x", + "matrix_path": "2.2.x/security-matrix.md", + "release_repository": "release-repo" + } + ], + "source_repositories": [ + { + "name": "backend", + "url": "https://github.com/example/backend", + "deployment_context": "customer-shipped" + } + ], + "vex_justification_field": "customfield_10665", + "upstream_affected_component_field": "customfield_10632", + "stream_field": "customfield_10832" + }, + "external_evidence": { + "mitre": { + "source_url": "https://example.com/mitre/CVE-2026-31812", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "cveMetadata": { + "cveId": "CVE-2026-31812" + }, + "containers": { + "cna": { + "affected": [ + { + "product": "quinn-proto", + "versions": [ + { + "status": "affected", + "lessThan": "0.11.14", + "versionType": "semver" + } + ] + } + ] + } + } + } + }, + "osv": { + "source_url": "https://example.com/osv/CVE-2026-31812", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "id": "CVE-2026-31812", + "affected": [ + { + "package": { + "ecosystem": "Cargo", + "name": "quinn-proto" + }, + "ranges": [ + { + "type": "SEMVER", + "events": [ + { + "introduced": "0" + }, + { + "fixed": "0.11.14" + } + ] + } + ] + } + ] + } + }, + "lifecycle": { + "source_url": "https://example.com/lifecycle", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "supported_streams": [ + "2.1.x", + "2.2.x" + ], + "eol_streams": [] + } + } + }, + "matrix": { + "streams": [ + { + "name": "2.1.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.1.x matrix snapshot", + "rows": [ + { + "version": "2.1.0", + "source_commits": { + "backend": "a003008" + }, + "retag_of": null + }, + { + "version": "2.1.1", + "source_commits": { + "backend": "a003012" + }, + "retag_of": null + } + ] + }, + { + "name": "2.2.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.2.x matrix snapshot", + "rows": [ + { + "version": "2.2.0", + "source_commits": { + "backend": "a004005" + }, + "retag_of": null + }, + { + "version": "2.2.1", + "source_commits": { + "backend": "a004008" + }, + "retag_of": null + }, + { + "version": "2.2.2", + "source_commits": { + "backend": "a004008" + }, + "retag_of": "2.2.1" + }, + { + "version": "2.2.3", + "source_commits": { + "backend": "a004011" + }, + "retag_of": null + }, + { + "version": "2.2.4", + "source_commits": { + "backend": "a004012" + }, + "retag_of": null + } + ] + } + ] + }, + "source_evidence": { + "lock_files": [ + { + "repository": "backend", + "ref": "a003008", + "path": "Cargo.lock", + "command": "git show a003008:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + }, + { + "repository": "backend", + "ref": "a003012", + "path": "Cargo.lock", + "command": "git show a003012:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + }, + { + "repository": "backend", + "ref": "a004005", + "path": "Cargo.lock", + "command": "git show a004005:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.9\"\n" + }, + { + "repository": "backend", + "ref": "a004008", + "path": "Cargo.lock", + "command": "git show a004008:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.12\"\n" + }, + { + "repository": "backend", + "ref": "a004011", + "path": "Cargo.lock", + "command": "git show a004011:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + }, + { + "repository": "backend", + "ref": "a004012", + "path": "Cargo.lock", + "command": "git show a004012:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + } + ], + "development_streams": [ + { + "repository": "backend", + "ref": "release/0.3.z", + "path": "Cargo.lock", + "command": "git show release/0.3.z:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + }, + { + "repository": "backend", + "ref": "release/0.4.z", + "path": "Cargo.lock", + "command": "git show release/0.4.z:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + } + ] + }, + "jira_metadata": { + "versions": [ + { + "id": "version-2-1-0", + "name": "RHTPA 2.1.0", + "released": true + }, + { + "id": "version-2-1-1", + "name": "RHTPA 2.1.1", + "released": true + }, + { + "id": "version-2-2-0", + "name": "RHTPA 2.2.0", + "released": true + }, + { + "id": "version-2-2-1", + "name": "RHTPA 2.2.1", + "released": true + }, + { + "id": "version-2-2-2", + "name": "RHTPA 2.2.2", + "released": true + }, + { + "id": "version-2-2-3", + "name": "RHTPA 2.2.3", + "released": true + }, + { + "id": "version-2-2-4", + "name": "RHTPA 2.2.4", + "released": true + }, + { + "id": "version-dev", + "name": "RHTPA 2.2.5", + "released": false + } + ], + "sibling_searches": [], + "related_issues": [] + }, + "idempotency": { + "action_markers": [ + "triage-security:tc-8001:assignment", + "triage-security:tc-8001:affects-versions", + "triage-security:tc-8001:labels", + "triage-security:tc-8001:status:assigned", + "triage-security:tc-8001:status:in-progress", + "triage-security:tc-8001:comment:summary", + "triage-security:tc-8001:remediation:upstream", + "triage-security:tc-8001:remediation:downstream", + "triage-security:tc-8001:link:depend:upstream", + "triage-security:tc-8001:link:blocks:upstream-downstream" + ], + "existing_remediation": [ + { + "key": "TC-8100", + "summary": "Remediate CVE-2026-31812: upstream quinn-proto >= 0.11.14 [rhtpa-2.2]", + "status": "In Progress", + "labels": [ + "ai-generated-jira", + "Security", + "CVE-2026-31812" + ], + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Existing upstream remediation task." + } + ] + } + ] + }, + "comments": [], + "links": [] + }, + { + "key": "TC-8101", + "summary": "Remediate CVE-2026-31812: downstream quinn-proto >= 0.11.14 [rhtpa-2.2]", + "status": "In Progress", + "labels": [ + "ai-generated-jira", + "Security", + "CVE-2026-31812" + ], + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Existing downstream remediation task." + } + ] + } + ] + }, + "comments": [ + { + "body": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "[sdlc-workflow] Description digest: sha256-adf:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa" + } + ] + } + ] + } + } + ], + "links": [] + } + ] + }, + "authorization": { + "mutation_authorized": true + } +} diff --git a/evals/triage-security/files/fullsend-eval-2-trusted-input.json b/evals/triage-security/files/fullsend-eval-2-trusted-input.json new file mode 100644 index 000000000..0d6ca7c96 --- /dev/null +++ b/evals/triage-security/files/fullsend-eval-2-trusted-input.json @@ -0,0 +1,336 @@ +{ + "schema_version": "1", + "issue": { + "key": "TC-8002", + "summary": "CVE-2026-28940 serde_json - Stack overflow on deeply nested input [rhtpa-2.2]", + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "A vulnerability was found in serde_json. Versions of serde_json before 1.0.135 are vulnerable to a stack overflow when deserializing deeply nested JSON input. An attacker can craft a JSON payload with thousands of nested arrays or objects that causes unbounded recursion during deserialization, leading to a stack overflow and process crash.\n\n**Affected package**: serde_json\n**Affected versions**: versions before 1.0.135\n**Fixed version**: 1.0.135\n**CVSS**: 5.3 (Medium)\n\nThe fix introduces a configurable recursion limit that defaults to 128 levels of nesting.\n\n### References\n\n- https://github.com/advisories/GHSA-2026-j9r2-m5vk\n- https://rustsec.org/advisories/RUSTSEC-2026-0019.html" + } + ] + } + ] + }, + "status": "New", + "labels": [ + "CVE-2026-28940", + "pscomponent:org/rhtpa-server" + ], + "versions": [], + "reporter": { + "account_id": "synthetic-reporter", + "display_name": "Synthetic Reporter" + }, + "comments": [], + "fields": { + "fixture_purpose": "SYNTHETIC TEST DATA — independent Fullsend counterpart of interactive eval 2; all identities, URLs, CVEs, commit pins and evidence are deliberate public test material, not live advisory evidence.", + "current_user": { + "accountId": "synthetic-engineer", + "displayName": "Synthetic Engineer" + }, + "ecosystem": "Cargo", + "affected_package": "serde_json", + "fixed_version": "1.0.135", + "affected_range": "< 1.0.135", + "affectsVersions": [ + { + "id": "version-stale", + "name": "RHTPA 2.0.0" + } + ], + "assignee": null, + "issuelinks": [], + "customfield_10632": "serde_json" + } + }, + "remote_links": [ + { + "url": "https://example.com/advisories/CVE-2026-28940", + "title": "Synthetic CVE-2026-28940 advisory" + } + ], + "configuration": { + "project_key": "TC", + "jira_version_prefix": "RHTPA", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "version_streams": [ + { + "name": "2.1.x", + "matrix_path": "2.1.x/security-matrix.md", + "release_repository": "release-repo" + }, + { + "name": "2.2.x", + "matrix_path": "2.2.x/security-matrix.md", + "release_repository": "release-repo" + } + ], + "source_repositories": [ + { + "name": "backend", + "url": "https://github.com/example/backend", + "deployment_context": "customer-shipped" + } + ], + "vex_justification_field": "customfield_10665", + "upstream_affected_component_field": "customfield_10632", + "stream_field": "customfield_10832" + }, + "external_evidence": { + "mitre": { + "source_url": "https://example.com/mitre/CVE-2026-28940", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "cveMetadata": { + "cveId": "CVE-2026-28940" + }, + "containers": { + "cna": { + "affected": [ + { + "product": "serde_json", + "versions": [ + { + "status": "affected", + "lessThan": "1.0.135", + "versionType": "semver" + } + ] + } + ] + } + } + } + }, + "osv": { + "source_url": "https://example.com/osv/CVE-2026-28940", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "id": "CVE-2026-28940", + "affected": [ + { + "package": { + "ecosystem": "Cargo", + "name": "serde_json" + }, + "ranges": [ + { + "type": "SEMVER", + "events": [ + { + "introduced": "0" + }, + { + "fixed": "1.0.135" + } + ] + } + ] + } + ] + } + }, + "lifecycle": { + "source_url": "https://example.com/lifecycle", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "supported_streams": [ + "2.1.x", + "2.2.x" + ], + "eol_streams": [] + } + } + }, + "matrix": { + "streams": [ + { + "name": "2.1.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.1.x matrix snapshot", + "rows": [ + { + "version": "2.1.0", + "source_commits": { + "backend": "a003008" + }, + "retag_of": null + }, + { + "version": "2.1.1", + "source_commits": { + "backend": "a003012" + }, + "retag_of": null + } + ] + }, + { + "name": "2.2.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.2.x matrix snapshot", + "rows": [ + { + "version": "2.2.0", + "source_commits": { + "backend": "a004005" + }, + "retag_of": null + }, + { + "version": "2.2.1", + "source_commits": { + "backend": "a004008" + }, + "retag_of": null + }, + { + "version": "2.2.2", + "source_commits": { + "backend": "a004008" + }, + "retag_of": "2.2.1" + }, + { + "version": "2.2.3", + "source_commits": { + "backend": "a004011" + }, + "retag_of": null + }, + { + "version": "2.2.4", + "source_commits": { + "backend": "a004012" + }, + "retag_of": null + } + ] + } + ] + }, + "source_evidence": { + "lock_files": [ + { + "repository": "backend", + "ref": "a003008", + "path": "Cargo.lock", + "command": "git show a003008:Cargo.lock", + "content": "[[package]]\nname = \"serde_json\"\nversion = \"1.0.137\"\n" + }, + { + "repository": "backend", + "ref": "a003012", + "path": "Cargo.lock", + "command": "git show a003012:Cargo.lock", + "content": "[[package]]\nname = \"serde_json\"\nversion = \"1.0.137\"\n" + }, + { + "repository": "backend", + "ref": "a004005", + "path": "Cargo.lock", + "command": "git show a004005:Cargo.lock", + "content": "[[package]]\nname = \"serde_json\"\nversion = \"1.0.138\"\n" + }, + { + "repository": "backend", + "ref": "a004008", + "path": "Cargo.lock", + "command": "git show a004008:Cargo.lock", + "content": "[[package]]\nname = \"serde_json\"\nversion = \"1.0.138\"\n" + }, + { + "repository": "backend", + "ref": "a004011", + "path": "Cargo.lock", + "command": "git show a004011:Cargo.lock", + "content": "[[package]]\nname = \"serde_json\"\nversion = \"1.0.139\"\n" + }, + { + "repository": "backend", + "ref": "a004012", + "path": "Cargo.lock", + "command": "git show a004012:Cargo.lock", + "content": "[[package]]\nname = \"serde_json\"\nversion = \"1.0.139\"\n" + } + ], + "development_streams": [ + { + "repository": "backend", + "ref": "release/0.3.z", + "path": "Cargo.lock", + "command": "git show release/0.3.z:Cargo.lock", + "content": "[[package]]\nname = \"serde_json\"\nversion = \"1.0.135\"\n" + }, + { + "repository": "backend", + "ref": "release/0.4.z", + "path": "Cargo.lock", + "command": "git show release/0.4.z:Cargo.lock", + "content": "[[package]]\nname = \"serde_json\"\nversion = \"1.0.135\"\n" + } + ] + }, + "jira_metadata": { + "versions": [ + { + "id": "version-2-1-0", + "name": "RHTPA 2.1.0", + "released": true + }, + { + "id": "version-2-1-1", + "name": "RHTPA 2.1.1", + "released": true + }, + { + "id": "version-2-2-0", + "name": "RHTPA 2.2.0", + "released": true + }, + { + "id": "version-2-2-1", + "name": "RHTPA 2.2.1", + "released": true + }, + { + "id": "version-2-2-2", + "name": "RHTPA 2.2.2", + "released": true + }, + { + "id": "version-2-2-3", + "name": "RHTPA 2.2.3", + "released": true + }, + { + "id": "version-2-2-4", + "name": "RHTPA 2.2.4", + "released": true + }, + { + "id": "version-dev", + "name": "RHTPA 2.2.5", + "released": false + } + ], + "sibling_searches": [], + "related_issues": [] + }, + "idempotency": { + "action_markers": [], + "existing_remediation": [] + }, + "authorization": { + "mutation_authorized": false + } +} diff --git a/evals/triage-security/files/fullsend-eval-3-trusted-input.json b/evals/triage-security/files/fullsend-eval-3-trusted-input.json new file mode 100644 index 000000000..d6009f170 --- /dev/null +++ b/evals/triage-security/files/fullsend-eval-3-trusted-input.json @@ -0,0 +1,368 @@ +{ + "schema_version": "1", + "issue": { + "key": "TC-8003", + "summary": "CVE-2026-31812 quinn-proto - Panic on large stream counts [rhtpa-2.2]", + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "A vulnerability was found in quinn-proto. The quinn-proto crate before version 0.11.14 allows a remote attacker to cause a panic by sending a QUIC transport frame that creates an excessive number of streams. This vulnerability is classified as a denial of service (DoS).\n\n**Affected package**: quinn-proto\n**Affected versions**: versions before 0.11.14\n**Fixed version**: 0.11.14\n**CVSS**: 7.5 (High)\n\nThe vulnerability exists because quinn-proto does not properly validate the number of streams requested in a STREAMS frame. An attacker can send a specially crafted frame that causes the server to allocate an unbounded number of stream state objects, leading to a panic when the allocation exceeds internal limits.\n\n### References\n\n- https://github.com/advisories/GHSA-2026-qp73-x4mq\n- https://rustsec.org/advisories/RUSTSEC-2026-0042.html\n\n### Sibling Issues (provided by JQL search)\n\nThe following sibling issue exists for the same CVE:\n\n- **TC-7999** — CVE-2026-31812 quinn-proto - Panic on large stream counts [rhtpa-2.2]\n - **Status**: In Progress\n - **Labels**: CVE-2026-31812, pscomponent:org/rhtpa-server\n - **Affects Versions**: RHTPA 2.2.0, RHTPA 2.2.1\n - **Stream suffix**: [rhtpa-2.2] (same stream as TC-8003)" + } + ] + } + ] + }, + "status": "New", + "labels": [ + "CVE-2026-31812", + "pscomponent:org/rhtpa-server" + ], + "versions": [], + "reporter": { + "account_id": "synthetic-reporter", + "display_name": "Synthetic Reporter" + }, + "comments": [], + "fields": { + "fixture_purpose": "SYNTHETIC TEST DATA — independent Fullsend counterpart of interactive eval 3; all identities, URLs, CVEs, commit pins and evidence are deliberate public test material, not live advisory evidence.", + "current_user": { + "accountId": "synthetic-engineer", + "displayName": "Synthetic Engineer" + }, + "ecosystem": "Cargo", + "affected_package": "quinn-proto", + "fixed_version": "0.11.14", + "affected_range": "< 0.11.14", + "affectsVersions": [ + { + "id": "version-stale", + "name": "RHTPA 2.0.0" + } + ], + "assignee": null, + "issuelinks": [], + "customfield_10632": "quinn-proto" + } + }, + "remote_links": [ + { + "url": "https://example.com/advisories/CVE-2026-31812", + "title": "Synthetic CVE-2026-31812 advisory" + } + ], + "configuration": { + "project_key": "TC", + "jira_version_prefix": "RHTPA", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "version_streams": [ + { + "name": "2.1.x", + "matrix_path": "2.1.x/security-matrix.md", + "release_repository": "release-repo" + }, + { + "name": "2.2.x", + "matrix_path": "2.2.x/security-matrix.md", + "release_repository": "release-repo" + } + ], + "source_repositories": [ + { + "name": "backend", + "url": "https://github.com/example/backend", + "deployment_context": "customer-shipped" + } + ], + "vex_justification_field": "customfield_10665", + "upstream_affected_component_field": "customfield_10632", + "stream_field": "customfield_10832" + }, + "external_evidence": { + "mitre": { + "source_url": "https://example.com/mitre/CVE-2026-31812", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "cveMetadata": { + "cveId": "CVE-2026-31812" + }, + "containers": { + "cna": { + "affected": [ + { + "product": "quinn-proto", + "versions": [ + { + "status": "affected", + "lessThan": "0.11.14", + "versionType": "semver" + } + ] + } + ] + } + } + } + }, + "osv": { + "source_url": "https://example.com/osv/CVE-2026-31812", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "id": "CVE-2026-31812", + "affected": [ + { + "package": { + "ecosystem": "Cargo", + "name": "quinn-proto" + }, + "ranges": [ + { + "type": "SEMVER", + "events": [ + { + "introduced": "0" + }, + { + "fixed": "0.11.14" + } + ] + } + ] + } + ] + } + }, + "lifecycle": { + "source_url": "https://example.com/lifecycle", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "supported_streams": [ + "2.1.x", + "2.2.x" + ], + "eol_streams": [] + } + } + }, + "matrix": { + "streams": [ + { + "name": "2.1.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.1.x matrix snapshot", + "rows": [ + { + "version": "2.1.0", + "source_commits": { + "backend": "a003008" + }, + "retag_of": null + }, + { + "version": "2.1.1", + "source_commits": { + "backend": "a003012" + }, + "retag_of": null + } + ] + }, + { + "name": "2.2.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.2.x matrix snapshot", + "rows": [ + { + "version": "2.2.0", + "source_commits": { + "backend": "a004005" + }, + "retag_of": null + }, + { + "version": "2.2.1", + "source_commits": { + "backend": "a004008" + }, + "retag_of": null + }, + { + "version": "2.2.2", + "source_commits": { + "backend": "a004008" + }, + "retag_of": "2.2.1" + }, + { + "version": "2.2.3", + "source_commits": { + "backend": "a004011" + }, + "retag_of": null + }, + { + "version": "2.2.4", + "source_commits": { + "backend": "a004012" + }, + "retag_of": null + } + ] + } + ] + }, + "source_evidence": { + "lock_files": [ + { + "repository": "backend", + "ref": "a003008", + "path": "Cargo.lock", + "command": "git show a003008:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.9\"\n" + }, + { + "repository": "backend", + "ref": "a003012", + "path": "Cargo.lock", + "command": "git show a003012:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.9\"\n" + }, + { + "repository": "backend", + "ref": "a004005", + "path": "Cargo.lock", + "command": "git show a004005:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.9\"\n" + }, + { + "repository": "backend", + "ref": "a004008", + "path": "Cargo.lock", + "command": "git show a004008:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.12\"\n" + }, + { + "repository": "backend", + "ref": "a004011", + "path": "Cargo.lock", + "command": "git show a004011:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + }, + { + "repository": "backend", + "ref": "a004012", + "path": "Cargo.lock", + "command": "git show a004012:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + } + ], + "development_streams": [ + { + "repository": "backend", + "ref": "release/0.3.z", + "path": "Cargo.lock", + "command": "git show release/0.3.z:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + }, + { + "repository": "backend", + "ref": "release/0.4.z", + "path": "Cargo.lock", + "command": "git show release/0.4.z:Cargo.lock", + "content": "[[package]]\nname = \"quinn-proto\"\nversion = \"0.11.14\"\n" + } + ] + }, + "jira_metadata": { + "versions": [ + { + "id": "version-2-1-0", + "name": "RHTPA 2.1.0", + "released": true + }, + { + "id": "version-2-1-1", + "name": "RHTPA 2.1.1", + "released": true + }, + { + "id": "version-2-2-0", + "name": "RHTPA 2.2.0", + "released": true + }, + { + "id": "version-2-2-1", + "name": "RHTPA 2.2.1", + "released": true + }, + { + "id": "version-2-2-2", + "name": "RHTPA 2.2.2", + "released": true + }, + { + "id": "version-2-2-3", + "name": "RHTPA 2.2.3", + "released": true + }, + { + "id": "version-2-2-4", + "name": "RHTPA 2.2.4", + "released": true + }, + { + "id": "version-dev", + "name": "RHTPA 2.2.5", + "released": false + } + ], + "sibling_searches": [ + { + "purpose": "same-cve-siblings", + "jql": "project = TC AND labels = \"CVE-2026-31812\" AND key != \"TC-8003\"", + "issues": [ + { + "key": "TC-7999", + "summary": "CVE-2026-31812 quinn-proto [rhtpa-2.2]", + "status": "In Progress", + "labels": [ + "CVE-2026-31812" + ], + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Same CVE, same 2.2.x stream; Affects Versions RHTPA 2.2.0, RHTPA 2.2.1." + } + ] + } + ] + }, + "comments": [], + "links": [] + } + ] + } + ], + "related_issues": [] + }, + "idempotency": { + "action_markers": [], + "existing_remediation": [] + }, + "authorization": { + "mutation_authorized": true + } +} diff --git a/evals/triage-security/files/fullsend-eval-4-trusted-input.json b/evals/triage-security/files/fullsend-eval-4-trusted-input.json new file mode 100644 index 000000000..b30800f63 --- /dev/null +++ b/evals/triage-security/files/fullsend-eval-4-trusted-input.json @@ -0,0 +1,336 @@ +{ + "schema_version": "1", + "issue": { + "key": "TC-8004", + "summary": "CVE-2026-33501 h2 - Memory exhaustion via CONTINUATION frames", + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "A vulnerability was found in the h2 crate. Versions of h2 before 0.4.8 are vulnerable to memory exhaustion caused by a peer sending an excessive number of CONTINUATION frames following a HEADERS frame. The h2 library accumulates all CONTINUATION frame data without enforcing a size limit on the accumulated header block, allowing an attacker to consume unbounded memory on the server.\n\n**Affected package**: h2\n**Affected versions**: versions before 0.4.8\n**Fixed version**: 0.4.8\n**CVSS**: 7.5 (High)\n\nThis issue is distinct from CVE-2024-2758 (httpd CONTINUATION flood) — this CVE specifically affects the Rust h2 library's header accumulation logic. The fix adds a configurable maximum header list size that defaults to 16 KiB.\n\nNote: This issue has NO stream suffix in the summary — it is unscoped and covers all streams. The version impact analysis should check all streams and create remediation only for actually affected streams.\n\n### References\n\n- https://github.com/advisories/GHSA-2026-kv8p-r3n7\n- https://rustsec.org/advisories/RUSTSEC-2026-0055.html" + } + ] + } + ] + }, + "status": "New", + "labels": [ + "CVE-2026-33501", + "pscomponent:org/rhtpa-server" + ], + "versions": [], + "reporter": { + "account_id": "synthetic-reporter", + "display_name": "Synthetic Reporter" + }, + "comments": [], + "fields": { + "fixture_purpose": "SYNTHETIC TEST DATA — independent Fullsend counterpart of interactive eval 4; all identities, URLs, CVEs, commit pins and evidence are deliberate public test material, not live advisory evidence.", + "current_user": { + "accountId": "synthetic-engineer", + "displayName": "Synthetic Engineer" + }, + "ecosystem": "Cargo", + "affected_package": "h2", + "fixed_version": "0.4.8", + "affected_range": "< 0.4.8", + "affectsVersions": [ + { + "id": "version-stale", + "name": "RHTPA 2.0.0" + } + ], + "assignee": null, + "issuelinks": [], + "customfield_10632": "h2" + } + }, + "remote_links": [ + { + "url": "https://example.com/advisories/CVE-2026-33501", + "title": "Synthetic CVE-2026-33501 advisory" + } + ], + "configuration": { + "project_key": "TC", + "jira_version_prefix": "RHTPA", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "version_streams": [ + { + "name": "2.1.x", + "matrix_path": "2.1.x/security-matrix.md", + "release_repository": "release-repo" + }, + { + "name": "2.2.x", + "matrix_path": "2.2.x/security-matrix.md", + "release_repository": "release-repo" + } + ], + "source_repositories": [ + { + "name": "backend", + "url": "https://github.com/example/backend", + "deployment_context": "customer-shipped" + } + ], + "vex_justification_field": "customfield_10665", + "upstream_affected_component_field": "customfield_10632", + "stream_field": "customfield_10832" + }, + "external_evidence": { + "mitre": { + "source_url": "https://example.com/mitre/CVE-2026-33501", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "cveMetadata": { + "cveId": "CVE-2026-33501" + }, + "containers": { + "cna": { + "affected": [ + { + "product": "h2", + "versions": [ + { + "status": "affected", + "lessThan": "0.4.8", + "versionType": "semver" + } + ] + } + ] + } + } + } + }, + "osv": { + "source_url": "https://example.com/osv/CVE-2026-33501", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "id": "CVE-2026-33501", + "affected": [ + { + "package": { + "ecosystem": "Cargo", + "name": "h2" + }, + "ranges": [ + { + "type": "SEMVER", + "events": [ + { + "introduced": "0" + }, + { + "fixed": "0.4.8" + } + ] + } + ] + } + ] + } + }, + "lifecycle": { + "source_url": "https://example.com/lifecycle", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "supported_streams": [ + "2.1.x", + "2.2.x" + ], + "eol_streams": [] + } + } + }, + "matrix": { + "streams": [ + { + "name": "2.1.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.1.x matrix snapshot", + "rows": [ + { + "version": "2.1.0", + "source_commits": { + "backend": "a003008" + }, + "retag_of": null + }, + { + "version": "2.1.1", + "source_commits": { + "backend": "a003012" + }, + "retag_of": null + } + ] + }, + { + "name": "2.2.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.2.x matrix snapshot", + "rows": [ + { + "version": "2.2.0", + "source_commits": { + "backend": "a004005" + }, + "retag_of": null + }, + { + "version": "2.2.1", + "source_commits": { + "backend": "a004008" + }, + "retag_of": null + }, + { + "version": "2.2.2", + "source_commits": { + "backend": "a004008" + }, + "retag_of": "2.2.1" + }, + { + "version": "2.2.3", + "source_commits": { + "backend": "a004011" + }, + "retag_of": null + }, + { + "version": "2.2.4", + "source_commits": { + "backend": "a004012" + }, + "retag_of": null + } + ] + } + ] + }, + "source_evidence": { + "lock_files": [ + { + "repository": "backend", + "ref": "a003008", + "path": "Cargo.lock", + "command": "git show a003008:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.5\"\n" + }, + { + "repository": "backend", + "ref": "a003012", + "path": "Cargo.lock", + "command": "git show a003012:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.5\"\n" + }, + { + "repository": "backend", + "ref": "a004005", + "path": "Cargo.lock", + "command": "git show a004005:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.8\"\n" + }, + { + "repository": "backend", + "ref": "a004008", + "path": "Cargo.lock", + "command": "git show a004008:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.8\"\n" + }, + { + "repository": "backend", + "ref": "a004011", + "path": "Cargo.lock", + "command": "git show a004011:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.9\"\n" + }, + { + "repository": "backend", + "ref": "a004012", + "path": "Cargo.lock", + "command": "git show a004012:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.9\"\n" + } + ], + "development_streams": [ + { + "repository": "backend", + "ref": "release/0.3.z", + "path": "Cargo.lock", + "command": "git show release/0.3.z:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.5\"\n" + }, + { + "repository": "backend", + "ref": "release/0.4.z", + "path": "Cargo.lock", + "command": "git show release/0.4.z:Cargo.lock", + "content": "[[package]]\nname = \"h2\"\nversion = \"0.4.8\"\n" + } + ] + }, + "jira_metadata": { + "versions": [ + { + "id": "version-2-1-0", + "name": "RHTPA 2.1.0", + "released": true + }, + { + "id": "version-2-1-1", + "name": "RHTPA 2.1.1", + "released": true + }, + { + "id": "version-2-2-0", + "name": "RHTPA 2.2.0", + "released": true + }, + { + "id": "version-2-2-1", + "name": "RHTPA 2.2.1", + "released": true + }, + { + "id": "version-2-2-2", + "name": "RHTPA 2.2.2", + "released": true + }, + { + "id": "version-2-2-3", + "name": "RHTPA 2.2.3", + "released": true + }, + { + "id": "version-2-2-4", + "name": "RHTPA 2.2.4", + "released": true + }, + { + "id": "version-dev", + "name": "RHTPA 2.2.5", + "released": false + } + ], + "sibling_searches": [], + "related_issues": [] + }, + "idempotency": { + "action_markers": [], + "existing_remediation": [] + }, + "authorization": { + "mutation_authorized": true + } +} diff --git a/evals/triage-security/files/fullsend-eval-5-trusted-input.json b/evals/triage-security/files/fullsend-eval-5-trusted-input.json new file mode 100644 index 000000000..b5df4b9d0 --- /dev/null +++ b/evals/triage-security/files/fullsend-eval-5-trusted-input.json @@ -0,0 +1,366 @@ +{ + "schema_version": "1", + "issue": { + "key": "TC-8005", + "summary": "CVE-2026-40215 openssl-libs - Buffer over-read in X.509 certificate verification [rhtpa-2.2]", + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "A vulnerability was found in openssl-libs. Versions of openssl before 3.0.7-28.el9_4 are vulnerable to a buffer over-read during X.509 certificate chain verification. A remote attacker can craft a certificate with a malformed extension that triggers an out-of-bounds read, potentially leaking sensitive memory contents or causing a crash.\n\n**Affected package**: openssl-libs\n**Affected versions**: versions before 3.0.7-28.el9_4\n**Fixed version**: 3.0.7-28.el9_4\n**CVSS**: 7.1 (High)\n\nThe vulnerability exists in the `X509_verify_cert()` code path where the extension parser does not properly validate the length field of a Subject Alternative Name extension. The fix adds bounds checking before reading extension data.\n\n### References\n\n- https://www.cve.org/CVERecord?id=CVE-2026-40215\n- https://access.redhat.com/errata/RHSA-2026:4021" + } + ] + } + ] + }, + "status": "New", + "labels": [ + "CVE-2026-40215", + "pscomponent:org/rhtpa-server" + ], + "versions": [], + "reporter": { + "account_id": "synthetic-reporter", + "display_name": "Synthetic Reporter" + }, + "comments": [], + "fields": { + "fixture_purpose": "SYNTHETIC TEST DATA — independent Fullsend counterpart of interactive eval 5; all identities, URLs, CVEs, commit pins and evidence are deliberate public test material, not live advisory evidence.", + "current_user": { + "accountId": "synthetic-engineer", + "displayName": "Synthetic Engineer" + }, + "ecosystem": "RPM", + "affected_package": "openssl-libs", + "fixed_version": "3.0.7-28.el9_4", + "affected_range": "< 3.0.7-28.el9_4", + "affectsVersions": [ + { + "id": "version-stale", + "name": "RHTPA 2.0.0" + } + ], + "assignee": null, + "issuelinks": [], + "customfield_10632": "openssl-libs", + "sbom_evidence": { + "source": "trusted pre-retrieved SBOM comparison", + "packages": [ + { + "name": "openssl-libs", + "version": "3.0.7-25.el9_3", + "release": "2.2.0" + }, + { + "name": "openssl-libs", + "version": "3.0.7-27.el9_4", + "release": "2.2.1" + }, + { + "name": "openssl-libs", + "version": "3.0.7-27.el9_4", + "release": "2.2.2" + }, + { + "name": "openssl-libs", + "version": "3.0.7-28.el9_4", + "release": "2.2.3" + }, + { + "name": "openssl-libs", + "version": "3.0.7-28.el9_4", + "release": "2.2.4" + } + ] + } + } + }, + "remote_links": [ + { + "url": "https://example.com/advisories/CVE-2026-40215", + "title": "Synthetic CVE-2026-40215 advisory" + } + ], + "configuration": { + "project_key": "TC", + "jira_version_prefix": "RHTPA", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "version_streams": [ + { + "name": "2.1.x", + "matrix_path": "2.1.x/security-matrix.md", + "release_repository": "release-repo" + }, + { + "name": "2.2.x", + "matrix_path": "2.2.x/security-matrix.md", + "release_repository": "release-repo" + } + ], + "source_repositories": [ + { + "name": "release-repo", + "url": "https://github.com/example/release-repo", + "deployment_context": "customer-shipped" + } + ], + "vex_justification_field": "customfield_10665", + "upstream_affected_component_field": "customfield_10632", + "stream_field": "customfield_10832" + }, + "external_evidence": { + "mitre": { + "source_url": "https://example.com/mitre/CVE-2026-40215", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "cveMetadata": { + "cveId": "CVE-2026-40215" + }, + "containers": { + "cna": { + "affected": [ + { + "product": "openssl-libs", + "versions": [ + { + "status": "affected", + "lessThan": "3.0.7-28.el9_4", + "versionType": "semver" + } + ] + } + ] + } + } + } + }, + "osv": { + "source_url": "https://example.com/osv/CVE-2026-40215", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "id": "CVE-2026-40215", + "affected": [ + { + "package": { + "ecosystem": "RPM", + "name": "openssl-libs" + }, + "ranges": [ + { + "type": "SEMVER", + "events": [ + { + "introduced": "0" + }, + { + "fixed": "3.0.7-28.el9_4" + } + ] + } + ] + } + ] + } + }, + "lifecycle": { + "source_url": "https://example.com/lifecycle", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "supported_streams": [ + "2.1.x", + "2.2.x" + ], + "eol_streams": [] + } + } + }, + "matrix": { + "streams": [ + { + "name": "2.1.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.1.x matrix snapshot", + "rows": [ + { + "version": "2.1.0", + "source_commits": { + "release-repo": "a003008" + }, + "retag_of": null + }, + { + "version": "2.1.1", + "source_commits": { + "release-repo": "a003012" + }, + "retag_of": null + } + ] + }, + { + "name": "2.2.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.2.x matrix snapshot", + "rows": [ + { + "version": "2.2.0", + "source_commits": { + "release-repo": "a004005" + }, + "retag_of": null + }, + { + "version": "2.2.1", + "source_commits": { + "release-repo": "a004008" + }, + "retag_of": null + }, + { + "version": "2.2.2", + "source_commits": { + "release-repo": "a004008" + }, + "retag_of": "2.2.1" + }, + { + "version": "2.2.3", + "source_commits": { + "release-repo": "a004011" + }, + "retag_of": null + }, + { + "version": "2.2.4", + "source_commits": { + "release-repo": "a004012" + }, + "retag_of": null + } + ] + } + ] + }, + "source_evidence": { + "lock_files": [ + { + "repository": "release-repo", + "ref": "a003008", + "path": "rpms.lock.yaml", + "command": "git show a003008:rpms.lock.yaml", + "content": "packages:\n - name: openssl-libs\n version: 3.0.7-24.el9\n" + }, + { + "repository": "release-repo", + "ref": "a003012", + "path": "rpms.lock.yaml", + "command": "git show a003012:rpms.lock.yaml", + "content": "packages:\n - name: openssl-libs\n version: 3.0.7-24.el9\n" + }, + { + "repository": "release-repo", + "ref": "a004005", + "path": "rpms.lock.yaml", + "command": "git show a004005:rpms.lock.yaml", + "content": "packages:\n - name: openssl-libs\n version: 3.0.7-25.el9_3\n" + }, + { + "repository": "release-repo", + "ref": "a004008", + "path": "rpms.lock.yaml", + "command": "git show a004008:rpms.lock.yaml", + "content": "packages:\n - name: openssl-libs\n version: 3.0.7-27.el9_4\n" + }, + { + "repository": "release-repo", + "ref": "a004011", + "path": "rpms.lock.yaml", + "command": "git show a004011:rpms.lock.yaml", + "content": "packages:\n - name: openssl-libs\n version: 3.0.7-28.el9_4\n" + }, + { + "repository": "release-repo", + "ref": "a004012", + "path": "rpms.lock.yaml", + "command": "git show a004012:rpms.lock.yaml", + "content": "packages:\n - name: openssl-libs\n version: 3.0.7-28.el9_4\n" + } + ], + "development_streams": [ + { + "repository": "release-repo", + "ref": "release/0.3.z", + "path": "rpms.lock.yaml", + "command": "git show release/0.3.z:rpms.lock.yaml", + "content": "packages:\n - name: openssl-libs\n version: 3.0.7-28.el9_4\n" + }, + { + "repository": "release-repo", + "ref": "release/0.4.z", + "path": "rpms.lock.yaml", + "command": "git show release/0.4.z:rpms.lock.yaml", + "content": "packages:\n - name: openssl-libs\n version: 3.0.7-28.el9_4\n" + } + ] + }, + "jira_metadata": { + "versions": [ + { + "id": "version-2-1-0", + "name": "RHTPA 2.1.0", + "released": true + }, + { + "id": "version-2-1-1", + "name": "RHTPA 2.1.1", + "released": true + }, + { + "id": "version-2-2-0", + "name": "RHTPA 2.2.0", + "released": true + }, + { + "id": "version-2-2-1", + "name": "RHTPA 2.2.1", + "released": true + }, + { + "id": "version-2-2-2", + "name": "RHTPA 2.2.2", + "released": true + }, + { + "id": "version-2-2-3", + "name": "RHTPA 2.2.3", + "released": true + }, + { + "id": "version-2-2-4", + "name": "RHTPA 2.2.4", + "released": true + }, + { + "id": "version-dev", + "name": "RHTPA 2.2.5", + "released": false + } + ], + "sibling_searches": [], + "related_issues": [] + }, + "idempotency": { + "action_markers": [], + "existing_remediation": [] + }, + "authorization": { + "mutation_authorized": false + } +} diff --git a/evals/triage-security/files/fullsend-eval-8-trusted-input.json b/evals/triage-security/files/fullsend-eval-8-trusted-input.json new file mode 100644 index 000000000..e8660f6d6 --- /dev/null +++ b/evals/triage-security/files/fullsend-eval-8-trusted-input.json @@ -0,0 +1,434 @@ +{ + "schema_version": "1", + "issue": { + "key": "TC-8010", + "summary": "CVE-2026-44492 axios - Server-Side Request Forgery via crafted URL [rhtpa-2.2]", + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "A vulnerability was found in axios. The axios package before version 1.8.2 is vulnerable to Server-Side Request Forgery (SSRF) via a crafted URL that bypasses hostname validation. An attacker can exploit this to make requests to internal services.\n\n**Affected package**: axios\n**Affected versions**: versions before 1.8.2\n**Fixed version**: 1.8.2\n**CVSS**: 8.1 (High)\n\nThe vulnerability exists because axios does not properly validate the hostname in URLs when following redirects. An attacker can craft a URL that initially resolves to an external host but redirects to an internal service.\n\n### References\n\n- https://github.com/advisories/GHSA-2026-ax91-r7pp\n\n### Related CVE Jiras (provided by JQL search on customfield_10632 = \"axios\")\n\nThe following related CVE Jira was found affecting the same upstream component:\n\n- **TC-8008** — CVE-2026-42035 axios - Prototype Pollution via header parsing [rhtpa-2.2]\n - **Status**: In Progress\n - **Labels**: CVE-2026-42035, pscomponent:org/rhtpa-ui\n - **customfield_10632**: axios\n - **customfield_10669**: pscomponent:org/rhtpa-ui\n - **customfield_10832**: rhtpa-2.2\n - **Issue Links**:\n - **Depend**: TC-8009 (remediation Task)\n - **Summary**: Bump axios to 1.9.0 in rhtpa-ui [rhtpa-2.2]\n - **Status**: In Progress\n - **Description excerpt**: \"Bump axios from 1.7.4 to 1.9.0 to resolve CVE-2026-42035. The fix requires axios >= 1.8.0.\"" + } + ] + } + ] + }, + "status": "New", + "labels": [ + "CVE-2026-44492", + "pscomponent:org/rhtpa-ui" + ], + "versions": [], + "reporter": { + "account_id": "synthetic-reporter", + "display_name": "Synthetic Reporter" + }, + "comments": [], + "fields": { + "fixture_purpose": "SYNTHETIC TEST DATA — independent Fullsend counterpart of interactive eval 8; all identities, URLs, CVEs, commit pins and evidence are deliberate public test material, not live advisory evidence.", + "current_user": { + "accountId": "synthetic-engineer", + "displayName": "Synthetic Engineer" + }, + "ecosystem": "npm", + "affected_package": "axios", + "fixed_version": "1.8.2", + "affected_range": "< 1.8.2", + "affectsVersions": [ + { + "id": "version-stale", + "name": "RHTPA 2.0.0" + } + ], + "assignee": null, + "issuelinks": [ + { + "type": { + "name": "Related" + }, + "outwardIssue": { + "key": "TC-8008" + } + } + ], + "customfield_10632": "axios" + } + }, + "remote_links": [ + { + "url": "https://example.com/advisories/CVE-2026-44492", + "title": "Synthetic CVE-2026-44492 advisory" + } + ], + "configuration": { + "project_key": "TC", + "jira_version_prefix": "RHTPA", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "version_streams": [ + { + "name": "2.1.x", + "matrix_path": "2.1.x/security-matrix.md", + "release_repository": "release-repo" + }, + { + "name": "2.2.x", + "matrix_path": "2.2.x/security-matrix.md", + "release_repository": "release-repo" + } + ], + "source_repositories": [ + { + "name": "ui", + "url": "https://github.com/example/ui", + "deployment_context": "customer-shipped" + } + ], + "vex_justification_field": "customfield_10665", + "upstream_affected_component_field": "customfield_10632", + "stream_field": "customfield_10832" + }, + "external_evidence": { + "mitre": { + "source_url": "https://example.com/mitre/CVE-2026-44492", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "cveMetadata": { + "cveId": "CVE-2026-44492" + }, + "containers": { + "cna": { + "affected": [ + { + "product": "axios", + "versions": [ + { + "status": "affected", + "lessThan": "1.8.2", + "versionType": "semver" + } + ] + } + ] + } + } + } + }, + "osv": { + "source_url": "https://example.com/osv/CVE-2026-44492", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "id": "CVE-2026-44492", + "affected": [ + { + "package": { + "ecosystem": "npm", + "name": "axios" + }, + "ranges": [ + { + "type": "SEMVER", + "events": [ + { + "introduced": "0" + }, + { + "fixed": "1.8.2" + } + ] + } + ] + } + ] + } + }, + "lifecycle": { + "source_url": "https://example.com/lifecycle", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "supported_streams": [ + "2.1.x", + "2.2.x" + ], + "eol_streams": [] + } + } + }, + "matrix": { + "streams": [ + { + "name": "2.1.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.1.x matrix snapshot", + "rows": [ + { + "version": "2.1.0", + "source_commits": { + "ui": "a003008" + }, + "retag_of": null + }, + { + "version": "2.1.1", + "source_commits": { + "ui": "a003012" + }, + "retag_of": null + } + ] + }, + { + "name": "2.2.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.2.x matrix snapshot", + "rows": [ + { + "version": "2.2.0", + "source_commits": { + "ui": "a004005" + }, + "retag_of": null + }, + { + "version": "2.2.1", + "source_commits": { + "ui": "a004008" + }, + "retag_of": null + }, + { + "version": "2.2.2", + "source_commits": { + "ui": "a004008" + }, + "retag_of": "2.2.1" + }, + { + "version": "2.2.3", + "source_commits": { + "ui": "a004011" + }, + "retag_of": null + }, + { + "version": "2.2.4", + "source_commits": { + "ui": "a004012" + }, + "retag_of": null + } + ] + } + ] + }, + "source_evidence": { + "lock_files": [ + { + "repository": "ui", + "ref": "a003008", + "path": "package-lock.json", + "command": "git show a003008:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/axios\": {\"version\": \"1.7.4\"}}}" + }, + { + "repository": "ui", + "ref": "a003012", + "path": "package-lock.json", + "command": "git show a003012:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/axios\": {\"version\": \"1.7.4\"}}}" + }, + { + "repository": "ui", + "ref": "a004005", + "path": "package-lock.json", + "command": "git show a004005:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/axios\": {\"version\": \"1.7.4\"}}}" + }, + { + "repository": "ui", + "ref": "a004008", + "path": "package-lock.json", + "command": "git show a004008:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/axios\": {\"version\": \"1.7.4\"}}}" + }, + { + "repository": "ui", + "ref": "a004011", + "path": "package-lock.json", + "command": "git show a004011:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/axios\": {\"version\": \"1.9.0\"}}}" + }, + { + "repository": "ui", + "ref": "a004012", + "path": "package-lock.json", + "command": "git show a004012:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/axios\": {\"version\": \"1.9.0\"}}}" + } + ], + "development_streams": [ + { + "repository": "ui", + "ref": "release/0.3.z", + "path": "package-lock.json", + "command": "git show release/0.3.z:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/axios\": {\"version\": \"1.8.2\"}}}" + }, + { + "repository": "ui", + "ref": "release/0.4.z", + "path": "package-lock.json", + "command": "git show release/0.4.z:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/axios\": {\"version\": \"1.8.2\"}}}" + } + ] + }, + "jira_metadata": { + "versions": [ + { + "id": "version-2-1-0", + "name": "RHTPA 2.1.0", + "released": true + }, + { + "id": "version-2-1-1", + "name": "RHTPA 2.1.1", + "released": true + }, + { + "id": "version-2-2-0", + "name": "RHTPA 2.2.0", + "released": true + }, + { + "id": "version-2-2-1", + "name": "RHTPA 2.2.1", + "released": true + }, + { + "id": "version-2-2-2", + "name": "RHTPA 2.2.2", + "released": true + }, + { + "id": "version-2-2-3", + "name": "RHTPA 2.2.3", + "released": true + }, + { + "id": "version-2-2-4", + "name": "RHTPA 2.2.4", + "released": true + }, + { + "id": "version-dev", + "name": "RHTPA 2.2.5", + "released": false + } + ], + "sibling_searches": [ + { + "purpose": "cross-cve-overlap", + "jql": "project = TC AND cf[10632] ~ \"axios\"", + "issues": [ + { + "key": "TC-8008", + "summary": "Other CVE in axios [rhtpa-2.2]", + "status": "In Progress", + "labels": [], + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Same upstream component axios, stream 2.2.x." + } + ] + } + ] + }, + "comments": [], + "links": [ + { + "type": "Depend", + "outward": "TC-8009" + } + ] + } + ] + } + ], + "related_issues": [ + { + "key": "TC-8008", + "summary": "Other CVE in axios [rhtpa-2.2]", + "status": "In Progress", + "labels": [], + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Same upstream component axios, stream 2.2.x." + } + ] + } + ] + }, + "comments": [], + "links": [ + { + "type": "Depend", + "outward": "TC-8009" + } + ] + }, + { + "key": "TC-8009", + "summary": "Bump axios to 1.9.0 in ui [rhtpa-2.2]", + "status": "In Progress", + "labels": [], + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Bump axios to 1.9.0; same component and 2.2.x stream." + } + ] + } + ] + }, + "comments": [], + "links": [] + } + ] + }, + "idempotency": { + "action_markers": [ + "triage-security:tc-8010:link:related:tc-8008" + ], + "existing_remediation": [] + }, + "authorization": { + "mutation_authorized": true + } +} diff --git a/evals/triage-security/files/fullsend-eval-9-trusted-input.json b/evals/triage-security/files/fullsend-eval-9-trusted-input.json new file mode 100644 index 000000000..c0823bb2d --- /dev/null +++ b/evals/triage-security/files/fullsend-eval-9-trusted-input.json @@ -0,0 +1,423 @@ +{ + "schema_version": "1", + "issue": { + "key": "TC-8011", + "summary": "CVE-2026-45678 webpack - Arbitrary Code Execution via loader chain [rhtpa-2.2]", + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "A vulnerability was found in webpack. The webpack package before version 5.98.0 allows arbitrary code execution through a specially crafted loader chain configuration. An attacker with control over a project's webpack configuration can execute arbitrary code during the build process.\n\n**Affected package**: webpack\n**Affected versions**: versions before 5.98.0\n**Fixed version**: 5.98.0\n**CVSS**: 7.8 (High)\n\nThe vulnerability exists because webpack does not properly sanitize loader paths when resolving the loader chain, allowing path traversal to execute arbitrary modules.\n\n### References\n\n- https://github.com/advisories/GHSA-2026-wk55-m3rr\n\n### Related CVE Jiras (provided by JQL search on customfield_10632 = \"webpack\")\n\nThe following related CVE Jira was found affecting the same upstream component:\n\n- **TC-8012** — CVE-2026-43210 webpack - ReDoS in chunk name validation [rhtpa-2.2]\n - **Status**: Closed (Done)\n - **Labels**: CVE-2026-43210, pscomponent:org/rhtpa-ui\n - **customfield_10632**: webpack\n - **customfield_10669**: pscomponent:org/rhtpa-ui\n - **customfield_10832**: rhtpa-2.2\n - **Issue Links**:\n - **Depend**: TC-8013 (remediation Task)\n - **Summary**: Bump webpack to 5.96.1 in rhtpa-ui [rhtpa-2.2]\n - **Status**: Closed (Done)\n - **Description excerpt**: \"Bump webpack from 5.95.0 to 5.96.1 to resolve CVE-2026-43210. The fix requires webpack >= 5.96.0.\"" + } + ] + } + ] + }, + "status": "New", + "labels": [ + "CVE-2026-45678", + "pscomponent:org/rhtpa-ui" + ], + "versions": [], + "reporter": { + "account_id": "synthetic-reporter", + "display_name": "Synthetic Reporter" + }, + "comments": [], + "fields": { + "fixture_purpose": "SYNTHETIC TEST DATA — independent Fullsend counterpart of interactive eval 9; all identities, URLs, CVEs, commit pins and evidence are deliberate public test material, not live advisory evidence.", + "current_user": { + "accountId": "synthetic-engineer", + "displayName": "Synthetic Engineer" + }, + "ecosystem": "npm", + "affected_package": "webpack", + "fixed_version": "5.98.0", + "affected_range": "< 5.98.0", + "affectsVersions": [ + { + "id": "version-stale", + "name": "RHTPA 2.0.0" + } + ], + "assignee": null, + "issuelinks": [], + "customfield_10632": "webpack" + } + }, + "remote_links": [ + { + "url": "https://example.com/advisories/CVE-2026-45678", + "title": "Synthetic CVE-2026-45678 advisory" + } + ], + "configuration": { + "project_key": "TC", + "jira_version_prefix": "RHTPA", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "version_streams": [ + { + "name": "2.1.x", + "matrix_path": "2.1.x/security-matrix.md", + "release_repository": "release-repo" + }, + { + "name": "2.2.x", + "matrix_path": "2.2.x/security-matrix.md", + "release_repository": "release-repo" + } + ], + "source_repositories": [ + { + "name": "ui", + "url": "https://github.com/example/ui", + "deployment_context": "customer-shipped" + } + ], + "vex_justification_field": "customfield_10665", + "upstream_affected_component_field": "customfield_10632", + "stream_field": "customfield_10832" + }, + "external_evidence": { + "mitre": { + "source_url": "https://example.com/mitre/CVE-2026-45678", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "cveMetadata": { + "cveId": "CVE-2026-45678" + }, + "containers": { + "cna": { + "affected": [ + { + "product": "webpack", + "versions": [ + { + "status": "affected", + "lessThan": "5.98.0", + "versionType": "semver" + } + ] + } + ] + } + } + } + }, + "osv": { + "source_url": "https://example.com/osv/CVE-2026-45678", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "id": "CVE-2026-45678", + "affected": [ + { + "package": { + "ecosystem": "npm", + "name": "webpack" + }, + "ranges": [ + { + "type": "SEMVER", + "events": [ + { + "introduced": "0" + }, + { + "fixed": "5.98.0" + } + ] + } + ] + } + ] + } + }, + "lifecycle": { + "source_url": "https://example.com/lifecycle", + "retrieved_at": "2026-09-30T12:00:00Z", + "status": 200, + "body": { + "supported_streams": [ + "2.1.x", + "2.2.x" + ], + "eol_streams": [] + } + } + }, + "matrix": { + "streams": [ + { + "name": "2.1.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.1.x matrix snapshot", + "rows": [ + { + "version": "2.1.0", + "source_commits": { + "ui": "a003008" + }, + "retag_of": null + }, + { + "version": "2.1.1", + "source_commits": { + "ui": "a003012" + }, + "retag_of": null + } + ] + }, + { + "name": "2.2.x", + "matrix_source": "SYNTHETIC TEST DATA — trusted 2.2.x matrix snapshot", + "rows": [ + { + "version": "2.2.0", + "source_commits": { + "ui": "a004005" + }, + "retag_of": null + }, + { + "version": "2.2.1", + "source_commits": { + "ui": "a004008" + }, + "retag_of": null + }, + { + "version": "2.2.2", + "source_commits": { + "ui": "a004008" + }, + "retag_of": "2.2.1" + }, + { + "version": "2.2.3", + "source_commits": { + "ui": "a004011" + }, + "retag_of": null + }, + { + "version": "2.2.4", + "source_commits": { + "ui": "a004012" + }, + "retag_of": null + } + ] + } + ] + }, + "source_evidence": { + "lock_files": [ + { + "repository": "ui", + "ref": "a003008", + "path": "package-lock.json", + "command": "git show a003008:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/webpack\": {\"version\": \"5.95.0\"}}}" + }, + { + "repository": "ui", + "ref": "a003012", + "path": "package-lock.json", + "command": "git show a003012:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/webpack\": {\"version\": \"5.95.0\"}}}" + }, + { + "repository": "ui", + "ref": "a004005", + "path": "package-lock.json", + "command": "git show a004005:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/webpack\": {\"version\": \"5.95.0\"}}}" + }, + { + "repository": "ui", + "ref": "a004008", + "path": "package-lock.json", + "command": "git show a004008:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/webpack\": {\"version\": \"5.95.0\"}}}" + }, + { + "repository": "ui", + "ref": "a004011", + "path": "package-lock.json", + "command": "git show a004011:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/webpack\": {\"version\": \"5.96.1\"}}}" + }, + { + "repository": "ui", + "ref": "a004012", + "path": "package-lock.json", + "command": "git show a004012:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/webpack\": {\"version\": \"5.96.1\"}}}" + } + ], + "development_streams": [ + { + "repository": "ui", + "ref": "release/0.3.z", + "path": "package-lock.json", + "command": "git show release/0.3.z:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/webpack\": {\"version\": \"5.98.0\"}}}" + }, + { + "repository": "ui", + "ref": "release/0.4.z", + "path": "package-lock.json", + "command": "git show release/0.4.z:package-lock.json", + "content": "{\"lockfileVersion\": 3, \"packages\": {\"node_modules/webpack\": {\"version\": \"5.96.1\"}}}" + } + ] + }, + "jira_metadata": { + "versions": [ + { + "id": "version-2-1-0", + "name": "RHTPA 2.1.0", + "released": true + }, + { + "id": "version-2-1-1", + "name": "RHTPA 2.1.1", + "released": true + }, + { + "id": "version-2-2-0", + "name": "RHTPA 2.2.0", + "released": true + }, + { + "id": "version-2-2-1", + "name": "RHTPA 2.2.1", + "released": true + }, + { + "id": "version-2-2-2", + "name": "RHTPA 2.2.2", + "released": true + }, + { + "id": "version-2-2-3", + "name": "RHTPA 2.2.3", + "released": true + }, + { + "id": "version-2-2-4", + "name": "RHTPA 2.2.4", + "released": true + }, + { + "id": "version-dev", + "name": "RHTPA 2.2.5", + "released": false + } + ], + "sibling_searches": [ + { + "purpose": "cross-cve-overlap", + "jql": "project = TC AND cf[10632] ~ \"webpack\"", + "issues": [ + { + "key": "TC-8012", + "summary": "Other CVE in webpack [rhtpa-2.2]", + "status": "In Progress", + "labels": [], + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Same upstream component webpack, stream 2.2.x." + } + ] + } + ] + }, + "comments": [], + "links": [ + { + "type": "Depend", + "outward": "TC-8013" + } + ] + } + ] + } + ], + "related_issues": [ + { + "key": "TC-8012", + "summary": "Other CVE in webpack [rhtpa-2.2]", + "status": "In Progress", + "labels": [], + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Same upstream component webpack, stream 2.2.x." + } + ] + } + ] + }, + "comments": [], + "links": [ + { + "type": "Depend", + "outward": "TC-8013" + } + ] + }, + { + "key": "TC-8013", + "summary": "Bump webpack to 5.96.1 in ui [rhtpa-2.2]", + "status": "In Progress", + "labels": [], + "description": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Bump webpack to 5.96.1; same component and 2.2.x stream." + } + ] + } + ] + }, + "comments": [], + "links": [] + } + ] + }, + "idempotency": { + "action_markers": [], + "existing_remediation": [] + }, + "authorization": { + "mutation_authorized": true + } +} diff --git a/evals/triage-security/files/fullsend-gate-interactive-config.md b/evals/triage-security/files/fullsend-gate-interactive-config.md new file mode 100644 index 000000000..229a19913 --- /dev/null +++ b/evals/triage-security/files/fullsend-gate-interactive-config.md @@ -0,0 +1,22 @@ + + +# Project Configuration + +## Repository Registry + +| Repository | Role | Serena Instance | Path | +|---|---|---|---| +| synthetic-project | gate fixture | — | ./ | + +## Jira Configuration + +- Project key: TC +- Cloud ID: synthetic-cloud-no-access + +## Code Intelligence + +No Serena instances configured. + +This deliberate fixture has no Security Configuration. The genuinely absent +Fullsend gate must enter interactive mode, read this file and stop at the existing +Step 0 missing-configuration guard before Jira initialization or credential reads. diff --git a/evals/triage-security/files/fullsend-idempotent-retry.md b/evals/triage-security/files/fullsend-idempotent-retry.md new file mode 100644 index 000000000..c4d0a1f9b --- /dev/null +++ b/evals/triage-security/files/fullsend-idempotent-retry.md @@ -0,0 +1,131 @@ + + +# Fullsend idempotent retry result + +This fixture models a rerun after the synthetic issue already has its labels, +summary comment, remediation task, and related link. The action plan therefore +must cause no duplicate Jira mutation. + +```json +{ + "result": { + "schema_version": "1", + "mode": "mutation-authorized", + "report": { + "issue": "TC-42", + "outcome": "affected", + "summary_markdown": "Synthetic retry result.", + "evidence": [ + { + "source": "https://evidence.invalid/fullsend/retry", + "detail": "Deliberate non-production retry evidence." + } + ] + }, + "actions": [ + { + "type": "field-edit", + "marker": "triage-security:synthetic-labels", + "issue": "TC-42", + "fields": { + "labels": ["ai-cve-triaged"] + } + }, + { + "type": "status-transition", + "marker": "triage-security:synthetic-status", + "issue": "TC-42", + "status": "In Progress" + }, + { + "type": "comment", + "marker": "triage-security:synthetic-summary", + "issue": "TC-42", + "body_adf": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Synthetic summary already posted." + } + ] + } + ] + } + }, + { + "type": "remediation-task", + "marker": "triage-security:synthetic-remediation", + "ref": "remediation", + "project": "TC", + "summary": "Fix synthetic CVE", + "description_adf": { + "type": "doc", + "version": 1, + "content": [ + { + "type": "paragraph", + "content": [ + { + "type": "text", + "text": "Synthetic remediation." + } + ] + } + ] + }, + "labels": [] + }, + { + "type": "link", + "marker": "triage-security:synthetic-link", + "link_type": "Depend", + "inward": "TC-42", + "outward": "{{remediation.key}}" + } + ] + }, + "trusted_input": { + "authorization": { + "mutation_authorized": true + }, + "idempotency": { + "action_markers": [ + "triage-security:synthetic-labels", + "triage-security:synthetic-summary", + "triage-security:synthetic-remediation", + "triage-security:synthetic-link" + ], + "existing_remediation": [ + { + "key": "TC-9001", + "summary": "Fix synthetic CVE", + "labels": ["ai-generated-jira"], + "comments": [ + { + "body": "[sdlc-workflow] Description digest: sha256-adf:synthetic" + } + ] + } + ] + }, + "issue": { + "key": "TC-42", + "status": "In Progress", + "fields": { + "labels": ["ai-cve-triaged"], + "comment": {"comments": [{"body": "Synthetic summary already posted."}]}, + "remediation": {"key": "TC-9001", "summary": "Fix synthetic CVE"}, + "issuelinks": [{ + "type": {"name": "Depend"}, + "outwardIssue": {"key": "TC-9001"} + }] + } + } + } +} +``` diff --git a/evals/triage-security/files/fullsend-invalid-bundle.md b/evals/triage-security/files/fullsend-invalid-bundle.md new file mode 100644 index 000000000..9a0452dab --- /dev/null +++ b/evals/triage-security/files/fullsend-invalid-bundle.md @@ -0,0 +1,40 @@ + + +# Fullsend invalid result + +This fixture deliberately combines a mutation-authorized mode with an empty +field edit. The trusted executor must reject it before attempting a Jira write. + +```json +{ + "result": { + "schema_version": "1", + "mode": "mutation-authorized", + "report": { + "issue": "TC-42", + "outcome": "affected", + "summary_markdown": "Synthetic evidence is contradictory and incomplete.", + "evidence": [ + { + "source": "https://evidence.invalid/fullsend/contradictory", + "detail": "Deliberate contradictory non-production evidence." + } + ] + }, + "actions": [ + { + "type": "field-edit", + "marker": "triage-security:synthetic-empty-edit", + "issue": "TC-42", + "fields": {} + } + ] + }, + "trusted_input": { + "authorization": { + "mutation_authorized": true + }, + "idempotency": {} + } +} +``` diff --git a/evals/triage-security/files/fullsend-invalid-trusted-input.md b/evals/triage-security/files/fullsend-invalid-trusted-input.md new file mode 100644 index 000000000..b6743b14e --- /dev/null +++ b/evals/triage-security/files/fullsend-invalid-trusted-input.md @@ -0,0 +1,12 @@ + + +# Mounted input + +The file at `/sandbox/workspace/.pre-script/triage-security-input.json` is +present but contains the following invalid JSON: + +```json +{"schema_version":"1","issue": +``` + +No other trusted input is available. diff --git a/evals/triage-security/files/fullsend-report-only-trusted-bundle.md b/evals/triage-security/files/fullsend-report-only-trusted-bundle.md new file mode 100644 index 000000000..5cafab0e3 --- /dev/null +++ b/evals/triage-security/files/fullsend-report-only-trusted-bundle.md @@ -0,0 +1,10 @@ + + +# Trusted bundle facts + +The mounted `triage-security-input.json` is already schema-valid. It contains +issue `TC-8101`, authorization `mutation_authorized: false`, and evidence for +these proposed operations: assignment, Affects Versions and label field edits, +an `In Progress` transition, a triage comment, a Depend link, and a remediation +task. The report must name each as withheld. No other files, credentials, or +external evidence are available. diff --git a/evals/triage-security/files/fullsend-report-only-trusted-input.json b/evals/triage-security/files/fullsend-report-only-trusted-input.json new file mode 100644 index 000000000..bef56e395 --- /dev/null +++ b/evals/triage-security/files/fullsend-report-only-trusted-input.json @@ -0,0 +1,12 @@ +{ + "schema_version": "1", + "issue": {"key": "TC-8101", "summary": "CVE-2026-8101", "description": {}, "status": "New", "labels": [], "versions": [], "reporter": {"account_id": "reporter-1", "display_name": "Reporter"}, "comments": [], "fields": {"fullsend_actions": {"withheld": ["assignment", "field edits", "transition", "comment", "link", "remediation task"]}}}, + "remote_links": [{"url": "https://example.com/evidence", "title": "Evidence"}], + "configuration": {"project_key": "TC", "jira_version_prefix": "RHTPA", "vulnerability_issue_type_id": "10016", "component_label_pattern": "pscomponent:", "version_streams": [{"name": "2.2.x", "matrix_path": "security-matrix.md", "release_repository": "release-repo"}], "source_repositories": [{"name": "release-repo", "url": "https://github.com/example/release-repo", "deployment_context": "internal"}]}, + "external_evidence": {"mitre": {"source_url": "https://example.com/mitre", "retrieved_at": "2026-09-21T12:00:00Z", "status": 200, "body": {}}, "osv": {"source_url": "https://example.com/osv", "retrieved_at": "2026-09-21T12:00:00Z", "status": 200, "body": {}}, "lifecycle": {"source_url": "https://example.com/lifecycle", "retrieved_at": "2026-09-21T12:00:00Z", "status": 200, "body": {}}}, + "matrix": {"streams": [{"name": "2.2.x", "matrix_source": "trusted", "rows": [{"version": "2.2.0", "source_commits": {"release-repo": "abcdef0"}, "retag_of": null}]}]}, + "source_evidence": {"lock_files": [{"repository": "release-repo", "ref": "abcdef0", "path": "rpms.lock.yaml", "command": "git show abcdef0:rpms.lock.yaml", "content": "openssl-libs: 3.0.7-1"}], "development_streams": [{"repository": "release-repo", "ref": "main", "path": "rpms.lock.yaml", "command": "git show main:rpms.lock.yaml", "content": "openssl-libs: 3.0.8-2"}]}, + "jira_metadata": {"versions": [], "sibling_searches": [], "related_issues": []}, + "idempotency": {"action_markers": [], "existing_remediation": []}, + "authorization": {"mutation_authorized": false} +} diff --git a/evals/triage-security/files/fullsend-report-only.md b/evals/triage-security/files/fullsend-report-only.md new file mode 100644 index 000000000..c17aecf03 --- /dev/null +++ b/evals/triage-security/files/fullsend-report-only.md @@ -0,0 +1,38 @@ + + +# Fullsend report-only result + +This fixture represents a valid sandbox result whose proposed outcome is reported +to the engineer without any Jira mutation authorization. + +```json +{ + "result": { + "schema_version": "1", + "mode": "report-only", + "report": { + "issue": "TC-42", + "outcome": "needs-review", + "summary_markdown": "Synthetic evidence requires engineer review.", + "evidence": [ + { + "source": "https://evidence.invalid/fullsend/report-only", + "detail": "Deliberate non-production report-only evidence." + } + ] + }, + "actions": [ + { + "type": "report-only", + "marker": "triage-security:synthetic-report-only" + } + ] + }, + "trusted_input": { + "authorization": { + "mutation_authorized": false + }, + "idempotency": {} + } +} +``` diff --git a/evals/triage-security/files/fullsend-rpm-trusted-bundle.md b/evals/triage-security/files/fullsend-rpm-trusted-bundle.md new file mode 100644 index 000000000..5263a939d --- /dev/null +++ b/evals/triage-security/files/fullsend-rpm-trusted-bundle.md @@ -0,0 +1,19 @@ + + +# Trusted bundle facts + +The mounted `triage-security-input.json` is already schema-valid for issue +`TC-8102`, with authorization `mutation_authorized: false`. Its trusted 2.2.x +matrix and `rpms.lock.yaml` evidence identify `openssl-libs` as an RPM package: + +| Release | Lock version | Assessment | +| --- | --- | --- | +| 2.2.0 | 3.0.7-1 | vulnerable | +| 2.2.1 | 3.0.7-1 | vulnerable | +| 2.2.2 | 3.0.7-1 | vulnerable | +| 2.2.3 | 3.0.8-2 | patched | +| 2.2.4 | 3.0.8-2 | patched | + +The same bundle includes an already retrieved SBOM record confirming the RPM +package classification. It supplies no local files, cosign access, Jira access, +git repository, or network access. diff --git a/evals/triage-security/files/fullsend-rpm-trusted-input.json b/evals/triage-security/files/fullsend-rpm-trusted-input.json new file mode 100644 index 000000000..bed5a9dae --- /dev/null +++ b/evals/triage-security/files/fullsend-rpm-trusted-input.json @@ -0,0 +1,12 @@ +{ + "schema_version": "1", + "issue": {"key": "TC-8102", "summary": "CVE-2026-8102 openssl-libs", "description": {}, "status": "New", "labels": [], "versions": [], "reporter": {"account_id": "reporter-1", "display_name": "Reporter"}, "comments": [], "fields": {"rpm_assessment": {"vulnerable": ["2.2.0", "2.2.1", "2.2.2"], "patched": ["2.2.3", "2.2.4"], "sbom": "already-retrieved RPM record"}}}, + "remote_links": [{"url": "https://example.com/evidence", "title": "Evidence"}], + "configuration": {"project_key": "TC", "jira_version_prefix": "RHTPA", "vulnerability_issue_type_id": "10016", "component_label_pattern": "pscomponent:", "version_streams": [{"name": "2.2.x", "matrix_path": "security-matrix.md", "release_repository": "release-repo"}], "source_repositories": [{"name": "release-repo", "url": "https://github.com/example/release-repo", "deployment_context": "internal"}]}, + "external_evidence": {"mitre": {"source_url": "https://example.com/mitre", "retrieved_at": "2026-09-21T12:00:00Z", "status": 200, "body": {}}, "osv": {"source_url": "https://example.com/osv", "retrieved_at": "2026-09-21T12:00:00Z", "status": 200, "body": {}}, "lifecycle": {"source_url": "https://example.com/lifecycle", "retrieved_at": "2026-09-21T12:00:00Z", "status": 200, "body": {}}}, + "matrix": {"streams": [{"name": "2.2.x", "matrix_source": "trusted", "rows": [{"version": "2.2.0", "source_commits": {"release-repo": "abcdef0"}, "retag_of": null}, {"version": "2.2.3", "source_commits": {"release-repo": "abcdef1"}, "retag_of": null}]}]}, + "source_evidence": {"lock_files": [{"repository": "release-repo", "ref": "abcdef0", "path": "rpms.lock.yaml", "command": "git show abcdef0:rpms.lock.yaml", "content": "openssl-libs: 3.0.7-1"}], "development_streams": [{"repository": "release-repo", "ref": "main", "path": "rpms.lock.yaml", "command": "git show main:rpms.lock.yaml", "content": "openssl-libs: 3.0.8-2"}]}, + "jira_metadata": {"versions": [], "sibling_searches": [], "related_issues": []}, + "idempotency": {"action_markers": [], "existing_remediation": []}, + "authorization": {"mutation_authorized": false} +} diff --git a/evals/verify-pr/evals.json b/evals/verify-pr/evals.json index 94fa96088..c7750c91d 100644 --- a/evals/verify-pr/evals.json +++ b/evals/verify-pr/evals.json @@ -42,7 +42,7 @@ }, { "id": 3, - "prompt": "Verify PR #744 for task TC-9103. The task description is in task-with-reviews.md, the PR diff is in pr-diff-with-reviews.md, the review comments are in pr-review-comments.md, and the target repository structure is in repo-backend.md. All CI checks pass. Write your verification report to outputs/report.md. For each review comment, write your classification reasoning to outputs/review-N.md (where N is the comment id). For each sub-task that would be created, write its full description to outputs/subtask-N.md following the task-description-template.md format.", + "prompt": "Verify PR #744 for task TC-9103. The task description is in task-with-reviews.md, the PR diff is in pr-diff-with-reviews.md, the review comments are in pr-review-comments.md, and the target repository structure is in repo-backend.md. All CI checks pass. Write your verification report to outputs/report.md. Inspect only the supplied Reviews JSON list in pr-review-comments.md for Eval Quality detection. Record each actual review's author check against github-actions[bot], body marker check for ## Eval Results, body footer check for sdlc-workflow/run-evals, the number of reviews matching all three checks, and the resulting Eval Quality verdict and its effect on the Test Quality combination in outputs/report.md. When no review matches all three checks, explain why Eval Quality is N/A and does not affect the Test Quality combination. Use the supplied fixture data; do not invent reviews, treat inline comments as review bodies, or call live GitHub/Jira. For each review comment, write your classification reasoning to outputs/review-N.md (where N is the comment id). For each sub-task that would be created, write its full description to outputs/subtask-N.md following the task-description-template.md format.", "expected_output": "A verification report where acceptance criteria mostly pass, and review feedback processing creates sub-tasks for code change requests. The transaction wrapping comment (id 30001) should be classified as a code change request with a sub-task created. The index comment (id 30002) should be classified as a suggestion with no sub-task — the reviewer uses suggestive language and no project convention backs an upgrade. The nit (id 30003) and question (id 30004) should be classified but NOT trigger sub-tasks. Sub-task descriptions must follow task-description-template.md and include Target PR and Review Context extension sections.", "files": ["files/task-with-reviews.md", "files/pr-diff-with-reviews.md", "files/pr-review-comments.md", "files/repo-backend.md"], "assertions": [ diff --git a/fullsend.md b/fullsend.md new file mode 100644 index 000000000..de1165e6d --- /dev/null +++ b/fullsend.md @@ -0,0 +1,616 @@ +# Running sdlc-workflow skills with fullsend + +Run sdlc-workflow skills inside secure sandboxes via +[fullsend](https://github.com/fullsend-ai/fullsend). The agent runs in an +isolated container with least-privilege network and filesystem policies, while +all issue-tracker and GitHub writes happen on the trusted runner — never inside +the sandbox. + +This guide documents the **native CLI path**: a standalone root-level +harness, the stock digest-pinned `fullsend-code` image, a single pinned-URL +registration, and the plugin referenced in place with zero duplication. The only +skill wired up today is `verify-pr`, which runs both on demand (`fullsend run`, +locally or in CI) and automatically on every pull request via a CI-gated +dispatch (see **CI deployment**). + +## How it works + +The setup has three moving parts, all in this repo: + +- **`harness/verify-pr.yaml`** (repo root) — a **standalone harness** with no + base composition. Placed at the repo root so that, when fullsend is invoked + with `--fullsend-dir` = repo root, its relative children resolve against the + repo root: `plugins/sdlc-workflow` is delivered **in place as a whole plugin** + and its sibling `shared/` resources resolve intact. It pins the stock + `fullsend-code` image by digest and declares the policy, provider, profile, + env mount, pre/post scripts, and validation loop. +- **`.fullsend/config.yaml`** — registers a single agent, `verify-pr`, whose + source is the **local composing child** `.fullsend/harness/verify-pr.yaml`. +- **`.fullsend/harness/verify-pr.yaml`** — the local composing child. It pins + the root harness by **raw URL at a commit SHA plus a `#sha256=`** on its + `base:` line, and adds the runner-local *absolute* mounts (GCP credential, + OIDC token, pre-script output) that fullsend v0.37.0 refuses to inherit from a + URL-sourced base. `host_files` are concatenated base + child (dedup by dest, + child wins). + +Because the base is pinned to one commit SHA, **every relative child of the +harness** — the pre/post scripts, schemas, agent prompt, policy, provider, +profile, and the whole `plugins/sdlc-workflow` directory — resolves at that same +SHA. The single `#sha256=` on the `base:` line verifies the base file itself; +`.fullsend/lock.yaml` then freezes the content hash of every transitively +resolved child. One pinned URL, one integrity anchor, everything else covered +transitively. + +There is **no custom image and no marketplace baking**. fullsend fabricates the +Claude Code marketplace cache from the in-place `plugins:` entry at runtime, so +the Claude Code marketplace structure is unchanged and the same plugin files +serve interactive Claude Code users and fullsend runs with zero duplication. + +### Split-trust I/O + +The Jira and GitHub tokens live **only on the runner**. The pre_script prefetches +everything the agent needs (the Jira issue and a GitHub read bundle) and the +post_script performs every write after the sandbox is destroyed. The sandbox +itself receives read-only context only — `JIRA_BASE_URL` (display links) — plus +the prefetched read bundle mounted read-only. The Jira key is not passed as an +env var: the pre_script derives it from the triggering PR URL and records it as +`task_id` inside `verify-pr-input.json`, which the agent reads. The **only** +credential that ever enters the sandbox is the Vertex AI service-account key, +because model inference runs in-sandbox (see below). + +## Credential delivery and tiers + +Credentials use the highest isolation tier possible. Two of the three services +never place a credential in the sandbox at all. + +| Service | Tier | Credential in sandbox? | How it is delivered | +|---|---|---|---| +| **Jira** | 1 | No | The pre_script prefetches the issue on the runner; the post_script posts comments with `fullsend issues post-comment --tracker jira`. The token stays in runner env only. | +| **GitHub** | 1 | No | The pre_script prefetches the read bundle (`gh pr diff` / `gh pr view` → diff, diffstat, reviews, comments, commits), the head-SHA **CI check-run outcomes** (`github.check_runs` — name/status/conclusion/details_url), and the **concatenated `--log-failed` output of every failed check-run** (written to `check-run-logs.txt`, mounted separately; only its path rides in the bundle so the large log text stays off the agent's context until Correctness Check 1b reads it on a FAIL), and records the PR head ref name and head commit SHA in it; the post_script writes via `fullsend issues post-comment --tracker github` and `gh` (PR reviews/replies via `gh api`). The token stays in runner env only. The PR-head working tree is the `--target-repo` clone, checked out at the head SHA before the run (see below) — the pre_script does not check out. | +| **Vertex AI** | 4 (fullsend-mandated) | **Yes** | The in-sandbox runtime does the model inference and reads `GOOGLE_APPLICATION_CREDENTIALS` from a file (`/tmp/.gcp-credentials.json`). Vertex auth requires local JWT signing, so tier 4 (file on the sandbox filesystem) is unavoidable. This is the one credential set in the sandbox. | + +Because Vertex is the only in-sandbox credential, the sandbox's **only network +egress is `*.googleapis.com`** (declared in +`plugins/sdlc-workflow/profiles/fullsend-vertex-ai.yaml`). There is no +atlassian, github, or anthropic egress from the sandbox — those are all handled +on the runner. + +> GitHub is a **tier-1** service — it needs no OpenShell provider. There is no +> Jira provider and no `github-ro` provider. The only provider is +> `plugins/sdlc-workflow/providers/vertex-ai.yaml`, which selects the Vertex +> egress profile (its credential still arrives via the host-file mount, not a +> proxied placeholder). + +## Prerequisites + +- [fullsend](https://github.com/fullsend-ai/fullsend) CLI installed (CI pins + **v0.43.0** — see **Bumping the fullsend version**; a local run needs a + compatible release). +- An OpenShell gateway running — fullsend uses OpenShell as its sandbox runtime. +- GCP credentials for Vertex AI — a service-account key JSON (local) or a WIF + external-account config (CI). Referenced by `GOOGLE_APPLICATION_CREDENTIALS`. +- A Jira API token and a GitHub token — used by the pre/post scripts on the + runner only. +- The Python `jsonschema[format]` package on the runner — the `validation_loop` + validates agent output against the JSON schema before the post_script runs, and + the triage-security pre-script enforces URI and date-time schema formats + (`pip install 'jsonschema[format]'`). + +## Running verify-pr + +The command is **identical locally and in CI** — same harness, same registration, +only the runtime environment differs (a local SA key vs. a CI WIF config; a CI +run additionally supplies an OIDC token file): + +```bash +fullsend run verify-pr \ + --fullsend-dir .fullsend \ + --target-repo /tmp/my-repo-clone \ + --env-file secrets.env \ + --env-file <(echo "FULLSEND_WORK_ITEM_URL=https://github.com/RHEcosystemAppEng/sdlc-plugins/pull/1234") +``` + +The Jira key is **not** an input — the pre_script resolves it from the triggering +PR URL (`FULLSEND_WORK_ITEM_URL`) by a JQL search on the Git Pull Request custom +field, gates it (status Review + the `ai-generated-jira` label), and records the +resolved key as `task_id` in `verify-pr-input.json`. In CI, `harness-run` exports +`FULLSEND_WORK_ITEM_URL`; a local run must export the same PR URL. + +`--target-repo` must be a **disposable clone**, not your working directory — +fullsend deletes and re-creates it after each run. It must already be **checked +out at the PR head** (`github.commit_sha` from the read bundle): neither the +pre_script nor the harness performs the checkout, so the caller establishes the +PR-head working tree. Locally, `gh pr checkout ` in the clone before +the run; in CI, the checkout step fetches the PR head. The sandbox then inspects +that tree directly (there is no `gh` CLI in the sandbox). + +A minimal `secrets.env` (never commit it): + +```bash +# Vertex AI (tier 4 — the only credential that enters the sandbox) +ANTHROPIC_VERTEX_PROJECT_ID=my-project +GOOGLE_CLOUD_PROJECT=my-project +CLOUD_ML_REGION=global +GOOGLE_APPLICATION_CREDENTIALS=/path/to/sa-key.json +# Jira (tier 1 — runner only) +JIRA_SERVER_URL=https://myorg.atlassian.net +JIRA_EMAIL=me@example.com +JIRA_API_TOKEN=my-jira-token +JIRA_PROJECT_KEY=TC +# GitHub (tier 1 — runner only) +GH_TOKEN=my-github-token +``` + +## CI deployment (per-repo GitHub App) + +The same pinned harness that runs locally is dispatched automatically in CI, once +per pull request, after the PR's other checks finish. Three one-time provisioning +steps stand this up (day-0, run by a repo/org admin); the recurring run then needs +no manual step. + +### Installation + +1. **`fullsend github setup`** — provisions the per-repo installation. It writes + the two-layer config (`.fullsend/config.base.yaml`, the vendor preset base + layer, and `.fullsend/config.yaml`, the repo overlay) and installs the vendor + shim workflow `.github/workflows/fullsend.yaml`. `install_mode: per-repo` + throughout. **verify-pr is registered only in `config.base.yaml`**, never in + `config.yaml`: the shim greps `config.yaml` alone for agents, so it never + dispatches verify-pr, while `fullsend dispatch` / `fullsend run` still resolve + it because `LoadConfig` merges `config.yaml` over `config.base.yaml` (ADR 0069). +2. **`fullsend inference provision`** — provisions hosted Vertex AI via Workload + Identity Federation and records the result as the inference override in + `.fullsend/config.yaml`: + ```yaml + inference: + project: it-gcp-tpa + wif_provider: projects/442181572212/locations/global/workloadIdentityPools/fullsend-inference/providers/gh-rhecosystemappeng-sdlc-plugin + ``` + In CI the WIF provider mints a short-lived Vertex credential per run — no + long-lived key is stored (contrast the local run, which uses an SA key file). +3. **Org GitHub App install** — install the fullsend GitHub App on the org/repo + with the **review** role. The App mints the review-role token via OIDC at + dispatch time (the `mint_url` input) for the PR review/report writes. This is an + org-admin action in GitHub, not a CLI step. + +The recurring dispatch is a dedicated workflow, +`.github/workflows/fullsend-verify-pr.yml` (not the vendor shim) — see +**Dispatch design** below. + +### Dispatch design + +verify-pr must run **after** a PR's other CI finishes, so it can treat CI results +as input data — never as a gate (it runs whether CI passed or failed). fullsend +has no CI-completion event to trigger on, which drives the design: + +- There is **no `check_suite` / `check_run` / `workflow_run` TransitionKind** in + fullsend, and GitHub's recursion guard suppresses `check_suite` / `check_run` + for suites created by Actions. A CI-completion-triggered dispatch is therefore + not reliably deliverable. +- The workaround is an **inline `pull_request_target` wait**: + `fullsend-verify-pr.yml` triggers on `pull_request_target` (`opened`, + `synchronize`, `reopened`, `labeled`) and its `wait-for-checks` job blocks on + [`lewagon/wait-on-check-action`](https://github.com/lewagon/wait-on-check-action) + (pinned by SHA) until every *other* check on the PR head reaches a terminal + state. Only then does the `verify-pr` job dispatch. This inline wait is the only + reliable "run after CI" trigger (TC-6180 NFRs). The trigger is + `pull_request_target` (not `pull_request`) so **fork-origin PRs** get the + upstream vars/secrets/OIDC the mint step needs — see **Fork PRs** below. +- **CI result is data, not a gate.** `allowed-conclusions` lists *all* terminal + conclusions (`success`, `failure`, `neutral`, `cancelled`, `skipped`, + `timed_out`, `action_required`, `stale`, `startup_failure`), so a failing check + still releases the wait and verify-pr still runs and reports (recording + `CI Status = FAIL`). + `fail-on-no-checks: false` lets a PR with no other checks proceed; + `timeout-minutes: 45` bounds a hung check (the dispatch is then skipped). +- One in-flight run per PR: a `concurrency` group keyed on the PR number with + `cancel-in-progress: true` supersedes a stale run on a new push. + +The dispatch job hands a pre-built single-entry matrix (agent `verify-pr`, role +`review`) to the vendor reusable workflow +`fullsend-ai/fullsend/.github/workflows/reusable-dispatch.yml@v0`, whose +`harness-run` job mints the review-role App token via OIDC and runs `fullsend run`. + +This `pull_request_target` trigger is the CI-side equivalent of the harness's own +**CEL trigger**, which drives the `fullsend dispatch` path. The trigger is declared on +the composing child `.fullsend/harness/verify-pr.yaml`: + +``` +event.entity.kind == "change_proposal" && +event.transition.kind in ["synchronized", "opened", "reopened"] +``` + +It **must** live on the child, not only on the base: `mergeBaseIntoChild` +(`internal/harness/compose.go`) carries `role` and other scalars from base→child +but intentionally does **not** carry `trigger`, so a trigger set only on the base +is inert (ADR 0061). + +Before the agent runs, the **pre_script gates on Jira** (see **Running +verify-pr**): it resolves the Jira key from the PR URL by a JQL search on the Git +Pull Request custom field and proceeds only when the issue is status `Review` with +the `ai-generated-jira` label (ADR 0072 skip otherwise). This scopes automated +review to sdlc-workflow-tracked PRs. + +#### Fork PRs (`pull_request_target` + `ok-to-test`) + +The workflow triggers on **`pull_request_target`**, not `pull_request` (TC-6331). +A `pull_request` run whose head branch lives on a **fork** gets no repo +vars/secrets and no OIDC (`id-token`), so `vars.FULLSEND_MINT_URL` resolves empty +and the mint step fails (`FULLSEND_MINT_URL is not set`). `pull_request_target` +runs the **base-branch** version of the workflow with full access to upstream +vars/secrets/OIDC, so a fork PR can dispatch (fullsend **ADR-0009**). + +This is only safe because the workflow **never checks out or executes fork code**. +It builds the dispatch matrix from `github.event` alone (`GITHUB_EVENT_PATH`, via +`jq --argjson` — no shell interpolation of PR-controlled strings) and hands off to +`reusable-dispatch.yml`, which does its own checkout/mint in base-repo context. +That closes the classic `pull_request_target` "pwn request" credential-exfiltration +hole. + +Defense-in-depth is an **`ok-to-test` maintainer-label gate**, modelled on +fullsend's `docs/guides/dev/e2e-testing.md`: + +- **Same-repo (upstream) PRs** need no label — they already carry secrets/OIDC — + and run on `opened` / `synchronize` / `reopened`. +- **Fork PRs** dispatch **only** when a maintainer applies the `ok-to-test` label + (the `labeled` trigger). A maintainer reviews the fork's code first, then labels. +- The `remove-ok-to-test` job strips the label on every new push (`synchronize`) + to a fork PR, so new commits force re-review before the label — and thus another + verify-pr run — can be re-applied. + +There is **no longer an "upstream head branch required" constraint**: fork +contributors open PRs normally, and a maintainer gates each run with the label. + +### Variables and secrets (day-2 management) + +CI configuration is split between GitHub Actions **variables** (non-secret, +`vars.*`) and **secrets** (`secrets.*`), all consumed by +`.github/workflows/fullsend-verify-pr.yml`. To change any value, edit it under the +repo's **Settings → Secrets and variables → Actions** (or with `gh variable set` / +`gh secret set`) — no code change is needed, and the next dispatch picks it up. + +| Name | Kind | Purpose | Update when | +|---|---|---|---| +| `JIRA_EMAIL` | secret | Jira Service Account email for the tier-1 Basic-auth credential (mapped to vendor input `JIRA_USER_EMAIL`); authors the report comment. | Changing the Jira Service Account. | +| `JIRA_API_TOKEN` | secret | Jira scoped API token for the same account (mapped to `JIRA_TOKEN`). | Rotating the SA token (tokens expire / are revoked). | +| `JIRA_BASE_URL` | variable | Jira site URL for display links inside the sandbox. | Jira site moves. | +| `FULLSEND_GCP_WIF_PROVIDER` | secret | Vertex Workload Identity Federation provider resource. | Re-provisioning inference (`fullsend inference provision`). | +| `FULLSEND_GCP_PROJECT_ID` | secret | GCP project hosting Vertex inference. | GCP project changes. | +| `FULLSEND_MINT_URL` | variable | OIDC mint endpoint for the review-role App token. | App / mint endpoint changes. | +| `FULLSEND_GCP_REGION` | variable | Vertex region. | Region changes. | +| `FULLSEND_PROJECT_NUMBER` | variable | GCP project number for OIDC. | GCP project changes. | +| `OTEL_EXPORTER_OTLP_HEADERS`, `OTEL_EXPORTER_OTLP_TRACES_HEADERS` | secret | OpenTelemetry export auth headers (optional observability). | Rotating the OTLP endpoint credential. | + +The Jira credential is a dedicated **Service Account** — email + scoped token, +Basic auth (fullsend's Jira tracker is Basic-auth only). It is a **tier-1** +credential: it stays in runner env and never enters the sandbox (see **Credential +delivery and tiers**). To rotate it, issue a new scoped token for the SA and +update `JIRA_API_TOKEN` (and `JIRA_EMAIL` if the account itself changes); the next +dispatch authenticates the post_script Jira comment with the new value. The same +pattern applies to every row above — update the value in repo settings, no code +change. + +## Releasing an update (this repo) + +The `.fullsend/` registration pins content by commit SHA, so **editing a file is +not enough** — you must re-pin and re-lock in the same change. The full loop: + +1. **Edit** the harness (`harness/verify-pr.yaml`) and/or any pinned plugin file + under `plugins/sdlc-workflow/` (see the re-pin rule below). +2. **Commit and push** to `RHEcosystemAppEng`. `main` is the release channel; + during feature work, push to the feature branch and merge to `main`. +3. **Re-pin** the `base:` line in `.fullsend/harness/verify-pr.yaml`: set the + commit SHA to the new commit and the `#sha256=` to the base file's hash: + ```bash + shasum -a 256 harness/verify-pr.yaml + ``` +4. **Re-lock** so every transitive child is re-resolved at the new SHA: + ```bash + fullsend lock verify-pr --fullsend-dir .fullsend + ``` + `fullsend lock --all --fullsend-dir .fullsend` locks every harness at once — + equivalent here since `verify-pr` is the only one (TC-5813 used `--all`). +5. **Commit `.fullsend/`** (the updated `lock.yaml` and re-pinned child) in the + same PR as the content edit. + +> **`fullsend agent update` does not apply to this repo.** The registered +> `verify-pr` agent is a *local path* (`source: harness/verify-pr.yaml`), so +> `fullsend agent update verify-pr` fails with *"agent 'verify-pr' is a local +> path — nothing to update"*. `agent update` re-pins **URL** agents only — it is +> for adopters (below), not for the self-hosted `.fullsend/` model. Here you +> re-pin by editing the `base:` SHA/hash and re-locking. + +### Bumping the fullsend version + +The CI dispatch pins the fullsend CLI to a released tag so it does not float +(the vendor default floats to `job.workflow_sha`). The pin lives in +`.github/workflows/fullsend-verify-pr.yml`: + +```yaml + fullsend_version: "v0.43.0" +``` + +The current pin is **v0.43.0** — bumped from v0.37.0 because v0.37.0's bundled +OpenShell never reached sandbox readiness (`not ready after 2m0s` → `Failed to +create sandbox`). To bump: + +1. Set `fullsend_version` to the new released tag (v-prefixed, matching the + vendor's git tags). +2. **Re-validate end to end**: open a throwaway qualifying PR and confirm the + dispatch reaches sandbox readiness and posts a report (an OpenShell/runtime + regression surfaces as `Failed to create sandbox`). Only merge once a live run + passes. +3. Update the local `## Prerequisites` version note to match if the local CLI + must track CI. + +The fullsend CLI version is independent of both the harness `base:` re-pin and the +plugin version below. + +### Keeping the two plugin-version files in sync + +The plugin/skill version is stored in **two** files that must always match: + +- `plugins/sdlc-workflow/.claude-plugin/plugin.json` — the plugin manifest + (required by CI validation). +- `.claude-plugin/marketplace.json` — the marketplace registry (required for + relative-path plugins). + +When releasing a skill/plugin change, bump **both** files in the same change (both +are currently `0.13.9`). This version bump is orthogonal to — but released +alongside — the `base:` re-pin + re-lock loop above: the re-pin/re-lock is what +makes the edited plugin content take effect at runtime, while the version bump is +the human-facing marker of that release. + +### Constraints and gotchas + +- **Merge with a merge commit — never squash or rebase.** The pinned raw URL in + `.fullsend/` (and the adopter `fullsend agent add …/blob/main/…` URL) resolves + to a specific commit SHA. A squash or rebase merge rewrites or drops that SHA, + so the pin — and any adopter fetch — 404s. This applies at **every** merge hop + up to `main`. +- **Re-pin + re-lock is mandatory after editing ANY pinned plugin file** — + `SKILL.md`, scripts, schemas, sub-skill templates, `plugin.json`, and so on. + Each such file is locked as a *member* of the `plugins[0]` directory dependency + in `.fullsend/lock.yaml`. An edit that is not paired with a re-pin/re-lock is + served **stale** at runtime: the pinned pre-edit content runs. Pair every + content edit with edit → commit → push → re-pin `base:` → `fullsend lock …` → + commit `.fullsend/` in the same PR. +- **Scope — what re-pinning does *not* affect.** The re-pin/re-lock loop governs + only the `fullsend run` path (local + CI). Interactive / marketplace installs + of the skill ignore `.fullsend/lock.yaml` entirely and use the plugin's + working-tree copy — so a missed re-pin never affects interactive users, only + fullsend runs. +- **Troubleshooting `cache integrity check failed`** at harness/lock load: caused + by a stale `.fullsend/.fullsend-cache` (and the sibling root `./.fullsend-cache`). + Fix: + ```bash + rm -rf .fullsend/.fullsend-cache .fullsend-cache + ``` + then re-run online so the pinned content re-fetches. + +## Adopting verify-pr (external fullsend users) + +Other repos do not need to clone sdlc-plugins or copy any files. Register the +harness by URL — no SHA needed, fullsend resolves `main` HEAD and pins it for +you: + +```bash +# 1. Register (one command, SHA-free — fullsend pins main HEAD) +fullsend agent add https://github.com/RHEcosystemAppEng/sdlc-plugins/blob/main/harness/verify-pr.yaml + +# 2. Lock and run +fullsend lock verify-pr --fullsend-dir .fullsend +fullsend run verify-pr --fullsend-dir .fullsend --target-repo /tmp/clone --env-file secrets.env + +# 3. Later, pull in a new release +fullsend agent update verify-pr --fullsend-dir .fullsend +``` + +`main` is the **release channel** — the URL above tracks it. There is no separate +`release` branch yet because `fullsend agent update` does not track a branch: +it re-pins to a commit SHA you resolve at update time, so a dedicated release +branch would add a maintenance hop without buying reproducibility that the SHA +pin does not already provide. The **merge-commit-only** rule above applies to +adopters too — a squash/rebase on any hop to `main` breaks the pinned URL the +adopter fetched. + +## File inventory + +All plugin paths are relative to `plugins/sdlc-workflow/`. + +| File | Purpose | +|---|---| +| `harness/verify-pr.yaml` (repo root) | Standalone harness — stock digest-pinned image, in-place plugin, policy, provider, profile, env mount, pre/post scripts, validation loop, split-trust env. | +| `.fullsend/config.yaml` | Registers the `verify-pr` agent (local source) and the single allowed remote-resource prefix. | +| `.fullsend/harness/verify-pr.yaml` | Local composing child — pins the root harness by raw URL (`base:` + `#sha256=`) and adds the runner-local absolute mounts. | +| `.fullsend/lock.yaml` | Generated by `fullsend lock` — freezes every transitive dependency URL + SHA256. Do not edit by hand. | +| `agents/verify-pr.md` | Agent prompt (YAML frontmatter). fullsend launches Claude Code with this as the system prompt; it reads `task_id` from the pre-fetched `verify-pr-input.json` and invokes the skill. | +| `policies/verify-pr.yaml` | Sandbox network/filesystem policy. | +| `profiles/fullsend-vertex-ai.yaml` | OpenShell egress profile — `*.googleapis.com:443` only. | +| `providers/vertex-ai.yaml` | Selects the Vertex egress profile (no proxied credential). | +| `env/gcp-vertex.env` | Vertex env template, expanded from the secrets file (`expand: true`); points `GOOGLE_APPLICATION_CREDENTIALS` at `/tmp/.gcp-credentials.json`. | +| `schemas/verify-pr-result.schema.json` | JSON Schema for the agent's structured output; enforced by `validation_loop`. | +| `scripts/pre-verify-pr.sh` | Pre_script — validates inputs, prefetches the Jira issue and the GitHub read bundle, and records the PR head ref name + head commit SHA in the bundle (it does **not** check out — the PR-head working tree comes from `--target-repo`). Delegates to `pre_verify_pr.py`. | +| `scripts/post-verify-pr.sh` | Post_script — finds `agent-result.json` and delegates to `execute-actions.py`. Runs on the trusted runner after the sandbox is destroyed. | +| `scripts/execute-actions.py` | Action executor — posts Jira/GitHub sticky comments via `fullsend issues post-comment`, and PR reviews/replies via `gh api`. | +| `scripts/validate-output-schema.sh` + `strip_extra_properties.py` | Strips benign agent-added metadata, then validates against the schema. | + +## Design decisions + +### Why a standalone root-level harness + +Placing the harness at the repo root lets fullsend resolve its relative children +against the repo root, so `plugins/sdlc-workflow` is delivered in place as a +whole plugin and its `shared/` resources resolve intact. No base composition +keeps the harness self-contained — the Go defaults already supply security +(`enabled: true`, `fail_mode: closed`) plus the sandbox hooks, so no explicit +`security:` block is needed. + +### Why the plugin is referenced in place (no duplication) + +fullsend fabricates the Claude Code marketplace cache from the `plugins:` entry +at runtime. The same plugin files back both interactive Claude Code installs and +fullsend runs, so there is no custom image, no Dockerfile, no bootstrap script, +and no second copy of the skills to keep in sync. + +### Why Jira reads/writes use native fullsend CLI (not MCP) in the sandbox + +MCP servers are not available inside the sandbox. All tracker and GitHub I/O is +handled on the runner: the pre_script prefetches with `gh` and the Jira REST +client, and the post_script posts sticky comments with +`fullsend issues post-comment`, which creates a marker-tagged comment on the +first run and edits it in place on re-runs (no comment flooding). + +### Why the stock digest-pinned image + +The harness pins `ghcr.io/fullsend-ai/fullsend-code` by `@sha256:` digest. The +stock image already carries Claude Code, `git`, the `gh` CLI, Python, and the +fullsend security tooling, so there is nothing to add — the sandbox *policy*, not +the image, is the enforcement layer. + +## Acceptance run + +The pinned agent was proven end-to-end with the exact command CI runs (TC-5815). +The command is identical locally and in CI (see **Running verify-pr**) — same +harness, same registration, same pinned content; only the runtime environment +differs (a local SA key vs. a CI WIF config). + +**Run** — `verify-pr` against **TC-6137 / PR #294** at commit `c0bbad9`, off the +base pin `6572360e`, on 2026-09-09: + +```bash +fullsend run verify-pr --fullsend-dir .fullsend \ + --target-repo /tmp/verify-pr-clone \ + --env-file --env-file \ + --keep-sandbox +``` + +`--target-repo` was a **disposable clone checked out at the PR head**, never the +working directory (fullsend deletes it after the run — see **Known issues**). + +**Results:** + +| Acceptance criterion | Result | +|---|---| +| Sandbox-mode run off the pin | ✅ model `claude-opus-4-6`, 1 iteration, 39 turns, 52 tool calls, 4 sub-agents dispatched | +| Sub-agents dispatched within 30 min | ✅ sandbox wall-clock ≈ 9.5 min | +| Schema-valid `agent-result.json` | ✅ agent exit 0, `Validation: passed` | +| Zero `*.atlassian.net` egress from sandbox | ✅ none | +| Zero `api.github.com` egress from sandbox | ✅ none | +| Real writes (full run, post-script on runner) | ✅ sticky report **edited in place** on PR #294 (idempotent, no flood) + posted to Jira TC-6137 | + +**DENIED endpoints — expected, none added.** The sandbox's only allowed egress is +`*.googleapis.com` (Vertex). The OCSF sandbox log records the least-privilege +policy correctly blocking every other attempted endpoint; these DENIED entries are +the control **working**, not failures, and no endpoint needed to be added to make +the run pass: + +| Blocked endpoint | Process | Why it is correct | +|---|---|---| +| `github.com:443` | `git-remote-http`, `tirith` | Split-trust — the sandbox never touches GitHub; the PR-head tree arrives via `--target-repo` and writes happen on the runner. | +| `raw.githubusercontent.com:443` | `tirith` | Harness/plugin content is delivered pre-resolved from the pin; no in-sandbox fetch. | +| `downloads.claude.ai:443` | `claude` | Claude CLI self-update/telemetry — irrelevant to the run and correctly denied. | + +The run's verdict on PR #294 was `FAIL` (verify-pr's assessment of that PR at +`c0bbad9`); the **run mechanics** above are what this acceptance proves, and they +all pass. This is one of several real production runs of the pinned agent during +Epic D — see the sticky `verify-pr` reports on PRs #292, #293, and #294. + +### E2E acceptance via CI dispatch (TC-6192) + +TC-5815 above proved the *agent*; TC-6192 proves the *CI-gated dispatch path* end +to end: a qualifying PR triggers `.github/workflows/fullsend-verify-pr.yml`, the +`wait-for-checks` job waits for the PR's other checks to reach a terminal state, +and the `verify-pr` reusable-dispatch job then mints the review-role App token via +OIDC and posts the report — whether CI passed or failed (TC-6180 Reqs 4, 6, 7). + +**Vehicle** — a throwaway PR **#300** (`tc-6192-e2e-acceptance` → +`verify-pr-fullsend`, head branch on the **upstream** repo so `pull_request` +secrets + OIDC are available — a fork head would get neither; this pre-TC-6331 +constraint is now lifted, see **Fork PRs**), qualified against +Jira task **TC-6254** (status `Review`, label `ai-generated-jira`, Git PR field = +PR #300 URL). The Jira report comment is authored by the tier-1 Jira account +configured for that run (**Marco Rizzi**). + +**Results — all five acceptance criteria proven:** + +| Acceptance criterion | Result | +|---|---| +| Qualifying PR → CI finishes → verify-pr dispatches | ✅ `wait-for-checks` released after `python-tests` reached terminal, then `verify-pr` ran | +| Report posts to the PR **and** Jira | ✅ sticky report on PR #300 + comment on Jira TC-6254, every run | +| Runs with **CI passing** | ✅ `python-tests` **34990552632 = success** → verify-pr **34990553318** posted report for commit `3271577` (CI Status PASS, Overall **WARN**) | +| Runs with **CI failing** | ✅ `python-tests` **34991800394 = failure** (all 4 matrix jobs, intentional `assert False`) → `wait-for-checks` still released → verify-pr **34991800865** posted report for commit `64b818c` (CI Status **FAIL**, Overall **FAIL**); CI-failure sub-task **TC-6255** auto-created | +| No duplicate reports on re-run of the same head | ✅ re-ran verify-pr **34991800865** on the **same** head `64b818c` (completed/success): PR #300 stayed at exactly **2** report comments (`3271577`, `64b818c` — the second edited in place, not re-posted), Jira TC-6254 stayed at exactly **1** comment (updated in place, both reports appended) | + +The **CI-fail** row is the crux: CI result is *data, not a gate*. The +`allowed-conclusions` on `wait-for-checks` includes `failure`, so a failing +`python-tests` still releases the wait and verify-pr runs and reports — it simply +records `CI Status = FAIL` in the report. Idempotency holds independently on each +surface: the PR uses one comment per distinct commit SHA (marker +``); Jira uses a single +comment updated in place per task. + +**DENIED endpoints — expected, none added.** As with TC-5815, the sandbox's +only egress is Vertex; the OCSF log's DENIED `github.com` / `raw.githubusercontent.com` +entries are the split-trust policy **working** (prefetch on the runner, writes on +the runner via the post-script), not failures — no endpoint was added to pass. + +The vehicle is disposable: PR #300, branch `tc-6192-e2e-acceptance`, test issues +TC-6254/TC-6255, and the scaffolding files (`docs/e2e/tc-6192-vehicle.md`, +`plugins/sdlc-workflow/scripts/test_tc6192_ci_fail.py`) are removed after +acceptance is recorded. + +### CI Status sourced from prefetched check-runs (TC-6257) + +TC-6192 above recorded a `CI Status` verdict, but the sandbox had **no** way to +retrieve it: with no `gh` CLI and no egress, the check ran off the diff and warned +*"CI status could not be retrieved in sandbox mode."* TC-6257 closes that gap — the +pre_script now prefetches the head-SHA check-run outcomes (`github.check_runs`) on +the trusted runner, and the concatenated `--log-failed` output of every failed +check-run into `check-run-logs.txt` (mounted separately; only its path rides in +the bundle, so the log text stays off the agent's context until Correctness +Check 1b reads it on a FAIL). This supersedes the TC-6192 "unavailable in sandbox" +caveat: `CI Status` is now sourced from real CI data. + +**Re-pin is load-bearing.** The children delivered to the sandbox are fetched from +the URL-pinned harness base. Advancing the prefetch code alone is inert until the +base pin moves to a commit that contains it: the base-file bytes are unchanged +(sha256 `3c9dc221…`), so only the pinned commit moves (`a9099de0` → `fa3f4b7d`) and +`fullsend lock` refreezes all 11 children at `fa3f4b7d`. Before the re-pin the +sandbox kept running the pre-TC-6257 skill, whose bundle carries no `check_runs`, +and `CI Status` fell back to the old WARN. + +**Vehicle** — throwaway PR **#303** (`tc-6257-e2e-vehicle` → `verify-pr-fullsend`, +**upstream** so `pull_request` gets secrets + OIDC — pre-TC-6331, now lifted), +qualified against Jira task **TC-6266** (status `Review`, label +`ai-generated-jira`, Git PR field = PR #303). + +**Results — both prefetch surfaces proven live:** + +| Acceptance criterion | Result | +|---|---| +| `CI Status` sourced from prefetched `check_runs` (not the WARN fallback) | ✅ verify-pr **35075966985** (commit `1185ae2`, post-re-pin) named the failed check by run ID + conclusion — data the pre-TC-6257 sandbox could not obtain | +| **CI failing** → prefetched log read by Check 1b | ✅ vehicle-only red commit `8865e11`: `Script Unit Tests` **35077403829 = failure** (deliberate canary) → `wait-for-checks` released → verify-pr **35077404421** reported `CI Status = FAIL` and Check 1b surfaced the **runtime-computed canary `cc08a78a24bc66f6`** — a value present **only** in the prefetched `--log-failed` output, never as a literal in the diff — proving the log was read, not inferred | +| CI-failure sub-task auto-created | ✅ **TC-6271** ("Fix failing CI check Script Unit Tests … intentional canary") created, blocks TC-6266, with the `details_url` for the full log | + +The canary is the crux: because its value is computed at test runtime and never +written in the source, the only way the report could quote `cc08a78a24bc66f6` is by +reading the prefetched `check-run-logs.txt` — end-to-end proof of the failed-CI log +prefetch path. (The live run also surfaced two real hardening follow-ups on the +prefetch code — treat `startup_failure`/`stale` as failed conclusions, and only +expose `check_run_logs_path` when the log fetch actually succeeds.) + +The vehicle is disposable: PR #303, branch `tc-6257-e2e-vehicle`, test issue +TC-6266 (and its sub-tasks), and the scaffolding file +(`plugins/sdlc-workflow/scripts/test_tc6257_ci_fail.py`, vehicle-only — never on +the TC-6257 deliverable branch) are removed after acceptance is recorded. + +## Known issues + +- **fullsend deletes the target repo directory** after each run — always pass a + disposable clone as `--target-repo`, never your working directory. +- **Markdown tables collapse to one line in Jira comments** posted via + `fullsend issues n --tracker jira` — the vendor's markdown→ADF conversion + doesn't support GFM tables (parser has no table extension; no table node + handler). GitHub renders the same body correctly. Tracked upstream: + . diff --git a/harness/triage-security.yaml b/harness/triage-security.yaml new file mode 100644 index 000000000..de283a2d5 --- /dev/null +++ b/harness/triage-security.yaml @@ -0,0 +1,69 @@ +# triage-security — standalone root-level Fullsend harness. +# +# The trusted runner gathers every Jira, CVE, lifecycle, and source-repository +# input before sandbox creation. The sandbox receives only the mounted evidence +# bundle and returns a validated action plan; TC-6209 is the only component that +# may execute that plan against Jira. +# +# TC-6208 creates pre-triage-security.sh and its bundle producer. TC-6209 creates +# post-triage-security.sh and its action executor. This contract task declares +# their integration points but deliberately does not implement either script. + +role: triage +agent: plugins/sdlc-workflow/agents/triage-security.md +model: claude-opus-4-8 +effort: high +image: ghcr.io/fullsend-ai/fullsend-code@sha256:9743bc7b6e451e0bcea25ae4a67e0c040c296f1fee04c08988ae80c53fafcfe6 +readonly_repo: true + +plugins: + - plugins/sdlc-workflow + +policy: plugins/sdlc-workflow/policies/triage-security.yaml + +providers: + - plugins/sdlc-workflow/providers/vertex-ai.yaml + +openshell: + profiles: + - plugins/sdlc-workflow/profiles/fullsend-vertex-ai.yaml + +# The pre-script writes the complete non-secret bundle before Fullsend creates +# the sandbox. The Vertex environment mount configures the sole sandbox +# credential used automatically for Vertex AI inference; it is not triage input. +host_files: + - src: plugins/sdlc-workflow/env/gcp-vertex.env + dest: /sandbox/workspace/.env.d/gcp-vertex.env + expand: true + - src: ${FULLSEND_RUN_DIR}/pre/triage-security-input.json + dest: /sandbox/workspace/.pre-script/triage-security-input.json + optional: true + +pre_script: plugins/sdlc-workflow/scripts/pre-triage-security.sh +post_script: plugins/sdlc-workflow/scripts/post-triage-security.sh + +validation_loop: + # Fail closed: validate the original, unmutated agent output against this + # schema before any post-processing. The validator must not strip or normalize + # properties; an unknown or invalid property rejects the iteration before the + # trusted post-script can consume an action plan. + script: plugins/sdlc-workflow/scripts/validate-output-schema.sh + schema: plugins/sdlc-workflow/schemas/triage-security-result.schema.json + max_iterations: 2 + +env: + runner: + PYTHONDONTWRITEBYTECODE: "1" + JIRA_SERVER_URL: "${JIRA_BASE_URL}" + JIRA_EMAIL: "${JIRA_USER_EMAIL}" + JIRA_API_TOKEN: "${JIRA_TOKEN}" + JIRA_PROJECT_KEY: "TC" + FULLSEND_INPUT_SCHEMA: "${TARGET_REPO_DIR}/plugins/sdlc-workflow/schemas/triage-security-input.schema.json" + FULLSEND_OUTPUT_SCHEMA: "${TARGET_REPO_DIR}/plugins/sdlc-workflow/schemas/triage-security-result.schema.json" + FULLSEND_OUTPUT_FILE: "agent-result.json" + sandbox: + # Display-only context; the sandbox has no Jira egress or credential. + JIRA_BASE_URL: "${JIRA_BASE_URL}" + +timeout_minutes: 30 +version: 1 diff --git a/harness/verify-pr.yaml b/harness/verify-pr.yaml new file mode 100644 index 000000000..fef04beee --- /dev/null +++ b/harness/verify-pr.yaml @@ -0,0 +1,152 @@ +# verify-pr — standalone root-level fullsend harness (no base composition). +# +# Runs /sdlc-workflow:verify-pr as a fullsend BYOA agent that checks a PR +# against its Jira task's acceptance criteria. Placed at the REPO ROOT so +# fullsend (invoked with --fullsend-dir = repo root) resolves this harness's +# relative children against the repo root: `plugins/sdlc-workflow` is delivered +# in place as a whole plugin (fullsend fabricates the marketplace cache) and its +# sibling `shared/` resources resolve intact. Confirmed against fullsend v0.37.0 +# — see docs/plans/notes/task1-findings.md (TC-5805). +# +# No base composition and no explicit `security:` block: the Go defaults already +# supply security (enabled: true, fail_mode: closed) plus the ADR-0090 sandbox +# hooks. Composing agents/review.yaml would only leak its review-agent semantics. +# +# Split-trust I/O: the Jira and GitHub tokens live ONLY on the runner (pre_script +# prefetch + post_script writes) and are NEVER exposed to the sandbox. The +# sandbox receives read-only context only (JIRA_BASE_URL for display links) and +# reads the Jira task + PR URL from the pre_script's verify-pr-input.json. +# +# Several referenced children are authored by later tasks in this epic and do not +# exist yet — providers/vertex-ai.yaml (TC-5808), policies/verify-pr.yaml +# (TC-5809), profiles/fullsend-vertex-ai.yaml + schemas/verify-pr-result.schema.json +# + scripts/pre-verify-pr.sh (TC-5810), scripts/post-verify-pr.sh (TC-5811). This +# task validates the harness syntactically only; full fullsend resolution is +# deferred to TC-5810/TC-5811. + +role: review + +# Least-privilege GitHub token levels per run-stage (fullsend ADR-0073). The +# review role's token is minted at the named level each stage actually needs, +# instead of a blanket write token for the whole run: +# pre_script read — GitHub prefetch is read-only (gh pr view/diff; reviews + +# comments via gh api). No writes. +# runtime read — the sandbox has NO api.github.com egress and no token at +# all (split-trust I/O), so it never needs write. The +# validation_loop inherits this stage's level. +# post_script write — the ONLY writer: posts PR review comments, inline +# replies, and the sticky report/summary +# (execute-actions.py), which needs PullRequests: write. +# `read` downgrades every `*:write` in the role to read; `write` is the role's +# full current permission set, so post_script's token is unchanged from today. +# ${GH_TOKEN} in env.runner keeps expanding to the stage-minted token (ADR-0073 +# sets the correct-level token before env expansion). Declaring this applies +# least-privilege now AND survives fullsend #6516 flipping the mint default from +# write to read (post_script keeps requesting write explicitly). +privilege_levels: + pre_script: read + runtime: read + post_script: write + +# CEL trigger (ADR 0061). Evaluated by `fullsend dispatch` over the normalized +# event (internal/normevent); root var `event`, must return bool. Fires when a +# PR change_proposal is opened, updated, or reopened. This is the dispatch-path +# equivalent of the CI poller's `on: pull_request` types in +# .github/workflows/fullsend-verify-pr.yml — kept in sync with it. Kinds stay +# within fullsend's TransitionKind vocabulary (no CI-completion kind exists). +trigger: | + event.entity.kind == "change_proposal" && + event.transition.kind in ["synchronized", "opened", "reopened"] + +agent: plugins/sdlc-workflow/agents/verify-pr.md +# Pinned to a concrete generation, not the floating `opus` alias: under fullsend +# CLI v0.43.0 that alias resolves to claude-opus-4-6, which the Red Hat Vertex +# deployment has not enabled (run 34951983374: "model claude-opus-4-6 is not +# available on your vertex deployment"). opus 4.8 is the fleet's target model. +model: claude-opus-4-8 +# Reasoning effort pinned to the fleet's implicit default: Claude Code runs at +# high effort on effort-capable models, and the fullsend pi runtime passes +# --thinking high when effort is unset. Revisit if `model:` moves to a +# generation with a different default effort. +effort: high +image: ghcr.io/fullsend-ai/fullsend-code@sha256:9743bc7b6e451e0bcea25ae4a67e0c040c296f1fee04c08988ae80c53fafcfe6 +readonly_repo: true + +# In-place whole-plugin delivery — resolved at the repo root. +plugins: + - plugins/sdlc-workflow + +policy: plugins/sdlc-workflow/policies/verify-pr.yaml + +# Vertex only — GitHub is tier-1 and needs no provider. +providers: + - plugins/sdlc-workflow/providers/vertex-ai.yaml + +openshell: + profiles: + - plugins/sdlc-workflow/profiles/fullsend-vertex-ai.yaml + +# URL-pinned base: only repo-relative host_files may live here. When this harness +# is consumed via a pinned raw URL (config.yaml / a child's `base:`), fullsend +# fetches these entries from the base URL and rewrites their paths (ADR-0038). +# Runner-local absolute mounts — the GCP credential/OIDC files and the pre_script +# prefetch output — CANNOT be URL-sourced (fullsend v0.37.0 rejects an absolute +# host_files.src that is inherited from a URL-sourced harness). They are declared +# instead in the local composing child (.fullsend/harness/verify-pr.yaml), whose +# host_files are concatenated with these (dedup by dest, child wins). +host_files: + - src: plugins/sdlc-workflow/env/gcp-vertex.env + dest: /sandbox/workspace/.env.d/gcp-vertex.env + expand: true + +pre_script: plugins/sdlc-workflow/scripts/pre-verify-pr.sh +post_script: plugins/sdlc-workflow/scripts/post-verify-pr.sh + +validation_loop: + # Custom validator: the result schema uses additionalProperties:false + # throughout to document the contract, so strip_extra_properties.py drops + # benign agent-added metadata before jsonschema validates (fullsend's + # native schema-only check would reject it). fullsend requires a script for + # validation_loop regardless — schema alone is not sufficient. + script: plugins/sdlc-workflow/scripts/validate-output-schema.sh + schema: plugins/sdlc-workflow/schemas/verify-pr-result.schema.json + max_iterations: 2 + +env: + # Tokens (JIRA_API_TOKEN, GH_TOKEN) live on the runner ONLY — the pre_script + # prefetch and post_script writes need them; they are never placed in + # env.sandbox. + runner: + # Disable Python bytecode writes for every runner-side Python invocation. + # The pre/post scripts execute in place inside fullsend's content-addressed + # cache tree; a stray __pycache__/*.pyc written there mutates the tree and + # trips fullsend's next-run integrity check. Defense in depth alongside + # execute-actions.py's own sys.dont_write_bytecode (TC-6112). + PYTHONDONTWRITEBYTECODE: "1" + # Map fullsend's dispatch-provided Jira host vars (JIRA_BASE_URL / + # JIRA_USER_EMAIL / JIRA_TOKEN — set by reusable-dispatch harness-run) onto + # the canonical names jira-client.py reads. A local run must export the same + # JIRA_BASE_URL / JIRA_USER_EMAIL / JIRA_TOKEN host vars. JIRA_ISSUE_ID is no + # longer an input: the pre_script derives the Jira key from the triggering PR + # URL (FULLSEND_WORK_ITEM_URL, read ambiently) by JQL on customfield_10875. + JIRA_SERVER_URL: "${JIRA_BASE_URL}" + JIRA_EMAIL: "${JIRA_USER_EMAIL}" + JIRA_API_TOKEN: "${JIRA_TOKEN}" + # Project key for post_script root-cause task creation. The CI dispatch path + # exports no project-key host var, so pin this repo's constant (CLAUDE.md + # Jira Configuration: Project key TC) as a literal rather than a ${ref} that + # would fail `fullsend run` env-validation on CI. + JIRA_PROJECT_KEY: "TC" + GH_TOKEN: "${GH_TOKEN}" + FULLSEND_OUTPUT_SCHEMA: "${FULLSEND_DIR}/plugins/sdlc-workflow/schemas/verify-pr-result.schema.json" + FULLSEND_OUTPUT_FILE: "agent-result.json" + # Read-only context only — no tokens, and no atlassian/github egress from the + # sandbox (network policy enforced in TC-5809). JIRA_BASE_URL is the base for + # display links only (no token), sourced from the dispatch-provided host var. + # The Jira key and PR URL reach the sandbox via verify-pr-input.json (the + # pre_script writes the JQL-derived key as task_id), so no JIRA_ISSUE_ID env + # var is needed here. + sandbox: + JIRA_BASE_URL: "${JIRA_BASE_URL}" + +timeout_minutes: 30 diff --git a/plugins/sdlc-workflow/agents/triage-security.md b/plugins/sdlc-workflow/agents/triage-security.md new file mode 100644 index 000000000..c517883ae --- /dev/null +++ b/plugins/sdlc-workflow/agents/triage-security.md @@ -0,0 +1,46 @@ +--- +name: triage-security +description: >- + Run triage-security from trusted runner evidence and emit a validated + report or Jira action plan without direct external access. +model: opus +--- + +# Triage Security Agent + +You run inside an OpenShell sandbox with read-only repository delivery. The +trusted runner has already prefetched the evidence needed by +`/sdlc-workflow:triage-security`; Jira credentials, external network access, and +all Jira mutations remain outside this sandbox. The Vertex AI credential is the +sole credential in the sandbox and is used automatically for model inference. +You must not inspect, copy, modify, or use it for any other purpose. + +## Startup procedure + +1. Read `/sandbox/workspace/.pre-script/triage-security-input.json`. Fail fast + if the file is missing or if its `issue.key` is absent. +2. Invoke the triage-security skill with that key: + ``` + /sdlc-workflow:triage-security + ``` +3. Use only the mounted evidence bundle for Jira, CVE, lifecycle, and + source-repository facts. Do not call a network API, use `gh`, use `curl`, or + inspect the Vertex AI credential. +4. Write `agent-result.json` to `$FULLSEND_OUTPUT_DIR` using + `triage-security-result.schema.json`. Use `report-only` unless runner + authorization in `authorization.mutation_authorized` is `true`. +5. Use stable `triage-security:` markers. Every placeholder such as + `{{remediation-1.key}}` must be introduced by a remediation-task or + resolve-reference action in the same plan. + +## Constraints + +- Do not modify repository files, push branches, or create pull requests. +- Do not call Jira directly or post comments. The trusted runner validates the + result and performs all authorized Jira actions after the sandbox exits. +- Do not emit an unknown action type or an action with a missing required field. +- Treat every placeholder as unresolved until the trusted TC-6209 executor + resolves it from its action registry; never emit a placeholder without a + matching remediation-task or resolve-reference action. +- Stop and emit a report-only blocked result if supplied evidence is incomplete + or inconsistent; never infer missing security evidence. diff --git a/plugins/sdlc-workflow/agents/verify-pr.md b/plugins/sdlc-workflow/agents/verify-pr.md new file mode 100644 index 000000000..bfcb47570 --- /dev/null +++ b/plugins/sdlc-workflow/agents/verify-pr.md @@ -0,0 +1,55 @@ +--- +name: verify-pr +description: >- + Verify a PR against its Jira task acceptance criteria using the + sdlc-workflow verify-pr skill inside an OpenShell sandbox. +model: opus +--- + +# Verify PR Agent + +You are a PR verification agent running inside an OpenShell sandbox. The +sdlc-workflow plugin is delivered in place, so the `/sdlc-workflow:verify-pr` +skill and its shared resources are available to you directly. + +The sandbox is read-only and has no Jira or GitHub write access: your Jira +context is pre-fetched onto disk and any Jira/GitHub side effects are performed +by the runner after you finish. Produce the structured output and nothing else. + +## Startup procedure + +1. Read the Jira issue ID from the pre-fetched input bundle. The `pre_script` + resolves the key from the triggering PR URL and writes it as `task_id` in + `verify-pr-input.json`, mounted read-only into the sandbox. `JIRA_ISSUE_ID` is + **not** provided to the sandbox — do not read it. Fail fast if the file is + missing or `task_id` is absent/empty (the skill cannot run without it): + ```bash + task_id=$(python3 -c 'import json,sys; print(json.load(open("/sandbox/workspace/.pre-script/verify-pr-input.json"))["task_id"])') + if [ -z "$task_id" ]; then + echo "ERROR: task_id missing from verify-pr-input.json" >&2 + exit 1 + fi + echo "$task_id" + ``` + +2. Invoke the verify-pr skill with that issue ID. The skill is available as + `/sdlc-workflow:verify-pr`. Example (substitute the resolved `$task_id`): + ``` + /sdlc-workflow:verify-pr + ``` + +3. The skill handles everything: reading the pre-fetched Jira task, identifying + the PR, dispatching sub-agents for analysis, and producing the output. + +4. After the skill completes, verify the output file exists: + ```bash + ls -la $FULLSEND_OUTPUT_DIR/agent-result.json + ``` + +## Constraints + +- Do not modify code. This agent only verifies. +- Do not push branches or create PRs. +- Do not call Jira write APIs directly — the skill writes structured JSON output. +- Do not post GitHub comments directly — the post_script handles this. +- Follow the skill's output — do not improvise verification steps. diff --git a/plugins/sdlc-workflow/env/gcp-vertex.env b/plugins/sdlc-workflow/env/gcp-vertex.env new file mode 100644 index 000000000..6eedfa648 --- /dev/null +++ b/plugins/sdlc-workflow/env/gcp-vertex.env @@ -0,0 +1,5 @@ +export CLAUDE_CODE_USE_VERTEX=1 +export ANTHROPIC_VERTEX_PROJECT_ID=${ANTHROPIC_VERTEX_PROJECT_ID} +export CLOUD_ML_REGION=${CLOUD_ML_REGION} +export GOOGLE_APPLICATION_CREDENTIALS=/tmp/.gcp-credentials.json +export GOOGLE_CLOUD_PROJECT=${GOOGLE_CLOUD_PROJECT} diff --git a/plugins/sdlc-workflow/plugin.json b/plugins/sdlc-workflow/plugin.json new file mode 100644 index 000000000..8292622b4 --- /dev/null +++ b/plugins/sdlc-workflow/plugin.json @@ -0,0 +1,3 @@ +{ + "_comment": "TEMPORARY workaround for fullsend bug — tracked upstream at https://github.com/fullsend-ai/fullsend/issues/7008. fullsend's URL plugin resolver (internal/harness/compose.go fetchBasePlugin/fetchBasePluginDir) requires a plugin.json at the plugin-directory ROOT before it sparse-checks-out the plugin over a pinned raw URL. fullsend only checks that this file EXISTS — it never parses it (and its runtime derives the plugin name from the directory basename, not this file). A root-level plugin.json is INVALID per the Claude Code plugin spec, which mandates .claude-plugin/plugin.json; that canonical manifest (name/description/version/author) is what Claude Code, local-path resolution, `claude plugin validate`, and the marketplace read — edit THAT one, not this. This marker intentionally carries no metadata, so there is nothing to keep in sync and no drift. A symlink is not an option: fullsend's gitfetch rejects symlinks. Remove this file once #7008 ships a fix that accepts .claude-plugin/plugin.json." +} diff --git a/plugins/sdlc-workflow/policies/triage-security.yaml b/plugins/sdlc-workflow/policies/triage-security.yaml new file mode 100644 index 000000000..729594c7f --- /dev/null +++ b/plugins/sdlc-workflow/policies/triage-security.yaml @@ -0,0 +1,52 @@ +# Sandbox policy for the triage-security Fullsend harness. +# +# Triage analysis consumes a trusted evidence bundle. Jira, GitHub, CVE +# databases, lifecycle pages, and source repositories remain runner-side, where +# credentials and network access are available. Only inference/telemetry egress +# is permitted from the sandbox. + +version: 1 + +filesystem_policy: + include_workdir: false + read_only: [/var/log, /usr, /lib, /lib64, /proc, /dev/urandom, /etc, /opt] + read_write: [/sandbox, /tmp, /dev/null] +landlock: + compatibility: best_effort +process: + run_as_user: sandbox + run_as_group: sandbox + +network_policies: + claude_code: + name: claude-code + endpoints: + - host: "api.anthropic.com" + port: 443 + protocol: rest + enforcement: enforce + access: read-write + - host: "*.googleapis.com" + port: 443 + protocol: rest + enforcement: enforce + access: read-write + - host: "platform.claude.com" + port: 443 + protocol: rest + enforcement: enforce + access: read-write + - host: "statsig.anthropic.com" + port: 443 + protocol: rest + enforcement: enforce + access: read-write + - host: "sentry.io" + port: 443 + protocol: rest + enforcement: enforce + access: read-write + binaries: + - path: "**/claude" + - path: "**/pi" + - path: "**/node" diff --git a/plugins/sdlc-workflow/policies/verify-pr.yaml b/plugins/sdlc-workflow/policies/verify-pr.yaml new file mode 100644 index 000000000..8d4108c94 --- /dev/null +++ b/plugins/sdlc-workflow/policies/verify-pr.yaml @@ -0,0 +1,72 @@ +version: 1 + +# Sandbox policy for the verify-pr fullsend harness (read-only agent). +# +# The agent checks a PR against its Jira task's acceptance criteria. It needs +# ONLY inference egress: Anthropic hosts + Vertex AI (`*.googleapis.com`). +# Format confirmed against fullsend v0.37.0 upstream policies +# (/Users/mrizzi/git/cloned/agents/policies/*): per-entry `name:` is still +# required, hosts support `*` wildcards, and binaries use `**/…` globs. +# +# Split-trust I/O — deliberately NO tier-1 egress from the sandbox: +# * NO `*.atlassian.net` (Jira): the Jira token lives only on the runner; +# issue context is prefetched host-side (pre_script, TC-5810) and mounted +# read-only, and all Jira writes run host-side (post_script, TC-5811). +# * NO `api.github.com`/`github.com` (GitHub): the GitHub token lives only on +# the runner; PR context is prefetched host-side and writes run host-side. +# A `gh` read reaching api.github.com from the sandbox is a leak to fix in +# the skill/prefetch — NOT a reason to add a github egress rule here. +# +# curl and gh are excluded from the binary allowlist to prevent raw HTTP +# egress with any injected token; only the fullsend runtimes (claude/pi/node) +# may open the network. `**/pi` is the runtime that brokers Vertex AI +# (`*.googleapis.com`) inference, so it must be present alongside claude/node. +# +# include_workdir: false — the fullsend dir is not mounted into the sandbox; +# the target repo is delivered read-only under /sandbox (readonly_repo: true). + +filesystem_policy: + include_workdir: false + read_only: [/var/log, /usr, /lib, /lib64, /proc, /dev/urandom, /etc, /opt] + read_write: [/sandbox, /tmp, /dev/null] +landlock: + compatibility: best_effort +process: + run_as_user: sandbox + run_as_group: sandbox + +network_policies: + claude_code: + name: claude-code + endpoints: + # Anthropic inference + Vertex AI (googleapis) — the only egress allowed. + - host: "api.anthropic.com" + port: 443 + protocol: rest + enforcement: enforce + access: read-write + - host: "*.googleapis.com" + port: 443 + protocol: rest + enforcement: enforce + access: read-write + - host: "platform.claude.com" + port: 443 + protocol: rest + enforcement: enforce + access: read-write + # Telemetry / error reporting (POST) — kept minimal. + - host: "statsig.anthropic.com" + port: 443 + protocol: rest + enforcement: enforce + access: read-write + - host: "sentry.io" + port: 443 + protocol: rest + enforcement: enforce + access: read-write + binaries: + - path: "**/claude" + - path: "**/pi" + - path: "**/node" diff --git a/plugins/sdlc-workflow/profiles/fullsend-vertex-ai.yaml b/plugins/sdlc-workflow/profiles/fullsend-vertex-ai.yaml new file mode 100644 index 000000000..153bdf14d --- /dev/null +++ b/plugins/sdlc-workflow/profiles/fullsend-vertex-ai.yaml @@ -0,0 +1,15 @@ +--- +id: fullsend-vertex-ai +display_name: Fullsend Vertex AI +description: Google Cloud APIs for Vertex AI inference +category: inference +endpoints: + - host: "*.googleapis.com" + port: 443 + protocol: rest + access: read-write + enforcement: enforce +binaries: + - "**/claude" + - "**/pi" + - "**/node" diff --git a/plugins/sdlc-workflow/providers/vertex-ai.yaml b/plugins/sdlc-workflow/providers/vertex-ai.yaml new file mode 100644 index 000000000..50ba5f207 --- /dev/null +++ b/plugins/sdlc-workflow/providers/vertex-ai.yaml @@ -0,0 +1,5 @@ +--- +name: vertex-ai +type: fullsend-vertex-ai +credentials: + _NOOP_VERTEX_AI: "" diff --git a/plugins/sdlc-workflow/schemas/triage-security-input.schema.json b/plugins/sdlc-workflow/schemas/triage-security-input.schema.json new file mode 100644 index 000000000..a51c581bf --- /dev/null +++ b/plugins/sdlc-workflow/schemas/triage-security-input.schema.json @@ -0,0 +1,349 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "triage-security-input.schema.json", + "title": "Triage Security Trusted Prefetch Input", + "description": "Complete trusted-runner evidence bundle for the tokenless triage-security sandbox.", + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "issue", + "remote_links", + "configuration", + "external_evidence", + "matrix", + "source_evidence", + "jira_metadata", + "idempotency", + "authorization" + ], + "properties": { + "schema_version": { + "type": "string", + "const": "1" + }, + "issue": { + "$ref": "#/$defs/issue" + }, + "remote_links": { + "type": "array", + "minItems": 1, + "items": { + "$ref": "#/$defs/remote_link" + } + }, + "configuration": { + "$ref": "#/$defs/configuration" + }, + "external_evidence": { + "$ref": "#/$defs/external_evidence" + }, + "matrix": { + "$ref": "#/$defs/matrix" + }, + "source_evidence": { + "$ref": "#/$defs/source_evidence" + }, + "jira_metadata": { + "$ref": "#/$defs/jira_metadata" + }, + "idempotency": { + "$ref": "#/$defs/idempotency" + }, + "authorization": { + "$ref": "#/$defs/authorization" + } + }, + "$defs": { + "issue_key": { + "type": "string", + "pattern": "^[A-Z][A-Z0-9]+-[0-9]+$" + }, + "url": { + "type": "string", + "format": "uri" + }, + "issue": { + "type": "object", + "additionalProperties": false, + "required": [ + "key", + "summary", + "description", + "status", + "labels", + "versions", + "reporter", + "comments", + "fields" + ], + "properties": { + "key": { "$ref": "#/$defs/issue_key" }, + "summary": { "type": "string", "minLength": 1 }, + "description": { "type": "object" }, + "status": { "type": "string", "minLength": 1 }, + "labels": { + "type": "array", + "items": { "type": "string" } + }, + "versions": { + "type": "array", + "items": { "$ref": "#/$defs/jira_version" } + }, + "reporter": { "$ref": "#/$defs/person" }, + "comments": { + "type": "array", + "items": { "type": "object" } + }, + "fields": { + "type": "object", + "description": "Raw issue fields needed to audit extracted triage values." + } + } + }, + "person": { + "type": "object", + "additionalProperties": false, + "required": ["account_id", "display_name"], + "properties": { + "account_id": { "type": "string", "minLength": 1 }, + "display_name": { "type": "string", "minLength": 1 } + } + }, + "jira_version": { + "type": "object", + "additionalProperties": false, + "required": ["id", "name", "released"], + "properties": { + "id": { "type": "string", "minLength": 1 }, + "name": { "type": "string", "minLength": 1 }, + "released": { "type": "boolean" }, + "archived": { "type": "boolean" } + } + }, + "remote_link": { + "type": "object", + "additionalProperties": false, + "required": ["url", "title"], + "properties": { + "url": { "$ref": "#/$defs/url" }, + "title": { "type": "string", "minLength": 1 } + } + }, + "configuration": { + "type": "object", + "additionalProperties": false, + "required": [ + "project_key", + "jira_version_prefix", + "vulnerability_issue_type_id", + "component_label_pattern", + "version_streams", + "source_repositories" + ], + "properties": { + "project_key": { "type": "string", "minLength": 1 }, + "jira_version_prefix": { "type": "string", "minLength": 1 }, + "vulnerability_issue_type_id": { "type": "string", "minLength": 1 }, + "component_label_pattern": { "type": "string", "minLength": 1 }, + "vex_justification_field": { "type": "string" }, + "upstream_affected_component_field": { "type": "string" }, + "ps_component_field": { "type": "string" }, + "stream_field": { "type": "string" }, + "prodsec_account_id": { "type": "string" }, + "embargo_policy_url": { "$ref": "#/$defs/url" }, + "version_streams": { + "type": "array", + "minItems": 1, + "items": { "$ref": "#/$defs/stream_config" } + }, + "source_repositories": { + "type": "array", + "minItems": 1, + "items": { "$ref": "#/$defs/source_repository" } + } + } + }, + "stream_config": { + "type": "object", + "additionalProperties": false, + "required": ["name", "matrix_path", "release_repository"], + "properties": { + "name": { "type": "string", "minLength": 1 }, + "matrix_path": { "type": "string", "minLength": 1 }, + "release_repository": { "type": "string", "minLength": 1 } + } + }, + "source_repository": { + "type": "object", + "additionalProperties": false, + "required": ["name", "url", "deployment_context"], + "properties": { + "name": { "type": "string", "minLength": 1 }, + "url": { "$ref": "#/$defs/url" }, + "deployment_context": { + "type": "string", + "enum": ["internal", "upstream", "customer-shipped"] + } + } + }, + "external_evidence": { + "type": "object", + "additionalProperties": false, + "required": ["mitre", "osv", "lifecycle"], + "properties": { + "mitre": { "$ref": "#/$defs/retrieved_evidence" }, + "osv": { "$ref": "#/$defs/retrieved_evidence" }, + "lifecycle": { "$ref": "#/$defs/retrieved_evidence" } + } + }, + "retrieved_evidence": { + "type": "object", + "additionalProperties": false, + "required": ["source_url", "retrieved_at", "status", "body"], + "properties": { + "source_url": { "$ref": "#/$defs/url" }, + "retrieved_at": { "type": "string", "format": "date-time" }, + "status": { "type": "integer", "minimum": 100, "maximum": 599 }, + "body": { "type": ["object", "array", "string"] } + } + }, + "matrix": { + "type": "object", + "additionalProperties": false, + "required": ["streams"], + "properties": { + "streams": { + "type": "array", + "minItems": 1, + "items": { "$ref": "#/$defs/matrix_stream" } + } + } + }, + "matrix_stream": { + "type": "object", + "additionalProperties": false, + "required": ["name", "matrix_source", "rows"], + "properties": { + "name": { "type": "string", "minLength": 1 }, + "matrix_source": { "type": "string", "minLength": 1 }, + "rows": { + "type": "array", + "minItems": 1, + "items": { "$ref": "#/$defs/matrix_row" } + } + } + }, + "matrix_row": { + "type": "object", + "additionalProperties": false, + "required": ["version", "source_commits", "retag_of"], + "properties": { + "version": { "type": "string", "minLength": 1 }, + "source_commits": { + "type": "object", + "additionalProperties": { "type": "string", "minLength": 7 } + }, + "retag_of": { "type": ["string", "null"] } + } + }, + "source_evidence": { + "type": "object", + "additionalProperties": false, + "required": ["lock_files", "development_streams"], + "properties": { + "lock_files": { + "type": "array", + "minItems": 1, + "items": { "$ref": "#/$defs/source_read" } + }, + "development_streams": { + "type": "array", + "minItems": 1, + "items": { "$ref": "#/$defs/source_read" } + } + } + }, + "source_read": { + "type": "object", + "additionalProperties": false, + "required": ["repository", "ref", "path", "command", "content"], + "properties": { + "repository": { "type": "string", "minLength": 1 }, + "ref": { "type": "string", "minLength": 1 }, + "path": { "type": "string", "minLength": 1 }, + "command": { "type": "string", "pattern": "^git show " }, + "content": { "type": "string" } + } + }, + "jira_metadata": { + "type": "object", + "additionalProperties": false, + "required": ["versions", "sibling_searches", "related_issues"], + "properties": { + "versions": { + "type": "array", + "items": { "$ref": "#/$defs/jira_version" } + }, + "sibling_searches": { + "type": "array", + "items": { "$ref": "#/$defs/jira_search" } + }, + "related_issues": { + "type": "array", + "items": { "$ref": "#/$defs/related_issue" } + } + } + }, + "jira_search": { + "type": "object", + "additionalProperties": false, + "required": ["purpose", "jql", "issues"], + "properties": { + "purpose": { "type": "string", "minLength": 1 }, + "jql": { "type": "string", "minLength": 1 }, + "issues": { + "type": "array", + "items": { "$ref": "#/$defs/related_issue" } + } + } + }, + "related_issue": { + "type": "object", + "additionalProperties": false, + "required": ["key", "summary", "status", "labels", "description", "comments", "links"], + "properties": { + "key": { "$ref": "#/$defs/issue_key" }, + "summary": { "type": "string" }, + "status": { "type": "string" }, + "labels": { "type": "array", "items": { "type": "string" } }, + "description": { "type": "object" }, + "comments": { "type": "array", "items": { "type": "object" } }, + "links": { "type": "array", "items": { "type": "object" } } + } + }, + "idempotency": { + "type": "object", + "additionalProperties": false, + "required": ["action_markers", "existing_remediation"], + "properties": { + "action_markers": { "type": "array", "items": { "type": "string" } }, + "existing_remediation": { + "type": "array", + "items": { "$ref": "#/$defs/related_issue" } + } + } + }, + "authorization": { + "type": "object", + "additionalProperties": false, + "required": ["mutation_authorized"], + "properties": { + "mutation_authorized": { + "type": "boolean", + "description": "Trusted runner authorization. False requires report-only sandbox output." + } + } + } + } +} diff --git a/plugins/sdlc-workflow/schemas/triage-security-result.schema.json b/plugins/sdlc-workflow/schemas/triage-security-result.schema.json new file mode 100644 index 000000000..3351ce035 --- /dev/null +++ b/plugins/sdlc-workflow/schemas/triage-security-result.schema.json @@ -0,0 +1,231 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "triage-security-result.schema.json", + "title": "Triage Security Agent Result", + "description": "Fail-closed sandbox output. Before any Jira mutation, TC-6209 independently reads trusted input authorization, rejects every mutating action when it is false, and resolves every action reference.", + "type": "object", + "additionalProperties": false, + "required": ["schema_version", "mode", "report", "actions"], + "properties": { + "schema_version": { + "type": "string", + "const": "1" + }, + "mode": { + "type": "string", + "enum": ["report-only", "mutation-authorized"] + }, + "report": { + "$ref": "#/$defs/report" + }, + "actions": { + "type": "array", + "minItems": 1, + "items": { "$ref": "#/$defs/action" } + } + }, + "allOf": [ + { + "if": { + "properties": { "mode": { "const": "report-only" } } + }, + "then": { + "properties": { + "actions": { + "items": { + "type": "object", + "properties": { "type": { "const": "report-only" } } + } + } + } + } + } + ], + "$defs": { + "issue_key": { + "type": "string", + "pattern": "^[A-Z][A-Z0-9]+-[0-9]+$" + }, + "reference_name": { + "type": "string", + "pattern": "^[a-z][a-z0-9-]*$" + }, + "issue_reference": { + "type": "string", + "pattern": "^(?:[A-Z][A-Z0-9]+-[0-9]+|\\{\\{[a-z][a-z0-9-]*\\.key\\}\\})$", + "description": "TC-6209 resolves placeholders against the action registry and rejects unknown or unresolved references before any Jira call." + }, + "stable_marker": { + "type": "string", + "pattern": "^triage-security:[a-z0-9][a-z0-9._:-]*$" + }, + "adf_document": { + "type": "object", + "additionalProperties": false, + "required": ["type", "version", "content"], + "properties": { + "type": { "const": "doc" }, + "version": { "const": 1 }, + "content": { + "type": "array", + "minItems": 1, + "items": { + "type": "object", + "required": ["type"], + "properties": { + "type": { "type": "string", "minLength": 1 } + } + } + } + } + }, + "report": { + "type": "object", + "additionalProperties": false, + "required": ["issue", "outcome", "summary_markdown", "evidence"], + "properties": { + "issue": { "$ref": "#/$defs/issue_key" }, + "outcome": { + "type": "string", + "enum": ["affected", "not-affected", "duplicate", "needs-review", "blocked"] + }, + "summary_markdown": { "type": "string", "minLength": 1 }, + "evidence": { + "type": "array", + "minItems": 1, + "items": { "$ref": "#/$defs/evidence" } + } + } + }, + "evidence": { + "type": "object", + "additionalProperties": false, + "required": ["source", "detail"], + "properties": { + "source": { "type": "string", "minLength": 1 }, + "detail": { "type": "string", "minLength": 1 }, + "url": { "type": "string", "format": "uri" } + } + }, + "action": { + "type": "object", + "required": ["type", "marker"], + "properties": { + "type": { + "type": "string", + "enum": [ + "report-only", + "field-edit", + "status-transition", + "comment", + "link", + "remediation-task", + "resolve-reference" + ] + }, + "marker": { "$ref": "#/$defs/stable_marker" } + }, + "allOf": [ + { + "if": { "properties": { "type": { "const": "report-only" } } }, + "then": { + "required": ["type", "marker"], + "properties": { + "type": {}, + "marker": {} + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "field-edit" } } }, + "then": { + "required": ["type", "marker", "issue", "fields"], + "properties": { + "type": {}, + "marker": {}, + "issue": { "$ref": "#/$defs/issue_reference" }, + "fields": { + "type": "object", + "minProperties": 1, + "additionalProperties": true + } + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "status-transition" } } }, + "then": { + "required": ["type", "marker", "issue", "status"], + "properties": { + "type": {}, + "marker": {}, + "issue": { "$ref": "#/$defs/issue_reference" }, + "status": { "type": "string", "minLength": 1 } + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "comment" } } }, + "then": { + "required": ["type", "marker", "issue", "body_adf"], + "properties": { + "type": {}, + "marker": {}, + "issue": { "$ref": "#/$defs/issue_reference" }, + "body_adf": { "$ref": "#/$defs/adf_document" } + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "link" } } }, + "then": { + "required": ["type", "marker", "link_type", "inward", "outward"], + "properties": { + "type": {}, + "marker": {}, + "link_type": { "type": "string", "enum": ["Blocks", "Depend", "Related"] }, + "inward": { "$ref": "#/$defs/issue_reference" }, + "outward": { "$ref": "#/$defs/issue_reference" } + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "remediation-task" } } }, + "then": { + "required": ["type", "marker", "ref", "project", "summary", "description_adf", "labels"], + "properties": { + "type": {}, + "marker": {}, + "ref": { "$ref": "#/$defs/reference_name" }, + "project": { "type": "string", "minLength": 1 }, + "summary": { "type": "string", "minLength": 1, "maxLength": 255 }, + "description_adf": { "$ref": "#/$defs/adf_document" }, + "labels": { "type": "array", "items": { "type": "string" } }, + "priority": { "type": "string", "minLength": 1 }, + "fix_versions": { "type": "array", "items": { "type": "string" } } + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "resolve-reference" } } }, + "then": { + "required": ["type", "marker", "ref", "issue"], + "properties": { + "type": {}, + "marker": {}, + "ref": { "$ref": "#/$defs/reference_name" }, + "issue": { "$ref": "#/$defs/issue_key" } + }, + "additionalProperties": false + } + } + ] + } + } +} diff --git a/plugins/sdlc-workflow/schemas/verify-pr-input.schema.json b/plugins/sdlc-workflow/schemas/verify-pr-input.schema.json new file mode 100644 index 000000000..014f7056d --- /dev/null +++ b/plugins/sdlc-workflow/schemas/verify-pr-input.schema.json @@ -0,0 +1,144 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "verify-pr-input.schema.json", + "title": "Verify PR Pre-Script Input", + "description": "Pre-fetched task data written by the pre_script and mounted into the sandbox. The skill reads this instead of calling the issue tracker or GitHub APIs directly.", + "type": "object", + "required": ["task_id", "task", "pr_url", "github", "idempotency"], + "additionalProperties": false, + "properties": { + "task_id": { + "type": "string", + "minLength": 1, + "description": "Issue tracker task ID (e.g., TC-4741, #42)" + }, + "task": { + "type": "object", + "required": ["summary", "description", "status", "labels"], + "properties": { + "summary": { "type": "string", "minLength": 1 }, + "description": { + "type": "object", + "description": "Task description in the issue tracker's native format (e.g., ADF for Jira)" + }, + "status": { "type": "string" }, + "labels": { + "type": "array", + "items": { "type": "string" } + }, + "issue_links": { + "type": "array", + "items": { + "type": "object", + "properties": { + "type": { "type": "string" }, + "direction": { "type": "string", "enum": ["inward", "outward"] }, + "key": { "type": "string" } + } + }, + "description": "Related issues (parent, blocks, relates)" + }, + "custom_fields": { + "type": "object", + "description": "Tracker-specific custom field values" + } + } + }, + "pr_url": { + "type": "string", + "description": "PR URL extracted from the task, or empty if not linked" + }, + "github": { + "type": "object", + "description": "GitHub tier-1 read bundle, prefetched on the trusted runner (where GH_TOKEN lives) so the sandbox needs no api.github.com egress. Every read the verify-pr skill performs against the PR is captured here.", + "required": ["pr_repo", "pr_number", "headRefName", "commit_sha", "diff", "stat", "reviews", "review_comments", "issue_comments", "commits", "check_runs", "check_run_logs_path"], + "additionalProperties": false, + "properties": { + "pr_repo": { + "type": "string", + "pattern": "^[a-zA-Z0-9._-]+/[a-zA-Z0-9._-]+$", + "description": "owner/repo parsed from pr_url" + }, + "pr_number": { "type": "integer", "minimum": 1 }, + "headRefName": { + "type": "string", + "description": "PR head branch name (gh pr view --json headRefName)" + }, + "commit_sha": { + "type": "string", + "pattern": "^[0-9a-f]{7,40}$", + "description": "OID of the PR head ref tip commit (gh pr view --json headRefOid --jq .headRefOid)" + }, + "diff": { + "type": "string", + "description": "Unified PR diff (gh pr diff)" + }, + "stat": { + "type": "string", + "description": "PR diffstat (git apply --stat on the fetched pr.diff; gh pr diff has no --stat flag)" + }, + "reviews": { + "type": "array", + "description": "PR reviews (gh api repos//pulls//reviews)" + }, + "review_comments": { + "type": "array", + "description": "PR review (inline code) comments (gh api repos//pulls//comments)" + }, + "issue_comments": { + "type": "array", + "description": "PR conversation comments (gh api repos//issues//comments)" + }, + "commits": { + "type": "array", + "description": "PR commits (gh pr view --json commits)" + }, + "check_runs": { + "type": "array", + "description": "Head-SHA CI check-run outcomes (gh api repos//commits//check-runs), reduced to name/status/conclusion/details_url. Read by correctness.md Check 1 (CI Status) in sandbox mode; empty when the PR head has no checks.", + "items": { "type": "object" } + }, + "check_run_logs_path": { + "type": "string", + "description": "Mounted sandbox path of the concatenated failed-check logs (gh run view --log-failed for each failed check-run), or \"\" when no check failed. correctness.md Check 1b reads this file only on a FAIL, so the (large) log text stays out of the bundle and off the sub-agent's context on the green path. The log content is deliberately NOT inlined here." + } + } + }, + "idempotency": { + "type": "object", + "description": "Related-issue metadata prefetched on the trusted runner so the sandbox can run the Steps 6d/6f/7c idempotency checks (duplicate sub-task and root-cause detection) without a Jira token or egress. Always present in a valid prefetch: the producer runs on the trusted runner with a Jira token, so it can always emit the bundle — with an empty related_issues array when the task has no sub-tasks or linked issues yet. Requiring it (rather than leaving it optional) closes the gap that let a schema-valid prefetch omit the bundle and silently skip dedup in the tokenless sandbox.", + "additionalProperties": false, + "required": ["related_issues"], + "properties": { + "related_issues": { + "type": "array", + "description": "The task's existing sub-tasks and linked issues (e.g., root-cause tasks), each with the fields the dedup checks inspect.", + "items": { + "type": "object", + "required": ["key", "summary", "labels", "description", "issuetype", "comments"], + "properties": { + "key": { "type": "string", "description": "Issue key (e.g., TC-4742)" }, + "summary": { "type": "string" }, + "labels": { "type": "array", "items": { "type": "string" } }, + "description": { "type": "object", "description": "Issue description in the tracker's native format (e.g., ADF)" }, + "issuetype": { "type": "string", "description": "Issue type name (e.g., Sub-task, Task)" }, + "comments": { + "type": "array", + "description": "Comment bodies in the tracker's native format; searched by Step 7c for the root-cause marker.", + "items": { "type": "object" } + } + } + } + } + } + }, + "source": { + "type": "object", + "description": "Raw issue tracker response for fields not captured above", + "properties": { + "tracker": { "type": "string", "description": "Issue tracker type (e.g., jira, github)" }, + "raw": { "type": "object", "description": "Full raw API response" } + } + } + } +} diff --git a/plugins/sdlc-workflow/schemas/verify-pr-result.schema.json b/plugins/sdlc-workflow/schemas/verify-pr-result.schema.json new file mode 100644 index 000000000..3c63813cb --- /dev/null +++ b/plugins/sdlc-workflow/schemas/verify-pr-result.schema.json @@ -0,0 +1,134 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "verify-pr-result.schema.json", + "title": "Verify PR Agent Result", + "description": "Structured output from the verify-pr agent. Validated by fullsend's validation_loop before the post_script runs.", + "type": "object", + "additionalProperties": false, + "required": ["report", "actions"], + "properties": { + "report": { "$ref": "#/$defs/report" }, + "actions": { + "type": "array", + "items": { "$ref": "#/$defs/action" } + } + }, + "$defs": { + "report": { + "type": "object", + "required": ["jira_issue_id", "pr_repo", "pr_number", "commit_sha", "overall", "table_md", "report_md", "report_adf", "plugin_version"], + "additionalProperties": false, + "properties": { + "jira_issue_id": { "type": "string", "pattern": "^[A-Z]+-[0-9]+$" }, + "pr_repo": { "type": "string", "pattern": "^[a-zA-Z0-9._-]+/[a-zA-Z0-9._-]+$" }, + "pr_number": { "type": "integer", "minimum": 1 }, + "commit_sha": { "type": "string", "pattern": "^[0-9a-f]{7,40}$" }, + "overall": { "type": "string", "enum": ["PASS", "WARN", "FAIL"] }, + "table_md": { "type": "string", "minLength": 1 }, + "report_md": { "type": "string", "minLength": 1 }, + "report_adf": { "type": "object" }, + "plugin_version": { "type": "string", "pattern": "^[0-9]+\\.[0-9]+\\.[0-9]+$" } + } + }, + "action": { + "type": "object", + "required": ["type"], + "properties": { + "type": { "type": "string", "enum": ["create_subtask", "create_link", "post_pr_reply", "post_pr_comment", "create_root_cause_task", "post_comment", "post_report"] } + }, + "allOf": [ + { + "if": { "properties": { "type": { "const": "create_subtask" } } }, + "then": { + "required": ["type", "ref", "parent", "summary", "labels", "description_adf"], + "properties": { + "type": {}, + "ref": { "type": "string", "pattern": "^[a-z0-9-]+$" }, + "parent": { "type": "string", "pattern": "^[A-Z]+-[0-9]+$" }, + "summary": { "type": "string", "minLength": 1, "maxLength": 255 }, + "labels": { "type": "array", "items": { "type": "string" } }, + "description_adf": { "type": "object" } + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "create_link" } } }, + "then": { + "required": ["type", "link_type", "inward", "outward"], + "properties": { + "type": {}, + "link_type": { "type": "string", "enum": ["Blocks", "Related"] }, + "inward": { "type": "string", "minLength": 1 }, + "outward": { "type": "string", "minLength": 1 } + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "post_pr_reply" } } }, + "then": { + "required": ["type", "repo", "pr_number", "comment_id", "body"], + "properties": { + "type": {}, + "repo": { "type": "string" }, + "pr_number": { "type": "integer", "minimum": 1 }, + "comment_id": { "type": "integer", "minimum": 1 }, + "body": { "type": "string", "minLength": 1 } + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "post_pr_comment" } } }, + "then": { + "required": ["type", "repo", "pr_number", "body"], + "properties": { + "type": {}, + "repo": { "type": "string" }, + "pr_number": { "type": "integer", "minimum": 1 }, + "body": { "type": "string", "minLength": 1 } + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "create_root_cause_task" } } }, + "then": { + "required": ["type", "ref", "summary", "labels", "description_adf"], + "properties": { + "type": {}, + "ref": { "type": "string", "pattern": "^[a-z0-9-]+$" }, + "summary": { "type": "string", "minLength": 1, "maxLength": 255 }, + "labels": { "type": "array", "items": { "type": "string" } }, + "description_adf": { "type": "object" } + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "post_comment" } } }, + "then": { + "required": ["type", "issue", "body_adf"], + "properties": { + "type": {}, + "issue": { "type": "string", "pattern": "^([A-Z]+-[0-9]+|\\{\\{[a-z0-9-]+\\.key\\}\\})$" }, + "body_adf": { "type": "object" } + }, + "additionalProperties": false + } + }, + { + "if": { "properties": { "type": { "const": "post_report" } } }, + "then": { + "required": ["type"], + "properties": { + "type": {} + }, + "additionalProperties": false + } + } + ] + } + } +} diff --git a/plugins/sdlc-workflow/scripts/execute-actions.py b/plugins/sdlc-workflow/scripts/execute-actions.py new file mode 100755 index 000000000..d963416d9 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/execute-actions.py @@ -0,0 +1,759 @@ +#!/usr/bin/env python3 +"""Execute verify-pr structured output actions. + +Reads agent-result.json, processes actions sequentially, resolves +{{ref.key}} and {{ref.url}} placeholders as entities are created. + +Comments — the verify-pr report on both the GitHub PR and the Jira issue, plus +standalone Jira analysis comments — are posted through the native +`fullsend issues post-comment` sticky-comment CLI (`--tracker github|jira`). +The irreducible Jira writes that have no native primitive yet — sub-tasks, +links, and root-cause tasks — call jira-client.py functions directly (imported +as a module). The remaining GitHub PR operations that have no native fullsend +command — replying to a review-comment thread and posting a per-item PR comment +— still use the gh CLI. Runs on the fullsend runner (trusted side), not inside +the sandbox. + +Not idempotent for entity creation: if an action fails mid-execution, +previously created Jira sub-tasks are not rolled back. Manual cleanup may be +needed after partial failures. Comment posting is idempotent: the Jira sticky +marker updates the existing comment instead of duplicating it, and the GitHub +report comment carries a commit-scoped marker so a retry after a partial +failure updates the same commit's report comment rather than posting a +duplicate. + +Usage: + execute-actions.py + +Required env vars: + JIRA_SERVER_URL, JIRA_EMAIL, JIRA_API_TOKEN — Jira credentials + GH_TOKEN — GitHub token + JIRA_PROJECT_KEY — Jira project key (for root-cause task creation) +""" + +import datetime +import importlib.util +import json +import os +import re +import subprocess +import sys +from typing import Any + +_script_dir = os.path.dirname(os.path.abspath(__file__)) +_spec = importlib.util.spec_from_file_location( + "jira_client", os.path.join(_script_dir, "jira-client.py") +) +_jira_mod = importlib.util.module_from_spec(_spec) +# fullsend executes this script in place inside its content-addressed cache tree. +# SourceFileLoader.exec_module writes __pycache__/*.pyc next to the source, which +# mutates that tree and breaks fullsend's next-run integrity check ("cache +# integrity check failed"). Disable bytecode writes before loading the sibling +# module so nothing is written into the cache tree (TC-6112). The harness also +# sets PYTHONDONTWRITEBYTECODE for defense in depth. +sys.dont_write_bytecode = True +_spec.loader.exec_module(_jira_mod) + +REF_PATTERN = re.compile(r"\{\{([a-z0-9-]+)\.(key|url)\}\}") + +# Stable marker so the native sticky-comment CLI updates the same Jira +# comment across re-runs instead of posting duplicates. Used by post_report +# (the verification report on the main task). +STICKY_COMMENT_MARKER = "" + +# Distinct sticky marker for standalone analysis comments (post_comment). The +# native CLI treats the marker as the sticky-comment identity, so a post_comment +# and a post_report targeting the SAME Jira issue in one run must not share one +# marker — otherwise the second post would overwrite the first. Each path keeps +# its own stable marker, so per-path re-run idempotency is preserved. +# The suffix is hyphenated (not "post_comment"): fullsend rejects a Jira +# --marker containing \*_`[]& because Jira's markdown round-trip escapes those +# characters on read-back, which would break marker re-detection on later runs. +POST_COMMENT_STICKY_MARKER = "" + +# The verify-pr report comment is posted via the native fullsend sticky-comment +# CLI (`fullsend issues post-comment --tracker github`), which prepends this +# marker as a hidden HTML comment and, on re-runs, finds the marked comment and +# edits it in place. A canonical (full-length) commit SHA is appended per post, +# scoping the sticky identity to a single verification run/commit: a retry for +# the same commit updates that comment instead of duplicating it, while a later +# commit gets a fresh comment — preserving the per-run verification history that +# verify-pr SKILL.md Step 9 posts. +GITHUB_REPORT_MARKER_PREFIX = "" + + # The CLI prepends the marker; drop the agent's embedded leading marker line + # (any commit-scoped variant, matched by prefix) so the comment does not open + # with two duplicate marker lines. + if report_md.startswith(GITHUB_REPORT_MARKER_PREFIX): + report_md = report_md.split("\n", 1)[1] if "\n" in report_md else "" + report_md = report_md.lstrip("\n") + + post_github_comment_native(repo, pr_number, report_md, marker) + print(f" Posted report to PR #{pr_number}") + + post_jira_comment_native(jira_issue_id, jira_body_md) + print(f" Posted report to Jira {jira_issue_id}") + + +EXECUTORS = { + "create_subtask": execute_create_subtask, + "create_link": execute_create_link, + "post_pr_reply": execute_post_pr_reply, + "post_pr_comment": execute_post_pr_comment, + "create_root_cause_task": execute_create_root_cause_task, + "post_comment": execute_post_comment, +} + + +def main(): + if len(sys.argv) != 2: + print(f"Usage: {sys.argv[0]} ", file=sys.stderr) + sys.exit(1) + + result_path = sys.argv[1] + with open(result_path) as f: + data = json.load(f) + + report = data["report"] + actions = data["actions"] + registry: dict[str, dict[str, str]] = {} + + print(f"Executing {len(actions)} actions for {report['jira_issue_id']}...") + + # post_report resolves {{ref.key}} placeholders in report_md against the + # registry, which the entity-creating actions (create_subtask, + # create_root_cause_task) populate as they run. Defer every post_report until + # after the actions loop so the registry is fully populated first — otherwise + # a post_report ordered before an action it references would raise an + # uncaught KeyError from resolve_refs and abort the run. verify-pr already + # emits post_report last, so this only removes an undocumented, + # order-dependent trap; it does not change the observed behavior. + deferred_reports: list[dict] = [] + + for i, action in enumerate(actions): + action_type = action["type"] + print(f"[{i + 1}/{len(actions)}] {action_type}") + + if action_type == "post_report": + deferred_reports.append(action) + elif action_type in EXECUTORS: + EXECUTORS[action_type](action, registry) + else: + print(f" Unknown action type: {action_type}", file=sys.stderr) + sys.exit(1) + + for action in deferred_reports: + execute_post_report(action, registry, report) + + print(f"Done. {len(actions)} actions executed successfully.") + + +if __name__ == "__main__": + main() diff --git a/plugins/sdlc-workflow/scripts/execute-triage-security-actions.py b/plugins/sdlc-workflow/scripts/execute-triage-security-actions.py new file mode 100644 index 000000000..d0d70b532 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/execute-triage-security-actions.py @@ -0,0 +1,420 @@ +#!/usr/bin/env python3 +"""Execute an authorized triage-security result on the trusted runner. + +The sandbox may propose an action plan, but only the runner-provided input bundle +authorizes mutations. This module therefore treats that bundle as the authority +for both mutation permission and retry state. +""" + +import copy +import hashlib +import importlib.util +import json +import os +import re +import sys +from typing import Any + +from jsonschema import SchemaError, ValidationError, validate + + +_SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__)) +_RESULT_SCHEMA_PATH = os.path.join(_SCRIPT_DIR, "..", "schemas", "triage-security-result.schema.json") + + +def _load_module(name: str, filename: str): + """Load a sibling script without making its hyphenated filename importable.""" + spec = importlib.util.spec_from_file_location(name, os.path.join(_SCRIPT_DIR, filename)) + if spec is None or spec.loader is None: + raise RuntimeError("could not load {}".format(filename)) + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +# Reuse the established placeholder renderer and the fail-fast Jira client. +_action_helpers = _load_module("execute_actions", "execute-actions.py") +_jira_mod = _load_module("jira_client", "jira-client.py") + + +class ActionError(ValueError): + """Raised when a sandbox plan is unsafe or incompatible with trusted input.""" + + +# Required fields and authorization targets share one action definition. +_ACTION_DEFINITIONS = { + "report-only": {"required": {"type", "marker"}, "targets": ()}, + "field-edit": {"required": {"type", "marker", "issue", "fields"}, "targets": ("issue",)}, + "status-transition": {"required": {"type", "marker", "issue", "status"}, "targets": ("issue",)}, + "comment": {"required": {"type", "marker", "issue", "body_adf"}, "targets": ("issue",)}, + "link": {"required": {"type", "marker", "link_type", "inward", "outward"}, "targets": ("inward", "outward")}, + "remediation-task": {"required": {"type", "marker", "ref", "project", "summary", "description_adf", "labels"}, "targets": ("project",)}, + "resolve-reference": {"required": {"type", "marker", "ref", "issue"}, "targets": ("issue",)}, +} +_REQUIRED_ACTION_FIELDS = { + action_type: definition["required"] + for action_type, definition in _ACTION_DEFINITIONS.items() +} + + +def _plugin_version() -> str: + """Read the installed plugin version used in triage comment footers.""" + manifest = os.path.join(_SCRIPT_DIR, "..", ".claude-plugin", "plugin.json") + with open(manifest, encoding="utf-8") as manifest_file: + return json.load(manifest_file)["version"] + + +def _browse_url(issue_key: str) -> str: + """Build a Jira browse URL for a generated issue reference.""" + base_url = os.environ.get("JIRA_SERVER_URL", "https://jira.invalid").rstrip("/") + return "{}/browse/{}".format(base_url, issue_key) + + +def _validate_action(action: Any) -> None: + """Reject malformed or unknown action objects before they can be executed.""" + if not isinstance(action, dict): + raise ActionError("action must be an object") + action_type = action.get("type") + required = _REQUIRED_ACTION_FIELDS.get(action_type) + if required is None: + raise ActionError("unknown action type: {}".format(action_type)) + missing = sorted(required - set(action)) + if missing: + raise ActionError("{} action is missing {}".format(action_type, ", ".join(missing))) + if not isinstance(action["marker"], str) or not action["marker"].startswith("triage-security:"): + raise ActionError("action marker must use the triage-security namespace") + + +def _validate_result(result: Any) -> None: + """Fail closed unless sandbox output satisfies the trusted result schema.""" + try: + with open(_RESULT_SCHEMA_PATH, encoding="utf-8") as schema_file: + schema = json.load(schema_file) + validate(instance=result, schema=schema) + except (OSError, json.JSONDecodeError, SchemaError, ValidationError) as error: + raise ActionError("result schema validation failed: {}".format(error)) from error + + +def _append_footer(document: dict[str, Any]) -> dict[str, Any]: + """Append the required triage-security comment footer to an ADF document.""" + if not isinstance(document, dict) or document.get("type") != "doc": + raise ActionError("comment body_adf must be an ADF document") + rendered = copy.deepcopy(document) + content = rendered.setdefault("content", []) + if not isinstance(content, list): + raise ActionError("comment body_adf.content must be a list") + content.extend([ + {"type": "rule"}, + { + "type": "paragraph", + "content": [ + {"type": "text", "text": "This comment was AI-generated by "}, + { + "type": "text", + "text": "sdlc-workflow/triage-security", + "marks": [{"type": "link", "attrs": {"href": "https://github.com/RHEcosystemAppEng/sdlc-plugins"}}], + }, + {"type": "text", "text": " v{}.".format(_plugin_version())}, + ], + }, + ]) + return rendered + + +def _post_adf_comment(issue_key: str, document: dict[str, Any]) -> None: + """Post sanitized ADF through the existing Jira client's request path.""" + _jira_mod.make_request( + "POST", "issue/{}/comment".format(issue_key), + {"body": _jira_mod.sanitize_adf(document)}, + ) + + +def _description_digest(document: dict[str, Any]) -> str: + """Compute the protocol's compact ADF SHA-256 digest for a created task.""" + canonical = json.dumps(document, separators=(",", ":")) + return "sha256-adf:{}".format(hashlib.sha256(canonical.encode("utf-8")).hexdigest()) + + +def _post_digest(issue_key: str) -> None: + """Digest Jira's stored remediation description before dependent actions.""" + created = _jira_mod.get_issue(issue_key) + description = created.get("fields", {}).get("description") + if not isinstance(description, dict): + raise ActionError("created remediation task has no ADF description") + _post_adf_comment(issue_key, { + "type": "doc", + "version": 1, + "content": [{ + "type": "paragraph", + "content": [{ + "type": "text", + "text": "[sdlc-workflow] Description digest: {}".format(_description_digest(description)), + }], + }], + }) + + +def _existing_remediation(action: dict[str, Any], trusted_input: dict[str, Any]) -> dict[str, Any] | None: + """Return a matching prior remediation task so retries reuse its reference.""" + for existing in trusted_input.get("idempotency", {}).get("existing_remediation", []): + labels = existing.get("labels") or [] + if existing.get("summary") == action["summary"] and "ai-generated-jira" in labels and existing.get("key"): + return existing + return None + + +def _has_description_digest(existing: dict[str, Any]) -> bool: + """Identify whether trusted remediation state records its standalone digest.""" + return "[sdlc-workflow] Description digest:" in json.dumps(existing.get("comments", [])) + + +def _register_existing_remediation(action: dict[str, Any], registry: dict[str, dict[str, str]], trusted_input: dict[str, Any]) -> bool: + """Register a prior remediation task and repair a digest lost during a partial run.""" + existing = _existing_remediation(action, trusted_input) + if existing is None: + return False + issue_key = existing["key"] + registry[action["ref"]] = {"key": issue_key, "url": _browse_url(issue_key)} + if not _has_description_digest(existing): + _post_digest(issue_key) + return True + + +def _resolve_action(action: dict[str, Any], registry: dict[str, dict[str, str]]) -> dict[str, Any]: + """Resolve placeholders in an action or raise instead of sending them to Jira.""" + try: + return _action_helpers.resolve_refs_in_obj(action, registry) + except KeyError as error: + raise ActionError("unresolved action reference: {}".format(error.args[0])) from error + + +def _trusted_targets(trusted_input: dict[str, Any]) -> set[str]: + """Collect existing targets from the runner's documented triage relationships.""" + issue = trusted_input.get("issue") + primary = issue.get("key") if isinstance(issue, dict) else None + if not isinstance(primary, str) or not re.fullmatch(r"[A-Z][A-Z0-9]+-[0-9]+", primary): + raise ActionError("mutation-authorized execution requires a valid trusted issue key") + metadata = trusted_input.get("jira_metadata") or {} + targets = {primary} + related = list(metadata.get("related_issues") or []) + related.extend((trusted_input.get("idempotency") or {}).get("existing_remediation") or []) + for search in metadata.get("sibling_searches") or []: + if search.get("purpose") in {"same-cve-siblings", "cross-cve-overlap", "preemptive-remediation"}: + related.extend(search.get("issues") or []) + for item in related: + key = item.get("key") if isinstance(item, dict) else None + if not isinstance(key, str) or not re.fullmatch(r"[A-Z][A-Z0-9]+-[0-9]+", key): + raise ActionError("trusted triage relationship requires a valid issue key") + targets.add(key) + return targets + + +def _preflight_plan(result: dict[str, Any], trusted_input: dict[str, Any]) -> None: + """Authorize every target and reference before Jira reads, writes, or retry repair.""" + targets = _trusted_targets(trusted_input) + if result["report"]["issue"] != trusted_input["issue"]["key"]: + raise ActionError("report issue does not match the trusted issue") + registry: dict[str, dict[str, str]] = {} + markers = set((trusted_input.get("idempotency") or {}).get("action_markers") or []) + for raw_action in result["actions"]: + _validate_action(raw_action) + action = _resolve_action(raw_action, registry) + action_type = action["type"] + target_fields = _ACTION_DEFINITIONS.get(action_type, {}).get("targets") + if not target_fields and action_type != "report-only": + raise ActionError("{} action has no declared target fields".format(action_type)) + for field in target_fields: + if field == "project": + project = (trusted_input.get("configuration") or {}).get("project_key") + if not isinstance(project, str) or not project or action[field] != project: + raise ActionError("remediation project does not match the trusted project") + elif action[field] not in targets: + raise ActionError("unauthorized action target: {}".format(action[field])) + if action_type not in {"resolve-reference", "remediation-task"}: + continue + ref = action["ref"] + if ref in registry: + raise ActionError("action reference cannot be rebound: {}".format(ref)) + if action_type == "resolve-reference": + key = action["issue"] + else: + # Keep generated identities symbolic until the existing executor creates them. + existing = _existing_remediation( + raw_action if raw_action["marker"] in markers else action, trusted_input) + if raw_action["marker"] in markers and existing is None: + raise ActionError("marked remediation task is absent from trusted retry state") + key = existing["key"] if existing else "{{" + ref + ".key}}" + targets.add(key) + registry[ref] = {"key": key, "url": _browse_url(key)} + + +def _field_value_matches(field_name: str, snapshot_value: Any, action_value: Any) -> bool: + """Compare an action field value with its compact or full Jira snapshot form.""" + if not isinstance(action_value, dict): + return snapshot_value == action_value + if not isinstance(snapshot_value, dict): + return False + if field_name == "assignee" and set(action_value) in ({"id"}, {"accountId"}): + assignee_id = action_value.get("id", action_value.get("accountId")) + return assignee_id in (snapshot_value.get("accountId"), snapshot_value.get("id")) + if field_name == "resolution" and set(action_value) == {"id"}: + return action_value["id"] == snapshot_value.get("id") + if field_name == "resolution" and set(action_value) == {"name"}: + return snapshot_value.get("name") == action_value["name"] + return snapshot_value == action_value + + +def _already_applied(action: dict[str, Any], trusted_input: dict[str, Any]) -> bool: + """Check the trusted Jira snapshot for a mutation already applied on a retry.""" + issue = trusted_input.get("issue", {}) + fields = issue.get("fields", {}) if isinstance(issue, dict) else {} + action_type = action["type"] + if action_type in {"field-edit", "status-transition"} and action["issue"] != issue.get("key"): + return False + if action_type == "field-edit": + action_fields = action["fields"] + return bool(action_fields) and all( + _field_value_matches(name, fields.get(name), value) + for name, value in action_fields.items() + ) + if action_type == "status-transition": + return issue.get("status") == action["status"] + if action_type == "link": + snapshot_key = issue.get("key") + if snapshot_key == action["inward"]: + far_end_key = action["outward"] + elif snapshot_key == action["outward"]: + far_end_key = action["inward"] + else: + return False + for link in fields.get("issuelinks", []) or []: + if (link.get("type", {}).get("name") == action["link_type"] and + (link.get("inwardIssue", {}).get("key") == far_end_key or + link.get("outwardIssue", {}).get("key") == far_end_key)): + return True + return False + + +def _execute_action(action: dict[str, Any], registry: dict[str, dict[str, str]], trusted_input: dict[str, Any]) -> None: + """Perform one already-authorized action and update generated references.""" + action_type = action["type"] + if action_type == "report-only": + return + if action_type == "field-edit": + _jira_mod.update_issue(action["issue"], action["fields"]) + return + if action_type == "status-transition": + transitions = _jira_mod.get_transitions(action["issue"]) + transition = next( + (item for item in transitions if item.get("to", {}).get("name") == action["status"]), + None, + ) + if transition is None and action["issue"] != trusted_input["issue"]["key"]: + issue = _jira_mod.get_issue(action["issue"], fields="status") + if issue.get("fields", {}).get("status", {}).get("name") == action["status"]: + return + if transition is None or not transition.get("id"): + raise ActionError("no transition named {} for {}".format(action["status"], action["issue"])) + _jira_mod.transition_issue(action["issue"], transition["id"]) + return + if action_type == "comment": + document = _append_footer(action["body_adf"]) + # The native sticky-comment command persists this marker as a hidden + # comment identity, so the pre-script can safely skip it on a retry. + _action_helpers.post_jira_comment_native( + action["issue"], + _action_helpers.adf_to_markdown(document), + marker="".format(action["marker"]), + ) + return + if action_type == "link": + _jira_mod.create_link(action["inward"], action["outward"], action["link_type"]) + return + if action_type == "resolve-reference": + registry[action["ref"]] = {"key": action["issue"], "url": _browse_url(action["issue"])} + return + if action_type == "remediation-task": + if _register_existing_remediation(action, registry, trusted_input): + return + labels = ["ai-generated-jira", *[label for label in action["labels"] if label != "ai-generated-jira"]] + created = _jira_mod.create_issue( + project_key=action["project"], + summary=action["summary"], + issue_type="Task", + labels=labels, + priority=action.get("priority"), + fix_versions=action.get("fix_versions"), + description_adf=action["description_adf"], + ) + issue_key = created.get("key") + if not issue_key: + raise ActionError("remediation task creation returned no issue key") + registry[action["ref"]] = {"key": issue_key, "url": _browse_url(issue_key)} + _post_digest(issue_key) + return + raise ActionError("unknown action type: {}".format(action_type)) + + +def execute_plan(result: dict[str, Any], trusted_input: dict[str, Any]) -> dict[str, dict[str, str]]: + """Execute a validated result only when the trusted input grants mutations.""" + if not isinstance(result, dict) or not isinstance(trusted_input, dict): + raise ActionError("result and trusted input must be objects") + _validate_result(result) + mode = result.get("mode") + actions = result.get("actions") + if mode not in {"report-only", "mutation-authorized"} or not isinstance(actions, list): + raise ActionError("result has an invalid mode or action list") + authorized = trusted_input.get("authorization", {}).get("mutation_authorized") is True + mutating = any(isinstance(action, dict) and action.get("type") != "report-only" for action in actions) + if mode == "report-only" and mutating: + raise ActionError("report-only mode cannot contain mutating actions") + if mode == "mutation-authorized" and not authorized: + raise ActionError("mutation-authorized result is not authorized by the trusted runner") + if mode == "report-only": + for action in actions: + _validate_action(action) + if action["type"] != "report-only": + raise ActionError("report-only mode cannot contain mutating actions") + return {} + + _preflight_plan(result, trusted_input) + markers = set(trusted_input.get("idempotency", {}).get("action_markers", [])) + registry: dict[str, dict[str, str]] = {} + for raw_action in actions: + _validate_action(raw_action) + if raw_action["type"] == "resolve-reference": + _execute_action(raw_action, registry, trusted_input) + continue + if raw_action["marker"] in markers: + if raw_action["type"] == "remediation-task": + if not _register_existing_remediation(raw_action, registry, trusted_input): + raise ActionError("marked remediation task is absent from trusted retry state") + continue + action = _resolve_action(raw_action, registry) + if _already_applied(action, trusted_input): + continue + _execute_action(action, registry, trusted_input) + return registry + + +def main(argv: list[str] | None = None) -> int: + """Load runner files and exit non-zero whenever action execution fails.""" + args = list(sys.argv[1:] if argv is None else argv) + if len(args) != 2: + print("Usage: {} ".format(sys.argv[0]), file=sys.stderr) + return 1 + try: + with open(args[0], encoding="utf-8") as result_file: + result = json.load(result_file) + with open(args[1], encoding="utf-8") as input_file: + trusted_input = json.load(input_file) + execute_plan(result, trusted_input) + except Exception as error: + print("ERROR: {}".format(error), file=sys.stderr) + return 1 + print("Triage-security actions completed successfully.") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/plugins/sdlc-workflow/scripts/jira-client.py b/plugins/sdlc-workflow/scripts/jira-client.py index 5f5feec71..1d7d8ccdc 100755 --- a/plugins/sdlc-workflow/scripts/jira-client.py +++ b/plugins/sdlc-workflow/scripts/jira-client.py @@ -443,15 +443,17 @@ def get_issue(issue_key: str, fields: str = "*all") -> Dict[str, Any]: def create_issue( project_key: str, summary: str, - description_md: str, - issue_type: str, + description_md: Optional[str] = None, + issue_type: str = "", labels: Optional[List[str]] = None, assignee_id: Optional[str] = None, priority: Optional[str] = None, fix_versions: Optional[List[str]] = None, - custom_fields: Optional[Dict[str, Any]] = None + custom_fields: Optional[Dict[str, Any]] = None, + description_adf: Optional[Dict[str, Any]] = None, + parent: Optional[str] = None ) -> Dict[str, Any]: - """Create JIRA issue with markdown description. + """Create JIRA issue with a markdown or pre-rendered ADF description. Args: project_key: Project key (e.g., TC) @@ -463,19 +465,47 @@ def create_issue( priority: Optional priority name (e.g., "Major") fix_versions: Optional list of fixVersion names custom_fields: Optional custom field values (field_id: value) + description_adf: Pre-rendered ADF description (takes precedence over description_md) + parent: Optional parent issue key (creates a sub-task under it) Returns: Created issue object with key and ID """ + # Fail fast on a missing issue type: an empty (or whitespace-only) value would + # otherwise serialize to {"name": ""} below and surface as an opaque Jira 400 + # instead of a clear programming error at the call site. + if not issue_type or not issue_type.strip(): + print( + "❌ Missing issue type: create_issue requires a non-empty issue_type", + file=sys.stderr, + ) + print( + "Pass an issue type ID (e.g., 10001) or name (e.g., Task, Sub-task).", + file=sys.stderr, + ) + sys.exit(1) + + # Prefer a pre-rendered ADF description; fall back to converting markdown, + # then to an empty document when neither is supplied. + if description_adf is not None: + description = description_adf + elif description_md is not None: + description = markdown_to_adf(description_md) + else: + description = {"type": "doc", "version": 1, "content": []} + data = { "fields": { "project": {"key": project_key}, "summary": summary, - "description": sanitize_adf(markdown_to_adf(description_md)), + "description": sanitize_adf(description), "issuetype": {"id": issue_type} if issue_type.isdigit() else {"name": issue_type}, } } + if parent: + data["fields"]["parent"] = {"key": parent} + if labels: data["fields"]["labels"] = labels @@ -550,26 +580,95 @@ def search_jql( jql: str, fields: Optional[str] = None, max_results: int = 50, - start_at: int = 0 + next_page_token: Optional[str] = None ) -> Dict[str, Any]: """Search JIRA issues using JQL. + Uses the enhanced search endpoint ``POST /rest/api/3/search/jql``. The legacy + ``/rest/api/3/search`` endpoint was removed by Atlassian and now returns + HTTP 410 Gone. The enhanced endpoint drops the ``total`` field and replaces + the ``startAt`` offset with opaque ``nextPageToken`` cursor pagination. + + POST (JSON body) is used rather than GET so ``fields`` is sent as a proper + JSON array. The GET variant takes ``fields`` as a comma-separated query + value, and URL-encoding the commas (``%2C``) can make the endpoint treat the + whole list as one unknown field name — the issues come back without the + requested custom fields, which silently breaks callers that read them. + Args: jql: JQL query string fields: Comma-separated field names (default: summary,status,assignee) max_results: Maximum results per page (max 50) - start_at: Pagination offset + next_page_token: Opaque cursor from a prior response's ``nextPageToken`` + (omit for the first page) Returns: - Search results with issues array and total count + Search results with an ``issues`` array; ``nextPageToken`` is present + when more pages remain (``isLast`` is false) """ if fields is None: fields = "summary,status,assignee,priority,issuetype,labels" - from urllib.parse import quote - jql_encoded = quote(jql) - endpoint = f"search?jql={jql_encoded}&fields={fields}&maxResults={max_results}&startAt={start_at}" - return make_request('GET', endpoint) + body: Dict[str, Any] = { + "jql": jql, + "fields": [f.strip() for f in fields.split(",") if f.strip()], + "maxResults": max_results, + } + if next_page_token: + body["nextPageToken"] = next_page_token + return make_request('POST', 'search/jql', body) + + +def search_jql_all( + jql: str, + fields: Optional[str] = None, + page_size: int = 50, + max_pages: int = 200, +) -> Dict[str, Any]: + """Search JIRA issues using JQL, following ``nextPageToken`` across all pages. + + ``search_jql`` returns a single page (at most 50 issues) and an opaque + ``nextPageToken`` when more remain, so a single-page call silently drops any + match beyond the first page. A broad ``~`` recall on the Git Pull Request + custom field can exceed one page, which would make ``resolve_gated_issue`` + miss an exact PR match and emit a false ADR-0072 skip. This helper follows + the cursor to the last page and returns one response whose ``issues`` array + holds every matched issue, so an exact-match verifier sees the full set. + + Args: + jql: JQL query string + fields: Comma-separated field names (default matches ``search_jql``) + page_size: Results per page (the endpoint caps this at 50 server-side) + max_pages: Safety bound on pages fetched, to avoid an unbounded loop if + the server keeps returning a cursor + + Returns: + A search-response-shaped dict with the aggregated ``issues`` array and + ``isLast`` true (``nextPageToken`` omitted). + + Raises: + SystemExit: propagated from ``make_request`` on any Jira/HTTP error, or + raised here (exit 1) if ``max_pages`` is exhausted while the server + still offers a cursor — failing loud rather than returning a + possibly-truncated result set (CONVENTIONS.md §Error Handling). + """ + all_issues: List[Dict[str, Any]] = [] + token: Optional[str] = None + for _ in range(max_pages): + page = search_jql( + jql, fields=fields, max_results=page_size, next_page_token=token + ) + all_issues.extend(page.get("issues", [])) + token = page.get("nextPageToken") + if not token: + return {"issues": all_issues, "isLast": True} + + print( + f"❌ search_jql_all: exceeded max_pages={max_pages} while paginating; " + "refusing to return a possibly-truncated result set", + file=sys.stderr, + ) + sys.exit(1) def create_link( @@ -718,8 +817,10 @@ def main(argv=None): search_parser = subparsers.add_parser('search_jql', help='Search issues with JQL') search_parser.add_argument('--jql', required=True, help='JQL query string') search_parser.add_argument('--fields', help='Comma-separated fields') - search_parser.add_argument('--max-results', type=int, default=50, help='Max results (default: 50)') - search_parser.add_argument('--start-at', type=int, default=0, help='Pagination offset') + search_parser.add_argument('--max-results', type=int, default=50, help='Max results per page (default: 50, endpoint cap)') + search_parser.add_argument('--next-page-token', help='Opaque cursor from a prior response nextPageToken (single-page pagination)') + search_parser.add_argument('--all', action='store_true', + help='Follow nextPageToken and return every matching issue (ignores --next-page-token)') # create_link link_parser = subparsers.add_parser('create_link', help='Create issue link') @@ -791,7 +892,10 @@ def main(argv=None): result = get_transitions(args.issue_key) elif args.command == 'search_jql': - result = search_jql(args.jql, args.fields, args.max_results, args.start_at) + if args.all: + result = search_jql_all(args.jql, args.fields, args.max_results) + else: + result = search_jql(args.jql, args.fields, args.max_results, args.next_page_token) elif args.command == 'create_link': create_link(args.inward, args.outward, args.link_type) @@ -812,8 +916,11 @@ def main(argv=None): elif args.command == 'get_versions': result = get_versions(args.project_key, args.unreleased_only) - # Print result as JSON - if result: + # Print result as JSON. Guard on ``is not None`` (not truthiness) so an empty + # collection — e.g. get_versions/get_remote_links/search_jql with no matches — + # is still emitted as valid JSON ([] or {}). Printing nothing would make JSON + # callers raise a misleading "invalid JSON" error on empty stdout. + if result is not None: print(json.dumps(result, indent=2)) diff --git a/plugins/sdlc-workflow/scripts/post-triage-security.sh b/plugins/sdlc-workflow/scripts/post-triage-security.sh new file mode 100755 index 000000000..2eb46be7b --- /dev/null +++ b/plugins/sdlc-workflow/scripts/post-triage-security.sh @@ -0,0 +1,58 @@ +#!/usr/bin/env bash +# post-triage-security.sh — execute a schema-validated triage-security plan. +# +# This runs only on the trusted Fullsend runner after the sandbox exits. It +# selects the final iteration result and passes it together with the trusted +# pre-script authorization bundle to the action executor. + +set -euo pipefail + +: "${FULLSEND_RUN_DIR:?FULLSEND_RUN_DIR is required}" + +RUN_DIR="$(cd "${FULLSEND_RUN_DIR}" && pwd -P)" +TRUSTED_INPUT="${RUN_DIR}/pre/triage-security-input.json" +if [[ ! -f "${TRUSTED_INPUT}" ]]; then + echo "ERROR: trusted triage-security input is missing: ${TRUSTED_INPUT}" >&2 + exit 1 +fi + +RESULT_FILE="" +while IFS= read -r iteration_dir; do + [[ -d "${RUN_DIR}/${iteration_dir}" ]] || continue + if [[ -f "${RUN_DIR}/${iteration_dir}/agent-result.json" ]]; then + RESULT_FILE="${iteration_dir}/agent-result.json" + elif [[ -f "${RUN_DIR}/${iteration_dir}/result.json" ]]; then + RESULT_FILE="${iteration_dir}/result.json" + fi +done < <(cd "${RUN_DIR}" && printf '%s\n' iteration-*/output | sort -V) + +if [[ -z "${RESULT_FILE}" ]]; then + echo "ERROR: no validated triage-security result found in an iteration output directory" >&2 + exit 1 +fi + +RESULT_FILE="${RUN_DIR}/${RESULT_FILE}" +RESULT_REAL="$(realpath "${RESULT_FILE}")" +INPUT_REAL="$(realpath "${TRUSTED_INPUT}")" +case "${RESULT_REAL}" in + "${RUN_DIR}"/*) ;; + *) + echo "ERROR: selected result path escapes FULLSEND_RUN_DIR" >&2 + exit 1 + ;; +esac +case "${INPUT_REAL}" in + "${RUN_DIR}"/*) ;; + *) + echo "ERROR: trusted input path escapes FULLSEND_RUN_DIR" >&2 + exit 1 + ;; +esac + +if ! jq empty "${RESULT_REAL}" >/dev/null 2>&1; then + echo "ERROR: selected triage-security result is not valid JSON" >&2 + exit 1 +fi + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +exec python3 "${SCRIPT_DIR}/execute-triage-security-actions.py" "${RESULT_REAL}" "${INPUT_REAL}" diff --git a/plugins/sdlc-workflow/scripts/post-verify-pr.sh b/plugins/sdlc-workflow/scripts/post-verify-pr.sh new file mode 100755 index 000000000..c6ae68abd --- /dev/null +++ b/plugins/sdlc-workflow/scripts/post-verify-pr.sh @@ -0,0 +1,57 @@ +#!/usr/bin/env bash +# post-verify-pr.sh — Execute verify-pr structured output actions. +# +# Runs on the fullsend runner AFTER the sandbox is destroyed. +# Working directory is the fullsend run output directory. +# +# Required env vars: +# JIRA_SERVER_URL — Jira instance URL +# JIRA_EMAIL — Jira user email +# JIRA_API_TOKEN — Jira API token +# JIRA_PROJECT_KEY — Jira project key (for root-cause task creation) +# GH_TOKEN — GitHub token +# +# The agent writes its output to output/agent-result.json (relative to +# the iteration directory). This script finds the most recent iteration's +# output and delegates to execute-actions.py. + +set -euo pipefail + +RESULT_FILE="" +# Iterate iteration directories in ascending numeric order so the +# highest-numbered iteration that has a result file wins. Plain glob order is +# lexicographic (iteration-9 sorts after iteration-20), which would select a +# stale iteration once there are >= 10 iterations. `sort -V` orders the +# embedded iteration numbers numerically; the `[[ -d ]]` guard skips the +# literal glob pattern when no iteration directory exists. +while IFS= read -r dir; do + [[ -d "${dir}" ]] || continue + # Prefer agent-result.json; fall back to result.json when it is absent, + # matching the precedence in validate-output-schema.sh (agents sometimes + # write "result.json" instead of "agent-result.json"). + if [[ -f "${dir}/agent-result.json" ]]; then + RESULT_FILE="${dir}/agent-result.json" + elif [[ -f "${dir}/result.json" ]]; then + RESULT_FILE="${dir}/result.json" + fi +done < <(printf '%s\n' iteration-*/output | sort -V) + +if [[ -z "${RESULT_FILE}" ]]; then + echo "ERROR: no agent-result.json or result.json found in any iteration output directory" + exit 1 +fi + +echo "Reading verify-pr result from: ${RESULT_FILE}" + +if ! jq empty "${RESULT_FILE}" 2>/dev/null; then + echo "ERROR: ${RESULT_FILE} is not valid JSON" + exit 1 +fi + +OVERALL=$(jq -r '.report.overall' "${RESULT_FILE}") +ACTION_COUNT=$(jq '.actions | length' "${RESULT_FILE}") +echo "Overall: ${OVERALL}" +echo "Actions: ${ACTION_COUNT}" + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +python3 "${SCRIPT_DIR}/execute-actions.py" "${RESULT_FILE}" diff --git a/plugins/sdlc-workflow/scripts/pre-triage-security.sh b/plugins/sdlc-workflow/scripts/pre-triage-security.sh new file mode 100755 index 000000000..26bef3425 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/pre-triage-security.sh @@ -0,0 +1,38 @@ +#!/usr/bin/env bash +# pre-triage-security.sh — collect one poller-dispatched security issue's +# credentialed evidence before Fullsend creates the token-free sandbox. + +set -euo pipefail + +# The Fullsend poller dispatches one Jira work item per run. The pre-script does +# not discover issues: it validates this URL, derives its Jira key, and gathers +# the complete evidence bundle for that one issue. +: "${FULLSEND_WORK_ITEM_URL:?FULLSEND_WORK_ITEM_URL (Jira work-item URL) is required}" +: "${JIRA_SERVER_URL:?JIRA_SERVER_URL is required}" +: "${JIRA_EMAIL:?JIRA_EMAIL is required}" +: "${JIRA_API_TOKEN:?JIRA_API_TOKEN is required}" + +if [[ ! "${FULLSEND_WORK_ITEM_URL}" =~ ^https?://[^/]+/browse/([A-Z][A-Z0-9]+-[0-9]+)/?$ ]]; then + echo "ERROR: FULLSEND_WORK_ITEM_URL must be a Jira issue URL" >&2 + exit 1 +fi +ISSUE_KEY="${BASH_REMATCH[1]}" + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +PROJECT_ROOT="${FULLSEND_PROJECT_ROOT:-$(pwd)}" +if [[ ! -f "${PROJECT_ROOT}/CLAUDE.md" ]]; then + echo "ERROR: target project CLAUDE.md is required at ${PROJECT_ROOT}/CLAUDE.md" >&2 + exit 1 +fi + +: "${FULLSEND_RUN_DIR:?FULLSEND_RUN_DIR is required}" +PRE_OUTPUT_DIR="${FULLSEND_RUN_DIR}/pre" +mkdir -p "${PRE_OUTPUT_DIR}" +OUTPUT_FILE="${PRE_OUTPUT_DIR}/triage-security-input.json" +TEMP_FILE="$(mktemp "${PRE_OUTPUT_DIR}/.${ISSUE_KEY}.XXXXXX")" +trap 'rm -f "${TEMP_FILE}"' EXIT + +python3 "${SCRIPT_DIR}/pre_triage_security.py" collect "${ISSUE_KEY}" "${PROJECT_ROOT}" > "${TEMP_FILE}" +mv "${TEMP_FILE}" "${OUTPUT_FILE}" + +echo "Pre-fetched triage-security evidence for ${ISSUE_KEY} to ${OUTPUT_FILE}" diff --git a/plugins/sdlc-workflow/scripts/pre-verify-pr.sh b/plugins/sdlc-workflow/scripts/pre-verify-pr.sh new file mode 100755 index 000000000..3d36a79ea --- /dev/null +++ b/plugins/sdlc-workflow/scripts/pre-verify-pr.sh @@ -0,0 +1,314 @@ +#!/usr/bin/env bash +# pre-verify-pr.sh — Derive + gate the Jira task from the triggering PR, then +# pre-fetch Jira + GitHub data. +# +# Runs on the fullsend runner BEFORE the sandbox is created, where the Jira +# and GitHub tokens live. The sandbox never sees a token — it reads only the +# JSON this script produces. +# +# The triggering PR URL is the SOLE entry point — the Jira key is derived, not +# supplied. fullsend's harness-run exports the PR URL as FULLSEND_WORK_ITEM_URL +# on the runner (a local run must export the same var). The Jira key is resolved +# by a JQL search on the Git Pull Request custom field (customfield_10875) and +# then re-verified to be in status Review AND carry the ai-generated-jira label, +# so a stale or unqualified PR is never reviewed. +# +# 1. Validates required env vars and the PR URL shape +# 2. Resolves + gates the Jira key by JQL on customfield_10875 == PR URL +# 3. If the gate fails (no/ambiguous match, status != Review, label absent): +# emits an ADR-0072 skip signal and exits 0 (nothing to verify) +# 4. Fetches the full Jira issue and prefetches the GitHub tier-1 read bundle +# (diff, stat, reviews, comments, commits, CI check-runs, failed-check logs, +# head ref + commit SHA) so the sandbox needs no api.github.com egress +# 5. Writes the tracker-agnostic verify-pr-input.json that host_files mounts +# into the sandbox +# +# Required env vars: +# FULLSEND_WORK_ITEM_URL — the triggering PR URL (harness-run / local export) +# JIRA_SERVER_URL — Jira instance URL +# JIRA_EMAIL — Jira user email +# JIRA_API_TOKEN — Jira API token +# GH_TOKEN — GitHub token (PR prefetch) +# +# Optional env vars: +# PRE_DIR — output directory (default: /tmp/fullsend-pre-output). +# The harness host_files src is the default path. +# FULLSEND_PRESCRIPT_OUTPUT — key=value skip-signal file created by fullsend run +# (pre-script output protocol v1). Guarded — older +# CLIs leave it unset. + +set -euo pipefail + +# 1. Validate required env vars are set +: "${FULLSEND_WORK_ITEM_URL:?FULLSEND_WORK_ITEM_URL (triggering PR URL) is required}" +: "${JIRA_SERVER_URL:?JIRA_SERVER_URL is required}" +: "${JIRA_EMAIL:?JIRA_EMAIL is required}" +: "${JIRA_API_TOKEN:?JIRA_API_TOKEN is required}" + +PR_URL="${FULLSEND_WORK_ITEM_URL}" + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +PRE_OUTPUT_DIR="${PRE_DIR:-/tmp/fullsend-pre-output}" +mkdir -p "${PRE_OUTPUT_DIR}" + +# request_skip REASON — emit an ADR-0072 skip signal and exit cleanly. +# The pre-script output protocol (v1) is line-based key=value; the guard on +# FULLSEND_PRESCRIPT_OUTPUT matches the scaffold's GITHUB_OUTPUT guard, so a +# missing variable (older CLI) fails open to a normal run rather than erroring. +request_skip() { + local reason="$1" + echo "SKIP: ${reason}" + if [[ -n "${FULLSEND_PRESCRIPT_OUTPUT:-}" ]]; then + { + echo "skipped=true" + echo "reason=${reason}" + } >> "${FULLSEND_PRESCRIPT_OUTPUT}" + fi + exit 0 +} + +# 2. Parse owner/repo/number from the PR URL. End-anchor the pattern (allowing +# only an optional trailing slash) so a malformed value like `.../pull/42abc` +# or `.../pull/42/extra` is rejected outright instead of silently truncating +# the pull number to 42 and prefetching the wrong PR. +if [[ ! "${PR_URL}" =~ ^https://github\.com/([^/]+/[^/]+)/pull/([0-9]+)/?$ ]]; then + echo "ERROR: PR URL '${PR_URL}' is not a github.com pull request URL" + exit 1 +fi +PR_REPO="${BASH_REMATCH[1]}" +PR_NUM="${BASH_REMATCH[2]}" +echo "PR: ${PR_REPO}#${PR_NUM} (${PR_URL})" + +# 3. Resolve + gate the Jira key by JQL on the Git Pull Request custom field. +# The JQL `~` recall is broad; resolve-gated-issue re-confirms the exact PR +# URL and enforces status=Review + the ai-generated-jira label in Python. +# `--all` follows nextPageToken across every page so an exact match beyond the +# first 50 recall results is never dropped (which would emit a false ADR-0072 +# skip); the Python exact-match verification remains the source of truth. +PR_JQL=$(python3 "${SCRIPT_DIR}/pre_verify_pr.py" build-pr-jql "${PR_URL}") +SEARCH_JSON=$(python3 "${SCRIPT_DIR}/jira-client.py" search_jql \ + --jql "${PR_JQL}" --fields "status,labels,customfield_10875" --all \ + 2>"/tmp/fullsend-pre-jira-stderr.txt") || { + JIRA_STDERR=$(cat /tmp/fullsend-pre-jira-stderr.txt 2>/dev/null || echo "") + if echo "${JIRA_STDERR}" | grep -qi "401\|unauthorized"; then + echo "ERROR: Jira authentication failed — check JIRA_EMAIL and JIRA_API_TOKEN" + elif echo "${JIRA_STDERR}" | grep -qi "403\|forbidden"; then + echo "ERROR: Jira permission denied — check that the API token can search issues" + else + echo "ERROR: Failed to search Jira for PR ${PR_URL}" + echo "${JIRA_STDERR}" + fi + rm -f /tmp/fullsend-pre-jira-stderr.txt + exit 1 +} +rm -f /tmp/fullsend-pre-jira-stderr.txt + +# resolve-gated-issue exits 3 when a gate fails (→ ADR-0072 skip), 0 with the +# resolved key on success, and 1 on an unexpected error. Capture without letting +# set -e abort on the non-zero gate/skip exit. +set +e +GATE_OUT=$(printf '%s\n' "${SEARCH_JSON}" | python3 "${SCRIPT_DIR}/pre_verify_pr.py" resolve-gated-issue "${PR_URL}") +GATE_RC=$? +set -e +if [[ ${GATE_RC} -eq 3 ]]; then + request_skip "${GATE_OUT}" +elif [[ ${GATE_RC} -ne 0 ]]; then + echo "ERROR: failed to resolve the Jira issue for PR ${PR_URL}" + echo "${GATE_OUT}" + exit 1 +fi +JIRA_ISSUE_ID="${GATE_OUT}" +echo "Jira issue resolved + gated: ${JIRA_ISSUE_ID} (status=Review, label=ai-generated-jira)" + +# 4. Fetch full issue details (pre-fetches the task context for the sandbox) +ISSUE_JSON=$(python3 "${SCRIPT_DIR}/jira-client.py" get_issue "${JIRA_ISSUE_ID}" --fields "*all" 2>"/tmp/fullsend-pre-jira-stderr.txt") || { + JIRA_STDERR=$(cat /tmp/fullsend-pre-jira-stderr.txt 2>/dev/null || echo "") + if echo "${JIRA_STDERR}" | grep -qi "401\|unauthorized"; then + echo "ERROR: Jira authentication failed — check JIRA_EMAIL and JIRA_API_TOKEN" + elif echo "${JIRA_STDERR}" | grep -qi "403\|forbidden"; then + echo "ERROR: Jira permission denied — check that the API token has access to ${JIRA_ISSUE_ID}" + elif echo "${JIRA_STDERR}" | grep -qi "404\|not found"; then + echo "ERROR: Jira issue ${JIRA_ISSUE_ID} not found" + else + echo "ERROR: Failed to fetch Jira issue ${JIRA_ISSUE_ID}" + echo "${JIRA_STDERR}" + fi + rm -f /tmp/fullsend-pre-jira-stderr.txt + exit 1 +} +rm -f /tmp/fullsend-pre-jira-stderr.txt + +echo "Jira issue verified: ${JIRA_ISSUE_ID}" + +# 5. GitHub tier-1 prefetch — runs on the trusted runner where GH_TOKEN lives. +# Every read the verify-pr skill performs against the PR is captured here so +# the sandbox needs no api.github.com egress. +: "${GH_TOKEN:?GH_TOKEN is required to prefetch PR ${PR_REPO}#${PR_NUM}}" + +HEAD_REF=$(gh pr view "${PR_NUM}" -R "${PR_REPO}" --json headRefName --jq .headRefName) +# Derive the head SHA from the ref tip OID, not from `.commits[-1].oid`: gh's +# `pr view` commits connection is bounded, so on a PR with more commits than the +# cap `.commits[-1]` is the last commit of a truncated page rather than the head +# — a silently-wrong SHA that still satisfies the result schema's hex pattern. +COMMIT_SHA=$(gh pr view "${PR_NUM}" -R "${PR_REPO}" --json headRefOid --jq .headRefOid) + +gh pr diff "${PR_NUM}" -R "${PR_REPO}" > "${PRE_OUTPUT_DIR}/pr.diff" +# `gh pr diff` has no --stat flag; derive the per-file diffstat from the patch we +# just fetched using a supported git command. Guard the empty-diff case: `git +# apply --stat` errors on an empty patch, which would abort under set -euo pipefail. +if [[ -s "${PRE_OUTPUT_DIR}/pr.diff" ]]; then + git apply --stat "${PRE_OUTPUT_DIR}/pr.diff" > "${PRE_OUTPUT_DIR}/pr.stat" +else + : > "${PRE_OUTPUT_DIR}/pr.stat" +fi +# GitHub REST returns ~30 items per page; without --paginate the reviews and +# comments are silently truncated on any active PR. --slurp aggregates the +# per-page arrays into an array-of-pages, which `jq 'add'` concatenates back +# into the single flat array that pre_verify_pr.py's build_github_bundle +# expects. (--slurp cannot be combined with gh's built-in --jq, so the merge +# uses a standalone jq.) pipefail makes a failed gh or jq abort the script. +gh api --paginate --slurp "repos/${PR_REPO}/pulls/${PR_NUM}/reviews" | jq 'add' > "${PRE_OUTPUT_DIR}/reviews.json" +gh api --paginate --slurp "repos/${PR_REPO}/pulls/${PR_NUM}/comments" | jq 'add' > "${PRE_OUTPUT_DIR}/review-comments.json" +gh api --paginate --slurp "repos/${PR_REPO}/issues/${PR_NUM}/comments" | jq 'add' > "${PRE_OUTPUT_DIR}/issue-comments.json" +# The bundled commits list uses gh's bounded `pr view` commits connection, so it +# may be truncated on a very large PR. The authoritative head SHA is taken from +# headRefOid above; this list is best-effort context for commit-traceability. If +# a consumer ever treats it as authoritative, switch to a paginated +# `gh api --paginate .../pulls/${PR_NUM}/commits` (which returns a different, +# `sha`-shaped object needing reshaping to match the `oid` contract). +gh pr view "${PR_NUM}" -R "${PR_REPO}" --json commits --jq .commits > "${PRE_OUTPUT_DIR}/commits.json" +# CI check-run outcomes for the head SHA. correctness.md Check 1 (CI Status) reads +# these in sandbox mode instead of shelling out to `gh pr checks`/`gh run view` +# (no gh CLI or egress in the sandbox). The check-runs endpoint returns an OBJECT +# per page ({total_count, check_runs:[...]}), so unlike the array-returning +# reviews/comments endpoints the merge flattens `.[].check_runs[]` across pages +# rather than `add`. The authoritative head SHA is COMMIT_SHA (headRefOid, above), +# and the wait-for-checks job in fullsend-verify-pr.yml guarantees terminal states +# before dispatch. Reduced to the fields the verdict needs (name/status/conclusion +# and details_url for the failure-log link) to keep the bundle small. A PR with no +# checks yields an empty array — consistent with the reviews/comments empties. +gh api --paginate --slurp "repos/${PR_REPO}/commits/${COMMIT_SHA}/check-runs" \ + | jq '[.[].check_runs[] | {name, status, conclusion, details_url}]' > "${PRE_OUTPUT_DIR}/check-runs.json" + +# Self-exclusion of verify-pr's OWN workflow check-runs (TC-6343). check-runs.json +# above lists EVERY check on the head SHA — including this workflow's own runs +# (the in-progress dispatch and any superseded prior attempt), which are +# non-terminal/failed at evaluation time and would drag correctness.md Check 1a's +# CI Status to a permanent self-referential WARN/FAIL. This is the CI-Status +# analogue of Step 1's `running-workflow-name: wait-for-checks` self-exclusion in +# fullsend-verify-pr.yml. Gather the check-run NAMES this workflow +# (`fullsend verify-pr`, .github/workflows/fullsend-verify-pr.yml) produced at the +# head SHA — across ALL of its runs at that SHA, so a superseded attempt with a +# different run ID is covered too — one per line so job names with spaces survive. +# The transform's pure Python filter (filter_own_check_runs) drops them before +# they reach github.check_runs, keeping the exclusion unit-testable. Best-effort: +# a token lacking actions:read (or an API hiccup) leaves the names file empty, +# degrading to no self-exclusion (prior behavior) with a warning rather than +# aborting the whole review — the fetch failure is non-fatal, like the per-run log +# fetch below. +OWN_WORKFLOW_FILE="fullsend-verify-pr.yml" +OWN_CHECK_NAMES_FILE="${PRE_OUTPUT_DIR}/own-check-names.txt" +: > "${OWN_CHECK_NAMES_FILE}" +set +e +OWN_RUN_IDS=$(gh api --paginate \ + "repos/${PR_REPO}/actions/workflows/${OWN_WORKFLOW_FILE}/runs?head_sha=${COMMIT_SHA}" \ + --jq '.workflow_runs[].id' 2>/dev/null) +for own_run_id in ${OWN_RUN_IDS}; do + gh api --paginate "repos/${PR_REPO}/actions/runs/${own_run_id}/jobs" \ + --jq '.jobs[].name' 2>/dev/null +done | sort -u > "${OWN_CHECK_NAMES_FILE}" +set -e +if [[ -s "${OWN_CHECK_NAMES_FILE}" ]]; then + echo "Self-exclusion: filtering $(grep -c . "${OWN_CHECK_NAMES_FILE}") own check-run name(s) from CI Status" +else + echo "WARNING: could not enumerate verify-pr's own check-runs for ${COMMIT_SHA}; CI Status self-exclusion is a no-op this run" +fi + +# Failed-check logs. correctness.md Check 1b needs the failure logs to analyse a +# red CI check, but the sandbox has no `gh` CLI or egress to run +# `gh run view --log-failed`. host_files mounts single files only (fullsend has +# no directory mount), and the set of failed checks is dynamic, so the logs are +# concatenated into ONE file mounted alongside verify-pr-input.json; the sub-agent +# reads it only when Check 1 is FAIL, keeping the (large) log text out of the +# input bundle/schema and off the agent's context on the common green path. The +# wait-for-checks job (fullsend-verify-pr.yml) guarantees terminal conclusions +# before dispatch, so a failed check's log is complete and fetchable here. The +# file is always created (empty when nothing failed) so its host_files mount is +# never missing. Distinct GitHub Actions run IDs are extracted from each FAILED +# check-run's details_url (.../actions/runs//...); non-Actions checks have +# no such URL and are skipped (their logs aren't reachable via `gh run view` — 1b +# falls back to the diff + details_url for those). A per-run fetch failure is +# non-fatal: a note is written and the run continues, since 1b can still fall back. +CHECK_LOGS_FILE="${PRE_OUTPUT_DIR}/check-run-logs.txt" +: > "${CHECK_LOGS_FILE}" +FAILED_RUN_IDS=$(jq -r ' + [ .[] + | select(.conclusion // "" | IN("failure", "timed_out", "cancelled", "action_required")) + | (.details_url // "") + | select(test("actions/runs/[0-9]+")) + | capture("actions/runs/(?[0-9]+)").id + ] | unique | .[]' "${PRE_OUTPUT_DIR}/check-runs.json") +for run_id in ${FAILED_RUN_IDS}; do + { + echo "===== CI run ${run_id} — failed steps =====" + gh run view "${run_id}" --log-failed -R "${PR_REPO}" 2>&1 \ + || echo "(log fetch failed for run ${run_id}; see its details_url in check_runs)" + echo + } >> "${CHECK_LOGS_FILE}" +done + +echo "GitHub read bundle prefetched to ${PRE_OUTPUT_DIR}" + +# 6. Idempotency prefetch — the sandbox has no Jira token, but Steps 6d/6f/7c +# dedupe against the task's existing sub-tasks and linked (e.g., root-cause) +# issues. Fetch each related issue here on the trusted runner (summary, +# labels, description, issuetype, and comments) so the sandbox can run those +# checks tokenlessly. No `|| true`: a related issue the token created should +# be readable, so a fetch failure is a real error surfaced under set -e. +REL_DIR="${PRE_OUTPUT_DIR}/related-issues" +rm -rf "${REL_DIR}" +mkdir -p "${REL_DIR}" +RELATED_KEYS=$(printf '%s\n' "${ISSUE_JSON}" | python3 "${SCRIPT_DIR}/pre_verify_pr.py" related-keys) +for key in ${RELATED_KEYS}; do + python3 "${SCRIPT_DIR}/jira-client.py" get_issue "${key}" \ + --fields "summary,labels,description,issuetype,comment" \ + > "${REL_DIR}/${key}.json" +done +echo "Idempotency read bundle prefetched to ${REL_DIR}" + +# 7. Re-validate the gate on the FULL issue actually used to build the sandbox +# input, immediately before the write. Step 3 gated the lightweight JQL search +# response; ISSUE_JSON came from a SECOND fetch (Step 4), so the issue may have +# left status Review, lost the ai-generated-jira label, or had its Git Pull +# Request field changed since — a TOCTOU gap. Re-run the exact PR-URL + status +# + label gate here and map a failure to the same ADR-0072 skip, so a stale +# successful gate can never launch a verification on a now-unqualified issue. +# Same capture pattern as Step 3: exit 3 → skip, exit != 0 → hard error. +set +e +REVAL_OUT=$(printf '%s\n' "${ISSUE_JSON}" | python3 "${SCRIPT_DIR}/pre_verify_pr.py" revalidate-gate "${PR_URL}") +REVAL_RC=$? +set -e +if [[ ${REVAL_RC} -eq 3 ]]; then + request_skip "${REVAL_OUT}" +elif [[ ${REVAL_RC} -ne 0 ]]; then + echo "ERROR: failed to re-validate the gate for ${JIRA_ISSUE_ID} before writing sandbox input" + echo "${REVAL_OUT}" + exit 1 +fi +echo "Jira issue re-gated on full fetch: ${JIRA_ISSUE_ID} (status=Review, label=ai-generated-jira)" + +# 8. Write pre-fetched data for sandbox consumption (tracker-agnostic format, +# with the GitHub bundle embedded under `github` and the idempotency +# related-issue metadata under `idempotency`). +printf '%s\n' "${ISSUE_JSON}" | python3 "${SCRIPT_DIR}/pre_verify_pr.py" transform \ + "${JIRA_ISSUE_ID}" "${PR_URL}" \ + --github-dir "${PRE_OUTPUT_DIR}" \ + --pr-repo "${PR_REPO}" \ + --pr-number "${PR_NUM}" \ + --head-ref "${HEAD_REF}" \ + --commit-sha "${COMMIT_SHA}" \ + --own-check-names-file "${OWN_CHECK_NAMES_FILE}" \ + --idempotency-dir "${REL_DIR}" > "${PRE_OUTPUT_DIR}/verify-pr-input.json" + +echo "Pre-fetched data written to ${PRE_OUTPUT_DIR}/verify-pr-input.json" +echo "Input validation passed" diff --git a/plugins/sdlc-workflow/scripts/pre_triage_security.py b/plugins/sdlc-workflow/scripts/pre_triage_security.py new file mode 100644 index 000000000..a6dfe016a --- /dev/null +++ b/plugins/sdlc-workflow/scripts/pre_triage_security.py @@ -0,0 +1,753 @@ +#!/usr/bin/env python3 +"""Build and validate trusted evidence bundles for triage-security.""" + +import argparse +import json +import os +import re +import subprocess +import sys +from datetime import datetime, timezone +from pathlib import Path +from urllib.error import HTTPError, URLError +from urllib.request import Request, urlopen + +from jsonschema import FormatChecker, ValidationError, validate + + +_ISSUE_KEY_RE = re.compile(r"^[A-Z][A-Z0-9]+-[0-9]+$") +_SCHEMA_PATH = Path(__file__).parent.parent / "schemas" / "triage-security-input.schema.json" +_EVIDENCE_USER_AGENT = "Mozilla/5.0 (compatible; sdlc-triage-security/1.0)" + + +class EvidenceError(ValueError): + """Raised when runner evidence is incomplete or inconsistent.""" + + +def _ref_token(cell): + """Extract a git ref/branch token from a matrix cell. + + Matrix cells may wrap the ref in backticks and append a human annotation + (e.g. ``release/0.6.z (pending re-point to 0.7.z)``). A git ref never + contains whitespace or parentheses, so take the ref at the start of the + cell after removing a leading prose annotation. + """ + original = cell + cell = (cell or "").strip().strip("`").strip() + if not cell: + return "" + prefix = re.match(r"(?:\([^)]*\)\s*|N/A\s*-\s*see\s+)", cell, re.IGNORECASE) + if prefix: + cell = cell[prefix.end():] + match = re.match(r"[A-Za-z0-9._/\-]+", cell.lstrip(" `")) + if not match: + raise EvidenceError("matrix cell has no ref token: {!r}".format(original)) + return match.group(0) + + +def _markdown_section(document, heading, level): + """Return a Markdown section body without its heading.""" + pattern = r"^{} {}\s*$\n?(.*?)(?=^#{{1,{}}}\s|\Z)".format( + "#" * level, re.escape(heading), level) + match = re.search(pattern, document, flags=re.MULTILINE | re.DOTALL) + if not match: + raise EvidenceError("missing {} heading".format(heading)) + return match.group(1) + + +def _markdown_table(section, heading): + """Parse a named Markdown table into dictionaries keyed by its headers.""" + body = _markdown_section(section, heading, 3) + rows = [ + [cell.strip() for cell in line.strip().strip("|").split("|")] + for line in body.splitlines() + if line.strip().startswith("|") + ] + if len(rows) < 3: + raise EvidenceError("{} must contain a data table".format(heading)) + headers = rows[0] + if not all(headers): + raise EvidenceError("{} has an invalid table header".format(heading)) + data_rows = [] + for row in rows[2:]: + if len(row) != len(headers) or any(not cell for cell in row): + raise EvidenceError("{} has an incomplete table row".format(heading)) + data_rows.append(dict(zip(headers, row))) + if not data_rows: + raise EvidenceError("{} must contain at least one row".format(heading)) + return data_rows + + +def _markdown_table_rows(section, heading, level=2): + """Parse a Markdown table while allowing empty informational cells.""" + body = _markdown_section(section, heading, level) + rows = [ + [cell.strip() for cell in line.strip().strip("|").split("|")] + for line in body.splitlines() + if line.strip().startswith("|") + ] + if len(rows) < 3: + raise EvidenceError("{} must contain a data table".format(heading)) + headers = rows[0] + if not all(headers): + raise EvidenceError("{} has an invalid table header".format(heading)) + values = [] + for row in rows[2:]: + if len(row) != len(headers): + raise EvidenceError("{} has an incomplete table row".format(heading)) + values.append(dict(zip(headers, row))) + if not values: + raise EvidenceError("{} must contain at least one row".format(heading)) + return headers, values + + +def parse_security_matrix(stream_name, matrix_path, content): + """Parse a stream matrix into schema rows and trusted read instructions.""" + if not isinstance(content, str) or not content: + raise EvidenceError("security matrix content is required") + headers, rows = _markdown_table_rows(content, "Supportability Matrix") + version_header = next((header for header in headers if "version" in header.lower()), None) + if not version_header: + raise EvidenceError("Supportability Matrix has no version column") + ignored = {version_header, "Build", "Build Date", "Notes"} + source_headers = [header for header in headers if header not in ignored] + if not source_headers: + raise EvidenceError("Supportability Matrix has no source commit columns") + + matrix_rows = [] + for row in rows: + version = row[version_header] + source_commits = { + header: _ref_token(row[header]) + for header in source_headers + if row[header].strip("`") + } + retag = re.search(r"\bretag of\s+([^\s|]+)", row.get("Notes", ""), re.IGNORECASE) + matrix_rows.append({ + "version": version, + "source_commits": source_commits, + "retag_of": retag.group(1) if retag else None, + }) + + _, ecosystem_rows = _markdown_table_rows(content, "Ecosystem Mappings") + mappings = [] + for row in ecosystem_rows: + try: + mappings.append({ + "ecosystem": row["Ecosystem"], + "repository": row["Repository"], + "lock_file": row["Lock File"].strip("`"), + "check_command": row["Check Command"].strip("`"), + "upstream_branch": _ref_token(row["Upstream Branch"]), + }) + except KeyError as error: + raise EvidenceError("Ecosystem Mappings is missing {}".format(error.args[0])) from error + return { + "name": stream_name, + "matrix_source": content, + "rows": matrix_rows, + }, mappings + + +def parse_security_configuration(claude_md): + """Parse target-project Jira and Security Configuration into schema fields.""" + if not isinstance(claude_md, str): + raise EvidenceError("CLAUDE.md content must be text") + jira = _markdown_section(claude_md, "Jira Configuration", 2) + security = _markdown_section(claude_md, "Security Configuration", 2) + lifecycle = _markdown_section(security, "Product Lifecycle", 3) + values = { + key: value.strip() + for key, value in re.findall(r"^- ([^:]+):\s*(.+?)\s*$", lifecycle, re.MULTILINE) + if not value.strip().startswith("{{") + } + project_key = re.search(r"^- Project key:\s*(\S+)\s*$", jira, re.MULTILINE) + if not project_key: + raise EvidenceError("Jira Configuration is missing Project key") + + required = { + "Product pages URL": "product lifecycle URL", + "Jira version prefix": "Jira version prefix", + "Vulnerability issue type ID": "Vulnerability issue type ID", + "Component label pattern": "Component label pattern", + } + for field, label in required.items(): + if not values.get(field): + raise EvidenceError("Security Configuration is missing {}".format(label)) + + streams = _markdown_table(security, "Version Streams") + # Parse Source Repositories with the blank-cell-tolerant reader so an empty + # (present but blank) Deployment Context cell can fall back to "upstream" + # rather than being rejected as an incomplete row before the default applies. + _, sources = _markdown_table_rows(security, "Source Repositories", level=3) + try: + version_streams = [{ + "name": row["Stream"], + "matrix_path": row["Security Matrix Path"], + "release_repository": row["Konflux Release Repo"], + } for row in streams] + source_repositories = [] + for row in sources: + name = row["Repository"] + url = row["URL"] + if not name or not url: + raise EvidenceError("Source Repositories row is missing Repository or URL") + source_repositories.append({ + "name": name, + "url": url, + "deployment_context": row.get("Deployment Context") or "upstream", + }) + except KeyError as error: + raise EvidenceError("Security Configuration table is missing {}".format(error.args[0])) from error + + configuration = { + "project_key": project_key.group(1), + "jira_version_prefix": values["Jira version prefix"], + "vulnerability_issue_type_id": values["Vulnerability issue type ID"], + "component_label_pattern": values["Component label pattern"], + "version_streams": version_streams, + "source_repositories": source_repositories, + } + optional_fields = { + "VEX Justification custom field": "vex_justification_field", + "Upstream Affected Component custom field": "upstream_affected_component_field", + "PS Component custom field": "ps_component_field", + "Stream custom field": "stream_field", + "ProdSec Jira account ID": "prodsec_account_id", + "Embargo policy URL": "embargo_policy_url", + } + for source, destination in optional_fields.items(): + if values.get(source): + configuration[destination] = values[source] + return configuration + + +def _runner_configuration(claude_md, project_root): + """Extract runner-only paths and URLs alongside the sandbox configuration.""" + configuration = parse_security_configuration(claude_md) + security = _markdown_section(claude_md, "Security Configuration", 2) + lifecycle = _markdown_section(security, "Product Lifecycle", 3) + lifecycle_url = re.search(r"^- Product pages URL:\s*(\S+)\s*$", lifecycle, re.MULTILINE) + if not lifecycle_url: + raise EvidenceError("Security Configuration is missing Product pages URL") + stream_rows = _markdown_table(security, "Version Streams") + _, registry_rows = _markdown_table_rows(claude_md, "Repository Registry", level=2) + paths = { + row["Repository"]: Path(row["Path"]) + for row in registry_rows + if "Repository" in row and "Path" in row + } + root = Path(project_root) + return configuration, lifecycle_url.group(1), stream_rows, { + name: path if path.is_absolute() else root / path + for name, path in paths.items() + } + + +def extract_cve_id(issue): + """Return the required CVE identifier from an issue's labels or summary.""" + issue = _require_mapping(issue, "issue") + fields = _require_mapping(issue.get("fields"), "issue.fields") + values = list(_require_list(fields.get("labels"), "issue.fields.labels", allow_empty=True)) + values.append(fields.get("summary", "")) + for value in values: + match = re.search(r"\bCVE-\d{4}-\d{4,}\b", value or "", re.IGNORECASE) + if match: + return match.group(0).upper() + raise EvidenceError("issue does not contain a CVE identifier") + + +def _require_mapping(value, name): + """Return a required mapping or raise a descriptive evidence error.""" + if not isinstance(value, dict): + raise EvidenceError("{} must be an object".format(name)) + return value + + +def _require_list(value, name, allow_empty=False): + """Return a required list or raise a descriptive evidence error.""" + if not isinstance(value, list) or (not allow_empty and not value): + qualifier = "a non-empty list" if not allow_empty else "a list" + raise EvidenceError("{} must be {}".format(name, qualifier)) + return value + + +def normalize_issue(issue): + """Extract the schema-required audit fields from a full Jira issue response.""" + issue = _require_mapping(issue, "issue") + key = issue.get("key", "") + if not isinstance(key, str) or not _ISSUE_KEY_RE.fullmatch(key): + raise EvidenceError("issue.key must be a Jira issue key") + + fields = _require_mapping(issue.get("fields"), "issue.fields") + status = _require_mapping(fields.get("status"), "issue.fields.status") + reporter = _require_mapping(fields.get("reporter"), "issue.fields.reporter") + comments = _require_mapping(fields.get("comment"), "issue.fields.comment") + description = fields.get("description") + if not isinstance(description, dict): + raise EvidenceError("issue.fields.description must be an ADF object") + + summary = fields.get("summary") + if not isinstance(summary, str) or not summary: + raise EvidenceError("issue.fields.summary is required") + status_name = status.get("name") + if not isinstance(status_name, str) or not status_name: + raise EvidenceError("issue.fields.status.name is required") + + account_id = reporter.get("accountId") + display_name = reporter.get("displayName") + if not isinstance(account_id, str) or not account_id: + raise EvidenceError("issue.fields.reporter.accountId is required") + if not isinstance(display_name, str) or not display_name: + raise EvidenceError("issue.fields.reporter.displayName is required") + + return { + "key": key, + "summary": summary, + "description": description, + "status": status_name, + "labels": _require_list(fields.get("labels"), "issue.fields.labels", allow_empty=True), + "versions": normalize_versions(fields.get("versions")), + "reporter": {"account_id": account_id, "display_name": display_name}, + "comments": _require_list(comments.get("comments"), "issue.fields.comment.comments", allow_empty=True), + "fields": fields, + } + + +def normalize_versions(versions): + """Reduce Jira version objects to the schema's stable audit fields.""" + versions = _require_list(versions, "versions", allow_empty=True) + normalized = [] + for index, version in enumerate(versions): + version = _require_mapping(version, "versions[{}]".format(index)) + for field in ("id", "name", "released"): + if field not in version: + raise EvidenceError("versions[{}].{} is required".format(index, field)) + entry = { + "id": str(version["id"]), + "name": version["name"], + "released": version["released"], + } + if "archived" in version: + entry["archived"] = version["archived"] + normalized.append(entry) + return normalized + + +def normalize_remote_links(remote_links): + """Convert Jira remote links into the narrow sandbox link contract.""" + links = _require_list(remote_links, "remote_links") + normalized = [] + for index, link in enumerate(links): + link = _require_mapping(link, "remote_links[{}]".format(index)) + source = link.get("object", link) + source = _require_mapping(source, "remote_links[{}].object".format(index)) + url = source.get("url") + title = source.get("title") + if not isinstance(url, str) or not url: + raise EvidenceError("remote_links[{}] is missing object.url".format(index)) + if not isinstance(title, str) or not title: + raise EvidenceError("remote_links[{}] is missing object.title".format(index)) + normalized.append({"url": url, "title": title}) + return normalized + + +def _validate_matrix(matrix): + """Reject a matrix that cannot support deterministic version analysis.""" + matrix = _require_mapping(matrix, "matrix") + for stream_index, stream in enumerate(_require_list(matrix.get("streams"), "matrix.streams")): + stream = _require_mapping(stream, "matrix.streams[{}]".format(stream_index)) + rows = _require_list(stream.get("rows"), "matrix.streams[{}].rows".format(stream_index)) + for row_index, row in enumerate(rows): + row = _require_mapping(row, "matrix.streams[{}].rows[{}]".format(stream_index, row_index)) + if not row.get("version") or not isinstance(row.get("source_commits"), dict): + raise EvidenceError("matrix row {}:{} is malformed".format(stream_index, row_index)) + if row.get("retag_of") is None: + if not row["source_commits"]: + raise EvidenceError("matrix row {}:{} has no source commits".format(stream_index, row_index)) + return matrix + + +def _validate_source_evidence(source_evidence): + """Reject missing lock-file or development-stream runner evidence.""" + source_evidence = _require_mapping(source_evidence, "source_evidence") + for name in ("lock_files", "development_streams"): + reads = _require_list(source_evidence.get(name), "source_evidence.{}".format(name)) + for index, read in enumerate(reads): + read = _require_mapping(read, "source_evidence.{}[{}]".format(name, index)) + for field in ("repository", "ref", "path", "command", "content"): + if not isinstance(read.get(field), str) or not read[field]: + raise EvidenceError("source_evidence.{}[{}].{} is required".format(name, index, field)) + if not read["command"].startswith("git show "): + raise EvidenceError("source evidence commands must use git show") + return source_evidence + + +# Schema formats whose constraints must be enforced on every bundle. jsonschema +# only checks a "format" when its backing validation library is installed (the +# jsonschema[format] extra: rfc3987 for uri, rfc3339-validator for date-time). +# Without the extra, FormatChecker silently treats every value as conforming, so +# these constraints would be skipped and malformed URLs/timestamps could enter a +# supposedly schema-validated bundle. +_REQUIRED_FORMATS = ("uri", "date-time") + + +def _format_checker(): + """Return a FormatChecker, failing closed if required formats are inactive. + + A trusted runner missing the jsonschema[format] extra must abort rather than + emit a bundle whose uri/date-time constraints went unenforced. + """ + checker = FormatChecker() + missing = [fmt for fmt in _REQUIRED_FORMATS if fmt not in checker.checkers] + if missing: + raise EvidenceError( + "jsonschema format validation is unavailable for {}; " + "install jsonschema[format] on the runner".format(", ".join(missing))) + return checker + + +def validate_bundle(bundle, schema_path=None): + """Validate a completed bundle against the sandbox's exact JSON schema.""" + path = Path(schema_path or os.environ.get("FULLSEND_INPUT_SCHEMA", _SCHEMA_PATH)) + format_checker = _format_checker() + try: + with path.open() as schema_file: + schema = json.load(schema_file) + validate(instance=bundle, schema=schema, format_checker=format_checker) + except (OSError, json.JSONDecodeError, ValidationError) as error: + raise EvidenceError("triage-security input validation failed: {}".format(error)) from error + + +def build_bundle(issue, remote_links, configuration, external_evidence, matrix, + source_evidence, jira_metadata, idempotency, + mutation_authorized): + """Build a schema-validated, credential-free triage-security input bundle.""" + configuration = _require_mapping(configuration, "configuration") + external_evidence = _require_mapping(external_evidence, "external_evidence") + jira_metadata = _require_mapping(jira_metadata, "jira_metadata") + idempotency = _require_mapping(idempotency, "idempotency") + + bundle = { + "schema_version": "1", + "issue": normalize_issue(issue), + "remote_links": normalize_remote_links(remote_links), + "configuration": configuration, + "external_evidence": external_evidence, + "matrix": _validate_matrix(matrix), + "source_evidence": _validate_source_evidence(source_evidence), + "jira_metadata": jira_metadata, + "idempotency": idempotency, + "authorization": {"mutation_authorized": bool(mutation_authorized)}, + } + validate_bundle(bundle) + return bundle + + +def _jira_client(command, *arguments): + """Run the existing Jira client and decode its JSON response.""" + client = Path(__file__).with_name("jira-client.py") + # List/argv form with shell=False: arguments reach execve as separate argv + # entries, so no shell parses them and shell metacharacters cannot inject a + # command. The program is fixed (sys.executable running the repo-local + # jira-client.py); command is an internal literal and *arguments are internal + # literals plus a validated issue key and _jql_escape'd JQL. Not injectable; + # verified false positive (Sourcery agreed on PR #311). + result = subprocess.run( # nosemgrep: python.lang.security.audit.dangerous-subprocess-use-audit + [sys.executable, str(client), command, *arguments], + check=True, + capture_output=True, + text=True, + ) + try: + return json.loads(result.stdout) + except json.JSONDecodeError as error: + raise EvidenceError("jira-client returned invalid JSON for {}".format(command)) from error + + +def _fetch_url(url, allow_missing=False): + """Fetch required external evidence and retain its retrieval provenance. + + When ``allow_missing`` is set, a 404 is treated as legitimate evidence (the + CVE is not tracked by this source, e.g. an OSV gap or a reserved/embargoed + MITRE record) and recorded with its status and body instead of aborting the + whole bundle. Every other non-2xx status still fails loudly. + """ + # access.redhat.com (and some other evidence hosts) sit behind bot + # protection that rejects the default "Python-urllib/x.y" User-Agent with + # HTTP 403. Send a browser-style UA so required evidence is retrievable. + request = Request(url, headers={"User-Agent": _EVIDENCE_USER_AGENT}) + try: + with urlopen(request, timeout=30) as response: + body = response.read().decode("utf-8") + status = response.status + except HTTPError as error: + body = error.read().decode("utf-8", errors="replace") + status = error.code + except URLError as error: + raise EvidenceError("could not retrieve {}: {}".format(url, error.reason)) from error + if status < 200 or status >= 300: + if not (allow_missing and status == 404): + raise EvidenceError("required evidence {} returned HTTP {}".format(url, status)) + try: + body = json.loads(body) + except json.JSONDecodeError: + pass + return { + "source_url": url, + "retrieved_at": datetime.now(timezone.utc).isoformat().replace("+00:00", "Z"), + "status": status, + "body": body, + } + + +def _resolve_ref(repository_path, ref): + """Resolve a matrix ref to a committish git can read, read-only. + + Development-branch refs (e.g. ``release/0.5.z``) resolve by bare name only + when a local branch or tag of that exact name exists. After a plain + ``git fetch`` a clone usually has just the remote-tracking copy + (``refs/remotes//release/0.5.z``), which the bare name will not + match. Fall back to that copy without mutating the repository, preferring a + canonical remote (upstream, then origin) so the same branch name carried by + two remotes never resolves ambiguously. Returns the committish to read, or + ``None`` when nothing matches (e.g. a genuinely missing commit SHA). + """ + verifies = subprocess.run( + ["git", "-C", str(repository_path), "rev-parse", "--verify", "--quiet", + "{}^{{commit}}".format(ref)], + capture_output=True, text=True, + ) + if verifies.returncode == 0: + return ref + + listing = subprocess.run( + ["git", "-C", str(repository_path), "for-each-ref", "--format=%(refname)", + "refs/remotes/"], + capture_output=True, text=True, + ) + matches = [] + for refname in listing.stdout.split(): + rest = refname[len("refs/remotes/"):] + remote, _, tail = rest.partition("/") + if tail == ref: # exact single-remote match, not a slashed branch suffix + matches.append((remote, refname)) + if not matches: + return None + preferred = {"upstream": 0, "origin": 1} + matches.sort(key=lambda m: (preferred.get(m[0], 2), m[0])) + return matches[0][1] + + +def _git_show(repository_path, ref, path): + """Read source evidence with the permitted read-only git show command.""" + resolved = _resolve_ref(repository_path, ref) + if resolved is None: + raise EvidenceError( + "git show failed for {}:{}: no local or remote-tracking ref resolves " + "{!r} in {}".format(ref, path, ref, repository_path) + ) + try: + result = subprocess.run( + ["git", "-C", str(repository_path), "show", "{}:{}".format(resolved, path)], + check=True, + capture_output=True, + text=True, + ) + except subprocess.CalledProcessError as error: + detail = error.stderr.strip() or error.stdout.strip() + raise EvidenceError("git show failed for {}:{}: {}".format(ref, path, detail)) from error + return { + "ref": ref, + "path": path, + "command": "git show {}:{}".format(resolved, path), + "content": result.stdout, + } + + +def _related_issue(issue): + """Normalize a fetched related Jira issue for audit and idempotency checks.""" + fields = _require_mapping(issue.get("fields"), "related issue fields") + status = _require_mapping(fields.get("status"), "related issue status") + return { + "key": issue.get("key", ""), + "summary": fields.get("summary", ""), + "status": status.get("name", ""), + "labels": fields.get("labels", []) or [], + "description": fields.get("description", {}) or {}, + "comments": (fields.get("comment") or {}).get("comments", []) or [], + "links": fields.get("issuelinks", []) or [], + } + + +def _related_keys(issue): + """Return each directly related Jira key once, in stable order.""" + fields = _require_mapping(issue.get("fields"), "issue.fields") + keys = {entry.get("key") for entry in fields.get("subtasks", []) if entry.get("key")} + for link in fields.get("issuelinks", []) or []: + related = link.get("inwardIssue") or link.get("outwardIssue") or {} + if related.get("key"): + keys.add(related["key"]) + return sorted(keys) + + +def _action_markers(issue): + """Extract stable triage action markers from the current issue's comments.""" + markers = [] + for comment in (issue.get("fields", {}).get("comment") or {}).get("comments", []) or []: + text = json.dumps(comment.get("body", {})) + markers.extend(re.findall(r"triage-security:[A-Za-z0-9_-]+", text)) + return sorted(set(markers)) + + +def _jql_escape(value): + """Escape a value embedded in a quoted JQL literal.""" + return value.replace("\\", "\\\\").replace('"', '\\"') + + +def collect_bundle(issue_key, project_root): + """Collect all credentialed evidence for one poller-dispatched security issue.""" + # Validate the key before it is interpolated into any JQL. A conforming key + # carries no JQL metacharacters, so this is a self-contained injection defense + # that does not rely on the pre-triage-security.sh regex gate upstream of us. + if not isinstance(issue_key, str) or not _ISSUE_KEY_RE.fullmatch(issue_key): + raise EvidenceError("issue_key must be a Jira issue key") + root = Path(project_root) + claude_path = root / "CLAUDE.md" + if not claude_path.is_file(): + raise EvidenceError("target project CLAUDE.md is required") + configuration, lifecycle_url, stream_rows, repository_paths = _runner_configuration( + claude_path.read_text(), root) + issue = _jira_client("get_issue", issue_key, "--fields", "*all") + cve_id = extract_cve_id(issue) + escaped_project_key = _jql_escape(configuration["project_key"]) + escaped_cve_id = _jql_escape(cve_id) + remote_links = _jira_client("get_remote_links", issue_key) + versions = _jira_client("get_versions", configuration["project_key"]) + + sibling_jql = ( + 'project = "{}" AND labels = "{}" AND issuetype = {} AND key != "{}"' + .format(escaped_project_key, escaped_cve_id, configuration["vulnerability_issue_type_id"], issue_key) + ) + searches = [("same-cve-siblings", sibling_jql)] + component_field = configuration.get("upstream_affected_component_field") + component = issue.get("fields", {}).get(component_field, "") if component_field else "" + component_field_number = re.search(r"(\d+)$", component_field or "") + if isinstance(component, str) and component and component_field_number: + overlap_jql = ( + 'project = "{}" AND issuetype = {} AND cf[{}] ~ "{}" AND key != "{}"' + .format( + escaped_project_key, + configuration["vulnerability_issue_type_id"], + component_field_number.group(1), + _jql_escape(component), + issue_key, + ) + ) + searches.append(("cross-cve-overlap", overlap_jql)) + preemptive_jql = ( + 'project = "{}" AND issuetype = Task AND labels = "security-preemptive" ' + 'AND labels = "{}" ORDER BY created DESC' + .format(escaped_project_key, escaped_cve_id) + ) + searches.append(("preemptive-remediation", preemptive_jql)) + + fetched = {} + search_results = [] + for purpose, jql in searches: + search = _jira_client("search_jql", "--jql", jql, "--fields", "*all", "--all") + items = search.get("issues", []) + for item in items: + if item.get("key"): + fetched[item["key"]] = item + search_results.append((purpose, jql, items)) + for key in _related_keys(issue): + fetched.setdefault(key, _jira_client("get_issue", key, "--fields", "*all")) + related = [_related_issue(fetched[key]) for key in sorted(fetched)] + + external_evidence = { + "mitre": _fetch_url("https://cveawg.mitre.org/api/cve/{}".format(cve_id), allow_missing=True), + "osv": _fetch_url("https://api.osv.dev/v1/vulns/{}".format(cve_id), allow_missing=True), + "lifecycle": _fetch_url(lifecycle_url), + } + + matrix_streams = [] + lock_files = [] + development_streams = [] + for stream_row in stream_rows: + try: + stream_name = stream_row["Stream"] + matrix_path = stream_row["Security Matrix Path"] + local_release = Path(stream_row["Local Path"]) + except KeyError as error: + raise EvidenceError( + "Version Streams table is missing {}".format(error.args[0])) from error + matrix_file = root / matrix_path + if matrix_file.is_file(): + matrix_content = matrix_file.read_text() + else: + matrix_content = _git_show(local_release, "main", matrix_path)["content"] + matrix, mappings = parse_security_matrix(stream_name, matrix_path, matrix_content) + matrix_streams.append(matrix) + for mapping in mappings: + repository = mapping["repository"] + if repository not in repository_paths: + raise EvidenceError("Repository Registry has no path for {}".format(repository)) + for row in matrix["rows"]: + ref = row["source_commits"].get(repository) + if ref: + read = _git_show(repository_paths[repository], ref, mapping["lock_file"]) + lock_files.append({"repository": repository, **read}) + read = _git_show(repository_paths[repository], mapping["upstream_branch"], mapping["lock_file"]) + development_streams.append({"repository": repository, **read}) + + jira_metadata = { + "versions": normalize_versions(versions), + "sibling_searches": [{ + "purpose": purpose, + "jql": jql, + "issues": [_related_issue(item) for item in items], + } for purpose, jql, items in search_results], + "related_issues": related, + } + return build_bundle( + issue=issue, + remote_links=remote_links, + configuration=configuration, + external_evidence=external_evidence, + matrix={"streams": matrix_streams}, + source_evidence={"lock_files": lock_files, "development_streams": development_streams}, + jira_metadata=jira_metadata, + idempotency={ + "action_markers": _action_markers(issue), + "existing_remediation": related, + }, + mutation_authorized=False, + ) + + +def main(argv=None): + """Run a requested deterministic transform or trusted collection command.""" + parser = argparse.ArgumentParser(prog="pre_triage_security.py") + subparsers = parser.add_subparsers(dest="command", required=True) + collect = subparsers.add_parser("collect") + collect.add_argument("issue_key") + collect.add_argument("project_root") + config = subparsers.add_parser("parse-configuration") + config.add_argument("claude_md") + args = parser.parse_args(argv) + try: + if args.command == "collect": + result = collect_bundle(args.issue_key, args.project_root) + else: + result = parse_security_configuration(Path(args.claude_md).read_text()) + except (EvidenceError, subprocess.CalledProcessError) as error: + print("ERROR: {}".format(error), file=sys.stderr) + return 1 + json.dump(result, sys.stdout, indent=2) + sys.stdout.write("\n") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/plugins/sdlc-workflow/scripts/pre_verify_pr.py b/plugins/sdlc-workflow/scripts/pre_verify_pr.py new file mode 100644 index 000000000..9d1bcc343 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/pre_verify_pr.py @@ -0,0 +1,550 @@ +#!/usr/bin/env python3 +"""Pre-verify-pr data extraction functions. + +Extracts PR URL from Jira custom fields, assembles the GitHub tier-1 read +bundle prefetched on the runner, and transforms Jira issue JSON into the +tracker-agnostic input schema used by the sandbox. + +CLI usage (called by pre-verify-pr.sh): + echo "$ISSUE_JSON" | python3 pre_verify_pr.py extract-pr-url + echo "$ISSUE_JSON" | python3 pre_verify_pr.py related-keys + echo "$ISSUE_JSON" | python3 pre_verify_pr.py transform TASK_ID PR_URL \\ + [--github-dir DIR --pr-repo REPO --pr-number N \\ + --head-ref REF --commit-sha SHA --idempotency-dir DIR] + +When the --github-* options are supplied, `transform` reads the raw GitHub +reads from DIR (pr.diff, pr.stat, reviews.json, review-comments.json, +issue-comments.json, commits.json, check-runs.json, check-run-logs.txt) and +embeds them under a `github` key. check-run-logs.txt is not inlined — only its +mounted sandbox path is embedded (empty when no check failed), so the large log +text stays off the input bundle and off the sub-agent's context until read. + +`related-keys` prints the task's sub-task and linked-issue keys (one per line) +so the shell can prefetch each on the runner. When --idempotency-dir is given, +`transform` reads those prefetched issue JSONs and embeds their summary/labels/ +description/comments under an `idempotency` key, giving the sandbox a tokenless +data source for the Steps 6d/6f/7c idempotency checks. +""" + +import argparse +import glob +import json +import os +import re +import sys + + +def commit_references_task(commit, task_id): + """True if the commit references ``task_id`` in its headline OR body. + + The Jira task ID conventionally sits in a trailer or body line (e.g. + ``Implements TC-5812``), not only the subject, so BOTH ``messageHeadline`` + and ``messageBody`` are scanned. Word boundaries keep ``TC-5812`` from + matching ``TC-58120`` or another ID like ``TC-5982``. This computes the + Commit Traceability fact deterministically on the runner so the tokenless + sandbox agent never has to re-derive it from ``git log`` (which, with + ``--oneline``/``%s``, would see subjects only and miss the trailer). + """ + if not task_id: + return False + text = "{}\n{}".format( + commit.get("messageHeadline", "") or "", + commit.get("messageBody", "") or "", + ) + return re.search(r"\b{}\b".format(re.escape(task_id)), text) is not None + + +_URL_RE = re.compile(r"https?://\S+") + + +def _first_url_in_adf(node): + """Return the first URL found in an ADF node, depth-first in document order. + + Covers every shape Jira uses to populate a URL/textarea custom field: + an ``inlineCard`` (smart link), a ``text`` node carrying a ``link`` mark, + or plain text that merely contains a bare URL. Returns "" when none is found. + """ + if isinstance(node, dict): + if node.get("type") == "inlineCard": + url = node.get("attrs", {}).get("url", "") + if url: + return url + if node.get("type") == "text": + for mark in node.get("marks", []): + if mark.get("type") == "link": + href = mark.get("attrs", {}).get("href", "") + if href: + return href + match = _URL_RE.search(node.get("text", "") or "") + if match: + return match.group(0).rstrip(".,);]") + for child in node.get("content", []): + url = _first_url_in_adf(child) + if url: + return url + elif isinstance(node, list): + for child in node: + url = _first_url_in_adf(child) + if url: + return url + return "" + + +def extract_pr_url(issue): + """Extract the PR URL from the Jira Git Pull Request custom field. + + The field is format-agnostic: it may be a plain string, an ADF + ``inlineCard`` smart link, an ADF ``text`` node with a ``link`` mark, or + plain ADF text holding a bare URL. All are handled so a linked PR is never + missed because of how the field happened to be populated. Returns "" when + the field is absent or contains no URL. + """ + field = issue.get("fields", {}).get("customfield_10875") + if not field: + return "" + if isinstance(field, str): + match = _URL_RE.search(field) + return match.group(0).rstrip(".,);]") if match else field.strip() + if isinstance(field, dict): + return _first_url_in_adf(field) + return "" + + +# Gates the verification: the JQL-resolved issue must be in this status AND +# carry this label, mirroring the interactive verify-pr entry conditions. Kept +# as module constants so the skip reasons and the tests reference one source. +GATE_STATUS = "Review" +GATE_LABEL = "ai-generated-jira" + + +def build_pr_jql(pr_url): + """JQL recalling issues whose Git Pull Request field references ``pr_url``. + + Uses the ``~`` (contains) operator on the custom field by id (``cf[10875]``): + an exact ``=`` match is unreliable for Jira URL/text fields, which tokenize + their stored value. Recall is intentionally broad — ``resolve_gated_issue`` + re-confirms the *exact* PR URL in Python, so a fuzzy ``~`` hit on a different + PR can never be accepted. The value is wrapped in a quoted JQL string literal + with backslashes and quotes escaped so a crafted URL cannot break out of it. + """ + escaped = pr_url.replace("\\", "\\\\").replace('"', '\\"') + return 'cf[10875] ~ "{}"'.format(escaped) + + +def resolve_gated_issue(search_result, pr_url): + """Resolve the single Jira issue that gates verification of ``pr_url``. + + Returns ``(key, None)`` when exactly one issue's Git Pull Request field + *exactly* equals ``pr_url`` AND that issue is in status ``Review`` AND + carries the ``ai-generated-jira`` label. Returns ``(None, reason)`` when any + gate fails, with a human-readable ``reason`` for the ADR-0072 skip signal: + + - no issue references ``pr_url`` (the broad ``~`` recall matched nothing exact) + - more than one issue references ``pr_url`` (ambiguous — refuse to guess) + - the matched issue is not in status ``Review`` + - the matched issue lacks the ``ai-generated-jira`` label + + The exact-URL re-match (``extract_pr_url`` per candidate) is the trust + anchor: the JQL ``~`` operator over-matches, so acceptance is decided here, + never by JQL alone. + """ + issues = search_result.get("issues", []) if isinstance(search_result, dict) else [] + matches = [issue for issue in issues if extract_pr_url(issue) == pr_url] + if not matches: + return None, ( + "no Jira issue links PR {} in its Git Pull Request field".format(pr_url) + ) + if len(matches) > 1: + keys = ", ".join(sorted(issue.get("key", "?") for issue in matches)) + return None, ( + "multiple Jira issues link PR {} ({}) — refusing to guess".format( + pr_url, keys) + ) + issue = matches[0] + key = issue.get("key", "") + fields = issue.get("fields", {}) + status = (fields.get("status") or {}).get("name", "") + if status != GATE_STATUS: + return None, "{} is in status '{}', not '{}'".format( + key, status, GATE_STATUS) + labels = fields.get("labels", []) or [] + if GATE_LABEL not in labels: + return None, "{} is missing the '{}' label".format(key, GATE_LABEL) + return key, None + + +def revalidate_gate(issue, pr_url): + """Re-apply the gate to the FULL issue actually used to build the input. + + ``resolve_gated_issue`` gates the lightweight JQL *search* response + (status/labels/customfield_10875 only). The full issue that populates + verify-pr-input.json is then fetched in a SECOND request — between the two, + the issue can leave status ``Review``, lose the ``ai-generated-jira`` label, + or have its Git Pull Request field changed, so a stale-but-successful gate + could hand the sandbox an issue that no longer qualifies (a TOCTOU gap). + + Re-run the exact-URL + status + label gate on the full issue here, + immediately before the write. Reuses ``resolve_gated_issue`` (wrapping the + single issue as a one-element search result) so the acceptance rule and skip + reasons stay defined in exactly one place. Returns ``(key, None)`` when the + full issue still exactly links ``pr_url`` AND is in status ``Review`` AND + carries the label; ``(None, reason)`` on any failure (ADR-0072 skip). + """ + return resolve_gated_issue({"issues": [issue]}, pr_url) + + +# Sandbox-side path of the concatenated failed-check logs. Mirrors the +# host_files dest in .fullsend/harness/verify-pr.yaml (the runner writes +# check-run-logs.txt next to verify-pr-input.json and mounts it here). Embedded +# in the bundle only when a check failed, so the sub-agent Reads it on demand. +SANDBOX_CHECK_RUN_LOGS_PATH = "/sandbox/workspace/.pre-script/check-run-logs.txt" + + +def filter_own_check_runs(check_runs, own_names): + """Drop verify-pr's own workflow check-runs from the prefetched list. + + ``own_names`` is the set of check-run names produced by verify-pr's own + workflow (``fullsend verify-pr``) at the PR head SHA, gathered on the trusted + runner across EVERY run of that workflow at the SHA so both the in-progress + dispatch and any superseded prior attempt are covered. Those own runs are + non-terminal (or a superseded attempt failed) at evaluation time, so + correctness.md Check 1a maps them to pending/failed and drags CI Status to a + permanent self-referential WARN/FAIL. Removing them here — the CI-Status + analogue of Step 1's ``running-workflow-name: wait-for-checks`` self-exclusion + — lets a PR whose substantive checks all pass report CI Status = PASS. + + Matches on the check-run ``name`` (an entry with no ``name`` is kept, since a + substantive check is never nameless in practice). Returns a NEW list; the + input is not mutated. An empty ``own_names`` returns the input unchanged, so a + run that could not enumerate its own check-runs simply keeps prior behavior. + """ + names = set(own_names) + return [c for c in check_runs if c.get("name") not in names] + + +def build_github_bundle(pr_repo, pr_number, head_ref, commit_sha, + diff, stat, reviews, review_comments, + issue_comments, commits, check_runs=None, + check_run_logs_path=""): + """Assemble the GitHub tier-1 read bundle embedded in the input. + + diff/stat are raw text; the reviews/review_comments/issue_comments/commits/ + check_runs arguments are already-parsed JSON (lists). ``check_runs`` carries + the head-SHA CI check-run outcomes (name/status/conclusion/details_url) so the + Correctness sub-agent's CI Status check reads real CI data in the tokenless + sandbox instead of shelling out to `gh`; it defaults to an empty list (a PR + with no checks) so every bundle carries the key. ``check_run_logs_path`` is + the mounted sandbox path of the concatenated failed-check logs (correctness.md + Check 1b reads it only on a FAIL); it is "" when no check failed, so the log + text never enters the bundle or the sub-agent's context on the green path. + Keys mirror the reads the verify-pr skill performs so the sandbox needs no + api.github.com egress. + """ + return { + "pr_repo": pr_repo, + "pr_number": int(pr_number), + "headRefName": head_ref, + "commit_sha": commit_sha, + "diff": diff, + "stat": stat, + "reviews": reviews, + "review_comments": review_comments, + "issue_comments": issue_comments, + "commits": commits, + "check_runs": check_runs if check_runs is not None else [], + "check_run_logs_path": check_run_logs_path, + } + + +def related_keys(issue): + """Keys of the task's sub-tasks and linked issues (idempotency targets). + + Steps 6d/6f/7c dedupe against the parent task's existing sub-tasks and its + linked (e.g., root-cause) issues. The pre_script fetches each of these keys + on the trusted runner so the sandbox can run those idempotency checks + without a Jira token. Returns a sorted, de-duplicated list. + """ + fields = issue.get("fields", {}) + keys = set() + for sub in fields.get("subtasks") or []: + key = sub.get("key") + if key: + keys.add(key) + for link in fields.get("issuelinks") or []: + related = link.get("inwardIssue") or link.get("outwardIssue") or {} + key = related.get("key") + if key: + keys.add(key) + return sorted(keys) + + +def build_idempotency_bundle(related_issue_jsons): + """Bundle related-issue metadata for the sandbox idempotency checks. + + Each item is a full issue JSON the runner fetched with fields summary, + labels, description, issuetype, and comment. The sandbox reads this instead + of calling Jira to dedupe sub-tasks (Steps 6d/6f) and root-cause tasks + (Step 7c). Descriptions and comment bodies are kept in the tracker's native + format (ADF for Jira) — the sandbox agent inspects them directly. + """ + related = [] + for ri in related_issue_jsons: + fields = ri.get("fields", {}) + comments = [ + c.get("body") or {} + for c in (fields.get("comment") or {}).get("comments", []) + ] + related.append({ + "key": ri.get("key", ""), + "summary": fields.get("summary", ""), + "labels": fields.get("labels", []), + # Coerce an explicit null description to {} so the value stays an + # object per verify-pr-input.schema.json (see transform_to_input). + "description": fields.get("description") or {}, + "issuetype": (fields.get("issuetype") or {}).get("name", ""), + "comments": comments, + }) + return {"related_issues": related} + + +def transform_to_input(issue, task_id, pr_url, github=None, idempotency=None): + """Transform Jira issue JSON to tracker-agnostic input schema. + + The ``idempotency`` bundle is always emitted (defaulting to an empty + ``related_issues`` list when none is supplied) so the tokenless sandbox + always has a data source for the Steps 6d/6f/7c dedup checks. This matches + verify-pr-input.schema.json, which requires ``idempotency.related_issues``: + a schema-valid prefetch can never omit the bundle and silently skip dedup. + """ + fields = issue.get("fields", {}) + result = { + "task_id": task_id, + "task": { + "summary": fields.get("summary", ""), + # Jira may return an explicit null description; `.get(key, {})` only + # defaults on an ABSENT key, so `or {}` also coerces null → {} to + # keep task.description an object per verify-pr-input.schema.json. + "description": fields.get("description") or {}, + "status": (fields.get("status") or {}).get("name", ""), + "labels": fields.get("labels", []), + "issue_links": [ + { + "type": (link.get("type") or {}).get("name", ""), + "direction": "inward" if "inwardIssue" in link else "outward", + "key": ( + link.get("inwardIssue") or link.get("outwardIssue") or {} + ).get("key", ""), + } + for link in fields.get("issuelinks", []) + ], + "custom_fields": { + k: v + for k, v in fields.items() + if k.startswith("customfield_") + }, + }, + "pr_url": pr_url, + "source": { + "tracker": "jira", + "raw": issue, + }, + } + if github is not None: + # Annotate each commit with the deterministic Commit Traceability fact + # (Check 3). commits.items is unconstrained in the input schema, so the + # extra key is schema-valid; the sandbox agent reads references_task_id + # instead of running its own subjects-only git log. + commits = github.get("commits") + if isinstance(commits, list): + for commit in commits: + if isinstance(commit, dict): + commit["references_task_id"] = commit_references_task( + commit, task_id) + result["github"] = github + result["idempotency"] = ( + idempotency if idempotency is not None else {"related_issues": []} + ) + return result + + +def _read_text(path): + with open(path) as f: + return f.read() + + +def _read_json(path): + with open(path) as f: + return json.load(f) + + +def _read_own_check_names(path): + """Read verify-pr's own workflow check-run names (one per line) from PATH. + + The shell gathers these on the trusted runner (via `gh`) and writes them to a + file — one name per line, so job names containing spaces survive intact — so + the pure ``filter_own_check_runs`` helper stays unit-testable off the file. + Returns an empty set when PATH is falsy (the option was not supplied) or the + file is empty; an empty set means no self-exclusion, keeping prior behavior. + """ + if not path: + return set() + with open(path) as f: + return {line.strip() for line in f if line.strip()} + + +def _github_from_dir(args): + """Assemble the github bundle from raw read files written by the shell. + + Filters verify-pr's own workflow check-runs (TC-6343) out of the prefetched + check-runs before embedding them, using the own-check-name set the shell + gathered on the runner. This is done here (not in the shell's jq reduction) so + the exclusion logic is a unit-testable pure function. + """ + d = args.github_dir.rstrip("/") + check_runs = _read_json(f"{d}/check-runs.json") + own_names = _read_own_check_names(getattr(args, "own_check_names_file", None)) + check_runs = filter_own_check_runs(check_runs, own_names) + return build_github_bundle( + pr_repo=args.pr_repo, + pr_number=args.pr_number, + head_ref=args.head_ref, + commit_sha=args.commit_sha, + diff=_read_text(f"{d}/pr.diff"), + stat=_read_text(f"{d}/pr.stat"), + reviews=_read_json(f"{d}/reviews.json"), + review_comments=_read_json(f"{d}/review-comments.json"), + issue_comments=_read_json(f"{d}/issue-comments.json"), + commits=_read_json(f"{d}/commits.json"), + check_runs=check_runs, + check_run_logs_path=_check_run_logs_path(f"{d}/check-run-logs.txt"), + ) + + +def _check_run_logs_path(path): + """Sandbox path for the failed-check logs, or "" when none were captured. + + The shell always creates check-run-logs.txt (empty when nothing failed) so + its host_files mount is never missing. A non-empty file means at least one + check failed and its log was fetched; return the mounted sandbox path so the + sub-agent can Read it. An empty (or absent) file means no failure logs — emit + "" so Check 1b skips the read entirely. + """ + try: + if os.path.getsize(path) > 0: + return SANDBOX_CHECK_RUN_LOGS_PATH + except OSError: + pass + return "" + + +def _idempotency_from_dir(path): + """Read every related-issue JSON the shell wrote and build the bundle. + + Globs ``/*.json`` (one file per related key, written by + pre-verify-pr.sh). An empty directory yields an empty related_issues list. + """ + files = sorted(glob.glob(os.path.join(path, "*.json"))) + return build_idempotency_bundle([_read_json(f) for f in files]) + + +def main(argv): + parser = argparse.ArgumentParser(prog="pre_verify_pr.py") + sub = parser.add_subparsers(dest="command", required=True) + + sub.add_parser("extract-pr-url") + sub.add_parser("related-keys") + + jql = sub.add_parser("build-pr-jql") + jql.add_argument("pr_url") + + gate = sub.add_parser("resolve-gated-issue") + gate.add_argument("pr_url") + + reval = sub.add_parser("revalidate-gate") + reval.add_argument("pr_url") + + t = sub.add_parser("transform") + t.add_argument("task_id") + t.add_argument("pr_url") + t.add_argument("--github-dir") + t.add_argument("--pr-repo") + t.add_argument("--pr-number", type=int) + t.add_argument("--head-ref") + t.add_argument("--commit-sha") + t.add_argument("--own-check-names-file") + t.add_argument("--idempotency-dir") + + args = parser.parse_args(argv) + + # build-pr-jql takes only an argument — it reads no stdin, so resolve it + # before the stdin-consuming commands. + if args.command == "build-pr-jql": + print(build_pr_jql(args.pr_url)) + return + + if args.command == "resolve-gated-issue": + # Exit 3 is the gate-failure contract the shell maps to an ADR-0072 + # skip; exit 1 (argparse/JSON errors) stays a hard failure. + search_result = json.load(sys.stdin) + key, reason = resolve_gated_issue(search_result, args.pr_url) + if reason is not None: + # Diagnostic to stderr (runner log only — stdout stays the clean + # skip reason). Reveals what the JQL search actually returned so a + # gate failure can be told apart from an empty/visibility-limited + # search. Only non-secret shape is logged: issue count, keys, and + # the PR URL extracted from each candidate's Git Pull Request field. + issues = ( + search_result.get("issues", []) + if isinstance(search_result, dict) else [] + ) + summary = ", ".join( + "{}=>{!r}".format(i.get("key", "?"), extract_pr_url(i)) + for i in issues + ) or "(none)" + print( + "resolve-gated-issue: search returned {} issue(s): {}".format( + len(issues), summary), + file=sys.stderr, + ) + print(reason) + sys.exit(3) + print(key) + return + + if args.command == "revalidate-gate": + # TOCTOU re-check on the full issue (read from stdin), used just before + # the sandbox input is written. Same exit contract as resolve-gated-issue: + # 3 → gate failed (shell maps to an ADR-0072 skip), 0 → still qualifies. + issue = json.load(sys.stdin) + key, reason = revalidate_gate(issue, args.pr_url) + if reason is not None: + print( + "revalidate-gate: full issue no longer satisfies the gate " + "before writing sandbox input: {}".format(reason), + file=sys.stderr, + ) + print(reason) + sys.exit(3) + print(key) + return + + issue = json.load(sys.stdin) + + if args.command == "extract-pr-url": + print(extract_pr_url(issue)) + elif args.command == "related-keys": + for key in related_keys(issue): + print(key) + elif args.command == "transform": + github = _github_from_dir(args) if args.github_dir else None + idempotency = ( + _idempotency_from_dir(args.idempotency_dir) + if args.idempotency_dir else None + ) + result = transform_to_input( + issue, args.task_id, args.pr_url, github, idempotency) + json.dump(result, sys.stdout, indent=2) + + +if __name__ == "__main__": + main(sys.argv[1:]) diff --git a/plugins/sdlc-workflow/scripts/strip_extra_properties.py b/plugins/sdlc-workflow/scripts/strip_extra_properties.py new file mode 100644 index 000000000..1a361a416 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/strip_extra_properties.py @@ -0,0 +1,101 @@ +#!/usr/bin/env python3 +"""Strip additional properties from JSON based on a JSON Schema. + +Recursively walks the schema tree and removes properties not declared +in `properties` or `allOf/if/then/properties` at every node where +`additionalProperties: false`. Works with any schema structure +including discriminated unions (allOf with if/then). + +CLI usage (called by validate-output-schema.sh): + python3 strip_extra_properties.py + +Strips the JSON file in-place and exits 0. The caller validates after. +""" + +import json +import sys + + +def _resolve_ref(ref, root): + node = root + for part in ref.lstrip("#/").split("/"): + node = node[part] + return node + + +def _deref(schema, root): + if "$ref" in schema: + return _resolve_ref(schema["$ref"], root) + return schema + + +def _matching_then(instance, branches): + """Find the allOf branch whose `if` matches the instance.""" + for branch in branches: + if_clause = branch.get("if", {}) + if_props = if_clause.get("properties", {}) + match = all( + instance.get(k) == v.get("const") + for k, v in if_props.items() + if "const" in v + ) + if match and "then" in branch: + return branch["then"] + return None + + +def strip(instance, schema, root): + """Recursively strip properties not allowed by the schema.""" + schema = _deref(schema, root) + + if not isinstance(instance, dict) or schema.get("type") not in ("object", None): + return instance + + then = _matching_then(instance, schema.get("allOf", [])) + has_strict = schema.get("additionalProperties") is False + then_strict = then is not None and then.get("additionalProperties") is False + + if has_strict or then_strict: + allowed = set(schema.get("properties", {}).keys()) + if then: + allowed |= set(then.get("properties", {}).keys()) + removed = [k for k in instance if k not in allowed] + if removed: + print(f" stripped: {removed}") + instance = {k: v for k, v in instance.items() if k in allowed} + + for key, prop_schema in schema.get("properties", {}).items(): + if key not in instance: + continue + prop_schema = _deref(prop_schema, root) + if isinstance(instance[key], dict): + instance[key] = strip(instance[key], prop_schema, root) + elif isinstance(instance[key], list): + items_schema = prop_schema.get("items", {}) + items_schema = _deref(items_schema, root) + instance[key] = [ + strip(item, items_schema, root) if isinstance(item, dict) else item + for item in instance[key] + ] + + return instance + + +def main(): + if len(sys.argv) < 3: + print("Usage: strip_extra_properties.py ", file=sys.stderr) + sys.exit(1) + + with open(sys.argv[1]) as f: + instance = json.load(f) + with open(sys.argv[2]) as f: + schema = json.load(f) + + instance = strip(instance, schema, schema) + + with open(sys.argv[1], "w") as f: + json.dump(instance, f, indent=2) + + +if __name__ == "__main__": + main() diff --git a/plugins/sdlc-workflow/scripts/test_execute_actions.py b/plugins/sdlc-workflow/scripts/test_execute_actions.py new file mode 100644 index 000000000..ba8480965 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/test_execute_actions.py @@ -0,0 +1,1065 @@ +#!/usr/bin/env python3 +"""Tests for execute-actions.py ref resolution and native Jira comment posting.""" + +import sys +import os +import json +import tempfile +import importlib.util + +from jsonschema import validate, ValidationError + +script_dir = os.path.dirname(os.path.abspath(__file__)) +spec = importlib.util.spec_from_file_location( + "execute_actions", + os.path.join(script_dir, "execute-actions.py"), +) +execute_actions = importlib.util.module_from_spec(spec) +spec.loader.exec_module(execute_actions) + +resolve_refs = execute_actions.resolve_refs + +# The result schema the fullsend validation_loop enforces before the post_script +# runs. Loaded once so the schema-validation tests below assert against the real +# shipped constraints rather than a reimplementation. +_RESULT_SCHEMA_PATH = os.path.join( + script_dir, "..", "schemas", "verify-pr-result.schema.json" +) +with open(_RESULT_SCHEMA_PATH) as _schema_f: + _RESULT_SCHEMA = json.load(_schema_f) + +# Validate a single action instance against the schema's action definition. The +# action def carries no external $refs, so wrapping it with the document's $defs +# and dialect lets `validate` exercise the post_comment if/then branch directly. +_ACTION_SCHEMA = { + "$schema": _RESULT_SCHEMA["$schema"], + "$defs": _RESULT_SCHEMA["$defs"], + "$ref": "#/$defs/action", +} + + +def test_resolve_refs_replaces_key(): + """A single {{ref.key}}/{{ref.url}} pair resolves to registered values.""" + registry = {"subtask-1": {"key": "TC-100", "url": "https://jira.example.com/browse/TC-100"}} + text = "Sub-task [{{subtask-1.key}}]({{subtask-1.url}}) created." + result = resolve_refs(text, registry) + assert result == "Sub-task [TC-100](https://jira.example.com/browse/TC-100) created.", f"Got: {result}" + + +def test_resolve_refs_no_placeholders(): + """Text without placeholders is returned unchanged.""" + registry = {} + text = "No placeholders here." + result = resolve_refs(text, registry) + assert result == "No placeholders here." + + +def test_resolve_refs_unknown_ref_raises(): + """An unregistered ref raises KeyError.""" + registry = {} + text = "{{unknown-ref.key}}" + try: + resolve_refs(text, registry) + assert False, "Should have raised KeyError" + except KeyError: + pass + + +def test_resolve_refs_in_adf(): + """resolve_refs_in_obj resolves placeholders nested inside an ADF doc.""" + registry = {"rc-1": {"key": "TC-200", "url": "https://jira.example.com/browse/TC-200"}} + adf = { + "type": "doc", + "content": [ + {"type": "text", "text": "Task {{rc-1.key}} created"} + ] + } + result = execute_actions.resolve_refs_in_obj(adf, registry) + assert result["content"][0]["text"] == "Task TC-200 created" + + +def test_resolve_refs_multiple_different_refs(): + """Multiple distinct refs in one string each resolve independently.""" + registry = { + "subtask-1": {"key": "TC-100", "url": "https://jira.example.com/browse/TC-100"}, + "rc-1": {"key": "TC-200", "url": "https://jira.example.com/browse/TC-200"}, + } + text = "Sub-task {{subtask-1.key}} and root-cause {{rc-1.key}} ({{rc-1.url}})." + result = resolve_refs(text, registry) + assert result == "Sub-task TC-100 and root-cause TC-200 (https://jira.example.com/browse/TC-200).", f"Got: {result}" + + +def test_resolve_refs_repeated_placeholder(): + """A placeholder repeated in one string resolves at every occurrence.""" + registry = {"subtask-1": {"key": "TC-100", "url": "https://jira.example.com/browse/TC-100"}} + text = "{{subtask-1.key}} depends on {{subtask-1.key}}." + result = resolve_refs(text, registry) + assert result == "TC-100 depends on TC-100.", f"Got: {result}" + + +def test_resolve_refs_mixed_key_url_same_ref(): + """The .key and .url fields of one ref resolve to their respective values.""" + registry = {"subtask-1": {"key": "TC-100", "url": "https://jira.example.com/browse/TC-100"}} + text = "See {{subtask-1.key}} at {{subtask-1.url}}; {{subtask-1.key}} must be done first." + result = resolve_refs(text, registry) + assert result == "See TC-100 at https://jira.example.com/browse/TC-100; TC-100 must be done first.", f"Got: {result}" + + +class _FakeCompleted: + """Stand-in for subprocess.CompletedProcess.""" + + def __init__(self, returncode=0, stderr="", stdout=""): + self.returncode = returncode + self.stderr = stderr + self.stdout = stdout + + +class _RunRecorder: + """Captures the argv/input/env of a single subprocess.run call.""" + + def __init__(self, returncode=0, stderr=""): + self.returncode = returncode + self.stderr = stderr + self.cmd = None + self.input = None + self.env = None + + def __call__(self, cmd, input=None, text=None, capture_output=None, env=None): + self.cmd = cmd + self.input = input + self.env = env + return _FakeCompleted(self.returncode, self.stderr) + + +_JIRA_ENV = { + "JIRA_SERVER_URL": "https://jira.example.com", + "JIRA_EMAIL": "bot@example.com", + "JIRA_API_TOKEN": "s3cr3t", +} + + +def _with_jira_env_and_recorder(recorder): + """Install a fake subprocess.run + Jira env; return a restore callback.""" + saved_run = execute_actions.subprocess.run + saved_env = {k: os.environ.get(k) for k in _JIRA_ENV} + execute_actions.subprocess.run = recorder + os.environ.update(_JIRA_ENV) + + def restore(): + execute_actions.subprocess.run = saved_run + for k, v in saved_env.items(): + if v is None: + os.environ.pop(k, None) + else: + os.environ[k] = v + + return restore + + +def test_post_jira_comment_native_builds_argv(): + """post_jira_comment_native calls the native CLI with marker, project, and number.""" + recorder = _RunRecorder() + restore = _with_jira_env_and_recorder(recorder) + try: + execute_actions.post_jira_comment_native("TC-321", "hello **world**") + finally: + restore() + + assert recorder.cmd[:4] == ["fullsend", "issues", "post-comment", "--tracker"], f"Got: {recorder.cmd}" + assert "jira" in recorder.cmd + assert "--project" in recorder.cmd and recorder.cmd[recorder.cmd.index("--project") + 1] == "TC" + assert "--number" in recorder.cmd and recorder.cmd[recorder.cmd.index("--number") + 1] == "321" + assert "--marker" in recorder.cmd + assert recorder.cmd[recorder.cmd.index("--marker") + 1] == execute_actions.STICKY_COMMENT_MARKER + assert "--result" in recorder.cmd and recorder.cmd[recorder.cmd.index("--result") + 1] == "-" + assert recorder.input == "hello **world**" + + +def test_post_jira_comment_native_maps_env(): + """The native CLI receives JIRA_BASE_URL/JIRA_USER_EMAIL/JIRA_TOKEN mapped from this script's vars.""" + recorder = _RunRecorder() + restore = _with_jira_env_and_recorder(recorder) + try: + execute_actions.post_jira_comment_native("TC-1", "body") + finally: + restore() + + assert recorder.env["JIRA_BASE_URL"] == "https://jira.example.com" + assert recorder.env["JIRA_USER_EMAIL"] == "bot@example.com" + assert recorder.env["JIRA_TOKEN"] == "s3cr3t" + + +def test_post_jira_comment_native_nonzero_exits(): + """A non-zero CLI exit aborts with sys.exit(1).""" + recorder = _RunRecorder(returncode=1, stderr="boom") + restore = _with_jira_env_and_recorder(recorder) + try: + execute_actions.post_jira_comment_native("TC-1", "body") + assert False, "Should have exited" + except SystemExit as e: + assert e.code == 1 + finally: + restore() + + +def test_post_jira_comment_native_invalid_key_exits(): + """A malformed issue key (no hyphen) aborts before invoking the CLI.""" + recorder = _RunRecorder() + restore = _with_jira_env_and_recorder(recorder) + try: + execute_actions.post_jira_comment_native("TC123", "body") + assert False, "Should have exited" + except SystemExit as e: + assert e.code == 1 + finally: + restore() + assert recorder.cmd is None, "CLI should not run for an invalid key" + + +def test_schema_post_comment_accepts_valid_jira_key(): + """The post_comment schema accepts a well-formed hyphenated Jira key so a + legitimate action still validates and routes to the native CLI.""" + # Given a post_comment action whose issue is a valid Jira key + action = {"type": "post_comment", "issue": "TC-5811", "body_adf": {}} + # When validating it against the result schema's action definition + # Then validation passes (validate raises ValidationError on failure) + validate(instance=action, schema=_ACTION_SCHEMA) + + +def test_schema_post_comment_accepts_ref_key_placeholder(): + """The post_comment schema accepts a {{.key}} placeholder so a comment + targeting an issue created by an earlier action (which execute_post_comment + resolves via resolve_refs) survives fullsend's validation_loop instead of + being rejected before the reference can be resolved.""" + # Given a post_comment action whose issue is a {{.key}} placeholder + action = {"type": "post_comment", "issue": "{{sub-1.key}}", "body_adf": {}} + # When validating it against the result schema's action definition + # Then validation passes (validate raises ValidationError on failure) + validate(instance=action, schema=_ACTION_SCHEMA) + + +def test_schema_post_comment_rejects_non_key_issue(): + """A schema-valid-string-but-non-key issue (numeric ID, URL, lowercase, + missing hyphen, or a non-.key placeholder) is rejected at validation, so it + can never pass the producer boundary only to hit the executor's rpartition + guard and sys.exit(1).""" + # Given post_comment actions whose issue is neither a hyphenated Jira key + # nor a {{.key}} placeholder + non_keys = [ + "12345", # numeric Jira ID + "https://jira.example.com/browse/TC-5811", # URL + "tc-5811", # lowercase project + "TC5811", # missing hyphen + "TC-", # missing number + "-5811", # missing project + "{{sub-1.url}}", # .url placeholder (not a key) + "{{SUB.key}}", # uppercase ref name + "{{sub-1.status}}", # unsupported placeholder attr + "sub-1.key", # missing braces + "prefix {{sub-1.key}}", # placeholder not anchored + ] + for issue in non_keys: + action = {"type": "post_comment", "issue": issue, "body_adf": {}} + # When validating each against the schema + # Then validation fails before the action can reach the executor + try: + validate(instance=action, schema=_ACTION_SCHEMA) + assert False, f"non-key issue should be rejected: {issue!r}" + except ValidationError: + pass + + +def test_adf_to_markdown_renders_blocks_and_marks(): + """adf_to_markdown renders headings, lists, code blocks, rules, and inline marks.""" + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "heading", "attrs": {"level": 2}, + "content": [{"type": "text", "text": "Title"}]}, + {"type": "paragraph", "content": [ + {"type": "text", "text": "See "}, + {"type": "text", "text": "TC-1", "marks": [{"type": "strong"}]}, + {"type": "text", "text": " and "}, + {"type": "text", "text": "run", "marks": [{"type": "code"}]}, + {"type": "text", "text": " at "}, + {"type": "text", "text": "here", + "marks": [{"type": "link", "attrs": {"href": "https://x.example/y"}}]}, + ]}, + {"type": "bulletList", "content": [ + {"type": "listItem", "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "first"}]}]}, + {"type": "listItem", "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "second"}]}]}, + ]}, + {"type": "rule"}, + {"type": "codeBlock", "attrs": {"language": "python"}, + "content": [{"type": "text", "text": "x = 1"}]}, + ], + } + result = execute_actions.adf_to_markdown(doc) + expected = ( + "## Title\n\n" + "See **TC-1** and `run` at [here](https://x.example/y)\n\n" + "- first\n- second\n\n" + "---\n\n" + "```python\nx = 1\n```" + ) + assert result == expected, f"Got: {result!r}" + + +def test_adf_to_markdown_renders_task_list(): + """adf_to_markdown renders taskList DONE/TODO items as - [x] / - [ ] markers, for both paragraph-wrapped and inline taskItem content.""" + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "taskList", "content": [ + {"type": "taskItem", "attrs": {"state": "DONE"}, "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "done item"}]}]}, + {"type": "taskItem", "attrs": {"state": "TODO"}, "content": [ + {"type": "text", "text": "todo item"}]}, + ]}, + ], + } + result = execute_actions.adf_to_markdown(doc) + assert result == "- [x] done item\n- [ ] todo item", f"Got: {result!r}" + + +def test_adf_to_markdown_escapes_markdown_active_chars_in_literal_text(): + """Literal markdown-active characters in a plain text node are backslash-escaped + so the native CLI renders them verbatim instead of reinterpreting them as + formatting.""" + # Given a paragraph whose literal text contains * _ [ ] and a backtick + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "paragraph", "content": [ + {"type": "text", "text": "a*b_c[d]e`f"}, + ]}, + ], + } + # When rendering the ADF to markdown + result = execute_actions.adf_to_markdown(doc) + # Then each active character is escaped with a leading backslash + assert result == "a\\*b\\_c\\[d\\]e\\`f", f"Got: {result!r}" + + +def test_adf_to_markdown_does_not_double_escape_marks_or_code(): + """Intentional marks (strong/link) and inline code render correctly: the mark + syntax the renderer adds is not escaped, inline-code content stays literal, and + link hrefs are not escaped.""" + # Given marked text, an inline-code span containing an asterisk, and a link + # whose href contains an underscore + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "paragraph", "content": [ + {"type": "text", "text": "bold", "marks": [{"type": "strong"}]}, + {"type": "text", "text": " and "}, + {"type": "text", "text": "a*b", "marks": [{"type": "code"}]}, + {"type": "text", "text": " see "}, + {"type": "text", "text": "here", + "marks": [{"type": "link", "attrs": {"href": "https://x.example/a_b"}}]}, + ]}, + ], + } + # When rendering the ADF to markdown + result = execute_actions.adf_to_markdown(doc) + # Then the ** stays, the code asterisk stays literal, and the href underscore + # is preserved (none are escaped) + assert result == "**bold** and `a*b` see [here](https://x.example/a_b)", f"Got: {result!r}" + + +def test_adf_to_markdown_escapes_line_leading_block_markers(): + """A paragraph whose literal text begins with a Markdown block marker + (heading/bullet/blockquote/ordered-list) has that marker backslash-escaped so + the native CLI renders it verbatim instead of re-parsing it as a block.""" + # Given paragraphs each starting with a different line-leading block marker + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "# not a heading"}]}, + {"type": "paragraph", "content": [{"type": "text", "text": "### also not"}]}, + {"type": "paragraph", "content": [{"type": "text", "text": "- not a bullet"}]}, + {"type": "paragraph", "content": [{"type": "text", "text": "+ not a bullet"}]}, + {"type": "paragraph", "content": [{"type": "text", "text": "> not a quote"}]}, + {"type": "paragraph", "content": [{"type": "text", "text": "1. not ordered"}]}, + {"type": "paragraph", "content": [{"type": "text", "text": "2) not ordered"}]}, + ], + } + # When rendering the ADF to markdown + result = execute_actions.adf_to_markdown(doc) + # Then the leading marker of each line is escaped (heading/bullet/quote escape + # the first char; ordered lists escape the . / ) separator) + assert result == ( + "\\# not a heading\n\n" + "\\### also not\n\n" + "\\- not a bullet\n\n" + "\\+ not a bullet\n\n" + "\\> not a quote\n\n" + "1\\. not ordered\n\n" + "2\\) not ordered" + ), f"Got: {result!r}" + + +def test_adf_to_markdown_escapes_line_leading_marker_after_hardbreak(): + """A block marker that starts a line *after* a hardBreak inside a paragraph is + escaped too, since it is at a real line start once rendered.""" + # Given a paragraph with a hardBreak followed by text starting with "# " + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "paragraph", "content": [ + {"type": "text", "text": "see:"}, + {"type": "hardBreak"}, + {"type": "text", "text": "# heading"}, + ]}, + ], + } + # When rendering the ADF to markdown + result = execute_actions.adf_to_markdown(doc) + # Then only the post-hardBreak line-leading marker is escaped + assert result == "see:\n\\# heading", f"Got: {result!r}" + + +def test_adf_to_markdown_does_not_escape_midline_or_non_marker_text(): + """Escaping is line-position-sensitive: a marker char mid-line, or a + marker-like prefix that does not actually form a block (no trailing space, a + heading start intentionally emitted by the heading renderer), is left alone.""" + # Given a heading node, a paragraph with a mid-line '#', and paragraphs whose + # leading chars do not form a block construct ("-5", "1.5" have no space) + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "heading", "attrs": {"level": 2}, + "content": [{"type": "text", "text": "Real Heading"}]}, + {"type": "paragraph", "content": [{"type": "text", "text": "not # a heading"}]}, + {"type": "paragraph", "content": [{"type": "text", "text": "-5 degrees"}]}, + {"type": "paragraph", "content": [{"type": "text", "text": "1.5 times"}]}, + ], + } + # When rendering the ADF to markdown + result = execute_actions.adf_to_markdown(doc) + # Then the real heading keeps its intentional prefix and nothing else is escaped + assert result == ( + "## Real Heading\n\n" + "not # a heading\n\n" + "-5 degrees\n\n" + "1.5 times" + ), f"Got: {result!r}" + + +def test_adf_to_markdown_renders_non_text_inline_nodes(): + """Each non-text inline leaf node (mention/emoji/inlineCard/date/status) + renders its attrs-sourced displayable value instead of being dropped to an + empty string.""" + # Given a paragraph containing one of each non-text inline leaf type, with + # an underscore in the inlineCard URL (URLs must not be escaped) and an + # epoch-millisecond date timestamp for 2021-01-01 UTC + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "paragraph", "content": [ + {"type": "mention", "attrs": {"id": "abc", "text": "@Marco Rizzi"}}, + {"type": "text", "text": " "}, + {"type": "emoji", "attrs": {"shortName": ":smile:", "text": "😄"}}, + {"type": "text", "text": " "}, + {"type": "inlineCard", "attrs": {"url": "https://example.com/a_b"}}, + {"type": "text", "text": " "}, + {"type": "date", "attrs": {"timestamp": "1609459200000"}}, + {"type": "text", "text": " "}, + {"type": "status", "attrs": {"text": "In Progress", "color": "yellow"}}, + ]}, + ], + } + # When rendering the ADF to markdown + result = execute_actions.adf_to_markdown(doc) + # Then every node contributes its attrs value (URL underscore preserved, date + # formatted as YYYY-MM-DD) and nothing is silently dropped + assert result == "@Marco Rizzi 😄 https://example.com/a_b 2021-01-01 In Progress", \ + f"Got: {result!r}" + + +def test_adf_to_markdown_renders_inline_nodes_in_task_item(): + """A taskItem whose inline content mixes text with a mention and an + inlineCard renders all of them — the inline nodes are not dropped in the + taskItem context.""" + # Given a TODO taskItem with inline mention and inlineCard nodes + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "taskList", "content": [ + {"type": "taskItem", "attrs": {"state": "TODO"}, "content": [ + {"type": "text", "text": "ping "}, + {"type": "mention", "attrs": {"text": "@dev"}}, + {"type": "text", "text": " re "}, + {"type": "inlineCard", "attrs": {"url": "https://example.com/pr/1"}}, + ]}, + ]}, + ], + } + # When rendering the ADF to markdown + result = execute_actions.adf_to_markdown(doc) + # Then the checklist item retains the mention and inlineCard values + assert result == "- [ ] ping @dev re https://example.com/pr/1", f"Got: {result!r}" + + +def test_adf_to_markdown_renders_table(): + """A table renders as a GFM table: first tableRow is the header (with a --- + separator), tableCell/tableHeader content is rendered, and literal pipes in a + cell are escaped so they do not break the column grid.""" + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "table", "content": [ + {"type": "tableRow", "content": [ + {"type": "tableHeader", "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "Name"}]}]}, + {"type": "tableHeader", "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "Note"}]}]}, + ]}, + {"type": "tableRow", "content": [ + {"type": "tableCell", "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "a"}]}]}, + {"type": "tableCell", "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "b|c"}]}]}, + ]}, + ]}, + ], + } + result = execute_actions.adf_to_markdown(doc) + assert result == ( + "| Name | Note |\n" + "| --- | --- |\n" + "| a | b\\|c |" + ), f"Got: {result!r}" + + +def test_adf_to_markdown_renders_blockquote_and_panel(): + """blockquote renders as > -prefixed lines; a panel renders as a quote with a + bold panelType label so its kind is preserved.""" + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "blockquote", "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "quoted"}]}]}, + {"type": "panel", "attrs": {"panelType": "info"}, "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "heads up"}]}]}, + ], + } + result = execute_actions.adf_to_markdown(doc) + assert result == ( + "> quoted\n" + "\n" + "> **info**\n" + ">\n" + "> heads up" + ), f"Got: {result!r}" + + +def test_adf_to_markdown_renders_media_instead_of_dropping(): + """media/mediaSingle render a non-empty image link (or [alt] placeholder when + no URL is present) rather than being silently dropped — media nodes have no + text content, so the pre-fix flattening fallback emitted nothing.""" + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "mediaSingle", "content": [ + {"type": "media", "attrs": {"url": "https://ex.com/i.png", "alt": "chart"}}]}, + {"type": "mediaSingle", "content": [ + {"type": "media", "attrs": {"type": "file", "id": "abc-123"}}]}, + ], + } + result = execute_actions.adf_to_markdown(doc) + assert result == "![chart](https://ex.com/i.png)\n\n[abc-123]", f"Got: {result!r}" + + +def test_adf_to_markdown_unknown_block_still_flattens(): + """A genuinely unknown container block still degrades gracefully by recursing + into its nested content (the preserved fallback), so the fix does not regress + forward-compat handling of block nodes it does not explicitly cover.""" + doc = { + "type": "doc", + "version": 1, + "content": [ + {"type": "someFutureBlock", "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "still here"}]}]}, + ], + } + result = execute_actions.adf_to_markdown(doc) + assert result == "still here", f"Got: {result!r}" + + +def test_render_adf_date_falls_back_on_out_of_range_timestamp(): + """An integer-parseable but out-of-range epoch-ms timestamp renders as its + literal string instead of raising, so a single malformed date node cannot + abort the whole post_script. datetime.fromtimestamp raises OverflowError/OSError + (or ValueError) for out-of-range values, and that call now sits inside the + guarded try; before the fix it was outside and any such exception propagated.""" + # Given an integer-parseable epoch-millisecond value far outside the + # representable datetime range + out_of_range = "99999999999999999" + # When rendering it as an ADF date node + result = execute_actions._render_adf_date(out_of_range) + # Then the literal value is returned and no exception propagates + assert result == out_of_range, f"Got: {result!r}" + + +# A report's tracker-native body. execute_post_report renders this (not the +# GitHub-only report_md) to markdown for Jira via adf_to_markdown, so the fixtures +# below give it text that renders to a string distinct from report_md — proving +# Jira receives the ADF-derived body while GitHub keeps report_md + its marker. +_REPORT_ADF = { + "type": "doc", + "version": 1, + "content": [ + {"type": "paragraph", "content": [{"type": "text", "text": "Jira native body."}]} + ], +} +_REPORT_ADF_MD = "Jira native body." + + +def test_execute_post_comment_routes_to_native(): + """execute_post_comment resolves refs in body_adf, renders to markdown, and posts it.""" + recorder = _RunRecorder() + restore = _with_jira_env_and_recorder(recorder) + registry = {"sub-1": {"key": "TC-500", "url": "https://jira.example.com/browse/TC-500"}} + body_adf = { + "type": "doc", + "version": 1, + "content": [ + {"type": "paragraph", "content": [ + {"type": "text", "text": "See "}, + {"type": "text", "text": "{{sub-1.key}}", "marks": [{"type": "strong"}]}, + {"type": "text", "text": " at "}, + {"type": "text", "text": "link", + "marks": [{"type": "link", "attrs": {"href": "{{sub-1.url}}"}}]}, + ]}, + ], + } + try: + execute_actions.execute_post_comment( + {"type": "post_comment", "issue": "{{sub-1.key}}", "body_adf": body_adf}, + registry, + ) + finally: + restore() + + assert recorder.cmd[recorder.cmd.index("--number") + 1] == "500" + assert recorder.input == "See **TC-500** at [link](https://jira.example.com/browse/TC-500)", \ + f"Got: {recorder.input!r}" + + +def test_post_comment_and_report_use_distinct_sticky_markers(): + """A post_comment and a post_report pass DIFFERENT --marker values to the + native CLI, so two sticky comments on the same Jira issue never share one + marker identity and clobber each other.""" + # Given the comment path + comment_recorder = _RunRecorder() + restore = _with_jira_env_and_recorder(comment_recorder) + try: + execute_actions.execute_post_comment( + {"type": "post_comment", "issue": "TC-777", + "body_adf": {"type": "doc", "version": 1, "content": []}}, + {}, + ) + finally: + restore() + comment_marker = comment_recorder.cmd[comment_recorder.cmd.index("--marker") + 1] + + # And the report path (native GitHub comment + native Jira comment) + report_calls = [] + + def fake_run(cmd, input=None, text=None, capture_output=None, env=None): + report_calls.append(cmd) + return _FakeCompleted(0, "") + + saved_run = execute_actions.subprocess.run + saved_env = {k: os.environ.get(k) for k in _JIRA_ENV} + execute_actions.subprocess.run = fake_run + os.environ.update(_JIRA_ENV) + try: + report = { + "pr_repo": "acme/widget", + "pr_number": 42, + "jira_issue_id": "TC-777", + "commit_sha": "946556e", + "report_md": "## Verify report", + "report_adf": _REPORT_ADF, + } + execute_actions.execute_post_report({"type": "post_report"}, {}, report) + finally: + execute_actions.subprocess.run = saved_run + for k, v in saved_env.items(): + if v is None: + os.environ.pop(k, None) + else: + os.environ[k] = v + jira_cmd = next(c for c in report_calls + if c[:3] == ["fullsend", "issues", "post-comment"] + and c[c.index("--tracker") + 1] == "jira") + report_marker = jira_cmd[jira_cmd.index("--marker") + 1] + + # Then the two markers differ, and each matches its dedicated constant + assert comment_marker == execute_actions.POST_COMMENT_STICKY_MARKER + assert report_marker == execute_actions.STICKY_COMMENT_MARKER + assert comment_marker != report_marker, "post_comment and post_report must use distinct markers" + + +def test_jira_bound_markers_have_no_forbidden_chars(): + """Markers posted to Jira must avoid characters Jira's markdown round-trip + escapes (\\*_`[]&) — fullsend rejects such a --marker because the escaping + would break sticky-comment re-detection on later runs. Guards against a + regression like the underscore in the original "post_comment" marker.""" + forbidden = set("\\*_`[]&") + for marker in (execute_actions.STICKY_COMMENT_MARKER, + execute_actions.POST_COMMENT_STICKY_MARKER): + offending = forbidden & set(marker) + assert not offending, f"{marker!r} contains forbidden char(s) {offending}" + + +def test_execute_post_report_posts_github_then_jira(): + """execute_post_report posts the report to the GitHub PR then to Jira, both via + the native `fullsend issues post-comment` sticky CLI (GitHub first).""" + calls = [] + + def fake_run(cmd, input=None, text=None, capture_output=None, env=None): + calls.append({"cmd": cmd, "input": input, "env": env}) + return _FakeCompleted(0, "") + + saved_run = execute_actions.subprocess.run + saved_env = {k: os.environ.get(k) for k in _JIRA_ENV} + execute_actions.subprocess.run = fake_run + os.environ.update(_JIRA_ENV) + try: + report = { + "pr_repo": "acme/widget", + "pr_number": 42, + "jira_issue_id": "TC-777", + "commit_sha": "946556e", + "report_md": "## Verify report\nAll good.", + "report_adf": _REPORT_ADF, + } + execute_actions.execute_post_report({"type": "post_report"}, {}, report) + finally: + execute_actions.subprocess.run = saved_run + for k, v in saved_env.items(): + if v is None: + os.environ.pop(k, None) + else: + os.environ[k] = v + + # SHA canonicalization (git rev-parse) + native GitHub + native Jira. + assert len(calls) == 3, f"Expected rev-parse + github + jira calls, got {len(calls)}" + native = [c for c in calls if c["cmd"][:3] == ["fullsend", "issues", "post-comment"]] + gh_call = next(c for c in native if c["cmd"][c["cmd"].index("--tracker") + 1] == "github") + jira_call = next(c for c in native if c["cmd"][c["cmd"].index("--tracker") + 1] == "jira") + # GitHub side: native sticky CLI targets the PR by number, the marker carries + # the commit SHA, and the body (stdin) is report_md with NO embedded marker + # (the CLI prepends it). + assert gh_call["cmd"][gh_call["cmd"].index("--project") + 1] == "acme/widget" + assert gh_call["cmd"][gh_call["cmd"].index("--number") + 1] == "42" + assert gh_call["cmd"][gh_call["cmd"].index("--marker") + 1] == \ + "" + assert gh_call["input"] == "## Verify report\nAll good." + assert "sdlc-workflow:verify-pr report commit:" not in gh_call["input"], \ + "marker must not be embedded in the GitHub body; the CLI prepends it" + # GitHub is posted before Jira. + assert native[0] is gh_call, "GitHub report must be posted before Jira" + # Jira side: sticky CLI, and the body is rendered from report_adf (the + # tracker-native content) via adf_to_markdown, NOT the GitHub-only report_md, + # which carries the commit marker Jira must never receive. + assert jira_call["cmd"][jira_call["cmd"].index("--number") + 1] == "777" + assert jira_call["input"] == _REPORT_ADF_MD + assert jira_call["input"] != report["report_md"], "Jira must not receive report_md" + assert "sdlc-workflow:verify-pr report commit:" not in jira_call["input"] + + +def test_execute_post_report_strips_embedded_leading_marker_from_github_body(): + """When the agent's report_md already begins with the commit-scoped marker + line, execute_post_report strips it before calling the native CLI (which + prepends the marker), so the GitHub comment never opens with two duplicate + marker lines.""" + calls = [] + + def fake_run(cmd, input=None, text=None, capture_output=None, env=None): + calls.append({"cmd": cmd, "input": input}) + return _FakeCompleted(0, "") + + saved_run = execute_actions.subprocess.run + saved_env = {k: os.environ.get(k) for k in _JIRA_ENV} + execute_actions.subprocess.run = fake_run + os.environ.update(_JIRA_ENV) + try: + marker_line = "" + report = { + "pr_repo": "acme/widget", + "pr_number": 42, + "jira_issue_id": "TC-777", + "commit_sha": "946556e", + "report_md": f"{marker_line}\n## Verify report\nAll good.", + "report_adf": _REPORT_ADF, + } + execute_actions.execute_post_report({"type": "post_report"}, {}, report) + finally: + execute_actions.subprocess.run = saved_run + for k, v in saved_env.items(): + if v is None: + os.environ.pop(k, None) + else: + os.environ[k] = v + + gh_call = next(c for c in calls + if c["cmd"][:3] == ["fullsend", "issues", "post-comment"] + and c["cmd"][c["cmd"].index("--tracker") + 1] == "github") + # The leading marker line is stripped; the body starts with the report heading + # and carries no embedded marker (the CLI prepends the single marker copy). + assert gh_call["input"] == "## Verify report\nAll good.", f"Got: {gh_call['input']!r}" + assert "sdlc-workflow:verify-pr report commit:" not in gh_call["input"] + + +def _run_post_report_with_sha_resolution(commit_sha, resolve): + """Run execute_post_report with a fake ``git rev-parse`` that maps each input + ref to a canonical full SHA via ``resolve``. Returns the recorded calls.""" + calls = [] + + def fake_run(cmd, input=None, text=None, capture_output=None, env=None): + calls.append({"cmd": cmd, "input": input}) + if cmd[:2] == ["git", "rev-parse"]: + # cmd[-1] is "^{commit}"; strip the peel suffix to look up. + ref = cmd[-1].split("^", 1)[0] + return _FakeCompleted(0, "", resolve.get(ref, "")) + return _FakeCompleted(0, "") + + saved_run = execute_actions.subprocess.run + saved_env = {k: os.environ.get(k) for k in _JIRA_ENV} + execute_actions.subprocess.run = fake_run + os.environ.update(_JIRA_ENV) + try: + report = { + "pr_repo": "acme/widget", + "pr_number": 42, + "jira_issue_id": "TC-777", + "commit_sha": commit_sha, + "report_md": "## Verify report", + "report_adf": _REPORT_ADF, + } + execute_actions.execute_post_report({"type": "post_report"}, {}, report) + finally: + execute_actions.subprocess.run = saved_run + for k, v in saved_env.items(): + if v is None: + os.environ.pop(k, None) + else: + os.environ[k] = v + return calls + + +def _github_marker_from_calls(calls): + """Return the --marker passed to the native GitHub post-comment call.""" + gh = next(c for c in calls + if c["cmd"][:3] == ["fullsend", "issues", "post-comment"] + and c["cmd"][c["cmd"].index("--tracker") + 1] == "github") + return gh["cmd"][gh["cmd"].index("--marker") + 1] + + +def test_execute_post_report_dedup_marker_invariant_to_sha_length(): + """A full-length SHA and an abbreviated SHA for the SAME commit produce the + SAME --marker on the native GitHub call, because git rev-parse canonicalizes + both forms to the same full SHA. The native CLI then deduplicates on that + shared marker, so a retry edits one comment instead of duplicating it.""" + full_sha = "946556e" + "a" * 33 # 40 hex chars + short_sha = "946556e" # 7-char abbreviation of the same commit + # git rev-parse resolves either form of this one commit to its full SHA. + resolve = {full_sha: full_sha, short_sha: full_sha} + + full_marker = _github_marker_from_calls( + _run_post_report_with_sha_resolution(full_sha, resolve)) + short_marker = _github_marker_from_calls( + _run_post_report_with_sha_resolution(short_sha, resolve)) + + assert full_marker == f"{execute_actions.GITHUB_REPORT_MARKER_PREFIX}{full_sha} -->", \ + f"marker should carry the canonical full SHA: {full_marker!r}" + assert full_marker == short_marker, \ + "full and abbreviated SHA of one commit must yield the same sticky marker" + + +def test_execute_post_report_distinct_commits_same_prefix_do_not_collide(): + """Two distinct commits sharing a 7-hex prefix get DISTINCT --marker values + (each resolved up to its full 40-char SHA), so the native CLI keeps them as + separate sticky comments instead of one overwriting the other.""" + full_a = "946556e" + "a" * 33 # commit A + full_b = "946556e" + "b" * 33 # commit B, same first 7 hex chars, distinct object + resolve = {full_a: full_a, full_b: full_b} + + marker_a = _github_marker_from_calls( + _run_post_report_with_sha_resolution(full_a, resolve)) + marker_b = _github_marker_from_calls( + _run_post_report_with_sha_resolution(full_b, resolve)) + + assert marker_a == f"{execute_actions.GITHUB_REPORT_MARKER_PREFIX}{full_a} -->" + assert marker_b == f"{execute_actions.GITHUB_REPORT_MARKER_PREFIX}{full_b} -->" + assert marker_a != marker_b, "distinct commits must get distinct sticky markers" + + +def test_normalize_commit_sha_falls_back_to_prefix_when_unresolvable(): + """When git cannot resolve the SHA (git missing or object absent), + _normalize_commit_sha falls back to the fixed-length prefix so dedup keeps + working instead of being disabled.""" + long_sha = "946556e" + "f" * 33 + + # git rev-parse reports failure (non-zero, empty stdout) → fallback to prefix. + def failing_run(cmd, input=None, text=None, capture_output=None, env=None): + return _FakeCompleted(1, "fatal: Needed a single revision", "") + + saved_run = execute_actions.subprocess.run + execute_actions.subprocess.run = failing_run + try: + result = execute_actions._normalize_commit_sha(long_sha) + finally: + execute_actions.subprocess.run = saved_run + assert result == long_sha[:execute_actions.COMMIT_SHA_MARKER_LENGTH], \ + f"Got: {result!r}" + + # git binary missing (OSError) → same fallback. + def raising_run(cmd, input=None, text=None, capture_output=None, env=None): + raise FileNotFoundError("git") + + execute_actions.subprocess.run = raising_run + try: + result = execute_actions._normalize_commit_sha(long_sha) + finally: + execute_actions.subprocess.run = saved_run + assert result == long_sha[:execute_actions.COMMIT_SHA_MARKER_LENGTH], \ + f"Got: {result!r}" + + +def test_post_report_resolves_ref_created_by_later_action(): + """report_md refs resolve regardless of action ordering: a post_report ordered + BEFORE the create_subtask it references still resolves, because main() defers + every post_report until after the actions loop populates the registry.""" + calls = [] + + def fake_run(cmd, input=None, text=None, capture_output=None, env=None): + calls.append({"cmd": cmd, "input": input}) + return _FakeCompleted(0, "") + + def fake_create_issue(**kwargs): + return {"key": "TC-999"} + + # Actions deliberately order post_report FIRST, then the create_subtask whose + # ref its report_md interpolates — the pre-fix inline order would KeyError. + data = { + "report": { + "pr_repo": "acme/widget", + "pr_number": 42, + "jira_issue_id": "TC-777", + "commit_sha": "946556e", + "report_md": "Filed sub-task {{sub-1.key}}.", + "report_adf": { + "type": "doc", + "version": 1, + "content": [ + {"type": "paragraph", + "content": [{"type": "text", "text": "Filed sub-task {{sub-1.key}}."}]} + ], + }, + }, + "actions": [ + {"type": "post_report"}, + { + "type": "create_subtask", + "ref": "sub-1", + "parent": "TC-100", + "summary": "A sub-task", + "labels": ["review-feedback"], + "description_adf": {"type": "doc", "version": 1, "content": []}, + }, + ], + } + + saved_run = execute_actions.subprocess.run + saved_create = execute_actions._jira_mod.create_issue + saved_argv = sys.argv + saved_env = {k: os.environ.get(k) for k in _JIRA_ENV} + execute_actions.subprocess.run = fake_run + execute_actions._jira_mod.create_issue = fake_create_issue + os.environ.update(_JIRA_ENV) + fd, path = tempfile.mkstemp(suffix=".json") + try: + with os.fdopen(fd, "w") as f: + json.dump(data, f) + sys.argv = ["execute-actions.py", path] + execute_actions.main() + finally: + execute_actions.subprocess.run = saved_run + execute_actions._jira_mod.create_issue = saved_create + sys.argv = saved_argv + os.remove(path) + for k, v in saved_env.items(): + if v is None: + os.environ.pop(k, None) + else: + os.environ[k] = v + + # The GitHub report body (native call stdin) carries the resolved key, not the + # raw placeholder. + native = [c for c in calls if c["cmd"][:3] == ["fullsend", "issues", "post-comment"]] + gh_call = next(c for c in native if c["cmd"][c["cmd"].index("--tracker") + 1] == "github") + assert "Filed sub-task TC-999." in gh_call["input"], f"ref not resolved: {gh_call['input']}" + assert "{{sub-1.key}}" not in gh_call["input"], "placeholder leaked into report body" + # The Jira body is rendered from report_adf, and its refs resolve too. + jira_call = next(c for c in native if c["cmd"][c["cmd"].index("--tracker") + 1] == "jira") + assert jira_call["input"] == "Filed sub-task TC-999.", f"ref not resolved in ADF: {jira_call['input']}" + assert "{{sub-1.key}}" not in jira_call["input"], "placeholder leaked into Jira body" + + +if __name__ == "__main__": + test_resolve_refs_replaces_key() + test_resolve_refs_no_placeholders() + test_resolve_refs_unknown_ref_raises() + test_resolve_refs_in_adf() + test_resolve_refs_multiple_different_refs() + test_resolve_refs_repeated_placeholder() + test_resolve_refs_mixed_key_url_same_ref() + test_post_jira_comment_native_builds_argv() + test_post_jira_comment_native_maps_env() + test_post_jira_comment_native_nonzero_exits() + test_post_jira_comment_native_invalid_key_exits() + test_schema_post_comment_accepts_valid_jira_key() + test_schema_post_comment_accepts_ref_key_placeholder() + test_schema_post_comment_rejects_non_key_issue() + test_adf_to_markdown_renders_blocks_and_marks() + test_adf_to_markdown_renders_task_list() + test_adf_to_markdown_escapes_markdown_active_chars_in_literal_text() + test_adf_to_markdown_does_not_double_escape_marks_or_code() + test_adf_to_markdown_escapes_line_leading_block_markers() + test_adf_to_markdown_escapes_line_leading_marker_after_hardbreak() + test_adf_to_markdown_does_not_escape_midline_or_non_marker_text() + test_adf_to_markdown_renders_non_text_inline_nodes() + test_adf_to_markdown_renders_inline_nodes_in_task_item() + test_adf_to_markdown_renders_table() + test_adf_to_markdown_renders_blockquote_and_panel() + test_adf_to_markdown_renders_media_instead_of_dropping() + test_adf_to_markdown_unknown_block_still_flattens() + test_render_adf_date_falls_back_on_out_of_range_timestamp() + test_execute_post_comment_routes_to_native() + test_post_comment_and_report_use_distinct_sticky_markers() + test_execute_post_report_posts_github_then_jira() + test_execute_post_report_strips_embedded_leading_marker_from_github_body() + test_execute_post_report_dedup_marker_invariant_to_sha_length() + test_execute_post_report_distinct_commits_same_prefix_do_not_collide() + test_normalize_commit_sha_falls_back_to_prefix_when_unresolvable() + test_post_report_resolves_ref_created_by_later_action() + print("All tests passed.") diff --git a/plugins/sdlc-workflow/scripts/test_execute_triage_security_actions.py b/plugins/sdlc-workflow/scripts/test_execute_triage_security_actions.py new file mode 100644 index 000000000..db78f32c7 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/test_execute_triage_security_actions.py @@ -0,0 +1,797 @@ +#!/usr/bin/env python3 +"""Tests for trusted triage-security action execution.""" + +import copy +import hashlib +import importlib.util +import json +import os +import subprocess + +import pytest + + +SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__)) +SPEC = importlib.util.spec_from_file_location( + "execute_triage_security_actions", + os.path.join(SCRIPT_DIR, "execute-triage-security-actions.py"), +) +executor = importlib.util.module_from_spec(SPEC) +SPEC.loader.exec_module(executor) + + +def _plan(actions, mode="mutation-authorized"): + """Build a minimally valid triage-security result plan.""" + return { + "schema_version": "1", + "mode": mode, + "report": { + "issue": "TC-42", + "outcome": "affected", + "summary_markdown": "Affected by the vulnerability.", + "evidence": [{"source": "test", "detail": "deliberate fixture"}], + }, + "actions": actions, + } + + +def _trusted_input(authorized=True, markers=None, remediation=None, issue="TC-42", related=None, searches=None): + """Build the trusted runner fields consumed by the executor.""" + return { + "issue": {"key": issue}, + "configuration": {"project_key": "TC"}, + "jira_metadata": {"related_issues": related or [], "sibling_searches": searches or []}, + "authorization": {"mutation_authorized": authorized}, + "idempotency": { + "action_markers": markers or [], + "existing_remediation": remediation or [], + }, + } + + +class _JiraRecorder: + """Capture Jira client calls without issuing network requests.""" + + def __init__(self): + self.calls = [] + self.created = 0 + + def update_issue(self, issue, fields): + """Record a Jira field update.""" + self.calls.append(("field-edit", issue, fields)) + + def get_transitions(self, issue): + """Provide a catalog whose action name differs from its target status.""" + self.calls.append(("get-transitions", issue)) + return [{"id": "31", "name": "Start Progress", "to": {"name": "In Progress"}}] + + def transition_issue(self, issue, transition): + """Record a Jira transition.""" + self.calls.append(("transition", issue, transition)) + + def create_issue(self, **kwargs): + """Record remediation creation and return a deterministic Jira key.""" + self.created += 1 + self.calls.append(("remediation-task", kwargs)) + return {"key": "TC-900{}".format(self.created), "self": "https://jira.example/issue/900{}".format(self.created)} + + def get_issue(self, issue): + """Return the stored description Jira would return after normalization.""" + self.calls.append(("get-issue", issue)) + return {"fields": {"description": { + "type": "doc", "version": 1, + "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Stored by Jira."}]}], + }}} + + def create_link(self, inward, outward, link_type): + """Record an issue-link operation.""" + self.calls.append(("link", inward, outward, link_type)) + + def make_request(self, method, path, body): + """Record direct ADF digest comments.""" + self.calls.append(("digest", method, path, body)) + return {"id": "1"} + + def post_native(self, issue, body, marker): + """Record a native sticky Jira comment.""" + self.calls.append(("comment", issue, body, marker)) + + +@pytest.fixture +def recorder(monkeypatch): + """Replace Jira client calls with a recorder for each test.""" + value = _JiraRecorder() + monkeypatch.setattr(executor._jira_mod, "update_issue", value.update_issue) + monkeypatch.setattr(executor._jira_mod, "get_transitions", value.get_transitions) + monkeypatch.setattr(executor._jira_mod, "transition_issue", value.transition_issue) + monkeypatch.setattr(executor._jira_mod, "create_issue", value.create_issue) + monkeypatch.setattr(executor._jira_mod, "get_issue", value.get_issue) + monkeypatch.setattr(executor._jira_mod, "create_link", value.create_link) + monkeypatch.setattr(executor._jira_mod, "make_request", value.make_request) + monkeypatch.setattr(executor._action_helpers, "post_jira_comment_native", value.post_native) + return value + + +def test_untrusted_field_target_is_rejected_before_any_jira_call(recorder): + """A runner grant for TC-8100 cannot authorize a field edit on TC-42.""" + # Given a schema-valid result whose target differs from its trusted identity + result = _plan([{ + "type": "field-edit", "marker": "triage-security:labels", "issue": "TC-42", + "fields": {"labels": ["ai-cve-triaged"]}, + }]) + result["report"]["issue"] = "TC-8100" + + # When the executor checks the trusted runner grant + with pytest.raises(executor.ActionError): + executor.execute_plan(result, _trusted_input(issue="TC-8100")) + + # Then even read-before-write Jira calls are absent + assert recorder.calls == [] + + +def test_report_only_plan_never_mutates_jira(recorder): + """A report-only result is preserved without invoking any Jira operation.""" + # Given a sandbox result that contains only the report-only sentinel + result = _plan([{"type": "report-only", "marker": "triage-security:report"}], "report-only") + trusted = _trusted_input(authorized=False) + del trusted["issue"] + + # When the trusted runner executes it without mutation authorization + executor.execute_plan(result, trusted) + + # Then no Jira operation was requested + assert recorder.calls == [] + + +@pytest.mark.parametrize("issue", [None, {}, {"key": None}, {"key": 42}, {"key": "invalid"}, {"key": "TC-43"}]) +def test_missing_malformed_or_mismatched_trusted_identity_is_rejected(issue, recorder): + """Mutation grants require a valid trusted identity matching the report.""" + # Given a valid action but absent, malformed, or contradictory trusted identity + result = _plan([{"type": "field-edit", "marker": "triage-security:labels", + "issue": "TC-42", "fields": {"labels": ["ai-cve-triaged"]}}]) + trusted = _trusted_input() + trusted["issue"] = issue + + # When validating the grant against the complete plan + with pytest.raises(executor.ActionError): + executor.execute_plan(result, trusted) + + # Then no Jira call occurs + assert recorder.calls == [] + + +def _remediation_action(): + """Build a synthetic remediation action for authorization regressions.""" + return { + "type": "remediation-task", "marker": "triage-security:remediation", + "ref": "remediation", "project": "TC", "summary": "Fix CVE", + "description_adf": {"type": "doc", "version": 1, "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Synthetic triage."}]}]}, "labels": [], + } + + +@pytest.mark.parametrize("marked", [False, True]) +@pytest.mark.parametrize("definition", [None, {"targets": ()}, {}]) +def test_new_action_without_target_declaration_fails_before_any_jira_call(definition, marked, recorder, monkeypatch): + """A newly supported action cannot silently bypass target authorization.""" + # Given a future action recognized by validation but lacking target metadata + monkeypatch.setitem(executor._REQUIRED_ACTION_FIELDS, "future-mutation", {"type", "marker", "issue"}) + if definition is not None: + monkeypatch.setitem(executor._ACTION_DEFINITIONS, "future-mutation", definition) + result = _plan([ + {"type": "status-transition", "marker": "triage-security:first", "issue": "TC-42", "status": "In Progress"}, + {"type": "future-mutation", "marker": "triage-security:future", "issue": "TC-43"}, + ]) + trusted = _trusted_input(markers=["triage-security:future"] if marked else []) + + # When the complete plan is preflighted before executing its valid prefix + with pytest.raises(executor.ActionError, match="target fields"): + executor._preflight_plan(result, trusted) + + # Then even read-before-write Jira calls are absent + assert recorder.calls == [] + + +@pytest.mark.parametrize("target_fields", [None, ()]) +def test_mutation_with_missing_or_empty_targets_rejects_before_valid_prefix(target_fields, recorder, monkeypatch): + """Even a schema-valid mutation must have an authorization target declaration.""" + # Given a known action whose target declaration was omitted or left empty + definition = copy.deepcopy(executor._ACTION_DEFINITIONS["field-edit"]) + if target_fields is None: + del definition["targets"] + else: + definition["targets"] = target_fields + monkeypatch.setitem(executor._ACTION_DEFINITIONS, "field-edit", definition) + result = _plan([ + {"type": "status-transition", "marker": "triage-security:first", "issue": "TC-42", "status": "In Progress"}, + {"type": "field-edit", "marker": "triage-security:labels", "issue": "TC-42", "fields": {"labels": ["ai-cve-triaged"]}}, + ]) + + # When execution preflights the complete schema-valid plan + with pytest.raises(executor.ActionError, match="target fields"): + executor.execute_plan(result, _trusted_input()) + + # Then the valid prefix did not read transitions or write Jira state + assert recorder.calls == [] + + +@pytest.mark.parametrize("issue", ["TC-42", "TC-43"]) +def test_new_action_uses_its_declared_target_fields(issue, recorder, monkeypatch): + """New action declarations authorize trusted targets and reject unrelated ones.""" + # Given a future action whose definition declares its mutation target + monkeypatch.setitem(executor._REQUIRED_ACTION_FIELDS, "future-mutation", {"type", "marker", "issue"}) + monkeypatch.setitem(executor._ACTION_DEFINITIONS, "future-mutation", {"targets": ("issue",)}) + result = _plan([{"type": "future-mutation", "marker": "triage-security:future", "issue": issue}]) + + # When preflight checks the new target declaration + if issue == "TC-42": + executor._preflight_plan(result, _trusted_input()) + else: + with pytest.raises(executor.ActionError, match="unauthorized action target: TC-43"): + executor._preflight_plan(result, _trusted_input()) + + # Then preflight performs no Jira operation in either case + assert recorder.calls == [] + + +@pytest.mark.parametrize("action", [ + {"type": "field-edit", "marker": "triage-security:untrusted", "issue": "TC-43", "fields": {"labels": ["ai-cve-triaged"]}}, + {"type": "status-transition", "marker": "triage-security:untrusted", "issue": "TC-43", "status": "In Progress"}, + {"type": "comment", "marker": "triage-security:untrusted", "issue": "TC-43", "body_adf": {"type": "doc", "version": 1, "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Synthetic triage."}]}]}}, + {"type": "resolve-reference", "marker": "triage-security:untrusted", "ref": "alias", "issue": "TC-43"}, + {"type": "link", "marker": "triage-security:untrusted", "link_type": "Related", "inward": "TC-43", "outward": "TC-42"}, + {"type": "link", "marker": "triage-security:untrusted", "link_type": "Related", "inward": "TC-42", "outward": "TC-43"}, + {**_remediation_action(), "marker": "triage-security:untrusted", "project": "OTHER"}, +]) +@pytest.mark.parametrize("marked", [False, True]) +def test_late_untrusted_actions_reject_the_entire_plan_before_retry_suppression(action, marked, recorder): + """Neither a valid first action nor a retry marker permits unrelated targets.""" + # Given a valid first mutation and a late action outside the trusted scope + result = _plan([ + {"type": "field-edit", "marker": "triage-security:first", "issue": "TC-42", "fields": {"labels": ["first-action"]}}, + action, + ]) + if action["type"] == "resolve-reference": + result["actions"].append({"type": "link", "marker": "triage-security:alias-link", + "link_type": "Related", "inward": "TC-42", "outward": "{{alias.key}}"}) + trusted = _trusted_input(markers=["triage-security:untrusted"] if marked else []) + trusted["issue"]["fields"] = {"labels": ["ai-cve-triaged"]} + trusted["issue"]["status"] = "In Progress" + + # When the executor preflights all actions before markers or state suppress them + with pytest.raises(executor.ActionError): + executor.execute_plan(result, trusted) + + # Then no mutation or read occurred for the otherwise valid prefix + assert recorder.calls == [] + + +@pytest.mark.parametrize("marked", [False, True]) +def test_invalid_follow_up_rejects_before_existing_remediation_digest_repair(marked, recorder): + """An invalid late link prevents digest repair on an earlier retry action.""" + # Given existing remediation missing its digest and a link to an unknown ref + result = _plan([_remediation_action(), { + "type": "link", "marker": "triage-security:link", "link_type": "Depend", + "inward": "TC-42", "outward": "{{unknown.key}}", + }]) + trusted = _trusted_input( + markers=["triage-security:remediation", "triage-security:link"] if marked else [], + remediation=[{"key": "TC-777", "summary": "Fix CVE", "labels": ["ai-generated-jira"], "comments": []}], + ) + + # When complete preflight discovers the unresolved reference + with pytest.raises(executor.ActionError): + executor.execute_plan(result, trusted) + + # Then even digest reads and repair writes are absent + assert recorder.calls == [] + + +@pytest.mark.parametrize("binding", [ + {"type": "resolve-reference", "marker": "triage-security:rebind", "ref": "remediation", "issue": "TC-43"}, + {**_remediation_action(), "marker": "triage-security:rebind"}, +]) +def test_reference_rebinding_cannot_redirect_generated_remediation(binding, recorder): + """Generated reference names cannot be rebound by a later sandbox action.""" + # Given a valid creation followed by a conflicting reference definition + result = _plan([_remediation_action(), binding, { + "type": "link", "marker": "triage-security:link", "link_type": "Depend", + "inward": "TC-42", "outward": "{{remediation.key}}", + }]) + + # When preflight tracks reference definitions in order + with pytest.raises(executor.ActionError): + executor.execute_plan(result, _trusted_input(related=[{"key": "TC-43"}])) + + # Then the invalid plan cannot create or digest its first task + assert recorder.calls == [] + + +@pytest.mark.parametrize("context", [ + {"related": [{"key": "TC-43"}]}, + {"searches": [{"purpose": "same-cve-siblings", "issues": [{"key": "TC-43"}]}]}, + {"searches": [{"purpose": "cross-cve-overlap", "issues": [{"key": "TC-43"}]}]}, + {"searches": [{"purpose": "preemptive-remediation", "issues": [{"key": "TC-43"}]}]}, + {"remediation": [{"key": "TC-43", "labels": ["ai-generated-jira"]}]}, +]) +def test_prefetched_triage_relationships_authorize_related_targets(context, recorder): + """Trusted related targets support reconciliation, comments, and both link ends.""" + # Given a prefetched companion or remediation and actions targeting it + result = _plan([ + {"type": "resolve-reference", "marker": "triage-security:related-ref", "ref": "related", "issue": "TC-43"}, + {"type": "field-edit", "marker": "triage-security:related-labels", "issue": "TC-43", "fields": {"labels": ["ai-cve-triaged"]}}, + {"type": "status-transition", "marker": "triage-security:related-status", "issue": "{{related.key}}", "status": "In Progress"}, + {"type": "comment", "marker": "triage-security:related-comment", "issue": "{{related.key}}", "body_adf": {"type": "doc", "version": 1, "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Synthetic triage."}]}]}}, + {"type": "link", "marker": "triage-security:related-link", "link_type": "Related", "inward": "{{related.key}}", "outward": "TC-42"}, + ]) + trusted = _trusted_input(**context) + trusted["issue"]["fields"] = {"labels": ["ai-cve-triaged"]} + trusted["issue"]["status"] = "In Progress" + + # When trusted relationships establish membership independently of the plan + executor.execute_plan(result, trusted) + + # Then reconciliation and cross-stream operations retain their targets and order + assert recorder.calls[0] == ("field-edit", "TC-43", {"labels": ["ai-cve-triaged"]}) + assert recorder.calls[1:3] == [("get-transitions", "TC-43"), ("transition", "TC-43", "31")] + assert recorder.calls[3][0:2] == ("comment", "TC-43") + assert recorder.calls[4] == ("link", "TC-43", "TC-42", "Related") + + +def test_related_status_transition_retry_skips_completed_transition_and_continues(recorder, monkeypatch): + """Replaying a completed related transition does not block later actions.""" + # Given an authorized related target whose transition disappears after success + result = _plan([ + {"type": "status-transition", "marker": "triage-security:related-status", "issue": "TC-43", "status": "In Progress"}, + {"type": "field-edit", "marker": "triage-security:labels", "issue": "TC-42", "fields": {"labels": ["ai-cve-triaged"]}}, + ]) + trusted = _trusted_input(related=[{"key": "TC-43"}]) + executor.execute_plan(result, trusted) + assert recorder.calls == [ + ("get-transitions", "TC-43"), ("transition", "TC-43", "31"), + ("field-edit", "TC-42", {"labels": ["ai-cve-triaged"]}), + ] + recorder.calls.clear() + + def get_transitions(issue): + """Return no transition once the related target has reached its status.""" + recorder.calls.append(("get-transitions", issue)) + return [] + + def get_issue(issue, fields="*all"): + """Return the related issue's current Jira status on replay.""" + recorder.calls.append(("get-issue", issue, fields)) + return {"fields": {"status": {"name": "In Progress"}}} + + monkeypatch.setattr(executor._jira_mod, "get_transitions", get_transitions) + monkeypatch.setattr(executor._jira_mod, "get_issue", get_issue) + + # When the identical unmarked plan is replayed + executor.execute_plan(result, trusted) + + # Then Jira confirms the target state, no transition is written, and execution continues + assert recorder.calls == [ + ("get-transitions", "TC-43"), ("get-issue", "TC-43", "status"), + ("field-edit", "TC-42", {"labels": ["ai-cve-triaged"]}), + ] + + +def test_related_unavailable_transition_in_different_status_still_fails(recorder, monkeypatch): + """A missing transition is not a no-op unless its target status is reached.""" + # Given an authorized target with no transition and a different current status + result = _plan([ + {"type": "status-transition", "marker": "triage-security:related-status", "issue": "TC-43", "status": "In Progress"}, + {"type": "field-edit", "marker": "triage-security:labels", "issue": "TC-42", "fields": {"labels": ["ai-cve-triaged"]}}, + ]) + + def get_transitions(issue): + """Record the unavailable transition lookup.""" + recorder.calls.append(("get-transitions", issue)) + return [] + + def get_issue(issue, fields="*all"): + """Return a related issue still awaiting the requested transition.""" + recorder.calls.append(("get-issue", issue, fields)) + return {"fields": {"status": {"name": "New"}}} + + monkeypatch.setattr(executor._jira_mod, "get_transitions", get_transitions) + monkeypatch.setattr(executor._jira_mod, "get_issue", get_issue) + + # When the missing transition cannot be justified by the current Jira state + with pytest.raises(executor.ActionError, match="no transition named In Progress for TC-43"): + executor.execute_plan(result, _trusted_input(related=[{"key": "TC-43"}])) + + # Then execution fails before any mutation or later action + assert recorder.calls == [ + ("get-transitions", "TC-43"), ("get-issue", "TC-43", "status"), + ] + + +def test_unrelated_search_purpose_cannot_expand_trusted_scope(recorder): + """An arbitrary prefetched search is not a documented triage relationship.""" + # Given a target present only in a search unrelated to triage + result = _plan([{"type": "field-edit", "marker": "triage-security:labels", + "issue": "TC-43", "fields": {"labels": ["ai-cve-triaged"]}}]) + trusted = _trusted_input(searches=[{"purpose": "all-project-issues", "issues": [{"key": "TC-43"}]}]) + + # When validating documented relationships + with pytest.raises(executor.ActionError): + executor.execute_plan(result, trusted) + + # Then the search does not authorize a mutation + assert recorder.calls == [] + + +@pytest.mark.parametrize("configuration", [None, {}, {"project_key": ""}, {"project_key": "OTHER"}]) +def test_remediation_creation_requires_the_trusted_project(configuration, recorder): + """Missing or contradictory project context cannot authorize task creation.""" + # Given a schema-valid creation with no matching trusted project + result = _plan([_remediation_action()]) + trusted = _trusted_input() + trusted["configuration"] = configuration + + # When the runner authorizes the proposed creation + with pytest.raises(executor.ActionError): + executor.execute_plan(result, trusted) + + # Then neither creation nor digest reads reach Jira + assert recorder.calls == [] + + +def test_cli_rejects_untrusted_targets_with_nonzero_status(tmp_path, recorder, capsys): + """Real CLI validation exits non-zero without Jira calls for invalid targets.""" + # Given readable runner files with an unauthorized mutation target + result = _plan([{"type": "field-edit", "marker": "triage-security:labels", + "issue": "TC-43", "fields": {"labels": ["ai-cve-triaged"]}}]) + result_path, input_path = tmp_path / "result.json", tmp_path / "input.json" + result_path.write_text(json.dumps(result)) + input_path.write_text(json.dumps(_trusted_input())) + + # When the CLI executes its real authorization path + status = executor.main([str(result_path), str(input_path)]) + + # Then it fails loudly without reporting success or making Jira calls + assert status == 1 + output = capsys.readouterr() + assert "unauthorized action target" in output.err + assert "successfully" not in output.out + assert recorder.calls == [] + + +@pytest.mark.parametrize("action", [ + { + "type": "link", + "marker": "triage-security:unsupported-link-type", + "link_type": "Unsupported", + "inward": "TC-42", + "outward": "TC-43", + }, + { + "type": "field-edit", + "marker": "triage-security:empty-field-edit", + "issue": "TC-42", + "fields": {}, + }, + { + "type": "comment", + "marker": "triage-security:malformed-adf-content", + "issue": "TC-42", + "body_adf": {"type": "doc", "version": 1, "content": [{}]}, + }, +]) +def test_schema_invalid_plan_is_rejected_before_any_jira_operation(action, recorder): + """The trusted runner rejects schema-invalid sandbox output before mutation.""" + # Given a result that passed only the sandbox's untrusted validation path + result = _plan([action]) + + # When the trusted runner receives the invalid result with authorization + with pytest.raises(executor.ActionError, match="result schema validation failed"): + executor.execute_plan(result, _trusted_input()) + + # Then it makes no Jira request, including read-before-write operations + assert recorder.calls == [] + + +def test_unauthorized_mutating_plan_fails_without_jira_calls(recorder): + """A mutating result cannot bypass an absent trusted authorization grant.""" + # Given a plan with a valid-looking mutation but no runner authorization + result = _plan([{ + "type": "field-edit", "marker": "triage-security:labels", "issue": "TC-42", + "fields": {"labels": ["ai-cve-triaged"]}, + }]) + + # When execution checks the trusted bundle + with pytest.raises(executor.ActionError, match="not authorized"): + executor.execute_plan(result, _trusted_input(authorized=False)) + + # Then authorization fails before any Jira operation + assert recorder.calls == [] + + +def test_authorized_actions_execute_in_order_with_footnoted_comments(recorder): + """Authorized field, transition, and comment actions retain plan order and footer.""" + # Given an authorized plan with all non-creation mutation forms + result = _plan([ + {"type": "field-edit", "marker": "triage-security:fields", "issue": "TC-42", "fields": {"assignee": {"id": "owner"}, "labels": ["ai-cve-triaged"], "resolution": {"name": "Done"}}}, + {"type": "status-transition", "marker": "triage-security:status", "issue": "TC-42", "status": "In Progress"}, + {"type": "comment", "marker": "triage-security:summary", "issue": "TC-42", "body_adf": {"type": "doc", "version": 1, "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Triage complete."}]}]}}, + ]) + + # When the trusted runner executes the plan + executor.execute_plan(result, _trusted_input()) + + # Then the operations run deterministically and the comment has its audit footer + assert [call[0] for call in recorder.calls] == ["field-edit", "get-transitions", "transition", "comment"] + comment = recorder.calls[-1] + assert comment[3] == "" + assert "---" in comment[2] + assert "sdlc-workflow/triage-security" in comment[2] + assert "v{}".format(executor._plugin_version()) in comment[2] + + +def test_remediation_digest_precedes_links_and_resolves_references(recorder): + """Created remediation tasks are registered and digested before dependent links.""" + # Given a remediation task followed by a link to its generated reference + result = _plan([ + {"type": "remediation-task", "marker": "triage-security:create-remediation", "ref": "remediation-1", "project": "TC", "summary": "Fix CVE", "description_adf": {"type": "doc", "version": 1, "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Fix CVE."}]}]}, "labels": ["security-preemptive"], "priority": "Major", "fix_versions": ["1.2"]}, + {"type": "link", "marker": "triage-security:link-remediation", "link_type": "Depend", "inward": "TC-42", "outward": "{{remediation-1.key}}"}, + ]) + + # When the authorized actions run + executor.execute_plan(result, _trusted_input()) + + # Then task creation carries inherited metadata, its digest is first, and its key resolves in the link + assert recorder.calls[0][0] == "remediation-task" + assert recorder.calls[0][1]["labels"] == ["ai-generated-jira", "security-preemptive"] + assert recorder.calls[0][1]["priority"] == "Major" + assert recorder.calls[0][1]["fix_versions"] == ["1.2"] + assert recorder.calls[1] == ("get-issue", "TC-9001") + assert recorder.calls[2][0] == "digest" + stored = recorder.get_issue("TC-9001")["fields"]["description"] + expected = hashlib.sha256(json.dumps(stored, separators=(",", ":")).encode("utf-8")).hexdigest() + assert "Description digest: sha256-adf:{}".format(expected) in json.dumps(recorder.calls[2][3]) + assert recorder.calls[3] == ("link", "TC-42", "TC-9001", "Depend") + + +def test_existing_marker_and_remediation_skip_retry_duplicates(recorder): + """Stable runner markers and existing remediation state make retries idempotent.""" + # Given actions already applied by a previous trusted runner attempt + result = _plan([ + {"type": "field-edit", "marker": "triage-security:fields", "issue": "TC-42", "fields": {"labels": ["ai-cve-triaged"]}}, + {"type": "remediation-task", "marker": "triage-security:create-remediation", "ref": "remediation-1", "project": "TC", "summary": "Fix CVE", "description_adf": {"type": "doc", "version": 1, "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Fix CVE."}]}]}, "labels": []}, + ]) + trusted = _trusted_input( + markers=["triage-security:fields"], + remediation=[{"key": "TC-777", "summary": "Fix CVE", "labels": ["ai-generated-jira"], "description": {"type": "doc", "version": 1, "content": []}, "comments": [{"body": "[sdlc-workflow] Description digest: sha256-adf:already"}]}], + ) + + # When the same plan is retried + registry = executor.execute_plan(result, trusted) + + # Then no duplicate writes occur and the existing task resolves the reference + assert recorder.calls == [] + assert registry["remediation-1"]["key"] == "TC-777" + + +def test_marked_remediation_still_populates_references_for_dependent_actions(recorder): + """A skipped remediation action still resolves its reference before later links.""" + # Given a retry marker and the prior task recorded in trusted idempotency state + result = _plan([ + {"type": "remediation-task", "marker": "triage-security:create-remediation", "ref": "remediation-1", "project": "TC", "summary": "Fix CVE", "description_adf": {"type": "doc", "version": 1, "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Fix CVE."}]}]}, "labels": []}, + {"type": "link", "marker": "triage-security:link", "link_type": "Depend", "inward": "TC-42", "outward": "{{remediation-1.key}}"}, + ]) + trusted = _trusted_input( + markers=["triage-security:create-remediation"], + remediation=[{"key": "TC-777", "summary": "Fix CVE", "labels": ["ai-generated-jira"], "description": {"type": "doc", "version": 1, "content": []}, "comments": [{"body": "[sdlc-workflow] Description digest: sha256-adf:already"}]}], + ) + + # When the marked plan is retried + registry = executor.execute_plan(result, trusted) + + # Then its generated reference remains usable by the dependent link + assert registry["remediation-1"]["key"] == "TC-777" + assert recorder.calls == [("link", "TC-42", "TC-777", "Depend")] + + +def test_marked_reference_binding_still_resolves_dependent_actions(recorder): + """A marker never suppresses an in-memory reference binding needed by later actions.""" + # Given a reference binding whose marker was persisted by an earlier attempt + result = _plan([ + {"type": "resolve-reference", "marker": "triage-security:known-task", "ref": "known-task", "issue": "TC-777"}, + {"type": "link", "marker": "triage-security:link", "link_type": "Related", "inward": "TC-42", "outward": "{{known-task.key}}"}, + ]) + + # When the retry skips persisted mutations + registry = executor.execute_plan(result, _trusted_input(markers=["triage-security:known-task"], related=[{"key": "TC-777"}])) + + # Then the binding survives and the dependent link uses the Jira key + assert registry["known-task"]["key"] == "TC-777" + assert recorder.calls == [("link", "TC-42", "TC-777", "Related")] + + +def test_existing_remediation_without_digest_is_digested_before_follow_up_actions(recorder): + """A retry repairs a partial creation by posting its missing digest before linking.""" + # Given a prior remediation task that was created before its digest could post + result = _plan([ + {"type": "remediation-task", "marker": "triage-security:create-remediation", "ref": "remediation-1", "project": "TC", "summary": "Fix CVE", "description_adf": {"type": "doc", "version": 1, "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Fix CVE."}]}]}, "labels": []}, + {"type": "link", "marker": "triage-security:link", "link_type": "Depend", "inward": "TC-42", "outward": "{{remediation-1.key}}"}, + ]) + trusted = _trusted_input(remediation=[{"key": "TC-777", "summary": "Fix CVE", "labels": ["ai-generated-jira"], "description": {"type": "doc", "version": 1, "content": []}, "comments": []}]) + + # When retry execution discovers the prior task + executor.execute_plan(result, trusted) + + # Then it posts the stored-description digest before the dependent link + assert [call[0] for call in recorder.calls] == ["get-issue", "digest", "link"] + + +def test_existing_issue_state_skips_retried_field_status_and_link_actions(recorder): + """A real-shaped Jira link snapshot prevents duplicate retry mutations.""" + # Given a later pre-script snapshot with existing state and Jira's single far-end link shape + result = _plan([ + {"type": "field-edit", "marker": "triage-security:fields", "issue": "TC-42", "fields": {"labels": ["ai-cve-triaged"]}}, + {"type": "status-transition", "marker": "triage-security:status", "issue": "TC-42", "status": "In Progress"}, + {"type": "link", "marker": "triage-security:link", "link_type": "Depend", "inward": "TC-42", "outward": "TC-9001"}, + ]) + trusted = _trusted_input(related=[{"key": "TC-9001"}]) + trusted["issue"] = { + "key": "TC-42", + "status": "In Progress", + "fields": { + "labels": ["ai-cve-triaged"], + "issuelinks": [{ + "type": {"name": "Depend"}, + "outwardIssue": {"key": "TC-9001"}, + }], + }, + } + + # When the identical plan is retried without markers + executor.execute_plan(result, trusted) + + # Then the current Jira state still prevents duplicate mutations + assert recorder.calls == [] + + +def test_existing_object_fields_skip_retried_field_edits(recorder): + """Full Jira field objects match compact assignee and resolution retry values.""" + # Given a retry action and its full Jira snapshot field representations + result = _plan([{ + "type": "field-edit", "marker": "triage-security:fields", "issue": "TC-42", + "fields": {"assignee": {"id": "owner"}, "resolution": {"name": "Done"}}, + }]) + trusted = _trusted_input() + trusted["issue"]["fields"] = { + "assignee": {"accountId": "owner", "displayName": "Owner"}, + "resolution": {"id": "10000", "name": "Done"}, + } + + # When the same object-valued field edit is retried + executor.execute_plan(result, trusted) + + # Then the executor avoids clobbering the current Jira field values + assert recorder.calls == [] + + +def test_empty_field_edit_is_not_treated_as_already_applied(): + """An empty field map never becomes idempotent through all([]).""" + # Given an otherwise valid field-edit action with no values to compare + action = {"type": "field-edit", "marker": "triage-security:empty", "issue": "TC-42", "fields": {}} + + # When idempotency evaluates the empty map + already_applied = executor._already_applied(action, _trusted_input()) + + # Then it remains eligible for validation rather than appearing applied + assert already_applied is False + + +def test_other_object_fields_require_an_exact_snapshot_match(): + """Compact matching does not hide changed values in unrelated object fields.""" + # Given an object-valued custom field whose value differs from the snapshot + action = { + "type": "field-edit", "marker": "triage-security:custom", "issue": "TC-42", + "fields": {"customfield_12345": {"name": "Risk", "value": "new"}}, + } + trusted = _trusted_input() + trusted["issue"] = {"fields": {"customfield_12345": {"name": "Risk", "value": "old"}}} + + # When idempotency compares the custom object field + already_applied = executor._already_applied(action, trusted) + + # Then the changed value remains eligible for an update + assert already_applied is False + + +def test_unresolved_or_malformed_actions_fail_before_writes(recorder): + """Unknown references and malformed actions cannot reach Jira as mutations.""" + # Given independently invalid plans + unresolved = _plan([{"type": "link", "marker": "triage-security:bad-link", "link_type": "Related", "inward": "TC-42", "outward": "{{unknown.key}}"}]) + malformed = _plan([{"type": "field-edit", "marker": "triage-security:bad-fields", "issue": "TC-42"}]) + + # When execution validates each action plan + with pytest.raises(executor.ActionError): + executor.execute_plan(unresolved, _trusted_input()) + with pytest.raises(executor.ActionError): + executor.execute_plan(malformed, _trusted_input()) + + # Then no partial mutation is emitted for invalid input + assert recorder.calls == [] + + +def test_partial_failure_propagates_without_success_message(recorder, monkeypatch, capsys): + """A failed Jira write aborts execution and never prints a completion claim.""" + # Given an authorized update whose Jira client operation fails + result = _plan([{"type": "field-edit", "marker": "triage-security:fields", "issue": "TC-42", "fields": {"labels": ["ai-cve-triaged"]}}]) + monkeypatch.setattr(executor._jira_mod, "update_issue", lambda *_args: (_ for _ in ()).throw(RuntimeError("network down"))) + + # When the executor receives the operational error + with pytest.raises(RuntimeError, match="network down"): + executor.execute_plan(result, _trusted_input()) + + # Then it does not incorrectly report success + assert "successfully" not in capsys.readouterr().out.lower() + + +def test_main_returns_nonzero_for_executor_failures(tmp_path, monkeypatch, capsys): + """The command entry point reports an executor failure with a non-zero status.""" + # Given readable runner files whose execution raises an operational error + result_path = tmp_path / "result.json" + input_path = tmp_path / "input.json" + result_path.write_text(json.dumps(_plan([{"type": "report-only", "marker": "triage-security:report"}], "report-only"))) + input_path.write_text(json.dumps(_trusted_input(False))) + monkeypatch.setattr(executor, "execute_plan", lambda *_args: (_ for _ in ()).throw(RuntimeError("Jira unavailable"))) + + # When the CLI runs the executor + status = executor.main([str(result_path), str(input_path)]) + + # Then it exposes failure rather than claiming completion + assert status == 1 + assert "Jira unavailable" in capsys.readouterr().err + + +def test_post_script_rejects_an_iteration_result_outside_its_run_directory(tmp_path): + """The trusted post-script refuses a symlinked result that escapes the run directory.""" + # Given a runner directory whose selected iteration result resolves elsewhere + script = os.path.join(SCRIPT_DIR, "post-triage-security.sh") + run_directory = tmp_path / "run" + (run_directory / "pre").mkdir(parents=True) + (run_directory / "pre" / "triage-security-input.json").write_text(json.dumps(_trusted_input(False))) + iteration_output = run_directory / "iteration-1" / "output" + iteration_output.mkdir(parents=True) + outside_result = tmp_path / "outside-result.json" + outside_result.write_text(json.dumps(_plan([{"type": "report-only", "marker": "triage-security:report"}], "report-only"))) + (iteration_output / "agent-result.json").symlink_to(outside_result) + + # When the runner attempts to select the post-validation result + result = subprocess.run( + ["bash", script], cwd=run_directory, + env={"PATH": os.environ["PATH"], "FULLSEND_RUN_DIR": str(run_directory)}, + capture_output=True, text=True, + ) + + # Then it fails before handing an untrusted path to the executor + assert result.returncode != 0 + assert "escapes" in result.stderr + + +def test_post_script_selects_an_in_run_result_from_a_different_working_directory(tmp_path): + """The post-script resolves validated output from FULLSEND_RUN_DIR, not its CWD.""" + # Given the validated output and matching pre-script authorization in one run + script = os.path.join(SCRIPT_DIR, "post-triage-security.sh") + run_directory = tmp_path / "run" + (run_directory / "pre").mkdir(parents=True) + (run_directory / "pre" / "triage-security-input.json").write_text(json.dumps(_trusted_input(False))) + iteration_output = run_directory / "iteration-1" / "output" + iteration_output.mkdir(parents=True) + (iteration_output / "agent-result.json").write_text( + json.dumps(_plan([{"type": "report-only", "marker": "triage-security:report"}], "report-only"))) + + # When the trusted post-script runs outside its Fullsend run directory + result = subprocess.run( + ["bash", script], cwd=tmp_path, + env={"PATH": os.environ["PATH"], "FULLSEND_RUN_DIR": str(run_directory)}, + capture_output=True, text=True, + ) + + # Then it completes without needing Jira credentials or performing a mutation + assert result.returncode == 0, result.stderr + assert "completed successfully" in result.stdout diff --git a/plugins/sdlc-workflow/scripts/test_fullsend_gate_eval.py b/plugins/sdlc-workflow/scripts/test_fullsend_gate_eval.py new file mode 100644 index 000000000..01c6a45e5 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/test_fullsend_gate_eval.py @@ -0,0 +1,853 @@ +"""Deterministic native-eval contracts; these tests never execute an agent.""" + +import hashlib +import importlib.metadata +import importlib.util +import io +import json +import os +from pathlib import Path +import shutil +import signal +import subprocess +import sys +import tarfile + +import pytest +import yaml + + +ROOT = Path(__file__).resolve().parents[3] +SUITE = ROOT / "evals/fullsend/triage-security" +FIXTURES = ROOT / "evals/triage-security/files" + + +def load_script(path): + """Import test tooling without starting the external model runtime.""" + assert path.is_file(), f"Missing native eval tooling: {path}" + spec = importlib.util.spec_from_file_location(path.stem.replace("-", "_"), path) + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def synthetic_fixture_root(path): + """SYNTHETIC TEST DATA — isolated unchanged retained inputs, never live data.""" + fixtures = path / "evals/triage-security/files" + fixtures.mkdir(parents=True) + for name in ["fullsend-gate-interactive-config.md", "fullsend-invalid-trusted-input.md", + "fullsend-report-only-trusted-input.json"]: + shutil.copy2(FIXTURES / name, fixtures / name) + return path + + +def synthetic_judge_summary(value=True): + """SYNTHETIC STATIC RECORDS — summary contract only, never live eval evidence.""" + cases = {} + for case, count in zip(["033-absent", "034-empty", "035-malformed", "036-valid"], [4, 5, 5, 7]): + cases[case] = {} + for index in range(1, 8): + condition = f'annotations.get("assertion_count", 0) > {index - 1}' + cases[case][f"assertion_{index}"] = { + "value": value if index <= count else None, + "rationale": "SYNTHETIC NOT A JUDGMENT" if index <= count else f"Skipped: condition '{condition}' is false", + "judge_type": "llm", + } + return {"run_id": "synthetic", "per_case": cases} + + +@pytest.mark.parametrize("value", [True, False]) +def test_summary_integrity_accepts_complete_boolean_results_without_grading(tmp_path, value): + """Completeness accepts actual False outcomes; upstream alone owns thresholds.""" + common = load_script(ROOT / "evals/fullsend/run.py") + # Given 21 explicitly synthetic Boolean outcomes and seven legitimate skips + path = tmp_path / "summary.yaml" + raw = yaml.safe_dump(synthetic_judge_summary(value)).encode() + path.write_bytes(raw) + # When checking integrity without any scorer or runtime invocation + common.validate_summary(tmp_path, "synthetic") + # Then raw outcomes/rationales remain byte-for-byte intact, including False + assert path.read_bytes() == raw + + +@pytest.mark.parametrize("defect", [ + "missing-case", "extra-case", "wrong-case", "missing-assertion", "extra-assertion", + "missing-value", "null", "integer", "string", "error", "applicable-skip", + "false-with-skip", "nonapplicable-boolean", "nonapplicable-error", "condition-error", + "wrong-run", "not-mapping", "case-not-mapping", "result-not-mapping", +]) +def test_summary_integrity_rejects_incomplete_or_ambiguous_results(tmp_path, defect): + """Static malformed summary records must never allow aggregate-only CI success.""" + common = load_script(ROOT / "evals/fullsend/run.py") + # Given adversarial synthetic metadata, never generated execution/judge evidence + summary = synthetic_judge_summary() + cases = summary["per_case"] + results = cases["033-absent"] + result = results["assertion_1"] + if defect == "missing-case": + del cases["036-valid"] + elif defect == "extra-case": + cases["unexpected"] = results + elif defect == "wrong-case": + cases["wrong-valid"] = cases.pop("036-valid") + elif defect == "missing-assertion": + del results["assertion_1"] + elif defect == "extra-assertion": + results["assertion_8"] = result + elif defect == "missing-value": + del result["value"] + elif defect in ["null", "integer", "string"]: + result["value"] = {"null": None, "integer": 1, "string": "true"}[defect] + elif defect == "error": + result["error"] = "SYNTHETIC scorer failure, even with a Boolean value" + elif defect == "applicable-skip": + results["assertion_1"] = dict(results["assertion_5"]) + elif defect == "false-with-skip": + result.update(value=False, rationale="Skipped: synthetic invalid applicable skip") + elif defect == "nonapplicable-boolean": + results["assertion_5"]["value"] = True + elif defect == "nonapplicable-error": + results["assertion_5"]["error"] = "SYNTHETIC condition failure" + elif defect == "condition-error": + results["assertion_5"]["rationale"] = "Condition error: SYNTHETIC failure" + elif defect == "wrong-run": + summary["run_id"] = "different-run" + elif defect == "not-mapping": + summary = [] + elif defect == "case-not-mapping": + cases["033-absent"] = [] + else: + results["assertion_1"] = [] + path = tmp_path / "summary.yaml" + raw = yaml.safe_dump(summary).encode() + path.write_bytes(raw) + # When validating the preserved upstream artifact, fail closed without rewriting + with pytest.raises(ValueError, match="summary"): + common.validate_summary(tmp_path, "synthetic") + assert path.read_bytes() == raw + + +@pytest.mark.parametrize("raw", [None, b"[invalid", b"run_id: synthetic\nrun_id: synthetic\n", + b"per_case:\n 033-absent: {}\n 033-absent: {}\n"]) +def test_summary_integrity_rejects_missing_invalid_or_duplicate_yaml(tmp_path, raw): + """Missing, unparsable and duplicate-key artifacts are not unambiguous results.""" + common = load_script(ROOT / "evals/fullsend/run.py") + # Given explicitly synthetic raw YAML or no summary artifact + path = tmp_path / "summary.yaml" + if raw is not None: + path.write_bytes(raw) + with pytest.raises(ValueError, match="summary"): + common.validate_summary(tmp_path, "synthetic") + assert path.read_bytes() == raw if raw is not None else not path.exists() + + +def test_ordinary_evals_preserve_baseline_and_exclude_native_cases(): + """Ordinary Claude evals must not accidentally execute native-only scenarios.""" + # Given the manifests consumed by the existing hosted run-evals + triage = json.loads((ROOT / "evals/triage-security/evals.json").read_text())["evals"] + verify = json.loads((ROOT / "evals/verify-pr/evals.json").read_text())["evals"] + # Then native cases are separate and every retained object is unchanged + assert [c["id"] for c in triage] == list(range(1, 33)) + assert sum(len(c["assertions"]) for c in triage) == 164 + assert len(verify) == 6 and sum(len(c["assertions"]) for c in verify) == 68 + # Bootstrap main and reviewed PR299 contain different pre-existing triage + # assertion objects. Accept only those two immutable baselines, never edit + # the active ordinary manifests to match the native branch's historical hash. + # Canonical complete-object digests keep this portable to a shallow checkout: + # main ab20bee6, native aa15d776, and PR299 d83ee90b (TC4636 prompt repair). + for cases, digests in [ + (triage, {"205eeca4b564c0483c919be0951e50b3d5510f981278c61fa1c47468af5fe76d", + "b3f9e9d4f4ab1eb92f053c0c0e4199a36eabb12ab9589e9d27e9c59509eee501"}), + (verify, {"251863edaed38f0b133c0cf0981ddffe80692f5d0655b51f7bfe214be38f69b1", + "cf587edf1e94e6a1a210d5c97788f33b3336a0818a86551e364ec056bd1a3be1"}), + ]: + assert hashlib.sha256(json.dumps(cases, sort_keys=True, separators=(",", ":")).encode()).hexdigest() in digests + assert not (FIXTURES / "fullsend-gate-tools.py").exists() + + +@pytest.mark.parametrize("scenario, fragment, fixture", [ + ("absent", "unset FULLSEND_OUTPUT_DIR", None), + ("empty", "export FULLSEND_OUTPUT_DIR=''", None), + ("malformed", None, "fullsend-invalid-trusted-input.md"), + ("valid", None, "fullsend-report-only-trusted-input.json"), +]) +def test_pre_script_prepares_native_mounts_without_live_fetch(tmp_path, scenario, fragment, fixture): + """A native pre-script must inject distinct states before runtime startup.""" + # Given only synthetic fixture paths and the native pre-script environment + script = SUITE / "prepare-fixture.py" + assert script.is_file(), "Native synthetic pre-script is missing" + root = synthetic_fixture_root(tmp_path) + environment = dict(os.environ, TC6677_SCENARIO=scenario, TC6677_REPO_ROOT=str(root)) + environment.pop("FULLSEND_RUN_DIR", None) # Native pre-script gets no such var. + # When preparing the host files (not running Fullsend or an agent) + result = subprocess.run([sys.executable, str(script)], env=environment, + capture_output=True, text=True, check=False) + assert result.returncode == 0, result.stderr + # Then exact input bytes and only the intended gate injection are mounted + gate = (tmp_path / "pre/tc-6677-gate.env").read_text() + assert gate.startswith("# SYNTHETIC TEST DATA") + assert gate.splitlines()[1:] == ([fragment] if fragment else []) + mounted = tmp_path / "pre/triage-security-input.json" + if fixture: + expected = (FIXTURES / fixture).read_bytes() + if fixture.endswith(".md"): + expected = expected.split(b"```json\n", 1)[1].split(b"\n```", 1)[0] + assert mounted.read_bytes() == expected + else: + assert not mounted.exists() + assert sorted(p.name for p in (tmp_path / "pre").iterdir()) == ( + ["tc-6677-gate.env", "triage-security-input.json"] if fixture else ["tc-6677-gate.env"]) + assert not result.stdout and not result.stderr + + +def test_pre_script_refuses_stale_inputs(tmp_path): + """Fixture failure must stay visible instead of silently reusing a prior run.""" + # Given a previously populated native pre directory + script = SUITE / "prepare-fixture.py" + assert script.is_file(), "Native synthetic pre-script is missing" + (tmp_path / "pre").mkdir() + (tmp_path / "pre/triage-security-input.json").write_bytes(b"stale") + environment = dict(os.environ, TC6677_SCENARIO="absent", TC6677_REPO_ROOT=str(tmp_path)) + environment.pop("FULLSEND_RUN_DIR", None) + # When an absent case would otherwise inherit a stale nonempty input + result = subprocess.run([sys.executable, str(script)], env=environment, + capture_output=True, text=True, check=False) + # Then it fails without rewriting that evidence + assert result.returncode != 0 and "stale" in result.stderr.lower() + assert (tmp_path / "pre/triage-security-input.json").read_bytes() == b"stale" + + +@pytest.mark.parametrize("failure", ["blocked-pre-directory", "missing-valid-input"]) +def test_pre_script_fails_before_mounts_when_required_fixture_cannot_be_prepared(tmp_path, failure): + """Deferred optional mounts cannot convert a real preparation failure into success.""" + # Given synthetic fixture failure, no FULLSEND_RUN_DIR and no agent/runtime + root = synthetic_fixture_root(tmp_path) + if failure == "blocked-pre-directory": + (root / "pre").write_text("SYNTHETIC TEST DATA — blocks required gate delivery") + else: + (root / "evals/triage-security/files/fullsend-report-only-trusted-input.json").unlink() + environment = dict(os.environ, TC6677_SCENARIO="valid", TC6677_REPO_ROOT=str(root)) + environment.pop("FULLSEND_RUN_DIR", None) + # When the actual fixture pre-script cannot produce its required files + result = subprocess.run([sys.executable, str(SUITE / "prepare-fixture.py")], env=environment, + capture_output=True, text=True, check=False) + # Then nonzero propagates to Fullsend's pre-script abort; no gate can be mounted + assert result.returncode == 1 and "preparation failed" in result.stderr + expected_path = root / ("pre" if failure == "blocked-pre-directory" else + "evals/triage-security/files/fullsend-report-only-trusted-input.json") + assert str(expected_path) in result.stderr + assert not (root / "pre/tc-6677-gate.env").exists() + + +def test_malformed_assertion_requires_raw_abort_and_host_retention_evidence(): + """Require actual rejection for both absent output and intentional host stripping.""" + # Given the real case contract, without generating runtime evidence or grading + annotations = yaml.safe_load((SUITE / "cases/035-malformed/annotations.yaml").read_text()) + # When inspecting the malformed output assertion's evidence requirements + assertion = annotations["assertions"][3] + # Then both output branches require actual rejection, not merely empty output + assert annotations["assertion_count"] == len(annotations["assertions"]) == 5 + for requirement in [ + "actual sdlc-workflow:triage-security invocation", "invalid JSON parser error", + "tool exit1", "no result file", "no recovery write is required", + "If an error-only file was written", "BEFORE host validation", + '{"error":"triage-security aborted: trusted input is missing, invalid JSON, or fails ' + 'triage-security-input.schema.json; no interactive fallback is available in the sandbox."}', + "validate-output-schema.sh", "strip_extra_properties.py", "stripped: ['error']", + "rejection of the success schema", "collected agent-result.json exactly {}", + "native Fullsend CLI exit nonzero", "complete ordered evidence chain", + "no successful analysis, fallback or actions", "empty output or nonzero alone", + "infrastructure/inference failure", "nonempty unexpected output or success report", + "narrated outcomes are FAIL", + "Output directory absence or a failed attempted abort write AFTER proven real Skill invalid JSON rejection", + "accepted on the no-result path", "not a disqualifying bootstrap/inference failure", + "disqualifying only when it prevents actual Skill input validation", + "Without genuine invalid JSON proof, FAIL", + ]: + assert requirement in assertion + assert "no output files OR sole agent-result.json containing {}" in annotations["assertions"][4] + + +def test_separate_suite_declares_all_strict_execution_assertions(): + """Native scenarios must retain 21 distinct execution requirements.""" + # Given the native framework's dataset rather than ordinary evals.json + assert (SUITE / "eval.yaml").is_file(), "Separate native suite is missing" + config = yaml.safe_load((SUITE / "eval.yaml").read_text()) + # Then the opaque CLI contract supplies every independent case to the framework + assert config["runner"]["type"] == "cli" + assert isinstance(config["runner"]["command"], list) + assert "{scenario}" in config["runner"]["command"] + assert not config.get("hooks") + assert config["outputs"] == [{"path": "output"}] + cases = sorted((SUITE / "cases").iterdir()) + assert [p.name for p in cases] == ["033-absent", "034-empty", "035-malformed", "036-valid"] + assert [len(yaml.safe_load((p / "annotations.yaml").read_text())["assertions"]) for p in cases] == [4, 5, 5, 7] + assert all(j["feedback_type"] == "bool" for j in config["judges"]) + assert all(t["min_pass_rate"] == 1.0 for t in config["thresholds"].values()) + + +@pytest.mark.parametrize("exit_code", [0, 7]) +@pytest.mark.parametrize("scenario", ["absent", "empty", "malformed", "valid"]) +def test_native_adapter_preserves_process_exit_and_artifacts(tmp_path, monkeypatch, exit_code, scenario): + """Only the external CLI is doubled: staging, arguments and raw retention are real.""" + # Given a synthetic CLI process, explicitly not model or Skill execution + adapter = load_script(SUITE / "run-fullsend.py") + workspace = tmp_path / "case with spaces" + workspace.mkdir() + output = workspace / "output" + output.mkdir() + observed = {} + native_bytes = b'{"synthetic":"NOT AGENT EXECUTION","total_cost_usd":2}\n' + actual_process = subprocess.run + + def fake_process(command, **kwargs): + """Stand in for Fullsend only; create unmistakably synthetic native files.""" + if command[0] != "/isolated/fullsend": + return actual_process(command, **kwargs) + observed["command"] = command + observed["cwd"] = kwargs["cwd"] + config_dir = Path(command[command.index("--fullsend-dir") + 1]) + observed["config_dir"] = config_dir + observed["config"] = yaml.safe_load((config_dir / "config.yaml").read_text()) + observed["harness"] = yaml.safe_load((config_dir / "harness/triage-security-gate.yaml").read_text()) + native = Path(command[command.index("--output-dir") + 1]) / "agent-triage-security-gate-synthetic" + (native / "iteration-1/transcripts").mkdir(parents=True) + (native / "iteration-1/transcripts/runtime.jsonl").write_bytes(b"SYNTHETIC NOT AGENT EXECUTION\n") + (native / "metrics.json").write_bytes(native_bytes) + return subprocess.CompletedProcess(command, exit_code) + + monkeypatch.setattr(adapter.subprocess, "run", fake_process) + monkeypatch.delenv("FULLSEND_MINT_URL", raising=False) + # When the native adapter stages its resources and delegates one command + actual = adapter.run_case(ROOT, workspace, output, scenario, "model-under-test", "high", + Path("/isolated/fullsend"), Path("/isolated/linux-fullsend")) + # Then raw failure is preserved and only native metrics are copied unchanged + assert actual == exit_code + command = observed["command"] + assert command[:3] == ["/isolated/fullsend", "run", "triage-security-gate"] + assert command[command.index("--model") + 1] == "model-under-test" + assert command[command.index("--effort") + 1] == "high" + assert command[command.index("--runtime") + 1] == "claude" + assert command[command.index("--fullsend-binary") + 1] == "/isolated/linux-fullsend" + assert "--env-file" not in command and "--status-number" not in command + assert "--no-post-script" in command + h = observed["harness"] + # Reviewed production contract from d83ee90b:harness/triage-security.yaml. + # Bootstrap ports companions only, so retain this small expected-value + # fixture rather than requiring an unpublished git object or live harness. + production = { + "image": "ghcr.io/fullsend-ai/fullsend-code@sha256:9743bc7b6e451e0bcea25ae4a67e0c040c296f1fee04c08988ae80c53fafcfe6", + "policy": "plugins/sdlc-workflow/policies/triage-security.yaml", + "providers": ["plugins/sdlc-workflow/providers/vertex-ai.yaml"], + "openshell": {"profiles": ["plugins/sdlc-workflow/profiles/fullsend-vertex-ai.yaml"]}, + "validation_loop": {"script": "plugins/sdlc-workflow/scripts/validate-output-schema.sh", + "schema": "plugins/sdlc-workflow/schemas/triage-security-result.schema.json"}, + } + assert h["image"] == production["image"] and h["readonly_repo"] is True + for field in ["policy", "agent", "pre_script"]: + assert Path(h[field]).is_absolute() + staged = observed["config_dir"] + assert h["plugins"] == [str(staged / "plugins/sdlc-workflow")] + assert h["policy"] == str(staged / production["policy"]) + assert h["providers"] == [str(staged / p) for p in production["providers"]] + assert h["openshell"]["profiles"] == [str(staged / p) for p in production["openshell"]["profiles"]] + assert h["validation_loop"]["script"] == str(staged / production["validation_loop"]["script"]) + assert h["validation_loop"]["schema"] == str(staged / production["validation_loop"]["schema"]) + assert h["host_files"][0]["src"] == str(staged / "plugins/sdlc-workflow/env/gcp-vertex.env") + assert h["env"]["runner"]["TC6677_REPO_ROOT"] == str(staged) + assert h["env"]["runner"]["FULLSEND_OUTPUT_SCHEMA"] == h["validation_loop"]["schema"] + # Then delivery resources are contained real copies, not links to outside code + plugin_files = [p for p in (ROOT / "plugins/sdlc-workflow").rglob("*") + if p.is_file() and "__pycache__" not in p.parts and p.suffix != ".pyc"] + for original in plugin_files: + copied = staged / original.relative_to(ROOT) + assert copied.resolve().is_relative_to(staged.resolve()) + assert copied.read_bytes() == original.read_bytes() + assert not list(staged.rglob("__pycache__")) + for name in ["agent.md", "prepare-fixture.py"]: + assert (staged / "evals/fullsend/triage-security" / name).read_bytes() == (SUITE / name).read_bytes() + fixture_names = [ + "fullsend-gate-interactive-config.md", "fullsend-invalid-trusted-input.md", "fullsend-report-only-trusted-input.json"] + assert sorted(p.name for p in (staged / "evals/triage-security/files").iterdir()) == fixture_names + for name in fixture_names: + assert (staged / "evals/triage-security/files" / name).read_bytes() == (FIXTURES / name).read_bytes() + assert h["validation_loop"]["max_iterations"] == 1 and "post_script" not in h + assert "sandbox" not in h["env"] and "JIRA_API_TOKEN" not in h["env"]["runner"] + assert h["host_files"][2]["dest"] == "/sandbox/workspace/.env.d/zz-tc-6677-gate.env" + assert h["host_files"][1]["dest"] == "/sandbox/workspace/.pre-script/triage-security-input.json" + assert [f["src"] for f in h["host_files"][1:3]] == ["pre/triage-security-input.json", "pre/tc-6677-gate.env"] + assert all(f["optional"] is True for f in h["host_files"][1:3]) + assert len(h["host_files"]) == 4 + assert h["host_files"][3] == { + "src": "${GOOGLE_APPLICATION_CREDENTIALS}", "dest": "/tmp/.gcp-credentials.json"} + target = Path(command[command.index("--target-repo") + 1]) + # Then the actual local Git fixture satisfies native copy/read-only setup, + # without a remote, commit, outside repository or extra project content. + assert (target / ".git").is_dir(), "Native read-only setup requires real Git metadata" + git_root = actual_process(["git", "-C", str(target), "rev-parse", "--show-toplevel"], + capture_output=True, text=True, check=True) + assert Path(git_root.stdout.strip()).resolve() == target.resolve() + assert actual_process(["git", "-C", str(target), "remote"], + capture_output=True, text=True, check=True).stdout == "" + assert actual_process(["git", "-C", str(target), "rev-parse", "--verify", "HEAD"], + capture_output=True, check=False).returncode != 0 + assert (target / ".git/info/exclude").is_file() + if scenario == "absent": + assert (target / "CLAUDE.md").read_bytes() == (FIXTURES / "fullsend-gate-interactive-config.md").read_bytes() + assert sorted(p.name for p in target.iterdir()) == [".git", "CLAUDE.md"] + else: + assert [p.name for p in target.iterdir()] == [".git"], "Noninteractive cases must not preload interactive configuration" + assert (output / "metrics.json").read_bytes() == native_bytes + assert sorted(p.name for p in output.iterdir()) == ["metrics.json", "native"] + + +def test_staged_layout_with_actual_pinned_fullsend_resolver(tmp_path, monkeypatch): + """Use native Go resolution, not a duplicate containment check or Skill execution.""" + # Given optional read-only source and cached Go deps; consumer setup needs neither + synthetic_adc = tmp_path / "external-synthetic-not-credentials.txt" + synthetic_adc.write_text("# SYNTHETIC TEST DATA — NOT CREDENTIALS; native path validation only\n") + source = os.environ.get("TC6677_FULLSEND_SOURCE") + if not source or not shutil.which("go"): + pytest.skip("Native resolver contract needs TC6677_FULLSEND_SOURCE and Go with cached dependencies") + pin = "d5f36921ac754705619f38c637ef692873809fbc" + archive = subprocess.check_output(["git", "-C", source, "archive", pin]) + snapshot = tmp_path / "pinned-source" + snapshot.mkdir() + with tarfile.open(fileobj=io.BytesIO(archive)) as package: + package.extractall(snapshot, filter="data") + probe = tmp_path / "resolver-probe" + probe.mkdir() + go_mod = (snapshot / "go.mod").read_text().replace( + "module github.com/fullsend-ai/fullsend\n", "module github.com/fullsend-ai/fullsend/tc6677-resolver-probe\n", 1) + (probe / "go.mod").write_text(go_mod + '\nrequire github.com/fullsend-ai/fullsend v0.0.0\nreplace github.com/fullsend-ai/fullsend => ' + json.dumps(str(snapshot)) + '\n') + shutil.copy2(snapshot / "go.sum", probe / "go.sum") + (probe / "main.go").write_text('''// SYNTHETIC TEST DATA — calls pinned native resource APIs only, no runtime +package main +import ( + "context" + "encoding/json" + "fmt" + "os" + "path/filepath" + "github.com/fullsend-ai/fullsend/internal/harness" + "github.com/fullsend-ai/fullsend/internal/resolve" +) +func main() { + root := os.Args[1] + h, _, err := harness.LoadWithBase(context.Background(), filepath.Join(root, "harness/triage-security-gate.yaml"), harness.ComposeOpts{WorkspaceRoot: root}) + if err == nil { err = h.ResolveRelativeTo(root) } + var result resolve.ResolveResult + if err == nil { result, err = resolve.ResolveHarness(context.Background(), h, resolve.ResolveOpts{WorkspaceRoot: root}) } + if err == nil && (len(result.Profiles) != 1 || len(result.Providers) != 1) { err = fmt.Errorf("expected native profile/provider records") } + // Actual early validation: no native-generated host variable exists yet. + if err == nil { err = h.ValidateRunnerEnvWith(os.LookupEnv) } + if err == nil { err = h.ValidateFilesExist() } + if err != nil { fmt.Fprintln(os.Stderr, err); os.Exit(1) } + if len(h.HostFiles) != 4 { fmt.Fprintln(os.Stderr, "missing native inference credential mount"); os.Exit(1) } + json.NewEncoder(os.Stdout).Encode(map[string]string{"gate_src": h.HostFiles[2].Src, "input_src": h.HostFiles[1].Src, "credential_src": h.HostFiles[3].Src, "credential_dest": h.HostFiles[3].Dest}) +} +''') + environment = dict(os.environ, GOPROXY="off", GOSUMDB="off", GOTOOLCHAIN="local", GOWORK="off", + GOCACHE=str(Path(os.environ.get("TC6677_GO_CACHE", str(tmp_path / "go-cache")))), + GOOGLE_APPLICATION_CREDENTIALS=str(synthetic_adc)) + binary = probe / "resolver" + built = subprocess.run(["go", "build", "-mod=mod", "-o", str(binary), "."], cwd=probe, + env=environment, capture_output=True, text=True, check=False) + assert built.returncode == 0, built.stderr + adapter = load_script(SUITE / "run-fullsend.py") + workspace = tmp_path / "case" + workspace.mkdir() + monkeypatch.delenv("FULLSEND_MINT_URL", raising=False) + observed = {} + + def resolve_only(command, **kwargs): + """Replace the inference CLI boundary with the actual source-pinned resolver.""" + setup = Path(command[command.index("--fullsend-dir") + 1]) + observed["setup"] = setup + result = subprocess.run([str(binary), str(setup)], env=environment, + capture_output=True, text=True, check=False) + assert result.returncode == 0, result.stderr + mounts = json.loads(result.stdout) + assert mounts == {"gate_src": str(setup / "pre/tc-6677-gate.env"), + "input_src": str(setup / "pre/triage-security-input.json"), + "credential_src": "${GOOGLE_APPLICATION_CREDENTIALS}", + "credential_dest": "/tmp/.gcp-credentials.json"} + assert not synthetic_adc.resolve().is_relative_to(setup.resolve()) + assert not list(setup.rglob(synthetic_adc.name)) + # Actual pinned validation must reject an unset mandatory source variable. + missing_adc = dict(environment) + missing_adc.pop("GOOGLE_APPLICATION_CREDENTIALS") + rejected = subprocess.run([str(binary), str(setup)], env=missing_adc, + capture_output=True, text=True, check=False) + assert rejected.returncode == 1 and "host variable GOOGLE_APPLICATION_CREDENTIALS is not set" in rejected.stderr + # The staged pre-script must consume the staged exact fixtures successfully. + h = yaml.safe_load((setup / "harness/triage-security-gate.yaml").read_text()) + fixture_env = dict(environment, TC6677_REPO_ROOT=str(setup), TC6677_SCENARIO="valid") + fixture_env.pop("FULLSEND_RUN_DIR", None) + prepared = subprocess.run([sys.executable, h["pre_script"]], env=fixture_env, + capture_output=True, text=True, check=False) + assert prepared.returncode == 0, prepared.stderr + assert Path(mounts["gate_src"]).read_bytes() == "# SYNTHETIC TEST DATA — deliberate native gate condition injection\n".encode() + assert Path(mounts["input_src"]).read_bytes() == (FIXTURES / "fullsend-report-only-trusted-input.json").read_bytes() + return result + + # Only the adapter's command launch is doubled; native resolver APIs really run + original_run = subprocess.run + monkeypatch.setattr(adapter.subprocess, "run", lambda command, **kwargs: + resolve_only(command, **kwargs) if command[0] == "/not-launched/fullsend" else original_run(command, **kwargs)) + # When staging production resources inside the configuration workspace + assert adapter.run_case(ROOT, workspace, workspace / "output", "valid", "unused", "high", + Path("/not-launched/fullsend"), Path("/not-launched/linux-fullsend")) == 0 + # Then the actual resolver still rejects an external profile and a symlink escape + setup = observed["setup"] + path = setup / "harness/triage-security-gate.yaml" + h = yaml.safe_load(path.read_text()) + external = ROOT / "plugins/sdlc-workflow/profiles/fullsend-vertex-ai.yaml" + for profile in [external, setup / "escaping-profile.yaml"]: + if profile != external: + profile.symlink_to(external) + h["openshell"]["profiles"] = [str(profile)] + path.write_text(yaml.safe_dump(h)) + rejected = subprocess.run([str(binary), str(setup)], env=environment, + capture_output=True, text=True, check=False) + assert rejected.returncode == 1 and "outside workspace root" in rejected.stderr + + +@pytest.mark.parametrize("failure", ["stale", "mint"]) +def test_native_adapter_rejects_stale_or_live_mint_configuration(tmp_path, monkeypatch, failure): + """Synthetic execution must never reuse old evidence or mint a live forge token.""" + adapter = load_script(SUITE / "run-fullsend.py") + workspace = tmp_path / "case" + workspace.mkdir() + output = workspace / "output" + output.mkdir() + monkeypatch.delenv("FULLSEND_MINT_URL", raising=False) + if failure == "stale": + (output / "native").mkdir() + else: + monkeypatch.setenv("FULLSEND_MINT_URL", "https://synthetic.invalid/no-access") + # When preparing a run before any external process could start + with pytest.raises(ValueError): + adapter.run_case(ROOT, workspace, output, "valid", "model", "high", + Path("/no-such-cli"), Path("/no-such-linux-cli")) + + +def test_common_entrypoint_resolves_framework_contract(tmp_path): + """The common entrypoint must resolve CLI placeholders and absolute dataset paths.""" + common = load_script(ROOT / "evals/fullsend/run.py") + # Given a Python path with spaces and explicit runtime/judge choices + config = common.resolved_config(Path("/cache with spaces/bin/python"), "skill-model", "judge-model", "high") + # Then framework-consumed paths and invocation agree without shell splitting + assert config["dataset"]["path"] == str(SUITE / "cases") + assert config["runner"]["command"][:2] == ["/cache with spaces/bin/python", str(SUITE / "run-fullsend.py")] + assert config["execution"]["skill"] == "triage-security-gate" + assert config["models"] == {"skill": "skill-model", "judge": "judge-model"} + assert config["runner"]["effort"] == "high" + + +@pytest.mark.parametrize("score_exit, summary_present", [(0, True), (0, False), (7, True)]) +def test_common_pipeline_collects_failures_before_upstream_judging(tmp_path, monkeypatch, score_exit, summary_present): + """Nonzero execution must preserve case evidence and still reach upstream scoring.""" + common = load_script(ROOT / "evals/fullsend/run.py") + # Given a synthetic framework boundary; no agent/judge/inference runs here + calls = [] + run_dir = tmp_path / "runs/triage-security-gate/synthetic" + workspace = tmp_path / "workspace" + + def fake_process(command, **kwargs): + """Only external framework phases are doubled; orchestration remains real.""" + phase = Path(command[1]).stem + calls.append((phase, command, kwargs)) + if phase == "execute": + for case in ["033-absent", "034-empty", "035-malformed", "036-valid"]: + p = run_dir / "cases" / case + p.mkdir(parents=True) + (p / "run_result.json").write_text('{"exit_code":7}') + return subprocess.CompletedProcess(command, 7) + if phase == "score": + if summary_present: + (run_dir / "summary.yaml").write_text(yaml.safe_dump(synthetic_judge_summary())) + return subprocess.CompletedProcess(command, score_exit) + return subprocess.CompletedProcess(command, 0) + + monkeypatch.setattr(common.subprocess, "run", fake_process) + # When the common pipeline delegates workspace/execute/collect/score + if not summary_present: + with pytest.raises(ValueError, match="summary"): + common.pipeline(Path("/venv/python"), Path("/harness"), tmp_path / "eval.yaml", + workspace, run_dir, "synthetic", dict(os.environ)) + else: + result = common.pipeline(Path("/venv/python"), Path("/harness"), tmp_path / "eval.yaml", + workspace, run_dir, "synthetic", dict(os.environ)) + assert result == score_exit + # Then expected case failures are not normalized or locally graded + assert [c[0] for c in calls] == ["workspace", "execute", "collect", "score"] + assert all(c[1][0] == "/venv/python" for c in calls) + assert calls[-1][1][2] == "judges" + assert all(json.loads(p.read_text())["exit_code"] == 7 for p in run_dir.glob("cases/*/run_result.json")) + + +def test_common_pipeline_refuses_missing_case_results(tmp_path, monkeypatch): + """Infrastructure failure must not become a vacuous zero-case grading success.""" + common = load_script(ROOT / "evals/fullsend/run.py") + called = [] + + def fake_process(command, **kwargs): + called.append(Path(command[1]).stem) + return subprocess.CompletedProcess(command, 1 if called[-1] == "execute" else 0) + + monkeypatch.setattr(common.subprocess, "run", fake_process) + with pytest.raises(ValueError, match="case results"): + common.pipeline(Path("/python"), Path("/harness"), tmp_path / "eval.yaml", + tmp_path / "ws", tmp_path / "run", "synthetic", dict(os.environ)) + assert called == ["workspace", "execute"] + + +def test_binary_setup_rejects_corrupted_release_before_install(tmp_path): + """A pinned release mismatch must fail before any executable can be installed.""" + common = load_script(ROOT / "evals/fullsend/run.py") + # Given downloaded bytes that do not match the approved release digest + (tmp_path / "fullsend-linux-amd64.tar.gz").write_bytes(b"SYNTHETIC corrupt release") + # When an isolated installation consumes them + with pytest.raises(ValueError, match="digest mismatch"): + common.install_binary(tmp_path, "linux-amd64", common.pins()["fullsend"]) + # Then neither execution nor a partial CLI install occurred + assert not (tmp_path / "bin/fullsend-linux-amd64").exists() + + +def test_setup_overrides_user_pip_install_location(tmp_path, monkeypatch): + """Host pip user-install defaults must not redirect the isolated dependency install.""" + common = load_script(ROOT / "evals/fullsend/run.py") + source = tmp_path / "agent-eval-harness" + source.mkdir() + (tmp_path / "venv/bin").mkdir(parents=True) + (tmp_path / "venv/bin/python").touch() + commands = [] + monkeypatch.setattr(common.sys, "version_info", (3, 12)) + monkeypatch.setattr(common, "verify_source", lambda *args: None) + monkeypatch.setattr(common, "install_binary", lambda *args: None) + + def fake_install(command, **kwargs): + commands.append(command) + return subprocess.CompletedProcess(command, 0) + + monkeypatch.setattr(common.subprocess, "run", fake_install) + # When setup prepares dependency installation without external processes + common.setup(tmp_path) + # Then explicit venv location wins over pip's host user-install configuration + assert all("--no-user" in command for command in commands) + assert all(command[command.index("--cache-dir") + 1] == str(tmp_path / "pip-cache") for command in commands) + assert "--require-hashes" in commands[0] + assert "--no-deps" in commands[1] and "--no-build-isolation" in commands[1] + + +def test_verified_archive_download_uses_host_transport(tmp_path, monkeypatch): + """Native host TLS transport and pinned archive verification both precede installation.""" + import io + import tarfile + common = load_script(ROOT / "evals/fullsend/run.py") + # Given an unmistakably synthetic release archive, never an agent executable + buffer = io.BytesIO() + with tarfile.open(fileobj=buffer, mode="w:gz") as package: + member = tarfile.TarInfo("release/fullsend") + data = b"SYNTHETIC NOT A CLI\n" + member.size = len(data) + package.addfile(member, io.BytesIO(data)) + archive = buffer.getvalue() + observed = [] + + def fake_download(command, **kwargs): + """Replace only curl's network transfer with synthetic bytes.""" + observed.append(command) + Path(command[command.index("--output") + 1]).write_bytes(archive) + return subprocess.CompletedProcess(command, 0) + + monkeypatch.setattr(common.subprocess, "run", fake_download) + # When setup consumes a download without using model/network credentials + dependency = {"version": "0.43.0", "archives": {"linux-amd64": hashlib.sha256(archive).hexdigest()}} + common.install_binary(tmp_path, "linux-amd64", dependency) + # Then curl retains TLS verification and checksum-verified payload bytes are installed + assert observed[0][:2] == ["curl", "--fail"] + assert "--insecure" not in observed[0] and "-k" not in observed[0] + assert (tmp_path / "bin/fullsend-linux-amd64").read_bytes() == data + + +@pytest.mark.parametrize("installed", ["wrong-version", None]) +def test_dependency_preflight_rejects_unlocked_or_missing_packages(tmp_path, monkeypatch, installed): + """A dependency-only preflight must detect drift before a model can be launched.""" + common = load_script(ROOT / "evals/fullsend/run.py") + (tmp_path / "requirements.lock").write_text('pyyaml==6.0.3 \\\n --hash=sha256:' + 'a' * 64 + '\n') + assert callable(getattr(common, "verify_locked_dependencies", None)), "Dependency lock verification is missing" + monkeypatch.setattr(importlib.metadata, "version", lambda package: installed) + with pytest.raises(ValueError, match="Locked dependency"): + common.verify_locked_dependencies(tmp_path / "requirements.lock") + + +def test_dependency_preflight_accepts_exact_locked_packages(tmp_path, monkeypatch): + """Exact version/hash entries remain acceptable without importing inference clients.""" + common = load_script(ROOT / "evals/fullsend/run.py") + (tmp_path / "requirements.lock").write_text('pyyaml==6.0.3 \\\n --hash=sha256:' + 'a' * 64 + '\n') + assert callable(getattr(common, "verify_locked_dependencies", None)), "Dependency lock verification is missing" + monkeypatch.setattr(importlib.metadata, "version", lambda package: "6.0.3") + common.verify_locked_dependencies(tmp_path / "requirements.lock") + + +def test_upstream_collection_retains_nested_native_bytes(tmp_path): + """Characterize the needed opaque CLI collection boundary, without execution/grading.""" + # Given installed pinned tooling and unmistakably synthetic native artifacts + cache = Path(os.environ.get("TC6677_EVAL_CACHE", "/tmp/tc-6677-eval-deps")).resolve() + python = cache / "venv/bin/python" + if not python.is_file(): + pytest.skip("Optional dependency contract: run the isolated setup first") + common = load_script(ROOT / "evals/fullsend/run.py") + common.verify_source(cache / "agent-eval-harness", common.pins()["harness"]) + workspace = tmp_path / "workspace" + native = workspace / "cases/036-valid/output/native/agent-synthetic/iteration-1/transcripts" + native.mkdir(parents=True) + raw = b'{"synthetic":"NOT AGENT EXECUTION"}\n{"partial":' + (native / "runtime.jsonl").write_bytes(raw) + output = tmp_path / "collected" + config = tmp_path / "eval.yaml" + config.write_text(yaml.safe_dump(common.resolved_config(python, "unused", "unused", "high"))) + # When the real upstream collector consumes our declared output path + process = subprocess.run([str(python), str(cache / "agent-eval-harness/skills/eval-run/scripts/collect.py"), + "--config", str(config), "--workspace", str(workspace), "--output", str(output)], + cwd=ROOT, capture_output=True, text=True, check=False) + # Then nested original bytes are retained, without repaired transcripts or verdicts + assert process.returncode == 0, process.stderr + assert (output / "cases/036-valid/output/native/agent-synthetic/iteration-1/transcripts/runtime.jsonl").read_bytes() == raw + assert not list(output.rglob("judge*")) and not list(output.rglob("agent-result.json")) + + +@pytest.mark.parametrize("termination", [signal.SIGTERM, signal.SIGKILL]) +def test_native_adapter_preserves_signal_termination(tmp_path, termination): + """A native CLI signal must not be rewritten to Python's unsigned exit code.""" + # Given a synthetic failing process, never Fullsend or inference + fake = tmp_path / "synthetic-cli" + fake.write_text(f'#!{sys.executable}\n# SYNTHETIC TEST DATA — process signal contract only\nimport os\nos.kill(os.getpid(),{int(termination)})\n') + fake.chmod(0o755) + workspace = tmp_path / "case" + workspace.mkdir() + environment = dict(os.environ, TC6677_FULLSEND_BIN=str(fake), TC6677_SANDBOX_FULLSEND_BIN=str(fake)) + environment.pop("FULLSEND_MINT_URL", None) + # When the real adapter delegates to that CLI double + process = subprocess.run([sys.executable, str(SUITE / "run-fullsend.py"), "--agent", "triage-security-gate", + "--workspace", str(workspace), "--output-dir", str(workspace / "output"), + "--scenario", "valid", "--model", "unused", "--effort", "high"], + env=environment, capture_output=True, text=True, check=False) + # Then the outer process reports the same actual signal to CliRunner + assert process.returncode == -termination + + +@pytest.mark.parametrize("failure_phase", ["workspace", "execute", "collect", "score", "summary", None]) +def test_pipeline_exports_safe_failure_phase_without_changing_execution(tmp_path, monkeypatch, failure_phase): + """Diagnostics preserve phase exits and grading rules without exporting subprocess text.""" + # Given synthetic framework failures and expected negative case exits + common = load_script(ROOT / "evals/fullsend/run.py") + run_dir = tmp_path / "synthetic" + diagnostics = {} + calls = [] + + def fake_process(command, **kwargs): + """Double framework execution; deliberately secret-shaped stderr stays private.""" + phase = Path(command[1]).stem + calls.append(phase) + if phase == "execute" and failure_phase != "execute": + for case in common.CASES: + directory = run_dir / "cases" / case + directory.mkdir(parents=True) + (directory / "run_result.json").write_text('{"exit_code":7}') + if phase == "score" and failure_phase != "summary": + (run_dir / "summary.yaml").write_text(yaml.safe_dump(synthetic_judge_summary())) + if phase == failure_phase: + kwargs["stderr"].write(b"SYNTHETIC SECRET bearer /private/credentials\n") + return subprocess.CompletedProcess(command, 7 if phase == failure_phase or phase == "execute" else 0) + + monkeypatch.setattr(common.subprocess, "run", fake_process) + # When the real pipeline runs and exports diagnostics to the existing safe report + if failure_phase in {"execute", "summary"}: + with pytest.raises(ValueError): + common.pipeline(Path("/python"), Path("/harness"), tmp_path / "config.yaml", + tmp_path / "ws", run_dir, "synthetic", {}, diagnostics) + result = 1 + else: + result = common.pipeline(Path("/python"), Path("/harness"), tmp_path / "config.yaml", + tmp_path / "ws", run_dir, "synthetic", {}, diagnostics) + source = {key: "a" * 40 for key in ["head_sha", "merge_sha", "base_sha", "trusted_sha", "eval_source_sha"]} + source["pr_number"] = 299 + common.publish_report(run_dir, tmp_path / "safe", source, result, diagnostics) + report = json.loads((tmp_path / "safe/native-result.json").read_text()) + # Then only fixed phase/category identifiers and actual numeric exits leave the host + assert report["diagnostics"]["phase"] == (failure_phase or "complete") + assert report["diagnostics"]["phase_exits"] == { + phase: 7 if phase == failure_phase or phase == "execute" else 0 for phase in calls} + expected_code = {"execute": "missing-case-results", "summary": "invalid-summary"}.get( + failure_phase, "phase-exit" if failure_phase else "none") + assert report["diagnostics"]["code"] == expected_code + assert "SECRET" not in json.dumps(report) + assert "credentials" not in json.dumps(report) + assert result == (0 if failure_phase is None else 1 if failure_phase in {"execute", "summary"} else 7) + + +@pytest.mark.parametrize("stderr,expected", [ + (b"SYNTHETIC HTTP 401 Unauthorized SECRET", "authentication-error"), + (b"SYNTHETIC PermissionDenied SECRET", "permission-error"), + (b"SYNTHETIC RESOURCE_EXHAUSTED SECRET", "quota-error"), + (b"SYNTHETIC deadline exceeded SECRET", "timeout"), + (b"SYNTHETIC connection refused SECRET", "connection-error"), + (b"SYNTHETIC arbitrary hostile text SECRET", "phase-exit"), +]) +def test_phase_failure_category_is_fixed_and_contains_no_upstream_text(tmp_path, monkeypatch, stderr, expected): + """Known error markers map to advisory categories; arbitrary bytes never enter reports.""" + common = load_script(ROOT / "evals/fullsend/run.py") + diagnostics = {} + def fake_process(command, **kwargs): + """SYNTHETIC TEST DATA — diagnostic text only, no real inference.""" + kwargs["stderr"].write(stderr) + return subprocess.CompletedProcess(command, 1) + monkeypatch.setattr(common.subprocess, "run", fake_process) + assert common.pipeline(Path("/python"), Path("/harness"), tmp_path / "config.yaml", + tmp_path / "ws", tmp_path / "run", "synthetic", {}, diagnostics) == 1 + assert diagnostics["code"] == expected + assert "SECRET" not in json.dumps(diagnostics) + + +def test_safe_diagnostics_reject_unknown_fields_and_untrusted_values(tmp_path): + """The safe report refuses diagnostic text or invented phase/category names.""" + common = load_script(ROOT / "evals/fullsend/run.py") + source = {key: "a" * 40 for key in ["head_sha", "merge_sha", "base_sha", "trusted_sha", "eval_source_sha"]} + source["pr_number"] = 299 + for diagnostic in [ + {"phase": "score", "code": "phase-exit", "phase_exits": {}, "message": "SECRET"}, + {"phase": "SECRET", "code": "phase-exit", "phase_exits": {}}, + {"phase": "score", "code": "SECRET", "phase_exits": {}}, + {"phase": "score", "code": "phase-exit", "phase_exits": {"score": "SECRET"}}, + ]: + with pytest.raises(ValueError, match="diagnostic"): + common.publish_report(tmp_path / "missing", tmp_path / "safe", source, 1, diagnostic) + assert not (tmp_path / "safe/native-result.json").exists() + + +def test_phase_stderr_is_streamed_with_bounded_diagnostic_reads(tmp_path, monkeypatch, capsys): + """Large private stderr never requires a whole-log allocation before safe reporting.""" + common = load_script(ROOT / "evals/fullsend/run.py") + reads = [] + class BoundedFile(io.BytesIO): + """SYNTHETIC TEST DATA — reject any unbounded spool read.""" + def read(self, size=-1): + """Record and constrain diagnostic read sizes.""" + assert 0 < size <= 65536 + reads.append(size) + return super().read(size) + monkeypatch.setattr(common.tempfile, "TemporaryFile", BoundedFile) + def fake_process(command, **kwargs): + """SYNTHETIC TEST DATA — large output remains in the private stderr stream.""" + kwargs["stderr"].write(b"HTTP 401 Unauthorized\n" + b"x" * 200000) + return subprocess.CompletedProcess(command, 1) + monkeypatch.setattr(common.subprocess, "run", fake_process) + diagnostics = {} + assert common.pipeline(Path("/python"), Path("/harness"), tmp_path / "config.yaml", + tmp_path / "ws", tmp_path / "run", "synthetic", {}, diagnostics) == 1 + assert len(reads) >= 4 + assert diagnostics["code"] == "authentication-error" + assert len(capsys.readouterr().err) == len(b"HTTP 401 Unauthorized\n") + 200000 diff --git a/plugins/sdlc-workflow/scripts/test_jira_client.py b/plugins/sdlc-workflow/scripts/test_jira_client.py index 2b957f589..0056634a3 100644 --- a/plugins/sdlc-workflow/scripts/test_jira_client.py +++ b/plugins/sdlc-workflow/scripts/test_jira_client.py @@ -22,6 +22,8 @@ sanitize_adf = jira_client.sanitize_adf get_versions = jira_client.get_versions create_issue = jira_client.create_issue +search_jql = jira_client.search_jql +search_jql_all = jira_client.search_jql_all def test_code_block_with_blank_lines(): @@ -412,6 +414,128 @@ def test_get_versions_unreleased_only_filters_correctly(): print("✓ get_versions unreleased_only filter test passed") +def test_search_jql_posts_fields_as_array(): + """search_jql must POST to search/jql with fields as a JSON array. + + Regression guard: the enhanced endpoint replaced /rest/api/3/search (410 + Gone). Passing fields as a URL-encoded comma string (GET) made the endpoint + return issues without the requested custom fields; POST with a fields array + avoids that. Also asserts the legacy startAt offset is gone. + """ + captured = {} + original_make_request = jira_client.make_request + + def fake_make_request(method, endpoint, data=None): + captured["method"] = method + captured["endpoint"] = endpoint + captured["data"] = data + return {"issues": [], "isLast": True} + + jira_client.make_request = fake_make_request + try: + search_jql( + 'cf[10875] ~ "https://example/pull/1"', + fields="status,labels,customfield_10875", + max_results=50, + ) + assert captured["method"] == "POST", captured["method"] + assert captured["endpoint"] == "search/jql", captured["endpoint"] + assert captured["data"]["fields"] == [ + "status", "labels", "customfield_10875", + ], captured["data"]["fields"] + assert captured["data"]["jql"] == 'cf[10875] ~ "https://example/pull/1"' + assert captured["data"]["maxResults"] == 50 + assert "startAt" not in captured["data"] + assert "nextPageToken" not in captured["data"] # omitted on first page + finally: + jira_client.make_request = original_make_request + + print("✓ search_jql posts fields as array test passed") + + +def test_search_jql_all_collects_issues_beyond_first_page(): + """search_jql_all follows nextPageToken so an exact match on a later page is + not dropped (the TC-6233 false ADR-0072 skip). Two pages, 60 issues total; + the exact PR match sits on page 2 (index 55, beyond the first 50). + """ + page1 = { + "issues": [{"key": f"TC-{i}"} for i in range(50)], + "nextPageToken": "PAGE2", + } + # Page 2 carries the exact match beyond the first 50 recall results; the last + # page omits nextPageToken, which is the loop's stop condition. + page2 = {"issues": [{"key": f"TC-{i}"} for i in range(50, 60)]} + pages = [page1, page2] + calls = [] + original_make_request = jira_client.make_request + + def fake_make_request(method, endpoint, data=None): + calls.append({"method": method, "endpoint": endpoint, "data": data}) + return pages[len(calls) - 1] + + jira_client.make_request = fake_make_request + try: + result = search_jql_all('cf[10875] ~ "https://example/pull/1"') + # All 60 issues aggregated, including the beyond-page-1 exact match + assert len(result["issues"]) == 60, len(result["issues"]) + assert {"key": "TC-55"} in result["issues"] + assert result["isLast"] is True + assert "nextPageToken" not in result + # Exactly two POSTs; page 1 omits the cursor, page 2 sends PAGE2 + assert len(calls) == 2, calls + assert calls[0]["endpoint"] == "search/jql" + assert "nextPageToken" not in calls[0]["data"] + assert calls[1]["data"]["nextPageToken"] == "PAGE2" + finally: + jira_client.make_request = original_make_request + + print("✓ search_jql_all collects issues beyond first page test passed") + + +def test_search_jql_all_single_page_stops_without_token(): + """A response without nextPageToken ends pagination after one request.""" + calls = [] + original_make_request = jira_client.make_request + + def fake_make_request(method, endpoint, data=None): + calls.append(endpoint) + return {"issues": [{"key": "TC-1"}], "isLast": True} + + jira_client.make_request = fake_make_request + try: + result = search_jql_all('cf[10875] ~ "x"') + assert result["issues"] == [{"key": "TC-1"}] + assert len(calls) == 1, calls + finally: + jira_client.make_request = original_make_request + + print("✓ search_jql_all single page test passed") + + +def test_search_jql_all_fails_loud_on_runaway_pagination(): + """A server that never stops returning a cursor must fail loud (exit 1), not + loop forever or silently truncate (CONVENTIONS.md §Error Handling). + """ + original_make_request = jira_client.make_request + + def fake_make_request(method, endpoint, data=None): + return {"issues": [{"key": "TC-x"}], "nextPageToken": "ALWAYS"} + + jira_client.make_request = fake_make_request + try: + raised = False + try: + search_jql_all('cf[10875] ~ "x"', max_pages=3) + except SystemExit as e: + raised = True + assert e.code == 1, e.code + assert raised, "expected SystemExit on runaway pagination" + finally: + jira_client.make_request = original_make_request + + print("✓ search_jql_all runaway-pagination guard test passed") + + def test_create_issue_priority_field_mapping(): """Verifies that create_issue maps priority parameter to correct Jira field structure.""" captured = {} @@ -556,6 +680,36 @@ def fake_make_request(method, endpoint, data=None): print("✓ create_issue fix_versions filters empty names test passed") +def test_create_issue_fails_fast_on_empty_issue_type(): + """Verifies create_issue with an empty issue_type fails fast without issuing a request.""" + called = {"made": False} + original_make_request = jira_client.make_request + + def fake_make_request(method, endpoint, data=None): + called["made"] = True + return {"key": "TEST-1", "id": "1"} + + jira_client.make_request = fake_make_request + try: + # When creating an issue with an empty issue_type + exit_code = None + try: + create_issue("TC", "Test", "desc", "") + raised = False + except SystemExit as e: + raised = True + exit_code = e.code + + # Then it fails fast (SystemExit 1) and issues no HTTP request + assert raised, "Expected create_issue to fail fast (SystemExit) on empty issue_type" + assert exit_code == 1, f"Expected exit code 1, got {exit_code}" + assert not called["made"], "Expected no HTTP request to be issued" + finally: + jira_client.make_request = original_make_request + + print("✓ create_issue fails fast on empty issue_type test passed") + + def run_all_tests(): """Run all tests and report results.""" tests = [ @@ -578,6 +732,7 @@ def run_all_tests(): test_create_issue_omits_fix_versions_when_none, test_create_issue_all_optional_fields_together, test_create_issue_fix_versions_filters_empty_names, + test_create_issue_fails_fast_on_empty_issue_type, ] failed = [] diff --git a/plugins/sdlc-workflow/scripts/test_native_fullsend_eval_ci.py b/plugins/sdlc-workflow/scripts/test_native_fullsend_eval_ci.py index dd7e681a1..772676e60 100644 --- a/plugins/sdlc-workflow/scripts/test_native_fullsend_eval_ci.py +++ b/plugins/sdlc-workflow/scripts/test_native_fullsend_eval_ci.py @@ -1,5 +1,9 @@ """Trusted workflow contracts; no native suite or model execution lives here.""" +import importlib.util +import re +import shutil +import sys import json import os from pathlib import Path @@ -27,6 +31,7 @@ def run_js(script, data, env=None): const data = JSON.parse(process.argv[1]); const outputs = {}, statuses = [], errors = [], reviews = [], updates = []; const storedReviews = data.existingReviews || []; +const comments = [], storedComments = data.existingComments || []; let prReads = 0; const delays = []; const setTimeout = (fn,ms)=>{delays.push(ms);fn();}; @@ -55,10 +60,13 @@ def run_js(script, data, env=None): createReview:async r=>{reviews.push(r);storedReviews.push({...r,id:storedReviews.length+1,user:{login:'github-actions[bot]'}});}, listFiles:async()=> (data.paths||[]).map(filename=>({filename}))}, git:{getCommit:async()=>({data:{parents:(data.parents||['b'.repeat(40),'a'.repeat(40)]).map(sha=>({sha}))}})}, + issues:{listComments:async()=>({data:storedComments}), + updateComment:async r=>{updates.push(r);storedComments.find(s=>s.id===r.comment_id).body=r.body;}, + createComment:async r=>{comments.push(r);storedComments.push({...r,id:storedComments.length+1,user:{login:'github-actions[bot]'}});}}, repos:{getCollaboratorPermissionLevel:async()=>({data:{permission:data.permission||'read'}}), getContent:async()=>({data:{}}),createCommitStatus:async s=>statuses.push(s)}}}; (async()=>{for(let i=0;i<(data.repeat || 1);i++){SCRIPT -}})().then(()=>process.stdout.write(JSON.stringify({outputs,statuses,errors,reviews,updates,prReads,delays}))) +}})().then(()=>process.stdout.write(JSON.stringify({outputs,statuses,errors,reviews,updates,comments,prReads,delays}))) .catch(e=>{process.stdout.write(JSON.stringify({outputs,statuses,errors:[...errors,e.message],reviews,updates}));}); '''.replace("SCRIPT", script) result = subprocess.run(["node", "-e", code, json.dumps(data)], @@ -90,12 +98,12 @@ def test_source_resolution_refuses_stale_or_unrelated_revisions(defect): (299, "verify-pr-fullsend", "plugins/sdlc-workflow/schemas/triage-security-input.schema.json", "true"), (299, "verify-pr-fullsend", "plugins/sdlc-workflow/providers/vertex-ai.yaml", "true"), (299, "verify-pr-fullsend", ".github/scripts/run-native-fullsend-evals.sh", "true"), - (300, "verify-pr-fullsend", "evals/fullsend/run.py", "false"), - (299, "wrong", "evals/fullsend/run.py", "false"), + (300, "verify-pr-fullsend", "evals/fullsend/run.py", "true"), + (299, "wrong", "evals/fullsend/run.py", "true"), (299, "verify-pr-fullsend", "README.md", "false"), ]) -def test_native_discovery_is_relevant_and_bootstrap_only(number, branch, path, expected): - """Only relevant changes on exact PR299/source branch activate bootstrap native CI.""" +def test_native_discovery_runs_for_relevant_prs_after_activation(number, branch, path, expected): + """Relevant changes activate native CI after normal rollout, for every approved PR.""" result = run_js(script_step("discover", "Discover changed skills")["with"]["script"], {"paths": [path]}, {"PR_NUMBER": str(number), "SOURCE_BRANCH": branch, "MERGE_SHA": "c" * 40}) assert result["errors"] == [] @@ -233,7 +241,7 @@ def test_all_publication_jobs_check_latest_run(job): assert job_data["permissions"]["actions"] == "read" for step in job_data["steps"]: script = step.get("with", {}).get("script", "") - if any(api in script for api in ["createCommitStatus(", "createReview(", "updateReview("]): + if any(api in script for api in ["createCommitStatus(", "createReview(", "updateReview(", "createComment(", "updateComment("]): assert any(f"steps.{name}.outputs.latest == 'true'" in step["if"] for name in ["publication", "gate-publication"]) @@ -261,12 +269,12 @@ def test_same_pr_head_cannot_publish_concurrently(): assert concurrency["cancel-in-progress"] is False -@pytest.mark.parametrize("native", [True, False]) +@pytest.mark.parametrize("native", [True]) @pytest.mark.parametrize("existing_kind", ["none", "matching", "wrong-head", "human", "other-suite", "later-page"]) def test_review_reruns_reuse_only_matching_bot_head_review(native, existing_kind): - """Both publishers create once, update reruns and leave unrelated reviews alone.""" + """Native publishing creates once, updates reruns and leaves unrelated reviews alone.""" job, name = ("report-status", "Publish native result alongside ordinary review") if native else ( - "run-evals", "Post eval results review") + "run-evals", "Post eval results comment") script = script_step(job, name)["with"]["script"] marker = "## Native Fullsend Eval Results" if native else "## Eval Results" existing = {"id": 888, "user": {"login": "github-actions[bot]"}, "commit_id": "a" * 40, "body": marker} @@ -490,6 +498,224 @@ def test_multiline_credentials_register_individual_nonempty_masks(tmp_path): assert "::warning::synthetic-second" not in lines +def test_host_validator_dependency_is_in_isolated_lock(): + """Fresh CI must install the jsonschema module used by trusted host validation.""" + requirements = (ROOT / "evals/fullsend/requirements.in").read_text() + lock = (ROOT / "evals/fullsend/requirements.lock").read_text() + assert "jsonschema" in re.findall(r"^([a-zA-Z0-9_-]+)", requirements, re.MULTILINE) + assert re.search(r"^jsonschema==[^\n]+", lock, re.MULTILINE) + + +@pytest.mark.parametrize("linked_part", ["plugin", "plugins"]) +def test_native_cli_rejects_root_and_ancestor_plugin_symlinks(tmp_path, linked_part): + """The CLI must reject PR path links before they redirect reads to the trusted host.""" + # Given a PR path pointing outside its checkout through either directory level + checkout = tmp_path / "pr-head" + checkout.mkdir() + trusted_plugin = ROOT / "plugins/sdlc-workflow" + if linked_part == "plugin": + (checkout / "plugins").mkdir() + (checkout / "plugins/sdlc-workflow").symlink_to(trusted_plugin, target_is_directory=True) + else: + (checkout / "plugins").symlink_to(trusted_plugin.parent, target_is_directory=True) + environment = dict(os.environ, TC6677_FULLSEND_BIN="/not-launched", TC6677_SANDBOX_FULLSEND_BIN="/not-launched") + environment.pop("FULLSEND_MINT_URL", None) + workspace = tmp_path / "workspace" + workspace.mkdir() + # When invoking the actual CLI with the lexical selected plugin path + result = subprocess.run([sys.executable, str(ROOT / "evals/fullsend/triage-security/run-fullsend.py"), + "--agent", "triage-security-gate", "--workspace", str(workspace), + "--output-dir", str(workspace / "output"), "--scenario", "valid", + "--model", "unused", "--effort", "high", "--plugin-root", + str(checkout / "plugins/sdlc-workflow")], env=environment, capture_output=True, text=True) + # Then rejection precedes staging and any native launch + assert result.returncode == 1 + assert "symlink" in result.stderr + assert not (workspace / "native-config").exists() + + +def test_ci_adapter_separates_trusted_validation_from_tested_plugin(tmp_path, monkeypatch): + """PR validator/pre-script/policy bytes remain sandbox data and never host commands.""" + path = ROOT / "evals/fullsend/triage-security/run-fullsend.py" + assert path.is_file(), "Native adapter is missing" + spec = importlib.util.spec_from_file_location("native_adapter", path) + adapter = importlib.util.module_from_spec(spec) + spec.loader.exec_module(adapter) + # Given a deliberately adversarial plugin and distinct synthetic ADC files + plugin = tmp_path / "untrusted-plugin" + shutil.copytree(ROOT / "plugins/sdlc-workflow", plugin) + for relative in ["scripts/validate-output-schema.sh", "scripts/strip_extra_properties.py", "policies/triage-security.yaml"]: + (plugin / relative).write_text("# ADVERSARIAL TEST FIXTURE — must never run on host\nUNTRUSTED\n") + host_adc, sandbox_adc = tmp_path / "host-adc", tmp_path / "sandbox-adc" + host_adc.write_text("SYNTHETIC HOST"); sandbox_adc.write_text("SYNTHETIC SANDBOX") + monkeypatch.setenv("GOOGLE_APPLICATION_CREDENTIALS", str(host_adc)) + monkeypatch.setenv("TC6726_SANDBOX_CREDENTIALS", str(sandbox_adc)) + monkeypatch.setenv("FULLSEND_GCP_OIDC_AUTH_FILE", "/synthetic/auth") + token = tmp_path / "oidc-token"; token.write_text("SYNTHETIC NOT A TOKEN") + monkeypatch.setenv("GCP_OIDC_TOKEN_FILE", str(token)) + monkeypatch.delenv("FULLSEND_MINT_URL", raising=False) + workspace = tmp_path / "case"; workspace.mkdir() + real_run = subprocess.run + observed = {} + + def native_boundary(command, **kwargs): + """Double only Fullsend execution; observe real staging and environment selection.""" + if command[0] != "/synthetic/fullsend": + return real_run(command, **kwargs) + setup = Path(command[command.index("--fullsend-dir") + 1]) + observed.update(setup=setup, env=kwargs["env"], harness=yaml.safe_load((setup / "harness/triage-security-gate.yaml").read_text())) + return subprocess.CompletedProcess(command, 7) + + monkeypatch.setattr(adapter.subprocess, "run", native_boundary) + # When the trusted adapter selects independent plugin and host resource sources + assert adapter.run_case(ROOT, workspace, workspace / "output", "valid", "unused", "high", + Path("/synthetic/fullsend"), Path("/synthetic/linux-fullsend"), plugin_root=plugin) == 7 + # Then trusted host scripts and policy retain reviewed bytes, host ADC is untouched + h = observed["harness"] + validator = Path(h["validation_loop"]["script"]) + assert validator.read_bytes() == (ROOT / "plugins/sdlc-workflow/scripts/validate-output-schema.sh").read_bytes() + assert Path(h["policy"]).read_bytes() == (ROOT / "plugins/sdlc-workflow/policies/triage-security.yaml").read_bytes() + assert not validator.is_relative_to(Path(h["plugins"][0])) + assert (Path(h["plugins"][0]) / "scripts/validate-output-schema.sh").read_bytes() == (plugin / "scripts/validate-output-schema.sh").read_bytes() + assert observed["env"]["GOOGLE_APPLICATION_CREDENTIALS"] == str(sandbox_adc) + assert observed["env"]["FULLSEND_GCP_OIDC_AUTH_FILE"] == "/synthetic/auth" + assert os.environ["GOOGLE_APPLICATION_CREDENTIALS"] == str(host_adc) + assert not list(observed["setup"].rglob("*adc*")) + assert h["host_files"][-1] == {"src": str(token), "dest": "/sandbox/workspace/.gcp-oidc-token"} + assert not list(observed["setup"].rglob("oidc-token")) + + +@pytest.mark.parametrize("variable", ["TC6726_SANDBOX_CREDENTIALS", "GCP_OIDC_TOKEN_FILE"]) +def test_prepared_credential_files_must_exist_before_native_cli(tmp_path, monkeypatch, variable): + """Missing prepared credential/token files cause failure rather than native skips.""" + spec = importlib.util.spec_from_file_location("native_adapter", ROOT / "evals/fullsend/triage-security/run-fullsend.py") + adapter = importlib.util.module_from_spec(spec); spec.loader.exec_module(adapter) + monkeypatch.setenv(variable, str(tmp_path / "missing")) + monkeypatch.delenv("FULLSEND_MINT_URL", raising=False) + workspace = tmp_path / "workspace"; workspace.mkdir() + with pytest.raises(ValueError, match="Missing prepared"): + adapter.run_case(ROOT, workspace, workspace / "output", "valid", "unused", "high", + Path("/not-launched"), Path("/not-launched")) + + +def load_common(): + """Import trusted entrypoint without running its CLI.""" + spec = importlib.util.spec_from_file_location("native_common", ROOT / "evals/fullsend/run.py") + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def test_plugin_argument_reaches_only_trusted_adapter(): + """The selected plugin is a data argument, never a PR-owned runner/config.""" + common = load_common() + plugin = Path("/synthetic/pr-head/plugins/sdlc-workflow") + config = common.resolved_config(Path("/trusted/python"), "skill", "judge", "high", plugin_root=plugin) + assert config["runner"]["command"][:2] == ["/trusted/python", str(ROOT / "evals/fullsend/triage-security/run-fullsend.py")] + assert config["runner"]["command"][-2:] == ["--plugin-root", str(plugin)] + assert config["dataset"]["path"] == str(ROOT / "evals/fullsend/triage-security/cases") + + +def test_safe_report_contains_boolean_outcomes_and_revision_without_raw_credentials(tmp_path): + """Publishing allowlists counts/Booleans, excluding arbitrary transcript/rationale bytes.""" + common = load_common() + # Given adversarial upstream rationale content and strict synthetic case outcomes + run = tmp_path / "runs/triage-security-gate/synthetic" + run.mkdir(parents=True) + per_case = {} + for case, count in common.ASSERTION_COUNTS.items(): + per_case[case] = {} + for index in range(1, 8): + per_case[case][f"assertion_{index}"] = {"value": True if index <= count else None, + "rationale": "SECRET bearer /tmp/gha-creds-evil" if index <= count else + f'''Skipped: condition 'annotations.get("assertion_count", 0) > {index - 1}' is false'''} + (run / "summary.yaml").write_text(yaml.safe_dump({"run_id": "synthetic", "per_case": per_case})) + source = {"head_sha": "a" * 40, "merge_sha": "c" * 40, "base_sha": "b" * 40, "trusted_sha": "e" * 40, "eval_source_sha": "f" * 40, "pr_number": 299} + # When exporting only the reviewed safe report contract + common.publish_report(run, tmp_path / "safe", source, 0) + result = json.loads((tmp_path / "safe/native-result.json").read_text()) + assert result["source"] == source + assert result["passed"] == result["total"] == 21 and result["exit_code"] == 0 + assert result["outcomes"]["033-absent"] == {f"assertion_{i}": True for i in range(1, 5)} + assert "SECRET" not in (tmp_path / "safe/native-result.json").read_text() + assert sorted(p.name for p in (tmp_path / "safe").iterdir()) == ["native-result.json"] + + +def test_safe_report_fails_closed_on_missing_upstream_summary(tmp_path): + """Infrastructure failure exports source-bound failure without invented outcomes.""" + common = load_common() + source = {"head_sha": "a" * 40, "merge_sha": "c" * 40, "base_sha": "b" * 40, "trusted_sha": "e" * 40, "eval_source_sha": "f" * 40, "pr_number": 299} + common.publish_report(tmp_path / "missing", tmp_path / "safe", source, 1) + result = json.loads((tmp_path / "safe/native-result.json").read_text()) + assert result["exit_code"] == 1 and result["complete"] is False and result["outcomes"] == {} + + +def test_native_plugin_symlinks_are_rejected_before_host_launch(tmp_path, monkeypatch): + """PR symlinks cannot make trusted staging read external host credential bytes.""" + spec = importlib.util.spec_from_file_location("native_adapter", ROOT / "evals/fullsend/triage-security/run-fullsend.py") + adapter = importlib.util.module_from_spec(spec); spec.loader.exec_module(adapter) + plugin = tmp_path / "plugin"; plugin.mkdir() + (plugin / "escape").symlink_to(tmp_path / "credentials") + monkeypatch.delenv("FULLSEND_MINT_URL", raising=False) + workspace = tmp_path / "workspace"; workspace.mkdir() + with pytest.raises(ValueError, match="symlinks"): + adapter.run_case(ROOT, workspace, workspace / "output", "valid", "unused", "high", + Path("/not-launched"), Path("/not-launched"), plugin_root=plugin) + + + +@pytest.mark.parametrize("existing_kind", ["none", "matching", "older-head", "human", "other-suite", "later-page"]) +def test_ordinary_reruns_reuse_only_matching_bot_sticky_comment(existing_kind): + """Ordinary reports reuse their bot comment across heads without duplicating reruns.""" + # Given a prior bot report, unrelated comment, or no report + marker = "" + existing = {"id": 888, "user": {"login": "github-actions[bot]"}, "body": marker} + if existing_kind == "older-head": existing["body"] += "\nSource head: " + "b" * 40 + elif existing_kind == "human": existing["user"]["login"] = "human" + elif existing_kind == "other-suite": existing["body"] = "## Native Fullsend Eval Results" + stored = [] if existing_kind == "none" else [existing] + script = script_step("run-evals", "Post eval results comment")["with"]["script"] + if existing_kind == "later-page": + stored = [{"id": i, "user": {"login": "human"}} for i in range(100)] + stored + assert "github.paginate(github.rest.issues.listComments" in script + # When the actual publisher runs twice for the current source + result = run_js(script, {"existingComments": stored, "repeat": 2}, { + "PR_NUMBER": "299", "HEAD_SHA": "a" * 40, "MERGE_SHA": "c" * 40, + "SKILLS_CSV": "triage-security"}) + # Then only the marked bot comment is reused and its source is refreshed + reuse = existing_kind in {"matching", "older-head", "later-page"} + assert result["errors"] == [] + assert len(result["comments"]) == (0 if reuse else 1) + assert len(result["updates"]) == (2 if reuse else 1) + assert all(r["body"].startswith(marker) and "a" * 40 in r["body"] + for r in result["comments"] + result["updates"]) + assert all((r["comment_id"] == 888) is reuse for r in result["updates"]) + + +def test_activation_preserves_rendering_and_trusted_main_suite_selection(): + """PR299 keeps deterministic rendering and selects suite code only from trusted main.""" + jobs = workflow()["jobs"] + assert workflow()["env"]["NATIVE_EVAL_SOURCE_SHA"] == "${{ github.sha }}" + render = script_step("run-evals", "Render eval results")["run"] + assert "aggregate_benchmark.py" in render and "render_summary.py" in render + assert jobs["run-evals"]["permissions"]["issues"] == "write" + + +def test_older_head_cannot_overwrite_current_sticky_report(): + """A completed old-head run cannot replace the shared report for a newer PR head.""" + # Given an existing current bot report and an obsolete publishing run + existing = {"id": 888, "user": {"login": "github-actions[bot]"}, + "body": "\nSource head: " + "d" * 40} + # When the actual publisher rechecks the PR head immediately before writing + result = run_js(script_step("run-evals", "Post eval results comment")["with"]["script"], + {"head": "d" * 40, "existingComments": [existing]}, + {"PR_NUMBER": "299", "HEAD_SHA": "a" * 40, "MERGE_SHA": "c" * 40, + "SKILLS_CSV": "triage-security"}) + # Then the current report is neither updated nor duplicated + assert result["errors"] == [] + assert result["updates"] == result["comments"] == [] + + def test_native_wrapper_passes_requested_judge_model(tmp_path): """The trusted wrapper overrides the reviewed runner's older judge default.""" output = ("GOOGLE_APPLICATION_CREDENTIALS=synthetic-sandbox-adc\n" diff --git a/plugins/sdlc-workflow/scripts/test_pre_triage_security.py b/plugins/sdlc-workflow/scripts/test_pre_triage_security.py new file mode 100644 index 000000000..c44631174 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/test_pre_triage_security.py @@ -0,0 +1,820 @@ +#!/usr/bin/env python3 +"""Tests for the trusted triage-security evidence transformer.""" + +import io +import json +import os +import subprocess +import sys +import time +from unittest.mock import MagicMock +from urllib.error import HTTPError +from urllib.request import Request + +import pytest +from jsonschema import FormatChecker, validate + + +SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__)) +sys.path.insert(0, SCRIPT_DIR) + +import pre_triage_security + + +# jsonschema only enforces uri/date-time formats when the jsonschema[format] +# extra is installed (rfc3987 / rfc3339-validator). validate_bundle now fails +# closed without it, so format-dependent tests are skipped rather than failed on +# an environment lacking the extra; CI installs it and exercises them fully. +_HAS_FORMAT_EXTRA = all( + fmt in FormatChecker().checkers for fmt in pre_triage_security._REQUIRED_FORMATS) +requires_format_extra = pytest.mark.skipif( + not _HAS_FORMAT_EXTRA, + reason="requires jsonschema[format] for uri/date-time format enforcement") + + +def test_parse_security_configuration_extracts_required_runner_values(): + """A complete project configuration becomes the sandbox configuration object.""" + # Given a target project's populated Security Configuration + claude_md = """# Project Configuration + +## Jira Configuration + +- Project key: TC + +## Security Configuration + +### Product Lifecycle + +- Product pages URL: https://example.com/lifecycle +- Jira version prefix: PRODUCT +- Vulnerability issue type ID: 10016 +- Component label pattern: pscomponent: +- ProdSec Jira account ID: account-1 + +### Version Streams + +| Stream | Konflux Release Repo | Local Path | Security Matrix Path | +|---|---|---|---| +| 1.0.x | release-repo | /repos/release | docs/matrix.md | + +### Source Repositories + +| Repository | URL | Deployment Context | +|---|---|---| +| component | https://github.com/org/component | customer-shipped | +""" + + # When the runner parses the configuration before fetching evidence + configuration = pre_triage_security.parse_security_configuration(claude_md) + + # Then all schema-required values and optional security settings are retained + assert configuration == { + "project_key": "TC", + "jira_version_prefix": "PRODUCT", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "prodsec_account_id": "account-1", + "version_streams": [{ + "name": "1.0.x", + "matrix_path": "docs/matrix.md", + "release_repository": "release-repo", + }], + "source_repositories": [{ + "name": "component", + "url": "https://github.com/org/component", + "deployment_context": "customer-shipped", + }], + } + + +def test_collect_bundle_escapes_configured_values_in_jql_literals(tmp_path, monkeypatch): + """JQL searches preserve quote and backslash-containing configured values.""" + # Given configured project and CVE values containing JQL metacharacters + project_key = 'TC"\\OPS' + cve_id = 'CVE-2026-"\\12345' + issue = { + "key": "TC-42", + "fields": { + "summary": "security issue", + "labels": [], + "comment": {"comments": []}, + "issuelinks": [], + "customfield_12345": "component", + }, + } + configuration = { + "project_key": project_key, + "vulnerability_issue_type_id": "10016", + "upstream_affected_component_field": "customfield_12345", + } + searched_jql = [] + (tmp_path / "CLAUDE.md").write_text("# Project Configuration\n") + + def jira_client(command, *arguments): + """Return minimal runner evidence while retaining each JQL search.""" + if command == "get_issue": + return issue + if command == "get_remote_links": + return [] + if command == "get_versions": + return [] + if command == "search_jql": + searched_jql.append(arguments[arguments.index("--jql") + 1]) + return {"issues": []} + raise AssertionError("unexpected Jira command: {}".format(command)) + + monkeypatch.setattr( + pre_triage_security, + "_runner_configuration", + lambda _content, _root: (configuration, "https://example.com/lifecycle", [], {}), + ) + monkeypatch.setattr(pre_triage_security, "_jira_client", jira_client) + monkeypatch.setattr(pre_triage_security, "extract_cve_id", lambda _issue: cve_id) + monkeypatch.setattr(pre_triage_security, "_fetch_url", lambda _url, **_kwargs: {}) + monkeypatch.setattr(pre_triage_security, "build_bundle", lambda **bundle: bundle) + + # When the trusted runner builds its JQL searches + pre_triage_security.collect_bundle("TC-42", tmp_path) + + # Then every quoted configured value is escaped before interpolation + escaped_project_key = 'TC\\"\\\\OPS' + escaped_cve_id = 'CVE-2026-\\"\\\\12345' + assert searched_jql == [ + 'project = "{}" AND labels = "{}" AND issuetype = 10016 AND key != "TC-42"'.format( + escaped_project_key, escaped_cve_id), + 'project = "{}" AND issuetype = 10016 AND cf[12345] ~ "component" AND key != "TC-42"'.format( + escaped_project_key), + 'project = "{}" AND issuetype = Task AND labels = "security-preemptive" ' + 'AND labels = "{}" ORDER BY created DESC'.format(escaped_project_key, escaped_cve_id), + ] + + +def test_pre_triage_script_rejects_missing_poller_work_item_url(): + """The runner fails before collection when the poller supplied no work item.""" + # Given runner credentials but no work item dispatched by the poller + script = os.path.join(SCRIPT_DIR, "pre-triage-security.sh") + environment = { + "PATH": os.environ["PATH"], + "JIRA_SERVER_URL": "https://jira.example.com", + "JIRA_EMAIL": "runner@example.com", + "JIRA_API_TOKEN": "token", + } + + # When the trusted pre-script starts + result = subprocess.run(["bash", script], env=environment, capture_output=True, text=True) + + # Then it rejects the missing poller contract before creating a bundle + assert result.returncode != 0 + assert "FULLSEND_WORK_ITEM_URL" in result.stderr + + +def test_pre_triage_script_rejects_missing_runner_credentials(): + """The runner fails before collection when its Jira credential is absent.""" + # Given a poller work item but no Jira API token on the trusted runner + script = os.path.join(SCRIPT_DIR, "pre-triage-security.sh") + environment = { + "PATH": os.environ["PATH"], + "FULLSEND_WORK_ITEM_URL": "https://jira.example.com/browse/TC-42", + "JIRA_SERVER_URL": "https://jira.example.com", + "JIRA_EMAIL": "runner@example.com", + } + + # When the trusted pre-script starts + result = subprocess.run(["bash", script], env=environment, capture_output=True, text=True) + + # Then no collection runs without the credential required by jira-client.py + assert result.returncode != 0 + assert "JIRA_API_TOKEN" in result.stderr + + +def test_pre_triage_script_isolates_bundles_per_fullsend_run(tmp_path, request): + """Separate Fullsend runs publish isolated bundles through their own handoff paths.""" + # Given a harmless collector and two poller-dispatched issues on one runner + script = os.path.join(SCRIPT_DIR, "pre-triage-security.sh") + project_root = tmp_path / "project" + project_root.mkdir() + (project_root / "CLAUDE.md").write_text("# Project Configuration\n") + fake_bin = tmp_path / "bin" + fake_bin.mkdir() + collector = fake_bin / "python3" + collector.write_text( + "#!/usr/bin/env bash\n" + "touch \"${COLLECTOR_READY_DIR}/$3\"\n" + "while [[ ! -f \"${COLLECTOR_RELEASE}\" ]]; do sleep 0.01; done\n" + "printf '{\"issue\":\"%s\"}\\n' \"$3\"\n" + ) + collector.chmod(0o755) + shared_output = tmp_path / "shared" + ready_directory = tmp_path / "ready" + ready_directory.mkdir() + release_file = tmp_path / "release" + processes = [] + + def release_collectors(): + release_file.touch() + for process in processes: + if process.poll() is None: + process.terminate() + process.communicate(timeout=5) + + request.addfinalizer(release_collectors) + + def environment_for(issue_key, run_directory): + return { + "PATH": "{}{}{}".format(fake_bin, os.pathsep, os.environ["PATH"]), + "FULLSEND_WORK_ITEM_URL": "https://jira.example.com/browse/{}".format(issue_key), + "FULLSEND_PROJECT_ROOT": str(project_root), + "FULLSEND_RUN_DIR": str(run_directory), + "PRE_DIR": str(shared_output), + "COLLECTOR_READY_DIR": str(ready_directory), + "COLLECTOR_RELEASE": str(release_file), + "JIRA_SERVER_URL": "https://jira.example.com", + "JIRA_EMAIL": "runner@example.com", + "JIRA_API_TOKEN": "token", + } + + first_run = tmp_path / "run-one" + second_run = tmp_path / "run-two" + + # When two runs collect their dispatched issues concurrently + first_process = subprocess.Popen( + ["bash", script], env=environment_for("TC-42", first_run), stdout=subprocess.PIPE, + stderr=subprocess.PIPE, text=True, + ) + processes.append(first_process) + second_process = subprocess.Popen( + ["bash", script], env=environment_for("TC-43", second_run), stdout=subprocess.PIPE, + stderr=subprocess.PIPE, text=True, + ) + processes.append(second_process) + first_output = first_run / "pre" / "triage-security-input.json" + second_output = second_run / "pre" / "triage-security-input.json" + deadline = time.monotonic() + 5 + while len(list(ready_directory.iterdir())) < 2 and time.monotonic() < deadline: + time.sleep(0.01) + + # Then neither blocked collector can expose a partial final bundle + assert len(list(ready_directory.iterdir())) == 2 + assert not first_output.exists() + assert not second_output.exists() + release_file.touch() + first_stdout, first_stderr = first_process.communicate(timeout=5) + second_stdout, second_stderr = second_process.communicate(timeout=5) + + # And both bundles are separately and atomically published for their run + assert first_process.returncode == 0, first_stderr + assert second_process.returncode == 0, second_stderr + assert first_output.read_text() == '{"issue":"TC-42"}\n' + assert second_output.read_text() == '{"issue":"TC-43"}\n' + assert sorted(path.name for path in first_output.parent.iterdir()) == ["triage-security-input.json"] + assert sorted(path.name for path in second_output.parent.iterdir()) == ["triage-security-input.json"] + assert not shared_output.exists() + with open(os.path.join(SCRIPT_DIR, "..", "..", "..", "harness", "triage-security.yaml")) as harness_file: + assert "src: ${FULLSEND_RUN_DIR}/pre/triage-security-input.json" in harness_file.read() + + +def test_normalize_issue_rejects_missing_reporter_metadata(): + """A Jira response without reporter evidence cannot enter the sandbox bundle.""" + # Given a malformed Jira issue response without its required reporter + issue = {"key": "TC-42", "fields": { + "summary": "CVE-2026-12345", "description": {}, "status": {"name": "New"}, + "labels": [], "versions": [], "comment": {"comments": []}, + }} + + # When the runner normalizes the issue for the schema + # Then it fails loudly rather than emit incomplete audit evidence + with pytest.raises(pre_triage_security.EvidenceError, match="reporter"): + pre_triage_security.normalize_issue(issue) + + +def test_normalize_remote_links_rejects_missing_evidence(): + """A security issue without remote-link evidence is rejected before sandbox creation.""" + # Given a Jira issue whose required remote-link prefetch returned no links + # When the runner normalizes the evidence + # Then it rejects the incomplete input rather than silently continue + with pytest.raises(pre_triage_security.EvidenceError, match="remote_links"): + pre_triage_security.normalize_remote_links([]) + + +def test_matrix_validation_rejects_non_retag_without_pinned_commit(): + """A released matrix row without a pinned source commit is rejected.""" + # Given a non-retag release row with no source commit evidence + matrix = {"streams": [{ + "name": "1.0.x", "matrix_source": "matrix", "rows": [{ + "version": "1.0.0", "source_commits": {}, "retag_of": None, + }], + }]} + + # When the runner validates the matrix before reading source repositories + # Then it fails instead of asking the sandbox to infer a ref + with pytest.raises(pre_triage_security.EvidenceError, match="no source commits"): + pre_triage_security._validate_matrix(matrix) + + +def test_source_evidence_accepts_rpm_system_package_reads(): + """RPM lock-file evidence is retained for system-package CVE analysis.""" + # Given read-only git-show evidence for a released and development RPM lock file + evidence = {"lock_files": [{ + "repository": "component", "ref": "abc1234", "path": "rpms.lock.yaml", + "command": "git show abc1234:rpms.lock.yaml", "content": "openssl: 3.0.7", + }], "development_streams": [{ + "repository": "component", "ref": "main", "path": "rpms.lock.yaml", + "command": "git show main:rpms.lock.yaml", "content": "openssl: 3.0.8", + }]} + + # When the runner validates the system-package evidence + result = pre_triage_security._validate_source_evidence(evidence) + + # Then it preserves the RPM reads for the sandbox rather than assuming Cargo or npm + assert result["lock_files"][0]["path"] == "rpms.lock.yaml" + assert result["development_streams"][0]["content"] == "openssl: 3.0.8" + + +def test_action_markers_collect_prior_trusted_runner_actions(): + """Existing triage action markers are preserved for idempotency decisions.""" + # Given comment history containing a prior trusted-runner action marker + issue = {"fields": {"comment": {"comments": [{"body": { + "type": "doc", "content": [{"type": "text", "text": "triage-security:created-remediation"}], + }}]}}} + + # When idempotency evidence is extracted + # Then the sandbox receives the prior action marker exactly once + assert pre_triage_security._action_markers(issue) == ["triage-security:created-remediation"] + + +def test_parse_security_matrix_preserves_rows_retags_and_ecosystem_commands(): + """Matrix parsing normalizes source refs and accepts empty branch cells.""" + # Given a configured stream matrix with populated and empty upstream branches + matrix_markdown = """## Supportability Matrix + +| PRODUCT Version | Build | component | Notes | +|---|---|---|---| +| 1.0.0 | build-1 | `abc1234` (retag) | | +| 1.0.1 | build-2 | `abc1234` | retag of 1.0.0 | + +## Ecosystem Mappings + +| Ecosystem | Repository | Lock File | Check Command | Upstream Branch | +|---|---|---|---|---| +| Cargo | component | `Cargo.lock` | grep library | `release/0.6.z` (pending re-point to 0.7.z) | +| RPM | component | `rpms.lock.yaml` | grep library | main | +| Go | component | go.sum | grep library | | +""" + + # When parsing the matrix on the trusted runner + matrix, mappings = pre_triage_security.parse_security_matrix( + stream_name="1.0.x", matrix_path="docs/matrix.md", content=matrix_markdown) + + # Then the sandbox receives both product rows and all configured read commands + assert matrix == { + "name": "1.0.x", + "matrix_source": matrix_markdown, + "rows": [ + {"version": "1.0.0", "source_commits": {"component": "abc1234"}, "retag_of": None}, + {"version": "1.0.1", "source_commits": {"component": "abc1234"}, "retag_of": "1.0.0"}, + ], + } + assert mappings == [ + {"ecosystem": "Cargo", "repository": "component", "lock_file": "Cargo.lock", "check_command": "grep library", "upstream_branch": "release/0.6.z"}, + {"ecosystem": "RPM", "repository": "component", "lock_file": "rpms.lock.yaml", "check_command": "grep library", "upstream_branch": "main"}, + {"ecosystem": "Go", "repository": "component", "lock_file": "go.sum", "check_command": "grep library", "upstream_branch": ""}, + ] + + +@pytest.mark.parametrize("cell, expected", [ + ("`release/0.6.z` (pending re-point to 0.7.z)", "release/0.6.z"), + ("release/0.5.z", "release/0.5.z"), + ("`main` — current stable branch", "main"), + ("(pending re-point) release/0.6.z", "release/0.6.z"), + ("N/A - see release/0.6.z", "release/0.6.z"), +]) +def test_ref_token_extracts_branch_from_annotated_matrix_cells(cell, expected): + """A matrix branch cell yields its ref token despite surrounding prose.""" + assert pre_triage_security._ref_token(cell) == expected + + +@pytest.mark.parametrize("cell", ["", None, " ", "``"]) +def test_ref_token_returns_empty_for_blank_cells(cell): + """A blank upstream-branch cell normalizes to an empty ref.""" + assert pre_triage_security._ref_token(cell) == "" + + +def _complete_bundle(): + """Build complete source-dependency evidence for bundle validation tests.""" + # Given a CVE issue and all evidence gathered by the trusted runner + issue = { + "key": "TC-42", + "fields": { + "summary": "CVE-2026-12345 component: library issue", + "description": {"type": "doc", "content": []}, + "status": {"name": "New"}, + "labels": ["CVE-2026-12345", "pscomponent:org/component"], + "versions": [{"id": "1", "name": "1.0", "released": False, "self": "https://jira.example.com/version/1"}], + "reporter": {"accountId": "reporter-1", "displayName": "Reporter"}, + "comment": {"comments": []}, + "issuelinks": [], + }, + } + configuration = { + "project_key": "TC", + "jira_version_prefix": "PRODUCT", + "vulnerability_issue_type_id": "10016", + "component_label_pattern": "pscomponent:", + "version_streams": [{ + "name": "1.0.x", + "matrix_path": "security-matrix.md", + "release_repository": "release-repo", + }], + "source_repositories": [{ + "name": "component", + "url": "https://github.com/org/component", + "deployment_context": "upstream", + }], + } + external_evidence = { + name: { + "source_url": url, + "retrieved_at": "2026-09-21T12:00:00Z", + "status": 200, + "body": {"id": "CVE-2026-12345"}, + } + for name, url in { + "mitre": "https://cveawg.mitre.org/api/cve/CVE-2026-12345", + "osv": "https://api.osv.dev/v1/vulns/CVE-2026-12345", + "lifecycle": "https://example.com/lifecycle", + }.items() + } + matrix = {"streams": [{ + "name": "1.0.x", + "matrix_source": "security-matrix.md", + "rows": [{"version": "1.0", "source_commits": {"component": "abc1234"}, "retag_of": None}], + }]} + source_evidence = { + "lock_files": [{ + "repository": "component", + "ref": "abc1234", + "path": "Cargo.lock", + "command": "git show abc1234:Cargo.lock", + "content": 'name = "library"\\nversion = "1.2.3"', + }], + "development_streams": [{ + "repository": "component", + "ref": "main", + "path": "Cargo.lock", + "command": "git show main:Cargo.lock", + "content": 'name = "library"\\nversion = "1.2.3"', + }], + } + + # When transforming the runner evidence for the sandbox + bundle = pre_triage_security.build_bundle( + issue=issue, + remote_links=[{"object": {"url": "https://github.com/org/component/pull/1", "title": "Fix"}}], + configuration=configuration, + external_evidence=external_evidence, + matrix=matrix, + source_evidence=source_evidence, + jira_metadata={"versions": [], "sibling_searches": [], "related_issues": []}, + idempotency={"action_markers": [], "existing_remediation": []}, + mutation_authorized=False, + ) + + # Then the exact sandbox contract accepts the resulting bundle + with open(os.path.join(SCRIPT_DIR, "..", "schemas", "triage-security-input.schema.json")) as schema_file: + schema = json.load(schema_file) + validate(instance=bundle, schema=schema) + assert bundle["issue"]["key"] == "TC-42" + assert bundle["issue"]["versions"] == [{"id": "1", "name": "1.0", "released": False}] + assert bundle["remote_links"] == [{"url": "https://github.com/org/component/pull/1", "title": "Fix"}] + return bundle + + +@requires_format_extra +def test_build_bundle_accepts_complete_source_dependency_evidence(): + """Complete source-dependency evidence produces schema-valid sandbox input.""" + assert _complete_bundle()["issue"]["key"] == "TC-42" + + +@requires_format_extra +def test_validate_bundle_rejects_malformed_uri(): + """A malformed remote-link URL cannot enter the sandbox bundle.""" + # Given an otherwise valid bundle with an invalid URI-format field + bundle = _complete_bundle() + bundle["remote_links"][0]["url"] = "not a URL" + + # When the bundle is schema-validated before the sandbox runs + # Then format validation rejects the malformed URL + with pytest.raises(pre_triage_security.EvidenceError, match="triage-security input validation failed"): + pre_triage_security.validate_bundle(bundle) + + +@requires_format_extra +def test_validate_bundle_rejects_malformed_retrieval_timestamp(): + """A malformed evidence retrieval timestamp cannot enter the sandbox bundle.""" + # Given an otherwise valid bundle with an invalid date-time-format field + bundle = _complete_bundle() + bundle["external_evidence"]["mitre"]["retrieved_at"] = "not-a-timestamp" + + # When the bundle is schema-validated before the sandbox runs + # Then format validation rejects the malformed timestamp + with pytest.raises(pre_triage_security.EvidenceError, match="triage-security input validation failed"): + pre_triage_security.validate_bundle(bundle) + + +def test_validate_bundle_fails_closed_without_format_extra(monkeypatch): + """A runner missing jsonschema[format] aborts instead of skipping format checks.""" + # Given a FormatChecker with no uri/date-time checkers (jsonschema[format] absent) + class _NoFormatChecker: + checkers = {"regex": None} + + monkeypatch.setattr(pre_triage_security, "FormatChecker", _NoFormatChecker) + + # When a bundle is validated on that misprovisioned runner + # Then it fails closed, naming the missing extra, rather than validating silently + with pytest.raises(pre_triage_security.EvidenceError, match=r"jsonschema\[format\]"): + pre_triage_security.validate_bundle({"any": "bundle"}) + + +def _http_error_opener(status, body=b'{"message": "gone"}'): + """Return a urlopen replacement that raises an HTTPError with the given status.""" + def opener(url, timeout=30): + raise HTTPError(url, status, "error", {}, io.BytesIO(body)) + return opener + + +def test_fetch_url_records_missing_external_evidence(monkeypatch): + """A 404 from an incomplete evidence source is recorded, not fatal, when allowed.""" + # Given an OSV/MITRE source that does not track this CVE (a routine 404) + monkeypatch.setattr(pre_triage_security, "urlopen", _http_error_opener(404)) + + # When the runner fetches it as tolerable-missing evidence + evidence = pre_triage_security._fetch_url( + "https://api.osv.dev/v1/vulns/CVE-2026-12345", allow_missing=True) + + # Then the 404 is preserved as evidence rather than aborting the whole bundle + assert evidence["status"] == 404 + assert evidence["body"] == {"message": "gone"} + + # And a 404 on required (non-missing-tolerant) evidence still fails loudly + with pytest.raises(pre_triage_security.EvidenceError, match="HTTP 404"): + pre_triage_security._fetch_url("https://api.osv.dev/v1/vulns/CVE-2026-12345") + + # And any non-404 failure stays fatal even when misses are tolerated + monkeypatch.setattr(pre_triage_security, "urlopen", _http_error_opener(500)) + with pytest.raises(pre_triage_security.EvidenceError, match="HTTP 500"): + pre_triage_security._fetch_url( + "https://api.osv.dev/v1/vulns/CVE-2026-12345", allow_missing=True) + + +def test_fetch_url_sends_custom_user_agent(monkeypatch): + """Evidence fetches send the configured User-Agent on a urllib Request.""" + # Given a successful HTTP response and an opener that records its request + response = MagicMock() + response.__enter__.return_value = response + response.status = 200 + response.read.return_value = b'{"evidence": "ok"}' + calls = [] + + def open_url(request, timeout): + calls.append((request, timeout)) + return response + + monkeypatch.setattr(pre_triage_security, "urlopen", open_url) + + # When evidence is fetched + evidence = pre_triage_security._fetch_url("https://access.redhat.com/advisory") + + # Then urllib receives the custom User-Agent and the response is decoded + request, timeout = calls[0] + assert isinstance(request, Request) + assert request.full_url == "https://access.redhat.com/advisory" + assert request.get_header("User-agent") == pre_triage_security._EVIDENCE_USER_AGENT + assert timeout == 30 + assert evidence["body"] == {"evidence": "ok"} + + +def _init_source_repository(path): + """Create a clean source repository for read-only git evidence tests.""" + subprocess.run( + ["git", "-C", str(path), "init", "--quiet", "--initial-branch=main"], + check=True, + ) + for key, value in (("user.name", "Test User"), ("user.email", "test@example.com")): + subprocess.run(["git", "-C", str(path), "config", key, value], check=True) + (path / "README.md").write_text("source evidence\n") + subprocess.run(["git", "-C", str(path), "add", "README.md"], check=True) + subprocess.run(["git", "-C", str(path), "commit", "--quiet", "-m", "fixture"], check=True) + return subprocess.check_output( + ["git", "-C", str(path), "rev-parse", "HEAD"], text=True).strip() + + +def test_resolve_ref_prefers_remotes_and_returns_none_when_missing(tmp_path): + """Bare branches resolve to remote refs in precedence order, or None.""" + # Given a source repository with the same branch on several remotes + repository = tmp_path / "source" + repository.mkdir() + commit = _init_source_repository(repository) + branch = "release/0.5.z" + for remote in ("backup-z", "backup-a", "origin", "upstream"): + subprocess.run([ + "git", "-C", str(repository), "update-ref", + "refs/remotes/{}/{}".format(remote, branch), commit, + ], check=True) + + # When the bare branch is resolved with every remote available + assert pre_triage_security._resolve_ref(repository, "main") == "main" + assert pre_triage_security._resolve_ref(repository, branch) == ( + "refs/remotes/upstream/{}".format(branch)) + + # Then origin wins after upstream is removed, followed by the first other remote + subprocess.run([ + "git", "-C", str(repository), "update-ref", "-d", + "refs/remotes/upstream/{}".format(branch), + ], check=True) + assert pre_triage_security._resolve_ref(repository, branch) == ( + "refs/remotes/origin/{}".format(branch)) + subprocess.run([ + "git", "-C", str(repository), "update-ref", "-d", + "refs/remotes/origin/{}".format(branch), + ], check=True) + assert pre_triage_security._resolve_ref(repository, branch) == ( + "refs/remotes/backup-a/{}".format(branch)) + assert pre_triage_security._resolve_ref(repository, "missing-branch") is None + + +def test_git_show_records_resolved_ref_without_mutating_source_repository(tmp_path): + """Remote-tracking evidence keeps the matrix ref and leaves the checkout intact.""" + # Given a clean source checkout with a remote-tracking development branch + repository = tmp_path / "source" + repository.mkdir() + commit = _init_source_repository(repository) + branch = "release/0.5.z" + subprocess.run([ + "git", "-C", str(repository), "update-ref", + "refs/remotes/upstream/{}".format(branch), commit, + ], check=True) + before_head = subprocess.check_output( + ["git", "-C", str(repository), "rev-parse", "HEAD"], text=True).strip() + before_status = subprocess.check_output( + ["git", "-C", str(repository), "status", "--porcelain"], text=True) + + # When evidence is read from the bare matrix branch + evidence = pre_triage_security._git_show(repository, branch, "README.md") + + # Then provenance records both refs and the source worktree remains unchanged + assert evidence == { + "ref": branch, + "path": "README.md", + "command": "git show refs/remotes/upstream/{}:README.md".format(branch), + "content": "source evidence\n", + } + assert subprocess.check_output( + ["git", "-C", str(repository), "rev-parse", "HEAD"], text=True).strip() == before_head + assert subprocess.check_output( + ["git", "-C", str(repository), "status", "--porcelain"], text=True) == before_status + + +def test_jira_client_decodes_empty_collection(monkeypatch): + """An empty Jira collection is decoded as [] rather than a misleading JSON error.""" + # Given jira-client.py that now prints an empty collection as valid JSON + class _Result: + stdout = "[]\n" + + monkeypatch.setattr(pre_triage_security.subprocess, "run", lambda *a, **k: _Result()) + + # When the runner reads a command with no results (e.g. get_versions) + # Then it returns the empty list, not an "invalid JSON" evidence error + assert pre_triage_security._jira_client("get_versions", "TC") == [] + + +def test_parse_security_configuration_defaults_blank_deployment_context(): + """A present-but-blank Deployment Context cell falls back to 'upstream'.""" + # Given a Source Repositories table whose Deployment Context cell is left blank + claude_md = """# Project Configuration + +## Jira Configuration + +- Project key: TC + +## Security Configuration + +### Product Lifecycle + +- Product pages URL: https://example.com/lifecycle +- Jira version prefix: PRODUCT +- Vulnerability issue type ID: 10016 +- Component label pattern: pscomponent: + +### Version Streams + +| Stream | Konflux Release Repo | Local Path | Security Matrix Path | +|---|---|---|---| +| 1.0.x | release-repo | /repos/release | docs/matrix.md | + +### Source Repositories + +| Repository | URL | Deployment Context | +|---|---|---| +| component | https://github.com/org/component | | +""" + + # When the runner parses the configuration + configuration = pre_triage_security.parse_security_configuration(claude_md) + + # Then the blank cell defaults instead of being rejected as an incomplete row + assert configuration["source_repositories"] == [{ + "name": "component", + "url": "https://github.com/org/component", + "deployment_context": "upstream", + }] + + +def test_parse_security_configuration_rejects_blank_repository(): + """A blank required Source Repositories cell still fails loudly.""" + # Given a Source Repositories table missing the Repository value + claude_md = """# Project Configuration + +## Jira Configuration + +- Project key: TC + +## Security Configuration + +### Product Lifecycle + +- Product pages URL: https://example.com/lifecycle +- Jira version prefix: PRODUCT +- Vulnerability issue type ID: 10016 +- Component label pattern: pscomponent: + +### Version Streams + +| Stream | Konflux Release Repo | Local Path | Security Matrix Path | +|---|---|---|---| +| 1.0.x | release-repo | /repos/release | docs/matrix.md | + +### Source Repositories + +| Repository | URL | Deployment Context | +|---|---|---| +| | https://github.com/org/component | upstream | +""" + + # When the runner parses the configuration + # Then it rejects the row rather than emit a nameless source repository + with pytest.raises(pre_triage_security.EvidenceError, match="Repository or URL"): + pre_triage_security.parse_security_configuration(claude_md) + + +def _minimal_jira_client(): + """Return a _jira_client stub sufficient to reach collect_bundle's stream loop.""" + def jira_client(command, *arguments): + if command == "get_issue": + return {"key": "TC-42", "fields": {}} + if command in ("get_remote_links", "get_versions"): + return [] + if command == "search_jql": + return {"issues": []} + raise AssertionError("unexpected Jira command: {}".format(command)) + return jira_client + + +def test_collect_bundle_reports_missing_version_streams_column(tmp_path, monkeypatch): + """A Version Streams table missing a column raises a clear error, not a traceback.""" + # Given a runner configuration whose stream row lacks the Local Path column + (tmp_path / "CLAUDE.md").write_text("# Project Configuration\n") + configuration = {"project_key": "TC", "vulnerability_issue_type_id": "10016"} + stream_rows = [{"Stream": "1.0.x", "Security Matrix Path": "docs/matrix.md"}] + monkeypatch.setattr( + pre_triage_security, "_runner_configuration", + lambda _content, _root: (configuration, "https://example.com/lifecycle", stream_rows, {})) + monkeypatch.setattr(pre_triage_security, "_jira_client", _minimal_jira_client()) + monkeypatch.setattr(pre_triage_security, "extract_cve_id", lambda _issue: "CVE-2026-12345") + monkeypatch.setattr(pre_triage_security, "_fetch_url", lambda *a, **k: {}) + + # When the runner reaches the stream loop + # Then the missing column surfaces as an EvidenceError naming it + with pytest.raises(pre_triage_security.EvidenceError, match="Local Path"): + pre_triage_security.collect_bundle("TC-42", tmp_path) + + +def test_related_issue_handles_null_comment_field(): + """A related issue whose comment field is JSON null yields no comments, not a crash.""" + # Given a related issue with restricted comment visibility (comment field is null) + issue = {"key": "TC-9", "fields": {"status": {"name": "New"}, "comment": None}} + + # When the runner normalizes it for audit and idempotency + # Then the null comment field degrades to an empty list rather than raising + assert pre_triage_security._related_issue(issue)["comments"] == [] + assert pre_triage_security._action_markers(issue) == [] + + +def test_collect_bundle_rejects_malformed_issue_key(tmp_path): + """A non-conforming issue key is rejected before any JQL is constructed.""" + # Given a poller-supplied value that is not a valid Jira issue key + (tmp_path / "CLAUDE.md").write_text("# Project Configuration\n") + + # When the runner begins collection + # Then it fails on the key before interpolating it into any JQL + with pytest.raises(pre_triage_security.EvidenceError, match="issue_key"): + pre_triage_security.collect_bundle("not-a-key", tmp_path) diff --git a/plugins/sdlc-workflow/scripts/test_pre_verify_pr.py b/plugins/sdlc-workflow/scripts/test_pre_verify_pr.py new file mode 100644 index 000000000..b4dd8d164 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/test_pre_verify_pr.py @@ -0,0 +1,1671 @@ +#!/usr/bin/env python3 +"""Tests for pre_verify_pr.py — PR URL extraction, GitHub bundle, transform.""" + +import json +import os +import re +import subprocess +import sys +import tempfile + +script_dir = os.path.dirname(os.path.abspath(__file__)) +sys.path.insert(0, script_dir) +import pre_verify_pr + + +# --- extract_pr_url --- + +def test_extract_pr_url_adf_inline_card(): + issue = {"fields": {"customfield_10875": { + "type": "doc", "version": 1, + "content": [{"type": "paragraph", "content": [ + {"type": "inlineCard", "attrs": {"url": "https://github.com/org/repo/pull/42"}} + ]}] + }}} + result = pre_verify_pr.extract_pr_url(issue) + assert result == "https://github.com/org/repo/pull/42", f"Got: {result}" + + +def test_extract_pr_url_plain_string(): + issue = {"fields": {"customfield_10875": "https://github.com/org/repo/pull/7"}} + result = pre_verify_pr.extract_pr_url(issue) + assert result == "https://github.com/org/repo/pull/7", f"Got: {result}" + + +def test_extract_pr_url_missing_field(): + issue = {"fields": {}} + result = pre_verify_pr.extract_pr_url(issue) + assert result == "", f"Expected empty string, got: {result}" + + +def test_extract_pr_url_null_field(): + issue = {"fields": {"customfield_10875": None}} + result = pre_verify_pr.extract_pr_url(issue) + assert result == "", f"Expected empty string, got: {result}" + + +def test_extract_pr_url_adf_no_inline_card(): + issue = {"fields": {"customfield_10875": { + "type": "doc", "version": 1, + "content": [{"type": "paragraph", "content": [ + {"type": "text", "text": "no link here"} + ]}] + }}} + result = pre_verify_pr.extract_pr_url(issue) + assert result == "", f"Expected empty string, got: {result}" + + +def test_extract_pr_url_adf_text_link_mark(): + """ADF text node carrying a link mark (how a manually-typed link is stored).""" + issue = {"fields": {"customfield_10875": { + "type": "doc", "version": 1, + "content": [{"type": "paragraph", "content": [ + {"type": "text", "text": "PR", "marks": [ + {"type": "link", "attrs": { + "href": "https://github.com/org/repo/pull/13"}} + ]} + ]}] + }}} + result = pre_verify_pr.extract_pr_url(issue) + assert result == "https://github.com/org/repo/pull/13", f"Got: {result}" + + +def test_extract_pr_url_adf_plain_text_url(): + """ADF plain text that merely contains a bare URL (no mark, no card).""" + issue = {"fields": {"customfield_10875": { + "type": "doc", "version": 1, + "content": [{"type": "paragraph", "content": [ + {"type": "text", "text": "see https://github.com/org/repo/pull/99 for details"} + ]}] + }}} + result = pre_verify_pr.extract_pr_url(issue) + assert result == "https://github.com/org/repo/pull/99", f"Got: {result}" + + +def test_extract_pr_url_string_with_surrounding_text(): + """Plain-string field whose value embeds a URL among other text.""" + issue = {"fields": {"customfield_10875": "PR: https://github.com/org/repo/pull/5."}} + result = pre_verify_pr.extract_pr_url(issue) + assert result == "https://github.com/org/repo/pull/5", f"Got: {result}" + + +# --- build_github_bundle --- + +def test_build_github_bundle(): + bundle = pre_verify_pr.build_github_bundle( + pr_repo="org/repo", pr_number="42", head_ref="feat/x", + commit_sha="abc1234", diff="diff --git a b", stat=" 1 file changed", + reviews=[{"id": 1}], review_comments=[{"id": 2}], + issue_comments=[{"id": 3}], commits=[{"oid": "abc1234"}], + check_runs=[{"name": "pytest", "status": "completed", + "conclusion": "success", "details_url": "https://ci/1"}], + ) + assert bundle["pr_repo"] == "org/repo" + assert bundle["pr_number"] == 42 # coerced to int + assert bundle["headRefName"] == "feat/x" + assert bundle["commit_sha"] == "abc1234" + assert bundle["diff"] == "diff --git a b" + assert bundle["stat"] == " 1 file changed" + assert bundle["reviews"] == [{"id": 1}] + assert bundle["review_comments"] == [{"id": 2}] + assert bundle["issue_comments"] == [{"id": 3}] + assert bundle["commits"] == [{"oid": "abc1234"}] + assert bundle["check_runs"] == [{"name": "pytest", "status": "completed", + "conclusion": "success", + "details_url": "https://ci/1"}] + # No check_run_logs_path argument → "" (green path, nothing to read) + assert bundle["check_run_logs_path"] == "" + + +def test_build_github_bundle_carries_check_run_logs_path(): + """When a check failed, the bundle carries the mounted log-file path. + + The log text itself is never inlined — only the sandbox path, so Check 1b can + Read it on demand on a FAIL. + """ + bundle = pre_verify_pr.build_github_bundle( + "o/r", 5, "b", "deadbee", "d", "s", [], [], [], [], + check_run_logs_path=pre_verify_pr.SANDBOX_CHECK_RUN_LOGS_PATH, + ) + assert bundle["check_run_logs_path"] == pre_verify_pr.SANDBOX_CHECK_RUN_LOGS_PATH + + +def test_build_github_bundle_check_runs_defaults_to_empty_list(): + """A PR head with no checks yields check_runs=[], never absent. + + The github object is additionalProperties:false with check_runs required, so + the key must always be present for the prefetch to validate against + verify-pr-input.schema.json. + """ + # Given a bundle built without an explicit check_runs argument + bundle = pre_verify_pr.build_github_bundle( + "o/r", 5, "b", "deadbee", "d", "s", [], [], [], [], + ) + + # Then check_runs is present and defaults to an empty list + assert bundle["check_runs"] == [] + + +# --- filter_own_check_runs (CI Status self-exclusion, TC-6343) --- + +# The head-SHA check-run names verify-pr's own workflow produces (wait-for-checks +# plus the reusable-dispatch harness job). Reused across the filter tests as the +# own-workflow name set the trusted runner gathers. +OWN_CHECK_NAMES = { + "wait-for-checks", + "verify-pr / harness-run (verify-pr, review)", +} + +# A matrix of substantive checks a real PR carries, all terminal + green. +SUBSTANTIVE_CHECK_RUNS = [ + {"name": "pytest", "status": "completed", "conclusion": "success", + "details_url": "https://ci/pytest"}, + {"name": "validate-plugins", "status": "completed", "conclusion": "success", + "details_url": "https://ci/validate"}, + {"name": "skillsaw", "status": "completed", "conclusion": "success", + "details_url": "https://ci/skillsaw"}, + {"name": "Sourcery", "status": "completed", "conclusion": "success", + "details_url": "https://ci/sourcery"}, +] + + +def test_filter_own_check_runs_reproduces_and_fixes_ci_status_bug(): + """Reproducer for TC-6332/TC-6343: own-harness runs are dropped, substantive kept. + + Before the fix the prefetch kept verify-pr's own workflow check-runs, whose + non-terminal/failed states dragged CI Status to a permanent self-referential + WARN/FAIL. The filter removes exactly those, leaving an all-passing set that + maps to CI Status = PASS under correctness.md Check 1a. + """ + # Given all-passing substantive checks mixed with verify-pr's OWN harness + # check-runs: one still in_progress (non-terminal) and one from a superseded + # prior attempt that failed. + own_in_progress = { + "name": "verify-pr / harness-run (verify-pr, review)", + "status": "in_progress", "conclusion": None, + "details_url": "https://ci/actions/runs/2/job/9"} + own_failed_superseded = { + "name": "wait-for-checks", "status": "completed", "conclusion": "failure", + "details_url": "https://ci/actions/runs/1/job/8"} + check_runs = SUBSTANTIVE_CHECK_RUNS + [own_in_progress, own_failed_superseded] + + # When filtering with the own-workflow check-run names gathered on the runner + filtered = pre_verify_pr.filter_own_check_runs(check_runs, OWN_CHECK_NAMES) + + # Then only the substantive checks remain — none of the own-harness entries — + # and every survivor is a terminal success (CI Status can now be PASS). + assert filtered == SUBSTANTIVE_CHECK_RUNS + assert own_in_progress not in filtered + assert own_failed_superseded not in filtered + assert all(c["conclusion"] == "success" for c in filtered) + + +def test_filter_own_check_runs_empty_input_yields_empty(): + """An empty check-run list filters to an empty list.""" + assert pre_verify_pr.filter_own_check_runs([], OWN_CHECK_NAMES) == [] + + +def test_filter_own_check_runs_no_own_runs_returned_unchanged(): + """A list with no own-workflow runs passes through unchanged.""" + # Given only substantive checks, none matching an own-workflow name + check_runs = list(SUBSTANTIVE_CHECK_RUNS) + + # When filtering, Then the list is returned unchanged (same items, same order) + assert pre_verify_pr.filter_own_check_runs(check_runs, OWN_CHECK_NAMES) == \ + SUBSTANTIVE_CHECK_RUNS + + +def test_filter_own_check_runs_only_own_runs_yields_empty(): + """A list containing only own-workflow runs filters to empty.""" + # Given a list made up solely of verify-pr's own harness check-runs + check_runs = [ + {"name": "wait-for-checks", "status": "completed", "conclusion": "success"}, + {"name": "verify-pr / harness-run (verify-pr, review)", + "status": "in_progress", "conclusion": None}, + ] + + # When filtering with those names, Then nothing survives + assert pre_verify_pr.filter_own_check_runs(check_runs, OWN_CHECK_NAMES) == [] + + +def test_filter_own_check_runs_empty_names_is_noop(): + """An empty own-name set (enumeration failed) keeps prior behavior: no removal.""" + check_runs = list(SUBSTANTIVE_CHECK_RUNS) + assert pre_verify_pr.filter_own_check_runs(check_runs, set()) == \ + SUBSTANTIVE_CHECK_RUNS + + +def test_filter_own_check_runs_preserves_fields_verbatim(): + """Surviving substantive check-runs keep name/status/conclusion/details_url intact.""" + # Given a matrix of substantive checks and a disjoint own-name set + # When filtering + filtered = pre_verify_pr.filter_own_check_runs( + SUBSTANTIVE_CHECK_RUNS, OWN_CHECK_NAMES) + + # Then every field of every survivor is preserved verbatim + assert filtered == SUBSTANTIVE_CHECK_RUNS + for original, kept in zip(SUBSTANTIVE_CHECK_RUNS, filtered): + assert kept["name"] == original["name"] + assert kept["status"] == original["status"] + assert kept["conclusion"] == original["conclusion"] + assert kept["details_url"] == original["details_url"] + + +def test_filter_own_check_runs_does_not_mutate_input(): + """The input list is not mutated; a new list is returned.""" + # Given an input list carrying an own-workflow run + check_runs = SUBSTANTIVE_CHECK_RUNS + [ + {"name": "wait-for-checks", "status": "completed", "conclusion": "failure"}] + before = list(check_runs) + + # When filtering + pre_verify_pr.filter_own_check_runs(check_runs, OWN_CHECK_NAMES) + + # Then the original list is untouched + assert check_runs == before + + +def test_filter_own_check_runs_keeps_nameless_entries(): + """A check-run with no name is kept (never treated as an own-workflow run).""" + check_runs = [{"status": "completed", "conclusion": "success"}] + assert pre_verify_pr.filter_own_check_runs(check_runs, OWN_CHECK_NAMES) == \ + check_runs + + +# --- _read_own_check_names (own-name file reader) --- + +def test_read_own_check_names_reads_names_one_per_line(): + """Names are read one per line, with surrounding whitespace stripped.""" + with tempfile.NamedTemporaryFile("w", suffix=".txt", delete=False) as f: + f.write("wait-for-checks\nverify-pr / harness-run (verify-pr, review)\n") + path = f.name + try: + # A job name containing spaces survives intact (line-based, not token-based) + assert pre_verify_pr._read_own_check_names(path) == OWN_CHECK_NAMES + finally: + os.unlink(path) + + +def test_read_own_check_names_none_path_yields_empty_set(): + """A falsy path (option not supplied) yields an empty set, no error.""" + assert pre_verify_pr._read_own_check_names(None) == set() + + +def test_read_own_check_names_empty_file_yields_empty_set(): + """An empty names file (enumeration produced nothing) yields an empty set.""" + with tempfile.NamedTemporaryFile("w", suffix=".txt", delete=False) as f: + path = f.name + try: + assert pre_verify_pr._read_own_check_names(path) == set() + finally: + os.unlink(path) + + +# --- transform_to_input --- + +def test_transform_basic(): + issue = {"fields": { + "summary": "Add feature X", + "description": {"type": "doc", "content": []}, + "status": {"name": "In Progress"}, + "labels": ["backend", "api"], + "issuelinks": [], + }} + result = pre_verify_pr.transform_to_input(issue, "TC-100", "https://github.com/o/r/pull/1") + assert result["task_id"] == "TC-100" + assert result["task"]["summary"] == "Add feature X" + assert result["task"]["status"] == "In Progress" + assert result["task"]["labels"] == ["backend", "api"] + assert result["task"]["issue_links"] == [] + assert result["pr_url"] == "https://github.com/o/r/pull/1" + assert result["source"]["tracker"] == "jira" + assert result["source"]["raw"] is issue + + +def test_transform_without_github_omits_key(): + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], "issuelinks": []}} + result = pre_verify_pr.transform_to_input(issue, "TC-1", "") + assert "github" not in result + + +def test_transform_with_github(): + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], "issuelinks": []}} + github = pre_verify_pr.build_github_bundle( + "o/r", 5, "b", "deadbee", "d", "s", [], [], [], [], + ) + result = pre_verify_pr.transform_to_input(issue, "TC-1", "https://github.com/o/r/pull/5", github) + assert result["github"]["pr_repo"] == "o/r" + assert result["github"]["pr_number"] == 5 + assert result["github"]["commit_sha"] == "deadbee" + + +def test_transform_embeds_check_runs_under_github(): + """The CI check-run outcomes are embedded under the github bundle. + + Mirrors the reviews/comments bundle tests: correctness.md Check 1 reads + github.check_runs in sandbox mode, so the transform must pass them through. + """ + # Given an issue and a github bundle carrying head-SHA CI check-run outcomes + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], "issuelinks": []}} + check_runs = [ + {"name": "pytest", "status": "completed", "conclusion": "success", + "details_url": "https://ci/pytest"}, + {"name": "skillsaw", "status": "completed", "conclusion": "failure", + "details_url": "https://ci/skillsaw"}, + ] + github = pre_verify_pr.build_github_bundle( + "o/r", 5, "b", "deadbee", "d", "s", [], [], [], [], check_runs, + ) + + # When transforming to the tracker-agnostic input + result = pre_verify_pr.transform_to_input( + issue, "TC-1", "https://github.com/o/r/pull/5", github) + + # Then the check-run outcomes are embedded verbatim under github.check_runs + assert result["github"]["check_runs"] == check_runs + + +# --- commit_references_task (Commit Traceability determinism) --- + +def test_commit_references_task_in_body_trailer(): + # The canonical failure mode: the ID lives only in the body trailer, far + # past where a subject-only or truncated read would look. + commit = { + "messageHeadline": "feat(verify-pr): re-sync sandbox dual-mode onto SKILL.md", + "messageBody": "Long body paragraph ...\n\nImplements TC-5812\n\nAssisted-by: x", + } + assert pre_verify_pr.commit_references_task(commit, "TC-5812") is True + + +def test_commit_references_task_in_headline(): + commit = {"messageHeadline": "TC-5812: fix scope", "messageBody": ""} + assert pre_verify_pr.commit_references_task(commit, "TC-5812") is True + + +def test_commit_references_task_absent(): + commit = {"messageHeadline": "fix: thing", "messageBody": "no id here"} + assert pre_verify_pr.commit_references_task(commit, "TC-5812") is False + + +def test_commit_references_task_word_boundary(): + # A superstring ID must not match, nor a different task in the same family. + assert pre_verify_pr.commit_references_task( + {"messageHeadline": "x", "messageBody": "see TC-58120"}, "TC-5812") is False + assert pre_verify_pr.commit_references_task( + {"messageHeadline": "x", "messageBody": "the TC-5982 fix"}, "TC-5812") is False + + +def test_commit_references_task_missing_body_key(): + assert pre_verify_pr.commit_references_task( + {"messageHeadline": "TC-5812: x"}, "TC-5812") is True + + +def test_transform_annotates_commit_references_task_id(): + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], "issuelinks": []}} + commits = [ + {"oid": "aaa", "messageHeadline": "feat: x", "messageBody": "Implements TC-5812"}, + {"oid": "bbb", "messageHeadline": "fix: y", "messageBody": "TC-6033 unrelated"}, + ] + github = pre_verify_pr.build_github_bundle( + "o/r", 5, "b", "aaa", "d", "s", [], [], [], commits, + ) + result = pre_verify_pr.transform_to_input(issue, "TC-5812", "", github) + annotated = result["github"]["commits"] + assert annotated[0]["references_task_id"] is True + assert annotated[1]["references_task_id"] is False + + +def test_transform_issue_links(): + issue = {"fields": { + "summary": "S", "description": {}, "status": {"name": "Open"}, + "labels": [], "issuelinks": [ + {"type": {"name": "Blocks"}, "outwardIssue": {"key": "TC-200"}}, + {"type": {"name": "Related"}, "inwardIssue": {"key": "TC-300"}}, + ], + }} + links = pre_verify_pr.transform_to_input(issue, "TC-100", "")["task"]["issue_links"] + assert len(links) == 2 + assert links[0] == {"type": "Blocks", "direction": "outward", "key": "TC-200"} + assert links[1] == {"type": "Related", "direction": "inward", "key": "TC-300"} + + +def test_transform_custom_fields(): + issue = {"fields": { + "summary": "S", "description": {}, "status": None, "labels": [], + "issuelinks": [], + "customfield_10875": "https://github.com/o/r/pull/5", + "customfield_99999": {"value": "something"}, + "priority": {"name": "High"}, + }} + cf = pre_verify_pr.transform_to_input(issue, "TC-1", "")["task"]["custom_fields"] + assert "customfield_10875" in cf + assert "customfield_99999" in cf + assert "priority" not in cf + + +def test_transform_empty_fields(): + issue = {"fields": {}} + result = pre_verify_pr.transform_to_input(issue, "TC-1", "") + assert result["task"]["summary"] == "" + assert result["task"]["status"] == "" + assert result["task"]["labels"] == [] + assert result["task"]["issue_links"] == [] + + +def test_transform_null_status(): + issue = {"fields": {"summary": "S", "status": None, "labels": [], "issuelinks": []}} + result = pre_verify_pr.transform_to_input(issue, "TC-1", "") + assert result["task"]["status"] == "" + + +def test_transform_null_description_coerced_to_object(): + """An explicit null description becomes {} so task.description stays an object. + + Regression for Sourcery id 3896899434: `fields.get("description", {})` only + defaults on an absent key, so an explicit JSON null (an issue with no + description) yielded task.description = null, violating the input schema. + """ + # Given a Jira issue whose description field is an explicit null + issue = {"fields": {"summary": "S", "description": None, "status": {"name": "Open"}, + "labels": [], "issuelinks": []}} + + # When transforming it to the tracker-agnostic input + result = pre_verify_pr.transform_to_input(issue, "TC-1", "") + + # Then the null is coerced to an empty object, not left as null + assert result["task"]["description"] == {}, \ + f"expected {{}}, got: {result['task']['description']!r}" + assert isinstance(result["task"]["description"], dict) + + +def test_null_description_input_validates_against_schema(): + """A produced input with a null-source description validates against the schema. + + Drives the full transform (task + github bundle) for an issue with a null + description and asserts the result satisfies verify-pr-input.schema.json — + the acceptance criterion for TC-5886. Without the null coercion the instance + would carry task.description = null and fail (description must be an object). + """ + from jsonschema import validate + + # Given an issue with a null description and the prefetched github bundle a + # real run embeds (the schema requires `github`, so a bare task won't do) + issue = {"fields": {"summary": "S", "description": None, "status": {"name": "Open"}, + "labels": [], "issuelinks": []}} + github = pre_verify_pr.build_github_bundle( + "o/r", 5, "feat/x", "deadbee", "diff", "stat", [], [], [], [], + ) + + # When producing the input and loading the input schema + result = pre_verify_pr.transform_to_input( + issue, "TC-1", "https://github.com/o/r/pull/5", github) + schema_path = os.path.join( + script_dir, "..", "schemas", "verify-pr-input.schema.json") + with open(schema_path) as f: + schema = json.load(f) + + # Then it validates cleanly (validate raises ValidationError on failure) + assert result["task"]["description"] == {} + validate(instance=result, schema=schema) + + +def test_transform_large_payload(): + """Regression test: large payloads must work via stdin, not argv.""" + issue = {"fields": { + "summary": "Large issue", + "description": "x" * 500_000, + "status": {"name": "Open"}, + "labels": [], + "issuelinks": [], + }} + payload = json.dumps(issue) + assert len(payload) > 500_000 + + result = subprocess.run( + [sys.executable, os.path.join(script_dir, "pre_verify_pr.py"), + "transform", "TC-BIG", "https://example.com/pr/1"], + input=payload, capture_output=True, text=True, + ) + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + output = json.loads(result.stdout) + assert output["task_id"] == "TC-BIG" + assert output["task"]["summary"] == "Large issue" + assert "github" not in output + + +def test_cli_transform_github_dir(): + """CLI transform reads the raw GitHub files and embeds the bundle.""" + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], "issuelinks": []}} + with tempfile.TemporaryDirectory() as d: + with open(os.path.join(d, "pr.diff"), "w") as f: + f.write("diff --git a b\n") + with open(os.path.join(d, "pr.stat"), "w") as f: + f.write(" 1 file changed\n") + for name, payload in [ + ("reviews.json", [{"id": 1, "state": "APPROVED"}]), + ("review-comments.json", [{"id": 2}]), + ("issue-comments.json", [{"id": 3}]), + ("commits.json", [{"oid": "abc1234def"}]), + ("check-runs.json", [{"name": "pytest", "status": "completed", + "conclusion": "success", + "details_url": "https://ci/1"}]), + ]: + with open(os.path.join(d, name), "w") as f: + json.dump(payload, f) + + result = subprocess.run( + [sys.executable, os.path.join(script_dir, "pre_verify_pr.py"), + "transform", "TC-9", "https://github.com/o/r/pull/9", + "--github-dir", d, "--pr-repo", "o/r", "--pr-number", "9", + "--head-ref", "feat/x", "--commit-sha", "abc1234def"], + input=json.dumps(issue), capture_output=True, text=True, + ) + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + output = json.loads(result.stdout) + gh = output["github"] + assert gh["pr_repo"] == "o/r" + assert gh["pr_number"] == 9 + assert gh["headRefName"] == "feat/x" + assert gh["commit_sha"] == "abc1234def" + assert gh["diff"] == "diff --git a b\n" + assert gh["stat"] == " 1 file changed\n" + assert gh["reviews"] == [{"id": 1, "state": "APPROVED"}] + assert gh["review_comments"] == [{"id": 2}] + assert gh["issue_comments"] == [{"id": 3}] + # transform annotates each commit with the deterministic traceability fact. + assert gh["commits"] == [{"oid": "abc1234def", "references_task_id": False}] + assert gh["check_runs"] == [{"name": "pytest", "status": "completed", + "conclusion": "success", + "details_url": "https://ci/1"}] + # No check-run-logs.txt written (green run) → empty path, nothing to read. + assert gh["check_run_logs_path"] == "" + + +def test_cli_transform_own_check_names_file_filters_github_check_runs(): + """CLI transform with --own-check-names-file drops own-harness check-runs (TC-6343). + + End-to-end through the real transform: the prefetched check-runs.json holds a + substantive success plus verify-pr's own harness run; the own-names file lists + the own run's name; github.check_runs must contain only the substantive check. + """ + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], + "issuelinks": []}} + with tempfile.TemporaryDirectory() as d: + with open(os.path.join(d, "pr.diff"), "w") as f: + f.write("d\n") + with open(os.path.join(d, "pr.stat"), "w") as f: + f.write("s\n") + for name in ["reviews.json", "review-comments.json", + "issue-comments.json", "commits.json"]: + with open(os.path.join(d, name), "w") as f: + json.dump([], f) + with open(os.path.join(d, "check-runs.json"), "w") as f: + json.dump([ + {"name": "pytest", "status": "completed", "conclusion": "success", + "details_url": "https://ci/1"}, + {"name": "verify-pr / harness-run (verify-pr, review)", + "status": "in_progress", "conclusion": None, + "details_url": "https://ci/2"}, + ], f) + own_names_file = os.path.join(d, "own-check-names.txt") + with open(own_names_file, "w") as f: + f.write("verify-pr / harness-run (verify-pr, review)\n") + + result = subprocess.run( + [sys.executable, os.path.join(script_dir, "pre_verify_pr.py"), + "transform", "TC-9", "https://github.com/o/r/pull/9", + "--github-dir", d, "--pr-repo", "o/r", "--pr-number", "9", + "--head-ref", "feat/x", "--commit-sha", "abc1234def", + "--own-check-names-file", own_names_file], + input=json.dumps(issue), capture_output=True, text=True, + ) + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + check_runs = json.loads(result.stdout)["github"]["check_runs"] + assert check_runs == [{"name": "pytest", "status": "completed", + "conclusion": "success", "details_url": "https://ci/1"}] + + +def test_cli_transform_without_own_check_names_file_keeps_all_check_runs(): + """Omitting --own-check-names-file is a no-op: all check-runs pass through. + + Guards the default path (older invocation / enumeration unavailable) so the + self-exclusion never silently drops a check when no names were gathered. + """ + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], + "issuelinks": []}} + with tempfile.TemporaryDirectory() as d: + with open(os.path.join(d, "pr.diff"), "w") as f: + f.write("d\n") + with open(os.path.join(d, "pr.stat"), "w") as f: + f.write("s\n") + for name in ["reviews.json", "review-comments.json", + "issue-comments.json", "commits.json"]: + with open(os.path.join(d, name), "w") as f: + json.dump([], f) + with open(os.path.join(d, "check-runs.json"), "w") as f: + json.dump([ + {"name": "pytest", "status": "completed", "conclusion": "success", + "details_url": "https://ci/1"}, + {"name": "wait-for-checks", "status": "completed", + "conclusion": "success", "details_url": "https://ci/2"}, + ], f) + + result = subprocess.run( + [sys.executable, os.path.join(script_dir, "pre_verify_pr.py"), + "transform", "TC-9", "https://github.com/o/r/pull/9", + "--github-dir", d, "--pr-repo", "o/r", "--pr-number", "9", + "--head-ref", "feat/x", "--commit-sha", "abc1234def"], + input=json.dumps(issue), capture_output=True, text=True, + ) + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + names = [c["name"] for c in json.loads(result.stdout)["github"]["check_runs"]] + assert names == ["pytest", "wait-for-checks"] + + +def test_cli_transform_embeds_check_run_logs_path_when_present(): + """A non-empty check-run-logs.txt makes transform embed the mounted path. + + The file content is never inlined — only the sandbox path — so Check 1b reads + the failure logs on demand on a FAIL. + """ + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], "issuelinks": []}} + with tempfile.TemporaryDirectory() as d: + with open(os.path.join(d, "pr.diff"), "w") as f: + f.write("d\n") + with open(os.path.join(d, "pr.stat"), "w") as f: + f.write("s\n") + for name in ["reviews.json", "review-comments.json", + "issue-comments.json", "commits.json", "check-runs.json"]: + with open(os.path.join(d, name), "w") as f: + json.dump([], f) + # A failed check left log text on the runner. + with open(os.path.join(d, "check-run-logs.txt"), "w") as f: + f.write("===== CI run 42 — failed steps =====\nE assert False\n") + + result = subprocess.run( + [sys.executable, os.path.join(script_dir, "pre_verify_pr.py"), + "transform", "TC-9", "https://github.com/o/r/pull/9", + "--github-dir", d, "--pr-repo", "o/r", "--pr-number", "9", + "--head-ref", "feat/x", "--commit-sha", "abc1234def"], + input=json.dumps(issue), capture_output=True, text=True, + ) + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + gh = json.loads(result.stdout)["github"] + assert gh["check_run_logs_path"] == pre_verify_pr.SANDBOX_CHECK_RUN_LOGS_PATH + + +def test_cli_transform_empty_check_run_logs_yields_empty_path(): + """An empty check-run-logs.txt (no failures) → empty path, no read.""" + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], "issuelinks": []}} + with tempfile.TemporaryDirectory() as d: + with open(os.path.join(d, "pr.diff"), "w") as f: + f.write("d\n") + with open(os.path.join(d, "pr.stat"), "w") as f: + f.write("s\n") + for name in ["reviews.json", "review-comments.json", + "issue-comments.json", "commits.json", "check-runs.json"]: + with open(os.path.join(d, name), "w") as f: + json.dump([], f) + open(os.path.join(d, "check-run-logs.txt"), "w").close() # empty + + result = subprocess.run( + [sys.executable, os.path.join(script_dir, "pre_verify_pr.py"), + "transform", "TC-9", "https://github.com/o/r/pull/9", + "--github-dir", d, "--pr-repo", "o/r", "--pr-number", "9", + "--head-ref", "feat/x", "--commit-sha", "abc1234def"], + input=json.dumps(issue), capture_output=True, text=True, + ) + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + gh = json.loads(result.stdout)["github"] + assert gh["check_run_logs_path"] == "" + + +# --- idempotency prefetch (related_keys, build_idempotency_bundle, transform) --- + +def test_related_keys_from_subtasks_and_links(): + """related_keys collects sub-task keys and linked-issue keys, deduped/sorted.""" + issue = {"fields": { + "subtasks": [{"key": "TC-201"}, {"key": "TC-202"}], + "issuelinks": [ + {"type": {"name": "Blocks"}, "inwardIssue": {"key": "TC-202"}}, + {"type": {"name": "Related"}, "outwardIssue": {"key": "TC-300"}}, + ], + }} + # TC-202 appears as both a sub-task and a link — deduped; result is sorted + assert pre_verify_pr.related_keys(issue) == ["TC-201", "TC-202", "TC-300"] + + +def test_related_keys_empty_when_no_relations(): + issue = {"fields": {"summary": "S"}} + assert pre_verify_pr.related_keys(issue) == [] + + +def test_build_idempotency_bundle_extracts_fields_and_comments(): + """The bundle carries the fields the dedup checks inspect, incl. comment bodies.""" + ri = { + "key": "TC-500", + "fields": { + "summary": "Fix eval-3 assertion failures", + "labels": ["ai-generated-jira", "eval-failure"], + "description": {"type": "doc", "content": []}, + "issuetype": {"name": "Sub-task"}, + "comment": {"comments": [ + {"body": {"type": "doc", "content": [{"type": "text"}]}}, + {"body": {"type": "doc"}}, + ]}, + }, + } + bundle = pre_verify_pr.build_idempotency_bundle([ri]) + entry = bundle["related_issues"][0] + assert entry["key"] == "TC-500" + assert entry["summary"] == "Fix eval-3 assertion failures" + assert entry["labels"] == ["ai-generated-jira", "eval-failure"] + assert entry["description"] == {"type": "doc", "content": []} + assert entry["issuetype"] == "Sub-task" + assert len(entry["comments"]) == 2 + assert entry["comments"][0] == {"type": "doc", "content": [{"type": "text"}]} + + +def test_build_idempotency_bundle_handles_missing_fields(): + """A related issue lacking comments/description yields empty defaults, not errors.""" + bundle = pre_verify_pr.build_idempotency_bundle([{"key": "TC-9", "fields": {}}]) + entry = bundle["related_issues"][0] + assert entry["summary"] == "" + assert entry["labels"] == [] + assert entry["description"] == {} # null/absent coerced to object + assert entry["issuetype"] == "" + assert entry["comments"] == [] + + +def test_build_idempotency_bundle_empty(): + assert pre_verify_pr.build_idempotency_bundle([]) == {"related_issues": []} + + +def test_transform_with_idempotency_attaches_key(): + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], "issuelinks": []}} + idem = pre_verify_pr.build_idempotency_bundle([{"key": "TC-1", "fields": {}}]) + result = pre_verify_pr.transform_to_input(issue, "TC-1", "", None, idem) + assert result["idempotency"]["related_issues"][0]["key"] == "TC-1" + + +def test_transform_without_idempotency_defaults_to_empty(): + """With no idempotency supplied, transform still emits an empty bundle. + + Option 1 of TC-6026: the key is always present so the tokenless sandbox + always has a data source for the Steps 6d/6f/7c dedup checks — a missing + argument degrades to "no known duplicates", never an omitted key. + """ + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], "issuelinks": []}} + result = pre_verify_pr.transform_to_input(issue, "TC-1", "") + assert result["idempotency"] == {"related_issues": []} + + +def test_cli_transform_idempotency_dir(): + """CLI transform globs the related-issue JSONs and embeds the idempotency bundle.""" + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], "issuelinks": []}} + with tempfile.TemporaryDirectory() as d: + with open(os.path.join(d, "TC-501.json"), "w") as f: + json.dump({"key": "TC-501", "fields": { + "summary": "Existing sub-task", "labels": ["review-feedback"], + "description": {"type": "doc"}, "issuetype": {"name": "Sub-task"}, + "comment": {"comments": [{"body": {"type": "doc"}}]}, + }}, f) + result = subprocess.run( + [sys.executable, os.path.join(script_dir, "pre_verify_pr.py"), + "transform", "TC-1", "https://github.com/o/r/pull/1", + "--idempotency-dir", d], + input=json.dumps(issue), capture_output=True, text=True, + ) + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + output = json.loads(result.stdout) + related = output["idempotency"]["related_issues"] + assert len(related) == 1 + assert related[0]["key"] == "TC-501" + assert related[0]["labels"] == ["review-feedback"] + assert related[0]["comments"] == [{"type": "doc"}] + + +def test_cli_transform_empty_idempotency_dir(): + """An empty related-issues dir yields an empty related_issues list, not an error.""" + issue = {"fields": {"summary": "S", "status": {"name": "Open"}, "labels": [], "issuelinks": []}} + with tempfile.TemporaryDirectory() as d: + result = subprocess.run( + [sys.executable, os.path.join(script_dir, "pre_verify_pr.py"), + "transform", "TC-1", "https://github.com/o/r/pull/1", + "--idempotency-dir", d], + input=json.dumps(issue), capture_output=True, text=True, + ) + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + assert json.loads(result.stdout)["idempotency"] == {"related_issues": []} + + +def test_cli_related_keys(): + """The related-keys subcommand prints sub-task and linked-issue keys.""" + issue = {"fields": { + "subtasks": [{"key": "TC-201"}], + "issuelinks": [{"type": {"name": "Related"}, "outwardIssue": {"key": "TC-300"}}], + }} + result = subprocess.run( + [sys.executable, os.path.join(script_dir, "pre_verify_pr.py"), "related-keys"], + input=json.dumps(issue), capture_output=True, text=True, + ) + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + assert result.stdout.split() == ["TC-201", "TC-300"] + + +def test_idempotency_input_validates_against_schema(): + """A produced input carrying the idempotency bundle validates against the schema.""" + from jsonschema import validate + + issue = {"fields": {"summary": "S", "description": {}, "status": {"name": "Open"}, + "labels": [], "issuelinks": []}} + github = pre_verify_pr.build_github_bundle( + "o/r", 5, "feat/x", "deadbee", "diff", "stat", [], [], [], []) + idem = pre_verify_pr.build_idempotency_bundle([{"key": "TC-9", "fields": { + "summary": "Existing", "labels": ["review-feedback"], + "description": {"type": "doc"}, "issuetype": {"name": "Sub-task"}, + "comment": {"comments": [{"body": {"type": "doc"}}]}, + }}]) + result = pre_verify_pr.transform_to_input( + issue, "TC-1", "https://github.com/o/r/pull/5", github, idem) + schema_path = os.path.join( + script_dir, "..", "schemas", "verify-pr-input.schema.json") + with open(schema_path) as f: + schema = json.load(f) + validate(instance=result, schema=schema) # raises on failure + + +def test_prefetch_omitting_idempotency_fails_schema_validation(): + """A prefetch lacking the idempotency bundle is rejected by the schema. + + Option 1 of TC-6026: idempotency is a top-level required key, so Step 0.7 + validation rejects a bundle that omits it *before* any consuming step runs — + the fail-fast contract from TC-5980/TC-5981. This closes the gap that let a + schema-valid prefetch skip the tokenless dedup checks silently. + """ + from jsonschema import validate + from jsonschema.exceptions import ValidationError + + # Given an otherwise-valid input (task + github) with no idempotency key + issue = {"fields": {"summary": "S", "description": {}, "status": {"name": "Open"}, + "labels": [], "issuelinks": []}} + github = pre_verify_pr.build_github_bundle( + "o/r", 5, "feat/x", "deadbee", "diff", "stat", [], [], [], []) + instance = pre_verify_pr.transform_to_input( + issue, "TC-1", "https://github.com/o/r/pull/5", github) + del instance["idempotency"] # simulate an older/hand-supplied bundle + + schema_path = os.path.join( + script_dir, "..", "schemas", "verify-pr-input.schema.json") + with open(schema_path) as f: + schema = json.load(f) + + # When validating it against the input schema, Then it is rejected + try: + validate(instance=instance, schema=schema) + assert False, "schema accepted a prefetch missing idempotency" + except ValidationError as e: + assert "idempotency" in str(e), f"unexpected error: {e}" + + +def test_incomplete_related_issue_fails_schema_validation(): + """A related-issue entry missing a per-item field is rejected by the schema. + + TC-6032: idempotency.related_issues.items requires the fields the dedup + checks consume (Steps 6d/6f/7c). Without a per-item `required` list, a + structurally incomplete entry passed Step 0.7 and dedup silently treated it + as non-matching — the item-level analogue of the TC-6026 array-level gap. + """ + from jsonschema import validate + from jsonschema.exceptions import ValidationError + + # Given an otherwise-valid input whose sole related issue omits `comments` + issue = {"fields": {"summary": "S", "description": {}, "status": {"name": "Open"}, + "labels": [], "issuelinks": []}} + github = pre_verify_pr.build_github_bundle( + "o/r", 5, "feat/x", "deadbee", "diff", "stat", [], [], [], []) + instance = pre_verify_pr.transform_to_input( + issue, "TC-1", "https://github.com/o/r/pull/5", github) + instance["idempotency"] = {"related_issues": [{ + "key": "TC-9", "summary": "Existing", "labels": [], + "description": {}, "issuetype": "Sub-task", + # `comments` intentionally omitted + }]} + + schema_path = os.path.join( + script_dir, "..", "schemas", "verify-pr-input.schema.json") + with open(schema_path) as f: + schema = json.load(f) + + # When validating it against the input schema, Then it is rejected + try: + validate(instance=instance, schema=schema) + assert False, "schema accepted a related issue missing a required field" + except ValidationError as e: + assert "comments" in str(e), f"unexpected error: {e}" + + +# --- stat production (pre-verify-pr.sh) --- + +pre_verify_sh = os.path.join(script_dir, "pre-verify-pr.sh") + + +def test_stat_produced_by_git_apply_stat(): + """The stat mechanism (git apply --stat) yields a diffstat matching the + downstream github.stat contract: a per-file line plus a summary line. + + Exercises the real command pre-verify-pr.sh runs, not a gh stub that + silently accepts the unsupported --stat flag. + """ + # Given a unified diff like the one gh pr diff writes to pr.diff + patch = ( + "diff --git a/foo.txt b/foo.txt\n" + "index 1111111..2222222 100644\n" + "--- a/foo.txt\n" + "+++ b/foo.txt\n" + "@@ -1,3 +1,3 @@\n" + " line1\n" + "-line2\n" + "+CHANGED\n" + " line3\n" + ) + with tempfile.TemporaryDirectory() as d: + with open(os.path.join(d, "pr.diff"), "w") as f: + f.write(patch) + + # When producing the stat with the exact command pre-verify-pr.sh uses. + # Run it from the (non-repo) temp dir: `git apply --stat` is CWD-sensitive + # — inside a repo subdirectory it scopes the patch to that subtree and + # reports "0 files changed", so cwd=d keeps this a pure textual diffstat + # independent of where the test runner is launched. + result = subprocess.run( + ["git", "apply", "--stat", "pr.diff"], + cwd=d, capture_output=True, text=True, + ) + + # Then it succeeds and emits a git diffstat downstream can consume + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + assert "foo.txt" in result.stdout, f"Got: {result.stdout!r}" + assert "1 file changed" in result.stdout, f"Got: {result.stdout!r}" + + +def test_pre_verify_sh_uses_supported_stat_command(): + """Regression guard: the prefetch derives the stat with git apply --stat and + never passes the unsupported --stat flag to gh pr diff (Sourcery id 3896899424). + """ + # Given the current pre-verify-pr.sh source + with open(pre_verify_sh) as f: + script = f.read() + + # Then no gh pr diff invocation uses --stat, and git apply --stat is present + for line in script.splitlines(): + if line.lstrip().startswith("#"): + continue # skip comments (which may mention the removed flag) + if "gh pr diff" in line: + assert "--stat" not in line, f"unsupported gh flag reintroduced: {line!r}" + assert "git apply --stat" in script, "expected git apply --stat stat mechanism" + + +# --- idempotency prefetch (pre-verify-pr.sh) --- + +def test_pre_verify_sh_prefetches_related_issues(): + """Regression guard: pre-verify-pr.sh fetches related-issue metadata for the + sandbox idempotency checks (TC-5982) — it lists related keys, fetches each + issue including its comments, and passes --idempotency-dir to transform. + """ + with open(pre_verify_sh) as f: + script = f.read() + + # It enumerates the task's related keys via pre_verify_pr.py related-keys + assert "related-keys" in script, "expected related-keys enumeration" + # It fetches each related issue including the comment field (for Step 7c) + non_comment = [ + line for line in script.splitlines() + if 'get_issue "${key}"' in line and not line.lstrip().startswith("#") + ] + assert non_comment, "expected a per-key get_issue fetch" + # The fields list feeding that fetch must include comment, description, labels + assert "summary,labels,description,issuetype,comment" in script, \ + "related-issue fetch must request comment/description/labels" + # And it wires the prefetched dir into transform + assert "--idempotency-dir" in script, "transform must receive --idempotency-dir" + + +# --- CI Status self-exclusion (pre-verify-pr.sh) --- + +def test_pre_verify_sh_gathers_own_check_names_and_wires_filter(): + """Regression guard (TC-6343): pre-verify-pr.sh enumerates verify-pr's own + workflow check-run names for the head SHA and routes them through the transform + filter, so CI Status self-excludes the harness's own runs. + """ + # Given the current pre-verify-pr.sh source + with open(pre_verify_sh) as f: + script = f.read() + + non_comment = [ + line for line in script.splitlines() if not line.lstrip().startswith("#") + ] + body = "\n".join(non_comment) + + # It enumerates this workflow's own runs at the head SHA by workflow name, + # covering superseded attempts (all runs at the SHA, not one run ID). + assert "fullsend-verify-pr.yml" in body, \ + "must enumerate own runs by the verify-pr workflow file" + assert "head_sha=${COMMIT_SHA}" in body, \ + "must scope the own-run enumeration to the head SHA" + assert "actions/runs/" in body and "/jobs" in body, \ + "must resolve own check-run names from each run's jobs" + # The gathered names are written to a file and passed into transform, where the + # pure Python filter drops them before they reach github.check_runs. + assert "--own-check-names-file" in body, \ + "transform must receive the own-check-names file" + + +# --- PR-URL-derived Jira gating (pre-verify-pr.sh) --- + +def test_pre_verify_sh_derives_and_gates_from_pr_url(): + """Regression guard (TC-6190): pre-verify-pr.sh takes the PR URL as its entry + point, derives the Jira key by JQL on the Git Pull Request field, and maps a + gate failure to the ADR-0072 skip signal — it never requires JIRA_ISSUE_ID. + """ + # Given the current pre-verify-pr.sh source + with open(pre_verify_sh) as f: + script = f.read() + + non_comment = [ + line for line in script.splitlines() if not line.lstrip().startswith("#") + ] + body = "\n".join(non_comment) + + # The triggering PR URL is the required input (from fullsend harness-run) + assert "FULLSEND_WORK_ITEM_URL" in body, "PR URL entry point must be required" + # It never validates or requires a supplied JIRA_ISSUE_ID as an input + assert "JIRA_ISSUE_ID:?" not in body, "JIRA_ISSUE_ID must not be a required input" + # The key is derived by JQL on the custom field then gated in Python + assert "build-pr-jql" in body, "must build the PR JQL" + assert "search_jql" in body, "must search Jira by JQL" + assert "resolve-gated-issue" in body, "must resolve + gate the issue" + # A gate failure (exit 3) is mapped to the ADR-0072 skip signal + assert "-eq 3" in body and "request_skip" in body, \ + "gate failure (exit 3) must trigger request_skip" + + +# --- paginated fetch aggregation (pre-verify-pr.sh) --- + +def test_paginated_pages_aggregate_into_flat_array(): + """Multi-page fetches merge into one complete array, not truncated at page 1. + + Exercises the exact merge pre-verify-pr.sh runs on the `gh api --paginate + --slurp` output — a per-page array-of-arrays piped through `jq 'add'` — and + asserts every page's items survive in order with their object shape intact. + """ + # Given the array-of-pages that `gh api --paginate --slurp` emits: three + # pages, so page-2 and page-3 items only appear if pagination is honored. + slurped_pages = [ + [{"id": 1}, {"id": 2}], + [{"id": 3}, {"id": 4}], + [{"id": 5}], + ] + + # When merged with the same standalone `jq 'add'` the script pipes through + result = subprocess.run( + ["jq", "add"], + input=json.dumps(slurped_pages), capture_output=True, text=True, + ) + + # Then the pages flatten into one array carrying items beyond the first page + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + merged = json.loads(result.stdout) + assert merged == [{"id": 1}, {"id": 2}, {"id": 3}, {"id": 4}, {"id": 5}], \ + f"expected flat concatenation, got: {merged!r}" + assert [o["id"] for o in merged] == [1, 2, 3, 4, 5] # order preserved + assert {"id": 5} in merged # last page (beyond first) not truncated + + +def test_pre_verify_sh_paginates_review_comment_fetches(): + """Regression guard: all three gh api review/comment fetches request every + page and merge with `jq 'add'` (Sourcery id 3896899429), keeping the stored + value a flat array for pre_verify_pr.py. + """ + # Given the current pre-verify-pr.sh source + with open(pre_verify_sh) as f: + script = f.read() + + # Then each of the three paginated endpoints is fetched with --paginate + # --slurp and merged via jq add on a non-comment line. + endpoints = [ + "/pulls/${PR_NUM}/reviews", + "/pulls/${PR_NUM}/comments", + "/issues/${PR_NUM}/comments", + ] + for endpoint in endpoints: + matches = [ + line for line in script.splitlines() + if endpoint in line and not line.lstrip().startswith("#") + ] + assert matches, f"no fetch line found for {endpoint}" + for line in matches: + if "gh api" not in line: + continue + assert "--paginate" in line, f"missing --paginate: {line!r}" + assert "--slurp" in line, f"missing --slurp: {line!r}" + assert "jq 'add'" in line, f"missing jq 'add' merge: {line!r}" + + +# --- PR-URL parsing regex (pre-verify-pr.sh) --- + +def _github_pr_url_regex(): + """Extract the github.com PR-URL match regex from pre-verify-pr.sh. + + The tests below run the *actual* regex the script ships (not a re-typed + copy), so an accidental loss of the end anchor is caught behaviorally. + """ + with open(pre_verify_sh) as f: + for line in f: + if "=~" in line and "github" in line and "/pull/" in line: + m = re.search(r"=~\s+(\S.*?)\s+\]\]", line) + if m: + return m.group(1) + raise AssertionError("github.com PR-URL regex not found in pre-verify-pr.sh") + + +def _match_pr_url(url): + """Run pre-verify-pr.sh's exact `[[ =~ ]]` test against url. + + Returns (matched, repo, number) using the same BASH_REMATCH groups the + script consumes downstream, so a truncating match surfaces as a wrong + `number` rather than a silent pass. + """ + regex = _github_pr_url_regex() + snippet = ( + 'r="$1"; u="$2"\n' + 'if [[ "$u" =~ $r ]]; then\n' + ' printf "MATCH\\t%s\\t%s" "${BASH_REMATCH[1]}" "${BASH_REMATCH[2]}"\n' + 'else\n' + ' printf "NOMATCH"\n' + 'fi\n' + ) + result = subprocess.run( + ["bash", "-c", snippet, "bash", regex, url], + capture_output=True, text=True, + ) + assert result.returncode == 0, f"bash error: {result.stderr}" + parts = result.stdout.split("\t") + if parts[0] == "MATCH": + return True, parts[1], parts[2] + return False, None, None + + +def test_pr_url_wellformed_accepted(): + """A canonical github.com PR URL parses to its owner/repo and PR number.""" + # Given a well-formed PR URL like a Jira inlineCard stores + # When matched by the script's regex + matched, repo, number = _match_pr_url("https://github.com/org/repo/pull/42") + + # Then it matches and captures the exact repo and number + assert matched, "well-formed URL should match" + assert repo == "org/repo", f"Got repo: {repo!r}" + assert number == "42", f"Got number: {number!r}" + + +def test_pr_url_trailing_slash_accepted(): + """A well-formed PR URL with a trailing slash still parses correctly.""" + # Given a PR URL with a trailing slash + # When matched by the script's regex + matched, repo, number = _match_pr_url("https://github.com/org/repo/pull/42/") + + # Then it matches with the same repo and number (the slash is tolerated) + assert matched, "trailing-slash URL should match" + assert repo == "org/repo", f"Got repo: {repo!r}" + assert number == "42", f"Got number: {number!r}" + + +def test_pr_url_nonnumeric_suffix_rejected(): + """A pull number with a trailing non-numeric suffix is rejected, not truncated. + + Regression for Sourcery id 3902059934: the un-anchored regex accepted + `.../pull/42abc` and truncated the number to 42, so the pre-script fetched + the wrong PR. The end-anchored regex must reject it outright. + """ + # Given a malformed URL with a non-numeric suffix on the pull number + # When matched by the script's regex + matched, _repo, number = _match_pr_url("https://github.com/org/repo/pull/42abc") + + # Then it does not match (rather than truncating to 42) + assert not matched, f"expected rejection, but matched with number={number!r}" + + +def test_pr_url_extra_path_segment_rejected(): + """An extra path segment after the pull number is rejected, not truncated. + + Regression for Sourcery id 3902059934: `.../pull/42/invalid` previously + matched and truncated to PR 42. The end anchor must reject it. + """ + # Given a malformed URL with an extra path segment after the number + # When matched by the script's regex + matched, _repo, number = _match_pr_url( + "https://github.com/org/repo/pull/42/invalid") + + # Then it does not match (rather than truncating to 42) + assert not matched, f"expected rejection, but matched with number={number!r}" + + +# --- COMMIT_SHA derivation (pre-verify-pr.sh) --- + +def _commit_sha_command(): + """Extract the COMMIT_SHA assignment command from pre-verify-pr.sh. + + The behavioral test below runs the *actual* command the script ships (not a + re-typed copy), so a regression back to the bounded commits connection is + caught by execution, not just by source inspection. + """ + with open(pre_verify_sh) as f: + for line in f: + if line.lstrip().startswith("COMMIT_SHA="): + return line.strip() + raise AssertionError("COMMIT_SHA assignment not found in pre-verify-pr.sh") + + +def test_commit_sha_derived_from_head_ref_oid(): + """The prefetched COMMIT_SHA is the PR head ref tip OID, not the last commit + of gh's bounded commits connection. + + Runs the exact COMMIT_SHA command pre-verify-pr.sh ships against a gh stub + whose headRefOid and commits[-1].oid disagree (simulating a large PR whose + commits connection is truncated below the head). Regression for Sourcery id + 3902563589: the old `.commits[-1].oid` read returns the truncated oid. + """ + # Given a gh stub where the head ref OID and the (truncated) commits + # connection's last oid disagree + head_oid = "a" * 40 + truncated_oid = "b" * 40 + with tempfile.TemporaryDirectory() as d: + gh_stub = os.path.join(d, "gh") + with open(gh_stub, "w") as f: + f.write( + "#!/usr/bin/env bash\n" + "for arg in \"$@\"; do\n" + f' if [[ "$arg" == headRefOid ]]; then echo {head_oid}; exit 0; fi\n' + f' if [[ "$arg" == commits ]]; then echo {truncated_oid}; exit 0; fi\n' + "done\n" + "echo UNEXPECTED >&2; exit 1\n" + ) + os.chmod(gh_stub, 0o755) + + # When running the exact COMMIT_SHA assignment the script ships, with the + # stub gh ahead on PATH and PR_NUM/PR_REPO supplied + command = _commit_sha_command() + env = {**os.environ, "PATH": d + os.pathsep + os.environ["PATH"], + "PR_NUM": "275", "PR_REPO": "o/r"} + result = subprocess.run( + ["bash", "-c", command + "\nprintf '%s' \"$COMMIT_SHA\""], + capture_output=True, text=True, env=env, + ) + + # Then the emitted commit SHA is the head ref OID, not the truncated last commit + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + assert result.stdout == head_oid, f"Got: {result.stdout!r} (stderr: {result.stderr!r})" + assert result.stdout != truncated_oid + + +def test_pre_verify_sh_derives_commit_sha_from_head_ref_oid(): + """Regression guard: COMMIT_SHA reads headRefOid and never the truncatable + commits connection (.commits[-1].oid) (Sourcery id 3902563589). + """ + # Given the current pre-verify-pr.sh source + with open(pre_verify_sh) as f: + script = f.read() + + # Then every COMMIT_SHA assignment reads headRefOid and none falls back to + # gh's bounded commits connection + commit_sha_lines = [ + line for line in script.splitlines() + if line.lstrip().startswith("COMMIT_SHA=") + ] + assert commit_sha_lines, "no COMMIT_SHA assignment found in pre-verify-pr.sh" + for line in commit_sha_lines: + assert "headRefOid" in line, f"COMMIT_SHA not derived from headRefOid: {line!r}" + assert ".commits[-1]" not in line, \ + f"COMMIT_SHA reintroduced the bounded commits read: {line!r}" + + +# --- build_pr_jql (PR-URL → JQL) --- + +PR_URL = "https://github.com/org/repo/pull/42" + + +def test_build_pr_jql_targets_custom_field_and_url(): + """The JQL queries the Git Pull Request custom field by id for the PR URL.""" + # Given a PR URL, when the JQL is built + jql = pre_verify_pr.build_pr_jql(PR_URL) + + # Then it matches the custom field (by id) against the URL with ~ recall + assert jql == 'cf[10875] ~ "https://github.com/org/repo/pull/42"', f"Got: {jql}" + + +def test_build_pr_jql_escapes_quotes(): + """A URL containing a double quote cannot break out of the JQL string literal.""" + # Given a hostile URL embedding a quote + jql = pre_verify_pr.build_pr_jql('https://x/"; DROP') + + # Then the quote is backslash-escaped inside the quoted literal + assert jql == 'cf[10875] ~ "https://x/\\"; DROP"', f"Got: {jql}" + + +# --- resolve_gated_issue (JQL result → gated key / skip reason) --- + +def _search_issue(key, pr_url, status="Review", labels=("ai-generated-jira",)): + """A single Jira search-result issue with the given gate-relevant fields.""" + return { + "key": key, + "fields": { + "status": {"name": status}, + "labels": list(labels), + "customfield_10875": pr_url, + }, + } + + +def test_resolve_gated_issue_all_gates_pass(): + """One issue matching the PR URL, in Review, with the label → resolves the key.""" + # Given a search result with exactly one qualifying issue + result = {"issues": [_search_issue("TC-6190", PR_URL)]} + + # When resolving against the PR URL + key, reason = pre_verify_pr.resolve_gated_issue(result, PR_URL) + + # Then the key resolves and there is no skip reason + assert key == "TC-6190", f"Got: {key}" + assert reason is None, f"Got: {reason}" + + +def test_resolve_gated_issue_no_pr_match_skips(): + """No issue's custom field exactly equals the PR URL → skip (PR-match gate).""" + # Given a search that recalled a different PR (broad ~ over-match) + result = {"issues": [_search_issue("TC-1", "https://github.com/org/repo/pull/99")]} + + # When resolving against the target PR URL + key, reason = pre_verify_pr.resolve_gated_issue(result, PR_URL) + + # Then no key resolves and the reason names the missing link + assert key is None, f"Got: {key}" + assert "no Jira issue links" in reason, f"Got: {reason}" + + +def test_resolve_gated_issue_multiple_matches_skips(): + """More than one issue links the same PR URL → skip (ambiguous, no guess).""" + # Given two issues both exactly linking the PR URL + result = {"issues": [_search_issue("TC-1", PR_URL), _search_issue("TC-2", PR_URL)]} + + # When resolving + key, reason = pre_verify_pr.resolve_gated_issue(result, PR_URL) + + # Then it refuses to guess and names both keys + assert key is None, f"Got: {key}" + assert "multiple Jira issues" in reason, f"Got: {reason}" + assert "TC-1" in reason and "TC-2" in reason, f"Got: {reason}" + + +def test_resolve_gated_issue_wrong_status_skips(): + """The matched issue is not in Review → skip (status gate).""" + # Given the sole match in the wrong status + result = {"issues": [_search_issue("TC-6190", PR_URL, status="In Progress")]} + + # When resolving + key, reason = pre_verify_pr.resolve_gated_issue(result, PR_URL) + + # Then it skips and the reason names the status gate + assert key is None, f"Got: {key}" + assert "In Progress" in reason and "Review" in reason, f"Got: {reason}" + + +def test_resolve_gated_issue_missing_label_skips(): + """The matched issue lacks the ai-generated-jira label → skip (label gate).""" + # Given the sole match in Review but without the gating label + result = {"issues": [_search_issue("TC-6190", PR_URL, labels=("other",))]} + + # When resolving + key, reason = pre_verify_pr.resolve_gated_issue(result, PR_URL) + + # Then it skips and the reason names the missing label + assert key is None, f"Got: {key}" + assert "ai-generated-jira" in reason, f"Got: {reason}" + + +def test_resolve_gated_issue_matches_adf_custom_field(): + """The PR-URL match works when the custom field is an ADF smart link, not a + plain string — extract_pr_url normalizes both before the equality check.""" + # Given an issue whose Git Pull Request field is an ADF inlineCard + issue = { + "key": "TC-6190", + "fields": { + "status": {"name": "Review"}, + "labels": ["ai-generated-jira"], + "customfield_10875": { + "type": "doc", "version": 1, + "content": [{"type": "paragraph", "content": [ + {"type": "inlineCard", "attrs": {"url": PR_URL}} + ]}], + }, + }, + } + + # When resolving against the plain PR URL + key, reason = pre_verify_pr.resolve_gated_issue({"issues": [issue]}, PR_URL) + + # Then the ADF-stored link still matches and the key resolves + assert key == "TC-6190", f"Got: {key} (reason: {reason})" + assert reason is None + + +def test_resolve_gated_issue_finds_exact_match_beyond_first_page(): + """TC-6233: a >50-result broad ~ recall where the exact PR match is not on the + first page still resolves — given the full page-aggregated result that + jira-client's `search_jql --all` produces (no false ADR-0072 skip). + """ + # Given 60 recalled issues (page 1 = 50, page 2 = 10, as search_jql_all would + # aggregate) where only the one at index 55 (beyond the first page) exactly + # links the target PR; the rest are broad ~ over-matches on other PRs. + # A distinct repo path so no generated recall URL can collide with PR_URL. + other = "https://github.com/org/other-repo/pull/{}" + issues = [_search_issue(f"TC-{i}", other.format(i)) for i in range(60)] + issues[55] = _search_issue("TC-6190", PR_URL) + result = {"issues": issues, "isLast": True} + + # When resolving against the target PR URL + key, reason = pre_verify_pr.resolve_gated_issue(result, PR_URL) + + # Then the beyond-first-page issue resolves with no skip + assert key == "TC-6190", f"Got: {key} (reason: {reason})" + assert reason is None + + +def test_pre_verify_sh_paginates_jql_search_with_all(): + """Regression guard (TC-6233): pre-verify-pr.sh runs the gating JQL search + with --all so a match beyond the first 50 recall results is never dropped. + """ + # Given the current pre-verify-pr.sh source + with open(pre_verify_sh) as f: + script = f.read() + + # Then the search_jql invocation passes --all on a non-comment line + search_lines = [ + line for line in script.splitlines() + if "search_jql" in line and not line.lstrip().startswith("#") + ] + assert search_lines, "no search_jql invocation found" + # The flag may sit on a continuation line of the same command; assert it is + # present in the search_jql command block (the --fields line carries it). + assert any("--all" in line for line in script.splitlines() + if "--fields" in line and not line.lstrip().startswith("#")), \ + "gating JQL search must use --all to paginate" + + +# --- CLI: build-pr-jql / resolve-gated-issue exit-code contract --- + +def _run_cli(args, stdin=None): + return subprocess.run( + [sys.executable, os.path.join(script_dir, "pre_verify_pr.py"), *args], + input=stdin, capture_output=True, text=True, + ) + + +def test_cli_build_pr_jql(): + """build-pr-jql takes the URL as an argument and reads no stdin.""" + # When invoked with only the PR URL + result = _run_cli(["build-pr-jql", PR_URL]) + + # Then it prints the JQL and exits 0 + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + assert result.stdout.strip() == 'cf[10875] ~ "%s"' % PR_URL + + +def test_cli_resolve_gated_issue_success_exit_0(): + """On a passing gate the CLI prints the key and exits 0.""" + # Given a qualifying search result on stdin + payload = json.dumps({"issues": [_search_issue("TC-6190", PR_URL)]}) + + # When resolving via the CLI + result = _run_cli(["resolve-gated-issue", PR_URL], stdin=payload) + + # Then it exits 0 with the resolved key + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + assert result.stdout.strip() == "TC-6190" + + +def test_cli_resolve_gated_issue_gate_fail_exit_3(): + """On a failing gate the CLI exits 3 (the shell's ADR-0072 skip signal).""" + # Given a search result whose sole match is in the wrong status + payload = json.dumps( + {"issues": [_search_issue("TC-6190", PR_URL, status="Closed")]}) + + # When resolving via the CLI + result = _run_cli(["resolve-gated-issue", PR_URL], stdin=payload) + + # Then it exits 3 and prints the skip reason + assert result.returncode == 3, f"Exit {result.returncode}: {result.stdout}" + assert "Closed" in result.stdout and "Review" in result.stdout + + +# --- revalidate_gate (TOCTOU re-check on the full issue before write) --- + +def test_revalidate_gate_passes_when_full_issue_still_qualifies(): + """The full issue still links the PR, in Review, with the label → key, no skip.""" + # Given the full issue fetched after resolution still satisfies the gate + issue = _search_issue("TC-6190", PR_URL) + + # When re-validating immediately before writing the input + key, reason = pre_verify_pr.revalidate_gate(issue, PR_URL) + + # Then it resolves the key with no skip reason + assert key == "TC-6190", f"Got: {key}" + assert reason is None, f"Got: {reason}" + + +def test_revalidate_gate_skips_when_status_changed_after_search(): + """Issue left Review between the search gate and the full fetch → skip.""" + # Given the full issue is no longer in Review (the TOCTOU window) + issue = _search_issue("TC-6190", PR_URL, status="Closed") + + # When re-validating the full issue + key, reason = pre_verify_pr.revalidate_gate(issue, PR_URL) + + # Then no key resolves and the skip names the status change + assert key is None, f"Got: {key}" + assert "Closed" in reason and "Review" in reason, f"Got: {reason}" + + +def test_revalidate_gate_skips_when_label_removed_after_search(): + """Issue lost the ai-generated-jira label after the search → skip.""" + # Given the full issue no longer carries the gate label + issue = _search_issue("TC-6190", PR_URL, labels=()) + + # When re-validating the full issue + key, reason = pre_verify_pr.revalidate_gate(issue, PR_URL) + + # Then no key resolves and the skip names the missing label + assert key is None, f"Got: {key}" + assert "ai-generated-jira" in reason, f"Got: {reason}" + + +def test_revalidate_gate_skips_when_pr_field_changed_after_search(): + """The Git Pull Request field was re-pointed after the search → skip.""" + # Given the full issue now links a different PR + issue = _search_issue("TC-6190", "https://github.com/org/repo/pull/99") + + # When re-validating against the originally-resolved PR URL + key, reason = pre_verify_pr.revalidate_gate(issue, PR_URL) + + # Then no key resolves and the skip names the broken PR link + assert key is None, f"Got: {key}" + assert "no Jira issue links" in reason, f"Got: {reason}" + + +def test_cli_revalidate_gate_success_exit_0(): + """On a still-qualifying full issue the CLI prints the key and exits 0.""" + # Given a qualifying full issue on stdin + payload = json.dumps(_search_issue("TC-6190", PR_URL)) + + # When re-validating via the CLI + result = _run_cli(["revalidate-gate", PR_URL], stdin=payload) + + # Then it exits 0 with the resolved key + assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr}" + assert result.stdout.strip() == "TC-6190" + + +def test_cli_revalidate_gate_gate_fail_exit_3(): + """A status change detected at re-validation exits 3 (ADR-0072 skip signal).""" + # Given a full issue that has left Review since the search + payload = json.dumps(_search_issue("TC-6190", PR_URL, status="Closed")) + + # When re-validating via the CLI + result = _run_cli(["revalidate-gate", PR_URL], stdin=payload) + + # Then it exits 3 and prints the skip reason + assert result.returncode == 3, f"Exit {result.returncode}: {result.stdout}" + assert "Closed" in result.stdout and "Review" in result.stdout + + +def test_pre_verify_sh_revalidates_gate_before_writing_input(): + """The shell re-gates the full issue via revalidate-gate before the transform.""" + # Given the pre-verify-pr.sh source + with open(pre_verify_sh) as f: + script = f.read() + + # Then it invokes revalidate-gate, maps a gate failure to request_skip, and + # does so BEFORE writing verify-pr-input.json (the transform step). + assert "revalidate-gate" in script, "re-validation step missing" + reval_idx = script.index("revalidate-gate") + transform_idx = script.index("pre_verify_pr.py\" transform") + assert reval_idx < transform_idx, \ + "revalidate-gate must run before the transform that writes the input" + # The re-validation feeds the FULL issue (ISSUE_JSON), not the search result. + reval_line = next( + ln for ln in script.splitlines() if "revalidate-gate" in ln) + assert "ISSUE_JSON" in reval_line, \ + f"revalidate-gate must re-check the full issue: {reval_line!r}" + # A gate failure at re-validation emits the ADR-0072 skip, like Step 3. + assert "request_skip \"${REVAL_OUT}\"" in script, \ + "a failed re-validation must map to request_skip" + + +# --- runner --- + +if __name__ == "__main__": + tests = [v for k, v in sorted(globals().items()) if k.startswith("test_")] + failed = [] + for t in tests: + try: + t() + print(f" ✓ {t.__name__}") + except AssertionError as e: + print(f" ✗ {t.__name__}: {e}") + failed.append(t.__name__) + print(f"{'=' * 60}") + if failed: + print(f"FAILED: {len(failed)}/{len(tests)} test(s) failed") + sys.exit(1) + else: + print(f"SUCCESS: All {len(tests)} tests passed") + sys.exit(0) diff --git a/plugins/sdlc-workflow/scripts/test_triage_security_fullsend.py b/plugins/sdlc-workflow/scripts/test_triage_security_fullsend.py new file mode 100644 index 000000000..9cb8c4757 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/test_triage_security_fullsend.py @@ -0,0 +1,452 @@ +#!/usr/bin/env python3 +"""Deterministic boundary contracts; hosted evals prove actual skill execution.""" + +import importlib.util +import json +import os +from pathlib import Path +import re +import subprocess + +import pytest +from jsonschema import FormatChecker + + +ROOT = Path(__file__).resolve().parents[3] +SCRIPT_DIR = Path(__file__).parent +FIXTURE_DIR = ROOT / "evals" / "triage-security" / "files" + + +def _load_module(name, filename): + """Load a hyphenated workflow script as an importable test module.""" + spec = importlib.util.spec_from_file_location(name, SCRIPT_DIR / filename) + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +executor = _load_module("triage_security_fullsend_executor", "execute-triage-security-actions.py") +pre_triage = _load_module("triage_security_fullsend_pre_triage", "pre_triage_security.py") +requires_format_extra = pytest.mark.skipif( + not all(fmt in FormatChecker().checkers for fmt in pre_triage._REQUIRED_FORMATS), + reason="requires jsonschema[format] for uri/date-time format enforcement") + + +def _fixture(name): + """Load the deliberate JSON contract data embedded in a synthetic fixture.""" + document = (FIXTURE_DIR / name).read_text() + return json.loads(document.split("```json\n", 1)[1].split("\n```", 1)[0]) + + +def _trusted_input(name): + """Load a mounted trusted-input fixture exactly as the sandbox receives it.""" + return json.loads((FIXTURE_DIR / name).read_text()) + + +def _skill_bash(heading, index=0): + """Read a literal delivered instruction block without duplicating its logic.""" + skill = (ROOT / "plugins/sdlc-workflow/skills/triage-security/SKILL.md").read_text() + section = skill.split(heading, 1)[1].split("\n##", 1)[0] + return re.findall(r"```bash\n(.*?)\n```", section, re.DOTALL)[index].replace( + "${CLAUDE_PLUGIN_ROOT}", str(ROOT / "plugins/sdlc-workflow")) + + +def _run_instruction(source, gate): + """Execute source contracts only, never an agent or a local skill eval.""" + environment = {key: value for key, value in os.environ.items() + if key != "FULLSEND_OUTPUT_DIR"} + if gate is not None: + environment["FULLSEND_OUTPUT_DIR"] = gate + return subprocess.run(["bash", "-c", source], env=environment, + capture_output=True, text=True, check=False) + + +class _JiraRecorder: + """Record boundary writes that a trusted runner would send to Jira.""" + + def __init__(self): + self.calls = [] + + def update_issue(self, issue, fields): + """Record a field mutation.""" + self.calls.append(("field-edit", issue, fields)) + + def get_transitions(self, issue): + """Return the transition used by the synthetic action plan.""" + self.calls.append(("get-transitions", issue)) + return [{"id": "31", "to": {"name": "In Progress"}}] + + def transition_issue(self, issue, transition): + """Record a status mutation.""" + self.calls.append(("status-transition", issue, transition)) + + def create_issue(self, **kwargs): + """Record creation of the deterministic remediation task.""" + self.calls.append(("remediation-task", kwargs)) + return {"key": "TC-9001"} + + def get_issue(self, issue): + """Return the stored remediation description for digest generation.""" + self.calls.append(("get-issue", issue)) + return {"fields": {"description": { + "type": "doc", "version": 1, + "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Stored."}]}], + }}} + + def create_link(self, inward, outward, link_type): + """Record an issue-link mutation.""" + self.calls.append(("link", inward, outward, link_type)) + + def make_request(self, method, path, body): + """Record the remediation description digest comment.""" + self.calls.append(("digest", method, path, body)) + return {"id": "1"} + + def post_native(self, issue, body, marker): + """Record a sticky summary comment.""" + self.calls.append(("comment", issue, body, marker)) + + +@pytest.fixture +def recorder(monkeypatch): + """Replace the Jira boundary with a recorder for each contract test.""" + value = _JiraRecorder() + monkeypatch.setattr(executor._jira_mod, "update_issue", value.update_issue) + monkeypatch.setattr(executor._jira_mod, "get_transitions", value.get_transitions) + monkeypatch.setattr(executor._jira_mod, "transition_issue", value.transition_issue) + monkeypatch.setattr(executor._jira_mod, "create_issue", value.create_issue) + monkeypatch.setattr(executor._jira_mod, "get_issue", value.get_issue) + monkeypatch.setattr(executor._jira_mod, "create_link", value.create_link) + monkeypatch.setattr(executor._jira_mod, "make_request", value.make_request) + monkeypatch.setattr(executor._action_helpers, "post_jira_comment_native", value.post_native) + return value + + +def test_report_only_fixture_never_reaches_the_jira_boundary(recorder): + """A report-only Fullsend result must not perform Jira mutations.""" + # Given a synthetic report-only result and unauthorized trusted input + contract = _fixture("fullsend-report-only.md") + + # When the trusted executor processes the result + executor.execute_plan(contract["result"], contract["trusted_input"]) + + # Then the Jira boundary remains untouched + assert recorder.calls == [] + + +def test_invalid_bundle_fixture_is_rejected_before_jira_writes(recorder): + """Malformed Fullsend output must fail closed before any Jira operation.""" + # Given a synthetic result with contradictory evidence and an invalid action + contract = _fixture("fullsend-invalid-bundle.md") + + # When the trusted executor validates the result + with pytest.raises(executor.ActionError, match="result schema validation failed"): + executor.execute_plan(contract["result"], contract["trusted_input"]) + + # Then validation prevents all Jira boundary calls + assert recorder.calls == [] + + +@requires_format_extra +def test_trusted_input_fixtures_validate_against_the_sandbox_schema(): + """Authorized, report-only, and RPM inputs are executable trusted bundles.""" + # Given the three JSON fixtures mounted for successful Fullsend paths + fixture_names = [ + "fullsend-authorized-trusted-input.json", + "fullsend-report-only-trusted-input.json", + "fullsend-rpm-trusted-input.json", + ] + + # When the same pre-triage validator checks each mounted input + for name in fixture_names: + pre_triage.validate_bundle(_trusted_input(name)) + + # Then their authorization values preserve their intended execution branches + assert _trusted_input(fixture_names[0])["authorization"]["mutation_authorized"] is True + assert all(not _trusted_input(name)["authorization"]["mutation_authorized"] + for name in fixture_names[1:]) + + +def test_invalid_trusted_input_fixture_fails_closed_before_analysis(tmp_path): + """The literal input-validation contract emits only the exact abort result.""" + # Given the deliberately incomplete mounted JSON fixture + document = (FIXTURE_DIR / "fullsend-invalid-trusted-input.md").read_text() + malformed = document.split("```json\n", 1)[1].split("\n```", 1)[0] + + input_path = tmp_path / "trusted-input.json" + input_path.write_text(malformed) + output_path = tmp_path / "output" + output_path.mkdir() + source = _skill_bash("### Step 0.7 – Load Trusted Fullsend Input").replace( + "/sandbox/workspace/.pre-script/triage-security-input.json", str(input_path)) + + # When executing the literal instruction, not an agent execution + result = _run_instruction(source, str(output_path)) + + # Then it writes only the prescribed failure result + assert result.returncode == 1 + assert "ERROR: trusted triage-security input is invalid JSON:" in result.stdout + assert json.loads((output_path / "agent-result.json").read_text()) == {"error": "triage-security aborted: trusted input is missing, invalid JSON, or fails triage-security-input.schema.json; no interactive fallback is available in the sandbox."} + assert sorted(path.name for path in output_path.iterdir()) == ["agent-result.json"] + + +@pytest.mark.parametrize("gate, exit_code, stdout, stderr", [ + (None, 0, "interactive mode\n", ""), + ("", 1, "", "ERROR: FULLSEND_OUTPUT_DIR is set but empty\n"), + ("/synthetic-output", 0, "sandbox mode: /synthetic-output\n", ""), +]) +def test_literal_gate_instruction_distinguishes_presence(gate, exit_code, stdout, stderr): + """Catch truthiness regressions in source; this is not proof the agent routed.""" + # Given the production instruction and separately supplied environment state + source = _skill_bash("### Step 0.6 – Fullsend Mode Detection") + + # When a credential-free shell executes that exact gate instruction + result = _run_instruction(source, gate) + + # Then presence, empty export and nonempty export have distinct outcomes + assert (result.returncode, result.stdout, result.stderr) == (exit_code, stdout, stderr) + + +def test_literal_valid_input_instruction_continues_without_abort(tmp_path): + """Valid input must not write the failure object; analysis remains untested here.""" + # Given the same retained trusted bundle used by the hosted success scenario + source = _skill_bash("### Step 0.7 – Load Trusted Fullsend Input").replace( + "/sandbox/workspace/.pre-script/triage-security-input.json", + str(FIXTURE_DIR / "fullsend-report-only-trusted-input.json")) + + # When the literal input-validation instruction executes + result = _run_instruction(source, str(tmp_path)) + + # Then it continues without writing any failure or fabricated analysis result + assert result.returncode == 0 + assert result.stdout == "Trusted triage-security input available\n" + assert list(tmp_path.iterdir()) == [] + + +@pytest.mark.parametrize("result_mode", ["valid", "invalid-json", "error-only"]) +def test_literal_final_validator_contract(tmp_path, result_mode): + """Check the inline validator's exits, independently of actual agent execution.""" + # Given deliberately synthetic output, not an agent-produced analysis + documents = { + "valid": json.dumps(_fixture("fullsend-report-only.md")["result"]), + "invalid-json": '{"schema_version":', + "error-only": '{"error":"synthetic abort"}', + } + (tmp_path / "agent-result.json").write_text(documents[result_mode]) + source = _skill_bash("### Fullsend final output (successful analysis only)", 1) + + # When the real inline instruction validates these contract inputs + result = _run_instruction(source, str(tmp_path)) + + # Then malformed/error outputs fail rather than falling back interactively + assert result.returncode == (0 if result_mode == "valid" else 1) + expected = ("Final Fullsend result validated" if result_mode == "valid" else + "ERROR: final Fullsend result failed JSON/schema validation:") + assert expected in result.stdout + assert sorted(path.name for path in tmp_path.iterdir()) == ["agent-result.json"] + + +def _authorized_action_plan(): + """Build a synthetic plan covering all Jira mutation categories.""" + return { + "schema_version": "1", + "mode": "mutation-authorized", + "report": { + "issue": "TC-42", + "outcome": "affected", + "summary_markdown": "Affected.", + "evidence": [{"source": "synthetic", "detail": "deliberate contract data"}], + }, + "actions": [ + {"type": "field-edit", "marker": "triage-security:labels", "issue": "TC-42", "fields": {"labels": ["ai-cve-triaged"]}}, + {"type": "status-transition", "marker": "triage-security:status", "issue": "TC-42", "status": "In Progress"}, + {"type": "comment", "marker": "triage-security:summary", "issue": "TC-42", "body_adf": {"type": "doc", "version": 1, "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Summary."}]}]}}, + {"type": "remediation-task", "marker": "triage-security:remediation", "ref": "remediation", "project": "TC", "summary": "Fix synthetic CVE", "description_adf": {"type": "doc", "version": 1, "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Fix."}]}]}, "labels": []}, + {"type": "link", "marker": "triage-security:link", "link_type": "Depend", "inward": "TC-42", "outward": "{{remediation.key}}"}, + ], + } + + +def test_authorized_action_plan_executes_each_mutation_category_in_order(recorder): + """Authorized plans perform field, status, comment, task, digest, and link actions.""" + # Given a complete authorized action plan and explicit trusted identity + result = _authorized_action_plan() + trusted_input = { + "issue": {"key": "TC-42"}, + "configuration": {"project_key": "TC"}, + "authorization": {"mutation_authorized": True}, "idempotency": {}, + } + + # When the trusted runner executes the plan + executor.execute_plan(result, trusted_input) + + # Then every mutation category occurs and the digest precedes the dependent link + assert [call[0] for call in recorder.calls] == [ + "field-edit", "get-transitions", "status-transition", "comment", + "remediation-task", "get-issue", "digest", "link", + ] + assert recorder.calls[-1] == ("link", "TC-42", "TC-9001", "Depend") + + +def test_authorized_contract_rejects_mismatched_runner_identity(recorder): + """A schema-valid Fullsend plan cannot redirect another issue's runner grant.""" + # Given the complete TC-42 plan under a trusted TC-8100 authorization grant + result = _authorized_action_plan() + trusted_input = { + "issue": {"key": "TC-8100"}, + "configuration": {"project_key": "TC"}, + "authorization": {"mutation_authorized": True}, "idempotency": {}, + } + + # When trusted identity differs from the sandbox report and its targets + with pytest.raises(executor.ActionError): + executor.execute_plan(result, trusted_input) + + # Then no mutation category or read reaches Jira + assert recorder.calls == [] + + +def test_idempotent_retry_fixture_skips_duplicate_mutations(recorder): + """A rerun with existing state must not recreate Fullsend Jira artifacts.""" + # Given a synthetic retry with existing markers, state, and remediation task + contract = _fixture("fullsend-idempotent-retry.md") + contract["trusted_input"]["configuration"] = {"project_key": "TC"} + + # When the trusted executor replays the same action plan + registry = executor.execute_plan(contract["result"], contract["trusted_input"]) + + # Then no duplicate Jira mutation occurs and the existing task remains resolvable + assert recorder.calls == [] + assert registry["remediation"]["key"] == "TC-9001" + + +def test_retry_snapshot_suppresses_existing_field_and_link_actions(): + """Trusted Jira state suppresses replayed field and resolved-link actions.""" + # Given a retry fixture with both markers and a populated Jira snapshot + contract = _fixture("fullsend-idempotent-retry.md") + field_action, _, _, _, link_action = contract["result"]["actions"] + resolved_link = {**link_action, "outward": "TC-9001"} + + # When marker suppression is bypassed for actions represented in the snapshot + field_snapshot = {**contract["trusted_input"], "idempotency": {"action_markers": [], "existing_remediation": []}} + + # Then _already_applied independently detects both existing operations + assert executor._already_applied(field_action, field_snapshot) + assert executor._already_applied(resolved_link, field_snapshot) + + +@pytest.mark.parametrize("source_id", [1, 2, 3, 4, 5, 8, 9, 11, 12, 18]) +@requires_format_extra +def test_conditional_fullsend_inputs_are_valid_and_subject_bound(source_id): + """Every retained Fullsend unit fixture has valid input, identity and authorization.""" + # Given independent trusted-input fixtures retained for deterministic unit coverage + bundle = _trusted_input("fullsend-eval-{}-trusted-input.json".format(source_id)) + + # When the production validator checks the input directly + pre_triage.validate_bundle(bundle) + + # Then identity and authorization belong to this scenario, not a generic sample + subjects = {1: "TC-8001", 2: "TC-8002", 3: "TC-8003", 4: "TC-8004", + 5: "TC-8005", 8: "TC-8010", 9: "TC-8011", 11: "TC-8021", + 12: "TC-8030", 18: "TC-8001"} + assert bundle["issue"]["key"] == subjects[source_id] + assert bundle["authorization"]["mutation_authorized"] is (source_id not in [2, 5, 12]) + assert "SYNTHETIC TEST DATA" in bundle["issue"]["fields"]["fixture_purpose"] + + +def test_conditional_retry_input_has_an_existing_task_without_a_digest(): + """The new partial retry is distinct from the retained fully triaged interactive case.""" + # Given a trusted snapshot of an interrupted remediation creation + bundle = _trusted_input("fullsend-eval-18-trusted-input.json") + existing = bundle["idempotency"]["existing_remediation"] + + # When existing remediation identity and ordinary markers are inspected + assert [item["key"] for item in existing] == ["TC-8100", "TC-8101"] + assert existing[0]["comments"] == [] + assert executor._has_description_digest(existing[0]) is False + assert executor._has_description_digest(existing[1]) is True + + # Then the digest repair path retains its stable task reference and retry markers + assert "triage-security:tc-8001:remediation:upstream" in bundle["idempotency"]["action_markers"] + assert bundle["issue"]["status"] == "In Progress" + assert "ai-cve-triaged" in bundle["issue"]["labels"] + + +def test_conditional_inputs_supply_each_scenarios_distinct_evidence(): + """Scenario inputs retain actual impact, duplicate, overlap, RPM and enrichment facts.""" + # Given independent JSON inputs rather than a shared generic evidence sample + bundles = {source_id: _trusted_input("fullsend-eval-{}-trusted-input.json".format(source_id)) + for source_id in [1, 2, 3, 4, 5, 8, 9, 11, 12, 18]} + + # When trusted pins are joined with their actual lock evidence + for bundle in bundles.values(): + reads = {(read["repository"], read["ref"]): read + for read in bundle["source_evidence"]["lock_files"]} + for stream in bundle["matrix"]["streams"]: + for row in stream["rows"]: + assert all((repo, ref) in reads for repo, ref in row["source_commits"].items()) + + # Then each scenario's decision is supported by its own supplied facts + assert bundles[1]["issue"]["fields"]["affected_package"] == "quinn-proto" + assert bundles[1]["issue"]["fields"]["current_user"]["accountId"] == "synthetic-engineer" + assert all('version = "1.0.13' in read["content"] + for read in bundles[2]["source_evidence"]["lock_files"]) + assert bundles[3]["jira_metadata"]["sibling_searches"][0]["issues"][0]["key"] == "TC-7999" + split_reads = bundles[4]["source_evidence"]["lock_files"] + assert [read["content"].split('version = "')[1].split('"')[0] for read in split_reads] == [ + "0.4.5", "0.4.5", "0.4.8", "0.4.8", "0.4.9", "0.4.9"] + assert bundles[5]["issue"]["fields"]["sbom_evidence"]["packages"] == [ + {"name": "openssl-libs", "version": version, "release": release} + for release, version in [("2.2.0", "3.0.7-25.el9_3"), ("2.2.1", "3.0.7-27.el9_4"), + ("2.2.2", "3.0.7-27.el9_4"), ("2.2.3", "3.0.7-28.el9_4"), + ("2.2.4", "3.0.7-28.el9_4")]] + assert "1.9.0" in bundles[8]["jira_metadata"]["related_issues"][1]["summary"] + assert bundles[8]["issue"]["fields"]["fixed_version"] == "1.8.2" + assert bundles[8]["idempotency"]["action_markers"] == ["triage-security:tc-8010:link:related:tc-8008"] + assert "5.96.1" in bundles[9]["jira_metadata"]["related_issues"][1]["summary"] + assert bundles[9]["issue"]["fields"]["fixed_version"] == "5.98.0" + preemptive = bundles[11]["jira_metadata"]["sibling_searches"][0]["issues"][0] + assert preemptive["key"] == "TC-8022" + assert "security-preemptive" in preemptive["labels"] + mitre = bundles[12]["external_evidence"]["mitre"]["body"] + assert mitre["containers"]["cna"]["affected"][0]["versions"][0]["lessThan"] == "0.4.8" + assert bundles[12]["external_evidence"]["osv"]["status"] == 503 + assert bundles[12]["external_evidence"]["osv"]["body"] == {} + + + + + +def test_conditional_retry_fixture_repairs_digest_before_resolving_new_link(recorder): + """The real partial-retry input repairs one digest and resolves its unapplied link.""" + # Given the mounted partial-retry scenario and its existing task identities + bundle = _trusted_input("fullsend-eval-18-trusted-input.json") + tasks = bundle["idempotency"]["existing_remediation"] + actions = [ + {"type": "remediation-task", "marker": "triage-security:tc-8001:remediation:" + ref, + "ref": ref, "project": "TC", "summary": task["summary"], + "description_adf": task["description"], "labels": task["labels"]} + for ref, task in zip(["upstream", "downstream"], tasks) + ] + actions.append({"type": "link", "marker": "triage-security:tc-8001:link:depend:downstream", + "link_type": "Depend", "inward": "TC-8001", "outward": "{{downstream.key}}"}) + result = {"schema_version": "1", "mode": "mutation-authorized", + "report": {"issue": "TC-8001", "outcome": "affected", "summary_markdown": "Partial retry.", + "evidence": [{"source": "trusted-input", "detail": "Existing tasks, one missing digest."}]}, + "actions": actions} + + # When the trusted executor processes the repair against a recorded Jira boundary + registry = executor.execute_plan(result, bundle) + + # Then task creation is skipped, one digest precedes the resolved downstream link + assert [call[0] for call in recorder.calls] == ["get-issue", "digest", "link"] + assert recorder.calls[0] == ("get-issue", "TC-8100") + assert recorder.calls[-1] == ("link", "TC-8001", "TC-8101", "Depend") + assert {ref: value["key"] for ref, value in registry.items()} == {"upstream": "TC-8100", "downstream": "TC-8101"} + + # Given a refreshed snapshot after the repair, the next retry performs no writes + tasks[0]["comments"] = [{"body": recorder.calls[1][3]}] + bundle["idempotency"]["action_markers"].append(actions[-1]["marker"]) + recorder.calls.clear() + assert executor.execute_plan(result, bundle) == registry + assert recorder.calls == [] diff --git a/plugins/sdlc-workflow/scripts/test_triage_security_staleness_evals.py b/plugins/sdlc-workflow/scripts/test_triage_security_staleness_evals.py new file mode 100644 index 000000000..2e9d034ee --- /dev/null +++ b/plugins/sdlc-workflow/scripts/test_triage_security_staleness_evals.py @@ -0,0 +1,82 @@ +"""Check staleness eval inputs; hosted evals verify the skill's actual outputs.""" + +from datetime import datetime, timedelta +import json +from pathlib import Path +import re + +import pytest + + +ROOT = Path(__file__).resolve().parents[3] +EVAL_DIR = ROOT / "evals" / "triage-security" + + +def _case(case_id): + """Load an existing runnable case without changing its fixture inputs.""" + cases = json.loads((EVAL_DIR / "evals.json").read_text())["evals"] + return next(case for case in cases if case["id"] == case_id) + + +def _timestamp(value): + """Parse the timezone-aware ISO timestamps supplied to the eval.""" + return datetime.fromisoformat(value.replace("Z", "+00:00")) + + +@pytest.mark.parametrize("evidence", [ + r"Last-Updated HTML comment", + r"parsed ISO 8601 (?:value|timestamp)", + r"(?:comparison|evaluation) clock", +]) +def test_stale_prompt_requests_extraction_evidence_in_output(evidence): + """Require extraction evidence in output instructions, not just supplied inputs.""" + # Given the actual stale case prompt, including its supplied timestamp + prompt = _case(19)["prompt"] + + # When isolating requests to record evidence in the grader-visible output + requests = re.findall(r"Record (.+?) in outputs/staleness-check\.md", prompt) + + # Then each asserted mechanism fact must be requested as output evidence + assert any(re.search(evidence, request) for request in requests), evidence + + +@pytest.mark.parametrize("case_id, days_old", [(19, 72), (20, 14)]) +def test_prompt_and_declared_clock_agree_with_matrix_state(case_id, days_old): + """Check executor input and declarations; graders receive output evidence only.""" + # Given the existing fresh/stale case and its unchanged matrix fixture + case = _case(case_id) + matrix = next(path for path in case["files"] if "security-matrix" in path) + updated = re.search(r"", (EVAL_DIR / matrix).read_text()) + + # When reading the executor prompt and declarative expected output + clocks = [re.search(r"Evaluation clock: (\S+)", case[field]) + for field in ("prompt", "expected_output")] + + # Then both declarations agree and the prompt requests grader-visible evidence + assert all(clocks), "prompt and expected_output need an explicit evaluation clock" + assert clocks[0][1] == clocks[1][1] == "2026-07-12T10:00:00Z" + assert _timestamp(clocks[0][1]) - _timestamp(updated[1]) == timedelta(days=days_old) + assert f"{days_old} days old" in case["expected_output"] + assert "Record the evaluation clock and computed matrix age in outputs/staleness-check.md" in case["prompt"] + + +@pytest.mark.parametrize("context", ["prompt", "expected_output"]) +@pytest.mark.parametrize("day_offset, days_old, stale", [(-1, 13, False), (0, 14, False), (1, 15, True)]) +def test_fresh_fixture_clock_straddles_strict_production_boundary(context, day_offset, days_old, stale): + """Check prompt/declaration boundary inputs, without executing or grading outputs.""" + # Given the actual fresh case, fixture and production threshold contract + case = _case(20) + matrix = next(path for path in case["files"] if "security-matrix" in path) + updated = re.search(r"", (EVAL_DIR / matrix).read_text()) + skill = (ROOT / "plugins/sdlc-workflow/skills/triage-security/SKILL.md").read_text() + step = skill.split("## Step 0.3 – Matrix Staleness Check", 1)[1].split("## Step 0.5", 1)[0] + threshold = re.search(r"older than \*\*(\d+) days\*\*", step) + assert threshold, "production contract must retain a strict older-than threshold" + + # When inspecting either side of the recorded clock without running a skill + clock = re.search(r"Evaluation clock: (\S+)", case[context]) + age = _timestamp(clock[1]) + timedelta(days=day_offset) - _timestamp(updated[1]) + + # Then the actual inputs imply fresh below/at equality and stale above it + assert age == timedelta(days=days_old) + assert (age > timedelta(days=int(threshold[1]))) is stale diff --git a/plugins/sdlc-workflow/scripts/test_verify_pr_eval_observability.py b/plugins/sdlc-workflow/scripts/test_verify_pr_eval_observability.py new file mode 100644 index 000000000..579a6520d --- /dev/null +++ b/plugins/sdlc-workflow/scripts/test_verify_pr_eval_observability.py @@ -0,0 +1,94 @@ +"""Check verify-pr evidence requests and fixtures; hosted evals grade actual outputs.""" + +import json +from pathlib import Path +import re + +import pytest + + +ROOT = Path(__file__).resolve().parents[3] +EVAL_DIR = ROOT / "evals" / "verify-pr" + + +def _case(): + """Load the existing review-feedback case without changing its inputs.""" + cases = json.loads((EVAL_DIR / "evals.json").read_text())["evals"] + return next(case for case in cases if case["id"] == 3) + + +@pytest.mark.parametrize("evidence", [ + r"author check against github-actions\[bot\]", + r"body marker check for ## Eval Results", + r"body footer check for sdlc-workflow/run-evals", + r"number of reviews matching all three checks", + r"resulting Eval Quality verdict and its effect on the Test Quality combination", +]) +def test_case3_requests_detection_evidence_in_report(evidence): + """Catch detection facts mentioned as inputs but never requested in the report.""" + # Given the actual review-feedback case prompt + prompt = _case()["prompt"] + + # When isolating requests to record evidence in the grader-visible report + requests = re.findall(r"Record (.+?) in outputs/report\.md", prompt) + + # Then every detection fact and its verdict impact is requested for actual reviews + assert any("each actual review" in request and re.search(evidence, request) + for request in requests), evidence + + +def test_case3_requests_fixture_based_detection_and_conditional_na(): + """Catch invented review evidence, inline/body confusion, and unconditional N/A.""" + # Given the actual prompt, rather than its declarative expected output + prompt = _case()["prompt"] + + # When inspecting the detection source and no-match reporting instructions + source = re.search(r"Inspect only (.+?) for Eval Quality detection", prompt) + no_match = re.search(r"When no review matches all three checks, (.+?)\.", prompt) + + # Then the report must derive evidence from the supplied Reviews list only + assert source, "the detection request must identify the supplied review inventory" + assert "supplied Reviews JSON list in pr-review-comments.md" in source[1] + assert no_match, "N/A reporting must be conditional on the actual match result" + assert "Eval Quality is N/A" in no_match[1] + assert "does not affect the Test Quality combination" in no_match[1] + assert "do not invent reviews, treat inline comments as review bodies, or call live GitHub/Jira" in prompt + + +def test_case3_fixture_supplies_review_inputs_with_no_qualifying_eval_review(): + """Verify the fixture's actual review author/body values imply no eval match.""" + # Given the review fixture mounted by the existing case + case = _case() + assert "files/pr-review-comments.md" in case["files"] + fixture = (EVAL_DIR / "files/pr-review-comments.md").read_text() + + # When reading the Reviews JSON, excluding the separate inline comments list + reviews_section = fixture.split("## Reviews (", 1)[1].split("\n## Review Comments (", 1)[0] + reviews = json.loads(re.search(r"```json\n(.+?)\n```", reviews_section, re.S)[1]) + inputs = [(review["id"], review["user"]["login"], review["body"]) for review in reviews] + checks = [(review_id, author == "github-actions[bot]", "## Eval Results" in body, + "sdlc-workflow/run-evals" in body) for review_id, author, body in inputs] + + # Then the supplied human review fails all three checks and yields no matches + assert inputs == [(20001, "reviewer-a", + "Good approach overall. A few things to address before we can merge.")] + assert checks == [(20001, False, False, False)] + assert [review_id for review_id, author, marker, footer in checks + if author and marker and footer] == [] + + +@pytest.mark.parametrize("criterion", [ + "github-actions[bot]", "## Eval Results", "sdlc-workflow/run-evals", +]) +def test_case3_assertion_agrees_with_production_detection_criteria(criterion): + """Catch criterion drift between the unchanged assertion and production contract.""" + # Given the existing case 3 detection assertion and production skill + assertion = _case()["assertions"][13] + skill = (ROOT / "plugins/sdlc-workflow/skills/verify-pr/SKILL.md").read_text() + + # When isolating the production three-condition detection contract + detection = skill.split("### Step 4a.1 – Detect Eval Result Reviews", 1)[1].split("### Step 4b", 1)[0] + + # Then both contracts retain each independently declared detection criterion + assert criterion in assertion + assert criterion in detection diff --git a/plugins/sdlc-workflow/scripts/validate-output-schema.sh b/plugins/sdlc-workflow/scripts/validate-output-schema.sh new file mode 100755 index 000000000..edef3c972 --- /dev/null +++ b/plugins/sdlc-workflow/scripts/validate-output-schema.sh @@ -0,0 +1,85 @@ +#!/usr/bin/env bash +# validate-output-schema.sh — Validate agent output against a JSON Schema. +# +# Generic script used by the harness validation_loop (ADR 0022). +# Works for any agent — the schema path is configured in the harness. +# +# Required env vars: +# FULLSEND_OUTPUT_SCHEMA — path to the JSON Schema file +# +# Optional env vars: +# FULLSEND_OUTPUT_FILE — filename to validate (default: agent-result.json) +# +# The script looks for the output file in the iteration output directory. +# The working directory is the iteration dir (set by run.go). + +set -euo pipefail + +: "${FULLSEND_OUTPUT_SCHEMA:?FULLSEND_OUTPUT_SCHEMA must be set}" + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" + +# Find the output JSON file in this iteration's output directory. +OUTPUT_DIR="output" +if [[ ! -d "${OUTPUT_DIR}" ]]; then + echo "FAIL: output directory not found" + exit 1 +fi + +_output_file="${FULLSEND_OUTPUT_FILE:-agent-result.json}" +_output_file="$(basename "${_output_file}")" +RESULT_FILE="${OUTPUT_DIR}/${_output_file}" +if [[ ! -f "${RESULT_FILE}" ]]; then + # Agents sometimes write "result.json" instead of "agent-result.json". + # Accept the common variant rather than burning a full retry iteration. + _fallback="${OUTPUT_DIR}/result.json" + if [[ "${_output_file}" == "agent-result.json" && -f "${_fallback}" ]]; then + echo "WARN: expected ${RESULT_FILE} but found ${_fallback} — using fallback" + RESULT_FILE="${_fallback}" + else + echo "FAIL: ${RESULT_FILE} not found" + exit 1 + fi +fi +echo "Validating: ${RESULT_FILE} against ${FULLSEND_OUTPUT_SCHEMA}" + +# Validate JSON is parseable. +if ! python3 -m json.tool "${RESULT_FILE}" > /dev/null 2>&1; then + echo "FAIL: ${RESULT_FILE} is not valid JSON" + exit 1 +fi + +# Validate against schema using Python's jsonschema. +# jsonschema is required — fail hard if not installed. +if ! python3 -c "import jsonschema" 2>/dev/null; then + echo "FAIL: python3 jsonschema package is not installed (required by ADR 0022)" + exit 1 +fi + +# Strip extra properties before validation. The schema stays strict +# (additionalProperties: false) to document the contract, but the +# stripping makes it forgiving for benign metadata the agent adds. +python3 "${SCRIPT_DIR}/strip_extra_properties.py" "${RESULT_FILE}" "${FULLSEND_OUTPUT_SCHEMA}" + +if ! python3 -c " +import json, sys +from jsonschema import validate, ValidationError + +with open(sys.argv[1]) as f: + instance = json.load(f) +with open(sys.argv[2]) as f: + schema = json.load(f) +try: + validate(instance=instance, schema=schema) + print('PASS: output validated against schema') +except ValidationError as e: + print(f'FAIL: schema validation error: {e.message}') + if e.path: + print(f' at: {\".\".join(str(p) for p in e.path)}') + if 'properties' in e.schema: + allowed = ', '.join(sorted(e.schema['properties'].keys())) + print(f' allowed properties: {allowed}') + sys.exit(1) +" "${RESULT_FILE}" "${FULLSEND_OUTPUT_SCHEMA}"; then + exit 1 +fi diff --git a/plugins/sdlc-workflow/skills/triage-security/SKILL.md b/plugins/sdlc-workflow/skills/triage-security/SKILL.md index 4b2f60a63..eaf2808cb 100644 --- a/plugins/sdlc-workflow/skills/triage-security/SKILL.md +++ b/plugins/sdlc-workflow/skills/triage-security/SKILL.md @@ -37,7 +37,9 @@ Do **not** use for: | 0 | Validate Configuration | CLAUDE.md | Project key, Cloud ID, Security Config | | 0.3 | Matrix Staleness Check | security-matrix.md timestamps | Staleness warning or proceed | | 0.5 | Jira Access | -- | MCP or REST API connection | -| 0.7 | Assign and Transition to Assigned | Vulnerability issue key | Issue assigned to current user, status Assigned | +| 0.6 | Fullsend Mode Detection | `FULLSEND_OUTPUT_DIR` presence | Sandbox or interactive mode | +| 0.7 | Load Trusted Fullsend Input | Mounted JSON bundle | Validated evidence or hard failure | +| 0.8 | Assign and Transition to Assigned | Vulnerability issue key | Issue assigned to current user, status Assigned | | 1 | Data Extraction | Vulnerability issue key | CVE ID, library, affected range, remote links | | 1.5 | External CVE Data Enrichment | CVE ID | Structured version ranges, cross-validated fix thresholds | | 1.7 | Embargo Check | Embargo policy URL, CVE severity | Confirmation to proceed (or stop) | @@ -50,25 +52,246 @@ Do **not** use for: | 7.5 | Release Jira Orchestration | Stream name, supportability matrix | Release Epic + Task, cross-CVE dedup | | 8 | Remediation | Impact analysis results | Remediation tasks or close recommendation | +## Fullsend dual-mode contract + +Run Steps **0.6** and **0.7** before every other numbered step. They determine the +execution mode; in interactive mode they leave the existing workflow unchanged, while +in Fullsend mode they prevent all untrusted and external reads. The literal +`${CLAUDE_PLUGIN_ROOT}` in this skill is substituted into the delivered skill body +before it runs; use it for bundled schemas and companion files, not a repo-relative +path or a shell environment variable. + +### Step 0.6 – Fullsend Mode Detection + +Detect the Fullsend gate by **presence**, not truthiness. An exported-but-empty gate +is a configuration error and must fail closed rather than reaching the credentialed +interactive workflow: + +```bash +if [ "${FULLSEND_OUTPUT_DIR+x}" = x ]; then + if [ -z "$FULLSEND_OUTPUT_DIR" ]; then + echo "ERROR: FULLSEND_OUTPUT_DIR is set but empty" >&2 + exit 1 + fi + echo "sandbox mode: $FULLSEND_OUTPUT_DIR" +else + echo "interactive mode" +fi +``` + +If `FULLSEND_OUTPUT_DIR` is genuinely absent, run the established interactive +workflow, including its configuration reads, confirmation gates, external lookups, +and permitted local matrix behavior. If it is present and non-empty, run Fullsend +sandbox mode: there are no Jira, GitHub, WebFetch, CVE, lifecycle, Git, cosign, or +other external calls; mounted repository evidence is read-only; and the sandbox must +never write `security-matrix.md`. + +### Step 0.7 – Load Trusted Fullsend Input (sandbox mode only) + +Skip this step in interactive mode. In Fullsend mode, read **only** +`/sandbox/workspace/.pre-script/triage-security-input.json` and validate it before +any analysis against +`${CLAUDE_PLUGIN_ROOT}/schemas/triage-security-input.schema.json`: + +```bash +if ! python3 - << 'PYEOF' +import json, sys +from jsonschema import validate, ValidationError + +INPUT = "/sandbox/workspace/.pre-script/triage-security-input.json" +SCHEMA = "${CLAUDE_PLUGIN_ROOT}/schemas/triage-security-input.schema.json" +try: + with open(INPUT) as f: + bundle = json.load(f) + with open(SCHEMA) as f: + schema = json.load(f) + validate(instance=bundle, schema=schema) + print("Trusted triage-security input available") +except FileNotFoundError: + print("ERROR: trusted triage-security input missing"); sys.exit(1) +except json.JSONDecodeError as e: + print(f"ERROR: trusted triage-security input is invalid JSON: {e}"); sys.exit(1) +except ValidationError as e: + path = ".".join(str(p) for p in e.path) or "" + print(f"ERROR: trusted triage-security input failed schema validation at {path}: {e.message}") + sys.exit(1) +PYEOF +then + cat > "$FULLSEND_OUTPUT_DIR/agent-result.json" << 'RESULT_EOF' +{ "error": "triage-security aborted: trusted input is missing, invalid JSON, or fails triage-security-input.schema.json; no interactive fallback is available in the sandbox." } +RESULT_EOF + exit 1 +fi +``` + +The error-object write is inside the `if ! …; then` failure branch: it runs only when +the input is missing, invalid JSON, or schema-invalid, then exits immediately without +an interactive fallback, external call, or later step. This object intentionally omits +the required `schema_version`, `mode`, `report`, and `actions` members of +`triage-security-result.schema.json`; the runner must reject it visibly as the hard +failure. A successful validation does **not** write this failure object and continues +to analysis. Initialize the successful result with its required fields: +`schema_version: "1"`, `mode` from +`authorization.mutation_authorized` (`report-only` when false, otherwise +`mutation-authorized`), an evidence-backed `report`, and ordered `actions`. A +report-only result contains only its `report-only` action. Accumulate every later +sandbox action in this result; never execute it in the sandbox. + +```json +{ + "schema_version": "1", + "mode": "report-only", + "report": { + "issue": "PROJ-123", + "outcome": "needs-review", + "summary_markdown": "Trusted-evidence triage in progress.", + "evidence": [ + { + "source": "trusted-input", + "detail": "Validated triage-security input bundle." + } + ] + }, + "actions": [ + { + "type": "report-only", + "marker": "triage-security:report-only" + } + ] +} +``` + +The shown shape is the complete report-only form. For mutation-authorized input, +replace `mode` with `mutation-authorized` and build the ordered schema actions from +the same trusted evidence; use `issue.key` from the validated bundle, never the +illustrative issue key above. + +### Fullsend action serialization and authorization + +This section applies to **every** Jira write described by this skill and its +companion procedures. It does not change interactive mode: when +`FULLSEND_OUTPUT_DIR` is absent, retain every existing engineer-confirmation prompt +and perform the confirmed Jira operation exactly as documented. + +In Fullsend mode, do not ask for confirmation and do not call Jira. Determine the +branch once from `authorization.mutation_authorized` after trusted-input validation: + +- When it is `false`, produce an evidence-backed `report-only` result with **exactly + one** `report-only` action. Its report must list every otherwise-proposed mutation + (assignment, field update, transition, comment, link, remediation task, and + reconciliation) with the trusted evidence and the reason it is withheld: trusted + runner authorization is required. Do not append a proposed mutation as an action. +- When it is `true`, append only schema-defined actions. Give each action a stable, + unique lowercase `triage-security:` marker derived from the current issue, + operation, and target. Consult `idempotency.action_markers`, `idempotency.existing_remediation`, + and trusted existing links/comments before appending. Omit an already-recorded + ordinary field, transition, comment, or link action, but **always emit each + planned `remediation-task` action with its stable marker and `ref`**. On a partial + retry the executor uses that action to re-register an existing task reference, + avoid duplicate creation, and repair a missing description digest before later + placeholder links or comments are resolved. + +| Interactive Jira operation | Mutation-authorized Fullsend action | +|---|---| +| Assignment, Affects Versions, VEX Justification, labels, resolution fields | `field-edit` | +| Assigned, In Progress, or Closed workflow change | `status-transition` | +| Triage, correction, overlap, reconciliation, cross-stream, and summary comments | `comment` with `body_adf`, preserving required ADF mentions and footnotes | +| Related, Depend, or Blocks relationship | `link` with the matching `link_type` | +| Remediation Task creation | `remediation-task` with `ref`, `description_adf`, labels, and optional priority/fix versions | +| Registering a separately trusted, pre-existing reference before dependent work | `resolve-reference` | + +Preserve the existing step order. For every newly planned remediation task, append +the `remediation-task` action before its dependent `link` actions and before later +comments (task-list, reconciliation, cross-stream, and post-triage summary). The +binding executor registers the task reference and posts its description digest +exactly once inside `remediation-task`, including when it reuses an existing task on +a partial retry. Do not serialize a separate digest `comment` or a +post-creation `resolve-reference` action. Use `{{ref.key}}` in later link/comment +actions where the result schema permits it; never invent a Jira key in the sandbox. +This digest-before-link rule applies to standard and preemptive tasks alike. + +### Trusted-evidence map for Fullsend mode + +After validation, the input bundle is authoritative. Do not supplement an absent or +empty trusted collection by reading the network, Jira, GitHub, local configuration, +or a repository. Use `configuration` for Step 0 and deployment settings; `issue` and +`remote_links` for Step 1; `external_evidence.mitre` and `.osv` for Step 1.5; +`configuration.embargo_policy_url` plus issue/evidence severity for Step 1.7; +`matrix.streams` and `source_evidence` for Step 2; `jira_metadata` for Steps 3–7; +`external_evidence.lifecycle` for Step 5; and `idempotency` for every duplicate, +existing-action, comment, link, and remediation check. `authorization` controls the +result mode. An empty schema-required collection means "no trusted match"; it never +permits a fallback read. + +### Fullsend final output (successful analysis only) + +After all Fullsend analysis has completed, serialize the completed accumulated result +to **exactly** `$FULLSEND_OUTPUT_DIR/agent-result.json`. This is the only allowed +sandbox-side write after a successful validation; it must contain the final +`schema_version`, `mode`, evidence-backed `report`, and ordered `actions`, not the +initial example or the validation-failure object. + +```bash +cat > "$FULLSEND_OUTPUT_DIR/agent-result.json" << 'RESULT_EOF' + +RESULT_EOF +``` + +Before finishing, validate that exact file both as JSON and against the delivered +result schema. Schema validation failure is a hard failure: do not attempt an +interactive or external fallback. + +```bash +python3 - << 'PYEOF' +import json, os, sys +from jsonschema import validate, ValidationError + +OUTPUT = os.path.join(os.environ["FULLSEND_OUTPUT_DIR"], "agent-result.json") +SCHEMA = "${CLAUDE_PLUGIN_ROOT}/schemas/triage-security-result.schema.json" +try: + with open(OUTPUT) as f: + result = json.load(f) + with open(SCHEMA) as f: + schema = json.load(f) + validate(instance=result, schema=schema) + print("Final Fullsend result validated") +except (OSError, json.JSONDecodeError, ValidationError) as e: + print(f"ERROR: final Fullsend result failed JSON/schema validation: {e}") + sys.exit(1) +PYEOF +``` + ## Guardrails -- **This skill is Jira-only for output**, with one exception: it may write to +- **Interactive mode only:** This skill is Jira-only for output, with one exception: it may write to local `security-matrix.md` files in the project working directory to populate - or update the supportability matrix (see Step 2.1). All other mutations go - through Jira. + or update the supportability matrix (see Step 2.1). All other mutations go through + Jira. In Fullsend mode, replace Jira mutations with ordered result actions and never + write `security-matrix.md`. - **Read-only source access.** Source repositories and Konflux release repos are accessed only via `git show :` for lock file inspection and fallback matrix reads. No checkouts, no branch switches, no file modifications - outside local `security-matrix.md` files. -- **Every Jira mutation requires confirmation.** Present the proposed change and rationale - to the engineer; wait for explicit approval before executing. Never perform bulk or - silent Jira writes. -- **Do NOT fabricate data.** Every version, commit hash, dependency version, and version - impact assessment must come from actual `git show` output or Jira API responses — never - invented or assumed. + outside interactive local `security-matrix.md` files or the narrow Fullsend output + exception below. +- **Interactive mode only: every Jira mutation requires confirmation.** Present the + proposed change and rationale to the engineer; wait for explicit approval before + executing. Never perform bulk or silent Jira writes. In Fullsend mode, do not ask + for confirmation: use `authorization.mutation_authorized` deterministically and + serialize the corresponding ordered actions (or the report-only result) instead. +- **Do NOT fabricate data.** In interactive mode, every version, commit hash, + dependency version, and version-impact assessment must come from actual `git show` + output or Jira API responses — never invented or assumed. In Fullsend mode, the + validated trusted input bundle is also accepted provenance: derive those values from + its `source_evidence` and `jira_metadata`, never infer or invent them. +- **Fullsend output exception:** In Fullsend mode, the only permitted sandbox-side file + write is `$FULLSEND_OUTPUT_DIR/agent-result.json`: write either the deliberately + schema-invalid validation failure object or the completed valid accumulated result. + Do not create, modify, or repair any other local file. - **Do NOT use Edit, Write, or Bash tools** to change files — except for local `security-matrix.md` files in the project working directory (see Step 2.1). - Only use Bash for read-only `git show` commands and JIRA REST API fallback scripts. + In Fullsend mode, the narrow output exception above also permits writing + `$FULLSEND_OUTPUT_DIR/agent-result.json`; otherwise only use Bash for read-only + `git show` commands and JIRA REST API fallback scripts. - If any step fails (e.g., Jira MCP unavailable, lock file not found, repo not cloned), stop and inform the user rather than attempting alternative actions. @@ -92,6 +315,12 @@ Follow the format in `shared/comment-footnote.md`, using skill name `triage-secu ## Step 0 – Validate Project Configuration +**Fullsend mode:** do not read `CLAUDE.md`. Use the validated bundle's +`configuration` member as the complete Security Configuration input. Its required +version streams and source repositories replace the interactive configuration read; +an optional configuration member that is absent remains absent rather than triggering +a local or external lookup. + Before proceeding, read the project's CLAUDE.md and verify that the following sections exist under `# Project Configuration`: @@ -148,6 +377,10 @@ Extract the following from the configuration for use in later steps: ## Step 0.3 – Matrix Staleness Check +**Fullsend mode:** the trusted `matrix.streams` evidence is the supplied matrix state. +Do not read a local matrix, repopulate it, repair it, or offer a refresh; matrix +staleness handling and every local matrix write are interactive/trusted-runner-only. + Before proceeding with triage, verify that each version stream's `security-matrix.md` has been updated recently enough to reflect the current release landscape. A stale matrix can cause triage to miss newly released versions or use outdated source commit @@ -189,6 +422,10 @@ updated `Last-Updated` timestamp), continue with the refreshed matrix. ## Step 0.5 – JIRA Access Initialization +**Fullsend mode:** skip JIRA access initialization and every REST fallback. The +validated `issue`, `remote_links`, `jira_metadata`, and `idempotency` evidence replace +Jira reads; planned mutations are accumulated in the result rather than called. + Follow the JIRA Access protocol in `shared/jira-access-strategy.md`. **REST API equivalents for this skill's operations:** @@ -204,13 +441,25 @@ Follow the JIRA Access protocol in `shared/jira-access-strategy.md`. **Exception for Bash tool:** When using REST API fallback, this skill may use `bash -c "python3 scripts/jira-client.py "` for JIRA operations only. -## Step 0.7 – Assign and Transition to Assigned +## Step 0.8 – Assign and Transition to Assigned Assign the CVE Vulnerability issue to the current user and transition it to Assigned status. This provides immediate visibility into who is actively triaging the issue and enables Step 7 (Concurrent Triage Detection) to reliably identify active work. +**Fullsend mode:** obtain the current issue state from `issue` and existing markers +from `idempotency`; do not retrieve a user, assignment, or transition from Jira. An +assignment `field-edit` is allowed only when the validated input contains the +current triager's account ID in trusted `issue.fields`. Never infer an account ID +from a name or look one up. If assignment is required but that trusted ID is absent, +withhold **only** the assignment and append a blocked, evidence-backed `report-only` +recommendation naming the missing trusted assignee identity and required trusted-input +configuration. Continue the mutation-authorized plan: serialize every other authorized +field edit, transition, comment, remediation task, and link in the existing step order. +If the ID is present and authorization permits it, append the assignment and transition +actions in order. + 1. **Retrieve the current user's Jira account ID:** ``` @@ -246,6 +495,10 @@ active work. ## Inputs +**Fullsend mode:** the validated `issue.key` is the invocation target. Do not run +discovery JQL, prompt for a missing issue key, or query Jira; schema validation has +already established the input shape. Interactive discovery mode below is unchanged. + The user provides a single Vulnerability issue key. Example: @@ -367,6 +620,12 @@ in project ." ## Step 1 – Data Extraction +**Fullsend mode:** do not fetch the issue or its remote links. Extract these fields +from the validated `issue` and `remote_links` bundle members, using `configuration` +for the component-label and stream mappings. If trusted evidence cannot support a +critical extraction, record a blocked/needs-review result; never ask Jira, GitHub, or +the user to fill the gap from an untrusted source. + Fetch the Vulnerability issue from Jira: ``` @@ -493,6 +752,12 @@ If the affected repository is not found in the Source Repositories table, defaul ## Step 1.5 – External CVE Data Enrichment +**Fullsend mode:** do not query MITRE or OSV. Use the captured +`external_evidence.mitre` and `external_evidence.osv` records, including their source +URLs, retrieval status, and bodies, for the same cross-validation and fix-threshold +precedence rules. A non-success record is the trusted evidence of unavailability; do +not retry it over the network. + After extracting data from the Jira description, query external CVE databases for structured vulnerability data. This is not a fallback — external sources are **always** queried to supplement and cross-validate the Jira description data. @@ -560,6 +825,11 @@ impact comparisons. ## Step 1.7 – Embargo Check +**Fullsend mode:** derive the configured policy from `configuration` and severity from +the trusted issue/external evidence. Do not fetch a policy URL or use an interactive +confirmation. Represent an embargo block or permitted next action in the accumulated +result according to trusted `authorization`. + This step is an advisory warning gate for high-severity vulnerabilities that may be under embargo. It does not enforce embargo procedures — it surfaces a warning and links to the organization's embargo policy for the engineer to verify. @@ -603,6 +873,12 @@ files at pinned commits, builds the version impact table, and checks for upstrea Read `version-impact-analysis.md` for the detailed procedures (Steps 2.1–2.5). +**Fullsend mode:** that procedure consumes only `matrix.streams`, +`source_evidence.lock_files`, and `source_evidence.development_streams`. It preserves +all-supported-version coverage, released pinned-commit evidence, development-head +evidence, retag propagation, dependency-chain analysis, and enriched fix-threshold +rules without running Git, cosign, or a matrix fallback. + **Sub-steps:** - **2.1** – Load the supportability matrix from local files (with Konflux repo fallback) - **2.2** – Detect the development stream via unreleased Jira versions @@ -622,6 +898,12 @@ version lifecycle status, and detecting already-fixed scenarios. Read `jira-triage-operations.md` for the detailed procedures. +**Fullsend mode:** use `jira_metadata.versions` for version discovery, +`jira_metadata.sibling_searches` and `.related_issues` for duplicate/sibling/overlap +analysis, `external_evidence.lifecycle` for lifecycle status, and `idempotency` for +already-fixed and existing-artifact checks. Do not issue JQL, fetch Jira records, or +query lifecycle pages; turn any proposed mutation into an ordered result action. + - **Step 3** – Affects Versions Correction: discover available Jira versions dynamically, compare against the version impact table, and correct with engineer confirmation @@ -670,6 +952,12 @@ flowchart TD **Important**: This skill never creates Vulnerability issues. PSIRT owns Vulnerability issue creation — the skill only creates remediation **Tasks**. +**Fullsend mode:** retain the existing decision tree and evidence thresholds, but +append remediation, field, status, comment, link, and reference operations to the +ordered result only. Authorization false produces the report-only action and names +each withheld mutation; authorization true still does not execute a sandbox-side Jira +call. + ## Step 7 – Concurrent triage detection Before proceeding to Case A/B/C branching, check whether another engineer is @@ -680,6 +968,10 @@ simultaneously. Follow the concurrent triage detection protocol in `jira-triage-operations.md` — Step 7. +**Fullsend mode:** determine concurrent or existing triage solely from the validated +`jira_metadata` and `idempotency` entries. Do not run JQL; report a missing trusted +match as no known match, not permission to search externally. + If the Upstream Affected Component custom field is not configured, skip this step entirely. diff --git a/plugins/sdlc-workflow/skills/triage-security/jira-triage-operations.md b/plugins/sdlc-workflow/skills/triage-security/jira-triage-operations.md index cd5f7610a..39fbf73ef 100644 --- a/plugins/sdlc-workflow/skills/triage-security/jira-triage-operations.md +++ b/plugins/sdlc-workflow/skills/triage-security/jira-triage-operations.md @@ -5,6 +5,38 @@ triage-security skill. These steps handle Affects Versions correction, duplicate and sibling detection, cross-CVE overlap detection, preemptive task reconciliation, version lifecycle checks, and already-fixed detection. +## Fullsend action mapping + +This file's interactive procedures and confirmation prompts are unchanged when +`FULLSEND_OUTPUT_DIR` is absent. In Fullsend mode, use only validated trusted input +and `authorization.mutation_authorized`; never call Jira or ask the engineer to +confirm a sandbox action. + +If authorization is false, do not serialize any mutation from Steps 3–7. Return the +top-level evidence-backed report-only result with exactly its one `report-only` +action, and name each withheld correction, link, closure, assignment, label update, +or comment in the report. If authorization is true, map each write one-for-one to +the existing result-schema types below. Every marker is stable and unique in the +form `triage-security:::` (using only +schema-valid marker characters), and existing `idempotency.action_markers` or trusted +existing artifacts mean the action is omitted on retry. + +| Procedure write | Fullsend action | +|---|---| +| Affects Versions correction, VEX value, assignment, add/remove label, resolution | `field-edit` | +| Assigned, In Progress, Closed | `status-transition` | +| Affects Versions, duplicate, overlap, lifecycle, already-fixed, reconciliation, and skip comments | `comment` with `body_adf` | +| Related, Depend, Blocks links | `link` with the same `link_type` | + +Build all comment bodies as ADF documents, retaining required Comment Footnotes and +ProdSec mentions. Step 4.4 reconciliation is specifically a `link` action for the +new `Depend` relationship and a `field-edit` action that removes +`security-preemptive`; it is never a direct Jira update in the sandbox. Keep the +skill's step order, and defer comments that list newly created tasks until the +executor-owned `remediation-task` action has registered each task reference and +posted its description digest exactly once, followed by that task's links. Do not +serialize a separate digest comment or post-creation `resolve-reference` action. + ## Step 3 – Affects Versions Correction ### 3.1 – Discover available Jira versions diff --git a/plugins/sdlc-workflow/skills/triage-security/remediation-templates.md b/plugins/sdlc-workflow/skills/triage-security/remediation-templates.md index 33ff3e762..b6a936baf 100644 --- a/plugins/sdlc-workflow/skills/triage-security/remediation-templates.md +++ b/plugins/sdlc-workflow/skills/triage-security/remediation-templates.md @@ -304,6 +304,42 @@ compatibility), omit the coordination guidance entirely — do not add the subse ## Jira Issue Creation +### Fullsend serialization + +The following creation pseudocode is interactive-only and retains its existing +confirmation behavior. In Fullsend mode, do not call `create_issue`, `add_comment`, +or `create_link`. Use the validated task description, labels, priority, fix-version, +and idempotency evidence to serialize schema actions instead. + +When `authorization.mutation_authorized` is false, no task, digest, reference, or +link action is permitted. The evidence-backed top-level report must name each +withheld remediation task and its intended digest/link work, and `actions` contains +exactly the single `report-only` action. + +When authorization is true, each standard or preemptive task has a stable, +idempotent `triage-security:` marker and a schema-valid reference name. Always emit +its `remediation-task` action with that marker and reference, even when trusted +`idempotency` evidence records the marker or matching existing remediation. This +lets the binding executor re-register the existing reference, avoid duplicate task +creation, and repair a missing description digest during a partial retry. Later +`{{ref.key}}` placeholders then resolve against that rehydrated reference. For each +task, append actions in this exact order: + +1. `remediation-task` with `ref`, project, summary, `description_adf`, labels, and + any supported priority or fix versions. The executor registers its reference and + posts the description digest exactly once as part of this action. +2. `link` actions: `Depend` for a task and its own CVE, `Related` for a preemptive + task and its originating CVE, and `Blocks` from upstream to downstream. +3. Only after all of that task's links, append later comments such as task-list, + cross-stream, reconciliation, or post-triage summary comments. + +Use `{{ref.key}}` in schema fields that accept an issue reference instead of +inventing the Jira key. Do not add a standalone digest `comment` or a +post-creation `resolve-reference`: the executor owns reference registration and +exactly-once digest posting for `remediation-task`. Serialize a preemptive task's +`security-preemptive` label in its `remediation-task`, and serialize later +reconciliation label removal as a `field-edit`; never update Jira from the sandbox. + After creating each remediation task, post a description digest comment per `shared/description-digest-protocol.md`. The digest comment MUST be posted before creating issue links or other comments on the task. diff --git a/plugins/sdlc-workflow/skills/triage-security/version-impact-analysis.md b/plugins/sdlc-workflow/skills/triage-security/version-impact-analysis.md index 9cbad9650..36ebd578e 100644 --- a/plugins/sdlc-workflow/skills/triage-security/version-impact-analysis.md +++ b/plugins/sdlc-workflow/skills/triage-security/version-impact-analysis.md @@ -4,8 +4,35 @@ This companion file contains the detailed procedures for Step 2 of the triage-security skill. It determines which supported product versions actually ship the vulnerable dependency by reading lock files at pinned source commits. +## Fullsend trusted-evidence substitution + +When Step 0.6 selected Fullsend mode, this procedure must consume only the +validated `triage-security-input.json` bundle. Use `configuration.version_streams` +for stream configuration, `matrix.streams[].rows` for every released-version row and +its `source_commits`/`retag_of`, `source_evidence.lock_files` for released pinned +reads, and `source_evidence.development_streams` for development-head reads. The +bundle's `external_evidence` preserves captured MITRE/OSV fix-threshold evidence and +lifecycle evidence; `jira_metadata` and `idempotency` preserve the accompanying Jira +evidence without a Jira read. + +The analysis rules do not weaken: enumerate **all** supported matrix rows, use the +released pinned commits (never a branch tip), preserve retag carry-forward behavior, +trace the dependency chain from supplied lock/manifest content, and apply the +cross-validated external fix threshold before the Jira prose fallback. A present but +empty trusted collection is evidence of no known match; it never authorizes an +external fallback. + +Fullsend makes no local matrix read, Git/GitHub/Jira/WebFetch/cosign/lifecycle call, +or repository mutation. Matrix fallback, format repair, on-demand population, and +all local `security-matrix.md` writes are interactive/trusted-runner-only; the +sandbox must never write security-matrix.md. + ## 2.1 – Load the supportability matrix +**Fullsend mode:** load the aggregated matrix directly from `matrix.streams`, matching +each entry to `configuration.version_streams`. Treat its rows as trusted, read-only +evidence; do not open the configured local `security-matrix.md` path. + For each row in the **Version Streams** table in Security Configuration, read the `security-matrix.md` file at the path given in the **Security Matrix Path** column, resolved relative to the project's working directory. Step 0.3 already verified @@ -24,6 +51,9 @@ identifies a stream, so no chaining is needed. ### Fallback to Konflux repos +**Fullsend mode:** this fallback is prohibited. The runner must provide the required +matrix rows in the validated bundle; do not invoke `git show` or save a local copy. + When a local `security-matrix.md` file does not exist at the configured path, fall back to reading from the Konflux release repo via `git show`: @@ -48,6 +78,10 @@ the user which streams to configure before proceeding. ### 2.1.1 — Validate matrix format +**Fullsend mode:** do not validate, repair, or rewrite a mounted/local matrix. Schema +validation of the trusted bundle is the entry gate; matrix format validation and its +interactive warnings remain interactive/trusted-runner-only. + After loading each stream's matrix file (from local path or Konflux fallback), validate its structure against the canonical template at `docs/templates/security-matrix.template.md` before proceeding to aggregation. @@ -125,6 +159,10 @@ Aggregate all versions from all streams into a single working matrix. ### On-demand matrix population +**Fullsend mode:** on-demand population is prohibited. Missing matrix coverage is a +trusted-input deficiency to report, not a reason to query repositories or write a +matrix. + If a stream's Supportability Matrix is empty or missing rows (e.g., the template was scaffolded but never populated), research and fill it in before proceeding: @@ -143,6 +181,11 @@ repositories are read-only, and Jira is the only other output channel. ## 2.2 – Detect the development stream +**Fullsend mode:** use `jira_metadata.versions` to identify unreleased versions and +the matching `configuration.version_streams` entry, then use the matching +`source_evidence.development_streams` record as the development-head evidence. Do not +call Jira or inspect a branch directly. + Query Jira for unreleased versions to identify the current development stream: 1. Call `getJiraIssueTypeMetaWithFields` for the Vulnerability issue type in the @@ -161,6 +204,14 @@ there is no released version yet. ## 2.3 – Extract dependency versions +**Fullsend mode:** select the matching `source_evidence.lock_files` record by +repository, pinned `ref`, and path for each released matrix row, and the matching +`source_evidence.development_streams` record for the development stream. Parse only +the captured `content` using its recorded `command`; do not execute that command, +`git show`, `which cosign`, or an SBOM download in the sandbox. Preserve the existing +lock-file, RPM-origin, retag, dependency-chain, and fix-threshold decisions using this +trusted content. + **Environment variable resolution:** Paths in the Version Streams table may contain environment variable references (e.g., `${TRUSTIFY_GL_PATH}`) in the **Konflux Release Repo** and **Local Path** columns. These must be expanded before using the @@ -238,6 +289,13 @@ For each version in the aggregated matrix (plus the development stream): ### 2.3.5 – Dependency chain context +**Fullsend mode:** derive direct/transitive classification, dependency paths, +profiles, introduction points, RPM origin, and any supplied SBOM comparison only from +the matching trusted `source_evidence` content. Do not inspect manifests, Dockerfiles, +or SBOMs by calling Git, cosign, or the network. If the bundle lacks evidence required +to classify a chain, record that limitation in the result rather than filling it with +a sandbox read. + For affected versions (where the vulnerable dependency is in the lock file and within the affected range), trace the dependency chain to give the engineer context about how the vulnerable package entered the tree. This information helps assess @@ -512,6 +570,11 @@ engineer can see both the impact and the remediation path at a glance. ## 2.5 – Upstream fix check +**Fullsend mode:** determine upstream-fix status from the matching captured +`source_evidence.development_streams` (and released `lock_files` where applicable). +The trusted runner has already performed the source read; do not run `git -C`, read a +branch, or otherwise query an upstream repository in the sandbox. + For each affected stream, check whether the upstream source repository has already fixed the vulnerability on the branch that feeds that stream. Read the **Upstream Branch** column from the stream's Ecosystem Mappings table. diff --git a/plugins/sdlc-workflow/skills/verify-pr/SKILL.md b/plugins/sdlc-workflow/skills/verify-pr/SKILL.md index 3069fad70..c5b7c9239 100644 --- a/plugins/sdlc-workflow/skills/verify-pr/SKILL.md +++ b/plugins/sdlc-workflow/skills/verify-pr/SKILL.md @@ -9,6 +9,30 @@ argument-hint: "[jira-issue-id]" You are an AI verification assistant that orchestrates PR verification through parallel domain sub-agents. You verify a pull request against its Jira task's acceptance criteria and deterministic guardrails. You classify PR review feedback, dispatch domain sub-agents for parallel analysis, aggregate their findings, create tracked Jira sub-tasks for required code fixes, and investigate root causes of implementation mistakes across the full workflow chain. You post findings to both GitHub and Jira, but you do **NOT** modify code and do **NOT** auto-merge. +## Resolving this skill's own files + +This skill reads several of its own bundled files (sub-skill instruction files, +dispatch/finding templates, JSON schemas, the plugin manifest). **Always resolve these +from `${CLAUDE_PLUGIN_ROOT}`** — never a repo-relative path like `plugins/sdlc-workflow/...`. +The CWD is **not** always the plugin's repository: when verify-pr reviews a PR against +`sdlc-plugins` itself, the CWD is the target checkout, so a repo-relative path would read +that branch's copy and let the PR under review **shadow** the pinned, stable skill actually +running. `${CLAUDE_PLUGIN_ROOT}` always points at the delivered plugin, in every mode — +interactive (Claude Code) and sandbox (fullsend). + +`${CLAUDE_PLUGIN_ROOT}` is a **textual token** Claude Code substitutes into this skill's +markdown body (this whole file, including fenced code blocks) before the skill runs — it +is **not** an OS environment variable. Write the literal token wherever you need the path: +prose, Read/Glob paths, and inside a `<< 'PYEOF'` heredoc alike — substitution happens +before execution, so even inside a single-quoted heredoc it is already the absolute path by +the time Python runs. Do **not** read it at runtime via `os.environ["CLAUDE_PLUGIN_ROOT"]` +or `$CLAUDE_PLUGIN_ROOT` in a shell — it is not exported to the shell or subprocesses, so +those resolve to empty / `KeyError`. + +Note: path patterns used to **filter the PR diff** (e.g. detecting changes under +`plugins/sdlc-workflow/skills/run-evals/`) stay repo-relative — those describe files +inside the PR being reviewed, not files this skill reads. + ## Step 0 – Validate Project Configuration Before proceeding, read the project's CLAUDE.md and verify that the following sections exist under `# Project Configuration`: @@ -58,57 +82,169 @@ Before attempting any JIRA operations throughout this skill, determine the acces Refer to `shared/jira-rest-fallback.md` for complete implementation details. -## Inputs - -The user will provide a Jira issue ID for a task that has an associated PR. +## Step 0.6 – Sandbox Mode Detection + +Detect whether the `FULLSEND_OUTPUT_DIR` environment variable is **present** in the +environment (exported at all), independently of whether it holds a value. Use +presence detection (`${VAR+x}`), not a default-value expansion (`${VAR:-...}`): the +`:-` operator treats an exported-but-empty value the same as unset, which would send +a runner that exported the gate with an empty value down the credentialed +interactive path — the opposite of the intended tokenless behavior. An +exported-but-empty value is a misconfiguration and must fail fast, not silently fall +back: + +```bash +if [ "${FULLSEND_OUTPUT_DIR+x}" = x ]; then + if [ -z "$FULLSEND_OUTPUT_DIR" ]; then + echo "ERROR: FULLSEND_OUTPUT_DIR is set but empty" >&2 + exit 1 + fi + echo "sandbox mode: $FULLSEND_OUTPUT_DIR" +else + echo "interactive mode" +fi +``` -Example: +If **present** (and non-empty), this skill is running inside a fullsend sandbox (no +Jira or GitHub tokens, no `gh` CLI, no network egress). Switch to **sandbox mode**: -/verify-pr PROJ-231 +- Do NOT call Jira write APIs (`create_issue`, `add_comment`, `create_issue_link`, + `transition_issue`) directly. +- Do NOT call GitHub read APIs (`gh pr view`/`gh pr diff`/`gh api`) — every PR read + is pre-fetched into `verify-pr-input.json` and the PR head is already checked out + (GitHub is tier-1; there is no `gh`/`GH_TOKEN` in the sandbox). +- Do NOT post GitHub PR comments or replies directly. +- Instead, accumulate every write operation as an action in an in-memory JSON + structure, and at the end of execution write the complete result to + `$FULLSEND_OUTPUT_DIR/agent-result.json` (see Step 9's **Sandbox Mode Output**). -## Comment Footnote +If **genuinely unset** (absent from the environment), execute in **interactive +mode** — the current behavior, calling Jira and GitHub APIs directly. Marketplace +users are unaffected. -Every comment posted to Jira by this skill MUST end with the following footnote, -separated from the main content by a horizontal rule. +Throughout the remaining steps, when you encounter a write operation: +- **Interactive mode:** execute it directly (existing behavior). +- **Sandbox mode:** append it to the actions array instead — each step below gives the + matching action shape. -Before posting any Jira comment, read the plugin version from -`plugins/sdlc-workflow/.claude-plugin/plugin.json` and extract the `version` field. -Use this value as `{version}` in the footer below. - -Use ADF `contentFormat` to ensure the rule and text render correctly: +Initialize the accumulator as an in-memory JSON structure: ```json { - "type": "rule" -}, + "report": {}, + "actions": [] +} +``` + +The `report` object and every action conform to +`${CLAUDE_PLUGIN_ROOT}/schemas/verify-pr-result.schema.json`; the runner's +`post_script` executes the accumulated actions after the sandbox exits. + +## Step 0.7 – Load Pre-Fetched Data (sandbox mode only) + +**Skip this step in interactive mode.** + +In sandbox mode the `pre_script` fetches all task and PR data on the trusted runner +(where the tokens live) and mounts it read-only into the sandbox. Validate it against +`${CLAUDE_PLUGIN_ROOT}/schemas/verify-pr-input.schema.json` before using it — a +valid-JSON but structurally wrong or incomplete prefetch (e.g. missing the `github` +bundle or `task` fields) must be treated as invalid, not fail deep inside a later step: + +```bash +python3 - << 'PYEOF' +import json, sys +from jsonschema import validate, ValidationError + +INPUT = "/sandbox/workspace/.pre-script/verify-pr-input.json" +# ${CLAUDE_PLUGIN_ROOT} below is a Claude Code body-substitution token: it is already +# the delivered plugin's absolute path by the time this command runs (NOT a shell/env var). +SCHEMA = "${CLAUDE_PLUGIN_ROOT}/schemas/verify-pr-input.schema.json" +try: + with open(INPUT) as f: + instance = json.load(f) + with open(SCHEMA) as f: + schema = json.load(f) + validate(instance=instance, schema=schema) + print("Pre-fetched data available") +except FileNotFoundError: + print("ERROR: Pre-fetched data missing"); sys.exit(1) +except json.JSONDecodeError as e: + print(f"ERROR: Pre-fetched data is not valid JSON: {e}"); sys.exit(1) +except ValidationError as e: + path = ".".join(str(p) for p in e.path) or "" + print(f"ERROR: Pre-fetched data failed schema validation at {path}: {e.message}") + sys.exit(1) +PYEOF +``` + +If the file validates against the schema: + +- Read `task` (`summary`, `description` in the tracker's native ADF, `status`, + `labels`, `issue_links`, `custom_fields`) and `pr_url`. **Skip Step 1 (Fetch and + Parse Jira Task) and Step 2 (Identify PR)** — parse the task sections from + `task.description` and take the PR URL from `pr_url`. +- Read the `github` bundle (`pr_repo`, `pr_number`, `headRefName`, `commit_sha`, + `diff`, `stat`, `reviews`, `review_comments`, `issue_comments`, `commits`, + `check_runs`, `check_run_logs_path`) and use the already-checked-out PR-head + tree. **Skip every GitHub read step** — Step 3 (checkout), Step 4a + (review/comment fetches), Step 5a (diff/stat/commits), and Step 9's HEAD-SHA + retrieval — using the bundle fields instead. `check_runs` carries the head-SHA + CI check-run outcomes (name/status/conclusion/details_url); pass it as the + Correctness sub-agent's **CI Status** dispatch input so its Check 1 reads real + CI data instead of calling `gh` (there is no `gh` CLI or egress in the sandbox). + `check_run_logs_path` is the mounted path of the concatenated failed-check logs + (empty when no check failed); pass it as the Correctness sub-agent's **CI + Failure Logs** input so Check 1b can Read the real failure logs on a FAIL + instead of inferring from the diff. +- Read the `idempotency.related_issues` array (each entry has `key`, `summary`, + `labels`, `description`, `issuetype`, `comments`) — the task's existing sub-tasks + and linked issues, prefetched on the runner. **Use it for every idempotency read** + (Steps 6d, 6f, 7c) instead of calling Jira; the sandbox has no token. The schema + **requires** this array, so a prefetch that passes validation always carries it — an + empty list (never absent) means the task has no sub-tasks/linked issues yet. Treat + present-but-empty as "no known duplicates"; never fall back to a Jira read for dedup. + +If the validation above fails (file missing, not valid JSON, or not conforming to the +schema), the handling depends on the mode: + +- **Sandbox mode (`FULLSEND_OUTPUT_DIR` is set):** do **not** fall back to Steps 1–3 or + direct Jira/GitHub reads — the sandbox has no tokens, no `gh` CLI, and no network + egress, so a credentialed fallback cannot succeed (it would fail obscurely or hang). + Fail fast and loud: write a structured failure result and stop the skill without any + further step or write operation. + +```bash +cat > "$FULLSEND_OUTPUT_DIR/agent-result.json" << 'RESULT_EOF' { - "type": "paragraph", - "content": [ - { - "type": "text", - "text": "This comment was AI-generated by " - }, - { - "type": "text", - "text": "sdlc-workflow/verify-pr", - "marks": [ - { - "type": "link", - "attrs": { - "href": "https://github.com/RHEcosystemAppEng/sdlc-plugins" - } - } - ] - }, - { - "type": "text", - "text": " v{version}." - } - ] + "error": "verify-pr aborted: pre-fetched input (/sandbox/workspace/.pre-script/verify-pr-input.json) is missing, unparseable, or does not conform to verify-pr-input.schema.json. The pre_script must produce a valid input bundle before the sandbox runs; no credentialed fallback is possible inside the sandbox." } +RESULT_EOF ``` -Append these two nodes at the end of the ADF document's `content` array. + This result intentionally omits the `report`/`actions` keys required by + `verify-pr-result.schema.json`, so the runner's output validation rejects it and + surfaces the `error` as a hard failure rather than silently attempting the + interactive path. **Stop execution here — do not run any subsequent step.** + +- **Interactive mode (`FULLSEND_OUTPUT_DIR` is unset):** log a warning and fall back to + Steps 1–3 and the direct GitHub reads (existing interactive behavior, unchanged). + +## Inputs + +The user will provide a Jira issue ID for a task that has an associated PR. + +Example: + +/verify-pr PROJ-231 + +## Comment Footnote + +Every comment posted to Jira by this skill MUST end with the footnote defined in +`shared/comment-footnote.md`, using skill name `verify-pr`. **Override** that doc's +version-path instruction: read the plugin version from +`${CLAUDE_PLUGIN_ROOT}/.claude-plugin/plugin.json` — never the repo-relative path (see +"Resolving this skill's own files"). Append the two ADF nodes (rule + paragraph) at the +end of the comment document's `content` array. ## Step 1 – Fetch and Parse Jira Task @@ -140,6 +276,10 @@ or not configured, ask the user to provide the PR URL. Extract the PR number and repository from the URL for use in subsequent `gh` commands. +**Sandbox mode:** skip the custom-field read — take the PR URL from `pr_url` and the +`owner/repo` + PR number from `github.pr_repo` / `github.pr_number` in the pre-fetched +data (Step 0.7). + ## Step 3 – Checkout PR Branch Ensure the PR branch is checked out locally so that subsequent steps inspect the correct code. @@ -157,17 +297,17 @@ gh pr view --json headRefName -R git branch --show-current ``` -3. If the branches match, proceed without action — the correct code is already available locally (e.g. the author running self-verification after `/implement-task`). +3. If the branches match, proceed — the correct code is already local (author self-verification after `/implement-task`; no checkout needed). -4. If they differ, check out the PR branch: +4. If they differ, check out the PR branch (reviewer/CI audit from an arbitrary branch): ``` gh pr checkout -R ``` -This step supports two use cases: -- **Author self-verification** — the contributor already has the PR branch checked out; no checkout needed. -- **Reviewer/CI audit** — another person or CI job runs `/verify-pr` from an arbitrary branch; the PR branch must be checked out first. +**Sandbox mode:** skip this step entirely — the runner has already checked out the PR +head tree (`github.headRefName` at `github.commit_sha`) and there is no `gh` CLI in +the sandbox. Inspect the working tree directly. ## Step 4 – Classify Review Feedback @@ -183,6 +323,10 @@ gh api repos//pulls//reviews gh api repos//pulls//comments ``` +**Sandbox mode:** do not call `gh api` — use `github.reviews`, `github.review_comments`, +and (for the enumeration below) `github.issue_comments` from the pre-fetched data +(Step 0.7). + Group inline comments into threads using the `in_reply_to_id` field. Each top-level comment (no `in_reply_to_id`) starts a thread; replies are grouped under their parent. @@ -241,12 +385,10 @@ Log both lists so the run output shows the full enumeration result, making it auditable that all items were considered. > **Why this matters:** Without mandatory enumeration, a re-run can check only for -> existing classification replies and conclude "nothing to do" — completely missing -> new items that arrived after the previous run. This caused a real failure where a -> bot review comment posted after `/implement-task` pushed a fix commit was missed -> by the subsequent `/verify-pr` re-run. The same gap allowed review body -> suggestions (e.g., from sourcery-ai) to be silently skipped because only inline -> comments were enumerated. +> existing classification replies, conclude "nothing to do", and miss new items that +> arrived after the previous run — a real failure mode (a bot review comment posted +> after an `/implement-task` fix commit was missed on re-run; review-body suggestions +> from sourcery-ai were skipped because only inline comments were enumerated). ### Step 4a.1 – Detect Eval Result Reviews @@ -276,9 +418,9 @@ inform classification decisions. with the absolute path from the Registry (e.g., `/CONVENTIONS.md`). If present, read its contents. This provides explicit, documented project conventions. -2. **Codebase convention cache:** This step does not perform exhaustive codebase analysis - yet — that happens in the Style/Conventions sub-agent. The goal here is only to load - CONVENTIONS.md once for reuse across comment classification and sub-agent dispatch. +2. **Scope:** load CONVENTIONS.md once here for reuse across comment classification and + sub-agent dispatch; exhaustive codebase analysis happens later in the + Style/Conventions sub-agent. If `CONVENTIONS.md` does not exist, proceed normally — the Style/Conventions sub-agent will check for implicit conventions demonstrated by codebase usage patterns. @@ -308,10 +450,18 @@ change requests before sub-task creation. Dispatch four domain sub-agents in parallel for comprehensive PR analysis. Each sub-agent performs focused checks and returns structured findings. The orchestrator constructs dispatch envelopes following the structure defined in -`plugins/sdlc-workflow/skills/verify-pr/dispatch-template.md`. +`${CLAUDE_PLUGIN_ROOT}/skills/verify-pr/dispatch-template.md`. ### Step 5a – Gather Dispatch Inputs +**Sandbox mode:** do not run the `gh pr diff`/`gh pr view` commands below — the full +diff, diffstat, and commits are pre-fetched as `github.diff`, `github.stat`, and +`github.commits` (Step 0.7). Use those values as the corresponding dispatch inputs. +The sandbox has no `gh` CLI; do NOT substitute a live repository read (e.g. +`git log`) for any pre-fetched value — for commit traceability in particular, +`git log --oneline`/`--format=%s` emit subject lines only and will miss a Jira +ID in a commit body/trailer, producing a false FAIL. + Collect all inputs needed for sub-agent dispatch envelopes: 1. **PR diff (full)** — for Security, Correctness, and Style/Conventions sub-agents: @@ -325,10 +475,16 @@ Collect all inputs needed for sub-agent dispatch envelopes: gh pr diff --stat -R ``` -3. **PR commits** — for Intent Alignment sub-agent: +3. **PR commits** — for Intent Alignment sub-agent. Pass each commit's `oid`, its + **full** `messageHeadline` and `messageBody` (do NOT truncate the body — a Jira + trailer such as `Implements PROJ-231` typically sits at the very end), and, in + sandbox mode, the runner-computed `references_task_id` boolean. In interactive + mode fetch with: ``` gh pr view --json commits --jq '.commits[] | {oid: .oid, messageHeadline: .messageHeadline, messageBody: .messageBody}' -R ``` + In sandbox mode use `github.commits` as-is (each item already carries + `references_task_id`); never re-slice or subject-only-summarize the bodies. 4. **Task specification sections** — extracted from the Jira task description parsed in Step 1: @@ -361,13 +517,13 @@ Collect all inputs needed for sub-agent dispatch envelopes: Read the following files to construct dispatch prompts: -1. **Dispatch template:** `plugins/sdlc-workflow/skills/verify-pr/dispatch-template.md` -2. **Finding template:** `plugins/sdlc-workflow/skills/verify-pr/finding-template.md` +1. **Dispatch template:** `${CLAUDE_PLUGIN_ROOT}/skills/verify-pr/dispatch-template.md` +2. **Finding template:** `${CLAUDE_PLUGIN_ROOT}/skills/verify-pr/finding-template.md` 3. **Sub-agent skill files:** - - `plugins/sdlc-workflow/skills/verify-pr/intent-alignment.md` - - `plugins/sdlc-workflow/skills/verify-pr/security.md` - - `plugins/sdlc-workflow/skills/verify-pr/correctness.md` - - `plugins/sdlc-workflow/skills/verify-pr/style-conventions.md` + - `${CLAUDE_PLUGIN_ROOT}/skills/verify-pr/intent-alignment.md` + - `${CLAUDE_PLUGIN_ROOT}/skills/verify-pr/security.md` + - `${CLAUDE_PLUGIN_ROOT}/skills/verify-pr/correctness.md` + - `${CLAUDE_PLUGIN_ROOT}/skills/verify-pr/style-conventions.md` ### Step 5c – Construct and Dispatch @@ -409,7 +565,7 @@ execute side effects (sub-task creation, PR comment replies). ### Step 6a – Collect Sub-Agent Results Parse the structured findings returned by each sub-agent using the format defined -in `plugins/sdlc-workflow/skills/verify-pr/finding-template.md`: +in `${CLAUDE_PLUGIN_ROOT}/skills/verify-pr/finding-template.md`: 1. **Extract verdicts** from each sub-agent's Verdicts table. Map sub-agent check names to report rows: @@ -502,6 +658,35 @@ After creating each sub-task, create a "Blocks" issue link from the sub-task to jira.create_issue_link(type="Blocks", inwardIssue=, outwardIssue=) +**Sandbox mode:** Instead of calling `jira.create_issue` and `jira.create_issue_link`, +append a `create_subtask` action and a `create_link` action. Use this pattern for **all** +sub-task categories below (review feedback, CI failure, eval failure): + +```json +{ + "type": "create_subtask", + "ref": "subtask-N", + "parent": "", + "summary": "", + "labels": ["ai-generated-jira", ""], + "description_adf": +} +``` + +```json +{ + "type": "create_link", + "link_type": "Blocks", + "inward": "{{subtask-N.key}}", + "outward": "" +} +``` + +Use incrementing refs (`subtask-1`, `subtask-2`, …). The `{{subtask-N.key}}` (and, in +replies, `{{subtask-N.url}}`) placeholders are resolved by the `post_script` once the +sub-task is created. Set `` to `review-feedback` for review-feedback and +CI-failure sub-tasks, or `eval-failure` for eval-failure sub-tasks. + #### CI failure sub-tasks Process `create-sub-task` actions from the Correctness sub-agent. For each action: @@ -509,7 +694,9 @@ Process `create-sub-task` actions from the Correctness sub-agent. For each actio 1. **Idempotency check:** check the parent task's existing sub-tasks (issue links) for sub-tasks with labels `["ai-generated-jira", "review-feedback"]` whose descriptions reference the same CI check name or failure. If a matching sub-task - already exists, skip creation for that failure. + already exists, skip creation for that failure. **Sandbox mode:** read the + candidate sub-tasks from `idempotency.related_issues` (Step 0.7) — match on + `labels` and `description` — instead of a live Jira read. 2. **Create sub-task:** create a Jira sub-task using the action's Title, Relevant files, and Root cause fields: @@ -529,6 +716,9 @@ Process `create-sub-task` actions from the Correctness sub-agent. For each actio 3. **Create issue link:** jira.create_issue_link(type="Blocks", inwardIssue=, outwardIssue=) +**Sandbox mode:** use the `create_subtask` + `create_link` pattern from the review +feedback sandbox-mode block above, with labels `["ai-generated-jira", "review-feedback"]`. + #### Eval failure sub-tasks Process eval assertion failures from the Style/Conventions sub-agent's Check 5 @@ -553,7 +743,9 @@ and sub-task creation below. 2. **Idempotency check:** Check the parent task's existing sub-tasks (issue links) for sub-tasks with labels `["ai-generated-jira", "eval-failure"]` whose summaries reference the same eval ID. If a matching sub-task already exists, skip creation - for that eval. + for that eval. **Sandbox mode:** read the candidate sub-tasks from + `idempotency.related_issues` (Step 0.7) — match on `labels` and `summary` — + instead of a live Jira read. 3. **Create sub-task:** For each failing eval, create a Jira sub-task: @@ -573,6 +765,9 @@ and sub-task creation below. 4. **Create issue link:** jira.create_issue_link(type="Blocks", inwardIssue=, outwardIssue=) +**Sandbox mode:** use the `create_subtask` + `create_link` pattern from the review +feedback sandbox-mode block above, with labels `["ai-generated-jira", "eval-failure"]`. + ### Step 6e – Reply to Review Comments Reply to **every** classified review comment thread with the classification label and @@ -585,74 +780,75 @@ gh api repos//pulls//comments//replies -f bod ``` **For suggestions upgraded to code change requests via convention check (Step 6b):** - -Include the convention evidence in the reply so the upgrade reasoning is transparent: +Include the convention evidence so the upgrade reasoning is transparent (e.g. "…this +matches project convention: 17 migrations use Index::create for FK columns; CONVENTIONS.md +§Indexes documents this pattern. Sub-task […] created…"): ``` gh api repos//pulls//comments//replies -f body="[sdlc-workflow/verify-pr] Classified as **code change request** (upgraded from suggestion) — this matches project convention: . Sub-task []() created to address this feedback." ``` -Example: `"[sdlc-workflow/verify-pr] Classified as **code change request** (upgraded from suggestion) — this matches project convention: 17 migrations use Index::create for FK columns; CONVENTIONS.md §Indexes documents this pattern. Sub-task [PROJ-456](https://redhat.atlassian.net/browse/PROJ-456) created to address this feedback."` - -**For all other classifications (suggestion, question, nit):** - -Reply with a brief explanation of the classification and why no sub-task was created: +**For all other classifications (suggestion, question, nit):** reply with a brief +explanation of the classification and why no sub-task was created (e.g. suggestion → "not +documented in CONVENTIONS.md and no established codebase pattern"; question → "asks for +clarification; no code change needed"; nit → "minor style, does not affect correctness"): ``` gh api repos//pulls//comments//replies -f body="[sdlc-workflow/verify-pr] Classified as **** — . No sub-task created." ``` -Example replies: -- `"[sdlc-workflow/verify-pr] Classified as **suggestion** — this proposes an alternative approach that is not documented in CONVENTIONS.md and has no established codebase pattern. No sub-task created."` -- `"[sdlc-workflow/verify-pr] Classified as **question** — this asks for clarification; no code change needed. No sub-task created."` -- `"[sdlc-workflow/verify-pr] Classified as **nit** — minor style feedback that does not affect correctness. No sub-task created."` - #### Review body items Review body items (identified by `review-body-*` synthetic IDs) lack a `comment_id` -and cannot receive threaded replies. Instead, post a **standalone PR comment** using -the issues API: +and cannot receive threaded replies. Post a **standalone PR comment** via the issues +API instead — use the **same classification body as the matching inline case above**, +but prefixed with `Re: @ review — ` right after the `[sdlc-workflow/verify-pr]` +tag (code change request → sub-task link; upgraded suggestion → convention evidence + +sub-task link; suggestion/question/nit → reasoning + "No sub-task created."): ``` -gh api repos//issues//comments -f body="[sdlc-workflow/verify-pr] Re: @ review — Classified as **** — . " +gh api repos//issues//comments -f body="[sdlc-workflow/verify-pr] Re: @ review — Classified as **** — " ``` -**For code change requests that resulted in a sub-task:** +When a review body contains multiple classified suggestions (sub-identifiers), post +a single standalone comment that lists all classifications together rather than one +comment per sub-identifier. -``` -gh api repos//issues//comments -f body="[sdlc-workflow/verify-pr] Re: @ review — Classified as **code change request** — sub-task []() created to address this feedback." -``` +**Sandbox mode:** Instead of calling `gh api`, append one action per reply. -**For suggestions upgraded via convention check (Step 6b):** +For **inline comment threads** (they have a `comment_id`), use `post_pr_reply`: -``` -gh api repos//issues//comments -f body="[sdlc-workflow/verify-pr] Re: @ review — Classified as **code change request** (upgraded from suggestion) — this matches project convention: . Sub-task []() created to address this feedback." +```json +{ + "type": "post_pr_reply", + "repo": "", + "pr_number": , + "comment_id": , + "body": "" +} ``` -**For all other classifications:** +For **review body items** (no `comment_id`), use `post_pr_comment`: +```json +{ + "type": "post_pr_comment", + "repo": "", + "pr_number": , + "body": "" +} ``` -gh api repos//issues//comments -f body="[sdlc-workflow/verify-pr] Re: @ review — Classified as **** — . No sub-task created." -``` - -When a review body contains multiple classified suggestions (sub-identifiers), post -a single standalone comment that lists all classifications together rather than one -comment per sub-identifier. ### Step 6f – Idempotency Guarantees -Idempotency is enforced **within** Step 4a's mandatory enumeration, not as a -separate gate before Steps 6d/6e. The enumeration in Step 4a partitions all -classifiable items (inline comment threads and review body items) into unclassified -and already-classified lists — only unclassified items proceed to Steps 4b–4c for -classification and then to Steps 6c–6e for side effects. This design ensures that: - -- Every classifiable item is always discovered (enumeration cannot be bypassed). -- Already-processed items are filtered out per-item (no duplicate replies or - sub-tasks). -- New items arriving between runs (e.g., from a bot re-analyzing code after a - fix commit, or a reviewer adding a new review) are always detected because the - enumeration is unconditional. +Idempotency is enforced **within** Step 4a's mandatory enumeration, not as a separate +gate before Steps 6d/6e: the enumeration partitions all classifiable items (inline +comment threads and review body items) into unclassified and already-classified lists — +only unclassified items proceed to classification (4b–4c) and side effects (6c–6e). +Every item is always discovered (enumeration is unconditional and cannot be bypassed), +already-processed items are filtered per-item (no duplicate replies/sub-tasks), and new +items arriving between runs (a bot re-analyzing after a fix commit, a new review) are +always detected. Do **not** treat this as a top-level gate that can skip enumeration. **Idempotency detection per item type:** - **Inline comment threads:** a reply containing `"[sdlc-workflow/verify-pr] Classified as"` @@ -665,10 +861,8 @@ check the parent task's issue links for existing sub-tasks whose descriptions reference the same review comment or review body. If a matching sub-task already exists, skip creation. This guards against edge cases where a classification reply was not posted (e.g., due to a network error) but the sub-task was created. - -Do **not** interpret this step as a top-level decision that can skip item -enumeration. The enumeration in Step 4a is always mandatory; this step only -documents the idempotency mechanisms embedded within that enumeration. +**Sandbox mode:** read these existing sub-tasks from `idempotency.related_issues` +(Step 0.7) — inspect each entry's `description` — instead of a live Jira read. ### Step 6g – Record Result @@ -711,22 +905,10 @@ The sub-agent receives these inputs: 3. **Review comments** — the code change requests that triggered sub-tasks in Step 6d 4. **Relevant code** — the files on the PR branch related to each flagged defect 5. **Project CONVENTIONS.md** — if it exists in the repository root -6. **Aggregated domain findings** — findings from all four domain sub-agents, - organized by source with clear attribution: - - ``` - ### From Intent Alignment - - - ### From Security - - - ### From Correctness - - - ### From Style/Conventions -