Skip to content

test(triage-security): exercise conditional Fullsend scenarios - #321

Merged
mrizzi merged 2 commits into
RHEcosystemAppEng:verify-pr-fullsendfrom
mrizzi:TC-6676
Oct 1, 2026
Merged

mrizzi merged 2 commits into
RHEcosystemAppEng:verify-pr-fullsendfrom
mrizzi:TC-6676

Conversation

@mrizzi

@mrizzi mrizzi commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Implements TC-6676.

  • Move ten conditional Fullsend assertions into separate scenarios with synthetic trusted input fixtures.
  • Preserve the original 36 interactive prompts and all existing assertion text.
  • Add fixture validation and mocked executor coverage, including interrupted retry and digest repair.
  • Export runner audit evidence and the result byte-for-byte into grader outputs, preserving the isolated sandbox write contract.

Validation

At repair head 49ddc78cd3ccafa6a9767443bb28881c7df2acb7:

  • Local script unit tests: 294 passed, including ten transport-only handoff regressions. These do not execute the skill or establish schema/runtime acceptance.
  • Skillsaw: zero errors, seven existing warnings. Plugin validation and git diff --check passed.
  • Original 36 cases are byte-preserved; every assertion is unchanged. Changes are limited to the ten Fullsend prompts and their test module.
  • Actual skill eval validation is pending. The earlier generated outputs and predetermined grades are withdrawn as eval evidence; they establish neither passes nor runtime regressions.

Repair awaiting review

The grader-evidence finding is addressed by a runner-owned handoff: outputs/invocation.json and outputs/agent-result.json are flat grader inputs; the separate FULLSEND_OUTPUT_DIR keeps only the skill result. No production or global run-evals code changed.

Next step

After the repair is reviewed, the human reviewer decides whether to merge into verify-pr-fullsend. Hosted eval feedback then comes through PR #299. Integration does not itself establish task completion.

Summary by Sourcery

Expand triage-security Fullsend evaluation coverage with independently validated conditional scenarios and retry behavior.

Enhancements:

  • Split ten conditional Fullsend contracts into independent scenarios with scenario-specific trusted input fixtures while preserving the existing interactive assertions.
  • Add coverage for runner evidence handoff, fixture validation, interrupted retries, digest repair, and idempotent link resolution.

Tests:

  • Add validation ensuring each conditional evaluation has executable inputs and distinct evidence supporting its expected decisions.

Implements TC-6676

Assisted-by: Claude Code
@sourcery-ai

sourcery-ai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Reviewer's Guide

This repair makes all ten conditional triage-security Fullsend scenarios independently executable with validated, scenario-specific trusted inputs and adds regression coverage for invocation coherence and interrupted-retry idempotency, without changing production schemas, skills, or trust boundaries. Review the new eval wiring and large fixture contents closely; the selected interactive eval suite remains partially failing in unrelated existing policy/output contracts.

File-Level Changes

Change Details Files
Converted ten conditional Fullsend assertions into independently executable, trusted-input eval cases while preserving the original interactive coverage.
  • Added Fullsend invocations for cases 37–46 with matching mounted JSON inputs and subject-specific prompts.
  • Added synthetic, schema-validated evidence covering authorization, impact, duplicate, overlap, RPM, enrichment, and external-evidence scenarios.
  • Added coherence tests to ensure every conditional case has exactly one executable Fullsend input and correct identity/authorization metadata.
evals/triage-security/evals.json
evals/triage-security/files/fullsend-eval-1-trusted-input.json
evals/triage-security/files/fullsend-eval-2-trusted-input.json
evals/triage-security/files/fullsend-eval-3-trusted-input.json
evals/triage-security/files/fullsend-eval-4-trusted-input.json
evals/triage-security/files/fullsend-eval-5-trusted-input.json
evals/triage-security/files/fullsend-eval-8-trusted-input.json
evals/triage-security/files/fullsend-eval-9-trusted-input.json
evals/triage-security/files/fullsend-eval-11-trusted-input.json
evals/triage-security/files/fullsend-eval-12-trusted-input.json
evals/triage-security/files/fullsend-eval-18-trusted-input.json
plugins/sdlc-workflow/scripts/test_triage_security_fullsend.py
Added focused coverage for an interrupted retry that repairs an incomplete remediation before resolving a new dependency link and is idempotent on subsequent execution.
  • Distinguished the partial retry fixture from the retained fully triaged interactive case using two existing tasks with only the upstream digest missing.
  • Verified the production executor performs digest repair, link resolution, and no task creation in the expected order.
  • Verified a refreshed retry performs zero writes and retains task identity and action markers.
evals/triage-security/files/fullsend-eval-18-trusted-input.json
plugins/sdlc-workflow/scripts/test_triage_security_fullsend.py

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've reviewed your changes and they look great!

Sourcery assessment

Approved.


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

@mrizzi

mrizzi commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author

Verification status — TC-6676

Evidence-handoff fix checked at 49ddc78cd3ccafa6a9767443bb28881c7df2acb7.

The grader evidence mismatch is fixed. Actual skill eval validation remains pending.

Established checks

  • CI script unit tests on Python 3.11–3.14, skill lint and plugin validation passed at the new head.
  • The follow-up changes only ten Fullsend prompts and their regression test module. The original 36 scenarios and all assertion text are unchanged.
  • Ten new transport regression cases check byte-for-byte evidence delivery and an unchanged sandbox result directory. These tests establish file transport, not skill execution.

Evidence handoff

The eval runner now copies invocation.json and agent-result.json into the grader's assigned outputs/ directory after skill execution. The separate FULLSEND_OUTPUT_DIR remains limited to the skill result. No production or global run-evals code changed.

Evidence correction

The earlier generated outputs and predetermined grades are withdrawn as eval evidence. They establish neither passes nor runtime regressions. No replacement skill eval result is claimed.

Next step

Review the updated repair. A human merge into verify-pr-fullsend allows PR #299 to provide hosted eval feedback. TC-6676 remains unverified until actual results support its acceptance criteria.


This updates the code-finding status after inspecting the follow-up diff and CI; it is not a new skill verification run.

Implements TC-6676

Assisted-by: Claude Code
@mrizzi
mrizzi merged commit 011efcc into RHEcosystemAppEng:verify-pr-fullsend Oct 1, 2026
10 checks passed
@mrizzi
mrizzi deleted the TC-6676 branch October 1, 2026 09:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant