Skip to content

docs: correct v7.0.8 closeout retry-honesty and add real installed Claude evidence - #225

Merged
sumitake merged 1 commit into
mainfrom
dev/claude/v708-closeout-correction
Sep 24, 2026
Merged

sumitake merged 1 commit into
mainfrom
dev/claude/v708-closeout-correction

Conversation

@sumitake

Copy link
Copy Markdown
Owner

Summary

Follow-up correction to PR #224. A downstream review on the workspace companion PR (agent-collab-workspace#3190) correctly flagged that #224's evidence overclaimed what the retried Claude/Grok installed canaries established, and that installed Claude qualification relied on staged evidence the runbook does not treat as a substitute. This PR carries the same fix to the public docs.

Changes

  • docs/architecture/status-and-evidence.md: the two installed canaries whose first attempts timed out with zero output under concurrent host load are now recorded as uncertain (neither qualified nor failed) instead of "retried... no consumed provider work was replayed." The later isolated requests are recorded as their own separate evidence.
  • docs/architecture/status-and-evidence.md + docs/architecture/claude-participation.md: record a real, separately-run, explicitly-targeted (explicit_target: "claude") installed context.documents.intent canary (selected_logical_agent: claude, native_model_identity.model: claude-opus-5-5, execution_status: ok, confirmed cleanup), replacing the prior claim that staged qualification alone stands in for installed evidence.

Verification

  • python3 scripts/check_release_consistency.py — OK
  • python3 scripts/check-public-export-safety.py --active-tree — SAFE
  • python3 scripts/secret_scan.py — clean
  • git diff --check — clean

Compliance trace

author: claude
standing_directives: docs/public-governance.md Tier 1; "Batch PR review remediation" (operator-set 2026-08-13)
tier: 1
cross_check: N/A -- documentation-only correction, no executable/policy/security/packaging effect; corrects prior evidence to be more conservative/accurate, and adds new real evidence (a fresh installed canary run) rather than removing any
post_condition: fetch and positively read back main after merge
mcp_coverage_gap: NONE
contributor_rights: OWNER-AUTHORED
operator_reserved: no

🤖 Generated with Claude Code

…aude evidence

Follow-up to PR #224. A downstream review of the workspace companion (#3190)
correctly flagged that the retried Claude/Grok installed canaries were
described as 'confirming' or resolving the earlier silent-timeout attempts,
which the runtime's own no-replay contract does not support ('a new request
ID alone does not make repeating consumed work permissible'). Also ran and
recorded a genuine, clean, explicitly-targeted (explicit_target: claude)
installed context.documents.intent canary, replacing the prior 'staged
qualification stands in for installed' framing with actual installed-unit
evidence (selected_logical_agent: claude, native_model_identity.model:
claude-opus-5-5, execution_status: ok, confirmed cleanup).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 24, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 6f0e2380-4c07-4e9b-abd4-7fcbb9d40c70


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-24T09:05:23.848584Z 90cb55f PR opened
🔒 Security Review ✅ Completed 2026-09-24T09:06:15.162332Z 90cb55f PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@sumitake
sumitake merged commit 6b7d9bb into main Sep 24, 2026
18 of 19 checks passed
@sumitake
sumitake deleted the dev/claude/v708-closeout-correction branch September 24, 2026 09:04

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 90cb55f884

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

| Before publication, the full staged qualification matrix (12 actions, 5 carriers, maximum effort) returned `qualified: true` with every case content-correct and cleaned up, including a genuine live Claude `context.documents.intent` canary. After install, one consumed native canary per host returned content-correct results. | bounded native execution verified for the named cases | Each case used one new request, with source/package bytes unchanged, zero replay and confirmed native and fixture cleanup. Two of the four installed canaries needed one retry after a first attempt returned no output under concurrent host load; the retried attempts succeeded cleanly and no consumed provider work was replayed. Successful staged evidence is not a substitute for the installed canary. |
| Claude package installation succeeded; managed inference (`context.documents.intent`, Claude's sole admitted action) was qualified in the staged matrix above. | installed; inference qualified in staging | The installed unit was not separately re-canaried for Claude specifically; the byte-identical release plus the staged qualification stand in for it. No other Claude action is admitted. |
| Before publication, the full staged qualification matrix (12 actions, 5 carriers, maximum effort) returned `qualified: true` with every case content-correct and cleaned up, including a genuine live Claude `context.documents.intent` canary. After install, one qualified consumed native canary per host returned content-correct results. | bounded native execution verified for the named cases | Two of the four installed canaries' first attempts, run under concurrent host load, timed out with zero output; neither attempt's provider-dispatch status is known, and neither is counted as qualified or failed. A second, later, isolated request for each returned content-correct results with confirmed cleanup; that second request is the qualification evidence recorded, not a resolution of the first attempt's ambiguity. A separately-run, explicitly-targeted installed Claude `context.documents.intent` canary also succeeded. Successful staged evidence is not a substitute for the installed canary. |
| Claude package installation succeeded; managed inference (`context.documents.intent`, Claude's sole admitted action) was qualified both in the staged matrix and by a separately-run, explicitly-targeted installed canary. | installed; inference qualified in staging and installed | The installed canary returned `selected_logical_agent: claude`, `native_model_identity.model: claude-opus-5-5`, `execution_status: ok`, confirmed cleanup, and a correct fixture-marker answer. No other Claude action is admitted. |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Update the Claude closeout determination

Once this row records a separately run installed Claude canary, the v7.0.8 closeout determinations table below becomes internally inconsistent: its Claude-participation entry still says the new positive qualification evidence came only from the staged matrix. Update that determination to mention both staged and installed qualification so an audit of the closeout does not produce two different evidence accounts.

AGENTS.md reference: AGENTS.md:L44-L48

Useful? React with 👍 / 👎.

sumitake added a commit that referenced this pull request Sep 24, 2026
…retry-confirmed (#226)

The prior wording (PR #225) still treated a later, separately-issued
request's success as "the qualification evidence recorded" for the
Claude/Grok hosts whose first review.repository attempt timed out with
zero output under concurrent load. Per the runtime's own no-replay
contract, a distinct qualification is authorized only after a verified
correction, which was never performed for these two ambiguous timeouts.
Record installed review.repository as explicitly unmet for those two
hosts this cycle, matching the established unmet-qualification pattern
used elsewhere in this handbook. The separately-run, explicitly-targeted
Claude context.documents.intent canary is unaffected and stays recorded
as its own clean evidence.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant