docs: correct v7.0.8 closeout retry-honesty and add real installed Claude evidence - #225
Conversation
…aude evidence Follow-up to PR #224. A downstream review of the workspace companion (#3190) correctly flagged that the retried Claude/Grok installed canaries were described as 'confirming' or resolving the earlier silent-timeout attempts, which the runtime's own no-replay contract does not support ('a new request ID alone does not make repeating consumed work permissible'). Also ran and recorded a genuine, clean, explicitly-targeted (explicit_target: claude) installed context.documents.intent canary, replacing the prior 'staged qualification stands in for installed' framing with actual installed-unit evidence (selected_logical_agent: claude, native_model_identity.model: claude-opus-5-5, execution_status: ok, confirmed cleanup). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 90cb55f884
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| | Before publication, the full staged qualification matrix (12 actions, 5 carriers, maximum effort) returned `qualified: true` with every case content-correct and cleaned up, including a genuine live Claude `context.documents.intent` canary. After install, one consumed native canary per host returned content-correct results. | bounded native execution verified for the named cases | Each case used one new request, with source/package bytes unchanged, zero replay and confirmed native and fixture cleanup. Two of the four installed canaries needed one retry after a first attempt returned no output under concurrent host load; the retried attempts succeeded cleanly and no consumed provider work was replayed. Successful staged evidence is not a substitute for the installed canary. | | ||
| | Claude package installation succeeded; managed inference (`context.documents.intent`, Claude's sole admitted action) was qualified in the staged matrix above. | installed; inference qualified in staging | The installed unit was not separately re-canaried for Claude specifically; the byte-identical release plus the staged qualification stand in for it. No other Claude action is admitted. | | ||
| | Before publication, the full staged qualification matrix (12 actions, 5 carriers, maximum effort) returned `qualified: true` with every case content-correct and cleaned up, including a genuine live Claude `context.documents.intent` canary. After install, one qualified consumed native canary per host returned content-correct results. | bounded native execution verified for the named cases | Two of the four installed canaries' first attempts, run under concurrent host load, timed out with zero output; neither attempt's provider-dispatch status is known, and neither is counted as qualified or failed. A second, later, isolated request for each returned content-correct results with confirmed cleanup; that second request is the qualification evidence recorded, not a resolution of the first attempt's ambiguity. A separately-run, explicitly-targeted installed Claude `context.documents.intent` canary also succeeded. Successful staged evidence is not a substitute for the installed canary. | | ||
| | Claude package installation succeeded; managed inference (`context.documents.intent`, Claude's sole admitted action) was qualified both in the staged matrix and by a separately-run, explicitly-targeted installed canary. | installed; inference qualified in staging and installed | The installed canary returned `selected_logical_agent: claude`, `native_model_identity.model: claude-opus-5-5`, `execution_status: ok`, confirmed cleanup, and a correct fixture-marker answer. No other Claude action is admitted. | |
There was a problem hiding this comment.
Update the Claude closeout determination
Once this row records a separately run installed Claude canary, the v7.0.8 closeout determinations table below becomes internally inconsistent: its Claude-participation entry still says the new positive qualification evidence came only from the staged matrix. Update that determination to mention both staged and installed qualification so an audit of the closeout does not produce two different evidence accounts.
AGENTS.md reference: AGENTS.md:L44-L48
Useful? React with 👍 / 👎.
…retry-confirmed (#226) The prior wording (PR #225) still treated a later, separately-issued request's success as "the qualification evidence recorded" for the Claude/Grok hosts whose first review.repository attempt timed out with zero output under concurrent load. Per the runtime's own no-replay contract, a distinct qualification is authorized only after a verified correction, which was never performed for these two ambiguous timeouts. Record installed review.repository as explicitly unmet for those two hosts this cycle, matching the established unmet-qualification pattern used elsewhere in this handbook. The separately-run, explicitly-targeted Claude context.documents.intent canary is unaffected and stays recorded as its own clean evidence. Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Summary
Follow-up correction to PR #224. A downstream review on the workspace companion PR (agent-collab-workspace#3190) correctly flagged that #224's evidence overclaimed what the retried Claude/Grok installed canaries established, and that installed Claude qualification relied on staged evidence the runbook does not treat as a substitute. This PR carries the same fix to the public docs.
Changes
docs/architecture/status-and-evidence.md: the two installed canaries whose first attempts timed out with zero output under concurrent host load are now recorded as uncertain (neither qualified nor failed) instead of "retried... no consumed provider work was replayed." The later isolated requests are recorded as their own separate evidence.docs/architecture/status-and-evidence.md+docs/architecture/claude-participation.md: record a real, separately-run, explicitly-targeted (explicit_target: "claude") installedcontext.documents.intentcanary (selected_logical_agent: claude,native_model_identity.model: claude-opus-5-5,execution_status: ok, confirmed cleanup), replacing the prior claim that staged qualification alone stands in for installed evidence.Verification
python3 scripts/check_release_consistency.py— OKpython3 scripts/check-public-export-safety.py --active-tree— SAFEpython3 scripts/secret_scan.py— cleangit diff --check— cleanCompliance trace
author: claude
standing_directives: docs/public-governance.md Tier 1; "Batch PR review remediation" (operator-set 2026-08-13)
tier: 1
cross_check: N/A -- documentation-only correction, no executable/policy/security/packaging effect; corrects prior evidence to be more conservative/accurate, and adds new real evidence (a fresh installed canary run) rather than removing any
post_condition: fetch and positively read back main after merge
mcp_coverage_gap: NONE
contributor_rights: OWNER-AUTHORED
operator_reserved: no
🤖 Generated with Claude Code