diff --git a/CHANGELOG.md b/CHANGELOG.md index f586149..15e941d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,53 @@ All notable changes to SkipHow 2.x and later appear in this file. Earlier release notes remain available on [GitHub Releases](https://github.com/mzored/SkipHow/releases). +## 4.8.4 (2026-09-16) + +A repeat of the same kind of failure is now evidence about a class. The always-loaded kernel says that an owner +reporting a symptom again, or a slow gate failing on one more stale expectation, calls for finding the other +members of that class with the cheapest check that covers all of them and fixing them together before returning +to the loop that found them. The bug workflow's sibling inspection now also covers the other places the owner +would see the same symptom, since a repeated symptom can have a second cause. No candidate-count limit, +mandatory local suite, sweep procedure, or polling rule was added. + +The skill description carries one conditional sentence: a session this skill was governing reopens the skill +after the host compacts its context. It is not a hook and executes nothing. It answers the reopen condition +recorded in the 3.0.0 reminder decision, which fired: in a private installed Codex session on 4.8.2 no text from any +package file survived any of fifteen compactions, only the skill descriptions the host keeps in its own skill list +did, and the kernel was never re-read across a delivery loop of almost four hours. Whether the sentence causes a reload +is `UNVERIFIED`. + +The motivating observations, prior-art reading and refused alternatives are recorded in +[evidence](docs/evidence.md#484-repeat-failures-and-continuity-after-compaction) and +[decisions](docs/decisions.md#the-484-class-evidence-and-continuity-correction). No private session content is +published and no paid behavioral comparison was run. This is a compatible wording patch. + +The full local package gate passed all 389 tests under the pinned dependencies in an isolated worktree, and +`git diff --check` passed. Both host schema validators passed on the candidate tree. Clean installation and model behavior were not retested for this +unchanged package structure. + +An isolated, read-only Codex review found one qualifying documentation defect: the release rationale said none +of the package's text survives Codex compaction while also calling the surviving description package-owned text. +The three sentences now distinguish the discarded file bodies from the descriptions the host keeps in its own +skill list. The reviewer confirmed that the kernel sentence contradicts neither the verification nor the diagnosis +reference, that the description sentence selects nothing for an ungoverned session and grants no authority, that +the bug workflow clause is consistent with the kernel, that all referenced anchors resolve, and that no active +version pin was missed. No findings were refused. The review transcript confirmed separate operating-system and +host homes with no personal instruction or installed-plugin contamination markers. Authentication referenced +the existing credential file; the owned review workspace was removed. + +| Capability | Local candidate evidence | +| --- | --- | +| Deterministic package gate | PASS, full local command and 389 tests | +| Codex schema validation | PASS | +| Claude schema validation | PASS | +| Clean Codex install | UNVERIFIED | +| Clean Claude install | UNVERIFIED | +| Explicit invocation | UNVERIFIED | +| Implicit activation | UNVERIFIED | +| Continuity | UNVERIFIED | +| Behavioral suite | UNVERIFIED, [evidence](docs/evidence.md) | + ## 4.8.3 (2026-09-13) Delegate briefs now account for context the host actually supplies instead of asserting that delegates diff --git a/VERSION b/VERSION index f99c658..57c4b30 100644 --- a/VERSION +++ b/VERSION @@ -1 +1 @@ -4.8.3 +4.8.4 diff --git a/docs/decisions.md b/docs/decisions.md index 146c72a..d3af5f1 100644 --- a/docs/decisions.md +++ b/docs/decisions.md @@ -4,10 +4,11 @@ This page records the choices that still matter when SkipHow changes. Read it be ## Current decisions -The live decisions, their premises, and what would reopen each. "Evidence" says what stands behind the decision today: `Observed` means a retained run showed it on the package that carried it, `Contract` means the shipped text encodes it and no run has tested that text, `Deterministic` means a check proves it on every run. Last reviewed 2026-09-13 against the 4.8.3 candidate. The [owner-outcome contract](outcome-contract.md) governs implementation choices. +The live decisions, their premises, and what would reopen each. "Evidence" says what stands behind the decision today: `Observed` means a retained run showed it on the package that carried it, `Contract` means the shipped text encodes it and no run has tested that text, `Deterministic` means a check proves it on every run. Last reviewed 2026-09-16 against the 4.8.4 candidate. The [owner-outcome contract](outcome-contract.md) governs implementation choices. | Decision | Active rationale | Premises | Evidence | Reopens when | | --- | --- | --- | --- | --- | +| A repeat of the same kind of failure is evidence about a class, found with the cheapest check that covers the whole class and fixed together before the loop that found it runs again; a session the skill was governing reopens the kernel after the host compacts its context | Fixing each recurrence as a new instance sent the owner hunting the next one and re-ran a half-hour gate once per stale expectation; Codex keeps no text from any package file across compaction and only the skill description in the host's own skill list survives | The second appearance is readable in the owner's words or the gate's output; the description is the one package-owned surface that survives on both hosts; the sentence conditions on the skill already governing | `Contract` (4.8.4); one private Codex session on 4.8.2 motivates both; effect `UNVERIFIED` | A run sweeps unrelated surfaces on a first report, a governed Codex session still never reopens the kernel after compaction, or the description sentence selects the skill for an ungoverned session | | A prepared plan distinguishes broad records from executable assignments and bounds review as well as implementation | A tracker group can name an outcome without establishing a feasible next assignment; source size and shared writes affect execution | Existing bounded-slice contract and an explicit owner request for dispatch-ready planning; later boundaries may depend on earlier results | Clarified `Contract` in 4.8.2; causal attribution and behavioral improvement `UNVERIFIED` | Receipts show infeasible assignments surviving review or the clarification causing speculative decomposition or unnecessary small-task ceremony | | Assess material technical risks and accumulating costs encountered in ordinary work | Passing requested behavior and checks can leave a narrow acceptance-only reading of CTO supervision | Concerns need evidence and foreseeable product, development, or operational consequences; recognition does not grant broader repair | Owner-requested clarification in 4.8.1; cause and behavioral benefit `UNVERIFIED` | A controlled receipt shows missed material concerns, speculative audit expansion, or routine work made heavier without benefit | | One accountable CTO kernel with optional workflow skills and shared references | The owner explicitly wants recurring work patterns while retaining CTO supervision; workflows declare a relative link to the existing kernel instead of copying it | Owner request for bug, plan, longrun, deploy-ready, and fast-fixes; ordinary requests remain supported; linking is not host-enforced loading | `Deterministic` for package shape and declared kernel links; workflow loading and behavior `UNVERIFIED` | Receipts show missing kernel loading, misselection, contradictory contracts, or a host provides a stronger dependency mechanism | @@ -20,7 +21,7 @@ The live decisions, their premises, and what would reopen each. "Evidence" says | Broad project-state reconciliation reconstructs before it mutates; status, record correction, and cleanup have request-scoped effects; only a non-reconstructible outcome-to-workspace association becomes durable | Tracker state, branch names, merge status, and CI can each be stale or refer to the wrong outcome, revision, or destination; a parallel lifecycle registry would create another claim to reconcile | Live Git, reviews, destinations, and revision-specific validation establish every transition in the current cases; workspace ownership cannot always be reconstructed after a session ends | `Contract` and deterministic corpus validation (4.6.0); behavior `UNVERIFIED` | A receipt finds live evidence insufficient to recover a material transition, the ownership association insufficient for safe recovery, or the ordering adding ceremony without preventing a failure | | Delegates are read-only without verified distinct isolation; the root serializes writes; lead and delegate model and effort are independent choices minimizing expected total cost within owner constraints; adequate current routes survive comparable work, and settings use actual host controls | One shared checkout has one index and one branch; a delegate's own account of its isolation is not proof; host controls change independently of this package; a review's independence and framing matter more than its level | The active host supplies the controls a run can use; no portable writer-isolation or absolute-level interface exists across supported hosts, so unverifiable isolation takes the safe fallback | Failures `Observed` on 2.x (five lanes in one checkout; a worktree that reported success into the shared tree); the portable fallback is `Contract`, not run; writer lanes remain host-specific and must be verified | A portable capability interface appears, a host makes isolation verifiable and default, or paired runs settle the routing cost question | | Prefer host-native execution; admit thin bindings when a demonstrated gap justifies them | Duplicated runtime state increases cost; optional skills reuse one kernel without a service or custom workers | Required outcomes survive implementation changes; new components need host schema, safety, and compatibility checks | Current package shape `Deterministic`; adapter benefits require receipts | A controlled check shows a host binding is needed to preserve an outcome | -| Default ordinary-language governance references the installed skill from an owned reversible block in the trusted user instruction file the host actually reads; the skill enables, checks, and disables itself through a packaged host-aware helper and a setup playbook; explicit invocation is the fallback; no hook ships | Host documentation establishes instruction loading, not correct selection; a block in a shadowed or unread file configures nothing, so the agent resolves the target and reports configured, available, and loaded separately | Codex reads a non-empty `AGENTS.override.md` over `AGENTS.md` in `CODEX_HOME`; Claude reads `CLAUDE.md` and unconditional `rules/` under `CLAUDE_CONFIG_DIR`; isolated Codex login works while isolated Claude login remains unavailable | Helper lifecycle and host resolution `Deterministic`; Codex loading from the block shown once each on 4.1.0 (`AGENTS.md`) and 4.2.0 (`AGENTS.override.md`); Claude persistent loading `UNVERIFIED`; a staged comparison of the pointer block, a kernel-printing `SessionStart` hook, and an `@import` is specified in the 4.3.0 audit record and not run | A host changes its discovery order, a clean receipt shows missed or false activation, or the specified comparison shows another mechanism passing the whole enable, disable, upgrade, and uninstall lifecycle | +| Default ordinary-language governance references the installed skill from an owned reversible block in the trusted user instruction file the host actually reads; the skill enables, checks, and disables itself through a packaged host-aware helper and a setup playbook; explicit invocation is the fallback; no hook ships | Host documentation establishes instruction loading, not correct selection; a block in a shadowed or unread file configures nothing, so the agent resolves the target and reports configured, available, and loaded separately | Codex reads a non-empty `AGENTS.override.md` over `AGENTS.md` in `CODEX_HOME`; Claude reads `CLAUDE.md` and unconditional `rules/` under `CLAUDE_CONFIG_DIR`; isolated Codex login works while isolated Claude login remains unavailable | Helper lifecycle and host resolution `Deterministic`; Codex loading from the block shown once each on 4.1.0 (`AGENTS.md`) and 4.2.0 (`AGENTS.override.md`); Claude persistent loading `UNVERIFIED`; a staged comparison of the pointer block, a kernel-printing `SessionStart` hook, and an `@import` is specified in the 4.3.0 audit record and not run; the skill description carries a conditional continuity sentence since 4.8.4, effect `UNVERIFIED` | A host changes its discovery order, a clean receipt shows missed or false activation, or the specified comparison shows another mechanism passing the whole enable, disable, upgrade, and uninstall lifecycle | | A failed merge is recovered by consequence: disposable failures may stay for diagnosis, a shared target others depend on is contained or restored, production restoration keeps its grant | Leaving every failed merge in place made a broken shared branch the default while diagnosis ran | Containment of a covered non-production destination is within the established workflow; production is an effect that keeps its own grant | Source review, 4.2.0; behavior `UNVERIFIED` | A run restores a target it should have left for diagnosis, or leaves a shared target broken | | Evaluation oracles name outcomes, not implementations; a fixture is preflighted against a registered expected state before any model spend | The continuity oracle banned a thin host binding the contract permits, and the canonical large-programme prompt named a branch its fixture never created | A grader that encodes the expected behavior independently of the code under test can grade any retained end state | `Deterministic` preflight and grader tests, 4.2.0 | A scenario needs an outcome the registry or grader cannot express | | Deterministic checks protect package, security, release, and corpus semantics only; presentation and wording are lint or unchecked | A check that pins a sentence or a topology froze editorial choices without protecting anything a host depends on | Spec 11 classification, applied 2026-09-04 | `Deterministic` | A host starts depending on a detail now treated as editorial | @@ -31,6 +32,40 @@ The live decisions, their premises, and what would reopen each. "Evidence" says The sections below are the history behind those rows: what each release tried, measured, and rejected. They are non-normative. Where a section and the index disagree, the index is current and the section records how it got there. +## The 4.8.4 class-evidence and continuity correction + +Reopened by one private installed Codex session on 4.8.2, described in +[`docs/evidence.md`](evidence.md#484-repeat-failures-and-continuity-after-compaction). + +Two defects were readable in the shipped text. The kernel treated recurring workarounds, timeouts and churn as +engineering-system signals but said nothing about a repeat of the same failure: an owner reporting a symptom a +second time, or a slow gate failing once more on the same kind of stale expectation. The bug workflow scoped the +class of a defect to the rule that failed and the sibling paths that rule governs, which is the right scope for a +cause and the wrong one for a symptom with a second cause. Adopted: one kernel sentence in the engineering-system +paragraph that makes the second appearance evidence about a class, asks for the cheapest check that covers the +whole class, and puts the class fix before the next pass of the loop that found it; and one clause in the bug +workflow adding the places the owner would see the same symptom. Both state outcomes; how the class is found stays +with the agent. + +The larger mechanism was loading, and it is the reopen clause of the 3.0.0 reminder decision below: a session +SkipHow was governing lost its continuity, because Codex retains no text from any package file across compaction and +nothing that survives tells the session to reload. The 4.0.0 removal of the hook stands: it executed and did not +load. Adopted instead: one conditional sentence in the kernel description, the only package-owned text measured to +survive Codex compaction because the host keeps it in its own skill list, telling a session this skill was governing to reopen it after compaction. The condition +is the one 3.0.0 chose, already governing, so the sentence selects nothing for a session that never used the +skill. It executes nothing and adds under two hundred characters to every session's skill list on both hosts. +Claude re-attaches invoked skills on its own, where the sentence is redundant and harmless. + +Not adopted: rewording the documented activation line in the user instruction file, since installed files carry +their own wording and the package does not control that surface; a hook, per 4.0.0; a mandatory local end-to-end +run or a cap on release candidates, which belong to a project's delivery contract; a delegate-writer change, +because that kernel text was plain and only absent; a polling cadence, which the diagnosis reference already +states. Prior art read as it stands offered nothing on either shape. + +Evidence for the wording defects is one session; the causal effect of both changes is `UNVERIFIED`. This reopens +if a governed Codex session still never re-reads after compaction, if selection receipts show the description +sentence pulling the skill into an ungoverned session, or if a run sweeps unrelated surfaces on a first report. + ## The 4.8.3 delegation context and privacy correction Keep the economic routing contract and the existing continuation rule. The inspected sessions explicitly @@ -390,7 +425,7 @@ Version 2.11.2 leaves the resume and compaction reminder unconditional but stops Version 3.0.0 rejects the premise that kept the resume and compaction reminder unconditional. That premise was stated here twice: the reminders apply after the skill is already active, so they cannot broaden discovery the way the startup reminder did. It is not true. A session is compacted because it grew long and resumed because somebody came back to it, and neither event says anything about what the session was doing. A reminder that tells any resumed session to load the owner kernel selects the skill for requests the description was narrowed in 2.11.1 to exclude, which is the same failure that release fixed at startup, arriving instead at the point where a long session has just lost the context that would contradict it. The reminder is now conditional on SkipHow already governing the request. Where it was, continuity is restored exactly as before; where it was not, nothing asks the session to load a kernel it never used. The 2.11.2 correction stands unchanged: the reminder still selects no continuation store. -Revisit this if evidence shows that technical fluency itself changes which side of the decision boundary a person should occupy, or that the selection description rejects fitting outcome-first project work or selects adoption questions for mandatory development workflows and runtime orchestrators, or if a resumed session that SkipHow was governing loses its continuity because the reminder no longer fires. A preference to approve the method is already a different product fit, regardless of fluency. +Revisit this if evidence shows that technical fluency itself changes which side of the decision boundary a person should occupy, or that the selection description rejects fitting outcome-first project work or selects adoption questions for mandatory development workflows and runtime orchestrators, or if a resumed session that SkipHow was governing loses its continuity because the reminder no longer fires. A preference to approve the method is already a different product fit, regardless of fluency. The last condition was met by a private installed Codex session on 4.8.2; the 4.8.4 section above records the answer, a conditional sentence in the skill description rather than a hook. ## Host-native execution diff --git a/docs/evidence.md b/docs/evidence.md index f1b148d..128ce45 100644 --- a/docs/evidence.md +++ b/docs/evidence.md @@ -2,6 +2,50 @@ This page separates package checks from observed model behavior. The full 2.0 evidence remains in the immutable [`v2.0.1` research snapshot](https://github.com/mzored/SkipHow/tree/1c811262e6acdbdc58a2ee862b54e0b8d3478eaa/docs/research/2026-08-27). +## 4.8.4 repeat failures and continuity after compaction + +One private installed Codex session on 4.8.2 (Codex Desktop 0.154.0-alpha.6.2, about fourteen hours, ninety +turns, fifteen compactions) was inspected after the owner reported fixes that did not hold across screens and a +final review-and-deploy request that took almost four hours. No private session content, identifiers or project +details are retained here. + +Loading. After every compaction the retained history held only the host's developer message, which lists each +installed skill with its description and path, the owner's instruction files, the most recent owner messages and +one encrypted summary. Searching that retained history for distinctive sentences of nine package files found none +after any compaction. Of twenty-four owner requests asking for a systemic fix, package text was in context for +four. The kernel was read at the start of the final delivery request, dropped at the next compaction thirteen +minutes later, and never re-read. This is the reopen condition written into the 3.0.0 reminder decision. On Claude +Code the same claim was retracted in the 2026-09-06 audit because that host re-attaches invoked skills after +compaction; the two hosts differ and both facts stand. + +Class of fixes. Six reports of one visual symptom arrived over two hours; each fix repaired a different shared +component with a regression test and no sweep of the other screens. In the one turn where the bug workflow was in +context the run followed its text: it repaired the layer owning the failed rule and inspected the sibling that rule +governed. The recurrences had different causes. That is a readable gap between the workflow's cause-scoped class +and the outcome the owner asked for, supported by one observation. + +Delivery loop. The final request spent 590 tool calls and about 102 million input tokens, 99 percent of them +cached. Review was proportionate: three parallel forked reviewers found four qualifying defects in under half an +hour, and a one-minute re-review covered the changed parts. The remaining three and a half hours went to seven +sequential release candidates of eleven to thirty minutes each, five of which failed on one more member of the +same drift class, expectations left stale by agreed product changes. Local checks covering that whole class +existed and were not run before a candidate. The project's own contract runs end-to-end tests only on a +candidate, so the expectation is partly the project's; package text was absent, so the cause of the choice is +`UNVERIFIED`. Three delegates re-tasked as concurrent writers in one worktree contradicted kernel text that was +out of context; the host forbids spawning unless asked, and no wording changed. 383 host wait polls at the host's +thirty-second limit are host mechanics the diagnosis reference already addresses. + +Prior art was read as it stands on 2026-09-16: Superpowers' [`systematic-debugging`](https://github.com/obra/superpowers/blob/main/skills/systematic-debugging/SKILL.md) +and [`verification-before-completion`](https://github.com/obra/superpowers/blob/main/skills/verification-before-completion/SKILL.md) +say nothing about other instances of a defect class, a second report of the same symptom, or sweeping before +re-running an expensive gate. Nothing was taken. + +Not adopted: a candidate-count limit or a mandatory local suite before release, which belong to a project's +delivery contract; a delegate-writer change, since that text was plain and merely absent; a polling rule; a new +evaluation case or synthetic receipt. Behavior under both changes is `UNVERIFIED`. What would confirm them: a +later governed Codex session reading a package file after a compaction, and a repeated symptom report followed by +a fix set larger than the reported instance. + ## 4.8.3 delegation context and publication privacy Two private installed sessions on 4.8.2 were inspected for routing cost. Their recorded dispatches selected diff --git a/evals/cases.json b/evals/cases.json index 1311001..9ef5439 100644 --- a/evals/cases.json +++ b/evals/cases.json @@ -1,6 +1,6 @@ { "corpus_version": 4, - "package_under_test": "4.8.3", + "package_under_test": "4.8.4", "purpose": "Synthetic cases for three separate instruments: activation, forced-activation CTO behavior, and host smoke. Every case names a positive success observable, the product result shared across comparison arms, and explicit required-absence events. Nothing here has been run.", "not_a_gate": "No model run gates a pull request. python scripts/check.py and the pytest suite validate shape and internal satisfiability and never start a model. A run happens only when the owner authorizes a paid receipt, under the limits recorded in run_limits. A deterministic check passing is never evidence of behavior.", "evidence_labels": { diff --git a/evals/cto-cases.json b/evals/cto-cases.json index aa979fa..2b4638d 100644 --- a/evals/cto-cases.json +++ b/evals/cto-cases.json @@ -1,6 +1,6 @@ { "instrument": "forced_activation_behavior", - "package_under_test": "4.8.3", + "package_under_test": "4.8.4", "suite_status": "not_run", "minimum_coverage": { "case_ids": [ diff --git a/evals/host-smoke.json b/evals/host-smoke.json index 9c945a9..c74b12a 100644 --- a/evals/host-smoke.json +++ b/evals/host-smoke.json @@ -1,6 +1,6 @@ { "instrument": "host_smoke", - "package_under_test": "4.8.3", + "package_under_test": "4.8.4", "scope": "external_candidate_receipts", "checks": { "clean_install": { diff --git a/plugins/skiphow/.claude-plugin/plugin.json b/plugins/skiphow/.claude-plugin/plugin.json index db11df6..c9104f2 100644 --- a/plugins/skiphow/.claude-plugin/plugin.json +++ b/plugins/skiphow/.claude-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "skiphow", - "version": "4.8.3", + "version": "4.8.4", "description": "Adaptive virtual CTO for founders and product owners using Claude Code or Codex. Describe the product outcome; SkipHow owns the technical lifecycle through verified completion.", "author": { "name": "mzored", diff --git a/plugins/skiphow/.codex-plugin/plugin.json b/plugins/skiphow/.codex-plugin/plugin.json index d168856..c73e8bd 100644 --- a/plugins/skiphow/.codex-plugin/plugin.json +++ b/plugins/skiphow/.codex-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "skiphow", - "version": "4.8.3", + "version": "4.8.4", "description": "Adaptive virtual CTO for founders and product owners using Claude Code or Codex. Describe the product outcome; SkipHow owns the technical lifecycle through verified completion.", "author": { "name": "mzored", diff --git a/plugins/skiphow/skills/skiphow-bug/SKILL.md b/plugins/skiphow/skills/skiphow-bug/SKILL.md index 313bba4..4dd3612 100644 --- a/plugins/skiphow/skills/skiphow-bug/SKILL.md +++ b/plugins/skiphow/skills/skiphow-bug/SKILL.md @@ -7,7 +7,7 @@ description: Investigate and repair a reported defect at its root cause, coverin Resolve the reported defect and the class of failure that caused it. Before consequential work, have the [SkipHow CTO kernel](../skiphow/SKILL.md) in context. Read it if absent. It governs authority, scope, delegation, review, and delivery throughout this workflow. -Use [diagnosis](../skiphow/references/diagnosis.md) to establish the original failure signal and test the proposed cause. Identify the rule that failed and inspect the sibling paths governed by it. Repair the layer that owns that rule within the authorized scope. A general repair does not require a repository-wide refactor; record a separable problem through the kernel's findings policy. +Use [diagnosis](../skiphow/references/diagnosis.md) to establish the original failure signal and test the proposed cause. Identify the rule that failed and inspect the sibling paths governed by it, and the other places the owner would see the same symptom, since a repeat of one symptom can have a second cause. Repair the layer that owns that rule within the authorized scope. A general repair does not require a repository-wide refactor; record a separable problem through the kernel's findings policy. Consult [technical design](../skiphow/references/technical-design.md) when the repair depends on external facts, introduces a dependency or abstraction, or could reuse an existing capability. Research the uncertainty that affects the repair. diff --git a/plugins/skiphow/skills/skiphow/SKILL.md b/plugins/skiphow/skills/skiphow/SKILL.md index fe00160..093b987 100644 --- a/plugins/skiphow/skills/skiphow/SKILL.md +++ b/plugins/skiphow/skills/skiphow/SKILL.md @@ -1,6 +1,6 @@ --- name: skiphow -description: Act as an adaptive virtual CTO for a founder or product owner. Use for any current-project outcome stated in ordinary language, including questions, research, reviews, bugs, ideas, features, iterations on something the owner will look at, lists, programmes, project status, unfinished work, cleanup, delivery, process problems, pauses, and resumes. The owner keeps product decisions; the agent owns the technical lifecycle through verified completion. Also use when the owner asks to enable, check, or disable SkipHow itself on this machine. Do not use for unrelated conversation. +description: Act as an adaptive virtual CTO for a founder or product owner. Use for any current-project outcome stated in ordinary language, including questions, research, reviews, bugs, ideas, features, iterations on something the owner will look at, lists, programmes, project status, unfinished work, cleanup, delivery, process problems, pauses, and resumes. The owner keeps product decisions; the agent owns the technical lifecycle through verified completion. Also use when the owner asks to enable, check, or disable SkipHow itself on this machine. If the host compacted the context while this skill was governing the session, reopen this file before the next consequential action. Do not use for unrelated conversation. --- # SkipHow @@ -71,7 +71,7 @@ Delegate only bounded work whose context isolation, independent judgment, or par Every change gets a fresh review of the final state. A small clear low-risk edit may use a cold self-review and targeted evidence; visibility and file count alone do not require delegation. Use an independent reviewer when substantive behavior, interacting changes, or dependency and integration risks make a shared blind spot consequential. Architecture, security, authentication, payments, privacy, migration, concurrency, or public-contract changes get stronger independent challenge. Confirm findings against the repository, fix qualifying defects, and rerun affected evidence. Re-review the changed parts after a fix. Stop when the remaining items are taste, lack evidence, or are explicitly reported as unresolved; use another broad reviewer only to resolve a high-consequence disagreement or contradictory evidence. -Treat activation, fixtures, CI, permissions, tools, hooks, worktrees, coordination, flaky checks, silent errors, repeated timeouts, recurring manual workarounds, and verification cost or maintenance materially disproportionate to the changed behavior as engineering-system signals. Recurring broad test churn, slow feedback, expensive setup, or poor failure localization call for diagnosis of the responsible layer, not a presumed test-type cause or an automatic broad refactor. Do not hide a process or environment defect by extending a timeout, adding retries, disabling checks, or weakening assertions. +Treat activation, fixtures, CI, permissions, tools, hooks, worktrees, coordination, flaky checks, silent errors, repeated timeouts, recurring manual workarounds, and verification cost or maintenance materially disproportionate to the changed behavior as engineering-system signals. Recurring broad test churn, slow feedback, expensive setup, or poor failure localization call for diagnosis of the responsible layer, not a presumed test-type cause or an automatic broad refactor. A second appearance of the same kind of failure, the owner reporting a symptom again or a slow gate failing on one more stale expectation, is evidence about a class rather than another instance: find the other members with the cheapest check that covers the whole class and fix them together before returning to the loop that found them. Do not hide a process or environment defect by extending a timeout, adding retries, disabling checks, or weakening assertions. ## Work you do not own and delegates diff --git a/site/evidence/index.html b/site/evidence/index.html index 13eeb29..aa36aab 100644 --- a/site/evidence/index.html +++ b/site/evidence/index.html @@ -56,11 +56,11 @@

Claims stop where the receipts stop.

-
+

What controlled 2.x runs showed.

-

Historical observations below: 2.x only. Current package: 4.8.3. Retained current-package CTO scenarios with Observed receipts: 0 of 12. Eight historical 4.0.1 Claude run records remain incomplete. Current CTO scenario behavior is UNVERIFIED; configuration and deterministic checks do not establish model behavior.

+

Historical observations below: 2.x only. Current package: 4.8.4. Retained current-package CTO scenarios with Observed receipts: 0 of 12. Eight historical 4.0.1 Claude run records remain incomplete. Current CTO scenario behavior is UNVERIFIED; configuration and deterministic checks do not establish model behavior.