Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
47 changes: 47 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,53 @@

All notable changes to SkipHow 2.x and later appear in this file. Earlier release notes remain available on [GitHub Releases](https://github.com/mzored/SkipHow/releases).

## 4.8.4 (2026-09-16)

A repeat of the same kind of failure is now evidence about a class. The always-loaded kernel says that an owner
reporting a symptom again, or a slow gate failing on one more stale expectation, calls for finding the other
members of that class with the cheapest check that covers all of them and fixing them together before returning
to the loop that found them. The bug workflow's sibling inspection now also covers the other places the owner
would see the same symptom, since a repeated symptom can have a second cause. No candidate-count limit,
mandatory local suite, sweep procedure, or polling rule was added.

The skill description carries one conditional sentence: a session this skill was governing reopens the skill
after the host compacts its context. It is not a hook and executes nothing. It answers the reopen condition
recorded in the 3.0.0 reminder decision, which fired: in a private installed Codex session on 4.8.2 no text from any
package file survived any of fifteen compactions, only the skill descriptions the host keeps in its own skill list
did, and the kernel was never re-read across a delivery loop of almost four hours. Whether the sentence causes a reload
is `UNVERIFIED`.

The motivating observations, prior-art reading and refused alternatives are recorded in
[evidence](docs/evidence.md#484-repeat-failures-and-continuity-after-compaction) and
[decisions](docs/decisions.md#the-484-class-evidence-and-continuity-correction). No private session content is
published and no paid behavioral comparison was run. This is a compatible wording patch.

The full local package gate passed all 389 tests under the pinned dependencies in an isolated worktree, and
`git diff --check` passed. Both host schema validators passed on the candidate tree. Clean installation and model behavior were not retested for this
unchanged package structure.

An isolated, read-only Codex review found one qualifying documentation defect: the release rationale said none
of the package's text survives Codex compaction while also calling the surviving description package-owned text.
The three sentences now distinguish the discarded file bodies from the descriptions the host keeps in its own
skill list. The reviewer confirmed that the kernel sentence contradicts neither the verification nor the diagnosis
reference, that the description sentence selects nothing for an ungoverned session and grants no authority, that
the bug workflow clause is consistent with the kernel, that all referenced anchors resolve, and that no active
version pin was missed. No findings were refused. The review transcript confirmed separate operating-system and
host homes with no personal instruction or installed-plugin contamination markers. Authentication referenced
the existing credential file; the owned review workspace was removed.

| Capability | Local candidate evidence |
| --- | --- |
| Deterministic package gate | PASS, full local command and 389 tests |
| Codex schema validation | PASS |
| Claude schema validation | PASS |
| Clean Codex install | UNVERIFIED |
| Clean Claude install | UNVERIFIED |
| Explicit invocation | UNVERIFIED |
| Implicit activation | UNVERIFIED |
| Continuity | UNVERIFIED |
| Behavioral suite | UNVERIFIED, [evidence](docs/evidence.md) |

## 4.8.3 (2026-09-13)

Delegate briefs now account for context the host actually supplies instead of asserting that delegates
Expand Down
2 changes: 1 addition & 1 deletion VERSION
Original file line number Diff line number Diff line change
@@ -1 +1 @@
4.8.3
4.8.4
41 changes: 38 additions & 3 deletions docs/decisions.md

Large diffs are not rendered by default.

44 changes: 44 additions & 0 deletions docs/evidence.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,50 @@

This page separates package checks from observed model behavior. The full 2.0 evidence remains in the immutable [`v2.0.1` research snapshot](https://github.com/mzored/SkipHow/tree/1c811262e6acdbdc58a2ee862b54e0b8d3478eaa/docs/research/2026-08-27).

## 4.8.4 repeat failures and continuity after compaction

One private installed Codex session on 4.8.2 (Codex Desktop 0.154.0-alpha.6.2, about fourteen hours, ninety
turns, fifteen compactions) was inspected after the owner reported fixes that did not hold across screens and a
final review-and-deploy request that took almost four hours. No private session content, identifiers or project
details are retained here.

Loading. After every compaction the retained history held only the host's developer message, which lists each
installed skill with its description and path, the owner's instruction files, the most recent owner messages and
one encrypted summary. Searching that retained history for distinctive sentences of nine package files found none
after any compaction. Of twenty-four owner requests asking for a systemic fix, package text was in context for
four. The kernel was read at the start of the final delivery request, dropped at the next compaction thirteen
minutes later, and never re-read. This is the reopen condition written into the 3.0.0 reminder decision. On Claude
Code the same claim was retracted in the 2026-09-06 audit because that host re-attaches invoked skills after
compaction; the two hosts differ and both facts stand.

Class of fixes. Six reports of one visual symptom arrived over two hours; each fix repaired a different shared
component with a regression test and no sweep of the other screens. In the one turn where the bug workflow was in
context the run followed its text: it repaired the layer owning the failed rule and inspected the sibling that rule
governed. The recurrences had different causes. That is a readable gap between the workflow's cause-scoped class
and the outcome the owner asked for, supported by one observation.

Delivery loop. The final request spent 590 tool calls and about 102 million input tokens, 99 percent of them
cached. Review was proportionate: three parallel forked reviewers found four qualifying defects in under half an
hour, and a one-minute re-review covered the changed parts. The remaining three and a half hours went to seven
sequential release candidates of eleven to thirty minutes each, five of which failed on one more member of the
same drift class, expectations left stale by agreed product changes. Local checks covering that whole class
existed and were not run before a candidate. The project's own contract runs end-to-end tests only on a
candidate, so the expectation is partly the project's; package text was absent, so the cause of the choice is
`UNVERIFIED`. Three delegates re-tasked as concurrent writers in one worktree contradicted kernel text that was
out of context; the host forbids spawning unless asked, and no wording changed. 383 host wait polls at the host's
thirty-second limit are host mechanics the diagnosis reference already addresses.

Prior art was read as it stands on 2026-09-16: Superpowers' [`systematic-debugging`](https://github.com/obra/superpowers/blob/main/skills/systematic-debugging/SKILL.md)
and [`verification-before-completion`](https://github.com/obra/superpowers/blob/main/skills/verification-before-completion/SKILL.md)
say nothing about other instances of a defect class, a second report of the same symptom, or sweeping before
re-running an expensive gate. Nothing was taken.

Not adopted: a candidate-count limit or a mandatory local suite before release, which belong to a project's
delivery contract; a delegate-writer change, since that text was plain and merely absent; a polling rule; a new
evaluation case or synthetic receipt. Behavior under both changes is `UNVERIFIED`. What would confirm them: a
later governed Codex session reading a package file after a compaction, and a repeated symptom report followed by
a fix set larger than the reported instance.

## 4.8.3 delegation context and publication privacy

Two private installed sessions on 4.8.2 were inspected for routing cost. Their recorded dispatches selected
Expand Down
2 changes: 1 addition & 1 deletion evals/cases.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"corpus_version": 4,
"package_under_test": "4.8.3",
"package_under_test": "4.8.4",
"purpose": "Synthetic cases for three separate instruments: activation, forced-activation CTO behavior, and host smoke. Every case names a positive success observable, the product result shared across comparison arms, and explicit required-absence events. Nothing here has been run.",
"not_a_gate": "No model run gates a pull request. python scripts/check.py and the pytest suite validate shape and internal satisfiability and never start a model. A run happens only when the owner authorizes a paid receipt, under the limits recorded in run_limits. A deterministic check passing is never evidence of behavior.",
"evidence_labels": {
Expand Down
2 changes: 1 addition & 1 deletion evals/cto-cases.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"instrument": "forced_activation_behavior",
"package_under_test": "4.8.3",
"package_under_test": "4.8.4",
"suite_status": "not_run",
"minimum_coverage": {
"case_ids": [
Expand Down
2 changes: 1 addition & 1 deletion evals/host-smoke.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"instrument": "host_smoke",
"package_under_test": "4.8.3",
"package_under_test": "4.8.4",
"scope": "external_candidate_receipts",
"checks": {
"clean_install": {
Expand Down
2 changes: 1 addition & 1 deletion plugins/skiphow/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "skiphow",
"version": "4.8.3",
"version": "4.8.4",
"description": "Adaptive virtual CTO for founders and product owners using Claude Code or Codex. Describe the product outcome; SkipHow owns the technical lifecycle through verified completion.",
"author": {
"name": "mzored",
Expand Down
2 changes: 1 addition & 1 deletion plugins/skiphow/.codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "skiphow",
"version": "4.8.3",
"version": "4.8.4",
"description": "Adaptive virtual CTO for founders and product owners using Claude Code or Codex. Describe the product outcome; SkipHow owns the technical lifecycle through verified completion.",
"author": {
"name": "mzored",
Expand Down
2 changes: 1 addition & 1 deletion plugins/skiphow/skills/skiphow-bug/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ description: Investigate and repair a reported defect at its root cause, coverin

Resolve the reported defect and the class of failure that caused it. Before consequential work, have the [SkipHow CTO kernel](../skiphow/SKILL.md) in context. Read it if absent. It governs authority, scope, delegation, review, and delivery throughout this workflow.

Use [diagnosis](../skiphow/references/diagnosis.md) to establish the original failure signal and test the proposed cause. Identify the rule that failed and inspect the sibling paths governed by it. Repair the layer that owns that rule within the authorized scope. A general repair does not require a repository-wide refactor; record a separable problem through the kernel's findings policy.
Use [diagnosis](../skiphow/references/diagnosis.md) to establish the original failure signal and test the proposed cause. Identify the rule that failed and inspect the sibling paths governed by it, and the other places the owner would see the same symptom, since a repeat of one symptom can have a second cause. Repair the layer that owns that rule within the authorized scope. A general repair does not require a repository-wide refactor; record a separable problem through the kernel's findings policy.

Consult [technical design](../skiphow/references/technical-design.md) when the repair depends on external facts, introduces a dependency or abstraction, or could reuse an existing capability. Research the uncertainty that affects the repair.

Expand Down
4 changes: 2 additions & 2 deletions plugins/skiphow/skills/skiphow/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: skiphow
description: Act as an adaptive virtual CTO for a founder or product owner. Use for any current-project outcome stated in ordinary language, including questions, research, reviews, bugs, ideas, features, iterations on something the owner will look at, lists, programmes, project status, unfinished work, cleanup, delivery, process problems, pauses, and resumes. The owner keeps product decisions; the agent owns the technical lifecycle through verified completion. Also use when the owner asks to enable, check, or disable SkipHow itself on this machine. Do not use for unrelated conversation.
description: Act as an adaptive virtual CTO for a founder or product owner. Use for any current-project outcome stated in ordinary language, including questions, research, reviews, bugs, ideas, features, iterations on something the owner will look at, lists, programmes, project status, unfinished work, cleanup, delivery, process problems, pauses, and resumes. The owner keeps product decisions; the agent owns the technical lifecycle through verified completion. Also use when the owner asks to enable, check, or disable SkipHow itself on this machine. If the host compacted the context while this skill was governing the session, reopen this file before the next consequential action. Do not use for unrelated conversation.
---

# SkipHow
Expand Down Expand Up @@ -71,7 +71,7 @@ Delegate only bounded work whose context isolation, independent judgment, or par

Every change gets a fresh review of the final state. A small clear low-risk edit may use a cold self-review and targeted evidence; visibility and file count alone do not require delegation. Use an independent reviewer when substantive behavior, interacting changes, or dependency and integration risks make a shared blind spot consequential. Architecture, security, authentication, payments, privacy, migration, concurrency, or public-contract changes get stronger independent challenge. Confirm findings against the repository, fix qualifying defects, and rerun affected evidence. Re-review the changed parts after a fix. Stop when the remaining items are taste, lack evidence, or are explicitly reported as unresolved; use another broad reviewer only to resolve a high-consequence disagreement or contradictory evidence.

Treat activation, fixtures, CI, permissions, tools, hooks, worktrees, coordination, flaky checks, silent errors, repeated timeouts, recurring manual workarounds, and verification cost or maintenance materially disproportionate to the changed behavior as engineering-system signals. Recurring broad test churn, slow feedback, expensive setup, or poor failure localization call for diagnosis of the responsible layer, not a presumed test-type cause or an automatic broad refactor. Do not hide a process or environment defect by extending a timeout, adding retries, disabling checks, or weakening assertions.
Treat activation, fixtures, CI, permissions, tools, hooks, worktrees, coordination, flaky checks, silent errors, repeated timeouts, recurring manual workarounds, and verification cost or maintenance materially disproportionate to the changed behavior as engineering-system signals. Recurring broad test churn, slow feedback, expensive setup, or poor failure localization call for diagnosis of the responsible layer, not a presumed test-type cause or an automatic broad refactor. A second appearance of the same kind of failure, the owner reporting a symptom again or a slow gate failing on one more stale expectation, is evidence about a class rather than another instance: find the other members with the cheapest check that covers the whole class and fix them together before returning to the loop that found them. Do not hide a process or environment defect by extending a timeout, adding retries, disabling checks, or weakening assertions.

## Work you do not own and delegates

Expand Down
4 changes: 2 additions & 2 deletions site/evidence/index.html
Original file line number Diff line number Diff line change
Expand Up @@ -56,11 +56,11 @@ <h1>Claims stop where the receipts stop.</h1>
</div>
</section>

<section class="section" aria-labelledby="matrix-title" data-observed-package-series="2.x" data-current-package="4.8.3" data-current-observed-scenarios="0">
<section class="section" aria-labelledby="matrix-title" data-observed-package-series="2.x" data-current-package="4.8.4" data-current-observed-scenarios="0">
<div class="shell">
<p class="section-label">01 &middot; Historical observations</p>
<h2 id="matrix-title">What controlled 2.x runs showed.</h2>
<p>Historical observations below: 2.x only. Current package: 4.8.3. Retained current-package CTO scenarios with Observed receipts: 0 of 12. Eight historical 4.0.1 Claude run records remain incomplete. Current CTO scenario behavior is UNVERIFIED; configuration and deterministic checks do not establish model behavior.</p>
<p>Historical observations below: 2.x only. Current package: 4.8.4. Retained current-package CTO scenarios with Observed receipts: 0 of 12. Eight historical 4.0.1 Claude run records remain incomplete. Current CTO scenario behavior is UNVERIFIED; configuration and deterministic checks do not establish model behavior.</p>
<div class="comparison-wrap wide-comparison" role="region" aria-label="Historical 2.x evidence table" tabindex="0">
<table>
<thead>
Expand Down