Repository navigation
world: first in-repo slice of the research lane - #7
Conversation
Creates the World layout at world/ with WORLD_VERSION 0.1.0 and seeds the research lane: four spacing claims and three BKT claims, every citation resolved against Crossref before being written, and per-skill-type BKT priors that link back to those claims. Everything enters at review: "unreviewed" with reviewer: null. Nothing here claims to be correct; it claims to be proposed. Adds scripts/validate-world.mjs, the layer-1 mechanical gate from docs/WORLD.md, and wires it into CI beside npm test. Adds the OpenClaw worktree hooks (.worktreeinclude, .openclaw/worktree-setup.sh) and the spacing-literature watcher with its source URL left as an explicit TODO. Pins two new boundaries in test/invariants.test.ts: no World data structure contains a learner id, and no pipeline code path can read a family record. Also lands CLAUDE.md at the root and the WORLD and OPENCLAW specs under docs/, which are the documents this slice implements. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The sha256-of-a-page primitive in the docs/OPENCLAW.md example fires on view counters, rotating sidebars, and session tokens. It would have woken a model run every week regardless of whether any literature appeared, which is the queue-nobody-reviews failure that file warns about, reached from the other direction. New DOIs since the last check are a real signal; changed bytes are not. Crossref needs no key and has from-created-date built for this. But a relevance query alone is unusable here: "spacing" and "retention" are cross-disciplinary homonyms, and the unrestricted query returns aerospace, civil engineering, and nurse-staffing papers. Measured against the live API, the open query matched 153,105 works and the same query restricted to an ISSN allowlist matched 45, all on topic. The allowlist covers the venues world/claims/spacing.md already cites, because a watcher blind to the journals its own claims came from would never notice those claims being superseded. Every ISSN was resolved through api.crossref.org/journals; two of them appear inside the DOIs already cited in the claims file. MIN_NEW is 1 against a measured rate of 33 works in 12 months. A higher threshold would leave the pipeline silent for months and indistinguishable from a broken one. First run seeds without firing, so the proposer never faces a batch it cannot turn into one coherent PR. A failed fetch fires with a message saying the watcher is broken and preserves the prior DOI set, because returning fresh state on failure would make the next run treat every known paper as new. No top-level return, so the script is valid whether the runtime evaluates it as a module or wraps it in a function. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The ISSN allowlist filters venue, not topic, so an off-topic paper in a watched journal will still wake the proposer. Verified against the live API: the trailing 30-day window returned one work, on primary-secondary school transitions, which matches "retention" in the ed-psych sense and has nothing to do with spacing. An isolated cron run is unattended by contract and its final reply must be the deliverable, which is exactly the pressure that turns a quiet week into a manufactured PR. Say plainly that an empty week is a normal outcome and that replying with one sentence and no PR is a valid deliverable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Pointing a pipeline agent's workspace at this repository, as docs/OPENCLAW.md specifies, makes OpenClaw bootstrap its own files into the working tree on first run: SOUL.md, IDENTITY.md, HEARTBEAT.md, USER.md, TOOLS.md, memory/, openclaw-workspace-state.json. Observed after the first real run of the spacing watcher. That is maintainer-side machinery and must never reach a release. memory/ is the pointed one: it accumulates agent session content, and this repository's whole argument is that private material does not travel. AGENTS.md is deliberately not ignored. docs/OPENCLAW.md expects a hand-written lane charter there, and ignoring it would make that charter hard to commit later. The file currently in the tree is OpenClaw's generated placeholder, not the charter. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The file in the tree was OpenClaw's generated placeholder, which told a pipeline agent to keep daily memory files, react with emoji in group chats, and use TTS for storytime. None of that is true of an agent whose entire job is to open one pull request against world/ and then stop. Four agents share this repository as their workspace, so there is one charter file and each agent reads the section for the id it runs as. The lane descriptions are copied from docs/WORLD.md so a lane cannot drift from the spec without the drift showing in a diff, and each carries its current status: research live, curriculum unbuilt, careers deliberately jobless per the build order, attacker live. A lane woken before its watcher exists is told to report the misconfiguration and do nothing. Both standing rules are stated first because they are the ones an unattended run is most likely to reason its way around under pressure: propose via PR only, and touch nothing outside world/ and the lane's fixtures. The corollaries are spelled out rather than implied. A proposal that seems to need a change in src/ is a finding for the human, not a diff. Promotion from world/curriculum/ to src/curriculum/ is a human act. Proposing nothing is a successful run, which is the same pressure release the spacing watcher already carries and which belongs in the charter too, not only in one watcher's message. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Layer 2 of the admission pipeline had all four agents and no job. This adds the watcher that wakes world-attacker, and changes two things from the sketch in docs/OPENCLAW.md. It selects pull requests by changed path rather than by the head:openclaw/ branch prefix. The failure modes in the attack brief belong to World content, not to agents: a miscited paper, a skill placed a grade early, an agreement count that dissolves because two sources share an upstream. A human pull request that seeds world/claims/ carries every one of them, and PR VedSoni-dev#7 is exactly that pull request, so a branch filter would skip the one PR definition-of-done step 4 names. Meanwhile an agent branch touching only src/ has nothing in it to attack. Verified against the live repository: the branch filter matches zero open pull requests today, the path filter matches VedSoni-dev#7 and lists the five World files it changes. .gitkeep is excluded, since a PR that only adds directory placeholders has moved no claim. It keys the seen-set on PR number plus head commit, so a proposal pushed again after an attack is attacked again. Keying on the number alone, as the sketch does, attacks the first revision and nothing after it, which is backwards: the revision most likely to be wrong is the one written in response to the last set of findings. The attacker holds no write tool, so its own comment never moves a head and never re-fires the watcher. Hourly, not weekly. A run that finds nothing costs the script budget alone, and a PR is worth attacking while its author is still looking at it. Each fired run is capped at three so that each PR gets a real read; the overflow stays unseen for the next run and the message says how many are waiting, because a silent cap reads as full coverage. A failed gh call fires with a broken-watcher message and preserves state, the same shape the spacing watcher uses, since expired auth otherwise fails silently forever. Registered against the gateway and left disabled. Enabling it posts a public comment on an open pull request within the hour, which is the maintainer's call to make, not this session's. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Attacked at head 65fb972. I resolved all seven DOIs against Crossref and each resolves to the paper the entry attributes to it: Cepeda et al. 2006 (doi:10.1037/0033-2909.132.3.354), Cepeda et al. 2008 (doi:10.1111/j.1467-9280.2008.02209.x), Roediger and Karpicke 2006 (doi:10.1111/j.1467-9280.2006.01693.x), Settles and Meeder 2016 (doi:10.18653/v1/P16-1174), Corbett and Anderson 1995 (doi:10.1007/BF01099821), Baker, Corbett, and Aleven 2008 (doi:10.1007/978-3-540-69132-7_44), and Beck and Chang 2007 (doi:10.1007/978-3-540-73078-1_17). The quantitative claims hold: Cepeda 2006 reports 317 experiments in 184 articles, Cepeda 2008 taught more than 1,350 individuals with the optimal-gap ridgeline rising with retention interval, and Settles and Meeder report the HLR error reduction the half-life claim leans on. Strength labels are calibrated to what the sources support. The half-life claim is marked moderate and flags its population as adult language learners, not children. The identifiability claim is marked strong and matches Beck and Chang's infinite-family result. The two agreement downcounts are conservative and correct. spacing/optimal-gap-grows counts separately from the 2006 meta-analysis despite shared authors rather than claiming agreement, and bkt/guess-two-choice lists two DOIs but counts agreement as one because Corbett is a shared author. No prior contradicts a claim it depends on: the selected_response, performance, and constructed_response guess priors cite bkt/guess-two-choice under its stated general principle that guess tracks the chance rate of the answer format, and 0.45 for two_choice sits under the 0.5 identifiability ceiling that bkt/identifiability asserts. No learner id and no demographic variation appears in any world/ file. The two grep hits on those terms are the doctrine restatement in world/README.md and the honest population-scoping note in the half-life claim. This proposal survives the attack. This comment admits nothing and is not a merge signal. |
The first attack on PR VedSoni-dev#7 was a forced run, and a forced run executes the payload without ever evaluating the watcher. Trigger state stayed empty, so the next scheduled evaluation would have found VedSoni-dev#7 unattacked and posted a second set of findings on a pull request that already carried them. Caught before it fired. A duplicate attack comment is the exact thing that teaches a reviewer to skim this layer, which is the only way layer 2 can fail quietly. The seen-set is now a cache in front of the durable answer rather than the answer. Every comment opens with `<!-- world-attacker head=<sha> -->` and the watcher reads existing PR comments, so state that is lost, reset by a job edit, or never written costs nothing. The already-posted comment on VedSoni-dev#7 is matched by its prose form, so the fix covers the comment that prompted it. The signature doubles as an attribution fix. The attacker posts through the maintainer's gh credentials, so its findings land under the maintainer's own GitHub account with nothing marking them machine-authored. The brief now requires it to name itself and the head it attacked in the first sentence. Layer 2 speaking is not the human speaking, and rule 7 only means something if the record shows which one did. Ten cases verified against a stubbed pull request list: fires on an unattacked head, suppressed by either marker form at that head, fires again when the head moves past a marker, skips drafts, skips PRs with no world/ content, skips gitkeep-only PRs, caps at three, and picks the overflow up on the next run rather than losing it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
I am world-attacker, and I attacked this PR at head 7f55113. This proposal survives the attack: I found no defect that should stop it reaching human review. What I checked and what held:
Two notes for the human reviewer, neither a defect in this PR:
I admit nothing and this comment is not a merge signal. The merge is the human's. |
Both watchers fired on the first failed check. The attacker watcher runs hourly, the laptop went offline for about sixteen hours on 1 August, and the result was eight fired runs and eight failure alerts. Every one of those runs then hung and aborted, between sixteen and twenty-three minutes each, because the missing network that broke the gh call also broke the agent's own provider call. The run was woken to diagnose a condition that prevents it from running. Verified against the job history: ten runs, two successes before the outage, then AbortError, FailoverError DNS lookup failed, and one model idle timeout. The distinction that fixes it is whether a response came back. A call that never completed is counted and stays quiet, because the most likely cause is the machine, not the source. A call that completed and returned the wrong shape is announced immediately, because a response arriving proves the network is up and the source contract is what changed. That keeps the property the original code was reaching for, which is that a broken watcher must not look like a quiet one, without paying a wasted model run for every wifi hiccup. Alerts back off by a factor of four after each firing. Twenty-four hourly polls with no network now produce two alerts instead of twenty-four, and a single blip produces none. Counters clear on the first successful check. Previous state is spread rather than replaced, so an outage cannot cost the attacker its seen-set or the research lane its DOI set, either of which would cause a duplicate attack or several hundred papers treated as new. Both jobs also get --timeout-seconds 600, against successful runs of 115 and 176 seconds. A run that loses the network does not fail, it hangs, and a ceiling is what turns a twenty-three minute hang into a bounded failure. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
world-attacker, layer 2 of the admission pipeline, attacked this proposal at head 7df447c. I read the full diff and resolved every citation the World content carries. The claims and priors are the attack surface; the docs, watchers, CI wiring, and validator are scaffolding this lane does not gate. Findings against the seven DOIs and the coherence links:
This proposal survives the attack. I admit nothing and this comment is not a merge signal; a named human owns the merge and the provenance |
What changed
world/exists, laid out perdocs/WORLD.md:claims/,priors/,curriculum/,lenses/,trajectories/, andWORLD_VERSIONat 0.1.0.scripts/validate-world.mjsimplements the layer-1 mechanical checks fromdocs/WORLD.md: schema conformance, prerequisite cycle detection, duplicate skill ids, provenance completeness, agreement bounds, reading level of child-facing prompt text, and assessable success criteria for every skill. It runs in CI besidenpm testand fails loudly..worktreeincludeand an executable.openclaw/worktree-setup.shsupport OpenClaw managed worktrees. The setup script runsnpm ciandnpm test, so a proposal can never start from a checkout whose tests do not pass.automation/watchers/spacing-literature.jsis the research-lane watcher. Its source URL is a named constant left as an explicit TODO for the operator. A failed fetch fires with a warning instead of going quiet.test/invariants.test.tspins two new boundaries: no World data structure contains a learner id, and no pipeline code path can read a family record.world/claims/spacing.md(four claims),world/claims/bkt-parameters.md(three claims), andworld/priors/bkt-priors.json(per-skill-type BKT and forgetting-curve priors, each linked back to the claims it rests on).CLAUDE.mdlands at the repository root and the WORLD and OPENCLAW specs land underdocs/. They are the documents this slice implements.Why
The research lane is the smallest artifact that exercises the whole admission pipeline, and it bridges the honest "principled defaults" label in SPEC.md toward "fitted to data". This PR is the in-repo half only: the layout, the mechanical gate, the hooks, and seed content. The operator-side half, agents, cron jobs, and the trigger gate, is deliberately not part of this PR. Those are decisions a human makes at a keyboard.
Provenance
Every citation was resolved against Crossref metadata on 2026-07-30 before being written down. Verification of a citation is not review of a claim: every entry carries
review: "unreviewed"andreviewer: null. Agreement counts are conservative; two sources that share an author are counted as one.Status
Everything here is proposed. Nothing in this PR claims to be correct, reviewed, or shipped. The test suite was green before the change and is green after it (186 tests, including the two new invariant pins), and
node scripts/validate-world.mjspasses on the seeded World.🤖 Generated with Claude Code