Skip to content

RFC: activating the permission system — first write actions and the gated approval loop #48

Description

@hsliuustc0106

RFC: activating the permission system — first write actions and the gated approval loop

Summary

The permission foundations shipped in #13 (modes, scoped grants, the
approvals surface, the no-external-write invariant) are dormant because
not a single write action exists. This RFC proposes how to activate them:
a write port (ports/writer.py) as a new explicit egress point,
PR comment, label, and review approval as the first write actions,
both write modes going live with auto as the default (pre-granted
capabilities only, never prompts; gated is the opt-in interactive loop,
readonly the explicit hard-off), a separated fine-grained write
token
so a default install remains physically incapable of writing, and
a hard write quota so write pressure can never flood a PR or
destabilize the watch loop.

Maintainer decisions from 2026-10-03 are recorded inline and in the
Decisions section (1–8). All open questions are resolved; the RFC is
converged and ready for P0 implementation breakdown.

Current state (main @ 32c214e)

  • core/permissions.py: Mode.READONLY/GATED/AUTO with GATED/AUTO
    as documented dormant placeholders ("write actions would pause for
    approval" / "pre-granted scoped capabilities only"). Grant and
    ApprovalRequest schemas persist in SQLite (request TTL 24h, states
    pending/approved/denied/expired). assert_allowed denies every external
    write; scope changes atomically revoke grants.
  • nanodot approvals lists requests and grants — always empty today.
  • native/github_client.py is read-only REST by construction;
    docs/design/egress.md states "no POST/PUT/PATCH/DELETE is ever issued"
    and names exactly two egress points.
  • The belt-and-suspenders design (Define nanodot product design and MVP scope #1 decision 4): the mode governs what
    nanodot attempts; the read-only PAT caps what it could do — "even a
    bug cannot produce a write."
  • Decision-log guarantee (Activity log as a decision log: per-poll observations and rule provenance #36): every stop/notify decision is replayable
    from the activity log. Write actions must join this audit model.

Motivation

  1. The watch loop is a spectator: it sees failures and terminal states but
    cannot act on them (leave the "retested, ignore flaky X" comment, label
    a PR watched-by-nanodot). The obvious next unit of value is a write
    that closes the loop the user is already watching.
  2. The permission schema was built "so they have somewhere to land"
    (Permissions: modes, scoped grants, approvals surface #13) — this is the landing, and doing it late would risk the schema
    drifting from real needs.
  3. Every dormant-safety feature is an untested one: the approval loop has
    never executed a real request lifecycle end to end.

Proposed goals

  1. Write port (ports/writer.py, decided): the only code path able to
    issue an authenticated non-GET request, admitted through the
    adapter-seam disciplines. The read client keeps its "cannot write"
    invariant syntactic.
  2. First write actions: PR comment, PR label, and PR review approval
    (decided; merge stays out — irreversible). All three are auditable on
    GitHub's side and recoverable if misfired; approve requests must carry
    enough context (PR, head SHA, evidence digest) for an informed
    approve/deny decision.
  3. Both write modes activate; auto is the default (decided).
    permission-mode defaults to auto: a write executes only against a
    pre-existing grant — no runtime prompting, ever; a grant-less write is
    skipped and recorded (skipped: no grant). gated is the opt-in
    interactive mode: a write without a grant pauses as a pending
    ApprovalRequest, the CLI gains nanodot approvals approve/deny ID,
    expiry keeps the request visible, denials are recorded and never
    re-asked verbatim. readonly remains the explicit hard-off. Silence
    is never approval in either mode. Without a write token, every mode is
    behaviorally readonly.
  4. Separated write credential. A distinct secret (e.g.
    github-write-token) holding a fine-grained PAT limited to the exact
    first-action permissions (PR comments + labels + PR reviews, single
    repository). No write token configured ⇒ gated mode cannot be enabled,
    and the read-only PAT remains a hard cap for every other path. A
    default install stays physically read-only.
  5. Audit: every write attempt — requested, approved/denied/expired,
    issued, failed, succeeded — appends to the activity log with the grant
    id and evidence shape, so the decision log remains the complete
    replayable record.
  6. Write quota (decided): the write port carries its own hard budget —
    per action, per watch, per day (decided: 3 comments / 3 labels /
    1 approval per watch per day, configurable per watch). Exhausting
    it is a fail-closed degrade, never an error path: the action stops
    being attempted for the day, the watch loop, reads, and raw
    notifications continue untouched.

Non-goals (proposed)

  • No local command execution. A ZCode-style executor has a completely
    different risk surface (arbitrary code, filesystem, network); it gets
    its own RFC if ever pursued.
  • No merge writes — merge is irreversible; revisit after the
    comment/label/approve loop is proven.
  • No external notification channels (Slack et al.) — decision 1 keeps
    them deferred; outbound messaging is a different egress category.
  • No Web UI for approvals; CLI first per decision 2.

Design sketch (for discussion)

Write action lifecycle. The runner reaches a state that wants a write
(e.g. a failure event with comment-on-failure enabled) → builds a
WriteIntent (action, target repo+PR, scope, evidence digest) → quota
check first → PermissionCenter.request_or_grant(): a matching live grant
executes the write; without one, behavior splits by mode — auto (the
default) skips and records skipped: no grant, gated persists a pending
request and notifies. nanodot approvals approve ID (gated) or
nanodot approvals grant --action … --target … --scope … [--expiry …]
(auto's pre-grant, same grant schema #13 shipped) creates the grant; the
next poll re-attempts under it. Pending write requests carry a 4h TTL
(decision 6). One approved grant per
(action, target, scope) with expiry — the matching semantics #13 shipped
and tested.

Repeat-request policy (decided). After an unanswered request expires,
an identical request may be re-asked exactly once; a second silence records
a denial-by-silence (terminal, never re-asked verbatim). Explicit denials
keep the existing #13 semantics unchanged.

Write quota (decided). Checked before any request is created: per
action, per watch, per day — 3 comments, 3 labels, 1 approval per watch
per day, configurable per watch. An exhausted budget records a
quota-exhausted activity entry and stops attempting that action for the
day — same fail-closed degrade path as any provider unavailability. The
notification dedup fingerprint guards against duplicate writes; the quota
guards against volume (a flapping PR must not generate an unbounded
comment/approve stream).

Where the write executes. Proposed: the runner performs writes only
when a grant exists at poll time — no long-lived "resume a half-finished
run" machinery; a denied/expired request simply never writes and the task
stays unblocked-but-noted (writes are enhancements, never gates — the
memory-subsystem precedent from #12).

Egress contract. egress.md gains a third section: the write port —
destination, exact request shapes (comment body, label name), credential
header, and the invariant that a request can only be constructed from a
persisted grant id + the whitelisted evidence that justified it. The
redactor scrubs the comment body like any outbound value.

Token acquisition. nanodot config set github-write-token (hidden
prompt / stdin, same UX as github-token); validation at set-time checks
it is not the same value as the read token and probes the fine-grained
permissions read-only endpoint, failing closed on mismatch with the
first-action set.

Testing. The no-external-write invariant test is inverted, not
removed: the write port becomes the only construction site for POST
requests (grep/invariant test), and the offline e2e suite gains a full
approval lifecycle scenario (request → expiry → re-request denial path →
approve → write against a fake GitHub → audit replay).

Decisions (maintainer, 2026-10-03)

  1. First wave includes approve — comment + label + review approval;
    merge stays out.
  2. Dedicated write port — ports/writer.py; the GitHub read client is
    never widened.
  3. Silence-expiry re-ask — an expired unanswered request is re-asked
    exactly once; a second silence is a recorded denial-by-silence.
  4. Write quota is in — hard per-action/watch/day budget, fail-closed
    exhaustion, mirroring the RFC: multi-provider inference — third-party model API access #46 provider-budget pattern; without it a
    flapping watch could flood the PR and destabilize the system.
  5. auto activates and is the default mode — pre-granted scoped
    capabilities only, zero runtime prompting; gated becomes the opt-in
    interactive loop; readonly the explicit hard-off. Safety holds
    because a fresh install has no write token (every mode inert) and no
    grants (nothing to execute); pre-granting requires an explicit CLI act.
  6. Write-request TTL is 4 hours — shortened from the 24h schema
    default: a late-approved comment/approve on a fast-moving PR references
    a stale head SHA and is noise at best; 4h keeps the request alive for a
    working-session turnaround while the re-ask-once policy covers
    overnight silences. Per-action TTL differentiation (e.g. tighter for
    approve) stays available as a P2 refinement if experience demands it.
  7. Explicit permission-mode key, default auto — mode is never
    derived from github-write-token presence: presence-derived modes are
    surprising, and the token gates capability while the mode states
    posture. readonly remains selectable as the explicit hard-off.
  8. Quota defaults: 3 comments, 3 labels, 1 approval per watch per day,
    configurable per watch — approvals are heavier social signals than
    comments, hence the tighter cap.

Status

All open questions are resolved (decisions 1–8, 2026-10-03). The RFC is
converged; next step is breaking P0 (write port, invariant inversion,
token plumbing, quota counters) into implementation issues.

Invariants (must hold)

  • A write request can only be constructed with: a live grant id, the
    write token present in the secret store, a payload built from
    whitelisted evidence, and remaining quota. Absent any one, no code path
    reaches the transport with a non-GET method.
  • The read-only PAT never authorizes a write; the write token never
    authorizes a read path (separate secrets, separate headers, same
    no-redirect transport policy).
  • Silence is never approval; denials are recorded, never re-asked
    verbatim; a silence-expired request is re-asked at most once; scope
    changes revoke grants atomically (the first three inherited from Permissions: modes, scoped grants, approvals surface #13
    and must keep their tests).
  • The default mode is auto, and it never prompts: grant-less writes are
    skipped and recorded. Without a write token, every mode is behaviorally
    readonly — mode changes what nanodot attempts, the token cap holds
    regardless.
  • Every write decision is reconstructable from the activity log.
  • Watch correctness never depends on writes: a denied/expired/
    quota-exhausted write leaves the watch loop's read path, notifications,
    and state machine untouched. Write volume is bounded by quota — the
    system's stability never rides on write pressure.

Relationship to existing work

Suggested phasing

  1. P0: write port + invariant inversion + separated token plumbing +
    quota counters (no actions yet; readonly behavior identical).
  2. P1: comment/label/approve actions behind both mode paths — the
    auto default (pre-grant CLI, skip-and-record) and the gated loop
    (request → approve/deny CLI → execute → audit, including the
    silence-expiry re-ask path), offline e2e scenario, egress.md update.
  3. P2: TTL/quota refinements from production experience (e.g.
    per-action TTL, per-watch quota tuning), user-guide documentation.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions