Skip to content

History, incident records, fleet triage, autonomous watch and incident drills - #9

Open
abhijitkrm wants to merge 7 commits into
mainfrom
feat/history-incidents
Open

abhijitkrm wants to merge 7 commits into
mainfrom
feat/history-incidents

Conversation

@abhijitkrm

@abhijitkrm abhijitkrm commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Items 1–4 from the SRE gap list: cometcli now remembers, compares the whole fleet, watches it without you, and can be measured on real faults.

1. History and incident records

  • Every node.triage run is saved to the node's history (30 days). Triage opens with what changed since the last check, for example node.peers 8→0, val.jailed false→true.
  • node.history shows each signal's first → last value, min/max, a trend line and when it changed. mon.record samples between triages.
  • incident.record saves a finished incident as JSON plus a Markdown postmortem. The timeline and transaction hashes come from the audit log. There's also incident.list / incident.show, and triage flags recurring problems.
  • /incident and /recover-jail now work from general mode: name the node (/incident val01 is down) and cometcli switches to it, or the agent picks it.

2. Fleet triage

  • fleet.triage triages every node at once and compares them. It reports:
    • a node behind the others
    • a chain-wide halt
    • version mismatches
    • two nodes running one consensus key
    • validators sharing a host
    • a validator whose sentries are down (new sentries profile field)
  • It's offered in general mode, and the agent uses it for fleet questions.

3. cometcli watch

  • Sweeps the fleet on an interval and turns a newly appearing critical/high problem into an incident. Three modes:
    • notify
    • diagnose (default): the agent works it read-only and reports
    • fix: every approval goes to Telegram with Approve / Deny buttons; an unanswered approval is denied, and only presses from the configured chat and users count
  • Transactions are refused unless watch.allow_tx is set.
  • It announces recoveries, waits out a cooldown before working the same problem again, and keeps its state across restarts. Reports go to Slack, Discord and Telegram.

4. cometcli drill

  • drill testnet up --image <evmd> builds a throwaway local 4-validator network with fast slashing. drill run breaks a node for real (stopped until jailed, broken config, halt-height, OOM limit, network partition) and scores whether triage named the right cause and whether the agent actually fixed it, verified on the node. Then it puts the node back.
  • Approvals are automatic only on drill profiles.
  • First run on claude-haiku-4-5: fixed 5/5, triage 4/5. The miss came from stale log evidence. Fixing it made the re-run of that scenario 1/1, and the agent wrote a postmortem for every incident.

Found by the drills

  • Cosmos-EVM mempool: evmd refuses to start unless config.toml [mempool] type = "app", but init-files writes flood. The drill net now sets it. The knowledge-base case was rewritten from evidence (a live node disproved the old rule), and the log scanner now recognizes the error line.
  • Config errors left over in the logs no longer count as a current config problem.

Tests

Unit tests for history, incidents, fleet findings, the watcher lifecycle, Telegram approvals (fake API, including a stranger's press and a timeout), the drill builder and its safety guards, plus end-to-end triage tests on the fake chain. go test -race ./..., vet and lint pass.

🤖 Generated with Claude Code

abhijitkrm and others added 7 commits October 8, 2026 17:22
cometcli now remembers:
- every node.triage run is saved to the node's signal history
  (~/.cometcli/history/<profile>/, 30 days); triage opens with what
  changed since the last check (node.peers 8→1, val.jailed false→true)
- node.history: first→last, min/max, trend line and change times per
  signal over a window ('since when?', 'is it getting worse?')
- mon.record: a sampler that records every interval between triages
- incident.record saves a finished incident (root cause, case, evidence,
  actions, outcome) as JSON + a Markdown postmortem, with the timeline and
  tx hashes rebuilt from the audit log; incident.list / incident.show
- triage names the node's past incidents and flags recurring cases
- /incident and the node prompt use them; both are core tools

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
'/incident val01 is down' answered 'needs node mode — /one <profile>
first'. A node-only command now switches to the profile its arguments
name; with none (or several) named, it runs and the model picks the
profile with use_node. It only errors when there are no profiles. The
welcome tip and /one hint say you can just name the node.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Triages all profiles in parallel (recording each node's history) and
compares them: a node behind the others on its chain, a chain-wide halt
(every node stalled), version mismatches, two nodes running one
consensus key (double-sign risk), validators sharing a host, and
validators whose sentries are down (new profile field: sentries).
Offered in general mode (tools can now declare FleetWide); the general
prompt sends fleet questions to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sweeps the fleet on an interval (fleet triage). A newly appearing
critical/high case, or a halt / double-sign / cut-off fleet finding,
becomes an incident: notify (report it), diagnose (default: the agent
works it read-only and reports), or fix (the agent may act; every
approval goes to Telegram with Approve/Deny buttons, unanswered within
watch.approval_timeout = denied, presses only from the configured chat
and users). Transactions are refused unless watch.allow_tx. Recoveries
are announced; an issue isn't worked again within the cooldown; state
survives restarts. Reports go to Slack, Discord and Telegram.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- drill testnet up --image <evmd>: a local validator network (fast
  slashing, 60s votes) with drill profiles (metadata drill=true,
  container signer, fallback gRPC); testnet down removes everything
- drill run: per scenario inject a real fault (node stopped until
  jailed, broken config, halt-height, OOM limit, network partition),
  score whether triage named an expected root case, let the agent work
  it via /incident with approvals auto-granted only on drill profiles,
  verify on the node, then recover it (files, limits, network, unjail)
- scoreboard (triage x/y, fixed x/y, steps, tokens, time) and JSON
  results in ~/.cometcli/drills/

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The first drill network wouldn't start: evmd requires config.toml
[mempool] type = "app" (configs mismatch: EVM mempool enabled …), and
init-files writes "flood". The drill net now sets it. A live node also
disproved the old rule (type app + max-txs -1 runs fine on evmd 0.7), so
config.mempool_mismatch now fires for Cosmos-EVM on flood/nop only when
the node is failing (config errors in the logs, or not running); the
case is Cosmos-EVM only and the log scanner knows the error line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
First drill run on claude-haiku-4-5: fixed 5/5, triage 4/5. The miss:
config errors from the previous scenario were still in the log window,
so triage said cfg-parse-error on a halt-height fault. cfg-parse-error
now needs the files themselves to be broken (log-only config errors are
node-exit-config-error, which requires a failing process). Re-run:
triage 1/1, fixed 1/1.

- drill results kept the scenario time and recovery error (named result)
- incident records: numbered actions aren't double-numbered; the
  timeline shows '/incident <what was asked>' instead of the expanded
  procedure
- docs: fleet triage and sentries, cometcli watch (modes, config,
  Telegram approvals), cometcli drill; README overview
- lint

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@abhijitkrm abhijitkrm changed the title Signal history and incident records History, incident records, fleet triage, autonomous watch and incident drills Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant