Repository navigation
History, incident records, fleet triage, autonomous watch and incident drills - #9
Open
abhijitkrm wants to merge 7 commits into
Open
abhijitkrm wants to merge 7 commits into
abhijitkrm wants to merge 7 commits into
Conversation
cometcli now remembers:
- every node.triage run is saved to the node's signal history
(~/.cometcli/history/<profile>/, 30 days); triage opens with what
changed since the last check (node.peers 8→1, val.jailed false→true)
- node.history: first→last, min/max, trend line and change times per
signal over a window ('since when?', 'is it getting worse?')
- mon.record: a sampler that records every interval between triages
- incident.record saves a finished incident (root cause, case, evidence,
actions, outcome) as JSON + a Markdown postmortem, with the timeline and
tx hashes rebuilt from the audit log; incident.list / incident.show
- triage names the node's past incidents and flags recurring cases
- /incident and the node prompt use them; both are core tools
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
'/incident val01 is down' answered 'needs node mode — /one <profile> first'. A node-only command now switches to the profile its arguments name; with none (or several) named, it runs and the model picks the profile with use_node. It only errors when there are no profiles. The welcome tip and /one hint say you can just name the node. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Triages all profiles in parallel (recording each node's history) and compares them: a node behind the others on its chain, a chain-wide halt (every node stalled), version mismatches, two nodes running one consensus key (double-sign risk), validators sharing a host, and validators whose sentries are down (new profile field: sentries). Offered in general mode (tools can now declare FleetWide); the general prompt sends fleet questions to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sweeps the fleet on an interval (fleet triage). A newly appearing critical/high case, or a halt / double-sign / cut-off fleet finding, becomes an incident: notify (report it), diagnose (default: the agent works it read-only and reports), or fix (the agent may act; every approval goes to Telegram with Approve/Deny buttons, unanswered within watch.approval_timeout = denied, presses only from the configured chat and users). Transactions are refused unless watch.allow_tx. Recoveries are announced; an issue isn't worked again within the cooldown; state survives restarts. Reports go to Slack, Discord and Telegram. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- drill testnet up --image <evmd>: a local validator network (fast slashing, 60s votes) with drill profiles (metadata drill=true, container signer, fallback gRPC); testnet down removes everything - drill run: per scenario inject a real fault (node stopped until jailed, broken config, halt-height, OOM limit, network partition), score whether triage named an expected root case, let the agent work it via /incident with approvals auto-granted only on drill profiles, verify on the node, then recover it (files, limits, network, unjail) - scoreboard (triage x/y, fixed x/y, steps, tokens, time) and JSON results in ~/.cometcli/drills/ Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The first drill network wouldn't start: evmd requires config.toml [mempool] type = "app" (configs mismatch: EVM mempool enabled …), and init-files writes "flood". The drill net now sets it. A live node also disproved the old rule (type app + max-txs -1 runs fine on evmd 0.7), so config.mempool_mismatch now fires for Cosmos-EVM on flood/nop only when the node is failing (config errors in the logs, or not running); the case is Cosmos-EVM only and the log scanner knows the error line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
First drill run on claude-haiku-4-5: fixed 5/5, triage 4/5. The miss: config errors from the previous scenario were still in the log window, so triage said cfg-parse-error on a halt-height fault. cfg-parse-error now needs the files themselves to be broken (log-only config errors are node-exit-config-error, which requires a failing process). Re-run: triage 1/1, fixed 1/1. - drill results kept the scenario time and recovery error (named result) - incident records: numbered actions aren't double-numbered; the timeline shows '/incident <what was asked>' instead of the expanded procedure - docs: fleet triage and sentries, cometcli watch (modes, config, Telegram approvals), cometcli drill; README overview - lint Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Items 1–4 from the SRE gap list: cometcli now remembers, compares the whole fleet, watches it without you, and can be measured on real faults.
1. History and incident records
node.triagerun is saved to the node's history (30 days). Triage opens with what changed since the last check, for examplenode.peers 8→0, val.jailed false→true.node.historyshows each signal's first → last value, min/max, a trend line and when it changed.mon.recordsamples between triages.incident.recordsaves a finished incident as JSON plus a Markdown postmortem. The timeline and transaction hashes come from the audit log. There's alsoincident.list/incident.show, and triage flags recurring problems./incidentand/recover-jailnow work from general mode: name the node (/incident val01 is down) and cometcli switches to it, or the agent picks it.2. Fleet triage
fleet.triagetriages every node at once and compares them. It reports:sentriesprofile field)3.
cometcli watchnotifydiagnose(default): the agent works it read-only and reportsfix: every approval goes to Telegram with Approve / Deny buttons; an unanswered approval is denied, and only presses from the configured chat and users countwatch.allow_txis set.4.
cometcli drilldrill testnet up --image <evmd>builds a throwaway local 4-validator network with fast slashing.drill runbreaks a node for real (stopped until jailed, broken config, halt-height, OOM limit, network partition) and scores whether triage named the right cause and whether the agent actually fixed it, verified on the node. Then it puts the node back.Found by the drills
[mempool] type = "app", butinit-fileswritesflood. The drill net now sets it. The knowledge-base case was rewritten from evidence (a live node disproved the old rule), and the log scanner now recognizes the error line.Tests
Unit tests for history, incidents, fleet findings, the watcher lifecycle, Telegram approvals (fake API, including a stranger's press and a timeout), the drill builder and its safety guards, plus end-to-end triage tests on the fake chain.
go test -race ./..., vet and lint pass.🤖 Generated with Claude Code