Skip to content

Rework cometcli into a Claude Code–style SRE terminal with validator incident handling - #3

Merged
abhijitkrm merged 48 commits into
mainfrom
feat/phase1-agent-core
Oct 7, 2026
Merged

abhijitkrm merged 48 commits into
mainfrom
feat/phase1-agent-core

Conversation

@abhijitkrm

Copy link
Copy Markdown
Owner

Reworks cometcli from a validator dashboard plus agent into a Claude Code–style SRE terminal. One engine runs in two scopes: cometcli on your own machine, and cometcli one <profile> for a node. It handles validator incidents from finding the problem to fixing it and checking the fix.

Agent and terminal

  • Two scopes, one engine. General mode gives the agent shell, file and web tools on your machine. Node mode adds the Cosmos tools, the validator system prompt and a live snapshot of the node. The agent switches to a node by itself when a question needs one (use_node), and you can switch with /one.
  • Works across providers. Anthropic, OpenAI, OpenRouter (including the free router), Groq, Gemini and OpenAI-compatible servers. Covers retries, prompt caching, streamed thinking, effort levels only where the model supports them, compaction, saved and resumable sessions, -p with text / json / stream-json output, and piped input.
  • Permission engine. Allow / ask / deny rules in Claude Code syntax, with ops, accept-edits, read-only and bypass modes. Transactions always ask, and key material and state resets are always refused. A shell classifier inspects commands, pipelines, scripts and docker exec before they run.
  • Extensibility. settings.json at user, project and local level, COMET.md memory, custom slash commands, hooks, MCP servers and subagents.
  • Inline terminal UI in the style of Claude Code. Startup art, ⏺ tool calls with output collapsed (ctrl+r expands it), collapsed thinking (ctrl+o shows it), approval dialogs, a slash-command menu, @ file mentions, ! to run a command in your own terminal, and messages typed mid-turn that reach the model at its next step (Esc sends them immediately).

Incidents

  • node.triage collects about 50 signals in parallel from the node, chain, governance, host, process, logs, configs and EVM endpoint. A source that's down is reported, not fatal.
  • Knowledge base of 77 cases for Cosmos SDK, CometBFT and Cosmos-EVM. Each has confirm, cause, fix, verify and never-do steps, plus cause → symptom links so the root cause ranks first. Every case's sample signals are tested to rank it in the top 3, and a healthy node matches nothing. You add your own with kb.add.
  • /incident runs the general procedure: triage, then the top case, then fix the root cause, wait, verify and report. /recover-jail is the ready-made procedure for a jailed validator.
  • Unjail checks before spending fees. Unjail is refused when: the jail period isn't over, the node isn't synced, the signer's saved state is ahead of the chain, the node runs a different consensus key than the one registered, self-delegation is below the minimum, or there's no fee balance.
  • Two signers. cometcli's own keyring, or the node container's keyring (the key never leaves the node), chosen when you sign. A local key is used only if it belongs to the validator's operator account.
  • Governance. gov.propose (text, software upgrade, arbitrary messages) and gov.deposit. The upgrade height is checked to be reachable within the voting period, and deposits are checked against the chain's minimums.

Fixes found along the way

  • Transactions were always signed with sequence 0.
  • Coin amounts, votes and validator edits are now validated.
  • Peer edits and statesync config patching rewritten; upgrade staging now verifies the binary.
  • fleet.shell went around the shell checks; it now goes through them.
  • Key management was callable by the agent; it's now operator-only.
  • The EVM client read malformed numbers as 0.
  • The CLI applied a flat 90s deadline even to long-running tools.

Testing

  • Unit and end-to-end tests run with the race detector in CI. They include a simulated chain (fake gRPC node, fake CometBFT RPC and fake docker) that runs a full jail recovery through the real agent loop.

  • Container signing was validated against the real primium-evm:v0.7.2 image.

  • Live drills on a throwaway 4-validator testnet with an extra RPC node:

    • downtime jail
    • crash loop from a broken config
    • OOM kill inside the container
    • network partition
    • halt height left in config
    • signing state ahead of the chain
    • wrong consensus key
    • full data disk
    • chain-wide upgrade halt and coordinated skip
    • governance proposal lifecycle
    • unjail attempted too early

    Each drill turned up bugs, and each fix has a regression test. The agent resolved the config and OOM incidents end to end on a free model.

  • The UI was checked in a real pseudo-terminal at 70 and 110 columns, not only with snapshot tests.

Not included

  • The serve web chat and the macOS app are frozen: they still compile but weren't extended.
  • Genesis-ceremony tooling.

Docs: docs/COMMANDS.md (incidents, triage, signing, governance) and docs/EXTENDING.md (settings, hooks, MCP, knowledge-base case format).

🤖 Generated with Claude Code

abhijitkrm and others added 30 commits October 6, 2026 16:32
- retries: 408/409/429/5xx/529 and connection errors with jittered
  backoff honoring Retry-After; announced as notices, timeouts not
  retried; --debug logs every API attempt
- normalized stop reasons: auto-continue after max_tokens, refusals
  surface as errors
- frozen per-session system prompt; later node snapshots ride on user
  turns so provider prompt caches stay warm
- token usage per provider (incl. cache reads), /cost
- compaction: automatic at a turn boundary near the context limit,
  on context-overflow errors, and /compact [focus]
- effort (--effort, /effort, agent.effort) mapped to Anthropic
  output_config.effort, OpenAI/Groq reasoning_effort, Gemini
  thinking_level; streamed reasoning shown as dimmed "∴" lines
- saved sessions in ~/.cometcli/sessions: -c/--continue, -r/--resume,
  cometcli sessions, /sessions, /resume
- Anthropic: verbatim replay of thinking blocks (dropped when the tool
  set changes), top-level prompt caching, eager tool input streaming,
  server-side refusal fallback, default model claude-opus-5-5
- head+tail, UTF-8-safe tool output truncation; max turns 16 -> 50

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Tools (agent-only: no CLI subcommand, not exported over MCP), all over
the profile's host transport so they work locally and over SSH:
- bash: persistent working directory, combined output capped head+tail
  at 30k, timeout (default 2m, max 10m) that kills the process group,
  stdin closed with non-interactive env
- read / write / edit: line-numbered reads, edits require a fresh read
  and show a diff for approval, file mode preserved
- glob / grep: newest-first file search; ripgrep or grep -E, with key
  files and keyring dirs excluded
- web_fetch: HTML reduced to text, each new domain approved, localhost
  and private addresses fetched directly, cross-host redirects reported
- todo_write: a visible checklist (EvTodos), saved with the session

Shell classification: every command line is lexed (quotes, &&/||/;/|,
redirects, $(…), heredocs, sudo/env/xargs/timeout wrappers, bash -c,
docker exec/run) and each simple command classified. Reads auto-run;
changes ask; transactions (evmd tx <module> anywhere, cast send,
eth_sendRawTransaction, or a script whose body does so) always ask;
consensus keys, mnemonics, key ceremonies and state resets are refused.

Permissions: allow/ask/deny rules in Claude Code syntax (bash(cmd:*),
read(./x/**), web_fetch(domain:x), node.*) from agent.permissions,
--allowedTools/--disallowedTools and /allow /ask /deny; deny > ask >
allow, and an allow must cover every segment of a compound command.
Modes: ops, accept-edits, readonly, bypass (--permission-mode, /mode).

Also: Context.Host falls back to the local machine without a profile;
clearer denial when there is no terminal to approve on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every request sent all 65 tool schemas (~6.5k tokens), over Groq's free
8k tokens-per-minute budget on its own. Requests now carry the general
tools plus node.status/health/logs and val.status/signing, the frozen
system prompt lists the rest by name, and the built-in tool_search
loads definitions on demand (select:<names> or keywords). Loaded tools
persist for the session and are saved with it. agent.tools: all
restores the old behavior. First request: ~30KB -> ~10KB.

Groq's 413 "request too large" now triggers compaction like a
context-window overflow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- bare `cometcli [prompt]` is general mode: an SRE agent on this
  machine with the general tools and a general system prompt (OS, host,
  cwd, available node profiles); no profile needed. An explicit
  --profile means node mode.
- `cometcli one [profile] [prompt]`: node mode — node tools (deferred),
  validator prompt, live snapshot, shell on the node's host. /one
  <profile> and /one off switch scope mid-session, keeping the
  conversation (replay state dropped, the model told on the next turn).
- -p/--print: answer once and exit; --output-format text|json|
  stream-json, --include-partial-messages, --verbose, --append-system-
  prompt, --max-turns. Headless runs never read approvals from stdin.
- piped stdin (pipes and redirected files only) is attached as <stdin>
  context, capped at 200k with head+tail kept.
- global `agent:` section in config.yaml, overridden per profile
  (AgentFor/MergeAgent); general mode falls back to the active profile's
  settings. `cometcli config show|set|unset agent.<key>`.
- single-word typos of subcommands (doctr) are reported, not sent.
- the loop streams whenever the provider can, rendered or not.
- ask/agent/ui unchanged (node mode on the active profile).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gpt-oss on Groq answered "summarize docker container health" with a
generic explainer and no tool calls. Both system prompts now say that
questions concern this machine/node and must be answered from evidence
gathered with tools, generic explanations only on request.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ts (phase 4)

Claude Code-compatible formats throughout:
- settings.json at user / project / local scope: model, effort,
  permissions (allow/ask/deny/defaultMode), env (exported before the
  provider reads keys), hooks, mcpServers; plus a project .mcp.json.
- trust: project hooks and MCP servers run commands, so they are
  ignored (with a notice) until `cometcli trust`; same for a project
  command's allowed-tools.
- COMET.md memory: user, project root down to cwd (AGENTS.md fallback,
  COMET.local.md), per-node memory in node mode, @imports (key material
  refused); /memory, /remember appends and tells the model at once.
- custom slash commands from .cometcli/commands/**.md: frontmatter
  (description, argument-hint, allowed-tools as one-turn allow rules),
  $ARGUMENTS/$1..$9, !`cmd` inlined for read-only commands only, @file;
  listed in /help, usable with -p.
- hooks: PreToolUse (exit 2 / permissionDecision deny|ask|allow — allow
  never skips a tx prompt), PostToolUse (feedback to the model),
  UserPromptSubmit (block / add context), Stop (keep working, max 3),
  SessionStart; stdin JSON contract, process-group kill on timeout.
- MCP client: stdio and streamable HTTP (SSE responses, session ids),
  pagination, ${VAR} expansion, server-request rejection; tools become
  mcp__server__tool, deferred behind tool_search, readOnlyHint runs
  freely, others ask, "mcp__server" rules cover a server.
  `cometcli mcp add|list [--check]|remove` (user/project/local scope);
  bare `cometcli mcp` still serves.
- subagents: a task tool running a fresh-context child agent (same
  tools/permissions, no recursion, events shown with ↳), plus custom
  agents from .cometcli/agents/*.md with tool restrictions.
- docs/EXTENDING.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mmands

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…se 5)

New internal/tui/chatui, used by bare `cometcli` and `cometcli one`:
- inline (no alt screen): finished messages print into scrollback;
  only streaming text, spinner, input, dialogs and the status line redraw
- banner, ● tool lines with ⎿ results (diffs colored), markdown answers,
  todo checklists, last answer shown on resume, resume hint on exit
- input: enter sends, \⏎ / ctrl+j newline, history, / command menu
  (builtins + custom commands), @file completion with content attached,
  # adds memory, messages typed during a turn are queued
- !cmd runs in the operator's real terminal via ExecProcess (prompts and
  hidden password input work, e.g. vote-upgrade.sh); output never
  reaches the model, which is told the command and exit code; over
  `ssh -t` in node mode with an SSH transport
- approvals: Yes / Yes, don't ask again for <suggested rule> (session +
  .cometcli/settings.local.json or user settings) / No (declines and
  interrupts); transactions offer only Yes/No
- shift+tab cycles ops → accept-edits → read-only (→ bypass when started
  with it); esc interrupts; ctrl+c clear/interrupt/double-exit; ctrl+d
- no terminal background query (dark unless COLORFGBG says light);
  zero-size pty reports ignored

Supporting: toolkit.SuggestRule + RuleHint on approval details (never
for transactions), Context.ToolName, Event.Output previews, /status and
/clear, agent.Commands/ExpandMentions/AddNote/AllowAlways/ProjectFiles/
ContextPercent; the y/N approver hides "_" hint keys. `cometcli ui`
keeps the full-screen dashboard.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Options.ForceSeq was documented as ">=0 overrides the queried sequence",
but its zero value is 0 and no caller set -1, so Build always signed
with sequence 0. Ordinary transactions only worked because simulation
failed with "expected N, got 0" and the single automatic retry used the
right sequence — costing an extra round trip and a spurious "sequence
drifted" message on every tx. Two cases were genuinely broken:
offline signing with --seq always signed sequence 0 (pinned accounts
are never retried), and a real sequence drift could not heal because
the bug had already spent the one retry.

The override is now an unexported pointer set via ForceSequence(n), so
the zero Options means "use the chain's sequence".

Found by the new fakenode test harness (internal/testutil/fakenode): an
in-process cosmos gRPC node that decodes every broadcast, verifies the
secp256k1 signature over the SIGN_MODE_DIRECT sign doc (keccak for
eth_secp256k1, sha256 for secp256k1) and enforces sequences with the
SDK's error text. New tests cover both key types, fee/gas math, the
approval gate (denied txs never reach the node; on-chain is never
auto-approved), drift healing with re-approval, pinned sequences,
CheckTx rejection, in-block failure, unconfirmed-as-pending, wrong
chain-id, and DecScaled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tatus

Found by new tests running the tx tools against fakenode:
- parseCoin accepted "1.5adex" (decimal amount), "1e18adex" (1 unit of
  denom "e18adex"), "10 adex", "0adex" and one-letter denoms. Amounts
  are now positive integers in base units with an SDK-valid denom, with
  specific messages for decimals and scientific notation.
- val.vote broadcast VOTE_OPTION_UNSPECIFIED for an unknown or
  upper-case option and proposal 0 when the id was missing; options are
  now case-insensitive and validated, the id required.
- val.edit with nothing to change broadcast a no-op edit (paying fees);
  it now refuses.
- val.status dereferenced Commission/Description without nil checks.

Tests: commission scaling + description preservation on edit, unjail,
withdraw with commission, vote options, send message contents, input
validation never reaching the node.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
net.add-peer / net.rm-peer:
- anchor on the `persistent_peers = "…"` line itself (a comment
  mentioning persistent_peers matched first and the edit failed; the
  longer persistent_peers_max_dial_period key was also a risk)
- validate peers as <40-hex node id>@host:port before asking
- removing an absent peer (or re-adding a present one) changes nothing
  and asks nothing, instead of rewriting the file and saying "removed"
- keep config.toml's file mode (writes reset it to 0644)

upgrade.prepare:
- find the binary anywhere in the archive (releases nest it under bin/)
- accept the checksum of the archive or of the binary — releases
  publish either; a correct archive checksum used to fail
- no checksum: the approval carries a warning and the result says
  UNVERIFIED with verified=false
- work dir cleaned up on exit

Tests for both, against a temp node home on the local host and a test
HTTP server serving real tarballs (three layouts × two checksum kinds,
mismatch, missing binary).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
snap.statesync --apply reported success but wrote nothing when config.toml
had no [statesync] section, never added missing keys when [statesync] was
the last section, matched keys by prefix (enable vs enable_*), and reset
the file mode to 0644. It now edits exact keys in place, appends missing
keys or the whole section, and keeps the mode.

Audit log tests: concurrent sessions sharing one file keep every JSONL line
intact, 0600/0700 permissions, session lookup across daily files (bad lines
skipped), first prompts per session, nil logger no-op.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…less resume note

E2E tests run the real root command against a scripted OpenAI-compatible
server: -p text/json/stream-json, general-mode tools and prompt, a real
bash call, headless denial of a write (and --allowedTools allowing it),
sessions/-c/-r, piped stdin, slash and custom commands in -p, config set,
typo detection, one without a profile.

Bugs found:
- the openai-compat provider (Ollama/vLLM/llama.cpp) had no name, so it
  reported itself as "openai" and was treated as hosted: it got
  max_completion_tokens and reasoning_effort, fields local servers may
  reject. It now identifies as openai-compat.
- cometcli -p -r <id> resumed silently; the resume note goes to stderr.
- a retry test asserted on timing; made robust.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fleet.shell ran any command on every host after one generic local-change
approval, bypassing the shell classifier: key material was not refused
(cat priv_validator_key.json across the fleet would reach the model),
an `evmd tx` was not treated as a transaction (in bypass mode it ran
without asking), and permission rules didn't apply. It now goes through
shell.Classify + Context.Check like bash; read-only commands still run
without a prompt, and a fleet-wide script run always asks (remote
scripts can't be inspected).

keys.add / keys.rm were agent-callable: a generated mnemonic went into
the tool result (scrubbed by redaction, so the operator never saw the
only copy) and --recover read stdin directly. New toolkit.OperatorOnly:
such tools are never advertised to the model, refused if called, hidden
from tool_search and MCP, and refused in runbook steps the agent runs.
They stay CLI subcommands.

Runbooks now validate every step before running any (an unknown tool in
step 5 surfaced after steps 1-4 ran). keys.convert validates 0x input.

Tests: fleet (reads across hosts, key material refused, tx and scripts
ask in bypass, rules), fleet.exec refusing mutating tools, keys round
trip + operator-only flags, runbook steps/optional/manual/validation,
host Local exec (combined output, closed stdin, non-interactive env,
head+tail cap, timeout killing children, file mode), agent refusing
operator-only tools.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Golden files (internal/tui/chatui/testdata, regenerate with -update) for the
idle prompt, slash menu, bash and transaction approval dialogs, a running
turn, the accept-edits footer and a transcript of tool calls, diffs,
subagent results, todos and markdown — layout changes show up as diffs.

CI runs go test -race with a coverage summary. Fixed a data race in the
approval test (shared variable polled across goroutines → channel).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
hexToUint64 returned 0 on a parse failure, so a malformed eth_blockNumber
read as block 0 and evm.parity / the monitor reported a huge false drift.
Quantities must be 0x-hex or the call fails; non-2xx responses (a proxy's
502 page) report the status instead of 'invalid character <'.

Tests: quantities, syncing false vs progress object, txpool, client
version, decimal-string rejection, RPC errors, HTTP errors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… at signing time

Transactions are now prepared → approved → signed → broadcast, with two
signers behind tx.Signer:
- LocalSigner: cometcli's ops keyring (previous behavior).
- ContainerSigner: the node's own keyring inside its container — the
  key never leaves the node. cometcli builds the unsigned tx (the
  --generate-only JSON, passed base64 in the command so stdin carries
  only a password), runs `evmd tx sign --offline --sign-mode direct` and
  `evmd tx encode` via `docker exec` (locally or over SSH), and
  broadcasts the result. Gas is simulated with the account's on-chain
  public key; without one it falls back to 300000 and the approval says
  so. A file keyring's password is asked for after approval only, never
  stored or sent to the model; wrong passwords, missing keys and stopped
  containers get clear errors.

Choosing: signer.mode local|container, else the only one available,
else the operator picks when both can sign (CLI numbered prompt, TUI
dialog); headless runs must set signer.mode. Keyring backend and key
address are detected when not configured (test keyring probed, else
file; address from metadata account/valoper or `keys show`).

New Context hooks: Chooser and Secret (CLI: shared stdin reader, /dev/tty
with echo off; TUI: choice dialog and masked input; headless: env
COMETCLI_CONTAINER_KEYRING_PASSWORD), carried through agent tool
contexts and scope switches. Host exec gained stdin (ExecIn).

Tests: a fake `docker` (the test binary on PATH) emulates evmd keys
show / tx sign / tx encode with real SIGN_MODE_DIRECT signing, verified
by fakenode: test and file keyrings, simulation with an on-chain pubkey
vs default gas, no password prompt for a declined tx, wrong password,
signer choice / headless refusal / signer.mode.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-flight

For running an incident to resolution (e.g. jailed → recovered):
- val.jail-check: jail reason (downtime / tombstoned) and when; evidence
  from the node's own logs in the window before the jailing (the
  x/slashing jail line, then peer loss, signer/privval errors, panics,
  OOM, disk, app-hash, clock patterns with counts and samples), container
  or systemd restarts/OOM kills, disk; then every unjail blocker with its
  fix: jail period not over, node not synced (it would be re-jailed),
  self-delegation below minimum, fee balance, tombstoned.
- val.consensus: bonded and unjailed, in the CometBFT validator set with
  voting power, our signature present in the last N commits, node synced.
- wait.until: synced | height | unjailable | bonded | signing |
  in-consensus | tx-committed | proposal-status — polled in-process
  (no model tokens while waiting), live progress (sync rate and ETA),
  done=false with the last state on timeout, cancellable. Also a CLI
  command: cometcli wait until --condition synced.
- val.unjail refuses on hard blockers before paying a fee; every tx
  failure carries a cause → next-step hint (still jailed, not jailed,
  self-delegation, fees, funds, gas, wrong key, chain-id, mempool).
- Progress hook on tool contexts → EvProgress; the TUI shows it as a
  live line under the spinner.
- jail-check, consensus and wait.until join the always-loaded core tools.

Test harness: fakenode gains a fake CometBFT JSON-RPC chain (status,
net_info, health, commit, validators) whose height, sync state, block
times and our signatures tests drive; delegation, bank balance and
downtime_jail_duration on the gRPC side; fake docker serves logs/inspect.
Fixed in wait.until while testing: a poll interval longer than the
remaining time ended the wait after one check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ulated chain

/recover-jail (built-in, node mode; a project/user command of the same
name overrides it) gives the agent the procedure: val.jail-check →
stop if tombstoned → fix the root cause from the evidence → wait.until
synced → wait.until unjailable → val.unjail (operator approves; signer
chosen at signing time) → on failure read the cause/hint, fix, retry
(max 3) → wait.until in-consensus (50 blocks, 95%) → val.consensus →
report with evidence and tx hashes. Built-ins show in /help and the
slash menu.

internal/scenario runs the whole incident with the real agent loop and
tool registry against a simulated chain that evolves while the agent
works: jailed for downtime with peer loss in the node's logs, node
behind and catching up, jail period ending in chain time, the first
unjail rejected at CheckTx for an insufficient fee (node's min gas price
raised), retry with a higher gas price, the chain re-admitting the
validator, signatures resuming. A scripted model reacts only to tool
output; the test checks the procedure order, the evidence, one rejection
+ one broadcast, two approvals, signer asked per attempt, and the final
"participating in consensus" (20/20 signed). Deterministic: the node
starts syncing when the agent starts waiting.

fakenode: CheckFailOnce and OnCommit hooks. Docs: incidents, signing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Running the container signer against primium-evm:v0.7.2 (a throwaway
container with fresh test keys) found that the image has no mktemp; the
sign script now uses per-process files under /tmp with umask 077.
With that, evmd accepts cometcli's unsigned tx JSON, signs it in both a
test and a password-protected file keyring, and the result verifies.

fakenode now accepts the 65-byte r||s||v signatures evmd produces for
eth_secp256k1 keys (the chain's ethsecp256k1 verifier drops v), as well
as 64-byte ones.

The validation is kept as a build-tagged test (realdocker) with the
container setup in its header.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
node.triage collects ~50 flat signals in parallel (comet, chain, gov,
host, process, logs, config, evm) with per-source timeouts and matches
them against a YAML knowledge base of 73 Cosmos SDK / CometBFT /
Cosmos-EVM cases (confirm, causes, tagged fix steps, verify, never-do).
Cases are embedded, extendable per user/project, and each carries a
fixture the test suite proves ranks it in the top 3; a healthy node
matches nothing.

- internal/kb: case schema, condition language, matcher, search
- internal/logscan: shared log categories (val forensics uses it too)
- kb.search / kb.show / kb.add tools; wait.until condition=signal
- /incident built-in and a generic incident method in the node prompt
- toolkit.Context.Derive for fan-out with shared clients

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- logscan: strip ANSI colors; count info/debug lines only for categories
  where they are the evidence (jail, double sign, upgrade halt); tighter
  app-hash pattern (a healthy handshake logs appHash=…)
- endpoints.fallback_grpc: chain queries go through another node while
  this one is down, so a dead validator's jail state is still visible;
  failed gRPC dials are no longer cached for the context's lifetime
- kb: advisory kind (hardening) listed separately; severity ranks first
- jail-check: log-gap detection (docker logs -t) shows when the node was
  silent and its last line — stopped from outside vs crashed; docker
  lifecycle events when still in the daemon's buffer

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- agent.provider openrouter: OpenAI-compatible, default model
  openrouter/free, key OPENROUTER_API_KEY; always requests the reasoning
  stream (reasoning.enabled) so thinking shows live
- reasoning_content deltas accepted (DeepSeek/vLLM-style servers)
- cometcli config set-key NAME: saves an API key from stdin to
  ~/.cometcli/credentials (0600); env vars still win
- logscan: skip cobra --help output, startup Error: lines, last error
  signal, jail lines filtered to our own validator
- kb: explains links — root causes rank first, symptoms grouped

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tx/block/app hashes are public and the operator needs them in reports;
unlabelled 64-hex still redacts (key-shaped). The agent guessed config
paths in the live drill — triage now shows config.file / app.file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- triage: container memory limit, usage vs limit, cgroup oom_kill count
  (docker's OOMKilled only covers PID 1; nodes run under wrapper shells)
- kb: container-memory-limit case (root of crash loops / jails)
- logscan: a signer refusing a height/round/step regression is the
  double-sign guard working, not a privval fault
- /incident + node prompt: apply [change]/[tx] steps by calling the tool
  (its prompt is the approval) instead of stopping to ask in text — the
  live drill showed a model stop at 'requires your approval'

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- triage: inside-the-container RPC probe and attached-network count — a
  detached or unreachable node is no longer mistaken for a dead one;
  kb cases node-network-detached, node-rpc-path-broken
- triage logs start 5m before the current process start (older errors
  are history and produced stale root causes)
- retries: a daily/credit-quota 429 fails immediately with its message
- docs: signer is chosen first, then the tx approved

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eouts

- val: SignerAhead — priv_validator_state ahead of the chain is a hard
  unjail blocker (the live drill unjailed a validator that couldn't sign
  yet and it was jailed again); fix says wait, never reset
- cli: tools run with their own Timeout() (wait.until, triage) instead of
  a flat 90s deadline
- common.ReadNodeFile shared by triage and val
- kb: 'no SignBytes found' documented under privval errors

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
abhijitkrm and others added 18 commits October 7, 2026 12:06
The vote drill died at 'exit -1': under a loaded Docker VM (loadavg 25+)
the container signer's evmd calls exceeded the flat 90s call deadline,
which also counts the operator reading the approval prompt.
toolkit.CallTimeout is shared by the CLI and the agent. Triage reads the
docker host's load from inside the container (Docker Desktop's VM is
otherwise invisible from macOS).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A just-restarted node dials fine but answers 'invalid height' while it
loads state — exactly when an operator tries to unjail. Gather retries
through endpoints.fallback_grpc, and without one says what's happening.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- docker nodes: df of <container home>/data (signer.container_home or
  the mount of the profile's home) — the disk that fills can be its own
  volume; the live disk-full drill showed the host disk at 77% while the
  node's volume was full
- jail pattern no longer matches store keys (store_key=…slashing)
- node-rpc-path-broken only after 30s uptime (probes raced a restart)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The vote drill offered tn0's 'ops' key on tn1..3 (shared keyring, same
signer.key): choosing it would have signed a valid tx as another
account. With metadata.valoper set, a local key must match its account.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- jail lines filtered by our valoper too (slash events name valopers)
- node-crash-loop needs recent restarts (uptime < 10m); restart counts
  are lifetime-cumulative
- upgrade-halt case: matches on the plan height too; documents the
  coordinated --unsafe-skip-upgrades path
- container signer surfaces cobra's 'Error:' line instead of --help tail
- signer choice is single-sourced: TxSigner is cached per context and
  common.Account uses it, so a message's address and its signature can't
  come from different keys (seen live: local ops key vs container node2)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Seen live: an RPC node syncing from genesis crash-looped at the upgrade
height the validators had skipped. Triage no longer runs validator
queries for rpc/sentry profiles (it only collects the upgrade plan).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- gov.propose: text, software upgrade and arbitrary-message proposals;
  deposit defaults to the chain minimum and is checked against both
  min_initial_deposit_ratio and min_deposit_ratio before approval; an
  upgrade height the vote can't reach in time (voting period at the
  measured block time) is refused; reports the new proposal id
- gov.deposit tops up exactly what's missing to the minimum
- nested messages packed with Cosmos type URLs ('/cosmos.…', not
  type.googleapis.com/ — chains rejected the latter, seen live)
- val.vote says when a proposal is closed or still in deposit instead of
  surfacing 'inactive proposal' from simulation
- chain gov shows the deposit-period end for proposals awaiting deposit

Drilled live: propose → 4 votes → PASSED; deposit top-up moved a
proposal into voting; upgrade and community-pool-spend proposals
submitted through the container signer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Checked against the real terminal (pty + emulator), not just goldens:
- welcome box 'Welcome to cometcli!' at a fixed width (long paths lose
  their middle instead of wrapping the box) + tips; printed after the
  first window size so it fits the terminal
- tool calls: a blinking ⏺ with '⎿ Running…' while executing (paused
  'Waiting for your answer…' during dialogs), then a green/red ⏺ with
  output collapsed to 3 lines, '… +N lines (ctrl+r to expand)'; Read
  shows 'Read N lines'
- thinking: the spinner says Thinking…, then '∴ Thought for Ns (ctrl+o
  to show thinking)'; ctrl+o prints it
- streaming text rendered as markdown live; '- ' bullets, inline code
  colored without padding, headings bold
- spinner verb varies per turn; '>' prompt; placeholder only before the
  first message; Interrupted · What should cometcli do instead?
- every printed line wrapped to the terminal width with a hanging
  indent (terminal-wrapped lines corrupted the live region redraw)
- /status shows the current setup in any mode (model, mode, scope, cwd,
  session, memory, MCP), plus the node snapshot in node mode; /status
  and /memory no longer build the node prompt (it probed the node and
  froze the UI when the node was down)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Like Claude Code's block-letter greeting: COMETCLI in the ANSI Shadow
font with a gold → coral → violet gradient, under a comet whose tail
flares from a glowing nucleus through cyan and blue into violet dust.
Shown on fresh sessions when the terminal is wide enough (≥ 74 cols);
narrower terminals get the welcome box alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Haiku 4.5 (and Sonnet 4.5) reject output_config.effort with a 400,
which failed every request once agent.effort was set. Effort now goes
only to Opus 4.5+, Sonnet 4.6+ and the 5 family; if a model still
rejects it, the request is retried without it and effort is dropped for
the rest of the session.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Asked 'is anyone jailed?' in general mode, the model ran 'which comet'
and ls, then asked the operator to type /one — the general prompt told
it to. Now:
- use_node(profile|off): a built-in, read-only tool that moves the
  session onto a node profile mid-turn (node tools, rules, snapshot)
  or back to general; the prompt lists the profiles and the active one
  and says to call it instead of asking
- chain.validators status=jailed lists only jailed validators, every
  line shows JAILED and the valoper; it's a core tool in node mode
- tool_search points at use_node instead of /one

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Typing 'okay lets stop' during a turn only queued it until the whole
turn ended. Like Claude Code now:
- Agent.Steer: a message typed mid-turn is delivered at the model's
  next step, appended to the latest tool result (valid message shape
  for every provider) with a note that it may change or stop the work
- if the turn ends before another step, the message starts the next
  turn immediately
- esc interrupts and sends the waiting message right away (it used to
  discard it); the live region shows that a message is waiting

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- the step limit is 200 (was 50) and reaching it pauses the turn with
  'say continue' instead of 'agent exceeded 50 iterations' (history is
  valid, so continuing picks up where it stopped)
- bash results that hand-roll polling (sleep && …) carry a tip to use
  wait.until, which waits without spending steps — the run that hit
  the limit was polling evmd query tx every 2s
- 'query tx <hash>', 'q tx', 'tx get <hash>' count as hash labels: the
  hash was shown as [REDACTED_HEX]

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- README: what cometcli is now (a Claude Code-style terminal for Cosmos
  SDK and Cosmos-EVM validators), quick start with Anthropic (low-cost
  Haiku) or OpenRouter free models and saved keys, the terminal and its
  keys, incidents (triage, knowledge base, /incident, unjail checks),
  transactions with both signers, governance, the command table
  (triage, gov, kb, wait), providers, safety, config layout
- COMMANDS.md: provider/key setup, ctrl+r / ctrl+o / esc and mid-turn
  messages, a Governance section, jail-check and unjail blockers,
  chain validators --status jailed

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@abhijitkrm
abhijitkrm merged commit 258cfcf into main Oct 7, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant