Repository navigation
Rework cometcli into a Claude Code–style SRE terminal with validator incident handling - #3
Merged
Merged
Conversation
- retries: 408/409/429/5xx/529 and connection errors with jittered backoff honoring Retry-After; announced as notices, timeouts not retried; --debug logs every API attempt - normalized stop reasons: auto-continue after max_tokens, refusals surface as errors - frozen per-session system prompt; later node snapshots ride on user turns so provider prompt caches stay warm - token usage per provider (incl. cache reads), /cost - compaction: automatic at a turn boundary near the context limit, on context-overflow errors, and /compact [focus] - effort (--effort, /effort, agent.effort) mapped to Anthropic output_config.effort, OpenAI/Groq reasoning_effort, Gemini thinking_level; streamed reasoning shown as dimmed "∴" lines - saved sessions in ~/.cometcli/sessions: -c/--continue, -r/--resume, cometcli sessions, /sessions, /resume - Anthropic: verbatim replay of thinking blocks (dropped when the tool set changes), top-level prompt caching, eager tool input streaming, server-side refusal fallback, default model claude-opus-5-5 - head+tail, UTF-8-safe tool output truncation; max turns 16 -> 50 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Tools (agent-only: no CLI subcommand, not exported over MCP), all over the profile's host transport so they work locally and over SSH: - bash: persistent working directory, combined output capped head+tail at 30k, timeout (default 2m, max 10m) that kills the process group, stdin closed with non-interactive env - read / write / edit: line-numbered reads, edits require a fresh read and show a diff for approval, file mode preserved - glob / grep: newest-first file search; ripgrep or grep -E, with key files and keyring dirs excluded - web_fetch: HTML reduced to text, each new domain approved, localhost and private addresses fetched directly, cross-host redirects reported - todo_write: a visible checklist (EvTodos), saved with the session Shell classification: every command line is lexed (quotes, &&/||/;/|, redirects, $(…), heredocs, sudo/env/xargs/timeout wrappers, bash -c, docker exec/run) and each simple command classified. Reads auto-run; changes ask; transactions (evmd tx <module> anywhere, cast send, eth_sendRawTransaction, or a script whose body does so) always ask; consensus keys, mnemonics, key ceremonies and state resets are refused. Permissions: allow/ask/deny rules in Claude Code syntax (bash(cmd:*), read(./x/**), web_fetch(domain:x), node.*) from agent.permissions, --allowedTools/--disallowedTools and /allow /ask /deny; deny > ask > allow, and an allow must cover every segment of a compound command. Modes: ops, accept-edits, readonly, bypass (--permission-mode, /mode). Also: Context.Host falls back to the local machine without a profile; clearer denial when there is no terminal to approve on. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every request sent all 65 tool schemas (~6.5k tokens), over Groq's free 8k tokens-per-minute budget on its own. Requests now carry the general tools plus node.status/health/logs and val.status/signing, the frozen system prompt lists the rest by name, and the built-in tool_search loads definitions on demand (select:<names> or keywords). Loaded tools persist for the session and are saved with it. agent.tools: all restores the old behavior. First request: ~30KB -> ~10KB. Groq's 413 "request too large" now triggers compaction like a context-window overflow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- bare `cometcli [prompt]` is general mode: an SRE agent on this machine with the general tools and a general system prompt (OS, host, cwd, available node profiles); no profile needed. An explicit --profile means node mode. - `cometcli one [profile] [prompt]`: node mode — node tools (deferred), validator prompt, live snapshot, shell on the node's host. /one <profile> and /one off switch scope mid-session, keeping the conversation (replay state dropped, the model told on the next turn). - -p/--print: answer once and exit; --output-format text|json| stream-json, --include-partial-messages, --verbose, --append-system- prompt, --max-turns. Headless runs never read approvals from stdin. - piped stdin (pipes and redirected files only) is attached as <stdin> context, capped at 200k with head+tail kept. - global `agent:` section in config.yaml, overridden per profile (AgentFor/MergeAgent); general mode falls back to the active profile's settings. `cometcli config show|set|unset agent.<key>`. - single-word typos of subcommands (doctr) are reported, not sent. - the loop streams whenever the provider can, rendered or not. - ask/agent/ui unchanged (node mode on the active profile). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gpt-oss on Groq answered "summarize docker container health" with a generic explainer and no tool calls. Both system prompts now say that questions concern this machine/node and must be answered from evidence gathered with tools, generic explanations only on request. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ts (phase 4) Claude Code-compatible formats throughout: - settings.json at user / project / local scope: model, effort, permissions (allow/ask/deny/defaultMode), env (exported before the provider reads keys), hooks, mcpServers; plus a project .mcp.json. - trust: project hooks and MCP servers run commands, so they are ignored (with a notice) until `cometcli trust`; same for a project command's allowed-tools. - COMET.md memory: user, project root down to cwd (AGENTS.md fallback, COMET.local.md), per-node memory in node mode, @imports (key material refused); /memory, /remember appends and tells the model at once. - custom slash commands from .cometcli/commands/**.md: frontmatter (description, argument-hint, allowed-tools as one-turn allow rules), $ARGUMENTS/$1..$9, !`cmd` inlined for read-only commands only, @file; listed in /help, usable with -p. - hooks: PreToolUse (exit 2 / permissionDecision deny|ask|allow — allow never skips a tx prompt), PostToolUse (feedback to the model), UserPromptSubmit (block / add context), Stop (keep working, max 3), SessionStart; stdin JSON contract, process-group kill on timeout. - MCP client: stdio and streamable HTTP (SSE responses, session ids), pagination, ${VAR} expansion, server-request rejection; tools become mcp__server__tool, deferred behind tool_search, readOnlyHint runs freely, others ask, "mcp__server" rules cover a server. `cometcli mcp add|list [--check]|remove` (user/project/local scope); bare `cometcli mcp` still serves. - subagents: a task tool running a fresh-context child agent (same tools/permissions, no recursion, events shown with ↳), plus custom agents from .cometcli/agents/*.md with tool restrictions. - docs/EXTENDING.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mmands Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…se 5) New internal/tui/chatui, used by bare `cometcli` and `cometcli one`: - inline (no alt screen): finished messages print into scrollback; only streaming text, spinner, input, dialogs and the status line redraw - banner, ● tool lines with ⎿ results (diffs colored), markdown answers, todo checklists, last answer shown on resume, resume hint on exit - input: enter sends, \⏎ / ctrl+j newline, history, / command menu (builtins + custom commands), @file completion with content attached, # adds memory, messages typed during a turn are queued - !cmd runs in the operator's real terminal via ExecProcess (prompts and hidden password input work, e.g. vote-upgrade.sh); output never reaches the model, which is told the command and exit code; over `ssh -t` in node mode with an SSH transport - approvals: Yes / Yes, don't ask again for <suggested rule> (session + .cometcli/settings.local.json or user settings) / No (declines and interrupts); transactions offer only Yes/No - shift+tab cycles ops → accept-edits → read-only (→ bypass when started with it); esc interrupts; ctrl+c clear/interrupt/double-exit; ctrl+d - no terminal background query (dark unless COLORFGBG says light); zero-size pty reports ignored Supporting: toolkit.SuggestRule + RuleHint on approval details (never for transactions), Context.ToolName, Event.Output previews, /status and /clear, agent.Commands/ExpandMentions/AddNote/AllowAlways/ProjectFiles/ ContextPercent; the y/N approver hides "_" hint keys. `cometcli ui` keeps the full-screen dashboard. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Options.ForceSeq was documented as ">=0 overrides the queried sequence", but its zero value is 0 and no caller set -1, so Build always signed with sequence 0. Ordinary transactions only worked because simulation failed with "expected N, got 0" and the single automatic retry used the right sequence — costing an extra round trip and a spurious "sequence drifted" message on every tx. Two cases were genuinely broken: offline signing with --seq always signed sequence 0 (pinned accounts are never retried), and a real sequence drift could not heal because the bug had already spent the one retry. The override is now an unexported pointer set via ForceSequence(n), so the zero Options means "use the chain's sequence". Found by the new fakenode test harness (internal/testutil/fakenode): an in-process cosmos gRPC node that decodes every broadcast, verifies the secp256k1 signature over the SIGN_MODE_DIRECT sign doc (keccak for eth_secp256k1, sha256 for secp256k1) and enforces sequences with the SDK's error text. New tests cover both key types, fee/gas math, the approval gate (denied txs never reach the node; on-chain is never auto-approved), drift healing with re-approval, pinned sequences, CheckTx rejection, in-block failure, unconfirmed-as-pending, wrong chain-id, and DecScaled. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tatus Found by new tests running the tx tools against fakenode: - parseCoin accepted "1.5adex" (decimal amount), "1e18adex" (1 unit of denom "e18adex"), "10 adex", "0adex" and one-letter denoms. Amounts are now positive integers in base units with an SDK-valid denom, with specific messages for decimals and scientific notation. - val.vote broadcast VOTE_OPTION_UNSPECIFIED for an unknown or upper-case option and proposal 0 when the id was missing; options are now case-insensitive and validated, the id required. - val.edit with nothing to change broadcast a no-op edit (paying fees); it now refuses. - val.status dereferenced Commission/Description without nil checks. Tests: commission scaling + description preservation on edit, unjail, withdraw with commission, vote options, send message contents, input validation never reaching the node. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
net.add-peer / net.rm-peer: - anchor on the `persistent_peers = "…"` line itself (a comment mentioning persistent_peers matched first and the edit failed; the longer persistent_peers_max_dial_period key was also a risk) - validate peers as <40-hex node id>@host:port before asking - removing an absent peer (or re-adding a present one) changes nothing and asks nothing, instead of rewriting the file and saying "removed" - keep config.toml's file mode (writes reset it to 0644) upgrade.prepare: - find the binary anywhere in the archive (releases nest it under bin/) - accept the checksum of the archive or of the binary — releases publish either; a correct archive checksum used to fail - no checksum: the approval carries a warning and the result says UNVERIFIED with verified=false - work dir cleaned up on exit Tests for both, against a temp node home on the local host and a test HTTP server serving real tarballs (three layouts × two checksum kinds, mismatch, missing binary). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
snap.statesync --apply reported success but wrote nothing when config.toml had no [statesync] section, never added missing keys when [statesync] was the last section, matched keys by prefix (enable vs enable_*), and reset the file mode to 0644. It now edits exact keys in place, appends missing keys or the whole section, and keeps the mode. Audit log tests: concurrent sessions sharing one file keep every JSONL line intact, 0600/0700 permissions, session lookup across daily files (bad lines skipped), first prompts per session, nil logger no-op. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…less resume note E2E tests run the real root command against a scripted OpenAI-compatible server: -p text/json/stream-json, general-mode tools and prompt, a real bash call, headless denial of a write (and --allowedTools allowing it), sessions/-c/-r, piped stdin, slash and custom commands in -p, config set, typo detection, one without a profile. Bugs found: - the openai-compat provider (Ollama/vLLM/llama.cpp) had no name, so it reported itself as "openai" and was treated as hosted: it got max_completion_tokens and reasoning_effort, fields local servers may reject. It now identifies as openai-compat. - cometcli -p -r <id> resumed silently; the resume note goes to stderr. - a retry test asserted on timing; made robust. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fleet.shell ran any command on every host after one generic local-change approval, bypassing the shell classifier: key material was not refused (cat priv_validator_key.json across the fleet would reach the model), an `evmd tx` was not treated as a transaction (in bypass mode it ran without asking), and permission rules didn't apply. It now goes through shell.Classify + Context.Check like bash; read-only commands still run without a prompt, and a fleet-wide script run always asks (remote scripts can't be inspected). keys.add / keys.rm were agent-callable: a generated mnemonic went into the tool result (scrubbed by redaction, so the operator never saw the only copy) and --recover read stdin directly. New toolkit.OperatorOnly: such tools are never advertised to the model, refused if called, hidden from tool_search and MCP, and refused in runbook steps the agent runs. They stay CLI subcommands. Runbooks now validate every step before running any (an unknown tool in step 5 surfaced after steps 1-4 ran). keys.convert validates 0x input. Tests: fleet (reads across hosts, key material refused, tx and scripts ask in bypass, rules), fleet.exec refusing mutating tools, keys round trip + operator-only flags, runbook steps/optional/manual/validation, host Local exec (combined output, closed stdin, non-interactive env, head+tail cap, timeout killing children, file mode), agent refusing operator-only tools. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Golden files (internal/tui/chatui/testdata, regenerate with -update) for the idle prompt, slash menu, bash and transaction approval dialogs, a running turn, the accept-edits footer and a transcript of tool calls, diffs, subagent results, todos and markdown — layout changes show up as diffs. CI runs go test -race with a coverage summary. Fixed a data race in the approval test (shared variable polled across goroutines → channel). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
hexToUint64 returned 0 on a parse failure, so a malformed eth_blockNumber read as block 0 and evm.parity / the monitor reported a huge false drift. Quantities must be 0x-hex or the call fails; non-2xx responses (a proxy's 502 page) report the status instead of 'invalid character <'. Tests: quantities, syncing false vs progress object, txpool, client version, decimal-string rejection, RPC errors, HTTP errors. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… at signing time Transactions are now prepared → approved → signed → broadcast, with two signers behind tx.Signer: - LocalSigner: cometcli's ops keyring (previous behavior). - ContainerSigner: the node's own keyring inside its container — the key never leaves the node. cometcli builds the unsigned tx (the --generate-only JSON, passed base64 in the command so stdin carries only a password), runs `evmd tx sign --offline --sign-mode direct` and `evmd tx encode` via `docker exec` (locally or over SSH), and broadcasts the result. Gas is simulated with the account's on-chain public key; without one it falls back to 300000 and the approval says so. A file keyring's password is asked for after approval only, never stored or sent to the model; wrong passwords, missing keys and stopped containers get clear errors. Choosing: signer.mode local|container, else the only one available, else the operator picks when both can sign (CLI numbered prompt, TUI dialog); headless runs must set signer.mode. Keyring backend and key address are detected when not configured (test keyring probed, else file; address from metadata account/valoper or `keys show`). New Context hooks: Chooser and Secret (CLI: shared stdin reader, /dev/tty with echo off; TUI: choice dialog and masked input; headless: env COMETCLI_CONTAINER_KEYRING_PASSWORD), carried through agent tool contexts and scope switches. Host exec gained stdin (ExecIn). Tests: a fake `docker` (the test binary on PATH) emulates evmd keys show / tx sign / tx encode with real SIGN_MODE_DIRECT signing, verified by fakenode: test and file keyrings, simulation with an on-chain pubkey vs default gas, no password prompt for a declined tx, wrong password, signer choice / headless refusal / signer.mode. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-flight For running an incident to resolution (e.g. jailed → recovered): - val.jail-check: jail reason (downtime / tombstoned) and when; evidence from the node's own logs in the window before the jailing (the x/slashing jail line, then peer loss, signer/privval errors, panics, OOM, disk, app-hash, clock patterns with counts and samples), container or systemd restarts/OOM kills, disk; then every unjail blocker with its fix: jail period not over, node not synced (it would be re-jailed), self-delegation below minimum, fee balance, tombstoned. - val.consensus: bonded and unjailed, in the CometBFT validator set with voting power, our signature present in the last N commits, node synced. - wait.until: synced | height | unjailable | bonded | signing | in-consensus | tx-committed | proposal-status — polled in-process (no model tokens while waiting), live progress (sync rate and ETA), done=false with the last state on timeout, cancellable. Also a CLI command: cometcli wait until --condition synced. - val.unjail refuses on hard blockers before paying a fee; every tx failure carries a cause → next-step hint (still jailed, not jailed, self-delegation, fees, funds, gas, wrong key, chain-id, mempool). - Progress hook on tool contexts → EvProgress; the TUI shows it as a live line under the spinner. - jail-check, consensus and wait.until join the always-loaded core tools. Test harness: fakenode gains a fake CometBFT JSON-RPC chain (status, net_info, health, commit, validators) whose height, sync state, block times and our signatures tests drive; delegation, bank balance and downtime_jail_duration on the gRPC side; fake docker serves logs/inspect. Fixed in wait.until while testing: a poll interval longer than the remaining time ended the wait after one check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ulated chain /recover-jail (built-in, node mode; a project/user command of the same name overrides it) gives the agent the procedure: val.jail-check → stop if tombstoned → fix the root cause from the evidence → wait.until synced → wait.until unjailable → val.unjail (operator approves; signer chosen at signing time) → on failure read the cause/hint, fix, retry (max 3) → wait.until in-consensus (50 blocks, 95%) → val.consensus → report with evidence and tx hashes. Built-ins show in /help and the slash menu. internal/scenario runs the whole incident with the real agent loop and tool registry against a simulated chain that evolves while the agent works: jailed for downtime with peer loss in the node's logs, node behind and catching up, jail period ending in chain time, the first unjail rejected at CheckTx for an insufficient fee (node's min gas price raised), retry with a higher gas price, the chain re-admitting the validator, signatures resuming. A scripted model reacts only to tool output; the test checks the procedure order, the evidence, one rejection + one broadcast, two approvals, signer asked per attempt, and the final "participating in consensus" (20/20 signed). Deterministic: the node starts syncing when the agent starts waiting. fakenode: CheckFailOnce and OnCommit hooks. Docs: incidents, signing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Running the container signer against primium-evm:v0.7.2 (a throwaway container with fresh test keys) found that the image has no mktemp; the sign script now uses per-process files under /tmp with umask 077. With that, evmd accepts cometcli's unsigned tx JSON, signs it in both a test and a password-protected file keyring, and the result verifies. fakenode now accepts the 65-byte r||s||v signatures evmd produces for eth_secp256k1 keys (the chain's ethsecp256k1 verifier drops v), as well as 64-byte ones. The validation is kept as a build-tagged test (realdocker) with the container setup in its header. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
node.triage collects ~50 flat signals in parallel (comet, chain, gov, host, process, logs, config, evm) with per-source timeouts and matches them against a YAML knowledge base of 73 Cosmos SDK / CometBFT / Cosmos-EVM cases (confirm, causes, tagged fix steps, verify, never-do). Cases are embedded, extendable per user/project, and each carries a fixture the test suite proves ranks it in the top 3; a healthy node matches nothing. - internal/kb: case schema, condition language, matcher, search - internal/logscan: shared log categories (val forensics uses it too) - kb.search / kb.show / kb.add tools; wait.until condition=signal - /incident built-in and a generic incident method in the node prompt - toolkit.Context.Derive for fan-out with shared clients Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- logscan: strip ANSI colors; count info/debug lines only for categories where they are the evidence (jail, double sign, upgrade halt); tighter app-hash pattern (a healthy handshake logs appHash=…) - endpoints.fallback_grpc: chain queries go through another node while this one is down, so a dead validator's jail state is still visible; failed gRPC dials are no longer cached for the context's lifetime - kb: advisory kind (hardening) listed separately; severity ranks first - jail-check: log-gap detection (docker logs -t) shows when the node was silent and its last line — stopped from outside vs crashed; docker lifecycle events when still in the daemon's buffer Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- agent.provider openrouter: OpenAI-compatible, default model openrouter/free, key OPENROUTER_API_KEY; always requests the reasoning stream (reasoning.enabled) so thinking shows live - reasoning_content deltas accepted (DeepSeek/vLLM-style servers) - cometcli config set-key NAME: saves an API key from stdin to ~/.cometcli/credentials (0600); env vars still win - logscan: skip cobra --help output, startup Error: lines, last error signal, jail lines filtered to our own validator - kb: explains links — root causes rank first, symptoms grouped Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tx/block/app hashes are public and the operator needs them in reports; unlabelled 64-hex still redacts (key-shaped). The agent guessed config paths in the live drill — triage now shows config.file / app.file. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- triage: container memory limit, usage vs limit, cgroup oom_kill count (docker's OOMKilled only covers PID 1; nodes run under wrapper shells) - kb: container-memory-limit case (root of crash loops / jails) - logscan: a signer refusing a height/round/step regression is the double-sign guard working, not a privval fault - /incident + node prompt: apply [change]/[tx] steps by calling the tool (its prompt is the approval) instead of stopping to ask in text — the live drill showed a model stop at 'requires your approval' Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- triage: inside-the-container RPC probe and attached-network count — a detached or unreachable node is no longer mistaken for a dead one; kb cases node-network-detached, node-rpc-path-broken - triage logs start 5m before the current process start (older errors are history and produced stale root causes) - retries: a daily/credit-quota 429 fails immediately with its message - docs: signer is chosen first, then the tx approved Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eouts - val: SignerAhead — priv_validator_state ahead of the chain is a hard unjail blocker (the live drill unjailed a validator that couldn't sign yet and it was jailed again); fix says wait, never reset - cli: tools run with their own Timeout() (wait.until, triage) instead of a flat 90s deadline - common.ReadNodeFile shared by triage and val - kb: 'no SignBytes found' documented under privval errors Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The vote drill died at 'exit -1': under a loaded Docker VM (loadavg 25+) the container signer's evmd calls exceeded the flat 90s call deadline, which also counts the operator reading the approval prompt. toolkit.CallTimeout is shared by the CLI and the agent. Triage reads the docker host's load from inside the container (Docker Desktop's VM is otherwise invisible from macOS). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A just-restarted node dials fine but answers 'invalid height' while it loads state — exactly when an operator tries to unjail. Gather retries through endpoints.fallback_grpc, and without one says what's happening. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- docker nodes: df of <container home>/data (signer.container_home or the mount of the profile's home) — the disk that fills can be its own volume; the live disk-full drill showed the host disk at 77% while the node's volume was full - jail pattern no longer matches store keys (store_key=…slashing) - node-rpc-path-broken only after 30s uptime (probes raced a restart) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The vote drill offered tn0's 'ops' key on tn1..3 (shared keyring, same signer.key): choosing it would have signed a valid tx as another account. With metadata.valoper set, a local key must match its account. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- jail lines filtered by our valoper too (slash events name valopers) - node-crash-loop needs recent restarts (uptime < 10m); restart counts are lifetime-cumulative - upgrade-halt case: matches on the plan height too; documents the coordinated --unsafe-skip-upgrades path - container signer surfaces cobra's 'Error:' line instead of --help tail - signer choice is single-sourced: TxSigner is cached per context and common.Account uses it, so a message's address and its signature can't come from different keys (seen live: local ops key vs container node2) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Seen live: an RPC node syncing from genesis crash-looped at the upgrade height the validators had skipped. Triage no longer runs validator queries for rpc/sentry profiles (it only collects the upgrade plan). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- gov.propose: text, software upgrade and arbitrary-message proposals;
deposit defaults to the chain minimum and is checked against both
min_initial_deposit_ratio and min_deposit_ratio before approval; an
upgrade height the vote can't reach in time (voting period at the
measured block time) is refused; reports the new proposal id
- gov.deposit tops up exactly what's missing to the minimum
- nested messages packed with Cosmos type URLs ('/cosmos.…', not
type.googleapis.com/ — chains rejected the latter, seen live)
- val.vote says when a proposal is closed or still in deposit instead of
surfacing 'inactive proposal' from simulation
- chain gov shows the deposit-period end for proposals awaiting deposit
Drilled live: propose → 4 votes → PASSED; deposit top-up moved a
proposal into voting; upgrade and community-pool-spend proposals
submitted through the container signer.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Checked against the real terminal (pty + emulator), not just goldens: - welcome box 'Welcome to cometcli!' at a fixed width (long paths lose their middle instead of wrapping the box) + tips; printed after the first window size so it fits the terminal - tool calls: a blinking ⏺ with '⎿ Running…' while executing (paused 'Waiting for your answer…' during dialogs), then a green/red ⏺ with output collapsed to 3 lines, '… +N lines (ctrl+r to expand)'; Read shows 'Read N lines' - thinking: the spinner says Thinking…, then '∴ Thought for Ns (ctrl+o to show thinking)'; ctrl+o prints it - streaming text rendered as markdown live; '- ' bullets, inline code colored without padding, headings bold - spinner verb varies per turn; '>' prompt; placeholder only before the first message; Interrupted · What should cometcli do instead? - every printed line wrapped to the terminal width with a hanging indent (terminal-wrapped lines corrupted the live region redraw) - /status shows the current setup in any mode (model, mode, scope, cwd, session, memory, MCP), plus the node snapshot in node mode; /status and /memory no longer build the node prompt (it probed the node and froze the UI when the node was down) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Like Claude Code's block-letter greeting: COMETCLI in the ANSI Shadow font with a gold → coral → violet gradient, under a comet whose tail flares from a glowing nucleus through cyan and blue into violet dust. Shown on fresh sessions when the terminal is wide enough (≥ 74 cols); narrower terminals get the welcome box alone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Haiku 4.5 (and Sonnet 4.5) reject output_config.effort with a 400, which failed every request once agent.effort was set. Effort now goes only to Opus 4.5+, Sonnet 4.6+ and the 5 family; if a model still rejects it, the request is retried without it and effort is dropped for the rest of the session. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Asked 'is anyone jailed?' in general mode, the model ran 'which comet' and ls, then asked the operator to type /one — the general prompt told it to. Now: - use_node(profile|off): a built-in, read-only tool that moves the session onto a node profile mid-turn (node tools, rules, snapshot) or back to general; the prompt lists the profiles and the active one and says to call it instead of asking - chain.validators status=jailed lists only jailed validators, every line shows JAILED and the valoper; it's a core tool in node mode - tool_search points at use_node instead of /one Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Typing 'okay lets stop' during a turn only queued it until the whole turn ended. Like Claude Code now: - Agent.Steer: a message typed mid-turn is delivered at the model's next step, appended to the latest tool result (valid message shape for every provider) with a note that it may change or stop the work - if the turn ends before another step, the message starts the next turn immediately - esc interrupts and sends the waiting message right away (it used to discard it); the live region shows that a message is waiting Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- the step limit is 200 (was 50) and reaching it pauses the turn with 'say continue' instead of 'agent exceeded 50 iterations' (history is valid, so continuing picks up where it stopped) - bash results that hand-roll polling (sleep && …) carry a tip to use wait.until, which waits without spending steps — the run that hit the limit was polling evmd query tx every 2s - 'query tx <hash>', 'q tx', 'tx get <hash>' count as hash labels: the hash was shown as [REDACTED_HEX] Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- README: what cometcli is now (a Claude Code-style terminal for Cosmos SDK and Cosmos-EVM validators), quick start with Anthropic (low-cost Haiku) or OpenRouter free models and saved keys, the terminal and its keys, incidents (triage, knowledge base, /incident, unjail checks), transactions with both signers, governance, the command table (triage, gov, kb, wait), providers, safety, config layout - COMMANDS.md: provider/key setup, ctrl+r / ctrl+o / esc and mid-turn messages, a Governance section, jail-check and unjail blockers, chain validators --status jailed Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reworks cometcli from a validator dashboard plus agent into a Claude Code–style SRE terminal. One engine runs in two scopes:
cometclion your own machine, andcometcli one <profile>for a node. It handles validator incidents from finding the problem to fixing it and checking the fix.Agent and terminal
use_node), and you can switch with/one.-pwith text / json / stream-json output, and piped input.docker execbefore they run.settings.jsonat user, project and local level,COMET.mdmemory, custom slash commands, hooks, MCP servers and subagents.⏺tool calls with output collapsed (ctrl+r expands it), collapsed thinking (ctrl+o shows it), approval dialogs, a slash-command menu,@file mentions,!to run a command in your own terminal, and messages typed mid-turn that reach the model at its next step (Esc sends them immediately).Incidents
node.triagecollects about 50 signals in parallel from the node, chain, governance, host, process, logs, configs and EVM endpoint. A source that's down is reported, not fatal.kb.add./incidentruns the general procedure: triage, then the top case, then fix the root cause, wait, verify and report./recover-jailis the ready-made procedure for a jailed validator.gov.propose(text, software upgrade, arbitrary messages) andgov.deposit. The upgrade height is checked to be reachable within the voting period, and deposits are checked against the chain's minimums.Fixes found along the way
fleet.shellwent around the shell checks; it now goes through them.Testing
Unit and end-to-end tests run with the race detector in CI. They include a simulated chain (fake gRPC node, fake CometBFT RPC and fake docker) that runs a full jail recovery through the real agent loop.
Container signing was validated against the real
primium-evm:v0.7.2image.Live drills on a throwaway 4-validator testnet with an extra RPC node:
Each drill turned up bugs, and each fix has a regression test. The agent resolved the config and OOM incidents end to end on a free model.
The UI was checked in a real pseudo-terminal at 70 and 110 columns, not only with snapshot tests.
Not included
serveweb chat and the macOS app are frozen: they still compile but weren't extended.Docs:
docs/COMMANDS.md(incidents, triage, signing, governance) anddocs/EXTENDING.md(settings, hooks, MCP, knowledge-base case format).🤖 Generated with Claude Code