A self-hosted LLM routing daemon with fallback chains, live quota intelligence, and an OpenAI-compatible HTTP API. Part of the clagentic suite.
- Routes LLM calls across multiple backends (Claude CLI, Codex CLI, Gemini CLI, Ollama, Anthropic API, OpenAI API, AWS Bedrock)
- Walks a fallback chain when backends are unavailable or rate-limited
- Scores backends by health, quota pressure, latency EMA, and cost weight; near-ties broken by jitter
- Tracks quota/rate-limit state persistently in SQLite; auto-recovers when windows reset
- Parses
rate_limit_eventfrom the Claude CLI stream — captures live utilization, reset time, and bucket type on every response; persists toquota_snapshotstable for historical analysis.openai_api/anthropic_apifeed the same table via a synthetic utilization computed from their rate-limit headers;gemini_clihas no proactive quota signal reachable through its JSON invocation path (verified live — the CLI's internal quota model exists but is not wired to--output-format json/stream-json, seegemini_cli.go's package doc);codex_clihas a verified proactive signal available (account/rateLimits/readon the codex app-server JSON-RPC protocol, EXPERIMENTAL) that is not yet wired in — the JSON-RPC transport differs from every other adapter's one-shot subprocess call and needs its own verified client (seecodex_cli.go's package doc,TODO(lr-c98c)). Both remain reactive-only (error-text parsing) inquota_snapshotstoday - Exposes an OpenAI-compatible
/v1/chat/completionsendpoint — any OpenAI SDK works without changes; also exposes Anthropic Messages (/v1/messages) and Bedrock InvokeModel-shaped endpoints - Delivers webhook alerts (HMAC-signed, exponential retry) on backend state changes
- Runs as a daemon on any Linux host; CLI adapters (
claude_cli,codex_cli,codex_subagent,gemini_cli) require OAuth sessions on that host; API adapters (anthropic_api,openai_api,bedrock_api) work anywhere, including containers
This README is an orientation and link hub, not the manual. Full documentation is split by audience so each stays focused:
| Doc | Audience | Covers |
|---|---|---|
| docs/OPERATOR-GUIDE.md | Human operator | Install, configure, add a backend, deploy (systemd/Docker/update), diagnose a failure, logging |
| docs/AGENT-REFERENCE.md | Agent/integrator calling the daemon | Full API surface, adapter capability matrix, wire-field semantics, error taxonomy, routing invariants |
| docs/BEDROCK.md | Either | Every AWS Bedrock path: CLI-adapter Bedrock auth, bedrock_api, Bedrock InvokeModel HTTP endpoints |
router.example.yaml |
Either | Every config key, annotated with defaults and examples |
| CLAUDE.md | Contributor editing this repo's Go source | Build-time contract: breadth principle, import graph, discovery-vs-hardcode rules, subprocess cwd/HOME contract |
| docs/smoke-test.md | Human operator | End-to-end validation procedure against a live daemon |
docs/AGENT-REFERENCE.md states explicitly where its scope ends and
CLAUDE.md's begins — read its "Boundary with CLAUDE.md" section before
adding to either, so the two contracts don't drift apart by duplicating
the same claim twice.
# 1. Build
make build
# 2. Configure
cp router.example.yaml router.yaml
$EDITOR router.yaml
# 3. Run
export CLAGENTIC_ROUTER_TOKEN=mysecret
./bin/clagentic-router serve --config router.yaml
# 4. Call it
export CLAGENTIC_ROUTER_TOKEN=mysecret
./bin/clagentic-router call --model claude-haiku --message "What is 2+2?"
# Or via any OpenAI SDK:
# base_url = "http://localhost:8765/v1"
# api_key = "mysecret"Requirements: Go 1.25+. No CGO — pure Go SQLite via modernc.org/sqlite.
Full install/configure/deploy walkthrough: docs/OPERATOR-GUIDE.md.
Clagentic: Router exists to route across heterogeneous LLM backends, not to serve one provider well. A feature that only works for one provider, one auth mode, or one host is treated as incomplete. In practice this means:
- Discover, don't hardcode. Model/provider/project identifiers are
resolved at runtime from the provider's own source of truth (e.g. the codex
CLI's local config and model catalog) rather than typed into
router.yaml. Static values remain available as explicit overrides. - Explicit config always wins. If you set a value, it is used byte-identically and discovery is never attempted for it — safe to layer onto an existing deployment.
- Named production paths stay stable. The Claude brand account
(
claude_cli) and ChatGPT-Plus (codex_cli) are load-bearing; changes that improve one backend must not regress the others.
See CLAUDE.md for the full principle and the reference implementations; see docs/AGENT-REFERENCE.md for the wire-visible consequences (discovery is invisible to a caller by design — this is repo-internal context, not something a client integrates against).
Clagentic: Router is a self-hosted daemon. It accepts OpenAI-compatible requests, scores and selects backends via a pluggable adapter layer, walks a configurable fallback chain on failure, and persists health/quota state in SQLite.
graph LR
subgraph Clients
SDK["OpenAI SDK"]
CLI["Clagentic: Router CLI"]
Console["Clagentic: Console"]
end
subgraph Daemon["Clagentic: Router Daemon"]
API["HTTP API\n/v1/chat/completions"]
Router["Router\n(score + fallback)"]
State["State Machine\n(SQLite)"]
Webhook["Webhook Delivery\n(HMAC + retry)"]
end
subgraph Backends["LLM Backends"]
ClaudeCLI["claude CLI\n(OAuth)"]
CodexCLI["codex CLI\n(OAuth)"]
GeminiCLI["gemini CLI\n(OAuth)"]
Ollama["Ollama HTTP"]
AnthropicAPI["Anthropic API"]
OpenAIAPI["OpenAI API"]
end
SDK -->|Bearer token| API
CLI -->|Bearer token| API
Console -->|Bearer token| API
API --> Router
Router --> State
Router --> ClaudeCLI
Router --> CodexCLI
Router --> GeminiCLI
Router --> Ollama
Router --> AnthropicAPI
Router --> OpenAIAPI
State --> Webhook
- Per-backend
timeout_seconds(default 180 s) is enforced byRouter.Routeas a context deadline around everyInvoke, for all adapters (claude_cli,codex_cli,codex_subagent,gemini_cli,bedrock_api,anthropic_api,openai_api,ollama_http). An expiry is recorded as atimeoutfailure and the chain advances while the request is still live. A client disconnect mid-call is not charged to the backend. CLI adapters also bound how long they wait for a killed subprocess's output pipes (backend.SubprocessWaitDelay, 3 s), so a grandchild process holding stdout open cannot hold a request past its deadline. Before this, only the three HTTP adapters honored it; a CLI orbedrock_apibackend with notimeout_secondswas unbounded and now gets the 180 s default, so settimeout_secondsexplicitly on any backend that legitimately runs longer. - Per-request write deadline on the LLM endpoints (
/v1/chat/completions,/v1/messages,/model/{id}/invoke[-with-response-stream]): routed requests get the sum, over chain entries, of the largest backend timeout in each entry, plus 30 s, capped byproxy.max_request_seconds(default 1800). Passthrough requests are bounded byproxy.max_request_secondsalone, as the total including the 30 s margin, measured from just after the request body is read. Work whose failure is reported by writing an error (routing; for passthrough, credential resolution, signing and the upstream call up to response headers) stops 30 s before the write deadline so the error can always be written (when the bound is 30 s or less, it stops at the midpoint instead). Once a passthrough response has started, the body relay runs until the write deadline, so a stream still flowing at 30 s before the bound is delivered, and one still flowing at the bound is cut. The server-wide 300 sWriteTimeoutremains a backstop for health/admin/metrics endpoints only. - Attribution. Every deadline is created with a cause, and an outcome is read
once from
context.Causeat the moment it fires, never from which context looks done afterwards or from the clock at report time. Per routed attempt: the backend's owntimeout_seconds(health penalty; the chain advances if the request is live), the chain budget running out (charged as atimeoutto the backend running at that moment, even if it ran less than its own full timeout, because it overran the budget the chain allotted it; no further tier), the operator'sproxy.max_request_secondscap (request_deadline, no penalty), or a client disconnect (cancelled, no penalty). The first event wins: a client that leaves after the backend's own deadline already fired does not turn the timeout into a cancel. - If the response write still fails after a successful route, a
response delivery failedwarning is logged withrequest_id,backend,elapsed_ms,failed_after_msand acausetaken from the write error itself:write_deadline(the connection's write deadline,os.ErrDeadlineExceeded) orwrite_failed(anything else, e.g. the client left). Passthrough relay and upstream failures log thecauseas well (work_deadline,write_deadline,client_cancelled,nonefor an upstream fault).call_logstill records the backend outcome.
config → (stdlib)
state → (stdlib)
store → state
backend → config
webhook → state, store
router → backend, config, state, store, webhook
server → router, state, store
cmd/clagentic-router → config, backend, router, server, store, webhook
Full contributor-facing detail (adding an import, error classification internals, scoring formula, quota-alert edge-trigger mechanics) is in CLAUDE.md.
make tidy # go mod tidy
make build # produces bin/clagentic-router
make install # installs to GOBIN
make test # go test ./...
make docker # builds Docker imageIf clagentic:router is useful to you: ko-fi.com/clagentic
Not affiliated with Anthropic or OpenAI. Claude is a trademark of Anthropic. Codex is a trademark of OpenAI. Provided "as is" without warranty. Users are responsible for complying with their AI provider's terms of service.
FSL-1.1-MIT — Functional Source License 1.1, with MIT as the Change License.
Free for personal, internal-business, evaluation, research, and non-commercial use. Not free for offering this tool (or a substantial fork) as a competing commercial product. Each release auto-converts to MIT on its second anniversary.
