You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: README.en.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -16,7 +16,7 @@ credential pool, unified scheduling, and per-user usage stats.
16
16
-**OpenCode Zen free tier** (third channel, `zen`): the free models at `opencode.ai/zen`, standard OpenAI protocol, no login. The upstream list mixes in paid models with no free/paid marker, so the gateway narrows by the `-free` suffix and probes each candidate, exposing **only the free models that actually answer anonymously** (fetched and probed live every time, no static allowlist). Requests automatically satisfy the free-tier gate; responses have the gate's injected pseudo-tool calls filtered out. Zen has no credentials — one virtual pool row lets it be scheduled, paused, and counted like any other channel.
17
17
-**Kilo Gateway free tier** (fourth channel, `kilo`): the free models at `api.kilo.ai/api/gateway`, standard OpenAI protocol, no login. Filtered by the authoritative per-model `isFree` flag (no probing, to conserve the small free quota); free models are marked **x0**. Same credential-less virtual-row model as Zen.
18
18
-**Qoder** (fifth channel, `qoder`): a **real-account** upstream ([Qoder](https://qoder.com)) reached via device-code PKCE login. The upstream speaks a private COSY protocol (custom Base64 body + envelope SSE); the gateway signs and unwraps it, so the surface stays standard OpenAI. Supports quota probing and daily check-in (credits accumulate). Since 2026-10 check-in is **campaign-based** (`/sash/api/v1/me/campaigns`, with a `Cosy-ClientType` header; the legacy `daily-check-in` endpoint now returns `DISABLED` and is kept only as a fallback). Qoder sometimes wraps its own inference-node failures in a 400 (`[FAIL]node:… msg:Execution failed`); the gateway classifies these as a **model-scoped transient fault** — it cools only that model and returns a "model temporarily unavailable" 503 instead of misreporting a missing model or an exhausted pool.
19
-
-**CodeArts** (sixth channel, `codearts`): a **real-account** upstream ([Huawei Cloud CodeArts](https://codearts.huaweicloud.com)) reached via OAuth2 PKCE → STS (AK/SK signing + DPoP refresh). The portal redirects the authorization code to the *user's*`127.0.0.1` callback, which a server cannot listen on, so the login flow asks the user to paste that callback URL back and exchanges the code server-side. The upstream emits cumulative-text SSE, which the gateway reduces to incremental events. Supports quota probing and benefit-token claiming. **No daily check-in** — the free quota is a **daily pool of 10M free tokens that resets at midnight (no rollover)**, so token auto-refresh is the keep-alive. Because that pool is use-it-or-lose-it, the day's remainder is registered as quota expiring at the next local midnight, so scheduling **burns it first** and falls back to other channels only once it is exhausted.
19
+
- **CodeArts** (sixth channel, `codearts`): a **real-account** upstream ([Huawei Cloud CodeArts](https://codearts.huaweicloud.com)) reached via OAuth2 PKCE → STS (AK/SK signing + DPoP refresh). The portal redirects the authorization code to the *user's* `127.0.0.1` callback, which a server cannot listen on, so the login flow asks the user to paste that callback URL back and exchanges the code server-side. The upstream emits cumulative-text SSE, which the gateway reduces to incremental events. Supports quota probing and benefit-token claiming. **No daily check-in** — the free quota is a **daily pool of 10M free tokens that resets at midnight (no rollover)**, so token auto-refresh is the keep-alive. Because that pool is use-it-or-lose-it, the day's remainder is registered as quota expiring at the next local midnight, so scheduling **burns it first** and falls back to other channels only once it is exhausted. Benefit models carry no rate (they consume the daily pool), so per-request usage records `credit = input + output tokens` (1:1 with the pool, marked ≈ as derived).
20
20
-**Expiry-aware scheduling**: among credits expiring within `QUOTA_EXPIRY_WINDOW_SECONDS` (default 36h), the largest balance is burned first; ties fall to `QUOTA_EXPIRY_SECONDARY_WINDOW_SECONDS` (default 7 days), so near-expiry quota is not wasted. CodeArts' daily token pool lands in this ladder every day (its remainder expires at midnight), so it is consumed before other channels; units differ per channel (CodeArts counts tokens, the rest credits)
21
21
-**Conversation stickiness**: a multi-turn conversation keeps one credential and only rotates on error. Identified by an explicit id when the client sends one (`conversation_id` / `conversationId` / `prompt_cache_key`, top-level or in `metadata`), otherwise by a message-prefix fingerprint. Pinned credentials always win
22
22
-**Cached-token accounting**: TRAE's `cache_read_input_tokens` / `cache_creation_input_tokens` are mapped to per-request `cached_tokens` and surfaced in stats; the "Token usage" card also shows a **cache hit rate** (cached ÷ input, 1 decimal), or `—` when cache was never reported / input is 0
0 commit comments