v0.12.0 feat: free models that stay up, and OpenCode as an engine - #4
Merged
Merged
Conversation
The agent loop takes a fallback chain. A model the gateway rejects (retired free id, missing provider key, quota) or that returns nothing hands the same request to the next model; the one that answers becomes the session's model, so a failure costs one request, not one per step. A dead gateway is still a connection error, and a provider that only rejects streaming still gets its non-streaming retry on the same model. When every model fails the error names each one and why.
OpenCode Zen serves its free models (Nemotron 3.5 Lightning, MiMo V2.5, Big Pickle, Ling 3.0 Flash, Nemotron 3 Ultra) only to OpenCode itself. So OmniWork runs OpenCode: one `opencode serve` per process, sessions scoped to the workspace with the official directory header, events (text, reasoning, tools, permissions, errors) mapped onto OmniWork's agent events by OpenCodeAgent, which has the same surface as Agent. Snapshots and the file watcher are off for engine sessions (a home-folder session went from 174 s to 4 s). The binary is found by absolute path wherever it landed: an env override, the packaged app's resources, a clone's node_modules, a download in the data dir, or an install the user already has; download() fetches OpenCode's official release archive. The server stops with the process, on exit or signal.
providers.js connects real free tiers through OmniRoute's loopback management API: OpenRouter via its PKCE flow (browser sign-in, no key to paste), Ollama Cloud, Kilo, Groq, Cerebras, NVIDIA, Gemini and Mistral by pasted key (exercised once; a rejected key is not kept), local Ollama / LM Studio / llama.cpp / vLLM by detection, and the OpenCode engine. headless.js resolves a session's model and chain live: an explicit choice, else the env, else whatever is connected, strongest coder first, local last; makeAgent() picks the engine for opencode/… models; noModelHint() tells a caller whose every model failed what to do.
…dead turn ACP: the gateway catalog and the engine's free models are a `model` config option (session/set_config_option, the older session/set_model, or _meta.model on session/new); auth methods `openrouter`, `local` and `opencode` connect a provider; when no gateway model answers the turn finishes on the OpenCode engine and the client sees the switch. MCP: delegate / delegate_parallel take model + fallback_models, list_models shows the engine section, list_providers / connect_provider manage providers (connect_provider never installs anything itself), and a delegation that got no model reruns on the engine and says so.
…back A 🆓 free models panel in the rail: one-click OpenRouter, paste-connect for the key providers, detected local servers, and the OpenCode engine (Download when missing). Connected providers become every session's fallback chain on the spot; opencode/… models appear in the model menu and run on the engine; a session whose gateway model fails finishes on the engine and writes the switch into the transcript.
…nd installers `npm run providers` (also `npx omniwork-providers`) shows and connects providers from a terminal. opencode-ai is an optional dependency so `npm install` brings the platform binary; the build hook stages OpenCode's release archive per target platform next to the bundled Node runtime and the installers exclude the npm copy; doctor reports where the binary was found. The Claude Code guidance teaches list_providers / connect_provider and the engine models.
From two adversarial reviews of the branch (Codex and a Claude subagent): - The OpenCode server answers only requests carrying a per-process secret (its own basic auth); loopback is not an authorization boundary. - OpenCode downloads come from the npm registry and are verified against the sha512 pinned in package-lock.json before anything runs; the build staging uses the same path. GitHub releases publish no checksums. - connect_provider over MCP takes no API key: a key supplied by a prompt-injected model would route the user's code through whoever supplied it. Keys go in the app or the terminal. - A new key is proven before the old one goes: old connections are paused for the check, restored on failure, removed on success; a key that lists no models is not kept. - The OAuth callback lives on a nonce path and the browser sees "connected" only after the key exchange; opener failures surface instead of a five-minute silence. - The fallback chain walks only on failures that are about the model (auth, quota, retired id, provider 5xx, a 400 that names the model), stops when two models fail identically, and treats an empty last reply as an error rather than a blank done. A 429 no longer triggers the engine switch; OMNIWORK_ENGINE_FALLBACK=off disables it. - A mid-session switch to the engine carries the recent conversation into its first prompt; ACP persists OpenCode's session id so a resumed session keeps its context. - Start timeouts kill the child; a replaced server's exit no longer clobbers the new one; event streams are ref-counted per workspace and closed by the last subscriber; concurrent installs share one download; signal handlers exit with the conventional code and are not installed under Electron. - Switching models across the engine boundary aborts a running agent first; delegate_parallel on an engine model runs one session per task.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…8.50) Free tiers are the same for everyone now, so the edge is tokens per finished task. electron/tuning.js plus wiring adds six levers, all on by default and all opt-outable: - Gateway RTK compression, enabled by the sidecar on boot. Rewrites tool-result text with cheap heuristics (no model call), keeps failed-command output verbatim. Measured 9,812 -> 3,170 tokens (68%) on a real grep result. - Housekeeping (titles, memory, compaction summaries) runs on auto/best-fast, never the session's model. - Step tiers: auto runs grunt work on the fast pool and escalates to the coding pool only when the fast model stalls (two failed steps, or six steps in). A pinned model never tiers. - 429 round-robin: a rate-limited model cools for a minute and the step rotates to the next connected provider, without a permanent switch, so stacked free tiers share load. Hard failures still switch. - MCP delegate ends with a cheap PASS/FAIL verdict from the utility model, so the orchestrator re-delegates only when the work fell short. - Stable x-session-id per session for gateway prompt-cache affinity. - Desktop opens the free-models panel once on first run when nothing is connected. Bundles the OmniRoute 3.8.50 dependency bump (two months of upstream fixes: model-catalog grouping, per-model lockouts instead of provider cooldowns). New test/tuning.js (17 checks) covers every lever against a fake gateway. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review pass over each lever, tightening the ones that were leaving tokens on the table or risking correctness: - Tool output: head-only truncation dropped the tail of long command output — where the error message and the [exit code N] marker live — so the model and the escalation heuristic lost the failure signal. Now head+tail, cap raised 30k -> 48k, and the exit-code marker always survives. - RTK compression: send the full documented rtkConfig, not just level=standard. applyToToolResults, dedup, grouping, strip comments but preserve docstrings, and rawOutputRetention=failures so failed commands stay verbatim. The rich->minimal fallback still covers any future schema drift. - Escalation: stop pulling every multi-step task onto the coding tier at step 6. Escalate on two failed steps (as before), on the step budget only when there has already been trouble, or as a hard backstop halfway through — so a fast model steadily editing many files stays cheap. - toolFailed: detect a non-zero exit code anywhere (tail included), and anchor the shell-error phrases to end-of-line so "No such file or directory" inside normal output no longer counts as a failure. - Delegate verifier: skip the extra model call on read-only delegations; only verify when the task changed something or its wording implies a write. test/tuning.js grows to 22 checks covering the tightened helpers. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
OmniWork's "free out of the box" rested on OmniRoute's keyless pool: six unofficial endpoints (an Augment CLI shim, a scraped Vercel app, DuckDuckGo's chat widget, Xiaomi, a Chipotle chatbot, a WordPress ajax page). On 2026-09-15 every one was dead and
autoreturned 503 after 29 attempts; a fresh gateway with a clean data dir reproduced it. This PR makes the durable kinds of free one click away and makes a dead model a hop instead of a dead turn.Fallback chain (
electron/agent.js) — a model that fails (auth, quota, retired id, provider 5xx, empty reply) hands the same request to the next; the one that answers sticks for the session; total failure names every model. Identical consecutive failures short-circuit (that's the gateway, not the model). A dead gateway is still a connection error.Free providers (
electron/providers.js) — OpenRouter via its PKCE flow (browser sign-in, no key to paste, nonce-bound callback), Ollama Cloud / Kilo / Groq / Cerebras / NVIDIA / Gemini / Mistral by pasted key (proven before the old key goes), local Ollama / LM Studio / llama.cpp / vLLM by detection. All through OmniRoute's loopback management API. Connected providers become the default fallback chain for MCP, ACP and every desktop session, live.OpenCode as an engine (
electron/opencode-engine.js) — Zen serves its free models (Nemotron 3.5 Lightning, MiMo V2.5, Big Pickle, Ling 3.0 Flash, Nemotron 3 Ultra) only to OpenCode itself, so OmniWork runs OpenCode:opencode serveonce per process, password-protected, sessions scoped to the workspace by the official directory header, events mapped onto OmniWork's (text, reasoning on its own lane, tools, permissions, errors).opencode/<model>is a model in the desktop menu, ACP config options anddelegate; when no gateway model answers, desktop sessions, ACP turns and MCP delegations finish on the engine, carrying the conversation, and say so.OMNIWORK_ENGINE_FALLBACK=offopts out.Getting OpenCode without PATH —
opencode-aiis an optional dependency (a clone'snpm installbrings the platform binary); the build hook stages the target platform's package next to the bundled Node runtime; otherwise one user gesture downloads the npm package (~45 MB) verified against the sha512 pinned inpackage-lock.json. Always launched by absolute path.Surfaces — desktop 🆓 free models panel;
npm run providers/npx omniwork-providers; MCPlist_providers,connect_provider(takes no API key, on purpose),list_models,model+fallback_modelsondelegate; ACPmodelconfig option (session/set_config_option,session/set_model,_meta.model) and auth methodsopenrouter/local/opencode.Verification
autofallthroughTests: 12 → 16 files (+4: fallback, providers, opencode, fake-opencode fixture).
Pre-landing review
Two adversarial reviews (Codex, Claude subagent). Fixed in
fix: pre-landing review fixes: loopback auth on the OpenCode server; verified downloads; no API keys through the MCP tool; prove-before-retire on re-pasted keys; nonce-bound OAuth callback with exchange-before-success; narrower fallback triggers, identical-failure short-circuit, empty-last-reply as error, 429 excluded from the engine switch; carried context and persisted engine session on switch; start-timeout leak, stale exit handler, ref-counted event streams, concurrent installs, signal exit codes; abort-before-swap; parallel engine delegation.Accepted as designed, called out: ACP
authenticate("opencode")downloads the engine when a client asks — in ACP that call is the user picking the method in their client. Local servers on the default ports are trusted when the user connects them, as any Ollama client does. Universal macOS builds stage the host arch (same as the Node runtime); the other half downloads on first use.Not in this PR
The Homebrew cask points at 0.11.1 until the 0.12.0 DMG exists (sha256 in the cask). Tag
v0.12.0after merge to run the release workflow, then bump the cask.🤖 Generated with Claude Code