Skip to content

v0.12.0 feat: free models that stay up, and OpenCode as an engine - #4

Merged
inquerium merged 11 commits into
mainfrom
feat/acp-agent
Sep 15, 2026
Merged

inquerium merged 11 commits into
mainfrom
feat/acp-agent

Conversation

@inquerium

Copy link
Copy Markdown
Collaborator

Summary

OmniWork's "free out of the box" rested on OmniRoute's keyless pool: six unofficial endpoints (an Augment CLI shim, a scraped Vercel app, DuckDuckGo's chat widget, Xiaomi, a Chipotle chatbot, a WordPress ajax page). On 2026-09-15 every one was dead and auto returned 503 after 29 attempts; a fresh gateway with a clean data dir reproduced it. This PR makes the durable kinds of free one click away and makes a dead model a hop instead of a dead turn.

Fallback chain (electron/agent.js) — a model that fails (auth, quota, retired id, provider 5xx, empty reply) hands the same request to the next; the one that answers sticks for the session; total failure names every model. Identical consecutive failures short-circuit (that's the gateway, not the model). A dead gateway is still a connection error.

Free providers (electron/providers.js) — OpenRouter via its PKCE flow (browser sign-in, no key to paste, nonce-bound callback), Ollama Cloud / Kilo / Groq / Cerebras / NVIDIA / Gemini / Mistral by pasted key (proven before the old key goes), local Ollama / LM Studio / llama.cpp / vLLM by detection. All through OmniRoute's loopback management API. Connected providers become the default fallback chain for MCP, ACP and every desktop session, live.

OpenCode as an engine (electron/opencode-engine.js) — Zen serves its free models (Nemotron 3.5 Lightning, MiMo V2.5, Big Pickle, Ling 3.0 Flash, Nemotron 3 Ultra) only to OpenCode itself, so OmniWork runs OpenCode: opencode serve once per process, password-protected, sessions scoped to the workspace by the official directory header, events mapped onto OmniWork's (text, reasoning on its own lane, tools, permissions, errors). opencode/<model> is a model in the desktop menu, ACP config options and delegate; when no gateway model answers, desktop sessions, ACP turns and MCP delegations finish on the engine, carrying the conversation, and say so. OMNIWORK_ENGINE_FALLBACK=off opts out.

Getting OpenCode without PATH — opencode-ai is an optional dependency (a clone's npm install brings the platform binary); the build hook stages the target platform's package next to the bundled Node runtime; otherwise one user gesture downloads the npm package (~45 MB) verified against the sha512 pinned in package-lock.json. Always launched by absolute path.

Surfaces — desktop 🆓 free models panel; npm run providers / npx omniwork-providers; MCP list_providers, connect_provider (takes no API key, on purpose), list_models, model + fallback_models on delegate; ACP model config option (session/set_config_option, session/set_model, _meta.model) and auth methods openrouter / local / opencode.

Verification

Check Result
test/opencode.js, providers.js, fallback.js, acp.js, mcp.js, modes, approvals, persist all pass
Live: real OpenCode 1.18.31 + Zen — ACP pinned turn, retired-id fallthrough, MCP file-writing delegation pass (4–17 s)
Live: desktop app driven with Playwright — panel, engine turn, auto fallthrough pass
Live: sha512-verified download from npm, build staging for darwin-arm64 pass
Real server refuses unauthenticated requests 401
Servers left behind after tests none

Tests: 12 → 16 files (+4: fallback, providers, opencode, fake-opencode fixture).

Pre-landing review

Two adversarial reviews (Codex, Claude subagent). Fixed in fix: pre-landing review fixes: loopback auth on the OpenCode server; verified downloads; no API keys through the MCP tool; prove-before-retire on re-pasted keys; nonce-bound OAuth callback with exchange-before-success; narrower fallback triggers, identical-failure short-circuit, empty-last-reply as error, 429 excluded from the engine switch; carried context and persisted engine session on switch; start-timeout leak, stale exit handler, ref-counted event streams, concurrent installs, signal exit codes; abort-before-swap; parallel engine delegation.

Accepted as designed, called out: ACP authenticate("opencode") downloads the engine when a client asks — in ACP that call is the user picking the method in their client. Local servers on the default ports are trusted when the user connects them, as any Ollama client does. Universal macOS builds stage the host arch (same as the Node runtime); the other half downloads on first use.

Not in this PR

The Homebrew cask points at 0.11.1 until the 0.12.0 DMG exists (sha256 in the cask). Tag v0.12.0 after merge to run the release workflow, then bump the cask.

🤖 Generated with Claude Code

inquerium and others added 11 commits September 15, 2026 16:56
The agent loop takes a fallback chain. A model the gateway rejects (retired free id, missing provider key, quota) or that returns nothing hands the same request to the next model; the one that answers becomes the session's model, so a failure costs one request, not one per step. A dead gateway is still a connection error, and a provider that only rejects streaming still gets its non-streaming retry on the same model. When every model fails the error names each one and why.
OpenCode Zen serves its free models (Nemotron 3.5 Lightning, MiMo V2.5, Big Pickle, Ling 3.0 Flash, Nemotron 3 Ultra) only to OpenCode itself. So OmniWork runs OpenCode: one `opencode serve` per process, sessions scoped to the workspace with the official directory header, events (text, reasoning, tools, permissions, errors) mapped onto OmniWork's agent events by OpenCodeAgent, which has the same surface as Agent. Snapshots and the file watcher are off for engine sessions (a home-folder session went from 174 s to 4 s). The binary is found by absolute path wherever it landed: an env override, the packaged app's resources, a clone's node_modules, a download in the data dir, or an install the user already has; download() fetches OpenCode's official release archive. The server stops with the process, on exit or signal.
providers.js connects real free tiers through OmniRoute's loopback management API: OpenRouter via its PKCE flow (browser sign-in, no key to paste), Ollama Cloud, Kilo, Groq, Cerebras, NVIDIA, Gemini and Mistral by pasted key (exercised once; a rejected key is not kept), local Ollama / LM Studio / llama.cpp / vLLM by detection, and the OpenCode engine. headless.js resolves a session's model and chain live: an explicit choice, else the env, else whatever is connected, strongest coder first, local last; makeAgent() picks the engine for opencode/… models; noModelHint() tells a caller whose every model failed what to do.
…dead turn

ACP: the gateway catalog and the engine's free models are a `model` config option (session/set_config_option, the older session/set_model, or _meta.model on session/new); auth methods `openrouter`, `local` and `opencode` connect a provider; when no gateway model answers the turn finishes on the OpenCode engine and the client sees the switch. MCP: delegate / delegate_parallel take model + fallback_models, list_models shows the engine section, list_providers / connect_provider manage providers (connect_provider never installs anything itself), and a delegation that got no model reruns on the engine and says so.
…back

A 🆓 free models panel in the rail: one-click OpenRouter, paste-connect for the key providers, detected local servers, and the OpenCode engine (Download when missing). Connected providers become every session's fallback chain on the spot; opencode/… models appear in the model menu and run on the engine; a session whose gateway model fails finishes on the engine and writes the switch into the transcript.
…nd installers

`npm run providers` (also `npx omniwork-providers`) shows and connects providers from a terminal. opencode-ai is an optional dependency so `npm install` brings the platform binary; the build hook stages OpenCode's release archive per target platform next to the bundled Node runtime and the installers exclude the npm copy; doctor reports where the binary was found. The Claude Code guidance teaches list_providers / connect_provider and the engine models.
From two adversarial reviews of the branch (Codex and a Claude subagent):
- The OpenCode server answers only requests carrying a per-process secret (its own basic auth); loopback is not an authorization boundary.
- OpenCode downloads come from the npm registry and are verified against the sha512 pinned in package-lock.json before anything runs; the build staging uses the same path. GitHub releases publish no checksums.
- connect_provider over MCP takes no API key: a key supplied by a prompt-injected model would route the user's code through whoever supplied it. Keys go in the app or the terminal.
- A new key is proven before the old one goes: old connections are paused for the check, restored on failure, removed on success; a key that lists no models is not kept.
- The OAuth callback lives on a nonce path and the browser sees "connected" only after the key exchange; opener failures surface instead of a five-minute silence.
- The fallback chain walks only on failures that are about the model (auth, quota, retired id, provider 5xx, a 400 that names the model), stops when two models fail identically, and treats an empty last reply as an error rather than a blank done. A 429 no longer triggers the engine switch; OMNIWORK_ENGINE_FALLBACK=off disables it.
- A mid-session switch to the engine carries the recent conversation into its first prompt; ACP persists OpenCode's session id so a resumed session keeps its context.
- Start timeouts kill the child; a replaced server's exit no longer clobbers the new one; event streams are ref-counted per workspace and closed by the last subscriber; concurrent installs share one download; signal handlers exit with the conventional code and are not installed under Electron.
- Switching models across the engine boundary aborts a running agent first; delegate_parallel on an engine model runs one session per task.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…8.50)

Free tiers are the same for everyone now, so the edge is tokens per finished
task. electron/tuning.js plus wiring adds six levers, all on by default and all
opt-outable:

- Gateway RTK compression, enabled by the sidecar on boot. Rewrites tool-result
  text with cheap heuristics (no model call), keeps failed-command output
  verbatim. Measured 9,812 -> 3,170 tokens (68%) on a real grep result.
- Housekeeping (titles, memory, compaction summaries) runs on auto/best-fast,
  never the session's model.
- Step tiers: auto runs grunt work on the fast pool and escalates to the coding
  pool only when the fast model stalls (two failed steps, or six steps in). A
  pinned model never tiers.
- 429 round-robin: a rate-limited model cools for a minute and the step rotates
  to the next connected provider, without a permanent switch, so stacked free
  tiers share load. Hard failures still switch.
- MCP delegate ends with a cheap PASS/FAIL verdict from the utility model, so
  the orchestrator re-delegates only when the work fell short.
- Stable x-session-id per session for gateway prompt-cache affinity.
- Desktop opens the free-models panel once on first run when nothing is
  connected.

Bundles the OmniRoute 3.8.50 dependency bump (two months of upstream fixes:
model-catalog grouping, per-model lockouts instead of provider cooldowns).

New test/tuning.js (17 checks) covers every lever against a fake gateway.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review pass over each lever, tightening the ones that were leaving tokens on
the table or risking correctness:

- Tool output: head-only truncation dropped the tail of long command output —
  where the error message and the [exit code N] marker live — so the model and
  the escalation heuristic lost the failure signal. Now head+tail, cap raised
  30k -> 48k, and the exit-code marker always survives.
- RTK compression: send the full documented rtkConfig, not just level=standard.
  applyToToolResults, dedup, grouping, strip comments but preserve docstrings,
  and rawOutputRetention=failures so failed commands stay verbatim. The
  rich->minimal fallback still covers any future schema drift.
- Escalation: stop pulling every multi-step task onto the coding tier at step 6.
  Escalate on two failed steps (as before), on the step budget only when there
  has already been trouble, or as a hard backstop halfway through — so a fast
  model steadily editing many files stays cheap.
- toolFailed: detect a non-zero exit code anywhere (tail included), and anchor
  the shell-error phrases to end-of-line so "No such file or directory" inside
  normal output no longer counts as a failure.
- Delegate verifier: skip the extra model call on read-only delegations; only
  verify when the task changed something or its wording implies a write.

test/tuning.js grows to 22 checks covering the tightened helpers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@inquerium
inquerium merged commit 579cba8 into main Sep 15, 2026
2 of 3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant