You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
CLM-v0.1-8B (Apache-2.0, released 2026-09-21) is an open System 1 decision model: a frozen Qwen3-8B encoder with state and action projection heads. It scores each candidate by the cosine similarity of the two projections. The authors report results on par with Jev on computer use and tool calling, up to 9× lower latency, and 13× faster than Jev at about 1,000 candidates when candidate embeddings are cached. Its authors also report state-of-the-art verification results (DeepSWE 81.6%, Terminal-Bench 2.1 87.6%). None of these numbers have been reproduced here.
The agent can reach Jev, LAYA and Cua-S1 Nano today. ThinkFlowLab/system1-omni#9 proposes CLM as the second served model after LAYA. This issue covers the agent side.
Scope
clm-serve answers POST /v1/systemone with the TypeSafe request and answer shape (src/clm/schema.py at bb42c6c). JevModel already builds this body. CLM runs as its own server because an 8B encoder is too large to load in the agent's process: a vllm serve Qwen/Qwen3-8B --runner pooling process plus clm-serve.
Add a clm backend that reuses the Jev wire code with its own endpoint setting and its own name, so that evaluation tables label it clm. Setting TYPESAFE_API_URL on the jev backend reaches the same server, but it labels every result jev and requires a placeholder API key. Coordinate the endpoint setting with the local-endpoint work in #20 so that both use one configuration pattern.
Test CLM in three roles:
pick and browser choice questions, where cached candidate embeddings matter most.
Rail noul checks, where the reported verifier results apply.
CLM's probabilities are relative to the candidates offered in each request, and score is an expected level index. Record whether the rails' thresholds, set against Jev, need separate values for CLM.
Acceptance criteria
Setup instructions specify the CLM and vLLM revisions, GPU and memory requirements, both server commands, and a reproducible smoke test.
Behavioural tests on a scripted transport cover the request body, choice, noul and score answers, invalid answers, and an unreachable server.
Why
CLM-v0.1-8B (Apache-2.0, released 2026-09-21) is an open System 1 decision model: a frozen Qwen3-8B encoder with state and action projection heads. It scores each candidate by the cosine similarity of the two projections. The authors report results on par with Jev on computer use and tool calling, up to 9× lower latency, and 13× faster than Jev at about 1,000 candidates when candidate embeddings are cached. Its authors also report state-of-the-art verification results (DeepSWE 81.6%, Terminal-Bench 2.1 87.6%). None of these numbers have been reproduced here.
The agent can reach Jev, LAYA and Cua-S1 Nano today. ThinkFlowLab/system1-omni#9 proposes CLM as the second served model after LAYA. This issue covers the agent side.
Scope
clm-serveanswersPOST /v1/systemonewith the TypeSafe request and answer shape (src/clm/schema.pyatbb42c6c).JevModelalready builds this body. CLM runs as its own server because an 8B encoder is too large to load in the agent's process: avllm serve Qwen/Qwen3-8B --runner poolingprocess plusclm-serve.Add a
clmbackend that reuses the Jev wire code with its own endpoint setting and its ownname, so that evaluation tables label itclm. SettingTYPESAFE_API_URLon thejevbackend reaches the same server, but it labels every resultjevand requires a placeholder API key. Coordinate the endpoint setting with the local-endpoint work in #20 so that both use one configuration pattern.Test CLM in three roles:
pickand browser choice questions, where cached candidate embeddings matter most.noulchecks, where the reported verifier results apply.CLM's probabilities are relative to the candidates offered in each request, and
scoreis an expected level index. Record whether the rails' thresholds, set against Jev, need separate values for CLM.Acceptance criteria
choice,noulandscoreanswers, invalid answers, and an unreachable server.CLM_ACTION_CACHE) and peak GPU memory.Dependencies and references
CLM code, model card, Jev backend, decision-model interface.
Serving: ThinkFlowLab/system1-omni#9. Endpoint configuration: #20. Evaluation: #13. Comparison table: #14. Verifier role: #16.