Skip to content

[Feat] Integrate and evaluate CLM-v0.1-8B through clm-serve #24

Description

@Yunaik

Why

CLM-v0.1-8B (Apache-2.0, released 2026-09-21) is an open System 1 decision model: a frozen Qwen3-8B encoder with state and action projection heads. It scores each candidate by the cosine similarity of the two projections. The authors report results on par with Jev on computer use and tool calling, up to 9× lower latency, and 13× faster than Jev at about 1,000 candidates when candidate embeddings are cached. Its authors also report state-of-the-art verification results (DeepSWE 81.6%, Terminal-Bench 2.1 87.6%). None of these numbers have been reproduced here.

The agent can reach Jev, LAYA and Cua-S1 Nano today. ThinkFlowLab/system1-omni#9 proposes CLM as the second served model after LAYA. This issue covers the agent side.

Scope

clm-serve answers POST /v1/systemone with the TypeSafe request and answer shape (src/clm/schema.py at bb42c6c). JevModel already builds this body. CLM runs as its own server because an 8B encoder is too large to load in the agent's process: a vllm serve Qwen/Qwen3-8B --runner pooling process plus clm-serve.

Add a clm backend that reuses the Jev wire code with its own endpoint setting and its own name, so that evaluation tables label it clm. Setting TYPESAFE_API_URL on the jev backend reaches the same server, but it labels every result jev and requires a placeholder API key. Coordinate the endpoint setting with the local-endpoint work in #20 so that both use one configuration pattern.

Test CLM in three roles:

  1. pick and browser choice questions, where cached candidate embeddings matter most.
  2. Rail noul checks, where the reported verifier results apply.
  3. Checks during replanning in [Feat] Recover from stalled browser and desktop tasks with bounded replanning #16, once that lands.

CLM's probabilities are relative to the candidates offered in each request, and score is an expected level index. Record whether the rails' thresholds, set against Jev, need separate values for CLM.

Acceptance criteria

  1. Setup instructions specify the CLM and vLLM revisions, GPU and memory requirements, both server commands, and a reproducible smoke test.
  2. Behavioural tests on a scripted transport cover the request body, choice, noul and score answers, invalid answers, and an unreachable server.
  3. Evaluation on the [Eval] Add repeatable browser and desktop task suites #13 suites reports completion rate, total task time, decision latency and failure reasons next to the models compared in [Feat] Integrate and evaluate Cua-S1 4B 0.2 #14, on the same tasks and budgets.
  4. Decision latency is reported separately for a cold and a warm candidate cache, together with the cache size (CLM_ACTION_CACHE) and peak GPU memory.
  5. Results state which of the three roles CLM is useful for and list configurations that failed or remain unsupported.

Dependencies and references

CLM code, model card, Jev backend, decision-model interface.

Serving: ThinkFlowLab/system1-omni#9. Endpoint configuration: #20. Evaluation: #13. Comparison table: #14. Verifier role: #16.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions