Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,13 @@ MODEL_NAME=google/gemini-2.5-flash
# CUA_S1_SUBFOLDER=text # the text-only checkpoint; the window is 256 bytes of state
# CUA_S1_DEVICE=auto # cpu, cuda, mps

# ---- OmniJev (the in-process vision decision model behind --model omnijev; needs `uv sync --extra omnijev`) ----
# OMNIJEV_REPO=/path/to/OmniJev # required: a clone of https://github.com/tinnel123666888/OmniJev
# OMNIJEV_CHECKPOINT=tinnel123/OmniJev-0.8B # or OmniJev-2B, OmniJev (4B); a Hugging Face id or a local directory
# OMNIJEV_REVISION=v1.1 # for a Hugging Face id
# OMNIJEV_BASE=Qwen/Qwen3.5-0.8B # the Qwen3.5 backbone of the same size
# OMNIJEV_PROMPT=full # full: text state + goal + rules; short: goal + ask over the screenshot

# ---- Cua Driver (the hands of the desktop agents, e.g. calculator; install: https://cua.ai/docs/cua-driver) ----
# CUA_DRIVER_BIN=/usr/local/bin/cua-driver # when cua-driver is not on PATH
# CUA_DRIVER_BIN=C:/path/to/cua-driver.exe # Windows: full path to the installed executable
Expand Down
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,12 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); ver

### Added

- `--model omnijev` on the browser agents: [OmniJev](https://github.com/tinnel123666888/OmniJev) (Apache-2.0), a
Qwen3.5 vision-language decision model, in process behind `uv sync --extra omnijev` and a local clone named by
`OMNIJEV_REPO`. It decides over a screenshot; `docs/decision-models.md`, `docs/configuration.md`.
- Browser front: a decision model that reads images (`supports_images`) gets a viewport PNG with every tick's
observation, captured through the probe's run-code executor; the tick records its `screenshot_ms`. Jev, Laya and
Cua-S1 read text only and see no change.
- The MCP `decide` tool accepts `model="jev"|"laya"|"cua"`, defaulting to `jev`. Local backends use their
optional extras and need no Jev API key.
- `docs/benchmarks.md`: the Google Flights driver comparison rerun on 2026-09-23 from Poland, every arm three times on
Expand Down
26 changes: 24 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,28 @@ wall clock. The other Allrecipes runs, longer games and the Google Flights drive
</tr>
</table>

### OmniJev: a decision model that reads the screen

[OmniJev](https://github.com/tinnel123666888/OmniJev) (Apache-2.0, Beijing Zhongguancun Academy, CASIA and Zevo) is a
System 1 decision model on Qwen3.5 vision-language backbones (0.8B, 2B, 4B): it answers the same typed questions as
Jev over a screenshot, a video or a robot camera. `--model omnijev` puts it in the slot of the browser agents, which
then send it a screenshot of the page at every step ([docs/decision-models.md](docs/decision-models.md)).

The clips below are OmniJev's own v1.1 demos with its 4B model: replays of recorded trajectories with the model's
probabilities, not s1a runs and not live control. They are shown from the OmniJev repository and keep their upstream
terms.

<table>
<tr>
<td width="50%"><img src="https://raw.githubusercontent.com/tinnel123666888/OmniJev/14dbec4f71e194852c8d7b88ab36ef639493f400/docs/media/v11/web.gif" alt="OmniJev on Mind2Web web tasks: a Central Park to JFK route and a nightstand comparison, with the model's probabilities per step"><br><sub>Web: Mind2Web test tasks (OmniJev v1.1 replay)</sub></td>
<td width="50%"><img src="https://raw.githubusercontent.com/tinnel123666888/OmniJev/14dbec4f71e194852c8d7b88ab36ef639493f400/docs/media/v11/phone.gif" alt="OmniJev on AndroidControl phone tasks: London weather, a 59-minute timer and a drawing tutorial"><br><sub>Phone: AndroidControl tasks (OmniJev v1.1 replay)</sub></td>
</tr>
<tr>
<td><img src="https://raw.githubusercontent.com/tinnel123666888/OmniJev/14dbec4f71e194852c8d7b88ab36ef639493f400/docs/media/v11/arcade.gif" alt="OmniJev on Atari recordings: Enduro, Skiing and Pong, next input and steering"><br><sub>Games: Enduro, Skiing and Pong (OmniJev v1.1 replay)</sub></td>
<td><img src="https://raw.githubusercontent.com/tinnel123666888/OmniJev/14dbec4f71e194852c8d7b88ab36ef639493f400/docs/media/v11/robot.gif" alt="OmniJev on a two-view Bridge robot recording folding a cloth: jog direction, gripper and move size"><br><sub>Robotics: folding a cloth, two camera views (OmniJev v1.1 replay)</sub></td>
</tr>
</table>

## Choose your path

### From Claude Code or Codex
Expand Down Expand Up @@ -124,8 +146,8 @@ The gates and the templates: [docs/skills.md](docs/skills.md#build-a-system-1-ag
- `game2048`, `millionaire`, `blackjack`: games with a score per episode.
- `injection_guard`: a rail that answers one question at a hook of a running agent and fails closed.

Every agent runs on `jev`, `laya` or `cua`, and on the chat model for the comparison. Flags, run commands and
extras: [docs/agents.md](docs/agents.md).
Every agent runs on `jev`, `laya` or `cua`, and on the chat model for the comparison; the browser agents also run on
`omnijev`, which decides over a screenshot. Flags, run commands and extras: [docs/agents.md](docs/agents.md).

## How it works

Expand Down
2 changes: 1 addition & 1 deletion docs/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ hook of a running agent. The injection guard rail fails closed: a decision error

Every tool agent takes `--model jev|laya|cua|llm|random|rule`, `--rethink on|off`, `--episodes N`, `--seed S`,
`--max-steps` and `--timeout`, and writes a Harbor-shaped job folder under `evals/results/<agent>/`. A browser agent
takes `--model jev|laya|cua|llm` and `--goal`. A rail takes `--model jev|laya`, the two models that answer `noul`.
takes `--model jev|laya|cua|omnijev|llm` and `--goal`. A rail takes `--model jev|laya`, the two models that answer `noul`.
`uv run python -m evals.table evals/results` aggregates every job folder per eval and model into one table.

Every `run` prints one JSON object on stdout and nothing else there; `s1a-mcp` serves the same agents over stdio
Expand Down
4 changes: 2 additions & 2 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@ the first import of both, routes the harness logs to files under `runs/logs` bef
MCP stdio protocol only.

Every `run` prints one JSON object on stdout: a tool agent's series summary with its `job_dir`, a browser agent's
answer, a rail's evaluation. A browser agent takes `--model jev|laya|cua|llm`; its policy switches are run-time flags:
answer, a rail's evaluation. A browser agent takes `--model jev|laya|cua|omnijev|llm`; its policy switches are run-time flags:
`--batch on|off`, `--prefetch on|off`, `--goal-values on|off`. `s1a-mcp` serves the same agents to an MCP host over stdio,
one Runner for the server's lifetime and one run at a time. `uv run python -m evals.table evals/results`
aggregates every job folder per eval and model into one table. `scripts/showcase.sh` plays one visual episode per
Expand All @@ -71,7 +71,7 @@ eval and model outside the matrix and `python -m evals.replay` renders a pair si

Every tool agent, `desktop` included, takes `--model jev|laya|cua|llm|random|rule`, `--rethink on|off`,
`--episodes N`, `--seed S`, `--max-steps`, `--timeout` and `--headed`, and writes a Harbor-shaped job folder under
`evals/results/<agent>/`. Every browser agent takes `--model jev|laya|cua|llm` and `--goal`. A rail takes
`evals/results/<agent>/`. Every browser agent takes `--model jev|laya|cua|omnijev|llm` and `--goal`. A rail takes
`--model jev|laya`, the two models that answer `noul`. `decide` and `probe` take `--model jev|laya|cua`. On a browser
agent `laya` needs `LAYA_MAX_LEN` raised to the page's size; `cua` reads a 256-byte context (header, goal, state,
then rules) and 96 bytes per option, a baseline on any page. Exit codes: 0 for a finished run, including one whose
Expand Down
5 changes: 5 additions & 0 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,11 @@ Variables can be exported in your shell or placed in a `.env` file at the root o
| `CUA_S1_CHECKPOINT` | `cua` model | `cua-ai/cua-s1-nano-0.1` | Hugging Face checkpoint ID or local directory for Cua-S1 Nano option scorer. |
| `CUA_S1_SUBFOLDER` | `cua` model | `text` | Subfolder within checkpoint directory containing text option scoring weights. |
| `CUA_S1_DEVICE` | `cua` model | `auto` | PyTorch device used for Cua-S1 Nano evaluation (`auto`, `cpu`, `cuda`, or `mps`). |
| `OMNIJEV_REPO` | `omnijev` model | *(unset, required)* | Local clone of the OmniJev repository; its `mso` package is imported from there. |
| `OMNIJEV_CHECKPOINT` | `omnijev` model | `tinnel123/OmniJev-0.8B` | Hugging Face ID or local directory of the OmniJev adapter and heads (`OmniJev-2B`, `OmniJev` for 4B). |
| `OMNIJEV_REVISION` | `omnijev` model | `v1.1` | Revision of a Hugging Face `OMNIJEV_CHECKPOINT`. |
| `OMNIJEV_BASE` | `omnijev` model | `Qwen/Qwen3.5-0.8B` | Hugging Face ID or local directory of the Qwen3.5 backbone of the same size as the checkpoint. |
| `OMNIJEV_PROMPT` | `omnijev` model | `full` | `full` sends the text state, goal and rules with each question; `short` sends the goal and the ask only, the screenshot carrying the page. |
| `CUA_DRIVER_BIN` | `desktop` agent | `cua-driver` | Path to the `cua-driver` executable on Windows or macOS when not located on `PATH`. |
| `CUA_DRIVER_PERMISSION_MODE` | `desktop` agent | `standard` | Permission mode passed to `cua-driver mcp` (`standard`, or `bounded` for restricted capability manifests). |
| `HF_HOME` | Hugging Face runtime | `~/.cache/huggingface` | Cache directory where Laya and Cua-S1 checkpoints are downloaded on first run. |
Expand Down
1 change: 1 addition & 0 deletions docs/decision-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,7 @@ shorthands; `warm()` and `close()` open and release the backend.
| `jev` | `JevModel(transport)` | `jev` | the request body every front sent before the layer existed, byte for byte; `from_env` picks TypeSafe or the OpenRouter proxy (see [configuration.md](configuration.md)) |
| `laya` | `LayaModel(agent, model=)` | `laya` | one forward pass per call on a thread; `MODEL_SERVICE_CONFIG_ERROR` when `input_tokens` fills the window (Laya cuts the state silently; see `LAYA_MAX_LEN` and `LAYA_HEAD_MAX_LEN` in [configuration.md](configuration.md)); `ValueError` and `RuntimeError` from the library become `MODEL_CALL_FAILED` |
| `cua` | `CuaS1Model(scorer, collator, model=, context_bytes=, option_bytes=)` | `cua` | Cua-S1 Nano, one `score_elements` pass per request on a thread; choice questions only, text only, deterministic; the context is header, state and rules; the checkpoint reads its first 256 bytes, and the first overflowing request logs one warning; `from_env` reads `CUA_S1_*` (see [configuration.md](configuration.md)) |
| `omnijev` | `OmniJevModel(agent, model=)` | `omnijev` | OmniJev (a Qwen3.5 vision-language decision model), one `MSO1.system_one` call per request on a thread; choice and noul questions over the observation's screenshot, deterministic; the text state, goal and rules go in front of each question (a browser state puts its recent actions first and caps the page text at 2,000 characters, so a long page cannot push the history out of the 6,000-character context), a browser element row becomes its label and value; a choice's probabilities, which OmniJev leaves `abstain` out of, are renormalised over the offered keys; an observation without an image is `MODEL_SERVICE_CONFIG_ERROR`; `from_env` reads `OMNIJEV_*` (see [configuration.md](configuration.md)) |
| `random` | `RandomModel(seed)` | `random` | uniform over the offered keys, confidence 0, one seeded stream per episode; choice questions only |
| `rule` | `RuleModel(name, rule)` | the rule's name | one-hot, confidence 1; a key outside the menu raises `RuntimeError`, a bug in the rule |

Expand Down
13 changes: 12 additions & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,17 @@ cua = [ # Cua-S1 Nano behind --model cua; pinned to the Cua PR that ships the c
"cua-s1 @ git+https://github.com/trycua/cua.git@aea61b6eb97e2d8c0f6f71eb804e5769fe910af4#subdirectory=libs/cua-s1/python",
"huggingface-hub>=0.24",
]
dev = ["pytest>=8", "pytest-asyncio>=0.24", "ruff>=0.6", "ty>=0.0.83"]
omnijev = [ # OmniJev behind --model omnijev; the model code itself is a clone named by OMNIJEV_REPO (not a package)
"torch>=2.4",
"torchvision>=0.15", # transformers' Qwen3-VL video processor imports it; OmniJev's requirements.txt omits it
"transformers>=5.0",
"peft>=0.15",
"accelerate",
"safetensors",
"pillow",
"huggingface-hub>=0.24",
]
dev =["pytest>=8", "pytest-asyncio>=0.24", "ruff>=0.6", "ty>=0.0.83"]

[build-system]
requires = ["hatchling"]
Expand Down Expand Up @@ -76,6 +86,7 @@ include = [
"s1a/agents/blackjack.py",
"s1a/decision_models/cua.py",
"s1a/decision_models/laya.py",
"s1a/decision_models/omnijev.py",
]

[tool.ty.overrides.rules]
Expand Down
7 changes: 4 additions & 3 deletions s1a/browser/browse.py
Original file line number Diff line number Diff line change
Expand Up @@ -38,8 +38,9 @@
"jev",
"laya",
"cua",
"omnijev",
"llm",
) # a decision model (Jev over HTTP, Laya or Cua-S1 in process) or the chat model
) # a decision model (Jev over HTTP, Laya, Cua-S1 or OmniJev in process) or the chat model
RUNS_DIR = HOME / "runs" / "browser"


Expand Down Expand Up @@ -164,7 +165,7 @@ async def browse(
workspace = str(logs_dir / "workspace") # the harness scaffolds SOUL.md, memory/ and friends here, not in the cwd
instance = BrowserInstanceConfig(launch_args=browser_launch_args(headless))
match model_name:
case "jev" | "laya" | "cua":
case "jev" | "laya" | "cua" | "omnijev":
if decision_model is None:
raise RuntimeError(f"--model {model_name} needs a decision model")
slot_model = BrowserDecisionModel(spec, policy, counted, decision_model=decision_model, value_model=None)
Expand Down Expand Up @@ -224,7 +225,7 @@ def parser(spec: BrowserAgentSpec) -> argparse.ArgumentParser:
"--model",
choices=BROWSER_MODEL_NAMES,
required=True,
help="who decides each browser step: jev (over HTTP), laya or cua (in process), or llm (the chat model in MODEL_NAME)",
help="who decides each browser step: jev (over HTTP), laya, cua or omnijev (in process; omnijev reads a screenshot), or llm (the chat model in MODEL_NAME)",
)
build.add_argument(
"--goal", default=spec.goal, required=spec.goal is None, help="the task; the spec's goal when it has one"
Expand Down
36 changes: 35 additions & 1 deletion s1a/browser/decision_model.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@
from __future__ import annotations

import asyncio
import base64
import json
import re
import statistics
Expand All @@ -41,7 +42,7 @@
top_probabilities,
)
from s1a.browser.probe_js import POLICY_PROBE_JS, STAMP_ATTRIBUTE
from s1a.decision_models import DecisionModel
from s1a.decision_models import DecisionModel, Image, Observation
from s1a.spec import BrowserAgentSpec

BROWSER_TURN_TOOL = "browser_click"
Expand Down Expand Up @@ -82,6 +83,14 @@
MAX_PROBE_SETTLE_MS = 1500 # keeps load(3s)+settle+1s JS lastResort >=1s under the 30s transport request timeout
WAIT_SETTLE_BUDGET_MS = 3000 # total in-page settle time one WAIT streak may spend before the step gives up
URL_RE = re.compile(r"https?://[^\s'\"<>]+")
# The viewport as PNG, for a decision model that reads images; one run-code call through the probe's own executor.
# A background tab renders no frames, so the page is brought to the front first; the capture waits 15 s at most.
SCREENSHOT_JS = (
"async (page) => { await page.bringToFront(); "
'const png = await page.screenshot({type: "png", timeout: 15000}); '
"return JSON.stringify({png: png.toString('base64')}); }"
)
_SCREENSHOT_RE = re.compile(r'png\\*"\s*:\s*\\*"([A-Za-z0-9+/=]+)')


@dataclass(frozen=True)
Expand Down Expand Up @@ -258,6 +267,11 @@ async def _decide_message(self, messages: Any) -> AssistantMessage:
run.values = []
values = run.values
observation = build_observation(space, snapshot, run.history)
screenshot_ms = None
if self._decision_model.supports_images:
screenshot, screenshot_ms = await self._screenshot()
if screenshot is not None:
observation = Observation(observation.state, images=(screenshot,))
questions = build_questions(
space, goal=run.goal, values=values, rules=self._spec.rules, language=self._language
)
Expand Down Expand Up @@ -287,6 +301,8 @@ async def _decide_message(self, messages: Any) -> AssistantMessage:
"probabilities": probabilities, # the replay's bars: the settled head's top keys
"candidates": candidates,
}
if screenshot_ms is not None: # an image-reading model: the capture is part of the step's cost
record["screenshot_ms"] = screenshot_ms
run.ticks.append(record)
if move.operation == "WAIT":
run.consecutive_waits += 1
Expand Down Expand Up @@ -397,6 +413,24 @@ async def _finish_probe(
self._prefetch_values(run, snapshot)
return snapshot, round((time.perf_counter() - started) * 1000)

async def _screenshot(self) -> tuple[Image | None, int]:
"""The viewport as PNG and the milliseconds the capture took. ``None`` when it fails: the step goes on
without the picture and the decision model says whether it can decide without one."""
started = time.perf_counter()
executor = getattr(self._runtime, "code_executor", None)
raw: Any = None
if callable(executor):
try:
raw = await executor(SCREENSHOT_JS)
except Exception: # noqa: BLE001 - a failed capture degrades to a text-only observation
logger.warning("[BrowserDecisionModel] screenshot failed", exc_info=True)
ms = round((time.perf_counter() - started) * 1000)
found = _SCREENSHOT_RE.search(json.dumps(raw, default=str)) if raw is not None else None
if found is None:
logger.warning("[BrowserDecisionModel] no screenshot in the run-code result")
return None, ms
return Image(base64.b64decode(found.group(1))), ms

async def _raw_probe(self, *, settle_ms: int, quiet_ms: int, after: dict[str, Any] | None) -> dict[str, Any]:
params = {
"stamp_attribute": STAMP_ATTRIBUTE,
Expand Down
2 changes: 2 additions & 0 deletions s1a/decision_models/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@
from s1a.decision_models.fakes import ScriptedModel, ScriptedTransport
from s1a.decision_models.jev import JevModel, jev_question
from s1a.decision_models.laya import LayaModel
from s1a.decision_models.omnijev import OmniJevModel
from s1a.decision_models.types import (
Answer,
Choice,
Expand Down Expand Up @@ -47,6 +48,7 @@
"LayaModel",
"Noul",
"NoulQuestion",
"OmniJevModel",
"Observation",
"Question",
"RandomModel",
Expand Down
6 changes: 5 additions & 1 deletion s1a/decision_models/factory.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,25 +8,29 @@
from s1a.decision_models.cua import CuaS1Model
from s1a.decision_models.jev import JevModel
from s1a.decision_models.laya import LayaModel
from s1a.decision_models.omnijev import OmniJevModel

DECISION_MODEL_NAMES = (
"jev",
"laya",
"cua",
"omnijev",
"random",
"rule",
) # the names that build a decision model; ``llm`` is not one


def build_model(model_name: str, *, seed: int = 0, rule: tuple[str, Rule] | None = None) -> DecisionModel:
"""``jev``, ``laya`` and ``cua`` from the environment, ``random`` from the seed, ``rule`` from the agent's baseline."""
"""``jev``, ``laya``, ``cua`` and ``omnijev`` from the environment, ``random`` from the seed, ``rule`` from the agent's baseline."""
match model_name:
case "jev":
return JevModel.from_env()
case "laya":
return LayaModel.from_env() # the laya import happens inside
case "cua":
return CuaS1Model.from_env() # the cua_s1 import happens inside
case "omnijev":
return OmniJevModel.from_env() # the OmniJev import happens inside; it decides over a screenshot
case "random":
return RandomModel(seed)
case "rule":
Expand Down
Loading