Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 20 additions & 14 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,34 +48,40 @@ try your own inputs and see its choices and probabilities.

### ⚡ Jev inside LLM agent loops

> **Key takeaway:** the LLM plans; the environment or LLM supplies a short menu;
> Jev picks the routine action; the LLM verifies and finishes the task.
**Key takeaways**

**1 · Terminal-Bench · `sqlite-db-truncate` — select one of three commands**
- **Task level:** T0 controlled → T4 open terminal. Task difficulty and delegation are separate.
- **Delegation level:** D0 LLM-only → D4 bounded subgoal. Increase delegation only when local choices remain easy to verify.
- **Delegate when:** the LLM can generate 2–4 valid, meaningfully different, reversible choices with immediate feedback.
- **Keep with the LLM:** planning, open-ended search, exact edits, recovery, high-risk actions, and completion.
- **Success test:** reward stays the same or improves while LLM calls, tokens, or time fall; otherwise return control to the LLM.

[![Jev selecting raw-page inspection from three Terminal-Bench commands before LLM parsing and verification](docs/demos/jev-agent-harness-sqlite.gif)](reports/JevAny_Tech_Report_Agent_Harness_Appendix.md#h1-sqlite-recovery-a-meaningful-three-way-decision)
**1 · WebShop — choose the exact color and size from LLM-generated menus**

**2 · WebShop — select the required color from the page actions**
[![The LLM generates three color and three size candidates, Jev selects the exact options, and the LLM completes the purchase](docs/demos/jev-decision-webshop-v2.gif)](docs/demos/jev-agent-harness-traces.json)

[![Jev selecting the required black product option, followed by the LLM choosing size 11.5 and completing the purchase](docs/demos/jev-agent-harness-webshop.gif)](docs/demos/jev-agent-harness-traces.json)
Reward `1→1` · LLM calls `9→4` · tokens `38,852→14,256` · time `18.54s→7.83s`

**3 · FrozenLake — compare four directions at every state**
**2 · FrozenLake — compare four directions at every state**

[![Jev choosing four state-dependent navigation actions after one LLM plan in FrozenLake](docs/demos/jev-agent-harness-frozen-lake.gif)](results/agent-harness-v1/formal-matrix.md)
[![Jev compares four directions, executes each selected move, reaches the goal, and reduces LLM calls](docs/demos/jev-decision-frozen-lake-v2.gif)](results/agent-harness-v1/formal-matrix.md)

Reward `1→1` · LLM calls `4→1` · tokens `2,338→663`

**3 · Terminal-Bench · `sqlite-db-truncate` — select one of three commands**

[![Jev compares three real Terminal-Bench commands, selects raw-page inspection, and the LLM recovers ten rows](docs/demos/jev-decision-terminal-v2.gif)](reports/JevAny_Tech_Report_Agent_Harness_Appendix.md#h1-sqlite-recovery-a-meaningful-three-way-decision)

Reward `1→1` · LLM calls `15→8` · time `187.9s→144.7s`

| Task | Success | Efficiency |
|---|---:|---:|
| GPT-5.6-sol · FrozenLake · 10 pairs | 100% → 100% | LLM calls −64.4% · tokens −63.1% · time −37.6% |
| WebShop · 10 pairs | 50% → 60% | LLM calls −7.7% · tokens −2.9% · time −6.1% |
| WebShop · LLM-generated menus · 3 pairs | 67% → 100% | LLM calls −21.4% · tokens −14.3% · time −15.0% |
| WebArena · 6 pairs | 50% → 50% | LLM calls +5.6% · tokens +28.2% · time −0.4% |
| Terminal-Bench · 6 pairs | 1/6 → 3/6 | LLM calls −9.0% |

**Use Jev for:** 2–4 bounded, reversible choices with an immediate observation.

**Keep with the LLM:** planning, exact edits, recovery, and the final answer.

[Full results](reports/JevAny_Tech_Report_Agent_Harness_Appendix.md) ·
[combined demo](docs/demos/jev-agent-harness.gif) ·
[task/delegation levels](docs/experiments/AGENT_HARNESS_FRONTIER_PROTOCOL.md) ·
[combined technical report](reports/JevAny_Tech_Report_with_Agent_Harness.pdf)

Expand Down
33 changes: 20 additions & 13 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,33 +45,40 @@

### ⚡ Jev 加入 LLM Agent 循环

> **关键结论:**LLM 负责规划,环境或 LLM 提供少量选项,Jev 选择常规动作,LLM 验证并完成任务。
**关键结论**

**1 · Terminal-Bench · `sqlite-db-truncate`——从三个命令中选择一个**
- **任务层级:**T0 受控环境 → T4 开放终端;任务难度与委托程度分开判断。
- **委托层级:**D0 仅 LLM → D4 有边界子目标;只有局部选择容易验证时才提高委托程度。
- **适合 Jev:**LLM 能生成 2–4 个有效、有明确差别、可回退且能立即看到反馈的选项。
- **保留给 LLM:**规划、开放搜索、精确修改、失败恢复、高风险动作和最终完成。
- **成功标准:**reward 不降,同时减少 LLM calls、tokens 或时间;否则立即交回 LLM。

[![Jev 从三个 Terminal-Bench 命令中选择原始页面检查,LLM 随后解析并验证](docs/demos/jev-agent-harness-sqlite.gif)](reports/JevAny_Tech_Report_Agent_Harness_Appendix.md#h1-sqlite-recovery-a-meaningful-three-way-decision)
**1 · WebShop——从 LLM 生成的颜色和尺寸候选中选择精确选项**

**2 · WebShop——从页面动作中选择任务要求的颜色**
[![LLM 分别生成三个颜色和三个尺寸候选,Jev 选中精确选项,LLM 完成购买](docs/demos/jev-decision-webshop-v2.gif)](docs/demos/jev-agent-harness-traces.json)

[![Jev 选择任务要求的黑色商品选项,LLM 随后选择 11.5 尺码并完成购买](docs/demos/jev-agent-harness-webshop.gif)](docs/demos/jev-agent-harness-traces.json)
Reward `1→1` · LLM calls `9→4` · tokens `38,852→14,256` · 时间 `18.54s→7.83s`

**3 · FrozenLake——每一步都比较四个方向**
**2 · FrozenLake——每一步都比较四个方向**

[![LLM 规划一次后,Jev 在 FrozenLake 中连续选择四个依赖当前状态的导航动作](docs/demos/jev-agent-harness-frozen-lake.gif)](results/agent-harness-v1/formal-matrix.md)
[![Jev 比较四个方向、执行选中的移动、到达目标并减少 LLM 调用](docs/demos/jev-decision-frozen-lake-v2.gif)](results/agent-harness-v1/formal-matrix.md)

Reward `1→1` · LLM calls `4→1` · tokens `2,338→663`

**3 · Terminal-Bench · `sqlite-db-truncate`——从三个真实命令中选择一个**

[![Jev 比较三个真实 Terminal-Bench 命令,选择原始页面检查,LLM 恢复十行数据](docs/demos/jev-decision-terminal-v2.gif)](reports/JevAny_Tech_Report_Agent_Harness_Appendix.md#h1-sqlite-recovery-a-meaningful-three-way-decision)

Reward `1→1` · LLM calls `15→8` · 时间 `187.9s→144.7s`

| 任务 | 成功率 | 效率 |
|---|---:|---:|
| GPT-5.6-sol · FrozenLake · 10 pairs | 100% → 100% | LLM calls −64.4% · tokens −63.1% · 时间 −37.6% |
| WebShop · 10 pairs | 50% → 60% | LLM calls −7.7% · tokens −2.9% · 时间 −6.1% |
| WebShop · LLM 生成候选 · 3 pairs | 67% → 100% | LLM calls −21.4% · tokens −14.3% · 时间 −15.0% |
| WebArena · 6 pairs | 50% → 50% | LLM calls +5.6% · tokens +28.2% · 时间 −0.4% |
| Terminal-Bench · 6 pairs | 1/6 → 3/6 | LLM calls −9.0% |

**适合交给 Jev:**2–4 个有边界、可回退、能立即观察结果的选择。

**保留给 LLM:**规划、精确修改、失败恢复和最终答案。

[完整结果](reports/JevAny_Tech_Report_Agent_Harness_Appendix.md) ·
[组合动画](docs/demos/jev-agent-harness.gif) ·
[任务与委托分级](docs/experiments/AGENT_HARNESS_FRONTIER_PROTOCOL.md) ·
[合并版技术报告](reports/JevAny_Tech_Report_with_Agent_Harness.pdf)

Expand Down
68 changes: 43 additions & 25 deletions docs/demos/jev-agent-harness-traces.json
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
{
"version": 2,
"version": 3,
"summary": {
"controlled_and_web_pairs": 96,
"terminal_pairs": 6,
"frozen_lake": {"success": "100% -> 100%", "llm_calls_change": "-64.4%"},
"webshop": {"success": "50% -> 60%"},
"webshop": {"success": "67% -> 100%", "llm_calls_change": "-21.4%", "pairs": 3},
"terminal_bench": {"success": "1/6 -> 3/6", "llm_calls_change": "-9.0%"}
},
"frozen_lake": {
Expand Down Expand Up @@ -92,43 +92,61 @@
}
},
"webshop": {
"seed": 3106,
"selection": "illustrative success-gain pair from the frozen 10-pair run",
"goal": "Men's lace-up boots, black, size 11.5, under $160",
"menu_provenance": "The full Jev candidate menu was not persisted; these four controls are confirmed by executed actions in the paired trace.",
"highlighted_actions": [
{"action": "click[black]", "evidence": "executed by Jev", "order": 1},
{"action": "click[11.5]", "evidence": "executed later by LLM", "order": 2},
{"action": "click[buy now]", "evidence": "executed later by LLM", "order": 3},
{"action": "click[features]", "evidence": "visible in paired baseline", "order": null}
"seed": 3107,
"selection": "positive pair from the LLM-authored candidate supplement",
"goal": "Women's long-sleeve blazer, z-dark green, small, under $80",
"candidate_groups": [
{
"subgoal": "Select z-dark green color",
"actions": ["click[z-dark green]", "click[z-army green]", "click[z-khaki]"],
"selected": "click[z-dark green]",
"confidence": 1.0,
"hidden_state": {"color": "z-dark green"}
},
{
"subgoal": "Select small size",
"actions": ["click[x-small]", "click[small]", "click[medium]"],
"selected": "click[small]",
"confidence": 1.0,
"hidden_state": {"color": "z-dark green", "size": "small"}
}
],
"baseline": {
"reward": 0,
"llm_calls": 11,
"tokens": 47360,
"wall_time_seconds": 20.9,
"reward": 1,
"llm_calls": 9,
"tokens": 38852,
"wall_time_seconds": 18.54,
"path": [
"search",
"open product",
"select color",
"select size",
"inspect features",
"back",
"search again",
"step budget ends"
"repeat color and size",
"buy now"
]
},
"jev_harness": {
"reward": 1,
"llm_calls": 5,
"tokens": 15705,
"wall_time_seconds": 9.9,
"llm_calls": 4,
"tokens": 14256,
"wall_time_seconds": 7.83,
"path": [
{"controller": "llm", "action": "search"},
{"controller": "llm", "action": "open matching product"},
{"controller": "jev", "action": "click[black]", "confidence": 0.89, "action_is_effective": false},
{"controller": "llm", "action": "click[11.5]"},
{"controller": "jev", "action": "click[z-dark green]", "confidence": 1.0, "action_is_valid": true},
{"controller": "jev", "action": "click[small]", "confidence": 1.0, "action_is_valid": true},
{"controller": "llm", "action": "click[buy now]"}
]
},
"source": "runs/agent-harness/webshop-opus-27b-holdout.json",
"source_sha256": "54cbceb65ebe2a81f5fe27ff3f634ff178918f999adcc1458459a309075dd96e"
"three_pair_result": {
"baseline_success": 0.6667,
"jev_harness_success": 1.0,
"llm_calls_change": -0.2143,
"tokens_change": -0.1430,
"wall_time_change": -0.1501
},
"source": "runs/agent-harness/webshop-llm-candidates-supplement-v2.json",
"source_sha256": "0363fb8d6fcc241fbe7e1d3eed5b0d4d9fc571a76fee204dfe24b9b3833bbd8e"
}
}
Binary file added docs/demos/jev-decision-frozen-lake-v2.gif
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/demos/jev-decision-terminal-v2.gif
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/demos/jev-decision-webshop-v2.gif
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
13 changes: 12 additions & 1 deletion docs/experiments/AGENT_HARNESS_FRONTIER_PROTOCOL.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,9 @@ the realized coverage must both be shown. The current Terminal-Bench candidate p
is in the D3 family: routine decisions go to Jev, precise non-routine edits can execute
directly, and composable execution is capped at two commands. The existing cross-domain
optional-delegation matrix is D2 because the frontier autonomously gates delegation.
D1 and D4 are proposed ablations and have no result yet.
D1 remains proposed. D4 has a three-pair WebShop supplement: the LLM generated bounded
color/size candidate groups, Jev executed the selections, and the LLM retained search,
verification, and `buy now`.

An **eligible decision** is a low-risk, reversible action for which the LLM produced at
least two locally valid options. Examples include inspection commands, waiting versus
Expand Down Expand Up @@ -152,6 +154,9 @@ Recommended showcase panels based on evidence available now:
- **Real-web mechanism:** WebArena task 264 popup-aware smoke, where both passed and the
observed trace reduced Opus calls 8 to 2 and tokens 76,674 to 13,407. Label it a
single-task smoke; the frozen six-task aggregate is 50% to 50% with worse calls/tokens.
- **Structured-interaction mechanism:** WebShop seed 3107, where the LLM generated
three color and three size candidates, Jev selected the exact options, reward stayed
at 1, and calls/tokens/time changed 9→4, 38,852→14,256, and 18.54s→7.83s.
- **Open-terminal candidate trace:** reserve this panel for a completed final-protocol
paired run with full candidate logging. The v3 `regex-log` run is an exploratory
quality-recovery candidate (0 to 1 reward and 30 to 18 calls), but tokens rose and v3
Expand All @@ -167,6 +172,12 @@ success but no efficiency win. WebShop improves from 50% to 60% in one ten-pair
but its confidence intervals cross zero. See
[`results/agent-harness-v1/formal-matrix.md`](../../results/agent-harness-v1/formal-matrix.md).

The later three-pair WebShop D4 supplement uses LLM-authored menus and excludes purchase,
navigation, and information tabs from Jev control. Success is 2/3→3/3, mean frontier
calls 9.33→7.33, total tokens 114,388→98,030, and mean time 16.76s→14.24s. Seed 3105
violated the 2–4 candidate bound and safely fell back to the LLM. This supplement is
stored separately and does not overwrite the ten-pair D2 matrix.

Terminal-Bench development history demonstrates why the delegation level must be
controlled rather than maximized:

Expand Down
Loading
Loading