Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions NOTICE
Original file line number Diff line number Diff line change
Expand Up @@ -12,3 +12,6 @@ This product includes software developed by third parties:
evals/2048/style/fonts are Clear Sans by Intel Corporation (Apache-2.0)
- evals/millionaire/questions.json, fetched at first run: Open Trivia Database
(https://opentdb.com), CC BY-SA 4.0
- s1a/agents/_sokoban.py: adapted Valen Sokoban simulator (Liuziyu77/Valen), Apache-2.0.
Sokoban levels and test references: Valen-Team/Valen-Eval-Game, Apache-2.0.
Pinned sources and modifications: THIRD_PARTY_LICENSES.md
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -121,7 +121,8 @@ The gates and the templates: [docs/skills.md](docs/skills.md#build-a-system-1-ag
- `desktop`: computer use, any Windows or macOS app window through [Cua Driver](https://cua.ai/docs/cua-driver).
- `ticket_router`: 30 labelled support tickets to five queues.
- `alfworld`: household tasks in text, with the AI2-THOR scene in the replays.
- `game2048`, `millionaire`, `blackjack`: games with a score per episode.
- `game2048`, `millionaire`, `blackjack`, `sokoban`: games with a score per episode.
Sokoban uses text boards from Valen’s selected evaluation levels; see [the protocol](evals/sokoban/README.md).
- `injection_guard`: a rail that answers one question at a hook of a running agent and fails closed.

Every agent runs on `jev`, `laya` or `cua`, and on the chat model for the comparison. Flags, run commands and
Expand Down
13 changes: 13 additions & 0 deletions THIRD_PARTY_LICENSES.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,3 +39,16 @@ Fonts under `evals/2048/style/fonts` are Clear Sans by Intel Corporation, Apache

`evals/millionaire/questions.json`, fetched on the first Millionaire run, holds questions from the
[Open Trivia Database](https://opentdb.com), Creative Commons Attribution-ShareAlike 4.0 International.

## Valen Sokoban

`s1a/agents/_sokoban.py` is adapted from
[Liuziyu77/Valen](https://github.com/Liuziyu77/Valen/blob/c96c4f736c6928d2cb1166bc105a24cebdfcdc40/evaluation/sokoban/env.py),
commit `c96c4f736c6928d2cb1166bc105a24cebdfcdc40`, Apache-2.0.
Changes: removed unused hash, inverse-action and clone helpers; applied repository formatting.

`s1a/agents/_data/sokoban_levels.jsonl` and `tests/fixtures/sokoban_reference.jsonl` are unchanged copies of
`eval_sokoban/levels.jsonl` and `eval_sokoban/reference.jsonl` from
[Valen-Team/Valen-Eval-Game](https://huggingface.co/datasets/Valen-Team/Valen-Eval-Game/tree/4f25038cd30d9a96aa377eea176a4d2e1680a35c),
revision `4f25038cd30d9a96aa377eea176a4d2e1680a35c`, declared Apache-2.0.
The full Apache-2.0 license text is this repository's `LICENSE`.
8 changes: 8 additions & 0 deletions docs/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@ hook of a running agent. The injection guard rail fails closed: a decision error
| `game2048` | tool | score and largest tile at a move cap | `s1a run game2048 --model jev --rethink on --episodes 10` |
| `millionaire` | tool | winnings on a 15-question quiz ladder | `s1a run millionaire --model jev --rethink off --episodes 5` |
| `blackjack` | tool | payoff per hand (RLCard) | `s1a run blackjack --model jev --rethink off --episodes 100` |
| `sokoban` | tool | solved fraction on 100 selected Valen levels, using text boards | `s1a run sokoban --model jev --rethink off --episodes 100` |
| `injection_guard` | rail | precision and recall on a labelled injection set | `s1a run injection_guard` |

Every tool agent takes `--model jev|laya|cua|llm|random|rule`, `--rethink on|off`, `--episodes N`, `--seed S`,
Expand Down Expand Up @@ -50,3 +51,10 @@ shipped set is 30 tickets in `s1a/agents/_data/ticket_router_eval.jsonl`, each w
order status; the label is read by the scorer alone. `--dataset` points the agent at a
JSONL of your own with the same fields, and `--batch-size` caps the tickets per episode. The rules text the model
reads is `RULES` in `s1a/agents/ticket_router.py`.

## Sokoban

`uv run s1a run sokoban --model jev --rethink off --episodes 100` plays the bundled 100 Valen levels as text
boards. No extra is needed. `--seed` is the zero-based level offset; the selected range must fit within 100
levels. The score is 1 for solved and 0 otherwise. Use `--model random` for an offline smoke run; there is no
rule baseline. Rethink must be off. See [protocol and provenance](../evals/sokoban/README.md).
1 change: 1 addition & 0 deletions evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ folder per episode). `uv run python -m evals.table evals/results` aggregates the

| eval | measures | run |
|---|---|---|
| Sokoban (text board, selected Valen levels) | solved fraction; text adaptation, not the upstream visual benchmark | `s1a run sokoban --model jev --rethink off --episodes 100` |
| Blackjack (RLCard) | payoff per hand | `s1a run blackjack --model jev --rethink off --episodes 100` |
| 2048 (the MIT game, self-hosted) | score and largest tile at a move cap | `s1a run game2048 --model jev --rethink on --episodes 10` |
| Millionaire (self-hosted quiz, Open Trivia DB) | winnings; the 50:50 lifeline is a candidate the model may pick; six ladders: `--seed` plus `--episodes` stays at or below 6 | `s1a run millionaire --model jev --rethink off --episodes 5` |
Expand Down
33 changes: 33 additions & 0 deletions evals/sokoban/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
# Sokoban (text-board adaptation)

```bash
uv run s1a run sokoban --model random --rethink off --episodes 1 --max-steps 10
uv run s1a run sokoban --model jev --rethink off --episodes 100
```

No additional dependency, browser, dataset download or GPU is needed for the environment.
The model receives an ASCII board and remaining step budget, with the legend in its rules.
The four directions remain available even when blocked; blocked moves consume a step.
There is no solver, legal-action mask, undo, rule baseline or rethink assistance.

`--seed` selects the zero-based level offset, with consecutive levels per episode; the
range must stay within the bundled 100 levels. The episode ends on success or after at
most 200 actions (`--max-steps` can lower this); `--timeout` also caps wall time.
Score is 1 for a solved board and 0 otherwise, so `mean_score` is the solved fraction
among non-error episodes. Always report errors separately. Episode metadata records the
level ID, text observation mode, pushes and whether the environment step limit was reached.
The standard episode records include the final board, steps, elapsed time and decisions.

These are Valen's **selected** 100 evaluation levels (35 easy, 65 medium), not an unbiased
sample. They retained successful 2B RLCD cases before filling the set. Results here use
symbolic text boards and must not be compared as equivalent to Valen's image-only evaluation.
No learned-model performance has been measured by this import.

The simulator comes from Valen, not vllm-jev's demo assets. The level definitions are
bundled without changes; evaluator-only reference solutions live under `tests/fixtures`
and are never loaded by the agent. Tests replay all 100 reference solutions to validate
transitions and level compatibility; this is not model accuracy.

See [third-party provenance](../../THIRD_PARTY_LICENSES.md#valen-sokoban) for pinned
code/data revisions and [the upstream dataset card](https://huggingface.co/datasets/Valen-Team/Valen-Eval-Game)
for the original protocol and selection details.
Loading
Loading