Repository navigation
feat(eval): one tagged 88-task dataset with runtime tasks, plus bashkit_generate eval - #2664
Merged
Merged
Conversation
One bashkit_bash eval; tasks carry mode (agent|runtime), difficulty (basic|repo|hard) and an optional per-task max_turns. Reference solutions live in data/solutions.jsonl and are required for every non-basic task.
The agent-loop eval hides how often a model's first script fails on bashkit. bashkit_generate gives the model the task plus a fixed sandbox description, takes ONE bash script from a single reply (no tools, no feedback), runs it once on build_task_bash and scores it with the shared expectations scorer. - generate.rs: subject, documented extraction rule (bash/sh/shell fence, else untagged fence, else whole reply), 60s run limit, metrics (script_found, extracted, script_bytes, exit_code, timing); no script scores 0, provider errors are infra errors - providers omit tools when none are offered - 15 tasks (basic/hard) in data/generate-tasks.jsonl with single-script references; tests require references to pass and `true` to fail - just eval-generate, knowledge/operations/eval.md Generate Eval section
Turso writes WAL-mode headers (bytes 18/19 = 2). The Memory backend persists only the checkpointed main file, so the image is a valid legacy-mode database; mark it as one. CPython's sqlite3 module in the WASI guest has no WAL and rejected the file as "file is not a database", so python3 could not open a db written by sqlite3. Also record known sqlite3 shell divergences (csv quoting, exit codes, .import) found by the eval runtime tasks.
…ltin An executable file with #!/usr/bin/env python3 (or #!/usr/bin/python3, #!/usr/bin/awk -f) run by path or PATH lookup had its shebang stripped and its body parsed as bash, so agent-built Python CLIs failed with a parse error. A shebang naming a registered non-shell builtin now runs that builtin with the script path as its argument (Linux semantics: one optional argument; env -S splits). bash/sh, unknown interpreters and shell-only builtins keep running the content as bash. Adds CPython integration tests for shebang scripts and for python3 reading/writing a database made by the sqlite builtin.
Every eval Bash now runs real CPython 3.14 (cpython feature) and the sqlite builtin (sqlite feature, opt-in env set); the tool prompt the model sees registers the same builtins, so it lists their hints instead of "python/python3 not available". bashkit-replay uses the same runtimes. 12 runtime tasks (tag runtime, just eval-runtime): python3 file processing (CSV rollup, JSONL flatten, sessionization, encodings, TOML render, a 12k-line log), a bash+python report script, sqlite3 (CSV load + report, a reusable migration file, python reading a sqlite3-built db) and state across calls (a CLI on PATH, an idempotent ingester). Tasks gain a hidden verify script, run after the agent in a fresh Bash on the same filesystem; its output (/.eval/verify.out) is scored, so a built tool is re-run on unseen input. Goldens come from the generator, references pass in bashkit and against real python 3.13, sqlite3 3.45 and bash 5.2. The untouched-fixture test covers runtime tasks.
…up by task count Claude-Session: https://claude.ai/code/session_015PWfrXwJG9NsMJqio8M1gW
Deploying with
|
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs |
bashkit | 40956d8 | Commit Preview URL Branch Preview URL |
Oct 09 2026, 11:42 PM |
The python3 setup took ~7s in debug and timed out (30s) under coverage instrumentation. The awk generator writes the byte-identical log (same md5 in bashkit and real bash) in under 1s. Claude-Session: https://claude.ai/code/session_015PWfrXwJG9NsMJqio8M1gW
Instrumented CPython runs ~10x slower; replaying rt_py_large_log's python3 reference exceeded the 30s shell and python3 limits under tarpaulin. cfg(tarpaulin) builds get 300s; eval runs keep defaults. Claude-Session: https://claude.ai/code/session_015PWfrXwJG9NsMJqio8M1gW
bashkit-eval now enables bashkit's cpython and sqlite features, so workspace feature unification adds the CPython/sqlite tests to the bashkit integration binary under coverage, which then ran past 300s. Claude-Session: https://claude.ai/code/session_015PWfrXwJG9NsMJqio8M1gW
This reverts commit e72bc40.
This reverts commit ac458aa.
bashkit-eval enables bashkit's cpython feature; under --workspace feature unification that pulled CPython tests into the instrumented bashkit integration binary, which then timed out. Excluding it keeps coverage on main's feature set; eval tests still run in CI's test job. Claude-Session: https://claude.ai/code/session_015PWfrXwJG9NsMJqio8M1gW
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
One tagged dataset. The basic, repo and hard tasks are merged into
bashkit_bash, 88 tasks in total. Each task is taggedmode=agent|runtimeanddifficulty=basic|repo|hard, and carries its ownmax_turns. Subsets run with--tag(just eval-repo,just eval-hard,just eval-runtime).12 runtime tasks (
rt_*). In these, bashkit is the agent's execution runtime:The eval
Bashnow has CPython and sqlite. A new hiddenverifyscript re-runs what the model built against unseen input.New
bashkit_generateeval. It has 15 one-shot tasks, run withjust eval-generate. The model replies with a single script, which bashkit runs once and the shared expectations scorer scores.Bashkit fixes found by the reference solutions:
#!/usr/bin/env python3) now runs that builtin.Results. The README, the eval README, the site and the knowledge log are updated. The homepage table now lists only runs on the newest run's task count, so 58-task and 88-task scores never mix.
Why
The old dataset only measured bash as an agent tool. Agents also use bashkit as a runtime (python3, sqlite, persistent VFS) and to execute generated scripts. Keeping one tagged dataset avoids splitting it into four tracks; generation gets its own eval because it is scored differently.
Before / After
Before, the 58-task set was saturated, with three models at 58/58.
After, on 2026-10-09:
GPT-6.1 Sol and GPT-6 Luna are not in this run because the OpenAI key ran out of credit. Their cases can be filled in later with
mira run ... --resume <run_id>.Risk
Checklist
bashkit_repo/bashkit_hardare replaced by tags)Generated by Claude Code