A benchmark and evaluation harness for LLM agents on scientific inverse problems in computational imaging.
Project Website | Tasks | Evaluation Harness | Evaluation Guide
Inverse-101 packages 57 computational imaging tasks and a unified evaluation harness for testing LLM agents. Each task asks an agent to reason about a scientific inverse problem, implement part or all of the imaging pipeline, and produce outputs that can be scored automatically.
The benchmark covers astronomy, biology, chemistry and materials, earth science, medicine, and physics. Tasks follow a common layout with README.md, plan/, src/, data/, and evaluation/ folders so agents can be compared consistently.
| Mechanism | CLI value | What it measures | Main score |
|---|---|---|---|
| Planning | planning or plan |
Can the model design the algorithm and module structure? | LLM-as-judge plan score |
| Function mode | function |
Can the model implement one module against unit tests? | pytest pass rate |
| End-to-end | end2end or end_to_end |
Can the model build a full reconstruction pipeline? | NCC / NRMSE |
After cloning and configuring credentials, choose a model, task, and mechanism:
python -m evaluation_harness run \
--model "$MODEL" \
--task ct_fan_beam \
--mechanism end2end \
--level L1 \
--base-url "$BASE_URL" \
--api-key "$API_KEY" \
--framework react \
--output results/end_to_endUse --dry-run to validate the command without calling an LLM:
python -m evaluation_harness run \
--model demo-model \
--task ct_fan_beam \
--mechanism end2end \
--dry-rungit clone https://github.com/starpacker/inverse-101.git
cd inverse-101
python -m pip install -r evaluation_harness/requirements.txt
python scripts/download_assets.py --task ct_fan_beamSet an OpenAI-compatible endpoint:
export MODEL="gpt-4o"
export BASE_URL="https://api.openai.com/v1"
export API_KEY="your-api-key"On Windows PowerShell:
$env:MODEL="gpt-4o"
$env:BASE_URL="https://api.openai.com/v1"
$env:API_KEY="your-api-key"Optional but recommended for isolated execution:
docker build -t imaging101-sandbox -f evaluation_harness/Dockerfile .If Docker is unavailable, the harness falls back to a local temporary workspace.
You do not need to install every task package into your main Python environment. During a real run, the harness installs tasks/<task>/requirements.txt inside the Docker sandbox; without Docker, it creates an isolated local workspace and venv for that run.
For a quick preflight before spending tokens:
python -m evaluation_harness check-env --task ct_fan_beamFor machine-readable automation:
python -m evaluation_harness check-env --task ct_fan_beam --jsonFor manual debugging or repeated local experiments, create a reusable task venv:
python -m evaluation_harness setup-env --task ct_fan_beam --venv .venvs/ct_fan_beamUse --dry-run to print the setup commands without installing packages.
Environment tiers:
| Tier | Meaning | What to do |
|---|---|---|
| Tier 1 | Standard pip dependencies such as NumPy/SciPy/Matplotlib | Run normally |
| Tier 2 | Additional or specialized Python packages | Run check-env; use Docker or setup-env if needed |
| Tier 3 | Task-specific environment files such as Dockerfile, environment.yml, or ENVIRONMENT.md |
Read the task notes before running |
python -m evaluation_harness run \
--task ct_fan_beam \
--mechanism planning \
--model "$MODEL" \
--base-url "$BASE_URL" \
--api-key "$API_KEY" \
--output results/planningThe model writes plan/approach.md and plan/design.md; the harness scores them against the task reference plan.
python -m evaluation_harness run \
--task ct_fan_beam \
--mechanism function \
--target-function physics_model \
--model "$MODEL" \
--base-url "$BASE_URL" \
--api-key "$API_KEY" \
--output results/function_modeThe model implements one module in src/; the harness runs the matching pytest file from evaluation/tests/.
python -m evaluation_harness run \
--task ct_fan_beam \
--mechanism end2end \
--level L1 \
--model "$MODEL" \
--base-url "$BASE_URL" \
--api-key "$API_KEY" \
--output results/end_to_endThe model builds the full pipeline and should produce output/reconstruction.npy. The harness scores reconstruction quality with NCC, NRMSE, PSNR, and SSIM where reference data are available.
End-to-end levels:
| Level | Agent receives |
|---|---|
L1 |
Task README, data, requirements |
L2 |
L1 plus plan/approach.md |
L3 |
L2 plus plan/design.md |
The repository keeps code, task definitions, and lightweight metadata in GitHub. Large task arrays, fixtures, and reference outputs are stored on Hugging Face at starpacker52/imaging-101 and tracked by assets_manifest.json.
Download assets for one task:
python scripts/download_assets.py --task ct_fan_beamDownload the full benchmark asset set:
python scripts/download_assets.py --allThe downloader restores files into the matching tasks/<task_name>/data/, tasks/<task_name>/evaluation/fixtures/, and tasks/<task_name>/evaluation/reference_outputs/ paths, then verifies SHA-256 hashes from the manifest.
inverse-101/
tasks/ 57 benchmark tasks
evaluation_harness/ CLI, agents, sandbox runners, scorers
docs/ evaluation and contribution guides
scripts/ batch runners and analysis utilities
assets_manifest.json Hugging Face asset index with size and SHA-256 hashes
config_llm.yaml example model endpoint configuration
tasks/<task_name>/
README.md problem statement and data description
requirements.txt task-specific Python dependencies
main.py full pipeline entry point
data/ observations and metadata
plan/ reference approach and design
src/ implementation modules
evaluation/ tests, metrics, fixtures, reference outputs
notebooks/ tutorial or exploratory notebook
New tasks should preserve the standard layout and include a clear problem statement, runnable pipeline, tests, and scoring assets. See docs/NEW_TASK_GUIDE.md.
@misc{inverse101,
title = {Inverse-101: Benchmarking LLM Agents for Scientific Computational Imaging Problems},
url = {https://github.com/starpacker/inverse-101},
year = {2026}
}See LICENSE if present in the distribution.