Skip to content

Repository files navigation

JevAny: Your Jev from Any Model to Any Application

Homepage Checkpoints API docs Examples Python 3.12+ Tests License

🇺🇸 English | 🇨🇳 简体中文
🎮 Results & Demos | ⚡ Quickstart | 🤗 Models | 📊 Benchmarks | 📚 Docs

JevAny is open infra for System 1 decision model training and deployment, covering data preparation, model adaptation and evaluation. Use a released model or train on your own data to route support tickets, select tools, or choose a robot's next action. One API takes the state, question and candidate options, then directly returns a choice and its probabilities.

JevAny infra for System 1 decision model training, deployment and application integration

🎮 Results and Demos

JevAny checkpoints and baselines compared on Transfer and JevBench in side-by-side bar charts

Full benchmark results and evaluation details.

The following 30 examples are archived replays from an earlier compatible JevAny checkpoint. The current default release is JevAny-Qwen3.8-27B. Explore the cases, or try the playground to inspect recorded actions and option probabilities.

JevAny choosing actions across robotics, browser, software, laboratory and mobility tasks

📑 Table of Contents

⚡ 1. Quickstart

Use Python 3.12 or newer. Clone the repository and install the lightweight package:

git clone https://github.com/SimpleJev/JevAny.git
cd JevAny
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -e .

Keep this environment active and work from the repository root. Choose Training to build your own model or Deployment to use a released checkpoint. For a preview on CPU, try the playground.

🛠️ 1.1 JevAny Training

Train your own System 1 model on the same state and questions you send at inference, with a label for each question. Start with the bundled synthetic support tickets, then train on your own labelled data. The starter recipe uses Qwen3.5-0.8B on CUDA with BF16 and writes runs/my-jev:

python -m pip install -e '.[train]'
jevany data init --out data/starter
jevany data validate data/starter/train.jsonl
jevany train --config recipes/sft.toml --dry-run
jevany train --config recipes/sft.toml

After training, try the checkpoint on the included ticket request:

jevany decide examples/request.json --checkpoint runs/my-jev

Pass --data to train on your own JSONL data, or use recipes/finetune.toml to adapt the released 27B model. See the training guide for CPU settings, multimodal data and standard torchrun launches. For image/video training or fine-tuning the released 27B model, install .[train,multimodal].

After SFT, you can continue with experimental RLCR, which rewards correctness and probability calibration:

jevany train --config recipes/rlcr.toml

🚀 1.2 JevAny Deployment

Install the serving dependencies and start the released Qwen 4B model on a CUDA GPU. See the hardware and loading guide for memory requirements.

python -m pip install -e '.[serve,multimodal]'
jevany serve --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
  --device cuda --dtype bf16 --port 8008

The default path favors reproducibility. CUDA deployments can opt into BF16 LoRA merging, SDPA and torch.compile; the useful settings differ between 4B and 27B. See the inference acceleration guide for commands, H200 measurements and accuracy caveats.

To serve your training output, replace the checkpoint ID with runs/my-jev. Keep the server running. In a Python session using the same environment, send a ticket and the departments that can handle it:

from jevany import Choice, JevClient

jev = JevClient("http://127.0.0.1:8008")
result = jev.system_one(
    state={"ticket": "I was charged twice. Please help."},
    questions={
        "department": Choice(
            instructions="Which team should handle this?",
            criteria={"billing": "Payment problems", "shipping": "Delivery problems"},
        ),
    },
)
answer = result["answers"]["department"]
print("Selected team:", answer["choice"])
print("Probabilities:", answer["probabilities"])

choice is one of the department names; probabilities maps each name to its probability. Your application can use these fields to route the ticket or ask for review when the decision is uncertain. Use Noul for yes/no questions, such as whether a ticket needs urgent review, and Score for ordered levels, such as low, normal and high priority. See the API reference for all three question types.

For in-process inference, load a model in Python and use the same interface. For image and video inputs, follow the media setup.

🤗 2. Pretrained Models

Model Readout Intended use
 JevAny-Gemma-4B Pointer Compact Gemma release
 JevAny-Qwen3.5-4B Pointer Compact, flexible choice count
 JevAny-Qwen3.5-4B-Direct-Token Direct-token Best released 4B JevBench accuracy
 JevAny-Qwen3.8-27B Pointer Default; highest released accuracy
 JevAny-Muse-Glimmer-30B Pointer Muse Glimmer alternative

These LoRA adapters were trained with SFT on 1,772,725 text records containing 2,180,242 labelled decisions; see training compute and experiments for the setup. Full-parameter SFT and further post-training improvements are planned.

The corresponding base model is loaded separately and its license and access terms apply. Allow roughly twice the base parameter count in bytes for BF16 weights, plus runtime memory. See the hardware and loading guide.

Pointer and direct-token models share the same API. Pointer supports up to 4,096 options within the context limit; direct-token supports up to 255. See readout choices for training and accuracy tradeoffs.

📊 3. Benchmark Results

JevAny-Qwen3.8-27B leads both benchmarks and has the lowest NLL and Brier. Among 4B releases, direct-token leads on JevBench; pointer leads on Transfer.

JevAny checkpoints and baselines ranked by mean accuracy on Transfer and JevBench

Model Transfer ↑ JevBench ↑ NLL ↓ Brier ↓ ECE ↓
 Kev-4B 74.19% 75.32% 0.858 0.380 0.125
 Kev-27B 82.31% 85.28% 0.533 0.265 0.050
 Jev 1.13.0 85.37% 86.58% 0.644 0.212 0.033
 Laya 52.29% 58.01% 1.264 0.615 0.127
JevAny releases
 JevAny-Gemma-4B 70.84% 77.49% 0.706 0.369 0.056
 JevAny-Qwen3.5-4B 78.68% 80.09% 0.587 0.297 0.035
 JevAny-Qwen3.5-4B-Direct-Token 78.20% 80.95% 0.564 0.291 0.029
 JevAny-Muse-Glimmer-30B 83.46% 87.45% 0.464 0.229 0.032
 JevAny-Qwen3.8-27B 86.04% 90.04% 0.388 0.195 0.026

NLL, Brier and ECE are measured on Transfer.

Full results and protocols · Machine-readable results · Method and ablation report

🕹️ 4. Examples & Test Environments

The playground includes the three environments below. These GIFs preserve historical model actions and option probabilities; run the current JevAny-Qwen3.8-27B checkpoint with the commands in the playground guide.

Use a Franka gripper to grasp, align and insert a peg, checked by PyBullet contact physics.

Robot browser replay showing the Franka arm inserting a peg, recorded model probabilities and physical success checks

Clear the final room by defeating the enemies on the left and right, then move forward. The environment uses ViZDoom and the included Freedoom assets.

Doom checkpoint replay: kill both enemies, then advance

Gather wood, craft tools and mine stone while managing health and supplies.

Crafter browser replay showing resource gathering, crafting actions and progress through four goal milestones

🎮 4.4 Try the playground

From the Quickstart environment, start the playground:

jevany demo

Open http://127.0.0.1:8090 and choose Replay to watch a recorded run. The bundled recordings play locally on CPU.

To run your model, keep the server from Deployment running. Stop the replay viewer with Ctrl+C, install the optional game engines, then restart the playground with the server address:

python -m pip install -e '.[demo]'
jevany demo --base-url http://127.0.0.1:8008 --text-only

Choose Run model in the browser, or Play yourself to control the game. Live control sends text state to the model. Robot control uses the .[robotics] extra. See the playground guide for platform requirements and environment APIs, or integrations to combine JevAny decisions with an LLM planner.

🧩 5. Supported Model Families

Model IDs, supported inputs and setup requirements.

26 supported models across Qwen, Gemma, Muse, Mistral, GLM, Nemotron and Llama

📚 6. Documentation and Contributing

Training · Deployment · API · Data · Evaluation · Contributing

To contribute a model adapter, evaluation or application example, start with the contribution guide. The technical report describes the model design, multimodal path, experiments and open questions.

Code and starter data are Apache-2.0. Some components are adapted from Kev; see NOTICE and ACKNOWLEDGEMENTS.md. Base models and upstream datasets retain their own terms.

About

Open infrastructure for training, evaluating, and deploying System 1 decision models across language and multimodal backbones.

Topics

Resources

Contributing

Stars

18 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages