Ruoxuan Feng*, Yutong Chen*, Ruihua Song, Huan Yang, Zhongyuan Wang, Guocai Yao, Di Hu✉
*Equal contribution ✉Corresponding author
Humans build an understanding of the physical world through an active process: when sensory evidence is insufficient, we decide what information is missing, how to acquire it, and when enough evidence has been obtained. Existing multi-sensory robot systems, in contrast, mostly integrate whatever sensory inputs they are given.
ROMA is an LLM-based system for Real-World Object-Centric Multi-Sensory Active Perception. It integrates vision, audio, touch, and force into a reasoning-interaction-feedback loop: the LLM identifies the missing evidence and selects the target object, the interaction (lift, press, collide, shake, rotate, squeeze), and the sensory modalities, while a physical interface executes the interaction and returns the multi-sensory feedback.
This repository accompanies the paper and contains:
- ROMA-7B, a multi-sensory LLM built on Qwen2.5-Omni with action / modality tokens, an audio branch, and an AnyTouch 2 tactile branch, trained with multi-sensory alignment followed by active-perception SFT.
- ROMI-2K, a real-world multi-sensory object interaction dataset covering nearly 2,000 objects and 6 atomic interactions with synchronized visual, audio, tactile, and force feedback. (Coming Soon!)
- ROMA Bench, 2,100 scene-level tasks (single-chain, multi-chain, and intent-driven) for evaluating active perception. (Coming Soon!)
- A local web demo that streams ROMA's active-perception process on a real tabletop scene.
| Model | ROMA Bench (total) | Real-world execution (total) |
|---|---|---|
| GPT-5.4 | 43.5 | 40.2 |
| Gemini 3.5 Flash | 53.0 | 41.7 |
| Qwen 2.5-Omni | 22.0 | 23.5 |
| Qwen 3-Omni | 19.7 | 23.5 |
| ROMA-7B | 72.9 | 61.4 |
Task success rate (%). See the paper for the full results.
The demo runs locally. Given the initial scene image and a question, ROMA reasons step by step; whenever it emits an action token, the corresponding recorded feedback (wrist image, tactile images, contact sound, gravity force) is loaded from resources/example_data/1 and fed back to the model, and generation continues until it answers. Everything is streamed to the page:
- every
graspis drawn on the scene image as a numbered, colored box; - each interaction is shown as an Interaction & Feedback card with the returned sensory evidence;
- the final answer is highlighted.
See Usage to launch it.
conda create -n roma python=3.10.20
conda activate roma
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt(Optional) Install FlashAttention for faster inference. If installed, enable it by uncommenting the attn_implementation="flash_attention_2" line in demo.py.
Before running, pull GeWu-Lab/ROMA from Hugging Face into the resources directory:
hf download GeWu-Lab/ROMA --repo-type dataset --local-dir resourcesThe ROMA-7B checkpoint directory (ROMA-Qwen2.5-Omni-7B) is expected to contain:
ROMA-Qwen2.5-Omni-7B/
├── xxx.safetensors # base Qwen2.5-Omni-7B model
├── anytouch2.pth # tactile encoder
├── audio.bin # fine-tuned audio adapter and encoder
├── tactile.bin # fine-tuned tactile adapter
└── ROMA-LLM.bin # ROMA LLM weights
The demo reads the checkpoint directory from MODEL_PATH_2_5 in config/global_name.py, which defaults to resources/models/ROMA-Qwen2.5-Omni-7B. If you downloaded the checkpoint elsewhere, edit this path accordingly.
python demo.py --gpu 0 --port 7860Then open http://localhost:7860, type a question (or pick an example), and press Ask ROMA. The base model and all adapters are loaded from the checkpoint directory in the same order as during training (base Qwen2.5-Omni, tactile encoder, new tokens, audio adapter, tactile adapter, ROMA LLM). Options: --gpu, --host, --port, --share (create a public Gradio link).
ROMA
├── demo.py # local web demo (Gradio)
├── qwen25omni/ # Qwen2.5-Omni modified with the tactile modality
├── anytouch2/ # tactile encoder
├── my_qwen_omni_utils/ # audio / vision / tactile input processing
├── config/, utils/ # arguments, DeepSpeed and checkpoint utilities
└── resources/ # checkpoints, datasets, and the demo example scene
- Release the local web demo
- Release ROMA-7B checkpoint and the ROMI-2K dataset
- Release the ROMA Bench evaluation code
- Release the training code (multi-sensory alignment and active-perception SFT)
- Release the physical interface (grasp interface and interaction execution)
If you find this work useful, please consider citing:
@article{feng2026roma,
title = {ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception},
author = {Feng, Ruoxuan and Chen, Yutong and Song, Ruihua and Yang, Huan and
Wang, Zhongyuan and Yao, Guocai and Hu, Di},
journal = {arXiv preprint arXiv:2610.06955},
url = {https://arxiv.org/abs/2610.06955},
year = {2026}
}ROMA builds on Qwen2.5-Omni, AnyTouch 2, and AnyGrasp. We thank the authors for their open-source contributions.


