Qianjun Xia
← All publications

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Guanxiong Chen*, Qianjun Xia*, Jiawei Peng*, Heng Zhang*, Pengyu Jing*, Bole Ma*, Justin Qian, Yixian Cheng, Ziyi Jiao, Bingyang Zhou, Yiduo Qu, Luoxin Ye, Kaifeng Zhang, Kunyi Wang, Weijia Zeng, Yunuo Chen, Pengzhi Yang, Ziqiu Zeng, Siyuan Luo, Huamin Wang, Chao Liu, Alan Yuille, Fan Shi, Changxi Zheng, Yunzhu Li, Chenfanfu Jiang, Peter Yichen Chen

PreprintarXiv preprint2026
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

* Equal technical contribution.

Converting a real interaction into a simulation usually means hand-tuning models and aligning coordinate frames by hand, once per scene. This work replaces that with vision-language agents that recover geometries, object states and physical parameters directly from the recording.

One pipeline covers rigid-object manipulation, deformable-object interaction and humanoid motion — cases that previously needed separate, specialised Real2Sim methods. An open-weight VLM backend reaches success rates comparable to frontier models, which makes the pipeline cheap enough to run at dataset scale. The resulting twins are used downstream for policy learning and for evaluating policies against something that behaves like the world did.

Videos, an interactive scene viewer and the 25-episode DROID gallery are on the project site.