English
Related papers

Related papers: Omni-WorldBench: Towards a Comprehensive Interacti…

200 papers

Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existing OLM benchmarks remain anchored to static, accuracy-centric tasks, leaving a critical gap in…

Artificial Intelligence · Computer Science 2026-03-18 Tianyu Xie , Jinfa Huang , Yuexiao Ma , Rongfang Luo , Yan Yang , Wang Chen , Yuhui Zeng , Ruize Fang , Yixuan Zou , Xiawu Zheng , Jiebo Luo , Rongrong Ji

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Ziqi Huang , Yinan He , Jiashuo Yu , Fan Zhang , Chenyang Si , Yuming Jiang , Yuanhan Zhang , Tianxing Wu , Qingyang Jin , Nattapol Chanpaisit , Yaohui Wang , Xinyuan Chen , Limin Wang , Dahua Lin , Yu Qiao , Ziwei Liu

While world models have emerged as a cornerstone of embodied intelligence by enabling agents to reason about environmental dynamics through action-conditioned prediction, their evaluation remains fragmented. Current evaluation of embodied…

Text-to-video (T2V) models have shown remarkable performance in generating visually reasonable scenes, while their capability to leverage world knowledge for ensuring semantic consistency and factual accuracy remains largely understudied.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Yubin Chen , Xuyang Guo , Zhenmei Shi , Zhao Song , Jiahao Zhang

Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. However, evaluations of unified multimodal models (UMMs) remain decoupled, assessing their understanding and generation…

Artificial Intelligence · Computer Science 2025-12-22 Kai Liu , Leyang Chen , Wenbo Li , Zhikai Chen , Zhixin Wang , Renjing Pei , Linghe Kong , Yulun Zhang

As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations, including uncontrollable task complexity, extensive manual…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Wendong Bu , Yang Wu , Qifan Yu , Minghe Gao , Bingchen Miao , Zhenkui Zhang , Kaihang Pan , Yunfei Li , Mengze Li , Wei Ji , Juncheng Li , Siliang Tang , Yueting Zhuang

Recent advances in large-scale video world models have enabled increasingly realistic future prediction, raising the prospect of using generated videos as scalable supervision for robot learning. However, for embodied manipulation,…

Recent advancements in multimodal large language models (MLLMs) have aimed to integrate and interpret data across diverse modalities. However, the capacity of these models to concurrently process and reason about multiple modalities remains…

World models are central to building AI agents capable of flexible reasoning and planning. Yet current evaluations (i) test only properties measurable from observed interactions, such as next-frame prediction or task return, and (ii) do not…

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Ziqi Huang , Fan Zhang , Xiaojie Xu , Yinan He , Jiashuo Yu , Ziyue Dong , Qianli Ma , Nattapol Chanpaisit , Chenyang Si , Yuming Jiang , Yaohui Wang , Xinyuan Chen , Ying-Cong Chen , Limin Wang , Dahua Lin , Yu Qiao , Ziwei Liu

Today's driving world models can generate remarkably realistic dash-cam videos, yet no single model excels universally. Some generate photorealistic textures but violate basic physics; others maintain geometric consistency but fail when…

Action-conditioned video prediction models (often referred to as world models) have shown strong potential for robotics applications, but existing approaches are often slow and struggle to capture physically consistent interactions over…

Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurately simulate physical phenomena remains a critical and unresolved challenge. This paper presents…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jing Gu , Xian Liu , Yu Zeng , Ashwin Nagarajan , Fangrui Zhu , Daniel Hong , Yue Fan , Qianqi Yan , Kaiwen Zhou , Ming-Yu Liu , Xin Eric Wang

Current world models lack a unified and controlled setting for systematic evaluation, making it difficult to assess whether they truly capture the underlying rules that govern environment dynamics. In this work, we address this open…

Machine Learning · Computer Science 2025-12-01 Xinyi Li , Zaishuo Xia , Weyl Lu , Chenjie Hao , Yubei Chen

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…

Multimodal Large Languages models have been progressing from uni-modal understanding toward unifying visual, audio and language modalities, collectively termed omni models. However, the correlation between uni-modal and omni-modal remains…

Computation and Language · Computer Science 2025-10-31 Chen Chen , ZeYang Hu , Fengjiao Chen , Liya Ma , Jiaxing Liu , Xiaoyu Li , Ziwen Wang , Xuezhi Cao , Xunliang Cai

Autonomous driving world models are expected to work effectively across three core dimensions: state, action, and reward. Existing models, however, are typically restricted to limited state modalities, short video sequences, imprecise…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Bohan Li , Zhuang Ma , Dalong Du , Baorui Peng , Zhujin Liang , Zhenqiang Liu , Chao Ma , Yueming Jin , Hao Zhao , Wenjun Zeng , Xin Jin

We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-visual streams. Unlike offline video understanding or text-prompted streaming QA,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Xudong Lu , Xueying Li , Annan Wang , Yang Bo , Jinpeng Chen , Zengliang Li , Nianzu Yang , Rui Liu , Xue Yang , Jingwen Hou , Hongsheng Li

Recent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based on inherent characteristics continues to hinder the progress…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Yiran Qin , Zhelun Shi , Jiwen Yu , Xijun Wang , Enshen Zhou , Lijun Li , Zhenfei Yin , Xihui Liu , Lu Sheng , Jing Shao , Lei Bai , Wanli Ouyang , Ruimao Zhang

The rapidly developing field of large multimodal models (LMMs) has led to the emergence of diverse models with remarkable capabilities. However, existing benchmarks fail to comprehensively, objectively and accurately evaluate whether LMMs…