English
Related papers

Related papers: Rethinking Video Generation Model for the Embodied…

200 papers

The advancement of embodied AI has unlocked significant potential for intelligent humanoid robots. However, progress in both Vision-Language-Action (VLA) models and world models is severely hampered by the scarcity of large-scale, diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Pei Yang , Hai Ci , Yiren Song , Mike Zheng Shou

Video generative models are increasingly used as world models for robotics, where a model generates a future visual rollout conditioned on the current observation and task instruction, and an inverse dynamics model (IDM) converts the…

Robotics · Computer Science 2026-03-25 Ruixiang Wang , Qingming Liu , Yueci Deng , Guiliang Liu , Zhen Liu , Kui Jia

Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for embodied intelligence. A major challenge in this setting is…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Yiren Song , Xiyao Deng , Pei Yang , Yihan Wang , Mike Zheng Shou

Generative video models are increasingly studied as implicit world models, yet evaluating whether they produce physically plausible 3D structure and motion remains challenging. Most existing video evaluation pipelines rely heavily on human…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Jiaxin Wu , Yihao Pi , Yinling Zhang , Yuheng Li , Xueyan Zou

Motion generation, the task of synthesizing realistic motion sequences from various conditioning inputs, has become a central problem in computer vision, computer graphics, and robotics, with applications ranging from animation and virtual…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Aliasghar Khani , Arianna Rampini , Bruno Roy , Larasika Nadela , Noa Kaplan , Evan Atherton , Derek Cheung , Jacky Bibliowicz

Simulating robot-world interactions is a cornerstone of Embodied AI. Recently, a few works have shown promise in leveraging video generations to transcend the rigid visual/physical constraints of traditional simulators. However, they…

Robotics · Computer Science 2026-03-18 Mutian Xu , Tianbao Zhang , Tianqi Liu , Zhaoxi Chen , Xiaoguang Han , Ziwei Liu

The training of controllable text-to-video (T2V) models relies heavily on the alignment between videos and captions, yet little existing research connects video caption evaluation with T2V generation assessment. This paper introduces…

Artificial Intelligence · Computer Science 2025-05-20 Xinlong Chen , Yuanxing Zhang , Chongling Rao , Yushuo Guan , Jiaheng Liu , Fuzheng Zhang , Chengru Song , Qiang Liu , Di Zhang , Tieniu Tan

Despite the remarkable progress in text-driven video editing, generating coherent non-rigid deformations remains a critical challenge, often plagued by physical distortion and temporal flicker. To bridge this gap, we propose NRVBench, the…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Bingzheng Qu , Kehai Chen , Xuefeng Bai , Jun Yu , Min Zhang

Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true multimodal reasoning? Existing benchmarks fail to address this question rigorously, as…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xiaotian Zhang , Jianhui Wei , Yuan Wang , Jie Tan , Yichen Li , Yan Zhang , Ziyi Chen , Daoan Zhang , Dezhi YU , Wei Xu , Songtao Jiang , Zuozhu Liu

The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consistency. However, despite recent breakthroughs such as Veo…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Harold Haodong Chen , Disen Lan , Wen-Jie Shu , Qingyang Liu , Zihan Wang , Sirui Chen , Wenkai Cheng , Kanghao Chen , Hongfei Zhang , Zixin Zhang , Rongjin Guo , Yu Cheng , Ying-Cong Chen

Text-to-video generative models have made significant strides in recent years, producing high-quality videos that excel in both aesthetic appeal and accurate instruction following, and have become central to digital art creation and user…

Machine Learning · Computer Science 2025-05-02 Xuyang Guo , Jiayan Huo , Zhenmei Shi , Zhao Song , Jiahao Zhang , Jiale Zhao

Multimodal Retrieval-Augmented Generation (MRAG) enables Multimodal Large Language Models (MLLMs) to generate responses with external multimodal evidence, and numerous video-based MRAG benchmarks have been proposed to evaluate model…

Computation and Language · Computer Science 2025-10-13 Kaiwen Wei , Xiao Liu , Jie Zhang , Zijian Wang , Ruida Liu , Yuming Yang , Xin Xiao , Xiao Sun , Haoyang Zeng , Changzai Pan , Yidan Zhang , Jiang Zhong , Peijin Wang , Yingchao Feng

Accurate robot segmentation is a fundamental capability for robotic perception. It enables precise visual servoing for VLA systems, scalable robot-centric data augmentation, accurate real-to-sim transfer, and reliable safety monitoring in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Haiyang Mei , Qiming Huang , Hai Ci , Mike Zheng Shou

Recent years have seen rapid advances in AI-driven image generation. Early diffusion models emphasized perceptual quality, while newer multimodal models like GPT-4o-image integrate high-level reasoning, improving semantic understanding and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Yifan Chang , Yukang Feng , Jianwen Sun , Jiaxin Ai , Chuanhao Li , S. Kevin Zhou , Kaipeng Zhang

Automatic identification of events and recurrent behavior analysis are critical for video surveillance. However, most existing content-based video retrieval benchmarks focus on scene-level similarity and do not evaluate the action…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Oriol Rabasseda , Zenjie Li , Kamal Nasrollahi , Sergio Escalera

The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual and comprehension abilities. Existing benchmarks, however,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Purui Bai , Tao Wu , Jiayang Sun , Xinyue Liu , Huaibo Huang , Ran He

The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Yuhang Yang , Ke Fan , Shangkun Sun , Hongxiang Li , Ailing Zeng , FeiLin Han , Wei Zhai , Wei Liu , Yang Cao , Zheng-Jun Zha

Unified video models exhibit strong capabilities in understanding and generation, yet they struggle with reason-informed visual editing even when equipped with powerful internal vision-language models (VLMs). We attribute this gap to two…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xinyu Liu , Hangjie Yuan , Yujie Wei , Jiazheng Xing , Yujin Han , Jiahao Pan , Yanbiao Ma , Chi-Min Chan , Kang Zhao , Shiwei Zhang , Wenhan Luo , Yike Guo

With the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models. However, most benchmarks predominantly assess…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Kunchang Li , Yali Wang , Yinan He , Yizhuo Li , Yi Wang , Yi Liu , Zun Wang , Jilan Xu , Guo Chen , Ping Luo , Limin Wang , Yu Qiao

Recent video multimodal large language models achieve impressive results across various benchmarks. However, current evaluations suffer from two critical limitations: (1) inflated scores can mask deficiencies in fine-grained visual…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Jiahao Meng , Tan Yue , Qi Xu , Haochen Wang , Zhongwei Ren , Weisong Liu , Yuhao Wang , Renrui Zhang , Yunhai Tong , Haodong Duan
‹ Prev 1 4 5 6 7 8 10 Next ›