中文
相关论文

相关论文: DynamicVerse: A Physically-Aware Multimodal Framew…

200 篇论文

View-predictive generative models provide strong priors for lifting object-centric images and videos into 3D and 4D through rendering and score distillation objectives. A question then remains: what about lifting complete multi-object…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Wen-Hsuan Chu , Lei Ke , Katerina Fragkiadaki

Generating dynamic and interactive 3D trees has wide applications in virtual reality, games, and world simulation. However, existing methods still face various challenges in generating structurally consistent and realistic 4D motion for…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Yaokun Li , Lihe Ding , Xiao Chen , Guang Tan , Tianfan Xue

Understanding 3D scenes requires flexible combinations of visual reasoning tasks, including depth estimation, novel view synthesis, and object manipulation, all of which are essential for perception and interaction. Existing approaches have…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Wanhee Lee , Klemen Kotar , Rahul Mysore Venkatesh , Jared Watrous , Honglin Chen , Khai Loong Aw , Daniel L. K. Yamins

Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses,…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Chia-Hsiang Kao , Cong Phuoc Huynh , Chien-Yi Wang , Noranart Vesdapunt , Stefan Stojanov , Bharath Hariharan , Oleksandr Obiednikov , Ning Zhou

Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Xuan Ju , Tianyu Wang , Yuqian Zhou , He Zhang , Qing Liu , Nanxuan Zhao , Zhifei Zhang , Yijun Li , Yuanhao Cai , Shaoteng Liu , Daniil Pakhomov , Zhe Lin , Soo Ye Kim , Qiang Xu

True understanding of videos comes from a joint analysis of all its modalities: the video frames, the audio track, and any accompanying text such as closed captions. We present a way to learn a compact multimodal feature representation that…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Vivek Sharma , Makarand Tapaswi , Rainer Stiefelhagen

World models play a crucial role in understanding and predicting the dynamics of the world, which is essential for video generation. However, existing world models are confined to specific scenarios such as gaming or driving, limiting their…

计算机视觉与模式识别 · 计算机科学 2024-01-19 Xiaofeng Wang , Zheng Zhu , Guan Huang , Boyuan Wang , Xinze Chen , Jiwen Lu

Video representation is an important and challenging task in the computer vision community. In this paper, we assume that image frames of a moving scene can be modeled as a Linear Dynamical System. We propose a sparse coding framework,…

计算机视觉与模式识别 · 计算机科学 2013-12-20 Xian Wei , Hao Shen , Martin Kleinsteuber

We introduce a novel framework for evaluating multimodal deep learning models with respect to their language understanding and generalization abilities. In this approach, artificial data is automatically generated according to the…

计算与语言 · 计算机科学 2017-04-18 Alexander Kuhnle , Ann Copestake

Despite recent progress, video diffusion models still struggle to synthesize realistic videos involving highly dynamic motions or requiring fine-grained motion controllability. A central limitation lies in the scarcity of such examples in…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Wonjoon Jin , Jiyun Won , Janghyeok Han , Qi Dai , Chong Luo , Seung-Hwan Baek , Sunghyun Cho

This paper tackles the challenge of robust reconstruction, i.e., the task of reconstructing a 3D scene from a set of inconsistent multi-view images. Some recent works have attempted to simultaneously remove image inconsistencies and perform…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Jin Cao , Hongrui Wu , Ziyong Feng , Hujun Bao , Xiaowei Zhou , Sida Peng

In this paper, we introduce \textbf{DimensionX}, a framework designed to generate photorealistic 3D and 4D scenes from just a single image with video diffusion. Our approach begins with the insight that both the spatial structure of a 3D…

计算机视觉与模式识别 · 计算机科学 2024-11-08 Wenqiang Sun , Shuo Chen , Fangfu Liu , Zilong Chen , Yueqi Duan , Jun Zhang , Yikai Wang

Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks rely on recognition-style protocols such as Visual Question Answering (VQA) and Violation of…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Jiarong Liang , Max Ku , Ka-Hei Hui , Ping Nie , Wenhu Chen

Despite rapid progress in multimodal large language models (MLLMs) and emerging omni-modal architectures, current benchmarks remain limited in scope and integration, suffering from incomplete modality coverage, restricted interaction to…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Yue Jiang , Dingkang Yang , Minghao Han , Jinghang Han , Zizhi Chen , Yizhou Liu , Mingcheng Li , Peng Zhai , Lihua Zhang

Generating flexible-view 3D scenes, including 360{\deg} rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework consisting of two key components: (1) a strong…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Luxi Chen , Zihan Zhou , Min Zhao , Yikai Wang , Ge Zhang , Wenhao Huang , Hao Sun , Ji-Rong Wen , Chongxuan Li

Understanding and reconstructing the complex geometry and motion of dynamic scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward model designed to efficiently…

Synthesizing high-fidelity videos from real-world multi-view input is challenging because of the complexities of real-world environments and highly dynamic motions. Previous works based on neural radiance fields have demonstrated…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Feng Wang , Sinan Tan , Xinghang Li , Zeyue Tian , Yafei Song , Huaping Liu

Physics-aware driving world model is essential for drive planning, out-of-distribution data synthesis, and closed-loop evaluation. However, existing methods often rely on a single diffusion model to directly map driving actions to videos,…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Zhenya Yang , Zhe Liu , Yuxiang Lu , Liping Hou , Chenxuan Miao , Siyi Peng , Bailan Feng , Xiang Bai , Hengshuang Zhao

Monocular dynamic reconstruction is a challenging and long-standing vision problem due to the highly ill-posed nature of the task. Existing approaches depend on templates, are effective only in quasi-static scenes, or fail to model 3D…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Qianqian Wang , Vickie Ye , Hang Gao , Weijia Zeng , Jake Austin , Zhengqi Li , Angjoo Kanazawa

Modern Large Multimodal Models (LMMs) have demonstrated extraordinary ability in static image and single-state spatial-temporal understanding. However, their capacity to comprehend the dynamic changes of objects within a shared spatial…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Kewei Wei , Bocheng Hu , Jie Cao , Xiaohan Chen , Zhengxi Lu , Wubing Xia , Weili Xu , Jiaao Wu , Junchen He , Mingyu Jia , Ciyun Zhao , Ye Sun , Yizhi Li , Zhonghan Zhao , Jian Zhang , Gaoang Wang