中文
相关论文

相关论文: 3D-Aware Implicit Motion Control for View-Adaptive…

200 篇论文

Current video models fail as world model as they lack fine-graiend control. General-purpose household robots require real-time fine motor control to handle delicate tasks and urgent situations. In this work, we introduce fine-grained…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Yichen Li , Antonio Torralba

Existing video generation models struggle to follow complex text prompts and synthesize multiple objects, raising the need for additional grounding input for improved controllability. In this work, we propose to decompose videos into visual…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Weixi Feng , Chao Liu , Sifei Liu , William Yang Wang , Arash Vahdat , Weili Nie

In recent years, Generative Adversarial Networks have achieved impressive results in photorealistic image synthesis. This progress nurtures hopes that one day the classical rendering pipeline can be replaced by efficient models that are…

计算机视觉与模式识别 · 计算机科学 2020-03-25 Yiyi Liao , Katja Schwarz , Lars Mescheder , Andreas Geiger

Action recognition is a relatively established task, where givenan input sequence of human motion, the goal is to predict its ac-tion category. This paper, on the other hand, considers a relativelynew problem, which could be thought of as…

计算机视觉与模式识别 · 计算机科学 2020-07-31 Chuan Guo , Xinxin Zuo , Sen Wang , Shihao Zou , Qingyao Sun , Annan Deng , Minglun Gong , Li Cheng

Recent advancements in diffusion models have significantly enhanced the data synthesis with 2D control. Yet, precise 3D control in street view generation, crucial for 3D perception tasks, remains elusive. Specifically, utilizing Bird's-Eye…

计算机视觉与模式识别 · 计算机科学 2024-05-06 Ruiyuan Gao , Kai Chen , Enze Xie , Lanqing Hong , Zhenguo Li , Dit-Yan Yeung , Qiang Xu

Recent advances in robot manipulation have leveraged pre-trained vision-language models (VLMs) and explored integrating 3D spatial signals into these models for effective action prediction, giving rise to the promising…

机器人学 · 计算机科学 2026-01-14 Zhenyang Liu , Yongchong Gu , Yikai Wang , Xiangyang Xue , Yanwei Fu

Controlling the movements of dynamic objects and the camera within generated videos is a meaningful yet challenging task. Due to the lack of datasets with comprehensive 6D pose annotations, existing text-to-video methods can not…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Xincheng Shuai , Henghui Ding , Zhenyuan Qin , Hao Luo , Xingjun Ma , Dacheng Tao

Computer vision is largely based on 2D techniques, with 3D vision still relegated to a relatively narrow subset of applications. However, by building on recent advances in 3D models such as neural radiance fields, some authors have shown…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Vadim Tschernezki , Diane Larlus , Iro Laina , Andrea Vedaldi

Recent Multimodal Large Language Models (MLLMs) have shown high potential for spatial reasoning within 3D scenes. However, they typically rely on computationally expensive 3D representations like point clouds or reconstructed Bird's-Eye…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Shuyao Shi , Kang G. Shin

Generative models have emerged as an essential building block for many image synthesis and editing tasks. Recent advances in this field have also enabled high-quality 3D or video content to be generated that exhibits either multi-view or…

计算机视觉与模式识别 · 计算机科学 2023-08-10 Sherwin Bahmani , Jeong Joon Park , Despoina Paschalidou , Hao Tang , Gordon Wetzstein , Leonidas Guibas , Luc Van Gool , Radu Timofte

Estimation of 3D human pose from monocular image has gained considerable attention, as a key step to several human-centric applications. However, generalizability of human pose estimation models developed using supervision on large-scale…

计算机视觉与模式识别 · 计算机科学 2020-06-26 Jogendra Nath Kundu , Siddharth Seth , Rahul M , Mugalodi Rakesh , R. Venkatesh Babu , Anirban Chakraborty

Physics-aware driving world model is essential for drive planning, out-of-distribution data synthesis, and closed-loop evaluation. However, existing methods often rely on a single diffusion model to directly map driving actions to videos,…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Zhenya Yang , Zhe Liu , Yuxiang Lu , Liping Hou , Chenxuan Miao , Siyi Peng , Bailan Feng , Xiang Bai , Hengshuang Zhao

Recent advances in text-to-video (T2V) and image-to-video (I2V) models, have enabled the creation of visually compelling and dynamic videos from simple textual descriptions or initial frames. However, these models often fail to provide an…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Thomas Ressler-Antal , Frank Fundel , Malek Ben Alaya , Stefan Andreas Baumann , Felix Krause , Ming Gui , Björn Ommer

Recent strides in Text-to-3D techniques have been propelled by distilling knowledge from powerful large text-to-image diffusion models (LDMs). Nonetheless, existing Text-to-3D approaches often grapple with challenges such as…

计算机视觉与模式识别 · 计算机科学 2023-08-23 Yiwen Chen , Chi Zhang , Xiaofeng Yang , Zhongang Cai , Gang Yu , Lei Yang , Guosheng Lin

Controllable generation is considered a potentially vital approach to address the challenge of annotating 3D data, and the precision of such controllable generation becomes particularly imperative in the context of data production for…

计算机视觉与模式识别 · 计算机科学 2025-07-16 Jinhua Zhang , Hualian Sheng , Sijia Cai , Bing Deng , Qiao Liang , Wen Li , Ying Fu , Jieping Ye , Shuhang Gu

Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse driving videos, while preserving the spatiotemporal…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Li-Heng Chen , Ke Cheng , Yahui Liu , Lei Shi , Shi-Sheng Huang , Hongbo Fu

Recent advancements in trajectory-guided video generation have achieved notable progress. However, existing models still face challenges in generating object motions with potentially changing 6D poses under wide-range rotations, due to…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Longbin Ji , Lei Zhong , Pengfei Wei , Changjian Li

Recent studies have shown remarkable advances in 3D human pose estimation from monocular images, with the help of large-scale in-door 3D datasets and sophisticated network architectures. However, the generalizability to different…

计算机视觉与模式识别 · 计算机科学 2019-03-28 Xipeng Chen , Kwan-Yee Lin , Wentao Liu , Chen Qian , Xiaogang Wang , Liang Lin

Dynamic human rendering from video sequences has achieved remarkable progress by formulating the rendering as a mapping from static poses to human images. However, existing methods focus on the human appearance reconstruction of every…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Tao Hu , Fangzhou Hong , Ziwei Liu

Autonomous driving is an exciting new industry, posing important research questions. Within the perception module, 3D human pose estimation is an emerging technology, which can enable the autonomous vehicle to perceive and understand the…

计算机视觉与模式识别 · 计算机科学 2022-12-16 Andrei Zanfir , Mihai Zanfir , Alexander Gorban , Jingwei Ji , Yin Zhou , Dragomir Anguelov , Cristian Sminchisescu