中文
相关论文

相关论文: Uni3C: Unifying Precisely 3D-Enhanced Camera and H…

200 篇论文

Camera control has been extensively studied in conditioned video generation; however, performing precisely altering the camera trajectories while faithfully preserving the video content remains a challenging task. The mainstream approach to…

计算机视觉与模式识别 · 计算机科学 2026-02-02 Dong-Yu Chen , Yixin Guo , Shuojin Yang , Tai-Jiang Mu , Shi-Min Hu

Hand motion plays a central role in human interaction, yet modeling realistic 4D hand motion (i.e., 3D hand pose sequences over time) remains challenging. Research in this area is typically divided into two tasks: (1) Estimation approaches…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Zhihao Sun , Tong Wu , Ruirui Tu , Daoguo Dong , Zuxuan Wu

Current 3D object detection models follow a single dataset-specific training and testing paradigm, which often faces a serious detection accuracy drop when they are directly deployed in another dataset. In this paper, we study the task of…

计算机视觉与模式识别 · 计算机科学 2023-05-01 Bo Zhang , Jiakang Yuan , Botian Shi , Tao Chen , Yikang Li , Yu Qiao

Human motion generation is a significant pursuit in generative computer vision with widespread applications in film-making, video games, AR/VR, and human-robot interaction. Current methods mainly utilize either diffusion-based generative…

计算机视觉与模式识别 · 计算机科学 2025-02-03 Canxuan Gang

We present I2V3D, a novel framework for animating static images into dynamic videos with precise 3D control, leveraging the strengths of both 3D geometry guidance and advanced generative models. Our approach combines the precision of a…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Zhiyuan Zhang , Dongdong Chen , Jing Liao

Achieving machine autonomy and human control often represent divergent objectives in the design of interactive AI systems. Visual generative foundation models such as Stable Diffusion show promise in navigating these goals, especially when…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Can Qin , Shu Zhang , Ning Yu , Yihao Feng , Xinyi Yang , Yingbo Zhou , Huan Wang , Juan Carlos Niebles , Caiming Xiong , Silvio Savarese , Stefano Ermon , Yun Fu , Ran Xu

The field of physics-based animation is gaining importance due to the increasing demand for realism in video games and films, and has recently seen wide adoption of data-driven techniques, such as deep reinforcement learning (RL), which…

图形学 · 计算机科学 2020-12-01 Tingwu Wang , Yunrong Guo , Maria Shugrina , Sanja Fidler

Transformers have emerged as a universal backbone across 3D perception, video generation, and world models for autonomous driving and embodied AI, where understanding camera geometry is essential for grounding visual observations in…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Cheng Zhang , Boying Li , Meng Wei , Yan-Pei Cao , Camilo Cruz Gambardella , Dinh Phung , Jianfei Cai

The scale diversity of point cloud data presents significant challenges in developing unified representation learning techniques for 3D vision. Currently, there are few unified 3D models, and no existing pre-training method is equally…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Ziyi Wang , Yanran Zhang , Jie Zhou , Jiwen Lu

Recent advancements in personalized Text-to-Video (T2V) generation have made significant strides in synthesizing character-specific content. However, these methods face a critical limitation: the inability to perform fine-grained control…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Haopeng Fang , Di Qiu , Binjie Mao , He Tang

Recent advances in generative video models have led to significant breakthroughs in high-fidelity video synthesis, specifically in controllable video generation where the generated video is conditioned on text and action inputs, e.g., in…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Zhiting Mei , Tenny Yin , Micah Baker , Ola Shorinwa , Anirudha Majumdar

We introduce UPose3D, a novel approach for multi-view 3D human pose estimation, addressing challenges in accuracy and scalability. Our method advances existing pose estimation frameworks by improving robustness and flexibility without…

计算机视觉与模式识别 · 计算机科学 2024-07-11 Vandad Davoodnia , Saeed Ghorbani , Marc-André Carbonneau , Alexandre Messier , Ali Etemad

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Jianhui Wei , Xiaotian Zhang , Yichen Li , Yuan Wang , Yan Zhang , Ziyi Chen , Zhihang Tang , Wei Xu , Zuozhu Liu

Despite the remarkable progress of 3D generation, achieving controllability, i.e., ensuring consistency between generated 3D content and input conditions like edge and depth, remains a significant challenge. Existing methods often struggle…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Hongbin Xu , Chaohui Yu , Feng Xiao , Jiazheng Xing , Hai Ci , Weitao Chen , Fan Wang , Ming Li

Reconstructing 3D scenes from monocular surgical videos can enhance surgeon's perception and therefore plays a vital role in various computer-assisted surgery tasks. However, achieving scale-consistent reconstruction remains an open…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Jiaxin Guo , Wenzhen Dong , Tianyu Huang , Hao Ding , Ziyi Wang , Haomin Kuang , Qi Dou , Yun-Hui Liu

Monocular 3D estimation is crucial for visual perception. However, current methods fall short by relying on oversimplified assumptions, such as pinhole camera models or rectified images. These limitations severely restrict their general…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Luigi Piccinelli , Christos Sakaridis , Mattia Segu , Yung-Hsu Yang , Siyuan Li , Wim Abbeloos , Luc Van Gool

Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geometric consistency. However, existing methods typically rely on fragmented geometric…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Hong Jiang , Wensong Song , Zongxing Yang , Ruijie Quan , Yi Yang

Loco-Manipulation for humanoid robots aims to enable robots to integrate mobility with upper-body tracking capabilities. Most existing approaches adopt hierarchical architectures that decompose control into isolated upper-body…

机器人学 · 计算机科学 2026-03-03 Wandong Sun , Luying Feng , Baoshi Cao , Yang Liu , Yaochu Jin , Zongwu Xie

Traditional and neural video codecs commonly encounter limitations in controllability and generality under ultra-low-bitrate coding scenarios. To overcome these challenges, we propose M3-CVC, a controllable video compression framework…

图像与视频处理 · 电气工程与系统科学 2024-12-30 Rui Wan , Qi Zheng , Yibo Fan

The need for automated real-time visual systems in applications such as smart camera surveillance, smart environments, and drones necessitates the improvement of methods for visual active monitoring and control. Traditionally, the active…

计算机视觉与模式识别 · 计算机科学 2021-07-29 Christos Kyrkou