English
Related papers

Related papers: 3DSPA: A 3D Semantic Point Autoencoder for Evaluat…

200 papers

In this paper, we introduce SPA, a novel representation learning framework that emphasizes the importance of 3D spatial awareness in embodied AI. Our approach leverages differentiable neural rendering on multi-view images to endow a vanilla…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Haoyi Zhu , Honghui Yang , Yating Wang , Jiange Yang , Limin Wang , Tong He

Current 3D human animation methods struggle to achieve photorealism: kinematics-based approaches lack non-rigid dynamics (e.g., clothing dynamics), while methods that leverage video diffusion priors can synthesize non-rigid motion but…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Qi Sun , Can Wang , Jiaxiang Shang , Yingchun Liu , Jing Liao

Reliable autonomous driving relies on large-scale, well-labeled data and robust models. However, manual data collection is resource-intensive, and traditional simulation suffers from a persistent reality gap. While recent generative…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Kaicong Huang , Talha Azfar , Weisong Shi , Ruimin Ke

The rapid advancement of generative AI enables highly realistic synthetic videos, posing significant challenges for content authentication and raising urgent concerns about misuse. Existing detection methods often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Christian Internò , Robert Geirhos , Markus Olhofer , Sunny Liu , Barbara Hammer , David Klindt

A key challenge of learning a visual representation for the 3D high fidelity geometry of dressed humans lies in the limited availability of the ground truth data (e.g., 3D scanned models), which results in the performance degradation of 3D…

Computer Vision and Pattern Recognition · Computer Science 2022-12-29 Yasamin Jafarian , Hyun Soo Park

Leading methods in the domain of action recognition try to distill information from both the spatial and temporal dimensions of an input video. Methods that reach State of the Art (SotA) accuracy, usually make use of 3D convolution layers…

Computer Vision and Pattern Recognition · Computer Science 2021-05-28 Gilad Sharir , Asaf Noy , Lihi Zelnik-Manor

Synthetic data has been a critical tool for training scene text detection and recognition models. On the one hand, synthetic word images have proven to be a successful substitute for real images in training scene text recognizers. On the…

Computer Vision and Pattern Recognition · Computer Science 2020-08-19 Shangbang Long , Cong Yao

Explorable 3D world generation from a single image or text prompt forms a cornerstone of spatial intelligence. Recent works utilize video model to achieve wide-scope and generalizable 3D world generation. However, existing approaches often…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Zhongqi Yang , Wenhang Ge , Yuqi Li , Jiaqi Chen , Haoyuan Li , Mengyin An , Fei Kang , Hua Xue , Baixin Xu , Yuyang Yin , Eric Li , Yang Liu , Yikai Wang , Hao-Xiang Guo , Yahui Zhou

This paper proposes a novel pretext task to address the self-supervised video representation learning problem. Specifically, given an unlabeled video clip, we compute a series of spatio-temporal statistical summaries, such as the spatial…

Computer Vision and Pattern Recognition · Computer Science 2021-02-01 Jiangliu Wang , Jianbo Jiao , Linchao Bao , Shengfeng He , Wei Liu , Yun-hui Liu

The introduction of neural radiance fields has greatly improved the effectiveness of view synthesis for monocular videos. However, existing algorithms face difficulties when dealing with uncontrolled or lengthy scenarios, and require…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Kaichen Zhou , Jia-Xing Zhong , Sangyun Shin , Kai Lu , Yiyuan Yang , Andrew Markham , Niki Trigoni

It has long been challenging to recover the underlying dynamic 3D scene representations from a monocular RGB video. Existing works formulate this problem into finding a single most plausible solution by adding various constraints such as…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Ziyang Song , Jinxi Li , Bo Yang

DUSt3R has recently shown that one can reduce many tasks in multi-view geometry, including estimating camera intrinsics and extrinsics, reconstructing the scene in 3D, and establishing image correspondences, to the prediction of a pair of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Edgar Sucar , Zihang Lai , Eldar Insafutdinov , Andrea Vedaldi

Deep-learning based face-swap videos, also known as deep fakes, are becoming more and more realistic and deceiving. The malicious usage of these face-swap videos has caused wide concerns. The research community has been focusing on the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Xianyun Sun , Beibei Dong , Caiyong Wang , Bo Peng , Jing Dong

Recent advances in interactive 3D segmentation from 2D images have demonstrated impressive performance. However, current models typically require extensive scene-specific training to accurately reconstruct and segment objects, which limits…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Yansong Guo , Jie Hu , Yansong Qu , Liujuan Cao

Visual understanding of the world goes beyond the semantics and flat structure of individual images. In this work, we aim to capture both the 3D structure and dynamics of real-world scenes from monocular real-world videos. Our Dynamic Scene…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Maximilian Seitzer , Sjoerd van Steenkiste , Thomas Kipf , Klaus Greff , Mehdi S. M. Sajjadi

Realistic simulation is critical for applications ranging from robotics to animation. Learned simulators have emerged as a possibility to capture real world physics directly from video data, but very often require privileged information…

Graphics · Computer Science 2025-08-12 Mikel Zhobro , Andreas René Geist , Georg Martius

Hand gestures are a natural means of interaction in Augmented Reality and Virtual Reality (AR/VR) applications. Recently, there has been an increased focus on removing the dependence of accurate hand gesture recognition on complex sensor…

Computer Vision and Pattern Recognition · Computer Science 2019-12-09 Varun Jain , Shivam Aggarwal , Suril Mehta , Ramya Hebbalaguppe

Emerging world models autoregressively generate video frames in response to actions, such as camera movements and text prompts, among other control signals. Due to limited temporal context window sizes, these models often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Tong Wu , Shuai Yang , Ryan Po , Yinghao Xu , Ziwei Liu , Dahua Lin , Gordon Wetzstein

Video summarization helps turn long videos into clear, concise representations that are easier to review, document, and analyze, especially in high-stakes domains like surgical training. Prior work has progressed from using basic visual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Shreya Rajpal , Michal Golovanevsky , Carsten Eickhoff

Aerial vehicles are revolutionizing the way film-makers can capture shots of actors by composing novel aerial and dynamic viewpoints. However, despite great advancements in autonomous flight technology, generating expressive camera…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Rogerio Bonatti , Arthur Bucker , Sebastian Scherer , Mustafa Mukadam , Jessica Hodgins