English
Related papers

Related papers: NVDS+: Towards Efficient and Versatile Neural Stab…

200 papers

Monocular depth estimation is fundamental for 3D scene understanding and downstream applications. However, even under the supervised setup, it is still challenging and ill-posed due to the lack of full geometric constraints. Although a…

Computer Vision and Pattern Recognition · Computer Science 2023-04-14 Luigi Piccinelli , Christos Sakaridis , Fisher Yu

We propose a smooth regularization technique that instills a strong temporal inductive bias in video recognition models, particularly benefiting lightweight architectures. Our method encourages smoothness in the intermediate-layer…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Gil Goldman , Raja Giryes , Mahadev Satyanarayanan

We introduce a novel geometry-guided online video view synthesis method with enhanced view and temporal consistency. Traditional approaches achieve high-quality synthesis from dense multi-view camera setups but require significant…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Hyunho Ha , Lei Xiao , Christian Richardt , Thu Nguyen-Phuoc , Changil Kim , Min H. Kim , Douglas Lanman , Numair Khan

Monocular depth estimation (MDE) plays a pivotal role in various computer vision applications, such as robotics, augmented reality, and autonomous driving. Despite recent advancements, existing methods often fail to meet key requirements…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Andrii Litvynchuk , Ivan Livinsky , Anand Ravi , Nima Kalantari , Andrii Tsarov

Dynamic novel view synthesis (NVS) is essential for creating immersive experiences. Existing approaches have advanced dynamic NVS by introducing 3D Gaussian Splatting (3DGS) with implicit deformation fields or indiscriminately assigned…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Kaizhe Zhang , Yijie Zhou , Weizhan Zhang , Caixia Yan , Haipeng Du , yugui xie , Yu-Hui Wen , Yong-Jin Liu

Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Recent works address this problem by applying a video…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Jiahao Lu , Tianyu Huang , Peng Li , Zhiyang Dou , Cheng Lin , Zhiming Cui , Zhen Dong , Sai-Kit Yeung , Wenping Wang , Yuan Liu

Human pose is a useful feature for fine-grained sports action understanding. However, pose estimators are often unreliable when run on sports video due to domain shift and factors such as motion blur and occlusions. This leads to poor…

Computer Vision and Pattern Recognition · Computer Science 2021-09-06 James Hong , Matthew Fisher , Michaël Gharbi , Kayvon Fatahalian

There are increasing interests of studying the video-to-depth (V2D) problem with machine learning techniques. While earlier methods directly learn a mapping from images to depth maps and camera poses, more recent works enforce multi-view…

Computer Vision and Pattern Recognition · Computer Science 2023-03-08 Xiaodong Gu , Weihao Yuan , Zuozhuo Dai , Siyu Zhu , Chengzhou Tang , Zilong Dong , Ping Tan

We propose a physically-motivated deep learning framework to solve a general version of the challenging indoor lighting estimation problem. Given a single LDR image with a depth map, our method predicts spatially consistent lighting at any…

Computer Vision and Pattern Recognition · Computer Science 2023-05-09 Zhengqin Li , Li Yu , Mikhail Okunev , Manmohan Chandraker , Zhao Dong

3D Gaussian Splatting (3DGS) has emerged as a powerful representation due to its efficiency and high-fidelity rendering. 3DGS training requires a known camera pose for each input view, typically obtained by Structure-from-Motion (SfM)…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Zhen-Hui Dong , Sheng Ye , Yu-Hui Wen , Nannan Li , Yong-Jin Liu

We address the problem of highlight detection from a 360 degree video by summarizing it both spatially and temporally. Given a long 360 degree video, we spatially select pleasantly-looking normal field-of-view (NFOV) segments from unlimited…

Computer Vision and Pattern Recognition · Computer Science 2018-02-01 Youngjae Yu , Sangho Lee , Joonil Na , Jaeyun Kang , Gunhee Kim

Encoding video content into compact latent tokens has become a fundamental step in video generation and understanding, driven by the need to address the inherent redundancy in pixel-level representations. Consequently, there is a growing…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Anni Tang , Tianyu He , Junliang Guo , Xinle Cheng , Li Song , Jiang Bian

Applying image processing algorithms independently to each video frame often leads to temporal inconsistency in the resulting video. To address this issue, we present a novel and general approach for blind video temporal consistency. Our…

Computer Vision and Pattern Recognition · Computer Science 2020-10-23 Chenyang Lei , Yazhou Xing , Qifeng Chen

Human image animation involves generating a video from a static image by following a specified pose sequence. Current approaches typically adopt a multi-stage pipeline that separately learns appearance and motion, which often leads to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Qilin Wang , Zhengkai Jiang , Chengming Xu , Jiangning Zhang , Yabiao Wang , Xinyi Zhang , Yun Cao , Weijian Cao , Chengjie Wang , Yanwei Fu

Dynamic stereo matching is the task of estimating consistent disparities from stereo videos with dynamic objects. Recent learning-based methods prioritize optimal performance on a single stereo pair, resulting in temporal inconsistencies.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Junpeng Jing , Ye Mao , Krystian Mikolajczyk

Video Instance Segmentation (VIS) fundamentally struggles with pervasive challenges including object occlusions, motion blur, and appearance variations during temporal association. To overcome these limitations, this work introduces…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Quanzhu Niu , Yikang Zhou , Shihao Chen , Tao Zhang , Shunping Ji

Video watermarking embeds a message into a cover video in an imperceptible manner, which can be retrieved even if the video undergoes certain modifications or distortions. Traditional watermarking methods are often manually designed for…

Multimedia · Computer Science 2021-04-27 Xiyang Luo , Yinxiao Li , Huiwen Chang , Ce Liu , Peyman Milanfar , Feng Yang

Deploying depth estimation networks in the real world requires high-level robustness against various adverse conditions to ensure safe and reliable autonomy. For this purpose, many autonomous vehicles employ multi-modal sensor systems,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Ukcheol Shin , Kyunghyun Lee , Jean Oh

The explosive growth of video data in recent years has brought higher demands for video analytics, where accuracy and efficiency remain the two primary concerns. Deep neural networks (DNNs) have been widely adopted to ensure accuracy;…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Shanjiang Tang , Rui Huang , Hsinyu Luo , Chunjiang Wang , Ce Yu , Yusen Li , Hao Fu , Chao Sun , and Jian Xiao

Novel View Synthesis (NVS) aims to generate unseen views of a 3D object given a limited number of known views. Existing methods often struggle to synthesize plausible views for unobserved regions, particularly under single-view input, and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Jinglin Liang , Zijian Zhou , Rui Huang , Shuangping Huang , Yichen Gong