English
Related papers

Related papers: Real-Time ESFP: Estimating, Smoothing, Filtering, …

200 papers

This paper presents a novel tightly-coupled monocular visual-inertial Simultaneous Localization and Mapping algorithm, which provides accurate and robust localization within the globally consistent map in real time on a standard CPU. This…

Robotics · Computer Science 2021-02-24 Meixiang Quan , Songhao Piao , Minglang Tan , Shi-Sheng Huang

Endoscopy is essential in medical imaging, used for diagnosis, prognosis and treatment. Developing a robust dynamic 3D reconstruction pipeline for endoscopic videos could enhance visualization, improve diagnostic accuracy, aid in treatment…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Laura Salort-Benejam , Antonio Agudo

Ultra-high-definition (UHD) image restoration is vital for applications demanding exceptional visual fidelity, yet existing methods often face a trade-off between restoration quality and efficiency, limiting their practical deployment. In…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Xin Su , Chen Wu , Zhuoran Zheng

In existing restoration-oriented Video Frame Interpolation (VFI) approaches, the motion estimation between neighboring frames plays a crucial role. However, the estimation accuracy in existing methods remains a challenge, primarily due to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Yan Han , Xiaogang Xu , Yingqi Lin , Jiafei Wu , Zhe Liu , Ming-Hsuan Yang

Video endoscopy represents a major advance in the investigation of gastrointestinal diseases. Reviewing endoscopy videos often involves frequent adjustments and reorientations to piece together a complete view, which can be both…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Juming Xiong , Muyang Li , Ruining Deng , Tianyuan Yao , Shunxing Bao , Regina N Tyree , Girish Hiremath , Yuankai Huo

While soft robot manipulators offer compelling advantages over rigid counterparts, including inherent compliance, safe human-robot interaction, and the ability to conform to complex geometries, accurate forward modeling from low-dimensional…

Robotics · Computer Science 2026-03-23 Ziyong Ma , Uksang Yoo , Jonathan Francis , Weiming Zhi , Jeffrey Ichnowski , Jean Oh

We present a novel end-to-end framework for facial performance capture given a monocular video of an actor's face. Our framework are comprised of 2 parts. First, to extract the information in the frames, we optimize a triplet loss to learn…

Graphics · Computer Science 2018-12-27 Hsien-Yu Meng , Tzu-heng Lin , Xiubao Jiang , Yao Lu , Jiangtao Wen

Estimating a scene's depth to achieve collision avoidance against moving pedestrians is a crucial and fundamental problem in the robotic field. This paper proposes a novel, low complexity network architecture for fast and accurate human…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Shan An , Fangru Zhou , Mei Yang , Haogang Zhu , Changhong Fu , Konstantinos A. Tsintotas

We present a real-time on-device hand tracking pipeline that predicts hand skeleton from single RGB camera for AR/VR applications. The pipeline consists of two models: 1) a palm detector, 2) a hand landmark model. It's implemented via…

Computer Vision and Pattern Recognition · Computer Science 2020-06-19 Fan Zhang , Valentin Bazarevsky , Andrey Vakunov , Andrei Tkachenka , George Sung , Chuo-Ling Chang , Matthias Grundmann

Recent advances in the masked autoencoder (MAE) paradigm have significantly propelled self-supervised skeleton-based action recognition. However, most existing approaches limit reconstruction targets to raw joint coordinates or their simple…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Shengkai Sun , Zefan Zhang , Jianfeng Dong , Zhiyong Cheng , Xiaojun Chang , Meng Wang

Although many video prediction methods have obtained good performance in low-resolution (64$\sim$128) videos, predictive models for high-resolution (512$\sim$4K) videos have not been fully explored yet, which are more meaningful due to the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Zheng Chang , Xinfeng Zhang , Shanshe Wang , Siwei Ma , Wen Gao

Robots operating in unstructured environments require a comprehensive understanding of their surroundings, necessitating geometric and semantic information from sensor data. Traditional RGB-D processing pipelines focus primarily on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Zhiwu Zheng , Lauren Mentzer , Berk Iskender , Michael Price , Colm Prendergast , Audren Cloitre

In order to safely and efficiently collaborate with humans, industrial robots need the ability to alter their motions quickly to react to sudden changes in the environment, such as an obstacle appearing across a planned trajectory. In…

Robotics · Computer Science 2022-07-19 Shohei Fujii , Quang-Cuong Pham

Real-time monocular 3D reconstruction is a challenging problem that remains unsolved. Although recent end-to-end methods have demonstrated promising results, tiny structures and geometric boundaries are hardly captured due to their…

Computer Vision and Pattern Recognition · Computer Science 2023-07-26 Chenyangguang Zhang , Zhiqiang Lou , Yan Di , Federico Tombari , Xiangyang Ji

We propose a novel geometric and photometric 3D mapping pipeline for accurate and real-time scene reconstruction from monocular images. To achieve this, we leverage recent advances in dense monocular SLAM and real-time hierarchical…

Computer Vision and Pattern Recognition · Computer Science 2022-10-26 Antoni Rosinol , John J. Leonard , Luca Carlone

The framework of dominant learned video compression methods is usually composed of motion prediction modules as well as motion vector and residual image compression modules, suffering from its complex structure and error propagation…

Image and Video Processing · Electrical Eng. & Systems 2021-04-14 Zhenhong Sun , Zhiyu Tan , Xiuyu Sun , Fangyi Zhang , Dongyang Li , Yichen Qian , Hao Li

There has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. While these tasks rely heavily on known camera poses, the problem of finding such poses using structure-from-motion…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Lily Goli , Sara Sabour , Mark Matthews , Marcus Brubaker , Dmitry Lagun , Alec Jacobson , David J. Fleet , Saurabh Saxena , Andrea Tagliasacchi

We describe our preliminary design of a real-time asynchronous event-based monocular odometry for planetary exploration. Operating under strict computational constraints, planetary rovers frequently encounter complex, unpredictable…

Robotics · Computer Science 2026-05-28 Benat Inigo , Florian Steidle , Wolfgang Stuerzl

Consecutive frames in a video contain redundancy, but they may also contain relevant complementary information for the detection task. The objective of our work is to leverage this complementary information to improve detection. Therefore,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-19 Noreen Anwar , Guillaume-Alexandre Bilodeau , Wassim Bouachir

We present an approach to estimating camera rotation in crowded, real-world scenes from handheld monocular video. While camera rotation estimation is a well-studied problem, no previous methods exhibit both high accuracy and acceptable…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Fabien Delattre , David Dirnfeld , Phat Nguyen , Stephen Scarano , Michael J. Jones , Pedro Miraldo , Erik Learned-Miller