English
Related papers

Related papers: A Deeper Look into DeepCap

200 papers

Combining sparse IMUs and a monocular camera is a new promising setting to perform real-time human motion capture. This paper proposes a diffusion-based solution to learn human motion priors and fuse the two modalities of signals together…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Shaohua Pan , Xinyu Yi , Yan Zhou , Weihua Jian , Yuan Zhang , Pengfei Wan , Feng Xu

Self-supervised monocular depth estimation has emerged as a promising approach since it does not rely on labeled training data. Most methods combine convolution and Transformer to model long-distance dependencies to estimate depth…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Xuezhi Xiang , Yao Wang , Lei Zhang , Denis Ombati , Himaloy Himu , Xiantong Zhen

This paper introduces a novel deep learning based approach for vision based single target tracking. We address this problem by proposing a network architecture which takes the input video frames and directly computes the tracking score for…

Computer Vision and Pattern Recognition · Computer Science 2016-07-12 Mengyao Zhai , Mehrsan Javan Roshtkhari , Greg Mori

We propose a CNN-based approach for 3D human body pose estimation from single RGB images that addresses the issue of limited generalizability of models trained solely on the starkly limited publicly available 3D pose data. Using only the…

Computer Vision and Pattern Recognition · Computer Science 2017-10-05 Dushyant Mehta , Helge Rhodin , Dan Casas , Pascal Fua , Oleksandr Sotnychenko , Weipeng Xu , Christian Theobalt

A long-standing challenge in scene analysis is the recovery of scene arrangements under moderate to heavy occlusion, directly from monocular video. While the problem remains a subject of active research, concurrent advances have been made…

Graphics · Computer Science 2019-07-19 Aron Monszpart , Paul Guerrero , Duygu Ceylan , Ersin Yumer , Niloy J. Mitra

We address the problem of 3D human pose estimation from 2D input images using only weakly supervised training data. Despite showing considerable success for 2D pose estimation, the application of supervised machine learning to 3D pose…

Computer Vision and Pattern Recognition · Computer Science 2018-07-31 Matteo Ruggero Ronchi , Oisin Mac Aodha , Robert Eng , Pietro Perona

We propose to leverage recent advances in reliable 2D pose estimation with Convolutional Neural Networks (CNN) to estimate the 3D pose of people from depth images in multi-person Human-Robot Interaction (HRI) scenarios. Our method is based…

Computer Vision and Pattern Recognition · Computer Science 2020-11-11 Angel Martínez-González , Michael Villamizar , Olivier Canévet , Jean-Marc Odobez

In this paper we propose a method based on deep learning that detects multiple people from a single overhead depth image with high reliability. Our neural network, called DPDnet, is based on two fully-convolutional encoder-decoder neural…

Computer Vision and Pattern Recognition · Computer Science 2020-06-02 David Fuentes-Jimenez , Roberto Martin-Lopez , Cristina Losada-Gutierrez , David Casillas-Perez , Javier Macias-Guarasa , Daniel Pizarro , Carlos A. Luna

Self-supervised monocular depth estimation networks are trained to predict scene depth using nearby frames as a supervision signal during training. However, for many applications, sequence information in the form of video frames is also…

Computer Vision and Pattern Recognition · Computer Science 2021-07-15 Jamie Watson , Oisin Mac Aodha , Victor Prisacariu , Gabriel Brostow , Michael Firman

Previous methods on estimating detailed human depth often require supervised training with `ground truth' depth data. This paper presents a self-supervised method that can be trained on YouTube videos without known depth, which makes…

Computer Vision and Pattern Recognition · Computer Science 2020-05-08 Feitong Tan , Hao Zhu , Zhaopeng Cui , Siyu Zhu , Marc Pollefeys , Ping Tan

Estimating 3D human poses from a monocular video is still a challenging task. Many existing methods' performance drops when the target person is occluded by other objects, or the motion is too fast/slow relative to the scale and speed of…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Cheng Yu , Bo Wang , Bo Yang , Robby T. Tan

Dense 3D reconstruction and tracking of dynamic scenes from monocular video remains an important open challenge in computer vision. Progress in this area has been constrained by the scarcity of high-quality datasets with dense, complete,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Zeren Jiang , Yushi Lan , Yihang Luo , Yufan Deng , Zihang Lai , Edgar Sucar , Christian Rupprecht , Iro Laina , Diane Larlus , Chuanxia Zheng , Andrea Vedaldi

We introduce a free-viewpoint rendering method -- HumanNeRF -- that works on a given monocular video of a human performing complex body motions, e.g. a video from YouTube. Our method enables pausing the video at any frame and rendering the…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Chung-Yi Weng , Brian Curless , Pratul P. Srinivasan , Jonathan T. Barron , Ira Kemelmacher-Shlizerman

Depth is a vital piece of information for autonomous vehicles to perceive obstacles. Due to the relatively low price and small size of monocular cameras, depth estimation from a single RGB image has attracted great interest in the research…

Robotics · Computer Science 2021-11-25 Xingshuai Dong , Matthew A. Garratt , Sreenatha G. Anavatti , Hussein A. Abbass

Human is one of the most essential classes in visual recognition tasks such as detection, segmentation, and pose estimation. Although much effort has been put into individual tasks, multi-task learning for these three tasks has been rarely…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Hyeongseok Son , Sangil Jung , Solae Lee , Seongeun Kim , Seung-In Park , ByungIn Yoo

We present a novel paradigm of building an animatable 3D human representation from a monocular video input, such that it can be rendered in any unseen poses and views. Our method is based on a dynamic Neural Radiance Field (NeRF) rigged by…

Computer Vision and Pattern Recognition · Computer Science 2022-08-19 Gusi Te , Xiu Li , Xiao Li , Jinglu Wang , Wei Hu , Yan Lu

Estimating the pose of an uncooperative spacecraft is an important computer vision problem for enabling the deployment of automatic vision-based systems in orbit, with applications ranging from on-orbit servicing to space debris removal.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Leo Pauly , Wassim Rharbaoui , Carl Shneider , Arunkumar Rathinam , Vincent Gaudilliere , Djamila Aouada

It is a classical compute vision problem to obtain real scene depth maps by using a monocular camera, which has been widely concerned in recent years. However, training this model usually requires a large number of artificially labeled…

Computer Vision and Pattern Recognition · Computer Science 2020-09-15 Chunlai Chai , Yukuan Lou , Shijin Zhang

Depth estimation from single monocular images is a key component of scene understanding and has benefited largely from deep convolutional neural networks (CNN) recently. In this article, we take advantage of the recent deep residual…

Computer Vision and Pattern Recognition · Computer Science 2017-08-14 Yuanzhouhan Cao , Zifeng Wu , Chunhua Shen

A reliable sense-and-avoid system is critical to enabling safe autonomous operation of unmanned aircraft. Existing sense-and-avoid methods often require specialized sensors that are too large or power intensive for use on small unmanned…

Computer Vision and Pattern Recognition · Computer Science 2021-11-04 John Mern , Kyle Julian , Rachael E. Tompa , Mykel J. Kochenderfer
‹ Prev 1 8 9 10 Next ›