English
Related papers

Related papers: Learning to compose 6-DoF omnidirectional videos u…

200 papers

Multi-focus image fusion is a technique for obtaining an all-in-focus image in which all objects are in focus to extend the limited depth of field (DoF) of an imaging system. Different from traditional RGB-based methods, this paper presents…

Computer Vision and Pattern Recognition · Computer Science 2018-06-06 Hang Liu , Hengyu Li , Jun Luo , Shaorong Xie , Yu Sun

Omnidirectional videos that capture the entire surroundings are employed in a variety of fields such as VR applications and remote sensing. However, their wide field of view often causes unwanted objects to appear in the videos. This…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Ryosuke Seshimo , Mariko Isogawa

6-DoF robotic grasping is a long-lasting but unsolved problem. Recent methods utilize strong 3D networks to extract geometric grasping representations from depth sensors, demonstrating superior accuracy on common objects but perform…

We propose a novel 3D-aware diffusion-based method for generating photorealistic talking head videos directly from a single identity image and explicit control signals (e.g., expressions). Our method generates Multiplane Images (MPIs) that…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Yuan Li , Ziqian Bai , Feitong Tan , Zhaopeng Cui , Sean Fanello , Yinda Zhang

We present a method for generating a full 360{\deg} orbit video around a person from a single input image. Existing methods typically adapt image-based diffusion models for multi-view synthesis, but yield inconsistent results across views…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Keito Suzuki , Kunyao Chen , Lei Wang , Bang Du , Runfa Blark Li , Peng Liu , Ning Bi , Truong Nguyen

We design a multiscopic vision system that utilizes a low-cost monocular RGB camera to acquire accurate depth estimation. Unlike multi-view stereo with images captured at unconstrained camera poses, the proposed system controls the motion…

Computer Vision and Pattern Recognition · Computer Science 2021-08-21 Weihao Yuan , Rui Fan , Michael Yu Wang , Qifeng Chen

We present a learning-based method to infer plausible high dynamic range (HDR), omnidirectional illumination given an unconstrained, low dynamic range (LDR) image from a mobile phone camera with a limited field of view (FOV). For training…

Computer Vision and Pattern Recognition · Computer Science 2019-04-03 Chloe LeGendre , Wan-Chun Ma , Graham Fyffe , John Flynn , Laurent Charbonnel , Jay Busch , Paul Debevec

Omnidirectional (or 360-degree) images and videos are emergent signals in many areas such as robotics and virtual/augmented reality. In particular, for virtual reality, they allow an immersive experience in which the user is provided with a…

This paper presents a novel end-to-end dynamic time-lapse video generation framework, named DTVNet, to generate diversified time-lapse videos from a single landscape image conditioned on normalized motion vectors. The proposed DTVNet…

Computer Vision and Pattern Recognition · Computer Science 2021-12-20 Jiangning Zhang , Chao Xu , Yong Liu , Yunliang Jiang

Advances in generative modeling have significantly enhanced digital content creation, extending from 2D images to complex 3D and 4D scenes. Despite substantial progress, producing high-fidelity and temporally consistent dynamic 4D content…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 DongFu Yin , Xiaotian Chen , Fei Richard Yu , Xuanchen Li , Xinhao Zhang

We introduce MVSplat360, a feed-forward approach for 360{\deg} novel view synthesis (NVS) of diverse real-world scenes, using only sparse observations. This setting is inherently ill-posed due to minimal overlap among input views and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Yuedong Chen , Chuanxia Zheng , Haofei Xu , Bohan Zhuang , Andrea Vedaldi , Tat-Jen Cham , Jianfei Cai

We propose a learning-based approach for novel view synthesis for multi-camera 360$^{\circ}$ panorama capture rigs. Previous work constructs RGBD panoramas from such data, allowing for view synthesis with small amounts of translation, but…

Computer Vision and Pattern Recognition · Computer Science 2020-08-06 Kai-En Lin , Zexiang Xu , Ben Mildenhall , Pratul P. Srinivasan , Yannick Hold-Geoffroy , Stephen DiVerdi , Qi Sun , Kalyan Sunkavalli , Ravi Ramamoorthi

This work aims to estimate 6Dof (6D) object pose in background clutter. Considering the strong occlusion and background noise, we propose to utilize the spatial structure for better tackling this challenging task. Observing that the 3D mesh…

Computer Vision and Pattern Recognition · Computer Science 2022-06-07 Jianhan Mei , Xudong Jiang , Henghui Ding

This paper presents an approach to estimating the continuous 6-DoF pose of an object from a single RGB image. The approach combines semantic keypoints predicted by a convolutional network (convnet) with a deformable shape model. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2022-04-13 Karl Schmeckpeper , Philip R. Osteen , Yufu Wang , Georgios Pavlakos , Kenneth Chaney , Wyatt Jordan , Xiaowei Zhou , Konstantinos G. Derpanis , Kostas Daniilidis

Omnidirectional depth sensing has its advantage over the conventional stereo systems since it enables us to recognize the objects of interest in all directions without any blind regions. In this paper, we propose a novel wide-baseline…

Computer Vision and Pattern Recognition · Computer Science 2019-08-19 Changhee Won , Jongbin Ryu , Jongwoo Lim

We address the problem of generating a 360-degree image from a single image with a narrow field of view by estimating its surroundings. Previous methods suffered from overfitting to the training resolution and deterministic generation. This…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Naofumi Akimoto , Yuhi Matsuo , Yoshimitsu Aoki

We present a system for multi-level scene awareness for robotic manipulation. Given a sequence of camera-in-hand RGB images, the system calculates three types of information: 1) a point cloud representation of all the surfaces in the scene,…

Robotics · Computer Science 2021-10-18 Yunzhi Lin , Jonathan Tremblay , Stephen Tyree , Patricio A. Vela , Stan Birchfield

To render a spherical (360 degree or omnidirectional) image on planar displays, a 2D image -- called as viewport -- must be obtained by projecting a sphere region on a plane, according to the users viewing direction and a predefined field…

Multimedia · Computer Science 2024-06-06 Falah Jabar , Joao Ascenso , Maria Paula Queluz

We propose an approach for 3D reconstruction and segmentation of a single object placed on a flat surface from an input video. Our approach is to perform dense depth map estimation for multiple views using a proposed objective function that…

Computer Vision and Pattern Recognition · Computer Science 2016-07-29 Tanmay Gupta , Daeyun Shin , Naren Sivagnanadasan , Derek Hoiem

This paper presents a new method to synthesize an image from arbitrary views and times given a collection of images of a dynamic scene. A key challenge for the novel view synthesis arises from dynamic scene reconstruction where epipolar…

Computer Vision and Pattern Recognition · Computer Science 2020-04-06 Jae Shin Yoon , Kihwan Kim , Orazio Gallo , Hyun Soo Park , Jan Kautz