中文
相关论文

相关论文: Monocular Multi-Layer Layout Estimation for Wareho…

200 篇论文

Monocular 3D object understanding has largely been cast as a 2D RoI-to-3D box lifting problem. However, emerging downstream applications require image-plane geometry (e.g., projected 3D box corners) which cannot be easily obtained without…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Changwoo Jeon , Rishi Upadhyay , Achuta Kadambi

Room layout estimation from multiple-perspective images is poorly investigated due to the complexities that emerge from multi-view geometry, which requires muti-step solutions such as camera intrinsic and extrinsic estimation, image…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Yaxuan Huang , Xili Dai , Jianan Wang , Xianbiao Qi , Yixing Yuan , Xiangyu Yue

A recent strand of work in view synthesis uses deep learning to generate multiplane images (a camera-centric, layered 3D representation) given two or more input images at known viewpoints. We apply this representation to single-view view…

计算机视觉与模式识别 · 计算机科学 2020-04-24 Richard Tucker , Noah Snavely

In this work, we aim to predict human eye fixation with view-free scenes based on an end-to-end deep learning architecture. Although Convolutional Neural Networks (CNNs) have made substantial improvement on human attention prediction, it is…

计算机视觉与模式识别 · 计算机科学 2018-03-26 Wenguan Wang , Jianbing Shen

In 3D shape recognition, multi-view based methods leverage human's perspective to analyze 3D shapes and have achieved significant outcomes. Most existing research works in deep learning adopt handcrafted networks as backbones due to their…

计算机视觉与模式识别 · 计算机科学 2020-12-11 Zhaoqun Li , Hongren Wang , Jinxing Li

Monocular normal estimation aims to estimate the normal map from a single RGB image of an object under arbitrary lights. Existing methods rely on deep models to directly predict normal maps. However, they often suffer from 3D misalignment:…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Zongrui Li , Xinhua Ma , Minghui Hu , Yunqing Zhao , Yingchen Yu , Qian Zheng , Chang Liu , Xudong Jiang , Song Bai

Monocular vision-based Simultaneous Localization and Mapping (SLAM) is used for various purposes due to its advantages in cost, simple setup, as well as availability in the environments where navigation with satellites is not effective.…

机器人学 · 计算机科学 2018-10-03 Young-Hee Lee , Chen Zhu , Gabriele Giorgi , Christoph Günther

Monocular depth estimation in the wild inherently predicts depth up to an unknown scale. To resolve scale ambiguity issue, we present a learning algorithm that leverages monocular simultaneous localization and mapping (SLAM) with…

计算机视觉与模式识别 · 计算机科学 2022-03-11 Jaehoon Choi , Dongki Jung , Yonghan Lee , Deokhwa Kim , Dinesh Manocha , Donghwan Lee

We consider the problem of depth estimation from a single monocular image in this work. It is a challenging task as no reliable depth cues are available, e.g., stereo correspondences, motions, etc. Previous efforts have been focusing on…

计算机视觉与模式识别 · 计算机科学 2015-10-01 Fayao Liu , Chunhua Shen , Guosheng Lin

We design a multiscopic vision system that utilizes a low-cost monocular RGB camera to acquire accurate depth estimation. Unlike multi-view stereo with images captured at unconstrained camera poses, the proposed system controls the motion…

计算机视觉与模式识别 · 计算机科学 2021-08-21 Weihao Yuan , Rui Fan , Michael Yu Wang , Qifeng Chen

Monocular 3D lane detection aims to estimate the 3D position of lanes from frontal-view (FV) images. However, existing methods are fundamentally constrained by the inherent ambiguity of single-frame input, which leads to inaccurate…

计算机视觉与模式识别 · 计算机科学 2025-11-06 Huan Zheng , Wencheng Han , Tianyi Yan , Cheng-zhong Xu , Jianbing Shen

Pooling layers (e.g., max and average) may overlook important information encoded in the spatial arrangement of pixel intensity and/or feature values. We propose a novel lacunarity pooling layer that aims to capture the spatial…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Akshatha Mohan , Joshua Peeples

Given a single RGB image of a complex outdoor road scene in the perspective view, we address the novel problem of estimating an occlusion-reasoned semantic scene layout in the top-view. This challenging problem not only requires an accurate…

计算机视觉与模式识别 · 计算机科学 2018-03-30 Samuel Schulter , Menghua Zhai , Nathan Jacobs , Manmohan Chandraker

The representation of geometry in real-time 3D perception systems continues to be a critical research issue. Dense maps capture complete surface shape and can be augmented with semantic labels, but their high dimensionality makes them…

计算机视觉与模式识别 · 计算机科学 2019-04-16 Michael Bloesch , Jan Czarnowski , Ronald Clark , Stefan Leutenegger , Andrew J. Davison

In this work, we present a robotic solution to automate the task of wall construction. To that end, we present an end-to-end visual perception framework that can quickly detect and localize bricks in a clutter. Further, we present a light…

机器人学 · 计算机科学 2021-07-28 Mohit Vohra , Ashish Kumar , Ravi Prakash , Laxmidhar Behera

Monocular depth estimation methods traditionally train deep models to infer depth directly from RGB pixels. This implicit learning often overlooks explicit monocular cues that the human visual system relies on, such as occlusion boundaries,…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Calin Teodor Ioan

Estimating the 3D position and orientation of objects in the environment with a single RGB camera is a critical and challenging task for low-cost urban autonomous driving and mobile robots. Most of the existing algorithms are based on the…

计算机视觉与模式识别 · 计算机科学 2021-02-02 Yuxuan Liu , Yuan Yixuan , Ming Liu

An ability to generalize unconstrained conditions such as severe occlusions and large pose variations remains a challenging goal to achieve in face alignment. In this paper, a multistage model based on deep neural networks is proposed which…

计算机视觉与模式识别 · 计算机科学 2020-02-05 Huabin Wang , Rui Cheng , Jian Zhou , Liang Tao , Hon Keung Kwan

Scene flow is the dense 3D reconstruction of motion and geometry of a scene. Most state-of-the-art methods use a pair of stereo images as input for full scene reconstruction. These methods depend a lot on the quality of the RGB images and…

计算机视觉与模式识别 · 计算机科学 2020-08-20 Rishav , Ramy Battrawy , René Schuster , Oliver Wasenmüller , Didier Stricker

Monocular 3D lane detection is challenging due to the difficulty in capturing depth information from single-camera images. A common strategy involves transforming front-view (FV) images into bird's-eye-view (BEV) space through inverse…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Dongxin Lyu , Han Huang , Cheng Tan , Zimu Li