English
Related papers

Related papers: SALVe: Semantic Alignment Verification for Floorpl…

200 papers

We present FLARE, a feed-forward model designed to infer high-quality camera poses and 3D geometry from uncalibrated sparse-view images (i.e., as few as 2-8 inputs), which is a challenging yet practical setting in real-world applications.…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Shangzhan Zhang , Jianyuan Wang , Yinghao Xu , Nan Xue , Christian Rupprecht , Xiaowei Zhou , Yujun Shen , Gordon Wetzstein

Prior panorama stitching approaches heavily rely on pairwise feature correspondences and are unable to leverage geometric consistency across multiple views. This leads to severe distortion and misalignment, especially in challenging scenes…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Zhengdong Zhu , Weiyi Xue , Zuyuan Yang , Wenlve Zhou , Zhiheng Zhou

Indoor navigation remains a critical challenge for people with visual impairments. The current solutions mainly rely on infrastructure-based systems, which limit their ability to navigate safely in dynamic environments. We propose a novel…

Artificial Intelligence · Computer Science 2026-05-13 Aydin Ayanzadeh , Tim Oates

Reconstructing 3D shape and pose of static objects from a single image is an essential task for various industries, including robotics, augmented reality, and digital content creation. This can be done by directly predicting 3D shape in…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Florian Langer , Ignas Budvytis , Roberto Cipolla

We describe a set of tools for analyzing, visualizing, and assessing architectural/construction progress with unordered photo collections and 3D building models. With our interface, a user guides the registration of the model in one of the…

Graphics · Computer Science 2019-12-30 Kevin Karsch , Mani Golparvar-Fard , David Forsyth

Inter-object relations underpin spatial intelligence, yet existing representations -- linguistic prepositions or object-level scene graphs -- are too coarse to specify which regions actually support, contain, or contact one another, leading…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yinuo Bai , Peijun Xu , Kuixiang Shao , Yuyang Jiao , Jingxuan Zhang , Kaixin Yao , Jiayuan Gu , Jingyi Yu

Our work introduces SAVeD (Semantically Aware Version Detection), a contrastive learning-based framework for identifying versions of structured datasets without relying on metadata, labels, or integration-based assumptions. SAVeD addresses…

Machine Learning · Computer Science 2026-01-13 Artem Frenk , Roee Shraga

We introduce Scan2Plan, a novel approach for accurate estimation of a floorplan from a 3D scan of the structural elements of indoor environments. The proposed method incorporates a two-stage approach where the initial stage clusters an…

Computer Vision and Pattern Recognition · Computer Science 2020-03-17 Ameya Phalak , Vijay Badrinarayanan , Andrew Rabinovich

Visual entailment (VE) is to recognize whether the semantics of a hypothesis text can be inferred from the given premise image, which is one special task among recent emerged vision and language understanding tasks. Currently, most of the…

Computer Vision and Pattern Recognition · Computer Science 2022-11-17 Biwei Cao , Jiuxin Cao , Jie Gui , Jiayun Shen , Bo Liu , Lei He , Yuan Yan Tang , James Tin-Yau Kwok

We address the task of estimating 6D camera poses from sparse-view image sets (2-8 images). This task is a vital pre-processing stage for nearly all contemporary (neural) reconstruction algorithms but remains challenging given sparse views,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Amy Lin , Jason Y. Zhang , Deva Ramanan , Shubham Tulsiani

Visual SLAM (Simultaneous Localization and Mapping) based on planar features has found widespread applications in fields such as environmental structure perception and augmented reality. However, current research faces challenges in…

Robotics · Computer Science 2024-02-15 Xinggang Hu , Yanmin Wu , Mingyuan Zhao , Linghao Yang , Xiangkui Zhang , Xiangyang Ji

We present an approach for recognizing all objects in a scene and estimating their full pose from an accurate 3D instance-aware semantic reconstruction using an RGB-D camera. Our framework couples convolutional neural networks (CNNs) and a…

Robotics · Computer Science 2019-10-01 Dinh-Cuong Hoang , Todor Stoyanov , Achim J. Lilienthal

Given only a few glimpses of an environment, how much can we infer about its entire floorplan? Existing methods can map only what is visible or immediately apparent from context, and thus require substantial movements through a space to…

Computer Vision and Pattern Recognition · Computer Science 2021-01-01 Senthil Purushwalkam , Sebastian Vicenc Amengual Gari , Vamsi Krishna Ithapu , Carl Schissler , Philip Robinson , Abhinav Gupta , Kristen Grauman

We present ANISE, a method that reconstructs a 3D~shape from partial observations (images or sparse point clouds) using a part-aware neural implicit shape representation. The shape is formulated as an assembly of neural implicit functions,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-10 Dmitry Petrov , Matheus Gadelha , Radomir Mech , Evangelos Kalogerakis

Reconstructing building floor plans from point cloud data is key for indoor navigation, BIM, and precise measurements. Traditional methods like geometric algorithms and Mask R-CNN-based deep learning often face issues with noise, limited…

Reconstruction of indoor surfaces with limited texture information or with repeated textures, a situation common in walls and ceilings, may be difficult with a monocular Structure from Motion system. We propose a Semantic Room Wireframe…

Computer Vision and Pattern Recognition · Computer Science 2022-06-02 David Gillsjö , Gabrielle Flood , Kalle Åström

The hierarchical structure of 3D scene graphs shows a high relevance for representations purposes, as it fits common patterns from man-made environments. But, additionally, the semantic and geometric information in such hierarchical…

Robotics · Computer Science 2025-10-06 Hriday Bavle , Jose Luis Sanchez-Lopez , Muhammad Shaheer , Javier Civera , Holger Voos

Two-view structure-from-motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM. Existing deep learning-based approaches formulate the problem by either recovering absolute pose scales from two consecutive frames or predicting…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Jianyuan Wang , Yiran Zhong , Yuchao Dai , Stan Birchfield , Kaihao Zhang , Nikolai Smolyanskiy , Hongdong Li

Indoor panorama typically consists of human-made structures parallel or perpendicular to gravity. We leverage this phenomenon to approximate the scene in a 360-degree image with (H)orizontal-planes and (V)ertical-planes. To this end, we…

Computer Vision and Pattern Recognition · Computer Science 2021-09-10 Cheng Sun , Chi-Wei Hsiao , Ning-Hsu Wang , Min Sun , Hwann-Tzong Chen

One practical approach to infer 3D scene structure from a single image is to retrieve a closely matching 3D model from a database and align it with the object in the image. Existing methods rely on supervised training with images and pose…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Pattaramanee Arsomngern , Sasikarn Khwanmuang , Matthias Nießner , Supasorn Suwajanakorn