English
Related papers

Related papers: Generalized Geometry Encoding Volume for Real-time…

200 papers

Image-based volumetric humans using pixel-aligned features promise generalization to unseen poses and identities. Prior work leverages global spatial encodings and multi-view geometric consistency to reduce spatial ambiguity. However,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Marko Mihajlovic , Aayush Bansal , Michael Zollhoefer , Siyu Tang , Shunsuke Saito

Estimating depth from a single image represents an attractive alternative to more traditional approaches leveraging multiple cameras. In this field, deep learning yielded outstanding results at the cost of needing large amounts of data…

Computer Vision and Pattern Recognition · Computer Science 2019-08-14 Lorenzo Andraghetti , Panteleimon Myriokefalitakis , Pier Luigi Dovesi , Belen Luque , Matteo Poggi , Alessandro Pieropan , Stefano Mattoccia

Modern cameras are equipped with a wide array of sensors that enable recording the geospatial context of an image. Taking advantage of this, we explore depth estimation under the assumption that the camera is geocalibrated, a problem we…

Computer Vision and Pattern Recognition · Computer Science 2021-09-22 Scott Workman , Hunter Blanton

Recently, neural implicit surfaces learning by volume rendering has become popular for multi-view reconstruction. However, one key challenge remains: existing approaches lack explicit multi-view geometry constraints, hence usually fail to…

Computer Vision and Pattern Recognition · Computer Science 2022-06-01 Qiancheng Fu , Qingshan Xu , Yew-Soon Ong , Wenbing Tao

While interest in models that generalize at test time to new compositions has risen in recent years, benchmarks in the visually-grounded domain have thus far been restricted to synthetic images. In this work, we propose COVR, a new test-bed…

Computation and Language · Computer Science 2021-09-23 Ben Bogin , Shivanshu Gupta , Matt Gardner , Jonathan Berant

Recovering the scene depth from a single image is an ill-posed problem that requires additional priors, often referred to as monocular depth cues, to disambiguate different 3D interpretations. In recent works, those priors have been learned…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Lam Huynh , Phong Nguyen-Ha , Jiri Matas , Esa Rahtu , Janne Heikkila

Training model to generate data has increasingly attracted research attention and become important in modern world applications. We propose in this paper a new geometry-based optimization approach to address this problem. Orthogonal to…

Machine Learning · Computer Science 2017-08-18 Trung Le , Hung Vu , Tu Dinh Nguyen , Dinh Phung

Due to the inherent ill-posed nature of 2D-3D projection, monocular 3D object detection lacks accurate depth recovery ability. Although the deep neural network (DNN) enables monocular depth-sensing from high-level learned features, the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-17 Qing Lian , Peiliang Li , Xiaozhi Chen

Deploying multimodal models in real-world scenarios requires generalization to new environments where recording conditions differ from training, a challenge known as multimodal domain generalization (MMDG). Standard architectures employ…

Machine Learning · Computer Science 2026-05-05 Yavuz Yarici , Ghassan AlRegib

Semantic segmentation algorithms require access to well-annotated datasets captured under diverse illumination conditions to ensure consistent performance. However, poor visibility conditions at varying illumination conditions result in…

Computer Vision and Pattern Recognition · Computer Science 2022-03-01 Pranjay Shyam , Antyanta Bangunharcana , Kuk-Jin Yoon , Kyung-Soo Kim

This work extends the generalized nearest neighbor decoding (GNND), originally developed as a receiver architecture for memoryless channels, to a vectorized GNND (Vec-GNND) suitable for in-block memory (IBM) channels. Leveraging the…

Information Theory · Computer Science 2026-05-18 Yuhao Liu , Xinwei Li , Shuqin Pang , Hao Wu , Wenyi Zhang

This paper proposes 3DGeoDet, a novel geometry-aware 3D object detection approach that effectively handles single- and multi-view RGB images in indoor and outdoor environments, showcasing its general-purpose applicability. The key challenge…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yi Zhang , Yi Wang , Yawen Cui , Lap-Pui Chau

This paper presents a stereo object matching method that exploits both 2D contextual information from images as well as 3D object-level information. Unlike existing stereo matching methods that exclusively focus on the pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2021-03-25 Jaesung Choe , Kyungdon Joo , Francois Rameau , In So Kweon

Accurate recovery of 3D geometrical surfaces from calibrated 2D multi-view images is a fundamental yet active research area in computer vision. Despite the steady progress in multi-view stereo reconstruction, most existing methods are still…

Computer Vision and Pattern Recognition · Computer Science 2016-01-20 Zhaoxin Li , Kuanquan Wang , Wangmeng Zuo , Deyu Meng , Lei Zhang

The generalized eigenvalue problem (GEP) serves as a cornerstone in a wide range of applications in numerical linear algebra and scientific computing. However, traditional approaches that aim to maximize the classical Rayleigh quotient…

Optimization and Control · Mathematics 2025-07-04 Xiaozhi Liu , Yong Xia

Spatial representation learning is essential for GeoAI applications such as urban analytics, enabling the encoding of shapes, locations, and spatial relationships (topological and distance-based) of geo-entities like points, polylines, and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Chen Chu , Cyrus Shahabi

Despite the significant advances in domain generalized stereo matching, existing methods still exhibit domain-specific preferences when transferring from synthetic to real domains, hindering their practical applications in complex and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Peng Xu , Zhiyu Xiang , Jingyun Fu , Tianyu Pu , Hanzhi Zhong , Eryun Liu

Recent advances in deep learning have made it possible to quantify urban metrics at fine resolution, and over large extents using street-level images. Here, we focus on measuring urban tree cover using Google Street View (GSV) images.…

Computer Vision and Pattern Recognition · Computer Science 2019-10-17 Bill Yang Cai , Xiaojiang Li , Ian Seiferling , Carlo Ratti

Despite the empirical success of Diffusion Models (DMs) and Variational Autoencoders (VAEs), their generalization performance remains theoretically underexplored, especially lacking a full consideration of the shared encoder-generator…

Machine Learning · Computer Science 2025-06-03 Qi Chen , Jierui Zhu , Florian Shkurti

This paper presents a novel iToF-RGB fusion framework designed to address the inherent limitations of indirect Time-of-Flight (iToF) depth sensing, such as low spatial resolution, limited field-of-view (FoV), and structural distortion in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Yansong Du , Yutong Deng , Yuting Zhou , Feiyu Jiao , Jian Song , Xun Guan