English
Related papers

Related papers: Monocular Building Height Estimation from PhiSat-2…

200 papers

This paper exploits the intrinsic features of urban-scene images and proposes a general add-on module, called height-driven attention networks (HANet), for improving semantic segmentation for urban-scene images. It emphasizes informative…

Computer Vision and Pattern Recognition · Computer Science 2020-04-08 Sungha Choi , Joanne T. Kim , Jaegul Choo

In this work, we propose a novel neural network focusing on semantic labeling of ALS point clouds, which investigates the importance of long-range spatial and channel-wise relations and is termed as global relation-aware attentional network…

Computer Vision and Pattern Recognition · Computer Science 2020-12-29 Rong Huang , Yusheng Xu , Uwe Stilla

The scarcity of comprehensive datasets in surveillance, identification, image retrieval systems, and healthcare poses a significant challenge for researchers in exploring new methodologies and advancing knowledge in these respective fields.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Pronay Debnath , Usafa Akther Rifa , Busra Kamal Rafa , Ali Haider Talukder Akib , Md. Aminur Rahman

In this paper, we address the challenge of land use and land cover classification using Sentinel-2 satellite images. The Sentinel-2 satellite images are openly and freely accessible provided in the Earth observation program Copernicus. We…

Computer Vision and Pattern Recognition · Computer Science 2019-02-04 Patrick Helber , Benjamin Bischke , Andreas Dengel , Damian Borth

Small object detection remains a significant challenge due to feature degradation from downsampling, mutual occlusion in dense clusters, and complex background interference. To address these issues, this paper proposes FSDETR, a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Jianchao Huang , Fengming Zhang , Haibo Zhu , Tao Yan

Relying on monocular image data for precise 3D object detection remains an open problem, whose solution has broad implications for cost-sensitive applications such as traffic monitoring. We present UrbanNet, a modular architecture for long…

Computer Vision and Pattern Recognition · Computer Science 2021-10-13 Juan Carrillo , Steven Waslander

Access to labeled reference data is one of the grand challenges in supervised machine learning endeavors. This is especially true for an automated analysis of remote sensing images on a global scale, which enables us to address global…

Crowd counting is an important yet challenging task in computer vision due to serious occlusions, complex background and large scale variations, etc. Multi-column architecture is widely adopted to overcome these challenges, yielding…

Computer Vision and Pattern Recognition · Computer Science 2020-07-29 Junhao Cheng , Zhuojun Chen , XinYu Zhang , Yizhou Li , Xiaoyuan Jing

In medical images, various types of lesions often manifest significant differences in their shape and texture. Accurate medical image segmentation demands deep learning models with robust capabilities in multi-scale and boundary feature…

Image and Video Processing · Electrical Eng. & Systems 2024-08-20 Zhenhuan Zhou , Along He , Yanlin Wu , Rui Yao , Xueshuo Xie , Tao Li

Head pose estimation is a challenging task that aims to solve problems related to predicting three dimensions vector, that serves for many applications in human-robot interaction or customer behavior. Previous researches have proposed some…

Computer Vision and Pattern Recognition · Computer Science 2021-11-16 Linh Nguyen Viet , Tuan Nguyen Dinh , Hoang Nguyen Viet , Duc Tran Minh , Long Tran Quoc

We propose MoGe-2, an advanced open-domain geometry estimation model that recovers a metric scale 3D point map of a scene from a single image. Our method builds upon the recent monocular geometry estimation approach, MoGe, which predicts…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Ruicheng Wang , Sicheng Xu , Yue Dong , Yu Deng , Jianfeng Xiang , Zelong Lv , Guangzhong Sun , Xin Tong , Jiaolong Yang

We propose a learning-based method that solves monocular stereo and can be extended to fuse depth information from multiple target frames. Given two unconstrained images from a monocular camera with known intrinsic calibration, our network…

Computer Vision and Pattern Recognition · Computer Science 2019-09-13 Kaixuan Wang , Shaojie Shen

Lightweight vision classification models such as MobileNet, ShuffleNet, and EfficientNet are increasingly deployed in mobile and embedded systems, yet their performance has been predominantly benchmarked on ImageNet. This raises critical…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Weidong Zhang , Pak Lun Kevin Ding , Huan Liu

We consider the problem of segmentation and classification of high-resolution and hyperspectral remote sensing images. Unlike conventional natural (RGB) images, the inherent large scale and complex structures of remote sensing images pose…

Computer Vision and Pattern Recognition · Computer Science 2020-08-21 Qingsong Xu , Xin Yuan , Chaojun Ouyang , Yue Zeng

This study uses the challenging and publicly available SpaceNet dataset to establish a performance baseline for a state-of-the-art object detector in satellite imagery. Specifically, we examine how various features of the data affect…

Computer Vision and Pattern Recognition · Computer Science 2018-08-08 Eliza Mace , Keith Manville , Monica Barbu-McInnis , Michael Laielli , Matthew Klaric , Samuel Dooley

This study investigates whether the geospatial and multimodal features encoded in \textit{Earth Embeddings} can effectively guide deep learning (DL) regression models for regional surface height mapping. In particular, we focused on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Alireza Hamoudzadeh , Valeria Belloni , Roberta Ravanelli

We present a multi-modal classification framework that fuses satellite and street-level imagery through a Perceiver IO architecture operating on spatial patch tokens from a shared DINOv2 backbone. The design naturally handles a variable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Niels Sombekke , Rob G. J. Wijnhoven , Martin R. Oswald

Accurate estimation of building heights using very high resolution (VHR) synthetic aperture radar (SAR) imagery is crucial for various urban applications. This paper introduces a Deep Learning (DL)-based methodology for automated building…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Babak Memar , Luigi Russo , Silvia Liberata Ullo , Paolo Gamba

Aerial imagery has been increasingly adopted in mission-critical tasks, such as traffic surveillance, smart cities, and disaster assistance. However, identifying objects from aerial images faces the following challenges: 1) objects of…

Computer Vision and Pattern Recognition · Computer Science 2020-01-24 Ziyang Tang , Xiang Liu , Guangyu Shen , Baijian Yang

Optical Coherence Tomography (OCT) imaging is pivotal in diagnosing ophthalmic conditions by providing detailed cross-sectional images of the anterior and posterior segments of the eye. Nonetheless, speckle noise and other imaging artifacts…

Image and Video Processing · Electrical Eng. & Systems 2024-09-26 Akkidas Noel Prakash , Jahnvi Sai Ganta , Ramaswami Krishnadas , Tin A. Tunc , Satish K Panda