English
Related papers

Related papers: DUNIA: Pixel-Sized Embeddings via Cross-Modal Alig…

200 papers

We explore simple methods for adapting a trained multi-task UNet which predicts canopy cover and height to a new geographic setting using remotely sensed data without the need of training a domain-adaptive classifier and extensive…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 John Francis , Stephen Law

We consider the problem of source-free unsupervised category-level pose estimation from only RGB images to a target domain without any access to source domain data or 3D annotations during adaptation. Collecting and annotating real-world 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-01-22 Prakhar Kaushik , Aayush Mishra , Adam Kortylewski , Alan Yuille

We introduce a new dataset, MELINDA, for Multimodal biomEdicaL experImeNt methoD clAssification. The dataset is collected in a fully automated distant supervision manner, where the labels are obtained from an existing curated database, and…

Computation and Language · Computer Science 2020-12-18 Te-Lin Wu , Shikhar Singh , Sayan Paul , Gully Burns , Nanyun Peng

We present the first work demonstrating that a pure Mamba block can achieve efficient Dense Global Fusion, meanwhile guaranteeing top performance for camera-LiDAR multi-modal 3D object detection. Our motivation stems from the observation…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Hanshi Wang , Jin Gao , Weiming Hu , Zhipeng Zhang

Unsupervised domain adaptation (UDA) in 3D segmentation tasks presents a formidable challenge, primarily stemming from the sparse and unordered nature of point cloud data. Especially for LiDAR point clouds, the domain discrepancy becomes…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Xidong Peng , Runnan Chen , Feng Qiao , Lingdong Kong , Youquan Liu , Yujing Sun , Tai Wang , Xinge Zhu , Yuexin Ma

Mapping standing dead trees is crucial for acquiring information on the effects of climate change on forests and forest biodiversity. However, leveraging high-quality aerial imagery for dead tree segmentation poses challenges due to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Mete Ahishali , Anis Ur Rahman , Einari Heinaro , Aysen Degerli , Samuli Junttila

In the context of pose-invariant object recognition and retrieval, we demonstrate that it is possible to achieve significant improvements in performance if both the category-based and the object-identity-based embeddings are learned…

Computer Vision and Pattern Recognition · Computer Science 2024-03-04 Rohan Sarkar , Avinash Kak

Partially-supervised multi-organ medical image segmentation aims to develop a unified semantic segmentation model by utilizing multiple partially-labeled datasets, with each dataset providing labels for a single class of organs. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Xixi Jiang , Dong Zhang , Xiang Li , Kangyi Liu , Kwang-Ting Cheng , Xin Yang

Automatic multi-class object detection in remote sensing images in unconstrained scenarios is of high interest for several applications including traffic monitoring and disaster management. The huge variation in object scale, orientation,…

Computer Vision and Pattern Recognition · Computer Science 2018-11-02 Seyed Majid Azimi , Eleonora Vig , Reza Bahmanyar , Marco Körner , Peter Reinartz

Learning with few labeled data is a key challenge for visual recognition, as deep neural networks tend to overfit using a few samples only. One of the Few-shot learning methods called metric learning addresses this challenge by first…

Computer Vision and Pattern Recognition · Computer Science 2022-03-28 Li Ke , Meng Pan , Weigao Wen , Dong Li

Learning a metric of natural image patches is an important tool for analyzing images. An efficient means is to train a deep network to map an image patch to a vector space, in which the Euclidean distance reflects patch similarity. Previous…

Computer Vision and Pattern Recognition · Computer Science 2018-07-10 Dov Danon , Hadar Averbuch-Elor , Ohad Fried , Daniel Cohen-Or

LiDAR (short for "Light Detection And Ranging" or "Laser Imaging, Detection, And Ranging") technology can be used to provide detailed three-dimensional elevation maps of urban and rural landscapes. To date, airborne LiDAR imaging has been…

Machine Learning · Computer Science 2022-08-03 Matthew Stevenson , Christophe Mues , Cristián Bravo

Due to their radiation hardness, kilohertz frame rates, and high dynamic range, hybrid pixel detectors have recently expanded their application range to electron diffraction and recently also electron imaging. However, these detectors…

Cross-view geo-localization aims to spot images of the same location shot from two platforms, e.g., the drone platform and the satellite platform. Existing methods usually focus on optimizing the distance between one embedding with others…

Computer Vision and Pattern Recognition · Computer Science 2022-11-11 Tingyu Wang , Zhedong Zheng , Zunjie Zhu , Yuhan Gao , Yi Yang , Chenggang Yan

In recent years, hyperspectral imaging, also known as imaging spectroscopy, has been paid an increasing interest in geoscience and remote sensing community. Hyperspectral imagery is characterized by very rich spectral information, which…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Danfeng Hong , Jing Yao , Xin Wu , Jocelyn Chanussot , Xiao Xiang Zhu

While the Earth observation community has witnessed a surge in high-impact foundation models and global Earth embedding datasets, a significant barrier remains in translating these academic assets into freely accessible tools. This tutorial…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Yijie Zheng , Weijie Wu , Bingyue Wu , Long Zhao , Guoqing Li , Mikolaj Czerkawski , Konstantin Klemmer

We propose a vision-based method that localizes a ground vehicle using publicly available satellite imagery as the only prior knowledge of the environment. Our approach takes as input a sequence of ground-level images acquired by the…

Robotics · Computer Science 2022-03-08 Dong-Ki Kim , Matthew R. Walter

Endeavors have been recently made to transfer knowledge from the labeled pinhole image domain to the unlabeled panoramic image domain via Unsupervised Domain Adaptation (UDA). The aim is to tackle the domain gaps caused by the style…

Computer Vision and Pattern Recognition · Computer Science 2023-08-11 Xu Zheng , Tianbo Pan , Yunhao Luo , Lin Wang

Learning transferable multimodal embeddings for urban environments is challenging because urban understanding is inherently spatial, yet existing datasets and benchmarks lack explicit alignment between street-view images and urban…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Jie Zhang , Xingtong Yu , Yuan Fang , Rudi Stouffs , Zdravko Trivic

Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creating misalignment between the two tasks and preventing fully end-to-end optimization from…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Zhiheng Liu , Weiming Ren , Xiaoke Huang , Shoufa Chen , Tianhong Li , Mengzhao Chen , Yatai Ji , Sen He , Jonas Schult , Belinda Zeng , Tao Xiang , Wenhu Chen , Ping Luo , Luke Zettlemoyer , Yuren Cong