中文
相关论文

相关论文: GEOBIND: Binding Text, Image, and Audio through Sa…

200 篇论文

Image retrieval enables an efficient search through vast amounts of satellite imagery and returns similar images to a query. Deep learning models can identify images across various semantic concepts without the need for annotations. This…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Benedikt Blumenstiel , Viktoria Moor , Romeo Kienzler , Thomas Brunschwiler

Satellite imaging has a central role in monitoring, detecting and estimating the intensity of key natural phenomena. One important feature of satellite images is the trade-off between spatial/spectral resolution and their revisiting time, a…

图像与视频处理 · 电气工程与系统科学 2022-04-28 Haoqing Li , Bhavia Duvviri , Ricardo Borsoi , Tales Imbiriba , Edward Beighley , Deniz Erdogmus , Pau Closas

As is expressed in the adage "a picture is worth a thousand words", when using spoken language to communicate visual information, brevity can be a challenge. This work describes a novel technique for leveraging machine-learned feature…

音频与语音处理 · 电气工程与系统科学 2021-08-27 Andrew Port , Chelhwon Kim , Mitesh Patel

Developments in three-dimensional real worlds promote the integration of geoinformation and building information models (BIM) known as GeoBIM in urban construction. Light detection and ranging (LiDAR) integrated with global navigation…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Jie Shao , Wei Yao , Puzuo Wang , Zhiyi He , Lei Luo

Multi-modal learning combines various modalities to provide a comprehensive understanding of real-world problems. A common strategy is to directly bind different modalities together in a specific joint embedding space. However, the…

机器学习 · 计算机科学 2026-02-09 Zhuo Huang , Runnan Chen , Bo Han , Gang Niu , Masashi Sugiyama , Tongliang Liu

Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimodal alignment. We extend the application of these models,…

计算机视觉与模式识别 · 计算机科学 2023-11-08 Sooyoung Park , Arda Senocak , Joon Son Chung

We propose a novel geometric approach for learning bilingual mappings given monolingual embeddings and a bilingual dictionary. Our approach decouples learning the transformation from the source language to the target language into (a)…

机器学习 · 计算机科学 2018-12-19 Pratik Jawanpuria , Arjun Balgovind , Anoop Kunchukuttan , Bamdev Mishra

One of the promising applications of satellite images is building construction monitoring. It allows to control the construction progress around the world even in the locations that are hard to reach. One of the main hurdles of this…

计算机视觉与模式识别 · 计算机科学 2022-10-03 Insaf Ashrapov , Dmitriy Malakhov , Anton Marchenkov , Anton Lulin , Dani El-Ayyass

A comprehensive understanding of vision and language and their interrelation are crucial to realize the underlying similarities and differences between these modalities and to learn more generalized, meaningful representations. In recent…

计算机视觉与模式识别 · 计算机科学 2021-12-10 Anindya Sundar Das , Sriparna Saha

Language grounding aims at linking the symbolic representation of language (e.g., words) into the rich perceptual knowledge of the outside world. The general approach is to embed both textual and visual information into a common space -the…

计算与语言 · 计算机科学 2021-09-15 Hassan Shahmohammadi , Hendrik P. A. Lensch , R. Harald Baayen

While integrating multiple modalities has the potential to improve environmental monitoring, current approaches struggle to combine data sources with heterogeneous formats or contents. A central difficulty arises when combining continuous…

计算与语言 · 计算机科学 2026-03-27 Valerie Zermatten , Chiara Vanalli , Gencer Sumbul , Diego Marcos , Devis Tuia

Satellite imagery is widely used in many application sectors, including agriculture, navigation, and urban planning. Frequently, satellite imagery involves both large numbers of images as well as high pixel counts, making satellite datasets…

计算机视觉与模式识别 · 计算机科学 2021-05-27 Joshua Abraham , Calden Wloka

This paper proposes a novel method for geo-tracking, i.e. continuous metric self-localization in outdoor environments by registering a vehicle's sensor information with aerial imagery of an unseen target region. Geo-tracking methods offer…

计算机视觉与模式识别 · 计算机科学 2022-09-12 Florian Fervers , Sebastian Bullinger , Christoph Bodensteiner , Michael Arens , Rainer Stiefelhagen

Previous approaches in singer identification have used one of monophonic vocal tracks or mixed tracks containing multiple instruments, leaving a semantic gap between these two domains of audio. In this paper, we present a system to learn a…

声音 · 计算机科学 2019-06-27 Kyungyun Lee , Juhan Nam

This paper proposes a method for learning joint embeddings of images and text using a two-branch neural network with multiple layers of linear projections followed by nonlinearities. The network is trained using a large margin objective…

计算机视觉与模式识别 · 计算机科学 2016-04-15 Liwei Wang , Yin Li , Svetlana Lazebnik

Solo piano music, despite being a single-instrument medium, possesses significant expressive capabilities, conveying rich semantic information across genres, moods, and styles. However, current general-purpose music representation models,…

声音 · 计算机科学 2025-09-05 Hayeon Bang , Eunjin Choi , Seungheon Doh , Juhan Nam

Satellite imagery differs fundamentally from natural images: its aerial viewpoint, very high resolution, diverse scale variations, and abundance of small objects demand both region-level spatial reasoning and holistic scene understanding.…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Emanuel Sánchez Aimar , Gulnaz Zhambulova , Fahad Shahbaz Khan , Yonghao Xu , Michael Felsberg

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing…

机器人学 · 计算机科学 2026-02-05 Rui Tang , Guankun Wang , Long Bai , Huxin Gao , Jiewen Lai , Chi Kit Ng , Jiazheng Wang , Fan Zhang , Hongliang Ren

Satellite imagery has long been an attractive data source that provides a wealth of information on human-inhabited areas. While super resolution satellite images are rapidly becoming available, little study has focused on how to extract…

计算机视觉与模式识别 · 计算机科学 2019-12-19 Sungwon Han , Donghyun Ahn , Hyunji Cha , Jeasurk Yang , Sungwon Park , Meeyoung Cha

Worldwide visual geo-localization aims to determine the geographic location of an image anywhere on Earth using only its visual content. Despite recent progress, learning expressive representations of geographic space remains challenging…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Angel Daruna , Nicholas Meegan , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar