中文
相关论文

相关论文: Sat2Sound: A Unified Framework for Zero-Shot Sound…

200 篇论文

Object-goal navigation in open-vocabulary settings requires agents to locate novel objects in unseen environments, yet existing approaches suffer from opaque decision-making processes and low success rate on locating unseen objects. To…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Wentao Xiang , Haokang Zhang , Tianhang Yang , Zedong Chu , Ruihang Chu , Shichao Xie , Yujian Yuan , Jian Sun , Zhining Gu , Junjie Wang , Xiaolong Wu , Mu Xu , Yujiu Yang

We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on visual generative…

Surficial geologic (SG) maps are essential for understanding surface processes and supporting infrastructure planning, but current workflows are labor-intensive and difficult to scale. We introduce EarthScape, an AI-ready multimodal dataset…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Matthew Massey , Nusrat Munia , Abdullah-Al-Zubaer Imran

Satellite imagery differs fundamentally from natural images: its aerial viewpoint, very high resolution, diverse scale variations, and abundance of small objects demand both region-level spatial reasoning and holistic scene understanding.…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Emanuel Sánchez Aimar , Gulnaz Zhambulova , Fahad Shahbaz Khan , Yonghao Xu , Michael Felsberg

How does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with the large gaps that often exist between sight and sound. We…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Andrew Owens , Tae-Hyun Oh

Worldwide visual geo-localization aims to determine the geographic location of an image anywhere on Earth using only its visual content. Despite recent progress, learning expressive representations of geographic space remains challenging…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Angel Daruna , Nicholas Meegan , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

Speech clarity and spatial audio immersion are the two most critical factors in enhancing remote conferencing experiences. Existing methods are often limited: either due to the lack of spatial information when using only one microphone, or…

声音 · 计算机科学 2025-07-14 Cheng Chi , Xiaoyu Li , Yuxuan Ke , Qunping Ni , Yao Ge , Xiaodong Li , Chengshi Zheng

Sound source localization (SSL) is the task of locating the source of sound within an image. Due to the lack of localization labels, the de facto standard in SSL has been to represent an image and audio as a single embedding vector each,…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Inho Kim , Youngkil Song , Jicheol Park , Won Hwa Kim , Suha Kwak

Accurate 3D volumetric mapping is critical for autonomous underwater vehicles operating in obstacle-rich environments. Vision-based perception provides high-resolution data but fails in turbid conditions, while sonar is robust to lighting…

机器人学 · 计算机科学 2026-03-17 Ivana Collado-Gonzalez , John McConnell , Brendan Englot

Worldwide image geo-localization aims to infer the geographic location of an image captured anywhere on Earth, spanning street, city, regional, national, and continental scales. Existing methods rely on visual features that are sensitive to…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Junchao Cui , Wenqi Shi , Shaoyong Du , Hang He , Xuanzi Ma , Hao Tang , Xiangyang Luo

Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR,…

We describe a novel metric-based learning approach that introduces a multimodal framework and uses deep audio and geophone encoders in siamese configuration to design an adaptable and lightweight supervised model. This framework eliminates…

声音 · 计算机科学 2021-11-16 Muhammad Shakeel , Katsutoshi Itoyama , Kenji Nishida , Kazuhiro Nakadai

Species distributions encode valuable ecological and environmental information, yet their potential for guiding representation learning in remote sensing remains underexplored. We introduce WildSAT, which pairs satellite images with…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Rangel Daroya , Elijah Cole , Oisin Mac Aodha , Grant Van Horn , Subhransu Maji

We present SOS-Match, a novel framework for detecting and matching objects in unstructured environments. Our system consists of 1) a front-end mapping pipeline using a zero-shot segmentation model to extract object masks from images and…

机器人学 · 计算机科学 2024-11-28 Annika Thomas , Jouko Kinnari , Parker Lusk , Kota Kondo , Jonathan P. How

The rapid development of generative audio raises ethical and security concerns stemming from forged data, making deepfake sound detection an important safeguard against the malicious use of such technologies. Although prior studies have…

声音 · 计算机科学 2025-09-29 Zeyu Xie , Yaoyun Zhang , Xuenan Xu , Yongkang Yin , Chenxing Li , Mengyue Wu , Yuexian Zou

We investigate an method for quantifying city characteristics based on impressions of a sound environment. The quantification of the city characteristics will be beneficial to government policy planning, tourism projects, etc. In this…

声音 · 计算机科学 2022-09-12 Yusuke Ono , Sunao Hara , Masanobu Abe

A large variety of geospatial data layers is available around the world ranging from remotely-sensed raster data like satellite imagery, digital elevation models, predicted land cover maps, and human-annotated data, to data derived from…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Arjun Rao , Esther Rolf

Visual events are usually accompanied by sounds in our daily lives. We pose the question: Can the machine learn the correspondence between visual scene and the sound, and localize the sound source only by observing sound and visual scene…

计算机视觉与模式识别 · 计算机科学 2019-02-18 Arda Senocak , Tae-Hyun Oh , Junsik Kim , Ming-Hsuan Yang , In So Kweon

Modeling and understanding the 3D world is crucial for various applications, from augmented reality to robotic navigation. Recent advancements based on 3D Gaussian Splatting have integrated semantic information from multi-view images into…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Xingrui Wang , Cuiling Lan , Hanxin Zhu , Zhibo Chen , Yan Lu

Image denoising is a fundamental problem in computer vision and medical imaging. However, real-world images are often degraded by structured noise with strong anisotropic correlations that existing methods struggle to remove. Most…

图像与视频处理 · 电气工程与系统科学 2025-10-03 Jianxu Wang , Ge Wang