中文
相关论文

相关论文: Learning Tri-modal Embeddings for Zero-Shot Sounds…

200 篇论文

Advancements in audio neural networks have established state-of-the-art results on downstream audio tasks. However, the black-box structure of these models makes it difficult to interpret the information encoded in their internal audio…

声音 · 计算机科学 2025-04-22 Alice Zhang , Edison Thomaz , Lie Lu

Recent advances have been witnessed in audio-language joint learning, such as CLAP, that shows much success in multi-modal understanding tasks. These models usually aggregate uni-modal local representations, namely frame or word features,…

音频与语音处理 · 电气工程与系统科学 2024-08-16 Yiming Li , Zhifang Guo , Xiangdong Wang , Hong Liu

This paper focuses on perceiving and navigating 3D environments using echoes and RGB image. In particular, we perform depth estimation by fusing RGB image with echoes, received from multiple orientations. Unlike previous works, we go beyond…

计算机视觉与模式识别 · 计算机科学 2024-02-12 Lingyu Zhu , Esa Rahtu , Hang Zhao

Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on…

There exists a correlation between geospatial activity temporal patterns and type of land use. A novel self-supervised approach is proposed to stratify landscape based on mobility activity time series. First, the time series signal is…

计算机视觉与模式识别 · 计算机科学 2024-01-18 Yi Cao , Swetava Ganguli , Vipul Pandey

In computer vision and machine learning for geographic data, out-of-domain generalization is a pervasive challenge, arising from uneven global data coverage and distribution shifts across geographic regions. Though models are frequently…

机器学习 · 计算机科学 2026-04-20 Haoran Zhang , Livia Betti , Konstantin Klemmer , Esther Rolf , David Alvarez-Melis

Modern cameras are equipped with a wide array of sensors that enable recording the geospatial context of an image. Taking advantage of this, we explore depth estimation under the assumption that the camera is geocalibrated, a problem we…

计算机视觉与模式识别 · 计算机科学 2021-09-22 Scott Workman , Hunter Blanton

Text-to-audio (TTA) systems have recently demonstrated strong performance in synthesizing monaural audio from text. However, the task of generating binaural spatial audio from text, which provides a more immersive auditory experience by…

音频与语音处理 · 电气工程与系统科学 2025-02-18 Linfeng Feng , Lei Zhao , Boyu Zhu , Xiao-Lei Zhang , Xuelong Li

Several animal species (e.g., bats, dolphins, and whales) and even visually impaired humans have the remarkable ability to perform echolocation: a biological sonar used to perceive spatial layout and locate objects in the world. We explore…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Ruohan Gao , Changan Chen , Ziad Al-Halah , Carl Schissler , Kristen Grauman

We address the problem of estimating depth with multi modal audio visual data. Inspired by the ability of animals, such as bats and dolphins, to infer distance of objects with echolocation, some recent methods have utilized echoes for depth…

计算机视觉与模式识别 · 计算机科学 2021-04-06 Kranti Kumar Parida , Siddharth Srivastava , Gaurav Sharma

We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production…

音频与语音处理 · 电气工程与系统科学 2025-09-04 Dyah A. M. G. Wisnu , Ryandhimas E. Zezario , Stefano Rini , Hsin-Min Wang , Yu Tsao

Worldwide geo-localization involves determining the exact geographic location of images captured globally, typically guided by geographic cues such as climate, landmarks, and architectural styles. Despite advancements in geo-localization…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Furong Jia , Lanxin Liu , Ce Hou , Fan Zhang , Xinyan Liu , Yu Liu

Understanding the geometric and semantic structure of environments is essential for embodied navigation and reasoning. Existing semantic mapping methods trade off between explicit geometry and multi-scale semantics, and lack a native…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Sixian Zhang , Yiyao Wang , Xinhang Song , Keming Zhang , Zijian Xu , Shuqiang Jiang

We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from…

Classical methods for acoustic scene mapping require the estimation of time difference of arrival (TDOA) between microphones. Unfortunately, TDOA estimation is very sensitive to reverberation and additive noise. We introduce an unsupervised…

音频与语音处理 · 电气工程与系统科学 2024-03-14 Idan Cohen , Ofir Lindenbaum , Sharon Gannot

A crucial ability of mobile intelligent agents is to integrate the evidence from multiple sensory inputs in an environment and to make a sequence of actions to reach their goals. In this paper, we attempt to approach the problem of…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Chuang Gan , Yiwei Zhang , Jiajun Wu , Boqing Gong , Joshua B. Tenenbaum

A soundscape is composed of three types of sound: biophony (sounds made by animals), geophony (natural abiotic sounds) and anthropophony (sounds made by humans). A key research question in the field of soundscape ecology is how these…

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronization. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment…

计算机视觉与模式识别 · 计算机科学 2020-11-05 Soo-Whan Chung , Joon Son Chung , Hong-Goo Kang

Environmental sound understanding in computational auditory scene analysis (CASA) is often formulated as an audio-only recognition problem. This formulation leaves a persistent drawback in multi-label audio tagging (AT): acoustic similarity…

音频与语音处理 · 电气工程与系统科学 2026-03-12 Yuanbo Hou , Yanru Wu , Qiaoqiao Ren , Shengchen Li , Stephen Roberts , Dick Botteldooren

Place recognition using SOund Navigation and Ranging (SONAR) images is an important task for simultaneous localization and mapping(SLAM) in underwater environments. This paper proposes a robust and efficient imaging SONAR based place…

机器人学 · 计算机科学 2024-03-12 Hogyun Kim , Gilhwan Kang , Seokhwan Jeong , Seungjun Ma , Younggun Cho