English
Related papers

Related papers: Learning Tri-modal Embeddings for Zero-Shot Sounds…

200 papers

Advancements in audio neural networks have established state-of-the-art results on downstream audio tasks. However, the black-box structure of these models makes it difficult to interpret the information encoded in their internal audio…

Sound · Computer Science 2025-04-22 Alice Zhang , Edison Thomaz , Lie Lu

Recent advances have been witnessed in audio-language joint learning, such as CLAP, that shows much success in multi-modal understanding tasks. These models usually aggregate uni-modal local representations, namely frame or word features,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-16 Yiming Li , Zhifang Guo , Xiangdong Wang , Hong Liu

This paper focuses on perceiving and navigating 3D environments using echoes and RGB image. In particular, we perform depth estimation by fusing RGB image with echoes, received from multiple orientations. Unlike previous works, we go beyond…

Computer Vision and Pattern Recognition · Computer Science 2024-02-12 Lingyu Zhu , Esa Rahtu , Hang Zhao

Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on…

There exists a correlation between geospatial activity temporal patterns and type of land use. A novel self-supervised approach is proposed to stratify landscape based on mobility activity time series. First, the time series signal is…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Yi Cao , Swetava Ganguli , Vipul Pandey

In computer vision and machine learning for geographic data, out-of-domain generalization is a pervasive challenge, arising from uneven global data coverage and distribution shifts across geographic regions. Though models are frequently…

Machine Learning · Computer Science 2026-04-20 Haoran Zhang , Livia Betti , Konstantin Klemmer , Esther Rolf , David Alvarez-Melis

Modern cameras are equipped with a wide array of sensors that enable recording the geospatial context of an image. Taking advantage of this, we explore depth estimation under the assumption that the camera is geocalibrated, a problem we…

Computer Vision and Pattern Recognition · Computer Science 2021-09-22 Scott Workman , Hunter Blanton

Text-to-audio (TTA) systems have recently demonstrated strong performance in synthesizing monaural audio from text. However, the task of generating binaural spatial audio from text, which provides a more immersive auditory experience by…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-18 Linfeng Feng , Lei Zhao , Boyu Zhu , Xiao-Lei Zhang , Xuelong Li

Several animal species (e.g., bats, dolphins, and whales) and even visually impaired humans have the remarkable ability to perform echolocation: a biological sonar used to perceive spatial layout and locate objects in the world. We explore…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Ruohan Gao , Changan Chen , Ziad Al-Halah , Carl Schissler , Kristen Grauman

We address the problem of estimating depth with multi modal audio visual data. Inspired by the ability of animals, such as bats and dolphins, to infer distance of objects with echolocation, some recent methods have utilized echoes for depth…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Kranti Kumar Parida , Siddharth Srivastava , Gaurav Sharma

We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-04 Dyah A. M. G. Wisnu , Ryandhimas E. Zezario , Stefano Rini , Hsin-Min Wang , Yu Tsao

Worldwide geo-localization involves determining the exact geographic location of images captured globally, typically guided by geographic cues such as climate, landmarks, and architectural styles. Despite advancements in geo-localization…

Computer Vision and Pattern Recognition · Computer Science 2025-09-08 Furong Jia , Lanxin Liu , Ce Hou , Fan Zhang , Xinyan Liu , Yu Liu

Understanding the geometric and semantic structure of environments is essential for embodied navigation and reasoning. Existing semantic mapping methods trade off between explicit geometry and multi-scale semantics, and lack a native…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Sixian Zhang , Yiyao Wang , Xinhang Song , Keming Zhang , Zijian Xu , Shuqiang Jiang

We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from…

Classical methods for acoustic scene mapping require the estimation of time difference of arrival (TDOA) between microphones. Unfortunately, TDOA estimation is very sensitive to reverberation and additive noise. We introduce an unsupervised…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-14 Idan Cohen , Ofir Lindenbaum , Sharon Gannot

A crucial ability of mobile intelligent agents is to integrate the evidence from multiple sensory inputs in an environment and to make a sequence of actions to reach their goals. In this paper, we attempt to approach the problem of…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Chuang Gan , Yiwei Zhang , Jiajun Wu , Boqing Gong , Joshua B. Tenenbaum

A soundscape is composed of three types of sound: biophony (sounds made by animals), geophony (natural abiotic sounds) and anthropophony (sounds made by humans). A key research question in the field of soundscape ecology is how these…

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronization. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment…

Computer Vision and Pattern Recognition · Computer Science 2020-11-05 Soo-Whan Chung , Joon Son Chung , Hong-Goo Kang

Environmental sound understanding in computational auditory scene analysis (CASA) is often formulated as an audio-only recognition problem. This formulation leaves a persistent drawback in multi-label audio tagging (AT): acoustic similarity…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-12 Yuanbo Hou , Yanru Wu , Qiaoqiao Ren , Shengchen Li , Stephen Roberts , Dick Botteldooren

Place recognition using SOund Navigation and Ranging (SONAR) images is an important task for simultaneous localization and mapping(SLAM) in underwater environments. This paper proposes a robust and efficient imaging SONAR based place…

Robotics · Computer Science 2024-03-12 Hogyun Kim , Gilhwan Kang , Seokhwan Jeong , Seungjun Ma , Younggun Cho
‹ Prev 1 3 4 5 6 7 10 Next ›