中文
相关论文

相关论文: Learning Spatially-Aware Language and Audio Embedd…

200 篇论文

Contrastive language--audio pretraining (CLAP) has achieved remarkable success as an audio--text embedding framework, but existing approaches are limited to monaural or single-source conditions and cannot fully capture spatial information.…

Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic information between the…

计算机视觉与模式识别 · 计算机科学 2020-06-15 Karren Yang , Bryan Russell , Justin Salamon

Spatial audio reasoning enables machines to interpret auditory scenes by understanding events and their spatial attributes. In this work, we focus on spatial audio understanding with an emphasis on reasoning about moving sources. First, we…

声音 · 计算机科学 2025-09-19 Arvind Krishna Sridhar , Yinyi Guo , Erik Visser

Learning from audio-visual data offers many possibilities to express correspondence between the audio and visual content, similar to the human perception that relates aural and visual information. In this work, we present a method for…

音频与语音处理 · 电气工程与系统科学 2022-11-23 Shanshan Wang , Archontis Politis , Annamaria Mesaros , Tuomas Virtanen

This paper explores enabling large language models (LLMs) to understand spatial information from multichannel audio, a skill currently lacking in auditory LLMs. By leveraging LLMs' advanced cognitive and inferential abilities, the aim is to…

声音 · 计算机科学 2024-06-17 Changli Tang , Wenyi Yu , Guangzhi Sun , Xianzhao Chen , Tian Tan , Wei Li , Jun Zhang , Lu Lu , Zejun Ma , Yuxuan Wang , Chao Zhang

Spatial audio understanding is essential for accurately perceiving and interpreting acoustic environments. However, existing audio-language models exhibit limitations in processing spatial audio and perceiving spatial acoustic scenes. To…

声音 · 计算机科学 2025-09-19 Jinbo Hu , Yin Cao , Ming Wu , Zhenbo Luo , Jun Yang

The acoustic cues used by humans and other animals to localise sounds are subtle, and change during and after development. This means that we need to constantly relearn or recalibrate the auditory spatial map throughout our lifetimes. This…

神经与进化计算 · 计算机科学 2025-04-18 Yang Chu , Wayne Luk , Dan Goodman

Automatic speech recognition (ASR) has benefited from advances in pretrained speech and language models, yet most systems remain constrained to monolingual settings and short, isolated utterances. While recent efforts in context-aware ASR…

计算与语言 · 计算机科学 2026-03-09 Yuchen Zhang , Haralambos Mouratidis , Ravi Shekhar

Audio Large Language Models (Audio LLMs) have demonstrated strong capabilities in integrating speech perception with language understanding. However, whether their internal representations align with human neural dynamics during…

声音 · 计算机科学 2026-02-04 Haoyun Yang , Xin Xiao , Jiang Zhong , Yu Tian , Dong Xiaohua , Yu Mao , Hao Wu , Kaiwen Wei

Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene…

音频与语音处理 · 电气工程与系统科学 2025-05-20 Zhisheng Zheng , Puyuan Peng , Ziyang Ma , Xie Chen , Eunsol Choi , David Harwath

Our goal is to develop a sound event localization and detection (SELD) system that works robustly in unknown environments. A SELD system trained on known environment data is degraded in an unknown environment due to environmental effects…

音频与语音处理 · 电气工程与系统科学 2022-02-21 Masahiro Yasuda , Yasunori Ohishi , Shoichiro Saito

Visual learning often occurs in a specific context, where an agent acquires skills through exploration and tracking of its location in a consistent environment. The historical spatial context of the agent provides a similarity signal for…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Lizhen Zhu , James Z. Wang , Wonseuk Lee , Brad Wyble

In recent years supervised representation learning has provided state of the art or close to the state of the art results in semantic analysis tasks including ranking and information retrieval. The core idea is to learn how to embed items…

计算与语言 · 计算机科学 2017-08-11 Dasha Bogdanova , Majid Yazdani

Integrating spatial context into large language models (LLMs) has the potential to revolutionize human-computer interaction, particularly in wearable devices. In this work, we present a novel system architecture that incorporates spatial…

声音 · 计算机科学 2025-04-28 Ayushi Mishra , Yang Bai , Priyadarshan Narayanasamy , Nakul Garg , Nirupam Roy

Large pre-trained speech models such as Whisper offer strong generalization but pose significant challenges for resource-efficient adaptation. Low-Rank Adaptation (LoRA) has become a popular parameter-efficient fine-tuning method, yet its…

声音 · 计算机科学 2026-01-23 Yujian Ma , Xikun Lu , Jinqiu Sang , Xianquan Jiang , Ruizhe Li

Sound event localization and detection (SELD) consists of two subtasks, which are sound event detection and direction-of-arrival estimation. While sound event detection mainly relies on time-frequency patterns to distinguish different sound…

音频与语音处理 · 电气工程与系统科学 2022-06-07 Thi Ngoc Tho Nguyen , Karn N. Watcharasupat , Ngoc Khanh Nguyen , Douglas L. Jones , Woon-Seng Gan

Machine hearing of the environmental sound is one of the important issues in the audio recognition domain. It gives the machine the ability to discriminate between the different input sounds that guides its decision making. In this work we…

声音 · 计算机科学 2022-07-20 Peter Ochieng , Dennis Kaburu

General accent recognition (AR) models tend to directly extract low-level information from spectrums, which always significantly overfit on speakers or channels. Considering accent can be regarded as a series of shifts relative to native…

声音 · 计算机科学 2022-07-04 Qijie Shao , Jinghao Yan , Jian Kang , Pengcheng Guo , Xian Shi , Pengfei Hu , Lei Xie

Auditory scene analysis (ASA) aims to retrieve information from the acoustic environment, by carrying out three main tasks: sound source location, separation, and classification. These tasks are traditionally executed with a linear data…

音频与语音处理 · 电气工程与系统科学 2025-08-21 Caleb Rascon , Luis Gato-Diaz , Eduardo García-Alarcón

This report presents our systems submitted to the audio-only and audio-visual tracks of the DCASE2025 Task 3 Challenge: Stereo Sound Event Localization and Detection (SELD) in Regular Video Content. SELD is a complex task that combines…

音频与语音处理 · 电气工程与系统科学 2025-07-08 Davide Berghi , Philip J. B. Jackson
‹ 上一页 1 2 3 10 下一页 ›