中文
相关论文

相关论文: Neural Ambisonic Encoding For Multi-Speaker Scenar…

200 篇论文

Spatial audio reasoning enables machines to interpret auditory scenes by understanding events and their spatial attributes. In this work, we focus on spatial audio understanding with an emphasis on reasoning about moving sources. First, we…

声音 · 计算机科学 2025-09-19 Arvind Krishna Sridhar , Yinyi Guo , Erik Visser

The increasing popularity of spatial audio in applications such as teleconferencing, entertainment, and virtual reality has led to the recent developments of binaural reproduction methods. However, only a few of these methods are…

音频与语音处理 · 电气工程与系统科学 2025-02-17 Ami Berger , Vladimir Tourbabin , Jacob Donley , Zamir Ben-Hur , Boaz Rafaely

This paper introduces an area-based source separation method designed for virtual meeting scenarios. The aim is to preserve speech signals from an unspecified number of sources within a defined spatial area in front of a linear microphone…

音频与语音处理 · 电气工程与系统科学 2024-08-20 Martin Strauss , Okan Köpüklü

Reasoning about spatial audio with large language models requires a spatial audio encoder as an acoustic front-end to obtain audio embeddings for further processing. Such an encoder needs to capture all information required to detect the…

音频与语音处理 · 电气工程与系统科学 2025-11-04 Kevin Wilkinghoff , Zheng-Hua Tan

Conventional far-field automatic speech recognition (ASR) systems typically employ microphone array techniques for speech enhancement in order to improve robustness against noise or reverberation. However, such speech enhancement techniques…

音频与语音处理 · 电气工程与系统科学 2021-12-23 Minhua Wu , Kenichi Kumatani , Shiva Sundaram , Nikko Strom , Bjorn Hoffmeister

We introduce BANC, a neural binaural audio codec designed for efficient speech compression in single and two-speaker scenarios while preserving the spatial location information of each speaker. Our key contributions are as follows: 1) The…

声音 · 计算机科学 2024-11-26 Anton Ratnarajah , Shi-Xiong Zhang , Dong Yu

Existing deep learning-based speech denoising approaches require clean speech signals to be available for training. This paper presents a deep learning-based approach to improve speech denoising in real-world audio environments by not…

音频与语音处理 · 电气工程与系统科学 2020-02-25 Nasim Alamdari , Arian Azarang , Nasser Kehtarnavaz

We propose to use neural networks for simultaneous detection and localization of multiple sound sources in human-robot interaction. In contrast to conventional signal processing techniques, neural network-based sound source localization…

声音 · 计算机科学 2018-09-18 Weipeng He , Petr Motlicek , Jean-Marc Odobez

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

声音 · 计算机科学 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

In this contribution, we present a novel online approach to multichannel speech enhancement. The proposed method estimates the enhanced signal through a filter-and-sum framework. More specifically, complex-valued masks are estimated by a…

音频与语音处理 · 电气工程与系统科学 2021-08-09 Mhd Modar Halimeh , Walter Kellermann

This paper presents an alternate representation framework to commonly used time-frequency representation for acoustic scene classification (ASC). A raw audio signal is represented using a pre-trained convolutional neural network (CNN) using…

音频与语音处理 · 电气工程与系统科学 2022-04-04 Arshdeep Singh

The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as…

声音 · 计算机科学 2022-11-04 You Jin Kim , Hee-Soo Heo , Jee-weon Jung , Youngki Kwon , Bong-Jin Lee , Joon Son Chung

Binaural audio provides human listeners with an immersive spatial sound experience, but most existing videos lack binaural audio recordings. We propose an audio spatialization method that draws on visual information in videos to convert…

计算机视觉与模式识别 · 计算机科学 2021-11-23 Rishabh Garg , Ruohan Gao , Kristen Grauman

Differential microphone arrays offer a promising solution for far-field acoustic signal acquisition due to their high spatial directivity and compact array structure. A key challenge lies in designing differential beamformers that are…

音频与语音处理 · 电气工程与系统科学 2026-02-27 Tiantian Xiong , Yongyi Deng , Kunlong Zhao , Jilu Jin , Xueqin Luo , Gongping Huang , Jingdong Chen , Jacob Benesty

We consider a semantic communication system for speech signals, named DeepSC-S. Motivated by the breakthroughs in deep learning (DL), we make an effort to recover the transmitted speech signals in the semantic communication systems, which…

音频与语音处理 · 电气工程与系统科学 2021-09-09 Zhenzi Weng , Zhijin Qin , Geoffrey Ye Li

Audio self-supervised learning (SSL) aims to learn general-purpose representations from large-scale unlabeled audio data. While recent advances have been driven mainly by generative reconstruction objectives, contrastive approaches remain…

机器学习 · 计算机科学 2026-05-15 Hanxun Huang , Qizhou Wang , Xingjun Ma , Cihang Xie , Christopher Leckie , Sarah Erfani

In this paper, we propose a deep learning based multi-speaker direction of arrival (DOA) estimation with audio and visual signals by using permutation-free loss function. We first collect a data set for multi-modal sound source localization…

音频与语音处理 · 电气工程与系统科学 2022-10-27 Qing Wang , Hang Chen , Ya Jiang , Zhe Wang , Yuyang Wang , Jun Du , Chin-Hui Lee

Traditionally, research in automated speech recognition has focused on local-first encoding of audio representations to predict the spoken phonemes in an utterance. Unfortunately, approaches relying on such hyper-local information tend to…

音频与语音处理 · 电气工程与系统科学 2022-09-19 David M. Chan , Shalini Ghosh , Debmalya Chakrabarty , Björn Hoffmeister

Time-division multiplexed (TDM) channel sounders, in which a single RF chain is connected sequentially via an electronic switch to different elements of an array, are widely used for the measurement of double-directional/MIMO propagation…

信息论 · 计算机科学 2018-05-22 Rui Wang , Olivier Renaudin , C. Umit Bas , Seun Sangodoyin , Andreas F. Molisch

Millimeter-wave (mmWave) massive Multiple Input Multiple Output (MIMO) systems encounter both spatial wideband spreading and temporal wideband effects in the communication channels of individual users. Accurate estimation of a user's…

信号处理 · 电气工程与系统科学 2025-04-22 Chandrashekhar Rai , Debarati Sen