中文
相关论文

相关论文: Designing a 9-channel location microphone from scr…

200 篇论文

This paper proposes sound event localization and detection methods from multichannel recording. The proposed system is based on two Convolutional Recurrent Neural Networks (CRNNs) to perform sound event detection (SED) and time difference…

音频与语音处理 · 电气工程与系统科学 2019-10-23 Francois Grondin , James Glass , Iwona Sobieraj , Mark D. Plumbley

In this paper we provide three contributions to the field of channel sounding waveform design in asynchronous Multi-user (MU) MIMO systems. The first contribution is a derivation of the asynchronous MU-MIMO model and the conditions that the…

信息论 · 计算机科学 2013-02-20 Zhenhua Yu , Robert J. Baxley , Brett T. Walkenhorst , G. Tong Zhou

While 3D human body modeling has received much attention in computer vision, modeling the acoustic equivalent, i.e. modeling 3D spatial audio produced by body motion and speech, has fallen short in the community. To close this gap, we…

计算机视觉与模式识别 · 计算机科学 2023-11-14 Xudong Xu , Dejan Markovic , Jacob Sandakly , Todd Keebler , Steven Krenn , Alexander Richard

Sound-tracking refers to the process of determining the direction from which a sound originates, making it a fundamental component of sound source localization. This capability is essential in a variety of applications, including security…

声音 · 计算机科学 2025-10-13 Mahdi Ali Pour , Zahra Habibzadeh

Binaural audio gives the listener the feeling of being in the recording place and enhances the immersive experience if coupled with AR/VR. But the problem with binaural audio recording is that it requires a specialized setup which is not…

声音 · 计算机科学 2021-08-12 Kranti Kumar Parida , Siddharth Srivastava , Neeraj Matiyali , Gaurav Sharma

We present a new method to capture the acoustic characteristics of real-world rooms using commodity devices, and use the captured characteristics to generate similar sounding sources with virtual models. Given the captured audio and an…

声音 · 计算机科学 2021-09-28 Zhenyu Tang , Nicholas J. Bryan , Dingzeyu Li , Timothy R. Langlois , Dinesh Manocha

This contribution introduces a dataset of 7th-order Ambisonic Room Impulse Responses (HOA-RIRs), created using the Image Source Method. By employing higher-order Ambisonics, our dataset enables precise spatial audio reproduction, a critical…

声音 · 计算机科学 2025-06-02 Shivam Saini , Jürgen Peissig

Acousto-optic sensing provides an alternative approach to traditional microphone arrays by shedding light on the interaction of light with an acoustic field. Sound field reconstruction is a fascinating and advanced technique used in…

声音 · 计算机科学 2025-07-01 Phuc Duc Nguyen , Kenji Ishikawa , Noboru Harada , Takehiro Moriya

Consider a microphone array, such as those present in Amazon Echos, conference phones, or self-driving cars. One of the goals of these arrays is to decode the angles in which acoustic signals arrive at them. This paper considers the problem…

声音 · 计算机科学 2021-09-28 Yu-Lin Wei , Romit Roy Choudhury

In this paper, two microphone based systems for audio zooming is proposed for the first time. The audio zooming application allows sound capture and enhancement from the front direction while attenuating interfering sources from all other…

音频与语音处理 · 电气工程与系统科学 2020-01-15 Anant Khandelwal , E. B. Goud , Y. Chand , L. Kumar , S. Prasad , N. Agarwala , R. Singh

In this paper, we present a novel interdisciplinary approach to study the relationship between diffusive surface structures and their acoustic performance. Using computational design, surface structures are iteratively generated and 3D…

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

声音 · 计算机科学 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

Room acoustics measurements are used in many areas of audio research, from physical acoustics modelling and speech enhancement to virtual reality applications. This paper documents the technical specifications and choices made in the…

音频与语音处理 · 电气工程与系统科学 2021-11-24 Thomas McKenzie , Leo McCormack , Christoph Hold

This paper considers the problem of simultaneous 2-D room shape reconstruction and self-localization without the requirement of any pre-established infrastructure. A mobile device equipped with co-located microphone and loudspeaker as well…

机器人学 · 计算机科学 2016-12-20 Tiexing Wang , Fangrong Peng , Biao Chen

Wireless systems are getting deployed in many new environments with different antenna heights, frequency bands and multipath conditions. This has led to an increasing demand for more channel measurements to understand wireless propagation…

信息论 · 计算机科学 2016-11-18 Muhammad Nazmul Islam , Byoung-Jo J. Kim , Paul Henry , Eric Rozner

Speaker diarization is the task of answering Who spoke and when? in an audio stream. Pipeline systems rely on speech segmentation to extract speakers' segments and achieve robust speaker diarization. This paper proposes a common framework…

声音 · 计算机科学 2023-06-08 Théo Mariotte , Anthony Larcher , Silvio Montrésor , Jean-Hugh Thomas

Automatic speech recognition (ASR) of multi-channel multi-speaker overlapped speech remains one of the most challenging tasks to the speech community. In this paper, we look into this challenge by utilizing the location information of…

声音 · 计算机科学 2021-11-23 Yiwen Shao , Shi-Xiong Zhang , Dong Yu

Indoor localization is a long-standing challenge in mobile computing, with significant implications for enabling location-aware and intelligent applications within smart environments such as homes, offices, and retail spaces. As AI…

音频与语音处理 · 电气工程与系统科学 2025-08-26 Amod K. Agrawal

Angle of arrival (AOA) is widely used to locate a wireless signal emitter in unmanned aerial vehicle (UAV) localization. Compared with received signal strength (RSS) and time of arrival (TOA), it has higher accuracy and is not sensitive to…

信号处理 · 电气工程与系统科学 2023-05-17 Baihua Shi , Yifan Li , Guilu Wu , Shihao Yan , Feng Shu

This paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end neural network manner, called directional automatic speech recognition (D-ASR), which explicitly models source speaker locations. In D-ASR, the…

音频与语音处理 · 电气工程与系统科学 2020-11-03 Aswin Shanmugam Subramanian , Chao Weng , Shinji Watanabe , Meng Yu , Yong Xu , Shi-Xiong Zhang , Dong Yu