English
Related papers

Related papers: Simultaneously Learning Speaker's Direction and He…

200 papers

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper, we tackle…

Recent studies have demonstrated that the representations of artificial neural networks (ANNs) can exhibit notable similarities to cortical representations when subjected to identical auditory sensory inputs. In these studies, the ability…

Neurons and Cognition · Quantitative Biology 2024-12-23 Taketo Akama , Zhuohao Zhang , Pengcheng Li , Kotaro Hongo , Hiroaki Kitano , Shun Minamikawa , Natalia Polouliakh

This contribution gives an overview of face recogni-tion algorithms, their implementation and practical uses. First, a training set of different persons' faces has to be collected and used to train a face recognizer. The resulting face…

Computer Vision and Pattern Recognition · Computer Science 2017-07-05 Johannes Reschke , Armin Sehr

Due to the high performance of multi-channel speech processing, we can use the outputs from a multi-channel model as teacher labels when training a single-channel model with knowledge distillation. To the contrary, it is also known that…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-10 Shota Horiguchi , Yuki Takashima , Shinji Watanabe , Paola Garcia

A new database of head-related transfer functions (HRTFs) for accurate sound source localization is presented through precise measurement and post-processing in terms of improved frequency bandwidth and causality of head-related impulse…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-07 Gyeong-Tae Lee , Sang-Min Choi , Byeong-Yun Ko , Yong-Hwa Park

Natural language processing (NLP) models trained on people-generated data can be unreliable because, without any constraints, they can learn from spurious correlations that are not relevant to the task. We hypothesize that enriching models…

Computation and Language · Computer Science 2022-03-18 Alissa Ostapenko , Shuly Wintner , Melinda Fricke , Yulia Tsvetkov

Sounds reach one microphone in a stereo pair sooner than the other, resulting in an interaural time delay that conveys their directions. Estimating a sound's time delay requires finding correspondences between the signals recorded by each…

Computer Vision and Pattern Recognition · Computer Science 2023-01-31 Ziyang Chen , David F. Fouhey , Andrew Owens

Personalized Head-Related Transfer Functions (HRTFs) are starting to be introduced in many commercial immersive audio applications and are crucial for realistic spatial audio rendering. However, one of the main hesitations regarding their…

Sound · Computer Science 2025-10-03 Xuyi Hu , Jian Li , Shaojie Zhang , Stefan Goetz , Lorenzo Picinali , Ozgur B. Akan , Aidan O. T. Hogg

Methods for extracting audio and speech features have been studied since pioneering work on spectrum analysis decades ago. Recent efforts are guided by the ambition to develop general-purpose audio representations. For example, deep neural…

We propose BeamTransformer, an efficient architecture to leverage beamformer's edge in spatial filtering and transformer's capability in context sequence modeling. BeamTransformer seeks to optimize modeling of sequential relationship among…

Sound · Computer Science 2021-09-10 Siqi Zheng , Shiliang Zhang , Weilong Huang , Qian Chen , Hongbin Suo , Ming Lei , Jinwei Feng , Zhijie Yan

Individualized head-related transfer functions (HRTFs) are crucial for accurate sound positioning in virtual auditory displays. As the acoustic measurement of HRTFs is resource-intensive, predicting individualized HRTFs using machine…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-28 Yutong Wen , You Zhang , Zhiyao Duan

Voice conversion (VC) using sequence-to-sequence learning of context posterior probabilities is proposed. Conventional VC using shared context posterior probabilities predicts target speech parameters from the context posterior…

Sound · Computer Science 2017-08-08 Hiroyuki Miyoshi , Yuki Saito , Shinnosuke Takamichi , Hiroshi Saruwatari

Speech production is a complex sequential process which involve the coordination of various articulatory features. Among them tongue being a highly versatile active articulator responsible for shaping airflow to produce targeted speech…

Sound · Computer Science 2025-04-28 Leena G Pillai , D. Muhammad Noorul Mubarak , Elizabeth Sherly

Hierarchical models are utilized in a wide variety of problems which are characterized by task hierarchies, where predictions on smaller subtasks are useful for trying to predict a final task. Typically, neural networks are first trained…

We explore the interpretation of sound for robot decision making, inspired by human speech comprehension. While previous methods separate sound processing unit and robot controller, we propose an end-to-end deep neural network which…

Robotics · Computer Science 2020-09-17 Peixin Chang , Shuijing Liu , Haonan Chen , Katherine Driggs-Campbell

The objective of Audio Augmented Reality (AAR) applications are to seamlessly integrate virtual sound sources within a real environment. It is critical for these applications that virtual sources are localised precisely at the intended…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-13 Vincent Martin , Lorenzo Picinali

We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to emphasize or suppress…

Speech perception involves storing and integrating sequentially presented items. Recent work in cognitive neuroscience has identified temporal and contextual characteristics in humans' neural encoding of speech that may facilitate this…

Computation and Language · Computer Science 2024-05-15 Oli Danyi Liu , Hao Tang , Naomi Feldman , Sharon Goldwater

This paper investigates the joint localization, detection, and tracking of sound events using a convolutional recurrent neural network (CRNN). We use a CRNN previously proposed for the localization and detection of stationary sources, and…

Sound · Computer Science 2019-04-30 Sharath Adavanne , Archontis Politis , Tuomas Virtanen

Automatic speech transcription and speaker recognition are usually treated as separate tasks even though they are interdependent. In this study, we investigate training a single network to perform both tasks jointly. We train the network in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-21 Siddharth Sigtia , Erik Marchi , Sachin Kajarekar , Devang Naik , John Bridle