English
Related papers

Related papers: Head Orientation Estimation with Distributed Micro…

200 papers

Accurate sound field reproduction in rooms is often limited by the lack of knowledge of the room characteristics. Information about the room shape or nearby reflecting boundaries can, in principle, be used to improve the accuracy of the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-04 Vincenzo Zaccà , Pablo Martinez-Nuevo , Martin Møller , Jorge Martínez , Richard Heusdens

The audio data is increasing day by day throughout the globe with the increase of telephonic conversations, video conferences and voice messages. This research provides a mechanism for identifying a speaker in an audio file, based on the…

Sound · Computer Science 2022-05-31 Syeda Rabia Arshad , Syed Mujtaba Haider , Abdul Basit Mughal

The objective of the present work is to propose a method to automatically detect polarity of the speech signals by estimating instants of significant excitation of the vocaltract and the cosine phase of the analytic signal representation.…

Sound · Computer Science 2014-07-18 D. Govind , Anju Susan Biju , Aguthu Smily

Speaker counting is the task of estimating the number of people that are simultaneously speaking in an audio recording. For several audio processing tasks such as speaker diarization, separation, localization and tracking, knowing the…

Sound · Computer Science 2021-01-07 Pierre-Amaury Grumiaux , Srdan Kitic , Laurent Girin , Alexandre Guérin

Speech evaluation measures a learners oral proficiency using automatic models. Corpora for training such models often pose sparsity challenges given that there often is limited scored data from teachers, in addition to the score…

Artificial Intelligence · Computer Science 2024-09-24 Huayun Zhang , Jeremy H. M. Wong , Geyu Lin , Nancy F. Chen

Conversations between a clinician and a patient, in natural conditions, are valuable sources of information for medical follow-up. The automatic analysis of these dialogues could help extract new language markers and speed-up the…

This paper presents a statistical modeling approach of the real-life user-induced randomness due to mobile phone orientations for different phone usage types. As well-known, the radiated performance of a wireless device depends on its…

Signal Processing · Electrical Eng. & Systems 2018-04-24 Andrés Alayón Glazunov , Per Hjalmar Lehne

Joint audio-visual speaker tracking requires that the locations of microphones and cameras are known and that they are given in a common coordinate system. Sensor self-localization algorithms, however, are usually separately developed for…

Sound · Computer Science 2015-04-14 Florian Jacob , Reinhold Haeb-Umbach

Formants are the spectral maxima that result from acoustic resonances of the human vocal tract, and their accurate estimation is among the most fundamental speech processing problems. Recent work has been shown that those frequencies can…

Sound · Computer Science 2022-06-24 Yosi Shrem , Felix Kreuk , Joseph Keshet

This paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end neural network manner, called directional automatic speech recognition (D-ASR), which explicitly models source speaker locations. In D-ASR, the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-03 Aswin Shanmugam Subramanian , Chao Weng , Shinji Watanabe , Meng Yu , Yong Xu , Shi-Xiong Zhang , Dong Yu

Sound-tracking refers to the process of determining the direction from which a sound originates, making it a fundamental component of sound source localization. This capability is essential in a variety of applications, including security…

Sound · Computer Science 2025-10-13 Mahdi Ali Pour , Zahra Habibzadeh

So far, several physical models have been proposed for the study of vocal fold oscillations during phonation. The parameters of these models, such as vocal fold elasticity, resistance, etc. are traditionally determined through the…

Sound · Computer Science 2020-02-13 Wenbo Zhao , Rita Singh

Self-supervised learning has been used to leverage unlabelled data, improving accuracy and generalisation of speech systems through the training of representation models. While many recent works have sought to produce effective…

Computation and Language · Computer Science 2023-10-18 Antoni Dimitriadis , Siqi Pan , Vidhyasaharan Sethu , Beena Ahmed

We propose BeamTransformer, an efficient architecture to leverage beamformer's edge in spatial filtering and transformer's capability in context sequence modeling. BeamTransformer seeks to optimize modeling of sequential relationship among…

Sound · Computer Science 2021-09-10 Siqi Zheng , Shiliang Zhang , Weilong Huang , Qian Chen , Hongbin Suo , Ming Lei , Jinwei Feng , Zhijie Yan

The estimation of room impulse responses (RIRs) between static loudspeaker and microphone locations can be done using a number of well-established measurement and inference procedures. While these procedures assume a time-invariant acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-14 Kathleen MacWilliam , Thomas Dietzen , Randall Ali , Toon van Waterschoot

In this paper, a machine learning method for predicting the evolution of a mobile communication channel based on a specific type of convolutional neural network is developed and evaluated in a simulated multipath transmission scenario. The…

Signal Processing · Electrical Eng. & Systems 2020-03-03 Julian Ahrens , Lia Ahrens , Hans D. Schotten

Self-supervised representations of speech are currently being widely used for a large number of applications. Recently, some efforts have been made in trying to analyze the type of information present in each of these representations. Most…

Sound · Computer Science 2023-09-22 Pablo Riera , Manuela Cerdeiro , Leonardo Pepino , Luciana Ferrer

Recently, stunning improvements on multi-channel speech separation have been achieved by neural beamformers when direction information is available. However, most of them neglect to utilize speaker's 2-dimensional (2D) location cues…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-05 Yanjie Fu , Meng Ge , Honglong Wang , Nan Li , Haoran Yin , Longbiao Wang , Gaoyan Zhang , Jianwu Dang , Chengyun Deng , Fei Wang

Speech-to-speech translation is a typical sequence-to-sequence learning task that naturally has two directions. How to effectively leverage bidirectional supervision signals to produce high-fidelity audio for both directions? Existing…

Computation and Language · Computer Science 2023-05-23 Xianchao Wu

Modern cars provide versatile tools to enhance speech communication. While an in-car communication (ICC) system aims at enhancing communication between the passengers by playing back desired speech via loudspeakers in the car, these…

Sound · Computer Science 2022-11-08 Kaspar Müller , Simon Doclo , Jan Østergaard , Tobias Wolff