English
Related papers

Related papers: Modified Group Delay Based MultiPitch Estimation i…

200 papers

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

Sound · Computer Science 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee

This paper proposes a new method for calculating joint-state posteriors of mixed-audio features using deep neural networks to be used in factorial speech processing models. The joint-state posterior information is required in factorial…

Sound · Computer Science 2017-07-11 Mahdi Khademian , Mohammad Mehdi Homayounpour

A novel approach for speech segmentation is proposed, based on Multilevel Hybrid (mean/min) Filters (MHF) with the following features: An accurate transition location. Good performance in noisy environments (gaussian and impulsive noise).…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-04 Marcos Faundez-Zanuy , Francesc Vallverdu-Bayes

The presence of rich scattering in indoor and urban radio propagation scenarios may cause a high arrival density of multipath components (MPCs). Often the MPCs arrive in clusters at the receiver, where MPCs within one cluster have similar…

Signal Processing · Electrical Eng. & Systems 2021-02-02 Tarik Kazaz , Jac Romme , Gerard J. M. Janssen , Alle-Jan van der Veen

In real-time listening enhancement applications, such as hearing aid signal processing, sounds must be processed with no more than a few milliseconds of delay to sound natural to the listener. Listening devices can achieve better…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-12 Ryan M. Corey , Naoki Tsuda , Andrew C. Singer

Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-17 Shota Horiguchi , Takanori Ashihara , Marc Delcroix , Atsushi Ando , Naohiro Tawara

In this paper, we investigate a deep learning approach for speech denoising through an efficient ensemble of specialist neural networks. By splitting up the speech denoising task into non-overlapping subproblems and introducing a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Aswin Sivaraman , Minje Kim

Recently, some single-step systems without onset detection have shown their effectiveness in automatic musical tempo estimation. Following the success of these systems, in this paper we propose a Multi-scale Grouped Attention Network to…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-06 Xiaoheng Sun , Qiqi He , Yongwei Gao , Wei Li

This work considers uplink asynchronous massive machine-type communications, where a large number of low-power and low-cost devices asynchronously transmit short packets to an access point equipped with multiple receive antennas. If…

Information Theory · Computer Science 2025-07-22 Z. Shao , X. Yuan , R. de Lamare

A filter for universal real-time prediction of band-limited signals is presented. The filter consists of multiple time-delayed feedback terms in order to accomplish anticipatory coupling, which again leads to a negative group delay for…

Sound · Computer Science 2017-11-29 Henning U. Voss

Much of the engineering behind current wireless systems has focused on designing an efficient and high-throughput downlink to support human-centric communication such as video streaming and internet browsing. This paper looks ahead to…

Signal Processing · Electrical Eng. & Systems 2024-12-06 Sandesh Rao Mattu , Imran Ali Khan , Venkatesh Khammammetti , Beyza Dabak , Saif Khan Mohammed , Krishna Narayanan , Robert Calderbank

Reverberation can severely degrade the quality of speech signals recorded using microphones in an enclosure. In acoustic sensor networks with spatially distributed microphones, a similar dereverberation performance may be achieved using…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-04 Anselm Lohmann , Toon van Waterschoot , Joerg Bitzer , Simon Doclo

In the task of speaker diarization, the number of small-scale meetings accounts for a large proportion. When microphone arrays are employed as a recording device, its spatial information is usually ignored by most researchers. In this…

Sound · Computer Science 2022-10-27 Yuxuan Du , Ruohua Zhou

This paper addresses the estimation of fractional delay and Doppler shifts in multipath channels that cause doubly selective fading-an essential task for integrated sensing and communication (ISAC) systems in high-mobility environments.…

Signal Processing · Electrical Eng. & Systems 2025-06-24 Yutaka Jitsumatsu , Liangchen Sun

We consider channel estimation within pulse-shaping multicarrier multiple-input multiple-output (MIMO) systems transmitting over doubly selective MIMO channels. This setup includes MIMO orthogonal frequency-division multiplexing (MIMO-OFDM)…

Information Theory · Computer Science 2016-08-03 Daniel Eiwen , Georg Tauboeck , Franz Hlawatsch , Hans Georg Feichtinger

While log-amplitude mel-spectrogram has widely been used as the feature representation for processing speech based on deep learning, the effectiveness of another aspect of speech spectrum, i.e., phase information, was shown recently for…

Sound · Computer Science 2022-05-02 Shunsuke Hidaka , Kohei Wakamiya , Tokihiko Kaburagi

We introduce a system capable of faithfully modifying the perceptual voice quality of creak while preserving the speaker's perceived identity. While it is well known that high creak probability is typically correlated with low pitch, it is…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-17 Frederik Rautenberg , Jana Wiechmann , Petra Wagner , Reinhold Haeb-Umbach

This study addresses the task of performing robust and reliable time-delay estimation in signals in noisy and reverberating environments. In contrast to the popular signal processing based methods, this paper proposes to transform the input…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Akshay Raina , Vipul Arora

Speech Segmentation is the process change point detection for partitioning an input audio stream into regions each of which corresponds to only one audio source or one speaker. One application of this system is in Speaker Diarization…

Artificial Intelligence · Computer Science 2012-05-09 Behrouz Abdolali , Hossein Sameti

Source-tract decomposition (or glottal flow estimation) is one of the basic problems of speech processing. For this, several techniques have been proposed in the literature. However studies comparing different approaches are almost…

Sound · Computer Science 2020-01-06 Thomas Drugman , Baris Bozkurt , Thierry Dutoit