English
Related papers

Related papers: Time-Frequency Phase Retrieval for Audio -- The Ef…

200 papers

Given a time series vector, how can we efficiently compute a specified part of Fourier coefficients? Fast Fourier transform (FFT) is a widely used algorithm that computes the discrete Fourier transform in many machine learning applications.…

Machine Learning · Computer Science 2020-08-31 Yong-chan Park , Jun-Gi Jang , U Kang

Accurate and reliable identification of the relative transfer functions (RTFs) between microphones with respect to a desired source is an essential component in the design of microphone array beamformers, specifically when applying the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-19 Daniel Levi , Amit Sofer , Sharon Gannot

Phase serves as a critical component of speech that influences the quality and intelligibility. Current speech enhancement algorithms are beginning to address phase distortions, but the algorithms focus on normal-hearing (NH) listeners. It…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-30 Zhuohuang Zhang , Donald S. Williamson , Yi Shen

Several methods have recently been proposed to analyze speech and automatically infer the personality of the speaker. These methods often rely on prosodic and other hand crafted speech processing features extracted with off-the-shelf…

Computer Vision and Pattern Recognition · Computer Science 2022-05-10 Marc-André Carbonneau , Eric Granger , Yazid Attabi , Ghyslain Gagnon

In this work, we explore the constant-Q transform (CQT) for speech emotion recognition (SER). The CQT-based time-frequency analysis provides variable spectro-temporal resolution with higher frequency resolution at lower frequencies. Since…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Premjeet Singh , Goutam Saha , Md Sahidullah

We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to estimate a time-frequency (T-F) mask in the short-time Fourier…

Audio and Speech Processing · Electrical Eng. & Systems 2019-03-22 Daiki Takeuchi , Kohei Yatabe , Yuma Koizumi , Yasuhiro Oikawa , Noboru Harada

We introduce a simple and linear SNR (strictly speaking, periodic to random power ratio) estimator (0dB to 80dB without additional calibration/linearization) for providing reliable descriptions of aperiodicity in speech corpus. The main…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-06 Hideki Kawahara , Ken-Ichi Sakakibara , Masanori Morise , Hideki Banno , Tomoki Toda

Traditional speech enhancement techniques modify the magnitude of a speech in time-frequency domain, and use the phase of a noisy speech to resynthesize a time domain speech. This work proposes a complex-valued Gaussian process latent…

Sound · Computer Science 2017-01-02 Sih-Huei Chen , Yuan-Shan Lee , Jia-Ching Wang

This paper proposes an end-to-end approach for single-channel speaker-independent multi-speaker speech separation, where time-frequency (T-F) masking, the short-time Fourier transform (STFT), and its inverse are represented as layers within…

Sound · Computer Science 2018-04-30 Zhong-Qiu Wang , Jonathan Le Roux , DeLiang Wang , John R. Hershey

This paper considers Pseudo-Relevance Feedback (PRF) methods for dense retrievers in a resource constrained environment such as that of cheap cloud instances or embedded systems (e.g., smartphones and smartwatches), where memory and CPU are…

Information Retrieval · Computer Science 2024-12-09 Hang Li , Chuting Yu , Ahmed Mourad , Bevan Koopman , Guido Zuccon

Deep learning has dramatically improved the performance of speech recognition systems through learning hierarchies of features optimized for the task at hand. However, true end-to-end learning, where features are learned directly from…

Computation and Language · Computer Science 2016-04-06 Zhenyao Zhu , Jesse H. Engel , Awni Hannun

Many audio synthesizers can produce the same signal given different parameter configurations, meaning the inversion from sound to parameters is an inherently ill-posed problem. We show that this is largely due to intrinsic symmetries of the…

Sound · Computer Science 2025-06-10 Ben Hayes , Charalampos Saitis , György Fazekas

Parameter-Efficient finetuning (PEFT) enhances model performance on downstream tasks by updating a minimal subset of parameters. Representation finetuning (ReFT) methods further improve efficiency by freezing model weights and optimizing…

Machine Learning · Computer Science 2025-11-17 Sirui Liang , Pengfei Cao , Jian Zhao , Cong Huang , Jun Zhao , Kang Liu

The Personal Alert Safety System (PASS) is an alarm signal device carried by firefighters to help rescuers locate and extricate downed firefighters. A fire creates temperature gradients and inhomogeneous time-varying temperature, density,…

Applied Physics · Physics 2022-04-06 Mustafa Z. Abbasi , Preston S. Wilson , Ofodike A. Ezekoye

Recently, a novel form of audio partial forgery has posed challenges to its forensics, requiring advanced countermeasures to detect subtle forgery manipulations within long-duration audio. However, existing countermeasures still serve a…

Multimedia · Computer Science 2024-07-24 Junyan Wu , Wei Lu , Xiangyang Luo , Rui Yang , Qian Wang , Xiaochun Cao

Due to its appearance in a remarkably wide field of applications, such as audio processing and coherent diffraction imaging, the short-time Fourier transform (STFT) phase retrieval problem has seen a great deal of attention in recent years.…

Functional Analysis · Mathematics 2025-05-06 Philipp Grohs , Lukas Liehr

The Fractional Fourier Transform (FRT) corresponds to an arbitrary-angle rotation in the phase space, e.g. the time-frequency (TF) space, and generalizes the fundamentally important Fourier Transform. FRT applications range from classical…

Optics · Physics 2024-03-06 Michał Lipka , Michał Parniak

The SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is sub-optimal, which…

Audio and Speech Processing · Electrical Eng. & Systems 2019-03-26 Chenglin Xu , Wei Rao , Eng Siong Chng , Haizhou Li

The main aim of this paper is to study quaternion phase retrieval (QPR), i.e., the recovery of quaternion signal from the magnitude of quaternion linear measurements. We show that all $d$-dimensional quaternion signals can be reconstructed…

Signal Processing · Electrical Eng. & Systems 2023-07-25 Junren Chen , Michael K. Ng

It is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with a single-channel speech enhancement (SE) front-end. This is generally attributed to the processing distortions caused by the nonlinear…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-24 Tsubasa Ochiai , Kazuma Iwamoto , Marc Delcroix , Rintaro Ikeshita , Hiroshi Sato , Shoko Araki , Shigeru Katagiri
‹ Prev 1 8 9 10 Next ›