English
Related papers

Related papers: Masks Fusion with Multi-Target Learning For Speech…

200 papers

Fine-tuning speech representation models can enhance performance on specific tasks but often compromises their cross-task generalization ability. This degradation is often caused by excessive changes in the representations, making it…

Computation and Language · Computer Science 2026-04-28 Tzu-Quan Lin , Wei-Ping Huang , Hao Tang , Hung-yi Lee

Self-supervised pre-training of a speech foundation model, followed by supervised fine-tuning, has shown impressive quality improvements on automatic speech recognition (ASR) tasks. Fine-tuning separate foundation models for many downstream…

Machine Learning · Computer Science 2022-11-08 Zhouyuan Huo , Khe Chai Sim , Bo Li , Dongseong Hwang , Tara N. Sainath , Trevor Strohman

Universal speech enhancement aims at handling inputs with various speech distortions and recording conditions. In this work, we propose a novel hybrid architecture that synergizes the signal fidelity of discriminative modeling with the…

Sound · Computer Science 2026-01-28 Yinghao Liu , Chengwei Liu , Xiaotao Liang , Haoyin Yan , Shaofei Xue , Zheng Xue

Performance degradation of an Automatic Speech Recognition (ASR) system is commonly observed when the test acoustic condition is different from training. Hence, it is essential to make ASR systems robust against various environmental…

Sound · Computer Science 2021-02-08 Ruizhi Li , Gregory Sell , Hynek Hermansky

Time-frequency masking or spectrum prediction computed via short symmetric windows are commonly used in low-latency deep neural network (DNN) based source separation. In this paper, we propose the usage of an asymmetric analysis-synthesis…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-23 Shanshan Wang , Gaurav Naithani , Archontis Politis , Tuomas Virtanen

Target speech separation refers to extracting a target speaker's voice from an overlapped audio of simultaneous talkers. Previously the use of visual modality for target speech separation has demonstrated great potentials. This work…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-26 Rongzhi Gu , Shi-Xiong Zhang , Yong Xu , Lianwu Chen , Yuexian Zou , Dong Yu

Statistical signal processing based speech enhancement methods adopt expert knowledge to design the statistical models and linear filters, which is complementary to the deep neural network (DNN) based methods which are data-driven. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-19 Wei Xue , Gang Quan , Chao Zhang , Guohong Ding , Xiaodong He , Bowen Zhou

DeepFake Audio, unlike DeepFake images and videos, has been relatively less explored from detection perspective, and the solutions which exist for the synthetic speech classification either use complex networks or dont generalize to…

Sound · Computer Science 2022-10-24 Vardhan Dongre , Abhinav Thimma Reddy , Nikhitha Reddeddy

Audio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker. Typical methods focus on predicting clean speech spectrum via a naive convolution neural network based…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-01 Xinmeng Xu , Yang Wang , Jie Jia , Binbin Chen , Dejun Li

Monaural speech enhancement has been widely studied using real networks in the time-frequency (TF) domain. However, the input and the target are naturally complex-valued in the TF domain, a fully complex network is highly desirable for…

Sound · Computer Science 2023-02-24 Shengkui Zhao , Bin Ma

End-to-end models in general, and Recurrent Neural Network Transducer (RNN-T) in particular, have gained significant traction in the automatic speech recognition community in the last few years due to their simplicity, compactness, and…

Computation and Language · Computer Science 2020-11-17 Duc Le , Gil Keren , Julian Chan , Jay Mahadeokar , Christian Fuegen , Michael L. Seltzer

For time-frequency (TF) domain speech enhancement (SE) methods, the overlap-and-add operation in the inverse TF transformation inevitably leads to an algorithmic delay equal to the window size. However, typical causal SE systems fail to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-22 Yuewei Zhang , Huanbin Zou , Jie Zhu

We study the problem of learning association between face and voice, which is gaining interest in the computer vision community lately. Prior works adopt pairwise or triplet loss formulations to learn an embedding space amenable for…

Computer Vision and Pattern Recognition · Computer Science 2021-12-21 Muhammad Saad Saeed , Muhammad Haris Khan , Shah Nawaz , Muhammad Haroon Yousaf , Alessio Del Bue

Deep learning has the potential to enhance speech signals and increase their intelligibility for users of hearing aids. Deep models suited for real-world application should feature a low computational complexity and low processing delay of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-31 Nils L. Westhausen , Hendrik Kayser , Theresa Jansen , Bernd T. Meyer

In this work, we propose a training algorithm for an audio-visual automatic speech recognition (AV-ASR) system using deep recurrent neural network (RNN).First, we train a deep RNN acoustic model with a Connectionist Temporal Classification…

Computer Vision and Pattern Recognition · Computer Science 2016-11-10 Abhinav Thanda , Shankar M Venkatesan

This paper aims to address two issues existing in the current speech enhancement methods: 1) the difficulty of phase estimations; 2) a single objective function cannot consider multiple metrics simultaneously. To solve the first problem, we…

Machine Learning · Statistics 2017-09-12 Szu-Wei Fu , Ting-yao Hu , Yu Tsao , Xugang Lu

Modern speaker verification models use deep neural networks to encode utterance audio into discriminative embedding vectors. During the training process, these networks are typically optimized to differentiate arbitrary speakers. This…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-09 Hua Shen , Yuguang Yang , Guoli Sun , Ryan Langman , Eunjung Han , Jasha Droppo , Andreas Stolcke

Detecting spoofed utterances is a fundamental problem in voice-based biometrics. Spoofing can be performed either by logical accesses like speech synthesis, voice conversion or by physical accesses such as replaying the pre-recorded…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-28 Mari Ganesh Kumar , Suvidha Rupesh Kumar , Saranya M , B. Bharathi , Hema A. Murthy

In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning input that specifies…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-24 Zhepei Wang , Ritwik Giri , Devansh Shah , Jean-Marc Valin , Michael M. Goodwin , Paris Smaragdis

Most of the deep learning based speech enhancement (SE) methods rely on estimating the magnitude spectrum of the clean speech signal from the observed noisy speech signal, either by magnitude spectral masking or regression. These methods…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-28 Raktim Gautam Goswami , Sivaganesh Andhavarapu , K Sri Rama Murty
‹ Prev 1 8 9 10 Next ›