English
Related papers

Related papers: McNet: Fuse Multiple Cues for Multichannel Speech …

200 papers

Dialogue separation involves isolating a dialogue signal from a mixture, such as a movie or a TV program. This can be a necessary step to enable dialogue enhancement for broadcast-related applications. In this paper, ConcateNet for dialogue…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-19 Mhd Modar Halimeh , Matteo Torcoli , Emanuël Habets

Recently, the end-to-end training approach for neural beamformer-supported multi-channel ASR has shown its effectiveness in multi-channel speech recognition. However, the integration of multiple modules makes it more difficult to perform…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-07 Keyu An , Zhijian Ou

This paper presents AMNet, an Acoustic Model Network designed to improve the performance of Mandarin speech synthesis by incorporating phrase structure annotation and local convolution modules. AMNet builds upon the FastSpeech 2…

Sound · Computer Science 2025-04-15 Yubing Cao , Yinfeng Yu , Yongming Li , Liejun Wang

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

Sound · Computer Science 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee

Time series forecasting is crucial in many fields, yet current deep learning models struggle with noise, data sparsity, and capturing complex multi-scale patterns. This paper presents MFF-FTNet, a novel framework addressing these challenges…

Machine Learning · Computer Science 2024-11-27 Yangyang Shi , Qianqian Ren , Yong Liu , Jianguo Sun

Audio-visual speech enhancement (AV-SE) methods utilize auxiliary visual cues to enhance speakers' voices. Therefore, technically they should be able to outperform the audio-only speech enhancement (SE) methods. However, there are few works…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-14 Zirun Zhu , Hemin Yang , Min Tang , Ziyi Yang , Sefik Emre Eskimez , Huaming Wang

Few-shot semantic segmentation is the task of learning to locate each pixel of the novel class in the query image with only a few annotated support images. The current correlation-based methods construct pair-wise feature correlations to…

Computer Vision and Pattern Recognition · Computer Science 2023-01-20 Huafeng Liu , Pai Peng , Tao Chen , Qiong Wang , Yazhou Yao , Xian-Sheng Hua

Transformers are powerful neural architectures that allow integrating different modalities using attention mechanisms. In this paper, we leverage the neural transformer architectures for multi-channel speech recognition systems, where the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Feng-Ju Chang , Martin Radfar , Athanasios Mouchtaris , Brian King , Siegfried Kunzmann

Multimodal speech emotion recognition aims to detect speakers' emotions from audio and text. Prior works mainly focus on exploiting advanced networks to model and fuse different modality information to facilitate performance, while…

Computation and Language · Computer Science 2023-04-11 Zhen Wu , Yizhe Lu , Xinyu Dai

Recent research has made significant progress in designing fusion modules for audio-visual speech separation. However, they predominantly focus on multi-modal fusion at a single temporal scale of auditory and visual features without…

Sound · Computer Science 2024-02-05 Kai Li , Runxuan Yang , Fuchun Sun , Xiaolin Hu

This paper proposes a speech enhancement method which exploits the high potential of residual connections in a Wide Residual Network architecture. This is supported on single dimensional convolutions computed alongside the time domain,…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-11 Jorge Llombart , Dayana Ribas , Antonio Miguel , Luis Vicente , Alfonso Ortega , Eduardo Lleida

One persistent challenge in Speech Emotion Recognition (SER) is the ubiquitous environmental noise, which frequently results in deteriorating SER performance in practice. In this paper, we introduce a Two-level Refinement Network, dubbed…

Sound · Computer Science 2024-09-04 Chengxin Chen , Pengyuan Zhang

Human beings have developed fantastic abilities to integrate information from various sensory sources exploring their inherent complementarity. Perceptual capabilities are therefore heightened, enabling, for instance, the well-known…

Computer Vision and Pattern Recognition · Computer Science 2021-04-14 Gustavo Assunção , Nuno Gonçalves , Paulo Menezes

In this paper, we investigate a deep learning approach for speech denoising through an efficient ensemble of specialist neural networks. By splitting up the speech denoising task into non-overlapping subproblems and introducing a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Aswin Sivaraman , Minje Kim

Disfluency, though originating from human spoken utterances, is primarily studied as a uni-modal text-based Natural Language Processing (NLP) task. Based on early-fusion and self-attention-based multimodal interaction between text and…

Computation and Language · Computer Science 2022-11-29 Sreyan Ghosh , Utkarsh Tyagi , Sonal Kumar , Manan Suri , Rajiv Ratn Shah

Over the past few years, speech enhancement methods based on deep learning have greatly surpassed traditional methods based on spectral subtraction and spectral estimation. Many of these new techniques operate directly in the the short-time…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Jean-Marc Valin , Umut Isik , Neerad Phansalkar , Ritwik Giri , Karim Helwani , Arvindh Krishnaswamy

The human auditory system has the ability to selectively focus on key speech elements in an audio stream while giving secondary attention to less relevant areas such as noise or distortion within the background, dynamically adjusting its…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-09 Nursadul Mamun , John H. L. Hansen

In this paper, we present TridentSE, a novel architecture for speech enhancement, which is capable of efficiently capturing both global information and local details. TridentSE maintains T-F bin level representation to capture details, and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-25 Dacheng Yin , Zhiyuan Zhao , Chuanxin Tang , Zhiwei Xiong , Chong Luo

To achieve continuous massive data transmission with significantly reduced data payload, the users can adopt semantic communication techniques to compress the redundant information by transmitting semantic features instead. However, current…

Signal Processing · Electrical Eng. & Systems 2024-01-30 Youcheng Zeng , Xinxin He , Xu Chen , Haonan Tong , Zhaohui Yang , Yijun Guo , Jianjun Hao

Emotion represents an essential aspect of human speech that is manifested in speech prosody. Speech, visual, and textual cues are complementary in human communication. In this paper, we study a hybrid fusion method, referred to as…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-10 Zexu Pan , Zhaojie Luo , Jichen Yang , Haizhou Li