English
Related papers

Related papers: Improving Visual Speech Enhancement Network by Lea…

200 papers

We propose a multi-stage framework for universal speech enhancement, designed for the Interspeech 2025 URGENT Challenge. Our system first employs a Sparse Compression Network to robustly separate sources and extract an initial clean speech…

Sound · Computer Science 2025-06-03 Nabarun Goswami , Tatsuya Harada

Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement.…

Sound · Computer Science 2022-10-31 Shulin He , Wei Rao , Jinjiang Liu , Jun Chen , Yukai Ju , Xueliang Zhang , Yannan Wang , Shidong Shang

Current Audio-Visual Source Separation methods primarily adopt two design strategies. The first strategy involves fusing audio and visual features at the bottleneck layer of the encoder, followed by processing the fused features through the…

Sound · Computer Science 2025-05-01 Yinfeng Yu , Shiyu Sun

Visual information can serve as an effective cue for target speaker extraction (TSE) and is vital to improving extraction performance. In this paper, we propose AV-SepFormer, a SepFormer-based attention dual-scale model that utilizes cross-…

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvements over audio-only…

Nowadays, pre-trained encoders are widely used in medical image segmentation due to their strong capability in extracting rich and generalized feature representations. However, existing methods often fail to fully leverage these features,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Xiaolin Gou , Chuanlin Liao , Jizhe Zhou , Fengshuo Ye , Yi Lin

A speech emotion recognition algorithm based on multi-feature and Multi-lingual fusion is proposed in order to resolve low recognition accuracy caused by lack of large speech dataset and low robustness of acoustic features in the…

Computation and Language · Computer Science 2020-01-17 Chunyi Wang

Multimodal analysis has recently drawn much interest in affective computing, since it can improve the overall accuracy of emotion recognition over isolated uni-modal approaches. The most effective techniques for multimodal emotion…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 R. Gnana Praveen , Eric Granger , Patrick Cardinal

Recently, referring image segmentation has aroused widespread interest. Previous methods perform the multi-modal fusion between language and vision at the decoding side of the network. And, linguistic feature interacts with visual feature…

Computer Vision and Pattern Recognition · Computer Science 2021-05-06 Guang Feng , Zhiwei Hu , Lihe Zhang , Huchuan Lu

Perceptually-inspired objective functions such as the perceptual evaluation of speech quality (PESQ), signal-to-distortion ratio (SDR), and short-time objective intelligibility (STOI), have recently been used to optimize performance of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-27 Khandokar Md. Nayem , Donald S. Williamson

This study aims to design and implement a laughter recognition system based on multimodal fusion and deep learning, leveraging image and audio processing technologies to achieve accurate laughter recognition and emotion analysis. First, the…

Sound · Computer Science 2024-08-01 Fuzheng Zhao , Yu Bai

This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy audio speech with a…

Sound · Computer Science 2022-07-14 Joanna Hong , Minsu Kim , Daehun Yoo , Yong Man Ro

Currently, end-to-end (E2E) speech recognition methods have achieved promising performance. However, auto speech recognition (ASR) models still face challenges in recognizing multi-accent speech accurately. We propose a layer-adapted fusion…

Sound · Computer Science 2024-07-04 Jinming Chen , Jingyi Fang , Yuanzhong Zheng , Yaoxuan Wang , Haojun Fei

In cooperative perception studies, there is often a trade-off between communication bandwidth and perception performance. While current feature fusion solutions are known for their excellent object detection performance, transmitting the…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Deyuan Qu , Qi Chen , Yongqi Zhu , Yihao Zhu , Sergei S. Avedisov , Song Fu , Qing Yang

We present an end-to-end multichannel speaker-attributed automatic speech recognition (MC-SA-ASR) system that combines a Conformer-based encoder with multi-frame crosschannel attention and a speaker-attributed Transformer-based decoder. To…

Computation and Language · Computer Science 2023-10-17 Can Cui , Imran Ahamad Sheikh , Mostafa Sadeghi , Emmanuel Vincent

Attention has become one of the most commonly used mechanisms in deep learning approaches. The attention mechanism can help the system focus more on the feature space's critical regions. For example, high amplitude regions can play an…

Sound · Computer Science 2022-08-24 Junghun Kim , Yoojin An , Jihie Kim

Recent studies have increasingly acknowledged the advantages of incorporating visual data into speech enhancement (SE) systems. In this paper, we introduce a novel audio-visual SE approach, termed DCUC-Net (deep complex U-Net with conformer…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-10 Shafique Ahmed , Chia-Wei Chen , Wenze Ren , Chin-Jou Li , Ernie Chu , Jun-Cheng Chen , Amir Hussain , Hsin-Min Wang , Yu Tsao , Jen-Cheng Hou

Auditory attention detection (AAD) aims to detect the target speaker in a multi-talker environment from brain signals, such as electroencephalography (EEG), which has made great progress. However, most AAD methods solely utilize attention…

Human-Computer Interaction · Computer Science 2025-05-22 Lu Li , Cunhang Fan , Hongyu Zhang , Jingjing Zhang , Xiaoke Yang , Jian Zhou , Zhao Lv

Speech quality and intelligibility are significantly degraded in noisy environments. This paper presents a novel transformer-based learning framework to address the single-channel noise suppression problem for real-time applications.…

Sound · Computer Science 2025-11-18 Behnaz Bahmei , Siamak Arzanpour , Elina Birmingham

In this paper, we propose an end-to-end speech recognition network based on Nvidia's previous QuartzNet model. We try to promote the model performance, and design three components: (1) Multi-Resolution Convolution Module, replaces the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-30 Jian Luo , Jianzong Wang , Ning Cheng , Guilin Jiang , Jing Xiao