English
Related papers

Related papers: Magnitude and Phase-based Feature Fusion Using Co-…

200 papers

Transformer-based speech recognition models have achieved great success due to the self-attention (SA) mechanism that utilizes every frame in the feature extraction process. Especially, SA heads in lower layers capture various phonetic…

Computation and Language · Computer Science 2022-07-13 Kyuhong Shim , Wonyong Sung

Due to the high computational complexity to model more frequency bands, it is still intractable to conduct real-time full-band speech enhancement based on deep neural networks. Recent studies typically utilize the compressed perceptually…

Sound · Computer Science 2022-06-16 Guochen Yu , Andong Li , Wenzhe Liu , Chengshi Zheng , Yutian Wang , Hui Wang

Facial Expression Recognition (FER) is a critical task within computer vision with diverse applications across various domains. Addressing the challenge of limited FER datasets, which hampers the generalization capability of expression…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Bach Nguyen-Xuan , Thien Nguyen-Hoang , Thanh-Huy Nguyen , Nhu Tai-Do

Target speech separation refers to extracting a target speaker's voice from an overlapped audio of simultaneous talkers. Previously the use of visual modality for target speech separation has demonstrated great potentials. This work…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-26 Rongzhi Gu , Shi-Xiong Zhang , Yong Xu , Lianwu Chen , Yuexian Zou , Dong Yu

Speech emotion recognition (SER) is crucial for enhancing affective computing and enriching the domain of human-computer interaction. However, the main challenge in SER lies in selecting relevant feature representations from speech signals…

Sound · Computer Science 2024-12-16 Niloy Kumar Kundu , Sarah Kobir , Md. Rayhan Ahmed , Tahmina Aktar , Niloya Roy

The audio-video based multimodal emotion recognition has attracted a lot of attention due to its robust performance. Most of the existing methods focus on proposing different cross-modal fusion strategies. However, these strategies…

Computer Vision and Pattern Recognition · Computer Science 2021-11-04 Ziwang Fu , Feng Liu , Hanyang Wang , Jiayin Qi , Xiangling Fu , Aimin Zhou , Zhibin Li

Transformer based end-to-end modelling approaches with multiple stream inputs have been achieved great success in various automatic speech recognition (ASR) tasks. An important issue associated with such approaches is that the intermediate…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-11 Jin Li , Rongfeng Su , Xurong Xie , Nan Yan , Lan Wang

Humans perceive the world by concurrently processing and fusing high-dimensional inputs from multiple modalities such as vision and audio. Machine perception models, in stark contrast, are typically modality-specific and optimised for…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Arsha Nagrani , Shan Yang , Anurag Arnab , Aren Jansen , Cordelia Schmid , Chen Sun

The joint training framework for speech enhancement and recognition methods have obtained quite good performances for robust end-to-end automatic speech recognition (ASR). However, these methods only utilize the enhanced feature as the…

Sound · Computer Science 2020-11-10 Cunhang Fan , Jiangyan Yi , Jianhua Tao , Zhengkun Tian , Bin Liu , Zhengqi Wen

Most of the prior studies in the spatial \ac{DoA} domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio…

Sound · Computer Science 2021-05-14 Xinyuan Qian , Maulik Madhavi , Zexu Pan , Jiadong Wang , Haizhou Li

The use of multi-modal data for deep machine learning has shown promise when compared to uni-modal approaches with fusion of multi-modal features resulting in improved performance in several applications. However, most state-of-the-art…

Machine Learning · Computer Science 2020-10-26 Darshana Priyasad , Tharindu Fernando , Simon Denman , Sridha Sridharan , Clinton Fookes

Initially developed for natural language processing (NLP), Transformer model is now widely used for speech processing tasks such as speaker recognition, due to its powerful sequence modeling capabilities. However, conventional…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-28 Rui Wang , Junyi Ao , Long Zhou , Shujie Liu , Zhihua Wei , Tom Ko , Qing Li , Yu Zhang

In this paper, a hierarchical attention network to generate utterance-level embeddings (H-vectors) for speaker identification is proposed. Since different parts of an utterance may have different contributions to speaker identities, the use…

Computation and Language · Computer Science 2019-10-22 Yanpei Shi , Qiang Huang , Thomas Hain

We propose a multi-stage framework for universal speech enhancement, designed for the Interspeech 2025 URGENT Challenge. Our system first employs a Sparse Compression Network to robustly separate sources and extract an initial clean speech…

Sound · Computer Science 2025-06-03 Nabarun Goswami , Tatsuya Harada

Sentiment analysis, mostly based on text, has been rapidly developing in the last decade and has attracted widespread attention in both academia and industry. However, the information in the real world usually comes from multiple…

Computation and Language · Computer Science 2019-12-12 Feiyang Chen , Ziqian Luo , Yanyan Xu , Dengfeng Ke

Target speaker extraction focuses on extracting a target speech signal from an environment with multiple speakers by leveraging an enrollment. Existing methods predominantly rely on speaker embeddings obtained from the enrollment,…

Sound · Computer Science 2025-02-13 Ke Xue , Rongfei Fan , Shanping Yu , Chang Sun , Jianping An

Recently, hyperspherical embeddings have established themselves as a dominant technique for face and voice recognition. Specifically, Euclidean space vector embeddings are learned to encode person-specific information in their direction…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-25 Nikita Kuzmin , Igor Fedorov , Alexey Sholokhov

This paper proposes attentive statistics pooling for deep speaker embedding in text-independent speaker verification. In conventional speaker embedding, frame-level features are averaged over all the frames of a single utterance to form an…

Audio and Speech Processing · Electrical Eng. & Systems 2019-02-27 Koji Okabe , Takafumi Koshinaka , Koichi Shinoda

Speech Emotion Recognition (SER) plays a key role in advancing human-computer interaction. Attention mechanisms have become the dominant approach for modeling emotional speech due to their ability to capture long-range dependencies and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Marc Casals-Salvador , Federico Costa , Rodolfo Zevallos , Javier Hernando

Identifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task. In this paper, a hierarchical attention network is proposed to solve a weakly labelled speaker identification problem. The use of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Yanpei Shi , Qiang Huang , Thomas Hain