English
Related papers

Related papers: Magnitude and Phase-based Feature Fusion Using Co-…

200 papers

Speaker Verification (SV) systems involve mainly two individual stages: feature extraction and classification. In this paper, we explore these two modules with the aim of improving the performance of a speaker verification system under…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-06 Kerlos Atia Abdalmalak , Ascensión Gallardo-Antol'in

Multi-view subspace clustering always performs well in high-dimensional data analysis, but is sensitive to the quality of data representation. To this end, a two stage fusion strategy is proposed to embed representation learning into the…

Signal Processing · Electrical Eng. & Systems 2022-01-07 Run-kun Lu , Jian-wei Liu , Ze-yu Liu , Jin-zhong Chen

Perceptual voice quality assessment plays a vital role in diagnosing and monitoring voice disorders. Traditional methods, such as the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) and the Grade, Roughness, Breathiness,…

Sound · Computer Science 2025-12-12 Whenty Ariyanti , Kuan-Yu Chen , Sabato Marco Siniscalchi , Hsin-Min Wang , Yu Tsao

To accomplish punctuation restoration, most existing methods focus on introducing extra information (e.g., part-of-speech) or addressing the class imbalance problem. Recently, large-scale transformer-based pre-trained language models (PLMS)…

Computation and Language · Computer Science 2022-11-10 Yangjun Wu , Kebin Fang , Yao Zhao , Hao Zhang , Lifeng Shi , Mengqi Zhang

The joint training of speech enhancement and speaker embedding networks for speaker recognition is widely adopted under noisy acoustic environments. While effective, this paradigm often fails to leverage the generalization and robustness…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-29 Chong-Xin Gan , Peter Bell , Man-Wai Mak , Zhe Li , Zezhong Jin , Zilong Huang , Kong Aik Lee

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

Sound · Computer Science 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee

Recently, deep neural network (DNN) based time-frequency (T-F) mask estimation has shown remarkable effectiveness for speech enhancement. Typically, a single T-F mask is first estimated based on DNN and then used to mask the spectrogram of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-29 Liangchen Zhou , Wenbin Jiang , Jingyan Xu , Fei Wen , Peilin Liu

Accurate recognition of sign language in healthcare communication poses a significant challenge, requiring frameworks that can accurately interpret complex multimodal gestures. To deal with this, we propose FusionEnsemble-Net, a novel…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Md. Milon Islam , Md Rezwanul Haque , S M Taslim Uddin Raju , Fakhri Karray

The objective of this paper is to combine multiple frame-level features into a single utterance-level representation considering pairwise relationship. For this purpose, we propose a novel graph attentive feature aggregation module by…

Sound · Computer Science 2021-12-24 Hye-jin Shim , Jungwoo Heo , Jae-han Park , Ga-hui Lee , Ha-Jin Yu

Recently, self-supervised learning (SSL) techniques have been introduced to solve the monaural speech enhancement problem. Due to the lack of using clean phase information, the enhancement performance is limited in most SSL methods.…

Sound · Computer Science 2021-12-22 Yi Li , Yang Sun , Syed Mohsen Naqvi

Audiovisual active speaker detection (ASD) in egocentric recordings is challenged by frequent occlusions, motion blur, and audio interference, which undermine the discernability of temporal synchrony between lip movement and speech.…

Multimedia · Computer Science 2025-08-15 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

Multi-modal learning has emerged as a crucial research direction, as integrating textual and visual information can substantially enhance performance in tasks such as classification, retrieval, and scene understanding. Despite advances with…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Md. Mithun Hossain , Md. Shakil Hossain , Sudipto Chaki , M. F. Mridha

In this study, we propose a novel multi-modal end-to-end neural approach for automated assessment of non-native English speakers' spontaneous speech using attention fusion. The pipeline employs Bi-directional Recurrent Convolutional Neural…

Computation and Language · Computer Science 2021-11-30 Manraj Singh Grover , Yaman Kumar , Sumit Sarin , Payman Vafaee , Mika Hama , Rajiv Ratn Shah

Despite improvements in automatic speaker verification (ASV), vulnerability against spoofing attacks remains a major concern. In this study, we investigate the integration of ASV and countermeasure (CM) subsystems into a modular spoof-aware…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-17 Oguzhan Kurnaz , Tomi Kinnunen , Cemal Hanilci

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

Sound · Computer Science 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

Majority of the recent approaches for text-independent speaker recognition apply attention or similar techniques for aggregation of frame-level feature descriptors generated by a deep neural network (DNN) front-end. In this paper, we…

Sound · Computer Science 2019-10-22 Sarthak Yadav , Atul Rai

In recent years, an association is established between faces and voices of celebrities leveraging large scale audio-visual information from YouTube. The availability of large scale audio-visual datasets is instrumental in developing speaker…

Sound · Computer Science 2023-02-28 Saqlain Hussain Shah , Muhammad Saad Saeed , Shah Nawaz , Muhammad Haroon Yousaf

Speech emotion recognition (SER) plays a critical role in building emotion-aware speech systems, but its performance degrades significantly under noisy conditions. Although speech enhancement (SE) can improve robustness, it often introduces…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-29 Jing-Tong Tzeng , Carlos Busso , Chi-Chun Lee

Natural Language Processing has recently made understanding human interaction easier, leading to improved sentimental analysis and behaviour prediction. However, the choice of words and vocal cues in conversations presents an underexplored…

Computers and Society · Computer Science 2022-06-24 Amna Anwar , Eiman Kanjo , Dario Ortega Anderez

The audio-video based emotion recognition aims to classify a given video into basic emotions. In this paper, we describe our approaches in EmotiW 2019, which mainly explores emotion features and feature fusion strategies for audio and…

Computer Vision and Pattern Recognition · Computer Science 2020-12-29 Hengshun Zhou , Debin Meng , Yuanyuan Zhang , Xiaojiang Peng , Jun Du , Kai Wang , Yu Qiao