English
Related papers

Related papers: PAEFF: Precise Alignment and Enhanced Gated Featur…

200 papers

Incremental improvements in accuracy of Convolutional Neural Networks are usually achieved through use of deeper and more complex models trained on larger datasets. However, enlarging dataset and models increases the computation and storage…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-24 Mahdi Hajibabaei , Dengxin Dai

In this paper, we propose a framework for disentangling the appearance and geometry representations in the face recognition task. To provide supervision for this aim, we generate geometrically identical faces by incorporating spatial…

Computer Vision and Pattern Recognition · Computer Science 2020-01-15 Ali Dabouei , Fariborz Taherkhani , Sobhan Soleymani , Jeremy Dawson , Nasser M. Nasrabadi

The joint training framework for speech enhancement and recognition methods have obtained quite good performances for robust end-to-end automatic speech recognition (ASR). However, these methods only utilize the enhanced feature as the…

Sound · Computer Science 2020-11-10 Cunhang Fan , Jiangyan Yi , Jianhua Tao , Zhengkun Tian , Bin Liu , Zhengqi Wen

Despite significant recent advances in the field of face recognition, implementing face verification and recognition efficiently at scale presents serious challenges to current approaches. In this paper we present a system, called FaceNet,…

Computer Vision and Pattern Recognition · Computer Science 2016-11-18 Florian Schroff , Dmitry Kalenichenko , James Philbin

Audio-visual speech enhancement system is regarded to be one of promising solutions for isolating and enhancing speech of desired speaker. Conventional methods focus on predicting clean speech spectrum via a naive convolution neural network…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-28 Xinmeng Xu , Jianjun Hao

Previous works on voice-face matching and voice-guided face synthesis demonstrate strong correlations between voice and face, but mainly rely on coarse semantic cues such as gender, age, and emotion. In this paper, we aim to investigate the…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Xiang Li , Yandong Wen , Muqiao Yang , Jinglu Wang , Rita Singh , Bhiksha Raj

In the domain of facial recognition security, multimodal Face Anti-Spoofing (FAS) is essential for countering presentation attacks. However, existing technologies encounter challenges due to modality biases and imbalances, as well as domain…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Yingjie Ma , Zitong Yu , Xun Lin , Weicheng Xie , Linlin Shen

In the field of face recognition, a model learns to distinguish millions of face images with fewer dimensional embedding features, and such vast information may not be properly encoded in the conventional model with a single branch. We…

Computer Vision and Pattern Recognition · Computer Science 2020-05-26 Yonghyun Kim , Wonpyo Park , Myung-Cheol Roh , Jongju Shin

Speech emotion recognition (SER) remains a challenging yet crucial task due to the inherent complexity and diversity of human emotions. To address this problem, researchers attempt to fuse information from other modalities via multimodal…

Sound · Computer Science 2024-12-10 Feng Li , Jiusong Luo , Wanjun Xia

Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained temporal correspondences…

Verifying the identity of a speaker is crucial in modern human-machine interfaces, e.g., to ensure privacy protection or to enable biometric authentication. Classical speaker verification (SV) approaches estimate a fixed-dimensional…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-29 Ahmad Aloradi , Wolfgang Mack , Mohamed Elminshawi , Emanuël A. P. Habets

Speech contains both acoustic and linguistic patterns that reflect cognitive decline, and therefore models describing only one domain cannot fully capture such complexity. This study investigates how early fusion (EF) of speech and its…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-02 Krystof Novotny , Laureano Moro-Velázquez , Jiri Mekyska

Detecting forgery videos is highly desirable due to the abuse of deepfake. Existing detection approaches contribute to exploring the specific artifacts in deepfake videos and fit well on certain data. However, the growing technique on these…

Computer Vision and Pattern Recognition · Computer Science 2022-06-14 Harry Cheng , Yangyang Guo , Tianyi Wang , Qi Li , Xiaojun Chang , Liqiang Nie

Effective fusion of multi-scale features is crucial for improving speaker verification performance. While most existing methods aggregate multi-scale features in a layer-wise manner via simple operations, such as summation or concatenation.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-04 Yafeng Chen , Siqi Zheng , Hui Wang , Luyao Cheng , Qian Chen , Jiajun Qi

Personal Voice Activity Detection (PVAD) is crucial for identifying target speaker segments in the mixture, yet its performance heavily depends on the quality of speaker embeddings. A key practical limitation is the short enrollment…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Fuyuan Feng , Wenbin Zhang , Yu Gao , Longting Xu , Xiaofeng Mou , Yi Xu

The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, audio-visual systems are among the most widely used multimodal systems. In the recent years, associating face and voice…

Blur in facial images significantly impedes the efficiency of recognition approaches. However, most existing blind deconvolution methods cannot generate satisfactory results due to their dependence on strong edges, which are sufficient in…

Computer Vision and Pattern Recognition · Computer Science 2019-04-19 Dayong Tian , Dacheng Tao

Approximately 1.2% of the world's population has impaired voice production. As a result, automatic dysphonic voice detection has attracted considerable academic and clinical interest. However, existing methods for automated voice assessment…

Sound · Computer Science 2023-01-27 Jianwei Zhang , Julie Liss , Suren Jayasuriya , Visar Berisha

Audio-visual navigation tasks require agents to locate and navigate toward continuously vocalizing targets using only visual observations and acoustic cues. However, existing methods mainly rely on simple feature concatenation or late…

Sound · Computer Science 2026-04-06 Shaohang Wu , Yinfeng Yu

Face Recognition (FR) tasks have made significant progress with the advent of Deep Neural Networks, particularly through margin-based triplet losses that embed facial images into high-dimensional feature spaces. During training, these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Pierrick Leroy , Antonio Mastropietro , Marco Nurisso , Francesco Vaccarino