English
Related papers

Related papers: Exploring Robust Face-Voice Matching in Multilingu…

200 papers

Large audio-language models (LALMs) show strong zero-shot ability on speech tasks, suggesting promise for speech emotion recognition (SER). However, SER in real-world deployments often fails under domain mismatch, where source data are…

Computation and Language · Computer Science 2025-09-26 Hsiao-Ying Huang , Yi-Cheng Lin , Hung-yi Lee

Recent advances in deep learning and automatic speech recognition (ASR) have enabled the end-to-end (E2E) ASR system and boosted the accuracy to a new level. The E2E systems implicitly model all conventional ASR components, such as the…

As artificial intelligence systems increasingly operate in Real-world environments, the integration of multi-modal data sources such as vision, language, and audio presents both unprecedented opportunities and critical challenges for…

Machine Learning · Computer Science 2025-07-01 Sree Bhargavi Balija

Cross-modal associations between voice and face from a person can be learnt algorithmically, which can benefit a lot of applications. The problem can be defined as voice-face matching and retrieval tasks. Much research attention has been…

Computer Vision and Pattern Recognition · Computer Science 2020-01-01 Chuyuan Xiong , Deyuan Zhang , Tao Liu , Xiaoyong Du

Leveraging both visual frames and audio has been experimentally proven effective to improve large-scale video classification. Previous research on video classification mainly focuses on the analysis of visual content among extracted video…

Computer Vision and Pattern Recognition · Computer Science 2018-10-01 Jinlai Liu , Zehuan Yuan , Changhu Wang

Far-field speech recognition is a challenging task that conventionally uses signal processing beamforming to attack noise and interference problem. But the performance has been found usually limited due to heavy reliance on environmental…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-08 Dongdi Zhao , Jianbo Ma , Lu Lu , Jinke Li , Xuan Ji , Lei Zhu , Fuming Fang , Ming Liu , Feijun Jiang

Recently, we have seen an increase in the global facial recognition market size. Despite significant advances in face recognition technology with the adoption of convolutional neural networks, there are still open challenges, such as when…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Marcus de Assis Angeloni , Helio Pedrini

In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing. Its aim is to improve speaker recognition performance when speech signals are corrupted by noise. Instead of…

Computation and Language · Computer Science 2020-05-25 Yanpei Shi , Qiang Huang , Thomas Hain

Factorizing speech as disentangled speech representations is vital to achieve highly controllable style transfer in voice conversion (VC). Conventional speech representation learning methods in VC only factorize speech as speaker and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-06 Jie Wang , Jingbei Li , Xintao Zhao , Zhiyong Wu , Shiyin Kang , Helen Meng

Visual language models (VLMs) have shown remarkable capabilities in multimodal tasks but face challenges in maintaining fairness across demographic groups, particularly when deployed in federated learning (FL) environments. This paper…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Chaomeng Chen , Zitong Yu , Junhao Dong , Sen Su , Linlin Shen , Shutao Xia , Xiaochun Cao

Speech forensic tasks (SFTs), such as automatic speaker recognition (ASR), speech emotion recognition (SER), gender recognition (GR), and age estimation (AE), find use in different security and biometric applications. Previous works have…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-18 Orchid Chetia Phukan , Devyani Koshal , Swarup Ranjan Behera , Arun Balaji Buduru , Rajesh Sharma

Deepfake speech detection presents a growing challenge as generative audio technologies continue to advance. We propose a hybrid training framework that advances detection performance through novel augmentation strategies. First, we…

Sound · Computer Science 2025-11-14 Inbal Rimon , Oren Gal , Haim Permuter

The multi-stream paradigm of audio processing, in which several sources are simultaneously considered, has been an active research area for information fusion. Our previous study offered a promising direction within end-to-end automatic…

Computation and Language · Computer Science 2019-10-24 Ruizhi Li , Gregory Sell , Xiaofei Wang , Shinji Watanabe , Hynek Hermansky

Phase-based features related to vocal source characteristics can be incorporated into magnitude-based speaker recognition systems to improve the system performance. However, traditional feature-level fusion methods typically ignore the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-20 Rongfeng Su , Mengjie Du , Xiaokang Liu , Lan Wang , Nan Yan

Emotion recognition plays an important role in human-computer interaction (HCI) and has been extensively studied for decades. Although tremendous improvements have been achieved for posed expressions, recognizing human emotions in…

Computer Vision and Pattern Recognition · Computer Science 2019-06-07 Jie Cai , Zibo Meng , Ahmed Shehab Khan , Zhiyuan Li , James O'Reilly , Shizhong Han , Ping Liu , Min Chen , Yan Tong

For recognizing speakers in video streams, significant research studies have been made to obtain a rich machine learning model by extracting high-level speaker's features such as facial expression, emotion, and gender. However, generating…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Ehsan Asali , Farzan Shenavarmasouleh , Farid Ghareh Mohammadi , Prasanth Sengadu Suresh , Hamid R. Arabnia

Voice-based interfaces rely on a wake-up word mechanism to initiate communication with devices. However, achieving a robust, energy-efficient, and fast detection remains a challenge. This paper addresses these real production needs by…

Sound · Computer Science 2023-10-18 Fernando López , Jordi Luque , Carlos Segura , Pablo Gómez

Face-voice association is widely studied in multimodal learning and is approached representing faces and voices with embeddings that are close for a same person and well separated from those of others. Previous work achieved this with loss…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Marta Moscati , Oleksandr Kats , Mubashir Noman , Muhammad Zaigham Zaheer , Yufang Hou , Markus Schedl , Shah Nawaz

Federated learning (FL) has attracted considerable interest in the medical domain due to its capacity to facilitate collaborative model training while maintaining data privacy. However, conventional FL methods typically necessitate multiple…

Machine Learning · Computer Science 2025-01-08 Naibo Wang , Yuchen Deng , Shichen Fan , Jianwei Yin , See-Kiong Ng

Recent years have witnessed the extraordinary development of automatic speaker verification (ASV). However, previous works show that state-of-the-art ASV models are seriously vulnerable to voice spoofing attacks, and the recently proposed…

Sound · Computer Science 2022-06-22 Haibin Wu , Jiawen Kang , Lingwei Meng , Yang Zhang , Xixin Wu , Zhiyong Wu , Hung-yi Lee , Helen Meng
‹ Prev 1 4 5 6 7 8 10 Next ›