English
Related papers

Related papers: Exploring Robust Face-Voice Matching in Multilingu…

200 papers

Emotional Mimicry Intensity (EMI) estimation plays a pivotal role in understanding human social behavior and advancing human-computer interaction. The core challenges lie in dynamic correlation modeling and robust fusion of multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Jun Yu , Lingsi Zhu , Yanjun Chi , Yunxiang Zhang , Yang Zheng , Yongqi Wang , Xilong Lu

Methods for building fair predictors often involve tradeoffs between fairness and accuracy and between different fairness criteria, but the nature of these tradeoffs varies. Recent work seeks to characterize these tradeoffs in specific…

Machine Learning · Statistics 2021-09-02 Alan Mishler , Edward Kennedy

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a…

Sound · Computer Science 2025-05-30 Zhaokai Sun , Li Zhang , Qing Wang , Pan Zhou , Lei Xie

State-of-the-art transformer models for Speech Emotion Recognition (SER) rely on temporal feature aggregation, yet advanced pooling methods remain underexplored. We systematically benchmark pooling strategies, including Multi-Query…

Detecting fake news in large datasets is challenging due to its diversity and complexity, with traditional approaches often focusing on textual features while underutilizing semantic and emotional elements. Current methods also rely heavily…

Computation and Language · Computer Science 2024-10-22 Xiaoman Xu , Xiangrun Li , Taihang Wang , Ye Jiang

With the availability of diverse sensor modalities (i.e., RGB, Depth, Infrared) and the success of multi-modal learning, multi-modal face anti-spoofing (FAS) has emerged as a prominent research focus. The intuition behind it is that…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Jingyi Yang , Xun Lin , Zitong Yu , Liepiao Zhang , Xin Liu , Hui Li , Xiaochen Yuan , Xiaochun Cao

Multimodal Emotion Recognition in Conversation (MERC) significantly enhances emotion recognition performance by integrating complementary emotional cues from text, audio, and visual modalities. While existing methods commonly utilize…

Multimedia · Computer Science 2026-02-12 Xinyi Che , Wenbo Wang , Jian Guan , Qijun Zhao

Domain Generalizable Face Anti-Spoofing (DGFAS) methods effectively capture domain-invariant features by aligning the directions (weights) of local decision boundaries across domains. However, the bias terms associated with these boundaries…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Seungjin Jung , Kanghee Lee , Yonghyun Jeong , Haeun Noh , Jungmin Lee , Jongwon Choi

Oral cancer is frequently diagnosed at later stages due to its similarity to other lesions. Existing research on computer aided diagnosis has made progress using deep learning; however, most approaches remain limited by small, imbalanced…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Joy Naoum , Revana Salama , Ali Hamdi

Recently, more and more personalized speech enhancement systems (PSE) with excellent performance have been proposed. However, two critical issues still limit the performance and generalization ability of the model: 1) Acoustic environment…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-23 Xiaofeng Ge , Jiangyu Han , Haixin Guan , Yanhua Long

Lack of large, well-annotated emotional speech corpora continues to limit the performance and robustness of speech emotion recognition (SER), particularly as models grow more complex and the demand for multimodal systems increases. While…

Sound · Computer Science 2026-02-13 Chung-Soo Ahn , Rajib Rana , Sunil Sivadas , Carlos Busso , Jagath C. Rajapakse

Meetings are a common activity in professional contexts, and it remains challenging to endow vocal assistants with advanced functionalities to facilitate meeting management. In this context, a task like active speaker detection can provide…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Lionel Pibre , Francisco Madrigal , Cyrille Equoy , Frédéric Lerasle , Thomas Pellegrini , Julien Pinquier , Isabelle Ferrané

In this paper, we present methods in deep multimodal learning for fusing speech and visual modalities for Audio-Visual Automatic Speech Recognition (AV-ASR). First, we study an approach where uni-modal deep networks are trained separately…

Computation and Language · Computer Science 2015-01-23 Youssef Mroueh , Etienne Marcheret , Vaibhava Goel

In this paper, we present an Adaptive Ensemble Learning framework that aims to boost the performance of deep neural networks by intelligently fusing features through ensemble learning techniques. The proposed framework integrates ensemble…

Artificial Intelligence · Computer Science 2023-04-07 Neelesh Mungoli

Despite recent advances in text-to-speech (TTS) models, audio-visual-to-audio-visual (AV2AV) translation still faces a critical challenge: maintaining speaker consistency between the original and translated vocal and facial features. To…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-31 Sungwoo Cho , Jeongsoo Choi , Sungnyun Kim , Se-Young Yun

Audio-Language Models (ALMs) are making strides in understanding speech and non-speech audio. However, domain-specialist Foundation Models (FMs) remain the best for closed-ended speech processing tasks such as Speech Emotion Recognition…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-25 Saurabh Kataria , Xiao Hu

Current video generation models struggle with identity preservation under large face poses, primarily facing two challenges: the difficulty in exploring an effective mechanism to integrate identity features into DiT architectures, and the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yuji Wang , Moran Li , Xiaobin Hu , Ran Yi , Jiangning Zhang , Chengming Xu , Weijian Cao , Yabiao Wang , Chengjie Wang , Lizhuang Ma

The field of natural language processing (NLP) has made significant strides in recent years, particularly in the development of large-scale vision-language models (VLMs). These models aim to bridge the gap between text and visual…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Sheng Shen , Zhewei Yao , Chunyuan Li , Trevor Darrell , Kurt Keutzer , Yuxiong He

Automatic Mean Opinion Score (MOS) prediction is crucial to evaluate the perceptual quality of the synthetic speech. While recent approaches using pre-trained self-supervised learning (SSL) models have shown promising results, they only…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-01 Hui Wang , Shiwan Zhao , Xiguang Zheng , Yong Qin

The performance of the keyword spotting (KWS) system based on audio modality, commonly measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. Therefore, audio-visual keyword spotting,…

Sound · Computer Science 2023-03-15 Ao Zhang , He Wang , Pengcheng Guo , Yihui Fu , Lei Xie , Yingying Gao , Shilei Zhang , Junlan Feng