English
Related papers

Related papers: GateFusion: Hierarchical Gated Cross-Modal Fusion …

200 papers

Exploring proper way to conduct multi-speech feature fusion for cross-corpus speech emotion recognition is crucial as different speech features could provide complementary cues reflecting human emotion status. While most previous approaches…

Sound · Computer Science 2024-06-14 Xueyu Liu , Jie Lin , Chao Wang

The use of deep neural networks (DNN) has dramatically elevated the performance of automatic speaker verification (ASV) over the last decade. However, ASV systems can be easily neutralized by spoofing attacks. Therefore, the Spoofing-Aware…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-14 Jungwoo Heo , Ju-ho Kim , Hyun-seo Shin

Fusing multiple modalities has proven effective for multimodal information processing. However, the incongruity between modalities poses a challenge for multimodal fusion, especially in affect recognition. In this study, we first analyze…

Computation and Language · Computer Science 2023-11-14 Yaoting Wang , Yuanchao Li , Paul Pu Liang , Louis-Philippe Morency , Peter Bell , Catherine Lai

Multi-Agent Debate (MAD) is a collaborative framework in which multiple agents iteratively refine solutions through the generation of reasoning and alternating critique cycles. Current work primarily optimizes intra-round topologies and…

Multiagent Systems · Computer Science 2026-04-14 Yiqing Liu , Hantao Yao , Wu Liu , Allen He , Yongdong Zhang

Current automated speaking assessment (ASA) systems for use in multi-aspect evaluations often fail to make full use of content relevance, overlooking image or exemplar cues, and employ superficial grammar analysis that lacks detailed error…

Computation and Language · Computer Science 2025-06-23 Hao-Chien Lu , Jhen-Ke Lin , Hong-Yun Lin , Chung-Chun Wang , Berlin Chen

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a…

Sound · Computer Science 2025-05-30 Zhaokai Sun , Li Zhang , Qing Wang , Pan Zhou , Lei Xie

Evaluating AI generated dubbed content is inherently multi-dimensional, shaped by synchronization, intelligibility, speaker consistency, emotional alignment, and semantic context. Human Mean Opinion Scores (MOS) remain the gold standard but…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-27 Ashwini Dasare , Nirmesh Shah , Ashishkumar Gudmalwar , Pankaj Wasnik

Depression, a common mental disorder, significantly influences individuals and imposes considerable societal impacts. The complexity and heterogeneity of the disorder necessitate prompt and effective detection, which nonetheless, poses a…

Sound · Computer Science 2023-08-25 Xiao Xu , Yang Wang , Xinru Wei , Fei Wang , Xizhe Zhang

We address the Ambivalence/Hesitancy (A/H) Video Recognition Challenge at the 10th ABAW Competition (CVPR 2026). We propose a divergence-based multimodal fusion that explicitly measures cross-modal conflict between visual, audio, and…

Hate speech in online videos is posing an increasingly serious threat to digital platforms, especially as video content becomes increasingly multimodal and context-dependent. Existing methods often struggle to effectively fuse the complex…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Shuonan Yang , Tailin Chen , Jiangbei Yue , Guangliang Cheng , Jianbo Jiao , Zeyu Fu

This paper presents a novel framework for joint speaker diarization (SD) and automatic speech recognition (ASR), named SLIDAR (sliding-window diarization-augmented recognition). SLIDAR can process arbitrary length inputs and can handle any…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-04 Samuele Cornell , Jee-weon Jung , Shinji Watanabe , Stefano Squartini

Auditory attention detection (AAD) aims to detect the target speaker in a multi-talker environment from brain signals, such as electroencephalography (EEG), which has made great progress. However, most AAD methods solely utilize attention…

Human-Computer Interaction · Computer Science 2025-05-22 Lu Li , Cunhang Fan , Hongyu Zhang , Jingjing Zhang , Xiaoke Yang , Jian Zhou , Zhao Lv

Speech-driven facial animation requires accurate correspondence between acoustic signals and facial motion, especially for articulation-related mouth movements. However, directly mapping speech audio to facial coefficients often overlooks…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Kai Zheng , Zejian Kang , Rui Mao , Hongyuan Zou , Yuanchen Fei , Xuanyang Xu , Xiangru Huang

Semantic segmentation generates comprehensive understanding of scenes through densely predicting the category for each pixel. High-level features from Deep Convolutional Neural Networks already demonstrate their effectiveness in semantic…

Computer Vision and Pattern Recognition · Computer Science 2020-02-25 Xiangtai Li , Houlong Zhao , Lei Han , Yunhai Tong , Kuiyuan Yang

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which…

Multimedia · Computer Science 2025-08-12 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Xinyi Yin , Danlei Huang , Fei Yu

As a crucial element of public security, video anomaly detection (VAD) aims to measure deviations from normal patterns for various events in real-time surveillance systems. However, most existing VAD methods rely on large-scale models to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Jiahao Lyu , Minghua Zhao , Xuewen Huang , Yifei Chen , Shuangli Du , Jing Hu , Cheng Shi , Zhiyong Lv

Deploying emotion recognition systems in real-world environments where devices must be small, low-power, and private remains a significant challenge. This is especially relevant for applications such as tension monitoring, conflict…

Speech activity detection (SAD) plays an important role in current speech processing systems, including automatic speech recognition (ASR). SAD is particularly difficult in environments with acoustic noise. A practical solution is to…

Computation and Language · Computer Science 2023-05-15 Fei Tao , Carlos Busso

Multimodal analysis has recently drawn much interest in affective computing, since it can improve the overall accuracy of emotion recognition over isolated uni-modal approaches. The most effective techniques for multimodal emotion…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 R. Gnana Praveen , Eric Granger , Patrick Cardinal

The adoption of multimodal interactions by Voice Assistants (VAs) is growing rapidly to enhance human-computer interactions. Smartwatches have now incorporated trigger-less methods of invoking VAs, such as Raise To Speak (RTS), where the…