中文
相关论文

相关论文: EMO-BOOST: Emotion-Augmented Audio-Visual Features…

200 篇论文

In the latest social networks, more and more people prefer to express their emotions in videos through text, speech, and rich facial expressions. Multimodal video emotion analysis techniques can help understand users' inner world…

计算机视觉与模式识别 · 计算机科学 2022-09-22 Qinglan Wei , Xuling Huang , Yuan Zhang

As generative artificial intelligence evolves, deepfake attacks have escalated from single-modality manipulations to complex, multimodal threats. Existing forensic techniques face a severe generalization bottleneck: by relying excessively…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Jingtong Dou , Chuancheng Shi , Jian Wang , Fei Shen , Zhiyong Wang , Tat-Seng Chua

The development of technologies for easily and automatically falsifying video has raised practical questions about people's ability to detect false information online. How vulnerable are people to deepfake videos? What technologies can be…

人机交互 · 计算机科学 2023-04-11 Emilie Josephs , Camilo Fosco , Aude Oliva

Deepfake speech detection presents a growing challenge as generative audio technologies continue to advance. We propose a hybrid training framework that advances detection performance through novel augmentation strategies. First, we…

声音 · 计算机科学 2025-11-14 Inbal Rimon , Oren Gal , Haim Permuter

Speech emotion recognition (SER) remains a challenging yet crucial task due to the inherent complexity and diversity of human emotions. To address this problem, researchers attempt to fuse information from other modalities via multimodal…

声音 · 计算机科学 2024-12-10 Feng Li , Jiusong Luo , Wanjun Xia

Emotions conveyed through voice and face shape engagement and context in human AI interaction. Despite rapid progress in omni modal large language models, the holistic evaluation of emotional reasoning with audiovisual cues remains limited.…

This paper presents a system for detecting fake audio-visual content (i.e., video deepfake), developed for Track 2 of the DDL Challenge. The proposed system employs a two-stage framework, comprising unimodal detection and multimodal score…

多媒体 · 计算机科学 2026-02-03 Qingcao Li , Miao He , Liang Yi , Qing Wen , Yitao Zhang , Hongshuo Jin , Peng Cheng , Zhongjie Ba , Li Lu , Kui Ren

Existing deepfake detection methods often exhibit bias, lack transparency, and fail to capture temporal information, leading to biased decisions and unreliable results across different demographic groups. In this paper, we propose a…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Akihito Yoshii , Ryosuke Sonoda , Ramya Srinivasan

Deep learning has emerged as a powerful alternative to hand-crafted methods for emotion recognition on combined acoustic and text modalities. Baseline systems model emotion information in text and acoustic modes independently using Deep…

音频与语音处理 · 电气工程与系统科学 2020-10-13 Darshana Priyasad , Tharindu Fernando , Simon Denman , Clinton Fookes , Sridha Sridharan

Face enhancement techniques are widely used to enhance facial appearance. However, they can inadvertently distort biometric features, leading to significant decrease in the accuracy of deepfake detectors. This study hypothesizes that these…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Muhammad Saad Saeed , Ijaz Ul Haq , Khalid Malik

Modern deepfake detection models have achieved strong performance even on the challenging cross-dataset task. However, detection performance under non-ideal conditions remains very unstable, limiting success on some benchmark datasets and…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Benedikt Hopf , Radu Timofte

The widespread adoption of complex machine learning models in high-stakes domains has brought the "black-box" problem to the forefront of responsible AI research. This paper aims at addressing this issue by improving the Explainable…

机器学习 · 计算机科学 2025-12-02 Isara Liyanage , Uthayasanker Thayasivam

The advancements of AI-synthesized human voices have introduced a growing threat of impersonation and disinformation. It is therefore of practical importance to developdetection methods for synthetic human voices. This work proposes a new…

声音 · 计算机科学 2023-04-28 Chengzhe Sun , Shan Jia , Shuwei Hou , Ehab AlBadawy , Siwei Lyu

Affective Image Manipulation (AIM) seeks to modify user-provided images to evoke specific emotional responses. This task is inherently complex due to its twofold objective: significantly evoking the intended emotion, while preserving the…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Jingyuan Yang , Jiawei Feng , Weibin Luo , Dani Lischinski , Daniel Cohen-Or , Hui Huang

Deep learning models perform best with abundant, high-quality labels, yet such conditions are rarely achievable in EEG-based emotion recognition. Electroencephalogram (EEG) signals are easily corrupted by artifacts and individual…

机器学习 · 计算机科学 2025-11-20 Hyo-Jeong Jang , Hye-Bin Shin , Kang Yin

Emotion recognition has become an important field of research in Human Computer Interactions as we improve upon the techniques for modelling the various aspects of behaviour. With the advancement of technology our understanding of emotions…

人工智能 · 计算机科学 2019-11-11 Samarth Tripathi , Sarthak Tripathi , Homayoon Beigi

Face identity provides a powerful signal for deepfake detection. Prior studies show that even when not explicitly modeled, classifiers often learn identity features implicitly. This has led to conflicting views: some suppress identity cues…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Younghun Kim , Minsuk Jang , Myung-Joon Kwon , Wonjun Lee , Changick Kim

We present the first large-scale open-set benchmark for multilingual audio-video deepfake detection. Our dataset comprises over 250 hours of real and fake videos across eight languages, with 60% of data being generated. For each language,…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Florinel-Alin Croitoru , Vlad Hondru , Marius Popescu , Radu Tudor Ionescu , Fahad Shahbaz Khan , Mubarak Shah

We study universal deepfake detection. Our goal is to detect synthetic images from a range of generative AI approaches, particularly from emerging ones which are unseen during training of the deepfake detector. Universal deepfake detection…

计算机视觉与模式识别 · 计算机科学 2024-01-18 Chandler Timm Doloriel , Ngai-Man Cheung

Multimodal emotion understanding requires effective integration of text, audio, and visual modalities for both discrete emotion recognition and continuous sentiment analysis. We present EGMF, a unified framework combining expert-guided…

计算与语言 · 计算机科学 2026-01-13 Jiaqi Qiao , Xiujuan Xu , Xinran Li , Yu Liu