中文
相关论文

相关论文: GLCF: A Global-Local Multimodal Coherence Analysis…

200 篇论文

Given the audio-visual clip of the speaker, facial reaction generation aims to predict the listener's facial reactions. The challenge lies in capturing the relevance between video and audio while balancing appropriateness, realism, and…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Jiaming Li , Sheng Wang , Xin Wang , Yitao Zhu , Honglin Xiong , Zixu Zhuang , Qian Wang

The rapid advancement of diffusion-based video generation models has led to increasingly realistic synthetic content, presenting new challenges for video forgery detection. Existing methods often struggle to capture fine-grained temporal…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Xi Xue , Kunio Suzuki , Nabarun Goswami , Takuya Shintate

The rapid advancement of generative adversarial networks (GANs) and diffusion models has enabled the creation of highly realistic deepfake content, posing significant threats to digital trust across audio-visual domains. While unimodal…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Chende Zheng , Ruiqi Suo , Zhoulin Ji , Jingyi Deng , Fangbin Yi , Chenhao Lin , Chao Shen

Most research efforts in the multimedia forensics domain have focused on detecting forgery audio-visual content and reached sound achievements. However, these works only consider deepfake detection as a classification task and ignore the…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Qilin Yin , Wei Lu , Xiangyang Luo , Xiaochun Cao

With the rapid advancement of speech generation technologies, the threat posed by speech deepfakes in real-time communication (RTC) scenarios has intensified. However, existing detection studies mainly focus on offline simulations and…

声音 · 计算机科学 2026-04-28 Jun Xue , Zhuolin Yi , Yihuan Huang , Yanzhen Ren , Yujie Chen , Cunhang Fan , Zicheng Su , Yonghong Zhang , Bo Cai

Synthesizing images from text descriptions has become an active research area with the advent of Generative Adversarial Networks. The main goal here is to generate photo-realistic images that are aligned with the input descriptions.…

计算机视觉与模式识别 · 计算机科学 2022-05-26 D. M. A. Ayanthi , Sarasi Munasinghe

Audio-driven talking head generation is a core component of digital avatars, and 3D Gaussian Splatting has shown strong performance in real-time rendering of high-fidelity talking heads. However, achieving precise control over fine-grained…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Shaoyang Xie , Xiaofeng Cong , Baosheng Yu , Zhipeng Gui , Jie Gui , Yuan Yan Tang , James Tin-Yau Kwok

The goal of this paper is to synthesise talking faces with controllable facial motions. To achieve this goal, we propose two key ideas. The first is to establish a canonical space where every face has the same motion patterns but different…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Youngjoon Jang , Kyeongha Rho , Jong-Bin Woo , Hyeongkeun Lee , Jihwan Park , Youshin Lim , Byeong-Yeol Kim , Joon Son Chung

Existing Text Image Forgery Localization (T-IFL) methods often suffer from poor generalization due to the limited scale of real-world datasets and the distribution gap caused by synthetic data that fails to capture the complexity of…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Zeqin Yu , Haotao Xie , Jian Zhang , Jiangqun Ni , Wenkan Su , Jiwu Huang

Multi-focus image fusion aims to generate an all-in-focus image from a sequence of partially focused input images. Existing fusion algorithms generally assume that, for every spatial location in the scene, there is at least one input image…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Xinzhe Xie , Buyu Guo , Bolin Li , Shuangyan He , Yanzhen Gu , Qingyan Jiang , Peiliang Li

We present SpeakingFaces as a publicly-available large-scale multimodal dataset developed to support machine learning research in contexts that utilize a combination of thermal, visual, and audio data streams; examples include…

This paper presents a new approach for the detection of fake videos, based on the analysis of style latent vectors and their abnormal behavior in temporal changes in the generated videos. We discovered that the generated facial videos…

计算机视觉与模式识别 · 计算机科学 2024-05-21 Jongwook Choi , Taehoon Kim , Yonghyun Jeong , Seungryul Baek , Jongwon Choi

The rapid advances in generative models have significantly lowered the barrier to producing convincing multimodal disinformation. Fabricated images and manipulated captions increasingly co-occur to create persuasive false narratives. While…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Gagandeep Singh , Samudi Amarsinghe , Priyanka Singh , Xue Li

The rise of highly convincing synthetic speech poses a growing threat to audio communications. Although existing Audio Deepfake Detection (ADD) methods have demonstrated good performance under clean conditions, their effectiveness drops…

音频与语音处理 · 电气工程与系统科学 2025-08-05 Haohan Shi , Xiyu Shi , Safak Dogan , Tianjin Huang , Yunxiao Zhang

Media forensics has attracted a lot of attention in the last years in part due to the increasing concerns around DeepFakes. Since the initial DeepFake databases from the 1st generation such as UADFV and FaceForensics++ up to the latest…

计算机视觉与模式识别 · 计算机科学 2020-11-13 Ruben Tolosana , Sergio Romero-Tapiador , Julian Fierrez , Ruben Vera-Rodriguez

Talking head video generation aims to animate a human face in a still image with dynamic poses and expressions using motion information derived from a target-driving video, while maintaining the person's identity in the source image.…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Fa-Ting Hong , Dan Xu

Deep generator technology can produce high-quality fake videos that are indistinguishable, posing a serious social threat. Traditional forgery detection methods directly centralized training on data and lacked consideration of information…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Decheng Liu , Zhan Dang , Chunlei Peng , Nannan Wang , Ruimin Hu , Xinbo Gao

Deep-learning-based technologies such as deepfakes ones have been attracting widespread attention in both society and academia, particularly ones used to synthesize forged face images. These automatic and professional-skill-free face…

计算机视觉与模式识别 · 计算机科学 2022-12-08 YuYang Sun , ZhiYong Zhang , Isao Echizen , Huy H. Nguyen , ChangZhen Qiu , Lu Sun

The emergence of artificial intelligence-generated content (AIGC) has raised concerns about the authenticity of multimedia content in various fields. However, existing research for forgery content detection has focused mainly on binary…

多媒体 · 计算机科学 2023-08-29 Rui Zhang , Hongxia Wang , Mingshan Du , Hanqing Liu , Yang Zhou , Qiang Zeng

With the continuous research on Deepfake forensics, recent studies have attempted to provide the fine-grained localization of forgeries, in addition to the coarse classification at the video-level. However, the detection and localization…

计算机视觉与模式识别 · 计算机科学 2022-10-31 Wu Haiwei , Zhou Jiantao , Zhang Shile , Tian Jinyu