中文
相关论文

相关论文: An Audio-Visual Attention Based Multimodal Network…

200 篇论文

Deepfakes offer great potential for innovation and creativity, but they also pose significant risks to privacy, trust, and security. With a vast Hindi-speaking population, India is particularly vulnerable to deepfake-driven misinformation…

声音 · 计算机科学 2024-11-26 Sukhandeep Kaur , Mubashir Buhari , Naman Khandelwal , Priyansh Tyagi , Kiran Sharma

Deepfake media is becoming widespread nowadays because of the easily available tools and mobile apps which can generate realistic looking deepfake videos/images without requiring any technical knowledge. With further advances in this field…

计算机视觉与模式识别 · 计算机科学 2022-08-12 Sohail Ahmed Khan , Duc-Tien Dang-Nguyen

Deepfake poses a serious threat to the reliability of judicial evidence and intellectual property protection. In spite of an urgent need for Deepfake identification, existing pixel-level detection methods are increasingly unable to resist…

计算机视觉与模式识别 · 计算机科学 2021-11-01 Maoyu Mao , Jun Yang

Talking face synthesis driven by audio is one of the current research hotspots in the fields of multidimensional signal processing and multimedia. Neural Radiance Field (NeRF) has recently been brought to this research field in order to…

计算机视觉与模式识别 · 计算机科学 2024-05-17 Chongke Bi , Xiaoxing Liu , Zhilei Liu

The rapid advancement of talking-head deepfake generation fueled by advanced generative models has elevated the realism of synthetic videos to a level that poses substantial risks in domains such as media, politics, and finance. However,…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Xinqi Xiong , Prakrut Patel , Qingyuan Fan , Amisha Wadhwa , Sarathy Selvam , Xiao Guo , Luchao Qi , Xiaoming Liu , Roni Sengupta

Automatic deception detection is an important task that has gained momentum in computational linguistics due to its potential applications. In this paper, we propose a simple yet tough to beat multi-modal neural model for deception…

计算与语言 · 计算机科学 2018-03-21 Gangeshwar Krishnamurthy , Navonil Majumder , Soujanya Poria , Erik Cambria

Lip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a textdependent…

计算机视觉与模式识别 · 计算机科学 2020-10-14 Jianrong Wang , Tong Wu , Shanyu Wang , Mei Yu , Qiang Fang , Ju Zhang , Li Liu

Multimedia data, particularly images and videos, is integral to various applications, including surveillance, visual interaction, biometrics, evidence gathering, and advertising. However, amateur or skilled counterfeiters can simulate them…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Kutub Uddin , Nusrat Tasnim , Byung Tae Oh

With the rapid advancement of generative AI, multimodal deepfakes, which manipulate both audio and visual modalities, have drawn increasing public concern. Currently, deepfake detection has emerged as a crucial strategy in countering these…

声音 · 计算机科学 2024-05-16 Yang Hou , Haitao Fu , Chuankai Chen , Zida Li , Haoyu Zhang , Jianjun Zhao

Existing DeepFake detection techniques primarily focus on facial manipulations, such as face-swapping or lip-syncing. However, advancements in text-to-video (T2V) and image-to-video (I2V) generative models now allow fully AI-generated…

计算机视觉与模式识别 · 计算机科学 2025-09-04 Rohit Kundu , Hao Xiong , Vishal Mohanty , Athula Balachandran , Amit K. Roy-Chowdhury

We introduce a novel approach for high-resolution talking head generation from a single image and audio input. Prior methods using explicit face models, like 3D morphable models (3DMM) and facial landmarks, often fall short in generating…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Sejong Yang , Seoung Wug Oh , Yang Zhou , Seon Joo Kim

Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standing challenge to be further investigated. Motivated by xxx, in…

计算机视觉与模式识别 · 计算机科学 2022-04-12 Ganglai Wang , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Today's Automatic Speech Recognition systems only rely on acoustic signals and often don't perform well under noisy conditions. Performing multi-modal speech recognition - processing acoustic speech signals and lip-reading video…

计算机视觉与模式识别 · 计算机科学 2018-03-14 Matthijs Van keirsbilck , Bert Moons , Marian Verhelst

Audio DeepFakes allow the creation of high-quality, convincing utterances and therefore pose a threat due to its potential applications such as impersonation or fake news. Methods for detecting these manipulations should be characterized by…

声音 · 计算机科学 2022-10-13 Piotr Kawa , Marcin Plata , Piotr Syga

We propose an end to end deep learning approach for generating real-time facial animation from just audio. Specifically, our deep architecture employs deep bidirectional long short-term memory network and attention mechanism to discover the…

机器学习 · 计算机科学 2019-05-28 Guanzhong Tian , Yi Yuan , Yong liu

Audio-visual speech enhancement system is regarded to be one of promising solutions for isolating and enhancing speech of desired speaker. Conventional methods focus on predicting clean speech spectrum via a naive convolution neural network…

音频与语音处理 · 电气工程与系统科学 2022-09-28 Xinmeng Xu , Jianjun Hao

Talking face video generation with arbitrary speech audio is a significant challenge within the realm of digital human technology. The previous studies have emphasized the significance of audio-lip synchronization and visual quality.…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Yifan Xie , Fei Ma , Yi Bin , Ying He , Fei Yu

The rise of social media and short-form video (SFV) has facilitated a breeding ground for misinformation. With the emergence of large language models, significant research has gone into curbing this misinformation problem with automatic…

机器学习 · 计算机科学 2024-10-23 Andrew Kan , Christopher Kan , Zaid Nabulsi

Facial forgery by deepfakes has caused major security risks and raised severe societal concerns. As a countermeasure, a number of deepfake detection methods have been proposed. Most of them model deepfake detection as a binary…

计算机视觉与模式识别 · 计算机科学 2023-10-12 Aakash Varma Nadimpalli , Ajita Rattani

Automatic Speaker Verification (ASV) systems, which identify speakers based on their voice characteristics, have numerous applications, such as user authentication in financial transactions, exclusive access control in smart devices, and…