中文
相关论文

相关论文: DiMoDif: Discourse Modality-information Differenti…

200 篇论文

The rise of AI-driven generative models has enabled the creation of highly realistic speech deepfakes - synthetic audio signals that can imitate target speakers' voices - raising critical security concerns. Existing methods for detecting…

声音 · 计算机科学 2025-03-25 Emma Coletta , Davide Salvi , Viola Negroni , Daniele Ugo Leonzio , Paolo Bestagini

Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to perturbations poses a significant threat to their reliability in real-world applications. Despite often being…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Jia Fu , Yongtao Wu , Yihang Chen , Kunyu Peng , Xiao Zhang , Volkan Cevher , Sepideh Pashami , Anders Holst

The recent renaissance in generative models, driven primarily by the advent of diffusion models and iterative improvement in GAN methods, has enabled many creative applications. However, each advancement is also accompanied by a rise in the…

计算机视觉与模式识别 · 计算机科学 2023-08-28 Sanjay Saha , Rashindrie Perera , Sachith Seneviratne , Tamasha Malepathirana , Sanka Rasnayaka , Deshani Geethika , Terence Sim , Saman Halgamuge

This paper presents DiffMoog - a differentiable modular synthesizer with a comprehensive set of modules typically found in commercial instruments. Being differentiable, it allows integration into neural networks, enabling automated sound…

音频与语音处理 · 电气工程与系统科学 2024-01-24 Noy Uzrad , Oren Barkan , Almog Elharar , Shlomi Shvartzman , Moshe Laufer , Lior Wolf , Noam Koenigstein

The proliferation of multi-modal fake news on social media poses a significant threat to public trust and social stability. Traditional detection methods, primarily text-based, often fall short due to the deceptive interplay between…

密码学与安全 · 计算机科学 2025-08-11 Junhao He , Tianyu Liu , Jingyuan Zhao , Benjamin Turner

Advancements in artificial intelligence and machine learning have significantly improved synthetic speech generation. This paper explores diffusion models, a novel method for creating realistic synthetic speech. We create a diffusion…

密码学与安全 · 计算机科学 2025-01-15 Anton Firc , Kamil Malinka , Petr Hanáček

Deepfake generation methods are evolving fast, making fake media harder to detect and raising serious societal concerns. Most deepfake detection and dataset creation research focuses on monolingual content, often overlooking the challenges…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Kartik Kuckreja , Parul Gupta , Injy Hamed , Thamar Solorio , Muhammad Haris Khan , Abhinav Dhall

The rapid advancement in deep learning makes the differentiation of authentic and manipulated facial images and video clips unprecedentedly harder. The underlying technology of manipulating facial appearances through deep generative…

计算机视觉与模式识别 · 计算机科学 2021-09-08 Sm Zobaed , Md Fazle Rabby , Md Istiaq Hossain , Ekram Hossain , Sazib Hasan , Asif Karim , Khan Md. Hasib

Most existing speech disfluency detection techniques only rely upon acoustic data. In this work, we present a practical multimodal disfluency detection approach that leverages available video data together with audio. We curate an…

计算与语言 · 计算机科学 2024-06-12 Payal Mohapatra , Shamika Likhite , Subrata Biswas , Bashima Islam , Qi Zhu

Deepfake techniques have been widely used for malicious purposes, prompting extensive research interest in developing Deepfake detection methods. Deepfake manipulations typically involve tampering with facial parts, which can result in…

多媒体 · 计算机科学 2023-05-11 Juan Hu , Xin Liao , Difei Gao , Satoshi Tsutsui , Qian Wang , Zheng Qin , Mike Zheng Shou

Deepfake techniques have been widely used for malicious purposes, prompting extensive research interest in developing Deepfake detection methods. Deepfake manipulations typically involve tampering with facial parts, which can result in…

计算机视觉与模式识别 · 计算机科学 2023-05-09 Juan Hu , Xin Liao , Difei Gao , Satoshi Tsutsui , Qian Wang , Zheng Qin , Mike Zheng Shou

Deepfake technology utilizes deep learning based face manipulation techniques to seamlessly replace faces in videos creating highly realistic but artificially generated content. Although this technology has beneficial applications in media…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Siddiqui Muhammad Yasir , Hyun Kim

This research explores the positive application of deepfake technology for upper body generation, specifically sign language for the Deaf and Hard of Hearing (DHoH) community. Given the complexity of sign language and the scarcity of…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Shahzeb Naeem , Muhammad Riyyan Khan , Usman Tariq , Abhinav Dhall , Carlos Ivan Colon , Hasan Al-Nashash

Recent advances in AIGC have exacerbated the misuse of malicious deepfake content, making the development of reliable deepfake detection methods an essential means to address this challenge. Although existing deepfake detection models…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Changtao Miao , Yi Zhang , Weize Gao , Zhiya Tan , Weiwei Feng , Man Luo , Jianshu Li , Ajian Liu , Yunfeng Diao , Qi Chu , Tao Gong , Zhe Li , Weibin Yao , Joey Tianyi Zhou

Text-to-Speech (TTS) and Voice Conversion (VC) models have exhibited remarkable performance in generating realistic and natural audio. However, their dark side, audio deepfake poses a significant threat to both society and individuals.…

密码学与安全 · 计算机科学 2024-09-17 Xinfeng Li , Kai Li , Yifan Zheng , Chen Yan , Xiaoyu Ji , Wenyuan Xu

A lip-syncing deepfake is a digitally manipulated video in which a person's lip movements are created convincingly using AI models to match altered or entirely new audio. Lip-syncing deepfakes are a dangerous type of deepfakes as the…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Soumyya Kanti Datta , Shan Jia , Siwei Lyu

Partial deepfake speech detection requires identifying manipulated regions that may occur within short temporal portions of an otherwise bona fide utterance, making the task particularly challenging for conventional utterance-level…

声音 · 计算机科学 2026-04-06 Inbal Rimon , Oren Gal , Haim Permuter

Visuomotor imitation learning policies enable robots to efficiently acquire manipulation skills from visual demonstrations. However, as scene complexity and visual distractions increase, policies that perform well in simple settings often…

Better generative models and larger datasets have led to more realistic fake videos that can fool the human eye but produce temporal and spatial artifacts that deep learning approaches can detect. Most current Deepfake detection methods…

计算机视觉与模式识别 · 计算机科学 2020-06-29 Oscar de Lima , Sean Franklin , Shreshtha Basu , Blake Karwoski , Annet George

As multimodal misinformation becomes more sophisticated, its detection and grounding are crucial. However, current multimodal verification methods, relying on passive holistic fusion, struggle with sophisticated misinformation. Due to…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Zizhao Chen , Ping Wei , Ziyang Ren , Huan Li , Xiangru Yin