中文
相关论文

相关论文: ICLAD: In-Context Learning with Comparison-Guidanc…

200 篇论文

This paper proposes an audio-visual deepfake detection approach that aims to capture fine-grained temporal inconsistencies between audio and visual modalities. To achieve this, both architectural and data synthesis strategies are…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Marcella Astrid , Enjie Ghorbel , Djamila Aouada

Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the SSL…

声音 · 计算机科学 2025-06-17 Tony Alex , Sara Ahmed , Armin Mustafa , Muhammad Awais , Philip JB Jackson

Most deepfake detection methods focus on detecting spatial and/or spatio-temporal changes in facial attributes and are centered around the binary classification task of detecting whether a video is real or fake. This is because available…

计算机视觉与模式识别 · 计算机科学 2023-07-18 Zhixi Cai , Shreya Ghosh , Abhinav Dhall , Tom Gedeon , Kalin Stefanov , Munawar Hayat

Deepfakes are synthetic media generated using deep generative algorithms and have posed a severe societal and political threat. Apart from facial manipulation and synthetic voice, recently, a novel kind of deepfakes has emerged with either…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Vinaya Sree Katamneni , Ajita Rattani

In this paper, we propose an enhanced audio-visual deep detection method. Recent methods in audio-visual deepfake detection mostly assess the synchronization between audio and visual features. Although they have shown promising results,…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Marcella Astrid , Enjie Ghorbel , Djamila Aouada

Speech deepfake detection has achieved remarkable success in clean environments but faces significant challenges in complex, real-world scenarios where speech is often mixed with background music or noise. Current state-of-the-art methods…

声音 · 计算机科学 2026-05-25 Qingcao Li , Yipeng Lin , Weichen Lian , Zhongjie Ba , Peng Cheng , Zhichao Lian

This paper presents the Multi-Language Audio Anti-Spoofing Dataset (MLAAD), version 10: a dataset of synthetic audio to train and evaluate audio deepfake detection models. It features 175 Text-to-Speech (TTS) models, comprising a total of…

This paper proposes a method for unsupervised anomalous sound detection (UASD) and captioning the reason for detection. While there is a method that captions the difference between given normal and anomalous sound pairs, it is assumed to be…

音频与语音处理 · 电气工程与系统科学 2024-10-30 Ryoya Ogura , Tomoya Nishida , Yohei Kawaguchi

Detecting Alzheimer's Disease (AD) from narrative transcripts remains a challenging task for large language models (LLMs), particularly under out-of-distribution (OOD) and data-scarce conditions. While in-context learning (ICL) provides a…

计算与语言 · 计算机科学 2025-11-11 Puzhen Su , Yongzhu Miao , Chunxi Guo , Jintao Tang , Shasha Li , Ting Wang

Fake audio attack becomes a major threat to the speaker verification system. Although current detection approaches have achieved promising results on dataset-specific scenarios, they encounter difficulties on unseen spoofing data.…

声音 · 计算机科学 2022-07-12 Haoxin Ma , Jiangyan Yi , Jianhua Tao , Ye Bai , Zhengkun Tian , Chenglong Wang

While the technologies empowering malicious audio deepfakes have dramatically evolved in recent years due to generative AI advances, the same cannot be said of global research into spoofing (deepfake) countermeasures. This paper highlights…

音频与语音处理 · 电气工程与系统科学 2026-03-16 Héctor Delgado , Giorgio Ramondetti , Emanuele Dalmasso , Gennady Karvitsky , Daniele Colibro , Haydar Talib

Audio Language Models (ALMs) have recently shown strong capabilities in unified reasoning over speech, sound, and natural language; yet they inherit behavioral issues observed in Large Language Models, including sycophancy--the tendency to…

Current DeepFake detection scenarios are mostly binary, yet data manipulation can vary across audio, video, or both, whose variability is not captured in binary settings. Four-class audio-visual formulations address this by discriminating…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Sharayu Nilesh Deshmukh , Kailash A. Hambarde , Joana C. Costa , Hugo Proença , Tiago Roxo

Autoregressive (AR) large audio language models (LALMs) such as Qwen-2.5-Omni have achieved strong performance on audio understanding and interaction, but scaling them remains costly in data and computation, and strictly sequential decoding…

声音 · 计算机科学 2026-02-02 Jiaming Zhou , Xuxin Cheng , Shiwan Zhao , Yuhang Jia , Cao Liu , Ke Zeng , Xunliang Cai , Yong Qin

Advances in Generative AI have made video-level deepfake detection increasingly challenging, exposing the limitations of current detection techniques. In this paper, we present HOLA, our solution to the Video-Level Deepfake Detection track…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Xuecheng Wu , Danlei Huang , Heli Sun , Xinyi Yin , Yifan Wang , Hao Wang , Jia Zhang , Fei Wang , Peihao Guo , Suyu Xing , Junxiao Xue , Liang He

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a…

声音 · 计算机科学 2025-05-30 Zhaokai Sun , Li Zhang , Qing Wang , Pan Zhou , Lei Xie

The state-of-the-art audio deepfake detectors leveraging deep neural networks exhibit impressive recognition performance. Nonetheless, this advantage is accompanied by a significant carbon footprint. This is mainly due to the use of…

声音 · 计算机科学 2024-03-22 Subhajit Saha , Md Sahidullah , Swagatam Das

Detecting forgery videos is highly desirable due to the abuse of deepfake. Existing detection approaches contribute to exploring the specific artifacts in deepfake videos and fit well on certain data. However, the growing technique on these…

计算机视觉与模式识别 · 计算机科学 2022-06-14 Harry Cheng , Yangyang Guo , Tianyi Wang , Qi Li , Xiaojun Chang , Liqiang Nie

Significant advancements made in the generation of deepfakes have caused security and privacy issues. Attackers can easily impersonate a person's identity in an image by replacing his face with the target person's face. Moreover, a new…

计算机视觉与模式识别 · 计算机科学 2021-09-08 Hasam Khalid , Minha Kim , Shahroz Tariq , Simon S. Woo

Deepfakes are AI-generated media in which an image or video has been digitally modified. The advancements made in deepfake technology have led to privacy and security issues. Most deepfake detection techniques rely on the detection of a…

计算机视觉与模式识别 · 计算机科学 2023-10-09 Sneha Muppalla , Shan Jia , Siwei Lyu