中文
相关论文

相关论文: 3-Tracer: A Tri-level Temporal-Aware Framework for…

200 篇论文

We present a unified and hardware efficient architecture for two stage voice trigger detection (VTD) and false trigger mitigation (FTM) tasks. Two stage VTD systems of voice assistants can get falsely activated to audio segments…

音频与语音处理 · 电气工程与系统科学 2021-05-17 Vineet Garg , Wonil Chang , Siddharth Sigtia , Saurabh Adya , Pramod Simha , Pranay Dighe , Chandra Dhir

The widespread emergence of face-swap Deepfake videos poses growing risks to digital security, privacy, and media integrity, necessitating effective forensic tools for identifying the source of such manipulations. Although most prior…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Wasim Ahmad , Yan-Tsung Peng , Yuan-Hao Chang

We address multimodal deepfake detection requiring both robustness and interpretability by proposing FakeHunter, a unified framework that combines memory guided retrieval, a structured Observation-Thought-Action reasoning loop, and adaptive…

多媒体 · 计算机科学 2025-09-11 Chen Chen , Runze Li , Zejun Zhang , Pukun Zhao , Fanqing Zhou , Longxiang Wang , Haojian Huang

Identifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task. This paper proposes a hierarchical network with transformer encoders and memory mechanism to address this problem. The proposed…

声音 · 计算机科学 2020-11-02 Yanpei Shi , Mingjie Chen , Qiang Huang , Thomas Hain

The rapid advancement of spoofing algorithms necessitates the development of robust detection methods capable of accurately identifying emerging fake audio. Traditional approaches, such as finetuning on new datasets containing these novel…

声音 · 计算机科学 2023-06-16 Xiaohui Zhang , Jiangyan Yi , Jianhua Tao , Chenlong Wang , Le Xu , Ruibo Fu

The proliferation of sophisticated AI-generated deepfakes poses critical challenges for digital media authentication and societal security. While existing detection methods perform well within specific generative domains, they exhibit…

计算机视觉与模式识别 · 计算机科学 2025-05-26 Naseem Khan , Tuan Nguyen , Amine Bermak , Issa Khalil

With the rapid development of generation model, AI-based face manipulation technology, which called DeepFakes, has become more and more realistic. This means of face forgery can attack any target, which poses a new threat to personal…

计算机视觉与模式识别 · 计算机科学 2021-11-16 Yuyang Sun , Zhiyong Zhang , Changzhen Qiu , Liang Wang , Zekai Wang

Freely available and easy-to-use audio editing tools make it straightforward to perform audio splicing. Convincing forgeries can be created by combining various speech samples from the same person. Detection of such splices is important…

声音 · 计算机科学 2024-05-06 Denise Moussa , Germans Hirsch , Christian Riess

As a sub-field of object detection, moving infrared small target detection presents significant challenges due to tiny target sizes and low contrast against backgrounds. Currently-existing methods primarily rely on the features extracted…

计算机视觉与模式识别 · 计算机科学 2024-09-06 Weiwei Duan , Luping Ji , Shengjia Chen , Sicheng Zhu , Mao Ye

This study introduces LENS-DF, a novel and comprehensive recipe for training and evaluating audio deepfake detection and temporal localization under complicated and realistic audio conditions. The generation part of the recipe outputs…

声音 · 计算机科学 2025-07-25 Xuechen Liu , Wanying Ge , Xin Wang , Junichi Yamagishi

Many fault diagnosis methods of rotating machines are based on discriminative features extracted from signals collected from the key components such as bearings. However, under complex operating conditions, periodic impulsive…

信号处理 · 电气工程与系统科学 2025-12-12 Yuhan Yuan , Xiaomo Jiang , Haibin Yang , Haixin Zhao , Shengbo Wang , Xueyu Cheng , Jigang Meng , Shuhua Yang

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried within complex…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dragos-Alexandru Boldisor , Stefan Smeu , Dan Oneata , Elisabeta Oneata

The technological advancements of deep learning have enabled sophisticated face manipulation schemes, raising severe trust issues and security concerns in modern society. Generally speaking, detecting manipulated faces and locating the…

计算机视觉与模式识别 · 计算机科学 2022-04-07 Chenqi Kong , Baoliang Chen , Haoliang Li , Shiqi Wang , Anderson Rocha , Sam Kwong

Audio-visual deepfakes have reached a level of realism that makes perceptual detection unreliable, threatening media integrity and biometric security. While multimodal detection has shown promise, most approaches are binary classification…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Wasim Ahmad , Wei Zhang , Xuerui Mao

Detecting deepfake videos is highly challenging given the complexity of characterizing spatio-temporal artifacts. Most existing methods rely on binary classifiers trained using real and fake image sequences, therefore hindering their…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Dat Nguyen , Marcella Astrid , Anis Kacem , Enjie Ghorbel , Djamila Aouada

Phoneme boundary detection plays an essential first step for a variety of speech processing applications such as speaker diarization, speech science, keyword spotting, etc. In this work, we propose a neural architecture coupled with a…

音频与语音处理 · 电气工程与系统科学 2020-02-18 Felix Kreuk , Yaniv Sheena , Joseph Keshet , Yossi Adi

The proliferation of fake news has emerged as a severe societal problem, raising significant interest from industry and academia. While existing deep-learning based methods have made progress in detecting fake news accurately, their…

计算与语言 · 计算机科学 2024-05-29 Hui Liu , Wenya Wang , Haoru Li , Haoliang Li

Online, real-time, and fine-grained 3D segmentation constitutes a fundamental capability for embodied intelligent agents to perceive and comprehend their operational environments. Recent advancements employ predefined object queries to…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Hanshi Wang , Zijian Cai , Jin Gao , Yiwei Zhang , Weiming Hu , Ke Wang , Zhipeng Zhang

Music learners can greatly benefit from tools that accurately detect errors in their practice. Existing approaches typically compare audio recordings to music scores using heuristics or learnable models. This paper introduces LadderSym, a…

With the proliferation of Large Language Model (LLM) based deepfake audio, there is an urgent need for effective detection methods. Previous deepfake audio generation methods typically involve a multi-step generation process, with the final…