中文
相关论文

相关论文: HM-Conformer: A Conformer-based audio deepfake det…

200 篇论文

To train transcriptor models that produce robust results, a large and diverse labeled dataset is required. Finding such data with the necessary characteristics is a challenging task, especially for languages less popular than English.…

声音 · 计算机科学 2026-05-01 Alexandre R. Ferreira , Cláudio E. C. Campelo

Detecting forgery videos is highly desirable due to the abuse of deepfake. Existing detection approaches contribute to exploring the specific artifacts in deepfake videos and fit well on certain data. However, the growing technique on these…

计算机视觉与模式识别 · 计算机科学 2022-06-14 Harry Cheng , Yangyang Guo , Tianyi Wang , Qi Li , Xiaojun Chang , Liqiang Nie

Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest…

End-to-end approaches to anti-spoofing, especially those which operate directly upon the raw signal, are starting to be competitive with their more traditional counterparts. Until recently, all such approaches consider only the learning of…

音频与语音处理 · 电气工程与系统科学 2021-10-07 Wanying Ge , Jose Patino , Massimiliano Todisco , Nicholas Evans

Neural speech synthesis techniques have enabled highly realistic speech deepfakes, posing major security risks. Speech deepfake detection is challenging due to distribution shifts across spoofing methods and variability in speakers,…

声音 · 计算机科学 2025-09-30 Pu Huang , Shouguang Wang , Siya Yao , Mengchu Zhou

In this paper, we propose a speaker verification method by an Attentive Multi-scale Convolutional Recurrent Network (AMCRN). The proposed AMCRN can acquire both local spatial information and global sequential information from the input…

音频与语音处理 · 电气工程与系统科学 2023-06-02 Yanxiong Li , Zhongjie Jiang , Wenchang Cao , Qisheng Huang

Different from traditional sentence-level audio deepfake detection (ADD), partial audio deepfake detection (PADD) requires frame-level positioning of the location of fake speech. While some progress has been made in this area, leveraging…

计算与语言 · 计算机科学 2025-09-05 Huhong Xian , Rui Liu , Berrak Sisman , Haizhou Li

Conformer models have achieved state-of-the-art(SOTA) results in end-to-end speech recognition. However Conformer mainly focuses on temporal modeling while pays less attention on time-frequency property of speech feature. In this paper we…

音频与语音处理 · 电气工程与系统科学 2022-07-01 Yongjun Jiang , Jian Yu , Wenwen Yang , Bihong Zhang , Yanfeng Wang

Diverse promising datasets have been designed to hold back the development of fake audio detection, such as ASVspoof databases. However, previous datasets ignore an attacking situation, in which the hacker hides some small fake clips in…

声音 · 计算机科学 2023-12-19 Jiangyan Yi , Ye Bai , Jianhua Tao , Haoxin Ma , Zhengkun Tian , Chenglong Wang , Tao Wang , Ruibo Fu

To address the issue of poor generalization ability in end-to-end speech recognition models within deep learning, this study proposes a new Conformer-based speech recognition model called "Conformer-R" that incorporates the R-drop…

声音 · 计算机科学 2023-06-16 Weidong Ji , Shijie Zan , Guohui Zhou , Xu Wang

The growing prominence of the field of audio deepfake detection is driven by its wide range of applications, notably in protecting the public from potential fraud and other malicious activities, prompting the need for greater attention and…

音频与语音处理 · 电气工程与系统科学 2024-12-12 Jiangyan Yi , Chu Yuan Zhang , Jianhua Tao , Chenglong Wang , Xinrui Yan , Yong Ren , Hao Gu , Junzuo Zhou

The performance of automatic speaker verification (ASV) systems could be degraded by voice spoofing attacks. Most existing works aimed to develop standalone spoofing countermeasure (CM) systems. Relatively little work targeted at developing…

音频与语音处理 · 电气工程与系统科学 2026-02-05 You Zhang , Ge Zhu , Zhiyao Duan

This paper addresses the challenge of developing a robust audio-visual deepfake detection model. In practical use cases, new generation algorithms are continually emerging, and these algorithms are not encountered during the development of…

声音 · 计算机科学 2024-08-20 Kyungbok Lee , You Zhang , Zhiyao Duan

Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the speech models tend to be oversimplified or the inference would…

计算与语言 · 计算机科学 2018-02-16 Kaizhi Qian , Yang Zhang , Shiyu Chang , Xuesong Yang , Dinei Florencio , Mark Hasegawa-Johnson

The recently proposed Conformer architecture has shown state-of-the-art performances in Automatic Speech Recognition by combining convolution with attention to model both local and global dependencies. In this paper, we study how to reduce…

音频与语音处理 · 电气工程与系统科学 2021-09-09 Maxime Burchi , Valentin Vielzeuf

With a recent influx of voice generation methods, the threat introduced by audio DeepFake (DF) is ever-increasing. Several different detection methods have been presented as a countermeasure. Many methods are based on so-called front-ends,…

声音 · 计算机科学 2023-06-05 Piotr Kawa , Marcin Plata , Michał Czuba , Piotr Szymański , Piotr Syga

Audio deepfake detection systems trained on one dataset often fail when deployed on data from different sources due to distributional shifts in recording conditions, synthesis methods, and acoustic environments. We present a modular…

声音 · 计算机科学 2026-03-10 Urawee Thani , Gagandeep Singh , Priyanka Singh

In this paper, we propose a method called Hodge and Podge for sound event detection. We demonstrate Hodge and Podge on the dataset of Detection and Classification of Acoustic Scenes and Events (DCASE) 2019 Challenge Task 4. This task aims…

声音 · 计算机科学 2020-02-17 Ziqiang Shi , Liu Liu , Huibin Lin , Rujie Liu

Speech deepfake detection (SDD) systems perform well on standard benchmarks datasets but often fail to generalize to expressive and emotional spoofing attacks. Many methods rely on spoof-heavy training data, learning dataset-specific…

音频与语音处理 · 电气工程与系统科学 2026-04-16 Aurosweta Mahapatra , Ismail Rasim Ulgen , Kong Aik Lee , Nicholas Andrews , Berrak Sisman

Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study…

声音 · 计算机科学 2026-01-09 Prajwal Chinchmalatpure , Suyash Chinchmalatpure , Siddharth Chavan
‹ 上一页 1 8 9 10 下一页 ›