中文
相关论文

相关论文: Genuine-Focused Learning using Mask AutoEncoder fo…

200 篇论文

The joint training framework for speech enhancement and recognition methods have obtained quite good performances for robust end-to-end automatic speech recognition (ASR). However, these methods only utilize the enhanced feature as the…

声音 · 计算机科学 2020-11-10 Cunhang Fan , Jiangyan Yi , Jianhua Tao , Zhengkun Tian , Bin Liu , Zhengqi Wen

This paper addresses the challenge of developing a robust audio-visual deepfake detection model. In practical use cases, new generation algorithms are continually emerging, and these algorithms are not encountered during the development of…

声音 · 计算机科学 2024-08-20 Kyungbok Lee , You Zhang , Zhiyao Duan

Many existing face anti-spoofing (FAS) methods focus on modeling the decision boundaries for some predefined spoof types. However, the diversity of the spoof samples including the unknown ones hinders the effective decision boundary…

计算机视觉与模式识别 · 计算机科学 2020-05-11 Haocheng Feng , Zhibin Hong , Haixiao Yue , Yang Chen , Keyao Wang , Junyu Han , Jingtuo Liu , Errui Ding

Visual generative models (e.g., diffusion models) typically operate in compressed latent spaces to balance training efficiency and sample quality. In parallel, there has been growing interest in leveraging high-quality pre-trained visual…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Yuan Gao , Chen Chen , Tianrong Chen , Jiatao Gu

Face anti-spoofing (FAS) has lately attracted increasing attention due to its vital role in securing face recognition systems from presentation attacks (PAs). As more and more realistic PAs with novel types spring up, traditional FAS…

计算机视觉与模式识别 · 计算机科学 2022-09-05 Zitong Yu , Yunxiao Qin , Xiaobai Li , Chenxu Zhao , Zhen Lei , Guoying Zhao

With various face presentation attacks arising under unseen scenarios, face anti-spoofing (FAS) based on domain generalization (DG) has drawn growing attention due to its robustness. Most existing methods utilize DG frameworks to align the…

计算机视觉与模式识别 · 计算机科学 2021-08-06 Shubao Liu , Ke-Yue Zhang , Taiping Yao , Mingwei Bi , Shouhong Ding , Jilin Li , Feiyue Huang , Lizhuang Ma

With the rapid development of speech synthesis and voice conversion technologies, Audio Deepfake has become a serious threat to the Automatic Speaker Verification (ASV) system. Numerous countermeasures are proposed to detect this type of…

音频与语音处理 · 电气工程与系统科学 2024-01-11 Yinlin Guo , Haofan Huang , Xi Chen , He Zhao , Yuehai Wang

This paper studies a simple extension of image-based Masked Autoencoders (MAE) to self-supervised representation learning from audio spectrograms. Following the Transformer encoder-decoder design in MAE, our Audio-MAE first encodes audio…

In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE) equipped with self-attention as a generative model to…

音频与语音处理 · 电气工程与系统科学 2023-05-23 Xiao-Min Zeng , Yan Song , Zhu Zhuo , Yu Zhou , Yu-Hong Li , Hui Xue , Li-Rong Dai , Ian McLoughlin

Deepfake audio presents a growing threat to digital security, due to its potential for social engineering, fraud, and identity misuse. However, existing detection models suffer from poor generalization across datasets, due to implicit…

声音 · 计算机科学 2025-05-13 Yasaman Ahmadiadli , Xiao-Ping Zhang , Naimul Khan

Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes…

Due to the successful development of deep image generation technology, visual data forgery detection would play a more important role in social and economic security. Existing forgery detection methods suffer from unsatisfactory…

计算机视觉与模式识别 · 计算机科学 2023-07-24 Decheng Liu , Tao Chen , Chunlei Peng , Nannan Wang , Ruimin Hu , Xinbo Gao

Recent advances in synthetic speech have made audio deepfakes increasingly realistic, posing significant security risks. Existing detection methods that rely on a single modality, either raw waveform embeddings or spectral based features,…

Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio Detection (FAD), as we expect that MOS can be used to…

音频与语音处理 · 电气工程与系统科学 2024-01-26 Wangjin Zhou , Zhengdong Yang , Chenhui Chu , Sheng Li , Raj Dabre , Yi Zhao , Tatsuya Kawahara

This paper introduces a semi-supervised contrastive learning framework and its application to text-independent speaker verification. The proposed framework employs generalized contrastive loss (GCL). GCL unifies losses from two different…

音频与语音处理 · 电气工程与系统科学 2020-06-09 Nakamasa Inoue , Keita Goto

Audio deepfake detection (ADD) models are commonly evaluated using datasets that combine multiple synthesizers, with performance reported as a single Equal Error Rate (EER). However, this approach disproportionately weights synthesizers…

声音 · 计算机科学 2025-09-12 Chin Yuen Kwok , Jia Qi Yip , Zhen Qiu , Chi Hung Chi , Kwok Yan Lam

In the rapidly evolving field of self-supervised learning on graphs, generative and contrastive methodologies have emerged as two dominant approaches. Our study focuses on masked feature reconstruction (MFR), a generative technique where a…

机器学习 · 计算机科学 2025-12-16 Jianyuan Bo , Yuan Fang

Deep learning-based fault diagnosis (FD) approaches require a large amount of training data, which are difficult to obtain since they are located across different entities. Federated learning (FL) enables multiple clients to collaboratively…

机器学习 · 计算机科学 2023-10-16 Jixuan Cui , Jun Li , Zhen Mei , Kang Wei , Sha Wei , Ming Ding , Wen Chen , Song Guo

Recent advances in speech deepfake detection (SDD) have significantly improved artifacts-based detection in spoofed speech. However, most models overlook speech naturalness, a crucial cue for distinguishing bona fide speech from spoofed…

音频与语音处理 · 电气工程与系统科学 2025-05-21 Taewoo Kim , Guisik Kim , Choongsang Cho , Young Han Lee

Neural front-ends are an appealing alternative to traditional, fixed feature extraction pipelines for automatic speech recognition (ASR) systems since they can be directly trained to fit the acoustic model. However, their performance often…

音频与语音处理 · 电气工程与系统科学 2025-10-01 Peter Vieting , Maximilian Kannen , Benedikt Hilmes , Ralf Schlüter , Hermann Ney