中文
相关论文

相关论文: Few-Shot Speech Deepfake Detection Adaptation with…

200 篇论文

Personalizing a speech synthesis system is a highly desired application, where the system can generate speech with the user's voice with rare enrolled recordings. There are two main approaches to build such a system in recent works: speaker…

声音 · 计算机科学 2022-08-01 Sung-Feng Huang , Chyi-Jiunn Lin , Da-Rong Liu , Yi-Chen Chen , Hung-yi Lee

3D Gaussian Splatting (3DGS) has become one of the most influential works in the past year. Due to its efficient and high-quality novel view synthesis capabilities, it has been widely adopted in many research fields and applications.…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Glenn Grubert , Florian Barthel , Anna Hilsmann , Peter Eisert

Deep neural networks often encounter significant performance drops while facing with domain shifts between training (source) and test (target) data. To address this issue, Test Time Adaptation (TTA) methods have been proposed to adapt…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Siqi Luo , Yi Xin , Yuntao Du , Tao Tan , Guangtao Zhai , Xiaohong Liu

Generalizability, the capacity of a robust model to perform effectively on unseen data, is crucial for audio deepfake detection due to the rapid evolution of text-to-speech (TTS) and voice conversion (VC) technologies. A promising approach…

声音 · 计算机科学 2025-04-16 Botao Zhao , Zuheng Kang , Yayun He , Xiaoyang Qu , Junqing Peng , Jing Xiao , Jianzong Wang

Audio Deepfake Detection (ADD) aims to detect the fake audio generated by text-to-speech (TTS), voice conversion (VC) and replay, etc., which is an emerging topic. Traditionally we take the mono signal as input and focus on robust feature…

声音 · 计算机科学 2023-05-29 Rui Liu , Jinhua Zhang , Guanglai Gao , Haizhou Li

Utilizing the large-scale unlabeled data from the target domain via pseudo-label clustering algorithms is an important approach for addressing domain adaptation problems in speaker verification tasks. In this paper, we propose a novel…

声音 · 计算机科学 2023-05-23 Zhuo Li , Jingze Lu , Zhenduo Zhao , Wenchao Wang , Pengyuan Zhang

Recently, synthesizing personalized speech by text-to-speech (TTS) application is highly demanded. But the previous TTS models require a mass of target speaker speeches for training. It is a high-cost task, and hard to record lots of…

声音 · 计算机科学 2022-05-25 Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

Generative speech enhancement (GSE) models show great promise in producing high-quality clean speech from noisy inputs, enabling applications such as curating noisy text-to-speech (TTS) datasets into high-quality ones. However, GSE models…

声音 · 计算机科学 2026-01-21 Kazuki Yamauchi , Masato Murata , Shogo Seki

The rapid advancement of deepfake generation techniques poses significant threats to public safety and causes societal harm through the creation of highly realistic synthetic facial media. While existing detection methods demonstrate…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Jianfeng Liao , Yichen Wei , Raymond Chan Ching Bon , Shulan Wang , Kam-Pui Chow , Kwok-Yan Lam

Existing approaches towards anomaly detection~(AD) often rely on a substantial amount of anomaly-free data to train representation and density models. However, large anomaly-free datasets may not always be available before the inference…

计算机视觉与模式识别 · 计算机科学 2024-03-01 Jingyi Liao , Xun Xu , Manh Cuong Nguyen , Adam Goodge , Chuan Sheng Foo

High-quality and intelligible speech is essential to text-to-speech (TTS) model training, however, obtaining high-quality data for low-resource languages is challenging and expensive. Applying speech enhancement on Automatic Speech…

音频与语音处理 · 电气工程与系统科学 2023-09-20 Zhaoheng Ni , Sravya Popuri , Ning Dong , Kohei Saijo , Xiaohui Zhang , Gael Le Lan , Yangyang Shi , Vikas Chandra , Changhan Wang

Existing face forgery detection usually follows the paradigm of training models in a single domain, which leads to limited generalization capacity when unseen scenarios and unknown attacks occur. In this paper, we elaborately investigate…

计算机视觉与模式识别 · 计算机科学 2024-07-01 Yingxin Lai , Zitong Yu , Jing Yang , Bin Li , Xiangui Kang , Linlin Shen

Most existing methods for audio classification assume that the vocabulary of audio classes to be classified is fixed. When novel (unseen) audio classes appear, audio classification systems need to be retrained with abundant labeled samples…

音频与语音处理 · 电气工程与系统科学 2023-06-01 Yanxiong Li , Wenchang Cao , Wei Xie , Jialong Li , Emmanouil Benetos

The advancements in generative AI have enabled the improvement of audio synthesis models, including text-to-speech and voice conversion. This raises concerns about its potential misuse in social manipulation and political interference, as…

密码学与安全 · 计算机科学 2024-09-12 Hong-Hanh Nguyen-Le , Van-Tuan Tran , Dinh-Thuc Nguyen , Nhien-An Le-Khac

Since the majority of audio DeepFake (DF) detection methods are trained on English-centric datasets, their applicability to non-English languages remains largely unexplored. In this work, we present a benchmark for the multilingual audio DF…

声音 · 计算机科学 2025-08-07 Bartłomiej Marek , Piotr Kawa , Piotr Syga

Deep Gaussian Processes (DGP) are hierarchical generalizations of Gaussian Processes (GP) that have proven to work effectively on a multiple supervised regression tasks. They combine the well calibrated uncertainty estimates of GPs with the…

We propose a few-shot learning method for spatial regression. Although Gaussian processes (GPs) have been successfully used for spatial regression, they require many observations in the target task to achieve a high predictive performance.…

机器学习 · 统计学 2020-10-12 Tomoharu Iwata , Yusuke Tanaka

Automatic speaker verification (ASV) technology is recently finding its way to end-user applications for secure access to personal data, smart services or physical facilities. Similar to other biometric technologies, speaker verification is…

声音 · 计算机科学 2016-09-16 Cemal Hanilci , Tomi Kinnunen , Md Sahidullah , Aleksandr Sizov

We develop Bayesian machine learning methods for mixed data sampling (MIDAS) regressions. This involves handling frequency mismatches and specifying functional relationships between many predictors and the dependent variable. We use…

计量经济学 · 经济学 2024-09-11 Niko Hauzenberger , Massimiliano Marcellino , Michael Pfarrhofer , Anna Stelzer

The transduction of sequence has been mostly done by recurrent networks, which are computationally demanding and often underestimate uncertainty severely. We propose a computationally efficient attention-based network combined with the…

机器学习 · 计算机科学 2021-02-11 Kuilin Chen , Chi-Guhn Lee