中文
相关论文

相关论文: SCDF: A Speaker Characteristics DeepFake Speech Da…

200 篇论文

Spoofing-robust speaker verification (SASV) combines the tasks of speaker and spoof detection to authenticate speakers under adversarial settings. Many SASV systems rely on fusion of speaker and spoof cues at embedding, score or decision…

音频与语音处理 · 电气工程与系统科学 2026-03-31 Oğuzhan Kurnaz , Jagabandhu Mishra , Tomi H. Kinnunen , Cemal Hanilçi

Speech signals are complex intermingling of various informative factors, and this information blending makes decoding any of the individual factors extremely difficult. A natural idea is to factorize each speech frame into independent…

声音 · 计算机科学 2017-06-27 Dong Wang , Lantian Li , Ying Shi , Yixiang Chen , Zhiyuan Tang

Speech technologies are deployed in high-stakes settings, yet fairness concerns remain fragmented across tasks and disciplines. Existing surveys either adopt a general machine-learning perspective that overlooks speech-specific properties…

音频与语音处理 · 电气工程与系统科学 2026-05-05 Yi-Cheng Lin , Yun-Shao Tsai , Kuan-Yu Chen , Hsiao-Ying Huang , Huang-Cheng Chou , Hung-yi Lee

The recent integration of generative neural strategies and audio processing techniques have fostered the widespread of synthetic speech synthesis or transformation algorithms. This capability proves to be harmful in many legal and…

声音 · 计算机科学 2022-10-07 Daniele Mari , Federica Latora , Simone Milani

Audio deepfakes pose a growing threat, already exploited in fraud and misinformation. A key challenge is ensuring detectors remain robust to unseen synthesis methods and diverse speakers, since generation techniques evolve quickly. Despite…

声音 · 计算机科学 2025-10-28 Jiyoung Hong , Yoonseo Chung , Seungyeon Oh , Juntae Kim , Jiyoung Lee , Sookyung Kim , Hyunsoo Cho

The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures (CMs) achieve promising results on public datasets, their…

声音 · 计算机科学 2025-08-15 Yuankun Xie , Ruibo Fu , Xiaopeng Wang , Zhiyong Wang , Ya Li , Zhengqi Wen , Haonnan Cheng , Long Ye

Speaker embeddings represent a means to extract representative vectorial representations from a speech signal such that the representation pertains to the speaker identity alone. The embeddings are commonly used to classify and discriminate…

音频与语音处理 · 电气工程与系统科学 2023-02-07 Adriana Stan

Recent Audio Multimodal Large Language Models (Audio MLLMs) demonstrate impressive performance on speech benchmarks, yet it remains unclear whether these models genuinely process acoustic signals or rely on text-based semantic inference. To…

人工智能 · 计算机科学 2026-03-23 Jiaqi Xiong , Yunjia Qi , Qi Cao , Yu Zheng , Yutong Zhang , Ziteng Wang , Ruofan Liao , Weisheng Xu , Sichen Liu

Semi-supervised dialogue summarization (SSDS) leverages model-generated summaries to reduce reliance on human-labeled data and improve the performance of summarization models. While addressing label noise, previous works on semi-supervised…

计算与语言 · 计算机科学 2024-03-08 Jianfeng He , Hang Su , Jason Cai , Igor Shalyminov , Hwanjun Song , Saab Mansour

Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information. This comprehensive nature of speech significantly impacts communication and is crucial for human-computer…

计算与语言 · 计算机科学 2025-01-17 Junyi Ao , Yuancheng Wang , Xiaohai Tian , Dekun Chen , Jun Zhang , Lu Lu , Yuxuan Wang , Haizhou Li , Zhizheng Wu

Speaker verification is the process by which a speakers claim of identity is tested against a claimed speaker by his or her voice. Speaker verification is done by the use of some parameters (features) from the speakers voice which can be…

声音 · 计算机科学 2019-08-16 Bhavana V. S , Pradip K. Das

Recent research has highlighted a key issue in speech deepfake detection: models trained on one set of deepfakes perform poorly on others. The question arises: is this due to the continuously improving quality of Text-to-Speech (TTS)…

声音 · 计算机科学 2024-06-13 Nicolas M. Müller , Nicholas Evans , Hemlata Tak , Philip Sperl , Konstantin Böttinger

As for other forms of AI, speech recognition has recently been examined with respect to performance disparities across different user cohorts. One approach to achieve fairness in speech recognition is to (1) identify speaker cohorts that…

Multichannel speech enhancement algorithms are essential for improving the intelligibility of speech signals in noisy environments. These algorithms are usually evaluated at the utterance level, but this approach overlooks the disparities…

声音 · 计算机科学 2025-06-24 Nasser-Eddine Monir , Paul Magron , Romain Serizel

This paper conducts a comprehensive layer-wise analysis of self-supervised learning (SSL) models for audio deepfake detection across diverse contexts, including multilingual datasets (English, Chinese, Spanish), partial, song, and…

音频与语音处理 · 电气工程与系统科学 2025-02-10 Yassine El Kheir , Youness Samih , Suraj Maharjan , Tim Polzehl , Sebastian Möller

Warning: This paper may contain texts with uncomfortable content. Large Language Models (LLMs) have achieved remarkable performance in various tasks, including those involving multimodal data like speech. However, these models often exhibit…

计算与语言 · 计算机科学 2025-05-22 Yi-Cheng Lin , Wei-Chih Chen , Hung-yi Lee

The SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is sub-optimal, which…

音频与语音处理 · 电气工程与系统科学 2019-03-26 Chenglin Xu , Wei Rao , Eng Siong Chng , Haizhou Li

Audio deepfake detection is well-studied as a binary problem, but partially manipulated speech, where a short synthesised segment is spliced into an otherwise genuine utterance, poses a harder and more realistic threat. Detecting such…

声音 · 计算机科学 2026-05-29 S. Sutharya , Remya K. Sasi

Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often undermined by evolving…

声音 · 计算机科学 2026-02-12 Qizhou Wang , Hanxun Huang , Guansong Pang , Sarah Erfani , Christopher Leckie

Recent advances in text-to-speech (TTS) systems, particularly those with voice cloning capabilities, have made voice impersonation readily accessible, raising ethical and legal concerns due to potential misuse for malicious activities like…

声音 · 计算机科学 2024-10-10 Hongbin Liu , Youzheng Chen , Arun Narayanan , Athula Balachandran , Pedro J. Moreno , Lun Wang