English
Related papers

Related papers: Towards Robust Speech Deepfake Detection via Human…

200 papers

Infants acquire words and phonemes from unsegmented speech signals using segmentation cues, such as distributional, prosodic, and co-occurrence cues. Many pre-existing computational models that represent the process tend to focus on…

Computation and Language · Computer Science 2023-01-18 Yasuaki Okuda , Ryo Ozaki , Tadahiro Taniguchi

Local explanation methods such as LIME (Ribeiro et al., 2016) remain fundamental to trustworthy AI, yet their application to NLP is limited by a reliance on random token masking. These heuristic perturbations frequently generate…

Computation and Language · Computer Science 2026-01-21 George Mihaila , Suleyman Olcay Polat , Poli Nemkova , Himanshu Sharma , Namratha V. Urs , Mark V. Albert

Recent advances in audio large language models (ALLMs) have made high-quality synthetic audio widely accessible, increasing the risk of malicious audio deepfakes across speech, environmental sounds, singing voice, and music. Real-world…

Sound · Computer Science 2026-01-07 Yuankun Xie , Xiaoxuan Guo , Jiayi Zhou , Tao Wang , Jian Liu , Ruibo Fu , Xiaopeng Wang , Haonan Cheng , Long Ye

Speaker verification systems are increasingly deployed in security-sensitive applications but remain highly vulnerable to adversarial perturbations. In this work, we propose the Mask Diffusion Detector (MDD), a novel adversarial detection…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-27 Yibo Bai , Sizhou Chen , Michele Panariello , Xiao-Lei Zhang , Massimiliano Todisco , Nicholas Evans

While supervised quality predictors for synthesized speech have demonstrated strong correlations with human ratings, their requirement for in-domain labeled training data hinders their generalization ability to new domains. Unsupervised…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-08 Erica Cooper , Takuma Okamoto , Yamato Ohtani , Tomoki Toda , Hisashi Kawai

Speech deepfake detection has achieved remarkable success in clean environments but faces significant challenges in complex, real-world scenarios where speech is often mixed with background music or noise. Current state-of-the-art methods…

Sound · Computer Science 2026-05-25 Qingcao Li , Yipeng Lin , Weichen Lian , Zhongjie Ba , Peng Cheng , Zhichao Lian

The advancements of AI-synthesized human voices have introduced a growing threat of impersonation and disinformation. It is therefore of practical importance to developdetection methods for synthetic human voices. This work proposes a new…

Sound · Computer Science 2023-04-28 Chengzhe Sun , Shan Jia , Shuwei Hou , Ehab AlBadawy , Siwei Lyu

As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited…

Sound · Computer Science 2025-07-30 Wen Huang , Yanmei Gu , Zhiming Wang , Huijia Zhu , Yanmin Qian

Dialogue serves as the most natural manner of human-computer interaction (HCI). Recent advancements in speech language models (SLM) have significantly enhanced speech-based conversational AI. However, these models are limited to turn-based…

Computation and Language · Computer Science 2024-08-06 Ziyang Ma , Yakun Song , Chenpeng Du , Jian Cong , Zhuo Chen , Yuping Wang , Yuxuan Wang , Xie Chen

Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). However, whether these models can…

Sound · Computer Science 2025-12-01 Chen Li , Peiji Yang , Yicheng Zhong , Jianxing Yu , Zhisheng Wang , Zihao Gou , Wenqing Chen , Jian Yin

Improving the accessibility of psychotherapy with the aid of Large Language Models (LLMs) is garnering a significant attention in recent years. Recognizing cognitive distortions from the interviewee's utterances can be an essential part of…

Computation and Language · Computer Science 2024-03-22 Sehee Lim , Yejin Kim , Chi-Hyun Choi , Jy-yong Sohn , Byung-Hoon Kim

Speaker recognition, recognizing speaker identities based on voice alone, enables important downstream applications, such as personalization and authentication. Learning speaker representations, in the context of supervised learning,…

Machine Learning · Computer Science 2022-07-13 Metehan Cekic , Ruirui Li , Zeya Chen , Yuguang Yang , Andreas Stolcke , Upamanyu Madhow

Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialized output training…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-09 Yuke Lin , Ming Cheng , Ze Li , Beilong Tang , Ming Li

Current DeepFake detection scenarios are mostly binary, yet data manipulation can vary across audio, video, or both, whose variability is not captured in binary settings. Four-class audio-visual formulations address this by discriminating…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Sharayu Nilesh Deshmukh , Kailash A. Hambarde , Joana C. Costa , Hugo Proença , Tiago Roxo

Recent advances in Text-to-Speech (TTS) systems have substantially increased the realism of synthetic speech, raising new challenges for audio deepfake detection. This work presents a comparative evaluation of three state-of-the-art TTS…

Research on speech processing has traditionally considered the task of designing hand-engineered acoustic features (feature engineering) as a separate distinct problem from the task of designing efficient machine learning (ML) models to…

Sound · Computer Science 2021-09-27 Siddique Latif , Rajib Rana , Sara Khalifa , Raja Jurdak , Junaid Qadir , Björn W. Schuller

The recent advancements in generative artificial speech models have made possible the generation of highly realistic speech signals. At first, it seems exciting to obtain these artificially synthesized signals such as speech clones or deep…

Sound · Computer Science 2022-03-09 Karan Bhatia , Ansh Agrawal , Priyanka Singh , Arun Kumar Singh

Speech understanding is essential for interpreting the diverse forms of information embedded in spoken language, including linguistic, paralinguistic, and non-linguistic cues that are vital for effective human-computer interaction. The…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-08 Jing Peng , Yucheng Wang , Bohan Li , Yiwei Guo , Hankun Wang , Yangui Fang , Yu Xi , Haoyu Li , Xu Li , Ke Zhang , Shuai Wang , Kai Yu

Recent developments in speech synthesis have produced systems capable of outcome intelligible speech, but now researchers strive to create models that more accurately mimic human voices. One such development is the incorporation of multiple…

Sound · Computer Science 2016-02-09 Marvin Coto-Jiménez , John Goddard-Close

With the proliferation of deepfake audio, there is an urgent need to investigate their attribution. Current source tracing methods can effectively distinguish in-distribution (ID) categories. However, the rapid evolution of deepfake…

Sound · Computer Science 2024-06-11 Yuankun Xie , Ruibo Fu , Zhengqi Wen , Zhiyong Wang , Xiaopeng Wang , Haonnan Cheng , Long Ye , Jianhua Tao
‹ Prev 1 8 9 10 Next ›