中文
相关论文

相关论文: Proactive Detection of Voice Cloning with Localize…

200 篇论文

Voice interfaces are quickly becoming a common way for people to interact with AI systems. This also brings new security risks, such as prompt injection, social engineering, and harmful voice commands. Traditional security methods rely on…

声音 · 计算机科学 2026-03-10 Sumit Ranjan , Sugandha Sharma , Ubaid Abbas , Puneeth N Ail

Recent advances in language and speech modelling have made it possible to build autonomous voice assistants that understand and generate human dialogue in real time. These systems are increasingly being deployed in domains such as customer…

人工智能 · 计算机科学 2025-09-08 Krittanon Kaewtawee , Wachiravit Modecrua , Krittin Pachtrachai , Touchapon Kraisingkorn

In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose…

音频与语音处理 · 电气工程与系统科学 2025-01-13 Han Yin , Jisheng Bai , Yang Xiao , Hui Wang , Siqi Zheng , Yafeng Chen , Rohan Kumar Das , Chong Deng , Jianfeng Chen

Given the high cost of large language model (LLM) training from scratch, safeguarding LLM intellectual property (IP) has become increasingly crucial. As the standard paradigm for IP ownership verification, LLM fingerprinting thus plays a…

密码学与安全 · 计算机科学 2026-03-20 Zixun Xiong , Gaoyi Wu , Qingyang Yu , Mingyu Derek Ma , Lingfeng Yao , Miao Pan , Xiaojiang Du , Hao Wang

The goal of LocaGen is to improve the localization performance of audio signals in the 2-D beam localization problem. LocaGen reduces sampling quantization errors through machine learning models trained on realistic synthetic data generated…

声音 · 计算机科学 2025-12-10 Ishaan Kunwar , Henry Cantor , Tyler Rizzo , Ayaan Qayyum

The rapid advancement of large language models (LLMs) has raised concerns regarding their potential misuse, particularly in generating fake news and misinformation. To address these risks, watermarking techniques for autoregressive language…

密码学与安全 · 计算机科学 2025-06-24 Koichi Nagatsuka , Terufumi Morishita , Yasuhiro Sogawa

Amid the burgeoning development of generative models like diffusion models, the task of differentiating synthesized audio from its natural counterpart grows more daunting. Deepfake detection offers a viable solution to combat this…

密码学与安全 · 计算机科学 2024-07-18 Weizhi Liu , Yue Li , Dongdong Lin , Hui Tian , Haizhou Li

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

计算与语言 · 计算机科学 2025-04-11 Lakshmipathi Balaji , Karan Singla

Watermarking enables GenAI providers to verify whether content was generated by their models. A watermark is a hidden signal in the content, whose presence can be detected using a secret watermark key. A core security threat are forgery…

密码学与安全 · 计算机科学 2026-05-12 Toluwani Aremu , Noor Hussein , Munachiso Nwadike , Samuele Poppi , Jie Zhang , Karthik Nandakumar , Neil Gong , Nils Lukas

In the speech signal, acoustic landmarks identify times when the acoustic manifestations of the linguistically motivated distinctive features are most salient. Acoustic landmarks have been widely applied in various domains, including speech…

音频与语音处理 · 电气工程与系统科学 2025-05-26 Xiangyu Zhang , Daijiao Liu , Tianyi Xiao , Cihan Xiao , Tuende Szalay , Mostafa Shahin , Beena Ahmed , Julien Epps

Current text-to-speech algorithms produce realistic fakes of human voices, making deepfake detection a much-needed area of research. While researchers have presented various techniques for detecting audio spoofs, it is often unclear exactly…

Compared with automatic speech recognition (ASR), the human auditory system is more adept at handling noise-adverse situations, including environmental noise and channel distortion. To mimic this adeptness, auditory models have been widely…

计算与语言 · 计算机科学 2016-09-16 Peng Dai , Xue Teng , Frank Rudzicz , Ing Yann Soon

Rapid advancements in generative modeling have made synthetic audio generation easy, making speech-based services vulnerable to spoofing attacks. Consequently, there is a dire need for robust countermeasures more than ever. Existing…

声音 · 计算机科学 2025-09-03 Arnab Das , Yassine El Kheir , Carlos Franzreb , Tim Herzig , Tim Polzehl , Sebastian Möller

Autoregressive next-token prediction with the Transformer decoder has become a de facto standard in large language models (LLMs), achieving remarkable success in Natural Language Processing (NLP) at scale. Extending this paradigm to audio…

音频与语音处理 · 电气工程与系统科学 2025-07-15 Shu-wen Yang , Byeonggeun Kim , Kuan-Po Huang , Qingming Tang , Huy Phan , Bo-Ru Lu , Harsha Sundar , Shalini Ghosh , Hung-yi Lee , Chieh-Chi Kao , Chao Wang

Synthetic-voice cloning technologies have seen significant advances in recent years, giving rise to a range of potential harms. From small- and large-scale financial fraud to disinformation campaigns, the need for reliable methods to…

声音 · 计算机科学 2023-09-28 Sarah Barrington , Romit Barua , Gautham Koorma , Hany Farid

The rapid progress of Generative Artificial Intelligence (GenAI) has enabled the effortless synthesis of high-quality visual content, while simultaneously raising pressing concerns about intellectual property protection, authenticity, and…

密码学与安全 · 计算机科学 2026-03-17 Jie Cao , Qi Li , Zelin Zhang , Jianbing Ni , Rongxing Lu

With the recent advances in voice synthesis, AI-synthesized fake voices are indistinguishable to human ears and widely are applied to produce realistic and natural DeepFakes, exhibiting real threats to our society. However, effective and…

音频与语音处理 · 电气工程与系统科学 2020-08-18 Run Wang , Felix Juefei-Xu , Yihao Huang , Qing Guo , Xiaofei Xie , Lei Ma , Yang Liu

Significant advancements made in the generation of deepfakes have caused security and privacy issues. Attackers can easily impersonate a person's identity in an image by replacing his face with the target person's face. Moreover, a new…

计算机视觉与模式识别 · 计算机科学 2021-09-08 Hasam Khalid , Minha Kim , Shahroz Tariq , Simon S. Woo

The indistinguishability of large language model (LLM) output from human-authored content poses significant challenges, raising concerns about potential misuse of AI-generated text and its influence on future model training. Watermarking…

密码学与安全 · 计算机科学 2026-04-16 Alexander Nemecek , Yuzhou Jiang , Erman Ayday

Audio-Visual Source Localization (AVSL) aims to localize the source of sound within a video. In this paper, we identify a significant issue in existing benchmarks: the sounding objects are often easily recognized based solely on visual…

多媒体 · 计算机科学 2024-09-12 Liangyu Chen , Zihao Yue , Boshen Xu , Qin Jin