English
Related papers

Related papers: Security and Non-Repudiation for Voice-Over-IP Con…

200 papers

Generating spoken dialogue is inherently more complex than monologue text-to-speech (TTS), as it demands both realistic turn-taking and the maintenance of distinct speaker timbres. While existing autoregressive (AR) models have made…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Han Zhu , Wei Kang , Liyong Guo , Zengwei Yao , Fangjun Kuang , Weiji Zhuang , Zhaoqing Li , Zhifeng Han , Dong Zhang , Xin Zhang , Xingchen Song , Lingxuan Ye , Long Lin , Daniel Povey

Semantic communication has emerged as a promising paradigm for enhancing communication efficiency in sixth-generation (6G) networks. However, the broadcast nature of wireless channels makes SemCom systems vulnerable to eavesdropping, which…

Information Theory · Computer Science 2025-04-01 Shunpu Tang , Yuhao Chen , Qianqian Yang , Ruichen Zhang , Dusit Niyato , Zhiguo Shi

The advancements in generative AI have enabled the improvement of audio synthesis models, including text-to-speech and voice conversion. This raises concerns about its potential misuse in social manipulation and political interference, as…

Cryptography and Security · Computer Science 2024-09-12 Hong-Hanh Nguyen-Le , Van-Tuan Tran , Dinh-Thuc Nguyen , Nhien-An Le-Khac

Voice Service Providers (VSPs) participating in VoIP peering frequently want to withhold their identity and related privacy-sensitive information from other parties during the VoIP communication. A number of existing documents on VoIP…

Networking and Internet Architecture · Computer Science 2008-07-09 Charles Shen , Henning Schulzrinne

With just a few speech samples, it is possible to perfectly replicate a speaker's voice in recent years, while malicious voice exploitation (e.g., telecom fraud for illegal financial gain) has brought huge hazards in our daily lives.…

Sound · Computer Science 2024-10-29 Zhisheng Zhang , Qianyi Yang , Derui Wang , Pengyang Huang , Yuxin Cao , Kai Ye , Jie Hao

Voice anonymization has been developed as a technique for preserving privacy by replacing the speaker's voice in a speech signal with that of a pseudo-speaker, thereby obscuring the original voice attributes from machine recognition and…

Sound · Computer Science 2024-11-13 Rui Wang , Liping Chen , Kong AiK Lee , Zhen-Hua Ling

We introduce a new paradigm for task-oriented dialogue systems: safety certification as a computational primitive for answer reuse. Current systems treat each turn independently, recomputing answers via retrieval or generation even when…

Artificial Intelligence · Computer Science 2026-03-24 Cosimo Spera

Language models now routinely produce text that is difficult to distinguish from human writing, raising the need for robust tools to verify content provenance. Watermarking has emerged as a promising countermeasure, with existing work…

Cryptography and Security · Computer Science 2026-02-18 Huijia Lin , Kameron Shahabi , Min Jae Song

In this article, the authors discuss the problem of forensic authentication of digital audio recordings. Although forensic audio has been addressed in several articles, the existing approaches are focused on analog magnetic recordings,…

Cryptography and Security · Computer Science 2022-03-15 Marcos Faundez-Zanuy , Jose Juan Lucena-Molina , Martin Hagmueller

Waku is a family of modular protocols that enable secure, censorship-resistant, and anonymous peer-to-peer communication. Waku protocols provide capabilities that make them suitable to run in resource-restricted environments e.g., mobile…

Cryptography and Security · Computer Science 2022-07-04 Oskar Thorén , Sanaz Taheri-Boshrooyeh , Hanno Cornelius

Digital signatures are the building blocks of modern communication to prevent masquerading by any party other than recipients, repudiation by signatory and forgery by any individual recipient. Digital signature scheme is said to be standard…

Quantum Physics · Physics 2015-12-31 Muhammad Nadeem , Xiaolin Wang

Audio CAPTCHAs are supposed to provide a strong defense for online resources; however, advances in speech-to-text mechanisms have rendered these defenses ineffective. Audio CAPTCHAs cannot simply be abandoned, as they are specifically named…

Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation is important, such as completing a recording, dubbing in a new language, or preserving the…

Sound · Computer Science 2026-05-28 Kaitlyn Zhou , Federico Bianchi , Martijn Bartelds , Anna Pot , Yongchan Kwon , James Zou

Audio platforms have evolved beyond entertainment. They have become central to public discourse, from podcasts and radio to WhatsApp voice notes and live streams. With millions of shows and hundreds of millions of listeners, audio platforms…

Computation and Language · Computer Science 2026-04-21 Chaewan Chun , Delvin Ce Zhang , Dongwon Lee

Spoken dialogue systems often rely on cascaded pipelines that transcribe, process, and resynthesize speech. While effective, this design discards paralinguistic cues and limits expressivity. Recent end-to-end methods reduce latency and…

The acoustic background plays a crucial role in natural conversation. It provides context and helps listeners understand the environment, but a strong background makes it difficult for listeners to understand spoken words. The appropriate…

Sound · Computer Science 2025-02-12 Leying Zhang , Wangyou Zhang , Zhengyang Chen , Yanmin Qian

Since trust measures in human-robot interaction are often subjective or not possible to implement real-time, we propose to use speech cues (on what, when and how the user talks) as an objective real-time measure of trust. This could be…

Human-Computer Interaction · Computer Science 2021-04-13 Ella Velner , Khiet P. Truong , Vanessa Evers

Packet loss is a major cause of voice quality degradation in VoIP transmissions with serious impact on intelligibility and user experience. This paper describes a system based on a generative adversarial approach, which aims to repair the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-31 Carlo Aironi , Samuele Cornell , Luca Serafini , Stefano Squartini

In recent advancements within speech processing, converting read speech to conversational speech has gained significant attention. The primary challenge in this domain is maintaining naturalness and intelligibility while minimizing…

Computation and Language · Computer Science 2026-05-19 Parshav Singla , Agnik Banerjee , Aaditya Arora , Shruti Aggarwal , Anil Kumar Verma , Vikram C M , Raj Prakash Gohil , Gopal Kumar Agarwal

Input validation is the first line of defense against malformed or malicious inputs. It is therefore critical that the validator (which is often part of the parser) is free of bugs. To build dependable input validators, we propose using…

Formal Languages and Automata Theory · Computer Science 2017-07-11 Pierre Ganty , Boris Köpf , Pedro Valero