English
Related papers

Related papers: INTERSPEECH 2022 Audio Deep Packet Loss Concealmen…

200 papers

This paper addresses the prevalent issue of incorrect speech output in audio-visual speech enhancement (AVSE) systems, which is often caused by poor video quality and mismatched training and test data. We introduce a post-processing…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-01 Wenze Ren , Kuo-Hsuan Hung , Rong Chao , YouJin Li , Hsin-Min Wang , Yu Tsao

This paper investigates physical layer security (PLS) in wireless interference networks. Specifically, we consider confidential transmission from a legitimate transmitter (Alice) to a legitimate receiver (Bob), in the presence of…

Information Theory · Computer Science 2023-06-01 Lin Hu , Jiabing Fan , Hong Wen , Jie Tang , Qianbin Chen

Caching popular contents is a promising way to offload the mobile data traffic in wireless networks, but so far the potential advantage of caching in improving physical layer security (PLS) is rarely considered. In this paper, we contribute…

Information Theory · Computer Science 2017-11-15 Wu Zhao , Zhiyong Chen , Kuikui Li , Bin Xia

Audio deepfake detection is an emerging topic in the artificial intelligence community. The second Audio Deepfake Detection Challenge (ADD 2023) aims to spur researchers around the world to build new innovative technologies that can further…

Speaker anonymization is an effective privacy protection solution that conceals the speaker's identity while preserving the linguistic content and paralinguistic information of the original speech. To establish a fair benchmark and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-05 Jixun Yao , Nikita Kuzmin , Qing Wang , Pengcheng Guo , Ziqian Ning , Dake Guo , Kong Aik Lee , Eng-Siong Chng , Lei Xie

Large audio-language models (LALMs) enhance traditional large language models by integrating audio perception capabilities, allowing them to tackle audio-related tasks. Previous research has primarily focused on assessing the performance of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Chun-Yi Kuan , Wei-Ping Huang , Hung-yi Lee

Text-to-Speech (TTS) and Voice Conversion (VC) models have exhibited remarkable performance in generating realistic and natural audio. However, their dark side, audio deepfake poses a significant threat to both society and individuals.…

Cryptography and Security · Computer Science 2024-09-17 Xinfeng Li , Kai Li , Yifan Zheng , Chen Yan , Xiaoyu Ji , Wenyuan Xu

This paper describes the submissions of team TalTech-IRIT-LIS to the DISPLACE 2024 challenge. Our team participated in the speaker diarization and language diarization tracks of the challenge. In the speaker diarization track, our best…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-18 Joonas Kalda , Tanel Alumäe , Martin Lebourdais , Hervé Bredin , Séverin Baroudi , Ricard Marxer

Current audio deepfake detectors cannot be trusted. While they excel on controlled benchmarks, they fail when tested in the real world. We introduce Perturbed Public Voices (P$^{2}$V), an IRB-approved dataset capturing three critical…

Sound · Computer Science 2025-08-18 Chongyang Gao , Marco Postiglione , Isabel Gortner , Sarit Kraus , V. S. Subrahmanian

The VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the voice anonymisation…

Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-17 Shota Horiguchi , Takanori Ashihara , Marc Delcroix , Atsushi Ando , Naohiro Tawara

Large Audio-Language Models (LALMs), such as GPT-4o, have recently unlocked audio dialogue capabilities, enabling direct spoken exchanges with humans. The potential of LALMs broadens their applicability across a wide range of practical…

Artificial Intelligence · Computer Science 2025-07-29 Kuofeng Gao , Shu-Tao Xia , Ke Xu , Philip Torr , Jindong Gu

The latest Audio Language Models (Audio LMs) process speech directly instead of relying on a separate transcription step. This shift preserves detailed information, such as intonation or the presence of multiple speakers, that would…

Sound · Computer Science 2025-09-10 Luxi He , Xiangyu Qi , Michel Liao , Inyoung Cheong , Prateek Mittal , Danqi Chen , Peter Henderson

Recently, Large Audio Language Models (LALMs) have progressed rapidly, demonstrating their strong efficacy in universal audio understanding through cross-modal integration. To evaluate LALMs' audio understanding performance, researchers…

Sound · Computer Science 2026-02-16 Han Yin , Jung-Woo Choi

Multilingual language models (LMs) promise broader NLP access, yet current systems deliver uneven performance across the world's languages. This survey examines why these gaps persist and whether they reflect intrinsic linguistic difficulty…

Computation and Language · Computer Science 2026-04-13 Chen Shani , Yuval Reif , Nathan Roll , Dan Jurafsky , Ekaterina Shutova

Integrated Sensing and Communication (ISAC) is emerging as a cornerstone technology for forthcoming 6G systems, significantly improving spectrum and energy efficiency. However, the commercial viability of ISAC hinges on addressing critical…

Audio Captioning (AC) plays a pivotal role in enhancing audio-text cross-modal understanding during the pretraining and finetuning of Multimodal LLMs (MLLMs). To strengthen this alignment, recent works propose Audio Difference Captioning…

Sound · Computer Science 2026-01-27 Yuhang Jia , Xu Zhang , Yujie Guo , Yang Chen , Shiwan Zhao

The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition,…

Due to great efficiency improvement in resource and hardware space, integrated sensing and communication (ISAC) has gained much attention. In the paper, the physical layer security (PLS) of ISAC system under communication eavesdropper…

Information Theory · Computer Science 2026-04-22 Tian Zhang , Zhirong Su , Yueyi Dong

This paper investigates the physical-layer security for an indoor visible light communication (VLC) network consisting of a transmitter, a legitimate receiver and an eavesdropper. Both the main channel and the wiretapping channel have…

Information Theory · Computer Science 2018-07-27 Jin-Yuan Wang , Cheng Liu , Jun-Bo Wang , Yongpeng Wu , Min Lin , Julian Cheng
‹ Prev 1 4 5 6 7 8 10 Next ›