中文
相关论文

相关论文: INTERSPEECH 2022 Audio Deep Packet Loss Concealmen…

200 篇论文

Recent speech and audio coding standards such as 3GPP Enhanced Voice Services match the foreseeable needs and requirements in transmission of speech and audio, when using current transmission infrastructure and applications. Trends in…

音频与语音处理 · 电气工程与系统科学 2018-11-15 Tom Bäckström

Audio is a fundamental modality for analyzing speech, music, and environmental sounds. Although pretrained audio models have significantly advanced audio understanding, they remain fragile in real-world settings where data distributions…

声音 · 计算机科学 2026-02-04 Chang Li , Kanglei Zhou , Liyuan Wang

The INTERSPEECH 2020 Deep Noise Suppression (DNS) Challenge is intended to promote collaborative research in real-time single-channel Speech Enhancement aimed to maximize the subjective (perceptual) quality of the enhanced speech. A typical…

Most intelligent reflecting surface (IRS)-aided indoor visible light communication (VLC) studies ignore the time delays introduced by reflected paths, even though these delays are inherent in practical wideband systems. In this work, we…

信息论 · 计算机科学 2026-03-10 Rashid Iqbal , Ahmed Zoha , Salama Ikki , Muhammad Ali Imran , Hanaa Abumarshoud

This paper describes that semi-supervised learning called peer collaborative learning (PCL) can be applied to the polyphonic sound event detection (PSED) task, which is one of the tasks in the Detection and Classification of Acoustic Scenes…

音频与语音处理 · 电气工程与系统科学 2022-03-31 Hayato Endo , Hiromitsu Nishizaki

This paper describes the TSUP team's submission to the ISCSLP 2022 conversational short-phrase speaker diarization (CSSD) challenge which particularly focuses on short-phrase conversations with a new evaluation metric called conversational…

声音 · 计算机科学 2023-10-26 Bowen Pang , Huan Zhao , Gaosheng Zhang , Xiaoyue Yang , Yang Sun , Li Zhang , Qing Wang , Lei Xie

Integrated sensing and communication (ISAC) systems have emerged as a promising solution to improve spectrum efficiency and enable functional convergence. However, ensuring secure information transmission while maintaining high-quality…

信号处理 · 电气工程与系统科学 2025-12-12 Yufei Wang , Qiang Li , Hongli Liu , Ying Zhang , Jingran Lin

The availability of smart devices leads to an exponential increase in multimedia content. However, advancements in deep learning have also enabled the creation of highly sophisticated Deepfake content, including speech Deepfakes, which pose…

声音 · 计算机科学 2025-07-16 Menglu Li , Yasaman Ahmadiadli , Xiao-Ping Zhang

In recent advancements within speech processing, converting read speech to conversational speech has gained significant attention. The primary challenge in this domain is maintaining naturalness and intelligibility while minimizing…

Existing Audio Deepfake Detection (ADD) systems often struggle to generalise effectively due to the significantly degraded audio quality caused by audio codec compression and channel transmission effects in real-world communication…

音频与语音处理 · 电气工程与系统科学 2026-05-12 Haohan Shi , Xiyu Shi , Safak Dogan , Saif Alzubi , Tianjin Huang , Yunxiao Zhang

This paper investigates physical layer security (PLS) in the intelligent reflecting surface (IRS)-assisted multiple-user uplink channel. Since the instantaneous eavesdropper's channel state information (CSI) is unavailable, the secrecy rate…

信息论 · 计算机科学 2022-05-24 Xiangrui Cheng , Yiliang Liu , Zhou Su , Wei Wang

In this paper, we explore the encoding/pooling layer and loss function in the end-to-end speaker and language recognition system. First, a unified and interpretable end-to-end system for both speaker and language recognition is developed.…

音频与语音处理 · 电气工程与系统科学 2018-04-17 Weicheng Cai , Jinkun Chen , Ming Li

In this paper, we consider the physical layer security (PLS) problem for integrated sensing and communication (ISAC) systems in the presence of hybrid-colluding eavesdroppers, where an active eavesdropper (AE) and a passive eavesdropper…

系统与控制 · 电气工程与系统科学 2025-02-10 Meiding Liu , Zhengchun Zhou , Qiao Shi , Guyue Li , Zilong Liu , Pingzhi Fan , Inkyu Lee

Widely-deployed encryption-based security prevents unauthorized decoding, but does not ensure undetectability of communication. However, covert, or low probability of detection/intercept (LPD/LPI) communication is crucial in many scenarios…

信息论 · 计算机科学 2015-06-02 Boulat A. Bash , Dennis Goeckel , Saikat Guha , Don Towsley

Audio data are widely exchanged over telecommunications networks. Due to the limitations of network resources, these data are typically compressed before transmission. Various methods are available for compressing audio data. To access such…

多媒体 · 计算机科学 2025-02-12 Farzane Jafari

Collaborative intelligence (CI) involves dividing an artificial intelligence (AI) model into two parts: front-end, to be deployed on an edge device, and back-end, to be deployed in the cloud. The deep feature tensors produced by the…

图像与视频处理 · 电气工程与系统科学 2023-07-06 Korcan Uyanik , S. Faegheh Yeganli , Ivan V. Bajić

Large-scale LLM training requires collective communication libraries to exchange data among distributed GPUs. As a company dedicated to building and operating large-scale GPU training clusters, we encounter several challenges when using…

Language-queried Audio Separation (LASS) employs linguistic queries to isolate target sounds based on semantic descriptions. However, existing methods face challenges in aligning complex auditory features with linguistic context while…

声音 · 计算机科学 2025-06-23 Jianyuan Feng , Guangzheng Li , Yangfei Xu

Privacy in speech and audio has many facets. A particularly under-developed area of privacy in this domain involves consideration for information related to content and context. Speech content can include words and their meaning or even…

音频与语音处理 · 电气工程与系统科学 2023-01-24 Jennifer Williams , Karla Pizzi , Shuvayanti Das , Paul-Gauthier Noe

We review current solutions and technical challenges for automatic speech recognition, keyword spotting, device arbitration, speech enhancement, and source localization in multidevice home environments to provide context for the INTERSPEECH…

音频与语音处理 · 电气工程与系统科学 2022-07-01 Gregory Ciccarelli , Jarred Barber , Arun Nair , Israel Cohen , Tao Zhang