中文
相关论文

相关论文: INTERSPEECH 2022 Audio Deep Packet Loss Concealmen…

200 篇论文

In high-noise environments such as factories, subways, and busy streets, capturing clear speech is challenging. Throat microphones can offer a solution because of their inherent noise-suppression capabilities; however, the passage of sound…

声音 · 计算机科学 2026-04-23 Yunsik Kim , Yonghun Song , Yoonyoung Chung

While natural language is the de facto communication medium for LLM-based agents, it presents a fundamental constraint. The process of downsampling rich, internal latent states into discrete tokens inherently limits the depth and nuance of…

机器学习 · 计算机科学 2026-04-17 Zhuoyun Du , Runze Wang , Huiyu Bai , Zouying Cao , Xiaoyong Zhu , Yu Cheng , Bo Zheng , Wei Chen , Haochao Ying

Audio large language models (Audio LLMs) demonstrate strong performance on speech understanding tasks, yet their ability to understand paralinguistic information remains limited. To systematically quantify this issue, we introduce…

声音 · 计算机科学 2026-05-28 Jiacheng Pang , Ashutosh Chaubey , Mohammad Soleymani

The cloud-based speech recognition/API provides developers or enterprises an easy way to create speech-enabled features in their applications. However, sending audios about personal or company internal information to the cloud, raises…

密码学与安全 · 计算机科学 2019-05-15 Shi-Xiong Zhang , Yifan Gong , Dong Yu

Covert channels can be used to circumvent system and network policies by establishing communications that have not been considered in the design of the computing system. We construct a covert channel between different computing systems that…

密码学与安全 · 计算机科学 2014-06-06 Michael Hanspach , Michael Goetz

Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance…

We revisit the problem of federated learning (FL) with private data from people who do not trust the server or other silos/clients. In this context, every silo (e.g. hospital) has data from several people (e.g. patients) and needs to…

机器学习 · 计算机科学 2024-09-10 Changyu Gao , Andrew Lowy , Xingyu Zhou , Stephen J. Wright

A novel private communication framework is proposed where privacy is induced by transmitting over a channel instances of linear inverse problems that are identifiable to the legitimate receiver but unidentifiable to an eavesdropper. The gap…

信息论 · 计算机科学 2024-07-24 Maxime Ferreira Da Costa , Jianxiu Li , Urbashi Mitra

Deep neural networks (DNNs) have shown promising results for acoustic echo cancellation (AEC). But the DNN-based AEC models let through all near-end speakers including the interfering speech. In light of recent studies on personalized…

声音 · 计算机科学 2022-07-01 Shimin Zhang , Ziteng Wang , Yukai Ju , Yihui Fu , Yueyue Na , Qiang Fu , Lei Xie

High sound pressure levels (SPL) pose notable risks in loud environments, particularly due to noise-induced hearing loss. Ill-fitting earplugs often lead to sound leakage, a phenomenon this study seeks to investigate. To validate our…

声音 · 计算机科学 2025-10-21 Haocheng Yu , Krishan K. Ahuja , Lakshmi N. Sankar , Spencer H. Bryngelson

The aim of this paper is to introduce a new schema, based on a Compressive Sampling technique, for the recovery of lost data in multimedia streaming. The audio streaming data are encapsuled in different packets by using an interleaving…

多媒体 · 计算机科学 2013-08-21 Angelo Ciaramella , Giulio Giunta

Contrastive predictive coding (CPC) aims to learn representations of speech by distinguishing future observations from a set of negative examples. Previous work has shown that linear classifiers trained on CPC features can accurately…

音频与语音处理 · 电气工程与系统科学 2021-08-03 Benjamin van Niekerk , Leanne Nortje , Matthew Baas , Herman Kamper

Accurately detecting dysfluencies in spoken language can help to improve the performance of automatic speech and language processing components and support the development of more inclusive speech and language technologies. Inspired by the…

The paper considers constrained linear systems with stochastic additive disturbances and noisy measurements transmitted over a lossy communication channel. We propose a model predictive control (MPC) law that minimizes a discounted cost…

最优化与控制 · 数学 2020-05-08 Shuhao Yan , Mark Cannon , Paul Goulart

This paper summarizes the outcomes from the ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge (ICSRC). We first address the necessity of the challenge and then introduce the associated dataset collected from a new-energy vehicle…

声音 · 计算机科学 2022-11-04 Ao Zhang , Fan Yu , Kaixun Huang , Lei Xie , Longbiao Wang , Eng Siong Chng , Hui Bu , Binbin Zhang , Wei Chen , Xin Xu

In this paper the secure performance for the visible light communication (VLC) system with multiple eavesdroppers is studied. By considering the practical amplitude constraint instead of an average power constraint in the VLC system, the…

信号处理 · 电气工程与系统科学 2019-10-17 Xiaodong Liu , Yuhao Wang , Fuhui Zhou , Zhenyu Deng , Rose Qingyang Hu

Privacy-preserving voice protection approaches primarily suppress privacy-related information derived from paralinguistic attributes while preserving the linguistic content. Existing solutions focus particularly on single-speaker scenarios.…

声音 · 计算机科学 2025-03-28 Xiaoxiao Miao , Ruijie Tao , Chang Zeng , Xin Wang

The performance of conventional interference management strategies degrades when interference power is comparable to signal power. We consider a new perspective on interference management using semantic communication. Specifically, a…

信号处理 · 电气工程与系统科学 2024-08-09 Zian Meng , Qiang Li , Ashish Pandharipande , Xiaohu Ge

Physical Layer Security (PLS) is an emerging concept in the field of secrecy for wireless communications that can be used alongside cryptography to prevent unauthorized devices from eavesdropping a legitimate transmission. It offers low…

信号处理 · 电气工程与系统科学 2025-04-24 Leonardo Barbosa da Silva , Evelio Martín Garcia Fernández , Ândrei Camponogara

Automatic speech recognition (ASR) is a key technology in many services and applications. This typically requires user devices to send their speech data to the cloud for ASR decoding. As the speech signal carries a lot of information about…

计算与语言 · 计算机科学 2019-11-13 Brij Mohan Lal Srivastava , Aurélien Bellet , Marc Tommasi , Emmanuel Vincent
‹ 上一页 1 8 9 10 下一页 ›