English
Related papers

Related papers: Generating Localized Audible Zones Using a Single-…

200 papers

Text prompts enable intuitive content creation but may fall short in achieving high precision for intricate tasks; knob or slider controls offer precise adjustments at the cost of increased complexity. To address the gap between knobs and…

Human-Computer Interaction · Computer Science 2025-08-15 Yuan-Yi Fan

The design of zero-delay Joint Source-Channel Coding (JSCC) schemes for the transmission of correlated information over fading Multiple Access Channels (MACs) is an interesting problem for many communication scenarios like Wireless Sensor…

Information Theory · Computer Science 2024-01-31 O. Fresnedo , P. Suárez-Casal , L. Castedo

Source-Free Domain Adaptation (SFDA) aims to train a target model without source data, and the key is to generate pseudo-labels using a pre-trained source model. However, we observe that the source model often produces highly uncertain…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Jie Cheng , Hao Zheng , Meiguang Zheng , Lei Wang , Hao Wu , Jian Zhang

Accurately and efficiently addressing the multiple source localization (MSL) problem in urban environments, particularly designing a general method adaptable to an arbitrary number of sources, plays a crucial role in various fields such as…

Signal Processing · Electrical Eng. & Systems 2025-12-19 Qilu Zhang , Hongying Tang , Wen Chen , Ziyi Song , Jiang Wang

We describe a private audio messaging system that uses echoes to unscramble messages at a few predetermined locations in a room. The system works by splitting the audio into short chunks and emitting them from different loudspeakers. The…

Sound · Computer Science 2018-09-18 Yu-Jeh Liu , Jonah Casebeer , Ivan Dokmanić

Traditional speech enhancement systems produce speech with compromised quality. Here we propose to use the high quality speech generation capability of neural vocoders for better quality speech enhancement. We term this parametric…

Sound · Computer Science 2019-11-15 Soumi Maiti , Michael I Mandel

Consistency regularization has prevailed in semi-supervised semantic segmentation and achieved promising performance. However, existing methods typically concentrate on enhancing the Image-augmentation based Prediction consistency and…

Multimedia · Computer Science 2025-03-25 Jianjian Yin , Tao Chen , Gensheng Pei , Yazhou Yao , Liqiang Nie , Xiansheng Hua

Spatial audio formats like Ambisonics are playback device layout-agnostic and well-suited for applications such as teleconferencing and virtual reality. Conventional Ambisonic encoding methods often rely on spherical microphone arrays for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Yue Qiao , Vinay Kothapally , Meng Yu , Dong Yu

While Separate Source-Channel Coding (SSCC) retains the practical benefits of modular system design, its effectiveness in noisy text transmission is fundamentally constrained by the fragility of autoregressive source decoding. In low-SNR…

Information Theory · Computer Science 2026-05-08 Ziqiong Wang , Rongpeng Li

Channel models are essential for the design, evaluation, and optimization of wireless communication systems. The emerging space-air-ground-sea integrated network (SAGSIN), characterized by diverse service applications and extended-spectrum…

Signal Processing · Electrical Eng. & Systems 2026-03-17 Nanhao Zhou , Chao Zou , Yu Zhou , Yanqun Tang , Xiaoying Zhang , Haoran Yin , Xuefeng Yin , Yuxiang Zhang , Dan Fei , Fan Jiang

Self-supervised learning (SSL) has emerged as a powerful paradigm for Chest X-ray (CXR) analysis under limited annotations. Yet, existing SSL strategies remain suboptimal for medical imaging. Masked image modeling allocates substantial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Wangyu Feng , Shawn Young , Lijian Xu

A new single-letter achievable rate region is proposed for the two-user discrete memoryless multiple-access channel(MAC) with noiseless feedback. The proposed region includes the Cover-Leung rate region [1], and it is shown that the…

Information Theory · Computer Science 2014-03-31 Ramji Venkataramanan , S. Sandeep Pradhan

We present the analysis of a single-carrier massive MIMO system for the frequency selective Gaussian multi-user channel, in both uplink and downlink directions. We develop expressions for the achievable sum rate when there is spatial…

Information Theory · Computer Science 2019-10-10 Nader Beigiparast , Gokhan M. Guvensen , Ender Ayanoglu

High harmonic generation (HHG) from semiconductors and insulators has become a very active area of research due to its great potential for developing compact HHG devices. Here we show that by growing monolayers (ML) of insulators on…

Atomic Physics · Physics 2016-12-28 Néstor F. Aguirre , Fernando Martín

This dissertation covers a single-processor approach to the speech processing pipeline of bilateral Cochlear Implants (CIs). The use of only a single processor to provide binaural stimulation signals overcomes the synchronization problem,…

Sound · Computer Science 2014-09-24 Taher Shahbazi Mirzahasanloo

Replay attacks remain a critical vulnerability for automatic speaker verification systems, particularly in real-time voice assistant applications. In this work, we propose acoustic maps as a novel spatial feature representation for replay…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Michael Neri , Tuomas Virtanen

The dual-stream transformer architecture-based joint audio-video generation method has become the dominant paradigm in current research. By incorporating pre-trained video diffusion models and audio diffusion models, along with a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Bingqi Ma , Linlong Lang , Ming Zhang , Dailan He , Xingtong Ge , Yi Zhang , Guanglu Song , Yu Liu

Multimodal large language models (MLLMs) achieve strong performance by jointly processing inputs from multiple modalities, such as vision, audio, and language. However, building such models or extending them to new modalities often requires…

Machine Learning · Computer Science 2026-03-24 Md Kaykobad Reza , Ameya Patil , Edward Ayrapetian , M. Salman Asif

The performance of automatic speech recognition (ASR) systems severely degrades when multi-talker speech overlap occurs. In meeting environments, speech separation is typically performed to improve the robustness of ASR systems. Recently,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-18 Hassan Taherian , DeLiang Wang

We consider spatially coupled low-density parity-check (SC-LDPC) codes within a non-orthogonal interleave division multiple access (IDMA) scheme to avoid cumbersome degree profile matching of the LDPC code components to the iterative…

Information Theory · Computer Science 2019-01-29 Sebastian Cammerer , Xiaojie Wang , Yingyan Ma , Stephan ten Brink