中文
相关论文

相关论文: A first-order DirAC-based parametric Ambisonic cod…

200 篇论文

Neural Speech Codecs face a fundamental trade-off at low bitrates: preserving acoustic fidelity often compromises semantic richness. To address this, we introduce SACodec, a novel codec built upon an asymmetric dual-quantizer that employs…

声音 · 计算机科学 2025-12-25 Zhongren Dong , Bin Wang , Jing Han , Haotian Guo , Xiaojun Mo , Yimin Cao , Zixing Zhang

Pretrained latent diffusion models have shown strong potential for lossy image compression, owing to their powerful generative priors. Most existing diffusion-based methods reconstruct images by iteratively denoising from random noise,…

图像与视频处理 · 电气工程与系统科学 2025-10-21 Jinpei Guo , Yifei Ji , Zheng Chen , Kai Liu , Min Liu , Wang Rao , Wenbo Li , Yong Guo , Yulun Zhang

Direction-of-arrival (DOA) estimation is an important task in microphone array processing and many downstream applications. The steered response power with phase transform (SRP-PHAT) method has been widely adopted for DOA estimation in…

音频与语音处理 · 电气工程与系统科学 2026-04-29 Ming Huang , Shuting Xu , Leying Yang , Huanzhang Hu , Yujie Zhang , Jiang Wang , Yu Liu , Hao Zhao , He Kong

Binaural rendering of ambisonic signals is of broad interest to virtual reality and immersive media. Conventional methods often require manually measured Head-Related Transfer Functions (HRTFs). To address this issue, we collect a paired…

声音 · 计算机科学 2022-11-07 Yin Zhu , Qiuqiang Kong , Junjie Shi , Shilei Liu , Xuzhou Ye , Ju-chiang Wang , Junping Zhang

Volumetric video based on Neural Radiance Field (NeRF) holds vast potential for various 3D applications, but its substantial data volume poses significant challenges for compression and transmission. Current NeRF compression lacks the…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Zihan Zheng , Houqiang Zhong , Qiang Hu , Xiaoyun Zhang , Li Song , Ya Zhang , Yanfeng Wang

Ambisonics, a popular format of spatial audio, is the spherical harmonic (SH) representation of the plane wave density function of a sound field. Many algorithms operate in the SH domain and utilize the Ambisonics as their input signal. The…

音频与语音处理 · 电气工程与系统科学 2024-03-01 Bar Shaybet , Anurag Kumar , Vladimir Tourbabin , Boaz Rafaely

We propose a new architecture for distributed image compression from a group of distributed data sources. The work is motivated by practical needs of data-driven codec design, low power consumption, robustness, and data privacy. The…

计算机视觉与模式识别 · 计算机科学 2020-06-11 Enmao Diao , Jie Ding , Vahid Tarokh

In this work, we address the challenge of encoding speech captured by a microphone array using deep learning techniques with the aim of preserving and accurately reconstructing crucial spatial cues embedded in multi-channel recordings. We…

声音 · 计算机科学 2024-07-10 Zhongweiyang Xu , Yong Xu , Vinay Kothapally , Heming Wang , Muqiao Yang , Dong Yu

We propose a low-resolution analog-to-digital converter (ADC) module assisted hybrid beamforming architecture for millimeter-wave (mmWave) communications. We prove that the proposed low-cost and flexible architecture can reduce the beam…

信息论 · 计算机科学 2018-03-28 Jie Yang , Xi Yang , Shi Jin , Chao-Kai Wen , Michalis Matthaiou

A codec for compression of music signals is proposed. The method belongs to the class of transform lossy compression. It is conceived to be applied in the high quality recovery range though. The transformation, endowing the codec with its…

声音 · 计算机科学 2015-12-15 Laura Rebollo-Neira

While molecular communication via diffusion experiences significant inter-symbol interference (ISI), recent work suggests that ISI can be mitigated via time differentiation pre-processing which achieves pulse narrowing. Herein, the approach…

信息论 · 计算机科学 2021-08-19 Mustafa Can Gursoy , Urbashi Mitra

Direction of arrival (DOA) estimation is a fundamental problem in array signal processing with applications spanning radar, sonar, wireless communications, and acoustic signal processing. This tutorial survey provides a comprehensive…

信号处理 · 电气工程与系统科学 2025-09-03 Amgad A. Salama

Audio watermarking is widely used for leaking source tracing. The robustness of the watermark determines the traceability of the algorithm. With the development of digital technology, audio re-recording (AR) has become an efficient and…

声音 · 计算机科学 2023-04-04 Chang Liu , Jie Zhang , Han Fang , Zehua Ma , Weiming Zhang , Nenghai Yu

Integrated sensing and communications (ISAC) has emerged as a means to efficiently utilize spectrum and thereby save cost and power. At the higher end of the spectrum, ISAC systems operate at wideband using large antenna arrays to meet the…

信号处理 · 电气工程与系统科学 2024-11-06 Ahmet M. Elbir , Abdulkadir Celik , Ahmed M. Eltawil

This paper investigates a practical partially-connected hybrid beamforming transmitter for integrated sensing and communication (ISAC) with distortion from nonlinear power amplification. For this ISAC system, we formulate a communication…

信号处理 · 电气工程与系统科学 2025-07-21 Zeyuan Zhang , Yue Xiu , Phee Lep Yeoh , Guangyi Liu , Zixing Wu , Ning Wei

This paper proposes a tensor-based parametric channel estimation technique for IRS-assisted communication systems with time-varying channel parameters. We exploit the multidimensional structure of the received signal by developing a…

信号处理 · 电气工程与系统科学 2026-05-29 Kenneth B. A. Benício , André L. F. de Almeida , Bruno Sokal , Fazal-E-Asim , Behrooz Makki , Gabor Fodor

Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, prevailing methods augment acoustic…

声音 · 计算机科学 2026-01-28 Xin Zhang , Lin Li , Xiangni Lu , Jianquan Liu , Kong Aik Lee

A depth image provides partial geometric information of a 3D scene, namely the shapes of physical objects as observed from a particular viewpoint. This information is important when synthesizing images of different virtual camera viewpoints…

多媒体 · 计算机科学 2016-12-26 Yuan Yuan , Gene Cheung , Patrick Le Callet , Pascal Frossard , Hong Vicky Zhao

Neural audio codecs have recently gained traction for their ability to compress high-fidelity audio and provide discrete tokens for generative modeling. However, leading approaches often rely on resource-intensive models and complex…

声音 · 计算机科学 2025-08-18 Linwei Zhai , Han Ding , Cui Zhao , fei wang , Ge Wang , Wang Zhi , Wei Xi

This paper introduces a novel method for transmitting video data over noisy wireless channels with high efficiency and controllability. The method derivates from model division multiple access (MDMA) to extract common semantic features from…

计算工程、金融与科学 · 计算机科学 2023-08-11 Zhicheng Bao , Haotai Liang , Chen Dong , Cong Li , Xiaodong Xu , Ping Zhang