English
Related papers

Related papers: Perceptually Relevant Preservation of Interaural T…

200 papers

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Zhenjie Liu , Jianzhang Lu , Renjie Lu , Cong Liang , Shangfei Wang

Astronomical Kinetic Inductance Detectors (KIDs), similar to quantum information devices, experience performance limiting noise from materials. In particular, 1/f (frequency) noise can be a dominant noise mechanism, which arises from…

Superconductivity · Physics 2023-12-14 N. Foroozani , B. Sarabi , S. H. Moseley , T. Stevenson , E. J. Wollack , O. Noroozian , K. D. Osborn

While different variants of perceptual losses have been employed in super-resolution literature to synthesize more realistic, appealing, and detailed high-resolution images, most are convolutional neural networks-based, causing information…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Shoaib Meraj Sami , Md Mahedi Hasan , Mohammad Saeed Ebrahimi Saadabadi , Jeremy Dawson , Nasser Nasrabadi , Raghuveer Rao

Current audio-visual separation methods share a standard architecture design where an audio encoder-decoder network is fused with visual encoding features at the encoder bottleneck. This design confounds the learning of multi-modal feature…

Computer Vision and Pattern Recognition · Computer Science 2022-12-09 Jiaben Chen , Renrui Zhang , Dongze Lian , Jiaqi Yang , Ziyao Zeng , Jianbo Shi

For computational acoustics, schemes need to have low-dispersion and low-dissipation properties in order to capture the amplitude and phase of the wave correctly. To improve the spectral properties of the scheme, the authors have previously…

Computational Physics · Physics 2021-11-15 Y. H. Li , Y. X. Ren , Y. T. Su

Hand-crafted spatial features, such as inter-channel intensity difference (IID) and inter-channel phase difference (IPD), play a fundamental role in recent deep learning based dual-microphone speech enhancement (DMSE) systems. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-04 Xinmeng Xu , Rongzhi Gu , Yuexian Zou

Optoelectronic systems based on multiple modes of light can often exceed the performance of their single-mode counterparts. However, multimode nonlinear interactions often introduce considerable amounts of noise, limiting the ultimate…

Noise performance is one of the most crucial aspects of any detector. Superconducting Microwave Kinetic Inductance Detectors (MKIDs) have an "excess" frequency noise that shows up as a small time dependent jitter of the resonance frequency…

Instrumentation and Detectors · Physics 2012-06-27 Omid Noroozian , Jiansong Gao , Jonas Zmuidzinas , Henry G. LeDuc , Benjamin A. Mazin

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

Sound · Computer Science 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

The end-to-end approach for single-channel speech separation has been studied recently and shown promising results. This paper extended the previous approach and proposed a new end-to-end model for multi-channel speech separation. The…

Sound · Computer Science 2019-05-29 Rongzhi Gu , Jian Wu , Shi-Xiong Zhang , Lianwu Chen , Yong Xu , Meng Yu , Dan Su , Yuexian Zou , Dong Yu

3D reconstruction techniques such as LiDAR scanning and photogrammetry have made it practical to build detailed geometric models of real-world environments. Such reconstructed models can potentially serve as the foundation for wireless…

Networking and Internet Architecture · Computer Science 2026-05-27 Haofan Lu , Yadi Cao , Wanghao Yi , Omid Abari

Acoustic beamforming models typically assume wide-sense stationarity of speech signals within short time frames. However, voiced speech is better modeled as a cyclostationary (CS) process, a random process whose mean and autocorrelation are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-16 Giovanni Bologni , Richard Heusdens , Richard C. Hendriks

Reviving natural ventilation (NV) for urban sustainability presents challenges for indoor acoustic comfort. Active control and interference-based noise mitigation strategies, such as the use of loudspeakers, offer potential solutions to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-13 Bhan Lam , Kelvin Chee Quan Lim , Kenneth Ooi , Zhen-Ting Ong , Dongyuan Shi , Woon-Seng Gan

The performance of conventional speech enhancement systems degrades sharply in extremely low signal-to-noise ratio (SNR) environments where air-conduction (AC) microphones are overwhelmed by ambient noise. Although bone-conduction (BC)…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-04 Yilei Wu , Changyan Zheng , Xingyu Zhang , Yakun Zhang , Chengshi Zheng , Shuang Yang , Ye Yan , Erwei Yin

Image-based diagnostic decision support systems (DDSS) utilizing deep learning have the potential to optimize clinical workflows. However, developing DDSS requires extensive datasets with expert annotations and is therefore costly.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Helen Schneider , Sebastian Nowak , Aditya Parikh , Yannik C. Layer , Maike Theis , Wolfgang Block , Alois M. Sprinkart , Ulrike Attenberger , Rafet Sifa

Industrial anomaly detection (IAD) increasingly benefits from integrating 2D and 3D data, but robust cross-modal fusion remains challenging. We propose a novel unsupervised framework, Multi-Modal Attention-Driven Fusion Restoration (MAFR),…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Usman Ali , Ali Zia , Abdul Rehman , Umer Ramzan , Zohaib Hassan , Talha Sattar , Jing Wang , Wei Xiang

Diffusion models have become a leading paradigm for image super-resolution (SR), but existing methods struggle to guarantee both the high-frequency perceptual quality and the low-frequency structural fidelity of generated images. Although…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Hexin Zhang , Dong Li , Jie Huang , Bingzhou Wang , Xueyang Fu , Zhengjun Zha

Unseen noise signal which is not considered in a model training process is difficult to anticipate and would lead to performance degradation. Various methods have been investigated to mitigate unseen noise. In our previous work, an…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-24 Donghyeon Kim , Gwantae Kim , Bokyeung Lee , Jeong-gi Kwak , David K. Han , Hanseok Ko

RNN-T models are widely used in ASR, which rely on the RNN-T loss to achieve length alignment between input audio and target sequence. However, the implementation complexity and the alignment-based optimization target of RNN-T loss lead to…

Sound · Computer Science 2024-11-28 Tian-Hao Zhang , Dinghao Zhou , Guiping Zhong , Jiaming Zhou , Baoxiang Li

This paper presents a computational methodology for analyzing intonation and deriving tuning systems in microtonal oral traditions, utilizing pitch histograms, Dynamic Time Warping (DTW), and optimization techniques, with a case study on a…

Sound · Computer Science 2025-08-29 Sepideh Shafiei , Shapour Hakam