中文
相关论文

相关论文: Perceptually Relevant Preservation of Interaural T…

200 篇论文

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Zhenjie Liu , Jianzhang Lu , Renjie Lu , Cong Liang , Shangfei Wang

Astronomical Kinetic Inductance Detectors (KIDs), similar to quantum information devices, experience performance limiting noise from materials. In particular, 1/f (frequency) noise can be a dominant noise mechanism, which arises from…

超导电性 · 物理学 2023-12-14 N. Foroozani , B. Sarabi , S. H. Moseley , T. Stevenson , E. J. Wollack , O. Noroozian , K. D. Osborn

While different variants of perceptual losses have been employed in super-resolution literature to synthesize more realistic, appealing, and detailed high-resolution images, most are convolutional neural networks-based, causing information…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Shoaib Meraj Sami , Md Mahedi Hasan , Mohammad Saeed Ebrahimi Saadabadi , Jeremy Dawson , Nasser Nasrabadi , Raghuveer Rao

Current audio-visual separation methods share a standard architecture design where an audio encoder-decoder network is fused with visual encoding features at the encoder bottleneck. This design confounds the learning of multi-modal feature…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Jiaben Chen , Renrui Zhang , Dongze Lian , Jiaqi Yang , Ziyao Zeng , Jianbo Shi

For computational acoustics, schemes need to have low-dispersion and low-dissipation properties in order to capture the amplitude and phase of the wave correctly. To improve the spectral properties of the scheme, the authors have previously…

计算物理 · 物理学 2021-11-15 Y. H. Li , Y. X. Ren , Y. T. Su

Hand-crafted spatial features, such as inter-channel intensity difference (IID) and inter-channel phase difference (IPD), play a fundamental role in recent deep learning based dual-microphone speech enhancement (DMSE) systems. However,…

音频与语音处理 · 电气工程与系统科学 2022-05-04 Xinmeng Xu , Rongzhi Gu , Yuexian Zou

Optoelectronic systems based on multiple modes of light can often exceed the performance of their single-mode counterparts. However, multimode nonlinear interactions often introduce considerable amounts of noise, limiting the ultimate…

Noise performance is one of the most crucial aspects of any detector. Superconducting Microwave Kinetic Inductance Detectors (MKIDs) have an "excess" frequency noise that shows up as a small time dependent jitter of the resonance frequency…

仪器与探测器 · 物理学 2012-06-27 Omid Noroozian , Jiansong Gao , Jonas Zmuidzinas , Henry G. LeDuc , Benjamin A. Mazin

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

声音 · 计算机科学 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

The end-to-end approach for single-channel speech separation has been studied recently and shown promising results. This paper extended the previous approach and proposed a new end-to-end model for multi-channel speech separation. The…

声音 · 计算机科学 2019-05-29 Rongzhi Gu , Jian Wu , Shi-Xiong Zhang , Lianwu Chen , Yong Xu , Meng Yu , Dan Su , Yuexian Zou , Dong Yu

3D reconstruction techniques such as LiDAR scanning and photogrammetry have made it practical to build detailed geometric models of real-world environments. Such reconstructed models can potentially serve as the foundation for wireless…

网络与互联网体系结构 · 计算机科学 2026-05-27 Haofan Lu , Yadi Cao , Wanghao Yi , Omid Abari

Acoustic beamforming models typically assume wide-sense stationarity of speech signals within short time frames. However, voiced speech is better modeled as a cyclostationary (CS) process, a random process whose mean and autocorrelation are…

音频与语音处理 · 电气工程与系统科学 2025-12-16 Giovanni Bologni , Richard Heusdens , Richard C. Hendriks

Reviving natural ventilation (NV) for urban sustainability presents challenges for indoor acoustic comfort. Active control and interference-based noise mitigation strategies, such as the use of loudspeakers, offer potential solutions to…

音频与语音处理 · 电气工程与系统科学 2023-07-13 Bhan Lam , Kelvin Chee Quan Lim , Kenneth Ooi , Zhen-Ting Ong , Dongyuan Shi , Woon-Seng Gan

The performance of conventional speech enhancement systems degrades sharply in extremely low signal-to-noise ratio (SNR) environments where air-conduction (AC) microphones are overwhelmed by ambient noise. Although bone-conduction (BC)…

音频与语音处理 · 电气工程与系统科学 2026-03-04 Yilei Wu , Changyan Zheng , Xingyu Zhang , Yakun Zhang , Chengshi Zheng , Shuang Yang , Ye Yan , Erwei Yin

Image-based diagnostic decision support systems (DDSS) utilizing deep learning have the potential to optimize clinical workflows. However, developing DDSS requires extensive datasets with expert annotations and is therefore costly.…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Helen Schneider , Sebastian Nowak , Aditya Parikh , Yannik C. Layer , Maike Theis , Wolfgang Block , Alois M. Sprinkart , Ulrike Attenberger , Rafet Sifa

Industrial anomaly detection (IAD) increasingly benefits from integrating 2D and 3D data, but robust cross-modal fusion remains challenging. We propose a novel unsupervised framework, Multi-Modal Attention-Driven Fusion Restoration (MAFR),…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Usman Ali , Ali Zia , Abdul Rehman , Umer Ramzan , Zohaib Hassan , Talha Sattar , Jing Wang , Wei Xiang

Diffusion models have become a leading paradigm for image super-resolution (SR), but existing methods struggle to guarantee both the high-frequency perceptual quality and the low-frequency structural fidelity of generated images. Although…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Hexin Zhang , Dong Li , Jie Huang , Bingzhou Wang , Xueyang Fu , Zhengjun Zha

Unseen noise signal which is not considered in a model training process is difficult to anticipate and would lead to performance degradation. Various methods have been investigated to mitigate unseen noise. In our previous work, an…

音频与语音处理 · 电气工程与系统科学 2022-10-24 Donghyeon Kim , Gwantae Kim , Bokyeung Lee , Jeong-gi Kwak , David K. Han , Hanseok Ko

RNN-T models are widely used in ASR, which rely on the RNN-T loss to achieve length alignment between input audio and target sequence. However, the implementation complexity and the alignment-based optimization target of RNN-T loss lead to…

声音 · 计算机科学 2024-11-28 Tian-Hao Zhang , Dinghao Zhou , Guiping Zhong , Jiaming Zhou , Baoxiang Li

This paper presents a computational methodology for analyzing intonation and deriving tuning systems in microtonal oral traditions, utilizing pitch histograms, Dynamic Time Warping (DTW), and optimization techniques, with a case study on a…

声音 · 计算机科学 2025-08-29 Sepideh Shafiei , Shapour Hakam