中文
相关论文

相关论文: iMagLS: Interaural Level Difference with Magnitude…

200 篇论文

In this work, we propose a deep beamforming framework for speech enhancement in dynamic acoustic environments. The framework learns time-varying beamformer weights from noisy multichannel signals via a deep neural network, guided by a…

音频与语音处理 · 电气工程与系统科学 2026-02-18 Ilai Zaidel , Sharon Gannot

End-to-end neural TTS training has shown improved performance in speech style transfer. However, the improvement is still limited by the training data in both target styles and speakers. Inadequate style transfer performance occurs when the…

声音 · 计算机科学 2021-06-21 Xiaochun An , Frank K. Soong , Lei Xie

Ambisonics is a method for capturing and rendering a sound field accurately, assuming that the acoustics of the playback room does not significantly influence the sound field. However, in practice, the acoustics of the playback room may…

音频与语音处理 · 电气工程与系统科学 2025-10-14 Ali Fallah , Shun Nakamura , Steven van de Par

Hallucination remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While Direct Preference Optimization (DPO) is a key alignment framework, existing approaches often rely heavily on costly external evaluators for…

机器学习 · 计算机科学 2026-02-04 Yuanshuai Li , Yuping Yan , Jirui Han , Fei Ming , Lingjuan Lv , Yaochu Jin

As spatial audio is enjoying a surge in popularity, data-driven machine learning techniques that have been proven successful in other domains are increasingly used to process head-related transfer function measurements. However, these…

音频与语音处理 · 电气工程与系统科学 2022-12-09 Johan Pauwels , Lorenzo Picinali

Hearing loss (HL) simulators, which allow normal hearing (NH) listeners to experience HL, have been used in speech intelligibility experiments, but not in sound quality experiments due to perceptible distortion. If they produced less…

音频与语音处理 · 电气工程与系统科学 2024-09-25 Toshio Irino , Shintaro Doan , Minami Ishikawa

In this work, we propose a two-dimensional Head-Related Transfer Function (HRTF)-based robust beamformer design for robot audition, which allows for explicit control of the beamformer response for the entire three-dimensional sound field…

声音 · 计算机科学 2017-03-10 Hendrik Barfuss , Michael Buerger , Jasper Podschus , Walter Kellermann

Remote sensing imagery suffers from clouds, haze, noise, resolution limits, and sensor heterogeneity. Existing restoration and fusion approaches train separate models per degradation type. In this work, we present Language-conditioned…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Yongchuan Cui , Peng Liu

The multichannel Wiener filter (MWF) and its variations have been extensively applied to binaural hearing aids. However, its major drawback is the distortion of the binaural cues of the residual noise, changing the original acoustic…

音频与语音处理 · 电气工程与系统科学 2019-11-12 Johnny Werner , Marcio H. Costa

Many efforts have been devoted to designing sampling, mining, and weighting strategies in high-level deep metric learning (DML) loss objectives. However, little attention has been paid to low-level but essential data transformation. In this…

计算机视觉与模式识别 · 计算机科学 2021-05-24 Zhiyuan Chen , Guang Yao , Wennan Ma , Lin Xu

Several individualization methods have recently been proposed to estimate a subject's Head-Related Transfer Function (HRTF) using convenient input modalities such as anthropometric measurements or pinnae photographs. There exists a need for…

音频与语音处理 · 电气工程与系统科学 2023-10-23 Etienne Thuillier , Craig Jin , Vesa Välimäki

Conventional automatic speech recognition (ASR) system uses second-order minkowski loss during inference time which is suboptimal as it incorporates only first order statistics in posterior estimation [2]. In this paper we have proposed…

音频与语音处理 · 电气工程与系统科学 2021-12-03 Vishwanath Pratap Singh , Shakti P. Rath , Abhishek Pandey

Auditory large language models (ALLMs) have demonstrated strong general capabilities in audio understanding and reasoning tasks. However, their reliability is still undermined by hallucination issues. Existing hallucination evaluation…

声音 · 计算机科学 2026-04-13 Qixuan Huang , Khalid Zaman , Masashi Unoki

AI-based neural decoding reconstructs visual perception by leveraging generative models to map brain activity, measured through functional MRI (fMRI), into latent hierarchical representations. Traditionally, ridge linear models transform…

图像与视频处理 · 电气工程与系统科学 2025-09-04 Lorenzo Veronese , Andrea Moglia , Luca Mainardi , Pietro Cerveri

We introduce HiFi-HARP, a large-scale dataset of 7th-order Higher-Order Ambisonic Room Impulse Responses (HOA-RIRs) consisting of more than 100,000 RIRs generated via a hybrid acoustic simulation in realistic indoor scenes. HiFi-HARP…

声音 · 计算机科学 2025-10-27 Shivam Saini , Jürgen Peissig

Spatial audio formats like Ambisonics are playback device layout-agnostic and well-suited for applications such as teleconferencing and virtual reality. Conventional Ambisonic encoding methods often rely on spherical microphone arrays for…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Yue Qiao , Vinay Kothapally , Meng Yu , Dong Yu

Implicit Neural Representations (INRs) are nowadays used to represent multimedia signals across various real-life applications, including image super-resolution, image compression, or 3D rendering. Existing methods that leverage INRs are…

机器学习 · 计算机科学 2023-06-21 Filip Szatkowski , Karol J. Piczak , Przemysław Spurek , Jacek Tabor , Tomasz Trzciński

In this work, we present a new multi-view depth estimation method that utilizes both conventional reconstruction and learning-based priors over the recently proposed neural radiance fields (NeRF). Unlike existing neural network based…

计算机视觉与模式识别 · 计算机科学 2021-10-06 Yi Wei , Shaohui Liu , Yongming Rao , Wang Zhao , Jiwen Lu , Jie Zhou

Retrospective MRI harmonization is limited by poor scalability across modalities and reliance on traveling subject datasets. To address these challenges, we introduce IHF-Harmony, a unified invertible hierarchy flow framework for…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Pengli Zhu , Yitao Zhu , Haowen Pang , Anqi Qiu

Many purely neural network based speech separation approaches have been proposed to improve objective assessment scores, but they often introduce nonlinear distortions that are harmful to modern automatic speech recognition (ASR) systems.…

音频与语音处理 · 电气工程与系统科学 2021-11-17 Zhuohuang Zhang , Yong Xu , Meng Yu , Shi-Xiong Zhang , Lianwu Chen , Donald S. Williamson , Dong Yu