English
Related papers

Related papers: iMagLS: Interaural Level Difference with Magnitude…

200 papers

In this work, we propose a deep beamforming framework for speech enhancement in dynamic acoustic environments. The framework learns time-varying beamformer weights from noisy multichannel signals via a deep neural network, guided by a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-18 Ilai Zaidel , Sharon Gannot

End-to-end neural TTS training has shown improved performance in speech style transfer. However, the improvement is still limited by the training data in both target styles and speakers. Inadequate style transfer performance occurs when the…

Sound · Computer Science 2021-06-21 Xiaochun An , Frank K. Soong , Lei Xie

Ambisonics is a method for capturing and rendering a sound field accurately, assuming that the acoustics of the playback room does not significantly influence the sound field. However, in practice, the acoustics of the playback room may…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-14 Ali Fallah , Shun Nakamura , Steven van de Par

Hallucination remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While Direct Preference Optimization (DPO) is a key alignment framework, existing approaches often rely heavily on costly external evaluators for…

Machine Learning · Computer Science 2026-02-04 Yuanshuai Li , Yuping Yan , Jirui Han , Fei Ming , Lingjuan Lv , Yaochu Jin

As spatial audio is enjoying a surge in popularity, data-driven machine learning techniques that have been proven successful in other domains are increasingly used to process head-related transfer function measurements. However, these…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-09 Johan Pauwels , Lorenzo Picinali

Hearing loss (HL) simulators, which allow normal hearing (NH) listeners to experience HL, have been used in speech intelligibility experiments, but not in sound quality experiments due to perceptible distortion. If they produced less…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-25 Toshio Irino , Shintaro Doan , Minami Ishikawa

In this work, we propose a two-dimensional Head-Related Transfer Function (HRTF)-based robust beamformer design for robot audition, which allows for explicit control of the beamformer response for the entire three-dimensional sound field…

Sound · Computer Science 2017-03-10 Hendrik Barfuss , Michael Buerger , Jasper Podschus , Walter Kellermann

Remote sensing imagery suffers from clouds, haze, noise, resolution limits, and sensor heterogeneity. Existing restoration and fusion approaches train separate models per degradation type. In this work, we present Language-conditioned…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Yongchuan Cui , Peng Liu

The multichannel Wiener filter (MWF) and its variations have been extensively applied to binaural hearing aids. However, its major drawback is the distortion of the binaural cues of the residual noise, changing the original acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-12 Johnny Werner , Marcio H. Costa

Many efforts have been devoted to designing sampling, mining, and weighting strategies in high-level deep metric learning (DML) loss objectives. However, little attention has been paid to low-level but essential data transformation. In this…

Computer Vision and Pattern Recognition · Computer Science 2021-05-24 Zhiyuan Chen , Guang Yao , Wennan Ma , Lin Xu

Several individualization methods have recently been proposed to estimate a subject's Head-Related Transfer Function (HRTF) using convenient input modalities such as anthropometric measurements or pinnae photographs. There exists a need for…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-23 Etienne Thuillier , Craig Jin , Vesa Välimäki

Conventional automatic speech recognition (ASR) system uses second-order minkowski loss during inference time which is suboptimal as it incorporates only first order statistics in posterior estimation [2]. In this paper we have proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-03 Vishwanath Pratap Singh , Shakti P. Rath , Abhishek Pandey

Auditory large language models (ALLMs) have demonstrated strong general capabilities in audio understanding and reasoning tasks. However, their reliability is still undermined by hallucination issues. Existing hallucination evaluation…

Sound · Computer Science 2026-04-13 Qixuan Huang , Khalid Zaman , Masashi Unoki

AI-based neural decoding reconstructs visual perception by leveraging generative models to map brain activity, measured through functional MRI (fMRI), into latent hierarchical representations. Traditionally, ridge linear models transform…

Image and Video Processing · Electrical Eng. & Systems 2025-09-04 Lorenzo Veronese , Andrea Moglia , Luca Mainardi , Pietro Cerveri

We introduce HiFi-HARP, a large-scale dataset of 7th-order Higher-Order Ambisonic Room Impulse Responses (HOA-RIRs) consisting of more than 100,000 RIRs generated via a hybrid acoustic simulation in realistic indoor scenes. HiFi-HARP…

Sound · Computer Science 2025-10-27 Shivam Saini , Jürgen Peissig

Spatial audio formats like Ambisonics are playback device layout-agnostic and well-suited for applications such as teleconferencing and virtual reality. Conventional Ambisonic encoding methods often rely on spherical microphone arrays for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Yue Qiao , Vinay Kothapally , Meng Yu , Dong Yu

Implicit Neural Representations (INRs) are nowadays used to represent multimedia signals across various real-life applications, including image super-resolution, image compression, or 3D rendering. Existing methods that leverage INRs are…

Machine Learning · Computer Science 2023-06-21 Filip Szatkowski , Karol J. Piczak , Przemysław Spurek , Jacek Tabor , Tomasz Trzciński

In this work, we present a new multi-view depth estimation method that utilizes both conventional reconstruction and learning-based priors over the recently proposed neural radiance fields (NeRF). Unlike existing neural network based…

Computer Vision and Pattern Recognition · Computer Science 2021-10-06 Yi Wei , Shaohui Liu , Yongming Rao , Wang Zhao , Jiwen Lu , Jie Zhou

Retrospective MRI harmonization is limited by poor scalability across modalities and reliance on traveling subject datasets. To address these challenges, we introduce IHF-Harmony, a unified invertible hierarchy flow framework for…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Pengli Zhu , Yitao Zhu , Haowen Pang , Anqi Qiu

Many purely neural network based speech separation approaches have been proposed to improve objective assessment scores, but they often introduce nonlinear distortions that are harmful to modern automatic speech recognition (ASR) systems.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-17 Zhuohuang Zhang , Yong Xu , Meng Yu , Shi-Xiong Zhang , Lianwu Chen , Donald S. Williamson , Dong Yu
‹ Prev 1 3 4 5 6 7 10 Next ›