English
Related papers

Related papers: iMagLS: Interaural Level Difference with Magnitude…

200 papers

High-quality audio is essential in a wide range of applications, including online communication, virtual assistants, and the multimedia industry. However, degradation caused by noise, compression, and transmission artifacts remains a major…

This paper investigates the inverse capabilities and broader utility of multimodal latent spaces within task-specific AI (Artificial Intelligence) models. While these models excel at their designed forward tasks (e.g., text-to-image…

Machine Learning · Computer Science 2025-08-01 Siwoo Park

In Self-Supervised Learning (SSL), various pretext tasks are designed for learning feature representations through contrastive loss. However, previous studies have shown that this loss is less tolerant to semantically similar samples due to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-07 Shanshan Wang , Soumya Tripathy , Annamaria Mesaros

Transformer based end-to-end modelling approaches with multiple stream inputs have been achieved great success in various automatic speech recognition (ASR) tasks. An important issue associated with such approaches is that the intermediate…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-11 Jin Li , Rongfeng Su , Xurong Xie , Nan Yan , Lan Wang

Synthesizing medical images while preserving their structural information is crucial in medical research. In such scenarios, the preservation of anatomical content becomes especially important. Although recent advances have been made by…

Image and Video Processing · Electrical Eng. & Systems 2024-11-14 Ziqi Yu , Botao Zhao , Shengjie Zhang , Xiang Chen , Jianfeng Feng , Tingying Peng , Xiao-Yong Zhang

Virtual sound synthesis is a technology that allows users to perceive spatial sound through headphones or earphones. However, accurate virtual sound requires an individual head-related transfer function (HRTF), which can be difficult to…

Sound · Computer Science 2023-10-24 Tatsuki Kobayashi , Yoshiko Maruyama , Isao Nambu , Shohei Yano , Yasuhiro Wada

Head-related transfer functions (HRTFs) describe the directional filtering of the incoming sound caused by the morphology of a listener's head and pinnae. When an accurate model of a listener's morphology exists, HRTFs can be calculated…

Numerical Analysis · Mathematics 2016-07-28 Harald Ziegelwanger , Wolfgang Kreuzer , Piotr Majdak

In the growing field of virtual auditory display, personalized head-related transfer functions (HRTFs) play a vital role in establishing an accurate sound image for mixed and augmented reality applications. In this work, we propose an HRTF…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-06 Yuxiang Wang , You Zhang , Zhiyao Duan , Mark Bocko

Users often possess a clear visual intent but struggle to articulate it precisely in language. This intention-expression gap makes aligning generated images with latent visual preferences a fundamental challenge in text-to-image diffusion…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Wenxi Wang , Hongbin Liu , Mingqian Li , Junyan Yuan , Junqi Zhang

Learning from a label distribution has achieved promising results on ordinal regression tasks such as facial age and head pose estimation wherein, the concept of adaptive label distribution learning (ALDL) has drawn lots of attention…

Computer Vision and Pattern Recognition · Computer Science 2022-04-04 Qiang Li , Jingjing Wang , Zhaoliang Yao , Yachun Li , Pengju Yang , Jingwei Yan , Chunmao Wang , Shiliang Pu

Clinical monitoring of functional decline in ALS relies on periodic assessments that may miss critical changes occurring between visits. To address this gap, semi-supervised regression models were developed to estimate rates of decline in a…

Machine Learning · Computer Science 2025-07-15 Noah Marchal , William E. Janes , Mihail Popescu , Xing Song

While Transformer has become the de-facto standard for speech, modeling upon the fine-grained frame-level features remains an open challenge of capturing long-distance dependencies and distributing the attention weights. We propose…

Computation and Language · Computer Science 2023-05-30 Chen Xu , Yuhao Zhang , Chengbo Jiao , Xiaoqian Liu , Chi Hu , Xin Zeng , Tong Xiao , Anxiang Ma , Huizhen Wang , JingBo Zhu

Large Audio-Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information, particularly in the multi-modal fusion layers of the Transformer architecture. This bias hinders their…

Sound · Computer Science 2025-09-24 Junyu Wang , Ziyang Ma , Zhengding Luo , Tianrui Wang , Meng Ge , Xiaobao Wang , Longbiao Wang

The estimation of glottal flow from a speech waveform is a key method for speech analysis and parameterization. Significant research effort has been made to dissociate the first vocal tract resonance from the glottal formant (the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-09 Olivier Perrotin , Ian Vince McLoughlin

Advanced auditory models are useful in designing signal-processing algorithms for hearing-loss compensation or speech enhancement. Such auditory models provide rich and detailed descriptions of the auditory pathway, and might allow for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-18 Peter Leer , Jesper Jensen , Zheng-Hua Tan , Jan Østergaard , Lars Bramsløw

Video-assisted transoral tracheal intubation (TI) necessitates using an endoscope that helps the physician insert a tracheal tube into the glottis instead of the esophagus. The growing trend of robotic-assisted TI would require a medical…

Artificial Intelligence · Computer Science 2023-07-31 Guankun Wang , Tian-Ao Ren , Jiewen Lai , Long Bai , Hongliang Ren

While the spatial directivity of multichannel speech enhancement algorithms improves with the number of microphones, fitting large capture arrays into real-world edge devices is typically limited by physical constraints. To overcome this…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-08 Dongheon Lee , Ashutosh Pandey , Sanjeel Parekh , Daniel Wong , Jacob Donley , Buye Xu , Juan Azcarreta

Ultrahigh field (UHF) Magnetic Resonance Imaging (MRI) provides a higher signal-to-noise ratio and, thereby, higher spatial resolution. However, UHF MRI introduces challenges such as transmit radiofrequency (RF) field (B1+) inhomogeneities,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-07 Zhengyi Lu , Hao Liang , Xiao Wang , Xinqiang Yan , Yuankai Huo

Implicit Neural Representations (INRs) have emerged as a paradigm in knowledge representation, offering exceptional flexibility and performance across a diverse range of applications. INRs leverage multilayer perceptrons (MLPs) to model…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Amer Essakine , Yanqi Cheng , Chun-Wun Cheng , Lipei Zhang , Zhongying Deng , Lei Zhu , Carola-Bibiane Schönlieb , Angelica I Aviles-Rivero

Portable, ultra-low-field (ULF) magnetic resonance imaging has the potential to expand access to neuroimaging but currently suffers from coarse spatial and angular resolutions and low signal-to-noise ratios. Diffusion tensor imaging (DTI),…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Mark D. Olchanyi , Annabel Sorby-Adams , John Kirsch , Brian L. Edlow , Ava Farnan , Renfei Liu , Matthew S. Rosen , Emery N. Brown , W. Taylor Kimberly , Juan Eugenio Iglesias