English
Related papers

Related papers: Interpreting Audiograms with Multi-stage Neural Ne…

200 papers

Audio agents extend large audio-language models (LALMs) by decomposing audio questions into tool calls, intermediate evidence, and iterative reasoning steps. However, as LALMs become stronger, the key challenge shifts from enabling tool use…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-28 Yucheng Wang , Jing Peng , Hanqi Li , Chenghao Wang , Wenming Tu , Yu Xi , Zhaokai Sun , Kai Yu , Shuai Wang

Brain signals constitute the information that are processed by millions of brain neurons (nerve cells and brain cells). These brain signals can be recorded and analyzed using various of non-invasive techniques such as the…

Neurons and Cognition · Quantitative Biology 2022-01-13 Almabrok Essa , Hari Kotte

Many visualizations have been developed for explainable AI (XAI), but they often require further reasoning by users to interpret. Investigating XAI for high-stakes medical diagnosis, we propose improving domain alignment with diagrammatic…

Artificial Intelligence · Computer Science 2025-02-27 Brian Y. Lim , Joseph P. Cahaly , Chester Y. F. Sng , Adam Chew

The proliferation of high-dimensional datasets in fields such as genomics, healthcare, and finance has created an urgent need for machine learning models that are both highly accurate and inherently interpretable. While traditional deep…

Machine Learning · Computer Science 2025-10-28 Rekha R Nair , Tina Babu , Alavikunhu Panthakkan , Hussain Al-Ahmad , Balamurugan Balusamy

Personalized speech intelligibility prediction is challenging. Previous approaches have mainly relied on audiograms, which are inherently limited in accuracy as they only capture a listener's hearing threshold for pure tones. Rather than…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Haoshuai Zhou , Changgeng Mo , Boxuan Cao , Linkai Li , Shan Xiang Wang

Advanced auditory models are useful in designing signal-processing algorithms for hearing-loss compensation or speech enhancement. Such auditory models provide rich and detailed descriptions of the auditory pathway, and might allow for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-18 Peter Leer , Jesper Jensen , Zheng-Hua Tan , Jan Østergaard , Lars Bramsløw

Mammogram mass detection is crucial for diagnosing and preventing the breast cancers in clinical practice. The complementary effect of multi-view mammogram images provides valuable information about the breast anatomical prior structure and…

Computer Vision and Pattern Recognition · Computer Science 2021-05-24 Yuhang Liu , Fandong Zhang , Chaoqi Chen , Siwen Wang , Yizhou Wang , Yizhou Yu

Recent advances in deep learning for tomographic reconstructions have shown great potential to create accurate and high quality images with a considerable speed-up. In this work we present a deep neural network that is specifically designed…

Computer Vision and Pattern Recognition · Computer Science 2020-09-07 Andreas Hauptmann , Felix Lucka , Marta Betcke , Nam Huynh , Jonas Adler , Ben Cox , Paul Beard , Sebastien Ourselin , Simon Arridge

Auditory attention decoding (AAD) is the process of identifying the attended speech in a multi-talker environment using brain signals, typically recorded through electroencephalography (EEG). Over the past decade, AAD has undergone…

Sound · Computer Science 2025-07-08 Nhan Duc Thanh Nguyen , Huy Phan , Simon Geirnaert , Kaare Mikkelsen , Preben Kidmose

Explainable Artificial Intelligence (XAI) has emerged as a critical tool for interpreting the predictions of complex deep learning models. While XAI has been increasingly applied in various domains within acoustics, its use in bioacoustics,…

Sound · Computer Science 2025-09-11 Zubair Faruqui , Mackenzie S. McIntire , Rahul Dubey , Jay McEntee

Magnetoencephalography (MEG) is a cutting-edge neuroimaging technique that measures the intricate brain dynamics underlying cognitive processes with an unparalleled combination of high temporal and spatial precision. MEG data analytics has…

Neurons and Cognition · Quantitative Biology 2025-05-20 Arthur Dehgan , Hamza Abdelhedi , Vanessa Hadid , Irina Rish , Karim Jerbi

We propose a novel approach for time-scale modification of audio signals. Unlike traditional methods that rely on the framing technique or the short-time Fourier transform to preserve the frequency during temporal stretching, our neural…

Sound · Computer Science 2023-10-09 Ernie Chu , Ju-Ting Chen , Chia-Ping Chen

Deep learning architectures have made significant progress in terms of performance in many research areas. The automatic speech recognition (ASR) field has thus benefited from these scientific and technological advances, particularly for…

Sound · Computer Science 2024-03-01 Quentin Raymondaud , Mickael Rouvier , Richard Dufour

Music auto-tagging is often handled in a similar manner to image classification by regarding the 2D audio spectrogram as image data. However, music auto-tagging is distinguished from image classification in that the tags are highly diverse…

Neural and Evolutionary Computing · Computer Science 2017-08-02 Jongpil Lee , Juhan Nam

Musicians and audio engineers sculpt and transform their sounds by connecting multiple processors, forming an audio processing graph. However, most deep-learning methods overlook this real-world practice and assume fixed graph settings. To…

Sound · Computer Science 2023-05-09 Sungho Lee , Jaehyun Park , Seungryeol Paik , Kyogu Lee

The insufficient supervision limit the performance of the deep supervised models for brain disease diagnosis. It is important to develop a learning framework that can capture more information in limited data and insufficient supervision. To…

Neurons and Cognition · Quantitative Biology 2024-10-10 Wenjing Gao , Yuanyuan Yang , Jianrui Wei , Xuntao Yin , Xinhan Di

Dysarthria is a disability that causes a disturbance in the human speech system and reduces the quality and intelligibility of a person's speech. Because of this effect, the normal speech processing systems can not work properly on impaired…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-22 Aref Farhadipour , Hadi Veisi

Ultrasound (US) imaging is widely used for biometric measurement and diagnosis of internal organs due to the advantages of being real-time and radiation-free. However, due to inter-operator variations, resulting images highly depend on the…

Robotics · Computer Science 2023-11-30 Zhongliang Jiang , Yuan Bi , Mingchuan Zhou , Ying Hu , Michael Burke , Nassir Navab

Reliable detection of the prodromal stages of Alzheimer's disease (AD) remains difficult even today because, unlike other neurocognitive impairments, there is no definitive diagnosis of AD in vivo. In this context, existing research has…

Machine Learning · Computer Science 2021-08-03 Amish Mittal , Sourav Sahoo , Arnhav Datar , Juned Kadiwala , Hrithwik Shalu , Jimson Mathew

The paper presents Multi-layer Auto Resonance Networks (ARN), a new neural model, for image recognition. Neurons in ARN, called Nodes, latch on to an incoming pattern and resonate when the input is within its 'coverage.' Resonance allows…

Computer Vision and Pattern Recognition · Computer Science 2020-10-12 Shilpa Mayannavar , Uday Wali , V M Aparanji