中文
相关论文

相关论文: Interpreting Audiograms with Multi-stage Neural Ne…

200 篇论文

Automated interpretation of electrocardiograms (ECG) has garnered significant attention with the advancements in machine learning methodologies. Despite the growing interest, most current studies focus solely on classification or regression…

信号处理 · 电气工程与系统科学 2023-11-07 Jielin Qiu , Jiacheng Zhu , Shiqi Liu , William Han , Jingqi Zhang , Chaojing Duan , Michael Rosenberg , Emerson Liu , Douglas Weber , Ding Zhao

Advancements in generative Artificial Intelligence (AI) hold great promise for automating radiology workflows, yet challenges in interpretability and reliability hinder clinical adoption. This paper presents an automated radiology report…

Seeing is believing, however, the underlying mechanism of how human visual perceptions are intertwined with our cognitions is still a mystery. Thanks to the recent advances in both neuroscience and artificial intelligence, we have been able…

图像与视频处理 · 电气工程与系统科学 2023-08-17 Yu-Ting Lan , Kan Ren , Yansen Wang , Wei-Long Zheng , Dongsheng Li , Bao-Liang Lu , Lili Qiu

In this paper we present our researches regarding automat parsing of audio recordings. These recordings are obtained from children with dyslalia and are necessary for an accurate identification of speech problems. We develop the ADM…

计算机与社会 · 计算机科学 2014-06-20 Ovidiu-Andrei Schipor , Titus-Marian Nestor

Replay attacks remain a critical vulnerability for automatic speaker verification systems, particularly in real-time voice assistant applications. In this work, we propose acoustic maps as a novel spatial feature representation for replay…

音频与语音处理 · 电气工程与系统科学 2026-05-21 Michael Neri , Tuomas Virtanen

Photoacoustic (PA) imaging has the potential to revolutionize functional medical imaging in healthcare due to the valuable information on tissue physiology contained in multispectral photoacoustic measurements. Clinical translation of the…

Deep neural networks have been used widely to learn the latent structure of datasets, across modalities such as images, shapes, and audio signals. However, existing models are generally modality-dependent, requiring custom architectures and…

机器学习 · 计算机科学 2021-11-12 Yilun Du , Katherine M. Collins , Joshua B. Tenenbaum , Vincent Sitzmann

Transformers have set new benchmarks in audio processing tasks, leveraging self-attention mechanisms to capture complex patterns and dependencies within audio data. However, their focus on pairwise interactions limits their ability to…

声音 · 计算机科学 2025-01-30 Shubhr Singh , Emmanouil Benetos , Huy Phan , Dan Stowell

MUltiple SIgnal Classification (MUSIC) and Estimation of signal parameters via rotational via rotational invariance (ESPRIT) has been widely used in super resolution direction of arrival estimation (DoA) in both Uniform Linear Arrays (ULA)…

信号处理 · 电气工程与系统科学 2020-02-04 Jianyuan Yu

Usage of the fast development of real-life digital applications in modern technology should guarantee novel and efficient way-outs of their protection. Encryption facilitates the data hiding. With the express development of technology,…

密码学与安全 · 计算机科学 2025-02-26 Sachith Dassanayaka

Optical Music Recognition (OMR) is concerned with transcribing sheet music into a machine-readable format. The transcribed copy should allow musicians to compose, play and edit music by taking a picture of a music sheet. Complete…

计算机视觉与模式识别 · 计算机科学 2020-06-23 Elona Shatri , György Fazekas

Perception of auditory events is inherently multimodal relying on both audio and visual cues. A large number of existing multimodal approaches process each modality using modality-specific models and then fuse the embeddings to encode the…

声音 · 计算机科学 2022-07-19 Amir Shirian , Krishna Somandepalli , Victor Sanchez , Tanaya Guha

Despite the recent success of machine learning algorithms, most models face drawbacks when considering more complex tasks requiring interaction between different sources, such as multimodal input data and logical time sequences. On the…

声音 · 计算机科学 2023-02-01 Leandro A. Passos , João Paulo Papa , Amir Hussain , Ahsan Adeel

Neuro-steered speaker extraction aims to extract the listener's brain-attended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical activity. This activity is usually recorded using…

音频与语音处理 · 电气工程与系统科学 2023-12-13 Zexu Pan , Gordon Wichern , Francois G. Germain , Sameer Khurana , Jonathan Le Roux

Machine learning has been developed dramatically and witnessed a lot of applications in various fields over the past few years. This boom originated in 2009, when a new model emerged, that is, the deep artificial neural network, which began…

计算机视觉与模式识别 · 计算机科学 2020-12-03 Changchun Yang , Hengrong Lan , Feng Gao , Fei Gao

Speech sounds are produced as the coordinated movement of the speaking organs. There are several available methods to model the relation of articulatory movements and the resulting speech signal. The reverse problem is often called as…

声音 · 计算机科学 2019-04-16 Dagoberto Porras , Alexander Sepúlveda-Sepúlveda , Tamás Gábor Csapó

Multi-channel acoustic signal processing is a well-established and powerful tool to exploit the spatial diversity between a target signal and non-target or noise sources for signal enhancement. However, the textbook solutions for optimal…

音频与语音处理 · 电气工程与系统科学 2025-01-14 Reinhold Haeb-Umbach , Tomohiro Nakatani , Marc Delcroix , Christoph Boeddeker , Tsubasa Ochiai

Early diagnosis and discovery of therapeutic drug targets are crucial objectives for effective management of Alzheimer's Disease (AD). Current approaches for AD diagnosis and treatment planning are based on radiological imaging and largely…

机器学习 · 计算机科学 2026-05-01 Maryam Khalid , Fadeel Sher Khan , John Broussard , Arko Barman

In clinical practice, human radiologists actually review medical images with high resolution monitors and zoom into region of interests (ROIs) for a close-up examination. Inspired by this observation, we propose a hierarchical graph neural…

图像与视频处理 · 电气工程与系统科学 2019-12-17 Hao Du , Jiashi Feng , Mengling Feng

As deep learning systems are scaled up to many billions of parameters, relating their internal structure to external behaviors becomes very challenging. Although daunting, this problem is not new: Neuroscientists and cognitive scientists…