中文
相关论文

相关论文: A Generalist Audio Foundation Model for Comprehens…

200 篇论文

Artificial intelligence (AI) that can effectively learn ultrasound representations by integrating multi-source data holds significant promise for advancing clinical care. However, the scarcity of large labeled datasets in real-world…

With the advancement of audio generation, generative models can produce highly realistic audios. However, the proliferation of deepfake general audio can pose negative consequences. Therefore, we propose a new task, deepfake general audio…

声音 · 计算机科学 2024-06-13 Zeyu Xie , Baihan Li , Xuenan Xu , Zheng Liang , Kai Yu , Mengyue Wu

Respiratory auscultation is crucial for early detection of pediatric pneumonia, a condition that can quickly worsen without timely intervention. In areas with limited physician access, effective auscultation is challenging. We present a…

人机交互 · 计算机科学 2025-04-23 Seung Gyu Jeong , Sung Woo Nam , Seong Kwan Jung , Seong-Eun Kim

General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separation models are either domain-specific, designed for fixed…

Respiratory diseases are among the most common causes of severe illness and death worldwide. Prevention and early diagnosis are essential to limit or even reverse the trend that characterizes the diffusion of such diseases. In this regard,…

音频与语音处理 · 电气工程与系统科学 2019-07-15 Diego Perna , Andrea Tagarelli

Universal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music. Yet, the potential of universal source separation is limited because most…

声音 · 计算机科学 2023-10-03 Jordi Pons , Xiaoyu Liu , Santiago Pascual , Joan Serrà

Listening to heart and lung sounds - auscultation - is one of the first and most fundamental steps in a clinical examination. Despite being fast and non-invasive, it demands years of experience to interpret subtle audio cues. Recent deep…

机器学习 · 计算机科学 2026-03-03 Yishan Wang , Tsai-Ning Wang , Mathias Funk , Aaqib Saeed

Ultrasound is widely used in clinical practice due to its affordability, portability, and safety. However, current AI research often overlooks combined disease prediction and tissue segmentation. We propose UniUSNet, a universal framework…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Zehui Lin , Zhuoneng Zhang , Xindi Hu , Zhifan Gao , Xin Yang , Yue Sun , Dong Ni , Tao Tan

This review is focused on the data-driven approaches applied in different applications of Acoustic-to-Articulatory Inversion (AAI) of speech. This review paper considered the relevant works published in the last ten years (2011-2021). The…

声音 · 计算机科学 2025-04-21 Leena G Pillai , D. Muhammad Noorul Mubarak

There has been a surge of interest in leveraging speech as a marker of health for a wide spectrum of conditions. The underlying premise is that any neurological, mental, or physical deficits that impact speech production can be objectively…

音频与语音处理 · 电气工程与系统科学 2024-10-30 Si-Ioi Ng , Lingfeng Xu , Ingo Siegert , Nicholas Cummins , Nina R. Benway , Julie Liss , Visar Berisha

Accurately interpreting cardiac auscultation signals plays a crucial role in diagnosing and managing cardiovascular diseases. However, the paucity of labelled data inhibits classification models' training. Researchers have turned to…

声音 · 计算机科学 2025-06-18 Leigh Abbott , Milan Marocchi , Matthew Fynn , Yue Rong , Sven Nordholm

Traditional echocardiographic parameters such as ejection fraction (EF) and global longitudinal strain (GLS) have limitations in the early detection of cardiac dysfunction. EF often remains normal despite underlying pathology, and GLS is…

机器学习 · 计算机科学 2025-07-21 Beka Begiashvili , Carlos J. Fernandez-Candel , Matías Pérez Paredes

IMPORTANCE: Modern ultrasound systems are universal diagnostic tools capable of imaging the entire body. However, current AI solutions remain fragmented into single-task tools. This critical gap between hardware versatility and software…

In the ever-evolving landscape of medical diagnostics, this study details the systematic design process and concept selection methodology for developing an advanced digital stethoscope, demonstrating the evolution from traditional acoustic…

人机交互 · 计算机科学 2024-12-20 Abraham G. Taye , Sador Yemane , Eshetu Negash , Yared Minwuyelet , Nebiha Tofik

The increase in cardiac and pulmonary diseases presents an alarming and pervasive health challenge on a global scale responsible for unexpected and premature mortalities. In spite of how serious these conditions are, existing methods of…

信号处理 · 电气工程与系统科学 2026-05-12 Hania Ghouse , Juveria Tanveen , Abdul Muqtadir Ahmed , Uma N. Dulhare

Masked token prediction has emerged as a powerful pre-training objective across language, vision, and speech, offering the potential to unify these diverse modalities through a single pre-training task. However, its application for general…

Body sounds provide rich information about the state of the human body and can be useful in many medical applications. Auscultation, the practice of listening to body sounds, has been used for centuries in respiratory and cardiac medicine…

人机交互 · 计算机科学 2020-08-13 Shyam A. Tailor , Jagmohan Chauhan , Cecilia Mascolo

Heart sound auscultation has been applied in clinical usage for early screening of cardiovascular diseases. Due to the high demand for auscultation expertise, automatic auscultation can help with auxiliary diagnosis and reduce the burden of…

声音 · 计算机科学 2024-05-14 Zhao Ren , Yi Chang , Thanh Tam Nguyen , Yang Tan , Kun Qian , Björn W. Schuller

Hypernasality is a common characteristic symptom across many motor-speech disorders. For voiced sounds, hypernasality introduces an additional resonance in the lower frequencies and, for unvoiced sounds, there is reduced articulatory…

音频与语音处理 · 电气工程与系统科学 2020-09-14 Michael Saxon , Ayush Tripathi , Yishan Jiao , Julie Liss , Visar Berisha

Respiratory diseases remain major global health challenges, and traditional auscultation is often limited by subjectivity, environmental noise, and inter-clinician variability. This study presents an explainable multimodal deep learning…

声音 · 计算机科学 2025-12-02 S M Asiful Islam Saky , Md Rashidul Islam , Md Saiful Arefin , Shahaba Alam
‹ 上一页 1 2 3 10 下一页 ›