English
Related papers

Related papers: GeHirNet: A Gender-Aware Hierarchical Model for Vo…

200 papers

Audio deepfake detection aims to detect real human voices from those generated by Artificial Intelligence (AI) and has emerged as a significant problem in the field of voice biometrics systems. With the ever-improving quality of synthetic…

Sound · Computer Science 2026-05-12 Aishwarya Fursule , Shruti Kshirsagar , Anderson R. Avila

Speech-based AI models are emerging as powerful tools for detecting depression and the presence of Post-traumatic stress disorder (PTSD), offering a non-invasive and cost-effective way to assess mental health. However, these models often…

Artificial Intelligence · Computer Science 2025-05-07 June-Woo Kim , Haram Yoon , Wonkyo Oh , Dawoon Jung , Sung-Hoon Yoon , Dae-Jin Kim , Dong-Ho Lee , Sang-Yeol Lee , Chan-Mo Yang

In recent years, deep-learning-based speech emotion recognition models have outperformed classical machine learning models. Previously, neural network designs, such as Multitask Learning, have accounted for variations in emotional…

Machine Learning · Computer Science 2021-09-10 Lance Ying , Amrit Romana , Emily Mower Provost

Audio-based disease prediction is emerging as a promising supplement to traditional medical diagnosis methods, facilitating early, convenient, and non-invasive disease detection and prevention. Multimodal fusion, which integrates features…

Alzheimer's disease and related dementias (ADRD) affect one in five adults over 60, yet more than half of individuals with cognitive decline remain undiagnosed. Speech-based assessments show promise for early detection, as phonetic motor…

Automated sleep stage classification typically employs a single population-agnostic model, disregarding established demographic variations in sleep architecture. Sleep patterns, however, differ substantially across gender, age, and…

Machine Learning · Computer Science 2026-05-05 S M Asif Hossain , Shruti Kshirsagar

Background: Twelve lead ECGs are a core diagnostic tool for cardiovascular diseases. Here, we describe and analyse an ensemble deep neural network architecture to classify 24 cardiac abnormalities from 12-lead ECGs. Method: We proposed a…

Signal Processing · Electrical Eng. & Systems 2022-04-13 Zhibin Zhao , Darcy Murphy , Hugh Gifford , Stefan Williams , Annie Darlington , Samuel D. Relton , Hui Fang , David C. Wong

Early and accurate detection systems for ear diseases, powered by deep learning, are essential for preventing hearing impairment and improving population health. However, the limited diversity of existing otoendoscopy datasets and the poor…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Feiyan Lu , Yubiao Yue , Zhenzhang Li , Meiping Zhang , Wen Luo , Fan Zhang , Tong Liu , Jingyong Shi , Guang Wang , Xinyu Zeng

In this paper we extend the x-vector framework for the task of speaker's age estimation and gender classification. In particular, we replace the baseline multilayer-TDNN architecture with QuartzNet, a convolutional architecture that has…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-04 Damian Kwasny , Daria Hemmerling

Background: Breast cancer has the highest prevalence in women globally. The classification and diagnosis of breast cancer and its histopathological images have always been a hot spot of clinical concern. In Computer-Aided Diagnosis (CAD),…

Computer Vision and Pattern Recognition · Computer Science 2022-05-11 Yuchao Zheng , Chen Li , Xiaomin Zhou , Haoyuan Chen , Hao Xu , Yixin Li , Haiqing Zhang , Xiaoyan Li , Hongzan Sun , Xinyu Huang , Marcin Grzegorzek

There is a strong need for automated systems to improve diagnostic quality and reduce the analysis time in histopathology image processing. Automated detection and classification of pathological tissue characteristics with computer-aided…

Computer Vision and Pattern Recognition · Computer Science 2019-11-22 Muhammed Talo

AI models for medical diagnosis often exhibit uneven performance across patient populations due to heterogeneity in disease prevalence, imaging appearance, and clinical risk profiles. Existing algorithmic fairness approaches typically seek…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Gelei Xu , Yuying Duan , Jun Xia , Ruining Deng , Wei Jin , Yiyu Shi

Molecular subtyping of breast cancer is crucial for personalized treatment and prognosis. Traditional classification approaches rely on either histopathological images or gene expression profiling, limiting their predictive power. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Amin Honarmandi Shandiz

Learning to classify time series with limited data is a practical yet challenging problem. Current methods are primarily based on hand-designed feature extraction rules or domain-specific data augmentation. Motivated by the advances in deep…

Machine Learning · Computer Science 2022-01-17 Chao-Han Huck Yang , Yun-Yun Tsai , Pin-Yu Chen

Capturing high-level structure in audio waveforms is challenging because a single second of audio spans tens of thousands of timesteps. While long-range dependencies are difficult to model directly in the time domain, we show that they can…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-05 Sean Vasquez , Mike Lewis

Recent progress has been made in detecting early stage dementia entirely through recordings of patient speech. Multimodal speech analysis methods were applied to the PROCESS challenge, which requires participants to use audio recordings of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-14 Lei Chi , Arav Sharma , Ari Gebhardt , Joseph T. Colonel

Advances in artificial intelligence (AI) show great potential in revealing underlying information from phonon microscopy (high-frequency ultrasound) data to identify cancerous cells. However, this technology suffers from the 'batch effect'…

Quantitative Methods · Quantitative Biology 2024-03-28 Yijie Zheng , Rafael Fuentes-Dominguez , Matt Clark , George S. D. Gordon , Fernando Perez-Cota

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. To improve robustness of speaker recognition system performance in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Yanpei Shi , Qiang Huang , Thomas Hain

The recent developments in technology have re-warded us with amazing audio synthesis models like TACOTRON and WAVENETS. On the other side, it poses greater threats such as speech clones and deep fakes, that may go undetected. To tackle…

Machine Learning · Computer Science 2021-07-27 Arun Kumar Singh , Priyanka Singh , Karan Nathwani

Accurate identification of breast cancer types plays a critical role in guiding treatment decisions and improving patient outcomes. This paper presents an artificial intelligence enabled tool designed to aid in the identification of breast…

Image and Video Processing · Electrical Eng. & Systems 2025-05-28 Neil Chaudhary , Zaynah Dhunny