English
Related papers

Related papers: RBA-FE: A Robust Brain-Inspired Audio Feature Extr…

200 papers

Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes…

Deepfake speech detection systems are often limited to binary classification tasks and struggle to generate interpretable reasoning or provide context-rich explanations for their decisions. These models primarily extract latent embeddings…

Sound · Computer Science 2026-04-01 Runkun Chen , Yixiong Fang , Pengyu Chang , Yuante Li , Massa Baali , Bhiksha Raj

Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domains. This trend has extended to respiratory and heart sound…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-28 Daisuke Niizumi , Daiki Takeuchi , Masahiro Yasuda , Binh Thien Nguyen , Yasunori Ohishi , Noboru Harada

Major depressive disorder (MDD) is a common mental disorder that typically affects a person's mood, cognition, behavior, and physical health. Resting-state functional magnetic resonance imaging (rs-fMRI) data are widely used for…

Image and Video Processing · Electrical Eng. & Systems 2024-06-10 Yunling Ma , Chaojun Zhang , Xiaochuan Wang , Qianqian Wang , Liang Cao , Limei Zhang , Mingxia Liu

The objective of this study is to derive functional networks for the autism spectrum disorder (ASD) population using the group ICA and dictionary learning model together and to classify ASD and typically developing (TD) participants using…

Neurons and Cognition · Quantitative Biology 2021-06-17 Xin Yang , Ning Zhang , Donglin Wang

Current audio deepfake detection has achieved remarkable performance using diverse deep learning architectures such as ResNet, and has seen further improvements with the introduction of large models (LMs) like Wav2Vec. The success of large…

Sound · Computer Science 2026-03-27 Yupei Li , Shuaijie Shao , Manuel Milling , Björn Schuller

Alzheimers disease is a fatal progressive brain disorder that worsens with time. It is high time we have inexpensive and quick clinical diagnostic techniques for early detection and care. In previous studies, various Machine Learning…

Computation and Language · Computer Science 2021-09-27 Akshay Valsaraj , Ithihas Madala , Nikhil Garg , Veeky Baths

In this paper, we introduce a framework ARBEx, a novel attentive feature extraction framework driven by Vision Transformer with reliability balancing to cope against poor class distributions, bias, and uncertainty in the facial expression…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Azmine Toushik Wasi , Karlo Šerbetar , Raima Islam , Taki Hasan Rafi , Dong-Kyu Chae

Preserving a patient's identity is a challenge for automatic, speech-based diagnosis of mental health disorders. In this paper, we address this issue by proposing adversarial disentanglement of depression characteristics and speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-08 Vijay Ravi , Jinhan Wang , Jonathan Flint , Abeer Alwan

Resting-state functional Magnetic Resonance Imaging (R-fMRI) holds the promise to reveal functional biomarkers of neuropsychiatric disorders. However, extracting such biomarkers is challenging for complex multi-faceted neuropatholo-gies,…

Environmental audio tagging aims to predict only the presence or absence of certain acoustic events in the interested acoustic scene. In this paper we make contributions to audio tagging in two parts, respectively, acoustic modeling and…

With the emergence of AI techniques for depression diagnosis, the conflict between high demand and limited supply for depression screening has been significantly alleviated. Among various modal data, audio-based depression diagnosis has…

Cryptography and Security · Computer Science 2026-03-27 Xintao Hu , Feng-Qi Cui

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additive noise. The…

Computation and Language · Computer Science 2015-03-31 Matthew Ager , Zoran Cvetkovic , Peter Sollich

Speaker extraction aims to extract target speech signal from a multi-talker environment with interference speakers and surrounding noise, given the target speaker's reference information. Most speaker extraction systems achieve satisfactory…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-12 Chengyun Deng , Shiqian Ma , Yi Zhang , Yongtao Sha , Hui Zhang , Hui Song , Xiangang Li

Operator learning is a data-driven approximation of mappings between infinite-dimensional function spaces, such as the solution operators of partial differential equations. Kernel-based operator learning can offer accurate, theoretically…

Machine Learning · Computer Science 2025-12-22 Xinyue Yu , Hayden Schaeffer

Acoustic scene classification (ASC) aims to identify the type of scene (environment) in which a given audio signal is recorded. The log-mel feature and convolutional neural network (CNN) have recently become the most popular time-frequency…

Sound · Computer Science 2021-08-12 Yuzhong Wu , Tan Lee

Improving the interpretability of deep neural networks has recently gained increased attention, especially when the power of deep learning is leveraged to solve problems in physics. Interpretability helps us understand a model's ability to…

Sound · Computer Science 2023-10-12 Karim Helwani , Erfan Soltanmohammadi , Michael M. Goodwin

In few-shot learning (FSL), the labeled samples are scarce. Thus, label errors can significantly reduce classification accuracy. Since label errors are inevitable in realistic learning tasks, improving the robustness of the model in the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Nan Xiang , Lifeng Xing , Dequan Jin

Recently, Meta-Black-Box-Optimization (MetaBBO) methods significantly enhance the performance of traditional black-box optimizers through meta-learning flexible and generalizable meta-level policies that excel in dynamic algorithm…

Neural and Evolutionary Computing · Computer Science 2025-03-25 Hongshu Guo , Sijie Ma , Zechuan Huang , Yuzhi Hu , Zeyuan Ma , Xinglin Zhang , Yue-Jiao Gong

We present AFEN (Audio Feature Ensemble Learning), a model that leverages Convolutional Neural Networks (CNN) and XGBoost in an ensemble learning fashion to perform state-of-the-art audio classification for a range of respiratory diseases.…

Sound · Computer Science 2024-05-10 Rahul Nadkarni , Emmanouil Nikolakakis , Razvan Marinescu