中文
相关论文

相关论文: RBA-FE: A Robust Brain-Inspired Audio Feature Extr…

200 篇论文

Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes…

Deepfake speech detection systems are often limited to binary classification tasks and struggle to generate interpretable reasoning or provide context-rich explanations for their decisions. These models primarily extract latent embeddings…

声音 · 计算机科学 2026-04-01 Runkun Chen , Yixiong Fang , Pengyu Chang , Yuante Li , Massa Baali , Bhiksha Raj

Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domains. This trend has extended to respiratory and heart sound…

音频与语音处理 · 电气工程与系统科学 2025-04-28 Daisuke Niizumi , Daiki Takeuchi , Masahiro Yasuda , Binh Thien Nguyen , Yasunori Ohishi , Noboru Harada

Major depressive disorder (MDD) is a common mental disorder that typically affects a person's mood, cognition, behavior, and physical health. Resting-state functional magnetic resonance imaging (rs-fMRI) data are widely used for…

图像与视频处理 · 电气工程与系统科学 2024-06-10 Yunling Ma , Chaojun Zhang , Xiaochuan Wang , Qianqian Wang , Liang Cao , Limei Zhang , Mingxia Liu

The objective of this study is to derive functional networks for the autism spectrum disorder (ASD) population using the group ICA and dictionary learning model together and to classify ASD and typically developing (TD) participants using…

神经元与认知 · 定量生物学 2021-06-17 Xin Yang , Ning Zhang , Donglin Wang

Current audio deepfake detection has achieved remarkable performance using diverse deep learning architectures such as ResNet, and has seen further improvements with the introduction of large models (LMs) like Wav2Vec. The success of large…

声音 · 计算机科学 2026-03-27 Yupei Li , Shuaijie Shao , Manuel Milling , Björn Schuller

Alzheimers disease is a fatal progressive brain disorder that worsens with time. It is high time we have inexpensive and quick clinical diagnostic techniques for early detection and care. In previous studies, various Machine Learning…

计算与语言 · 计算机科学 2021-09-27 Akshay Valsaraj , Ithihas Madala , Nikhil Garg , Veeky Baths

In this paper, we introduce a framework ARBEx, a novel attentive feature extraction framework driven by Vision Transformer with reliability balancing to cope against poor class distributions, bias, and uncertainty in the facial expression…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Azmine Toushik Wasi , Karlo Šerbetar , Raima Islam , Taki Hasan Rafi , Dong-Kyu Chae

Preserving a patient's identity is a challenge for automatic, speech-based diagnosis of mental health disorders. In this paper, we address this issue by proposing adversarial disentanglement of depression characteristics and speaker…

音频与语音处理 · 电气工程与系统科学 2023-06-08 Vijay Ravi , Jinhan Wang , Jonathan Flint , Abeer Alwan

Resting-state functional Magnetic Resonance Imaging (R-fMRI) holds the promise to reveal functional biomarkers of neuropsychiatric disorders. However, extracting such biomarkers is challenging for complex multi-faceted neuropatholo-gies,…

Environmental audio tagging aims to predict only the presence or absence of certain acoustic events in the interested acoustic scene. In this paper we make contributions to audio tagging in two parts, respectively, acoustic modeling and…

With the emergence of AI techniques for depression diagnosis, the conflict between high demand and limited supply for depression screening has been significantly alleviated. Among various modal data, audio-based depression diagnosis has…

密码学与安全 · 计算机科学 2026-03-27 Xintao Hu , Feng-Qi Cui

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additive noise. The…

计算与语言 · 计算机科学 2015-03-31 Matthew Ager , Zoran Cvetkovic , Peter Sollich

Speaker extraction aims to extract target speech signal from a multi-talker environment with interference speakers and surrounding noise, given the target speaker's reference information. Most speaker extraction systems achieve satisfactory…

音频与语音处理 · 电气工程与系统科学 2022-08-12 Chengyun Deng , Shiqian Ma , Yi Zhang , Yongtao Sha , Hui Zhang , Hui Song , Xiangang Li

Operator learning is a data-driven approximation of mappings between infinite-dimensional function spaces, such as the solution operators of partial differential equations. Kernel-based operator learning can offer accurate, theoretically…

机器学习 · 计算机科学 2025-12-22 Xinyue Yu , Hayden Schaeffer

Acoustic scene classification (ASC) aims to identify the type of scene (environment) in which a given audio signal is recorded. The log-mel feature and convolutional neural network (CNN) have recently become the most popular time-frequency…

声音 · 计算机科学 2021-08-12 Yuzhong Wu , Tan Lee

Improving the interpretability of deep neural networks has recently gained increased attention, especially when the power of deep learning is leveraged to solve problems in physics. Interpretability helps us understand a model's ability to…

声音 · 计算机科学 2023-10-12 Karim Helwani , Erfan Soltanmohammadi , Michael M. Goodwin

In few-shot learning (FSL), the labeled samples are scarce. Thus, label errors can significantly reduce classification accuracy. Since label errors are inevitable in realistic learning tasks, improving the robustness of the model in the…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Nan Xiang , Lifeng Xing , Dequan Jin

Recently, Meta-Black-Box-Optimization (MetaBBO) methods significantly enhance the performance of traditional black-box optimizers through meta-learning flexible and generalizable meta-level policies that excel in dynamic algorithm…

神经与进化计算 · 计算机科学 2025-03-25 Hongshu Guo , Sijie Ma , Zechuan Huang , Yuzhi Hu , Zeyuan Ma , Xinglin Zhang , Yue-Jiao Gong

We present AFEN (Audio Feature Ensemble Learning), a model that leverages Convolutional Neural Networks (CNN) and XGBoost in an ensemble learning fashion to perform state-of-the-art audio classification for a range of respiratory diseases.…

声音 · 计算机科学 2024-05-10 Rahul Nadkarni , Emmanouil Nikolakakis , Razvan Marinescu