English
Related papers

Related papers: Complex Cepstrum-based Decomposition of Speech for…

200 papers

This paper is concerned with reconstructing an acoustic obstacle and its excitation sources from the phaseless near-field measurements. By supplementing some artificial sources to the inverse scattering system, this co-inversion problem can…

Numerical Analysis · Mathematics 2022-12-20 Deyue Zhang , Yue Wu , Yukun Guo

Zero-shot semantic segmentation (ZS3) aims to segment the novel categories that have not been seen in the training. Existing works formulate ZS3 as a pixel-level zeroshot classification problem, and transfer semantic knowledge from seen…

Computer Vision and Pattern Recognition · Computer Science 2022-04-18 Jian Ding , Nan Xue , Gui-Song Xia , Dengxin Dai

Speech-based depression detection tools could aid early screening. Here, we propose an interpretable speech foundation model approach to enhance the clinical applicability of such tools. We introduce a speech-level Audio Spectrogram…

Sound · Computer Science 2026-03-26 Qingkun Deng , Saturnino Luz , Sofia de la Fuente Garcia

This article focuses on covariance estimation for multi-study data. Popular approaches employ factor-analytic terms with shared and study-specific loadings that decompose the variance into (i) a shared low-rank component, (ii)…

Methodology · Statistics 2026-01-26 Lorenzo Mauri , Niccolò Anceschi , David B. Dunson

Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-art end-to-end text-to-speech synthesis system, to uncover…

Computation and Language · Computer Science 2018-08-07 Daisy Stanton , Yuxuan Wang , RJ Skerry-Ryan

Electrocardiogram (ECG) interpretation is essential for cardiovascular disease diagnosis, but current automated systems often struggle with transparency and generalization to unseen conditions. To address this, we introduce ZETA, a…

Machine Learning · Computer Science 2025-10-27 Jialu Tang , Hung Manh Pham , Ignace De Lathauwer , Henk S. Schipper , Yuan Lu , Dong Ma , Aaqib Saeed

The electromyogram (EMG) is an important tool for assessing the activity of a muscle and thus also a valuable measure for the diagnosis and control of respiratory support. In this article we propose convolutive blind source separation (BSS)…

Signal Processing · Electrical Eng. & Systems 2019-04-09 Herbert Buchner , Eike Petersen , Marcus Eger , Philipp Rostalski

Prior studies in the automatic classification of voice quality have mainly studied the use of the acoustic speech signal as input. Recently, a few studies have been carried out by jointly using both speech and neck surface accelerometer…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-08 Sudarsana Reddy Kadiri , Farhad Javanmardi , Paavo Alku

Heart sound signals, phonocardiography (PCG) signals, allow for the automatic diagnosis of potential cardiovascular pathology. Such classification task can be tackled using the bidirectional long short-term memory (biLSTM) network, trained…

Sound · Computer Science 2026-04-16 Mahmoud Fakhry , Abeer FathAllah Brery

This paper proposes an approach for optimizing a Convolutional BeamFormer (CBF) that can jointly perform denoising (DN), dereverberation (DR), and source separation (SS). First, we develop a blind CBF optimization algorithm that requires no…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-05 Tomohiro Nakatani , Rintaro Ikeshita , Keisuke Kinoshita , Hiroshi Sawada , Shoko Araki

In a range of recent works, object-centric architectures have been shown to be suitable for unsupervised scene decomposition in the vision domain. Inspired by these methods we present AudioSlots, a slot-centric generative model for blind…

Sound · Computer Science 2023-05-10 Pradyumna Reddy , Scott Wisdom , Klaus Greff , John R. Hershey , Thomas Kipf

Synthetic data generation represents a significant advancement in boosting the performance of machine learning (ML) models, particularly in fields where data acquisition is challenging, such as echocardiography. The acquisition and labeling…

Machine Learning · Computer Science 2025-08-28 Nima Kondori , Hanwen Liang , Hooman Vaseli , Bingyu Xie , Christina Luong , Purang Abolmaesumi , Teresa Tsang , Renjie Liao

With the development of computer -systems that can collect and analyze enormous volumes of data, the medical profession is establishing several non-invasive tools. This work attempts to develop a non-invasive technique for identifying…

Sound · Computer Science 2023-03-16 Hafsa Gulzar , Jiyun Li , Arslan Manzoor , Sadaf Rehmat , Usman Amjad , Hadiqa Jalil Khan

Several speaker identification systems are giving good performance with clean speech but are affected by the degradations introduced by noisy audio conditions. To deal with this problem, we investigate the use of complementary information…

Sound · Computer Science 2014-07-03 Imen Trabelsi , Dorra Ben Ayed

Cardiovascular system diseases can be identified by using a specialized diagnostic process utilizing a digital stethoscope. Digital stethoscopes provide phonocardiography (PCG) recordings for further inspection, besides filtering and…

Signal Processing · Electrical Eng. & Systems 2024-02-21 Ibrahim Ozkan , Atila Yilmaz

Given a set of mixtures, blind source separation attempts to retrieve the source signals without or with very little information of the the mixing process. We present a geometric approach for blind separation of nonnegative linear mixtures…

Numerical Analysis · Mathematics 2013-01-04 P. Yin , Y. Sun , J. Xin

Decomposition of an audio mixture into harmonic and percussive components, namely harmonic/percussive source separation (HPSS), is a useful pre-processing tool for many audio applications. Popular approaches to HPSS exploit the distinctive…

Audio and Speech Processing · Electrical Eng. & Systems 2019-03-14 Yoshiki Masuyama , Kohei Yatabe , Yasuhiro Oikawa

The Continuous Wavelet Transform (CWT) is an effective tool for feature extraction in acoustic recognition using Convolutional Neural Networks (CNNs), particularly when applied to non-stationary audio. However, its high computational cost…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-01 Dang Thoai Phan , Tuan Anh Huynh , Van Tuan Pham , Cao Minh Tran , Van Thuan Mai , Ngoc Quy Tran

Given recent advances in deep music source separation, we propose a feature representation method that combines source separation with a state-of-the-art representation learning technique that is suitably repurposed for computer audition…

Sound · Computer Science 2020-12-08 Gabriel Mersy , Jin Hong Kuan

Systems based on automatic speech recognition (ASR) technology can provide important functionality in computer assisted language learning applications. This is a young but growing area of research motivated by the large number of students…

Sound · Computer Science 2016-02-29 Zhenhao Ge , Sudhendu R. Sharma , Mark J. T. Smith
‹ Prev 1 4 5 6 7 8 10 Next ›