English
Related papers

Related papers: Causal-Anticausal Decomposition of Speech using Co…

200 papers

General-purpose Large Language Models (LLMs) are frequently fine-tuned through supervised fine-tuning (SFT) to enhance performance in specific domains. Better results can be achieved by distilling the chain-of-thought of a larger model at…

Machine Learning · Computer Science 2026-03-24 Andrey Goncharov , Daniil Vyazhev , Petr Sychev , Edvard Khalafyan , Alexey Zaytsev

Causal generative modeling is essential for developing reliable and transparent AI systems capable of counterfactual reasoning. While existing approaches focus on integrating causal constraints during the training of generative models, they…

Machine Learning · Computer Science 2026-05-25 Aneesh Komanduri , Xintao Wu

Estimating causal effects from observational data has become increasingly critical in diverse fields including healthcare, economics, and social policy. The fundamental challenge in causal inference arises from the missing counterfactuals…

Machine Learning · Computer Science 2026-05-08 Yifei Xie , Jian Huang

Current state-of-the-art speech recognition systems build on recurrent neural networks for acoustic and/or language modeling, and rely on feature extraction pipelines to extract mel-filterbanks or cepstral coefficients. In this paper we…

Computation and Language · Computer Science 2019-04-10 Neil Zeghidour , Qiantong Xu , Vitaliy Liptchinsky , Nicolas Usunier , Gabriel Synnaeve , Ronan Collobert

In this study, we propose a novel multi-modal end-to-end neural approach for automated assessment of non-native English speakers' spontaneous speech using attention fusion. The pipeline employs Bi-directional Recurrent Convolutional Neural…

Computation and Language · Computer Science 2021-11-30 Manraj Singh Grover , Yaman Kumar , Sumit Sarin , Payman Vafaee , Mika Hama , Rajiv Ratn Shah

The recently proposed semi-blind source separation (SBSS) method for nonlinear acoustic echo cancellation (NAEC) outperforms adaptive NAEC in attenuating the nonlinear acoustic echo. However, the multiplicative transfer function (MTF)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-18 Guoliang Cheng , Lele Liao , Kai Chen , Yuxiang Hu , Changbao Zhu , Jing Lu

Over the past years, semantic segmentation, as many other tasks in computer vision, benefited from the progress in deep neural networks, resulting in significantly improved performance. However, deep architectures trained with…

Computer Vision and Pattern Recognition · Computer Science 2022-02-02 Guanglei Yang , Enrico Fini , Dan Xu , Paolo Rota , Mingli Ding , Hao Tang , Xavier Alameda-Pineda , Elisa Ricci

We introduce DecompSR, decomposed spatial reasoning, a large benchmark dataset (over 5m datapoints) and generation framework designed to analyse compositional spatial reasoning ability. The generation of DecompSR allows users to…

Artificial Intelligence · Computer Science 2026-04-15 Lachlan McPheat , Navdeep Kaur , Robert Blackwell , Alessandra Russo , Anthony G. Cohn , Pranava Madhyastha

Multimodal respiratory sound classification offers promise for early pulmonary disease detection by integrating bioacoustic signals with patient metadata. Nevertheless, current approaches remain vulnerable to spurious correlations from…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-28 Heejoon Koo , Miika Toikkanen , Yoon Tae Kim , Soo Yong Kim , June-Woo Kim

Recent studies have demonstrated that correntropy is an efficient tool for analyzing higher-order statistical moments in nonGaussian noise environments. Although correntropy has been used with complex data, no theoretical study was pursued…

Information Theory · Computer Science 2016-08-19 João Paulo Ferreira Guimarães

This study aims to capture aerodynamic causality from snapshot data with a time-varying mode decomposition technique referred to as information-theoretic machine learning. The current approach extracts time-dependent informative vortical…

Fluid Dynamics · Physics 2026-05-19 Ryo Koshikawa , Ryo Araki , Qiong Liu , Kai Fukami

In speech-related classification tasks, frequency-domain acoustic features such as logarithmic Mel-filter bank coefficients (FBANK) and cepstral-domain acoustic features such as Mel-frequency cepstral coefficients (MFCC) are often used.…

Sound · Computer Science 2022-06-20 Yikang Wang , Hiromitsu Nishizaki

A novel text-independent speaker identification (SI) method is proposed. This method uses the Mel-frequency Cepstral coefficients (MFCCs) and the dynamic information among adjacent frames as feature sets to capture speaker's…

Sound · Computer Science 2020-02-04 Zhanyu Ma , Hong Yu

Goal: Numerous studies had successfully differentiated normal and abnormal voice samples. Nevertheless, further classification had rarely been attempted. This study proposes a novel approach, using continuous Mandarin speech instead of a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-23 Syu-Siang Wang , Chi-Te Wang , Chih-Chung Lai , Yu Tsao , Shih-Hau Fang

This paper proposes a low algorithmic latency adaptation of the deep clustering approach to speaker-independent speech separation. It consists of three parts: a) the usage of long-short-term-memory (LSTM) networks instead of their…

Sound · Computer Science 2019-02-20 Shanshan Wang , Gaurav Naithani , Tuomas Virtanen

Action Quality Assessment (AQA) predicts fine-grained execution scores from action videos and is widely applied in sports, rehabilitation, and skill evaluation. Long-term AQA, as in figure skating or rhythmic gymnastics, is especially…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Ruisheng Han , Kanglei Zhou , Shuang Chen , Amir Atapour-Abarghouei , Hubert P. H. Shum

A comprehensive understanding of molecular clumps is essential for investigating star formation. We present an algorithm for molecular clump detection, called FacetClumps. This algorithm uses a morphological approach to extract signal…

Instrumentation and Methods for Astrophysics · Physics 2023-08-09 Yu Jiang , Zhiwei Chen , Sheng Zheng , Zhibo Jiang , Yao Huang , Shuguang Zeng , Xiangyun Zeng , Xiaoyu Luo

Early identification of respiratory irregularities is critical for improving lung health and reducing global mortality rates. The analysis of respiratory sounds plays a significant role in characterizing the respiratory system's condition…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-12 Loredana Daria Mang , Francisco David Gonzalez Martinez , Damian Martinez Munoz , Sebastian Garcia Galan , Raquel Cortina

In recent works, a flow-based neural vocoder has shown significant improvement in real-time speech generation task. The sequence of invertible flow operations allows the model to convert samples from simple distribution to audio samples.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-18 Hyun-Wook Yoon , Sang-Hoon Lee , Hyeong-Rae Noh , Seong-Whan Lee

For time-frequency (TF) domain speech enhancement (SE) methods, the overlap-and-add operation in the inverse TF transformation inevitably leads to an algorithmic delay equal to the window size. However, typical causal SE systems fail to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-22 Yuewei Zhang , Huanbin Zou , Jie Zhu