English
Related papers

Related papers: Causal-Anticausal Decomposition of Speech using Co…

200 papers

This paper proposes a deep denoising auto-encoder technique to extract better acoustic features for speech synthesis. The technique allows us to automatically extract low-dimensional features from high dimensional spectral features in a…

Sound · Computer Science 2015-06-18 Zhenzhou Wu , Shinji Takaki , Junichi Yamagishi

This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-based vocoder without requiring any additional training or…

Computational morphology has the potential to support language documentation through tasks like morphological segmentation and the generation of Interlinear Glossed Text (IGT). However, our research outputs have seen limited use in…

Computation and Language · Computer Science 2025-09-25 Enora Rice , Katharina von der Wense , Alexis Palmer

In this work, we incorporated acoustically derived source features, aperiodicity, periodicity and pitch as additional targets to an acoustic-to-articulatory speech inversion (SI) system. We also propose a Temporal Convolution based SI…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-01 Yashish M. Siriwardena , Carol Espy-Wilson

The widespread availability of complex time series data in various domains such as environmental science, epidemiology, and economics demands robust causal discovery methods that can identify intricate contemporaneous and lagged…

Machine Learning · Computer Science 2026-05-12 Omar Faruque , Sahara Ali , Xue Zheng , Jianwu Wang

In recent years, the continuous wavelet transform (CWT) has been employed as a spectral feature extractor for acoustic recognition tasks in conjunction with machine learning and deep learning models. However, applying the CWT to each…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-01 Dang Thoai Phan

End-to-end acoustic speech recognition has quickly gained widespread popularity and shows promising results in many studies. Specifically the joint transformer/CTC model provides very good performance in many tasks. However, under noisy and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-20 Wentao Yu , Steffen Zeiler , Dorothea Kolossa

The goal of this work is to train strong models for visual speech recognition without requiring human annotated ground truth data. We achieve this by distilling from an Automatic Speech Recognition (ASR) model that has been trained on a…

Computer Vision and Pattern Recognition · Computer Science 2020-04-01 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Current interpretability methods focus on explaining a particular model's decision through present input features. Such methods do not inform the user of the sufficient conditions that alter these decisions when they are not desirable.…

Machine Learning · Computer Science 2023-01-20 Julia El Zini , Mohammad Mansour , Mariette Awad

How to usefully encode compositional task structure has long been a core challenge in AI. Recent work in chain of thought prompting has shown that for very large neural language models (LMs), explicitly demonstrating the inferential steps…

Computation and Language · Computer Science 2022-10-25 Victor S. Bursztyn , David Demeter , Doug Downey , Larry Birnbaum

We propose the use of parameter-efficient fine-tuning (PEFT) of foundation models for cleft lip and palate (CLP) detection and severity classification. In CLP, nasalization increases with severity due to the abnormal passage between the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-22 Susmita Bhattacharjee , Jagabandhu Mishra , H. S. Shekhawat , S. R. Mahadeva Prasanna

Accurate classification of articulatory-phonological features plays a vital role in understanding human speech production and developing robust speech technologies, particularly in clinical contexts where targeted phonemic analysis and…

Integrating various data modalities brings valuable insights into underlying phenomena. Multimodal factor analysis (FA) uncovers shared axes of variation underlying different simple data modalities, where each sample is represented by a…

Machine Learning · Computer Science 2025-04-29 Małgorzata Łazęcka , Ewa Szczurek

Spotforming is a target-speaker extraction technique that uses multiple microphone arrays. This method applies beamforming (BF) to each microphone array, and the common components among the BF outputs are estimated as the target source.…

Sound · Computer Science 2024-07-15 Shoma Ayano , Li Li , Shogo Seki , Daichi Kitamura

Manner of articulation detection using deep neural networks require a priori knowledge of the attribute discriminative features or the decent phoneme alignments. However generating an appropriate phoneme alignment is complex and its…

Computation and Language · Computer Science 2018-11-20 Pradeep Rangan , Sreenivasa Rao K

The use of hearing aids will increase in the coming years due to demographic change. One open problem that remains to be solved by a new generation of hearing aids is the cocktail party problem. A possible solution is…

Signal Processing · Electrical Eng. & Systems 2026-02-27 René Pallenberg , Fabrice Katzberg , Alfred Mertins , Marco Maass

Revealing the large-scale structure from the 21cm intensity mapping surveys is only possible after the foreground cleaning. However, most current cleaning techniques relying on the smoothness of the foreground spectrum lead to a severe side…

Cosmology and Nongalactic Astrophysics · Physics 2024-07-18 Zhenyuan Wang , Donghui Jeong

In this work, we focus on front-end design for speech deepfake detectors, the component that determines the discriminative acoustic cues provided to the classifier. Existing approaches are primarily categorized into two types. Hand-crafted…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-01 Xi Xuan , Davide Carbone , Wenxin Zhang , Ruchi Pandey , Tomi H. Kinnunen

Many of the existing TTS systems cannot accurately synthesize text containing a variety of numerical formats, resulting in reduced intelligibility of the synthesized speech. This research aims to develop a numerical format classifier that…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-03 Yaser Darwesh , Lit Wei Wern , Mumtaz Begum Mustafa

Research about brain activities involving spoken word production is considerably underdeveloped because of the undiscovered characteristics of speech artifacts, which contaminate electroencephalogram (EEG) signals and prevent the inspection…

Sound · Computer Science 2022-06-02 Holy Lovenia , Hiroki Tanaka , Sakriani Sakti , Ayu Purwarianti , Satoshi Nakamura