English
Related papers

Related papers: Glottal source estimation robustness: A comparison…

200 papers

An inverse scattering problem is analyzed for vowel articulation in the human vocal tract. When a unit amplitude, monochromatic, sinusoidal volume velocity is sent from the glottis towards the lips, various types of scattering data are used…

Mathematical Physics · Physics 2007-05-23 Tuncay Aktosun

Underwater acoustic monitoring systems record many hours of audio data for marine research, making fast and reliable non-causal signal detection paramount. Such detectors assist in reducing the amount of labor required for signal…

Signal Processing · Electrical Eng. & Systems 2023-02-07 Marco W. Rademan , Daniel J. Versfeld , Johan A. du Preez

This paper introduces a novel (HDAG - Harmonic Detection for Auditory Gain) method for speech intelligibility enhancement in noisy scenarios. In the proposed scheme, a series of selective Gammachirp filters are adopted to emphasize the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-23 A. Queiroz , R. Coelho

Due to the subjective nature of current clinical evaluation, the need for automatic severity evaluation in dysarthric speech has emerged. DNN models outperform ML models but lack user-friendly explainability. ML models offer explainable…

Sound · Computer Science 2024-12-06 Yerin Choi , Jeehyun Lee , Myoung-Wan Koo

Despite improved performances of the latest Automatic Speech Recognition (ASR) systems, transcription errors are still unavoidable. These errors can have a considerable impact in critical domains such as healthcare, when used to help with…

Computation and Language · Computer Science 2022-07-25 Nimshi Venkat Meripo , Sandeep Konam

Traditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking. However, recent generative models may introduce…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-19 Paul A. Bereuter , Benjamin Stahl , Mark D. Plumbley , Alois Sontacchi

The modeling of speech production often relies on a source-filter approach. Although methods parameterizing the filter have nowadays reached a certain maturity, there is still a lot to be gained for several speech processing applications in…

Sound · Computer Science 2020-01-07 Thomas Drugman , Thierry Dutoit

Face frontalization consists of synthesizing a frontally-viewed face from an arbitrarily-viewed one. The main contribution of this paper is a frontalization methodology that preserves non-rigid facial deformations in order to boost the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Zhiqi Kang , Mostafa Sadeghi , Radu Horaud , Xavier Alameda-Pineda

While Transformer has become the de-facto standard for speech, modeling upon the fine-grained frame-level features remains an open challenge of capturing long-distance dependencies and distributing the attention weights. We propose…

Computation and Language · Computer Science 2023-05-30 Chen Xu , Yuhao Zhang , Chengbo Jiao , Xiaoqian Liu , Chi Hu , Xin Zeng , Tong Xiao , Anxiang Ma , Huizhen Wang , JingBo Zhu

Despite significant advances in ASR, the specific acoustic cues models rely on remain unclear. Prior studies have examined such cues on a limited set of phonemes and outdated models. In this work, we apply a feature attribution technique to…

Computation and Language · Computer Science 2025-06-04 Dennis Fucci , Marco Gaido , Matteo Negri , Mauro Cettolo , Luisa Bentivogli

Real-time Magnetic Resonance Imaging (rtMRI) visualizes vocal tract action, offering a comprehensive window into speech articulation. However, its signals are high dimensional and noisy, hindering interpretation. We investigate compact…

Image and Video Processing · Electrical Eng. & Systems 2026-01-30 Jay Park , Hong Nguyen , Sean Foley , Jihwan Lee , Yoonjeong Lee , Dani Byrd , Shrikanth Narayanan

Ear biometric is considered as one of the most reliable and invariant biometrics characteristics in line with iris and fingerprint characteristics. In many cases, ear biometrics can be compared with face biometrics regarding many…

Computer Vision and Pattern Recognition · Computer Science 2010-02-03 Dakshina Ranjan Kisku , Hunny Mehrotra , Phalguni Gupta , Jamuna Kanta Sing

This paper is devoted to improve automatic emotion recognition from speech by incorporating rhythm and temporal features. Research on automatic emotion recognition so far has mostly been based on applying features like MFCCs, pitch and…

Computer Vision and Pattern Recognition · Computer Science 2013-03-08 Mayank Bhargava , Tim Polzehl

Clipping or saturation in audio signals is a very common problem in signal processing, for which, in the severe case, there is still no satisfactory solution. In such case, there is a tremendous loss of information, and traditional methods…

Current accent conversion (AC) systems do not disentangle the two main sources of non-native accent: segmental and prosodic characteristics. Being able to manipulate a non-native speaker's segmental and/or prosodic channels independently is…

Computation and Language · Computer Science 2024-08-21 Waris Quamer , Ricardo Gutierrez-Osuna

Previous approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciation issues, which…

Sound · Computer Science 2025-03-05 Tuan Nam Nguyen , Seymanur Akti , Ngoc Quan Pham , Alexander Waibel

For multi-channel speech recognition, speech enhancement techniques such as denoising or dereverberation are conventionally applied as a front-end processor. Deep learning-based front-ends using such techniques require aligned clean and…

Sound · Computer Science 2020-07-28 Hyeongju Kim , Hyeonseung Lee , Woo Hyun Kang , Hyung Yong Kim , Nam Soo Kim

Punctuation and Segmentation are key to readability in Automatic Speech Recognition (ASR), often evaluated using F1 scores that require high-quality human transcripts and do not reflect readability well. Human evaluation is expensive,…

Computation and Language · Computer Science 2022-10-28 Piyush Behre , Sharman Tan , Amy Shah , Harini Kesavamoorthy , Shuangyu Chang , Fei Zuo , Chris Basoglu , Sayan Pathak

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which…

Multimedia · Computer Science 2025-08-12 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Xinyi Yin , Danlei Huang , Fei Yu

In speech-related classification tasks, frequency-domain acoustic features such as logarithmic Mel-filter bank coefficients (FBANK) and cepstral-domain acoustic features such as Mel-frequency cepstral coefficients (MFCC) are often used.…

Sound · Computer Science 2022-06-20 Yikang Wang , Hiromitsu Nishizaki