中文
相关论文

相关论文: Glottal source estimation robustness: A comparison…

200 篇论文

An inverse scattering problem is analyzed for vowel articulation in the human vocal tract. When a unit amplitude, monochromatic, sinusoidal volume velocity is sent from the glottis towards the lips, various types of scattering data are used…

数学物理 · 物理学 2007-05-23 Tuncay Aktosun

Underwater acoustic monitoring systems record many hours of audio data for marine research, making fast and reliable non-causal signal detection paramount. Such detectors assist in reducing the amount of labor required for signal…

信号处理 · 电气工程与系统科学 2023-02-07 Marco W. Rademan , Daniel J. Versfeld , Johan A. du Preez

This paper introduces a novel (HDAG - Harmonic Detection for Auditory Gain) method for speech intelligibility enhancement in noisy scenarios. In the proposed scheme, a series of selective Gammachirp filters are adopted to emphasize the…

音频与语音处理 · 电气工程与系统科学 2024-01-23 A. Queiroz , R. Coelho

Due to the subjective nature of current clinical evaluation, the need for automatic severity evaluation in dysarthric speech has emerged. DNN models outperform ML models but lack user-friendly explainability. ML models offer explainable…

声音 · 计算机科学 2024-12-06 Yerin Choi , Jeehyun Lee , Myoung-Wan Koo

Despite improved performances of the latest Automatic Speech Recognition (ASR) systems, transcription errors are still unavoidable. These errors can have a considerable impact in critical domains such as healthcare, when used to help with…

计算与语言 · 计算机科学 2022-07-25 Nimshi Venkat Meripo , Sandeep Konam

Traditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking. However, recent generative models may introduce…

音频与语音处理 · 电气工程与系统科学 2025-11-19 Paul A. Bereuter , Benjamin Stahl , Mark D. Plumbley , Alois Sontacchi

The modeling of speech production often relies on a source-filter approach. Although methods parameterizing the filter have nowadays reached a certain maturity, there is still a lot to be gained for several speech processing applications in…

声音 · 计算机科学 2020-01-07 Thomas Drugman , Thierry Dutoit

Face frontalization consists of synthesizing a frontally-viewed face from an arbitrarily-viewed one. The main contribution of this paper is a frontalization methodology that preserves non-rigid facial deformations in order to boost the…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Zhiqi Kang , Mostafa Sadeghi , Radu Horaud , Xavier Alameda-Pineda

While Transformer has become the de-facto standard for speech, modeling upon the fine-grained frame-level features remains an open challenge of capturing long-distance dependencies and distributing the attention weights. We propose…

计算与语言 · 计算机科学 2023-05-30 Chen Xu , Yuhao Zhang , Chengbo Jiao , Xiaoqian Liu , Chi Hu , Xin Zeng , Tong Xiao , Anxiang Ma , Huizhen Wang , JingBo Zhu

Despite significant advances in ASR, the specific acoustic cues models rely on remain unclear. Prior studies have examined such cues on a limited set of phonemes and outdated models. In this work, we apply a feature attribution technique to…

计算与语言 · 计算机科学 2025-06-04 Dennis Fucci , Marco Gaido , Matteo Negri , Mauro Cettolo , Luisa Bentivogli

Real-time Magnetic Resonance Imaging (rtMRI) visualizes vocal tract action, offering a comprehensive window into speech articulation. However, its signals are high dimensional and noisy, hindering interpretation. We investigate compact…

图像与视频处理 · 电气工程与系统科学 2026-01-30 Jay Park , Hong Nguyen , Sean Foley , Jihwan Lee , Yoonjeong Lee , Dani Byrd , Shrikanth Narayanan

Ear biometric is considered as one of the most reliable and invariant biometrics characteristics in line with iris and fingerprint characteristics. In many cases, ear biometrics can be compared with face biometrics regarding many…

计算机视觉与模式识别 · 计算机科学 2010-02-03 Dakshina Ranjan Kisku , Hunny Mehrotra , Phalguni Gupta , Jamuna Kanta Sing

This paper is devoted to improve automatic emotion recognition from speech by incorporating rhythm and temporal features. Research on automatic emotion recognition so far has mostly been based on applying features like MFCCs, pitch and…

计算机视觉与模式识别 · 计算机科学 2013-03-08 Mayank Bhargava , Tim Polzehl

Clipping or saturation in audio signals is a very common problem in signal processing, for which, in the severe case, there is still no satisfactory solution. In such case, there is a tremendous loss of information, and traditional methods…

Current accent conversion (AC) systems do not disentangle the two main sources of non-native accent: segmental and prosodic characteristics. Being able to manipulate a non-native speaker's segmental and/or prosodic channels independently is…

计算与语言 · 计算机科学 2024-08-21 Waris Quamer , Ricardo Gutierrez-Osuna

Previous approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciation issues, which…

声音 · 计算机科学 2025-03-05 Tuan Nam Nguyen , Seymanur Akti , Ngoc Quan Pham , Alexander Waibel

For multi-channel speech recognition, speech enhancement techniques such as denoising or dereverberation are conventionally applied as a front-end processor. Deep learning-based front-ends using such techniques require aligned clean and…

声音 · 计算机科学 2020-07-28 Hyeongju Kim , Hyeonseung Lee , Woo Hyun Kang , Hyung Yong Kim , Nam Soo Kim

Punctuation and Segmentation are key to readability in Automatic Speech Recognition (ASR), often evaluated using F1 scores that require high-quality human transcripts and do not reflect readability well. Human evaluation is expensive,…

计算与语言 · 计算机科学 2022-10-28 Piyush Behre , Sharman Tan , Amy Shah , Harini Kesavamoorthy , Shuangyu Chang , Fei Zuo , Chris Basoglu , Sayan Pathak

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which…

多媒体 · 计算机科学 2025-08-12 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Xinyi Yin , Danlei Huang , Fei Yu

In speech-related classification tasks, frequency-domain acoustic features such as logarithmic Mel-filter bank coefficients (FBANK) and cepstral-domain acoustic features such as Mel-frequency cepstral coefficients (MFCC) are often used.…

声音 · 计算机科学 2022-06-20 Yikang Wang , Hiromitsu Nishizaki