中文
相关论文

相关论文: Boosting Local Spectro-Temporal Features for Speec…

200 篇论文

Clinical characterization and interpretation of respiratory sound symptoms have remained a challenge due to the similarities in the audio properties that manifest during auscultation in medical diagnosis. The misinterpretation and…

系统与控制 · 电气工程与系统科学 2021-10-18 Chinazunwa Uwaoma , Gunjan Mansingh

Sonorant sounds are characterized by regions with prominent formant structure, high energy and high degree of periodicity. In this work, the vocal-tract system, excitation source and suprasegmental features derived from the speech signal…

声音 · 计算机科学 2021-07-02 Bidisha Sharma , S. R. Mahadeva Prasanna

Phone level localization of mis-articulation is a key requirement for an automatic articulation error assessment system. A robust phone segmentation technique is essential to aid in real-time assessment of phone level mis-articulations of…

音频与语音处理 · 电气工程与系统科学 2020-05-01 Bhavik Vachhani , Chitralekha Bhat , Sunil Kopparapu

Segmentation for continuous Automatic Speech Recognition (ASR) has traditionally used silence timeouts or voice activity detectors (VADs), which are both limited to acoustic features. This segmentation is often overly aggressive, given that…

The effects of adding pitch and voice quality features such as jitter and shimmer to a state-of-the-art CNN model for Automatic Speech Recognition are studied in this work. Pitch features have been previously used for improving classical…

音频与语音处理 · 电气工程与系统科学 2020-11-11 Guillermo Cámbara , Jordi Luque , Mireia Farrús

ASR has been shown to achieve great performance recently. However, most of them rely on massive paired data, which is not feasible for low-resource languages worldwide. This paper investigates how to learn directly from unpaired phone…

声音 · 计算机科学 2022-08-01 Da-rong Liu , Po-chun Hsu , Yi-chen Chen , Sung-feng Huang , Shun-po Chuang , Da-yi Wu , Hung-yi Lee

Feature selection is crucial for pinpointing relevant features in high-dimensional datasets, mitigating the 'curse of dimensionality,' and enhancing machine learning performance. Traditional feature selection methods for classification use…

机器学习 · 计算机科学 2025-04-08 Rittwika Kansabanik , Adrian Barbu

Cardiovascular system diseases can be identified by using a specialized diagnostic process utilizing a digital stethoscope. Digital stethoscopes provide phonocardiography (PCG) recordings for further inspection, besides filtering and…

信号处理 · 电气工程与系统科学 2024-02-21 Ibrahim Ozkan , Atila Yilmaz

Precise elevation perception in binaural audio remains a challenge, despite extensive research on head-related transfer functions (HRTFs) and spectral cues. While prior studies have advanced our understanding of sound localization cues, the…

信号处理 · 电气工程与系统科学 2025-03-17 Juan Antonio De Rus , Mario Montagud , Jesus Lopez-Ballester , Francesc J. Ferri , Maximo Cobos

Audio classification is an important task of mapping audio samples into their corresponding labels. Recently, the transformer model with self-attention mechanisms has been adopted in this field. However, existing audio transformers require…

声音 · 计算机科学 2022-02-03 Ke Chen , Xingjian Du , Bilei Zhu , Zejun Ma , Taylor Berg-Kirkpatrick , Shlomo Dubnov

We propose a learning method with feature selection for Locality-Sensitive Hashing. Locality-Sensitive Hashing converts feature vectors into bit arrays. These bit arrays can be used to perform similarity searches and personal…

机器学习 · 计算机科学 2012-10-12 Makiko Konoshima , Yui Noma

Acoustic recognition has emerged as a prominent task in deep learning research, frequently utilizing spectral feature extraction techniques such as the spectrogram from the Short-Time Fourier Transform and the scalogram from the Wavelet…

音频与语音处理 · 电气工程与系统科学 2025-12-01 Dang Thoai Phan

Previous work has shown that it is possible to improve speech recognition by learning acoustic features from paired acoustic-articulatory data, for example by using canonical correlation analysis (CCA) or its deep extensions. One limitation…

计算与语言 · 计算机科学 2018-03-21 Qingming Tang , Weiran Wang , Karen Livescu

From a machine learning perspective, the human ability localize sounds can be modeled as a non-parametric and non-linear regression problem between binaural spectral features of sound received at the ears (input) and their sound-source…

声音 · 计算机科学 2015-02-12 Yuancheng Luo , Dmitry N. Zotkin , Ramani Duraiswami

In artificial-intelligence-aided signal processing, existing deep learning models often exhibit a black-box structure, and their validity and comprehensibility remain elusive. The integration of topological methods, despite its relatively…

机器学习 · 计算机科学 2023-11-28 Pingyao Feng , Siheng Yi , Qingrui Qu , Zhiwang Yu , Yifei Zhu

Speech foundation models (SFMs) have demonstrated strong performance across a variety of downstream tasks, including speech intelligibility prediction for hearing-impaired people (SIP-HI). However, optimizing SFMs for SIP-HI has been…

人工智能 · 计算机科学 2025-05-14 Haoshuai Zhou , Boxuan Cao , Changgeng Mo , Linkai Li , Shan Xiang Wang

Speaker verification is the process by which a speakers claim of identity is tested against a claimed speaker by his or her voice. Speaker verification is done by the use of some parameters (features) from the speakers voice which can be…

声音 · 计算机科学 2019-08-16 Bhavana V. S , Pradip K. Das

Embedding acoustic information into fixed length representations is of interest for a whole range of applications in speech and audio technology. Two novel unsupervised approaches to generate acoustic embeddings by modelling of acoustic…

计算与语言 · 计算机科学 2021-02-08 Yanpei Shi , Thomas Hain

Frequency modulation features capture the fine structure of speech formants that constitute beneficial and supplementary to the traditional energy-based cepstral features. Improvements have been demonstrated mainly in GMM-HMM systems for…

声音 · 计算机科学 2019-09-04 Isidoros Rodomagoulakis , Petros Maragos

In this paper, we proposed a novel pipeline for image-level classification in the hyperspectral images. By doing this, we show that the discriminative spectral information at image-level features lead to significantly improved performance…

计算机视觉与模式识别 · 计算机科学 2016-05-12 Vivek Sharma , Luc Van Gool