中文
相关论文

相关论文: Pronunciation recognition of English phonemes /\te…

200 篇论文

A major hurdle in data-driven research on typology is having sufficient data in many languages to draw meaningful conclusions. We present VoxClamantis v1.0, the first large-scale corpus for phonetic typology, with aligned segments and…

In this paper, we propose to improve emotion recognition by combining acoustic information and conversation transcripts. On the one hand, an LSTM network was used to detect emotion from acoustic features like f0, shimmer, jitter, MFCC, etc.…

音频与语音处理 · 电气工程与系统科学 2019-11-04 Jaejin Cho , Raghavendra Pappagari , Purva Kulkarni , Jesus Villalba , Yishay Carmiel , Najim Dehak

Deep learning has enabled highly realistic synthetic speech, raising concerns about fraud, impersonation, and disinformation. Despite rapid progress in neural detectors, transparent baselines are needed to reveal which acoustic cues…

音频与语音处理 · 电气工程与系统科学 2026-04-16 Faheem Ahmad , Ajan Ahmed , Masudul Imtiaz

The expressive nature of the voice provides a powerful medium for communicating sonic ideas, motivating recent research on methods for query by vocalisation. Meanwhile, deep learning methods have demonstrated state-of-the-art results for…

多媒体 · 计算机科学 2018-02-15 Adib Mehrabi , Keunwoo Choi , Simon Dixon , Mark Sandler

An ideal audio retrieval system efficiently and robustly recognizes a short query snippet from an extensive database. However, the performance of well-known audio fingerprinting systems falls short at high signal distortion levels. This…

音频与语音处理 · 电气工程与系统科学 2024-11-22 Anup Singh , Kris Demuynck , Vipul Arora

This study addresses unsupervised subword modeling, i.e., learning acoustic feature representations that can distinguish between subword units of a language. We propose a two-stage learning framework that combines self-supervised learning…

音频与语音处理 · 电气工程与系统科学 2021-06-08 Siyuan Feng , Odette Scharenborg

Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a…

音频与语音处理 · 电气工程与系统科学 2026-03-13 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

This paper presents a new approach for classification of dysfluent and fluent speech using Mel-Frequency Cepstral Coefficient (MFCC). The speech is fluent when person's speech flows easily and smoothly. Sounds combine into syllable,…

声音 · 计算机科学 2013-01-10 P. Mahesha , D. S. Vinod

In this work, we conduct an extensive comparison of various approaches to speech based emotion recognition systems. The analyses were carried out on audio recordings from Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS).…

声音 · 计算机科学 2019-12-24 Kannan Venkataramanan , Haresh Rengaraj Rajamohan

In clinical voice signal analysis, mishandling of subharmonic voicing may cause an acoustic parameter to signal false negatives. As such, the ability of a fundamental frequency estimator to identify speaking fundamental frequency is…

音频与语音处理 · 电气工程与系统科学 2025-01-10 Takeshi Ikuma , Melda Kunduk , Andrew J. McWhorter

A novel text-independent speaker identification (SI) method is proposed. This method uses the Mel-frequency Cepstral coefficients (MFCCs) and the dynamic information among adjacent frames as feature sets to capture speaker's…

声音 · 计算机科学 2020-02-04 Zhanyu Ma , Hong Yu

While Word2Vec represents words (in text) as vectors carrying semantic information, audio Word2Vec was shown to be able to represent signal segments of spoken words as vectors carrying phonetic structure information. Audio Word2Vec can be…

计算与语言 · 计算机科学 2018-08-08 Yu-Hsuan Wang , Hung-yi Lee , Lin-shan Lee

Current state-of-the-art speech recognition systems build on recurrent neural networks for acoustic and/or language modeling, and rely on feature extraction pipelines to extract mel-filterbanks or cepstral coefficients. In this paper we…

计算与语言 · 计算机科学 2019-04-10 Neil Zeghidour , Qiantong Xu , Vitaliy Liptchinsky , Nicolas Usunier , Gabriel Synnaeve , Ronan Collobert

We propose a weakly-supervised model for word-level mispronunciation detection in non-native (L2) English speech. To train this model, phonetically transcribed L2 speech is not required and we only need to mark mispronounced words. The lack…

音频与语音处理 · 电气工程与系统科学 2021-06-08 Daniel Korzekwa , Jaime Lorenzo-Trueba , Thomas Drugman , Shira Calamaro , Bozena Kostek

Cochlear implant (CI) users have considerable difficulty in understanding speech in reverberant listening environments. Time-frequency (T-F) masking is a common technique that aims to improve speech intelligibility by multiplying…

音频与语音处理 · 电气工程与系统科学 2021-06-01 Kevin M. Chu , Leslie M. Collins , Boyla O. Mainsah

Speaker verification (SV) systems are currently being used to make sensitive decisions like giving access to bank accounts or deciding whether the voice of a suspect coincides with that of the perpetrator of a crime. Ensuring that these…

音频与语音处理 · 电气工程与系统科学 2025-11-18 Mariel Estevez , Luciana Ferrer

Emotion recognition from audio signals has been regarded as a challenging task in signal processing as it can be considered as a collection of static and dynamic classification tasks. Recognition of emotions from speech data has been…

声音 · 计算机科学 2020-09-21 Soham Chattopadhyay , Arijit Dey , Hritam Basak

In speaker verification, the extraction of voice representations is mainly based on the Residual Neural Network (ResNet) architecture. ResNet is built upon convolution layers which learn filters to capture local spatial patterns along all…

音频与语音处理 · 电气工程与系统科学 2021-09-14 Mickael Rouvier , Pierre-Michel Bousquet

In this article, we provide a model to estimate a real-valued measure of the intelligibility of individual speech segments. We trained regression models based on Convolutional Neural Networks (CNN) for stop consonants…

音频与语音处理 · 电气工程与系统科学 2021-03-29 Ali Abavisani , Mark Hasegawa-Johnson

Speech recognition systems have made tremendous progress since the last few decades. They have developed significantly in identifying the speech of the speaker. However, there is a scope of improvement in speech recognition systems in…

计算与语言 · 计算机科学 2021-10-19 Pierre Berjon , Avishek Nag , Soumyabrata Dev