中文
相关论文

相关论文: Turkish Voice Commands based Chess Game using Gamm…

200 篇论文

This paper proposes an approach to detect emotion from human speech employing majority voting technique over several machine learning techniques. The contribution of this work is in two folds: firstly it selects those features of speech…

声音 · 计算机科学 2018-07-12 Md. Kamruzzaman Sarker , Kazi Md. Rokibul Alam , Md. Arifuzzaman

Decision making is an important component in a speaker verification system. For the conventional GMM-UBM architecture, the decision is usually conducted based on the log likelihood ratio of the test utterance against the GMM of the claimed…

声音 · 计算机科学 2016-09-28 Lantian Li , Renyu Wang , Gang Wang , Caixia Wang , Thomas Fang Zheng

Language models have shown unprecedented capabilities, sparking debate over the source of their performance. Is it merely the outcome of learning syntactic patterns and surface level statistics, or do they extract semantics and a world…

机器学习 · 计算机科学 2024-07-16 Adam Karvonen

Automatic Speech Recognition involves mainly two steps; feature extraction and classification . Mel Frequency Cepstral Coefficient is used as one of the prominent feature extraction techniques in ASR. Usually, the set of all 12 MFCC…

计算与语言 · 计算机科学 2015-05-14 Sarika Hegde , K. K. Achary , Surendra Shetty

Major advances in Question Answering technology were needed for IBM Watson to play Jeopardy! at championship level -- the show requires rapid-fire answers to challenging natural language questions, broad general knowledge, high precision,…

人工智能 · 计算机科学 2014-02-05 Gerald Tesauro , David C. Gondek , Jonathan Lenchner , James Fan , John M. Prager

Novel text-to-speech systems can generate entirely new voices that were not seen during training. However, it remains a difficult task to efficiently create personalized voices from a high-dimensional speaker space. In this work, we use…

One of the major parts of the voice recognition field is the choice of acoustic features which have to be robust against the variability of the speech signal, mismatched conditions, and noisy environments. Thus, different speech feature…

音频与语音处理 · 电气工程与系统科学 2021-10-26 Zhor Benhafid , Kawthar Yasmine Zergat , Abderrahmane Amrouche

As the most important human-machine interfacing tool, an insignificant amount of work has been carried out on Bangla Speech Recognition compared to the English language. Motivated by this, in this work, the performance of…

音频与语音处理 · 电气工程与系统科学 2025-08-12 Dipayan Bhadra , Mehrab Hosain , Fatema Alam

In recent years, End-to-End speech recognition technology based on deep learning has developed rapidly. Due to the lack of Turkish speech data, the performance of Turkish speech recognition system is poor. Firstly, this paper studies a…

声音 · 计算机科学 2023-03-23 Zeyu Ren , Nurmement Yolwas , Huiru Wang , Wushour Slamu

For researching effects of gamification in foreign language learning for children in the "Say It Again, Kid!" project we developed a feedback paradigm that can drive gameplay in pronunciation learning games. We describe our scoring system…

音频与语音处理 · 电气工程与系统科学 2019-05-08 Reima Karhila , Anna-Riikka Smolander , Sari Ylinen , Mikko Kurimo

This paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two…

音频与语音处理 · 电气工程与系统科学 2021-09-09 Saida Mussakhojayeva , Aigerim Janaliyeva , Almas Mirzakhmetov , Yerbolat Khassanov , Huseyin Atakan Varol

This paper presents a new approach for classification of dysfluent and fluent speech using Mel-Frequency Cepstral Coefficient (MFCC). The speech is fluent when person's speech flows easily and smoothly. Sounds combine into syllable,…

声音 · 计算机科学 2013-01-10 P. Mahesha , D. S. Vinod

Post-traumatic stress disorder (PTSD) is a mental disorder that can be developed after witnessing or experiencing extremely traumatic events. PTSD can affect anyone, regardless of ethnicity, or culture. An estimated one in every eleven…

声音 · 计算机科学 2024-03-29 Mamadou Dia , Ghazaleh Khodabandelou , Alice Othmani

Professional athletes increasingly use automated analysis of meta- and signal data to improve their training and game performance. As in other related human-to-human research fields, signal data, in particular, contain important…

声音 · 计算机科学 2022-02-21 Lukas Stappen , Manuel Milling , Valentin Munst , Korakot Hoffmann , Bjorn W. Schuller

Extracting features from the speech is the most critical process in speech signal processing. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used features in the majority of the speaker and speech recognition applications,…

声音 · 计算机科学 2025-10-31 Rinku Sebastian , Simon O'Keefe , Martin Trefzer

We describe a preliminary investigation into learning a Chess player's style from game records. The method is based on attempting to learn features of a player's individual evaluation function using the method of temporal differences, with…

人工智能 · 计算机科学 2009-04-20 Mark Levene , Trevor Fenner

English speakers use probabilistic phrases such as likely to communicate information about the probability or likelihood of events. Communication is successful to the extent that the listener grasps what the speaker means to convey and, if…

神经元与认知 · 定量生物学 2023-11-28 Laurence T Maloney , Maria F Dal Martello , Vivian Fei , Valerie Ma

One of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major speech recognition improvements on the representative…

Speech Emotion Recognition (SER) traditionally relies on auditory data analysis for emotion classification. Several studies have adopted different methods for SER. However, existing SER methods often struggle to capture subtle emotional…

声音 · 计算机科学 2026-01-23 HyeYoung Lee , Muhammad Nadeem

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only…

音频与语音处理 · 电气工程与系统科学 2020-04-09 Shaojin Ding , Quan Wang , Shuo-yiin Chang , Li Wan , Ignacio Lopez Moreno