English
Related papers

Related papers: Improving Computer Assisted Speech Therapy Through…

200 papers

Vietnamese Speech Emotion Recognition (SER) remains challenging due to ambiguous acoustic patterns and the lack of reliable annotated data, especially in real-world conditions where emotional boundaries are not clearly separable. To address…

Computation and Language · Computer Science 2026-04-03 Truc Nguyen , Then Tran , Binh Truong , Phuoc Nguyen T. H

Emotion and intent recognition from speech is essential and has been widely investigated in human-computer interaction. The rapid development of social media platforms, chatbots, and other technologies has led to a large volume of speech…

Sound · Computer Science 2025-07-11 Zhao Ren , Rathi Adarshi Rammohan , Kevin Scheck , Sheng Li , Tanja Schultz

Recent studies have outlined the accessibility challenges faced by blind or visually impaired, and less-literate people, in interacting with social networks, in-spite of facilitating technologies such as monotone text-to-speech (TTS) screen…

Social and Information Networks · Computer Science 2024-10-28 Suparna De , Ionut Bostan , Nishanth Sastry

Recent innovations in self-supervised representation learning have led to remarkable advances in natural language processing. That said, in the speech processing domain, self-supervised representation learning-based systems are not yet…

Computation and Language · Computer Science 2022-03-02 Hagai Aronowitz , Itai Gat , Edmilson Morais , Weizhong Zhu , Ron Hoory

Large-scale audio language models (ALMs), such as Qwen2-Audio, are capable of comprehending diverse audio signal, performing audio analysis and generating textual responses. However, in speech emotion recognition (SER), ALMs often suffer…

Sound · Computer Science 2025-12-30 Zhixian Zhao , Xinfa Zhu , Xinsheng Wang , Shuiyuan Wang , Xuelong Geng , Wenjie Tian , Lei Xie

In this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. First, we investigated…

Computation and Language · Computer Science 2024-10-28 Ziyang Ma , Wen Wu , Zhisheng Zheng , Yiwei Guo , Qian Chen , Shiliang Zhang , Xie Chen

Automatic Speech Recognition (ASR) is a technology that converts spoken words into text, facilitating interaction between humans and machines. One of the most common applications of ASR is Speech-To-Text (STT) technology, which simplifies…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-02 Jaeyoung Huh , Sangjoon Park , Jeong Eun Lee , Jong Chul Ye

Speech therapy is essential for rehabilitating speech disorders caused by neurological impairments such as stroke. However, traditional manual and computer-assisted systems are limited in real-time accessibility and articulatory motion…

Sound · Computer Science 2025-11-03 Yudong Yang , Xiaokang Liu , Shaofeng zhao , Rongfeng Su , Nan Yan , Lan Wang

Speech emotion recognition (SER) has been a challenging problem in spoken language processing research, because it is unclear how human emotions are connected to various components of sounds such as pitch, loudness, and energy. This paper…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Tai Vu

The study proposes and tests a technique for automated emotion recognition through mouth detection via Convolutional Neural Networks (CNN), meant to be applied for supporting people with health disorders with communication skills issues…

Computer Vision and Pattern Recognition · Computer Science 2022-12-13 Giulio Biondi , Valentina Franzoni , Osvaldo Gervasi , Damiano Perri

Cognitive behavioral therapy (CBT) is a widely used therapeutic method for guiding individuals toward restructuring their thinking patterns as a means of addressing anxiety, depression, and other challenges. We developed a large language…

The goal of our research is to automatically retrieve the satisfaction and the frustration in real-life call-center conversations. This study focuses an industrial application in which the customer satisfaction is continuously tracked down…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-10 Manon Macary , Marie Tahon , Yannick Estève , Daniel Luzzati

Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features such as fundamental frequency, intensity, and temporal…

Speech Large Language Models (LLMs) show great promise for speech emotion recognition (SER) via generative interfaces. However, shifting from closed-set classification to open text generation introduces zero-shot stochasticity, making…

Sound · Computer Science 2026-03-11 Hezhao Zhang , Huang-Cheng Chou , Shrikanth Narayanan , Thomas Hain

Automatic speech emotion recognition (SER) by a computer is a critical component for more natural human-machine interaction. As in human-human interaction, the capability to perceive emotion correctly is essential to take further steps in a…

Sound · Computer Science 2022-10-27 Bagus Tris Atmaja , Masato Akagi

In this paper, we propose to improve emotion recognition by combining acoustic information and conversation transcripts. On the one hand, an LSTM network was used to detect emotion from acoustic features like f0, shimmer, jitter, MFCC, etc.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-04 Jaejin Cho , Raghavendra Pappagari , Purva Kulkarni , Jesus Villalba , Yishay Carmiel , Najim Dehak

Self-supervised learning (SSL)-based speech models are extensively used for full-stack speech processing. However, it has been observed that improving SSL-based speech representations using unlabeled speech for content-related tasks is…

Computation and Language · Computer Science 2024-06-14 Amit Meghanani , Thomas Hain

Emotion recognition from speech plays a vital role in the development of empathetic human-computer interaction systems. This paper presents a comparative analysis of lightweight transformer-based models, DistilHuBERT and PaSST, by…

Sound · Computer Science 2025-11-04 Lucky Onyekwelu-Udoka , Md Shafiqul Islam , Md Shahedul Hasan

The primary objective is to teach a machine about human emotions, which has become an essential requirement in the field of social intelligence, also expedites the progress of human-machine interactions. The ability of a machine to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-23 Sai Nikhil Chennoor , B. R. K. Madhur , Moujiz Ali , T. Kishore Kumar

Speech emotion recognition is a challenging task for three main reasons: 1) human emotion is abstract, which means it is hard to distinguish; 2) in general, human emotion can only be detected in some specific moments during a long…

Sound · Computer Science 2019-05-03 Yuanyuan Zhang , Jun Du , Zirui Wang , Jianshu Zhang