English
Related papers

Related papers: Breathing and Semantic Pause Detection and Exertio…

200 papers

Affective computing is very important in the relationship between man and machine. In this paper, a system for speech emotion recognition (SER) based on speech signal is proposed, which uses new techniques in different stages of processing.…

Sound · Computer Science 2021-11-16 Fatemeh Daneshfar , Seyed Jahanshah Kabudian

For speech classification tasks, deep learning models often achieve high accuracy but exhibit shortcomings in calibration, manifesting as classifiers exhibiting overconfidence. The significance of calibration lies in its critical role in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-27 Yaqian Hao , Chenguang Hu , Yingying Gao , Shilei Zhang , Junlan Feng

Paraphasias are speech errors that are often characteristic of aphasia and they represent an important signal in assessing disease severity and subtype. Traditionally, clinicians manually identify paraphasias by transcribing and analyzing…

Sound · Computer Science 2023-12-19 Matthew Perez , Duc Le , Amrit Romana , Elise Jones , Keli Licata , Emily Mower Provost

This paper describes our approach for the Detecting Stance in Tweets task (SemEval-2016 Task 6). We utilized recent advances in short text categorization using deep learning to create word-level and character-level models. The choice…

Computation and Language · Computer Science 2016-06-21 Prashanth Vijayaraghavan , Ivan Sysoev , Soroush Vosoughi , Deb Roy

Achieving super-human performance in recognizing human speech has been a goal for several decades, as researchers have worked on increasingly challenging tasks. In the 1990's it was discovered, that conversational speech between two humans…

Computer Vision and Pattern Recognition · Computer Science 2021-07-28 Thai-Son Nguyen , Sebastian Stueker , Alex Waibel

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

Sound · Computer Science 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee

Audio-based stuttering systems to date have been trained for detection -- what disfluency is present now -- leaving prediction, the capability needed for closed-loop intervention, unstudied at deployable scale. We train a 616K-parameter CNN…

Sound · Computer Science 2026-05-01 Nazar Kozak

Digital screening and monitoring applications can aid providers in the management of behavioral health conditions. We explore deep language models for detecting depression, anxiety, and their co-occurrence from conversational speech…

Computation and Language · Computer Science 2024-12-31 Tomasz Rutowski , Elizabeth Shriberg , Amir Harati , Yang Lu , Piotr Chlebek , Ricardo Oliveira

This article describes a system for analyzing acoustic data to assist in the diagnosis and classification of children's speech sound disorders (SSDs) using a computer. The analysis concentrated on identifying and categorizing four distinct…

Sound · Computer Science 2022-07-07 Yao-Ming Kuo , Shanq-Jang Ruan , Yu-Chin Chen , Ya-Wen Tu

Public speaking and presentation competence plays an essential role in many areas of social interaction in our educational, professional, and everyday life. Since our intention during a speech can differ from what is actually understood by…

Computer Vision and Pattern Recognition · Computer Science 2021-05-07 Ömer Sümer , Cigdem Beyan , Fabian Ruth , Olaf Kramer , Ulrich Trautwein , Enkelejda Kasneci

In this work, we explore a multimodal semi-supervised learning approach for punctuation prediction by learning representations from large amounts of unlabelled audio and text data. Conventional approaches in speech processing typically use…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-04 Monica Sunkara , Srikanth Ronanki , Dhanush Bekal , Sravan Bodapati , Katrin Kirchhoff

Stuttering is a complex speech disorder that can be identified by repetitions, prolongations of sounds, syllables or words, and blocks while speaking. Severity assessment is usually done by a speech therapist. While attempts at automated…

Quantitative Methods · Quantitative Biology 2020-06-17 Sebastian P. Bayerl , Florian Hönig , Joelle Reister , Korbinian Riedhammer

This paper presents and explores a robust deep learning framework for auscultation analysis. This aims to classify anomalies in respiratory cycles and detect disease, from respiratory sound recordings. The framework begins with front-end…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-04 Lam Pham , Huy Phan , Ramaswamy Palaniappan , Alfred Mertins , Ian McLoughlin

Coughing is a typical symptom of COVID-19. To detect and localize coughing sounds remotely, a convolutional neural network (CNN) based deep learning model was developed in this work and integrated with a sound camera for the visualization…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-16 Gyeong-Tae Lee , Hyeonuk Nam , Seong-Hu Kim , Sang-Min Choi , Youngkey Kim , Yong-Hwa Park

This paper presents an unobtrusive solution that can automatically identify deep breath when a person is walking past the global depth camera. Existing non-contact breath assessments achieve satisfactory results under restricted conditions…

Computer Vision and Pattern Recognition · Computer Science 2020-10-23 Yunlu Wang , Cheng Yang , Menghan Hu , Jian Zhang , Qingli Li , Guangtao Zhai , Xiao-Ping Zhang

End-to-end approaches for sequence tasks are becoming increasingly popular. Yet for complex sequence tasks, like speech translation, systems that cascade several models trained on sub-tasks have shown to be superior, suggesting that the…

Computation and Language · Computer Science 2021-05-04 Siddharth Dalmia , Brian Yan , Vikas Raunak , Florian Metze , Shinji Watanabe

In this paper, we are interested in exploiting textual and acoustic data of an utterance for the speech emotion classification task. The baseline approach models the information from audio and text independently using two deep neural…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-02 Seunghyun Yoon , Seokhyun Byun , Subhadeep Dey , Kyomin Jung

In this paper, we propose a deep convolutional neural network-based acoustic word embedding system on code-switching query by example spoken term detection. Different from previous configurations, we combine audio data in two languages for…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-26 Murong Ma , Haiwei Wu , Xuyang Wang , Lin Yang , Junjie Wang , Ming Li

Accurate classification of respiratory sounds requires deep learning models that effectively capture fine-grained acoustic features and long-range temporal dependencies. Convolutional Neural Networks (CNNs) are well-suited for extracting…

Sound · Computer Science 2025-07-29 Nouhaila Fraihi , Ouassim Karrakchou , Mounir Ghogho

We introduce the SEER (Span-based Emotion Evidence Retrieval) Benchmark to test Large Language Models' (LLMs) ability to identify the specific spans of text that express emotion. Unlike traditional emotion recognition tasks that assign a…

Computation and Language · Computer Science 2025-10-29 Aneesha Sampath , Oya Aran , Emily Mower Provost