English
Related papers

Related papers: asya: Mindful verbal communication using deep lear…

200 papers

Deep learning models have been used widely for various purposes in recent years in object recognition, self-driving cars, face recognition, speech recognition, sentiment analysis, and many others. However, in recent years it has been shown…

Computation and Language · Computer Science 2020-06-16 Aminul Huq , Mst. Tasnim Pervin

Following the success of spoken dialogue systems (SDS) in smartphone assistants and smart speakers, a number of communicative robots are developed and commercialized. Compared with the conventional SDSs designed as a human-machine…

Computation and Language · Computer Science 2021-05-04 Tatsuya Kawahara , Koji Inoue , Divesh Lala

Nowadays, speech is becoming a more common, if not standard, interface to technology. This can be seen in the trend of technology changes over the years. Increasingly, voice is used to control programs, appliances and personal devices…

Human-Computer Interaction · Computer Science 2019-09-10 Abraham Glasser

Mindfulness meditation is a validated means of helping people manage stress. Voice-based virtual assistants (VAs) in smart speakers, smartphones, and smart environments can assist people in carrying out mindfulness meditation through guided…

Human-Computer Interaction · Computer Science 2023-04-25 Bonhee Ku , Tatsuya Itagaki , Katie Seaborn

In the era of advanced artificial intelligence and human-computer interaction, identifying emotions in spoken language is paramount. This research explores the integration of deep learning techniques in speech emotion recognition, offering…

Sound · Computer Science 2023-10-20 Hanan Hamza , Fiza Gafoor , Fathima Sithara , Gayathri Anil , V. S. Anoop

Active speaker detection (ASD) is a multi-modal task that aims to identify who, if anyone, is speaking from a set of candidates. Current audio-visual approaches for ASD typically rely on visually pre-extracted face tracks (sequences of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-08 Davide Berghi , Adrian Hilton , Philip J. B. Jackson

Humans are surrounded by audio signals that include both speech and non-speech sounds. The recognition and understanding of speech and non-speech audio events, along with a profound comprehension of the relationship between them, constitute…

Sound · Computer Science 2023-12-12 Yuan Gong , Alexander H. Liu , Hongyin Luo , Leonid Karlinsky , James Glass

For people with noise sensitivity, everyday soundscapes can be overwhelming. Existing tools such as active noise cancellation reduce discomfort by suppressing the entire acoustic environment, often at the cost of awareness of surrounding…

Sound · Computer Science 2026-04-02 Jeremy Zhengqi Huang , Emani Hicks , Sidharth , Gillian R. Hayes , Dhruv Jain

Automated speaking assessment (ASA) typically involves automatic speech recognition (ASR) and hand-crafted feature extraction from the ASR transcript of a learner's speech. Recently, self-supervised learning (SSL) has shown stellar…

Sound · Computer Science 2025-03-04 Tien-Hong Lo , Fu-An Chao , Tzu-I Wu , Yao-Ting Sung , Berlin Chen

This paper addresses the problem of modeling textual conversations and detecting emotions. Our proposed model makes use of 1) deep transfer learning rather than the classical shallow methods of word embedding; 2) self-attention mechanisms…

Computation and Language · Computer Science 2019-06-18 Waleed Ragheb , Jérôme Azé , Sandra Bringay , Maximilien Servajean

Audio Deepfake Detection (ADD) aims to detect spoof speech from bonafide speech. Most prior studies assume that stronger correlations within or across acoustic and emotional features imply authenticity, and thus focus on enhancing or…

Sound · Computer Science 2026-01-21 Jinhua Zhang , Zhenqi Jia , Rui Liu

As social robots and other intelligent machines enter the home, artificial emotional intelligence (AEI) is taking center stage to address users' desire for deeper, more meaningful human-machine interaction. To accomplish such efficacious…

Computation and Language · Computer Science 2022-06-16 Benjamin Wortman , James Z. Wang

Transforming sound insights into actionable streams of data, this abstract leverages findings from degree thesis research to enhance automotive system intelligence, enabling us to address road type [1].By extracting and interpreting…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-16 Renjith Rajagopal , Peter Winzell , Sladjana Strbac , Konstantin Lindström , Petter Hörling , Faisal Kohestani , Niloofar Mehrzad

Multimodal question answering tasks can be used as proxy tasks to study systems that can perceive and reason about the world. Answering questions about different types of input modalities stresses different aspects of reasoning such as…

Computation and Language · Computer Science 2019-11-22 Haytham M. Fayek , Justin Johnson

The problem of audio-to-text alignment has seen significant amount of research using complete supervision during training. However, this is typically not in the context of long audio recordings wherein the text being queried does not appear…

Computation and Language · Computer Science 2023-10-11 Piyush Singh Pasi , Karthikeya Battepati , Preethi Jyothi , Ganesh Ramakrishnan , Tanmay Mahapatra , Manoj Singh

Articulatory-to-acoustic (forward) mapping is a technique to predict speech using various articulatory acquisition techniques as input (e.g. ultrasound tongue imaging, MRI, lip video). The advantage of lip video is that it is easily…

Computer Vision and Pattern Recognition · Computer Science 2021-04-30 Frigyes Viktor Arthur , Tamás Gábor Csapó

Deep neural networks typically rely on the representation produced by their final hidden layer to make predictions, implicitly assuming that this single vector fully captures the semantics encoded across all preceding transformations.…

Machine Learning · Computer Science 2025-11-18 Gennaro Vessio

Deep Audio Analyzer is an open source speech framework that aims to simplify the research and the development process of neural speech processing pipelines, allowing users to conceive, compare and share results in a fast and reproducible…

Sound · Computer Science 2023-10-31 Valerio Francesco Puglisi , Oliver Giudice , Sebastiano Battiato

This study introduces an integrated approach to recognizing Arabic Sign Language (ArSL) using state-of-the-art deep learning models such as MobileNetV3, ResNet50, and EfficientNet-B2. These models are further enhanced by explainable AI…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Mazen Balat , Rewaa Awaad , Ahmed B. Zaky , Salah A. Aly

The main objective of this paper is to propose an approach for developing an Artificial Intelligence (AI)-powered Language Assessment (LA) tool. Such tools can be used to assess language impairments associated with dementia in older adults.…

Computation and Language · Computer Science 2022-09-27 Mahboobeh Parsapoor , Muhammad Raisul Alam , Alex Mihailidis