English
Related papers

Related papers: Age Group Classification with Speech and Metadata …

200 papers

The interest in demographic information retrieval based on text data has increased in the research community because applications have shown success in different sectors such as security, marketing, heath-care, and others. Recognition and…

Computation and Language · Computer Science 2021-07-07 Daniel Escobar-Grisales , Juan Camilo Vasquez-Correa , Juan Rafael Orozco-Arroyave

An utterance that contains speech from multiple languages is known as a code-switched sentence. In this work, we propose a novel technique to predict whether given audio is mono-lingual or code-switched. We propose a multi-modal learning…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-05 Krishna D N

In diarization, the PLDA is typically used to model an inference structure which assumes the variation in speech segments be induced by various speakers. The speaker variation is then learned from the training data. However, human…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-01 Jiamin Xie , Suzanna Sia , Paola Garcia , Daniel Povey , Sanjeev Khudanpur

The existing computational visual attention systems have focused on the objective to basically simulate and understand the concept of visual attention system in adults. Consequently, the impact of observer's age in scene viewing behavior…

Computer Vision and Pattern Recognition · Computer Science 2019-04-30 Onkar Krishna , Kiyoharu Aizawa , Go Irie

End-to-end acoustic speech recognition has quickly gained widespread popularity and shows promising results in many studies. Specifically the joint transformer/CTC model provides very good performance in many tasks. However, under noisy and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-20 Wentao Yu , Steffen Zeiler , Dorothea Kolossa

This paper explores the development of a multimodal sentiment analysis model that integrates text, audio, and visual data to enhance sentiment classification. The goal is to improve emotion detection by capturing the complex interactions…

Computation and Language · Computer Science 2025-01-15 Hui Lee , Singh Suniljit , Yong Siang Ong

Achieving super-human performance in recognizing human speech has been a goal for several decades, as researchers have worked on increasingly challenging tasks. In the 1990's it was discovered, that conversational speech between two humans…

Computer Vision and Pattern Recognition · Computer Science 2021-07-28 Thai-Son Nguyen , Sebastian Stueker , Alex Waibel

Today, data collection has improved in various areas, and the medical domain is no exception. Auscultation, as an important diagnostic technique for physicians, due to the progress and availability of digital stethoscopes, lends itself well…

This research report presents a proof-of-concept study on the application of machine learning techniques to video and speech data collected during diagnostic cognitive assessments of children with a neurodevelopmental disorder. The study…

Human-Computer Interaction · Computer Science 2023-12-20 Giulio Falcioni , Alexandra Georgescu , Emilia Molimpakis , Lev Gottlieb , Taylor Kuhn , Stefano Goria

Spatially dense self-supervised learning is a rapidly growing problem domain with promising applications for unsupervised segmentation and pretraining for dense downstream tasks. Despite the abundance of temporal data in the form of videos,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Mohammadreza Salehi , Efstratios Gavves , Cees G. M. Snoek , Yuki M. Asano

Active Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Yu Wang , Juhyung Ha , Frangil M. Ramirez , Yuchen Wang , David J. Crandall

There is a natural correlation between the visual and auditive elements of a video. In this work we leverage this connection to learn general and effective models for both audio and video analysis from self-supervised temporal…

Computer Vision and Pattern Recognition · Computer Science 2018-11-13 Bruno Korbar , Du Tran , Lorenzo Torresani

Voice disorders negatively impact the quality of daily life in various ways. However, accurately recognizing the category of pathological features from raw audio remains a considerable challenge due to the limited dataset. A promising…

Sound · Computer Science 2024-10-08 Lipeng Shen , Yifan Xiong , Dongyue Guo , Wei Mo , Lingyu Yu , Hui Yang , Yi Lin

Embodied conversational agents benefit from being able to accompany their speech with gestures. Although many data-driven approaches to gesture generation have been proposed in recent years, it is still unclear whether such systems can…

Human-Computer Interaction · Computer Science 2022-01-17 Taras Kucherenko , Rajmund Nagy , Michael Neff , Hedvig Kjellström , Gustav Eje Henter

Children's speech recognition is considered a low-resource task mainly due to the lack of publicly available data. There are several reasons for such data scarcity, including expensive data collection and annotation processes, and data…

Computation and Language · Computer Science 2024-06-25 Vrunda N. Sukhadia , Shammur Absar Chowdhury

Videoconferencing is now a frequent mode of communication in both professional and informal settings, yet it often lacks the fluidity and enjoyment of in-person conversation. This study leverages multimodal machine learning to predict…

Machine Learning · Computer Science 2025-03-11 Andrew Chang , Viswadruth Akkaraju , Ray McFadden Cogliano , David Poeppel , Dustin Freeman

Depression has been the leading cause of mental-health illness worldwide. Major depressive disorder (MDD), is a common mental health disorder that affects both psychologically as well as physically which could lead to loss of lives. Due to…

Computer Vision and Pattern Recognition · Computer Science 2019-09-05 Anupama Ray , Siddharth Kumar , Rutvik Reddy , Prerana Mukherjee , Ritu Garg

In recent years identity-vector (i-vector) based speaker verification (SV) systems have become very successful. Nevertheless, environmental noise and speech duration variability still have a significant effect on degrading the performance…

Sound · Computer Science 2016-08-09 Ali Khodabakhsh , Seyyed Saeed Sarfjoo , Umut Uludag , Osman Soyyigit , Cenk Demiroglu

Speech sound disorders are a common communication impairment in childhood. Because speech disorders can negatively affect the lives and the development of children, clinical intervention is often recommended. To help with diagnosis and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-02 Manuel Sam Ribeiro , Joanne Cleland , Aciel Eshky , Korin Richmond , Steve Renals

While demographic factors like age and gender change the way people talk, and in particular, the way people talk to machines, there is little investigation into how large pre-trained language models (LMs) can adapt to these changes. To…

Computation and Language · Computer Science 2024-02-06 Anthony Sicilia , Jennifer C. Gates , Malihe Alikhani
‹ Prev 1 8 9 10 Next ›