English
Related papers

Related papers: Age Group Classification with Speech and Metadata …

200 papers

Deep learning algorithms for predicting neuroimaging data have shown considerable promise in various applications. Prior work has demonstrated that deep learning models that take advantage of the data's 3D structure can outperform standard…

Image and Video Processing · Electrical Eng. & Systems 2023-03-07 Yuda Bi , Anees Abrol , Zening Fu , Jiayu Chen , Jingyu Liu , Vince Calhoun

Video face re-aging deals with altering the apparent age of a person to the target age in videos. This problem is challenging due to the lack of paired video datasets maintaining temporal consistency in identity and age. Most re-aging…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Abdul Muqeet , Kyuchul Lee , Bumsoo Kim , Yohan Hong , Hyungrae Lee , Woonggon Kim , KwangHee Lee

This work seeks the possibility of generating the human face from voice solely based on the audio-visual data without any human-labeled annotations. To this end, we propose a multi-modal learning framework that links the inference stage and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-14 Hyeong-Seok Choi , Changdae Park , Kyogu Lee

Currently, every 1 in 54 children have been diagnosed with Autism Spectrum Disorder (ASD), which is 178% higher than it was in 2000. An early diagnosis and treatment can significantly increase the chances of going off the spectrum and…

Computer Vision and Pattern Recognition · Computer Science 2021-04-05 Spencer He , Ryan Liu

With the rapid advancement of conversational and diffusion-based AI, there is a growing adoption of AI in educational services, ranging from grading and assessment tools to personalized learning systems that provide targeted support for…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-09 Rohan Sharma , Dancheng Liu , Jingchen Sun , Shijie Zhou , Jiayu Qin , Jinjun Xiong , Changyou Chen

Exploiting both audio and visual modalities for video classification is a challenging task, as the existing methods require large model architectures, leading to high computational complexity and resource requirements. Smaller…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Mahrukh Awan , Asmar Nadeem , Muhammad Junaid Awan , Armin Mustafa , Syed Sameed Husain

Automatically generating a natural language sentence to describe the content of an input video is a very challenging problem. It is an essential multimodal task in which auditory and visual contents are equally important. Although audio…

Computer Vision and Pattern Recognition · Computer Science 2018-12-10 Yapeng Tian , Chenxiao Guan , Justin Goodman , Marc Moore , Chenliang Xu

Early detection of autism, a neurodevelopmental disorder marked by social communication challenges, is crucial for timely intervention. Recent advancements have utilized naturalistic home videos captured via the mobile application…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Marie Huynh , Aaron Kline , Saimourya Surabhi , Kaitlyn Dunlap , Onur Cezmi Mutlu , Mohammadmahdi Honarmand , Parnian Azizian , Peter Washington , Dennis P. Wall

The potential of multimodal generative artificial intelligence (mAI) to replicate human grounded language understanding, including the pragmatic, context-rich aspects of communication, remains to be clarified. Humans are known to use…

Automatic prediction of age and gender from face images has drawn a lot of attention recently, due it is wide applications in various facial analysis problems. However, due to the large intra-class variation of face images (such as…

Computer Vision and Pattern Recognition · Computer Science 2020-12-09 Amirali Abdolrashidi , Mehdi Minaei , Elham Azimi , Shervin Minaee

Audio Descriptions (ADs) convey essential on-screen information, allowing visually impaired audiences to follow videos. To be effective, ADs must form a coherent sequence that helps listeners to visualise the unfolding scene, rather than…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Eshika Khandelwal , Junyu Xie , Tengda Han , Max Bain , Arsha Nagrani , Andrew Zisserman , Gül Varol , Makarand Tapaswi

Face recognition for infants and toddlers presents unique challenges due to rapid facial morphology changes, high inter-class similarity, and limited dataset availability. This study evaluates the performance of four deep learning-based…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Afzal Hossain , Mst Rumana Sumi , Stephanie Schuckers

Automatic speaker naming is the problem of localizing as well as identifying each speaking character in a TV/movie/live show video. This is a challenging problem mainly attributes to its multimodal nature, namely face cue alone is…

Computer Vision and Pattern Recognition · Computer Science 2015-07-20 Yongtao Hu , Jimmy Ren , Jingwen Dai , Chang Yuan , Li Xu , Wenping Wang

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

Sound · Computer Science 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

Mild cognitive impairment (MCI) is a major public health concern due to its high risk of progressing to dementia. This study investigates the potential of detecting MCI with spontaneous voice assistant (VA) commands from 35 older adults in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-08 Nana Lin , Youxiang Zhu , Xiaohui Liang , John A. Batsis , Caroline Summerour

In this paper, we present a detailed analysis on extracting soft biometric traits, age and gender, from ear images. Although there have been a few previous work on gender classification using ear images, to the best of our knowledge, this…

Computer Vision and Pattern Recognition · Computer Science 2018-06-18 Dogucan Yaman , Fevziye Irem Eyiokur , Nurdan Sezgin , Hazım Kemal Ekenel

Public speaking and presentation competence plays an essential role in many areas of social interaction in our educational, professional, and everyday life. Since our intention during a speech can differ from what is actually understood by…

Computer Vision and Pattern Recognition · Computer Science 2021-05-07 Ömer Sümer , Cigdem Beyan , Fabian Ruth , Olaf Kramer , Ulrich Trautwein , Enkelejda Kasneci

We address the need for a large-scale database of children's faces by using generative adversarial networks (GANs) and face age progression (FAP) models to synthesize a realistic dataset referred to as HDA-SynChildFaces. To this end, we…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Magnus Falkenberg , Anders Bensen Ottsen , Mathias Ibsen , Christian Rathgeb

Given a gallery of face images of missing children, state-of-the-art face recognition systems fall short in identifying a child (probe) recovered at a later age. We propose a feature aging module that can age-progress deep face features…

Computer Vision and Pattern Recognition · Computer Science 2020-03-20 Debayan Deb , Divyansh Aggarwal , Anil K. Jain

Natural Language Processing has recently made understanding human interaction easier, leading to improved sentimental analysis and behaviour prediction. However, the choice of words and vocal cues in conversations presents an underexplored…

Computers and Society · Computer Science 2022-06-24 Amna Anwar , Eiman Kanjo , Dario Ortega Anderez