English
Related papers

Related papers: Age Group Classification with Speech and Metadata …

200 papers

During the first years of life, infant vocalizations change considerably, as infants develop the vocalization skills that enable them to produce speech sounds. Characterizations based on specific acoustic features, protophone categories, or…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-27 Silvia Pagliarini , Sara Schneider , Christopher T. Kello , Anne S. Warlaumont

Second-order statistical methods show very good results for automatic speaker identification in controlled recording conditions. These approaches are generally used on the entire speech material available. In this paper, we study the…

Information Retrieval · Computer Science 2024-02-27 Ivan Magrin-Chagnolleau , Jean François Bonastre , Frédéric Bimbot

This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability technique has been…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Puneet Kumar , Sarthak Malik , Balasubramanian Raman

Retrieval-augmented generation can improve audio captioning by incorporating relevant audio-text pairs from a knowledge base. Existing methods typically rely solely on the input audio as a unimodal retrieval query. In contrast, we propose…

Sound · Computer Science 2025-06-11 Choi Changin , Lim Sungjun , Rhee Wonjong

Generative models are a popular choice for adult-to-adult voice conversion (VC) because of their efficient way of modelling unlabelled data. To this point their usefulness in producing children speech and in particular adult to child VC has…

Sound · Computer Science 2025-12-16 Protima Nomo Sudro , Anton Ragni , Thomas Hain

Audio-recordings collected with a child-worn device are a fundamental tool in child language research. Long-form recordings collected over whole days promise to capture children's input and production with minimal observer bias, and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Loann Peurey , Marvin Lavechin , Tarek Kunze , Manel Khentout , Lucas Gautheron , Emmanuel Dupoux , Alejandrina Cristia

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

This paper provides a comprehensive evaluation of demographic and linguistic biases in omnimodal language models that process text, images, audio, and video within a single framework. Although these models are being widely deployed, their…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Alaa Elobaid

We present Attend-Fusion, a novel and efficient approach for audio-visual fusion in video classification tasks. Our method addresses the challenge of exploiting both audio and visual modalities while maintaining a compact model…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Mahrukh Awan , Asmar Nadeem , Armin Mustafa

Generally, facial age variations affect gender classification accuracy significantly, because facial shape and skin texture change as they grow old. This requires re-examination on the gender classification system to consider facial age…

Computer Vision and Pattern Recognition · Computer Science 2018-09-10 Jun Beom Kho

Flexible laryngoscopy is commonly performed by otolaryngologists to detect laryngeal diseases and to recognize potentially malignant lesions. Recently, researchers have introduced machine learning techniques to facilitate automated…

Computer Vision and Pattern Recognition · Computer Science 2023-05-29 Tianxiao Zhang , Andrés M. Bur , Shannon Kraft , Hannah Kavookjian , Bryan Renslo , Xiangyu Chen , Bo Luo , Guanghui Wang

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovisual fusion by…

Children are one of the most under-represented groups in speech technologies, as well as one of the most vulnerable in terms of privacy. Despite this, anonymization techniques targeting this population have received little attention. In…

Computers and Society · Computer Science 2025-06-05 Ajinkya Kulkarni , Francisco Teixeira , Enno Hermann , Thomas Rolland , Isabel Trancoso , Mathew Magimai Doss

We propose a novel method to use both audio and a low-resolution image to perform extreme face super-resolution (a 16x increase of the input size). When the resolution of the input image is very low (e.g., 8x8 pixels), the loss of…

Computer Vision and Pattern Recognition · Computer Science 2020-04-03 Givi Meishvili , Simon Jenni , Paolo Favaro

This paper presents results on Speaker Recognition (SR) for children's speech, using the OGI Kids corpus and GMM-UBM and GMM-SVM SR systems. Regions of the spectrum containing important speaker information for children are identified by…

With the increase in video-sharing platforms across the internet, it is difficult for humans to moderate the data for explicit content. Hence, an automated pipeline to scan through video data for explicit content has become the need of the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-22 Shaunak Joshi , Raghav Gaggar

This paper presents an improved framework for character-aware audio-visual subtitling in TV shows. Our approach integrates speech recognition, speaker diarisation, and character recognition, utilising both audio and visual cues. This…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Jaesung Huh , Andrew Zisserman

In this paper, we present multimodal deep neural network frameworks for age and gender classification, which take input a profile face image as well as an ear image. Our main objective is to enhance the accuracy of soft biometric trait…

Computer Vision and Pattern Recognition · Computer Science 2019-07-25 Dogucan Yaman , Fevziye Irem Eyiokur , Hazım Kemal Ekenel

Humans perceive the world by concurrently processing and fusing high-dimensional inputs from multiple modalities such as vision and audio. Machine perception models, in stark contrast, are typically modality-specific and optimised for…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Arsha Nagrani , Shan Yang , Anurag Arnab , Aren Jansen , Cordelia Schmid , Chen Sun

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extraction has been proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-03 Hiroshi Sato , Tsubasa Ochiai , Keisuke Kinoshita , Marc Delcroix , Tomohiro Nakatani , Shoko Araki