中文
相关论文

相关论文: Age Group Classification with Speech and Metadata …

200 篇论文

During the first years of life, infant vocalizations change considerably, as infants develop the vocalization skills that enable them to produce speech sounds. Characterizations based on specific acoustic features, protophone categories, or…

音频与语音处理 · 电气工程与系统科学 2022-04-27 Silvia Pagliarini , Sara Schneider , Christopher T. Kello , Anne S. Warlaumont

Second-order statistical methods show very good results for automatic speaker identification in controlled recording conditions. These approaches are generally used on the entire speech material available. In this paper, we study the…

信息检索 · 计算机科学 2024-02-27 Ivan Magrin-Chagnolleau , Jean François Bonastre , Frédéric Bimbot

This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability technique has been…

计算机视觉与模式识别 · 计算机科学 2023-01-10 Puneet Kumar , Sarthak Malik , Balasubramanian Raman

Retrieval-augmented generation can improve audio captioning by incorporating relevant audio-text pairs from a knowledge base. Existing methods typically rely solely on the input audio as a unimodal retrieval query. In contrast, we propose…

声音 · 计算机科学 2025-06-11 Choi Changin , Lim Sungjun , Rhee Wonjong

Generative models are a popular choice for adult-to-adult voice conversion (VC) because of their efficient way of modelling unlabelled data. To this point their usefulness in producing children speech and in particular adult to child VC has…

声音 · 计算机科学 2025-12-16 Protima Nomo Sudro , Anton Ragni , Thomas Hain

Audio-recordings collected with a child-worn device are a fundamental tool in child language research. Long-form recordings collected over whole days promise to capture children's input and production with minimal observer bias, and…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Loann Peurey , Marvin Lavechin , Tarek Kunze , Manel Khentout , Lucas Gautheron , Emmanuel Dupoux , Alejandrina Cristia

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

This paper provides a comprehensive evaluation of demographic and linguistic biases in omnimodal language models that process text, images, audio, and video within a single framework. Although these models are being widely deployed, their…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Alaa Elobaid

We present Attend-Fusion, a novel and efficient approach for audio-visual fusion in video classification tasks. Our method addresses the challenge of exploiting both audio and visual modalities while maintaining a compact model…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Mahrukh Awan , Asmar Nadeem , Armin Mustafa

Generally, facial age variations affect gender classification accuracy significantly, because facial shape and skin texture change as they grow old. This requires re-examination on the gender classification system to consider facial age…

计算机视觉与模式识别 · 计算机科学 2018-09-10 Jun Beom Kho

Flexible laryngoscopy is commonly performed by otolaryngologists to detect laryngeal diseases and to recognize potentially malignant lesions. Recently, researchers have introduced machine learning techniques to facilitate automated…

计算机视觉与模式识别 · 计算机科学 2023-05-29 Tianxiao Zhang , Andrés M. Bur , Shannon Kraft , Hannah Kavookjian , Bryan Renslo , Xiangyu Chen , Bo Luo , Guanghui Wang

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovisual fusion by…

Children are one of the most under-represented groups in speech technologies, as well as one of the most vulnerable in terms of privacy. Despite this, anonymization techniques targeting this population have received little attention. In…

计算机与社会 · 计算机科学 2025-06-05 Ajinkya Kulkarni , Francisco Teixeira , Enno Hermann , Thomas Rolland , Isabel Trancoso , Mathew Magimai Doss

We propose a novel method to use both audio and a low-resolution image to perform extreme face super-resolution (a 16x increase of the input size). When the resolution of the input image is very low (e.g., 8x8 pixels), the loss of…

计算机视觉与模式识别 · 计算机科学 2020-04-03 Givi Meishvili , Simon Jenni , Paolo Favaro

This paper presents results on Speaker Recognition (SR) for children's speech, using the OGI Kids corpus and GMM-UBM and GMM-SVM SR systems. Regions of the spectrum containing important speaker information for children are identified by…

With the increase in video-sharing platforms across the internet, it is difficult for humans to moderate the data for explicit content. Hence, an automated pipeline to scan through video data for explicit content has become the need of the…

计算机视觉与模式识别 · 计算机科学 2023-11-22 Shaunak Joshi , Raghav Gaggar

This paper presents an improved framework for character-aware audio-visual subtitling in TV shows. Our approach integrates speech recognition, speaker diarisation, and character recognition, utilising both audio and visual cues. This…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Jaesung Huh , Andrew Zisserman

In this paper, we present multimodal deep neural network frameworks for age and gender classification, which take input a profile face image as well as an ear image. Our main objective is to enhance the accuracy of soft biometric trait…

计算机视觉与模式识别 · 计算机科学 2019-07-25 Dogucan Yaman , Fevziye Irem Eyiokur , Hazım Kemal Ekenel

Humans perceive the world by concurrently processing and fusing high-dimensional inputs from multiple modalities such as vision and audio. Machine perception models, in stark contrast, are typically modality-specific and optimised for…

计算机视觉与模式识别 · 计算机科学 2022-12-02 Arsha Nagrani , Shan Yang , Anurag Arnab , Aren Jansen , Cordelia Schmid , Chen Sun

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extraction has been proposed…

音频与语音处理 · 电气工程与系统科学 2021-02-03 Hiroshi Sato , Tsubasa Ochiai , Keisuke Kinoshita , Marc Delcroix , Tomohiro Nakatani , Shoko Araki