English
Related papers

Related papers: Audio-visual child-adult speaker classification in…

200 papers

Isolating the voice of a specific person while filtering out other voices or background noises is challenging when video is shot in noisy environments. We propose audio-visual methods to isolate the voice of a single speaker and eliminate…

Computer Vision and Pattern Recognition · Computer Science 2018-02-13 Aviv Gabbay , Ariel Ephrat , Tavi Halperin , Shmuel Peleg

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker's face. However, when multiple candidate speakers are visible this traditionally requires…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Otavio Braga , Olivier Siohan

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

Children's early speech often bears little resemblance to that of adults, and yet parents and other caregivers are able to interpret that speech and react accordingly. Here we investigate how these adult inferences as listeners reflect…

Computation and Language · Computer Science 2023-03-17 Stephan C. Meylan , Ruthe Foushee , Nicole H. Wong , Elika Bergelson , Roger P. Levy

Speaker diarization consists of assigning speech signals to people engaged in a dialogue. An audio-visual spatiotemporal diarization model is proposed. The model is well suited for challenging scenarios that consist of several participants…

Computer Vision and Pattern Recognition · Computer Science 2018-10-15 Israel D. Gebru , Silèye Ba , Xiaofei Li , Radu Horaud

Silent speech interfaces have been recently proposed as a way to enable communication when the acoustic signal is not available. This introduces the need to build visual speech recognition systems for silent and whispered speech. However,…

Computer Vision and Pattern Recognition · Computer Science 2018-02-20 Stavros Petridis , Jie Shen , Doruk Cetin , Maja Pantic

Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction. The absence of a large,…

A speaker extraction algorithm seeks to extract the speech of a target speaker from a multi-talker speech mixture when given a cue that represents the target speaker, such as a pre-enrolled speech utterance, or an accompanying video track.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-25 Zexu Pan , Ruijie Tao , Chenglin Xu , Haizhou Li

Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies ``who spoke when'', is an essential component of the automated analysis. However, publicly available…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-13 Anfeng Xu , Tiantian Feng , Helen Tager-Flusberg , Catherine Lord , Shrikanth Narayanan

Acoustic analyses of infant vocalizations are valuable for research on speech development as well as applications in sound classification. Previous studies have focused on measures of acoustic features based on theories of speech…

Sound · Computer Science 2020-05-27 Mohammad K. Ebrahimpour , Sara Schneider , David C. Noelle , Christopher T. Kello

Conventional audio-visual approaches for active speaker detection (ASD) typically rely on visually pre-extracted face tracks and the corresponding single-channel audio to find the speaker in a video. Therefore, they tend to fail every time…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-22 Davide Berghi , Philip J. B. Jackson

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localization without…

Computer Vision and Pattern Recognition · Computer Science 2021-12-23 Di Hu , Yake Wei , Rui Qian , Weiyao Lin , Ruihua Song , Ji-Rong Wen

Recordings gathered with child-worn devices promised to revolutionize both fundamental and applied speech sciences by allowing the effortless capture of children's naturalistic speech environment and language production. This promise hinges…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Tarek Kunze , Marianne Métais , Hadrien Titeux , Lucas Elbert , Joseph Coffey , Emmanuel Dupoux , Alejandrina Cristia , Marvin Lavechin

The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and acoustic scene…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-14 Abhinav Shukla , Stavros Petridis , Maja Pantic

Learning to understand speech appears almost effortless for typically developing infants, yet from an information-processing perspective, acquiring a language from acoustic speech is an enormous challenge. This chapter reviews recent…

Computation and Language · Computer Science 2026-03-12 Okko Räsänen

Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SSL) models perform well on adult speech tasks, their ability…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-15 Abhijit Sinha , Harishankar Kumar , Mohit Joshi , Hemant Kumar Kathania , Shrikanth Narayanan , Sudarsana Reddy Kadiri

This paper presents an improved framework for character-aware audio-visual subtitling in TV shows. Our approach integrates speech recognition, speaker diarisation, and character recognition, utilising both audio and visual cues. This…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Jaesung Huh , Andrew Zisserman

The accuracy of modern automatic speaker verification (ASV) systems, when trained exclusively on adult data, drops substantially when applied to children's speech. The scarcity of children's speech corpora hinders fine-tuning ASV systems…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-26 Vishwanath Pratap Singh , Md Sahidullah , Tomi Kinnunen

The range of potential applications of acoustic analysis is wide. Classification of sounds, in particular, is a typical machine learning task that received a lot of attention in recent years. The most common approaches to sound…