English
Related papers

Related papers: Visualizations of Complex Sequences of Family-Infa…

200 papers

To perform automatic family audio analysis, past studies have collected recordings using phone, video, or audio-only recording devices like LENA, investigated supervised learning methods, and used or fine-tuned general-purpose embeddings…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-12 Jialu Li , Mark Hasegawa-Johnson , Nancy L. McElwain

The assessment of children at risk of autism typically involves a clinician observing, taking notes, and rating children's behaviors. A machine learning model that can label adult and child audio may largely save labor in coding children's…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-07 Jialu Li , Mark Hasegawa-Johnson , Karrie Karahalios

We design a framework for studying prelinguistic child voicefrom 3 to 24 months based on state-of-the-art algorithms in di-arization. Our system consists of a time-invariant feature ex-tractor, a context-dependent embedding generator, and a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-06 Junzhe Zhu , Mark Hasegawa-Johnson , Nancy McElwain

Acoustic analyses of infant vocalizations are valuable for research on speech development as well as applications in sound classification. Previous studies have focused on measures of acoustic features based on theories of speech…

Sound · Computer Science 2020-05-27 Mohammad K. Ebrahimpour , Sara Schneider , David C. Noelle , Christopher T. Kello

Recent findings show that pre-trained wav2vec 2.0 models are reliable feature extractors for various speaker characteristics classification tasks. We show that latent representations extracted at different layers of a pre-trained wav2vec…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-02 Ilja Baumann , Dominik Wagner , Franziska Braun , Sebastian P. Bayerl , Elmar Nöth , Korbinian Riedhammer , Tobias Bocklet

Spontaneous conversations in real-world settings such as those found in child-centered recordings have been shown to be amongst the most challenging audio files to process. Nevertheless, building speech processing models handling such a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Marvin Lavechin , Ruben Bousbib , Hervé Bredin , Emmanuel Dupoux , Alejandrina Cristia

During the first years of life, infant vocalizations change considerably, as infants develop the vocalization skills that enable them to produce speech sounds. Characterizations based on specific acoustic features, protophone categories, or…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-27 Silvia Pagliarini , Sara Schneider , Christopher T. Kello , Anne S. Warlaumont

Speech technology systems struggle with many downstream tasks for child speech due to small training corpora and the difficulties that child speech pose. We apply a novel dataset, SpeechMaturity, to state-of-the-art transformer models to…

Computation and Language · Computer Science 2025-06-11 Theo Zhang , Madurya Suresh , Anne S. Warlaumont , Kasia Hitczenko , Alejandrina Cristia , Margaret Cychosz

Studying early speech development at scale requires automatic tools, yet automatic phoneme recognition, especially for young children, remains largely unsolved. Building on decades of data collection, we curate TinyVox, a corpus of more…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Marvin Lavechin , Elika Bergelson , Roger Levy

We present ChildVox, a novel benchmark for characterizing the diverse acoustic signals through which children communicate. Specifically, ChildVox follows the full developmental trajectory from birth through school age, covering…

Interactions involving children span a wide range of important domains from learning to clinical diagnostic and therapeutic contexts. Automated analyses of such interactions are motivated by the need to seek accurate insights and offer…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-13 Anfeng Xu , Kevin Huang , Tiantian Feng , Helen Tager-Flusberg , Shrikanth Narayanan

Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-25 Alkis Koudounas , Moreno La Quatra , Marco Sabato Siniscalchi , Elena Baralis

Wav2Vec2.0 is a state-of-the-art model which learns speech representations through unlabeled speech data, aka, self supervised learning. The pretrained model is then fine tuned on small amounts of labeled data to use it for speech-to-text…

Sound · Computer Science 2022-02-15 Santosh Gondi

Earlier research has suggested that human infants might use statistical dependencies between speech and non-linguistic multimodal input to bootstrap their language learning before they know how to segment words from running speech. However,…

Computation and Language · Computer Science 2019-06-25 Okko Räsänen , Khazar Khorrami

Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SSL) models perform well on adult speech tasks, their ability…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-15 Abhijit Sinha , Harishankar Kumar , Mohit Joshi , Hemant Kumar Kathania , Shrikanth Narayanan , Sudarsana Reddy Kadiri

We show for the first time that learning powerful representations from speech audio alone followed by fine-tuning on transcribed speech can outperform the best semi-supervised methods while being conceptually simpler. wav2vec 2.0 masks the…

Computation and Language · Computer Science 2020-10-23 Alexei Baevski , Henry Zhou , Abdelrahman Mohamed , Michael Auli

Wav2vec 2.0 is a recently proposed self-supervised framework for speech representation learning. It follows a two-stage training process of pre-training and fine-tuning, and performs well in speech recognition tasks especially ultra-low…

Sound · Computer Science 2021-01-15 Zhiyun Fan , Meng Li , Shiyu Zhou , Bo Xu

Naturalistic recordings capture audio in real-world environments where participants behave naturally without interference from researchers or experimental protocols. Naturalistic long-form recordings extend this concept by capturing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Jialu Li , Marvin Lavechin , Xulin Fan , Nancy L. McElwain , Alejandrina Cristia , Paola Garcia-Perera , Mark Hasegawa-Johnson

The studies of predicting affective states from human voices have relied heavily on speech. This study, indeed, explores the recognition of humans' affective state from their vocal burst, a short non-verbal vocalization. Borrowing the idea…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-27 Bagus Tris Atmaja , Akira Sasou

Certain environmental noises have been associated with negative developmental outcomes for infants and young children. Though classifying or tagging sound events in a domestic environment is an active research area, previous studies focused…

‹ Prev 1 2 3 10 Next ›