English
Related papers

Related papers: Infant Vocal Tract Development Analysis and Diagno…

200 papers

Existing methods for speaker age estimation usually treat it as a multi-class classification or a regression problem. However, precise age identification remains a challenge due to label ambiguity, \emph{i.e.}, utterances from adjacent age…

Sound · Computer Science 2022-02-24 Shijing Si , Jianzong Wang , Junqing Peng , Jing Xiao

This paper integrates a voice activity detection (VAD) function with end-to-end automatic speech recognition toward an online speech interface and transcribing very long audio recordings. We focus on connectionist temporal classification…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-16 Takenori Yoshimura , Tomoki Hayashi , Kazuya Takeda , Shinji Watanabe

This article describes a system for analyzing acoustic data to assist in the diagnosis and classification of children's speech sound disorders (SSDs) using a computer. The analysis concentrated on identifying and categorizing four distinct…

Sound · Computer Science 2022-07-07 Yao-Ming Kuo , Shanq-Jang Ruan , Yu-Chin Chen , Ya-Wen Tu

In this research, we model and analyze the vocal tract under normal and stressful talking conditions. This research answers the question of the degradation in the recognition performance of text-dependent speaker identification under…

Sound · Computer Science 2017-07-04 Ismail Shahin , Nazeih Botros

In this work, a Bayesian approach to speaker normalization is proposed to compensate for the degradation in performance of a speaker independent speech recognition system. The speaker normalization method proposed herein uses the technique…

Sound · Computer Science 2016-10-20 Dhananjay Ram , Debasis Kundu , Rajesh M. Hegde

Teaching with the cooperation of expert teacher and assistant teacher, which is the so-called "double-teachers classroom", i.e., the course is giving by the expert online and presented through projection screen at the classroom, and the…

Sound · Computer Science 2021-06-01 Lu Ma , Xintian Wang , Song Yang , Yaguang Gong , Zhongqin Wu

Neonatal respiratory distress is a common condition that if left untreated, can lead to short- and long-term complications. This paper investigates the usage of digital stethoscope recorded chest sounds taken within 1min post-delivery, to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-13 Ethan Grooby , Chiranjibi Sitaula , Kenneth Tan , Lindsay Zhou , Arrabella King , Ashwin Ramanathan , Atul Malhotra , Guy A. Dumont , Faezeh Marzbanrad

Voice disorders affect patients profoundly, and acoustic tools can potentially measure voice function objectively. Nonetheless, existing tools are limited to analysing voices displaying near periodicity, and do not account for inherent…

Cellular Automata and Lattice Gases · Physics 2019-10-23 Max A Little , Patrick E McSharry , Stephen J Roberts , Declan AE Costello , Irene M Moroz

Recent findings show that pre-trained wav2vec 2.0 models are reliable feature extractors for various speaker characteristics classification tasks. We show that latent representations extracted at different layers of a pre-trained wav2vec…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-02 Ilja Baumann , Dominik Wagner , Franziska Braun , Sebastian P. Bayerl , Elmar Nöth , Korbinian Riedhammer , Tobias Bocklet

Fetal alcohol spectrum disorder (FASD) is a syndrome whose only difference compared to other children's conditions is the mother's alcohol consumption during pregnancy. An earlier diagnosis of FASD improving the quality of life of children…

Machine Learning · Computer Science 2021-06-01 Vannessa de J. Duarte , Paul Leger , Sergio Contreras , Hiroaki Fukuda

Preterm infants (born between 28 and 37 weeks of gestation) face elevated risks of neurodevelopmental delays, making early identification crucial for timely intervention. While deep learning-based volumetric segmentation of brain MRI scans…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Lexin Ren , Jiamiao Lu , Weichuan Zhang , Benqing Wu , Tuo Wang , Yi Liao , Jiapan Guo , Changming Sun , Liang Guo

Mispronunciation detection tools could increase treatment access for speech sound disorders impacting, e.g., /r/. We show age-and-sex normalized formant estimation outperforms cepstral representation for detection of fully rhotic vs.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-22 Nina R Benway , Jonathan L Preston , Asif Salekin , Yi Xiao , Harshit Sharma , Tara McAllister

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Zhenjie Liu , Jianzhang Lu , Renjie Lu , Cong Liang , Shangfei Wang

Recommendations for common outcome measures following pediatric traumatic brain injury (TBI) support the integration of instrumental measurements alongside perceptual assessment in recovery and treatment plans. A comprehensive set of…

Rodents employ a broad spectrum of ultrasonic vocalizations (USVs) for social communication. As these vocalizations offer valuable insights into affective states, social interactions, and developmental stages of animals, various deep…

Digital biomarkers for depression have largely relied on static acoustic descriptors, pooled summary statistics, or conventional machine learning representations. Such approaches may miss nonlinear temporal organization embedded in…

Sound · Computer Science 2026-04-30 Himadri S Samanta

Congenital anomalies arising as a result of a defect in the structure of the heart and great vessels are known as congenital heart diseases or CHDs. A PCG can provide essential details about the mechanical conduction system of the heart and…

Multi-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Rahil Parikh , Nadee Seneviratne , Ganesh Sivaraman , Shihab Shamma , Carol Espy-Wilson

The aim of this paper was the detection of pathologies through respiratory sounds. The ICBHI (International Conference on Biomedical and Health Informatics) Benchmark was used. This dataset is composed of 920 sounds of which 810 are of…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 María Teresa García-Ordás , José Alberto Benítez-Andrades , Isaías García-Rodríguez , Carmen Benavides , Héctor Alaiz-Moretón

This paper presents results on Speaker Recognition (SR) for children's speech, using the OGI Kids corpus and GMM-UBM and GMM-SVM SR systems. Regions of the spectrum containing important speaker information for children are identified by…

‹ Prev 1 8 9 10 Next ›