English
Related papers

Related papers: Pre-Trained Foundation Model representations to un…

200 papers

Recently proposed self-supervised learning approaches have been successful for pre-training speech representation models. The utility of these learned representations has been observed empirically, but not much has been studied about the…

Computation and Language · Computer Science 2022-12-06 Ankita Pasad , Ju-Chieh Chou , Karen Livescu

Using mobile phone video of the fingertip as a data source for estimating vital signs such as heart rate (HR) and respiratory rate (RR) during daily life has long been suggested. While existing literature indicates that these estimates are…

Signal Processing · Electrical Eng. & Systems 2025-07-01 Ibne Farabi Shihab

Respiratory rate (RR) is a vital sign with significant diagnostic value. Existing RR monitors often suffer from baseline drift over time, breaths can be occluded by limb or body movements, and many systems struggle to resolve shallow or…

For conversational large-vocabulary continuous speech recognition (LVCSR) tasks, up to about two thousand hours of audio is commonly used to train state of the art models. Collection of labeled conversational audio however, is prohibitively…

Computation and Language · Computer Science 2017-05-30 Shane Walker , Morten Pedersen , Iroro Orife , Jason Flaks

Breathing rate (BR), minute ventilation (VE), and other respiratory parameters are essential for real-time patient monitoring in many acute health conditions, such as asthma. The clinical standard for measuring respiration, namely…

Signal Processing · Electrical Eng. & Systems 2020-11-26 Ridwan Alam , David B. Peden , John C. Lach

Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven…

Image and Video Processing · Electrical Eng. & Systems 2024-09-25 Hong Nguyen , Sean Foley , Kevin Huang , Xuan Shi , Tiantian Feng , Shrikanth Narayanan

Video-based respiratory rate (RR) estimation is often unreliable due to inconsistent signal quality across extraction methods. We present a predictive, quality-aware framework that integrates heterogeneous signal sources with dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Nhi Nguyen , Constantino Álvarez Casado , Le Nguyen , Manuel Lage Cañellas , Miguel Bordallo López

The exponential rise in wearable sensors has garnered significant interest in assessing the physiological parameters during day-to-day activities. Respiration rate is one of the vital parameters used in the performance assessment of…

Signal Processing · Electrical Eng. & Systems 2022-02-10 Kapil Singh Rathore , Sricharan Vijayarangan , Preejith SP , Mohanasankar Sivaprakasam

Respiratory rate (RR) is a clinical sign representing ventilation. An abnormal change in RR is often the first sign of health deterioration as the body attempts to maintain oxygen delivery to its tissues. There has been a growing interest…

Machine Learning · Computer Science 2021-08-03 Seyed Amir Hossein Aqajari , Rui Cao , Amir Hosein Afandizadeh Zargari , Amir M. Rahmani

Generative models have gained more and more attention in recent years for their remarkable success in tasks that required estimating and sampling data distribution to generate high-fidelity synthetic data. In speech, text-to-speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-27 Alexander H. Liu , Matt Le , Apoorv Vyas , Bowen Shi , Andros Tjandra , Wei-Ning Hsu

This work explores speech as a biomarker and investigates the detection of respiratory insufficiency (RI) by analyzing speech samples. Previous work \cite{spira2021} constructed a dataset of respiratory insufficiency COVID-19 patient…

Sound · Computer Science 2022-10-26 Marcelo Matheus Gauy , Marcelo Finger

Both speech and sensor time series data encode information in both the time- and frequency- domains, like spectral powers and waveform shapelets. We show that speech foundation models learn representations that generalize beyond the speech…

Machine Learning · Computer Science 2025-11-25 Jaya Narain , Zakaria Aldeneh , Shirley Ren

Speech quality in online conferencing applications is typically assessed through human judgements in the form of the mean opinion score (MOS) metric. Since such a labor-intensive approach is not feasible for large-scale speech quality…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-04 Bastiaan Tamm , Helena Balabin , Rik Vandenberghe , Hugo Van hamme

Masked speech modeling (MSM) methods such as wav2vec2 or w2v-BERT learn representations over speech frames which are randomly masked within an utterance. While these methods improve performance of Automatic Speech Recognition (ASR) systems,…

This paper introduces our approaches for the Mask and Breathing Sub-Challenge in the Interspeech COMPARE Challenge 2020. For the mask detection task, we train deep convolutional neural networks with filter-bank energies, gender-aware…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-17 Haiwei Wu , Lin Zhang , Lin Yang , Xuyang Wang , Junjie Wang , Dong Zhang , Ming Li

Developing Text-to-Speech (TTS) systems that can synthesize natural breath is essential for human-like voice agents but requires extensive manual annotation of breath positions in training data. To this end, we propose a self-training…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-17 Dong Yang , Tomoki Koriyama , Yuki Saito

Real-time magnetic resonance imaging (rtMRI) of speech production enables non-invasive visualization of dynamic vocal-tract motion and is valuable for speech science and clinical assessment. However, rtMRI is fundamentally constrained by…

Recent advancements in foundation models have sparked interest in respiratory audio foundation models. However, the effectiveness of applying conventional pre-training schemes to datasets that are small-sized and lack diversity has not been…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-22 Daisuke Niizumi , Daiki Takeuchi , Masahiro Yasuda , Binh Thien Nguyen , Yasunori Ohishi , Noboru Harada

Respiratory-gated radiation therapy (RGRT) has been used to minimize the dose to normal tissue in lung-cancer radiotherapy. The present research aims to improve the regularity of respiration in RGRT using a video coached respiration guiding…

Medical Physics · Physics 2015-03-13 Hyun Jeong Lee , Ji Woon Yea , Se An Oh

Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-25 Alkis Koudounas , Moreno La Quatra , Marco Sabato Siniscalchi , Elena Baralis