English
Related papers

Related papers: Pre-Trained Foundation Model representations to un…

200 papers

Automatic speech recognition (ASR) models are normally trained to operate over single utterances, with a short duration of less than 30 seconds. This choice has been made in part due to computational constraints, but also reflects a common,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-11 Robert Flynn , Anton Ragni

An experimental testbed has been constructed to assess the capabilities of Light-Wave Sensing, a promising new vitals monitoring approach. A Light-Wave Sensing apparatus utilizes infrared radiation to contactlessly monitor the subtle…

Medical Physics · Physics 2023-11-08 Brenden Martin , Md Zobaer Islam , Carly Gotcher , Tyler Martinez , Sabit Ekin , John F. O'Hara

Speech recognition in noisy and channel distorted scenarios is often challenging as the current acoustic modeling schemes are not adaptive to the changes in the signal distribution in the presence of noise. In this work, we develop a novel…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-03 Purvi Agrawal , Sriram Ganapathy

The generation of co-speech gestures for digital humans is an emerging area in the field of virtual human creation. Prior research has made progress by using acoustic and semantic information as input and adopting classify method to…

Sound · Computer Science 2024-04-16 Fan Zhang , Naye Ji , Fuxing Gao , Siyuan Zhao , Zhaohan Wang , Shunman Li

Infants, adults, non-human primates and non-primates all learn patterns implicitly, and they do so across modalities. The biological evidence supports the hypothesis that the mechanism for this learning is general but computationally local.…

Neurons and Cognition · Quantitative Biology 2021-08-16 John Rohrlich , Randall C. O'Reilly

Voice Activity Detection (VAD) refers to the task of identification of regions of human speech in digital signals such as audio and video. While VAD is a necessary first step in many speech processing systems, it poses challenges when there…

Machine Learning · Computer Science 2020-08-24 Arnab Kumar Mondal , Prathosh A. P

As deepfake audio becomes more realistic and diverse, developing generalizable countermeasure systems has become crucial. Existing detection methods primarily depend on XLS-R front-end features to improve generalization. Nonetheless, their…

Sound · Computer Science 2026-02-17 Zhe Ye , Xiangui Kang , Jiayi He , Chengxin Chen , Wei Zhu , Kai Wu , Yin Yang , Jiwu Huang

Visual speech recognition (VSR) systems decode spoken words from an input sequence using only the video data. Practical applications of such systems include medical assistance as well as human-machine interactions. A VSR system is typically…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Iason Ioannis Panagos , Giorgos Sfikas , Christophoros Nikou

We introduce a new approach for speech pre-training named SPIRAL which works by learning denoising representation of perturbed data in a teacher-student framework. Specifically, given a speech utterance, we first feed the utterance to a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-08 Wenyong Huang , Zhenhe Zhang , Yu Ting Yeung , Xin Jiang , Qun Liu

We propose a Perceiver-based sequence classifier to detect abnormalities in speech reflective of several neurological disorders. We combine this classifier with a Universal Speech Model (USM) that is trained (unsupervised) on 12 million…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-23 Hagen Soltau , Izhak Shafran , Alex Ottenwess , Joseph R. JR Duffy , Rene L. Utianski , Leland R. Barnard , John L. Stricker , Daniela Wiepert , David T. Jones , Hugo Botha

Auscultation of respiratory sounds is the primary tool for screening and diagnosing lung diseases. Automated analysis, coupled with digital stethoscopes, can play a crucial role in enabling tele-screening of fatal lung diseases. Deep neural…

Sound · Computer Science 2021-05-10 Siddhartha Gairola , Francis Tom , Nipun Kwatra , Mohit Jain

This paper addresses the performance of systems which use commercial wireless devices to make bistatic RF channel measurements for non-contact respiration sensing. Published research has typically presented results from short controlled…

Inspired by the recent success of sequence modeling in RL and the use of masked language model for pre-training, we propose a masked model for pre-training in RL, RePreM (Representation Pre-training with Masked Model), which trains the…

Machine Learning · Computer Science 2023-03-06 Yuanying Cai , Chuheng Zhang , Wei Shen , Xuyun Zhang , Wenjie Ruan , Longbo Huang

A key barrier to making phonetic studies scalable and replicable is the need to rely on subjective, manual annotation. To help meet this challenge, a machine learning algorithm was developed for automatic measurement of a widely used…

Machine Learning · Statistics 2017-03-08 Yossi Adi , Joseph Keshet , Emily Cibelli , Erin Gustafson , Cynthia Clopper , Matthew Goldrick

Respiration waveforms are increasingly recognized as important biomarkers, offering insights beyond simple respiration rates, such as detecting breathing irregularities for disease diagnosis or monitoring breath patterns to guide…

Signal Processing · Electrical Eng. & Systems 2025-03-17 Ziqi Wang , Derek Hua , Wenjun Jiang , Tianwei Xing , Xun Chen , Mani Srivastava

Spoken communication occurs in a "noisy channel" characterized by high levels of environmental noise, variability within and between speakers, and lexical and syntactic ambiguity. Given these properties of the received linguistic input,…

Computation and Language · Computer Science 2021-01-26 Stephan C. Meylan , Sathvik Nair , Thomas L. Griffiths

Modern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging to read due to disfluency, filter words, and other errata…

Computation and Language · Computer Science 2021-02-23 Junwei Liao , Yu Shi , Ming Gong , Linjun Shou , Sefik Eskimez , Liyang Lu , Hong Qu , Michael Zeng

In this paper, we present a novel multi-modal deep neural network architecture that uses speech and text entanglement for learning phonetically sound spoken-word representations. STEPs-RL is trained in a supervised manner to predict the…

Computation and Language · Computer Science 2020-11-24 Prakamya Mishra

Tracking organ motion is important in image-guided interventions, but motion annotations are not always easily available. Thus, we propose Repetitive Motion Estimation Network (RMEN) to recover cardiac and respiratory signals. It learns the…

Computer Vision and Pattern Recognition · Computer Science 2018-11-09 Xiaoxiao Li , Vivek Singh , Yifan Wu , Klaus Kirchberg , James Duncan , Ankur Kapoor

Diffusion models have shown a great ability at bridging the performance gap between predictive and generative approaches for speech enhancement. We have shown that they may even outperform their predictive counterparts for non-additive…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-13 Jean-Marie Lemercier , Julius Richter , Simon Welker , Timo Gerkmann
‹ Prev 1 4 5 6 7 8 10 Next ›